← Back to All Work

Nashaat

DEPLOYED · 2024
Qatar Ministry of Education / IMS · Government Analytics & Data Engineering

Ministerial analytics and automated reporting pipeline across 100+ schools in Qatar.

Executive Summary

ROLE & OWNERSHIP

Sole Backend & Data Pipeline Engineer — responsible for end-to-end data ingestion architecture, schema normalization, PostgreSQL database design, background processing workers, and Django REST analytical endpoints.

CORE STACK

Python, Django, PostgreSQL, Pandas, Celery, Redis, Docker, Chart.js

PRIMARY OUTCOME

Reduced ministerial reporting turnaround from 4 weeks of manual tabulation to under 3 minutes.

1. The Problem Space & Hard Constraints

Over 100 schools submitted heterogeneous quarterly evaluation surveys across disconnected formats. Central ministry administrators faced 4 weeks of manual spreadsheet aggregation per reporting cycle, resulting in data entry discrepancies, delayed strategic oversight, and zero real-time cohort visibility.

OPERATIONAL CONSTRAINTS

Strict data privacy compliance for state educational records, handling Arabic dialectal variations in free-text fields, and ensuring zero-downtime report generation during peak end-of-term submission windows.

2. Engineered Architecture & Data Pipeline

Architected an automated ETL pipeline with Arabic text normalization, scheduled asynchronous batch processing, robust validation schemas, and real-time dashboard APIs. Aggregates multi-school metrics across thousands of data points into instant executive visual analytics and exportable compliance reports.

End-to-End System Pipeline
Data Sources
100+ School Feeds & CSV/Excel Uploads
Ingestion & ETL
Pandas Normalization & Arabic Text Cleaning
Relational Storage
PostgreSQL with JSONB Meta & Indexed Views
Analytics Layer
Django APIs with Cached Query Aggregations
Executive UI
Real-Time Ministerial Dashboard & PDF Reports

3. Key Decisions & Hard Bottlenecks Overcome

Architectural Decision

Chose PostgreSQL with native JSONB columns for flexible school-specific question attributes rather than managing 50+ relational join tables. Leveraged Celery task chunks to process 10,000+ row survey uploads without locking HTTP request workers.

Bottleneck Solved

Initial pandas in-memory joins during report generation consumed 1.8GB RAM under peak load. Re-architected calculation logic into PostgreSQL database-level aggregation queries, dropping memory footprint to <120MB and execution time by 88%.

4. Verified Measurable Outcomes

  • Reduced ministerial reporting turnaround from 4 weeks of manual tabulation to under 3 minutes.
  • Eliminated data entry discrepancies with 100% schema validation across 100+ participating educational institutions.
  • Built automated PDF/Excel executive summary generation utilized directly by regional education directors.
NEXT CASE STUDY

BrandMinder

Read Next Case Study →