
Adel Abu Hashim
Verified Expert in Engineering
Data Engineer Developer
Riyadh, Riyadh Province, Saudi Arabia
Toptal member since April 8, 2024
Adel is a data architect and senior data engineer with 6+ years in telecom and fintech, designing warehouses, lakehouses, and pipelines on Teradata, Oracle, Cloudera, AWS, and Databricks. At STC, he rebuilt a national telecom platform's archival layer, reclaiming 250+ TB and cutting query latency 80%; at Paymob, he cut Redshift cost 20% while doubling query speed. He holds CDMP, AWS Data Engineer, and Microsoft Fabric certifications, and has delivered 120+ freelance projects at a 5-star average.
Portfolio
Experience
- Windows - 10 years
- Python 3 - 5 years
- SQL - 5 years
- Redshift - 3 years
- ETL - 3 years
- Pandas - 3 years
- DataHub - 2 years
- Apache Airflow - 2 years
Preferred Environment
SQL, Python, Apache Airflow, Cloudera, AWS IoT, Azure, Google Cloud Platform (GCP), Databricks, Oracle, Teradata
The most amazing...
...thing I've built is an engine that reverse-engineers Informatica and DataStage workflows into machine-readable lineage, cutting migration time by 85%.
Work Experience
Senior Data Engineer
STC
- Architected an automated Oracle offloading engine using multiple techniques to reduce storage consumption by 30%+, combining retention policy design with pipeline automation.
- Offloaded Jawwy's ODS to Hadoop via Sqoop with built-in retry and validation logic, freeing 6x storage capacity while preserving full data integrity.
- Designed a modular Teradata deletion framework unifying historical and incremental offload pipelines, reclaiming 50+ TB of storage.
- Scaled the deep analysis profiling phase from 20% to 80% table coverage across 100+ tables through automated vertical profiling.
- Integrated Trino with Teradata via TIBCO DV views, cutting query latency by 80% and eliminating data duplication.
- Archived 200+ TB of cold data from C5 cluster logs to MinIO, materially improving cluster storage performance.
- Led access management coordination with STC stakeholders, resolving 10+ critical requests to minimize operational downtime.
Airtable Developer
Verde Valley Resources, LLC
- Replaced a legacy CSV-to-Airtable workflow with a production-grade PostgreSQL system, processing 39,000+ mineral rights records with complete ownership history tracking.
- Designed a normalized schema (5 core tables, 6 semantic views) paired with a 3-stage Python ETL pipeline and hash-based deduplication, eliminating duplicate records at scale.
- Built a 24-assertion automated test suite with adaptive Python-based ID resolution, driving pipeline failure rates to zero and delivering a clean Airtable sync layer for business users.
Data Engineer
Paymob
- Migrated 5M+ records from Paxstore and Gonisight, enabling in-house geolocation and removing dependency on Google Maps.
- Architected ETL and Reverse-ETL pipelines across 10+ projects, streamlining data flow across systems.
- Optimized Redshift clusters (Spark, Hive, S3, Presto), cutting infrastructure costs by 20% and doubling query speed.
- Launched automated CleverTap campaign triggers reaching 1,500+ POS users, reducing support complaints by 25%.
Big Data Engineer
Etisalat Egypt
- Centralized fragmented metadata into DataHub, improving discoverability and integration for 100+ engineers.
- Built a Python-based data quality engine with automated KPIs, cutting data errors by 50%.
- Automated metadata extraction from Informatica/DataStage, reducing migration time by 85% and enabling a full transition to Airflow/NiFi.
- Developed scalable Python/SQL pipelines that saved 100+ hours of manual data preparation monthly.
Data Engineer and Analyst
Worldie
- Engineered distributed data systems and pipelines, reducing processing time by 90%.
- Applied NLP and text analytics on large-scale social data, uncovering 100+ trends that informed research and strategic decisions.
- Led data science team, implementing version control practices that enhanced collaboration and code documentation standards.
Data Engineer & Analyst
Freelancing Agency
- Delivered 120+ end-to-end data projects with a 100% success rate and 5-star average rating across 95+ verified reviews, serving clients across the US, UK, Australia, Belgium, and Nigeria.
- Built end-to-end ETL and reverse-ETL pipelines in Python, SQL, Spark, and Hive, cutting processing time 2x through optimized data modeling and query tuning.
- Built a custom data quality engine improving reliability by 50%, eliminating manual validation overhead for a high-volume analytics client.
- Delivered 15+ interactive dashboards (Power BI, Tableau, Plotly Dash) covering financial, marketing, sports, and social media analytics for international clients.
- Built multi-page Power BI dashboards with DAX measures, drill-through pages, cross-filtering, and custom visuals—enabling non-technical stakeholders to self-serve insights across 5+ KPI categories.
- Built multi-page Power BI dashboards with DAX measures, drill-throughs, and cross-filtering, enabling self-serve insight across 5+ KPI categories.
- Developed advanced Tableau workbooks with LOD expressions, blended sources, and dynamic parameters for segment-level and time-series analysis.
- Built production-grade Plotly Dash and Matplotlib visualization tools, replacing costly third-party subscriptions and saving one client $500+/month.
- Applied NLP (sentiment analysis, topic modeling, entity extraction) to millions of records from YouTube, Instagram, Reddit, and Twitter, generating structured BI reports.
- Delivered deep learning solutions (CNNs, RNNs) with Amazon SageMaker deployment, bridging data pipelines into ML model serving.
Physics and Mathematics Teacher
Town School
- Helped students achieve high grades through structured lessons and problem-solving techniques.
- Developed personalized learning strategies that significantly improved students' understanding and performance.
- Taught over 100 students in mathematics and physics, covering topics from basic to advanced levels.
Experience
Data Warehouse for Sparkify
https://github.com/adelabuhashim/DataWarehouseAs a data engineer, this project involved building an ETL pipeline that extracts data from Amazon S3, stages it in Redshift, and transforms data into a set of dimensional tables for analytics team to continue finding insights into what songs their users are listening to. The project also involved testing the database and ETL pipeline by running SQL queries.
Education
Bachelor's Degree in Computer Engineering
Zagazig University - Zagazig, Egypt
Certifications
Fabric Data Engineer (DP-700)
Microsoft
Certified Data Management Professional (CDMP)
DAMA International®
AWS Certified Data Engineer
Amazon Web Services
Astronomer Certification for Apache Airflow Fundamentals
Astronomer
Astronomer Certification DAG Authoring for Apache Airflow
Astronomer
Data Architect Nanodegree
Udacity
Data Engineering Nanodegree
Udacity
Data Analyst Nanodegree
Udacity
Skills
Libraries/APIs
Pandas, PySpark, REST APIs, NumPy, Matplotlib, X (formerly Twitter) API, Stripe
Tools
Apache Airflow, DataHub, PyCharm, Terminal, AWS Glue, Microsoft Excel, Microsoft Power BI, Tableau, GitHub, AWS Transfer Family, Amazon Elastic MapReduce (EMR), GitLab, Impala, Cloudera, Plotly, Git, Microsoft Access, Looker, BigQuery, Stitch Data, Pytest, Claude Code, Apache Impala, AWS Command Line Interface (CLI), AWS IAM, IBM InfoSphere (DataStage), Apache Sqoop, IBM Watson, Amazon Kinesis Data Firehose, Amazon CloudWatch, Amazon EKS, Amazon Redshift Spectrum, AWS Step Functions, Amazon Athena, Amazon QuickSight, Amazon Simple Notification Service (SNS), Amazon Simple Queue Service (SQS), AWS DataSync, Power Query, Consent Management Platforms
Languages
Python 3, SQL, Python, Snowflake, Arabic, HTML, Stored Procedure, Bash Script
Paradigms
ETL, Automation, Database Design, Business Intelligence (BI), Role-based Access Control (RBAC), Dimensional Modeling, Load Testing, OLAP
Platforms
Windows, Visual Studio Code (VS Code), Amazon Web Services (AWS), Azure, Databricks, Talend, Linux, Apache Kafka, Docker, AWS Lambda, Google Cloud Platform (GCP), MacOS, Jupyter Notebook, Oracle, Kubernetes, Microsoft Fabric, Google Dataform, Fabric Lakehouse, AWS IoT
Storage
Redshift, PostgreSQL, Amazon S3 (AWS S3), Data Lakes, Data Pipelines, JSON, Relational Databases, Database Management, Data Integration, Database Architecture, Databases, Apache Hive, HDFS, Amazon DynamoDB, SQLite, Distributed Databases, Data Validation, Google Cloud SQL, Google Cloud Storage, SQL Server Reporting Services (SSRS), Google Cloud, OLTP, Teradata, SQL Stored Procedures, Amazon Aurora, NoSQL, Operational Data Store (ODS), Elasticsearch, Microsoft SQL Server, Master Data Management (MDM), MongoDB, Microsoft Entra ID
Frameworks
Spark, Hadoop, Apache Spark, Presto, Trino, Streamlit, Data Lakehouse
Other
Data Modeling, Data Warehousing, Big Data, Cloud, Data Engineering, API Integration, Data Build Tool (dbt), Dashboards, Amazon Redshift, Data Mesh, Data Processing, APIs, Google BigQuery, Airtable, Relational Database Design, Amazon Managed Workflows for Apache Airflow (MWAA), Data, ELT, CI/CD Pipelines, Migration, Data Analytics, Artificial Intelligence (AI), Performance Optimization, ETL Pipelines, CSV File Processing, Large Data Sets, Database Schema Design, Analytics, Database Normalization, Cloud Storage, EMR, AWS Database Migration Service (DMS), Data Governance, Data Transformation, MDM, System Administration, Data Analysis, Data Migration, Data & Backup Management, Data Visualization, Azure Data Factory (ADF), Azure Databricks, Star Schema, Apache Superset, Reports, Reporting, Data Cleaning, Architecture, Data Marts, Data Quality, Key Performance Indicators (KPIs), Airtable Automations, Machine Learning, Medallion Architecture, Query Optimization, AI Tools, Data Classification, Data-driven Decision-making, Social Networks, Identity & Access Management (IAM), Enterprise Data Warehouse (EDW), Data Warehouse Design, Financial Data, Finance, Observability, Process Mining, Data Mining, ETL Testing, Microsoft Azure, Software Engineering, Computer Networking, Computer Engineering, Designing Data Systems, Apache Cassandra, Normalization, Data Lineage, Data Architecture, MinIO, Mathematics, Physics, IT Service Management (ITSM), TIBCO, NiFi, Informatica, Data Migration Testing, Data Science, Large Language Models (LLMs), DAG Authoring, Scheduling, Orchestration, Directed Acrylic Graphs (DAG), Data Management, Statistics, Data Wrangling, Natural Language Processing (NLP), Sentiment Analysis, Text Analytics, Research, Amazon Kinesis, Lambda Functions, Amazon RDS, Amazon MSK, Amazon EventBridge, Amazon AppFlow, Amazon MemoryDB for Redis, Data Center Infrastructure, SOX, Business Analysis, Product Lifecycle Management (PLM), Metadata, Data Strategy, Big Data Architecture, Reverse-ETL, Dash, Looker Studio, Deep Learning, Data Security, DAX, Security, Polars, Regression Testing, activator, Streaming, OpenAI
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring