Mohamed Ahmed, Developer in Berlin, Germany
Mohamed is available for hire
Hire Mohamed

Mohamed Ahmed

Big Data Architect and Back-end Developer

Berlin, Germany

Toptal member since November 4, 2022

Bio

Mohamed is a senior data platform architect and staff data engineer with 19+ years of experience designing and scaling cloud data platforms for product companies. He specializes in lakehouse/medallion architectures, distributed processing, and governed data products, treating the platform as a product with clear contracts and great DX. He works fully remote with distributed teams, bridging strategy and execution as a hands-on architect and mentor.

Portfolio

mobile.de
Apache Spark, Scala, Apache Kafka, BigQuery, Google Cloud Platform (GCP)...
Careem Networks FZ
Apache Spark, Apache Airflow, Scala, Python, Amazon Elastic MapReduce (EMR)...
Searchmetrics gmbh
Amazon Web Services (AWS), RabbitMQ, Apache Spark, Apache Zeppelin, Java...

Experience

  • SQL - 11 years
  • Data Engineering - 7 years
  • Scala - 7 years
  • Apache Spark - 7 years
  • BigQuery - 5 years
  • Apache Kafka - 5 years
  • Python - 4 years
  • Google Cloud Platform (GCP) - 3 years

Preferred Environment

Google Cloud Platform (GCP), Amazon Web Services (AWS), Big Data Architecture, Apache Spark, Data Build Tool (dbt), BigQuery, Snowflake, Airbyte, Databricks

The most amazing...

...thing was designing and leading a GDPR-aligned, privacy-by-design data lake at eBay, unifying dozens of products into a GCP lakehouse with six-figure savings.

Work Experience

Big Data Architect

2020 - 2024
mobile.de
  • Architected 500+ TB multi-cloud data platform (GCP → AWS) across 25 product teams and 44 projects; designed the data lake, query engine, BI layer, and cloud cost governance with FinOps frameworks.
  • Designed ML infrastructure for distributed training and inference on GPU clusters; built GDPR-compliant data privacy governance and cloud cost FinOps frameworks.
  • Pioneered domain-driven data mesh architecture, built real-time/batch data infrastructure, and chaired cross-company architecture councils; established self-service platform blueprints adopted enterprise-wide.
Technologies: Apache Spark, Scala, Apache Kafka, BigQuery, Google Cloud Platform (GCP), Python, Delta Lake, General Data Protection Regulation (GDPR), Apache Airflow, Kubernetes, Apache Cassandra, PostgreSQL, Linux, Data Pipelines, GitHub, Data Modeling, Data Architecture, Data Governance, Big Data, Big Data Architecture, Database Architecture, Solution Architecture, Data Build Tool (dbt), Cloud Migration, Data Migration, Infrastructure as Code (IaC), Data Lake Design, Docker, Terraform

Staff Data Engineer

2018 - 2020
Careem Networks FZ
  • Developed an ETL framework that automates data processing pipelines and runs hundreds of ETL jobs daily.
  • Optimized an ETL pipeline, which saved hundreds of thousands of dollars yearly.
  • Established a data academy to help full-stack engineers grow their data engineering knowledge.
Technologies: Apache Spark, Apache Airflow, Scala, Python, Amazon Elastic MapReduce (EMR), Apache Hive, Data Pipelines, ELT, Amazon S3 (AWS S3), SQL, GitHub, Data Architecture, Amazon Web Services (AWS), Data Orchestration

Senior Software Engineer and Big Data

2017 - 2018
Searchmetrics gmbh
  • Designed and developed a challengeable data pipeline for billions of messages and records.
  • Devised the ETL framework to work with many sources and sinks.
  • Presented new technologies and discussed them with my team.
Technologies: Amazon Web Services (AWS), RabbitMQ, Apache Spark, Apache Zeppelin, Java, Apache Kafka, Spark Streaming, Apache Flink, MySQL, Hadoop, Apache Hive, Scala, Spring Boot, Data Processing, Linux, Data Engineering, Data Pipelines, ETL, ELT, Amazon S3 (AWS S3), Data Structures, Amazon Elastic MapReduce (EMR), SQL, Jira, GitHub, Data Modeling, Spark, User-defined Functions (UDF), Big Data, Data Integration, Microservices, NoSQL, Database Architecture, Apache Maven, EMR, RESTful Microservices, AWS Certified Solution Architect, Amazon EC2, Amazon Virtual Private Cloud (VPC), DevOps, Data Lake Design, Security, Docker, ETL Tools, Distributed Systems, Data Lakes, Performance Optimization, ETL Pipelines, Data Orchestration, CI/CD Pipelines, Distributed Databases, Application Architecture, Scalability, Orchestration, Automation, Enterprise, DataOps, Cloud Architecture, Databases

Senior Software Engineer and Big Data

2015 - 2017
Agoda
  • Developed and tuned recommendation microservices to accept millions of requests per second with a success rate of 99.99 in 20 milliseconds for the whole request trip.
  • Designed and developed a reactive DAG framework to build any logical flow over Akka actors and futures.
  • Assisted the data scientist team with ETL pipelines to apply ML offline training.
Technologies: Apache Spark, Hadoop, Scala, Akka, Data Processing, Apache Cassandra, PostgreSQL, Linux, Data Engineering, Data Pipelines, ETL, ELT, Data Structures, SQL, Jira, Data Modeling, Spark, User-defined Functions (UDF), Big Data, APIs, Data Integration, Microservices, Database Architecture, Apache Maven, Oozie, RESTful Microservices, Machine Learning Operations (MLOps), Security, ETL Tools, Back-end, Machine Learning, Distributed Systems, Performance Optimization, ETL Pipelines, Data Orchestration, Distributed Databases, Application Architecture, Scalability, Orchestration, Enterprise, Cloudera, DataOps, Databases

Back-end Specialist

2014 - 2015
CIT global
  • Developed a logging service that tracks all app actions on MongoDB with AspectJ.
  • Developed an e-payment workflow using Mule ESB that controls payment steps.
  • Created a wallet payment microservice that transfers payments across bank accounts.
Technologies: Java, MongoDB, Hibernate, Oracle Database, SQL, Apache Cassandra, Stored Procedure, APIs, Microservices, Database Architecture, Apache Maven, RESTful Microservices, Back-end Development, Back-end, Distributed Systems, Databases, Full-stack Development

Senior Java Developer

2010 - 2014
E-Finance
  • Developed back-end and front-end payment services using multiple frameworks; ADF, Struts, and ICEfaces.
  • Created business reports using the Jasper Reporting tool.
  • Built and automated administration pages created from a DB ER diagram.
Technologies: Oracle PL/SQL, Oracle Database, SQL, Web Services, Data Modeling, Stored Procedure, APIs, Apache Maven, Back-end Development, Back-end, Full-stack Development

Service Information Developer

2009 - 2010
HP Inc
  • Built the endpoint of the sales (EPOS) client app validator using Servlet and JSP, which can validate big XML files and return invalid tags.
  • Wrote a user tutorial that guided users to new features and increased the customer acceptance rate.
  • Contributed to the internal development community that helped new users get familiar with internal tools.
Technologies: Hibernate, Web Services, APIs, Back-end Development, Back-end, Full-stack Development

Java Developer

2007 - 2009
Networks Valley
  • Created a custom payroll desktop app that handled complicated payroll logic and generated company payroll reports.
  • Devised an innovative home service that monitored smart homes and sent mobile notifications to homeowners.
  • Built a PCL interface app that controlled devices in an electricity plant.
Technologies: Hibernate, Microsoft SQL Server, Java, SQL, Back-end Development, Back-end, Enterprise Resource Planning (ERP), Full-stack Development

Experience

On-premises to Public Cloud (GCP) Migration

We experienced various limitations in the private cloud, so we decided to move to the public cloud (GCP). I took the initiative to design, lead, and build migration across the company and be an excellent example for new data platform infrastructure.

I collected and discussed pain points with stakeholders, created general architecture ADR, reviewed the new design with my team and stakeholders, and collected feedback. Next, I estimated the budget and discussed it with the head of technology, removed obstacles to implementation, and modified open-source frameworks to fit our needs; for example, I added a new feature to the Atlas data catalog framework to support delta-lake. I reviewed the road map with my team and broke it down into epics and parallel stories, jumped in to help when blocks arose, and discussed best practices with different teams in the company from a data point of view.

Real-time Analytics Service

This real-time analytics tool extracts user-tracking metrics from event streams depending on configurable input using Kafka and Spark-streaming frameworks.

I designed a real-time solution that fulfilled stakeholders' requirements, tuned the reader service to achieve <100 milliseconds latency in the 99.99 success rate percentile, and introduced network solutions as the project ran in a hybrid cloud environment.

Building Marketplace Data Platform

I worked across teams in three countries to define common pain points and introduce new solutions. This included creating the POC/RFC for new solutions, discussing solutions with teams, collaborating with teams and project managers to set the execution plan, mentoring data academy members, and building standard tools that accelerate development time.

Keyword Ranking

I designed and developed innovative stream and batch data processing projects to improve the ranking of our clients with millions of keywords for many search engines in many countries with many languages.

I created a changeable data pipeline for billions of messages and records and designed the ETL framework to work with many sources and sinks. I presented new technologies and discussed them with the team. I tuned the jobs to fit our cluster, reviewed the code, and took ownership of the project.

Reactive Framework (Jarvis)

Jarvis is a reactive DAG framework to build logical flow over Akka actors and futures. I reviewed the requirements with all teams involved to simplify and remove repeated functions. I designed and built a DAG solution to streamline and service our business logic units in reactive mode. I presented the solution and helped teams use it.

Hotel Recommendation

I designed and developed an innovative project to rank hotels based on user preferences. I assisted data scientists in collecting the data to apply the offline training and designed and built a solution replicating the ALS model to five data centers. I created a distributed and local cache solution to hold the ML model in memory. This achieved a four-millisecond response time in the worst-case scenarios. I developed a load balancer between servers in the same data center, applied our DAG framework (Jarvis) to build our ranking service, and tuned the microservice to accept millions of requests with a success rate of 99.99. I then audited the customer interaction with the microservice to use it for model evaluation and configured the deployment scripts for production and stage servers.

Bidding Channel ROI Manager

This is a back-end project to manage the bidding channel jobs for companies such as Google and TripAdvisor. I reviewed the design with the software architect, and built the dynamic implementation for channels, accounts, sync data, and the Oozie and HDFS clients. I built the migration scripts and configured the deployment environment for production.

Education

2002 - 2007

Bachelor's Degree in Electrical Engineering

Fayoum University - Egypt

Certifications

DECEMBER 2018 - PRESENT

Algorithms on Graphs

Coursera

JUNE 2018 - PRESENT

Deep Learning Specialisation

Coursera

FEBRUARY 2017 - PRESENT

Data Structures

Coursera

JANUARY 2017 - PRESENT

Algorithmic Toolbox

Coursera

JULY 2014 - PRESENT

OCE Java EE 6 EJB 3.x (1Z0-895)

Oracle

MAY 2014 - PRESENT

OCE Java EE 6 Web Service (1Z0-897)

Oracle

NOVEMBER 2013 - PRESENT

OCE Java Persistence API 2.0 - EE 6 (1Z0-898)

Oracle

MAY 2008 - PRESENT

Sun Certified Web Component Developer SCWCD 5 (310-083)

Sun Microsystems

AUGUST 2006 - PRESENT

Sun Java5 Certified SCJP 5 (310-055)

Sun Microsystems

Skills

Libraries/APIs

PySpark, Spark Streaming, Pandas, Cloud Key Management Service (KMS)

Tools

BigQuery, Apache Airflow, Apache Maven, Composer, Kafka Streams, Jira, GitHub, Apache Zeppelin, Amazon Elastic MapReduce (EMR), Oozie, Terraform, Apache Iceberg, Cloudera, RabbitMQ, AWS Glue, Amazon Virtual Private Cloud (VPC), dbt Cloud

Languages

Java, Scala, Python, SQL, Snowflake, Rust, Stored Procedure, Go

Frameworks

Apache Spark, Spark, Hadoop, Akka, Presto, Spring Boot, Hibernate

Paradigms

ETL, Application Architecture, Automation, Real-time Systems, Role-based Access Control (RBAC), Data-driven Design, DevOps, Microservices

Platforms

Google Cloud Platform (GCP), Apache Kafka, Amazon Web Services (AWS), Docker, Kubernetes, Linux, Jupyter Notebook, Apache Flink, Oracle Database, AWS Lambda, Amazon EC2, Cloud Run, Databricks, Airbyte

Storage

Data Pipelines, NoSQL, Google Cloud, Data Lake Design, Google Cloud Storage, Data Lakes, Databases, Google Cloud SQL, MySQL, PostgreSQL, Apache Hive, MongoDB, Oracle PL/SQL, Microsoft SQL Server, Amazon S3 (AWS S3), Data Integration, Amazon DynamoDB, Database Architecture, Data Validation, Distributed Databases, Database Modeling

Other

Data Processing, General Data Protection Regulation (GDPR), Data Engineering, ELT, Data Architecture, Big Data, Big Data Architecture, Architecture, Solution Architecture, Google BigQuery, Cloud Migration, Data Migration, Data Analysis, Large-scale Data Migration, ETL Tools, Distributed Systems, Software Architecture, Performance Optimization, ETL Pipelines, Data Orchestration, Dataproc, Medallion Architecture, Scalability, Data Transformation, Orchestration, Enterprise, DataOps, Systems Design, Technical Documentation, Cloud Architecture, Web Services, Algorithms, Data Structures, Graph Algorithms, Delta Lake, Google Cloud Functions, Apache Cassandra, Apache Livy, Pub/Sub, Data Modeling, User-defined Functions (UDF), Data Governance, Data Strategy, APIs, EMR, RESTful Microservices, AWS Certified Solution Architect, Back-end Development, Machine Learning Operations (MLOps), Security, Back-end, Real-time Data, CI/CD Pipelines, Database Schema Design, Data Warehousing, Enterprise Resource Planning (ERP), Amazon Redshift, Streaming, Full-stack Development, Apache Flume, Technical Program Management, Data Build Tool (dbt), Infrastructure as Code (IaC), Amazon Kinesis, Machine Learning, CTO

Collaboration That Works

How to Work with Toptal

Toptal matches you directly with global industry experts from our network in hours—not weeks or months.

1

Share your needs

Discuss your requirements and refine your scope in a call with a Toptal domain expert.
2

Choose your talent

Get a short list of expertly matched talent within 24 hours to review, interview, and choose from.
3

Start your risk-free talent trial

Work with your chosen talent on a trial basis for up to two weeks. Pay only if you decide to hire them.

Top talent is in high demand.

Start hiring