
Faris Hambo
Verified Expert in Engineering
Data Engineer and Developer
Sarajevo, Bosnia and Herzegovina
Toptal member since July 7, 2026
Faris is a senior data engineer and data scientist with 7 years of experience delivering production cloud-based ETL pipelines and ML models for enterprise clients in healthcare, automotive, and SAP/ERP sectors. He specializes in Python, AWS, and SQL. While at Symphony, Faris designed and deployed ML models and ETL pipelines processing 150+ terabytes of data for US healthcare and EU automotive clients.
Portfolio
Experience
- Python - 8 years
- Machine Learning - 8 years
- ETL - 6 years
- SQL - 6 years
- AWS Glue - 5 years
- Data Engineering - 5 years
- PySpark - 5 years
- Terraform - 4 years
Preferred Environment
Terraform, Python, SQL, PySpark, ETL, Applied Mathematics, Data Engineering, Machine Learning Operations (MLOps)
The most amazing...
...work I've done is design and deliver an e2e AWS-based ETL pipeline processing 150+ terabytes with data quality checks, orchestration, testing, and monitoring.
Work Experience
Senior MLOps Engineer
ELFA Data
- Built a GenAI-powered enterprise data assistant platform enabling business users to query governed data using natural language.
- Worked across agent orchestration, semantic data layers, MCP-based services, vector databases, and end-to-end deployment of LLM systems.
- Focused on reliability, AI observability, scalable production delivery, and enterprise-grade GenAI architecture.
- Developed retrieval-augmented generation (RAG) pipelines and integrated enterprise AI services, including Amazon Bedrock and Anthropic Claude.
Teaching Associate
Politehnički fakultet Univerziteta u Zenici
- Delivered lectures and graded assignments for undergraduate and graduate students across 4 courses: Introduction to Object-Oriented Programming, Data Mining, Data Structures and Algorithms, and Introduction to Computational Geometry.
- Mentored student projects that bridge mathematical theory and applied data science, helping students move from coursework to practical ML implementations.
- Mentored and guided student research, resulting in the publication of a research paper.
Senior Data Engineer | Data Scientist
Symphony
- Designed and deployed scalable ML models and ETL pipelines on AWS (SageMaker, Glue, Lambda, S3, Redshift, RDS) for US healthcare, processing 150+ terabytes of data and powering analytics.
- Led data and ML engineering on automotive predictive models (car parts replacement prediction). Applied the GenAI approach for generating synthetic data, increasing coverage from 15% to 95%.
- Automated infrastructure with Terraform and orchestrated complex ETL workflows via AWS Step Functions and Apache Airflow, cutting deployment and pipeline-recovery time and enabling reliable enterprise SAP/ERP data ingestion.
- Mentored 5+ junior data scientists and engineers across multiple projects.
Data Scientist
Infinity Mesh
- Built Azure-based ETL pipelines, dashboards, and visualizations (Azure ML Studio, Azure Data Factory), improving data-processing efficiency by 20-25% and enabling faster analytics workflows.
- Developed and deployed a fraud-detection ML model for a production application handling approximately 15,000 concurrent users, reducing false positives and supporting real-time risk decisions.
- Collaborated with engineering, product, and business teams to deliver operational monitoring and data-driven features end-to-end.
Data Scientist | Analyst
E387
- Built and deployed forecasting ML models predicting electric-vehicle station availability with an over 85% accuracy, improving utilization forecasts.
- Designed scalable data pipelines and operational dashboards used for real-time monitoring and reporting.
- Delivered analytics and forecasting features adopted by a platform serving 200,000+ end users, improving user experience and operational planning.
Experience
AI-DLC
https://github.com/FarisHambo-Node/AI-DLCI designed a queue-based flow that moves work from requirements through implementation, testing, security review, deployment, and post-deployment incident feedback. The repository includes an initial Python harness that validates task contracts, loads task-specific skills and context, routes model calls, evaluates acceptance criteria, and applies hard guardrails around protected branches, CI failures, and production approvals. I built the project to explore auditable, human-gated agent workflows and a clear separation between LLM judgment and deterministic execution. It is an early-stage technical prototype rather than a production deployment.
Data Orchestration Comparator
https://github.com/FarisHambo-Node/Symplhy-OrdeqThe project includes a classical ML workflow covering data preparation, feature engineering, train/test splitting, model training, evaluation, persisted metrics, and visual outputs. It also includes an LLM text-classification workflow with dataset preparation, Hugging Face model loading, inference, prediction analysis, and unit tests. I implemented reusable data catalogs and datasets, custom model IO, execution timing and logging hooks, command-line entrypoints, Mermaid/Kedro visualization, and isolated tests for pipeline nodes. The project was built as a technical evaluation of maintainable ML pipeline patterns rather than as a distributed production orchestrator.
Education
PhD in Applied Mathematics and Computer Science
University of Sarajevo - Sarajevo, Bosnia and Herzegovina
Master's Degree in Applied Mathematics
University of Sarajevo - Sarajevo, Bosnia and Herzegovina
Bachelor's Degree in Theoretical Mathematics
University of Sarajevo - Sarajevo, Bosnia and Herzegovina
Skills
Libraries/APIs
PySpark
Tools
AWS Glue, Jupyter, Terraform, Amazon SageMaker, AWS Step Functions, Amazon CloudWatch, GitHub, Apache Airflow, Plotly, BigQuery
Languages
Python, SQL, C++, Snowflake
Paradigms
ETL, HIPAA Compliance, Model Context Protocol (MCP), Object-oriented Programming (OOP)
Platforms
Jupyter Notebook, AWS Lambda, Google Cloud Platform (GCP)
Storage
Data Validation, Data Pipelines, Relational Databases, Amazon S3 (AWS S3), Redshift, Databases, Google Cloud
Industry Expertise
Healthcare
Frameworks
Chainlit, Bedrock, LangGraph, Kedro
Other
Machine Learning, Data Engineering, Calculus, Workflow Automation, Database Normalization, ETL Pipelines, Amazon RDS, Big Data, Data Mining, Statistics, Numerical Methods, Applied Mathematics, Machine Learning Operations (MLOps), Linear Algebra, Metaheuristics, Deep Learning, Analytics, API Integration, Data Analytics, EMR, Data Cleaning, Data Transformation, Query Optimization, Data Modeling, Database Schema Design, Audits, Data Warehouse Design, HIPAA, Reporting, Dashboards, Electronic Health Records (EHR), Microsoft Office, Large Language Models (LLMs), Amazon Bedrock, Qdrant, DuckDB, Retrieval-augmented Generation (RAG), Prompt Engineering, Step Functions, SAP, Agent Orchestration, Observability, Model Validation, Stochastic Differential Equations, Fourier Analysis, Dynamic Systems Modeling, Partial Differential Equations, Integral Transformations, Spectral Theory, Graph Theory, Time Scale Calculus, Combinatorics, Probability Theory, Topology, Differential Equations, Algebra, MLflow, Fraud Detection, LangChain, GitHub Actions, Queue Management, Modeling, Time Series Analysis, Ordeq, Data Analysis, Financial Reporting Dashboards, Google BigQuery, Zoho
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring