
Igor Gorbenko
Verified Expert in Engineering
Database and Back-end Developer
Dubai, United Arab Emirates
Toptal member since October 18, 2021
Igor is a CTO, Data/ML Architect, and engineering leader with 16+ years of experience building high-load data, AI, and cloud platforms. He has designed systems processing 2+ PB of data, 10+ TB daily, and up to 300K RPS, and led teams of up to 30 engineers. At TangoMe, his recommendation and anti-fraud platforms drove significant revenue growth. Today, he architects enterprise AI and compliance platforms with LLMs, RAG, Kubernetes, and cloud technologies, from strategy to production delivery
Portfolio
Experience
- Data Pipelines - 16 years
- SQL - 13 years
- Python - 10 years
- Amazon Web Services (AWS) - 8 years
- Big Data Architecture - 6 years
- Big Data - 6 years
- Google Cloud Platform (GCP) - 5 years
- Machine Learning Operations (MLOps) - 4 years
Preferred Environment
PyCharm, Slack, Linux, Git
The most amazing...
...thing I've built: a 2PB lakehouse processing 10TB daily and powering a high-load AdTech platform handling up to 300K requests per second
Work Experience
Chief Technology Officer
Compliance Control
- Architected and led the development of a multi-tenant compliance platform supporting both SaaS and enterprise on-premise deployments.
- Designed and delivered a private AI Agent platform and AI Compliance Copilot for secure LLM usage, intelligent document search, evidence analysis, and audit automation.
- Built and managed engineering teams across backend, frontend, and DevOps, owning platform architecture, infrastructure, security, and engineering standards.
- Partnered with founders and enterprise customers to translate complex regulatory and business requirements into scalable technical solutions.
- Delivered the initial MVP within three months and launched both SaaS and enterprise deployment models.
Head of Data | Software Architect
Omniverse
- Built the Data Department and defined the company-wide data strategy for an AdTech platform processing 2+ PB of data and up to 300K requests per second.
- Architected DWH, DSP, and DMP platforms on AWS and implemented a scalable Data Lakehouse based on Apache Iceberg and Apache Spark.
- Reduced Snowflake costs from $30K to $10K per month through architecture and workload optimization.
- Built a Data Intelligence Service supporting complex investigations, anomaly detection, and relationship discovery across large-scale datasets.
- Designed and deployed a fraud detection platform that reduced operational costs by approximately 20%.
- Established Data Engineering and MLOps architecture, standards, and engineering practices across the organization.
Head of ML Engineering | Big Data Architect
Tango
- Co-founded the Data Department and led ML Engineering, scaling the organization to up to 30 engineers.
- Architected and delivered a real-time recommendation platform responsible for significant revenue growth across the product.
- Built high-throughput real-time AI inference pipelines for large-scale recommendation systems.
- Developed anti-fraud systems that reduced false-positive rates by more than 80%.
- Built a real-time video content labeling platform supporting 10K+ concurrent video streams.
- Established Data and MLOps engineering practices and participated in the company Architecture Committee.
Big Data Architect
Netwrix
- Designed and developed a User and Entity Behavior Analytics (UEBA) platform for anomaly detection that became a key product capability.
- Migrated large-scale anomaly detection workloads from standalone Docker containers to distributed Apache Spark on AWS EMR, significantly improving processing performance.
- Reduced AWS infrastructure costs severalfold by implementing dynamic EMR cluster provisioning and workload-aware resource allocation.
- Designed monitoring, reporting, and alerting capabilities for production data and ML workloads.
- Established infrastructure automation and CI/CD practices using Terraform and cloud-native tooling.
Lead Big Data Developer
First Line Software
- Designed and implemented end-to-end ETL pipelines for transforming healthcare data into the OMOP Common Data Model (CDM).
- Built an automated data-transformation framework using Python, SQL, and Apache Spark.
- Developed data quality and visualization tooling for analyzing converted healthcare datasets.
- Worked across AWS and GCP environments with large-scale analytical datasets.
Senior Software Developer
Fujitsu Global
- Designed and developed automated incident-routing and operational analytics solutions for enterprise environments.
- Built project tracking and reporting systems used to improve operational visibility and workflow management.
- Migrated enterprise billing reporting workloads to Microsoft SQL Server Reporting Services.
- Worked with large enterprise systems using SQL Server, Informix, Linux, Bash, and data-processing technologies.
Deputy Head of Analytical Department
Gazprombank
- Led the development of analytical and management reporting systems supporting business and financial decision-making.
- Designed and implemented an automated retail FX pricing system, improving exchange-rate management and reducing currency risk.
- Built planning, forecasting, and performance-monitoring systems for business operations.
- Combined hands-on software development with ownership of analytical systems and technical leadership.
Experience
Compliance App - AI-Powered Compliance and Audit Platform
I serve as CTO and Software Architect and own the technical architecture of the platform across application development, infrastructure, security, data, and AI capabilities. I designed a multi-tenant architecture supporting both SaaS and on-premise enterprise deployments. The platform includes audit management, evidence collection and validation, asset and risk management, vulnerability management, workflow automation, reporting, and integrations with external enterprise systems.
One of the key areas I designed is the platform's AI layer. It includes a private AI assistant based on LLM and RAG technologies that can work with regulatory documentation, audit evidence, internal policies, and platform data. The architecture is designed to support private and on-premise deployments where sensitive customer data cannot be transferred to public AI services.
My role combines CTO responsibilities, solution architecture, AI architecture, engineering management, and direct collaboration with customers on complex technical and regulatory requirements.
Enterprise Data Lakehouse for Omniverse
The architecture was built on Amazon S3 and Apache Iceberg, with Apache Spark as the primary distributed processing engine. The platform processed up to 10 TB of new data per day and supported both batch and near-real-time ingestion patterns.
I designed two major ingestion paths. The batch layer handled high-volume historical processing, transformation, and aggregation workloads, while the streaming layer consumed continuously generated events and made fresh data available for analytics and downstream services.
Apache Iceberg provided a scalable table layer with schema evolution, partition management, and transactional data operations, while Spark was used for complex transformations and large-scale distributed computation.
As Head of Data and Software Architect, I was responsible not only for the technical architecture but also for the overall data strategy, engineering standards, infrastructure decisions, and integration of the Lakehouse with analytical and operational systems.
Real-Time Recommendation Platform for Tango
https://www.tango.me/live/recommendedThe system continuously analyzed user behavior and content signals to select and rank the most relevant live content for each user. The architecture was designed for low-latency online inference and high-throughput event processing.
The platform was built primarily on Google Cloud and used Bigtable, BigQuery, Dataflow, Redis, and Python-based services for feature processing, candidate generation, ranking, and real-time recommendation delivery.
I owned the architecture and engineering delivery across the data, machine learning, and cloud layers and led the ML Engineering organization responsible for the platform.
In addition to the recommendation platform, my team developed real-time anti-fraud and video content analysis systems. The recommendation initiatives contributed materially to Tango's revenue growth, while the anti-fraud platform reduced false-positive rates by more than 80%.
My responsibilities included architecture, engineering leadership, MLOps practices, production reliability, team development, and collaboration with product and business stakeholders.
Healthcare Data Transformation Platform Based on OMOP CDM
https://www.ohdsi.org/data-standardization/the-common-data-model/The main engineering challenge was supporting datasets originating from multiple technologies and storage systems, including Amazon S3, Google Cloud Storage, Hadoop HDFS, PostgreSQL, Amazon Redshift, and other relational and distributed data platforms.
We developed a reusable framework that automated the preparation and execution of large-scale ETL pipelines using Python, Apache Spark, and SQL.
The framework reduced the amount of manual work required to onboard new datasets and provided standardized orchestration of transformation stages. It also supported automated post-processing steps such as data validation, unit tests, statistics generation, and quality reports.
As Technical Lead, I designed core framework components, developed Python services, reviewed code, helped define ETL architecture, and participated in running and troubleshooting production data pipelines.
Education
Master's Degree in Information Technologies
Kazan National Research Technical University - Kazan, Russia
Certifications
AWS Certified Machine Learning - Specialty
AWS
SnowPro Core Certification
Snowflake
AWS Certified Solutions Architect Associate
AWS
Professional Cloud Architect
Google Cloud
Professional Data Engineer
Google Cloud
Associate Cloud Engineer
Google Cloud
AWS Certified Developer
PSI
AWS Certified Cloud Practitioner
PSI
Skills
Libraries/APIs
Pandas, Complex SQL Queries, PySpark, NumPy, Dropbox API, Google APIs
Tools
PyCharm, Git, Apache Airflow, Terraform, BigQuery, AWS Glue, Tableau, Microsoft Power BI, Apache Beam, Postman, Slack, Grafana, Amazon Cognito, Cloud Dataflow, GitLab, Apache NiFi, Google Kubernetes Engine (GKE), Spark SQL, Amazon Athena, Google Cloud Dataproc, Apache Iceberg, Amazon SageMaker, dbt Cloud
Languages
SQL, Bash, Python, Snowflake, Cypher, Scala, C#.NET, Excel VBA, Python 3
Paradigms
REST, ETL, Database Design, Role-based Access Control (RBAC)
Platforms
Linux, Amazon Web Services (AWS), Google Cloud Platform (GCP), Docker, Databricks, AWS Lambda, Azure, Apache Kafka, New Relic, Oracle, Cloud Run, Kubernetes, Yandex Cloud
Storage
Redshift, PostgreSQL, Microsoft SQL Server, Amazon DynamoDB, Data Pipelines, JSON, Databases, Amazon S3 (AWS S3), Google Cloud, NoSQL, Google Bigtable, IBM Informix, Cloud Firestore, ClickHouse, On-premise
Frameworks
Flask, Apache Spark, Django, Locust, Spark, Trino
Other
IT Systems Architecture, Google BigQuery, Big Data, Big Data Architecture, Data Architecture, Data Engineering, Analytics, Cloud, Data Analysis, Cloud Platforms, Data Visualization, Data Warehousing, Data Analytics, Foundry, Palantir, Amazon RDS, Database Schema Design, Distributed Systems, Data Migration, Data Modeling, Data Reporting, Large Language Models (LLMs), Artificial Intelligence (AI), Data Governance, Query Optimization, FinOps, FastAPI, Redis Clusters, Machine Learning Operations (MLOps), Machine Learning, Data Build Tool (dbt), Pub/Sub, Investments, Stock Market, Google Cloud Functions, EMR, Data Science, Workday, Enterprise Cybersecurity, RAG Architecture, RAG Systems, Engineering Management, Team Leadership, Software Architecture, MinIO
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring