
Vishant Thareja
Verified Expert in Engineering
Data Engineer and Developer
Dubai, United Arab Emirates
Toptal member since April 29, 2026
Vishant has 12+ years of experience building enterprise data platforms, lakehouses, real-time streaming, and RAG pipelines that make business data talk to LLMs. He started with Oracle and PL/SQL, moved through Cloudera and Hadoop, and is now deep in Azure Databricks, Snowflake, DBT, Delta Lake, Kafka, AWS Glue, and Databricks Vector Search. Vishant built data teams from scratch and delivered across 8+ business verticals.
Portfolio
Experience
- Azure Databricks - 10 years
- PySpark - 10 years
- Data Warehousing - 9 years
- Azure Data Factory (ADF) - 5 years
- Python - 5 years
- Spark Streaming - 3 years
- AWS Glue - 2 years
- NoSQL - 2 years
Preferred Environment
PySpark, Azure Data Factory (ADF), AWS Glue, Python, Apache Kafka, Spark Streaming, Apache Airflow, Databricks, SQL, Data Build Tool (dbt)
The most amazing...
...thing I’ve done is define data strategy and optimize jobs and cluster utilization to generate significant cost savings for companies.
Work Experience
Data Engineering Specialist
Al Futtaim Group
- Delivered RAG (retrieval-augmented generation) pipeline on Databricks. Implemented semantic chunking of enterprise documents, generated text embeddings via a managed embedding model, and indexed them in DB Vector Search for faster LLM search.
- Built a technical strategy to be followed for the best data solutions for the client and business requirements. Designed end-to-end data solutions for minimal error and cost-optimized utilization.
- Architected a unified enterprise Lakehouse on Azure Data Factory and Databricks (Delta Lake, Unity Catalog, Delta Live Tables) serving 8+ business verticals — reduced time-to-insight for analytics teams from days to under four hours.
- Built customer 360 segmentation pipelines unifying siloed data across Al-Futtaim's retail, automotive, and real-estate divisions; directly enabled personalized marketing campaigns estimated to lift conversion rates.
- Slashed infrastructure spend by around 20% by auditing idle Databricks clusters and redundant ADF pipelines, consolidating jobs into optimized Databricks Workflows with auto-scaling policies.
- Automated data quality and schema enforcement across all ingestion layers via Great Expectations + Databricks Workflows; cut data-quality incidents reaching downstream BI.
- Led a squad of engineers across ingestion, transformation, and BI domains; introduced structured sprint ceremonies and code-review gates that eliminated critical bugs in production.
- Served as primary technical interface to business stakeholders — translated business requirements into actionable architecture blueprints with agreed SLAs.
- Developed a cost-effective and efficient solution for recurring tasks like data promotion from one layer to another. Saved time by creating generic code reusable across diverse data engineering teams.
Senior Data Engineer
Avrioc
- Built scalable batch and real-time data pipelines on Azure (ADF, Databricks, Snowflake, and Synapse), integrating Kafka, APIs, and multiple databases to deliver reliable, high-quality data for business and operational use cases.
- Optimized existing ETL workloads by tuning performance, right-sizing resources, and eliminating unused components, resulting in significant cost savings while maintaining high system performance and stability.
- Built the data engineering function from zero — defined team charter, hired engineers, and delivered production data infrastructure.
- Designed a real-time anti-cheat detection pipeline using Spark Streaming + Kafka for the MyWhoosh racing app; system processes 2M+ events/hour with sub-second detection latency across 10,000+ concurrent users.
- Reduced Kafka topic enrichment processing time by 30% by introducing KSQL DB for dimensional lookups, replacing brittle PySpark micro-batch workarounds.
- Integrated 6 heterogeneous data sources (MongoDB, Kafka, Elasticsearch, MySQL, REST APIs, flat files) into a unified gold layer on Azure Databricks using Python, PySpark, SQL, Azure Data Factory, and Databricks. enabling the SD team to deliver models.
- Used Snowflake with virtual warehouse isolation to separate team workloads, and dbt to manage Silver/Gold transformations — incremental models, source freshness tests, and lineage that meant data issues got caught in the pipeline.
Senior Data Engineer
Mashreq
- Built and managed scalable data pipelines using Azure Data Factory, Snowflake, and Databricks to ingest and transform large-scale enterprise data, implementing Unity Catalog for centralized governance, access control, and improved data security.
- Developed distributed data processing workflows using PySpark and Spark Streaming, enabling efficient handling of large datasets while ensuring governed, secure, and high-performance data transformations.
- Designed and maintained real-time streaming solutions using Spark Streaming, improving data freshness and enabling faster decision-making through near real-time processing of business-critical data.
- Integrated and transformed data from Oracle, SQL Server, and CSV sources across Customers, Accounts, Revenue, Risk, and Cards to support analytics and reporting.
Senior Data Engineer
Dubai Municipality
- Designed and implemented scalable PySpark data ingestion pipelines to migrate and process data from SQL Server, Oracle, NoSQL, and raw files into Hadoop and data lake environments, ensuring reliable, structured data flow.
- Built end-to-end transformation workflows in Databricks using PySpark and Python, delivering curated datasets for reporting, analytics, and data science teams, enabling faster insights and model development.
- Automated data ingestion and transformation processes to significantly reduce manual intervention and improve reliability, consistency, and operational efficiency of data pipelines.
- Optimized large-scale data processing across multiple services by applying performance tuning techniques in PySpark, improving execution speed and ensuring efficient handling of high-volume enterprise data.
Data Engineer
Cognizant
- Designed and migrated enterprise data workloads from on-premises Cloudera to Azure Databricks, analyzing architectures and dependencies and optimizing solutions to improve performance and reduce resource utilization.
- Built and modernized scalable data pipelines using PySpark, Hive, and the Cloudera ecosystem, transforming legacy PL/SQL and Informatica-based logic into distributed processing frameworks for analytics and data science use cases.
- Led data transformation and migration by rewriting complex PL/SQL logic into PySpark and Hive, enabling efficient data preparation for reporting, insights, and machine learning.
- Improved system performance and reliability through optimization of legacy ETL processes, complex SQL/PLSQL tuning, and the redesign of data flows across multiple integrated enterprise systems.
Database Developer
Infosys
- Developed and implemented complex business logic using SQL and PL/SQL to support critical enterprise data processing and reporting requirements.
- Enhanced and optimized existing database procedures and functions based on evolving business needs, improving efficiency and maintainability of core data operations.
- Contributed actively to integration testing and defect resolution, ensuring data accuracy, system stability, and smooth delivery of releases.
- Collaborated with stakeholders and end-users to troubleshoot and resolve production issues, ensuring timely support and minimal business impact in live environments.
Experience
Realtime Cheating Detection in a Racing App
I developed real-time data pipelines to detect and prevent live cheaters in the MyWhoosh racing environment.
Education
Master's Degree in Computer Science
Maharishi Arvind Institute of Science and Management - Jaipur, India
Skills
Libraries/APIs
PySpark, Spark Streaming
Tools
AWS Glue, Apache Airflow, GitLab, ksqlDB, Cloudera, Apache Sqoop, Hue, Oozie, Subversion (SVN), Kibana, Microsoft Power BI, dbt Cloud, Apache Iceberg
Languages
SQL, Python, Snowflake
Frameworks
Data Lakehouse, Apache Spark, Spark Structured Streaming, Hadoop
Paradigms
ETL, Automation, Azure DevOps, Agile
Platforms
Databricks, Azure, Apache Kafka, Azure Synapse, Docker, Oracle, Azure Event Hubs, Cloudera Data Platform, Azure Data Lake Storage, Hortonworks Data Platform (HDP)
Storage
Data Pipelines, MongoDB, NoSQL, Cassandra, PostgreSQL, Apache Hive, Oracle SQL, Oracle PL/SQL, Elasticsearch, Databases, Data Lakes
Other
Azure Data Factory (ADF), Azure Databricks, Big Data, Data Warehousing, Data Engineering, Delta Lake, Metadata, Data Architecture, ELT, Orchestration, Scalability, SQL Server, Designing for Data, Unity Catalog, Data Modeling, RESTFul APIs, CI/CD Pipelines, Data Build Tool (dbt), AI-assisted Development, Enterprise, Data Transformation, Geospatial Data, Data Governance, Computer Science, APIs, aws glue, Informatica, Unix Shell Scripting, KSQL, Medallion Architecture, Big Data Architecture, Data Warehouse Design, Design, Architecture, Azure Data Lake, hive metastore, Data Taxonomy
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring