
Vishnu Vardhan Reddy
Verified Expert in Engineering
Data Engineer and Developer
Chicago, IL, United States
Toptal member since April 22, 2026
Vishnu is a senior data engineer with over eight years of experience building scalable cloud data solutions on Azure and AWS. He specializes in Databricks, Snowflake, Synapse Analytics, Azure Data Factory (ADF), AWS Glue, Redshift, PySpark, Python, and SQL. With expertise in ETL/ELT pipelines, metadata-driven frameworks, and data modeling, Vishnu focuses on delivering high-quality, analytics-ready data and optimizing large-scale processing.
Portfolio
Experience
- Data Engineering - 8 years
- SQL - 8 years
- Python 3 - 8 years
- PySpark - 6 years
- Databricks - 6 years
- Azure Data Factory (ADF) - 5 years
- Snowflake - 5 years
- Data Warehousing - 4 years
Preferred Environment
PySpark, Databricks, Snowflake, Data Warehousing, Microsoft Power BI, SQL, Python 3, Azure Data Factory (ADF), Data Engineering, Tableau
The most amazing...
...solutions I've built were scalable ETL pipelines on Databricks, ADF, and Snowflake to streamline multi-vendor ingestion and deliver analytics-ready data.
Work Experience
Senior Data Engineer
Northern Trust
- Designed and built scalable, metadata-driven data pipelines on Databricks and Snowflake to process high-volume banking and customer datasets.
- Developed optimized PySpark transformations for large-scale data processing, including joins and aggregations.
- Orchestrated end-to-end workflows using Apache Airflow and Azure Data Factory (ADF), improving pipeline reliability and reducing manual effort.
- Implemented Delta Lake (DLT) and Autoloader for incremental data ingestion and schema evolution, enabling efficient batch and streaming pipelines.
- Designed and implemented Snowflake dimensional data models to support customer analytics and reporting, improving query performance by approximately 35%.
- Integrated REST APIs and multiple enterprise data sources to enrich customer data and build a unified customer view.
- Applied Spark optimization techniques such as partitioning, caching, and broadcast joins, reducing processing time by approximately 30% and lowering compute costs by approximately 25%.
- Automated infrastructure provisioning and deployments using Terraform and CI/CD pipelines in Azure DevOps, ensuring consistent and reliable multi-environment releases.
Senior Data Engineer
Elevance Health
- Built ETL/ELT workflows using Databricks and PySpark to process claims, provider, and member data for analytics and reporting use cases.
- Modeled data in Snowflake using dimensional modeling and star schema design to support business dashboards and downstream consumption.
- Developed metadata-driven frameworks to enable faster onboarding of new data sources and standardize pipeline development.
- Optimized Spark jobs using partitioning and tuning techniques, improving processing performance by approximately 30% and reducing compute costs.
- Delivered curated datasets and dashboards that improved reporting efficiency and reduced manual effort for business teams.
- Used Amazon Athena for ad-hoc querying and validation of large datasets directly on S3, enabling faster data exploration.
- Implemented data enrichment pipelines by joining raw datasets with reference and master data to improve data usability for analytics.
- Built modular, reusable data transformation models using dbt, enabling scalable and maintainable ELT pipelines in Snowflake.
- Implemented dbt tests, documentation, and CI/CD workflows to ensure data quality, lineage visibility, and reliable deployments.
Cloud Data Engineer
Lincoln Financial Group
- Architected real-time ingestion pipelines using Apache Kafka and Spark Structured Streaming in Scala on Azure Databricks to process high-volume transactional and customer event data.
- Developed scalable Spark streaming applications in Scala with windowing, watermarking, and stateful aggregations for near real-time analytics.
- Integrated Kafka streams with Delta Lake on Databricks for continuous ingestion, ensuring low-latency and ACID-compliant data storage.
- Orchestrated hybrid pipelines using Azure Data Factory for batch ingestion and Databricks workflows for streaming, enabling unified data processing.
- Implemented CDC pipelines using Kafka Connect with Debezium to capture changes from relational databases and stream them into the lakehouse.
- Managed schema evolution and data consistency using Avro and Schema Registry across streaming pipelines.
- Built fault-tolerant streaming jobs with checkpointing and exactly-once processing semantics to ensure reliability.
- Optimized streaming performance using partitioning, caching, and tuning techniques, reducing latency and improving throughput.
Big Data Developer
Arcient technologies
- Worked on a big data ecosystem in Hadoop using MapReduce, Spark, Hive, Pig, Sqoop, HBase, Oozie, and Impala for large-scale data processing.
- Installed and configured Hadoop, including HDFS and MapReduce, and developed MapReduce jobs for data cleansing and preprocessing.
- Built optimized Hive and SQL consumption views on top of metrics to improve the performance of complex queries.
- Performed data migration from Teradata to Snowflake by writing SQL scripts for data validation and mismatch analysis.
- Generated custom SQL queries to analyze dependencies across daily, weekly, and monthly batch jobs.
- Contributed to end-to-end testing, including functional, integration, regression, smoke, and performance testing.
- Tested Hadoop-based pipelines developed using Python, Pig, and Hive for data accuracy and reliability.
- Managed defect tracking and resolution using Jira, coordinating with development teams.
- Analyzed campaign performance, including product listing ads (PLAs), using statistical techniques and BI tools to drive ROI improvements.
- Performed data analysis using regression, data cleaning, Excel functions including VLOOKUP and histograms, and TOAD client.
Experience
Enterprise Risk Data Analytics Platform (Northern Trust)
I developed optimized PySpark transformations and implemented Delta Lake (DLT) with Autoloader for incremental ingestion and schema evolution across batch and streaming pipelines. I also orchestrated workflows using Airflow and ADF to improve reliability.
I designed Snowflake data models for risk reporting and analytics, improving query performance by approximately 35%. I applied Spark optimizations to reduce processing time by approximately 30% and enabled Power BI dashboards for key risk insights.
Skills
Libraries/APIs
PySpark
Tools
Microsoft Power BI, Tableau, Git, GitHub, AWS Glue, Amazon EKS, Azure MFA, Kafka Streams, Kafka Connect, Apache Sqoop
Languages
Snowflake, Python 3, SQL, Python
Platforms
Databricks, Azure, Azure Data Lake Storage, Azure Synapse Analytics, Amazon Web Services (AWS), Amazon EC2, AWS Lambda, Apache Kafka
Frameworks
Delta Live Tables (DLT), Spark, Hadoop
Paradigms
Role-based Access Control (RBAC), ETL, MapReduce
Storage
PostgreSQL, Redshift, Data Pipelines, Azure SQL, Apache Hive
Other
Azure Data Factory (ADF), Data Engineering, Azure Databricks, Data Warehousing, Delta Tables, Data Modeling, CI/CD Pipelines, Amazon Neptune, Amazon RDS, Azure IoT, Data Build Tool (dbt)
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring