
Rahul Valupadasu
Verified Expert in Engineering
Azure Data Engineer and Developer
Toronto, ON, Canada
Toptal member since December 12, 2024
Rahul is an expert Azure data engineer with over five years of experience delivering big data solutions across the Azure ecosystem. He specializes in Azure Data Factory, Databricks, Delta Lake, Microsoft Fabric, and Spark, including PySpark, Spark SQL, and Scala, for real-time and batch data processing. Rahul designs and optimizes high-performance data systems with robust data governance, ensuring actionable insights and business success.
Portfolio
Experience
- SQL - 6 years
- Spark - 5 years
- PySpark - 5 years
- Azure Data Factory (ADF) - 4 years
- Azure - 4 years
- Azure Databricks - 4 years
- Microsoft Fabric - 2 years
- Microsoft Power BI - 2 years
Preferred Environment
Azure Databricks, Azure Data Factory (ADF), Azure Data Lake Storage, Azure Synapse Analytics, Delta Lake, Spark, PySpark, Data Fabric, Microsoft Fabric, Azure, Model Context Protocol (MCP)
The most amazing...
...thing I've built is a production-grade modular ingestion and MDM framework on Azure, handling CDC, SCD2, streaming, and batch with enterprise-level logging.
Work Experience
Senior Databricks Data Engineer
Darkstar Development, LLC
- Architected and developed an end-to-end insurance data platform on Azure Databricks using PySpark, Delta Lake, and Medallion Architecture (Bronze to SSOT to Silver to Gold), standardizing heterogeneous carrier data into a unified analytical model.
- Designed and implemented a Single Source of Truth (SSOT) canonicalization framework that reconciled 600+ standardized business attributes across multiple insurance carriers, enabling consistent downstream reporting and AI-driven schema mapping.
- Built scalable batch ingestion pipelines capable of processing schema-drifting CSV, Excel, and Parquet files into Delta Lake while preserving raw lineage, ingestion metadata, and full reprocessing capabilities.
- Engineered a configurable data quality framework comprising 29 validation rules across null, format, boundary, referential integrity, cross-column, and trend-based validations using a registry-driven dispatcher architecture in PySpark.
- Developed a hybrid rule engine leveraging explicit function catalogs and metadata registries to execute validation logic dynamically, improving maintainability while supporting client-owned post-production enhancements.
- Developed self-service Databricks notebooks and Genie-powered natural language analytics capabilities, allowing business users to investigate run health, data quality metrics, and carrier-level anomalies without writing SQL.
- Designed dimensional data models and conformed lookup tables for carriers, agencies, lines of business, and reporting periods, supporting scalable Power BI reporting and enterprise analytics.
- Diagnosed and resolved complex production data reconciliation issues by correcting aggregation grain, run-selection logic, and carry-forward calculations, achieving 100% parity between computed and independently validated results.
Software Developer
Scale AI - Gen AI
- Authored evaluation tasks for AI agents and large language models, designing scenario-based prompts that tested tool use, instruction following, and multi-step reasoning across simulated user environments.
- Built grading rubrics and reference outputs used to compare model performance, aligning prompts, expected outcomes, and scoring criteria so evaluations reliably surfaced differences between stronger and weaker models.
- Designed task difficulty to expose common LLM failure modes such as hallucination, time zone and date errors, and mishandling of conflicting sources, ensuring evaluations produced meaningful signals on real agent weaknesses.
Data Engineer
Zoetis - Data and Digital Solutions
- Designed and delivered end-to-end Azure data pipelines using Azure Data Factory, Azure Databricks, and ADLS Gen2 to support analytics dashboards consumed by commercial and sales teams across multiple business units.
- Integrated curated datasets with middleware applications by building Python-based REST APIs using FastAPI and Django, transforming data extracts into structured object models for application consumption.
- Built scalable PySpark and Spark SQL transformation layers processing high-volume datasets, optimizing partitioning and file sizing to improve query performance and reduce execution time in production workloads.
- Implemented a metadata-driven data extract framework that automated daily, weekly, and monthly deliveries, reducing manual extract effort and improving on-time data availability for downstream analytics.
- Developed reusable, parameterized ETL components to handle frequent ad-hoc business requests, enabling rapid turnaround of custom data extracts without duplicating pipeline logic.
- Enforced data quality controls, including schema validation, null checks, record counts, and reconciliation logic, to ensure accuracy and consistency across curated analytical datasets.
- Collaborated with product owners, analytics teams, and application developers to translate business requirements into scalable data models, transformation logic, and API contracts.
- Implemented robust error handling, logging, and alerting mechanisms across pipelines to improve production reliability and reduce mean time to resolution for data issues.
- Leveraged internal AI tooling (Zen AI) to improve developer productivity across Databricks-based data pipelines, accelerating PySpark development, validation logic, and unit test creation while maintaining production-grade standards.
Data Engineer
Scotiabank
- Designed and deployed scalable ETL pipelines using Azure Data Factory and Databricks, improving data ingestion efficiency by 30% for critical business processes.
- Implemented DLT in Databricks, streamlining real-time and batch data processing workflows, reducing latency, and improving data pipeline reliability.
- Migrated data solutions from on-premises to the Azure ecosystem, leveraging Azure Data Factory, Azure Synapse Analytics, and ADLS to improve scalability and reduce storage costs by 20%.
- Optimized Spark-based workflows in Azure Databricks, reducing data processing times by 40% while ensuring high performance for large-scale datasets.
- Implemented Unity Catalog for centralized data governance, enabling streamlined data access management and ensuring regulatory compliance across the enterprise.
- Built real-time streaming pipelines using Azure Event Hubs, Kafka, and Azure Databricks, enabling seamless data integration for near-instant analytics and reporting.
- Enhanced data performance through partitioning and optimization techniques, achieving 30% faster query execution in Delta Lake for reporting and analysis.
Data Engineer
Inkresults-Outsourcing
- Developed PySpark scripts to perform comprehensive data profiling, validation, and quality checks on Azure Data Lake Storage, enhancing data reliability and trust.
- Migrated legacy SSIS packages to Azure Data Factory, modernizing ETL processes and achieving a 30% improvement in scalability and flexibility.
- Coordinated ETL workflows using Apache Airflow and Azure Databricks, ensuring reliable and automated data pipelines with 95% accuracy.
- Built PySpark-based ETL workflows to automate complex transformations, enhancing data processing efficiency by 25% and improving data quality checks.
- Collaborated with stakeholders to develop actionable Power BI visualizations, enabling data-driven decision-making and improving business insights.
- Reduced ETL processing time by 20% by optimizing Azure Databricks pipelines and implementing config-driven solutions for flexible workflows.
- Optimized SQL queries and data models, leading to a 50% improvement in query performance for complex analytical queries.
- Developed Spark programs in PySpark for data transformation and quality checks, ensuring consistency and integrity across multiple datasets.
- Implemented strong data governance practices with role-based access control (RBAC) and data encryption, ensuring compliance with GDPR and industry data privacy standards.
- Implemented data transformation strategies using Python (Pandas) and Azure Data Factory, reducing data cleansing and enrichment time by 30%.
Experience
Enterprise Data Platform Modernization
• Building robust ETL pipelines using Azure Data Factory and Databricks to ingest and process structured and unstructured data from multiple financial systems.
• Implementing Delta Live Tables and PySpark in Databricks for efficient real-time data transformations and incremental loads from bronze to silver layers following the medallion architecture.
• Integrating Unity Catalog to centralize data security, manage access, and enforce compliance across the enterprise.
• Leveraging Azure Data Lake Storage and Azure Synapse Analytics for scalable storage and querying of large datasets, enhancing performance and business insights.
• Achieving a 30% reduction in processing time by optimizing Spark-based workflows and improving data pipeline performance.
This project enhanced data reliability, faster insights, and improved governance for critical financial reporting and decision-making.
Revenue Cycle Management Data Integration Project
• Designing and implementing ETL pipelines in Azure Data Factory to extract data from on-premises systems, cloud databases, and APIs.
• Utilizing PySpark and Databricks for advanced data cleansing, enrichment, and validation, ensuring 99% data accuracy.
• Optimizing SQL queries and applied partitioning techniques to improve the performance of complex analytical workloads by 50%.
• Ensuring HIPAA compliance by implementing RBAC and encrypting sensitive data.
• Integrating Power BI dashboards with Azure SQL Database to deliver actionable insights into patient revenue trends, improving decision-making for healthcare administrators.
This project improved operational efficiency, enabled real-time revenue tracking, and reduced reporting latency by 20%.
Enterprise Azure to Microsoft Fabric Migration & Data Platform Modernization – CMHC
Education
Postgraduate Degree in Computer Science
Lambton College - Toronto, Canada
Certifications
Microsoft Certified: Fabric Data Engineer Associate
Microsoft
Microsoft Certified Azure Data Associate
Microsoft
Skills
Libraries/APIs
PySpark, Fabric, Pandas
Tools
Spark SQL, Microsoft Power BI, Claude, Microsoft Copilot, Azure Monitor, Apache Airflow, Control-M, Tableau, Jira, Terraform
Languages
Python, SQL, Snowflake
Frameworks
Spark, Apache Spark, Delta Live Tables (DLT), Data Lakehouse, Hadoop, Windows PowerShell, Data Fabric
Paradigms
ETL, Role-based Access Control (RBAC), Synthetic Data Generation, Business Intelligence (BI), Model Context Protocol (MCP)
Platforms
Databricks, Microsoft Fabric, Azure Data Lake Storage, Azure, Linux, Azure Functions, Azure Service Fabric, Azure SQL Data Warehouse, Microsoft Copilot Studio, Azure Synapse Analytics, Jupyter Notebook
Storage
Azure SQL Databases, IBM Db2, Databases, Database Performance, Microsoft SQL Server, Data Pipelines, Database Administration (DBA), SQL Server DBA, SQL Performance, SQL Server Integration Services (SSIS), JSON, Relational Databases, Azure Cosmos DB
Other
Azure Databricks, Data Engineering, Data Migration, Database Partitioning, Microsoft Azure, Pipelines, Data Warehouse Implementation, Row-level Security (RLS), Data Quality Analysis, CSV, ELT, Data Analysis, Data Analytics, Medallion Architecture, Azure Data Factory (ADF), Delta Lake, Unity Catalog, Data Cleansing, Data Profiling, Data Quality, Big Data, Distributed Systems, Data Cleaning, Data Conversion, Data Modeling, Fact Tables, Data Warehousing, Data Architecture, Azure Data Lake, Architecture, API Integration, CI/CD Pipelines, Data Visualization, Data Build Tool (dbt), Migration, FTP, OneLake, Star Schema, Artificial Intelligence (AI), AI Modeling, Large Language Models (LLMs), AI Agents
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring