
Saipriya Puram
Verified Expert in Engineering
Data Engineer and Developer
San Jose, CA, United States
Toptal member since July 24, 2026
Saipriya has spent over 10 years building scalable data architectures and deploying machine learning models across technology and healthcare. Her toolkit centers on Databricks, AWS, and advanced analytics. While at Tesla, Saipriya achieved a 74% increase in processing speed by innovating data processing techniques.
Portfolio
Experience
- Python - 13 years
- ETL - 13 years
- GitHub - 12 years
- Data Engineering - 12 years
- CI/CD Pipelines - 12 years
- AWS IoT - 10 years
- Artificial Intelligence (AI) - 7 years
- Databricks - 7 years
Preferred Environment
AWS IoT, Azure, Databricks, Apache Airflow, Terraform, Data Build Tool (dbt), Jira, Confluence, GitHub
The most amazing...
...fraud detection model I've deployed reduced fraudulent activities by 30% for Socure using Databricks MLflow and advanced analytics.
Work Experience
Senior Lead Data Engineer/Architect
Lovelytics
- Delivered end-to-end migration of SAS/SQL scoring models such as T65, T66, Consumer Direct, and Supportive into Databricks with PySpark and Unity Catalog.
- Built scalable ETL pipelines replacing Teradata, Oracle, and Redshift, leveraging Delta Lake bronze, silver, and gold layers for governance and consistency.
- Optimized Spark jobs using repartitioning strategies, reducing scoring runtime from approximately three hours to just seven minutes on 50+ million records.
- Enriched datasets with crosswalk demographics and prospect metadata, enabling targeted Medicare campaigns.
- Documented workflows in Confluence and Jira and collaborated with marketing and governance teams for Salesforce Marketing Cloud integration.
- Worked on the migration of critical data sources from Snowflake to Databricks, enhancing fraud detection capabilities.
- Implemented Databricks Lakehouse architecture, including bronze, silver, and gold layers, to streamline ETL workflows and enable advanced analytics.
- Deployed machine learning models for fraud detection using Databricks MLflow and automated CI/CD pipelines with Git and Terraform.
- Directed the migration of Hadoop workloads to Databricks, ensuring data model accuracy and efficient orchestration using Airflow.
- Optimized performance bottlenecks during data migration and established automated validation frameworks.
Senior Data Engineer/Architect
FIT:MATCH.ai
- Led the adoption of Intel's 3D edge LADAR sensors for accurate body measurements in retail, linking this data with AI applications.
- Established AWS-based data pipelines to manage the flow from 3D sensors through to analysis, ensuring streamlined data handling.
- Created scalable ETL processes in Databricks for efficient transformation of detailed 3D measurement data.
Senior Data Engineering Manager of Data Engineering
Alchemee
- Designed and implemented end-to-end data pipelines using AWS services such as Lambda, S3, and Glue to extract, transform, and load data from 3rd-party APIs, including TikTok, Facebook, and Amazon.
- Built and managed AWS Identity and Access Management (IAM) roles and permissions to ensure secure access to AWS services and 3rd-party APIs.
- Utilized dbt in conjunction with Databricks to refine analytical processing of 3D measurement data, bypassing traditional storage in Redshift.
- Mentored junior data engineers on best practices for managing and analyzing 3D sensor data, emphasizing the innovative use of Databricks and dbt.
Senior Data Engineer
Tesla
- Engineered and managed over 70 reporting dashboards, delivering critical insights to stakeholders and streamlining data analysis and pipelines with SQL and Python.
- Integrated data across eight systems into a centralized analytics framework, handling 100 million records from fourteen sources, significantly improving data accessibility in a high-volume manufacturing setting.
- Innovated data processing by developing JSON format transformations and a cloud-first data ingestion process with Python, SQL, and Spark, achieving a 74% increase in processing speed.
Data Scientist
Georgia Pacific
- Applied predictive analytics to reduce robot failure rates from 80% to 28%, enhancing operational reliability.
- Drove a strategic migration from Oracle to Redshift, leveraging Amazon Athena and S3, resulting in substantial cost savings and improved system performance.
- Developed a Python library to streamline data processing from external vendors, achieving a 12% reduction in data errors.
Experience
Enterprise Healthcare Data Platform Modernization
Enterprise Healthcare Data Platform Modernization
Education
Master's Degree in Electrical Engineering
Northwestern Polytechnic University - Fremont, CA, USA
Bachelor's Degree in Electronic and Communication Engineering
Jawaharlal Nehru Technological University (JNTU) - Hyderabad, India
Certifications
Databricks Certified Machine Learning Associate
Databricks
Databricks Certified Data Engineer Associate
Databricks
Azure Data Engineer
Microsoft
AWS Certified Solutions Architect – Professional
Amazon Web Services
Databricks Certified Advanced Professional Data Engineer
Databricks
IBM Professional Data Scientist
IBM
Lean Six Sigma Black Belt
GE
Skills
Libraries/APIs
PySpark
Tools
Git, Terraform, AWS Glue, Amazon Athena, GitHub, Confluence, Jira, Apache Airflow
Languages
SQL, Python
Frameworks
Spark, Apache Spark
Paradigms
ETL
Platforms
Databricks, AWS IoT, AWS Lambda, Azure, Amazon Web Services (AWS)
Storage
Redshift, Amazon S3 (AWS S3), Amazon DynamoDB
Other
Unity Catalog, Delta Lake, MLflow, Data Build Tool (dbt), Data Engineering, CI/CD Pipelines, AWS Certified Solution Architect, Machine Learning, Computer Science, Artificial Intelligence (AI), Data Science, Six Sigma Black Belt Certified, Cryptography, Advance data engineering
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring