
Satish Satish
Verified Expert in Engineering
Data Engineer and Developer
Bengaluru, Karnataka, India
Toptal member since May 8, 2026
Satish is a senior data, analytics, and AI engineer with 6+ years of experience designing scalable Snowflake data warehouses, enterprise RAG systems, and LLM-driven applications across AWS, GCP, and Databricks. He specializes in Azure OpenAI integrations, dbt modeling, and Airflow orchestration, delivering high-performance Medallion architectures and GenAI knowledge base agents. He handled over 45GB of daily data and thousands of enterprise documents with maximum reliability and optimized cost.
Portfolio
Experience
- PySpark - 6 years
- SQL - 5 years
- AWS Glue - 5 years
- Python - 5 years
- Redshift - 5 years
- ETL - 5 years
- Azure Databricks - 3 years
- Snowflake - 3 years
Preferred Environment
Azure Databricks, AWS Glue, Snowflake, Google BigQuery, Python, SQL, Data Build Tool (dbt), AWS Lambda, Redshift, RAG Pipelines
The most amazing...
...thing: I reduced a 3‑hour pipeline to 15 minutes by implementing a robust data quality framework, optimizing ETL pipelines, and improving decision-making speed.
Work Experience
AI Engineer I
Publicis Sapient
- Architected a production RAG system for GE Vernova, allowing employees to query thousands of finance documents using Azure OpenAI GPT-4.
- Engineered a robust multi-format document parser and OCR processing pipeline, implementing advanced table extraction, bug-resolution mechanisms, and comprehensive test suites to ensure 100% data integrity.
- Architected a production GenAI RAG agent utilizing hybrid search, Cohere reranking, and automated LLM-evaluation gates to optimize enterprise-grade retrieval relevance and response accuracy.
Senior Data Engineer
Publicis Sapient
- Architected and automated Databricks ETL pipelines across bronze, silver, and gold layers using PySpark to unify listener identities into a single "Golden Record" 360 profile.
- Engineered a high-throughput migration of 17+ million records from Amazon Redshift to Salesforce via AWS Glue and Bulk APIs, achieving a less than 1% error rate for production delivery.
- Developed an automated batch data pipeline on GCP utilizing Cloud Functions and BigQuery, which optimized query performance and significantly reduced data processing costs.
- Implemented robust data quality frameworks and Terraform-driven infrastructure deployments, ensuring 99.9% pipeline reliability and stable, scalable production environments.
Data Engineer
Happiest Minds
- Built and operationalized a DataOps framework across multiple environments, ensuring high data quality and reliable ML deployments for transport-sector applications.
- Automated end-to-end ETL pipelines using AWS Glue and PySpark, delivering curated datasets for parts back order analysis and improving inventory planning accuracy.
- Administered Amazon Redshift environments and developed automated billing dashboards in Power BI using CloudWatch metrics to optimize cloud spend and usage tracking.
- Developed back-end APIs via Amazon API Gateway and Lambda to expose curated Redshift data to real-time dashboards, accelerating data-driven decision-making for stakeholders.
Data Engineer
Webtouch Software Development
- Architected event-driven ETL pipelines using S3, Lambda, and AWS Glue to ingest real-time eCommerce feeds into Snowflake, improving data availability for the health sector.
- Engineered an automated sentiment analysis pipeline for banking data, transforming raw customer feedback into ML-ready datasets via AWS Glue and S3.
- Designed and implemented a scalable Snowflake data warehouse, integrating Salesforce into a unified analytics platform.
- Built modular dbt transformation layers (staging and marts) with incremental models, ensuring optimized performance, lineage, and maintainability.
- Developed API-driven Salesforce data extraction pipelines with schema standardization and transformation logic for cross-system data consolidation.
- Implemented automated data quality and reconciliation frameworks with validation, referential integrity checks, and anomaly detection safeguards.
Experience
GenAI Finance Knowledge Base Agent — GE Vernova
My Core Contributions:
• Multi-Format Pipeline: Built a zero-downtime serverless Lambda pipeline supporting 12+ file formats (PDF, DOC/X, PPTX, XLSX/S/B, images) via Textract OCR.
• Lambda Optimization: Engineered a zero-dependency OLE2 parser for Lambda, bypassing a 300MB LibreOffice binary constraint to recover legacy data.
• Table Extraction: Parsed binary .doc markers using olefile to generate structured markdown tables inside an isolated runtime.
• Hybrid Retrieval: Combined 3072-dim pgvector embeddings, tsvector text search, RRF ranking, and Cohere reranker to hit a 3.88/5 relevance score.
• CI/CD Quality Gates: Automated evaluation runs on deployment using RAGAS to enforce KPI thresholds (Correctness, Groundedness) and eliminate regression.
• SME-Validated Impact: Secured a 5.0/5 domain accuracy rating from Subject Matter Experts for answer correctness and completeness.
Golden Record Audience Platform
My primary responsibility was architecting a Medallion data architecture (bronze, silver, and gold layers) using Databricks and PySpark on AWS. I engineered high-performance ETL pipelines that processed millions of records daily, ensuring that disparate data points from streaming services and social platforms were resolved into a single, high-fidelity user identity. To ensure production-grade reliability, I automated the infrastructure deployment using Terraform and integrated AWS Lambda for event-driven processing. I also implemented a robust data quality framework that performed automated validation at the silver layer, reducing data errors by over 20%.
This project directly enabled the client to perform advanced predictive modeling and hyper-personalized marketing, resulting in a significant uplift in listener engagement and data-driven revenue growth.
Automotive Supply Chain | Parts Back-order and Inventory Analytics
My role involved designing and implementing a serverless ETL architecture using Amazon S3, AWS Glue (PySpark), and Amazon Redshift. I also integrated Amazon Athena to enable rapid ad hoc validation of large datasets. To bridge the gap between data and action, I built a secure API layer using AWS API Gateway and Lambda, which exposed curated Redshift data to real-time executive dashboards and dealership applications.
This solution significantly accelerated the insight-to-decision cycle, ensuring critical parts were available when needed and reducing operational bottlenecks across the dealership network.
Enterprise Snowflake Data Warehouse with QA Automation and dbt Modeling
I developed modular dbt models with incremental processing to optimize performance and ensure efficient data transformations. I implemented a robust data quality framework combining dbt tests, SQL validations, and Python-based automation to enforce completeness, accuracy, and referential integrity across pipelines.
My work included orchestrating end-to-end workflows using Airflow, enabling automated ingestion, transformation, testing, and monitoring. I integrated CI/CD pipelines for version control and automated deployments, ensuring reliable and maintainable production workflows. I also built monitoring and alerting mechanisms to track pipeline health, detect anomalies, and prevent data issues from propagating downstream.
Finally, I delivered trusted, analytics-ready datasets that improved reporting efficiency, reduced data quality issues, and enabled faster, data-driven decision-making across business teams.
Skills
Libraries/APIs
PySpark, SQLAlchemy, REST APIs, Bulk API, PyPDF
Tools
AWS Glue, Apache Airflow, Amazon Athena, AWS Step Functions, Amazon CloudWatch, GitHub, BigQuery, Cloudera, Pytest, dbt Cloud, Amazon Elastic Container Service (ECS), Terraform, Salesforce Sales Cloud, Microsoft Power BI, Amazon Textract, Azure OpenAI Service
Languages
Python, SQL, Snowflake
Frameworks
Apache Spark, Data Lakehouse, Spark, Flask, Hadoop, Django REST Framework
Paradigms
ETL, Role-based Access Control (RBAC)
Platforms
Databricks, Amazon Web Services (AWS), AWS Lambda, Amazon EC2, Azure, Oracle, Kubernetes, Docker
Storage
Amazon S3 (AWS S3), Data Pipelines, Apache Hive, MySQL, AWS Data Pipeline Service, Redshift, Microsoft Entra ID, Databases, Data Lakes, Amazon Aurora, Database Architecture, PostgreSQL, Database Performance, Data Validation
Other
Data Engineering, Data Cleansing, ETL Pipelines, Data Warehousing, Performance Optimization, Data Governance, AWS Secrets Manager, Performance Tuning, Query Optimization, Big Data, Large Data Sets, Large-scale Data Processing, Azure Databricks, Google BigQuery, Amazon API Gateway, Data Warehouse Design, Amazon RDS, EMR, Delta Lake, CI/CD Pipelines, AI-assisted Development, Medallion Architecture, Azure Data Lake, Data Management, Data Quality, DataOps, Financial Data, Monitoring, Observability, IT Service Management (ITSM), Data Modeling, Star Schema, Geospatial Data, Quality Assurance (QA), Performance, Concurrency, Database Optimization, Azure Data Factory (ADF), FastAPI, Artificial Intelligence (AI), APIs, Architecture, Data Architecture, Product Management.Agile Product Management, Project Management.Project Delivery, Product Management.Relationship Management.Stakeholder Management, Product Management.Technical Product Management, Looker Studio, Machine Learning, Data Build Tool (dbt), ETL/ELT Pipelines, Data Modeling (Bronze/Silver/Gold), Data Quality & Validation, CI/CD (Git-based Workflows), Data Warehouse Architecture, RAG Pipelines, LangChain, langfus, Amazon Bedrock AgentCore, Vector Databases, Knowledge Graphs, Large Language Models (LLMs), Prompt Engineering, Retrieval-augmented Generation (RAG)
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring