
Utkarsh Gupta
Verified Expert in Engineering
ML Engineer and Developer
Bengaluru, India
Toptal member since August 3, 2026
Utkarsh is a senior ML/AI engineer with 11 years of experience turning complex AI problems into reliable, measurable products. He builds production systems across GenAI, RAG, NLP, Vision, MCP, and LLM-based agents, and works the end-to-end ML lifecycle, including data, models, deployment, monitoring, and retraining. His edge is in what happens after the demo: rigorous evaluation, governance, and architecture that holds up at scale.
Portfolio
Experience
- Python - 10 years
- Artificial Intelligence (AI) - 10 years
- Data Science - 10 years
- Machine Learning - 10 years
- Deep Learning - 8 years
- Amazon Managed Workflows for Apache Airflow (MWAA) - 3 years
- AWS Glue - 3 years
- Apache Airflow - 3 years
Preferred Environment
Python, Artificial Intelligence (AI), Large Language Models (LLMs), Amazon Web Services (AWS), Classification, Generative Artificial Intelligence (GenAI), Model Context Protocol (MCP), LangChain, LangGraph, Spring Boot
The most amazing...
...project I've worked on was end-to-end agentic fraud investigation platform that pairs LLM reasoning with data retrieval and governance to investigate fraud.
Work Experience
Senior ML Engineer
Twilio
- Led the retraining pipeline for Twilio's account takeover fraud model, spanning data science, feature engineering, and MLOps across Amazon SageMaker, Airflow, Glue, and ClickHouse.
- Built an end-to-end CatBoost pipeline on Amazon SageMaker for voice fraud detection, including feature engineering, training, and batch and real-time inference via a containerized inference service.
- Designed and shipped an automated model-monitoring system using PSI, KS-test, chi-squared tests with Datadog alerting, Evidently reports, and a Streamlit dashboard backed by a 38-test suite.
- Maintained production MWAA/Airflow DAGs for fraud metrics, ATO detection, and Kafka event publishing with OpenTelemetry observability into Twilio's Springer/Grafana platform.
- Built a Terraform-managed AWS Glue job powering an automated fraud-alerting pipeline.
- Completed hands-on ML and data science upskilling projects covering Python, applied ML, and ML infrastructure as code.
Lead Algorithm Developer
Applied Materials India Pvt Ltd.
- Developed a defect classification pipeline using AutoML and scikit-learn, improving defect filtering by 35%.
- Deployed containerized ML pipelines on-site using Docker, enabling plug-and-play inference and training.
- Implemented ONNX-based graph models for production usage with high throughput.
- Designed an image processing module for defect attribute extraction, achieving 95% faster processing per tile.
- Delivered segmentation modules for patterned images, enhancing region-of-interest identification.
Senior Research Engineer
Blaize Inc.
- Directed a team of three to evaluate and fine-tune transfer learning backbones for semantic segmentation.
- Built and deployed joint BERT-based NLP pipelines for intent and entity extraction as a stateless REST API.
- Developed a dialogue state tracking (DST) system using PyTorch as a stateful REST API with session history.
- Served as an architect for the knowledge management system supporting smart search and LSI-based retrieval.
- Designed and deployed microservices using Flask, Redis, and MongoDB, integrating with the Azure cloud for scalable access.
Research Engineer
The Hi-Tech Robotic Systemz Limited
- Delivered real-time face, eye, and head gaze detection modules with over 95% accuracy using CNN plus Kalman Filters.
- Optimized C++/DLib code for embedded ARM hardware, improving frame rate from 16 frames per second to 22 frames per second.
- Developed an eyewear detection CNN model, achieving 96% accuracy on the evaluation set.
- Managed a team of seven to architect and deploy a data aggregation pipeline for the fleet management platform using Android, AWS, and web.
- Performed roles of requirement gathering and task assignment to the team.
- Designed the flowchart, test cases, probable failure cause, and traffic handling reports.
Guest Researcher
(DFKI) German Research Centre for Artificial Intelligence
- Researched sparse localized deformation components for improved 3D reconstruction.
- Built a high-speed camera setup and a smart eye tracker with calibrated multi-camera configurations.
- Ported and optimized a Python 3D reconstruction model to C++ using Boost, Eigen, and OpenMP, significantly accelerating processing performance while achieving minimal average error.
Software Engineer
Motherson Group
- Integrated TFS into an ASP.NET web portal using WCF Web Services to centralize version control and streamline development workflows.
- Migrated legacy web applications to modern frameworks and delivered ongoing maintenance, ensuring seamless business continuity and client satisfaction.
- Engineered Linux scripts and configured server environments to automate routine file management, operational monitoring, and maintenance.
- Designed and deployed custom EDI data maps to extract, transform, and integrate critical business data across enterprise systems.
Experience
Agentic Fraud Investigation & Governance Platform
Agentic fraud investigation platform combining LLM reasoning, multi-source retrieval, dynamic skill dispatch, and automated governance.
• Two-phase workflow in LangGraph—stateful nodes and conditional edges keep escalation auditable and replayable; low-risk events gated deterministically, ambiguous cases escalated to the LLM branch, controlling cost and latency.
• Registry-driven skill architecture enabling dynamic selection and parallel fan-out across Presto, ClickHouse, and S3, with reducer-based state merge for concurrent outputs.
• Grounded Amazon Bedrock/Kimi K2 reasoning in real-time evidence, enforcing structured JSON contracts, score calibration, credential isolation, and audit-ready traces.
• Spring Boot microservice fronting the orchestration layer - owns REST contract, auth, credential brokering, and auditing, decoupling the agent layer from the enterprise layer.
• Evaluation harness: 90% investigation accuracy vs. 78–83% for isolated skills; orchestration resolved 71.4% of conflicting skill assessments.
• Token, latency, execution, and retrieval-level observability for cost governance and production monitoring.
Real-time Voice Fraud Detection & Automated MLOps
Led the ML/MLOps lifecycle for an account-takeover fraud detection platform - feature engineering, training, real-time inference, monitoring, and retraining.
• Built an end-to-end CatBoost pipeline on SageMaker covering feature engineering, training, batch inference, and containerized real-time inference.
Led the retraining pipeline across MWAA/Airflow, AWS Glue, SageMaker, and ClickHouse.
• Designed automated model/data monitoring using PSI, Kolmogorov–Smirnov, and χ² tests, with Datadog alerting, Evidently reports, Streamlit dashboards, and a 38-test validation suite.
• Maintained production MWAA workflows for fraud metrics, ATO detection, and Kafka event publishing, with OpenTelemetry observability.
• Used AWS MCP servers with Claude to accelerate deployment and ops — inspecting SageMaker endpoints, Glue jobs, and MWAA DAG state through a schema-defined tool interface instead of ad-hoc scripting.
• Exposed ClickHouse fraud metrics via MCP for conversational drift and performance queries during incident triage, with scoped read-only credentials keeping access auditable.
Semiconductor Defect Intelligence & AutoML Platform
• Developed production ML and computer-vision solutions for nanometer-scale semiconductor defect inspection, automating defect filtering, classification, and attribute extraction across high-volume wafer inspection data.
• Built an AutoML/scikit-learn defect classification pipeline, improving defect filtering performance by 35%.
• Engineered image-processing pipelines for automated defect attribute extraction, achieving around 95% faster processing per tile.
• Developed segmentation modules for patterned semiconductor images to improve region-of-interest identification and downstream defect analysis.
• Deployed containerized ML pipelines using Docker, enabling plug-and-play training and inference in customer/on-site environments.
• Implemented ONNX-based production inference graphs optimized for high-throughput execution.
Production Conversational AI & RAG Platform
Designed and developed production-oriented RAG and conversational AI systems combining knowledge retrieval, LLM generation, prompt engineering, and multi-step conversational workflows.
- Built the end-to-end RAG lifecycle covering data ingestion, extraction, chunking, normalization, embeddings, retrieval, context construction, and LLM generation.
- Designed conversational state-management strategies to preserve relevant context across multi-turn interactions while minimizing context pollution and hallucination.
- Implemented retrieval and prompt optimization strategies to improve grounding, response quality, and knowledge utilization.
- Integrated knowledge retrieval with conversational workflows to select the appropriate knowledge sources and tools based on interaction state.
- Applied production-oriented evaluation and monitoring approaches to assess retrieval quality, hallucination, and response reliability.
OneData: Large-scale AI/Data Platform
• Designed and optimized a cloud-based data platform processing approximately 90 billion rows per month to provide scalable, reusable data products for analytics and ML workloads.
• Built large-scale ETL pipelines for ingestion, normalization, metric computation, and downstream ML/analytics consumption.
• Optimized AWS Glue/Spark processing to reduce DPU costs by ~40–50%.
Improved data-read performance by approximately 5× and reduced query execution time by ~70% through storage and data-model optimization.
• Worked across AWS Glue, S3, Snowflake, Presto/Athena, Redshift, and ClickHouse.
• Designed reusable data pipelines supporting a single-source-of-truth architecture for business and technical metrics.
Active Learning for High-volume Defect Classification
• Designed an active-learning workflow to reduce the manual labeling burden associated with hundreds of thousands of semiconductor inspection defects.
• Developed an intelligent sample-selection strategy to prioritize the most informative defects for expert labeling instead of relying on random sampling.
• Reduced the effective labeling workload from approximately 500,000 candidate defects to around 500 targeted samples.
• Reduced mislabeled samples by approximately 40% while improving system accuracy by around 25%.
• Created an iterative feedback loop between model predictions and expert annotations to continuously improve defect classification quality.
Education
Master's Degree in Artificial Intelligence
Technical University of Kaiserslautern - Kaiserslautern, Germany
Bachelor's Degree in Computer Science Engineering
Krishna Engineering College - Uttar Pradesh, India
Certifications
Machine Learning
Coursera
Skills
Libraries/APIs
Claude API, XGBoost, CatBoost, Keras, Scikit-learn, PyTorch, Dlib
Tools
Claude, Apache Airflow, AWS Glue, AI Prompts, Microsoft Copilot, Claude Code, Grafana, Terraform, Open Neural Network Exchange (ONNX), MATLAB, Slack, Mermaid
Languages
Python, SQL, Java, C++, Lua
Platforms
Amazon Web Services (AWS), Docker, AWS IoT, Kubernetes, Azure, Android, Web, PagerDuty
Storage
Datadog, MySQL, Data Pipelines, ClickHouse, Redis, MongoDB
Frameworks
Streamlit, Flask, C4 Model, LangGraph, Spring Boot
Paradigms
Automation, Model Context Protocol (MCP)
Other
Machine Learning, Deep Learning, Artificial Intelligence (AI), Data Science, Large Language Models (LLMs), Retrieval-augmented Generation (RAG), Prompt Engineering, LangChain, Software Engineering, Back-end, Back-end Development, Software Architecture, Personally Identifiable Information (PII), Statistics, Causal Inference, Classification, Communication, Feature Engineering, Model Development, Model Evaluation, Monitoring, Stakeholder Management, Model Deployment, Linear Regression, Predictive Modeling, Generative Artificial Intelligence (GenAI), APIs, Amazon Managed Workflows for Apache Airflow (MWAA), Kafka, Machine Learning Operations (MLOps), CI/CD Pipelines, Computer Vision, Analytics, Data Analysis, Agentic AI, Data Engineering, AI Agents, RAG Systems, Agentic RAG Systems, Conversational AI, Knowledge Graphs, Memory Management, Conversational Design, Natural Language Processing (NLP), Document Processing, Data Extraction, Time Series Forecasting, Demand Forecasting, Combinatorial Optimization, OpenTelemetry, BERT, Convolutional Neural Networks (CNNs), ARM, Support Vector Machines (SVM), Image Processing, 2D Image Processing, Distributed Systems, Computer Science, Defect Management, Web Services, Electronic Data Interchange (EDI), Engineering Software, Amazon SageMaker Pipelines, Data Analytics, Governance, Risk, Compliance, System Design, Agent Skills, Random Forests, ONNX Runtime
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring