
Shadab Arif Ansari
Verified Expert in Engineering
ML Engineer and Developer
Bengaluru, Karnataka, India
Toptal member since October 9, 2025
Shadab is a lead AI engineer with 8+ years of experience building production AI systems for healthcare and fintech enterprises. He specializes in agentic AI, multi-agent orchestration (LangGraph), and advanced RAG/GraphRAG architectures powered by Claude, OpenAI, and Llama. Shadab has deployed secure, high-throughput ML platforms on AWS and Azure—from real-time lending decisioning at scale to enterprise member-risk prediction. He's an IIT Kharagpur alumnus and a top 3 LLM engineer at Turing.
Portfolio
Experience
- Python - 9 years
- FastAPI - 6 years
- LangChain - 4 years
- Retrieval-augmented Generation (RAG) - 4 years
- Generative Artificial Intelligence (GenAI) - 4 years
- LangGraph - 4 years
- Chatbots - 4 years
- Agentic AI - 3 years
Preferred Environment
Amazon Web Services (AWS), Azure, LangGraph, LangChain, PyTorch, Python 3, Pinecone, FAISS, Agentic RAG Systems, Agentic AI
The most amazing...
...: I shipped a multi-agent customer support RAG platform end-to-end—back-end, AI, and front-end—that served 10,000+ customers and cut support workload by 44%.
Work Experience
Lead AI Engineer
UnitedHealth Group
- Developed a GenAI-powered agent to automate claim intake, policy verification, document review, fraud detection, and claim summarization, improving processing speed and operational efficiency.
- Built a RAG-based AI agent that answers policyholder queries, explains coverage, exclusions, premiums, and claim eligibility using natural language, improving customer support efficiency.
- Created an LLM-based document processing solution to extract, classify, summarize, and query information from clinical notes, discharge summaries, and insurance documents using RAG and vector search.
- Built an ML-based member risk prediction model to identify high-risk members using healthcare and behavioral data, enabling proactive interventions, improved risk stratification, and data-driven care management.
Lead AI Engineer
RVIN
- Led development of a multi-agent RAG system using LangGraph, Azure OpenAI, Azure AI Speech, AKS, and Azure AI Search for Arabic voice/text support, reducing support tickets by 44% and improving customer resolution rates by 28% overall.
- Built an autonomous LangGraph scheduling agent with MCP tool integrations for booking, rescheduling, and cancellations—integrated with Databricks-backed data workflows and deployed on Azure and Kubernetes for scalable, zero-touch workflow automation.
- Developed a GraphRAG platform using Neo4j, LangGraph, and vector search to connect policies, SOPs, and enterprise documents, enabling multi-hop reasoning and contextual AI search across large-scale enterprise knowledge bases.
- Built an LLM pipeline to extract organizations, job roles, locations, and relationships from unstructured documents and webpages using structured prompting, validation, and evaluation workflows for high-accuracy data generation.
- Created an AI-powered email reply agent using Microsoft Copilot that reads your inbox, understands context, and auto-drafts smart responses — boosting productivity with LLMs and Gmail API.
- Developed a multimodal AI chatbot with Llama, Claude APIs, Whisper, TTS, image generation, vector memory, intent detection, and context-aware personalized responses.
- Built a personalized AI companion using LLMs, RAG, Pinecone, Rasa, voice AI, semantic memory, prompt caching, RL-based personalization, and scalable data pipelines.
Senior AI Engineer
Digital Exchange
- Designed and deployed a LangGraph and Claude AI agent for patient intake, query resolution, appointment workflows, and medical form extraction — delivered end-to-end independently for clinical use.
- Built and deployed an AI assistant to analyze lab reports and clinical documents, enabling healthcare teams to retrieve insights and summaries instantly — shipped independently for a live client.
- Created a healthcare support AI agent using LangGraph and Claude to answer patient queries, retrieve treatment information, and assist appointment workflows—designed, built, and delivered the full solution independently.
- Built a personalized recommendation engine using user behavior and transaction data; implemented feature pipelines, affinity scoring, and real-time inference, achieving ~80% AUC with strong precision and recall, improving recommendation relevance.
Senior AI Engineer
An Online Freelance Agency
- Built a LangGraph + Azure OpenAI email agent integrated with Outlook via Microsoft Graph API; context-aware drafting at 3s latency, reducing manual email effort by 40% in production.
- Fine-tuned LLMs on domain-specific datasets using LoRA and QLoRA, improving contextual accuracy by 28% for compliance-critical applications across legal, medical, and finance verticals.
- Architected a LangGraph multi-agent chatbot with STT and context-aware recommendations across 10,000+ products — increased customer engagement by 35% in production.
- Delivered a LLaMA + RAG invoice agent, achieving 90% field extraction accuracy across 10,000+ invoices — reduced manual processing effort by 65% in live workflows.
- Designed a LangGraph + RAG claims agent that ingests policy docs and medical bills, verifies coverage, flags fraud risks, and autonomously drafts claim decisions end-to-end.
- Fine-tuned Qwen2.5-VL-3B with QLoRA (4-bit NF4 + LoRA r=16) on invoice receipt dataset for structured JSON extraction, achieving 90.2% strict / 97.8% normalized macro exact-match with 100% JSON parse rate on held-out test set.
Senior Data Scientist
Paytm
- Built a financial document intelligence platform using hybrid RAG to extract and validate data from statements, PDFs, and spreadsheets with provenance tracking, audit trails, and deterministic calculations.
- Designed a customer engagement scoring machine learning model using app usage, recharge/payment frequency, transaction diversity, and retention behavior to improve campaign targeting and personalized recommendations.
- Built a fraud risk machine learning model with lightgbm, using transaction velocity, device signals, location patterns, and behavioral anomalies to identify suspicious merchants and payment activities in near real-time.
- Built stacked XGBoost and BERT-based models on GCP Vertex AI, leveraging over 35 million records to achieve a 75–76% AUC. Powered real-time risk alerts, loan underwriting, and portfolio intervention at Paytm scale.
- Architected comprehensive ML pipelines on SageMaker and EC2, incorporating CI/CD, MLflow tracking, Docker, and hyperparameter tuning. Established a production-grade training and deployment infrastructure for credit risk models.
- Built merchant clustering and segmentation models using transaction trends, business categories, geography, and growth signals to support targeted marketing, cross-sell opportunities, and business prioritization strategies.
Senior Data Scientist
Protium
- Built a credit card churn prediction model using XGBoost and behavioral transaction features to identify high-risk inactive borrowers and improve retention campaign targeting.
- Built a customer retention model for lending users using transaction, repayment, and engagement signals to predict repeat borrowing propensity and improve renewal conversion rates.
- Developed XGBoost and random forest models on GCP Vertex AI across 2+ million records, achieving a 76% AUC, reducing bad rates by 5%, increasing approvals by 10%, and cutting defaults by 3% in production.
- Built a reconstruction-error autoencoder on normal loan data, eliminating label dependency, reducing manual fraud reviews by 35%, lowering fraud losses by 12%, and accelerating loan approvals.
- Deployed AWS monitoring pipeline with Airflow and MLflow tracking AUC, PSI, and drift across 6 models; automated retraining triggers reduced performance incidents by 30%.
Data Scientist
L&T India
- Built and fine-tuned ResNet and EfficientNet CNN models in PyTorch for construction defect detection, achieving 92% accuracy; automated inspections, reducing manual review time by 40% and improving on-site quality control.
- Trained LSTM and Transformer models on historical project data for cost and delay prediction, improving time estimation accuracy by 18% and enabling proactive risk mitigation in large-scale construction projects.
- Built a Mask R-CNN defect segmentation model in PyTorch for construction site images, achieving 90% IoU; enabled automated crack and structural anomaly detection, significantly improving on-site inspection accuracy and safety compliance.
- Performed advanced EDA using Pandas, NumPy, and PySpark on 500,000+ construction records to analyze material usage, costs, and timelines; improved forecasting accuracy by 15% and reduced budget overruns through data-driven planning.
Python Developer
Freelance.com
- Developed Python back-end services using Flask/Django, creating REST APIs for recommendation engines, sentiment analysis, and anomaly detection, integrating ML models (scikit-learn, NLTK) with MySQL/Redis for scalable data processing.
- Engineered back-end systems for data-driven applications, including URL shorteners, email classification, and reporting tools, leveraging REST APIs, database optimization, and basic machine learning for intelligent decision-making.
- Built microservices in Python for log analysis, task automation, and user management, implementing background jobs with Celery, caching via Redis, and exposing APIs for real-time insights and system monitoring.
Experience
Agentic RAG Outlook Assistant (Apollo Global via Turing)
FastAPI back-end containerized on Docker, deployed to Azure App Service with Application Insights for latency tracing and error alerting. Structured output parsing enforced reply format consistency across varied email intents. Achieved around 3s end-to-end latency, 90% reduction in manual drafting workload, and ranked top three company-wide project at Turing.
Multilingual Agentic Customer Support Platform (RVIN)
RAGAS continuously evaluated faithfulness, relevance, and context precision to detect retrieval drift post-updates. Redis cached top-K chunks per session, reducing redundant vector lookups on follow-up turns. Reduced hallucination by 38%, improved retrieval precision by 52%, and maintained sub-60ms latency across Egyptian, Khaleeji, and Levantine query patterns.
Multi-agent Conversational Recommendation System (Turing)
https://github.com/shadabansari794/tripgearagent2Conversation memory managed via LangChain's buffer window memory maintained a multi-turn context for follow-up queries and preference tracking. A recommendation agent fused retrieved candidates with the user session history using a re-ranking prompt to surface personalized results. Hybrid dense-sparse retrieval improved recall on ambiguous or short queries common in voice input. Improved product discovery efficiency and increased user engagement by 35% through contextually grounded, conversational shopping experiences with sub-2s response latency.
Multi-agent AI Customer Support Platform using Amazon Bedrock
The platform integrates vector retrieval from documents stored in Amazon S3, contextual conversation memory, and scalable serverless infrastructure using AWS Lambda, API Gateway, and DynamoDB. The solution enables accurate knowledge-grounded responses and automated workflows, significantly improving response time and reducing manual support workload.
Credit Card Marketing Mix Model (MMM)
Integrated media spend, promotional calendars, and seasonal signals (Fourier terms for cyclical patterns) as exogenous regressors to improve forecast accuracy. Estimated channel-level ROI with posterior credible intervals to quantify uncertainty in budget recommendations. Built a constrained optimization layer using SciPy to simulate budget reallocation scenarios under spend caps, surfacing optimal channel mix for maximum acquisition yield. Model validation via MAPE and out-of-sample backtesting ensured reliability across campaign cycles.
Credit Risk & Loan Default Prediction System
Deployed as a real-time FastAPI scoring endpoint with sub-100ms latency, containerized on Docker with CI/CD versioned rollouts. Integrated SHAP explainability for per-prediction attribution to meet regulatory interpretability requirements. Monitored production stability via PSI for covariate drift, CSI for score distribution shifts, and AUC decay alerting. Score-cutoff optimization aligned to business risk appetite improved loan approval rates while reducing bad rates across borrower segments.
TelegramRAGBot
https://github.com/shadabansari794/TelegramRAGBotI designed and developed the complete pipeline, from data ingestion and cleaning to vector embedding and retrieval using FAISS. The system stores channel messages, converts them into semantic embeddings, and retrieves the most relevant context at query time for answer generation.
I implemented prompt engineering techniques to reduce hallucinations and enhance response relevance while optimizing for latency and scalability. The bot was deployed using FastAPI with asynchronous message handling, enabling near real-time interactions.
This project demonstrates production-ready RAG design, modular agent architecture, and efficient integration between LLMs, vector databases, and messaging APIs.
Credit Card Customer Churn Prediction Model
Automated Billing & Budget Workflow Platform
Leveraged n8n as the workflow orchestration engine to automate budget approvals, invoice generation, payment tracking, exception handling, and stakeholder notifications. Integrated QuickBooks for accounting operations, Airtable for operational planning, Google Workspace, and Slack/email notifications to create a seamless end-to-end workflow.
Implemented role-based access controls, configurable approval chains, validation rules, and audit logging to ensure compliance and data accuracy. Developed automated reconciliation workflows that synchronized financial records across systems, reduced manual effort by over 70%, and improved visibility into budget utilization and billing status. Deployed the solution on AWS with monitoring, logging, and CI/CD pipelines, enabling scalable and reliable operations across multiple business teams.
Multimodal Conversational AI System
My work included integrating Whisper for speech-to-text, TTS for voice responses, and generative image models for multimodal experiences. I built scalable data pipelines for chat datasets, embeddings, user profiles, and model evaluation, exposing capabilities via FastAPI endpoints with Rasa NLU for intent classification. I also benchmarked third-party LLM APIs against fine-tuned open-source models on latency, cost, and reasoning quality.
Finally, I applied neural networks, tree-based models, and reinforcement learning to improve personalization and response selection. I also established LLMOps workflows covering evaluation, monitoring, safety guardrails, and production deployment.
Title Production Spend Forecasting & Rigorous Time-series Backtesting
Education
Master's Degree in Mechanical Engineering
IIT Kharagpur - Kharagpur, India
Bachelor's Degree in Mechanical Engineering
Haldia Institute of Technology - India
Certifications
Academy Accreditation - Generative AI Fundamentals
Databricks
AWS Certified Cloud Practitioner
AWS
Knowledge Graphs for RAG
DeepLearning.AI
MCP: Build Rich-context AI Apps with Anthropic
DeepLearning.AI
Serverless Agentic Workflows with Amazon Bedrock
DeepLearning.AI
LangChain for LLM Application Development
DeepLearning.AI
Introducing Multimodal Llama 3.2
DeepLearning.AI
Data Structure and Algorithms
Udemy
Time Series Analysis
Coursera
Machine Learning
Coursera
Skills
Libraries/APIs
PySpark, XGBoost, CatBoost, PyTorch, Scikit-learn, REST APIs, Pandas, NumPy, SciPy, OpenAI API, TensorFlow, Pydantic, Auth, API Development, Beautiful Soup, WhatsApp API, React, Asyncio, Hugging Face Transformers, Imbalanced-learn, Node.js, Rasa NLU, OpenAPI, OpenCV, Claude API, GraphQL API, vLLM, Matplotlib
Tools
Apache Airflow, Excel 2010, ChatGPT, Microsoft PowerPoint, Pytest, AWS Glue, Microsoft Copilot, Amazon Textract, ARIMA, StatsModels, BigQuery, GitHub, Microsoft Excel, Claude Code, Claude, AI SDK, Amazon ElastiCache, Claude Agent SDK, n8n, Azure Kubernetes Service (AKS), Azure Machine Learning, Azure ML Studio, Codex, Git, Amazon EKS, Whisper, Rasa.ai, Docker Swarm, Azure OpenAI Service, Terraform, Observability Tools, You Only Look Once (YOLO), Visual Language Models (VLMs), Amazon SageMaker, Open Neural Network Exchange (ONNX), GraphRAG, GIS, RingCentral, Grafana, GitLab CI/CD, Seaborn
Languages
Python, SQL, Python 3, SAS, Snowflake, JavaScript, TypeScript, Go, C#, Java
Frameworks
Flask, LangGraph, LightGBM, Agentic Frameworks, Spark, Optuna, TensorFlow Lite, AutoGen, Django, LlamaIndex, Streamlit, Next.js, .NET, ASP.NET
Paradigms
Model Context Protocol (MCP), Automation, Microservices, ETL, Testing, Rule-based Programming, Asynchronous Programming, Event-driven Design (EDD), Synthetic Data Generation, High-performance Computing (HPC), HIPAA Compliance, Event-driven Architecture, Business Intelligence (BI), DevOps
Platforms
AWS Lambda, Amazon Web Services (AWS), Azure, Google Cloud Platform (GCP), Docker, Kubernetes, Amazon EC2, Jupyter Notebook, Azure AI Search, Microsoft Copilot Studio, Azure AI Studio, Vertex AI, AWS IoT, LangSmith, Twilio, Vercel, Ollama, Harness, Langfuse, Palantir Foundry, Azure Functions, Replit, Cloud Run, Databricks, Microsoft Power Platform, CrewAI, Observable Framework, Kubeflow, Cortex
Storage
Amazon S3 (AWS S3), Data Pipelines, PostgreSQL, Redis, Neo4j, Graph Databases, MongoDB, Datadog, JSON, Data Lakes, Azure Cosmos DB
Industry Expertise
Banking & Finance, Project Management, Healthcare, Marketing, Bioinformatics
Other
Machine Learning, Deep Learning, Time Series Analysis, Statistics, Retrieval-augmented Generation (RAG), Agentic AI, Chatbots, Natural Language Processing (NLP), Generative Artificial Intelligence (GenAI), Model Monitoring, Feature Engineering, Metabase, Data Analytics, Data Visualization, Modeling, Regression, Random Forests, Linear Regression, Logistic Regression, Neural Networks, Data Analysis, Computer Vision, YOLOv5, Optical Character Recognition (OCR), Convolutional Neural Networks (CNNs), FastAPI, LangChain, Prompt Engineering, Large Language Models (LLMs), OpenAI GPT-4 API, Conversational AI, Forecasting, Meta Llama, OpenAI GPT-3 API, Knowledge Graphs, API Integration, OpenAI, Artificial Intelligence (AI), Data Science, Machine Learning Operations (MLOps), Pinecone, APIs, A/B Testing, Data Analytics (Marketing), Marketing Analytics, Hypothesis Testing, Predictive Analytics, Marketing Mix Modeling, Data Collection, Web Scraping, Predictive Maintenance, Architecture, AI Modeling, AI Model Training, Vector Databases, eCommerce, Anthropic, Sentiment Analysis, AI Automation, ChatGPT API, Data Modeling, AI Tools, AI Chatbots, FAISS, Document Processing, Large Language Model Operations (LLMOps), Financial Systems, Document Parsing, Statistical Modeling, MLflow, Hugging Face, Financial Markets, Algorithms, Time Series Forecasting, Credit Underwriting, Underwriting, Analytics, Predictive Modeling, Marketing Attribution, Communication, Performance Marketing, Text Classification, Minimum Viable Product (MVP), Classification, Decision Trees, Gradient Boosting, Time Series, K-means Clustering, Dimensionality Reduction, CI/CD Pipelines, Data Engineering, Data Handling, Workflow, Deployment, Llama 3, Object Detection, Business Analysis, AI Consulting, AI Design, OAuth, Generative Pre-trained Transformers (GPT), Tesseract, Chatbot Conversation Design, Open-source LLMs, Speech-to-Text (STT), Text-to-Speech (TTS), RAG Pipelines, AI Assistants, Agentic RAG Systems, Deep Neural Networks (DNNs), Natural Language Understanding (NLU), NLU, RAG Architecture, ETL Pipelines, Model Validation, AI Copilots, Financial Modeling, AI-generated Video, Workflows, Data Privacy, ML Pipelines, Scalability, Semantic Search, Data Governance, Data Scientist, AI Architecture, Risk Modeling, Risk Models, Website Data Scraping, ETL Tools, RAG Systems, Custom Models, Credit Risk, Financial Data, Financial Data Analytics, RESTFul APIs, Foundry, Low Code, Financial Analysis, AI Pipeline, AI Agents, Debugging, API Design, Platform Design, Scraping, Probabilistic Modeling, Logistics & Supply Chain, Consulting, Scalable Vector Databases, Hyperparameter Tuning, Benchmarking, Cloud Platforms, Multi-agent Systems, Model Evaluation, System Architecture, Supabase, ElevenLabs Solutions, Cloud, Product Development, Regression Modeling, Statistical Analysis, Statistical Methods, AI Agents, Agentic AI, Artificial Intelligence (AI), Generative Artificial Intelligence (GenAI), Large Language Models (LLMs), Python, Machine Learning, Product Management, Amazon Web Services (AWS), Azure, Product Development, System Design, OpenAI SDK, Platform Engineering, Technical Writing, AI Marketing, Coding, Gemini, LoRa, QLoRA, BERT, SHAP, EDA, Leadership, Team Leadership, SDKs, Finance, Vector Search, Unity Catalog, Agentic AI Systems, Governance, Regression Testing, Cloud Architecture, Data Architecture, Enterprise Architecture, Solution Architecture, Stakeholder Management, Decision Modeling, Clustering Algorithms, Recommendation Systems, Prediction Markets, Data Warehousing, Embedding Models, ETL Development, Azure AI Services, AI Agent Orchestration, Light LLMs, Pgvector, UiPath, Bayesian Statistics, Multiagent Generative Systems (MAGs), LLM Integration, Semantic Code, Churn Analysis, Large-scale Data Processing, Cloud Governance, API Gateways, Agentic Coding, Audio Processing, Automatic Speech Recognition (ASR), Distributed Systems, Monitoring, Cursor AI, Industrials, Azure Function App, Amazon Bedrock AgentCore, Full-stack, Application State Management, GeoPandas, Data Security, Quantization, AI Hallucinations Management, Agentic Workflow Design, HIPAA, Qwen, AI platform engineer, Machine Learning Infrastructure Engineer, Demand Forecasting, Model Deployment, Model Development, Datasets, LLM Fine-tuning, Documentation, LLM Reasoning, Image Generation, Multimodal GenAI, Small Language Models (SLMs), Technical Architecture, Data Quality, Delta Lake, Apollo, Clay, Text to Image, Text to Image AI, Software Architecture, Technical Leadership, Back-end, Data Extraction, Google Document AI, Handwriting Recognition, Applied AI, Early-stage Startups, SaaS, Startups, Combinatorial Optimization, Cost Modeling, LLM Agents, OpenAI Agents SDK, AI Model Integration, Image Classification, Multimodal Models, Reinforcement Learning, Scientific Data Analysis, Object Recognition, Workflow Automation, Reliability, Spatial Analysis, AI Voice Agents, Google Calendar, Real-time Data, Fraud Detection, Reporting, Observability, CRM APIs, Workflow Automation & System Integration, WhatsApp, Image Segmentation, Webhooks, Qdrant, ColBERT, Pricing Elasticity, Pricing Models, Artificial Intelligence as a Service (AIaaS), iGaming, Clinical Research, ONNX Runtime, Graphics Processing Unit (GPU), Fine-tuning, Azure Databricks, Telemetry, Product Engineering, Content, Azure AI Document Intelligence, Streaming Data, Semantic Kernel (SK), NVIDIA TensorRT, Airtable, Transformers
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring