
Sushant Shambharkar
Verified Expert in Engineering
AI Architect and Developer
Pune, Maharashtra, India
Toptal member since June 23, 2026
Sushant is an AI architect and tech lead with 10+ years of experience building and scaling production AI platforms across finance, adtech, analytics, and cloud ecosystems. Specializing in GenAI, agentic AI, LLM infrastructure, RAG, and distributed inference, he has a proven track record of transforming complex business workflows into intelligent products and autonomous AI capabilities. Sushant is a technical leader who drives AI strategy, platform architecture, and engineering excellence.
Portfolio
Experience
- Machine Learning - 12 years
- Machine Learning Operations (MLOps) - 12 years
- Python - 12 years
- Docker - 12 years
- Natural Language Processing (NLP) - 8 years
- RAG Architecture - 2 years
- LangGraph - 2 years
- Agentic AI - 2 years
Preferred Environment
Linux, Docker, Kubernetes, Terraform, Ansible, Apache Airflow, Amazon Web Services (AWS), Google Cloud Platform (GCP), Databricks, Git
The most amazing...
...platform I've architected is a multi-agent wealth management system that improved engagement by 35%.
Work Experience
AI Architect
AngelOne
- Architected and deployed a multi-agent wealth management platform enabling autonomous customer servicing, portfolio intelligence, and investment advisory workflows across multiple business functions.
- Designed planner-executor-reviewer agent workflows with persistent memory, RAG, and real-time market intelligence, reducing complex customer query resolution time by 70%.
- Built an enterprise AI platform standardizing agent orchestration, memory management, tool integration, evaluation, and governance, accelerating AI application delivery by 3x.
- Improved customer engagement by 35% through AI-driven portfolio insights, personalized investment recommendations, and proactive servicing experiences.
- Reduced advisor response latency by 60% by automating portfolio analysis, customer query resolution, and investment recommendation workflows using autonomous agents.
- Increased portfolio diversification by 28% through risk-aware allocation models, market intelligence pipelines, and AI-generated portfolio rebalance recommendations.
- Reduced relationship manager onboarding time by 50% by delivering AI-powered copilot capabilities, automated knowledge retrieval, and contextual client intelligence.
Tech Lead
Amagi
- Architected a distributed audience intelligence platform combining edge ML inference and centralized AI services, enabling real-time viewer segmentation and ad personalization across large-scale streaming workloads.
- Built low-latency edge inference pipelines using PyTorch and ONNX Runtime, reducing model serving latency by 60% while supporting millions of daily audience classification requests.
- Implemented NLP and LLM-powered content understanding services that automated content classification, audience insight generation, and campaign optimization workflows, improving operational efficiency by 40%.
- Designed and deployed cloud-native machine learning (ML) infrastructure on Kubernetes with MLflow-based experimentation and model lifecycle management, accelerating model deployment cycles by 3x.
- Optimized advertising relevance through AI-driven audience targeting and behavioral analytics, increasing click-through rates by 30% and improving viewer retention by 20%.
- Engineered scalable data pipelines on Snowflake and Delta Lake to unify audience, campaign, and engagement datasets, reducing analytics processing time by 50%.
- Spearheaded a cross-functional team of engineers and data scientists to deliver production AI capabilities for streaming advertisers, improving campaign performance, and accelerating feature delivery.
Lead Machine Learning Engineer
Slintel
- Pioneered an AI-powered Analytics Copilot platform that transformed business users into self-service data consumers through natural language querying, automated dashboard generation, and intelligent analytics workflows.
- Built NLP-to-SQL capabilities leveraging fine-tuned language models and semantic query understanding, reducing dependence on data teams and accelerating analytics delivery across business functions.
- Improved dashboard creation efficiency by 200% through automated insight generation, metadata-driven recommendations, and AI-assisted visualization workflows.
- Reduced dashboard load times by 3x by designing dynamic aggregation strategies that analyzed Snowflake query history and automatically routed workloads to optimized aggregate tables.
- Developed intelligent query optimization and workload management systems that reduced compute consumption while improving analytics responsiveness for enterprise users.
- Engineered large-scale data and ML pipelines on Snowflake, Spark, and cloud-native infrastructure to process billions of sales intelligence and technology adoption signals.
- Architected entity extraction and enrichment pipelines using NLP techniques to improve company intelligence, technology detection accuracy, and downstream AI model performance.
- Drove cross-functional initiatives spanning data engineering, machine learning, and platform architecture, accelerating delivery of AI-powered analytics capabilities across multiple product lines.
Senior Machine Learning Engineer
- Developed an AI-powered software delivery intelligence platform that analyzed source code, deployment telemetry, and operational signals to predict release risk and improve production reliability.
- Built machine learning models leveraging historical code changes, incident data, rollback events, and deployment outcomes to identify high-risk releases before production deployment.
- Reduced code-induced production failures by 25% through predictive risk scoring, deployment conflict detection, and automated release quality recommendations.
- Implemented CodeBERT-based code intelligence capabilities to understand code changes, identify risky modifications, and generate actionable deployment insights for engineering teams.
- Designed canary rollout recommendation systems that analyzed historical deployment patterns and failure signals to improve release safety and reduce rollback frequency.
- Built large-scale data pipelines to process software development, deployment, and operational telemetry data to support machine learning-driven engineering decisions.
- Minimized hotfix cycles through proactive release risk assessment, enabling engineering teams to detect and mitigate production issues before customer impact occurred.
- Collaborated with cross-functional engineering teams to integrate machine learning capabilities into software delivery workflows, improving deployment confidence and operational efficiency.
Senior Machine Learning Engineer
Goldman Sachs
- Developed an AI-assisted client communication platform leveraging BERT-based language models to automate response recommendations, semantic search, and intelligent message prioritization.
- Trained and fine-tuned enterprise NLP models on large-scale internal communication datasets and SME-curated golden labels, significantly improving recommendation relevance and response quality.
- Reduced client response latency by automating information retrieval, response drafting, and communication triage workflows for customer-facing teams.
- Built semantic search capabilities that enabled users to discover relevant historical communications, knowledge assets, and client interactions using natural language queries.
- Designed machine learning pipelines for continuous model training, evaluation, deployment, and monitoring, improving model performance and reducing operational overhead.
- Implemented intelligent prioritization systems that analyzed communication context, urgency signals, and historical engagement patterns to optimize client servicing workflows.
- Established automated experimentation and model governance frameworks, enabling rapid iteration and reliable deployment of production NLP solutions.
- Collaborated with business stakeholders and subject-matter experts to translate domain knowledge into machine learning features, training datasets, and measurable AI outcomes.
Machine Learning Engineer
Persistent Systems
- Engineered predictive workforce optimization systems for route planning, task duration estimation, and resource allocation, improving field workforce productivity and operational efficiency.
- Developed machine learning models that predicted job completion times, workforce availability, and scheduling constraints, enabling more accurate planning and execution.
- Built optimization pipelines that improved technician routing efficiency, reduced travel overhead, and increased daily task completion rates across distributed field operations.
- Migrated legacy SAS-based analytics workloads to scalable PySpark machine learning pipelines, reducing processing latency by 10x while eliminating licensing costs.
- Designed large-scale data processing frameworks capable of handling operational telemetry, workforce events, and scheduling data for predictive decision-making.
- Automated workforce planning workflows through predictive analytics and data-driven recommendations, reducing manual intervention and improving scheduling accuracy.
- Collaborated with product and engineering teams to translate business requirements into production-grade machine learning solutions deployed across enterprise customers.
- Implemented end-to-end model development, validation, deployment, and monitoring processes, establishing scalable machine learning engineering practices within the organization.
Experience
AgentForge — Enterprise Agent Platform
The system supports single-agent and multi-agent workflows, persistent memory, retrieval-augmented generation (RAG), tool orchestration, and policy-driven governance. AgentForge provides reusable agent templates, workflow builders, evaluation pipelines, tracing, and operational dashboards that help organizations accelerate AI adoption while maintaining reliability and compliance.
Key capabilities include agent lifecycle management, memory management, tool registries, automated benchmarking, guardrails, observability, and integrations with enterprise systems through APIs, databases, messaging platforms, and Model Context Protocol (MCP).
As the principal architect and developer, I designed the platform architecture, implemented orchestration frameworks, built evaluation and observability capabilities, and established scalable deployment patterns for enterprise AI workloads.
AgentPulse Platform
Platform captures agent interactions, tool usage, decision paths, costs, latency, and outcomes to create a comprehensive view of agent performance. It provides automated evaluation pipelines, KPI attribution, failure analysis, experiment tracking, and observability dashboards that help organizations identify which agents deliver measurable business value.
AgentPulse supports both single-agent and multi-agent systems, enabling teams to compare agent versions, detect regressions, analyze failure patterns, and optimize workflows using real-world outcome data. The platform bridges the gap between technical agent metrics and executive business reporting.
As the principal architect and developer, I designed the evaluation framework, implemented observability and KPI attribution systems, built experiment tracking capabilities, and developed analytics workflows that connect AI performance to business impact.
ProofMind Verification Platform
The platform combines workflow analysis, state machine validation, policy enforcement, and tool permission verification to identify potential failure paths, unsafe actions, and governance violations. ProofMind enables teams to define behavioral constraints, validate execution plans, and verify that agent workflows conform to organizational and regulatory requirements.
The system provides automated verification pipelines, policy testing, workflow simulation, audit trails, and compliance reporting for both single-agent and multi-agent architectures. By detecting risks before production deployment, ProofMind helps organizations improve trust, safety, and operational reliability.
As the principal architect and developer, I designed the verification framework, implemented workflow validation engines, developed policy enforcement capabilities, and built governance tooling for enterprise AI systems operating in regulated environments.
Education
Master's Degree in Computer Science
IIT Bombay - Mumbai, India
Bachelor's Degree in Computer Science
Nagpur University - Nagpur, India
Certifications
Python (Intermediate) Certificate
HackerRank
Problem Solving (Basic) Certificate
HackerRank
Python (Basic) Certificate
HackerRank
Java (Basic) Certificate
HackerRank
Certified MBA (CMBA)
Udemy
Skills
Libraries/APIs
Pydantic, Claude API, REST APIs, PyTorch, TensorFlow, SQLAlchemy, React, PySpark, vLLM
Tools
Git, Claude, Claude Code, n8n, ChatGPT, Flink, Terraform, Ansible, Apache Airflow, Grafana, Vault, Microsoft Copilot, Oracle Cloud Infrastructure (OCI) Generative AI
Languages
Python, SQL, TypeScript, HTML, Java, JavaScript, Snowflake, SAS, Bash, Python 3, Python 2, Rust, C++
Frameworks
LangGraph, Agentic Frameworks, Spark, LlamaIndex, MXNet, Ray
Paradigms
Model Context Protocol (MCP), Business Intelligence (BI), DevOps, ETL
Platforms
Docker, Databricks, LangSmith, Kubernetes, Amazon Web Services (AWS), Google Cloud Platform (GCP), Langfuse, Azure, Apache Flink, Linux, Apache Kafka, CrewAI
Storage
Databases, Data Pipelines, PostgreSQL, Elasticsearch, Redis
Other
Machine Learning Operations (MLOps), LangChain, FastAPI, Machine Learning, Artificial Intelligence (AI), Agentic AI, Natural Language Processing (NLP), AWS Bedrock AgentCore, Large Language Models (LLMs), Generative Artificial Intelligence (GenAI), AI Tools, AI Integration, APIs, Data Modeling, Agentic Workflow Design, Prompt Engineering, Minimum Viable Product (MVP), Large Data Sets, Cloud Infrastructure, Full-stack, API Integration, Full-stack Development, Architecture, Low Latency, Trading, API Design, Web Application Design, Agentic Coding, Software Architecture, Algorithms, AI Copilots, Agentic RAG Systems, User Experience (UX), Workflow Automation, Code Review, Communication, Team Leadership, Screeners, Interviewing, LLM Integration, System Architecture, Infrastructure, AI-generated Code, AI-assisted Development, Software Engineering, Back-end, AWS Cloud Architecture, Cloud Architecture, Solution Architecture, Retrieval-augmented Generation (RAG), Kafka, CI/CD Pipelines, AI Agents, Agentic AI Systems, RAG Architecture, Technical Leadership, Devin, Delta Lake, Hugging Face, LoRa, Reinforcement Learning from Human Feedback (RLHF), TensorRT-LLM, LiteLLM, Qdrant, Temporal, MLflow, OpenTelemetry, Prometheus, Platforms, AI Pipeline, ML Pipelines, Google ADK, Distributed Systems, ONNX Runtime, Multi-agent Systems, Large Language Model Operations (LLMOps), Vector Databases, A2A, Evaluation, GuardRails, Data Engineering, Analytics Engineering, Data Warehousing, ELT, Dashboards, Deep Learning, CodeBERT, Predictive Analytics, Cloud Computing, BERT, Custom BERT, Information Retrieval, Semantic Search, AI Model Training, Model Evaluation, Natural Language Understanding (NLU), Enterprise Search, Big Data, Business, AgentOp, AgentOps, AI Evaluation, Observability, Analytics, Key Performance Indicators (KPIs), Experimental Design, Formal Verification, Governance, State Machines, Risk Models, Team Management
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring