
Abhilash V J
Verified Expert in Engineering
ML Engineer and Developer
Kerala, India
Toptal member since September 8, 2026
Abhilash is an ML engineer with 9+ years of experience building production AI systems across generative AI (GenAI), RAG, agentic workflows, NLP, computer vision, and document intelligence. He specializes in Python, LangGraph, FastAPI, PostgreSQL/pgvector, and Kubernetes, and has built multi-agent and source-grounded systems for regulated domains. Abhilash's recent work reduced the pharmaceutical compliance review effort by around 50% and improved RAG relevance across HR and legal use cases.
Portfolio
Experience
- Natural Language Processing (NLP) - 10 years
- Machine Learning - 10 years
- Deep Learning - 9 years
- FastAPI - 6 years
- Large Language Models (LLMs) - 4 years
- Retrieval-augmented Generation (RAG) - 3 years
- Agentic RAG Systems - 2 years
- Agentic AI - 2 years
Preferred Environment
Docker, PostgreSQL, Pgvector, Agentic RAG Systems, Agentic AI, Python 3, Natural Language Processing (NLP), LangGraph, FastAPI, RESTFul APIs
The most amazing...
...source-grounded regulatory AI platform I've architected combines hybrid RAG, reranking, and evidence citations to reduce manual compliance-review effort by 50%.
Work Experience
AI Engineer
Turing
- Architected an enterprise architecture-review platform using Python, FastAPI, LangGraph, and domain-specific business, technology, information, and data/AI agents. Generated source-grounded reports.
- Engineered deterministic and resumable LangGraph workflows with shared state, conditional routing, retries, pass/fail/escalation paths, parallel agent dispatch, background execution, and human-in-the-loop checkpoints.
- Built multimodal ingestion for PowerPoint, PDF, image, spreadsheet, SharePoint, and Confluence content using python-pptx, LibreOffice, PyMuPDF, and Pillow. Indexed text and visual evidence in PostgreSQL with pgvector with traceable citations.
- Developed asynchronous FastAPI, SQLAlchemy, and PostgreSQL services for initiative intake, file processing, review execution, run-status tracking, structured result persistence, and evidence-grounded streaming chat.
- Developed a source-grounded pharmaceutical document-verification system that cross-checked regulatory content against retrieved evidence and reduced manual compliance-review time by approximately 50%.
- Implemented hybrid RAG using dense embeddings, BM25 sparse retrieval, Reciprocal Rank Fusion (RRF), HyDE-style query expansion, adaptive chunking, cross-encoder reranking, and provenance tracking over PostgreSQL with pgvector.
- Deployed a GPU-optimized large language model (LLM) serving with vLLM on Kubernetes and integrated cost-aware routing and fallbacks across OpenAI, Azure OpenAI, Anthropic Claude, Google Gemini, and DeepSeek models.
- Built Azure Document Intelligence pipelines for complex PDFs and evaluation workflows with expert-validated synthetic data, regression tests, structured schemas, guardrails, tracing, and latency and cost monitoring.
- Improved RAG answer relevance by 35% for HR and legal assistants through retrieval, reranking, chunking, and prompt optimization, validated against curated expert evaluation datasets.
- Developed LLM-powered medical document extraction that reduced manual data entry by 75%. Built document-similarity classification, OCR, and domain-specific LLM fine-tuning workflows.
Senior Software Engineer
Hexaware Technologies
- Designed a planogram compliance system using object detection and spatial matrix comparison against master shelf configurations, automating retail shelf audits previously performed by manual store inspection.
- Developed a question-answering and information retrieval system using Elasticsearch with Hugging Face BERT models for passage ranking and answer selection over internal document collections.
- Developed a content-moderation proof of concept (PoC) for digital communication platforms, training a toxic-comment classifier to identify abusive and inappropriate messages using supervised ML and NLP pipelines.
- Engineered REST APIs for model inference and lifecycle management in Python using Django, Flask, and FastAPI, standardizing how models were served across client projects.
- Established MLOps practices covering containerized deployment model versioning, automated release workflows, and built monitoring dashboards for model usage and prediction quality.
AI Engineer
Innovation Incubator Advisory
- Designed a convolutional CRNN optical character recognition model with connectionist temporal classification loss in Keras, replacing fully connected layers with 1D convolutions to shrink model size and speed up inference. Achieved 90% accuracy.
- Built an online model update pipeline on Amazon S3 with a key-value store enabling zero-downtime model deployments without taking the inference API offline.
- Implemented OCR post-processing with regex rules and ML-based correction for dates, amounts, names, and financial fields, where specialized correction models delivered a critical 2-3% accuracy gain on financial fields.
- Developed a structured data extraction pipeline for form identity documents and cheques using word-level bounding box coordinates combined with rule-based field matching.
- Built a knowledge-graph FAQ chatbot on the Grakn graph database. Entity-driven graph queries answered correctly across all evaluated cases where entity extraction succeeded, outperforming a BiLSTM question-answering baseline.
- Deployed Rasa and Dialogflow chatbots integrated with the knowledge graph.
- Developed collaborative filtering recommender systems plus text field classification models for the automotive domain.
Associate Consultant, AI Research and Development
Applexus Technologies
- Developed face detection, recognition, image retrieval, image segmentation, and video analytics systems using convolutional neural networks and U-Net architectures in Keras, TensorFlow, OpenCV, and dlib.
- Built NLP systems for pattern matching, sentiment analysis, chatbots, including recurrent neural network and LSTM models for sequence classification, language modeling, and question-answer selection.
- Deployed streaming analytics infrastructure on Amazon EC2 running Kafka, Spark, Cassandra, with ingestion pipelines on Amazon Kinesis, Spark on EMR, and DynamoDB.
- Delivered end-to-end data science workflows spanning feature engineering, exploratory data analysis, modeling, evaluation using XGBoost, LightGBM, bagging, and boosting ensembles for forecasting and classification.
- Automated deployment of Python Flask applications to Amazon EC2 using boto3, replacing manual release steps with scripted provisioning.
Automation Engineer, AI Research and Development
Curvelogics Advanced Technology Solutions
- Built Facebook Messenger chatbots using spaCy NLTK for intent handling and entity extraction.
- Developed face-recognition object-detection object-classification systems using deep learning in Python.
- Implemented occupancy-grid mapping, particle-filter localization, robotic-arm motion planning using rapidly-exploring Random Trees, MATLAB, and ROS.
- Presented a deep learning session at the Computer Society of India workshop "Data Science: From Sales Forecast to Cognitive Computing" in August 2017.
Experience
Source-grounded Regulatory Document Verification System
I implemented hybrid retrieval with dense embeddings, BM25, RRF, HyDE-style query expansion, adaptive chunking, cross-encoder reranking, and provenance tracking on PostgreSQL with pgvector. I also integrated Azure Document Intelligence for complex PDFs, structured LLM outputs, guardrails, evaluation datasets, regression tests, tracing, and latency/cost monitoring. GPU-optimized model serving used vLLM on Kubernetes, with routing and fallbacks across multiple commercial and open-source model providers.
The solution reduced manual compliance-review effort by approximately 50%.
Modular Multi-domain RAG and Document Intelligence Platform
The system separates the ingestion, parsing, chunking, embeddings, sparse retrieval, vector search, query expansion, reranking, generation, evaluation, and model-serving layers, so that encoders, LLMs, rankers, prompts, and retrieval strategies can be swapped per domain. I used hybrid dense and BM25 retrieval, RRF, cross-encoder reranking, pgvector/Elasticsearch-style indexes, structured evaluation, and domain-specific prompts and schemas. Retrieval, chunking, reranking, and prompt optimization were employed to improve answer relevance for HR and legal assistants, while related medical-document extraction workflows reduced manual data entry time.
Enterprise Agentic Architecture Review and Compliance Platform
The platform produces structured summaries, risks, clarification questions, alignment findings, and source-grounded review results while maintaining auditability and human oversight.
Education
Bachelor's Degree in Electronics and Communication Engineering
Cochin University of Science and Technology - Pathanamthitta, India
Certifications
Kaggle Competitions Expert
Kaggle
Deep Learning Specialization
DeepLearning.AI
Robotics Specialization
Coursera
Machine Learning Specialization
University of Washington | via Coursera
Java Programming and Software Engineering Fundamentals
Coursera
Python for Everybody
Coursera
Skills
Libraries/APIs
TensorFlow, PyTorch, Python-pptx, PiLLoW, SQLAlchemy, vLLM, Keras, OpenCV, Dlib, SpaCy, Hugging Face Transformers, XGBoost, Scikit-learn, CatBoost, REST APIs
Tools
LibreOffice, Rasa.ai, Dialogflow, Jira, Confluence, MATLAB, You Only Look Once (YOLO), Google AI Platform, GraphRAG
Languages
Python, SQL, JavaScript, TypeScript, VHDL, Python 3, HTML, CSS, Java
Frameworks
LangGraph, Agentic Frameworks, Django, Flask, LightGBM, Apache Spark, LlamaIndex
Storage
On-premise, PostgreSQL, Elasticsearch, Neo4j, MongoDB, Cassandra, JSON, Relational Databases, Redis
Paradigms
Machine-learned Ranking (MLR), Text Retrieval, Database Design, Model Context Protocol (MCP)
Platforms
Kubernetes, Amazon EC2, Docker, Apache Kafka, CrewAI
Other
FastAPI, Machine Learning, Artificial Intelligence (AI), Large Language Models (LLMs), Agentic AI, AI Agents, LLM Agents, Optical Character Recognition (OCR), AI Architecture, RAG Architecture, Communication, Document Processing, Fine-tuning, Model Evaluation, Computer Vision, Natural Language Processing, LangChain, Multi-agent Systems, Agentic AI Systems, Multi-agent Orchestration, AI Systems, Conversational AI, Context Engineering, Knowledge Graphs, RAG Systems, Architecture, Legal Technology (Legaltech), Hugging Face, Startups, Data Science, Pgvector, BM25, OpenAI, Google Gemini, Object Detection, Convolutional Neural Networks, NLTK, Deep Learning, Microsoft Azure, MLflow, Tokenization, Text Classification, Reranking, Image Segmentation, Document AI, Pinecone, Milvus, CI/CD Pipelines, Electronics, PID Controllers, Digital Electronics, Convolutional Neural Networks (CNNs), Natural Language Processing (NLP), Long Short-term Memory (LSTM), hyper parameter tuning, Machine Learning Operations (MLOps), Image to Vector, Image Search, Feature Engineering, hyper prameter tuning, Ensemble Modeling, model stacking, Exploratory Data Analysis, Information Retrieval, Retrieval-augmented Generation (RAG), Vector Search, query expansion, Robotics, Motion Planning, Pose Estimation, Statistical Modeling, Control Systems, Bayesian Statistics, Software Engineering, Debugging, Algorithms, Computer Programming, Encryption, Data Structures, RESTFul APIs, RESTful API Design, File I/O, Program Development, Embedding Models, Hybrid Search, Agentic RAG Systems, Vector Databases, Prompt Engineering, Azure AI Document Intelligence, System Architecture Design, OpenAI Agents SDK, Recommendation Systems, Ontologies, Weaviate, Software Architecture, AI Governance, AI Security
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring