
Muhammad Talha Zubair
Verified Expert in Engineering
Artificial Intelligence Developer
Lisbon, Portugal
Toptal member since June 6, 2023
Muhammad is an accomplished AI developer with over 7 years of experience in GenAI, data science, ML, and computer vision. He's made substantial contributions to several high-profile projects. His versatile tech stack includes expertise in AI agent development, GenAI Integrations, LLM fine-tuning, proficiency in deep learning frameworks, and experience with MLOps and AWS services. Muhammad's competencies extend to both back-end and front-end development.
Portfolio
Experience
- Amazon Web Services (AWS) - 6 years
- Docker - 4 years
- Vector Databases - 3 years
- Retrieval-augmented Generation (RAG) - 3 years
- Web Development - 3 years
- LangGraph - 2 years
- Agentic AI Systems - 2 years
- Model Context Protocol (MCP) - 1 year
Preferred Environment
Linux, Agentic AI, Agentic RAG Systems, Model Context Protocol (MCP), AI Agents, AI Development, Retrieval-augmented Generation (RAG), Software Development, Amazon Web Services (AWS), Google Cloud Platform (GCP)
The most amazing...
...project I've developed was a smart adtech recommendation system with adaptive product suggestions and real-time budget allocation based on live rate updates.
Work Experience
Senior AI Engineer
Accenture
- Architected, developed, and deployed an enterprise AI assistant platform supporting 500+ retail stores across Canada, enabling operational workflows, enterprise knowledge retrieval, and decision support for store managers and corporate teams.
- Designed and implemented scalable multi-agent AI systems using LangGraph with agent, node, and tool-based architectures, integrating MCP servers and FastAPI for contextual task execution, workflow orchestration, and dynamic agent routing.
- Designed and integrated FastMCP-based MCP servers for an internal financial and operational insights platform, enabling LangGraph agents to securely access enterprise data, invoke business tools, and execute contextual workflows.
- Improved AI application performance by reducing end-to-end pipeline latency by 10–15% through asynchronous processing, parallel execution, optimized service orchestration, and workflow optimization.
- Built advanced retrieval-augmented generation (RAG) pipelines by integrating OpenAI and Google Gemini models, delivering accurate, context-aware enterprise search and conversational AI capabilities.
- Created end-to-end AI infrastructure on Google Cloud Platform (GCP) using Vertex AI (Gemini models for inference and embeddings), BigQuery, Firestore, Cloud Storage, Cloud Composer (Apache Airflow), Google Kubernetes Engine (GKE), Artifact Registry.
- Orchestrated AI and data pipelines using Apache Airflow DAGs, automating document ingestion, embedding generation, scheduled workflows, model evaluation, and operational task execution.
- Implemented centralized LLM observability and prompt management using Langfuse, including prompt versioning, experiment tracking, trace logging, agent monitoring, and evaluation metrics.
- Implemented enterprise AI security controls, including LLM guardrails, content moderation, prompt safety, and PII detection/redaction, to ensure secure, compliant, and responsible AI deployments.
Senior Data Scientist
Turing
- Built a real-time financial insights platform for investment advisors, integrating OpenAI Whisper-based live transcriptions, sentiment analysis, and live stock data, reducing manual research time by approximately 40% per advisor session.
- Designed and deployed the platform on Microsoft Azure, leveraging Azure OpenAI Service, Azure AI Search, Azure Blob Storage, Azure Container Apps/AKS, and Azure Monitor to deliver a scalable, secure, and enterprise-grade AI infrastructure.
- Developed a RAG pipeline with a Django REST and WebSocket back end for intelligent financial data querying and conversational insights, supporting 100+ concurrent advisor sessions with low-latency responses.
- Engineered a smart recommendation system for an AdTech platform by combining classical machine learning models with LLM-driven intelligence, deployed on Vertex AI and integrated with BigQuery for large-scale audience segmentation.
- Built automated data pipelines for user behavior analysis and campaign optimization, improving estimated click-through rates by 15% through targeted high-conversion audience recommendations.
- Built PDF document processing and handling using Azure Document Intelligence.
GenAI Developer
Turing
- Designed and developed back-end services using Django with WebSocket support to enable real-time interaction between front-end applications and external data APIs.
- Implemented live transcription pipelines using OpenAI Whisper models to process streaming audio data in real time.
- Integrated retrieval-augmented generation (RAG) pipelines to allow intelligent interaction with large knowledge bases, supporting complex querying and document updates.
- Developed dynamic sentiment analysis modules focused on keyword-driven insights extracted from live or historical data.
- Created robust, scalable architectures capable of simulating real-time data ingestion and processing using cached datasets for enhanced testing and demo readiness.
- Built modular, extensible systems to easily plug in additional AI models, financial data sources, or SOP document repositories based on evolving project requirements.
AI/ML Expert
Kometsoft
- Developed a full-stack solution with a React front end and a FastAPI back end, ensuring a seamless transcription workflow for Malaysian Parliament YouTube videos.
- Integrated AWS services, including AWS Transcribe for Malaysian language speech-to-text conversion, AWS S3 for audio and text storage, AWS RDS for structured data management, and AWS Elastic Beanstalk for scalable back-end deployment.
- Leveraged AWS Bedrock's Nova model to integrate large language model (LLM) capabilities, enabling interactive user experiences such as keyword suggestions and dynamic text processing.
- Designed and implemented video transcription features, including URL input for YouTube videos, embedded video playback, Start/Stop transcription controls, real-time scrolling transcription, and click-to-play audio chunks stored in S3.
- Engineered the real-time transcription pipeline to initiate speech-to-text processing upon user command, with dynamic text updates and synchronized audio playback, achieving a smooth user experience with minimal latency.
- Built a keyword management system allowing users to highlight predefined keywords in the transcription, as well as add, edit, and delete custom keywords through an intuitive user interface.
AI Engineer
Libertify SAS
- Enhanced the overall retrieval-augmented generation (RAG) pipeline through prompt fine-tuning and decoupling of the LLM service from OpenAI, enabling service-agnostic functionality.
- Integrated PGVector, which improves retrieval accuracy and increases retrieval speed by 20% within the RAG pipeline, optimizing performance for large-scale document collections.
- Enhanced RAG accuracy by refining the PDF parser and integrating state-of-the-art solutions like Docling. This improved markdown extraction, enabling more accurate and efficient retrieval-augmented generation workflows.
- Deployed deep learning models on Google Cloud Platform (GCP), providing both public and private endpoints for scalable access.
- Integrated Llama as a large language model (LLM) service within the existing pipeline, broadening the range of supported AI capabilities.
- Enhanced the language detection system to accurately support Cantonese and Mandarin, significantly improving the chatbot’s multilingual proficiency.
- Enhanced voice generation using Elevanlabs and AWS TTS.
- Integrated Amazon Transcribe and Google Text-to-speech for speech-to-text conversion, audio-based document processing, and Amazon Textract for document handling, enhancing accessibility and automation in AI-driven applications.
Python Data Science Developer
An Online Freelance Agency
- Validated and ranked AI model responses for user queries across platforms like Llama, X.ai, and Character.ai.
- Developed and analyzed adversarial conversations with AI models.
- Collaborated on enhancing AI user experience and dialogue reliability.
- Analyzed and documented AI model failure points for development insights.
Machine Learning Engineer
Elm
- Established a real-time video streaming pipeline integrating Apache Kafka, PySpark, and Flask, processing with AI model computations for enhanced scalability.
- Developed a human gaze estimation system using a custom dataset and achieved an 80% + accuracy with multimodal and depth estimation algorithms.
- Designed and implemented a real-time video analysis pipeline leveraging detection, gender classification, and re-identification models, reducing processing time by 25%.
- Optimized deep learning models by post-quantization and framework conversion (e.g., PyTorch to ONNX), increasing inference speed by 30% and reducing the model size by 40%.
Associate Data Scientist | ML Engineer
TenX
- Reduced detection time for harmful bacteria and cells like Salmonella and Coccidia in chicken meat from 2 days to 15 minutes using microscopic imagery and advanced algorithms.
- Developed a tracking-over-segmentation pipeline with SAM and AOT tracker, enabling live data ingestion from Google Drive and automating uploads to an S3 bucket, improving data handling efficiency by 50%.
- Implemented computer vision algorithms to analyze drivers' behaviors, such as lane changes and hard braking, and detect road surface conditions. Deployed the system on AWS cloud and iOS platforms using CoreML, achieving an 85% accuracy rate.
Machine Learning Engineer
TeReSol
- Worked on both service-based and product-based streams, delivering facial recognition and vehicle detection and tracking solutions, respectively.
- Developed facial recognition solutions using algorithms like Haar cascades, MTCNN, and FaceNet in TensorFlow and deployed them on web platforms.
- Designed and deployed a vehicle object detection and tracking solution for thermal imagery using NVIDIA Jetson Tk1, enhancing thermal vision capabilities by 30% .
- Built and tested a feature-based tracking method using a normalized cross-correlation (NCC) template matching algorithm, SIFT feature selection, and the Kalman Filter.
- Managed a team of four, bridging the gap between hardware and software teams, and guided the annotation team for accurate object annotation in various videos.
AI/ML Engineer
Codistan
- Worked on Madhunt, an augmented reality game inspired by Pokemon GO, incorporating real-time object detection using YOLO and TensorFlow word2vec for finding related elements.
- Designed a deep reinforcement learning-based recommendation algorithm tailored specifically for custom users playing the game using Python.
- Developed reward functions within the reinforcement learning framework and seamlessly integrated the algorithms into the existing game structure.
- Handled queries from Firebase and AWS using Python using Firebase SDK and the Boto library.
- Implemented a YOLO image recognition algorithm for object detection within the game, analyzing images captured during gameplay.
- Conducted research and development on state-of-the-art recommendation systems based on reinforcement learning techniques. Explored existing recommendation systems based on machine learning algorithms, particularly collaborative filtering systems.
Experience
Voice Cloning
• Collected and refined reference voice samples using Amazon Transcribe from a former team member’s YouTube content.
• Utilized curated samples for voice cloning with industry-leading services, including ElevenLabs, Voice AI, and Speechify.
• Explored open-source models such as F5-TTS and OpenVoice, though their performance did not meet the required standards.
AI Ready Media
https://www.libertify.com/• Integrated PGVector as the vector database, optimizing similarity search and embedding management for improved scalability and efficiency.
• Designed and implemented an LLM-agnostic service, enabling seamless integration with various large language models, including OpenAI's GPT, ensuring flexibility for diverse client needs.
• Enhanced the PDF parsing mechanism, enabling accurate extraction of both structured and unstructured data for better AI-driven insights.
• Integrated Amazon Transcribe and Google Text to Speech to convert speech to text, allowing the system to process voice-based inputs and improve accessibility for document analysis.
• Gained expertise in back-end development, database optimization, and AI-based SaaS solutions, particularly in regulated domains, delivering impactful results in document-based knowledge transformation.
Advisor Insights at Scale
• Designed and developed back-end services using Django with WebSocket support to enable real-time interaction between front-end applications and the FinnHub API.
• Implemented live audio transcription feature using OpenAI Whisper models to transcribe earnings calls in real time.
• Developed sentiment analysis pipelines focused on dynamically detecting sentiment based on specific financial keywords mentioned during earnings calls.
• Built investor reaction modules to capture and display market sentiment indicators during live and recorded calls.
• Integrated a Retrieval-Augmented Generation (RAG) pipeline to enable intelligent querying and interaction over financial data, enhancing decision-making insights for advisors.
• Created a robust, scalable architecture simulating real-time data streams using cached datasets, ensuring demo flexibility and realism.
• Ensured seamless multimodal data processing by combining live voice transcriptions, textual financial reports, and real-time stock market data into a unified, interactive user experience.
GenAI-powered Pharma SOP Automation
https://v0-image-analysis-pdlb81.vercel.app/RAG Performance Improvement | Question-answering System
• Utilized GPT-3.5 and FAISS for efficient document retrieval from a diverse corpus, including technical manuals, academic papers, business reports, and historical documents.
• Stored extracted information in FAISS to ensure quick access to the most relevant documents.
• Addressed challenges related to handling large data volumes and ensuring high-quality, contextually relevant retrieval.
• Optimized the pipeline to reduce latency by refining the integration of retrieval and generation steps.
• Integrated ChromaDB for persistent storage, acting as a cache to complement FAISS and further reduce response times.
• Achieved improved accuracy and contextual relevance in answers, increasing user satisfaction with the system.
Pathogen Detection in Microscopic Imagery
https://www.ancera.com/• Built a custom Keras data generator for efficient dataset loading and preprocessing.
• Utilized OpenCV, k-means, PCA, and ResNet-50 for noise removal and image enhancement.
• Conducted data analysis using PostgreSQL and Pandas to extract meaningful insights from microscopy images.
• Implemented an MLOps pipeline with wandB for seamless dataset and model tracking.
• Integrated GitHub pre-commit hooks to enforce code formatting and maintain code quality.
• Deployed solutions using Docker and optimized version control and management with Azure DevOps.
• Enhanced Linux-based workflows by implementing bash scripts for automation.
• Established CI/CD pipelines and automated scheduling using cron jobs.
• Managed Docker containers with Kubernetes and tested deployments via Flask APIs.
Dash Cam Analytics Platform
• Converted the Torch YOLOv5 model to the Core ML format for efficient inference on iOS platforms.
• Used DeepLab for segmenting different road parts and applied image up-sampling techniques via cv2 superRes.
• Developed Bash scripts and Docker pipelines to automate training and testing processes.
• Collected and managed data using PostgreSQL, with comprehensive data analysis conducted through Pandas.
• Worked extensively with Amazon EC2 and S3 to ensure efficient deployment and storage.
• Implemented unit tests and integrated them with GitHub Actions for continuous testing and deployment.
• Prioritized code optimization and maintained version control using Git for high-quality development and collaboration.
• Tested Flask API implementations rigorously to ensure reliable back-end performance and functionality.
Vehicle Detection and Tracking over Thermal Imagery
• Trained PyTorch's YOLOv3-tiny model on a custom dataset to enhance detection accuracy.
• Created a Python-based feature-tracking pipeline using normalized cross-correlation (NCC), SIFT feature selection, and the Kalman Filter for optimal estimation.
• Analyzed detector and tracker performance, extracting valuable insights for system optimization.
• Led a four-person team, fostering collaboration between hardware and software teams and ensuring accurate object annotation for various videos.
• Deployed the Python solution on NVIDIA Jetson TK1, utilizing OpenCV for pipeline development and Numpy for efficient array operations.
• Improved performance with Cython and Numba, enabling real-time capabilities.
• Managed code versioning, reporting, and requirement tracking using Azure DevOps.
• Developed Linux-based frameworks and automated workflows using bash scripts for efficient pipeline execution.
Recommendation System Based on Reinforcement Learning
I implemented the YOLO image recognition algorithm to facilitate real-time object detection during gameplay while researching and developing cutting-edge recommendation systems based on reinforcement learning. Exploring existing recommendation systems led to a focus on machine learning algorithms, such as collaborative filtering systems.
I also designed and implemented facial recognition solutions using algorithms like Haar cascades, MTCNN, and FaceNet in TensorFlow, deploying these on web platforms. I used Flask and NGINX for efficient web hosting and server-side functionality and Azure DevOps for code versioning, reporting, and requirements management.
To ensure the highest quality standards, I prioritized code optimization and version control using Git. Finally, I handled CRUD operations on a MySQL database to manage daily user interactions with the application.
Gaze Estimation in Real-time Images
• Trained the gaze estimation model using a custom dataset developed in PyTorch.
• Utilized WandB as an MLOps tool to manage and monitor machine learning experiments effectively.
• Conducted model training and testing on Google Cloud Platform (GCP) instances, ensuring scalable compute resources for efficient performance.
• Optimized the training pipeline to enhance accuracy and reliability in gaze estimation tasks.
Custom Optical Character Recognition Systems
• Collaborated with a cross-functional team of software engineers and data scientists to optimize OCR workflows, reducing processing time by 30%.
• Conducted a comprehensive analysis of text data and improved OCR accuracy using preprocessing techniques such as tokenization, stemming, and lemmatization.
• Integrated advanced NLP techniques, including named entity recognition (NER) and sentiment analysis, to extract valuable insights from documents.
• Trained and evaluated deep learning models, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to enhance OCR performance on complex document types.
Evaluation Metrics for Stable Diffusion Generated Images
• Leveraged Detectron2 for object detection and analysis within the generated images.
• Implemented a supervised graph approach, incorporating a relative size graph to evaluate object realism.
• Enhanced the assessment process by ensuring accurate determination of object proportions and visual fidelity.
Developing a Speaker Diarization Algorithm Using LSTM
• Preprocessed audio data by extracting key acoustic features, normalizing speech signals, and reducing background noise for improved diarization accuracy.
• Fine-tuned diarization models to segment and label multiple speakers in parliamentary debates and call center conversations.
• Optimized speaker clustering methods by handling overlapping speech, adjusting model parameters, and refining segmentation accuracy.
• Developed a user-friendly interface to display speaker-segmented transcriptions, facilitating easier analysis of structured discussions.
• Contributed to government debate analysis by implementing diarization models that classified speech based on ministries, improving the categorization of legislative discussions.
Education
Bachelor's Degree in Computer Science
National University of Science and Technology - Islamabad, Pakistan
Certifications
Claude Certified Developer
Anthropic
Neural Networks and Deep Learning
DeepLearning.AI | via Coursera
Introduction to TensorFlow for Artificial Intelligence, Machine Learning, and Deep Learning
DeepLearning.AI | via Coursera
SQL for Data Science
UC Davis | via Coursera
Machine Learning
Stanford University | via Coursera
Skills
Libraries/APIs
TensorFlow, PyTorch, OpenCV, Pandas, Keras, Scikit-learn, NumPy, Google Sheets API, Google Vision API, JSON API, LSTM, Flask-RESTful, PyPDF2, OpenAI API, API Development, React, XGBoost
Tools
You Only Look Once (YOLO), Azure OpenAI Service, Amazon Textract, GitHub, Atlassian SDK, Open Neural Network Exchange (ONNX), Azure Machine Learning, Google Sheets, Claude, Amazon EKS, Slack, Jira, MATLAB, AWS Command Line Interface (CLI), NGINX, Amazon SageMaker, ChatGPT, AWS Step Functions, Docling, Amazon Transcribe, Docker Compose, Whisper, Claude Agent SDK
Languages
SQL, Python 3, Bash, Python, Bash Script
Platforms
Docker, Linux, Amazon EC2, Google Cloud Platform (GCP), Amazon Web Services (AWS), Firebase, Kubernetes, NVIDIA CUDA, Azure, AWS Lambda, Vertex AI, Langfuse
Frameworks
Flask, FastMCP, LangGraph, Django
Paradigms
Azure DevOps, DevOps, Test-driven Development (TDD), Unit Testing, Agile, Model Context Protocol (MCP)
Storage
PostgreSQL, MySQL, Amazon S3 (AWS S3), Data Pipelines, Google Cloud, Amazon DynamoDB
Industry Expertise
Project Management, Healthcare
Other
Machine Learning, Computer Vision, Facial Recognition, Fine-tuning, Back-end, Development, Meetings, Deep Learning, GitHub Actions, Optical Character Recognition (OCR), BERT, Natural Language Processing (NLP), Mobile Vision, Data Analytics, Image Recognition, NVIDIA Jetson TK1, Reinforcement Learning, Deep Reinforcement Learning, Research, Machine Learning Operations (MLOps), Image Processing, Computer Vision Algorithms, Convolutional Neural Networks (CNNs), Artificial Intelligence (AI), Speech-to-Text (STT), Language Models, Hugging Face, Videos, Graphics Processing Unit (GPU), Image Analysis, Frameworks, Object Detection, Medical Diagnostics, Architecture, FastAPI, Containerization, Motion Tracking, Large Language Models (LLMs), AI Chatbots, OpenAI, Generative Artificial Intelligence (GenAI), Cloud, Data Extraction, PDF, PDF Scraping, Retrieval-augmented Generation (RAG), LangChain, Site Reliability Engineering (SRE), Workflow Automation & System Integration, ChatGPT API, Amazon API Gateway, Amazon Bedrock AgentCore, Gemini, Tesseract, Chatbots, AI Agents, AI Content Creation, Voice, Vector Databases, Prompt Engineering, Open-source LLMs, Voice Analysis, Audio Analysis, Full-stack, Web Development, Agentic AI Systems, Agentic RAG Systems, Agentic Workflow Design, Leadership, MCP Servers, Software Architecture, Classification, Azure AI Document Intelligence, Document Processing, Software Development, Networks, Data Analysis, Machine Vision, Generative Adversarial Networks (GANs), Stable Diffusion, Instance Segmentation, Deep Neural Networks (DNNs), Variational Autoencoders (VAEs), Speech Recognition, Speaker Identification (SI), Speaker Diarization, Audio Streaming, Ray.io, Security, Data Science, Document Parsing, Vectorization, Llama 3, OpenAI GPT-3 API, FAISS, ChromaDB, PDF Splitter, Mistral AI, Pulumi, Pgvector, Meta Llama, ML Pipelines, Infrastructure as Code (IaC), CI/CD Pipelines, CGI, Text-to-Speech (TTS), ElevenLabs Solutions, Voice Cloning, Google Text-to-Speech, Automatic Speech Recognition (ASR), Gemini API, Finnhub, Recommendation Systems, Content-based Filtering, Regression, Random Forests, Google BigQuery, EDA, Agentic AI, AI Agent Orchestration, LLM Integration, AG-UI, Deep Agents, AI Development, AI Voice Agents, AI Avatars, RAG Architecture, Scalable Vector Databases, Agentic SDK
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring