
Vazgen Tadevosyan
Verified Expert in Engineering
ML Engineer and Developer
Yerevan, Armenia
Toptal member since March 17, 2026
Vazgen is a senior ML/Python engineer with 7+ years of experience across LLMs, GenAI, RAG, computer vision, and agentic/multi-agent systems. He architects MCP-powered AI solutions with custom orchestration, hybrid retrieval, and scalable pipelines. An expert in Python and cloud (AWS, Azure, GCP), he has led teams and delivered production-grade systems grounded in research and real-world impact. He brings strong expertise in back-end, MLOps, and end-to-end system design from data to deployment.
Portfolio
Experience
- Machine Learning - 8 years
- Natural Language Processing (NLP) - 8 years
- Python - 7 years
- PyTorch - 7 years
- Deep Learning - 7 years
- Retrieval-augmented Generation (RAG) - 4 years
- AWS IoT - 4 years
- Generative Artificial Intelligence (GenAI) - 4 years
Preferred Environment
Python, PyTorch, Retrieval-augmented Generation (RAG), Generative Artificial Intelligence (GenAI), Databases, Model Context Protocol (MCP), Artificial Intelligence (AI), Agentic AI, Amazon Web Services (AWS), Pinecone
The most amazing...
...thing I've built is an MCP-powered multi-agent orchestration system that assists pharmacologists as an intelligent assistant with tool access.
Work Experience
Senior ML Engineer
Comm-IT
- Architected and delivered end-to-end AI products, including chatbots, agents, semantic search, and document-intelligence workflows, across travel, hospitality, insurance, finance, and healthcare using LLMs, vector databases, and multimodal pipelines.
- Developed Redshift-backed knowledge bases for semantic search and information retrieval, incorporating relational databases. Managed model deployment, prompt tuning, benchmarking, and embedding optimization with FAISS and Pinecone.
- Built production-grade RAG systems for chatbots and search, translating academic ideas into scalable code and leveraging AWS Bedrock for managed foundation-model inference and rapid model iteration.
- Built a serverless AWS stack with sub-second latency using Lambda, API Gateway, DynamoDB, and S3, incorporating CloudWatch and AWS Powertools for full observability and achieving auto-scaling to thousands of concurrent requests.
- Led and mentored a 5-engineer ML team, instituting code-review standards, MLOps best practices, and agile workflows that cut release cycles by 30 %.
Senior ML Engineer
Cognaize
- Developed advanced RAG pipelines using SPLADE-v2 implemented from scratch in PyTorch, ReAct agents, and domain-specific LLM fine-tuning for financial document analysis.
- Built a multi-agent AI-powered search engine to identify financial document types, integrating external search APIs like Bing and Brave for web-scale augmentation.
- Implemented LoRA-based instruction fine-tuning, boosting accuracy on specialized tasks.
- Engineered memory-aware conversational AI with persistent chat context for enterprise clients.
- Designed a knowledge-base chatbot covering 1,000+ financial documents, including 10-Ks, annual reports, and ESG, delivering accurate retrieval and summarization with persistent chat context.
- Conducted hands-on DevOps work with Docker and serverless Google Cloud services, including Cloud Functions (Fission), Cloud Run, and Vertex AI.
- Supervised 4 engineers, conducting weekly sync-ups, code reviews, and architecture discussions.
Data Scientist
Plat.AI
- Recommended a set of attributes to maximize model accuracy using statistical analysis.
- Identified clients likely to repay loans using random forest and logistic regression models with a 0.85 AUC score.
- Developed an ML deployment pipeline using Docker and REST API, thus automating the data preprocessing, data cleaning, and feature extraction, and optimizing prediction in real time.
Graduate Research Assistant
Rochester Institute of Technology
- Conducted research on ML in cybersecurity and examined the impact of imbalanced data on the performance of intrusion detection systems.
- Proposed solutions to address the imbalance problem, including undersampling and oversampling using SMOTE.
- Used preprocessing techniques, including PCA, and applied clustering techniques, such as DBSCAN, to merge minority classes.
ML Engineer
SoftConstruct
- Reduced traffic load by 30% by finding sessions with not-human-like behavior, applying hierarchical clustering methods.
- Applied supervised tree-based models such as XGBoost and random forest to predict bots with an accuracy of 93%.
- Conducted data analysis of user sessions using Elasticsearch and Kibana, providing reports on user groups that overloaded network traffic.
Teaching Associate of Programming for Data Science Courses
American University of Armenia
- Created a set of code slides and problem sets as extracurricular materials for the Programming for Data Science Course.
- Assisted students in building Shiny dashboards and optimizing code for their course projects using Python and R.
- Led problem-solving sessions with classes of 30+ students and organized additional office hours to discuss topics.
Experience
Multimodal Representation Learning
https://github.com/paligonshik/CoMM_IMPLEMENTATIONKEY ACHIEVEMENTS
• Built the complete pipeline with BLIP-2 vision/ text encoders, modality-specific converters, transformer fusion, and InfoNCE-based MI objectives.
• Reproduced the paper's MM-IMDb results and verified the model's ability to capture redundancy, uniqueness, and synergy across modalities.
• Designed training/augmentation setup, critic modules, and multimodal evaluation following the theoretical formulation in the paper.
Capstone: Siamese Networks for Open Set Recognition
KEY ACHIEVEMENTS
• Compared contrastive and triplet loss. The contrastive loss trained faster and generalized better on unseen categories.
• Proposed an effective thresholding method to decide when an input should be rejected as "unknown."
• Evaluated Euclidean vs. Cosine similarity in embedding space. Euclidean distance proved consistently more stable and reliable.
• Demonstrated that prototype-based embeddings improved inference speed by approximately 25 times while maintaining around 96% classification accuracy.
Armenian TTS | VITS-2 Fine-tuning
https://github.com/paligonshik/vits2_pytorch_hyKEY ACHIEVEMENTS
• Fine-tuned VITS-2 on around 9 hours of Armenian speech, producing high-quality natural audio.
• Trained it using distributed PyTorch across doubled H100 GPUs with mixed precision and a custom tokenizer design.
• Explored both phoneme-based and character-based tokenization strategies for low-resource TTS.
Duplicate Pull Request Detection | Text and Code Similarity Modeling
https://github.com/paligonshik/PullRequest-Detection/blob/main/Refining_Duplicate_Contribution_Detection_in_Pull_Based_Project.pdfKEY ACHIEVEMENTS
• Implemented customized TF-IDF NLP pipeline (lemmatization and stop-word removal) and multiple code-level similarity metrics, including file-path matching, Jaccard for removed lines, and token-level comparison for added lines.
• Designed a weighted similarity aggregation method outperforming prior work (Li et al.) in recall across 16 GitHub projects.
• Demonstrated that combining heterogeneous similarity measures significantly improves Recall@K and robustness across repositories.
Sparse Lexical Expansion | SPLADE-v2 Inference Implementation
KEY ACHIEVEMENTS
• Integrated the model into the company's production RAG pipeline for financial documents.
• Benchmarked SPLADE-v2 against dense retrievers, showing clear gains in precision and retrieval robustness.
• Identified practical gaps between academic IR models and enterprise production constraints, including latency, memory, and stability.
Education
Master's Degree in Data Science
Rochester Institute of Technology - Rochester, NY, USA
Master's Degree in Industrial Engineering and Systems of Management
American University of Armenia - Yerevan, Armenia
Bachelor's Degree in Economics
Armenian State University of Armenia - Yerevan, Armenia
Certifications
Generative AI with Large Language Models
Coursera
Deep Learning
Coursera
Skills
Libraries/APIs
PyTorch, REST APIs, Amazon Rekognition, Bing API, LSTM, Pandas, vLLM, Scikit-learn, XGBoost, TensorFlow, WhatsApp API
Tools
Amazon CloudWatch, Amazon Textract, Git, GitHub, Azure OpenAI Service, Claude Code, Kibana, Kafka Streams, Microsoft Copilot, GitLab CI/CD, Grafana, Terraform
Languages
Python, SQL, Snowflake, GraphQL, Java
Frameworks
Django, Agentic Frameworks, Apache Spark, Streamlit, Spark
Paradigms
Object-oriented Programming (OOP), Model Context Protocol (MCP), Siamese Neural Networks, Automation
Platforms
AWS IoT, Docker, AWS Lambda, LangSmith, Ollama, Amazon Web Services (AWS), Google Cloud Platform (GCP), Vertex AI, Apache Kafka, Linux, Weights & Biases, Kubernetes
Storage
Amazon DynamoDB, Redshift, MongoDB, Databases, Data Pipelines, Elasticsearch
Other
Retrieval-augmented Generation (RAG), Generative Artificial Intelligence (GenAI), Machine Learning, Natural Language Processing (NLP), Software Engineering, Computer Vision, Distributed Systems, Reinforcement Learning, Data Science, Data Structures, Algorithms, Data Scraping, Business Analytics, Data Mining and Predictive Analytics, Amazon Bedrock AgentCore, AI Agents, ReAct Agents, FAISS, ChromaDB, Pinecone, FastAPI, Qdrant, AWS personlaize, Tesseract, LangChain, Search Engines, Llama 3, LoRa, Principal Component Analysis (PCA), Regression, Decision Trees, Hypothesis Testing, A/B Testing, Support Vector Machines (SVM), Linear Discriminant Analysis (LDA), K-NN, K-means Clustering, Conjoint Analysis, SMOTE, DBSCAN, Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Transformers, Hierarchical Clustering, Tf-idf, Linear Algebra, Calculus, Statistics, Mathematics, Torch, Text-to-Speech (TTS), AI Voice Agents, Audio Processing, Hugging Face, Large Language Models (LLMs), Chatbots, Machine Translation, Game Theory, Time Series, Artificial Intelligence (AI), Agentic RAG Systems, RESTFul APIs, RAG Architecture, GPU Computing, Graphics Processing Unit (GPU), Large Language Model Operations (LLMOps), Architecture, Agentic AI, AI Architecture, Training, AI Model Training, MLflow, Machine Learning Operations (MLOps), API Integration, AI Automation, Web Scraping, Website Data Scraping, AI Programming, AI Consulting, Probabilistic Modeling, Forecasting, Time Series Forecasting, APIs, Optical Character Recognition (OCR), Google Cloud Functions, R Programming, Elastic Net, Speech Recognition, Solution Architecture, CI/CD Pipelines, Prometheus, System Design, Logistics & Supply Chain, Deep Learning
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring