
Stanko Kuveljic
Verified Expert in Engineering
ML Engineer and Developer
Novi Sad, Serbia
Toptal member since July 29, 2026
Stanko is a staff-level ML engineer with 10 years of experience building end-to-end production AI systems, focusing on search, retrieval, and LLMs. His impact includes increasing eCommerce CTR by 17%, reducing model training time from 20h to 1h (cutting cloud costs by 90%), and automating regulatory workflows with 91% recall and 70% fewer manual reviews. Stanko has led ML teams and owns the entire lifecycle—from data to MLOps.
Portfolio
Experience
- Python - 10 years
- Machine Learning - 10 years
- Software Engineering - 10 years
- Machine Learning Operations (MLOps) - 6 years
- Semantic Search - 5 years
- Artificial Intelligence (AI) - 4 years
- Technical Leadership - 4 years
- AI Agents - 2 years
Preferred Environment
Machine Learning Operations (MLOps), Software Engineering, Machine Learning, AI Engineering, Semantic Search, AI Agents, Large Language Models (LLMs)
The most amazing...
...system I've built is an enterprise semantic search platform using custom embeddings and rerankers; the company behind the tech was later acquired by Shopify.
Work Experience
Senior Machine Learning Engineer
OLX Global
- Built an automated fine-tuning pipeline for embedding models using PyTorch, MLflow, and Amazon SageMaker, supporting parallel training and repeatable experimentation.
- Optimized model training on GPUs by reducing VRAM and CPU overhead and improving the efficiency of the training pipeline.
- Deployed hybrid search models as production APIs using a centralized model registry, LitServe, and Kubernetes.
- Worked across model training, experimentation, packaging, and production serving for semantic and hybrid search systems.
Founding AI Engineer
The ReelNews
- Owned the AI and relevance architecture for a real-time mobile news product, covering feed generation, personalization, vector retrieval, ranking, clustering, and LLM-powered content workflows.
- Built the foundations of a personalized feed using Qdrant vector search, user-following signals, notification signals, and custom ranking logic.
- Designed a news clustering approach combining semantic embeddings, time decay, and keyword extraction to group related stories as news develops.
- Built repeatable evaluation workflows for LLM prompts used in summarization and clustering.
- Worked directly with the founders to translate product requirements into AI architecture and production implementations.
Staff Machine Learning Engineer
SmartCat.io
- Architected scalable semantic search and ranking systems for Vantage Discovery (later acquired by Shopify) using embeddings, rerankers, and Databricks, streamlining MLOps workflows, reducing experimentation cycles, and improving CTR by 17%.
- Designed a graph-based RAG and multi-agent reasoning architecture using Neo4j, LangGraph, and Claude to orchestrate US CFR compliance workflows, achieving 91% recall in identifying applicable regulations and reducing manual review time by over 70%.
- Directed the ML engineering department, mentored teams, and established robust MLOps standards, driving the successful delivery of scalable AI systems across multiple enterprise client environments.
Machine Learning Engineer
SmartCat
- Engineered a real-time computer vision and IoT analytics platform using YOLO and TensorFlow, building end-to-end data pipelines that achieved 98% accuracy for production object tracking.
- Engineered a real-time ticket recommendation system for an online betting client using XGBoost and Apache Spark, driving a 20% conversion rate on recommended tickets.
- Delivered the entire MLOps infrastructure for the recommendation engine, including streaming data integration, a centralized feature store, automated model retraining workflows, and a high-performance FastAPI inference layer for low-latency serving.
Experience
Real-time AI News Feed
Enterprise Semantic Search and Ranking Platform
Graph-based Multi-agent Regulatory Compliance System
Scalable Hybrid Search and ML Pipeline Optimization
To further improve efficiency, I executed a deep optimization of model training on T4 GPUs. By applying advanced memory management techniques to minimize VRAM and CPU usage, I successfully condensed model training time from 20 hours down to just 20 minutes. This massive performance leap simultaneously drove a 90% reduction in associated AWS compute costs.
Finally, I deployed these assets to production, setting up the hybrid search models as highly scalable APIs. This was achieved by leveraging a centralized model registry alongside LitServe, orchestrated within a Kubernetes environment to ensure reliability and speed.
AI-native Music School Platform
I designed and implemented a deterministic, state-driven workflow using LangGraph, combining slot-filling mechanisms with dynamic UI elements and free-text inputs. I built a multi-agent system using OpenAI models and custom tools capable of handling both domain-specific Q&A and transactional scheduling operations. I also integrated FastAPI and PostgreSQL for the back-end service layer, alongside Langfuse for LLM observability and performance evaluation.
Enterprise Next Best Action ML Platform for Pharmaceutical Sales
I designed the core NBA architecture, built feature store pipelines in Snowflake, and established automated MLOps pipelines using AWS SageMaker for continuous model training and daily scheduled batch inference. To tackle low-engagement signal sparsity, I engineered recency-weighted interaction aggregations and cross-channel composite features. Implemented XGBoost models that provided daily actionable outreach recommendations directly to sales teams.
Enterprise Data Mesh Platform & POC Implementations
I designed and implemented end-to-end data product prototypes leveraging Databricks for scalable data processing and Apache Kafka for real-time data streaming. I also developed custom integration scripts and data pipelines in Python to connect client infrastructure to the platform, enabling self-serve data sharing, MCPs, and automated federated governance across business units.
Real-time Space Utilization & Sensor Analytics Platform
I contributed to scalable back-end services and high-performance API endpoints using Python and FastAPI. I improved spatial count accuracy by implementing data cleaning and processing algorithms with NumPy to filter signal noise and harmonize discrepancies between raw sensor counts and active bookings. I also built integration pipelines to ingest external sensor data feeds and managed scalable data storage pipelines using AWS S3.
Computer Vision Space Analytics Platform for Fish-eye Cameras
Built dual real-time and batch processing architectures. Implemented serverless inference on AWS Lambda for live stream processing, real-time occupancy counting, and automated event alerting. For historical analysis, built data processing pipelines using AWS Batch, S3, DynamoDB, and Amazon Athena to power custom Power BI reports of space analytics. Standardized MLOps workflows using MLflow and LakeFS for data and model versioning, accelerating onboarding time for new physical sites.
Education
Master's Degree in Electrical and Computer Engineering
University of Novi Sad - Novi Sad, Serbia
Bachelor's Degree in Electrical and Computer Engineering
University of Novi Sad - Novi Sad, Serbia
Certifications
Speaker: AI_dev: Open Source GenAI & ML Summit Europe 2025
The Linux Foundation
Skills
Libraries/APIs
REST APIs, PyTorch, API Development, Scikit-learn, XGBoost, NumPy, OpenCV, Pandas
Tools
Claude, Amazon SageMaker, You Only Look Once (YOLO), AWS Batch, Claude Code
Languages
Python, Snowflake, Java
Frameworks
LangGraph, Spark, Apache Spark
Paradigms
Machine-learned Ranking (MLR), Model Context Protocol (MCP)
Platforms
Docker, Databricks, Langfuse, Amazon Web Services (AWS), Kubernetes, AWS Lambda, Azure
Storage
Data Pipelines, Amazon S3 (AWS S3), Amazon DynamoDB, PostgreSQL, Neo4j
Other
Semantic Search, LLM Applications, AI Agents, Machine Learning Operations (MLOps), Software Engineering, Machine Learning, AI Engineering, Search, Artificial Intelligence (AI), LLM Integration, Machine Learning (ML) APIs, Back-end, Agentic AI Systems, AI Architecture, Large Language Models (LLMs), Retrieval-augmented Generation (RAG), MLflow, LangChain, Vector Databases, Technical Leadership, Reranking, Personalization, Generative Artificial Intelligence (GenAI), Data Science, AI Model Training, Prompt Engineering, Agentic AI, Recommendation Systems, RAG Systems, Conversational AI, Evaluation, Natural Language Processing (NLP), Sentiment Analysis, Multi-agent Systems, Context Engineering, Knowledge Graphs, Qdrant, System Design, Sentence Transformers, FastAPI, Hybrid Search, Embedding Models, Deployment, Vector Search, Distributed Systems, OpenAI, Feature Engineering, Data Mesh, Internet of Things (IoT), Anthropic, PDF Scraping, PDF, Computer Vision, Observability, Large Language Model Operations (LLMOps), Image Processing, Multimodal GenAI, Multimodal Models
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring