
Viktor Cikojevic
Verified Expert in Engineering
Machine Learning Engineer and Developer
Split, Croatia
Toptal member since September 3, 2026
Viktor is a senior ML research engineer with over eight years of experience in generative modeling, LLM systems, and production deployment. At Salient Predictions, he co-led the design and training of GEM-2, a 275-million-parameter generative model that cut training cost by 100x and inference cost by 10x. Viktor brings deep expertise in PyTorch, distributed multi-GPU training, and GCP for AI-driven startups.
Portfolio
Experience
- Python - 6 years
- Machine Learning - 5 years
- GCP - 5 years
- Computer Vision - 4 years
- XGBoost - 3 years
- TensorFlow - 2 years
- AI Agents - 2 years
- Diffusion Models - 2 years
Preferred Environment
Docker, GCP, BigQuery, Firebase, Vercel, GitHub, Git, PyTorch, Pydantic, Python
The most amazing...
...generative modeling system I've built is GEM-2, which cut training cost by 100x and inference cost by 10x compared to prior diffusion ensembles.
Work Experience
Machine Learning Engineer
Salient Predictions
- Acted as a major contributor to Salient’s generative weather models, under the GEM family (GEM-1, GEM-2, GEM-3): architecture design, training-data preparation, distributed multi-GPU training, and evaluation systems.
- Co-led the technical design and trained and deployed the diffusion and flow-matching generative models powering GEM-1.
- Co-led the technical design of GEM-2 (equal contribution, arXiv:2601.03753).
- Co-led the technical design of GEM-3 (arXiv:2608.06241).
- Co-invented 2 US AI patents (one for GEM-1 and one for GEM-2).
- Delivered a prior-generation production model: achieved a 20% skill score improvement over the previous baseline.
Data Scientist
Bellabeat
- Developed a multi-task transformer for menstrual health predictions.
- Deployed PyTorch models on-device for the iOS and Android apps, converting to Core ML (iOS) and ONNX (Android) for low-latency, offline inference.
- Shipped user-facing health features end-to-end. Performed data engineering and production pipelines, monitoring, and governance on GCP through to user-facing report design, collaborating with product and design to put model outputs in front of users.
- Owned product and business analytics: analyzed app usage, built onboarding and subscription funnels, created retention and engagement cohorts, and implemented experiment readouts that shaped product and growth decisions.
Computer Vision Engineer
Necogi by Codeasy
- Built an aerial imagery analysis pipeline in PyTorch for object detection and classification.
- Helped set the project from its idea stage to a viable MVP.
- Trained the AI model, evaluated it, and deployed it as an MVP.
Experience
GEM Global Weather Models
https://arxiv.org/abs/2601.03753• GEM-1 used diffusion and flow matching.
• GEM-2 (https://arxiv.org/abs/2601.03753), which I co-led with equal contribution on the paper, collapsed the slow multi-step sampling into a single forward pass, while beating operational forecast systems.
• GEM-3 (https://arxiv.org/abs/2608.06241), which I co-authored, trains one set of weights whose forecast timestep is chosen at inference.
I also co-invented two US AI patents, one for GEM-1 and one for GEM-2.
LLM Agents and RAG on a Single 16 GB GPU
https://www.kaggle.com/competitions/llm-20-questionsIn the LLM Science Exam, the questions were written by GPT-3.5, and only 200 labels were given, so we generated around 160,000 synthetic training questions from clustered Wikipedia pages and fine-tuned DeBERTa answer models behind a FAISS retrieval pipeline. To add a much bigger judge, I ran a 70B Llama on one T4 by streaming it layer by layer from disk, reusing the KV cache across the five answer options.
Transcribevoice.app, a Transcription SaaS
https://transcribevoice.appThe GPU worker went through eight versions on the RunPod serverless. Speaker labels come from my own diarization pipeline: NVIDIA Sortformer on overlapping 8-minute chunks, stitched with embedding clustering so labels stay consistent over hours of audio.
The product side is Next.js on Vercel: Firebase auth, Stripe credit packs, live transcript streaming into the UI, Claude-powered summaries and chat, 25 languages, and a free tier that runs Whisper entirely in the visitor's browser.
On-device Menstrual Cycle Prediction Transformer
https://period-diary.com/scientific-article/period-ovulation-tracking-advanced-menstrual-tracking-algorithms/The results are described in a published article on the tracker's site. I took the model to production on-device, converting PyTorch to Core ML (iOS) and ONNX (Android) so predictions run offline, and health data stays on the phone. I built the data pipelines, monitoring, and governance on GCP.
Lumbar Spine MRI Grading
https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classificationOur pipeline had three stages. U-Net ensembles first locate the anatomical keypoints on each MRI series. Then, instead of training a model to tell left from right or assign vertebral levels, we computed that geometrically from DICOM metadata, which maps every pixel into 3D patient space. Finally, a shared CNN with three specialized heads grades the severity of crops around each keypoint, weighted to match the competition metric.
We validated by re-implementing the official metric locally over a 5-fold cross-validation, and shipped the whole thing as a fully offline Kaggle notebook.
Education
PhD in Physics
UPC Barcelona - Barcelona, Spain
Master's Degree in Physics
University of Split - Split, Croatia
Skills
Libraries/APIs
PyTorch, XGBoost, CatBoost, Pandas, NumPy, TensorFlow, Keras, Scikit-learn, Dask, Pydantic, PyTorch Lightning, Python Asyncio, Stripe, Hugging Face Transformers, OpenCV
Tools
BigQuery, ChatGPT, Open Neural Network Exchange (ONNX), Git, GitLab, Whisper
Languages
Python, SQL, C++, C, Bash
Platforms
Google Cloud Platform (GCP), Cloud Run, Ollama, NVIDIA NeMo, Firebase, Vertex AI
Storage
Data Pipelines
Frameworks
Core ML, Next.js, DSPy, Chainlit, Hydra, LightGBM, Optuna
Paradigms
Quantitative Research, Synthetic Data Generation
Other
Diffusion Models, GCP, Computer Vision, Artificial Intelligence (AI), Distributed Training, Probabilistic Forecasting, Deep Learning, Machine Learning, Generative Artificial Intelligence (GenAI), Applied AI, Data Science, Model Development, Model Evaluation, Linear Regression, Demand Forecasting, Computer Vision Algorithms, Transformers, Time Series Forecasting, AI Agents, Communication, Model Deployment, Data Engineering, Monitoring, Image Processing, Code Review, Calibration, Convolutional Neural Networks (CNNs), LSTMs, Natural Language Processing (NLP), LangChain, Xarray, WandB, Diffusion-based AI Models, Video Transformers, Machine Learning Operations (MLOps), Feature Engineering, Speech-to-Text (STT), LLM Agents, Light LLMs, Retrieval-augmented Generation (RAG), LLM Fine-tuning, LoRa, FAISS, Research, Science, Cloud Computing, Writing, University Teaching, Prompt Engineering, Medical Imaging, DICOM, Image Segmentation, Time Series Analysis, Signal Processing, 3D, Stakeholder Management
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring