
Janos Horvath
Verified Expert in Engineering
Research Engineer and AI Developer
Santa Clara, CA, United States
Toptal member since February 17, 2025
Janos is an experienced research engineer and academic specializing in machine learning, computer vision, and video processing. His innovative work at Dolby Laboratories and Purdue University drives advancements in Dolby Vision and satellite image forensics. As a prolific author and patent holder, Janos brings a unique blend of technical expertise and visionary leadership, fostering collaborative breakthroughs in high-impact technology.
Portfolio
Experience
- C - 12 years
- Linux - 10 years
- Python - 9 years
- Computer Vision - 8 years
- AI Research - 8 years
- Machine Learning - 8 years
- Video Coding - 8 years
- Image Processing - 8 years
Preferred Environment
Python, Linux, n8n, Agentic AI, AI Chatbots, Interactive Voice Response (IVR), Image Classification, DeepSeek, Bash, Bash Script, Git, SSH, Terminal, Edge Computing, Google Vision API, Drones, YOLOv8, GitHub Actions, NVIDIA CUDA, Correlational Analysis, Feature Engineering, Statistical Analysis, Statistical Modeling, Funnel Analysis, Churn Analysis, Voice Chat, TypeScript, Data Extraction, Roboflow, Reinforcement Learning from Human Feedback (RLHF), LoRa, Supervised Learning, Open-source LLMs, OpenAI GPT-4 API, Text to Image, Graph Databases, GitHub, Statistics, Data Scraping, Statistical Data Analysis, Statistical Methods, AI Compliance Agents, Audio Processing, Audio, Digital Signal Processing, Real-time Audio Processing, Object-oriented Programming (OOP), Sound, Pure Data, Max/MSP/Jitter, Fraud Audits, Fraud Detection, Fraud Prevention, Python 3, AI Prompts, LlamaIndex, Cloud Run, PySpark, Solution Architecture, AI-generated Code, Azure, Databricks, Front-end Development, Search Engines, User Interface (UI), User Experience (UX), GitOps, DevOps, Predictive Modeling, Model Context Protocol (MCP), Ontologies, Data Pipelines, Microsoft Excel, Vertex AI, Google Cloud, Mathematical Modeling, User Behavioral Analytics (UBA), Sentiment Analysis, Pattern Analysis, Data, Mathematics, Java, AI Integration, REST APIs, Applied Statistics, Actuarial, R, Capacitor, Geofencing, Geofencing & Geotargeting, Geofencing API, Geographic Information Systems, Google Maps, Google Maps API, Google Maps SDK, Maps, Web GIS, AI Voice Agents, ElevenLabs Solutions, React, Minimum Viable Product (MVP), Full-stack, Retrieval-augmented Generation (RAG), VAPI, Vapi, Data Mining, Research, Leadership, Executive Consulting, Business Analysis, Functional Analysis, Microsoft Copilot, GeoJSON
The most amazing...
...thing I've done is pioneer a DARPA-funded project on satellite image forensics that revolutionized detection accuracy, fueling my passion for tech innovation.
Work Experience
AI and LLM/SLM Expert | Team Lead
Circana - Main
- Led LLM and SLM architecture strategy for AI-driven enterprise apps, using latency, accuracy, and cost benchmarks to guide model selection.
- Diagnosed fine-tuned model failures across specific business data segments and defined retraining, validation, and error-analysis workflows.
- Optimized fine-tuning pipelines for smaller language models, including dataset prep, training runs, evaluation metrics, and automation steps.
- Designed intent-parsing workflows that converted natural-language questions into standardized back-end query instructions instead of free-form answers.
- Developed evaluation sets for top 10 rankings, brand comparisons, and sales-analysis prompts to measure intent accuracy and output consistency.
- Guided local hardware SLM deployment planning to reduce cloud dependency while preserving business data privacy and acceptable response speed.
- Mentored engineering teams on fine-tuning best practices, model diagnostics, prompt design, and production-readiness for enterprise AI systems.
- Evaluated Azure AI Foundry and alternative model training options to support scalable experimentation, deployment, and governance decisions.
RAG Engineer
HighRes Biosolutions, Inc.
- Developed multi-RAG pipelines to extract, chunk, and index technical content from Confluence, SharePoint, and other enterprise data sources.
- Optimized semantic chunking and metadata extraction workflows to improve retrieval quality across complex biotech and lab automation documents.
- Designed data models and ontology structures to organize technical knowledge for production-ready RAG and agentic LLM workflows.
- Evaluated vector database and graph database approaches to determine the best architecture for structured, semi-structured, and semantic retrieval.
- Refined an advanced proof of concept into a production-grade RAG application with improved pipelines, prompts, and retrieval reliability.
- Integrated agentic LLM workflows with data pipelines to support automated reasoning, document classification, and technical knowledge retrieval.
- Implemented prompt engineering strategies to improve answer accuracy, source grounding, and user interaction quality in a co-pilot application.
- Documented architecture decisions, pipeline risks, and production-readiness requirements to support long-term deployment and future scaling.
Senior ML Engineer
Zero Electric Inc
- Developed computer vision models for electricity-grid mapping, transforming geospatial and image data into features for MVP insights and weekly releases.
- Fine-tuned time-series forecasting models to predict grid load trends using existing datasets and support stakeholder decision-making.
- Engineered Databricks pipelines for data preparation, ML training, model evaluation, and repeatable experimentation across grid datasets.
- Built FastAPI-ready ML components to integrate forecasting and computer vision outputs into a production-oriented mapping product.
- Supported weekly and bi-weekly feature delivery by collaborating closely with a 2-person core team on model design, testing, and deployment priorities.
- Designed geospatial ML workflows compatible with Kepler GL visualizations to help users interpret electricity-grid patterns and anomalies.
- Evaluated model performance using data-driven metrics for accuracy, forecast reliability, and readiness for MVP-stage product deployment.
- Documented model assumptions, data requirements, training workflows, and technical risks to support a 3-month engagement and future product scaling.
Senior AI Engineer/Consultant
Vix Media Group LLC
- Reviewed the existing AI-agent architecture and identified scalability, API-integration, and multi-LLM gaps for a 2–4 week development roadmap.
- Designed a modular agent deployment framework enabling users to configure agents for business, crypto, research, and social media workflows.
- Led technical scoping for Python, LangChain, LLM, and API-based agent features across a 5–40 hour per week remote consulting engagement.
- Developed customizable agent behavior specifications, including prompt workflows, tool access, memory handling, and user-controlled configuration options.
- Integrated multi-LLM architecture planning to support flexible model selection, external APIs, and domain-specific agent deployment.
- Guided international development teams on implementation priorities, architecture corrections, and feature delivery for an AI-agent platform.
- Evaluated competitor capabilities and market requirements to inform technical strategy, platform positioning, and accelerated development priorities.
- Documented architecture risks, implementation gaps, and recommended next steps to support rapid delivery within a 2–4 week project timeline.
AI Engineer
SportsMedAnalytics, LLC
- Designed an AI-driven discovery workflow to sync iPhone 4K and 1080p video with external USB microphone audio while preserving source resolution and export quality.
- Prototyped automated video and audio synchronization methods to align high-quality external audio with recorded footage for a streamlined editing pipeline.
- Evaluated AI-based techniques for detecting script timing cues and overlaying relevant Twitter/X thread screenshots at the right moments.
- Developed a proof-of-concept workflow for trimming predefined intro, outro, and content gaps to reduce manual post-production effort.
- Benchmarked output quality requirements to ensure edited videos maintained 4K and 1080p resolution and minimized compression or quality loss.
- Designed an independent audio-export capability so finalized voice tracks could be delivered separately from the edited video output.
- Documented technical feasibility, workflow risks, and development requirements to support a 2–4 week discovery phase and future full-scale build.
Senior Research Engineer
Dolby Laboratories
- Developed a space- and time-efficient denoising and super-resolution method that significantly reduced processing time while enhancing image clarity.
- Engineered a TPB-based compression method that lowered storage requirements while maintaining high video quality.
- Pioneered a new 360 video codec that improved streaming efficiency and decreased latency in real-time applications.
- Spearheaded floor plan construction for multiple perspective videos using object-based latent vector aggregation, enhancing reconstruction accuracy and performance.
- Implemented advanced deep learning models for time-series forecasting (RNN, LSTM, GRU, CNN, and Transformer-based models), achieving improved accuracy through metrics like MAPE, RMSE, MAE, SMAPE, R², and log loss.
- Gained insights into EV charging trends through industry research, analyzing factors like time of day, weather, and location while identifying grid management challenges.
Experience
PhD Thesis
Working under the guidance of Professor Edward J. Delp in the video and image processing (VIPER) laboratory, I created advanced detection algorithms that included a fusion-based method for forensic splicing localization and a data-driven approach for panchromatic imagery copy-paste localization. By integrating state-of-the-art techniques such as vision transformers, deep belief networks, and nested attention U-Nets, I enhanced manipulation detection capabilities and set new benchmarks in digital forensics research. My work has been featured in prominent conferences, including SI22 SPIE Defense + Commercial Sensing, CVPRW, and the International Conference on Acoustics, Speech, and Signal Processing, highlighting its impact on advancing the field.
Education
PhD in Electrical and Computer Engineering
Purdue University - West Lafayette, IN, USA
Skills
Libraries/APIs
PyTorch, TensorFlow, Keras, Matplotlib, NumPy, OpenAI API, Hugging Face Transformers, OpenCV, Pandas, Scikit-learn, LSTM, WebRTC, Google Speech API, Google Vision API, Dask, REST APIs, Geofencing API, Google Maps, Google Maps API, Google Maps SDK, PySpark, Azure Computer Vision API, React, Kepler.gl
Tools
ChatGPT, Mathematica, You Only Look Once (YOLO), Algorithm Design, Whisper, Git, Terminal, GitHub, AI Prompts, Microsoft Excel, Microsoft Copilot, AutoML, Amazon Transcribe, n8n, DJI SDK, Claude, GraphRAG, Web GIS, DeepSeek, Tekton, GIS, Capacitor
Languages
Python, C, Bash, Bash Script, TypeScript, Python 3, C++, SQL, JavaScript, Java, Max/MSP/Jitter, R
Frameworks
LlamaIndex, Agentic Frameworks, LangGraph
Paradigms
Business Intelligence (BI), Synthetic Data Generation, Object-oriented Programming (OOP), Functional Analysis, Automation, Mobile Development, Mobile Design, Agile Software Development, ETL, DevOps, Model Context Protocol (MCP), User Behavioral Analytics (UBA)
Platforms
Linux, Kubernetes, Jupyter Notebook, NVIDIA CUDA, Docker, Windows, LiveKit, Amazon Web Services (AWS), AWS Lambda, AWS Cloud Computing Services, iOS, Google Cloud Platform (GCP), Kubeflow, Azure, Databricks, Vertex AI, Cloud Run
Storage
Data Integration, Data Pipelines, Amazon S3 (AWS S3), Neo4j, Graph Databases, Google Cloud
Industry Expertise
Applied Statistics, Formulation, Healthcare
Other
AI Research, Computer Vision, Image Processing, Machine Learning, Dolby Vision, Video Coding, API Integration, Data Engineering, Natural Language Processing (NLP), Artificial Intelligence (AI), Data Classification, Data Science, Data Analytics, Generative Artificial Intelligence (GenAI), Large Language Models (LLMs), Hugging Face, Speech Recognition, Architecture, AI Model Training, Diffusion Models, Image Segmentation, Deep Learning, Convolutional Neural Networks (CNNs), Forecasting, MAPE, RMSE, LangChain, OpenAI, OpenAI GPT-3 API, AI Agents, Reinforcement Learning, Transformers, Quantization, Pipedrive, Prompt Engineering, Geospatial Analytics, AI Programming, Geospatial Data, Automatic Speech Recognition (ASR), BERT, Optical Character Recognition (OCR), Video & Audio Processing, Technical Analysis, Facial Recognition, Text-to-Speech (TTS), Video Analysis, Machine Learning Operations (MLOps), Generative Pre-trained Transformers (GPT), Audio Analysis, Video Transcoding, Data Analysis, Data Build Tool (dbt), Demand Forecasting, Data Visualization, Neural Networks, Technical Leadership, Algorithms, AI Data Classification, Data Processing, Agentic AI, Machine Learning Algorithms, AI Chatbots, Interactive Voice Response (IVR), Software Architecture, AI Model Integration, Image Generation, Conversational AI, Retrieval-augmented Generation (RAG), Speech-to-Text (STT), Multimodal Models, Real-time Data, Time Series Analysis, Time Series Forecasting, Image Classification, Xarray, SSH, Drones, YOLOv8, Fine-tuning, GitHub Actions, LSTM Networks, Correlational Analysis, Feature Engineering, Data Scientist, Statistical Analysis, Statistical Modeling, Data Extraction, Reinforcement Learning from Human Feedback (RLHF), LoRa, Supervised Learning, Unsupervised Learning, Open-source LLMs, OpenAI GPT-4 API, CI/CD Pipelines, Statistics, Data Scraping, Statistical Data Analysis, Statistical Methods, Audio, Digital Signal Processing, Sound, AI-generated Code, Front-end Development, GitOps, Predictive Modeling, Pattern Analysis, Game Analytics, Data, Mathematics, AI Integration, Data Architecture, PDF Scraping, ElevenLabs Solutions, Minimum Viable Product (MVP), Data Mining, Research, Leadership, Business Analysis, GeoJSON, Scraping, Audio Processing, APIs, LLM Integration, Windows UI Automation, Multithreading, Recurrent Neural Networks (RNNs), Pinecone, Vector Databases, Object Detection, Chatbot Conversation Design, Chatbots, Video Transformers, Video Editing, Large Language Model Operations (LLMOps), Detectron2, Exploratory Data Analysis, AI Modeling, Cloud, FastAPI, Graphics, Stable Diffusion, Fashion, Electronic Health Records (EHR), Back-end, Web Development, Multi GPU training, Deep Reinforcement Learning, Text Generation Inference, NVIDIA TensorRT, Supervised Fine-tuning (SFT), LLM-as-a-judge, LLM Evaluation BLEU - ROUGE, LM Evaluation Harness, Bayesian Inference & Modeling, AWS SSH Keys, Edge Computing, Google Cloud Functions, Health, Medical Software, Funnel Analysis, Churn Analysis, Voice Chat, Anthropic, Roboflow, Small Language Models (SLMs), AI Content Creation, Text to Image, AI Compliance Agents, Real-time Audio Processing, Fraud Audits, Fraud Detection, Fraud Prevention, Search Engines, User Interface (UI), User Experience (UX), Real World Data, Healthcare Data Science, Biostatistics, Ontologies, Mathematical Modeling, Sentiment Analysis, Time Series, Geofencing, Geofencing & Geotargeting, Geographic Information Systems, Maps, AI Voice Agents, Full-stack, Vapi, Executive Consulting, Web Scraping, IPC (Inter-Process Communication), Memory Optimization, Project Scoping, Materials Science, Manufacturing, LLM inference, Memory Management, Pure Data, Solution Architecture, QA Automation, Actuarial, Infrastructure as Code (IaC)
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring