
Ingyu Koh
Verified Expert in Engineering
C++ Developer
Donghae-si, Gangwon-do, South Korea
Toptal member since September 4, 2026
Ingyu is a quantitative researcher and systems builder with 30+ years of experience in computational and statistical research, including a decade in low-latency trading and scientific AI. He analyzes problems before proposing solutions and is clear about what evidence can support. His expertise spans market microstructure, GPU backtesting, and deep learning for scientific and medical imaging.
Portfolio
Experience
- C++ - 20 years
- Computational Physics - 19 years
- Python - 14 years
- Fintech - 12 years
- Facial Recognition - 10 years
- Git - 10 years
- Quantitative Research - 8 years
- PyTorch - 5 years
Preferred Environment
Python, PyTorch, PyTorch Geometric (PyG), Git, Linux, CUDA Kernel, C++, Machine Learning, Feature Engineering, Data Science, Model Development, Model Evaluation
The most amazing...
...research system I've built is a GPU-accelerated backtesting platform that enabled large-scale model training for high-frequency trading.
Work Experience
Independent Applied AI Researcher — ML, API Security and AWS
Self-employed
- Implemented an October 2026 bounded SDLC lab with LangGraph and Guardrails AI, using exact patch digests, 120-second human approvals, and atomic single-use claims. Tested forbidden tools and changed, expired, and replayed approvals.
- Corrected the Business Decision Context Lab’s bypassable keyword check with code-owned output; rejected 30 authored invalid probes and accepted 10 valid controls. These scoped regression cases are separate from the SDLC lab.
- Deployed the public SDLC lab on AWS Lambda and DynamoDB with scoped IAM, minimized logs, and request limits. Passed 61 container tests and 23 live HTTP assertions; restored a waiting review after worker replacement and rejected replay.
Independent Microsoft Fabric Analytics Developer
Self-employed
- Built an independent Microsoft Fabric prototype that ingested and transformed data, stored it in OneLake/Lakehouse, and prepared it for analysis.
- Queried and modeled the transformed data for analytics, then exposed the results for downstream reporting and analysis through Power BI.
- Orchestrated pipeline and notebook steps and checked that ingestion, storage, analytics, and Power BI reporting connected into a usable end-to-end flow.
Independent Agentic AI and Graph ML Developer
Self-employed
- Built the graph neural network and data pipeline for the personal graph neural network project, keeping model training separate from agent execution experiments.
- Integrated an agentic workflow with Amazon Bedrock AgentCore Runtime to test deployment and execution around the graph/ML system.
- Tested managed agent infrastructure for controlled tool access, session handling, and operational tracing, and inspected workflow failures during hands-on experiments.
Independent AI Researcher
Self-employed
- Built and evaluated public LangGraph agents for recruiting and SEC filing retrieval using MCP tools, LangChain adapters, and an optional OpenAI API provider; published code and offline evaluations on GitHub.
- Built and evaluated deep-learning segmentation models on scientific and medical imagery, reaching Dice coefficients above 0.85 on held-out data.
- Applied computational physics and statistical inference methods to modern deep-learning problems, bridging first-principles modeling with data.
- Built 2026 Federal Spending Graph reports and dashboards across agencies, sub-agencies, recipients, programs, geographies, and obligations, showing comparative models, risk, and validation metrics.
- Normalized agencies, sub-agencies, recipients, programs, NAICS, and PSC classifications into consistent graph entities for project-level master-data analysis; did not administer an enterprise MDM platform.
- Designed 2026 Sales Forecast and AMAT QA procedures using automated tests, leakage and train/validation checks, baseline comparisons, reproducibility checks, failure-case review, and acceptance criteria.
- Reviewed Federal Spending, Sales Forecast, and AMAT outputs for data quality and metric correctness; rejected unsupported conclusions and wrote reports on methods, results, risks, and next actions.
Freelance AI Researcher
Freelance Clients
- Developed a Python pharmaceutical research prototype that organized biomedical entities and drug–target–disease relationships from public literature into a knowledge graph.
- Implemented graph-based retrieval to give LLM answers relevant to the relationship context and source evidence, supporting traceable responses to pharmaceutical research questions.
- Evaluated retrieval relevance, citation accuracy, and unsupported claims to measure answer quality and identify failures requiring further review.
- Managed three team members while working as a freelance AI researcher on the pharmaceutical knowledge graph and LLM research prototype.
Quantitative Research and Low-latency Systems Engineer
Freelance Clients
- Researched short-horizon signals from tick and limit-order-book data, including spread behavior, quote and depth imbalance, and short-horizon price formation.
- Built and optimized research and backtesting infrastructure for large tick-data histories, turning market hypotheses into repeatable experiments.
- Worked on execution systems for exchange-colocated environments where deterministic microsecond-scale latency was central to system design.
- Applied GPU acceleration to increase throughput for backtesting and model training across large-scale market data.
- Improved research capacity and execution-system reliability through kernel bypass, cache-aware critical paths, lock-free data structures, and jitter control.
Independent Biomedical Research Developer — Peptide Evidence Lab
Self-employed
- Built an independent Python peptide-literature demo in October 2026 with AI coding assistance, using six public studies, Europe PMC ingestion, BM25 retrieval, and source-linked evidence checks.
- Reproduced a simple keyword classifier’s error on mouse-islet study PMID 25830090; used reviewed species and population labels to exclude it from human-only searches while retaining mechanistic evidence.
- Deployed the public AWS Lambda demo with decision traces, source hashes, and CSV export; passed 17 regression tests. Kept this six-study illustration distinct from a live LLM or scheduled monitoring pipeline.
Senior Researcher
Naver
- Applied deep-learning models to production recognition tasks, balancing accuracy against inference cost under real serving constraints.
- Contributed to production pattern-recognition systems serving large-scale consumer traffic, focusing on model accuracy and inference efficiency.
- Evaluated recognition models against production data distributions, identifying failure modes that differed from benchmark performance.
Facial-recognition Research and Development Expert
Confidential Israeli Computer-vision Startup
- Applied computer-vision and pattern-recognition methods to facial identity matching in a startup environment.
- Developed facial identity-matching methods and evaluated them against held-out benchmarks, tuning the trade-off between false accepts and false rejects.
- Adapted computer-vision pipelines to real-world capture conditions, addressing variation in pose, illumination, and image quality.
Full Professor, Department of Physics
Korea Advanced Institute of Science and Technology
- Led research in grand unified theories, monopoles and dyons, Kaluza-Klein supergravity, string theory, conformal field theory, affine Toda theory, stochastic processes, and quantum cryptography.
- Built a 94-record international publication portfolio and maintained collaborations through visiting appointments in Europe, North America, and Japan.
- Supervised graduate students through to doctoral completion, training researchers who went on to academic and industrial positions in Korea and abroad.
- Taught graduate and undergraduate physics for 28 years, developing courses spanning quantum field theory, statistical mechanics, and mathematical methods.
Founder
KOTECH SYSTEM
- Founded an applied pattern-recognition firm and worked on deployable systems for passport OCR, vehicle license-plate recognition, online Korean and Chinese handwriting recognition, and object tracking.
- Combined statistical modeling, image processing, and production software to translate recognition methods into operational systems.
- Led the firm for ten years, taking pattern-recognition research from prototype through to deployed products across document, vehicle, and handwriting recognition.
- Delivered recognition systems into operational use for identity documents, vehicle plates, and handwritten input, running on constrained production hardware.
Associate Professor, Department of Physics
Sogang University
- Conducted research in theoretical particle physics, publishing on gauge theories and their underlying mathematical structures in international journals.
- Taught undergraduate and graduate physics, covering classical mechanics, electromagnetism, quantum mechanics, and mathematical methods.
- Established international research collaborations early in an academic career, presenting results at conferences in Asia, Europe, and North America.
Experience
Forensic License Plate Recognition in High-speed Video
Independent Statistical and Mathematical Audit of Research Claims
The deliverable separated three things most reviews conflate: what the audit could confirm (mathematical consistency and correct calculation), what it could not (whether the model corresponds to reality), and where the methodology itself was the weak link. Explicitly stating that boundary is what makes an audit useful rather than merely reassuring. I delivered a written report that the client used to revise the work.
GPU-accelerated Backtesting Platform for High-frequency Trading Research
The platform enforced the discipline that matters more than speed: strict point-in-time data access so no future information can leak into a signal, realistic fill and latency assumptions, and rolling-origin evaluation instead of a single train/test split. Signals studied included spread behavior, quote and depth imbalance, and short-horizon price formation.
The result was a shorter path from question to evidence, and fewer results that looked good only because the test was wrong.
DIVEROID — iOS Development for an Underwater Smartphone System
https://www.diveroid.com/MCP Recruiting Agent — Guarded Multi-turn Assistant
https://github.com/ingyukoh/mcp-recruiting-agentSEC Filing RAG Auditor — Verified Financial Retrieval
https://github.com/ingyukoh/agentic-rag-evalEvidence-first Agentic RAG Evaluation Platform
https://github.com/ingyukoh/evidence-first-agentic-ragLegal Contract Review Lab — LoRA, Source Evidence and AWS
https://np6ibbw5l6tdgn77ihpgmbcj7u0majwu.lambda-url.us-east-1.on.aws/Kalshi Trade-Tape Quoting Lab — Fill Models, Inventory Replay, and AWS
https://pz5l45g4wb77gg7q6idoazwaeq0eqvxn.lambda-url.us-east-1.on.aws/Business Decision Context Lab — Predictive ML, Semantic Layer, and LLM
https://gvk2rzatfezmfj5rk6givmxgue0vsrto.lambda-url.us-east-1.on.aws/Approval-gated SDLC Lab — Guardrails AI, LangGraph, and AWS
https://kjigfrfabegyxerw5jagfvhyza0kveeb.lambda-url.us-east-1.on.aws/Education
PhD in Theoretical Physics
Korea Advanced Institute of Science and Technology - South Korea
Certifications
Financial Risk Manager (FRM)
Global Association of Risk Professionals (GARP)
Certified Public Accountant (CPA)
Delaware Board of Accountancy
Skills
Libraries/APIs
TensorFlow, XGBoost, Claude API, React, OpenAI API, PyTorch, PyTorch Geometric (PyG), API Development, REST APIs
Tools
Claude, AI Prompts, ChatGPT, Git, Microsoft Power BI, GraphRAG
Languages
Python, C++, SQL, Swift
Frameworks
LangGraph, Agentic Frameworks
Paradigms
Quantitative Research, Model Context Protocol (MCP), DevOps, ETL
Platforms
Linux, iOS, Amazon Web Services (AWS), Harness, Microsoft Fabric
Storage
PostgreSQL, Data Pipelines, Master Data Management (MDM)
Other
Optical Character Recognition (OCR), Statistical Modeling, Facial Recognition, Handwriting Recognition, Statistical Inference, Computational Physics, Fintech, Signal Analysis, Machine Learning, Time Series Forecasting, Feature Engineering, Data Science, Model Development, Model Evaluation, Data Engineering, Deep Learning, Artificial Intelligence (AI), Training Workshops, Large Language Models (LLMs), Predictive Modeling, Data Analysis, Generative Artificial Intelligence (GenAI), iPhone, Bug Fixes, User Interface (UI), Full-stack, Agentic AI, FastAPI, AI Agents, Answer Engine Optimization (AEO), ChatGPT Prompts, ChatGPT API, Prompt Engineering, Solution Architecture, Computer Vision, Image Processing, AI Model Training, AI Model Integration, Web Applications, HDR Photography, LangChain, LLM Agents, AI Workflow, Amazon Bedrock AgentCore, Statistics, System Architecture, Training, Quantitative Analysis, Financial Data, Object Tracking, Communication, Demand Forecasting, Monitoring, Employee Upskilling, Workshop Facilitation, Supabase, Document Processing, Startups, CUDA Kernel, Statistical Methods, GPU Computing, Final Accounts, Stakeholder Management, Combinatorial Optimization, Multi-agent Systems, Google+, BI Reports, Dashboards, Quality Assurance (QA), Quality Control (QC), Reports, Knowledge Graphs, Hugging Face, Fine-tuning, Natural Language Processing (NLP), Data Security, Legal Technology (Legaltech), Crypto, AI Evaluation, Models, Agentic AI Systems, AI Security, RAG Systems, AI Safety, Cloud Security, CI/CD Pipelines, Machine Learning Operations (MLOps), Retrieval-augmented Generation (RAG)
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring