Caio Theodoro, Developer in Curitiba - State of Paraná, Brazil
Caio is available for hire
Hire Caio

Caio Theodoro

Software Engineer and Developer

Curitiba - State of Paraná, Brazil

Toptal member since August 19, 2026

Bio

For the past seven years, Caio has built agent infrastructure and production LLM systems. He has worked extensively across fintech, education, and supply-chain industries. His work centers on Python, TypeScript, MLops, and AWS for clients including Adopt AI, Avenza, and Acorns.

Portfolio

Adopt AI
Temporal, Agent Orchestration, AI Engineering, AI Architecture, AI Guardrails...
Avenza
Agent Orchestration, AI Engineering, AI Architecture, Amazon Web Services (AWS)...
Acorns
Rails, Node.js, React, Next.js, GraphQL, Software Engineering, Microservices...

Experience

  • Microservices - 7 years
  • TypeScript - 7 years
  • Software Engineering - 7 years
  • AI Architecture - 3 years
  • Retrieval-augmented Generation (RAG) - 3 years
  • LLM Evaluation - 3 years
  • AI Guardrails - 2 years
  • AI Engineering - 2 years

Preferred Environment

Docker, Amazon Web Services (AWS), GCP, Azure, Kubernetes, Terraform

The most amazing...

...thing I've built is Harness, including most of its components, which serves big companies like Automation Anywhere, UHY, and Beyond Risk.

Work Experience

Senior AI Engineer

2026 - PRESENT
Adopt AI
  • Architected the core agent execution harness, implementing turn and message lifecycles, idempotent claim semantics, and replay-safe Temporal orchestration to support all agent capabilities on the platform.
  • Built the live agent evaluation dashboard that scores conversation quality in production, surfacing degradation early enough for the team to intervene before customers feel it.
  • Shipped the human-in-the-loop intervention system end to end, including a signal-driven resolve state machine with optimistic locking and tenant isolation, a 3-action resolution model, and a Temporal signal contract for escalations.
  • Designed the generative UI architecture for tax and accounting workflows using a registry of approved components, reducing dashboard prompts and saving 90% on multi-turn follow-up edits with a JSON-Patch delta.
  • Built a policy engine for agent guardrails with policy-version pinning and full audit logging, and integrated AWS Bedrock for multi-model inference behind one selection interface, cutting latency and cost through provider routing and prompt caching.
Technologies: Temporal, Agent Orchestration, AI Engineering, AI Architecture, AI Guardrails, Amazon S3 (AWS S3), Software Engineering, Machine Learning, Retrieval-augmented Generation (RAG), Microservices, TypeScript, LLM Evaluation, LLM Workflows, Multi-Agent Systems, Tool Calling, Model Context Protocol (MCP), LangGraph, Prompt Engineering, Context Engineering, Observability, CI/CD Pipelines, REST, Temporal Workflows, NumPy, Pandas, TensorFlow, PyTorch, Machine Learning Operations (MLOps), Inference Optimization, Reinforcement Learning from Human Feedback (RLHF), LLM-as-Judge, Embeddings & Semantic Search, Vector Search, LangChain, Responsible AI, Explainability, Calibration, Quantization, OpenAI, Google Gemini, Hugging Face, Natural Language Processing (NLP)

Founder and Principal Engineer

2025 - PRESENT
Avenza
  • Directed a 7-engineer team across client discovery, solution architecture, delivery, and engineering standards, delivering systems tied to roughly $2.8 million in client revenue over eight months.
  • Architected production AI agents across logistics, customer support, operational workflows, and computer vision using structured outputs, retrieval, model-backed classification, and human review loops.
  • Delivered end-to-end software and automation for a foreign-exchange house covering operational workflows, back-office processing, and third-party system integrations.
Technologies: Agent Orchestration, AI Engineering, AI Architecture, Amazon Web Services (AWS), AI Guardrails, Python, TypeScript, Software Engineering, Machine Learning, Retrieval-augmented Generation (RAG), Microservices, LLM Evaluation, LLM Workflows, Multi-Agent Systems, Tool Calling, Model Context Protocol (MCP), LangGraph, Prompt Engineering, Context Engineering, Observability, CI/CD Pipelines, REST, Temporal Workflows, Go, NumPy, Pandas, TensorFlow, PyTorch, Machine Learning Operations (MLOps), Inference Optimization, Reinforcement Learning from Human Feedback (RLHF), LLM-as-Judge, Embeddings & Semantic Search, Vector Search, LangChain, Responsible AI, Calibration, Natural Language Processing (NLP)

Senior Software Engineer

2025 - 2026
Acorns
  • Shipped production changes across the full stack (Rails, Node services, React, Next.js, GraphQL APIs) throughout a phased monolith-to-microservices migration.
  • Built GraphQL-driven admin and support tooling that gave internal operations direct visibility into servicing workflows and resolved schema and data-consistency failures in the gateway introduced by splitting Rails-owned data across newly indep.
  • Shipped ephemeral per-PR environments to production, closing the staging/production parity gap that had been a recurring source of regressions.
Technologies: Rails, Node.js, React, Next.js, GraphQL, Software Engineering, Microservices, TypeScript, Tool Calling, Model Context Protocol (MCP), LangGraph, Prompt Engineering, Observability, CI/CD Pipelines, REST, Kafka, Go, NumPy, Pandas, TensorFlow, PyTorch, Vector Search

Software Engineer

2022 - 2025
MB Labs
  • Built production systems on AWS (ECS, RDS, ElastiCache, S3, CloudFront, ALB) with Cloudflare and Redis for low-latency fintech workloads serving millions of users.
  • Designed microservices and APIs for high-traffic fintech clients; connection pooling and regional caching cut latency while right-sizing compute brought infrastructure spend down.
  • Architected event-driven data pipelines (SQS, async workers) for compliance-grade fintech operations.
Technologies: Amazon Web Services (AWS), Cloudflare, Redis, Amazon Simple Queue Service (SQS), Software Engineering, Microservices, TypeScript, Prompt Engineering, Observability, CI/CD Pipelines, REST, Kafka, Rust, Go, NumPy, Pandas, TensorFlow, Machine Learning Operations (MLOps), Recommendation Systems

Software Engineer

2021 - 2022
Mactus Informática LTDA
  • Modernized legacy business applications and improved stability for critical operational workflows.
  • Designed cloud-ready services with clearer performance, maintenance, and cost boundaries.
  • Led new products discovery and implementation inside the company.
Technologies: JavaScript, TypeScript, Python, Software Engineering, Observability, CI/CD Pipelines, REST, Recommendation Systems

Software Engineer Intern

2020 - 2021
Atla
  • Built Python scrapers processing 5,000+ records per day and REST APIs for internal education data workflows.
  • Built backoffice softwares that empowered the core main product.
  • Implemented Agile methods inside the squad to improve efficiency.
Technologies: Python, Software Engineering, TypeScript, Scraping, REST, Recommendation Systems

Software Engineer

2019 - 2021
Haken
  • Delivered web and mobile software for early clients while owning project scope, delivery plans, and client communication.
  • Established effective communication channels between technical teams and stakeholders.
  • Implemented RESTful APIs to streamline data processing workflows.
Technologies: JavaScript, TypeScript, Python, Software Engineering, REST

Experience

ReconForge

https://github.com/caiotheodoro/reconforge
I fine-tuned a 1.7-billion-parameter model with LoRA on a laptop to detect financial reconciliation exceptions. I optimized for severity-weighted recall, the metric tied to dollar risk, reaching 0.913 versus 0.872 for a frontier model, with 100% recall on high-severity exceptions and zero API cost. I built the event pipeline and reconciliation state on Kafka, Temporal, Postgres, and Neo4j, so exceptions are caught, orchestrated, and traced end-to-end without relying on a hosted LLM API.

LossBench

https://github.com/caiotheodoro/lossbench
I built an evaluation and control plane that scores AI agents on expected operational loss instead of raw accuracy: severity-weighted, calibrated, and replayable. Every decision runs through a record-calibrate-decide-escalate-replay pipeline backed by a hash-chained decision ledger, so outcomes can be audited and reproduced after the fact. Ships with LangGraph and DeepSeek Harness adapters, making it drop-in for agent stacks that already use those frameworks and need loss-aware evaluation rather than accuracy alone.

Calibrated Evaluation Foundry

https://github.com/caiotheodoro/substrate
I reproduced ARC-AGI-3's benchmark methodology as a repeatable pipeline, then used it to stress-test its own parameters against real solvers and a real LLM judge rather than trusting the published methodology at face value. The goal was to find where the benchmark's assumptions break down: which parameter choices meaningfully change solver rankings, and where an LLM judge disagrees with or diverges from solver-based scoring. It was built as a foundry so the calibration pipeline can be rerun against new solvers or judges as they change.

Education

2018 - 2025

Bachelor's Degree in Computer Science

Technological Federal University of Paraná - Brazil

Certifications

SEPTEMBER 2024 - PRESENT

Cloud Technical Essentials

AWS

JANUARY 2024 - PRESENT

DevOps and Software Engineering

IBM

JANUARY 2024 - PRESENT

Machine Learning

IBM

JANUARY 2024 - PRESENT

AWS Fundamentals

AWS

JANUARY 2024 - PRESENT

Cloud Solutions Architect

AWS

AUGUST 2023 - PRESENT

Full-stack Software Developer

IBM

AUGUST 2023 - PRESENT

Data Analytics

Google

Skills

Libraries/APIs

PyTorch, TensorFlow, Pandas, NumPy, Node.js, React, OpenCV

Tools

GraphRAG, Amazon Simple Queue Service (SQS), Terraform

Languages

Python, TypeScript, JavaScript, SQL, GraphQL, Go, Rust

Paradigms

Microservices, REST, Automation, Model Context Protocol (MCP)

Platforms

Harness, Docker, Apache Kafka, Cloud Native, Amazon Web Services (AWS), Azure, Kubernetes

Frameworks

LangGraph, Next.js

Storage

PostgreSQL, Neo4j, Amazon S3 (AWS S3), Redis

Other

AI Engineering, AI Architecture, Machine Learning, LLM Workflows, Agent Orchestration, Structured Outputs, AI Guardrails, Retrieval-augmented Generation (RAG), Embeddings & Semantic Search, Vector Search, Inference Optimization, Temporal Workflows, Multi-Agent Systems, Tool Calling, LangChain, Prompt Engineering, Machine Learning Operations (MLOps), OpenAI, Computer Vision, Natural Language Processing (NLP), Kafka, CI/CD Pipelines, Software Engineering, Scraping, LoRa, LLM Evaluation, Data Cleansing, Cloud Computing, Temporal, Rails, Cloudflare, GCP, Context Engineering, Responsible AI, Explainability, LLM-as-Judge, Calibration, Reinforcement Learning from Human Feedback (RLHF), Quantization, Google Gemini, Hugging Face, Predictive Analytics, Recommendation Systems, Observability, Infrastructure as a Service (IaaS)

Collaboration That Works

How to Work with Toptal

Toptal matches you directly with global industry experts from our network in hours—not weeks or months.

1

Share your needs

Discuss your requirements and refine your scope in a call with a Toptal domain expert.
2

Choose your talent

Get a short list of expertly matched talent within 24 hours to review, interview, and choose from.
3

Start your risk-free talent trial

Work with your chosen talent on a trial basis for up to two weeks. Pay only if you decide to hire them.

Top talent is in high demand.

Start hiring