
Jose Tomas Lavin Leon
Verified Expert in Engineering
Data Scientist, Machine Learning Engineer, and Developer
Santiago, Chile
Toptal member since July 29, 2026
José is a data scientist and machine learning engineer with over 5 years of experience spanning data engineering, credit risk modeling, and end-to-end MLOps. He has built data pipelines and warehouses on AWS and GCP that feed the ML infrastructure running on top of them, primarily for the financial sector. José reduced portfolio risk by 25% while at Aplazame.
Portfolio
Experience
- Machine Learning - 5 years
- Docker - 5 years
- Apache Airflow - 5 years
- SQL - 5 years
- Python - 5 years
- AWS IoT - 5 years
- MLflow - 3 years
- Scikit-learn - 3 years
Preferred Environment
AWS IoT, Google Cloud Platform (GCP), GitHub Actions, Docker, Tableau, Apache Airflow, Amazon CloudWatch, Sentry, BigQuery, Amazon S3 (AWS S3)
The most amazing...
...ML observability system I've architected integrated CloudWatch, Airflow, and Dash to monitor thousands of daily scoring requests.
Work Experience
Data Scientist | ML Engineer
Aplazame
- Led the end-to-end MLOps lifecycle for production credit scoring models (XGBoost, LightGBM, CatBoost), including experiment tracking with MLflow, CI/CD via GitHub Actions, Docker, and AWS ECS to deploy the serving layer with FastAPI.
- Led the development and architected a 3-layer ML observability system, leveraging Airflow, Kubernetes, and MSF Teams for data-quality checks, feature drift, and model-performance monitoring.
- Built an internal app (Dash) for a stakeholder dashboard with model explainability (SHAP and segment drill-down).
Consultant
Panel Ciudadano
- Designed and built a data warehouse in GCP to enable customers to query their own longitudinal survey data through an AI assistant.
- Built a GenAI-powered tool enabling customers to query their own data through MCP and a semantic layer over BigQuery.
- Used dbt for data modeling and deployed a FastAPI MCP server on Cloud Run with JWT-scoped tenant isolation.
Data Engineer
Aplazame
- Co-designed and built Aplazame's first data warehouse from scratch using a medallion architecture on S3, with ETL pipelines in AWS Glue and Airflow.
- Established the foundational data infrastructure later used to train and serve all credit-scoring models.
- Designed and maintained Airflow DAG orchestration for ingestion, transformation, and quality-assurance pipelines across heterogeneous sources (S3, SFTP, DocumentDB, ElasticSearch), guaranteeing data reliability for downstream ML model development.
- Developed a visualization layer for stakeholders in Tableau.
Experience
ML Observability & Drift Monitoring System
Before this, there was no automated way to catch model degradation. Issues surfaced only when someone noticed a problem downstream. Now, data quality issues are flagged within minutes, model performance is reviewed monthly, and SHAP explanations are ready for any model version, which has also sped up internal audits. It's become the reference monitoring architecture that Aplazame expects new models to ship with.
Generative BI: AI Agent Access to Business Data via MCP
Panel Ciudadano is both the client and the primary user its internal team needed a way to get answers from survey data without writing SQL or routing every question through a data request queue. I made the core architecture decisions myself, from the pipeline and semantic layer design to the structure of the MCP layer. The platform is currently in its improvement phase, with the goal of letting anyone at the company ask a question like "what's our completion rate for survey X" in natural language and receive a consistent, accurate answer instead of waiting on a ticket.
Credit Scoring Models: Portfolio Risk Reduction
The client is Aplazame itself, and the end users are the risk team, who rely on these models to approve or decline transactions in real time, and, indirectly, Aplazame's retail customers, whose checkout experience depends on fast, accurate decisions.
Across the models I contributed to, the combined impact was a 25% annual reduction in portfolio risk while maintaining the existing acceptance rate, meaning fewer bad loans approved without declining more good customers, alongside models that are simpler to maintain and retrain than the legacy monolithic approach.
Education
Master's Degree in Data Science
Universidad De Navarra - Madrid, Spain
Bachelor's Degree in Commercial Engineering
Universidad Católica De Chile - Santiago, Chile
Skills
Libraries/APIs
XGBoost, CatBoost, Scikit-learn
Tools
Apache Airflow, AWS Glue, Sentry, Tableau, BigQuery, Amazon CloudWatch, MATLAB, Claude
Languages
Python, SQL, R
Platforms
Docker, AWS IoT, Google Cloud Platform (GCP), Cloud Run, Kubernetes
Storage
Amazon S3 (AWS S3), Elasticsearch
Frameworks
LightGBM, Optuna, FastMCP
Other
MLflow, Machine Learning, Economics, Econometrics, GitHub Actions, FastAPI, Dash, SHAP, Data Build Tool (dbt), Data Analysis, Data Engineering, SFTP, DocumentDB, Deep Learning, ECS
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring