
James Irie
Verified Expert in Engineering
Data Scientist, Machine Learning Engineer, and Developer
Austin, United States
Toptal member since August 6, 2026
James is a data scientist and machine learning engineer with 7+ years of experience building practical ML and analytics solutions from complex, multi-source datasets. He specializes in predictive modeling, anomaly detection, and production-oriented data pipelines using Python, SQL, and cloud-based workflows. James is particularly strong at turning ambiguous business problems into structured analyses, reliable models, and clear recommendations for technical and non-technical stakeholders.
Portfolio
Experience
- Python - 8 years
- Data Science - 8 years
- Machine Learning - 8 years
- Churn Analysis - 8 years
- Amazon Web Services (AWS) - 8 years
- SQL - 7 years
- model managements - 6 years
- model performance analysis - 6 years
Preferred Environment
Git, MLflow, Python, Pandas, Amazon Web Services (AWS), SQL, PostgreSQL
The most amazing...
...thing I'm proud of is building an early churn model that turned product usage data into clear retention signals for customer success teams.
Work Experience
Senior Data Scientist
AffiniPay
- Led the end-to-end churn analysis from an ambiguous business brief, independently defining the problem scope, identifying relevant data sources, and building the analytical foundation in an environment with no existing DS infrastructure.
- Translated churn findings into a new Customer Success communication channel for at-risk customers. Early outreach increased product usage among targeted customers by approximately 40% and achieved email response rates above 50%.
- Identified key drivers of early customer churn by synthesizing behavioral data with qualitative signals from customer transcripts and support logs; presented findings to VP and Director-level stakeholders in sales and customer success.
- Evaluated whether LLMs could reliably extract customer sentiment and product-feature interest from sales and onboarding transcripts. Designed a repeatable validation framework to quantify hallucination and consistency risk.
- Demonstrated that LLM classifications became unreliable when supporting evidence was absent and recommended a human-in-the-loop workflow using LLMs for targeted evidence retrieval.
- Conducted behavioral cohort analyses to identify customer segments with distinct product usage and feature-adoption patterns, establishing a segmentation framework to inform targeted product development, customer campaigns, and acquisition strategies.
- Established reproducible documentation and analytical standards as one of the company’s first data scientists, helping formalize team practices and creating reusable resources for future data scientists and cross-functional partners.
Machine Learning Engineer
Civitas Learning
- Developed institution-specific graduation prediction models estimating students’ likelihood of graduating within 2, 4, and 6-year windows across 30+ partner institutions, supporting long-range enrollment planning and financial forecasting.
- Helped rebuild production ML infrastructure from a Scala/Spark-based system to Python using Scikit-learn, Metaflow, and AWS Batch after identifying that many institution-level workloads did not require distributed compute.
- The new architecture aligned with the team’s Python expertise, reduced dependency on difficult-to-maintain Scala workflows, enabled workload-specific EC2 sizing, and lowered daily scoring runtime by approximately 10% and AWS compute costs by 6–8%.
- Led the retraining and production rollout of 30+ institution-specific models during the ML infrastructure migration, coordinating validation, configuration, and deployment with Data Quality Analysts within two months without production disruption.
- Designed data warehouse schemas, ingestion logic, and incremental update tables in a SQL-based ETL framework to integrate model features and predictions with upstream and downstream production pipelines.
- Partnered with software engineers, DE, product managers, and data quality analysts to translate modeling requirements into production workflows, resolve data and validation issues, and deliver ML capabilities through agile development processes.
- Maintained production model performance across institution-specific deployments by defining validation thresholds, parameter settings, and monitoring standards, and documenting procedures to support consistent model review and maintenance.
Data Scientist
Civitas Learning
- Applied propensity score matching to evaluate whether student outreach interventions improved retention, helping internal teams and partner institutions measure campaign effectiveness through platform reporting and custom impact analyses.
- Built a shared Python library for anomaly detection, bias analysis, and model diagnostics, standardizing production model and feature monitoring across the DS team and providing reusable components for automated performance and drift checks.
- Designed and built a QuickSight dashboard to automate model performance checks. Enabled the DS team to quickly distinguish raw data quality issues from genuine feature drift across production models and reducing investigation time.
- Responded to significant shifts in production data by diagnosing model and segment-level performance changes, adjusting model configurations and feature usage as needed, and validating updates before deployment to maintain reliable predictions.
Geophysicist
Marathon Oil (currently ConocoPhillips)
- Managed multiple projects related specifically to subsurface imaging to guide future oil exploration projects (both onshore and offshore).
- Mitigated risks of not hitting the target geological layer (Oil reserves) by identifying any uncertainty that exists within project areas, such as seismic faults and small layers of geological layers that may interfere with subsurface images.
- Helped design seismic testing at upcoming drilling sites to enable the company to acquire a subsurface image of sufficient quality in a cost-effective way.
Experience
Product Churn Analysis
The strongest predictor was prolonged login inactivity, which identified a high-risk segment but did not fully explain why customers churned. To make the results actionable, I combined model findings with exit survey review and manual case analysis of inactive churned users across firm sizes and practice areas.
Two consistent drivers emerged: a lack of effective product training and unresolved support issues prior to churn. Based on these findings, Sales and Customer Success launched an outreach campaign to about 50 recently acquired but inactive users, achieving a response rate above 50% and validating the need for improved onboarding, training, and support.
ML Infrastructure Rebuild
I built and maintained the training and scoring workflows for production models, including Redshift feature and scoring tables under dedicated data science schemas, using a strict SQL and YAML-based framework with GitHub review and historical prediction logging for model traceability. I also played a key role in retraining 30+ institution-specific models during the infrastructure migration, managing retrain scheduling, collaborating with the data quality team on performance validation, and overseeing configuration updates to ensure a smooth rollout within a two-month timeline.
Beyond the technical work, I partnered closely with data engineers, software engineers, and product managers in an agile environment and worked directly with the DQA team to ensure they could operate the new retraining process independently.
LLM Performance Evaluation
Automated ML Model Monitoring Dashboard
Education
Certification in Data Science
Galvanize - Austin, TX, USA
Master's Degree in Geophysics
University of Texas at El Paso - El Paso, TX, USA
Bachelor's Degree in Geophysics
Cornell University - Ithaca, NY, USA
Skills
Libraries/APIs
Scikit-learn, Pandas, NumPy, Matplotlib, XGBoost, SciPy, ArcGIS
Tools
AWS Batch, Amazon QuickSight, Seaborn, Amazon SageMaker, GitHub, dbt Cloud, Tableau, Git, Amazon Athena, Amazon CloudWatch, MATLAB, AWS Step Functions, Domo, AI Prompts
Languages
Python, SQL, Batch
Frameworks
Metaflow, Spark
Paradigms
Anomaly Detection
Platforms
Jupyter Notebook, Amazon Web Services (AWS), Docker, AWS Lambda
Storage
PostgreSQL, Redshift, Amazon S3 (AWS S3), Data Pipelines, MySQL
Other
Excel 365, Data Processing, Microsoft Office, Machine Learning, Statistics, Churn Analysis, model managements, ml-pipelines, Data Science, model depoyments, model performance analysis, Data Collection, Research, MLflow, Data Analysis, Statistical Learning, Data Products, model management, Model Maintenance, Random Forests, Logistic Regression, Product Analytics, Performance Analysis, Model Evaluation, Statistical Analysis, Model Monitoring, Data Scientist, Predictive Modeling, Signal Processing, Signal Analysis, subsurface imaging, Experimental Research, Workflow, Machine Learning Operations (MLOps), Data Management, Dashboards, Natural Language Processing (NLP), Step Functions, seismology, Data Build Tool (dbt), ETL Maintenance, Forecasting, Data Warehouse Design, bias detection, Experimental Design, Batch File Processing, Time Series Forecasting, Causal Inference, Time Series, Prompt Engineering, Large Language Models (LLMs), Sentiment Analysis, Deep Learning, Time Series Analysis, Time Series Data
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring