
Mehdi Sattari
Verified Expert in Engineering
Data Scientist and Developer
Plano, TX, United States
Toptal member since March 2, 2026
Mehdi is an experienced data scientist who specializes in scalable machine learning, predictive modeling, and end-to-end data solutions. He thrives across fintech and industrial environments. Additionally, he has a strong background in customer segmentation, churn prediction, and credit risk analysis.
Portfolio
Experience
- Python - 8 years
- Scikit-learn - 8 years
- XGBoost - 7 years
- SQL - 7 years
- Machine Learning - 7 years
- Deep Learning - 4 years
- Snowflake - 3 years
- Time Series Forecasting - 2 years
Preferred Environment
JupyterLab, Visual Studio Code (VS Code)
The most amazing...
...long-standing churn issue I've resolved at iKhokha involved embedding process control logic into transactional data, driving a 31% retention uplift.
Work Experience
Data Scientist II
Elevate Credit
- Delivered 4th-generation Rise and Elastic subprime credit risk models targeting 6-month charge-off, improving AUC from 0.58 to 0.66 across originations and restoring portfolio performance.
- Enhanced ETL and ML pipelines using multiprocessing, nested CV, and XGBoost on Azure VMs with Docker, achieving around 80% faster feature selection and around 40% faster model training through automated, reproducible workflows.
- Implemented Boruta feature selection and Optuna/Bayesian tuning, reducing 4,000+ Clarity and Experian Trended and Premier attributes to less than 100, strengthening parsimony, stability, and regulatory interpretability in credit-risk models.
- Built a transaction-level risk POC using Prism Insight cashflow data, outperforming Prism CashScore Extend by 0.09+ AUC and improving early-delinquency separation and portfolio risk stratification.
Data Scientist
iKhokha
- Developed and deployed a Snowpark churn model forecasting 30-day-ahead attrition and integrated RFM with process control chart triggers to detect sudden activity shifts, expanding proactive early-warning coverage.
- Designed and executed controlled A/B tests using Pingouin to validate churn-driven retention strategies, achieving around 31% uplift in merchant retention with 95% statistical significance and strong effect sizes.
- Built and automated LTV forecasting pipelines using SARIMA, LSTM, and boosted trees, eliminating manual Excel-based forecasting and saving the finance team around 2 weeks per reporting cycle.
- Built interactive Microsoft Power BI and Tableau dashboards to track predicted versus actual churn, recovered revenue, and cohort trends, supporting real-time performance monitoring and product decisions.
- Developed segmentation frameworks using dimensionality reduction, clustering, and behavioral features, driving differentiated engagement strategies based on merchant risk and value.
Senior Data Scientist
Energy and Combustion Services
- Engineered Optimizer, a decision-analysis framework combining distribution fitting, outlier detection, equivalence testing, and Monte Carlo simulation to quantify intervention impact and optimize haul trucks' fuel consumption.
- Developed computer vision models for road defect detection using EfficientDet and YOLO, improving accuracy by 37%, and built sensor-based braking-event and friction risk models to enhance automated road quality assessment.
- Consulted for an Anglo-American mining operation, applying regression-based risk modeling within the Dataiku platform to forecast energy consumption and operational variability.
- Led technical mentorship for data analysts in Python and advanced anomaly detection techniques, including deep autoencoders and the PyOD package, strengthening team capability in statistical modeling and outlier detection.
Data Scientist
The Virtual Agent
- Built and deployed real estate propensity-to-buy and lead-scoring models, and developed logistic regression scorecards with in-database scoring in Microsoft SQL Server, achieving 95.9% classification accuracy for a first-time buyer scorecard.
- Developed an XGBoost model for the Ford Comprehensive insurance campaign, increasing conversion rates by 17% through improved targeting and uplift-driven audience selection.
- Led a small data team and delivered scalable AWS EC2–based ETL pipelines with NLP URL categorization using Word2Vec (82% accuracy), while communicating risk rank-ordering, cutoffs, and model performance to non-technical stakeholders.
Experience
Churn Prediction Framework
The XGBoost model generates monthly scores for the entire active merchant base, while the "Triggers" system provides daily risk updates. I deployed the entire pipeline as Snowpark UDFs and Stored Procedures within Snowflake, ensuring high-performance, in-database execution and automated scheduling.
To drive business impact, I built interactive Power BI and Tableau dashboards that track predicted vs. actual churn, recovered revenue, and cohort-level trends. This allows stakeholders to monitor intervention effectiveness in real-time and make data-driven decisions.
I designed a retention-led A/B test validating a 31% uplift in merchant retention. This experiment proved that my dual-strategy framework significantly outperforms the traditional baseline.
Subprime Credit-Risk Models
The models were ultimately validated rigorously via smoke tests, then containerized and deployed via Azure Docker containers. To ensure long-term stability, I implemented automated drift detection and Microsoft Power BI dashboards for real-time PSI and AUC monitoring. This provided a stable, scalable foundation for the organization's strategy, ensuring the models remained robust against economic shifts while meeting strict regulatory standards.
Education
PhD in Chemical Engineering
University of KwaZulu-Natal - Durban, South Africa
Certifications
Advanced Data Science with IBM
IBM
Applied AI with Deep Learning
IBM
Fundamentals of Scalable Data Science
IBM
Skills
Libraries/APIs
Snowpark, Scikit-learn, NumPy, XGBoost, Pandas, Joblib, Keras, TensorFlow, PyTorch, OpenCV, Natural Language Toolkit (NLTK), Beautiful Soup, PySpark
Tools
StatsModels, Plotly, ARIMA, SARIMA, Jira, MATLAB, Git
Languages
Python, Snowflake, SQL
Frameworks
Optuna, Apache Spark
Platforms
Azure, Docker, Visual Studio Code (VS Code)
Other
Boruta, JupyterLab, Bayesian Statistics, Time Series Forecasting, Clustering, Dimensionality Reduction, Image Classification, Numba, Deep Learning, Machine Learning, Pingouin, Data Build Tool (dbt), YOLOv5, Object Detection, Web Scraping
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring