
Dauren Baitursyn
Verified Expert in Engineering
Natural Language Processing (NLP) Developer
Dubai, United Arab Emirates
Toptal member since November 18, 2021
Dauren is an AI engineer with 8+ years of experience building production ML systems, from LLM-powered applications to distributed training pipelines. He specializes in NLP (BERT, GPT, spaCy), conversational AI (Rasa, RAG), and MLOps (MLflow, Ray). He's built systems processing millions of records with LLM classification, semantic search, and ensemble forecasting. With a CS degree from KAIST (top 50 QS), Dauren delivers end-to-end AI solutions from research to production.
Portfolio
Experience
- Natural Language Processing (NLP) - 8 years
- Python - 8 years
- Machine Learning - 8 years
- SQL - 8 years
- ETL - 6 years
- Generative Pre-trained Transformers (GPT) - 6 years
- Deep Learning - 5 years
- Data Transformation - 5 years
Preferred Environment
Jupyter, Git, Linux, AWS IoT, Visual Studio Code (VS Code), Docker, Databricks, MLflow, PostgreSQL, Elasticsearch
The most amazing...
...thing was achieving 7x F1 improvement (0.11→0.81) on a 46-topic text classifier using BERT fine-tuning and transfer learning, and Power BI to visualize results.
Work Experience
Senior AI Engineer
Caterpillar
- Migrated CI/CD infrastructure for a multi-repo AI/ML platform from Azure DevOps to GitHub Actions, modernizing authentication to OIDC and cutting deployment configuration complexity by roughly 75%.
- Led a large-scale AWS Lambda runtime upgrade across dozens of functions and multiple teams, ensuring compliance ahead of a critical vendor deprecation deadline while closing all outstanding security vulnerabilities flagged by automated scanning.
- Designed and built new cloud automation pipelines—including video processing and LLM prompt optimization workflows—spanning AWS and Azure services with secure, production-grade infrastructure as code.
Senior AI Engineer
Amgreat North America
- Developed LLM-powered classification system using OpenAI GPT for hierarchical product categorization (5 categories, 30 sub-categories) with structured JSON output parsing and prompt engineering.
- Built a topic modeling pipeline using GSDMM and spaCy NLP for unsupervised customer insight extraction from 100,000+ social media posts.
- Architected a distributed ML training system with Dagster orchestration, Ray parallel computing, and MLflow experiment tracking.
- Implemented ensemble forecasting combining XGBoost, CatBoost, and Random Forest with Optuna hyperparameter optimization (100+ trials per model).
- Integrated Google Generative AI for multi-language video content categorization.
ML Engineer
Eduworks Corporation
- Led the development of a production conversational AI system using Rasa with a custom NLU pipeline, achieving 10+ intents and complex multi-turn dialogue management.
- Built a RAG-style semantic search system combining BM25 lexical retrieval with dense vector embeddings in Elasticsearch, improving retrieval relevance by 30%+.
- Implemented sentence transformer embeddings for document vectorization, enabling semantic similarity matching across 50,000+ documents.
- Achieved significant retrieval improvements: Top-1 accuracy +9.2%, Top-3 +30.5%, OOS recall +65.4%.
ML Engineer
Philip Morris International
- Developed a BERT-based multi-label text classification system (46 categories), achieving 7x F1 improvement (0.11 to 0.81) through transfer learning, fine-tuning, and data augmentation strategies.
- Migrated static reports to dynamic dashboards using Power BI. Built, maintained, and improved dashboards with complex data models with more than 30 entities.
- Built complex ETL pipelines using Python to transform and extract data from different data sources into data models (STAR schema) in Power BI.
- Communicated findings and insights to stakeholders in data-story presentations.
ML Engineer
One Technologies
- Migrated the chatbot service from Dialogflow to a Rasa open source chatbot platform.
- Achieved a baseline NLU model performance on production data: F1 score of 0.732.
- Implemented a CRUD microservice for a chatbot platform.
- Created pipelines for data cleaning, data transformation, and data checks for the chatbot.
- Automated tests for the chatbot data consistency and integrity.
Data Scientist
PrimeSource
- Built and deployed statistical models using SPSS modelers like loan pre-approval and top-up models for targeted consumer campaigns.
- Designed and implemented data views and tables for campaign data panels using SQL and built and maintained ETL processes for the temporary tables needed for statistical models.
- Performed the analysis for a diverse set of ad hoc requests from internal stakeholders, measured the effectiveness of rolled-out campaigns, and communicated the results with data-story presentations.
- Led a team of two data analysts, conducted daily stand-ups for check-ins and progress on tasks, and communicated the projects' status to the principal data scientist.
Assistant Researcher
Graduate School of Knowledge Service Engineering | KAIST
- Implemented an information retrieval framework for clinical decision support in cancer diagnosis and treatment.
- Submitted the paper to the TREC 2017 Precision Medicine Track conference, but it has not been accepted.
- Assisted with the research in the field of precision medicine.
Experience
NewsAgg
https://github.com/biddy1618/newsProjectThe project has a search engine to retrieve the articles based on the cosine similarity of the frequency-inverse document frequency (TF-IDF) representation of the queries and articles.
Kaggle Alice Competition
https://github.com/biddy1618/alicekagglecompetitionHere we will try to identify a user on the internet by tracking their sequence of attended web pages. The algorithm to be built will take a webpage session—a series of web pages attended consequently by the same person—and predict whether it belongs to Alice or somebody else.
As for 23.12.18, my current LB standing is 87th out of 2,000ish.
For this competition, the data was time-series session information regarding user browser history. We had to distinguish between a regular user and an intruder user based on sites visited and time of visit information.
I performed the exploratory data analysis (EDA), data wrangling, and feature engineering to achieve a ROC-AUC score of 0.95856.
IMPLEMENTED FEATURES
• Dummy hour feature of the session start time (from now on start time).
• Sin and cos transformation of the start time.
• Active start hours of the intruder.
• Dummy weekday feature of the start time.
• Active weekday of intruder feature.
• Dummy month feature of the start time.
• Sin and cos transformation of the year and day feature.
• Session length.
• Session sites stay standard deviation.
Prediction of Churn | Telco Customer Churn Sample
https://github.com/biddy1618/churn-rateWe managed to train two models; logistic regression and random forest. The motivation behind including these statistical methods is based on the fact that logistic regression is a classical method for classification that gives good interpretability and is more or less stable. In contrast, random forest is robust and works well with small datasets due to its bagging sampling.
Final ROC-AUC scores on test data are as follows:
• For logistic regression - 0.652
• For random forest - 0.989
The top five major features for logistic regression are state of the state of residence, age, highest education acquired, use of internet services, and several complaints.
Random forest showed much better performance measures than logistic regression, and its major features make much more sense.
From a business perspective, age, unpaid balance, annual income (and other major features) seem to be valid features for churn rate.
Complete ML Project - From Getting Raw Data to Deployment
https://github.com/biddy1618/udacity-mldevops-3-project-mlmodel-fastapi-herokuThis project showcases CI/CD pipeline implementation using GitHub Actions and Heroku deployment.
CI includes PyTest tests and PEP8 code correspondence. DVC encapsulates different training components into stages.
This project is well-documented and has a model card description.
Education
Bachelor's Degree in Computer Science
Korea Advanced Institute of Science and Technology (KAIST) - Daejeon, South Korea
Exchange Program Specialized in Computer Science
Innopolis - Kazan, Republic of Tatarstan, Russia
Exchange Program Specialized in Computer Engineering
Middle East Technical University - Ankara, Turkey
Certifications
Machine Learning DevOps Engineer Nanodegree
Udacity
Natural Language Processing
Coursera
Introduction to Deep Learning
Coursera
Mathematics for Machine Learning Specialization | PCA
Coursera
IBM Certified Specialist | SPSS Modeler Professional V3
IBM
Machine Learning Specialization
Coursera
Machine Learning
Coursera
Skills
Libraries/APIs
Pandas, SQLAlchemy, PyTorch, Scikit-learn, NumPy, SciPy, Rasa NLU, TensorFlow, REST APIs, SpaCy, LSTM, OpenAI API
Tools
Jupyter, Microsoft Power BI, SPSS Modeler, Rasa.ai, IBM SPSS, Git, Pytest, GitLab CI/CD, Docker Compose, GitLab, Logging, Plotly, Tableau, AWS CloudFormation, AWS Step Functions, AWS IAM, Amazon Simple Notification Service (SNS)
Languages
Python, SQL, Java, JavaScript, Scala
Paradigms
Object-oriented Programming (OOP), ETL, Clean Code, DevOps, Functional Programming, Testing, Human-computer Interaction (HCI), CRUD, Azure DevOps
Platforms
Visual Studio Code (VS Code), Jupyter Notebook, Linux, Docker, Amazon EC2, Heroku, Amazon Web Services (AWS), Apache Kafka, Databricks, AWS IoT, AWS Lambda
Storage
Data Pipelines, PostgreSQL, Elasticsearch, Databases, Relational Databases, MySQL, Amazon S3 (AWS S3)
Frameworks
Flask, DeepPavlov, GraphLab, Spark, OAuth 2, Ray, Optuna
Other
Algorithms, Data Structures, Machine Learning, Data Transformation, Data Science, Predictive Modeling, Web Crawlers, Web Scraping, Data Visualization, Data Analysis, Code Versioning, Natural Language Processing (NLP), Deep Learning, Hugging Face, GitHub Actions, Deployment, DevOps Engineer, Machine Learning Automation, GitOps, Experiment Tracking, Generative Pre-trained Transformers (GPT), Artificial Intelligence (AI), BERT, MLflow, Data Versioning, FastAPI, CI/CD Pipelines, APIs, Flake8, Probability Theory, Linear Algebra, Operating Systems, System Programming, Discrete Mathematics, Linear Optimization, OOP Designs, Computer Science, Information Retrieval, Live Chat, Chatbots, Search Engines, Containers, User Intent Scoring, Conda, Data, Data Cleaning, Dashboards, Data Modeling, Reporting, Statistical Methods, Data Extraction, Authorization, Data Migration, Scripting, Ad Campaigns, Customer Segmentation, Market Segmentation, Time Series, Time Series Analysis, Machine Learning Operations (MLOps), Data Mining, Large Language Models (LLMs), AWS ECS Fargate, Amazon RDS, Topic Modeling, Prompt Engineering, Dagster, OpenID Connect (OIDC), ARM, Azure AI Foundry, AWS Secrets Manager, JFrog
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring