
Aditya Andra
Verified Expert in Engineering
Analyst and Developer
Hyderabad, Telangana, India
Toptal member since June 8, 2020
Aditya is a data scientist with around 12 years of experience developing and leading generative AI (GenAI) and machine learning (ML) solutions. He specializes in large language models (LLMs), retrieval-augmented generation (RAG) pipelines, and scalable MLOps for financial and operational analytics. Aditya has led the creation of a high-impact advanced analytics team in business finance at Novo Nordisk.
Portfolio
Experience
- Data Science - 7 years
- Machine Learning - 7 years
- Quantitative Research - 6 years
- Python - 6 years
- Time Series Analysis - 6 years
- Risk Modeling - 6 years
- Generative Pre-trained Transformers (GPT) - 3 years
- Natural Language Processing (NLP) - 3 years
Preferred Environment
Machine Learning, Python, Git, Jupyter, Data Science
The most amazing...
...project I've developed involved building a custom trial optimization model for a pharmaceutical company, which outperformed all existing ML models.
Work Experience
Lead AI ML Engineer
JPMorgan Chase
- Built a self-improving (dynamic knowledge base) AI chat solution for logical data model generation.
- Developed an LDM similarity solution using TigerGraph.
- Built multi-agentic systems that interact with each other to generate RL outcomes.
Senior Data Scientist
Novo Nordisk
- Architected RAG pipelines with vector embeddings, hybrid search, caching, and vector databases, enabling real-time analysis of unstructured data for decision-making in financial workflows.
- Developed a novel cosine similarity-based forecasting model for clinical trials and long-term sales forecasting of key products with limited historical data, achieving ±3% accuracy, and extracted insights from unstructured financial datasets.
- Built time series forecasting models using SOTA deep learning algorithms like N-HiTS and N-BEATS, which outperformed traditional ARIMA and Holt-Winters ES models.
- Built classification models to predict patient behavior. Used XGBoost and a multitude of data science techniques to optimize model performance.
- Developed a domain-specific financial commentary LLM using QLoRA fine-tuning, enabling context-aware insights for business applications.
Machine Learning Developer
Zvoid
- Created a tweet listener capable of listening to the tweets from a given list of authors and making the data ready for the decision engine.
- Built the automated trading capacity using the Alpaca API.
- Developed the end-end analysis of a particular Twitter IPO hypothesis.
- Worked on the decision engine using a random forest regressor that accepts the tweet and the stock price and gives out a stock buying or selling recommendation.
Senior Data Scientist
COGNIZER AI
- Developed a BERT-based conversational AI solution based on business requirements.
- Converted natural language queries into SQL queries using BERT-based deep-learning architecture.
- Contributed to significant parts of the back-end flow and took ownership of those flows.
- Extracted various fields from contract PDFs using regex and deep learning models and optimized the models to increase processing speed using TensorRT.
- Put the DL models into production using APIs and Docker. Used AWS and GCP to enable autoscaling features.
Data Scientist | Researcher
Freelance
- Built data pipelines for data coming from multiple sources like the Quandl API and a SQL database.
- Performed an exploratory data analysis on the built dataset, derived insights, and presented it to the stakeholders on Jupyter Notebook and Tableau.
- Modeled the data using decision tree-based regression models.
CTO
WiseLike
- Competed at the IE Business School's startup lab and won the investors' choice award and the most innovative project award.
- Developed the whole machine learning pipeline from scratch, starting with a web scraper for pictures, extracting properties of a picture, and training the model using the data.
- Served the model using a REST API (Flask) on the website wiselike.pythonanywhere.com.
- Performed A/B and hypothesis testing to test the validity of the model.
Quantitative Analyst
Futures First
- Performed an exploratory data analysis on large-scale financial datasets and derived insights that led to tradable strategies, using Python and visualizing data through dashboards in Tableau.
- Implemented a time series analysis (SARIMA and GARCH) of prices in commodity markets, considering CFTC reports and external factors like currency.
- Developed regression-based mean-reverting strategies in fixed-income markets of the US and Brazil.
- Deployed ETL pipelines and ML pipelines working on GCP.
- Performed backtesting and forward testing of strategies by tracking their Sharpe ratios.
- Performed hypothesis testing and evaluated the risk for strategies based on Monte Carlo simulations and historical value at risk.
- Built natural language pipelines to track news sentiment.
Research Intern
Next Sapiens
- Developed a novel 4D (degrees of freedom) solution for the simultaneous localization and mapping of an unmanned aerial vehicle to reduce the computation cost and published research on the same (Leeexplore.ieee.org/document/6461785).
- Combined location data from various sources like LIDAR, proximity sensors, inertial measurement units, and camera using extended Kalman filters to update the state information of the robot.
- Developed a fuzzy logic-based PID controller for the unmanned aerial vehicle to maintain stability during flight.
Experience
Churn Prediction for a Book Publisher
Stock Suggestions | Distributed System with PySpark
Word Recommendation System for Movie and Series Reviews
SQL Database for North American Oil and Gas and Visualization through Tableau
Machine Learning Model to Suggest Better Pictures for Social Media
Generating Insights in Stock Market Data
Predicting the Probability of a Default of a Company to Make Loan Decisions
https://github.com/MBD-RiskandFraud/fintech_platform_ieLive Tweet Sentiment Tracking
Cancer Prediction Using VOC Data
https://github.com/adia4/voc_cancer_predictionVOC database with labeled cancer data. The results are deployed using a Flask API which predicts the kind of cancer based on the VOC content.
Sales Forecast Model for FMCG, Taking the COVID Scenario Into Account
Time Series Forecasting
End-to-end NLP Model Deployment
Built APIs to allow its interaction with external modules.
Dockerized the whole application.
Connected it with AWS and GCP solutions like Lambda, container registry, etc., to achieve autoscaling of the API.
Education
Accelerated General Management Program in General Management
IIM Ahmedabad - Ahmedabad, India
Master's Degree in Business Analytics and Big Data
IE Business School - Madrid, Spain
Bachelor of Technology Degree in Electrical Engineering
Indian Institute of Technology (ISM), Dhanbad - Dhanbad, India
Certifications
Data Engineering, Big Data, and Machine Learning on GCP Specialization
Coursera
Certification in Quantitative Finance
Fitch Learning
Skills
Libraries/APIs
Pandas, NumPy, Scikit-learn, Keras, REST APIs, Spark ML, TensorFlow, Natural Language Toolkit (NLTK), Spark Streaming, Bloomberg API, PySpark
Tools
ARIMA, Tableau, DataViz, Spark SQL, Jupyter, Git, MATLAB, Bloomberg, Reuters Eikon, Azure Machine Learning
Languages
Python, SQL, R, C++, Python 3, SAS, Embedded C, Excel VBA
Paradigms
Functional Programming, Quantitative Research, Object-oriented Programming (OOP), ETL, Linear Programming
Frameworks
Flask, Spark, Apache Spark, LangGraph
Platforms
Linux, Amazon Web Services (AWS), Windows, Amazon EC2, Apache Kafka, Google Cloud Platform (GCP), Docker, Pentaho, Jupyter Notebook, Databricks
Storage
MongoDB, Redshift, NoSQL, PostgreSQL, Azure SQL Databases, MySQL, Databases
Industry Expertise
Social Media
Other
Quantitative Modeling, Statistics, Finance, Time Series, Mathematics, Natural Language Processing (NLP), Financial Modeling, Machine Learning, Time Series Analysis, Risk Modeling, Automated Trading Software, Statistical Analysis, Data Science, Quantitative Analysis, Predictive Analytics, Data Analysis, Data Analytics, Statistical Modeling, Regression Modeling, Hypothesis Testing, Data Visualization, Forecasting, Deep Learning, Generative Pre-trained Transformers (GPT), Recommendation Systems, Multivariate Statistical Modeling, Data Modeling, Machine Learning Operations (MLOps), Data Warehousing, Algorithms, Dashboards, Web Scraping, Neural Networks, Computer Vision, Data Warehouse Design, Websites, Social Media Marketing (SMM), A/B Testing, Trading, Data Engineering, Big Data, APIs, Derivatives, Decision Trees, Signal Analysis, Custom BERT, Quantitative Finance, General Management, Supply Chain Optimization, Autoscaling, Gunicorn, Generative Artificial Intelligence (GenAI), RAG Pipelines, Agentic AI, Reinforcement Learning
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring