
Saikat Banerjee
Verified Expert in Engineering
Quantitative Machine Learning Developer
Ridgefield, CT, United States
Toptal member since July 29, 2022
Saikat is a quantitative ML researcher with 10+ years of experience developing probabilistic models and optimization algorithms to separate weak signals from noise in high-dimensional data. He specializes in interpretable, uncertainty-aware Bayesian modeling and turns complex results into decision-ready evidence. His edge comes from a rare physics–statistics–genomics career path, always focused on weak signals, massive noise, correlated variables, sparse ground truth, and high overfitting risk.
Portfolio
Experience
- Bayesian Statistics - 12 years
- Machine Learning - 10 years
- Regression - 10 years
- Predictive Modeling - 10 years
- Generalized Linear Model (GLM) - 8 years
- Convex Optimization - 5 years
- Dimensionality Reduction - 5 years
- Agentic Coding - 1 year
Preferred Environment
Python, C++, Claude Code, Codex, Snakemake, NumPy, SciPy, PyTorch, Pandas, MySQL
The most amazing...
...method I built is Clorinn: convex optimization yields reproducible low-rank latent structure from noisy, high-dimensional data — trustworthy for real decisions.
Work Experience
Staff Scientist
New York Genome Center
- Developed fast convex optimization algorithms based on Frank–Wolfe methods to separate weak latent signal from noise in high-dimensional correlated data–convexity provided stable optimization, reproducible solutions, and interpretable latent structure.
- Released open-source Python package Clorinn with modular solver/objective design, clear documentation, version-controlled development, and extensible APIs for applying the method to new noisy high-dimensional datasets.
- Led collaborative analyses to identify shared and distinct signals across 100+ heterogeneous disease traits, translating latent structures into interpretable quantitative factors.
- Designed reproducible Python/Snakemake workflows for model selection, cross-validation, stability analysis, and visualization to evaluate the robustness of learned latent representations.
- Modeled dependency structure across heterogeneous data sources using structural equation modeling, causal graph/network approaches (Mendelian randomization), and LLM-assisted clustering to assess robustness of inferred relationships.
- Extended B-LORE to predict outcomes after feature-selection from sparse, high-dimensional data, isolating the few relevant signals from large, noisy backgrounds.
- Coordinated cross-institutional collaborations aligning analysis strategy, deliverables, and interpretation across stakeholders.
Postdoctoral Scientist
The University of Chicago
- Variational inference (VI) for Bayesian multiple regression updates one feature at a time – I reformulated the problem to inherit the whole optimization toolkit: quasi-Newton, fast matrix-vector products, autodiff, and trivial prior swaps.
- Improved predictive accuracy and recovery of relevant features from noisy, high-dimensional data by using an adaptive shrinkage prior in variational empirical Bayes inference for sparse multiple regression.
- Developed a modular Python package, GradVI, which provides a framework for gradient-based variational inference, adaptable to all prior families with analytically tractable Normal Means representations.
Postdoctoral Scientist
Max Planck Society
- Led development of Tejaas, a scalable “reverse regression” framework for discovering weak signals and inferring networks in high-dimensional expression data; resulted in a corresponding author publication.
- Improved speed/accuracy of variable selection in a variational empirical Bayes framework for multiple logistic regression by introducing a quasi-Laplace approximation for the analytical treatment of otherwise intractable integrals in non-linear models.
- Established and led a research team — supervised a master's thesis and mentored three internship students. Mentored one postdoc.
- Collaborated with clinicians and multidisciplinary teams within the e:AtheroSysMed consortium.
- Presented our work at the 2019 International Society for Computational Biology conference and 2020 e:Med. Invited for a presentation at the University of Göttingen.
Experience
Patient Subgroup Discovery for Differential Drug Response
Working with a genomics startup, I aligned preprocessing across heterogeneous data sources, selected the molecular and quantitative features most predictive of differential response, and built and evaluated classification models to stratify patients into response subgroups. Beyond the modeling, I designed the analysis strategy and translated the results into decision-ready evidence that non-technical stakeholders could act on for translational use cases.
Trans-eQTL Discovery from GTEx Data
Our goal was to develop a reliable method of identifying trans-eQTLs. We proposed a new model and created open-source software. Applying our method to eQTL data from the Genotype-Tissue Expression Project (GTEx) demonstrated that its performance is significantly better than that of the state-of-the-art.
Bayesian Multiple Logistic Regression
https://doi.org/10.1371/journal.pgen.1007856We proposed a methodology using the point-normal prior for faster and more accurate Bayesian multiple logistic regression, developing open-source software for the project. Applying our method to human genetics data, we proved it outperforms state-of-the-art variable selection and prediction for sparse multiple logistic regression problems of high dimension (n >> p problems.)
Education
PhD in Computational Biophysics
Indian Institute of Science - Bangalore, India
Master's Degree in Chemistry
Indian Institute of Science - Bangalore, India
Skills
Libraries/APIs
NumPy, SciPy, Scikit-learn, Matplotlib, MPI, OpenMP, PyTorch, Claude API, Pandas
Tools
Jupyter, Shell, Claude, GitHub, Claude Code, Codex, Snakemake
Languages
Python, Bash, C++, Fortran, SQL
Platforms
Ubuntu, Linux, Debian
Paradigms
Parallel Programming
Storage
MySQL
Industry Expertise
Transcriptomics, Bioinformatics
Other
Bayesian Statistics, Statistical Methods, Linear Regression, Logistic Regression, Biostatistics, Predictive Modeling, Machine Learning, Data Analysis, Research, Generalized Linear Model, Regression, Dimensionality Reduction, Convex Optimization, Workflows, Biophysics, Generalized Linear Model (GLM), Mixed-effects Models, Computational Biological Physics, Agentic Coding, Large Language Models (LLMs), Data Science, Natural Language Processing (NLP), Numerical Optimization, Bayesian Inference & Modeling, Causal Inference, Computational Biophysics, Linux HPC, Differential Equations, Equilibrium and Non-equilibrium Statistical Mechanics, Bayesian Networks, Time Series Analysis, Model Regularization, Genomics, Mathematics, Random Forests, Bayesian Classifiers, Data Classification, AI Consulting, Consulting
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring