Saikat Banerjee, Developer in Ridgefield, CT, United States
Saikat is available for hire
Hire Saikat

Saikat Banerjee

Quantitative Machine Learning Developer

Ridgefield, CT, United States

Toptal member since July 29, 2022

Bio

Saikat is a quantitative ML researcher with 10+ years of experience developing probabilistic models and optimization algorithms to separate weak signals from noise in high-dimensional data. He specializes in interpretable, uncertainty-aware Bayesian modeling and turns complex results into decision-ready evidence. His edge comes from a rare physics–statistics–genomics career path, always focused on weak signals, massive noise, correlated variables, sparse ground truth, and high overfitting risk.

Portfolio

New York Genome Center
Machine Learning, Agentic Coding, Claude Code, Codex, Dimensionality Reduction...
The University of Chicago
Statistical Methods, Bayesian Statistics, Linear Regression...
Max Planck Society
Bayesian Statistics, Statistical Methods, Linear Regression...

Experience

  • Bayesian Statistics - 12 years
  • Machine Learning - 10 years
  • Regression - 10 years
  • Predictive Modeling - 10 years
  • Generalized Linear Model (GLM) - 8 years
  • Convex Optimization - 5 years
  • Dimensionality Reduction - 5 years
  • Agentic Coding - 1 year

Preferred Environment

Python, C++, Claude Code, Codex, Snakemake, NumPy, SciPy, PyTorch, Pandas, MySQL

The most amazing...

...method I built is Clorinn: convex optimization yields reproducible low-rank latent structure from noisy, high-dimensional data — trustworthy for real decisions.

Work Experience

Staff Scientist

2023 - PRESENT
New York Genome Center
  • Developed fast convex optimization algorithms based on Frank–Wolfe methods to separate weak latent signal from noise in high-dimensional correlated data–convexity provided stable optimization, reproducible solutions, and interpretable latent structure.
  • Released open-source Python package Clorinn with modular solver/objective design, clear documentation, version-controlled development, and extensible APIs for applying the method to new noisy high-dimensional datasets.
  • Led collaborative analyses to identify shared and distinct signals across 100+ heterogeneous disease traits, translating latent structures into interpretable quantitative factors.
  • Designed reproducible Python/Snakemake workflows for model selection, cross-validation, stability analysis, and visualization to evaluate the robustness of learned latent representations.
  • Modeled dependency structure across heterogeneous data sources using structural equation modeling, causal graph/network approaches (Mendelian randomization), and LLM-assisted clustering to assess robustness of inferred relationships.
  • Extended B-LORE to predict outcomes after feature-selection from sparse, high-dimensional data, isolating the few relevant signals from large, noisy backgrounds.
  • Coordinated cross-institutional collaborations aligning analysis strategy, deliverables, and interpretation across stakeholders.
Technologies: Machine Learning, Agentic Coding, Claude Code, Codex, Dimensionality Reduction, Regression, Predictive Modeling, Causal Inference, Bayesian Inference & Modeling, Convex Optimization, Ubuntu, Python, Bayesian Statistics, Linear Regression, Logistic Regression, Data Analysis, NumPy, SciPy, Data Science, Research, Transcriptomics, Genomics, Claude, Claude API, Large Language Models (LLMs), Workflows

Postdoctoral Scientist

2020 - 2023
The University of Chicago
  • Variational inference (VI) for Bayesian multiple regression updates one feature at a time – I reformulated the problem to inherit the whole optimization toolkit: quasi-Newton, fast matrix-vector products, autodiff, and trivial prior swaps.
  • Improved predictive accuracy and recovery of relevant features from noisy, high-dimensional data by using an adaptive shrinkage prior in variational empirical Bayes inference for sparse multiple regression.
  • Developed a modular Python package, GradVI, which provides a framework for gradient-based variational inference, adaptable to all prior families with analytically tractable Normal Means representations.
Technologies: Statistical Methods, Bayesian Statistics, Linear Regression, Logistic Regression, Predictive Modeling, Machine Learning, Generalized Linear Model (GLM), Bayesian Inference & Modeling, Time Series Analysis, Convex Optimization, Regression, Ubuntu, Python, NumPy, SciPy, Data Science, Research, Causal Inference, Bayesian Networks, Transcriptomics, Genomics

Postdoctoral Scientist

2015 - 2020
Max Planck Society
  • Led development of Tejaas, a scalable “reverse regression” framework for discovering weak signals and inferring networks in high-dimensional expression data; resulted in a corresponding author publication.
  • Improved speed/accuracy of variable selection in a variational empirical Bayes framework for multiple logistic regression by introducing a quasi-Laplace approximation for the analytical treatment of otherwise intractable integrals in non-linear models.
  • Established and led a research team — supervised a master's thesis and mentored three internship students. Mentored one postdoc.
  • Collaborated with clinicians and multidisciplinary teams within the e:AtheroSysMed consortium.
  • Presented our work at the 2019 International Society for Computational Biology conference and 2020 e:Med. Invited for a presentation at the University of Göttingen.
Technologies: Bayesian Statistics, Statistical Methods, Linear Regression, Logistic Regression, Predictive Modeling, Machine Learning, Bayesian Inference & Modeling, Generalized Linear Model (GLM), Bayesian Networks, Regression, Ubuntu, Python, Data Analysis, NumPy, SciPy, Data Science, Research, Causal Inference, Transcriptomics, Genomics, Workflows

Experience

Patient Subgroup Discovery for Differential Drug Response

When response to a treatment varies across a population, the commercially valuable question is which subgroups respond differently and which features drive that difference. This is hard for the usual reasons: the discriminating signal is weak, the candidate features are numerous and correlated across several data modalities, and labeled ground truth is scarce — so naive models overfit and identify subgroups that don't replicate.
Working with a genomics startup, I aligned preprocessing across heterogeneous data sources, selected the molecular and quantitative features most predictive of differential response, and built and evaluated classification models to stratify patients into response subgroups. Beyond the modeling, I designed the analysis strategy and translated the results into decision-ready evidence that non-technical stakeholders could act on for translational use cases.

Trans-eQTL Discovery from GTEx Data

Genetic variants regulating distant target genes are called trans-acting expression quantitative trait loci (trans-eQTLs). Many genetic variants are believed to mediate disease risk via the trans-eQTLs. It is crucial to discover trans-eQTLs and understand their mechanism to reveal the link between genetic variants and disease phenotypes. It is challenging to identify trans-eQTLs due to small effect sizes, tissue specificity, and a severe multiple-testing burden.

Our goal was to develop a reliable method of identifying trans-eQTLs. We proposed a new model and created open-source software. Applying our method to eQTL data from the Genotype-Tissue Expression Project (GTEx) demonstrated that its performance is significantly better than that of the state-of-the-art.

Bayesian Multiple Logistic Regression

https://doi.org/10.1371/journal.pgen.1007856
Logistic regression is the method of choice to analyze binary outcomes. Multiple logistic regression uses numerous variables in a logistic model. Bayesian multiple logistic regression offers several benefits, including variable selection, prediction, easier interpretation of results, and leveraging prior information. However, Bayesian multiple logistic regression requires costly and technically challenging Markov Chain Monte Carlo (MCMC) sampling or approximations that significantly reduce the logistic model's flexibility.

We proposed a methodology using the point-normal prior for faster and more accurate Bayesian multiple logistic regression, developing open-source software for the project. Applying our method to human genetics data, we proved it outperforms state-of-the-art variable selection and prediction for sparse multiple logistic regression problems of high dimension (n >> p problems.)

Education

2010 - 2015

PhD in Computational Biophysics

Indian Institute of Science - Bangalore, India

2007 - 2010

Master's Degree in Chemistry

Indian Institute of Science - Bangalore, India

Skills

Libraries/APIs

NumPy, SciPy, Scikit-learn, Matplotlib, MPI, OpenMP, PyTorch, Claude API, Pandas

Tools

Jupyter, Shell, Claude, GitHub, Claude Code, Codex, Snakemake

Languages

Python, Bash, C++, Fortran, SQL

Platforms

Ubuntu, Linux, Debian

Paradigms

Parallel Programming

Storage

MySQL

Industry Expertise

Transcriptomics, Bioinformatics

Other

Bayesian Statistics, Statistical Methods, Linear Regression, Logistic Regression, Biostatistics, Predictive Modeling, Machine Learning, Data Analysis, Research, Generalized Linear Model, Regression, Dimensionality Reduction, Convex Optimization, Workflows, Biophysics, Generalized Linear Model (GLM), Mixed-effects Models, Computational Biological Physics, Agentic Coding, Large Language Models (LLMs), Data Science, Natural Language Processing (NLP), Numerical Optimization, Bayesian Inference & Modeling, Causal Inference, Computational Biophysics, Linux HPC, Differential Equations, Equilibrium and Non-equilibrium Statistical Mechanics, Bayesian Networks, Time Series Analysis, Model Regularization, Genomics, Mathematics, Random Forests, Bayesian Classifiers, Data Classification, AI Consulting, Consulting

Collaboration That Works

How to Work with Toptal

Toptal matches you directly with global industry experts from our network in hours—not weeks or months.

1

Share your needs

Discuss your requirements and refine your scope in a call with a Toptal domain expert.
2

Choose your talent

Get a short list of expertly matched talent within 24 hours to review, interview, and choose from.
3

Start your risk-free talent trial

Work with your chosen talent on a trial basis for up to two weeks. Pay only if you decide to hire them.

Top talent is in high demand.

Start hiring