
Anish Chakravartty
Verified Expert in Engineering
ETL Developer
Kolkata, West Bengal, India
Toptal member since February 9, 2022
Anish is an IT professional with over 13 years of experience in various verticals and knowledge of several technologies. His experience includes working at several service provider companies and freelancing. As a professional with vast working experience with state-of-the-art and modern technologies in data engineering, cloud solutions, and data analytics, Anish excels in full-stack languages and tools, including JavaScript and MEVN stack.
Portfolio
Experience
- Business Intelligence (BI) - 11 years
- Google BigQuery - 6 years
- Spark - 4 years
- Azure Databricks - 4 years
- Cloud Dataflow - 4 years
- Snowflake - 3 years
- Retrieval-augmented Generation (RAG) - 2 years
- Geospatial Analytics - 2 years
Preferred Environment
PostgreSQL, BigQuery, Google Cloud, Node.js, Azure, Microsoft Power BI, Azure Databricks, Databricks
The most amazing...
...thing I've worked on is designing the ETL framework for IBM InfoSphere DataStage batch jobs, enabling over 3,000 ETL processes to load data across the project.
Work Experience
Senior Data Engineer
Battle Creek Games
- Built scalable ETL pipelines to collect and integrate ad performance data from IronSource, AppsFlyer, and Google Analytics, including metrics such as impressions, clicks, installs, revenue, and cost.
- Automated the ingestion and transformation of cohort and attribution data, enabling real-time insights into campaign performance across multiple game titles.
- Engineered data quality checks and normalization logic to reconcile discrepancies between different ad platforms, ensuring consistency in key performance metrics.
- Implemented creative-to-campaign mapping logic, enabling the identification of underperforming ad creatives and supporting actionable optimizations by the marketing team.
- Created a unified reporting layer in Snowflake, blending cost, revenue, and LTV data across channels to support ROI analysis and budget allocation decisions.
- Reduced reporting latency by over 70% by migrating to a modular, batched ETL framework with retry and fallback mechanisms for unstable external APIs.
- Collaborated with game designers and UA teams to develop custom dashboards and metrics tailored to user acquisition, retention, and monetization strategies.
- Enabled downstream modeling for predictive LTV and churn analysis, laying the groundwork for smarter ad spend optimization and audience segmentation.
Senior Data Engineer
Citian
- Led the development of a spatial analytics pipeline to integrate crash and event datasets with geospatial features (e.g., segments and intersections) using Snowflake and GeoPandas, enabling location-based analysis across transportation networks.
- Performed large-scale spatial joins between point-level incident data and geometry-based segment layers, resulting in a 30% improvement in spatial attribution accuracy for traffic events.
- Built reusable Python modules to process and enrich geospatial data from Snowflake, including automatic handling of WKT and GeoJSON formats, and wrote results back into dynamic Snowflake tables with complete audit trails.
- Implemented segment- and intersection-level aggregation logic, allowing city planners to identify high-risk areas based on crash density, severity, and time patterns.
- Created statistical dashboards and summary tables leveraging Snowflake dynamic tables and streamlining data consumption for internal analysts and DOT partners.
- Automated the end-to-end ETL pipeline for crash/event ingestion, spatial enrichment, and metric generation using Python and scheduled orchestrations, reducing manual intervention by 90%.
- Developed and maintained fallback logic for incomplete or unmatched spatial data points, ensuring robust and comprehensive reporting even with inconsistent source feeds.
- Collaborated with data governance and infrastructure teams to optimize Snowflake storage, table partitioning, and compute usage, achieving a 40% reduction in overall query cost for heavy geospatial workloads.
Azure Database Developer
FluidityIQ, LLC
- Designed and implemented a scalable retrieval-augmented generation (RAG) system for a patent search platform by integrating a SentenceTransformer-based pipeline, replacing a slower Databricks-native GTE embedding API.
- Increased vector generation throughput by 5x by migrating to GPU-accelerated clusters and optimizing embedding logic, achieving over 500 embeddings per second at peak load.
- Extracted, cleaned, and transformed over 100 million patent documents, ensuring high-quality, chunked textual input optimized for embedding and semantic search.
- Built and deployed a robust microbatching architecture to handle Pinecone rate limits, enabling efficient and reliable ingestion of 100,000 vectors per batch with automatic retries and cooldown intervals.
- Reduced overall embedding pipeline latency by over 60% by replacing UDF-based inference with native Spark-NLP transformer integration and custom batching logic.
- Automated the end-to-end RAG pipeline orchestration using Databricks Jobs API, including dependency tracking, error logging, and batch-level recovery mechanisms.
- Enabled domain-specific semantic search by combining embeddings with metadata-aware filtering, improving top-k result relevance by 40% in internal evaluations.
- Collaborated with ML and product teams to align vector search outputs with user expectations, iterating on prompt engineering and chunking strategies to boost retrieval quality.
IoT Data Engineer - Azure & Looker Studio
Regent Climate Connect Knowledge Solutions Pvt Ltd
- Created a warehouse on BigQuery to store SCADA data from various power plants, handling modeling, ETL, and DevOps tasks. Created ETL jobs on Cloud Dataflow to feed data to downstream systems.
- Assisted the Data Science team in deploying their data models on Vertex AI and created a dashboard on Looker on the model outputs, automating the entire process.
- Performed several enhancements and bug fixes on an existing Data Lake on Azure and the ETL framework on Databricks and Data Factory. Set up a fresh dashboard on Power BI based on the Data Lake.
- Helped set up and manage the data team for the company.
Data Engineers (Toptal Teams)
Compliance Made Easy Inc. d/b/a CertifyOS
- Created an ETL framework using Apache Beam (Cloud Dataflow) to migrate data from Firestore to PostgreSQL.
- Employed Cloud Run to execute Dockerized Python jobs essential for the migration process.
- Used Cloud Workflow to orchestrate the entire ETL framework.
Senior Developer
Freelance Clients
- Developed a Wix plugin dealing with abandoned carts as a full-stack developer.
- Acted as a back-end developer for a mobile app providing banking services.
- Contributed to a bot creation project for a pharmaceutical test clinic as the primary developer.
- Developed a plugin for the front integration platform as a full-stack developer.
- Created multiple chatbots built on a varied number of platforms.
Assistant Consultant
Tata Consultancy Services
- Designed an ETL framework that enables 3,000+ batch ETL jobs to load data across various subject areas.
- Acted as a senior developer and architect for migrating on-prem databases to the cloud.
- Designed and developed a data-fix process in the data warehouse and ODS layers when source systems were merged during the cloud migration.
- Contributed to the SME projects using IBM InfoSphere DataStage and Oracle PL/SQL.
- Conducted technical interviews for relevant technologies while onboarding new people across the project.
- Received recognition and an award as the contextual master of the entire business unit and numerous appreciations from senior client management for successfully implementing several projects.
ETL Specialist
IBM
- Designed the ETL architecture to load data into a data warehouse for an insurance client.
- Created the architecture of data marts for several service lines.
- Acted as a senior developer for a data migration project to SAP.
- Performed efficient multiple data recoveries and various spontaneous actions for which I received appreciation emails from business users.
CRM Consultant
Cognizant
- Loaded data into Siebel CRM as an ETL subject matter expert.
- Tuned the data loading tasks with low performance, which saved a lot of time.
- Provided training to new joiners on DataStage development and application-based knowledge transfer.
Project Engineer
Wipro
- Acted as the SPOC for data warehousing applications for EMEA and APAC regions.
- Implemented performance enhancement and new geography accommodation for existing applications successfully. Reduced run time by more than half as needed for older tasks.
- Trained and guided new team members on DataStage job development and application-based knowledge transfer.
- Received the AIM Star of The Year and two awards for individual performance in different projects.
Experience
Accelerating Patent Discovery with a Scalable RAG Pipeline and GPU-optimized Embeddings
Integrated the pipeline with Pinecone using a custom microbatching architecture that ingested 100,000 vectors per batch with rate-limit handling, retries, and resume logic. Enriched all embeddings with metadata tags such as section, jurisdiction, and kind code to enable highly filtered semantic search.
End-to-end orchestration was automated using Databricks Jobs API, ensuring fault tolerance and batch-level traceability. This project enabled faster and more accurate prior art discovery, improving search relevance by 40% and reducing indexing latency by over 60%, directly aligning with the client’s goal to enhance innovation analysis and user productivity.
Geospatial Crash Analytics Platform Using Snowflake and Python for Location-based Insights
Designed and implemented functions to aggregate crash metrics at both segment and intersection levels, incorporating attributes such as crash severity, time-of-day, and roadway conditions. Created fallback logic to handle unmatched points, ensuring data completeness across the pipeline.
Engineered a robust ETL workflow to extract, process, and write enriched results back into Snowflake dynamic tables, supporting downstream analysis and reporting. Optimized spatial processing to reduce compute time and improve query performance for heavy geospatial workloads.
This solution empowered transportation analysts and planners to identify high-risk areas, evaluate safety countermeasures, and support data-driven decision-making for infrastructure improvements across city and state networks.
Data Warehouse for a Renewable Energy Provider on GCP
After loading the data into the warehouse, I created various materialized views with different aggregations to optimize performance. These views were utilized in a dashboard built in Looker for easy visualization and analysis. Additionally, I developed a web app dashboard that utilized data from these materialized views for real-time insights.
Additionally, I constructed a batch ETL framework using Apache Beam deployed on Cloud Dataflow, generating files on Cloud Storage for data models deployed on Vertex AI. The output of these models was leveraged on Looker dashboards and the web app.
I created the architecture of the entire system and worked on its implementation. Also, I set up and managed the initial phases of a data engineering team for the clients.
Adtech Data Integration and Campaign Optimization for a Mobile Gaming Platform
Engineered transformation logic to normalize and reconcile cross-platform data, accounting for attribution differences and enhancing data accuracy. Implemented creative-to-campaign mapping, allowing the team to identify underperforming creatives and iterate rapidly on ad design and targeting.
Designed a unified reporting layer in Snowflake, combining cohort data and LTV projections to support ROI calculations and budget decisions. Reduced pipeline latency by over 70% through a modular architecture with retry handling and efficient batching for API calls.
This project empowered the marketing and UA teams to optimize ad spend, increase campaign ROI, and scale user acquisition strategies with confidence, directly impacting revenue and player engagement.
ETL Framework to Enable InfoSphere DataStage Batch Loads
I created the architecture of the new framework and worked as the technical lead in implementing the new framework and migrating jobs. The completed project enabled DataStage jobs migration without any changes, which resulted in all-around applause and appreciation.
Abandoned Cart Recovery Wix Plugin
I contributed to this project as a full-stack developer creating the front end for the Wix plugin and subscription modal script in Vue 2. The back end was hosted as multiple Google Cloud functions using Node.js and Express.js, which followed the REST architecture. We used Cloud Firestore as the operational database.
WhatsApp Bot and ODS Development for Pharmacy Test Clinic
The second part of the project consisted of designing an ODS in BigQuery that used data from the WhatsApp bot, client website, and physical store data. Cloud functions were used to consume data in the form of files, REST API calls, and BigQuery stored procedures wrapped in pub/sub cloud functions for loading data into the ODS.
Education
Bachelor's Degree in Electronics and Instrumentation Engineering
West Bengal University of Technology - Kolkata, India
Skills
Libraries/APIs
Node.js, PySpark, Vue, Pandas, Hugging Face Transformers, Rasa NLU, Vue 2
Tools
BigQuery, IBM InfoSphere (DataStage), Dialogflow, Apache Beam, Microsoft Power BI, Looker, Cloud Dataflow, Apache Airflow, Amazon Elastic Container Service (ECS), Docker Compose, Google Analytics, ManyChat, Terraform
Languages
JavaScript, SQL, Python, Snowflake, C, C++, Embedded C
Frameworks
Express.js, Spark, Botpress.io, LangGraph, Unity
Paradigms
ETL, Business Intelligence (BI), Microservices, DevOps
Platforms
Oracle, Google Cloud Platform (GCP), Databricks, Amazon Web Services (AWS), AWS Lambda, Apache Kafka, Docker, Azure, LangSmith, Chatfuel, Kubernetes, Firebase, AWS IoT, Vertex AI, AppsFlyer
Storage
PL/SQL, Oracle PL/SQL, Databases, Data Pipelines, Data Lakes, PostgreSQL, Google Cloud, IBM Db2, SQL Server 2010, Cloud Firestore, MongoDB, Siebel EIM, Teradata, Database Migration, NoSQL
Other
PL/SQL Tuning, Google Cloud Functions, Data Engineering, Data Architecture, Data Warehousing, Data Modeling, Google Data Studio, API Integration, Pub/Sub, Google BigQuery, Azure Databricks, Data Analysis, Pinecone, Retrieval-augmented Generation (RAG), Amazon Managed Workflows for Apache Airflow (MWAA), Full-stack, Azure Data Lake, Azure Data Factory (ADF), CI/CD Pipelines, LangChain, GeoPandas, Geospatial Data, Geospatial Analytics, Metabase, Shell Scripting, Microprocessors, Internet of Things (IoT), Looker Studio, BERT, Hugging Face, Advertising Technology (Adtech)
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring