
Vishwajeet Thakur
Verified Expert in Engineering
Senior Full-stack Developer
Ballia, Uttar Pradesh, India
Toptal member since April 28, 2026
Vishwajeet is a software developer specializing in Python, FastAPI, SQL/PostgreSQL, and production GenAI systems. He has built audited RAG platforms, agentic workflows, and high-throughput back-end services on AWS using LangChain, LangGraph, LangSmith, Pinecone, Docker, and ClearML to help enterprise teams search better, investigate faster, and deploy AI with confidence. Vishwajeet will be a great addition to any team.
Portfolio
Experience
- Python - 6 years
- PostgreSQL - 5 years
- Databases - 5 years
- Web Development - 5 years
- FastAPI - 4 years
- LangChain - 3 years
- AI Agents - 3 years
- RAG Architecture - 3 years
Preferred Environment
PostgreSQL, FastAPI, Django, LangChain, LangGraph, RAG Architecture, AI Agents, Docker, Pinecone, CSS, Privacy, Security, API Design, HTML5, Communication, Role-based Access Control (RBAC), CrewAI, Google Cloud Platform (GCP), BigQuery, Terraform Cloud, Software Engineering, API Management, DevOps
The most amazing...
...project I've done was turning a log-analysis workflow into an audited RAG platform that gave analysts cited, structured answers instead of raw logs.
Work Experience
Senior Back-end GenAI Engineer
Pfizer
- Optimized back-end APIs handling 500,000+ daily transactions, reducing response time from around 800ms to under 300ms, which is around 60% improvement through query tuning and caching.
- Implemented RBAC with OAuth and JWT, resolving audit vulnerabilities and ensuring compliance with enterprise healthcare data security standards.
- Built REST APIs by integrating 5+ internal systems, enabling real-time data flow and reducing manual processing effort by 37%.
- Increased test coverage to over 85%, reducing post-release defects in critical modules by 26%.
- Built Pfizer RAG for operational-log search using LangChain hybrid retrieval (BM25 and Pinecone vectors) and cross-encoder reranking, returning cited snippets with metadata filters for team, system, and time window.
- Developed LangGraph agent workflows to read PDFs, images, and tables, call Bedrock and Pinecone tools, and produce structured RCA summaries, and added pytest evaluations for Recall@k, MRR@10, groundedness, and parse-fail rates.
Software Engineer
Charter Communications
- Reduced device-ingestion time from 219 seconds to 4.07 seconds (98% faster) by redesigning the pipeline with multiprocessing, concurrent futures, and connection pooling for high-volume vendor data ingestion.
- Built a greenfield device-operations platform on AWS with 45+ versioned REST endpoints, Dockerized FastAPI services on EKS, PostgreSQL on RDS, and automated SSH workflows for artifact retrieval, image installation, verification, and health checks.
- Delivered a production RAG assistant for NOC runbooks and ticket triage using Pinecone, hybrid BM25 with vector retrieval, cross-encoder reranking, and Pydantic-validated outputs to produce more precise, source-grounded answers.
- Implemented vendor adapters for Ciena, Nokia, and other providers, handling authentication, pagination, and payload normalization. Scheduled reconciliation jobs with idempotent upserts and full audit trails across shipment and inventory data.
Data Engineer
Charles Schwab
- Built versioned Django REST APIs with ViewSets, Routers, serializers, and optimized ORM queries to support predictable latency for batch-triggered data workflows.
- Owned 20+ Airflow DAGs with sensors, retries, SLA alerts, backfills, and environment-based parameters for nightly ETL orchestration.
- Developed PySpark jobs on EMR to ingest, cleanse, and conform datasets using schema enforcement, partitioning, joins, and window functions.
- Produced optimized Parquet datasets on S3 with deterministic partition layouts for reliable downstream analytics and reprocessing.
- Modeled Snowflake star schemas with multiple fact tables and around 18 dimensions, including SCD type 2 and MERGE-based upserts.
- Implemented Snowflake data pipelines using Streams, Tasks, MERGE, and Time Travel to support governed warehouse processing.
- Added Airflow-integrated data quality gates for referential integrity, range checks, SLA failures, and deterministic reruns through checkpointed stages.
- Tuned Spark shuffles, partitions, file sizes, and commit frequency to meet nightly batch SLAs and reduce Snowflake warehouse load cost.
Experience
Hybrid RAG System for Investigative Log Analytics for High-recall with Cross-encoder Reranking
KEY CONTRIBUTIONS
• Developed an async FastAPI back end with strict Pydantic schemas powering semantic search, investigation sessions, and audit APIs, handling 500-700 queries per day with consistent structured outputs.
• Implemented hybrid retrieval with dense vectors, metadata filters, and keyword constraints to improve recall across noisy logs, followed by cross-encoder reranking on top-k candidates, improving top-5 relevance/grounding by 20-35%, which was measured by offline evaluation.
• Built ingestion pipelines for log normalization, chunking, and embedding, processing millions of log lines per day with high-throughput bulk upserts and transactional guarantees using SQLAlchemy.
• Orchestrated retrieval and generation via LangChain with enforced structured outputs (root-cause hypotheses and next actions).
• Integrated LangSmith to track latency, retrieval quality, and failure modes, reducing debugging time and iteration cycles.
The application was deployed on EKS.
Autonomous Browser Agent for Authenticated External Portals
Combined browser use for semantic interaction with unfamiliar pages and Playwright for deterministic control of navigation, frames, accessibility locators, cookies, session state, downloads, screenshots, and network traffic. The agent automated login, account selection, date filters, pagination, iframes, and statement retrieval through an observe plan, act, verify loop.
The system supported approved MFA paths and detected CAPTCHA, WAF, region blocks, expired sessions, and rate limits; and resumed failed jobs from durable checkpoints. Downloads were validated through file signatures, parser checks, account metadata, statement periods, and HTML error detection. I added idempotency keys, SHA-256 deduplication, transactional upserts, tenant-isolated contexts, secret redaction, read-only controls, and audit correlation IDs.
Education
Bachelor's Degree in Computer Science
National Institute of Technology, Arunachal Pradesh - Jote, India
Skills
Libraries/APIs
React, Pydantic, SQLAlchemy, Node.js, API Development, Stripe, REST APIs, Claude API, PySpark, NumPy, Pandas, TensorFlow, Amazon Marketplace Web Service (MWS), Playwright
Tools
GitHub, Jira, Terraform, Claude Code, Claude, BigQuery, RabbitMQ, Apache Airflow, Amazon EKS, Amazon SageMaker, Pytest
Languages
Python, SQL, JavaScript, TypeScript, GraphQL, CSS, HTML5, HTML, Clojure, Snowflake
Frameworks
Django, LangGraph, Django REST Framework, Next.js, Tailwind CSS, Material UI, ClojureScript, Agentic Frameworks, OAuth 2
Paradigms
REST, Microservices, ETL, Synthetic Data Generation, Model Context Protocol (MCP), HIPAA Compliance, Automated Testing, Automation, Role-based Access Control (RBAC), Functional Programming, Event-driven Design (EDD), DevOps, Event-driven Architecture, Real-time Systems
Platforms
Docker, Amazon Web Services (AWS), Kubernetes, Clerk, AWS Lambda, Jupyter Notebook, Google Cloud Platform (GCP), LangSmith, Apache Kafka, Azure, Vercel, CrewAI
Storage
PostgreSQL, Databases, Amazon S3 (AWS S3), NoSQL, Data Integration, Redis, MongoDB, MySQL, Data Pipelines, Database Management
Industry Expertise
Healthcare
Other
FastAPI, LangChain, RAG Architecture, AI Agents, Data Structures, Algorithms, Web Development, APIs, Hybrid Retrieval, Vector Search, Embeddings, Agentic AI, Large Language Models (LLMs), Artificial Intelligence (AI), Back-end, Software Architecture, LLM Integration, Data Engineering, ETL Pipelines, Prompt Engineering, Full-stack, Supabase, OAuth, OpenAI, AI Agent Orchestration, Data Privacy, API Integration, Generative Artificial Intelligence (GenAI), Retrieval-augmented Generation (RAG), Vector Databases, CI/CD Pipelines, Full-stack Development, Webhooks, Privacy, Security, Technical Leadership, Architecture, Integration, Third-party Integration, HIPAA, UI Development, User Interface (UI), Optical Character Recognition (OCR), Minimum Viable Product (MVP), Medical Imaging, Software as a Service (SaaS), Health, API Design, Web Application Design, Cloud Architecture, Site Reliability Engineering (SRE), Front-end Development, RESTful Web Services, Multi-agent Orchestration, Multi-agent Systems, Workflow, Front-end, Client Communication, Communication, Code Review, Screeners, Team Leadership, Interviewing, Conda, Authentication, Cloud Storage, Charts, Data Visualization, Web Scraping, AI-assisted Development, Anthropic, Data Extraction, Scraping, API Gateways, RAG Pipelines, Temporal, Agentic Coding, Agentic Workflow Design, Distributed Software, Distributed Systems, Machine Learning, Cloud Infrastructure, Life Science, Amazon Bedrock AgentCore, Agentic AI Systems, Terraform Cloud, Software Engineering, API Management, Infrastructure, Infrastructure as Code (IaC), Orchestration, Observability, LLM Agents, Bun, RAG Systems, Data Processing, Scalability, Integrations, Workflow Automation, Pinecone, Networking, Cloud, Machine Learning Operations (MLOps), EMR, Cross-Encoder Reranking, Fine-tuning, Financial Data, Cloud Networking, Large Language Model Operations (LLMOps), Fintech, DNS, Customer Support, Data Analytics, Agentic RAG Systems, browser-use, Chromium
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring