Mayank Sharma

About Me

Senior Data Scientist / Applied AI Engineer with 5+ years building production AI — NLP pipelines, RAG systems, healthcare AI. The question I keep coming back to is how do you make this actually work in production, not just in a notebook.

M.Tech CSE - Data Science, IIT Jammu · India

My Story: Building Production AI Systems

My M.Tech in Data Science at IIT Jammu taught me to think rigorously about evidence. At nference.ai, I built production NLP and multimodal AI for healthcare. What I learned: a model that's technically impressive but slow, expensive, or unreliable in production isn't actually done.

At Lamipak, I designed and built LamiOps (RAG-powered assistant, 10K+ daily users) and LamiTracker (200+ sources, 18+ countries) — architecture, retrieval design, evaluation frameworks, and the backend underneath all of it. That's the work I want to keep doing at a senior technical level: owning hard AI systems end to end, not handing the interesting parts to someone else.

My Journey: Data Science Foundations → Production AI Systems → Senior AI/ML Engineering

From rigorous data science foundations to owning production AI systems end to end, at increasing scale and complexity.

Data Science Foundations

M.Tech in Data Science, IIT Jammu (CGPA: 8.70/10). Built deep foundations in NLP, Deep Learning, LLM, and GenAI. Mentored 100+ students as Teaching Assistant.

Production AI Systems

At nference.ai, built NLP pipelines for healthcare - knowledge distillation (8× smaller, 9× faster), clinical extraction with BioClinicalBERT, multimodal AI. Learned: technical impressiveness means nothing if users don't want it.

Owning Systems End-to-End

At Lamipak, built LamiOps (10K+ daily users) and LamiTracker (200+ sources, 18+ countries) - architecture, AI pipeline, backend, deployment. This is where I proved I could own a hard technical system solo, at production scale.

Impact & Achievements

10K+ Daily Users

LamiOps - RAG-powered enterprise assistant, owned end-to-end. Shipped with a quality eval framework (Groundedness, Hallucination Rate, MRR). Retrieval time: hours → seconds.

200+ Sources, 18+ Countries

LamiTracker - Regulatory Intelligence platform. LLM-based extraction, full AI + backend build. Reduced manual R&D tracking effort by 95%.

Evidence-Based Decisions

Design metrics that reveal what's actually happening, not just what looks good in a report. Published research (AAAI 2024, NEJM AI).

Full-Stack Technical Depth

Built the systems myself, end to end - model, retrieval, backend, deployment - so I know exactly what's feasible, what's costly, and where the real trade-offs are.

Problem-First Engineering

Good systems start from real constraints, not impressive technology for its own sake. This shaped everything from healthcare AI at nference to the architecture decisions at Lamipak.

Teaching & Mentorship

Teaching Assistant at IIT Jammu (2019–2021), mentoring 100+ students. Explaining clearly is as important as building well.

Beyond the Code

Built on hard work, powered by a touch of talent, sprinkled with humor, and fueled by infinite nerdiness.

Staying Active

When I'm not debugging models, you'll find me balancing code with cardio: gym sessions for strength, running for endurance, and cycling for exploring new trails. Movement keeps the mind sharp and the ideas flowing.

Fledgling Bookworm

A growing passion for reading keeps me curious beyond technical papers. Whether it's exploring new ideas, learning from diverse perspectives, or simply unwinding with a good story, books are becoming an essential part of my routine.

Grounded in Mindfulness

Meditation helps me stay centered amidst the fast-paced world of AI. It's not just about relaxation, it's about clarity, focus, and maintaining perspective when tackling complex problems.

One thing to remember about me: my humility isn't flattery, it's a value shaped by a humble upbringing.

Skills & Expertise

ML Systems Architecture

RAG & Retrieval Architecture Hybrid Search & Vector Retrieval Model Serving & Inference Optimization Multilingual LLM Systems Latency/Cost/Accuracy Trade-offs Evaluation Framework Design Scalable Backend Design System Design for Production AI

AI & Data Engineering

Data Pipeline Design LLM & NLP Pipelines Knowledge Distillation Model Fine-tuning Python & Production Code Backend & API Development MLOps Fundamentals SQL & Data Analysis

AI & Data Science Expertise

Deep Learning & AI Systems Natural Language Processing Clinical NLP (BioClinicalBERT, SciBERT) Model Evaluation & Benchmarking Data Architecture & Pipelines Python & Data Science Tools Research & Publication (AAAI, NEJM AI) Statistical Rigor & Experimentation