Skip to content
Nikhil Kumar Reddy

Nikhil Kumar Reddy

AI/ML Engineer · Washington, DC

nikhilkumarreddy0@gmail.com+1 (703) 906-6508

Download PDF ↓

Summary

AI/ML Engineer with 3+ years building production LLM systems, RAG pipelines and multi-agent applications across healthcare and enterprise. Track record of taking projects from prototype to production — fine-tuning BERT and Llama models, designing evaluation frameworks with RAGAS, and shipping systems that clinical and engineering teams use daily.

Experience

AI/ML EngineerMilken Institute School of Public Health, GWU

Jan 2025 – May 2026

  • Built a LangChain RAG pipeline over 15K behavioural health records using GPT-4, measuring retrieval quality with RAGAS rather than eyeballing outputs — a 34% improvement over the keyword baseline, which let clinical teams handle 3x more patient queries.
  • Fine-tuned BERT via LoRA on 5K patient survey responses for clinical text classification, taking held-out accuracy from 71% to 87%, then sat with clinical staff to validate outputs and document the edge cases the metric hid.
  • Containerised FastAPI inference endpoints on AWS SageMaker with Docker, adding input validation and output safety constraints before anything reached a decision-support workflow. Average response latency fell 45%.
  • Automated retraining with MLflow, drift detection and CI/CD guardrails, replacing an ad-hoc process and sustaining 78% production uptime across model versions.

Data Scientist (NLP/ML)Data Science for Sustainable Development

Aug 2024 – Dec 2024

  • Deployed FAISS semantic search with anomaly validators across JSON, CSV and Parquet census data, tracing cross-format quality gaps with the data engineering team until the pipeline reached full compliance.
  • Designed automated feature engineering workflows in PySpark on Databricks, using LangGraph agents to surface and test candidate features — 35% more modelling throughput for a policy research team.
  • Shipped Streamlit dashboards backed by Azure OpenAI that translated model output into plain-language summaries for non-technical leadership, cutting reporting prep by 40%.

Data ScientistCogno AI

May 2022 – Jul 2024

  • Built a production RAG system on OpenAI embeddings and Pinecone across 10K e-commerce SKUs, taking average query latency from 3.2s to 0.8s — the difference between search that felt broken and search that felt instant.
  • Trained an XGBoost demand forecasting model on retail POS data, 17% more accurate than the ARIMA baseline, reducing annual inventory overstock.
  • Set up the team's MLOps foundation — MLflow, Git, Docker, experiment tracking, A/B model evaluation and observability dashboards — cutting the iteration cycle by 45% and becoming the standard the team kept.

Skills

Ingestion & retrieval
LangChain · LlamaIndex · FAISS · Pinecone · ChromaDB · Qdrant · BM25 hybrid search · BGE / MiniLM embeddings · Cohere Rerank · PySpark · Databricks · Snowflake · dbt
Modelling & generation
PyTorch · TensorFlow · Hugging Face Transformers · BERT · Llama · LoRA / QLoRA / PEFT · Claude API · OpenAI API · LangGraph · CrewAI · AutoGen · MCP · XGBoost · scikit-learn
Safety & guardrails
Input validation · Prompt-injection filtering · PII redaction · Output verification · Hallucination checks · Sandboxed execution (E2B) · Hardcoded escalation rules
Evaluation & observability
RAGAS · DeepEval · LangSmith · MLflow · A/B testing · Drift detection · SHAP · Held-out evaluation sets · Model monitoring
Deployment & the human loop
Docker · Kubernetes · FastAPI · AWS SageMaker · Azure ML · GCP Vertex AI · Airflow · CI/CD · Railway · PostgreSQL · MongoDB

Selected projects

  • Post-Discharge Voice Agent

    Calls patients 72 hours after discharge — and escalates on hardcoded clinical rules the AI is never allowed to touch.

    Sensitivity across 92 evaluation cases: 100% · False-escalation rate: 0% · Passing tests: 418

  • Inbound Clinic Voice Agent

    A real phone number that answers, books an appointment end to end, and hangs up.

    Task completion: 88% → 100% · Evaluation cases passing: 19/19 · End-to-end turn latency: P50 4.1s · P95 5.2s

  • SEC Finance RAG

    Citation-backed answers over 25,000 SEC filings, with the guardrails that make them trustworthy.

    Median query latency, cache miss: <500ms · RAGAS faithfulness: >0.80 · DeepEval hallucination rate: <0.20

  • Vectorless RAG

    Retrieval over SEC filings with no embeddings and no vector database at all.

    Sections retrieved per query: top-K → 1–3 whole · Response time: 9–12s · Filings indexed: 6 (FY2020–25)

  • AnalystAI

    Seven specialised agents that take a spreadsheet and a plain-English question to a boardroom answer.

    Specialised agents in the pipeline: 7 · Generated code isolation: E2B sandbox · Supported sources: CSV · Excel · SQL · Sheets · REST

Education

  • MS, Data AnalyticsThe George Washington University · Aug 2024 – May 2026
  • BTech, Computer Science & EngineeringDayananda Sagar University