Anusha Pundir
A sunlit meadow of wildflowers with a small retro computer and the name anusha pundir stitched across the grassAnusha Pundir

AI engineer building agentic LLM systems, and the products that make them useful.

Studying computer science with an AI specialization at Bennett University, graduating in 2027. Previously an AI engineering intern at JettyAI. Open to AI and full-stack internships.

  • Built an agent that reads every email, call note and meeting on a deal and proposes CRM changes (stage, amount, close date, next step, contacts, risks), each with a reason and the exact sentence it rests on. Nothing is saved until a person approves it.
  • Wrote the citation check as plain code, not another model call: a proposed change whose quote does not appear word for word in the activity it cites is struck through and cannot be approved.
  • Evaluated on 12 hand-written deals, each hiding a trap (a delay buried in a P.S., sarcasm, two people named Sam): 14 of 15 updates correct, no made-up quotes and every risk caught, against 10 of 15 for a plain prompt and 3 of 15 for keyword rules. In-sample, single run.
  • Exposed the same agent as an MCP server with tools to list deals, read one, and propose updates.

Next.js, TypeScript, Claude, MCP, Zod

  • Built a pipeline that reads a company's homepage, finds about 10 real buyers with web search, researches one recent reason to reach out to each, and drafts a short first email in the founder's voice, streamed live into an approval queue. Nothing is sent.
  • Enforced a verified-source rule: a signal reaches an email only if its URL is one the search tool actually returned, checked against the API's search results rather than the model's text, with one retry before the signal is dropped and counted on screen.
  • Kept runs bounded and safe with capped searches per step, zod-typed outputs, and a homepage fetch that blocks private addresses, follows at most 3 redirects, and stops at 1 MB or 8 seconds. A run of 10 companies takes about 100 seconds and $1.

Next.js, TypeScript, Claude, Web Search, Zod

  • Built an agent that works out what a company sells and who buys it, then ranks 10 companies that sell something different to those same buyers, each with a fit score, a one-line reason and exact quotes as evidence. Look-alikes are flagged as competitors instead of recommended.
  • Built the eval from partnerships listed on 10 real companies' public pages: Wingman found 56 of 154 real partners with no competitors in its top 10s, against 12 partners and 25 competitors for a TF-IDF similarity baseline.
  • Dropped and counted any unknown id or any quote that is not an exact substring, rather than silently fixing it, and added a live mode that ranks partners for any company URL, also available over MCP.

Next.js, TypeScript, Claude, MCP, LLM Evaluation

  • Built a tool-using agent that reads real SEC-filed contracts clause by clause, applies the five steps of ASC 606 from the seller's side, and cites the exact clauses and guidance paragraphs behind every conclusion.
  • Forced retrieval over recall: the agent starts with only a clause index and must search (a standard-library BM25), read clauses, and look up guidance before it concludes. It can only cite clauses it actually read, and invented clause ids are kept and struck through rather than silently dropped.
  • Scored the agent against hand-labelled answers and a keyword-rule baseline on six CUAD contracts, lifting the overall score from 0.78 to 0.94 and obligation F1 from 0.51 to 0.83, with grounding multiplied into the score so an answer built on invented clauses cannot score well.
  • Drafts an auditor-style Word memo with the five-step analysis and an appendix quoting every cited clause, served through a FastAPI UI that lights up each cited clause in the contract.

Claude, Tool Use, LLM Evaluation, FastAPI, Python

  • Built an agent that reads each account's usage, support, and billing timeline, recommends one action (rescue call, fix billing, upsell, expand seats, and more), and cites the exact events behind it.
  • Graded it with an eval harness against a spreadsheet-style rule baseline on 40 seeded accounts: 100% accuracy vs 62.5%, catching all $152,900 of at-risk MRR vs $133,500, including traps like a holiday usage dip and a failed card that was already fixed.
  • Forced a single zod-validated tool call per account and moved any cited event ids not on the account into a tracked hallucination count, so evidence is checked, not trusted.
  • Pushes recommendations to HubSpot as CRM tasks and exposes the whole flow as an MCP server any MCP client can call.

Next.js 16, TypeScript, Claude, MCP, LLM Evaluation

  • Built and shipped a full-stack AI voice agent that phones customers about their appointment, negotiates a new time, and commits the rebooking through live tool calls before the call ends.
  • Made double-booking impossible at the database layer with a PostgreSQL exclusion constraint (btree_gist) over time ranges. Two concurrent calls can never both win the same slot, and that is proven by concurrent-write tests rather than check-then-write application logic.
  • Hardened the webhook path with HMAC-SHA256 signature verification, idempotent event dedupe, and fast-200 async processing, so duplicate or forged deliveries cannot corrupt appointment state.
  • Ran post-call extraction through Claude Haiku with strict JSON output and a retry-then-degrade path, keeping tool-committed outcomes authoritative so a parse failure never overwrites a booking.
  • Deployed on Cloud Run and Cloud SQL via multi-stage Docker and Cloud Build, with Secret Manager credentials, Clerk auth, Google OAuth, and versioned Drizzle migrations.

Next.js 16, TypeScript, PostgreSQL, Retell AI, Google Cloud

  • Built a self-improving agent on a LangGraph state machine that grounds answers in a local RAG knowledge base, scores its own output, then critiques and revises in a closed loop with guaranteed termination.
  • Designed a hybrid harness that runs free deterministic checks (grounding, coverage) on the full set and reserves a sampled LLM-as-judge for correctness, completeness, and clarity. Validated the judge against human-labelled data first, reaching 1.00 ranking accuracy with clean separation between strong and weak answers (0.93 vs 0.30 mean score).
  • Routed high-volume traffic to a free local model (Ollama qwen2.5) and kept a cached, sampled Claude Haiku judge as the only paid path, running the default pipeline at near-zero cost. Covered by 88 passing tests.

LangGraph, RAG, LLM-as-Judge, Ollama, Claude

  • Authored an MCP server from scratch (FastMCP, stdio JSON-RPC) exposing 11 typed tools that let any MCP-compatible app search and fetch tech-community threads across Hacker News, Lobsters, and Reddit.
  • Added a cross-source synthesis tool that fans out concurrently, dedupes cross-posts by normalized URL, and merges sources into one ranked digest that tolerates partial source failure.
  • Orchestrated a LangGraph ReAct agent that consumes the server as a second MCP client over stdio, with a provider-swappable model layer (Claude Haiku or local Ollama) and an LLM-as-judge that caught the local 7B model fabricating URLs and dates where Claude stayed grounded.

MCP, FastMCP, LangGraph, ReAct, JSON-RPC

  • Trained EfficientNet-B0 on the HAM10000 dataset (10,000+ images, 7 classes) with two-stage fine-tuning, reaching 75% validation macro-F1 on an imbalanced set using Test-Time Augmentation.
  • Built a FAISS RAG pipeline over sentence-transformers to retrieve medically relevant knowledge, and fine-tuned TinyLlama 1.1B Chat with QLoRA (4-bit, roughly 2% trainable parameters) to generate structured clinical reports.
  • Deployed backend inference on Hugging Face Spaces and the Next.js frontend on Vercel.

PyTorch, EfficientNet, QLoRA, FAISS, Next.js

  • Built a full-stack AI SaaS platform with authentication, secure multi-tenant data models, and a real-time analytics dashboard.
  • Engineered a structured Google Gemini inference pipeline that turns unstructured dream text into validated JSON for symbols, emotions, and personalized insights, then surfaces it on an anonymous social feed with likes and comments.
  • Designed type-safe Next.js API routes over shared TypeScript models, backed by a Supabase PostgreSQL database with Row-Level Security enforcing per-user data isolation.

Next.js, Supabase, Google Gemini, PostgreSQL, TypeScript

  • Designed and trained an EfficientNet classifier that sorts disaster images into none, mild, and severe, reaching 84% accuracy.
  • Built the full severity-prediction pipeline (preprocessing, augmentation, inference) and integrated it into the team's disaster-management app for real-time hazard assessment.
  • Placed 2nd of 150 teams in the AI/ML track of a college showcase.

EfficientNet, PyTorch, Computer Vision

  • Fine-tuned a pretrained ResNet18 in PyTorch to classify brain MRI scans into glioma, meningioma, pituitary tumor, or no tumor.
  • Shipped it as a Streamlit app that takes an uploaded scan and shows the predicted class with its confidence and the full probability distribution.

PyTorch, ResNet, Transfer Learning, Streamlit

  • Designed and built a social platform where builders post ideas, showcase projects, and meet collaborators through swipe-based matching and casual virtual coffee chats.
  • Built the front end in React, TypeScript, and Vite with Framer Motion animations and a themeable styled-components design system, including a full dark mode.

React, TypeScript, Vite, Framer Motion