AI Engineer · GenAI & Agentic Systems · LLM Evaluation
The model proposes.
My code decides.
Five production systems shipped across healthcare, fintech, and e-commerce, plus hands-on LLM evaluation and observability work on a live agentic RAG product — each built so the model can reason and draft, but the decision that matters runs through deterministic, auditable code.
LLM evaluation and answer-quality work for a live agentic RAG learning platform. Reviewed ~250 cases by hand, redesigned an accuracy check that was measuring the wrong part of the pipeline, and traced a duplicate-content retrieval bug confirmed across 8 production sessions.
Multi-agent pipeline — scrape → critique → prescribe — that audits product listings and returns specific description fixes. Scoring stays in code, not the LLM, and work is split into Inngest steps to beat a hard 10-second serverless timeout.
Flags suspicious transactions and explains why in language a human reviewer can act on, instead of returning a black-box risk score. A three-agent LangGraph workflow — detection, investigation, decision — with risk scoring kept rule-based and separate from the LLM for deterministic, auditable verdicts.
Voice-guided mental-performance companion that adapts its coaching tone to how the user is feeling. A single GPT-4o call — not a multi-agent chain — keeps every response under a hard 10-second timeout, with a validation layer that rejects unsafe or off-tone output before it reaches the user.
LLM evaluation & observability for a live agentic RAG product.
Backend and architecture lead across 3 concurrent production AI systems.
Conversational AI systems on production FastAPI backends.
HIPAA-compliant audio-to-record pipeline on GCP. MedGemma extracts structured clinical fields from consultation audio; Gemini summarizes. ~60-second turnaround on a 15-minute recording.
6-stage financial ML pipeline: news ingestion → FinBERT sentiment → Z-score anomaly → GPT-4 alert summaries. JWT auth, watchlist API, React/TypeScript frontend.
6-layer agentic system — Planner / Executor / Synthesizer agents via LangGraph. ChromaDB RAG + Tavily web search. Built bottom-to-top in 13 days solo.
Production-quality Q&A bot with MongoDB Atlas vector storage, SentenceTransformers (384-dim), Azure OpenAI GPT-3.5 Turbo. 100% test accuracy, 69.2% avg token overlap.
What people I've worked with say
He developed our evaluation system from the ground up, creating an agent-as-judge pipeline that analyzes traces for tool usage, latency, cost, and failure reasoning, and incorporates human-reviewed golden datasets and in-depth investigations.
Jibin has been the leader for entire LLM development for Mindgym and has demonstrated strong ownership and capability to develop backend from zero to 1 and has been awarded Tier 3: AI Rising Star Award!
He was instrumental in architecting our 'Data Integrity Layer,' moving the project beyond simple LLM prompts into a strictly typed, contract-first system.
Open to full-time AI Engineer roles and interesting collaborations.
Answers are drawn only from this portfolio's structured project and experience data.