AI engineer ~ I build LLM systems that prove their work. Prefers primitives >>> frameworks
🟢 status — 🚧 building a review-gated PR risk agent · writing at byraj.dev · India, remote-friendly
Upload an agreement and its amendments; get obligations extracted (owners, deadlines, penalties), an explained diff, and every change mapped to the clauses it touches. Each statement cites its source text, low-confidence items queue for human review, and every LLM call's cost is persisted per tenant.
1.0 citation validity · 0.90 extraction precision · 0.96 recall@10
FastAPI pgvector + Postgres FTS (RRF) Postgres job queue (SKIP LOCKED) AWS ECS/Fargate Terraform Next.js
Standalone OTP timer for React Native.
- No citation, no claim. LLM output must trace to source text; invalid items get one repair pass, then drop.
- Evals before vibes. Precision/recall on frozen sets; CI fails when quality regresses.
- Budgets are product features. Per-run caps, per-tenant limits, cost persisted per call.
- Humans hold the write. Agents draft; people approve.
- Agent harnesses — tool loops, checkpointing, interrupt/resume when the model misbehaves.
- Multi-agent systems — where committees beat one loop, and where they just multiply the bill.
- Harder evals — frozen query splits, Recall@5 / MRR, regression gates instead of vibes.
- Multimodal retrieval — video/images as a searchable index
- Local-first agent spend — metering and capping coding agents across vendors with a SQLite ledger and hooks.
AI/LLM ~~> RAG (pgvector + FTS, RRF) · evals · structured outputs · LangGraph agents · MCP · LLM cost tracking
Backend ~~> Python · FastAPI · Pydantic v2 · SQLAlchemy · PostgreSQL · Node.js
Frontend ~~> React · Next.js · TypeScript · Redux Toolkit · TanStack Query · Tailwind
Infra ~~> AWS (ECS/Fargate, RDS, S3) · Terraform · Docker · GitHub Actions · Sentry · CloudWatch
Full-stack at product startups (Innovaccer, Geeks Invention). Now building AI systems end to end.
Off the clock: singing and tactical shooters <valorant,CS 🎮>



