PORTFOLIO 2026

Samarth
Agarwal

AI Engineer · LLM Evaluation · Edge Systems

Samarth.
Hi, I'm

Samarth
Agarwal.

AI Engineer Backend · Data Pipelines · LLM Systems

I build software, data, and AI systems that hold up in production — reliable backends, measurable pipelines, and applications that actually ship.

2+
Years Experience
6+
Projects Shipped
AI + Data
Specialist
Open to opportunities — MS Data Science, Stony Brook ’26 📍 New York, USA
Samarth Agarwal
Who I Am

About Me

I'm an engineer working across AI, software, and data science, with an MS in Data Science from Stony Brook University and 2+ years of enterprise software experience at Genpact building Salesforce CRM architectures for global clients.

My focus is building systems that are measurable and production-ready — whether that's a backend service, a data pipeline, or an LLM application. I've shipped evaluation frameworks benchmarking GPT-4, Gemini, and LLaMA across normalized metrics, edge-native apps with recursive failover, real-time serverless dashboards, and autonomous multi-agent systems. I care about the parts that decide whether software survives contact with production: clear metrics, sane architecture, and reliability under load.

MS Data Science, Stony Brook University — New York
Ex-Genpact: Salesforce CRM for GE Renewable Energy & Huntsman
GenAI Prompt Engineer Intern @ ARMA AI Labs, California
GSolve.ai ML Hackathon 2023 Winner
Selected Work

Projects

Backend services, data pipelines, and AI systems — each with metrics to back it up.

PrepGenie interview coach interface
01 — Edge AI

PrepGenie: LLM Interview Coach

Edge-native interview platform with resume-aware questions and a Gemini → Groq → Mistral recursive failover chain. 120ms average TTFB, sub-100ms cold starts, 99.9% uptime.

Pyodide/WASMEdge NativeLLM-as-JudgeFailover
LLM evaluation framework dashboard
02 — LLM Pipeline

LLM Content Evaluation Framework

Unified DeepEval pipeline benchmarking GPT-4, Gemini 2.5, and LLaMA 3-70B across 7 normalized metrics — revealing 96% GPT-4 accuracy and 40x Gemini cost-efficiency for dynamic routing.

DeepEvalPythonBLEU/BERTScoreG-Eval
AI penetration testing framework
03 — Security AI

Autonomous AI Penetration Testing

Multi-agent cybersecurity framework orchestrating recon, vulnerability assessment, and exploitation agents. 95% SQLi / 90% XSS detection, saving $10K+/yr in scanner costs.

MCPLlama 3.3 70BOWASPDocker
Real-time stocks dashboard
04 — Serverless

Real-Time Stocks Dashboard

Global serverless financial dashboard on V8 isolates with on-demand caching — sub-60s data latency, 90% fewer API calls, zero idle compute cost.

React 18Cloudflare D1/R2D3.jsFinnhub
GestureSynth hand gesture to MIDI
05 — Computer Vision

GestureSynth: Gestures to MIDI

Vision-based system translating hand gestures into live MIDI at 30 FPS with <80ms latency, plus an LLM-powered adaptive tuning module that speeds up learning by 40%.

MediaPipeOpenCVMIDIPyQt5
Manifesto Explainer NLP platform
06 — NLP

Manifesto Explainer

NLP platform turning dense political manifestos into accessible insights — sentiment analysis, Llama-3 summarization, party symbol detection, and multilingual auto-translation.

Spring BootNLTKLlama-3Groq
Where I've Worked

Experience

2025

Generative AI Prompt Engineer Intern

ARMA AI Labs — California, USA

  • Built a DeepEval-powered evaluation framework establishing reproducible QA pipelines across 7 normalized metrics.
  • Benchmarked GPT-4, Gemini, and LLaMA on 20+ real-world articles, revealing accuracy vs. cost-efficiency trade-offs.
  • Delivered a production model-selection pipeline with dynamic LLM routing.
2022–24

Senior Associate Software Developer

Genpact — India

  • Reduced manual data entry 60% for enterprise clients (GE Renewable Energy, Huntsman) with scalable Salesforce CRM architectures.
  • Integrated Salesforce with ERP systems over SFTP for low-latency global data consistency.
  • Shipped Lightning Web Components dashboards and resolved production defects.
2021

Machine Learning Engineer Intern

TannMann Foundation — Remote

  • Achieved 96% validation accuracy in brand-logo detection with a TensorFlow pipeline using layer freezing and augmentation.
  • Mitigated early-epoch overfitting for production-ready performance under varying lighting.
Proof of Work

Credentials

🏆
SAP Certified — Generative AI Developer
SAP
Verify on Credly →
☁️
Salesforce Certified AI Associate
Salesforce
Verify on Trailblazer →
🥇
GSolve.ai ML Hackathon 2023 — Winner
Genpact Internal Hackathon
📊
AWS Solutions Architecture · BCG Data Science
Virtual Experience Programs
📈
Purchase Pattern Prediction — Rank 123 / 2000+
Univ.ai Competition
🎓
KPMG Data Analytics · Goldman Sachs SWE
Virtual Experience Programs