I'm an AI/ML Engineer specializing in Generative AI, LLM systems, Agentic AI, RAG, and MLOps, building production-grade AI solutions across enterprise environments. My work sits at the intersection of AI engineering, software engineering, cloud infrastructure, and measurable business impact.
At Uber, I contribute to real-time ML inference infrastructure supporting ride-demand forecasting using AWS SageMaker, Kubernetes, Python, and LangGraph. I also develop model-routing workflows that evaluate quality, latency, and inference cost across foundation and fine-tuned models.
Previously at KPMG, I delivered enterprise AI systems that processed 40,000+ legal clause types, reclaimed approximately 1,900 attorney-hours per quarter, reduced auditor search time from 45 minutes to under 4 seconds, lowered model degradation incidents from 9 per quarter to 1, and reduced deployment lead time from 8 days to 18 hours.
My projects focus on Agentic AI, LLM evaluation, observability, NL-to-SQL, OCR automation, forecasting, and computer vision. I'm especially interested in evaluation, hallucination mitigation, model routing, latency, cost optimization, responsible AI, and reliable production deployment.
When I'm not building things: cricket, gym, and the occasional sketch.
June 2026 – Current | USA (Remote)
- Contributing to a real-time ML inference platform supporting ride-demand forecasting workloads using AWS SageMaker and Kubernetes.
- Developing model-routing workflows in Python and LangGraph to evaluate quality, latency, and inference cost across foundation and fine-tuned models.
June 2022 – June 2024 | India
- Delivered NLP, RAG, MLOps, anomaly detection, and LLM evaluation solutions across enterprise environments.
- Built systems using BERT, GPT-3, Azure OpenAI, Pinecone, Azure ML, MLflow, XGBoost, Docker, and GitHub Actions.
- Helped reduce manual review effort, search time, model degradation incidents, and deployment lead time through production AI automation.
| Area | Tools |
|---|---|
| Large Language Models | GPT-4 · LLaMA · Mistral · Google Gemma · Hugging Face Transformers · Ollama |
| Agentic AI | LangGraph · LangChain · LlamaIndex · Multi-Agent Systems · Tool-Augmented LLMs |
| RAG & Retrieval | Pinecone · Weaviate · FAISS · pgvector · Embedding Models · Hybrid Search · Semantic Search |
| LLM Engineering | Fine-Tuning · SFT · RLHF · LoRA · QLoRA · Prompt Engineering · LLM Evaluation · LLM-as-a-Judge |
| Deep Learning & ML | PyTorch · TensorFlow · Scikit-learn · XGBoost · BERT · Transformers · NLP · Time-Series Forecasting |
| Computer Vision | YOLOv8 · Deep SORT · OpenCV · dlib · MTCNN · EasyOCR · face-recognition |
| Cloud — GCP | Vertex AI · Cloud Run · BigQuery · Artifact Registry · Secret Manager |
| Cloud — AWS | SageMaker · Bedrock · S3 · IAM · ECR |
| Cloud — Azure | Azure Machine Learning · Azure OpenAI |
| MLOps | MLflow · Kubeflow · Docker · Kubernetes · Terraform · GitHub Actions · Model Monitoring |
| Backend | FastAPI · Streamlit · REST APIs · Microservices |
| Databases & Data | PostgreSQL · Apache Spark · ETL Pipelines · SQL |
| Languages | Python · SQL · Bash · YAML · JavaScript |
| AI Governance | Responsible AI · Bias Detection · Hallucination Mitigation · AI Safety · Guardrails |
|
A LangGraph state machine that handles the full GitHub issue lifecycle without human intervention — classification, assignment, SLA enforcement, audit logging, and auto-close on stale issues. Business impact: Returns 5–10 hrs/week of engineering overhead back to the team.
|
Polls LLM trace data from Arize every minute, evaluates each response with a Vertex AI judge model, and deploys the full pipeline on GCP Cloud Run via Terraform. Multi-cloud: built on AWS SageMaker, runs on GCP. Business impact: Catches model degradation in minutes, not when a customer complains.
|
|
Schema-grounded RAG pipeline that lets anyone query a PostgreSQL database in plain English. The LLM receives live table definitions, foreign keys, and business rules before generating SQL. Runs entirely locally with Ollama + Gemma — no data leaves the machine. Business impact: Answers in 30s what used to require a Jira ticket and a day's wait.
|
Full-stack vending operations platform. OCR pipeline reads supplier receipts and updates inventory automatically. Rolling 7-day demand forecast flags machines before they stock out. Profit calculated from actual invoice costs, not estimates. Business impact: Eliminates manual data entry + targets 15–25% revenue lost to stockouts.
|
|
Upgraded an open-source tracking system — replaced Darknet/TF 1.14 with YOLOv8, fixed a concurrency bug by giving each camera its own Deep SORT instance, and added vehicle intelligence: color detection, plate OCR, and type classification. Business impact: One operator monitoring 10+ live feeds with automated event detection.
|
Two detection pipelines (Haar Cascade + MTCNN), LBPH recognition, and real-time color-coded alerts based on criminal record lookup. Undergraduate thesis published at BVRITHCON-2023 (Springer). Research: Compared traditional vs deep learning face detection on live video.
|
Ongoing Research: CNN-Based Autism Detection via 4D fMRI — 3D CNN on resting-state neuroimaging to classify ASD vs. neurotypical subjects. Research project at UNT. (Repo coming soon)
| Institution | Period | Result | |
|---|---|---|---|
| MS Computer Science | University of North Texas, TX | 2024 – Present | GPA 3.8 / 4.0 |
| BE CSE (AI & ML) | GIET, JNTU Kakinada, India | 2020 – 2024 | CGPA 8.0 / 10 |
Published Research
Efficient Person Identification using Artificial Neural Networks International Conference BVRITHCON-2023 · Published by Springer doi.org/10.1007/978-981-95-0144-1_25
Certifications
- Microsoft Azure AI Engineer Associate — Microsoft (Apr 2023)
- AWS Academy: Machine Learning Foundations — Amazon Web Services (Jan 2023)
- AWS Academy: Cloud Architecting — Amazon Web Services (Jan 2023)
- AWS Academy: Cloud Foundations — Amazon Web Services (Nov 2022)
- Python for Data Science — IBM (Jun 2023)
- MTA: Introduction to Programming Using Python — Microsoft (Jun 2022)
I'm actively looking for AI/ML engineering roles — full-time or internship. If you're working on something in LLMs, agentic AI, computer vision, or MLOps, I'd love to connect.


