I am a Generative AI Engineer with 2.5 years of production experience architecting enterprise-grade LLM applications, multimodal AI pipelines, and high-performance inference platforms. My work spans fine-tuning foundation models (QLoRA/SFT), optimizing latency and compute costs via AutoRound W4A16 quantization and vLLM multi-GPU serving, and building autonomous agentic workflows.
- Current Focus: Fine-tuning VLMs/LLMs, accelerating high-throughput vLLM serving, and optimizing Graph-RAG architectures.
- Impact Highlight: Creator of a quantized model on Hugging Face with 100K+ downloads.
- Core Philosophy: "Optimization is not just about speed; it's about making advanced AI architecturally sustainable."
- Fine-Tuning: Domain adaptation using QLoRA, SFT, Unsloth, and Hugging Face PEFT.
- Quantization: W4A16/INT4 model compression using AutoRound, AWQ, and GPTQ.
- High-Throughput Serving: Production multi-GPU deployment via vLLM, cutting P95 latency by 50%.
- Hybrid Retrieval: Orchestrating BM25, semantic vector search, and Reciprocal Rank Fusion (RRF) in Qdrant and Neo4j.
- Graph-RAG Architectures: Implementing LightRAG and property graph strategies using Qdrant and Neo4j for dual-level knowledge extraction.
- Observability & Evals: Production tracing, prompt evaluation, and hallucination tracking with Langfuse and Ragas.
- Stateful Agents: Designing autonomous multi-agent loops, self-healing tool execution, and session memory using LangGraph.
- Multimodal Pipelines: Real-time object detection (YOLOv12) and speech-to-text processing (Whisper ASR) for automated verification platforms.
- Agentic LightRAG for Executive Audit Review: Built a LangGraph multi-agent loop with LightRAG, Qdrant and Neo4j acting as a document, vector, and property graph store for long-form audit analysis.
- Agentic Hybrid RAG for Regulatory Compliance: Engineered a self-healing RAG platform in Qdrant with BM25 + Vector RRF search, intercepting DB exceptions for auto-correction and 99.9% uptime.
- AI-Driven vKYC Automation System: Built an end-to-end multimodal pipeline featuring YOLOv12 object detection, Whisper ASR, and FastAPI to streamline document and audio verification.
- Published the
gemma-4-E4B-it-W4A16-AutoRound-GPTQquantized model, surpassing 100K+ downloads.
- Contributed official integrations for the Google Gemini API and PostgreSQL multi-tenant workspace isolation.
- PRs: #2538 | #2556 | #2615
- Identified critical W4A16 serialization bugs in AWQ export for Qwen3-VL models to improve vLLM inference stability.
- Issue: intel/auto-round #1377
| Category | Tools & Technologies |
|---|---|
| GenAI & Agents | LangGraph, LangChain, LightRAG, Ragas, Langfuse |
| Fine-Tuning & Quant | vLLM, AutoRound, QLoRA, PEFT, Unsloth, Hugging Face |
| Vector DBs & Search | Qdrant, Neo4j, Graph-RAG, BM25 + Semantic (RRF) |
| Multimodal & CV | YOLOv12, Whisper ASR, OpenCV, PyTorch, CUDA |
| Backend & Cloud | Python, FastAPI, Docker, AWS (EC2, ECS, ECR, S3), RunPod |
- LinkedIn: vishva-r
- Instagram: @justt_vishva
- Portfolio: GitHub Repositories
⭐️ Building the infrastructure that makes AI smarter, faster, and more accessible.


