Solutions Architect @ NVIDIA · HPC · AI Infrastructure · Energy Efficiency
Santa Clara, CA
I design and optimize large-scale AI systems at NVIDIA, focusing on:
- 🏗️ HPC + AI infrastructure design (multi-node, scheduling, deployment)
- ⚡ LLM inference optimization (vLLM · TensorRT-LLM · DeepSpeed)
- 🔥 Performance-per-watt benchmarking across GPU clusters
How fast can we run — and at what energy cost?
- Multi-node GPU benchmarking across:
- vLLM · TensorRT-LLM · DeepSpeed · HuggingFace
- Throughput · Latency · Memory · Power efficiency
- Performance-per-watt analysis across workloads
- Sustainable AI system design
- Trade-offs: performance vs energy cost
- Slurm-based HPC scheduling
- Kubernetes-based AI deployment
- Hybrid cluster orchestration
- Multi-node GPU scaling
AI / ML
GPU / Systems
Infrastructure
Languages
- LLM inference engine benchmarking
- GPU power and energy-efficiency profiling
- Multi-node AI cluster performance analysis
- HPC scheduling and AI infrastructure design
- Sustainable AI system optimization
Website · LinkedIn · Google Scholar · ORCID

