AI infrastructure engineer — I make large language models run fast on hardware that shouldn't be able to run them.
9 years building and deploying mission-critical systems. Currently focused on LLM inference optimization: quantization, tensor parallelism, and serving 30B-class models fully offline on constrained GPUs.
| Project | What it is |
|---|---|
| enterprise-airgapped-llm | Production self-hosted LLM platform for air-gapped environments — 30B MoE (AWQ 4-bit) across dual Turing GPUs via tensor parallelism, 80–120 tok/s, sub-500 ms TTFT, zero cloud dependency. |
| YaYan-AI | Offline multi-dialect speech-intelligence system — 22 Chinese dialects + 40 languages, speaker diarization, character-level timestamps, LLM-assisted correction. |
| yolov5-parallel | Multi-GPU parallelized YOLOv5 training and inference pipeline. |
| MultiBodyCuboids | M.S. thesis — multi-body motion segmentation for arbitrary numbers of disordered 3D point sets. |
| openclaw-agents | Autonomous LLM agents — OODA decision support, scheduled briefings, human-in-the-loop safety gates. |
LLM Inference vLLM · quantization (AWQ / INT4) · tensor parallelism · KV-cache tuning · FlashInfer · TensorRT-LLM
ML & Languages PyTorch · HuggingFace · CUDA · Python · C++
Infrastructure Docker · Linux · Kafka · Elasticsearch · PostgreSQL · AWS
M.S. CS @ NYCU · AWS Solutions Architect – Associate · CEH · CompTIA Security+ / Network+ · HITCON white-hat · CVE research
[email protected] — open to AI infrastructure / LLM inference / ML systems roles (2027)
