Skip to content
View wu840407's full-sized avatar
💭
AI infra & speech-AI engineer · self-hosted LLM · ASR · PyTorch/CUDA
💭
AI infra & speech-AI engineer · self-hosted LLM · ASR · PyTorch/CUDA

Block or report wu840407

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
wu840407/README.md

Hi, I'm Roger (ChengRung Wu) 👋

AI infrastructure engineer — I make large language models run fast on hardware that shouldn't be able to run them.

9 years building and deploying mission-critical systems. Currently focused on LLM inference optimization: quantization, tensor parallelism, and serving 30B-class models fully offline on constrained GPUs.


🚀 Featured Projects

Project What it is
enterprise-airgapped-llm Production self-hosted LLM platform for air-gapped environments — 30B MoE (AWQ 4-bit) across dual Turing GPUs via tensor parallelism, 80–120 tok/s, sub-500 ms TTFT, zero cloud dependency.
YaYan-AI Offline multi-dialect speech-intelligence system — 22 Chinese dialects + 40 languages, speaker diarization, character-level timestamps, LLM-assisted correction.
yolov5-parallel Multi-GPU parallelized YOLOv5 training and inference pipeline.
MultiBodyCuboids M.S. thesis — multi-body motion segmentation for arbitrary numbers of disordered 3D point sets.
openclaw-agents Autonomous LLM agents — OODA decision support, scheduled briefings, human-in-the-loop safety gates.

🛠 Tech Stack

LLM Inference vLLM · quantization (AWQ / INT4) · tensor parallelism · KV-cache tuning · FlashInfer · TensorRT-LLM

ML & Languages PyTorch · HuggingFace · CUDA · Python · C++

Infrastructure Docker · Linux · Kafka · Elasticsearch · PostgreSQL · AWS

🎓 Background

M.S. CS @ NYCU · AWS Solutions Architect – Associate · CEH · CompTIA Security+ / Network+ · HITCON white-hat · CVE research

📫 Contact

[email protected] — open to AI infrastructure / LLM inference / ML systems roles (2027)

GitHub stats

Pinned Loading

  1. YaYan-AI YaYan-AI Public

    Fully-offline multi-dialect speech-intelligence system — 22 Chinese dialects + 40 languages, speaker diarization, character-level timestamps, LLM-assisted correction.

    Python 1

  2. enterprise-airgapped-llm enterprise-airgapped-llm Public

    Production reference architecture for self-hosted LLM in air-gapped enterprise environments. Dell R740 + Turing GPUs + vLLM + Qwen3-Coder + AD/LDAPS.

    Shell

  3. openclaw-agents openclaw-agents Public

    Four standalone autonomous LLM agents — OODA strategy support, scheduled briefings, content pipelines with human-in-the-loop safety gates.

    Python

  4. MultiBodyCuboids MultiBodyCuboids Public

    M.S. thesis — multi-body motion-segmentation architecture for arbitrary numbers of disordered 3D point sets (PyTorch, C++).

    Python

  5. yolov5-parallel yolov5-parallel Public

    Multi-GPU parallelized YOLOv5 training and inference pipeline — data-parallel scaling and throughput tuning.

    Python 2

  6. wu840407.github.io wu840407.github.io Public

    Personal site — projects, writing, and contact.

    HTML