DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
-
Updated
Aug 11, 2026 - Python
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Optimizing inference proxy for LLMs
Decentralized deep learning in PyTorch. Built to train models on thousands of volunteers across the world.
Run Mixtral-8x7B models in Colab or consumer desktops
【TMM 2025🔥】 Mixture-of-Experts for Large Vision-Language Models
PyTorch Re-Implementation of "The Sparsely-Gated Mixture-of-Experts Layer" by Noam Shazeer et al. https://arxiv.org/abs/1701.06538
Codebase for Aria - an Open Multimodal Native MoE
Tutel MoE: Optimized Mixture-of-Experts Library, Support GptOss/DeepSeek/Kimi-K2/Qwen3 using FP8/NVFP4/MXFP4
⛷️ LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training (EMNLP 2024)
SMT: The Surrogate Modeling Toolbox
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
A Pytorch implementation of Sparsely-Gated Mixture of Experts, for massively increasing the parameter count of language models
From scratch implementation of a sparse mixture of experts language model inspired by Andrej Karpathy's makemore :)
A TensorFlow Keras implementation of "Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts" (KDD 2018)
中文Mixtral混合专家大模型(Chinese Mixtral MoE LLMs)
A library for easily merging multiple LLM experts, and efficiently train the merged LLM.
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
Add a description, image, and links to the mixture-of-experts topic page so that developers can more easily learn about it.
To associate your repository with the mixture-of-experts topic, visit your repo's landing page and select "manage topics."