Skip to content
#

rtx-3090

Here are 47 public repositories matching this topic...

SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.

  • Updated Sep 13, 2026
  • Python

llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): TurboQuant KV cache, MTP speculative decoding with a 64K draft-vocabulary shortlist, custom SM86 + Qwen kernels. 90 tok/s over a 100K-token generation at temperature 1.

  • Updated Sep 18, 2026
  • C++

llama.cpp speculative decoding measured on one RTX 3090, Qwen3.6-35B-A3B UD-Q4_K_XL, commit 3737e4137. Published figures are re-derived from the committed data by a checker that fails on drift, a coverage probe reports how many of them it actually covers, and ERRATA.md lists this study's own retracted claims.

  • Updated Sep 17, 2026
  • Python

Add this topic to your repo

To associate your repository with the rtx-3090 topic, visit your repo's landing page and select "manage topics."

Learn more