Skip to content

Popular repositories Loading

  1. distill-kura distill-kura Public

    蒸留蔵 — distilled long-term memory for agents: recall by meaning, writing gated by evidence, one kura per agent mode. Ships as a DeepSeek Harness plugin and an MCP server.

    Python 49 7

  2. blackwell-geforce-nvfp4-gemm blackwell-geforce-nvfp4-gemm Public

    NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.

    Python 25 2

  3. flash-next-8gb flash-next-8gb Public

    Run a 177B model (Qwen3.8-Flash-Next) on an 8 GB laptop GPU. Measured: 6.6 GiB VRAM, 47.8 GiB RAM, 34-35 tok/s. The n-gram table stays on disk; the experts run on CPU.

    Python 20 5

  4. lna-es lna-es Public

    あらゆるジャンルのテキストをLLMを使いNeo4Jグラフ化して、グラフのみのデータから意味的復元をするシステムのスターター(MCP対応予定)

    16 1

  5. gemma4-12b-vllm-sm120 gemma4-12b-vllm-sm120 Public

    Reproducible recipe: serve abliterated Gemma-4-12B (gemma4_unified) at 50-118 tok/s on no-NVLink Blackwell (SM120) via vLLM nightly + ModelOpt FP8/NVFP4 + MTP spec-decode.

    Python 14

  6. GGUF-to-NVFP4-SM120 GGUF-to-NVFP4-SM120 Public

    Lna-Lab production pipeline: GGUF -> modelopt-format NVFP4 + working MTP head for vLLM on RTX PRO 6000 Blackwell (SM120). Stages 2 (NVFP4) and 3 (MTP graft) are Lna-Lab originals; stage 1 (GGUF->bf…

    Python 10 3

Repositories

Showing 10 of 22 repositories
  • lna-lab/DeepSeek-V4.1-Flash-EXL3-DGX-Spark-recipe's past year of commit activity
    Python 0 5 0 0 Updated Sep 15, 2026
  • vllm-exl3 Public Forked from vcruz305/vllm-exl3

    Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights

    lna-lab/vllm-exl3's past year of commit activity
    Python 0 9 0 0 Updated Sep 14, 2026
  • distill-kura Public

    蒸留蔵 — distilled long-term memory for agents: recall by meaning, writing gated by evidence, one kura per agent mode. Ships as a DeepSeek Harness plugin and an MCP server.

    lna-lab/distill-kura's past year of commit activity
    Python 49 MIT 7 0 0 Updated Sep 6, 2026
  • dsv4-carve Public

    DeepSeek-V4-Flash-Vision 305B EXL3 on 8x16GB: 380K context, 4 streams, DSpark3 — Lna-Lab serving recipe

    lna-lab/dsv4-carve's past year of commit activity
    Python 0 0 0 0 Updated Sep 4, 2026
  • flash-next-8gb Public

    Run a 177B model (Qwen3.8-Flash-Next) on an 8 GB laptop GPU. Measured: 6.6 GiB VRAM, 47.8 GiB RAM, 34-35 tok/s. The n-gram table stays on disk; the experts run on CPU.

    lna-lab/flash-next-8gb's past year of commit activity
    Python 20 MIT 5 1 0 Updated Sep 1, 2026
  • lna-lab/Ashigaru-Search's past year of commit activity
    Python 10 Apache-2.0 1 0 0 Updated Aug 10, 2026
  • lna-lab/DS4-combination's past year of commit activity
    C 0 MIT 0 0 0 Updated Jun 16, 2026
  • LNAIME Public
    lna-lab/LNAIME's past year of commit activity
    Python 1 0 0 0 Updated Jun 10, 2026
  • gemma4-12b-vllm-sm120 Public

    Reproducible recipe: serve abliterated Gemma-4-12B (gemma4_unified) at 50-118 tok/s on no-NVLink Blackwell (SM120) via vLLM nightly + ModelOpt FP8/NVFP4 + MTP spec-decode.

    lna-lab/gemma4-12b-vllm-sm120's past year of commit activity
    Python 14 Apache-2.0 0 0 0 Updated Jun 7, 2026
  • LLM-Formura1 Public
    lna-lab/LLM-Formura1's past year of commit activity
    Python 0 0 0 0 Updated May 29, 2026

People

This organization has no public members. You must be a member to see who’s a part of this organization.