Skip to content
#

llm-benchmarks

Here are 28 public repositories matching this topic...

AI 大模型世界,把 556 个大模型拟人化成像素小人的可视化站点。进来就能看到此刻谁最聪明、谁最会写代码、谁最便宜、谁刚发布,往下是国内与国外分区的厂商广场、完整的发布时间线和多维排行榜。搜索认模型名、厂商和能力,输入「多模态」会直接列出全部多模态模型。数据取自 Epoch AI、models.dev、LiveBench 与 Hugging Face,每小时自动同步,所有文案由真实数据生成,不调用任何 LLM。Next.js 静态导出,零后端。

  • Updated Sep 19, 2026
  • TypeScript

This project aims to address this gap by conducting a systematic, controlled study of human versus LLM-generated text detectability using paired question–answer datasets. Rather than proposing a novel detection architecture, the focus is on analyzing detection robustness, failure modes, and the impact of adversarial humanization strategies.

  • Updated Mar 19, 2026
  • Jupyter Notebook

A local LLM benchmarking framework designed to evaluate model performance across multiple backends (LM Studio, Ollama, etc.), including metrics for speed, quality, and instruction adherence. Supports structured test runs, result analysis, and reproducible evaluations.

  • Updated May 3, 2026
  • HTML

A living, evidence-based catalog of benchmarks for personalized LLMs and AI agents—covering preference alignment, long-term memory, tool use, safety, privacy, and multimodal adaptation.

  • Updated Jul 17, 2026
  • Python

Add this topic to your repo

To associate your repository with the llm-benchmarks topic, visit your repo's landing page and select "manage topics."

Learn more