研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

19 个 · 评测集 · AI 核心

排序 Stars 周增
Andyyyy64/whichllm
Python · 2026-08-05 评测基准 评测集 生产可用 Stars 6205 周增 +21

找到在你的硬件上真正能跑且性能最优的本地 LLM。排名基于真实且时新的基准测试,而非参数量。一条命令,即刻运行。Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

llm-infraevaluation
open-compass/VLMEvalKit
Python · 2026-08-17 评测基准 评测集 研究原型 Stars 4345 周增 +9

大型多模态模型 LMM 的开源评估工具包,支持 220+ LMM 与 80+ 基准测试。Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

multimodalevaluationllm-infra
camel-ai/crab
Python · 2026-09-30 Agent 智能体 评测集 研究原型 Stars 427 周增 +0

🦀️ CRAB: 面向多模态语言模型 Agent 的跨环境 Agent 基准。https://crab.camel-ai.org/🦀️ CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents. https://crab.camel-ai.org/

agentmultimodalevaluationllm-infra
bowen-upenn/PersonaMem
Python · 2026-09-06 评测基准 评测集 实验 Stars 193 周增 +0

[COLM 2025] 论文 Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale——大语言模型在动态用户画像与大规模个性化响应任务上的基准测试。[COLM 2025] Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale

evaluationllm-infra
asimsinan/LLM-Research
Python · 2026-08-13 评测基准 评测集 实验 Stars 67 周增 +0

一个 LLM 相关论文、学位论文、工具、数据集、课程与基准的合集。A collection of LLM related papers, thesis, tools, datasets, courses, benchmarks

evaluationllm-infra
zjunlp/MemBase
Python · 2026-08-11 评测基准 评测集 实验 Stars 45 周增 +0

面向长时对话记忆层的综合基准测试框架A Comprehensive Benchmarking Framework for Long-Term Conversational Memory Layers

agentevaluationllm-infra
linny006/vector-db-live
Python · 2026-09-24 评测基准 评测集 实验 Stars 3 周增 +0

实时更新的向量数据库项目、集成和基准评测全景图——每……刷新。Live-updating landscape of vector database projects, integrations, and benchmarks — refreshed every

ragevaluationdatabasellm-infra
MrPeppersDev/agent-infrastructure-landscape
HTML · 2026-08-11 Agent 智能体 评测集 实验 Stars 2 周增 +0

AI agent memory 与基础设施全景——912 个系统 × 68 列的对比目录,覆盖记忆层、agent 框架、运行时、vector store、知识图谱、MCP server、benchmark。支持按类型化边、谱系、引用进行检索。AI agent memory & infrastructure landscape — comparative catalog of 912 systems × 68 columns covering memory layers, agent frameworks, runtimes, vector stores, knowledge graphs, MCP servers, benchmarks. Searchable with typed edges, lineages, citations.

agentragevaluationdatabase
kunal4040/hybrid-search-eval
Python · 2026-09-20 评测基准 评测集 实验 Stars 1 周增 +0

🔍 在 Weaviate 混合检索中评测 embedding 模型,基于自有数据或 MTEB 数据集评估 MRR@K、Hit@K、延迟与内存占用🔍 Benchmark embedding models in hybrid search with Weaviate. Evaluate MRR@K, Hit@K, latency, and memory using your data or MTEB datasets.

ragllm-infraevaluationdatabase
seaofokhotskquakerism746/dabench-rlm-eval
Python · 2026-08-28 评测基准 评测集 实验 Stars 0 周增 +0

在 DABench 数据分析任务上对 DSPy RLM(Recursive Language Models)进行基准测试,使用自动评分实现基于代码的迭代评估Benchmark DSPy Recursive Language Models on DABench data analysis tasks with automated scoring for iterative code-based evaluation

agentragllm-infraevaluation
Sai-Kartheek-Reddy/HateMirage-ICON2026
CSS · 2026-08-19 评测基准 评测集 研究原型 Stars 0 周增 +0

HateMirage——可解释的伪仇恨检测与多维推理,ICON 2026 共享任务。在 HateMirage 语料库(4,530 条带标注的伪仇恨评论)上进行目标识别 + 意图与隐含意义生成。HateMirage - Explainable Faux Hate Detection and Multi-Dimensional Reasoning - Shared Task @ ICON 2026. Target identification + Intent and Implication generation over the HateMirage corpus (4,530 annotated Faux Hate comments).

ragevaluationllm-infra
naturalmoods/clawclones
TypeScript · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

追踪并比较每一个严肃的 OpenClaw 替代方案——基于实测仓库数据与 AI 撰写的决策支持,且刻意保持独立。Track and compare every serious OpenClaw alternative — measured repo data, AI-written decision support, kept apart on purpose.

agent
friedpotato04/CUDA-L2
Cuda · 2026-09-10 LLM 基础设施 评测集 实验 Stars 0 周增 +0

🚀 使用强化学习优化半精度通用矩阵乘法(HGEMM)CUDA kernel,性能超越 cuBLAS 及其他基准。🚀 Optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels using reinforcement learning, surpassing cuBLAS and other benchmarks with superior performance.

evaluationllm-infra
DanceNitra/ramr
Python · 2026-08-12 RAG 检索增强 评测集 实验 Stars 0 周增 +0

RAMR —— 检索增强记忆可靠性:面向 Agentic-RAG / 记忆系统的抗污染合成基准(附方法与发现)。RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)

agentragevaluationllm-infra
Bot87Ever/thunderbolt-ai
Python · 2026-08-21 RAG 检索增强 评测集 实验 Stars 0 周增 +0

本地 LLM 基准测试、RAG、语音交互与 AI 助手Local LLM benchmarking, RAG, voice interaction and AI assistant

ragevaluationllm-infra
avrsnramasamy/AI-DESIGN-BENCHMARK
HTML · 2026-09-20 评测基准 评测集 实验 Stars 0 周增 +0

🎨 利用 AI 实时生成高保真 UI 设计,对比多版本方案,并跨多个模型导出可直接使用的代码。🎨 Generate high-fidelity UI designs in real-time with AI, compare variations, and export ready-to-use code across multiple models.

agentevaluationllm-infra
AnimeshShaw/GenIaC-SecBench
Python · 2026-08-31 评测基准 评测集 实验 Stars 0 周增 +0

基于人类基准的 LLM 生成 Infrastructure-as-Code 安全基准测试。100 场景 × 12 模型配置 = 1,196 个工件,由 Checkov/Trivy/KICS 扫描,并与 634 个人工编写模板对比。所有模型的漏洞密度均为人工的 3.2–3.9 倍。arXiv:2608.28021Human-anchored security benchmark for LLM-generated Infrastructure-as-Code. 100 scenarios × 12 model configs = 1,196 artifacts scanned by Checkov/Trivy/KICS, compared against 634 human-written templates. Every model: 3.2–3.9× human vulnerability density. arXiv:2608.28021

evaluationriskllm-infra
aasimmalikin/agentic-qa
Python · 2026-08-11 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向文档溯源 QA Agent 的生产级 Harness:检索、起草、自评、重写、升级循环,配套 LLM-as-judge 评测套件、分级权限工具、对破坏性操作的人工审批,以及一键容器化部署。A production-grade harness for a document-grounded QA agent: a retrieve, draft, self-score, re-draft, escalate loop with an LLM-as-judge eval suite, permission-tiered tools, human-in-the-loop approval for destructive actions, and a one-command container deploy.

agentragengineeringllm-infra
0vertake/jetpacker
Kotlin · 2026-08-17 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向 AI coding agents 的 token 预算 context pack,基于编译器解析的 Kotlin 结构(Analysis API/PSI)构建,并附带衡量其是否优于 chunk RAG 的 benchmark。Token-budgeted context packs for AI coding agents, built from compiler-resolved Kotlin structure (Analysis API/PSI) — with the benchmark that measures whether it beats chunk RAG

agentragevaluationllm-infra