研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

49 个 · 评测基准 · AI 核心

排序 Stars 周增
bytedance/deer-flow
Python · 2026-08-11 评测基准 工具 生产可用 Stars 79697 周增 +161

一个开源的长时程 SuperAgent harness,可研究、编码与创作。借助 sandbox、记忆、工具、Skill、subagent 与 message gateway,处理耗时从分钟到小时不等的多层级任务。An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.

agentragllm-infra
tensorzero/tensorzero
Rust · 2026-06-11 评测基准 框架 生产可用 Stars 11724 周增 +0

TensorZero 是一个开源 LLMOps 平台,统一了 LLM gateway、可观测性、评估、优化与实验TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.

evaluationengineeringllm-infra
Arize-ai/phoenix
Python · 2026-08-11 评测基准 应用 生产可用 Stars 10982 周增 +49

AI 可观测性与评估。AI Observability & Evaluation

agentevaluationllm-infra
evidentlyai/evidently
Jupyter Notebook · 2026-08-05 评测基准 框架 生产可用 Stars 7797 周增 +14

Ev​​idently 是一个开源的 ML 和 LLM 可观测性框架,评估、测试和监控任何 AI 驱动的系统或数据 pipeline,覆盖从表格数据到 Gen AI 场景,提供 100+ 指标。Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.

engineeringllm-infra
Andyyyy64/whichllm
Python · 2026-08-05 评测基准 评测集 生产可用 Stars 6205 周增 +21

找到在你的硬件上真正能跑且性能最优的本地 LLM。排名基于真实且时新的基准测试,而非参数量。一条命令,即刻运行。Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

llm-infraevaluation
Helicone/helicone
TypeScript · 2026-08-31 评测基准 框架 生产可用 Stars 6116 周增 +23

🧊 开源 LLM 可观测性平台,一行代码即可实现监控、评估与实验。YC W23 🍓🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓

agentevaluationllm-infra
open-compass/VLMEvalKit
Python · 2026-08-17 评测基准 评测集 研究原型 Stars 4345 周增 +9

大型多模态模型 LMM 的开源评估工具包,支持 220+ LMM 与 80+ 基准测试。Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

multimodalevaluationllm-infra
Agenta-AI/agenta
TypeScript · 2026-07-18 评测基准 框架 研究原型 Stars 4303 周增 +21

开源 LLMOps 平台:集成 prompt playground、prompt 管理、LLM 评估和 LLM 可观测性。The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

agentragevaluationllm-infra
ianarawjo/ChainForge
TypeScript · 2026-09-27 评测基准 工具 研究原型 Stars 3030 周增 +2

用于对 LLM prompt 进行对抗测试的开源可视化编程环境An open-source visual programming environment for battle-testing prompts to LLMs.

evaluationllm-infra
Gentleman-Programming/gentle-pi
TypeScript · 2026-09-09 评测基准 应用 研究原型 Stars 666 周增 +0

将 Pi 打造成 el Gentleman:一套高级架构师级开发 harness,集成 SDD/OpenSpec、子 Agent、严格 TDD 证据、review guardrail 与 skill 发现机制。Turn Pi into el Gentleman: a senior-architect development harness with SDD/OpenSpec, subagents, strict TDD evidence, review guardrails, and skill discovery.

agentllm-infra
codelion/adaptive-classifier
Python · 2026-10-08 评测基准 库 实验 Stars 574 周增 +0

一个面向动态文本分类的灵活、自适应分类系统A flexible, adaptive classification system for dynamic text classification

llm-infra
PerryLink/dsh-auto-review
TypeScript · 2026-10-05 评测基准 模型 实验 Stars 233 周增 +0

DeepSeek Harness 审批请求的 second-model AI 自动复核:只读复核子 Agent 返回结构化 allow/deny 判定及理由,默认 fail-closed,会话日志(approval/asked → autoReview/verdict → approval/decided)全程可审计。Second-model AI auto-review for DeepSeek Harness approval requests: a read-only reviewer subagent returns structured allow/deny verdicts with reasons, fail-closed by default, fully auditable from the session log (approval/asked -> autoReview/verdict -> approval/decided).

agentriskllm-infra
PerryLink/dsh-research-report
TypeScript · 2026-10-05 评测基准 应用 实验 Stars 214 周增 +0

DeepSeek Harness 的可验证研究报告引擎:内容寻址的 evidence ledger(claim-snapshot 绑定、防篡改),以及版本化 sealed 报告,含逐条 claim 的验证判定与 manifest-sealed 目录。Verifiable research-report engine for DeepSeek Harness: content-addressed evidence ledger (claim-snapshot binding, tamper-evident) plus versioned sealed reports with per-claim verification verdicts and a manifest-sealed directory.

agent
bowen-upenn/PersonaMem
Python · 2026-09-06 评测基准 评测集 实验 Stars 193 周增 +0

[COLM 2025] 论文 Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale——大语言模型在动态用户画像与大规模个性化响应任务上的基准测试。[COLM 2025] Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale

evaluationllm-infra
thenicolas1894/awesome-claude-fable-5-prompt-vault
HTML · 2026-10-07 评测基准 收藏榜 实验 Stars 142 周增 +0

Claude Fable 5 终极指南 2026:使用场景、集成与基准测试Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks

agentevaluationllm-infra
YanpengQi7/ai-reliability-copilot
TypeScript · 2026-09-30 评测基准 应用 实验 Stars 83 周增 +0

将生产事故转化为结构化的 9 段式 LLM 响应(严重等级、根因、缓解措施、事后复盘)。附带 5 个场景的回归测试集 + LLM-as-judge 评估流水线。Turn a production incident into a structured 9-section LLM response (severity, root cause, mitigation, postmortem). Ships with a 5-scenario regression suite + LLM-as-judge eval pipeline.

ragevaluationengineeringdatabase
asimsinan/LLM-Research
Python · 2026-08-13 评测基准 评测集 实验 Stars 67 周增 +0

一个 LLM 相关论文、学位论文、工具、数据集、课程与基准的合集。A collection of LLM related papers, thesis, tools, datasets, courses, benchmarks

evaluationllm-infra
PerryLink/dsh-doublecheck
TypeScript · 2026-10-05 评测基准 应用 实验 Stars 55 周增 +0

发布前双重检查:审视需求、测试实现、证明交付。面向 DeepSeek Harness 的工程纪律工具集。Double-check before you ship: grill the requirements, test the implementation, prove the delivery. An engineering-discipline bundle for DeepSeek Harness.

agent
zjunlp/MemBase
Python · 2026-08-11 评测基准 评测集 实验 Stars 45 周增 +0

面向长时对话记忆层的综合基准测试框架A Comprehensive Benchmarking Framework for Long-Term Conversational Memory Layers

agentevaluationllm-infra
mohsenhariri/scorio
Python · 2026-10-05 评测基准 应用 实验 Stars 23 周增 +0

Bayes@N [ICLR'26]、Ranking LLMs [ACL'26 Main]:大语言模型的统计评估、比较与排序。Bayes@N [ICLR'26], Ranking LLMs [ACL'26 Main]: Statistical evaluation, comparison, and ranking of Large Language Models

llm-infraevaluation
yyh-001/llm-value-rankings
CSS · 2026-08-27 评测基准 工具 实验 Stars 18 周增 +3

每日 LLM 价值排行榜——基于智能、速度、价格对比 300+ 模型。OpenRouter + Artificial Analysis。大模型性价比排行榜Daily LLM value rankings - compare 300+ models by intelligence, speed and price. OpenRouter + Artificial Analysis. 大模型性价比排行榜

evaluationllm-infra
DigitalHarborFoundation/FlexEval
Python · 2026-10-01 评测基准 工具 实验 Stars 17 周增 +0

FlexEval 是一个面向实际量化分析的 LLM 评估工具。FlexEval is an LLM evaluation tool designed for practical quantitative analysis.

evaluationllm-infra
ELM-Research/ECG-Language-Models
Python · 2026-08-11 评测基准 框架 实验 Stars 16 周增 +0

面向 ECG-语言模型(ELM)的研究型训练与评估框架A research-oriented training and evaluation framework for ECG-Language Models (ELMs)

multimodalevaluationllm-infra
KCNyu/clawock
Python · 2026-09-03 评测基准 应用 实验 Stars 14 周增 +2

AI 争论,代码结算,亏损留在账面上。一个由 Agent 驱动、面向香港及美国市场的真实券商账户:每个交易决策都必须经过辩论,并由模型无法触碰的代码完成结算。将同一决策工作流安装到你的 Agent:OpenClaw、Claude Code、Codex 或 DeepSeek Harness。AI argues. Code settles. The losses stay on the page. A real HK + US brokerage account run by agents that must debate every call, settled by code the model never touches. Install the same decision workflow into your own agent: OpenClaw, Claude Code, Codex, or DeepSeek Harness.

agentragrisk
zjunlp/Mechanist
Python · 2026-08-12 评测基准 框架 研究原型 Stars 13 周增 +0

AI 系统作为理解智能机制的科学仪器AI Systems as Scientific Instruments for Understanding the Mechanisms of Intelligence

agentllm-infra
MARKTECHPOST-AI-MEDIA-INC/LLMs-Tutorials-Projects
未知语言 · 2026-08-11 评测基准 教程 实验 Stars 9 周增 +0

微调、评估、提示工程、开源模型Fine-tuning, evaluation, prompting, open-source models

evaluationllm-infra
saurabhr/psychscanner
HTML · 2026-08-13 评测基准 工具 研究原型 Stars 3 周增 +0

自动化并规模化 "LLMs as a participant",将 LLM 作为研究参与者Automates and scales "LLMs as a participant."

evaluationllm-infra
linny006/vector-db-live
Python · 2026-09-24 评测基准 评测集 实验 Stars 3 周增 +0

实时更新的向量数据库项目、集成和基准评测全景图——每……刷新。Live-updating landscape of vector database projects, integrations, and benchmarks — refreshed every

ragevaluationdatabasellm-infra
harshtiwari01/llm-heatmap-visualizer
Jupyter Notebook · 2026-10-03 评测基准 工具 实验 Stars 3 周增 +0

用于生成基于 Transformer 的 LLM 完整注意力头热力图的一组脚本。A set of scripts to generate full attention-head heatmaps for transformer-based LLMs

llm-infraevaluation
Zhuchen00123/Verified-Executable-Search
Python · 2026-08-11 评测基准 工具 实验 Stars 1 周增 +0

面向 AI 生成可执行方案的 verifier-first 运行时,支持独立验证、对比与检索。A verifier-first runtime for independently verifying, comparing, and searching AI-generated executable solutions.

agentriskllm-infra
yyxcnasd/amadeus-for-dsh
JavaScript · 2026-08-15 评测基准 应用 实验 Stars 1 周增 +0

Amadeus(来自《Steins;Gate 0》的 AI 助手)适配 DeepSeek Harness。Amadeus (AI assistant from Steins;Gate 0) for DeepSeek Harness

agent
Raina-Xin/PA_Score
Python · 2026-10-08 评测基准 库 研究原型 Stars 1 周增 +0

[NeurIPS 2026] 基于释义感知评分的 LLM 不确定性量化 conformal prediction[NeurIPS 2026] Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification.

llm-infra
pt9912/ai-harness-init
Go · 2026-09-14 评测基准 工具 实验 Stars 1 周增 +0

一个用 AI Harness 流程引导 Git 仓库的 CLI —— 基于固定课程骨架提供模板、文档门禁与特定语言的代码门禁。CLI, die ein Git-Repo mit dem AI-Harness-Prozess bootstrappt — Templates, Doc-Gates und sprachspezifische Code-Gates aus gepinnten Kurs-Skeletten.

agentllm-infra
NinjaSln-labs/dsh-plugins
TypeScript · 2026-08-16 评测基准 应用 实验 Stars 1 周增 +0

DeepSeek Harness 个人自研插件集:上下文罗盘 / 跨会话知识 / 子代理模型路由 / AI 生图(Personally developed plugins for DeepSeek Harness)

agentllm-infra
kunal4040/hybrid-search-eval
Python · 2026-09-20 评测基准 评测集 实验 Stars 1 周增 +0

🔍 在 Weaviate 混合检索中评测 embedding 模型,基于自有数据或 MTEB 数据集评估 MRR@K、Hit@K、延迟与内存占用🔍 Benchmark embedding models in hybrid search with Weaviate. Evaluate MRR@K, Hit@K, latency, and memory using your data or MTEB datasets.

ragllm-infraevaluationdatabase
chrisliu298/awesome-rubric-rewards
未知语言 · 2026-08-11 评测基准 收藏榜 生产可用 Stars 1 周增 +0

精选的评分量表、检查清单、评分标准集、原则和评分指南汇总,用于对现代生成模型进行评分、排名、验证、过滤或训练。A curated list of rubrics, checklists, criteria sets, principles, and scoring guides used to score, rank, verify, filter, or train modern generative models.

evaluationriskllm-infra