研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

12 个 · Agent 智能体 · 评测集

排序 Stars 周增
tirth8205/code-review-graph
Python · 2026-08-02 Agent 智能体 评测集 生产可用 Stars 29735 周增 +357

Local-first 代码智能图谱,面向 MCP 与 CLI。为代码库构建持久化映射,使 AI 编程工具只读取关键内容,在代码评审与大仓库工作流中实现可基准测试的上下文缩减。Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

ragevaluationllm-infra
rohitg00/agentmemory
TypeScript · 2026-08-10 Agent 智能体 评测集 生产可用 Stars 26852 周增 +133

基于真实场景基准测试的、面向 AI 编程 Agent 的 #1 持久化记忆方案#1 Persistent memory for AI coding agents based on real-world benchmarks

agentevaluation
IBM/AssetOpsBench
Python · 2026-08-11 Agent 智能体 评测集 研究原型 Stars 2122 周增 +42

AssetOpsBench - Industry 4.0:面向工业4.0资产运维的领域AI Agent构建、编排与评估统一基准与框架,包含460+场景、5类专业Agent(IoT、FMSR、TSFM、工单等),以及基于MCP的多Agent编排蓝图(MetaAgent、AgentHive)AssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-specific AI agents for Industry 4.0 asset operations and maintenance, with 460+ scenarios, 5 specialist agents (IoT, FMSR, TSFM, Work Order,...), and multi-agent orchestration blueprints (MetaAgent, AgentHive) over MCP.

agentevaluationllm-infra
uber/ADR
Python · 2026-09-16 Agent 智能体 评测集 研究原型 Stars 1569 周增 +0

ADR 通过可观测性、安全基准测试与威胁检测,为企业级 AI Agent 提供安全保障。已部署于 Uber。ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

agentevaluationriskllm-infra
Ar9av/PaperOrchestra
Python · 2026-09-21 Agent 智能体 评测集 研究原型 Stars 662 周增 +4

基于 Google PaperOrchestra 论文实现的全自动 AI 研究论文写作器,通过技能-基准测试 + 自动评分器,配合任意编码 Agent(Claude Code、Cursor、Antigravity、Cline、Aider)。无需 API Key,无需 LLM SDK。An automated AI research-paper writer based off Google's PaperOrchestra paper's implementation through a skills - benchmark + autoraters using any coding agent (Claude Code, Cursor, Antigravity, Cline, Aider). No API keys, no LLM SDKs.

agentevaluationllm-infraengineering
camel-ai/crab
Python · 2026-09-30 Agent 智能体 评测集 研究原型 Stars 427 周增 +0

🦀️ CRAB: 面向多模态语言模型 Agent 的跨环境 Agent 基准。https://crab.camel-ai.org/🦀️ CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents. https://crab.camel-ai.org/

agentmultimodalevaluationllm-infra
omnilink-tech/omnisim
Python · 2026-10-03 Agent 智能体 评测集 实验 Stars 186 周增 +0

面向 coding agent 的开源机器人仿真器:HTTP/JSON + MCP 控制、Newton 物理、wgpu 渲染、ROS 2,以及可复现的 benchmark。Open-source robotics simulator for coding agents: HTTP/JSON + MCP control, Newton physics, wgpu rendering, ROS 2, and reproducible benchmarks.

agentevaluation
GamePhanes/GamePhanes
JavaScript · 2026-08-22 Agent 智能体 评测集 实验 Stars 104 周增 +0

面向 Godot 的开源游戏编程 Agent 环境与基准An open-source game coding agent environment and benchmark for Godot.

agentevaluation
MrPeppersDev/agent-infrastructure-landscape
HTML · 2026-08-11 Agent 智能体 评测集 实验 Stars 2 周增 +0

AI agent memory 与基础设施全景——912 个系统 × 68 列的对比目录,覆盖记忆层、agent 框架、运行时、vector store、知识图谱、MCP server、benchmark。支持按类型化边、谱系、引用进行检索。AI agent memory & infrastructure landscape — comparative catalog of 912 systems × 68 columns covering memory layers, agent frameworks, runtimes, vector stores, knowledge graphs, MCP servers, benchmarks. Searchable with typed edges, lineages, citations.

agentragevaluationdatabase
dandovdub/residoo
JavaScript · 2026-09-03 Agent 智能体 评测集 实验 Stars 1 周增 +0

查找你的 AI coding Agent 泄露到磁盘的敏感信息。免费、MIT 许可、零依赖、零网络调用。在我们公开发布的基准测试中击败 TruffleHog 和 GitGuardian 的引擎。Find secrets your AI coding agent leaked to disk. Free, MIT, zero deps, zero network calls. Beat TruffleHog and GitGuardian's engine on our own published benchmark.

agentevaluationrisk
aasimmalikin/agentic-qa
Python · 2026-08-11 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向文档溯源 QA Agent 的生产级 Harness:检索、起草、自评、重写、升级循环,配套 LLM-as-judge 评测套件、分级权限工具、对破坏性操作的人工审批,以及一键容器化部署。A production-grade harness for a document-grounded QA agent: a retrieve, draft, self-score, re-draft, escalate loop with an LLM-as-judge eval suite, permission-tiered tools, human-in-the-loop approval for destructive actions, and a one-command container deploy.

agentragengineeringllm-infra
0vertake/jetpacker
Kotlin · 2026-08-17 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向 AI coding agents 的 token 预算 context pack,基于编译器解析的 Kotlin 结构(Analysis API/PSI)构建,并附带衡量其是否优于 chunk RAG 的 benchmark。Token-budgeted context packs for AI coding agents, built from compiler-resolved Kotlin structure (Analysis API/PSI) — with the benchmark that measures whether it beats chunk RAG

agentragevaluationllm-infra