研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

5 个 · Agent 智能体 · 评测集 · Skill/MCP

排序 Stars 周增
rohitg00/agentmemory
TypeScript · 2026-08-10 Agent 智能体 评测集 生产可用 Stars 26852 周增 +133

基于真实场景基准测试的、面向 AI 编程 Agent 的 #1 持久化记忆方案#1 Persistent memory for AI coding agents based on real-world benchmarks

agentevaluation
IBM/AssetOpsBench
Python · 2026-08-11 Agent 智能体 评测集 研究原型 Stars 2122 周增 +42

AssetOpsBench - Industry 4.0:面向工业4.0资产运维的领域AI Agent构建、编排与评估统一基准与框架,包含460+场景、5类专业Agent(IoT、FMSR、TSFM、工单等),以及基于MCP的多Agent编排蓝图(MetaAgent、AgentHive)AssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-specific AI agents for Industry 4.0 asset operations and maintenance, with 460+ scenarios, 5 specialist agents (IoT, FMSR, TSFM, Work Order,...), and multi-agent orchestration blueprints (MetaAgent, AgentHive) over MCP.

agentevaluationllm-infra
uber/ADR
Python · 2026-09-16 Agent 智能体 评测集 研究原型 Stars 1569 周增 +0

ADR 通过可观测性、安全基准测试与威胁检测,为企业级 AI Agent 提供安全保障。已部署于 Uber。ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

agentevaluationriskllm-infra
omnilink-tech/omnisim
Python · 2026-10-03 Agent 智能体 评测集 实验 Stars 186 周增 +0

面向 coding agent 的开源机器人仿真器:HTTP/JSON + MCP 控制、Newton 物理、wgpu 渲染、ROS 2,以及可复现的 benchmark。Open-source robotics simulator for coding agents: HTTP/JSON + MCP control, Newton physics, wgpu rendering, ROS 2, and reproducible benchmarks.

agentevaluation
dandovdub/residoo
JavaScript · 2026-09-03 Agent 智能体 评测集 实验 Stars 1 周增 +0

查找你的 AI coding Agent 泄露到磁盘的敏感信息。免费、MIT 许可、零依赖、零网络调用。在我们公开发布的基准测试中击败 TruffleHog 和 GitGuardian 的引擎。Find secrets your AI coding agent leaked to disk. Free, MIT, zero deps, zero network calls. Beat TruffleHog and GitGuardian's engine on our own published benchmark.

agentevaluationrisk