Repositories · organized/repo_cards

仓库/Skill 库

45 个 · 评测集

排序 Stars 周增
MemPalace/mempalace
Python · 2026-08-08 评测基准 评测集 生产可用 Stars 58289 周增 +70

经过最佳基准测试的开源 AI 记忆系统,而且是免费的。The best-benchmarked open-source AI memory system. And it's free.

evaluationllm-infra
tirth8205/code-review-graph
Python · 2026-08-02 Agent 智能体 评测集 生产可用 Stars 29735 周增 +357

Local-first 代码智能图谱,面向 MCP 与 CLI。为代码库构建持久化映射,使 AI 编程工具只读取关键内容,在代码评审与大仓库工作流中实现可基准测试的上下文缩减。Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

ragevaluationllm-infra
rohitg00/agentmemory
TypeScript · 2026-08-10 Agent 智能体 评测集 生产可用 Stars 26852 周增 +133

基于真实场景基准测试的、面向 AI 编程 Agent 的 #1 持久化记忆方案#1 Persistent memory for AI coding agents based on real-world benchmarks

agentevaluation
Andyyyy64/whichllm
Python · 2026-08-05 评测基准 评测集 生产可用 Stars 6205 周增 +21

找到在你的硬件上真正能跑且性能最优的本地 LLM。排名基于真实且时新的基准测试,而非参数量。一条命令,即刻运行。Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

llm-infraevaluation
open-compass/VLMEvalKit
Python · 2026-08-17 评测基准 评测集 研究原型 Stars 4345 周增 +9

大型多模态模型 LMM 的开源评估工具包,支持 220+ LMM 与 80+ 基准测试。Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

multimodalevaluationllm-infra
modelscope/evalscope
Python · 2026-08-10 评测基准 评测集 研究原型 Stars 3218 周增 -7

一个精简且可定制的高效大模型(LLM、VLM、AIGC)评估与性能基准测试框架。A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

ragevaluationllm-infra
IBM/AssetOpsBench
Python · 2026-08-11 Agent 智能体 评测集 研究原型 Stars 2122 周增 +42

AssetOpsBench - Industry 4.0:面向工业4.0资产运维的领域AI Agent构建、编排与评估统一基准与框架,包含460+场景、5类专业Agent(IoT、FMSR、TSFM、工单等),以及基于MCP的多Agent编排蓝图(MetaAgent、AgentHive)AssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-specific AI agents for Industry 4.0 asset operations and maintenance, with 460+ scenarios, 5 specialist agents (IoT, FMSR, TSFM, Work Order,...), and multi-agent orchestration blueprints (MetaAgent, AgentHive) over MCP.

agentevaluationllm-infra
Sushegaad/Claude-Skills-Governance-Risk-and-Compliance
HTML · 2026-07-20 评测基准 评测集 研究原型 Stars 814 周增 +7

面向治理、风险与合规(GRC)的 Claude Skills:针对 ISO 27001、SOC 2、FedRAMP、GDPR、HIPAA、NIST CSF、PCI DSS、EU AI Act、ISO 42001、ISO 27701、DORA、CSRD、印度 DPDPA、CMMC 2.0、NIST AI Risk、SWIFT、澳大利亚 ISM、EU NIS2、CCPA/CPRA 等的专家级合规指导。使用 skills 基准 97%,不使用 81%。Claude Skills for Governance, Risk, & Compliance (GRC): Expert-level compliance guidance for ISO 27001, SOC 2, FedRAMP, GDPR, HIPAA, NIST CSF, PCI DSS, EU AI Act, ISO 42001, ISO 27701, DORA, CSRD, India's DPDPA, CMMC 2.0, NIST AI Risk, SWIFT, Australia's ISM, EU NIS2, CCPA/CPRA, and others. Benchmark 97% (with skills) vs 81% (without skills).

evaluationrisk
Ar9av/PaperOrchestra
Python · 2026-08-09 Agent 智能体 评测集 研究原型 Stars 636 周增 +7

基于 Google PaperOrchestra 论文实现的全自动 AI 研究论文写作器,通过技能-基准测试 + 自动评分器,配合任意编码 Agent(Claude Code、Cursor、Antigravity、Cline、Aider)。无需 API Key,无需 LLM SDK。An automated AI research-paper writer based off Google's PaperOrchestra paper's implementation through a skills - benchmark + autoraters using any coding agent (Claude Code, Cursor, Antigravity, Cline, Aider). No API keys, no LLM SDKs.

agentevaluationllm-infraengineering
syv-ai/qwen38-27b-rtx3090
Python · 2026-08-21 LLM 基础设施 评测集 研究原型 Stars 345 周增 +0

Qwen3.8-27B 在单卡 RTX 3090 上使用 vLLM 部署:64 并发下约 1,000 tok/s(int8 张量核心 GEMM、fp16 DeltaNet 状态),默认采样下单用户约 114 tok/s/贪心约 124 tok/s(MTP 草稿、自输出草稿词表、校准 int4 lm_head、split-KV 校验注意力),150k–262k 上下文;附带补丁、重新量化脚本与基准测试Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

llm-infraevaluation
GamePhanes/GamePhanes
JavaScript · 2026-08-22 Agent 智能体 评测集 实验 Stars 104 周增 +0

面向 Godot 的开源游戏编程 Agent 环境与基准An open-source game coding agent environment and benchmark for Godot.

agentevaluation
asimsinan/LLM-Research
Python · 2026-08-13 评测基准 评测集 实验 Stars 67 周增 +0

一个 LLM 相关论文、学位论文、工具、数据集、课程与基准的合集。A collection of LLM related papers, thesis, tools, datasets, courses, benchmarks

evaluationllm-infra
LitLLM/litllms-for-literature-review-tmlr
Python · 2025-04-20 评测基准 评测集 研究原型 Stars 61 周增 -7

论文 LitLLMs, LLMs for Literature Review: Are we there yet?(TMLR 2025)的代码仓库。Code for LitLLMs, LLMs for Literature Review: Are we there yet? (TMLR 2025)

llm-infra
YZCU/OOTB
C · 2025-07-11 多模态 评测集 实验 Stars 53 周增 +0

[ISPRS 2024] 卫星视频单目标跟踪:系统综述与定向目标跟踪基准[ISPRS 2024] Satellite Video Single Object Tracking: A Systematic Review and An Oriented Object Tracking Benchmark

multimodalevaluation
zjunlp/MemBase
Python · 2026-08-11 评测基准 评测集 实验 Stars 45 周增 +0

面向长时对话记忆层的综合基准测试框架A Comprehensive Benchmarking Framework for Long-Term Conversational Memory Layers

agentevaluationllm-infra
linny006/vector-db-live
Python · 2026-08-25 评测基准 评测集 实验 Stars 3 周增 +0

实时更新的向量数据库项目、集成和基准评测全景图——每……刷新。Live-updating landscape of vector database projects, integrations, and benchmarks — refreshed every

ragevaluationdatabasellm-infra
SamInMotion/Medical-intervention-text-classification
Python · 2026-08-13 评测基准 评测集 实验 Stars 2 周增 +0

硕士论文(UiB,计算语言学,2023)及扩展工作。针对系统综述筛选的文本分类。原始 11 工作流分析结合 NEO ontology 集成,以及 Cohen 等人(2006)药物类别基准的扩展,比较 BoW 与 BiomedBERT。Master's thesis (UiB, Computational Linguistics, 2023) and extension work. Text classification for systematic review screening. Original 11-workflow analysis with NEO ontology integration plus Cohen et al. (2006) drug-class benchmark extension comparing BoW and BiomedBERT.

evaluation
MrPeppersDev/agent-infrastructure-landscape
HTML · 2026-08-11 Agent 智能体 评测集 实验 Stars 2 周增 +0

AI agent memory 与基础设施全景——912 个系统 × 68 列的对比目录,覆盖记忆层、agent 框架、运行时、vector store、知识图谱、MCP server、benchmark。支持按类型化边、谱系、引用进行检索。AI agent memory & infrastructure landscape — comparative catalog of 912 systems × 68 columns covering memory layers, agent frameworks, runtimes, vector stores, knowledge graphs, MCP servers, benchmarks. Searchable with typed edges, lineages, citations.

agentragevaluationdatabase
kunal4040/hybrid-search-eval
Python · 2026-08-13 评测基准 评测集 实验 Stars 1 周增 +0

🔍 在 Weaviate 混合检索中评测 embedding 模型,基于自有数据或 MTEB 数据集评估 MRR@K、Hit@K、延迟与内存占用🔍 Benchmark embedding models in hybrid search with Weaviate. Evaluate MRR@K, Hit@K, latency, and memory using your data or MTEB datasets.

ragllm-infraevaluationdatabase
HmZ9874/umd-planetary-memory
Python · 2026-08-12 评测基准 评测集 实验 Stars 1 周增 +0

UMD 行星长期记忆算法、公式、基准、SDK 与可复现研究。UMD planetary long-term memory algorithms, formulas, benchmarks, SDKs, and reproducible research

ragevaluationllm-infra
suanlab/pinns-neural-operators-review
Python · 2026-08-25 评测基准 评测集 实验 Stars 0 周增 +0

论文《Physics-Informed Neural Networks and Neural Operators for PDEs: A Unified Taxonomy and Systematic Review》(Neural Networks)的补充材料:PRISMA 数据集、PDE 复杂度评分标准以及概念验证的统一基准。Supplementary materials for 'Physics-Informed Neural Networks and Neural Operators for PDEs: A Unified Taxonomy and Systematic Review' (Neural Networks): PRISMA datasets, PDE complexity rubric, and proof-of-concept unified benchmark

evaluation
soheylfalahzade/geometric-spanners-lab
Python · 2026-08-16 评测基准 评测集 实验 Stars 0 周增 +0

可复现研究基准实验:二维度量空间下的贪心 t-Spanner 构造算法A reproducible research benchmarking lab for Greedy t-Spanner Construction Algorithms in 2D Metric Spaces.

evaluation
sharing-123/ai-myopia-diagnostic-accuracy-systematic-review
Python · 2026-08-17 评测基准 评测集 研究原型 Stars 0 周增 +0

AI 诊断近视准确性的系统综述:完整数据、提取流程、代码与稿件。Systematic review of AI diagnostic accuracy for myopia detection: full data, extraction, code, and manuscript

seaofokhotskquakerism746/dabench-rlm-eval
Python · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

在 DABench 数据分析任务上对 DSPy RLM(Recursive Language Models)进行基准测试,使用自动评分实现基于代码的迭代评估Benchmark DSPy Recursive Language Models on DABench data analysis tasks with automated scoring for iterative code-based evaluation

agentragllm-infraevaluation
savch1102/Systematic-review-vectorborne-models
R · 2026-08-22 评测基准 评测集 研究原型 Stars 0 周增 +0
Sai-Kartheek-Reddy/HateMirage-ICON2026
CSS · 2026-08-19 评测基准 评测集 研究原型 Stars 0 周增 +0

HateMirage——可解释的伪仇恨检测与多维推理,ICON 2026 共享任务。在 HateMirage 语料库(4,530 条带标注的伪仇恨评论)上进行目标识别 + 意图与隐含意义生成。HateMirage - Explainable Faux Hate Detection and Multi-Dimensional Reasoning - Shared Task @ ICON 2026. Target identification + Intent and Implication generation over the HateMirage corpus (4,530 annotated Faux Hate comments).

ragevaluationllm-infra
Ruixixu-hub/2026-surf-american-risk-surfaces
Python · 2026-08-20 评测基准 评测集 实验 Stars 0 周增 +0

SURF2026 计算金融项目,研究感知自由边界的美式期权风险曲面,使用 CN/PSOR 基准、文献综述、实验报告以及 Codex 辅助的分步研究规划SURF2026 computational finance project on free-boundary-aware American option risk surfaces, using CN/PSOR benchmarks, literature review, experiment reports, and step-by-step Codex-assisted research planning.

evaluationrisk
QRSocietyTMU/AlphaProject
Jupyter Notebook · 2026-08-19 评测基准 评测集 实验 Stars 0 周增 +0

可复现研究:基于基准、统计模型与走步前向验证,检验可解释的市场信号能否预测 SPY 的五日方向。Reproducible research testing whether interpretable market signals can forecast SPY’s five-day direction using benchmarks, statistical models, and walk-forward validation.

evaluation
naturalmoods/clawclones
TypeScript · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

追踪并比较每一个严肃的 OpenClaw 替代方案——基于实测仓库数据与 AI 撰写的决策支持,且刻意保持独立。Track and compare every serious OpenClaw alternative — measured repo data, AI-written decision support, kept apart on purpose.

agent
luizalober/dengue-nn-systematic-review
Jupyter Notebook · 2026-08-14 评测基准 评测集 研究原型 Stars 0 周增 +0

用于复现 https://arxiv.org/abs/2106.12905 中图示所使用代码与数据的仓库。A repository to story the code and data used to create the figures shown in https://arxiv.org/abs/2106.12905

kabila5h/Shadow-API-scanner
Python · 2026-08-22 评测基准 评测集 实验 Stars 0 周增 +0

Shadow / Unmanaged API 的发现、分类、安全验证与可复现研究基准。Shadow / Unmanaged API discovery, classification, security validation, and reproducible research benchmarks

evaluationrisk
jan-steen/takeoff
TeX · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

无人机系统(UAS)在生态学研究中应用的系统综述。Systematic Review of UAS use in ecological research

InvestmentMDideas/GLASS-Data-Extraction-Benchmark
Python · 2026-08-16 评测基准 评测集 实验 Stars 0 周增 +0

面向系统综述数据抽取、证据定位与偏倚风险评估的冻结式(frozen)防泄漏基准。Frozen, leakage-aware benchmark for systematic-review data extraction, evidence localization, and risk-of-bias support

evaluationrisk
fsy2004/MetaWingman
Python · 2026-08-24 评测基准 评测集 实验 Stars 0 周增 +0

以问题为先、步骤可验证、可自我进化的 agent skill,用于系统综述与 Meta 分析。A question-first, step-verified, self-improving agent skill for systematic reviews and meta-analysis

agentllm-infra
friedpotato04/CUDA-L2
Cuda · 2026-08-22 LLM 基础设施 评测集 实验 Stars 0 周增 +0

🚀 使用强化学习优化半精度通用矩阵乘法(HGEMM)CUDA kernel,性能超越 cuBLAS 及其他基准。🚀 Optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels using reinforcement learning, surpassing cuBLAS and other benchmarks with superior performance.

evaluationllm-infra
DanceNitra/ramr
Python · 2026-08-12 RAG 检索增强 评测集 实验 Stars 0 周增 +0

RAMR —— 检索增强记忆可靠性:面向 Agentic-RAG / 记忆系统的抗污染合成基准(附方法与发现)。RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)

agentragevaluationllm-infra