Repositories · organized/repo_cards

仓库/Skill 库

40 个 · 评测基准 · 工具

排序 Stars 周增
bytedance/deer-flow
Python · 2026-08-11 评测基准 工具 生产可用 Stars 79697 周增 +161

一个开源的长时程 SuperAgent harness,可研究、编码与创作。借助 sandbox、记忆、工具、Skill、subagent 与 message gateway,处理耗时从分钟到小时不等的多层级任务。An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.

agentragllm-infra
ianarawjo/ChainForge
TypeScript · 2026-06-10 评测基准 工具 研究原型 Stars 3023 周增 -7

用于对 LLM prompt 进行对抗测试的开源可视化编程环境An open-source visual programming environment for battle-testing prompts to LLMs.

evaluationllm-infra
yb2460/harness-anything
Python · 2026-07-28 评测基准 工具 研究原型 Stars 1419 周增 +14

Harness Anything - AI agent 控制中枢:支持 WPS、MS Office、Zotero、Photoshop、47 个 CLI 命令、27 项学术技能、SVG 转 PPTXHarness Anything - AI agent control hub: WPS, MS Office, Zotero, Photoshop, 47 CLI commands, 27 academic skills, SVG-to-PPTX

agentengineering
zouyuxuan122/Deepseek-Harness-EAC
JavaScript · 2026-08-19 评测基准 工具 研究原型 Stars 875 周增 +0

DeepSeek Harness (dsh) Windows / Linux 桌面客户端 — 内置 Node.js + dsh CLI,一键启动,10 套内置 UI 皮肤。EAC:Embracing All Creation 揽尽万象DeepSeek Harness (dsh) Windows / Linux desktop client - bundled Node.js + dsh CLI, one-click launch, 10 built-in UI skins. EAC: Embracing All Creation 揽尽万象

agent
0xwilliamortiz/humanizer-cli
JavaScript · 2026-08-04 评测基准 工具 实验 Stars 588 周增 +7

33 种在终端中识别 AI 生成文本的方法。提供 before/after 示例、草稿检查器,零依赖。33 ways to spot AI-written text, right in your terminal. Before/after examples, draft checker, zero dependencies.

AMAP-ML/LongHorizon-Harness
Python · 2026-08-11 评测基准 工具 研究原型 Stars 571 周增 +0

长时程计算机使用 harness。让 AI Agent 在桌面应用与 CLI 中长时间运行,同时保持任务状态并在复杂工作流中可靠推进。具备全新上下文执行、持久化已验证状态、独立审计、可恢复进度以及原生 Claude Code / Codex / OpenClaw 集成。The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.

agentllm-infra
myYangyunfan/dsh_desktop
JavaScript · 2026-08-18 评测基准 工具 研究原型 Stars 463 周增 +315

DeepSeek Harness (dsh) Windows 桌面客户端:内置 Node.js + dsh CLI,一键启动DeepSeek Harness (dsh) Windows desktop client - bundled Node.js + dsh CLI, one-click launch

agent
clover-Saber/Review-Helper
JavaScript · 2026-01-25 评测基准 工具 实验 Stars 127 周增 +0

Review Helper 是一款 AI 工具,帮助研究者快速摘要、整理和评估学术论文。它通过自动化关键任务来简化文献综述流程,使研究更高效、更准确。Review Helper is an AI tool that helps researchers quickly summarize, organize, and evaluate academic papers. It streamlines literature reviews by automating key tasks, making research more efficient and accurate.

PerryLink/dsh-memento
JavaScript · 2026-08-18 评测基准 工具 实验 Stars 57 周增 +0

为 DeepSeek Harness 提供有界、分层、需审批、可审计的跨会话记忆(能力接缝:ctx.memory + SQLite provider + memory tool + 冻结快照注入)。Bounded, layered, approval-gated, auditable cross-session memory for DeepSeek Harness (capability seam: ctx.memory + SQLite provider + memory tool + frozen snapshot injection)

agentdatabasellm-infra
DigitalHarborFoundation/FlexEval
Python · 2026-08-13 评测基准 工具 实验 Stars 16 周增 +0

FlexEval 是一个面向实际量化分析的 LLM 评估工具。FlexEval is an LLM evaluation tool designed for practical quantitative analysis.

evaluationllm-infra
yyh-001/llm-value-rankings
CSS · 2026-08-11 评测基准 工具 实验 Stars 12 周增 +0

每日 LLM 价值排行榜——基于智能、速度、价格对比 300+ 模型。OpenRouter + Artificial Analysis。大模型性价比排行榜Daily LLM value rankings - compare 300+ models by intelligence, speed and price. OpenRouter + Artificial Analysis. 大模型性价比排行榜

evaluationllm-infra
bgcarlisle/Numbat
PHP · 2026-08-20 评测基准 工具 生产可用 Stars 8 周增 +0

Numbat Systematic Review ManagerNumbat Systematic Review Manager

saurabhr/psychscanner
HTML · 2026-08-13 评测基准 工具 研究原型 Stars 3 周增 +0

自动化并规模化 "LLMs as a participant",将 LLM 作为研究参与者Automates and scales "LLMs as a participant."

evaluationllm-infra
harshtiwari01/llm-heatmap-visualizer
Jupyter Notebook · 2026-08-11 评测基准 工具 实验 Stars 3 周增 +0

用于生成基于 Transformer 的 LLM 完整注意力头热力图的一组脚本。A set of scripts to generate full attention-head heatmaps for transformer-based LLMs

llm-infraevaluation
Bardakor/The-De-margining-Artifact
Python · 2026-08-18 评测基准 工具 研究原型 Stars 3 周增 +0

可复现研究代码:检验足球预测优势是否依赖于博彩公司去水方法。Reproducible research code testing whether football forecasting edges depend on bookmaker de-margining methods.

Zhuchen00123/Verified-Executable-Search
Python · 2026-08-11 评测基准 工具 实验 Stars 1 周增 +0

面向 AI 生成可执行方案的 verifier-first 运行时,支持独立验证、对比与检索。A verifier-first runtime for independently verifying, comparing, and searching AI-generated executable solutions.

agentriskllm-infra
herbertkokholm/attest
Python · 2026-08-25 评测基准 工具 研究原型 Stars 1 周增 +0

针对 LLM 集成证据筛选(系统综述标题/摘要筛选)的筛选-自验证内核。Screening-and-self-validation kernel for LLM-ensemble evidence screening (systematic review title/abstract screening).

llm-infra
bristlepine/ilri-climate-adaptation-effectiveness
HTML · 2026-08-19 评测基准 工具 研究原型 Stars 1 周增 +0

ILRI 农业-食品系统气候适应证据综合与系统综述咨询仓库,聚焦"衡量关键:追踪小农户气候适应的有效性"。A repository for the ILRI consultancy on evidence synthesis and systematic reviews of climate adaptation in agri-food systems, focused on “Measuring what matters: tracking the effectiveness of climate adaptation for smallholder producers.”

wolfenix/llm-math-reasoning-analysis
HTML · 2026-08-11 评测基准 工具 实验 Stars 0 周增 +0

🔍 使用多种 prompt 技巧在多步数学问题上分析 Mistral-7B 模型的数学推理能力。🔍 Analyze the mathematical reasoning abilities of the Mistral-7B model using diverse prompting techniques on multi-step math problems.

evaluationllm-infra
WENHA0ZHANG/text_mining-based_literature_reviewer
Jupyter Notebook · 2026-08-20 评测基准 工具 研究原型 Stars 0 周增 +0

由文本挖掘驱动的科学文献综述。A text mining-driven review of scientific literature.

sumit7194/Phronesis
HTML · 2026-08-16 评测基准 工具 实验 Stars 0 周增 +0

在小型 LLM 中植入认知美德的激活引导实验——三项发现与一个失败模式数据集(DOI 见 README)Activation-steering experiments on installing epistemic virtues in small LLMs — three findings + a failure-mode dataset (DOIs in README).

riskllm-infra
sm0704/abstractScreening
Python · 2026-08-21 评测基准 工具 研究原型 Stars 0 周增 +0

系统综述工具:面向学生自我关怀与学业功能的全文筛选与 Covidence 数据提取。Systematic-review tooling: full-text screening and Covidence data extraction for self-compassion and academic functioning in students

sdwurg180507280211/metersphere
Java · 2026-08-19 评测基准 工具 生产可用 Stars 0 周增 +0

MeterSphere v2.10 个人开发分支 — 一站式开源持续测试平台。工作流引擎(Flowable 7) · 需求池(从0到1) · 微前端(qiankun→micro-app) · AI知识库 · 测试跟踪,含个人笔记与创意项目。

rag
russellmiller49/Systematic_review
TypeScript · 2026-08-21 评测基准 工具 实验 Stars 0 周增 +0
researchmaredi/honours-project
R · 2026-08-22 评测基准 工具 研究原型 Stars 0 周增 +0

使用 R 进行荣誉研究项目的源码与统计分析。Source code and statistical analysis for my honours research project using R.

Pegi1727/Prompt-Skill-Matters-L2-Academic-Writing
Python · 2026-08-09 评测基准 工具 研究原型 Stars 0 周增 +0

论文《Prompt Skill Matters: Examining GenAI Interaction and Writing Quality in L2 Academic Writing》的复现材料、统计分析脚本与补充数据,AI 介导的 L2 写作研究开放科学资源。Replication materials, statistical analysis scripts, and supplementary data for the study 'Prompt Skill Matters: Examining GenAI Interaction and Writing Quality in L2 Academic Writing'. Open science resource for AI-mediated L2 writing research.

engineering
nicoleoyetunji/Inter-rator-Reliability
Jupyter Notebook · 2026-08-21 评测基准 工具 研究原型 Stars 0 周增 +0

计算两位独立评分者在系统评价所用 5 项方法学严谨性标准(检测性能、理论基础、计算效率、可解释性、生态相关性)上的 Cohen's kappa(非加权与二次加权)、一致率、平均绝对差及 Spearman 相关系数。Computes Cohen's kappa (unweighted & quadratic-weighted), percent agreement, mean absolute difference, and Spearman correlation between two independent rators' scores across 5 methodological rigour criteria (detection performance, theoretical foundation, computational efficiency,interpretability, ecological relevance) used in the systematic review.

MilSmo/https---github.com-MilSmo-systematic-review-audit
Python · 2026-08-16 评测基准 工具 研究原型 Stars 0 周增 +0
medria/stomp-annotation
Python · 2026-08-15 评测基准 工具 生产可用 Stars 0 周增 +0

一个基于 Web 的标注平台,用于手工标注与公共空间相关的 STOMP 文章。该应用支持人工标注员对文章进行分类、回答结构化研究问题,并导出标注结果以用于计算社会科学研究。A web-based annotation platform for manually labeling STOMP articles related to common spaces. The application enables human annotators to classify articles, answer structured research questions and export annotations for computational social science research.

klzbmmt/EMI-USAS-scoring-fairness
Python · 2026-08-23 评测基准 工具 研究原型 Stars 0 周增 +0

使用 EMI Corpus of Student Academic Writing(Gablasova et al., 2024)中的 449 篇文章,检验写作者 L1 背景是否调节 22 项 USAS 语义域频率与作文分数之间的关系。How to use 449 essays from the EMI Corpus of Student Academic Writing (Gablasova et al., 2024) to test whether a writer's L1 background moderates the relationship between 22 USAS semantic-domain frequencies and essay scores.

engineering
kautum/msc-cyberattack-detection-llm
Jupyter Notebook · 2026-08-11 评测基准 工具 研究原型 Stars 0 周增 +0

硕士论文(KCL):在攻击者-机器留出划分下重新衡量 IoT 入侵检测,并测试基于 LLM 的报文预测是否能提升性能。移除泄露后,macro-F1 从 0.90 降至 0.60。MSc dissertation (KCL): re-measuring IoT intrusion detection under an attacker-machine holdout, and testing whether LLM-based packet prediction improves it. macro-F1 0.90 -> 0.60 once leakage is removed.

riskllm-infra
j-vaught/tex-complexity
Python · 2026-08-14 评测基准 工具 研究原型 Stars 0 周增 +0

在 LaTeX 文档中分析英文散文,评估可读性与结构复杂度。Analyze English prose in LaTeX documents for readability and structural complexity.

engineering
g0dswer/audit-scientific-papers
Python · 2026-08-20 评测基准 工具 研究原型 Stars 0 周增 +0

一个可复用的 skill,用于对临床研究、系统综述和 meta-analyses 进行严谨且有来源支持的评估与分析。A reusable skill for rigorous, source-backed appraisal and analysis of clinical studies, systematic reviews, and meta-analyses.

Elaine-764/Statistical-Properties-of-Financial-Market-Returns
MATLAB · 2026-08-21 评测基准 工具 研究原型 Stars 0 周增 +0

研究问题:流动性金融资产的相关性与波动率随时间有多稳定?常用的高斯性质在真实市场数据中是否成立?Research Questions: How stable are correlations and volatility over time in liquid financial assets? Do commonly assumed Gaussian properties hold for real market data?

dzweben/ABCD-screen-media-review
HTML · 2026-08-13 评测基准 工具 研究原型 Stars 0 周增 +0

ABCD Study 屏幕媒介出版物的系统综述:检索、筛选与效应量审计。Systematic review of ABCD Study screen media publications: search, screening, and effect size audit

dionysia-debug/RGEP-Systematic-review-and-meta-analysis
未知语言 · 2026-08-17 评测基准 工具 研究原型 Stars 0 周增 +0

本仓库包含用于系统综述和 meta 分析的全部代码,数据来自用 IMU 数据构建 ML 模型预测帕金森症状的论文。This repository contains all the codes used to systematically review and meta-analyse the data reported in papers where ML models have been built to predict Parkinson's symptoms using IMU data.