一个开源的长时程 SuperAgent harness,可研究、编码与创作。借助 sandbox、记忆、工具、Skill、subagent 与 message gateway,处理耗时从分钟到小时不等的多层级任务。An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
仓库/Skill 库
49 个 · 评测基准 · 工具
用于对 LLM prompt 进行对抗测试的开源可视化编程环境An open-source visual programming environment for battle-testing prompts to LLMs.
Harness Anything - AI agent 控制中枢:支持 WPS、MS Office、Zotero、Photoshop、47 个 CLI 命令、27 项学术技能、SVG 转 PPTXHarness Anything - AI agent control hub: WPS, MS Office, Zotero, Photoshop, 47 CLI commands, 27 academic skills, SVG-to-PPTX
DeepSeek Harness (dsh) Windows / Linux 桌面客户端 — 内置 Node.js + dsh CLI,一键启动,10 套内置 UI 皮肤。EAC:Embracing All Creation 揽尽万象DeepSeek Harness (dsh) Windows / Linux desktop client - bundled Node.js + dsh CLI, one-click launch, 10 built-in UI skins. EAC: Embracing All Creation 揽尽万象
33 种在终端中识别 AI 生成文本的方法。提供 before/after 示例、草稿检查器,零依赖。33 ways to spot AI-written text, right in your terminal. Before/after examples, draft checker, zero dependencies.
长时程计算机使用 harness。让 AI Agent 在桌面应用与 CLI 中长时间运行,同时保持任务状态并在复杂工作流中可靠推进。具备全新上下文执行、持久化已验证状态、独立审计、可恢复进度以及原生 Claude Code / Codex / OpenClaw 集成。The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.
DeepSeek Harness (dsh) Windows 桌面客户端:内置 Node.js + dsh CLI,一键启动DeepSeek Harness (dsh) Windows desktop client - bundled Node.js + dsh CLI, one-click launch
Review Helper 是一款 AI 工具,帮助研究者快速摘要、整理和评估学术论文。它通过自动化关键任务来简化文献综述流程,使研究更高效、更准确。Review Helper is an AI tool that helps researchers quickly summarize, organize, and evaluate academic papers. It streamlines literature reviews by automating key tasks, making research more efficient and accurate.
为 DeepSeek Harness 提供有界、分层、需审批、可审计的跨会话记忆(能力接缝:ctx.memory + SQLite provider + memory tool + 冻结快照注入)。Bounded, layered, approval-gated, auditable cross-session memory for DeepSeek Harness (capability seam: ctx.memory + SQLite provider + memory tool + frozen snapshot injection)
系统综述的数据提取,直接引用自论文。向试验报告及其补充材料提交你的提取表单或 RoB 2、ROBINS-I、QUADAS-2、TIDieR 模板;Jev 指向具体行,每条答案均为带页码的逐字引用,由你核对后导出表格。文件保留在浏览器本地。Data extraction for systematic reviews, quoted from the papers. Ask a trial report and its supplements your extraction form or a RoB 2, ROBINS-I, QUADAS-2 or TIDieR template; Jev points at the lines, every answer is a verbatim quote with its page, you check it and export the table. Files stay in your browser.
每日 LLM 价值排行榜——基于智能、速度、价格对比 300+ 模型。OpenRouter + Artificial Analysis。大模型性价比排行榜Daily LLM value rankings - compare 300+ models by intelligence, speed and price. OpenRouter + Artificial Analysis. 大模型性价比排行榜
FlexEval 是一个面向实际量化分析的 LLM 评估工具。FlexEval is an LLM evaluation tool designed for practical quantitative analysis.
自动化并规模化 "LLMs as a participant",将 LLM 作为研究参与者Automates and scales "LLMs as a participant."
用于生成基于 Transformer 的 LLM 完整注意力头热力图的一组脚本。A set of scripts to generate full attention-head heatmaps for transformer-based LLMs
可复现研究代码:检验足球预测优势是否依赖于博彩公司去水方法。Reproducible research code testing whether football forecasting edges depend on bookmaker de-margining methods.
面向 AI 生成可执行方案的 verifier-first 运行时,支持独立验证、对比与检索。A verifier-first runtime for independently verifying, comparing, and searching AI-generated executable solutions.
一个用 AI Harness 流程引导 Git 仓库的 CLI —— 基于固定课程骨架提供模板、文档门禁与特定语言的代码门禁。CLI, die ein Git-Repo mit dem AI-Harness-Prozess bootstrappt — Templates, Doc-Gates und sprachspezifische Code-Gates aus gepinnten Kurs-Skeletten.
针对 LLM 集成证据筛选(系统综述标题/摘要筛选)的筛选-自验证内核。Screening-and-self-validation kernel for LLM-ensemble evidence screening (systematic review title/abstract screening).
ILRI 农业-食品系统气候适应证据综合与系统综述咨询仓库,聚焦"衡量关键:追踪小农户气候适应的有效性"。A repository for the ILRI consultancy on evidence synthesis and systematic reviews of climate adaptation in agri-food systems, focused on “Measuring what matters: tracking the effectiveness of climate adaptation for smallholder producers.”
🔍 使用多种 prompt 技巧在多步数学问题上分析 Mistral-7B 模型的数学推理能力。🔍 Analyze the mathematical reasoning abilities of the Mistral-7B model using diverse prompting techniques on multi-step math problems.
由文本挖掘驱动的科学文献综述。A text mining-driven review of scientific literature.
对两篇文章所获引用的质量审计:可复现的方法、数据、分析与报告Auditoria da qualidade das citações recebidas por dois artigos: método, dados, análises e relatório reprodutíveis
在小型 LLM 中植入认知美德的激活引导实验——三项发现与一个失败模式数据集(DOI 见 README)Activation-steering experiments on installing epistemic virtues in small LLMs — three findings + a failure-mode dataset (DOIs in README).
系统综述工具:面向学生自我关怀与学业功能的全文筛选与 Covidence 数据提取。Systematic-review tooling: full-text screening and Covidence data extraction for self-compassion and academic functioning in students
MeterSphere v2.10 个人开发分支 — 一站式开源持续测试平台。工作流引擎(Flowable 7) · 需求池(从0到1) · 微前端(qiankun→micro-app) · AI知识库 · 测试跟踪,含个人笔记与创意项目。
使用 R 进行荣誉研究项目的源码与统计分析。Source code and statistical analysis for my honours research project using R.
论文《Prompt Skill Matters: Examining GenAI Interaction and Writing Quality in L2 Academic Writing》的复现材料、统计分析脚本与补充数据,AI 介导的 L2 写作研究开放科学资源。Replication materials, statistical analysis scripts, and supplementary data for the study 'Prompt Skill Matters: Examining GenAI Interaction and Writing Quality in L2 Academic Writing'. Open science resource for AI-mediated L2 writing research.
计算两位独立评分者在系统评价所用 5 项方法学严谨性标准(检测性能、理论基础、计算效率、可解释性、生态相关性)上的 Cohen's kappa(非加权与二次加权)、一致率、平均绝对差及 Spearman 相关系数。Computes Cohen's kappa (unweighted & quadratic-weighted), percent agreement, mean absolute difference, and Spearman correlation between two independent rators' scores across 5 methodological rigour criteria (detection performance, theoretical foundation, computational efficiency,interpretability, ecological relevance) used in the systematic review.
一个基于 Web 的标注平台,用于手工标注与公共空间相关的 STOMP 文章。该应用支持人工标注员对文章进行分类、回答结构化研究问题,并导出标注结果以用于计算社会科学研究。A web-based annotation platform for manually labeling STOMP articles related to common spaces. The application enables human annotators to classify articles, answer structured research questions and export annotations for computational social science research.
使用 EMI Corpus of Student Academic Writing(Gablasova et al., 2024)中的 449 篇文章,检验写作者 L1 背景是否调节 22 项 USAS 语义域频率与作文分数之间的关系。How to use 449 essays from the EMI Corpus of Student Academic Writing (Gablasova et al., 2024) to test whether a writer's L1 background moderates the relationship between 22 USAS semantic-domain frequencies and essay scores.