一个开源的长时程 SuperAgent harness,可研究、编码与创作。借助 sandbox、记忆、工具、Skill、subagent 与 message gateway,处理耗时从分钟到小时不等的多层级任务。An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
仓库/Skill 库
49 个 · 评测基准 · AI 核心
TensorZero 是一个开源 LLMOps 平台,统一了 LLM gateway、可观测性、评估、优化与实验TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
Evidently 是一个开源的 ML 和 LLM 可观测性框架,评估、测试和监控任何 AI 驱动的系统或数据 pipeline,覆盖从表格数据到 Gen AI 场景,提供 100+ 指标。Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
找到在你的硬件上真正能跑且性能最优的本地 LLM。排名基于真实且时新的基准测试,而非参数量。一条命令,即刻运行。Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
🧊 开源 LLM 可观测性平台,一行代码即可实现监控、评估与实验。YC W23 🍓🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
大型多模态模型 LMM 的开源评估工具包,支持 220+ LMM 与 80+ 基准测试。Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
开源 LLMOps 平台:集成 prompt playground、prompt 管理、LLM 评估和 LLM 可观测性。The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.
用于对 LLM prompt 进行对抗测试的开源可视化编程环境An open-source visual programming environment for battle-testing prompts to LLMs.
将 Pi 打造成 el Gentleman:一套高级架构师级开发 harness,集成 SDD/OpenSpec、子 Agent、严格 TDD 证据、review guardrail 与 skill 发现机制。Turn Pi into el Gentleman: a senior-architect development harness with SDD/OpenSpec, subagents, strict TDD evidence, review guardrails, and skill discovery.
一个面向动态文本分类的灵活、自适应分类系统A flexible, adaptive classification system for dynamic text classification
DeepSeek Harness 审批请求的 second-model AI 自动复核:只读复核子 Agent 返回结构化 allow/deny 判定及理由,默认 fail-closed,会话日志(approval/asked → autoReview/verdict → approval/decided)全程可审计。Second-model AI auto-review for DeepSeek Harness approval requests: a read-only reviewer subagent returns structured allow/deny verdicts with reasons, fail-closed by default, fully auditable from the session log (approval/asked -> autoReview/verdict -> approval/decided).
DeepSeek Harness 的可验证研究报告引擎:内容寻址的 evidence ledger(claim-snapshot 绑定、防篡改),以及版本化 sealed 报告,含逐条 claim 的验证判定与 manifest-sealed 目录。Verifiable research-report engine for DeepSeek Harness: content-addressed evidence ledger (claim-snapshot binding, tamper-evident) plus versioned sealed reports with per-claim verification verdicts and a manifest-sealed directory.
[COLM 2025] 论文 Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale——大语言模型在动态用户画像与大规模个性化响应任务上的基准测试。[COLM 2025] Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
Claude Fable 5 终极指南 2026:使用场景、集成与基准测试Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks
将生产事故转化为结构化的 9 段式 LLM 响应(严重等级、根因、缓解措施、事后复盘)。附带 5 个场景的回归测试集 + LLM-as-judge 评估流水线。Turn a production incident into a structured 9-section LLM response (severity, root cause, mitigation, postmortem). Ships with a 5-scenario regression suite + LLM-as-judge eval pipeline.
一个 LLM 相关论文、学位论文、工具、数据集、课程与基准的合集。A collection of LLM related papers, thesis, tools, datasets, courses, benchmarks
发布前双重检查:审视需求、测试实现、证明交付。面向 DeepSeek Harness 的工程纪律工具集。Double-check before you ship: grill the requirements, test the implementation, prove the delivery. An engineering-discipline bundle for DeepSeek Harness.
面向长时对话记忆层的综合基准测试框架A Comprehensive Benchmarking Framework for Long-Term Conversational Memory Layers
Bayes@N [ICLR'26]、Ranking LLMs [ACL'26 Main]:大语言模型的统计评估、比较与排序。Bayes@N [ICLR'26], Ranking LLMs [ACL'26 Main]: Statistical evaluation, comparison, and ranking of Large Language Models
每日 LLM 价值排行榜——基于智能、速度、价格对比 300+ 模型。OpenRouter + Artificial Analysis。大模型性价比排行榜Daily LLM value rankings - compare 300+ models by intelligence, speed and price. OpenRouter + Artificial Analysis. 大模型性价比排行榜
FlexEval 是一个面向实际量化分析的 LLM 评估工具。FlexEval is an LLM evaluation tool designed for practical quantitative analysis.
面向 ECG-语言模型(ELM)的研究型训练与评估框架A research-oriented training and evaluation framework for ECG-Language Models (ELMs)
AI 争论,代码结算,亏损留在账面上。一个由 Agent 驱动、面向香港及美国市场的真实券商账户:每个交易决策都必须经过辩论,并由模型无法触碰的代码完成结算。将同一决策工作流安装到你的 Agent:OpenClaw、Claude Code、Codex 或 DeepSeek Harness。AI argues. Code settles. The losses stay on the page. A real HK + US brokerage account run by agents that must debate every call, settled by code the model never touches. Install the same decision workflow into your own agent: OpenClaw, Claude Code, Codex, or DeepSeek Harness.
AI 系统作为理解智能机制的科学仪器AI Systems as Scientific Instruments for Understanding the Mechanisms of Intelligence
微调、评估、提示工程、开源模型Fine-tuning, evaluation, prompting, open-source models
自动化并规模化 "LLMs as a participant",将 LLM 作为研究参与者Automates and scales "LLMs as a participant."
实时更新的向量数据库项目、集成和基准评测全景图——每……刷新。Live-updating landscape of vector database projects, integrations, and benchmarks — refreshed every
用于生成基于 Transformer 的 LLM 完整注意力头热力图的一组脚本。A set of scripts to generate full attention-head heatmaps for transformer-based LLMs
面向 AI 生成可执行方案的 verifier-first 运行时,支持独立验证、对比与检索。A verifier-first runtime for independently verifying, comparing, and searching AI-generated executable solutions.
Amadeus(来自《Steins;Gate 0》的 AI 助手)适配 DeepSeek Harness。Amadeus (AI assistant from Steins;Gate 0) for DeepSeek Harness
[NeurIPS 2026] 基于释义感知评分的 LLM 不确定性量化 conformal prediction[NeurIPS 2026] Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification.
一个用 AI Harness 流程引导 Git 仓库的 CLI —— 基于固定课程骨架提供模板、文档门禁与特定语言的代码门禁。CLI, die ein Git-Repo mit dem AI-Harness-Prozess bootstrappt — Templates, Doc-Gates und sprachspezifische Code-Gates aus gepinnten Kurs-Skeletten.
DeepSeek Harness 个人自研插件集:上下文罗盘 / 跨会话知识 / 子代理模型路由 / AI 生图(Personally developed plugins for DeepSeek Harness)
🔍 在 Weaviate 混合检索中评测 embedding 模型,基于自有数据或 MTEB 数据集评估 MRR@K、Hit@K、延迟与内存占用🔍 Benchmark embedding models in hybrid search with Weaviate. Evaluate MRR@K, Hit@K, latency, and memory using your data or MTEB datasets.
精选的评分量表、检查清单、评分标准集、原则和评分指南汇总,用于对现代生成模型进行评分、排名、验证、过滤或训练。A curated list of rubrics, checklists, criteria sets, principles, and scoring guides used to score, rank, verify, filter, or train modern generative models.