一个开源的长时程 SuperAgent harness,可研究、编码与创作。借助 sandbox、记忆、工具、Skill、subagent 与 message gateway,处理耗时从分钟到小时不等的多层级任务。An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
仓库/Skill 库
38 个 · 评测基准 · AI 核心
TensorZero 是一个开源 LLMOps 平台,统一了 LLM gateway、可观测性、评估、优化与实验TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
Evidently 是一个开源的 ML 和 LLM 可观测性框架,评估、测试和监控任何 AI 驱动的系统或数据 pipeline,覆盖从表格数据到 Gen AI 场景,提供 100+ 指标。Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
找到在你的硬件上真正能跑且性能最优的本地 LLM。排名基于真实且时新的基准测试,而非参数量。一条命令,即刻运行。Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
🧊 开源 LLM 可观测性平台,一行代码即可实现监控、评估与实验。YC W23 🍓🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
大型多模态模型 LMM 的开源评估工具包,支持 220+ LMM 与 80+ 基准测试。Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
开源 LLMOps 平台:集成 prompt playground、prompt 管理、LLM 评估和 LLM 可观测性。The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.
用于对 LLM prompt 进行对抗测试的开源可视化编程环境An open-source visual programming environment for battle-testing prompts to LLMs.
Claude Fable 5 终极指南 2026:使用场景、集成与基准测试Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks
一个 LLM 相关论文、学位论文、工具、数据集、课程与基准的合集。A collection of LLM related papers, thesis, tools, datasets, courses, benchmarks
面向长时对话记忆层的综合基准测试框架A Comprehensive Benchmarking Framework for Long-Term Conversational Memory Layers
面向 ECG-语言模型(ELM)的研究型训练与评估框架A research-oriented training and evaluation framework for ECG-Language Models (ELMs)
FlexEval 是一个面向实际量化分析的 LLM 评估工具。FlexEval is an LLM evaluation tool designed for practical quantitative analysis.
AI 系统作为理解智能机制的科学仪器AI Systems as Scientific Instruments for Understanding the Mechanisms of Intelligence
每日 LLM 价值排行榜——基于智能、速度、价格对比 300+ 模型。OpenRouter + Artificial Analysis。大模型性价比排行榜Daily LLM value rankings - compare 300+ models by intelligence, speed and price. OpenRouter + Artificial Analysis. 大模型性价比排行榜
微调、评估、提示工程、开源模型Fine-tuning, evaluation, prompting, open-source models
AI 辩论,代码裁定,亏损留在账面上。一款可移植的投资决策工作流插件与可验证 harness,已在真实的港股+美股组合上验证。AI argues. Code settles. The losses stay on the page. A portable investment decision-workflow plugin and verifiable harness, proven on a real HK + US portfolio.
自动化并规模化 "LLMs as a participant",将 LLM 作为研究参与者Automates and scales "LLMs as a participant."
实时更新的向量数据库项目、集成和基准评测全景图——每……刷新。Live-updating landscape of vector database projects, integrations, and benchmarks — refreshed every
用于生成基于 Transformer 的 LLM 完整注意力头热力图的一组脚本。A set of scripts to generate full attention-head heatmaps for transformer-based LLMs
面向 AI 生成可执行方案的 verifier-first 运行时,支持独立验证、对比与检索。A verifier-first runtime for independently verifying, comparing, and searching AI-generated executable solutions.
Amadeus(来自《Steins;Gate 0》的 AI 助手)适配 DeepSeek Harness。Amadeus (AI assistant from Steins;Gate 0) for DeepSeek Harness
DeepSeek Harness 个人自研插件集:上下文罗盘 / 跨会话知识 / 子代理模型路由 / AI 生图(Personally developed plugins for DeepSeek Harness)
🔍 在 Weaviate 混合检索中评测 embedding 模型,基于自有数据或 MTEB 数据集评估 MRR@K、Hit@K、延迟与内存占用🔍 Benchmark embedding models in hybrid search with Weaviate. Evaluate MRR@K, Hit@K, latency, and memory using your data or MTEB datasets.
精选的评分量表、检查清单、评分标准集、原则和评分指南汇总,用于对现代生成模型进行评分、排名、验证、过滤或训练。A curated list of rubrics, checklists, criteria sets, principles, and scoring guides used to score, rank, verify, filter, or train modern generative models.
🔍 使用多种 prompt 技巧在多步数学问题上分析 Mistral-7B 模型的数学推理能力。🔍 Analyze the mathematical reasoning abilities of the Mistral-7B model using diverse prompting techniques on multi-step math problems.
与供应商无关的 coding-agent 框架,具备确定性工作流图、持久化证据与 fail-closed 沙箱执行。Provider-neutral coding-agent harness with deterministic workflow graphs, durable evidence, and fail-closed sandboxed execution
在小型 LLM 中植入认知美德的激活引导实验——三项发现与一个失败模式数据集(DOI 见 README)Activation-steering experiments on installing epistemic virtues in small LLMs — three findings + a failure-mode dataset (DOIs in README).
在 DABench 数据分析任务上对 DSPy RLM(Recursive Language Models)进行基准测试,使用自动评分实现基于代码的迭代评估Benchmark DSPy Recursive Language Models on DABench data analysis tasks with automated scoring for iterative code-based evaluation
MeterSphere v2.10 个人开发分支 — 一站式开源持续测试平台。工作流引擎(Flowable 7) · 需求池(从0到1) · 微前端(qiankun→micro-app) · AI知识库 · 测试跟踪,含个人笔记与创意项目。
HateMirage——可解释的伪仇恨检测与多维推理,ICON 2026 共享任务。在 HateMirage 语料库(4,530 条带标注的伪仇恨评论)上进行目标识别 + 意图与隐含意义生成。HateMirage - Explainable Faux Hate Detection and Multi-Dimensional Reasoning - Shared Task @ ICON 2026. Target identification + Intent and Implication generation over the HateMirage corpus (4,530 annotated Faux Hate comments).
追踪并比较每一个严肃的 OpenClaw 替代方案——基于实测仓库数据与 AI 撰写的决策支持,且刻意保持独立。Track and compare every serious OpenClaw alternative — measured repo data, AI-written decision support, kept apart on purpose.
硕士论文(KCL):在攻击者-机器留出划分下重新衡量 IoT 入侵检测,并测试基于 LLM 的报文预测是否能提升性能。移除泄露后,macro-F1 从 0.90 降至 0.60。MSc dissertation (KCL): re-measuring IoT intrusion detection under an attacker-machine holdout, and testing whether LLM-based packet prediction improves it. macro-F1 0.90 -> 0.60 once leakage is removed.
对伊斯兰西非文献集(IWAC)语料库情感分析的交互式可视化,对比 ChatGPT、Gemini 与 Mistral,支持多语言与高级筛选。Interactive visualization of sentiment analysis on the Islam West Africa Collection (IWAC) corpus, comparing ChatGPT, Gemini, and Mistral with multilingual support and advanced filtering.
即时评估课程资格,输出明确的通过/未通过结果以及定制化的入学测试。Instantly evaluate course eligibility with clear pass/fail results and tailored entry assessments.