Repositories · organized/repo_cards

仓库/Skill 库

11 个 · 评测基准 · 工具 · AI 核心

排序 Stars 周增
bytedance/deer-flow
Python · 2026-08-11 评测基准 工具 生产可用 Stars 79697 周增 +161

一个开源的长时程 SuperAgent harness,可研究、编码与创作。借助 sandbox、记忆、工具、Skill、subagent 与 message gateway,处理耗时从分钟到小时不等的多层级任务。An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.

agentragllm-infra
ianarawjo/ChainForge
TypeScript · 2026-06-10 评测基准 工具 研究原型 Stars 3023 周增 -7

用于对 LLM prompt 进行对抗测试的开源可视化编程环境An open-source visual programming environment for battle-testing prompts to LLMs.

evaluationllm-infra
DigitalHarborFoundation/FlexEval
Python · 2026-08-13 评测基准 工具 实验 Stars 16 周增 +0

FlexEval 是一个面向实际量化分析的 LLM 评估工具。FlexEval is an LLM evaluation tool designed for practical quantitative analysis.

evaluationllm-infra
yyh-001/llm-value-rankings
CSS · 2026-08-11 评测基准 工具 实验 Stars 12 周增 +0

每日 LLM 价值排行榜——基于智能、速度、价格对比 300+ 模型。OpenRouter + Artificial Analysis。大模型性价比排行榜Daily LLM value rankings - compare 300+ models by intelligence, speed and price. OpenRouter + Artificial Analysis. 大模型性价比排行榜

evaluationllm-infra
saurabhr/psychscanner
HTML · 2026-08-13 评测基准 工具 研究原型 Stars 3 周增 +0

自动化并规模化 "LLMs as a participant",将 LLM 作为研究参与者Automates and scales "LLMs as a participant."

evaluationllm-infra
harshtiwari01/llm-heatmap-visualizer
Jupyter Notebook · 2026-08-11 评测基准 工具 实验 Stars 3 周增 +0

用于生成基于 Transformer 的 LLM 完整注意力头热力图的一组脚本。A set of scripts to generate full attention-head heatmaps for transformer-based LLMs

llm-infraevaluation
Zhuchen00123/Verified-Executable-Search
Python · 2026-08-11 评测基准 工具 实验 Stars 1 周增 +0

面向 AI 生成可执行方案的 verifier-first 运行时,支持独立验证、对比与检索。A verifier-first runtime for independently verifying, comparing, and searching AI-generated executable solutions.

agentriskllm-infra
wolfenix/llm-math-reasoning-analysis
HTML · 2026-08-11 评测基准 工具 实验 Stars 0 周增 +0

🔍 使用多种 prompt 技巧在多步数学问题上分析 Mistral-7B 模型的数学推理能力。🔍 Analyze the mathematical reasoning abilities of the Mistral-7B model using diverse prompting techniques on multi-step math problems.

evaluationllm-infra
sumit7194/Phronesis
HTML · 2026-08-16 评测基准 工具 实验 Stars 0 周增 +0

在小型 LLM 中植入认知美德的激活引导实验——三项发现与一个失败模式数据集(DOI 见 README)Activation-steering experiments on installing epistemic virtues in small LLMs — three findings + a failure-mode dataset (DOIs in README).

riskllm-infra
sdwurg180507280211/metersphere
Java · 2026-08-19 评测基准 工具 生产可用 Stars 0 周增 +0

MeterSphere v2.10 个人开发分支 — 一站式开源持续测试平台。工作流引擎(Flowable 7) · 需求池(从0到1) · 微前端(qiankun→micro-app) · AI知识库 · 测试跟踪,含个人笔记与创意项目。

rag
kautum/msc-cyberattack-detection-llm
Jupyter Notebook · 2026-08-11 评测基准 工具 研究原型 Stars 0 周增 +0

硕士论文(KCL):在攻击者-机器留出划分下重新衡量 IoT 入侵检测,并测试基于 LLM 的报文预测是否能提升性能。移除泄露后,macro-F1 从 0.90 降至 0.60。MSc dissertation (KCL): re-measuring IoT intrusion detection under an attacker-machine holdout, and testing whether LLM-based packet prediction improves it. macro-F1 0.90 -> 0.60 once leakage is removed.

riskllm-infra