Repositories · organized/repo_cards

仓库/Skill 库

7 个 · 评测基准 · 应用 · AI 核心

排序 Stars 周增
Arize-ai/phoenix
Python · 2026-08-11 评测基准 应用 生产可用 Stars 10982 周增 +49

AI 可观测性与评估。AI Observability & Evaluation

agentevaluationllm-infra
KCNyu/clawock
Python · 2026-08-11 评测基准 应用 实验 Stars 8 周增 +0

AI 辩论,代码裁定,亏损留在账面上。一款可移植的投资决策工作流插件与可验证 harness,已在真实的港股+美股组合上验证。AI argues. Code settles. The losses stay on the page. A portable investment decision-workflow plugin and verifiable harness, proven on a real HK + US portfolio.

agentriskllm-infra
yyxcnasd/amadeus-for-dsh
JavaScript · 2026-08-15 评测基准 应用 实验 Stars 1 周增 +0

Amadeus(来自《Steins;Gate 0》的 AI 助手)适配 DeepSeek Harness。Amadeus (AI assistant from Steins;Gate 0) for DeepSeek Harness

agent
NinjaSln-labs/dsh-plugins
TypeScript · 2026-08-16 评测基准 应用 实验 Stars 1 周增 +0

DeepSeek Harness 个人自研插件集:上下文罗盘 / 跨会话知识 / 子代理模型路由 / AI 生图(Personally developed plugins for DeepSeek Harness)

agentllm-infra
synaptiai/flow-harness
TypeScript · 2026-08-20 评测基准 应用 实验 Stars 0 周增 +0

与供应商无关的 coding-agent 框架,具备确定性工作流图、持久化证据与 fail-closed 沙箱执行。Provider-neutral coding-agent harness with deterministic workflow graphs, durable evidence, and fail-closed sandboxed execution

agentllm-infra
fmadore/IWAC-sentiment-analysis
Svelte · 2026-08-12 评测基准 应用 实验 Stars 0 周增 +0

对伊斯兰西非文献集(IWAC)语料库情感分析的交互式可视化,对比 ChatGPT、Gemini 与 Mistral,支持多语言与高级筛选。Interactive visualization of sentiment analysis on the Islam West Africa Collection (IWAC) corpus, comparing ChatGPT, Gemini, and Mistral with multilingual support and advanced filtering.

evaluationllm-infra
Dedebanded912/course-eligibility-hub
HTML · 2026-08-24 评测基准 应用 实验 Stars 0 周增 +0

即时评估课程资格,输出明确的通过/未通过结果以及定制化的入学测试。Instantly evaluate course eligibility with clear pass/fail results and tailored entry assessments.

multimodaldatabaseriskllm-infra