AI 可观测性与评估。AI Observability & Evaluation
仓库/Skill 库
12 个 · 评测基准 · 应用 · AI 核心
将 Pi 打造成 el Gentleman:一套高级架构师级开发 harness,集成 SDD/OpenSpec、子 Agent、严格 TDD 证据、review guardrail 与 skill 发现机制。Turn Pi into el Gentleman: a senior-architect development harness with SDD/OpenSpec, subagents, strict TDD evidence, review guardrails, and skill discovery.
DeepSeek Harness 的可验证研究报告引擎:内容寻址的 evidence ledger(claim-snapshot 绑定、防篡改),以及版本化 sealed 报告,含逐条 claim 的验证判定与 manifest-sealed 目录。Verifiable research-report engine for DeepSeek Harness: content-addressed evidence ledger (claim-snapshot binding, tamper-evident) plus versioned sealed reports with per-claim verification verdicts and a manifest-sealed directory.
将生产事故转化为结构化的 9 段式 LLM 响应(严重等级、根因、缓解措施、事后复盘)。附带 5 个场景的回归测试集 + LLM-as-judge 评估流水线。Turn a production incident into a structured 9-section LLM response (severity, root cause, mitigation, postmortem). Ships with a 5-scenario regression suite + LLM-as-judge eval pipeline.
发布前双重检查:审视需求、测试实现、证明交付。面向 DeepSeek Harness 的工程纪律工具集。Double-check before you ship: grill the requirements, test the implementation, prove the delivery. An engineering-discipline bundle for DeepSeek Harness.
Bayes@N [ICLR'26]、Ranking LLMs [ACL'26 Main]:大语言模型的统计评估、比较与排序。Bayes@N [ICLR'26], Ranking LLMs [ACL'26 Main]: Statistical evaluation, comparison, and ranking of Large Language Models
AI 争论,代码结算,亏损留在账面上。一个由 Agent 驱动、面向香港及美国市场的真实券商账户:每个交易决策都必须经过辩论,并由模型无法触碰的代码完成结算。将同一决策工作流安装到你的 Agent:OpenClaw、Claude Code、Codex 或 DeepSeek Harness。AI argues. Code settles. The losses stay on the page. A real HK + US brokerage account run by agents that must debate every call, settled by code the model never touches. Install the same decision workflow into your own agent: OpenClaw, Claude Code, Codex, or DeepSeek Harness.
Amadeus(来自《Steins;Gate 0》的 AI 助手)适配 DeepSeek Harness。Amadeus (AI assistant from Steins;Gate 0) for DeepSeek Harness
DeepSeek Harness 个人自研插件集:上下文罗盘 / 跨会话知识 / 子代理模型路由 / AI 生图(Personally developed plugins for DeepSeek Harness)
与供应商无关的 coding-agent 框架,具备确定性工作流图、持久化证据与 fail-closed 沙箱执行。Provider-neutral coding-agent harness with deterministic workflow graphs, durable evidence, and fail-closed sandboxed execution
交互式仪表盘,对比五个 LLM(GPT-5.6 Luna、Mistral Small 4、DeepSeek v4 Flash、Gemma 4 31B 与 Qwen3.8 27B)对 12,349 篇西非法语区新闻文章的标注结果,包含评分者一致性、差异分析与盲法仲裁。法/英双语;此前的三轮三模型标注活动保留归档。Interactive dashboard comparing how five LLMs — GPT-5.6 Luna, Mistral Small 4, DeepSeek v4 Flash, Gemma 4 31B and Qwen3.8 27B — annotated 12,349 francophone West African press articles, with inter-rater agreement, discrepancy analysis and blind arbitration. Bilingual FR/EN; the earlier three-model campaign stays archived.
即时评估课程资格,输出明确的通过/未通过结果以及定制化的入学测试。Instantly evaluate course eligibility with clear pass/fail results and tailored entry assessments.