面向 AI agent 的基于证据的评估——将每条断言与 agent 真实工具输出进行核对(受约束、基于证据的模型判断,而非整体式 LLM 评判的猜测),并附带置信区间。Evidence-grounded evaluation for AI agents — verifies each claim against the agent's real tool outputs (constrained, evidence-grounded model judgment, not holistic LLM-judge guesswork), with confidence intervals.
仓库/Skill 库
163 个 · 模型
邮件的大脑:MailFathom 将 IMAP 邮箱转变为自托管、AI-native 的服务。邮件同步至你自己的 PostgreSQL,建立索引以支持搜索与检索,并通过 Model Context Protocol 提供给 AI agent。具备语义检索、问答与受控的写入工具。基于 .NET 10,AGPL-3.0-only 许可。A brain for your mail: MailFathom turns IMAP mailboxes into a self-hosted, AI-native service. Mail synchronizes into your own PostgreSQL, is indexed for search and retrieval, and is served to AI agents over the Model Context Protocol. Semantic retrieval, answering, and gated write tools. .NET 10, AGPL-3.0-only.
UniArticles-An MCP (Model Context Protocol) tool that queries and curates new research papers from Scopus, arXiv and more using APIs/official libraries. 亿文通——一个MCP(模型上下文协议)工具,通过 API/官方库聚合检索并整理新文献。
面向律所的自托管、GDPR 合规 AI Agent。零云依赖、哈希链审计追溯(符合 AI Act 第 12 条)、自带模型、支持 9 种语言版本。willchen96/mike 的硬分叉,AGPL-3.0 许可。Self-hosted, GDPR-safe AI agent for law firms. Zero-cloud, hash-chained audit trail (AI Act art. 12), bring-your-own-model, 9 language editions. Hard fork of willchen96/mike, AGPL-3.0.
本地 Windows 桌面监视器,用于查看正在运行的 Claude Code agent —— 实时状态、预估成本、模型、主机以及一键聚焦,按项目分组。只读且完全离线。Local Windows desktop monitor for your running Claude Code agents - live status, estimated cost, model, host and one-click focus, grouped by project. Read-only and fully offline.
SharpAI 是基于 llama.cpp(通过 LlamaSharp)构建的可嵌入 Embedding、补全与模型管理平台,内置 Ollama 兼容的 Web 服务。SharpAI is an embeddable embeddings, completions, and model management platform using llama.cpp via LlamaSharp, with a built-in Ollama-compatible webserver.
运行 Ollama 本地 LLM 服务的 Docker 镜像。默认安全,所有 API 请求需 Bearer token(首次启动时自动生成)。OpenAI 兼容 API。支持首次启动模型预拉取、NVIDIA GPU (CUDA) 加速和持久化模型存储。多架构:amd64、arm64。Docker image to run an Ollama local LLM server. Secure by default, all API requests require a Bearer token (auto-generated on first start). OpenAI-compatible API. Supports first-start model pre-pull, NVIDIA GPU (CUDA) acceleration, and persistent model storage. Multi-arch: amd64, arm64.
Autobots | 由 NVIDIA NIM 驱动的去中心化模型 agent 集群。通过专用 6 文件控制架构编排高精度、端到端可鉴权软件开发流程,具备自动化安全审计、无冗余任务路由与自主状态管理能力。Autobots | A decentralized model agentic swarm powered by NVIDIA NIM. Orchestrating high-precision, end-to-end authenticated software development through a specialized 6-file control architecture. Features automated security auditing, non-redundant task routing, and autonomous state management.
Hearting —— 面向 Claude Code、Codex 与 OpenCode 的 agent 设置。一份可移植契约覆盖三者,支持能力路由、跨运行时密封调度、节点级模型分层与实时 Fleet 视图,前身为 agent_settingHearting — the agent setting for Claude Code, Codex, and OpenCode. One portable contract projected onto all three, with routed capabilities, sealed cross-harness dispatch, per-node model tiers, and a live Fleet view. Formerly agent_setting.
你的 agent 能写出变更——却无法告诉你还有哪些内容依赖于它正在修改的代码。本工具从你自己的代码构建产品模型,然后对每个任务打分:变更内容、影响范围、波及的用户旅程与敏感数据——让运算结果直观呈现在屏幕上。本地运行,零依赖,无遥测。包含 MCP。Your agent can write the change — it can't tell you what else depends on the code it's touching. This builds a model of your product from your own code, then scores every task: what it changes, what that reaches, which user journeys and sensitive data are in the blast radius — arithmetic on screen. Local, zero deps, no telemetry. MCP included.
可移植的 Rust Agent Client Protocol (ACP) server,具备模型路由、工具、权限、沙箱、会话与 MCP 集成。Portable Rust Agent Client Protocol (ACP) server with model routing, tools, permissions, sandboxing, sessions, and MCP integration.
Spring AI 的隐私护栏:跨模型/工具调用的请求作用域 PII 令牌化与最小特权披露,基于 Presidio 与 JVM 本地 OpenNLP。Privacy guardrails for Spring AI: request-scoped PII tokenization and least-privilege disclosure across model/tool calls, with Presidio and JVM-local OpenNLP.
CloserAI — 基于 DeepSeek Harness 构建的本地优先、模型无关、权限透明的桌面 AI 工作台。CloserAI - a local-first, model-agnostic, permission-transparent desktop AI workbench built on DeepSeek Harness.
面向 Model Context Protocol 服务器的 OWASP MCP Top 10 安全扫描器OWASP MCP Top 10 security scanner for Model Context Protocol servers
确定性的 MCP 安全架构。以 FrozenNamespace 作为 Model Context Protocol 工具验证的可信根(Root of Trust)。Deterministic MCP Security Architecture. FrozenNamespace as Root of Trust for Model Context Protocol tool verification
自托管、多用户转录平台:录制或上传音频。支持说话人标记、带时间戳的转录、跨录音识别说话人、摘要、提取行动项,并可使用自有 OpenAI 兼容 LLM 与转录内容对话。你的音频、你的服务器、你的模型。已在笔记本 RTX4070、台式机 RTX3090 与 RTX5090 上测试。Self-hosted, multi-user transcription platform: record or upload audio. Speaker-labeled, timestamped transcripts, Recognize speakers across recordings, Summarize, extract action items and chat over your transcripts with your own OpenAI-compatible LLM. Your Audio, your Server, your Model. Tested on Laptop RTX4070, Desktop RTX3090 and RTX5090
从零开始构建开源预训练 LLM,附带已发布的数据、训练代码、消融实验与结果,以高效提升模型质量。Build open pretraining LLMs from scratch with released data, training code, ablations, and results to improve model quality efficiently
用知识图谱刻画 AI/ML 模型从创建到部署的完整生命周期。Knowledge Graph to capture AI/ML model lifecycle from creation through deployments.
Model Context Protocol(MCP)的 Apache-2.0 安全扩展提案。包含对象级签名、新鲜度/重放保护、mTLS 传输、委托授权以及普通 MCP 服务器的 sidecar 保护。Rust workspace;安全审计发布于 docs/security/Apache-2.0 security extension proposal for the Model Context Protocol (MCP). Object-level signing, freshness/replay protection, mTLS transport, delegated authorization, sidecar protection of ordinary MCP servers. Rust workspace; published security audits under docs/security/.
Tox21 多终点毒性预测:可复现的研究与推理制品(冻结模型、FastAPI 服务、已审计的安全机制)Tox21 multi-endpoint toxicity prediction: a reproducible research and inference artifact (frozen model, FastAPI service, audited security)
你的 AI Agent 无法掏空的钱包。一个非托管的 Soroban 资金库,限定自主 Agent 的支出上限——由合约而非模型强制执行。The wallet your AI agent can't drain. A non-custodial Soroban treasury that bounds what an autonomous agent may spend - enforced by the contract, not the model.
自构建的 AI Agent 框架——多模型 ReAct + 分层缓存友好的上下文 + 异步多 Agent + XML 工作流。CLI 与 WebUI,零 LangChain。An AI agent framework that builds itself — multi-model ReAct + tiered cache-friendly context + async multi-agent + XML workflows. CLI & WebUI, zero LangChain.
面向生产环境的、有内存约束的跨模型 KV-cache 传输,具备受保护的回退机制与可复现研究工具链。Production-oriented, memory-bounded cross-model KV-cache transfer with guarded fallback and reproducible research tooling.
混合神经符号 AI:Llama 3.2 1B + 精确数学 + 经验证事实,807 MB 即时 CPU 原生回答,无需 GPU。封装 .aef 分发。Hybrid neuro-symbolic AI: Llama 3.2 1B + exact math + verified facts — instant CPU-native answers in 807 MB, no GPU. Sealed .aef distribution
论文《Controllable molecular graph generation from natural-language chemical constraints》(基于自然语言化学约束的可控分子图生成)的代码仓库。This is repository for "Controllable molecular graph generation from natural-language chemical constraints"
🤖 通过 Model Context Protocol 服务器在 IDE 与 CLI 工具间共享记忆并协调任务,增强 AI Agent 之间的协作🤖 Enhance collaboration among AI agents with a Model Context Protocol server that shares memory and coordinates tasks across IDEs and CLI tools.
AI it yourself——一款原生 iOS 聊天应用,Agent 可在其中构建你所需的小工具,沙箱化且立即可用。无服务器、无账号:自带 API key、本地机器或在 iPhone 上运行的模型。AI it yourself — a native iOS chat where agents build the small tools you need, sandboxed and usable immediately. No server, no account: bring your own API key, your own machine, or a model running on the iPhone.
阅读网络 meta 分析的证据结构。在浏览器中拟合图论模型并以八种方式绘制:精度几何、试验、证据流、贡献、带符号的研究级重构、不一致性的 Hodge 分解、扩散与弹簧图。Read the evidence structure of a network meta-analysis. Fits the graph-theoretical model in your browser and draws it eight ways: precision geometry, trials, evidence flow, contributions, a signed study-level reconstruction, a Hodge split of inconsistency, diffusion and springs.
开源的托管型 agent harness:持久化的 LLM agent session、每 session 一个 sandbox、支持无人值守运行——基础设施由你拥有,model key 由你掌握。An open-source managed agent harness: durable LLM agent sessions, a sandbox per session, and unattended runs — on infrastructure you own, with model keys you hold.
一款将模型命令沙箱化,并使用可编辑状态机处理长时间任务的编码 AgentA coding agent that jails model commands and uses editable state machines for long-running tasks
arXiv 论文的端到端推荐引擎:流式接入开放快照,使用量化 ONNX 模型对标题与摘要进行 embedding,构建 FAISS 向量索引,通过 FastAPI 服务提供推荐,并在 Streamlit 面板中交互探索。End-to-end recommendation engine for arXiv research papers: stream an open snapshot, embed titles + abstracts with a quantized ONNX model, build a FAISS vector index, serve recommendations through a FastAPI service, and explore them in a Streamlit dashboard.
Same Targets, Different Computation 的代码与 artifact:后训练如何在模型各层之间划分计算Code and artifacts for Same Targets, Different Computation: how post-training divides work across model layers.
将 Claude-English 重写为通俗英语。基于开源模型的 MCP 服务器、REST API 和 Claude Code hookRewrites Claude-English into plain English. MCP server, REST API and Claude Code hook, on an open-source model.
一个为 Model Context Protocol 提供策略执行能力的 gateway。聚合 MCP server,强制工具调用策略,扫描结果中的 prompt injection 与凭证泄露,并审计一切。Rust 实现。A policy-enforcing gateway for the Model Context Protocol. Aggregates MCP servers, enforces tool-call policy, scans results for prompt injection and leaked credentials, and audits everything. Rust.
多语言 AI agent 服务 — 基于单一 schema 驱动的 WebSocket 协议,提供知识对话、工具调用、持久检查点、人机协同与多方会话。基于 smooth-operator-core 引擎构建(5 种语言特性一致)。可部署到 Kubernetes、AWS serverless 或本地运行。托管于 lom.smoo.ai。Polyglot AI agent service — knowledge chat, tools, durable checkpoints, human-in-the-loop, and multi-participant conversations over one schema-driven WebSocket protocol. Built on the smooth-operator-core engine (5-language parity). Deploy to Kubernetes, AWS serverless, or run locally. Hosted at lom.smoo.ai.
面向大学规章的问答:引用具体条款、规章未规定时如实承认、两节冲突时发出告警。基于 FastAPI + 本地 embeddings + 免费 OpenRouter 模型。QA over a university rulebook that cites clauses, admits when the rulebook is silent, and flags when two sections disagree. FastAPI + local embeddings + free OpenRouter model.