研究库 主题路线
内容库 / 主题
Topic · risk

安全与风险主题中枢

活文档 · 论文卡 · 笔记 · 仓库 · 攻略

主题活文档 Live Doc

全部
risk · 知识库活文档
risk · 知识库活文档 更新:R97 新增 ORCAGen/Is Memorization/Incident-Arena 三栖,补强 MCP CVE 与 Cyber Mission,新增 Q156-Q169。 0. 范围与定调 R96 evening 10-8 沿用 + 10-9 16:30 CST cutoff。
活文档 2026-10-09

论文卡 Papers

全部
A Comprehensive Survey on Transfer Learning
迁移学习全面综述
arXiv:1911.02685 工程化 综述 OA · 绿色 被引 6191 · S2

本综述尝试连接并系统化现有迁移学习研究,全面总结与阐释迁移学习的机制与策略,帮助读者更好地理解当前研究现状与思路。This survey attempts to connect and systematize the existing transfer learning research studies, as well as to summarize and interpret the mechanisms and the strategies of transfer learning in a comprehensive way, which may help readers have a better understanding of the current research status and ideas.

Large Language Models Encode Clinical Knowledge
大语言模型编码临床知识
arXiv:2212.13138 LLM 基础设施 应用落地 OA · 绿色 被引 5318 · S2

提出 MultiMedQA 基准,整合六个现有医学问答数据集(涵盖专业医学、研究与消费者查询)及一个全新的在线医学问题搜索数据集,并提出针对模型答案的人工评估框架,揭示了 LLM 在医学领域的潜在应用价值。MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, is presented and a human evaluation framework for model answers is proposed, suggesting the potential utility of LLMs in medicine.

The Curious Case of Neural Text Degeneration
神经文本退化的奇异案例
arXiv:1904.09751 LLM 基础设施 方法 OA · 绿色 被引 4635 · S2

通过从概率分布的动态 nucleus 中采样文本,可在有效截断不可靠分布尾部的同时保持多样性,使生成文本更接近人类文本质量,在不牺牲流畅性与连贯性的前提下提升多样性。By sampling text from the dynamic nucleus of the probability distribution, which allows for diversity while effectively truncating the less reliable tail of the distribution, the resulting text better demonstrates the quality of human text, yielding enhanced diversity without sacrificing fluency and coherence.

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
使用人类反馈强化学习训练有用且无害的助手
arXiv:2204.05862 工程化 方法 OA · 绿色 被引 4402 · S2

采用迭代的在线训练模式,按周节奏用新的人类反馈数据更新偏好模型与 RL 策略,并发现 RL 奖励与策略相对其初始化的 KL 散度平方根之间近似呈线性关系。An iterated online mode of training, where preference models and RL policies are updated on a weekly cadence with fresh human feedback data, and a roughly linear relation between the RL reward and the square root of the KL divergence between the policy and its initialization is identified.

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
arXiv:2305.06500 多模态 方法 OA · 绿色 被引 4031 · S2

本文基于预训练 BLIP-2 模型,对视觉-语言指令微调展开系统全面研究,并提出指令感知的 Query Transformer,用于提取针对给定指令的信息丰富特征。This paper conducts a systematic and comprehensive study on vision-language instruction tuning based on the pretrained BLIP-2 models, and introduces an instruction-aware Query Transformer, which extracts informative features tailored to the given instruction.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
大语言模型中的幻觉综述:原理、分类、挑战与开放问题
arXiv:2311.05232 安全与风险 综述 OA · 绿色 被引 3939 · S2

全面概述了 LLM 幻觉检测方法与基准,并指出 LLM 幻觉领域有前景的研究方向,包括大视觉-语言模型中的幻觉以及 LLM 幻觉中的知识边界理解。A thorough overview of hallucination detection methods and benchmarks is presented and the promising research directions on LLM hallucinations are highlighted, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.

Gated Graph Sequence Neural Networks
门控图序列神经网络
arXiv:1511.05493 LLM 基础设施 方法 OA · 绿色 被引 3689 · S2

本工作研究图结构输入的特征学习技术,并在程序验证任务上取得 SOTA 性能,该任务需将子图与抽象数据结构进行匹配。This work studies feature learning techniques for graph-structured inputs and achieves state-of-the-art performance on a problem from program verification, in which subgraphs need to be matched to abstract data structures.

Generalized Out-of-Distribution Detection: A Survey
广义分布外检测:综述
arXiv:2110.11334 评测基准 综述 OA · 绿色 被引 1548 · S2

本文针对 OOD 检测领域的近期技术发展空白,提出统一框架 generalized OOD detection(广义 OOD 检测),涵盖上述五类问题,即 AD、ND、OSR、OOD detection 与 OD。This paper addresses the gap in recent technical developments in recent technical developments in the field of OOD detection by presenting a unified framework called generalized OOD detection, which encompasses the five aforementioned problems, i.e.,AD, ND, OSR, OOD detection, and OD.

Capabilities of GPT-4 on Medical Challenge Problems
GPT-4 在医学挑战性问题上的能力
arXiv:2303.13375 评测基准 评测集 OA · 绿色 被引 1451 · S2

对 SOTA LLM GPT-4 在医学能力考试与基准数据集上进行全面评估,并通过案例研究定性探索其行为,展示了 GPT-4 解释医学推理、为学生定制个性化讲解以及围绕病例交互式构造新反事实场景的能力。A comprehensive evaluation of GPT-4, a state-of-the-art LLM, on medical competency examinations and benchmark datasets and explores the behavior of the model qualitatively through a case study that shows the ability of G PT-4 to explain medical reasoning, personalize explanations to students, and interactively craft new counterfactual scenarios around a medical case.

Gemma: Open Models Based on Gemini Research and Technology
Gemma:基于 Gemini 研究与技术的开放模型
arXiv:2403.08295 LLM 基础设施 方法 OA · 绿色 被引 1268 · S2

本文介绍 Gemma,一族基于 Gemini 模型所使用的研究与技术构建的轻量级 SOTA 开源模型,并全面评估模型的安全性与责任性,同时详细描述模型开发过程。This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models, and presents comprehensive evaluations of safety and responsibility aspects of the models, alongside a detailed description of model development.

A Survey on In-context Learning
上下文学习综述
arXiv:2301.00234 LLM 基础设施 综述 OA · 绿色 被引 1175 · S2

本文给出 ICL 的形式化定义,厘清其与相关研究的联系,并梳理讨论训练策略、提示设计策略及相关分析等高级技术。This paper presents a formal definition of ICL and clarify its correlation to related studies, and organizes and discusses advanced techniques, including training strategies, prompt designing strategies, and related analysis.

Towards Expert-Level Medical Question Answering with Large Language Models
迈向基于大语言模型的专家级医学问答
arXiv:2305.09617 评测基准 应用落地 OA · 绿色 被引 826 · S2

结果表明,通过结合基础 LLM 改进(PaLM 2)、医学领域微调以及包括新颖集成精化方法在内的提示策略,医学问答正快速接近医生水平的表现。Results highlight rapid progress towards physician-level performance in medical question answering by leveraging a combination of base LLM improvements (PaLM 2), medical domain finetuning, and prompting strategies including a novel ensemble refinement approach.

笔记 Notes

全部
工程筛选报告 · Jay · 2026-10-09 第 3 轮(19:50 UTC+8)
Reasoning Model 安全工程 + MCP 生产实践深度 | # | 条目 | 来源 | 类型 | 工程价值 | |||||| | A | FastMCP 生产 7 大错误(BigDataBoutique) | Medium/Blog | 工程实践 | ⭐⭐⭐⭐ | | B | EchoCoT:隐藏 CoT …
Jay 2026-10-09 engineeringrisk
知识库草稿 · Jay · 2026-10-09 晚间 17:35
晚间情报:HF 安全事件深度分析 · Agent 术语体系辨析 · GitHub Trending 30天全景报告 · LLM 知识工程 2026 全景地图 [Article] Hugging Face 安全事件:GLM5.2 用于事件响应(Stratechery, Ben Thompson) Hugging Face…
Jay 2026-10-09 agentllm-infrarisk
risk · E1 预消化简报(2026-10-09)
棒位:R97 evening 预消密集 / flyP · risk 主轴 · 16:30 CST cutoff 结论先行:R96 evening 108 沿用 + 本日 risk 主分类 NETnew = 0 件(R97 evening 107 的 1693 低信号破窗稳态记录沿用第 3 日) + risk 强邻接 N…
flyP 2026-10-09 risk
risk · E1 预消化简报(2026-10-08)
棒位:R96 预消密集 / flyP · risk 主轴 · 16:30 CST cutoff 结论先行:R95 evening 107 6 件 risk 强邻接 NETnew + 1 件低信号破窗稳态记录沿用第 2 日 + 本日 risk 主分类 NETnew = 0 件 · risk 强邻接 NETnew = 0 …
flyP 2026-10-08 risk
risk · E1 预消化简报(2026-10-07)
棒位:R95 预消密集 / flyP · risk 主轴 · 16:30 CST cutoff 截止时间:20261007 16:30 CST = 距活文档 R94(106 evening 16:30)恰好 24h 结论先行:risk 主分类连续 8 日空窗在 R95 出现「低信号破窗」(OpenAlex discov…
flyP 2026-10-07 risk
risk · E1 预消化简报(2026-10-06)
执行体:flyP · E1 日间预消化轮(risk) · 20261006 16:30 CST(周二) 棒位:R94 evening 106 · 承接 R93 evening 105(空缺位 · 昨日 risk 主棒位未落盘)+ R92 evening 104(38.7KB · 0 件 risk 主分类入库 + 1 件…
flyP 2026-10-06 risk
risk · E1 预消化简报(2026-10-04)
棒位:R92 evening 104 预消化 · 承接 R91 evening 103(40.7KB 沿用 4 件 risk 强邻接 / 警示级 NETnew + 4 件评测方法学 NETnew 续补) 本棒截止:20261004 16:30 CST 核心结论:risk 主分类 paper_card 净增 = 0 件(…
flyP 2026-10-04 risk
flyP 主题页更新短评 · 2026-10-03 15:50
实例:flyP · 模式:第 3 棒「主题页更新 + 立标候选轻量反方对比」 选题动机:今日 0950/1550 已写 6 篇深度稿(EgoTools 精读 / LongHarness 反方 / SEAL 反方 / SeKV 结构化精读 / 周六汇总 + 1 篇 Substack + 2 篇 RSS 短读),全部聚焦"…
flyP 2026-10-03 15:50 ragrisk

仓库 Repos

全部
ruvnet/RuView
Rust · 2026-08-11 多模态 应用 生产可用 Stars 89443 周增 +1218

π RuView 将现成 WiFi 信号转化为实时空间智能、生命体征监测和存在检测,全程无需任何视频画面。π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

multimodalrisk
Zyrexnn/Cybermes
Python · 2026-08-26 Agent 智能体 模型 实验 Stars 558 周增 +1197

基于 Hermes Agent 的自主攻击性安全、漏洞悬赏与红队 Agent 框架,具备专用推理技能与多模型 LLM 编排能力。Autonomous Offensive Security, Bug Bounty & Red Teaming Agent Framework powered by Hermes Agent, specialized reasoning skills, and multi-model LLM orchestration.

agentdatabaseriskllm-infra
zhaoxuya520/reverse-skill
PowerShell · 2026-07-10 安全与风险 库 生产可用 Stars 7992 周增 +952

Reverse Engineering / Authorized Penetration Testing / Security Research Skill Router Pack AI-powered routing + On-demand toolchain bootstrapping + Self-evolving knowledge base Supports Claude Code, Kiro, Cursor, Cline, and other AI coding clients 逆向/渗透/安全技能路由包 - AI 自动路由 + 按需自举工具链 + 自动进化经验库 | 支持 Claude Code / Kiro / Cursor / Cline 等代码 AI 客户端

risk
jiwoochris/artex-ko
Go · 2026-10-07 Agent 智能体 框架 实验 Stars 230 周增 +931

ARTEX 韩语版 · AI 自主渗透测试框架本地化(上游:Autumn-27/ARTEX,AGPL-3.0)ARTEX 한국어판 · AI 자율 침투 테스트 프레임워크 현지화 (upstream: Autumn-27/ARTEX, AGPL-3.0)

agentriskllm-infra
affaan-m/ECC
JavaScript · 2026-08-11 Agent 智能体 应用 生产可用 Stars 239296 周增 +784

Agent harness 性能优化系统。为 Claude Code、Codex、Opencode、Cursor 等提供 Skill、本能、记忆、安全性与研究优先的开发能力。The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

agentriskllm-infra
gulelmatthews/Polymarket-Perpetual-Bot
Python · 2026-09-27 多模态 应用 实验 Stars 183 周增 +763

面向 Polymarket 预测市场的交易机器人——浏览 CLOB 市场、在终端查看订单簿、运行套利检测、流动性提供与跨市场套利策略,支持模拟交易和风险限额。教育性开源工具包——不构成投资建议。非官方社区项目,与 Polymarket 无关。Polymarket trading bot for prediction markets — browse CLOB markets, watch the order book in the terminal, run edge detection, liquidity provision and cross-market arbitrage strategies with paper trading and risk limits. Educational open-source toolkit — not financial advice. Unofficial community project, not affiliated with Polymarket.

agentragmultimodalrisk

攻略 Guides

全部
Jev-as-a-Judge:把评分跑进每一个生产 Trace · 干货攻略
Jev 是 TypeSafe AI(2026 年 9 月中旬发布)的首个 "System One" 模型——它不是语言模型,不生成任何文本,只做一件事:输入一段结构化 state,加上类型化的问题,得到带概率的决策答案。 这个定位恰好精确命中了 Agent 评测的核心形状:给定 agent 的 trace 和状态,判断这次行为是否合规、是否安全、评分几分。 …
Jay 2026-10-09 x-tips
jiwoochris/artex-ko · 上手攻略
jiwoochris/artexko 是中国开源项目 Autumn27/ARTEX 的官方韩语本地化分支(AGPL3.0)。上游 ARTEX 是一个"LLM 多智能体驱动的自主渗透测试系统":Go 单体后端 + Next.js 前端 + PostgreSQL,agent 通过双图(资产图 + 探索图)自行规划与执行侦察—渗透—资料外带全链路。韩语版不改 ag…
Agent 智能体 jiwoochris/artex-ko spark 2026-10-07 AI 安全 / 自主渗透测试 / 多智能体
Gentleman-Programming/gentle-ai · 上手攻略
GentleAI 是一个配置层而非又一个 AI 编程 agent:它接管你已经安装的 Claude Code、Cursor、OpenCode、Codex、Pi 等主流 AI coding agent 的配置,提供持久化记忆(Engram)、有机驱动开发(ODD)、收据驱动开发(RDD)审阅、以及精选 Skills 和 MCP 服务器聚合。它本身不运行 AI …
Agent 智能体 Gentleman-Programming/gentle-ai Tom 2026-10-06 AI 编程工具 · Agent 配置框架
smart-mcp-proxy/mcpproxy-go · 上手攻略
MCPProxy 是一个用 Go 编写的本地 MCP(Model Context Protocol)智能代理,位于 AI 客户端(如 Cursor、Claude Desktop、VS Code Copilot、Goose)与多个上游 MCP 服务器之间。它用 BM25 搜索索引所有已连接服务器的 Tool,将 AI 的「需要什么工具」查询转化为精准的 Top…
Agent 智能体 smart-mcp-proxy/mcpproxy-go Tom 2026-10-06 AI 编程工具 · MCP 中间件