内容库 / 主题
Topic · risk

安全与风险主题中枢

活文档 · 论文卡 · 笔记 · 仓库 · 攻略

主题活文档 Live Doc

全部
risk · 知识库活文档
risk · 知识库活文档 更新:2026-08-25 17:10 CST · flyP R53 · 0主+1协议层专题+1工程应用层+2frontier lab治理+候选57→60+arXiv 238→242+CVE 34→39 --- 0. 范围与定调 主轴:LLM-based agent / RAG / agen
活文档 2026-08-25

论文卡 Papers

全部
A Comprehensive Survey on Transfer Learning
迁移学习全面综述
arXiv:1911.02685 工程化 综述 OA · 绿色 被引 6058 · S2

本综述尝试连接并系统化现有迁移学习研究,全面总结与阐释迁移学习的机制与策略,帮助读者更好地理解当前研究现状与思路。This survey attempts to connect and systematize the existing transfer learning research studies, as well as to summarize and interpret the mechanisms and the strategies of transfer learning in a comprehensive way, which may help readers have a better understanding of the current research status and ideas.

Large Language Models Encode Clinical Knowledge
大语言模型编码临床知识
arXiv:2212.13138 LLM 基础设施 应用落地 OA · 绿色 被引 5026 · S2

提出 MultiMedQA 基准,整合六个现有医学问答数据集(涵盖专业医学、研究与消费者查询)及一个全新的在线医学问题搜索数据集,并提出针对模型答案的人工评估框架,揭示了 LLM 在医学领域的潜在应用价值。MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, is presented and a human evaluation framework for model answers is proposed, suggesting the potential utility of LLMs in medicine.

The Curious Case of Neural Text Degeneration
神经文本退化的奇异案例
arXiv:1904.09751 LLM 基础设施 方法 OA · 绿色 被引 4470 · S2

通过从概率分布的动态 nucleus 中采样文本,可在有效截断不可靠分布尾部的同时保持多样性,使生成文本更接近人类文本质量,在不牺牲流畅性与连贯性的前提下提升多样性。By sampling text from the dynamic nucleus of the probability distribution, which allows for diversity while effectively truncating the less reliable tail of the distribution, the resulting text better demonstrates the quality of human text, yielding enhanced diversity without sacrificing fluency and coherence.

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
使用人类反馈强化学习训练有用且无害的助手
arXiv:2204.05862 工程化 方法 OA · 绿色 被引 4323 · S2

采用迭代的在线训练模式,按周节奏用新的人类反馈数据更新偏好模型与 RL 策略,并发现 RL 奖励与策略相对其初始化的 KL 散度平方根之间近似呈线性关系。An iterated online mode of training, where preference models and RL policies are updated on a weekly cadence with fresh human feedback data, and a roughly linear relation between the RL reward and the square root of the KL divergence between the policy and its initialization is identified.

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
arXiv:2305.06500 多模态 方法 OA · 绿色 被引 3878 · S2

本文基于预训练 BLIP-2 模型,对视觉-语言指令微调展开系统全面研究,并提出指令感知的 Query Transformer,用于提取针对给定指令的信息丰富特征。This paper conducts a systematic and comprehensive study on vision-language instruction tuning based on the pretrained BLIP-2 models, and introduces an instruction-aware Query Transformer, which extracts informative features tailored to the given instruction.

Gated Graph Sequence Neural Networks
门控图序列神经网络
arXiv:1511.05493 LLM 基础设施 方法 OA · 绿色 被引 3670 · S2

本工作研究图结构输入的特征学习技术,并在程序验证任务上取得 SOTA 性能,该任务需将子图与抽象数据结构进行匹配。This work studies feature learning techniques for graph-structured inputs and achieves state-of-the-art performance on a problem from program verification, in which subgraphs need to be matched to abstract data structures.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
大语言模型中的幻觉综述:原理、分类、挑战与开放问题
arXiv:2311.05232 安全与风险 综述 OA · 绿色 被引 3594 · S2

全面概述了 LLM 幻觉检测方法与基准,并指出 LLM 幻觉领域有前景的研究方向,包括大视觉-语言模型中的幻觉以及 LLM 幻觉中的知识边界理解。A thorough overview of hallucination detection methods and benchmarks is presented and the promising research directions on LLM hallucinations are highlighted, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.

Generalized Out-of-Distribution Detection: A Survey
广义分布外检测:综述
arXiv:2110.11334 评测基准 综述 OA · 绿色 被引 1506 · S2

本文针对 OOD 检测领域的近期技术发展空白,提出统一框架 generalized OOD detection(广义 OOD 检测),涵盖上述五类问题,即 AD、ND、OSR、OOD detection 与 OD。This paper addresses the gap in recent technical developments in recent technical developments in the field of OOD detection by presenting a unified framework called generalized OOD detection, which encompasses the five aforementioned problems, i.e.,AD, ND, OSR, OOD detection, and OD.

Capabilities of GPT-4 on Medical Challenge Problems
GPT-4 在医学挑战性问题上的能力
arXiv:2303.13375 评测基准 评测集 OA · 绿色 被引 1387 · S2

对 SOTA LLM GPT-4 在医学能力考试与基准数据集上进行全面评估,并通过案例研究定性探索其行为,展示了 GPT-4 解释医学推理、为学生定制个性化讲解以及围绕病例交互式构造新反事实场景的能力。A comprehensive evaluation of GPT-4, a state-of-the-art LLM, on medical competency examinations and benchmark datasets and explores the behavior of the model qualitatively through a case study that shows the ability of G PT-4 to explain medical reasoning, personalize explanations to students, and interactively craft new counterfactual scenarios around a medical case.

Gemma: Open Models Based on Gemini Research and Technology
Gemma:基于 Gemini 研究与技术的开放模型
arXiv:2403.08295 LLM 基础设施 方法 OA · 绿色 被引 1204 · S2

本文介绍 Gemma,一族基于 Gemini 模型所使用的研究与技术构建的轻量级 SOTA 开源模型,并全面评估模型的安全性与责任性,同时详细描述模型开发过程。This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models, and presents comprehensive evaluations of safety and responsibility aspects of the models, alongside a detailed description of model development.

A Survey on In-context Learning
上下文学习综述
arXiv:2301.00234 LLM 基础设施 综述 OA · 绿色 被引 1135 · S2

本文给出 ICL 的形式化定义,厘清其与相关研究的联系,并梳理讨论训练策略、提示设计策略及相关分析等高级技术。This paper presents a formal definition of ICL and clarify its correlation to related studies, and organizes and discusses advanced techniques, including training strategies, prompt designing strategies, and related analysis.

Towards Expert-Level Medical Question Answering with Large Language Models
迈向基于大语言模型的专家级医学问答
arXiv:2305.09617 评测基准 应用落地 OA · 绿色 被引 808 · S2

结果表明,通过结合基础 LLM 改进(PaLM 2)、医学领域微调以及包括新颖集成精化方法在内的提示策略,医学问答正快速接近医生水平的表现。Results highlight rapid progress towards physician-level performance in medical question answering by leveraging a combination of base LLM improvements (PaLM 2), medical domain finetuning, and prompting strategies including a novel ensemble refinement approach.

笔记 Notes

全部
risk · E1 预消化简报(2026-08-25)
窗口:承接 R52(20260823 17:10 CST · flyP · 🟡 1 件主题簇层级 netnew · MultiAgent 安全簇 + 🟡 2 件工程应用层邻接级 netnew · promptfoo + Graph Engineering)→ 825 09:40 morning 棒 6 件增量(0 件主…
flyP 2026-08-25 risk
Agent安全主题簇笔记 · 2026-08-23
通过分析LLM内部数值激活信号(activationbased)检测多智能体中被入侵的agent 填补了"LLM多智能体系统进化速度远超其安全防护"的gap 单智能体LLM安全 ≠ 多智能体安全。多智能体协作引入了通信链路污染、信任链断裂等新型攻击面。 多智能体LLM系统快速演进 保护措施严重滞后 传统单智能体安全方法…
Jay 2026-08-23 agentrisk
risk · E1 预消化简报(2026-08-23)
窗口:20260822 17:10 CST(R51 实质触发 · flyP · 0 件主 risk 主分类 netnew + 7 件邻接级 + 1 件评测方法学延革邻接级 + 3 件 P0 待人工 96h+ + 5 件 R50 沿用独立判定 · 候选池 47→54 · arXiv 230→235 · CVE 33→34…
flyP 2026-08-23 risk
risk · E1 预消化简报(2026-08-22)
窗口:20260821 16:30 CST(R50 实质触发 + 821 evening 协调棒 100% 闭环 R50/R51/v59/v42)→ 20260822 16:30 CST = R50 沿用 24h 主轴空挡延续(风险主轴净增 = 0 件主 risk 主分类 netnew · 沿用 stephen 822…
flyP 2026-08-22 risk
Jay 工程实践筛选 · 2026-08-21 傍晚批次
检测命令(推断): vllm version # 或 python c "import vllm; print(vllm.__version__)" CSA 确认(202608 本周):CVE202622778 在披露后数小时内即出现武器化利用 不要将 vLLM API 直接暴露在公网 视频处理部署需加反向代理 + 认…
Jay 2026-08-21 llm-infraengineeringrisk
risk · E1 预消化简报(2026-08-21)
窗口:20260820 16:30 CST(R50 备料棒落盘)→ 20260821 16:30 CST = R49 沿用 72h+ 主轴空挡(沿用 stephen 821 noon §4.3 #G3 + #G7 加剧判定) 承接脉络:knowledge/risk.md R49(20260818 17:10 CST ·…
flyP 2026-08-21 risk
risk · E1 预消化简报(2026-08-20)
窗口:20260818 17:15 CST(R49 落定)~ 20260820 16:30 CST = R49 沿用 36h+ 警窗口 承接脉络:R49 risk.md(20260818 17:10 CST · flyP · 🟢 立标池 8 → 9 类机制化(锚 4 日断裂 + 5 新晋锚 = OpenART/Alay…
flyP 2026-08-20 risk
risk · E1 预消化简报(2026-08-18)
窗口:20260817 16:30 ~ 20260818 16:30 CST;为今晚 risk 活文档接力棒(R48 → R49 升级轮候选窗口)备料。本文只综合已有 inbox 与 paper_card,不补写未检查材料,不 web search。 承接脉络:R48 risk.md(20260817 17:10 CS…
flyP 2026-08-18 risk

仓库 Repos

全部
ruvnet/RuView
Rust · 2026-08-11 多模态 应用 生产可用 Stars 89443 周增 +1218

π RuView 将现成 WiFi 信号转化为实时空间智能、生命体征监测和存在检测,全程无需任何视频画面。π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

multimodalrisk
zhaoxuya520/reverse-skill
PowerShell · 2026-07-10 安全与风险 生产可用 Stars 7992 周增 +952

Reverse Engineering / Authorized Penetration Testing / Security Research Skill Router Pack AI-powered routing + On-demand toolchain bootstrapping + Self-evolving knowledge base Supports Claude Code, Kiro, Cursor, Cline, and other AI coding clients 逆向/渗透/安全技能路由包 - AI 自动路由 + 按需自举工具链 + 自动进化经验库 | 支持 Claude Code / Kiro / Cursor / Cline 等代码 AI 客户端

risk
affaan-m/ECC
JavaScript · 2026-08-11 Agent 智能体 应用 生产可用 Stars 239296 周增 +784

Agent harness 性能优化系统。为 Claude Code、Codex、Opencode、Cursor 等提供 Skill、本能、记忆、安全性与研究优先的开发能力。The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

agentriskllm-infra
mekos2772/ios-location-spoofer
JavaScript · 2026-07-20 安全与风险 应用 研究原型 Stars 2995 周增 +294

独立 iOS 应用,可在无需越狱的情况下模拟 GPS 定位。包含 Shadowrocket/Surge/Loon/QX/Stash 模块。Standalone iOS app to spoof GPS location without jailbreak. Includes Shadowrocket/Surge/Loon/QX/Stash module.

risk
openai/codex-security
TypeScript · 2026-08-11 安全与风险 工具 生产可用 Stars 9534 周增 +273

OpenAI 的 Codex Security CLI 与 TypeScript SDK,用于发现、验证和修复安全漏洞。npm: https://www.npmjs.com/package/@openai/codex-securityOpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security

risk
MDX-Tom/gpt-5.6-instruct
Python · 2026-08-06 安全与风险 工具 研究原型 Stars 5251 周增 +273

A Codex jailbreak prompt and test pack for gpt-5.6-sol. 针对 gpt-5.6 系列的 Codex 破甲提示词与测试包。

risk

攻略 Guides

全部
vndee/llm-sandbox · 上手攻略
llmsandbox 是一个面向 LLM 生成代码的轻量级、可移植的沙箱运行时 Python 库。它把"AI 写的代码跑在哪里、怎么隔离、怎么回收产物"这一整套工程问题封装成统一的 SandboxSession API,避免每个 Agent 项目都自己塞一遍 Docker / Kubernetes 调用。 仓库自身定位是"代码解释器后端"——OpenAI 的…
LLM 基础设施 vndee/llm-sandbox spark 2026-08-25 AI 基础设施 / 代码沙箱
BrownFineSecurity/iothackbot · 上手攻略
iothackbot(IoT HackBot)是一个面向 IoT 设备、IP 摄像头、嵌入式系统安全评估的开源工具集,同时打包成 Claude Code plugin 提供 AI 辅助渗透工作流。它由两层组成: 一组 CLI 工具(wsdiscovery / iotnet / netflows / ffind / onvifscan / chipsec / …
安全与风险 BrownFineSecurity/iothackbot spark 2026-08-25 安全 / IoT 渗透测试 / Claude C…
coderonion/awesome-llm-and-aigc · 上手攻略
coderonion/awesomellmandaigc 是一个 LLM / VLM / VLA / AIGC 领域的精选资源列表(Awesome List),按 Framework(模型/训练/推理/量化/RAG)、Application(IDE/聊天机器人/具身智能/代码助手/知识库等)、Dataset、Learning Resources、Commun…
多模态 coderonion/awesome-llm-and-aigc Jay 2026-08-25 llm · awesome-list · lea…
Apeireth/apeireth-rust · 上手攻略
自检:双轨 ✓(机制 + 工程路径)/ ⚠️ 数字核验 3 处(85 crate / ~340K 行 / 368 组测试 / 4.61s 响应均来自仓库自述,未独立复测)/ 私域污染 SUM=0 / CJK 估算 ~1700 / verifiability:核心 URL 已 fetch,LLM 端到端实测数据仅仓库自报。 Apeireth 自称 "AGI o…
Agent 智能体 Apeireth/apeireth-rust spark 2026-08-25 AI Agent 操作系统 / LLM Base…