Papers · organized/paper_cards

论文

21 张论文卡片 · 安全与风险

开放获取 全部 绿色 · 677
SSGM框架(Stability and Safety-Governed Memory)
3. SSGM框架(Stability and Safety-Governed Memory)
arXiv:2603.11768 安全与风险 观点 OA · 绿色 被引 13 · S2

通过形式化分析与架构分解,展示 SSGM 如何缓解拓扑引发的知识泄漏(敏感上下文被固化到长期存储),以及有助于防止语义漂移(知识在迭代摘要中退化)。Through formal analysis and architectural decomposition, it is shown how SSGM can mitigate topology-induced knowledge leakage where sensitive contexts are solidified into long-term storage, and help prevent semantic drift where knowledge degrades through iterative summarization.

arXiv-3:A First Look at the Security Issues in the Model Context Protocol Ecosystem
arXiv-3:初探Model Context Protocol生态中的安全问题
arXiv:2510.16558 安全与风险 方法 被引 6 · S2

本文分析了六个公共注册表中共计67,057个服务器,识别出可导致服务器劫持与调用操控的普遍隐患,并实现了MCPInspect——一款集成前分析工具,可检测误导性的工具元数据与可利用的代码漏洞。This paper analyzes 67,057 servers across six public registries and identifies widespread conditions enabling server hijacking and invocation manipulation, and implements MCPInspect, a pre-integration analysis tool that detects misleading tool metadata and exploitable code vulnerabilities.

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
潜在世界模型中的决策度量对齐:用于 MPC 规划的诊断方法与动作条件目标
arXiv:2608.18746 安全与风险 方法 被引 0 · S2

动作条件目标改善了基于欧几里得代价与 CEM 的潜在 MPC 所使用的几何结构,DA-LeWM 在 LeWM 基础上增加了逆动力学和演示条件的目标-动作头,加速了收敛并取得比 LeWM 更高的在线成功率Action-conditioned objectives improve the geometry used by Euclidean-cost, CEM-based latent MPC, and DA-LeWM augments LeWM with inverse-dynamics and demonstration-conditioned goal-action heads, and accelerates convergence and achieves higher online success than LeWM.

1. Data Flow Control(DFC):AI Agent 数据安全策略的内核级执行框架
arXiv:2606.05679 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文将数据安全形式化为 provenance monomials 上的聚合谓词,并提出 Passant——一个无需物化 provenance 即可强制执行 DFC 策略的可移植查询重写层。This paper formalizes data safety as aggregate predicates over provenance monomials and presents Passant, a portable query rewriting layer that enforces DFC policies without materializing provenance.

4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety
4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety(⭐⭐⭐⭐⭐)
arXiv:2602.16666 安全与风险 方法 Open MIND OA · 绿色 被引 45 · S2

本工作提出 12 个具体指标,从一致性、鲁棒性、可预测性和安全性四个关键维度分解 Agent 可靠性,可与传统评估互补,并提供用于分析 Agent 表现、退化与失败方式的工具。This work proposes twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety, which complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
arXiv:2606.26015 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Tatoxa,一种面向鞑靼语文本去毒的 SOTA 系统;对比实验表明,该方法在关键质量指标上优于现有开源及商用闭源 LLM。Tatoxa is presented, a novel state-of-the-art system for text detoxification in the Tatar language, and comparative experiments show that the proposed approach outperforms existing open source and proprietary commercial LLMs on key quality metrics.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Qwen-RobotManip 技术报告:对齐释放机器人操作基础模型的规模化潜力
arXiv:2606.17846 安全与风险 方法 OA · 绿色 被引 25 · S2

Qwen-RobotManip 在所有 OOD 场景下大幅超越包括 π0.5 在内的已有 SOTA 模型,在 RoboChallenge 中排名第一,相对改进 20%,并在 AgileX ALOHA、Franka、UR、ARX 等真实机器人平台上完成验证。Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $\pi$0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.

IsabeLLM: Automated Theorem Proving Applied to Formally Verifying Consensus
IsabeLLM:将自动定理证明应用于共识协议的形式化验证
arXiv:2606.18098 安全与风险 方法 OA · 绿色 被引 1 · S2

实现了一个 RAG 框架,包含面向 LLM 的错误追踪与反例生成以提供更优上下文,并兼容最新版 Isabelle 与 Sledgehammer 以提升效率。A Retrieval-Augmented Generation framework, Error tracing and counterexample generation for improved context supplied to the Large Language Model, and Compatibility with the latest version of Isabelle and Sledgehammer is implemented for improved efficiency.

Pareto Optimal Re-ranking with Semi-Automated Content Credibility Detection
基于半自动化内容可信度检测的帕累托最优重排序
arXiv:2606.18031 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种基于优化的方法,通过精炼现有内容排序来提升社交媒体信息流中新闻内容的可信度;同时构建一条鲁棒的半自动化流水线,基于检索增强打分与人工事实核查的混合方式为内容赋予可信度分数。An optimization-based method to improve the credibility of news content on social media feeds by refining existing content rankings is presented and a robust semi-automated pipeline for assigning credibility scores to content based on a mixture of retrieval-augmented score assignments and human-generated fact-checks is proposed.

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
arXiv:2607.05910 安全与风险 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文提出 PolicyShiftGuard,一个紧凑的策略条件护栏,采用结合随机策略 SFT(RP-SFT)与边界对策略适配(BP-Adapt)的两阶段训练方案,并验证匹配的通过/拒绝边界对是稳定策略适配的关键。This work proposes PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt), and confirms that matched pass/block boundary pairs are essential for stable policy adaptation.

DeepLoop: Depth Scaling for Looped Transformers
DeepLoop:循环 Transformer 的深度扩展
arXiv:2607.13491 安全与风险 方法 OA · 绿色 被引 1 · S2

结果表明稳定的循环深度需要计入参数访问次数(而非仅名义层数)的残差缩放规则;DeepLoop 在不存在物理块被重复访问时表现为中性,一旦启用循环深度则改善验证损失与下游准确率。The results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count, and DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated.

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment
SUFLECA:面向 CAD-to-image 对齐的特征学习规模化方法
arXiv:2607.15058 安全与风险 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

SUFLECA (Scaling Up Feature LEarning for CAD Alignment),一个用于零样本 CAD 对齐的弱监督框架,贡献有二,并提出一种几何一致的匹配算法,可建立可靠的 CAD 到图像的一一对应关系。SUFLECA (Scaling Up Feature LEarning for CAD Alignment), a weakly-supervised framework for zero-shot CAD alignment with two key contributions, and a geometrically consistent matching algorithm that establishes reliable one-to-one CAD-to-image correspondences.

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
AI 海洋中的海妖之歌:大语言模型幻觉问题综述
arXiv:2309.01219 安全与风险 综述 OA · 绿色 被引 1128 · S2

本文给出了 LLM 幻觉现象与评估基准的分类体系,分析了现有缓解 LLM 幻觉的方法,并讨论了未来研究的潜在方向。This paper presents taxonomies of the LLM hallucination phenomena and evaluation benchmarks, analyzes existing approaches aiming at mitigating LLm hallucination, and discusses potential directions for future research.

Consciousness in Artificial Intelligence: Insights from the Science of Consciousness
人工智能中的意识:来自意识科学的洞察
arXiv:2308.08708 安全与风险 观点 OA · 绿色 被引 258 · S2

该报告主张并例证了一种严谨且基于经验的方法来研究 AI 意识:依据获得最佳支持的神经科学意识理论,详细评估现有 AI 系统。This report argues for, and exemplifies, a rigorous and empirically grounded approach to AI consciousness: assessing existing AI systems in detail, in light of best-supported neuroscientific theories of consciousness.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
大语言模型中的幻觉综述:原理、分类、挑战与开放问题
arXiv:2311.05232 安全与风险 综述 OA · 绿色 被引 3594 · S2

全面概述了 LLM 幻觉检测方法与基准,并指出 LLM 幻觉领域有前景的研究方向,包括大视觉-语言模型中的幻觉以及 LLM 幻觉中的知识边界理解。A thorough overview of hallucination detection methods and benchmarks is presented and the promising research directions on LLM hallucinations are highlighted, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.

Deep Visual-Semantic Alignments for Generating Image Descriptions
用于生成图像描述的深度视觉-语义对齐
arXiv:1412.2306 安全与风险 方法 OA · 绿色 被引 6111 · S2

提出一个模型,基于图像区域上的 CNN、句子上的双向 RNN 以及通过多模态嵌入对齐两种模态的结构化目标,生成图像及其区域的自然语言描述。A model that generates natural language descriptions of images and their regions based on a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding is presented.

Constitutional Midtraining: Content Presence Drives Alignment Gains
宪法式中训练:内容存在驱动对齐收益
arXiv:2607.26654 安全与风险 方法 被引 0 · S2

训练后对齐往往较浅,会在微调中被侵蚀。而中训练干预能否在干净隔离于训练后的情况下产生持久对齐,此前未经检验。我们通过宪法式中训练来测试:在 120B 规模上,插入基于原则与价值观的内容,与仅做回放的对照组进行对比。我们基于 Anthropic 的 Constitution 构建了 394M token 的宪法语料,并采用 2×2 析因设计(课程顺序 × 审慎推理),形成四种宪法式中训练条件与一组对照,随后在自生成与既有...Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and establis

MemSFT: Mitigating Alignment Tax with an External Parametric Memory
标题 -> 标题中文:MemSFT:借助外部参数化记忆缓解对齐税
arXiv:2607.25614 安全与风险 方法 被引 1 · S2

将 LLM 适配到专用领域常会带来对齐税:针对领域特定任务进行微调会导致灾难性遗忘,并显著降低在通用任务上的表现。我们提出 MemSFT,通过将领域专业化与主干参数更新解耦,以即插即用的参数化记忆来缓解对齐税。该记忆被训练为模仿在领域数据上运作的非参数化检索器,从而记住原本需通过检索获取的知识与模式。一旦在Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We propose MemSFT, which mitigates the alignment tax by decoupling domain specialization from backbone parameter updates through a plug-and-play parametric memory. The memory is trained to imitate the behavior of a non-parametric retriever operating over domain data, thereby memorizing knowledge and patterns that would otherwise be accessed through retrieval. Once trained on

Invariant Risk Minimization
不变风险最小化
arXiv:1907.02893 安全与风险 方法 OA · 绿色 被引 2940 · S2

本文提出了不变风险最小化(IRM),一种用于在多个训练分布上估计不变相关性的学习范式,并展示了 IRM 学到的不变性如何与数据的因果结构相关,从而实现分布外泛化。This work introduces Invariant Risk Minimization, a learning paradigm to estimate invariant correlations across multiple training distributions and shows how the invariances learned by IRM relate to the causal structures governing the data and enable out-of-distribution generalization.

Stealing Reasoning Traces from Proprietary LLM APIs
从商用 LLM API 窃取推理轨迹
arXiv:2608.09867 安全与风险 方法 被引 2 · S2

本文识别出一种绕过 anti-distillation 机制、允许攻击者窃取专有模型推理能力的架构漏洞,并提出具体的密码学与系统级缓解措施以保障客户端推理安全。An architectural vulnerability is identified that circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as well as proposing concrete cryptographic and system-level mitigations to secure client-side reasoning.

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
从不可听输入到模型失败:LALM 中的低频安全风险
arXiv:2608.09158 安全与风险 方法 被引 0 · S2

提出 Intermittent Low-Frequency Lockout(ILL),一种基于通用波形模板在黑盒设置下评估该风险的不可听 red teaming 方法;同时提出 Distributional Requery Guard(DRG),用于缓解该风险。This paper proposes Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting and proposes Distributional Requery Guard (DRG), an inaudible red teaming method to mitigate this risk.