Papers · organized/paper_cards

论文

15 张论文卡片 · 安全与风险 · 方法

开放获取 全部 绿色 · 724
arXiv-3:A First Look at the Security Issues in the Model Context Protocol Ecosystem
arXiv-3:初探Model Context Protocol生态中的安全问题
arXiv:2510.16558 安全与风险 方法 被引 6 · S2

本文分析了六个公共注册表中共计67,057个服务器,识别出可导致服务器劫持与调用操控的普遍隐患,并实现了MCPInspect——一款集成前分析工具,可检测误导性的工具元数据与可利用的代码漏洞。This paper analyzes 67,057 servers across six public registries and identifies widespread conditions enabling server hijacking and invocation manipulation, and implements MCPInspect, a pre-integration analysis tool that detects misleading tool metadata and exploitable code vulnerabilities.

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
潜在世界模型中的决策度量对齐:用于 MPC 规划的诊断方法与动作条件目标
arXiv:2608.18746 安全与风险 方法 被引 0 · S2

动作条件目标改善了基于欧几里得代价与 CEM 的潜在 MPC 所使用的几何结构,DA-LeWM 在 LeWM 基础上增加了逆动力学和演示条件的目标-动作头,加速了收敛并取得比 LeWM 更高的在线成功率Action-conditioned objectives improve the geometry used by Euclidean-cost, CEM-based latent MPC, and DA-LeWM augments LeWM with inverse-dynamics and demonstration-conditioned goal-action heads, and accelerates convergence and achieves higher online success than LeWM.

1. Data Flow Control(DFC):AI Agent 数据安全策略的内核级执行框架
arXiv:2606.05679 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文将数据安全形式化为 provenance monomials 上的聚合谓词,并提出 Passant——一个无需物化 provenance 即可强制执行 DFC 策略的可移植查询重写层。This paper formalizes data safety as aggregate predicates over provenance monomials and presents Passant, a portable query rewriting layer that enforces DFC policies without materializing provenance.

4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety
4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety(⭐⭐⭐⭐⭐)
arXiv:2602.16666 安全与风险 方法 Open MIND OA · 绿色 被引 45 · S2

本工作提出 12 个具体指标,从一致性、鲁棒性、可预测性和安全性四个关键维度分解 Agent 可靠性,可与传统评估互补,并提供用于分析 Agent 表现、退化与失败方式的工具。This work proposes twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety, which complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
arXiv:2606.26015 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Tatoxa,一种面向鞑靼语文本去毒的 SOTA 系统;对比实验表明,该方法在关键质量指标上优于现有开源及商用闭源 LLM。Tatoxa is presented, a novel state-of-the-art system for text detoxification in the Tatar language, and comparative experiments show that the proposed approach outperforms existing open source and proprietary commercial LLMs on key quality metrics.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Qwen-RobotManip 技术报告:对齐释放机器人操作基础模型的规模化潜力
arXiv:2606.17846 安全与风险 方法 OA · 绿色 被引 25 · S2

Qwen-RobotManip 在所有 OOD 场景下大幅超越包括 π0.5 在内的已有 SOTA 模型,在 RoboChallenge 中排名第一,相对改进 20%,并在 AgileX ALOHA、Franka、UR、ARX 等真实机器人平台上完成验证。Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $\pi$0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.

IsabeLLM: Automated Theorem Proving Applied to Formally Verifying Consensus
IsabeLLM:将自动定理证明应用于共识协议的形式化验证
arXiv:2606.18098 安全与风险 方法 OA · 绿色 被引 1 · S2

实现了一个 RAG 框架,包含面向 LLM 的错误追踪与反例生成以提供更优上下文,并兼容最新版 Isabelle 与 Sledgehammer 以提升效率。A Retrieval-Augmented Generation framework, Error tracing and counterexample generation for improved context supplied to the Large Language Model, and Compatibility with the latest version of Isabelle and Sledgehammer is implemented for improved efficiency.

Pareto Optimal Re-ranking with Semi-Automated Content Credibility Detection
基于半自动化内容可信度检测的帕累托最优重排序
arXiv:2606.18031 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种基于优化的方法,通过精炼现有内容排序来提升社交媒体信息流中新闻内容的可信度;同时构建一条鲁棒的半自动化流水线,基于检索增强打分与人工事实核查的混合方式为内容赋予可信度分数。An optimization-based method to improve the credibility of news content on social media feeds by refining existing content rankings is presented and a robust semi-automated pipeline for assigning credibility scores to content based on a mixture of retrieval-augmented score assignments and human-generated fact-checks is proposed.

DeepLoop: Depth Scaling for Looped Transformers
DeepLoop:循环 Transformer 的深度扩展
arXiv:2607.13491 安全与风险 方法 OA · 绿色 被引 1 · S2

结果表明稳定的循环深度需要计入参数访问次数(而非仅名义层数)的残差缩放规则;DeepLoop 在不存在物理块被重复访问时表现为中性,一旦启用循环深度则改善验证损失与下游准确率。The results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count, and DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated.

Deep Visual-Semantic Alignments for Generating Image Descriptions
用于生成图像描述的深度视觉-语义对齐
arXiv:1412.2306 安全与风险 方法 OA · 绿色 被引 6111 · S2

提出一个模型,基于图像区域上的 CNN、句子上的双向 RNN 以及通过多模态嵌入对齐两种模态的结构化目标,生成图像及其区域的自然语言描述。A model that generates natural language descriptions of images and their regions based on a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding is presented.

Constitutional Midtraining: Content Presence Drives Alignment Gains
宪法式中训练:内容存在驱动对齐收益
arXiv:2607.26654 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

训练后对齐往往较浅,会在微调中被侵蚀。而中训练干预能否在干净隔离于训练后的情况下产生持久对齐,此前未经检验。我们通过宪法式中训练来测试:在 120B 规模上,插入基于原则与价值观的内容,与仅做回放的对照组进行对比。我们基于 Anthropic 的 Constitution 构建了 394M token 的宪法语料,并采用 2×2 析因设计(课程顺序 × 审慎推理),形成四种宪法式中训练条件与一组对照,随后在自生成与既有...Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and establis

MemSFT: Mitigating Alignment Tax with an External Parametric Memory
标题 -> 标题中文:MemSFT:借助外部参数化记忆缓解对齐税
arXiv:2607.25614 安全与风险 方法 被引 1 · S2

将 LLM 适配到专用领域常会带来对齐税:针对领域特定任务进行微调会导致灾难性遗忘,并显著降低在通用任务上的表现。我们提出 MemSFT,通过将领域专业化与主干参数更新解耦,以即插即用的参数化记忆来缓解对齐税。该记忆被训练为模仿在领域数据上运作的非参数化检索器,从而记住原本需通过检索获取的知识与模式。一旦在Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We propose MemSFT, which mitigates the alignment tax by decoupling domain specialization from backbone parameter updates through a plug-and-play parametric memory. The memory is trained to imitate the behavior of a non-parametric retriever operating over domain data, thereby memorizing knowledge and patterns that would otherwise be accessed through retrieval. Once trained on

Invariant Risk Minimization
不变风险最小化
arXiv:1907.02893 安全与风险 方法 OA · 绿色 被引 2940 · S2

本文提出了不变风险最小化(IRM),一种用于在多个训练分布上估计不变相关性的学习范式,并展示了 IRM 学到的不变性如何与数据的因果结构相关,从而实现分布外泛化。This work introduces Invariant Risk Minimization, a learning paradigm to estimate invariant correlations across multiple training distributions and shows how the invariances learned by IRM relate to the causal structures governing the data and enable out-of-distribution generalization.

Stealing Reasoning Traces from Proprietary LLM APIs
从商用 LLM API 窃取推理轨迹
arXiv:2608.09867 安全与风险 方法 被引 2 · S2

本文识别出一种绕过 anti-distillation 机制、允许攻击者窃取专有模型推理能力的架构漏洞,并提出具体的密码学与系统级缓解措施以保障客户端推理安全。An architectural vulnerability is identified that circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as well as proposing concrete cryptographic and system-level mitigations to secure client-side reasoning.

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
从不可听输入到模型失败:LALM 中的低频安全风险
arXiv:2608.09158 安全与风险 方法 被引 0 · S2

提出 Intermittent Low-Frequency Lockout(ILL),一种基于通用波形模板在黑盒设置下评估该风险的不可听 red teaming 方法;同时提出 Distributional Requery Guard(DRG),用于缓解该风险。This paper proposes Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting and proposes Distributional Requery Guard (DRG), an inaudible red teaming method to mitigate this risk.