研究库 论文知识库
Papers · organized/paper_cards

论文

38 张论文卡片 · 安全与风险

开放获取 全部 绿色 · 1640
SSGM框架(Stability and Safety-Governed Memory)
3. SSGM框架(Stability and Safety-Governed Memory)
arXiv:2603.11768 安全与风险 观点 OA · 绿色 被引 21 · S2

通过形式化分析与架构分解,展示 SSGM 如何缓解拓扑引发的知识泄漏(敏感上下文被固化到长期存储),以及有助于防止语义漂移(知识在迭代摘要中退化)。Through formal analysis and architectural decomposition, it is shown how SSGM can mitigate topology-induced knowledge leakage where sensitive contexts are solidified into long-term storage, and help prevent semantic drift where knowledge degrades through iterative summarization.

arXiv-3:A First Look at the Security Issues in the Model Context Protocol Ecosystem
arXiv-3:初探Model Context Protocol生态中的安全问题
arXiv:2510.16558 安全与风险 方法 被引 10 · S2

本文分析了六个公共注册表中共计67,057个服务器,识别出可导致服务器劫持与调用操控的普遍隐患,并实现了MCPInspect——一款集成前分析工具,可检测误导性的工具元数据与可利用的代码漏洞。This paper analyzes 67,057 servers across six public registries and identifies widespread conditions enabling server hijacking and invocation manipulation, and implements MCPInspect, a pre-integration analysis tool that detects misleading tool metadata and exploitable code vulnerabilities.

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
潜在世界模型中的决策度量对齐:用于 MPC 规划的诊断方法与动作条件目标
arXiv:2608.18746 安全与风险 方法 OA · 绿色 被引 7 · S2

动作条件目标改善了基于欧几里得代价与 CEM 的潜在 MPC 所使用的几何结构,DA-LeWM 在 LeWM 基础上增加了逆动力学和演示条件的目标-动作头,加速了收敛并取得比 LeWM 更高的在线成功率Action-conditioned objectives improve the geometry used by Euclidean-cost, CEM-based latent MPC, and DA-LeWM augments LeWM with inverse-dynamics and demonstration-conditioned goal-action heads, and accelerates convergence and achieves higher online success than LeWM.

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
StepGuard:基于可扩展监督与安全-效用平衡的步骤级护栏学习
arXiv:2608.24777 安全与风险 方法 OA · 绿色 被引 5 · S2

提出 StepGuard,一种 step-level guard model,可对已完成的 agent trajectory 进行审计并在工具动作执行前进行检查;并引入 StepGen,一种自动数据引擎,能在风险步生成上下文相同但动作不同的安全与不安全 trajectory。This work proposes StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed, and introduces StepGen, an automatic data engine that generates safe and unsafe trajectories with the same context but different actions at the risky step.

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization
对齐中的语言链:跨语言排序偏好优化
arXiv:2608.23149 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Cross-lingual Ranking Preference Optimization (CRPO),一种新框架,利用来自英语的鲁棒偏好知识来促进目标语言的偏好对齐,从而增强语言适应性与输出质量。This paper proposes Cross-lingual Ranking Preference Optimization~ (CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language, thereby enhancing language adaptation and output quality.

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques
OntoAligner-Ensemble:基于投票的异构本体对齐技术融合
arXiv:2608.31137 安全与风险 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

结果揭示了集成组合对精度–召回权衡的直接影响:异构跨范式集成通常提升精度,而同构 LLM 集成更常取得更高的整体 F1。It is revealed that ensemble composition directly affects the precision-recall trade-off: heterogeneous cross-paradigm ensembles generally improve precision, whereas homogeneous LLM ensembles more often achieve higher overall F1-scores.

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
SafeAtlas-VL:超越二元判断的大规模多模态安全数据与防护模型
arXiv:2608.29098 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该论文提出 SafeAtlas-VL,一个包含 1.5M 训练实例的数据集,将图像、请求和响应级判断置于五级有序量表上,并通过 target-conditioned tuning 训练 SafeAtlas Guard 系列模型,用于多模态安全检测。This paper introduces SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale, and trains the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection.

Safin-1: Safety from Within through Memory-Native State Evolution
Safin-1:通过记忆原生状态演化实现内在安全
arXiv:2609.00092 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

路由状态接口在模型的原生计算中统一了上下文记忆与持久的能力适配,将记忆从对历史上下文的被动记录重塑为维持与演化模型行为的主动基质。The routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior.

Recursive Criticality of AI Self-Improvement
AI 自我改进的递归临界性
arXiv:2609.00137 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该框架识别了 AI R&D 系统的若干可测量属性,可用于区分递归放大与其他来源驱动的快速进展,包括递归反馈的强度、改进向后续系统传播的有效性、周期时长,以及进一步取得进展的难度递增。The framework identifies measurable properties of AI R\&D systems that can help distinguish recursive amplification from rapid progress driven by other sources, including the strength of recursive feedback, how effectively improvements propagate into successor systems, cycle duration, and the increasing difficulty of further progress.

Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations
无需跨资产收益协方差的投资组合风险边界:来自语言模型表征的分布场
arXiv:2608.29692 安全与风险 方法 被引 2 · S2

投资组合风险评估通常依赖可靠的跨资产收益协方差估计,而在短、高维面板中难以获得。我们表明公司层面的分布型特征可提供投资组合风险的单边证书。在"从特征到系统性暴露、从暴露到收益"的既定联系下,多公司 Wasserstein-2 离散度对系统性投资组合方差给出紧上界,并对标准化收益给出相应边界。加权成对松弛产生一个目标函数……Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are difficult to obtain in short, high-dimensional panels. We show that firm-level distribution-valued characteristics can instead provide one-sided certificates of portfolio risk. Under maintained links from characteristics to systematic exposures and from exposures to returns, multi-firm Wasserstein-2 dispersion yields a sharp upper bound on systematic portfolio variance and a corresponding bound for standardized returns. A weighted pairwise relaxation produces an objective that is conv

Enoki: Efficient Multi-Level Hallucination Detection
Enoki:高效多层级幻觉检测
arXiv:2609.00581 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Enoki,一个面向多级幻觉检测的开放信息抽取框架,支持基于 LLM、基于编码器与基于规则的三类抽取模式,通过统一接口平衡准确率与推理成本。This work proposes Enoki, an Open Information Extraction framework for multi-level hallucination detection that supports LLM-based, encoder-based, and rule-based extraction regimes, balancing accuracy and inference cost through a common interface.

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Refuse without Refusal:面向语言模型安全调优响应的结构性分析以减少误拒
arXiv:2609.04714 安全与风险 方法 OA · 绿色 被引 2 · S2

本文将安全微调数据集中的回复拆分为两个独立部分:模板化的拒答声明与解释拒答的理由,并表明拒答声明会诱导模型依赖表层线索,从而妨碍对有害与良性查询的准确区分。This paper decomposes a response in the safety-tuning dataset into two distinct components: a boilerplate refusal statement and a rationale explaining the refusal, and shows that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues.

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
为谁安全?用于可控 LLM 安全拒绝的边界感知自蒸馏
arXiv:2609.04482 安全与风险 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明数据组成控制安全性与可用性的权衡,且安全对齐应在预期拒答边界的两侧进行评估。The results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
A*-Thought-V2:通过 LLM 几何动力学实现高效潜在推理
arXiv:2609.07821 安全与风险 方法 OA · 绿色 被引 1 · S2

提出 A*-Thought-V2,一个由 LLM 引导的几何动力学框架,将 CoT 建模为隐状态轨迹,并以显式-隐式交错潜在架构替代硬删除,引入更广义的软目标以促进更丰富的步骤级特征学习。A*-Thought-V2 is presented, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture that reflects broader soft targets that encourage richer step-level feature learning.

A Zeroth-Order Paradigm for LLM Preference Alignment
一种面向 LLM 偏好对齐的零阶范式
arXiv:2609.19144 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出并分析 Comparison-based Preference Optimization (ComPO),一种基于比较预言机 (oracle) 的零阶对齐方法,并在光滑性、梯度稀疏性以及 oracle 与潜在目标相容的条件下,为其基础离线方案建立了收敛性保证。This paper proposes and analyzes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles, and establishes a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective.

Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models
价值几何:语言模型伦理偏好对齐中的任务向量组合
arXiv:2609.21094 安全与风险 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

一项基于任务向量迁移的实验:在计算了价值偏好方向对应的任务向量后,将其相对于通用指令跟随向量进行正交化处理,该方法能够有效隔离出特定价值偏好的方向,从而通过任务算术获得具有相反立场的模型。A task vector transfer based experiment where after computing the task vectors for a direction of value preference the authors orthogonalize it with respect to the general instruction following vector shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.

1. Data Flow Control(DFC):AI Agent 数据安全策略的内核级执行框架
arXiv:2606.05679 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文将数据安全形式化为 provenance monomials 上的聚合谓词,并提出 Passant——一个无需物化 provenance 即可强制执行 DFC 策略的可移植查询重写层。This paper formalizes data safety as aggregate predicates over provenance monomials and presents Passant, a portable query rewriting layer that enforces DFC policies without materializing provenance.

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
直接问 Jev:校准决策的强化学习作为 AI 对齐失败的零样本检测器
arXiv:2609.29429 安全与风险 评测集 OA · 绿色 被引 4 · S2

本文提出 RLCDAlignBench,在十类对齐失败上对 Jev 进行基准测试:谄媚、越狱、欺骗、提示注入、幻觉、隐私侵犯、社会偏见、奖励黑客、不确定性隐瞒与权力寻求。RLCDAlignBench is presented, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking.

When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
模型内部何时有用?探究表征工程在 LLM 安全中的作用
arXiv:2609.34771 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

总体而言,representation engineering 并不能普遍替代行为安全护栏,但在特定条件下具有实际优势,并可带来互补的安全收益。Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.

Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
迈向可信的 AI 开发:支持可验证声明的机制
arXiv:2004.07213 安全与风险 观点 OA · 绿色 被引 507 · S2

本报告建议不同利益相关方可采取多种措施,提升关于 AI 系统及其相关开发流程声明的可验证性,重点是为 AI 系统的安全性、安保、公平性与隐私保护提供证据。This report suggests various steps that different stakeholders can take to improve the verifiability of claims made about AI systems and their associated development processes, with a focus on providing evidence about the safety, security, fairness, and privacy protection of AI systems.

4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety
4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety(⭐⭐⭐⭐⭐)
arXiv:2602.16666 安全与风险 方法 Open MIND OA · 绿色 被引 78 · S2

本工作提出 12 个具体指标,从一致性、鲁棒性、可预测性和安全性四个关键维度分解 Agent 可靠性,可与传统评估互补,并提供用于分析 Agent 表现、退化与失败方式的工具。This work proposes twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety, which complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.

Cyber Security Awareness Campaigns: Why do they fail to change behaviour?
网络安全意识宣传活动:为何它们未能改变行为?
arXiv:1901.02672 安全与风险 综述 OA · 绿色 被引 465 · S2

综述了包括广泛使用的"恐惧诉求"在内的说服技巧的适用性,并提炼出意识宣传活动的基本要素以及导致活动成败的关键因素。The suitability of persuasion techniques, including the widely used 'fear appeals', are reviewed, and essential components for an awareness campaign as well as factors which can lead to a campaign's success or failure are extracted.

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
arXiv:2606.26015 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Tatoxa,一种面向鞑靼语文本去毒的 SOTA 系统;对比实验表明,该方法在关键质量指标上优于现有开源及商用闭源 LLM。Tatoxa is presented, a novel state-of-the-art system for text detoxification in the Tatar language, and comparative experiments show that the proposed approach outperforms existing open source and proprietary commercial LLMs on key quality metrics.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Qwen-RobotManip 技术报告:对齐释放机器人操作基础模型的规模化潜力
arXiv:2606.17846 安全与风险 方法 OA · 绿色 被引 49 · S2

Qwen-RobotManip 在所有 OOD 场景下大幅超越包括 π0.5 在内的已有 SOTA 模型,在 RoboChallenge 中排名第一,相对改进 20%,并在 AgileX ALOHA、Franka、UR、ARX 等真实机器人平台上完成验证。Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $\pi$0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.

IsabeLLM: Automated Theorem Proving Applied to Formally Verifying Consensus
IsabeLLM:将自动定理证明应用于共识协议的形式化验证
arXiv:2606.18098 安全与风险 方法 OA · 绿色 被引 1 · S2

实现了一个 RAG 框架,包含面向 LLM 的错误追踪与反例生成以提供更优上下文,并兼容最新版 Isabelle 与 Sledgehammer 以提升效率。A Retrieval-Augmented Generation framework, Error tracing and counterexample generation for improved context supplied to the Large Language Model, and Compatibility with the latest version of Isabelle and Sledgehammer is implemented for improved efficiency.

Pareto Optimal Re-ranking with Semi-Automated Content Credibility Detection
基于半自动化内容可信度检测的帕累托最优重排序
arXiv:2606.18031 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种基于优化的方法,通过精炼现有内容排序来提升社交媒体信息流中新闻内容的可信度;同时构建一条鲁棒的半自动化流水线,基于检索增强打分与人工事实核查的混合方式为内容赋予可信度分数。An optimization-based method to improve the credibility of news content on social media feeds by refining existing content rankings is presented and a robust semi-automated pipeline for assigning credibility scores to content based on a mixture of retrieval-augmented score assignments and human-generated fact-checks is proposed.

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
arXiv:2607.05910 安全与风险 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文提出 PolicyShiftGuard,一个紧凑的策略条件护栏,采用结合随机策略 SFT(RP-SFT)与边界对策略适配(BP-Adapt)的两阶段训练方案,并验证匹配的通过/拒绝边界对是稳定策略适配的关键。This work proposes PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt), and confirms that matched pass/block boundary pairs are essential for stable policy adaptation.

DeepLoop: Depth Scaling for Looped Transformers
DeepLoop:循环 Transformer 的深度扩展
arXiv:2607.13491 安全与风险 方法 OA · 绿色 被引 5 · S2

结果表明稳定的循环深度需要计入参数访问次数(而非仅名义层数)的残差缩放规则;DeepLoop 在不存在物理块被重复访问时表现为中性,一旦启用循环深度则改善验证损失与下游准确率。The results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count, and DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated.

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment
SUFLECA:面向 CAD-to-image 对齐的特征学习规模化方法
arXiv:2607.15058 安全与风险 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

SUFLECA(Scaling Up Feature LEarning for CAD-to-image Alignment),一种用于零样本 CAD 对齐的弱监督框架,含两项关键贡献,并提出一种几何一致的匹配算法以建立可靠的 CAD-图像对应。SUFLECA (Scaling Up Feature LEarning for CAD-to-image Alignment), a weakly supervised framework for zero-shot CAD alignment with two key contributions, and proposes a geometrically consistent matching algorithm that establishes reliable CAD-to-image correspondences.

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
AI 海洋中的海妖之歌:大语言模型幻觉问题综述
arXiv:2309.01219 安全与风险 综述 OA · 绿色 被引 1186 · S2

本文给出了 LLM 幻觉现象与评估基准的分类体系,分析了现有缓解 LLM 幻觉的方法,并讨论了未来研究的潜在方向。This paper presents taxonomies of the LLM hallucination phenomena and evaluation benchmarks, analyzes existing approaches aiming at mitigating LLm hallucination, and discusses potential directions for future research.

Consciousness in Artificial Intelligence: Insights from the Science of Consciousness
人工智能中的意识:来自意识科学的洞察
arXiv:2308.08708 安全与风险 观点 OA · 绿色 被引 274 · S2

该报告主张并例证了一种严谨且基于经验的方法来研究 AI 意识:依据获得最佳支持的神经科学意识理论,详细评估现有 AI 系统。This report argues for, and exemplifies, a rigorous and empirically grounded approach to AI consciousness: assessing existing AI systems in detail, in light of best-supported neuroscientific theories of consciousness.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
大语言模型中的幻觉综述:原理、分类、挑战与开放问题
arXiv:2311.05232 安全与风险 综述 OA · 绿色 被引 3939 · S2

全面概述了 LLM 幻觉检测方法与基准,并指出 LLM 幻觉领域有前景的研究方向,包括大视觉-语言模型中的幻觉以及 LLM 幻觉中的知识边界理解。A thorough overview of hallucination detection methods and benchmarks is presented and the promising research directions on LLM hallucinations are highlighted, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.

Deep Visual-Semantic Alignments for Generating Image Descriptions
用于生成图像描述的深度视觉-语义对齐
arXiv:1412.2306 安全与风险 方法 OA · 绿色 被引 6152 · S2

提出一个模型,基于图像区域上的 CNN、句子上的双向 RNN 以及通过多模态嵌入对齐两种模态的结构化目标,生成图像及其区域的自然语言描述。A model that generates natural language descriptions of images and their regions based on a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding is presented.

Constitutional Midtraining: Content Presence Drives Alignment Gains
宪法式中训练:内容存在驱动对齐收益
arXiv:2607.26654 安全与风险 方法 OA · 绿色 被引 2 · S2

训练后对齐往往较浅,会在微调中被侵蚀。而中训练干预能否在干净隔离于训练后的情况下产生持久对齐,此前未经检验。我们通过宪法式中训练来测试:在 120B 规模上,插入基于原则与价值观的内容,与仅做回放的对照组进行对比。我们基于 Anthropic 的 Constitution 构建了 394M token 的宪法语料,并采用 2×2 析因设计(课程顺序 × 审慎推理),形成四种宪法式中训练条件与一组对照,随后在自生成与既有...Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and establis

MemSFT: Mitigating Alignment Tax with an External Parametric Memory
标题 -> 标题中文:MemSFT:借助外部参数化记忆缓解对齐税
arXiv:2607.25614 安全与风险 方法 OA · 绿色 被引 2 · S2

将 LLM 适配到专用领域常会带来对齐税:针对领域特定任务进行微调会导致灾难性遗忘,并显著降低在通用任务上的表现。我们提出 MemSFT,通过将领域专业化与主干参数更新解耦,以即插即用的参数化记忆来缓解对齐税。该记忆被训练为模仿在领域数据上运作的非参数化检索器,从而记住原本需通过检索获取的知识与模式。一旦在Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We propose MemSFT, which mitigates the alignment tax by decoupling domain specialization from backbone parameter updates through a plug-and-play parametric memory. The memory is trained to imitate the behavior of a non-parametric retriever operating over domain data, thereby memorizing knowledge and patterns that would otherwise be accessed through retrieval. Once trained on

Invariant Risk Minimization
不变风险最小化
arXiv:1907.02893 安全与风险 方法 OA · 绿色 被引 3013 · S2

本文提出了不变风险最小化(IRM),一种用于在多个训练分布上估计不变相关性的学习范式,并展示了 IRM 学到的不变性如何与数据的因果结构相关,从而实现分布外泛化。This work introduces Invariant Risk Minimization, a learning paradigm to estimate invariant correlations across multiple training distributions and shows how the invariances learned by IRM relate to the causal structures governing the data and enable out-of-distribution generalization.