研究库 论文知识库
Papers · organized/paper_cards

论文

5 张论文卡片 · 安全与风险 · 观点 · OA 绿色

开放获取 全部 绿色 · 1640
SSGM框架(Stability and Safety-Governed Memory)
3. SSGM框架(Stability and Safety-Governed Memory)
arXiv:2603.11768 安全与风险 观点 OA · 绿色 被引 21 · S2

通过形式化分析与架构分解,展示 SSGM 如何缓解拓扑引发的知识泄漏(敏感上下文被固化到长期存储),以及有助于防止语义漂移(知识在迭代摘要中退化)。Through formal analysis and architectural decomposition, it is shown how SSGM can mitigate topology-induced knowledge leakage where sensitive contexts are solidified into long-term storage, and help prevent semantic drift where knowledge degrades through iterative summarization.

Recursive Criticality of AI Self-Improvement
AI 自我改进的递归临界性
arXiv:2609.00137 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该框架识别了 AI R&D 系统的若干可测量属性,可用于区分递归放大与其他来源驱动的快速进展,包括递归反馈的强度、改进向后续系统传播的有效性、周期时长,以及进一步取得进展的难度递增。The framework identifies measurable properties of AI R\&D systems that can help distinguish recursive amplification from rapid progress driven by other sources, including the strength of recursive feedback, how effectively improvements propagate into successor systems, cycle duration, and the increasing difficulty of further progress.

Enoki: Efficient Multi-Level Hallucination Detection
Enoki:高效多层级幻觉检测
arXiv:2609.00581 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Enoki,一个面向多级幻觉检测的开放信息抽取框架,支持基于 LLM、基于编码器与基于规则的三类抽取模式,通过统一接口平衡准确率与推理成本。This work proposes Enoki, an Open Information Extraction framework for multi-level hallucination detection that supports LLM-based, encoder-based, and rule-based extraction regimes, balancing accuracy and inference cost through a common interface.

Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
迈向可信的 AI 开发:支持可验证声明的机制
arXiv:2004.07213 安全与风险 观点 OA · 绿色 被引 507 · S2

本报告建议不同利益相关方可采取多种措施,提升关于 AI 系统及其相关开发流程声明的可验证性,重点是为 AI 系统的安全性、安保、公平性与隐私保护提供证据。This report suggests various steps that different stakeholders can take to improve the verifiability of claims made about AI systems and their associated development processes, with a focus on providing evidence about the safety, security, fairness, and privacy protection of AI systems.

Consciousness in Artificial Intelligence: Insights from the Science of Consciousness
人工智能中的意识:来自意识科学的洞察
arXiv:2308.08708 安全与风险 观点 OA · 绿色 被引 274 · S2

该报告主张并例证了一种严谨且基于经验的方法来研究 AI 意识:依据获得最佳支持的神经科学意识理论,详细评估现有 AI 系统。This report argues for, and exemplifies, a rigorous and empirically grounded approach to AI consciousness: assessing existing AI systems in detail, in light of best-supported neuroscientific theories of consciousness.