研究库 论文知识库
Papers · organized/paper_cards

论文

38 张论文卡片 · 安全与风险

开放获取 全部 绿色 · 1640
Stealing Reasoning Traces from Proprietary LLM APIs
从商用 LLM API 窃取推理轨迹
arXiv:2608.09867 安全与风险 方法 OA · 绿色 被引 11 · S2

本文识别出一种绕过 anti-distillation 机制、允许攻击者窃取专有模型推理能力的架构漏洞,并提出具体的密码学与系统级缓解措施以保障客户端推理安全。An architectural vulnerability is identified that circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as well as proposing concrete cryptographic and system-level mitigations to secure client-side reasoning.

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
从不可听输入到模型失败:LALM 中的低频安全风险
arXiv:2608.09158 安全与风险 方法 被引 0 · S2

提出 Intermittent Low-Frequency Lockout(ILL),一种基于通用波形模板在黑盒设置下评估该风险的不可听 red teaming 方法;同时提出 Distributional Requery Guard(DRG),用于缓解该风险。This paper proposes Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting and proposes Distributional Requery Guard (DRG), an inaudible red teaming method to mitigate this risk.