SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

  • 类型:arxiv
  • 标识:2608.29098
  • 链接:https://arxiv.org/abs/2608.29098
  • 主分类:risk
  • 形态:method
  • TLDR:Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world
  • 副分类:multimodal
  • 待LLM分类:否
  • 标题中文:SafeAtlas-VL:超越二元判断的大规模多模态安全数据与防护模型
  • TLDR中文:多模态安全审核需要区分源自视觉内容、用户意图与助手行为的风险。然而现有防护模型通常只针对单一判断目标训练,将安全评估简化为二元决策,导致风险难以在多模态交互中跨类型比较,模糊案例也被掩盖。我们提出 SafeAtlas-VL,一个包含 150 万训练实例的数据集,将图像、请求、回答三个层级判断统一在五级有序量表上,并从真实世界
  • 来源文件
  • /inbox/tom/_candidates/2026-09-01-agent-rag-longcontext-candidates.json