Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

  • 类型:arxiv
  • 标识:2607.21072
  • 链接:https://arxiv.org/abs/2607.21072
  • 主分类:evaluation
  • 形态:benchmark
  • 被引:1
  • 被引来源:Semantic Scholar
  • S2被引:1
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and text-based reasoning.
  • OpenAlex ID:W7170501523
  • OpenAlex DOI:10.48550/arxiv.2607.21072
  • DOI:10.48550/arxiv.2607.21072
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2607.21072
  • OpenAlex更新:2026-08-24
  • 待LLM分类:否
  • 标题中文:展示而非讲述:在生成像素而非 LLM 文本中评估空间认知
  • TLDR中文:提出 ProVisE(Protocolized Visual Evaluation),一个与基准无关的框架,通过受协议约束的视觉问答从图像生成模型中抽取答案,并将其解析为与原始指标兼容的结构化预测,揭示了像素空间表达与基于文本推理的互补优势。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-24-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]