Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

  • 类型:arxiv
  • 标识:2608.20438
  • 链接:https://arxiv.org/abs/2608.20438
  • 主分类:agent
  • 形态:benchmark
  • TLDR:Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model-by-topic-by-seed blocks. Relative to a topic-only control, a feed of previous-round peer posts ranked by peer-generated likes increases final-round lexical similarity in both the
  • 待LLM分类:否
  • 标题中文:基于同行投票的 LLM-Agent 压力测试:发现信息流引发的词汇趋同,但分布式来源未呈现可靠的等曝光优势
  • TLDR中文:大语言模型(LLM)agent 的群体行为无法用单 agent 基准刻画。我们提出 PV-SST,一个基于同行投票的社交平台测试平台,并报告一项独立冻结、预先注册的等曝光实验,涵盖四个话题、四个未使用种子、四个开源权重模型家族以及三个预设的更大模型变体。该实验包含 448 次试验和 112 个完整的"模型 × 话题 × 种子"区组。相对于仅话题对照条件,由同行生成的点赞排序的上一轮同行帖子信息流,会提升最终轮在两端的词汇相似度
  • 来源文件
  • /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json
  • /inbox/tom/_candidates/2026-08-25-agent-memory-tool-use-candidates.json