RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

  • 类型:arxiv
  • 标识:2607.26991
  • 链接:https://arxiv.org/abs/2607.26991
  • 主分类:multimodal
  • 形态:position
  • 被引:0
  • 被引来源:Semantic Scholar + OpenAlex
  • S2被引:0
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:This work introduces an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents, and discovers that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely.
  • OpenAlex ID:W7171842867
  • OpenAlex DOI:10.48550/arxiv.2607.26991
  • DOI:10.48550/arxiv.2607.26991
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2607.26991
  • OpenAlex更新:2026-08-26
  • 待LLM分类:否
  • 标题中文:RL^2-VLA:面向 Vision-Language-Action 模型的自适应强化学习潜在组合引导与测试时缩放
  • TLDR中文:提出一种自适应推理时引导框架,利用VLA Latents上的强化学习,发现推理时引导在成功与失败状态下遵循根本不同的scaling laws:动作多样性在基础VLA可能失败时最为有益,但在成功可能性高时可能不必要地扰动已准确的动作。
  • 副分类:llm-infra
  • 来源文件
  • /inbox/tom/_candidates/2026-08-03-agent-rag-longcontext-candidates.json
  • /inbox/tom/_candidates/2026-08-04-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-08-04-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]