Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model

  • 类型:arxiv
  • 标识:2609.07154
  • 链接:https://arxiv.org/abs/2609.07154
  • 主分类:multimodal
  • 形态:method
  • TLDR:We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in the <=2B parameter division with 0.8279 on the held-out test set. Our system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass; It is obtained by distilling the junior perception module of a tool-using agentic pipeline, not the agent itself into a small student, using teacher traces filtered to those that answered correctly. it reaches 89% of the accuracy of the large agentic pipeline using 1.1% of
  • 副分类:agent
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-15-agent-rag-longcontext-candidates.json