MAGIC 是一个四阶段 pipeline,能将单一自然语言提示转化为可运行的多场景游戏项目,相比 LLM 基线与 Holodeck 可恢复更多真实 portal,并生成显著更可导航的布局。MAGIC is a four-stage pipeline that turns a single natural-language prompt into a runnable multi-scene game project that recovers more ground-truth portals and yields markedly more navigable layouts than an LLM baseline and Holodeck.
论文
471 张论文卡片 · 方法 · OA 绿色
本研究将朴素搜索的根因追溯到生成器特有的、可演化的知识边界——即生成器经训练可内化的内容与必须保留于外部上下文的内容之间的鸿沟,并表明该边界可通过"先教后搜"协同训练框架被有效发现。This work traces the root cause of naive search to a generator-specific, evolving knowledge boundary: the divide between what a generator can internalize through training and what must remain in external context, and shows that it is discoverable through a teach-then-search co-training framework.
本研究考察了 GitHub 项目在各自引入首个 bot 前后两年间的情况,发现变化集中在采纳时点附近,而非逐渐累积,这与一种特定解读一致:可预测、基于规则的 Agent 能够成为社区社交基础设施的一部分。This work examines GitHub projects for two years before and after each adopted its first bot, finding changes cluster around adoption rather than accumulating gradually, consistent with a specific interpretation: predictable, rule-based agents can become part of a community's social infrastructure.
Hallo4D 提出"生成-检测-修正"范式,利用大型多模态语言模型(LMMs)从多视角与多帧渲染中识别并归纳时空不一致性,为一致性感知的内容生成提供了一种可扩展且可泛化的方案。Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings, providing a scalable and generalizable solution for consistency-aware content generation.
研究表明,通过更强的多模态编码器、Agentic prompt 改写及相关技术来增强 Boogu-Image 系统的理解能力,并结合数据质量、训练流程和 Agentic 推理时扩展的改进,即使在计算预算极为受限的条件下,也能显著提升生成与编辑性能。It is demonstrated that strengthening the understanding capability of the Boogu-Image system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets.
论文提出 Vinci2,一个主动式的第一人称视频协助系统,将端侧助手 Vinci 由被动响应推进到主动协助;以及免训练、记忆增强的 Agent EgoMemo,维护三种互补的记忆表征:多尺度时间摘要、语义知识图谱与视觉嵌入档案。Vinci2 is presented, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity and EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives.
论文提出 OAT,将该问题建模为基于神经受控微分方程的单类学习,在潜空间中刻画成功轨迹的动力学模式;实验表明其比基于 prompt 的基线更快,并在领域内和分布外数据集上均稳定优于基线。OAT is proposed, which casts this problem as one-class learning with neural controlled differential equations, modeling the dynamical pattern of successful trajectories in latent space, and is shown to be faster than prompting-based baselines and consistently outperforms them in both in-domain and out-of-distribution datasets.
STRACE(Structural TRajectory Analysis and Causal Extraction)是一个用于构建高信噪比优化上下文的框架,旨在对长周期 Agent 实施更精确、更有效的优化。STRACE (Structural TRajectory Analysis and Causal Extraction) is a framework that constructs high signal-noise optimization contexts for more precise and effective optimization of long-horizon agents.
本研究表明 DiT 与 ViT 在一个关键方面存在差异:DiT 不会出现 patch-token 异常值,但仍能受益于 registers;并且 registers 在像素空间 DiT 中比在潜空间 DiT 中效果更显著。This work shows that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers, and finds that registers are more effective in pixel-space DiTs than in latent-space DiTs.
本文提出 Earthquaker-AI,一个混合式教育框架,在已有教育机器人项目基础上集成基于 RAG 的对话式 AI 助手,旨在提升小学生的地震应急准备与主动行动意识。该系统将曾获奖的 STEM 项目 Earthquaker 从 Lego WeDo2 的机械模拟拓展至认知与元认知层面:机器人组件利用 Lego WeDo2 自动化模拟地震响应,使学生能够与传感器和执行器进行交互。This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a conversational AI assistant based on Retrieval-Augmented Generation. It aims to enhance earthquake preparedness and conscious action among primary-school students. The system extends the award-winning STEM project Earthquaker moving from mechanical simulation with Lego WeDo2 to cognitive and metacognitive processing. The robotics component uses Lego WeDo2 automation to simulate seismic response, letting students interact with sensors and actuat
面向第 11 届 ABAW 挑战赛的多任务学习系统,在标准确定性架构基础上扩展条件 Rectified Flow 头,建模真实场景下面部行为固有的模糊性,借助蒙特卡洛采样实现不确定性感知的一对多预测。A multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling.
压缩可减少推理 token 并保持大部分选择题准确率,同时提示影响接近基线——在一个前沿条件下,降低推理成本移除的证据比单纯缩短轨迹所预期的更多。Compression reduces reasoning tokens and preserves most multiple choice accuracy, while hint influence remains near baseline, in a frontier where reducing reasoning costs removes more evidence than shorter traces alone would predict.
研究表明模型具备相关子群体知识,但难以稳定传递到聚合估计中,这一差距使统计自一致性成为评估 LLM 的尚未饱和、无需参考的准则。It is suggested that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates, and this gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.
本文对 OPD 的作用、病态与调控进行系统研究,厘清 OPD 作为探索催化剂的角色,并证实调控良好的信号质量(而非单纯的教师模型规模)才是 OPD 中成功探索的主导因素。A systematic study examining the role, pathologies, and regulations of OPD, clarifying the role of OPD as an exploration catalyst and confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.
结果表明稳定的循环深度需要计入参数访问次数(而非仅名义层数)的残差缩放规则;DeepLoop 在不存在物理块被重复访问时表现为中性,一旦启用循环深度则改善验证损失与下游准确率。The results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count, and DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated.
本文提出 UniVR,这是首个从纯视觉演示中同时学习复杂推理、细粒度物理动力学和长期规划的研究,并配套首个在纯视觉协议下评估这些异质能力的综合评测套件。UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations, is introduced, the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol.
提出一个多 Agent 框架,通过结合监督微调、直接偏好优化和检索增强生成 (RAG) 来调和事实基础与意识形态对齐,产生稳定的胜者和排名,且以宣言为锚的谱系能可靠预测现实世界中的实现,而幻觉内容则不能。A multi-agent framework that reconciles factual grounding with ideological alignment by combining Supervised Fine-Tuning, Direct Preference Optimization, and Retrieval-Augmented Generation is presented, which yields a stable winner and ranking, and manifesto-anchored lineage reliably predicts real-world materialization whereas hallucinated content does not.
提出 GRASP,一个用于训练智能体在多步推理过程中自适应协调互补检索工具的强化学习(RL)框架,并指出学会协调检索信号与上下文粒度对智能体的正确推理至关重要。GRASP is introduced, a reinforcement learning (RL) framework for training agents to adaptively coordinate complementary retrieval tools during multi-step reasoning, and it is suggested that learning to coordinate retrieval signals and context granularity is critical for agent's correct reasoning.
本文提出 HDR (Hierarchical Denoising for Visual Reasoning),一个将层级潜变量集成到因果视频生成中以进行多步推理的统一框架,并引入一个含分布外情况的层级化多步视频推理基准。This work proposes HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning and introduces a level-stratified multi-step video reasoning benchmark with out-of-distribution cases.
提出 Hy-Embodied-RxBrain,一个具备语言-视觉联合推理与想象的具身认知基础模型,并将其扩展到连续机器人动作生成,在无需大规模动作数据预训练的情况下展现出可观的真实机器人性能。Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination, is introduced and extended to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining.
提出 Chat2Scenic,是首个以领域特定语言 (DSL) 生成场景脚本的迭代式检索增强框架,并构建了一个涵盖 NHTSA、联合国车辆法规及其他来源共 123 个场景的开源场景生成基准。Chat2Scenic is presented, the first iterative retrieval-augmented framework to generate scenario scripts in Domain Specific Language (DSL) and proposes an open benchmark for scenario generation comprising 123 scenarios from various regulations, including NHTSA and United Nations Vehicle Regulations, as well as other sources.
发现对 GPT 语言模型进行重复采样,是为困难 prompt 生成可用解的一种出乎意料的有效策略,并讨论了部署强大代码生成技术在安全性、安全保障与经济等方面的潜在更广泛影响。It is found that repeated sampling from the GPT language model is a surprisingly effective strategy for producing working solutions to difficult prompts, and the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics are discussed.
提出 Flamingo,一个 VLM 系列,能够桥接强大的纯视觉与纯语言预训练模型,处理任意交错排列的图文序列,并无缝接收图像或视频作为输入。This work introduces Flamingo, a family of Visual Language Models (VLM) with this ability to bridge powerful pretrained vision-only and language-only models, handle sequences of arbitrarily interleaved visual and textual data, and seamlessly ingest images or videos as inputs.
提出 BloombergGPT,一个 500 亿参数的语言模型,在广泛的金融数据上训练而成,并基于 Bloomberg 丰富的数据源构建了包含 3630 亿 token 的数据集,可能是迄今最大的领域专用数据集。This work presents BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data, and constructs a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet.
Gemini 1.5 在跨模态长上下文检索任务上取得近乎完美的召回率,在长文档 QA、长视频 QA 与长上下文 ASR 上刷新 SOTA,并在广泛基准上达到或超越 Gemini 1.0 Ultra 的 SOTA 表现。Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks.
Reflexion 是一个通过语言反馈而非更新权重来强化语言 agent 的新框架,在多种任务(序贯决策、编程、语言推理)上相较基线 agent 取得显著提升。Reflexion is a novel framework to reinforce language agents not by updating weights, but instead through linguistic feedback, which obtains significant improvements over a baseline agent across diverse tasks (sequential decision-making, coding, language reasoning).
HuggingGPT 是一个由 LLM 驱动的 Agent,利用 LLM(如 ChatGPT)连接机器学习社区中的各种 AI 模型以解决 AI 任务,能够处理跨模态、跨领域的大量复杂 AI 任务。HuggingGPT is an LLM-powered agent that leverages LLMs (e.g., ChatGPT) to connect various AI models in machine learning communities to solve AI tasks and can tackle a wide range of sophisticated AI tasks spanning different modalities and domains.
实验表明,与语言模型类似,视觉模型也会学习利用全局捷径,从而无法在任务长度或复杂度上泛化,但本文证明基于严格局部感知的循环视觉策略可以缓解这些失败,从而使模型在这些任务上具备泛化能力。The experiments reveal that, similar to language models, vision models can learn to exploit global shortcuts and thereby fail to generalize over task length or complexity, but it is shown that recurrent vision policies based on strictly local perception can mitigate these failures, thereby allowing models to generalize on these tasks.
综合评估表明,DeepSeek-V3 优于其他开源模型,并达到与领先闭源模型相当的性能。Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models.
本文在交通微观仿真器 SUMO 中应用现代深度强化学习方法构建一个真正自适应的交通信号控制智能体,并采用一种新的状态空间——离散交通状态编码——其信息密度较高。This work applies modern deep reinforcement learning methods to build a truly adaptive traffic signal control agent in the traffic microsimulator SUMO, using a new state space, the discrete traffic state encoding, which is information dense.
本文回顾现有方法,并融合多种技术从数据与模型规模两方面扩展预训练,提出一条自动化流水线以构建专用、多样且经过筛选的图像数据集,替代自监督文献中常用的未筛选数据。This work revisits existing approaches and combines different techniques to scale the pretraining in terms of data and model size, and proposes an automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data, as typically done in the self-supervised literature.
BLIP-2 在多种视觉-语言任务上取得 SOTA 性能,可训练参数远少于现有方法,并展现出遵循自然语言指令进行零样本图生文的新兴能力。BLIP-2 achieves state-of-the-art performance on various vision-language tasks, despite having significantly fewer trainable parameters than existing methods, and is demonstrated's emerging capabilities of zero-shot image-to-text generation that can follow natural language instructions.
一种面向语言模型推理的新框架 Tree of Thoughts (ToT),推广了流行的 Chain of Thought 提示方法,允许在作为问题求解中间步骤的连贯文本单元(thoughts)上进行探索。A new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving.
本文提出 Toolformer,训练其决定调用哪些 API、何时调用、传入什么参数,以及如何将结果最佳地融入后续 token 预测,在多种下游任务上显著提升零样本性能。This paper introduces Toolformer, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction, which achieves substantially improved zero-shot performance across a variety of downstream tasks.
在七个基准上,TimeLens2-2B 在所有基准上均优于规模相当的所有基线,4B 和 8B 变体则取得了 SOTA 性能,超越了参数量高达 397B 的开源模型。Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters.