研究库 论文知识库
Papers · organized/paper_cards

论文

1173 张论文卡片 · 方法

开放获取 全部 绿色 · 1640
KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora
[标题中文] KUPAS MASTER:将资深从业者的隐性专业知识蒸馏为 Agent 可用的经验语料
arXiv:2609.37673 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

KUPAS MASTER 是一个围绕九层认知语料构建的经验工程平台,将异构的工作记录与从业者访谈转化为 agent 可用的可追溯、可复用经验语料,提供了从个体隐性经验到组织知识与 agent 能力的可行路径。KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction, turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents, and provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.

Improved Distributional Diffusion Models
[标题中文] 改进的分布式扩散模型
arXiv:2609.37147 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种随机少步生成器,其 FID 在采样预算从 4 增加到 50 NFE 时不会下降,且相同配方可迁移到 text-to-image 生成。A stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation.

LLMs are General Asynchronous Agents
[标题中文] LLM 是通用异步 Agent
arXiv:2609.35427 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种异步 LLM 框架,允许用户(或 Agent 自身)定义具有重叠 memory 状态的推理协程,并展示 Qwen 3.x 模型无需任务专属训练即可在流式视频理解、电子游戏和监控中实现异步运行。This work develops an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states and showcases that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.

AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
AutoRef:面向 Agentic 多参考图生成的 Harness 优化
arXiv:2609.35530 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 AutoRef,在保持两侧模型冻结的前提下自动优化 harness:一个 coding Agent 迭代重写 harness 代码,并发现 AutoRef-Harness,可改进开源权重模型 FLUX。This work proposes AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code, and discovers AutoRef-Harness, which improves the open-weight FLUX.

Reasoning with Image Generation
结合图像生成的推理
arXiv:2609.16409 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在包括多视角空间推理和碰撞预测在内的六项多样化视觉推理任务中,ReImaGin 始终优于纯文本推理和专用 vision-tool 基线,提升幅度高达 25%,证明了灵活、可生成的视觉推理的优势。Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.

What Makes Recurrence Effective in Looped Language Models?
循环语言模型中的循环机制为何有效?
arXiv:2609.36636 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通道级 history-state 注入结合 timestep 条件化构成了一种低成本且更有效的设计,能够在更长 unrolling 下更好地保留知识,同时提升在不同推理预算下的鲁棒性。It is shown that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets.

Fractional State Space Transition for Long Sequence Modeling
用于长序列建模的分数阶状态空间转移
arXiv:2609.36314 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FRAC,一种源自分数阶动力学的选择性 SSM 架构,用幂律长记忆替代指数衰减,在长上下文性能上持续优于 SOTA SSM 基线,同时在短上下文上保持竞争力。FRAC is introduced, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory and consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context.

Org-Agent: Beyond Personal Assistants Towards Organizational Agents
Org-Agent:从个人助手走向组织级智能体
arXiv:2609.34392 Agent 智能体 方法 被引 0 · S2

本文提出 Org-Agent,一种以约束为中心的统一推理框架,将任务执行组织为三个阶段,并将任务分解为原子子任务,构建编码其依赖关系的任务依赖图。Org-Agent is introduced, a unified constraint-centric reasoning framework that organizes task execution in three stages and decomposes a task into atomic subtasks and constructs a task dependency graph whose edges encode the dependencies among them.

Scaling Properties of Same-Family On-Policy Distillation
同家族同策略蒸馏的缩放特性
arXiv:2609.32722 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

研究发现,早期 OPD 训练动态均呈现一种规律的 *useful-transfer* 区间,其中留出准确率(即 *gold score*, $G$)随 $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$(学生初始化在 token 级反向 KL 散度的平方根)近似线性上升。It is found that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization.

4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety
4.2 Reliability 不等于成功率:12 指标拆出 consistency / robustness / predictability / safety(⭐⭐⭐⭐⭐)
arXiv:2602.16666 安全与风险 方法 Open MIND OA · 绿色 被引 78 · S2

本工作提出 12 个具体指标,从一致性、鲁棒性、可预测性和安全性四个关键维度分解 Agent 可靠性,可与传统评估互补,并提供用于分析 Agent 表现、退化与失败方式的工具。This work proposes twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety, which complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
UniEvo-VL:一种用于多模态模型自我改进的 On-policy 自蒸馏训练方案
arXiv:2609.38721 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 UniEvo-VL,一种面向多模态模型的自演化框架,可在 test-time compute 阶段从这种建设性的自纠错反馈中学习,在无外部监督或指导的情况下提升用户使用多模态模型的体验。UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute, is introduced to enhance the user experience when using multimodal models without external supervision or guidance.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Mid-Harness:在 Model 与 Harness 之间扩展终端 Agent 的动作规模
arXiv:2609.39982 评测基准 方法 OA · 绿色 被引 1 · S2

本文提出 Mid-Harness,在执行前对候选动作进行采样和验证,同时保持 generator 和 harness 不变,并将 action scaling 识别为终端 Agent 中 test-time compute scaling 的一个有前景的目标。Mid-Harness is introduced, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged, and identifies action scaling as a promising target for test-time compute scaling in terminal agents.

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
False Frontiers:诊断与缓解自演化搜索 Agent 中的共作弊问题
arXiv:2609.39102 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出多样本验证(MSV),对同一模型在有源信息和无源信息下各查询三次,以决定任务接纳并替换不可靠的伪标签,这部分减少了虚假一致性,但仍残留大量 co-cheating。This work introduces multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels, which partially reduces false agreement but leaves substantial residual co-cheating.

Scaling Laws for Looped Mixture of Experts
循环化 Mixture of Experts 的扩展定律
arXiv:2609.40316 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Loop Scaling Laws,是首个将 recurrence 和 sparsity 与模型规模、数据联合建模的 scaling law,为在算力和显存约束下设计 looped MoE 模型提供了原则性基础。Loop Scaling Laws are introduced, the first scaling law to jointly model recurrence and sparsity alongside model size and data, and provide a principled foundation for designing looped MoE models under compute and memory constraints.

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
打破巴别塔:面向长字幕翻译的自进化多 Agent 系统
arXiv:2609.38660 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

SMART,面向长篇字幕翻译的自演化多 Agent 系统,构建持久化的剧集级 memory,并通过动态 router 和 Mixture-of-Agents 层翻译部分句子,配合术语验证、字幕约束校验和上下文检索工具。SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation, builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval.

RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
RoPE 走到尽头了吗?长上下文失效的理论、诊断与缓解
arXiv:2609.39929 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

基于 RoPE 的语言模型出现长上下文失效,根源在于 RoPE 在维持稳定 token 偏好与区分相近位置之间存在固有权衡。要判断应处理哪种弱点以及如何处理,需要更精确地刻画 RoPE 在不同上下文长度下训练后模型中的行为。我们通过允许 RoPE 各频率下 query-key 尺度不一致,弥补了先前理论的一个关键局限,使之与实际经验观测高度吻合。我们的理论使得上述脆弱性对单个注意力头与输入可测量,并量化了高频Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-freque

4.1 LogicalRAG:把 Agentic RAG 的重点从“更重 backend”转向“更强 retrieval control”
arXiv:2605.27123 RAG 检索增强 方法 OA · 绿色 被引 3 · S2

本文提出一个 Agentic RAG 框架,使 LLM 能够使用逻辑表达式构建检索意图,同时将检索后端简化为基于倒排索引的系统,并表明将检索过程锚定在逻辑查询上可显著降低生成响应中的幻觉。This paper proposes an agentic RAG framework that enables LLMs to formulate retrieval intents using logical expressions while simplifying the retrieval backend to an inverted-index-based system, and shows that anchoring the retrieval process in logical queries substantially reduces hallucinations in generated responses.

Safety of Latent Communication in Multi-Agent Systems
多 Agent 系统中潜在通信的安全
arXiv:2609.39788 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

即使是良性的链接训练也会在底层安全对齐 agent 不变的情况下,相对于基于文本的通信增加有害合规性;安全对齐需要将多智能体系统作为整体来考虑。This work shows that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged, and shows that safety alignment requires considering the multi-agent system as a whole.

ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
ATLAS:面向可靠世界模型规划的对齐潜在结构传输
arXiv:2609.36333 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ATLAS(Aligned Transport of Latent Structure),一种在显式保持关系几何结构的同时校准全局潜空间分布的训练目标,通过一维 Wasserstein-2 传输进行 Wasserstein 嵌入匹配来校准其边缘分布。This work introduces Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution and uses Wasserstein embedding matching to calibrate its marginal through one-dimensional Wasserstein-2 transport.

It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
关键不在图像本身:无关上下文会扰动 VLM 评判且不提供有效信息
arXiv:2609.37863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MIST(Misleading-Image Stress Test):200 个英文句子,每句围绕一个可作比喻或字面理解的短语,配以对齐图像(描绘其读法)、误导图像(描绘相反读法)或无图像三种条件。MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.

Generative Moment Matching Networks
生成矩匹配网络
arXiv:1502.02761 多模态 方法 OA · 绿色 被引 954 · S2

本文提出一种方法,通过多层感知机的一次前馈传播生成独立样本(与近期提出的 GAN 类似),并使用 MMD 学习生成可被解码为样本的 codes。This work forms a method that generates an independent sample via a single feedforward pass through a multilayer perceptron, as in the recently proposed generative adversarial networks, using MMD to learn to generate codes that can then be decoded to produce samples.

DAGent: Evaluate-then-Grow Planning for Deep Research Agents
DAGent:面向深度研究 Agent 的 Evaluate-then-Grow 规划
arXiv:2609.39154 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

DAGent 是一个基于 DAG 的多智能体框架,采用先评估再生长的增量规划:Orchestrator 逐批扩展任务图,每一步扩展都以已完成节点的置信度与不确定性信号为条件;证据条件化规划在更低的每任务 token、工具调用和步骤开销下达到了比 Plan-then-Patch 更高的准确率。DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes, which shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart.

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
MILO:通过编排式多 Agent 进化实现自动化 Harness 发现
arXiv:2609.38349 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MILO(Meta-evolutionary Island Orchestration),一个共同演化 agent harness 及其发现策略的框架,使用前沿模型(Opus 4.8)与开源权重模型(gpt-oss-120b)超越了八个 SOTA harness 与六种搜索方法。This work introduces MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them and outperforms eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models.

2.3 本轮补充公开检索
arXiv:2606.14589 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本文对一个自 2026 年 3 月起持续运行的个人助理 Agent 运行时中的静默失败进行纵向研究,该系统包含约 40 个定时任务、8 个 LLM 提供商、一个工具治理代理以及一个知识库记忆层,由 4,286 个单元测试和 827 项治理检查守护。A longitudinal study of silent failures in a personal-assistant agent runtime in continuous production since March 2026, with roughly 40 scheduled jobs, 8 LLM providers, a tool-governance proxy, and a knowledge-base memory plane, defended by 4,286 unit tests and 827 governance checks is presented.

Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Tacit-TTS:从自回归解码到掩码预测的高效免转录语音克隆
arXiv:2609.38658 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Tacit-TTS 是一个从 IndexTTS2 蒸馏而来的高效无需转录的零样本语音克隆系统,用掩码非自回归生成替换自回归的文本到语义解码,引入无需训练的声学长度估计,并通过 ReFlow 蒸馏加速流匹配渲染器。Tacit-TTS is presented, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2 that replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation.

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
PixelUMM:无编码器的统一图像与视频理解与生成
arXiv:2609.38597 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

统一多模态模型(UMM)通常依赖独立的视觉表示分别完成理解与生成,这增加了视觉上下文长度,并使其难以与既有视觉-语言预训练流程集成。近期 pixel-space modeling 的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,统一视觉接口的设计仍是开放问题。本文提出 PixelUMM,一种用于统一图像与视频理解的无编码器模型……Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding

Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
用奖励加权传输蒸馏对齐单步生成模型
arXiv:2609.30840 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

理论分析表明 RWTD 的不动点分布在参考策略的 off-policy 奖励倾斜与当前模型的 on-policy 倾斜之间插值,提供了一种在奖励适配与保留先验知识之间取得平衡的原则性方法。Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.

Walking the Embedding Space: Datastore Extraction from Multimodal RAG
漫步嵌入空间:来自多模态 RAG 的数据存储提取
arXiv:2610.01871 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种自适应、自动化的数据提取攻击流程,在黑盒设置下针对 MRAG(其中检索到的视觉产物本身就是答案)发起攻击,表明亟需专门面向多模态数据设计的安全防护。An adaptive and automatic data extraction attack procedure operating in a black box setting against MRAG, a configuration in which the retrieved visual artifact is itself the response, and shows the urgent need for safeguards specifically designed for multimodal data.

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
InterEvolve:用于人形机器人 loco-manipulation 的奖励程序测试时演化
arXiv:2610.02196 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

InterEvolve 提出了一个物体感知的前向-后向行为基础模型,其在冻结身体先验上的物体残差可在测试时将关于身体或物体的奖励转化为 loco-manipulation 行为,并将任务以奖励程序的形式指定:带完成条件与可调常数的分阶段奖励。InterEvolve develops an object-aware forward-backward behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time, and specifies tasks as reward programs: staged rewards with completion conditions and tunable constants.

Memorizon: Training World Models Beyond Their Context Window
Memorizon:在上下文窗口之外训练世界模型
arXiv:2610.00544 工程化 方法 被引 0 · S2

基于 Rec 的检索在所有划分上都提升了回访一致性;当上下文跨度足以覆盖每次返回的首次访问时,可再带来 24% 到 30% 的提升,但会牺牲一定图像质量;超过该跨度后,继续增加长度不再带来收益。Rec retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps.

Smaller Models, Better Rejects: Preference Distillation Scaling
更小的模型,更好的拒绝:偏好蒸馏的规模扩展
arXiv:2609.38987 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,有效的拒绝能够保留任务结构,同时限制与参考策略的耦合,并且较小的冻结模型可以低成本地提供此类拒绝能力。The results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.

2.3 本轮补充公开检索
arXiv:2606.14061 Agent 智能体 方法 OA · 绿色 被引 5 · S2

结果表明,纯视觉设置会降低准确率并增加 token 成本,因为 Agent 缺乏足够的符号化细节,需通过重复的视觉查询进行补偿;研究指向一种面向下一代编码 Agent 的实用文本与视觉混合设计。The results show that a strictly vision-only setup degrades accuracy and increases token cost, because agents lack sufficient symbolic detail and compensate with repeated visual queries, and point to a practical hybrid text-and-vision design for next-generation coding agents.

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
面向异构多 Agent LLM 的免预填充跨系列 KV Cache 迁移
arXiv:2609.32259 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,HeteroFold 能够在接收端无需预填充的前提下实现高效的跨系列 KV 复用,并在全部四个长上下文基准和大多数短上下文设置上取得最佳的 cache 迁移性能。Results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill, and achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.

DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
DataMagic:通过声明式多 Agent 编排制作数据可视化视频
arXiv:2609.33403 Agent 智能体 方法 OA · 绿色 被引 2 · S2

DataMagic 通过声明式多 Agent 编排,从原始表格数据自动生成数据可视化视频,在完全自动化与细粒度人工控制之间架起桥梁,并提升了创作效率、降低了感知认知负荷。DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration, bridging full automation with fine-grained human control, and improves creation efficiency and reduces perceived cognitive load.

LOCI: Spatial Linear Memory for Streaming World Models
LOCI:面向流式世界模型的空间线性记忆
arXiv:2609.40222 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LOCI,一种混合的空间记忆架构,同时维护键值缓存与循环记忆两种表示,在重访内容的复现上比代表性世界模型以及同配置的 full-softmax 模型都更为忠实。LOCI is introduced, a hybrid spatial-memory architecture that keeps both representations of key-value caches and recurrent memory that reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model.

Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans
Architect-Ant:可编辑的建筑平面图自动家具布置
arXiv:2606.10953 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 AntPlan,一个包含 505 张真实专业建筑平面图、覆盖 92 类物体和十类住宅房间且具有密集家具标注的精选数据集,以及 Architect-Ant,一个用于生成家具布局的框架,可在不依赖高成本迭代式 Agent 推理的情况下直接进行约束感知的布局生成。AntPlan, a curated dataset of 505 real professional architectural floor plans with dense furniture annotations spanning 92 object classes and ten residential room categories, and Architect-Ant, a framework for generating furniture layouts are introduced, enabling direct constraint-aware layout generation without relying on costly iterative agentic inference.