Papers · organized/paper_cards

论文

81 张论文卡片 · 评测基准 · OA 绿色

开放获取 全部 绿色 · 724
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
LakeQuest:面向跨数据湖根植式问答的三领域基准
arXiv:2607.12310 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 LakeQuest,一个经人工校验的基准,包含 9,846 个 QA 对,用于在真实数据湖场景下评估端到端检索与综合 pipeline,并揭示现代问答系统中的关键失效模式。LakeQuest is introduced, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes and exposes critical failure modes in modern QA systems.

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
LLMs 是否准备好进行科学发现?面向 AI 科学家的能力导向基准
arXiv:2607.11079 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SDABench,一个围绕六项能力(描述性、探索性、推断性、推断性、预测性、因果性、机制性)跨五大领域(生物、化学、环境、地理、物理)重新组织评估的基准。SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
SynthDocBench:面向长上下文视觉文档理解的受控基准
arXiv:2607.10400 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

视觉语言模型(VLMs)在 DocVQA、ChartQA、MMLongBench-Doc 等视觉文档理解基准上表现强劲。但真实文档融合长度、布局复杂度、模态、问题难度等多因素,难以将模型失败归因于具体原因。我们提出 SynthDocBench,一个全合成的长上下文视觉文档理解基准,系统性控制文档长度、布局结构、模态构成与问题类型等因子。该基准的构建……Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constr

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
arXiv:2607.13705 评测基准 评测集 OA · 绿色 被引 1 · S2

AgentCompass 被提出,它是一个开源、轻量且可扩展的面向 LLM-based Agent 的评估基础设施,将评估流程围绕三个独立组件组织,从而在不重新实现复杂执行逻辑的前提下支持灵活配置。AgentCompass is introduced, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents that organizes the evaluation process around three independent components, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
从受控环境到真实世界:面向实际场景的渗透测试 Agent 评估
arXiv:2605.10834 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种实用评估协议,将评估重点从任务完成转向经过验证的漏洞发现,可在涵盖多种攻击面与漏洞类型的足够复杂目标上开展评估,并结合结构化真值标注与基于 LLM 的语义匹配来识别漏洞。This paper presents a practical evaluation protocol that shifts assessment from task completion to validated vulnerability discovery, allowing evaluation in sufficiently complex targets spanning multiple attack surfaces and vulnerability classes, and combines structured ground-truth with LLM-based semantic matching to identify vulnerabilities.

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Self in Space:面向 UAV 具身智能的自我意识与空间认知基准
arXiv:2607.12477 评测基准 评测集 OA · 绿色 被引 3 · S2

本文提出 SIS-Bench,一个在统一 self-in-space 表述下评估 UAV 场景具身空间智能的基准,并探索了一种融合光流与视觉特征的运动感知表征,以纳入与自身相关的动态信息。SIS-Bench is introduced, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation, and a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion is explored.

Length Penalties Make Chain-of-Thought Less Monitorable
长度惩罚使 Chain-of-Thought 更难被监控
arXiv:2607.09786 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

压缩可减少推理 token 并保持大部分选择题准确率,同时提示影响接近基线——在一个前沿条件下,降低推理成本移除的证据比单纯缩短轨迹所预期的更多。Compression reduces reasoning tokens and preserves most multiple choice accuracy, while hint influence remains near baseline, in a frontier where reducing reasoning costs removes more evidence than shorter traces alone would predict.

UniVR: Thinking in Visual Space for Unified Visual Reasoning
UniVR:在视觉空间中思考以实现统一视觉推理
arXiv:2607.12800 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 UniVR,这是首个从纯视觉演示中同时学习复杂推理、细粒度物理动力学和长期规划的研究,并配套首个在纯视觉协议下评估这些异质能力的综合评测套件。UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations, is introduced, the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol.

Rethinking the Evaluation of Harness Evolution for Agents
Rethinking the Evaluation of Harness Evolution for Agents(重新审视 Agent 的 Harness 演化评估)
arXiv:2607.12227 评测基准 评测集 OA · 绿色 被引 9 · S2

对 LLM Agent 自动 harness 演化进行了广泛评估,在相当的反馈和推理预算下,将 harness 演化与简单的测试时扩展及发现类基线进行比较,并在留出任务上评估演化后的 harness 以判断其发现的改进是否具有泛化性。An extensive evaluation of automatic harness evolution for LLM agents is conducted, comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluating evolved harnesses on held-out tasks to assess whether the discovered improvements generalize.

Sparks of Artificial General Intelligence: Early experiments with GPT-4
Sparks of Artificial General Intelligence: Early experiments with GPT-4
arXiv:2303.12712 评测基准 评测集 OA · 绿色 被引 4391 · S2

认为(该早期版本的)GPT-4 属于新一代具备更通用智能的 LLM(与 ChatGPT、谷歌 PaLM 等并列),并讨论了这些模型不断增强的能力及其影响。It is argued that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models, and the rising capabilities and implications of these models are discussed.

Evaluating Large Language Models Trained on Code
Evaluating Large Language Models Trained on Code
arXiv:2107.03374 评测基准 方法 OA · 绿色 被引 11151 · S2

发现对 GPT 语言模型进行重复采样,是为困难 prompt 生成可用解的一种出乎意料的有效策略,并讨论了部署强大代码生成技术在安全性、安全保障与经济等方面的潜在更广泛影响。It is found that repeated sampling from the GPT language model is a surprisingly effective strategy for producing working solutions to difficult prompts, and the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics are discussed.

A Survey on Evaluation of Large Language Models
大语言模型评估综述
arXiv:2307.03109 评测基准 综述 OA · 绿色 被引 3721 · S2

本文对 LLM 的评估方法进行了全面综述,围绕三个关键维度展开:评估什么、在何处评估、如何评估,并为 LLM 评估领域的研究者提供了宝贵洞见。This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate, and offers invaluable insights to researchers in the realm of LLMs evaluation.

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Beyond the Imitation Game:语言模型能力的量化与外推
arXiv:2206.04615 评测基准 评测集 OA · 绿色 被引 2608 · S2

在 BIG-bench 上对 OpenAI 的 GPT 模型、Google 内部稠密 Transformer 架构及 Switch 风格稀疏 Transformer 进行评估,模型规模跨越百万至千亿参数,结果显示性能与校准均随规模提升而改善,但绝对水平仍然欠佳。Evaluation of OpenAI's GPT models, Google-internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters finds that model performance and calibration both improve with scale, but are poor in absolute terms.

Capabilities of GPT-4 on Medical Challenge Problems
GPT-4 在医学挑战性问题上的能力
arXiv:2303.13375 评测基准 评测集 OA · 绿色 被引 1387 · S2

对 SOTA LLM GPT-4 在医学能力考试与基准数据集上进行全面评估,并通过案例研究定性探索其行为,展示了 GPT-4 解释医学推理、为学生定制个性化讲解以及围绕病例交互式构造新反事实场景的能力。A comprehensive evaluation of GPT-4, a state-of-the-art LLM, on medical competency examinations and benchmark datasets and explores the behavior of the model qualitatively through a case study that shows the ability of G PT-4 to explain medical reasoning, personalize explanations to students, and interactively craft new counterfactual scenarios around a medical case.

GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
arXiv:2303.10130 评测基准 应用落地 OA · 绿色 被引 577 · S2

分析表明,借助 LLM,美国约 15% 的工作任务可在保持同等质量的前提下显著提速完成,意味着 LLM 驱动的软件将对底层模型经济影响的规模化产生实质性作用。The analysis suggests that, with access to an LLM, about 15% of all worker tasks in the US could be completed significantly faster at the same level of quality, implying that LLM-powered software will have a substantial effect on scaling the economic impacts of the underlying models.

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
arXiv:2607.18144 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

研究揭示了 LLM 空间能力中的清晰规律:尽管其仍落后于 SOTA 方法,但具有潜力并能同时处理多种空间约束,从而可扩展到异构场景。A clear pattern is revealed in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System
AILQA:面向印度法律体系的 AI 驱动法律问答系统评测
arXiv:2607.18825 评测基准 方法 OA · 绿色 被引 1 · S2

在该研究的评估协议下,部分 AI 生成回答获得了高于可用参考答案的评分,尤其当其包含准确且相关的支撑细节时尤为明显。Under the study's evaluation protocol, some AI-generated responses received higher ratings than the available reference answers, particularly when they contained accurate and relevant supporting details, particularly when they contained accurate and relevant supporting details.

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
转写策略即潜变量:通过词级时序激活可控逐字 ASR
arXiv:2607.18934 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Verbatimize 这一新任务,能够以高质量规范化逐字转录对语音语料库进行可扩展的创建与扩充,并通过有监督的 cross-attention 微调,将不流畅语音上的词级时间戳表现提升至超过强制对齐 baseline。Verbatimize is proposed, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions and supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines.

What do we need to build explainable AI systems for the medical domain?
构建医疗领域可解释 AI 系统,我们需要什么?
arXiv:1712.09923 评测基准 观点 OA · 绿色 被引 964 · S2

本文认为,可解释 AI 研究总体上有助于推动 AI/ML 在医疗领域的落地,并特别有助于增强透明性与信任。It is argued that research in explainable-AI would generally help to facilitate the implementation of AI/ML in the medical domain, and specifically help to facilitates transparency and trust.

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
FinMMEval 2026 任务 2 概述:多语言金融短答问答
arXiv:2607.19867 评测基准 方法 OA · 绿色 被引 1 · S2

FinMMEval 2026 Task 2 围绕多语言证据评估金融领域的短答问答,采用文档 RAG、跨语言证据处理、结构化提示、答案压缩与验证策略。FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence over multilingual evidence using document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
DocOps:面向复杂文档操作的自主 Agent 可验证基准
arXiv:2607.19865 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 DocOps,一种确定性可验证的评估框架,基于分层分类法,将受真实实践启发的文档操作分解为原子维度与逐级递增的工作流复杂度,从而揭示 Agent 在维护全局文档一致性方面的能力边界。DocOps is introduced, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities that exposes the capability boundaries of agents in maintaining global document consistency.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
展示而非讲述:在生成像素而非 LLM 文本中评估空间认知
arXiv:2607.21072 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 ProVisE(Protocolized Visual Evaluation),一个与基准无关的框架,通过受协议约束的视觉问答从图像生成模型中抽取答案,并将其解析为与原始指标兼容的结构化预测,揭示了像素空间表达与基于文本推理的互补优势。ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and text-based reasoning.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
K12-KGraph:用于教育 LLM 基准测试与训练的课程对齐知识图谱
arXiv:2605.09635 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 K12-KGraph,一个从人民教育出版社官方教材中提取的、与课程对齐的知识图谱,覆盖小学、初中和高中阶段的数学、物理、化学和生物,并由此衍生出 K12-Bench,一个包含 23,640 道题的多选基准,涵盖五类任务族,表明文本监督与视觉监督具有互补性。This work introduces K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school, and derives K12-Bench, a 23,640-question multi-select benchmark with five task families, showing that textual and visual supervision are complementary.

Mathematical Capabilities of ChatGPT
ChatGPT 的数学能力
arXiv:2301.13867 评测基准 评测集 OA · 绿色 被引 603 · S2

研究发现,ChatGPT 最成功地被用作数学助手,用于查询事实、充当数学搜索引擎和知识库接口;GPT-4 还可被用于本科水平的数学问题,但在研究生难度的题目上表现不佳。It is found that ChatGPT can be used most successfully as a mathematical assistant for querying facts, acting as a Mathematical search engine and knowledge base interface, and GPT-4 can additionally be used for undergraduate-level mathematics but fails on graduate-level difficulty.

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
ChatGPT 生成的代码真的正确吗?面向代码生成的大型语言模型严格评估
arXiv:2305.01210 评测基准 评测集 OA · 绿色 被引 2078 · S2

EvalPlus——一个用于严格基准测试 LLM 生成代码功能正确性的代码合成评估框架,通过 LLM 与基于 mutation 的策略驱动的自动测试输入生成器,为给定评估数据集补充大量新生成的测试用例。EvalPlus -- a code synthesis evaluation framework to rigorously benchmark the functional correctness of LLM-synthesized code and augments a given evaluation dataset with large amounts of test-cases newly produced by an automatic test input generator, powered by both LLM and mutation-based strategies.

Towards Expert-Level Medical Question Answering with Large Language Models
迈向基于大语言模型的专家级医学问答
arXiv:2305.09617 评测基准 应用落地 OA · 绿色 被引 808 · S2

结果表明,通过结合基础 LLM 改进(PaLM 2)、医学领域微调以及包括新颖集成精化方法在内的提示策略,医学问答正快速接近医生水平的表现。Results highlight rapid progress towards physician-level performance in medical question answering by leveraging a combination of base LLM improvements (PaLM 2), medical domain finetuning, and prompting strategies including a novel ensemble refinement approach.

Characterizing Warp Divergence from Pascal to Blackwell
从 Pascal 到 Blackwell 的 Warp 分歧特性分析
arXiv:2607.23402 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

即使 NVIDIA 控制流 ISA 与重汇聚(reconvergence)机制持续演进,Divergence 仍能保持稳定且可预期的性能开销。Divergence retains a stable and predictable performance cost even as NVIDIA's control-flow ISA and reconvergence mechanisms continue to evolve.

Codifying the Judge: Scalable Evaluation via Program Distillation
Codifying the Judge:通过程序蒸馏实现可扩展评估
arXiv:2607.22561 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PAJAMA,该系统将程序合成为 judge,将其决策聚合为联合裁决(joint verdict),并通过 fallback 机制选择性地将低置信度用例升级交由 LLM 处理。PAJAMA is introduced, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM.

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
Keep It InMind:基准测试 Agent 记忆中的隐式关联盲点
arXiv:2607.24368 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 InMind,一项包含 125 个任务、经专家校验、跨越十个生活领域的基准,其中 113 个任务基于可引用的公开来源;本文将此类失效模式命名为 implicit-association blind spot,并提出一种极简诊断探针(diagnostic probe),在 query 到达前保持 memory 可见即可恢复大部分性能差距。This work introduces InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources, and calls this failure mode the implicit-association blind spot, and introduces a minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
PerceptionBench:评估多模态大语言模型的原子级视觉感知
arXiv:2607.24957 评测基准 评测集 OA · 绿色 被引 1 · S2

PerceptionBench 通过诊断前沿 MLLMs 在 42 项现有基准上响应中的最早失效点,并构建一个感知分支定义十种原子感知能力的错误分类法,为衡量与诊断 MLLMs 视觉感知边界提供了一项能力级(capability-level)标准。PerceptionBench provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs, by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks and constructing an error taxonomy whose perception branch defines ten atomic perceptual capabilities.

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models
RepoReasoner:长上下文语言模型仓库级代码推理能力评估
arXiv:2607.25996 评测基准 评测集 OA · 绿色 被引 1 · S2

介绍 RepoReasoner,一个用于评估仓库级代码推理的 benchmark,评估两种互补能力:Output Prediction,衡量跨文件的细粒度、有状态执行推理;Call Chain Prediction,在噪声上下文下评估高层架构依赖理解。RepoReasoner is introduced, a benchmark for evaluating repository-level code reasoning that assesses two complementary abilities: Output Prediction, which measures fine-grained, stateful execution reasoning across files, and Call Chain Prediction, which evaluates high-level architectural dependency understanding under noisy context.

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
DecoEvo: 文本空间中 Solver 与 Rubric 生成器技能的分数解耦协同演化
arXiv:2607.25675 评测基准 方法 OA · 绿色 被引 2 · S2

介绍 DecoEvo(Decoupled Co-Evolution),在解耦目标下协同进化一个求解器 skill 和一个 rubric 生成器 skill,优化过程中不使用 gold rubric。DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization, is introduced.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
CoRT: 用于 token 级 rubric 引导策略优化的反事实回放
arXiv:2607.25659 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 CoRT,一种用于 rubric 条件化 GRPO 的 token 级 credit 加权方法,通过反事实回放在原始 rubric 条件化 prompt 和匹配的免准则 prompt 下对同一样本响应重新打分,并表明策略内部反事实似然对比为响应内 credit 分配提供了有效的训练信号。CoRT is proposed, a token-level credit weighting method for rubric-conditioned GRPO that uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt, and suggests that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation.

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
超越借用历史:面向交互式角色扮演评估的个体对齐用户模拟
arXiv:2607.27816 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PALATE(Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation),一个基于用户模拟器的可扩展 RPA benchmark,可针对具体 user-RPA 对给出可解释的评估,避免将系统压缩为单一、与用户无关的排名。PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation with Tailored Evaluation), a scalable RPA benchmark built on user simulators, produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Fairness Pruning:通过差异激活定位 GLU-MLP 层中的人口统计偏差
arXiv:2607.28319 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

对 Fairness Pruning 的实证评估表明,群体偏置处理与模型能力运行在可分离的电路上,奠定了从盲目零化向定向行为调制过渡的方法论基础。Empirical evaluation of Fairness Pruning empirically confirm that demographic bias processing and model capabilities operate on dissociable circuits, establishing the methodological foundations for transitioning from blind zeroing toward directional behavior modulation.

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
ExtractBench:一个面向模式引导的企业文档抽取基准
arXiv:2607.29677 评测基准 评测集 OA · 绿色 被引 1 · S2

LlamaExtract Agentic Plus在三项指标上均排名第一,准确度可与coding agent相媲美而成本仅为其一小部分,是首个同时在大规模下对数值准确性、记录完整性、grounding与实测成本进行打分的方法。LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.