研究库 论文知识库
Papers · organized/paper_cards

论文

287 张论文卡片 · 评测集

开放获取 全部 绿色 · 1640
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
arXiv:2607.20092 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ENTRAP-VL(ENTRainment Assessment Probe for Vision and Language),一个由人工策展的 1,500 条数据的数据集,涵盖八个类别,按一个跨双轴的分类体系组织,并划分为文本诱发流和视觉诱发流。ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes and split into a textual-entrainment stream and a visual-entrainment stream, is introduced.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
展示而非讲述:在生成像素而非 LLM 文本中评估空间认知
arXiv:2607.21072 评测基准 评测集 OA · 绿色 被引 2 · S2

提出 ProVisE(Protocolized Visual Evaluation),一个与基准无关的框架,通过受协议约束的视觉问答从图像生成模型中抽取答案,并将其解析为与原始指标兼容的结构化预测,揭示了像素空间表达与基于文本推理的互补优势。ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and text-based reasoning.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
K12-KGraph:用于教育 LLM 基准测试与训练的课程对齐知识图谱
arXiv:2605.09635 评测基准 评测集 OA · 绿色 被引 2 · S2

本文提出 K12-KGraph,一个从人民教育出版社官方教材中提取的、与课程对齐的知识图谱,覆盖小学、初中和高中阶段的数学、物理、化学和生物,并由此衍生出 K12-Bench,一个包含 23,640 道题的多选基准,涵盖五类任务族,表明文本监督与视觉监督具有互补性。This work introduces K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school, and derives K12-Bench, a 23,640-question multi-select benchmark with five task families, showing that textual and visual supervision are complementary.

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
FinanceComplexQA:面向工业级金融文档的智能体推理基准测试
arXiv:2607.19238 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

本文设计 Finance-LaTeX SKILL,一个基于专家知识合成复杂版面金融文档的 skill,并提出 FinanceComplexQA,一个全面、贴近真实场景的金融文档开放式生成基准。This work designs Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge, and introduces FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios.

Mathematical Capabilities of ChatGPT
ChatGPT 的数学能力
arXiv:2301.13867 评测基准 评测集 OA · 绿色 被引 621 · S2

研究发现,ChatGPT 最成功地被用作数学助手,用于查询事实、充当数学搜索引擎和知识库接口;GPT-4 还可被用于本科水平的数学问题,但在研究生难度的题目上表现不佳。It is found that ChatGPT can be used most successfully as a mathematical assistant for querying facts, acting as a Mathematical search engine and knowledge base interface, and GPT-4 can additionally be used for undergraduate-level mathematics but fails on graduate-level difficulty.

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
ChatGPT 生成的代码真的正确吗?面向代码生成的大型语言模型严格评估
arXiv:2305.01210 评测基准 评测集 OA · 绿色 被引 2333 · S2

EvalPlus——一个用于严格基准测试 LLM 生成代码功能正确性的代码合成评估框架,通过 LLM 与基于 mutation 的策略驱动的自动测试输入生成器,为给定评估数据集补充大量新生成的测试用例。EvalPlus -- a code synthesis evaluation framework to rigorously benchmark the functional correctness of LLM-synthesized code and augments a given evaluation dataset with large amounts of test-cases newly produced by an automatic test input generator, powered by both LLM and mutation-based strategies.

Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education
Teachy Mini:面向高等教育的知识驱动生成式社交机器人开发与初步评估
arXiv:2607.22345 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究在 Reachy Mini 机器人平台上,通过系统提示、检索增强生成 (RAG) 和有状态的提示编排,将选定的 KBD 需求落地实现,表明 KBD 可以塑造负责任的机器人行为,并有望提升机器人辅助学习中的学习效果。This study operationalized selected KBD requirements in the Reachy Mini robot platform through system prompting, retrieval-augmented generation, and stateful prompt orchestration, indicating that KBD can shape responsible robot behavior and potentially increase learning effectiveness in robot-supported learning.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
DataPrep-Bench: 将 LLM 作为训练数据准备器的基准测试
arXiv:2607.20465 工程化 评测集 OA · 绿色 被引 1 · S2

提出 DataPrep-Bench,首个统一基准,在共享的下游任务 grounding 协议下,对 LLM 驱动的数据准备在六个领域、多种 base model 上的两类能力进行联合评估。DataPrep-Bench is introduced, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains and multiple base models of LLM-driven data preparation.

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
无坐标与区域标签的视觉文档理解中的证据归因
arXiv:2607.24651 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究考察是否存在一条实用路径,在没有坐标界面、且无需高成本区域级监督的条件下提升归因效果,并指出了这样一条可行路径。A study investigates whether there is a practical path to improve attribution without a coordinate interface and without costly region-level supervision, and indicates a practical path to improve attribution without a coordinate interface and without costly region-level supervision.

Characterizing Warp Divergence from Pascal to Blackwell
从 Pascal 到 Blackwell 的 Warp 分歧特性分析
arXiv:2607.23402 评测基准 评测集 OA · 绿色 被引 1 · S2

即使 NVIDIA 控制流 ISA 与重汇聚(reconvergence)机制持续演进,Divergence 仍能保持稳定且可预期的性能开销。Divergence retains a stable and predictable performance cost even as NVIDIA's control-flow ISA and reconvergence mechanisms continue to evolve.

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
IndicTalk:面向印度语系的大规模人格化多语言对话语料库
arXiv:2607.23242 LLM 基础设施 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IndicTalk,目前最大的多语言印度语码混合(code-mixed)会话语料库,包含超过 13,28,604 段事件驱动的多轮对话,覆盖 9 种印度语言的 18 种语言变体,将公开发布以支持低资源印度语多语言会话 AI 的开发与评估。IndicTalk is presented, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages and will be released to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages.

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
Keep It InMind:基准测试 Agent 记忆中的隐式关联盲点
arXiv:2607.24368 评测基准 评测集 OA · 绿色 被引 1 · S2

本文提出 InMind,一项包含 125 个任务、经专家校验、跨越十个生活领域的基准,其中 113 个任务基于可引用的公开来源;本文将此类失效模式命名为 implicit-association blind spot,并提出一种极简诊断探针(diagnostic probe),在 query 到达前保持 memory 可见即可恢复大部分性能差距。This work introduces InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources, and calls this failure mode the implicit-association blind spot, and introduces a minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
PerceptionBench:评估多模态大语言模型的原子级视觉感知
arXiv:2607.24957 评测基准 评测集 OA · 绿色 被引 4 · S2

PerceptionBench 通过诊断前沿 MLLMs 在 42 项现有基准上响应中的最早失效点,并构建一个感知分支定义十种原子感知能力的错误分类法,为衡量与诊断 MLLMs 视觉感知边界提供了一项能力级(capability-level)标准。PerceptionBench provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs, by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks and constructing an error taxonomy whose perception branch defines ten atomic perceptual capabilities.

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models
RepoReasoner:长上下文语言模型仓库级代码推理能力评估
arXiv:2607.25996 评测基准 评测集 OA · 绿色 被引 6 · S2

介绍 RepoReasoner,一个用于评估仓库级代码推理的 benchmark,评估两种互补能力:Output Prediction,衡量跨文件的细粒度、有状态执行推理;Call Chain Prediction,在噪声上下文下评估高层架构依赖理解。RepoReasoner is introduced, a benchmark for evaluating repository-level code reasoning that assesses two complementary abilities: Output Prediction, which measures fine-grained, stateful execution reasoning across files, and Call Chain Prediction, which evaluates high-level architectural dependency understanding under noisy context.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
HANDBOOK.md:面向长上下文 Agent 指令遵循的基准
arXiv:2607.25398 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

介绍 HANDBOOK_md,一个包含 65 个 agentic 任务的 benchmark,模拟员工遵循公司手册的方式;每项任务对 10 份基础手册之一进行修改,变动评分所依赖的具体规则和阈值,因此没有任何两个任务共享同一套策略。HANDBOOK_md is presented, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks, and every task modifies one of 10 base handbooks, altering the specific rules and thresholds on which grading depends, so no two tasks share the same set of policies.

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
Agent Retrieval Bench:面向编码 Agent 的仓库上下文检索评测
arXiv:2607.24882 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

一项受控的种子干预实验发现,与随机非 gold 上下文相比,由检索得到的初始上下文以更少的后续探索取得更高的文件 F1,而 oracle gold 上下文仍存在可观的提升空间。A controlled seed-intervention pilot finds that retrieval-derived initial context yields higher file F1 with less post-seed exploration than random non-gold context, while oracle gold context shows substantial remaining headroom.

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
CLBench-V: 评估多模态上下文学习——从 grounding 到知识获取
arXiv:2607.25294 多模态 评测集 OA · 绿色 被引 1 · S2

本文介绍 CLBench-V,一个多模态上下文学习 benchmark,围绕三个维度组织任务——上下文 grounding、新信息应用与新知识学习——以解决定位上下文使用失效位置的难题。This work introduces CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
StealthBench: 衡量自主攻击性安全 Agent 的操作隐蔽性
arXiv:2607.26314 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

介绍 StealthBench,一个跨越六个 OPSEC 维度衡量自主攻击性安全 agent 操作隐蔽性的 benchmark,并以公共 benchmark 形式发布,以支持隐蔽感知 agent 的开发及自主攻击性安全部署中的自动化 OPSEC 监控。StealthBench, a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions, is introduced and released as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments.

KAMR: Grounding Generation via Knowledge-Aligned Multi-hop Retrieval
KAMR: 基于知识对齐多跳检索的接地生成
arXiv:2607.27136 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出知识对齐的多跳检索器 KAMR,区分受 query 强约束的 anchor triplet 和弱对齐但在结构上与 anchor 相连的 connected triplet,持续提升多跳检索及下游问答性能。A knowledge-aligned multi-hop retriever, KAMR, which distinguishes anchor triplets that are strongly constrained by the query from connected triplets that are weakly aligned yet structurally linked to the anchors, which consistently improves multi-hop retrieval and downstream question answering performance.

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
SpecFirst: 将行为规约获取作为基于 Agent 从零程序合成中的一等步骤
arXiv:2607.27167 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SpecFirst,一个两阶段框架,在代码合成前强制进行需求 elicitation,并证明显式需求工程阶段是从零构建程序的有效范式。This work presents SpecFirst, a two-stage framework that forces the specification elicitation before code synthesis, and demonstrates that an explicit requirements-engineering phase is an effective paradigm for from-scratch program construction.

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
超越借用历史:面向交互式角色扮演评估的个体对齐用户模拟
arXiv:2607.27816 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PALATE(Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation),一个基于用户模拟器的可扩展 RPA benchmark,可针对具体 user-RPA 对给出可解释的评估,避免将系统压缩为单一、与用户无关的排名。PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation with Tailored Evaluation), a scalable RPA benchmark built on user simulators, produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.

See2Think: Do Multimodal Models Really Use Intermediate Visual States?
See2Think:多模态模型真的使用了中间视觉状态吗?
arXiv:2607.26769 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

对代表性闭源与开源多模态模型的评测表明,视觉推理强依赖于模型与环境,没有任何单一设置能在所有任务上持续占优。Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.

Educating the Agentic Engineer: Curricula, Collaboration, and Continuous Learning in the AI Era
培养 Agentic 工程师:AI 时代的课程、协作与持续学习
arXiv:2607.29610 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

教育agentic工程师需要系统性变革而非增量式课程改革:教学必须从产出工件转向对日益自主的社会-技术系统进行判断。It is concluded that educating the agentic engineer requires systemic transformation rather than incremental curricular change: instruction must shift from producing artifacts to exercising judgment over increasingly autonomous socio-technical systems.

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
ExtractBench:一个面向模式引导的企业文档抽取基准
arXiv:2607.29677 评测基准 评测集 OA · 绿色 被引 1 · S2

LlamaExtract Agentic Plus在三项指标上均排名第一,准确度可与coding agent相媲美而成本仅为其一小部分,是首个同时在大规模下对数值准确性、记录完整性、grounding与实测成本进行打分的方法。LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
更少澄清,更优代码:面向编程助手的跨会话个性化歧义自适应基准测试
arXiv:2607.26611 评测基准 评测集 OA · 绿色 被引 1 · S2

CAPA通过六种机制刻画个性化编码歧义,并使用受控的三阶段生成流程将这些机制注入无歧义的可执行任务,为开发长期编码助手奠定基础,使其生成的代码更好地对齐用户意图并减少反复澄清。CAPA, which characterizes personalized coding ambiguity through six mechanisms and injects these mechanisms into unambiguous executable tasks using a controlled three-stage generation pipeline, provides a foundation for developing long-term coding assistants that better align generated code with user intent while reducing repeated clarification.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
模型还是 Harness?面向 Agent 失败定位的以交互为中心的分类法
arXiv:2607.28802 评测基准 评测集 OA · 绿色 被引 5 · S2

本文提出以交互为中心的分类法,将失败定位到其起源的交互并识别责任组件,将 41 种失败模式归到两个组件之间的边及指示修复归属的故障侧This work introduces an interaction-centric taxonomy that localizes failures to the interactions in which they originate and identifies the responsible component, and organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs.

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
PAST-Bench:面向个人 Agent 递归自我改进基础的基准评测
arXiv:2608.04003 Agent 智能体 评测集 OA · 绿色 被引 4 · S2

提出 PAST-Bench 基准,用于隔离评估持久 agent 从经验保留到系统性改进的能力;开发 Hermes+,提升了来自保留经验的平均增益,并提供更清晰的路径证据。The PAST-Bench benchmark is introduced, a benchmark designed to isolate how persistent agents can progress from retaining experience to systematically improving through it, and Hermes+ is developed, which raises the average gain from retained experience and provides clearer pathway evidence.

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
当 Agent 学会成为你:Persona Skill 中隐私泄漏、冒充风险与防御的基准评测
arXiv:2608.03700 评测基准 评测集 OA · 绿色 被引 3 · S2

提出 AntiSkillBench,一个端到端的基准,用于评估 persona-skill 流水线中的风险与防御;实验表明 persona-skill 风险在不同的 agent backbone 和蒸馏协议下持续存在,从显式属性扩展到沟通风格与个性特征。AntiSkillBench is introduced, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline, and experiments show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits.

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
LLM 的金融推理可信吗?基于长周期财报的真实世界检验
arXiv:2607.28661 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

FinIndices 是一个大规模基准,在未裁剪的财务报表(最长 32K tokens)上评估数据处理保真度,带来显著的零提示增益,验证通过以数据为中心的对齐可部分恢复结构化逻辑。FinIndices is a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens) and yields substantial zero-hint gains, validating that structured logic can be partially restored via data-centric alignment.

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
ChronoLens: 跨时间、语言与语言学层面的语言变化测量
arXiv:2608.03507 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ChronoLens,结合冻结的多语言语言模型、特征对齐的 crosscoder 与事后语言学干预,应用于来自五个议会传统、跨越 1803–2026 年的 4498 万篇文档和约 172 亿 tokens,表明历史语言变化是一个结构化的、多维度的过程。ChronoLens is introduced, a framework that combines frozen multilingual language models, feature-aligned crosscoders, and post-hoc linguistic interventions, and applies it to 44.98 million documents and approximately 17.2 billion tokens from five parliamentary traditions spanning 1803--2026, showing that historical language change is a structured, multidimensional process.

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
Teaching Nemotron Greek:挖掘语料、适配检索并为现代希腊语在专业领域提供有据生成的全文
arXiv:2608.05138 RAG 检索增强 评测集 OA · 绿色 被引 1 · S2

研究表明在希腊语专业语料上,无参数 BM25 基线优于多个现成的多语言稠密检索模型,并推出首个大规模希腊语 RAG 基准 HERA,同时发布适配模型与基准以支撑未来希腊语 RAG 系统研究。This study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora and introduces HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and releases adapted models and benchmark to support future research on Greek-language RAG systems.

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
AVE-Compass:迈向音视频编辑能力的整体评估
arXiv:2607.24821 多模态 评测集 OA · 绿色 被引 1 · S2

提出 AVE-Agent,一种模块化 agent 框架,将复杂指令分解为相互依赖的子任务,并通过自我反思与评估器反馈迭代改进编辑结果,在联合编辑中提升指令执行、保真度保持以及音视频对齐,同时保持有竞争力的感知质量。AVE-Agent is proposed, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback, and improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
GDPevo:面向真实业务任务的 agent 自进化评估
arXiv:2608.03764 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 GDPevo,一种基于 GDP 相关企业工作流的 evolution-native 基准,并配套完全自动化的数据生成流水线;表现最佳的进化 agent 仍远低于全信息 oracle 上限,表明当前 agent 的自进化能力远未充分实现。GDPevo is presented, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it, and the best evolved agents remain far below the fully informed oracle ceiling, indicating that the self-evolution ability of current agents remains far from fully realized.

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
SIGNPOST-Bench:面向多模态大语言模型中文本-视觉冲突消解的基准评测
arXiv:2608.04244 多模态 评测集 OA · 绿色 被引 1 · S2

这些结果将视觉地理定位确立为场景文本仲裁的连续诊断手段,并提供了一个受控框架,用于评估 MLLMs 如何解决冲突的多模态证据。These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.

FinanceHarness: Autonomous Financial Deep Research Framework
FinanceHarness:面向金融领域的自主深度研究框架
arXiv:2607.27853 评测基准 评测集 OA · 绿色 被引 1 · S2

FinanceHarness 是一个运行金融工具与从业者引导工作流的框架,端到端自动化金融深度研究:环境与数据构建、Agent 执行循环以及奖励建模;FinanceGym 包含论点驱动的研报问题与评分标准,结合 pre-cutoff 与 post-cutoff 准则。FinanceHarness is presented, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling, and FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria.

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
MameLoshnLM:意第绪语语言模型与评估基准
arXiv:2608.05850 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

MameLoshnLM 是首个专为意第绪语构建的 8B 参数开源语言模型,既为意第绪语 NLP 提供基础,也为历史悠久但数字化程度不足的语言建模开发提供可复用的实践模板。MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish, is presented, providing both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.