研究库 论文知识库
Papers · organized/paper_cards

论文

9 张论文卡片 · 评测基准 · 应用落地 · OA 绿色

开放获取 全部 绿色 · 1640
Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems
音乐上下文保留评估:面向音乐编辑系统的多面框架
arXiv:2512.14629 评测基准 应用落地 OA · 绿色 被引 1 · S2

提出首个 MuseCP 评估框架,涵盖四类音乐 facet,使用细粒度且量身定制的指标来捕捉音乐属性的细微变化,并希望为开发更有效、更可靠、具备强大 MuseCP 能力的音乐编辑策略提供实践指导。The first MuseCP evaluation framework is introduced that covers four categories of music facets with fine-grained and well-tailored metrics to capture nuanced changes in music attributes and hopes it can offer practical guidance for developing more effective and reliable music editing strategies with strong MuseCP capability.

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
分层自改进:面向任务特定可进化 Agent Harness 的框架
arXiv:2608.08466 评测基准 应用落地 OA · 绿色 被引 5 · S2

结果表明,任务特定的 harness 进化是改进冻结 LLM Agent 的一条可行路径,但其效果存在明确的经验性边界。Results demonstrate task-specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits.

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
EvoSafeHarness:面向 Agent 安全防护的演进式模型与领域专属 Harness
arXiv:2609.05903 评测基准 应用落地 OA · 绿色 被引 2 · S2

EvoSafeHarness 是一个面向安全性的优化框架,针对目标领域中的冻结模型合成可部署的 harness,由模型行为、领域规范和新鲜上下文对抗审查共同引导,以拒绝基准特定规则。EvoSafeHarness is a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules.

"I think this is the most disruptive technology": Exploring Sentiments of ChatGPT Early Adopters using Twitter Data
"我认为这是最具颠覆性的技术":基于 Twitter 数据探索 ChatGPT 早期采用者的情感
arXiv:2212.05856 评测基准 应用落地 OA · 绿色 被引 284 · S2

基于 10,732 条早期 ChatGPT 用户推文的混合方法研究,对每个主题进行深入定性情感分析,结果显示大多数早期采用者在软件开发颠覆性、娱乐与创意发挥等主题上表达了压倒性的积极情感。A mixed-method study using 10,732 tweets from early ChatGPT users to conduct an in-depth qualitative sentiment analysis of each topic, showing that the majority of the early adopters have expressed overwhelmingly positive sentiments related to topics such as Disruptions to software development, Entertainment and exercising creativity.

Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks
面向 IIoT 网络的轻量级入侵检测模型的跨域泛化失效
arXiv:2607.00553 评测基准 应用落地 OA · 绿色 被引 2 · S2

应在真实类别分布下使用跨网络评估来判断部署就绪度,而非仅依赖域内准确率。Deployment readiness should be assessed using cross-network evaluation under realistic class distributions, rather than within-domain accuracy alone, to suggest deployment readiness should be assessed using cross-network evaluation under realistic class distributions, rather than within-domain accuracy alone.

GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
arXiv:2303.10130 评测基准 应用落地 OA · 绿色 被引 600 · S2

分析表明,借助 LLM,美国约 15% 的工作任务可在保持同等质量的前提下显著提速完成,意味着 LLM 驱动的软件将对底层模型经济影响的规模化产生实质性作用。The analysis suggests that, with access to an LLM, about 15% of all worker tasks in the US could be completed significantly faster at the same level of quality, implying that LLM-powered software will have a substantial effect on scaling the economic impacts of the underlying models.

Towards Expert-Level Medical Question Answering with Large Language Models
迈向基于大语言模型的专家级医学问答
arXiv:2305.09617 评测基准 应用落地 OA · 绿色 被引 826 · S2

结果表明,通过结合基础 LLM 改进(PaLM 2)、医学领域微调以及包括新颖集成精化方法在内的提示策略,医学问答正快速接近医生水平的表现。Results highlight rapid progress towards physician-level performance in medical question answering by leveraging a combination of base LLM improvements (PaLM 2), medical domain finetuning, and prompting strategies including a novel ensemble refinement approach.

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools
基于分类法的开源 AI 风险缓解工具分析
arXiv:2608.07446 评测基准 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种结构化协议,通过对开源 LLM 评估与安全工具的分类驱动分析来自动化 AI 风险缓解,并给出一个可同时适用于开源与商用方案的分类驱动框架。This paper proposes a structured protocol to automate AI risk mitigation through a taxonomy-driven analysis of open-source LLM evaluation and security tools, and presents a taxonomy-driven framework applicable to open-source and proprietary solutions.

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
LLMRouter:面向 LLM 路由器开发、评估与部署的统一基础设施
arXiv:2608.06867 评测基准 应用落地 OA · 绿色 被引 1 · S2

本文给出了 LLM routing 的统一形式化,将其刻画为由五个组件构成的序贯决策过程:context 编码器、模型编码器、评分函数、决策规则和学习信号,涵盖单轮、多轮和个性化 routing。This work presents a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing.