arXiv:2609.34274 · Agent 智能体
BIABench: Evaluating AI agents on real-world bioimage analysis tasks
BIABench: 在真实生物图像分析任务上评估 AI agent
BIABench: Evaluating AI agents on real-world bioimage analysis tasks
- 类型:arxiv
- 标识:2609.34274
- 链接:https://arxiv.org/abs/2609.34274
- 主分类:agent
- 形态:benchmark
- TLDR:Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so an agent must choose and run an analysis through code, specialized software and rendered views. Published studies make this capability testable, because each pairs raw images with a peer-reviewed result. We introduce BIABench, a benchmark of 16 tasks reconstructed from published biological studies
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:BIABench: 在真实生物图像分析任务上评估 AI agent
- TLDR中文:人工智能 (AI) agent 有望自动化生物图像分析,但尚无基准评估其能否完成真实场景下的端到端分析。此类分析对 agent 而言具有难度,因为二维图像、三维体数据与时序序列通常过大、无法直接读入上下文,agent 必须通过代码、专业软件与渲染视图来选择并执行分析。已发表的研究使该能力可被测试,因为每一篇都包含原始图像与同行评审结果的配对。我们提出 BIABench,一个由 16 项任务构成、基于已发表研究重建的基准。
- 来源文件:
- /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json