论文 Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
论文 PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
论文 1️⃣2️⃣ arXiv · RAGPerf: End-to-End RAG Benchmarking Framework(⭐⭐⭐ 参考)
论文 AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
论文 Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
论文 EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?