Last Translation Benchmark

  • 类型:arxiv
  • 标识:2609.04173
  • 链接:https://arxiv.org/abs/2609.04173
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvemen
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-05-agent-rag-longcontext-candidates.json