Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

  • 类型:arxiv
  • 标识:2608.05138
  • 链接:http://arxiv.org/abs/2608.05138v1
  • 主分类:rag
  • 形态:benchmark
  • 被引:1
  • 被引来源:Semantic Scholar
  • S2被引:1
  • 影响力被引:0
  • TLDR:This study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora and introduces HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and releases adapted models and benchmark to support future research on Greek-language RAG systems.
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:Teaching Nemotron Greek:挖掘语料、适配检索并为现代希腊语在专业领域提供有据生成的全文
  • TLDR中文:研究表明在希腊语专业语料上,无参数 BM25 基线优于多个现成的多语言稠密检索模型,并推出首个大规模希腊语 RAG 基准 HERA,同时发布适配模型与基准以支撑未来希腊语 RAG 系统研究。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-06-agent-rag-longcontext-candidates.json
  • /inbox/tom/_candidates/2026-08-07-agent-rag-longcontext-candidates.json
  • [S2 enrich]