Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
- 类型:arxiv
- 标识:2608.05138
- 链接:http://arxiv.org/abs/2608.05138v1
- 主分类:rag
- 形态:benchmark
- 被引:1
- 被引来源:Semantic Scholar
- S2被引:1
- 影响力被引:0
- TLDR:This study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora and introduces HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and releases adapted models and benchmark to support future research on Greek-language RAG systems.
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:Teaching Nemotron Greek:挖掘语料、适配检索并为现代希腊语在专业领域提供有据生成的全文
- TLDR中文:研究表明在希腊语专业语料上,无参数 BM25 基线优于多个现成的多语言稠密检索模型,并推出首个大规模希腊语 RAG 基准 HERA,同时发布适配模型与基准以支撑未来希腊语 RAG 系统研究。
- 来源文件:
- /inbox/tom/_candidates/2026-08-06-agent-rag-longcontext-candidates.json
- /inbox/tom/_candidates/2026-08-07-agent-rag-longcontext-candidates.json
- [S2 enrich]