arXiv:2609.15991 · RAG 检索增强
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
The Functionalizer:面向子词分词的无损函数分解
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
- 类型:arxiv
- 标识:2609.15991
- 链接:https://arxiv.org/abs/2609.15991
- 主分类:rag
- 形态:position
- TLDR:Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and Héllo) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization. We present the Functionalizer, a lossless pre-tokenizer framework that factors orthographic and structural variations into a compositional opcode/operand prefix stream before tokenization: a canonical base token (operand) prefixed by parametric transformation operators (opcodes) encoded in the Unicode Private Use Area. We introduce operators covering c
- 待LLM分类:否
- 标题中文:The Functionalizer:面向子词分词的无损函数分解
- TLDR中文:标准子词分词器要么将同一词汇的每种拼写变体(如 hello、Hello、HELLO、Héllo)视为互不相关的词表条目,从而割裂嵌入空间,要么通过有损归一化丢弃这些变体。我们提出 Functionalizer,一种无损预分词框架,在分词之前将拼写与结构变体分解为合成的操作码/操作数前缀流:由 Unicode 私有使用区编码的形参变换操作符(操作码)作为前缀,后接规范化基底 token(操作数)。我们引入覆盖大小写
- 来源文件:
- /inbox/tom/_candidates/2026-09-23-agent-rag-longcontext-candidates.json