Towards a Densing Law for User Representation Learning at Billion-Scale Capacity
- 类型:arxiv
- 标识:2608.23392
- 链接:https://arxiv.org/abs/2608.23392
- 主分类:engineering
- 形态:method
- TLDR:User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance exhibit diminishing performance gains with larger-scale raw text user behavioral input, which can be mitigated by tokenization. (ii) Lack of quantitative analysis of how tokenization configurations should scale with data size. In this report, we propose User Behavioral Densing Law for characterizing the quantitative rela
- 待LLM分类:是
- 标题中文:迈向十亿级容量用户表示学习的稠密定律
- TLDR中文:真实工业场景中的用户表示学习通常通过增加用户数量、行为序列长度和模型规模来扩展。然而现有方法面临两个挑战:(i) 十亿级容量下原始数据扩展的瓶颈,因为随着更大规模原始文本用户行为输入的增加,性能增益呈现递减趋势,这可以通过 tokenization 缓解;(ii) 缺乏对 tokenization 配置应如何随数据规模扩展的定量分析。本报告中,我们提出 User Behavioral Densing Law 来刻画定量关系
- 来源文件:
- /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json