Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

  • 类型:arxiv
  • 标识:2608.04378
  • 链接:https://arxiv.org/abs/2608.04378
  • 主分类:agent
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:A hierarchical self-supervised ``world model'' for symbolic music is presented, using a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music-theory vocabulary.
  • 副分类:rag
  • 待LLM分类:否
  • 标题中文:帮助音乐共创 Agent "听懂":用于理解与生成的分层自监督世界模型
  • TLDR中文:提出一种用于符号音乐的分层自监督"世界模型",采用 2.55M 参数的 Swin V2 编码器,在 MIDI 钢琴卷帘图像上以 JEPA 风格目标(音高与时间平移等变性、掩码嵌入预测以及分布正则化)训练,无需标签与乐理词汇。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-07-agent-rag-longcontext-candidates.json
  • [S2 enrich]