Lior-Baruch/Thesis_PTO_GRPO

  • 类型:github
  • 标识:Lior-Baruch/Thesis_PTO_GRPO
  • 链接:https://github.com/Lior-Baruch/Thesis_PTO_GRPO
  • 主题:llm-infra
  • 主分类:engineering
  • 形态:app
  • 分类:ai
  • Stars:0
  • 周增:+0
  • 语言:Jupyter Notebook
  • 许可:MIT
  • 最近提交:2026-10-09
  • 简介:Training small LLM therapists for Motivational Interviewing with look-ahead rewards: Preference Tree Optimization (PTO) vs GRPO, with simulated patients and LLM judges. Master's thesis, Reichman University.
  • 上次采集:2026-10-09
  • 首次采集:2026-10-09
  • 待LLM分类:否
  • 成熟度:experimental
  • 简介中文:用前瞻性奖励训练面向动机式访谈的小型 LLM 治疗师:Preference Tree Optimization (PTO) vs GRPO,配合模拟患者和 LLM 评委。里赫曼大学硕士论文。
  • 来源文件:
  • [GitHub Search]