Lior-Baruch/Thesis_PTO_GRPO
- 类型:github
- 标识:Lior-Baruch/Thesis_PTO_GRPO
- 链接:https://github.com/Lior-Baruch/Thesis_PTO_GRPO
- 主题:llm-infra
- 主分类:engineering
- 形态:app
- 分类:ai
- Stars:0
- 周增:+0
- 语言:Jupyter Notebook
- 许可:MIT
- 最近提交:2026-10-09
- 简介:Training small LLM therapists for Motivational Interviewing with look-ahead rewards: Preference Tree Optimization (PTO) vs GRPO, with simulated patients and LLM judges. Master's thesis, Reichman University.
- 上次采集:2026-10-09
- 首次采集:2026-10-09
- 待LLM分类:否
- 成熟度:experimental
- 简介中文:用前瞻性奖励训练面向动机式访谈的小型 LLM 治疗师:Preference Tree Optimization (PTO) vs GRPO,配合模拟患者和 LLM 评委。里赫曼大学硕士论文。
- 来源文件:
- [GitHub Search]