Nathan Lambert 发布课程答疑视频,详细讲解 on-policy distillation 和 reward model 推导中的常见错误与修正,并提供额外资源帮助深入理解 · AIWatch