# Nathan Lambert 发布课程答疑视频，详细讲解 on-policy distillation 和 reward model 推导中的常见错误与修正，并提供额外资源帮助深入理解

- 来源：Nathan Lambert
- 发布时间：2026-07-02 04:24
- AIWatch 分数：63
- AIWatch 标记：未精选
- AIWatch 链接：https://aiwatch.icu/events/evt_01kx4dx2jtt3tke99rf48d1m12
- 原文链接：https://x.com/natolambert/status/2072416196142789095

## 精选理由

常规快讯，保留列表

## AI 摘要

Nathan Lambert 发布课程答疑视频，详细讲解 on-policy distillation 和 reward model 推导中的常见错误与修正，并提供额外资源帮助深入理解。

## 正文

Nathan Lambert 发布课程答疑视频，详细讲解 on-policy distillation 和 reward model 推导中的常见错误与修正，并提供额外资源帮助深入理解。
