Nathan Lambert launches free RLHF and post-training course
TL;DR
- Nathan Lambert has released a free companion course to his RLHF Book, opening with a welcome video and Lectures 1 through 4.
- The core series is planned as eight lectures mapped chapter-by-chapter to the book, plus Q&A sessions and a prerequisite ML foundations video.
- The course is aimed at early AI PhD or master's students but is designed to be accessible without prior RL or language modeling background.
Nathan Lambert has turned his RLHF Book into a free video course, and the shape of it is more interesting than another one-off lecture drop. The launch announcement, posted on X and pointed at from the course page on rlhfbook.com, bundles a welcome video with the first four lectures on YouTube: an overview of RLHF and post-training, then IFT, reward models and rejection sampling, then RL math, then RL implementation.
The structure is what makes this useful. The course overview lays out a core series of eight lectures mapped chapter by chapter to the book, plus Q&A sessions and a prerequisite ML foundations video for people who need a math refresher. Later lectures are planned to cover the rise of reasoning models, direct preference optimization, synthetic data and modern post-training methods, and preference data itself. Each lecture comes with slides, PDFs and source markdown alongside the video, which is unusually complete packaging for a free course in this area.
The stated audience is early AI PhD or master's students, but Lambert explicitly designs it to be accessible to anyone willing to put in the work, with no prior RL or language modeling background assumed. For applied teams that keep hiring people into post-training roles without an obvious onboarding path, that framing matters more than the raw content, because it gives you a single artifact to point new engineers at.
The honest caveat is that the announcement does not commit to a release cadence for the remaining lectures, and post-training is a fast-moving area where a chapter on DPO variants or reasoning recipes can date within months. What the reporting does not give you is any signal on how the material will be maintained, whether Q&A will be live or asynchronous, or how guest lectures will be selected. Treat the four released videos as the concrete thing on offer, and the rest of the roadmap as intent.
The upside, if the series lands as planned, is a shared reference curriculum for a subfield that has mostly been taught through scattered blog posts and paper reading groups. That is worth watching whether you are a student, a hiring manager, or someone writing your own internal training track.
Shared on Bluesky by 1 AI expert
Originally reported by youtube.com
Read the original article →Original headline: Welcome to The RLHF Book & Post-Training Course