Scott Geng on turning academic DPO into Olmo 3 practice
TL;DR
- Conversation 2 of Nathan Lambert's RLHF & Post-Training Course features Scott Geng discussing DPO implementation for Olmo 3.
- Scott Geng is a University of Washington PhD student advised by Pang Wei Koh and is listed as part of the Olmo Team.
- The framing is that 'building models is much more than the idea, it's making it fit in a more complex puzzle.'
A conversation on YouTube inside Nathan Lambert's RLHF & Post-Training Course does the thing most alignment content skips: it sits with a working researcher, Scott Geng, and walks through what happens when a well-studied technique like Direct Preference Optimization has to survive contact with a real near-frontier model build, in this case Olmo 3. Geng is a University of Washington PhD student advised by Pang Wei Koh and is listed as part of the Olmo Team, so he sits on both sides of the seam Lambert wants to examine.
The framing the course page uses is direct: 'building models is much more than the idea, it's making it fit in a more complex puzzle.' That is the whole segment. DPO has an entire lecture of its own earlier in Lambert's course, covering its 'derivation, variants, and practical application', and yet, per the course page, this conversation is about 'what it takes to land a well-grounded academic result into a near-frontier model.' That is a different kind of work than writing the paper.
Why this matters if you are not training models yourself: open post-training has plenty of papers and comparatively few first-person accounts of what actually broke on the way to a shipped model. Lambert is separately building a 'How We Built a Leading Reasoning Model (Olmo 3)' arc around this same release, so this conversation slots into a broader open post-mortem rather than a standalone chat.
The honest caveat is that the source here is a conversation attached to a course, not a technical paper. Expect intuitions and framing rather than ablations, hyperparameter tables, or head-to-head comparisons against newer RL-style alignment methods. The course page does not tell you which DPO variant made it into Olmo 3, what data mixes were used, or where the academic recipe had to be bent, so treat the specifics as reported, not settled.
For researchers who want their work to actually ship, and for small teams trying to close the gap on more closed labs, the useful thing on offer is the narration of the messy middle, the part that usually stays inside a lab.
Shared on Bluesky by 1 AI expert
Originally reported by youtube.com
Read the original article →Original headline: From Academic Research to a Frontier LLM: A Case Study in DPO | RLHF Book Course, Conversation 2