My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.
Thinking Machines just released with a ~1T param, 41B active, apache-2 model Benchmarks are a clear step up from Nemotron Ultra (55B active), new best American model, and omni input. A bit behind GLM 5.2 on agentic benches, and Kimi K 2.6 on multi modal Super exciting!!
I think what is pretty clear is that the Chinese labs are far more capital efficient. In a world where scaling labs are intelligence is proportional to effective capital (buys compute, data, & talent) that may be the greatest strength your AI industry could ever have.
Anthropic's political pressure on distillation is regulatory capture and most of the employees are blind to it under their veil of safety. Or their paycheck helped them buy into safety, is only human nature, I don't even fault them that much.
Being out of SF has lowered my information proximity but with the big upside of giving me space to cultivate my own beliefs and values around ai. We need more people zagging in AI, the monoculture just helps the incumbents win at this point.
Kimi K3 with more likes than downloads on HuggingFace is definitely showing us a glimpse of the future on open models. It's way less about individual access, and more of a distributed platform layer for companies.
Making talks with AI agents is awesome. I just told Fable to make a slide with real data on the KL distance from one of our reference Olmo 2 models and it made this with the wandb api (I edited text slightly).
Claude Fable is another big step in being able to make nice lectures based on existing educational content. Much better than Opus. GPT 5.6 is still very far off here. Is a good example of where Claude Code being a bit easier to work across different knowledge work tasks.
It's been a great effort by the early and growing American open-model labs since last June to put the US much more back on the map. We were getting totally owned last June. Nvidia, Ai2, Arcee, Gemma, GPT-OSS and a few others will be seen as saving American open AI.
The real comparison to Moore's law for AI isn't scaling laws, but rather the intelligence efficiency that we gain year-over-year using models to get better at training models.