LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work. So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560 We made ModSleuth to trace this. 🧵
Today we're introducing a preview of TutorMoments, a framework that measures whether AI tutors can make one of the hardest calls in teaching: when to step in and help a student, & when to hold back and let them do the heavy thinking. 🧵
We built SciArena to test how well AI models handle scientific literature questions, as judged by researchers. It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes. Here's what they told us. 🧵
AI image generators don't "draw"—they follow a compass: the score function, which points toward more probable images. The same compass drives Bayesian sampling and plasma physics. We built DiScoFormer to estimate the score far better when data gets complex. 🧵
When a model writes, where do its words come from? Are they new, or do they match exactly with language it saw in training? An AI-writing detector can't tell you. @tuhinchakr.bsky.social's group at Stony Brook has been dissecting AI-generated prose with our infini-gram engine. 🧵
Two updates to Asta, our ecosystem of AI agents for science: a one-click handoff from AutoDiscovery to Asta’s data analysis tools, & paper search that evaluates its own results + searches again when they fall short. 🧵
Building an LLM means evaluating it over & over as it changes. Tweak a hyperparameter or scale the model up, & every new checkpoint sends you back through the same benchmarking loop. We're releasing olmo-eval, a workbench built for this kind of iterative model development. 🧵
As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.
What does it actually take to build cutting-edge AI systems? On July 30 during #SeattleTechWeek, the researchers behind Ai2's open models sit down to talk through the deep technical work behind them. 🧵
What can you build with a fully open robotics model in a weekend? 🤖 Robotics engineer @0xbinh.bsky.social used MolmoAct 2, our open vision-language-action model, in the voice-controlled robot that won @southparkcommons.bsky.social's AI hackathon. Watch our interview with him ↓ 🎥