Ai2

Why they matter

Tracked through public AI activity and peer connections inside the directory.

AI signals
11
past 30d
Sources
5
distinct domains
Discusiones
0
past 30d
Latest signal
3d ago
View every signal from Ai2 →
Breakthrough AI to solve the world's biggest problems. › Join us: http://allenai.org/careers › Get our newsletter: https://share.hsforms.com/1uJkWs5aDRHWhiky3aHooIg3ioxm

Articles & links

Our fully open releases give researchers the data, code, checkpoints, and methods they need to inspect claims, reproduce findings, and advance new science. Read more about why that’s so important to us. ⬇️ allenai.org/blog/who-get...

Who gets to understand AI? | Ai2 allenai.org
AI Weekly's analysis
  • Ai2 argues meaningful AI transparency requires not just open weights but training data, code, methods, checkpoints, evaluations, and documentation.
  • Post cites three studies enabled by open Olmo releases, covering clinical demographic bias, benchmark inflation, and how models reason about drug names.
  • Without that access, Ai2 warns, technical direction of the field risks becoming concentrated inside a small number of companies.
Read full analysis →
View on Bluesky · ♥ 4 ↻ 1 ↩ 0 · 2 from the directory shared this · 45d ago

Our coupled emulator could help screen ideas before running E3SM itself & enable large ensembles that cheaply sample climate variability. Next, we’ll explore historical runs where greenhouse gases + other climate drivers change over time. 🤗 Download: buff.ly/FGKIHZL

allenai/SamudrACE-E3SMv3 · Hugging Face huggingface.co
View on Bluesky · ♥ 2 ↻ 0 ↩ 0 · 3d ago

BenchMIRT gives researchers a clearer view of what benchmarks really measure—and could help build evals that are smaller, more focused, & easier to interpret. We’re releasing it openly so others can build on it: 💻 buff.ly/FKrFkrC 📄 buff.ly/kMz7Gs8

GitHub - allenai/BenchMIRT github.com
View on Bluesky · ♥ 3 ↻ 1 ↩ 0 · 5d ago

These ideas came out of presentations and a panel at our Aug. 27 event with @healthyalgo.bsky.social, @hoifungpoon.bsky.social, Kyle Travaglini (@alleninstitute.org), Stephen Salerno (@washu.edu), Bodhisattwa Majumder, Sasha Stanton, & Kelly Paulson. Learn more: buff.ly/iLeuFa8

The hard parts of AI-assisted science | Ai2 allenai.org
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 6d ago

Try olmOCR 2 in the Ai2 Playground, check out our blog for more info, & download the weights and data from Hugging Face: ▶️ Playground: buff.ly/pJhoWXN 📝 Blog: buff.ly/8uUabID 🤗 Model & data: buff.ly/EUYHIuN

playground.allenai.org
View on Bluesky · ♥ 1 ↻ 1 ↩ 0 · 59d ago

Recent commentary

LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work. So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560 We made ModSleuth to trace this. 🧵

View on Bluesky · ♥ 53 ↻ 11 ↩ 1 · 88d ago

Today we're introducing a preview of TutorMoments, a framework that measures whether AI tutors can make one of the hardest calls in teaching: when to step in and help a student, & when to hold back and let them do the heavy thinking. 🧵

View on Bluesky · ♥ 51 ↻ 6 ↩ 2 · 31d ago

We built SciArena to test how well AI models handle scientific literature questions, as judged by researchers. It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes. Here's what they told us. 🧵

View on Bluesky · ♥ 9 ↻ 5 ↩ 2 · 54d ago

AI image generators don't "draw"—they follow a compass: the score function, which points toward more probable images. The same compass drives Bayesian sampling and plasma physics. We built DiScoFormer to estimate the score far better when data gets complex. 🧵

View on Bluesky · ♥ 14 ↻ 1 ↩ 2 · 70d ago

When a model writes, where do its words come from? Are they new, or do they match exactly with language it saw in training? An AI-writing detector can't tell you. @tuhinchakr.bsky.social's group at Stony Brook has been dissecting AI-generated prose with our infini-gram engine. 🧵

View on Bluesky · ♥ 13 ↻ 1 ↩ 1 · 37d ago

Two updates to Asta, our ecosystem of AI agents for science: a one-click handoff from AutoDiscovery to Asta’s data analysis tools, & paper search that evaluates its own results + searches again when they fall short. 🧵

View on Bluesky · ♥ 10 ↻ 2 ↩ 1 · 48d ago

Building an LLM means evaluating it over & over as it changes. Tweak a hyperparameter or scale the model up, & every new checkpoint sends you back through the same benchmarking loop. We're releasing olmo-eval, a workbench built for this kind of iterative model development. 🧵

View on Bluesky · ♥ 8 ↻ 3 ↩ 1 · 87d ago

As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.

View on Bluesky · ♥ 9 ↻ 1 ↩ 1 · 45d ago

What does it actually take to build cutting-edge AI systems? On July 30 during #SeattleTechWeek, the researchers behind Ai2's open models sit down to talk through the deep technical work behind them. 🧵

View on Bluesky · ♥ 9 ↻ 0 ↩ 2 · 52d ago

What can you build with a fully open robotics model in a weekend? 🤖 Robotics engineer @0xbinh.bsky.social used MolmoAct 2, our open vision-language-action model, in the voice-controlled robot that won @southparkcommons.bsky.social's AI hackathon. Watch our interview with him ↓ 🎥

View on Bluesky · ♥ 13 ↻ 0 ↩ 0 · 61d ago

In Ai2's orbit

Center = Ai2. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Ai2? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/ai2-bsky-social)