brookings.edu web signal

Brookings pushes rapid-cycle R&D for classroom AI tools

TL;DR

  • Nearly two-thirds of teachers now report using AI in their work, but only 1 in 5 edtech products in classrooms have evidence they improve outcomes.
  • The authors argue randomized controlled trials misfit AI because developers constantly update models and users evolve how they use them.
  • They propose 'implementation R&D': staged, rapid-cycle testing that studies how a tool is used before asking whether it works.

The uncomfortable number in a new Brookings commentary on classroom AI is the gap between how fast teachers are adopting the tools and how little evidence anyone has that they work. Nearly two-thirds of teachers now report using AI in their work, the piece notes, while only 1 in 5 edtech products in classrooms have any evidence that they can improve teaching and learning outcomes.

Stacey Alicea, executive director of the Research Partnership for Professional Learning, and Meghan McCormick, senior officer of research and impact at the Overdeck Family Foundation, argue that the standard evaluation playbook does not fit what is actually being deployed. Randomized controlled trials, they write, assume a stable intervention applied consistently across contexts, and that assumption does not hold for most early-stage AI tools, because developers constantly update models and refine prompts while users evolve their interactions as they learn what works. By the time an RCT ships a verdict, the version under study is no longer the version in classrooms.

Their proposed alternative is what they call implementation R&D, a mode of research traditionally used to study early-stage products and interventions that focuses on how a tool is designed, delivered, and used in real settings before a large-scale evaluation. They lay out three principles: build evidence in stages, ask how it works before asking whether it works, and let the questions drive the method. The Brookings essay points to the Institute of Education Sciences' generative AI R&D centers, Leanlab Education, Boston University's EVAL initiative, and Teaching Lab as examples of the rapid-cycle work they have in mind.

The honest caveat is that this is a commentary, not a study. The authors are making a methodological argument, not presenting data that shows implementation R&D produces better classroom outcomes than the RCT approach it is meant to supplement. What the reporting does not give you is a sense of how districts, funders, or product teams should actually make the trade-off between speed and rigor when they choose which studies to back, or who absorbs the cost of the extra iteration cycles.

The bet worth watching is whether the evidence base for classroom AI catches up faster if researchers are willing to study moving targets on their own terms rather than waiting for the tools to sit still.

Shared on Bluesky by 2 AI experts