runway.com web signal

Runway unveils Solaris, an 'interface world model' for live UIs

TL;DR

  • Runway's Solaris generates application interfaces frame by frame, using a language model to decide behavior and its Gen-4.5 video model to render each frame.
  • In a 250-participant, 7,500-comparison study, Solaris beat coded interfaces 61% to 24% on instruction following and 71% to 21% on natural behavior.
  • Runway lists unresolved gaps in legible text, trust grounding, long-session coherence, and screen-reader support; access is limited to a partner early-access form.

Runway's new research release, Solaris, generates application and website interfaces frame by frame as a user interacts with them, replacing the code layer that normally sits between a design and a running app. In a research post, the company calls it "the first model in a new family of AI systems" it labels Interface World Models, built on top of its Gen-4.5 video model.

The pitch is that there is no intermediate representation at all. "A single world model generates every frame and every response to user input, eliminating the need for an intermediate representation," Runway writes. A paired language model decides what the interface should do next; the world model draws the result.

The evidence Runway leans on is a preference study of 250 participants across 30 interaction examples, collecting nearly 7,500 pairwise judgments. Compared against coded interfaces, Solaris was preferred 61% of the time on instruction following (versus 24% for the coded version, with 13% called equivalent) and 71% on natural behavior (versus 21%). Runway ran the study on its own model and does not publish per-example scores.

Solaris is not being released as a public product. Runway says it is "working with key partners to launch Solaris publicly" and is collecting requests through an early-access form. The post also lists four problems the model does not yet handle: legible text, grounded trust for instructional or commercial use, coherence over long sessions, and integration with screen readers and other assistive technologies. On the first, the company is blunt: "Stable, legible text remains one of the hardest problems in video generation, yet interfaces depend on it more than almost any other visual domain." It lands in a busy stretch of multimodal releases we've tracked this summer, alongside work like Lucida's system for rebuilding indoor scenes as editable assets from video.

Shared on Bluesky by 1 AI expert