Runway unveils Solaris, an 'interface world model' for live UIs
TL;DR
- Solaris runs on Runway's Gen-4.5 video model targeting 720p at sub-500ms per-frame latency, with no DOM or accessibility APIs in the output.
- NVIDIA's hardware collaboration is a structural dependency for real-time generation, not background sponsorship.
- No independent benchmarks existed at announcement; AlphaSignal's 250-participant, 7,500-judgment study is the sole performance evidence.
Runway's new research release, Solaris, generates application and website interfaces frame by frame as a user interacts with them, replacing the code layer that normally sits between a design and a running app. In a research post, the company calls it "the first model in a new family of AI systems" it labels Interface World Models, built on top of its Gen-4.5 video model.
The pitch is that there is no intermediate representation at all. "A single world model generates every frame and every response to user input, eliminating the need for an intermediate representation," Runway writes. A paired language model decides what the interface should do next; the world model draws the result.
The evidence Runway leans on is a preference study of 250 participants across 30 interaction examples, collecting nearly 7,500 pairwise judgments. Compared against coded interfaces, Solaris was preferred 61% of the time on instruction following (versus 24% for the coded version, with 13% called equivalent) and 71% on natural behavior (versus 21%). Runway ran the study on its own model and does not publish per-example scores.
Solaris is not being released as a public product. Runway says it is "working with key partners to launch Solaris publicly" and is collecting requests through an early-access form. The post also lists four problems the model does not yet handle: legible text, grounded trust for instructional or commercial use, coherence over long sessions, and integration with screen readers and other assistive technologies. On the first, the company is blunt: "Stable, legible text remains one of the hardest problems in video generation, yet interfaces depend on it more than almost any other visual domain." It lands in a busy stretch of multimodal releases we've tracked this summer, alongside work like Lucida's system for rebuilding indoor scenes as editable assets from video.
What others are reporting
-
CryptoBriefing Read →
Flags absence of independent benchmarks at launch and identifies NVIDIA hardware collaboration as a structural dependency, not promotional context.
The AI handles rendering and interaction jointly, which means there are no intermediate representations to slow things down or introduce bugs.
-
The New Stack Read →
Developer-platform framing: positions Solaris as a runtime model where generation happens at interaction time, not at compile or deploy time.
Shared on Bluesky by 2 AI experts
Originally reported by runway.com
Read the original article →Original headline: Runway Unveils Solaris, Its First 'Interface World Model' That Generates UIs Frame-by-Frame in Real Time