← All posts / Models

Runway's Solaris Renders Software Itself: Inside the First 'Interface World Model'

Runway's Solaris generates interactive app and website interfaces frame by frame with no code, pairing a world model renderer with an LLM reasoner — and beat Claude-coded interfaces 61% to 24% on instruction-following in a 250-person study.

Runway's Solaris Renders Software Itself: Inside the First 'Interface World Model'

Every piece of software ever shipped has hidden a translation step: a designer’s visual idea must be converted into code — an intermediate representation — before it can do anything. On August 31, Runway published research on a model that deletes that step entirely. Solaris, the first of a new family the company calls Interface World Models, generates interactive app and website interfaces in real time, frame by frame, with no underlying code at all.

The framing question Runway poses is deceptively simple: what happens when an operating system generates apps and websites as you use them? Every OS from early terminals to macOS dictates what is rendered on screen and what happens when a person acts on it. Applications get built on top and stay frozen until a developer pushes an update. Solaris instead renders that layer directly — every frame is synthesized as you interact, so the interface responds continuously to what you do.

Why the translation step matters

Runway’s argument is that today’s software pipeline is a “lossy compression of the space of possible interactions, frozen before any user arrives.” Pixel-perfect mockups and image models can already generate entire screens nearly indistinguishable from finished products — but images don’t run. Every behavior still has to be explicitly defined and implemented ahead of time, and once a design is reduced to code, the interface responds quickly only by giving up much of the original design’s richness.

The company measured this loss directly. Using a reconstruction benchmark built on structural similarity (SSIM) and DINOv3 region-matching, Runway showed that even state-of-the-art multimodal language models degrade sharply when reconstructing interfaces as visual complexity increases — rich visual detail simply cannot be represented accurately in language. Solaris, by operating directly on the visual interface, preserves the complete visual and semantic state from the very first frame.

How it works

Solaris builds on Runway’s Gen-4.5 video generation model, adapted along the path opened by GWM-1, its general world model. Three design pieces make it work:

Learning interaction. User input — clicks, drags, typing — is treated as conditioning for the next frame, exactly the way text and image prompts are. Because the model only ever sees interactions that have already happened (never future ones), it learns the causal relationship between user actions and visual outcomes, without any interaction being explicitly programmed.

Running in real time. Standard video diffusion refines an entire clip over dozens of denoising steps — fine for content creation, hopeless for an interface. Runway converted Solaris into a real-time engine in three stages: it taught the model to generate frames autoregressively (each frame depending only on what came before), distilled the many-step denoising process down to just a few steps, and then trained the fast model on its own outputs so visual quality stays stable over long interactions. The same work, Runway says, made it orders of magnitude cheaper to run than a standard video diffusion model.

Reasoning and rendering, separated. This is the architectural heart of the system. A language model decides what the application should do next — interpreting user requests, deciding whether an interaction should modify the current scene or transition to a new one, defining the behaviors that make the world feel alive, and producing the prompts that guide rendering. The world model then generates how that behavior appears, at interactive speeds. One model reasons; the other renders.

The interaction model itself is striking: there are no predefined screens and no templates. You provide a starting state — a brand environment, a product scene — and the model streams frames. Text prompts define what clicks and drags mean in a particular scene. Runway calls this “redefining the mouse”: click on a cat, and your next clicks apply its fur color and texture to whatever you touch; click on a painting, and you might begin drawing with that style.

The numbers: Solaris vs. coded interfaces

The most concrete result is a head-to-head against a state-of-the-art language model — Claude Opus 5 — generating coded interfaces. Both systems started from the same image and received the same interaction requests. Runway then ran a user study with 250 participants across 30 interaction examples, collecting nearly 7,500 pairwise judgments.

Participants preferred Solaris on both measures:

  • Instruction-following: Solaris preferred in 61% of comparisons vs. 24% for the coded result (13% rated equivalent)
  • Natural behavior within the scene: Solaris preferred in 71% vs. 21% (6% equivalent)

The gap on naturalness is the telling one. A coded interface can often reproduce a requested change, but it treats each UI action as an isolated update. Because the world model already understands how objects, materials and environments behave, Solaris generates interactions that stay coherent with the whole scene — the couch changes color and the lighting responds.

A new way to train agents

Buried in the announcement is arguably its most consequential implication: agent training. Even the best LLMs today struggle with basic computer-use tasks like booking a hotel or ordering groceries, because text-based models trained on coded interfaces learn the specific layout they were trained on and fail to adapt to a slightly different one. By collapsing the space between action and response, Solaris lets agents train against interfaces that are constantly changing and layouts that may never have existed — a far richer curriculum than static web scrapes.

What it can’t do yet

Runway is unusually candid about the gaps. Stable, legible text remains one of the hardest problems in video generation — and interfaces depend on it more than almost any other visual domain; the suggested interim path is hybrid systems where image models render text-heavy views during pauses. Trust is unresolved: for instructional or commercial experiences, a convincing wrong answer is worse than no answer, so Solaris stays anchored to real product imagery while grounding on verified context remains an active research focus. Long-session coherence — maintaining visual and semantic consistency over extended open-ended use — is still open. And accessibility: a generated interface still has to work with screen readers and accessibility APIs, or flexibility comes at the expense of usability.

Solaris currently targets 720p quality with three engineering focuses: real-time interaction (the threshold where interactions stop feeling interactive sits around half a second of delay), coherence over an entire session, and visual quality that holds.

If interfaces become generated

The endpoint Runway sketches is an operating layer where the app stops being the unit you interact with. Today, getting something done means opening the pre-built app made for it. If the OS can generate useful interfaces on demand, the fixed catalog of apps starts to look like a legacy constraint — what you need simply shows up, customized to you. A storefront becomes a generated environment that preserves brand identity while reshaping around each individual; tutorials render the next step in your own context and recover naturally when you go off script.

Runway expects the trajectory to mirror image and video generation: each model generation faster, more coherent, more controllable. The company is working with key partners to launch Solaris publicly, with early access available via request form — and the challenges that once made generated interfaces seem impractical now look, in its words, “increasingly like solvable engineering problems.”

Whether Solaris itself becomes the product or merely the proof of concept, the conceptual move is significant: it treats software as something a model can be, rather than something a model helps you write. The gap between “an AI that codes your app” and “an AI that is your app” just got its first serious benchmark.