Back to Top

OUI-1: The Generative UI Model That Taught Itself

Most AI models generate text. In contrast, OUI-1 generative UI model from Thesys, generates something different: entire user interfaces. Give it a plain-language brief, and it writes back a working screen — a dashboard, a form, a card layout — as structured code, ready to render in React, Vue, or Svelte. Thesys calls it the first open-weight model built specifically for generative UI, and its training story is as interesting as the model itself.

OUI-1 generative UI

Image source: OpenUI — Introducing OUI-1

What Is OUI-1 generative UI?

OUI-1 is a fine-tune of Google’s DiffusionGemma 26B-A4B, a 26-billion-parameter diffusion language model that only activates about 4 billion parameters per token. Specifically, Thesys built it to write OpenUI Lang, the declarative language behind their OpenUI framework. Released on September 8, 2026, under Apache 2.0, it runs on a single consumer GPU — Thesys, for instance, specifically points to an RTX 5090 at FP8 precision.

Instead of one general-purpose model trying to do everything, OUI-1 does one job well. In fact, as the Hugging Face model card puts it plainly: “Not a general chat model.” Rather, it exists purely to turn a brief and a component library into a valid, renderable screen.

Why Thesys Built OUI-1 generative UI

Agent-driven interfaces — screens an AI assembles on the fly, rather than ones a developer hand-codes in advance — depend on three things happening at once. First, generation has to finish in under a second. Additionally, the result has to be reliable enough to actually ship as software. And finally, the model doing the generating has to be small enough to run locally, not in a data center.

Thesys had already tested this idea with AppLess, a “no-app phone” demo where every screen gets generated on demand instead of loaded from an installed app. Early versions ran on Gemma 4 via Cerebras hardware — fast, but tied to specialized cloud infrastructure. Consequently, moving that experience onto an actual device meant finding a model that kept the speed without the specialized hardware.

How OUI-1 generative UI Generates a Screen

DiffusionGemma gave Thesys the speed profile they needed. Specifically, rather than writing tokens one at a time like a typical chat model, it denoises a 256-token block all at once, using bidirectional attention to commit each token as soon as it becomes confident. As a result, Google reports over 1,000 tokens per second on a single H100, and over 700 on a consumer RTX 5090.

However, speed alone doesn’t make working software. The base DiffusionGemma model, for example, scored only 13.0% on Thesys’s own Generative UI Benchmark — fast, but unreliable. OpenUI Lang’s own parser made the failure modes easy to name: schema errors (an invented component, a wrong enum value) and wiring errors (a section defined but never attached to the screen’s root). Therefore, closing that gap, without losing the speed, became the actual engineering problem.

The Training Story: When Fine-Tuning Made Things Worse

Here’s where it gets genuinely interesting. Thesys’s first attempt — a straightforward supervised fine-tune on roughly 700 hand-written OpenUI Lang examples — backfired in two ways at once.

First, the model’s two error types moved like a see-saw. One training run would reduce wiring errors while increasing schema errors; then, the next run would reverse the trade rather than fixing both together. Second, and less expected, the model got slower. Specifically, generation time on light briefs rose from 1.6 seconds to 4.3 seconds, because the fine-tuned model wrote longer, more specific outputs and needed roughly twice as many denoising steps to commit each token.

The breakthrough, however, came from a simple realization: OpenUI Lang has a verifiable reward. The parser can tell, mechanically, whether a generated screen is structurally valid, and it can point to the exact defect when it isn’t. As a result, that turned the model into its own teacher. Thesys calls the approach rejection-sampled self-training with repair:

  1. The model generates a batch of OpenUI Lang programs.
  2. The parser keeps the ones that pass cleanly.
  3. Near-misses go through a targeted repair pass — fixing only the reported defect, never rewriting freely.
  4. A judge checks whether each surviving program actually matches its original brief.
  5. The survivors become the training set for the next round.

This loop, repeated across 500-step training runs on a single A100, is what actually closed the gap. As a result, speed came back — down to 1.9 seconds per output — and, crucially, the see-saw stopped. Consequently, schema errors and wiring errors dropped together for the first time, instead of trading off against each other.

Results

After extending the same recipe across 27 different component libraries, the final model — OUI-1 — reached 71.7% on the Generative UI Benchmark, up from DiffusionGemma’s 13.0%. In other words, that’s a 5.5x improvement.

ModelActive ParamsGenerative UI Score
Qwen3.8 27B (dense)27B78.8%
OUI-14B71.7%
Qwen3.6 27B (dense)27B68.5%
Qwen3.6 35B-A3B3B61.4%
Gemma 4 31B (dense)31B46.7%
Phi-4 14B (dense)14B44.0%
Gemma 4 26B-A4B4B29.9%
DiffusionGemma (base)4B13.0%

Notably, no other model at 4B active parameters or below scored anywhere close. The only model that beat OUI-1 outright, Qwen3.8 27B, is a dense model using 27 billion parameters on every single token — nearly seven times OUI-1’s active parameter count.

Furthermore, the gain held up outside the benchmark, too. On AppLess’s own component library — a different library from training, with 60 prompts the model had never seen — OUI-1 produced 55 valid outputs against DiffusionGemma’s 23.

From Research to Real Product: AppLess

OUI-1 isn’t just a benchmark exercise. As of mid-September 2026, for example, Thesys moved AppLess itself onto OUI-1, replacing the Cerebras-hosted Gemma 4 setup entirely. The pitch is a phone with no installed apps: every screen — checking weather, tracking a package, viewing a bank balance — gets generated live from a natural-language request instead of opening a pre-built app.

On top of that, Thesys open-sourced both the AppLess client and the OpenUI framework under MIT, so developers can fork the demo and build their own agent-driven interface experiments on top of it.

Running OUI-1 generative UI Yourself

For developers who want to self-host, OUI-1 is designed to fit realistic hardware budgets:

  • FP8 via vLLM 0.24+: roughly 25.8 GiB of GPU memory — light enough to share a GPU with other workloads.
  • bf16 via Transformers: about 52 GiB, fitting comfortably on a single A100 80GB or H100.
  • Tool calling: works out of the box through Gemma 4’s native tool-call format, letting OUI-1 pull real data (a weather API, a stock price) before writing the screen that displays it.

One quirk worth knowing: temperature and seed get accepted by the API but silently ignored, since the diffusion sampler runs its own fixed, entropy-bound schedule rather than standard autoregressive sampling. As a result, two identical requests can return differently worded — though structurally similar — screens.

Broader Implications

OUI-1’s approach points to a few larger shifts worth watching:

  • Verifiable rewards make self-distillation practical. Any domain with a mechanical way to check correctness — a parser, a compiler, a schema validator — can potentially apply the same rejection-sampled self-training loop Thesys used here.
  • Small, specialized models can beat much larger general ones. At 4B active parameters, OUI-1 outperforms dense models seven times its size on the one task it was built for.
  • Generative UI moves from server to device. Thesys explicitly frames OUI-1 as a step toward interfaces generated locally, in under a second, without a round trip to specialized cloud hardware.

Conclusion

OUI-1 makes a strong case that specialization, not scale, was the missing ingredient for reliable generative UI. Specifically, by turning a parser’s pass/fail signal into a training reward, Thesys took a 4-billion-active-parameter model from 13% to 71.7% accuracy — beating every other model near its size class, and closing in on models many times larger. Ultimately, for developers building agent-driven interfaces that need to run fast, reliably, and on hardware they actually control, OUI-1 is a genuinely new option, not just an incremental one.

Ready to bring intelligent AI capabilities closer to your users? Your journey starts at Webkul.

. . .

Leave a Comment

Your email address will not be published. Required fields are marked*


Be the first to comment.

Back to Top

Message Sent!

If you have more details or questions, you can reply to the received confirmation email.

Back to Home