roerohan/slate: Teaching AI to See Its Own UI Bugs

How a render-and-compare reinforcement learning loop uses computer vision to train Vision-Language Models for pixel-perfect Tailwind generation.

7 min read · roerohan/slate

An illustration of a robot eye looking through a magnifying glass at a web page layout, with red lines highlighting differences from a blueprint.
Slate gives Vision-Language Models a visual feedback loop, allowing them to "see" the UI they generate and correct their own mistakes.
Key Takeaways

The Blindness of Code Generation

Before diving in, a quick disambiguation: this article covers roerohan/slate, an experimental reinforcement learning framework. It is entirely unrelated to the massive React rich-text editor, ianstormtaylor/slate, despite the name collision. Our focus is on the former—a project aimed at solving a fundamental flaw in how AI writes user interfaces.

When a standard Large Language Model (LLM) generates a UI component, it is flying blind. It predicts text tokens (HTML and CSS classes) based on a prompt, but it never actually sees the result. It treats complex spatial layouts, padding, and alignment as flat text strings. This often leads to code that compiles but looks visually broken, because the model lacks spatial awareness of the rendered DOM.

The Render-and-Compare Engine

Slate tackles this blindness by giving the AI "eyes." It implements a render-and-compare pipeline that fundamentally changes the reward signal used during training. Instead of relying solely on text-based next-token prediction loss, Slate compiles the generated HTML, takes a screenshot, and uses computer vision to calculate a mathematical reward.

The visual reward cycle converts generated code into a mathematical score based on pixel-perfect similarity.

The core of this logic lives in rl/reward.py. Slate uses the Structural Similarity Index Measure (SSIM) to compare the model's rendered screenshot against a target reference image. This visual similarity is then converted into a scalar reward signal.

However, relying purely on visual similarity opens the door to "reward hacking." If the target reference is a minimalist, mostly white design, a model could achieve a high SSIM score simply by generating a completely blank white page. To prevent this, Slate implements a "Content Gate." If the text or color similarity falls to zero, the SSIM score is heavily penalized, ensuring the model actually generates the required content, not just a visually similar background.

An illustration of a heavy bank vault door locked by a mechanism resembling a scale, balancing a feather against a dense block of typography.
The Content Gate mechanism prevents models from cheating the SSIM metric by generating blank pages.

GRPO for Visual Layouts

To optimize the model based on this visual reward, Slate utilizes Group Relative Policy Optimization (GRPO), the same algorithm popularized by DeepSeek-R1. This approach eliminates the need for a massive, separate "critic" model during training.

Instead, for every screenshot in the training data, the model generates a group of different completions (usually four). Slate renders all four attempts and calculates the advantage by comparing each completion's reward to the mathematical average of the group. This simulates an evolutionary design process, where the model learns to favor the layout that performs best relative to its peers.

An illustration of an art studio where four painters work on identical canvases, while a master instructor measures one painting against a transparent reference sketch.
GRPO evaluates multiple attempts against a group mean, eliminating the need for a separate critic model.

The Red-Line Diff and Self-Correction

The most sophisticated aspect of Slate is its multi-turn training logic, found in multi_turn.py. It doesn't just teach the model to write code on the first try; it teaches the model to debug.

When the model makes an attempt, Slate uses NumPy to generate a diff image highlighting pixel-wise discrepancies in stark red. This image, along with a feedback prompt, is fed back into the model's context window. The AI is forced to look at its own visual failures and issue a correction, effectively simulating a senior developer reviewing a pull request with visual red-lines.

FeatureStandard LLM UI GenerationSlate Visual RL Pipeline
Feedback LoopText-based compiler errorsPixel-perfect visual SSIM diffs
OptimizationNext-token prediction lossGroup Relative Policy Optimization (GRPO)
CorrectionZero-shot guessingMulti-turn visual red-line debugging
Hardware ProfileMassive cluster fine-tuningLoRA Rank 32 via Tinker SDK (consumer viable)

Orchestrating the Edge

Running this pipeline requires robust data engineering. The src/ directory contains a TypeScript orchestration layer that prepares the HuggingFace WebSight dataset for training.

This pipeline strictly enforces a 4096-character token budget and utilizes Cloudflare Browser Rendering to test every HTML snippet before it enters the training set. If a snippet produces a blank or broken visual, it is discarded. This ensures that the Python RL suite only spends compute on viable, renderable examples.