answer-me-with-html: When the Model Writes Less and the Page Does More

A skill that turns LLM answers into polished HTML explainers by splitting content from presentation, cutting token bloat, and extending the same draft into narrated video.

7-8 min read • View on GitHub • More from QingYunA

A wide editorial scene shows a cluttered stack of handwritten Markdown pages on one side and a disciplined machine on the other, producing a single clean page with sections and diagrams. It explains the project's main idea: the model writes the draft, while deterministic tools handle layout and presentation.
The repository's thesis is a workflow split. Let the model draft content, then let code do the parts that are easier to standardize.
Key Takeaways

Why long answers fail

The problem this repo attacks is not model intelligence. It is reader throughput. A better answer still fails if it arrives as a wall of text that takes too long to scan, too long to trust, and too long to turn into something useful.

That is why answer-me-with-html matters. It treats presentation as part of the answer, but not as part of the model's job.

LLMs are getting faster and cheaper, but human reading speed is a constant. We need higher-bandwidth ways to consume their output than just scrolling text.

Andrej Karpathy, Notable AI Researcher (Influencer) · Andrej Karpathy on X

Separation of labor is the whole trick

The central move is simple and slightly radical. The model writes a compact Markdown draft, then a deterministic CLI turns that draft into a finished page with layout, styling, and visuals.

That inversion matters because it removes the hardest low-value work from the model. HTML, CSS, and SVG are great for machines to render and annoying for models to invent from scratch.

One draft enters. A deterministic CLI decides how it should appear. That is the design boundary the repo depends on.

A close-up mechanical press takes one Markdown draft and splits it into three finished outputs: an HTML page, a diagram sheet, and a video waveform with a playhead. It explains that the same lightweight content source can become multiple formats without asking the model to recompose each one.
The same draft can become a page, a diagram, or a narrated video. The repo is not just a formatter. It is a fan-out pipeline.

What the skill actually does

User-facing, the repo tries to make a better default response for explain-heavy questions. Instead of dumping a long answer into the chat window, it can produce a clean one-page explainer with sections, diagrams, and tighter visual hierarchy.

That makes it useful for architecture writeups, product explanations, and any answer where structure matters as much as prose. The point is not that every answer becomes an elaborate artifact. The point is that the model can choose a richer format when the question deserves it.

# Draft

## Problem
Readers cannot efficiently consume a wall of text.

## Approach
Write the content in Markdown, then let a CLI render HTML, diagrams, and video assets.

## Outcome
A compact draft becomes a polished explainer page.

How the pipeline works

The mechanics are deliberately boring in the best way. Draft in Markdown. Send that draft through the bundled CLI. Render to HTML, CSS, and SVG with deterministic code instead of model-generated boilerplate.

That design also makes the output more portable. The same authoring path can support a web page, a diagram, and a narrated video stack without asking the model to rewrite itself for each target.

Token savings are the business case

The practical argument is cost and latency. If the model does not have to write full HTML, then it produces fewer tokens, finishes sooner, and wastes less attention on layout details that code can handle better.

ApproachWho handles layoutToken costDeterminismOutput breadth
Raw HTML generationThe modelHighLowPage only
Claude Artifacts style renderingPlatform UIMediumMediumDepends on host
Mermaid-only workflowThe model plus a diagram rendererLow to mediumMediumDiagrams only
answer-me-with-htmlDeterministic CLILowHighPage, diagrams, and narrated video

That is why the repo feels less like a generator and more like an execution plan. It shifts effort from expensive text synthesis to predictable composition.

Why the video feature matters

The video output is not a gimmick. It extends the same source draft into another consumption mode, which is exactly what the project is optimizing for: clearer explanation with less reader effort.

Once the content is separated from presentation, narration becomes another deterministic downstream product. That broadens the repo from HTML formatting into a general explainer pipeline.

What this suggests about agent interfaces

The deeper implication is that the next useful agent interface may not be a smarter prompt. It may be a better division of labor between model and tools.

Let the model reason and draft. Let code own structure, styling, and repeated presentation rules. That is the cleanest way to get richer answers without paying to regenerate the same scaffolding every time.


Sources