answer-me-with-html: When the Model Writes Less and the Page Does More
A skill that turns LLM answers into polished HTML explainers by splitting content from presentation, cutting token bloat, and extending the same draft into narrated video.
- answer-me-with-html treats layout as a deterministic tool problem, which keeps the model focused on drafting instead of typesetting.
- The repo's real win is not prettier HTML, but lower token cost and faster turns on the same explanation.
- Its pipeline can fan one short Markdown draft into HTML, diagrams, and narrated video without asking the model to do all three jobs.
- The project points toward agent interfaces where composition lives in code and language models stay on the content path.
Why long answers fail
The problem this repo attacks is not model intelligence. It is reader throughput. A better answer still fails if it arrives as a wall of text that takes too long to scan, too long to trust, and too long to turn into something useful.
That is why answer-me-with-html matters. It treats presentation as part of the answer, but not as part of the model's job.
LLMs are getting faster and cheaper, but human reading speed is a constant. We need higher-bandwidth ways to consume their output than just scrolling text.
Separation of labor is the whole trick
The central move is simple and slightly radical. The model writes a compact Markdown draft, then a deterministic CLI turns that draft into a finished page with layout, styling, and visuals.
That inversion matters because it removes the hardest low-value work from the model. HTML, CSS, and SVG are great for machines to render and annoying for models to invent from scratch.
What the skill actually does
User-facing, the repo tries to make a better default response for explain-heavy questions. Instead of dumping a long answer into the chat window, it can produce a clean one-page explainer with sections, diagrams, and tighter visual hierarchy.
That makes it useful for architecture writeups, product explanations, and any answer where structure matters as much as prose. The point is not that every answer becomes an elaborate artifact. The point is that the model can choose a richer format when the question deserves it.
# Draft
## Problem
Readers cannot efficiently consume a wall of text.
## Approach
Write the content in Markdown, then let a CLI render HTML, diagrams, and video assets.
## Outcome
A compact draft becomes a polished explainer page.
How the pipeline works
The mechanics are deliberately boring in the best way. Draft in Markdown. Send that draft through the bundled CLI. Render to HTML, CSS, and SVG with deterministic code instead of model-generated boilerplate.
That design also makes the output more portable. The same authoring path can support a web page, a diagram, and a narrated video stack without asking the model to rewrite itself for each target.
Token savings are the business case
The practical argument is cost and latency. If the model does not have to write full HTML, then it produces fewer tokens, finishes sooner, and wastes less attention on layout details that code can handle better.
| Approach | Who handles layout | Token cost | Determinism | Output breadth |
|---|---|---|---|---|
| Raw HTML generation | The model | High | Low | Page only |
| Claude Artifacts style rendering | Platform UI | Medium | Medium | Depends on host |
| Mermaid-only workflow | The model plus a diagram renderer | Low to medium | Medium | Diagrams only |
| answer-me-with-html | Deterministic CLI | Low | High | Page, diagrams, and narrated video |
That is why the repo feels less like a generator and more like an execution plan. It shifts effort from expensive text synthesis to predictable composition.
Why the video feature matters
The video output is not a gimmick. It extends the same source draft into another consumption mode, which is exactly what the project is optimizing for: clearer explanation with less reader effort.
Once the content is separated from presentation, narration becomes another deterministic downstream product. That broadens the repo from HTML formatting into a general explainer pipeline.
What this suggests about agent interfaces
The deeper implication is that the next useful agent interface may not be a smarter prompt. It may be a better division of labor between model and tools.
Let the model reason and draft. Let code own structure, styling, and repeated presentation rules. That is the cleanest way to get richer answers without paying to regenerate the same scaffolding every time.