Syntax-Highlighter: The Python Project That Turns Parsing Into a Portable Highlighting Contract
A core/plugin architecture, a single span model, and a graceful no-parse fallback make this repo more interesting than its size suggests.
- Syntax-Highlighter is interesting because it treats highlighting as a contract problem, not just a parsing problem.
- The same Span model carries meaning from parser output to ANSI and HTML renderers, so the output layer stays stable.
- When parsing fails, the repo falls back to lexical highlighting instead of dropping the user into a broken experience.
- The core/plugin split keeps language support extensible without making the renderer care which language produced the spans.
The first thing worth knowing about this repo is not that it highlights code. It is that it refuses to stop highlighting when the parser breaks. That makes the project feel less like a demo and more like a small, disciplined argument about software boundaries.
The repository describes itself as SagharRamezani/Syntax-Highlighter, a modular syntax highlighter built for a university project using a Core/Plugin architecture. The public footprint is tiny, which is part of the charm: the code is doing the talking.
Why a broken file still gets colored
Most highlighters quietly assume the source is good enough to parse, or they collapse into a purely lexical path from the start. This project does something more deliberate. It keeps a no-parse fallback, so broken code still gets a useful visual treatment instead of a blank screen or a hard failure.
That is the article’s key idea in one sentence: the parser is an enrichment step, not a single point of failure. In a lot of tools, parsing is the whole game. Here, parsing is one possible source of spans.
A modular syntax highlighter built for a university project using a Core/Plugin architecture.
The abstraction that holds everything together
The project’s most important move is the Span model. Once highlighting is reduced to a list of spans with start, end, and role, the rendering layer stops caring whether those spans came from a parser, a lexer, or a recovery path.
@dataclass(frozen=True)
class Span:
start: int
end: int
role: str
# parser, fallback lexer, and both renderers
# all speak in spans
That separation is the whole design. It means the core can normalize spans, sort them, and render them without knowing anything about Haskell, Rust, or Dart.
It also means the project can keep the same conceptual contract across terminal and web output. ANSI and HTML are different surfaces, but they are fed by the same internal shape.
Inside the core/plugin split
The core/plugin split is the repo’s other clean boundary. The engine lives under src/highlighter/core/, while language packages sit under src/highlighter/languages/. That keeps the core language-agnostic and makes new languages feel like additions, not surgery.
| Dimension | Syntax-Highlighter | Typical regex-first highlighters |
|---|---|---|
| Language support | Small set of language plugins, designed to be extended | Broad catalog, usually optimized for coverage |
| Internal contract | One Span model for all output targets | Lexer output often tied more tightly to a specific renderer |
| Failure handling | Can fall back to lexical highlighting when parsing fails | Often assumes lexical highlighting is the primary mode |
| Architecture | Core plus plugins, with explicit registry boundaries | Usually flatter and more directly coupled to highlighting rules |
| Teaching value | Shows how to separate parsing, tokenization, and rendering | Shows practical highlighting, but less about pipeline design |
That table is the point. This repo is not trying to beat Pygments or Prism.js on breadth. It is demonstrating a cleaner interface between language support and output generation.
How the parser stays useful when code is messy
The parser side uses a recursive descent approach with recovery. That matters because a syntax highlighter lives in the real world, where files are often half-written, broken, or intentionally incomplete.
Instead of treating malformed input as a dead end, the parser can salvage enough structure to keep producing useful spans. And when even that is too much, the fallback path still keeps the output alive.
This is the part that feels mature. The project does not confuse correctness with usefulness. It aims for both, but it knows usefulness has to win when the input is imperfect.
Why the renderers are boring on purpose
The ANSI and HTML renderers are intentionally plain. That is a good sign. Their job is not to invent new behavior. Their job is to preserve the span contract cleanly, after normalization, without leaking rendering quirks back into the core.
def render(spans, source):
spans = _normalize_spans(spans)
# turn spans into ANSI escapes or HTML tags
return output
Normalization is where a lot of quiet engineering lives. If spans overlap or arrive out of order, the output can get ugly fast. By sorting and cleaning them first, the renderers stay predictable.
That predictability is what makes snapshot testing worthwhile. The repo is protecting rendered output as a stable interface, not as a side effect.
How it compares to the usual suspects
Compared with the big names, this project is tiny. But the comparison is useful because it clarifies what the repo is actually trying to prove.
| Project | Main strength | Main trade-off | What Syntax-Highlighter does differently |
|---|---|---|---|
| Pygments | Huge language coverage and many formats | Broadness can hide architectural ideas | Keeps the pipeline small enough to study |
| Prism.js | Lightweight browser-side highlighting | Mostly browser-centric and lexer-driven | Shows the same contract across terminal and web |
| Highlight.js | Convenient auto-detection and broad adoption | Less focused on explicit parser boundaries | Separates parsing, spans, and rendering more aggressively |
| Shiki | High-quality themed output from TextMate grammars | Integration model is different and heavier | Uses a simpler educational core with explicit fallback |
| Tree-sitter | Powerful parsing infrastructure | Not a drop-in highlighter by itself | Turns parsing into just one stage in a highlighting pipeline |
So the real contrast is not feature count. It is intent. The mature tools optimize for coverage, quality, or ecosystem fit. This repo optimizes for clarity of architecture.
What the tests are really protecting
The testing story matters because it proves the design is not just neat on paper. Snapshot tests lock down ANSI and HTML output, which means the project is defending the contract all the way to the edge of the system.
That is exactly the right thing to test when the architecture revolves around stable spans. If the contract changes, the visible output changes. The tests make that visible immediately.
For a university project, that is a strong signal. It shows a builder who understands that clean boundaries are only valuable if they stay clean over time.
The lesson
This repo’s most useful lesson is bigger than syntax highlighting. If you want a system to survive malformed input, multiple output targets, and plugin growth, the critical decision is the contract you choose early.
Here, that contract is the span. Everything else, from parser recovery to ANSI and HTML rendering, hangs off that choice. That is why a small Python project can still teach a large architectural lesson.