Syntax-Highlighter: The Python Project That Turns Parsing Into a Portable Highlighting Contract

A core/plugin architecture, a single span model, and a graceful no-parse fallback make this repo more interesting than its size suggests.

7-9 min read • View on GitHub • More from SagharRamezani

A broken code file enters a compact parsing machine on the left, then the same colored output leaves through two different finishers on the right: a terminal printer and a web page. In the center, a stream of span tokens keeps moving even as one part of the machine jams, showing that highlighting survives parser failure.
The repo’s core idea is simple: keep one output contract intact, even when parsing falls back to lexical highlighting.
Key Takeaways

The first thing worth knowing about this repo is not that it highlights code. It is that it refuses to stop highlighting when the parser breaks. That makes the project feel less like a demo and more like a small, disciplined argument about software boundaries.

The repository describes itself as SagharRamezani/Syntax-Highlighter, a modular syntax highlighter built for a university project using a Core/Plugin architecture. The public footprint is tiny, which is part of the charm: the code is doing the talking.


Why a broken file still gets colored

Most highlighters quietly assume the source is good enough to parse, or they collapse into a purely lexical path from the start. This project does something more deliberate. It keeps a no-parse fallback, so broken code still gets a useful visual treatment instead of a blank screen or a hard failure.

The same highlighting contract survives both success and failure. Parsing can enrich the output, but it does not own the output.

That is the article’s key idea in one sentence: the parser is an enrichment step, not a single point of failure. In a lot of tools, parsing is the whole game. Here, parsing is one possible source of spans.

A modular syntax highlighter built for a university project using a Core/Plugin architecture.

SagharRamezani, Author/Maintainer · SagharRamezani/Syntax-Highlighter

The abstraction that holds everything together

The project’s most important move is the Span model. Once highlighting is reduced to a list of spans with start, end, and role, the rendering layer stops caring whether those spans came from a parser, a lexer, or a recovery path.

@dataclass(frozen=True)
class Span:
    start: int
    end: int
    role: str

# parser, fallback lexer, and both renderers
# all speak in spans

That separation is the whole design. It means the core can normalize spans, sort them, and render them without knowing anything about Haskell, Rust, or Dart.

It also means the project can keep the same conceptual contract across terminal and web output. ANSI and HTML are different surfaces, but they are fed by the same internal shape.


Inside the core/plugin split

A close-up cabinet shows three language drawers labeled Haskell, Rust, and Dart. Each drawer has different token shapes inside, but every drawer feeds into one shared rail marked Span before the output splits again toward terminal and web renderers.
Languages vary, but the contract does not. Plugins supply language-specific intelligence, then hand off to the same core pipeline.

The core/plugin split is the repo’s other clean boundary. The engine lives under src/highlighter/core/, while language packages sit under src/highlighter/languages/. That keeps the core language-agnostic and makes new languages feel like additions, not surgery.

DimensionSyntax-HighlighterTypical regex-first highlighters
Language supportSmall set of language plugins, designed to be extendedBroad catalog, usually optimized for coverage
Internal contractOne Span model for all output targetsLexer output often tied more tightly to a specific renderer
Failure handlingCan fall back to lexical highlighting when parsing failsOften assumes lexical highlighting is the primary mode
ArchitectureCore plus plugins, with explicit registry boundariesUsually flatter and more directly coupled to highlighting rules
Teaching valueShows how to separate parsing, tokenization, and renderingShows practical highlighting, but less about pipeline design

That table is the point. This repo is not trying to beat Pygments or Prism.js on breadth. It is demonstrating a cleaner interface between language support and output generation.

How the parser stays useful when code is messy

The parser side uses a recursive descent approach with recovery. That matters because a syntax highlighter lives in the real world, where files are often half-written, broken, or intentionally incomplete.

Instead of treating malformed input as a dead end, the parser can salvage enough structure to keep producing useful spans. And when even that is too much, the fallback path still keeps the output alive.

This is the part that feels mature. The project does not confuse correctness with usefulness. It aims for both, but it knows usefulness has to win when the input is imperfect.

Why the renderers are boring on purpose

The ANSI and HTML renderers are intentionally plain. That is a good sign. Their job is not to invent new behavior. Their job is to preserve the span contract cleanly, after normalization, without leaking rendering quirks back into the core.

def render(spans, source):
    spans = _normalize_spans(spans)
    # turn spans into ANSI escapes or HTML tags
    return output

Normalization is where a lot of quiet engineering lives. If spans overlap or arrive out of order, the output can get ugly fast. By sorting and cleaning them first, the renderers stay predictable.

That predictability is what makes snapshot testing worthwhile. The repo is protecting rendered output as a stable interface, not as a side effect.

How it compares to the usual suspects

Compared with the big names, this project is tiny. But the comparison is useful because it clarifies what the repo is actually trying to prove.

ProjectMain strengthMain trade-offWhat Syntax-Highlighter does differently
PygmentsHuge language coverage and many formatsBroadness can hide architectural ideasKeeps the pipeline small enough to study
Prism.jsLightweight browser-side highlightingMostly browser-centric and lexer-drivenShows the same contract across terminal and web
Highlight.jsConvenient auto-detection and broad adoptionLess focused on explicit parser boundariesSeparates parsing, spans, and rendering more aggressively
ShikiHigh-quality themed output from TextMate grammarsIntegration model is different and heavierUses a simpler educational core with explicit fallback
Tree-sitterPowerful parsing infrastructureNot a drop-in highlighter by itselfTurns parsing into just one stage in a highlighting pipeline

So the real contrast is not feature count. It is intent. The mature tools optimize for coverage, quality, or ecosystem fit. This repo optimizes for clarity of architecture.

What the tests are really protecting

The testing story matters because it proves the design is not just neat on paper. Snapshot tests lock down ANSI and HTML output, which means the project is defending the contract all the way to the edge of the system.

That is exactly the right thing to test when the architecture revolves around stable spans. If the contract changes, the visible output changes. The tests make that visible immediately.

For a university project, that is a strong signal. It shows a builder who understands that clean boundaries are only valuable if they stay clean over time.


The lesson

This repo’s most useful lesson is bigger than syntax highlighting. If you want a system to survive malformed input, multiple output targets, and plugin growth, the critical decision is the contract you choose early.

Here, that contract is the span. Everything else, from parser recovery to ANSI and HTML rendering, hangs off that choice. That is why a small Python project can still teach a large architectural lesson.