`img2svg-bench`: The Benchmark That Turns SVG Taste Into Numbers

A local-first harness for comparing raster-to-vector pipelines, stripping away background noise, extracting fast metrics, and making “better SVG” less subjective.

9 min read • View on GitHub • More from Sunwood-ai-labs

A lab bench where a raster image is cleaned in a transparent chamber and then exits as a stack of crisp SVG path sheets. The scene explains the project’s core idea: benchmark inputs are normalized first, then vectorization outputs are compared on objective signals.
Before the race starts, the image gets cleaned. That makes the comparison less about stray background artifacts and more about the vectorizer itself.
Key Takeaways

Why SVG conversion needs a benchmark

Raster-to-vector conversion sits in a frustrating middle ground. You can see when an SVG looks better, but you cannot always explain why, and you definitely cannot trust a quick eyeball test across dozens of presets. That is the gap `img2svg-bench` tries to close.

The repo treats SVG quality as a measurement problem. Instead of asking people to argue over taste, it compares tools on file size, path count, timing, and reportable output, then keeps the whole process local so private datasets never need to leave the machine.

ApproachWhat you getWhat you lose
Manual comparisonFast intuition on a few examplesSubjective, hard to reproduce, slow at scale
`img2svg-bench`Repeatable runs, CSV summaries, side-by-side reportsLess room for ad hoc human judgment
Raw converter defaultsA quick first passNo way to know if a better preset exists

What `img2svg-bench` actually does

The repository is not a converter. It is a harness. It organizes datasets, runs multiple presets, writes outputs into a local workspace, and generates reports that let you compare results without rebuilding the experiment each time.

The repo is closer to an experiment loop than a single script. Clean input, run presets, collect metrics, then publish the results in a form humans can scan quickly.

Why `img2svg-bench`? Because not all SVG converters are created equal. Our tool helps you objectively compare performance and quality, so you can choose the right tool for the job. Local-first, open-source, and easy to use. #DeveloperTools #OpenSource #ImageProcessing

Sunwood AI Labs, Project Creator/Maintainer · X Post by Sunwood AI Labs

The trick that makes the benchmark honest

The sharpest idea in the repo is not the reporting layer. It is the background remover. `remove_edge_background.py` samples corners and edges, estimates the border color, then uses a flood fill from the outside in to make the background transparent before vectorization.

A close-up of a flood-fill frontier moving from the image edges toward the center while the border pixels are absorbed into transparency. The illustration explains how the benchmark removes background noise before vectorization so the converter can focus on the subject.
Benchmark quality starts before the converter runs. If the background is dirty, the comparison gets noisy too.
# Conceptual shape of the preprocessing step
# 1. Sample corners and edges
# 2. Estimate the background color
# 3. Flood fill contiguous matching pixels
# 4. Convert them to transparent RGBA

background = estimate_border_color(image)
mask = flood_fill_from_edges(image, background, threshold)
image.putalpha(mask)

That choice matters because vectorizers are sensitive to noise. A bright border can inflate path counts, create tiny artifacts, and make a clean subject look worse than it is. The benchmark does not just measure output. It shapes the input so the result is fairer.

How the harness measures SVG output

`vtracer_experiments.py` is the workhorse. It defines presets as structured parameters, reads images through Pillow, runs conversions, then extracts lightweight metrics like `` counts and unique fills with regex rather than a heavy XML parser.

MetricWhy it helpsWhy it is enough here
File sizeTracks output compactnessSmall files often reflect simpler vector structure
Path countShows complexity directlyEnough to compare presets at scale
TimingExposes throughput tradeoffsFast feedback matters in benchmarking

This is a practical tradeoff. Regex is not the purest way to inspect SVG, but it is fast, portable, and good enough for a benchmark whose job is to compare many runs, not validate every XML nuance.

Why regex beats heavy parsing here

A benchmark does not need to become a parser project. If the question is, “Did this preset produce a simpler SVG, and how quickly did it do it?” then counting paths and fills with small string rules is a reasonable answer. The repo prefers narrow tools that keep the loop tight.

MethodStrengthWeakness
Regex over SVG textFast and simpleCan miss edge-case structure
Full XML parsingMore formalSlower and more code than this use case needs
Human inspection onlyRich judgmentHard to scale or reproduce
path_count = len(re.findall(r"<path\b", svg_text))
fill_count = len(set(re.findall(r'fill="([^"]+)"', svg_text)))

From VTracer to OmniSVG

The benchmark gets more interesting when it stops being about one tracer. Scripts like `build_omnisvg_report.py` show the repo widening its scope toward AI-generated SVG workflows, which fail differently from classical tracers and need different comparisons.

That shift changes the question. Classical tools often fail by creating too much geometry. AI-style generation can fail by drifting structurally, missing token groups, or producing SVGs that are harder to reason about. The benchmark becomes a way to compare eras, not just tools.

PipelineTypical strengthTypical failure mode
Potrace and similar tracersStable geometric tracingCan become dense or overly stylized
VTracer-style toolsModern preset controlStill sensitive to preprocessing and tuning
AI SVG generationFlexible synthesisCan lose structural predictability

That is why the repo feels larger than its size. It is not just documenting a tool. It is tracking a changing category and giving people a common yardstick for it.

🎨 `img2svg-bench` update: We've added new presets for `potrace` and `vtracer`! Now you can compare even more options to find the perfect balance between quality and file size. Check out the latest benchmark reports in the repo. #SVG #VectorGraphics #Benchmarking

Sunwood AI Labs, Project Creator/Maintainer · X Post by Sunwood AI Labs

Why the project feels bigger than its size

The codebase is small, but the thinking is mature. It uses modern Python patterns, keeps the workflow local-first, and pairs scripts with a polished docs site so the benchmark is not trapped in the terminal.

That combination makes `img2svg-bench` unusually legible. It is focused enough to answer one hard question well, and broad enough to absorb new converters as the field changes. The result is a tiny repo with the shape of a lasting measuring tool.