Evading the PyTorch Dependency Tax: Inside chroma-core/onnx-embedding
How the ChromaDB team stripped a massive Transformer down to its bare mathematical essentials to build a frictionless, zero-config vector database.

We do this because sentence-transformers introduces a lot of transitive dependencies that we don't want to have to install in the chromadb and some of those also don't work on newer python versions.
- ChromaDB abandoned the standard sentence-transformers library to protect users from gigabytes of fragile PyTorch dependencies.
- The team achieved a dependency-free deployment by compiling the all-MiniLM-L6-v2 model into a static ONNX graph.
- Without high-level AI libraries, Chroma engineers had to manually rebuild Transformer mean pooling and normalization logic using raw Numpy.
- This lean architecture prioritizes day-one developer experience over raw batched inference speed.
The "Batteries Included" Dilemma
ChromaDB built its reputation on a flawless developer experience. You run a simple install command, and you immediately have a working vector database. No API keys are required to get started. However, shipping a default embedding model creates a massive logistical headache. The standard approach relies on the sentence-transformers library.
That library is powerful but heavy. It brings along a fragile dependency tree anchored by PyTorch binaries. These dependencies regularly exceed two gigabytes and frequently break on newer Python releases. For a database trying to remain lightweight, this dependency tax was simply unacceptable.
The ONNX Escape Hatch
The Chroma team needed a way to run a complex Transformer model without the baggage. Their solution lives in a specialized utility repository called chroma-core/onnx-embedding. Using Hugging Face's optimum library, the team authored a script to freeze the popular all-MiniLM-L6-v2 model into a static ONNX (Open Neural Network Exchange) graph.
This conversion severed the tie to PyTorch entirely. The runtime requirement dropped from gigabytes of machine learning frameworks to just a few megabytes of onnxruntime. The database could now ship with a "batteries-included" default model that installed instantly.
Rebuilding the Black Box in Numpy
Dropping the standard AI stack came with a hidden cost. Without sentence-transformers, the team lost the high-level API that magically turns token outputs into a single searchable vector. They had to perform open-heart surgery on the model's output layer.
Inside run_onnx.py, the developers manually implemented Mean Pooling and L2 Normalization. They used raw numpy.broadcast_to and numpy.expand_dims to apply attention masks to hidden states. It is a masterclass in low-level AI engineering, proving that you do not need a massive framework to execute complex tensor operations.
The Parity Check
A compressed model is useless if it loses semantic accuracy. The team had to guarantee that their ONNX port behaved exactly like the original PyTorch model. They built a rigorous validation suite in compare_onnx.py to test this.
The script runs both models through the GLUE STSb benchmark. It enforces strict dot-product checks, ensuring the embeddings match with an epsilon variance of 1e-6. This guarantees that users get the same predictive power without the bloat.
The Speed vs. Bloat Trade-off
Engineering is about trade-offs. While the ONNX implementation is a massive win for installation size and cold starts, it is not optimized for heavy enterprise workloads.
The benchmark suggests ONNX is an order of magnitude slower than SentenceTransformers.
Because the ONNX setup relies heavily on CPU-bound Numpy operations, it struggles with large batched inference compared to hardware-accelerated PyTorch. The Chroma team accepted this. The repository exists to provide a frictionless starting point. When users eventually hit scale, they can easily swap out the default function for a dedicated embedding provider.
| Feature | sentence-transformers | chroma-core/onnx-embedding |
|---|---|---|
| Core Dependency | PyTorch (~2GB+) | ONNX Runtime (~20MB) |
| Target Audience | Enterprise / Production Scale | Local Dev / Zero-Config |
| Pooling Logic | Abstracted High-Level API | Manual Numpy Broadcasting |
| Batch Inference | Highly Optimized (GPU/MPS) | CPU-bound (Standard) |