molecular-similarity-search: When Chemistry Becomes a Matrix Problem

A small Flask app that turns SMILES strings into fingerprints, fingerprints into HDF5, and similarity search into one fast NumPy operation.

8 min read • View on GitHub • More from ashishp563

A chemical note on the left becomes a compact grid of fingerprint bits on the right, with a numerical beam connecting the query to a larger stored matrix. The scene explains how the repo turns a molecule into vector data before any search happens.
The important move is not the web app. It is the translation from chemical structure into a numerical representation that can be scored in one pass.
Key Takeaways

The trick: a chemical library is really a matrix

The cleanest thing about ashishp563/molecular-similarity-search is that it refuses to make chemistry mystical. A query molecule arrives as SMILES, gets converted into a fingerprint, and then meets a database that has already been flattened into numbers. Once that happens, the search problem stops looking like chemistry software and starts looking like matrix math.

That is the real story here. The repo does not loop through molecules and compare them one by one in Python. It precomputes the library, stores the result in HDF5, and lets NumPy do the heavy lifting at query time.

The pipeline is simple to describe and easy to miss: one query becomes one fingerprint, then that fingerprint is compared against a whole matrix in a single numerical pass.

From SMILES to fingerprints to scores

The pipeline starts with the standard cheminformatics move. RDKit parses the SMILES string into a molecule, then builds a Morgan fingerprint, which captures local structural neighborhoods as a bit vector. In this repo, that fingerprint is not just a representation. It is the search index.

from rdkit import Chem
from rdkit.Chem import AllChem
import numpy as np

mol = Chem.MolFromSmiles(smiles)
fp = AllChem.GetMorganFingerprintAsBitVect(mol, radius=2, nBits=1024)
arr = np.frombuffer(fp.ToBitString().encode("utf-8"), dtype=np.uint8) - ord("0")

Once the bit vectors exist, similarity becomes a Tanimoto score. Conceptually, that means measuring how much two fingerprints overlap relative to their combined size. In practice, the repo turns that into vectorized arithmetic so the database can be scored as a batch instead of as a Python loop.

A close-up of a mechanical similarity engine compares one query pattern against many stored patterns at once. The left side suggests manual pairwise checking, while the right side shows parallel scoring collapsing into a ranked result.
The implementation choice matters more than the UI: vectorized scoring turns a potentially tedious pairwise comparison into one batched calculation.

Why the index lives in HDF5

The storage choice is understated and smart. HDF5 is built for dense numerical data, which makes it a good fit for a fingerprint matrix and the precomputed bit counts that the search engine needs. There is no need to wrap the index in a heavier database layer when the entire problem is already numeric.

That matters because this repo is trying to keep the moving parts legible. The data is prepared once, loaded once, and then reused by the Flask app without a lot of ceremony. The result feels closer to a scientific instrument than a service with a lot of infrastructure hanging off it.

Why the search core is the real product

The most interesting file is search_engine.py. That is where the repo stops being a demo of RDKit and becomes a lesson in implementation discipline. The index is precomputed, the query is vectorized, and the similarity score is derived from operations that NumPy can do efficiently in one shot.

c = db_bits @ query_bits
sims = c / (query_count + db_counts - c + 1e-8)

That tiny formula is the whole trick. The dot product gives the overlap count across the entire database, and the denominator turns overlap into Tanimoto similarity. This is why the project feels fast even though the conceptual model is simple.

ApproachBest forWhat it hidesWhat this repo teaches
Loop-based searchTiny toy datasetsAlmost nothingHow expensive pairwise comparison feels in Python
RDKit aloneFingerprint generation and chemistry primitivesThe broader search pipelineHow fingerprints are created
FaissLarge-scale vector retrievalIndexing infrastructure and approximate search choicesHow similarity search becomes a matrix problem
PostgreSQL with chem extensionsProduction data systemsDatabase operations and indexing complexityHow storage and search can be separated cleanly

The UI makes the system feel finished

The frontend is not the star, but it does important work. The dashboard styling, async fetch flow, and table-first result view make the app feel like a tool a scientist could actually use, not a script someone wrapped in a browser tab. That polish changes the project’s perceived maturity.

This is the difference between a technical demo and a product-shaped repo. The backend explains the algorithm, but the interface tells you how the author wants it to be experienced: fast, direct, and ready for repeated use.

How it compares to the bigger tools

This repo is not trying to outrun RDKit, Faiss, or a production PostgreSQL stack. It sits underneath them conceptually. RDKit gives you the chemistry primitives, Faiss gives you large-scale vector retrieval ideas, and PostgreSQL gives you durable infrastructure. This project shows the small, readable implementation pattern that sits in the middle.

ProjectAbstraction levelMolecule-specificEngineering weight
molecular-similarity-searchReference implementationYesLight
RDKitChemistry libraryYesMedium
FaissVector search libraryNoMedium to heavy
PostgreSQL with chem extensionsDatabase platformYesHeavy

That is why the repo is worth reading. It shows the shape of a real optimization problem without burying it in production machinery. Precompute aggressively. Store numerically. Compute in batches. Keep the interface thin.

What this repo teaches better than a tutorial

A lot of tutorials show the happy path and stop there. This project goes one step further by making the data model and the search model line up. That is the transferable lesson: if the domain can be reduced to vectors, the pipeline gets simpler, faster, and easier to reason about.

In that sense, the repo is more than a molecular search app. It is a compact example of how to build around representation. Once the system thinks in fingerprints instead of strings, the rest of the design becomes obvious.