molecular-similarity-search: When Chemistry Becomes a Matrix Problem
A small Flask app that turns SMILES strings into fingerprints, fingerprints into HDF5, and similarity search into one fast NumPy operation.
- The repo’s main idea is to stop treating molecules as strings and start treating them as vectors.
- Its speed comes from precomputing fingerprints, storing them numerically, and scoring the whole library with batched linear algebra.
- HDF5 is the quiet design choice that keeps the index simple, portable, and fast to load.
- The polished Flask UI matters because it makes a technical search pipeline feel like a finished lab tool.
The trick: a chemical library is really a matrix
The cleanest thing about ashishp563/molecular-similarity-search is that it refuses to make chemistry mystical. A query molecule arrives as SMILES, gets converted into a fingerprint, and then meets a database that has already been flattened into numbers. Once that happens, the search problem stops looking like chemistry software and starts looking like matrix math.
That is the real story here. The repo does not loop through molecules and compare them one by one in Python. It precomputes the library, stores the result in HDF5, and lets NumPy do the heavy lifting at query time.
From SMILES to fingerprints to scores
The pipeline starts with the standard cheminformatics move. RDKit parses the SMILES string into a molecule, then builds a Morgan fingerprint, which captures local structural neighborhoods as a bit vector. In this repo, that fingerprint is not just a representation. It is the search index.
from rdkit import Chem
from rdkit.Chem import AllChem
import numpy as np
mol = Chem.MolFromSmiles(smiles)
fp = AllChem.GetMorganFingerprintAsBitVect(mol, radius=2, nBits=1024)
arr = np.frombuffer(fp.ToBitString().encode("utf-8"), dtype=np.uint8) - ord("0")
Once the bit vectors exist, similarity becomes a Tanimoto score. Conceptually, that means measuring how much two fingerprints overlap relative to their combined size. In practice, the repo turns that into vectorized arithmetic so the database can be scored as a batch instead of as a Python loop.
Why the index lives in HDF5
The storage choice is understated and smart. HDF5 is built for dense numerical data, which makes it a good fit for a fingerprint matrix and the precomputed bit counts that the search engine needs. There is no need to wrap the index in a heavier database layer when the entire problem is already numeric.
That matters because this repo is trying to keep the moving parts legible. The data is prepared once, loaded once, and then reused by the Flask app without a lot of ceremony. The result feels closer to a scientific instrument than a service with a lot of infrastructure hanging off it.
Why the search core is the real product
The most interesting file is search_engine.py. That is where the repo stops being a demo of RDKit and becomes a lesson in implementation discipline. The index is precomputed, the query is vectorized, and the similarity score is derived from operations that NumPy can do efficiently in one shot.
c = db_bits @ query_bits
sims = c / (query_count + db_counts - c + 1e-8)
That tiny formula is the whole trick. The dot product gives the overlap count across the entire database, and the denominator turns overlap into Tanimoto similarity. This is why the project feels fast even though the conceptual model is simple.
| Approach | Best for | What it hides | What this repo teaches |
|---|---|---|---|
| Loop-based search | Tiny toy datasets | Almost nothing | How expensive pairwise comparison feels in Python |
| RDKit alone | Fingerprint generation and chemistry primitives | The broader search pipeline | How fingerprints are created |
| Faiss | Large-scale vector retrieval | Indexing infrastructure and approximate search choices | How similarity search becomes a matrix problem |
| PostgreSQL with chem extensions | Production data systems | Database operations and indexing complexity | How storage and search can be separated cleanly |
The UI makes the system feel finished
The frontend is not the star, but it does important work. The dashboard styling, async fetch flow, and table-first result view make the app feel like a tool a scientist could actually use, not a script someone wrapped in a browser tab. That polish changes the project’s perceived maturity.
This is the difference between a technical demo and a product-shaped repo. The backend explains the algorithm, but the interface tells you how the author wants it to be experienced: fast, direct, and ready for repeated use.
How it compares to the bigger tools
This repo is not trying to outrun RDKit, Faiss, or a production PostgreSQL stack. It sits underneath them conceptually. RDKit gives you the chemistry primitives, Faiss gives you large-scale vector retrieval ideas, and PostgreSQL gives you durable infrastructure. This project shows the small, readable implementation pattern that sits in the middle.
| Project | Abstraction level | Molecule-specific | Engineering weight |
|---|---|---|---|
| molecular-similarity-search | Reference implementation | Yes | Light |
| RDKit | Chemistry library | Yes | Medium |
| Faiss | Vector search library | No | Medium to heavy |
| PostgreSQL with chem extensions | Database platform | Yes | Heavy |
That is why the repo is worth reading. It shows the shape of a real optimization problem without burying it in production machinery. Precompute aggressively. Store numerically. Compute in batches. Keep the interface thin.
What this repo teaches better than a tutorial
A lot of tutorials show the happy path and stop there. This project goes one step further by making the data model and the search model line up. That is the transferable lesson: if the domain can be reduced to vectors, the pipeline gets simpler, faster, and easier to reason about.
In that sense, the repo is more than a molecular search app. It is a compact example of how to build around representation. Once the system thinks in fingerprints instead of strings, the rest of the design becomes obvious.