LensVault: The Self-Hosted Photo App That Turns Raw Media Into a Private, Searchable Archive

A FastAPI and React stack that deduplicates uploads, extracts metadata, clusters faces locally, and serves Google Photos-style convenience without sending your library to the cloud.

8 min read • View on GitHub • More from ShubhamSinghRawat10

A single printed photograph enters a mechanical sorting apparatus and exits as neatly organized folders, thumbnail cards, and clustered face markers. The scene explains LensVault’s core idea: one upload becomes a structured archive instead of a pile of files.
LensVault treats ingestion as transformation, not just storage. A raw file becomes a deduped original, a browseable timeline item, and a candidate for local intelligence.
Key Takeaways

What LensVault Actually Does to a Photo

LensVault is most interesting when you follow a single upload through the system. A file comes in, gets hashed, checked for duplication, parsed for EXIF and GPS, written into a human-readable year/month path, then fanned out to background workers for thumbnails and AI enrichment.

That sequence matters because it reframes the app. LensVault is not a gallery wrapped around storage. It is a media pipeline that converts raw uploads into a private archive with structure, search, and intelligence baked in.

The pipeline matters more than any single feature. LensVault is built so ingestion, enrichment, and browsing are one continuous system.


Why Self-Hosting Matters Here

Photo libraries are not disposable data. They hold family history, location traces, and the kinds of search patterns that reveal a lot about a person. Once a cloud platform becomes your memory layer, convenience starts to look like lock-in.

LensVault answers that problem with a simple thesis: the intelligence can stay local. The stack gives you most of the useful behaviors people expect from modern photo apps, but it does it without shipping the library, the embeddings, or the metadata to someone else’s cloud.

DimensionCloud photosLensVault
OwnershipProvider controls storage and inferenceYou control the stack and the data
SearchConvenient, but platform dependentStructured queries over local metadata
AI processingUsually remote and opaqueFace clustering and tagging run locally
Storage disciplineHidden behind product abstractionsContent hash, year/month paths, and dedupe logic
Best fitHands-off consumer convenienceSelf-hosted users who want control and insight

The Upload Path Is the Product

The core endpoint, upload_photo, shows what the project values. It hashes first, stores second, enriches third, and hands the expensive work to workers after the request is already done. That is a clean separation between user interaction and media processing.

def upload_photo(file, user_id):
    file_hash = sha256(file)
    if photo_exists(user_id, file_hash):
        return existing_photo_record()

    metadata = extract_exif(file)
    path = save_original(file, metadata)
    record = create_photo_record(user_id, file_hash, path, metadata)
    queue_thumbnail_job(record.id)
    queue_ai_enrichment_job(record.id)
    return record

The implementation is pragmatic. FastAPI handles the API edge, Celery and Redis absorb the heavy lifting, and the database becomes the source of truth for everything the UI later surfaces.

Several portrait tiles are pinned on a board and connected by dotted lines into a few tight clusters. One tile is pulled toward a local locked box, showing that LensVault keeps face intelligence on-device instead of routing it to the cloud.
The face layer is local and asynchronous. InsightFace builds embeddings, DBSCAN groups them, and the app turns those clusters into a browsable people view.

The Local Intelligence Stack

LensVault’s AI story is not generic. It uses InsightFace for detection and embeddings, DBSCAN for clustering, and CLIP for semantic tagging. That combination is important because each step solves a different problem: who is in the photo, which photos belong together, and what is in the scene.

The interesting part is not that those models exist. It is that the app treats them as background infrastructure. Uploading a photo does not block on inference, and browsing does not wait on a model call. The pipeline is designed so intelligence accumulates in the background while the user keeps moving.

Why the Search Feels Smarter Than a Folder Browser

LensVault’s search is more than filename lookup. The query syntax supports structured filters like camera, date ranges, and metadata-aware retrieval, then falls back to text search when the query is more open-ended. That makes the interface feel closer to an archive system than a folder tree.

Query styleWhat it does wellWhat it misses
Folder browsingSimple, predictable, familiarBreaks down once a library gets large
Filename searchFast for known filesIgnores the actual content of the media
LensVault structured searchUses metadata, time, and capture contextDepends on good extraction and indexing

The Architecture Bet: FastAPI, Celery, and Self-Healing Migrations

This is a small but serious architecture. FastAPI handles request traffic, Celery handles expensive jobs, PostgreSQL stores the record of truth, and Redis moves work through the queue. The split is obvious, but that is the point: media apps live or die by whether they can keep ingest fast while work piles up behind the scenes.

The most opinionated choice is the startup behavior. LensVault includes schema patching on launch, so the app can try to heal older databases instead of forcing every self-hosted operator to manually babysit migrations. That is not elegant in a purity sense, but it is very friendly to the reality of home deployments.

🖼️ LensVault — Hackathon build Smart photo organizer that clusters faces (InsightFace buffalo_l + DBSCAN) and auto-tags objects/scenes using OpenAI CLIP

Shubham Singh Rawat, Lead Developer · ShubhamSinghRawat10 GitHub Profile

Where LensVault Fits in the Self-Hosted Photo Market

LensVault is not trying to beat every incumbent on depth. It is sharper than that. Its identity is pipeline-first, local-first, and compact enough to feel like a focused implementation of a strong idea rather than a sprawling suite.

ProjectPrimary focusStackAI approachStrengthTrade-offBest fit
LensVaultPrivate media pipelineFastAPI, React, Celery, PostgreSQLLocal face clustering and semantic taggingClear ingestion-to-intelligence storyLess mature than the biggest incumbentsBuilders who want control and clarity
ImmichMobile-first photo backupGo backend, mobile appsBroad ML featuresPolished product experienceHeavier surface areaUsers who want a fuller replacement
PhotoPrismBrowsing and classificationGo and related servicesClassification and indexingEstablished browsing modelLess pipeline-centricUsers who prioritize browsing and indexing
LibrePhotosSelf-hosted photo managementPython stackFace recognition and browsing featuresSimilar privacy thesisLess sharply differentiated pipelineUsers who want a familiar open-source path

What This Project Suggests About the Next Wave of Personal Media Tools

The bigger pattern is clear. Personal media apps are shifting from passive viewers to active systems that organize, infer, and enrich. The winning products will not just show you files. They will understand enough about those files to make them searchable, sortable, and useful without giving up ownership.

LensVault is compelling because it shows how far that can go with a modern self-hosted stack. It is small enough to read, opinionated enough to learn from, and ambitious enough to hint at what a private media operating system could look like.