You Don't Need an AI Database: Inside Eventual-Inc/reference-multimodal-datalake

How a minimalist blueprint uses standard Parquet, S3, and the Daft query engine to process millions of images without a specialized vector store.

8 min read • View on GitHub • More from Eventual-Inc

A massive, glowing vault bypassed by a simple conveyor belt carrying crates to an open warehouse, illustrating the circumvention of complex AI databases.
The reference architecture rejects specialized AI databases in favor of standard object storage and Parquet files.

This repo shows an example of building an effective **Multimodal** data warehouse using nothing but Daft, good old Parquet files, and URLs pointing to files in AWS S3 (or any object store of your choice, really).

jaychia (via repository README), Primary Contributor · Eventual-Inc/reference-multimodal-datalake - GitHub
Key Takeaways

The AI Storage Trap

The current ecosystem pushes teams toward specialized vector databases or proprietary AI storage formats to handle multimodal data. Embedding 10GB videos or millions of raw images directly into a database is an architectural dead end that destroys interoperability. The Eventual-Inc/reference-multimodal-datalake repository offers a counter-proposal: standard Parquet files, simple S3 URLs, and zero re-encoding.

Hedcut portrait of jaychia.

The Pointer Paradigm

The repository's ingestion notebooks detail how raw web-scraped data is normalized. Instead of storing image bytes, the system extracts critical metadata and stores it alongside a lightweight S3 pointer in Parquet. This allows data scientists to run heavy analytics using only metadata, without ever loading a single image into memory.

The Pointer Architecture separates metadata queries from raw blob processing.

Stateful UDFs and the Compute Layer

A datalake requires a way to process the data. The project uses Daft, a distributed Python query engine, to handle the heavy lifting. It relies on stateful User Defined Functions (UDFs) to load ML models (like CLIP) into GPU memory exactly once per worker, allowing massive-scale embedding calculations without prohibitive overhead.

A close-up of a mechanical workbench where a stationary, complex iron stamp presses down on a continuous feed of paper tape, illustrating stateful UDFs.
Stateful UDFs initialize a heavy model once, rather than lifting and placing it for every single row of data.

Bridging the ML Divide

Data engineers prefer SQL and Parquet, while ML engineers favor Python and PyTorch tensors. The reference datalake eliminates this traditional impedance mismatch. Because the data remains in standard formats, it can be queried by DuckDB for analytics and streamed directly into PyTorch dataloaders for training without complex ETL pipelines.

Standard formats bridge the gap between data engineering and machine learning workflows.

The Lakehouse vs. The Vector Store

Compared to specialized alternatives like Deep Lake and LanceDB, the reference-multimodal-datalake approach prioritizes pure interoperability. While purpose-built tools offer extreme optimization for specific edge cases, the Parquet/S3/Daft stack avoids vendor lock-in entirely.

Featurereference-multimodal-datalakeDeep LakeLanceDB
Storage FormatStandard Parquet + S3 URLsCustom tensor storage formatLance format
Data DuplicationZero re-encodingRe-encodes dataHandles blobs internally or via references
ML Framework IntegrationDaft to PyTorchNative PyTorch/TFNative PyTorch
InteroperabilityHigh (DuckDB, Trino)LowMedium