You Don't Need an AI Database: Inside Eventual-Inc/reference-multimodal-datalake
How a minimalist blueprint uses standard Parquet, S3, and the Daft query engine to process millions of images without a specialized vector store.

This repo shows an example of building an effective **Multimodal** data warehouse using nothing but Daft, good old Parquet files, and URLs pointing to files in AWS S3 (or any object store of your choice, really).
- The reference-multimodal-datalake architecture avoids specialized AI databases by storing multimodal blobs in standard S3 buckets and maintaining lightweight pointers in Parquet files.
- Daft's stateful User Defined Functions (UDFs) enable efficient machine learning inference over large datasets without prohibitive memory overhead.
- By relying on standard formats like Parquet, the architecture ensures high interoperability with traditional SQL engines like DuckDB and Trino.
- This approach eliminates the impedance mismatch between data engineering and machine learning, allowing seamless transitions from analytics to PyTorch dataloading.
The AI Storage Trap
The current ecosystem pushes teams toward specialized vector databases or proprietary AI storage formats to handle multimodal data. Embedding 10GB videos or millions of raw images directly into a database is an architectural dead end that destroys interoperability. The Eventual-Inc/reference-multimodal-datalake repository offers a counter-proposal: standard Parquet files, simple S3 URLs, and zero re-encoding.
The Pointer Paradigm
The repository's ingestion notebooks detail how raw web-scraped data is normalized. Instead of storing image bytes, the system extracts critical metadata and stores it alongside a lightweight S3 pointer in Parquet. This allows data scientists to run heavy analytics using only metadata, without ever loading a single image into memory.
Stateful UDFs and the Compute Layer
A datalake requires a way to process the data. The project uses Daft, a distributed Python query engine, to handle the heavy lifting. It relies on stateful User Defined Functions (UDFs) to load ML models (like CLIP) into GPU memory exactly once per worker, allowing massive-scale embedding calculations without prohibitive overhead.
Bridging the ML Divide
Data engineers prefer SQL and Parquet, while ML engineers favor Python and PyTorch tensors. The reference datalake eliminates this traditional impedance mismatch. Because the data remains in standard formats, it can be queried by DuckDB for analytics and streamed directly into PyTorch dataloaders for training without complex ETL pipelines.
The Lakehouse vs. The Vector Store
Compared to specialized alternatives like Deep Lake and LanceDB, the reference-multimodal-datalake approach prioritizes pure interoperability. While purpose-built tools offer extreme optimization for specific edge cases, the Parquet/S3/Daft stack avoids vendor lock-in entirely.
| Feature | reference-multimodal-datalake | Deep Lake | LanceDB |
|---|---|---|---|
| Storage Format | Standard Parquet + S3 URLs | Custom tensor storage format | Lance format |
| Data Duplication | Zero re-encoding | Re-encodes data | Handles blobs internally or via references |
| ML Framework Integration | Daft to PyTorch | Native PyTorch/TF | Native PyTorch |
| Interoperability | High (DuckDB, Trino) | Low | Medium |