Eventual-Inc/Daft: The DataFrame Built for the Multimodal Era
How a Rust-powered execution engine bypasses the JVM tax to treat images, video, and audio as first-class data citizens.
- Daft replaces the traditional JVM-based data processing stack with a Rust execution core, eliminating expensive serialization overhead for AI workloads.
- By utilizing Apache Arrow as its foundational memory format, Daft enables zero-copy data handoffs directly to Python machine learning models.
- A native multimodal type system allows Daft to treat images and audio as structured data, pushing operations like resizing down to the optimization layer.
- The engine scales seamlessly from single-node execution to massive distributed clusters by integrating natively with Ray.
The Starving GPU Problem
Modern AI pipelines spend an exorbitant amount of compute just shuffling bytes. Moving data from a Spark cluster running on the Java Virtual Machine to a PyTorch training loop in Python requires expensive serialization and memory copying.
When that data involves millions of high-resolution images or audio files, the CPU becomes the bottleneck. Expensive GPUs sit idle waiting for data because legacy data pipelines choke on unstructured data and memory translation.
Bypassing the JVM Tax
Daft solves this architectural mismatch by leaving the JVM behind entirely. The codebase is roughly 75 percent Rust and 25 percent Python. By utilizing Apache Arrow as the foundational memory format and PyO3 as the bridge, Daft ensures that data remains in a zero-copy C++ and Rust memory space.
When the data is finally handed to a machine learning model, no translation is required. The handoff is instantaneous.
SQL for Pixels
Traditional databases handle unstructured data by casting it as an opaque binary large object. The engine knows nothing about what is inside. Daft introduces a multimodal type system with native media and image classifications.
`daft.File` brings file-native handling into Daft's lazy, distributed execution model.
This allows the engine to understand the data and dispatch specialized Rust kernels for operations like resizing or cropping before the data ever reaches the application layer.
import daft
# Load images directly from cloud storage
df = daft.from_glob_path("s3://bucket/images/*.jpg")
# The resize operation is pushed down to the Rust execution engine
df = df.with_column("resized", df["image"].image.resize(224, 224))
df.show()
Lazy Execution at Scale
Users write Python, but Daft executes a Rust logical plan. The DataFrame API is a lazy expression builder. When a user filters or joins data, Daft builds an expression tree, optimizes it, and maps the work into micro-partitions.
For distributed execution, it natively hooks into Ray to scale across thousands of nodes.
The Post-Spark Landscape
The data engineering ecosystem is fracturing into specialized tools. Polars dominates single-node Rust performance. Spark remains the enterprise standard for tabular data. Ray Data handles raw unstructured compute.
Daft carves out a distinct layer. It offers the ergonomics of a Pandas DataFrame, the scale of Spark, and the multimodal awareness of an AI-native tool.
| Feature | Daft | Apache Spark | Polars | Ray Data |
|---|---|---|---|---|
| Execution Core | Rust | JVM (Scala) | Rust | C++ / Python |
| Primary Focus | Distributed Multimodal | Distributed Tabular | Single-Node Tabular | Distributed Compute |
| Memory Format | Apache Arrow | JVM Objects / Tungsten | Apache Arrow | Apache Arrow |
| Image/Audio Awareness | Native Types | Opaque BLOBs | Opaque BLOBs | Basic Support |
| Scaling Boundary | Multi-node (Ray) | Multi-node (YARN/K8s) | Single Machine | Multi-node (Ray) |