Eventual-Inc/Daft: The DataFrame Built for the Multimodal Era

How a Rust-powered execution engine bypasses the JVM tax to treat images, video, and audio as first-class data citizens.

8 min read • View on GitHub • More from Eventual-Inc

A massive metallic funnel compressing various raw materials like text pages, film reels, and photographs into a single uniform glowing geometric beam. This illustrates Daft's ability to unify multimodal data into a single structured pipeline.
Daft unifies disparate, unstructured data types into a single high-performance execution pipeline.
Key Takeaways

The Starving GPU Problem

Modern AI pipelines spend an exorbitant amount of compute just shuffling bytes. Moving data from a Spark cluster running on the Java Virtual Machine to a PyTorch training loop in Python requires expensive serialization and memory copying.

When that data involves millions of high-resolution images or audio files, the CPU becomes the bottleneck. Expensive GPUs sit idle waiting for data because legacy data pipelines choke on unstructured data and memory translation.

Bypassing the JVM Tax

Daft solves this architectural mismatch by leaving the JVM behind entirely. The codebase is roughly 75 percent Rust and 25 percent Python. By utilizing Apache Arrow as the foundational memory format and PyO3 as the bridge, Daft ensures that data remains in a zero-copy C++ and Rust memory space.

A split scene comparison. On the left, a clunky wooden conveyor belt with workers manually packing heavy crates represents the JVM serialization overhead. On the right, a sleek magnetic levitation rail where items glide smoothly without containers represents zero-copy Arrow memory.
Moving from the JVM to Python requires heavy serialization. Daft uses Apache Arrow to keep data in a shared, zero-copy memory space.

When the data is finally handed to a machine learning model, no translation is required. The handoff is instantaneous.

SQL for Pixels

Traditional databases handle unstructured data by casting it as an opaque binary large object. The engine knows nothing about what is inside. Daft introduces a multimodal type system with native media and image classifications.

`daft.File` brings file-native handling into Daft's lazy, distributed execution model.

daft.ai, Project Blog · Introducing daft.File

This allows the engine to understand the data and dispatch specialized Rust kernels for operations like resizing or cropping before the data ever reaches the application layer.

import daft

# Load images directly from cloud storage
df = daft.from_glob_path("s3://bucket/images/*.jpg")

# The resize operation is pushed down to the Rust execution engine
df = df.with_column("resized", df["image"].image.resize(224, 224))

df.show()

Lazy Execution at Scale

Users write Python, but Daft executes a Rust logical plan. The DataFrame API is a lazy expression builder. When a user filters or joins data, Daft builds an expression tree, optimizes it, and maps the work into micro-partitions.

Daft's lazy execution model builds a logical plan, optimizes it by pushing down expensive operations, and shatters the workload into micro-partitions.

For distributed execution, it natively hooks into Ray to scale across thousands of nodes.

The Post-Spark Landscape

The data engineering ecosystem is fracturing into specialized tools. Polars dominates single-node Rust performance. Spark remains the enterprise standard for tabular data. Ray Data handles raw unstructured compute.

Daft carves out a distinct layer. It offers the ergonomics of a Pandas DataFrame, the scale of Spark, and the multimodal awareness of an AI-native tool.

FeatureDaftApache SparkPolarsRay Data
Execution CoreRustJVM (Scala)RustC++ / Python
Primary FocusDistributed MultimodalDistributed TabularSingle-Node TabularDistributed Compute
Memory FormatApache ArrowJVM Objects / TungstenApache ArrowApache Arrow
Image/Audio AwarenessNative TypesOpaque BLOBsOpaque BLOBsBasic Support
Scaling BoundaryMulti-node (Ray)Multi-node (YARN/K8s)Single MachineMulti-node (Ray)