Sail: The Rust Engine That Pulls Spark Out of the JVM

A drop-in Spark replacement that keeps PySpark and Spark SQL intact, then rebuilds execution, catalogs, and session state in native code.

10 min read • View on GitHub • More from lakehq

A Spark notebook and SQL query window sit in the foreground, linked through a narrow protocol tunnel to a native Rust backend made of gears, Arrow columns, and planner blocks. Off to the side, an idle JVM machine room is unplugged and abandoned, showing that the interface stayed while the engine changed.
Sail keeps the Spark front door and swaps out the machinery behind it.
Key Takeaways

Sail is interesting because it does not try to make Spark slightly less painful. It tries to move the pain out of the picture entirely. The trick is simple to describe and hard to pull off: keep the Spark interface, replace the implementation, and do it in Rust.

Spark Without the JVM

The headline claim is not that Sail is faster Spark. It is that Sail can preserve the PySpark and Spark SQL experience while dropping the Scala and Java runtime underneath it. That is a much bigger bet than an accelerator plugin.

Sail is the only solution we are aware of that is designed for the Big Data and AI unification. ... As the first milestone, Sail can now be used as a drop-in replacement for Apache Spark in single-process settings.

Why the Protocol Matters More Than the Engine

Sail’s central insight is that Spark Connect is more valuable as a contract than the old Spark runtime is as a codebase. Once the client talks over a stable protocol, the backend can be replaced without forcing every user to rewrite their code.

Spark Connect becomes the membrane that separates compatibility from execution.

That matters because compatibility is no longer welded to the JVM. The client can still feel like Spark, while the server can behave like a native system designed for startup speed, memory efficiency, and ephemeral workloads.

Inside Sail’s Rust Workspace

The repository is organized like a product with clear seams, not like a single monolith. The core pieces split into crates for Spark Connect, catalog backends, DataFusion extensions, CLI startup, and build-time code generation.

crates/
  sail-cli/
  sail-spark-connect/
  sail-catalog/
  sail-catalog-glue/
  sail-catalog-hms/
  sail-catalog-iceberg/
  sail-catalog-unity/
  sail-catalog-onelake/
  sail-common-datafusion/
  sail-build-scripts/

That structure tells you what the project thinks is hard. It is not just query execution. It is the whole surface area around it: catalog discovery, session state, protocol translation, and the parts of Python integration that break if you ignore them.

A process fork splits into two clear branches. One branch becomes a Python interpreter with a RUN_PYTHON path, while the other continues as the Rust server. The split shows how Sail preserves Python process behavior instead of treating Python as a thin client only.
Sail respects Python process semantics, not just SQL semantics.

The Hidden Trick: Python Still Works Like Python

The most revealing file in the repo is not a planner or a parser. It is the CLI entry point. The process can fork into a plain Python interpreter when the environment calls for it, which means Sail is trying to preserve multiprocessing and resource tracking behavior, not just query results.

Both Blaze and DataFusion Comet operate as Spark accelerators. They replace Spark physical plans with DataFusion ones when feasible, but fallback to the Spark Java implementation in other situations. ... Sail takes a different approach. Sail implements distributed processing from the ground up in Rust.

lake_sail (Heran Lin), Maintainer · Introducing Distributed Processing with Sail v0.2

That is the difference between a faster engine and a more faithful one. Sail is not just chasing SQL parity. It is trying to behave like Spark from the operating system outward.

How Sail Keeps Spark Semantics

Sail leans on a DataFusion extension layer to make native execution feel Spark-like. Session state, catalogs, temporary views, UDFs, and SQL behavior all need a place to live if the user is going to trust the system after the first query.

AreaSpark expectationSail approach
Session stateLives in the JVM runtimeLives in Rust session extensions
CatalogsTied to Spark’s built-in catalog modelManaged through modular catalog crates
Temporary views and UDFsResolved inside Spark’s processTracked and resolved in native catalog state
ExecutionJVM executor plus Python boundaryDataFusion execution with Arrow transfer

The important point is not that every Spark behavior is reproduced exactly by magic. The point is that Sail moves those behaviors into explicit native components, where they can be reasoned about, tested, and swapped.

What Sail Replaces, and What It Still Leans On

Sail is a stack of dependencies with a clear center of gravity. DataFusion does the heavy lifting for execution, Arrow carries columnar data, tonic handles gRPC, PyO3 bridges Python, and object storage connectors reach into cloud-backed data sources.

ComponentRole in SailWhy it matters
DataFusionPlanner and execution coreGives Sail a native query engine instead of a JVM wrapper
ArrowColumnar data transportCuts copy overhead and keeps transfers efficient
tonicgRPC transportImplements the Spark Connect server boundary
PyO3Python embeddingPreserves Python-side workflows and process behavior
object_storeStorage integrationConnects the engine to cloud object stores without extra shims

That mix explains why the project feels larger than a Spark plugin and smaller than a general-purpose platform. Sail is opinionated about the boundary that matters most: the one between the client protocol and the execution engine.

Why Sail Is Different From Comet, Gluten, and Ballista

This is where the category shift becomes obvious. Comet and Gluten accelerate Spark from inside the old house. Ballista and Ray offer new houses with new doors. Sail is trying to keep the Spark door and rebuild the house behind it.

ProjectArchitectureCompatibility modelWhere the JVM still existsWhat makes it different
SailStandalone Rust engineSpark Connect front doorNot in the execution pathReplaces the backend rather than accelerating it
CometSpark pluginNative operators with fallbackYes, as the fallback runtimeSpeeds up Spark without leaving Spark
GlutenSpark pluginNative execution via VeloxYes, for coordination and fallbackKeeps Spark semantics but stays hybrid
BallistaDistributed Rust engineCustom APINo JVM dependencyRust-native, but not Spark compatible
Ray or DaskDistributed Python systemsCustom APIsNo JVM dependencyGreat if you can rewrite the workload

That table is the thesis in compressed form. Sail is not competing on acceleration alone. It is competing on how much of Spark can survive once the JVM is gone.

The Bet Behind the Project

Sail has the feel of a serious project that is still proving the edge cases. The repo shows strict linting, a focused core team, and enough surface area to suggest real ambition rather than a demo.

That is also the risk. The broader the compatibility promise, the more invisible work is required in catalogs, UDFs, SQL semantics, and process behavior. Sail’s design is convincing because it knows exactly where the hard parts live.

If the project keeps tightening those seams, it could become more than a faster Spark. It could become a template for how a legacy data system survives after its protocol outgrows its original engine.