daft-runner-test: Daft Runner Test: The One-Way Door Hidden Inside Daft’s Execution Model
A tiny diagnostic repo reveals a sharp boundary between inference and initialization, and why Daft treats runner selection as a process-level commitment.
- Daft splits runner selection into a safe inference step and a irreversible creation step, and that split is the whole story.
- This repo works as executable documentation because it isolates a process-level rule that would be easy to miss in normal use.
- The second call fails by design, which makes runner choice feel like a contract rather than a configuration knob.
- Daft’s model favors clarity about side effects over mid-process flexibility, especially when distributed execution is on the table.
A Runner Choice You Only Get to Make Once
The surprise in Eventual-Inc/daft-runner-test is not that Daft supports both Native and Ray. It is that the choice is not treated like a toggle. Daft separates get_or_infer_runner_type() from get_or_create_runner(), and that boundary is the entire point of the repo.
One call asks what would happen. The other makes it happen and locks the process into that path. In practice, that means the runner is not a session preference. It is a commitment that lives for the life of the Python process.
What This Repo Actually Proves
This is not a big codebase pretending to be small. It is a small codebase that isolates one behavior cleanly enough to make it visible. The repository has two scripts, a pinned Daft version, a pinned Ray version, and just enough logic to prove a point.
# runner_check.py
import daft
print(daft.get_or_infer_runner_type())
runner = daft.get_or_create_runner()
print(runner)
# A second creation attempt is expected to fail.
try:
daft.get_or_create_runner()
except Exception as e:
print(type(e).__name__, e)
# runner_check_ray.py
import daft
print(daft.get_or_infer_runner_type())
daft.set_runner_ray()
runner = daft.get_or_create_runner()
print(runner)
| Library | Default execution posture | Can backend switch mid-process? | What the model optimizes for | What Daft does differently |
|---|---|---|---|---|
| pandas | Single-process, in-memory | No meaningful backend switch | Familiar local data manipulation | Daft makes backend selection explicit instead of implicit |
| Polars | Single-node first | Not as a process-level runtime choice | Fast local execution | Daft frames runner selection as a committed execution contract |
| Dask | Lazy distributed graph | Often reconfigured by a new client/session | Flexible parallel scheduling | Daft separates inference from creation more sharply |
| Spark | Cluster session centered | Usually requires a new session | Large-scale cluster orchestration | Daft keeps the choice visible inside the Python process itself |
| Ray Datasets | Distributed runtime first | Possible through new runtime setup | Distributed task execution | Daft can sit on top of Ray, but still treats runner creation as a one-way door |
Inference vs. Initialization
The cleanest mental model in the repo is simple. get_or_infer_runner_type() is a read. get_or_create_runner() is a write.
That distinction matters because the first call is safe to repeat. It can inspect environment state without starting a cluster or committing the process. The second call changes the world, which is why the next attempt can fail with a DaftCoreException instead of quietly picking a different path.
This is not just API hygiene. It is a boundary around side effects. Daft is telling you that observation and initialization belong to different phases, and that the phase boundary is enforced by the process, not by convention.
Why the Second Call Fails
The failure mode is the proof. In runner_check.py, a second attempt to set or create the runner triggers DaftCoreException. That is the library making an architectural rule visible instead of hiding it in a silent no-op.
The practical lesson is straightforward: if a process has already committed to Native or Ray, you need a new process to explore the other path. The repo’s scripts are separate for a reason. They are not just examples. They are isolated experiments around a one-runner-per-process rule.
Native by Default, Ray by Intent
The default posture is Native. That matters because it tells you how Daft thinks about ordinary use: local execution first, distribution when requested. Ray is supported, but it is not the ambient assumption.
That design is visible in the repo’s split scripts. One checks the default and the locked Native path. The other forces Ray. Together they show that distributed execution is not a casual runtime switch. It is an explicit mode selection.
| Aspect | Native | Ray |
|---|---|---|
| Default expectation | Yes | No |
| Setup cost | Low | Higher |
| Process state | Single local engine | Distributed runtime bound to the process |
| What the repo demonstrates | The lock after inference | The explicit path when Ray is chosen |
| Why it matters | Fast local work without surprises | Distributed work only after a deliberate commit |
Why Daft Would Design It This Way
The constraint looks strict until you treat it as a safeguard. Runner creation can carry real side effects. Starting the wrong backend, or switching underneath active objects, would create ambiguity that is much harder to debug than a hard failure at the boundary.
That is the architectural trade-off this repo exposes. Daft appears to prefer predictability over convenience once execution begins. For a data engine that can run locally or on Ray, that is a sensible way to keep performance, state, and debugging aligned.
How It Compares to Other Data Engines
The comparison is not about raw capability. It is about how much freedom a library leaves you after startup. Daft is unusually explicit about that boundary.
| Library | Execution posture | Runtime switching model | Takeaway |
|---|---|---|---|
| pandas | Local-first | Not the point | Configuration stays light because execution stays local |
| Polars | High-performance local engine | Session-level changes are not the story | Speed matters more than backend choreography |
| Dask | Graph-driven parallelism | Often configured through a client or scheduler setup | Flexibility comes from lazy task planning |
| Spark | Cluster session model | New session for major changes | The session is the unit of commitment |
| Daft | Native or Ray runner selected per process | Inference is separate from creation | Observation and commitment are deliberately split |
That is why this tiny repo is more revealing than a broad feature demo. It shows that Daft is not trying to hide process boundaries. It is surfacing them.
What This Tiny Repo Teaches Library Authors
Good libraries draw a bright line between looking and doing. Good diagnostics make that line testable. And good execution engines do not pretend that backend selection is harmless once the process has already started work.
That is the lesson here. A tiny repository can still act like a design memo. In this case, the memo is clear: inspect first, initialize once, never switch.