walgit: The Git Server Where the Bucket Is the Repository

A Rust-based Git host that turns object storage into the source of truth, treats local disk as disposable cache, and rebuilds itself from a write-ahead log.

8 to 10 min read • View on GitHub • More from tobi

A small stateless Git server on wheels sits beside a massive warehouse of object storage, connected by a thin stream of packfiles and manifests. The image explains that the server is temporary, while the bucket holds the durable repository state.
walgit in one picture: the server is a delivery truck, not the warehouse.
Key Takeaways

Most Git servers assume the machine is the thing you protect. walgit flips that assumption. The server is just a process pointing at S3 or GCS, while the repository’s durable state lives in the bucket as a write-ahead log and manifest.

The repository is not on the server

That is the move. Instead of keeping Git history on local disk or in a database, walgit stores the authoritative state in object storage and treats the node as a cache that can be rebuilt at any time. The practical effect is simple: if a box dies, you do not repair the box. You start another one and let it catch up.

DimensionTraditional Git serverwalgit
Source of truthLocal disk, database, or tightly coupled storageObject storage bucket plus WAL
Local disk rolePersistent state that must survive failuresDisposable cache
Crash recoveryRestore the machine or rebuild from backupsReplay from published sequence to head sequence
Scaling modelKeep the server alive and make disks biggerAdd stateless serving nodes and separate maintenance
Maintenance burdenBackups, disk migration, state repairStorage connectivity, cache warmup, WAL replay
Large repository behaviorHeavy reliance on local capacity and careful opsRemote reader and range fetches reduce disk pressure
Operational complexityTied to one box or one shared filesystemTied to object-store semantics and protocol correctness

How walgit turns pushes into an append-only log

The write path is the core story. A push lands as an immutable packfile, metadata is appended to the WAL, and then the manifest advances with a compare-and-swap step. That CAS gate matters because it is the point where durable intent becomes visible state.

The write path is a log, not a mutable server state machine. Recovery is just catch-up.

A close-up of an immutable packfile sliding into storage, followed by a manifest checkbox flipping forward through a compare-and-swap step. A small crash event sits to the side, but the recovery cursor continues moving. The image explains how durability, consistency, and recovery are separated into distinct steps.
A push becomes durable state only after the manifest advances by CAS.

The recovery path is just as important as the write path. walgit tracks a published cursor and can replay from that point to the head of the log, which means a replacement node does not need a handcrafted repair ritual. It reads the bucket, catches up, and rejoins the cluster.

Why the request path stays fast

A durable system is not automatically a slow one. walgit keeps Smart HTTP in the hot path, with v0 and v2 support, capability advertising, and a ref advertisement cache so it does not recompute the same answers for every fetch.

// Simplified shape of the serving path
async fn smart_http(state: AppState, req: Request) -> Response {
    let refs = state.caches.ref_advert.get_v0(&state.repo_id).await?;

    match req.git_protocol_version() {
        GitProtocolVersion::V2 => handle_v2(req, refs, &state).await,
        GitProtocolVersion::V0 => handle_v0(req, refs, &state).await,
    }
}

// Heavy work stays off the request path
spawn_background_task(maintain::run_loop(state.clone()));

That split matters. Fetch and ls-refs stay responsive, while compaction and bundle building move into background maintenance. The server stops trying to do everything at once.

Walgit, from @tobi, is exactly what @martinkl was talking about with the classic line: "the log is the database, and everything else is cache or materialized views." In walgit, the git repository is a write-ahead log in a bucket, and the SoT is the log itself, and the machines

Tigris Data, Community comment · @TigrisData on X

Heavy Git work gets pushed out of the hot path

walgit draws a sharp line between serving and maintenance. Serving nodes answer requests and keep caches warm. Maintainer nodes handle geometric compaction and bundle generation, which are the expensive chores that would otherwise poison latency for everyone else.

That separation makes the system feel like infrastructure rather than a single-purpose daemon. You can scale the reader side for traffic and scale the maintainer side for housekeeping, without pretending those are the same problem.

The remote reader changes the disk math

The other payoff is storage efficiency. The remote reader can use HTTP range requests against the bucket, which means a node does not need to keep every object locally to be useful. It can serve large repositories with a much smaller working set on disk.

Disk questionConventional serverwalgit with remote reader
How much history must stay local?Often most or all of itOnly what the cache and request mix demand
What happens to blobs?They usually occupy server diskThey can stay in object storage and be fetched by range
What does a small node look like?A constrained server that risks evictionA thin cache that still serves real traffic
How do you grow?Bigger disks, more careful backupsMore nodes, smarter caching, same bucket

That is why the architecture is compelling for large monorepos. It turns disk from a hard ceiling into a performance hint.

What walgit is really competing with

The real comparison is not Git versus Git. It is persistent-state hosting versus bucket-backed hosting. Traditional stacks optimize around keeping a machine healthy. walgit optimizes around keeping a storage substrate reachable.

QuestionPersistent Git hostingBucket-backed walgit
Where do you trust the truth?A server, filesystem, or databaseObject storage plus WAL
What fails first?The box, then the repoA node, which can be replaced
What is the ops burden?Patch, back up, migrate, restoreKeep storage available and caches warm
Who pays for scale?The machine and its disksThe object store and request layer
Best fitStateful legacy deploymentsCloud-native Git infrastructure with large repos

That is a conservative win, not a universal one. If you want a small self-contained Git box, this is probably too much architecture. If you want a clean durability model for serious hosting, the design starts to look obvious.

Why this repo feels unusually sharp

The codebase signals production intent everywhere: modern Rust edition, strict lints, clear crate boundaries, metrics and tracing, and a split between protocol handling, storage, bundle logic, and maintenance. Nothing about it reads like a demo.

That is the final tell. walgit is interesting because the idea is elegant, but it is credible because the implementation respects the ugly parts of real systems: retries, caches, recovery, and the cost of heavy Git operations.