walgit: The Git Server Where the Bucket Is the Repository
A Rust-based Git host that turns object storage into the source of truth, treats local disk as disposable cache, and rebuilds itself from a write-ahead log.
- walgit makes object storage the repository, so durability lives in the bucket instead of on a pet server.
- Its write path is an append-only log with compare-and-swap manifest updates, which makes crash recovery a replay problem.
- The serving path stays fast by caching refs and pushing heavy Git work into separate maintenance roles.
- The project competes less with Git itself than with operational habits built around persistent disks, backups, and tightly coupled servers.
Most Git servers assume the machine is the thing you protect. walgit flips that assumption. The server is just a process pointing at S3 or GCS, while the repository’s durable state lives in the bucket as a write-ahead log and manifest.
The repository is not on the server
That is the move. Instead of keeping Git history on local disk or in a database, walgit stores the authoritative state in object storage and treats the node as a cache that can be rebuilt at any time. The practical effect is simple: if a box dies, you do not repair the box. You start another one and let it catch up.
| Dimension | Traditional Git server | walgit |
|---|---|---|
| Source of truth | Local disk, database, or tightly coupled storage | Object storage bucket plus WAL |
| Local disk role | Persistent state that must survive failures | Disposable cache |
| Crash recovery | Restore the machine or rebuild from backups | Replay from published sequence to head sequence |
| Scaling model | Keep the server alive and make disks bigger | Add stateless serving nodes and separate maintenance |
| Maintenance burden | Backups, disk migration, state repair | Storage connectivity, cache warmup, WAL replay |
| Large repository behavior | Heavy reliance on local capacity and careful ops | Remote reader and range fetches reduce disk pressure |
| Operational complexity | Tied to one box or one shared filesystem | Tied to object-store semantics and protocol correctness |
How walgit turns pushes into an append-only log
The write path is the core story. A push lands as an immutable packfile, metadata is appended to the WAL, and then the manifest advances with a compare-and-swap step. That CAS gate matters because it is the point where durable intent becomes visible state.
The recovery path is just as important as the write path. walgit tracks a published cursor and can replay from that point to the head of the log, which means a replacement node does not need a handcrafted repair ritual. It reads the bucket, catches up, and rejoins the cluster.
Why the request path stays fast
A durable system is not automatically a slow one. walgit keeps Smart HTTP in the hot path, with v0 and v2 support, capability advertising, and a ref advertisement cache so it does not recompute the same answers for every fetch.
// Simplified shape of the serving path
async fn smart_http(state: AppState, req: Request) -> Response {
let refs = state.caches.ref_advert.get_v0(&state.repo_id).await?;
match req.git_protocol_version() {
GitProtocolVersion::V2 => handle_v2(req, refs, &state).await,
GitProtocolVersion::V0 => handle_v0(req, refs, &state).await,
}
}
// Heavy work stays off the request path
spawn_background_task(maintain::run_loop(state.clone()));
That split matters. Fetch and ls-refs stay responsive, while compaction and bundle building move into background maintenance. The server stops trying to do everything at once.
Walgit, from @tobi, is exactly what @martinkl was talking about with the classic line: "the log is the database, and everything else is cache or materialized views." In walgit, the git repository is a write-ahead log in a bucket, and the SoT is the log itself, and the machines
Heavy Git work gets pushed out of the hot path
walgit draws a sharp line between serving and maintenance. Serving nodes answer requests and keep caches warm. Maintainer nodes handle geometric compaction and bundle generation, which are the expensive chores that would otherwise poison latency for everyone else.
That separation makes the system feel like infrastructure rather than a single-purpose daemon. You can scale the reader side for traffic and scale the maintainer side for housekeeping, without pretending those are the same problem.
The remote reader changes the disk math
The other payoff is storage efficiency. The remote reader can use HTTP range requests against the bucket, which means a node does not need to keep every object locally to be useful. It can serve large repositories with a much smaller working set on disk.
| Disk question | Conventional server | walgit with remote reader |
|---|---|---|
| How much history must stay local? | Often most or all of it | Only what the cache and request mix demand |
| What happens to blobs? | They usually occupy server disk | They can stay in object storage and be fetched by range |
| What does a small node look like? | A constrained server that risks eviction | A thin cache that still serves real traffic |
| How do you grow? | Bigger disks, more careful backups | More nodes, smarter caching, same bucket |
That is why the architecture is compelling for large monorepos. It turns disk from a hard ceiling into a performance hint.
What walgit is really competing with
The real comparison is not Git versus Git. It is persistent-state hosting versus bucket-backed hosting. Traditional stacks optimize around keeping a machine healthy. walgit optimizes around keeping a storage substrate reachable.
| Question | Persistent Git hosting | Bucket-backed walgit |
|---|---|---|
| Where do you trust the truth? | A server, filesystem, or database | Object storage plus WAL |
| What fails first? | The box, then the repo | A node, which can be replaced |
| What is the ops burden? | Patch, back up, migrate, restore | Keep storage available and caches warm |
| Who pays for scale? | The machine and its disks | The object store and request layer |
| Best fit | Stateful legacy deployments | Cloud-native Git infrastructure with large repos |
That is a conservative win, not a universal one. If you want a small self-contained Git box, this is probably too much architecture. If you want a clean durability model for serious hosting, the design starts to look obvious.
Why this repo feels unusually sharp
The codebase signals production intent everywhere: modern Rust edition, strict lints, clear crate boundaries, metrics and tracing, and a split between protocol handling, storage, bundle logic, and maintenance. Nothing about it reads like a demo.
That is the final tell. walgit is interesting because the idea is elegant, but it is credible because the implementation respects the ugly parts of real systems: retries, caches, recovery, and the cost of heavy Git operations.