MiniTrue: The Leaderless IoT Database That Gossips Its Way to Consistency

A look at how `aphrodoe/minitrue` combines gossip, Merkle trees, consistent hashing, and custom columnar storage into a compact distributed time-series system.

8 min read • View on GitHub • More from aphrodoe

A ring of equal database nodes receives the same stream of sensor signals while small gossip pulses pass between neighbors. In the center, a state fragment splits into matching halves, showing how the cluster reconciles drift without a single leader. It explains why MiniTrue’s main story is coordination, not just storage.
MiniTrue’s core idea is simple to say and hard to do: let peers stay in sync without elevating one node above the rest.
Key Takeaways

Most small databases pick a lane. They either get the ingestion path right or they get the distributed story right. MiniTrue tries to do both, and that is what makes it worth a closer look.

Built in Go, the repo treats cluster membership, routing, storage, and querying as one connected system. That matters because telemetry workloads punish weak seams: if peers cannot agree on ownership, the write path slows down. If the storage format is generic, the query path pays for it later.

Why MiniTrue Is Interesting

The unusual part is not that MiniTrue stores time-series data. It is that it does so in a leaderless peer model, where nodes gossip with each other, compare Merkle roots, and repair drift without waiting on a master. That puts it closer to the design instincts of mature distributed systems than most compact hobby databases.

The other half of the project is storage. MiniTrue does not lean on a generic row store for telemetry. It batches, separates fields into columns, and compresses timestamps and values with a time-series-friendly encoding strategy, which is exactly the kind of specialization that makes sensor data cheaper to keep and faster to read.

A Cluster That Keeps Itself Honest

MiniTrue’s membership layer is built around a simple idea: peers should continuously check one another, not wait for a central registry to declare truth. Nodes exchange gossip messages, track heartbeats, and mark members as suspect or down when they stop hearing from them.

The interesting part of the control plane is that no single node has to coordinate the whole cluster. Gossip keeps everyone updated, and Merkle trees make repair selective instead of expensive.

The Merkle tree is the clever part. Instead of shipping the entire cluster state over and over, nodes compare root hashes first. If they differ, the tree opens only where the divergence lives, which keeps synchronization bandwidth low and avoids turning every membership check into a full-state dump.

Consistent hashing sits beside that machinery. MiniTrue uses it to decide who owns a key and who should hold replicas, so routing and repair share the same map. That is a good sign in a distributed system: fewer separate rules means fewer hidden inconsistencies.

A telemetry record moves through a tight pipeline from MQTT message to batch buffer to separated columns, then into a compact storage block. On the right, the same block is queried for one metric and only the relevant pieces are decoded. It explains how the storage engine turns ingest into compressed analytical access.
MiniTrue’s storage story is not just “writes to disk.” It is a pipeline that reshapes telemetry into something cheaper to compress and faster to query.

How a Write Becomes a Row, Then a Column

Telemetry enters through MQTT, which is a practical fit for IoT devices because it is lightweight and widely supported. From there, MiniTrue batches records before flushing them to storage, which reduces disk churn and makes the write path more efficient than forcing every event straight to disk.

The interesting move is the format itself. Instead of storing each record as a mixed row, MiniTrue separates timestamps, device IDs, metric names, and values into columns. That is the standard trick behind analytical time-series systems, because similar values compress better when they sit together.

The repo’s custom `.parq` format pushes that idea further with Gorilla-style compression. Timestamps can be delta-encoded, and numeric values can be compressed by exploiting how little they change between adjacent samples. The result is a storage layer tuned for repetitive telemetry rather than general-purpose document storage.

// Simplified from the storage path
batch := storage.NewBatch(10)
for msg := range mqttStream {
    record := storage.Record{
        Timestamp: msg.Timestamp,
        DeviceID:  msg.DeviceID,
        Metric:    msg.Metric,
        Value:     msg.Value,
    }
    batch.Add(record)

    if batch.Ready() {
        engine.Flush(batch)
        batch.Reset()
    }
}

That tradeoff shows up immediately. Columnar storage is great for aggregation and scan-heavy queries, but it makes some mutations harder. The repo’s own shape suggests that deletes and restart edge cases are still part of the rougher operational surface, which is normal for an early-stage distributed engine.

The Query Path: One Request, Many Peers

MiniTrue’s read path follows the same decentralized logic as its write path. When a query arrives, the receiving node does not act like a permanent coordinator. It becomes a temporary coordinator, hashes the key, figures out which peers own the relevant data, and fans out requests to them.

The partial results come back as aggregates such as sum, count, min, and max, then the node merges them into a single answer. That is a clean fit for time-series analytics, because many dashboard questions are naturally reducible to small summaries rather than full row scans.