Rivet and the Architecture of the Infinite Memory

How a 0.6KB Actor model is eliminating the database round-trip for AI agents and real-time apps.

8 min read • View on GitHub • More from rivet-dev

A tiny, exhausted messenger carrying a massive backpack labeled CONTEXT at a toll booth labeled DATABASE HANDSHAKE. Represents the overhead of stateless architecture.
The modern web request cycle pays a heavy latency tax just to remind the server what it was doing.
Portrait of Nathan Flurry

Every API call to a cloud provider is a dependency you can't control.

Key Takeaways

The Stateless Tax

Modern web architecture is defined by the stateless struggle. A request comes in, the server wakes up, and it immediately faces amnesia. To do any useful work, it must fetch state from a remote database, process the logic, save the result back across the network, and die.

Every interaction pays this latency tax. When building AI agents that require constant context updates or real-time collaborative applications, this database round-trip becomes a severe bottleneck. Developers spend more time orchestrating external state than writing core business logic.

The 0.6KB Sovereign

Rivet flips this model by making the Actor the primary unit of both compute and storage. An actor is a lightweight, stateful process. It holds its memory locally and hibernates when idle.

Instead of a monolithic server juggling thousands of stateless requests, Rivet provisions a dedicated, microscopic server for every single user or session. These actors are remarkably dense. A single instance consumes roughly 0.6KB of memory. They achieve cold starts in 20 milliseconds and offer zero-millisecond read latency because the data is already in memory.

A single translucent marble containing a tiny coiled clockwork mechanism and a microscopic library of books, sitting on a velvet cushion.
The Rivet Actor encapsulates compute and local state into a single, highly dense atomic unit.
A flow diagram showing the Wake-on-Request cycle in Rivet. It features five main nodes: Client

Gasoline: Execution That Never Forgets

At the core of Rivet is a durable workflow engine called Gasoline. When an actor wakes up, it does not start from scratch. It replays its history.

Gasoline uses a coordinate-based location system to track execution progress. Think of it as a Dewey Decimal system for code steps. If a process crashes or a server is restarted, Gasoline rewinds to the exact coordinate of the failure. It fast-forwards through previously completed, cached steps and resumes execution flawlessly.

A vertical timeline diagram illustrating Gasoline's Coordinate Replay system. The timeline shows steps 1.0

Guard: Routing Without a Map

Managing millions of tiny, transient actors requires a radically different approach to networking. Rivet uses a gateway called Guard to handle routing and proxying.

Instead of maintaining a massive, globally synchronized map of every actor's location, Guard uses optimistic routing. It assumes its local cache is correct. It tunnels the request directly to the expected node. If it hits a failure, Guard uses that signal to trigger fresh service discovery. This eliminates a massive amount of overhead.

The Agent's Body

AI agents are the killer application for this architecture. A stateless LLM call is just a prompt. An agent is a long-running process with state, memory, and agency.

Rivet gives the agent a body. The state and the compute live in the same isolated environment. The agent can hibernate for days, wake up instantly when a new event occurs, and access its entire memory without querying an external Postgres instance.

Rivet Actors solve this by unifying state and logic in a single actor definition. Instead of coordinating between databases and APIs, state and behavior live together

Nathan Flurry, Co-founder & CTO. AI & User Generated Backends with Rivet
Feature Stateless (Lambda + DB) Stateful (Rivet Actor)
Architecture Request-Response Actor-Based
State Latency 50-200ms (Network) 0ms (Local SQLite)
Persistence Manual (SQL/ORM) Automatic (Hibernation)
Scaling Per-Request Per-Instance

Sources: