observe: the open-source observability stack that measures AI margin, not just tokens

It turns model calls, Stripe revenue, and cached responses into one question: which customers and features actually pay for themselves?

8 min read • View on GitHub • More from katrinalaszlo

A large balance scale sits at the center of a ledger. On one side, streams of AI traces and token counters fall into a cost column, while on the other side Stripe receipts and revenue slips gather into a matching column. The image explains that this project treats observability as a profitability problem, not just a debugging problem.
observe changes the unit of analysis from usage to unit economics. The interesting number is not only what a trace costs, but whether it leaves enough margin behind.
Key Takeaways

The dashboard that asks if your AI feature is profitable

Most AI observability tools stop at tokens, latency, and errors. observe asks a harsher question: did this trace earn more than it cost? That shift matters for SaaS teams that now treat model calls as a line item, not a novelty.

The project’s core move is simple and unusual. It puts usage, model pricing, cache behavior, and revenue data into the same frame, so a product team can see customer profitability, margin-compressing features, and expensive flows instead of just raw traffic.

What observe measures that other tools miss

This is not just LLM monitoring with a billing tab bolted on. The codebase is built around revenue-aware analysis, which means the dashboard can answer questions that ordinary trace tools leave open: which customers are unprofitable, which features burn margin, and which requests only look healthy until you attach revenue to them.

Product typePrimary questionData sourcesWhat it misses
observeWhich traces and customers make money?Proxy, SDK, CSV, pricing, StripeNothing in the billing loop
Token observability toolWhat did the model call cost?LLM traces, latency, tokensRevenue, customer margin, pricing context
Generic observability stackIs the system healthy?Logs, metrics, tracesBusiness outcome and unit economics
Warehouse + BIHow profitable is the feature?Billing exports, logs, warehouse tablesSlow setup and no live trace context
Three different input streams, a proxy request lane, a manual SDK event lane, and a CSV backfill stack, all feed into the same ledger page. The image explains that observe normalizes different sources into one economics model instead of keeping them in separate tools.
Proxy traffic, SDK events, and CSV backfills are not separate products here. They are three entrances into the same ledger, which is what makes margin analysis possible.

How usage becomes a margin number

The mechanics are more interesting than the headline. In server/routes/proxy.ts, the proxy layer catches model traffic and can attach metadata before the request moves on. In packages/sdk, teams can instrument events manually when they want more control than a transparent proxy gives them.

Three ingestion paths, one economics ledger. The diagram shows how observe merges proxy traffic, SDK events, and CSV backfills before calculating cost, revenue, and margin.

From there, server/model-pricing.ts assigns cost, including override support for negotiated rates, while cache accounting tracks the savings from repeat calls. The final join happens when revenue data enters the same model, so the system can compute gross margin per trace instead of stopping at spend.

Why the product feels like a SaaS control plane

The repo does not read like a telemetry toy. server/billing.ts includes feature access logic, usage limits, and bonus credit mechanics, which means the product is thinking about customer lifecycle and monetization as part of the core system, not as an afterthought.

That makes sense for the category. If the point is to help teams decide whether an AI feature pays for itself, then billing, access control, sample data, and onboarding all belong in the same product surface. The whole stack starts to look less like an observability tool and more like a control plane for an AI SaaS business.

The left side shows a generic dashboard with token counts, latency charts, and red error dots. The right side shows a margin dashboard with customer profitability, feature cost, and a clear red flag on an expensive flow. The image explains why the project shifts the conversation from system health to business outcome.
The competitive difference is not a bigger chart. It is a different question. On one side, you monitor the model. On the other, you monitor the business result.

Where it sits in the landscape

Compared with token observability tools, observe is answering a different question. Compared with generic observability platforms, it is narrowing the field to the economics of AI features. Compared with a warehouse plus BI stack, it gets to the answer faster because the trace, the cost, and the revenue can meet before they disappear into separate systems.

That is the real differentiator. The project is not trying to become the broadest telemetry platform. It is trying to make the profit math legible while the product is still running.

What the codebase says about the team

The structure is unusually coherent for a broad project. A TypeScript-heavy backend, a Vue frontend, a dedicated SDK, and an orchestration package all point in the same direction: one team is building the data path, the UI, and the business rules together instead of stitching them in after the fact.

The commit pattern suggests a small, fast-moving core with occasional assisted contributions, which often produces the kind of sharp product opinion this repo has. That shows up in the details. Sample data is built to teach the product, not just fill an empty screen. Pricing is flexible enough for enterprise overrides. Feature access is part of the monetization story. The result is a repo that feels complete because the business logic is part of the architecture.