Prexo: The Zero-Marginal-Cost Architecture for AI Support Agents
How an opinionated TypeScript monorepo uses Edge functions, Redis streams, and free-tier model routing to bypass the LLM API tax.

We’re developing AI agents tailored for sales and customer support. If you follow our documentation closely, you can operate this service almost entirely for free—99% free, in fact!
- Prexo solves the linear scaling costs of AI customer support by routing requests to high-quality free-tier models via OpenRouter.
- The backend relies on Vercel Edge functions and Hono to maintain a lightweight, low-latency footprint.
- Instead of just returning text, Prexo's agents use tool calling to render functional React components directly in the chat UI.
- Heavy vectorization tasks are offloaded to asynchronous Upstash Redis streams to prevent Edge runtime timeouts.
The API Tax on Customer Support
The fundamental flaw in most AI customer service bots is economic, not technical. Every user interaction incurs a compute cost that eats directly into margins. If a bot gets wildly popular, the company goes broke paying OpenAI.
Prexo is an architectural solution to this exact trap. It is a multi-tenant platform designed to deploy AI agents that handle customer interactions without the crushing overhead.
Orchestrating the Edge
To achieve near-zero operating costs, Prexo abandons heavy Node.js servers. The backend is built on Hono and deployed to the Vercel Edge. This ensures sub-millisecond routing and minimal infrastructure overhead.
A custom dynamic cost algorithm acts as the traffic cop. It tracks request windows and thresholds, routing traffic through Unkey for rate-limiting and Upstash Redis for caching. The actual heavy lifting is offloaded to free-tier LLMs like DeepSeek via OpenRouter.
The UI is the Tool
Most chatbots are confined to text bubbles. Prexo treats the user interface as a tool the LLM can wield. In the sdk/ai.ts implementation, the agent is granted capabilities like sendCreateProjectForm.
The output is not a string of markdown. It is a structured JSON command that Next.js intercepts to render a functional React component directly in the user's chat window.
The Asynchronous Vectorizer
Retrieval-Augmented Generation ingestion is computationally expensive. Running it synchronously on the Edge guarantees timeouts. Prexo solves this with an asynchronous worker architecture.
When new knowledge base data arrives, a webhook pushes it into an Upstash Redis Stream named vectorizer-stream. This protects the fragile Edge runtime, allowing a separate background process to consume the stream and perform the heavy embedding tasks safely.
The Vertical Scalpel vs. The Swiss Army Knife
The AI framework ecosystem is dominated by sprawling generalists. LangChain attempts to be a toolkit for every conceivable use case. Botpress offers massive visual flow builders for enterprise teams.
Prexo takes the opposite approach. It is a highly opinionated, vertical monorepo designed to do exactly one thing: deploy cheap, UI-rich support agents.
| Feature | Prexo | LangChain | Botpress |
|---|---|---|---|
| Focus | Vertical (Support/Sales) | Horizontal (Anything) | Enterprise Bots |
| Architecture | Edge + Monorepo | Heavy Server | Hosted SaaS |
| Cost Model | Aggressive Free-Tier Routing | Bring-Your-Own-Keys | Tiered Subscription |
| UI Paradigm | Generative React Components | Text & JSON | Visual Flow Builder |