vercel-labs/gemini-chatbot: When the chatbot becomes the interface
A Vercel Labs template that uses Gemini, tools, and typed React components to turn chat into a task flow instead of a transcript.
- gemini-chatbot treats the model as an interface orchestrator that can move a user from chat into a task-specific UI.
- The repo's real trick is the loop between structured tool output and React components that can trigger the next message.
- Its stack is opinionated on purpose, because Next.js, Gemini, the AI SDK, Zod, Drizzle, Postgres, Blob, and NextAuth form one contract.
- It is most useful as a blueprint for workflow-heavy products where the right answer is a state change, not a paragraph.
Most chatbot templates still think in transcripts. This one thinks in states. The user starts with a question, but the app's job is to decide when to answer in text, when to render a flight list, and when to push the conversation back into the model.
The chat box is only the doorway
That is the useful inversion here. In a normal chatbot, the chat surface is the product. In gemini-chatbot, it is just the entry point to a workflow. The template is built around a simple claim: an LLM is better at choosing the next interface than it is at writing every answer as prose.
The demo speaks the language of flights because booking is inherently multi-step. Search, choose, reserve, pay. A transcript is the wrong shape for that work, which is why this repo matters as a pattern rather than as a flight app.
How gemini-chatbot turns model output into interface state
The control loop runs through app/(chat)/api/chat/route.ts. The route streams a model response, exposes tools, and keeps the system prompt focused on a flow rather than open-ended chat. When a tool resolves, the UI does not stringify the result and move on. It maps the result to a React component in components/custom/message.tsx, then hands the user another action surface.
const result = streamText({
model: google('gemini-2.5-pro'),
messages,
tools: {
searchFlights,
getWeather,
createReservation
}
});
if (toolInvocation.state === 'result') {
return <ListFlights results={toolInvocation.result} append={append} />;
}
That split matters. The model is not asked to invent UI markup. It returns structured intent, Zod keeps the shape honest, and React renders the right component. The conversation stays conversational, but the interface becomes task-aware.
Why the stack is so opinionated
This repo is a Vercel Labs project in the best sense. It chooses a stack that makes the handoff reliable: Next.js App Router for routing and streaming, the Vercel AI SDK for chat primitives, Gemini for model calls, Drizzle and Postgres for persistence, Blob for attachments, and NextAuth for identity. None of those pieces are flashy alone. Together they make generative UI feel like a normal web app instead of a science project.
prepareTools in @ai-sdk/google drops all function tools when provider-defined tools (e.g., googleSearch) are present. The function returns early with only provider tools, so custom function declarations are never sent to the Gemini API.
The project also exposes a practical model split. Use Gemini Pro when the app needs orchestration and decision-making. Use Gemini Flash when the job is structured object generation, like mock flight data. That is not a gimmick. It is cost control and latency control baked into the architecture.
The visual metaphor is the point. The baton is not just a prompt. It is a contract. One model decides what should happen next, and another model populates the UI with the typed data that makes that decision usable.
The clever part is not the model. It is the handoff
The best line in the repo is not written in prose, it is expressed in behavior. A flight card is not a dead end. It is an input device. Click it, and the app can append a new message, preserving the conversation while changing the user's state of play.
That feedback loop is what turns the app from a demo into a pattern. The user sees a flight, clicks it, and the UI generates the next message back into the chat history. The interface is doing orchestration work, not just presentation work.
This is a hard block for any application that needs both Google Search grounding and custom function calling in the same agent.
That matters because orchestration is only as good as the tools beneath it. The repo is opinionated partly because the surrounding stack still has rough edges, and it is cleaner to show one coherent path than to pretend every model provider behaves the same.
What this template replaces
| Option | What it gives you | Best for | Trade-off |
|---|---|---|---|
| gemini-chatbot | A full Next.js template with streaming chat, auth, storage, and generative UI wiring. | Teams that want a working reference instead of a blank slate. | Opinionated around the Vercel stack and the flight-booking demo. |
| Vercel AI SDK core | The primitives for chat, tools, and streaming without the surrounding app shell. | Builders who want maximum control over architecture. | You still have to assemble the UI, auth, persistence, and patterns. |
| Mastra | An agent framework with workflow abstractions around tools and model calls. | Teams centered on agent coordination. | Less of a drop-in UI template, more of an orchestration framework. |
| Plain text chatbot | The fastest way to get a conversational demo live. | Simple Q&A and support flows. | It stops at the transcript, so the user still has to act elsewhere. |
So who should start here? Teams that want a full-stack reference with Vercel-native defaults and a clear generative UI pattern. If you only need the chat primitives, use the AI SDK directly. If you are centered on agent workflows more than UI composition, an agent framework may fit better. If you want the simplest possible support bot, a transcript-first template is enough.
Why it matters beyond flights
Flights are a narrow demo, but the pattern is wider. Any product where the answer is a choice, a form, a filtered list, or a reservation can benefit from this shape. The model should not only speak. It should route the user into the right interface, then hand them back when the task needs another decision.
gemini-chatbot makes that transition feel ordinary, which is exactly why it is useful. The future of conversational software is not fewer UI components. It is more deliberate ones, each summoned by the model only when text is the wrong tool.