vercel-labs/gemini-chatbot: When the chatbot becomes the interface

A Vercel Labs template that uses Gemini, tools, and typed React components to turn chat into a task flow instead of a transcript.

10 min read • View on GitHub • More from vercel-labs

A chat bubble opens like a theater curtain to reveal a flight-booking interface underneath. The image shows that the conversation is only the entry point, while the real product is the task-specific UI that appears after the model decides what to do next.
The chat surface is not the destination. It is the doorway to a workflow.
Key Takeaways

Most chatbot templates still think in transcripts. This one thinks in states. The user starts with a question, but the app's job is to decide when to answer in text, when to render a flight list, and when to push the conversation back into the model.

The chat box is only the doorway

That is the useful inversion here. In a normal chatbot, the chat surface is the product. In gemini-chatbot, it is just the entry point to a workflow. The template is built around a simple claim: an LLM is better at choosing the next interface than it is at writing every answer as prose.

The demo speaks the language of flights because booking is inherently multi-step. Search, choose, reserve, pay. A transcript is the wrong shape for that work, which is why this repo matters as a pattern rather than as a flight app.

How gemini-chatbot turns model output into interface state

The app is a loop, not a one-shot prompt. User input becomes orchestration, orchestration becomes a tool call, and the rendered component can send the next message back into chat.

The control loop runs through app/(chat)/api/chat/route.ts. The route streams a model response, exposes tools, and keeps the system prompt focused on a flow rather than open-ended chat. When a tool resolves, the UI does not stringify the result and move on. It maps the result to a React component in components/custom/message.tsx, then hands the user another action surface.

const result = streamText({
  model: google('gemini-2.5-pro'),
  messages,
  tools: {
    searchFlights,
    getWeather,
    createReservation
  }
});

if (toolInvocation.state === 'result') {
  return <ListFlights results={toolInvocation.result} append={append} />;
}

That split matters. The model is not asked to invent UI markup. It returns structured intent, Zod keeps the shape honest, and React renders the right component. The conversation stays conversational, but the interface becomes task-aware.

Why the stack is so opinionated

This repo is a Vercel Labs project in the best sense. It chooses a stack that makes the handoff reliable: Next.js App Router for routing and streaming, the Vercel AI SDK for chat primitives, Gemini for model calls, Drizzle and Postgres for persistence, Blob for attachments, and NextAuth for identity. None of those pieces are flashy alone. Together they make generative UI feel like a normal web app instead of a science project.

prepareTools in @ai-sdk/google drops all function tools when provider-defined tools (e.g., googleSearch) are present. The function returns early with only provider tools, so custom function declarations are never sent to the Gemini API.

isac322, GitHub Issue Author · vercel/ai issue #13911

The project also exposes a practical model split. Use Gemini Pro when the app needs orchestration and decision-making. Use Gemini Flash when the job is structured object generation, like mock flight data. That is not a gimmick. It is cost control and latency control baked into the architecture.

A relay race between two engines, one large and deliberate, the other compact and fast, passes a baton across a clean white field. The image explains how the template routes reasoning to one model and structured object generation to another.
Different jobs get different models. The expensive model reasons, the faster one fills in structured details.

The visual metaphor is the point. The baton is not just a prompt. It is a contract. One model decides what should happen next, and another model populates the UI with the typed data that makes that decision usable.

The clever part is not the model. It is the handoff

The best line in the repo is not written in prose, it is expressed in behavior. A flight card is not a dead end. It is an input device. Click it, and the app can append a new message, preserving the conversation while changing the user's state of play.

A close-up of a hand selecting a flight card while a second hand types into a chat box connected by a taut thread. The image shows how a UI action becomes the next chat message, which is the core feedback loop in the template.
The user does not leave the interface to continue the task. The interface feeds the next message back into the conversation.

That feedback loop is what turns the app from a demo into a pattern. The user sees a flight, clicks it, and the UI generates the next message back into the chat history. The interface is doing orchestration work, not just presentation work.

This is a hard block for any application that needs both Google Search grounding and custom function calling in the same agent.

isac322, GitHub Issue Author · vercel/ai issue #13911

That matters because orchestration is only as good as the tools beneath it. The repo is opinionated partly because the surrounding stack still has rough edges, and it is cleaner to show one coherent path than to pretend every model provider behaves the same.

What this template replaces

OptionWhat it gives youBest forTrade-off
gemini-chatbotA full Next.js template with streaming chat, auth, storage, and generative UI wiring.Teams that want a working reference instead of a blank slate.Opinionated around the Vercel stack and the flight-booking demo.
Vercel AI SDK coreThe primitives for chat, tools, and streaming without the surrounding app shell.Builders who want maximum control over architecture.You still have to assemble the UI, auth, persistence, and patterns.
MastraAn agent framework with workflow abstractions around tools and model calls.Teams centered on agent coordination.Less of a drop-in UI template, more of an orchestration framework.
Plain text chatbotThe fastest way to get a conversational demo live.Simple Q&A and support flows.It stops at the transcript, so the user still has to act elsewhere.

So who should start here? Teams that want a full-stack reference with Vercel-native defaults and a clear generative UI pattern. If you only need the chat primitives, use the AI SDK directly. If you are centered on agent workflows more than UI composition, an agent framework may fit better. If you want the simplest possible support bot, a transcript-first template is enough.

Why it matters beyond flights

Flights are a narrow demo, but the pattern is wider. Any product where the answer is a choice, a form, a filtered list, or a reservation can benefit from this shape. The model should not only speak. It should route the user into the right interface, then hand them back when the task needs another decision.

gemini-chatbot makes that transition feel ordinary, which is exactly why it is useful. The future of conversational software is not fewer UI components. It is more deliberate ones, each summoned by the model only when text is the wrong tool.