Adding a chat assistant to a web product looks easy in a demo: call a model API, print the answer. In production, users expect text to appear as it is generated, the assistant to look things up in your system, answers to survive a page reload, and the whole thing not to bankrupt you when a bot discovers the endpoint. This guide walks through building an AI chat in a Next.js App Router application with the Vercel AI SDK, from the first streamed token to a hardened production feature. The same patterns apply if you call provider SDKs directly.
Architecture in one picture
Browser (React, useChat)
│ POST /api/chat (messages)
▼
Next.js Route Handler (server)
├─ auth + rate limit
├─ load context (user, RAG)
├─ streamText(model, system, messages, tools)
│ └─ tools → your DB / APIs (with user's permissions)
└─ stream response (SSE) ──► Browser renders tokens as they arrive
Key principle: the model API key and all tool execution stay on the server. The browser only sends messages and receives a stream.
Why streaming matters
A typical answer of 300 tokens takes several seconds to generate. Without streaming, users stare at a spinner; with streaming, the first words appear within a second and perceived latency drops dramatically. Streaming over HTTP uses a long-lived response with chunks, commonly formatted as server-sent events. Next.js route handlers support streaming responses natively via the Web ReadableStream API (Next.js route handlers), and the AI SDK handles the protocol for you.
Streaming also affects your infrastructure: proxies and load balancers must not buffer responses, and serverless functions need a timeout long enough for the full answer.
The server: a route handler
// app/api/chat/route.ts
import { streamText, convertToModelMessages, stepCountIs, type UIMessage } from "ai"
import { anthropic } from "@ai-sdk/anthropic"
import { auth } from "@/lib/auth"
import { rateLimit } from "@/lib/rate-limit"
import { orderTools } from "@/lib/ai/tools"
export const maxDuration = 60 // seconds, for serverless platforms
export async function POST(req: Request) {
const session = await auth()
if (!session) return new Response("Unauthorized", { status: 401 })
const { success } = await rateLimit(`chat:${session.user.id}`, { limit: 30, window: "10m" })
if (!success) return new Response("Too many requests", { status: 429 })
const { messages }: { messages: UIMessage[] } = await req.json()
const result = streamText({
model: anthropic(process.env.CHAT_MODEL!), // model id from config, not hard-coded
system: SYSTEM_PROMPT,
messages: await convertToModelMessages(messages.slice(-20)), // cap history
tools: orderTools(session.user.id),
stopWhen: stepCountIs(5), // max tool-call rounds
maxOutputTokens: 800,
abortSignal: req.signal, // stop generating if the user navigates away
})
return result.toUIMessageStreamResponse()
}
Details that matter:
- Authentication first. An unauthenticated chat endpoint is an open proxy to a paid API.
- Rate limiting per user, and globally, before any model call.
- History cap. Sending the entire conversation on every turn grows cost quadratically; keep the last N messages or summarise older ones; see context engineering.
- Step limit on tool-calling loops and max output tokens.
- Abort signal so a closed tab stops generation and billing.
- Model ID in configuration so you can switch models without a deploy; see choosing an LLM.
The client: useChat
// app/(app)/assistant/chat.tsx
"use client"
import { useChat } from "@ai-sdk/react"
import { useState } from "react"
export function Chat() {
const { messages, sendMessage, status, stop, error } = useChat()
const [input, setInput] = useState("")
const busy = status === "submitted" || status === "streaming"
return (
<section aria-label="Assistant">
<ol aria-live="polite" className="space-y-4">
{messages.map((m) => (
<li key={m.id} data-role={m.role}>
{m.parts.map((part, i) =>
part.type === "text" ? <p key={i}>{part.text}</p> : null,
)}
</li>
))}
</ol>
{error && <p role="alert">Something went wrong. Please try again.</p>}
<form
onSubmit={(e) => {
e.preventDefault()
if (!input.trim() || busy) return
sendMessage({ text: input })
setInput("")
}}
>
<label htmlFor="chat-input" className="sr-only">Your question</label>
<textarea id="chat-input" value={input} onChange={(e) => setInput(e.target.value)} />
{busy ? (
<button type="button" onClick={stop}>Stop</button>
) : (
<button type="submit">Send</button>
)}
</form>
</section>
)
}
Messages are made of typed parts — text, tool calls, tool results, reasoning, files — so you can render tool activity ("Looking up order #4821…") instead of leaving users waiting in silence. A Stop button is not optional; users need control over long generations.
Rendering model output safely
Model output is untrusted input to your UI. If you render Markdown:
- Use a Markdown renderer that does not allow raw HTML, or sanitise its output.
- Do not auto-load images from model-provided URLs — this is a known data exfiltration channel in prompt injection attacks.
- Open links with
rel="noopener noreferrer"and consider an allow-list of domains.
These defences are part of the broader picture in AI agent security and the classic web risks in OWASP Top 10.
Tools: letting the assistant act on your data
Tools are server-side functions the model can call. Define them with Zod schemas and enforce permissions in code:
// lib/ai/tools.ts
import { tool } from "ai"
import { z } from "zod"
import { db } from "@/lib/db"
export function orderTools(userId: string) {
return {
getOrderStatus: tool({
description:
"Get status, items and delivery estimate for one of the CURRENT user's orders. " +
"Use when the user asks where their order is or what it contains.",
inputSchema: z.object({
orderNumber: z.string().regex(/^\d{4,8}$/).describe("Order number, digits only"),
}),
execute: async ({ orderNumber }) => {
const order = await db.order.findFirst({
where: { number: orderNumber, customerId: userId }, // permission check
select: { number: true, status: true, eta: true, items: { select: { name: true, qty: true } } },
})
return order ?? { error: `Order ${orderNumber} not found for this account.` }
},
}),
}
}
Note that userId comes from the session, never from the model. Tool design — names, descriptions, compact outputs, actionable errors — decides how well this works; see designing tools for LLM agents.
Structured output for UI features
Not every AI feature is a chat. For forms, filters or summaries that drive UI, generate typed objects instead of text:
import { generateObject } from "ai"
import { z } from "zod"
const { object } = await generateObject({
model: anthropic(process.env.FAST_MODEL!),
schema: z.object({
category: z.enum(["bug", "billing", "feature_request", "other"]),
summary: z.string().max(200),
urgency: z.number().int().min(1).max(3),
}),
prompt: `Classify this support message:\n<message>${text}</message>`,
})
streamObject streams partial objects for progressive UI. Schema design and validation are covered in structured outputs from LLMs.
Persistence and resumability
Users expect conversations to survive reloads and device switches:
- Store messages server-side (Postgres works fine) keyed by conversation and user, including parts and tool results.
- Save on completion using the SDK's finish callback, and save the user message immediately.
- Load initial messages in a Server Component and pass them to the client.
- Resumable streams: if the connection drops mid-answer, either resume from a server-side buffer or regenerate. For most products, showing the partial answer with a "Regenerate" button is enough.
Apply retention policies: chat logs contain personal data; see privacy for LLM apps.
Grounding answers in your content
A product assistant should answer from your documentation, not from the model's general knowledge. Add retrieval before generation — search your docs with the user's question and include the top passages in the system prompt or as a tool the model can call. The full RAG pipeline is in RAG for business.
Reliability and cost
- Retries and fallbacks: provider outages and 429s happen; retry with backoff and fall back to another model or provider for critical features; see LLM API reliability.
- Prompt caching: keep the system prompt and tool definitions stable and first to maximise cache hits.
- Model routing: a fast, cheap model for classification and short answers; a stronger one when tools or reasoning are needed.
- Budgets: per-user daily token limits and global alerts; see LLM cost optimization.
- Observability: trace every request with tokens, latency and tool calls; see LLM observability.
Accessibility and UX
AI chat is still a web UI, and the European Accessibility Act applies to many consumer products:
- Use an
aria-live="polite"region for new messages, but avoid announcing every token — announce on completion or in sentence-sized chunks. - Make sure the input has a visible or screen-reader label, and that Send and Stop are keyboard accessible.
- Show clear status: thinking, using a tool, generating, error.
- Tell users they are talking to an AI and what it can and cannot do — this is also an EU AI Act transparency requirement; see EU AI Act guide.
Watch performance too: rendering a re-parsed Markdown tree on every token can hurt INP. Throttle updates or memoise completed messages.
Deployment notes
- Edge vs Node runtime: Node is the safer default for database access and SDK compatibility.
- Buffering: if you use Nginx or another reverse proxy, disable response buffering for the chat route, otherwise streaming arrives all at once.
- Timeouts: align serverless
maxDuration, proxy timeouts and the model's maximum generation time. - Bot protection: chat endpoints attract abuse; combine auth, rate limits and, for public chats, a challenge for anonymous users.
FAQ
Do we have to use the Vercel AI SDK? No. It removes boilerplate for streaming, tools and React state, and supports many providers. Calling provider SDKs directly works fine; you then implement the streaming protocol and client state yourself.
Can we use this with a self-hosted model? Yes. Use an OpenAI-compatible provider pointed at your vLLM or Ollama endpoint; see self-hosting LLMs.
Should the chat run in a Server Action instead of a route handler? Route handlers are the better fit for streaming chat and are easier to rate-limit and observe. Server Actions work well for one-shot generation tied to forms.
How do we test it? Unit-test tools and permission checks like any server code, and run an eval set of real questions against the whole endpoint on every prompt or model change; see LLM evals.
Sources
- Vercel. AI SDK documentation.
- Next.js. Route Handlers (route.js).
- MDN. Server-sent events and ReadableStream.
- Anthropic. Streaming messages.
- MDN. ARIA live regions.
- OWASP. Top 10 for LLM Applications 2025.