← Back to blog
[nextjs]October 5, 2026· 4 min read

How to Stream LLM Answers in Next.js 15 Without Breaking the UI

Learn practical patterns for streaming LLM responses in Next.js 15, handling data streams, UI state, and edge‑case bugs that trip up most developers.

#streaming#llm#react#nextjs15#performance

Why streaming matters and where it breaks

Modern AI apps are judged by the latency you expose to the user. A spinner that lasts 5‑10 seconds feels like a bug, while a word‑by‑word reveal feels instant. In Next.js 15 the new app router lets you return a ReadableStream directly, but the devil is in the details: back‑pressure, partial JSON, and React Server Components (RSC) all interact in subtle ways.

Below I walk through a production‑ready setup that avoids the three most common pitfalls:

  • Returning a raw Response instead of a NextResponse that strips the body.
  • Mixing server‑only streams with client‑side state that expects a complete payload.
  • Neglecting error handling inside the stream, which leaves the UI hanging.

Server‑side route: keep it pure

The API route lives under app/api/chat/route.ts. It receives a JSON payload, forwards it to the AI SDK, and pipes the result back as a ReadableStream. The key is to return result.toDataStreamResponse() directly—no extra wrapper, no json() call.

import { openai } from "@ai-sdk/openai";
import { streamText, convertToCoreMessages } from "ai";

export async function POST(req: Request) {
  const { messages } = await req.json();

  const result = streamText({
    model: openai("gpt-4o-mini"),
    messages: convertToCoreMessages(messages),
    onFinish: async ({ text }) => {
      // Persist the final answer for history
      await saveMessage({ role: "assistant", content: text });
    },
  });

  // Return the raw stream; Next.js will set the proper headers
  return result.toDataStreamResponse();
}

Notice the onFinish hook runs after the stream ends. If you need to update a DB mid‑stream you have to do it manually; the SDK only calls the hook once.

Client component: consume the stream safely

On the client you can't just fetch and response.json() because the body never resolves. Instead, read the ReadableStream and append chunks to state. React's useEffect is perfect for this, but you must guard against component unmounts.

import { useState, useEffect } from "react";

export default function ChatBox({ messages }) {
  const [answer, setAnswer] = useState("");
  const [loading, setLoading] = useState(false);

  useEffect(() => {
    let abort = false;
    const fetchStream = async () => {
      setLoading(true);
      const res = await fetch("/api/chat", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ messages }),
      });
      const reader = res.body?.getReader();
      const decoder = new TextDecoder();
      while (reader) {
        const { done, value } = await reader.read();
        if (done) break;
        if (abort) return;
        setAnswer((prev) => prev + decoder.decode(value, { stream: true }));
      }
      setLoading(false);
    };
    fetchStream();
    return () => { abort = true; };
  }, [messages]);

  return (
    <div>
      {loading && <p>Thinking…</p>}
      <p>{answer}</p>
    </div>
  );
}

The abort flag prevents state updates after the component unmounts, a subtle bug that otherwise throws "Can't perform a React state update on an unmounted component".

Handling partial JSON and errors

Some LLM providers stream data: {"content":"..."} lines. If you forward that verbatim, the client sees raw JSON fragments. The easiest fix is to strip the prefix server‑side:

const result = streamText({
  // ...
  // transform each chunk before sending
  onChunk: ({ chunk }) => chunk.replace(/^data: /, ""),
});

If the provider aborts early, the stream ends without onFinish. Wrap the fetch in a try/catch and send an SSE‑style error payload so the UI can show a fallback.

Performance knobs you shouldn't ignore

Streaming is cheap, but you still pay for the round‑trip each time you hit the API. Cache the messages array in a useRef and only send new user input. Also, enable compression in next.config.js to gzip the stream; browsers decompress on the fly without extra latency.

Finally, remember to set Cache-Control: no-store on the response to avoid CDN caching a partially‑filled stream.

Takeaway

Streaming LLM output in Next.js 15 is straightforward once you keep three rules in mind:

  1. Return the raw ReadableStream from the route.
  2. Read the stream chunk‑by‑chunk on the client, guarding against unmounts.
  3. Normalize provider‑specific payloads and surface errors early.

Follow these patterns and the UI will feel like a live conversation, not a loading screen.

// related