Your host, proxy, or React Native setup buffers or breaks streamed responses. Or you want one plain response per run. Send the run as one JSON body. useChat works the same: messages, tool calls, approvals, and errors.
Run the chat to its end with stream: false, then send the result with toJsonResponse():
import { chat, chatParamsFromRequest, toJsonResponse } from "@tanstack/ai";
import { openaiText } from "@tanstack/ai-openai";
export async function POST(request: Request) {
const { messages, threadId, runId } = await chatParamsFromRequest(request);
const result = await chat({
adapter: openaiText("gpt-5.6"),
messages,
threadId,
runId,
stream: false,
});
console.log(result.text);
return toJsonResponse(result);
}result.text is the full reply. result.chunks holds every chunk of the run, and the client reads them from the response.
You can also pass the stream from chat(). Then pass request.signal, so that the run stops when the client leaves:
import { chat, chatParamsFromRequest, toJsonResponse } from "@tanstack/ai";
import { openaiText } from "@tanstack/ai-openai";
export async function POST(request: Request) {
const { messages, threadId, runId } = await chatParamsFromRequest(request);
const stream = chat({
adapter: openaiText("gpt-5.6"),
messages,
threadId,
runId,
});
return toJsonResponse(stream, { signal: request.signal });
}The body is { chunks, offset?, done }. Without durability, done is always true, and the body holds the whole run.
Pass fetchJson() to useChat:
import { useChat, fetchJson } from "@tanstack/ai-react";
export function Chat() {
const { messages, sendMessage, isLoading } = useChat({
connection: fetchJson("/api/chat"),
});
return (
<div>
{messages.map((message) => (
<p key={message.id}>
{message.parts.map((part) => (part.type === "text" ? part.content : ""))}
</p>
))}
<button disabled={isLoading} onClick={() => void sendMessage("Hello")}>
Send
</button>
</div>
);
}fetchJson sends the same POST body as fetchHttpStream. It takes the same options: headers, body, credentials, fetchClient, and dynamic functions. See Request Options.
A long agent run can take more time than your host lets one request live. Add durability, and the server replies before the timeout. The client then asks for the rest.
import {
chat,
chatParamsFromRequest,
memoryStream,
resumeJsonResponse,
toJsonResponse,
} from "@tanstack/ai";
import { openaiText } from "@tanstack/ai-openai";
export async function POST(request: Request) {
const { messages, threadId, runId } = await chatParamsFromRequest(request);
const stream = chat({ adapter: openaiText("gpt-5.6"), messages, threadId, runId });
return toJsonResponse(stream, {
durability: { adapter: memoryStream(request) },
signal: request.signal,
maxWaitMs: 10_000,
});
}
// fetchJson polls GET ?runId=...&offset=... until the reply has done: true.
export async function GET(request: Request) {
return resumeJsonResponse({ adapter: memoryStream(request) });
}What happens:
resumeJsonResponse waits at most 1000 ms for new chunks before it replies. It never calls the model.
The client from step 2 needs no change. To set the wait between polls (default 250 ms), pass pollIntervalMs:
import { fetchJson } from "@tanstack/ai-react";
const connection = fetchJson("/api/chat", { pollIntervalMs: 500 });Pass this connection to useChat.
memoryStream(request) keeps the log in one process. If your requests run on many processes, use a production adapter. See Resumable Streams.
Send a message. When the run ends, the full reply appears in messages, with no stream between your server and the browser.