Skip to content

Rendering a streamed LLM answer without freezing the page

· 6 min read

Every AI chat app streams its answer. The text arrives a few characters at a time, and the frontend has to turn that into formatted markdown while it's still coming in.

The simplest version is to keep the text in state and render it with a markdown component. It works. I wanted to know what it costs, so I built a small lab to measure it: streaming-ui-lab.

The setupLink to section: The setup

There is no real model behind it. A fake stream sends a 10,000-character markdown answer (headings, lists, tables, code blocks) 5 characters every 20 ms. Every run gets the same input, and it costs nothing.

The same answer is rendered four ways:

  1. Naive. Every chunk goes into state. react-markdown renders the whole text again.
  2. Throttle. Chunks collect in a ref and state updates every 50 ms. This is what useChat({ throttle: 50 }) does in the AI SDK.
  3. Memoize. The text is split into blocks on blank lines (never inside a code fence). Each block is a React.memo component, so finished blocks don't render again. Only the last one does.
  4. Streamdown. Vercel's package for streaming markdown, which also memoizes blocks.

The page measures itself while it streams, so you don't need DevTools open:

  • Commits and render time, from React's Profiler.
  • Total blocking time, from the Long Animation Frames API. How long the page was frozen.
  • Input delay, from the Event Timing API, while I type into a text field during the stream.
  • FPS, from a requestAnimationFrame counter.

React Compiler is off. With it on, the naive version would get memoized automatically and stop being naive.

On a fast laptopLink to section: On a fast laptop

Three runs per mode, averaged:

ModeCommitsRender totalBlocking timeFPS (avg / min)
Naive196914.6 s16 ms60 / 58
Throttle7899.0 s0 ms60 / 59
Memoize19694.3 s0 ms60 / 60
Streamdown25926.3 s0 ms60 / 60

Memoize does a third of the work of naive. But look at the last two columns. The page never freezes and stays at 60 FPS in every mode. On my machine, a user would not notice any difference.

With the CPU slowed down 4xLink to section: With the CPU slowed down 4x

Then I turned on 4x CPU slowdown in DevTools, which is closer to a mid-range phone:

ModeRender totalBlocking timeWorst input eventFPS (avg / min)Time to show the full answer
Naive147.3 s68.9 s232 ms21 / 6168.6 s
Throttle125.5 s61.1 s72 ms23 / 6148.5 s
Memoize20.2 s1 ms64 ms60 / 5939.8 s
Streamdown49.0 s0.6 s96 ms46 / 464.3 s

The naive version falls apart. The page is frozen for 69 seconds in total, FPS drops to 6, and a keystroke takes 232 ms to show up, above the 200 ms that INP calls "good".

The lab during a naive run with 4x CPU slowdown, 30 seconds in: 971 commits, 19.7 s of render time, FPS 54 average and 42 minimum.
Naive mode, 4x CPU slowdown, 30 seconds in.

The last column needs a note. The answer takes 39 seconds to stream, but the naive version needs 168 seconds to show all of it. My fake stream runs on the same main thread, so it slows down too. With a real API the stream wouldn't slow down. The text would pile up and the screen would fall further and further behind the model.

React also started logging "Maximum update depth exceeded" in the naive mode. The AI SDK has a troubleshooting page for exactly this error, so it isn't just my demo.

Memoize barely notices the slowdown. Same 39.8 seconds as the stream, 60 FPS, 1 ms of blocking.

The lab during a memoize run with 4x CPU slowdown, 30 seconds in: 1492 commits, 4.8 s of render time, FPS 60 average and 58 minimum.
Memoize mode, same settings, also 30 seconds in.

Compare the commit counts in the two screenshots. Each commit is one chunk on the screen. After the same 30 seconds, naive has shown 971 chunks and memoize 1492. Naive is already a third of the way behind, and it has spent four times as long rendering to get there.

Why throttling isn't enoughLink to section: Why throttling isn't enough

Throttle is the usual advice, and it does help a little: fewer commits, and typing feels better. But the page still freezes almost as long as naive.

Each update is cheaper only in count, not in size. Every time state changes, react-markdown still parses the whole answer from the first character. A 10,000-character answer gets parsed hundreds of times. Throttling lowers how often you pay that cost. It doesn't make the cost smaller.

Memoize changes the cost itself. A finished paragraph's string never changes, so React.memo skips it. Each update only parses the block that is still being written, which is a few hundred characters at most.

The core of it is small:

tsx
const MemoBlock = memo(function MemoBlock({ content }: { content: string }) {
  return <Markdown>{content}</Markdown>;
});

export function MemoizeRenderer({ text }: { text: string }) {
  return splitBlocks(text).map((block, i) => <MemoBlock key={i} content={block} />);
}

splitBlocks is the only careful part. It must not split inside a code block, even one that contains blank lines.

StreamdownLink to section: Streamdown

I expected Streamdown to match my memoize version, since it uses the same idea. It was slower, and FPS dipped to 4 at the worst point. Streamdown also does things my version doesn't, like fixing half-finished markdown while it streams and hardening links. My guess is that this extra work is the difference, but I haven't profiled it in detail.

For most apps it's still the easier choice. You get the memoization, plus those fixes, without writing a block splitter yourself.

Screen readersLink to section: Screen readers

One more thing that's easy to miss. If the output is an aria-live region, a screen reader tries to announce every chunk, and the answer becomes noise.

The lab puts the output in role="log" with aria-live="off", so it is not announced while it streams. A separate, visually hidden role="status" element says "Response streaming" and then "Response ready". The user hears the state once and reads the answer when it's done.

What I took from itLink to section: What I took from it

  • On a fast laptop, rendering the whole answer on every chunk looks fine. Test with CPU slowdown before deciding it is.
  • Throttling cuts the number of renders, not the cost of each one.
  • Splitting the answer into memoized blocks fixed it, with a few lines of code.

These numbers come from development mode, so the absolute values are higher than in production. The gap between the modes is what matters, and it's large enough that I don't expect it to close.

The code is on GitHub and you can run the lab yourself at streaming-ui-lab.vercel.app.