Streaming Responses with fetch in the Service Worker

Stream LLM and search responses in an MV3 extension: read fetch bodies incrementally, parse server-sent events without EventSource, forward chunks over a port, cancel on close and stay inside lifetime limits.

Published October 2, 2026 Updated October 2, 2026 8 min read
Table of Contents

The extension asks an AI or search API for an answer and the user stares at a spinner for twenty seconds, then sees the whole answer at once — or, worse, nothing, because the worker was terminated while waiting for a response that took too long to start. Streaming APIs solve both problems: they start sending within a second and keep sending until done. In a page you would reach for EventSource; in an extension service worker it does not exist, and the streaming has to be done by hand with fetch. This guide shows how, and how to keep it inside MV3’s lifetime rules. It belongs to network requests and backend sync.

Why streaming matters more in a worker

Chrome terminates an extension service worker when a fetch takes longer than about thirty seconds to begin returning its response, and an event’s total work is expected to finish within five minutes. A non-streaming request to a slow model can easily take longer than thirty seconds to produce its first byte, because the server buffers the entire answer before sending headers. A streaming request returns headers immediately and then delivers data continuously, so the “response started” condition is met at once and every arriving chunk counts as activity. Streaming is therefore not just a UX nicety in MV3 — for slow endpoints it is the difference between working and being killed.

Buffered versus streamed response in the workerA buffered response sends nothing for 35 seconds and the worker is terminated at 30; a streamed response sends headers at 0.4 seconds and chunks continuously until done at 35 seconds.request sent35 sHeaders0.4 sChunks arrivingeach one is activityDone[DONE] e…first chunk resets the idle timer30 s: a buffered request would be cut off here
Same total time — but only the streamed request ever reaches the user.

Step-by-step: stream from worker to UI

1. Open a port from the UI

The UI needs to receive many messages for one request, which is what ports are for.

1// sidepanel.js
2const port = chrome.runtime.connect({ name: "answer" });
3const out = document.querySelector("#answer");
4port.onMessage.addListener((m) => {
5  if (m.type === "chunk") out.append(m.text);
6  if (m.type === "done") out.dataset.state = "done";
7  if (m.type === "error") out.dataset.state = "error";
8});
9port.postMessage({ type: "ask", prompt: document.querySelector("#q").value });

Execution context: the side panel (or popup). append with a string inserts a text node, so streamed content can never inject markup. When the panel closes, the port disconnects automatically — the signal the worker uses to cancel. Firefox and Safari support the same port API.

2. Start the streaming fetch in the worker

 1// sw.js
 2chrome.runtime.onConnect.addListener((port) => {
 3  if (port.name !== "answer") return;
 4  const ctrl = new AbortController();
 5  port.onDisconnect.addListener(() => ctrl.abort());
 6  port.onMessage.addListener(async (m) => {
 7    if (m.type !== "ask") return;
 8    try {
 9      await streamAnswer(m.prompt, port, ctrl.signal);
10      port.postMessage({ type: "done" });
11    } catch (err) {
12      if (err.name !== "AbortError") port.postMessage({ type: "error", message: err.message });
13    }
14  });
15});

Execution context: the service worker. Tying an AbortController to the port means closing the panel cancels the upstream request — important when the API bills per token. An open port from an extension page also keeps the worker alive for the duration of the stream.

3. Read the body as a stream of text

 1async function streamAnswer(prompt, port, signal) {
 2  const res = await fetch("https://api.acme.example/v1/answer", {
 3    method: "POST",
 4    signal,
 5    headers: { "Content-Type": "application/json", Accept: "text/event-stream",
 6               Authorization: `Bearer ${await getAccessToken()}` },
 7    body: JSON.stringify({ prompt, stream: true }),
 8  });
 9  if (!res.ok || !res.body) throw new Error(`HTTP ${res.status}`);
10  const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
11  for await (const event of sseEvents(reader)) {
12    if (event.data === "[DONE]") break;
13    const delta = JSON.parse(event.data).delta ?? "";
14    if (delta) port.postMessage({ type: "chunk", text: delta });
15  }
16}

Execution context: the service worker. TextDecoderStream handles multi-byte UTF-8 characters split across network chunks — decoding each chunk separately with new TextDecoder().decode(chunk) corrupts them. The [DONE] sentinel and the delta field follow a common LLM API convention; adapt to your API’s shape.

The streaming pipelineNetwork bytes pass through a TextDecoderStream into a server-sent-events parser that yields events; each event's delta is posted over the port to the side panel, which appends it as text.res.bodyUint8Array chunksTextDecoderStreamUTF-8 safesseEvents()split on blank linesone message per eventJSON.parse(data)delta textport.postMessage{type:'chunk'}append(text)no innerHTML
Decode as a stream, parse as a stream, render as text — nothing is buffered whole.

4. Parse server-sent events yourself

EventSource is unavailable in service workers and cannot send POST bodies anyway. The format is simple: events are separated by a blank line, and each data: line contributes to the event’s data.

 1async function* sseEvents(reader) {
 2  let buffer = "";
 3  for (;;) {
 4    const { value, done } = await reader.read();
 5    if (done) break;
 6    buffer += value;
 7    let sep;
 8    while ((sep = buffer.search(/\r?\n\r?\n/)) !== -1) {
 9      const raw = buffer.slice(0, sep);
10      buffer = buffer.slice(sep).replace(/^\r?\n\r?\n/, "");
11      const event = { event: "message", data: "" };
12      for (const line of raw.split(/\r?\n/)) {
13        if (line.startsWith(":")) continue;                 // comment / keep-alive
14        const [field, ...rest] = line.split(":");
15        const val = rest.join(":").replace(/^ /, "");
16        if (field === "data") event.data += (event.data ? "\n" : "") + val;
17        else if (field === "event") event.event = val;
18        else if (field === "id") event.id = val;
19      }
20      if (event.data) yield event;
21    }
22  }
23}

Execution context: the service worker. The buffer handles events split across chunks — the normal case, not an edge case. Lines starting with : are keep-alive comments that many servers send during long pauses; ignoring them is correct, and their arrival still counts as network activity that keeps the fetch alive.

5. Guard against long silences and huge answers

1const IDLE_LIMIT_MS = 25_000;
2let lastChunk = Date.now();
3const watchdog = setInterval(() => {
4  if (Date.now() - lastChunk > IDLE_LIMIT_MS) ctrl.abort(new Error("stream stalled"));
5}, 5_000);
6// update lastChunk on every event; clearInterval(watchdog) when finished

Execution context: the service worker. If the server pauses mid-stream for longer than the worker’s limits allow, it is better to abort cleanly and tell the user than to be terminated silently. Cap total output, too — a runaway response held in the side panel’s DOM can grow to megabytes. Persist the final answer to chrome.storage.session so reopening the panel shows it without a new request.

6. Render incrementally without thrashing the UI

A fast model can emit hundreds of small deltas per second. Appending each one as its own text node is safe, but re-rendering Markdown or re-running syntax highlighting on every delta makes the panel stutter and burns CPU in both the worker and the page. Batch on the UI side with requestAnimationFrame, and format only when the stream is done.

 1// sidepanel.js
 2let pending = "";
 3let scheduled = false;
 4port.onMessage.addListener((m) => {
 5  if (m.type !== "chunk") return;
 6  pending += m.text;
 7  if (!scheduled) {
 8    scheduled = true;
 9    requestAnimationFrame(() => { out.append(pending); pending = ""; scheduled = false; });
10  }
11});

Execution context: the side panel, which has a document and therefore requestAnimationFrame. At most one DOM update per frame keeps scrolling smooth however fast chunks arrive. When the done message arrives, run any expensive formatting once over the final text — and sanitise it if the formatter produces HTML, as described in preventing XSS in extension pages.

Cross-browser variation

  • Chrome / Edge: streaming bodies, TextDecoderStream, async iteration over readers via the generator above, and AbortController are all available in extension service workers. The thirty-second first-response rule makes streaming essential for slow endpoints.
  • Firefox: supports the same streaming primitives in its MV3 background. Its event page is less strict about response-start timing, so non-streaming requests may appear to work in Firefox and fail in Chrome.
  • Safari: supports streaming fetch and TextDecoderStream in recent versions; background contexts are suspended aggressively, so keep a port from an open extension page for the whole stream.
Streaming building blocks by engineAvailability of readable response bodies, TextDecoderStream, EventSource and AbortSignal in the background contexts of Chrome, Firefox and Safari extensions.PrimitiveChromeFirefoxSafarires.body streamYesYesYesTextDecoderStreamYesYesYes (recent)EventSource in backgroundNo (worker)Event page onlyNoAbort on port closeYesYesYes
EventSource is the one piece missing everywhere it matters — parse SSE by hand.

Verification

  1. Open the side panel and ask a question that produces a long answer. Text should begin appearing within about a second.
  2. In the worker’s DevTools Network panel, the request should show a growing size while “Pending”, and an EventStream tab listing events.
  3. Close the side panel mid-answer. The request should change to “(canceled)” within a moment.
  4. Point the client at a test endpoint that sends one event and then pauses for 40 seconds; confirm the watchdog aborts and the panel shows an error rather than hanging.

FAQ

Can I stream to a content script instead of a side panel?

Yes — open the port from the content script. Render with text nodes only; streamed model output is untrusted and the page’s DOM is not yours to inject markup into.

Why not do the fetch in the side panel directly?

You can, if the panel has host permission for the API. Doing it in the worker centralises tokens and lets a request continue if you later support detaching the UI. Either works; the parsing code is identical.

Does streaming reduce cost?

Not by itself, but cancelling on close does. Without the abort, a user who closes the panel still pays for the full answer.

Other Core APIs & Cross-Browser Data Management Resources