Handling Offline and Retrying Requests
Make extension requests resilient to flaky networks: classify retryable failures, exponential backoff with jitter on chrome.alarms, Retry-After, offline detection in a service worker and user-visible sync state.
Table of Contents
Laptops sleep, trains go through tunnels, captive portals intercept requests, and APIs return 503 during deploys. An extension that treats every failed request as final loses data; one that retries everything immediately hammers a struggling server and drains the battery. In MV3 there is an extra complication: the setTimeout that would schedule a retry in a web page does not survive the worker being evicted, so a retry scheduled for “in five minutes” usually never happens. This guide builds a retry policy that works with the service worker lifecycle. It sits under network requests and backend sync.
Why retries need durable scheduling
A retry has three parts: deciding whether a failure is worth retrying, deciding when, and actually running the retry at that time. The first two are pure logic. The third depends on the runtime surviving until the chosen moment. A web page usually does; an extension service worker usually does not — after thirty seconds without events it is terminated, and every pending timer disappears with it. The only scheduler that survives termination is chrome.alarms, which persists in the browser profile and wakes the worker when it fires. Its minimum period in packed Chrome builds is thirty seconds for one-shot when alarms in recent versions and one minute for periodic ones, so short backoffs (one to ten seconds) can use timers while the worker is demonstrably awake, and anything longer must use an alarm.
Step-by-step: a lifecycle-aware retry policy
1. Classify failures into retry decisions
1// net/classify.js
2export function classify(errOrResponse) {
3 if (errOrResponse instanceof Response) {
4 const s = errOrResponse.status;
5 if (s === 401) return { kind: "auth" };
6 if (s === 408 || s === 429 || s >= 500) {
7 const ra = errOrResponse.headers.get("Retry-After");
8 const after = ra ? (Number(ra) * 1000 || Date.parse(ra) - Date.now()) : null;
9 return { kind: "retry", after };
10 }
11 return { kind: "fatal", status: s };
12 }
13 if (errOrResponse?.name === "AbortError") return { kind: "cancelled" };
14 if (errOrResponse?.name === "TimeoutError" || errOrResponse instanceof TypeError) return { kind: "retry" };
15 return { kind: "fatal" };
16}
Execution context: a pure module used by the service worker, and easy to unit test in Node. fetch rejects with a TypeError for network-level failures — no connection, DNS failure, CORS rejection — so a TypeError is retryable unless your CORS setup is wrong, which you will find in development. Retry-After may be seconds or an HTTP date; both are handled.
2. Compute backoff with full jitter
1export function backoff(attempt, { base = 2_000, cap = 30 * 60_000 } = {}) {
2 const exp = Math.min(cap, base * 2 ** attempt);
3 return Math.round(Math.random() * exp); // "full jitter"
4}
Execution context: pure logic. Full jitter — a random delay between zero and the exponential ceiling — spreads retries from thousands of clients that all failed at the same moment, such as during your API’s deploy, so they do not return in synchronized waves. The cap keeps a long outage from pushing the next attempt days into the future.
3. Schedule short waits with timers and long waits with alarms
1// net/retry.js
2export async function scheduleRetry(jobId, delayMs) {
3 const { retries = {} } = await chrome.storage.local.get("retries");
4 retries[jobId] = { at: Date.now() + delayMs, attempt: (retries[jobId]?.attempt ?? 0) + 1 };
5 await chrome.storage.local.set({ retries });
6
7 if (delayMs < 20_000) {
8 setTimeout(() => runRetry(jobId), delayMs); // worker is awake now; good enough
9 }
10 // Always arm an alarm as the durable fallback
11 await chrome.alarms.create(`retry:${jobId}`, { when: Date.now() + Math.max(delayMs, 30_000) });
12}
13
14chrome.alarms.onAlarm.addListener(({ name }) => {
15 if (name.startsWith("retry:")) runRetry(name.slice(6));
16});
Execution context: the service worker, with the alarm listener registered at the top level. The persisted retries record is the source of truth; the timer is an optimisation for quick blips while the worker is awake, and the alarm guarantees the retry happens even if it is not. runRetry must be idempotent — check the record and skip if the job already succeeded — because both the timer and the alarm may fire. Firefox has no thirty-second floor on one-shot alarms; Safari may fire them late.
4. Detect offline cheaply and wait for connectivity
1export async function isProbablyOnline() {
2 if (!navigator.onLine) return false; // definitely offline
3 try {
4 const r = await fetch("https://api.acme.example/v1/ping", {
5 method: "HEAD", cache: "no-store", signal: AbortSignal.timeout(4_000),
6 });
7 return r.ok;
8 } catch { return false; } // captive portal, DNS, etc.
9}
Execution context: the service worker. navigator.onLine is available in workers, and false is reliable; true only means a network interface is up, which includes hotel Wi-Fi that intercepts every request. A cheap probe against your own API confirms real reachability. Do not register online/offline event listeners in the worker and expect to be woken by them — they are not extension events and will not start an evicted worker. Instead, have retries check reachability before running and re-arm themselves if offline.
5. Show the user what is pending
1export async function updateSyncIndicator() {
2 const { outbox = [], retries = {} } = await chrome.storage.local.get(["outbox", "retries"]);
3 const pending = outbox.length;
4 await chrome.action.setBadgeText({ text: pending ? "…" : "" });
5 await chrome.action.setTitle({
6 title: pending ? `${pending} change(s) waiting to sync` : "All changes synced",
7 });
8}
Execution context: the service worker, called after every enqueue and every successful send. Users tolerate offline gracefully when they can see that their work is safe and waiting; silent queues produce support tickets asking whether data was lost. Popups and side panels should show the same state from storage. After a long run of failures — say, eight attempts — surface a persistent message and stop automatic retries until the user acts or the next startup.
Cross-browser variation
- Chrome / Edge: one-shot alarms have a thirty-second minimum delay in recent versions; periodic alarms one minute.
navigator.onLineworks in the service worker. - Firefox: alarms are not clamped as aggressively, so short retries may arrive early in testing. Network errors also reject with
TypeError. - Safari: alarms may fire minutes late when the system is idle, and the background context is suspended quickly. Trigger a retry pass whenever an extension page opens, so a user who opens the popup sees work flushed immediately.
Verification
- Disconnect the network and save an item. Confirm a
retriesentry and aretry:*alarm inawait chrome.alarms.getAll(). - Stop the worker from
chrome://serviceworker-internals. - Reconnect. Within the alarm’s delay, the worker should start, probe, send the request and clear the retry record.
- Make the API return 503 with
Retry-After: 120and confirm the next attempt is scheduled about two minutes later, not sooner. - Make it return 400 and confirm no retry is scheduled and the error is surfaced.
FAQ
Why not just retry three times immediately?
Immediate retries only help with the briefest blips and make outages worse by tripling load at the worst moment. Backoff with jitter handles both blips and outages.
Should I retry POST requests?
Only if they are idempotent — carry an idempotency key the server deduplicates on, as in syncing extension data with a backend API. Otherwise a retry after a lost response creates duplicates.
How many retries are enough?
Enough to cover a typical outage — eight attempts with a thirty-minute cap spans a few hours — plus a retry on every startup. After that, tell the user rather than retrying forever.
Related
- Retrying failed background jobs with backoff — the same policy for non-network jobs.
- Minimum alarm period and throttling — the floor on alarm delays.
- Popup loading and empty states — showing pending state in the UI.
- Network requests and backend sync — the parent topic.