Retrying Failed Background Jobs with Backoff
Retry failed MV3 background jobs safely: classify transient and permanent failures, exponential backoff with jitter on chrome.alarms, a durable attempt record, caps, dead-letter handling and user-visible status.
Table of Contents
The nightly sync failed because the laptop had no network, the export job hit a rate limit, the cache rebuild threw on a malformed record. A job that simply fails leaves data stale until its next scheduled run, which may be a day away; a job that retries in a tight loop hammers your server and the user’s battery. Background jobs in an extension need the same disciplined retry policy as any distributed system, implemented on top of the one scheduler that survives service worker termination — chrome.alarms. This guide builds a small, reusable retry layer for any job. It belongs to alarms and scheduled background jobs.
Why retries must be durable
A retry is a promise to run again later. In a web page, setTimeout(retry, 60_000) keeps that promise as long as the tab is open. In an MV3 service worker, it usually breaks it: thirty seconds of inactivity evicts the worker and every pending timer with it. The retry’s existence has to live outside the worker — an attempt record in chrome.storage.local saying which job failed, how many times, and when to try next — and the wake-up has to come from an alarm, which the browser persists and fires even if the worker is not running. With that split, the worker can die at any point and the retry still happens.
Step-by-step: a reusable retry wrapper
1. Make the job idempotent
1// jobs/export.js
2export async function exportJob() {
3 const { exportCursor = 0 } = await chrome.storage.local.get("exportCursor");
4 const items = await loadItemsSince(exportCursor);
5 for (const chunk of chunks(items, 100)) {
6 await api("/v1/export", { method: "PUT", body: { items: chunk } }); // PUT by id: safe to repeat
7 await chrome.storage.local.set({ exportCursor: chunk.at(-1).seq });
8 }
9}
Execution context: the service worker. A job that may run several times must produce the same result each time — upserting by id rather than appending, advancing a cursor only after each chunk succeeds. Without idempotency, retries turn one failure into duplicated data.
2. Classify failures
1// jobs/classify.js
2export function classify(err) {
3 if (err?.status === 429 || err?.status === 503) return { retry: true, after: err.retryAfterMs };
4 if (err?.status >= 500) return { retry: true };
5 if (err instanceof TypeError || err?.name === "TimeoutError") return { retry: true }; // network
6 if (err?.name === "QuotaExceededError") return { retry: false, reason: "quota" };
7 return { retry: false, reason: err?.message ?? "unknown" };
8}
Execution context: a pure module, easy to unit test. Network failures, timeouts, rate limits and server errors are transient — time may fix them. Validation errors, missing permissions and quota exhaustion are permanent — retrying produces the same failure and wastes resources. Honour Retry-After when the server sends it.
3. Wrap the job with durable attempt tracking
1// jobs/retry.js
2const BASE_MS = 60_000, CAP_MS = 6 * 3600_000, MAX_ATTEMPTS = 8;
3
4export async function runWithRetry(name, job) {
5 const key = `retry:${name}`;
6 const { [key]: rec = { attempt: 0 } } = await chrome.storage.local.get(key);
7 try {
8 await job();
9 await chrome.storage.local.remove(key);
10 await chrome.alarms.clear(key);
11 return "ok";
12 } catch (err) {
13 const c = classify(err);
14 const attempt = rec.attempt + 1;
15 if (!c.retry || attempt >= MAX_ATTEMPTS) {
16 await chrome.storage.local.set({ [key]: { ...rec, attempt, dead: true, lastError: String(err?.message ?? err) } });
17 return "dead";
18 }
19 const delay = c.after ?? Math.random() * Math.min(CAP_MS, BASE_MS * 2 ** attempt);
20 const nextAt = Date.now() + Math.max(30_000, delay);
21 await chrome.storage.local.set({ [key]: { attempt, nextAt, lastError: String(err?.message ?? err) } });
22 await chrome.alarms.create(key, { when: nextAt });
23 return "scheduled";
24 }
25}
Execution context: the service worker. The attempt record is written before the alarm is created, so even if the worker dies between the two calls, a startup check (step 5) can re-arm the alarm from the record. Full jitter — a random delay up to the exponential ceiling — prevents thousands of installs that failed at the same moment from retrying in lockstep. The thirty-second floor respects Chrome’s minimum for one-shot alarm delays.
4. Route retry alarms back to their jobs
1const JOBS = { export: exportJob, reindex: reindexJob, sync: syncJob };
2
3chrome.alarms.onAlarm.addListener(({ name }) => {
4 if (name.startsWith("retry:")) {
5 const job = JOBS[name.slice(6)];
6 if (job) runWithRetry(name.slice(6), job);
7 } else if (JOBS[name]) {
8 runWithRetry(name, JOBS[name]); // the job's normal schedule
9 }
10});
Execution context: the service worker, top-level listener. Both the regular schedule and the retries go through runWithRetry, so a regular run that succeeds clears any pending retry, and a regular run that fails starts the backoff sequence. One registry of job functions keeps the mapping from alarm names to code in one place.
5. Re-arm retries on startup
1chrome.runtime.onStartup.addListener(async () => {
2 const all = await chrome.storage.local.get(null);
3 for (const [key, rec] of Object.entries(all)) {
4 if (!key.startsWith("retry:") || rec.dead) continue;
5 const existing = await chrome.alarms.get(key);
6 if (!existing) await chrome.alarms.create(key, { when: Math.max(Date.now() + 30_000, rec.nextAt) });
7 }
8});
Execution context: the service worker. Alarms normally survive restarts, but an extension update clears them, and a worker that died between writing the record and creating the alarm leaves an orphaned record. Re-arming from records closes both gaps. Retries whose time passed while the browser was closed run shortly after startup.
6. Surface dead jobs instead of hiding them
1export async function deadJobs() {
2 const all = await chrome.storage.local.get(null);
3 return Object.entries(all)
4 .filter(([k, v]) => k.startsWith("retry:") && v.dead)
5 .map(([k, v]) => ({ job: k.slice(6), attempts: v.attempt, error: v.lastError }));
6}
Execution context: the service worker, called from the options page or popup via a message. A job that exhausted its retries or failed permanently should be visible: a badge, a line in the popup (“Export failed: storage quota exceeded — free up space or turn off export”), and a “Try again” button that clears the dead record and runs the job. Permanent failures usually need the user to act; hiding them guarantees they never will.
Common mistakes
setTimeoutfor retries. The worker is evicted and the retry never happens.- Retrying everything. Permanent errors retried eight times are eight wasted requests and eight error reports.
- No jitter. After your API’s outage, every install retries at the same instant and causes a second outage.
- No cap. Unbounded exponential growth schedules the next attempt days away.
- Non-idempotent jobs. Retries create duplicates. Fix the job before adding retries.
Cross-browser variation
- Chrome / Edge: one-shot alarms have a 30-second minimum delay in recent versions; periodic alarms a one-minute minimum in packed builds.
- Firefox: alarms are not clamped as aggressively; retries may fire sooner in testing than in Chrome. The same code works.
- Safari: alarms can be delayed significantly when the system is idle or on battery. Trigger a retry check when an extension page opens, so a user who looks at the extension sees fresh state.
Verification
- Make the API return 503 and run the job: a
retry:exportrecord withattempt: 1and a matching alarm appear. - Stop the worker; when the alarm fires, the job runs again and the attempt count increases.
- Make the API succeed: the record and alarm are removed.
- Make the job throw a permanent error and confirm the record is marked
deadand no alarm is created.
1await chrome.storage.local.get("retry:export");
2await chrome.alarms.get("retry:export");
Execution context: the service worker console.
FAQ
How does this differ from retrying network requests?
The same principles, at job granularity. Request-level retries inside a job handle quick blips; job-level retries handle longer outages. See handling offline and retrying requests.
Should the retry count reset after a success?
Yes — success removes the record entirely, so the next failure starts from attempt one.
Can several jobs retry at once?
Yes; each has its own record and alarm. If they share a backend, consider a shared circuit breaker that pauses all jobs during an outage.
Related
- Persisting job progress across worker restarts — resuming instead of restarting.
- Auditing scheduled alarms with getAll — keeping retry alarms tidy.
- Minimum alarm period and throttling — the floor on retry delays.
- Alarms and scheduled background jobs — the parent topic.