Persisting Job Progress Across Worker Restarts
Make long MV3 background jobs resumable: checkpoint cursors in chrome.storage, process in bounded batches, use leases to prevent overlap, and resume from the last checkpoint after the service worker is terminated.
Table of Contents
A job processes 20,000 bookmarks to rebuild a search index. Around item 12,000 the service worker is terminated — an idle timeout between batches, the five-minute event limit, a browser restart — and the next run starts again from item 1. On a slow machine it never finishes. The problem is not the worker’s lifetime, which you cannot change, but the job’s assumption that it runs from start to finish in one go. Long jobs in MV3 must be written as a sequence of small, resumable steps that record their progress durably. This guide shows the pattern. It belongs to alarms and scheduled background jobs.
Why a job loses its place
The service worker gives each extension event a time budget and terminates the worker after about thirty seconds without events. A job that awaits a long chain of storage reads, network requests and computations may outlive the event that started it, at which point the worker can be torn down mid-loop. Everything in local variables — the loop index, the partially built index, the list of remaining items — vanishes. Nothing errors; the job just stops. The fix is to treat the job as a state machine whose state lives in storage: a cursor marking how far it got, a batch size small enough to finish comfortably within one event’s budget, and a scheduler that keeps calling “do the next batch” until the cursor reaches the end.
Step-by-step: checkpoint, batch, resume
1. Describe the job as durable state
1// jobs/reindex-state.js
2export const initialState = () => ({
3 id: crypto.randomUUID(),
4 phase: "scan", // scan → build → publish
5 cursor: null, // opaque position within the phase
6 processed: 0,
7 startedAt: Date.now(),
8});
Execution context: a shared module. Everything the job needs to continue lives in this object, stored under one key. A job id lets you distinguish a resumed job from a new one and ignore stale work. Phases split a job whose steps have different shapes — scanning bookmarks, building an index, swapping it in — so each can have its own cursor semantics.
2. Process one bounded batch per step
1// jobs/reindex.js
2const BATCH = 500;
3
4export async function step() {
5 const { reindexJob: job } = await chrome.storage.local.get("reindexJob");
6 if (!job) return "idle";
7
8 if (job.phase === "scan") {
9 const { items, next } = await readBookmarks(job.cursor, BATCH);
10 const partial = await buildPartialIndex(items);
11 const { indexDraft = {} } = await chrome.storage.local.get("indexDraft");
12 mergeInto(indexDraft, partial);
13 const done = next === null;
14 await chrome.storage.local.set({
15 indexDraft,
16 reindexJob: { ...job, cursor: next, processed: job.processed + items.length, phase: done ? "publish" : "scan" },
17 });
18 return "continue";
19 }
20
21 if (job.phase === "publish") {
22 const { indexDraft } = await chrome.storage.local.get("indexDraft");
23 await chrome.storage.local.set({ index: indexDraft, indexBuiltAt: Date.now() });
24 await chrome.storage.local.remove(["indexDraft", "reindexJob"]);
25 return "done";
26 }
27}
Execution context: the service worker. Writing the partial results and the advanced cursor in one chrome.storage.local.set call keeps them consistent: either both are saved or neither is, so a termination never leaves results that the cursor does not account for. Choose a batch size that finishes in a few seconds on a slow machine; the next batch costs only a storage read to resume. Building into a draft and publishing at the end means readers never see a half-built index.
3. Drive steps from events, not from one long loop
1// sw.js
2export async function pump(trigger) {
3 if (!(await acquireLease("reindex", 60_000))) return; // another worker instance is running it
4 try {
5 for (let i = 0; i < 10; i++) { // a few batches per wake-up
6 const r = await step();
7 if (r !== "continue") return;
8 }
9 await chrome.alarms.create("reindex-step", { when: Date.now() + 30_000 }); // more later
10 } finally {
11 await releaseLease("reindex");
12 }
13}
14
15chrome.alarms.onAlarm.addListener(({ name }) => name === "reindex-step" && pump("alarm"));
16chrome.runtime.onStartup.addListener(() => pump("startup"));
Execution context: the service worker. Each wake-up runs a handful of batches, then hands off to an alarm for the next round, so no single event approaches the five-minute limit. Startup resumes an interrupted job automatically. Popups or messages can also call pump to make progress while the user is watching.
4. Prevent overlapping runs with a lease
1export async function acquireLease(name, ttlMs) {
2 const key = `lease:${name}`;
3 const now = Date.now();
4 const { [key]: lease } = await chrome.storage.local.get(key);
5 if (lease && lease.expires > now) return false;
6 const mine = { owner: crypto.randomUUID(), expires: now + ttlMs };
7 await chrome.storage.local.set({ [key]: mine });
8 const { [key]: check } = await chrome.storage.local.get(key);
9 return check?.owner === mine.owner;
10}
11
12export async function releaseLease(name) {
13 await chrome.storage.local.remove(`lease:${name}`);
14}
Execution context: the service worker. Within one worker, an in-memory flag would suffice; the lease protects against overlap across a quick restart, or between a spanning and split instance, where two contexts might both try to advance the job. The expiry means a worker that died holding the lease cannot block the job forever. This is a cooperative lock, not a strict one — keep steps idempotent so a rare overlap is harmless.
5. Report progress to the UI
1// popup.js
2const { reindexJob } = await chrome.storage.local.get("reindexJob");
3if (reindexJob) {
4 progress.value = reindexJob.processed;
5 progress.max = await chrome.runtime.sendMessage({ type: "bookmarks:count" });
6}
7chrome.storage.onChanged.addListener((c, area) => {
8 if (area === "local" && c.reindexJob?.newValue) progress.value = c.reindexJob.newValue.processed;
9});
Execution context: the popup or options page. Because progress lives in storage, any page can display it and update live through storage.onChanged, with no messaging protocol needed. Opening the popup can also call pump so the job advances faster while the user is waiting for it.
6. Handle schema changes between versions
1chrome.runtime.onInstalled.addListener(async ({ reason }) => {
2 if (reason !== "update") return;
3 const { reindexJob } = await chrome.storage.local.get("reindexJob");
4 if (reindexJob && reindexJob.version !== JOB_VERSION) {
5 await chrome.storage.local.remove(["reindexJob", "indexDraft"]); // restart under new code
6 await chrome.storage.local.set({ reindexJob: { ...initialState(), version: JOB_VERSION } });
7 }
8});
Execution context: the service worker after an update. A job started by the previous version may have a cursor format the new code does not understand. Tag job state with a version and restart incompatible jobs rather than misreading their cursors.
Common mistakes
- One long
forloop over everything. It works on fast machines and silently stops on slow ones. - Saving the cursor before the results. A termination between the two writes skips a batch forever. Write both in one call.
- Batches sized for your laptop. Measure on a slow machine; aim for a few seconds per batch.
- No lease. Two instances advancing the same job double-process batches or corrupt the draft.
- Publishing partial results. Build into a draft and swap at the end so readers see either the old or the new data.
Cross-browser variation
- Chrome / Edge: strict worker lifetimes make this pattern necessary;
storage.localwrites of a few hundred kilobytes per batch are fast. - Firefox: the event page lives longer, so a long loop may appear to work — and then fail under memory pressure. The pattern is still correct.
- Safari: background contexts are suspended aggressively; smaller batches and more frequent checkpoints help, and alarms may be delayed, so also pump from extension pages.
Verification
- Start the job and, mid-way, stop the worker from
chrome://serviceworker-internals. - Trigger the next step (wait for the alarm or open the popup) and confirm
processedcontinues from where it stopped. - Compare the final index with one built in a single uninterrupted run: identical.
- Start two triggers simultaneously and confirm the lease allows only one to process.
FAQ
How big can a batch be?
Small enough to finish well within thirty seconds even on slow hardware, and with writes small enough not to stall storage. A few hundred items per batch is typical.
Can I use an offscreen document to avoid all this?
Offscreen documents have their own lifetime rules tied to their stated reason and are not a general way to run long jobs. Make the job resumable instead.
What if the job’s input changes while it runs?
Use a cursor based on a stable ordering (sequence numbers, creation time) and accept that items added during the run are picked up by the next run.
Related
- Chaining alarms for long-running jobs — scheduling the next step.
- Retrying failed background jobs with backoff — handling steps that fail.
- Rebuilding in-memory state after termination — the same idea for caches.
- Alarms and scheduled background jobs — the parent topic.