Building a Minimal Error Collection Endpoint
Collect extension errors without a third-party service: a small HTTP endpoint that validates, rate-limits and stores reports, a client that batches and sends from the service worker, privacy-safe payloads, and simple querying and alerting.
Table of Contents
The team wants to know when the extension breaks for users, but the privacy policy promises that no data goes to third parties, and a full error-monitoring SaaS is more than a small extension needs. A minimal self-hosted endpoint — one HTTP route that accepts a small JSON report, validates it, rate-limits it and stores it — plus a client that batches and sends reports from the service worker covers most of the value: you learn which errors happen, how often, in which version and browser. This guide builds both halves, with the privacy and abuse protections a public endpoint needs. It belongs to error monitoring and crash reporting.
The design in one paragraph
Every context forwards errors to the service worker, which normalises them, deduplicates, and queues them in chrome.storage.local. An alarm flushes the queue every few minutes as one batched POST to your endpoint. The endpoint accepts only a strict schema with size limits, rejects unknown fields, rate-limits per anonymous installation ID and per IP, aggregates identical errors into counters, and stores them in a database you can query. No page URLs, page content or user identifiers are collected. Because the extension is public, assume the endpoint will receive junk and abuse — validation and limits are not optional.
Step-by-step: endpoint and client
1. Define a strict report schema
1// shared/report-schema.js
2export const LIMITS = { maxReports: 50, message: 300, stack: 4000, field: 40 };
3
4export function validReport(r) {
5 return r && typeof r === "object"
6 && typeof r.fp === "string" && r.fp.length <= 128
7 && typeof r.message === "string" && r.message.length <= LIMITS.message
8 && typeof r.stack === "string" && r.stack.length <= LIMITS.stack
9 && ["sw", "popup", "options", "sidepanel", "content", "offscreen"].includes(r.context)
10 && /^\d+(\.\d+){0,3}$/.test(r.version)
11 && ["chrome", "edge", "firefox", "safari", "other"].includes(r.browser)
12 && Number.isInteger(r.count) && r.count >= 1 && r.count <= 10_000
13 && Object.keys(r).every((k) => ["fp", "message", "stack", "context", "version", "browser", "count"].includes(k));
14}
Execution context: a module shared by client and server. A closed schema — only known fields, with bounded sizes and enumerated values — rejects junk and prevents the endpoint from becoming a general-purpose data dump. No URL, user or page fields exist, so they cannot be sent by mistake.
2. Queue reports in the service worker
1// sw/error-queue.js
2export async function enqueue(err, context) {
3 const report = {
4 fp: await sha256(fingerprint(err, context)),
5 message: normaliseMessage(err.message).slice(0, 300),
6 stack: normaliseStack(err.stack).slice(0, 4000),
7 context,
8 version: chrome.runtime.getManifest().version,
9 browser: detectBrowser(),
10 count: 1,
11 };
12 const { errQueue = {} } = await chrome.storage.local.get("errQueue");
13 const existing = errQueue[report.fp];
14 errQueue[report.fp] = existing ? { ...existing, count: Math.min(existing.count + 1, 10_000) } : report;
15 if (Object.keys(errQueue).length <= 50) await chrome.storage.local.set({ errQueue });
16}
17
18chrome.runtime.onMessage.addListener((m, sender) => {
19 if (m.type === "report-error") enqueue(m.error, m.context);
20});
Execution context: the service worker. Keying the queue by fingerprint deduplicates on the client: a thousand identical errors become one report with count: 1000. Capping the queue at 50 distinct errors bounds storage. Normalisation removes extension IDs, URLs and long numbers, as in grouping and deduplicating extension errors.
3. Flush batches with an alarm
1chrome.alarms.create("flush-errors", { periodInMinutes: 15 });
2chrome.alarms.onAlarm.addListener(async ({ name }) => {
3 if (name !== "flush-errors") return;
4 const { errQueue = {}, installId } = await chrome.storage.local.get(["errQueue", "installId"]);
5 const reports = Object.values(errQueue);
6 if (!reports.length) return;
7 try {
8 const res = await fetch("https://errors.readable.example/v1/reports", {
9 method: "POST",
10 headers: { "Content-Type": "application/json" },
11 body: JSON.stringify({ install: installId, reports }),
12 credentials: "omit",
13 });
14 if (res.ok || res.status === 400) await chrome.storage.local.remove("errQueue"); // 400 = never retry bad data
15 } catch { /* offline — keep for next flush */ }
16});
Execution context: the service worker. Batching every 15 minutes keeps requests rare. installId is a random UUID generated at install — anonymous, but stable enough for per-install rate limiting and “how many installs are affected” counts. credentials: "omit" ensures no cookies are sent. Failed sends stay queued; a 400 means the data will never be accepted, so drop it rather than retrying forever. Respect the user’s choice: if they opted out of diagnostics, never enqueue. See reporting errors without breaking your privacy policy.
4. Implement the endpoint
1// server.js (Node, any framework; shown with plain http + SQLite)
2import http from "node:http";
3import Database from "better-sqlite3";
4import { validReport } from "./shared/report-schema.js";
5
6const db = new Database("errors.db");
7db.exec(`CREATE TABLE IF NOT EXISTS errors (fp TEXT, version TEXT, browser TEXT, context TEXT, message TEXT, stack TEXT,
8 day TEXT, events INTEGER, installs INTEGER, PRIMARY KEY (fp, version, browser, day))`);
9const upsert = db.prepare(`INSERT INTO errors VALUES (@fp,@version,@browser,@context,@message,@stack,@day,@count,1)
10 ON CONFLICT DO UPDATE SET events = events + @count, installs = installs + 1`);
11
12const buckets = new Map(); // naive in-memory limiter: key → {n, reset}
13const limited = (key, max) => { const now = Date.now(), b = buckets.get(key) ?? { n: 0, reset: now + 3600e3 };
14 if (now > b.reset) { b.n = 0; b.reset = now + 3600e3; } b.n++; buckets.set(key, b); return b.n > max; };
15
16http.createServer(async (req, res) => {
17 if (req.method !== "POST" || req.url !== "/v1/reports") return res.writeHead(404).end();
18 let size = 0, chunks = [];
19 for await (const c of req) { size += c.length; if (size > 65_536) return res.writeHead(413).end(); chunks.push(c); }
20 let body; try { body = JSON.parse(Buffer.concat(chunks)); } catch { return res.writeHead(400).end(); }
21 if (typeof body.install !== "string" || !/^[0-9a-f-]{36}$/.test(body.install) || !Array.isArray(body.reports) || body.reports.length > 50
22 || !body.reports.every(validReport)) return res.writeHead(400).end();
23 if (limited(`i:${body.install}`, 12) || limited(`ip:${req.socket.remoteAddress}`, 120)) return res.writeHead(429).end();
24 const day = new Date().toISOString().slice(0, 10);
25 db.transaction(() => body.reports.forEach((r) => upsert.run({ ...r, day })))();
26 res.writeHead(204).end();
27}).listen(8080);
Execution context: a small server behind HTTPS (a reverse proxy or platform TLS). The endpoint stores aggregates per fingerprint, version, browser and day — events and an approximate installs count — rather than raw reports, which bounds growth and avoids keeping per-install histories. The in-memory limiter is enough for a single instance; use a shared store if you scale out. Do not log request bodies or IP addresses beyond what rate limiting needs.
5. Query what matters
1-- Top errors in the latest version, last 7 days
2SELECT fp, message, context, browser, SUM(events) AS events, SUM(installs) AS installs
3FROM errors WHERE version = '2.4.0' AND day >= date('now', '-7 day')
4GROUP BY fp, browser ORDER BY installs DESC LIMIT 20;
Execution context: the database. Ranking by affected installs rather than raw events surfaces bugs that hit many users over loops that hit one. Comparing versions shows regressions; see alerting on error spikes after a release.
6. Delete old data automatically
Run a daily job that deletes rows older than your retention period (for example 90 days), and state that period in your privacy policy. Aggregates older than that rarely help debugging.
7. Declare the endpoint
Add the endpoint’s origin to host_permissions or rely on CORS (Access-Control-Allow-Origin for extension origins). Mention error reporting, what it contains and how to turn it off in your privacy policy and store listing’s data disclosures. See writing a privacy policy for an extension.
Common mistakes
- Accepting arbitrary JSON. The endpoint becomes a dumping ground.
- Sending each error immediately. Many requests; batch instead.
- Retrying 400s forever. Drop invalid data.
- Storing raw reports with IPs. Privacy risk and unbounded growth.
- No opt-out. Users and reviewers expect one.
Cross-browser variation
- Chrome / Edge: alarms flush at a minimum 30-second interval;
fetchfrom the worker with host permission or CORS. - Firefox: same approach;
moz-extension://origins differ per install, so allow them by scheme pattern in CORS rather than by exact origin. - Safari: same client code; Safari may restrict background activity more, so also flush when the popup opens.
Verification
- Throw the same error 100 times and confirm one report with
count: 100is sent. - Post a report with an extra field and confirm a 400.
- Exceed the per-install limit and confirm 429s.
- Confirm no URLs or user data appear in stored rows.
FAQ
Is a self-hosted endpoint more private than a SaaS?
It can be, because you control collection and retention — but only if you design it that way.
How do I get readable stack traces?
Store normalised stacks and symbolicate them with your private source maps when viewing, as in debugging production builds with source maps.
Can attackers send fake errors?
Yes; the endpoint is public. Validation and rate limits keep that from harming you, and aggregates by installs reduce the influence of any single sender.
When should I move to a hosted service?
When you need features like release health, alerting workflows and team triage at scale. The client-side normalisation and privacy rules carry over unchanged.
Related
- Capturing uncaught errors in every context — feeding the queue.
- Grouping and deduplicating extension errors — fingerprints.
- Privacy-preserving usage metrics — the same principles for metrics.
- Error monitoring and crash reporting — the parent topic.