Alerting on Error Spikes After a Release
Detect bad extension releases fast: comparing error rates per version, normalising by active installs, new-issue and spike alerts, using staged rollouts to limit damage, and a runbook for halting or rolling back a release.
Table of Contents
Version 2.4.0 reaches users on Tuesday afternoon. By Wednesday morning, the support inbox has thirty emails saying highlights disappeared, the store rating has dropped half a star, and the error dashboard — which nobody was watching — shows a new error affecting 18% of installs since the update. Extensions auto-update silently, so a bad release reaches the whole user base within days, and there is no “roll back” button in most stores. The only defence is to notice within hours and stop the rollout. This guide sets up release-aware error alerts and the runbook to act on them. It belongs to error monitoring and crash reporting.
What a release spike looks like
After a release, the new version’s share of installs grows over a day or two as browsers check for updates (roughly every few hours in Chrome). Raw error counts for the new version will rise simply because more users run it, so counts alone mislead. The useful signals are rates: errors (or, better, affected installs) per active install of each version, compared with the previous version over the same period. A healthy release has a similar rate; a bad one has a jump, or a new issue — a fingerprint never seen before — affecting a meaningful share of the new version’s installs. Alerts on those two signals, scoped to the newest version, catch almost every bad release.
Step-by-step: release-aware alerting
1. Tag every error and heartbeat with the version
1const VERSION = chrome.runtime.getManifest().version;
2Sentry.init({ release: `readable@${VERSION}` }); // or your own endpoint's version field
Execution context: every reporting context. Version tagging is the foundation; without it, there is nothing to compare. Use exactly the manifest version so it matches store data and uploaded source maps. See uploading source maps for readable stack traces.
2. Count active installs per version
1// sw.js — one anonymous heartbeat per day, if the user allows diagnostics
2chrome.alarms.create("heartbeat", { periodInMinutes: 24 * 60, delayInMinutes: 5 });
3chrome.alarms.onAlarm.addListener(async ({ name }) => {
4 if (name !== "heartbeat" || !(await diagnosticsAllowed())) return;
5 await fetch("https://metrics.readable.example/v1/heartbeat", {
6 method: "POST", credentials: "omit",
7 body: JSON.stringify({ version: VERSION, browser: detectBrowser() }),
8 });
9});
Execution context: the service worker. An error rate needs a denominator. A daily anonymous heartbeat with only version and browser gives active installs per version without identifying anyone. Hosted error services often provide this as “release health” or session tracking; use that if you already have it. See counting active users without tracking.
3. Compute rates per version
1-- installs reporting errors / active installs, per version, last 24 h
2WITH active AS (SELECT version, COUNT(*) AS installs FROM heartbeats WHERE at > now() - interval '24 hours' GROUP BY version),
3 erroring AS (SELECT version, SUM(installs) AS affected FROM errors WHERE day >= current_date - 1 GROUP BY version)
4SELECT a.version, a.installs, COALESCE(e.affected, 0) AS affected,
5 ROUND(100.0 * COALESCE(e.affected, 0) / NULLIF(a.installs, 0), 2) AS pct_affected
6FROM active a LEFT JOIN erroring e USING (version) ORDER BY a.version DESC;
Execution context: your metrics database (the self-hosted endpoint from building a minimal error collection endpoint stores exactly these aggregates). The percentage of installs affected is robust against error loops on a few machines. Hosted services compute “crash-free sessions/users” per release the same way.
4. Define alerts with minimum volumes
1// scripts/check-release-health.mjs — run every 30 minutes for 72 h after each release
2const [latest, previous] = await versionRates();
3if (latest.installs < 200) process.exit(0); // too early to judge
4const ratio = latest.pct / Math.max(previous.pct, 0.1);
5const newIssues = await newFingerprints(latest.version, { minPctInstalls: 1 });
6if (ratio >= 3 || latest.pct >= 3 || newIssues.length) {
7 await notify(`Release ${latest.version}: ${latest.pct}% installs erroring (prev ${previous.pct}%), new issues: ${newIssues.map((i) => i.message).join("; ")}`);
8}
Execution context: a scheduled job (CI cron or your server). Minimum volumes avoid alerting on the first ten installs, where one user’s bad day looks like a 10% rate. Thresholds should reflect your baseline — start loose, tighten as you learn. Route alerts to wherever the team actually looks (chat, pager, email). Limit the active window to the first few days after a release, when spikes matter most.
5. Use staged rollouts to limit the blast radius
The Chrome Web Store lets publishers with enough users release to a percentage of users and increase it later; the API supports a deployPercentage. Release to 5–10%, wait for the health check to pass over a day, then go to 100%. A bad release then affects a fraction of users while you fix it. See staged rollouts in the Chrome Web Store.
6. Write the runbook before you need it
1runbook-bad-release.md
21. Confirm: open the top new issue; reproduce on the released build if possible.
32. Halt: Chrome — set rollout percentage to current (no increase) or unpublish the draft; Edge/AMO — cancel pending submission if still in review.
43. Mitigate: flip the remote feature flag for the affected feature (no code release needed).
54. Fix: hotfix branch from the release tag → version +0.0.1 → full CI → staged release.
65. Communicate: store listing "known issue" note; reply template for support.
76. Review: add a test for the bug; adjust alert thresholds if the alert was late.
Execution context: the team’s docs. There is no rollback in extension stores — users who updated keep the new version until a newer one ships — so the fastest mitigations are halting the rollout and disabling the faulty feature with a remote flag that changes configuration, not code. See remote feature flags without remote code.
7. Watch non-error signals too
Some bad releases do not throw: a feature silently does nothing. Add a few success counters for core actions (highlights created per active install) to the same dashboard and alert on sharp drops, again comparing versions.
Common mistakes
- Alerting on raw counts. They rise with adoption.
- No denominator. Rates need active installs per version.
- Alerts on tiny samples. Ten installs produce noise.
- No staged rollout. A bad release reaches everyone.
- No runbook. The first hour is spent deciding what to do.
Cross-browser variation
- Chrome / Edge: Chrome Web Store supports percentage rollouts for larger extensions; Edge Add-ons offers fewer rollout controls.
- Firefox: AMO publishes to all users once approved; use feature flags to limit exposure.
- Safari: App Store phased release for macOS apps can spread an update over seven days and be paused.
Verification
- Simulate a release in staging with injected errors and confirm the alert fires once volume passes the minimum.
- Confirm the alert names the version, rates and top new issue.
- Rehearse the runbook: halt a staged rollout and flip a feature flag.
- Confirm the job stops alerting after the post-release window.
FAQ
How quickly do users get updates?
Chrome checks for updates every few hours; most active users update within a day or two.
Can I force users back to the previous version?
No. Publish a new version with the old code (and a higher version number).
What baseline error rate is normal?
It depends on the extension; measure your own over several releases and alert on relative change.
Should the health check run for every release?
Yes, including small ones. Many bad releases are “trivial” changes — a dependency bump, a manifest tweak — that nobody expected to matter.
Related
- Grouping and deduplicating extension errors — clean issues to alert on.
- Staged rollouts in the Chrome Web Store — limiting exposure.
- Building a minimal error collection endpoint — the data source.
- Error monitoring and crash reporting — the parent topic.