Privacy-Preserving Usage Metrics
Collect extension usage metrics users and reviewers can trust: an explicit event schema, on-device aggregation, coarse timestamps, no URLs, consent gating, incognito exclusion and minimum-count thresholds.
Table of Contents
The product team wants to know which features get used. The privacy policy promises the extension does not track browsing. The Chrome Web Store’s data disclosure form asks exactly what is collected. Those three constraints are compatible, but only if analytics are designed from the privacy side first: decide what you will never collect, then find the smallest set of signals that answers the product questions within that boundary. Extensions have a uniquely privileged view of users’ browsing, so the bar is higher than for a website. This guide sets out a pattern that clears it. It belongs to usage analytics and feature flags.
Why extension telemetry needs a stricter design
A website’s analytics see one site. An extension’s analytics, if careless, can see every site: the content script runs on page after page, the service worker observes tab URLs, and any event that carries context — “highlighted text on example.com”, “saved article titled …” — turns usage analytics into a browsing log. Combine that with a persistent install id and precise timestamps, and the data set can re-identify individuals even with names removed. The defences are structural rather than promissory: an allow-listed schema that has no field for page data, aggregation on the device so individual actions are never sent, coarse time, thresholds that suppress rare values, and a hard rule that private windows produce nothing.
Step-by-step: metrics that cannot leak browsing
1. Write down what you will never collect
1analytics/NEVER.md
2- URLs, hostnames or titles of pages the user visits
3- Any text from pages, selections, forms or the omnibox
4- Search terms, file names, bookmark titles
5- Anything from incognito / private windows
6- Account identifiers, emails, IP-derived location beyond country
7- Precise timestamps (finer than one day for counts, one hour for errors)
Execution context: a document in the repository, referenced by code review and by the privacy policy. Writing the exclusions first gives reviewers — internal and store — something concrete to check against, and it settles arguments about “just one more field” before they start.
2. Encode the allowed events as a closed schema
1// analytics/schema.js
2export const SCHEMA = Object.freeze({
3 feature_used: { feature: ["save", "summarise", "highlight", "read_aloud"] },
4 onboarding_step: { step: [1, 2, 3, 4] },
5 setting_toggled: { setting: ["autoHighlight", "darkMode", "syncEnabled"], on: [true, false] },
6});
7
8export function validate(name, props) {
9 const spec = SCHEMA[name];
10 if (!spec) return null;
11 const out = {};
12 for (const [k, allowed] of Object.entries(spec)) {
13 if (!allowed.includes(props[k])) return null; // unknown value → drop the event
14 out[k] = props[k];
15 }
16 return out;
17}
Execution context: a shared module used by the service worker. Every property is an enumeration, not free text, so there is no way to smuggle a URL or a title through a field. An event with an unexpected value is dropped rather than truncated. New events or values require a schema change, which is a reviewable diff.
3. Aggregate on the device
Instead of sending each action, keep daily counters and send totals.
1// sw.js
2export async function count(name, props) {
3 const clean = validate(name, props);
4 if (!clean) return;
5 const day = new Date().toISOString().slice(0, 10); // YYYY-MM-DD
6 const key = `${name}|${Object.values(clean).join("|")}`;
7 const { counters = {} } = await chrome.storage.local.get("counters");
8 counters[day] ??= {};
9 counters[day][key] = (counters[day][key] ?? 0) + 1;
10 await chrome.storage.local.set({ counters });
11}
Execution context: the service worker. The server receives “feature_used|save: 7 on 2026-10-01” rather than seven timestamped events, which removes the sequence of actions — a powerful re-identification signal — entirely. Product questions about usage frequency and adoption are still answerable. Prune counters older than the last successful send.
4. Gate everything on consent and context
1chrome.runtime.onMessage.addListener((msg, sender) => {
2 if (msg?.type !== "metrics:count") return;
3 if (sender.tab?.incognito || chrome.extension.inIncognitoContext) return;
4 chrome.storage.local.get("consent").then(({ consent }) => {
5 if (consent === "granted") count(msg.name, msg.props);
6 });
7});
Execution context: the service worker. Checking sender.tab.incognito covers content scripts and pages in private windows when the extension runs in spanning mode; inIncognitoContext covers split-mode instances. Consent defaults to absent, which means nothing is counted. Do the check at collection time, not only at send time, so that withdrawing consent stops accumulation immediately.
5. Send without a persistent identifier where possible
1chrome.alarms.onAlarm.addListener(async ({ name }) => {
2 if (name !== "metrics-send") return;
3 const { counters = {}, consent } = await chrome.storage.local.get(["counters", "consent"]);
4 if (consent !== "granted") return;
5 const today = new Date().toISOString().slice(0, 10);
6 const days = Object.keys(counters).filter((d) => d < today); // complete days only
7 if (!days.length) return;
8 const payload = { v: chrome.runtime.getManifest().version, days: days.map((d) => ({ d, c: counters[d] })) };
9 const res = await fetch("https://t.readable.example/v1/counts", { method: "POST", body: JSON.stringify(payload) });
10 if (res.ok) { for (const d of days) delete counters[d]; await chrome.storage.local.set({ counters }); }
11});
Execution context: the service worker. This payload has no install id at all: the server sums counts across all reports. You lose per-install retention analysis but gain a data set that cannot be linked to any individual — and active-user counts can be obtained separately with the technique in counting active users without tracking. Only complete days are sent, so partial-day counts never reveal “used it at 9am”.
6. Apply thresholds on the server
1-- Report only values seen from enough distinct reports
2SELECT event_key, SUM(count) AS total
3FROM daily_counts
4WHERE day = :day
5GROUP BY event_key
6HAVING COUNT(*) >= 20; -- suppress rare combinations
Execution context: your analytics backend. Rare combinations — a feature used by two people in one locale on one version — can identify those people when crossed with other knowledge. Suppressing groups below a threshold, and dropping IP addresses at the edge before storage, keeps reports useful for decisions while removing the long tail where re-identification lives.
Common mistakes
- Free-text properties.
{ feature: name }with an unchecked string becomes a URL or a title the first time a developer passes the wrong variable. - Collecting “just the hostname”. A hostname per event is a browsing history at domain resolution.
- Precise timestamps. Sequences of millisecond timestamps fingerprint individual behaviour.
- Counting in private windows. Users opened incognito precisely to avoid this.
- A privacy policy that drifts from the code. Generate the policy’s data list from the schema so they cannot disagree.
Cross-browser variation
- Chrome / Edge: the Chrome Web Store’s privacy practices form must list the categories collected; aggregated, non-identifying usage data is still “collected” for the form’s purposes.
- Firefox: AMO policy requires opt-in for this kind of collection, and recent Firefox versions surface the manifest’s data-collection declaration at install.
- Safari: the containing app’s privacy label must declare analytics, even aggregated.
Verification
- Search the analytics code for
location,url,titleandtextContent: none should reachcount(). - Use the extension in an incognito window and confirm counters do not change.
- Withdraw consent and confirm counters stop incrementing and the next send does not happen.
- Inspect a captured payload: no identifier, only enumerated keys, day-resolution dates.
FAQ
Can I still measure retention without an install id?
Not per install. Aggregate retention can be estimated from active-user counts by install cohort if the client reports its install week as a coarse value. Decide whether that question justifies the extra field.
Is a random install id personal data?
In many jurisdictions a persistent identifier tied to a device can be personal data even if random. Avoid it unless the analysis truly needs it, and disclose it if you use it.
Do error reports follow the same rules?
Similar principles, different needs: error reports may need a stack trace and a version, but never page content. See reporting errors without breaking your privacy policy.
Related
- Counting active users without tracking — the active-user number without ids.
- Sending analytics from a service worker with GA4 — if you need a hosted analytics tool.
- Writing a privacy policy for an extension — describing this design to users.
- Usage analytics and feature flags — the parent topic.