Running A/B Tests in an Extension

Run an A/B test in an MV3 extension: stable on-device assignment, variants shipped in the package, exposure logging through privacy-preserving counts, sample size, guardrail metrics and ending the test cleanly.

Published October 2, 2026 Updated October 2, 2026 7 min read
Table of Contents

You have two designs for the onboarding page, or two default settings for a feature, and opinions in the team are split. An A/B test settles it with data — if it is run correctly. Extensions add constraints web experiments do not have: both variants must be in the reviewed package, assignment has to survive worker restarts and updates, metrics must be collected without tracking browsing, and the audience is small enough that many tests cannot reach a meaningful result at all. This guide covers a design that works within those limits, and how to tell when a test is not worth running. It belongs to usage analytics and feature flags.

How an extension experiment fits together

An experiment is three things: an assignment that puts each installation in a variant and keeps it there, an exposure record noting that the installation actually saw the variant, and an outcome metric compared between variants. On the web, all three usually live in a third-party SDK. In an extension, assignment is a random number stored on install, exposure and outcomes are counts aggregated on the device per variant, and the experiment’s existence and allocation are controlled by the same flags document described in remote feature flags without remote code. Both variants ship in the package; the flags only choose.

The pieces of an extension A/B testA stored random seed assigns each install to a variant; the flags document turns the experiment on and sets allocation; the extension shows the packaged variant, records exposure and outcome counts per variant, and sends aggregates.flags.jsonexperiment on, 50/50Stored seedper experimentVariant A or Bpackaged coderecord what happened, per variantExposure countsaw the variantOutcome countcompleted onboardingDaily aggregatesno identifiers
Assignment is local, variants are packaged, results leave the device only as counts.

Step-by-step: one experiment, end to end

1. State the hypothesis and the metric first

1experiments/onboarding-v2.md
2Hypothesis:   A two-step onboarding (B) completes more often than the four-step one (A).
3Unit:         installation
4Primary:      onboarding_completed / onboarding_exposed, within 1 day of install
5Guardrails:   uninstall within 7 days; error_shown rate
6Allocation:   50/50 of new installs only
7Minimum detectable effect: +5 percentage points
8Stop:         at planned sample size, or guardrail breach

Execution context: a document in the repository. Writing the decision rule before data arrives prevents the most common experimental error — looking at the numbers daily and stopping when they look good. The guardrail metrics catch a variant that “wins” on the primary metric while harming users elsewhere.

2. Check that the test can reach a result

1// Rough sample size per variant for a difference in proportions (two-sided, 95%, 80% power)
2function sampleSizePerArm(p1, p2) {
3  const z = 1.96 + 0.84;
4  const pbar = (p1 + p2) / 2;
5  return Math.ceil((z ** 2 * 2 * pbar * (1 - pbar)) / (p2 - p1) ** 2);
6}
7sampleSizePerArm(0.60, 0.65);   // ≈ 1,470 new installs per variant

Execution context: a Node script or the browser console. If the extension gets 200 new installs a week, a 50/50 test needing 1,470 per arm runs for about fifteen weeks — longer than most product cycles. Knowing that up front lets you choose a bigger effect to look for, a more sensitive metric, or a qualitative method instead of an underpowered test whose result is noise.

Weeks to reach sample size by weekly new installsEstimated duration of a 50/50 test needing about 1,470 installs per variant, for extensions receiving 100, 250, 500, 1,000 and 5,000 new installs per week.100 installs/week29.4 weeks250 installs/week11.8 weeks500 installs/week5.9 weeks1,000 installs/week2.9 weeks5,000 installs/week0.6 weeks
Small extensions should test big changes or not test at all.

3. Assign once, on install, and persist

 1// sw.js
 2chrome.runtime.onInstalled.addListener(async ({ reason }) => {
 3  if (reason !== "install") return;                  // new installs only for this experiment
 4  const { experiments = {} } = await chrome.storage.local.get("experiments");
 5  experiments["onboarding-v2"] ??= { bucket: Math.floor(Math.random() * 100), enrolledAt: Date.now() };
 6  await chrome.storage.local.set({ experiments });
 7});
 8
 9export async function variant(expId) {
10  const { experiments = {}, flagsDoc } = await chrome.storage.local.get(["experiments", "flagsDoc"]);
11  const cfg = flagsDoc?.flags?.experiments?.[expId];
12  const me = experiments[expId];
13  if (!cfg?.on || !me) return "A";                   // control when off or not enrolled
14  return me.bucket < cfg.percentB ? "B" : "A";
15}

Execution context: the service worker for assignment; any context for reading. A per-experiment bucket stored on install keeps the assignment stable across restarts and updates. Restricting enrolment to new installs keeps existing users’ habits from confounding an onboarding test. When the experiment is turned off remotely, everyone falls back to the control, which is the safe default.

4. Log exposure where the variant is actually shown

1// onboarding.js
2const v = await chrome.runtime.sendMessage({ type: "experiment:variant", id: "onboarding-v2" });
3renderOnboarding(v);                                           // both variants are in the package
4chrome.runtime.sendMessage({ type: "metrics:count", name: "exp_exposed", props: { exp: "onboarding-v2", v } });
5
6document.querySelector("#finish").addEventListener("click", () => {
7  chrome.runtime.sendMessage({ type: "metrics:count", name: "exp_outcome", props: { exp: "onboarding-v2", v, outcome: "completed" } });
8});

Execution context: the onboarding page. Counting exposure at render time, not at assignment, means installs that never opened onboarding do not dilute the comparison. The counts flow through the same consent-gated, schema-validated, on-device aggregation described in privacy-preserving usage metrics; add exp, v and outcome to the schema as enumerations.

Assignment, exposure and outcome for one installOn install the worker stores a bucket; when onboarding opens it asks for the variant, renders B, records exposure; when the user finishes it records the outcome; daily aggregates carry both counts per variant.Service workerOnboarding pageCollectoronInstalled: bucket = 37variant('onboarding-v2')"B" (37 < 50)count exp_exposed {v:B}count exp_outcome {v:B, completed}daily totals per variant
The server only ever sees totals per variant.

5. Analyse with the planned rule, then end the test

1SELECT v,
2       SUM(CASE WHEN name = 'exp_outcome' THEN count END) * 1.0 /
3       SUM(CASE WHEN name = 'exp_exposed' THEN count END) AS completion_rate
4FROM daily_counts
5WHERE exp = 'onboarding-v2' AND day BETWEEN :start AND :end
6GROUP BY v;

Execution context: your analytics backend. Compare the rates with the test you planned (a two-proportion z-test for this design) once the sample size is reached, and check the guardrails. Then end the experiment: ship the winner as the only code path in the next release, delete the loser’s code, and remove the experiment from the flags document. Experiments left running indefinitely become permanent complexity.

Common mistakes

  • Fetching variant code remotely. Both variants must be in the package; remote variants are remote code.
  • Re-randomising on each check. Users flip between variants and every metric becomes noise.
  • Counting assignment as exposure. Installs that never saw the variant dilute the difference toward zero.
  • Peeking and stopping early. Daily checks with a “stop when significant” rule produce false winners. Use the planned sample size.
  • Underpowered tests. A test that cannot reach its sample size in reasonable time should not be run as an A/B test.

Cross-browser variation

  • Chrome / Edge: onInstalled with reason: "install" enrols new users; Edge installs come from a separate store and can be analysed as their own segment.
  • Firefox: same enrolment logic. AMO’s opt-in consent requirement means only consenting users contribute counts; check that consent rates are similar between variants.
  • Safari: works the same within the extension; because Safari users often reach the extension through the containing app’s onboarding, consider whether the experiment belongs in the app instead.

Verification

  1. Install fresh several times (new profiles) and confirm roughly half receive each variant, and that a given profile always gets the same one.
  2. Turn the experiment off in the flags document and confirm every install shows the control.
  3. Confirm exposure counts increment only when onboarding renders, and outcome counts only on completion.
  4. Run the analysis query on staging data and confirm both variants appear with plausible rates.

FAQ

Can I test changes that need different permissions?

Not cleanly. Permissions are declared in the manifest for everyone. Test permission-related flows with optional permissions requested in one variant only.

How do I test on existing users?

Enrol them at a defined moment — for example, the first startup after a release — rather than at install, and record that cohort separately.

Is an A/B test a policy concern?

Not if both variants are reviewed code, data collection is disclosed and consented, and variants do not differ in what data they collect.

Other Testing, Debugging & Performance Optimization Resources