Feature flags vs A/B tests: related, not the same
The confusion is understandable, because from ten feet away the two look identical: an if statement decides which of two code paths a user gets. Flag vendors sell experimentation add-ons, experiment platforms ship flag features, and teams end up asking whether they need one tool or two.
The clean way to separate them is by the question each answers. A feature flag answers "who should get this code right now?" An A/B test answers "what does this change cause?" One is about control, the other about knowledge. Everything else — tooling, statistics, cleanup habits — follows from that split.
What feature flags are actually for
- Decoupling deploy from release. Merge code to main behind a flag, deploy it dark, turn it on when ready. This is the core value and it has nothing to do with measurement.
- Progressive rollouts. Enable for 1%, then 10%, then everyone, watching error rates as you go. The goal is safety, not learning — you already decided to ship.
- Kill switches. When the new payment integration misbehaves at 2am, flipping a flag beats rolling back a deploy.
- Entitlements. Plan-based access, beta cohorts, internal-only features. These are permanent flags that behave like configuration.
Notice what is missing: a control group and a metric. A 10% rollout tells you the feature does not crash. It does not tell you whether the feature improved activation, because the 10% were not compared against anyone under randomized conditions with statistics deciding when the evidence is sufficient.
What an A/B test adds on top
An experiment takes that same switch and wraps discipline around it. Assignment must be random and sticky, so the groups differ only by the change. Exposure must be logged, so you know exactly who was in the test. A primary metric and guardrails must be declared up front. And an analysis method decides when the difference is evidence rather than noise.
This is why "we can A/B test, our flag tool does percentage rollouts" is only half true. Percentage rollout gives you the traffic split. It does not give you exposure tracking joined to your funnel, a pre-registered decision rule, or valid statistics — and teams that eyeball a dashboard during a rollout are running an uncontrolled experiment with extra steps.
Where the overlap gets real: graduating a flag into a test
The happy path between the two looks like this: a feature ships dark behind a flag, rolls out to a small slice for stability, and then — before the team declares victory — the flag graduates into an experiment. Fifty-fifty randomized assignment, a primary metric, guardrails, and a run until the statistics call it.
- Ship dark. Merge behind the flag, verify in production with internal users.
- Roll out for safety. A small percentage, watching errors and latency. This step is about not breaking things.
- Randomize for learning. Convert to a 50/50 experiment with a declared decision metric — we argue for exactly one decision metric per test.
- Decide and delete. Ship the winner to 100%, then remove the flag and the dead branch. A flag that lingers after its decision is tech debt with a nice UI.
Tooling implications: one tool or two?
| Need | Flag tool is enough | You need experimentation |
|---|---|---|
| Deploy dark, kill fast | Yes | — |
| Gradual rollout on error metrics | Yes | — |
| Plan-based entitlements | Yes | — |
| Know if a change moved conversion | No | Randomization + stats |
| Compare two designs of a flow | No | Randomization + stats |
| Decide when evidence is sufficient | No | Sequential or fixed-horizon analysis |
Small teams often do fine with a flag tool plus honest self-restraint about what rollouts can and cannot tell them. Once you genuinely need causal answers, the choice is between an experimentation platform layered over your flags or a tool that owns the loop end to end — we compare the categories in feature flag tools vs experiment platforms.
There is also a lifecycle difference worth planning for. Entitlement flags live forever by design. Experiment switches must die: the moment a test concludes, the losing branch is dead code. This is a place where automation honestly helps — Trevo treats the experiment switch as temporary by construction, writing the variant as a pull request when the test starts and opening a cleanup PR that removes the losing path when it concludes, so decided experiments do not accumulate as flag debt. The pattern matters more than the tool; we wrote up why in cleanup PRs and experiment debt.
Frequently asked questions
Can I use feature flags for A/B testing?
A flag can supply the traffic split, but a valid A/B test also needs sticky randomized assignment, exposure logging joined to your analytics, a pre-declared primary metric, and a statistical method for deciding. If your flag tool provides those, you are experimenting; if you are eyeballing dashboards during a rollout, you are guessing with percentages.
What is the difference between a rollout and an A/B test?
A rollout gradually increases the share of users on new code to catch breakage safely — the decision to ship is already made. An A/B test holds a randomized control group and compares metrics to make the decision. Rollouts protect stability; tests produce causal evidence. Mature teams do both, in that order.
Should feature flags be removed after an experiment?
Yes. Once an experiment concludes, the losing code path is dead weight: it confuses readers, complicates testing, and invites accidental re-enabling. Ship the winner, delete the loser and the switch. Long-lived flags are legitimate only for entitlements, ops kill switches, and configuration — not for decided experiments.
Do feature flag tools include statistics?
Some ship experimentation add-ons with stats engines; others provide only percentage rollouts and targeting. Check whether assignment is sticky and logged, whether results join to your product metrics rather than just flag evaluations, and what statistical method backs the significance claims — and whether any of it costs extra.