The novelty effect: why your variant looks great for a week
A team redesigns a navigation menu and launches the test. By day five the variant is up double digits and the dashboard is glowing. By day twenty the lift has drained to nothing, and by day thirty the variant is flat — or slightly behind. Nothing broke. Nobody made a statistics error. The team simply measured the one effect every change carries for free: being new.
The novelty effect is the temporary behavior change users exhibit toward anything different, independent of whether it is better. It systematically flatters early results, it wears off on its own schedule, and it is the reason "the graph looked amazing all week" and "the graph was telling the truth" are different claims. Every experiment on a surface with returning users has to reckon with it.
How novelty works — and its evil twin, primacy
Returning users have your current interface cached in muscle memory. Change something and the change itself grabs attention: the new banner gets clicked because it was not there yesterday, the moved button gets explored because it moved. That attention converts into whatever your metric counts, for a while. The mechanism is curiosity about difference, so it cannot last — once the new thing is the normal thing, the extra engagement evaporates, and what remains is the variant's true effect, which may be positive, zero, or negative.
Primacy is the mirror image. Habits are efficient: users who could navigate your checkout half-asleep suddenly have to look for things, and their performance temporarily drops. A genuinely better redesign can lose its first two weeks because it broke everyone's autopilot. Novelty makes bad changes look good; primacy makes good changes look bad. Both are lies with an expiry date, and both push hardest exactly when the data is thinnest — early.
The diagnostic: split new users from returning users
The cleanest test for novelty is built into your traffic. New users have never seen your old interface, so nothing is novel to them — the variant is simply the product. Returning users carry the memory that makes change feel like change. Compare the treatment effect in the two groups and the story tells itself.
| Pattern | Reading |
|---|---|
| Variant wins with new users and returning users | Likely a real improvement — ship with confidence |
| Variant wins only with returning users | Novelty suspect: the lift lives in the users for whom it is a change |
| Variant wins with new users, loses with returning users | Likely a real improvement masked by primacy — expect the returning-user gap to close |
| Variant flat with new users, fading spike with returning | The textbook novelty signature — the "win" is surprise wearing off |
How long do novelty effects take to settle?
There is no universal constant — the honest answer is "watch it settle" rather than "wait N days". But the shape is predictable: the effect decays with exposures, not calendar time. What that means practically: surfaces visited daily see novelty burn off in days; surfaces visited monthly can stay "new" to their audience for months; and the more dramatic the visual change, the taller the spike and the longer the tail. Watch the variant's effect by user cohort — users first exposed in week one versus week three — and you can see the decay directly: if week-three joiners show a smaller effect than week-one joiners did in their first week, the effect is fading, not durable.
- High-frequency surfaces (nav, dashboards, feeds): habits are strong, so primacy hits hard and novelty burns fast. Expect the noisiest early data on exactly these tests.
- Checkout and conversion flows: many users hit them rarely, so effects settle slowly — a stable first fortnight means less than it seems.
- Mostly-new-user surfaces (landing pages, signup): the one happy case. If your test audience is dominated by first-time visitors, novelty barely applies — there is no "old version" in anyone's head.
Guarding decisions against novelty
- Run past the spike. Full weeks, and for habit-heavy surfaces, enough weeks that returning users have had repeated exposures. The general calendar math is in how long to run an A/B test.
- Pre-register the new/returning split. It is the difference between a diagnostic and an excuse.
- Watch the trend of the effect, not just its size. A stable lift and a decaying lift can have identical week-two averages. The slope is the tell.
- Weight new-user results for the long term. Every future user is a new user. If the variant is flat for people without the old habit, its steady-state value is probably flat.
- Distrust dramatic early wins on redesigns proportionally to their drama. The changes that trigger the most curiosity are the big visual ones — the same ones teams most want to call early.
A note on statistics: sequential methods let you legitimately stop a test the moment evidence crosses a boundary — but the guarantee is about noise, not about time-varying effects. A novelty spike is real while it lasts, so sequential testing will honestly report that something happened; it cannot know the something is temporary. Statistical rigor and novelty-awareness are separate defenses, and you need both. This is also how Trevo treats it: experiments run on always-valid mSPRT statistics, and habit-heavy surfaces get runtimes that let effects settle before conclusions harden.
The novelty effect is not an argument against testing — it is an argument against impatience. The variant that survives week four earned its lift twice: once against the control, and once against its own opening-week theater.
Frequently asked questions
What is the novelty effect in A/B testing?
It is the temporary lift a variant receives simply because it is different: returning users notice the change and engage out of curiosity rather than preference. The extra engagement inflates early results and fades as the new design becomes familiar, leaving the variant's true effect — which may be zero — once the surprise wears off.
How do I know if my A/B test result is just novelty?
Split the analysis by new versus returning users, ideally pre-registered before launch. New users have never seen the old version, so nothing is novel to them. If the lift appears only among returning users, or if the effect visibly decays week over week among early cohorts, you are likely measuring surprise rather than improvement.
How long does the novelty effect last?
It decays with exposures rather than calendar time, so there is no universal duration. Surfaces users visit daily burn through novelty in days; surfaces visited monthly can stay novel to their audience for months. Compare the effect for users first exposed in different weeks — when later cohorts show weaker effects, the lift is fading.
What is the primacy effect in A/B testing?
The mirror image of novelty: a change disrupts habits, so returning users temporarily perform worse — they must search for controls they used to hit automatically. Primacy makes genuinely better redesigns look bad in their first weeks. A variant that wins with new users while losing with returning users is often a primacy case that will recover.