CUPED: How to Make Your A/B Test Smarter Without Finding More Users
An intuitive, no-heavy-math guide to one of the most useful ideas in product experimentation to not get DUPED.
An intuitive, no-heavy-math guide to one of the most useful ideas in product experimentation to not get DUPED.
Imagine you're a Product Analyst/Data Scientist at any experimentation-led company.
The Product team comes to you excitedly:
"We've redesigned checkout. Let's A/B test it."
You do the usual work.
Control: 20.0% conversion
Treatment: 21.0% conversion
+1 percentage point!
But then you calculate the confidence interval. And suddenly…
"Hmm. We can't confidently say the improvement is real."
The Product Manager asks the question every experimentation analyst eventually hears:
"Can we do something to make the experiment more sensitive?"
You have enough traffic. You have a properly randomized experiment. Your MDE, alpha and power were all planned correctly.
So what now?
First, What Problem Is CUPED Actually Solving?
Let's start with something more intuitive than statistics.
Imagine you're comparing two students after a new teaching program.
Student A scores 95. Student B scores 60. Clearly, A performed better.
But wait. What were their scores before the program?
Student A: 94. Student B: 59. Both improved by exactly 1 point. Suddenly the story is very different.
The problem wasn't the students. The problem was that we were looking at their final scores without considering where they started.
Product experiments have exactly the same problem.
Some users are naturally:
- Heavy buyers
- Frequent visitors
- Highly engaged
- High spenders
Others are naturally:
- Infrequent users
- Low spenders
- Occasional buyers
That natural difference creates noise. And noise makes it harder to see the effect of your experiment.
The A/B Test Problem
Suppose you're testing a new recommendation algorithm. Your KPI is:
GMV per user (Gross Merchandising Value per user)
Consider two users.

If you only look at experiment GMV, Key looks dramatically more valuable than Peele. But that's not necessarily because of your new recommendation algorithm.
Key was already a high-value user. Peele was already a low-value user. Their historical behavior explains part of what you're seeing today.
And here's the important realization:
We already have that historical information.
Why not use it?
This Is Where CUPED Comes In
CUPED stands for:
Controlled-experiment Using Pre-Experiment Data
The name sounds intimidating. The idea isn't. Here's the simplest definition:
CUPED uses what we already know about users before an experiment to remove predictable noise from what we observe during the experiment.
Or, even simpler:
Don't make the experiment rediscover what you already knew about your users.
But How Do We Know What Historical Data to Use?
This is where CUPED becomes interesting. Suppose our experiment KPI is GMV per user. What information from before the experiment might tell us something about future GMV?
Potential candidates:
- Previous GMV
- Previous orders
- Previous sessions
- Previous product views
- Previous active days
- Previous engagement
We don't blindly choose one. We ask:
"Does this historical behavior help explain variation in my experiment KPI?"

There's a clear relationship. Users who historically spend more tend to spend more in the future. That's useful information.
How Do We Check That?
You can start very simply. Take X = historical behavior and Y = experiment KPI, then examine their relationship — for example, historical GMV vs experiment-period GMV.
You can look at:
- Correlation
- Regression
- R²
- Variance reduction
If historical GMV has a strong relationship with future GMV, it's a good CUPED candidate. If the relationship is essentially random, CUPED won't help much.
A Very Important Rule
Your CUPED variable must be pre-treatment. That means:
- Good — GMV during the 30 days before the experiment.
- Good — orders during the previous month.
- Good — historical sessions.
- Bad — orders during the experiment.
- Very bad — clicks generated by the treatment.
Why? Because treatment may have influenced those variables. And once treatment influences your covariate, you've contaminated your adjustment.
Okay, So What Actually Happens to the Data?
Imagine your experiment produces this:

Normally, you'd simply compare average experiment GMV — treatment vs control. CUPED adds one more step. It asks:
"How much of each user's experiment-period outcome could we have predicted from their pre-experiment behavior?"
We estimate that relationship and statistically adjust the outcome. Conceptually:
Adjusted outcome = Observed outcome − predictable component
The exact implementation involves estimating a coefficient for the relationship between the pre-period covariate and the outcome. But you don't need the formula to understand the idea. You're essentially removing:
"This user behaved this way because that's how this user normally behaves."
What's left contains less predictable noise.
And Here's the Magic
Suppose your normal A/B test gives a treatment effect of +1.0%, with a 95% confidence interval of [-0.2%, +2.2%].
The point estimate is positive. But the uncertainty is large — look at that range.
Now apply CUPED. You might still get a treatment effect of +1.0%. But now the 95% confidence interval is [+0.4%, +1.6%].
Wait. What just happened? Did CUPED make the product better? No. Did the treatment effect magically increase? No.
What changed? The noise got smaller.
Why Can Statistical Significance Change?
This is probably the most important statistical intuition behind CUPED.
Suppose your observed treatment effect is +1%, but your standard error is 0.5%. The effect is only about two standard errors away from zero.
Now suppose CUPED reduces the standard error to 0.25%. The treatment effect is now about four standard errors away from zero.
The effect didn't change. The signal-to-noise ratio improved. And because the estimate is now more precise:
- The confidence interval gets narrower.
- The p-value can get smaller.
- A previously inconclusive result can become statistically significant.
That's CUPED.
So Is CUPED Only for Small Samples?
No. Suppose you've calculated: α = 5%, power = 80%, MDE = 1%, required sample = 100,000 — and you have 120,000 users. You already have enough sample. You don't need CUPED.
But if you have strong historical information that predicts your KPI, you may still use it. Why? Because:
Sample size tells you whether you have enough users. CUPED tells you how efficiently you can use those users.
Think: more users = more information. Less noise = more information per user.
But Should We Always Use CUPED?
Not necessarily. Imagine three experiments.
Experiment A
Required sample: 100K. Available: 120K. Historical KPI correlation: very weak → CUPED probably isn't worth the complexity.
Experiment B
Required sample: 100K. Available: 120K. Historical KPI correlation: strong → CUPED could materially improve precision.
Experiment C
Required sample: 100K. Available: 500K. Historical correlation: strong. You can use CUPED. But now ask: is the incremental precision worth the additional complexity? Maybe yes. Maybe no.
What If I Don't Use CUPED?
Nothing catastrophic happens. Your A/B test can still be completely valid. The tradeoff is simply more uncertainty. For example:
- Without CUPED — +1.0% lift, CI: [-0.2%, +2.2%]
- With CUPED — +1.0% lift, CI: [+0.4%, +1.6%]
CUPED didn't change what happened. It changed how clearly we can see what happened.
When Is CUPED Particularly Useful?
- High-variance KPI — revenue, GMV, orders/user, session duration.
- Existing users — you have historical behavioral data.
- Strong user-level persistence — users' past behavior is a good indicator of future behavior.
That's the sweet spot.
When Is CUPED Less Useful?
Consider a brand-new user experiment. You're testing onboarding for people who just created an account. What was their behavior before the experiment? Probably nothing.
There's little historical signal to exploit, so CUPED may provide little benefit. Similarly, if your historical variable has almost no relationship with your outcome, there's little variance to remove.
The Full CUPED Workflow
Here's the process I'd actually use in a real-world scenario:

The One Question I Now Ask Before Every Experiment
When designing an experiment, don't just ask:
"Do we have enough users?"
Also ask:
"What do we already know about these users before the experiment that could help us separate signal from noise?"
That's the question that leads you to CUPED.
And the One Sentence to Remember
If you forget everything else:
CUPED doesn't make the treatment effect bigger. It makes the noise smaller, so the treatment effect becomes easier to see.
That's why one of the most powerful tools in experimentation isn't necessarily about finding more users. Sometimes, it's about making the users you already have more informative.
If you're learning experimentation, CUPED is also a beautiful gateway into a broader idea: variance reduction. Once you understand why reducing variance improves statistical power, techniques like stratification, covariate adjustment, regression adjustment and pre-experiment blocking start to feel like variations of the same fundamental idea.

Written by
Faisal Siddique
Embracing the Magic of Analytics & Life
I write about product analytics, experimentation, data, AI, and the ideas that shape how we think and grow. The goal is simple: make sense of complexity, share what I learn, and hopefully leave you with something worth thinking about.