
Marketers: Stop Waiting Months. Size Test Groups With MDE and Power

A test group, also called an experimental group, treatment group, or variant group, is the set of units randomly assigned to receive a new experience while a control group keeps the original version. Comparing the two lets you measure whether the change actually caused a difference in behavior, not just correlated with one.
TL;DR:
- Random assignment using a stable, repeatable method like hashing on a user ID is essential to ensure reliable comparison and prevent contamination.
- Small effects require exponentially larger sample sizes, so detecting minor improvements on low-traffic pages may take months or be impractical.
- Guard against selection bias, contamination, and measurement errors by verifying tracking and bucketing logic before running the test.
- Combining qualitative focus groups with quantitative test groups helps generate hypotheses and measure actual behavioral impact accurately.
- Using lightweight, no-code tools can facilitate quick setup and maintain experiment validity without extensive engineering resources.
Table of Contents
- Definition, Synonyms, and Where You'll See the Term
- How Test Groups Are Created and Assigned
- Sizing a Test Group: MDE, Power, and Sample Size Basics
- Common Pitfalls and Validity Threats
- When You Need a Test Group Versus a Focus Group
- Practical Checklist: Setting Up a Test-Group Experiment
- Why Lightweight Tooling Matters for Test-Group Accuracy
- Get a Test Group Running Without the Engineering Backlog
- Sources
- FAQ
Definition, Synonyms, and Where You'll See the Term
A test group is the group exposed to whatever you're testing: a new headline, a redesigned checkout flow, a different price, or a new drug. The Stanford Graduate School of Business explainer defines it as the set of units that receives the treatment, with the term used interchangeably alongside experimental group, treatment group, and variant group depending on the field and the source.
You'll run into the phrase in a few recurring contexts:
- A/B tests: the test group sees variant B while the control sees variant A.
- Product testing: a batch of users trials a new feature before wider release.
- Software rollouts: a feature flag exposes a subset of accounts to new functionality.
The term gets confused with a couple of neighbors. A focus group is a small discussion panel used to gather opinions, not a randomized sample measuring behavior. A control group is the test group's counterpart, the baseline that did not receive the change. Test group and treatment group are functionally the same thing, just different names preferred by different disciplines.
How Test Groups Are Created and Assigned
Random assignment is what separates a real experiment from a comparison you can't trust. When units land in the test group by preference, geography, or time of day instead of chance, any difference you measure could reflect who ended up there rather than what you changed. The Stanford explainer is direct about this: randomization is what makes the treatment and control groups comparable enough to support a causal claim, and comparisons should run concurrently rather than against a historical baseline that carries its own seasonality and mix effects.
Setting this up in practice means working through a few decisions in order.
- Pick the experimentation unit: user, session, or device, depending on where the change lives and how people interact with it.
- Assign that unit to a group using a random, repeatable method such as hashing on a user ID.
- Persist the assignment so the same unit sees the same variant every time it returns.
- Verify the assignment logic before any real data collection starts.
Persistent assignment matters because a user who bounces between the test and control experience contaminates the comparison. Hashing a stable identifier, rather than assigning randomly on every page load, keeps exposure consistent.
Pro Tip: Run a quick check on your bucketing logic with a handful of test accounts before launch, confirming each one lands in the same group across sessions and devices.
Sizing a Test Group: MDE, Power, and Sample Size Basics
Before you launch anything, you need to decide how big the test group has to be, and that depends on three linked choices. The minimum detectable effect (MDE) is the smallest change you actually care about catching. Alpha is your tolerance for a false positive, the risk of declaring a win that isn't real. Power is your chance of detecting a true effect when one exists. Fix any two of these and the sample size is determined.

Detecting smaller effects gets expensive fast. Holding alpha and power constant, Stanford's experiment design lecture notes show that catching an effect half as large typically requires about four times the sample size, because sample size scales inverse-quadratically with the MDE. Halving your MDE again means another fourfold jump, which is why chasing tiny lifts on low-traffic pages can stall a test for months.
A practical sizing routine looks like this:
- Choose a single overall evaluation criterion (OEC) that reflects the outcome you actually care about.
- Set alpha and power targets; Stanford's guidance uses 5% alpha and 80% power as common planning examples, not universal rules.
- Estimate your current baseline conversion rate from existing data.
- Run a power calculation, using a step-by-step sample size walkthrough or a similar calculator, to get the sample size the test group needs.
Common Pitfalls and Validity Threats
Most broken experiments trace back to a handful of repeat offenders. Selection bias tops the list: any non-random path into the test group, even something as subtle as excluding users who churned early, can make the two groups incomparable before the test even starts.
- Contamination: a user who sees both variants, often through account switching or shared devices, muddies the comparison.
- Assignment drift: bucketing logic that changes mid-test after a code deploy silently reshuffles who's in which group.
- Instrumentation errors: a tracking pixel that fires inconsistently between variants will bias the metric before you ever look at significance.
- Peeking and multiple comparisons: checking results daily and stopping the moment you see significance inflates your false-positive rate far beyond the alpha you set.
Guard against the last one with a pre-registered analysis plan: decide your OEC, alpha, and sample size before launch, then watch guardrail metrics (page load time, error rates, unsubscribe rates) alongside the primary metric so a short-term win doesn't mask a longer-term cost.
Pro Tip: Run an A/A test occasionally, comparing identical experiences against each other, to confirm your platform's false-positive rate matches the alpha you're setting.
When You Need a Test Group Versus a Focus Group
Focus groups and test groups answer different questions. A focus group surfaces attitudes, language, and early hypotheses, but group dynamics can push participants toward conformity, so treat the output as exploratory rather than conclusive. A test group measures actual behavior against an OEC and, with proper randomization, can support a causal claim.
The two work best in sequence:
- Start with qualitative discovery to surface a hypothesis worth testing.
- Turn that hypothesis into a controlled test-group experiment to measure real behavioral impact.
- Follow up with more qualitative work to understand why the winning variant worked.
This loop, explored further in our comparison of qualitative and quantitative methods, keeps you from mistaking strong opinions for proven impact.
Practical Checklist: Setting Up a Test-Group Experiment
Running a valid test comes down to sequencing these steps correctly and not skipping the boring parts.
- Write a clear hypothesis and choose one OEC to judge success.
- Pick the assignment unit and confirm persistence across sessions.
- Set alpha, target power, and your MDE before collecting any data.
- Calculate the required sample size and randomize eligible units.
- Verify tracking is firing correctly, then run the test to its pre-planned end date.
- Analyze results, checking guardrail metrics and key segments before declaring a winner.
Our best practices roundup and this CRO tactics guide both dig deeper into the instrumentation checks worth running before step 4.
Why Lightweight Tooling Matters for Test-Group Accuracy
A slow script is its own validity threat. If the code assigning visitors to a test group adds noticeable lag, you're measuring the lag alongside whatever you intended to test, and the two effects blur together. That's the case for keeping experimentation scripts as light as possible and for verifying that assignment logic runs before the page renders, not after.

A no-code visual editor removes another source of error: manually wiring up variants in code invites the kind of inconsistent bucketing that contaminates a test group before day one. Persistent assignment handled automatically, rather than patched together per project, is what keeps a test group clean across sessions and devices.
For readers building out their own sample size and power calculations, our guide to A/B test significance walks through the same MDE and alpha tradeoffs covered above in more detail.
— Juan
Get a Test Group Running Without the Engineering Backlog
Setting up a properly randomized, persistent test group usually means engineering time you don't have on a small team. Gostellar's no-code visual editor lets you build and launch a variant without writing bucketing logic yourself, while a lightweight script keeps page performance out of the equation entirely. Plans run from a free Sandbox tier for lower-traffic sites up through Plus at $29 per month, Pro at $99 per month, and Scale at $399 per month, with an Enterprise tier available on request, all detailed on the Gostellar pricing page.
If you're testing a social proof message as part of your next experiment, the free FOMO popup generator is a quick way to create a variant worth putting in front of a test group. Check plan details and get started at Gostellar.
Sources
- Explainer: What is A/B testing? | Stanford Graduate School of Business
- Estimating power and sample size (Stanford guidance)
- Fundamentals of experiment design (Stanford lecture notes)
- Focus groups as a qualitative research method | Nielsen Norman Group
FAQ
What is another word for test group?
A test group is also called an experimental group, treatment group, or variant group, and the terms are used interchangeably depending on the field. In an A/B test, the test group is often referred to specifically as the group seeing variant B.
What does test group mean?
A test group is the set of people, sessions, or products randomly assigned to receive a new treatment or experience so its effect can be measured against a control. Dictionary definitions describe it as the group exposed to a treatment and compared with a control group that did not receive it.
What's the difference between a test group and a control group?
The test group receives the new treatment or variant, while the control group experiences the original, unchanged version. Comparing outcomes between the two, with both running at the same time, is what lets you attribute any difference to the change itself rather than to outside factors.
How do I decide how big my test group needs to be?
Sample size depends on your minimum detectable effect, your false-positive tolerance (alpha), and your target statistical power, with any two of the three determining the required size. Stanford's power and sample size guidance recommends running the calculation before launch rather than checking results as they come in.
Can I use a focus group instead of a test group?
A focus group can surface early opinions and language, but it can't measure actual behavioral impact or establish causality the way a randomized test group can. Research on focus group dynamics recommends treating their findings as hypotheses to validate with a behavioral experiment, not as a substitute for one.
Recommended
Published: 9/30/2026