The new version of the offer generated higher sales than the old one – but was this due to the offer itself, the season, competitors’ campaigns, or the random distribution of customers? Experimental research is best suited to answering such questions. It is an approach in which the researcher controls a variable and observes its effect on the outcome. The following explains how to design an A/B test so that its result provides a reliable basis for drawing conclusions about causation rather than coincidence.
How does experimental research differ from a standard comparison of results?
Most marketing analyses are observational in nature: sales results are compared before and after a change, customer segments are contrasted, and co-occurrences between phenomena are identified. The problem is that such data always involve multiple factors at once – seasonality, competitor activity, price changes, promotional activity, and even weather. Observation shows that two phenomena occur together, but it usually does not determine which one is the cause.
Experimental research solves this problem by deliberately introducing a difference between groups and randomly assigning units to those groups. If the test group and the control group are formed at random, and test execution does not introduce additional differences between them, then before the experiment begins they differ only by chance – in terms of age, loyalty, basket value, or propensity to buy. The only planned systematic difference between them is the stimulus introduced by the researcher. Therefore, the difference in outcomes can be attributed to that stimulus.
It is worth broadening the perspective beyond digital channels from the outset. An A/B test is the best-known form of experiment, but not the only one. In research practice, the following are tested:
- offer and pricing structure variants (two package versions, two discount thresholds),
- advertising message and sales argument variants,
- changes at the point of sale: displays, POS materials, shelf layout, sales scripts used in stores,
- after-sales service variants and sequences of customer contact,
- packaging and label variants in retail or simulated purchase settings.
When an experiment takes place under natural market conditions – in selected stores, regions, branches, or dedicated customer lists – it is referred to as a field experiment. Its advantage is high external validity: participants behave in a natural purchasing context, often without knowing that they are taking part in a study. Its drawback is less control over confounding factors, which must be offset through a better measurement design. A laboratory experiment (for example, a concept test in a controlled setting) reverses this arrangement: it provides greater control at the expense of the situation’s naturalness.
For managers, one implication is crucial. Correlational analysis answers the question, “what happens together with what?” Experimental research makes it possible to reliably answer the question, “what will happen if we do X?” – provided that randomization and study execution are correct – and to support causal conclusions.
How to design an A/B test step by step?
Well-designed experimental research is decided at the planning stage, not during analysis. Below is the sequence of decisions that must be made before launching an A/B test.
1. A hypothesis and one manipulated variable
The hypothesis must be directional and measurable: “Offer variant B will increase the share of orders containing the extended package compared with variant A.” If the variants differ in price, visual layout, and headline at the same time, the result will not indicate which element worked – it will apply to the entire package of changes. In an A/B test, when the goal is to attribute an effect to a specific element, one element is changed at a time. When several factors need to be examined simultaneously, a factorial design is used (several variants in a crossed arrangement), but this requires interaction analysis to be planned in advance and usually a larger sample, especially to detect those interactions.
2. Primary metric and guardrail metrics
Before the test begins, one primary metric is identified as the basis for the decision – ideally as close as possible to the business objective (purchase, order, signed contract), rather than a proxy for it (click). Guardrail metrics are defined alongside it to capture side effects: basket value, return rate, number of complaints, and cancellations. A variant that increases conversion but reduces margin or increases returns does not necessarily constitute the winning variant.
3. Randomization and the unit of randomization
Randomization is the heart of an experiment. Assignment to groups must be random, rather than based on convenience (“we will show the new offer to regular customers”) or a salesperson’s decision. Defining the unit of randomization is equally important. In a digital channel, this is usually the user, not the session – otherwise, the same person may see both variants and the effect will be diluted. In a field experiment, the unit may be a store, branch, region, or mailing list; in that case, the analysis must account for clustering of observations, which usually substantially increases the required sample size.
4. Sample size determined before launch
Sample size is determined by four parameters: the baseline level of the metric, the minimum effect that is practically meaningful, the assumed significance level, and statistical power. The smaller the effect to be detected, the larger the sample required. This calculation must be completed before the test is launched, not afterward. If the expected number of observations cannot be reached within a reasonable time, it is better to change the design (for example, test a larger difference between variants) than to run an experiment that is unlikely to produce a conclusive result.
5. Duration and behavioral cycles
The test should cover full purchasing cycles: usually full weeks, to balance differences between weekdays and weekends, and in categories with a longer decision-making process, also the time needed to close the sale. Measurement that is too short may capture a novelty effect rather than a lasting change in behavior. It is also worth avoiding periods that are heavily disrupted by holidays, major promotions, or broad-reaching brand campaigns.
As Hume’s Institute experts point out, a common mistake in experimental research is ending a test when the result appears favorable – duration and sample size should be determined before launch, not while watching the chart. Looking at results and stopping a test at a random peak leads to inflated effects; such a “winning” variant may stop working after implementation because it won against noise rather than against the alternative.
6. Implementation control and analysis plan
Before launch, it is worth documenting an analysis plan: which statistical test will be used, which segments will be reported, and how outliers and cases of incomplete exposure to the stimulus will be handled. Implementation should be monitored separately – whether variant B actually reached the test group, whether point-of-sale staff used the assigned script, and whether the system assigned users correctly. An experiment in which the manipulation was not carried out as planned may measure something other than what was intended.
Which errors most often invalidate an experiment’s results?
Even properly planned experimental research can produce an indefensible result if one of the typical flaws occurs during implementation. Below are those that occur most often in research practice.
- Stopping the test too early. Checking significance after every day and ending the test at the first “green” result increases the risk of a false positive conclusion. If in-test monitoring is necessary, methods designed for repeated looks at the data should be used (sequential analysis with adjusted thresholds).
- Contamination of groups. People in the control group come into contact with the stimulus intended for the test group: a customer sees both versions of the offer, a salesperson uses the new script in all conversations, or test and control stores serve the same customers. The effect may be diluted, and the test may show no difference where one actually exists.
- No genuine control group. A “before and after” comparison without a parallel control group does not separate the effect of the change from the effect of time. This is a common substitute for an experiment and a source of incorrect conclusions.
- Pseudorandomization. Assignment based on the first letter of a surname or assigning stores with better performance to the test group may be correlated with characteristics that affect the metric. Even assigning every other customer requires a random starting point and confidence that customer order is not related to the outcome. Randomization must be random and verifiable, and its effectiveness should be checked by comparing groups on baseline variables.
- Post hoc analysis of multiple segments. If the test did not show an overall effect, it is tempting to look for a group in which it “worked after all.” With a sufficient number of breakdowns, something will always be found. Segments for analysis should be declared before launch, and exploratory findings should be treated as hypotheses for separate testing.
- Novelty effect and habituation effect. Short tests of an interface, display, or message may measure the reaction to change itself. Longer measurement makes it possible to check whether the effect persists.
- Confusing statistical significance with practical significance. With very large samples, even a minimal difference may be statistically significant. The question is whether its magnitude justifies the cost of implementation – which is why the minimum practically meaningful effect is defined in advance.
The limitations of the method itself are a separate issue. Experimental research measures effects well, but does not necessarily explain the mechanism on its own: it shows that variant B performed better, without always explaining why. It is therefore often combined with a qualitative component – post-exposure interviews, usability testing, or research into responses to sales arguments. An experiment is also not suitable where variants cannot be separated across audiences (a mass-media campaign without the possibility of geographic segmentation), where the number of units is very small (a few key B2B customers), or where differentiating conditions raises ethical or legal concerns. In such situations, quasi-experimental designs are used: selecting a comparison group through matching, difference-in-differences analysis, or geographic tests with synthetic control. They usually provide a weaker basis for causal inference than randomization, but when their assumptions are met, they may be stronger than a simple “before and after” comparison.
When is it worth using an experiment, and when another method?
The list below organizes situations in which experimental research is the right tool, as well as those in which another research approach is better suited.
- Experiment – when the question concerns the impact of a specific, controlled change on behavior: a price variant, package structure, message content, display format, or contact sequence.
- Field experiment – when the change is to be implemented in a real sales environment and it is important for measurement to reflect actual purchasing behavior rather than declarations.
- Declarative research (survey, conjoint) – when the tested element does not yet exist in a form that can be brought to market, or when many combinations of attributes need to be compared at once.
- Qualitative research – when an understanding of customer motivations, language, and barriers is needed, meaning material for developing variants that will later be used in an A/B test.
- Historical data analysis – when an experiment is impossible; it can be used to generate hypotheses and, with an appropriate quasi-experimental design, also for cautious causal inference.
In Hume’s Institute projects, the best results come from a sequence in which qualitative research develops hypotheses, experiments verify them, and continuous measurement monitors the durability of the effect after implementation. A single test resolves one question – the value of the method increases when testing becomes a routine rather than a one-off undertaking.
Frequently asked questions
How does an experiment differ from correlational research?
In correlational research, the researcher observes the co-occurrence of variables in data that they do not control – they may demonstrate an association, but usually not its direction or whether it results from a third, unaccounted-for factor. In an experiment, the researcher introduces a change and randomly assigns units to conditions, so the groups should differ systematically only in terms of the stimulus being studied. Therefore, a properly conducted experiment provides a stronger basis for drawing causal conclusions.
How long should an A/B test last?
Duration is determined by the required sample size and the length of the purchasing cycle, not by the project calendar. Usually, the test should cover full weeks to balance variation between days of the week and, in categories with a longer decision-making process, a period covering the typical time from first contact to purchase. The end date is set before launch and should not be shortened after a favorable result is observed.
When is an A/B test result statistically significant?
It is statistically significant when the p-value, meaning the probability of obtaining, assuming there is no real effect, a result at least as consistent with the alternative hypothesis as the observed result, is lower than the pre-specified threshold (most often 0.05), given the analysis planned in advance. Significance depends on the effect size, number of observations, and variability of the metric – with large samples, even small differences may be significant. Therefore, effect size and the confidence interval are reported alongside significance, and the implementation decision is assessed against the previously defined threshold of practical significance.
Ask about designing an experiment for your offer or communication
If you are facing a decision about changing your offer, price, message, or approach to customer service, it is worth resolving it through a test rather than intuition. Hume’s Institute helps design experiments – from the hypothesis and randomization, through sample size, to results analysis; get in touch to discuss the scope of the study.