Experimentation is evidence to determine whether something works. Intent decides who receives what, when (or whether they need it at all). “Does this work.” Think of this as 80% exploration, and 20% exploitation.
The first works towards the best single answer for everybody to prove a hypothesis. The second works out where a single answer was never going to be enough; scaling a proven hypothesis for true commercial gain. “For whom does it work best, when should it be shown, and do a proportion of users even need it at all?” Think of this as 20% exploration, and 80% exploitation. Where testing tells you whether an experience works, but Intent decides who receives it, when, and whether they need it at all.
In this blog post, we'll identify the differences between Intent and experimentation, so you can understand when to utilise each strategy to further your business goals.
Defining A/B testing
An A/B test answers one single bounded question.
Does this change (a treatment) move the metric we care about, across our traffic, better than the alternative does (a control)?
It's a hypothesis, and what comes out the other end is 80% knowledge (learnings) and 20% value (hopefully positive gain). A primary "does this prove or disprove our hypothesis" binary answer.
Our intent engine answers a different question, and it answers it again every few seconds throughout a users session:
Given what this visitor is doing right now, which of the experiences we already know works suits them at this moment in time? (If any)
What comes out of that is not just a single decision, but hundreds of them a second, all trying to move the agents primary goal, or reward. The open question stopped being does this work a long time ago. What's left is who needs it, and when.
For example: if a visitor starts showing signs of basket abandonment. You can put a returns message in front of them, or a discount, or a trust signal, because you already trust all three of those mechanics. What you don't know is which one that particular person needs, at what point, or whether they need any of them at all. That's the decision the agent is making.

There are several reasons why experimentation (read: split testing) might not be an appropriate mechanism; here are four reasons:
But first, a note on the word experimentation. Throughout this set, experimentation means the whole practice of using evidence to drive ideas: user research, usability testing, prototyping, quasi-experiments and switchbacks where a split isn't possible, and A/B tests where it is. The split test is one instrument inside it, usually the last step rather than the whole of it.
1. Averages can lie.
An A/B test reports an average treatment effect which can hide segments or audiences that have a detrimental or negative impact.
An A/B test reports an average treatment effect. That's the whole point of it. Randomising across your population is what buys you a causal claim in the first place. The price you pay is that the average swallows everything inside it. A headline +2% can be +12% for hesitant first-timers and -4% for loyal returners. Same test. Same green tick. Two completely different stories.
Statisticians call these heterogeneous treatment effects, and there's good evidence that acting on them beats a blanket rollout. In one worked simulation of a ranking change, treating only the 69% of users predicted to benefit delivered a 2.2% gain in revenue per user against shipping it to everybody. A CUNY study of financial aid nudges captured roughly 75% of the total benefit while treating half the population.

Experimentation is deliberately built to stop you trusting the segment breakdown. Pre-registration, power calculations, no post-hoc slicing. Those rules exist because sliced results are where false positives breed. Kohavi, Deng and Vermeer's A/B Testing Intuition Busters (KDD 2022) found that underpowered analyses inflate the effects you observe by 25 to 50%. At the power levels most mid-market programmes actually run at, over half of your significant results are false positives. So the question you have to answer to deploy well, for whom, is the question your test is worst at answering. That's a boundary, and a sensible one. Intent sits on the other side of it. It predicts for every visitor in advance, rather than carving up a sample after the fact.
Where Intent solves this problem: The experience is being delivered across hundreds of contexts, reallocating to where the experience "works" (where it's best seen and best felt), excluding it from where "it doesn't work." In theory, it only looks and focuses on the good, excluding the bad.
2. Testing assumes the trigger.
Every A/B test starts from a fixed, assumed rule (usually on page load). The trigger "when to show it" is often assumed, arbitrary, and based on a website proxy like "3 page views" or "PDP" rather than genuine user intent.
Nearly every A/B test starts from a fixed rule. Show this on the PDP after three pageviews. Fire this on exit intent. Serve this to returning visitors. Then it tests which creative performs best inside that rule.
The treatment gets examined. The trigger gets assumed.
There's a practical reason for that. The trigger space is far too big to write out by hand. Nobody can enumerate every combination of buying stage, purchase confidence, abandon risk, intent trend and product affinity, let alone power a test across them. So teams pick a proxy or an average and move on.
Proxies are usually where the damage happens. These are largely based on website attributes as opposed to personalised signals. Returning visitor stands in for higher intent. Mobile stands in for lower intent. Pageview count stands in for engagement, when a confused shopper racks up far more pageviews than a decisive one, and high intent visitors often convert inside three.
Our own Intent Gap research found 10% of visitors on checkout pages aren't ready to buy. And 34% never get past browsing, whatever page they happen to be on. There's a reason we say stages, not pages.
Where Intent solves this problem: Intent agentic delivery serves the individual at a moment in time. A user might need an experience on the 3rd page view after 30 seconds, a different user might need it after 50 seconds, a different user might need it after 9 page views and 1 second. The continuous prediction modelling gives a threshold to meet first, and then a mechanism to automate a decision second.
3. Statistical significance is a search for one answer.
The search for statistical significance can a) slow you down and b) inhibit your ability to personalise to different segments.
Significance exists to license a claim about a single population, and the claim is singular by design where B beats A: one answer, for everyone. That is exactly what costs time.
You need a sample big enough to speak for the whole population with a minimum detectable effect. If you want the claim to be about a segment instead (what some call "personalisation") you need that sample inside the segment too, so the wait multiplies. Most claim they don't have the traffic to personalise because of this.
Experimentation vendor's posterior is a belief about a variant. One distribution per arm, computed across everybody who saw it, resolving to a single winner for all your traffic. Better maths than a p-value, and the shape of the answer is identical. One number per variant, and a threshold you wait to cross before you act.
Our posterior at Made with Intent is a belief about a variant given a context. Where their question is whether B beats A, ours is whether B beats A for a visitor in this intent state, at this point in their session. The model predicts reward conditional on context, so what comes out is a policy; hundreds of answers, running at once.

Put it this way: instead of a human designing three segments and running three underpowered tests, the winner goes to the agent alongside its alternatives and the segment level optimisation runs continuously. Three mechanisms make that more traffic-efficient than a segmented test. Worth understanding rather than taking on trust.
1. Allocation is adaptive. Thompson Sampling draws plausible performance rates from each variant's posterior and routes traffic accordingly, so your exploration cost falls as confidence climbs rather than sitting at 50% for the duration. In practice agentic campaigns push 80 to 90% of traffic to the winner inside two to three weeks. A fixed-allocation A/B test takes around six. Optimised Control shrinks the control group automatically as confidence grows.
2. Without waiting for statistical significance. It borrows statistical strength across slices. This is the important one. Made With Intent runs a contextual bandit, not a multi-armed one. A multi-armed bandit only learns from the data each arm receives, so every segment needs its own volume, which is the same problem as a segmented A/B test in different clothing. A contextual bandit trains a model that predicts reward given context, so it can estimate performance for a combination it has never directly seen. Show it mobile traffic, low intent traffic and a particular page separately and it can predict for "mobile, low intent, that page" by generalising from how each component behaves. A hand-built segment can't borrow like that. Every cell starts at zero.
3. You're not assuming which segment to slice or review. The agent can build up to 500 contextual combinations and up to 75 moment-based triggers per campaign, and it trims its own feature selection to keep each slice sample-rich. Nobody has to guess which distinctions matter before there's evidence about which ones do.
Why we don't report per-segment significance
This surprises experimentation-native teams more than anything else in the product. It's a deliberate methodological choice.
Run an independent significance test on every intent combination and multiple comparisons swallow the results. At α = 0.05, roughly 14 independent tests give you about a 50% chance of at least one false positive. A hundred tests will throw up around five false positives by chance alone, and across hundreds of combinations you'd be manufacturing findings.
We trialled per-segment significance reporting and pulled it, because it surfaced spurious and inverse correlations. For example, users traverse through different stages of intent in their journey, at what point is a user in low intent and when did they see the treatment?
The per-segment decisions get justified differently. Is the model calibrated, meaning when it says 70% does it convert around 70% of the time (measured by expected calibration error)? Is it discriminating, meaning can it rank likely converters above unlikely ones (AUC around 0.84)? A well calibrated probability is a legitimate basis for acting on a slice. An underpowered significance test on that same slice is not.

This calls back to something we talked about earlier.The reason you shouldn't chase per-segment significance in your testing tool is the same reason we don't chase it in ours. Prediction is the right instrument for the for whom question, retrospective slicing isn't.
Why we don't report on single variant uplifts
You get one clean causal number, control against the agent's allocation, measured the way any experimentation lead would want it measured. You give up per-arm inference, and in exchange the arm is adaptive. That's the trade. An A/B test gives you clean inference on every arm and one winner for everybody. An agentic campaign gives you clean inference on one arm, and a different winner in every context.
Load five experiences into an agentic campaign and it looks like a five-arm test. It isn't. It's actually a two-arm experiment, the same shape as any A/B test you've run. The difference being:
• Control. A baseline you define.
• Variant. The agent's allocation across all five experiences.
A traditional A/B test asks which variant wins with a blanket rule. An agentic campaign asks a different question with an adaptive rule (it's a reason why we call them campaigns, not experiments)
Does allocating these experiences by intent beat applying a blanket rule?
We don't report on single variant uplifts eg: Variant B is better by 5% because:
1. There's no single number to give you. Experience B might be the winner for high intent, budget-conscious mobile visitors and the loser for low intent desktop. That's the entire point of running it this way. Collapsing it to one figure puts back exactly the average this campaign exists to get rid of.
2. The arms were never randomised against control. Traffic reaching experience B was chosen by the agent, on context and accumulated evidence. It isn't a random slice of your audience so there's no effective control. Comparing a deliberately selected group against everybody is confounded by construction, and nothing in the data separates the effect of the experience from the effect of the selection.
3. You'd be reading a state the system has already left. If experience B was struggling in a context, the agent would have noticed days ago and moved traffic away (hence: reallocation). The number in front of you is an average across a period in which allocation was actively changing, and the campaign has moved on since. Acting on it means acting on history the agent has already corrected for.
That distinction does three things, in order.
1. It changes what you can personalise. A per-variant posterior can only ever produce one answer for everyone, however good the statistics behind it are. Varying what people get requires probability conditional on the person. That's the whole game, and it matters far more than the frequentist argument ever did.
2. It changes when you can act. Certainty becomes a dial rather than a gate. At 62% probability, Thompson Sampling tilts the allocation and updates again tonight. Nobody declares anything. A stopping rule stops being necessary once allocation is continuous.
3. It changes how fast you get there. Nothing is waiting for a threshold, so allocation improves from the first night onward. There's no finish line to reach before the work starts paying.
Where Intent solves this problem: Our posterior at Made with Intent is a belief about a variant given a context. Where experimentation question is whether B beats A, ours is whether B beats A for a visitor in this intent state, at this point in their session. The model predicts reward conditional on context, so what comes out is a policy; hundreds of answers, running at once based on a series of probabilities using a contextual bandit.
4. Limited to a single answer for only one point in time.
An A/B test only tells you what was true of your traffic during the weeks the test ran.
An A/B test tells you what was true of your traffic during the weeks the test ran; a shelf life. Nothing about that answer refreshes itself so brands often end up "re-testing" the same thing a few years later, hoping there's no change from a previous positive test. The world moves on, results degrade as users become "used" to the treatment, competitors copy; not to mention that your purchase lifecycle varies by days, weeks, months, years.

For example: a Microsoft experiment on MSN.com saw replacing one button with another produce a 4.7% increase in overall clicks. The daily breakdown post launch showed the difference decreasing rapidly day over day as users learned the change, leading to the team shut the experiment down mid-way. The conclusion was that the observed treatment effects "are not always permanently stable, sometimes revealing increasing or decreasing patterns over time."
There are a few reasons why this happens.
1. Novelty and primacy wear off. A new element gets attention because it's new. Returning visitors are briefly worse off because it isn't what they knew (which is why tests are often split, somewhat arbitrarily, between new and returning users)
2. Buying patterns differ throughout the year. December traffic behaves nothing like March traffic, different intent, different price sensitivity, different tolerance for being interrupted (especially for retailers)
3. The population itself drifts. Mix of paid traffic allocation shifts, category mixes alter or a competitor changes their delivery proposition. Macro-economic factors alter the interaction effect.
Does a mature site slowly accumulate a layer of decisions that were correct once, are serving everybody by default, yet are answerable to nobody? It is entirely possible to have a well-governed testing function and a site full of expired answers at the same time.
Where Intent solves this problem: Continuous allocation doesn't have this failure mode, because it never declares anything. The posterior is a live belief rather than a verdict, continually reallocating and therefore responding to changes, not static. Allocation is re-scored against current behaviour and retrained nightly, and standing exploration keeps a small share of traffic asking whether the current answer is still the right one. When February stops behaving like November, the agent finds out because it never stopped looking.
Intent vs experimentation: In depth
We've done a quick reference table on the differences between Intent and experimentation. Take a look, send it to your colleagues. You're welcome.
Experimentation and Intent aren't solving the same problem, and treating them as substitutes is where most testing programmes stall. An A/B test proves whether something works for everybody, on average. But it can't tell you who needs a returns message versus a discount versus nothing at all, and it becomes redundant the moment the traffic that validated it moves on.
Intent picks up where that boundary sits, deciding who an experience is for, when, and whether they need it at all, re-scored every few seconds instead of declared once and left to expire. Most eCommerce teams already have the first instrument. The second is what turns that evidence into revenue, visitor by visitor, instead of one rollout for everybody.
Book a demo to see how Made With Intent can help you deliver more appropriate experiences to your customers.
Become an Intent Insider
Get subscriber-only insights we don't publish anywhere else and event invites before anyone else.
By submitting this form you agree to our (more than fair) terms.
















