This is the third of a three part series where we talk about how Experimentation and Intent should work together. In this series, we cover theory, practical examples and when Intent should and shouldn't be used.
We'll pop the links to the other pieces in the series below, so you can jump between each blog post. Again, I'll hand you over to David:
Part 1 covers the differences between experimentation and Intent.
Part 2 talks about when you should use Intent vs experimentation.
Introducing the Four Intent Modes
IWe've already established when using Intent is a candidate for an experience. Asking two simple questions: → Will the effect vary by person and moment? and → Do you have evidence that it works? tends to do the trick There are four different modes based on that candidate selection that help you determine whether something is fit for Intent vs a split test.
1. Just do it
2. Test it
3. Sequence it (most popular)
4. Hand it straight to Intent (best for commercial impact)
So, what are suitable candidates for Intent?
The two core questions that determine whether an experience or tactic should be classified as a candidate for Intent can be seen in the below 2x2 chart:
• Will the effect vary by person and moment?
• Do you have evidence that it works?
Mode 4: "Hand it straight to Intent" is the most popular and the one that we'd recommend. Why?
• we operate largely within commercial teams for the purpose of value maximising
• think of agentic experience delivery as "getting the most of what you've already done". Squeezing the juice out the lemon
• most teams already know what works, it's about who receives that and when they receive that message. A useful message eg. free shipping is far more impactful to a targeted or focussed segment, than served statically and generically where it's fighting for attention and real-estate against competing messages.
Mode 1. Just do it
In short. Not everything deserves a test. Fix what's broken and ship what's obviously better or genuinely unambiguous. Mode 1 is the narrowest of the four (not to be mistaken for the most convenient).
In theory, everything is better off being measured. In practice, that's not the case.
We stand by the ethos that some changes don't need permission from a p-value. If something's broken, fix it. If customers keep asking where the delivery information is, put it where they're looking. If legal says the copy changes, the copy changes.
Some common scenarios that determine whether Mode 1: just do it is the right approach:
• You'd ship it whatever the result. If the decision is strategic, or brand-led, or the answer wouldn't change what you do next, the test is nothing but decoration.
• The test would be underpowered or the downside is small. At low power most of your significant results are noise. Low-traffic pages make conventional split testing impractical because significance can take months to reach. Kohavi, Deng and Vermeer explain that in most mid-market programmes over half of significant results are false positives.
• When the change is obviously better. If payments are failing and you know an explanatory error message helps, a test keeps half your visitors on the broken experience while you gather evidence for something you already believe. Just do it. (As Nike would say.)
A lack of resource within the team should not be the reason for not doing something. This is a more common argument amongst less mature experimentation programs.
As these programme mature, releases are wrapped in feature flags which costs almost nothing extra to measure them. In other words when measurement is part of the development workflow rather than the marketing one, resource is not often an excuse, and this mode becomes less relevant.

Mode 2. Test it.
In short. A split test gives evidence behind a claim, proving the direction of travel, giving a binary answer to the question "does this work?".
If your aim is to learn first, where the intent effect should not vary by people, buying stages, or intent and therefore the aggregate impact or learning matters more than the specific buying stage that the user is at; split test away.
The split test is incredibly powerful; helping brands understand whether a treatment is better = yes or no. The output is a binary effect designed to prove or disprove a hypothesis "does this work"? Note: it doesn't tell you by how much the treatment is different.
This gives brands controlled learnings, usually at an aggregated level: on average, our customers preferred A rather than B and here's the evidence. If the point of the test is to understand your customers well enough to plan next quarter, a clean split test is the right tool to use and Intent isn't appropriate.
The output is one singular answer, everybody gets it, and particularly when you want a causal claim before you commit engineering time to it. This is why we often like asking the question "Will the effect vary by person and moment?" If the answer is no, what we'll find is one treatment will impact most the same way (assumed, or otherwise)
Mode 3. Sequence it: Test, then let Intent scale
Most popular. Recommended for process and workflow orientation to reduce risk.
In short. Sequence the process by split testing first and then amplifying using agentic experience delivery. The first two steps are largely exploratory (80% exploration, 20% exploitation), the second two steps are largely exploitation (20% exploration, 80% exploitation)
1. Test broadly
2. Examine the result
3. Allocate it agentically
4. Measure it commercially.
One customer did exactly this with their next-day delivery countdown. Tested live they saw value. Then tested agentically across hundreds of different intent context against that blanket version, and saw even more.
One third of the reason why Personalisation doesn't work is because the metrics for success (eg. transactions) are largely incorrect. If personalisation is the act of being personal, measuring that on whether someone transacts is folly. There's nothing quite more impersonal than personalising for revenue purposes.
Another third, is that personalisation, to date, tests on historic website attributes, not in-moment personal ones. Being retrospective isn't really the biggest issue here being honest. A visitor that did something isn't as powerful as a visitor who is doing something, sure.
There's nothing more indicative of a users level of intent than what they are doing in moment. That being said, the majority of these personalisation efforts have been based on website attributes. A page view, a channel type, an exited session. That's not how the visitor is feeling or what they want, it's simply a series of proxies that try to infer a state of mind. Poorly, might we add.
The final third why personalisation has fallen over is because it is inherently not scalable. Rules tend to dominate the industry. Have you ever thought why the only successful form of personalisation to date has been recommendations? It's because it's algorithmic and therefore scalable.
1. Define your segments
2. Run a test inside each one
3. Ship each segment its own winner.
Not only is this manual and therefore not scalable.
Not only is this guess work (Why that segment? Why not another segment?).
But most sites don't have the traffic to test like this, so most revert to testing on the average. Splitting traffic across three segments means that every variant gets a third of the volume, so each test takes roughly three times as long to reach the same power.

To that end, we recommend you test first amongst the general population. That way you can add evidence across the average to give a causal claim. It answers the second question we'd want answering before you test: Do you have evidence that it works?
The price you pay, however, is that the average swallows everything inside it. Where "winners" can be hidden within "losers" and "losers" hidden within "winners". A headline +2% can be +12% for hesitant first-timers and −4% for loyal returners. It's the same test, the same green tick but two completely different stories. Why? Because user intent varies and oscillates across the session (that's why we exist).
Once proven through testing, then you expand.
Agentic experience delivery identifies what works for who and when, allocating traffic to where the test is best seen and felt. Even better, it's based on the context of continuously reading user intent. To that end, it maximises the impact of a test already run. So once the average is proven, you can then amplify.
From here, we recommend a four stage sequence. TEAM.
1. Test. Add evidence behind your hypothesis and establish that the experience (the "what") works at an aggregated level. Note: Read the T broadly because it is not only the split test, but the research, the ideation and the design that came before it. The test measures the idea but does not supply it.
2. Examine. Analyse, read and learn at an aggregated level for iterative hypotheses formation.
3. Allocate. Make the experience adaptive, agentically delivered on someone's intent, for the dual purpose of scaling and exploiting it i.e use Intent.
4. Measure. Prove it out commercially, against a baseline that saw nothing.
Simply put, where the first two "T" and "E" are about exploration, the "A" and "M" are more geared towards exploitation.
Mode 4. Hand it straight to Intent
Recommended for commercial impact
In short. This is a segmentation decision. You already know what you need as the experience or tactic because its already proven in its form. Most brands are sitting on a pile of interventions they already believe in, all firing flat (what we call: static and generic). What you are still learning is who and when. Here, the thing under test is the allocation, the timing, the triggering.
This is still a test, per se. There is still a control; nobody is being asked to take the result on faith. That being said, the variables being tested are the timing mechanism and the individual who sees it; unconventionally rare within experimentation.
This is our sweet-spot (obviously).
You already run tactics like basket recovery, email capture, social proof, delivery countdowns and free shipping thresholds. Nobody needs a test to prove those mechanics can work because they're staples within the ecommerce armoury (and the odds are you ran that test years ago)
What you don't know is who needs them or when they need it. Indeed, whether showing it is a distraction or friction point to them, preventing a conversion. That's a segmentation decision, the green shoots of personalisation, and it's the one you'd rather not keep making by hand. That rules-based pre-segment approach is not scalable.
There are two reasons Mode 4 is safer than it sounds.
1. The control never switches off. A 10% global holdback gives you a causal read without running a testing period at all. There's no "let's test it first" step, because the test doesn't end.
2. Suppression pays for itself. Even where the average effect is known and positive, finding the visitors who don't need the intervention is worth real money. One customer recovered 42% of the margin they'd been giving away by not serving the discount to everyone. Where a flat split test would suggest that the test "won", not everyone needed that tactic in order to convert.

Let's summarise this blog post
Testing answers does this work. Intent answers for whom, at what moment, and do they even need it. The two hold different purposes and therefore the "candidates for each" should also be treated differently.
The first question (the split test) has an owner in almost every ecommerce team (experimentation function, CRO function, product, engineering or marketing are usual go to's)
The second usually doesn't have a single owner given it's uniqueness. It's use cases are spread throughout trading, CRM, merchandising and, of course, experience (in what ever form that's called in the organisation).
The most common scenario is that the split test validates the "what" and Intent can then finish the job by scaling and amplifying to the right people at the right time. Scalable personalisation; that's mode 3.
Mode 4 (Hand it straight to Intent) is more commonly our de-facto as the majority of the time, you already know what you need as the experience or tactic because its already proven in its form. Especially if the answer to the below questions is "yes"
• Will the effect vary by person and moment?
• Do you have evidence that it works?
Of course, keep investing in experimentation to find what works. This is about amplifying experiences, messages, tactics that have evidence behind them; more commercial impact. Once that evidence is identified or known, add Intent so that what works reaches the people it was built for, at the moment it matters, and stops reaching the people who were always going to buy anyway.
So, there you have it, the third and final part of our experimentation and Intent series. You can catch up on the first two here.
Part 1 covers the differences between experimentation and Intent.
Part 2 talks about when you should use Intent vs experimentation.
If you've found this blog post interesting, and want to learn more about Intent, book a demo and our sales team will be in touch.
Become an Intent Insider
Get subscriber-only insights we don't publish anywhere else and event invites before anyone else.
By submitting this form you agree to our (more than fair) terms.




