How experimentation and Intent should work together (Part 3): The Four Modes

David is still writing. This time, he's come with his patented "Intent Modes". This is a good one, so click to learn more.
Read Time:
On this page

This is the third of a three part series where we talk about how Experimentation and Intent should work together. In this series, we cover theory, practical examples and when Intent should and shouldn't be used.

We'll pop the links to the other pieces in the series below, so you can jump between each blog post. Again, I'll hand you over to David:

Part 1 covers the differences between experimentation and Intent.

Part 2 talks about when you should use Intent vs experimentation.

Introducing the Four Intent Modes

IWe've already established when using Intent is a candidate for an experience. Asking two simple questions: → Will the effect vary by person and moment? and → Do you have evidence that it works? tends to do the trick There are four different modes based on that candidate selection that help you determine whether something is fit for Intent vs a split test.

1. Just do it

2. Test it

3. Sequence it (most popular)

4. Hand it straight to Intent (best for commercial impact)

So, what are suitable candidates for Intent?

The two core questions that determine whether an experience or tactic should be classified as a candidate for Intent can be seen in the below 2x2 chart:

• Will the effect vary by person and moment?

• Do you have evidence that it works?

Effect is much the same for everyone Effect varies by person and moment
You don't yet have evidence that it works Test it (Mode 2) Sequence it: Test, then let Intent scale (Mode 3)
You already have evidence that it works Just do it. (Mode 1) Hand it straight to Intent (Mode 4)

Mode 4: "Hand it straight to Intent" is the most popular and the one that we'd recommend. Why?

• we operate largely within commercial teams for the purpose of value maximising

• think of agentic experience delivery as "getting the most of what you've already done". Squeezing the juice out the lemon

• most teams already know what works, it's about who receives that and when they receive that message. A useful message eg. free shipping is far more impactful to a targeted or focussed segment, than served statically and generically where it's fighting for attention and real-estate against competing messages.

Mode 1. Just do it

In short. Not everything deserves a test. Fix what's broken and ship what's obviously better or genuinely unambiguous. Mode 1 is the narrowest of the four (not to be mistaken for the most convenient).

Use it when It's broken, or obviously right, or you'd ship it whatever a test said 😂
What you get out of it Speed. (albeit with a lack of evidence or learnings)
A good question to ask yourself "If you were to test this, what would you do with a neutral result?"
Examples
  • Trustpilot is broken
  • Filters don't work effectively
  • Navigation needs trimming down
  • Free shipping threshold not communicated

In theory, everything is better off being measured. In practice, that's not the case.

We stand by the ethos that some changes don't need permission from a p-value. If something's broken, fix it. If customers keep asking where the delivery information is, put it where they're looking. If legal says the copy changes, the copy changes.

Some common scenarios that determine whether Mode 1: just do it is the right approach:

• You'd ship it whatever the result. If the decision is strategic, or brand-led, or the answer wouldn't change what you do next, the test is nothing but decoration.

• The test would be underpowered or the downside is small. At low power most of your significant results are noise. Low-traffic pages make conventional split testing impractical because significance can take months to reach. Kohavi, Deng and Vermeer explain that in most mid-market programmes over half of significant results are false positives.

• When the change is obviously better. If payments are failing and you know an explanatory error message helps, a test keeps half your visitors on the broken experience while you gather evidence for something you already believe. Just do it. (As Nike would say.)

A lack of resource within the team should not be the reason for not doing something. This is a more common argument amongst less mature experimentation programs.

As these programme mature, releases are wrapped in feature flags which costs almost nothing extra to measure them. In other words when measurement is part of the development workflow rather than the marketing one, resource is not often an excuse, and this mode becomes less relevant.

Mode 2. Test it.

In short. A split test gives evidence behind a claim, proving the direction of travel, giving a binary answer to the question "does this work?".

If your aim is to learn first, where the intent effect should not vary by people, buying stages, or intent and therefore the aggregate impact or learning matters more than the specific buying stage that the user is at; split test away.

Use it when Structural change everybody experiences identically.
What you get out of it Proof. A causal claim before you commit engineering
A good question to ask yourself "Does this thing work?"
Examples
  • Does adding product benefits help or hinder?
  • Is it better to show brand imagery or product imagery?
  • Does customers respond better to a one page checkout flow, or a 3 step checkout flow?

The split test is incredibly powerful; helping brands understand whether a treatment is better = yes or no. The output is a binary effect designed to prove or disprove a hypothesis "does this work"? Note: it doesn't tell you by how much the treatment is different.

This gives brands controlled learnings, usually at an aggregated level: on average, our customers preferred A rather than B and here's the evidence. If the point of the test is to understand your customers well enough to plan next quarter, a clean split test is the right tool to use and Intent isn't appropriate.

The output is one singular answer, everybody gets it, and particularly when you want a causal claim before you commit engineering time to it. This is why we often like asking the question "Will the effect vary by person and moment?" If the answer is no, what we'll find is one treatment will impact most the same way (assumed, or otherwise)

Mode 3. Sequence it: Test, then let Intent scale

Most popular. Recommended for process and workflow orientation to reduce risk.

In short. Sequence the process by split testing first and then amplifying using agentic experience delivery. The first two steps are largely exploratory (80% exploration, 20% exploitation), the second two steps are largely exploitation (20% exploration, 80% exploitation)

1. Test broadly

2. Examine the result

3. Allocate it agentically

4. Measure it commercially.

One customer did exactly this with their next-day delivery countdown. Tested live they saw value. Then tested agentically across hundreds of different intent context against that blanket version, and saw even more.

Use it when A net-new intervention you don't trust yet, on decent traffic
What you get out of it Maximisation. Proof it works, then proof it works for the right people at the right time
A good question to ask yourself "Does or will the experience or tactic differ by context of the user i.e their buying stage, intent?"
Examples
  • Should we serve email capture actively? And then to who, and when?
  • Should we showcase product recommendations after the user adds to basket? Does this work for everyone?
  • Should we promote USP 1, USP 2 or USP 3 generically or at different stages of the user journey?

One third of the reason why Personalisation doesn't work is because the metrics for success (eg. transactions) are largely incorrect. If personalisation is the act of being personal, measuring that on whether someone transacts is folly. There's nothing quite more impersonal than personalising for revenue purposes.

Another third, is that personalisation, to date, tests on historic website attributes, not in-moment personal ones. Being retrospective isn't really the biggest issue here being honest. A visitor that did something isn't as powerful as a visitor who is doing something, sure.

There's nothing more indicative of a users level of intent than what they are doing in moment. That being said, the majority of these personalisation efforts have been based on website attributes. A page view, a channel type, an exited session. That's not how the visitor is feeling or what they want, it's simply a series of proxies that try to infer a state of mind. Poorly, might we add.

The final third why personalisation has fallen over is because it is inherently not scalable. Rules tend to dominate the industry. Have you ever thought why the only successful form of personalisation to date has been recommendations? It's because it's algorithmic and therefore scalable.

1. Define your segments

2. Run a test inside each one

3. Ship each segment its own winner.

Not only is this manual and therefore not scalable.

Not only is this guess work (Why that segment? Why not another segment?).

But most sites don't have the traffic to test like this, so most revert to testing on the average. Splitting traffic across three segments means that every variant gets a third of the volume, so each test takes roughly three times as long to reach the same power.

To that end, we recommend you test first amongst the general population. That way you can add evidence across the average to give a causal claim. It answers the second question we'd want answering before you test: Do you have evidence that it works?

The price you pay, however, is that the average swallows everything inside it. Where "winners" can be hidden within "losers" and "losers" hidden within "winners". A headline +2% can be +12% for hesitant first-timers and −4% for loyal returners. It's the same test, the same green tick but two completely different stories. Why? Because user intent varies and oscillates across the session (that's why we exist).

Once proven through testing, then you expand.

Agentic experience delivery identifies what works for who and when, allocating traffic to where the test is best seen and felt. Even better, it's based on the context of continuously reading user intent. To that end, it maximises the impact of a test already run. So once the average is proven, you can then amplify.

From here, we recommend a four stage sequence. TEAM.

1. Test. Add evidence behind your hypothesis and establish that the experience (the "what") works at an aggregated level. Note: Read the T broadly because it is not only the split test, but the research, the ideation and the design that came before it. The test measures the idea but does not supply it.

2. Examine. Analyse, read and learn at an aggregated level for iterative hypotheses formation.

3. Allocate. Make the experience adaptive, agentically delivered on someone's intent, for the dual purpose of scaling and exploiting it i.e use Intent.

4. Measure. Prove it out commercially, against a baseline that saw nothing.

Simply put, where the first two "T" and "E" are about exploration, the "A" and "M" are more geared towards exploitation.

Stage The question Owner What governs it
1 Test Does this work at all? Experimentation Pre-registration, power calculation, fixed duration
2 Examine Segment analysis; who did it help, who did it hurt? Experimentation and Experience, jointly Post-segment analysis, determining if the effect varied amongst user groups
3 Allocate Which visitors get it, and at what moment? Experience or Ecommerce, with Trading Continuous reallocation using bayesian contextual allocation
4 Measure What did the whole programme actually cause? Experimentation owns the method. Commercial consumes the output. Global 10% holdback

Mode 4. Hand it straight to Intent

Recommended for commercial impact

In short. This is a segmentation decision. You already know what you need as the experience or tactic because its already proven in its form. Most brands are sitting on a pile of interventions they already believe in, all firing flat (what we call: static and generic). What you are still learning is who and when. Here, the thing under test is the allocation, the timing, the triggering.

Use it when The mechanic is already validated. Your real question is who and when. What is the value of allocating this test?
What you get out of it Speed, real commercial uplift for only those users that require it, shoots of personalisation
A good question to ask yourself "Who or when should we serve this tactic or experience to, to maximise it's impact?"
Examples
  • Showing a next day delivery timer to maximise impact
  • Not showing discounts or email captures to everyone
  • Serving product-level messaging sparingly so they each hold impact

This is still a test, per se. There is still a control; nobody is being asked to take the result on faith. That being said, the variables being tested are the timing mechanism and the individual who sees it; unconventionally rare within experimentation.

This is our sweet-spot (obviously).

You already run tactics like basket recovery, email capture, social proof, delivery countdowns and free shipping thresholds. Nobody needs a test to prove those mechanics can work because they're staples within the ecommerce armoury (and the odds are you ran that test years ago)

What you don't know is who needs them or when they need it. Indeed, whether showing it is a distraction or friction point to them, preventing a conversion. That's a segmentation decision, the green shoots of personalisation, and it's the one you'd rather not keep making by hand. That rules-based pre-segment approach is not scalable.

There are two reasons Mode 4 is safer than it sounds.

1. The control never switches off. A 10% global holdback gives you a causal read without running a testing period at all. There's no "let's test it first" step, because the test doesn't end.

2. Suppression pays for itself. Even where the average effect is known and positive, finding the visitors who don't need the intervention is worth real money. One customer recovered 42% of the margin they'd been giving away by not serving the discount to everyone. Where a flat split test would suggest that the test "won", not everyone needed that tactic in order to convert.

Let's summarise this blog post

Testing answers does this work. Intent answers for whom, at what moment, and do they even need it. The two hold different purposes and therefore the "candidates for each" should also be treated differently.

The first question (the split test) has an owner in almost every ecommerce team (experimentation function, CRO function, product, engineering or marketing are usual go to's)

The second usually doesn't have a single owner given it's uniqueness. It's use cases are spread throughout trading, CRM, merchandising and, of course, experience (in what ever form that's called in the organisation).

The most common scenario is that the split test validates the "what" and Intent can then finish the job by scaling and amplifying to the right people at the right time. Scalable personalisation; that's mode 3.

Mode 4 (Hand it straight to Intent) is more commonly our de-facto as the majority of the time, you already know what you need as the experience or tactic because its already proven in its form. Especially if the answer to the below questions is "yes"

• Will the effect vary by person and moment?

• Do you have evidence that it works?

Of course, keep investing in experimentation to find what works. This is about amplifying experiences, messages, tactics that have evidence behind them; more commercial impact. Once that evidence is identified or known, add Intent so that what works reaches the people it was built for, at the moment it matters, and stops reaching the people who were always going to buy anyway.

So, there you have it, the third and final part of our experimentation and Intent series. You can catch up on the first two here.

Part 1 covers the differences between experimentation and Intent.

Part 2 talks about when you should use Intent vs experimentation.

If you've found this blog post interesting, and want to learn more about Intent, book a demo and our sales team will be in touch.

// the intent insider

Become an Intent Insider

Get subscriber-only insights we don't publish anywhere else and event invites before anyone else.

You're in. Welcome. Expect an insider-only email soon.
Oops. Looks like Something went wrong. Try again?
 No spam  No inappropriateness  Unsubscribe anytime

By submitting this form you agree to our (more than fair) terms.

See how brands are finding new growth with intent
Closing the Intent Gap playbook: cover and inside pages about using intent data for ecommerce growth
By submitting this you agree to our (more than fair) terms
Check your email. The playbook is on its way.
Oops! Something went wrong while submitting the form.
This site uses essential cookies to run properly and optional cookies to improve your experience. Optional cookies only run if you accept them. Privacy Policy here.