How experimentation and Intent should work together (Part 2): Candidates

David dusts off his pen (fingers) and tells you how experimentation and Intent work together. Part two of three in a series.
Read Time:
On this page

This is the second of a three part series where we talk about how Experimentation and Intent should work together. In this series, we cover theory, practical examples and when Intent should and shouldn't be used.

We'll pop the links to the other pieces in the series below, so you can jump between each blog post. For now, I'll hand you over to David.

Part 1 covers the differences between experimentation and Intent.

Part 3 is all about a model we've devised on how to get the most out of Intent.

Experimentation is evidence to determine whether something works. It answers the question "does this work", identifying the best single answer for everybody, designed to prove or disprove a hypothesis. Think of this as 80% exploration of experiences, and 20% exploitation of experiences. Intent decides who receives what, when (or whether they need it at all).

It answers the question "for whom does it work best, when should it be shown, and do a proportion of users even need it at all?". This moves us towards a state where a single answer was never going to be enough for true commercial gain, scaling beyond a single hypothesis.

Think of this as 20% exploration, and 80% exploitation. You might call this personalisation. We would too. But the core difference is that it's personalisation that is a) genuinely scalable given it's agentic infrastructure and b) designed for the visitor, not the website.

Experimentation and agentic experience delivery with Intent sit hand in hand. Sure, there's some overlap, but passing one off to the other can be a thing of beauty.

Keep doing the research. Keep solving the big problems for everybody. Keep experimenting to add further evidence. Intent is for the layer underneath that. A layer that says the same thing is right for some people and wrong for others. Not everything works for everyone, right? It's the motion which no team has ever had the hours to work through by hand, giving you the ability to scale and amplify messages, components, divs and elements that are already proven to work.

To that end, we don't necessarily ask where Intent sits within the organisation. From experience, it sits across departments and is completely dependent on the culture, approach to testing, the centralisation of experimentation, who bought Intent in the first place and the reason for buying it.

It's therefore less about where it sits, and who owns it, and it's more about:

1. at a top level; purpose (see the differences between experimentation and Intent)

2. at a lower down level; individual selection. Asking "what are the candidates for Intent?" The relationship between when to experiment is vs when to use agentic experience delivery (Intent).

This narrative will set out for answering the latter.

What, who, when, whether

The differences between experimentation and agentic experience delivery are evident. Their purposes differ. We outlined that in our previous blog post "What are the differences between experimentation and Intent?".

A good way to look at it is the 4x W's principal: what, who, when, whether.

You know how they say personalisation is the right message, for the right person, at the right time? Well, that's Intent. Our unique ability is that you aren't just testing the what, but you're testing the who and the when. All based on the users Intent; the most powerful personalisation context point there is.

The trigger mechanism that sits behind your experiences is often assumed as static and generic; its the same thing for everyone at the same time.

It's rare that the variable of time is even measured, let alone tested.

Decision Owner In practice
What Message, mechanic, creative, layout, page structure. Experimentation Vendor A hypothesis discipline to determine whether something is correct for your user base in the first place
Who Which visitor gets it. Intent Allocation on predicted buying stage, purchase confidence, abandon risk and intent trend
When At what point in the session. Intent A timing decision made on live behaviour
Whether Should anybody see this right now. Intent Suppression. Not everyone needs everything in order to convert

Which decisions are candidates for Intent

We don't ask where Intent sits in your organisation because it's so dependant on the relationship with split testing.

Experimentation is now so ingrained that it has become a cultural trait rather than a situational judgement. Culture defines how you test and, therefore, how you ship experiences. Some organisations test by instinct, some by reflex, most sit somewhere in between. In other words, sometimes the output is Intent-based, sometimes it is a split test, sometimes it's some other form of evidence like UX research or customer support queries.

While we try to identify what questions you might ask at a global level, the environment and culture of your brand ("how you think about experience optimisation") often outweighs the modes that sit below.

That's why we ask why experiences are candidates for Intent instead of "where it sits" in the org.

In short. Ask just 2x questions to determine whether an experience is a candidate for Intent or not: → Will the effect vary by person and moment? → Do you have evidence that it works?

Not everything on a roadmap is a candidate for Intent.

Here are some examples where we wouldn't recommend using Intent:

• New features and significant journey changes

• Proving that a singular thing works

• Fundamental UX and accessibility fixes

• The core journey. If every visitor has to pass through it, tune it for every visitor.

On the contrary, where Intent candidates tend to cluster:

• Messaging of every kind. Service messaging, USP messaging, reassurance, urgency, delivery information. "who" and "when" to serve the message

• Trading and merchandising decisions. Incentives, thresholds, promotional emphasis. "who", "when" and "whether" to serve the incentive

• Social proof, urgency, persuasive tactics, reassurance, trust and recommendations, where appetite genuinely varies by person.

There are certain conditions that we run through to determine which route to take, and when using Intent is appropriate. To find those candidates, we generally ask two core questions, with five smaller questions underneath:

→ Will the effect vary by person and moment?

When research shows an intervention works for some people and against others, you have found something a single rolled-out answer cannot serve. That is the moment a decision becomes a candidate for Intent. It's the difference between having a "static, generic" experience vs a "dynamic, individual" one.

For example: you could test a USP bar at the top of your page, but you know that the USP bar is both static and generic. The message of "free shipping" will impact different people at different stages of their journey.

→ Do you have evidence that it works?

Not do you believe it: "Do you have something to point at?" Sometimes you genuinely have nothing. Sometimes you have years of it, yours and everybody else's, and the only open question is one of aim. Evidence here is broader than a split test: research, or a live rollout you have already measured, or a mechanic the whole category settled long ago all count.

There are some other variables to take into account that are worth asking, too. These are the 5x smaller questions that ultimately determine candidates for Intent.

Timing

Q1. Does the moment for making the decision sit within the page, not just at page load?

We want you to think about the timing of when an experience is shown, not just the content. Who is shown to, and when it is shown, can sometimes be more important than what is shown.

The moments of decision often sits inside the page itself, in between events. Momentum dips, hesitation occurs, purchase confidence might decrease sharply all within the same page and without continuous intent signals being read, you might have no idea.

If the decision only needs making once, at page load, you don't need continuous allocation to make it. But more often than not the moments that matter are not at page load. A good question to therefore ask on whether something is a candidate for Intent is also "does the moment for making the decision sit within the page, not just at page load?"

Value

Q2. Do you want the money now oppose to proving the value uplift later?"

‍Q3. Is resource preventative to your testing roadmap?‍

‍Q4. Does the experience have the potential to be detrimental to our margin and incremental?

Given that your candidate selection is often based on culture rather than situation, we'll be up front and simply ask one question: "do you want the money now oppose to proving the value uplift later?"

I know, that's a loaded question. Sorry. But it's because split tests answer a question; bandits capture value. Remember the weighting between exploration and exploitation? 80% exploration for experimentation (vs 20% exploitation), and 80% exploitation for Intent (vs 20% exploration)

The continuous allocation is there to squeeze all the juice of the lemon and serve experiences to where they are best seen and felt. Simply put; if you want incremental value from an experience now, Intent would likely be your best option.

The second question (Q3): from a situational standpoint where are you in your testing roadmap? Do you have more good ideas than test availability? i.e is resource a hinderance to going faster? This is often a result of client side testing where whomever owns that function is the one on the side opposed to within. Any programme capped at four to six experiments a month is usually rationing test, not prioritising, and a core symptom of being "resource-strapped".

Also, that third question (Q4) there are often times where serving the experience can cost you something. While there might be evidence to suggest something works, testing the impact of not doing it has not been proven.

Sure there is a control, but if you measure abandonment emails, for example, by conversions gained, it'll look positive. But you can't really measure them by sales lost. Or by the people who are now less likely to buy. A winning test that finds "discount [x] lifts conversion" hands you a mandate to discount, but it doesn't hand you a mandate to discount selectively or appreciate the negative connotations between discounting. Sometimes, that's worth considerably more than just a do vs don't do binary state.

The problem is, the incremental upside can be quantified but the downside cannot. Not unless you can understand visitor intent and continuously reallocate to the right users at the right time. Other examples inc.

1. Margin on a discount. Serving a discount (global, individual, segmented) works, but costs the business margin; how necessary is that?

2. Attention on an interruption. Serving an email popup might increase newsletter signups, but be detrimental to conversion.

3. Goodwill on an interruption nobody needed. Serving a "other people also bought"

Time to value

Q5. How quick do we need to extract value?

It's not that bandits are quicker than a straight split test (fixed traffic allocations are optimal for reaching statistical significance), but there are times where serving experiences using Intent can speed up the process of finding value.

Ultimately this comes down to how quick you want the outcome based on what you're trying to achieve.

Because statistical significance is a search for one single answer, the posterior is a belief about a variant. It exists to license a claim about a single population, and the claim is therefore singular by design; where B beats A.

Our posterior at Intent is a belief about a variant given a context. In other words, B beats A for a visitor in this intent state, at this point in their session. Our model predicts reward that is conditional on the visitor's context, so what comes out is a policy: hundreds of answers, running at once, continually being' reallocated to where the experience is best seen and best felt.

This might sound "faster", but actually it's not. An even 50/50 split is mathematically optimal for statistical power. Ignoring the fact that you're testing against a single average that can hide losers in winners, and winners in losers, the moment you push traffic toward the leader you have unequal groups. Here, the smaller one dominates the standard error, and you need more total traffic to reach the same confidence.

But, it's how the time to value differs because each are answering different questions. The traditional split test asks which individual variant wins, where agentic experience delivery asks a completely different question: "Does allocating these experiences by intent beat applying a blanket rule?". This means:

1. It changes what you can personalise. A per-variant posterior can only ever produce one answer for everyone, however good the statistics behind it are. Varying what people get requires probability, conditional on the person.

2. It changes when you can act. Certainty becomes a dial rather than a barrier. At 62% probability, our Thompson Sampling tilts the allocation and updates again tonight. Nobody declares anything.

3. It changes how fast you get there. Nothing is waiting for a threshold, so allocation improves from the first night onward. There's no finish line to reach before the work starts paying.

The short version

Identifying what is a candidate for intent is perhaps more important than where Intent sits.

We've established that we simply ask 2x questions (even expanding out to seven) that determine whether an experience or iteration should be a candidate for Intent. Simply ask:

"Will the effect vary by person and moment?" "Do you have evidence that it works?"

In our next chapter, we look at the different types of modes that exist when you're experimenting and using Intent to help further answer what the relationship is between exploration and exploitation.

Effect is much the same for everyone Effect varies by person and moment
You don't yet have evidence that it works Test it (Mode 2) Sequence it: Test, then let Intent scale (Mode 3)
You already have evidence that it works Just do it. (Mode 1) Hand it straight to Intent (Mode 4)

So, there you have it, the second in our experimentation and intent series. You can catch up or carry on reading here.

Part 1 covers the differences between experimentation and Intent.

Part 3 is all about a model we've devised on how to get the most out of Intent.

In the meantime, if you're still considering Intent and would like to know more, book a demo, and our team will walk you through our tool.

// the intent insider

Become an Intent Insider

Get subscriber-only insights we don't publish anywhere else and event invites before anyone else.

You're in. Welcome. Expect an insider-only email soon.
Oops. Looks like Something went wrong. Try again?
 No spam  No inappropriateness  Unsubscribe anytime

By submitting this form you agree to our (more than fair) terms.

See how brands are finding new growth with intent
Closing the Intent Gap playbook: cover and inside pages about using intent data for ecommerce growth
By submitting this you agree to our (more than fair) terms
Check your email. The playbook is on its way.
Oops! Something went wrong while submitting the form.
This site uses essential cookies to run properly and optional cookies to improve your experience. Optional cookies only run if you accept them. Privacy Policy here.