The Made With Intent blog

"The website is good, it probably could be even better, no doubts. But it's "good enough". It does the job. It's stable. It's mobile optimised. It's got a good conversion rate relative to others and our expectations. What else can we do? Is it the right thing to do to have a team purely focused on conversion rate optimisation, or should we look more into product and trading? What even is the potential of our site?"
That was a thought-provoking quote direct from a prospect of ours in a recent sales call.
He's talking honestly about the idea of prioritisation and diminishing returns. A very well-known brand, decent infrastructure, good content, a site that works very well; probably even over-indexes on conversion rate efficiency if you're to compare it to competitors. He'd looked at the size of the remaining prize from making the digital experience better and quietly concluded it might be smaller than the prize from doing something else entirely.
This is a question of "where do you place your bets?"
I think they have reached a local maximum. For anyone who hasn't sat through the optimisation lecture: it's the top of a hill that isn't the top of the mountain. Every small step available from where you're standing leads downwards, so you stop climbing. Not necessarily because you've reached the highest point there is, but because you've reached the highest point reachable in small steps. Which is a fairly precise description of what a decade of testing does to an already-decent website. In other words, getting to a higher peak means changing direction.
He's asking the right question in my opinion and I think the honest answer is uncomfortable for most of the industry I've spent my entire career in, especially the purists.
This brand hasn't necessarily hit the limit of what's possible on their website. But instead has hit the ceiling of the average and anything further sees diminishing returns where the effort doesn't necessarily equate to the value. That, or he's potentially bored with the same-same solutions that are out there. Homogeneity is the killer of excitement.

I empathise. I got bored too.
I founded User Conversion; one of UK's most successful (read: largest?) independent conversion rate optimisation agencies. We did well, working with some of the biggest brand names the UK had to offer.
But like the above brand, over time, I grew more and more skeptical. A lot of our recommendations lacked creativity, they were all addressing similar problems with the same solutions. "Moving deck chairs on the Titanic" is what someone once put to me.
Conversion rate optimisation is a process of problem-solving with evidence based solutions. Learning, uncovering opportunities and problems, and fixing those problems; usually through AB testing (well, that's the outcome that most cared about; because it's sexy). And I'm not suggesting that the learning and the opportunities dissipate, but the solution often lacks impact because of the law of diminishing returns, their heterogeneity and the aggregated nature of them.
The evolution resolution beyond CRO
Conversion Rate Optimisation, for stakeholders at least, is a way to make the website earn more; and there are four ways to do that. Most of us have treated them as a maturity ladder. You graduate from one to the next, and the last one is the good one.
That's not quite right. It's less an evolution and I now see them more as levels of resolution where each one narrows the unit of decision.
Level one: conversion rate optimisation. The unit of decision is the average visitor. You look at where people struggle, you fix it, and the fix applies to everybody. This is genuinely valuable and I'd never argue otherwise but it is a) practically often an exercise in usability improvements and b) definitionally serving the mean.
The first statement encompasses this idea that the majority of solutions are things that don't change behaviour, they facilitate existing behaviour. Usability improvements. Small changes that ill-advised vendors purporting marketing promoting statistics have convinced us are worth the effort. We've all seen them. The famed 500% uplifts. Sticky add to cart buttons, adding trust signals under a call to action, that sort of thing. Read: deck chairs on the Titanic.
The second reinforces the statement that the mean doesn't exist. Instead, it is a continuously moving combination of different intent levels. Our own research found that, say, 10% of visitors sitting on checkout pages aren't ready to buy yet. Or that 34% of users never get past browsing, whatever page they happen to land on. There is no average shopper to optimise for, there are different jobs to be done. There's a distribution we've been flattening for twenty years because flattening it was the only thing we could do. And now we're used to that, we lack creativity of how to proceed.
Level two: experimentation. It's the same unit of decision, the average, but now you've proven it. This matters enormously and it's the most rigorous thing most organisations do.
From experience, it often comes in two flavours:
1. The immature version lives inside the CRO team or the marketing function, running client-side tests and it caps out at about four to six experiments a month. That ceiling is a resource limit, often not a statistical significant limit; people, build time, roadmap slots. You'll find here that you'll max out at a certain number of tests usually and your impact is limited to avoid cross-contamination of tests. Ever found yourself saying "we can't run a test on a PDP because we have something running there already?"
2. The more mature version is server-side experimentation, owned by engineering, decentralised across product teams, with testing built into the release process rather than bolted onto it. That version is genuinely near-limitless in cadence, and if you can get there you should; it's fantastic.
But note what even the mature version is for. It's designed to prove or disprove a claim about the population. One answer, for everybody, with confidence attached. That's the instrument working exactly as intended; but it's still an answer about the average. It's also an answer about the website, not the visitor (more on that later).
My prospect's line was exactly this: "For the years we have experimented, we did not see huge benefits. That's probably also what hindered further investment." I've heard that sentence in some form from almost every brand I've worked with.
Level three: personalisation. Here the unit of decision finally narrows to the segment. And here is where the industry has spent a decade making promises it couldn't keep. Unfortunately, to the extent where we now all hold PTSD; personalisation traumatic stress disorder.
Personalisation didn't fail (that's right, I said it failed) because it was a bad idea, every boardroom still talks about it to this day. Trust me, I literally wrote the book on it: The Person in Personalisation.

It failed because it doesn't scale, for two reasons, both of which compound.
1. First, you have to guess the segments before you have any evidence about which distinctions matter. Those pre-defined segments like "returning visitor," "paid traffic," "landed on a PDP" are website attributes, not people attributes. That's not person-alisation that's website-alisation, isn't it? Pageview count stands in for engagement, when a confused shopper racks up far more pageviews than a decisive one i.e. it's not true person-alisation.
2. Second, the arithmetic defeats you. Split traffic three ways and every test takes three times as long to reach significance. Try to prove the segments genuinely differ and you're chasing an interaction effect that needs roughly four times the sample again. A three-week test becomes a quarter-long project, and most teams call it early and ship an artefact, or just don't have the traffic (and therefore patience) i.e. it's not scalable.
So personalisation became a small number of hand-built manual rules, maintained by someone who'd rather be doing something else, delivering less than it promised. Ever wondered why recommendations was the only successful personalisation that brands have achieved? Because it's autonomous; in other words, scalable.
Level four: agentic delivery, with intent as the context. This is where we, Made with Intent, sit. The unit of decision becomes the person and what they're trying to do in the moment. Not the segment on retrospective data. Also, not the average. And critically, nobody writes the rule.
You give the system a strategy and a set of experiences that are already evidenced. It works out which of them suits which state of intent, person by person, at the time that it matters, and it keeps working it out. Nobody writes the rule.
Take one of our customers, Diamonds Factory who had a single basket-abandonment tactic: 25% off, to everybody. They gave the agent four options instead and let it choose between them. Most people, it turned out, didn't need the full discount to convert. And 15% needed no intervention at all.
That last number is the one that matters, because no level below four can produce it. A test has no vocabulary for show this to nobody as a good outcome. Not even the control of an experiment can show you that because a user is never bucketed into both the control and the treatment. A rule-based segment can't discover it. It only appears when something is allowed to decide, per person, whether to act at all, in the moment that it matters.
The Future of Personalisation
I think we treated this as four evolutionary components within a single ladder when, in fact, they're two.
• Levels one and two are about how sure you are. "Does this work." Think of this as 80% exploration, and 20% exploitation.
• Levels three and four are about how precisely you aim. "For whom does it work best, when should it be shown, and do a proportion of users even need it at all?" Think of this as 20% exploration, and 80% exploitation.
Conflating them is why so many personalisation programmes were run by people optimising for certainty, and why so many experimentation programmes never escaped one answer for everybody. Confidence and aim are different problems. You need both, and the tools for each are not the same tool.
The industry has worked out that continuous contextual allocation beats a fixed split. That's 50% of personalisation; serving "the right person, at the right time, with the right message".

The other 50% are the attributes that determine whether something is personal; and for us that's their intent. Their context. Their job to be done. The differentiator is what you put in the context window. Not a website attribute like device or location, because they describe who someone appears to be. Intent describes what they're about to do. One of those is a proxy and one of them is the thing itself.
What I'd actually tell my prospect
Not "invest more in CRO." And not "stop doing CRO," either, because each of these levels still holds a purpose and pretending otherwise is how vendors lose credibility. No, CRO is not dead.
But its purpose has changed. There's more.

Sure, if something is broken, fix it. If a step in the core journey confuses everyone like a login, a checkout, a filter that doesn't work, then everyone passes through it, the fix helps or hurts them all in the same direction, and the right answer is one answer. Optimising the average isn't a failure of ambition there. An experimentation partner of ours put it well: you don't stop the research, you don't stop the UX work, and if you don't have people designing properly for their users you're finished as a business regardless of what any agent does on top.
But once the site is good enough; once you've fixed what's broken and the remaining UX gains are genuinely marginal, I think the question changes. It stops being how do we make this better for everyone and becomes which of the things we already have should this particular person see, and when, and should they see anything at all.
Essentially the argument of diminishing returns. Unless your site is broken, terrible UX, or hard to navigate; the best bet is dynamic, trading-related experiences which capitalise on serving the right content at the right time. Sure I'm biased, but this is my arc within this industry over the past 15 years. I've seen what's possible and I'd like to share it with the world. A TLDR;
There are different ways to optimise, but just note that I've seen first hand that scalability is the biggest constraint to success. Not just that, but doing the same as everyone else, particularly "moving deck chairs on the Titanic" won't get you very far unless the baseline is so low.
The opportunity is vast. It amplifies existing experiences by autonomously serving those only to where the experience is best seen and best felt, excluding where it's not. Because no one in the organisation owns it, or because we are so accustomed to the way things work currently; the opportunity is still sitting there.
My prospect's instinct was right. He should probably move effort away from optimising the average. Just not away from the website.
If you're interested in how CRO is evolving, and want to learn more about intent-based personalisation, get in contact with our team here.
Latest articles

"The website is good, it probably could be even better, no doubts. But it's "good enough". It does the job. It's stable. It's mobile optimised. It's got a good conversion rate relative to others and our expectations. What else can we do? Is it the right thing to do to have a team purely focused on conversion rate optimisation, or should we look more into product and trading? What even is the potential of our site?"
That was a thought-provoking quote direct from a prospect of ours in a recent sales call.
He's talking honestly about the idea of prioritisation and diminishing returns. A very well-known brand, decent infrastructure, good content, a site that works very well; probably even over-indexes on conversion rate efficiency if you're to compare it to competitors. He'd looked at the size of the remaining prize from making the digital experience better and quietly concluded it might be smaller than the prize from doing something else entirely.
This is a question of "where do you place your bets?"
I think they have reached a local maximum. For anyone who hasn't sat through the optimisation lecture: it's the top of a hill that isn't the top of the mountain. Every small step available from where you're standing leads downwards, so you stop climbing. Not necessarily because you've reached the highest point there is, but because you've reached the highest point reachable in small steps. Which is a fairly precise description of what a decade of testing does to an already-decent website. In other words, getting to a higher peak means changing direction.
He's asking the right question in my opinion and I think the honest answer is uncomfortable for most of the industry I've spent my entire career in, especially the purists.
This brand hasn't necessarily hit the limit of what's possible on their website. But instead has hit the ceiling of the average and anything further sees diminishing returns where the effort doesn't necessarily equate to the value. That, or he's potentially bored with the same-same solutions that are out there. Homogeneity is the killer of excitement.

I empathise. I got bored too.
I founded User Conversion; one of UK's most successful (read: largest?) independent conversion rate optimisation agencies. We did well, working with some of the biggest brand names the UK had to offer.
But like the above brand, over time, I grew more and more skeptical. A lot of our recommendations lacked creativity, they were all addressing similar problems with the same solutions. "Moving deck chairs on the Titanic" is what someone once put to me.
Conversion rate optimisation is a process of problem-solving with evidence based solutions. Learning, uncovering opportunities and problems, and fixing those problems; usually through AB testing (well, that's the outcome that most cared about; because it's sexy). And I'm not suggesting that the learning and the opportunities dissipate, but the solution often lacks impact because of the law of diminishing returns, their heterogeneity and the aggregated nature of them.
The evolution resolution beyond CRO
Conversion Rate Optimisation, for stakeholders at least, is a way to make the website earn more; and there are four ways to do that. Most of us have treated them as a maturity ladder. You graduate from one to the next, and the last one is the good one.
That's not quite right. It's less an evolution and I now see them more as levels of resolution where each one narrows the unit of decision.
Level one: conversion rate optimisation. The unit of decision is the average visitor. You look at where people struggle, you fix it, and the fix applies to everybody. This is genuinely valuable and I'd never argue otherwise but it is a) practically often an exercise in usability improvements and b) definitionally serving the mean.
The first statement encompasses this idea that the majority of solutions are things that don't change behaviour, they facilitate existing behaviour. Usability improvements. Small changes that ill-advised vendors purporting marketing promoting statistics have convinced us are worth the effort. We've all seen them. The famed 500% uplifts. Sticky add to cart buttons, adding trust signals under a call to action, that sort of thing. Read: deck chairs on the Titanic.
The second reinforces the statement that the mean doesn't exist. Instead, it is a continuously moving combination of different intent levels. Our own research found that, say, 10% of visitors sitting on checkout pages aren't ready to buy yet. Or that 34% of users never get past browsing, whatever page they happen to land on. There is no average shopper to optimise for, there are different jobs to be done. There's a distribution we've been flattening for twenty years because flattening it was the only thing we could do. And now we're used to that, we lack creativity of how to proceed.
Level two: experimentation. It's the same unit of decision, the average, but now you've proven it. This matters enormously and it's the most rigorous thing most organisations do.
From experience, it often comes in two flavours:
1. The immature version lives inside the CRO team or the marketing function, running client-side tests and it caps out at about four to six experiments a month. That ceiling is a resource limit, often not a statistical significant limit; people, build time, roadmap slots. You'll find here that you'll max out at a certain number of tests usually and your impact is limited to avoid cross-contamination of tests. Ever found yourself saying "we can't run a test on a PDP because we have something running there already?"
2. The more mature version is server-side experimentation, owned by engineering, decentralised across product teams, with testing built into the release process rather than bolted onto it. That version is genuinely near-limitless in cadence, and if you can get there you should; it's fantastic.
But note what even the mature version is for. It's designed to prove or disprove a claim about the population. One answer, for everybody, with confidence attached. That's the instrument working exactly as intended; but it's still an answer about the average. It's also an answer about the website, not the visitor (more on that later).
My prospect's line was exactly this: "For the years we have experimented, we did not see huge benefits. That's probably also what hindered further investment." I've heard that sentence in some form from almost every brand I've worked with.
Level three: personalisation. Here the unit of decision finally narrows to the segment. And here is where the industry has spent a decade making promises it couldn't keep. Unfortunately, to the extent where we now all hold PTSD; personalisation traumatic stress disorder.
Personalisation didn't fail (that's right, I said it failed) because it was a bad idea, every boardroom still talks about it to this day. Trust me, I literally wrote the book on it: The Person in Personalisation.

It failed because it doesn't scale, for two reasons, both of which compound.
1. First, you have to guess the segments before you have any evidence about which distinctions matter. Those pre-defined segments like "returning visitor," "paid traffic," "landed on a PDP" are website attributes, not people attributes. That's not person-alisation that's website-alisation, isn't it? Pageview count stands in for engagement, when a confused shopper racks up far more pageviews than a decisive one i.e. it's not true person-alisation.
2. Second, the arithmetic defeats you. Split traffic three ways and every test takes three times as long to reach significance. Try to prove the segments genuinely differ and you're chasing an interaction effect that needs roughly four times the sample again. A three-week test becomes a quarter-long project, and most teams call it early and ship an artefact, or just don't have the traffic (and therefore patience) i.e. it's not scalable.
So personalisation became a small number of hand-built manual rules, maintained by someone who'd rather be doing something else, delivering less than it promised. Ever wondered why recommendations was the only successful personalisation that brands have achieved? Because it's autonomous; in other words, scalable.
Level four: agentic delivery, with intent as the context. This is where we, Made with Intent, sit. The unit of decision becomes the person and what they're trying to do in the moment. Not the segment on retrospective data. Also, not the average. And critically, nobody writes the rule.
You give the system a strategy and a set of experiences that are already evidenced. It works out which of them suits which state of intent, person by person, at the time that it matters, and it keeps working it out. Nobody writes the rule.
Take one of our customers, Diamonds Factory who had a single basket-abandonment tactic: 25% off, to everybody. They gave the agent four options instead and let it choose between them. Most people, it turned out, didn't need the full discount to convert. And 15% needed no intervention at all.
That last number is the one that matters, because no level below four can produce it. A test has no vocabulary for show this to nobody as a good outcome. Not even the control of an experiment can show you that because a user is never bucketed into both the control and the treatment. A rule-based segment can't discover it. It only appears when something is allowed to decide, per person, whether to act at all, in the moment that it matters.
The Future of Personalisation
I think we treated this as four evolutionary components within a single ladder when, in fact, they're two.
• Levels one and two are about how sure you are. "Does this work." Think of this as 80% exploration, and 20% exploitation.
• Levels three and four are about how precisely you aim. "For whom does it work best, when should it be shown, and do a proportion of users even need it at all?" Think of this as 20% exploration, and 80% exploitation.
Conflating them is why so many personalisation programmes were run by people optimising for certainty, and why so many experimentation programmes never escaped one answer for everybody. Confidence and aim are different problems. You need both, and the tools for each are not the same tool.
The industry has worked out that continuous contextual allocation beats a fixed split. That's 50% of personalisation; serving "the right person, at the right time, with the right message".

The other 50% are the attributes that determine whether something is personal; and for us that's their intent. Their context. Their job to be done. The differentiator is what you put in the context window. Not a website attribute like device or location, because they describe who someone appears to be. Intent describes what they're about to do. One of those is a proxy and one of them is the thing itself.
What I'd actually tell my prospect
Not "invest more in CRO." And not "stop doing CRO," either, because each of these levels still holds a purpose and pretending otherwise is how vendors lose credibility. No, CRO is not dead.
But its purpose has changed. There's more.

Sure, if something is broken, fix it. If a step in the core journey confuses everyone like a login, a checkout, a filter that doesn't work, then everyone passes through it, the fix helps or hurts them all in the same direction, and the right answer is one answer. Optimising the average isn't a failure of ambition there. An experimentation partner of ours put it well: you don't stop the research, you don't stop the UX work, and if you don't have people designing properly for their users you're finished as a business regardless of what any agent does on top.
But once the site is good enough; once you've fixed what's broken and the remaining UX gains are genuinely marginal, I think the question changes. It stops being how do we make this better for everyone and becomes which of the things we already have should this particular person see, and when, and should they see anything at all.
Essentially the argument of diminishing returns. Unless your site is broken, terrible UX, or hard to navigate; the best bet is dynamic, trading-related experiences which capitalise on serving the right content at the right time. Sure I'm biased, but this is my arc within this industry over the past 15 years. I've seen what's possible and I'd like to share it with the world. A TLDR;
There are different ways to optimise, but just note that I've seen first hand that scalability is the biggest constraint to success. Not just that, but doing the same as everyone else, particularly "moving deck chairs on the Titanic" won't get you very far unless the baseline is so low.
The opportunity is vast. It amplifies existing experiences by autonomously serving those only to where the experience is best seen and best felt, excluding where it's not. Because no one in the organisation owns it, or because we are so accustomed to the way things work currently; the opportunity is still sitting there.
My prospect's instinct was right. He should probably move effort away from optimising the average. Just not away from the website.
If you're interested in how CRO is evolving, and want to learn more about intent-based personalisation, get in contact with our team here.

Experimentation is evidence to determine whether something works. Intent decides who receives what, when (or whether they need it at all). “Does this work.” Think of this as 80% exploration, and 20% exploitation.
The first works towards the best single answer for everybody to prove a hypothesis. The second works out where a single answer was never going to be enough; scaling a proven hypothesis for true commercial gain. “For whom does it work best, when should it be shown, and do a proportion of users even need it at all?” Think of this as 20% exploration, and 80% exploitation. Where testing tells you whether an experience works, but Intent decides who receives it, when, and whether they need it at all.
In this blog post, we'll identify the differences between Intent and experimentation, so you can understand when to utilise each strategy to further your business goals.
Defining A/B testing
An A/B test answers one single bounded question.
Does this change (a treatment) move the metric we care about, across our traffic, better than the alternative does (a control)?
It's a hypothesis, and what comes out the other end is 80% knowledge (learnings) and 20% value (hopefully positive gain). A primary "does this prove or disprove our hypothesis" binary answer.
Our intent engine answers a different question, and it answers it again every few seconds throughout a users session:
Given what this visitor is doing right now, which of the experiences we already know works suits them at this moment in time? (If any)
What comes out of that is not just a single decision, but hundreds of them a second, all trying to move the agents primary goal, or reward. The open question stopped being does this work a long time ago. What's left is who needs it, and when.
For example: if a visitor starts showing signs of basket abandonment. You can put a returns message in front of them, or a discount, or a trust signal, because you already trust all three of those mechanics. What you don't know is which one that particular person needs, at what point, or whether they need any of them at all. That's the decision the agent is making.

There are several reasons why experimentation (read: split testing) might not be an appropriate mechanism; here are four reasons:
But first, a note on the word experimentation. Throughout this set, experimentation means the whole practice of using evidence to drive ideas: user research, usability testing, prototyping, quasi-experiments and switchbacks where a split isn't possible, and A/B tests where it is. The split test is one instrument inside it, usually the last step rather than the whole of it.
1. Averages can lie.
An A/B test reports an average treatment effect which can hide segments or audiences that have a detrimental or negative impact.
An A/B test reports an average treatment effect. That's the whole point of it. Randomising across your population is what buys you a causal claim in the first place. The price you pay is that the average swallows everything inside it. A headline +2% can be +12% for hesitant first-timers and -4% for loyal returners. Same test. Same green tick. Two completely different stories.
Statisticians call these heterogeneous treatment effects, and there's good evidence that acting on them beats a blanket rollout. In one worked simulation of a ranking change, treating only the 69% of users predicted to benefit delivered a 2.2% gain in revenue per user against shipping it to everybody. A CUNY study of financial aid nudges captured roughly 75% of the total benefit while treating half the population.

Experimentation is deliberately built to stop you trusting the segment breakdown. Pre-registration, power calculations, no post-hoc slicing. Those rules exist because sliced results are where false positives breed. Kohavi, Deng and Vermeer's A/B Testing Intuition Busters (KDD 2022) found that underpowered analyses inflate the effects you observe by 25 to 50%. At the power levels most mid-market programmes actually run at, over half of your significant results are false positives. So the question you have to answer to deploy well, for whom, is the question your test is worst at answering. That's a boundary, and a sensible one. Intent sits on the other side of it. It predicts for every visitor in advance, rather than carving up a sample after the fact.
Where Intent solves this problem: The experience is being delivered across hundreds of contexts, reallocating to where the experience "works" (where it's best seen and best felt), excluding it from where "it doesn't work." In theory, it only looks and focuses on the good, excluding the bad.
2. Testing assumes the trigger.
Every A/B test starts from a fixed, assumed rule (usually on page load). The trigger "when to show it" is often assumed, arbitrary, and based on a website proxy like "3 page views" or "PDP" rather than genuine user intent.
Nearly every A/B test starts from a fixed rule. Show this on the PDP after three pageviews. Fire this on exit intent. Serve this to returning visitors. Then it tests which creative performs best inside that rule.
The treatment gets examined. The trigger gets assumed.
There's a practical reason for that. The trigger space is far too big to write out by hand. Nobody can enumerate every combination of buying stage, purchase confidence, abandon risk, intent trend and product affinity, let alone power a test across them. So teams pick a proxy or an average and move on.
Proxies are usually where the damage happens. These are largely based on website attributes as opposed to personalised signals. Returning visitor stands in for higher intent. Mobile stands in for lower intent. Pageview count stands in for engagement, when a confused shopper racks up far more pageviews than a decisive one, and high intent visitors often convert inside three.
Our own Intent Gap research found 10% of visitors on checkout pages aren't ready to buy. And 34% never get past browsing, whatever page they happen to be on. There's a reason we say stages, not pages.
Where Intent solves this problem: Intent agentic delivery serves the individual at a moment in time. A user might need an experience on the 3rd page view after 30 seconds, a different user might need it after 50 seconds, a different user might need it after 9 page views and 1 second. The continuous prediction modelling gives a threshold to meet first, and then a mechanism to automate a decision second.
3. Statistical significance is a search for one answer.
The search for statistical significance can a) slow you down and b) inhibit your ability to personalise to different segments.
Significance exists to license a claim about a single population, and the claim is singular by design where B beats A: one answer, for everyone. That is exactly what costs time.
You need a sample big enough to speak for the whole population with a minimum detectable effect. If you want the claim to be about a segment instead (what some call "personalisation") you need that sample inside the segment too, so the wait multiplies. Most claim they don't have the traffic to personalise because of this.
Experimentation vendor's posterior is a belief about a variant. One distribution per arm, computed across everybody who saw it, resolving to a single winner for all your traffic. Better maths than a p-value, and the shape of the answer is identical. One number per variant, and a threshold you wait to cross before you act.
Our posterior at Made with Intent is a belief about a variant given a context. Where their question is whether B beats A, ours is whether B beats A for a visitor in this intent state, at this point in their session. The model predicts reward conditional on context, so what comes out is a policy; hundreds of answers, running at once.

Put it this way: instead of a human designing three segments and running three underpowered tests, the winner goes to the agent alongside its alternatives and the segment level optimisation runs continuously. Three mechanisms make that more traffic-efficient than a segmented test. Worth understanding rather than taking on trust.
1. Allocation is adaptive. Thompson Sampling draws plausible performance rates from each variant's posterior and routes traffic accordingly, so your exploration cost falls as confidence climbs rather than sitting at 50% for the duration. In practice agentic campaigns push 80 to 90% of traffic to the winner inside two to three weeks. A fixed-allocation A/B test takes around six. Optimised Control shrinks the control group automatically as confidence grows.
2. Without waiting for statistical significance. It borrows statistical strength across slices. This is the important one. Made With Intent runs a contextual bandit, not a multi-armed one. A multi-armed bandit only learns from the data each arm receives, so every segment needs its own volume, which is the same problem as a segmented A/B test in different clothing. A contextual bandit trains a model that predicts reward given context, so it can estimate performance for a combination it has never directly seen. Show it mobile traffic, low intent traffic and a particular page separately and it can predict for "mobile, low intent, that page" by generalising from how each component behaves. A hand-built segment can't borrow like that. Every cell starts at zero.
3. You're not assuming which segment to slice or review. The agent can build up to 500 contextual combinations and up to 75 moment-based triggers per campaign, and it trims its own feature selection to keep each slice sample-rich. Nobody has to guess which distinctions matter before there's evidence about which ones do.
Why we don't report per-segment significance
This surprises experimentation-native teams more than anything else in the product. It's a deliberate methodological choice.
Run an independent significance test on every intent combination and multiple comparisons swallow the results. At α = 0.05, roughly 14 independent tests give you about a 50% chance of at least one false positive. A hundred tests will throw up around five false positives by chance alone, and across hundreds of combinations you'd be manufacturing findings.
We trialled per-segment significance reporting and pulled it, because it surfaced spurious and inverse correlations. For example, users traverse through different stages of intent in their journey, at what point is a user in low intent and when did they see the treatment?
The per-segment decisions get justified differently. Is the model calibrated, meaning when it says 70% does it convert around 70% of the time (measured by expected calibration error)? Is it discriminating, meaning can it rank likely converters above unlikely ones (AUC around 0.84)? A well calibrated probability is a legitimate basis for acting on a slice. An underpowered significance test on that same slice is not.

This calls back to something we talked about earlier.The reason you shouldn't chase per-segment significance in your testing tool is the same reason we don't chase it in ours. Prediction is the right instrument for the for whom question, retrospective slicing isn't.
Why we don't report on single variant uplifts
You get one clean causal number, control against the agent's allocation, measured the way any experimentation lead would want it measured. You give up per-arm inference, and in exchange the arm is adaptive. That's the trade. An A/B test gives you clean inference on every arm and one winner for everybody. An agentic campaign gives you clean inference on one arm, and a different winner in every context.
Load five experiences into an agentic campaign and it looks like a five-arm test. It isn't. It's actually a two-arm experiment, the same shape as any A/B test you've run. The difference being:
• Control. A baseline you define.
• Variant. The agent's allocation across all five experiences.
A traditional A/B test asks which variant wins with a blanket rule. An agentic campaign asks a different question with an adaptive rule (it's a reason why we call them campaigns, not experiments)
Does allocating these experiences by intent beat applying a blanket rule?
We don't report on single variant uplifts eg: Variant B is better by 5% because:
1. There's no single number to give you. Experience B might be the winner for high intent, budget-conscious mobile visitors and the loser for low intent desktop. That's the entire point of running it this way. Collapsing it to one figure puts back exactly the average this campaign exists to get rid of.
2. The arms were never randomised against control. Traffic reaching experience B was chosen by the agent, on context and accumulated evidence. It isn't a random slice of your audience so there's no effective control. Comparing a deliberately selected group against everybody is confounded by construction, and nothing in the data separates the effect of the experience from the effect of the selection.
3. You'd be reading a state the system has already left. If experience B was struggling in a context, the agent would have noticed days ago and moved traffic away (hence: reallocation). The number in front of you is an average across a period in which allocation was actively changing, and the campaign has moved on since. Acting on it means acting on history the agent has already corrected for.
That distinction does three things, in order.
1. It changes what you can personalise. A per-variant posterior can only ever produce one answer for everyone, however good the statistics behind it are. Varying what people get requires probability conditional on the person. That's the whole game, and it matters far more than the frequentist argument ever did.
2. It changes when you can act. Certainty becomes a dial rather than a gate. At 62% probability, Thompson Sampling tilts the allocation and updates again tonight. Nobody declares anything. A stopping rule stops being necessary once allocation is continuous.
3. It changes how fast you get there. Nothing is waiting for a threshold, so allocation improves from the first night onward. There's no finish line to reach before the work starts paying.
Where Intent solves this problem: Our posterior at Made with Intent is a belief about a variant given a context. Where experimentation question is whether B beats A, ours is whether B beats A for a visitor in this intent state, at this point in their session. The model predicts reward conditional on context, so what comes out is a policy; hundreds of answers, running at once based on a series of probabilities using a contextual bandit.
4. Limited to a single answer for only one point in time.
An A/B test only tells you what was true of your traffic during the weeks the test ran.
An A/B test tells you what was true of your traffic during the weeks the test ran; a shelf life. Nothing about that answer refreshes itself so brands often end up "re-testing" the same thing a few years later, hoping there's no change from a previous positive test. The world moves on, results degrade as users become "used" to the treatment, competitors copy; not to mention that your purchase lifecycle varies by days, weeks, months, years.

For example: a Microsoft experiment on MSN.com saw replacing one button with another produce a 4.7% increase in overall clicks. The daily breakdown post launch showed the difference decreasing rapidly day over day as users learned the change, leading to the team shut the experiment down mid-way. The conclusion was that the observed treatment effects "are not always permanently stable, sometimes revealing increasing or decreasing patterns over time."
There are a few reasons why this happens.
1. Novelty and primacy wear off. A new element gets attention because it's new. Returning visitors are briefly worse off because it isn't what they knew (which is why tests are often split, somewhat arbitrarily, between new and returning users)
2. Buying patterns differ throughout the year. December traffic behaves nothing like March traffic, different intent, different price sensitivity, different tolerance for being interrupted (especially for retailers)
3. The population itself drifts. Mix of paid traffic allocation shifts, category mixes alter or a competitor changes their delivery proposition. Macro-economic factors alter the interaction effect.
Does a mature site slowly accumulate a layer of decisions that were correct once, are serving everybody by default, yet are answerable to nobody? It is entirely possible to have a well-governed testing function and a site full of expired answers at the same time.
Where Intent solves this problem: Continuous allocation doesn't have this failure mode, because it never declares anything. The posterior is a live belief rather than a verdict, continually reallocating and therefore responding to changes, not static. Allocation is re-scored against current behaviour and retrained nightly, and standing exploration keeps a small share of traffic asking whether the current answer is still the right one. When February stops behaving like November, the agent finds out because it never stopped looking.
Intent vs experimentation: In depth
We've done a quick reference table on the differences between Intent and experimentation. Take a look, send it to your colleagues. You're welcome.
Experimentation and Intent aren't solving the same problem, and treating them as substitutes is where most testing programmes stall. An A/B test proves whether something works for everybody, on average. But it can't tell you who needs a returns message versus a discount versus nothing at all, and it becomes redundant the moment the traffic that validated it moves on.
Intent picks up where that boundary sits, deciding who an experience is for, when, and whether they need it at all, re-scored every few seconds instead of declared once and left to expire. Most eCommerce teams already have the first instrument. The second is what turns that evidence into revenue, visitor by visitor, instead of one rollout for everybody.
Book a demo to see how Made With Intent can help you deliver more appropriate experiences to your customers.

Made With Intent is an on-site intent engine for eCommerce businesses. It reads buying intent in real time and decides which experiences go to which visitors, and when.
Dynamic Yield is an experience optimization and personalization platform. It helps businesses tailor digital customer journeys across websites, mobile apps, and email.
This post explains where the two platforms work together, and where you'd still use Dynamic Yield instead.
If you can't wait until the end, here's a TLDR;
Pick Dynamic Yield if:
• Your priority is consistent personalisation across web, mobile app, email, and in-store kiosk
• You've got the team and budget to do a weeks long enterprise deployment
• Recommendation capabilities matters more than first-page view intent coverage.
Made With Intent is for you if:
• You want to read buying intent in real-time
◦ Then, serve the right content, the right experience at the appropriate moment
• Want to benefit from a 30-60 minute install time
And when we're thinking about Made With Intent and Dynamic Yield working together, here's how we recommend thinking about it:
Made With Intent decides what experiences you should serve, to who and when, and Dynamic Yield delivers those experiences.

How Made With Intent is different to Dynamic Yield
There's three things that Made With Intent does different to Dynamic Yield. Let's dig into them below:
1. Made With Intent reads intent in real time
Dynamic Yield's intent workflow utilises two routes. First, Empathic Personalization classifies visitors into four inferred states, which are Curious, Interested, Focused and Satisfied.
Its Audience Hub lets teams hand-build "low / medium / high intent" audiences from rules. Dynamic Yield's Primary Audiences framework shows a leading golf retailer defining low intent as fewer than 12 page views per session and high intent as more than 24.
That's a useful starting estimate, but page views aren't intent. A hesitant shopper racks up more pages than a decisive one. A high-intent visitor often converts inside three.
The more common version isn't pageview counts — it's event proxies. Add to cart equals high intent. Wishlist equals consideration. But an add to cart is as often a price check, a size comparison or a shipping-cost probe as it is a purchase signal. The event tells you what happened. It doesn't tell you what it meant.
Now, this is pretty fundamental: Dynamic Yield's model relies on behaviours your shoppers have already exhibited. It's a snapshot backwards in time, and like many what we call "rules-based" personalisation tools, you're reacting to things that have already have happened. Unlike Made With Intent.
Made With Intent reads buying intent directly — multi-dimensional, second-by-second signals predicting where the visitor is in their decision right now.
We break downs shopper behaviour into six stages, and these are: Intent stage. Intent signals. Intent trends. Purchase confidence. Abandon risk. Shopper mindset.
All updated every three–five seconds during a live session, on a model trained across 150+ retailers and 50 billion+ events.
2. Allocation across hundreds of intent combinations and not post-test segmentation
It takes eCommerce teams considerable time to build and maintain audience rules in Dynamic Yield. For instance, coding things such as "low intent equals fewer than 12 page views, high intent equals more than 24," then QA-ing those cohorts and redesigning them after each test.
Made With Intent replaces that with continuous intent prediction and delivery. The model decides who sees what, in real time, across hundreds of intent combinations. The actual grunt work is handled by our agent. Your team focuses on strategy, creative, and proof.
Our agent retrains daily, so the experience keeps allocating toward the intent combinations where impact is actually felt. Always on, always learning, always improving on where it started.

3. First-pageview coverage — no fallback needed
Behavioural data takes time to accrue; for anonymous or first-time visitors, Dynamic Yield falls back to geo-based predictive targeting and contextual signals.
Made With Intent has no fallback by design. It doesn't need one. We've trained it across 50bn+ events, in a range of contexts, giving the model day-one predictive power on every visitor, anonymous, identified, first-time, returning.
Made With Intent and Dynamic Yield: In depth
There are lots of areas of cross-over between Made With Intent and Dynamic Yield. The way our technologies work is similar in principle, but different in its practical implementation.
Let's start first with how Dynamic Yield's prediction capabilities work:
How Dynamic Yield's prediction works
While we're talking about AdaptML, we have to say, it's a really sophisticated bit of engineering. It utilises recurrent neural networks and NLP models to get smarter and smarter as it consume more data its got on a specific retailer's visitors.
However, it predicts what you'd expect (affinity and relevance), not what's happening in a visitor's decision right now.
Our model is a driven by a single purpose: it reads buying intent. Sure, it's a narrower job, but it's the one that determines whether a visitor converts, hesitates, or leaves.
Made With Intent vs Dynamic Yield

When you should pick Dynamic Yield
Hey, we're not here to blindly sell you Made With Intent. Sometimes, our tool just isn't right for your business. So, unlike loads of other SaaS vendors, let us tell you when we wouldn't be a good fit:
• If you're trying to do cross-channel, Dynamic Yield is the right choice versus Made With Intent. We're web-first, and for multi-channel personalisation programmes Dynamic Yield is the right option.
• Recommendations: NextML and AffinityML are purpose built models for product and content recommendations. Made With Intent adds an intent decison later but it is not a replacement for NextML or AffinityML.
• Sophisticated email personalisation: Klaviyo-grade dynamic content across email and ad placements out of the box. If email is a key part of what you do, Dynamic Yield is a top choice.
• Security of a big vendor: Dynamic Yield is an eight-time Gartner Magic Quadrant Leader for personalisation engines. If you're enterprise org and need to run a formal RFP, this helps your procurement team make a decision.
• Strong experience builder. Sections, Page Contexts, and Selectors give users very fine control. But only after you've climbed the initial learning curve.
Where Made With Intent wins
Of course, this is an article designed to help you pick between Dynamic Yield and Made With Intent. From our table above, it's clear there's plenty of areas of overlap, but there's some stuff we do that, we don't mind saying, makes us a better choice. Have a read:
• Acts on the anonymous majority: Every visitor gets a multi-dimensional intent read from the first pageview. No prior data, no behavioural accrual, no geo fallback.
• Multi-dimensional intent prediction: The way we predict intent is comprehensive. We use six measures: stage, signals, trends, purchase confidence, abandon risk and shopper mindset. These are all updated every three-five seconds, unlike Dynamic Yield.
• Autonomously runs campaigns and makes decisions under directives: Agentic Campaigns can dynamically allocate experiences across hundreds of intent combinations during a campaign you'll run. Dynamic Yields's models are sophisticated. The targeting layer above them is still rule-driven. That's simply how the tool is built. But it means what experiences are served is decided before the campaign runs, and not during it.
• Cross-merchant model: 50 billion+ events across 150+ retailers. Day-one predictive power for our customers.
• Start within 30 minutes: Single 7kb GTM tag, no PII and we're ISO 27001 accredited.

So, it's time to make a decision: Let's pick one. Or both?
It's crunch time. We've given you all the facts, but it's time to wrap up and make a judgment call.
Pick Dynamic Yield if the you want consistent cross-channel personalisation — content and recommendations spanning web, mobile app, email, and kiosk. And if you have the technical resources to operate it.
Dynamic Yield's breadth and recommendation algorithm depth is class-leading, and trying to replicate the scale of that capability with Made With intent, plus some integrations would be a worse outcome for you.
Choose Made With Intent if you want a on-site decisioning, intent-driven experiences, and proven incrementality. Our tool delivers experiences natively, allocates them dynamically across hundreds of intent combinations, and proves causal lift on every one.
Now, something to think about. If you're an enterprise customer, you actually could consider both. And here's why:
• Dynamic Yield gives you the breadth and means of channels and mediums to serve personalised content to
• Made With Intent helps you accurately decide who, and when that content should be served
That's brought us to the end of the comparison blog post. If you're still unsure of the differences between Dynamic Yield and Made With Intent, the best thing for you to do is talk with one of our team.
You can book a demo here.
Disclaimer: This comparison is based on publicly available information from Dynamic Yield's documentation, marketing site, and customer reviews as of April 2026. Both products evolve continuously. If anything looks out of date, get in touch and we'll sort it.

Your A/B testing programme just found a winner. A new discount banner lifted conversion by a healthy margin, the test ran long enough to be statistically significant, and the result is sitting in a dashboard somewhere with a green checkmark next to it. Mission accomplished, right?
For most eCommerce teams, the default answer is the same: roll it out to 100% of visitors.
That decision gets made without a second thought, because it doesn't feel like a decision at all. It feels like the natural conclusion of a successful test. You've gone and validated your hypothesis. And when you've got a winner, there's a paper trail to an uplift.
It isn't. It's a choice, and for most brands it's the wrong one.
A/B testing tells you which experience wins across your traffic as a whole. It doesn't tell you which specific visitor should see that winner, or when. That's a different question, that's answered in an entirely different way. It's a decision that requires you needing to know something about the person you're serving an experience to in the first place.
Let's be clear — A/B testing definitely has a place
A/B testing tools do exactly what they're built to do, and they do it well. They replace opinion with evidence. A test that runs to statistical significance tells you, with confidence, that variant B outperforms variant A on your website or mobile app.
Made With Intent customers continue to use tools like Convert, VWO, and AB Tasty alongside our tool, because testing answers a question nothing else can: does this specific change move the number we care about?
Think about what a test actually proves. It proves that, across hundreds or thousands of visitors, variant B beat variant A. It says nothing about any single visitor in front of you right now. Statistical significance is for large volumes of people not a prediction about the individual standing at your checkout, which is a limitation CRO metrics share more broadly.
The difference between A/B testing tools and Made With Intent
Should 100% of people get your A/B test winner?
Once a test determines a winner, most teams roll that experience out broadly, often to every visitor who fits the original test's targeting. Many testing tools do offer segment targeting: device type, geography, new versus returning.
Some brands go further and layer in basic rules, like showing a winning banner only after 30 seconds on site, or only to visitors on their third page view. These rules feel like personalisation. They're really just a slightly more granular version of the same default: a fixed condition, checked once, applied broadly.
But look closely at how it gets set up. Teams choose segments once, at test design time, and then leave them alone. Even a well-built intent-based segmentation programme needs continuous application, not just at launch, or it drifts back into the same static-rule problem.
A rule that says "show this to returning visitors" doesn't change three minutes later when a returning visitor starts behaving like someone about to abandon their basket, or like someone who was going to buy anyway with or without an incentive.
The real problem is narrower and more difficult to solve. The targeting decision is frozen at the moment the test launches, while the person in front of it keeps changing their mind. (Their intent)
A discount that lifts conversion in aggregate can still be the wrong call for two specific people in your test group. A first-time visitor who's price-sensitive and hesitating needs it. A loyal customer who's already three products deep into checkout, on their fourth order this year, does not.
It's the same tension at the heart of whether discounting is a race to the bottom or a genuine growth lever. There's no problem with offering discounts, but it can be a blunt instrument or entirely inappropriate depending on where your customer is at in their buying journey.

How reading intent allows you to respond immediately to changing demands
A/B testing tools make their decision once, when the page loads or the session starts. Whatever segment a visitor falls into at that moment determines what they see for the rest of that test.
But purchase intent isn't fixed at page load. A visitor can arrive browsing casually, spend four minutes comparing two products, add one to their basket, then hesitate at the price for 30 seconds before starting to type a discount code into Google in another tab. That's not a static segment. That's a person, moving in real time, and most testing infrastructure was never built to react to this situation as it unfolds.
Some testing and CRO tools react to single behavioural triggers: an exit-intent popup fired on mouse movement toward the browser bar, a scroll-based banner, a basic engagement score. We think those are great, and definitely worth using. However, they're single-trigger rules, checked once a condition is met, and not a continuous read on a person as they use your website and apps.
What real-time behavioural data adds
Made With Intent's platform tracks over 900+ behavioural signals through each visitor's session, including scroll speed, click sequences, time spent, and product comparison patterns, updating its prediction of purchase intent continuously rather than once. When intent shifts, the response can shift with it.
None of this means every session tells a dramatic story. Plenty of visitors browse, leave, and never come close to converting in that session, and no amount of real-time signal turns a casual browser into a buyer. The claim isn't that every visitor is secretly ready to buy. It's that the visitors who are ready look different from the ones who aren't, and a rule set once at test launch can't tell the difference between them three minutes into the session.

What real-time intent-based deployment looks like
Picture the same winning discount from earlier. Instead of firing it to every visitor who matches a fixed segment, an intent-based system asks a narrower question: Is this specific visitor showing signals of hesitation or exit risk right now, or are they already moving smoothly toward checkout?
Made With Intent's Intent Scoring analyses those signals continuously and its Timing Engine decides the moment to act, rather than firing on a page-load rule or a fixed timer. A hesitant first-time visitor comparing prices across tabs might see the discount. A returning customer already at the payment page, showing no hesitation signals at all, doesn't need it, and doesn't get it. This is what discounting with intent looks like in practice: the same offer, reserved for the visitors who actually need it to convert.
The winning experience from your test doesn't change. Who receives it, and when, does.
Take a product page test that proved a scarcity message ("only 3 left") increased add-to-basket rate. Fired at everyone, it also lands on a loyal customer who's bought the same product line four times before and knows exactly what they want. To that visitor, it reads as pressure, not information. Deployed by intent, the same message is reserved for the visitors actually showing comparison and hesitation behaviour, the ones it was designed to nudge in the first place.
How an actual company used Made With Intent alongside A/B testing
Appliances Direct, a UK home appliances retailer, ran into a version of this problem. Their discounting was broad and largely untargeted; offers went out early in the visitor journey, without real-time context, regardless of whether a given shopper actually needed the incentive to convert.
Using Made With Intent, they began reserving discounts for visitors showing genuine exit and abandonment signals, rather than firing them to everyone who matched a static segment. They left everyone else to convert at full price, as they were always going to.
In the end, they made a 42% saving in margin previously given away to visitors who would have converted regardless, without a drop in conversions from the visitors who needed the discount to purchase.
But don't take our word for it. Take a look at our Appliances Direct case study and make up your mind yourself.
That figure lines up with a broader pattern. Made With Intent's own Intent Gap research found that 83% of shoppers have used a discount code despite being ready to pay full price.

Why new test winners are rare
That scarcity of new winners isn't unique to Appliances Direct. Research from Microsoft's own experimentation team, drawn from thousands of controlled experiments at Bing and elsewhere, found that only 10 to 20% of test ideas move the metric they were designed to move, and that a genuine breakthrough, the kind of result that reshapes a programme rather than nudging it, shows up in roughly one in 500 experiments. The findings come from Kohavi et al.'s Seven Rules of Thumb for Web Site Experimenters (2014). Most of what a mature testing programme finds after the early wins are used up is small, single-digit-percentage gains, not new step changes.
Testing tells you what works. Intent tells you for whom, right now.
None of this is an argument against testing. It's an argument for finishing the job testing starts.
Your A/B testing programme answers one question well: does this experience outperform the alternative? What it was never built to answer is a second, quieter question that gets decided by default instead of by design: who should see this winning experience, and at what moment in their session?
Testing finds your best experience. Making sure it reaches the right person, at the right moment, without giving away margin to people who didn't need it, is a different problem, and it needs a different kind of data to solve.
If you want to see what your existing winning experiences would look like deployed by real-time intent instead of a blanket rule, book a demo and we'll walk through it against your own traffic.

Made With Intent is an on-site decision engine for eCommerce businesses. It reads people's buying intent in real time and allocates experiences to the visitors most likely to respond to them
Optimizely is a digital experience platform. It does web and feature experimentation, rule-based personalisation, content recommendations. It's designed for teams running structured test-and-learn programmes at scale.
This blog post explains where the two platforms complement each other and when you would use both versus just one on its own. But if you're too impatient to read to the end, we've got you:
Made With Intent and Optimizely serve different functions. Optimizely handles experimentation and rule-based personalisation.
Made With Intent reads visitor intent in real time and decides who sees an experience and when. The two work really well together:
Optimizely executes what, Made with Intent decides who and when. We serve your Optimizely experiences when it matters (the right time within their session) and to whom, all based on what really matters; their intent.
How Made With Intent is different to Optimizely

1. Respond to users in real-time with experiences tailored to what their digital body language tells you
Optimizely's strongest targeting comes from rule-based audiences in Web Experimentation and Personalization, boolean logic on geo, device, behavioural events, URL targeting, and page Tags, plus optional machine learning (ML) layers:
- Adaptive Audiences (interest categories inferred from content engagement)
- Content Recommendations (NLP-driven topic affinity per visitor)
- Optimizely Data Platform (ODP) real-time audiences
The Stats Engine inside experiments is genuinely best-in-class — sequential testing with always-valid p-values — but the targeting decision still asks the marketer to define which audience rule a visitor fits into.
Made With Intent flips the model. It reads buying intent in real-time from the first pageview. Stage, signals, trends, purchase confidence, abandon risk, shopper mindset, and re-scores every three-five seconds.
Targeting is driven by what's happening now, in real-time, by a human, not by which rule-defined audience or topic interest a visitor has been mapped into. Nor by behaviour that's already happened. This means you can respond to the signals that sit between events and pageviews; what we call moments that matter.
You can read more about How Made With Intent Works, and the moments that matter, by clicking the link.
2. Segment the right experience to your customers automatically, without fiddling with rule trees manually
Your team currently builds and maintains audiences in Optimizely's Audience Builder, Dynamic Customer Profiles, Adaptive Audiences, and ODP. All with decision rule trees, content-tagging taxonomies, attribute conditions, event conditions.
Even with contextual bandits reallocating traffic within a defined audience, the audience definition itself stays manual: design the rules, QA the segments, redesign as the catalogue and content evolve. These segments are website-attributes, too. Not human attributes, not intent based (the most human of all attributes). That's where personalisation really succeeds.
Made With Intent replaces all of the above with continuous intent prediction and delivery.
Our agent decides who sees what, in real time, across hundreds of intent combinations. You'll focus on strategy, creative and proof, instead of tweaking and analysing rules or segments all the time.
The agent retrains daily, which means the experience compounds toward intent combinations where impact is seen and felt. In an always on state, always learning, always getting better.
Where Optimizely's multi-armed bandits (if used) allocate traffic across the variants of a single experiment to maximise one fixed metric within one defined audience, Made With Intent allocates across hundreds of audiences (intent combinations) — adding the layer of to whom and when an experience should be served, and measuring it against a holdback rather than just exploiting the winner.
Made With Intent and Optimizely: In depth
Here is how the two platforms compare on the things that matter most for eCommerce teams:
Where does Made With Intent integrate with Optimizely?
We've broken this section down into three parts. We want to be honest about where you'll gain functionality by utilising Made With Intent with Optimizely, where we augment it, and things we simply don't do, or Optimizely does better.
How Made With Intent adds new functionality to Optimizely
Understand and act on every visitor (including anonymous) from the first pageview. Optimizely's Personalization is rule-driven. It works once a visitor matches an audience rule. Content Recommendations builds a per-visitor interest profile from content engagement, which strengthens as the session progresses.
Made With Intent's model, which collates 50bn+ monthly events across 150+ retailers, reads continuous intent from pageview one, every few seconds. Anonymous, identified, first-time, returning. The opening moments of a session get the same intent read as the tenth pageview.
Automatically identify and serve the best experiences to people without guessing. Optimizely's contextual multi-armed bandit (CMAB, powered by Opal) is the closest thing in their stack, and it's genuinely good, so it's worth being precise about what it does.
A CMAB picks the best-performing variation for each visitor based on context (device, geo, behavioural history) to maximise one primary metric, within a single experiment.
Three design choices define it: the context attributes are declared up front and can't be added or removed once it starts (even paused); it optimises exclusively to a single primary metric fixed at launch; and it begins with a 100% exploration phase, randomly serving variations until it has gathered enough data before it shifts to exploiting the winner.
It's a smarter way to split traffic across the variations you built for the audience you defined.
But we'd like to go into detail on what that context is.
Device, geo and behavioural history are proxies for a person. They describe who a visitor appears to be, not what they want or how close they are to buying. They're arbitrary website attributes that correlate with conversion only loosely, and a bandit optimising over them is tuning against a weak signal.
Intent — buying stage, momentum, hesitation, purchase confidence — is the proximate driver of what a visitor actually does next. The proxies describe identity; intent describes decision, and decision is what moves the metric. Optimising the allocation over the wrong variable caps how much a CMAB can ever find.
Made With Intent's agentic campaigns work a level up.

Rather than splitting traffic across the variations of one experiment to maximise one metric, it allocates across hundreds of intent combinations.
Diamonds Factory, a Made With Intent customer, ran 560+ on a single abandonment use case, where the 'context' is live, multi-dimensional intent (stage, signal, trend, purchase confidence, abandon risk, mindset) that updates every 3–5 seconds and is discovered by the agent, not enumerated by the team upfront. Learn more about basket abandonment here.
There's no per-experiment exploration tax, because the model is trained across 50bn+ events, 150+ retailers and live from page view one.
And because a bandit is built to shift traffic toward winners, it has no standing no-treatment baseline — Made With Intent keeps a holdback on every experience, so it can answer "did this cause incremental orders," the question a metric-maximising bandit structurally can't.
Causal incrementality on every experience. Optimizely's Stats Engine reports variant lift with always-valid p-values; genuinely strong for variant comparison.
Made With Intent runs a holdback group on every experience by default and reports incremental orders against a no-intervention baseline.
The difference is between "variant A beat variant B" and "this experience caused X orders that wouldn't have happened otherwise."
Where Made With Intent improves Optimizely
These are things Optimizely does that get better with Made With Intent on top:
- Audience Builder and ODP audiences become intent-aware
Instead of boolean rules on attributes and events, or content-engagement interest categories, Optimizely audiences can take Made With Intent's live intent attributes and target on signals that actually predict conversion. ODP's segment builder picks these up as attribute conditions; Web Experimentation picks them up as Tags. - Personalization variants get the right routing
Made With Intent decides which Optimizely-created variant a visitor should see based on their current intent state, replacing rule-based audience routing. The marketer keeps the variant production; our agent handles the allocation. - Web Experimentation tests get causal lift on top of variant performance. Run Optimizely tests with Stats Engine as you do today. Add Made With Intent's holdback measurement on the experience itself to answer "would users have purchased regardless of any variant?"
- Content Recommendations recommend to live intent, not just topic affinity. Made With Intent lets Content Recommendations reflect what's happening in this session, not just historical topic engagement. Particularly useful for anonymous visitors where you’re not sure what their behaviour is telling you.
What Made With Intent doesn't do
- Feature flagging and server-side experimentation. Optimizely Feature Experimentation (SDKs in 19+ languages, Edge Workers, Agent microservice) is a category leader. We don’t offer anything here.
- Sequential testing methodology for variant comparison. The Stats Engine's mSPRT-based always-valid p-values are best-in-class for inferring variant winners under continuous monitoring. Made With Intent uses Bayesian A/B with holdback.
- CMS-native content personalisation across the wider DXP. Optimizely Content Cloud, Content Marketing Platform and the broader DXP integration are all things that are not within Made With Intent's scope.
- Made With Intent isn’t going to change your recommendations process Optimizely's NLP-driven content recommendations and ecommerce product recommendations are well-established. Made With Intent adds an intent layer; we don’t replace what’s powering your recommendations process.
- Mobile app personalisation. Optimizely Feature Experimentation has native SDKs for iOS, Android, React Native. Made With Intent is web-first.
- Edge experimentation. Optimizely's Edge Worker integrations (Cloudflare, Akamai, Fastly) sit outside Made With Intent's use case

How to get started with Optimizely and Made With Intent
- Use Made With Intent to create intent-ready audiences in Optimizely: Pass Made With Intent's live intent attributes — purchase confidence, abandon risk, buying stage, intent trend — into ODP as attribute conditions, or into Web Experimentation as Tags.
Build audiences like "high purchase confidence + declining intent trend" (needs reassurance), "medium confidence + high abandon risk" (needs timely intervention), "low confidence + active comparison signals" (needs guidance, not a discount). - Execute experiences in Optimizely using those audiences. Use Web Experimentation and Personalization for what they do well, so things like variant production, Stats Engine analysis, content variations across the DXP.
- Serve the experience (or a variant) through Made With Intent's decision agent, where the experience is served based on visitor intent rather than rule-defined audiences.
- Use Made With Intent to prove the incrementality of "Optimizely + intent". Holdback groups on top of the Optimizely experience answer the CFO question: did this cause incremental orders, or did we just personalise for visitors who would have bought anyway? Particularly valuable for discounting — Made With Intent surfaces which high-intent visitors needed no incentive.

Neve Jewels Group, the luxury jewellery brands Austen Blake and Sacet, moved from universal promotions to intent-level targeting across their basket abandonment strategy.
Rather than applying the same discount to every abandoning visitor, Made With Intent segmented interventions across four intent levels, targeting what Director of Customer Experience Jo Homer described as "the nudge moment."
The result: 13% conversion uplift on basket abandonment, and £2.4m in annual revenue uplift overall, 4.8x the original business case, paid back within a single experience. You can learn more here.
What sort of businesses work best with Made With Intent?
This all depends on the size and structure of your eCommerce operation.
Mid-market retailers (£20m–£100m online revenue)
Made With Intent often leads here. Optimizely's Intelligence Cloud commonly lands at £50–80k+/year for this band, based on G2 reviews and procurement data. Made With Intent's session-based pricing, 30 to 60 minute install, and agentic layer fit eCommerce teams of two to ten people who need to ship and prove things quickly.
If you are running Optimizely for the Stats Engine and feature experimentation and those are load-bearing, keep them. Add Made With Intent for intent targeting and causal measurement of your on-site experiences.
Enterprise retailers with dedicated personalisation teams
Both, with clear role separation. Optimizely for the experimentation surface, Stats Engine credibility, server-side feature flagging, and DXP integration. Made With Intent for on-site intent-driven decisioning and causal measurement.
Try Made With Intent today
There you have it. Made With Intent and Optimizely are designed to complement one-another, not compete with one another. Many of our clients use Optimizely to continue A/B testing, while using Made With Intent to serve experience to customers on a 1:1 basis.
If you're interested in learning more about how Made With Intent works, and how you can use it in your tool stack, book a demo here.
Disclaimer: This comparison is based on publicly available information from Optimizely's documentation, marketing site, and customer reviews as of July 2026. Both products evolve continuously. If anything looks out of date, get in touch.

Editor's note: this article is adapted from a recent Intent Live session with Jack Simkins, Digital Product Manager at Golfbreaks.com. Watch the full recording here.
A test that lifted time on site by 14% still spent its first week looking like a failure. Conversion rate was down. Under a traditional A/B testing programme, Golfbreaks.com would have pulled the experiment.
They didn't. And the reason why says a lot about what happens when a genuinely high-consideration purchase journey meets a testing programme built for one-session eCommerce.

Golfbreaks.com is a golf tour operator based in Windsor, with offices in Copenhagen and Charleston, though the Made with Intent account focuses on the US and UK. It sends golfers on trips ranging from a single night in the UK to a week in Spain or Portugal, across a lot of different golfer segments. It's a lead generation business first. Visitors don't check out online in one sitting, they enquire, then a sales agent works out flights, transfers, accommodation, and course access, and gets them to a booking over the phone, sometimes weeks later.
That's not unusual for travel, where research and comparison typically happen across several separate visits and sites before a decision gets made. Optimising a single-session conversion rate for a purchase that actually plays out over weeks measures the wrong moment entirely.
A booking journey that can't be forced into one session
Jack has spent seven years at Golfbreaks.com, the last couple focused on conversion rate optimisation. "We've got quite a unique scenario whereby we're trying to encourage that inquiry," he said. "Particularly in travel, in the industry in general, it's quite an unusual thing to not be able to book entirely online."
A small portionof trips get booked online. But most go through a sales agent, because a golf trip has too many moving parts (courses, transfers, flights, accommodation, and group logistics) for most visitors to configure and commit to in one sitting.
Colin Spooner, Principal Value Consultant at Made with Intent, put his finger on why that matters: "It's not our traditional eCommerce brand where it's a pure purchase journey. But that almost plays into the hands of intent, where you need to think about that considered purchase and how to get people through the funnel before even thinking about the booking, weeks and months down the line."
Measuring what happens before conversion
"You are what you measure" is a phrase the Golfbreaks.com team has adopted internally. If a new visitor is unlikely to enquire on their first visit, optimising purely for enquiry rate on that visit measures the wrong thing.
So alongside enquiry rate, the team tracks bounce rate and time on site together (a new visitor who bounces immediately clearly hasn't been given a reason to stay), pages viewed per session as a depth-of-exploration signal, and, specifically, movement from low to building intent, the kind of signals behind Made with Intent's content prioritisation and messaging use cases. None of these are vanity metrics here. They're proxies for whether a visitor is progressing through a decision process that runs across several sessions, not on a single visit.
Before building any experience, the team asks these questions to get in their customers shoes:
- What is a brand-new visitor actually trying to work out?
- Who are Golfbreaks.com?
- Can we be trusted?
- Do you have to pay full price up front?
- Can you book online at all?
That last one is really important. Because Golfbreaks.com can't be booked entirely online, setting that expectation early avoids disappointment later in the funnel, right when a visitor is closest to converting. As Colin put it, getting that messaging right up front was "a huge realisation" for how the whole experience needed to be built.
What agentic campaigns change about testing
Golfbreaks.com's testing programme runs on Made with Intent's agentic campaigns. In standard A/B testing, you decide up front which segment sees which variant, based on a hypothesis about who will respond to what. Agentic campaigns invert that: you define the strategy (the moment you're trying to influence, and the goal, whether that's enquiries, conversions, or a secondary metric) and hand the agent your set of tactics. It tests them against real segments and works out which one performs best, for whom, and when, using the same intent signals that power the rest of the platform.
For Jack, a self-described non-developer, the practical benefit was speed. "The tool allows me to get these experiences up much faster," he said. "My concept-to-live process is significantly shorter... it means the agents have got time to learn."
But the deeper change is what gets removed. "No longer am I having to set up those individual segments, or serve experiences to segments that I think will benefit from them," he said. "It's in the hands of the agent to then work out what segments it would benefit from... It's a much wider net." A message built for low-intent visitors might also help a segment already building toward a decision. A manual test is only as good as the human guess behind who it's shown to. An agent testing against a hundred segments simultaneously doesn't have that blind spot.
The "do nothing" variant is a genuine conversion tool
One of the more counter-intuitive parts of Golfbreaks.com's setup is what Jack calls the "do nothing" variant. In standard A/B testing, every visitor sees a control or one of several variants. Agentic campaigns add an option where the visitor sees nothing added or changed at all.
"The do nothing essentially sits within those variants as a copy of the control," Jack explained. "It's a safety net because it prevents us from showing negative experiences to customers that don't need to see it."
"In most experience it's always about adding things onto your site," Colin observed. "Having a version where actually sometimes the best thing is leaving the customer alone to progress, or even suppressing things on site, is a nice alternative to what we've experienced over the last 10, 20 years in experimentation." As the data below shows, it's frequently the top performer, because some visitors don't need an intervention. They're already progressing on their own, or they arrived with enough context that added messaging just gets in the way.

When the agent's early data looks wrong
Here's where the seven-day lesson from the top of this article comes back in. Jack's team built two welcome-visit experiences using the same messaging, one for the homepage, one for a location page such as a product discovery landing point for someone who searched "golf breaks in England."
The homepage experience delivered a 14% uplift in time on site. But conversion rate showed a negative trend for the first seven days. "With traditional AB testing, I would have perhaps turned it off," Jack said. "I would have panicked when I saw negative 14%, and I would have said, this isn't working."
He didn't, because the agent was still learning. That's consistent with what independent testing research shows more broadly: short test windows are more susceptible to random variation, and stopping early on interim results is one of the most common ways a genuinely winning test gets killed before it proves itself. Once the agent had enough data, performance turned around.
Same messaging, two different visitors
The location-page test surfaced something else: identical messaging performed in opposite ways depending on where a visitor arrived. On the homepage, a "how to book guide" message performed best, evidence of a genuinely low-intent visitor who needs some hand-holding.
On the location page, the top performer was a trust-building "number one tour operator" message and the do-nothing variant. Jack's take is that a visitor who searched "golf breaks in England" already has affinity toward the destination, closer to a returning visit mindset than a cold product visit. They don't need the basics explained.
A message about Golfbreaks.com's customisable packages underperformed with brand-new homepage visitors. It's true and important, but it's the wrong message at the wrong moment, the same lesson behind Made with Intent's discounting use case: showing a message before a visitor is ready for it does more harm than good.

The surprise that only showed up in the data
Asked what surprised him most, Jack pointed to something that had been sitting in plain sight. An early "ready to plan?" message aimed at brand-new visitors looked like a reasonable nudge. But in the data, it actually came across as overbearing.
"If you think about a new user landing on the site and asking, are you ready to plan? It's probably a bit overbearing," Jack said. "At the time, when you're setting up those tests, it's like, right, I'm going to use the same messages for the homepage, same message for the location page. They're surely going to work." They didn't, for every segment.
Colin's read: "The amount of times we see customers who have a predefined view of what will work, and it's completely different. That point around being subjective comes to life when you start to see the way the agent starts to make decisions." That tracks with the broader shift in shopper expectations, most consumers now expect a personalised experience and notice sharply when they don't get one, a generic message is no longer neutral, it actively reads as a miss.
When a message doesn't resonate with any segment, the fix is simple. Delete it, let the agent relearn, and add a new tactic later if needed. You don't need to manually re-segment.

Measure the journey, not just the moment
None of this is unique to golf holidays. Any purchase with a real consideration cycle shares the same shape, including B2B software, where roughly 70% of the buying journey typically happens before a buyer contacts a vendor at all. A first-time visitor is rarely the same as a returning, further-along one, and treating them identically wastes the message on the visitor least ready to act on it.
Three things carry over regardless of your industry. You don't need to be a developer to find early wins, a visual editor is enough to start. If your purchase journey has any real consideration cycle in it, top-of-funnel testing should measure more than conversion. And build a tactic sheet of messages at a global level, then let the data show which ones resonate with which visitors, rather than deciding that yourself up front.
If your own funnel has visitors who aren't ready to buy on visit one, the same logic applies. See the intent framework behind it. Or book a demo to see it against your own traffic.
If you've enjoyed this write up of our latest Intent Live session, why don't you join our next one?
No results found.
Subscribe to our newsletter
Written by an actual human our monthly newsletter is your deep-dive into everything Intent-related and eCommerce. Go on, give it a go.




