Social proof is powerful in ecommerce, but overuse makes it noise. Learn why intent matters, how timing changes impact, and how to align social proof with customer mindset to build trust and drive results.
Social proof should be one of the most powerful tools in ecommerce. At its core, it’s the influence that the actions, choices or approvals of others have on an individual’s behaviour.
People look to others when they’re uncertain about what to choose, who to trust, or whether to act. In ecommerce, that influence can appear anywhere in the journey. As reassurance that a brand is worth buying from, or as urgency to act before missing out.
It comes in many formats: scarcity messages (“only 3 left”), activity indicators (add to baskets, recent views, recent purchases), reviews and ratings, and trending or bestseller labels. Used well, these cues can reassure, create urgency, and help people find what’s popular or trusted.
The problem is, social proof has become one of the most overused and underthought tactics in the game. It’s often deployed as a blanket message to everyone, with little thought about whether it fits their mindset or the brand experience.
Retailers love it because it’s quick to turn on and almost always delivers an aggregate uplift. But those uplifts are often driven by a smaller group, and the negative effects on others are hidden in the averages.
The status quo of social proof
Most ecommerce teams apply it generically, showing the same messages to everyone – often on every product page. The most common use is as a conversion-driving technique late in the journey, but there’s a growing trend to apply it earlier in discovery (e.g., “bestseller” on PLPs).
Its popularity comes from being considered “best practice,” easy vendor implementation, and the reliable ROI it shows on aggregate. But those aggregate numbers are disproportionately influenced by high-intent visitors, which hides the harm it can cause to others.
What works for one mindset can actively put another off. As part of our research for The Intent Gap Report, we found:
“Trending” overlays on PLPs positively impact low-intent browsers.
“X sold last week” overlays on checkout pages deliver an average +5% conversion lift for high-intent visitors but cause a -1% drop for low-intent visitors.
Luxury and exclusivity-driven brands often avoid generic social proof entirely. In high-consideration categories, it can feel out of place – an engagement ring buyer doesn’t want to hear that “20 others bought this today,” and a £3000 jacket doesn’t need a flashing urgency tag over carefully curated imagery. In these cases, overlays can jar with the brand and undermine the premium feel.
When social proof is everywhere, it stops providing reassurance or focus. The message becomes noise, prompting the question: why stick with this approach?
Because most retailers rely on page-type triggers (e.g., PDP = ready to buy). But many PDP visitors are still browsing. Without behavioural context, tactics are based on where someone is, not how they’re behaving. That one-size-fits-all approach ignores timing and mindset. And that’s exactly why it needs a rethink.
Social proof with intent
Social proof can reassure early in the journey or create urgency later, but timing and fit are critical. Softer cues like “bestseller” or “trending” help those still discovering products. Urgency or scarcity works best when someone has decided what they want and just needs a final nudge. Use it too soon, and it risks creating anxiety or distraction.
Think of walking into a DIY store paint aisle: if you’re browsing, you don’t want someone saying, “Only three tins left – buy now!” before you’ve chosen a colour. But if you’re holding the exact tin you want, that message might spur you to buy. The same logic applies online.
Or picture a luxury sales assistant with a £3000 jacket. They wouldn’t start with “20 people bought this today.” They’d focus on its quality, heritage, or popular combinations, tailoring the message to the moment.
Real-time intent signals mean you can:
Show discovery-style social proof to those exploring
Reserve urgency and scarcity for visitors with strong product interest or signs of hesitation
Avoid showing it altogether to those it might deter
When you match the message to the moment, social proof stops being background noise and starts driving action.
The path to better social proof
While we’ll cover how to move from generic application to something more intent-based in a follow up, the core steps are:
Analyse performance by visitor mindset, not just aggregate.
Exclude audiences where a message harms conversion.
Adapt style and timing to fit both brand tone and visitor context.
The benefits? Higher incremental gains, reduced brand risk, and interactions that build trust.
Social proof works – but not for everyone, not everywhere, and not all the time. The more you align it with intent, the more it delivers.
Social proof should be one of the most powerful tools in ecommerce. At its core, it’s the influence that the actions, choices or approvals of others have on an individual’s behaviour.
People look to others when they’re uncertain about what to choose, who to trust, or whether to act. In ecommerce, that influence can appear anywhere in the journey. As reassurance that a brand is worth buying from, or as urgency to act before missing out.
It comes in many formats: scarcity messages (“only 3 left”), activity indicators (add to baskets, recent views, recent purchases), reviews and ratings, and trending or bestseller labels. Used well, these cues can reassure, create urgency, and help people find what’s popular or trusted.
The problem is, social proof has become one of the most overused and underthought tactics in the game. It’s often deployed as a blanket message to everyone, with little thought about whether it fits their mindset or the brand experience.
Retailers love it because it’s quick to turn on and almost always delivers an aggregate uplift. But those uplifts are often driven by a smaller group, and the negative effects on others are hidden in the averages.
The status quo of social proof
Most ecommerce teams apply it generically, showing the same messages to everyone – often on every product page. The most common use is as a conversion-driving technique late in the journey, but there’s a growing trend to apply it earlier in discovery (e.g., “bestseller” on PLPs).
Its popularity comes from being considered “best practice,” easy vendor implementation, and the reliable ROI it shows on aggregate. But those aggregate numbers are disproportionately influenced by high-intent visitors, which hides the harm it can cause to others.
What works for one mindset can actively put another off. As part of our research for The Intent Gap Report, we found:
“Trending” overlays on PLPs positively impact low-intent browsers.
“X sold last week” overlays on checkout pages deliver an average +5% conversion lift for high-intent visitors but cause a -1% drop for low-intent visitors.
Luxury and exclusivity-driven brands often avoid generic social proof entirely. In high-consideration categories, it can feel out of place – an engagement ring buyer doesn’t want to hear that “20 others bought this today,” and a £3000 jacket doesn’t need a flashing urgency tag over carefully curated imagery. In these cases, overlays can jar with the brand and undermine the premium feel.
When social proof is everywhere, it stops providing reassurance or focus. The message becomes noise, prompting the question: why stick with this approach?
Because most retailers rely on page-type triggers (e.g., PDP = ready to buy). But many PDP visitors are still browsing. Without behavioural context, tactics are based on where someone is, not how they’re behaving. That one-size-fits-all approach ignores timing and mindset. And that’s exactly why it needs a rethink.
Social proof with intent
Social proof can reassure early in the journey or create urgency later, but timing and fit are critical. Softer cues like “bestseller” or “trending” help those still discovering products. Urgency or scarcity works best when someone has decided what they want and just needs a final nudge. Use it too soon, and it risks creating anxiety or distraction.
Think of walking into a DIY store paint aisle: if you’re browsing, you don’t want someone saying, “Only three tins left – buy now!” before you’ve chosen a colour. But if you’re holding the exact tin you want, that message might spur you to buy. The same logic applies online.
Or picture a luxury sales assistant with a £3000 jacket. They wouldn’t start with “20 people bought this today.” They’d focus on its quality, heritage, or popular combinations, tailoring the message to the moment.
Real-time intent signals mean you can:
Show discovery-style social proof to those exploring
Reserve urgency and scarcity for visitors with strong product interest or signs of hesitation
Avoid showing it altogether to those it might deter
When you match the message to the moment, social proof stops being background noise and starts driving action.
The path to better social proof
While we’ll cover how to move from generic application to something more intent-based in a follow up, the core steps are:
Analyse performance by visitor mindset, not just aggregate.
Exclude audiences where a message harms conversion.
Adapt style and timing to fit both brand tone and visitor context.
The benefits? Higher incremental gains, reduced brand risk, and interactions that build trust.
Social proof works – but not for everyone, not everywhere, and not all the time. The more you align it with intent, the more it delivers.
"The website is good, it probably could be even better, no doubts. But it's "good enough". It does the job. It's stable. It's mobile optimised. It's got a good conversion rate relative to others and our expectations. What else can we do? Is it the right thing to do to have a team purely focused on conversion rate optimisation, or should we look more into product and trading? What even is the potential of our site?"
That was a thought-provoking quote direct from a prospect of ours in a recent sales call.
He's talking honestly about the idea of prioritisation and diminishing returns. A very well-known brand, decent infrastructure, good content, a site that works very well; probably even over-indexes on conversion rate efficiency if you're to compare it to competitors. He'd looked at the size of the remaining prize from making the digital experience better and quietly concluded it might be smaller than the prize from doing something else entirely.
This is a question of "where do you place your bets?"
I think they have reached a local maximum. For anyone who hasn't sat through the optimisation lecture: it's the top of a hill that isn't the top of the mountain. Every small step available from where you're standing leads downwards, so you stop climbing. Not necessarily because you've reached the highest point there is, but because you've reached the highest point reachable in small steps. Which is a fairly precise description of what a decade of testing does to an already-decent website. In other words, getting to a higher peak means changing direction.
He's asking the right question in my opinion and I think the honest answer is uncomfortable for most of the industry I've spent my entire career in, especially the purists.
This brand hasn't necessarily hit the limit of what's possible on their website. But instead has hit the ceiling of the average and anything further sees diminishing returns where the effort doesn't necessarily equate to the value. That, or he's potentially bored with the same-same solutions that are out there. Homogeneity is the killer of excitement.
I empathise. I got bored too.
I founded User Conversion; one of UK's most successful (read: largest?) independent conversion rate optimisation agencies. We did well, working with some of the biggest brand names the UK had to offer.
But like the above brand, over time, I grew more and more skeptical. A lot of our recommendations lacked creativity, they were all addressing similar problems with the same solutions. "Moving deck chairs on the Titanic" is what someone once put to me.
Conversion rate optimisation is a process of problem-solving with evidence based solutions. Learning, uncovering opportunities and problems, and fixing those problems; usually through AB testing (well, that's the outcome that most cared about; because it's sexy). And I'm not suggesting that the learning and the opportunities dissipate, but the solution often lacks impact because of the law of diminishing returns, their heterogeneity and the aggregated nature of them.
The evolution resolution beyond CRO
Conversion Rate Optimisation, for stakeholders at least, is a way to make the website earn more; and there are four ways to do that. Most of us have treated them as a maturity ladder. You graduate from one to the next, and the last one is the good one.
That's not quite right. It's less an evolution and I now see them more as levels of resolution where each one narrows the unit of decision.
Level one: conversion rate optimisation. The unit of decision is the average visitor. You look at where people struggle, you fix it, and the fix applies to everybody. This is genuinely valuable and I'd never argue otherwise but it is a) practically often an exercise in usability improvements and b) definitionally serving the mean.
The first statement encompasses this idea that the majority of solutions are things that don't change behaviour, they facilitate existing behaviour. Usability improvements. Small changes that ill-advised vendors purporting marketing promoting statistics have convinced us are worth the effort. We've all seen them. The famed 500% uplifts. Sticky add to cart buttons, adding trust signals under a call to action, that sort of thing. Read: deck chairs on the Titanic.
The second reinforces the statement that the mean doesn't exist. Instead, it is a continuously moving combination of different intent levels. Our own research found that, say, 10% of visitors sitting on checkout pages aren't ready to buy yet. Or that 34% of users never get past browsing, whatever page they happen to land on. There is no average shopper to optimise for, there are different jobs to be done. There's a distribution we've been flattening for twenty years because flattening it was the only thing we could do. And now we're used to that, we lack creativity of how to proceed.
Level two: experimentation. It's the same unit of decision, the average, but now you've proven it. This matters enormously and it's the most rigorous thing most organisations do.
From experience, it often comes in two flavours:
1. The immature version lives inside the CRO team or the marketing function, running client-side tests and it caps out at about four to six experiments a month. That ceiling is a resource limit, often not a statistical significant limit; people, build time, roadmap slots. You'll find here that you'll max out at a certain number of tests usually and your impact is limited to avoid cross-contamination of tests. Ever found yourself saying "we can't run a test on a PDP because we have something running there already?"
2. The more mature version is server-side experimentation, owned by engineering, decentralised across product teams, with testing built into the release process rather than bolted onto it. That version is genuinely near-limitless in cadence, and if you can get there you should; it's fantastic.
But note what even the mature version is for. It's designed to prove or disprove a claim about the population. One answer, for everybody, with confidence attached. That's the instrument working exactly as intended; but it's still an answer about the average. It's also an answer about the website, not the visitor (more on that later).
My prospect's line was exactly this: "For the years we have experimented, we did not see huge benefits. That's probably also what hindered further investment." I've heard that sentence in some form from almost every brand I've worked with.
Level three: personalisation. Here the unit of decision finally narrows to the segment. And here is where the industry has spent a decade making promises it couldn't keep. Unfortunately, to the extent where we now all hold PTSD; personalisation traumatic stress disorder.
Personalisation didn't fail (that's right, I said it failed) because it was a bad idea, every boardroom still talks about it to this day. Trust me, I literally wrote the book on it: The Person in Personalisation.
It failed because it doesn't scale, for two reasons, both of which compound.
1. First, you have to guess the segments before you have any evidence about which distinctions matter. Those pre-defined segments like "returning visitor," "paid traffic," "landed on a PDP" are website attributes, not people attributes. That's not person-alisation that's website-alisation, isn't it? Pageview count stands in for engagement, when a confused shopper racks up far more pageviews than a decisive one i.e. it's not true person-alisation.
2. Second, the arithmetic defeats you. Split traffic three ways and every test takes three times as long to reach significance. Try to prove the segments genuinely differ and you're chasing an interaction effect that needs roughly four times the sample again. A three-week test becomes a quarter-long project, and most teams call it early and ship an artefact, or just don't have the traffic (and therefore patience) i.e. it's not scalable.
So personalisation became a small number of hand-built manual rules, maintained by someone who'd rather be doing something else, delivering less than it promised. Ever wondered why recommendations was the only successful personalisation that brands have achieved? Because it's autonomous; in other words, scalable.
Level four: agentic delivery, with intent as the context. This is where we, Made with Intent, sit. The unit of decision becomes the person and what they're trying to do in the moment. Not the segment on retrospective data. Also, not the average. And critically, nobody writes the rule.
You give the system a strategy and a set of experiences that are already evidenced. It works out which of them suits which state of intent, person by person, at the time that it matters, and it keeps working it out. Nobody writes the rule.
Take one of our customers, Diamonds Factory who had a single basket-abandonment tactic: 25% off, to everybody. They gave the agent four options instead and let it choose between them. Most people, it turned out, didn't need the full discount to convert. And 15% needed no intervention at all.
That last number is the one that matters, because no level below four can produce it. A test has no vocabulary for show this to nobody as a good outcome. Not even the control of an experiment can show you that because a user is never bucketed into both the control and the treatment. A rule-based segment can't discover it. It only appears when something is allowed to decide, per person, whether to act at all, in the moment that it matters.
The Future of Personalisation
I think we treated this as four evolutionary components within a single ladder when, in fact, they're two.
• Levels one and two are about how sure you are. "Does this work." Think of this as 80% exploration, and 20% exploitation.
• Levels three and four are about how precisely you aim. "For whom does it work best, when should it be shown, and do a proportion of users even need it at all?" Think of this as 20% exploration, and 80% exploitation.
Conflating them is why so many personalisation programmes were run by people optimising for certainty, and why so many experimentation programmes never escaped one answer for everybody. Confidence and aim are different problems. You need both, and the tools for each are not the same tool.
The industry has worked out that continuous contextual allocation beats a fixed split. That's 50% of personalisation; serving "the right person, at the right time, with the right message".
The other 50% are the attributes that determine whether something is personal; and for us that's their intent. Their context. Their job to be done. The differentiator is what you put in the context window. Not a website attribute like device or location, because they describe who someone appears to be. Intent describes what they're about to do. One of those is a proxy and one of them is the thing itself.
What I'd actually tell my prospect
Not "invest more in CRO." And not "stop doing CRO," either, because each of these levels still holds a purpose and pretending otherwise is how vendors lose credibility. No, CRO is not dead.
But its purpose has changed. There's more.
Sure, if something is broken, fix it. If a step in the core journey confuses everyone like a login, a checkout, a filter that doesn't work, then everyone passes through it, the fix helps or hurts them all in the same direction, and the right answer is one answer. Optimising the average isn't a failure of ambition there. An experimentation partner of ours put it well: you don't stop the research, you don't stop the UX work, and if you don't have people designing properly for their users you're finished as a business regardless of what any agent does on top.
But once the site is good enough; once you've fixed what's broken and the remaining UX gains are genuinely marginal, I think the question changes. It stops being how do we make this better for everyone and becomes which of the things we already have should this particular person see, and when, and should they see anything at all.
Essentially the argument of diminishing returns. Unless your site is broken, terrible UX, or hard to navigate; the best bet is dynamic, trading-related experiences which capitalise on serving the right content at the right time. Sure I'm biased, but this is my arc within this industry over the past 15 years. I've seen what's possible and I'd like to share it with the world. A TLDR;
There are different ways to optimise, but just note that I've seen first hand that scalability is the biggest constraint to success. Not just that, but doing the same as everyone else, particularly "moving deck chairs on the Titanic" won't get you very far unless the baseline is so low.
The opportunity is vast. It amplifies existing experiences by autonomously serving those only to where the experience is best seen and best felt, excluding where it's not. Because no one in the organisation owns it, or because we are so accustomed to the way things work currently; the opportunity is still sitting there.
My prospect's instinct was right. He should probably move effort away from optimising the average. Just not away from the website.
If you're interested in how CRO is evolving, and want to learn more about intent-based personalisation, get in contact with our team here.
Experimentation is evidence to determine whether something works. Intent decides who receives what, when (or whether they need it at all). “Does this work.” Think of this as 80% exploration, and 20% exploitation.
The first works towards the best single answer for everybody to prove a hypothesis. The second works out where a single answer was never going to be enough; scaling a proven hypothesis for true commercial gain. “For whom does it work best, when should it be shown, and do a proportion of users even need it at all?” Think of this as 20% exploration, and 80% exploitation. Where testing tells you whether an experience works, but Intent decides who receives it, when, and whether they need it at all.
In this blog post, we'll identify the differences between Intent and experimentation, so you can understand when to utilise each strategy to further your business goals.
Defining A/B testing
An A/B test answers one single bounded question.
Does this change (a treatment) move the metric we care about, across our traffic, better than the alternative does (a control)?
It's a hypothesis, and what comes out the other end is 80% knowledge (learnings) and 20% value (hopefully positive gain). A primary "does this prove or disprove our hypothesis" binary answer.
Our intent engine answers a different question, and it answers it again every few seconds throughout a users session:
Given what this visitor is doing right now, which of the experiences we already know works suits them at this moment in time? (If any)
What comes out of that is not just a single decision, but hundreds of them a second, all trying to move the agents primary goal, or reward. The open question stopped being does this work a long time ago. What's left is who needs it, and when.
For example: if a visitor starts showing signs of basket abandonment. You can put a returns message in front of them, or a discount, or a trust signal, because you already trust all three of those mechanics. What you don't know is which one that particular person needs, at what point, or whether they need any of them at all. That's the decision the agent is making.
There are several reasons why experimentation (read: split testing) might not be an appropriate mechanism; here are four reasons:
But first, a note on the word experimentation. Throughout this set, experimentation means the whole practice of using evidence to drive ideas: user research, usability testing, prototyping, quasi-experiments and switchbacks where a split isn't possible, and A/B tests where it is. The split test is one instrument inside it, usually the last step rather than the whole of it.
1. Averages can lie.
An A/B test reports an average treatment effect which can hide segments or audiences that have a detrimental or negative impact.
An A/B test reports an average treatment effect. That's the whole point of it. Randomising across your population is what buys you a causal claim in the first place. The price you pay is that the average swallows everything inside it. A headline +2% can be +12% for hesitant first-timers and -4% for loyal returners. Same test. Same green tick. Two completely different stories.
Statisticians call these heterogeneous treatment effects, and there's good evidence that acting on them beats a blanket rollout. In one worked simulation of a ranking change, treating only the 69% of users predicted to benefit delivered a 2.2% gain in revenue per user against shipping it to everybody. A CUNY study of financial aid nudges captured roughly 75% of the total benefit while treating half the population.
Experimentation is deliberately built to stop you trusting the segment breakdown. Pre-registration, power calculations, no post-hoc slicing. Those rules exist because sliced results are where false positives breed. Kohavi, Deng and Vermeer's A/B Testing Intuition Busters (KDD 2022) found that underpowered analyses inflate the effects you observe by 25 to 50%. At the power levels most mid-market programmes actually run at, over half of your significant results are false positives. So the question you have to answer to deploy well, for whom, is the question your test is worst at answering. That's a boundary, and a sensible one. Intent sits on the other side of it. It predicts for every visitor in advance, rather than carving up a sample after the fact.
Where Intent solves this problem: The experience is being delivered across hundreds of contexts, reallocating to where the experience "works" (where it's best seen and best felt), excluding it from where "it doesn't work." In theory, it only looks and focuses on the good, excluding the bad.
2. Testing assumes the trigger.
Every A/B test starts from a fixed, assumed rule (usually on page load). The trigger "when to show it" is often assumed, arbitrary, and based on a website proxy like "3 page views" or "PDP" rather than genuine user intent.
Nearly every A/B test starts from a fixed rule. Show this on the PDP after three pageviews. Fire this on exit intent. Serve this to returning visitors. Then it tests which creative performs best inside that rule.
The treatment gets examined. The trigger gets assumed.
There's a practical reason for that. The trigger space is far too big to write out by hand. Nobody can enumerate every combination of buying stage, purchase confidence, abandon risk, intent trend and product affinity, let alone power a test across them. So teams pick a proxy or an average and move on.
Proxies are usually where the damage happens. These are largely based on website attributes as opposed to personalised signals. Returning visitor stands in for higher intent. Mobile stands in for lower intent. Pageview count stands in for engagement, when a confused shopper racks up far more pageviews than a decisive one, and high intent visitors often convert inside three.
Our own Intent Gap research found 10% of visitors on checkout pages aren't ready to buy. And 34% never get past browsing, whatever page they happen to be on. There's a reason we say stages, not pages.
Where Intent solves this problem: Intent agentic delivery serves the individual at a moment in time. A user might need an experience on the 3rd page view after 30 seconds, a different user might need it after 50 seconds, a different user might need it after 9 page views and 1 second. The continuous prediction modelling gives a threshold to meet first, and then a mechanism to automate a decision second.
3. Statistical significance is a search for one answer.
The search for statistical significance can a) slow you down and b) inhibit your ability to personalise to different segments.
Significance exists to license a claim about a single population, and the claim is singular by design where B beats A: one answer, for everyone. That is exactly what costs time.
You need a sample big enough to speak for the whole population with a minimum detectable effect. If you want the claim to be about a segment instead (what some call "personalisation") you need that sample inside the segment too, so the wait multiplies. Most claim they don't have the traffic to personalise because of this.
Experimentation vendor's posterior is a belief about a variant. One distribution per arm, computed across everybody who saw it, resolving to a single winner for all your traffic. Better maths than a p-value, and the shape of the answer is identical. One number per variant, and a threshold you wait to cross before you act.
Our posterior at Made with Intent is a belief about a variant given a context. Where their question is whether B beats A, ours is whether B beats A for a visitor in this intent state, at this point in their session. The model predicts reward conditional on context, so what comes out is a policy; hundreds of answers, running at once.
Put it this way: instead of a human designing three segments and running three underpowered tests, the winner goes to the agent alongside its alternatives and the segment level optimisation runs continuously. Three mechanisms make that more traffic-efficient than a segmented test. Worth understanding rather than taking on trust.
1. Allocation is adaptive. Thompson Sampling draws plausible performance rates from each variant's posterior and routes traffic accordingly, so your exploration cost falls as confidence climbs rather than sitting at 50% for the duration. In practice agentic campaigns push 80 to 90% of traffic to the winner inside two to three weeks. A fixed-allocation A/B test takes around six. Optimised Control shrinks the control group automatically as confidence grows.
2. Without waiting for statistical significance. It borrows statistical strength across slices. This is the important one. Made With Intent runs a contextual bandit, not a multi-armed one. A multi-armed bandit only learns from the data each arm receives, so every segment needs its own volume, which is the same problem as a segmented A/B test in different clothing. A contextual bandit trains a model that predicts reward given context, so it can estimate performance for a combination it has never directly seen. Show it mobile traffic, low intent traffic and a particular page separately and it can predict for "mobile, low intent, that page" by generalising from how each component behaves. A hand-built segment can't borrow like that. Every cell starts at zero.
3. You're not assuming which segment to slice or review. The agent can build up to 500 contextual combinations and up to 75 moment-based triggers per campaign, and it trims its own feature selection to keep each slice sample-rich. Nobody has to guess which distinctions matter before there's evidence about which ones do.
Why we don't report per-segment significance
This surprises experimentation-native teams more than anything else in the product. It's a deliberate methodological choice.
Run an independent significance test on every intent combination and multiple comparisons swallow the results. At α = 0.05, roughly 14 independent tests give you about a 50% chance of at least one false positive. A hundred tests will throw up around five false positives by chance alone, and across hundreds of combinations you'd be manufacturing findings.
We trialled per-segment significance reporting and pulled it, because it surfaced spurious and inverse correlations. For example, users traverse through different stages of intent in their journey, at what point is a user in low intent and when did they see the treatment?
The per-segment decisions get justified differently. Is the model calibrated, meaning when it says 70% does it convert around 70% of the time (measured by expected calibration error)? Is it discriminating, meaning can it rank likely converters above unlikely ones (AUC around 0.84)? A well calibrated probability is a legitimate basis for acting on a slice. An underpowered significance test on that same slice is not.
This calls back to something we talked about earlier.The reason you shouldn't chase per-segment significance in your testing tool is the same reason we don't chase it in ours. Prediction is the right instrument for the for whom question, retrospective slicing isn't.
Why we don't report on single variant uplifts
You get one clean causal number, control against the agent's allocation, measured the way any experimentation lead would want it measured. You give up per-arm inference, and in exchange the arm is adaptive. That's the trade. An A/B test gives you clean inference on every arm and one winner for everybody. An agentic campaign gives you clean inference on one arm, and a different winner in every context.
Load five experiences into an agentic campaign and it looks like a five-arm test. It isn't. It's actually a two-arm experiment, the same shape as any A/B test you've run. The difference being:
• Control. A baseline you define.
• Variant. The agent's allocation across all five experiences.
A traditional A/B test asks which variant wins with a blanket rule. An agentic campaign asks a different question with an adaptive rule (it's a reason why we call them campaigns, not experiments)
Does allocating these experiences by intent beat applying a blanket rule?
We don't report on single variant uplifts eg: Variant B is better by 5% because:
1. There's no single number to give you. Experience B might be the winner for high intent, budget-conscious mobile visitors and the loser for low intent desktop. That's the entire point of running it this way. Collapsing it to one figure puts back exactly the average this campaign exists to get rid of.
2. The arms were never randomised against control. Traffic reaching experience B was chosen by the agent, on context and accumulated evidence. It isn't a random slice of your audience so there's no effective control. Comparing a deliberately selected group against everybody is confounded by construction, and nothing in the data separates the effect of the experience from the effect of the selection.
3. You'd be reading a state the system has already left. If experience B was struggling in a context, the agent would have noticed days ago and moved traffic away (hence: reallocation). The number in front of you is an average across a period in which allocation was actively changing, and the campaign has moved on since. Acting on it means acting on history the agent has already corrected for.
That distinction does three things, in order.
1. It changes what you can personalise. A per-variant posterior can only ever produce one answer for everyone, however good the statistics behind it are. Varying what people get requires probability conditional on the person. That's the whole game, and it matters far more than the frequentist argument ever did.
2. It changes when you can act. Certainty becomes a dial rather than a gate. At 62% probability, Thompson Sampling tilts the allocation and updates again tonight. Nobody declares anything. A stopping rule stops being necessary once allocation is continuous.
3. It changes how fast you get there. Nothing is waiting for a threshold, so allocation improves from the first night onward. There's no finish line to reach before the work starts paying.
Where Intent solves this problem: Our posterior at Made with Intent is a belief about a variant given a context. Where experimentation question is whether B beats A, ours is whether B beats A for a visitor in this intent state, at this point in their session. The model predicts reward conditional on context, so what comes out is a policy; hundreds of answers, running at once based on a series of probabilities using a contextual bandit.
4. Limited to a single answer for only one point in time.
An A/B test only tells you what was true of your traffic during the weeks the test ran.
An A/B test tells you what was true of your traffic during the weeks the test ran; a shelf life. Nothing about that answer refreshes itself so brands often end up "re-testing" the same thing a few years later, hoping there's no change from a previous positive test. The world moves on, results degrade as users become "used" to the treatment, competitors copy; not to mention that your purchase lifecycle varies by days, weeks, months, years.
For example: a Microsoft experiment on MSN.com saw replacing one button with another produce a 4.7% increase in overall clicks. The daily breakdown post launch showed the difference decreasing rapidly day over day as users learned the change, leading to the team shut the experiment down mid-way. The conclusion was that the observed treatment effects "are not always permanently stable, sometimes revealing increasing or decreasing patterns over time."
There are a few reasons why this happens.
1. Novelty and primacy wear off. A new element gets attention because it's new. Returning visitors are briefly worse off because it isn't what they knew (which is why tests are often split, somewhat arbitrarily, between new and returning users)
2. Buying patterns differ throughout the year. December traffic behaves nothing like March traffic, different intent, different price sensitivity, different tolerance for being interrupted (especially for retailers)
3. The population itself drifts. Mix of paid traffic allocation shifts, category mixes alter or a competitor changes their delivery proposition. Macro-economic factors alter the interaction effect.
Does a mature site slowly accumulate a layer of decisions that were correct once, are serving everybody by default, yet are answerable to nobody? It is entirely possible to have a well-governed testing function and a site full of expired answers at the same time.
Where Intent solves this problem: Continuous allocation doesn't have this failure mode, because it never declares anything. The posterior is a live belief rather than a verdict, continually reallocating and therefore responding to changes, not static. Allocation is re-scored against current behaviour and retrained nightly, and standing exploration keeps a small share of traffic asking whether the current answer is still the right one. When February stops behaving like November, the agent finds out because it never stopped looking.
Intent vs experimentation: In depth
We've done a quick reference table on the differences between Intent and experimentation. Take a look, send it to your colleagues. You're welcome.
Experimentation
Intent-driven agentic delivery
Purpose
80% Exploration, 20% exploitation. Tells you whether an experience works
80% exploitation, 20% exploration. Tells you who receives an experience, when, and whether they need it at all.
What it decides
What to do, and whether it works. Proving a hypothesis.
Who receives it and when (or whether anyone should right now)
Unit of analysis
Entire population with an average effect.
The individual visitor, in session, re-scored continuously
What comes out
Evidence you can act on, value at an aggregate level.
A continuous daily allocation to where each experience is best seen and felt (i.e what works)
What counts as success
A trustworthy answer, including a learning of no effect
Incremental orders and protected margin
Best answer it can give
One answer, for everybody
A different value per person and per moment
How it's governed
Statistically. Power, statistical significance
Commercially. Probabilities, predictive modelling
What the probability is over
A variant. One posterior per arm, across everyone exposed. Frequentist or Bayesian, the shape of the answer is the same.
A variant given a context. A posterior per experience per intent state, so the output is a policy rather than a winner.
Arms under test
N variants, each measured against control. One winner, rolled out to everyone.
Two. Control against the agent's allocation policy. The policy contains N experiences, allocated per context.
Hit rate
10 to 20% of ideas move the target metric. Roughly 1 in 500 is a breakthrough. (Kohavi et al., 2014)
77.6% of agentic campaigns produce a result, against the under 20% of tests that beat their control. Different denominators, worth saying out loud: a campaign counts when it finds the right answer for some contexts, a test counts only when one variant beats another across everybody.
Time horizon
Fixed at the test window. Degrades silently after rollout; only a re-test reveals it, and nothing prompts one
Nothing is declared, so nothing expires. Reallocated nightly against current behaviour, with standing exploration watching for drift.
Experimentation and Intent aren't solving the same problem, and treating them as substitutes is where most testing programmes stall. An A/B test proves whether something works for everybody, on average. But it can't tell you who needs a returns message versus a discount versus nothing at all, and it becomes redundant the moment the traffic that validated it moves on.
Intent picks up where that boundary sits, deciding who an experience is for, when, and whether they need it at all, re-scored every few seconds instead of declared once and left to expire. Most eCommerce teams already have the first instrument. The second is what turns that evidence into revenue, visitor by visitor, instead of one rollout for everybody.
Book a demo to see how Made With Intent can help you deliver more appropriate experiences to your customers.
Made With Intent is an on-site intent engine for eCommerce businesses. It reads buying intent in real time and decides which experiences go to which visitors, and when.
Dynamic Yield is an experience optimization and personalization platform. It helps businesses tailor digital customer journeys across websites, mobile apps, and email.
This post explains where the two platforms work together, and where you'd still use Dynamic Yield instead.
If you can't wait until the end, here's a TLDR;
Pick Dynamic Yield if:
• Your priority is consistent personalisation across web, mobile app, email, and in-store kiosk
• You've got the team and budget to do a weeks long enterprise deployment
• Recommendation capabilities matters more than first-page view intent coverage.
Made With Intent is for you if:
• You want to read buying intent in real-time
◦ Then, serve the right content, the right experience at the appropriate moment
• Want to benefit from a 30-60 minute install time
And when we're thinking about Made With Intent and Dynamic Yield working together, here's how we recommend thinking about it:
Made With Intent decides what experiences you should serve, to who and when, and Dynamic Yield delivers those experiences.
How Made With Intent is different to Dynamic Yield
There's three things that Made With Intent does different to Dynamic Yield. Let's dig into them below:
1. Made With Intent reads intent in real time
Dynamic Yield's intent workflow utilises two routes. First, Empathic Personalization classifies visitors into four inferred states, which are Curious, Interested, Focused and Satisfied.
Its Audience Hub lets teams hand-build "low / medium / high intent" audiences from rules. Dynamic Yield's Primary Audiences framework shows a leading golf retailer defining low intent as fewer than 12 page views per session and high intent as more than 24.
That's a useful starting estimate, but page views aren't intent. A hesitant shopper racks up more pages than a decisive one. A high-intent visitor often converts inside three.
The more common version isn't pageview counts — it's event proxies. Add to cart equals high intent. Wishlist equals consideration. But an add to cart is as often a price check, a size comparison or a shipping-cost probe as it is a purchase signal. The event tells you what happened. It doesn't tell you what it meant.
Now, this is pretty fundamental: Dynamic Yield's model relies on behaviours your shoppers have already exhibited. It's a snapshot backwards in time, and like many what we call "rules-based" personalisation tools, you're reacting to things that have already have happened. Unlike Made With Intent.
We break downs shopper behaviour into six stages, and these are: Intent stage. Intent signals. Intent trends. Purchase confidence. Abandon risk. Shopper mindset.
All updated every three–five seconds during a live session, on a model trained across 150+ retailers and 50 billion+ events.
2. Allocation across hundreds of intent combinations and not post-test segmentation
It takes eCommerce teams considerable time to build and maintain audience rules in Dynamic Yield. For instance, coding things such as "low intent equals fewer than 12 page views, high intent equals more than 24," then QA-ing those cohorts and redesigning them after each test.
Made With Intent replaces that with continuous intent prediction and delivery. The model decides who sees what, in real time, across hundreds of intent combinations. The actual grunt work is handled by our agent. Your team focuses on strategy, creative, and proof.
Our agent retrains daily, so the experience keeps allocating toward the intent combinations where impact is actually felt. Always on, always learning, always improving on where it started.
3. First-pageview coverage — no fallback needed
Behavioural data takes time to accrue; for anonymous or first-time visitors, Dynamic Yield falls back to geo-based predictive targeting and contextual signals.
Made With Intent has no fallback by design. It doesn't need one. We've trained it across 50bn+ events, in a range of contexts, giving the model day-one predictive power on every visitor, anonymous, identified, first-time, returning.
Made With Intent and Dynamic Yield: In depth
There are lots of areas of cross-over between Made With Intent and Dynamic Yield. The way our technologies work is similar in principle, but different in its practical implementation.
Let's start first with how Dynamic Yield's prediction capabilities work:
Predictions update every 3–5 seconds during a live session; model retrains nightly
Training data
Per-merchant historical patterns
50bn+ events across 150+ retailers
Validation
DY-published — confident but not externally benchmarked
AUC ~0.83–0.84 on conversion/exit/return/add-to-cart heads, calibrated probabilities
While we're talking about AdaptML, we have to say, it's a really sophisticated bit of engineering. It utilises recurrent neural networks and NLP models to get smarter and smarter as it consume more data its got on a specific retailer's visitors.
However, it predicts what you'd expect (affinity and relevance), not what's happening in a visitor's decision right now.
Our model is a driven by a single purpose: it reads buying intent. Sure, it's a narrower job, but it's the one that determines whether a visitor converts, hesitates, or leaves.
Light: agentic campaigns reduce manual segmentation work after setup
Agentic campaigns
No agentic layer
Yes — strategy + tactics in, agent allocates dynamically
When you should pick Dynamic Yield
Hey, we're not here to blindly sell you Made With Intent. Sometimes, our tool just isn't right for your business. So, unlike loads of other SaaS vendors, let us tell you when we wouldn't be a good fit:
• If you're trying to do cross-channel, Dynamic Yield is the right choice versus Made With Intent. We're web-first, and for multi-channel personalisation programmes Dynamic Yield is the right option.
• Recommendations: NextML and AffinityML are purpose built models for product and content recommendations. Made With Intent adds an intent decison later but it is not a replacement for NextML or AffinityML.
• Sophisticated email personalisation: Klaviyo-grade dynamic content across email and ad placements out of the box. If email is a key part of what you do, Dynamic Yield is a top choice.
• Security of a big vendor: Dynamic Yield is an eight-time Gartner Magic Quadrant Leader for personalisation engines. If you're enterprise org and need to run a formal RFP, this helps your procurement team make a decision.
• Strong experience builder. Sections, Page Contexts, and Selectors give users very fine control. But only after you've climbed the initial learning curve.
Where Made With Intent wins
Of course, this is an article designed to help you pick between Dynamic Yield and Made With Intent. From our table above, it's clear there's plenty of areas of overlap, but there's some stuff we do that, we don't mind saying, makes us a better choice. Have a read:
• Acts on the anonymous majority: Every visitor gets a multi-dimensional intent read from the first pageview. No prior data, no behavioural accrual, no geo fallback.
• Multi-dimensional intent prediction: The way we predict intent is comprehensive. We use six measures: stage, signals, trends, purchase confidence, abandon risk and shopper mindset. These are all updated every three-five seconds, unlike Dynamic Yield.
• Autonomously runs campaigns and makes decisions under directives:Agentic Campaigns can dynamically allocate experiences across hundreds of intent combinations during a campaign you'll run. Dynamic Yields's models are sophisticated. The targeting layer above them is still rule-driven. That's simply how the tool is built. But it means what experiences are served is decided before the campaign runs, and not during it.
• Cross-merchant model: 50 billion+ events across 150+ retailers. Day-one predictive power for our customers.
• Start within 30 minutes: Single 7kb GTM tag, no PII and we're ISO 27001 accredited.
So, it's time to make a decision: Let's pick one. Or both?
It's crunch time. We've given you all the facts, but it's time to wrap up and make a judgment call.
Pick Dynamic Yield if the you want consistent cross-channel personalisation — content and recommendations spanning web, mobile app, email, and kiosk. And if you have the technical resources to operate it.
Dynamic Yield's breadth and recommendation algorithm depth is class-leading, and trying to replicate the scale of that capability with Made With intent, plus some integrations would be a worse outcome for you.
Choose Made With Intent if you want a on-site decisioning, intent-driven experiences, and proven incrementality. Our tool delivers experiences natively, allocates them dynamically across hundreds of intent combinations, and proves causal lift on every one.
Now, something to think about. If you're an enterprise customer, you actually could consider both. And here's why:
• Dynamic Yield gives you the breadth and means of channels and mediums to serve personalised content to
• Made With Intent helps you accurately decide who, and when that content should be served
That's brought us to the end of the comparison blog post. If you're still unsure of the differences between Dynamic Yield and Made With Intent, the best thing for you to do is talk with one of our team.
Disclaimer: This comparison is based on publicly available information from Dynamic Yield's documentation, marketing site, and customer reviews as of April 2026. Both products evolve continuously. If anything looks out of date, get in touch and we'll sort it.
August 10, 2026
Become an Intent Insider
Get subscriber-only insights straight to your inbox. No spam. No inappropriateness.
You're in. Welcome. Expect an insider-only email soon.
Oops! Something went wrong while submitting the form.
By submitting this you agree to our (more than fair) terms.
This site uses essential cookies to run properly and optional cookies to improve your experience. Optional cookies only run if you accept them. Privacy Policy here.