Why contextual bandits are the future of personalisation

This is a good read. David cobbled together some context on contextual bandits. See what I did there?
Read Time:
On this page

I've spent the last 10 to 15 years thinking about why personalisation never quite worked. I even wrote a book about it (available at all good bookstores for £11.99, he says, shamelessly).

The short version is the arithmetic failed us.

To personalise anything, somebody had to decide who gets what. Rules. For most of my career that somebody was a person. People wrote the decision down as a rule, maintained it for a while, then forgot about it. Which is fine at three rules, maybe even ten, but at three hundred it falls over. Ecommerce sites have far more moments than any team could ever write rules for, don’t they? So, most personalisation programmes ended up as a handful of hand-built segments, slowly going stale in a tool nobody logged into anymore.

That’s segmentation, not necessarily personalisation. And there's a difference between segmentation and personalisation. You can be personal within a segment. You can't do it the other way around.

A contextual bandit takes the person out of the decision loop (which visitor sees what, and when they see it). People still own the strategy and the creative, sure, but the machine makes the per-visitor call over and over again (and keeps on making it). It's what makes personalisation truly scalable, and therefore finally achievable.

Ironically, a contextual bandit puts the person in personalisation, by taking the person out of personalisation. Kinda unexpected, right?

Recommendations are the proof

If you ask 100 people in a room which company does personalisation really well I bet most would say “Netflix” or “Spotify” or “Amazon”.

But do they really do personalisation? Or is it recommendations?

Have you ever sat down and wondered why recommendations were the only personalisation brands ever made work?

Nobody was writing the rules. They are scalable, defined by an algorithm not by humans.

A model decided what to show, per person, continuously, and got better as it went, so it scaled. Everything else in personalisation stayed manual and stayed small.

Contextual bandits take the thing that made recommendations work (a machine making the per-person call, on repeat) and point it at messaging, reassurance, urgency, incentives, timing, and the decision to leave someone alone. That’s what we do at Made with Intent.

So, I expect this to become the default now that the arithmetic finally works.

From A/B tests to bandits, briefly

How traffic is splitWhat you getWhat it still can't do
A/B testFixed and even (50/50).One trustworthy answer about all your users.You get one answer for everybody, hiding pockets of losers within the winners.
Multi-armed banditMoves towards whichever option is winning while the campaign runs.The same answer, sooner, with less traffic wasted getting there.It's still hunting a single winner (a variant) for everybody.
Contextual banditMoves towards whichever option is winning for this kind of visitor.Different answers for different people, decided as it goes.It can only find differences that exist inside the context you give it.

So a plain bandit asks which variation wins. Where a contextual bandit asks which variation wins for whom. From there, it keeps on asking it because the answer moves. Be that by hour, by day, by week, or by month (your results from an A/B test, for example, will differ from March to November, especially in eCommerce). Imagine a world where the contextual bandit is always on; always learning, continually reallocating traffic to where the experience is best seen and felt.

The mechanics are actually a lot simpler than the name suggests: it tries things (that's exploration), notices what worked for which kind of visitor and sends more of those visitors to the option that suits them (that’s exploitation). It never stops doing a bit of both, and that's how it keeps up when behaviour shifts (it does, constantly).

There's no winner to declare and roll out because the allocation is the output.

Everyone has a bandit now, so look at what they're feeding it

The maths behind contextual bandits is better understood in this day and age. What differs (enormously) is what the model is allowed to know about the person. The context.

Most implementations use attributes: device, location, traffic source, new or returning, and in B2B tools things like industry, company size or revenue band. It's data you already hold, it's stable and it's easy to hand to a model, so of course everyone starts there.

Those are attributes of the website, and they tell you very little about the human doing the visiting. Ironic when we’re talking about personal-isation, not websit-isation.

The alternative is Intent, meaning what someone is doing right now: whether they're comparing or committing, confident or hesitating, picking up momentum or about to leave. None of that lives in your database. It has to be inferred from behaviour, continuously, while the visit is still happening.

I always come back to the shop assistant. A good one doesn't decide how to help you based on which door you came through or what car you parked outside, they're reading you. Browsing or hunting? Stuck? About to give up and walk out? The door you came in by is context too (it just doesn't change for the whole visit, while everything the assistant is reading does).

That context is so important for what ends up being the outcome.

A model can only find differences its context can describe. If the real reason someone abandoned was hesitation, and all your model knows is device and traffic source, it'll never find hesitation. It'll find whatever correlation happens to sit closest (mobile converts less, say) and optimise very confidently against a proxy, even though device doesn't cause hesitation at all.

On top of that, attribute context is usually picked in advance by a person. You choose the fields at setup from a list somebody maintains, so you're guessing which distinctions matter before you've got any evidence (the same guess that sank rule-based personalisation in the first place). The bandit has automated the allocation and left the guess exactly where it was (I find that quite funny, in a painful sort of way).

Inferred context gets round this because nobody writes the list. Our model reads around 800 signals per interaction, updated every three to five seconds, and works out which states matter before the bandit allocates across them.

So there are really two jobs here, perceiving and deciding, and most contextual bandits only do the deciding. (Sure I'm biased. I also think I'm right.)

What a bandit won't do for you

A contextual bandit doesn't replace experimentation. See my other article on why. It shouldn’t. It won't give you statistical significance (it isn't trying to). Traffic is deliberately unequal and that breaks the maths a test relies on, so if you need a defensible causal claim about your whole population, for a roadmap decision or a board paper, run a properly powered test.

It'll optimise precisely what you point it at, so aim it at a click near the experience and you'll get lots of those clicks, whether they were worth anything or not (choose the metric carefully, it'll take you literally).

Then there's exploration, which never stops and costs a little, because a small share of traffic is always trying alternatives. That's the price of not going stale and I reckon it's a good trade (it isn't free, mind).

And it can't rescue a broken experience. Allocation decides which of your things somebody sees, so if the thing itself is slow or confusing, a better-aimed version of it is still slow and confusing. Fix that first, or you're moving deck chairs on the Titanic with a very clever algorithm.

The evolution of experimentation

You're going to see a lot more bandits either way (there's more data about, and far more people who know how to build a model than there were ten years ago).

A contextual bandit still has to learn, so it starts out knowing nothing and needs enough of the right visitors through it before it gets any good. It asks for the same discipline as an A/B test (just less of it), and that discipline is exactly what we're trying to get away from.

At some point people are going to say "I don't want to wait for the data. I don't have it." And even when they do have it, it only goes so far. What if I want to run something for two days, or on a page that hardly anyone visits?

That's why I think of this as the evolution of experimentation, with each step asking you to wait a little less for the data than the one before. And the next step looks like it asks for almost none. Language already has its general purpose models (your ChatGPTs and your Claudes) and decision models are starting to turn up that work in a similar way. Jev, TypeSafe's new model, is a good example.

You give it a state (a description of what's going on) and a question, and it hands back a probability for each possible answer that software can act on. It's pre-trained, so it has a view before you've given it any history at all. On its own it could act as a pre-trained contextual bandit (an unvalidated one, mind).

We're already looking at how something like this gets our agentic campaigns past the cold start, because nobody gets to train a model on Black Friday before Black Friday, do they?

If you're interested in learning more about contextual bandits, and well, basically, Made With Intent (because that's literally what we do) book a demo with our team.

// the intent insider

Become an Intent Insider

Get subscriber-only insights we don't publish anywhere else and event invites before anyone else.

You're in. Welcome. Expect an insider-only email soon.
Oops. Looks like Something went wrong. Try again?
 No spam  No inappropriateness  Unsubscribe anytime

By submitting this form you agree to our (more than fair) terms.

See how brands are finding new growth with intent
Closing the Intent Gap playbook: cover and inside pages about using intent data for ecommerce growth
By submitting this you agree to our (more than fair) terms
Check your email. The playbook is on its way.
Oops! Something went wrong while submitting the form.
This site uses essential cookies to run properly and optional cookies to improve your experience. Optional cookies only run if you accept them. Privacy Policy here.