In 2000, Netflix offered to sell itself to Blockbuster for $50 million. Blockbuster declined. The decision looked reasonable at the time. Blockbuster had more than 9,000 stores, 84,000 employees, and a business model that worked. Netflix was a loss-making startup mailing DVDs. Why buy the problem child when you own the category?
Blockbuster spent the next decade perfecting what it already did. It refined its inventory systems. It optimised its rental pricing. It squeezed more revenue out of late fees. Every dollar went into the machine that was already printing money. Meanwhile Netflix pulled three new arms: subscription DVD-by-mail, then streaming, then original content production. Each was a gamble with no guaranteed payoff. Each, eventually, displaced the model Blockbuster had been optimising.
Blockbuster filed for bankruptcy in 2010. Netflix is now worth hundreds of billions.
This is the explore-exploit dilemma, and it is the oldest problem in adaptive decision-making. Every system that learns faces it. The question is always the same: do you exploit what you know works, or do you explore what might work better? You cannot do both with the same resources at the same time. Every dollar spent refining a proven process is a dollar not spent discovering a better one.
In 1991, the organisational theorist James March formalised this tension in a paper for Organization Science. Exploration, he wrote, covers search, variation, risk-taking, experimentation, play, discovery, innovation. Exploitation covers refinement, efficiency, selection, implementation, execution. Adaptive systems need both. But the two compete for the same scarce resources.
March’s insight was sharper than the framing. Exploitation is self-reinforcing. When an organisation exploits a successful approach and it works, the organisation gets better at it. Competence builds. Returns are measurable. So the organisation invests more in exploitation, which builds more competence, which generates more measurable returns. The loop tightens. Meanwhile, exploration produces no comparable returns. Experiments fail more often than they succeed. Their payoff is uncertain and delayed. In any head-to-head budget battle, exploration loses. Not because it is less valuable, but because its value is harder to measure on the timeline that budgets run on.
Then March’s model turns uncomfortable. He found that faster learning makes the trap worse. Organisations that learn rapidly from experience lock in to a narrower band of competence faster than organisations that learn slowly. Speed feels like an advantage. It is, for exploitation. But it accelerates the narrowing of what the organisation knows how to do. The better you get at what works, the faster you stop looking for what might work better.
AI as an exploitation amplifier
This is the problem with how most organisations deploy AI. Recommendation engines are exploitation machines. They take what you already know about a customer and serve more of it. Predictive analytics takes historical patterns and extends them. Process optimisation tools take existing workflows and squeeze out inefficiency. Every one of these applications refines the known. None of them discover the unknown.
The illusion is that AI feels like exploration. Companies run thousands of A/B tests and call it experimentation. But A/B testing within an existing product is not exploration. It is fine-grained exploitation. You are testing variations of something you already do, not testing whether you should be doing something fundamentally different. The recommendation engine that serves you more of what you already like is narrowing the range of outcomes, not expanding it. It optimises the current arm of the bandit. It never asks whether a different arm exists.
March’s model predicts exactly what we see. AI accelerates learning from existing data. That is the definition of exploitation. The faster an organisation learns from its data, the faster it locks in to its current competence. AI does not broaden the organisation’s adaptive range. It sharpens the one it already has. The dashboard shows improvement. Revenue ticks up. Costs tick down. Everyone feels progress.
What is actually happening is a narrowing.
The horizon problem
The multi-armed bandit problem, formalised by Herbert Robbins in 1952, gives this tension a mathematical structure. Imagine a row of slot machines. Each has a different payout rate, and you do not know any of them in advance. Every pull teaches you something about that machine’s average payout. Every pull of the best machine you have found so far is a pull not spent learning about the others. The optimal strategy depends on one thing above all: how many more pulls you have left.
If you have a thousand pulls ahead of you, exploration is worth a lot. The knowledge you gain compounds across hundreds of future decisions. If you have three pulls left, exploit. Squeeze what you know. This is the horizon problem, and it is where most organisations get their AI strategy backwards.
When a market is stable and your current model has a long life ahead of it, exploitation is rational. Optimise. Refine. Squeeze. But when disruption is coming, the horizon of your current model is shrinking. The arm you have been pulling is about to stop paying. That is precisely when you need to have been exploring, because exploration takes time to produce usable knowledge. You need new arms before the old one runs dry.
And here is the catch. Disruption feels like the worst time to experiment. Budgets tighten. Boards demand proof of return. Risk tolerance collapses. The organisation doubles down on exploitation exactly when the math says it should be exploring. This is not a failure of intelligence. It is the exploitation feedback loop doing what March described, accelerated by AI-driven measurement that makes exploration look even worse by comparison.
The epsilon strategy
The bandit literature offers a way out. It is less elegant than perfection but more durable. It is called epsilon-decreasing exploration. You spend the vast majority of your effort exploiting what works. But you set aside a fraction of your resources for genuine exploration. As you learn what works, you shrink that fraction. Exploration gets rarer. But it never hits zero.
The organisational equivalent is what Charles O’Reilly and Michael Tushman called ambidexterity. In a 2004 Harvard Business Review paper, “The Ambidextrous Organization,” they argued that the organisations which survive disruption are the ones that can exploit their current business and explore new ones at the same time. Not by asking the same team to do both. By building separate structures, each with its own culture, its own metrics, its own timeline. The exploitation engine keeps the lights on. The exploration engine keeps the future open.
Applied to AI, this means resisting the pull to spend every AI dollar on optimisation. It means allocating a deliberate fraction of AI investment to projects that have no immediate return. Generative AI that prototypes business models you might enter in three years. Simulation tools that test markets you do not currently serve. Internal platforms that let employees experiment with capabilities unrelated to their current jobs. None of these will show up on next quarter’s dashboard.
That is the point. Exploration that shows up on next quarter’s dashboard is not exploration. It is exploitation wearing a disguise.
The discipline is in the word “decreasing.” Early in an organisation’s AI journey, when nobody knows what works, the exploration fraction should be high. Most effort should go to discovery. As patterns emerge and proven applications surface, shift toward exploitation. But fix a floor. Some fraction of AI investment stays in the exploration budget permanently. Not because it produces returns this year. Because the arm you are optimising today is not guaranteed to pay out forever.
The uncomfortable trade-off
Blockbuster did not fail because its leaders were stupid. They failed because they were rational inside a system that rewards exploitation until the moment it doesn’t. Late fees were profitable. Store optimisation was measurable. The known model had higher returns than the experimental one. Every incentive pointed toward refining what worked.
The incentives were not wrong. They were local. And local optimisation is the definition of exploitation.
The uncomfortable truth about exploration is that it looks like waste right up until the exploited option stops working. The experiments that fail cost real money. The capability you build for a market you might never enter looks like a distraction. The team exploring adjacent possibilities looks like it isn’t pulling its weight. None of this registers as value on any dashboard that measures the current quarter.
But the bandit math is unforgiving. If you never explore, you are betting that the arm you are pulling now will pay out forever. In a market being reshaped by AI, that is a bet against the evidence. The organisations that last are not the ones that optimise hardest. They are the ones that keep one hand on a different machine, pulling it often enough to know what it pays, even when the current machine is still printing money.
