Unlock the Future with AI Innovations

Feedback Loops: Why Your AI Investments Are Either Compounding or Collapsing

Every AI investment your organisation makes is on one of two paths. Either it is compounding, gathering data and capability in a self-reinforcing spiral that grows harder for competitors to…

Every AI investment your organisation makes is on one of two paths. Either it is compounding, gathering data and capability in a self-reinforcing spiral that grows harder for competitors to match. Or it is collapsing, oscillating between overreaction and correction, burning capital without converging on anything useful. The difference between these outcomes is not the quality of the model, the size of the dataset, or the talent of the engineering team. It is whether the people running the investment understand feedback loops, and specifically whether they understand the one variable that decides whether a loop compounds or destroys: delay.

This is not a new problem. It was worked out seventy years ago by people thinking about radar guidance and factory inventory, not neural networks. Almost every organisation now deploying AI at scale has ignored their insights, and the costs of that ignorance are compounding faster than the technology itself.

The Reinforcing Loop: Where Compounding Lives

Norbert Wiener coined the term cybernetics in 1948, borrowing from the Greek kybernetes, meaning steersman. His subject was control through feedback: the principle that a system steers itself by measuring the gap between where it is and where it wants to be, then adjusting. Wiener’s specific interest was negative feedback, the self-correcting loops that keep a system stable. But the framework he built applies just as well to positive, or reinforcing, loops. The ones that amplify rather than correct.

A reinforcing feedback loop is straightforward: the output of a process feeds back as input, and the system grows. More users generate more data, which trains better models, which attract more users. This is the data flywheel, and it is the structural reason that Google’s search quality, Amazon’s recommendation accuracy, and Meta’s ad targeting have proven so difficult to challenge. Each iteration of the loop makes the next one cheaper relative to the value it produces. The system grows, and it grows more efficiently as it grows.

Reinforcement learning from human feedback, the technique behind most production large language models, is a reinforcing loop compressed into a training pipeline. Models generate outputs. Human evaluators rank them. A reward model encodes those judgements. The next model iteration is optimised against that reward. The loop closes, and each cycle produces a measurably better system. The speed of this loop, the tightness between action and consequence, is what makes RLHF effective. It is also what makes it dangerous. Tight loops with no balancing mechanism do not converge. They run away.

The Balancing Loop: Why Nothing Grows Forever

Every reinforcing loop eventually meets a constraint. Donella Meadows, in her posthumous 2008 primer Thinking in Systems, called these balancing loops. A balancing loop is the feedback structure that resists change: it pulls a system back toward a target, a limit, or a set point. A thermostat is a balancing loop. So is regulatory enforcement. So is the finite supply of human attention, which caps how many platforms any individual can meaningfully engage with.

Balancing loops are not obstacles to be removed. They are what keep reinforcing loops from consuming everything. Without them, compounding systems overshoot their resource base and collapse. Meadows spent her career demonstrating this, most famously in The Limits to Growth (1972), the system dynamics model that showed how exponential industrial growth, colliding with finite resources and pollution sinks, produces not steady expansion but overshoot and decline.

The problem for organisations investing in AI is that their reinforcing loops are running much faster than their balancing loops. Data infrastructure scales overnight. Governance, regulation, and ethical review scale on the timescale of committees, consultation periods, and quarterly board meetings. The gap between the speed of the reinforcing loop and the speed of the balancing loop is where damage accumulates. Content moderation systems optimise engagement in milliseconds; the societal consequences of that optimisation surface in election cycles. The mismatch is structural, not accidental.

Delay: The Variable That Kills

The most dangerous property of any feedback loop is delay, the time gap between an action and the signal that tells the system whether the action was correct. Jay Forrester identified this in 1961 in Industrial Dynamics, the founding text of system dynamics. Working with managers at General Electric, Forrester discovered that employment instability at GE’s appliance plants followed a roughly three-year cycle. The instinctive explanation was the business cycle, external demand fluctuation. Forrester’s models showed otherwise: the oscillation was generated entirely by the internal structure of the company’s own hiring and inventory decisions, combined with delays in information flow. The system was oscillating because of its own feedback dynamics, not because the environment was unstable.

Forrester called this phenomenon the bullwhip effect. A small change in customer demand, say five percent at the retail level, propagates upstream through a supply chain and amplifies at each stage. By the time it reaches the manufacturer, that five percent shift has become a forty percent swing in perceived demand. Each actor in the chain is behaving rationally given the information available, but the delay between order and delivery causes each to overcorrect, and their overcorrections compound.

John Sterman operationalised this at MIT with the Beer Distribution Game, a simulation where participants manage a supply chain with deliberately built-in delays. The results are consistent and damning. MBA students, experienced managers, supply chain professionals, all produce massive oscillation. They over-order when they see shortages, then drown in excess inventory when their delayed orders finally arrive. Sterman’s 1989 analysis showed that the failure was not irrationality but systematic misperception of feedback: participants consistently underestimated how long their actions would take to produce effects, and consistently attributed the resulting oscillation to external forces rather than their own decisions.

This is the exact failure mode now playing out in AI investments across every industry.

How AI Accelerates the Loop and the Blindness

Artificial intelligence compresses operational feedback loops to near-zero latency. Recommendation engines update in real time. Algorithmic trading executes in microseconds. Inventory systems adjust automatically to demand signals. This feels like progress, and at the operational level, it is. The problem is that strategic consequences operate on a completely different timescale. A recommendation algorithm that optimises for engagement today changes user behaviour over weeks. Market position shifts over quarters. Regulatory exposure builds over years. The operational loop runs in seconds. The strategic loop runs in fiscal periods. The delay between them is invisible because the operational loop feels so fast and responsive.

The COVID-19 supply chain collapse was a live demonstration of AI-accelerated bullwhip dynamics. Algorithmic inventory systems detected demand spikes and automatically increased orders. Suppliers, reading those orders through their own algorithmic lenses, amplified the signal further. By the time the overcorrection became visible, manufacturers had committed to production runs that would arrive as excess inventory months later, precisely when demand had normalised. The bullwhip cracked, and the cost was measured in billions in stranded inventory. The technology did not fail. It executed its feedback loops faithfully. What it lacked was any mechanism for accounting for the delay between its own responsiveness and the slower-moving consequences of that responsiveness.

The Leader’s Real Job

Stafford Beer, in his 1972 Brain of the Firm, formalised this into the Viable System Model, a cybernetic description of what any organisation needs to remain viable. His fourth Principle of Organization was explicit: the operation of the system’s feedback channels “must be cyclically maintained without delays.” Beer understood that delays are not inconveniences. They are pathologies. A system that cannot close its feedback loops quickly enough to respond to environmental change is not viable, no matter how sophisticated its individual components.

The uncomfortable truth for organisations investing in AI is this: most are not managing feedback loops. They are being managed by them. They have accelerated their reinforcing loops with machine learning, starved their balancing loops of resources and authority, and remained blind to the delays between their operational decisions and their strategic consequences. They are Sterman’s beer game participants with better dashboards and the same fundamental misunderstanding of how systems behave over time.

The leaders who will compound their AI investments are not the ones with the fastest models or the most data. They are the ones who can look at their organisation and answer three questions: Which loops are reinforcing? Which loops are balancing? And where are the delays that will turn our speed into overshoot? Until those questions are asked deliberately, every AI investment is a bet on a loop the organisation does not understand.