Attention is the last paywall – (Part III/III)
When anything can be written instantly, the scarce thing is the person reading it
This essay is part III of a trilogy of related concepts in languages, information theory, geometry, metaphysics, relativity, and the information economy.
Part I — Cavemen Created the First Large Language Model — was published on Sunday 9 August, and Part II — The Shape of Existence — on Wednesday 12 August.
This final part asks what becomes scarce when information itself becomes almost free.
The firehose has learned how to write
The modern world is an unrelenting firehose of information directed straight at your face. It never stops. Devices buzz. Screens blare. Noise intrudes. Instant messages demand. Podcasts accumulate faster than they can be heard. Red notification badges sit there like tiny administrative haemorrhoids, making us squirm in our seat.
Broadcasting information has never been so cheap. Receiving and processing information, on the other hand, is becoming very expensive.
Herbert Simon saw this coming in 1971. What information consumes, he pointed out, is the attention of its recipients, so a wealth of information produces a poverty of attention. He was describing an organisation with one mainframe in the basement. The mechanism has not changed. Only the volume has.
Civilisation has spent thousands of years progressively reducing the cost of creating, reproducing, compressing and transmitting information.
Speech allowed one human brain to transmit part of its internal model to another. Writing separated information from the person speaking it. Printing reduced the cost of copying. Telecommunications reduced the cost of distance. Broadcasting reduced the cost of reaching millions of people at once. The internet placed a printing press, television studio, filing cabinet and global distribution network on billions of desks.
Then social media made information transmission so cheap that people began broadcasting everything from their breakfast to their bum cheeks.
Generative artificial intelligence now attacks the remaining expensive part: creation itself. Producing a passable paragraph once required a human being to possess some combination of knowledge, motivation and time. Producing a thousand passable paragraphs now requires a text box and a sufficiently assertive relationship with the Return key.
This is a remarkable expansion of productive capacity. It is also an enormous increase in the potential supply of things demanding to be understood. Even if what demands to be understood is total, abject shite.
Every previous information technology reduced friction somewhere between the mind of the sender and the mind of the recipient. Generative AI removes so much friction from the sender’s side that the imbalance becomes impossible to ignore.
A person can now produce in seconds more text than another person can responsibly assess in an afternoon.
Motivated by outrage
Nor does the expanding supply of information necessarily leave its recipients better informed. Information has no inherent value simply because it reduces uncertainty. Knowing precisely which brand of baked beans a cabinet minister prefers is information. It is also knowledge you might reasonably have hoped to die without.
Shannon set that question aside on purpose, and was right to. Whether a message is worth having was never an engineering problem. It is ours.
The internet promised to make political life more transparent. It has certainly made more of it visible. We now receive a continuous video feed of parliamentary fragments, doorstep confrontations, podcast monologues, facial expressions and remarks recorded before their owner had fully decided whether to think them.
Political figures once appeared intermittently through long form interviews and thought-through opinion editorials. They now leak into public consciousness throughout the day, providing an apparently inexhaustible supply of content, perhaps without a corresponding supply of meaning.
This can create the impression of extraordinary political knowledge. We know which politician rolled their eyes, fluffed a sentence, looked uncomfortable beside a flag or produced the day’s most algorithmically efficient expression of outrage.
We may know far less about the administrative capacity of the state, the incentives shaping a policy or whether the proposed solution could survive three minutes of contact with reality.
The political information stream is therefore capable of increasing both information and ignorance at once. It supplies more observable data while directing attention towards the signals best adapted for transmission rather than those most useful for understanding. Social media also biases the new in news, the immediacy of the moment, regardless of whether it contains signal for long term value.
This is measurable. Markus Prior found that as media choice expanded, political knowledge and turnout did not rise. They diverged. People who preferred news learned more, and people who preferred entertainment, now able to avoid news altogether, learned less. More available information widened the gap instead of closing it.
Social platforms perform a kind of natural selection upon political communication. Statements that provoke immediate recognition, anger or tribal affirmation travel further than qualified explanations whose meaning depends on several paragraphs and a chart.
Research finds that social feedback can reinforce expressions of moral outrage, teaching users which emotional displays are rewarded within their networks. A large cross-platform study published in Science found that misinformation associated with outrage received more engagement, and that outrage increased people’s willingness to share — with outrage-evoking links more often shared without being read first.

More recent field research has also shown that reducing the prominence of content expressing partisan animosity and antidemocratic attitudes can alter what users subsequently encounter and report feeling. The distribution system is not a neutral pipe. Ranking choices influence the emotional and political composition of the information stream.
This does not invariably reward the stupid, but it creates unusually favourable atmospheric conditions for them. The result is the propulsion of imbeciles to improbable heights in political discourse.
Rupert Lowe, for example, has achieved a level of social-media reach that considerably more established politicians might envy. The Financial Times reported in May 2026 that ten of Lowe’s posts had reached at least ten million views since the launch of Restore Britain, while none of Nigel Farage’s had crossed the same threshold over that period.
Lowe had a third of Farage’s followers, but received repeated promotion from Elon Musk — who shared at least eight posts naming the party, all but one of which passed ten million views — and produced material unusually well adapted to the platform’s digestive system. Whether all of that reach was human is another question: the FT noted it is unclear how much of it represents genuine engagement.
The Financial Times has separately reported that Restore’s online ecosystem has been assisted by commercial misinformation networks using stolen and AI-generated content — 43 Facebook pages traced to a single network in Bangladesh and Pakistan, monetised through clicks and advertising.
One such page, having gathered a following as a Lowe support group, then reinvented itself as a vendor of television-streaming devices, to the considerable annoyance of its existing members. Virality is clearly evidence of communicative fitness. Whether it is valuable or contributes positively to society is another question entirely.
Lowe might be nasty piece of work, but the point is not solely ideological. The same system can propel nonsense from almost any political direction, provided it is immediate, emotionally legible and easy to transmit. It is a feature of an information economy that assigns value through attention and emotion.
Nuance is expensive to create, demanding to decompress and difficult to circulate. Vagary, accusation and emotional certainty arrive pre-compressed.
The engagement gauge does not distinguish drinking water from effluent. It simply records engagement with the pipeline.
The nervous system has no progress bar
Nobody has ever handed you a model of the world. What arrives, every time, is compressed instructions for building one, and the building is done by you, at your own expense, out of parts you already happen to own. That was always the arrangement, and it was bearable for exactly as long as producing the instructions was expensive.
The recipient is not completely defenceless. Search engines compress discovery. Specialism reduces the need to learn generalist skills. Recommendation systems compress choice. Language models compress documents, meetings and email chains.
Apple Intelligence summarises and prioritises notifications, allowing users to scan what the system regards as the important details without opening every message. Apple’s technical account of its foundation models explicitly identifies notification prioritisation and summarisation as intended uses.

The receiver no longer has to read every original word one at a time. Technology is increasingly capable of compressing the expanded.
But this does not remove the bottleneck so much as move it. A human must still decide whether the summary is accurate, whether omitted information matters, whether the source is trustworthy and whether the compressed result requires action.
If the subject is consequential, compression may create a second task: read the summary, become suspicious of it, then open the original anyway.
Apple’s notification summaries demonstrated this problem in miniature when misleading summaries of news alerts prompted the company to suspend the feature for news and entertainment applications while changes were made. Summaries had successfully reduced the character count, without carrying forward sufficient fidelity.

Nor is shorter communication necessarily cheaper to process. A clear hundred-word explanation may be easier to understand than an ambiguous ten-word instruction.
An ill-thought through wall of text on Slack optimises for the speed of asking, not understanding or answering. Five recipients might now need to establish what is actually wanted, whether it is wanted of them, whether somebody has already done it, and whether ‘we should probably’ includes the sender.
The sender might have saved five minutes of thinking by sending a high entropy message. But message recipients could now lose five minutes each trying to reconstruct original intentions. An exemplary little productivity gain.
The thinking cost never disappeared. It has been transferred from the producer to the recipient, and multiplied.
And the recipient’s capacity is seemingly finite. Working memory retains only a small amount of information in a readily accessible form, allowing it to support comprehension, reasoning and planning. The precise capacity depends on how information is grouped and what the person already knows, but the existence of a severe limit is well established.
Psychologist Nelson Cowan proposed that the focus of attention often holds around four meaningful chunks, rather than the roughly seven chunks associated with older estimates. The relevant unit is not a word or pixel but a chunk whose size depends on the structures already available to the recipient.
Existing knowledge therefore expands effective capacity by grouping information into larger, meaningful units. An expert can glance at a chart, equation or paragraph and recognise a structure that a novice must reconstruct piece by piece.
Chess masters do it with a board, and the advantage largely disappears when the pieces are scattered at random: what they are recalling is structure, not squares. Compression allows more information to pass through this limited aperture. It does not make the aperture infinite.
An entire discipline is built on that aperture. Cognitive load theory treats working memory as the binding constraint on learning, and treats the design of the material — what to leave out, what to group, what order to reveal it in — as the way around the constraint. Its useful finding is that the load is not a property of the subject. It is also a property of the presentation.
Every input must still compete for attention and, if it is to become useful, be integrated into some model of the world. A summary, notification or recommendation may lower the immediate processing cost, but the meaning it represents can still entail uncertainty, conflict, emotional consequences and further decisions.

There may be no clean subjective alarm as we approach our processing limits. The human nervous system does not display a pleasant progress bar announcing: ‘You have processed 93% of today’s maximum viable information. Consider becoming briefly unavailable.’
The signs may arrive indirectly: reduced concentration, fatigue, irritability, avoidance, compulsive checking, diffuse anxiety or an increasingly powerful desire to throw Slack into the sea.
A comprehensive review of information overload describes its association with stress, impaired decision-making, reduced performance and avoidance when informational demands exceed the recipient’s available processing resources. Research into ‘technostress’ has similarly found that digital demands can produce cognitive overload and greater distraction from goal-relevant information.
This is where uncertainty becomes psychologically important. In Claude Shannon’s equation, uncertainty is a neutral mathematical quantity. In human experience, unresolved uncertainty frequently induces anxiety, and motivates a demand for attention.
Who messaged? What happened? Does it concern me? Have I forgotten something? Is the apparently urgent email genuinely urgent, or has somebody marked a routine request urgent because they would personally enjoy an answer?
A notification is a particularly efficient uncertainty generator. It tells you that some information exists while often withholding enough of it to prevent completion.
A vibration in your pocket is not always a harbinger of meaningful, important information. But it creates a question.

A message preview provides more information, but may create three further questions. A machine-generated summary can answer those questions, unless it answers them incorrectly, in which case it has compressed the original uncertainty into a smaller but more concentrated form.
Research on attention and incomplete tasks does not justify a simple equation between uncertainty and anxiety, but it supports the broader idea that unresolved goals and interrupted activity can remain cognitively active. Bluma Zeigarnik’s original experiments in 1927 examined the greater accessibility of interrupted tasks, and later work found that unfinished goals intrude on unrelated activity — until the person makes a specific plan for them, at which point the intrusions stop.

The smartphone is therefore not merely an information-delivery device. It is a machine for creating unfinished acts of interpretation.
Perhaps this helps explain why information overload may not be experienced simply as the conscious thought I know too many things. It may instead be felt as agitation: a background sensation that something remains unresolved, some input has not been processed, some task is waiting just beyond recall.
The mind becomes surrounded by partially opened ZIP files, each requesting extraction, evaluation and possible action.
An iPhone can compress notifications. It cannot decide which relationships you care about, which obligations are morally important, or whether ‘Can we talk?’ is a routine administrative request or the beginning of the end of a relationship.
The other meaning of decompression
The word decompression already contains both sides of the problem. In computing, it describes the reconstruction of compressed information. In human life, it means releasing pressure: moving from a state of cognitive or emotional demand into one in which experience can be absorbed and organised. This may be more than a convenient linguistic coincidence.
Human decompression is partly the processing of information after its arrival. A quiet walk, an idle train journey or an evening without a stream of new inputs allows recent experiences to be revisited, connected and incorporated into longer-term models. What feels like doing nothing may be the mind completing work that incoming information repeatedly interrupts.
Evidence here should be stated carefully. Rest is not guaranteed to perform a magical nightly defragmentation of the personality. But the effect is real: a 2012 experiment found that ten minutes of quiet rest after hearing a story improved recall of it seven days later, and two recent meta-analyses pooling several hundred comparisons between them put the benefit of waking rest at around a third of a standard deviation — still detectable a week afterwards.
Neural research has also found replay of recently learned sequences during wakeful rest — at roughly twenty times the speed of the original behaviour — supporting the idea that apparently idle periods can contain active reorganisation of recent experience.
What the evidence does not support is the tidier version of the story. The benefit is largest in amnesia patients, moderate in older adults and smallest in healthy young ones, and it does not obviously depend on keeping the interval empty.
A 2025 experiment comparing waking rest with social-media use after learning found that rest helped — and that scrolling helped too, with no measurable difference between them. The mind does appear to need offline time to file the day away. That the offline time must be spent staring out of a window is not established.

This also helps explain the value of serendipity in a world designed around informational efficiency. Compression exploits patterns that are already known. Exploration encounters patterns for which no convenient code yet exists.
Recommendation systems, personalised feeds, search engines and AI assistants are excellent at guiding us rapidly through familiar conceptual territory. They infer what we are likely to want and shorten the route towards it.
But a life composed entirely of optimised routes risks becoming informationally narrow. The unfamiliar book, accidental conversation or diversion can present as a risk more readily than an opportunity. This even as the familiar offers only marginal gains, while the transformational upside gains lie in exploration.
Research into recommendation systems has repeatedly raised concerns about the tension between accuracy and discovery: optimising only for immediate relevance can reduce diversity, novelty and serendipity.
Serendipity is difficult to optimise because the value lies partly in encountering something the existing model did not know to request. Efficiency uses the decompressor you already have. Serendipity improves it.
A person exposed to more varied ideas, places and people accumulates a larger library of patterns through which future information can be understood. Exploration is not necessarily a failure of optimisation. It can be an investment in future enjoyment or domain compression.
This is also why the answer to overload cannot simply be better summaries. A summary helps us move more quickly through information somebody has already decided is relevant. It does not determine whether the information should have occupied our attention in the first place, nor does it create the silence in which its implications can be integrated.
We increasingly possess tools for compressing almost everything except the time human beings need to understand it.
Knowledge flows downhill
It is worth knowing that this problem has a solvable version, because we solve it daily for machines and rarely for each other.
When you invoke a metaphor in an LLM prompt, you reshape the representational landscape. ‘Explain X as if it were Y’ activates features and relationships associated with both concepts, encouraging the model to align the source structure of Y with the target structure of X.
The model does not open a cabinet, retrieve a neatly laminated metaphor card and begin photocopying. It constructs a context in which some continuations become more probable than others.
The architecture enabling this contextual transformation was introduced in Attention Is All You Need. Through repeated attention and feed-forward operations, token representations are updated in relation to other tokens before the model produces a distribution over possible continuations. The metaphor in a prompt becomes part of the context from which those probabilities are computed.
It may be useful to imagine generation as movement across a terrain of slopes, mountains, hills, valleys and flatlands. The prompt reshapes the landscape, creating shorter or steeper routes towards particular forms of explanation while leaving others uphill, distant or inaccessible. The model then produces one token at a time according to a probability landscape that is itself continually changed by the tokens already produced.
This is not literally a marble rolling towards one global minimum. Language generation involves a succession of probability distributions, sampling decisions and contextual states.
Nor has mechanistic interpretability reduced the operation of metaphor prompting to one known projection matrix. The landscape is a model of the conceptual effect, not a wiring diagram.
But it captures something important: a good metaphor does not dictate the words the model must produce. It changes the gradient of the terrain.
‘Explain transformer attention like a library index-card system’ gives the model a useful slope. It activates a dense collection of compatible relationships: books as stored information, catalogue cards as indexes, queries as searches and retrieved cards as relevant context. You have not specified the answer, but you have made many bad answers inconveniently uphill.
‘Explain quantum computing as a hen party’ activates superposition, entanglement, uncertain states and an observer whose intervention changes the outcome. Unfortunately, it also activates sashes, penis straws and somebody being sick outside Turtle Bay after too many espresso martinis. The model has found a slope, but not necessarily one tilted towards physics.

A poor analogy increases the number of improbable connections the model must construct. A good one supplies a short vector route through a well-mapped region of meaning. You have given it the starting postcode and pointed it downhill.
Attention is the last paywall
Historically, creating and distributing information was expensive. Access was scarce, and institutions such as libraries, universities, newspapers and broadcasters possessed enormous power because they requiried capital, and controlled the machinery through which knowledge travelled.

That scarcity is moving downstream. The difficult task is increasingly not obtaining information, but deciding which information deserves the cost of understanding.
Search engines reduced the cost of finding an answer. Social media reduced the cost of publishing one. Large language models reduce the cost of manufacturing one that sounds plausible. None automatically eliminates the cost of distinguishing a robust model from a stopped clock enjoying one of its two daily triumphs.
Indeed, the better machines become at fluency, the less fluency itself tells us. Clear, grammatical prose was once a useful – if unreliable – signal of effort, education or editorial scrutiny. It can now be produced instantaneously by a WiFi-enabled fridge.
This is a signal collapsing, not a standard slipping. Michael Spence’s account of signalling requires the signal to be expensive, and expensive specifically for whoever lacks the quality it is meant to indicate. Fluency met that condition for as long as writing well was hard.
This makes judgement more valuable, not less. A human recipient must still interpret information, compare it with an existing model, estimate its reliability, decide whether it matters and determine what should change as a result. These tasks can be assisted, delegated and compressed, but somebody eventually has to absorb the consequences of being wrong.
Research on cognitive offloading describes the genuine value of delegating mental operations to external tools, while also emphasising that offloading changes, rather than abolishes the cognitive task: the user must decide what to delegate, when to trust the external representation and how to act on it. A recent review of AI use frames this explicitly as a tension between cognitive offloading and cognitive overload.
This is the defining asymmetry of the information economy. The marginal cost of producing, compressing and transmitting information is collapsing towards zero, while the biological and epistemic cost of making sense of it remains stubbornly human.
Shapiro and Varian’s formula for information goods was costly to produce, cheap to reproduce. Generative models attack the first half. What is left is an economy in which nothing is expensive except being understood.
We have automated the creation of files. We have begun automating their compression. We have not automated the decision of which files deserve extraction, which summaries deserve trust or which ideas deserve incorporation into our model of reality.
Which returns us, finally, to compression itself. If the scarce resource is now the receiver’s decompression – their attention, their working memory, their evenings – then compression tuned to the receiver is no longer merely a stylistic virtue.
This is where the trilogy began. Relevance theory held that a listener weighs what an utterance achieves against the effort of understanding it, and that every act of communication works this way. It was a claim about conversation. At this volume it becomes a claim about the economy.
It is the ethical act of the information economy: the sender paying a cost so that the reader does not have to. Every carefully chosen metaphor, every paragraph that became a sentence, every message that arrived with its context attached is a small transfer of burden from the many readers to the one writer. The machines can generate. Only a writer who has modelled the reader can compress.
Civilisation has spent millennia solving the supply of information. In doing so, it has exposed the scarcity of understanding.
The information age is outrunning human decompression. Information has never been cheaper. Understanding it has never been more expensive.
This completes The Shape of Existence Itself — a trilogy. Part I began with language as compression; Part II asked whether the geometries that structure information might also structure reality; Part III ends with the scarcity exposed when the supply of information becomes nearly limitless: the time and attention required to understand it.
Part I — Cavemen Created the First Large Language Model — was published on Sunday 9 August, and Part II — The Shape of Existence — on Wednesday 12 August.


