Cavemen Created the First Large Language Model - (Part I/III)
It was trained on all human experience, it compresses whole worlds into portable code, and it hallucinates when pushed beyond its training data. Sounds familiar?
This essay is part I of a trilogy of related concepts in languages, information theory, geometry, metaphysics, relativity, and the information economy.
The release of part II will follow on Wednesday 12 August, and part III on Sunday 16 August.
Metaphors are, in essence, linguistic compression algorithms. We use metaphors to communicate complex ideas and concepts in fewer words by projecting information onto a domain we already understand. Metaphors are ZIP files: they compress information and, upon extraction, communicate a high-fidelity representation of the original information before it was compressed.
A bad metaphor is hydraulic fluid: it does not compress, and here it does not faithfully describe the original concept being communicated. The idea has not been shrunk into fewer bits and, upon decompression, the extraction does not accurately represent the original information.
ZIP files work by embedding information contained within source files in a more efficient way, encoding patterns found within the data. If a pattern exists, it can be stored as a pattern with a guide to its extraction, rather than by precisely describing each individual component every time it appears.
Common ZIP compression uses a format called DEFLATE, which combines two complementary tricks. LZ77 identifies sequences that have appeared before and replaces them with directions back to the earlier version. Huffman coding assigns short codes to common symbols and longer codes to rare ones. One removes repetition; the other avoids extravagantly spending bits on things that happen all the time.
The best writers do something similar. Complex ideas, ideologies and emotions are conveyed in their essence, distilling many hours of research, experience and cognition into a comparatively small collection of black squiggles. Words and letters are part of linguistic compression algorithms too – embedding the vast variance of sounds, concepts, experiences, feelings and ideas of humanity itself in a systematised and orderly collection of little shapes.
Not all compression is lossless. And the opening conceit has a confession to make: ZIP is lossless, and metaphors probably never are. A metaphor is closer to a JPEG — convincing at a glance, with the fine detail quietly discarded and the flaws showing up only when you look closely, or extrapolate.

But there is another similarity. ZIP compression is lossless, but a ZIP file still needs decoding software when it arrives. A metaphor needs a decoder too. It must point towards something the other person already understands — and if they cannot decode it, nothing unpacks at all. It’s all Greek to me. Ich verstehe nur Bahnhof.

The shortest route between two minds
The metaphor ‘the market is a jungle’ linguistically encodes many varied concepts: competition, predation, unpredictability, survival of the fittest, and emergent rather than designed order. Five words, but a great deal of information – communicated as much through visual invocation as through the words themselves.
And the extraction? It does not preserve every nuance of market economics, but it rebuilds the big picture with remarkable fidelity for the price of five words. That is an enormous ratio at little cost to meaning. If your account of how markets work cannot be distilled to something like it, the problem is more likely your model than the metaphor.
What the jungle leaves out is the institutional scaffolding: property rights, contract law, courts, regulators. That is not a defect. It is what lossy compression is for — throwing away the detail that does not have immediate purpose for conveying meaning. Artefacts only matter when extrapolation from compressed information is unfaithful to the original.

This is the function of a good metaphor. It identifies structural patterns within one well-understood concept and uses them to encode another. The jungle supplies a ready-made bundle of relationships. The reader probably already understands predators, prey, scarce resources, adaptation, danger and ecological niches. The metaphor does not need to explain each of these from first principles. It invokes a source domain of existing knowledge and projects it onto the target where structural patterns align.

This compression is not confined to metaphor. Every word is a compact retrieval instruction. The word ‘tree’ does not contain bark, roots, leaves, shade, growth rings, branches, nests, autumn, timber or the distant childhood memory of falling out of one.
It is four characters that activate some contextually useful subset of an enormous stored model. A picture may be worth a thousand words, but a well-chosen word can summon several thousand pictures, smells, textures and memories.
Efficient communication is not about brevity in bits alone. It is the optimal route between concepts already shared between sender and recipient. ‘It’s happening again’ can contain almost nothing when said to a stranger, but represent something quite different, known and complicated to a friend. The words are identical; the decompression algorithm is not.
Research into metaphor has long treated it as more than decorative substitution. Conceptual metaphor theory argues that metaphor allows the structure of one conceptual domain to organise another, while work on structure-mapping in analogy describes comparison as the alignment of relational systems rather than a hunt for superficial resemblance.
Relevance theory makes the widest claim of all: that every act of communication works this way – minimal cues, maximal inference, with the receiver’s existing model of the world doing most of the reconstruction.
A good writer therefore models the reader. They intuit which structures already exist in the recipient’s mind, which must be supplied, and which can be invoked with a nudge. A bad writer entertains effortful redundancy without purpose, or simply removes words and declares the result concise.
Every language compresses the same world
English, Russian, Spanish, Mandarin and Tagalog are different linguistic embeddings of largely shared human experience: historically accumulated systems for carving experienced phenomena into reusable patterns and transmitting them between minds.
Embeddings in language overlap enormously because their speakers inhabit the same physical world, possess similar bodies and continually need to discuss food, family, politics, danger, sex, weather and why the dishwasher has once again been loaded by a psychopath.

But languages distribute their descriptive resources differently. They make some distinctions compulsory, others readily available and others irritatingly circumlocutory. English may require an entire phrase to describe a highly specific administrative misery.German is more likely to assemble the relevant nouns into one immense linguistic freight train and drive it directly through the sentence into France.
Russian, meanwhile, conventionally divides what English gathers under blue into two basic categories, commonly transliterated as siniy and goluboy. In experiments, Russian speakers were quicker to distinguish colours that crossed this linguistic boundary than colours that sat within one category. The language did not furnish Russian eyes with additional photoreceptors; it provided a readily accessible distinction through which colour experience could be organised.
Research published in the Proceedings of the National Academy of Sciences found that the effect disappeared under verbal interference, strengthening the argument that language – not a difference in vision – was involved. Our own linguistic embedding spaces shape our experience of the world just as much as they describe it. Would you rather make love in Finnish, or Italian?
The divisions are not arbitrary, though. Berlin and Kay surveyed ninety-eight languages and found colour vocabularies growing in a near-fixed order — light and dark first, then red, then green and yellow, then blue — and clustering around the same focal shades however many terms a language had. Languages carve the spectrum differently. They do not carve it freely.

Languages can also structurally demand different information from their speakers. Evidentiality is the linguistic encoding of how somebody knows what they claim to know: whether an event was witnessed, inferred from evidence or reported by somebody else. In Turkish, for example, source information can be marked grammatically rather than left entirely to optional wording.
English lets the whole question of how you know something dissolve into ‘apparently’ — the source information is simply discarded, which makes gossip fast to transmit and easy to disown.
British English contains phrases whose extracted meaning bears only a passing relationship to the transmitted words. ‘Not ideal’ could describe anything from a marginally delayed train to fire raining down during the Rapture. The recipient is expected to infer the severity from the tone and any available context.
A language is not a mirror held neutrally against reality. It is a culturally inherited compression dictionary for navigating it. Each embodies decisions – most of them made gradually and passively by nobody in particular – about which patterns recur often enough to deserve a convenient code.
This does not require the strong claim that language imprisons thought. Speakers can describe concepts for which their language has no single word, just as a ZIP file can contain an unfamiliar format. Contemporary accounts of linguistic relativity generally concern influences on habitual attention, categorisation and processing rather than an absolute inability to conceive whatever one’s grammar declines to provide.
Some paths are simply shorter than others: a distinction named, grammaticalised and repeatedly reinforced is more readily available than one that must be reconstructed from scratch every time. Languages are therefore not merely collections of vocabulary. They are different ways of embedding low-cost routes that describe the human experience.
Knowing means being wrong less often
It is worth asking at this point: what is knowledge? Arguably, the simplest definition of knowledge is structured information with low prediction error. It does not merely reproduce an answer it has seen before (also known as overfitting); it captures a pattern that remains useful when inputs change.
That is a form of compression, and equating it with comprehension has a serious lineage, from Gregory Chaitin’s algorithmic information theory through the Hutter Prize‘s premise that compressing human knowledge is equivalent to understanding it, to recent work demonstrating that large language models function as competitive general-purpose compressors.

A student who has comprehensively studied a subject will probably perform well in an exam. Their internal model of the domain will be strong and should extract faithfully when challenged by unfamiliar questions. A peer who has spent more time in the bar than the library is more likely to answer incorrectly. Their responses will be more chaotic, less reliably connected to reality and more vulnerable to confident improvisation.
In another domain, we would call this ‘hallucination‘.
The gap between large-language-model hallucination and a student with an underdeveloped understanding of an exam topic is conceptually smaller than popular culture might imagine. Failure takes the same shape: output assembled from a region the model is probabilistically unlikely to cover faithfully.
What differs is the scope of the gap. The student’s model is thin more or less everywhere across the domain, because they did not do the reading, and an examiner should be able to anticipate roughly how and where it will fail.
Modern LLMs cover an enormous scope and – judging by the lucidity and quality of most generated text – it is thin only in patches, which makes its lapses far harder to anticipate: it can be genuinely authoritative about a subject in one sentence and invent a mistruth and attribute it to a fabricated source in the next.
Nothing in the surrounding competence tells you which patch you are standing on. In both cases the fluency is real and the grounding is not – which is precisely why fluency can be such a treacherous signal of knowledge.

A speaker who intuits a grammatical rule can apply it to sentences they have never encountered before. They do not merely remember that one verb takes a particular form; they recognise the pattern and extend it to valid sentence forumlations. Their knowledge is not a list of approved sentences, but a model capable of producing new ones.
You know that the sentence ‘I live at the wooden red big house’ is grammatically out of order, even if you hadn’t yet come to realise that English grammar expects a specific order to its adjectives. The same adjective order convention applies across English sentences.
Meanwhile, a stopped clock is not knowledgeable, because the prediction error across its relevant domain is enormous. Two correct outputs per day do not rescue 1,438 wrong ones. It has not learned an efficient, predictive or descriptive model of time.
The same applies to memorisation without understanding. A student who remembers one exact answer but cannot use it outside the circumstances in which it was learned possesses a brittle, crystalline compression that fractures under pressure. The information extracts correctly only when the question reproduces the original path.
Generative knowledge is therefore not only the capacity to produce a correct answer - that could just be lookup or recollection. It is the compression of observations into a model that remains predictively useful beyond the particular data from which it was constructed.
This aligns with the broad logic of predictive-processing theories. The brain is modelled as continually generating expectations, comparing them with incoming signals and using prediction errors to revise its internal representations. Prediction error also plays a central role in reinforcement learning, and has been used to explain how people learn not only from their own outcomes but from the actions and rewards of others.
More recent experimental work using magnetoencephalography found evidence that predictive learning reshapes the representational geometry of the human brain. Regularities learned across acoustic sequences changed how the brain represented those inputs, connecting prediction error with the organisation of internal representations rather than merely the correction of individual mistakes.
Learning, on this account, is the iterative refinement of a compression algorithm.
Claude Shannon counts surprise
This was the revolutionary insight formalised by the 20th-century mathematician Claude Shannon: information can be understood through uncertainty.
For each possible outcome, take how likely it is and multiply it by how surprising it would be to see it – where surprise is measured as the negative logarithm of the outcome’s probability, so rare events carry more of it. Add those values across all outcomes and you have the average surprise, measured in bits.
A coin that always lands heads communicates no new information when it lands heads. A fair coin communicates one bit because either result was equally plausible before the toss.
A sentence in which every word is completely predetermined is informationally inert. One in which every word is wholly random is maximally surprising. Useful communication sits between the two. It contains enough predictability to be interpretable and enough surprise to justify having interrupted you.

Shannon’s theory deliberately set aside the semantic value or truth of a message. The opening to his paper notes that the engineering problem of communication concerns reproducing a selected message, regardless of whether that message is meaningful. That boundary is vital: information theory can quantify the uncertainty resolved by a signal without determining whether anybody was better off receiving it.
If a message is predictable – if it contains redundancy – it can be represented in fewer bits. Huffman coding in ZIP files does exactly this: it assigns short codes to frequently occurring symbols and longer codes to rare ones.
The compression system spends its smallest codes on the things it expects to see most often. Find the redundancy. Exploit it. Say less without losing more. Thus, a great metaphor is a Huffman code for meaning.
English already works this way. Shannon estimated its redundancy by asking people to guess each next character of a passage and recording how many guesses it took. With long-range context included, he put ordinary written English at roughly 75 per cent redundant — about one bit of new information per character. George Zipf had noticed the consequence years earlier. The shortest words in English are the most load-bearing: pronouns, articles, prepositions and conjunctions — it, a, he, she, at, the, and, of. Nothing required in almost every sentence is permitted to be long. The language spends its briefest codes on its commonest meanings, exactly as a compressor would.
You can even observe the progression of compression through etymology. God be with you was already being used to say farewell by the late fifteenth century. Over the next hundred years it appears in increasingly compact forms: God be wy you, God b’uy, God buoye, godbwye.
Then God became good, probably pulled sideways by good day and good night. Goodbye eventually lost another syllable and can just be bye. Four words became one. The theology disappeared; the social function survived. Nobody saying bye is invoking the Almighty. English kept the useful bit and deleted the rest.
The rate of compression is related to usage. Omnibus, Latin for ‘for all’, became bus once everyone was catching one. The taximeter cab was contracted almost at once: taxicab was in print by March 1907 and taxi by November. English had compressed the new machine almost as quickly as it arrived. And comparing recordings of New Zealand English across generations, words whose usage rose over time became measurably quicker to say, while words falling out of use grew longer.
The refinement may go further. One study covering eleven languages found that contextual predictability tracked word length better than raw frequency, although later analyses have not consistently replicated that ordering. Language clearly shortens what is common, and may also shorten what the listener can already half-supply.
We ask people: ‘Could you explain this complicated concept to your grandparents?’ We request an ‘elevator pitch,’ or an ‘executive summary.’ Or perhaps: ‘Explain it like I’m five years old.’ TL;DR. These are cultural embeddings of linguistic compression algorithms. They also encode the wider value of concept refinement.
That paragraph you just sent to your colleague: could it have been a sentence? Would they understand what you were trying to tell them better if you gave them fewer bits to process?
Good communication is not the transmission of the largest available quantity of information. It is the selection of the smallest quantity capable of rebuilding the right model inside its recipient(s).
The shortest possible thought
There is a related concept – Kolmogorov complexity – which asks: what is the shortest possible description of something that fully specifies it?
More formally, it is the length of the shortest program that causes a specified universal computing system to reproduce an object. The wording ‘shortest possible description’ is a standard accessible rendering because the program functions as a complete generative description.
A thousand repetitions of ‘AB’ require two thousand characters if written individually, but very few if described as ‘print AB one thousand times’. The digits of π appear disordered if written out one by one, but can be produced from a comparatively short generative instruction.

long double z_re = 0, z_im = 0;
int iter = 0;
while (iter < 50000) {
long double z_re2 = z_re * z_re, z_im2 = z_im * z_im;
if (z_re2 + z_im2 >= 40000) break; // this point escapes
z_im = 2 * z_re * z_im + c_im; // z ← z² + c
z_re = z_re2 - z_im2 + c_re;
iter++;
}A perfect metaphor aspires to minimum Kolmogorov complexity for a target concept: the shortest possible linguistic instruction capable of reconstructing the idea with high fidelity.
There is an important human variation. Formal Kolmogorov complexity assumes a specified computational system. Human recipients are not identical machines.
The shortest explanation capable of reconstructing an economic concept in the mind of an economist may be incomprehensible to a child (just as likely a politician). The economist already possesses a large library of relevant patterns; the interlocutor requires the supporting files.
The minimum description of an idea therefore depends partly on the person receiving it.
This resembles the principle behind minimum description length, a model-selection framework which favours the account that gives the shortest combined description of a model and the data left unexplained by it.
An explanation that is very short only because it leaves everything important unresolved has not achieved elegant compression. It has hidden the remaining information in the error term.
This is why jargon can be elegant inside a profession and unbearable outside it. A technical term can compress an entire chapter for one reader, while presenting another with an indecipherable challenge.
A great metaphor, therefore, finds a short generative instruction that works across a large number of disparate minds: the philosopher’s stone of linguistic compression.
Nobody designed it; everybody trained it
Here is a system trained on everything its users have ever needed to say: what is safe to eat, who is not to be trusted, whether that noise was anything, and precisely what was meant by fine.
It has no specification, no version number, no maintainer, no issue tracker, no pull requests desperately vying for a review from @codeowners, and no roadmap. It was assembled over tens of thousands of years by people who were, for the most part, simply trying to get through the day with a full belly and empty balls.
This is not a description of a large language model. It is a description of English, and of Mandarin, Russian and Tagalog, and every other language spoken by humans. Nobody designed it; everybody trained it.
The resemblance between language and artificial intelligence is more than a clever parallel. It poses a question. We have spent a decade asking what machines might do to human thought, while paying far less attention to the older system already inside it: a generative compression dictionary inherited rather than chosen.
Every time you reach for a word, you follow a path worn into language by ancestors who never knew you would need it for the sentence you are formulating, nor the purpose of you uttering it.
Cavemen and their descendants created the first large language model, and they didn’t even know they were doing it.

A path worn into grass tells you something about the ground underneath it. If language is a compression of the world, its shape is evidence about the shape of the world.
Follow that evidence far enough and the question stops being about the shape of language. It starts to become: What is the shape of existence itself?
Part II — The Shape of Existence — follows on Wednesday 12 August. It begins with six trillion random proteins and four that bind ATP, then proceeds, with admirable restraint, to whether mathematics merely describes the machinery or might be the machinery.
Part III — Attention Is the Last Paywall — follows on Sunday 16 August.

