How a Language Model Learns
What is inside a language model, how it learns from text, why it grows more capable, how the models have evolved since 2017, and how all this compares with a human brain.
The companion explainer, How an AI Agent Works, describes the machinery around a model: the agent loop, the testing, the safeguards. This one opens the model itself. It explains what it is made of, how it learns, why each generation is more capable, how fast that has happened, and where the comparison with a human brain holds and where it fails. It assumes no technical background, simplifies where it must, and says so.
Numbers, layers, and the trick called attention
A language model is a neural network: a very large arrangement of simple units loosely inspired by neurons. Each unit takes in numbers, multiplies each by a weight, adds them up, passes the result through a simple bend — a function that, roughly, ignores small signals and passes strong ones — and hands it on. On its own a unit does almost nothing. Stacked by the million in layers, each layer working on the output of the one before, they can represent remarkably complex patterns. The weights are the strengths of all these connections, and there are hundreds of billions to trillions of them. They are the only thing that changes when a model learns.
Text has to become numbers first. It is cut into tokens — words or pieces of words — and each token is turned into a long list of numbers, its embedding, a position in a space of thousands of dimensions. In that space, related meanings end up near each other: training arranges it so that “Lisbon” sits close to “Porto” and far from “photosynthesis”.
The design used by almost every current model is the transformer, published by Google researchers in 2017 under the title “Attention Is All You Need”.[1] Its key idea is attention. In the sentence “The runner passed the bottle to the volunteer because she was thirsty”, working out who “she” is requires looking back at other words. Attention lets every word, at every layer, weigh every other word in the text and pull in information from the ones that matter: “she” draws on “runner” and “thirsty”. A transformer is many layers that alternate between this mixing across words and a processing step for each word, refining the representation of the whole text a little at each layer. At the top, the last layer gives a probability for every possible next token (Figure 1).
Walking downhill in fog
A new model’s weights are random, and its predictions are nonsense. Learning means adjusting the weights so the predictions improve. The procedure has three parts.
The usual image is a walker in fog on a vast hilly landscape, trying to reach the lowest valley: unable to see far, they feel which way the ground slopes under their feet and take a step that way (Figure 2). The landscape has as many directions as the model has weights, which is hard to picture and, as it turns out, helpful: with so many directions there is nearly always some way down.
The striking thing is what this simple goal produces. To predict the next word of text written by people across every subject, a model has to absorb a great deal: spelling and grammar, facts about the world, the shape of an argument, the steps of a calculation, the conventions of computer code. Nobody tells it any of this. It is the cheapest way to get the error down. The training data for a frontier model is now measured in trillions of tokens: Meta’s Llama 3.1, for example, was trained on about 15.6 trillion.[3]
Everything after this first stage — teaching the model to follow instructions, to refuse harmful requests, to solve problems step by step — uses the same machinery with different data and rewards. How an AI Agent Works describes those stages.
Generalising, and learning from the prompt
A model that only memorised its training text would be useless on anything new. What makes these models valuable is generalisation: they handle questions and texts they have never seen, because what they learned are patterns, not copies. They are not perfect at it — they also memorise, and they can be thrown by questions that differ only slightly from familiar ones, the “jagged” behaviour discussed in After the Warning Shot — but the ability is real and grows with scale.
Two findings surprised researchers. The first, from OpenAI’s GPT-3 in 2020, is in-context learning: a large model can pick up a new task from a few examples written into the prompt, with no change to its weights at all.[4] The weights are frozen; the “learning” happens inside the processing of the text, and is lost when the conversation ends.
The second is the appearance of abilities at scale. In 2022 Google researchers described “emergent” abilities, “not present in smaller models but … present in larger models”, which could not be predicted from smaller ones.[5] A year later Stanford researchers argued that many such jumps reflect “the researcher’s choice of metric rather than … fundamental changes in model behavior”: scored all-or-nothing, a skill seems to switch on suddenly; scored by partial credit, it improves smoothly.[6] Both can be true. The underlying improvement is often gradual, but the moment it becomes useful — or dangerous — can still arrive abruptly, which is why evaluations try to measure skills before they are complete.
Four engines of improvement
The regularity of the first three is what researchers call scaling laws: in 2020 OpenAI found that a model’s error falls smoothly and predictably as size, data and computing grow, over many orders of magnitude.[12] It is the main reason companies keep building larger models, and the main reason for disagreement about the future: nobody knows how long the curves continue.
Why the transformer won, and whether design still matters
Before 2017 the leading language models read text one word at a time, in order, carrying a running summary forward. That made them slow to train: each step had to wait for the last. The transformer’s attention looks at all the words at once, so the work can be split across thousands of graphics chips running in parallel.[1] That fit with the hardware is the main reason it won: it let computing, the first engine above, be used at enormous scale.
Two design ideas matter for today’s models. The context window is how much text a model can attend to at once: a few thousand tokens in early models, hundreds of thousands to millions now, which is what lets an agent keep a long task in view. And mixture of experts splits a model into many specialised sub-networks and uses only a few for each token: DeepSeek-V3, for example, has 671 billion weights but uses only about 37 billion for each token, so it has the knowledge of a very large model at the running cost of a much smaller one.[13]
Does design still matter, or only scale? Both. Scaling laws hold for a given design, but better designs shift the whole curve, and the eight-month halving of required computing comes partly from them.[10] What has not happened since 2017 is a change as large as the transformer itself; most progress has come from making it bigger, feeding it more and training it better.
From 2017 to 2026 in one chart
Figure 3 shows the computing used to train notable language models, from Epoch AI’s database. The scale is logarithmic: each gridline is a hundred times the one below. In nine years the largest training runs grew about a hundred million times, from the original transformer to GPT-6 Astra.[3] What that bought in capability is shown, on METR’s measure of the length of task agents can complete, in Figure 1 of After the Warning Shot: from seconds for GPT-2 to about twelve hours for the best models of early 2026.[14]
Inspired by neurons, unlike a brain
The words “neural network” and “learning” invite the comparison. It helps in places and misleads in others.
The useful lesson runs in both directions. Models show that a great deal of what looks like understanding can be learned from prediction alone, at enormous scale. Brains show how much more efficient learning can be, and how much of human intelligence comes from things models lack: a body, continuous learning, and a life among other people.
The terms used here — neural network, layer, embedding, attention, loss, gradient descent, backpropagation, context window, mixture of experts, in-context learning and others — are defined in the series glossary in How an AI Agent Works.
This piece was written collaboratively with Claude Opus 5.5 (Anthropic): human specification, editorial direction and critical review; machine research and drafting. It is a general explainer and simplifies: the description of a neuron, of attention and of training leaves out much that matters to specialists. Figures 1 and 2 are schematic diagrams; Figure 3 is drawn from Epoch AI’s published database, whose figures for closed models are estimates. Facts and quotations were checked against the cited sources on 3 October 2026. The cover is computed by scripts/how_model_learns_cover.py. Technical detail in this series follows one standard: it explains why a control failed and what that shows, but gives no reproducible procedure. Two AI tools were used, and both makers have a stake in the subject: Claude (Anthropic) for the three essays and the two explainers, and Codex (OpenAI) for the story; Anthropic and OpenAI both appear in the evidence.
Authored by: Luis Matos Ferreira — Physicist, Developer, Writer
Last updated 3 October 2026. The events described are still unfolding; facts are as known on that date.
- No corrections so far.
Six pieces. The Answer Key and The Warning Shot tell the same story at two lengths: read one or the other. Two routes through the rest:
- The Answer Key — the short account of the July 2026 incident.
- The Warning Shot — the long account: the test, the waves, the debate.
- After the Warning Shot — what more capable systems may bring: threats, evidence, sceptics, defences.
- How an AI Agent Works — the explainer on agents, testing and safeguards, with the series glossary.
- How a Language Model Learns (this piece) — the explainer on what is inside a model, how it learns and how it compares with a brain.
- The Completion — a short story set in Lisbon in 2027.
- A. Vaswani et al., “Attention Is All You Need”, arXiv 1706.03762, June 2017, arxiv.org.
- D. E. Rumelhart, G. E. Hinton and R. J. Williams, “Learning representations by back-propagating errors”, Nature 323, 1986, nature.com.
- Epoch AI, “Notable AI Models” database, accessed 3 October 2026, epoch.ai.
- T. Brown et al., “Language Models are Few-Shot Learners”, arXiv 2005.14165, May 2020, arxiv.org.
- J. Wei et al., “Emergent Abilities of Large Language Models”, arXiv 2206.07682, 2022, arxiv.org.
- R. Schaeffer, B. Miranda and S. Koyejo, “Are Emergent Abilities of Large Language Models a Mirage?”, arXiv 2304.15004, 2023, arxiv.org.
- Epoch AI, “Trends in artificial intelligence”, updated 5 February 2026, epoch.ai.
- J. Hoffmann et al., “Training Compute-Optimal Large Language Models”, arXiv 2203.15556, March 2022, arxiv.org.
- P. Villalobos et al., “Will we run out of data? Limits of LLM scaling based on human-generated data”, arXiv 2211.04325, revised June 2024, arxiv.org.
- A. Ho et al., “Algorithmic progress in language models”, arXiv 2403.05812, March 2024, arxiv.org.
- OpenAI, “Learning to reason with LLMs”, 12 September 2024, openai.com.
- J. Kaplan et al., “Scaling Laws for Neural Language Models”, arXiv 2001.08361, January 2020, arxiv.org.
- DeepSeek-AI, “DeepSeek-V3 Technical Report”, arXiv 2412.19437, December 2024, arxiv.org.
- METR, “Task-completion time horizons of frontier AI models”, time horizon 1.1, data file benchmark_results_1_1.yaml, updated 8 May 2026, metr.org.
- F. A. C. Azevedo et al., “Equal numbers of neuronal and nonneuronal cells make the human brain an isometrically scaled-up primate brain”, Journal of Comparative Neurology 513(5), 2009, pubmed.ncbi.nlm.nih.gov.
- Y. Tang, J. R. Nyengaard, D. M. De Groot and H. J. Gundersen, “Total regional and global number of synapses in the human brain neocortex”, Synapse 41(3), 2001, pubmed.ncbi.nlm.nih.gov.
- M. E. Raichle and D. A. Gusnard, “Appraising the brain’s energy budget”, PNAS 99(16), 2002, ncbi.nlm.nih.gov.
- A. Warstadt et al., “Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora”, arXiv 2504.08165, 2025, arxiv.org.
Comentários
Enviar um comentário