How a Language Model Learns

Computed figure: columns of small dots connected by thin lines, like the layers of a neural network, with a few paths through it drawn in magenta and growing stronger from left to right

What is inside a language model, how it learns from text, why it grows more capable, how the models have evolved since 2017, and how all this compares with a human brain.

The companion explainer, How an AI Agent Works, describes the machinery around a model: the agent loop, the testing, the safeguards. This one opens the model itself. It explains what it is made of, how it learns, why each generation is more capable, how fast that has happened, and where the comparison with a human brain holds and where it fails. It assumes no technical background, simplifies where it must, and says so.

The structure

Numbers, layers, and the trick called attention

A language model is a neural network: a very large arrangement of simple units loosely inspired by neurons. Each unit takes in numbers, multiplies each by a weight, adds them up, passes the result through a simple bend — a function that, roughly, ignores small signals and passes strong ones — and hands it on. On its own a unit does almost nothing. Stacked by the million in layers, each layer working on the output of the one before, they can represent remarkably complex patterns. The weights are the strengths of all these connections, and there are hundreds of billions to trillions of them. They are the only thing that changes when a model learns.

Text has to become numbers first. It is cut into tokens — words or pieces of words — and each token is turned into a long list of numbers, its embedding, a position in a space of thousands of dimensions. In that space, related meanings end up near each other: training arranges it so that “Lisbon” sits close to “Porto” and far from “photosynthesis”.

The design used by almost every current model is the transformer, published by Google researchers in 2017 under the title “Attention Is All You Need”.[1] Its key idea is attention. In the sentence “The runner passed the bottle to the volunteer because she was thirsty”, working out who “she” is requires looking back at other words. Attention lets every word, at every layer, weigh every other word in the text and pull in information from the ones that matter: “she” draws on “runner” and “thirsty”. A transformer is many layers that alternate between this mixing across words and a processing step for each word, refining the representation of the whole text a little at each layer. At the top, the last layer gives a probability for every possible next token (Figure 1).

A transformer, much simplified Left to right: text, cut into tokens; each token turned into a list of numbers, its embedding; a stack of repeated layers, each with attention, where every word looks at the others, and a processing step for each word; at the top, a probability for every possible next token. A transformer, much simplified “the runner passed the…” tokens → lists of numbers attention: each word looks at the others processing step for each word one layer · repeated dozens of times probability of each next token
Fig. 1 — A transformer, much simplified. Text becomes tokens, tokens become lists of numbers, and those pass through many repeated layers, each combining attention (every word looking at the others) with a processing step for each word. The output is a probability for each possible next token. Real models have dozens to over a hundred layers. Schematic.
Learning

Walking downhill in fog

A new model’s weights are random, and its predictions are nonsense. Learning means adjusting the weights so the predictions improve. The procedure has three parts.

1 · Measure the error
Show the model a piece of real text, hide the next token, and see what probability the model gave to the right one. A single number, the loss, measures how wrong it was: low if it gave the right token a high probability, high if not.
2 · Find which way is downhill
For every one of the billions of weights, work out whether nudging it up or down would have reduced the error, and by how much. A method called backpropagation, popularised in 1986, does this efficiently by passing the error backwards through the layers.[2]
3 · Take a small step
Nudge every weight a tiny amount in the helpful direction. This is gradient descent. Then repeat with the next piece of text, trillions of times.

The usual image is a walker in fog on a vast hilly landscape, trying to reach the lowest valley: unable to see far, they feel which way the ground slopes under their feet and take a step that way (Figure 2). The landscape has as many directions as the model has weights, which is hard to picture and, as it turns out, helpful: with so many directions there is nearly always some way down.

Gradient descent in one dimension A curve shows the error as one weight changes, lowest near the right of centre. Dots starting high on the left step down the slope, in shrinking steps, to the bottom of the valley. Gradient descent, in one dimension value of one weight error (loss) start: random weights, large error small steps near the bottom
Fig. 2 — Gradient descent, in one dimension. The curve is the error as one weight changes; each step moves the weight a little in the direction that lowers the error. A real model takes such steps in billions of dimensions at once. Schematic.

The striking thing is what this simple goal produces. To predict the next word of text written by people across every subject, a model has to absorb a great deal: spelling and grammar, facts about the world, the shape of an argument, the steps of a calculation, the conventions of computer code. Nobody tells it any of this. It is the cheapest way to get the error down. The training data for a frontier model is now measured in trillions of tokens: Meta’s Llama 3.1, for example, was trained on about 15.6 trillion.[3]

Everything after this first stage — teaching the model to follow instructions, to refuse harmful requests, to solve problems step by step — uses the same machinery with different data and rewards. How an AI Agent Works describes those stages.

What emerges

Generalising, and learning from the prompt

A model that only memorised its training text would be useless on anything new. What makes these models valuable is generalisation: they handle questions and texts they have never seen, because what they learned are patterns, not copies. They are not perfect at it — they also memorise, and they can be thrown by questions that differ only slightly from familiar ones, the “jagged” behaviour discussed in After the Warning Shot — but the ability is real and grows with scale.

Two findings surprised researchers. The first, from OpenAI’s GPT-3 in 2020, is in-context learning: a large model can pick up a new task from a few examples written into the prompt, with no change to its weights at all.[4] The weights are frozen; the “learning” happens inside the processing of the text, and is lost when the conversation ends.

The second is the appearance of abilities at scale. In 2022 Google researchers described “emergent” abilities, “not present in smaller models but … present in larger models”, which could not be predicted from smaller ones.[5] A year later Stanford researchers argued that many such jumps reflect “the researcher’s choice of metric rather than … fundamental changes in model behavior”: scored all-or-nothing, a skill seems to switch on suddenly; scored by partial credit, it improves smoothly.[6] Both can be true. The underlying improvement is often gradual, but the moment it becomes useful — or dangerous — can still arrive abruptly, which is why evaluations try to measure skills before they are complete.

Why power grows

Four engines of improvement

1 · More computing for training
The amount of computation used to train frontier models has grown about five times a year since 2020.[7] It is measured in FLOP, the number of basic arithmetic operations: GPT-2, in 2019, used about 2×1021; GPT-4, in 2023, about 2×1025, ten thousand times more; the largest model of 2026 in Epoch AI’s database, about 1027.[3]
2 · More data
Bigger models need more text to be trained well. Researchers at DeepMind showed in 2022 that earlier models had been too large for their data, and that size and data should grow together.[8] That creates a limit: Epoch AI projects that models will be trained on datasets as large as the whole stock of public human text “between 2026 and 2032”, after which progress must come from synthetic data, other kinds of data, or better use of what exists.[9]
3 · Better methods
Improvements in design and training make each unit of computing go further. Epoch AI estimated in 2024 that the computing needed to reach a given performance halves roughly every eight months, faster than chips improve — though the growth in raw computing still contributed more.[10]
4 · More thinking per answer
Since 2024 a new route: letting a model reason at length before answering, and training it to reason well. OpenAI reported in 2024 that its o1 model’s performance “consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)”.[11] Capability now grows not only with the size of the model but with how long it is allowed to work on a problem.

The regularity of the first three is what researchers call scaling laws: in 2020 OpenAI found that a model’s error falls smoothly and predictably as size, data and computing grow, over many orders of magnitude.[12] It is the main reason companies keep building larger models, and the main reason for disagreement about the future: nobody knows how long the curves continue.

Architecture

Why the transformer won, and whether design still matters

Before 2017 the leading language models read text one word at a time, in order, carrying a running summary forward. That made them slow to train: each step had to wait for the last. The transformer’s attention looks at all the words at once, so the work can be split across thousands of graphics chips running in parallel.[1] That fit with the hardware is the main reason it won: it let computing, the first engine above, be used at enormous scale.

Two design ideas matter for today’s models. The context window is how much text a model can attend to at once: a few thousand tokens in early models, hundreds of thousands to millions now, which is what lets an agent keep a long task in view. And mixture of experts splits a model into many specialised sub-networks and uses only a few for each token: DeepSeek-V3, for example, has 671 billion weights but uses only about 37 billion for each token, so it has the knowledge of a very large model at the running cost of a much smaller one.[13]

Does design still matter, or only scale? Both. Scaling laws hold for a given design, but better designs shift the whole curve, and the eight-month halving of required computing comes partly from them.[10] What has not happened since 2017 is a change as large as the transformer itself; most progress has come from making it bigger, feeding it more and training it better.

The record

From 2017 to 2026 in one chart

Figure 3 shows the computing used to train notable language models, from Epoch AI’s database. The scale is logarithmic: each gridline is a hundred times the one below. In nine years the largest training runs grew about a hundred million times, from the original transformer to GPT-6 Astra.[3] What that bought in capability is shown, on METR’s measure of the length of task agents can complete, in Figure 1 of After the Warning Shot: from seconds for GPT-2 to about twelve hours for the best models of early 2026.[14]

Training computing for notable language models, 2017 to 2026 Logarithmic scale in FLOP. Milestones: GPT-6 Astra (2026-09-03) 1.0e+27; Grok 4 (2025-07-09) 5.0e+26; DeepSeek-V3 (2024-12-24) 3.3e+24; Llama 3.1 (2024-07-23) 3.8e+25; GPT-4 (2023-03-15) 2.1e+25; GPT-3 (2020-05-28) 3.1e+23; GPT-2 (2019-02-14) 1.9e+21; Transformer (2017-06-12) 7.4e+18. In nine years the largest training runs grew about a hundred million times. Computing used to train language models, FLOP 1018 1020 1022 1024 1026 1028 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 GPT-6 Astra Grok 4 DeepSeek-V3 Llama 3.1 GPT-4 GPT-3 GPT-2 Transformer each gridline is 100 times the one below · grey: other language models
Fig. 3 — Training computing for notable language models, 2017–2026, in FLOP (basic arithmetic operations), logarithmic scale. Grey dots are all language models in the database; magenta dots, labelled, are milestones. Figures for closed models are Epoch AI’s estimates, with stated uncertainty. Data: Epoch AI, Notable AI Models, accessed 3 October 2026.
Model and brain

Inspired by neurons, unlike a brain

The words “neural network” and “learning” invite the comparison. It helps in places and misleads in others.

Units
An adult human brain has about 86 billion neurons, each a living cell with complex chemistry, connected by a far larger number of synapses: a careful count found about 164 trillion in the neocortex alone, the brain’s outer layer, and there is no equally careful count for the whole brain.[15][16] A model has up to trillions of weights, each a single number. The counts are not comparable: a synapse is not a weight, and a neuron is far richer than an artificial unit.
Energy
The brain is about 2 per cent of body weight but uses about 20 per cent of the body’s energy, roughly 20 watts.[17] Training a frontier model draws tens to hundreds of megawatts for months; Epoch AI estimates GPT-6 Astra’s training drew about 230 megawatts.[3]
Data
“Children can acquire language from less than 100 million words of input”, while language models “typically require 3 or 4 orders of magnitude more data”, in the words of the BabyLM researchers; today’s frontier models are trained on trillions of tokens.[18][3] Children are vastly more efficient learners — with the help of a body, a world and other people.
Learning
A brain learns all the time, from each experience. A model learns in a training phase and is then frozen; it does not remember yesterday’s conversation unless something stores it and puts it back in its context. And the brain is not known to learn by backpropagation in the form models use.
Grounding
A brain sits in a body that acts in the world, with needs and senses. A language model learns mostly from text about the world, which is part of why it can be fluent and wrong at once.
Experience
Whether there is anything it is like to be a model is an open question that this blog does not try to settle. Nothing in how models are built or trained requires it, and fluent talk about feelings is exactly what training on human text would produce either way.

The useful lesson runs in both directions. Models show that a great deal of what looks like understanding can be learned from prediction alone, at enormous scale. Brains show how much more efficient learning can be, and how much of human intelligence comes from things models lack: a body, continuous learning, and a life among other people.

Glossary

The terms used here — neural network, layer, embedding, attention, loss, gradient descent, backpropagation, context window, mixture of experts, in-context learning and others — are defined in the series glossary in How an AI Agent Works.

On method and tools

This piece was written collaboratively with Claude Opus 5.5 (Anthropic): human specification, editorial direction and critical review; machine research and drafting. It is a general explainer and simplifies: the description of a neuron, of attention and of training leaves out much that matters to specialists. Figures 1 and 2 are schematic diagrams; Figure 3 is drawn from Epoch AI’s published database, whose figures for closed models are estimates. Facts and quotations were checked against the cited sources on 3 October 2026. The cover is computed by scripts/how_model_learns_cover.py. Technical detail in this series follows one standard: it explains why a control failed and what that shows, but gives no reproducible procedure. Two AI tools were used, and both makers have a stake in the subject: Claude (Anthropic) for the three essays and the two explainers, and Codex (OpenAI) for the story; Anthropic and OpenAI both appear in the evidence.

Authored by: Luis Matos Ferreira — Physicist, Developer, Writer

Updates and corrections

Last updated 3 October 2026. The events described are still unfolding; facts are as known on that date.

  1. No corrections so far.
The series on the July 2026 incident

Six pieces. The Answer Key and The Warning Shot tell the same story at two lengths: read one or the other. Two routes through the rest:

  1. The Answer Key — the short account of the July 2026 incident.
  2. The Warning Shot — the long account: the test, the waves, the debate.
  3. After the Warning Shot — what more capable systems may bring: threats, evidence, sceptics, defences.
  4. How an AI Agent Works — the explainer on agents, testing and safeguards, with the series glossary.
  5. How a Language Model Learns (this piece) — the explainer on what is inside a model, how it learns and how it compares with a brain.
  6. The Completion — a short story set in Lisbon in 2027.
Sources
  1. A. Vaswani et al., “Attention Is All You Need”, arXiv 1706.03762, June 2017, arxiv.org.
  2. D. E. Rumelhart, G. E. Hinton and R. J. Williams, “Learning representations by back-propagating errors”, Nature 323, 1986, nature.com.
  3. Epoch AI, “Notable AI Models” database, accessed 3 October 2026, epoch.ai.
  4. T. Brown et al., “Language Models are Few-Shot Learners”, arXiv 2005.14165, May 2020, arxiv.org.
  5. J. Wei et al., “Emergent Abilities of Large Language Models”, arXiv 2206.07682, 2022, arxiv.org.
  6. R. Schaeffer, B. Miranda and S. Koyejo, “Are Emergent Abilities of Large Language Models a Mirage?”, arXiv 2304.15004, 2023, arxiv.org.
  7. Epoch AI, “Trends in artificial intelligence”, updated 5 February 2026, epoch.ai.
  8. J. Hoffmann et al., “Training Compute-Optimal Large Language Models”, arXiv 2203.15556, March 2022, arxiv.org.
  9. P. Villalobos et al., “Will we run out of data? Limits of LLM scaling based on human-generated data”, arXiv 2211.04325, revised June 2024, arxiv.org.
  10. A. Ho et al., “Algorithmic progress in language models”, arXiv 2403.05812, March 2024, arxiv.org.
  11. OpenAI, “Learning to reason with LLMs”, 12 September 2024, openai.com.
  12. J. Kaplan et al., “Scaling Laws for Neural Language Models”, arXiv 2001.08361, January 2020, arxiv.org.
  13. DeepSeek-AI, “DeepSeek-V3 Technical Report”, arXiv 2412.19437, December 2024, arxiv.org.
  14. METR, “Task-completion time horizons of frontier AI models”, time horizon 1.1, data file benchmark_results_1_1.yaml, updated 8 May 2026, metr.org.
  15. F. A. C. Azevedo et al., “Equal numbers of neuronal and nonneuronal cells make the human brain an isometrically scaled-up primate brain”, Journal of Comparative Neurology 513(5), 2009, pubmed.ncbi.nlm.nih.gov.
  16. Y. Tang, J. R. Nyengaard, D. M. De Groot and H. J. Gundersen, “Total regional and global number of synapses in the human brain neocortex”, Synapse 41(3), 2001, pubmed.ncbi.nlm.nih.gov.
  17. M. E. Raichle and D. A. Gusnard, “Appraising the brain’s energy budget”, PNAS 99(16), 2002, ncbi.nlm.nih.gov.
  18. A. Warstadt et al., “Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora”, arXiv 2504.08165, 2025, arxiv.org.

Comentários

Mensagens populares deste blogue

Work, Time and Money

The Warning Shot

Novos Desafios

The Duty to Work

EMUM - Eco Madeira Ultra Maratona 2016

The Fifteen-Hour Week

Provas Insanas - Westfield Sydney to Melbourne Ultramarathon 1983