Many Ways to Be Clever

Many Ways to Be Clever

The AI series · 8 October 2026

The biggest brain on Earth belongs to a whale that is not the cleverest animal. In brains and in AI models alike, what a system can do depends less on its size than on where the capacity goes and what it was built for.

The whale’s brain

A sperm whale’s brain weighs up to about 9 kilograms, some six times a human one, and it is the heaviest brain known.[1] Nobody thinks sperm whales are six times cleverer than we are. Something about size does matter, though, or there would be no reason for brains to be expensive. The human brain is about 2 per cent of the body’s weight but uses about 20 per cent of its energy,[2] roughly 20 watts.[3]

How a Language Model Learns ended with one brain and one model and the warning that their counts are not comparable. This piece asks a wider question. Different brains solve different problems in different ways, and so do different AI models. Looking at the variety, rather than at one comparison, shows what size buys, what it does not, and why “how big is it?” is usually the wrong first question.

Size has to be corrected for

The first thing a whale’s brain has to do is run a whale. Larger bodies have more skin to feel, more muscle to move and more space to map, and brains grow with them. Across mammal species, brain mass rises roughly with body mass to the power of three-quarters: a body ten times heavier carries a brain a bit under six times heavier.[4] The exact exponent depends on the method and on the group of animals, so treat it as a rule of thumb rather than a law.

This is the same move as in The Arithmetic of Bigness: before asking whether a number is big, ask what it should be for something of that size. By that standard the whale’s brain is about what a whale needs. The human brain is the outlier: unusually large for a primate of our body size.

But even the corrected number is crude, because a kilogram of brain is not one thing. It matters how many neurons it holds and where they sit.

Where the neurons are

The most careful count of a human brain found about 86 billion neurons. About 16 billion of them are in the cerebral cortex, the folded outer layer most associated with thinking, and about 69 billion in the cerebellum, a smaller structure at the back that coordinates movement and timing.[5]

The same laboratory later counted an African elephant’s brain. It holds about 257 billion neurons, three times as many as ours. But 97.5 per cent of them, some 251 billion, are in the cerebellum. The elephant’s cortex has about 5.6 billion, roughly a third of the human figure. The authors suggest the giant cerebellum is there mainly to control the trunk.[6]

Where the neurons are: human and African elephant Stacked bars of neuron counts. Human: 86 billion in total, 16.3 billion in the cerebral cortex and 69.0 billion in the cerebellum. African elephant: 257 billion in total, 5.6 billion in the cortex and 250.7 billion in the cerebellum. Fewer than one billion are elsewhere in each brain. Cerebral cortexCerebellum 0 50 100 150 200 250 billions of neurons Human Human: cerebral cortex 16.3 billion Human: cerebellum 69.0 billion 86 bn cortex 16.3 bn African elephant African elephant: cerebral cortex 5.6 billion African elephant: cerebellum 250.7 billion 257 bn cortex 5.6 bn
Fig. 1 — Where the neurons are. The elephant has three times as many neurons as a human, but almost all are in the cerebellum; its cerebral cortex has about a third as many as ours. Fewer than one billion neurons sit elsewhere in each brain. One adult male elephant; four adult human men. Data: Azevedo et al. 2009; Herculano-Houzel et al. 2014.

So the elephant’s neurons are spent mainly on fine control of a body part, and ours mainly on the cortex. Neither animal has “more brain” in a simple sense; each spends its brain on different work.

Whales complicate the picture in an instructive way. A study of ten long-finned pilot whales counted about 37 billion neurons in the neocortex, more than twice the human figure. The neurons are less densely packed, and the total comes from a very large brain. The authors themselves say the count does not explain human cognitive abilities.[7] Neuron number is a better guide than brain mass, but it is not the whole answer either.

Birds push from the other side. Parrots and songbirds pack neurons far more densely than mammals: on average twice as many as a primate brain of the same mass, and in large parrots and crows as many forebrain neurons as primates with much larger brains.[8] A raven’s walnut-sized forebrain holds about 1.2 billion neurons. A chimpanzee’s cortex holds about 7.4 billion. On several cognitive tests, ravens and keas perform at the level of great apes anyway.[9]

Built for the job

Look at what brains are specialised for and the variety gets sharper.

A barn owl hunts in the dark by sound. To tell where a mouse is, it compares when a rustle reaches each ear, differences of millionths of a second. In its brainstem, fibres from the two ears enter a nucleus from opposite sides and act as delay lines, and neurons fire most when the delays cancel: a timing circuit built in tissue, matching a mechanism first proposed in theory in 1948.[10]

A horseshoe bat navigates by echoes. About half the neurons in one of its auditory centres are tuned to a narrow band around the frequency of its own echo, an “acoustic fovea” like the sharp centre of our vision. The bat even shifts the pitch of its calls as it flies so that the Doppler-shifted echo lands in that band.[11]

Birds that hide food and find it months later tend to have a larger hippocampus, the structure involved in spatial memory, relative to brain and body size, though how robust that link is remains argued.[12] Scrub jays remember not only where they cached food but what and when: they returned to perishable worms after a short delay and to peanuts after a long one, once the worms would have rotted.[13]

An octopus spreads its brain around its body. It has roughly 500 million neurons,[14] and more of them are in the arms than in the brain.[15] An arm cut off from the brain and stimulated can still produce a nearly normal reaching movement, so part of the motor program runs locally.[16] The central brain chooses and adjusts; the arms do much of the detailed work. One brain and eight very capable arms is a better description than eight minds.

In each case the solution is shaped by the problem: timing for the owl, frequency for the bat, memory for the jay, distributed control for an animal with no skeleton.

Small brains, real tricks

At the other end of the scale, tiny brains do things that once seemed to need large ones.

A honeybee has fewer than a million neurons in a brain of about a cubic millimetre.[17] Bees can learn the abstract rule “same” or “different” and transfer it from colours to smells.[18] Trained to choose the smaller of two numbers of items, they placed an empty display below one without having been trained on it, treating nothing as a quantity.[19]

Portia, a jumping spider that hunts other spiders, will study a route to its prey, then set off on a detour that first leads away from the target and takes it out of sight, and more often than not choose a detour that actually arrives.[20] No one has counted its neurons precisely, but its brain is tiny by any measure.

The complete wiring diagram of an adult fruit fly, published in 2024, maps 139,255 neurons and about 54.5 million synapses.[21] The roundworm C. elegans manages its life with 302 neurons, the first animal whose wiring was mapped completely.[22]

Small brains get by on shortcuts: specialised circuits, built-in assumptions about their world, and behaviour that does part of the computing for them. They are good at a narrow set of things, often extremely good.

Different routes to the same answer

Crows and apes are separated by hundreds of millions of years of evolution, and their brains are organised very differently: birds lack the layered cortex of mammals. Yet both make and use tools, plan and reason about cause and effect. New Caledonian crows used a short stick to retrieve a longer one and then used the long one to reach meat; four of seven succeeded on their first trial.[23] Researchers who reviewed the evidence argued that complex cognition evolved several times “in distantly related species with vastly different brain structures”.[24]

That is the strongest point against equating intelligence with one design or one number. There are several ways to build a clever animal. What they seem to share is enough neurons in the right places, connected in ways that suit the problems the animal faces.

The same questions for models

AI models raise the same questions, and the answers rhyme.

Within one design, size predicts ability very reliably. Train the same kind of model bigger, on more data, with more computing, and its error falls along smooth, predictable curves; these “scaling laws” are why laboratories keep building bigger.[25] Animal brains have no equivalent: you cannot scale up a whale and get a cleverer whale. Here the analogy with brains is weakest, and it is worth saying plainly.

Across designs and training methods, though, size is a poor guide. In 2022 DeepMind showed that most large models had been trained on too little data for their size. Its Chinchilla model, with 70 billion parameters trained on 1.4 trillion tokens, beat its own 280-billion-parameter Gopher, which had used the same computing.[26] Two years later, Meta’s Llama 3 model with 8 billion parameters scored 66.6 per cent on a standard test of knowledge across 57 subjects, where the 175-billion-parameter GPT-3 had scored 43.9 per cent in 2020.[27][28] The comparison is indicative rather than exact: the tests were run differently, and newer models may have seen similar questions in training. Smaller models trained on the outputs of larger ones go further still: a 7-billion-parameter model distilled from DeepSeek-R1 scored 55.5 per cent on a hard mathematics competition where GPT-4o scored 9.3.[29] The parrot beats the larger, less densely wired brain.

Built-in assumptions

Like brains, models come with assumptions built into their structure, and the assumptions decide what they learn easily.

Convolutional networks, for years the standard design for images, look at small patches and reuse the same detectors across the whole picture, so they assume that what matters is local and that an object is the same object wherever it appears. Vision transformers drop those assumptions. Their authors found that they “do not generalize well when trained on insufficient amounts of data”, and only overtook convolutional networks when trained on hundreds of millions of images: “large scale training trumps inductive bias”.[30] The built-in assumptions are worth a great deal of data; enough data can replace them.

AlphaFold 2, which predicts the shape of proteins, was built with knowledge of geometry. One part of it reasons in “triangles” of distances, because three points in space must obey the triangle inequality, and another treats the protein backbone as a set of rigid frames.[31] At the 2020 protein-structure competition its median error was under one ångström, against 2.8 for the next-best method. Then AlphaFold 3, in 2024, removed those geometric assumptions and worked directly on atom positions; its authors note that the cost was occasional invented structure in regions that have none.[32] The owl’s delay lines and the bat’s acoustic fovea are biology’s version of the same choice: build the assumption in, or learn it at a price.

Thinking ahead, or knowing at a glance

Some problems reward search, others recognition. Chess programs show the trade-off clearly.

Stockfish, the strongest conventional chess engine, examines about 60 million positions a second. AlphaZero, which learned chess by playing itself, examined about 60 thousand, a thousand times fewer, and beat it. Its neural network judged positions so well that it needed to look at far fewer of them.[33] The hardware was different, so the comparison is not exact, but the principle stands: a better sense of which moves matter replaces brute search. Stockfish answered in 2020 by adding a small, fast neural network to judge positions, and gained strength immediately.[34]

Portia is the animal version. A small brain cannot compute everything, so it surveys the scene, picks a route and commits, combining a little planning with very good built-in judgement about what is worth attention.

What happens inside

The interior of a model is easier to study than a brain, and what researchers find there is less tidy than the word “experts” suggests.

Many large models use a mixture of experts: many sub-networks, with a router choosing a few for each word. DeepSeek-V3 has 671 billion parameters but uses 37 billion for each token.[35] It sounds like a brain with specialised regions, but when the makers of Mixtral checked, they found no clear assignment of experts by topic: scientific papers, medical abstracts and philosophy were routed much alike. Routing followed syntax and token type more than subject.[36] The specialisation is real but not the specialisation a human would design.

Interpretability research has started to show how models solve particular problems. Anthropic traced how a Claude model adds 36 and 59. It runs two paths in parallel: one estimates roughly, “near 92”, and the other works out precisely that the last digit is 5. Asked how it did the sum, the model described the carry-the-one method taught at school, apparently unaware of its own strategy. Writing rhyming verse, the same model chose the rhyme word before writing the line that ends on it.[37] A small model trained only on addition around a clock of 113 hours did something stranger: it learned to turn numbers into rotations and add the angles, using trigonometric identities nobody had taught it.[38]

These are found strategies, not designed ones, much as the owl’s circuit was found by evolution. And there are hints that different models trained on different data come to represent the world in increasingly similar ways as they grow, though the evidence for that convergence is still preliminary.[39] Two-Way Traffic describes the related finding that models trained for practical tasks end up resembling parts of the brain.

Doing it outside the head

The octopus lets its arms do part of the computing. The crow uses a stick. Models can do the same with tools.

A 6.7-billion-parameter model taught to call a calculator when it needed one scored four to six times higher on three sets of word problems than the same model without it, and well above the 175-billion-parameter GPT-3.[40] Offloading is one of the oldest tricks in biology, and it is the basis of AI agents: How an AI Agent Works describes models that act mainly through the tools their harness supplies. It is also why, as that series argues, an agent’s reach depends on its tools and permissions as much as on the model.

The cost of thinking

The brain does all of this on about 20 watts.[3] A single data-centre chip of the kind used to train and run large models draws up to 700 watts, and a frontier model runs on thousands of them.[41] The comparison is not exact, since the brain’s figure is its whole metabolism, but the gap is several orders of magnitude, and it is one reason researchers keep looking to biology for more efficient designs.

What the comparison teaches

Size matters, in brains and in models, but it is rarely the first thing to ask about. The better questions are where the capacity is spent, what assumptions are built in, what can be offloaded to the body, the environment or a tool, and what the system was shaped to do.

The whale’s 9 kilograms run a whale. The elephant’s 257 billion neurons mostly run a trunk. The bee’s million neurons run a forager that can count to zero. A small, well-trained model can outscore an old giant, and a large model can do arithmetic by a route its own designers did not expect.

The comparison also has limits. Brains learn continuously, run on very little power and come with bodies; models learn in a separate phase, run on enormous power and act through tools. Within one design, a bigger model really is more capable, which is not true of whales. There is more than one way to be clever, in nature and in machines, and the useful question is always which way, for what.

Sources and method

Prepared with Claude under the author’s editorial direction on 8 October 2026. Anthropic makes Claude, and one of the interpretability studies discussed here is Anthropic’s. Facts were checked against the cited papers on 8 October 2026; where a journal page could not be opened, the abstract or a full-text archive was used. Neuron counts come from small samples (one elephant, four human brains, ten pilot whales), and comparisons of benchmark scores across years are indicative, since tests were run differently. The figure is computed by scripts/many_ways_fig.py from the published counts. No new cover image was commissioned.

  1. Guinness World Records, “Heaviest brain”
  2. M. E. Raichle and D. A. Gusnard, “Appraising the brain’s energy budget”, PNAS 99, 2002
  3. V. Balasubramanian, “Brain power”, PNAS 118(32), 2021
  4. R. D. Martin and K. Isler, “The maternal energy hypothesis of brain evolution: an update”, in The Human Brain Evolving, Stone Age Institute Press
  5. F. A. C. Azevedo et al., “Equal numbers of neuronal and nonneuronal cells make the human brain an isometrically scaled-up primate brain”, Journal of Comparative Neurology 513, 2009
  6. S. Herculano-Houzel et al., “The elephant brain in numbers”, Frontiers in Neuroanatomy 8:46, 2014
  7. H. S. Mortensen et al., “Quantitative relationships in delphinid neocortex”, Frontiers in Neuroanatomy 8:132, 2014
  8. S. Olkowicz et al., “Birds have primate-like numbers of neurons in the forebrain”, PNAS 113, 2016
  9. O. Güntürkün, F. Ströckens, D. Scarf and M. Colombo, “Apes, feathered apes, and pigeons: differences and similarities”, Current Opinion in Behavioral Sciences, 2017
  10. C. E. Carr and M. Konishi, “A circuit for detection of interaural time differences in the brain stem of the barn owl”, Journal of Neuroscience 10, 1990
  11. G. Schuller and G. Pollak, “Disproportionate frequency representation in the inferior colliculus of Doppler-compensating greater horseshoe bats”, Journal of Comparative Physiology A 132, 1979
  12. J. R. Krebs et al., “Hippocampal specialization of food-storing birds”, PNAS 86, 1989
  13. N. S. Clayton and A. Dickinson, “Episodic-like memory during cache recovery by scrub jays”, Nature 395, 1998
  14. G. Levy and B. Hochner, “Embodied organization of Octopus vulgaris morphology, vision, and locomotion”, Frontiers in Physiology 8:164, 2017
  15. C. S. Olson, N. G. Schulz and C. W. Ragsdale, “Neuronal segmentation in cephalopod arms”, preprint, 2024
  16. G. Sumbre et al., “Control of octopus arm extension by a peripheral motor program”, Science 293, 2001
  17. BioNumbers: number of neurons in the honeybee brain, from R. Menzel and M. Giurfa, Trends in Cognitive Sciences 5, 2001
  18. M. Giurfa et al., “The concepts of 'sameness’ and 'difference' in an insect”, Nature 410, 2001
  19. S. R. Howard et al., “Numerical ordering of zero in honey bees”, Science 360, 2018
  20. M. S. Tarsitano and R. R. Jackson, “Araneophagic jumping spiders discriminate between detour routes that do and do not lead to prey”, Animal Behaviour 53, 1997
  21. S. Dorkenwald et al. (FlyWire Consortium), “Neuronal wiring diagram of an adult brain”, Nature 634, 2024
  22. J. G. White et al., “The structure of the nervous system of the nematode Caenorhabditis elegans”, Philosophical Transactions of the Royal Society B 314, 1986
  23. A. H. Taylor et al., “Spontaneous metatool use by New Caledonian crows”, Current Biology 17, 2007
  24. N. J. Emery and N. S. Clayton, “The mentality of crows: convergent evolution of intelligence in corvids and apes”, Science 306, 2004
  25. J. Kaplan et al., “Scaling Laws for Neural Language Models”, arXiv 2001.08361, 2020
  26. J. Hoffmann et al., “Training Compute-Optimal Large Language Models”, arXiv 2203.15556, 2022
  27. D. Hendrycks et al., “Measuring Massive Multitask Language Understanding”, ICLR 2021
  28. Meta, Llama 3 model card, April 2024
  29. DeepSeek-AI, DeepSeek-R1 model card and distilled models, January 2025
  30. A. Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, ICLR 2021
  31. J. Jumper et al., “Highly accurate protein structure prediction with AlphaFold”, Nature 596, 2021
  32. J. Abramson et al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3”, Nature, 2024
  33. Google DeepMind, “AlphaZero: Shedding new light on chess, shogi, and Go”, December 2018
  34. Stockfish, “Introducing NNUE Evaluation”, August 2020
  35. DeepSeek-AI, “DeepSeek-V3 Technical Report”, arXiv 2412.19437, 2024
  36. A. Q. Jiang et al., “Mixtral of Experts”, arXiv 2401.04088, 2024
  37. Anthropic, “Tracing the thoughts of a large language model”, March 2025
  38. N. Nanda et al., “Progress measures for grokking via mechanistic interpretability”, ICLR 2023
  39. M. Huh et al., “The Platonic Representation Hypothesis”, ICML 2024
  40. T. Schick et al., “Toolformer: Language Models Can Teach Themselves to Use Tools”, 2023
  41. NVIDIA, H100 Tensor Core GPU specifications

Authored by: Luis Matos Ferreira — Physicist, Developer, Writer

Return to AI, Agents and the Warning Shot for the reading guide.

Comentários

Mensagens populares deste blogue

How Trust Becomes Access

Where The Schooling Went

The Stalled Hour

The Completion