The Frozen Accident






Essay · biology · September 2026

Physics trusts beauty; Crick warned that biology cannot. Yet the genetic code is one in a million, the eye has been invented forty times, and the stripes on a fish obey an equation written in 1952. How much of life is found, and how much is the tape happening to run that way.

The question

In 1988 Francis Crick, who had spent the first half of his career finding the structure of the molecule of heredity and the second half trying to do the same for the brain, wrote a short memoir with a warning in it for physicists. Occam's razor, he said, is a useful tool in physics and a dangerous one in biology. Physicists can reasonably expect the deep laws to be simple and elegant, because they are; biologists cannot, because what they study is the product of a long, blind, opportunistic history that used whatever was to hand. It is very rash, he wrote, to use simplicity and elegance as a guide in biological research.[1]

That is the exact negation of Dirac's rule that it is more important for equations to be beautiful than to fit experiment, and it comes from a man who had done his best work by building a beautiful model. Between the two sentences sits the question a series on this blog has been asking about mathematics and physics: how much of what we find is structure that was there, and how much is a story we told and then could not stop telling. Biology is the field where the story seems to win. Nothing in a cell is the way it is because it had to be; it is the way it is because it happened, and then happened not to be undone. Israel Gelfand is supposed to have said that the only thing more unreasonable than the effectiveness of mathematics in physics is its ineffectiveness in biology, and most biologists nod.

This essay is about where that is true and where it is not. It turns out that the places in biology where mathematics does work are exactly the places where something was found rather than made — and that the map of one is the map of the other.

Contingency

The tape, replayed

Stephen Jay Gould's thought experiment, from Wonderful Life in 1989, has become the standard statement of the “made” position. Rewind the tape of life to the Cambrian, let it run again, and nothing like us appears; the survivors of the Burgess Shale were not the best designs but the lucky ones, and a second run would pick different winners. History, on Gould's account, is not converging on anything. It is one path through a space of paths, and the path is the explanation.[2]

The remarkable thing is that the experiment has been done. In February 1988 Richard Lenski, at Michigan State, put twelve genetically identical populations of Escherichia coli into twelve flasks of a minimal glucose medium and began transferring a sample of each into fresh medium every day. The transfers have continued, with interruptions, ever since; by 2025 the populations had passed eighty thousand generations. Every five hundred generations a sample of each is frozen, so that any earlier point on any of the twelve tapes can be thawed and run again.

For most of the experiment the twelve populations did the same things. They grew larger cells, they grew faster, they lost functions they no longer needed, and they did so by mutations in largely the same genes. Then, around generation 31,500, one population — and only one — began to consume the citrate in the medium, which E. coli cannot do in the presence of oxygen and which had been sitting there unused for fifteen years. Zachary Blount, Christina Borland and Lenski reported it in 2008 and then did the obvious thing: they thawed that population at earlier points and ran it forward, many times. Samples from before generation 20,000 never re-evolved the trait. Samples from after it did so repeatedly. Something had happened between the two — one or more “potentiating” mutations that did nothing visible on their own — without which the later innovation was unreachable. Eleven populations, on the same medium for the same time, never found it.[3]

That is Gould's tape, replayed under laboratory conditions, and it came out Gould's way: the innovation depended on a prior history that was invisible at the time and different in every flask. But the same experiment carries the other result too. The twelve populations converged — on cell size, on growth rate, on which genes to break — far more than they diverged, and where they diverged, the divergence was in what they found, not in what was there to find. The citrate was in every flask. One tape reached it.

Rigidity

One in a million

The first essay in the companion series argued that the complex numbers, although invented, were invented into a space with almost no room in it: keep arithmetic that behaves like arithmetic and there are three division algebras, then nothing. The nearest thing biology has to that fact is the genetic code, and the story is stranger, because the code was not found to be unique. It was found to be excellent.

The code is the table that assigns each of the sixty-four three-letter words of DNA to one of twenty amino acids or to “stop”. It is, with minor variations, the same in every organism on Earth, which is the strongest evidence there is that every organism on Earth is related. Crick's 1968 explanation for its universality was that it is a frozen accident: any change to the code alters every protein in the cell at once, so once a code was in use it could not be improved, and the one we have is simply the one that was in use when the freezing happened. Nothing about the table, on this view, would need to make sense.[4]

Most of it does not. But in 1998 Stephen Freeland and Laurence Hurst asked a question Crick had not: how good is the table at limiting the damage of mistakes? A mutation that changes one letter of a codon changes the amino acid; the damage depends on how different the new amino acid is from the old. They generated a million random codes by shuffling the assignments and scored each on how badly single-letter errors would change the chemistry of the resulting proteins. The natural code was better than all but one of the million.[5]

It is worth being precise about what that score measures, because it is easy to hear it as a claim that the code somehow suppresses mutation, and it is the opposite. Mutations happen at a rate set by the copying machinery — the polymerase, the proof-reading, the repair enzymes — and the code has no say in how many there are. What the code decides is what each one costs. In a table where neighbouring codons encode chemically similar amino acids, a single-letter error usually produces a near-synonym rather than nonsense: the protein is slightly different, not broken. Evolution keeps all its raw material and loses most of its catastrophes. In the language Sewall Wright introduced in 1932, the code is part of the map from genotype to fitness, and a good code makes the landscape smoother — more small survivable steps between one working protein and the next, fewer cliffs.[6] That is not a brake on evolution. It is the condition under which the mostly neutral drift Motoo Kimura described in 1968 can happen at all, and Freeland and his colleagues argued in 2003 that error-tolerance and evolvability are the same property of the code seen from two sides.[7][5]

Distribution of error-damage scores for random genetic codes with the natural code marked A bell-shaped curve of random codes with a marker far to the left, at the low-damage end, labelled as the natural code. less damage more damage a million random codes the natural code better than all but one of them average chemical change caused by a single-letter mutation how many codes score here
Fig. 1 — Freeland and Hurst's experiment, schematically. Shuffle the assignments of codons to amino acids a million times and score each table on how much a single-letter error changes the chemistry of the amino acid; random codes pile up in the middle. The code every organism uses sits in the far tail. The curve is a sketch of the published distribution, not the data.

That is not the same as being unique, and the difference matters for the argument. The space of possible codes is astronomically large, and the natural code occupies one point in the best tail of it. The realist's reading is that selection, over a long period before the freeze, was climbing toward a small region of very good codes and got most of the way there; the code is found in the sense that the good region was there to be found. The historian's reading is that the code's structure is a fossil of the order in which amino acids were added to it — the biosynthetic pathways of related amino acids share codon families — and that the error-robustness is a side effect. Both are probably right. A made thing was pushed, by a blind process, into the found region, and then frozen there.

It has since been unfrozen. In 2013 George Church's group replaced every instance of one stop codon in E. coli with another and deleted the machinery that read it; in 2019 Jason Chin's group in Cambridge rebuilt the entire four-million-letter genome of E. coli with a sixty-one-codon code, three of the sixty-four words removed and reassigned. The bacteria grew.[8] The code is not frozen because it cannot be changed. It is frozen because nothing in a history of blind steps could change it, and a history of deliberate ones can.

Convergence

Forty eyes

If the tape has structure, replaying it should produce the same solutions from different starting points, and it does, so often that the phenomenon has a name and a literature. The eye is the standard case. Luitfried Salvini-Plawen and Ernst Mayr estimated in 1977 that image-forming eyes had evolved independently between forty and sixty-five times across the animal kingdom. The camera eye of a vertebrate and the camera eye of an octopus are built by different developmental routes from different tissues — the vertebrate retina is inside-out, with the photoreceptors behind the wiring, and the octopus retina is the sensible way round — and they arrive at the same design, with a lens, an iris and a focusing mechanism, because the physics of forming an image with light admits very few solutions.[9]

The molecular cases are more surprising. Bats and toothed whales both echolocate, and both do it with modifications to a protein called prestin in the hair cells of the inner ear. When Yang Liu, Ying Li and colleagues sequenced the gene in 2010, the bat and dolphin versions grouped together on a tree of prestin sequences — two lineages that had separated tens of millions of years earlier, and had each acquired the same amino acid changes in the same positions, because those were the changes that worked. A 2013 study of whole genomes found the same signature in around two hundred genes.[10] The crows of the fourth essay in the companion series, with their number neurons in a brain that has no cortex, are another entry in the same catalogue.

Simon Conway Morris, who did the reconstruction of the Burgess Shale animals that Gould's book celebrates, drew the opposite conclusion from Gould, and Life's Solution in 2003 is a four-hundred-page list of convergences offered as evidence that the space of viable designs is small and evolution keeps finding the same points in it.[11] The disagreement between the two is the found-or-made argument with fossils. Gould says the path is the explanation; Conway Morris says the destination is. Lenski's flasks say: both, and the interesting question is which features are which.

The convergences tell you where the landscape has shape. An eye is found because optics constrains it; echolocation is found because the physics of hearing high frequencies constrains it; number sense is found because the statistics of the world constrain it. The inside-out vertebrate retina is made, a historical accident that no engineer would choose and that every vertebrate has carried for half a billion years because the cost of reversing it was never payable. Biology's answer to the question is not one answer. It is a sorting.

Form

Where the equations are

The companion series ended each essay on the same division: the vocabulary is made, the structure is found, and the structure is what mathematics describes. Biology has a version of that division, and it runs between the parts and the form.

D'Arcy Wentworth Thompson's On Growth and Form, from 1917, is the founding text of the found side. Thompson argued that a great deal of biological shape is not the product of selection acting on heredity but of physics acting on matter: cells pack like soap bubbles because surface tension makes them, shells spiral logarithmically because growth at a constant rate on a fixed margin produces that spiral and no other, and one fish's outline can be turned into another's by a smooth distortion of the coordinate grid. Selection chooses among forms; it does not invent the menu.[12]

The sharpest entry on the menu was written by Alan Turing in 1952, in the last years of his life. The Chemical Basis of Morphogenesis asks how a ball of identical cells, in a chemically uniform environment, can break its own symmetry and produce a pattern. Turing's answer is a pair of reactions and two rates of diffusion: an activator that makes more of itself and of an inhibitor, and an inhibitor that spreads faster. Under those conditions a uniform state is unstable and the system falls into spots, stripes or whorls whose spacing is set by the chemistry alone. The paper contains no biology. It is a theorem about which patterns are possible, and it took sixty years to find them in an animal.[13]

Shigeru Kondo and Rihito Asai did it in 1995 with angelfish. Turing's equations predict not just stripes but how stripes on a growing animal should rearrange as the animal gets larger — splitting, inserting, branching in a specific way — and the stripes on Pomacanthus did exactly that, over months, as photographed.[14] In 2012 two groups closed the case for mammals: Rushikesh Sheth and colleagues showed that removing Hox genes from a mouse limb increases the number of digits in the way a Turing mechanism with a shorter wavelength requires, and Andrew Economou and colleagues found the ridges on the roof of the mouse's mouth forming by the same instability, with the two chemicals identified. Selection tunes the wavelength. The physics decides that there will be a wavelength.[15]

“It is thus very rash to use simplicity and elegance as a guide in biological research.” Francis Crick, What Mad Pursuit, 1988

The other place mathematics works in biology is at the scale of whole organisms, and it has been embarrassing theory for ninety years. Max Kleiber found in 1932 that the resting metabolic rate of mammals, from mouse to steer, rises with body mass to the power of three-quarters — not two-thirds, as surface-to-volume reasoning predicts, and not one, as counting cells would predict. The law holds across twenty-one orders of magnitude if you include bacteria and whales, and the exponent's numerator and denominator show up again in heart rate, lifespan, and the branching of blood vessels.[16] In 1997 Geoffrey West, James Brown and Brian Enquist offered a derivation: a body is supplied by a branching network, the network must reach every cell and minimise the energy of transport, and the optimal such network in three dimensions has fractal properties that force the exponent to be three-quarters.[17] The derivation is contested and the exponent is argued over to the second decimal place. But the regularity is not contested, and it is precisely the kind of thing Crick said not to expect: a simple law, from a clean argument, that a mouse and a whale both obey.

Metabolic rate against body mass on logarithmic axes with Kleiber's three-quarters line A log-log chart with seven labelled animals on a straight line of slope 0.75 and a dashed comparison line of slope 0.67. 1 101 102 103 104 105 10 g 100 g 1 kg 10 kg 100 kg 1 t 10 t mouse rat cat dog human cow elephant mass3/4 (Kleiber) mass2/3 (surface law) body mass resting metabolic rate, kcal per day
Fig. 2 — Kleiber's law. On logarithmic axes, resting metabolic rate against body mass is a straight line of slope three-quarters from mouse to elephant, not the two-thirds that surface-to-volume reasoning predicts. The animals are placed on the fitted line (70 · mass3/4 kcal/day); real measurements scatter around it by a few tens of per cent.
Pictures

The helix and the tree

A companion essay on this blog argued that Feynman's diagrams are the clearest case of a picture that became a formalism, and the clearest warning about what a picture can make you believe. Biology has one of each.

The double helix is the first kind. Watson and Crick did not derive the structure of DNA; they built it, out of cardboard and metal, constrained by Rosalind Franklin's diffraction photographs and Erwin Chargaff's base ratios, and the model was the argument. The paper of April 1953 is barely a page, and its famous sentence — it has not escaped our notice that the specific pairing we have postulated immediately suggests a possible copying mechanism — is the picture doing the explaining. Nothing in the chemistry says how heredity works. The shape does: two strands, each the template for the other.[18] It was a structural picture in exactly the sense the fifth essay in the companion series describes, grasped before it was written down and then written down because it had been grasped, and it has never misled anyone, because it is the shape of the thing.

The tree of life is the second kind. Darwin drew one in his notebook in 1837, above the words “I think”, and it is the only figure in The Origin of Species.[19] It organised evolutionary biology for a century and a half: descent with modification, branching, every organism at the tip of a lineage that goes back to one root. Then, in the 1990s, whole genomes started to be sequenced, and the tree stopped being drawable. Bacteria and archaea exchange genes sideways, across lineages, routinely; a bacterial genome is a mosaic of pieces with different histories; and W. Ford Doolittle argued in 1999, in a paper called “Phylogenetic Classification and the Universal Tree”, that for most of life's history there was no tree, only a net.[20] The picture was not wrong the way Kempe's proof was wrong. It was right about animals and plants, which is where it had been drawn, and it had been taken for a picture of everything. Like the Feynman diagram that looks like a movie of a collision, it organised the calculation around a story — one ancestor, branching — and the story was the made part.

Crick's own arrow diagram is the instructive middle case. The central dogma, drawn in 1958 as arrows from DNA to RNA to protein, is the most reproduced picture in molecular biology and is routinely said to have been overturned in 1970 by the discovery of reverse transcriptase, which copies RNA back into DNA. Crick wrote a slightly irritated note to Nature that year pointing out that the picture had said no such thing: his claim had been that information does not flow from protein back to nucleic acid, and the arrow from RNA to DNA had been in his original diagram as a possibility.[4] The picture had been simplified by the people who copied it, and the simplification was what got overturned. The structure survived. The drawing of it did not.

Certificates

Two hundred million answers

The third essay in the companion series argued that proof has split into two jobs, certifying and explaining, and that machines now do the first without the second. Biology has just lived through the same split, on a problem that had been open for fifty years.

A protein is a chain of amino acids that folds into a shape, and the shape is what it does. Cyrus Levinthal pointed out in 1969 that a chain of a hundred residues has more possible conformations than the universe has had time to sample, so the fold cannot be found by search; it must be reached by a route, and predicting the fold from the sequence means understanding the route.[21] Attempts to do that were scored, from 1994, in a biennial competition called CASP, in which groups predict structures that have been solved experimentally but not yet published. For twenty-five years the best predictions were useful for easy cases and poor for hard ones.

In November 2020 DeepMind's AlphaFold 2 entered CASP14 and produced structures for the hard cases that were, in most instances, as accurate as the experimental ones. The paper came out in July 2021; a year later the database it was used to build held predicted structures for some two hundred million proteins — essentially every sequence then known. The protein-folding problem, in the sense in which it had been posed since Levinthal, was over.[22]

Except that nothing about the route had been understood. AlphaFold does not simulate folding; it does not know why a sequence adopts a shape; it has learned, from the roughly two hundred thousand structures that decades of crystallography had produced, a mapping from sequence to shape that it cannot explain and its authors cannot extract. It is the four colour theorem: a certified answer with no explanation in it, the most useful object in the field, and a proof of nothing about mechanism. The explanation job — how the chain actually moves, why the physics picks that basin — is exactly where it was in 2019, and researchers who work on it now do so with two hundred million certified answers to check against. Whether that is what the machine-assisted mathematics of the third essay will look like is not a rhetorical question. It is what happened.

Where it lands

Made by history, found by physics

Biology's answer to the question the companion series asked is not that life is found or that it is made. It is a sorting, and the sorting has a principle.

The parts are made. The genetic code, the inside-out retina, the mosaic bacterial genome, the fifty proteins of the ribosome, the particular potentiating mutation in one flask at Michigan State: these are history, blind and unrepeatable, and Crick's warning about them is correct. Nobody will derive the ribosome from a principle of elegance, because it was not built by one. The Complete Parts List essay on this blog argued that biology can enumerate almost everything and predict almost nothing, and that is true of the parts.

The forms are found. The eye, the stripe, the spiral shell, the three-quarters exponent, the small good region of code space, the double helix: these are what physics and optimisation allow, and history keeps arriving at them from different directions because there is nowhere else to arrive. Turing's equations knew about the angelfish forty years before the angelfish was photographed; Kleiber's law was true of whales before anyone weighed one. Where biology has equations, it has them because the thing described was found. Gelfand's remark is right about the parts and wrong about the forms, and the line between the two is the line the whole series has been drawing.

What the sorting explains is something the physics essays could not. In mathematics the constraints were mysteriously tight and nobody could say why. In biology the tightness varies, and one can see what sets it: the more physics there is in a feature, the fewer ways there are to do it, and the more history there is, the more. Convergence measures the first; contingency the second; and Lenski's flasks, with their eleven populations that never found the citrate and their twelve that all grew larger cells, are the cleanest available measurement of both at once. The tape, replayed, does not tell you whether life is found or made. It tells you which parts are which.

Open threads

Where this could go

The net, drawn properly. Doolittle's argument that the tree of life is a net was a claim about pictures. The picture that replaces it exists: the ancestral recombination graph, which records every branching and every sideways exchange in a set of genomes as a single structure, and which is now being inferred from hundreds of thousands of human genomes. It is the biological analogue of the shape under the Feynman diagrams — the found object under the made picture — and a piece explaining what it is and why it is hard to compute would be the natural sequel to the “pictures” section here.

Code space, mapped. Freeland and Hurst sampled a million codes. The space has something like 1084 members, and the modern question is not whether the natural code is good but what the landscape of codes looks like — how many good regions, how connected, whether a blind process starting anywhere would find one. The synthetic recoded organisms are the experimental instrument, and the question is answerable in a way it was not in 1998.

Turing, quantitatively. The 2012 mouse results establish the mechanism; nobody has yet measured, across many species, how much of visible pattern is set by reaction–diffusion physics and how much by selection tuning its parameters. A survey of where the wavelengths come from — fish, cats, seashells, fingerprints — would put a number on the found-versus-made ratio for one class of feature.

The ribosome as a frozen accident. If the code is the frozen accident that turned out to be near-optimal, the ribosome is the one that did not: a machine of fifty proteins and three RNAs whose core is a relic of a world before proteins. It is the best case for Crick's side of the argument, and it deserves the treatment the code got here.

AlphaFold, the explanation job. The certificates exist. Whether anyone can now extract a theory of folding from two hundred million predicted structures — or whether, as with the four colour theorem, the certified answer simply sits there while the mechanism stays open — is the live test of the third companion essay's worry, and it will be settled within a few years.

Where scaling laws stop. Kleiber's exponent is disputed at the second decimal, and the West–Brown–Enquist derivation more so. A piece on what the metabolic scaling controversy actually turns on, and what an exponent of 0.75 versus 0.67 would mean for the found-versus-made argument, would be the biology essay in which mathematics is argued about the way physicists argue.

On method and tools

This piece was written collaboratively with Claude Fable 5.1 (Anthropic): human specification, editorial direction and critical review; machine synthesis, drafting and figure generation. It is a companion to the Found or Made series and was produced the same way.

The two figures are schematic. Fig. 1 sketches the shape of Freeland and Hurst's distribution rather than reproducing their histogram; Fig. 2 places the animals on Kleiber's fitted line rather than at measured values, and says so in the caption. Both are drawings of a result, not the result.

Where a claim is a demonstrated finding it is cited to the primary paper. Where it is contested, that is said in the text rather than left to the reader: the prime-cycle explanation of the cicadas, the West–Brown–Enquist derivation of the three-quarters exponent, and the reading of the genetic code's error-tolerance as evolvability are argued positions, not settled ones. The Lenski generation count is approximate and should be checked against the experiment's own site before being quoted.

Nothing here was verified against a laboratory. The essay's argument — that biology's forms are found where physics constrains them and its parts are made where history does — is an interpretation of the cited work, and the people who did that work may not share it.

Authored by: Luis Matos Ferreira — Physicist, Developer, Writer

Related essays on this blog
  1. The Complete Parts List — open problems in biology, and why an inventory is not a theory.
  2. Invented, Then Unavoidable — Part I of Found or Made: imaginary numbers, rigidity, and beauty.
  3. What a Proof Is Now — Part III: certificates, explanations, and proofs no one has read.
  4. Where Number Comes From — Part IV: the number sense, crows, and the world doing mathematics.
  5. The Diagram That Did the Sum — Feynman's diagrams, and the shape underneath them.
Sources
  1. Crick, What Mad Pursuit, 1988.
  2. Gould, Wonderful Life, 1989.
  3. Blount, Borland & Lenski, “Historical contingency and the evolution of a key innovation in an experimental population of Escherichia coli”, PNAS, 2008.
  4. Crick, “On protein synthesis”, 1958; “The origin of the genetic code”, Journal of Molecular Biology, 1968; “Central dogma of molecular biology”, Nature, 1970.
  5. Freeland & Hurst, “The genetic code is one in a million”, Journal of Molecular Evolution, 1998; Freeland, Wu & Keulmann, “The case for an error minimizing standard genetic code”, Origins of Life and Evolution of Biospheres, 2003.
  6. Wright, “The roles of mutation, inbreeding, crossbreeding and selection in evolution”, 1932.
  7. Kimura, “Evolutionary rate at the molecular level”, Nature, 1968.
  8. Lajoie et al., “Genomically recoded organisms expand biological functions”, Science, 2013; Fredens et al., “Total synthesis of Escherichia coli with a recoded genome”, Nature, 2019.
  9. Salvini-Plawen & Mayr, “On the evolution of photoreceptors and eyes”, Evolutionary Biology, 1977.
  10. Liu, Cotton, Shen et al., “Convergent sequence evolution between echolocating bats and dolphins”, Current Biology, 2010; Parker et al., “Genome-wide signatures of convergent evolution in echolocating mammals”, Nature, 2013.
  11. Conway Morris, Life's Solution, 2003.
  12. Thompson, On Growth and Form, 1917.
  13. Turing, “The chemical basis of morphogenesis”, Philosophical Transactions of the Royal Society B, 1952.
  14. Kondo & Asai, “A reaction–diffusion wave on the skin of the marine angelfish Pomacanthus”, Nature, 1995.
  15. Sheth et al., “Hox genes regulate digit patterning by controlling the wavelength of a Turing-type mechanism”, Science, 2012; Economou et al., “Periodic stripe formation by a Turing mechanism operating at growth zones in the mammalian palate”, Nature Genetics, 2012.
  16. Kleiber, “Body size and metabolism”, Hilgardia, 1932.
  17. West, Brown & Enquist, “A general model for the origin of allometric scaling laws in biology”, Science, 1997.
  18. Watson & Crick, “Molecular structure of nucleic acids”, Nature, 1953.
  19. Darwin, Notebook B, 1837; On the Origin of Species, 1859.
  20. Doolittle, “Phylogenetic classification and the universal tree”, Science, 1999.
  21. Levinthal, “How to fold graciously”, 1969.
  22. Jumper et al., “Highly accurate protein structure prediction with AlphaFold”, Nature, 2021; Varadi et al., AlphaFold Protein Structure Database, 2022.

Comentários

Mensagens populares deste blogue

Provas Insanas - Westfield Sydney to Melbourne Ultramarathon 1983

Manuel das Corridas Atleta vs Manuel das Corridas Dirigente Associativo

ITRA Performance Index - Everything You Always Wanted to Know But Were Afraid to Ask

The Ministry of Doubt