The Second Derivative, Part IV: The Grid and the Cloud
The Grid and the Cloud
Three essays argued about what the climate is doing. None of them explained where the numbers come from — or how much of each one is instrument and how much is machinery.
Here is the fact that organises everything else in this essay, and it is the one most people are never told.
A climate model cannot see a cloud. Not because the physics is unknown — cloud physics is understood in considerable detail — but because of arithmetic. The models used for the last international assessment divide the atmosphere into boxes roughly 100 kilometres on a side. A cumulus cloud is about one kilometre across. The cloud is not merely smaller than the box; it is smaller than the box by a factor of a hundred in each horizontal direction, which means about ten thousand of them could sit inside a single grid cell, and the model has exactly one number to describe all of them at once.
This is not a scandal, and it is not a secret. It is a consequence of computing budgets and the fact that the atmosphere has structure at every scale from the planetary down to the millimetre. But it explains a remarkable amount: why models disagree about how much warming a doubling of CO&sub2; produces, why some of them run too hot, why the observed energy imbalance in Part II’s Figure 2 can exceed the modelled one, and why the single biggest engineering push in the field right now is simply to make the boxes smaller.
I. There is no such thing as “the climate model”
The phrase “climate models say” hides a ladder of quite different tools, built for different jobs, with different failure modes. Almost every public argument about models is really an argument about which rung someone is standing on.
II. What is actually inside an Earth system model
Strip away the acronyms and a climate model is a fluid-dynamics program. It solves the Navier–Stokes equations for a thin layer of gas on a rotating sphere, coupled to an ocean doing the same thing with a denser fluid, plus sea ice, a land surface, and increasingly the carbon and nitrogen cycles. It steps forward roughly every twenty minutes of simulated time and it does that for a century or more.
The useful distinction is not between what the model does and does not include. It is between what it resolves and what it parameterises.
| Resolved — computed from first principles | Parameterised — represented by rules |
|---|---|
| Large-scale circulation: jet streams, Hadley cells, storm tracks | Cloud formation, and cloud reflectivity |
| Ocean basin-scale currents and gyres | Deep convection — thunderstorms |
| Radiative transfer through the atmospheric column | Ocean mixing by eddies below grid scale |
| Conservation of mass, momentum and energy | Aerosol–cloud interaction |
| Seasonal cycle, and the response to orbital and solar forcing | Land-surface exchange, vegetation response |
Read the right-hand column again and notice something: it contains almost every process that has ever caused an argument about climate sensitivity. Clouds are the single largest source of spread between models, and they are on the wrong side of the line for the simple reason given in Figure 1.
The one thing models are structurally good at
The left-hand column is dominated by conservation laws, and conservation laws are not approximations. A model can be wrong about where the rain falls in the Sahel and still be trustworthy about the planetary energy budget, because energy conservation is enforced rather than estimated. This is why the strongest claims in Part II — the energy imbalance, ocean heat uptake — are also the ones where models and observations have least room to disagree in principle, and why the disagreement in Figure 2 of that essay is genuinely interesting rather than routine.
III. Tuning, stated plainly
The most common objection to climate models is that they are simply fitted to the observed warming and therefore prove nothing. It deserves a straight answer, and the straight answer is: partly true, and less damaging than it sounds — but the reason it is less damaging is not obvious and is rarely explained.
Tuning is real. Every model has free parameters inside its parameterisations — a cloud droplet fall speed, a mixing coefficient — that are not fixed by theory, and modelling groups adjust them so the model produces a realistic planet. Frédéric Hourdin and colleagues wrote the paper that dragged this into the open, arguing that tuning had been treated as “an unavoidable but dirty part of climate modelling, more engineering than science,” and that it should be documented rather than hidden.[1]
Two facts from that same work blunt the objection considerably. First, roughly half the modelling groups surveyed deliberately exclude the twentieth-century warming trend from their tuning targets, precisely so that reproducing it remains an out-of-sample test rather than a fitted result. Second, and more persuasive: the spread of simulated twentieth-century trends across the models is too wide to be consistent with everyone having tuned to match the observations. If they had all been fitted to the answer, they would agree with the answer, and with each other. They do not.[1]
If the models had all been fitted to the observed warming, they would agree with it. Their disagreement is the evidence that they were not.
IV. The hot model problem, and the thing nobody knows about the IPCC
Now the part that surprised me most while writing this, and which I suspect will surprise most readers.
When the current generation of models was assembled, a subset came back distinctly more sensitive than their predecessors. About 18–20% of CMIP6 models produce a warming above 5 °C for a doubling of CO&sub2;, and roughly a fifth fall outside the range the IPCC assesses as very likely.[2] These became known as the hot models. The cause is traced largely to changes in how the newer models handle cloud microphysics — the right-hand column of that table again.
Here is the consequence, and it is the single most useful thing in this essay for reading climate journalism: the IPCC does not average its models. It never simply took the ensemble mean and published it. For its sensitivity assessment it combined process understanding, the historical instrumental record, palaeoclimate evidence and emergent constraints, and arrived at a range narrower than the models produce, deliberately excluding the tails the models generate.[3]
Hausfather and colleagues argued the rest of the field should follow: stop treating the model ensemble as a democracy in which every model gets a vote, and weight them by how well they reproduce independent evidence.[2] Screening the ensemble by transient response discards about 40% of the models; screening by sensitivity discards about 60%.[4]
Why this cuts against both sides of the usual argument
Against the sceptical reading: the models being wide is not evidence they are useless, because the assessed number does not come from the models alone. Palaeoclimate and the instrumental record are doing real work, and they pull the range in.
Against the alarmed reading: the scariest numbers you see quoted — 5 °C and up per doubling — come from precisely the models the assessment process down-weighted. Quoting them as “what the models say” is technically true and materially misleading.
V. Two frontiers, one promising and one not yet
Making the boxes smaller
The direct assault on Figure 1 is simply to compute more. European efforts — nextGEMS, EERIE, and the Destination Earth digital twin — now run global models at roughly 5–10 kilometres, against the 100 or so of the standard generation.[5][6] Thirty-year simulations covering 2020–2049 have been produced at that resolution, at a throughput of roughly 500 simulated days per day on a large supercomputer.[5]
The reason this matters is specific rather than general. Somewhere between one and nine kilometres, deep convective systems — thunderstorms — start to be resolved rather than approximated, which means the deep convection parameterisation can be switched off entirely.[6] An entire rule-of-thumb is replaced by computed physics. That is not a refinement; it is the removal of one of the two or three largest sources of structural uncertainty in the field.
Letting a neural network do it
The other frontier is machine learning, and here I want to be careful, because it is the most over-claimed area in the field right now.
For weather, the results are genuine and startling: learned models now match or beat traditional forecasting systems at a fraction of the computational cost. For climate, the situation is different in a way that is easy to state and hard to fix. A learned model is fitted to the past. A climate projection is a question about conditions that have not occurred. Machine-learning systems extrapolate badly outside the range they were trained on, and reviews of the field are consistent about it: these models fail when exposed to thermodynamic or radiative conditions substantially outside their training distribution.[7] That is not an incidental flaw to be engineered away; it is a description of what fitting a function to data does.
The trap, stated as plainly as I can
A machine-learning climate model can reproduce the twentieth century beautifully and still be worthless for the twenty-first, because reproducing the twentieth century is what it was optimised to do. A physics-based model that reproduces the twentieth century has done something harder, since it was constrained by conservation laws it could not bend.
This is the same distinction the whole series has been drawing: fitting is not the same as explaining, and a curve that goes through the points is not automatically a curve that continues correctly.
The promising middle ground is hybrid: physics for the resolved dynamics, learned components for the parameterisations, which is roughly what systems like NeuralGCM attempt. Work on making learned parameterisations “climate-invariant” — by feeding them variables rescaled so the relationships hold under a different climate — is the most serious attempt to solve the extrapolation problem rather than route around it.[7] It is not solved yet.
VI. What is being asked of the next generation
CMIP7 is under way. Its scenario design was published in April 2026, simulations began during 2026, and the results feed the next IPCC assessment expected in 2028–29.[8] Two changes are worth recording.
The first is that the highest scenario, the one usually labelled RCP8.5 or SSP5-8.5 and endlessly misused in the press as “business as usual,” has been retired. The new set spans roughly 1.5 to 3.5 °C by 2100.[8] This is quiet but consequential: a great deal of alarming journalism over fifteen years rested on a scenario the modelling community no longer considers plausible enough to include.
The second is that tipping points are now one of four questions explicitly motivating the whole exercise, alongside sea-surface-temperature patterns, changing weather, and the water–carbon–climate nexus.[9] Part III spent an entire essay on the fact that thresholds are the worst-constrained thing in climate science. The modelling community has evidently reached the same conclusion and made it a headline target for the next decade of work.
VII. What models are actually for
The popular picture is that climate models are prediction machines: you set them running and they tell you the future. That is not what they are, and holding the wrong picture makes both trusting and distrusting them incoherent.
A climate model is a consistency engine. It answers: given the physics we believe we understand, what follows? Its value lies in the fact that it cannot be argued with once it is running. You cannot wish away energy conservation in a simulation the way you can in a paragraph.
Which leads to the deeper reason models are unavoidable rather than merely useful. There is no control planet. Every attribution statement — including the 0.27 °C per decade of human-caused warming in Part II’s Figure 9 — is a comparison between the world we observe and a world without our emissions. That second world does not exist and cannot be measured. It exists only inside a model. Anyone who rejects model-based reasoning outright is not being cautious; they are declining to answer the question at all.
The tests models were not allowed to fail
The persuasive evidence is not that models fit the past — that could be tuning. It is the cases where they got something right that they were not fitted to:
- Old projections, checked later. Hausfather and colleagues took climate projections published between 1970 and 2007 and scored them against what subsequently happened. Most were skillful: the majority projected warming statistically indistinguishable from observations, once you correct for the fact that actual emissions differed from what was assumed.[10] These were genuine out-of-sample forecasts, made with far cruder models than today’s.
- Pinatubo. The 1991 eruption gave the field an unplanned controlled experiment. Models predicted the magnitude and duration of the resulting global cooling in advance, from radiative physics.
- The stratospheric fingerprint. Greenhouse warming makes the lower atmosphere warm while the stratosphere cools. Solar forcing would warm both. The observed pattern is the greenhouse one — a signature no competing explanation reproduces, and one nobody tuned for.
Where measurement and model interleave — a worked example
Part II claimed Earth’s energy imbalance is confirmed by two independent methods. The truth is more layered, and more interesting. CERES measures radiation at the top of the atmosphere superbly in relative terms, but cannot resolve a one-watt residual between two fluxes of about 340 watts on absolute calibration alone. So a one-time offset is applied to the satellite product so that its mean over July 2005–June 2015 matches the in-situ ocean-heat estimate of 0.71 W/m².[11]
Crucially, that offset does not touch the trend or the year-to-year variability, which remain independently satellite-derived. So the level is anchored to Argo and the change is not. Since the entire acceleration argument is about the change, the claim survives — but “two independent instruments agree” is a simplification of something that deserves the extra sentence. I have added it to Part II.
VIII. Where models have been wrong
Every previous essay in this series had a section arguing against itself, and it would be poor form to stop now. Here is the honest ledger — and its shape is not what either side of the public argument expects.
The individual entries, briefly. Arctic summer sea ice declined faster than the ensembles projected, a mismatch documented as far back as 2007 under the title “Arctic sea ice decline: faster than forecast.”[12] Ice-sheet dynamics were so poorly represented that earlier assessments explicitly excluded rapid dynamic loss from their sea-level projections; observations then tracked the upper end. The energy imbalance is running at roughly twice the modelled value.[13] The AMOC entry is the subtle one: many models carry a bias in a freshwater-transport diagnostic that makes their overturning too stable, which is precisely the argument van Westen and colleagues use to justify a physics-based indicator over the models’ own behaviour.[14] And the Antarctic regime shift of 2016 was not projected by anyone.[15]
There is a pattern here worth naming. Models are built to be numerically stable over long integrations and are evaluated against a historical record that contains no abrupt transitions. It would be surprising if such tools were not biased towards smooth, well-behaved futures. That is a structural conservatism, not a conspiracy — and it means the errors are more likely to be pleasant than unpleasant.
IX. What would change my mind
- Convergence at kilometre scale. If several km-scale models, with deep convection resolved rather than parameterised, converge on a narrower cloud feedback, the largest uncertainty in the field is genuinely shrinking. If they resolve convection and still disagree, the problem was never really resolution and much of the current optimism is misplaced.
- Sensitivity, finally narrowing. The plausible range for climate sensitivity has been roughly 1.5–4.5 °C since 1979, tightened only modestly in AR6. If CMIP7 plus new observational constraints does not narrow it further, that is a substantive fact about the limits of the approach, not a delay.
- Machine learning, tested honestly. Train a learned climate model on data ending in 2000, then score it against 2000–2025 without further tuning. Until that test is passed publicly and repeatedly, learned climate projections should be treated as research, not evidence.
- The imbalance gap. If newer models reproduce the observed energy imbalance without being retuned to it, the divergence flagged in Part II was a model deficiency now corrected. If the gap persists, it is telling us something about the real planet.
X. The bottom line, for all four parts
The series began with a viral graphic and a suspicion. It ends somewhere less dramatic and, I think, more useful.
Every number in these four essays is a mixture of instrument and inference, and the single most valuable habit is asking what the ratio is. The CO&sub2; growth rate is almost pure instrument. The carbon budget is almost pure inference. Sea level and ocean heat sit in between, closer to the instrument end. Tipping thresholds are almost pure inference with the widest ranges in the entire field — which is exactly what Part III found by a completely different route.
The question is never “do you believe the models?” It is “how much of this particular number is model?”
And the conclusion that survives all four essays is not a model output at all. The energy imbalance is measured. The ocean heat is measured. The sea-level acceleration is measured. Glacier loss is measured, by people with stakes and radar. None of that requires you to trust a simulation. The models are needed for the harder questions — how much of it is us, what happens next, where the thresholds sit — and on those they give ranges rather than answers, which is the honest thing for them to do.
What the models add, in the end, is not precision. It is direction, and near-unanimity about it, across four generations of increasingly independent construction, by groups who compete with each other and would very much enjoy proving one another wrong. That agreement is not proof. But it is the strongest thing we have, and waiting for something stronger has a cost that compounds — which is where these essays came in.
References
- Hourdin, F. et al. (2017). “The Art and Science of Climate Model Tuning.” Bulletin of the American Meteorological Society 98, 589–602. Source of the survey finding that roughly half of modelling groups withhold twentieth-century trends from tuning. ametsoc.org
- Hausfather, Z., Marvel, K., Schmidt, G. A., Nielsen-Gammon, J. W. & Zelinka, M. (2022). “Climate simulations: recognize the ‘hot model’ problem.” Nature 605, 26–29. nature.com
- IPCC (2021). Sixth Assessment Report, Working Group I: The Physical Science Basis. Assessed likely ranges: 2.5–4.0 °C equilibrium climate sensitivity, 1.4–2.2 °C transient climate response, from multiple lines of evidence. ipcc.ch
- Carbon Brief (2022). “Guest post: How climate scientists should handle ‘hot models’.” On screening by transient response versus sensitivity, and how much of the ensemble each discards. carbonbrief.org
- nextGEMS consortium (2025). “nextGEMS: entering the era of kilometer-scale Earth system modeling.” Geoscientific Model Development 18, 7735. gmd.copernicus.org
- “The Destination Earth digital twin for climate change adaptation.” Geoscientific Model Development 19, 2821 (2026). ICON, IFS-NEMO and IFS-FESOM at 5–10 km; deep convection partially resolved below ~9 km. gmd.copernicus.org
- On the limits of learned models: “Imitation or identification: limitations of deep learning in extrapolating to future climate–carbon cycle change” (2025); and Beucler, T. et al., “Climate-invariant machine learning,” Science Advances (2024), on rescaling inputs so learned parameterisations transfer across climates. science.org
- “The Scenario Model Intercomparison Project for CMIP7 (ScenarioMIP-CMIP7).” Geoscientific Model Development 19, 2627 (2026). Published April 2026; SSP5-8.5 retired; scenarios span roughly 1.5–3.5 °C by 2100. gmd.copernicus.org
- WCRP (2025–26). CMIP7 design and Fast Track, Geoscientific Model Development 18, 6671 (2025), and the CMIP7 scenario explainer. Tipping points named as one of four motivating research questions. wcrp-cmip.org
- Hausfather, Z., Drake, H. F., Abbott, T. & Schmidt, G. A. (2020). “Evaluating the Performance of Past Climate Model Projections.” Geophysical Research Letters 47. Projections from 1970–2007 scored against subsequent observations. agupubs.onlinelibrary.wiley.com
- Loeb, N. G. et al. (2018). “Clouds and the Earth’s Radiant Energy System (CERES) Energy Balanced and Filled (EBAF) Top-of-Atmosphere (TOA) Edition-4.0 Data Product.” Journal of Climate 31, 895–918. Source of the one-time offset anchoring net TOA flux to the in-situ value of 0.71 W/m² for July 2005–June 2015. ametsoc.org
- Stroeve, J. et al. (2007). “Arctic sea ice decline: Faster than forecast.” Geophysical Research Letters 34, L09501. agupubs.onlinelibrary.wiley.com
- Mauritsen, T. et al. (2025). “Earth’s Energy Imbalance More Than Doubled in Recent Decades.” AGU Advances. agupubs.onlinelibrary.wiley.com
- van Westen, R. M., Kliphuis, M. & Dijkstra, H. A. (2024). “Physics-based early warning signal shows that AMOC is on tipping course.” Science Advances 10, eadk1189. science.org
- Hobbs, W. et al. (2024). “Observational Evidence for a Regime Shift in Summer Antarctic Sea Ice.” Journal of Climate 37, 2263–2275. ametsoc.org
This article was researched and written collaboratively with Claude (Anthropic): human specification and critical review, machine research synthesis and drafting, iterative refinement through structured dialogue. The research phase covered the CMIP7 design papers, the kilometre-scale modelling literature, the machine-learning-for-climate literature and its critics, the model-tuning literature, and the CERES EBAF data documentation.
Five hand-built SVG figures. Figure 1 is drawn to scale in the sense that matters — the cloud really is that small relative to the grid cell. Figures 2 and 5 are qualitative: positions and bar lengths convey relationships, not measurements, and both captions say so. Figures 3 and 4 use published values with approximate ensemble bounds, flagged in place.
This essay exists because a reader asked what kind of modelling underpins the first three, and because the answer exposed a simplification in Part II worth correcting rather than leaving. That correction has been made in Part II and is explained in Section VII here.
Physicist & Developer
Comentários
Enviar um comentário