Small errors grow until they swallow the forecast
The atmosphere is a chaotic fluid. Two almost identical starting states drift apart over time, which means even a perfect model would eventually go wrong if its starting picture was slightly off.
This simple relationship explains the whole history of forecasting. If errors double every day and a half, halving the starting error adds a day and a half of useful forecast. A tenfold improvement adds about five days. Nothing short of a perfect description of every gust and eddy could extend it indefinitely, and that is impossible. Worse, small features like thunderstorms grow their errors in hours, and those errors cascade upwards to larger weather systems. That cascade is why theory puts a ceiling of about two weeks on forecasting day-to-day weather in the mid-latitudes, however good the models become.
The same arithmetic is encouraging, too. Today’s forecasts still sit several days short of that ceiling, so there is real room left, and every gain in observations and assimilation converts directly into usable days.
A thunderstorm, a weather system and a season have different clocks
There is no single expiry date for a forecast. A local thunderstorm is sensitive to small-scale details that become uncertain within hours. The track of a large pressure system remains useful for days. Slower signals from the ocean, soil moisture and large-scale circulation can support probabilities for wetter or warmer weeks beyond the familiar two-week horizon. ECMWF explicitly distinguishes these scales. A forecast of increased seasonal flood risk can be useful without predicting which street will flood on a particular afternoon. ECMWF’s forecast error guide; ECMWF’s forecast skill horizon study
The logarithm in the model explains diminishing returns. If errors double every 1.5 days, making the starting error ten times smaller buys 1.5 × log₂(10), or 4.98 days. It does not buy ten times the forecast range. This is an illustration of error growth, not a universal law: doubling times depend on scale and flow, while errors generated inside the model also matter. Near the intrinsic limit, small-scale uncertainty spreads upward even when the initial large-scale picture is unusually accurate. Selz et al., Journal of the Atmospheric Sciences, 2022
Roughly two weeks describes detailed midlatitude weather under particular skill measures. It is not a ceiling on all probabilistic, subseasonal or seasonal information.
Why does better data buy only a few more days?
Small errors in today’s weather double every day or two until the forecast is no better than guessing from the climate. Each halving of the starting error therefore buys one more doubling time, never more.
6.4 error doublings before the forecast stops being useful
- Gain from halving the starting error
- +1.5 days
- Gain from a tenfold better starting state
- +5.0 days
- Starting error, share of saturation
- 0.7%
The doubling times and starting errors are illustrative, chosen so the presets reproduce published behaviour: about ten days of useful skill today, roughly four to five fewer in the 1980s, and up to five more days with a tenfold reduction in starting error, as Zhang and colleagues estimated in 2019. Thunderstorms grow errors in hours, which is why they are predictable for hours, not days.
Calculation & assumptions
Useful lead time = doubling time × log₂(usefulness threshold ÷ starting error). Errors are expressed as a share of saturation, the size of error a forecast would have if it knew nothing but the climate. Real error growth is not perfectly exponential: small scales saturate fast and pass their errors up to larger scales, which is what sets the roughly two-week intrinsic limit for mid-latitude weather. The model ignores model error, which better physics and machine learning reduce independently of observations.
A forecast is a chain, and value is lost at every link
A forecast only matters if it changes a decision. That requires every link, from a weather balloon to an evacuation order, to work.
Observations
Satellites, weather balloons, aircraft, ships, buoys, radar and ground stations measure the atmosphere as it is right now.
- Measure
- Observations assimilated per day · coverage gaps
- Failure boundary
- Where few observations exist, as over much of Africa, every forecast starts from a blurrier picture.
Where the frontier moves
Funding sustained surface and balloon networks in the countries with the largest gaps.
Data assimilation
Millions of scattered, imperfect measurements are combined with a short forecast to produce the best estimate of the whole atmosphere’s current state.
- Measure
- Error of the starting analysis
- Failure boundary
- A forecast can never be better than its starting point; errors here grow fastest.
Where the frontier moves
Machine-learning assimilation that uses more satellite data, faster.
The forecast model
Physics-based models solve the equations of fluid motion on a grid; AI models learn the same evolution from decades of past weather.
- Measure
- Error at each lead time · compute per forecast
- Failure boundary
- AI models are trained on past climate and may handle unprecedented extremes less reliably.
Where the frontier moves
AI models that match or beat physics models at a thousandth of the energy, freeing compute for larger ensembles.
Ensembles
Running many slightly different forecasts shows the range of possible outcomes and how confident the forecast is.
- Measure
- Members per forecast · reliability of stated probabilities
- Failure boundary
- A single forecast hides uncertainty; a poorly spread ensemble gives false confidence.
Where the frontier moves
Cheap AI ensembles with hundreds of members, improving the odds attached to rare extremes.
Warnings and communication
National weather services turn forecasts into warnings that people understand and trust.
- Measure
- Share of population covered by early warnings
- Failure boundary
- A perfect forecast that never reaches a farmer or a fishing village saves no one.
Where the frontier moves
Mobile alerts, impact-based warnings and local-language communication.
Action
Evacuations, harvest timing, grid planning and flight routing: the decisions that turn a forecast into value.
- Measure
- Lives and losses avoided per warning
- Failure boundary
- Warnings that are too frequent, or wrong too often, train people to ignore them.
Where the frontier moves
Pre-agreed actions and finance that trigger automatically when a forecast crosses a threshold.
A forecast that runs on one chip in minutes
Between 2022 and 2025, machine-learning weather models went from research curiosity to operational service.
What changed
Physics-based models need a supercomputer for hours. AI models, once trained, produce a global forecast in minutes on a single processor. They learned from ERA5, a reconstruction of the atmosphere back to 1940, built by combining historical observations with a physics model. That dependence matters: the AI models are only as good as that training data, they still need a physics-based analysis to start from, and they have less experience of extremes the climate has not produced before. The likely future is not AI replacing physics but the two working together, with the savings spent on far larger ensembles.
Hurricanes show the gains in human terms
In 2024 the U.S. National Hurricane Center’s track forecasts were the most accurate in its history at every lead time from 12 hours to five days. Its intensity forecasts have improved more slowly: storms that strengthen rapidly are still hard to anticipate, though forecasters’ average underestimate of such storms fell from 26 knots in 2010–14 to 16 knots in 2020–24. Better track forecasts mean smaller evacuation zones, fewer unnecessary evacuations, and more trust when an evacuation is ordered.
The forecast exists; the warning often does not
The world’s best forecasts are now available almost everywhere. The weakest links are observations in poorer countries and the systems that turn forecasts into warnings people act on.
Africa’s surface weather stations are about an eighth as dense as the World Meteorological Organization recommends, and more stations report from Germany than from the whole African continent. That starves every model of data over the regions where the population is most exposed to droughts, floods and heat. The UN-backed Systematic Observations Financing Facility aims to fix part of this, supporting the least developed countries and small island states to meet basic observing standards, which would multiply internationally shared balloon data more than tenfold.
Warnings are improving fast. In 2025, 119 countries, 60% of the total, reported having a multi-hazard early warning system, more than double the number ten years earlier. The WMO estimates that investment in early warnings could deliver benefits of at least $162 billion a year, about ten times their cost, and that better forecasts and warnings could save around 23,000 lives a year.
Cheap forecasts still need expensive observations and trusted warnings
ECMWF put its deterministic AIFS into operational service on 25 February 2025, reporting roughly a thousandfold reduction in energy for generating a forecast. That boundary matters: it excludes the entire observing network, preparation of the initial atmospheric analysis and training the model. AI learns from reanalysis built with physics and observations and still needs an up-to-date starting state. It makes one part of the forecasting chain dramatically cheaper; it does not remove the rest. ECMWF’s AIFS operational announcement
GenCast’s published evaluation beat ECMWF’s comparison ensemble on 97.2% of 1,320 evaluated targets. Those targets and the historical test period define the result. A practical warning system also needs calibrated probabilities, local rainfall detail and dependable performance on unusual events. Cheap ensembles can help estimate rare-event risk, but more members do not fix a shared bias. Public value comes from a chain: observe the hazard, estimate its probability, connect it to local exposure, communicate an understandable warning, and make evacuation or protective action possible. Price et al., Nature, GenCast; WMO’s 2025 early-warning report
The delivery chain is editorial synthesis. A country reporting an early-warning system is not evidence that every resident can receive and act on a warning.
Observe the empty places
Much of Africa, the oceans and the upper atmosphere are still thinly observed. Better starting data helps every model, physical or AI.
Keep physics and AI together
AI models learned from reanalyses built with physics models. Both are needed: physics to create the training data and handle the unprecedented, AI for speed and scale.
Spend the savings on ensembles
If an AI forecast costs a thousandth as much energy, the right use of the savings is many more ensemble members and better odds on extremes.
Forecast impacts, not weather
People need to know what a storm will do to their roof or field, not its wind speed in knots.
Close the last mile
In 2025, 40% of countries still reported no multi-hazard early warning system. That gap matters more than the next day of skill.
Respect the limit
Beyond about two weeks, day-to-day weather is unpredictable in principle. Progress there means better probabilities for weeks and seasons, not daily detail.
Who is building what
Forecast centres, AI labs and observation programmes. Search the record, or filter by role.
ECMWFAIFS Single and AIFS ENSMachine-learning forecasts run operationally alongside the physics-based IFS
- Reported evidence
- AIFS ENS, a 51-member AI ensemble, became operational in July 2025 with gains of up to 20% on some measures and about 1,000 times less energy per forecast.
- Announced next step
- Combined AI and physics forecasting and AI-based data assimilation.
- Unresolved risk
- Behaviour in unprecedented extremes; dependence on reanalysis training data.
Google DeepMindGraphCast and GenCastGraph neural network and diffusion-model forecasts trained on ERA5
- Reported evidence
- GenCast outperformed ECMWF’s ensemble on 97.2% of 1,320 targets and produces a 15-day ensemble in about eight minutes (Nature, 2024).
- Announced next step
- Operational use by forecasting agencies and in products.
- Unresolved risk
- Benchmarks are measured against reanalysis; operational performance on rare extremes needs time to establish.
MicrosoftAuroraA large foundation model for the atmosphere, fine-tuned for weather, air quality and ocean waves
- Reported evidence
- Published in Nature in 2025 with results matching or exceeding operational systems on several tasks.
- Announced next step
- Broader Earth-system forecasting from one model.
- Unresolved risk
- Training cost and the need for high-quality data in each new domain.
HuaweiPangu-WeatherA 3D transformer trained on decades of reanalysis
- Reported evidence
- Published in Nature in 2023 as one of the first AI models to beat a leading physics model on key medium-range scores.
- Announced next step
- Operational services through partner agencies.
- Unresolved risk
- Deterministic forecasts can be overconfident without ensembles.
NVIDIAEarth-2 / FourCastNetAI weather models and tools for running large ensembles on GPUs
- Reported evidence
- Open models used by research groups and agencies to generate very large ensembles cheaply.
- Announced next step
- High-resolution regional and climate-scale AI simulation.
- Unresolved risk
- Value depends on agencies adopting and verifying the outputs.
U.S. National Hurricane CenterOfficial track and intensity forecastsForecasters combining physics and AI model guidance
- Reported evidence
- Set accuracy records for Atlantic track forecasts at every lead time in 2024.
- Announced next step
- Better forecasts of rapid intensification.
- Unresolved risk
- Rapid intensification remains hard; budget and staffing cuts can affect observations.
WMO / UNDP / UNEPSystematic Observations Financing FacilityGrants to help the least developed countries and small island states meet basic observing standards
- Reported evidence
- Supporting 68 countries, aiming for more than a tenfold increase in shared balloon data and twentyfold in surface station data.
- Announced next step
- Sustained compliance with the Global Basic Observing Network.
- Unresolved risk
- Stations must be maintained for years after installation; funding must be sustained.
United NationsEarly Warnings for AllGlobal initiative to cover everyone with early warning systems
- Reported evidence
- In 2025, 119 countries, 60% of the total, reported a multi-hazard early warning system, up 113% in ten years.
- Announced next step
- Universal coverage by 2027.
- Unresolved risk
- Small island states and least developed countries lag; warnings must reach people and trigger action.
Benchmark scores compare models against reanalysis on chosen variables and periods. Operational value is shown by sustained performance in real time, including on rare extremes.
An optimistic view, with conditions
Better odds, further out, for everyone
Day-to-day forecasts can still gain several days before they reach the atmosphere’s limit. Cheap AI ensembles can give better probabilities for rare extremes. And closing the observation and warning gaps could deliver more value than any model improvement.
Bigger, cheaper ensembles
AI forecasts run in minutes, so centres can afford hundreds of ensemble members and sharper odds on extremes.
Observe the gaps
Fill the balloon and station gaps over Africa, the oceans and small islands, which improves forecasts for everyone downstream.
Weeks and seasons
Beyond two weeks, value comes from probabilities: a wetter-than-normal month or a likely heatwave, good enough to plan harvests and power grids.
Four numbers to watch
First, the lead time at which forecasts lose useful skill, tracked each year by major centres. Second, the performance of AI models on record-breaking extremes, not just averages. Third, the number of weather balloon and surface stations reporting from the least observed regions. Fourth, the share of people covered by early warnings, which measures whether better forecasts reach anyone.
Sources, method, and boundaries
Skill trends come from the European Centre’s published verification and the 2015 review by Bauer, Thorpe and Brunet. AI benchmark results are from peer-reviewed papers and ECMWF announcements. Predictability limits are from model experiments and remain estimates. The interactive model is an idealised error-growth calculation tuned to reproduce published behaviour, not a weather model.
- Forecast skill
- How much better a forecast is than a simple reference, such as the long-term average for that date.
- Ensemble
- A set of forecasts started from slightly different conditions, used to estimate the range of outcomes.
- Reanalysis
- A reconstruction of past weather made by combining historical observations with a modern forecast model.
Read More
Essays, books, talks, and research that shaped this field’s arguments. Influence is not endorsement; company communications and advocacy are labeled. Some publisher links require a subscription.
- Paper
Bauer, Thorpe & Brunet — The quiet revolution of numerical weather prediction (Nature, 2015)
The standard account of how forecasts gained a day of skill per decade.
- Paper
Edward Lorenz — Deterministic Nonperiodic Flow (1963)
The founding paper of chaos theory and the reason forecasts have a horizon.
- Paper
Price et al. — Probabilistic weather forecasting with machine learning (Nature, 2024)
The GenCast paper, in which an AI ensemble beat the leading physics-based ensemble on most targets.
- Paper
Zhang et al. — What is the predictability limit of midlatitude weather? (2019)
Model experiments estimating how much further forecasts can go.
Across the fields: learning curves, deployment, and rebound
- Theodore Wright — Factors Affecting the Cost of Airplanes (1936)
The original experience-curve formulation: costs change with accumulated production.
- Kenneth Arrow — The Economic Implications of Learning by Doing (1962)
The economics of productivity improvement through production experience.
- William Stanley Jevons — The Coal Question (1865)
The classic rebound argument: lower effective costs can expand total demand.
- Arnulf Grübler — The costs of the French nuclear scale-up: A case of negative learning by doing (2010)
A counterexample to assuming that greater deployment always lowers costs.

.jpg)

















