How Far Ahead Can We Forecast the Weather?

Forecasts gain about one day of useful range every decade, so a five-day forecast today is roughly as reliable as a two-day forecast was thirty years ago. AI models have just made forecasts faster and cheaper to run. But detailed midlatitude weather has a predictability horizon of roughly two weeks, and the real gap now is getting good forecasts to the people who need them.

Last updated October 2026
Figure 1 · The quiet revolution

One more day of useful forecast every decade

Weather forecasting has improved steadily for half a century, mostly out of public view. Three measurements show how far it has come and where it stands now.

~1 day / decadeLong-run gain in forecast skill since the 1970s: a forecast for day five today is as good as one for day four ten years ago (Measured, ECMWF).
Up to 75% smallerReduction in the National Hurricane Center’s Atlantic track errors compared with a few decades ago. In 2024 it set accuracy records at every lead time (Measured).
~10 → ~14 daysToday’s practical limit for useful day-to-day forecasts, against an estimated intrinsic limit of about two weeks (Measured and Projected).

Skill gains vary by variable, region and measure. “One day per decade” is the long-run trend for large-scale weather patterns in the global models of the European Centre for Medium-Range Weather Forecasts.

The story in one paragraph

Every forecast starts from an imperfect picture of the atmosphere right now, and the atmosphere amplifies small errors: a mistake in today’s wind doubles in a day or two, and keeps doubling until the forecast is no better than guessing from the season. That is why forecasting has improved steadily rather than suddenly. Better satellites, better ways of combining observations and faster computers have each shrunk the starting error a little, buying roughly one extra day of useful forecast every ten years. In 2023 and 2024 something new arrived: AI models trained on past weather began matching or beating leading physics-based systems on specified benchmarks. ECMWF’s operational AIFS later reported roughly a thousandfold reduction in forecast-generation energy. That is a genuine leap in cost and speed. It does not break the atmosphere’s limit, which theory puts at about two weeks for everyday weather. The remaining gains include better probabilities as well as wider access: better observations where the world is thinly measured, and warnings that actually reach the billions of people who still lack them.

  • Forecast skill has increased by about one day per decade for 40 years, through better observations, data assimilation, models and ensembles (Measured, Bauer, Thorpe and Brunet, 2015).
  • Google DeepMind’s GenCast beat the European Centre’s operational ensemble on 97.2% of 1,320 test targets, and generates one 15-day ensemble member in about eight minutes on a TPU (Measured, Nature 2024).
  • ECMWF’s deterministic AIFS became operational in February 2025, with approximately 1,000 times less forecast-generation energy; its AI ensemble followed in July 2025 (Reported operational milestones).
  • Reducing today’s starting errors tenfold would add up to about five days of useful forecast, pushing the limit for mid-latitude weather towards two weeks (Projected, Zhang et al., 2019).

Benchmark results, operational performance, modelled limits and targets are labelled. A model winning a benchmark is not the same as a forecast reaching a person in time.

Part I: Why weather is hard to predict

Small errors grow until they swallow the forecast

The atmosphere is a chaotic fluid. Two almost identical starting states drift apart over time, which means even a perfect model would eventually go wrong if its starting picture was slightly off.

error doubling time×log₂(useful error ÷ starting error)=forecast horizon

This simple relationship explains the whole history of forecasting. If errors double every day and a half, halving the starting error adds a day and a half of useful forecast. A tenfold improvement adds about five days. Nothing short of a perfect description of every gust and eddy could extend it indefinitely, and that is impossible. Worse, small features like thunderstorms grow their errors in hours, and those errors cascade upwards to larger weather systems. That cascade is why theory puts a ceiling of about two weeks on forecasting day-to-day weather in the mid-latitudes, however good the models become.

The same arithmetic is encouraging, too. Today’s forecasts still sit several days short of that ceiling, so there is real room left, and every gain in observations and assimilation converts directly into usable days.

What the horizon means

A thunderstorm, a weather system and a season have different clocks

There is no single expiry date for a forecast. A local thunderstorm is sensitive to small-scale details that become uncertain within hours. The track of a large pressure system remains useful for days. Slower signals from the ocean, soil moisture and large-scale circulation can support probabilities for wetter or warmer weeks beyond the familiar two-week horizon. ECMWF explicitly distinguishes these scales. A forecast of increased seasonal flood risk can be useful without predicting which street will flood on a particular afternoon. ECMWF’s forecast error guide; ECMWF’s forecast skill horizon study

The logarithm in the model explains diminishing returns. If errors double every 1.5 days, making the starting error ten times smaller buys 1.5 × log₂(10), or 4.98 days. It does not buy ten times the forecast range. This is an illustration of error growth, not a universal law: doubling times depend on scale and flow, while errors generated inside the model also matter. Near the intrinsic limit, small-scale uncertainty spreads upward even when the initial large-scale picture is unusually accurate. Selz et al., Journal of the Atmospheric Sciences, 2022

Roughly two weeks describes detailed midlatitude weather under particular skill measures. It is not a ceiling on all probabilistic, subseasonal or seasonal information.

Figure 2 · Interactive model

Why does better data buy only a few more days?

Small errors in today’s weather double every day or two until the forecast is no better than guessing from the climate. Each halving of the starting error therefore buys one more doubling time, never more.

Global models today9.6days of useful forecast

6.4 error doublings before the forecast stops being useful

Gain from halving the starting error
+1.5 days
Gain from a tenfold better starting state
+5.0 days
Starting error, share of saturation
0.7%

The doubling times and starting errors are illustrative, chosen so the presets reproduce published behaviour: about ten days of useful skill today, roughly four to five fewer in the 1980s, and up to five more days with a tenfold reduction in starting error, as Zhang and colleagues estimated in 2019. Thunderstorms grow errors in hours, which is why they are predictable for hours, not days.

Calculation & assumptions

Useful lead time = doubling time × log₂(usefulness threshold ÷ starting error). Errors are expressed as a share of saturation, the size of error a forecast would have if it knew nothing but the climate. Real error growth is not perfectly exponential: small scales saturate fast and pass their errors up to larger scales, which is what sets the roughly two-week intrinsic limit for mid-latitude weather. The model ignores model error, which better physics and machine learning reduce independently of observations.

Figure 2: An editorial error-growth model illustrating diminishing returns. It does not impose an intrinsic horizon; extreme parameter choices can exceed the physically realistic range for detailed weather.
Part II: From measurement to warning

A forecast is a chain, and value is lost at every link

A forecast only matters if it changes a decision. That requires every link, from a weather balloon to an evacuation order, to work.

01

Observations

Satellites, weather balloons, aircraft, ships, buoys, radar and ground stations measure the atmosphere as it is right now.

Measure
Observations assimilated per day · coverage gaps
Failure boundary
Where few observations exist, as over much of Africa, every forecast starts from a blurrier picture.
Where the frontier moves

Funding sustained surface and balloon networks in the countries with the largest gaps.

02

Data assimilation

Millions of scattered, imperfect measurements are combined with a short forecast to produce the best estimate of the whole atmosphere’s current state.

Measure
Error of the starting analysis
Failure boundary
A forecast can never be better than its starting point; errors here grow fastest.
Where the frontier moves

Machine-learning assimilation that uses more satellite data, faster.

03

The forecast model

Physics-based models solve the equations of fluid motion on a grid; AI models learn the same evolution from decades of past weather.

Measure
Error at each lead time · compute per forecast
Failure boundary
AI models are trained on past climate and may handle unprecedented extremes less reliably.
Where the frontier moves

AI models that match or beat physics models at a thousandth of the energy, freeing compute for larger ensembles.

04

Ensembles

Running many slightly different forecasts shows the range of possible outcomes and how confident the forecast is.

Measure
Members per forecast · reliability of stated probabilities
Failure boundary
A single forecast hides uncertainty; a poorly spread ensemble gives false confidence.
Where the frontier moves

Cheap AI ensembles with hundreds of members, improving the odds attached to rare extremes.

05

Warnings and communication

National weather services turn forecasts into warnings that people understand and trust.

Measure
Share of population covered by early warnings
Failure boundary
A perfect forecast that never reaches a farmer or a fishing village saves no one.
Where the frontier moves

Mobile alerts, impact-based warnings and local-language communication.

06

Action

Evacuations, harvest timing, grid planning and flight routing: the decisions that turn a forecast into value.

Measure
Lives and losses avoided per warning
Failure boundary
Warnings that are too frequent, or wrong too often, train people to ignore them.
Where the frontier moves

Pre-agreed actions and finance that trigger automatically when a forecast crosses a threshold.

Part III: The AI turn

A forecast that runs on one chip in minutes

Between 2022 and 2025, machine-learning weather models went from research curiosity to operational service.

What changed

97.2%Share of 1,320 targets on which GenCast beat the leading physics-based ensemble (Measured, Nature 2024).
~1,000×Lower energy use of the European Centre’s AI system than its physics model, per forecast (Reported by ECMWF).
51 membersSize of the AI ensemble the European Centre made operational in July 2025 (Measured).

Physics-based models need a supercomputer for hours. AI models, once trained, produce a global forecast in minutes on a single processor. They learned from ERA5, a reconstruction of the atmosphere back to 1940, built by combining historical observations with a physics model. That dependence matters: the AI models are only as good as that training data, they still need a physics-based analysis to start from, and they have less experience of extremes the climate has not produced before. The likely future is not AI replacing physics but the two working together, with the savings spent on far larger ensembles.

Hurricanes show the gains in human terms

In 2024 the U.S. National Hurricane Center’s track forecasts were the most accurate in its history at every lead time from 12 hours to five days. Its intensity forecasts have improved more slowly: storms that strengthen rapidly are still hard to anticipate, though forecasters’ average underestimate of such storms fell from 26 knots in 2010–14 to 16 knots in 2020–24. Better track forecasts mean smaller evacuation zones, fewer unnecessary evacuations, and more trust when an evacuation is ordered.

Part IV: The real bottleneck

The forecast exists; the warning often does not

The world’s best forecasts are now available almost everywhere. The weakest links are observations in poorer countries and the systems that turn forecasts into warnings people act on.

Africa’s surface weather stations are about an eighth as dense as the World Meteorological Organization recommends, and more stations report from Germany than from the whole African continent. That starves every model of data over the regions where the population is most exposed to droughts, floods and heat. The UN-backed Systematic Observations Financing Facility aims to fix part of this, supporting the least developed countries and small island states to meet basic observing standards, which would multiply internationally shared balloon data more than tenfold.

Warnings are improving fast. In 2025, 119 countries, 60% of the total, reported having a multi-hazard early warning system, more than double the number ten years earlier. The WMO estimates that investment in early warnings could deliver benefits of at least $162 billion a year, about ten times their cost, and that better forecasts and warnings could save around 23,000 lives a year.

From benchmark to public value

Cheap forecasts still need expensive observations and trusted warnings

ECMWF put its deterministic AIFS into operational service on 25 February 2025, reporting roughly a thousandfold reduction in energy for generating a forecast. That boundary matters: it excludes the entire observing network, preparation of the initial atmospheric analysis and training the model. AI learns from reanalysis built with physics and observations and still needs an up-to-date starting state. It makes one part of the forecasting chain dramatically cheaper; it does not remove the rest. ECMWF’s AIFS operational announcement

GenCast’s published evaluation beat ECMWF’s comparison ensemble on 97.2% of 1,320 evaluated targets. Those targets and the historical test period define the result. A practical warning system also needs calibrated probabilities, local rainfall detail and dependable performance on unusual events. Cheap ensembles can help estimate rare-event risk, but more members do not fix a shared bias. Public value comes from a chain: observe the hazard, estimate its probability, connect it to local exposure, communicate an understandable warning, and make evacuation or protective action possible. Price et al., Nature, GenCast; WMO’s 2025 early-warning report

The delivery chain is editorial synthesis. A country reporting an early-warning system is not evidence that every resident can receive and act on a warning.

Observe the empty places

Much of Africa, the oceans and the upper atmosphere are still thinly observed. Better starting data helps every model, physical or AI.

Keep physics and AI together

AI models learned from reanalyses built with physics models. Both are needed: physics to create the training data and handle the unprecedented, AI for speed and scale.

Spend the savings on ensembles

If an AI forecast costs a thousandth as much energy, the right use of the savings is many more ensemble members and better odds on extremes.

Forecast impacts, not weather

People need to know what a storm will do to their roof or field, not its wind speed in knots.

Close the last mile

In 2025, 40% of countries still reported no multi-hazard early warning system. That gap matters more than the next day of skill.

Respect the limit

Beyond about two weeks, day-to-day weather is unpredictable in principle. Progress there means better probabilities for weeks and seasons, not daily detail.

Who is building what

Forecast centres, AI labs and observation programmes. Search the record, or filter by role.

8 programmes
ECMWFAIFS Single and AIFS ENSMachine-learning forecasts run operationally alongside the physics-based IFS
Reported evidence
AIFS ENS, a 51-member AI ensemble, became operational in July 2025 with gains of up to 20% on some measures and about 1,000 times less energy per forecast.
Announced next step
Combined AI and physics forecasting and AI-based data assimilation.
Unresolved risk
Behaviour in unprecedented extremes; dependence on reanalysis training data.
Google DeepMindGraphCast and GenCastGraph neural network and diffusion-model forecasts trained on ERA5
Reported evidence
GenCast outperformed ECMWF’s ensemble on 97.2% of 1,320 targets and produces a 15-day ensemble in about eight minutes (Nature, 2024).
Announced next step
Operational use by forecasting agencies and in products.
Unresolved risk
Benchmarks are measured against reanalysis; operational performance on rare extremes needs time to establish.
MicrosoftAuroraA large foundation model for the atmosphere, fine-tuned for weather, air quality and ocean waves
Reported evidence
Published in Nature in 2025 with results matching or exceeding operational systems on several tasks.
Announced next step
Broader Earth-system forecasting from one model.
Unresolved risk
Training cost and the need for high-quality data in each new domain.
HuaweiPangu-WeatherA 3D transformer trained on decades of reanalysis
Reported evidence
Published in Nature in 2023 as one of the first AI models to beat a leading physics model on key medium-range scores.
Announced next step
Operational services through partner agencies.
Unresolved risk
Deterministic forecasts can be overconfident without ensembles.
NVIDIAEarth-2 / FourCastNetAI weather models and tools for running large ensembles on GPUs
Reported evidence
Open models used by research groups and agencies to generate very large ensembles cheaply.
Announced next step
High-resolution regional and climate-scale AI simulation.
Unresolved risk
Value depends on agencies adopting and verifying the outputs.
U.S. National Hurricane CenterOfficial track and intensity forecastsForecasters combining physics and AI model guidance
Reported evidence
Set accuracy records for Atlantic track forecasts at every lead time in 2024.
Announced next step
Better forecasts of rapid intensification.
Unresolved risk
Rapid intensification remains hard; budget and staffing cuts can affect observations.
WMO / UNDP / UNEPSystematic Observations Financing FacilityGrants to help the least developed countries and small island states meet basic observing standards
Reported evidence
Supporting 68 countries, aiming for more than a tenfold increase in shared balloon data and twentyfold in surface station data.
Announced next step
Sustained compliance with the Global Basic Observing Network.
Unresolved risk
Stations must be maintained for years after installation; funding must be sustained.
United NationsEarly Warnings for AllGlobal initiative to cover everyone with early warning systems
Reported evidence
In 2025, 119 countries, 60% of the total, reported a multi-hazard early warning system, up 113% in ten years.
Announced next step
Universal coverage by 2027.
Unresolved risk
Small island states and least developed countries lag; warnings must reach people and trigger action.

Benchmark scores compare models against reanalysis on chosen variables and periods. Operational value is shown by sustained performance in real time, including on rare extremes.

An optimistic view, with conditions

Better odds, further out, for everyone

Day-to-day forecasts can still gain several days before they reach the atmosphere’s limit. Cheap AI ensembles can give better probabilities for rare extremes. And closing the observation and warning gaps could deliver more value than any model improvement.

60%Countries reporting a multi-hazard early warning system in 2025, up 113% in ten years (Measured).
2027UN target year for early-warning coverage of every person on Earth (Target).
+5 daysPotential gain from a tenfold cut in starting error (Projected).
Now

Bigger, cheaper ensembles

AI forecasts run in minutes, so centres can afford hundreds of ensemble members and sharper odds on extremes.

Next scale test

Observe the gaps

Fill the balloon and station gaps over Africa, the oceans and small islands, which improves forecasts for everyone downstream.

Deep change

Weeks and seasons

Beyond two weeks, value comes from probabilities: a wetter-than-normal month or a likely heatwave, good enough to plan harvests and power grids.

Four numbers to watch

First, the lead time at which forecasts lose useful skill, tracked each year by major centres. Second, the performance of AI models on record-breaking extremes, not just averages. Third, the number of weather balloon and surface stations reporting from the least observed regions. Fourth, the share of people covered by early warnings, which measures whether better forecasts reach anyone.

Sources, method, and boundaries

Skill trends come from the European Centre’s published verification and the 2015 review by Bauer, Thorpe and Brunet. AI benchmark results are from peer-reviewed papers and ECMWF announcements. Predictability limits are from model experiments and remain estimates. The interactive model is an idealised error-growth calculation tuned to reproduce published behaviour, not a weather model.

Forecast skill
How much better a forecast is than a simple reference, such as the long-term average for that date.
Ensemble
A set of forecasts started from slightly different conditions, used to estimate the range of outcomes.
Reanalysis
A reconstruction of past weather made by combining historical observations with a modern forecast model.

Read More

Essays, books, talks, and research that shaped this field’s arguments. Influence is not endorsement; company communications and advocacy are labeled. Some publisher links require a subscription.

  1. Paper

    Bauer, Thorpe & Brunet — The quiet revolution of numerical weather prediction (Nature, 2015)

    The standard account of how forecasts gained a day of skill per decade.

  2. Paper

    Edward Lorenz — Deterministic Nonperiodic Flow (1963)

    The founding paper of chaos theory and the reason forecasts have a horizon.

  3. Paper

    Price et al. — Probabilistic weather forecasting with machine learning (Nature, 2024)

    The GenCast paper, in which an AI ensemble beat the leading physics-based ensemble on most targets.

  4. Paper

    Zhang et al. — What is the predictability limit of midlatitude weather? (2019)

    Model experiments estimating how much further forecasts can go.

Across the fields: learning curves, deployment, and rebound