Build Your Own Error Bars — Temperature

Weather archives hold two kinds of number for the same place and month: what a thermometer recorded, and what a weather model reconstructs. Here you measure the gap between them yourself — then test whether the gap you measured in one place tells you anything about another.

Pick a station. Put its record next to the model's reconstruction for the same spot. Build the distribution of their disagreements, freeze a band around it, and carry that band to a second station to see how it holds up. There is a sibling lab that runs the same idea on sunshine instead of temperature — see Build Your Own Error Bars — Sunshine (Lab 04).

  • What a reanalysis is, and how it differs from a station's own record.
  • How to turn two overlapping series into a difference distribution and an empirical error band.
  • How to test — rather than assume — whether an error band built in one place transfers to another.

What's being compared

For every station in this lab there are two monthly temperature series. One is the station's own record: monthly means built from thermometer readings, published in the NOAA GHCN monthly dataset. The other is a reanalysis estimate — a weather model's reconstruction of past conditions at the station's coordinates, from the ERA5 dataset.

Both series claim to describe the same place and the same months, from 1940 to the present. This lab never treats one as the answer key for the other; it only measures how much, and when, they disagree — and puts that measurement in your hands.

Rather than compare raw temperatures, the lab compares each series as an anomaly — its departure from its own 1991–2020 average, measured separately for each calendar month. That strips out a fixed gap between the two: a station sitting above or below the average height of ERA5's grid cell reads a degree or two apart from it for reasons that have nothing to do with climate. Set that offset aside and what remains is the thing this lab is about — the shape of the disagreement over time.

Why anomalies, and not raw temperatures?

An anomaly is a value measured against a fixed reference average instead of against zero degrees. Here every series — station and reanalysis alike — is expressed as its departure from its own 1991–2020 monthly average (the current WMO climate normal), using the same thirty-year window for every station.

The reason is that a reanalysis value is a grid cell's average, tied to the average elevation of the terrain across that cell. A station that sits higher or lower than that average reads a roughly constant amount cooler or warmer than the reanalysis — a fixed offset set by geometry, not by any disagreement about the climate. Left in, that offset dominates the chart and hides what is worth looking at. Subtracting each series' own baseline removes it, so the two start from a common zero and any gap that remains is a real difference in how they vary over time. The offset itself is neither error nor signal; it is expected, and the anomaly view simply sets it aside.

What is a reanalysis?

A reanalysis runs a modern weather-forecast model over the past, continuously nudged toward millions of historical observations — weather balloons, ships, aircraft, satellites, and surface stations. The output is a physically consistent estimate of the atmosphere everywhere on a grid, including places and times nothing was measured. ERA5, produced by the European Centre for Medium-Range Weather Forecasts, covers 1940 to the present.

A reanalysis value at a station's coordinates is not the station's reading fed back out: it is the model's estimate for a grid cell around that point, shaped by everything the model assimilated and by the model's own physics.

How is the model's series made comparable to the station's?

GHCN monthly means are built from daily maximum and minimum readings: each day contributes (Tmax+Tmin)/2, and the month averages those. The reanalysis series here is aggregated the same way — ERA5's daily maximum and minimum at the station's coordinates, combined as (Tmax+Tmin)/2 and averaged over each month (months missing more than 10% of days are dropped). Comparing like-for-like on definition matters: whether (Tmax+Tmin)/2 is itself a good stand-in for a day's true mean temperature is a separate question, deliberately outside this lab.

Is either series the "true" temperature?

This lab does not treat either as ground truth. Station records carry their own histories — siting, instrument changes, moves, time-of-observation changes. Reanalysis estimates carry model assumptions, a grid cell's worth of terrain standing in for a point, and an observing system that changed over the decades. The one thing that can be measured directly is how much the two disagree; what to make of that is yours to decide.

Choose a calibration station

This is where you will build your error band. The stations span a wide range of surrounding built-up land, from very little local development to dense urban areas — built-up context is shown on each card so you can consider it as part of your choice rather than having it hidden. The comparison window starts in 1940, when the reanalysis begins.

How were these stations chosen?

Every station here is a reference-grade long record. Most are stations in recognised reference networks — the GCOS Surface Network and the US Historical Climatology Network — chosen for long, well-documented records. A few are co-located "gold" sites, where a long record carries through a documented same-place instrument handover with no break in continuity (Stillwater, kept on university land through its move, is one). They are chosen to be trustworthy first, then to span surrounding built-up land from open country to city, so you can weigh the built-up context yourself — in either direction. Reference grade is a quality criterion fixed before any comparison, not chosen to favour a result, and the non-gold stations are drawn at random within each built-up band, in a random order, so no station reads as the "right" place to start.

Stations are shown in a random order to avoid nudging your choice.

Select a station above to continue.

Questions you might like to ask

Written before we computed any of the answers ourselves — they are prompts, not hints, and the lab takes no position on any of them.

  • Does the size of the disagreement depend on the season? On the decade?
  • Does the picture change when you switch the averaging from monthly to annual?
  • Does the adjusted (QCF) version of a record sit differently against the reanalysis than the unadjusted (QCU) version?
  • Does a band built at one station hold at its stated rate elsewhere? Does it matter how far away the second station is, or how different its built-up surroundings are?
  • Does the gap look different at heavily built-up stations than at barely developed ones?
  • If you build the band at a different calibration station, do your conclusions survive?
Data sources & methodology
Station temperature records — NOAA GHCNm v4
Monthly mean temperatures from the NOAA Global Historical Climatology Network Monthly v4, in both its unadjusted (QCU) and adjusted (QCF) versions. Menne et al., 2018. Series are fetched per station from klymot.com's published data mirror (www.klymot.com/data/{qcu,qcf}/<station-id>.csv), the same files behind the main site's station explorer.
Reanalysis temperature — ERA5 via Open-Meteo
ERA5 is produced by the European Centre for Medium-Range Weather Forecasts: Hersbach et al., 2020. Daily 2 m maximum and minimum temperature at each station's coordinates are fetched from the Open-Meteo Historical Weather API and aggregated to monthly means of (Tmax+Tmin)/2 — the GHCNm definition — keeping months with at least 90% of days present. Open-Meteo downscales ERA5's coarse grid to the station's own recorded elevation, so the model and the thermometer are compared at the same height rather than at whatever elevation a terrain model happens to assign the grid cell.
Anomalies — 1991–2020 baseline
Every series, station and reanalysis alike, is shown as an anomaly: its departure from its own average for each calendar month over 1991–2020 — the current WMO climate normal — the same thirty-year window applied to every station. Working per calendar month subtracts out the seasonal cycle, so the monthly and seasonal views show departures rather than the seasons themselves. This also removes the near-constant offset between a station and ERA5 that comes from the station's height relative to the grid cell's mean elevation (geometry, not climate), leaving a comparison of how the two vary rather than of a fixed gap. Any month with no data inside the baseline window is omitted.
Station pool
Reference-grade GHCN long records only: stations in the GCOS Surface Network (GSN) or the US Historical Climatology Network (HCN) — external NOAA/WMO quality flags, not our own selection — each with a near-continuous annual record from 1940 to today. Added to these are co-located "gold" reference sites, where a long COOP record is spliced onto its co-located successor: the newer reference station anchors the recent era unchanged, and the older record is shifted onto its level by the offset measured over their overlap (estimated per calendar month and documented in the data file). For these the record-version choice is not QCU vs QCF but Spliced (offset-corrected) vs Raw (the same records joined with no offset applied), so you can see what the correction does rather than take it on trust. From the eligible stations, one is drawn per built-up band so the pool spans open country to city. ("Gold" here means unbroken record continuity through the handover, not low development — Stillwater, for instance, is a gold reference that still sits in a town.)
Built-up context — GHSL
Built-up land fractions and imagery around each station are from the Global Human Settlement Layer's built-up surface data for 2020, as published on klymot.com's station explorer. European Commission JRC GHSL.
Bands and the transfer test
Differences are station minus reanalysis over periods where both have values, averaged over identical months. The band spans the central quantiles of those differences at your chosen coverage (empirical, no distribution assumed; at least 30 differences required). The transfer test counts the share of another station's values falling inside the reanalysis series plus your frozen band.