Build Your Own Error Bars — Temperature
Weather archives hold two kinds of number for the same place and month: what a thermometer recorded, and what a weather model reconstructs. Here you measure the gap between them yourself — then test whether the gap you measured in one place tells you anything about another.
Pick a station. Put its record next to the model's reconstruction for the same spot. Build the distribution of their disagreements, freeze a band around it, and carry that band to a second station to see how it holds up. There is a sibling lab that runs the same idea on sunshine instead of temperature — see Build Your Own Error Bars — Sunshine (Lab 04).
- What a reanalysis is, and how it differs from a station's own record.
- How to turn two overlapping series into a difference distribution and an empirical error band.
- How to test — rather than assume — whether an error band built in one place transfers to another.
What's being compared
For every station in this lab there are two monthly temperature series. One is the station's own record: monthly means built from thermometer readings, published in the NOAA GHCN monthly dataset. The other is a reanalysis estimate — a weather model's reconstruction of past conditions at the station's coordinates, from the ERA5 dataset.
Both series claim to describe the same place and the same months, from 1940 to the present. This lab never treats one as the answer key for the other; it only measures how much, and when, they disagree — and puts that measurement in your hands.
Rather than compare raw temperatures, the lab compares each series as an anomaly — its departure from its own 1991–2020 average, measured separately for each calendar month. That strips out a fixed gap between the two: a station sitting above or below the average height of ERA5's grid cell reads a degree or two apart from it for reasons that have nothing to do with climate. Set that offset aside and what remains is the thing this lab is about — the shape of the disagreement over time.
Choose a calibration station
This is where you will build your error band. The stations span a wide range of surrounding built-up land, from very little local development to dense urban areas — built-up context is shown on each card so you can consider it as part of your choice rather than having it hidden. The comparison window starts in 1940, when the reanalysis begins.
Stations are shown in a random order to avoid nudging your choice.
Select a station above to continue.
Choose which version of the station record to use
Every GHCN station has two published versions of its monthly record. Both are real, maintained datasets; they differ in whether adjustments for station changes have been applied. Pick the one you want to compare against the model — you can switch later, though switching clears the difference, band, and scorecard you've built, because they belong to one version of one record.
The two versions are shown in a random order to avoid nudging your choice.
Choose a version above to continue.
Look at the two series together
Both series are drawn as anomalies — each one's departure from its own 1991–2020 average — with equal weight; which one gets which colour, and which is listed first, is randomized on each visit, so neither is presented as the baseline the other is checked against. With the constant offset removed, the two lines share a common zero and you can watch where they part company. Change the averaging to see how much the comparison depends on it: a single calendar month, a single season, or whole years. When you compare by month or season, a second control picks which one — so like is always compared with like, rather than mixing times of year that can disagree with the model by very different amounts.
The starting averaging level and time-of-year filter are randomized on each visit; both are yours to change at any time.
Loading…
Construct the difference
So far you've seen the two series side by side. The next step subtracts one from the other, period by period — station minus reanalysis — and collects every one of those differences into a distribution you can put a band around.
Subtracts the reanalysis from the station record at your chosen averaging.
The difference, period by period
Each point is the station's anomaly minus the reanalysis anomaly for the same period. Because the fixed offset between them has already been removed, zero marks their typical 1991–2020 relationship: points above zero are periods where the station ran warmer relative to the model than it did on average across the baseline; points below, cooler. The averaging control above still applies.
Your difference distribution
The same differences, tallied — for the month, season, or year the controls above select. The shaded band marks the central share of them — you pick how central. A 50% band that misses half the points is behaving exactly as labelled; the label is the claim being tested later, so it stays on screen.
The starting coverage level is randomized on each visit; it's yours to change at any time.
Freezing locks the band so you can test it at another station.
Test your band somewhere else
Every time a model output stands in for a missing measurement, an assumption is being made: that the model's error where it can be checked resembles its error where it can't. Here you get to check it. Pick another station; your frozen band is drawn around the reanalysis series there, with the station's own measurements hidden until you choose to reveal them.
Stations are shown in a random order to avoid nudging your choice.
Select a station above to draw your band there.
Loading…
Pick a station above first.
Your scorecard
Each station you reveal adds a row: how many of its measurements your band caught, against the share the band held at your calibration station. Built-up land around each station is listed so you can look for any relation to built-up context yourself — in either direction.
| Station | BU in 1 km | Inside band | Share inside | Band claimed |
|---|