Close gaps found in a review of the documentation: - Track data/metrics/ and data/metric_sources.json in git so the data checker passes on a fresh clone (finding 23; decision logged) - State the Alaska and Hawaii coverage gap in the README limitations and extend finding 12 - File findings 24-26: the data-sources doc lacks gridMET and several pipeline commands; wettest/driest month are computed twice; solar GHI is written by the extreme-temperature apply step - Add the stale "fallback values" note to filter 2's tasks Tidy the document system: - Add docs/reviews/README.md with the numbering rules and a finding index - Rename koppen-mixed-display-plan.md to koppen-mixed-display.md and fix its stale Puerto Rico and "stage 5" text - Add the precipitation-month step to the README enrichment list - Describe the Current method / Previous method pattern in plan section 5 - Add CLAUDE.md with the project guardrails and doc layout Format filter-calculations.md so it renders on GitHub and in VS Code: inline math uses $...$, ranges use en dashes, and implementation references name functions instead of line numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
20 KiB
County Filter Calculations
This document records the equations, implementation behavior, and review notes for the 12 county filters currently exposed by the climate explorer.
Scope and notation
The current metric list is defined in app.js under METRICS. Unless noted
otherwise, long-term climate metrics use the 1991–2020 reference period.
| Symbol | Meaning |
|---|---|
c |
County |
i |
Raster or model grid cell |
m |
Calendar month |
d |
Calendar day |
h |
Hour or NSRDB time row |
y |
Year |
G_c |
Valid grid cells assigned to county c |
V_c |
Valid observations for county c |
\mathbf{1}[A] |
1 when condition A is true; otherwise 0 |
\operatorname{clip}(x,a,b) |
Restrict x to the interval [a,b] |
Missing values are omitted from means unless a metric-specific rule below says otherwise.
Browser filter predicate
For every numeric metric, a county is active when its value is finite and falls inside the selected range, including both endpoints:
\operatorname{passes}(c) \iff L \le x_c \le U.
For a categorical metric:
\operatorname{passes}(c) \iff
\left(F=\text{All}\right) \lor \left(x_c=F\right).
The Köppen-Geiger filter also matches Mixed counties by their top two classes; see §1.
Counties with missing categorical values, null numeric values, or non-finite
numeric values do not pass. The numeric slider limits are derived from the
loaded data and rounded outward by each metric's configured boundsStep.
Implementation: shouldFeaturePassFilter and getUniqueCategoryValuesInData
in app.js.
Filter inventory
| Group | UI label | Data key | Type | Display unit |
|---|---|---|---|---|
| Köppen-Geiger Classification | Köppen-Geiger Climate Class | koppenZone |
Categorical | Class |
| Temperature & Extremes | Annual Avg Temperature (Normals) | avgTempF |
Numeric | °F |
| Temperature & Extremes | Diurnal Temperature Range | avgDiurnalTempRangeF |
Numeric | °F difference |
| Temperature & Extremes | Annual Extreme Temperature Days | absoluteExtremeDays |
Numeric | Days/year |
| Temperature & Extremes | Annual 90 F+ Heat Index Days | humidHeatDays |
Numeric | Days/year |
| Precipitation & Moisture | Annual Precipitation (Normals) | annualPrecipIn |
Numeric | Inches/year |
| Precipitation & Moisture | Seasonality Index | seasonalityIndex |
Numeric | 0–100 index |
| Precipitation & Moisture | Wettest Month | wettestPrecipMonth |
Categorical | Month |
| Precipitation & Moisture | Driest Month | driestPrecipMonth |
Categorical | Month |
| Precipitation & Moisture | Summer Specific Humidity | avgSummerSpecificHumidityGKg |
Numeric | g/kg |
| Solar Resource | Mean Daily Global Horizontal Radiation (GHI) | meanDailyGlobalHorizontalRadiationKwhM2Day |
Numeric | kWh/m²/day |
| Solar Resource | Clear-Sky GHI Reduction Index | clearSkyGhiReductionIndex |
Numeric | Ratio |
1. Köppen-Geiger Climate Class
Data keys: koppenZone; koppenPrimaryClass and koppenSecondaryClass
for Mixed counties
There is no continuous numerical score. Each county is classified from the
share of its land covered by each Köppen class. The rule was adopted on
2026-09-12 and applied to data/climate-data.csv on 2026-09-13.
Let s_{c,k} be the share of county $c$'s land area covered by class k.
Each valid raster cell is weighted by the area of the cell that lies inside the
county, a_{c,i}; ocean and no-data cells are excluded:
s_{c,k}
=
\frac{\sum_{i}a_{c,i}\,\mathbf{1}[K_i=k]}{\sum_{i}a_{c,i}}.
Rank the classes so that s_{c,(1)}\ge s_{c,(2)}\ge\cdots, with
s_{c,(2)}=0 when only one class is present. A county is predominantly
class k_{(1)}, shown as a single color, if and only if both conditions hold:
K_c=
\begin{cases}
k_{(1)}, & s_{c,(1)}\ge0.50 \;\text{and}\; s_{c,(1)}-s_{c,(2)}\ge0.05,\\
\text{Mixed}, & \text{otherwise}.
\end{cases}
The gap is measured in percentage points. A county that fails either condition
is classified as Mixed (shown as "Mixed Climate"). For Mixed counties,
p_c=k_{(1)} and q_c=k_{(2)} are stored in koppenPrimaryClass and
koppenSecondaryClass, and the map draws the county with stripes of those two
classes (see koppen-mixed-display.md).
Both columns are blank for predominant counties. A county with no valid raster
cells is left blank; none in the 50 states and DC is.
Each county is read from a small raster window around its polygon. Each cell is
split into 16 × 16 sub-cells to estimate the fraction inside the county, and
scaled by \cos(\text{latitude}) for its true surface area. A county whose
polygon crosses the 180th meridian (Aleutians West, AK) is split into one piece
on each side, and each piece is read from its own window.
Filter. Choosing a class F shows counties that are predominantly that
class and Mixed counties where it is the primary or secondary class; the
"Mixed Climate" option shows every Mixed county:
\operatorname{passes}(c) \iff
\left(F=\text{All}\right) \lor \left(K_c=F\right) \lor
\left(K_c=\text{Mixed} \land F\in\{p_c,q_c\}\right).
Implementation: scripts/build_county_koppen_metric.py (rank_class_shares,
classify, build_koppen_records) writes data/metrics/koppen.csv; area
weighting, raster windows, and the 180th-meridian split are in
scripts/common/county_zonal_stats.py;
scripts/apply_koppen_metric_to_climate_data.py copies the three columns into
data/climate-data.csv; the filter is shouldFeaturePassFilter in app.js.
Rationale. No published standard defines when an area is predominantly one Köppen class; the classification is defined per grid cell. The 50% condition means the label describes more than half of the county's land. The 5-point gap condition catches near 50/50 splits, such as Schenectady, NY (Dfb 50.1%, Dfa 49.9%), where a single label would rest on a margin of a few tenths of a point. Map-unit purity standards from other fields were considered and rejected: the FAO Land Cover Classification System treats a unit as single only above 80%, and USDA soil survey consociations allow roughly 15–25% dissimilar inclusions. Applied to counties, those thresholds would mark about 40–47% of the map area as mixed. A published county-level Köppen dataset (Audirac, Harvard Dataverse, 2024) uses the plurality class and reports the share of every class, without a threshold.
Results. Using the Beck et al. 2023 1991–2020 1 km raster, for the 3,143 counties in the 50 states and DC:
| Classification | Counties | Share of counties | Share of map area |
|---|---|---|---|
| Predominant (single class) | 3,010 | 95.8% | 85.6% |
| Mixed climate | 133 | 4.2% | 14.4% |
Of the 133 mixed counties, 111 have no class covering 50% or more (37 of these also have a gap under 5 points), and 22 have a majority class whose runner-up is within 5 points. They are concentrated in the mountain West and Alaska: Colorado and Alaska (13 each), California and Montana (12 each), Washington (10), Idaho (9), and Utah (8). No county's top three classes fall within 2 points of one another, so no additional mixed categories are needed.
Compared with the previous method (below). Every county whose label changed became Mixed; no county moved to a different single class. The touched-cell counts and the area-weighted shares pick a different plurality winner in only 3 counties (Denver, CO; Schenectady, NY; Hood River, OR), all of which are Mixed under this rule. Applying this rule to touched-cell counts instead of area-weighted shares would classify 9 counties differently, because touched-cell counts give full weight to boundary cells that are mostly outside the county.
Boundary cases. Five counties lie within 0.25 points of a cutoff that decides their outcome: Sanpete, UT (Dfb 49.93%, Mixed), Giles, VA (49.95%, Mixed), Placer, CA (Csa 50.22%), Albany, WY (Dfb 50.24%), and Park, MT (gap 5.23 points, Dfb). Their classification depends on the precision of the area weighting. Aleutians West, AK (02016) spans the antimeridian, so its shares were computed without sub-cell sampling. Puerto Rico is not covered by these figures; it has 9 more Mixed counties and is not shown in the app.
Previous method
Until 2026-09-13, koppenZone was the most frequent valid Köppen raster code
among the cells touched by the county polygon:
K_c = \underset{k}{\arg\max}\;
\sum_{i\in G_c}\mathbf{1}[K_i=k].
Cells were equally weighted, regardless of how much of each cell lay inside the
county; an exact tie went to the smallest raster code, and a county with no
valid cell was assigned Cfa. scripts/build_county_climate_data.py still
computes this value until that script is retired (pipeline plan, Phase 3), so
the Köppen apply step must run after it.
2. Annual Avg Temperature (Normals)
Data key: avgTempF
The annual value is a 1991–2020 climatological standard normal: the mean of
the 12 monthly normals of NOAA nClimGrid-Monthly average temperature, averaged
over each county's area. The definition was adopted on 2026-09-15. It is not
yet applied to data/climate-data.csv, whose values still use the current
method below.
Source. nClimGrid-Monthly is a 1/24° (about 5 km) grid of monthly values
interpolated from GHCN station data, covering the contiguous U.S. from 1895 to
the present. Its average temperature, tavg, is the mean of maximum and
minimum temperature, (Tmax + Tmin)/2, not a 24-hour mean. Alaska and Hawaii
are outside the grid, so their counties are blank.
Cell normals. For each grid cell i and calendar month m, the
normal is the mean over the 1991–2020 years with a valid value:
T_{i,m}^{\mathrm{norm}}
=
\frac{1}{Y_{i,m}}
\sum_{y\in V_{i,m}}T_{i,m,y}.
This is NCEI's own method for its gridded normals, which it describes as "a simple 30-year average of monthly grids" (Rennie and Palecki, U.S. Monthly Gridded Precipitation and Temperature Climate Normals). Our copy of nClimGrid is a later version than the June 2021 data NCEI used, so values can differ slightly from NCEI's published grids. The result is not NCEI's station-based U.S. Climate Normals product.
County monthly values. Each cell is weighted by the area of the cell that
lies inside the county, a_{c,i}, estimated as for Köppen (§1). Cells
without data, such as ocean, are excluded:
T_{c,m}
=
\frac{\sum_{i}a_{c,i}\,T_{i,m}^{\mathrm{norm}}}{\sum_{i}a_{c,i}}.
Annual value. Every month has equal weight. The value is defined only when
all 12 monthly values T_{c,m} exist; otherwise the county is blank:
T_c(^\circ\mathrm{F})
=
\left(
\frac{1}{12}\sum_{m=1}^{12}T_{c,m}
\right)\frac{9}{5}+32.
Values are kept at full precision until the stored value is rounded to 0.1 °F.
Rationale. The method follows the WMO rules for annual normals, which NOAA also applies to its 1991–2020 Normals:
- Equal month weights. For a mean, WMO-No. 1203 §4.3.3(a) defines the annual normal as "the mean of the monthly normals", and its footnote says weighting months by their number of days "is not recommended for internationally exchanged products" (WMO Guidelines on the Calculation of Climate Normals, 2017). NOAA's methodology states that "each month is treated equally in calculating seasonal and annual averages; they are not weighted by the length of month" (Normals Calculation Methodology 2020, p. 2). Day-length weighting would have raised values by only 0.03 to 0.13 °F at five sample points.
- From monthly normals. WMO §4.3.3: annual normals "should be calculated from the monthly normals, and not from the individual annual values."
- Completeness. WMO §4.3.3: "If the monthly normal for any of the constituent months of the period of interest is missing, then the multimonth normal should also be considered as missing."
- Area weighting. Equal weights for every touched cell give full weight to cells that lie mostly outside the county, which matters most for small, narrow, and coastal counties. Area weighting matches the Köppen shares (§1).
Implementation: pending; see reviews/02-annual-avg-temperature.md.
Current method
Until the new values are applied, the cell normals are computed as above, but the two later steps differ. Each county month is the unweighted mean of every valid cell the county polygon touches:
T_{c,m}
=
\frac{1}{|G_{c,m}|}
\sum_{i\in G_{c,m}}T_{i,m}^{\mathrm{norm}},
and the annual value averages whichever of the M_c monthly values are
available, normally 12:
T_c(^\circ\mathrm{F})
=
\left(
\frac{1}{M_c}\sum_{m\in V_c}T_{c,m}
\right)\frac{9}{5}+32.
Implementation: _as_monthly_climatology, _zonal_mean, and
build_county_records in scripts/build_county_climate_data.py.
3. Diurnal Temperature Range
Data key: avgDiurnalTempRangeF
For each day with paired county Tmax and Tmin:
DTR_{c,d}=T^{\max}_{c,d}-T^{\min}_{c,d}.
Days with a missing input or a negative range are excluded. The final metric is the mean across all retained days in 1991–2020, followed by conversion of a Celsius temperature difference to a Fahrenheit difference:
\overline{DTR}_c(^\circ\mathrm{F})
=
\frac{9}{5}
\left(
\frac{1}{N_c}\sum_{d\in V_c}DTR_{c,d}
\right).
There is correctly no +32 term when converting a temperature difference.
The output artifact stores two decimal places.
Implementation: build_diurnal_temperature_range, c_delta_to_f_delta, and
write_diurnal_temperature_range_csv in
scripts/build_county_diurnal_temperature_range.py.
4. Annual Extreme Temperature Days
Data key: absoluteExtremeDays
A valid paired day is counted once when either the hot or cold absolute threshold is met:
E_{c,y}
=
\sum_{d\in V_{c,y}}
\mathbf{1}\!\left[
T^{\max}_{c,d}\ge95^\circ\mathrm{F}
\;\lor\;
T^{\min}_{c,d}\le0^\circ\mathrm{F}
\right].
The county filter is the arithmetic mean of the yearly counts:
E_c=\frac{1}{Y_c}\sum_{y\in V_c}E_{c,y}.
The checked-in data uses 1991–2025. The Boolean OR means that a hypothetical day meeting both conditions is still counted only once. A year is included when an annual record exists; counts are not normalized to 365 or 366 valid days.
Implementation: build_annual_counts, _average_or_none, and
write_comparison_csv in scripts/build_county_locally_extreme_data.py.
5. Annual 90 F+ Heat Index Days
Data key: humidHeatDays
The daily proxy pairs NOAA nClimGrid-Daily county Tmax, T, with the gridMET
county daily minimum relative humidity, R. Relative humidity is clipped to
[0,100].
The NWS simple Heat Index estimate is calculated in two steps:
S=\frac{1}{2}\left[T+61+1.2(T-68)+0.094R\right],
HI_s=\frac{S+T}{2}.
When HI_s\ge80^\circ\mathrm{F}, the Rothfusz regression is used:
\begin{aligned}
HI_r={}&-42.379+2.04901523T+10.14333127R-0.22475541TR\\
&-0.00683783T^2-0.05481717R^2+0.00122874T^2R\\
&+0.00085282TR^2-0.00000199T^2R^2.
\end{aligned}
For R<13 and 80\le T\le112, subtract:
A_{low}
=
\frac{13-R}{4}
\sqrt{\max\left(\frac{17-|T-95|}{17},0\right)}.
For R>85 and 80\le T\le87, add:
A_{high}=\frac{R-85}{10}\frac{87-T}{5}.
Thus:
HI(T,R)=
\begin{cases}
HI_s, & HI_s<80,\\
HI_r-A_{low}, & HI_s\ge80 \text{ and the low-RH condition holds},\\
HI_r+A_{high}, & HI_s\ge80 \text{ and the high-RH condition holds},\\
HI_r, & \text{otherwise}.
\end{cases}
The yearly and long-term values are:
H_{c,y}=\sum_{d\in V_{c,y}}
\mathbf{1}[HI(T^{\max}_{c,d},R^{\min}_{c,d})\ge90],
H_c=\frac{1}{Y_c}\sum_{y\in V_c}H_{c,y}.
The period is 1991–2020. A year with at least one valid paired day contributes equally to the final average; there is no completeness adjustment.
Implementation: _heat_index_f and summarize in
scripts/summarize_county_gridmet_humidity.py.
6. Annual Precipitation (Normals)
Data key: annualPrecipIn
Each monthly county total, P_{c,m}, is the unweighted mean of valid raster
cells touched by the county. Annual precipitation is the sum of available
monthly totals, converted from millimeters to inches:
P_c(\mathrm{in})
=
\frac{1}{25.4}
\sum_{m\in V_c}P_{c,m}(\mathrm{mm}).
The stored value is rounded to 0.1 inch. The implementation only requires one valid month, so missing months produce a partial annual sum rather than a blank.
Implementation: _zonal_mean and build_county_records in
scripts/build_county_climate_data.py.
7. Seasonality Index
Data key: seasonalityIndex
The filter is the coefficient of variation of available monthly precipitation totals. First calculate the monthly mean and population standard deviation:
\mu_c=\frac{1}{M_c}\sum_{m\in V_c}P_{c,m},
\sigma_c=
\sqrt{
\frac{1}{M_c}
\sum_{m\in V_c}(P_{c,m}-\mu_c)^2
}.
Then:
SI_c=
\operatorname{round}\!\left(
\operatorname{clip}\!\left(
100\frac{\sigma_c}{\mu_c},0,100
\right)
\right).
If \mu_c\le0, the index is set to zero. The value is stored as an integer.
Implementation: build_county_records in
scripts/build_county_climate_data.py.
8. Wettest Month
Data key: wettestPrecipMonth
W_c=\underset{m\in V_c}{\arg\max}\;P_{c,m}.
Missing monthly values are ignored. An exact tie resolves to the earliest tied
month because numpy.nanargmax returns the first occurrence.
Implementation: build_precip_month_lookup in
scripts/apply_precipitation_month_metrics_to_climate_data.py.
9. Driest Month
Data key: driestPrecipMonth
D_c=\underset{m\in V_c}{\arg\min}\;P_{c,m}.
Missing monthly values are ignored. An exact tie likewise resolves to the earliest tied month.
Implementation: build_precip_month_lookup in
scripts/apply_precipitation_month_metrics_to_climate_data.py.
10. Summer Specific Humidity
Data key: avgSummerSpecificHumidityGKg
gridMET cells whose centers fall inside a county are weighted by the cosine of their latitude to approximate their relative surface areas on a latitude–longitude grid:
q_{c,d}
=
\frac{
\sum_{i\in G_{c,d}}q_{i,d}\cos(\phi_i)
}{
\sum_{i\in G_{c,d}}\cos(\phi_i)
}.
The metric averages the valid daily county values for June, July, and August, then converts kg/kg to g/kg:
q_c^{\mathrm{summer}}
=
1000\left(
\frac{1}{N_c}
\sum_{d\in V_c,\;m(d)\in\{6,7,8\}}q_{c,d}
\right).
If no grid-cell center falls inside a county, the nearest grid cell to an interior representative point is used.
Implementation: _build_county_grid_map, _county_means_chunk, and
summarize in scripts/summarize_county_gridmet_humidity.py.
11. Mean Daily Global Horizontal Radiation (GHI)
Data key: meanDailyGlobalHorizontalRadiationKwhM2Day
For each NSRDB site, the current 60-minute, non-leap-year TMY data is summarized as:
G_s
=
\frac{\sum_h GHI_{s,h}}{1000\times365}
\quad\mathrm{kWh/m^2/day}.
For a county with polygon archive coverage:
G_c
=
\frac{\sum_s A_{c,s}G_s}{\sum_s A_{c,s}},
where A_{c,s} is the estimated overlap area between the county geometry and
the 4 km square grid cell centered on site s. A representative-point value
is used when a polygon summary is unavailable.
Implementation: site_average_daily_ghi and area_weighted_average in
scripts/summarize_nsrdb_county_polygon_archives.py.
12. Clear-Sky GHI Reduction Index
Data key: clearSkyGhiReductionIndex
Only rows with valid observed and clear-sky GHI and
CSGHI_{s,h}\ge50\;\mathrm{W/m^2} are treated as daylight rows. For each
retained row:
r_{s,h}
=
\operatorname{clip}\!\left(
\frac{GHI_{s,h}}{CSGHI_{s,h}},0,1
\right).
The site reduction index is:
R_s=1-\frac{1}{N_s}\sum_{h\in V_s}r_{s,h}.
The polygon county value is overlap-area-weighted:
R_c=\frac{\sum_s A_{c,s}R_s}{\sum_s A_{c,s}}.
A representative-point index is used where a polygon summary is unavailable.
This definition is the mean of time-row ratios; it is not generally equal to
1-\sum GHI/\sum CSGHI.
Implementation: site_cloud_metrics, weighted_metric, and
summarize_archives in
scripts/summarize_nsrdb_county_polygon_cloud_archives.py.
Review findings
Review findings and their status are kept with the filter reviews in reviews/: findings specific to one filter in that filter's file, and findings that affect several filters in reviews/00-cross-filter.md. Findings keep their original numbers; reviews/README.md lists every one.