Files
Climate-Mood-Analysis/docs/filter-calculations.md
T
KnouandClaude Opus 5 4d2b3e3d44 Complete Köppen-Geiger filter review with Mixed climate class
Classify each county by area-weighted Köppen class shares: a county is
predominantly its top class when that class covers at least 50% of its
land and leads the runner-up by at least 5 percentage points; otherwise
it is Mixed (133 of 3,143 counties in the 50 states and DC).

- Add build_county_koppen_metric.py (writes data/metrics/koppen.csv) and
  apply_koppen_metric_to_climate_data.py (writes koppenZone plus
  koppenPrimaryClass/koppenSecondaryClass for Mixed counties).
- Move shared helpers into scripts/common/ (county loading, Köppen
  legend, area-weighted raster shares); fix the 180th-meridian raster
  window for Aleutians West.
- Add check_climate_data.py to validate the app CSV.
- Draw Mixed counties in app.js as diagonal stripes of their top two
  classes, fixed to the ground and following the map at every zoom, with
  a crossfade only when the stripe size changes. Filtering a class also
  matches Mixed counties where it is primary or secondary.
- Document the rule, display, and pipeline plan in docs/ and update the
  README and data-source notes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 02:54:25 -04:00

20 KiB
Raw Blame History

County Filter Calculations

This document records the equations, implementation behavior, and review notes for the 12 county filters currently exposed by the climate explorer.

Scope and notation

The current metric list is defined in app.js under METRICS. Unless noted otherwise, long-term climate metrics use the 1991--2020 reference period.

Symbol Meaning
(c) County
(i) Raster or model grid cell
(m) Calendar month
(d) Calendar day
(h) Hour or NSRDB time row
(y) Year
(G_c) Valid grid cells assigned to county (c)
(V_c) Valid observations for county (c)
(\mathbf{1}[A]) 1 when condition (A) is true; otherwise 0
(\operatorname{clip}(x,a,b)) Restrict (x) to the interval ([a,b])

Missing values are omitted from means unless a metric-specific rule below says otherwise.

Browser filter predicate

For every numeric metric, a county is active when its value is finite and falls inside the selected range, including both endpoints:


\operatorname{passes}(c) \iff L \le x_c \le U.

For a categorical metric:


\operatorname{passes}(c) \iff
\left(F=\text{All}\right) \lor \left(x_c=F\right).

The Köppen-Geiger filter also matches Mixed counties by their top two classes; see §1.

Counties with missing categorical values, null numeric values, or non-finite numeric values do not pass. The numeric slider limits are derived from the loaded data and rounded outward by each metric's configured boundsStep.

Implementation: shouldFeaturePassFilter and getUniqueCategoryValuesInData in app.js.

Filter inventory

Group UI label Data key Type Display unit
Climate Classification Köppen-Geiger Climate Class koppenZone Categorical Class
Temperature & Extremes Annual Avg Temperature (Normals) avgTempF Numeric °F
Temperature & Extremes Diurnal Temperature Range avgDiurnalTempRangeF Numeric °F difference
Temperature & Extremes Annual Extreme Temperature Days absoluteExtremeDays Numeric Days/year
Temperature & Extremes Annual 90 F+ Heat Index Days humidHeatDays Numeric Days/year
Precipitation & Moisture Annual Precipitation (Normals) annualPrecipIn Numeric Inches/year
Precipitation & Moisture Seasonality Index seasonalityIndex Numeric 0--100 index
Precipitation & Moisture Wettest Month wettestPrecipMonth Categorical Month
Precipitation & Moisture Driest Month driestPrecipMonth Categorical Month
Precipitation & Moisture Summer Specific Humidity avgSummerSpecificHumidityGKg Numeric g/kg
Solar Resource Mean Daily Global Horizontal Radiation (GHI) meanDailyGlobalHorizontalRadiationKwhM2Day Numeric kWh/m²/day
Solar Resource Clear-Sky GHI Reduction Index clearSkyGhiReductionIndex Numeric Ratio

1. Köppen-Geiger Climate Class

Data keys: koppenZone; koppenPrimaryClass and koppenSecondaryClass for Mixed counties

There is no continuous numerical score. Each county is classified from the share of its land covered by each Köppen class. The rule was adopted on 2026-09-12 and applied to data/climate-data.csv on 2026-09-13.

Let (s_{c,k}) be the share of county (c)'s land area covered by class (k). Each valid raster cell is weighted by the area of the cell that lies inside the county, (a_{c,i}); ocean and no-data cells are excluded:


s_{c,k}
=
\frac{\sum_{i}a_{c,i}\,\mathbf{1}[K_i=k]}{\sum_{i}a_{c,i}}.

Rank the classes so that (s_{c,(1)}\ge s_{c,(2)}\ge\cdots), with (s_{c,(2)}=0) when only one class is present. A county is predominantly class (k_{(1)}), shown as a single color, if and only if both conditions hold:


K_c=
\begin{cases}
k_{(1)}, & s_{c,(1)}\ge0.50 \;\text{and}\; s_{c,(1)}-s_{c,(2)}\ge0.05,\\
\text{Mixed}, & \text{otherwise}.
\end{cases}

The gap is measured in percentage points. A county that fails either condition is classified as Mixed (shown as "Mixed Climate"). For Mixed counties, (p_c=k_{(1)}) and (q_c=k_{(2)}) are stored in koppenPrimaryClass and koppenSecondaryClass, and the map draws the county with stripes of those two classes (see koppen-mixed-display-plan.md). Both columns are blank for predominant counties. A county with no valid raster cells is left blank; none in the 50 states and DC is.

Each county is read from a small raster window around its polygon. Each cell is split into 16 × 16 sub-cells to estimate the fraction inside the county, and scaled by (\cos(\text{latitude})) for its true surface area. A county whose polygon crosses the 180th meridian (Aleutians West, AK) is split into one piece on each side, and each piece is read from its own window.

Filter. Choosing a class (F) shows counties that are predominantly that class and Mixed counties where it is the primary or secondary class; the "Mixed Climate" option shows every Mixed county:


\operatorname{passes}(c) \iff
\left(F=\text{All}\right) \lor \left(K_c=F\right) \lor
\left(K_c=\text{Mixed} \land F\in\{p_c,q_c\}\right).

Implementation: scripts/build_county_koppen_metric.py (rank_class_shares, classify, build_koppen_records) writes data/metrics/koppen.csv; area weighting, raster windows, and the 180th-meridian split are in scripts/common/county_zonal_stats.py; scripts/apply_koppen_metric_to_climate_data.py copies the three columns into data/climate-data.csv; the filter is shouldFeaturePassFilter in app.js.

Rationale. No published standard defines when an area is predominantly one Köppen class; the classification is defined per grid cell. The 50% condition means the label describes more than half of the county's land. The 5-point gap condition catches near 50/50 splits, such as Schenectady, NY (Dfb 50.1%, Dfa 49.9%), where a single label would rest on a margin of a few tenths of a point. Map-unit purity standards from other fields were considered and rejected: the FAO Land Cover Classification System treats a unit as single only above 80%, and USDA soil survey consociations allow roughly 15--25% dissimilar inclusions. Applied to counties, those thresholds would mark about 40--47% of the map area as mixed. A published county-level Köppen dataset (Audirac, Harvard Dataverse, 2024) uses the plurality class and reports the share of every class, without a threshold.

Results. Using the Beck et al. 2023 1991--2020 1 km raster, for the 3,143 counties in the 50 states and DC:

Classification Counties Share of counties Share of map area
Predominant (single class) 3,010 95.8% 85.6%
Mixed climate 133 4.2% 14.4%

Of the 133 mixed counties, 111 have no class covering 50% or more (37 of these also have a gap under 5 points), and 22 have a majority class whose runner-up is within 5 points. They are concentrated in the mountain West and Alaska: Colorado and Alaska (13 each), California and Montana (12 each), Washington (10), Idaho (9), and Utah (8). No county's top three classes fall within 2 points of one another, so no additional mixed categories are needed.

Compared with the previous method (below). Every county whose label changed became Mixed; no county moved to a different single class. The touched-cell counts and the area-weighted shares pick a different plurality winner in only 3 counties (Denver, CO; Schenectady, NY; Hood River, OR), all of which are Mixed under this rule. Applying this rule to touched-cell counts instead of area-weighted shares would classify 9 counties differently, because touched-cell counts give full weight to boundary cells that are mostly outside the county.

Boundary cases. Five counties lie within 0.25 points of a cutoff that decides their outcome: Sanpete, UT (Dfb 49.93%, Mixed), Giles, VA (49.95%, Mixed), Placer, CA (Csa 50.22%), Albany, WY (Dfb 50.24%), and Park, MT (gap 5.23 points, Dfb). Their classification depends on the precision of the area weighting. Aleutians West, AK (02016) spans the antimeridian, so its shares were computed without sub-cell sampling. Puerto Rico is not covered by these figures; it has 9 more Mixed counties and is not shown in the app.

Previous method

Until 2026-09-13, koppenZone was the most frequent valid Köppen raster code among the cells touched by the county polygon:


K_c = \underset{k}{\arg\max}\;
\sum_{i\in G_c}\mathbf{1}[K_i=k].

Cells were equally weighted, regardless of how much of each cell lay inside the county; an exact tie went to the smallest raster code, and a county with no valid cell was assigned Cfa. scripts/build_county_climate_data.py still computes this value until that script is retired (pipeline plan, Phase 3), so the Köppen apply step must run after it.

2. Annual Avg Temperature (Normals)

Data key: avgTempF

If the NOAA source is a historical monthly series, the script first forms a 1991--2020 climatology for each calendar month and cell:


T_{i,m}^{\mathrm{norm}}
=
\frac{1}{Y_{i,m}}
\sum_{y\in V_{i,m}}T_{i,m,y}.

The monthly county value is an unweighted mean of valid touched raster cells:


T_{c,m}
=
\frac{1}{|G_{c,m}|}
\sum_{i\in G_{c,m}}T_{i,m}^{\mathrm{norm}}.

The annual value is the equally weighted mean of the available monthly values, converted from Celsius to Fahrenheit:


T_c(^\circ\mathrm{F})
=
\left(
\frac{1}{M_c}\sum_{m\in V_c}T_{c,m}
\right)\frac{9}{5}+32.

Normally (M_c=12). The stored value is rounded to 0.1 °F.

Implementation: scripts/build_county_climate_data.py:124--166, scripts/build_county_climate_data.py:206--221, and scripts/build_county_climate_data.py:478--514.

3. Diurnal Temperature Range

Data key: avgDiurnalTempRangeF

For each day with paired county Tmax and Tmin:


DTR_{c,d}=T^{\max}_{c,d}-T^{\min}_{c,d}.

Days with a missing input or a negative range are excluded. The final metric is the mean across all retained days in 1991--2020, followed by conversion of a Celsius temperature difference to a Fahrenheit difference:


\overline{DTR}_c(^\circ\mathrm{F})
=
\frac{9}{5}
\left(
\frac{1}{N_c}\sum_{d\in V_c}DTR_{c,d}
\right).

There is correctly no (+32) term when converting a temperature difference. The output artifact stores two decimal places.

Implementation: scripts/build_county_diurnal_temperature_range.py:46--104 and scripts/build_county_diurnal_temperature_range.py:132--151.

4. Annual Extreme Temperature Days

Data key: absoluteExtremeDays

A valid paired day is counted once when either the hot or cold absolute threshold is met:


E_{c,y}
=
\sum_{d\in V_{c,y}}
\mathbf{1}\!\left[
T^{\max}_{c,d}\ge95^\circ\mathrm{F}
\;\lor\;
T^{\min}_{c,d}\le0^\circ\mathrm{F}
\right].

The county filter is the arithmetic mean of the yearly counts:


E_c=\frac{1}{Y_c}\sum_{y\in V_c}E_{c,y}.

The checked-in data uses 1991--2025. The Boolean OR means that a hypothetical day meeting both conditions is still counted only once. A year is included when an annual record exists; counts are not normalized to 365 or 366 valid days.

Implementation: scripts/build_county_locally_extreme_data.py:473--540, scripts/build_county_locally_extreme_data.py:669--674, and scripts/build_county_locally_extreme_data.py:705--722.

5. Annual 90 F+ Heat Index Days

Data key: humidHeatDays

The daily proxy pairs NOAA nClimGrid-Daily county Tmax, (T), with the gridMET county daily minimum relative humidity, (R). Relative humidity is clipped to ([0,100]).

The NWS simple Heat Index estimate is calculated in two steps:


S=\frac{1}{2}\left[T+61+1.2(T-68)+0.094R\right],

HI_s=\frac{S+T}{2}.

When (HI_s\ge80^\circ\mathrm{F}), the Rothfusz regression is used:


\begin{aligned}
HI_r={}&-42.379+2.04901523T+10.14333127R-0.22475541TR\\
&-0.00683783T^2-0.05481717R^2+0.00122874T^2R\\
&+0.00085282TR^2-0.00000199T^2R^2.
\end{aligned}

For (R<13) and (80\le T\le112), subtract:


A_{low}
=
\frac{13-R}{4}
\sqrt{\max\left(\frac{17-|T-95|}{17},0\right)}.

For (R>85) and (80\le T\le87), add:


A_{high}=\frac{R-85}{10}\frac{87-T}{5}.

Thus:


HI(T,R)=
\begin{cases}
HI_s, & HI_s<80,\\
HI_r-A_{low}, & HI_s\ge80 \text{ and the low-RH condition holds},\\
HI_r+A_{high}, & HI_s\ge80 \text{ and the high-RH condition holds},\\
HI_r, & \text{otherwise}.
\end{cases}

The yearly and long-term values are:


H_{c,y}=\sum_{d\in V_{c,y}}
\mathbf{1}[HI(T^{\max}_{c,d},R^{\min}_{c,d})\ge90],

H_c=\frac{1}{Y_c}\sum_{y\in V_c}H_{c,y}.

The period is 1991--2020. A year with at least one valid paired day contributes equally to the final average; there is no completeness adjustment.

Implementation: scripts/summarize_county_gridmet_humidity.py:439--474 and scripts/summarize_county_gridmet_humidity.py:547--617.

6. Annual Precipitation (Normals)

Data key: annualPrecipIn

Each monthly county total, (P_{c,m}), is the unweighted mean of valid raster cells touched by the county. Annual precipitation is the sum of available monthly totals, converted from millimeters to inches:


P_c(\mathrm{in})
=
\frac{1}{25.4}
\sum_{m\in V_c}P_{c,m}(\mathrm{mm}).

The stored value is rounded to 0.1 inch. The implementation only requires one valid month, so missing months produce a partial annual sum rather than a blank.

Implementation: scripts/build_county_climate_data.py:401--412 and scripts/build_county_climate_data.py:478--511.

7. Seasonality Index

Data key: seasonalityIndex

The filter is the coefficient of variation of available monthly precipitation totals. First calculate the monthly mean and population standard deviation:


\mu_c=\frac{1}{M_c}\sum_{m\in V_c}P_{c,m},

\sigma_c=
\sqrt{
\frac{1}{M_c}
\sum_{m\in V_c}(P_{c,m}-\mu_c)^2
}.

Then:


SI_c=
\operatorname{round}\!\left(
\operatorname{clip}\!\left(
100\frac{\sigma_c}{\mu_c},0,100
\right)
\right).

If (\mu_c\le0), the index is set to zero. The value is stored as an integer.

Implementation: scripts/build_county_climate_data.py:487--521.

8. Wettest Month

Data key: wettestPrecipMonth


W_c=\underset{m\in V_c}{\arg\max}\;P_{c,m}.

Missing monthly values are ignored. An exact tie resolves to the earliest tied month because numpy.nanargmax returns the first occurrence.

Implementation: scripts/apply_precipitation_month_metrics_to_climate_data.py:56--90.

9. Driest Month

Data key: driestPrecipMonth


D_c=\underset{m\in V_c}{\arg\min}\;P_{c,m}.

Missing monthly values are ignored. An exact tie likewise resolves to the earliest tied month.

Implementation: scripts/apply_precipitation_month_metrics_to_climate_data.py:56--90.

10. Summer Specific Humidity

Data key: avgSummerSpecificHumidityGKg

gridMET cells whose centers fall inside a county are weighted by the cosine of their latitude to approximate their relative surface areas on a latitude--longitude grid:


q_{c,d}
=
\frac{
\sum_{i\in G_{c,d}}q_{i,d}\cos(\phi_i)
}{
\sum_{i\in G_{c,d}}\cos(\phi_i)
}.

The metric averages the valid daily county values for June, July, and August, then converts kg/kg to g/kg:


q_c^{\mathrm{summer}}
=
1000\left(
\frac{1}{N_c}
\sum_{d\in V_c,\;m(d)\in\{6,7,8\}}q_{c,d}
\right).

If no grid-cell center falls inside a county, the nearest grid cell to an interior representative point is used.

Implementation: scripts/summarize_county_gridmet_humidity.py:297--380, scripts/summarize_county_gridmet_humidity.py:409--434, and scripts/summarize_county_gridmet_humidity.py:555--605.

11. Mean Daily Global Horizontal Radiation (GHI)

Data key: meanDailyGlobalHorizontalRadiationKwhM2Day

For each NSRDB site, the current 60-minute, non-leap-year TMY data is summarized as:


G_s
=
\frac{\sum_h GHI_{s,h}}{1000\times365}
\quad\mathrm{kWh/m^2/day}.

For a county with polygon archive coverage:


G_c
=
\frac{\sum_s A_{c,s}G_s}{\sum_s A_{c,s}},

where (A_{c,s}) is the estimated overlap area between the county geometry and the 4 km square grid cell centered on site (s). A representative-point value is used when a polygon summary is unavailable.

Implementation: scripts/summarize_nsrdb_county_polygon_archives.py:189--192 and scripts/summarize_nsrdb_county_polygon_archives.py:335--378.

12. Clear-Sky GHI Reduction Index

Data key: clearSkyGhiReductionIndex

Only rows with valid observed and clear-sky GHI and (CSGHI_{s,h}\ge50;\mathrm{W/m^2}) are treated as daylight rows. For each retained row:


r_{s,h}
=
\operatorname{clip}\!\left(
\frac{GHI_{s,h}}{CSGHI_{s,h}},0,1
\right).

The site reduction index is:


R_s=1-\frac{1}{N_s}\sum_{h\in V_s}r_{s,h}.

The polygon county value is overlap-area-weighted:


R_c=\frac{\sum_s A_{c,s}R_s}{\sum_s A_{c,s}}.

A representative-point index is used where a polygon summary is unavailable. This definition is the mean of time-row ratios; it is not generally equal to (1-\sum GHI/\sum CSGHI).

Implementation: scripts/summarize_nsrdb_county_polygon_cloud_archives.py:138--218 and scripts/summarize_nsrdb_county_polygon_cloud_archives.py:222--304.

Calculation review findings

  1. Annual temperature weights months equally. February has the same weight as January or July. If the intended label means an average across all days, monthly normals should instead be weighted by the number of days in each month.

  2. Base NOAA aggregation is not area-weighted. Every touched raster cell receives equal weight, including cells that intersect only a small portion of a county. This can matter most for small or narrow counties and along coastlines. Resolved for Köppen on 2026-09-13: class shares are now area-weighted (§1).

  3. Partial precipitation years are accepted. One valid monthly precipitation value is sufficient to produce annualPrecipIn; absent months silently lower the annual sum. Requiring all 12 months, or recording completeness, would be safer.

  4. The Köppen fallback can create false data. A county with no valid raster cells is labeled Cfa instead of missing. A null value plus an audit flag would distinguish missing coverage from a genuine humid-subtropical class. Resolved on 2026-09-13: the Köppen builder leaves such counties blank. The fallback remains only in build_county_climate_data.py, whose Köppen value is replaced by the apply step.

  5. Extreme-day counts are not completeness-normalized. A partially observed year contributes a raw count and receives the same weight as a complete year. Consider requiring a minimum number of valid days or annualizing partial counts explicitly.

  6. The absolute-extreme metric depends on unrelated percentile thresholds. build_annual_counts skips a county when its retired local p95/p05 thresholds are missing, even though the active 95 °F / 0 °F calculation does not require those percentiles. The absolute calculation should be separated from that prerequisite.

  7. Heat Index days are a daily-extrema proxy. Daily Tmax and daily minimum relative humidity are paired even though their observation times may differ. The result should not be described as an observed hourly maximum Heat Index.

  8. Heat-year completeness is permissive. Any year with at least one valid Tmax/RH pair is included in the equal-year average. A minimum valid-day rule would reduce low-biased partial-year counts.

  9. The GHI formula assumes hourly, 365-day input. It is correct for the current 60-minute, leap_day=false requests. If the request interval changes, the energy sum needs an interval-hours multiplier; leap-day handling would also need to change the divisor.

  10. Clear-sky reduction averages ratios rather than energy totals. This is a valid but specific definition. It gives each retained time row equal weight, rather than weighting rows by available clear-sky energy. The label and documentation should retain this distinction.

  11. Spatial weighting is inconsistent across metric families. Base NOAA normals use equal touched-cell weights, Köppen uses area-weighted class shares, gridMET humidity uses (\cos(\phi)) weights on cell centers, and NSRDB polygon metrics use estimated overlap areas. Cross-metric comparisons should account for these different county aggregation methods.

Verification status

This reference was derived from the checked-in calculation and merge scripts, not solely from UI descriptions. No calculation code was changed. The automated test suite was not executed during this review because pytest is not installed in either the system Python environment or the project virtual environment.

Section 1 and findings 2, 4, and 11 were updated on 2026-09-14, after the Köppen classification was reworked and applied.