Refine project documentation and track metric files

Close gaps found in a review of the documentation:
- Track data/metrics/ and data/metric_sources.json in git so the data
  checker passes on a fresh clone (finding 23; decision logged)
- State the Alaska and Hawaii coverage gap in the README limitations and
  extend finding 12
- File findings 24-26: the data-sources doc lacks gridMET and several
  pipeline commands; wettest/driest month are computed twice; solar GHI
  is written by the extreme-temperature apply step
- Add the stale "fallback values" note to filter 2's tasks

Tidy the document system:
- Add docs/reviews/README.md with the numbering rules and a finding index
- Rename koppen-mixed-display-plan.md to koppen-mixed-display.md and fix
  its stale Puerto Rico and "stage 5" text
- Add the precipitation-month step to the README enrichment list
- Describe the Current method / Previous method pattern in plan section 5
- Add CLAUDE.md with the project guardrails and doc layout

Format filter-calculations.md so it renders on GitHub and in VS Code:
inline math uses $...$, ranges use en dashes, and implementation
references name functions instead of line numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-15 17:31:26 -04:00
co-authored by Claude Opus 5
parent e855d583e3
commit 867a07cecb
20 changed files with 3519 additions and 121 deletions
+73 -74
View File
@@ -6,20 +6,20 @@ for the 12 county filters currently exposed by the climate explorer.
## Scope and notation
The current metric list is defined in `app.js` under `METRICS`. Unless noted
otherwise, long-term climate metrics use the 1991--2020 reference period.
otherwise, long-term climate metrics use the 1991–2020 reference period.
| Symbol | Meaning |
| --- | --- |
| \(c\) | County |
| \(i\) | Raster or model grid cell |
| \(m\) | Calendar month |
| \(d\) | Calendar day |
| \(h\) | Hour or NSRDB time row |
| \(y\) | Year |
| \(G_c\) | Valid grid cells assigned to county \(c\) |
| \(V_c\) | Valid observations for county \(c\) |
| \(\mathbf{1}[A]\) | 1 when condition \(A\) is true; otherwise 0 |
| \(\operatorname{clip}(x,a,b)\) | Restrict \(x\) to the interval \([a,b]\) |
| $c$ | County |
| $i$ | Raster or model grid cell |
| $m$ | Calendar month |
| $d$ | Calendar day |
| $h$ | Hour or NSRDB time row |
| $y$ | Year |
| $G_c$ | Valid grid cells assigned to county $c$ |
| $V_c$ | Valid observations for county $c$ |
| $\mathbf{1}[A]$ | 1 when condition $A$ is true; otherwise 0 |
| $\operatorname{clip}(x,a,b)$ | Restrict $x$ to the interval $[a,b]$ |
Missing values are omitted from means unless a metric-specific rule below says
otherwise.
@@ -60,7 +60,7 @@ in `app.js`.
| Temperature & Extremes | Annual Extreme Temperature Days | `absoluteExtremeDays` | Numeric | Days/year |
| Temperature & Extremes | Annual 90 F+ Heat Index Days | `humidHeatDays` | Numeric | Days/year |
| Precipitation & Moisture | Annual Precipitation (Normals) | `annualPrecipIn` | Numeric | Inches/year |
| Precipitation & Moisture | Seasonality Index | `seasonalityIndex` | Numeric | 0--100 index |
| Precipitation & Moisture | Seasonality Index | `seasonalityIndex` | Numeric | 0–100 index |
| Precipitation & Moisture | Wettest Month | `wettestPrecipMonth` | Categorical | Month |
| Precipitation & Moisture | Driest Month | `driestPrecipMonth` | Categorical | Month |
| Precipitation & Moisture | Summer Specific Humidity | `avgSummerSpecificHumidityGKg` | Numeric | g/kg |
@@ -76,9 +76,9 @@ There is no continuous numerical score. Each county is classified from the
share of its land covered by each Köppen class. The rule was adopted on
2026-09-12 and applied to `data/climate-data.csv` on 2026-09-13.
Let \(s_{c,k}\) be the share of county \(c\)'s land area covered by class \(k\).
Let $s_{c,k}$ be the share of county $c$'s land area covered by class $k$.
Each valid raster cell is weighted by the area of the cell that lies inside the
county, \(a_{c,i}\); ocean and no-data cells are excluded:
county, $a_{c,i}$; ocean and no-data cells are excluded:
$$
s_{c,k}
@@ -86,9 +86,9 @@ s_{c,k}
\frac{\sum_{i}a_{c,i}\,\mathbf{1}[K_i=k]}{\sum_{i}a_{c,i}}.
$$
Rank the classes so that \(s_{c,(1)}\ge s_{c,(2)}\ge\cdots\), with
\(s_{c,(2)}=0\) when only one class is present. A county is **predominantly**
class \(k_{(1)}\), shown as a single color, if and only if both conditions hold:
Rank the classes so that $s_{c,(1)}\ge s_{c,(2)}\ge\cdots$, with
$s_{c,(2)}=0$ when only one class is present. A county is **predominantly**
class $k_{(1)}$, shown as a single color, if and only if both conditions hold:
$$
K_c=
@@ -100,19 +100,19 @@ $$
The gap is measured in percentage points. A county that fails either condition
is classified as **Mixed** (shown as "Mixed Climate"). For Mixed counties,
\(p_c=k_{(1)}\) and \(q_c=k_{(2)}\) are stored in `koppenPrimaryClass` and
$p_c=k_{(1)}$ and $q_c=k_{(2)}$ are stored in `koppenPrimaryClass` and
`koppenSecondaryClass`, and the map draws the county with stripes of those two
classes (see [koppen-mixed-display-plan.md](koppen-mixed-display-plan.md)).
classes (see [koppen-mixed-display.md](koppen-mixed-display.md)).
Both columns are blank for predominant counties. A county with no valid raster
cells is left blank; none in the 50 states and DC is.
Each county is read from a small raster window around its polygon. Each cell is
split into 16 × 16 sub-cells to estimate the fraction inside the county, and
scaled by \(\cos(\text{latitude})\) for its true surface area. A county whose
scaled by $\cos(\text{latitude})$ for its true surface area. A county whose
polygon crosses the 180th meridian (Aleutians West, AK) is split into one piece
on each side, and each piece is read from its own window.
**Filter.** Choosing a class \(F\) shows counties that are predominantly that
**Filter.** Choosing a class $F$ shows counties that are predominantly that
class and Mixed counties where it is the primary or secondary class; the
"Mixed Climate" option shows every Mixed county:
@@ -136,13 +136,13 @@ condition catches near 50/50 splits, such as Schenectady, NY (Dfb 50.1%,
Dfa 49.9%), where a single label would rest on a margin of a few tenths of a
point. Map-unit purity standards from other fields were considered and
rejected: the FAO Land Cover Classification System treats a unit as single
only above 80%, and USDA soil survey consociations allow roughly 15--25%
only above 80%, and USDA soil survey consociations allow roughly 15–25%
dissimilar inclusions. Applied to counties, those thresholds would mark about
40--47% of the map area as mixed. A published county-level Köppen dataset
40–47% of the map area as mixed. A published county-level Köppen dataset
(Audirac, Harvard Dataverse, 2024) uses the plurality class and reports the
share of every class, without a threshold.
**Results.** Using the Beck et al. 2023 1991--2020 1 km raster, for the 3,143
**Results.** Using the Beck et al. 2023 1991–2020 1 km raster, for the 3,143
counties in the 50 states and DC:
| Classification | Counties | Share of counties | Share of map area |
@@ -194,7 +194,7 @@ the Köppen apply step must run after it.
**Data key:** `avgTempF`
The annual value is a 1991--2020 climatological standard normal: the mean of
The annual value is a 1991–2020 climatological standard normal: the mean of
the 12 monthly normals of NOAA nClimGrid-Monthly average temperature, averaged
over each county's area. The definition was adopted on 2026-09-15. It is not
yet applied to `data/climate-data.csv`, whose values still use the current
@@ -206,8 +206,8 @@ the present. Its average temperature, `tavg`, is the mean of maximum and
minimum temperature, (Tmax + Tmin)/2, not a 24-hour mean. Alaska and Hawaii
are outside the grid, so their counties are blank.
**Cell normals.** For each grid cell \(i\) and calendar month \(m\), the
normal is the mean over the 1991--2020 years with a valid value:
**Cell normals.** For each grid cell $i$ and calendar month $m$, the
normal is the mean over the 1991–2020 years with a valid value:
$$
T_{i,m}^{\mathrm{norm}}
@@ -224,7 +224,7 @@ values can differ slightly from NCEI's published grids. The result is not
NCEI's station-based U.S. Climate Normals product.
**County monthly values.** Each cell is weighted by the area of the cell that
lies inside the county, \(a_{c,i}\), estimated as for Köppen (§1). Cells
lies inside the county, $a_{c,i}$, estimated as for Köppen (§1). Cells
without data, such as ocean, are excluded:
$$
@@ -234,7 +234,7 @@ T_{c,m}
$$
**Annual value.** Every month has equal weight. The value is defined only when
all 12 monthly values \(T_{c,m}\) exist; otherwise the county is blank:
all 12 monthly values $T_{c,m}$ exist; otherwise the county is blank:
$$
T_c(^\circ\mathrm{F})
@@ -248,7 +248,7 @@ Values are kept at full precision until the stored value is rounded to
0.1 °F.
**Rationale.** The method follows the WMO rules for annual normals, which
NOAA also applies to its 1991--2020 Normals:
NOAA also applies to its 1991–2020 Normals:
- *Equal month weights.* For a mean, WMO-No. 1203 §4.3.3(a) defines the annual
normal as "the mean of the monthly normals", and its footnote says weighting
@@ -286,7 +286,7 @@ T_{c,m}
\sum_{i\in G_{c,m}}T_{i,m}^{\mathrm{norm}},
$$
and the annual value averages whichever of the \(M_c\) monthly values are
and the annual value averages whichever of the $M_c$ monthly values are
available, normally 12:
$$
@@ -297,9 +297,8 @@ T_c(^\circ\mathrm{F})
\right)\frac{9}{5}+32.
$$
Implementation: `scripts/build_county_climate_data.py:124--166`,
`scripts/build_county_climate_data.py:206--221`, and
`scripts/build_county_climate_data.py:478--514`.
Implementation: `_as_monthly_climatology`, `_zonal_mean`, and
`build_county_records` in `scripts/build_county_climate_data.py`.
## 3. Diurnal Temperature Range
@@ -312,7 +311,7 @@ DTR_{c,d}=T^{\max}_{c,d}-T^{\min}_{c,d}.
$$
Days with a missing input or a negative range are excluded. The final metric is
the mean across all retained days in 1991--2020, followed by conversion of a
the mean across all retained days in 1991–2020, followed by conversion of a
Celsius temperature *difference* to a Fahrenheit difference:
$$
@@ -324,11 +323,12 @@ $$
\right).
$$
There is correctly no \(+32\) term when converting a temperature difference.
There is correctly no $+32$ term when converting a temperature difference.
The output artifact stores two decimal places.
Implementation: `scripts/build_county_diurnal_temperature_range.py:46--104`
and `scripts/build_county_diurnal_temperature_range.py:132--151`.
Implementation: `build_diurnal_temperature_range`, `c_delta_to_f_delta`, and
`write_diurnal_temperature_range_csv` in
`scripts/build_county_diurnal_temperature_range.py`.
## 4. Annual Extreme Temperature Days
@@ -354,21 +354,20 @@ $$
E_c=\frac{1}{Y_c}\sum_{y\in V_c}E_{c,y}.
$$
The checked-in data uses 1991--2025. The Boolean OR means that a hypothetical
The checked-in data uses 1991–2025. The Boolean OR means that a hypothetical
day meeting both conditions is still counted only once. A year is included when
an annual record exists; counts are not normalized to 365 or 366 valid days.
Implementation: `scripts/build_county_locally_extreme_data.py:473--540`,
`scripts/build_county_locally_extreme_data.py:669--674`, and
`scripts/build_county_locally_extreme_data.py:705--722`.
Implementation: `build_annual_counts`, `_average_or_none`, and
`write_comparison_csv` in `scripts/build_county_locally_extreme_data.py`.
## 5. Annual 90 F+ Heat Index Days
**Data key:** `humidHeatDays`
The daily proxy pairs NOAA nClimGrid-Daily county Tmax, \(T\), with the gridMET
county daily minimum relative humidity, \(R\). Relative humidity is clipped to
\([0,100]\).
The daily proxy pairs NOAA nClimGrid-Daily county Tmax, $T$, with the gridMET
county daily minimum relative humidity, $R$. Relative humidity is clipped to
$[0,100]$.
The NWS simple Heat Index estimate is calculated in two steps:
@@ -380,7 +379,7 @@ $$
HI_s=\frac{S+T}{2}.
$$
When \(HI_s\ge80^\circ\mathrm{F}\), the Rothfusz regression is used:
When $HI_s\ge80^\circ\mathrm{F}$, the Rothfusz regression is used:
$$
\begin{aligned}
@@ -390,7 +389,7 @@ HI_r={}&-42.379+2.04901523T+10.14333127R-0.22475541TR\\
\end{aligned}
$$
For \(R<13\) and \(80\le T\le112\), subtract:
For $R<13$ and $80\le T\le112$, subtract:
$$
A_{low}
@@ -399,7 +398,7 @@ A_{low}
\sqrt{\max\left(\frac{17-|T-95|}{17},0\right)}.
$$
For \(R>85\) and \(80\le T\le87\), add:
For $R>85$ and $80\le T\le87$, add:
$$
A_{high}=\frac{R-85}{10}\frac{87-T}{5}.
@@ -428,17 +427,17 @@ $$
H_c=\frac{1}{Y_c}\sum_{y\in V_c}H_{c,y}.
$$
The period is 1991--2020. A year with at least one valid paired day contributes
The period is 1991–2020. A year with at least one valid paired day contributes
equally to the final average; there is no completeness adjustment.
Implementation: `scripts/summarize_county_gridmet_humidity.py:439--474` and
`scripts/summarize_county_gridmet_humidity.py:547--617`.
Implementation: `_heat_index_f` and `summarize` in
`scripts/summarize_county_gridmet_humidity.py`.
## 6. Annual Precipitation (Normals)
**Data key:** `annualPrecipIn`
Each monthly county total, \(P_{c,m}\), is the unweighted mean of valid raster
Each monthly county total, $P_{c,m}$, is the unweighted mean of valid raster
cells touched by the county. Annual precipitation is the sum of available
monthly totals, converted from millimeters to inches:
@@ -452,8 +451,8 @@ $$
The stored value is rounded to 0.1 inch. The implementation only requires one
valid month, so missing months produce a partial annual sum rather than a blank.
Implementation: `scripts/build_county_climate_data.py:401--412` and
`scripts/build_county_climate_data.py:478--511`.
Implementation: `_zonal_mean` and `build_county_records` in
`scripts/build_county_climate_data.py`.
## 7. Seasonality Index
@@ -485,9 +484,10 @@ SI_c=
\right).
$$
If \(\mu_c\le0\), the index is set to zero. The value is stored as an integer.
If $\mu_c\le0$, the index is set to zero. The value is stored as an integer.
Implementation: `scripts/build_county_climate_data.py:487--521`.
Implementation: `build_county_records` in
`scripts/build_county_climate_data.py`.
## 8. Wettest Month
@@ -500,8 +500,8 @@ $$
Missing monthly values are ignored. An exact tie resolves to the earliest tied
month because `numpy.nanargmax` returns the first occurrence.
Implementation:
`scripts/apply_precipitation_month_metrics_to_climate_data.py:56--90`.
Implementation: `build_precip_month_lookup` in
`scripts/apply_precipitation_month_metrics_to_climate_data.py`.
## 9. Driest Month
@@ -514,15 +514,15 @@ $$
Missing monthly values are ignored. An exact tie likewise resolves to the
earliest tied month.
Implementation:
`scripts/apply_precipitation_month_metrics_to_climate_data.py:56--90`.
Implementation: `build_precip_month_lookup` in
`scripts/apply_precipitation_month_metrics_to_climate_data.py`.
## 10. Summer Specific Humidity
**Data key:** `avgSummerSpecificHumidityGKg`
gridMET cells whose centers fall inside a county are weighted by the cosine of
their latitude to approximate their relative surface areas on a latitude--longitude
their latitude to approximate their relative surface areas on a latitude–longitude
grid:
$$
@@ -550,9 +550,8 @@ $$
If no grid-cell center falls inside a county, the nearest grid cell to an
interior representative point is used.
Implementation: `scripts/summarize_county_gridmet_humidity.py:297--380`,
`scripts/summarize_county_gridmet_humidity.py:409--434`, and
`scripts/summarize_county_gridmet_humidity.py:555--605`.
Implementation: `_build_county_grid_map`, `_county_means_chunk`, and
`summarize` in `scripts/summarize_county_gridmet_humidity.py`.
## 11. Mean Daily Global Horizontal Radiation (GHI)
@@ -576,19 +575,19 @@ G_c
\frac{\sum_s A_{c,s}G_s}{\sum_s A_{c,s}},
$$
where \(A_{c,s}\) is the estimated overlap area between the county geometry and
the 4 km square grid cell centered on site \(s\). A representative-point value
where $A_{c,s}$ is the estimated overlap area between the county geometry and
the 4 km square grid cell centered on site $s$. A representative-point value
is used when a polygon summary is unavailable.
Implementation: `scripts/summarize_nsrdb_county_polygon_archives.py:189--192`
and `scripts/summarize_nsrdb_county_polygon_archives.py:335--378`.
Implementation: `site_average_daily_ghi` and `area_weighted_average` in
`scripts/summarize_nsrdb_county_polygon_archives.py`.
## 12. Clear-Sky GHI Reduction Index
**Data key:** `clearSkyGhiReductionIndex`
Only rows with valid observed and clear-sky GHI and
\(CSGHI_{s,h}\ge50\;\mathrm{W/m^2}\) are treated as daylight rows. For each
$CSGHI_{s,h}\ge50\;\mathrm{W/m^2}$ are treated as daylight rows. For each
retained row:
$$
@@ -613,11 +612,11 @@ $$
A representative-point index is used where a polygon summary is unavailable.
This definition is the mean of time-row ratios; it is not generally equal to
\(1-\sum GHI/\sum CSGHI\).
$1-\sum GHI/\sum CSGHI$.
Implementation:
`scripts/summarize_nsrdb_county_polygon_cloud_archives.py:138--218` and
`scripts/summarize_nsrdb_county_polygon_cloud_archives.py:222--304`.
Implementation: `site_cloud_metrics`, `weighted_metric`, and
`summarize_archives` in
`scripts/summarize_nsrdb_county_polygon_cloud_archives.py`.
## Review findings
@@ -625,4 +624,4 @@ Review findings and their status are kept with the filter reviews in
[reviews/](reviews/): findings specific to one filter in that filter's file,
and findings that affect several filters in
[reviews/00-cross-filter.md](reviews/00-cross-filter.md). Findings keep their
original numbers.
original numbers; [reviews/README.md](reviews/README.md) lists every one.