Document restructuring and the beginnings of Filter 2 changes

Split the pipeline documentation by purpose so each fact has one home:
- docs/pipeline-plan.md keeps the plan, checklist, tracker, and guardrails
- docs/decisions.md holds open decisions and the dated decision log
- docs/reviews/ holds findings and tasks: one file per filter, plus
  00-cross-filter.md for findings that span filters
- scripts/common/README.md holds the shared-helper rules (formerly Phase 2)
- filter-calculations.md now describes calculations only

Filed findings 12-22 from a consistency audit of the app, docs, and scripts.

Filter 1 (Köppen-Geiger): use "Köppen" with the umlaut in all prose, labels,
docstrings, help text, and checker messages (finding 21), and correct the
base build's "majority" docstring (finding 22).

Filter 2 (annual avg temperature): record the adopted definition in
filter-calculations.md §2: equally weighted 1991-2020 monthly normals, per
WMO-No. 1203 and NOAA's 2020 methodology; area-weighted county means; blank
unless all 12 months exist. Code changes for this filter are still pending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-15 16:13:11 -04:00
co-authored by Claude Opus 5
parent a9e722791d
commit e855d583e3
27 changed files with 807 additions and 242 deletions
+121
View File
@@ -0,0 +1,121 @@
# Cross-filter review
Findings that affect more than one filter. Each finding lives in exactly one
review file; the filter reviews it affects link here and record only how it
applied to them. Findings keep their numbers across all review files, and new
findings take the next number wherever they are filed. A finding that turns out
to affect other filters moves here, leaving a link behind.
## Findings
**2. Base NOAA aggregation is not area-weighted.** Every touched raster cell
receives equal weight, including cells that intersect only a small portion
of a county. This can matter most for small or narrow counties and along
coastlines. *Resolved for Köppen on 2026-09-13: class shares are now
area-weighted ([filter-calculations.md](../filter-calculations.md) §1).*
Affects filters 1 (resolved), 2, 6, and 7.
**11. Spatial weighting is inconsistent across metric families.** Base NOAA
normals use equal touched-cell weights, Köppen uses area-weighted class
shares, gridMET humidity uses
\(\cos(\phi)\) weights on cell centers, and NSRDB polygon metrics use
estimated overlap areas. Cross-metric comparisons should account for these
different county aggregation methods.
Affects filters 1, 2, 6, 7, 10, 11, and 12. The planned fix is item 3 of the
target design in [pipeline-plan.md](../pipeline-plan.md): one shared
county-aggregation module used by every raster-based metric.
**12. Alaska and Hawaii are not covered by the NOAA and gridMET sources.**
NOAA nClimGrid and gridMET cover the contiguous U.S. only; the nClimGrid grid
spans latitude 24.56 to 49.35 and longitude −124.69 to −67.02. All 29 Alaska
and 5 Hawaii counties are therefore blank in filters 2–10. Solar GHI and
clear-sky reduction (filters 11 and 12) cover both states.
`check_climate_data.py` allows these blanks (`OUTSIDE_CONUS`). The notes on
the base build in `scripts/county_data_sources.md` say "fallback values are
applied" for counties outside NOAA coverage, but `build_county_climate_data.py`
leaves them blank (lines 499–520).
Affects filters 2–10. Open decision: "Alaska and Hawaii coverage" in
[decisions.md](../decisions.md).
**13. Lexington, VA (51678) is missing from the NOAA daily county files.**
Diurnal temperature range and extreme temperature days are blank for it, and
`check_climate_data.py` allows those blanks (`NOAA_DAILY_MISSING_FIPS`).
Heat-index days instead use the surrounding Rockbridge County (51163) as a
proxy, recorded in `humidHeatSourceFips` and `humidHeatFipsAdjustment`. The
three filters handle the same gap differently.
Affects filters 3, 4, and 5.
**14. Helper functions are duplicated across scripts, and some copies have
drifted.** A 2026-09-13 survey found 19 functions with identical copies in
several scripts and 19 with copies that have drifted apart. Most identical
copies are NSRDB helpers, which belong in an NSRDB module rather than
`common/`; `read_csv_rows` (4 identical copies in `apply_*` scripts) is
cross-source. Drifted copies need a decision on which version is correct
before merging. Notable drifts: `summarize_county_gridmet_humidity.py` has its
own county loader and FIPS normalizer, and the state FIPS table is also copied
in `build_county_representative_points.py`,
`summarize_county_gridmet_humidity.py`, and
`request_nsrdb_county_polygon_archives.py`.
Affects every script that holds a copy. The rules for moving shared code are
in [scripts/common/README.md](../../scripts/common/README.md).
**15. The app's source text for three NOAA metrics names the wrong product.**
In `app.js`, `avgTempF`, `annualPrecipIn`, and `seasonalityIndex` credit
"NOAA NCEI 1991-2020 U.S. Climate Normals" and link the station-based Normals
page. Their values are computed from the nClimGrid-Monthly series averaged
over 1991–2020, which is NCEI's gridded-normals method rather than the station
product. Source 2 in `scripts/county_data_sources.md` likewise says the build
reads monthly normals files, while the build command passes
`data/noaa/nclimgrid/nclimgrid_tavg.nc` and `nclimgrid_prcp.nc`. Found during
the filter 2 review (2026-09-15).
Affects filters 2, 6, and 7. Filter 2's task list covers the `avgTempF` text.
**16. Source links point to retired or missing pages.** Checked 2026-09-15.
In `app.js`, the "NOAA nClimGrid Monthly" link used by wettest and driest month
(`https://www.ncei.noaa.gov/products/land-based-station/nclimgrid`) returns
404, and the "NREL National Solar Radiation Database" link used by solar GHI
and clear-sky GHI reduction (`https://nsrdb.nrel.gov/`) no longer resolves.
NREL's sites have moved to `nlr.gov`: `https://nsrdb.nlr.gov/` loads, and the
NSRDB scripts already call `developer.nlr.gov`. In Source 4 of
`scripts/county_data_sources.md`, `https://developer.nrel.gov/docs/solar/nsrdb/`
also fails (its `developer.nlr.gov` counterpart loads), and
`https://www.nrel.gov/gis/solar-resource-maps` fails with no counterpart at the
same path on `nlr.gov`. Every other link in `app.js`, `README.md`, and
`scripts/county_data_sources.md` loads.
Affects filters 8, 9, 11, and 12.
**17. The data-sources metric list is out of date.** The "Metric definitions
in generated output" list in `scripts/county_data_sources.md` describes
`extremeDays` / `oldExtremeDays` as an audit column kept in the app CSV, but
neither column is in `data/climate-data.csv`. The list has no entry for
`avgDiurnalTempRangeF` or `absoluteExtremeDays`, 2 of the 12 app metrics.
Affects filters 3 and 4.
## Decisions
In [decisions.md](../decisions.md): Shared helpers (2026-09-13), and the open
decision on Alaska and Hawaii coverage.
## Tasks
Fixes for these findings are carried out in the affected filters' reviews.
## Original review
`filter-calculations.md` was derived from the checked-in calculation and merge
scripts, not solely from UI descriptions. No calculation code was changed. The
automated test suite was not executed during this review because `pytest` is
not installed in either the system Python environment or the project virtual
environment. The suite uses `unittest` and runs without pytest:
`python -m unittest discover -s tests`.
Section 1 and findings 2, 4, and 11 were updated on 2026-09-14, after the
Köppen classification was reworked and applied.
+78
View File
@@ -0,0 +1,78 @@
# Filter 1: Köppen-Geiger class
**Status:** Done (2026-09-14): rule applied, Mixed display built, documentation
updated.
**Data keys:** `koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`.
Calculation: [filter-calculations.md](../filter-calculations.md) §1. Display
design and history:
[koppen-mixed-display-plan.md](../koppen-mixed-display-plan.md).
## Findings
**4. The Köppen fallback can create false data.** A county with no valid raster
cells is labeled `Cfa` instead of missing. A null value plus an audit flag
would distinguish missing coverage from a genuine humid-subtropical class.
*Resolved on 2026-09-13: the Köppen builder leaves such counties blank. The
fallback remains only in `build_county_climate_data.py`, whose Köppen
value is replaced by the apply step.*
**21. The documented Köppen labels differ from the app.** `filter-calculations.md`
lists the group as "Climate Classification" and the filter as "Köppen-Geiger
Climate Class"; `app.js` shows "Koppen-Geiger Classification" and
"Koppen-Geiger Climate Class". The app spells Koppen without the umlaut
throughout, while the docs use Köppen. *Resolved on 2026-09-15: every
prose and display use of the name now reads "Köppen", in `app.js`,
`README.md`, `scripts/county_data_sources.md`, script docstrings, help text,
and checker messages (with the matching test expectations), and
`filter-calculations.md` quotes the app's group label, "Köppen-Geiger
Classification". Code identifiers, file names, and data paths such as
`koppenZone`, `normalizeKoppenCode`, and `koppen_geiger_tif/` keep the ASCII
spelling.*
**22. The base build's docstring calls its Köppen value a majority class.**
`build_county_climate_data.py` lists `koppenZone` as the "majority
Koppen-Geiger class", but it assigns the most common class among touched
cells, which need not be a majority. That value is replaced by the Köppen apply
step, and the script is retired in Phase 3. *Resolved on 2026-09-15: the
docstring now reads "largest-share Köppen-Geiger class among touched cells
(replaced by the Köppen apply step)".*
Cross-filter finding 2 (touched-cell aggregation) was resolved for Köppen on
2026-09-13; see [00-cross-filter.md](00-cross-filter.md).
## Decisions
In [decisions.md](../decisions.md): Köppen audit columns, Köppen no-data
fallback, and Applying Köppen to the app CSV (2026-09-12); Metric files and
Mixed climate display (2026-09-13).
## Tasks
Adopted rule: a county is predominantly its top class if and only if that class
covers at least 50% of the county's land and leads the runner-up by at least
5 percentage points; otherwise it is Mixed climate. Expected result for the
50 states and DC: 3,010 predominant, 133 Mixed.
- [x] Investigate low-majority counties and adopt the rule.
- [x] Windowed raster reads and 180th-meridian split.
- [x] Area-weighted class shares (16 × 16 sub-cells per raster cell, scaled by
cos(latitude)) in `scripts/common/county_zonal_stats.py`.
- [x] Apply the 50% / 5-point rule (`scripts/build_county_koppen_metric.py`;
counties with no valid cells are left blank).
- [x] Run the builder to write `data/metrics/koppen.csv` and confirm the
expected 3,010 predominant / 133 Mixed (2026-09-13).
- [x] `koppenZone`-only apply step
(`scripts/apply_koppen_metric_to_climate_data.py`, with `--dry-run`).
- [x] Apply to `data/climate-data.csv` (2026-09-13; 142 counties changed to
Mixed, with `koppenPrimaryClass` and `koppenSecondaryClass` added).
- [x] Allow `Mixed` in `check_climate_data.py`.
- [x] Add a Mixed climate category to `app.js`, drawn as stripes of the
county's top two classes; see
[koppen-mixed-display-plan.md](../koppen-mixed-display-plan.md).
- [x] Replace the plurality description in `filter-calculations.md` §1 and mark
review findings 2 and 4 resolved for Köppen (2026-09-14).
- [x] Update the Köppen descriptions and script lists in `README.md` and
`scripts/county_data_sources.md` (2026-09-14).
- [x] Tests for shares, the rule, boundary cases, and the apply step
(`tests/test_koppen_metric.py`).
+81
View File
@@ -0,0 +1,81 @@
# Filter 2: Annual avg temperature
**Status:** In progress: method decided (2026-09-15).
**Data key:** `avgTempF`. Calculation:
[filter-calculations.md](../filter-calculations.md) §2.
## Findings
**1. Annual temperature weights months equally.** February has the same weight
as January or July. If the intended label means an average across all days,
monthly normals should instead be weighted by the number of days in each
month. *Resolved on 2026-09-15: equal weighting is the WMO and NOAA standard
for annual normals and is kept; see "Annual avg temperature month weighting"
in [decisions.md](../decisions.md).*
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 2 and 11 (county aggregation),
12 (Alaska and Hawaii not covered), and 15 (source text cites the station
Normals).
## Decisions
In [decisions.md](../decisions.md), all 2026-09-15: Annual avg temperature
month weighting, aggregation, completeness, and in Alaska and Hawaii.
## Tasks
Adopted definition: the equally weighted mean of the 12 monthly 1991–2020
normals of nClimGrid-Monthly `tavg`, area-weighted to each county; blank
unless all 12 months are present. The per-cell normals follow NCEI's own
gridded-normals method, "a simple 30-year average of monthly grids"
(Rennie and Palecki, *U.S. Monthly Gridded Precipitation and Temperature
Climate Normals*). Alaska and Hawaii stay blank; see
[decisions.md](../decisions.md).
- [x] Verify the calculation against the source data and the WMO and NOAA
normals definitions (2026-09-15).
- [x] Decide month weighting, aggregation, completeness, and Alaska/Hawaii
handling (2026-09-15; see [decisions.md](../decisions.md)).
- [x] Rewrite `filter-calculations.md` §2 with the adopted definition, citing
WMO-No. 1203 §4.3.3 and NOAA's 1991–2020 Normals methodology. State that
average temperature is (Tmax + Tmin)/2 and that the grid covers the
contiguous U.S. only (2026-09-15; the current touched-cell method is kept
as "Current method" until the new values are applied).
- [ ] Add a continuous-value area-weighted mean to
`scripts/common/county_zonal_stats.py`. It takes an array and transform,
because nClimGrid is read from netCDF rather than a rasterio file, and
reuses `cell_coverage_fractions`, the cos(latitude) scaling, and the
padded window. NaN cells are excluded.
- [ ] Build `scripts/build_county_avg_temp_metric.py`, writing
`data/metrics/avg_temp.csv` with `avgTempF` and a valid-month count. The
1991–2020 climatology step is NOAA-specific, so it stays with the NOAA
code rather than `common/`.
- [ ] Single-column apply step
`scripts/apply_avg_temp_metric_to_climate_data.py` with `--dry-run`.
Move `read_csv_rows` (4 identical copies in `apply_*` scripts) into
`common/` as part of this step.
- [ ] `check_climate_data.py`: keep the 20–85 °F rule and the Alaska/Hawaii
blank allowance until that open decision in
[decisions.md](../decisions.md) is made; confirm no new blanks in the
contiguous U.S.
- [ ] Tests (`tests/test_avg_temp_metric.py`): partial-cell weights,
cos(latitude), NaN exclusion, the 12-month rule, equal month weights,
Celsius-to-Fahrenheit conversion and rounding, and the apply step.
- [ ] Apply to `data/climate-data.csv`, run the checker, and compare: the
spread of changes, the largest shifts (expected in small, narrow, and
coastal counties), and no new blanks among the 3,109 contiguous-U.S.
counties.
- [ ] Mark finding 2 resolved for this metric in
[00-cross-filter.md](00-cross-filter.md).
- [ ] Correct the `avgTempF` source text in `app.js`, which cites the
station-based U.S. Climate Normals; describe it as 1991–2020 normals
computed from nClimGrid-Monthly and link nClimGrid. The "(Normals)"
label stays.
- [ ] Correct `scripts/county_data_sources.md`: Source 2 says the build reads
monthly normals files, but it reads the nClimGrid monthly series and
averages 1991–2020. Update the `avgTempF` definition line to match.
- [ ] Add the area-weighted `avgTempF` to the §7 guardrail on rerunning
`build_county_climate_data.py` in
[pipeline-plan.md](../pipeline-plan.md).
@@ -0,0 +1,23 @@
# Filter 3: Diurnal temperature range
**Status:** Not started.
**Data key:** `avgDiurnalTempRangeF`. Calculation:
[filter-calculations.md](../filter-calculations.md) §3.
## Findings
No filter-specific findings yet.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered),
13 (Lexington, VA blank), and 17 (missing from the data-sources metric list).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
@@ -0,0 +1,40 @@
# Filter 4: Extreme temperature days
**Status:** Not started.
**Data key:** `absoluteExtremeDays`. Calculation:
[filter-calculations.md](../filter-calculations.md) §4.
## Findings
**5. Extreme-day counts are not completeness-normalized.** A partially observed
year contributes a raw count and receives the same weight as a complete year.
Consider requiring a minimum number of valid days or annualizing partial
counts explicitly.
**6. The absolute-extreme metric depends on unrelated percentile thresholds.**
`build_annual_counts` skips a county when its retired local p95/p05 thresholds
are missing, even though the active 95 °F / 0 °F calculation does not require
those percentiles. The absolute calculation should be separated from that
prerequisite.
**18. The data-sources doc describes the retired locally extreme metric.**
Source 5 in `scripts/county_data_sources.md` is headed `locallyExtremeDays`
and says `apply_locally_extreme_metric_to_climate_data.py` writes the locally
extreme average into `extremeDays`, keeps `oldExtremeDays`, and adds eight
detail and audit columns. None of those columns is in `data/climate-data.csv`,
and the script's docstring says locally percentile-based metrics are not
included in the app CSV.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered),
13 (Lexington, VA blank), and 17 (missing from the data-sources metric list).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
+30
View File
@@ -0,0 +1,30 @@
# Filter 5: 90 °F+ heat-index days
**Status:** Not started.
**Data keys:** `humidHeatDays`, with the audit columns `humidHeatSourceFips`
and `humidHeatFipsAdjustment`. Calculation:
[filter-calculations.md](../filter-calculations.md) §5.
## Findings
**7. Heat Index days are a daily-extrema proxy.** Daily Tmax and daily minimum
relative humidity are paired even though their observation times may differ.
The result should not be described as an observed hourly maximum Heat Index.
**8. Heat-year completeness is permissive.** Any year with at least one valid
Tmax/RH pair is included in the equal-year average. A minimum valid-day rule
would reduce low-biased partial-year counts.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered)
and 13 (Lexington, VA uses Rockbridge County as a proxy).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
+27
View File
@@ -0,0 +1,27 @@
# Filter 6: Annual precipitation
**Status:** Not started.
**Data key:** `annualPrecipIn`. Calculation:
[filter-calculations.md](../filter-calculations.md) §6.
## Findings
**3. Partial precipitation years are accepted.** One valid monthly precipitation
value is sufficient to produce `annualPrecipIn`; absent months silently lower
the annual sum. Requiring all 12 months, or recording completeness, would be
safer.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 2 and 11 (county aggregation),
12 (Alaska and Hawaii not covered), and 15 (source text cites the station
Normals).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
+24
View File
@@ -0,0 +1,24 @@
# Filter 7: Seasonality index
**Status:** Not started.
**Data key:** `seasonalityIndex`. Calculation:
[filter-calculations.md](../filter-calculations.md) §7.
## Findings
No filter-specific findings yet.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 2 and 11 (county aggregation),
12 (Alaska and Hawaii not covered), and 15 (source text cites the station
Normals).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
+24
View File
@@ -0,0 +1,24 @@
# Filter 8: Wettest month
**Status:** Not started.
**Data key:** `wettestPrecipMonth`. Calculation:
[filter-calculations.md](../filter-calculations.md) §8.
## Findings
No numbered findings yet. Known issue to review: computed in both the base
build and the precipitation-month script.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered)
and 16 (the app's nClimGrid source link returns 404).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
+25
View File
@@ -0,0 +1,25 @@
# Filter 9: Driest month
**Status:** Not started.
**Data key:** `driestPrecipMonth`. Calculation:
[filter-calculations.md](../filter-calculations.md) §9.
## Findings
No numbered findings yet. Known issue to review: same as wettest month, which
is computed in both the base build and the precipitation-month script; see
[08-wettest-month.md](08-wettest-month.md).
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered)
and 16 (the app's nClimGrid source link returns 404).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
@@ -0,0 +1,24 @@
# Filter 10: Summer specific humidity
**Status:** Not started.
**Data key:** `avgSummerSpecificHumidityGKg`. Calculation:
[filter-calculations.md](../filter-calculations.md) §10.
## Findings
No filter-specific findings yet. Known issue to review: cell-center
cos(latitude) aggregation differs from other metrics (finding 11).
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 11 (inconsistent aggregation) and
12 (Alaska and Hawaii not covered).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
+41
View File
@@ -0,0 +1,41 @@
# Filter 11: Solar GHI
**Status:** Not started.
**Data key:** `meanDailyGlobalHorizontalRadiationKwhM2Day`. Calculation:
[filter-calculations.md](../filter-calculations.md) §11.
## Findings
**9. The GHI formula assumes hourly, 365-day input.** It is correct for the
current 60-minute, `leap_day=false` requests. If the request interval changes,
the energy sum needs an interval-hours multiplier; leap-day handling would
also need to change the divisor.
**19. The data-sources doc recommends a GHI raster the pipeline does not use.**
Source 4 in `scripts/county_data_sources.md` says county means should come
from a gridded annual GHI raster passed with `--solar-ghi-raster`. The app
values come from NSRDB polygon archive summaries, with representative points
as fallback, applied by `apply_locally_extreme_metric_to_climate_data.py`.
**20. The base build labels any solar CSV as representative-point.** In
`build_county_climate_data.py`, the `--solar-ghi-csv` help text calls it a
representative-point fallback, and the per-row `source` tag is always
`solar-ghi-representative-point`, but the documented build command passes
`data/nrel/county_polygon_ghi_summary.csv`. The apply step later replaces both
the value and the tag, so only the base build's output is mislabeled.
Known issue to review: finalized by the extreme-temperature apply script.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 11 (inconsistent aggregation) and
16 (the app's NSRDB source link no longer resolves).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.
@@ -0,0 +1,26 @@
# Filter 12: Clear-sky GHI reduction
**Status:** Not started.
**Data key:** `clearSkyGhiReductionIndex`. Calculation:
[filter-calculations.md](../filter-calculations.md) §12.
## Findings
**10. Clear-sky reduction averages ratios rather than energy totals.** This is
a valid but specific definition. It gives each retained time row equal weight,
rather than weighting rows by available clear-sky energy. The label and
documentation should retain this distinction.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 11 (inconsistent aggregation) and
16 (the app's NSRDB source link no longer resolves).
## Decisions
None yet.
## Tasks
Not started; follow the per-filter review checklist in
[pipeline-plan.md](../pipeline-plan.md) §5.