Document restructuring and the beginnings of Filter 2 changes

Split the pipeline documentation by purpose so each fact has one home:
- docs/pipeline-plan.md keeps the plan, checklist, tracker, and guardrails
- docs/decisions.md holds open decisions and the dated decision log
- docs/reviews/ holds findings and tasks: one file per filter, plus
  00-cross-filter.md for findings that span filters
- scripts/common/README.md holds the shared-helper rules (formerly Phase 2)
- filter-calculations.md now describes calculations only

Filed findings 12-22 from a consistency audit of the app, docs, and scripts.

Filter 1 (Köppen-Geiger): use "Köppen" with the umlaut in all prose, labels,
docstrings, help text, and checker messages (finding 21), and correct the
base build's "majority" docstring (finding 22).

Filter 2 (annual avg temperature): record the adopted definition in
filter-calculations.md §2: equally weighted 1991-2020 monthly normals, per
WMO-No. 1203 and NOAA's 2020 methodology; area-weighted county means; blank
unless all 12 months exist. Code changes for this filter are still pending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-15 16:13:11 -04:00
co-authored by Claude Opus 5
parent a9e722791d
commit e855d583e3
27 changed files with 807 additions and 242 deletions
+30 -122
View File
@@ -3,7 +3,10 @@
This plan describes how the county data pipeline will move from scripts that
edit one shared CSV in place to per-metric outputs assembled into the app CSV.
It is a working reference for the filter-by-filter review. Calculation details
for each filter live in [filter-calculations.md](filter-calculations.md).
for each filter live in [filter-calculations.md](filter-calculations.md);
findings and tasks live in [reviews/](reviews/), one file per filter plus
[00-cross-filter.md](reviews/00-cross-filter.md); open and past decisions live
in [decisions.md](decisions.md).
**Started:** 2026-09-12
@@ -64,10 +67,8 @@ The README's enrichment list starts at step 2 and omits steps 1 and 3.
`scripts/county_data_sources.md`, and in each script's assumptions.
3. **Column ownership is unclear.** GHI is finalized by the extreme-temperature
apply script; wettest/driest month are computed in two places.
4. **County aggregation is inconsistent.** NOAA uses touched raster cells,
Köppen uses area-weighted shares (since 2026-09-13), gridMET uses cell
centers with cos(latitude) weights, and NSRDB uses overlap areas (finding 11
in `filter-calculations.md`).
4. **County aggregation is inconsistent.** See finding 11 in
[reviews/00-cross-filter.md](reviews/00-cross-filter.md).
5. **No single entry point or final check.** A new user must piece together
about 20 scripts, several large downloads, and an NSRDB API key.
@@ -125,50 +126,10 @@ existing CSV keeps working until Phase 3.
### Phase 2 — Shared helpers in `scripts/common/`
`scripts/common/` holds code used by more than one data source (NOAA, gridMET,
NSRDB, Köppen). Scripts import from it, for example
`from common.counties import load_counties`; nothing in it is run directly.
Rules for `common/`:
- **Cross-source only.** A helper goes in only if metrics from more than one
data source use it. Code shared by scripts of a single data source stays with
that source, for example a future NSRDB module for the NSRDB prompt,
redaction, and error-log helpers.
- **One topic per module.** Each module is named for its topic and has a
docstring. No catch-all `utils.py`.
- **Keep it small.** Before adding a helper, ask why it does not belong to any
one data source.
Current modules:
| Module | Contents | Why it is in `common/` |
| --- | --- | --- |
| `county_zonal_stats.py` | Raster windows, the 180th-meridian split, area-weighted class shares | Used by any raster-based metric |
| `counties.py` | County polygon loading, FIPS normalization, the state FIPS table | County identity is shared by nearly every pipeline |
| `koppen_legend.py` | The Köppen code map and legend loader | Temporary: also used by `build_county_climate_data.py`; moves next to the Köppen code in Phase 3 |
**Rationale.** Helper functions make each step of a computation explicit,
avoid repeated code, and can be tested separately
([Brown CSCI 0111, "Helper Functions"](https://cs.brown.edu/courses/csci0111/fall2018/lectures/helper-functions.html)).
Shared helper folders, however, tend to lose cohesion and collect unrelated
code; the recommended alternative is to keep code with the part of the system
it belongs to, allowing a shared folder only if it stays small and documented
([Helpers and Utils Folders in Software Architecture](https://dev.to/knzt/helpers-and-utils-folders-in-software-architecture-3f8h)).
The rules above follow both: shared functions, organized by topic and limited
to code that crosses data sources.
Other shared code moves when its filter is reviewed, so each move is tested
alongside that filter. A 2026-09-13 survey found 19 functions with identical
copies in several scripts and 19 with copies that have drifted apart. Most
identical copies are NSRDB helpers, which belong in an NSRDB module rather
than `common/`; `read_csv_rows` (4 identical copies in `apply_*` scripts) is
cross-source. Drifted copies need a decision on which version is correct
before merging. Notable drifts: `summarize_county_gridmet_humidity.py` has its
own county loader and FIPS normalizer, and the state FIPS table is also copied
in `build_county_representative_points.py`,
`summarize_county_gridmet_humidity.py`, and
`request_nsrdb_county_polygon_archives.py`.
Rules, current modules, and rationale are in
[scripts/common/README.md](../scripts/common/README.md). Duplicated helpers
found by the 2026-09-13 survey are finding 14 in
[reviews/00-cross-filter.md](reviews/00-cross-filter.md).
### Phase 3 — Assemble and restructure (after all 12 filters are clean)
@@ -195,8 +156,10 @@ in `build_county_representative_points.py`,
For each filter:
1. Verify the calculation against the source data and document findings.
2. Decide any rule or method changes with the project owner.
1. Verify the calculation against the source data and document findings in the
filter's review file under `docs/reviews/`.
2. Decide any rule or method changes with the project owner, and record them in
`decisions.md`.
3. Record the adopted definition in `filter-calculations.md`.
4. Implement the calculation, writing `data/metrics/<metric>.csv`.
5. Add a single-column apply step for the current CSV.
@@ -208,53 +171,24 @@ For each filter:
## 6. Filter tracker
Known issues come from `filter-calculations.md` ("Calculation review findings")
and this review; none beyond Köppen have been investigated yet.
Each filter's findings, decisions, and tasks are in its review file. Findings
that affect more than one filter are in
[00-cross-filter.md](reviews/00-cross-filter.md).
| # | Filter | Status | Known issues to review |
| # | Filter | Status | Review |
| --- | --- | --- | --- |
| 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | See tasks below |
| 2 | Annual avg temperature | Not started | Months weighted equally (finding 1); touched-cell aggregation (finding 2) |
| 3 | Diurnal temperature range | Not started | Lexington, VA (51678) blank, while heat-index days use Rockbridge County as a proxy |
| 4 | Extreme temperature days | Not started | Partial years not normalized (finding 5); depends on retired percentile thresholds (finding 6); Lexington, VA blank |
| 5 | 90 °F+ heat-index days | Not started | Daily-extrema proxy (finding 7); permissive year completeness (finding 8) |
| 6 | Annual precipitation | Not started | Partial-year sums accepted (finding 3); touched-cell aggregation (finding 2) |
| 7 | Seasonality index | Not started | Touched-cell aggregation (finding 2) |
| 8 | Wettest month | Not started | Computed in both the base build and the precipitation-month script |
| 9 | Driest month | Not started | Same as wettest month |
| 10 | Summer specific humidity | Not started | Cell-center cos(latitude) aggregation differs from other metrics (finding 11) |
| 11 | Solar GHI | Not started | Finalized by the extreme-temperature apply script; hourly, 365-day assumption (finding 9) |
| 12 | Clear-sky GHI reduction | Not started | Mean of ratios rather than energy totals (finding 10) |
### Köppen-Geiger tasks
Adopted rule: a county is predominantly its top class if and only if that class
covers at least 50% of the county's land and leads the runner-up by at least
5 percentage points; otherwise it is Mixed climate. Expected result for the
50 states and DC: 3,010 predominant, 133 Mixed.
- [x] Investigate low-majority counties and adopt the rule.
- [x] Windowed raster reads and 180th-meridian split.
- [x] Area-weighted class shares (16 × 16 sub-cells per raster cell, scaled by
cos(latitude)) in `scripts/common/county_zonal_stats.py`.
- [x] Apply the 50% / 5-point rule (`scripts/build_county_koppen_metric.py`;
counties with no valid cells are left blank).
- [x] Run the builder to write `data/metrics/koppen.csv` and confirm the
expected 3,010 predominant / 133 Mixed (2026-09-13).
- [x] `koppenZone`-only apply step
(`scripts/apply_koppen_metric_to_climate_data.py`, with `--dry-run`).
- [x] Apply to `data/climate-data.csv` (2026-09-13; 142 counties changed to
Mixed, with `koppenPrimaryClass` and `koppenSecondaryClass` added).
- [x] Allow `Mixed` in `check_climate_data.py`.
- [x] Add a Mixed climate category to `app.js`, drawn as stripes of the
county's top two classes; see
[koppen-mixed-display-plan.md](koppen-mixed-display-plan.md).
- [x] Replace the plurality description in `filter-calculations.md` §1 and mark
review findings 2 and 4 resolved for Köppen (2026-09-14).
- [x] Update the Köppen descriptions and script lists in `README.md` and
`scripts/county_data_sources.md` (2026-09-14).
- [x] Tests for shares, the rule, boundary cases, and the apply step
(`tests/test_koppen_metric.py`).
| 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | [01-koppen.md](reviews/01-koppen.md) |
| 2 | Annual avg temperature | In progress: method decided (2026-09-15) | [02-annual-avg-temperature.md](reviews/02-annual-avg-temperature.md) |
| 3 | Diurnal temperature range | Not started | [03-diurnal-temperature-range.md](reviews/03-diurnal-temperature-range.md) |
| 4 | Extreme temperature days | Not started | [04-extreme-temperature-days.md](reviews/04-extreme-temperature-days.md) |
| 5 | 90 °F+ heat-index days | Not started | [05-heat-index-days.md](reviews/05-heat-index-days.md) |
| 6 | Annual precipitation | Not started | [06-annual-precipitation.md](reviews/06-annual-precipitation.md) |
| 7 | Seasonality index | Not started | [07-seasonality-index.md](reviews/07-seasonality-index.md) |
| 8 | Wettest month | Not started | [08-wettest-month.md](reviews/08-wettest-month.md) |
| 9 | Driest month | Not started | [09-driest-month.md](reviews/09-driest-month.md) |
| 10 | Summer specific humidity | Not started | [10-summer-specific-humidity.md](reviews/10-summer-specific-humidity.md) |
| 11 | Solar GHI | Not started | [11-solar-ghi.md](reviews/11-solar-ghi.md) |
| 12 | Clear-sky GHI reduction | Not started | [12-clear-sky-ghi-reduction.md](reviews/12-clear-sky-ghi-reduction.md) |
## 7. Guardrails until Phase 3
@@ -263,29 +197,3 @@ covers at least 50% of the county's land and leads the runner-up by at least
`extremeDays`, and overwrite the Mixed classification in `koppenZone`.
- Run `check_climate_data.py` after every apply step.
- Change one filter at a time, and compare its before and after values.
## 8. Open decisions
| Decision | Options | Needed by |
| --- | --- | --- |
| Committing large intermediates | Commit metric files only, or also source summaries | Phase 3 |
### Decided
- **Köppen audit columns (2026-09-12):** `koppen.csv` stores the top class and
share and the runner-up class and share alongside `koppenZone`.
- **Köppen no-data fallback (2026-09-12):** a county with no valid raster cells
is left blank, not assigned `Cfa`.
- **Applying Köppen to the app CSV (2026-09-12):** wait until the app supports
the Mixed class. Done 2026-09-13.
- **Metric files (2026-09-13):** one CSV per metric under `data/metrics/`,
starting with `koppen.csv`.
- **Mixed climate display (2026-09-13):** diagonal stripes of each Mixed
county's top two classes; see
[koppen-mixed-display-plan.md](koppen-mixed-display-plan.md).
- **Puerto Rico (2026-09-13):** off the map and out of every filter. The app
already drops state FIPS 72; the data files keep the rows.
- **Shared helpers (2026-09-13):** `scripts/common/` holds only code used by
more than one data source, one topic per module; code shared within one data
source stays with that source. Duplicates move during their own filter's
review.