Split the pipeline documentation by purpose so each fact has one home: - docs/pipeline-plan.md keeps the plan, checklist, tracker, and guardrails - docs/decisions.md holds open decisions and the dated decision log - docs/reviews/ holds findings and tasks: one file per filter, plus 00-cross-filter.md for findings that span filters - scripts/common/README.md holds the shared-helper rules (formerly Phase 2) - filter-calculations.md now describes calculations only Filed findings 12-22 from a consistency audit of the app, docs, and scripts. Filter 1 (Köppen-Geiger): use "Köppen" with the umlaut in all prose, labels, docstrings, help text, and checker messages (finding 21), and correct the base build's "majority" docstring (finding 22). Filter 2 (annual avg temperature): record the adopted definition in filter-calculations.md §2: equally weighted 1991-2020 monthly normals, per WMO-No. 1203 and NOAA's 2020 methodology; area-weighted county means; blank unless all 12 months exist. Code changes for this filter are still pending. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
79 lines
4.0 KiB
Markdown
79 lines
4.0 KiB
Markdown
# Filter 1: Köppen-Geiger class
|
||
|
||
**Status:** Done (2026-09-14): rule applied, Mixed display built, documentation
|
||
updated.
|
||
|
||
**Data keys:** `koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`.
|
||
Calculation: [filter-calculations.md](../filter-calculations.md) §1. Display
|
||
design and history:
|
||
[koppen-mixed-display-plan.md](../koppen-mixed-display-plan.md).
|
||
|
||
## Findings
|
||
|
||
**4. The Köppen fallback can create false data.** A county with no valid raster
|
||
cells is labeled `Cfa` instead of missing. A null value plus an audit flag
|
||
would distinguish missing coverage from a genuine humid-subtropical class.
|
||
*Resolved on 2026-09-13: the Köppen builder leaves such counties blank. The
|
||
fallback remains only in `build_county_climate_data.py`, whose Köppen
|
||
value is replaced by the apply step.*
|
||
|
||
**21. The documented Köppen labels differ from the app.** `filter-calculations.md`
|
||
lists the group as "Climate Classification" and the filter as "Köppen-Geiger
|
||
Climate Class"; `app.js` shows "Koppen-Geiger Classification" and
|
||
"Koppen-Geiger Climate Class". The app spells Koppen without the umlaut
|
||
throughout, while the docs use Köppen. *Resolved on 2026-09-15: every
|
||
prose and display use of the name now reads "Köppen", in `app.js`,
|
||
`README.md`, `scripts/county_data_sources.md`, script docstrings, help text,
|
||
and checker messages (with the matching test expectations), and
|
||
`filter-calculations.md` quotes the app's group label, "Köppen-Geiger
|
||
Classification". Code identifiers, file names, and data paths such as
|
||
`koppenZone`, `normalizeKoppenCode`, and `koppen_geiger_tif/` keep the ASCII
|
||
spelling.*
|
||
|
||
**22. The base build's docstring calls its Köppen value a majority class.**
|
||
`build_county_climate_data.py` lists `koppenZone` as the "majority
|
||
Koppen-Geiger class", but it assigns the most common class among touched
|
||
cells, which need not be a majority. That value is replaced by the Köppen apply
|
||
step, and the script is retired in Phase 3. *Resolved on 2026-09-15: the
|
||
docstring now reads "largest-share Köppen-Geiger class among touched cells
|
||
(replaced by the Köppen apply step)".*
|
||
|
||
Cross-filter finding 2 (touched-cell aggregation) was resolved for Köppen on
|
||
2026-09-13; see [00-cross-filter.md](00-cross-filter.md).
|
||
|
||
## Decisions
|
||
|
||
In [decisions.md](../decisions.md): Köppen audit columns, Köppen no-data
|
||
fallback, and Applying Köppen to the app CSV (2026-09-12); Metric files and
|
||
Mixed climate display (2026-09-13).
|
||
|
||
## Tasks
|
||
|
||
Adopted rule: a county is predominantly its top class if and only if that class
|
||
covers at least 50% of the county's land and leads the runner-up by at least
|
||
5 percentage points; otherwise it is Mixed climate. Expected result for the
|
||
50 states and DC: 3,010 predominant, 133 Mixed.
|
||
|
||
- [x] Investigate low-majority counties and adopt the rule.
|
||
- [x] Windowed raster reads and 180th-meridian split.
|
||
- [x] Area-weighted class shares (16 × 16 sub-cells per raster cell, scaled by
|
||
cos(latitude)) in `scripts/common/county_zonal_stats.py`.
|
||
- [x] Apply the 50% / 5-point rule (`scripts/build_county_koppen_metric.py`;
|
||
counties with no valid cells are left blank).
|
||
- [x] Run the builder to write `data/metrics/koppen.csv` and confirm the
|
||
expected 3,010 predominant / 133 Mixed (2026-09-13).
|
||
- [x] `koppenZone`-only apply step
|
||
(`scripts/apply_koppen_metric_to_climate_data.py`, with `--dry-run`).
|
||
- [x] Apply to `data/climate-data.csv` (2026-09-13; 142 counties changed to
|
||
Mixed, with `koppenPrimaryClass` and `koppenSecondaryClass` added).
|
||
- [x] Allow `Mixed` in `check_climate_data.py`.
|
||
- [x] Add a Mixed climate category to `app.js`, drawn as stripes of the
|
||
county's top two classes; see
|
||
[koppen-mixed-display-plan.md](../koppen-mixed-display-plan.md).
|
||
- [x] Replace the plurality description in `filter-calculations.md` §1 and mark
|
||
review findings 2 and 4 resolved for Köppen (2026-09-14).
|
||
- [x] Update the Köppen descriptions and script lists in `README.md` and
|
||
`scripts/county_data_sources.md` (2026-09-14).
|
||
- [x] Tests for shares, the rule, boundary cases, and the apply step
|
||
(`tests/test_koppen_metric.py`).
|