Document restructuring and the beginnings of Filter 2 changes

Split the pipeline documentation by purpose so each fact has one home:
- docs/pipeline-plan.md keeps the plan, checklist, tracker, and guardrails
- docs/decisions.md holds open decisions and the dated decision log
- docs/reviews/ holds findings and tasks: one file per filter, plus
  00-cross-filter.md for findings that span filters
- scripts/common/README.md holds the shared-helper rules (formerly Phase 2)
- filter-calculations.md now describes calculations only

Filed findings 12-22 from a consistency audit of the app, docs, and scripts.

Filter 1 (Köppen-Geiger): use "Köppen" with the umlaut in all prose, labels,
docstrings, help text, and checker messages (finding 21), and correct the
base build's "majority" docstring (finding 22).

Filter 2 (annual avg temperature): record the adopted definition in
filter-calculations.md §2: equally weighted 1991-2020 monthly normals, per
WMO-No. 1203 and NOAA's 2020 methodology; area-weighted county means; blank
unless all 12 months exist. Code changes for this filter are still pending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-15 16:13:11 -04:00
co-authored by Claude Opus 5
parent a9e722791d
commit e855d583e3
27 changed files with 807 additions and 242 deletions
+39
View File
@@ -0,0 +1,39 @@
# `scripts/common/`
`scripts/common/` holds code used by more than one data source (NOAA, gridMET,
NSRDB, Köppen). Scripts import from it, for example
`from common.counties import load_counties`; nothing in it is run directly.
Rules for `common/`:
- **Cross-source only.** A helper goes in only if metrics from more than one
data source use it. Code shared by scripts of a single data source stays with
that source, for example a future NSRDB module for the NSRDB prompt,
redaction, and error-log helpers.
- **One topic per module.** Each module is named for its topic and has a
docstring. No catch-all `utils.py`.
- **Keep it small.** Before adding a helper, ask why it does not belong to any
one data source.
Other shared code moves when its filter is reviewed, so each move is tested
alongside that filter. Duplicated and drifted helpers found by the 2026-09-13
survey are finding 14 in
[docs/reviews/00-cross-filter.md](../../docs/reviews/00-cross-filter.md).
Current modules:
| Module | Contents | Why it is in `common/` |
| --- | --- | --- |
| `county_zonal_stats.py` | Raster windows, the 180th-meridian split, area-weighted class shares | Used by any raster-based metric |
| `counties.py` | County polygon loading, FIPS normalization, the state FIPS table | County identity is shared by nearly every pipeline |
| `koppen_legend.py` | The Köppen code map and legend loader | Temporary: also used by `build_county_climate_data.py`; moves next to the Köppen code in Phase 3 |
**Rationale.** Helper functions make each step of a computation explicit,
avoid repeated code, and can be tested separately
([Brown CSCI 0111, "Helper Functions"](https://cs.brown.edu/courses/csci0111/fall2018/lectures/helper-functions.html)).
Shared helper folders, however, tend to lose cohesion and collect unrelated
code; the recommended alternative is to keep code with the part of the system
it belongs to, allowing a shared folder only if it stays small and documented
([Helpers and Utils Folders in Software Architecture](https://dev.to/knzt/helpers-and-utils-folders-in-software-architecture-3f8h)).
The rules above follow both: shared functions, organized by topic and limited
to code that crosses data sources.
+2 -2
View File
@@ -1,4 +1,4 @@
"""Koppen-Geiger raster codes and the Beck et al. legend loader."""
"""Köppen-Geiger raster codes and the Beck et al. legend loader."""
from __future__ import annotations
@@ -42,7 +42,7 @@ DEFAULT_KOPPEN_CODE_MAP = {
def load_koppen_legend(legend_path: Path | None) -> Dict[int, str]:
"""Load Koppen raster codes, using defaults when no legend exists."""
"""Load Köppen raster codes, using defaults when no legend exists."""
if legend_path is None:
return DEFAULT_KOPPEN_CODE_MAP