Document restructuring and the beginnings of Filter 2 changes

Split the pipeline documentation by purpose so each fact has one home:
- docs/pipeline-plan.md keeps the plan, checklist, tracker, and guardrails
- docs/decisions.md holds open decisions and the dated decision log
- docs/reviews/ holds findings and tasks: one file per filter, plus
  00-cross-filter.md for findings that span filters
- scripts/common/README.md holds the shared-helper rules (formerly Phase 2)
- filter-calculations.md now describes calculations only

Filed findings 12-22 from a consistency audit of the app, docs, and scripts.

Filter 1 (Köppen-Geiger): use "Köppen" with the umlaut in all prose, labels,
docstrings, help text, and checker messages (finding 21), and correct the
base build's "majority" docstring (finding 22).

Filter 2 (annual avg temperature): record the adopted definition in
filter-calculations.md §2: equally weighted 1991-2020 monthly normals, per
WMO-No. 1203 and NOAA's 2020 methodology; area-weighted county means; blank
unless all 12 months exist. Code changes for this filter are still pending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-15 16:13:11 -04:00
co-authored by Claude Opus 5
parent a9e722791d
commit e855d583e3
27 changed files with 807 additions and 242 deletions
@@ -1,4 +1,4 @@
"""Apply the county Koppen-Geiger metric to the app CSV.
"""Apply the county Köppen-Geiger metric to the app CSV.
Replaces the koppenZone column of data/climate-data.csv with the values in
data/metrics/koppen.csv and writes koppenPrimaryClass and koppenSecondaryClass:
@@ -71,7 +71,7 @@ def load_koppen_values(koppen_metric: Path) -> Dict[str, Dict[str, str]]:
def apply_koppen_metric(
climate_data: Path, koppen_metric: Path, out: Path, dry_run: bool = False
) -> Tuple[List[Change], List[str]]:
"""Update the Koppen columns; return the value changes and any columns that were added."""
"""Update the Köppen columns; return the value changes and any columns that were added."""
fields, rows = read_csv_rows(climate_data)
if ZONE_FIELD not in fields:
raise ValueError(f"{climate_data} has no {ZONE_FIELD} column.")
@@ -101,9 +101,9 @@ def apply_koppen_metric(
def parse_args() -> argparse.Namespace:
"""Define and parse command-line options for this apply step."""
parser = argparse.ArgumentParser(description="Apply the county Koppen metric to the app CSV.")
parser = argparse.ArgumentParser(description="Apply the county Köppen metric to the app CSV.")
parser.add_argument("--climate-data", type=Path, default=DEFAULT_CLIMATE_DATA, help="App climate CSV to update.")
parser.add_argument("--koppen-metric", type=Path, default=DEFAULT_KOPPEN_METRIC, help="Koppen metric CSV.")
parser.add_argument("--koppen-metric", type=Path, default=DEFAULT_KOPPEN_METRIC, help="Köppen metric CSV.")
parser.add_argument("--out", type=Path, default=None, help="Output path; defaults to updating --climate-data in place.")
parser.add_argument("--dry-run", action="store_true", help="Report changes without writing.")
return parser.parse_args()
+3 -3
View File
@@ -6,7 +6,7 @@ Outputs a CSV file compatible with the browser app:
climate-data.csv
Metrics produced per county:
- koppenZone: majority Koppen-Geiger class
- koppenZone: largest-share Köppen-Geiger class among touched cells (replaced by the Köppen apply step)
- avgTempF: annual mean temperature from NOAA 1991-2020 gridded normals
- annualPrecipIn: annual total precipitation from NOAA 1991-2020 gridded normals
- seasonalityIndex: precipitation seasonality, coefficient of variation of monthly totals (%)
@@ -256,7 +256,7 @@ def _touched_raster_values(source, geometry: BaseGeometry, split_antimeridian: b
def _zonal_majority_class(koppen_raster: Path, counties: gpd.GeoDataFrame, code_map: Dict[int, str]) -> List[str]:
"""Assign each county its most common Koppen-Geiger class."""
"""Assign each county its most common Köppen-Geiger class."""
classes: List[str] = []
with rasterio.open(koppen_raster) as source:
raster_counties = counties
@@ -597,7 +597,7 @@ def parse_args() -> argparse.Namespace:
"""Define and parse command-line options for this generator."""
parser = argparse.ArgumentParser(description="Generate county climate records for the web app.")
parser.add_argument("--counties-geojson", type=Path, required=True, help="County polygon GeoJSON path.")
parser.add_argument("--koppen-raster", type=Path, required=True, help="Koppen-Geiger raster TIFF path.")
parser.add_argument("--koppen-raster", type=Path, required=True, help="Köppen-Geiger raster TIFF path.")
parser.add_argument("--koppen-legend", type=Path, default=None, help="Optional legend.txt mapping integer codes.")
parser.add_argument(
"--monthly-tavg-nc",
+5 -5
View File
@@ -1,4 +1,4 @@
"""Build the county Koppen-Geiger metric file from area-weighted class shares.
"""Build the county Köppen-Geiger metric file from area-weighted class shares.
Writes data/metrics/koppen.csv. A county is predominantly its top class when
that class covers at least 50% of the county's land and leads the runner-up by
@@ -58,7 +58,7 @@ def rank_class_shares(weights: Dict[int, float], code_map: Dict[int, str]) -> Li
return []
unknown = sorted(code for code in weights if code not in code_map)
if unknown:
raise ValueError(f"Raster codes {unknown} are not in the Koppen legend.")
raise ValueError(f"Raster codes {unknown} are not in the Köppen legend.")
ranked = sorted(weights.items(), key=lambda item: (-item[1], item[0]))
return [(code_map[code], weight / total) for code, weight in ranked]
@@ -117,7 +117,7 @@ def build_koppen_records(
def write_records(records: List[dict], out_file: Path) -> None:
"""Write the Koppen metric rows."""
"""Write the Köppen metric rows."""
out_file.parent.mkdir(parents=True, exist_ok=True)
with out_file.open("w", encoding="utf-8", newline="") as csv_file:
writer = csv.DictWriter(csv_file, fieldnames=FIELDS)
@@ -127,9 +127,9 @@ def write_records(records: List[dict], out_file: Path) -> None:
def parse_args() -> argparse.Namespace:
"""Define and parse command-line options for this builder."""
parser = argparse.ArgumentParser(description="Build the county Koppen-Geiger metric file.")
parser = argparse.ArgumentParser(description="Build the county Köppen-Geiger metric file.")
parser.add_argument("--counties-geojson", type=Path, default=DEFAULT_COUNTIES_GEOJSON, help="County polygon GeoJSON path.")
parser.add_argument("--koppen-raster", type=Path, default=DEFAULT_KOPPEN_RASTER, help="Koppen-Geiger raster TIFF path.")
parser.add_argument("--koppen-raster", type=Path, default=DEFAULT_KOPPEN_RASTER, help="Köppen-Geiger raster TIFF path.")
parser.add_argument("--koppen-legend", type=Path, default=DEFAULT_KOPPEN_LEGEND, help="legend.txt mapping raster codes.")
parser.add_argument("--subcells", type=int, default=DEFAULT_SUBCELLS, help="Sub-cells per raster cell edge.")
parser.add_argument("--out", type=Path, default=DEFAULT_OUT, help="Output metric CSV path.")
+5 -5
View File
@@ -110,7 +110,7 @@ METRIC_RULES: Dict[str, MetricRule] = {
}
IDENTITY_COLUMNS = ("countyFips", "countyName", "state")
# Stripe classes the app draws for Mixed Koppen counties; blank otherwise.
# Stripe classes the app draws for Mixed Köppen counties; blank otherwise.
KOPPEN_STRIPE_COLUMNS = ("koppenPrimaryClass", "koppenSecondaryClass")
AUDIT_COLUMNS = ("humidHeatSourceFips", "humidHeatFipsAdjustment", "source")
EXPECTED_COLUMNS = IDENTITY_COLUMNS + tuple(METRIC_RULES) + KOPPEN_STRIPE_COLUMNS + AUDIT_COLUMNS
@@ -341,13 +341,13 @@ def check_cross_fields(report: Report, rows: List[Dict[str, str]]) -> None:
secondary = row.get("koppenSecondaryClass") or ""
if zone == MIXED_KOPPEN_CLASS:
if not (primary and secondary):
report.add(CHECK_CROSS, f"{fips}: Mixed Koppen county needs koppenPrimaryClass and koppenSecondaryClass")
report.add(CHECK_CROSS, f"{fips}: Mixed Köppen county needs koppenPrimaryClass and koppenSecondaryClass")
elif primary not in KOPPEN_CODES or secondary not in KOPPEN_CODES:
report.add(CHECK_CROSS, f"{fips}: invalid Koppen stripe classes {primary!r}/{secondary!r}")
report.add(CHECK_CROSS, f"{fips}: invalid Köppen stripe classes {primary!r}/{secondary!r}")
elif primary == secondary:
report.add(CHECK_CROSS, f"{fips}: Koppen stripe classes are both {primary}")
report.add(CHECK_CROSS, f"{fips}: Köppen stripe classes are both {primary}")
elif primary or secondary:
report.add(CHECK_CROSS, f"{fips}: Koppen stripe classes are set but koppenZone is {zone or 'blank'}, not Mixed")
report.add(CHECK_CROSS, f"{fips}: Köppen stripe classes are set but koppenZone is {zone or 'blank'}, not Mixed")
if not (row.get("source") or "").strip():
report.add(CHECK_CROSS, f"{fips}: source is blank")
+39
View File
@@ -0,0 +1,39 @@
# `scripts/common/`
`scripts/common/` holds code used by more than one data source (NOAA, gridMET,
NSRDB, Köppen). Scripts import from it, for example
`from common.counties import load_counties`; nothing in it is run directly.
Rules for `common/`:
- **Cross-source only.** A helper goes in only if metrics from more than one
data source use it. Code shared by scripts of a single data source stays with
that source, for example a future NSRDB module for the NSRDB prompt,
redaction, and error-log helpers.
- **One topic per module.** Each module is named for its topic and has a
docstring. No catch-all `utils.py`.
- **Keep it small.** Before adding a helper, ask why it does not belong to any
one data source.
Other shared code moves when its filter is reviewed, so each move is tested
alongside that filter. Duplicated and drifted helpers found by the 2026-09-13
survey are finding 14 in
[docs/reviews/00-cross-filter.md](../../docs/reviews/00-cross-filter.md).
Current modules:
| Module | Contents | Why it is in `common/` |
| --- | --- | --- |
| `county_zonal_stats.py` | Raster windows, the 180th-meridian split, area-weighted class shares | Used by any raster-based metric |
| `counties.py` | County polygon loading, FIPS normalization, the state FIPS table | County identity is shared by nearly every pipeline |
| `koppen_legend.py` | The Köppen code map and legend loader | Temporary: also used by `build_county_climate_data.py`; moves next to the Köppen code in Phase 3 |
**Rationale.** Helper functions make each step of a computation explicit,
avoid repeated code, and can be tested separately
([Brown CSCI 0111, "Helper Functions"](https://cs.brown.edu/courses/csci0111/fall2018/lectures/helper-functions.html)).
Shared helper folders, however, tend to lose cohesion and collect unrelated
code; the recommended alternative is to keep code with the part of the system
it belongs to, allowing a shared folder only if it stays small and documented
([Helpers and Utils Folders in Software Architecture](https://dev.to/knzt/helpers-and-utils-folders-in-software-architecture-3f8h)).
The rules above follow both: shared functions, organized by topic and limited
to code that crosses data sources.
+2 -2
View File
@@ -1,4 +1,4 @@
"""Koppen-Geiger raster codes and the Beck et al. legend loader."""
"""Köppen-Geiger raster codes and the Beck et al. legend loader."""
from __future__ import annotations
@@ -42,7 +42,7 @@ DEFAULT_KOPPEN_CODE_MAP = {
def load_koppen_legend(legend_path: Path | None) -> Dict[int, str]:
"""Load Koppen raster codes, using defaults when no legend exists."""
"""Load Köppen raster codes, using defaults when no legend exists."""
if legend_path is None:
return DEFAULT_KOPPEN_CODE_MAP
+3 -3
View File
@@ -12,9 +12,9 @@ The browser blocks `fetch("data/climate-data.csv")` when `index.html` is opened
Then open [http://localhost:8000/](http://localhost:8000/). This keeps the app CSV-only while allowing the map and filters to load normally.
## Source 1: Koppen-Geiger classes (`koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`)
## Source 1: Köppen-Geiger classes (`koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`)
- Dataset: Beck et al. updated 1-km Koppen-Geiger climate classes (historical + future windows)
- Dataset: Beck et al. updated 1-km Köppen-Geiger climate classes (historical + future windows)
- Landing page: [https://www.gloh2o.org/koppen/](https://www.gloh2o.org/koppen/)
- Primary paper for updated release: [https://www.nature.com/articles/s41597-023-02549-6](https://www.nature.com/articles/s41597-023-02549-6)
- Coverage: 1901-2099 (use historical 1991-2020 layer for this project to align with NOAA baselines)
@@ -266,7 +266,7 @@ Then update the app CSV. Polygon archive GHI is used first; representative-point
## Metric definitions in generated output
- `koppenZone`: the county's predominant Koppen-Geiger class, meaning the class covering at least 50% of the county's land area and leading the runner-up by at least 5 percentage points; otherwise `Mixed`. Shares are area-weighted, with ocean and no-data cells excluded.
- `koppenZone`: the county's predominant Köppen-Geiger class, meaning the class covering at least 50% of the county's land area and leading the runner-up by at least 5 percentage points; otherwise `Mixed`. Shares are area-weighted, with ocean and no-data cells excluded.
- `koppenPrimaryClass` / `koppenSecondaryClass`: for Mixed counties only, the top and runner-up classes, drawn as stripes on the map; blank for predominant counties.
- `avgTempF`: mean of 12 monthly county mean temperatures, converted C -> F.
- `annualPrecipIn`: sum of 12 monthly county mean precipitation totals, converted mm -> inches.