Document restructuring and the beginnings of Filter 2 changes
Split the pipeline documentation by purpose so each fact has one home: - docs/pipeline-plan.md keeps the plan, checklist, tracker, and guardrails - docs/decisions.md holds open decisions and the dated decision log - docs/reviews/ holds findings and tasks: one file per filter, plus 00-cross-filter.md for findings that span filters - scripts/common/README.md holds the shared-helper rules (formerly Phase 2) - filter-calculations.md now describes calculations only Filed findings 12-22 from a consistency audit of the app, docs, and scripts. Filter 1 (Köppen-Geiger): use "Köppen" with the umlaut in all prose, labels, docstrings, help text, and checker messages (finding 21), and correct the base build's "majority" docstring (finding 22). Filter 2 (annual avg temperature): record the adopted definition in filter-calculations.md §2: equally weighted 1991-2020 monthly normals, per WMO-No. 1203 and NOAA's 2020 methodology; area-weighted county means; blank unless all 12 months exist. Code changes for this filter are still pending. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,4 +1,4 @@
|
||||
"""Apply the county Koppen-Geiger metric to the app CSV.
|
||||
"""Apply the county Köppen-Geiger metric to the app CSV.
|
||||
|
||||
Replaces the koppenZone column of data/climate-data.csv with the values in
|
||||
data/metrics/koppen.csv and writes koppenPrimaryClass and koppenSecondaryClass:
|
||||
@@ -71,7 +71,7 @@ def load_koppen_values(koppen_metric: Path) -> Dict[str, Dict[str, str]]:
|
||||
def apply_koppen_metric(
|
||||
climate_data: Path, koppen_metric: Path, out: Path, dry_run: bool = False
|
||||
) -> Tuple[List[Change], List[str]]:
|
||||
"""Update the Koppen columns; return the value changes and any columns that were added."""
|
||||
"""Update the Köppen columns; return the value changes and any columns that were added."""
|
||||
fields, rows = read_csv_rows(climate_data)
|
||||
if ZONE_FIELD not in fields:
|
||||
raise ValueError(f"{climate_data} has no {ZONE_FIELD} column.")
|
||||
@@ -101,9 +101,9 @@ def apply_koppen_metric(
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
"""Define and parse command-line options for this apply step."""
|
||||
parser = argparse.ArgumentParser(description="Apply the county Koppen metric to the app CSV.")
|
||||
parser = argparse.ArgumentParser(description="Apply the county Köppen metric to the app CSV.")
|
||||
parser.add_argument("--climate-data", type=Path, default=DEFAULT_CLIMATE_DATA, help="App climate CSV to update.")
|
||||
parser.add_argument("--koppen-metric", type=Path, default=DEFAULT_KOPPEN_METRIC, help="Koppen metric CSV.")
|
||||
parser.add_argument("--koppen-metric", type=Path, default=DEFAULT_KOPPEN_METRIC, help="Köppen metric CSV.")
|
||||
parser.add_argument("--out", type=Path, default=None, help="Output path; defaults to updating --climate-data in place.")
|
||||
parser.add_argument("--dry-run", action="store_true", help="Report changes without writing.")
|
||||
return parser.parse_args()
|
||||
|
||||
@@ -6,7 +6,7 @@ Outputs a CSV file compatible with the browser app:
|
||||
climate-data.csv
|
||||
|
||||
Metrics produced per county:
|
||||
- koppenZone: majority Koppen-Geiger class
|
||||
- koppenZone: largest-share Köppen-Geiger class among touched cells (replaced by the Köppen apply step)
|
||||
- avgTempF: annual mean temperature from NOAA 1991-2020 gridded normals
|
||||
- annualPrecipIn: annual total precipitation from NOAA 1991-2020 gridded normals
|
||||
- seasonalityIndex: precipitation seasonality, coefficient of variation of monthly totals (%)
|
||||
@@ -256,7 +256,7 @@ def _touched_raster_values(source, geometry: BaseGeometry, split_antimeridian: b
|
||||
|
||||
|
||||
def _zonal_majority_class(koppen_raster: Path, counties: gpd.GeoDataFrame, code_map: Dict[int, str]) -> List[str]:
|
||||
"""Assign each county its most common Koppen-Geiger class."""
|
||||
"""Assign each county its most common Köppen-Geiger class."""
|
||||
classes: List[str] = []
|
||||
with rasterio.open(koppen_raster) as source:
|
||||
raster_counties = counties
|
||||
@@ -597,7 +597,7 @@ def parse_args() -> argparse.Namespace:
|
||||
"""Define and parse command-line options for this generator."""
|
||||
parser = argparse.ArgumentParser(description="Generate county climate records for the web app.")
|
||||
parser.add_argument("--counties-geojson", type=Path, required=True, help="County polygon GeoJSON path.")
|
||||
parser.add_argument("--koppen-raster", type=Path, required=True, help="Koppen-Geiger raster TIFF path.")
|
||||
parser.add_argument("--koppen-raster", type=Path, required=True, help="Köppen-Geiger raster TIFF path.")
|
||||
parser.add_argument("--koppen-legend", type=Path, default=None, help="Optional legend.txt mapping integer codes.")
|
||||
parser.add_argument(
|
||||
"--monthly-tavg-nc",
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
"""Build the county Koppen-Geiger metric file from area-weighted class shares.
|
||||
"""Build the county Köppen-Geiger metric file from area-weighted class shares.
|
||||
|
||||
Writes data/metrics/koppen.csv. A county is predominantly its top class when
|
||||
that class covers at least 50% of the county's land and leads the runner-up by
|
||||
@@ -58,7 +58,7 @@ def rank_class_shares(weights: Dict[int, float], code_map: Dict[int, str]) -> Li
|
||||
return []
|
||||
unknown = sorted(code for code in weights if code not in code_map)
|
||||
if unknown:
|
||||
raise ValueError(f"Raster codes {unknown} are not in the Koppen legend.")
|
||||
raise ValueError(f"Raster codes {unknown} are not in the Köppen legend.")
|
||||
ranked = sorted(weights.items(), key=lambda item: (-item[1], item[0]))
|
||||
return [(code_map[code], weight / total) for code, weight in ranked]
|
||||
|
||||
@@ -117,7 +117,7 @@ def build_koppen_records(
|
||||
|
||||
|
||||
def write_records(records: List[dict], out_file: Path) -> None:
|
||||
"""Write the Koppen metric rows."""
|
||||
"""Write the Köppen metric rows."""
|
||||
out_file.parent.mkdir(parents=True, exist_ok=True)
|
||||
with out_file.open("w", encoding="utf-8", newline="") as csv_file:
|
||||
writer = csv.DictWriter(csv_file, fieldnames=FIELDS)
|
||||
@@ -127,9 +127,9 @@ def write_records(records: List[dict], out_file: Path) -> None:
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
"""Define and parse command-line options for this builder."""
|
||||
parser = argparse.ArgumentParser(description="Build the county Koppen-Geiger metric file.")
|
||||
parser = argparse.ArgumentParser(description="Build the county Köppen-Geiger metric file.")
|
||||
parser.add_argument("--counties-geojson", type=Path, default=DEFAULT_COUNTIES_GEOJSON, help="County polygon GeoJSON path.")
|
||||
parser.add_argument("--koppen-raster", type=Path, default=DEFAULT_KOPPEN_RASTER, help="Koppen-Geiger raster TIFF path.")
|
||||
parser.add_argument("--koppen-raster", type=Path, default=DEFAULT_KOPPEN_RASTER, help="Köppen-Geiger raster TIFF path.")
|
||||
parser.add_argument("--koppen-legend", type=Path, default=DEFAULT_KOPPEN_LEGEND, help="legend.txt mapping raster codes.")
|
||||
parser.add_argument("--subcells", type=int, default=DEFAULT_SUBCELLS, help="Sub-cells per raster cell edge.")
|
||||
parser.add_argument("--out", type=Path, default=DEFAULT_OUT, help="Output metric CSV path.")
|
||||
|
||||
@@ -110,7 +110,7 @@ METRIC_RULES: Dict[str, MetricRule] = {
|
||||
}
|
||||
|
||||
IDENTITY_COLUMNS = ("countyFips", "countyName", "state")
|
||||
# Stripe classes the app draws for Mixed Koppen counties; blank otherwise.
|
||||
# Stripe classes the app draws for Mixed Köppen counties; blank otherwise.
|
||||
KOPPEN_STRIPE_COLUMNS = ("koppenPrimaryClass", "koppenSecondaryClass")
|
||||
AUDIT_COLUMNS = ("humidHeatSourceFips", "humidHeatFipsAdjustment", "source")
|
||||
EXPECTED_COLUMNS = IDENTITY_COLUMNS + tuple(METRIC_RULES) + KOPPEN_STRIPE_COLUMNS + AUDIT_COLUMNS
|
||||
@@ -341,13 +341,13 @@ def check_cross_fields(report: Report, rows: List[Dict[str, str]]) -> None:
|
||||
secondary = row.get("koppenSecondaryClass") or ""
|
||||
if zone == MIXED_KOPPEN_CLASS:
|
||||
if not (primary and secondary):
|
||||
report.add(CHECK_CROSS, f"{fips}: Mixed Koppen county needs koppenPrimaryClass and koppenSecondaryClass")
|
||||
report.add(CHECK_CROSS, f"{fips}: Mixed Köppen county needs koppenPrimaryClass and koppenSecondaryClass")
|
||||
elif primary not in KOPPEN_CODES or secondary not in KOPPEN_CODES:
|
||||
report.add(CHECK_CROSS, f"{fips}: invalid Koppen stripe classes {primary!r}/{secondary!r}")
|
||||
report.add(CHECK_CROSS, f"{fips}: invalid Köppen stripe classes {primary!r}/{secondary!r}")
|
||||
elif primary == secondary:
|
||||
report.add(CHECK_CROSS, f"{fips}: Koppen stripe classes are both {primary}")
|
||||
report.add(CHECK_CROSS, f"{fips}: Köppen stripe classes are both {primary}")
|
||||
elif primary or secondary:
|
||||
report.add(CHECK_CROSS, f"{fips}: Koppen stripe classes are set but koppenZone is {zone or 'blank'}, not Mixed")
|
||||
report.add(CHECK_CROSS, f"{fips}: Köppen stripe classes are set but koppenZone is {zone or 'blank'}, not Mixed")
|
||||
|
||||
if not (row.get("source") or "").strip():
|
||||
report.add(CHECK_CROSS, f"{fips}: source is blank")
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
# `scripts/common/`
|
||||
|
||||
`scripts/common/` holds code used by more than one data source (NOAA, gridMET,
|
||||
NSRDB, Köppen). Scripts import from it, for example
|
||||
`from common.counties import load_counties`; nothing in it is run directly.
|
||||
|
||||
Rules for `common/`:
|
||||
|
||||
- **Cross-source only.** A helper goes in only if metrics from more than one
|
||||
data source use it. Code shared by scripts of a single data source stays with
|
||||
that source, for example a future NSRDB module for the NSRDB prompt,
|
||||
redaction, and error-log helpers.
|
||||
- **One topic per module.** Each module is named for its topic and has a
|
||||
docstring. No catch-all `utils.py`.
|
||||
- **Keep it small.** Before adding a helper, ask why it does not belong to any
|
||||
one data source.
|
||||
|
||||
Other shared code moves when its filter is reviewed, so each move is tested
|
||||
alongside that filter. Duplicated and drifted helpers found by the 2026-09-13
|
||||
survey are finding 14 in
|
||||
[docs/reviews/00-cross-filter.md](../../docs/reviews/00-cross-filter.md).
|
||||
|
||||
Current modules:
|
||||
|
||||
| Module | Contents | Why it is in `common/` |
|
||||
| --- | --- | --- |
|
||||
| `county_zonal_stats.py` | Raster windows, the 180th-meridian split, area-weighted class shares | Used by any raster-based metric |
|
||||
| `counties.py` | County polygon loading, FIPS normalization, the state FIPS table | County identity is shared by nearly every pipeline |
|
||||
| `koppen_legend.py` | The Köppen code map and legend loader | Temporary: also used by `build_county_climate_data.py`; moves next to the Köppen code in Phase 3 |
|
||||
|
||||
**Rationale.** Helper functions make each step of a computation explicit,
|
||||
avoid repeated code, and can be tested separately
|
||||
([Brown CSCI 0111, "Helper Functions"](https://cs.brown.edu/courses/csci0111/fall2018/lectures/helper-functions.html)).
|
||||
Shared helper folders, however, tend to lose cohesion and collect unrelated
|
||||
code; the recommended alternative is to keep code with the part of the system
|
||||
it belongs to, allowing a shared folder only if it stays small and documented
|
||||
([Helpers and Utils Folders in Software Architecture](https://dev.to/knzt/helpers-and-utils-folders-in-software-architecture-3f8h)).
|
||||
The rules above follow both: shared functions, organized by topic and limited
|
||||
to code that crosses data sources.
|
||||
@@ -1,4 +1,4 @@
|
||||
"""Koppen-Geiger raster codes and the Beck et al. legend loader."""
|
||||
"""Köppen-Geiger raster codes and the Beck et al. legend loader."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -42,7 +42,7 @@ DEFAULT_KOPPEN_CODE_MAP = {
|
||||
|
||||
|
||||
def load_koppen_legend(legend_path: Path | None) -> Dict[int, str]:
|
||||
"""Load Koppen raster codes, using defaults when no legend exists."""
|
||||
"""Load Köppen raster codes, using defaults when no legend exists."""
|
||||
if legend_path is None:
|
||||
return DEFAULT_KOPPEN_CODE_MAP
|
||||
|
||||
|
||||
@@ -12,9 +12,9 @@ The browser blocks `fetch("data/climate-data.csv")` when `index.html` is opened
|
||||
|
||||
Then open [http://localhost:8000/](http://localhost:8000/). This keeps the app CSV-only while allowing the map and filters to load normally.
|
||||
|
||||
## Source 1: Koppen-Geiger classes (`koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`)
|
||||
## Source 1: Köppen-Geiger classes (`koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`)
|
||||
|
||||
- Dataset: Beck et al. updated 1-km Koppen-Geiger climate classes (historical + future windows)
|
||||
- Dataset: Beck et al. updated 1-km Köppen-Geiger climate classes (historical + future windows)
|
||||
- Landing page: [https://www.gloh2o.org/koppen/](https://www.gloh2o.org/koppen/)
|
||||
- Primary paper for updated release: [https://www.nature.com/articles/s41597-023-02549-6](https://www.nature.com/articles/s41597-023-02549-6)
|
||||
- Coverage: 1901-2099 (use historical 1991-2020 layer for this project to align with NOAA baselines)
|
||||
@@ -266,7 +266,7 @@ Then update the app CSV. Polygon archive GHI is used first; representative-point
|
||||
|
||||
## Metric definitions in generated output
|
||||
|
||||
- `koppenZone`: the county's predominant Koppen-Geiger class, meaning the class covering at least 50% of the county's land area and leading the runner-up by at least 5 percentage points; otherwise `Mixed`. Shares are area-weighted, with ocean and no-data cells excluded.
|
||||
- `koppenZone`: the county's predominant Köppen-Geiger class, meaning the class covering at least 50% of the county's land area and leading the runner-up by at least 5 percentage points; otherwise `Mixed`. Shares are area-weighted, with ocean and no-data cells excluded.
|
||||
- `koppenPrimaryClass` / `koppenSecondaryClass`: for Mixed counties only, the top and runner-up classes, drawn as stripes on the map; blank for predominant counties.
|
||||
- `avgTempF`: mean of 12 monthly county mean temperatures, converted C -> F.
|
||||
- `annualPrecipIn`: sum of 12 monthly county mean precipitation totals, converted mm -> inches.
|
||||
|
||||
Reference in New Issue
Block a user