Complete Köppen-Geiger filter review with Mixed climate class

Classify each county by area-weighted Köppen class shares: a county is
predominantly its top class when that class covers at least 50% of its
land and leads the runner-up by at least 5 percentage points; otherwise
it is Mixed (133 of 3,143 counties in the 50 states and DC).

- Add build_county_koppen_metric.py (writes data/metrics/koppen.csv) and
  apply_koppen_metric_to_climate_data.py (writes koppenZone plus
  koppenPrimaryClass/koppenSecondaryClass for Mixed counties).
- Move shared helpers into scripts/common/ (county loading, Köppen
  legend, area-weighted raster shares); fix the 180th-meridian raster
  window for Aleutians West.
- Add check_climate_data.py to validate the app CSV.
- Draw Mixed counties in app.js as diagonal stripes of their top two
  classes, fixed to the ground and following the map at every zoom, with
  a crossfade only when the stripe size changes. Filtering a class also
  matches Mixed counties where it is primary or secondary.
- Document the rule, display, and pipeline plan in docs/ and update the
  README and data-source notes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-14 02:54:25 -04:00
co-authored by Claude Opus 5
parent 92fbbfb2e9
commit 4d2b3e3d44
20 changed files with 6400 additions and 3450 deletions
@@ -0,0 +1,138 @@
"""Apply the county Koppen-Geiger metric to the app CSV.
Replaces the koppenZone column of data/climate-data.csv with the values in
data/metrics/koppen.csv and writes koppenPrimaryClass and koppenSecondaryClass:
the two stripe classes the app draws for Mixed counties. Both are blank for
predominant counties. The two columns are added after koppenZone if missing.
Every other column, and the column order, is left unchanged. Use --dry-run to
report the changes without writing.
Run:
.venv\\Scripts\\python.exe scripts\\apply_koppen_metric_to_climate_data.py --dry-run
"""
from __future__ import annotations
import argparse
import csv
from collections import Counter
from pathlib import Path
from typing import Dict, List, Tuple
REPO_ROOT = Path(__file__).resolve().parents[1]
DEFAULT_CLIMATE_DATA = REPO_ROOT / "data" / "climate-data.csv"
DEFAULT_KOPPEN_METRIC = REPO_ROOT / "data" / "metrics" / "koppen.csv"
ZONE_FIELD = "koppenZone"
PRIMARY_FIELD = "koppenPrimaryClass"
SECONDARY_FIELD = "koppenSecondaryClass"
STRIPE_FIELDS = [PRIMARY_FIELD, SECONDARY_FIELD]
MIXED_CLASS = "Mixed"
METRIC_FIELDS = ["countyFips", ZONE_FIELD, "koppenTopClass", "koppenSecondClass"]
# (countyFips, column, old value, new value)
Change = Tuple[str, str, str, str]
def read_csv_rows(path: Path) -> Tuple[List[str], List[dict]]:
"""Read a CSV while preserving the source field order."""
with path.open("r", encoding="utf-8-sig", newline="") as csv_file:
reader = csv.DictReader(csv_file)
if reader.fieldnames is None:
raise ValueError(f"{path} has no CSV header.")
return list(reader.fieldnames), list(reader)
def fieldnames_with_stripe_columns(fields: List[str]) -> List[str]:
"""Place the stripe-class columns directly after koppenZone."""
base = [field for field in fields if field not in STRIPE_FIELDS]
insert_at = base.index(ZONE_FIELD) + 1
return base[:insert_at] + STRIPE_FIELDS + base[insert_at:]
def load_koppen_values(koppen_metric: Path) -> Dict[str, Dict[str, str]]:
"""Return koppenZone and the two stripe classes for each county in the metric file."""
fields, rows = read_csv_rows(koppen_metric)
missing_fields = [field for field in METRIC_FIELDS if field not in fields]
if missing_fields:
raise ValueError(f"{koppen_metric} is missing columns {missing_fields}.")
values: Dict[str, Dict[str, str]] = {}
for row in rows:
is_mixed = row[ZONE_FIELD] == MIXED_CLASS
primary = row["koppenTopClass"] if is_mixed else ""
secondary = row["koppenSecondClass"] if is_mixed else ""
if is_mixed and not (primary and secondary):
raise ValueError(f"{row['countyFips']} is Mixed but has no top or second class in {koppen_metric}.")
values[row["countyFips"]] = {ZONE_FIELD: row[ZONE_FIELD], PRIMARY_FIELD: primary, SECONDARY_FIELD: secondary}
return values
def apply_koppen_metric(
climate_data: Path, koppen_metric: Path, out: Path, dry_run: bool = False
) -> Tuple[List[Change], List[str]]:
"""Update the Koppen columns; return the value changes and any columns that were added."""
fields, rows = read_csv_rows(climate_data)
if ZONE_FIELD not in fields:
raise ValueError(f"{climate_data} has no {ZONE_FIELD} column.")
lookup = load_koppen_values(koppen_metric)
missing = [row["countyFips"] for row in rows if row["countyFips"] not in lookup]
if missing:
raise ValueError(f"{len(missing)} counties are missing from {koppen_metric}, e.g. {missing[:5]}")
added_columns = [field for field in STRIPE_FIELDS if field not in fields]
changes: List[Change] = []
for row in rows:
new_values = lookup[row["countyFips"]]
for field in [ZONE_FIELD, *STRIPE_FIELDS]:
old_value = row.get(field) or ""
if old_value != new_values[field]:
changes.append((row["countyFips"], field, old_value, new_values[field]))
row[field] = new_values[field]
if not dry_run:
with out.open("w", encoding="utf-8", newline="") as csv_file:
writer = csv.DictWriter(csv_file, fieldnames=fieldnames_with_stripe_columns(fields))
writer.writeheader()
writer.writerows(rows)
return changes, added_columns
def parse_args() -> argparse.Namespace:
"""Define and parse command-line options for this apply step."""
parser = argparse.ArgumentParser(description="Apply the county Koppen metric to the app CSV.")
parser.add_argument("--climate-data", type=Path, default=DEFAULT_CLIMATE_DATA, help="App climate CSV to update.")
parser.add_argument("--koppen-metric", type=Path, default=DEFAULT_KOPPEN_METRIC, help="Koppen metric CSV.")
parser.add_argument("--out", type=Path, default=None, help="Output path; defaults to updating --climate-data in place.")
parser.add_argument("--dry-run", action="store_true", help="Report changes without writing.")
return parser.parse_args()
def main() -> None:
args = parse_args()
changes, added_columns = apply_koppen_metric(
args.climate_data, args.koppen_metric, args.out or args.climate_data, args.dry_run
)
verb = "would" if args.dry_run else "did"
if added_columns:
print(f"Columns added after {ZONE_FIELD} ({verb} write): {', '.join(added_columns)}")
zone_changes = [change for change in changes if change[1] == ZONE_FIELD]
kinds = Counter(
"to Mixed" if new == MIXED_CLASS else "to blank" if not new else "to another class"
for _, _, _, new in zone_changes
)
stripe_values = sum(1 for _, field, _, new in changes if field in STRIPE_FIELDS and new)
print(f"{len(zone_changes)} {ZONE_FIELD} values {'would change' if args.dry_run else 'changed'}: {dict(kinds)}")
print(f"{stripe_values} stripe-class values {'would be set' if args.dry_run else 'set'}.")
for fips, _, old, new in zone_changes[:10]:
print(f" {fips}: {old or '(blank)'} -> {new or '(blank)'}")
if len(zone_changes) > 10:
print(f" ... and {len(zone_changes) - 10} more")
if args.dry_run:
print("Dry run: nothing was written.")
if __name__ == "__main__":
main()
@@ -18,10 +18,10 @@ from build_county_climate_data import (
MONTH_NAMES,
_as_monthly_climatology,
_extract_grid_2d,
_load_counties,
_select_data_var,
_zonal_mean,
)
from common.counties import load_counties
REPO_ROOT = Path(__file__).resolve().parents[1]
DEFAULT_CLIMATE_DATA = REPO_ROOT / "data" / "climate-data.csv"
@@ -61,7 +61,7 @@ def build_precip_month_lookup(
climatology_end_year: int,
) -> dict[str, tuple[str, str]]:
"""Calculate wettest and driest precipitation month for each county."""
counties = _load_counties(counties_geojson)
counties = load_counties(counties_geojson)
monthly_prcp = xr.open_dataset(monthly_prcp_nc, decode_times=True)
try:
prcp_var = _select_data_var(monthly_prcp, "mlyprcp_norm")
+30 -199
View File
@@ -31,102 +31,11 @@ import rasterio
import xarray as xr
from affine import Affine
from rasterio.features import geometry_mask
from shapely.geometry.base import BaseGeometry
DEFAULT_COUNTIES_GEOJSON_URL = "https://raw.githubusercontent.com/plotly/datasets/master/geojson-counties-fips.json"
STATE_FIPS_TO_ABBR = {
"01": "AL",
"02": "AK",
"04": "AZ",
"05": "AR",
"06": "CA",
"08": "CO",
"09": "CT",
"10": "DE",
"11": "DC",
"12": "FL",
"13": "GA",
"15": "HI",
"16": "ID",
"17": "IL",
"18": "IN",
"19": "IA",
"20": "KS",
"21": "KY",
"22": "LA",
"23": "ME",
"24": "MD",
"25": "MA",
"26": "MI",
"27": "MN",
"28": "MS",
"29": "MO",
"30": "MT",
"31": "NE",
"32": "NV",
"33": "NH",
"34": "NJ",
"35": "NM",
"36": "NY",
"37": "NC",
"38": "ND",
"39": "OH",
"40": "OK",
"41": "OR",
"42": "PA",
"44": "RI",
"45": "SC",
"46": "SD",
"47": "TN",
"48": "TX",
"49": "UT",
"50": "VT",
"51": "VA",
"53": "WA",
"54": "WV",
"55": "WI",
"56": "WY",
"60": "AS",
"66": "GU",
"69": "MP",
"72": "PR",
"78": "VI",
}
# Beck et al legend key is expected as text file, but this default handles common codes.
DEFAULT_KOPPEN_CODE_MAP = {
1: "Af",
2: "Am",
3: "Aw",
4: "BWh",
5: "BWk",
6: "BSh",
7: "BSk",
8: "Csa",
9: "Csb",
10: "Csc",
11: "Cwa",
12: "Cwb",
13: "Cwc",
14: "Cfa",
15: "Cfb",
16: "Cfc",
17: "Dsa",
18: "Dsb",
19: "Dsc",
20: "Dsd",
21: "Dwa",
22: "Dwb",
23: "Dwc",
24: "Dwd",
25: "Dfa",
26: "Dfb",
27: "Dfc",
28: "Dfd",
29: "ET",
30: "EF",
}
from common.counties import load_counties, normalize_fips
from common.county_zonal_stats import geometry_window, split_at_antimeridian
from common.koppen_legend import load_koppen_legend
MONTH_NAMES = [
"January",
@@ -144,94 +53,6 @@ MONTH_NAMES = [
]
def _normalize_fips(value: object, width: int) -> str:
"""Return a zero-padded FIPS code with the requested width."""
text = str(value).strip()
digits = "".join(ch for ch in text if ch.isdigit())
if not digits:
return ""
return digits.zfill(width)[-width:]
def _load_counties(counties_geojson: Path) -> gpd.GeoDataFrame:
"""Load county polygons and normalize fields used downstream."""
if not counties_geojson.exists():
try:
print(
f"County GeoJSON not found at {counties_geojson}. "
f"Attempting download from {DEFAULT_COUNTIES_GEOJSON_URL}..."
)
gdf = gpd.read_file(DEFAULT_COUNTIES_GEOJSON_URL)
counties_geojson.parent.mkdir(parents=True, exist_ok=True)
# Cache the downloaded file for subsequent runs.
gdf.to_file(counties_geojson, driver="GeoJSON")
print(f"Downloaded and cached county GeoJSON to {counties_geojson}")
except Exception as exc:
raise FileNotFoundError(
f"County GeoJSON not found at {counties_geojson}, and download from "
f"{DEFAULT_COUNTIES_GEOJSON_URL} failed. Download the file manually "
"and rerun with --counties-geojson pointing to it."
) from exc
gdf = gpd.read_file(counties_geojson)
if gdf.crs is None:
gdf = gdf.set_crs("EPSG:4326")
else:
gdf = gdf.to_crs("EPSG:4326")
feature_id = None
if "id" in gdf.columns:
feature_id = gdf["id"]
elif "GEOID" in gdf.columns:
feature_id = gdf["GEOID"]
elif "GEOID10" in gdf.columns:
feature_id = gdf["GEOID10"]
elif "fips" in gdf.columns:
feature_id = gdf["fips"]
else:
raise ValueError("Unable to locate county FIPS identifier column in county polygons.")
gdf["county_fips"] = feature_id.map(lambda value: _normalize_fips(value, 5))
gdf = gdf[gdf["county_fips"] != ""].copy()
if "NAME" in gdf.columns:
gdf["county_name"] = gdf["NAME"].fillna("").astype(str).str.strip()
elif "name" in gdf.columns:
gdf["county_name"] = gdf["name"].fillna("").astype(str).str.strip()
else:
gdf["county_name"] = gdf["county_fips"].map(lambda value: f"County {value}")
gdf["state_fips"] = gdf["county_fips"].str.slice(0, 2)
gdf["state"] = gdf["state_fips"].map(lambda code: STATE_FIPS_TO_ABBR.get(code, f"S{code}"))
gdf = gdf.sort_values("county_fips").reset_index(drop=True)
return gdf
def _load_koppen_legend(legend_path: Path | None) -> Dict[int, str]:
"""Load Koppen raster codes, using defaults when no legend exists."""
if legend_path is None:
return DEFAULT_KOPPEN_CODE_MAP
mapping: Dict[int, str] = {}
for line in legend_path.read_text(encoding="utf-8").splitlines():
text = line.strip()
if not text or text.startswith("#"):
continue
# Handles patterns like:
# "1: Af ..." or "1 = Af" or "1 Af"
import re
match = re.match(r"^(\d+)\s*[:=]?\s*([A-Za-z]{2,3})\b", text)
if not match:
continue
key = int(match.group(1))
value = match.group(2)
mapping[key] = value
return mapping if mapping else DEFAULT_KOPPEN_CODE_MAP
def _select_data_var(dataset: xr.Dataset, preferred: str) -> str:
"""Choose the best matching climate variable from a dataset."""
if preferred in dataset.data_vars:
@@ -283,7 +104,7 @@ def _load_solar_ghi_csv(solar_ghi_csv: Path, counties: gpd.GeoDataFrame) -> List
solar_by_fips: Dict[str, float] = {}
for row in reader:
county_fips = _normalize_fips(row.get(fips_field, ""), 5)
county_fips = normalize_fips(row.get(fips_field, ""), 5)
raw_value = str(row.get("meanDailyGlobalHorizontalRadiationKwhM2Day", "")).strip()
if not county_fips or not raw_value:
continue
@@ -414,6 +235,26 @@ def _zonal_mean_raster(raster_path: Path, counties: gpd.GeoDataFrame) -> List[fl
return _zonal_mean(values, source.transform, raster_counties)
def _touched_raster_values(source, geometry: BaseGeometry, split_antimeridian: bool) -> np.ndarray:
"""Read every raster cell a geometry touches, using a small window per piece."""
pieces = split_at_antimeridian(geometry) if split_antimeridian else [geometry]
selected: List[np.ndarray] = []
for piece in pieces:
window = geometry_window(source, piece)
values = source.read(1, window=window, masked=True).filled(0)
if source.nodata is not None:
values = np.where(values == source.nodata, 0, values)
mask = geometry_mask(
[piece.__geo_interface__],
out_shape=values.shape,
transform=source.window_transform(window),
invert=True,
all_touched=True,
)
selected.append(values[mask])
return np.concatenate(selected) if selected else np.array([], dtype=np.int64)
def _zonal_majority_class(koppen_raster: Path, counties: gpd.GeoDataFrame, code_map: Dict[int, str]) -> List[str]:
"""Assign each county its most common Koppen-Geiger class."""
classes: List[str] = []
@@ -421,21 +262,11 @@ def _zonal_majority_class(koppen_raster: Path, counties: gpd.GeoDataFrame, code_
raster_counties = counties
if source.crs is not None and counties.crs is not None and counties.crs != source.crs:
raster_counties = counties.to_crs(source.crs)
data = source.read(1, masked=True)
values = np.asarray(data.filled(0))
if source.nodata is not None:
values = np.where(values == source.nodata, 0, values)
# The 180-degree split only makes sense for longitude/latitude rasters.
split_antimeridian = source.crs is None or source.crs.is_geographic
for geometry in raster_counties.geometry:
mask = geometry_mask(
[geometry.__geo_interface__],
out_shape=values.shape,
transform=source.transform,
invert=True,
all_touched=True,
)
selected = values[mask]
selected = _touched_raster_values(source, geometry, split_antimeridian)
selected = selected[selected != 0]
if selected.size == 0:
classes.append("Cfa")
@@ -545,9 +376,9 @@ def build_county_records(
solar_ghi_csv: Path | None,
) -> Dict[str, dict]:
"""Build county climate records consumed by the web app."""
counties = _load_counties(counties_geojson)
counties = load_counties(counties_geojson)
koppen_classes = _zonal_majority_class(koppen_raster, counties, _load_koppen_legend(koppen_legend))
koppen_classes = _zonal_majority_class(koppen_raster, counties, load_koppen_legend(koppen_legend))
monthly_tavg = xr.open_dataset(monthly_tavg_nc, decode_times=True)
monthly_prcp = xr.open_dataset(monthly_prcp_nc, decode_times=True)
+155
View File
@@ -0,0 +1,155 @@
"""Build the county Koppen-Geiger metric file from area-weighted class shares.
Writes data/metrics/koppen.csv. A county is predominantly its top class when
that class covers at least 50% of the county's land and leads the runner-up by
at least 5 percentage points; otherwise it is "Mixed". Counties with no valid
raster cells are left blank. See docs/filter-calculations.md, section 1.
Run:
.venv\\Scripts\\python.exe scripts\\build_county_koppen_metric.py
"""
from __future__ import annotations
import argparse
import csv
from collections import Counter
from pathlib import Path
from typing import Dict, List, Tuple
import geopandas as gpd
import rasterio
from common.counties import load_counties
from common.county_zonal_stats import DEFAULT_SUBCELLS, area_weighted_class_weights
from common.koppen_legend import load_koppen_legend
PROJECT_ROOT = Path(__file__).resolve().parents[1]
DEFAULT_COUNTIES_GEOJSON = PROJECT_ROOT / "data" / "geojson-counties-fips.json"
DEFAULT_KOPPEN_RASTER = PROJECT_ROOT / "data" / "koppen_geiger_tif" / "1991_2020" / "koppen_geiger_0p00833333.tif"
DEFAULT_KOPPEN_LEGEND = PROJECT_ROOT / "data" / "koppen_geiger_tif" / "legend.txt"
DEFAULT_OUT = PROJECT_ROOT / "data" / "metrics" / "koppen.csv"
PREDOMINANT_MIN_SHARE = 0.50
PREDOMINANT_MIN_GAP = 0.05
MIXED_CLASS = "Mixed"
# Keeps shares that sit exactly on a cutoff from failing on floating-point error.
RULE_TOLERANCE = 1e-9
FIELDS = [
"countyFips",
"countyName",
"state",
"koppenZone",
"koppenTopClass",
"koppenTopShare",
"koppenSecondClass",
"koppenSecondShare",
]
def rank_class_shares(weights: Dict[int, float], code_map: Dict[int, str]) -> List[Tuple[str, float]]:
"""Convert per-code area weights to class shares, largest first.
Exact ties go to the smaller raster code, matching the previous build.
"""
total = sum(weights.values())
if total <= 0:
return []
unknown = sorted(code for code in weights if code not in code_map)
if unknown:
raise ValueError(f"Raster codes {unknown} are not in the Koppen legend.")
ranked = sorted(weights.items(), key=lambda item: (-item[1], item[0]))
return [(code_map[code], weight / total) for code, weight in ranked]
def classify(ranked: List[Tuple[str, float]]) -> str:
"""Return the predominant class, "Mixed", or blank when there is no data."""
if not ranked:
return ""
top_share = ranked[0][1]
second_share = ranked[1][1] if len(ranked) > 1 else 0.0
has_majority = top_share >= PREDOMINANT_MIN_SHARE - RULE_TOLERANCE
has_clear_lead = top_share - second_share >= PREDOMINANT_MIN_GAP - RULE_TOLERANCE
return ranked[0][0] if has_majority and has_clear_lead else MIXED_CLASS
def _format_share(ranked: List[Tuple[str, float]], index: int) -> Tuple[str, str]:
"""Return the class and 4-decimal share at a rank, or blanks."""
if index >= len(ranked):
return "", ""
code, share = ranked[index]
return code, f"{share:.4f}"
def build_koppen_records(
counties: gpd.GeoDataFrame,
koppen_raster: Path,
code_map: Dict[int, str],
subcells: int = DEFAULT_SUBCELLS,
) -> List[dict]:
"""Classify every county and return its metric row."""
records: List[dict] = []
with rasterio.open(koppen_raster) as source:
raster_counties = counties
if source.crs is not None and counties.crs is not None and counties.crs != source.crs:
raster_counties = counties.to_crs(source.crs)
geographic = source.crs is None or source.crs.is_geographic
for county, geometry in zip(counties.itertuples(), raster_counties.geometry):
weights = area_weighted_class_weights(source, geometry, geographic=geographic, subcells=subcells)
ranked = rank_class_shares(weights, code_map)
top_class, top_share = _format_share(ranked, 0)
second_class, second_share = _format_share(ranked, 1)
records.append(
{
"countyFips": county.county_fips,
"countyName": county.county_name,
"state": county.state,
"koppenZone": classify(ranked),
"koppenTopClass": top_class,
"koppenTopShare": top_share,
"koppenSecondClass": second_class,
"koppenSecondShare": second_share,
}
)
return records
def write_records(records: List[dict], out_file: Path) -> None:
"""Write the Koppen metric rows."""
out_file.parent.mkdir(parents=True, exist_ok=True)
with out_file.open("w", encoding="utf-8", newline="") as csv_file:
writer = csv.DictWriter(csv_file, fieldnames=FIELDS)
writer.writeheader()
writer.writerows(records)
def parse_args() -> argparse.Namespace:
"""Define and parse command-line options for this builder."""
parser = argparse.ArgumentParser(description="Build the county Koppen-Geiger metric file.")
parser.add_argument("--counties-geojson", type=Path, default=DEFAULT_COUNTIES_GEOJSON, help="County polygon GeoJSON path.")
parser.add_argument("--koppen-raster", type=Path, default=DEFAULT_KOPPEN_RASTER, help="Koppen-Geiger raster TIFF path.")
parser.add_argument("--koppen-legend", type=Path, default=DEFAULT_KOPPEN_LEGEND, help="legend.txt mapping raster codes.")
parser.add_argument("--subcells", type=int, default=DEFAULT_SUBCELLS, help="Sub-cells per raster cell edge.")
parser.add_argument("--out", type=Path, default=DEFAULT_OUT, help="Output metric CSV path.")
return parser.parse_args()
def main() -> None:
args = parse_args()
counties = load_counties(args.counties_geojson)
records = build_koppen_records(counties, args.koppen_raster, load_koppen_legend(args.koppen_legend), args.subcells)
write_records(records, args.out)
outcomes = Counter(
"blank" if not r["koppenZone"] else "Mixed" if r["koppenZone"] == MIXED_CLASS else "predominant" for r in records
)
print(
f"Wrote {len(records)} counties to {args.out}: {outcomes['predominant']} predominant, "
f"{outcomes['Mixed']} Mixed, {outcomes['blank']} blank."
)
if __name__ == "__main__":
main()
+427
View File
@@ -0,0 +1,427 @@
"""Validate data/climate-data.csv before the browser app loads it.
Checks file format, columns, county keys, the county geometry join, per-metric
types and ranges, allowed blank values, and cross-field consistency. Exits with
status 1 when any check fails.
Run:
.venv\\Scripts\\python.exe scripts\\check_climate_data.py
"""
from __future__ import annotations
import argparse
import csv
import json
import math
import re
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Dict, FrozenSet, List, Optional, Sequence, Tuple
PROJECT_ROOT = Path(__file__).resolve().parents[1]
DEFAULT_CLIMATE_CSV = PROJECT_ROOT / "data" / "climate-data.csv"
DEFAULT_COUNTIES_GEOJSON = PROJECT_ROOT / "data" / "geojson-counties-fips.json"
DEFAULT_METRIC_SOURCES = PROJECT_ROOT / "data" / "metric_sources.json"
FIPS_PATTERN = re.compile(r"\d{5}")
STATE_PATTERN = re.compile(r"[A-Z]{2}")
# The distinct-value check only runs on full-size files, not small fixtures.
COLLAPSE_CHECK_MIN_ROWS = 100
KOPPEN_CODES = (
"Af", "Am", "Aw", "BWh", "BWk", "BSh", "BSk",
"Csa", "Csb", "Csc", "Cwa", "Cwb", "Cwc", "Cfa", "Cfb", "Cfc",
"Dsa", "Dsb", "Dsc", "Dsd", "Dwa", "Dwb", "Dwc", "Dwd",
"Dfa", "Dfb", "Dfc", "Dfd", "ET", "EF",
)
# Counties with no predominant class; see docs/filter-calculations.md, section 1.
MIXED_KOPPEN_CLASS = "Mixed"
MONTH_NAMES = (
"January", "February", "March", "April", "May", "June",
"July", "August", "September", "October", "November", "December",
)
# NOAA nClimGrid and gridMET cover the contiguous U.S. only.
OUTSIDE_CONUS = frozenset({"AK", "HI", "PR"})
# NSRDB summaries were not requested for Puerto Rico.
NSRDB_NOT_REQUESTED = frozenset({"PR"})
# NOAA county daily files do not include Lexington city, VA.
NOAA_DAILY_MISSING_FIPS = frozenset({"51678"})
@dataclass(frozen=True)
class MetricRule:
"""Validation rule for one metric column."""
kind: str
minimum: Optional[float] = None
maximum: Optional[float] = None
integer: bool = False
categories: Tuple[str, ...] = ()
blank_states: FrozenSet[str] = frozenset()
blank_fips: FrozenSet[str] = frozenset()
min_distinct: int = 2
def _numeric(
minimum: float,
maximum: float,
*,
integer: bool = False,
blank_states: FrozenSet[str] = frozenset(),
blank_fips: FrozenSet[str] = frozenset(),
) -> MetricRule:
"""Build a numeric rule with a plausible physical range."""
return MetricRule(
"numeric",
minimum=minimum,
maximum=maximum,
integer=integer,
blank_states=blank_states,
blank_fips=blank_fips,
min_distinct=10,
)
def _categorical(categories: Sequence[str], *, blank_states: FrozenSet[str] = frozenset()) -> MetricRule:
"""Build a categorical rule limited to an allowed value list."""
return MetricRule("categorical", categories=tuple(categories), blank_states=blank_states)
# Ranges are physical plausibility limits, deliberately wider than the current
# data. They are not the app's color-scale bounds.
METRIC_RULES: Dict[str, MetricRule] = {
"koppenZone": _categorical(KOPPEN_CODES + (MIXED_KOPPEN_CLASS,)),
"avgTempF": _numeric(20, 85, blank_states=OUTSIDE_CONUS),
"avgDiurnalTempRangeF": _numeric(5, 40, blank_states=OUTSIDE_CONUS, blank_fips=NOAA_DAILY_MISSING_FIPS),
"annualPrecipIn": _numeric(1, 150, blank_states=OUTSIDE_CONUS),
"seasonalityIndex": _numeric(0, 100, integer=True, blank_states=OUTSIDE_CONUS),
"wettestPrecipMonth": _categorical(MONTH_NAMES, blank_states=OUTSIDE_CONUS),
"driestPrecipMonth": _categorical(MONTH_NAMES, blank_states=OUTSIDE_CONUS),
"absoluteExtremeDays": _numeric(0, 366, blank_states=OUTSIDE_CONUS, blank_fips=NOAA_DAILY_MISSING_FIPS),
"meanDailyGlobalHorizontalRadiationKwhM2Day": _numeric(1.5, 7.5, blank_states=NSRDB_NOT_REQUESTED),
"clearSkyGhiReductionIndex": _numeric(0, 1, blank_states=NSRDB_NOT_REQUESTED),
"avgSummerSpecificHumidityGKg": _numeric(1, 25, blank_states=OUTSIDE_CONUS),
"humidHeatDays": _numeric(0, 366, blank_states=OUTSIDE_CONUS),
}
IDENTITY_COLUMNS = ("countyFips", "countyName", "state")
# Stripe classes the app draws for Mixed Koppen counties; blank otherwise.
KOPPEN_STRIPE_COLUMNS = ("koppenPrimaryClass", "koppenSecondaryClass")
AUDIT_COLUMNS = ("humidHeatSourceFips", "humidHeatFipsAdjustment", "source")
EXPECTED_COLUMNS = IDENTITY_COLUMNS + tuple(METRIC_RULES) + KOPPEN_STRIPE_COLUMNS + AUDIT_COLUMNS
CHECK_FORMAT = "file format"
CHECK_COLUMNS = "columns"
CHECK_COUNTIES = "county keys"
CHECK_GEOMETRY = "map geometry join"
CHECK_BLANKS = "blank values"
CHECK_VALUES = "value types and ranges"
CHECK_CROSS = "cross-field consistency"
CHECK_SOURCES = "metric sources file"
@dataclass
class Report:
"""Problems found per check, in the order checks ran."""
checks: Dict[str, List[str]] = field(default_factory=dict)
notes: List[str] = field(default_factory=list)
row_count: int = 0
column_count: int = 0
def start(self, check: str) -> None:
self.checks.setdefault(check, [])
def add(self, check: str, message: str) -> None:
self.checks.setdefault(check, []).append(message)
@property
def problem_count(self) -> int:
return sum(len(messages) for messages in self.checks.values())
@property
def ok(self) -> bool:
return self.problem_count == 0
def read_climate_csv(path: Path) -> Tuple[List[str], List[List[str]]]:
"""Read the raw header and rows without DictReader's silent padding."""
with path.open("r", encoding="utf-8", newline="") as handle:
reader = csv.reader(handle)
headers = next(reader, [])
rows = [row for row in reader if row]
return headers, rows
def check_structure(
report: Report, headers: List[str], raw_rows: List[List[str]]
) -> Tuple[List[str], List[Dict[str, str]]]:
"""Check encoding, header names, and row widths; return rows as dicts."""
report.start(CHECK_FORMAT)
report.start(CHECK_COLUMNS)
if headers and headers[0].startswith("\ufeff"):
report.add(CHECK_FORMAT, "file starts with a UTF-8 BOM; the app would not find the first column")
headers = [headers[0].lstrip("\ufeff")] + headers[1:]
seen = set()
for name in headers:
if name in seen:
report.add(CHECK_COLUMNS, f"duplicate column {name!r}")
seen.add(name)
for name in EXPECTED_COLUMNS:
if name not in seen:
report.add(CHECK_COLUMNS, f"missing column {name!r}")
for name in headers:
if name not in EXPECTED_COLUMNS:
report.add(CHECK_COLUMNS, f"unexpected column {name!r} (add it to check_climate_data.py if intended)")
rows: List[Dict[str, str]] = []
for index, raw in enumerate(raw_rows):
if len(raw) != len(headers):
report.add(CHECK_FORMAT, f"line {index + 2}: {len(raw)} fields, expected {len(headers)}")
continue
rows.append(dict(zip(headers, raw)))
report.row_count = len(raw_rows)
report.column_count = len(headers)
return headers, rows
def check_counties(report: Report, rows: List[Dict[str, str]]) -> None:
"""Check FIPS format and uniqueness, names, and state/FIPS agreement."""
report.start(CHECK_COUNTIES)
seen = set()
prefix_states: Dict[str, set] = {}
state_prefixes: Dict[str, set] = {}
for row in rows:
fips = row.get("countyFips") or ""
if not FIPS_PATTERN.fullmatch(fips):
report.add(CHECK_COUNTIES, f"invalid countyFips {fips!r}")
continue
if fips in seen:
report.add(CHECK_COUNTIES, f"duplicate countyFips {fips}")
seen.add(fips)
if not (row.get("countyName") or "").strip():
report.add(CHECK_COUNTIES, f"{fips}: countyName is blank")
state = row.get("state") or ""
if not STATE_PATTERN.fullmatch(state):
report.add(CHECK_COUNTIES, f"{fips}: invalid state {state!r}")
continue
prefix_states.setdefault(fips[:2], set()).add(state)
state_prefixes.setdefault(state, set()).add(fips[:2])
for prefix, states in sorted(prefix_states.items()):
if len(states) > 1:
report.add(CHECK_COUNTIES, f"state FIPS {prefix} is labeled as {sorted(states)}")
for state, prefixes in sorted(state_prefixes.items()):
if len(prefixes) > 1:
report.add(CHECK_COUNTIES, f"state {state} spans FIPS prefixes {sorted(prefixes)}")
def _feature_fips(feature: dict) -> str:
"""Return the 5-digit county FIPS stored on a GeoJSON feature."""
props = feature.get("properties") or {}
raw = feature.get("id") or props.get("id") or str(props.get("GEO_ID") or "")[-5:]
text = str(raw or "").strip()
return text.zfill(5) if text.isdigit() else text
def check_geometry(report: Report, rows: List[Dict[str, str]], geojson_path: Path) -> None:
"""Check that CSV counties and map polygons match one-to-one."""
report.start(CHECK_GEOMETRY)
if not geojson_path.exists():
report.add(CHECK_GEOMETRY, f"county geometry file not found: {geojson_path}")
return
features = json.loads(geojson_path.read_text(encoding="utf-8")).get("features", [])
geo_fips = set()
for feature in features:
fips = _feature_fips(feature)
if not FIPS_PATTERN.fullmatch(fips):
report.add(CHECK_GEOMETRY, f"map feature without a valid county FIPS: {fips!r}")
continue
if fips in geo_fips:
report.add(CHECK_GEOMETRY, f"map has duplicate polygons for {fips}")
geo_fips.add(fips)
state_fips = str((feature.get("properties") or {}).get("STATE") or "").strip()
if state_fips and state_fips.zfill(2) != fips[:2]:
report.add(CHECK_GEOMETRY, f"map feature {fips} has STATE {state_fips}")
csv_fips = {row.get("countyFips") or "" for row in rows}
for fips in sorted(csv_fips - geo_fips):
report.add(CHECK_GEOMETRY, f"{fips} is in the CSV but has no map polygon")
for fips in sorted(geo_fips - csv_fips):
report.add(CHECK_GEOMETRY, f"{fips} has a map polygon but no CSV row")
def check_metrics(report: Report, headers: List[str], rows: List[Dict[str, str]]) -> None:
"""Check each metric's blanks, types, ranges, and categories."""
report.start(CHECK_BLANKS)
report.start(CHECK_VALUES)
for key, rule in METRIC_RULES.items():
if key not in headers:
continue
present: List[object] = []
for row in rows:
fips = row.get("countyFips") or "?"
state = row.get("state") or ""
value = row.get(key) or ""
if value == "":
if state not in rule.blank_states and fips not in rule.blank_fips:
report.add(CHECK_BLANKS, f"{fips} ({state}): {key} is blank")
continue
if value != value.strip():
report.add(CHECK_VALUES, f"{fips}: {key}={value!r} has surrounding whitespace")
continue
if rule.kind == "categorical":
if value in rule.categories:
present.append(value)
else:
report.add(CHECK_VALUES, f"{fips}: {key}={value!r} is not an allowed category")
continue
try:
number = float(value)
except ValueError:
report.add(CHECK_VALUES, f"{fips}: {key}={value!r} is not a number")
continue
if not math.isfinite(number):
report.add(CHECK_VALUES, f"{fips}: {key}={value!r} is not finite")
continue
if rule.integer and not number.is_integer():
report.add(CHECK_VALUES, f"{fips}: {key}={value} should be a whole number")
if not rule.minimum <= number <= rule.maximum:
report.add(CHECK_VALUES, f"{fips}: {key}={value} outside {rule.minimum:g}..{rule.maximum:g}")
present.append(number)
if len(rows) >= COLLAPSE_CHECK_MIN_ROWS and len(set(present)) < rule.min_distinct:
report.add(
CHECK_VALUES,
f"{key} has only {len(set(present))} distinct values; the column may have been overwritten",
)
def check_cross_fields(report: Report, rows: List[Dict[str, str]]) -> None:
"""Check relationships between columns within each row."""
report.start(CHECK_CROSS)
for row in rows:
fips = row.get("countyFips") or "?"
wet = row.get("wettestPrecipMonth") or ""
dry = row.get("driestPrecipMonth") or ""
if bool(wet) != bool(dry):
report.add(CHECK_CROSS, f"{fips}: only one of wettestPrecipMonth/driestPrecipMonth is set")
elif wet and wet == dry:
report.add(CHECK_CROSS, f"{fips}: wettest and driest month are both {wet}")
heat = row.get("humidHeatDays") or ""
heat_source = row.get("humidHeatSourceFips") or ""
if bool(heat) != bool(heat_source):
report.add(CHECK_CROSS, f"{fips}: humidHeatDays and humidHeatSourceFips must both be set or both blank")
if heat_source and not FIPS_PATTERN.fullmatch(heat_source):
report.add(CHECK_CROSS, f"{fips}: invalid humidHeatSourceFips {heat_source!r}")
elif heat_source and heat_source != fips and not (row.get("humidHeatFipsAdjustment") or "").strip():
report.add(CHECK_CROSS, f"{fips}: uses proxy county {heat_source} without a humidHeatFipsAdjustment note")
zone = row.get("koppenZone") or ""
primary = row.get("koppenPrimaryClass") or ""
secondary = row.get("koppenSecondaryClass") or ""
if zone == MIXED_KOPPEN_CLASS:
if not (primary and secondary):
report.add(CHECK_CROSS, f"{fips}: Mixed Koppen county needs koppenPrimaryClass and koppenSecondaryClass")
elif primary not in KOPPEN_CODES or secondary not in KOPPEN_CODES:
report.add(CHECK_CROSS, f"{fips}: invalid Koppen stripe classes {primary!r}/{secondary!r}")
elif primary == secondary:
report.add(CHECK_CROSS, f"{fips}: Koppen stripe classes are both {primary}")
elif primary or secondary:
report.add(CHECK_CROSS, f"{fips}: Koppen stripe classes are set but koppenZone is {zone or 'blank'}, not Mixed")
if not (row.get("source") or "").strip():
report.add(CHECK_CROSS, f"{fips}: source is blank")
def check_metric_sources(report: Report, path: Path) -> None:
"""Check that the metric metadata file parses and names real metrics."""
report.start(CHECK_SOURCES)
if not path.exists():
report.add(CHECK_SOURCES, f"{path.name} not found")
return
try:
data = json.loads(path.read_text(encoding="utf-8"))
except json.JSONDecodeError as exc:
report.add(CHECK_SOURCES, f"{path.name} is not valid JSON: {exc}")
return
if not isinstance(data, dict) or not isinstance(data.get("metrics"), dict):
report.add(CHECK_SOURCES, f'{path.name} must be a JSON object with a "metrics" object')
return
documented = set(data["metrics"])
for key in sorted(documented - set(METRIC_RULES)):
report.add(CHECK_SOURCES, f"{path.name} describes unknown metric {key!r}")
covered = len(documented & set(METRIC_RULES))
report.notes.append(f"{path.name} documents {covered} of {len(METRIC_RULES)} metrics.")
def run_checks(csv_path: Path, geojson_path: Path, metric_sources_path: Path) -> Report:
"""Run every check and return the combined report."""
report = Report()
headers, raw_rows = read_climate_csv(csv_path)
headers, rows = check_structure(report, headers, raw_rows)
check_counties(report, rows)
check_geometry(report, rows, geojson_path)
check_metrics(report, headers, rows)
check_cross_fields(report, rows)
check_metric_sources(report, metric_sources_path)
return report
def print_report(report: Report, csv_path: Path, max_examples: int) -> None:
"""Print a pass/fail line per check with example problems."""
print(f"Checked {csv_path}: {report.row_count} rows, {report.column_count} columns")
for check, messages in report.checks.items():
status = "PASS" if not messages else f"FAIL ({len(messages)})"
print(f" {status:<11}{check}")
for message in messages[:max_examples]:
print(f" - {message}")
if len(messages) > max_examples:
print(f" ... and {len(messages) - max_examples} more")
for note in report.notes:
print(f" note: {note}")
print("Result: PASS" if report.ok else f"Result: FAIL ({report.problem_count} problems)")
def parse_args() -> argparse.Namespace:
"""Define and parse command-line options for this checker."""
parser = argparse.ArgumentParser(description="Validate the browser app's county climate CSV.")
parser.add_argument("--csv", type=Path, default=DEFAULT_CLIMATE_CSV, help="Climate data CSV path.")
parser.add_argument("--geojson", type=Path, default=DEFAULT_COUNTIES_GEOJSON, help="County GeoJSON path.")
parser.add_argument("--metric-sources", type=Path, default=DEFAULT_METRIC_SOURCES, help="Metric metadata JSON path.")
parser.add_argument("--max-examples", type=int, default=10, help="Problems to print per failing check.")
return parser.parse_args()
def main() -> int:
args = parse_args()
if not args.csv.exists():
print(f"Climate data CSV not found: {args.csv}", file=sys.stderr)
return 2
report = run_checks(args.csv, args.geojson, args.metric_sources)
print_report(report, args.csv, args.max_examples)
return 0 if report.ok else 1
if __name__ == "__main__":
sys.exit(main())
+5
View File
@@ -0,0 +1,5 @@
"""Shared helpers imported by the county data pipeline scripts.
Modules here are not run directly. Scripts in the parent folder import them,
for example ``from common.counties import load_counties``.
"""
+131
View File
@@ -0,0 +1,131 @@
"""Load county polygons and normalize county identifiers."""
from __future__ import annotations
from pathlib import Path
import geopandas as gpd
DEFAULT_COUNTIES_GEOJSON_URL = "https://raw.githubusercontent.com/plotly/datasets/master/geojson-counties-fips.json"
STATE_FIPS_TO_ABBR = {
"01": "AL",
"02": "AK",
"04": "AZ",
"05": "AR",
"06": "CA",
"08": "CO",
"09": "CT",
"10": "DE",
"11": "DC",
"12": "FL",
"13": "GA",
"15": "HI",
"16": "ID",
"17": "IL",
"18": "IN",
"19": "IA",
"20": "KS",
"21": "KY",
"22": "LA",
"23": "ME",
"24": "MD",
"25": "MA",
"26": "MI",
"27": "MN",
"28": "MS",
"29": "MO",
"30": "MT",
"31": "NE",
"32": "NV",
"33": "NH",
"34": "NJ",
"35": "NM",
"36": "NY",
"37": "NC",
"38": "ND",
"39": "OH",
"40": "OK",
"41": "OR",
"42": "PA",
"44": "RI",
"45": "SC",
"46": "SD",
"47": "TN",
"48": "TX",
"49": "UT",
"50": "VT",
"51": "VA",
"53": "WA",
"54": "WV",
"55": "WI",
"56": "WY",
"60": "AS",
"66": "GU",
"69": "MP",
"72": "PR",
"78": "VI",
}
def normalize_fips(value: object, width: int) -> str:
"""Return a zero-padded FIPS code with the requested width."""
text = str(value).strip()
digits = "".join(ch for ch in text if ch.isdigit())
if not digits:
return ""
return digits.zfill(width)[-width:]
def load_counties(counties_geojson: Path) -> gpd.GeoDataFrame:
"""Load county polygons and normalize fields used downstream."""
if not counties_geojson.exists():
try:
print(
f"County GeoJSON not found at {counties_geojson}. "
f"Attempting download from {DEFAULT_COUNTIES_GEOJSON_URL}..."
)
gdf = gpd.read_file(DEFAULT_COUNTIES_GEOJSON_URL)
counties_geojson.parent.mkdir(parents=True, exist_ok=True)
# Cache the downloaded file for subsequent runs.
gdf.to_file(counties_geojson, driver="GeoJSON")
print(f"Downloaded and cached county GeoJSON to {counties_geojson}")
except Exception as exc:
raise FileNotFoundError(
f"County GeoJSON not found at {counties_geojson}, and download from "
f"{DEFAULT_COUNTIES_GEOJSON_URL} failed. Download the file manually "
"and rerun with --counties-geojson pointing to it."
) from exc
gdf = gpd.read_file(counties_geojson)
if gdf.crs is None:
gdf = gdf.set_crs("EPSG:4326")
else:
gdf = gdf.to_crs("EPSG:4326")
feature_id = None
if "id" in gdf.columns:
feature_id = gdf["id"]
elif "GEOID" in gdf.columns:
feature_id = gdf["GEOID"]
elif "GEOID10" in gdf.columns:
feature_id = gdf["GEOID10"]
elif "fips" in gdf.columns:
feature_id = gdf["fips"]
else:
raise ValueError("Unable to locate county FIPS identifier column in county polygons.")
gdf["county_fips"] = feature_id.map(lambda value: normalize_fips(value, 5))
gdf = gdf[gdf["county_fips"] != ""].copy()
if "NAME" in gdf.columns:
gdf["county_name"] = gdf["NAME"].fillna("").astype(str).str.strip()
elif "name" in gdf.columns:
gdf["county_name"] = gdf["name"].fillna("").astype(str).str.strip()
else:
gdf["county_name"] = gdf["county_fips"].map(lambda value: f"County {value}")
gdf["state_fips"] = gdf["county_fips"].str.slice(0, 2)
gdf["state"] = gdf["state_fips"].map(lambda code: STATE_FIPS_TO_ABBR.get(code, f"S{code}"))
gdf = gdf.sort_values("county_fips").reset_index(drop=True)
return gdf
+128
View File
@@ -0,0 +1,128 @@
"""Shared county zonal statistics for raster-based metrics.
Area weighting estimates how much of each raster cell lies inside a county by
rasterizing the county on a finer grid of sub-cells, then scales each cell by
the cosine of its latitude so cells count by their true surface area. Counties
that cross the 180th meridian are split so each side is read from its own small
raster window.
"""
from __future__ import annotations
import math
from typing import Dict, List
import numpy as np
import shapely
from affine import Affine
from rasterio.features import rasterize
from rasterio.windows import Window, from_bounds
from shapely.affinity import translate
from shapely.geometry import box, mapping
from shapely.geometry.base import BaseGeometry
from shapely.ops import unary_union
DEFAULT_SUBCELLS = 16
# Largest sub-cell grid rasterized for one county piece; bigger pieces use a coarser grid.
SUBCELL_BUDGET = 80_000_000
def split_at_antimeridian(geometry: BaseGeometry) -> List[BaseGeometry]:
"""Split a lon/lat geometry into pieces that each stay on one side of 180 degrees.
A county such as Aleutians West, AK has islands at both +179 and -179
degrees longitude. Its bounding box then spans nearly the whole globe, so
each side is returned as its own piece.
"""
minx, _, maxx, _ = geometry.bounds
if maxx - minx <= 180.0:
return [geometry]
positive: List[BaseGeometry] = []
negative: List[BaseGeometry] = []
for part in getattr(geometry, "geoms", [geometry]):
part_minx, _, part_maxx, _ = part.bounds
if part_maxx - part_minx > 180.0:
# One outline crosses the line: unwrap to 0..360, cut at 180, rewrap.
unwrapped = shapely.transform(
part,
lambda xy: np.column_stack((np.where(xy[:, 0] < 0, xy[:, 0] + 360.0, xy[:, 0]), xy[:, 1])),
)
positive.append(unwrapped.intersection(box(0.0, -90.0, 180.0, 90.0)))
negative.append(translate(unwrapped.intersection(box(180.0, -90.0, 360.0, 90.0)), xoff=-360.0))
elif part_minx >= 0:
positive.append(part)
else:
negative.append(part)
pieces = [unary_union(group) for group in (positive, negative) if group]
return [piece for piece in pieces if not piece.is_empty]
def geometry_window(source, geometry: BaseGeometry) -> Window:
"""Return the raster window covering a geometry, padded by one cell on each side."""
window = from_bounds(*geometry.bounds, transform=source.transform)
col_start = math.floor(window.col_off) - 1
row_start = math.floor(window.row_off) - 1
col_stop = math.ceil(window.col_off + window.width) + 1
row_stop = math.ceil(window.row_off + window.height) + 1
padded = Window(col_start, row_start, col_stop - col_start, row_stop - row_start)
return padded.intersection(Window(0, 0, source.width, source.height))
def subcells_for(shape: tuple[int, int], requested: int) -> int:
"""Return the finest sub-cell count, up to the request, that fits the budget."""
rows, cols = shape
subcells = requested
while subcells > 1 and rows * cols * subcells * subcells > SUBCELL_BUDGET:
subcells //= 2
return subcells
def cell_coverage_fractions(
shape: tuple[int, int], transform: Affine, geometry: BaseGeometry, subcells: int
) -> np.ndarray:
"""Estimate the fraction of each raster cell covered by a geometry."""
rows, cols = shape
fine = rasterize(
[mapping(geometry)],
out_shape=(rows * subcells, cols * subcells),
transform=transform * Affine.scale(1.0 / subcells),
fill=0,
default_value=1,
dtype="uint8",
)
return fine.reshape(rows, subcells, cols, subcells).mean(axis=(1, 3))
def area_weighted_class_weights(
source,
geometry: BaseGeometry,
*,
geographic: bool = True,
subcells: int = DEFAULT_SUBCELLS,
) -> Dict[int, float]:
"""Return the area inside a geometry covered by each value of a categorical raster.
Weights are relative surface areas: the fraction of each cell inside the
geometry, times cos(latitude) for geographic rasters. Cells equal to 0 or
the raster's nodata value are excluded.
"""
weights: Dict[int, float] = {}
pieces = split_at_antimeridian(geometry) if geographic else [geometry]
for piece in pieces:
window = geometry_window(source, piece)
values = source.read(1, window=window, masked=True).filled(0)
if source.nodata is not None:
values = np.where(values == source.nodata, 0, values)
transform = source.window_transform(window)
cell_weights = cell_coverage_fractions(values.shape, transform, piece, subcells_for(values.shape, subcells))
if geographic:
row_lat = transform.f + (np.arange(values.shape[0]) + 0.5) * transform.e
cell_weights = cell_weights * np.cos(np.radians(row_lat))[:, None]
counted = (values != 0) & (cell_weights > 0)
for value in np.unique(values[counted]):
code = int(value)
weights[code] = weights.get(code, 0.0) + float(cell_weights[counted & (values == value)].sum())
return weights
+64
View File
@@ -0,0 +1,64 @@
"""Koppen-Geiger raster codes and the Beck et al. legend loader."""
from __future__ import annotations
import re
from pathlib import Path
from typing import Dict
# Beck et al legend key is expected as text file, but this default handles common codes.
DEFAULT_KOPPEN_CODE_MAP = {
1: "Af",
2: "Am",
3: "Aw",
4: "BWh",
5: "BWk",
6: "BSh",
7: "BSk",
8: "Csa",
9: "Csb",
10: "Csc",
11: "Cwa",
12: "Cwb",
13: "Cwc",
14: "Cfa",
15: "Cfb",
16: "Cfc",
17: "Dsa",
18: "Dsb",
19: "Dsc",
20: "Dsd",
21: "Dwa",
22: "Dwb",
23: "Dwc",
24: "Dwd",
25: "Dfa",
26: "Dfb",
27: "Dfc",
28: "Dfd",
29: "ET",
30: "EF",
}
def load_koppen_legend(legend_path: Path | None) -> Dict[int, str]:
"""Load Koppen raster codes, using defaults when no legend exists."""
if legend_path is None:
return DEFAULT_KOPPEN_CODE_MAP
mapping: Dict[int, str] = {}
for line in legend_path.read_text(encoding="utf-8").splitlines():
text = line.strip()
if not text or text.startswith("#"):
continue
# Handles patterns like:
# "1: Af ..." or "1 = Af" or "1 Af"
match = re.match(r"^(\d+)\s*[:=]?\s*([A-Za-z]{2,3})\b", text)
if not match:
continue
key = int(match.group(1))
value = match.group(2)
mapping[key] = value
return mapping if mapping else DEFAULT_KOPPEN_CODE_MAP
+15 -2
View File
@@ -12,13 +12,24 @@ The browser blocks `fetch("data/climate-data.csv")` when `index.html` is opened
Then open [http://localhost:8000/](http://localhost:8000/). This keeps the app CSV-only while allowing the map and filters to load normally.
## Source 1: Koppen-Geiger classes (`koppenZone`)
## Source 1: Koppen-Geiger classes (`koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`)
- Dataset: Beck et al. updated 1-km Koppen-Geiger climate classes (historical + future windows)
- Landing page: [https://www.gloh2o.org/koppen/](https://www.gloh2o.org/koppen/)
- Primary paper for updated release: [https://www.nature.com/articles/s41597-023-02549-6](https://www.nature.com/articles/s41597-023-02549-6)
- Coverage: 1901-2099 (use historical 1991-2020 layer for this project to align with NOAA baselines)
- License shown on dataset page: CC BY 4.0
- Local files: `data/koppen_geiger_tif/1991_2020/koppen_geiger_0p00833333.tif` and `data/koppen_geiger_tif/legend.txt`
Build the county metric and apply it to the app CSV:
```powershell
.venv\Scripts\python.exe scripts\build_county_koppen_metric.py
.venv\Scripts\python.exe scripts\apply_koppen_metric_to_climate_data.py --dry-run
.venv\Scripts\python.exe scripts\apply_koppen_metric_to_climate_data.py
```
The builder writes `data/metrics/koppen.csv` with each county's class, top and runner-up classes, and their area-weighted shares. The apply step writes `koppenZone`, `koppenPrimaryClass`, and `koppenSecondaryClass`; `--dry-run` reports the changes without writing. Run it after `build_county_climate_data.py`, which still writes an older largest-share `koppenZone`. The classification rule is documented in [`docs/filter-calculations.md`](../docs/filter-calculations.md) §1.
## Source 2: NOAA 1991-2020 gridded normals (`avgTempF`, `annualPrecipIn`, `seasonalityIndex`, previous `extremeDays`)
@@ -255,7 +266,8 @@ Then update the app CSV. Polygon archive GHI is used first; representative-point
## Metric definitions in generated output
- `koppenZone`: majority class within county polygon from Koppen raster.
- `koppenZone`: the county's predominant Koppen-Geiger class, meaning the class covering at least 50% of the county's land area and leading the runner-up by at least 5 percentage points; otherwise `Mixed`. Shares are area-weighted, with ocean and no-data cells excluded.
- `koppenPrimaryClass` / `koppenSecondaryClass`: for Mixed counties only, the top and runner-up classes, drawn as stripes on the map; blank for predominant counties.
- `avgTempF`: mean of 12 monthly county mean temperatures, converted C -> F.
- `annualPrecipIn`: sum of 12 monthly county mean precipitation totals, converted mm -> inches.
- `seasonalityIndex`: coefficient of variation of monthly precipitation totals, scaled to 0-100.
@@ -295,6 +307,7 @@ python scripts/build_county_climate_data.py `
Notes:
- This writes only the base columns and an older largest-share `koppenZone`. Do not run it over the live `data/climate-data.csv`: it would drop the columns added by later stages. After a full rebuild, run the Köppen build and apply steps (Source 1) and the enrichment stages.
- This computes all counties in your geometry file, not just the sample records.
- For counties outside CONUS coverage in NOAA gridded files, fallback values are applied by the script when no valid grid values intersect.
- For physically-based daily `extremeDays`, provide true daily grids and set `--extreme-days-mode require-daily`.