Complete Köppen-Geiger filter review with Mixed climate class
Classify each county by area-weighted Köppen class shares: a county is predominantly its top class when that class covers at least 50% of its land and leads the runner-up by at least 5 percentage points; otherwise it is Mixed (133 of 3,143 counties in the 50 states and DC). - Add build_county_koppen_metric.py (writes data/metrics/koppen.csv) and apply_koppen_metric_to_climate_data.py (writes koppenZone plus koppenPrimaryClass/koppenSecondaryClass for Mixed counties). - Move shared helpers into scripts/common/ (county loading, Köppen legend, area-weighted raster shares); fix the 180th-meridian raster window for Aleutians West. - Add check_climate_data.py to validate the app CSV. - Draw Mixed counties in app.js as diagonal stripes of their top two classes, fixed to the ground and following the map at every zoom, with a crossfade only when the stripe size changes. Filtering a class also matches Mixed counties where it is primary or secondary. - Document the rule, display, and pipeline plan in docs/ and update the README and data-source notes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,138 @@
|
||||
"""Apply the county Koppen-Geiger metric to the app CSV.
|
||||
|
||||
Replaces the koppenZone column of data/climate-data.csv with the values in
|
||||
data/metrics/koppen.csv and writes koppenPrimaryClass and koppenSecondaryClass:
|
||||
the two stripe classes the app draws for Mixed counties. Both are blank for
|
||||
predominant counties. The two columns are added after koppenZone if missing.
|
||||
Every other column, and the column order, is left unchanged. Use --dry-run to
|
||||
report the changes without writing.
|
||||
|
||||
Run:
|
||||
.venv\\Scripts\\python.exe scripts\\apply_koppen_metric_to_climate_data.py --dry-run
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import csv
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
from typing import Dict, List, Tuple
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_CLIMATE_DATA = REPO_ROOT / "data" / "climate-data.csv"
|
||||
DEFAULT_KOPPEN_METRIC = REPO_ROOT / "data" / "metrics" / "koppen.csv"
|
||||
|
||||
ZONE_FIELD = "koppenZone"
|
||||
PRIMARY_FIELD = "koppenPrimaryClass"
|
||||
SECONDARY_FIELD = "koppenSecondaryClass"
|
||||
STRIPE_FIELDS = [PRIMARY_FIELD, SECONDARY_FIELD]
|
||||
MIXED_CLASS = "Mixed"
|
||||
METRIC_FIELDS = ["countyFips", ZONE_FIELD, "koppenTopClass", "koppenSecondClass"]
|
||||
|
||||
# (countyFips, column, old value, new value)
|
||||
Change = Tuple[str, str, str, str]
|
||||
|
||||
|
||||
def read_csv_rows(path: Path) -> Tuple[List[str], List[dict]]:
|
||||
"""Read a CSV while preserving the source field order."""
|
||||
with path.open("r", encoding="utf-8-sig", newline="") as csv_file:
|
||||
reader = csv.DictReader(csv_file)
|
||||
if reader.fieldnames is None:
|
||||
raise ValueError(f"{path} has no CSV header.")
|
||||
return list(reader.fieldnames), list(reader)
|
||||
|
||||
|
||||
def fieldnames_with_stripe_columns(fields: List[str]) -> List[str]:
|
||||
"""Place the stripe-class columns directly after koppenZone."""
|
||||
base = [field for field in fields if field not in STRIPE_FIELDS]
|
||||
insert_at = base.index(ZONE_FIELD) + 1
|
||||
return base[:insert_at] + STRIPE_FIELDS + base[insert_at:]
|
||||
|
||||
|
||||
def load_koppen_values(koppen_metric: Path) -> Dict[str, Dict[str, str]]:
|
||||
"""Return koppenZone and the two stripe classes for each county in the metric file."""
|
||||
fields, rows = read_csv_rows(koppen_metric)
|
||||
missing_fields = [field for field in METRIC_FIELDS if field not in fields]
|
||||
if missing_fields:
|
||||
raise ValueError(f"{koppen_metric} is missing columns {missing_fields}.")
|
||||
|
||||
values: Dict[str, Dict[str, str]] = {}
|
||||
for row in rows:
|
||||
is_mixed = row[ZONE_FIELD] == MIXED_CLASS
|
||||
primary = row["koppenTopClass"] if is_mixed else ""
|
||||
secondary = row["koppenSecondClass"] if is_mixed else ""
|
||||
if is_mixed and not (primary and secondary):
|
||||
raise ValueError(f"{row['countyFips']} is Mixed but has no top or second class in {koppen_metric}.")
|
||||
values[row["countyFips"]] = {ZONE_FIELD: row[ZONE_FIELD], PRIMARY_FIELD: primary, SECONDARY_FIELD: secondary}
|
||||
return values
|
||||
|
||||
|
||||
def apply_koppen_metric(
|
||||
climate_data: Path, koppen_metric: Path, out: Path, dry_run: bool = False
|
||||
) -> Tuple[List[Change], List[str]]:
|
||||
"""Update the Koppen columns; return the value changes and any columns that were added."""
|
||||
fields, rows = read_csv_rows(climate_data)
|
||||
if ZONE_FIELD not in fields:
|
||||
raise ValueError(f"{climate_data} has no {ZONE_FIELD} column.")
|
||||
lookup = load_koppen_values(koppen_metric)
|
||||
|
||||
missing = [row["countyFips"] for row in rows if row["countyFips"] not in lookup]
|
||||
if missing:
|
||||
raise ValueError(f"{len(missing)} counties are missing from {koppen_metric}, e.g. {missing[:5]}")
|
||||
|
||||
added_columns = [field for field in STRIPE_FIELDS if field not in fields]
|
||||
changes: List[Change] = []
|
||||
for row in rows:
|
||||
new_values = lookup[row["countyFips"]]
|
||||
for field in [ZONE_FIELD, *STRIPE_FIELDS]:
|
||||
old_value = row.get(field) or ""
|
||||
if old_value != new_values[field]:
|
||||
changes.append((row["countyFips"], field, old_value, new_values[field]))
|
||||
row[field] = new_values[field]
|
||||
|
||||
if not dry_run:
|
||||
with out.open("w", encoding="utf-8", newline="") as csv_file:
|
||||
writer = csv.DictWriter(csv_file, fieldnames=fieldnames_with_stripe_columns(fields))
|
||||
writer.writeheader()
|
||||
writer.writerows(rows)
|
||||
return changes, added_columns
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
"""Define and parse command-line options for this apply step."""
|
||||
parser = argparse.ArgumentParser(description="Apply the county Koppen metric to the app CSV.")
|
||||
parser.add_argument("--climate-data", type=Path, default=DEFAULT_CLIMATE_DATA, help="App climate CSV to update.")
|
||||
parser.add_argument("--koppen-metric", type=Path, default=DEFAULT_KOPPEN_METRIC, help="Koppen metric CSV.")
|
||||
parser.add_argument("--out", type=Path, default=None, help="Output path; defaults to updating --climate-data in place.")
|
||||
parser.add_argument("--dry-run", action="store_true", help="Report changes without writing.")
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
changes, added_columns = apply_koppen_metric(
|
||||
args.climate_data, args.koppen_metric, args.out or args.climate_data, args.dry_run
|
||||
)
|
||||
verb = "would" if args.dry_run else "did"
|
||||
|
||||
if added_columns:
|
||||
print(f"Columns added after {ZONE_FIELD} ({verb} write): {', '.join(added_columns)}")
|
||||
zone_changes = [change for change in changes if change[1] == ZONE_FIELD]
|
||||
kinds = Counter(
|
||||
"to Mixed" if new == MIXED_CLASS else "to blank" if not new else "to another class"
|
||||
for _, _, _, new in zone_changes
|
||||
)
|
||||
stripe_values = sum(1 for _, field, _, new in changes if field in STRIPE_FIELDS and new)
|
||||
print(f"{len(zone_changes)} {ZONE_FIELD} values {'would change' if args.dry_run else 'changed'}: {dict(kinds)}")
|
||||
print(f"{stripe_values} stripe-class values {'would be set' if args.dry_run else 'set'}.")
|
||||
for fips, _, old, new in zone_changes[:10]:
|
||||
print(f" {fips}: {old or '(blank)'} -> {new or '(blank)'}")
|
||||
if len(zone_changes) > 10:
|
||||
print(f" ... and {len(zone_changes) - 10} more")
|
||||
if args.dry_run:
|
||||
print("Dry run: nothing was written.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -18,10 +18,10 @@ from build_county_climate_data import (
|
||||
MONTH_NAMES,
|
||||
_as_monthly_climatology,
|
||||
_extract_grid_2d,
|
||||
_load_counties,
|
||||
_select_data_var,
|
||||
_zonal_mean,
|
||||
)
|
||||
from common.counties import load_counties
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_CLIMATE_DATA = REPO_ROOT / "data" / "climate-data.csv"
|
||||
@@ -61,7 +61,7 @@ def build_precip_month_lookup(
|
||||
climatology_end_year: int,
|
||||
) -> dict[str, tuple[str, str]]:
|
||||
"""Calculate wettest and driest precipitation month for each county."""
|
||||
counties = _load_counties(counties_geojson)
|
||||
counties = load_counties(counties_geojson)
|
||||
monthly_prcp = xr.open_dataset(monthly_prcp_nc, decode_times=True)
|
||||
try:
|
||||
prcp_var = _select_data_var(monthly_prcp, "mlyprcp_norm")
|
||||
|
||||
@@ -31,102 +31,11 @@ import rasterio
|
||||
import xarray as xr
|
||||
from affine import Affine
|
||||
from rasterio.features import geometry_mask
|
||||
from shapely.geometry.base import BaseGeometry
|
||||
|
||||
DEFAULT_COUNTIES_GEOJSON_URL = "https://raw.githubusercontent.com/plotly/datasets/master/geojson-counties-fips.json"
|
||||
|
||||
|
||||
STATE_FIPS_TO_ABBR = {
|
||||
"01": "AL",
|
||||
"02": "AK",
|
||||
"04": "AZ",
|
||||
"05": "AR",
|
||||
"06": "CA",
|
||||
"08": "CO",
|
||||
"09": "CT",
|
||||
"10": "DE",
|
||||
"11": "DC",
|
||||
"12": "FL",
|
||||
"13": "GA",
|
||||
"15": "HI",
|
||||
"16": "ID",
|
||||
"17": "IL",
|
||||
"18": "IN",
|
||||
"19": "IA",
|
||||
"20": "KS",
|
||||
"21": "KY",
|
||||
"22": "LA",
|
||||
"23": "ME",
|
||||
"24": "MD",
|
||||
"25": "MA",
|
||||
"26": "MI",
|
||||
"27": "MN",
|
||||
"28": "MS",
|
||||
"29": "MO",
|
||||
"30": "MT",
|
||||
"31": "NE",
|
||||
"32": "NV",
|
||||
"33": "NH",
|
||||
"34": "NJ",
|
||||
"35": "NM",
|
||||
"36": "NY",
|
||||
"37": "NC",
|
||||
"38": "ND",
|
||||
"39": "OH",
|
||||
"40": "OK",
|
||||
"41": "OR",
|
||||
"42": "PA",
|
||||
"44": "RI",
|
||||
"45": "SC",
|
||||
"46": "SD",
|
||||
"47": "TN",
|
||||
"48": "TX",
|
||||
"49": "UT",
|
||||
"50": "VT",
|
||||
"51": "VA",
|
||||
"53": "WA",
|
||||
"54": "WV",
|
||||
"55": "WI",
|
||||
"56": "WY",
|
||||
"60": "AS",
|
||||
"66": "GU",
|
||||
"69": "MP",
|
||||
"72": "PR",
|
||||
"78": "VI",
|
||||
}
|
||||
|
||||
# Beck et al legend key is expected as text file, but this default handles common codes.
|
||||
DEFAULT_KOPPEN_CODE_MAP = {
|
||||
1: "Af",
|
||||
2: "Am",
|
||||
3: "Aw",
|
||||
4: "BWh",
|
||||
5: "BWk",
|
||||
6: "BSh",
|
||||
7: "BSk",
|
||||
8: "Csa",
|
||||
9: "Csb",
|
||||
10: "Csc",
|
||||
11: "Cwa",
|
||||
12: "Cwb",
|
||||
13: "Cwc",
|
||||
14: "Cfa",
|
||||
15: "Cfb",
|
||||
16: "Cfc",
|
||||
17: "Dsa",
|
||||
18: "Dsb",
|
||||
19: "Dsc",
|
||||
20: "Dsd",
|
||||
21: "Dwa",
|
||||
22: "Dwb",
|
||||
23: "Dwc",
|
||||
24: "Dwd",
|
||||
25: "Dfa",
|
||||
26: "Dfb",
|
||||
27: "Dfc",
|
||||
28: "Dfd",
|
||||
29: "ET",
|
||||
30: "EF",
|
||||
}
|
||||
from common.counties import load_counties, normalize_fips
|
||||
from common.county_zonal_stats import geometry_window, split_at_antimeridian
|
||||
from common.koppen_legend import load_koppen_legend
|
||||
|
||||
MONTH_NAMES = [
|
||||
"January",
|
||||
@@ -144,94 +53,6 @@ MONTH_NAMES = [
|
||||
]
|
||||
|
||||
|
||||
def _normalize_fips(value: object, width: int) -> str:
|
||||
"""Return a zero-padded FIPS code with the requested width."""
|
||||
text = str(value).strip()
|
||||
digits = "".join(ch for ch in text if ch.isdigit())
|
||||
if not digits:
|
||||
return ""
|
||||
return digits.zfill(width)[-width:]
|
||||
|
||||
|
||||
def _load_counties(counties_geojson: Path) -> gpd.GeoDataFrame:
|
||||
"""Load county polygons and normalize fields used downstream."""
|
||||
if not counties_geojson.exists():
|
||||
try:
|
||||
print(
|
||||
f"County GeoJSON not found at {counties_geojson}. "
|
||||
f"Attempting download from {DEFAULT_COUNTIES_GEOJSON_URL}..."
|
||||
)
|
||||
gdf = gpd.read_file(DEFAULT_COUNTIES_GEOJSON_URL)
|
||||
counties_geojson.parent.mkdir(parents=True, exist_ok=True)
|
||||
# Cache the downloaded file for subsequent runs.
|
||||
gdf.to_file(counties_geojson, driver="GeoJSON")
|
||||
print(f"Downloaded and cached county GeoJSON to {counties_geojson}")
|
||||
except Exception as exc:
|
||||
raise FileNotFoundError(
|
||||
f"County GeoJSON not found at {counties_geojson}, and download from "
|
||||
f"{DEFAULT_COUNTIES_GEOJSON_URL} failed. Download the file manually "
|
||||
"and rerun with --counties-geojson pointing to it."
|
||||
) from exc
|
||||
|
||||
gdf = gpd.read_file(counties_geojson)
|
||||
if gdf.crs is None:
|
||||
gdf = gdf.set_crs("EPSG:4326")
|
||||
else:
|
||||
gdf = gdf.to_crs("EPSG:4326")
|
||||
|
||||
feature_id = None
|
||||
if "id" in gdf.columns:
|
||||
feature_id = gdf["id"]
|
||||
elif "GEOID" in gdf.columns:
|
||||
feature_id = gdf["GEOID"]
|
||||
elif "GEOID10" in gdf.columns:
|
||||
feature_id = gdf["GEOID10"]
|
||||
elif "fips" in gdf.columns:
|
||||
feature_id = gdf["fips"]
|
||||
else:
|
||||
raise ValueError("Unable to locate county FIPS identifier column in county polygons.")
|
||||
|
||||
gdf["county_fips"] = feature_id.map(lambda value: _normalize_fips(value, 5))
|
||||
gdf = gdf[gdf["county_fips"] != ""].copy()
|
||||
|
||||
if "NAME" in gdf.columns:
|
||||
gdf["county_name"] = gdf["NAME"].fillna("").astype(str).str.strip()
|
||||
elif "name" in gdf.columns:
|
||||
gdf["county_name"] = gdf["name"].fillna("").astype(str).str.strip()
|
||||
else:
|
||||
gdf["county_name"] = gdf["county_fips"].map(lambda value: f"County {value}")
|
||||
|
||||
gdf["state_fips"] = gdf["county_fips"].str.slice(0, 2)
|
||||
gdf["state"] = gdf["state_fips"].map(lambda code: STATE_FIPS_TO_ABBR.get(code, f"S{code}"))
|
||||
gdf = gdf.sort_values("county_fips").reset_index(drop=True)
|
||||
return gdf
|
||||
|
||||
|
||||
def _load_koppen_legend(legend_path: Path | None) -> Dict[int, str]:
|
||||
"""Load Koppen raster codes, using defaults when no legend exists."""
|
||||
if legend_path is None:
|
||||
return DEFAULT_KOPPEN_CODE_MAP
|
||||
|
||||
mapping: Dict[int, str] = {}
|
||||
for line in legend_path.read_text(encoding="utf-8").splitlines():
|
||||
text = line.strip()
|
||||
if not text or text.startswith("#"):
|
||||
continue
|
||||
# Handles patterns like:
|
||||
# "1: Af ..." or "1 = Af" or "1 Af"
|
||||
import re
|
||||
|
||||
match = re.match(r"^(\d+)\s*[:=]?\s*([A-Za-z]{2,3})\b", text)
|
||||
if not match:
|
||||
continue
|
||||
|
||||
key = int(match.group(1))
|
||||
value = match.group(2)
|
||||
mapping[key] = value
|
||||
|
||||
return mapping if mapping else DEFAULT_KOPPEN_CODE_MAP
|
||||
|
||||
|
||||
def _select_data_var(dataset: xr.Dataset, preferred: str) -> str:
|
||||
"""Choose the best matching climate variable from a dataset."""
|
||||
if preferred in dataset.data_vars:
|
||||
@@ -283,7 +104,7 @@ def _load_solar_ghi_csv(solar_ghi_csv: Path, counties: gpd.GeoDataFrame) -> List
|
||||
|
||||
solar_by_fips: Dict[str, float] = {}
|
||||
for row in reader:
|
||||
county_fips = _normalize_fips(row.get(fips_field, ""), 5)
|
||||
county_fips = normalize_fips(row.get(fips_field, ""), 5)
|
||||
raw_value = str(row.get("meanDailyGlobalHorizontalRadiationKwhM2Day", "")).strip()
|
||||
if not county_fips or not raw_value:
|
||||
continue
|
||||
@@ -414,6 +235,26 @@ def _zonal_mean_raster(raster_path: Path, counties: gpd.GeoDataFrame) -> List[fl
|
||||
return _zonal_mean(values, source.transform, raster_counties)
|
||||
|
||||
|
||||
def _touched_raster_values(source, geometry: BaseGeometry, split_antimeridian: bool) -> np.ndarray:
|
||||
"""Read every raster cell a geometry touches, using a small window per piece."""
|
||||
pieces = split_at_antimeridian(geometry) if split_antimeridian else [geometry]
|
||||
selected: List[np.ndarray] = []
|
||||
for piece in pieces:
|
||||
window = geometry_window(source, piece)
|
||||
values = source.read(1, window=window, masked=True).filled(0)
|
||||
if source.nodata is not None:
|
||||
values = np.where(values == source.nodata, 0, values)
|
||||
mask = geometry_mask(
|
||||
[piece.__geo_interface__],
|
||||
out_shape=values.shape,
|
||||
transform=source.window_transform(window),
|
||||
invert=True,
|
||||
all_touched=True,
|
||||
)
|
||||
selected.append(values[mask])
|
||||
return np.concatenate(selected) if selected else np.array([], dtype=np.int64)
|
||||
|
||||
|
||||
def _zonal_majority_class(koppen_raster: Path, counties: gpd.GeoDataFrame, code_map: Dict[int, str]) -> List[str]:
|
||||
"""Assign each county its most common Koppen-Geiger class."""
|
||||
classes: List[str] = []
|
||||
@@ -421,21 +262,11 @@ def _zonal_majority_class(koppen_raster: Path, counties: gpd.GeoDataFrame, code_
|
||||
raster_counties = counties
|
||||
if source.crs is not None and counties.crs is not None and counties.crs != source.crs:
|
||||
raster_counties = counties.to_crs(source.crs)
|
||||
|
||||
data = source.read(1, masked=True)
|
||||
values = np.asarray(data.filled(0))
|
||||
if source.nodata is not None:
|
||||
values = np.where(values == source.nodata, 0, values)
|
||||
# The 180-degree split only makes sense for longitude/latitude rasters.
|
||||
split_antimeridian = source.crs is None or source.crs.is_geographic
|
||||
|
||||
for geometry in raster_counties.geometry:
|
||||
mask = geometry_mask(
|
||||
[geometry.__geo_interface__],
|
||||
out_shape=values.shape,
|
||||
transform=source.transform,
|
||||
invert=True,
|
||||
all_touched=True,
|
||||
)
|
||||
selected = values[mask]
|
||||
selected = _touched_raster_values(source, geometry, split_antimeridian)
|
||||
selected = selected[selected != 0]
|
||||
if selected.size == 0:
|
||||
classes.append("Cfa")
|
||||
@@ -545,9 +376,9 @@ def build_county_records(
|
||||
solar_ghi_csv: Path | None,
|
||||
) -> Dict[str, dict]:
|
||||
"""Build county climate records consumed by the web app."""
|
||||
counties = _load_counties(counties_geojson)
|
||||
counties = load_counties(counties_geojson)
|
||||
|
||||
koppen_classes = _zonal_majority_class(koppen_raster, counties, _load_koppen_legend(koppen_legend))
|
||||
koppen_classes = _zonal_majority_class(koppen_raster, counties, load_koppen_legend(koppen_legend))
|
||||
|
||||
monthly_tavg = xr.open_dataset(monthly_tavg_nc, decode_times=True)
|
||||
monthly_prcp = xr.open_dataset(monthly_prcp_nc, decode_times=True)
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
"""Build the county Koppen-Geiger metric file from area-weighted class shares.
|
||||
|
||||
Writes data/metrics/koppen.csv. A county is predominantly its top class when
|
||||
that class covers at least 50% of the county's land and leads the runner-up by
|
||||
at least 5 percentage points; otherwise it is "Mixed". Counties with no valid
|
||||
raster cells are left blank. See docs/filter-calculations.md, section 1.
|
||||
|
||||
Run:
|
||||
.venv\\Scripts\\python.exe scripts\\build_county_koppen_metric.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import csv
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
from typing import Dict, List, Tuple
|
||||
|
||||
import geopandas as gpd
|
||||
import rasterio
|
||||
|
||||
from common.counties import load_counties
|
||||
from common.county_zonal_stats import DEFAULT_SUBCELLS, area_weighted_class_weights
|
||||
from common.koppen_legend import load_koppen_legend
|
||||
|
||||
PROJECT_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_COUNTIES_GEOJSON = PROJECT_ROOT / "data" / "geojson-counties-fips.json"
|
||||
DEFAULT_KOPPEN_RASTER = PROJECT_ROOT / "data" / "koppen_geiger_tif" / "1991_2020" / "koppen_geiger_0p00833333.tif"
|
||||
DEFAULT_KOPPEN_LEGEND = PROJECT_ROOT / "data" / "koppen_geiger_tif" / "legend.txt"
|
||||
DEFAULT_OUT = PROJECT_ROOT / "data" / "metrics" / "koppen.csv"
|
||||
|
||||
PREDOMINANT_MIN_SHARE = 0.50
|
||||
PREDOMINANT_MIN_GAP = 0.05
|
||||
MIXED_CLASS = "Mixed"
|
||||
# Keeps shares that sit exactly on a cutoff from failing on floating-point error.
|
||||
RULE_TOLERANCE = 1e-9
|
||||
|
||||
FIELDS = [
|
||||
"countyFips",
|
||||
"countyName",
|
||||
"state",
|
||||
"koppenZone",
|
||||
"koppenTopClass",
|
||||
"koppenTopShare",
|
||||
"koppenSecondClass",
|
||||
"koppenSecondShare",
|
||||
]
|
||||
|
||||
|
||||
def rank_class_shares(weights: Dict[int, float], code_map: Dict[int, str]) -> List[Tuple[str, float]]:
|
||||
"""Convert per-code area weights to class shares, largest first.
|
||||
|
||||
Exact ties go to the smaller raster code, matching the previous build.
|
||||
"""
|
||||
total = sum(weights.values())
|
||||
if total <= 0:
|
||||
return []
|
||||
unknown = sorted(code for code in weights if code not in code_map)
|
||||
if unknown:
|
||||
raise ValueError(f"Raster codes {unknown} are not in the Koppen legend.")
|
||||
ranked = sorted(weights.items(), key=lambda item: (-item[1], item[0]))
|
||||
return [(code_map[code], weight / total) for code, weight in ranked]
|
||||
|
||||
|
||||
def classify(ranked: List[Tuple[str, float]]) -> str:
|
||||
"""Return the predominant class, "Mixed", or blank when there is no data."""
|
||||
if not ranked:
|
||||
return ""
|
||||
top_share = ranked[0][1]
|
||||
second_share = ranked[1][1] if len(ranked) > 1 else 0.0
|
||||
has_majority = top_share >= PREDOMINANT_MIN_SHARE - RULE_TOLERANCE
|
||||
has_clear_lead = top_share - second_share >= PREDOMINANT_MIN_GAP - RULE_TOLERANCE
|
||||
return ranked[0][0] if has_majority and has_clear_lead else MIXED_CLASS
|
||||
|
||||
|
||||
def _format_share(ranked: List[Tuple[str, float]], index: int) -> Tuple[str, str]:
|
||||
"""Return the class and 4-decimal share at a rank, or blanks."""
|
||||
if index >= len(ranked):
|
||||
return "", ""
|
||||
code, share = ranked[index]
|
||||
return code, f"{share:.4f}"
|
||||
|
||||
|
||||
def build_koppen_records(
|
||||
counties: gpd.GeoDataFrame,
|
||||
koppen_raster: Path,
|
||||
code_map: Dict[int, str],
|
||||
subcells: int = DEFAULT_SUBCELLS,
|
||||
) -> List[dict]:
|
||||
"""Classify every county and return its metric row."""
|
||||
records: List[dict] = []
|
||||
with rasterio.open(koppen_raster) as source:
|
||||
raster_counties = counties
|
||||
if source.crs is not None and counties.crs is not None and counties.crs != source.crs:
|
||||
raster_counties = counties.to_crs(source.crs)
|
||||
geographic = source.crs is None or source.crs.is_geographic
|
||||
|
||||
for county, geometry in zip(counties.itertuples(), raster_counties.geometry):
|
||||
weights = area_weighted_class_weights(source, geometry, geographic=geographic, subcells=subcells)
|
||||
ranked = rank_class_shares(weights, code_map)
|
||||
top_class, top_share = _format_share(ranked, 0)
|
||||
second_class, second_share = _format_share(ranked, 1)
|
||||
records.append(
|
||||
{
|
||||
"countyFips": county.county_fips,
|
||||
"countyName": county.county_name,
|
||||
"state": county.state,
|
||||
"koppenZone": classify(ranked),
|
||||
"koppenTopClass": top_class,
|
||||
"koppenTopShare": top_share,
|
||||
"koppenSecondClass": second_class,
|
||||
"koppenSecondShare": second_share,
|
||||
}
|
||||
)
|
||||
return records
|
||||
|
||||
|
||||
def write_records(records: List[dict], out_file: Path) -> None:
|
||||
"""Write the Koppen metric rows."""
|
||||
out_file.parent.mkdir(parents=True, exist_ok=True)
|
||||
with out_file.open("w", encoding="utf-8", newline="") as csv_file:
|
||||
writer = csv.DictWriter(csv_file, fieldnames=FIELDS)
|
||||
writer.writeheader()
|
||||
writer.writerows(records)
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
"""Define and parse command-line options for this builder."""
|
||||
parser = argparse.ArgumentParser(description="Build the county Koppen-Geiger metric file.")
|
||||
parser.add_argument("--counties-geojson", type=Path, default=DEFAULT_COUNTIES_GEOJSON, help="County polygon GeoJSON path.")
|
||||
parser.add_argument("--koppen-raster", type=Path, default=DEFAULT_KOPPEN_RASTER, help="Koppen-Geiger raster TIFF path.")
|
||||
parser.add_argument("--koppen-legend", type=Path, default=DEFAULT_KOPPEN_LEGEND, help="legend.txt mapping raster codes.")
|
||||
parser.add_argument("--subcells", type=int, default=DEFAULT_SUBCELLS, help="Sub-cells per raster cell edge.")
|
||||
parser.add_argument("--out", type=Path, default=DEFAULT_OUT, help="Output metric CSV path.")
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
counties = load_counties(args.counties_geojson)
|
||||
records = build_koppen_records(counties, args.koppen_raster, load_koppen_legend(args.koppen_legend), args.subcells)
|
||||
write_records(records, args.out)
|
||||
|
||||
outcomes = Counter(
|
||||
"blank" if not r["koppenZone"] else "Mixed" if r["koppenZone"] == MIXED_CLASS else "predominant" for r in records
|
||||
)
|
||||
print(
|
||||
f"Wrote {len(records)} counties to {args.out}: {outcomes['predominant']} predominant, "
|
||||
f"{outcomes['Mixed']} Mixed, {outcomes['blank']} blank."
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,427 @@
|
||||
"""Validate data/climate-data.csv before the browser app loads it.
|
||||
|
||||
Checks file format, columns, county keys, the county geometry join, per-metric
|
||||
types and ranges, allowed blank values, and cross-field consistency. Exits with
|
||||
status 1 when any check fails.
|
||||
|
||||
Run:
|
||||
.venv\\Scripts\\python.exe scripts\\check_climate_data.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import csv
|
||||
import json
|
||||
import math
|
||||
import re
|
||||
import sys
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import Dict, FrozenSet, List, Optional, Sequence, Tuple
|
||||
|
||||
PROJECT_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_CLIMATE_CSV = PROJECT_ROOT / "data" / "climate-data.csv"
|
||||
DEFAULT_COUNTIES_GEOJSON = PROJECT_ROOT / "data" / "geojson-counties-fips.json"
|
||||
DEFAULT_METRIC_SOURCES = PROJECT_ROOT / "data" / "metric_sources.json"
|
||||
|
||||
FIPS_PATTERN = re.compile(r"\d{5}")
|
||||
STATE_PATTERN = re.compile(r"[A-Z]{2}")
|
||||
|
||||
# The distinct-value check only runs on full-size files, not small fixtures.
|
||||
COLLAPSE_CHECK_MIN_ROWS = 100
|
||||
|
||||
KOPPEN_CODES = (
|
||||
"Af", "Am", "Aw", "BWh", "BWk", "BSh", "BSk",
|
||||
"Csa", "Csb", "Csc", "Cwa", "Cwb", "Cwc", "Cfa", "Cfb", "Cfc",
|
||||
"Dsa", "Dsb", "Dsc", "Dsd", "Dwa", "Dwb", "Dwc", "Dwd",
|
||||
"Dfa", "Dfb", "Dfc", "Dfd", "ET", "EF",
|
||||
)
|
||||
# Counties with no predominant class; see docs/filter-calculations.md, section 1.
|
||||
MIXED_KOPPEN_CLASS = "Mixed"
|
||||
|
||||
MONTH_NAMES = (
|
||||
"January", "February", "March", "April", "May", "June",
|
||||
"July", "August", "September", "October", "November", "December",
|
||||
)
|
||||
|
||||
# NOAA nClimGrid and gridMET cover the contiguous U.S. only.
|
||||
OUTSIDE_CONUS = frozenset({"AK", "HI", "PR"})
|
||||
# NSRDB summaries were not requested for Puerto Rico.
|
||||
NSRDB_NOT_REQUESTED = frozenset({"PR"})
|
||||
# NOAA county daily files do not include Lexington city, VA.
|
||||
NOAA_DAILY_MISSING_FIPS = frozenset({"51678"})
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class MetricRule:
|
||||
"""Validation rule for one metric column."""
|
||||
|
||||
kind: str
|
||||
minimum: Optional[float] = None
|
||||
maximum: Optional[float] = None
|
||||
integer: bool = False
|
||||
categories: Tuple[str, ...] = ()
|
||||
blank_states: FrozenSet[str] = frozenset()
|
||||
blank_fips: FrozenSet[str] = frozenset()
|
||||
min_distinct: int = 2
|
||||
|
||||
|
||||
def _numeric(
|
||||
minimum: float,
|
||||
maximum: float,
|
||||
*,
|
||||
integer: bool = False,
|
||||
blank_states: FrozenSet[str] = frozenset(),
|
||||
blank_fips: FrozenSet[str] = frozenset(),
|
||||
) -> MetricRule:
|
||||
"""Build a numeric rule with a plausible physical range."""
|
||||
return MetricRule(
|
||||
"numeric",
|
||||
minimum=minimum,
|
||||
maximum=maximum,
|
||||
integer=integer,
|
||||
blank_states=blank_states,
|
||||
blank_fips=blank_fips,
|
||||
min_distinct=10,
|
||||
)
|
||||
|
||||
|
||||
def _categorical(categories: Sequence[str], *, blank_states: FrozenSet[str] = frozenset()) -> MetricRule:
|
||||
"""Build a categorical rule limited to an allowed value list."""
|
||||
return MetricRule("categorical", categories=tuple(categories), blank_states=blank_states)
|
||||
|
||||
|
||||
# Ranges are physical plausibility limits, deliberately wider than the current
|
||||
# data. They are not the app's color-scale bounds.
|
||||
METRIC_RULES: Dict[str, MetricRule] = {
|
||||
"koppenZone": _categorical(KOPPEN_CODES + (MIXED_KOPPEN_CLASS,)),
|
||||
"avgTempF": _numeric(20, 85, blank_states=OUTSIDE_CONUS),
|
||||
"avgDiurnalTempRangeF": _numeric(5, 40, blank_states=OUTSIDE_CONUS, blank_fips=NOAA_DAILY_MISSING_FIPS),
|
||||
"annualPrecipIn": _numeric(1, 150, blank_states=OUTSIDE_CONUS),
|
||||
"seasonalityIndex": _numeric(0, 100, integer=True, blank_states=OUTSIDE_CONUS),
|
||||
"wettestPrecipMonth": _categorical(MONTH_NAMES, blank_states=OUTSIDE_CONUS),
|
||||
"driestPrecipMonth": _categorical(MONTH_NAMES, blank_states=OUTSIDE_CONUS),
|
||||
"absoluteExtremeDays": _numeric(0, 366, blank_states=OUTSIDE_CONUS, blank_fips=NOAA_DAILY_MISSING_FIPS),
|
||||
"meanDailyGlobalHorizontalRadiationKwhM2Day": _numeric(1.5, 7.5, blank_states=NSRDB_NOT_REQUESTED),
|
||||
"clearSkyGhiReductionIndex": _numeric(0, 1, blank_states=NSRDB_NOT_REQUESTED),
|
||||
"avgSummerSpecificHumidityGKg": _numeric(1, 25, blank_states=OUTSIDE_CONUS),
|
||||
"humidHeatDays": _numeric(0, 366, blank_states=OUTSIDE_CONUS),
|
||||
}
|
||||
|
||||
IDENTITY_COLUMNS = ("countyFips", "countyName", "state")
|
||||
# Stripe classes the app draws for Mixed Koppen counties; blank otherwise.
|
||||
KOPPEN_STRIPE_COLUMNS = ("koppenPrimaryClass", "koppenSecondaryClass")
|
||||
AUDIT_COLUMNS = ("humidHeatSourceFips", "humidHeatFipsAdjustment", "source")
|
||||
EXPECTED_COLUMNS = IDENTITY_COLUMNS + tuple(METRIC_RULES) + KOPPEN_STRIPE_COLUMNS + AUDIT_COLUMNS
|
||||
|
||||
CHECK_FORMAT = "file format"
|
||||
CHECK_COLUMNS = "columns"
|
||||
CHECK_COUNTIES = "county keys"
|
||||
CHECK_GEOMETRY = "map geometry join"
|
||||
CHECK_BLANKS = "blank values"
|
||||
CHECK_VALUES = "value types and ranges"
|
||||
CHECK_CROSS = "cross-field consistency"
|
||||
CHECK_SOURCES = "metric sources file"
|
||||
|
||||
|
||||
@dataclass
|
||||
class Report:
|
||||
"""Problems found per check, in the order checks ran."""
|
||||
|
||||
checks: Dict[str, List[str]] = field(default_factory=dict)
|
||||
notes: List[str] = field(default_factory=list)
|
||||
row_count: int = 0
|
||||
column_count: int = 0
|
||||
|
||||
def start(self, check: str) -> None:
|
||||
self.checks.setdefault(check, [])
|
||||
|
||||
def add(self, check: str, message: str) -> None:
|
||||
self.checks.setdefault(check, []).append(message)
|
||||
|
||||
@property
|
||||
def problem_count(self) -> int:
|
||||
return sum(len(messages) for messages in self.checks.values())
|
||||
|
||||
@property
|
||||
def ok(self) -> bool:
|
||||
return self.problem_count == 0
|
||||
|
||||
|
||||
def read_climate_csv(path: Path) -> Tuple[List[str], List[List[str]]]:
|
||||
"""Read the raw header and rows without DictReader's silent padding."""
|
||||
with path.open("r", encoding="utf-8", newline="") as handle:
|
||||
reader = csv.reader(handle)
|
||||
headers = next(reader, [])
|
||||
rows = [row for row in reader if row]
|
||||
return headers, rows
|
||||
|
||||
|
||||
def check_structure(
|
||||
report: Report, headers: List[str], raw_rows: List[List[str]]
|
||||
) -> Tuple[List[str], List[Dict[str, str]]]:
|
||||
"""Check encoding, header names, and row widths; return rows as dicts."""
|
||||
report.start(CHECK_FORMAT)
|
||||
report.start(CHECK_COLUMNS)
|
||||
|
||||
if headers and headers[0].startswith("\ufeff"):
|
||||
report.add(CHECK_FORMAT, "file starts with a UTF-8 BOM; the app would not find the first column")
|
||||
headers = [headers[0].lstrip("\ufeff")] + headers[1:]
|
||||
|
||||
seen = set()
|
||||
for name in headers:
|
||||
if name in seen:
|
||||
report.add(CHECK_COLUMNS, f"duplicate column {name!r}")
|
||||
seen.add(name)
|
||||
|
||||
for name in EXPECTED_COLUMNS:
|
||||
if name not in seen:
|
||||
report.add(CHECK_COLUMNS, f"missing column {name!r}")
|
||||
for name in headers:
|
||||
if name not in EXPECTED_COLUMNS:
|
||||
report.add(CHECK_COLUMNS, f"unexpected column {name!r} (add it to check_climate_data.py if intended)")
|
||||
|
||||
rows: List[Dict[str, str]] = []
|
||||
for index, raw in enumerate(raw_rows):
|
||||
if len(raw) != len(headers):
|
||||
report.add(CHECK_FORMAT, f"line {index + 2}: {len(raw)} fields, expected {len(headers)}")
|
||||
continue
|
||||
rows.append(dict(zip(headers, raw)))
|
||||
|
||||
report.row_count = len(raw_rows)
|
||||
report.column_count = len(headers)
|
||||
return headers, rows
|
||||
|
||||
|
||||
def check_counties(report: Report, rows: List[Dict[str, str]]) -> None:
|
||||
"""Check FIPS format and uniqueness, names, and state/FIPS agreement."""
|
||||
report.start(CHECK_COUNTIES)
|
||||
seen = set()
|
||||
prefix_states: Dict[str, set] = {}
|
||||
state_prefixes: Dict[str, set] = {}
|
||||
|
||||
for row in rows:
|
||||
fips = row.get("countyFips") or ""
|
||||
if not FIPS_PATTERN.fullmatch(fips):
|
||||
report.add(CHECK_COUNTIES, f"invalid countyFips {fips!r}")
|
||||
continue
|
||||
if fips in seen:
|
||||
report.add(CHECK_COUNTIES, f"duplicate countyFips {fips}")
|
||||
seen.add(fips)
|
||||
|
||||
if not (row.get("countyName") or "").strip():
|
||||
report.add(CHECK_COUNTIES, f"{fips}: countyName is blank")
|
||||
|
||||
state = row.get("state") or ""
|
||||
if not STATE_PATTERN.fullmatch(state):
|
||||
report.add(CHECK_COUNTIES, f"{fips}: invalid state {state!r}")
|
||||
continue
|
||||
prefix_states.setdefault(fips[:2], set()).add(state)
|
||||
state_prefixes.setdefault(state, set()).add(fips[:2])
|
||||
|
||||
for prefix, states in sorted(prefix_states.items()):
|
||||
if len(states) > 1:
|
||||
report.add(CHECK_COUNTIES, f"state FIPS {prefix} is labeled as {sorted(states)}")
|
||||
for state, prefixes in sorted(state_prefixes.items()):
|
||||
if len(prefixes) > 1:
|
||||
report.add(CHECK_COUNTIES, f"state {state} spans FIPS prefixes {sorted(prefixes)}")
|
||||
|
||||
|
||||
def _feature_fips(feature: dict) -> str:
|
||||
"""Return the 5-digit county FIPS stored on a GeoJSON feature."""
|
||||
props = feature.get("properties") or {}
|
||||
raw = feature.get("id") or props.get("id") or str(props.get("GEO_ID") or "")[-5:]
|
||||
text = str(raw or "").strip()
|
||||
return text.zfill(5) if text.isdigit() else text
|
||||
|
||||
|
||||
def check_geometry(report: Report, rows: List[Dict[str, str]], geojson_path: Path) -> None:
|
||||
"""Check that CSV counties and map polygons match one-to-one."""
|
||||
report.start(CHECK_GEOMETRY)
|
||||
if not geojson_path.exists():
|
||||
report.add(CHECK_GEOMETRY, f"county geometry file not found: {geojson_path}")
|
||||
return
|
||||
|
||||
features = json.loads(geojson_path.read_text(encoding="utf-8")).get("features", [])
|
||||
geo_fips = set()
|
||||
for feature in features:
|
||||
fips = _feature_fips(feature)
|
||||
if not FIPS_PATTERN.fullmatch(fips):
|
||||
report.add(CHECK_GEOMETRY, f"map feature without a valid county FIPS: {fips!r}")
|
||||
continue
|
||||
if fips in geo_fips:
|
||||
report.add(CHECK_GEOMETRY, f"map has duplicate polygons for {fips}")
|
||||
geo_fips.add(fips)
|
||||
state_fips = str((feature.get("properties") or {}).get("STATE") or "").strip()
|
||||
if state_fips and state_fips.zfill(2) != fips[:2]:
|
||||
report.add(CHECK_GEOMETRY, f"map feature {fips} has STATE {state_fips}")
|
||||
|
||||
csv_fips = {row.get("countyFips") or "" for row in rows}
|
||||
for fips in sorted(csv_fips - geo_fips):
|
||||
report.add(CHECK_GEOMETRY, f"{fips} is in the CSV but has no map polygon")
|
||||
for fips in sorted(geo_fips - csv_fips):
|
||||
report.add(CHECK_GEOMETRY, f"{fips} has a map polygon but no CSV row")
|
||||
|
||||
|
||||
def check_metrics(report: Report, headers: List[str], rows: List[Dict[str, str]]) -> None:
|
||||
"""Check each metric's blanks, types, ranges, and categories."""
|
||||
report.start(CHECK_BLANKS)
|
||||
report.start(CHECK_VALUES)
|
||||
|
||||
for key, rule in METRIC_RULES.items():
|
||||
if key not in headers:
|
||||
continue
|
||||
present: List[object] = []
|
||||
for row in rows:
|
||||
fips = row.get("countyFips") or "?"
|
||||
state = row.get("state") or ""
|
||||
value = row.get(key) or ""
|
||||
|
||||
if value == "":
|
||||
if state not in rule.blank_states and fips not in rule.blank_fips:
|
||||
report.add(CHECK_BLANKS, f"{fips} ({state}): {key} is blank")
|
||||
continue
|
||||
if value != value.strip():
|
||||
report.add(CHECK_VALUES, f"{fips}: {key}={value!r} has surrounding whitespace")
|
||||
continue
|
||||
|
||||
if rule.kind == "categorical":
|
||||
if value in rule.categories:
|
||||
present.append(value)
|
||||
else:
|
||||
report.add(CHECK_VALUES, f"{fips}: {key}={value!r} is not an allowed category")
|
||||
continue
|
||||
|
||||
try:
|
||||
number = float(value)
|
||||
except ValueError:
|
||||
report.add(CHECK_VALUES, f"{fips}: {key}={value!r} is not a number")
|
||||
continue
|
||||
if not math.isfinite(number):
|
||||
report.add(CHECK_VALUES, f"{fips}: {key}={value!r} is not finite")
|
||||
continue
|
||||
if rule.integer and not number.is_integer():
|
||||
report.add(CHECK_VALUES, f"{fips}: {key}={value} should be a whole number")
|
||||
if not rule.minimum <= number <= rule.maximum:
|
||||
report.add(CHECK_VALUES, f"{fips}: {key}={value} outside {rule.minimum:g}..{rule.maximum:g}")
|
||||
present.append(number)
|
||||
|
||||
if len(rows) >= COLLAPSE_CHECK_MIN_ROWS and len(set(present)) < rule.min_distinct:
|
||||
report.add(
|
||||
CHECK_VALUES,
|
||||
f"{key} has only {len(set(present))} distinct values; the column may have been overwritten",
|
||||
)
|
||||
|
||||
|
||||
def check_cross_fields(report: Report, rows: List[Dict[str, str]]) -> None:
|
||||
"""Check relationships between columns within each row."""
|
||||
report.start(CHECK_CROSS)
|
||||
for row in rows:
|
||||
fips = row.get("countyFips") or "?"
|
||||
|
||||
wet = row.get("wettestPrecipMonth") or ""
|
||||
dry = row.get("driestPrecipMonth") or ""
|
||||
if bool(wet) != bool(dry):
|
||||
report.add(CHECK_CROSS, f"{fips}: only one of wettestPrecipMonth/driestPrecipMonth is set")
|
||||
elif wet and wet == dry:
|
||||
report.add(CHECK_CROSS, f"{fips}: wettest and driest month are both {wet}")
|
||||
|
||||
heat = row.get("humidHeatDays") or ""
|
||||
heat_source = row.get("humidHeatSourceFips") or ""
|
||||
if bool(heat) != bool(heat_source):
|
||||
report.add(CHECK_CROSS, f"{fips}: humidHeatDays and humidHeatSourceFips must both be set or both blank")
|
||||
if heat_source and not FIPS_PATTERN.fullmatch(heat_source):
|
||||
report.add(CHECK_CROSS, f"{fips}: invalid humidHeatSourceFips {heat_source!r}")
|
||||
elif heat_source and heat_source != fips and not (row.get("humidHeatFipsAdjustment") or "").strip():
|
||||
report.add(CHECK_CROSS, f"{fips}: uses proxy county {heat_source} without a humidHeatFipsAdjustment note")
|
||||
|
||||
zone = row.get("koppenZone") or ""
|
||||
primary = row.get("koppenPrimaryClass") or ""
|
||||
secondary = row.get("koppenSecondaryClass") or ""
|
||||
if zone == MIXED_KOPPEN_CLASS:
|
||||
if not (primary and secondary):
|
||||
report.add(CHECK_CROSS, f"{fips}: Mixed Koppen county needs koppenPrimaryClass and koppenSecondaryClass")
|
||||
elif primary not in KOPPEN_CODES or secondary not in KOPPEN_CODES:
|
||||
report.add(CHECK_CROSS, f"{fips}: invalid Koppen stripe classes {primary!r}/{secondary!r}")
|
||||
elif primary == secondary:
|
||||
report.add(CHECK_CROSS, f"{fips}: Koppen stripe classes are both {primary}")
|
||||
elif primary or secondary:
|
||||
report.add(CHECK_CROSS, f"{fips}: Koppen stripe classes are set but koppenZone is {zone or 'blank'}, not Mixed")
|
||||
|
||||
if not (row.get("source") or "").strip():
|
||||
report.add(CHECK_CROSS, f"{fips}: source is blank")
|
||||
|
||||
|
||||
def check_metric_sources(report: Report, path: Path) -> None:
|
||||
"""Check that the metric metadata file parses and names real metrics."""
|
||||
report.start(CHECK_SOURCES)
|
||||
if not path.exists():
|
||||
report.add(CHECK_SOURCES, f"{path.name} not found")
|
||||
return
|
||||
try:
|
||||
data = json.loads(path.read_text(encoding="utf-8"))
|
||||
except json.JSONDecodeError as exc:
|
||||
report.add(CHECK_SOURCES, f"{path.name} is not valid JSON: {exc}")
|
||||
return
|
||||
if not isinstance(data, dict) or not isinstance(data.get("metrics"), dict):
|
||||
report.add(CHECK_SOURCES, f'{path.name} must be a JSON object with a "metrics" object')
|
||||
return
|
||||
|
||||
documented = set(data["metrics"])
|
||||
for key in sorted(documented - set(METRIC_RULES)):
|
||||
report.add(CHECK_SOURCES, f"{path.name} describes unknown metric {key!r}")
|
||||
covered = len(documented & set(METRIC_RULES))
|
||||
report.notes.append(f"{path.name} documents {covered} of {len(METRIC_RULES)} metrics.")
|
||||
|
||||
|
||||
def run_checks(csv_path: Path, geojson_path: Path, metric_sources_path: Path) -> Report:
|
||||
"""Run every check and return the combined report."""
|
||||
report = Report()
|
||||
headers, raw_rows = read_climate_csv(csv_path)
|
||||
headers, rows = check_structure(report, headers, raw_rows)
|
||||
check_counties(report, rows)
|
||||
check_geometry(report, rows, geojson_path)
|
||||
check_metrics(report, headers, rows)
|
||||
check_cross_fields(report, rows)
|
||||
check_metric_sources(report, metric_sources_path)
|
||||
return report
|
||||
|
||||
|
||||
def print_report(report: Report, csv_path: Path, max_examples: int) -> None:
|
||||
"""Print a pass/fail line per check with example problems."""
|
||||
print(f"Checked {csv_path}: {report.row_count} rows, {report.column_count} columns")
|
||||
for check, messages in report.checks.items():
|
||||
status = "PASS" if not messages else f"FAIL ({len(messages)})"
|
||||
print(f" {status:<11}{check}")
|
||||
for message in messages[:max_examples]:
|
||||
print(f" - {message}")
|
||||
if len(messages) > max_examples:
|
||||
print(f" ... and {len(messages) - max_examples} more")
|
||||
for note in report.notes:
|
||||
print(f" note: {note}")
|
||||
print("Result: PASS" if report.ok else f"Result: FAIL ({report.problem_count} problems)")
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
"""Define and parse command-line options for this checker."""
|
||||
parser = argparse.ArgumentParser(description="Validate the browser app's county climate CSV.")
|
||||
parser.add_argument("--csv", type=Path, default=DEFAULT_CLIMATE_CSV, help="Climate data CSV path.")
|
||||
parser.add_argument("--geojson", type=Path, default=DEFAULT_COUNTIES_GEOJSON, help="County GeoJSON path.")
|
||||
parser.add_argument("--metric-sources", type=Path, default=DEFAULT_METRIC_SOURCES, help="Metric metadata JSON path.")
|
||||
parser.add_argument("--max-examples", type=int, default=10, help="Problems to print per failing check.")
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
if not args.csv.exists():
|
||||
print(f"Climate data CSV not found: {args.csv}", file=sys.stderr)
|
||||
return 2
|
||||
report = run_checks(args.csv, args.geojson, args.metric_sources)
|
||||
print_report(report, args.csv, args.max_examples)
|
||||
return 0 if report.ok else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,5 @@
|
||||
"""Shared helpers imported by the county data pipeline scripts.
|
||||
|
||||
Modules here are not run directly. Scripts in the parent folder import them,
|
||||
for example ``from common.counties import load_counties``.
|
||||
"""
|
||||
@@ -0,0 +1,131 @@
|
||||
"""Load county polygons and normalize county identifiers."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import geopandas as gpd
|
||||
|
||||
DEFAULT_COUNTIES_GEOJSON_URL = "https://raw.githubusercontent.com/plotly/datasets/master/geojson-counties-fips.json"
|
||||
|
||||
STATE_FIPS_TO_ABBR = {
|
||||
"01": "AL",
|
||||
"02": "AK",
|
||||
"04": "AZ",
|
||||
"05": "AR",
|
||||
"06": "CA",
|
||||
"08": "CO",
|
||||
"09": "CT",
|
||||
"10": "DE",
|
||||
"11": "DC",
|
||||
"12": "FL",
|
||||
"13": "GA",
|
||||
"15": "HI",
|
||||
"16": "ID",
|
||||
"17": "IL",
|
||||
"18": "IN",
|
||||
"19": "IA",
|
||||
"20": "KS",
|
||||
"21": "KY",
|
||||
"22": "LA",
|
||||
"23": "ME",
|
||||
"24": "MD",
|
||||
"25": "MA",
|
||||
"26": "MI",
|
||||
"27": "MN",
|
||||
"28": "MS",
|
||||
"29": "MO",
|
||||
"30": "MT",
|
||||
"31": "NE",
|
||||
"32": "NV",
|
||||
"33": "NH",
|
||||
"34": "NJ",
|
||||
"35": "NM",
|
||||
"36": "NY",
|
||||
"37": "NC",
|
||||
"38": "ND",
|
||||
"39": "OH",
|
||||
"40": "OK",
|
||||
"41": "OR",
|
||||
"42": "PA",
|
||||
"44": "RI",
|
||||
"45": "SC",
|
||||
"46": "SD",
|
||||
"47": "TN",
|
||||
"48": "TX",
|
||||
"49": "UT",
|
||||
"50": "VT",
|
||||
"51": "VA",
|
||||
"53": "WA",
|
||||
"54": "WV",
|
||||
"55": "WI",
|
||||
"56": "WY",
|
||||
"60": "AS",
|
||||
"66": "GU",
|
||||
"69": "MP",
|
||||
"72": "PR",
|
||||
"78": "VI",
|
||||
}
|
||||
|
||||
|
||||
def normalize_fips(value: object, width: int) -> str:
|
||||
"""Return a zero-padded FIPS code with the requested width."""
|
||||
text = str(value).strip()
|
||||
digits = "".join(ch for ch in text if ch.isdigit())
|
||||
if not digits:
|
||||
return ""
|
||||
return digits.zfill(width)[-width:]
|
||||
|
||||
|
||||
def load_counties(counties_geojson: Path) -> gpd.GeoDataFrame:
|
||||
"""Load county polygons and normalize fields used downstream."""
|
||||
if not counties_geojson.exists():
|
||||
try:
|
||||
print(
|
||||
f"County GeoJSON not found at {counties_geojson}. "
|
||||
f"Attempting download from {DEFAULT_COUNTIES_GEOJSON_URL}..."
|
||||
)
|
||||
gdf = gpd.read_file(DEFAULT_COUNTIES_GEOJSON_URL)
|
||||
counties_geojson.parent.mkdir(parents=True, exist_ok=True)
|
||||
# Cache the downloaded file for subsequent runs.
|
||||
gdf.to_file(counties_geojson, driver="GeoJSON")
|
||||
print(f"Downloaded and cached county GeoJSON to {counties_geojson}")
|
||||
except Exception as exc:
|
||||
raise FileNotFoundError(
|
||||
f"County GeoJSON not found at {counties_geojson}, and download from "
|
||||
f"{DEFAULT_COUNTIES_GEOJSON_URL} failed. Download the file manually "
|
||||
"and rerun with --counties-geojson pointing to it."
|
||||
) from exc
|
||||
|
||||
gdf = gpd.read_file(counties_geojson)
|
||||
if gdf.crs is None:
|
||||
gdf = gdf.set_crs("EPSG:4326")
|
||||
else:
|
||||
gdf = gdf.to_crs("EPSG:4326")
|
||||
|
||||
feature_id = None
|
||||
if "id" in gdf.columns:
|
||||
feature_id = gdf["id"]
|
||||
elif "GEOID" in gdf.columns:
|
||||
feature_id = gdf["GEOID"]
|
||||
elif "GEOID10" in gdf.columns:
|
||||
feature_id = gdf["GEOID10"]
|
||||
elif "fips" in gdf.columns:
|
||||
feature_id = gdf["fips"]
|
||||
else:
|
||||
raise ValueError("Unable to locate county FIPS identifier column in county polygons.")
|
||||
|
||||
gdf["county_fips"] = feature_id.map(lambda value: normalize_fips(value, 5))
|
||||
gdf = gdf[gdf["county_fips"] != ""].copy()
|
||||
|
||||
if "NAME" in gdf.columns:
|
||||
gdf["county_name"] = gdf["NAME"].fillna("").astype(str).str.strip()
|
||||
elif "name" in gdf.columns:
|
||||
gdf["county_name"] = gdf["name"].fillna("").astype(str).str.strip()
|
||||
else:
|
||||
gdf["county_name"] = gdf["county_fips"].map(lambda value: f"County {value}")
|
||||
|
||||
gdf["state_fips"] = gdf["county_fips"].str.slice(0, 2)
|
||||
gdf["state"] = gdf["state_fips"].map(lambda code: STATE_FIPS_TO_ABBR.get(code, f"S{code}"))
|
||||
gdf = gdf.sort_values("county_fips").reset_index(drop=True)
|
||||
return gdf
|
||||
@@ -0,0 +1,128 @@
|
||||
"""Shared county zonal statistics for raster-based metrics.
|
||||
|
||||
Area weighting estimates how much of each raster cell lies inside a county by
|
||||
rasterizing the county on a finer grid of sub-cells, then scales each cell by
|
||||
the cosine of its latitude so cells count by their true surface area. Counties
|
||||
that cross the 180th meridian are split so each side is read from its own small
|
||||
raster window.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
from typing import Dict, List
|
||||
|
||||
import numpy as np
|
||||
import shapely
|
||||
from affine import Affine
|
||||
from rasterio.features import rasterize
|
||||
from rasterio.windows import Window, from_bounds
|
||||
from shapely.affinity import translate
|
||||
from shapely.geometry import box, mapping
|
||||
from shapely.geometry.base import BaseGeometry
|
||||
from shapely.ops import unary_union
|
||||
|
||||
DEFAULT_SUBCELLS = 16
|
||||
# Largest sub-cell grid rasterized for one county piece; bigger pieces use a coarser grid.
|
||||
SUBCELL_BUDGET = 80_000_000
|
||||
|
||||
|
||||
def split_at_antimeridian(geometry: BaseGeometry) -> List[BaseGeometry]:
|
||||
"""Split a lon/lat geometry into pieces that each stay on one side of 180 degrees.
|
||||
|
||||
A county such as Aleutians West, AK has islands at both +179 and -179
|
||||
degrees longitude. Its bounding box then spans nearly the whole globe, so
|
||||
each side is returned as its own piece.
|
||||
"""
|
||||
minx, _, maxx, _ = geometry.bounds
|
||||
if maxx - minx <= 180.0:
|
||||
return [geometry]
|
||||
|
||||
positive: List[BaseGeometry] = []
|
||||
negative: List[BaseGeometry] = []
|
||||
for part in getattr(geometry, "geoms", [geometry]):
|
||||
part_minx, _, part_maxx, _ = part.bounds
|
||||
if part_maxx - part_minx > 180.0:
|
||||
# One outline crosses the line: unwrap to 0..360, cut at 180, rewrap.
|
||||
unwrapped = shapely.transform(
|
||||
part,
|
||||
lambda xy: np.column_stack((np.where(xy[:, 0] < 0, xy[:, 0] + 360.0, xy[:, 0]), xy[:, 1])),
|
||||
)
|
||||
positive.append(unwrapped.intersection(box(0.0, -90.0, 180.0, 90.0)))
|
||||
negative.append(translate(unwrapped.intersection(box(180.0, -90.0, 360.0, 90.0)), xoff=-360.0))
|
||||
elif part_minx >= 0:
|
||||
positive.append(part)
|
||||
else:
|
||||
negative.append(part)
|
||||
|
||||
pieces = [unary_union(group) for group in (positive, negative) if group]
|
||||
return [piece for piece in pieces if not piece.is_empty]
|
||||
|
||||
|
||||
def geometry_window(source, geometry: BaseGeometry) -> Window:
|
||||
"""Return the raster window covering a geometry, padded by one cell on each side."""
|
||||
window = from_bounds(*geometry.bounds, transform=source.transform)
|
||||
col_start = math.floor(window.col_off) - 1
|
||||
row_start = math.floor(window.row_off) - 1
|
||||
col_stop = math.ceil(window.col_off + window.width) + 1
|
||||
row_stop = math.ceil(window.row_off + window.height) + 1
|
||||
padded = Window(col_start, row_start, col_stop - col_start, row_stop - row_start)
|
||||
return padded.intersection(Window(0, 0, source.width, source.height))
|
||||
|
||||
|
||||
def subcells_for(shape: tuple[int, int], requested: int) -> int:
|
||||
"""Return the finest sub-cell count, up to the request, that fits the budget."""
|
||||
rows, cols = shape
|
||||
subcells = requested
|
||||
while subcells > 1 and rows * cols * subcells * subcells > SUBCELL_BUDGET:
|
||||
subcells //= 2
|
||||
return subcells
|
||||
|
||||
|
||||
def cell_coverage_fractions(
|
||||
shape: tuple[int, int], transform: Affine, geometry: BaseGeometry, subcells: int
|
||||
) -> np.ndarray:
|
||||
"""Estimate the fraction of each raster cell covered by a geometry."""
|
||||
rows, cols = shape
|
||||
fine = rasterize(
|
||||
[mapping(geometry)],
|
||||
out_shape=(rows * subcells, cols * subcells),
|
||||
transform=transform * Affine.scale(1.0 / subcells),
|
||||
fill=0,
|
||||
default_value=1,
|
||||
dtype="uint8",
|
||||
)
|
||||
return fine.reshape(rows, subcells, cols, subcells).mean(axis=(1, 3))
|
||||
|
||||
|
||||
def area_weighted_class_weights(
|
||||
source,
|
||||
geometry: BaseGeometry,
|
||||
*,
|
||||
geographic: bool = True,
|
||||
subcells: int = DEFAULT_SUBCELLS,
|
||||
) -> Dict[int, float]:
|
||||
"""Return the area inside a geometry covered by each value of a categorical raster.
|
||||
|
||||
Weights are relative surface areas: the fraction of each cell inside the
|
||||
geometry, times cos(latitude) for geographic rasters. Cells equal to 0 or
|
||||
the raster's nodata value are excluded.
|
||||
"""
|
||||
weights: Dict[int, float] = {}
|
||||
pieces = split_at_antimeridian(geometry) if geographic else [geometry]
|
||||
for piece in pieces:
|
||||
window = geometry_window(source, piece)
|
||||
values = source.read(1, window=window, masked=True).filled(0)
|
||||
if source.nodata is not None:
|
||||
values = np.where(values == source.nodata, 0, values)
|
||||
transform = source.window_transform(window)
|
||||
cell_weights = cell_coverage_fractions(values.shape, transform, piece, subcells_for(values.shape, subcells))
|
||||
if geographic:
|
||||
row_lat = transform.f + (np.arange(values.shape[0]) + 0.5) * transform.e
|
||||
cell_weights = cell_weights * np.cos(np.radians(row_lat))[:, None]
|
||||
|
||||
counted = (values != 0) & (cell_weights > 0)
|
||||
for value in np.unique(values[counted]):
|
||||
code = int(value)
|
||||
weights[code] = weights.get(code, 0.0) + float(cell_weights[counted & (values == value)].sum())
|
||||
return weights
|
||||
@@ -0,0 +1,64 @@
|
||||
"""Koppen-Geiger raster codes and the Beck et al. legend loader."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from pathlib import Path
|
||||
from typing import Dict
|
||||
|
||||
# Beck et al legend key is expected as text file, but this default handles common codes.
|
||||
DEFAULT_KOPPEN_CODE_MAP = {
|
||||
1: "Af",
|
||||
2: "Am",
|
||||
3: "Aw",
|
||||
4: "BWh",
|
||||
5: "BWk",
|
||||
6: "BSh",
|
||||
7: "BSk",
|
||||
8: "Csa",
|
||||
9: "Csb",
|
||||
10: "Csc",
|
||||
11: "Cwa",
|
||||
12: "Cwb",
|
||||
13: "Cwc",
|
||||
14: "Cfa",
|
||||
15: "Cfb",
|
||||
16: "Cfc",
|
||||
17: "Dsa",
|
||||
18: "Dsb",
|
||||
19: "Dsc",
|
||||
20: "Dsd",
|
||||
21: "Dwa",
|
||||
22: "Dwb",
|
||||
23: "Dwc",
|
||||
24: "Dwd",
|
||||
25: "Dfa",
|
||||
26: "Dfb",
|
||||
27: "Dfc",
|
||||
28: "Dfd",
|
||||
29: "ET",
|
||||
30: "EF",
|
||||
}
|
||||
|
||||
|
||||
def load_koppen_legend(legend_path: Path | None) -> Dict[int, str]:
|
||||
"""Load Koppen raster codes, using defaults when no legend exists."""
|
||||
if legend_path is None:
|
||||
return DEFAULT_KOPPEN_CODE_MAP
|
||||
|
||||
mapping: Dict[int, str] = {}
|
||||
for line in legend_path.read_text(encoding="utf-8").splitlines():
|
||||
text = line.strip()
|
||||
if not text or text.startswith("#"):
|
||||
continue
|
||||
# Handles patterns like:
|
||||
# "1: Af ..." or "1 = Af" or "1 Af"
|
||||
match = re.match(r"^(\d+)\s*[:=]?\s*([A-Za-z]{2,3})\b", text)
|
||||
if not match:
|
||||
continue
|
||||
|
||||
key = int(match.group(1))
|
||||
value = match.group(2)
|
||||
mapping[key] = value
|
||||
|
||||
return mapping if mapping else DEFAULT_KOPPEN_CODE_MAP
|
||||
@@ -12,13 +12,24 @@ The browser blocks `fetch("data/climate-data.csv")` when `index.html` is opened
|
||||
|
||||
Then open [http://localhost:8000/](http://localhost:8000/). This keeps the app CSV-only while allowing the map and filters to load normally.
|
||||
|
||||
## Source 1: Koppen-Geiger classes (`koppenZone`)
|
||||
## Source 1: Koppen-Geiger classes (`koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`)
|
||||
|
||||
- Dataset: Beck et al. updated 1-km Koppen-Geiger climate classes (historical + future windows)
|
||||
- Landing page: [https://www.gloh2o.org/koppen/](https://www.gloh2o.org/koppen/)
|
||||
- Primary paper for updated release: [https://www.nature.com/articles/s41597-023-02549-6](https://www.nature.com/articles/s41597-023-02549-6)
|
||||
- Coverage: 1901-2099 (use historical 1991-2020 layer for this project to align with NOAA baselines)
|
||||
- License shown on dataset page: CC BY 4.0
|
||||
- Local files: `data/koppen_geiger_tif/1991_2020/koppen_geiger_0p00833333.tif` and `data/koppen_geiger_tif/legend.txt`
|
||||
|
||||
Build the county metric and apply it to the app CSV:
|
||||
|
||||
```powershell
|
||||
.venv\Scripts\python.exe scripts\build_county_koppen_metric.py
|
||||
.venv\Scripts\python.exe scripts\apply_koppen_metric_to_climate_data.py --dry-run
|
||||
.venv\Scripts\python.exe scripts\apply_koppen_metric_to_climate_data.py
|
||||
```
|
||||
|
||||
The builder writes `data/metrics/koppen.csv` with each county's class, top and runner-up classes, and their area-weighted shares. The apply step writes `koppenZone`, `koppenPrimaryClass`, and `koppenSecondaryClass`; `--dry-run` reports the changes without writing. Run it after `build_county_climate_data.py`, which still writes an older largest-share `koppenZone`. The classification rule is documented in [`docs/filter-calculations.md`](../docs/filter-calculations.md) §1.
|
||||
|
||||
## Source 2: NOAA 1991-2020 gridded normals (`avgTempF`, `annualPrecipIn`, `seasonalityIndex`, previous `extremeDays`)
|
||||
|
||||
@@ -255,7 +266,8 @@ Then update the app CSV. Polygon archive GHI is used first; representative-point
|
||||
|
||||
## Metric definitions in generated output
|
||||
|
||||
- `koppenZone`: majority class within county polygon from Koppen raster.
|
||||
- `koppenZone`: the county's predominant Koppen-Geiger class, meaning the class covering at least 50% of the county's land area and leading the runner-up by at least 5 percentage points; otherwise `Mixed`. Shares are area-weighted, with ocean and no-data cells excluded.
|
||||
- `koppenPrimaryClass` / `koppenSecondaryClass`: for Mixed counties only, the top and runner-up classes, drawn as stripes on the map; blank for predominant counties.
|
||||
- `avgTempF`: mean of 12 monthly county mean temperatures, converted C -> F.
|
||||
- `annualPrecipIn`: sum of 12 monthly county mean precipitation totals, converted mm -> inches.
|
||||
- `seasonalityIndex`: coefficient of variation of monthly precipitation totals, scaled to 0-100.
|
||||
@@ -295,6 +307,7 @@ python scripts/build_county_climate_data.py `
|
||||
|
||||
Notes:
|
||||
|
||||
- This writes only the base columns and an older largest-share `koppenZone`. Do not run it over the live `data/climate-data.csv`: it would drop the columns added by later stages. After a full rebuild, run the Köppen build and apply steps (Source 1) and the enrichment stages.
|
||||
- This computes all counties in your geometry file, not just the sample records.
|
||||
- For counties outside CONUS coverage in NOAA gridded files, fallback values are applied by the script when no valid grid values intersect.
|
||||
- For physically-based daily `extremeDays`, provide true daily grids and set `--extreme-days-mode require-daily`.
|
||||
|
||||
Reference in New Issue
Block a user