Files
Climate-Mood-Analysis/docs/pipeline-plan.md
T
KnouandClaude Opus 5 e855d583e3 Document restructuring and the beginnings of Filter 2 changes
Split the pipeline documentation by purpose so each fact has one home:
- docs/pipeline-plan.md keeps the plan, checklist, tracker, and guardrails
- docs/decisions.md holds open decisions and the dated decision log
- docs/reviews/ holds findings and tasks: one file per filter, plus
  00-cross-filter.md for findings that span filters
- scripts/common/README.md holds the shared-helper rules (formerly Phase 2)
- filter-calculations.md now describes calculations only

Filed findings 12-22 from a consistency audit of the app, docs, and scripts.

Filter 1 (Köppen-Geiger): use "Köppen" with the umlaut in all prose, labels,
docstrings, help text, and checker messages (finding 21), and correct the
base build's "majority" docstring (finding 22).

Filter 2 (annual avg temperature): record the adopted definition in
filter-calculations.md §2: equally weighted 1991-2020 monthly normals, per
WMO-No. 1203 and NOAA's 2020 methodology; area-weighted county means; blank
unless all 12 months exist. Code changes for this filter are still pending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 16:13:11 -04:00

11 KiB

Data Pipeline Improvement Plan

This plan describes how the county data pipeline will move from scripts that edit one shared CSV in place to per-metric outputs assembled into the app CSV. It is a working reference for the filter-by-filter review. Calculation details for each filter live in filter-calculations.md; findings and tasks live in reviews/, one file per filter plus 00-cross-filter.md; open and past decisions live in decisions.md.

Started: 2026-09-12

Guiding decision: review and fix each of the 12 filters one at a time, confirm each works on its own, and restructure data/climate-data.csv only after all filters are clean. No large rewrite happens up front.

1. Current pipeline

Where each filter comes from

Filter Written into climate-data.csv by Upstream scripts
Köppen-Geiger class (plus the two stripe-class columns) apply_koppen_metric_to_climate_data.py build_county_koppen_metric.py → data/metrics/koppen.csv
Annual avg temperature build_county_climate_data.py —
Annual precipitation build_county_climate_data.py —
Seasonality index build_county_climate_data.py —
Wettest / driest month Base build, then overwritten by apply_precipitation_month_metrics_to_climate_data.py —
Diurnal temperature range apply_diurnal_temperature_range_to_climate_data.py build_county_diurnal_temperature_range.py
Extreme temperature days apply_locally_extreme_metric_to_climate_data.py build_county_locally_extreme_data.py
Summer specific humidity apply_gridmet_humidity_metric_to_climate_data.py download_gridmet_data.py → summarize_county_gridmet_humidity.py
90 °F+ heat-index days (plus 2 source-FIPS columns) apply_gridmet_humidity_metric_to_climate_data.py Same as summer humidity
Solar GHI Base build (optional), then replaced by apply_locally_extreme_metric_to_climate_data.py Point: build_county_representative_points.py → fetch_nsrdb_representative_point_ghi.py. Polygon: request_nsrdb_county_polygon_ghi_archives.py → download_nsrdb_county_polygon_ghi_archives.py → summarize_nsrdb_county_polygon_archives.py
Clear-sky GHI reduction apply_nsrdb_cloud_metric_to_climate_data.py Point: fetch_nsrdb_representative_point_cloud_metrics.py. Polygon: request_nsrdb_county_polygon_cloud_archives.py → download_nsrdb_county_polygon_cloud_archives.py → summarize_nsrdb_county_polygon_cloud_archives.py

Supporting scripts: request_nsrdb_county_polygon_archives.py and download_nsrdb_county_polygon_archives.py are the shared engines behind the GHI and cloud wrappers; rebuild_nsrdb_representative_point_ghi_summary.py rebuilds the point GHI summary from cache; check_climate_data.py validates the final CSV. Shared helpers live in scripts/common/ (Phase 2). The base build still writes an old largest-share koppenZone, so the Köppen apply step must run after it.

Current full-rebuild order

  1. build_county_climate_data.py
  2. build_county_koppen_metric.py → apply_koppen_metric_to_climate_data.py
  3. apply_precipitation_month_metrics_to_climate_data.py
  4. build_county_locally_extreme_data.py → apply_locally_extreme_metric_to_climate_data.py
  5. build_county_diurnal_temperature_range.py → apply_diurnal_temperature_range_to_climate_data.py
  6. summarize_county_gridmet_humidity.py → apply_gridmet_humidity_metric_to_climate_data.py
  7. apply_nsrdb_cloud_metric_to_climate_data.py

The README's enrichment list starts at step 2 and omits steps 1 and 3.

Problems

  1. Rerunning a step can destroy data. The base build writes 12 columns, including the retired extremeDays. Later scripts delete, overwrite, or add columns until the live CSV has 20. Rerunning the base build drops 9 live columns (koppenPrimaryClass, koppenSecondaryClass, avgDiurnalTempRangeF, absoluteExtremeDays, clearSkyGhiReductionIndex, avgSummerSpecificHumidityGKg, humidHeatDays, humidHeatSourceFips, humidHeatFipsAdjustment), restores extremeDays, and rewrites koppenZone with the old largest-share method.
  2. Order is implicit. The sequence lives in the README, in scripts/county_data_sources.md, and in each script's assumptions.
  3. Column ownership is unclear. GHI is finalized by the extreme-temperature apply script; wettest/driest month are computed in two places.
  4. County aggregation is inconsistent. See finding 11 in reviews/00-cross-filter.md.
  5. No single entry point or final check. A new user must piece together about 20 scripts, several large downloads, and an NSRDB API key.

2. Target design

  1. One metric, one file. Each metric pipeline writes a county-level file under data/metrics/, for example data/metrics/koppen.csv, containing countyFips, the app value, and any audit columns for that metric.
  2. One assemble step. A single script joins the metric files into data/climate-data.csv, using data/metric_sources.json for the column list and per-metric source notes, then runs check_climate_data.py.
    • Run order no longer matters; rerunning one metric cannot damage others.
    • Every column has exactly one owner.
    • The per-row source column moves into metric_sources.json.
    • Audit columns stay in the metric files rather than the app CSV.
  3. One shared county-aggregation module. Area-weighted zonal statistics, including the 180th-meridian split, used by every raster-based metric.
  4. One runner. For example python scripts/pipeline.py --only koppen --skip-download, with stages for fetch, build metrics, assemble, and check. Cached downloads are reused by default.

3. Reproduction tiers

The "Reproducing the data" guide (Phase 4) will be organized by how deep a user needs to go:

Tier What the user does Needs
1. Run the app .\serve.ps1 with the committed CSV Nothing else
2. Reassemble Rebuild climate-data.csv from committed metric files Python environment only
3. Regenerate one metric Download one source, rebuild one metric file, reassemble That metric's source data
4. Full rebuild Everything All sources; NSRDB API key; large downloads (the NOAA monthly temperature file alone is about 5 GB) and hours of paced NSRDB requests

The guide will list each dataset's size, download location, API-key needs, and approximate run time, and the Python requirements will be pinned.

4. Roadmap

Phase 0 — Groundwork (done)

  • scripts/check_climate_data.py validates the app CSV (8 checks) with tests in tests/test_check_climate_data.py.
  • data/metric_sources.json created as an empty skeleton.
  • Köppen raster reads use a padded window per county, and polygons that cross the 180th meridian are split (split_at_antimeridian in scripts/common/county_zonal_stats.py); output verified identical for all 3,221 counties; tests in tests/test_koppen_antimeridian.py.

Phase 1 — Filter-by-filter review (in progress)

Each filter goes through the checklist in section 5. Each fix delivers that metric's own file in data/metrics/ plus a single-column apply step, so the existing CSV keeps working until Phase 3.

Phase 2 — Shared helpers in scripts/common/

Rules, current modules, and rationale are in scripts/common/README.md. Duplicated helpers found by the 2026-09-13 survey are finding 14 in reviews/00-cross-filter.md.

Phase 3 — Assemble and restructure (after all 12 filters are clean)

  • Assemble script that builds climate-data.csv from data/metrics/.
  • Populate metric_sources.json; remove the per-row source column.
  • Move audit columns out of the app CSV.
  • Point the app's Sources panel at metric_sources.json.
  • Retire or rewrite build_county_climate_data.py as per-metric builders.
  • Move common/koppen_legend.py next to the Köppen code once nothing outside Köppen imports it.

Phase 4 — Runner and reproduction guide

  • scripts/pipeline.py runner with --only and --skip-download.
  • "Reproducing the data" guide organized by the tiers in section 3.
  • Pinned requirements.
  • End-to-end smoke test on a small synthetic county fixture.
  • Organize scripts by data source (noaa/, gridmet/, nsrdb/, koppen/), each holding its own helpers, with common/ keeping only cross-source code. Scripts in subfolders are run through the runner or as modules (python -m), and the README and data-source commands are updated to match.

5. Per-filter review checklist

For each filter:

  1. Verify the calculation against the source data and document findings in the filter's review file under docs/reviews/.
  2. Decide any rule or method changes with the project owner, and record them in decisions.md.
  3. Record the adopted definition in filter-calculations.md.
  4. Implement the calculation, writing data/metrics/<metric>.csv.
  5. Add a single-column apply step for the current CSV.
  6. Update the rules in check_climate_data.py.
  7. Add or update unit tests.
  8. Apply to the CSV, run check_climate_data.py, and compare changed counties against expectations.
  9. Update the app if the value set or display changes.

6. Filter tracker

Each filter's findings, decisions, and tasks are in its review file. Findings that affect more than one filter are in 00-cross-filter.md.

# Filter Status Review
1 Köppen-Geiger class Done (2026-09-14): rule applied, Mixed display built, documentation updated 01-koppen.md
2 Annual avg temperature In progress: method decided (2026-09-15) 02-annual-avg-temperature.md
3 Diurnal temperature range Not started 03-diurnal-temperature-range.md
4 Extreme temperature days Not started 04-extreme-temperature-days.md
5 90 °F+ heat-index days Not started 05-heat-index-days.md
6 Annual precipitation Not started 06-annual-precipitation.md
7 Seasonality index Not started 07-seasonality-index.md
8 Wettest month Not started 08-wettest-month.md
9 Driest month Not started 09-driest-month.md
10 Summer specific humidity Not started 10-summer-specific-humidity.md
11 Solar GHI Not started 11-solar-ghi.md
12 Clear-sky GHI reduction Not started 12-clear-sky-ghi-reduction.md

7. Guardrails until Phase 3

  • Do not rerun build_county_climate_data.py against data/climate-data.csv. It would drop 9 live columns, restore extremeDays, and overwrite the Mixed classification in koppenZone.
  • Run check_climate_data.py after every apply step.
  • Change one filter at a time, and compare its before and after values.