# Data Pipeline Improvement Plan This plan describes how the county data pipeline will move from scripts that edit one shared CSV in place to per-metric outputs assembled into the app CSV. It is a working reference for the filter-by-filter review. Calculation details for each filter live in [filter-calculations.md](filter-calculations.md); findings and tasks live in [reviews/](reviews/), one file per filter plus [00-cross-filter.md](reviews/00-cross-filter.md); open and past decisions live in [decisions.md](decisions.md). **Started:** 2026-09-12 **Guiding decision:** review and fix each of the 12 filters one at a time, confirm each works on its own, and restructure `data/climate-data.csv` only after all filters are clean. No large rewrite happens up front. ## 1. Current pipeline ### Where each filter comes from | Filter | Written into `climate-data.csv` by | Upstream scripts | | --- | --- | --- | | Köppen-Geiger class (plus the two stripe-class columns) | `apply_koppen_metric_to_climate_data.py` | `build_county_koppen_metric.py` → `data/metrics/koppen.csv` | | Annual avg temperature | `build_county_climate_data.py` | — | | Annual precipitation | `build_county_climate_data.py` | — | | Seasonality index | `build_county_climate_data.py` | — | | Wettest / driest month | Base build, then overwritten by `apply_precipitation_month_metrics_to_climate_data.py` | — | | Diurnal temperature range | `apply_diurnal_temperature_range_to_climate_data.py` | `build_county_diurnal_temperature_range.py` | | Extreme temperature days | `apply_locally_extreme_metric_to_climate_data.py` | `build_county_locally_extreme_data.py` | | Summer specific humidity | `apply_gridmet_humidity_metric_to_climate_data.py` | `download_gridmet_data.py` → `summarize_county_gridmet_humidity.py` | | 90 °F+ heat-index days (plus 2 source-FIPS columns) | `apply_gridmet_humidity_metric_to_climate_data.py` | Same as summer humidity | | Solar GHI | Base build (optional), then replaced by `apply_locally_extreme_metric_to_climate_data.py` | Point: `build_county_representative_points.py` → `fetch_nsrdb_representative_point_ghi.py`. Polygon: `request_nsrdb_county_polygon_ghi_archives.py` → `download_nsrdb_county_polygon_ghi_archives.py` → `summarize_nsrdb_county_polygon_archives.py` | | Clear-sky GHI reduction | `apply_nsrdb_cloud_metric_to_climate_data.py` | Point: `fetch_nsrdb_representative_point_cloud_metrics.py`. Polygon: `request_nsrdb_county_polygon_cloud_archives.py` → `download_nsrdb_county_polygon_cloud_archives.py` → `summarize_nsrdb_county_polygon_cloud_archives.py` | Supporting scripts: `request_nsrdb_county_polygon_archives.py` and `download_nsrdb_county_polygon_archives.py` are the shared engines behind the GHI and cloud wrappers; `rebuild_nsrdb_representative_point_ghi_summary.py` rebuilds the point GHI summary from cache; `check_climate_data.py` validates the final CSV. Shared helpers live in `scripts/common/` (Phase 2). The base build still writes an old largest-share `koppenZone`, so the Köppen apply step must run after it. ### Current full-rebuild order 1. `build_county_climate_data.py` 2. `build_county_koppen_metric.py` → `apply_koppen_metric_to_climate_data.py` 3. `apply_precipitation_month_metrics_to_climate_data.py` 4. `build_county_locally_extreme_data.py` → `apply_locally_extreme_metric_to_climate_data.py` 5. `build_county_diurnal_temperature_range.py` → `apply_diurnal_temperature_range_to_climate_data.py` 6. `summarize_county_gridmet_humidity.py` → `apply_gridmet_humidity_metric_to_climate_data.py` 7. `apply_nsrdb_cloud_metric_to_climate_data.py` ### Problems 1. **Rerunning a step can destroy data.** The base build writes 12 columns, including the retired `extremeDays`. Later scripts delete, overwrite, or add columns until the live CSV has 20. Rerunning the base build drops 9 live columns (`koppenPrimaryClass`, `koppenSecondaryClass`, `avgDiurnalTempRangeF`, `absoluteExtremeDays`, `clearSkyGhiReductionIndex`, `avgSummerSpecificHumidityGKg`, `humidHeatDays`, `humidHeatSourceFips`, `humidHeatFipsAdjustment`), restores `extremeDays`, and rewrites `koppenZone` with the old largest-share method. 2. **Order is implicit.** The sequence lives in the README, in `scripts/county_data_sources.md`, and in each script's assumptions. 3. **Column ownership is unclear.** GHI is finalized by the extreme-temperature apply script; wettest/driest month are computed in two places. See findings 25 and 26 in [reviews/00-cross-filter.md](reviews/00-cross-filter.md). 4. **County aggregation is inconsistent.** See finding 11 in [reviews/00-cross-filter.md](reviews/00-cross-filter.md). 5. **No single entry point or final check.** A new user must piece together about 20 scripts, several large downloads, and an NSRDB API key. ## 2. Target design 1. **One metric, one file.** Each metric pipeline writes a county-level file under `data/metrics/`, for example `data/metrics/koppen.csv`, containing `countyFips`, the app value, and any audit columns for that metric. 2. **One assemble step.** A single script joins the metric files into `data/climate-data.csv`, using `data/metric_sources.json` for the column list and per-metric source notes, then runs `check_climate_data.py`. - Run order no longer matters; rerunning one metric cannot damage others. - Every column has exactly one owner. - The per-row `source` column moves into `metric_sources.json`. - Audit columns stay in the metric files rather than the app CSV. 3. **One shared county-aggregation module.** Area-weighted zonal statistics, including the 180th-meridian split, used by every raster-based metric. 4. **One runner.** For example `python scripts/pipeline.py --only koppen --skip-download`, with stages for fetch, build metrics, assemble, and check. Cached downloads are reused by default. ## 3. Reproduction tiers The "Reproducing the data" guide (Phase 4) will be organized by how deep a user needs to go: | Tier | What the user does | Needs | | --- | --- | --- | | 1. Run the app | `.\serve.ps1` with the committed CSV | Nothing else | | 2. Reassemble | Rebuild `climate-data.csv` from committed metric files | Python environment only | | 3. Regenerate one metric | Download one source, rebuild one metric file, reassemble | That metric's source data | | 4. Full rebuild | Everything | All sources; NSRDB API key; large downloads (the NOAA monthly temperature file alone is about 5 GB) and hours of paced NSRDB requests | The guide will list each dataset's size, download location, API-key needs, and approximate run time, and the Python requirements will be pinned. ## 4. Roadmap ### Phase 0 — Groundwork (done) - [x] `scripts/check_climate_data.py` validates the app CSV (8 checks) with tests in `tests/test_check_climate_data.py`. - [x] `data/metric_sources.json` created as an empty skeleton. - [x] Köppen raster reads use a padded window per county, and polygons that cross the 180th meridian are split (`split_at_antimeridian` in `scripts/common/county_zonal_stats.py`); output verified identical for all 3,221 counties; tests in `tests/test_koppen_antimeridian.py`. ### Phase 1 — Filter-by-filter review (in progress) Each filter goes through the checklist in section 5. Each fix delivers that metric's own file in `data/metrics/` plus a single-column apply step, so the existing CSV keeps working until Phase 3. ### Phase 2 — Shared helpers in `scripts/common/` Rules, current modules, and rationale are in [scripts/common/README.md](../scripts/common/README.md). Duplicated helpers found by the 2026-09-13 survey are finding 14 in [reviews/00-cross-filter.md](reviews/00-cross-filter.md). ### Phase 3 — Assemble and restructure (after all 12 filters are clean) - Assemble script that builds `climate-data.csv` from `data/metrics/`. - Populate `metric_sources.json`; remove the per-row `source` column. - Move audit columns out of the app CSV. - Point the app's Sources panel at `metric_sources.json`. - Retire or rewrite `build_county_climate_data.py` as per-metric builders. - Move `common/koppen_legend.py` next to the Köppen code once nothing outside Köppen imports it. ### Phase 4 — Runner and reproduction guide - `scripts/pipeline.py` runner with `--only` and `--skip-download`. - "Reproducing the data" guide organized by the tiers in section 3. - Pinned requirements. - End-to-end smoke test on a small synthetic county fixture. - Organize scripts by data source (`noaa/`, `gridmet/`, `nsrdb/`, `koppen/`), each holding its own helpers, with `common/` keeping only cross-source code. Scripts in subfolders are run through the runner or as modules (`python -m`), and the README and data-source commands are updated to match. ## 5. Per-filter review checklist For each filter: 1. Verify the calculation against the source data and document findings in the filter's review file under `docs/reviews/`. 2. Decide any rule or method changes with the project owner, and record them in `decisions.md`. 3. Record the adopted definition in `filter-calculations.md`. Until the new values are applied, keep the old definition below it under "Current method". 4. Implement the calculation, writing `data/metrics/.csv`. 5. Add a single-column apply step for the current CSV. 6. Update the rules in `check_climate_data.py`. 7. Add or update unit tests. 8. Apply to the CSV, run `check_climate_data.py`, and compare changed counties against expectations. Then rename "Current method" to "Previous method" in `filter-calculations.md`, as §1 does. 9. Update the app if the value set or display changes. ## 6. Filter tracker Each filter's findings, decisions, and tasks are in its review file. Findings that affect more than one filter are in [00-cross-filter.md](reviews/00-cross-filter.md), and [reviews/README.md](reviews/README.md) indexes every finding number. | # | Filter | Status | Review | | --- | --- | --- | --- | | 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | [01-koppen.md](reviews/01-koppen.md) | | 2 | Annual avg temperature | In progress: method decided (2026-09-15) | [02-annual-avg-temperature.md](reviews/02-annual-avg-temperature.md) | | 3 | Diurnal temperature range | Not started | [03-diurnal-temperature-range.md](reviews/03-diurnal-temperature-range.md) | | 4 | Extreme temperature days | Not started | [04-extreme-temperature-days.md](reviews/04-extreme-temperature-days.md) | | 5 | 90 °F+ heat-index days | Not started | [05-heat-index-days.md](reviews/05-heat-index-days.md) | | 6 | Annual precipitation | Not started | [06-annual-precipitation.md](reviews/06-annual-precipitation.md) | | 7 | Seasonality index | Not started | [07-seasonality-index.md](reviews/07-seasonality-index.md) | | 8 | Wettest month | Not started | [08-wettest-month.md](reviews/08-wettest-month.md) | | 9 | Driest month | Not started | [09-driest-month.md](reviews/09-driest-month.md) | | 10 | Summer specific humidity | Not started | [10-summer-specific-humidity.md](reviews/10-summer-specific-humidity.md) | | 11 | Solar GHI | Not started | [11-solar-ghi.md](reviews/11-solar-ghi.md) | | 12 | Clear-sky GHI reduction | Not started | [12-clear-sky-ghi-reduction.md](reviews/12-clear-sky-ghi-reduction.md) | ## 7. Guardrails until Phase 3 - **Do not rerun `build_county_climate_data.py`** against `data/climate-data.csv`. It would drop 9 live columns, restore `extremeDays`, and overwrite the Mixed classification in `koppenZone`. - Run `check_climate_data.py` after every apply step. - Change one filter at a time, and compare its before and after values.