# Data Pipeline Improvement Plan This plan describes how the county data pipeline will move from scripts that edit one shared CSV in place to per-metric outputs assembled into the app CSV. It is a working reference for the filter-by-filter review. Calculation details for each filter live in [filter-calculations.md](filter-calculations.md). **Started:** 2026-09-12 **Guiding decision:** review and fix each of the 12 filters one at a time, confirm each works on its own, and restructure `data/climate-data.csv` only after all filters are clean. No large rewrite happens up front. ## 1. Current pipeline ### Where each filter comes from | Filter | Written into `climate-data.csv` by | Upstream scripts | | --- | --- | --- | | Köppen-Geiger class (plus the two stripe-class columns) | `apply_koppen_metric_to_climate_data.py` | `build_county_koppen_metric.py` → `data/metrics/koppen.csv` | | Annual avg temperature | `build_county_climate_data.py` | — | | Annual precipitation | `build_county_climate_data.py` | — | | Seasonality index | `build_county_climate_data.py` | — | | Wettest / driest month | Base build, then overwritten by `apply_precipitation_month_metrics_to_climate_data.py` | — | | Diurnal temperature range | `apply_diurnal_temperature_range_to_climate_data.py` | `build_county_diurnal_temperature_range.py` | | Extreme temperature days | `apply_locally_extreme_metric_to_climate_data.py` | `build_county_locally_extreme_data.py` | | Summer specific humidity | `apply_gridmet_humidity_metric_to_climate_data.py` | `download_gridmet_data.py` → `summarize_county_gridmet_humidity.py` | | 90 °F+ heat-index days (plus 2 source-FIPS columns) | `apply_gridmet_humidity_metric_to_climate_data.py` | Same as summer humidity | | Solar GHI | Base build (optional), then replaced by `apply_locally_extreme_metric_to_climate_data.py` | Point: `build_county_representative_points.py` → `fetch_nsrdb_representative_point_ghi.py`. Polygon: `request_nsrdb_county_polygon_ghi_archives.py` → `download_nsrdb_county_polygon_ghi_archives.py` → `summarize_nsrdb_county_polygon_archives.py` | | Clear-sky GHI reduction | `apply_nsrdb_cloud_metric_to_climate_data.py` | Point: `fetch_nsrdb_representative_point_cloud_metrics.py`. Polygon: `request_nsrdb_county_polygon_cloud_archives.py` → `download_nsrdb_county_polygon_cloud_archives.py` → `summarize_nsrdb_county_polygon_cloud_archives.py` | Supporting scripts: `request_nsrdb_county_polygon_archives.py` and `download_nsrdb_county_polygon_archives.py` are the shared engines behind the GHI and cloud wrappers; `rebuild_nsrdb_representative_point_ghi_summary.py` rebuilds the point GHI summary from cache; `check_climate_data.py` validates the final CSV. Shared helpers live in `scripts/common/` (Phase 2). The base build still writes an old largest-share `koppenZone`, so the Köppen apply step must run after it. ### Current full-rebuild order 1. `build_county_climate_data.py` 2. `build_county_koppen_metric.py` → `apply_koppen_metric_to_climate_data.py` 3. `apply_precipitation_month_metrics_to_climate_data.py` 4. `build_county_locally_extreme_data.py` → `apply_locally_extreme_metric_to_climate_data.py` 5. `build_county_diurnal_temperature_range.py` → `apply_diurnal_temperature_range_to_climate_data.py` 6. `summarize_county_gridmet_humidity.py` → `apply_gridmet_humidity_metric_to_climate_data.py` 7. `apply_nsrdb_cloud_metric_to_climate_data.py` The README's enrichment list starts at step 2 and omits steps 1 and 3. ### Problems 1. **Rerunning a step can destroy data.** The base build writes 12 columns, including the retired `extremeDays`. Later scripts delete, overwrite, or add columns until the live CSV has 20. Rerunning the base build drops 9 live columns (`koppenPrimaryClass`, `koppenSecondaryClass`, `avgDiurnalTempRangeF`, `absoluteExtremeDays`, `clearSkyGhiReductionIndex`, `avgSummerSpecificHumidityGKg`, `humidHeatDays`, `humidHeatSourceFips`, `humidHeatFipsAdjustment`), restores `extremeDays`, and rewrites `koppenZone` with the old largest-share method. 2. **Order is implicit.** The sequence lives in the README, in `scripts/county_data_sources.md`, and in each script's assumptions. 3. **Column ownership is unclear.** GHI is finalized by the extreme-temperature apply script; wettest/driest month are computed in two places. 4. **County aggregation is inconsistent.** NOAA uses touched raster cells, Köppen uses area-weighted shares (since 2026-09-13), gridMET uses cell centers with cos(latitude) weights, and NSRDB uses overlap areas (finding 11 in `filter-calculations.md`). 5. **No single entry point or final check.** A new user must piece together about 20 scripts, several large downloads, and an NSRDB API key. ## 2. Target design 1. **One metric, one file.** Each metric pipeline writes a county-level file under `data/metrics/`, for example `data/metrics/koppen.csv`, containing `countyFips`, the app value, and any audit columns for that metric. 2. **One assemble step.** A single script joins the metric files into `data/climate-data.csv`, using `data/metric_sources.json` for the column list and per-metric source notes, then runs `check_climate_data.py`. - Run order no longer matters; rerunning one metric cannot damage others. - Every column has exactly one owner. - The per-row `source` column moves into `metric_sources.json`. - Audit columns stay in the metric files rather than the app CSV. 3. **One shared county-aggregation module.** Area-weighted zonal statistics, including the 180th-meridian split, used by every raster-based metric. 4. **One runner.** For example `python scripts/pipeline.py --only koppen --skip-download`, with stages for fetch, build metrics, assemble, and check. Cached downloads are reused by default. ## 3. Reproduction tiers The "Reproducing the data" guide (Phase 4) will be organized by how deep a user needs to go: | Tier | What the user does | Needs | | --- | --- | --- | | 1. Run the app | `.\serve.ps1` with the committed CSV | Nothing else | | 2. Reassemble | Rebuild `climate-data.csv` from committed metric files | Python environment only | | 3. Regenerate one metric | Download one source, rebuild one metric file, reassemble | That metric's source data | | 4. Full rebuild | Everything | All sources; NSRDB API key; large downloads (the NOAA monthly temperature file alone is about 5 GB) and hours of paced NSRDB requests | The guide will list each dataset's size, download location, API-key needs, and approximate run time, and the Python requirements will be pinned. ## 4. Roadmap ### Phase 0 — Groundwork (done) - [x] `scripts/check_climate_data.py` validates the app CSV (8 checks) with tests in `tests/test_check_climate_data.py`. - [x] `data/metric_sources.json` created as an empty skeleton. - [x] Köppen raster reads use a padded window per county, and polygons that cross the 180th meridian are split (`split_at_antimeridian` in `scripts/common/county_zonal_stats.py`); output verified identical for all 3,221 counties; tests in `tests/test_koppen_antimeridian.py`. ### Phase 1 — Filter-by-filter review (in progress) Each filter goes through the checklist in section 5. Each fix delivers that metric's own file in `data/metrics/` plus a single-column apply step, so the existing CSV keeps working until Phase 3. ### Phase 2 — Shared helpers in `scripts/common/` `scripts/common/` holds code used by more than one data source (NOAA, gridMET, NSRDB, Köppen). Scripts import from it, for example `from common.counties import load_counties`; nothing in it is run directly. Rules for `common/`: - **Cross-source only.** A helper goes in only if metrics from more than one data source use it. Code shared by scripts of a single data source stays with that source, for example a future NSRDB module for the NSRDB prompt, redaction, and error-log helpers. - **One topic per module.** Each module is named for its topic and has a docstring. No catch-all `utils.py`. - **Keep it small.** Before adding a helper, ask why it does not belong to any one data source. Current modules: | Module | Contents | Why it is in `common/` | | --- | --- | --- | | `county_zonal_stats.py` | Raster windows, the 180th-meridian split, area-weighted class shares | Used by any raster-based metric | | `counties.py` | County polygon loading, FIPS normalization, the state FIPS table | County identity is shared by nearly every pipeline | | `koppen_legend.py` | The Köppen code map and legend loader | Temporary: also used by `build_county_climate_data.py`; moves next to the Köppen code in Phase 3 | **Rationale.** Helper functions make each step of a computation explicit, avoid repeated code, and can be tested separately ([Brown CSCI 0111, "Helper Functions"](https://cs.brown.edu/courses/csci0111/fall2018/lectures/helper-functions.html)). Shared helper folders, however, tend to lose cohesion and collect unrelated code; the recommended alternative is to keep code with the part of the system it belongs to, allowing a shared folder only if it stays small and documented ([Helpers and Utils Folders in Software Architecture](https://dev.to/knzt/helpers-and-utils-folders-in-software-architecture-3f8h)). The rules above follow both: shared functions, organized by topic and limited to code that crosses data sources. Other shared code moves when its filter is reviewed, so each move is tested alongside that filter. A 2026-09-13 survey found 19 functions with identical copies in several scripts and 19 with copies that have drifted apart. Most identical copies are NSRDB helpers, which belong in an NSRDB module rather than `common/`; `read_csv_rows` (4 identical copies in `apply_*` scripts) is cross-source. Drifted copies need a decision on which version is correct before merging. Notable drifts: `summarize_county_gridmet_humidity.py` has its own county loader and FIPS normalizer, and the state FIPS table is also copied in `build_county_representative_points.py`, `summarize_county_gridmet_humidity.py`, and `request_nsrdb_county_polygon_archives.py`. ### Phase 3 — Assemble and restructure (after all 12 filters are clean) - Assemble script that builds `climate-data.csv` from `data/metrics/`. - Populate `metric_sources.json`; remove the per-row `source` column. - Move audit columns out of the app CSV. - Point the app's Sources panel at `metric_sources.json`. - Retire or rewrite `build_county_climate_data.py` as per-metric builders. - Move `common/koppen_legend.py` next to the Köppen code once nothing outside Köppen imports it. ### Phase 4 — Runner and reproduction guide - `scripts/pipeline.py` runner with `--only` and `--skip-download`. - "Reproducing the data" guide organized by the tiers in section 3. - Pinned requirements. - End-to-end smoke test on a small synthetic county fixture. - Organize scripts by data source (`noaa/`, `gridmet/`, `nsrdb/`, `koppen/`), each holding its own helpers, with `common/` keeping only cross-source code. Scripts in subfolders are run through the runner or as modules (`python -m`), and the README and data-source commands are updated to match. ## 5. Per-filter review checklist For each filter: 1. Verify the calculation against the source data and document findings. 2. Decide any rule or method changes with the project owner. 3. Record the adopted definition in `filter-calculations.md`. 4. Implement the calculation, writing `data/metrics/.csv`. 5. Add a single-column apply step for the current CSV. 6. Update the rules in `check_climate_data.py`. 7. Add or update unit tests. 8. Apply to the CSV, run `check_climate_data.py`, and compare changed counties against expectations. 9. Update the app if the value set or display changes. ## 6. Filter tracker Known issues come from `filter-calculations.md` ("Calculation review findings") and this review; none beyond Köppen have been investigated yet. | # | Filter | Status | Known issues to review | | --- | --- | --- | --- | | 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | See tasks below | | 2 | Annual avg temperature | Not started | Months weighted equally (finding 1); touched-cell aggregation (finding 2) | | 3 | Diurnal temperature range | Not started | Lexington, VA (51678) blank, while heat-index days use Rockbridge County as a proxy | | 4 | Extreme temperature days | Not started | Partial years not normalized (finding 5); depends on retired percentile thresholds (finding 6); Lexington, VA blank | | 5 | 90 °F+ heat-index days | Not started | Daily-extrema proxy (finding 7); permissive year completeness (finding 8) | | 6 | Annual precipitation | Not started | Partial-year sums accepted (finding 3); touched-cell aggregation (finding 2) | | 7 | Seasonality index | Not started | Touched-cell aggregation (finding 2) | | 8 | Wettest month | Not started | Computed in both the base build and the precipitation-month script | | 9 | Driest month | Not started | Same as wettest month | | 10 | Summer specific humidity | Not started | Cell-center cos(latitude) aggregation differs from other metrics (finding 11) | | 11 | Solar GHI | Not started | Finalized by the extreme-temperature apply script; hourly, 365-day assumption (finding 9) | | 12 | Clear-sky GHI reduction | Not started | Mean of ratios rather than energy totals (finding 10) | ### Köppen-Geiger tasks Adopted rule: a county is predominantly its top class if and only if that class covers at least 50% of the county's land and leads the runner-up by at least 5 percentage points; otherwise it is Mixed climate. Expected result for the 50 states and DC: 3,010 predominant, 133 Mixed. - [x] Investigate low-majority counties and adopt the rule. - [x] Windowed raster reads and 180th-meridian split. - [x] Area-weighted class shares (16 × 16 sub-cells per raster cell, scaled by cos(latitude)) in `scripts/common/county_zonal_stats.py`. - [x] Apply the 50% / 5-point rule (`scripts/build_county_koppen_metric.py`; counties with no valid cells are left blank). - [x] Run the builder to write `data/metrics/koppen.csv` and confirm the expected 3,010 predominant / 133 Mixed (2026-09-13). - [x] `koppenZone`-only apply step (`scripts/apply_koppen_metric_to_climate_data.py`, with `--dry-run`). - [x] Apply to `data/climate-data.csv` (2026-09-13; 142 counties changed to Mixed, with `koppenPrimaryClass` and `koppenSecondaryClass` added). - [x] Allow `Mixed` in `check_climate_data.py`. - [x] Add a Mixed climate category to `app.js`, drawn as stripes of the county's top two classes; see [koppen-mixed-display-plan.md](koppen-mixed-display-plan.md). - [x] Replace the plurality description in `filter-calculations.md` §1 and mark review findings 2 and 4 resolved for Köppen (2026-09-14). - [x] Update the Köppen descriptions and script lists in `README.md` and `scripts/county_data_sources.md` (2026-09-14). - [x] Tests for shares, the rule, boundary cases, and the apply step (`tests/test_koppen_metric.py`). ## 7. Guardrails until Phase 3 - **Do not rerun `build_county_climate_data.py`** against `data/climate-data.csv`. It would drop 9 live columns, restore `extremeDays`, and overwrite the Mixed classification in `koppenZone`. - Run `check_climate_data.py` after every apply step. - Change one filter at a time, and compare its before and after values. ## 8. Open decisions | Decision | Options | Needed by | | --- | --- | --- | | Committing large intermediates | Commit metric files only, or also source summaries | Phase 3 | ### Decided - **Köppen audit columns (2026-09-12):** `koppen.csv` stores the top class and share and the runner-up class and share alongside `koppenZone`. - **Köppen no-data fallback (2026-09-12):** a county with no valid raster cells is left blank, not assigned `Cfa`. - **Applying Köppen to the app CSV (2026-09-12):** wait until the app supports the Mixed class. Done 2026-09-13. - **Metric files (2026-09-13):** one CSV per metric under `data/metrics/`, starting with `koppen.csv`. - **Mixed climate display (2026-09-13):** diagonal stripes of each Mixed county's top two classes; see [koppen-mixed-display-plan.md](koppen-mixed-display-plan.md). - **Puerto Rico (2026-09-13):** off the map and out of every filter. The app already drops state FIPS 72; the data files keep the rows. - **Shared helpers (2026-09-13):** `scripts/common/` holds only code used by more than one data source, one topic per module; code shared within one data source stays with that source. Duplicates move during their own filter's review.