Close gaps found in a review of the documentation: - Track data/metrics/ and data/metric_sources.json in git so the data checker passes on a fresh clone (finding 23; decision logged) - State the Alaska and Hawaii coverage gap in the README limitations and extend finding 12 - File findings 24-26: the data-sources doc lacks gridMET and several pipeline commands; wettest/driest month are computed twice; solar GHI is written by the extreme-temperature apply step - Add the stale "fallback values" note to filter 2's tasks Tidy the document system: - Add docs/reviews/README.md with the numbering rules and a finding index - Rename koppen-mixed-display-plan.md to koppen-mixed-display.md and fix its stale Puerto Rico and "stage 5" text - Add the precipitation-month step to the README enrichment list - Describe the Current method / Previous method pattern in plan section 5 - Add CLAUDE.md with the project guardrails and doc layout Format filter-calculations.md so it renders on GitHub and in VS Code: inline math uses $...$, ranges use en dashes, and implementation references name functions instead of line numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
12 KiB
Data Pipeline Improvement Plan
This plan describes how the county data pipeline will move from scripts that edit one shared CSV in place to per-metric outputs assembled into the app CSV. It is a working reference for the filter-by-filter review. Calculation details for each filter live in filter-calculations.md; findings and tasks live in reviews/, one file per filter plus 00-cross-filter.md; open and past decisions live in decisions.md.
Started: 2026-09-12
Guiding decision: review and fix each of the 12 filters one at a time,
confirm each works on its own, and restructure data/climate-data.csv only
after all filters are clean. No large rewrite happens up front.
1. Current pipeline
Where each filter comes from
| Filter | Written into climate-data.csv by |
Upstream scripts |
|---|---|---|
| Köppen-Geiger class (plus the two stripe-class columns) | apply_koppen_metric_to_climate_data.py |
build_county_koppen_metric.py → data/metrics/koppen.csv |
| Annual avg temperature | build_county_climate_data.py |
— |
| Annual precipitation | build_county_climate_data.py |
— |
| Seasonality index | build_county_climate_data.py |
— |
| Wettest / driest month | Base build, then overwritten by apply_precipitation_month_metrics_to_climate_data.py |
— |
| Diurnal temperature range | apply_diurnal_temperature_range_to_climate_data.py |
build_county_diurnal_temperature_range.py |
| Extreme temperature days | apply_locally_extreme_metric_to_climate_data.py |
build_county_locally_extreme_data.py |
| Summer specific humidity | apply_gridmet_humidity_metric_to_climate_data.py |
download_gridmet_data.py → summarize_county_gridmet_humidity.py |
| 90 °F+ heat-index days (plus 2 source-FIPS columns) | apply_gridmet_humidity_metric_to_climate_data.py |
Same as summer humidity |
| Solar GHI | Base build (optional), then replaced by apply_locally_extreme_metric_to_climate_data.py |
Point: build_county_representative_points.py → fetch_nsrdb_representative_point_ghi.py. Polygon: request_nsrdb_county_polygon_ghi_archives.py → download_nsrdb_county_polygon_ghi_archives.py → summarize_nsrdb_county_polygon_archives.py |
| Clear-sky GHI reduction | apply_nsrdb_cloud_metric_to_climate_data.py |
Point: fetch_nsrdb_representative_point_cloud_metrics.py. Polygon: request_nsrdb_county_polygon_cloud_archives.py → download_nsrdb_county_polygon_cloud_archives.py → summarize_nsrdb_county_polygon_cloud_archives.py |
Supporting scripts: request_nsrdb_county_polygon_archives.py and
download_nsrdb_county_polygon_archives.py are the shared engines behind the
GHI and cloud wrappers; rebuild_nsrdb_representative_point_ghi_summary.py
rebuilds the point GHI summary from cache; check_climate_data.py validates the
final CSV. Shared helpers live in scripts/common/ (Phase 2). The base build
still writes an old largest-share koppenZone, so the Köppen apply step must
run after it.
Current full-rebuild order
build_county_climate_data.pybuild_county_koppen_metric.py→apply_koppen_metric_to_climate_data.pyapply_precipitation_month_metrics_to_climate_data.pybuild_county_locally_extreme_data.py→apply_locally_extreme_metric_to_climate_data.pybuild_county_diurnal_temperature_range.py→apply_diurnal_temperature_range_to_climate_data.pysummarize_county_gridmet_humidity.py→apply_gridmet_humidity_metric_to_climate_data.pyapply_nsrdb_cloud_metric_to_climate_data.py
Problems
- Rerunning a step can destroy data. The base build writes 12 columns,
including the retired
extremeDays. Later scripts delete, overwrite, or add columns until the live CSV has 20. Rerunning the base build drops 9 live columns (koppenPrimaryClass,koppenSecondaryClass,avgDiurnalTempRangeF,absoluteExtremeDays,clearSkyGhiReductionIndex,avgSummerSpecificHumidityGKg,humidHeatDays,humidHeatSourceFips,humidHeatFipsAdjustment), restoresextremeDays, and rewriteskoppenZonewith the old largest-share method. - Order is implicit. The sequence lives in the README, in
scripts/county_data_sources.md, and in each script's assumptions. - Column ownership is unclear. GHI is finalized by the extreme-temperature apply script; wettest/driest month are computed in two places. See findings 25 and 26 in reviews/00-cross-filter.md.
- County aggregation is inconsistent. See finding 11 in reviews/00-cross-filter.md.
- No single entry point or final check. A new user must piece together about 20 scripts, several large downloads, and an NSRDB API key.
2. Target design
- One metric, one file. Each metric pipeline writes a county-level file
under
data/metrics/, for exampledata/metrics/koppen.csv, containingcountyFips, the app value, and any audit columns for that metric. - One assemble step. A single script joins the metric files into
data/climate-data.csv, usingdata/metric_sources.jsonfor the column list and per-metric source notes, then runscheck_climate_data.py.- Run order no longer matters; rerunning one metric cannot damage others.
- Every column has exactly one owner.
- The per-row
sourcecolumn moves intometric_sources.json. - Audit columns stay in the metric files rather than the app CSV.
- One shared county-aggregation module. Area-weighted zonal statistics, including the 180th-meridian split, used by every raster-based metric.
- One runner. For example
python scripts/pipeline.py --only koppen --skip-download, with stages for fetch, build metrics, assemble, and check. Cached downloads are reused by default.
3. Reproduction tiers
The "Reproducing the data" guide (Phase 4) will be organized by how deep a user needs to go:
| Tier | What the user does | Needs |
|---|---|---|
| 1. Run the app | .\serve.ps1 with the committed CSV |
Nothing else |
| 2. Reassemble | Rebuild climate-data.csv from committed metric files |
Python environment only |
| 3. Regenerate one metric | Download one source, rebuild one metric file, reassemble | That metric's source data |
| 4. Full rebuild | Everything | All sources; NSRDB API key; large downloads (the NOAA monthly temperature file alone is about 5 GB) and hours of paced NSRDB requests |
The guide will list each dataset's size, download location, API-key needs, and approximate run time, and the Python requirements will be pinned.
4. Roadmap
Phase 0 — Groundwork (done)
scripts/check_climate_data.pyvalidates the app CSV (8 checks) with tests intests/test_check_climate_data.py.data/metric_sources.jsoncreated as an empty skeleton.- Köppen raster reads use a padded window per county, and polygons that
cross the 180th meridian are split (
split_at_antimeridianinscripts/common/county_zonal_stats.py); output verified identical for all 3,221 counties; tests intests/test_koppen_antimeridian.py.
Phase 1 — Filter-by-filter review (in progress)
Each filter goes through the checklist in section 5. Each fix delivers that
metric's own file in data/metrics/ plus a single-column apply step, so the
existing CSV keeps working until Phase 3.
Phase 2 — Shared helpers in scripts/common/
Rules, current modules, and rationale are in scripts/common/README.md. Duplicated helpers found by the 2026-09-13 survey are finding 14 in reviews/00-cross-filter.md.
Phase 3 — Assemble and restructure (after all 12 filters are clean)
- Assemble script that builds
climate-data.csvfromdata/metrics/. - Populate
metric_sources.json; remove the per-rowsourcecolumn. - Move audit columns out of the app CSV.
- Point the app's Sources panel at
metric_sources.json. - Retire or rewrite
build_county_climate_data.pyas per-metric builders. - Move
common/koppen_legend.pynext to the Köppen code once nothing outside Köppen imports it.
Phase 4 — Runner and reproduction guide
scripts/pipeline.pyrunner with--onlyand--skip-download.- "Reproducing the data" guide organized by the tiers in section 3.
- Pinned requirements.
- End-to-end smoke test on a small synthetic county fixture.
- Organize scripts by data source (
noaa/,gridmet/,nsrdb/,koppen/), each holding its own helpers, withcommon/keeping only cross-source code. Scripts in subfolders are run through the runner or as modules (python -m), and the README and data-source commands are updated to match.
5. Per-filter review checklist
For each filter:
- Verify the calculation against the source data and document findings in the
filter's review file under
docs/reviews/. - Decide any rule or method changes with the project owner, and record them in
decisions.md. - Record the adopted definition in
filter-calculations.md. Until the new values are applied, keep the old definition below it under "Current method". - Implement the calculation, writing
data/metrics/<metric>.csv. - Add a single-column apply step for the current CSV.
- Update the rules in
check_climate_data.py. - Add or update unit tests.
- Apply to the CSV, run
check_climate_data.py, and compare changed counties against expectations. Then rename "Current method" to "Previous method" infilter-calculations.md, as §1 does. - Update the app if the value set or display changes.
6. Filter tracker
Each filter's findings, decisions, and tasks are in its review file. Findings that affect more than one filter are in 00-cross-filter.md, and reviews/README.md indexes every finding number.
| # | Filter | Status | Review |
|---|---|---|---|
| 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | 01-koppen.md |
| 2 | Annual avg temperature | In progress: method decided (2026-09-15) | 02-annual-avg-temperature.md |
| 3 | Diurnal temperature range | Not started | 03-diurnal-temperature-range.md |
| 4 | Extreme temperature days | Not started | 04-extreme-temperature-days.md |
| 5 | 90 °F+ heat-index days | Not started | 05-heat-index-days.md |
| 6 | Annual precipitation | Not started | 06-annual-precipitation.md |
| 7 | Seasonality index | Not started | 07-seasonality-index.md |
| 8 | Wettest month | Not started | 08-wettest-month.md |
| 9 | Driest month | Not started | 09-driest-month.md |
| 10 | Summer specific humidity | Not started | 10-summer-specific-humidity.md |
| 11 | Solar GHI | Not started | 11-solar-ghi.md |
| 12 | Clear-sky GHI reduction | Not started | 12-clear-sky-ghi-reduction.md |
7. Guardrails until Phase 3
- Do not rerun
build_county_climate_data.pyagainstdata/climate-data.csv. It would drop 9 live columns, restoreextremeDays, and overwrite the Mixed classification inkoppenZone. - Run
check_climate_data.pyafter every apply step. - Change one filter at a time, and compare its before and after values.