Close gaps found in a review of the documentation: - Track data/metrics/ and data/metric_sources.json in git so the data checker passes on a fresh clone (finding 23; decision logged) - State the Alaska and Hawaii coverage gap in the README limitations and extend finding 12 - File findings 24-26: the data-sources doc lacks gridMET and several pipeline commands; wettest/driest month are computed twice; solar GHI is written by the extreme-temperature apply step - Add the stale "fallback values" note to filter 2's tasks Tidy the document system: - Add docs/reviews/README.md with the numbering rules and a finding index - Rename koppen-mixed-display-plan.md to koppen-mixed-display.md and fix its stale Puerto Rico and "stage 5" text - Add the precipitation-month step to the README enrichment list - Describe the Current method / Previous method pattern in plan section 5 - Add CLAUDE.md with the project guardrails and doc layout Format filter-calculations.md so it renders on GitHub and in VS Code: inline math uses $...$, ranges use en dashes, and implementation references name functions instead of line numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
203 lines
12 KiB
Markdown
203 lines
12 KiB
Markdown
# Data Pipeline Improvement Plan
|
|
|
|
This plan describes how the county data pipeline will move from scripts that
|
|
edit one shared CSV in place to per-metric outputs assembled into the app CSV.
|
|
It is a working reference for the filter-by-filter review. Calculation details
|
|
for each filter live in [filter-calculations.md](filter-calculations.md);
|
|
findings and tasks live in [reviews/](reviews/), one file per filter plus
|
|
[00-cross-filter.md](reviews/00-cross-filter.md); open and past decisions live
|
|
in [decisions.md](decisions.md).
|
|
|
|
**Started:** 2026-09-12
|
|
|
|
**Guiding decision:** review and fix each of the 12 filters one at a time,
|
|
confirm each works on its own, and restructure `data/climate-data.csv` only
|
|
after all filters are clean. No large rewrite happens up front.
|
|
|
|
## 1. Current pipeline
|
|
|
|
### Where each filter comes from
|
|
|
|
| Filter | Written into `climate-data.csv` by | Upstream scripts |
|
|
| --- | --- | --- |
|
|
| Köppen-Geiger class (plus the two stripe-class columns) | `apply_koppen_metric_to_climate_data.py` | `build_county_koppen_metric.py` → `data/metrics/koppen.csv` |
|
|
| Annual avg temperature | `build_county_climate_data.py` | — |
|
|
| Annual precipitation | `build_county_climate_data.py` | — |
|
|
| Seasonality index | `build_county_climate_data.py` | — |
|
|
| Wettest / driest month | Base build, then overwritten by `apply_precipitation_month_metrics_to_climate_data.py` | — |
|
|
| Diurnal temperature range | `apply_diurnal_temperature_range_to_climate_data.py` | `build_county_diurnal_temperature_range.py` |
|
|
| Extreme temperature days | `apply_locally_extreme_metric_to_climate_data.py` | `build_county_locally_extreme_data.py` |
|
|
| Summer specific humidity | `apply_gridmet_humidity_metric_to_climate_data.py` | `download_gridmet_data.py` → `summarize_county_gridmet_humidity.py` |
|
|
| 90 °F+ heat-index days (plus 2 source-FIPS columns) | `apply_gridmet_humidity_metric_to_climate_data.py` | Same as summer humidity |
|
|
| Solar GHI | Base build (optional), then replaced by `apply_locally_extreme_metric_to_climate_data.py` | Point: `build_county_representative_points.py` → `fetch_nsrdb_representative_point_ghi.py`. Polygon: `request_nsrdb_county_polygon_ghi_archives.py` → `download_nsrdb_county_polygon_ghi_archives.py` → `summarize_nsrdb_county_polygon_archives.py` |
|
|
| Clear-sky GHI reduction | `apply_nsrdb_cloud_metric_to_climate_data.py` | Point: `fetch_nsrdb_representative_point_cloud_metrics.py`. Polygon: `request_nsrdb_county_polygon_cloud_archives.py` → `download_nsrdb_county_polygon_cloud_archives.py` → `summarize_nsrdb_county_polygon_cloud_archives.py` |
|
|
|
|
Supporting scripts: `request_nsrdb_county_polygon_archives.py` and
|
|
`download_nsrdb_county_polygon_archives.py` are the shared engines behind the
|
|
GHI and cloud wrappers; `rebuild_nsrdb_representative_point_ghi_summary.py`
|
|
rebuilds the point GHI summary from cache; `check_climate_data.py` validates the
|
|
final CSV. Shared helpers live in `scripts/common/` (Phase 2). The base build
|
|
still writes an old largest-share `koppenZone`, so the Köppen apply step must
|
|
run after it.
|
|
|
|
### Current full-rebuild order
|
|
|
|
1. `build_county_climate_data.py`
|
|
2. `build_county_koppen_metric.py` → `apply_koppen_metric_to_climate_data.py`
|
|
3. `apply_precipitation_month_metrics_to_climate_data.py`
|
|
4. `build_county_locally_extreme_data.py` → `apply_locally_extreme_metric_to_climate_data.py`
|
|
5. `build_county_diurnal_temperature_range.py` → `apply_diurnal_temperature_range_to_climate_data.py`
|
|
6. `summarize_county_gridmet_humidity.py` → `apply_gridmet_humidity_metric_to_climate_data.py`
|
|
7. `apply_nsrdb_cloud_metric_to_climate_data.py`
|
|
|
|
### Problems
|
|
|
|
1. **Rerunning a step can destroy data.** The base build writes 12 columns,
|
|
including the retired `extremeDays`. Later scripts delete, overwrite, or add
|
|
columns until the live CSV has 20. Rerunning the base build drops
|
|
9 live columns (`koppenPrimaryClass`, `koppenSecondaryClass`,
|
|
`avgDiurnalTempRangeF`, `absoluteExtremeDays`,
|
|
`clearSkyGhiReductionIndex`, `avgSummerSpecificHumidityGKg`,
|
|
`humidHeatDays`, `humidHeatSourceFips`, `humidHeatFipsAdjustment`),
|
|
restores `extremeDays`, and rewrites `koppenZone` with the old
|
|
largest-share method.
|
|
2. **Order is implicit.** The sequence lives in the README, in
|
|
`scripts/county_data_sources.md`, and in each script's assumptions.
|
|
3. **Column ownership is unclear.** GHI is finalized by the extreme-temperature
|
|
apply script; wettest/driest month are computed in two places. See findings
|
|
25 and 26 in [reviews/00-cross-filter.md](reviews/00-cross-filter.md).
|
|
4. **County aggregation is inconsistent.** See finding 11 in
|
|
[reviews/00-cross-filter.md](reviews/00-cross-filter.md).
|
|
5. **No single entry point or final check.** A new user must piece together
|
|
about 20 scripts, several large downloads, and an NSRDB API key.
|
|
|
|
## 2. Target design
|
|
|
|
1. **One metric, one file.** Each metric pipeline writes a county-level file
|
|
under `data/metrics/`, for example `data/metrics/koppen.csv`, containing
|
|
`countyFips`, the app value, and any audit columns for that metric.
|
|
2. **One assemble step.** A single script joins the metric files into
|
|
`data/climate-data.csv`, using `data/metric_sources.json` for the column list
|
|
and per-metric source notes, then runs `check_climate_data.py`.
|
|
- Run order no longer matters; rerunning one metric cannot damage others.
|
|
- Every column has exactly one owner.
|
|
- The per-row `source` column moves into `metric_sources.json`.
|
|
- Audit columns stay in the metric files rather than the app CSV.
|
|
3. **One shared county-aggregation module.** Area-weighted zonal statistics,
|
|
including the 180th-meridian split, used by every raster-based metric.
|
|
4. **One runner.** For example
|
|
`python scripts/pipeline.py --only koppen --skip-download`, with stages for
|
|
fetch, build metrics, assemble, and check. Cached downloads are reused by
|
|
default.
|
|
|
|
## 3. Reproduction tiers
|
|
|
|
The "Reproducing the data" guide (Phase 4) will be organized by how deep a
|
|
user needs to go:
|
|
|
|
| Tier | What the user does | Needs |
|
|
| --- | --- | --- |
|
|
| 1. Run the app | `.\serve.ps1` with the committed CSV | Nothing else |
|
|
| 2. Reassemble | Rebuild `climate-data.csv` from committed metric files | Python environment only |
|
|
| 3. Regenerate one metric | Download one source, rebuild one metric file, reassemble | That metric's source data |
|
|
| 4. Full rebuild | Everything | All sources; NSRDB API key; large downloads (the NOAA monthly temperature file alone is about 5 GB) and hours of paced NSRDB requests |
|
|
|
|
The guide will list each dataset's size, download location, API-key needs, and
|
|
approximate run time, and the Python requirements will be pinned.
|
|
|
|
## 4. Roadmap
|
|
|
|
### Phase 0 — Groundwork (done)
|
|
|
|
- [x] `scripts/check_climate_data.py` validates the app CSV (8 checks) with
|
|
tests in `tests/test_check_climate_data.py`.
|
|
- [x] `data/metric_sources.json` created as an empty skeleton.
|
|
- [x] Köppen raster reads use a padded window per county, and polygons that
|
|
cross the 180th meridian are split (`split_at_antimeridian` in
|
|
`scripts/common/county_zonal_stats.py`); output verified identical for
|
|
all 3,221 counties; tests in `tests/test_koppen_antimeridian.py`.
|
|
|
|
### Phase 1 — Filter-by-filter review (in progress)
|
|
|
|
Each filter goes through the checklist in section 5. Each fix delivers that
|
|
metric's own file in `data/metrics/` plus a single-column apply step, so the
|
|
existing CSV keeps working until Phase 3.
|
|
|
|
### Phase 2 — Shared helpers in `scripts/common/`
|
|
|
|
Rules, current modules, and rationale are in
|
|
[scripts/common/README.md](../scripts/common/README.md). Duplicated helpers
|
|
found by the 2026-09-13 survey are finding 14 in
|
|
[reviews/00-cross-filter.md](reviews/00-cross-filter.md).
|
|
|
|
### Phase 3 — Assemble and restructure (after all 12 filters are clean)
|
|
|
|
- Assemble script that builds `climate-data.csv` from `data/metrics/`.
|
|
- Populate `metric_sources.json`; remove the per-row `source` column.
|
|
- Move audit columns out of the app CSV.
|
|
- Point the app's Sources panel at `metric_sources.json`.
|
|
- Retire or rewrite `build_county_climate_data.py` as per-metric builders.
|
|
- Move `common/koppen_legend.py` next to the Köppen code once nothing outside
|
|
Köppen imports it.
|
|
|
|
### Phase 4 — Runner and reproduction guide
|
|
|
|
- `scripts/pipeline.py` runner with `--only` and `--skip-download`.
|
|
- "Reproducing the data" guide organized by the tiers in section 3.
|
|
- Pinned requirements.
|
|
- End-to-end smoke test on a small synthetic county fixture.
|
|
- Organize scripts by data source (`noaa/`, `gridmet/`, `nsrdb/`, `koppen/`),
|
|
each holding its own helpers, with `common/` keeping only cross-source code.
|
|
Scripts in subfolders are run through the runner or as modules
|
|
(`python -m`), and the README and data-source commands are updated to match.
|
|
|
|
## 5. Per-filter review checklist
|
|
|
|
For each filter:
|
|
|
|
1. Verify the calculation against the source data and document findings in the
|
|
filter's review file under `docs/reviews/`.
|
|
2. Decide any rule or method changes with the project owner, and record them in
|
|
`decisions.md`.
|
|
3. Record the adopted definition in `filter-calculations.md`. Until the new
|
|
values are applied, keep the old definition below it under "Current
|
|
method".
|
|
4. Implement the calculation, writing `data/metrics/<metric>.csv`.
|
|
5. Add a single-column apply step for the current CSV.
|
|
6. Update the rules in `check_climate_data.py`.
|
|
7. Add or update unit tests.
|
|
8. Apply to the CSV, run `check_climate_data.py`, and compare changed counties
|
|
against expectations. Then rename "Current method" to "Previous method" in
|
|
`filter-calculations.md`, as §1 does.
|
|
9. Update the app if the value set or display changes.
|
|
|
|
## 6. Filter tracker
|
|
|
|
Each filter's findings, decisions, and tasks are in its review file. Findings
|
|
that affect more than one filter are in
|
|
[00-cross-filter.md](reviews/00-cross-filter.md), and
|
|
[reviews/README.md](reviews/README.md) indexes every finding number.
|
|
|
|
| # | Filter | Status | Review |
|
|
| --- | --- | --- | --- |
|
|
| 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | [01-koppen.md](reviews/01-koppen.md) |
|
|
| 2 | Annual avg temperature | In progress: method decided (2026-09-15) | [02-annual-avg-temperature.md](reviews/02-annual-avg-temperature.md) |
|
|
| 3 | Diurnal temperature range | Not started | [03-diurnal-temperature-range.md](reviews/03-diurnal-temperature-range.md) |
|
|
| 4 | Extreme temperature days | Not started | [04-extreme-temperature-days.md](reviews/04-extreme-temperature-days.md) |
|
|
| 5 | 90 °F+ heat-index days | Not started | [05-heat-index-days.md](reviews/05-heat-index-days.md) |
|
|
| 6 | Annual precipitation | Not started | [06-annual-precipitation.md](reviews/06-annual-precipitation.md) |
|
|
| 7 | Seasonality index | Not started | [07-seasonality-index.md](reviews/07-seasonality-index.md) |
|
|
| 8 | Wettest month | Not started | [08-wettest-month.md](reviews/08-wettest-month.md) |
|
|
| 9 | Driest month | Not started | [09-driest-month.md](reviews/09-driest-month.md) |
|
|
| 10 | Summer specific humidity | Not started | [10-summer-specific-humidity.md](reviews/10-summer-specific-humidity.md) |
|
|
| 11 | Solar GHI | Not started | [11-solar-ghi.md](reviews/11-solar-ghi.md) |
|
|
| 12 | Clear-sky GHI reduction | Not started | [12-clear-sky-ghi-reduction.md](reviews/12-clear-sky-ghi-reduction.md) |
|
|
|
|
## 7. Guardrails until Phase 3
|
|
|
|
- **Do not rerun `build_county_climate_data.py`** against
|
|
`data/climate-data.csv`. It would drop 9 live columns, restore
|
|
`extremeDays`, and overwrite the Mixed classification in `koppenZone`.
|
|
- Run `check_climate_data.py` after every apply step.
|
|
- Change one filter at a time, and compare its before and after values.
|