Files
Climate-Mood-Analysis/docs/pipeline-plan.md
T
KnouandClaude Opus 5 867a07cecb Refine project documentation and track metric files
Close gaps found in a review of the documentation:
- Track data/metrics/ and data/metric_sources.json in git so the data
  checker passes on a fresh clone (finding 23; decision logged)
- State the Alaska and Hawaii coverage gap in the README limitations and
  extend finding 12
- File findings 24-26: the data-sources doc lacks gridMET and several
  pipeline commands; wettest/driest month are computed twice; solar GHI
  is written by the extreme-temperature apply step
- Add the stale "fallback values" note to filter 2's tasks

Tidy the document system:
- Add docs/reviews/README.md with the numbering rules and a finding index
- Rename koppen-mixed-display-plan.md to koppen-mixed-display.md and fix
  its stale Puerto Rico and "stage 5" text
- Add the precipitation-month step to the README enrichment list
- Describe the Current method / Previous method pattern in plan section 5
- Add CLAUDE.md with the project guardrails and doc layout

Format filter-calculations.md so it renders on GitHub and in VS Code:
inline math uses $...$, ranges use en dashes, and implementation
references name functions instead of line numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 17:31:26 -04:00

203 lines
12 KiB
Markdown

# Data Pipeline Improvement Plan
This plan describes how the county data pipeline will move from scripts that
edit one shared CSV in place to per-metric outputs assembled into the app CSV.
It is a working reference for the filter-by-filter review. Calculation details
for each filter live in [filter-calculations.md](filter-calculations.md);
findings and tasks live in [reviews/](reviews/), one file per filter plus
[00-cross-filter.md](reviews/00-cross-filter.md); open and past decisions live
in [decisions.md](decisions.md).
**Started:** 2026-09-12
**Guiding decision:** review and fix each of the 12 filters one at a time,
confirm each works on its own, and restructure `data/climate-data.csv` only
after all filters are clean. No large rewrite happens up front.
## 1. Current pipeline
### Where each filter comes from
| Filter | Written into `climate-data.csv` by | Upstream scripts |
| --- | --- | --- |
| Köppen-Geiger class (plus the two stripe-class columns) | `apply_koppen_metric_to_climate_data.py` | `build_county_koppen_metric.py` → `data/metrics/koppen.csv` |
| Annual avg temperature | `build_county_climate_data.py` | — |
| Annual precipitation | `build_county_climate_data.py` | — |
| Seasonality index | `build_county_climate_data.py` | — |
| Wettest / driest month | Base build, then overwritten by `apply_precipitation_month_metrics_to_climate_data.py` | — |
| Diurnal temperature range | `apply_diurnal_temperature_range_to_climate_data.py` | `build_county_diurnal_temperature_range.py` |
| Extreme temperature days | `apply_locally_extreme_metric_to_climate_data.py` | `build_county_locally_extreme_data.py` |
| Summer specific humidity | `apply_gridmet_humidity_metric_to_climate_data.py` | `download_gridmet_data.py` → `summarize_county_gridmet_humidity.py` |
| 90 °F+ heat-index days (plus 2 source-FIPS columns) | `apply_gridmet_humidity_metric_to_climate_data.py` | Same as summer humidity |
| Solar GHI | Base build (optional), then replaced by `apply_locally_extreme_metric_to_climate_data.py` | Point: `build_county_representative_points.py` → `fetch_nsrdb_representative_point_ghi.py`. Polygon: `request_nsrdb_county_polygon_ghi_archives.py` → `download_nsrdb_county_polygon_ghi_archives.py` → `summarize_nsrdb_county_polygon_archives.py` |
| Clear-sky GHI reduction | `apply_nsrdb_cloud_metric_to_climate_data.py` | Point: `fetch_nsrdb_representative_point_cloud_metrics.py`. Polygon: `request_nsrdb_county_polygon_cloud_archives.py` → `download_nsrdb_county_polygon_cloud_archives.py` → `summarize_nsrdb_county_polygon_cloud_archives.py` |
Supporting scripts: `request_nsrdb_county_polygon_archives.py` and
`download_nsrdb_county_polygon_archives.py` are the shared engines behind the
GHI and cloud wrappers; `rebuild_nsrdb_representative_point_ghi_summary.py`
rebuilds the point GHI summary from cache; `check_climate_data.py` validates the
final CSV. Shared helpers live in `scripts/common/` (Phase 2). The base build
still writes an old largest-share `koppenZone`, so the Köppen apply step must
run after it.
### Current full-rebuild order
1. `build_county_climate_data.py`
2. `build_county_koppen_metric.py` → `apply_koppen_metric_to_climate_data.py`
3. `apply_precipitation_month_metrics_to_climate_data.py`
4. `build_county_locally_extreme_data.py` → `apply_locally_extreme_metric_to_climate_data.py`
5. `build_county_diurnal_temperature_range.py` → `apply_diurnal_temperature_range_to_climate_data.py`
6. `summarize_county_gridmet_humidity.py` → `apply_gridmet_humidity_metric_to_climate_data.py`
7. `apply_nsrdb_cloud_metric_to_climate_data.py`
### Problems
1. **Rerunning a step can destroy data.** The base build writes 12 columns,
including the retired `extremeDays`. Later scripts delete, overwrite, or add
columns until the live CSV has 20. Rerunning the base build drops
9 live columns (`koppenPrimaryClass`, `koppenSecondaryClass`,
`avgDiurnalTempRangeF`, `absoluteExtremeDays`,
`clearSkyGhiReductionIndex`, `avgSummerSpecificHumidityGKg`,
`humidHeatDays`, `humidHeatSourceFips`, `humidHeatFipsAdjustment`),
restores `extremeDays`, and rewrites `koppenZone` with the old
largest-share method.
2. **Order is implicit.** The sequence lives in the README, in
`scripts/county_data_sources.md`, and in each script's assumptions.
3. **Column ownership is unclear.** GHI is finalized by the extreme-temperature
apply script; wettest/driest month are computed in two places. See findings
25 and 26 in [reviews/00-cross-filter.md](reviews/00-cross-filter.md).
4. **County aggregation is inconsistent.** See finding 11 in
[reviews/00-cross-filter.md](reviews/00-cross-filter.md).
5. **No single entry point or final check.** A new user must piece together
about 20 scripts, several large downloads, and an NSRDB API key.
## 2. Target design
1. **One metric, one file.** Each metric pipeline writes a county-level file
under `data/metrics/`, for example `data/metrics/koppen.csv`, containing
`countyFips`, the app value, and any audit columns for that metric.
2. **One assemble step.** A single script joins the metric files into
`data/climate-data.csv`, using `data/metric_sources.json` for the column list
and per-metric source notes, then runs `check_climate_data.py`.
- Run order no longer matters; rerunning one metric cannot damage others.
- Every column has exactly one owner.
- The per-row `source` column moves into `metric_sources.json`.
- Audit columns stay in the metric files rather than the app CSV.
3. **One shared county-aggregation module.** Area-weighted zonal statistics,
including the 180th-meridian split, used by every raster-based metric.
4. **One runner.** For example
`python scripts/pipeline.py --only koppen --skip-download`, with stages for
fetch, build metrics, assemble, and check. Cached downloads are reused by
default.
## 3. Reproduction tiers
The "Reproducing the data" guide (Phase 4) will be organized by how deep a
user needs to go:
| Tier | What the user does | Needs |
| --- | --- | --- |
| 1. Run the app | `.\serve.ps1` with the committed CSV | Nothing else |
| 2. Reassemble | Rebuild `climate-data.csv` from committed metric files | Python environment only |
| 3. Regenerate one metric | Download one source, rebuild one metric file, reassemble | That metric's source data |
| 4. Full rebuild | Everything | All sources; NSRDB API key; large downloads (the NOAA monthly temperature file alone is about 5 GB) and hours of paced NSRDB requests |
The guide will list each dataset's size, download location, API-key needs, and
approximate run time, and the Python requirements will be pinned.
## 4. Roadmap
### Phase 0 — Groundwork (done)
- [x] `scripts/check_climate_data.py` validates the app CSV (8 checks) with
tests in `tests/test_check_climate_data.py`.
- [x] `data/metric_sources.json` created as an empty skeleton.
- [x] Köppen raster reads use a padded window per county, and polygons that
cross the 180th meridian are split (`split_at_antimeridian` in
`scripts/common/county_zonal_stats.py`); output verified identical for
all 3,221 counties; tests in `tests/test_koppen_antimeridian.py`.
### Phase 1 — Filter-by-filter review (in progress)
Each filter goes through the checklist in section 5. Each fix delivers that
metric's own file in `data/metrics/` plus a single-column apply step, so the
existing CSV keeps working until Phase 3.
### Phase 2 — Shared helpers in `scripts/common/`
Rules, current modules, and rationale are in
[scripts/common/README.md](../scripts/common/README.md). Duplicated helpers
found by the 2026-09-13 survey are finding 14 in
[reviews/00-cross-filter.md](reviews/00-cross-filter.md).
### Phase 3 — Assemble and restructure (after all 12 filters are clean)
- Assemble script that builds `climate-data.csv` from `data/metrics/`.
- Populate `metric_sources.json`; remove the per-row `source` column.
- Move audit columns out of the app CSV.
- Point the app's Sources panel at `metric_sources.json`.
- Retire or rewrite `build_county_climate_data.py` as per-metric builders.
- Move `common/koppen_legend.py` next to the Köppen code once nothing outside
Köppen imports it.
### Phase 4 — Runner and reproduction guide
- `scripts/pipeline.py` runner with `--only` and `--skip-download`.
- "Reproducing the data" guide organized by the tiers in section 3.
- Pinned requirements.
- End-to-end smoke test on a small synthetic county fixture.
- Organize scripts by data source (`noaa/`, `gridmet/`, `nsrdb/`, `koppen/`),
each holding its own helpers, with `common/` keeping only cross-source code.
Scripts in subfolders are run through the runner or as modules
(`python -m`), and the README and data-source commands are updated to match.
## 5. Per-filter review checklist
For each filter:
1. Verify the calculation against the source data and document findings in the
filter's review file under `docs/reviews/`.
2. Decide any rule or method changes with the project owner, and record them in
`decisions.md`.
3. Record the adopted definition in `filter-calculations.md`. Until the new
values are applied, keep the old definition below it under "Current
method".
4. Implement the calculation, writing `data/metrics/<metric>.csv`.
5. Add a single-column apply step for the current CSV.
6. Update the rules in `check_climate_data.py`.
7. Add or update unit tests.
8. Apply to the CSV, run `check_climate_data.py`, and compare changed counties
against expectations. Then rename "Current method" to "Previous method" in
`filter-calculations.md`, as §1 does.
9. Update the app if the value set or display changes.
## 6. Filter tracker
Each filter's findings, decisions, and tasks are in its review file. Findings
that affect more than one filter are in
[00-cross-filter.md](reviews/00-cross-filter.md), and
[reviews/README.md](reviews/README.md) indexes every finding number.
| # | Filter | Status | Review |
| --- | --- | --- | --- |
| 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | [01-koppen.md](reviews/01-koppen.md) |
| 2 | Annual avg temperature | In progress: method decided (2026-09-15) | [02-annual-avg-temperature.md](reviews/02-annual-avg-temperature.md) |
| 3 | Diurnal temperature range | Not started | [03-diurnal-temperature-range.md](reviews/03-diurnal-temperature-range.md) |
| 4 | Extreme temperature days | Not started | [04-extreme-temperature-days.md](reviews/04-extreme-temperature-days.md) |
| 5 | 90 °F+ heat-index days | Not started | [05-heat-index-days.md](reviews/05-heat-index-days.md) |
| 6 | Annual precipitation | Not started | [06-annual-precipitation.md](reviews/06-annual-precipitation.md) |
| 7 | Seasonality index | Not started | [07-seasonality-index.md](reviews/07-seasonality-index.md) |
| 8 | Wettest month | Not started | [08-wettest-month.md](reviews/08-wettest-month.md) |
| 9 | Driest month | Not started | [09-driest-month.md](reviews/09-driest-month.md) |
| 10 | Summer specific humidity | Not started | [10-summer-specific-humidity.md](reviews/10-summer-specific-humidity.md) |
| 11 | Solar GHI | Not started | [11-solar-ghi.md](reviews/11-solar-ghi.md) |
| 12 | Clear-sky GHI reduction | Not started | [12-clear-sky-ghi-reduction.md](reviews/12-clear-sky-ghi-reduction.md) |
## 7. Guardrails until Phase 3
- **Do not rerun `build_county_climate_data.py`** against
`data/climate-data.csv`. It would drop 9 live columns, restore
`extremeDays`, and overwrite the Mixed classification in `koppenZone`.
- Run `check_climate_data.py` after every apply step.
- Change one filter at a time, and compare its before and after values.