Complete Köppen-Geiger filter review with Mixed climate class
Classify each county by area-weighted Köppen class shares: a county is predominantly its top class when that class covers at least 50% of its land and leads the runner-up by at least 5 percentage points; otherwise it is Mixed (133 of 3,143 counties in the 50 states and DC). - Add build_county_koppen_metric.py (writes data/metrics/koppen.csv) and apply_koppen_metric_to_climate_data.py (writes koppenZone plus koppenPrimaryClass/koppenSecondaryClass for Mixed counties). - Move shared helpers into scripts/common/ (county loading, Köppen legend, area-weighted raster shares); fix the 180th-meridian raster window for Aleutians West. - Add check_climate_data.py to validate the app CSV. - Draw Mixed counties in app.js as diagonal stripes of their top two classes, fixed to the ground and following the map at every zoom, with a crossfade only when the stripe size changes. Filtering a class also matches Mixed counties where it is primary or secondary. - Document the rule, display, and pipeline plan in docs/ and update the README and data-source notes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,291 @@
|
||||
# Data Pipeline Improvement Plan
|
||||
|
||||
This plan describes how the county data pipeline will move from scripts that
|
||||
edit one shared CSV in place to per-metric outputs assembled into the app CSV.
|
||||
It is a working reference for the filter-by-filter review. Calculation details
|
||||
for each filter live in [filter-calculations.md](filter-calculations.md).
|
||||
|
||||
**Started:** 2026-09-12
|
||||
|
||||
**Guiding decision:** review and fix each of the 12 filters one at a time,
|
||||
confirm each works on its own, and restructure `data/climate-data.csv` only
|
||||
after all filters are clean. No large rewrite happens up front.
|
||||
|
||||
## 1. Current pipeline
|
||||
|
||||
### Where each filter comes from
|
||||
|
||||
| Filter | Written into `climate-data.csv` by | Upstream scripts |
|
||||
| --- | --- | --- |
|
||||
| Köppen-Geiger class (plus the two stripe-class columns) | `apply_koppen_metric_to_climate_data.py` | `build_county_koppen_metric.py` → `data/metrics/koppen.csv` |
|
||||
| Annual avg temperature | `build_county_climate_data.py` | — |
|
||||
| Annual precipitation | `build_county_climate_data.py` | — |
|
||||
| Seasonality index | `build_county_climate_data.py` | — |
|
||||
| Wettest / driest month | Base build, then overwritten by `apply_precipitation_month_metrics_to_climate_data.py` | — |
|
||||
| Diurnal temperature range | `apply_diurnal_temperature_range_to_climate_data.py` | `build_county_diurnal_temperature_range.py` |
|
||||
| Extreme temperature days | `apply_locally_extreme_metric_to_climate_data.py` | `build_county_locally_extreme_data.py` |
|
||||
| Summer specific humidity | `apply_gridmet_humidity_metric_to_climate_data.py` | `download_gridmet_data.py` → `summarize_county_gridmet_humidity.py` |
|
||||
| 90 °F+ heat-index days (plus 2 source-FIPS columns) | `apply_gridmet_humidity_metric_to_climate_data.py` | Same as summer humidity |
|
||||
| Solar GHI | Base build (optional), then replaced by `apply_locally_extreme_metric_to_climate_data.py` | Point: `build_county_representative_points.py` → `fetch_nsrdb_representative_point_ghi.py`. Polygon: `request_nsrdb_county_polygon_ghi_archives.py` → `download_nsrdb_county_polygon_ghi_archives.py` → `summarize_nsrdb_county_polygon_archives.py` |
|
||||
| Clear-sky GHI reduction | `apply_nsrdb_cloud_metric_to_climate_data.py` | Point: `fetch_nsrdb_representative_point_cloud_metrics.py`. Polygon: `request_nsrdb_county_polygon_cloud_archives.py` → `download_nsrdb_county_polygon_cloud_archives.py` → `summarize_nsrdb_county_polygon_cloud_archives.py` |
|
||||
|
||||
Supporting scripts: `request_nsrdb_county_polygon_archives.py` and
|
||||
`download_nsrdb_county_polygon_archives.py` are the shared engines behind the
|
||||
GHI and cloud wrappers; `rebuild_nsrdb_representative_point_ghi_summary.py`
|
||||
rebuilds the point GHI summary from cache; `check_climate_data.py` validates the
|
||||
final CSV. Shared helpers live in `scripts/common/` (Phase 2). The base build
|
||||
still writes an old largest-share `koppenZone`, so the Köppen apply step must
|
||||
run after it.
|
||||
|
||||
### Current full-rebuild order
|
||||
|
||||
1. `build_county_climate_data.py`
|
||||
2. `build_county_koppen_metric.py` → `apply_koppen_metric_to_climate_data.py`
|
||||
3. `apply_precipitation_month_metrics_to_climate_data.py`
|
||||
4. `build_county_locally_extreme_data.py` → `apply_locally_extreme_metric_to_climate_data.py`
|
||||
5. `build_county_diurnal_temperature_range.py` → `apply_diurnal_temperature_range_to_climate_data.py`
|
||||
6. `summarize_county_gridmet_humidity.py` → `apply_gridmet_humidity_metric_to_climate_data.py`
|
||||
7. `apply_nsrdb_cloud_metric_to_climate_data.py`
|
||||
|
||||
The README's enrichment list starts at step 2 and omits steps 1 and 3.
|
||||
|
||||
### Problems
|
||||
|
||||
1. **Rerunning a step can destroy data.** The base build writes 12 columns,
|
||||
including the retired `extremeDays`. Later scripts delete, overwrite, or add
|
||||
columns until the live CSV has 20. Rerunning the base build drops
|
||||
9 live columns (`koppenPrimaryClass`, `koppenSecondaryClass`,
|
||||
`avgDiurnalTempRangeF`, `absoluteExtremeDays`,
|
||||
`clearSkyGhiReductionIndex`, `avgSummerSpecificHumidityGKg`,
|
||||
`humidHeatDays`, `humidHeatSourceFips`, `humidHeatFipsAdjustment`),
|
||||
restores `extremeDays`, and rewrites `koppenZone` with the old
|
||||
largest-share method.
|
||||
2. **Order is implicit.** The sequence lives in the README, in
|
||||
`scripts/county_data_sources.md`, and in each script's assumptions.
|
||||
3. **Column ownership is unclear.** GHI is finalized by the extreme-temperature
|
||||
apply script; wettest/driest month are computed in two places.
|
||||
4. **County aggregation is inconsistent.** NOAA uses touched raster cells,
|
||||
Köppen uses area-weighted shares (since 2026-09-13), gridMET uses cell
|
||||
centers with cos(latitude) weights, and NSRDB uses overlap areas (finding 11
|
||||
in `filter-calculations.md`).
|
||||
5. **No single entry point or final check.** A new user must piece together
|
||||
about 20 scripts, several large downloads, and an NSRDB API key.
|
||||
|
||||
## 2. Target design
|
||||
|
||||
1. **One metric, one file.** Each metric pipeline writes a county-level file
|
||||
under `data/metrics/`, for example `data/metrics/koppen.csv`, containing
|
||||
`countyFips`, the app value, and any audit columns for that metric.
|
||||
2. **One assemble step.** A single script joins the metric files into
|
||||
`data/climate-data.csv`, using `data/metric_sources.json` for the column list
|
||||
and per-metric source notes, then runs `check_climate_data.py`.
|
||||
- Run order no longer matters; rerunning one metric cannot damage others.
|
||||
- Every column has exactly one owner.
|
||||
- The per-row `source` column moves into `metric_sources.json`.
|
||||
- Audit columns stay in the metric files rather than the app CSV.
|
||||
3. **One shared county-aggregation module.** Area-weighted zonal statistics,
|
||||
including the 180th-meridian split, used by every raster-based metric.
|
||||
4. **One runner.** For example
|
||||
`python scripts/pipeline.py --only koppen --skip-download`, with stages for
|
||||
fetch, build metrics, assemble, and check. Cached downloads are reused by
|
||||
default.
|
||||
|
||||
## 3. Reproduction tiers
|
||||
|
||||
The "Reproducing the data" guide (Phase 4) will be organized by how deep a
|
||||
user needs to go:
|
||||
|
||||
| Tier | What the user does | Needs |
|
||||
| --- | --- | --- |
|
||||
| 1. Run the app | `.\serve.ps1` with the committed CSV | Nothing else |
|
||||
| 2. Reassemble | Rebuild `climate-data.csv` from committed metric files | Python environment only |
|
||||
| 3. Regenerate one metric | Download one source, rebuild one metric file, reassemble | That metric's source data |
|
||||
| 4. Full rebuild | Everything | All sources; NSRDB API key; large downloads (the NOAA monthly temperature file alone is about 5 GB) and hours of paced NSRDB requests |
|
||||
|
||||
The guide will list each dataset's size, download location, API-key needs, and
|
||||
approximate run time, and the Python requirements will be pinned.
|
||||
|
||||
## 4. Roadmap
|
||||
|
||||
### Phase 0 — Groundwork (done)
|
||||
|
||||
- [x] `scripts/check_climate_data.py` validates the app CSV (8 checks) with
|
||||
tests in `tests/test_check_climate_data.py`.
|
||||
- [x] `data/metric_sources.json` created as an empty skeleton.
|
||||
- [x] Köppen raster reads use a padded window per county, and polygons that
|
||||
cross the 180th meridian are split (`split_at_antimeridian` in
|
||||
`scripts/common/county_zonal_stats.py`); output verified identical for
|
||||
all 3,221 counties; tests in `tests/test_koppen_antimeridian.py`.
|
||||
|
||||
### Phase 1 — Filter-by-filter review (in progress)
|
||||
|
||||
Each filter goes through the checklist in section 5. Each fix delivers that
|
||||
metric's own file in `data/metrics/` plus a single-column apply step, so the
|
||||
existing CSV keeps working until Phase 3.
|
||||
|
||||
### Phase 2 — Shared helpers in `scripts/common/`
|
||||
|
||||
`scripts/common/` holds code used by more than one data source (NOAA, gridMET,
|
||||
NSRDB, Köppen). Scripts import from it, for example
|
||||
`from common.counties import load_counties`; nothing in it is run directly.
|
||||
|
||||
Rules for `common/`:
|
||||
|
||||
- **Cross-source only.** A helper goes in only if metrics from more than one
|
||||
data source use it. Code shared by scripts of a single data source stays with
|
||||
that source, for example a future NSRDB module for the NSRDB prompt,
|
||||
redaction, and error-log helpers.
|
||||
- **One topic per module.** Each module is named for its topic and has a
|
||||
docstring. No catch-all `utils.py`.
|
||||
- **Keep it small.** Before adding a helper, ask why it does not belong to any
|
||||
one data source.
|
||||
|
||||
Current modules:
|
||||
|
||||
| Module | Contents | Why it is in `common/` |
|
||||
| --- | --- | --- |
|
||||
| `county_zonal_stats.py` | Raster windows, the 180th-meridian split, area-weighted class shares | Used by any raster-based metric |
|
||||
| `counties.py` | County polygon loading, FIPS normalization, the state FIPS table | County identity is shared by nearly every pipeline |
|
||||
| `koppen_legend.py` | The Köppen code map and legend loader | Temporary: also used by `build_county_climate_data.py`; moves next to the Köppen code in Phase 3 |
|
||||
|
||||
**Rationale.** Helper functions make each step of a computation explicit,
|
||||
avoid repeated code, and can be tested separately
|
||||
([Brown CSCI 0111, "Helper Functions"](https://cs.brown.edu/courses/csci0111/fall2018/lectures/helper-functions.html)).
|
||||
Shared helper folders, however, tend to lose cohesion and collect unrelated
|
||||
code; the recommended alternative is to keep code with the part of the system
|
||||
it belongs to, allowing a shared folder only if it stays small and documented
|
||||
([Helpers and Utils Folders in Software Architecture](https://dev.to/knzt/helpers-and-utils-folders-in-software-architecture-3f8h)).
|
||||
The rules above follow both: shared functions, organized by topic and limited
|
||||
to code that crosses data sources.
|
||||
|
||||
Other shared code moves when its filter is reviewed, so each move is tested
|
||||
alongside that filter. A 2026-09-13 survey found 19 functions with identical
|
||||
copies in several scripts and 19 with copies that have drifted apart. Most
|
||||
identical copies are NSRDB helpers, which belong in an NSRDB module rather
|
||||
than `common/`; `read_csv_rows` (4 identical copies in `apply_*` scripts) is
|
||||
cross-source. Drifted copies need a decision on which version is correct
|
||||
before merging. Notable drifts: `summarize_county_gridmet_humidity.py` has its
|
||||
own county loader and FIPS normalizer, and the state FIPS table is also copied
|
||||
in `build_county_representative_points.py`,
|
||||
`summarize_county_gridmet_humidity.py`, and
|
||||
`request_nsrdb_county_polygon_archives.py`.
|
||||
|
||||
### Phase 3 — Assemble and restructure (after all 12 filters are clean)
|
||||
|
||||
- Assemble script that builds `climate-data.csv` from `data/metrics/`.
|
||||
- Populate `metric_sources.json`; remove the per-row `source` column.
|
||||
- Move audit columns out of the app CSV.
|
||||
- Point the app's Sources panel at `metric_sources.json`.
|
||||
- Retire or rewrite `build_county_climate_data.py` as per-metric builders.
|
||||
- Move `common/koppen_legend.py` next to the Köppen code once nothing outside
|
||||
Köppen imports it.
|
||||
|
||||
### Phase 4 — Runner and reproduction guide
|
||||
|
||||
- `scripts/pipeline.py` runner with `--only` and `--skip-download`.
|
||||
- "Reproducing the data" guide organized by the tiers in section 3.
|
||||
- Pinned requirements.
|
||||
- End-to-end smoke test on a small synthetic county fixture.
|
||||
- Organize scripts by data source (`noaa/`, `gridmet/`, `nsrdb/`, `koppen/`),
|
||||
each holding its own helpers, with `common/` keeping only cross-source code.
|
||||
Scripts in subfolders are run through the runner or as modules
|
||||
(`python -m`), and the README and data-source commands are updated to match.
|
||||
|
||||
## 5. Per-filter review checklist
|
||||
|
||||
For each filter:
|
||||
|
||||
1. Verify the calculation against the source data and document findings.
|
||||
2. Decide any rule or method changes with the project owner.
|
||||
3. Record the adopted definition in `filter-calculations.md`.
|
||||
4. Implement the calculation, writing `data/metrics/<metric>.csv`.
|
||||
5. Add a single-column apply step for the current CSV.
|
||||
6. Update the rules in `check_climate_data.py`.
|
||||
7. Add or update unit tests.
|
||||
8. Apply to the CSV, run `check_climate_data.py`, and compare changed counties
|
||||
against expectations.
|
||||
9. Update the app if the value set or display changes.
|
||||
|
||||
## 6. Filter tracker
|
||||
|
||||
Known issues come from `filter-calculations.md` ("Calculation review findings")
|
||||
and this review; none beyond Köppen have been investigated yet.
|
||||
|
||||
| # | Filter | Status | Known issues to review |
|
||||
| --- | --- | --- | --- |
|
||||
| 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | See tasks below |
|
||||
| 2 | Annual avg temperature | Not started | Months weighted equally (finding 1); touched-cell aggregation (finding 2) |
|
||||
| 3 | Diurnal temperature range | Not started | Lexington, VA (51678) blank, while heat-index days use Rockbridge County as a proxy |
|
||||
| 4 | Extreme temperature days | Not started | Partial years not normalized (finding 5); depends on retired percentile thresholds (finding 6); Lexington, VA blank |
|
||||
| 5 | 90 °F+ heat-index days | Not started | Daily-extrema proxy (finding 7); permissive year completeness (finding 8) |
|
||||
| 6 | Annual precipitation | Not started | Partial-year sums accepted (finding 3); touched-cell aggregation (finding 2) |
|
||||
| 7 | Seasonality index | Not started | Touched-cell aggregation (finding 2) |
|
||||
| 8 | Wettest month | Not started | Computed in both the base build and the precipitation-month script |
|
||||
| 9 | Driest month | Not started | Same as wettest month |
|
||||
| 10 | Summer specific humidity | Not started | Cell-center cos(latitude) aggregation differs from other metrics (finding 11) |
|
||||
| 11 | Solar GHI | Not started | Finalized by the extreme-temperature apply script; hourly, 365-day assumption (finding 9) |
|
||||
| 12 | Clear-sky GHI reduction | Not started | Mean of ratios rather than energy totals (finding 10) |
|
||||
|
||||
### Köppen-Geiger tasks
|
||||
|
||||
Adopted rule: a county is predominantly its top class if and only if that class
|
||||
covers at least 50% of the county's land and leads the runner-up by at least
|
||||
5 percentage points; otherwise it is Mixed climate. Expected result for the
|
||||
50 states and DC: 3,010 predominant, 133 Mixed.
|
||||
|
||||
- [x] Investigate low-majority counties and adopt the rule.
|
||||
- [x] Windowed raster reads and 180th-meridian split.
|
||||
- [x] Area-weighted class shares (16 × 16 sub-cells per raster cell, scaled by
|
||||
cos(latitude)) in `scripts/common/county_zonal_stats.py`.
|
||||
- [x] Apply the 50% / 5-point rule (`scripts/build_county_koppen_metric.py`;
|
||||
counties with no valid cells are left blank).
|
||||
- [x] Run the builder to write `data/metrics/koppen.csv` and confirm the
|
||||
expected 3,010 predominant / 133 Mixed (2026-09-13).
|
||||
- [x] `koppenZone`-only apply step
|
||||
(`scripts/apply_koppen_metric_to_climate_data.py`, with `--dry-run`).
|
||||
- [x] Apply to `data/climate-data.csv` (2026-09-13; 142 counties changed to
|
||||
Mixed, with `koppenPrimaryClass` and `koppenSecondaryClass` added).
|
||||
- [x] Allow `Mixed` in `check_climate_data.py`.
|
||||
- [x] Add a Mixed climate category to `app.js`, drawn as stripes of the
|
||||
county's top two classes; see
|
||||
[koppen-mixed-display-plan.md](koppen-mixed-display-plan.md).
|
||||
- [x] Replace the plurality description in `filter-calculations.md` §1 and mark
|
||||
review findings 2 and 4 resolved for Köppen (2026-09-14).
|
||||
- [x] Update the Köppen descriptions and script lists in `README.md` and
|
||||
`scripts/county_data_sources.md` (2026-09-14).
|
||||
- [x] Tests for shares, the rule, boundary cases, and the apply step
|
||||
(`tests/test_koppen_metric.py`).
|
||||
|
||||
## 7. Guardrails until Phase 3
|
||||
|
||||
- **Do not rerun `build_county_climate_data.py`** against
|
||||
`data/climate-data.csv`. It would drop 9 live columns, restore
|
||||
`extremeDays`, and overwrite the Mixed classification in `koppenZone`.
|
||||
- Run `check_climate_data.py` after every apply step.
|
||||
- Change one filter at a time, and compare its before and after values.
|
||||
|
||||
## 8. Open decisions
|
||||
|
||||
| Decision | Options | Needed by |
|
||||
| --- | --- | --- |
|
||||
| Committing large intermediates | Commit metric files only, or also source summaries | Phase 3 |
|
||||
|
||||
### Decided
|
||||
|
||||
- **Köppen audit columns (2026-09-12):** `koppen.csv` stores the top class and
|
||||
share and the runner-up class and share alongside `koppenZone`.
|
||||
- **Köppen no-data fallback (2026-09-12):** a county with no valid raster cells
|
||||
is left blank, not assigned `Cfa`.
|
||||
- **Applying Köppen to the app CSV (2026-09-12):** wait until the app supports
|
||||
the Mixed class. Done 2026-09-13.
|
||||
- **Metric files (2026-09-13):** one CSV per metric under `data/metrics/`,
|
||||
starting with `koppen.csv`.
|
||||
- **Mixed climate display (2026-09-13):** diagonal stripes of each Mixed
|
||||
county's top two classes; see
|
||||
[koppen-mixed-display-plan.md](koppen-mixed-display-plan.md).
|
||||
- **Puerto Rico (2026-09-13):** off the map and out of every filter. The app
|
||||
already drops state FIPS 72; the data files keep the rows.
|
||||
- **Shared helpers (2026-09-13):** `scripts/common/` holds only code used by
|
||||
more than one data source, one topic per module; code shared within one data
|
||||
source stays with that source. Duplicates move during their own filter's
|
||||
review.
|
||||
Reference in New Issue
Block a user