Refine project documentation and track metric files

Close gaps found in a review of the documentation:
- Track data/metrics/ and data/metric_sources.json in git so the data
  checker passes on a fresh clone (finding 23; decision logged)
- State the Alaska and Hawaii coverage gap in the README limitations and
  extend finding 12
- File findings 24-26: the data-sources doc lacks gridMET and several
  pipeline commands; wettest/driest month are computed twice; solar GHI
  is written by the extreme-temperature apply step
- Add the stale "fallback values" note to filter 2's tasks

Tidy the document system:
- Add docs/reviews/README.md with the numbering rules and a finding index
- Rename koppen-mixed-display-plan.md to koppen-mixed-display.md and fix
  its stale Puerto Rico and "stage 5" text
- Add the precipitation-month step to the README enrichment list
- Describe the Current method / Previous method pattern in plan section 5
- Add CLAUDE.md with the project guardrails and doc layout

Format filter-calculations.md so it renders on GitHub and in VS Code:
inline math uses $...$, ranges use en dashes, and implementation
references name functions instead of line numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-15 17:31:26 -04:00
co-authored by Claude Opus 5
parent e855d583e3
commit 867a07cecb
20 changed files with 3519 additions and 121 deletions
+4 -1
View File
@@ -15,10 +15,13 @@ Thumbs.db
!.env.example
# Downloaded and generated climate datasets are too large for Git.
# Keep only the assets required by the browser application.
# Keep the assets required by the browser application, the per-metric county
# files, and the metric metadata the data checker reads.
data/*
!data/climate-data.csv
!data/geojson-counties-fips.json
!data/metrics/
!data/metric_sources.json
# Runtime logs
*.log
+45
View File
@@ -0,0 +1,45 @@
# Project notes for Claude
US County Climate Explorer: a static Leaflet app (`index.html`, `app.js`,
`styles.css`) over `data/climate-data.csv`, built by the Python scripts in
`scripts/`. The 12 filters are being reviewed one at a time before the CSV is
restructured; see [docs/pipeline-plan.md](docs/pipeline-plan.md).
## Guardrails
- Never rerun `scripts/build_county_climate_data.py` over
`data/climate-data.csv`. It drops live columns and overwrites `koppenZone`
(pipeline plan §7).
- Change one filter at a time. After every apply step, run the checker and
compare changed counties against expectations.
- Run scripts from the project root with the virtual environment:
- `.venv\Scripts\python.exe scripts\check_climate_data.py`
- `.venv\Scripts\python.exe -m unittest discover -s tests`
- When `app.js` or the data files change, bump `APP_ASSET_VERSION` in `app.js`
(it also versions the CSV and GeoJSON URLs) and the matching `app.js?v=`
query in `index.html`. `styles.css` has its own query in `index.html`.
- The map and every filter cover the 50 states and DC. Puerto Rico rows stay in
the data files but are never shown.
- Do not start the Phase 3 CSV restructure until all 12 filters are reviewed.
## Writing
- Write "Köppen" with the umlaut in prose, UI text, help text, and messages.
Use ASCII `koppen` only in identifiers, file names, and data paths.
## Documentation layout
Each fact has one home:
- `docs/pipeline-plan.md`: current pipeline, target design, phases, the
per-filter checklist (§5), the filter tracker (§6), and guardrails (§7).
Keep it about 200 lines.
- `docs/reviews/`: one file per filter with its findings, decision links, and
tasks; `00-cross-filter.md` for findings that affect several filters.
`docs/reviews/README.md` has the numbering conventions and the finding index.
- `docs/decisions.md`: open decisions and a dated log of decided ones. An item
leaves "Open decisions" only when the project owner decides it.
- `docs/filter-calculations.md`: what each calculation is. No review status.
- `scripts/common/README.md`: rules for shared helpers.
Keep the tracker status in the plan and the review-file checkboxes in sync.
+10 -4
View File
@@ -46,7 +46,7 @@ calendar year. Definitions and time periods are shown in the application's
### Requirements
- A modern web browser
- Python 3
- Python 3 (the data pipeline targets Python 3.11)
- PowerShell for the included convenience script
- An internet connection for Leaflet, map tiles, and hosted fonts
@@ -167,6 +167,7 @@ are:
```powershell
.venv\Scripts\python.exe scripts\build_county_koppen_metric.py
.venv\Scripts\python.exe scripts\apply_koppen_metric_to_climate_data.py
.venv\Scripts\python.exe scripts\apply_precipitation_month_metrics_to_climate_data.py
.venv\Scripts\python.exe scripts\build_county_locally_extreme_data.py --skip-download
.venv\Scripts\python.exe scripts\apply_locally_extreme_metric_to_climate_data.py
.venv\Scripts\python.exe scripts\build_county_diurnal_temperature_range.py
@@ -176,8 +177,9 @@ are:
.venv\Scripts\python.exe scripts\apply_nsrdb_cloud_metric_to_climate_data.py
```
The order matters when rebuilding from scratch: run the Köppen apply step after
the base build, generate the locally extreme comparison and solar summaries
These stages follow the base build, `build_county_climate_data.py`, which must
not be rerun over the live CSV. The order matters when rebuilding from scratch:
run the Köppen and precipitation-month apply steps after the base build, generate the locally extreme comparison and solar summaries
before running their apply step, and summarize gridMET or NSRDB downloads
before merging them. See the data-source document linked above for acquisition
commands, expected artifacts, FIPS handling, and the representative-point and
@@ -193,7 +195,7 @@ After updating the CSV, validate it with:
Run the current automated tests with:
```powershell
python -m unittest discover -s tests
.venv\Scripts\python.exe -m unittest discover -s tests
```
## Planned Mood Analysis
@@ -223,6 +225,10 @@ causes changes in mood.
represent every location inside a county.
- Source datasets use different methods and, in some cases, different time
periods.
- Alaska and Hawaii have values only for the Köppen class and the two solar
metrics. The NOAA and gridMET sources behind the other nine metrics cover the
contiguous U.S. only, so those counties show "No data". Puerto Rico is not
shown.
- Some solar metrics use representative-point values where county polygon
summaries are unavailable.
- The current interface is an exploratory visualization, not a completed
+4
View File
@@ -0,0 +1,4 @@
{
"schemaVersion": 1,
"metrics": {}
}
File diff suppressed because it is too large Load Diff
+6 -2
View File
@@ -8,7 +8,7 @@ filter-by-filter review. The findings that led to them are in
| Decision | Options | Needed by |
| --- | --- | --- |
| Committing large intermediates | Commit metric files only, or also source summaries | Phase 3 |
| Committing source summaries | Also commit the source summaries that metric files are built from (for example `data/nrel/county_polygon_ghi_summary.csv`), or keep them out of git. Metric files are committed; see "Committing metric files" below | Phase 3 |
| Alaska and Hawaii coverage | Leave blank with a stated coverage gap, or add a source that covers them (for example Daymet or TerraClimate). NOAA nClimGrid and gridMET cover the contiguous U.S. only, so 29 AK and 5 HI counties are blank in filters 2–10; see finding 12 in [reviews/00-cross-filter.md](reviews/00-cross-filter.md) | Before Phase 3 |
## Decided
@@ -23,7 +23,7 @@ filter-by-filter review. The findings that led to them are in
starting with `koppen.csv`.
- **Mixed climate display (2026-09-13):** diagonal stripes of each Mixed
county's top two classes; see
[koppen-mixed-display-plan.md](koppen-mixed-display-plan.md).
[koppen-mixed-display.md](koppen-mixed-display.md).
- **Puerto Rico (2026-09-13):** off the map and out of every filter. The app
already drops state FIPS 72; the data files keep the rows.
- **Shared helpers (2026-09-13):** `scripts/common/` holds only code used by
@@ -43,3 +43,7 @@ filter-by-filter review. The findings that led to them are in
unless all 12 monthly values are present (WMO-No. 1203 §4.3.3).
- **Annual avg temperature in Alaska and Hawaii (2026-09-15):** left blank for
now; tracked as the cross-filter "Alaska and Hawaii coverage" open decision.
- **Committing metric files (2026-09-15):** `data/metrics/` and
`data/metric_sources.json` are tracked in git, so the data checker passes on a
fresh clone and reproduction tier 2 is possible. Closes finding 23. Whether
to also commit source summaries stays open.
+73 -74
View File
@@ -6,20 +6,20 @@ for the 12 county filters currently exposed by the climate explorer.
## Scope and notation
The current metric list is defined in `app.js` under `METRICS`. Unless noted
otherwise, long-term climate metrics use the 1991--2020 reference period.
otherwise, long-term climate metrics use the 1991–2020 reference period.
| Symbol | Meaning |
| --- | --- |
| \(c\) | County |
| \(i\) | Raster or model grid cell |
| \(m\) | Calendar month |
| \(d\) | Calendar day |
| \(h\) | Hour or NSRDB time row |
| \(y\) | Year |
| \(G_c\) | Valid grid cells assigned to county \(c\) |
| \(V_c\) | Valid observations for county \(c\) |
| \(\mathbf{1}[A]\) | 1 when condition \(A\) is true; otherwise 0 |
| \(\operatorname{clip}(x,a,b)\) | Restrict \(x\) to the interval \([a,b]\) |
| $c$ | County |
| $i$ | Raster or model grid cell |
| $m$ | Calendar month |
| $d$ | Calendar day |
| $h$ | Hour or NSRDB time row |
| $y$ | Year |
| $G_c$ | Valid grid cells assigned to county $c$ |
| $V_c$ | Valid observations for county $c$ |
| $\mathbf{1}[A]$ | 1 when condition $A$ is true; otherwise 0 |
| $\operatorname{clip}(x,a,b)$ | Restrict $x$ to the interval $[a,b]$ |
Missing values are omitted from means unless a metric-specific rule below says
otherwise.
@@ -60,7 +60,7 @@ in `app.js`.
| Temperature & Extremes | Annual Extreme Temperature Days | `absoluteExtremeDays` | Numeric | Days/year |
| Temperature & Extremes | Annual 90 F+ Heat Index Days | `humidHeatDays` | Numeric | Days/year |
| Precipitation & Moisture | Annual Precipitation (Normals) | `annualPrecipIn` | Numeric | Inches/year |
| Precipitation & Moisture | Seasonality Index | `seasonalityIndex` | Numeric | 0--100 index |
| Precipitation & Moisture | Seasonality Index | `seasonalityIndex` | Numeric | 0–100 index |
| Precipitation & Moisture | Wettest Month | `wettestPrecipMonth` | Categorical | Month |
| Precipitation & Moisture | Driest Month | `driestPrecipMonth` | Categorical | Month |
| Precipitation & Moisture | Summer Specific Humidity | `avgSummerSpecificHumidityGKg` | Numeric | g/kg |
@@ -76,9 +76,9 @@ There is no continuous numerical score. Each county is classified from the
share of its land covered by each Köppen class. The rule was adopted on
2026-09-12 and applied to `data/climate-data.csv` on 2026-09-13.
Let \(s_{c,k}\) be the share of county \(c\)'s land area covered by class \(k\).
Let $s_{c,k}$ be the share of county $c$'s land area covered by class $k$.
Each valid raster cell is weighted by the area of the cell that lies inside the
county, \(a_{c,i}\); ocean and no-data cells are excluded:
county, $a_{c,i}$; ocean and no-data cells are excluded:
$$
s_{c,k}
@@ -86,9 +86,9 @@ s_{c,k}
\frac{\sum_{i}a_{c,i}\,\mathbf{1}[K_i=k]}{\sum_{i}a_{c,i}}.
$$
Rank the classes so that \(s_{c,(1)}\ge s_{c,(2)}\ge\cdots\), with
\(s_{c,(2)}=0\) when only one class is present. A county is **predominantly**
class \(k_{(1)}\), shown as a single color, if and only if both conditions hold:
Rank the classes so that $s_{c,(1)}\ge s_{c,(2)}\ge\cdots$, with
$s_{c,(2)}=0$ when only one class is present. A county is **predominantly**
class $k_{(1)}$, shown as a single color, if and only if both conditions hold:
$$
K_c=
@@ -100,19 +100,19 @@ $$
The gap is measured in percentage points. A county that fails either condition
is classified as **Mixed** (shown as "Mixed Climate"). For Mixed counties,
\(p_c=k_{(1)}\) and \(q_c=k_{(2)}\) are stored in `koppenPrimaryClass` and
$p_c=k_{(1)}$ and $q_c=k_{(2)}$ are stored in `koppenPrimaryClass` and
`koppenSecondaryClass`, and the map draws the county with stripes of those two
classes (see [koppen-mixed-display-plan.md](koppen-mixed-display-plan.md)).
classes (see [koppen-mixed-display.md](koppen-mixed-display.md)).
Both columns are blank for predominant counties. A county with no valid raster
cells is left blank; none in the 50 states and DC is.
Each county is read from a small raster window around its polygon. Each cell is
split into 16 × 16 sub-cells to estimate the fraction inside the county, and
scaled by \(\cos(\text{latitude})\) for its true surface area. A county whose
scaled by $\cos(\text{latitude})$ for its true surface area. A county whose
polygon crosses the 180th meridian (Aleutians West, AK) is split into one piece
on each side, and each piece is read from its own window.
**Filter.** Choosing a class \(F\) shows counties that are predominantly that
**Filter.** Choosing a class $F$ shows counties that are predominantly that
class and Mixed counties where it is the primary or secondary class; the
"Mixed Climate" option shows every Mixed county:
@@ -136,13 +136,13 @@ condition catches near 50/50 splits, such as Schenectady, NY (Dfb 50.1%,
Dfa 49.9%), where a single label would rest on a margin of a few tenths of a
point. Map-unit purity standards from other fields were considered and
rejected: the FAO Land Cover Classification System treats a unit as single
only above 80%, and USDA soil survey consociations allow roughly 15--25%
only above 80%, and USDA soil survey consociations allow roughly 15–25%
dissimilar inclusions. Applied to counties, those thresholds would mark about
40--47% of the map area as mixed. A published county-level Köppen dataset
40–47% of the map area as mixed. A published county-level Köppen dataset
(Audirac, Harvard Dataverse, 2024) uses the plurality class and reports the
share of every class, without a threshold.
**Results.** Using the Beck et al. 2023 1991--2020 1 km raster, for the 3,143
**Results.** Using the Beck et al. 2023 1991–2020 1 km raster, for the 3,143
counties in the 50 states and DC:
| Classification | Counties | Share of counties | Share of map area |
@@ -194,7 +194,7 @@ the Köppen apply step must run after it.
**Data key:** `avgTempF`
The annual value is a 1991--2020 climatological standard normal: the mean of
The annual value is a 1991–2020 climatological standard normal: the mean of
the 12 monthly normals of NOAA nClimGrid-Monthly average temperature, averaged
over each county's area. The definition was adopted on 2026-09-15. It is not
yet applied to `data/climate-data.csv`, whose values still use the current
@@ -206,8 +206,8 @@ the present. Its average temperature, `tavg`, is the mean of maximum and
minimum temperature, (Tmax + Tmin)/2, not a 24-hour mean. Alaska and Hawaii
are outside the grid, so their counties are blank.
**Cell normals.** For each grid cell \(i\) and calendar month \(m\), the
normal is the mean over the 1991--2020 years with a valid value:
**Cell normals.** For each grid cell $i$ and calendar month $m$, the
normal is the mean over the 1991–2020 years with a valid value:
$$
T_{i,m}^{\mathrm{norm}}
@@ -224,7 +224,7 @@ values can differ slightly from NCEI's published grids. The result is not
NCEI's station-based U.S. Climate Normals product.
**County monthly values.** Each cell is weighted by the area of the cell that
lies inside the county, \(a_{c,i}\), estimated as for Köppen (§1). Cells
lies inside the county, $a_{c,i}$, estimated as for Köppen (§1). Cells
without data, such as ocean, are excluded:
$$
@@ -234,7 +234,7 @@ T_{c,m}
$$
**Annual value.** Every month has equal weight. The value is defined only when
all 12 monthly values \(T_{c,m}\) exist; otherwise the county is blank:
all 12 monthly values $T_{c,m}$ exist; otherwise the county is blank:
$$
T_c(^\circ\mathrm{F})
@@ -248,7 +248,7 @@ Values are kept at full precision until the stored value is rounded to
0.1 °F.
**Rationale.** The method follows the WMO rules for annual normals, which
NOAA also applies to its 1991--2020 Normals:
NOAA also applies to its 1991–2020 Normals:
- *Equal month weights.* For a mean, WMO-No. 1203 §4.3.3(a) defines the annual
normal as "the mean of the monthly normals", and its footnote says weighting
@@ -286,7 +286,7 @@ T_{c,m}
\sum_{i\in G_{c,m}}T_{i,m}^{\mathrm{norm}},
$$
and the annual value averages whichever of the \(M_c\) monthly values are
and the annual value averages whichever of the $M_c$ monthly values are
available, normally 12:
$$
@@ -297,9 +297,8 @@ T_c(^\circ\mathrm{F})
\right)\frac{9}{5}+32.
$$
Implementation: `scripts/build_county_climate_data.py:124--166`,
`scripts/build_county_climate_data.py:206--221`, and
`scripts/build_county_climate_data.py:478--514`.
Implementation: `_as_monthly_climatology`, `_zonal_mean`, and
`build_county_records` in `scripts/build_county_climate_data.py`.
## 3. Diurnal Temperature Range
@@ -312,7 +311,7 @@ DTR_{c,d}=T^{\max}_{c,d}-T^{\min}_{c,d}.
$$
Days with a missing input or a negative range are excluded. The final metric is
the mean across all retained days in 1991--2020, followed by conversion of a
the mean across all retained days in 1991–2020, followed by conversion of a
Celsius temperature *difference* to a Fahrenheit difference:
$$
@@ -324,11 +323,12 @@ $$
\right).
$$
There is correctly no \(+32\) term when converting a temperature difference.
There is correctly no $+32$ term when converting a temperature difference.
The output artifact stores two decimal places.
Implementation: `scripts/build_county_diurnal_temperature_range.py:46--104`
and `scripts/build_county_diurnal_temperature_range.py:132--151`.
Implementation: `build_diurnal_temperature_range`, `c_delta_to_f_delta`, and
`write_diurnal_temperature_range_csv` in
`scripts/build_county_diurnal_temperature_range.py`.
## 4. Annual Extreme Temperature Days
@@ -354,21 +354,20 @@ $$
E_c=\frac{1}{Y_c}\sum_{y\in V_c}E_{c,y}.
$$
The checked-in data uses 1991--2025. The Boolean OR means that a hypothetical
The checked-in data uses 1991–2025. The Boolean OR means that a hypothetical
day meeting both conditions is still counted only once. A year is included when
an annual record exists; counts are not normalized to 365 or 366 valid days.
Implementation: `scripts/build_county_locally_extreme_data.py:473--540`,
`scripts/build_county_locally_extreme_data.py:669--674`, and
`scripts/build_county_locally_extreme_data.py:705--722`.
Implementation: `build_annual_counts`, `_average_or_none`, and
`write_comparison_csv` in `scripts/build_county_locally_extreme_data.py`.
## 5. Annual 90 F+ Heat Index Days
**Data key:** `humidHeatDays`
The daily proxy pairs NOAA nClimGrid-Daily county Tmax, \(T\), with the gridMET
county daily minimum relative humidity, \(R\). Relative humidity is clipped to
\([0,100]\).
The daily proxy pairs NOAA nClimGrid-Daily county Tmax, $T$, with the gridMET
county daily minimum relative humidity, $R$. Relative humidity is clipped to
$[0,100]$.
The NWS simple Heat Index estimate is calculated in two steps:
@@ -380,7 +379,7 @@ $$
HI_s=\frac{S+T}{2}.
$$
When \(HI_s\ge80^\circ\mathrm{F}\), the Rothfusz regression is used:
When $HI_s\ge80^\circ\mathrm{F}$, the Rothfusz regression is used:
$$
\begin{aligned}
@@ -390,7 +389,7 @@ HI_r={}&-42.379+2.04901523T+10.14333127R-0.22475541TR\\
\end{aligned}
$$
For \(R<13\) and \(80\le T\le112\), subtract:
For $R<13$ and $80\le T\le112$, subtract:
$$
A_{low}
@@ -399,7 +398,7 @@ A_{low}
\sqrt{\max\left(\frac{17-|T-95|}{17},0\right)}.
$$
For \(R>85\) and \(80\le T\le87\), add:
For $R>85$ and $80\le T\le87$, add:
$$
A_{high}=\frac{R-85}{10}\frac{87-T}{5}.
@@ -428,17 +427,17 @@ $$
H_c=\frac{1}{Y_c}\sum_{y\in V_c}H_{c,y}.
$$
The period is 1991--2020. A year with at least one valid paired day contributes
The period is 1991–2020. A year with at least one valid paired day contributes
equally to the final average; there is no completeness adjustment.
Implementation: `scripts/summarize_county_gridmet_humidity.py:439--474` and
`scripts/summarize_county_gridmet_humidity.py:547--617`.
Implementation: `_heat_index_f` and `summarize` in
`scripts/summarize_county_gridmet_humidity.py`.
## 6. Annual Precipitation (Normals)
**Data key:** `annualPrecipIn`
Each monthly county total, \(P_{c,m}\), is the unweighted mean of valid raster
Each monthly county total, $P_{c,m}$, is the unweighted mean of valid raster
cells touched by the county. Annual precipitation is the sum of available
monthly totals, converted from millimeters to inches:
@@ -452,8 +451,8 @@ $$
The stored value is rounded to 0.1 inch. The implementation only requires one
valid month, so missing months produce a partial annual sum rather than a blank.
Implementation: `scripts/build_county_climate_data.py:401--412` and
`scripts/build_county_climate_data.py:478--511`.
Implementation: `_zonal_mean` and `build_county_records` in
`scripts/build_county_climate_data.py`.
## 7. Seasonality Index
@@ -485,9 +484,10 @@ SI_c=
\right).
$$
If \(\mu_c\le0\), the index is set to zero. The value is stored as an integer.
If $\mu_c\le0$, the index is set to zero. The value is stored as an integer.
Implementation: `scripts/build_county_climate_data.py:487--521`.
Implementation: `build_county_records` in
`scripts/build_county_climate_data.py`.
## 8. Wettest Month
@@ -500,8 +500,8 @@ $$
Missing monthly values are ignored. An exact tie resolves to the earliest tied
month because `numpy.nanargmax` returns the first occurrence.
Implementation:
`scripts/apply_precipitation_month_metrics_to_climate_data.py:56--90`.
Implementation: `build_precip_month_lookup` in
`scripts/apply_precipitation_month_metrics_to_climate_data.py`.
## 9. Driest Month
@@ -514,15 +514,15 @@ $$
Missing monthly values are ignored. An exact tie likewise resolves to the
earliest tied month.
Implementation:
`scripts/apply_precipitation_month_metrics_to_climate_data.py:56--90`.
Implementation: `build_precip_month_lookup` in
`scripts/apply_precipitation_month_metrics_to_climate_data.py`.
## 10. Summer Specific Humidity
**Data key:** `avgSummerSpecificHumidityGKg`
gridMET cells whose centers fall inside a county are weighted by the cosine of
their latitude to approximate their relative surface areas on a latitude--longitude
their latitude to approximate their relative surface areas on a latitude–longitude
grid:
$$
@@ -550,9 +550,8 @@ $$
If no grid-cell center falls inside a county, the nearest grid cell to an
interior representative point is used.
Implementation: `scripts/summarize_county_gridmet_humidity.py:297--380`,
`scripts/summarize_county_gridmet_humidity.py:409--434`, and
`scripts/summarize_county_gridmet_humidity.py:555--605`.
Implementation: `_build_county_grid_map`, `_county_means_chunk`, and
`summarize` in `scripts/summarize_county_gridmet_humidity.py`.
## 11. Mean Daily Global Horizontal Radiation (GHI)
@@ -576,19 +575,19 @@ G_c
\frac{\sum_s A_{c,s}G_s}{\sum_s A_{c,s}},
$$
where \(A_{c,s}\) is the estimated overlap area between the county geometry and
the 4 km square grid cell centered on site \(s\). A representative-point value
where $A_{c,s}$ is the estimated overlap area between the county geometry and
the 4 km square grid cell centered on site $s$. A representative-point value
is used when a polygon summary is unavailable.
Implementation: `scripts/summarize_nsrdb_county_polygon_archives.py:189--192`
and `scripts/summarize_nsrdb_county_polygon_archives.py:335--378`.
Implementation: `site_average_daily_ghi` and `area_weighted_average` in
`scripts/summarize_nsrdb_county_polygon_archives.py`.
## 12. Clear-Sky GHI Reduction Index
**Data key:** `clearSkyGhiReductionIndex`
Only rows with valid observed and clear-sky GHI and
\(CSGHI_{s,h}\ge50\;\mathrm{W/m^2}\) are treated as daylight rows. For each
$CSGHI_{s,h}\ge50\;\mathrm{W/m^2}$ are treated as daylight rows. For each
retained row:
$$
@@ -613,11 +612,11 @@ $$
A representative-point index is used where a polygon summary is unavailable.
This definition is the mean of time-row ratios; it is not generally equal to
\(1-\sum GHI/\sum CSGHI\).
$1-\sum GHI/\sum CSGHI$.
Implementation:
`scripts/summarize_nsrdb_county_polygon_cloud_archives.py:138--218` and
`scripts/summarize_nsrdb_county_polygon_cloud_archives.py:222--304`.
Implementation: `site_cloud_metrics`, `weighted_metric`, and
`summarize_archives` in
`scripts/summarize_nsrdb_county_polygon_cloud_archives.py`.
## Review findings
@@ -625,4 +624,4 @@ Review findings and their status are kept with the filter reviews in
[reviews/](reviews/): findings specific to one filter in that filter's file,
and findings that affect several filters in
[reviews/00-cross-filter.md](reviews/00-cross-filter.md). Findings keep their
original numbers.
original numbers; [reviews/README.md](reviews/README.md) lists every one.
@@ -158,7 +158,7 @@ After changing `app.js`, bump `APP_ASSET_VERSION` and the `app.js?v=` query in
and only if `koppenZone` is `Mixed`, with valid and different Köppen codes.
- **Tests:** `tests/test_koppen_metric.py` and `tests/test_check_climate_data.py`.
## 6. Documentation (stage 5, done 2026-09-14)
## 6. Documentation (done 2026-09-14)
- `filter-calculations.md` §1 describes the applied rule, the stripe columns,
and the filter; the old method is kept as a short "Previous method" note.
@@ -195,5 +195,5 @@ Decisions that were reversed or refined during the browser review:
The app already leaves Puerto Rico off the map: `prepareCountyFeature` in
`app.js` drops counties with state FIPS 72, and the dropdown and legend are built
from the counties on the map. The rows remain in `data/climate-data.csv` and
`data/metrics/koppen.csv`. Whether to also remove them from the data files is a
separate decision.
`data/metrics/koppen.csv`, as decided on 2026-09-13; see "Puerto Rico" in
[decisions.md](decisions.md).
+9 -6
View File
@@ -50,8 +50,6 @@ run after it.
6. `summarize_county_gridmet_humidity.py` → `apply_gridmet_humidity_metric_to_climate_data.py`
7. `apply_nsrdb_cloud_metric_to_climate_data.py`
The README's enrichment list starts at step 2 and omits steps 1 and 3.
### Problems
1. **Rerunning a step can destroy data.** The base build writes 12 columns,
@@ -66,7 +64,8 @@ The README's enrichment list starts at step 2 and omits steps 1 and 3.
2. **Order is implicit.** The sequence lives in the README, in
`scripts/county_data_sources.md`, and in each script's assumptions.
3. **Column ownership is unclear.** GHI is finalized by the extreme-temperature
apply script; wettest/driest month are computed in two places.
apply script; wettest/driest month are computed in two places. See findings
25 and 26 in [reviews/00-cross-filter.md](reviews/00-cross-filter.md).
4. **County aggregation is inconsistent.** See finding 11 in
[reviews/00-cross-filter.md](reviews/00-cross-filter.md).
5. **No single entry point or final check.** A new user must piece together
@@ -160,20 +159,24 @@ For each filter:
filter's review file under `docs/reviews/`.
2. Decide any rule or method changes with the project owner, and record them in
`decisions.md`.
3. Record the adopted definition in `filter-calculations.md`.
3. Record the adopted definition in `filter-calculations.md`. Until the new
values are applied, keep the old definition below it under "Current
method".
4. Implement the calculation, writing `data/metrics/<metric>.csv`.
5. Add a single-column apply step for the current CSV.
6. Update the rules in `check_climate_data.py`.
7. Add or update unit tests.
8. Apply to the CSV, run `check_climate_data.py`, and compare changed counties
against expectations.
against expectations. Then rename "Current method" to "Previous method" in
`filter-calculations.md`, as §1 does.
9. Update the app if the value set or display changes.
## 6. Filter tracker
Each filter's findings, decisions, and tasks are in its review file. Findings
that affect more than one filter are in
[00-cross-filter.md](reviews/00-cross-filter.md).
[00-cross-filter.md](reviews/00-cross-filter.md), and
[reviews/README.md](reviews/README.md) indexes every finding number.
| # | Filter | Status | Review |
| --- | --- | --- | --- |
+61 -8
View File
@@ -1,10 +1,7 @@
# Cross-filter review
Findings that affect more than one filter. Each finding lives in exactly one
review file; the filter reviews it affects link here and record only how it
applied to them. Findings keep their numbers across all review files, and new
findings take the next number wherever they are filed. A finding that turns out
to affect other filters moves here, leaving a link behind.
Findings that affect more than one filter. The numbering conventions and an
index of every finding are in [README.md](README.md).
## Findings
@@ -19,7 +16,7 @@ Affects filters 1 (resolved), 2, 6, and 7.
**11. Spatial weighting is inconsistent across metric families.** Base NOAA
normals use equal touched-cell weights, Köppen uses area-weighted class
shares, gridMET humidity uses
\(\cos(\phi)\) weights on cell centers, and NSRDB polygon metrics use
$\cos(\phi)$ weights on cell centers, and NSRDB polygon metrics use
estimated overlap areas. Cross-metric comparisons should account for these
different county aggregation methods.
@@ -35,7 +32,12 @@ clear-sky reduction (filters 11 and 12) cover both states.
`check_climate_data.py` allows these blanks (`OUTSIDE_CONUS`). The notes on
the base build in `scripts/county_data_sources.md` say "fallback values are
applied" for counties outside NOAA coverage, but `build_county_climate_data.py`
leaves them blank (lines 499–520).
leaves them blank (`build_county_records`).
The README's "Current Limitations" states the gap (added 2026-09-15), but the
app's county panel shows only "No data". If these counties stay blank, the app
should give the reason, for example "Not covered: source data is contiguous
U.S. only".
Affects filters 2–10. Open decision: "Alaska and Hawaii coverage" in
[decisions.md](../decisions.md).
@@ -99,6 +101,56 @@ neither column is in `data/climate-data.csv`. The list has no entry for
Affects filters 3 and 4.
**23. The metric files and `metric_sources.json` were not in git.**
`.gitignore` excluded everything under `data/` except the app CSV and the
county GeoJSON, so `data/metrics/koppen.csv` and `data/metric_sources.json`
existed only locally. On a fresh clone, `check_climate_data.py` failed its
metric sources check, and reproduction tier 2 in
[pipeline-plan.md](../pipeline-plan.md) (reassemble from committed metric
files) was impossible. Found 2026-09-15. *Resolved on 2026-09-15: `.gitignore`
now keeps both; see "Committing metric files" in
[decisions.md](../decisions.md).*
Affects every filter, since each adds a metric file.
**24. The data-sources doc does not cover several pipeline steps.**
`scripts/county_data_sources.md` has source sections for Köppen, the NOAA
gridded normals, county geometry, NSRDB solar, and nClimGrid-Daily, but none
for gridMET: no dataset link, license, variables, or commands for
`download_gridmet_data.py` → `summarize_county_gridmet_humidity.py` →
`apply_gridmet_humidity_metric_to_climate_data.py`. gridMET appears only in two
metric-definition lines. The doc also has no commands for
`build_county_diurnal_temperature_range.py` /
`apply_diurnal_temperature_range_to_climate_data.py` or
`apply_precipitation_month_metrics_to_climate_data.py`. Found 2026-09-15.
Affects filters 3, 5, 8, 9, and 10; each filter's review adds its section or
commands.
**25. Wettest and driest month are computed in two places.**
`build_county_climate_data.py` writes `wettestPrecipMonth` and
`driestPrecipMonth` (`_precip_month_extremes`), and
`apply_precipitation_month_metrics_to_climate_data.py` recomputes and
overwrites them (`build_precip_month_lookup`), reusing the base build's
climatology and zonal-mean helpers. The live values are correct only if the
second script runs after the base build. Noted in the original review;
numbered 2026-09-15.
Affects filters 8 and 9. This is part of problem 3 in
[pipeline-plan.md](../pipeline-plan.md) §1.
**26. Solar GHI is finalized by the extreme-temperature apply step.**
`build_county_climate_data.py` can write
`meanDailyGlobalHorizontalRadiationKwhM2Day`, but the live value is written by
`apply_locally_extreme_metric_to_climate_data.py`, which prefers polygon GHI
and falls back to representative-point GHI. Updating GHI therefore means
rerunning the extreme-temperature apply step, and nothing in that script's name
says it owns the solar column. Noted in the original review; numbered
2026-09-15.
Affects filters 4 and 11. This is part of problem 3 in
[pipeline-plan.md](../pipeline-plan.md) §1.
## Decisions
In [decisions.md](../decisions.md): Shared helpers (2026-09-13), and the open
@@ -115,7 +167,8 @@ scripts, not solely from UI descriptions. No calculation code was changed. The
automated test suite was not executed during this review because `pytest` is
not installed in either the system Python environment or the project virtual
environment. The suite uses `unittest` and runs without pytest:
`python -m unittest discover -s tests`.
`.venv\Scripts\python.exe -m unittest discover -s tests`. As of 2026-09-15,
all 74 tests pass.
Section 1 and findings 2, 4, and 11 were updated on 2026-09-14, after the
Köppen classification was reworked and applied.
+2 -2
View File
@@ -6,7 +6,7 @@ updated.
**Data keys:** `koppenZone`, `koppenPrimaryClass`, `koppenSecondaryClass`.
Calculation: [filter-calculations.md](../filter-calculations.md) §1. Display
design and history:
[koppen-mixed-display-plan.md](../koppen-mixed-display-plan.md).
[koppen-mixed-display.md](../koppen-mixed-display.md).
## Findings
@@ -69,7 +69,7 @@ covers at least 50% of the county's land and leads the runner-up by at least
- [x] Allow `Mixed` in `check_climate_data.py`.
- [x] Add a Mixed climate category to `app.js`, drawn as stripes of the
county's top two classes; see
[koppen-mixed-display-plan.md](../koppen-mixed-display-plan.md).
[koppen-mixed-display.md](../koppen-mixed-display.md).
- [x] Replace the plurality description in `filter-calculations.md` §1 and mark
review findings 2 and 4 resolved for Köppen (2026-09-14).
- [x] Update the Köppen descriptions and script lists in `README.md` and
@@ -76,6 +76,9 @@ Climate Normals*). Alaska and Hawaii stay blank; see
- [ ] Correct `scripts/county_data_sources.md`: Source 2 says the build reads
monthly normals files, but it reads the nClimGrid monthly series and
averages 1991–2020. Update the `avgTempF` definition line to match.
Also remove the "Run the generator" note that fallback values are
applied to counties outside NOAA coverage; they are left blank
(finding 12).
- [ ] Add the area-weighted `avgTempF` to the §7 guardrail on rerunning
`build_county_climate_data.py` in
[pipeline-plan.md](../pipeline-plan.md).
+2 -1
View File
@@ -11,7 +11,8 @@ No filter-specific findings yet.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered),
13 (Lexington, VA blank), and 17 (missing from the data-sources metric list).
13 (Lexington, VA blank), 17 (missing from the data-sources metric list), and
24 (no build or apply commands in the data-sources doc).
## Decisions
+2 -1
View File
@@ -28,7 +28,8 @@ included in the app CSV.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered),
13 (Lexington, VA blank), and 17 (missing from the data-sources metric list).
13 (Lexington, VA blank), 17 (missing from the data-sources metric list), and
26 (this filter's apply step also writes solar GHI).
## Decisions
+3 -2
View File
@@ -17,8 +17,9 @@ Tmax/RH pair is included in the equal-year average. A minimum valid-day rule
would reduce low-biased partial-year counts.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered)
and 13 (Lexington, VA uses Rockbridge County as a proxy).
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered),
13 (Lexington, VA uses Rockbridge County as a proxy), and 24 (no gridMET
section in the data-sources doc).
## Decisions
+5 -4
View File
@@ -7,12 +7,13 @@
## Findings
No numbered findings yet. Known issue to review: computed in both the base
build and the precipitation-month script.
No filter-specific findings yet.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered)
and 16 (the app's nClimGrid source link returns 404).
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered),
16 (the app's nClimGrid source link returns 404), 24 (no precipitation-month
command in the data-sources doc), and 25 (computed in both the base build and
the precipitation-month script).
## Decisions
+5 -5
View File
@@ -7,13 +7,13 @@
## Findings
No numbered findings yet. Known issue to review: same as wettest month, which
is computed in both the base build and the precipitation-month script; see
[08-wettest-month.md](08-wettest-month.md).
No filter-specific findings yet.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered)
and 16 (the app's nClimGrid source link returns 404).
[00-cross-filter.md](00-cross-filter.md): 12 (Alaska and Hawaii not covered),
16 (the app's nClimGrid source link returns 404), 24 (no precipitation-month
command in the data-sources doc), and 25 (computed in both the base build and
the precipitation-month script).
## Decisions
+4 -4
View File
@@ -7,12 +7,12 @@
## Findings
No filter-specific findings yet. Known issue to review: cell-center
cos(latitude) aggregation differs from other metrics (finding 11).
No filter-specific findings yet.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 11 (inconsistent aggregation) and
12 (Alaska and Hawaii not covered).
[00-cross-filter.md](00-cross-filter.md): 11 (inconsistent aggregation),
12 (Alaska and Hawaii not covered), and 24 (no gridMET section in the
data-sources doc).
## Decisions
+3 -4
View File
@@ -25,11 +25,10 @@ representative-point fallback, and the per-row `source` tag is always
`data/nrel/county_polygon_ghi_summary.csv`. The apply step later replaces both
the value and the tag, so only the base build's output is mislabeled.
Known issue to review: finalized by the extreme-temperature apply script.
Cross-filter findings that affect this filter, in
[00-cross-filter.md](00-cross-filter.md): 11 (inconsistent aggregation) and
16 (the app's NSRDB source link no longer resolves).
[00-cross-filter.md](00-cross-filter.md): 11 (inconsistent aggregation),
16 (the app's NSRDB source link no longer resolves), and 26 (finalized by the
extreme-temperature apply step).
## Decisions
+53
View File
@@ -0,0 +1,53 @@
# Filter reviews
Each filter has one review file (`01`–`12`) holding its findings, links to its
decisions, and its task checklist. [00-cross-filter.md](00-cross-filter.md)
holds findings that affect more than one filter. Review status is tracked in
[pipeline-plan.md](../pipeline-plan.md) §6, and every review follows the
checklist in §5.
## Conventions
- Each finding lives in exactly one review file. The other filter reviews it
affects link to it and record only how it applied to them.
- Findings keep one number across all review files. A new finding takes the
next number (27 as of 2026-09-15) wherever it is filed, and is added to the
index below.
- A finding that turns out to affect other filters moves to `00`, leaving a
link behind.
- A resolved finding stays where it is, with its resolution in italics.
- A finding accepted as a limitation becomes a caveat in
[filter-calculations.md](../filter-calculations.md).
## Finding index
Status is kept in each finding, not here.
| # | Finding | File |
| --- | --- | --- |
| 1 | Annual temperature weights months equally | [02](02-annual-avg-temperature.md) |
| 2 | Base NOAA aggregation is not area-weighted | [00](00-cross-filter.md) |
| 3 | Partial precipitation years are accepted | [06](06-annual-precipitation.md) |
| 4 | The Köppen fallback can create false data | [01](01-koppen.md) |
| 5 | Extreme-day counts are not completeness-normalized | [04](04-extreme-temperature-days.md) |
| 6 | The absolute-extreme metric depends on unrelated percentile thresholds | [04](04-extreme-temperature-days.md) |
| 7 | Heat Index days are a daily-extrema proxy | [05](05-heat-index-days.md) |
| 8 | Heat-year completeness is permissive | [05](05-heat-index-days.md) |
| 9 | The GHI formula assumes hourly, 365-day input | [11](11-solar-ghi.md) |
| 10 | Clear-sky reduction averages ratios rather than energy totals | [12](12-clear-sky-ghi-reduction.md) |
| 11 | Spatial weighting is inconsistent across metric families | [00](00-cross-filter.md) |
| 12 | Alaska and Hawaii are not covered by the NOAA and gridMET sources | [00](00-cross-filter.md) |
| 13 | Lexington, VA (51678) is missing from the NOAA daily county files | [00](00-cross-filter.md) |
| 14 | Helper functions are duplicated across scripts | [00](00-cross-filter.md) |
| 15 | The app's source text for three NOAA metrics names the wrong product | [00](00-cross-filter.md) |
| 16 | Source links point to retired or missing pages | [00](00-cross-filter.md) |
| 17 | The data-sources metric list is out of date | [00](00-cross-filter.md) |
| 18 | The data-sources doc describes the retired locally extreme metric | [04](04-extreme-temperature-days.md) |
| 19 | The data-sources doc recommends a GHI raster the pipeline does not use | [11](11-solar-ghi.md) |
| 20 | The base build labels any solar CSV as representative-point | [11](11-solar-ghi.md) |
| 21 | The documented Köppen labels differ from the app | [01](01-koppen.md) |
| 22 | The base build's docstring calls its Köppen value a majority class | [01](01-koppen.md) |
| 23 | The metric files and `metric_sources.json` were not in git | [00](00-cross-filter.md) |
| 24 | The data-sources doc does not cover several pipeline steps | [00](00-cross-filter.md) |
| 25 | Wettest and driest month are computed in two places | [00](00-cross-filter.md) |
| 26 | Solar GHI is finalized by the extreme-temperature apply step | [00](00-cross-filter.md) |