Complete Köppen-Geiger filter review with Mixed climate class
Classify each county by area-weighted Köppen class shares: a county is predominantly its top class when that class covers at least 50% of its land and leads the runner-up by at least 5 percentage points; otherwise it is Mixed (133 of 3,143 counties in the 50 states and DC). - Add build_county_koppen_metric.py (writes data/metrics/koppen.csv) and apply_koppen_metric_to_climate_data.py (writes koppenZone plus koppenPrimaryClass/koppenSecondaryClass for Mixed counties). - Move shared helpers into scripts/common/ (county loading, Köppen legend, area-weighted raster shares); fix the 180th-meridian raster window for Aleutians West. - Add check_climate_data.py to validate the app CSV. - Draw Mixed counties in app.js as diagonal stripes of their top two classes, fixed to the ground and following the map at every zoom, with a crossfade only when the stripe size changes. Filtering a class also matches Mixed counties where it is primary or secondary. - Document the rule, display, and pipeline plan in docs/ and update the README and data-source notes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,620 @@
|
||||
# County Filter Calculations
|
||||
|
||||
This document records the equations, implementation behavior, and review notes
|
||||
for the 12 county filters currently exposed by the climate explorer.
|
||||
|
||||
## Scope and notation
|
||||
|
||||
The current metric list is defined in `app.js` under `METRICS`. Unless noted
|
||||
otherwise, long-term climate metrics use the 1991--2020 reference period.
|
||||
|
||||
| Symbol | Meaning |
|
||||
| --- | --- |
|
||||
| \(c\) | County |
|
||||
| \(i\) | Raster or model grid cell |
|
||||
| \(m\) | Calendar month |
|
||||
| \(d\) | Calendar day |
|
||||
| \(h\) | Hour or NSRDB time row |
|
||||
| \(y\) | Year |
|
||||
| \(G_c\) | Valid grid cells assigned to county \(c\) |
|
||||
| \(V_c\) | Valid observations for county \(c\) |
|
||||
| \(\mathbf{1}[A]\) | 1 when condition \(A\) is true; otherwise 0 |
|
||||
| \(\operatorname{clip}(x,a,b)\) | Restrict \(x\) to the interval \([a,b]\) |
|
||||
|
||||
Missing values are omitted from means unless a metric-specific rule below says
|
||||
otherwise.
|
||||
|
||||
## Browser filter predicate
|
||||
|
||||
For every numeric metric, a county is active when its value is finite and falls
|
||||
inside the selected range, including both endpoints:
|
||||
|
||||
$$
|
||||
\operatorname{passes}(c) \iff L \le x_c \le U.
|
||||
$$
|
||||
|
||||
For a categorical metric:
|
||||
|
||||
$$
|
||||
\operatorname{passes}(c) \iff
|
||||
\left(F=\text{All}\right) \lor \left(x_c=F\right).
|
||||
$$
|
||||
|
||||
The Köppen-Geiger filter also matches Mixed counties by their top two classes;
|
||||
see §1.
|
||||
|
||||
Counties with missing categorical values, null numeric values, or non-finite
|
||||
numeric values do not pass. The numeric slider limits are derived from the
|
||||
loaded data and rounded outward by each metric's configured `boundsStep`.
|
||||
|
||||
Implementation: `shouldFeaturePassFilter` and `getUniqueCategoryValuesInData`
|
||||
in `app.js`.
|
||||
|
||||
## Filter inventory
|
||||
|
||||
| Group | UI label | Data key | Type | Display unit |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Climate Classification | Köppen-Geiger Climate Class | `koppenZone` | Categorical | Class |
|
||||
| Temperature & Extremes | Annual Avg Temperature (Normals) | `avgTempF` | Numeric | °F |
|
||||
| Temperature & Extremes | Diurnal Temperature Range | `avgDiurnalTempRangeF` | Numeric | °F difference |
|
||||
| Temperature & Extremes | Annual Extreme Temperature Days | `absoluteExtremeDays` | Numeric | Days/year |
|
||||
| Temperature & Extremes | Annual 90 F+ Heat Index Days | `humidHeatDays` | Numeric | Days/year |
|
||||
| Precipitation & Moisture | Annual Precipitation (Normals) | `annualPrecipIn` | Numeric | Inches/year |
|
||||
| Precipitation & Moisture | Seasonality Index | `seasonalityIndex` | Numeric | 0--100 index |
|
||||
| Precipitation & Moisture | Wettest Month | `wettestPrecipMonth` | Categorical | Month |
|
||||
| Precipitation & Moisture | Driest Month | `driestPrecipMonth` | Categorical | Month |
|
||||
| Precipitation & Moisture | Summer Specific Humidity | `avgSummerSpecificHumidityGKg` | Numeric | g/kg |
|
||||
| Solar Resource | Mean Daily Global Horizontal Radiation (GHI) | `meanDailyGlobalHorizontalRadiationKwhM2Day` | Numeric | kWh/m²/day |
|
||||
| Solar Resource | Clear-Sky GHI Reduction Index | `clearSkyGhiReductionIndex` | Numeric | Ratio |
|
||||
|
||||
## 1. Köppen-Geiger Climate Class
|
||||
|
||||
**Data keys:** `koppenZone`; `koppenPrimaryClass` and `koppenSecondaryClass`
|
||||
for Mixed counties
|
||||
|
||||
There is no continuous numerical score. Each county is classified from the
|
||||
share of its land covered by each Köppen class. The rule was adopted on
|
||||
2026-09-12 and applied to `data/climate-data.csv` on 2026-09-13.
|
||||
|
||||
Let \(s_{c,k}\) be the share of county \(c\)'s land area covered by class \(k\).
|
||||
Each valid raster cell is weighted by the area of the cell that lies inside the
|
||||
county, \(a_{c,i}\); ocean and no-data cells are excluded:
|
||||
|
||||
$$
|
||||
s_{c,k}
|
||||
=
|
||||
\frac{\sum_{i}a_{c,i}\,\mathbf{1}[K_i=k]}{\sum_{i}a_{c,i}}.
|
||||
$$
|
||||
|
||||
Rank the classes so that \(s_{c,(1)}\ge s_{c,(2)}\ge\cdots\), with
|
||||
\(s_{c,(2)}=0\) when only one class is present. A county is **predominantly**
|
||||
class \(k_{(1)}\), shown as a single color, if and only if both conditions hold:
|
||||
|
||||
$$
|
||||
K_c=
|
||||
\begin{cases}
|
||||
k_{(1)}, & s_{c,(1)}\ge0.50 \;\text{and}\; s_{c,(1)}-s_{c,(2)}\ge0.05,\\
|
||||
\text{Mixed}, & \text{otherwise}.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
The gap is measured in percentage points. A county that fails either condition
|
||||
is classified as **Mixed** (shown as "Mixed Climate"). For Mixed counties,
|
||||
\(p_c=k_{(1)}\) and \(q_c=k_{(2)}\) are stored in `koppenPrimaryClass` and
|
||||
`koppenSecondaryClass`, and the map draws the county with stripes of those two
|
||||
classes (see [koppen-mixed-display-plan.md](koppen-mixed-display-plan.md)).
|
||||
Both columns are blank for predominant counties. A county with no valid raster
|
||||
cells is left blank; none in the 50 states and DC is.
|
||||
|
||||
Each county is read from a small raster window around its polygon. Each cell is
|
||||
split into 16 × 16 sub-cells to estimate the fraction inside the county, and
|
||||
scaled by \(\cos(\text{latitude})\) for its true surface area. A county whose
|
||||
polygon crosses the 180th meridian (Aleutians West, AK) is split into one piece
|
||||
on each side, and each piece is read from its own window.
|
||||
|
||||
**Filter.** Choosing a class \(F\) shows counties that are predominantly that
|
||||
class and Mixed counties where it is the primary or secondary class; the
|
||||
"Mixed Climate" option shows every Mixed county:
|
||||
|
||||
$$
|
||||
\operatorname{passes}(c) \iff
|
||||
\left(F=\text{All}\right) \lor \left(K_c=F\right) \lor
|
||||
\left(K_c=\text{Mixed} \land F\in\{p_c,q_c\}\right).
|
||||
$$
|
||||
|
||||
Implementation: `scripts/build_county_koppen_metric.py` (`rank_class_shares`,
|
||||
`classify`, `build_koppen_records`) writes `data/metrics/koppen.csv`; area
|
||||
weighting, raster windows, and the 180th-meridian split are in
|
||||
`scripts/common/county_zonal_stats.py`;
|
||||
`scripts/apply_koppen_metric_to_climate_data.py` copies the three columns into
|
||||
`data/climate-data.csv`; the filter is `shouldFeaturePassFilter` in `app.js`.
|
||||
|
||||
**Rationale.** No published standard defines when an area is predominantly one
|
||||
Köppen class; the classification is defined per grid cell. The 50% condition
|
||||
means the label describes more than half of the county's land. The 5-point gap
|
||||
condition catches near 50/50 splits, such as Schenectady, NY (Dfb 50.1%,
|
||||
Dfa 49.9%), where a single label would rest on a margin of a few tenths of a
|
||||
point. Map-unit purity standards from other fields were considered and
|
||||
rejected: the FAO Land Cover Classification System treats a unit as single
|
||||
only above 80%, and USDA soil survey consociations allow roughly 15--25%
|
||||
dissimilar inclusions. Applied to counties, those thresholds would mark about
|
||||
40--47% of the map area as mixed. A published county-level Köppen dataset
|
||||
(Audirac, Harvard Dataverse, 2024) uses the plurality class and reports the
|
||||
share of every class, without a threshold.
|
||||
|
||||
**Results.** Using the Beck et al. 2023 1991--2020 1 km raster, for the 3,143
|
||||
counties in the 50 states and DC:
|
||||
|
||||
| Classification | Counties | Share of counties | Share of map area |
|
||||
| --- | --- | --- | --- |
|
||||
| Predominant (single class) | 3,010 | 95.8% | 85.6% |
|
||||
| Mixed climate | 133 | 4.2% | 14.4% |
|
||||
|
||||
Of the 133 mixed counties, 111 have no class covering 50% or more (37 of
|
||||
these also have a gap under 5 points), and 22 have a majority class whose
|
||||
runner-up is within 5 points. They are concentrated in the mountain West and
|
||||
Alaska: Colorado and Alaska (13 each), California and Montana (12 each),
|
||||
Washington (10), Idaho (9), and Utah (8). No county's top three classes fall
|
||||
within 2 points of one another, so no additional mixed categories are needed.
|
||||
|
||||
**Compared with the previous method** (below). Every county whose label changed
|
||||
became Mixed; no county moved to a different single class. The touched-cell
|
||||
counts and the area-weighted shares pick a different plurality winner in only
|
||||
3 counties (Denver, CO; Schenectady, NY; Hood River, OR), all of which are
|
||||
Mixed under this rule. Applying this rule to touched-cell counts instead of
|
||||
area-weighted shares would classify 9 counties differently, because
|
||||
touched-cell counts give full weight to boundary cells that are mostly
|
||||
outside the county.
|
||||
|
||||
**Boundary cases.** Five counties lie within 0.25 points of a cutoff that
|
||||
decides their outcome: Sanpete, UT (Dfb 49.93%, Mixed), Giles, VA (49.95%,
|
||||
Mixed), Placer, CA (Csa 50.22%), Albany, WY (Dfb 50.24%), and Park, MT (gap
|
||||
5.23 points, Dfb). Their classification depends on the precision of the
|
||||
area weighting. Aleutians West, AK (02016) spans the antimeridian, so its
|
||||
shares were computed without sub-cell sampling. Puerto Rico is not covered by
|
||||
these figures; it has 9 more Mixed counties and is not shown in the app.
|
||||
|
||||
### Previous method
|
||||
|
||||
Until 2026-09-13, `koppenZone` was the most frequent valid Köppen raster code
|
||||
among the cells touched by the county polygon:
|
||||
|
||||
$$
|
||||
K_c = \underset{k}{\arg\max}\;
|
||||
\sum_{i\in G_c}\mathbf{1}[K_i=k].
|
||||
$$
|
||||
|
||||
Cells were equally weighted, regardless of how much of each cell lay inside the
|
||||
county; an exact tie went to the smallest raster code, and a county with no
|
||||
valid cell was assigned `Cfa`. `scripts/build_county_climate_data.py` still
|
||||
computes this value until that script is retired (pipeline plan, Phase 3), so
|
||||
the Köppen apply step must run after it.
|
||||
|
||||
## 2. Annual Avg Temperature (Normals)
|
||||
|
||||
**Data key:** `avgTempF`
|
||||
|
||||
If the NOAA source is a historical monthly series, the script first forms a
|
||||
1991--2020 climatology for each calendar month and cell:
|
||||
|
||||
$$
|
||||
T_{i,m}^{\mathrm{norm}}
|
||||
=
|
||||
\frac{1}{Y_{i,m}}
|
||||
\sum_{y\in V_{i,m}}T_{i,m,y}.
|
||||
$$
|
||||
|
||||
The monthly county value is an unweighted mean of valid touched raster cells:
|
||||
|
||||
$$
|
||||
T_{c,m}
|
||||
=
|
||||
\frac{1}{|G_{c,m}|}
|
||||
\sum_{i\in G_{c,m}}T_{i,m}^{\mathrm{norm}}.
|
||||
$$
|
||||
|
||||
The annual value is the equally weighted mean of the available monthly values,
|
||||
converted from Celsius to Fahrenheit:
|
||||
|
||||
$$
|
||||
T_c(^\circ\mathrm{F})
|
||||
=
|
||||
\left(
|
||||
\frac{1}{M_c}\sum_{m\in V_c}T_{c,m}
|
||||
\right)\frac{9}{5}+32.
|
||||
$$
|
||||
|
||||
Normally \(M_c=12\). The stored value is rounded to 0.1 °F.
|
||||
|
||||
Implementation: `scripts/build_county_climate_data.py:124--166`,
|
||||
`scripts/build_county_climate_data.py:206--221`, and
|
||||
`scripts/build_county_climate_data.py:478--514`.
|
||||
|
||||
## 3. Diurnal Temperature Range
|
||||
|
||||
**Data key:** `avgDiurnalTempRangeF`
|
||||
|
||||
For each day with paired county Tmax and Tmin:
|
||||
|
||||
$$
|
||||
DTR_{c,d}=T^{\max}_{c,d}-T^{\min}_{c,d}.
|
||||
$$
|
||||
|
||||
Days with a missing input or a negative range are excluded. The final metric is
|
||||
the mean across all retained days in 1991--2020, followed by conversion of a
|
||||
Celsius temperature *difference* to a Fahrenheit difference:
|
||||
|
||||
$$
|
||||
\overline{DTR}_c(^\circ\mathrm{F})
|
||||
=
|
||||
\frac{9}{5}
|
||||
\left(
|
||||
\frac{1}{N_c}\sum_{d\in V_c}DTR_{c,d}
|
||||
\right).
|
||||
$$
|
||||
|
||||
There is correctly no \(+32\) term when converting a temperature difference.
|
||||
The output artifact stores two decimal places.
|
||||
|
||||
Implementation: `scripts/build_county_diurnal_temperature_range.py:46--104`
|
||||
and `scripts/build_county_diurnal_temperature_range.py:132--151`.
|
||||
|
||||
## 4. Annual Extreme Temperature Days
|
||||
|
||||
**Data key:** `absoluteExtremeDays`
|
||||
|
||||
A valid paired day is counted once when either the hot or cold absolute threshold
|
||||
is met:
|
||||
|
||||
$$
|
||||
E_{c,y}
|
||||
=
|
||||
\sum_{d\in V_{c,y}}
|
||||
\mathbf{1}\!\left[
|
||||
T^{\max}_{c,d}\ge95^\circ\mathrm{F}
|
||||
\;\lor\;
|
||||
T^{\min}_{c,d}\le0^\circ\mathrm{F}
|
||||
\right].
|
||||
$$
|
||||
|
||||
The county filter is the arithmetic mean of the yearly counts:
|
||||
|
||||
$$
|
||||
E_c=\frac{1}{Y_c}\sum_{y\in V_c}E_{c,y}.
|
||||
$$
|
||||
|
||||
The checked-in data uses 1991--2025. The Boolean OR means that a hypothetical
|
||||
day meeting both conditions is still counted only once. A year is included when
|
||||
an annual record exists; counts are not normalized to 365 or 366 valid days.
|
||||
|
||||
Implementation: `scripts/build_county_locally_extreme_data.py:473--540`,
|
||||
`scripts/build_county_locally_extreme_data.py:669--674`, and
|
||||
`scripts/build_county_locally_extreme_data.py:705--722`.
|
||||
|
||||
## 5. Annual 90 F+ Heat Index Days
|
||||
|
||||
**Data key:** `humidHeatDays`
|
||||
|
||||
The daily proxy pairs NOAA nClimGrid-Daily county Tmax, \(T\), with the gridMET
|
||||
county daily minimum relative humidity, \(R\). Relative humidity is clipped to
|
||||
\([0,100]\).
|
||||
|
||||
The NWS simple Heat Index estimate is calculated in two steps:
|
||||
|
||||
$$
|
||||
S=\frac{1}{2}\left[T+61+1.2(T-68)+0.094R\right],
|
||||
$$
|
||||
|
||||
$$
|
||||
HI_s=\frac{S+T}{2}.
|
||||
$$
|
||||
|
||||
When \(HI_s\ge80^\circ\mathrm{F}\), the Rothfusz regression is used:
|
||||
|
||||
$$
|
||||
\begin{aligned}
|
||||
HI_r={}&-42.379+2.04901523T+10.14333127R-0.22475541TR\\
|
||||
&-0.00683783T^2-0.05481717R^2+0.00122874T^2R\\
|
||||
&+0.00085282TR^2-0.00000199T^2R^2.
|
||||
\end{aligned}
|
||||
$$
|
||||
|
||||
For \(R<13\) and \(80\le T\le112\), subtract:
|
||||
|
||||
$$
|
||||
A_{low}
|
||||
=
|
||||
\frac{13-R}{4}
|
||||
\sqrt{\max\left(\frac{17-|T-95|}{17},0\right)}.
|
||||
$$
|
||||
|
||||
For \(R>85\) and \(80\le T\le87\), add:
|
||||
|
||||
$$
|
||||
A_{high}=\frac{R-85}{10}\frac{87-T}{5}.
|
||||
$$
|
||||
|
||||
Thus:
|
||||
|
||||
$$
|
||||
HI(T,R)=
|
||||
\begin{cases}
|
||||
HI_s, & HI_s<80,\\
|
||||
HI_r-A_{low}, & HI_s\ge80 \text{ and the low-RH condition holds},\\
|
||||
HI_r+A_{high}, & HI_s\ge80 \text{ and the high-RH condition holds},\\
|
||||
HI_r, & \text{otherwise}.
|
||||
\end{cases}
|
||||
$$
|
||||
|
||||
The yearly and long-term values are:
|
||||
|
||||
$$
|
||||
H_{c,y}=\sum_{d\in V_{c,y}}
|
||||
\mathbf{1}[HI(T^{\max}_{c,d},R^{\min}_{c,d})\ge90],
|
||||
$$
|
||||
|
||||
$$
|
||||
H_c=\frac{1}{Y_c}\sum_{y\in V_c}H_{c,y}.
|
||||
$$
|
||||
|
||||
The period is 1991--2020. A year with at least one valid paired day contributes
|
||||
equally to the final average; there is no completeness adjustment.
|
||||
|
||||
Implementation: `scripts/summarize_county_gridmet_humidity.py:439--474` and
|
||||
`scripts/summarize_county_gridmet_humidity.py:547--617`.
|
||||
|
||||
## 6. Annual Precipitation (Normals)
|
||||
|
||||
**Data key:** `annualPrecipIn`
|
||||
|
||||
Each monthly county total, \(P_{c,m}\), is the unweighted mean of valid raster
|
||||
cells touched by the county. Annual precipitation is the sum of available
|
||||
monthly totals, converted from millimeters to inches:
|
||||
|
||||
$$
|
||||
P_c(\mathrm{in})
|
||||
=
|
||||
\frac{1}{25.4}
|
||||
\sum_{m\in V_c}P_{c,m}(\mathrm{mm}).
|
||||
$$
|
||||
|
||||
The stored value is rounded to 0.1 inch. The implementation only requires one
|
||||
valid month, so missing months produce a partial annual sum rather than a blank.
|
||||
|
||||
Implementation: `scripts/build_county_climate_data.py:401--412` and
|
||||
`scripts/build_county_climate_data.py:478--511`.
|
||||
|
||||
## 7. Seasonality Index
|
||||
|
||||
**Data key:** `seasonalityIndex`
|
||||
|
||||
The filter is the coefficient of variation of available monthly precipitation
|
||||
totals. First calculate the monthly mean and population standard deviation:
|
||||
|
||||
$$
|
||||
\mu_c=\frac{1}{M_c}\sum_{m\in V_c}P_{c,m},
|
||||
$$
|
||||
|
||||
$$
|
||||
\sigma_c=
|
||||
\sqrt{
|
||||
\frac{1}{M_c}
|
||||
\sum_{m\in V_c}(P_{c,m}-\mu_c)^2
|
||||
}.
|
||||
$$
|
||||
|
||||
Then:
|
||||
|
||||
$$
|
||||
SI_c=
|
||||
\operatorname{round}\!\left(
|
||||
\operatorname{clip}\!\left(
|
||||
100\frac{\sigma_c}{\mu_c},0,100
|
||||
\right)
|
||||
\right).
|
||||
$$
|
||||
|
||||
If \(\mu_c\le0\), the index is set to zero. The value is stored as an integer.
|
||||
|
||||
Implementation: `scripts/build_county_climate_data.py:487--521`.
|
||||
|
||||
## 8. Wettest Month
|
||||
|
||||
**Data key:** `wettestPrecipMonth`
|
||||
|
||||
$$
|
||||
W_c=\underset{m\in V_c}{\arg\max}\;P_{c,m}.
|
||||
$$
|
||||
|
||||
Missing monthly values are ignored. An exact tie resolves to the earliest tied
|
||||
month because `numpy.nanargmax` returns the first occurrence.
|
||||
|
||||
Implementation:
|
||||
`scripts/apply_precipitation_month_metrics_to_climate_data.py:56--90`.
|
||||
|
||||
## 9. Driest Month
|
||||
|
||||
**Data key:** `driestPrecipMonth`
|
||||
|
||||
$$
|
||||
D_c=\underset{m\in V_c}{\arg\min}\;P_{c,m}.
|
||||
$$
|
||||
|
||||
Missing monthly values are ignored. An exact tie likewise resolves to the
|
||||
earliest tied month.
|
||||
|
||||
Implementation:
|
||||
`scripts/apply_precipitation_month_metrics_to_climate_data.py:56--90`.
|
||||
|
||||
## 10. Summer Specific Humidity
|
||||
|
||||
**Data key:** `avgSummerSpecificHumidityGKg`
|
||||
|
||||
gridMET cells whose centers fall inside a county are weighted by the cosine of
|
||||
their latitude to approximate their relative surface areas on a latitude--longitude
|
||||
grid:
|
||||
|
||||
$$
|
||||
q_{c,d}
|
||||
=
|
||||
\frac{
|
||||
\sum_{i\in G_{c,d}}q_{i,d}\cos(\phi_i)
|
||||
}{
|
||||
\sum_{i\in G_{c,d}}\cos(\phi_i)
|
||||
}.
|
||||
$$
|
||||
|
||||
The metric averages the valid daily county values for June, July, and August,
|
||||
then converts kg/kg to g/kg:
|
||||
|
||||
$$
|
||||
q_c^{\mathrm{summer}}
|
||||
=
|
||||
1000\left(
|
||||
\frac{1}{N_c}
|
||||
\sum_{d\in V_c,\;m(d)\in\{6,7,8\}}q_{c,d}
|
||||
\right).
|
||||
$$
|
||||
|
||||
If no grid-cell center falls inside a county, the nearest grid cell to an
|
||||
interior representative point is used.
|
||||
|
||||
Implementation: `scripts/summarize_county_gridmet_humidity.py:297--380`,
|
||||
`scripts/summarize_county_gridmet_humidity.py:409--434`, and
|
||||
`scripts/summarize_county_gridmet_humidity.py:555--605`.
|
||||
|
||||
## 11. Mean Daily Global Horizontal Radiation (GHI)
|
||||
|
||||
**Data key:** `meanDailyGlobalHorizontalRadiationKwhM2Day`
|
||||
|
||||
For each NSRDB site, the current 60-minute, non-leap-year TMY data is summarized
|
||||
as:
|
||||
|
||||
$$
|
||||
G_s
|
||||
=
|
||||
\frac{\sum_h GHI_{s,h}}{1000\times365}
|
||||
\quad\mathrm{kWh/m^2/day}.
|
||||
$$
|
||||
|
||||
For a county with polygon archive coverage:
|
||||
|
||||
$$
|
||||
G_c
|
||||
=
|
||||
\frac{\sum_s A_{c,s}G_s}{\sum_s A_{c,s}},
|
||||
$$
|
||||
|
||||
where \(A_{c,s}\) is the estimated overlap area between the county geometry and
|
||||
the 4 km square grid cell centered on site \(s\). A representative-point value
|
||||
is used when a polygon summary is unavailable.
|
||||
|
||||
Implementation: `scripts/summarize_nsrdb_county_polygon_archives.py:189--192`
|
||||
and `scripts/summarize_nsrdb_county_polygon_archives.py:335--378`.
|
||||
|
||||
## 12. Clear-Sky GHI Reduction Index
|
||||
|
||||
**Data key:** `clearSkyGhiReductionIndex`
|
||||
|
||||
Only rows with valid observed and clear-sky GHI and
|
||||
\(CSGHI_{s,h}\ge50\;\mathrm{W/m^2}\) are treated as daylight rows. For each
|
||||
retained row:
|
||||
|
||||
$$
|
||||
r_{s,h}
|
||||
=
|
||||
\operatorname{clip}\!\left(
|
||||
\frac{GHI_{s,h}}{CSGHI_{s,h}},0,1
|
||||
\right).
|
||||
$$
|
||||
|
||||
The site reduction index is:
|
||||
|
||||
$$
|
||||
R_s=1-\frac{1}{N_s}\sum_{h\in V_s}r_{s,h}.
|
||||
$$
|
||||
|
||||
The polygon county value is overlap-area-weighted:
|
||||
|
||||
$$
|
||||
R_c=\frac{\sum_s A_{c,s}R_s}{\sum_s A_{c,s}}.
|
||||
$$
|
||||
|
||||
A representative-point index is used where a polygon summary is unavailable.
|
||||
This definition is the mean of time-row ratios; it is not generally equal to
|
||||
\(1-\sum GHI/\sum CSGHI\).
|
||||
|
||||
Implementation:
|
||||
`scripts/summarize_nsrdb_county_polygon_cloud_archives.py:138--218` and
|
||||
`scripts/summarize_nsrdb_county_polygon_cloud_archives.py:222--304`.
|
||||
|
||||
## Calculation review findings
|
||||
|
||||
1. **Annual temperature weights months equally.** February has the same weight
|
||||
as January or July. If the intended label means an average across all days,
|
||||
monthly normals should instead be weighted by the number of days in each
|
||||
month.
|
||||
|
||||
2. **Base NOAA aggregation is not area-weighted.** Every touched raster cell
|
||||
receives equal weight, including cells that intersect only a small portion
|
||||
of a county. This can matter most for small or narrow counties and along
|
||||
coastlines. *Resolved for Köppen on 2026-09-13: class shares are now
|
||||
area-weighted (§1).*
|
||||
|
||||
3. **Partial precipitation years are accepted.** One valid monthly precipitation
|
||||
value is sufficient to produce `annualPrecipIn`; absent months silently lower
|
||||
the annual sum. Requiring all 12 months, or recording completeness, would be
|
||||
safer.
|
||||
|
||||
4. **The Köppen fallback can create false data.** A county with no valid raster
|
||||
cells is labeled `Cfa` instead of missing. A null value plus an audit flag
|
||||
would distinguish missing coverage from a genuine humid-subtropical class.
|
||||
*Resolved on 2026-09-13: the Köppen builder leaves such counties blank. The
|
||||
fallback remains only in `build_county_climate_data.py`, whose Köppen
|
||||
value is replaced by the apply step.*
|
||||
|
||||
5. **Extreme-day counts are not completeness-normalized.** A partially observed
|
||||
year contributes a raw count and receives the same weight as a complete year.
|
||||
Consider requiring a minimum number of valid days or annualizing partial
|
||||
counts explicitly.
|
||||
|
||||
6. **The absolute-extreme metric depends on unrelated percentile thresholds.**
|
||||
`build_annual_counts` skips a county when its retired local p95/p05 thresholds
|
||||
are missing, even though the active 95 °F / 0 °F calculation does not require
|
||||
those percentiles. The absolute calculation should be separated from that
|
||||
prerequisite.
|
||||
|
||||
7. **Heat Index days are a daily-extrema proxy.** Daily Tmax and daily minimum
|
||||
relative humidity are paired even though their observation times may differ.
|
||||
The result should not be described as an observed hourly maximum Heat Index.
|
||||
|
||||
8. **Heat-year completeness is permissive.** Any year with at least one valid
|
||||
Tmax/RH pair is included in the equal-year average. A minimum valid-day rule
|
||||
would reduce low-biased partial-year counts.
|
||||
|
||||
9. **The GHI formula assumes hourly, 365-day input.** It is correct for the
|
||||
current 60-minute, `leap_day=false` requests. If the request interval changes,
|
||||
the energy sum needs an interval-hours multiplier; leap-day handling would
|
||||
also need to change the divisor.
|
||||
|
||||
10. **Clear-sky reduction averages ratios rather than energy totals.** This is a
|
||||
valid but specific definition. It gives each retained time row equal weight,
|
||||
rather than weighting rows by available clear-sky energy. The label and
|
||||
documentation should retain this distinction.
|
||||
|
||||
11. **Spatial weighting is inconsistent across metric families.** Base NOAA
|
||||
normals use equal touched-cell weights, Köppen uses area-weighted class
|
||||
shares, gridMET humidity uses
|
||||
\(\cos(\phi)\) weights on cell centers, and NSRDB polygon metrics use
|
||||
estimated overlap areas. Cross-metric comparisons should account for these
|
||||
different county aggregation methods.
|
||||
|
||||
## Verification status
|
||||
|
||||
This reference was derived from the checked-in calculation and merge scripts,
|
||||
not solely from UI descriptions. No calculation code was changed. The automated
|
||||
test suite was not executed during this review because `pytest` is not installed
|
||||
in either the system Python environment or the project virtual environment.
|
||||
|
||||
Section 1 and findings 2, 4, and 11 were updated on 2026-09-14, after the
|
||||
Köppen classification was reworked and applied.
|
||||
@@ -0,0 +1,198 @@
|
||||
# Mixed Climate Display
|
||||
|
||||
**Status (2026-09-14):** Complete. The app draws Mixed counties as stripes, the
|
||||
pipeline writes the two stripe-class columns, and `data/climate-data.csv` has
|
||||
been updated (142 counties Mixed, 133 of them in the 50 states and DC). The
|
||||
display was reviewed in the browser and refined (see the history in section 7),
|
||||
and the project documentation was updated (section 6).
|
||||
|
||||
**Scope:** Show Köppen-Geiger "Mixed" counties as striped on the map, filter them
|
||||
by their top two classes, and carry the two stripe classes from the pipeline into
|
||||
the app. The classification rule itself is described in
|
||||
[filter-calculations.md](filter-calculations.md) §1; overall progress is tracked in
|
||||
[pipeline-plan.md](pipeline-plan.md).
|
||||
|
||||
## 1. Design
|
||||
|
||||
- **Only Mixed counties are striped.** A county is Mixed when its top class covers
|
||||
less than 50% of its land or leads the runner-up by less than 5 percentage
|
||||
points. Every other county keeps its solid class color.
|
||||
- **Stripe colors** are the county's top class (primary) and runner-up class
|
||||
(secondary), using the existing Köppen class colors.
|
||||
- **Stripes are diagonal at 45°**, from upper left to lower right, and run
|
||||
unbroken to the county boundary. Neighboring Mixed counties share the same
|
||||
angle and spacing, so stripes line up across borders. The county border is
|
||||
drawn on top.
|
||||
- **Stripes follow the map.** The density is exact at the default view
|
||||
(`DEFAULT_COUNTRY_VIEW`, the project's single reference scale). From there the
|
||||
stripe pair width doubles with each zoom-in step, all the way to the map's
|
||||
deepest zoom (19), so the number of stripes in a county stays the same at
|
||||
every zoom.
|
||||
- **Far-out style.** When stripes would drop below 5 pixels per pair (zoom 4 and
|
||||
below), the map-following width is doubled until it reaches 5 pixels (at
|
||||
least once, so there are half as many lines), giving 8.53-pixel pairs split
|
||||
70/30 so the secondary stays visible.
|
||||
- **Stripes are fixed to the ground.** Each pattern is anchored to the map
|
||||
projection's fixed origin, so a stripe stays on the same ground at every zoom
|
||||
from 5 to 19. At the far-out zooms every other stripe stays put while the
|
||||
crossfade blends the rest.
|
||||
- **Fades happen only when the stripe size changes** (the switch between the
|
||||
standard and far-out styles, and zoom steps within the far-out range). The fade
|
||||
is a 250 ms true crossfade that runs just after Leaflet's zoom animation: the
|
||||
old stripes, exactly as the animation left them, blend into the new ones with
|
||||
no jump and no gap.
|
||||
- **Filtering.** Choosing a class shows counties that are predominantly that
|
||||
class plus Mixed counties where it is the primary or secondary class. Third and
|
||||
lower classes do not match. One **"Mixed Climate"** option shows all Mixed
|
||||
counties; there are no per-pair Mixed options.
|
||||
- **Every class on the map is listed.** Classes that appear only as stripe
|
||||
colors (today Dsc and Dwc, both in Alaska) are listed in the dropdown and
|
||||
legend as regular classes.
|
||||
- **Legend.** The "Mixed Climate" row has a swatch of two neutral grays in the
|
||||
standard split (lighter gray primary, darker gray secondary), generated from
|
||||
the split setting so it always matches the map, and a note that stripes show
|
||||
each county's top two classes.
|
||||
- **Deferred:** a donut chart of every class share (like Europa Universalis 5).
|
||||
|
||||
### Settings
|
||||
|
||||
All stripe settings sit together near the top of `app.js`:
|
||||
|
||||
| Setting | Value | Controls |
|
||||
| --- | --- | --- |
|
||||
| `KOPPEN_MIXED_PRIMARY_STRIPE_FRACTION` | 0.8 | Primary share of each stripe pair, standard style (zoom 5 and closer) |
|
||||
| `KOPPEN_MIXED_LINE_PAIRS_PER_100_MILES` | 5 | Stripe pairs per 100 miles, measured across the stripes at the default view |
|
||||
| `KOPPEN_MIXED_MIN_PAIR_WIDTH_PX` | 5 | Below this pair width, stripes switch to the far-out style |
|
||||
| `KOPPEN_MIXED_FAR_PRIMARY_STRIPE_FRACTION` | 0.7 | Primary share of each stripe pair, far-out style |
|
||||
| `KOPPEN_MIXED_STYLE_FADE_MS` | 250 | Length of the crossfade |
|
||||
| `KOPPEN_MIXED_STRIPE_ROTATION_DEGREES` | -45 | Stripe angle (SVG rotation; -45 runs upper left to lower right) |
|
||||
| `KOPPEN_MIXED_LEGEND_PRIMARY_GRAY`, `KOPPEN_MIXED_LEGEND_SECONDARY_GRAY` | `#d4d8df`, `#6b7383` | Legend swatch colors |
|
||||
|
||||
## 2. Data
|
||||
|
||||
`data/climate-data.csv` has two columns directly after `koppenZone`:
|
||||
|
||||
| Column | Filled when | Value |
|
||||
| --- | --- | --- |
|
||||
| `koppenPrimaryClass` | `koppenZone` is `Mixed` | Top class, e.g. `Csb` |
|
||||
| `koppenSecondaryClass` | `koppenZone` is `Mixed` | Runner-up class, e.g. `Dsb` |
|
||||
|
||||
Both are blank for predominant counties. The values come from `koppenTopClass`
|
||||
and `koppenSecondClass` in `data/metrics/koppen.csv`.
|
||||
|
||||
## 3. How the stripes are drawn
|
||||
|
||||
The map uses Leaflet 1.9.4 with its default SVG renderer. Each ordered
|
||||
primary/secondary pair gets one SVG `<pattern>`, kept in a small hidden SVG added
|
||||
to the page; fill references such as `url(#id)` resolve anywhere in the page, so
|
||||
Leaflet's own SVG elements are not modified. The 133 Mixed counties in the
|
||||
50 states and DC use 50 pairs. Order matters: Csb/Dsb and Dsb/Csb give the
|
||||
primary share to different colors.
|
||||
|
||||
- **Pattern contents.** A primary-color background and a group of
|
||||
secondary-color bands. Normally the group has one band per stripe pair; during
|
||||
a crossfade it holds shared, old-only, and new-only bands.
|
||||
- **Coordinates.** `patternUnits="userSpaceOnUse"`, so every county is filled in
|
||||
the map's coordinate space and stripes align across borders.
|
||||
- **Anchoring.** `patternTransform="rotate(-45) translate(x y)"`. Leaflet's SVG
|
||||
coordinates start from a pixel origin that moves on every zoom, so the
|
||||
translate places the pattern's origin at the map projection's fixed origin,
|
||||
reduced to within one tile to avoid precision loss at deep zooms.
|
||||
- **Fill.** A Mixed county's style sets
|
||||
`fillColor: "url(#koppen-mixed-<primary>-<secondary>)"`.
|
||||
|
||||
**Pair width at the default view.** Leaflet uses Web Mercator, where
|
||||
|
||||
$$
|
||||
\text{meters per pixel} = \frac{156{,}543.03 \times \cos(\text{latitude})}{2^{\text{zoom}}}.
|
||||
$$
|
||||
|
||||
At the default view (latitude 39.5°, zoom 5) that is about 3,775 m per pixel,
|
||||
so 100 miles (160,934 m) is about 42.6 pixels. At 5 pairs per 100 miles, one
|
||||
stripe pair is about 8.5 pixels: about 6.8 pixels of primary color and
|
||||
1.7 pixels of secondary color. The app computes this from the settings.
|
||||
|
||||
| Zoom | Style | Pair width | Split |
|
||||
| --- | --- | --- | --- |
|
||||
| 2–4 | Far-out | 8.53 px | 70/30 |
|
||||
| 5 (default) | Standard | 8.53 px | 80/20 |
|
||||
| 6 | Standard | 17.05 px | 80/20 |
|
||||
| 7 | Standard | 34.11 px | 80/20 |
|
||||
| 8 | Standard | 68.2 px | 80/20 |
|
||||
| 9–19 | Standard | Doubling each step, to 139,704 px at zoom 19 | 80/20 |
|
||||
|
||||
**Zooming.** During Leaflet's zoom animation the whole SVG layer is scaled with
|
||||
the map, so the stripes stretch with it. When the zoom ends, the stripes are
|
||||
recomputed. Because every width is the reference width times a power of two, the
|
||||
old stripes (previous width × 2 to the power of the zoom change) and the new
|
||||
stripes fit one pattern tile. If the size is unchanged, the new layout is drawn
|
||||
directly; if it changed, bands secondary in both layouts stay solid, old-only
|
||||
bands fade out, and new-only bands fade in.
|
||||
|
||||
## 4. App code (`app.js`)
|
||||
|
||||
| Area | Functions and settings |
|
||||
| --- | --- |
|
||||
| Stripe settings | The `KOPPEN_MIXED_*` constants near the top of the file (section 1) |
|
||||
| Mixed class and parsing | `KOPPEN_CLASS_META` (`Mixed`, labelled "Mixed Climate", sorted last), `normalizeKoppenCode`, `sanitizeOverrideRecord` (reads the two stripe-class columns) |
|
||||
| Stripe styles | `getKoppenMixedStripePairWidthPx`, `getKoppenStripeStyle`, `getKoppenStripeClasses`, `getKoppenStripePatternId` |
|
||||
| Patterns | `createKoppenStripePatterns`, `buildKoppenStripePattern`, `drawKoppenStripeBands`, `applyKoppenStripeStyle`, `anchorKoppenStripePatterns` |
|
||||
| Zooming and crossfade | `updateKoppenStripesForZoom`, `crossfadeKoppenStripes`, `getKoppenSecondaryBands`, `classifyKoppenStripeBands`, `canCrossfadeKoppenStripes` |
|
||||
| Fill, filter, and labels | `fillForCounty`, `shouldFeaturePassFilter`, `getUniqueCategoryValuesInData`, `getCategoricalDisplayLabel`, `formatCountyMetricValue` |
|
||||
| Legend | `updateLegend`, `buildKoppenMixedLegendSwatchBackground` |
|
||||
|
||||
The hover tooltip shows only the county name, and `styles.css` is unchanged.
|
||||
After changing `app.js`, bump `APP_ASSET_VERSION` and the `app.js?v=` query in
|
||||
`index.html` so browsers load the new files.
|
||||
|
||||
## 5. Pipeline
|
||||
|
||||
- **`scripts/build_county_koppen_metric.py`** writes `data/metrics/koppen.csv`
|
||||
with each county's class, top and runner-up class, and their shares.
|
||||
- **`scripts/apply_koppen_metric_to_climate_data.py`** writes `koppenZone`,
|
||||
`koppenPrimaryClass`, and `koppenSecondaryClass` into `data/climate-data.csv`,
|
||||
adding the two stripe columns after `koppenZone` if missing and filling them
|
||||
only for Mixed counties. `--dry-run` reports changes without writing.
|
||||
- **`scripts/check_climate_data.py`** requires the two stripe columns: filled if
|
||||
and only if `koppenZone` is `Mixed`, with valid and different Köppen codes.
|
||||
- **Tests:** `tests/test_koppen_metric.py` and `tests/test_check_climate_data.py`.
|
||||
|
||||
## 6. Documentation (stage 5, done 2026-09-14)
|
||||
|
||||
- `filter-calculations.md` §1 describes the applied rule, the stripe columns,
|
||||
and the filter; the old method is kept as a short "Previous method" note.
|
||||
Review findings 2 and 4 are marked resolved for Köppen.
|
||||
- `README.md` and `scripts/county_data_sources.md` describe the new
|
||||
classification and list the Köppen build, apply, and check commands.
|
||||
|
||||
## 7. History
|
||||
|
||||
Decisions that were reversed or refined during the browser review:
|
||||
|
||||
- **2026-09-13 — Stripes follow the map.** Stripes were first a fixed width on
|
||||
screen; the owner wanted the number of stripes in a county not to change with
|
||||
zoom.
|
||||
- **2026-09-13 — True crossfade.** The first fade faded the stripes out and back
|
||||
in, which briefly left no stripes; it was replaced by a crossfade from the old
|
||||
stripes to the new ones.
|
||||
- **2026-09-13 — Crossfade after the zoom animation.** Starting zoom-in fades
|
||||
together with the animation was tried and reverted: the perceived slowness had
|
||||
been the map's own zoom animation.
|
||||
- **2026-09-13/14 — Near style removed.** A thicker 70/30 split from zoom 9 was
|
||||
added, then made instant (fades happen only when the stripe size changes), set
|
||||
to 80/20 by the owner, and finally removed.
|
||||
- **2026-09-14 — Stripes fixed to the ground.** Stripes slid across the land on
|
||||
each zoom because patterns were anchored to Leaflet's per-zoom pixel origin.
|
||||
- **2026-09-14 — No deepest-zoom limit.** A 60-pixel cap, later a 1,100-pixel
|
||||
safety limit with halving, was removed: any cap forces stripes to thin or
|
||||
subdivide on screen, and tests showed no performance cost without one.
|
||||
- **2026-09-14 — Zoom response setting removed.** The fixed-to-the-ground
|
||||
anchoring and the crossfade both require stripes to follow the map exactly.
|
||||
|
||||
## 8. Puerto Rico
|
||||
|
||||
The app already leaves Puerto Rico off the map: `prepareCountyFeature` in
|
||||
`app.js` drops counties with state FIPS 72, and the dropdown and legend are built
|
||||
from the counties on the map. The rows remain in `data/climate-data.csv` and
|
||||
`data/metrics/koppen.csv`. Whether to also remove them from the data files is a
|
||||
separate decision.
|
||||
@@ -0,0 +1,291 @@
|
||||
# Data Pipeline Improvement Plan
|
||||
|
||||
This plan describes how the county data pipeline will move from scripts that
|
||||
edit one shared CSV in place to per-metric outputs assembled into the app CSV.
|
||||
It is a working reference for the filter-by-filter review. Calculation details
|
||||
for each filter live in [filter-calculations.md](filter-calculations.md).
|
||||
|
||||
**Started:** 2026-09-12
|
||||
|
||||
**Guiding decision:** review and fix each of the 12 filters one at a time,
|
||||
confirm each works on its own, and restructure `data/climate-data.csv` only
|
||||
after all filters are clean. No large rewrite happens up front.
|
||||
|
||||
## 1. Current pipeline
|
||||
|
||||
### Where each filter comes from
|
||||
|
||||
| Filter | Written into `climate-data.csv` by | Upstream scripts |
|
||||
| --- | --- | --- |
|
||||
| Köppen-Geiger class (plus the two stripe-class columns) | `apply_koppen_metric_to_climate_data.py` | `build_county_koppen_metric.py` → `data/metrics/koppen.csv` |
|
||||
| Annual avg temperature | `build_county_climate_data.py` | — |
|
||||
| Annual precipitation | `build_county_climate_data.py` | — |
|
||||
| Seasonality index | `build_county_climate_data.py` | — |
|
||||
| Wettest / driest month | Base build, then overwritten by `apply_precipitation_month_metrics_to_climate_data.py` | — |
|
||||
| Diurnal temperature range | `apply_diurnal_temperature_range_to_climate_data.py` | `build_county_diurnal_temperature_range.py` |
|
||||
| Extreme temperature days | `apply_locally_extreme_metric_to_climate_data.py` | `build_county_locally_extreme_data.py` |
|
||||
| Summer specific humidity | `apply_gridmet_humidity_metric_to_climate_data.py` | `download_gridmet_data.py` → `summarize_county_gridmet_humidity.py` |
|
||||
| 90 °F+ heat-index days (plus 2 source-FIPS columns) | `apply_gridmet_humidity_metric_to_climate_data.py` | Same as summer humidity |
|
||||
| Solar GHI | Base build (optional), then replaced by `apply_locally_extreme_metric_to_climate_data.py` | Point: `build_county_representative_points.py` → `fetch_nsrdb_representative_point_ghi.py`. Polygon: `request_nsrdb_county_polygon_ghi_archives.py` → `download_nsrdb_county_polygon_ghi_archives.py` → `summarize_nsrdb_county_polygon_archives.py` |
|
||||
| Clear-sky GHI reduction | `apply_nsrdb_cloud_metric_to_climate_data.py` | Point: `fetch_nsrdb_representative_point_cloud_metrics.py`. Polygon: `request_nsrdb_county_polygon_cloud_archives.py` → `download_nsrdb_county_polygon_cloud_archives.py` → `summarize_nsrdb_county_polygon_cloud_archives.py` |
|
||||
|
||||
Supporting scripts: `request_nsrdb_county_polygon_archives.py` and
|
||||
`download_nsrdb_county_polygon_archives.py` are the shared engines behind the
|
||||
GHI and cloud wrappers; `rebuild_nsrdb_representative_point_ghi_summary.py`
|
||||
rebuilds the point GHI summary from cache; `check_climate_data.py` validates the
|
||||
final CSV. Shared helpers live in `scripts/common/` (Phase 2). The base build
|
||||
still writes an old largest-share `koppenZone`, so the Köppen apply step must
|
||||
run after it.
|
||||
|
||||
### Current full-rebuild order
|
||||
|
||||
1. `build_county_climate_data.py`
|
||||
2. `build_county_koppen_metric.py` → `apply_koppen_metric_to_climate_data.py`
|
||||
3. `apply_precipitation_month_metrics_to_climate_data.py`
|
||||
4. `build_county_locally_extreme_data.py` → `apply_locally_extreme_metric_to_climate_data.py`
|
||||
5. `build_county_diurnal_temperature_range.py` → `apply_diurnal_temperature_range_to_climate_data.py`
|
||||
6. `summarize_county_gridmet_humidity.py` → `apply_gridmet_humidity_metric_to_climate_data.py`
|
||||
7. `apply_nsrdb_cloud_metric_to_climate_data.py`
|
||||
|
||||
The README's enrichment list starts at step 2 and omits steps 1 and 3.
|
||||
|
||||
### Problems
|
||||
|
||||
1. **Rerunning a step can destroy data.** The base build writes 12 columns,
|
||||
including the retired `extremeDays`. Later scripts delete, overwrite, or add
|
||||
columns until the live CSV has 20. Rerunning the base build drops
|
||||
9 live columns (`koppenPrimaryClass`, `koppenSecondaryClass`,
|
||||
`avgDiurnalTempRangeF`, `absoluteExtremeDays`,
|
||||
`clearSkyGhiReductionIndex`, `avgSummerSpecificHumidityGKg`,
|
||||
`humidHeatDays`, `humidHeatSourceFips`, `humidHeatFipsAdjustment`),
|
||||
restores `extremeDays`, and rewrites `koppenZone` with the old
|
||||
largest-share method.
|
||||
2. **Order is implicit.** The sequence lives in the README, in
|
||||
`scripts/county_data_sources.md`, and in each script's assumptions.
|
||||
3. **Column ownership is unclear.** GHI is finalized by the extreme-temperature
|
||||
apply script; wettest/driest month are computed in two places.
|
||||
4. **County aggregation is inconsistent.** NOAA uses touched raster cells,
|
||||
Köppen uses area-weighted shares (since 2026-09-13), gridMET uses cell
|
||||
centers with cos(latitude) weights, and NSRDB uses overlap areas (finding 11
|
||||
in `filter-calculations.md`).
|
||||
5. **No single entry point or final check.** A new user must piece together
|
||||
about 20 scripts, several large downloads, and an NSRDB API key.
|
||||
|
||||
## 2. Target design
|
||||
|
||||
1. **One metric, one file.** Each metric pipeline writes a county-level file
|
||||
under `data/metrics/`, for example `data/metrics/koppen.csv`, containing
|
||||
`countyFips`, the app value, and any audit columns for that metric.
|
||||
2. **One assemble step.** A single script joins the metric files into
|
||||
`data/climate-data.csv`, using `data/metric_sources.json` for the column list
|
||||
and per-metric source notes, then runs `check_climate_data.py`.
|
||||
- Run order no longer matters; rerunning one metric cannot damage others.
|
||||
- Every column has exactly one owner.
|
||||
- The per-row `source` column moves into `metric_sources.json`.
|
||||
- Audit columns stay in the metric files rather than the app CSV.
|
||||
3. **One shared county-aggregation module.** Area-weighted zonal statistics,
|
||||
including the 180th-meridian split, used by every raster-based metric.
|
||||
4. **One runner.** For example
|
||||
`python scripts/pipeline.py --only koppen --skip-download`, with stages for
|
||||
fetch, build metrics, assemble, and check. Cached downloads are reused by
|
||||
default.
|
||||
|
||||
## 3. Reproduction tiers
|
||||
|
||||
The "Reproducing the data" guide (Phase 4) will be organized by how deep a
|
||||
user needs to go:
|
||||
|
||||
| Tier | What the user does | Needs |
|
||||
| --- | --- | --- |
|
||||
| 1. Run the app | `.\serve.ps1` with the committed CSV | Nothing else |
|
||||
| 2. Reassemble | Rebuild `climate-data.csv` from committed metric files | Python environment only |
|
||||
| 3. Regenerate one metric | Download one source, rebuild one metric file, reassemble | That metric's source data |
|
||||
| 4. Full rebuild | Everything | All sources; NSRDB API key; large downloads (the NOAA monthly temperature file alone is about 5 GB) and hours of paced NSRDB requests |
|
||||
|
||||
The guide will list each dataset's size, download location, API-key needs, and
|
||||
approximate run time, and the Python requirements will be pinned.
|
||||
|
||||
## 4. Roadmap
|
||||
|
||||
### Phase 0 — Groundwork (done)
|
||||
|
||||
- [x] `scripts/check_climate_data.py` validates the app CSV (8 checks) with
|
||||
tests in `tests/test_check_climate_data.py`.
|
||||
- [x] `data/metric_sources.json` created as an empty skeleton.
|
||||
- [x] Köppen raster reads use a padded window per county, and polygons that
|
||||
cross the 180th meridian are split (`split_at_antimeridian` in
|
||||
`scripts/common/county_zonal_stats.py`); output verified identical for
|
||||
all 3,221 counties; tests in `tests/test_koppen_antimeridian.py`.
|
||||
|
||||
### Phase 1 — Filter-by-filter review (in progress)
|
||||
|
||||
Each filter goes through the checklist in section 5. Each fix delivers that
|
||||
metric's own file in `data/metrics/` plus a single-column apply step, so the
|
||||
existing CSV keeps working until Phase 3.
|
||||
|
||||
### Phase 2 — Shared helpers in `scripts/common/`
|
||||
|
||||
`scripts/common/` holds code used by more than one data source (NOAA, gridMET,
|
||||
NSRDB, Köppen). Scripts import from it, for example
|
||||
`from common.counties import load_counties`; nothing in it is run directly.
|
||||
|
||||
Rules for `common/`:
|
||||
|
||||
- **Cross-source only.** A helper goes in only if metrics from more than one
|
||||
data source use it. Code shared by scripts of a single data source stays with
|
||||
that source, for example a future NSRDB module for the NSRDB prompt,
|
||||
redaction, and error-log helpers.
|
||||
- **One topic per module.** Each module is named for its topic and has a
|
||||
docstring. No catch-all `utils.py`.
|
||||
- **Keep it small.** Before adding a helper, ask why it does not belong to any
|
||||
one data source.
|
||||
|
||||
Current modules:
|
||||
|
||||
| Module | Contents | Why it is in `common/` |
|
||||
| --- | --- | --- |
|
||||
| `county_zonal_stats.py` | Raster windows, the 180th-meridian split, area-weighted class shares | Used by any raster-based metric |
|
||||
| `counties.py` | County polygon loading, FIPS normalization, the state FIPS table | County identity is shared by nearly every pipeline |
|
||||
| `koppen_legend.py` | The Köppen code map and legend loader | Temporary: also used by `build_county_climate_data.py`; moves next to the Köppen code in Phase 3 |
|
||||
|
||||
**Rationale.** Helper functions make each step of a computation explicit,
|
||||
avoid repeated code, and can be tested separately
|
||||
([Brown CSCI 0111, "Helper Functions"](https://cs.brown.edu/courses/csci0111/fall2018/lectures/helper-functions.html)).
|
||||
Shared helper folders, however, tend to lose cohesion and collect unrelated
|
||||
code; the recommended alternative is to keep code with the part of the system
|
||||
it belongs to, allowing a shared folder only if it stays small and documented
|
||||
([Helpers and Utils Folders in Software Architecture](https://dev.to/knzt/helpers-and-utils-folders-in-software-architecture-3f8h)).
|
||||
The rules above follow both: shared functions, organized by topic and limited
|
||||
to code that crosses data sources.
|
||||
|
||||
Other shared code moves when its filter is reviewed, so each move is tested
|
||||
alongside that filter. A 2026-09-13 survey found 19 functions with identical
|
||||
copies in several scripts and 19 with copies that have drifted apart. Most
|
||||
identical copies are NSRDB helpers, which belong in an NSRDB module rather
|
||||
than `common/`; `read_csv_rows` (4 identical copies in `apply_*` scripts) is
|
||||
cross-source. Drifted copies need a decision on which version is correct
|
||||
before merging. Notable drifts: `summarize_county_gridmet_humidity.py` has its
|
||||
own county loader and FIPS normalizer, and the state FIPS table is also copied
|
||||
in `build_county_representative_points.py`,
|
||||
`summarize_county_gridmet_humidity.py`, and
|
||||
`request_nsrdb_county_polygon_archives.py`.
|
||||
|
||||
### Phase 3 — Assemble and restructure (after all 12 filters are clean)
|
||||
|
||||
- Assemble script that builds `climate-data.csv` from `data/metrics/`.
|
||||
- Populate `metric_sources.json`; remove the per-row `source` column.
|
||||
- Move audit columns out of the app CSV.
|
||||
- Point the app's Sources panel at `metric_sources.json`.
|
||||
- Retire or rewrite `build_county_climate_data.py` as per-metric builders.
|
||||
- Move `common/koppen_legend.py` next to the Köppen code once nothing outside
|
||||
Köppen imports it.
|
||||
|
||||
### Phase 4 — Runner and reproduction guide
|
||||
|
||||
- `scripts/pipeline.py` runner with `--only` and `--skip-download`.
|
||||
- "Reproducing the data" guide organized by the tiers in section 3.
|
||||
- Pinned requirements.
|
||||
- End-to-end smoke test on a small synthetic county fixture.
|
||||
- Organize scripts by data source (`noaa/`, `gridmet/`, `nsrdb/`, `koppen/`),
|
||||
each holding its own helpers, with `common/` keeping only cross-source code.
|
||||
Scripts in subfolders are run through the runner or as modules
|
||||
(`python -m`), and the README and data-source commands are updated to match.
|
||||
|
||||
## 5. Per-filter review checklist
|
||||
|
||||
For each filter:
|
||||
|
||||
1. Verify the calculation against the source data and document findings.
|
||||
2. Decide any rule or method changes with the project owner.
|
||||
3. Record the adopted definition in `filter-calculations.md`.
|
||||
4. Implement the calculation, writing `data/metrics/<metric>.csv`.
|
||||
5. Add a single-column apply step for the current CSV.
|
||||
6. Update the rules in `check_climate_data.py`.
|
||||
7. Add or update unit tests.
|
||||
8. Apply to the CSV, run `check_climate_data.py`, and compare changed counties
|
||||
against expectations.
|
||||
9. Update the app if the value set or display changes.
|
||||
|
||||
## 6. Filter tracker
|
||||
|
||||
Known issues come from `filter-calculations.md` ("Calculation review findings")
|
||||
and this review; none beyond Köppen have been investigated yet.
|
||||
|
||||
| # | Filter | Status | Known issues to review |
|
||||
| --- | --- | --- | --- |
|
||||
| 1 | Köppen-Geiger class | Done (2026-09-14): rule applied, Mixed display built, documentation updated | See tasks below |
|
||||
| 2 | Annual avg temperature | Not started | Months weighted equally (finding 1); touched-cell aggregation (finding 2) |
|
||||
| 3 | Diurnal temperature range | Not started | Lexington, VA (51678) blank, while heat-index days use Rockbridge County as a proxy |
|
||||
| 4 | Extreme temperature days | Not started | Partial years not normalized (finding 5); depends on retired percentile thresholds (finding 6); Lexington, VA blank |
|
||||
| 5 | 90 °F+ heat-index days | Not started | Daily-extrema proxy (finding 7); permissive year completeness (finding 8) |
|
||||
| 6 | Annual precipitation | Not started | Partial-year sums accepted (finding 3); touched-cell aggregation (finding 2) |
|
||||
| 7 | Seasonality index | Not started | Touched-cell aggregation (finding 2) |
|
||||
| 8 | Wettest month | Not started | Computed in both the base build and the precipitation-month script |
|
||||
| 9 | Driest month | Not started | Same as wettest month |
|
||||
| 10 | Summer specific humidity | Not started | Cell-center cos(latitude) aggregation differs from other metrics (finding 11) |
|
||||
| 11 | Solar GHI | Not started | Finalized by the extreme-temperature apply script; hourly, 365-day assumption (finding 9) |
|
||||
| 12 | Clear-sky GHI reduction | Not started | Mean of ratios rather than energy totals (finding 10) |
|
||||
|
||||
### Köppen-Geiger tasks
|
||||
|
||||
Adopted rule: a county is predominantly its top class if and only if that class
|
||||
covers at least 50% of the county's land and leads the runner-up by at least
|
||||
5 percentage points; otherwise it is Mixed climate. Expected result for the
|
||||
50 states and DC: 3,010 predominant, 133 Mixed.
|
||||
|
||||
- [x] Investigate low-majority counties and adopt the rule.
|
||||
- [x] Windowed raster reads and 180th-meridian split.
|
||||
- [x] Area-weighted class shares (16 × 16 sub-cells per raster cell, scaled by
|
||||
cos(latitude)) in `scripts/common/county_zonal_stats.py`.
|
||||
- [x] Apply the 50% / 5-point rule (`scripts/build_county_koppen_metric.py`;
|
||||
counties with no valid cells are left blank).
|
||||
- [x] Run the builder to write `data/metrics/koppen.csv` and confirm the
|
||||
expected 3,010 predominant / 133 Mixed (2026-09-13).
|
||||
- [x] `koppenZone`-only apply step
|
||||
(`scripts/apply_koppen_metric_to_climate_data.py`, with `--dry-run`).
|
||||
- [x] Apply to `data/climate-data.csv` (2026-09-13; 142 counties changed to
|
||||
Mixed, with `koppenPrimaryClass` and `koppenSecondaryClass` added).
|
||||
- [x] Allow `Mixed` in `check_climate_data.py`.
|
||||
- [x] Add a Mixed climate category to `app.js`, drawn as stripes of the
|
||||
county's top two classes; see
|
||||
[koppen-mixed-display-plan.md](koppen-mixed-display-plan.md).
|
||||
- [x] Replace the plurality description in `filter-calculations.md` §1 and mark
|
||||
review findings 2 and 4 resolved for Köppen (2026-09-14).
|
||||
- [x] Update the Köppen descriptions and script lists in `README.md` and
|
||||
`scripts/county_data_sources.md` (2026-09-14).
|
||||
- [x] Tests for shares, the rule, boundary cases, and the apply step
|
||||
(`tests/test_koppen_metric.py`).
|
||||
|
||||
## 7. Guardrails until Phase 3
|
||||
|
||||
- **Do not rerun `build_county_climate_data.py`** against
|
||||
`data/climate-data.csv`. It would drop 9 live columns, restore
|
||||
`extremeDays`, and overwrite the Mixed classification in `koppenZone`.
|
||||
- Run `check_climate_data.py` after every apply step.
|
||||
- Change one filter at a time, and compare its before and after values.
|
||||
|
||||
## 8. Open decisions
|
||||
|
||||
| Decision | Options | Needed by |
|
||||
| --- | --- | --- |
|
||||
| Committing large intermediates | Commit metric files only, or also source summaries | Phase 3 |
|
||||
|
||||
### Decided
|
||||
|
||||
- **Köppen audit columns (2026-09-12):** `koppen.csv` stores the top class and
|
||||
share and the runner-up class and share alongside `koppenZone`.
|
||||
- **Köppen no-data fallback (2026-09-12):** a county with no valid raster cells
|
||||
is left blank, not assigned `Cfa`.
|
||||
- **Applying Köppen to the app CSV (2026-09-12):** wait until the app supports
|
||||
the Mixed class. Done 2026-09-13.
|
||||
- **Metric files (2026-09-13):** one CSV per metric under `data/metrics/`,
|
||||
starting with `koppen.csv`.
|
||||
- **Mixed climate display (2026-09-13):** diagonal stripes of each Mixed
|
||||
county's top two classes; see
|
||||
[koppen-mixed-display-plan.md](koppen-mixed-display-plan.md).
|
||||
- **Puerto Rico (2026-09-13):** off the map and out of every filter. The app
|
||||
already drops state FIPS 72; the data files keep the rows.
|
||||
- **Shared helpers (2026-09-13):** `scripts/common/` holds only code used by
|
||||
more than one data source, one topic per module; code shared within one data
|
||||
source stays with that source. Duplicates move during their own filter's
|
||||
review.
|
||||
Reference in New Issue
Block a user