Classify each county by area-weighted Köppen class shares: a county is predominantly its top class when that class covers at least 50% of its land and leads the runner-up by at least 5 percentage points; otherwise it is Mixed (133 of 3,143 counties in the 50 states and DC). - Add build_county_koppen_metric.py (writes data/metrics/koppen.csv) and apply_koppen_metric_to_climate_data.py (writes koppenZone plus koppenPrimaryClass/koppenSecondaryClass for Mixed counties). - Move shared helpers into scripts/common/ (county loading, Köppen legend, area-weighted raster shares); fix the 180th-meridian raster window for Aleutians West. - Add check_climate_data.py to validate the app CSV. - Draw Mixed counties in app.js as diagonal stripes of their top two classes, fixed to the ground and following the map at every zoom, with a crossfade only when the stripe size changes. Filtering a class also matches Mixed counties where it is primary or secondary. - Document the rule, display, and pipeline plan in docs/ and update the README and data-source notes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
230 lines
12 KiB
Markdown
230 lines
12 KiB
Markdown
# Temperature-Based Analysis
|
|
|
|
Temperature-Based Analysis is an exploratory project for studying how local
|
|
climate conditions vary across the United States and, eventually, whether those
|
|
conditions show any relationship with mood-based metrics.
|
|
|
|
The project currently provides the **US County Climate Explorer**, an
|
|
interactive county-level map for viewing and filtering climate data. The mood
|
|
dataset and climate-to-mood correlation analysis are **not implemented yet**.
|
|
At this stage, the project is focused on building, validating, and presenting
|
|
the climate side of the analysis.
|
|
|
|
## Current Features
|
|
|
|
- Interactive Leaflet map with county boundaries and county selection
|
|
- Metric groups for climate classification, temperature, precipitation,
|
|
moisture, and solar resource
|
|
- Numeric range filters and categorical filters
|
|
- County detail panel showing all available climate metrics
|
|
- Legends and in-app source descriptions
|
|
- Local CSV-based data loading with no application backend or build step
|
|
- Offline Python scripts for assembling and updating county climate data
|
|
|
|
The browser currently loads 3,221 county-level records from
|
|
`data/climate-data.csv`.
|
|
|
|
## Available Climate Metrics
|
|
|
|
| Group | Metrics |
|
|
| --- | --- |
|
|
| Koppen-Geiger Classification | Predominant county climate class (covering at least 50% of the county's land and leading the runner-up by at least 5 points), or Mixed, drawn as stripes of the county's top two classes |
|
|
| Temperature & Extremes | Annual average temperature, diurnal temperature range, annual extreme temperature days, and annual 90 F+ heat-index days |
|
|
| Precipitation & Moisture | Annual precipitation, precipitation seasonality, wettest month, driest month, and summer specific humidity |
|
|
| Solar Resource | Mean daily global horizontal radiation (GHI) and clear-sky GHI reduction index |
|
|
|
|
Most long-term climate metrics use a 1991-2020 reference period. In the
|
|
checked-in CSV, the annual extreme-temperature-days metric uses 1991-2025 data
|
|
and is the average annual count of days with `Tmax >= 95 F` or `Tmin <= 0 F`.
|
|
The build script defaults its analysis end year to the latest likely complete
|
|
calendar year. Definitions and time periods are shown in the application's
|
|
**Sources** dialog and documented in more detail in
|
|
[`scripts/county_data_sources.md`](scripts/county_data_sources.md).
|
|
|
|
## Run the Explorer
|
|
|
|
### Requirements
|
|
|
|
- A modern web browser
|
|
- Python 3
|
|
- PowerShell for the included convenience script
|
|
- An internet connection for Leaflet, map tiles, and hosted fonts
|
|
|
|
No Node.js installation or frontend build is required.
|
|
|
|
From the project root, run:
|
|
|
|
```powershell
|
|
.\serve.ps1
|
|
```
|
|
|
|
Then open:
|
|
|
|
```text
|
|
http://localhost:8000/
|
|
```
|
|
|
|
To use another port:
|
|
|
|
```powershell
|
|
.\serve.ps1 -Port 8080
|
|
```
|
|
|
|
The site must be served over HTTP because the browser loads the climate CSV and
|
|
county GeoJSON with `fetch()`. Opening `index.html` directly as a local file
|
|
will not work.
|
|
|
|
## Using the Map
|
|
|
|
1. Choose a metric group and metric from the control panel.
|
|
2. Adjust the value range or category filter to highlight matching counties.
|
|
3. Click a county to zoom to it and inspect all available metrics.
|
|
4. Use **Reset Country View** or **Clear Selection** to return to the broader
|
|
map.
|
|
5. Open **Sources** for metric definitions and data provenance.
|
|
|
|
## Project Structure
|
|
|
|
| Path | Purpose |
|
|
| --- | --- |
|
|
| `index.html` | Application structure and controls |
|
|
| `styles.css` | Layout and visual styling |
|
|
| `app.js` | Map rendering, filtering, county details, and source metadata |
|
|
| `serve.ps1` | Local HTTP server launcher |
|
|
| `data/climate-data.csv` | Browser-ready county climate records |
|
|
| `data/metrics/` | Per-metric county outputs, starting with `koppen.csv` |
|
|
| `data/geojson-counties-fips.json` | County geometry keyed by FIPS code |
|
|
| `scripts/` | Climate-data download, aggregation, and update tools |
|
|
| `scripts/county_data_sources.md` | Detailed metric definitions, data provenance, and pipeline examples |
|
|
| `scripts/requirements_county_etl.txt` | Python dependencies for the offline data pipeline |
|
|
| `docs/` | Filter calculations, the pipeline plan, and design notes |
|
|
| `tests/` | Tests for the Köppen metric, the climate CSV check, NSRDB request/download wrappers, cloud-metric merging, and gridMET heat-index calculations |
|
|
|
|
## Climate Data Pipeline
|
|
|
|
The checked-in browser assets are the final outputs needed to run the explorer.
|
|
The scripts directory contains the larger offline workflow used to derive those
|
|
outputs from sources including:
|
|
|
|
- Beck et al. Koppen-Geiger climate classification data
|
|
- NOAA NCEI Climate Normals and nClimGrid data
|
|
- gridMET humidity data
|
|
- NREL National Solar Radiation Database data
|
|
- Plotly's county GeoJSON, keyed by U.S. Census county FIPS codes
|
|
|
|
The Python tooling is configured for Python 3.11. To work on the data pipeline,
|
|
create a virtual environment and install the ETL dependencies:
|
|
|
|
```powershell
|
|
python -m venv .venv
|
|
.\.venv\Scripts\Activate.ps1
|
|
python -m pip install -r scripts\requirements_county_etl.txt
|
|
```
|
|
|
|
Some data-generation workflows download large files or require NREL/NSRDB API
|
|
credentials. Generated source datasets are intentionally excluded from Git;
|
|
only the browser-ready CSV and GeoJSON are tracked.
|
|
|
|
Run script commands from the project root so their default `data/...` paths
|
|
resolve correctly. The pipeline is split into a base generator, source-specific
|
|
builders, and small scripts that merge the resulting metrics into
|
|
`data/climate-data.csv`. Most `apply_*.py` commands update that CSV in place by
|
|
default.
|
|
|
|
### Script Inventory
|
|
|
|
| Script | Current role |
|
|
| --- | --- |
|
|
| `build_county_climate_data.py` | Builds the base app CSV from county geometry, Koppen-Geiger data, NOAA temperature/precipitation data, and optional solar inputs. Its older `extremeDays` output is replaced by the current absolute-threshold stage below, and its largest-share `koppenZone` by the Köppen stage. It writes only the base columns, so do not run it over the live CSV. |
|
|
| `build_county_koppen_metric.py` / `apply_koppen_metric_to_climate_data.py` | Builds area-weighted Köppen class shares per county in `data/metrics/koppen.csv`, classifies each county as predominant or Mixed, and writes `koppenZone` plus the two stripe-class columns for Mixed counties. |
|
|
| `apply_precipitation_month_metrics_to_climate_data.py` | Recomputes and merges the 1991-2020 wettest- and driest-month categories from monthly nClimGrid precipitation. |
|
|
| `build_county_locally_extreme_data.py` | Downloads or reads cached nClimGrid-Daily county Tmax/Tmin files, calculates county-percentile diagnostics, and calculates the app-facing absolute 95 F / 0 F day counts. |
|
|
| `apply_locally_extreme_metric_to_climate_data.py` | Writes `absoluteExtremeDays`, removes retired locally extreme/legacy fields, and selects polygon GHI with representative-point GHI as fallback. The filename is retained from the earlier pipeline. |
|
|
| `build_county_diurnal_temperature_range.py` / `apply_diurnal_temperature_range_to_climate_data.py` | Builds the 1991-2020 county mean daily Tmax-minus-Tmin artifact and merges `avgDiurnalTempRangeF`. |
|
|
| `download_gridmet_data.py` | Downloads 1991-2020 `sph`, `rmax`, and `rmin` NetCDF files by default. |
|
|
| `summarize_county_gridmet_humidity.py` / `apply_gridmet_humidity_metric_to_climate_data.py` | Produces county summer specific humidity and a 90 F+ Heat Index day proxy, then merges those metrics and FIPS audit fields. |
|
|
| `build_county_representative_points.py` | Creates interior county points used by the lightweight NSRDB workflows. |
|
|
| `fetch_nsrdb_representative_point_ghi.py` / `rebuild_nsrdb_representative_point_ghi_summary.py` | Fetches point-based NSRDB GHI or rebuilds its summary from cached responses without another API call. |
|
|
| `fetch_nsrdb_representative_point_cloud_metrics.py` | Fetches point-based GHI, clear-sky GHI, and cloud type, then summarizes the clear-sky GHI reduction index. |
|
|
| `request_nsrdb_county_polygon_archives.py` | Shared, resumable NSRDB polygon-request engine with tiling, site-count checks, pacing, and Polar fallback. |
|
|
| `request_nsrdb_county_polygon_ghi_archives.py` / `request_nsrdb_county_polygon_cloud_archives.py` | Recommended wrappers around the shared request engine, with separate GHI and cloud attributes and output paths. |
|
|
| `download_nsrdb_county_polygon_archives.py` | Shared state-machine downloader for completed NSRDB archive jobs. |
|
|
| `download_nsrdb_county_polygon_ghi_archives.py` / `download_nsrdb_county_polygon_cloud_archives.py` | Recommended wrappers around the shared downloader, keeping GHI and cloud archives separate. |
|
|
| `summarize_nsrdb_county_polygon_archives.py` | Combines county/tile GHI archives into area-weighted county summaries. |
|
|
| `summarize_nsrdb_county_polygon_cloud_archives.py` | Combines county/tile cloud archives into area-weighted clear-sky GHI reduction summaries. |
|
|
| `apply_nsrdb_cloud_metric_to_climate_data.py` | Merges the clear-sky GHI reduction index, preferring polygon summaries and falling back to representative points. |
|
|
| `check_climate_data.py` | Validates `data/climate-data.csv`, including Köppen codes and the stripe-class columns. |
|
|
| `common/` | Shared helpers: county loading, the Köppen legend, and area-weighted raster shares. |
|
|
|
|
The metric-specific NSRDB request and download wrappers are the normal entry
|
|
points. The shared engines remain available for custom attributes or artifact
|
|
paths. Representative-point results provide a faster first pass; polygon
|
|
summaries are the preferred county-area result when available.
|
|
|
|
Common local enrichment stages, after their source files have been downloaded,
|
|
are:
|
|
|
|
```powershell
|
|
.venv\Scripts\python.exe scripts\build_county_koppen_metric.py
|
|
.venv\Scripts\python.exe scripts\apply_koppen_metric_to_climate_data.py
|
|
.venv\Scripts\python.exe scripts\build_county_locally_extreme_data.py --skip-download
|
|
.venv\Scripts\python.exe scripts\apply_locally_extreme_metric_to_climate_data.py
|
|
.venv\Scripts\python.exe scripts\build_county_diurnal_temperature_range.py
|
|
.venv\Scripts\python.exe scripts\apply_diurnal_temperature_range_to_climate_data.py
|
|
.venv\Scripts\python.exe scripts\summarize_county_gridmet_humidity.py --years 1991-2020
|
|
.venv\Scripts\python.exe scripts\apply_gridmet_humidity_metric_to_climate_data.py
|
|
.venv\Scripts\python.exe scripts\apply_nsrdb_cloud_metric_to_climate_data.py
|
|
```
|
|
|
|
The order matters when rebuilding from scratch: run the Köppen apply step after
|
|
the base build, generate the locally extreme comparison and solar summaries
|
|
before running their apply step, and summarize gridMET or NSRDB downloads
|
|
before merging them. See the data-source document linked above for acquisition
|
|
commands, expected artifacts, FIPS handling, and the representative-point and
|
|
polygon NSRDB workflows. The full rebuild order is in
|
|
[`docs/pipeline-plan.md`](docs/pipeline-plan.md).
|
|
|
|
After updating the CSV, validate it with:
|
|
|
|
```powershell
|
|
.venv\Scripts\python.exe scripts\check_climate_data.py
|
|
```
|
|
|
|
Run the current automated tests with:
|
|
|
|
```powershell
|
|
python -m unittest discover -s tests
|
|
```
|
|
|
|
## Planned Mood Analysis
|
|
|
|
The longer-term goal is to add mood-based metrics and investigate whether
|
|
patterns in those metrics are associated with climate characteristics such as
|
|
temperature, sunlight availability, clear-sky GHI reduction, humidity,
|
|
precipitation, or extreme-weather frequency.
|
|
|
|
That phase still requires decisions about:
|
|
|
|
- How mood data will be collected or sourced
|
|
- Geographic and temporal granularity
|
|
- Privacy, consent, and aggregation requirements
|
|
- Confounding variables and missing-data handling
|
|
- Appropriate statistical and visualization methods
|
|
|
|
No mood records, mood visualizations, or correlation results are currently part
|
|
of the application. Any future relationship found by the project should be
|
|
treated as an association to investigate, not evidence that climate alone
|
|
causes changes in mood.
|
|
|
|
## Current Limitations
|
|
|
|
- The application presents long-term county summaries rather than live weather.
|
|
- Climate values aggregate conditions within county boundaries and do not
|
|
represent every location inside a county.
|
|
- Source datasets use different methods and, in some cases, different time
|
|
periods.
|
|
- Some solar metrics use representative-point values where county polygon
|
|
summaries are unavailable.
|
|
- The current interface is an exploratory visualization, not a completed
|
|
climate-and-mood research analysis.
|