Files
Climate-Mood-Analysis/README.md
T
KnouandClaude Opus 5 4d2b3e3d44 Complete Köppen-Geiger filter review with Mixed climate class
Classify each county by area-weighted Köppen class shares: a county is
predominantly its top class when that class covers at least 50% of its
land and leads the runner-up by at least 5 percentage points; otherwise
it is Mixed (133 of 3,143 counties in the 50 states and DC).

- Add build_county_koppen_metric.py (writes data/metrics/koppen.csv) and
  apply_koppen_metric_to_climate_data.py (writes koppenZone plus
  koppenPrimaryClass/koppenSecondaryClass for Mixed counties).
- Move shared helpers into scripts/common/ (county loading, Köppen
  legend, area-weighted raster shares); fix the 180th-meridian raster
  window for Aleutians West.
- Add check_climate_data.py to validate the app CSV.
- Draw Mixed counties in app.js as diagonal stripes of their top two
  classes, fixed to the ground and following the map at every zoom, with
  a crossfade only when the stripe size changes. Filtering a class also
  matches Mixed counties where it is primary or secondary.
- Document the rule, display, and pipeline plan in docs/ and update the
  README and data-source notes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 02:54:25 -04:00

230 lines
12 KiB
Markdown

# Temperature-Based Analysis
Temperature-Based Analysis is an exploratory project for studying how local
climate conditions vary across the United States and, eventually, whether those
conditions show any relationship with mood-based metrics.
The project currently provides the **US County Climate Explorer**, an
interactive county-level map for viewing and filtering climate data. The mood
dataset and climate-to-mood correlation analysis are **not implemented yet**.
At this stage, the project is focused on building, validating, and presenting
the climate side of the analysis.
## Current Features
- Interactive Leaflet map with county boundaries and county selection
- Metric groups for climate classification, temperature, precipitation,
moisture, and solar resource
- Numeric range filters and categorical filters
- County detail panel showing all available climate metrics
- Legends and in-app source descriptions
- Local CSV-based data loading with no application backend or build step
- Offline Python scripts for assembling and updating county climate data
The browser currently loads 3,221 county-level records from
`data/climate-data.csv`.
## Available Climate Metrics
| Group | Metrics |
| --- | --- |
| Koppen-Geiger Classification | Predominant county climate class (covering at least 50% of the county's land and leading the runner-up by at least 5 points), or Mixed, drawn as stripes of the county's top two classes |
| Temperature & Extremes | Annual average temperature, diurnal temperature range, annual extreme temperature days, and annual 90 F+ heat-index days |
| Precipitation & Moisture | Annual precipitation, precipitation seasonality, wettest month, driest month, and summer specific humidity |
| Solar Resource | Mean daily global horizontal radiation (GHI) and clear-sky GHI reduction index |
Most long-term climate metrics use a 1991-2020 reference period. In the
checked-in CSV, the annual extreme-temperature-days metric uses 1991-2025 data
and is the average annual count of days with `Tmax >= 95 F` or `Tmin <= 0 F`.
The build script defaults its analysis end year to the latest likely complete
calendar year. Definitions and time periods are shown in the application's
**Sources** dialog and documented in more detail in
[`scripts/county_data_sources.md`](scripts/county_data_sources.md).
## Run the Explorer
### Requirements
- A modern web browser
- Python 3
- PowerShell for the included convenience script
- An internet connection for Leaflet, map tiles, and hosted fonts
No Node.js installation or frontend build is required.
From the project root, run:
```powershell
.\serve.ps1
```
Then open:
```text
http://localhost:8000/
```
To use another port:
```powershell
.\serve.ps1 -Port 8080
```
The site must be served over HTTP because the browser loads the climate CSV and
county GeoJSON with `fetch()`. Opening `index.html` directly as a local file
will not work.
## Using the Map
1. Choose a metric group and metric from the control panel.
2. Adjust the value range or category filter to highlight matching counties.
3. Click a county to zoom to it and inspect all available metrics.
4. Use **Reset Country View** or **Clear Selection** to return to the broader
map.
5. Open **Sources** for metric definitions and data provenance.
## Project Structure
| Path | Purpose |
| --- | --- |
| `index.html` | Application structure and controls |
| `styles.css` | Layout and visual styling |
| `app.js` | Map rendering, filtering, county details, and source metadata |
| `serve.ps1` | Local HTTP server launcher |
| `data/climate-data.csv` | Browser-ready county climate records |
| `data/metrics/` | Per-metric county outputs, starting with `koppen.csv` |
| `data/geojson-counties-fips.json` | County geometry keyed by FIPS code |
| `scripts/` | Climate-data download, aggregation, and update tools |
| `scripts/county_data_sources.md` | Detailed metric definitions, data provenance, and pipeline examples |
| `scripts/requirements_county_etl.txt` | Python dependencies for the offline data pipeline |
| `docs/` | Filter calculations, the pipeline plan, and design notes |
| `tests/` | Tests for the Köppen metric, the climate CSV check, NSRDB request/download wrappers, cloud-metric merging, and gridMET heat-index calculations |
## Climate Data Pipeline
The checked-in browser assets are the final outputs needed to run the explorer.
The scripts directory contains the larger offline workflow used to derive those
outputs from sources including:
- Beck et al. Koppen-Geiger climate classification data
- NOAA NCEI Climate Normals and nClimGrid data
- gridMET humidity data
- NREL National Solar Radiation Database data
- Plotly's county GeoJSON, keyed by U.S. Census county FIPS codes
The Python tooling is configured for Python 3.11. To work on the data pipeline,
create a virtual environment and install the ETL dependencies:
```powershell
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -r scripts\requirements_county_etl.txt
```
Some data-generation workflows download large files or require NREL/NSRDB API
credentials. Generated source datasets are intentionally excluded from Git;
only the browser-ready CSV and GeoJSON are tracked.
Run script commands from the project root so their default `data/...` paths
resolve correctly. The pipeline is split into a base generator, source-specific
builders, and small scripts that merge the resulting metrics into
`data/climate-data.csv`. Most `apply_*.py` commands update that CSV in place by
default.
### Script Inventory
| Script | Current role |
| --- | --- |
| `build_county_climate_data.py` | Builds the base app CSV from county geometry, Koppen-Geiger data, NOAA temperature/precipitation data, and optional solar inputs. Its older `extremeDays` output is replaced by the current absolute-threshold stage below, and its largest-share `koppenZone` by the Köppen stage. It writes only the base columns, so do not run it over the live CSV. |
| `build_county_koppen_metric.py` / `apply_koppen_metric_to_climate_data.py` | Builds area-weighted Köppen class shares per county in `data/metrics/koppen.csv`, classifies each county as predominant or Mixed, and writes `koppenZone` plus the two stripe-class columns for Mixed counties. |
| `apply_precipitation_month_metrics_to_climate_data.py` | Recomputes and merges the 1991-2020 wettest- and driest-month categories from monthly nClimGrid precipitation. |
| `build_county_locally_extreme_data.py` | Downloads or reads cached nClimGrid-Daily county Tmax/Tmin files, calculates county-percentile diagnostics, and calculates the app-facing absolute 95 F / 0 F day counts. |
| `apply_locally_extreme_metric_to_climate_data.py` | Writes `absoluteExtremeDays`, removes retired locally extreme/legacy fields, and selects polygon GHI with representative-point GHI as fallback. The filename is retained from the earlier pipeline. |
| `build_county_diurnal_temperature_range.py` / `apply_diurnal_temperature_range_to_climate_data.py` | Builds the 1991-2020 county mean daily Tmax-minus-Tmin artifact and merges `avgDiurnalTempRangeF`. |
| `download_gridmet_data.py` | Downloads 1991-2020 `sph`, `rmax`, and `rmin` NetCDF files by default. |
| `summarize_county_gridmet_humidity.py` / `apply_gridmet_humidity_metric_to_climate_data.py` | Produces county summer specific humidity and a 90 F+ Heat Index day proxy, then merges those metrics and FIPS audit fields. |
| `build_county_representative_points.py` | Creates interior county points used by the lightweight NSRDB workflows. |
| `fetch_nsrdb_representative_point_ghi.py` / `rebuild_nsrdb_representative_point_ghi_summary.py` | Fetches point-based NSRDB GHI or rebuilds its summary from cached responses without another API call. |
| `fetch_nsrdb_representative_point_cloud_metrics.py` | Fetches point-based GHI, clear-sky GHI, and cloud type, then summarizes the clear-sky GHI reduction index. |
| `request_nsrdb_county_polygon_archives.py` | Shared, resumable NSRDB polygon-request engine with tiling, site-count checks, pacing, and Polar fallback. |
| `request_nsrdb_county_polygon_ghi_archives.py` / `request_nsrdb_county_polygon_cloud_archives.py` | Recommended wrappers around the shared request engine, with separate GHI and cloud attributes and output paths. |
| `download_nsrdb_county_polygon_archives.py` | Shared state-machine downloader for completed NSRDB archive jobs. |
| `download_nsrdb_county_polygon_ghi_archives.py` / `download_nsrdb_county_polygon_cloud_archives.py` | Recommended wrappers around the shared downloader, keeping GHI and cloud archives separate. |
| `summarize_nsrdb_county_polygon_archives.py` | Combines county/tile GHI archives into area-weighted county summaries. |
| `summarize_nsrdb_county_polygon_cloud_archives.py` | Combines county/tile cloud archives into area-weighted clear-sky GHI reduction summaries. |
| `apply_nsrdb_cloud_metric_to_climate_data.py` | Merges the clear-sky GHI reduction index, preferring polygon summaries and falling back to representative points. |
| `check_climate_data.py` | Validates `data/climate-data.csv`, including Köppen codes and the stripe-class columns. |
| `common/` | Shared helpers: county loading, the Köppen legend, and area-weighted raster shares. |
The metric-specific NSRDB request and download wrappers are the normal entry
points. The shared engines remain available for custom attributes or artifact
paths. Representative-point results provide a faster first pass; polygon
summaries are the preferred county-area result when available.
Common local enrichment stages, after their source files have been downloaded,
are:
```powershell
.venv\Scripts\python.exe scripts\build_county_koppen_metric.py
.venv\Scripts\python.exe scripts\apply_koppen_metric_to_climate_data.py
.venv\Scripts\python.exe scripts\build_county_locally_extreme_data.py --skip-download
.venv\Scripts\python.exe scripts\apply_locally_extreme_metric_to_climate_data.py
.venv\Scripts\python.exe scripts\build_county_diurnal_temperature_range.py
.venv\Scripts\python.exe scripts\apply_diurnal_temperature_range_to_climate_data.py
.venv\Scripts\python.exe scripts\summarize_county_gridmet_humidity.py --years 1991-2020
.venv\Scripts\python.exe scripts\apply_gridmet_humidity_metric_to_climate_data.py
.venv\Scripts\python.exe scripts\apply_nsrdb_cloud_metric_to_climate_data.py
```
The order matters when rebuilding from scratch: run the Köppen apply step after
the base build, generate the locally extreme comparison and solar summaries
before running their apply step, and summarize gridMET or NSRDB downloads
before merging them. See the data-source document linked above for acquisition
commands, expected artifacts, FIPS handling, and the representative-point and
polygon NSRDB workflows. The full rebuild order is in
[`docs/pipeline-plan.md`](docs/pipeline-plan.md).
After updating the CSV, validate it with:
```powershell
.venv\Scripts\python.exe scripts\check_climate_data.py
```
Run the current automated tests with:
```powershell
python -m unittest discover -s tests
```
## Planned Mood Analysis
The longer-term goal is to add mood-based metrics and investigate whether
patterns in those metrics are associated with climate characteristics such as
temperature, sunlight availability, clear-sky GHI reduction, humidity,
precipitation, or extreme-weather frequency.
That phase still requires decisions about:
- How mood data will be collected or sourced
- Geographic and temporal granularity
- Privacy, consent, and aggregation requirements
- Confounding variables and missing-data handling
- Appropriate statistical and visualization methods
No mood records, mood visualizations, or correlation results are currently part
of the application. Any future relationship found by the project should be
treated as an association to investigate, not evidence that climate alone
causes changes in mood.
## Current Limitations
- The application presents long-term county summaries rather than live weather.
- Climate values aggregate conditions within county boundaries and do not
represent every location inside a county.
- Source datasets use different methods and, in some cases, different time
periods.
- Some solar metrics use representative-point values where county polygon
summaries are unavailable.
- The current interface is an exploratory visualization, not a completed
climate-and-mood research analysis.