README.md has been updated to reflect project changes.

This commit is contained in:
2026-09-11 16:30:15 -04:00
parent 4e2d878e5b
commit 92fbbfb2e9
+67 -11
View File
@@ -33,11 +33,13 @@ The browser currently loads 3,221 county-level records from
| Precipitation & Moisture | Annual precipitation, precipitation seasonality, wettest month, driest month, and summer specific humidity | | Precipitation & Moisture | Annual precipitation, precipitation seasonality, wettest month, driest month, and summer specific humidity |
| Solar Resource | Mean daily global horizontal radiation (GHI) and clear-sky GHI reduction index | | Solar Resource | Mean daily global horizontal radiation (GHI) and clear-sky GHI reduction index |
Most long-term climate metrics use a 1991-2020 reference period. The annual Most long-term climate metrics use a 1991-2020 reference period. In the
extreme-temperature-days metric currently uses 1991-2025 data and includes days checked-in CSV, the annual extreme-temperature-days metric uses 1991-2025 data
meeting either the extreme heat or extreme cold threshold. Definitions and time and is the average annual count of days with `Tmax >= 95 F` or `Tmin <= 0 F`.
periods are shown in the application's **Sources** dialog and documented in more The build script defaults its analysis end year to the latest likely complete
detail in [`scripts/county_data_sources.md`](scripts/county_data_sources.md). calendar year. Definitions and time periods are shown in the application's
**Sources** dialog and documented in more detail in
[`scripts/county_data_sources.md`](scripts/county_data_sources.md).
## Run the Explorer ## Run the Explorer
@@ -92,7 +94,9 @@ will not work.
| `data/climate-data.csv` | Browser-ready county climate records | | `data/climate-data.csv` | Browser-ready county climate records |
| `data/geojson-counties-fips.json` | County geometry keyed by FIPS code | | `data/geojson-counties-fips.json` | County geometry keyed by FIPS code |
| `scripts/` | Climate-data download, aggregation, and update tools | | `scripts/` | Climate-data download, aggregation, and update tools |
| `tests/` | Tests for the NSRDB polygon request and download workflow | | `scripts/county_data_sources.md` | Detailed metric definitions, data provenance, and pipeline examples |
| `scripts/requirements_county_etl.txt` | Python dependencies for the offline data pipeline |
| `tests/` | Tests for NSRDB request/download wrappers, cloud-metric merging, and gridMET heat-index calculations |
## Climate Data Pipeline ## Climate Data Pipeline
@@ -104,10 +108,10 @@ outputs from sources including:
- NOAA NCEI Climate Normals and nClimGrid data - NOAA NCEI Climate Normals and nClimGrid data
- gridMET humidity data - gridMET humidity data
- NREL National Solar Radiation Database data - NREL National Solar Radiation Database data
- U.S. Census Bureau county geometry - Plotly's county GeoJSON, keyed by U.S. Census county FIPS codes
To work on the Python data pipeline, create a virtual environment and install The Python tooling is configured for Python 3.11. To work on the data pipeline,
the ETL dependencies: create a virtual environment and install the ETL dependencies:
```powershell ```powershell
python -m venv .venv python -m venv .venv
@@ -119,6 +123,58 @@ Some data-generation workflows download large files or require NREL/NSRDB API
credentials. Generated source datasets are intentionally excluded from Git; credentials. Generated source datasets are intentionally excluded from Git;
only the browser-ready CSV and GeoJSON are tracked. only the browser-ready CSV and GeoJSON are tracked.
Run script commands from the project root so their default `data/...` paths
resolve correctly. The pipeline is split into a base generator, source-specific
builders, and small scripts that merge the resulting metrics into
`data/climate-data.csv`. Most `apply_*.py` commands update that CSV in place by
default.
### Script Inventory
| Script | Current role |
| --- | --- |
| `build_county_climate_data.py` | Builds the base app CSV from county geometry, Koppen-Geiger data, NOAA temperature/precipitation data, and optional solar inputs. Its older `extremeDays` output is replaced by the current absolute-threshold stage below. |
| `apply_precipitation_month_metrics_to_climate_data.py` | Recomputes and merges the 1991-2020 wettest- and driest-month categories from monthly nClimGrid precipitation. |
| `build_county_locally_extreme_data.py` | Downloads or reads cached nClimGrid-Daily county Tmax/Tmin files, calculates county-percentile diagnostics, and calculates the app-facing absolute 95 F / 0 F day counts. |
| `apply_locally_extreme_metric_to_climate_data.py` | Writes `absoluteExtremeDays`, removes retired locally extreme/legacy fields, and selects polygon GHI with representative-point GHI as fallback. The filename is retained from the earlier pipeline. |
| `build_county_diurnal_temperature_range.py` / `apply_diurnal_temperature_range_to_climate_data.py` | Builds the 1991-2020 county mean daily Tmax-minus-Tmin artifact and merges `avgDiurnalTempRangeF`. |
| `download_gridmet_data.py` | Downloads 1991-2020 `sph`, `rmax`, and `rmin` NetCDF files by default. |
| `summarize_county_gridmet_humidity.py` / `apply_gridmet_humidity_metric_to_climate_data.py` | Produces county summer specific humidity and a 90 F+ Heat Index day proxy, then merges those metrics and FIPS audit fields. |
| `build_county_representative_points.py` | Creates interior county points used by the lightweight NSRDB workflows. |
| `fetch_nsrdb_representative_point_ghi.py` / `rebuild_nsrdb_representative_point_ghi_summary.py` | Fetches point-based NSRDB GHI or rebuilds its summary from cached responses without another API call. |
| `fetch_nsrdb_representative_point_cloud_metrics.py` | Fetches point-based GHI, clear-sky GHI, and cloud type, then summarizes the clear-sky GHI reduction index. |
| `request_nsrdb_county_polygon_archives.py` | Shared, resumable NSRDB polygon-request engine with tiling, site-count checks, pacing, and Polar fallback. |
| `request_nsrdb_county_polygon_ghi_archives.py` / `request_nsrdb_county_polygon_cloud_archives.py` | Recommended wrappers around the shared request engine, with separate GHI and cloud attributes and output paths. |
| `download_nsrdb_county_polygon_archives.py` | Shared state-machine downloader for completed NSRDB archive jobs. |
| `download_nsrdb_county_polygon_ghi_archives.py` / `download_nsrdb_county_polygon_cloud_archives.py` | Recommended wrappers around the shared downloader, keeping GHI and cloud archives separate. |
| `summarize_nsrdb_county_polygon_archives.py` | Combines county/tile GHI archives into area-weighted county summaries. |
| `summarize_nsrdb_county_polygon_cloud_archives.py` | Combines county/tile cloud archives into area-weighted clear-sky GHI reduction summaries. |
| `apply_nsrdb_cloud_metric_to_climate_data.py` | Merges the clear-sky GHI reduction index, preferring polygon summaries and falling back to representative points. |
The metric-specific NSRDB request and download wrappers are the normal entry
points. The shared engines remain available for custom attributes or artifact
paths. Representative-point results provide a faster first pass; polygon
summaries are the preferred county-area result when available.
Common local enrichment stages, after their source files have been downloaded,
are:
```powershell
.venv\Scripts\python.exe scripts\build_county_locally_extreme_data.py --skip-download
.venv\Scripts\python.exe scripts\apply_locally_extreme_metric_to_climate_data.py
.venv\Scripts\python.exe scripts\build_county_diurnal_temperature_range.py
.venv\Scripts\python.exe scripts\apply_diurnal_temperature_range_to_climate_data.py
.venv\Scripts\python.exe scripts\summarize_county_gridmet_humidity.py --years 1991-2020
.venv\Scripts\python.exe scripts\apply_gridmet_humidity_metric_to_climate_data.py
.venv\Scripts\python.exe scripts\apply_nsrdb_cloud_metric_to_climate_data.py
```
The order matters when rebuilding from scratch: generate the locally extreme
comparison and solar summaries before running their apply step, and summarize
gridMET or NSRDB downloads before merging them. See the data-source document
linked above for acquisition commands, expected artifacts, FIPS handling, and
the representative-point and polygon NSRDB workflows.
Run the current automated tests with: Run the current automated tests with:
```powershell ```powershell
@@ -129,8 +185,8 @@ python -m unittest discover -s tests
The longer-term goal is to add mood-based metrics and investigate whether The longer-term goal is to add mood-based metrics and investigate whether
patterns in those metrics are associated with climate characteristics such as patterns in those metrics are associated with climate characteristics such as
temperature, sunlight availability, clear-sky GHI reduction, humidity, precipitation, or extreme-weather temperature, sunlight availability, clear-sky GHI reduction, humidity,
frequency. precipitation, or extreme-weather frequency.
That phase still requires decisions about: That phase still requires decisions about: