Refine climate metrics and data pipeline
This commit is contained in:
@@ -45,14 +45,14 @@ If you have monthly nClimGrid history files (for example `nclimgrid_tavg.nc`) ra
|
||||
- Original Plotly county geometry source: [https://raw.githubusercontent.com/plotly/datasets/master/geojson-counties-fips.json](https://raw.githubusercontent.com/plotly/datasets/master/geojson-counties-fips.json)
|
||||
- Official county geometry reference (Census TIGER/Line): [https://www.census.gov/geographies/mapping-files/time-series/geo/tiger-line-file.html](https://www.census.gov/geographies/mapping-files/time-series/geo/tiger-line-file.html)
|
||||
|
||||
## Source 4: Solar resource (`avgSolarGhiKwhM2Day`)
|
||||
## Source 4: Solar resource (`meanDailyGlobalHorizontalRadiationKwhM2Day`)
|
||||
|
||||
- Recommended dataset: NREL National Solar Radiation Database (NSRDB)
|
||||
- Data/API page: [https://developer.nrel.gov/docs/solar/nsrdb/](https://developer.nrel.gov/docs/solar/nsrdb/)
|
||||
- Maps/geospatial data page: [https://www.nrel.gov/gis/solar-resource-maps](https://www.nrel.gov/gis/solar-resource-maps)
|
||||
- Fallback point API option: NASA POWER `ALLSKY_SFC_SW_DWN` (surface shortwave downwelling radiation), [https://power.larc.nasa.gov/docs/tutorials/service-data-request/api/](https://power.larc.nasa.gov/docs/tutorials/service-data-request/api/)
|
||||
|
||||
The app metric is designed for annual average daily global horizontal irradiance (GHI), in `kWh/m2/day`.
|
||||
The app metric is designed for mean daily global horizontal radiation (GHI), in `kWh/m2/day`.
|
||||
For county means, use a gridded annual GHI raster and pass it to the generator with `--solar-ghi-raster`.
|
||||
|
||||
### First test: NSRDB county representative points
|
||||
@@ -80,12 +80,12 @@ Outputs:
|
||||
|
||||
- `data/nrel/county_representative_points.csv`: county FIPS, name, state, latitude, and longitude.
|
||||
- `data/nrel/representative_point_csv/`: cached raw NSRDB CSV responses by county FIPS.
|
||||
- `data/nrel/county_representative_point_ghi_summary.csv`: summarized `avgSolarGhiKwhM2Day` values.
|
||||
- `data/nrel/county_representative_point_ghi_summary.csv`: summarized `meanDailyGlobalHorizontalRadiationKwhM2Day` values.
|
||||
|
||||
The calculation is:
|
||||
|
||||
```text
|
||||
avgSolarGhiKwhM2Day = sum(hourly GHI) / 1000 / 365
|
||||
meanDailyGlobalHorizontalRadiationKwhM2Day = sum(hourly GHI) / 1000 / 365
|
||||
```
|
||||
|
||||
If the raw cache is complete but the summary CSV only contains the last fetched batch, rebuild the summary from cached files without calling the API:
|
||||
@@ -96,9 +96,9 @@ If the raw cache is complete but the summary CSV only contains the last fetched
|
||||
|
||||
This point-based workflow is easier to validate and resume than full polygon downloads, but it is an approximation of county sunlight rather than an area-weighted county mean.
|
||||
|
||||
### First cloud-cover pass: NSRDB representative points
|
||||
### First clear-sky GHI reduction pass: NSRDB representative points
|
||||
|
||||
For a fast cloud-cover input layer, fetch representative-point NSRDB CSVs with observed GHI, Clearsky GHI, and Cloud Type:
|
||||
For a fast clear-sky GHI reduction input layer, fetch representative-point NSRDB CSVs with observed GHI, Clearsky GHI, and Cloud Type:
|
||||
|
||||
```powershell
|
||||
.venv\Scripts\python.exe scripts\fetch_nsrdb_representative_point_cloud_metrics.py --limit 10
|
||||
@@ -113,18 +113,18 @@ Once the first batch looks right, fetch every county:
|
||||
Outputs:
|
||||
|
||||
- `data/nrel/representative_point_cloud_csv/`: cached raw NSRDB CSV responses with `ghi,clearsky_ghi,cloud_type`.
|
||||
- `data/nrel/county_representative_point_cloud_summary.csv`: summarized cloud-cover proxy fields.
|
||||
- `data/nrel/county_representative_point_cloud_summary.csv`: summarized clear-sky GHI reduction fields.
|
||||
- `data/nrel/county_representative_point_cloud_error_log.csv`: failed county requests with redacted API context.
|
||||
|
||||
The primary calculation uses daylight rows where Clearsky GHI is at least 50 W/m2:
|
||||
|
||||
```text
|
||||
cloudinessIndexPct = 1 - mean(clamped(GHI / Clearsky GHI, 0, 1))
|
||||
clearSkyGhiReductionIndex = 1 - mean(clamped(GHI / Clearsky GHI, 0, 1))
|
||||
```
|
||||
|
||||
The same summary also stores daylight row counts, observed-to-clear-sky ratio, and broad Cloud Type frequency buckets. This is faster than the polygon archive workflow, but it remains a representative-point county approximation.
|
||||
|
||||
Apply the representative-point cloudiness metric to the browser app CSV:
|
||||
Apply the representative-point clear-sky GHI reduction metric to the browser app CSV:
|
||||
|
||||
```powershell
|
||||
.venv\Scripts\python.exe scripts\apply_nsrdb_cloud_metric_to_climate_data.py
|
||||
@@ -134,7 +134,7 @@ The current cloud summary covers the 3,143 county representative points and leav
|
||||
|
||||
### County-average target: NSRDB polygon cloud archive requests
|
||||
|
||||
For a less noisy county-level cloudiness layer, submit county polygons to the NSRDB archive workflow with `ghi,clearsky_ghi,cloud_type`. This uses the same polygon tiling and pacing logic as the GHI archive workflow, but writes separate cloud manifests, response JSON files, archives, and summaries.
|
||||
For a less noisy county-level clear-sky GHI reduction layer, submit county polygons to the NSRDB archive workflow with `ghi,clearsky_ghi,cloud_type`. This uses the same polygon tiling and pacing logic as the GHI archive workflow, but writes separate cloud manifests, response JSON files, archives, and summaries.
|
||||
|
||||
Start with a dry run:
|
||||
|
||||
@@ -166,7 +166,7 @@ Download completed archives:
|
||||
.venv\Scripts\python.exe scripts\download_nsrdb_county_polygon_cloud_archives.py
|
||||
```
|
||||
|
||||
Summarize downloaded archives into area-weighted county cloudiness:
|
||||
Summarize downloaded archives into area-weighted county clear-sky GHI reduction:
|
||||
|
||||
```powershell
|
||||
.venv\Scripts\python.exe scripts\summarize_nsrdb_county_polygon_cloud_archives.py --reuse-existing-output
|
||||
@@ -178,7 +178,7 @@ Outputs:
|
||||
- `data/nrel/county_polygon_cloud_error_log.csv`: cloud archive request errors.
|
||||
- `data/nrel/polygon_cloud_request_responses/`: raw NSRDB cloud acknowledgement JSON files.
|
||||
- `data/nrel/polygon_cloud_archives/`: downloaded cloud ZIP archives.
|
||||
- `data/nrel/county_polygon_cloud_summary.csv`: area-weighted county cloudiness summaries.
|
||||
- `data/nrel/county_polygon_cloud_summary.csv`: area-weighted county clear-sky GHI reduction summaries.
|
||||
|
||||
Then update the browser app CSV. Polygon area-weighted values are used first when `data/nrel/county_polygon_cloud_summary.csv` exists; representative-point values remain the fallback:
|
||||
|
||||
@@ -228,7 +228,7 @@ Outputs:
|
||||
- `data/nrel/county_polygon_ghi_error_log.csv`: counties that need retrying or tiling.
|
||||
- `data/nrel/polygon_request_responses/`: raw API acknowledgement JSON files.
|
||||
- `data/nrel/polygon_archives/`: downloaded county or tiled county ZIP archives.
|
||||
- `data/nrel/county_polygon_ghi_summary.csv`: summarized polygon archive GHI, including `avgSolarGhiKwhM2Day`.
|
||||
- `data/nrel/county_polygon_ghi_summary.csv`: summarized polygon archive GHI, including `meanDailyGlobalHorizontalRadiationKwhM2Day`.
|
||||
|
||||
Download completed GHI archives:
|
||||
|
||||
@@ -262,8 +262,10 @@ Then update the app CSV. Polygon archive GHI is used first; representative-point
|
||||
- `wettestPrecipMonth`: month with the highest 1991-2020 county mean precipitation total.
|
||||
- `driestPrecipMonth`: month with the lowest 1991-2020 county mean precipitation total.
|
||||
- previous `extremeDays` / `oldExtremeDays`: count of daily-normal or monthly-proxy days where county mean `tmax >= 95F` or `tmin <= 32F` (thresholds configurable in script). This is preserved only as an audit column after the NOAA nClimGrid-Daily metrics are applied.
|
||||
- `avgSolarGhiKwhM2Day`: county mean annual average daily GHI, in `kWh/m2/day`. The current app CSV uses NSRDB polygon archive area-weighted values where available, with representative-point values kept as fallback.
|
||||
- `cloudinessIndexPct`: representative-point NSRDB daylight cloudiness proxy derived from observed GHI divided by Clearsky GHI. Higher values mean observed irradiance is lower relative to modeled clear-sky irradiance. Despite the legacy field name, values are stored on a 0-1 scale.
|
||||
- `humidHeatDays`: average annual count of days where estimated Heat Index is at least 90 F, the lower bound of the NWS Extreme Caution category. The NWS Heat Index algorithm is applied to NOAA nClimGrid-Daily county `tmax` and gridMET daily minimum relative humidity (`rmin`) for 1991-2020. Because the inputs are paired daily extrema rather than coincident hourly observations, this is an estimated daily-peak proxy.
|
||||
- `avgSummerSpecificHumidityGKg`: county-cell-weighted mean gridMET specific humidity for June-August 1991-2020, converted from kg/kg to g/kg.
|
||||
- `meanDailyGlobalHorizontalRadiationKwhM2Day`: county mean daily global horizontal radiation (GHI), in `kWh/m2/day`. The current app CSV uses NSRDB polygon archive area-weighted values where available, with representative-point values kept as fallback.
|
||||
- `clearSkyGhiReductionIndex`: NSRDB daylight clear-sky GHI reduction index derived from observed GHI divided by modeled clear-sky GHI. Higher values mean observed irradiance is lower relative to clear-sky conditions. Values are stored on a 0-1 scale.
|
||||
|
||||
## Run the generator
|
||||
|
||||
@@ -297,7 +299,7 @@ Notes:
|
||||
- For counties outside CONUS coverage in NOAA gridded files, fallback values are applied by the script when no valid grid values intersect.
|
||||
- For physically-based daily `extremeDays`, provide true daily grids and set `--extreme-days-mode require-daily`.
|
||||
- If `--counties-geojson` does not exist locally, the script will try to download the county GeoJSON automatically from the Plotly URL above and cache it at that path.
|
||||
- `--solar-ghi-csv` is optional and can load county-keyed solar summaries into `avgSolarGhiKwhM2Day`.
|
||||
- `--solar-ghi-csv` is optional and can load county-keyed solar summaries into `meanDailyGlobalHorizontalRadiationKwhM2Day`.
|
||||
- `--solar-ghi-raster` is optional and takes precedence over `--solar-ghi-csv`. If you have a gridded annual GHI raster, add `--solar-ghi-raster path/to/annual_ghi_kwh_m2_day.tif` for a true county-area raster mean.
|
||||
|
||||
## Source 5: NOAA nClimGrid-Daily county area averages (`locallyExtremeDays`)
|
||||
|
||||
Reference in New Issue
Block a user