Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
26 commits
Select commit Hold shift + click to select a range
4a9f7b7
Version which implements the greedy, largest-smallest allocation rout…
ddahawkins-TUDelft Jul 7, 2026
1f79a5f
Updated EIA target capacity allocation routine
ddahawkins-TUDelft Jul 7, 2026
5700a53
Working version with cleaned code
ddahawkins-TUDelft Jul 8, 2026
8b8c7df
Added some comments to explain
ddahawkins-TUDelft Jul 8, 2026
5f914f5
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 8, 2026
6dfe258
Capacity Profile Method Now Addresses Retired Assets
ddahawkins-TUDelft Jul 9, 2026
ed6753b
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 9, 2026
3a33748
Updated figures to include retirement profiles
ddahawkins-TUDelft Jul 9, 2026
557632a
Working version
ddahawkins-TUDelft Jul 21, 2026
998b702
Working version with Planned Statues and addresses original PR concerns
ddahawkins-TUDelft Jul 21, 2026
9875124
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 21, 2026
33ea138
Added schema validation for start year and end year imputation sources
ddahawkins-TUDelft Jul 21, 2026
9796752
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 21, 2026
4bef4fd
Commit removes excessive diagnostics and leaves only the 'Capacity Da…
ddahawkins-TUDelft Jul 21, 2026
46ff3ae
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 21, 2026
0733140
Updated capacity stock calls
ddahawkins-TUDelft Jul 30, 2026
1d179be
Update Pixi Snakemake environment export task
ddahawkins-TUDelft Jul 30, 2026
8c1a9a1
Addressed PR feedback
ddahawkins-TUDelft Jul 30, 2026
c06dbad
Adjusted plot colours and fixed dodgy fucntion call bug
ddahawkins-TUDelft Jul 30, 2026
4ce0e98
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 30, 2026
8afbc1e
Small non-code fixes from review 2
irm-codebase Jul 31, 2026
1555409
Merge branch 'main' into feature/deterministic-age-imputation
irm-codebase Jul 31, 2026
6994c30
Fix time imputation edge cases
ddahawkins-TUDelft Jul 31, 2026
f144acb
Remove unused time imputation colour helper
ddahawkins-TUDelft Jul 31, 2026
8736400
revert arrangement of functions (for review)
irm-codebase Aug 3, 2026
e38d53f
Small fixes to schema checks
irm-codebase Aug 3, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
78 changes: 78 additions & 0 deletions config/config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,7 @@ imputation:
utility pv: "land"
time:
scenario: "construction" # Options: historical->construction->pre_construction->announced
start_year_imputation_method: capacity_profile
lifetime_years:
# General lifetime assumptions. Adjust as needed!
bioenergy turbine: 25
Expand Down Expand Up @@ -161,3 +162,80 @@ imputation:
smr reactor: 50
steam turbine: 5
utility pv: 2
planned_commissioning_year_windows:
bioenergy turbine:
construction: [1, 3]
pre-construction: [2, 5]
announced: [4, 8]
ccgt:
construction: [1, 3]
pre-construction: [2, 5]
announced: [4, 8]
coal turbine:
construction: [1, 3]
pre-construction: [2, 5]
announced: [4, 8]
concentrating solar power:
construction: [1, 4]
pre-construction: [3, 7]
announced: [5, 10]
high enthalpy geothermal:
construction: [2, 5]
pre-construction: [3, 7]
announced: [5, 10]
igcc:
construction: [2, 5]
pre-construction: [4, 8]
announced: [6, 12]
low enthalpy geothermal:
construction: [1, 4]
pre-construction: [3, 6]
announced: [4, 9]
nuclear reactor:
construction: [3, 7]
pre-construction: [5, 13]
announced: [8, 18]
ocgt:
construction: [1, 3]
pre-construction: [2, 5]
announced: [4, 8]
offshore:
construction: [2, 5]
pre-construction: [3, 8]
announced: [5, 12]
onshore:
construction: [1, 3]
pre-construction: [2, 6]
announced: [3, 10]
pumped storage:
construction: [3, 8]
pre-construction: [5, 12]
announced: [8, 18]
reciprocating engine:
construction: [1, 3]
pre-construction: [2, 5]
announced: [4, 8]
reservoir:
construction: [3, 8]
pre-construction: [5, 12]
announced: [8, 18]
rooftop pv:
construction: [1, 2]
pre-construction: [2, 3]
announced: [3, 5]
run of river:
construction: [2, 6]
pre-construction: [4, 10]
announced: [6, 15]
smr reactor:
construction: [2, 6]
pre-construction: [4, 10]
announced: [7, 15]
steam turbine:
construction: [1, 3]
pre-construction: [2, 5]
announced: [4, 8]
utility pv:
construction: [1, 3]
pre-construction: [2, 5]
announced: [3, 8]
52 changes: 52 additions & 0 deletions workflow/internal/config.schema.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,43 @@ $defs:
type: integer
minimum: 1

planned_commissioning_year_window:
type: array
description: >
Inclusive minimum and maximum commissioning-year offsets relative
to the dataset year.
minItems: 2
maxItems: 2
items:
type: integer
minimum: 1

planned_commissioning_status_windows:
type: object
additionalProperties: false
required:
- construction
- pre-construction
- announced
properties:
construction:
$ref: "#/$defs/planned_commissioning_year_window"
pre-construction:
$ref: "#/$defs/planned_commissioning_year_window"
announced:
$ref: "#/$defs/planned_commissioning_year_window"

planned_commissioning_technology_windows:
title: Planned commissioning year windows
description: |
Technology- and status-specific year-offset windows used to impute
missing commissioning years for planned powerplants. Each window is
expressed as [minimum_year_offset, maximum_year_offset] relative to
the dataset year (inclusive).
type: object
additionalProperties:
$ref: "#/$defs/planned_commissioning_status_windows"

technology_mapping_base:
title: Technology Mapping
description: >
Expand Down Expand Up @@ -423,8 +460,10 @@ properties:
additionalProperties: false
required:
- scenario
- start_year_imputation_method
- lifetime_years
- retirement_delay_years
- planned_commissioning_year_windows
properties:
scenario:
title: Scenario
Expand All @@ -444,3 +483,16 @@ properties:
$ref: "#/$defs/year_map"
retirement_delay_years:
$ref: "#/$defs/year_map"
start_year_imputation_method:
title: Start year imputation method
description: |
Method used to impute missing start years when backfilling from known
end years is not possible.

- 'capacity_profile' assigns missing start years deterministically using
capacity-weighted commissioning profiles from annual commissioning statistics.
type: string
enum:
- capacity_profile
planned_commissioning_year_windows:
$ref: "#/$defs/planned_commissioning_technology_windows"
5 changes: 5 additions & 0 deletions workflow/report/impute_time_profile.rst
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
Annual commissioning-capacity imputation profile.

Bars show commissioning capacity by start year, split into observed and imputed
powerplants. The line shows the target annual commissioning profile used during
imputation.
21 changes: 20 additions & 1 deletion workflow/rules/_utils.smk
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ IMPUTED_CAT_WITHOUT_ADJUSTMENT = {"large_solar"}


def additional_config_validation():
"""Ensures technology mapping and lifetime-related names match."""
"""Ensure technology mappings and time-imputation settings match."""
lifetime_set = set(config["imputation"]["time"]["lifetime_years"].keys())
mismatch = lifetime_set ^ set(
config["imputation"]["time"]["retirement_delay_years"]
Expand All @@ -54,6 +54,25 @@ def additional_config_validation():
f"Technology mapping does not match lifetime technologies for {mismatch}"
)

planned_windows = config["imputation"]["time"]["planned_commissioning_year_windows"]
mismatch = tech_map_set ^ set(planned_windows)
if mismatch:
raise ValueError(
"Technology mapping does not match planned commissioning windows "
f"for {mismatch}"
)
reversed_windows = {
f"{technology}/{status}": window
for technology, status_windows in planned_windows.items()
for status, window in status_windows.items()
if window[0] > window[1]
}
if reversed_windows:
raise ValueError(
"Planned commissioning windows must be ordered [minimum, maximum]: "
f"{reversed_windows}"
)


def get_excluded_powerplant_ids(category):
"""Handle cases where the naming between files and configuration mismatch."""
Expand Down
15 changes: 14 additions & 1 deletion workflow/rules/impute.smk
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,10 @@ rule impute_location:
rule impute_time:
input:
relocated=rules.impute_location.output.relocated,
category_capacity=(
"<resources>/automatic/shapes/{shapes}/"
"statistics/category_capacity.parquet"
),
output:
aged=workflow.pathvars.apply("<powerplants>").format(
shapes="{shapes}",
Expand All @@ -63,10 +67,19 @@ rule impute_time:
category="Powerplants module",
subcategory="{category}",
),
capacity_date_events=(
"<resources>/automatic/shapes/{shapes}/impute_time/{category}_capacity_date_events.parquet"
),
capacity_date_plot=report(
"<results>/{shapes}/powerplants/unadjusted/{category}_capacity_date_events.pdf",
caption="../report/impute_time_profile.rst",
category="Powerplants module",
subcategory="{category}",
),
log:
"<logs>/{shapes}/{category}/impute_time.log",
wildcard_constraints:
dataset="|".join(IMPUTED_CAT),
category="|".join(IMPUTED_CAT),
conda:
"../envs/module.yaml"
params:
Expand Down
24 changes: 24 additions & 0 deletions workflow/scripts/_plots.py
Original file line number Diff line number Diff line change
@@ -1,10 +1,13 @@
"""Plot functions used in one or more rules."""

from collections.abc import Collection

import _schemas
import _utils
import geopandas as gpd
import numpy as np
import pandas as pd
from cmap import Colormap
from matplotlib import pyplot as plt
from matplotlib import ticker as mticker
from matplotlib.axes import Axes
Expand Down Expand Up @@ -238,3 +241,24 @@ def plot_capacity_aggregation(
ax.set_axis_off()
ax.set_title(title + f" in year {agg.attrs['year']}")
fig.savefig(output_file, bbox_inches="tight")


def get_colour_dict(
sources: Collection[str],
colormap: str,
*,
value_range: tuple[float, float] = (0.0, 1.0),
) -> dict[str, str]:
"""Return deterministic colours for a collection of source types."""
sorted_sources = sorted(sources)

if not sorted_sources:
return {}

cmap = Colormap(colormap).to_mpl()
values = np.linspace(value_range[0], value_range[1], len(sorted_sources))

return {
source: colour
for source, colour in zip(sorted_sources, cmap(values), strict=True)
}
9 changes: 9 additions & 0 deletions workflow/scripts/_schemas.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@
# ruff: noqa: UP007
from typing import Literal

import _utils
import pandera.pandas as pa
from pandera.pandas import DataFrameModel, Field, check
from pandera.typing.geopandas import GeoSeries
Expand Down Expand Up @@ -87,6 +88,14 @@ class Config:
"Expected decommissioning year."
status: Series[str]
"Known state of the project."
start_year_source_type: Series[str] | None = Field(
isin=_utils.date_source_types_for("start_year")
)
"Source/provenance label for the start year."
end_year_source_type: Series[str] | None = Field(
isin=_utils.date_source_types_for("end_year")
)
"Source/provenance label for the end year."
# Location / size
geometry: GeoSeries[Point] = Field()
"Powerplant point data."
Expand Down
68 changes: 68 additions & 0 deletions workflow/scripts/_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,74 @@ def listify(item) -> list:
}
EIA_CAT_MAPPING = {k: listify(v) for k, v in EIA_CAT_MAPPING.items()}

DATE_SOURCE_METADATA = {
"observed": {
Comment on lines +52 to +53

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could be removed with some smart helpers. See below.

"label": "Observed date (from powerplant data)",
"applies_to": {"start_year", "end_year"},
},
"derived_from_end_year": {
"label": "Start date derived from observed end date",
"applies_to": {"start_year"},
},
"imputed_capacity_profile": {
"label": "Start date imputed from historical commissioning-profile",
"applies_to": {"start_year"},
},
"imputed_construction_window": {
"label": "Start date imputed within construction window",
"applies_to": {"start_year"},
},
"imputed_pre_construction_window": {
"label": "Start date imputed within pre-construction window",
"applies_to": {"start_year"},
},
"imputed_announced_window": {
"label": "Start date imputed within announced window",
"applies_to": {"start_year"},
},
"derived_from_imputed_retirement_end_year": {
"label": "Start date derived from retirement-profile end date",
"applies_to": {"start_year"},
},
"derived_from_start_year_lifetime": {
"label": "End date derived from start date and lifetime",
"applies_to": {"end_year"},
},
"imputed_retirement_capacity_profile": {
"label": "End date imputed from retirement-profile",
"applies_to": {"end_year"},
},
"derived_from_start_year_lifetime_capped_to_retired_status": {
"label": "End date derived from start date but capped to retired status",
"applies_to": {"end_year"},
},
"derived_from_start_year_lifetime_with_retirement_delay": {
"label": "End date derived from start date, lifetime, and retirement delay",
"applies_to": {"end_year"},
},
"observed_adjusted_with_retirement_delay": {
"label": "Observed end date adjusted with retirement delay",
"applies_to": {"end_year"},
},
}


def date_source_types_for(year_column: str) -> set[str]:
"""Return source types valid for a date column."""
return {
source_type
for source_type, metadata in DATE_SOURCE_METADATA.items()
if year_column in metadata["applies_to"]
}


def date_source_labels() -> dict[str, str]:
"""Return date-source display labels."""
return {
source_type: metadata["label"]
for source_type, metadata in DATE_SOURCE_METADATA.items()
}
Comment on lines +113 to +118

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: this kind of label boilerplate code can be avoided with the help of the inflection library, which can be installed like so:

pixi add --feature module inflection  # installs the lib for the module
pixi run export-snakemake-env module  # export the environment so rules can use it

Then just run:

from inflection import humanize

print(humanize("imputed_capacity_profile")
# Imputed capacity profile

This will make the the name easy to read while matching the dataset naming, which helps avoid confusion. It also forces the dev to come up with concise but clear naming too 😉

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I understand what you are asking, but humanize() doesnt add all the context that might helper a user unless i make all the variable names very very long. So I would prefer to keep my DATE_SOURCE_METADATA as-is .... it provides a "single source of truth' generally for this rule so it would have to exist anyway.



def get_eia_stats_in_cat_yr(
stats: pd.DataFrame, year: int, category: str
Expand Down
Loading