Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions docs/_ext/data_request_tables.py
Original file line number Diff line number Diff line change
Expand Up @@ -214,9 +214,9 @@ def _experiments_table(rows: list[dict[str, str]]) -> list[str]:
_plain(row['start_year_min']),
_plain(row['start_year_max']),
_plain(row['end_year']),
# `historical` carries -1, which is the file's way of saying
# that its duration follows from the start year the modeler
# chose rather than being fixed by the protocol.
# `historical` and `ocx` carry -1, which is the file's way of
# saying that the duration follows from the start year the
# modeler chose rather than being fixed by the protocol.
'set by the start year'
if row['duration'].strip() == '-1'
else _plain(row['duration']),
Expand Down
2 changes: 1 addition & 1 deletion docs/dev/generating-test-files.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@ ismip7-generate-test-files --grid GrIS_16000m --scenario ctrl \
- Time encoding: `days since 1850-01-01`, `calendar='standard'`.
- State (ST) variables: timestamp is Jan 1 of year N+1 (end-of-year snapshot). No `time_bounds`.
- Flux (FL) variables: timestamp is Jul 1 of year N (mid-year), with `time_bounds` = [Jan 1 of N, Jan 1 of N+1].
- `x,y,z,t` variables (e.g. `litemp`): ST snapshots at a sparse set of nominal years. For `historical`: first year of run, 1900 (if in range), last year of run. For projection scenarios: 2100, 2200, 2300 (each if within the simulation year range). The filename year range reflects the full simulation period, not the first/last snapshot year. The set is `CENTURY_SNAPSHOT_YEARS` in `isschecker.checker`, imported here so the generator and the checker cannot disagree; note that no snapshot is written at 2000, which the checker accepts but does not require (see {doc}`../user/time-encoding`).
- `x,y,z,t` variables (e.g. `litemp`): ST snapshots at a sparse set of nominal years. For `historical` and `ocx`: first year of run, 1900 (if in range), last year of run. For projection scenarios: 2100, 2200, 2300 (each if within the simulation year range). The filename year range reflects the full simulation period, not the first/last snapshot year. The set is `CENTURY_SNAPSHOT_YEARS` in `isschecker.checker`, imported here so the generator and the checker cannot disagree; note that no snapshot is written at 2000, which the checker accepts but does not require (see {doc}`../user/time-encoding`).
- An analytic ice sheet, so that the files agree with one another. Most variables are still drawn at random within the min/max range the data request allows them, but seven cannot be: `sftgif`, `sftgrf`, `sftflf`, `lithk`, `topg`, `base` and `orog` are computed from a radially symmetric dome on a bowl-shaped bed, because the checker compares them against each other and a random thickness does not agree with a random mask about where the ice is. The identities hold exactly rather than to a tolerance: grounded plus floating fraction is the ice fraction, thickness is positive exactly where the ice fraction is, surface elevation is base plus thickness, and the ice base rests on the bed where the ice is grounded and above it where it floats. The geometry is deliberately analytic rather than taken from BedMachine or Bedmap: a synthetic file's job is to be predictable, and depending on one dataset would invite depending on datasets for everything else.
- Missing values follow each variable's `fill_policy`. A variable the request defines only where there is ice is written as the fill value wherever the ice fraction is zero, and likewise for grounded-only and floating-only variables; the ones defined only inside the computational domain are missing outside it; and the rest are defined everywhere.
- Single precision (`float32`) for all variables and time.
Expand Down
11 changes: 11 additions & 0 deletions docs/user/checks.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,17 @@ lithk_GrIS_VUW_PISM1_m001_CESM2-WACCM_f001_ctrl_C001_2015-2300.nc
Each field is checked. What the year range *means* is checked under
[Time](#4-time).

The ocx experiment is not forced by an ESM. Its files still have the ESM
field, but any name of letters, digits and hyphens is accepted there, such as
NONE or the reanalysis you used:

```
lithk_GrIS_VUW_PISM1_m001_NONE_f001_ocx_C011_1979-2025.nc
lithk_GrIS_VUW_PISM1_m001_ERA5_f001_ocx_C011_1979-2025.nc
```

Experiment names are lower case: ocx, not OCX as in the forcing directories.

Inside the file, the variable the name promises is present with the
dimensions the data request asks for, in (time, z, y, x) order. Nothing else
is in the file beyond the coordinates and the companion variables CF lets
Expand Down
6 changes: 3 additions & 3 deletions docs/user/data-request.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,9 +67,9 @@ everywhere; see [Value ranges](errors-and-warnings.md#value-ranges).
## Experiments

An experiment's row fixes the nominal years its files may cover, and with
them the time axis every annual file must carry. The historical run is the
one experiment whose start year the modeler chooses; every projection runs
from 2015 to a fixed end year.
them the time axis every annual file must carry. The modeler chooses the
start year of historical and ocx, within the range given; every projection
runs from 2015 to a fixed end year.

```{include} ../_generated/experiments.md
```
Expand Down
1 change: 1 addition & 0 deletions docs/user/time-encoding.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ series. Which nominal years are required depends on the experiment:
|---|---|
| historical | first year of the run, 1900 (if in range), last year of the run (2014) |
| projection (ssp585, ctrl, ...) | 2100, 2200, 2300 (each if within the run) |
| ocx | first year of the run, last year of the run (2025) |

Together, a historical run and a projection give snapshots at the first
historical year, 1900, 2014, 2100, 2200 and 2300. A projection's initial
Expand Down
55 changes: 40 additions & 15 deletions isschecker/checker.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,9 @@
# An unrecognised region costs only the checks that need it (value range,
# grid extent and resolution, crs); the rest of the file is still checked.
# - ISM member id (field 4) matches format mNNN (e.g. m001).
# - ESM name (field 5) is a recognised CMIP6/CMIP7 model name.
# - ESM name (field 5) is a recognised CMIP6/CMIP7 model name. An
# experiment not forced by an ESM (ocx) may put any name of letters,
# digits and hyphens there.
# - Forcing member id (field 6) matches format fNNN (e.g. f001).
# - Set counter (field 8) matches format [C|E|P]NNN (e.g. C001, E041, P132).
# - Year range (field 9) matches format YYYY-YYYY and start <= end. What the
Expand Down Expand Up @@ -73,7 +75,7 @@
# cadence reasoning cannot see.
# - For x,y,z,t snapshot variables the axis is instead checked against the
# required set of snapshot nominal years: the run's last year, the century
# marks inside it, and (for historical only) the run's first year. A
# marks inside it, and (for historical and ocx only) the run's first year. A
# missing snapshot is an error; an unasked-for one is a warning, since the
# request specifies the snapshots as a minimum set. That includes a
# snapshot at 2000, which the data request does not ask for -- see
Expand Down Expand Up @@ -253,6 +255,14 @@ def _read_data_csv(filename: str, **kwargs) -> pd.DataFrame:
"UKESM2-0-LL",
}

# Experiments not forced by an ESM. OCX is forced by reanalysis (ERA5 for the
# atmosphere, EN4 for the Greenland ocean) and by expert judgement (the
# Antarctic ocean), and groups may choose other products, so no one name
# describes it. Its files still carry the ESM field, so that every file name
# has the same fields, but any name of letters, digits and hyphens will do:
# a placeholder such as NONE, the forcing's source, or even a CMIP model name.
EXPERIMENTS_WITHOUT_ESM = frozenset({"ocx"})


class Reporter:
"""Writes the log and counts what it wrote, by severity and category.
Expand Down Expand Up @@ -698,10 +708,10 @@ def _timestamp_to_nominal_year(timestamp, var_type: str) -> int:
def _expected_nominal_years(exp: dict, start_year: int) -> list[int]:
"""The nominal simulation years a file for this experiment should carry.

Every experiment but 'historical' pins its start year in
Every experiment but 'historical' and 'ocx' pins its start year in
experiments_ismip7.csv, so the whole axis follows from the table alone.
'historical' may start anywhere in [start_year_min, start_year_max] (which
is what a duration of -1 records), so it needs one number from the file --
Those two may start anywhere in [start_year_min, start_year_max] (which is
what a duration of -1 records), so they need one number from the file --
see :func:`_axis_start_year`.
"""
if exp["duration"] != -1:
Expand Down Expand Up @@ -866,9 +876,9 @@ def _required_snapshot_years(exp: dict, run_years: list[int]) -> set[int]:

The final year of the run is always reported, as are the century marks
inside it. The *first* year is only required where the protocol leaves it
open -- that is, for 'historical', whose start year is the modeller's choice
(duration -1) and so is not recorded anywhere else. A projection's initial
state is the historical run's final state, already reported as historical's
open -- that is, for 'historical' and 'ocx', whose start year is the
modeller's choice (duration -1) and so is not recorded anywhere else. A
projection's initial state is the historical run's final state, already reported as historical's
last-year snapshot, so requiring it again in every projection would ask for
the same field twice.
"""
Expand Down Expand Up @@ -1279,12 +1289,21 @@ def _process_single_experiment(
reporter.write(" ** Experiment: " + experiment_name + "\n ")
reporter.write("**********************************************************\n")
reporter.write("\n ")
experiment_names = [exp["experiment"] for exp in experiments]
# The forcing directories are named OCX, so a group copying them gets
# the case wrong; experiment names are lower case, like ssp585.
case_hint = (
f" Experiment names are lower case: '{experiment_name.lower()}'."
if experiment_name.lower() in experiment_names
else ""
)
naming_reporter.error(
"The compliance check is ignored for experiment "
+ experiment_name
+ " as it is not in "
+ str([exp["experiment"] for exp in experiments])
+ str(experiment_names)
+ "."
+ case_hint
)
report_naming_issues.append(
"Compliance check ignored : experiment "
Expand Down Expand Up @@ -1493,7 +1512,13 @@ def _check_naming(
)

esm_name = parts[ISMIP7_FILENAME_ESM_IDX]
if esm_name not in VALID_ESM_NAMES:
experiment_name = parts[ISMIP7_FILENAME_EXPERIMENT_IDX]
if experiment_name in EXPERIMENTS_WITHOUT_ESM:
if not re.fullmatch(r"[A-Za-z0-9-]+", esm_name):
reporter.error(
f"ESM name '{esm_name}' (field {ISMIP7_FILENAME_ESM_IDX}) must be a name of letters, digits and hyphens for experiment '{experiment_name}' (e.g. NONE or ERA5)."
)
elif esm_name not in VALID_ESM_NAMES:
reporter.error(
f"ESM name '{esm_name}' (field {ISMIP7_FILENAME_ESM_IDX}) is not a recognised CMIP6/CMIP7 model name."
)
Expand Down Expand Up @@ -2743,11 +2768,11 @@ def _check_time(
def _axis_start_year(exp: dict, filename_years, actual: list, var_type: str) -> int:
"""The nominal year the expected time axis should begin at.

Only 'historical' has a say in the matter -- every other experiment pins its
start year in experiments_ismip7.csv -- and for it the file name is what
decides: it is the file's own declared claim about its contents, and any
start year in [start_year_min, start_year_max] is permitted, so there is
nothing else to measure the file against.
Only 'historical' and 'ocx' have a say in the matter -- every other
experiment pins its start year in experiments_ismip7.csv -- and for them the
file name is what decides: it is the file's own declared claim about its
contents, and any start year in [start_year_min, start_year_max] is
permitted, so there is nothing else to measure the file against.

Falls back to the axis when the file name cannot supply a usable year, so
that the cadence and the end year are still checked rather than the file
Expand Down
1 change: 1 addition & 0 deletions isschecker/data/experiments_ismip7.csv
Original file line number Diff line number Diff line change
Expand Up @@ -4,3 +4,4 @@ ssp370;2015;2015;2100;86
ssp126;2015;2015;2300;286
ssp585;2015;2015;2300;286
ctrl;2015;2015;2300;286
ocx;1950;2015;2025;-1
2 changes: 1 addition & 1 deletion isschecker/generate.py
Original file line number Diff line number Diff line change
Expand Up @@ -649,7 +649,7 @@ def create_netcdf_file(output_file, grid_name='GrIS_16000m', scenario='ctrl', st
snap_set = {end_year} | {
y for y in CENTURY_SNAPSHOT_YEARS if start_year <= y <= end_year
}
if scenario == 'historical':
if scenario in ('historical', 'ocx'):
snap_set.add(start_year)
snapshot_years = sorted(snap_set)
origin = datetime(1850, 1, 1).date()
Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"

[project]
name = "isschecker"
version = "0.5.1"
version = "0.6.0"
description = "ISMIP7 compliance checker for ice sheet model simulation NetCDF datasets"
readme = "README.md"
license = "MIT"
Expand Down
129 changes: 129 additions & 0 deletions tests/test_compliance_checker.py
Original file line number Diff line number Diff line change
Expand Up @@ -639,6 +639,135 @@ def test_checker_reports_historical_time_range_violation(case_dir):
)


def generate_ocx_run(root: Path, start_year: int, include_xyt: bool) -> Path:
"""An OCX run ending in 2025, with NONE in the ESM field."""
generate_test_files.create_netcdf_file(
None,
grid_name="GrIS_16000m",
scenario="ocx",
start_year=start_year,
nyears=2025 - start_year + 1,
include_scalars=not include_xyt,
include_xyt=include_xyt,
include_non_mandatory=include_xyt,
esm_id="NONE",
set_counter="C011",
output_root=root,
)
return root / "GrIS" / "ISMIP7" / "SYNTH1" / "CORE" / "C011"


def test_checker_accepts_an_ocx_run(tmp_path):
core_dir = generate_ocx_run(tmp_path / "gen", 1950, include_xyt=False)

summary = run_checker(core_dir)

assert summary["total_errors"] == 0, summary["log_text"]
assert (
"covering nominal years 1950-2025, as experiment 'ocx' requires: OK"
in summary["log_text"]
)


@pytest.mark.parametrize("esm_name", ["ERA5", "CESM2-WACCM"])
def test_checker_accepts_any_esm_name_in_ocx(tmp_path, esm_name):
core_dir = generate_ocx_run(tmp_path / "gen", 2015, include_xyt=False)
for file_path in sorted(core_dir.glob("*.nc")):
rename_file_part(file_path, checker.ISMIP7_FILENAME_ESM_IDX, esm_name)

summary = run_checker(core_dir)

assert summary["total_errors"] == 0, summary["log_text"]


@pytest.mark.parametrize("esm_name", ["", "RACMO2.3p2-ERA"])
def test_checker_reports_a_malformed_esm_name_in_ocx(tmp_path, esm_name):
core_dir = generate_ocx_run(tmp_path / "gen", 2015, include_xyt=False)
rename_file_part(
first_dataset(core_dir), checker.ISMIP7_FILENAME_ESM_IDX, esm_name
)

summary = run_checker(core_dir)

assert summary["total_errors"] == 1
assert summary["total_naming_errors"] == 1
assert (
f"ESM name '{esm_name}' (field 5) must be a name of letters, digits and"
" hyphens for experiment 'ocx' (e.g. NONE or ERA5)." in summary["log_text"]
)


def test_checker_reports_an_ocx_file_without_an_esm_field(tmp_path):
core_dir = generate_ocx_run(tmp_path / "gen", 2015, include_xyt=False)
file_path = dataset_for_variable(core_dir, "lim")
file_path.rename(file_path.with_name(file_path.name.replace("_NONE_", "_")))

summary = run_checker(core_dir)

assert summary["total_errors"] == 2
assert (
"In experiment ocx, these mandatory variable(s) is (are) missing: ['lim']"
in summary["log_text"]
)


def test_checker_accepts_ocx_snapshots_at_the_first_and_last_year(tmp_path):
"""OCX does not branch from historical, so its initial state is required."""
core_dir = generate_ocx_run(tmp_path / "gen", 2010, include_xyt=True)

summary = run_xyt_checker(core_dir)

assert summary["total_errors"] == 0, summary["log_text"]

set_time_axis(
dataset_for_variable(core_dir, "litemp"), state_timestamps([2025])
)
summary = run_xyt_checker(core_dir)

assert summary["total_time_errors"] == 1
assert (
"required snapshot nominal year(s) missing: 2010" in summary["log_text"]
)


def test_checker_reports_ocx_starting_before_1950(tmp_path):
core_dir = generate_ocx_run(tmp_path / "gen", 1949, include_xyt=False)

summary = run_checker(core_dir)

assert (
"the file name starts at year 1949, but experiment 'ocx' must be"
" between 1950 and 2015." in summary["log_text"]
)


def test_checker_reports_ocx_ending_before_2025(tmp_path):
core_dir = generate_ocx_run(tmp_path / "gen", 2015, include_xyt=False)
target_file = dataset_for_variable(core_dir, "lim")
set_time_axis(target_file, state_timestamps(range(2015, 2025)))
rename_file_part(
target_file, checker.ISMIP7_FILENAME_YEAR_RANGE_IDX, "2015-2024.nc"
)

summary = run_checker(core_dir)

assert summary["total_naming_errors"] == 0
assert (
"the file name ends at year 2024, but experiment 'ocx' must end at 2025."
in summary["log_text"]
)


def test_checker_says_experiment_names_are_lower_case(tmp_path):
"""The forcing directories are named OCX; the experiment is ocx."""
core_dir = generate_ocx_run(tmp_path / "gen", 2015, include_xyt=False)
for file_path in sorted(core_dir.glob("*.nc")):
rename_file_part(file_path, checker.ISMIP7_FILENAME_EXPERIMENT_IDX, "OCX")

summary = run_checker(core_dir)

assert "Experiment names are lower case: 'ocx'." in summary["log_text"]

def test_checker_reports_time_axis_hollowed_out_between_correct_endpoints(
ctrl_case_dir,
):
Expand Down
Loading