Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions changelog.d/canonical-bundle-6-0-0.changed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Move the API bundle to PolicyEngine 6.0.0, which binds policyengine-core 3.32.5, policyengine-us 2.2.1, policyengine-uk 2.90.2 and spm-calculator 1.0.0, and raise the separate spm-calculator pin that both pip build paths install to 1.0.0 so the calculator stays compatible with the country model. Qualify the SPM housing cap against households that actually receive housing assistance, and record that ordinary household income counts the real award rather than the capped SPM resource.
2 changes: 1 addition & 1 deletion docker/Dockerfile
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
FROM python:3.12
# Match the API bundle in pyproject.toml; the bundle updater changes both pins.
# Exact bundled model requirements prevent pip from silently backtracking.
RUN pip install "policyengine[models]==5.2.0" spm-calculator==0.3.1 ipython
RUN pip install "policyengine[models]==6.0.0" spm-calculator==1.0.0 ipython
24 changes: 21 additions & 3 deletions docs/canonical-spm.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,20 @@ or read its provenance. SPM-dependent requests validate the required primitives
when the country calculates them. `/calculate-full` and stored household replay
request the full output set, which includes SPM dependencies.

Ordinary `household_benefits` and `household_net_income` count actual
`housing_assistance`; they do not use the SPM housing cap. A request for only
these household outputs can include a positive housing award without supplying
SPM geography. Changing a valid SPM selection does not change those outputs or
create measurement-year receipts. Other programs still use their own required
geographic inputs.

`spm_unit_net_income` is the SPM resource aggregate and retains
`spm_unit_capped_housing_subsidy`. The country allocates household housing awards
to SPM units before applying this cap. Units receiving a positive allocation
need the SPM geography and composition used by the cap. A zero-allocation unit
returns a zero housing resource without evaluating the SPM measurement; this
does not exempt an explicit threshold or poverty calculation from its inputs.

### Choosing a measurement, or not

A measurement is chosen by a request that sends `spm`, or by the household whose
Expand Down Expand Up @@ -118,8 +132,11 @@ inherit the certified defaults. On any other country the same routes reject an
`spm` key at all, null included, with `SPM_SETTINGS_UNSUPPORTED`, as
`POST /{country}/simulation` already did: there is no shape of it to correct.

Tax-only calculations can use periods outside the artifact's measurement years,
including a valid metro selection, without generating SPM receipts. When an SPM
Tax-only and ordinary household-income calculations can use periods outside the
artifact's measurement years, when their own formulas support those periods,
including a valid metro selection, without evaluating SPM amounts or adding
measurement-year receipts. This also holds for ordinary income with a positive
housing award. When an SPM
dependency actually executes for an unsupported year, a request that chose the
measurement returns a structured failure with the calculator's typed
`SPM_YEAR_UNAVAILABLE` code; a request that chose nothing leaves that year's
Expand Down Expand Up @@ -163,7 +180,8 @@ receipt beside its null cells. Read the values to learn which of them a
measurement produced. Provenance comes from the actual simulation and includes artifact,
scenario, years, geography, composition/storage methods and runtime versions.
It remains in JSON form through stored replay and cache hits. A tax-only receipt
may have empty `years` and `geographies` because no SPM measurement was requested.
may have empty `years` and `geographies` because no SPM measurement was requested;
the same is true for an ordinary household-income-only calculation.
Clients saving simulation outputs must retain this full response envelope.

`POST /us/simulation` takes `population_id`, `population_type` (`household` or
Expand Down
4 changes: 2 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -42,10 +42,10 @@ dependencies = [
"policyengine_canada==0.96.3",
"policyengine-ng==0.5.1",
"policyengine-il==0.1.0",
"policyengine[models]==5.2.0",
"policyengine[models]==6.0.0",
# Cloud Run installs with pip, which does not read uv.lock. Keep the
# calculator compatible with the country model in this bundle.
"spm-calculator==0.3.1",
"spm-calculator==1.0.0",
"pydantic",
"pymysql",
"python-dotenv",
Expand Down
15 changes: 15 additions & 0 deletions tests/contract/test_simulation_gateway_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@
MOCK_SIMULATION_PAYLOAD_WITH_TELEMETRY,
MOCK_SUBMIT_RESPONSE_SUCCESS,
)
from tests.fixtures.spm import worker_versions_document


@pytest.fixture(autouse=True)
Expand Down Expand Up @@ -56,6 +57,13 @@ def test_gateway_comparison_submit_and_poll_contract(monkeypatch):
monkeypatch.setenv("OLD_SIMULATION_GATEWAY_URL", "https://simulation.test")
client = _client_for(
{
# A certified bundle reads the worker's advertised SPM
# capability before it submits anything, so the registry read is
# part of the submission contract rather than a separate one.
("GET", "/versions"): _response(
status_code=200,
json_data=worker_versions_document(app_name=MOCK_RESOLVED_APP_NAME),
),
("POST", "/simulate/economy/comparison"): _response(
status_code=202,
json_data=MOCK_SUBMIT_RESPONSE_SUCCESS,
Expand Down Expand Up @@ -88,6 +96,13 @@ def test_gateway_budget_window_submit_and_poll_contract(monkeypatch):
monkeypatch.setenv("OLD_SIMULATION_GATEWAY_URL", "https://simulation.test")
client = _client_for(
{
# A certified bundle reads the worker's advertised SPM
# capability before it submits anything, so the registry read is
# part of the submission contract rather than a separate one.
("GET", "/versions"): _response(
status_code=200,
json_data=worker_versions_document(app_name=MOCK_RESOLVED_APP_NAME),
),
(
"POST",
"/simulate/economy/budget-window",
Expand Down
65 changes: 64 additions & 1 deletion tests/fixtures/services/economy_service.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,14 @@
)
from policyengine_api.data.v1_models import ReformImpact

from tests.fixtures.spm import (
INSTALLED_SPM_SELECTION,
spm_options,
spm_options_hash_segment,
spm_result_fields,
worker_spm_capability,
)

# Mock data constants
MOCK_COUNTRY_ID = "us"
MOCK_POLICY_ID = 123
Expand All @@ -21,10 +29,20 @@
MOCK_TIME_PERIOD = "2025"
MOCK_API_VERSION = "1.0"
MOCK_OPTIONS = {"option1": "value1", "option2": "value2"}
# A certified bundle resolves a measurement into the request options before
# anything is hashed, cached or submitted, so the doubles below carry it
# wherever the service would have written it. The selection itself is read from
# the installed bundle manifest, not written out, because its artifact hash
# belongs to the pinned bundle rather than to these tests.
MOCK_SPM_SELECTION = INSTALLED_SPM_SELECTION
MOCK_RESOLVED_OPTIONS = spm_options(MOCK_OPTIONS)
MOCK_SPM_OPTIONS_HASH_SEGMENT = spm_options_hash_segment()
MOCK_SPM_RESULT_FIELDS = spm_result_fields(years=[MOCK_TIME_PERIOD])
MOCK_DATA_VERSION = "faux-populace-us-2099-test-release"
MOCK_LOOKUP_OPTIONS_HASH = (
"[option1=value1&option2=value2"
"&dataset=hf://policyengine/faux-populace-us/faux_populace_us_2099.h5@"
+ MOCK_SPM_OPTIONS_HASH_SEGMENT
+ "&dataset=hf://policyengine/faux-populace-us/faux_populace_us_2099.h5@"
"faux-populace-us-2099-test-release"
"&model_version=1.2.3&target=general"
"&data_version=faux-populace-us-2099-test-release"
Expand Down Expand Up @@ -57,10 +75,15 @@
"poverty_impact": {"baseline": 0.12, "reform": 0.10},
"budget_impact": {"baseline": 1000, "reform": 1200},
"inequality_impact": {"baseline": 0.45, "reform": 0.42},
# A canonical worker returns what it measured alongside what it computed,
# and the service refuses a result it cannot certify against the selection
# that was submitted.
**MOCK_SPM_RESULT_FIELDS,
}

MOCK_SIM_CONFIG = {
"country": MOCK_COUNTRY_ID,
"spm": MOCK_SPM_SELECTION,
"reform": json.loads(MOCK_REFORM_POLICY_JSON),
"baseline": json.loads(MOCK_BASELINE_POLICY_JSON),
"region": MOCK_REGION,
Expand Down Expand Up @@ -151,6 +174,7 @@ def mock_simulation_entrypoint():
mock_api.get_execution_result.return_value = MOCK_REFORM_IMPACT_DATA
mock_api.run_budget_window_batch.return_value = mock_batch_execution
mock_api.get_budget_window_batch_by_id.return_value = mock_batch_execution
mock_api.get_spm_capability.return_value = worker_spm_capability()

with patch(
"policyengine_api.services.economy_service.simulation_entrypoint", mock_api
Expand Down Expand Up @@ -285,6 +309,44 @@ def create_mock_modal_execution(
return mock_execution


def create_mock_simulation_gateway():
"""A gateway double that certifies the bundle this API actually runs.

An unconfigured `MagicMock` answers `get_spm_capability` with an attribute
rather than a capability, which a certified bundle reads as a worker that
cannot run the measurement it resolved.
"""
gateway = MagicMock()
gateway.get_spm_capability.return_value = worker_spm_capability()
return gateway


def create_mock_budget_window_annual_impact(year, **fields):
"""One year of a budget-window worker result, with its own receipt."""
return {"year": str(year), **fields, **spm_result_fields(years=[year])}


def create_mock_budget_window_result(years, totals=None, **row_fields):
"""A complete budget-window worker result for exactly the given years.

A canonical worker returns one certified receipt per year and the API
refuses a window whose receipts do not cover the years it submitted, so a
partial `annualImpacts` list is a broken result rather than a small one.
"""
years = [str(year) for year in years]
return {
"kind": "budgetWindow",
"startYear": years[0],
"endYear": years[-1],
"windowSize": len(years),
"annualImpacts": [
create_mock_budget_window_annual_impact(year, **row_fields)
for year in years
],
"totals": {} if totals is None else totals,
}


def create_mock_budget_window_batch_execution(
batch_job_id=MOCK_MODAL_JOB_ID,
status=MODAL_EXECUTION_STATUS_SUBMITTED,
Expand Down Expand Up @@ -329,6 +391,7 @@ def mock_simulation_entrypoint_legacy():
mock_api.get_execution_by_id.return_value = mock_execution
mock_api.get_execution_status.return_value = MODAL_EXECUTION_STATUS_RUNNING
mock_api.get_execution_result.return_value = MOCK_REFORM_IMPACT_DATA
mock_api.get_spm_capability.return_value = worker_spm_capability()

with patch(
"policyengine_api.services.economy_service.simulation_entrypoint", mock_api
Expand Down
154 changes: 154 additions & 0 deletions tests/fixtures/spm.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,154 @@
"""Doubles for the canonical SPM boundary a certified bundle activates.

A certified bundle resolves an SPM selection for every US request. The API then
refuses to submit that request unless the selected worker advertises a matching
capability, and refuses to serve a result that carries no receipt for what was
measured. Doubles written while the pinned bundle was uncertified answer
neither, because the boundary was dormant and never asked them.

Every value here is read from the installed bundle manifest rather than written
down, so a double follows the pinned bundle instead of becoming a second copy
of it that the next bump would silently contradict. On a bundle whose US model
predates the canonical constructor the selection is `None`, the boundary stays
dormant, and every helper degrades to the legacy shapes.

Resolving the selection at import time also imports `policyengine_us` during
collection, before any test patches `datetime.datetime`, so the country import
inside the boundary is a `sys.modules` hit rather than a fresh import running
under a patched standard library.
"""

from __future__ import annotations

import pytest

from policyengine_api import spm
from policyengine_api.constants import POLICYENGINE_VERSION

# The contract identifier `worker_spm.validate_worker_spm` requires.
SPM_CONTRACT_VERSION = "canonical-spm-v1"

# What the installed bundle certifies for an unqualified US request, or None on
# a bundle whose model predates the contract.
INSTALLED_SPM_SELECTION = spm.normalize_spm_selection("us", None)


def worker_spm_capability(selection: dict | None = None) -> dict | None:
"""What a worker running this API's own bundle advertises to `/versions`."""
resolved = INSTALLED_SPM_SELECTION if selection is None else selection
if resolved is None:
return None
return {"contract_version": SPM_CONTRACT_VERSION, "defaults": dict(resolved)}


def worker_versions_document(
*,
bundle_version: str = POLICYENGINE_VERSION,
app_name: str = "test-worker-app",
country: str = "us",
country_version: str | None = None,
selection: dict | None = None,
) -> dict:
"""The gateway registry document `get_spm_capability` reads.

Keyed by wrapper version, as the deployed registry is: the capability
belongs to the bundle the worker runs, not to the country route that
resolves to its application.
"""
document: dict = {
"policyengine": {bundle_version: app_name, "latest": bundle_version},
}
if country_version is not None:
document[country] = {country_version: app_name, "latest": country_version}
capability = worker_spm_capability(selection)
if capability is not None:
document["spm_capabilities"] = {bundle_version: capability}
return document


def spm_receipt(*, years, selection: dict | None = None) -> dict:
"""One country provenance receipt covering the given calculation years."""
resolved = INSTALLED_SPM_SELECTION if selection is None else selection
return {
"forecast_id": "test-only",
"forecast_sha256": resolved["forecast_content_sha256"],
"scenario": resolved["scenario"],
"geography_kind": resolved["geography_kind"],
"runtime_versions": {"policyengine-us": "test-only"},
"years": {str(year): {"status": "forecast"} for year in years},
"geographies": [],
"composition_method": "classified-inputs",
"storage_method": "formula",
}


def spm_result_fields(*, years, selection: dict | None = None) -> dict:
"""The SPM half of a worker result, or nothing on an uncertified bundle."""
resolved = INSTALLED_SPM_SELECTION if selection is None else selection
if resolved is None:
return {}
receipt = spm_receipt(years=years, selection=resolved)
return {
"spm_config": dict(resolved),
"spm_provenance": {
"baseline": [dict(receipt)],
"reform": [dict(receipt)],
},
}


def household_receipt_fields(*, years, selection: dict | None = None) -> dict:
"""The SPM half of a cached household calculation.

A household carries one receipt for the calculation it is, where a worker
result carries a baseline and reform list for the comparison it ran.
"""
resolved = INSTALLED_SPM_SELECTION if selection is None else selection
if resolved is None:
return {}
return {
"spm_config": dict(resolved),
"spm_provenance": spm_receipt(years=years, selection=resolved),
}


def country_calculate_kwargs(
*, requested: bool = False, selection: dict | None = None
) -> dict:
"""The measurement arguments the service passes to a country package."""
resolved = INSTALLED_SPM_SELECTION if selection is None else selection
if resolved is None:
return {}
return {"spm": dict(resolved), "spm_requested": requested}


def spm_options(options: dict, selection: dict | None = None) -> dict:
"""Request options as the service resolves them against the bundle."""
resolved = INSTALLED_SPM_SELECTION if selection is None else selection
if resolved is None:
return dict(options)
return {**options, "spm": dict(resolved)}


def spm_options_hash_segment(selection: dict | None = None) -> str:
"""The segment a resolved selection contributes to a cache identity.

Options are serialized in sorted-key order, so `spm` follows every option a
caller sent. Derived rather than written out because the artifact hash it
carries belongs to the pinned bundle.
"""
resolved = INSTALLED_SPM_SELECTION if selection is None else selection
return "" if resolved is None else f"&spm={resolved}"


@pytest.fixture
def legacy_bundle(monkeypatch):
"""Pin a bundle that predates the canonical SPM contract.

Use this in tests that assert legacy behaviour by name: the resolution and
the transport contract both differ on an uncertified bundle, so leaving the
regime to whatever happens to be installed makes the assertion mean
whichever thing the environment chose.
"""
monkeypatch.setattr(spm, "_current_bundle", dict)
monkeypatch.setattr(spm, "_installed_country_implements_spm", lambda _: False)
6 changes: 6 additions & 0 deletions tests/integration/test_budget_window_in_flight_dedupe.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,8 @@
from flask import Flask
from policyengine_api.runtime_cache.fake import InMemoryCacheBackend

from tests.fixtures.spm import worker_spm_capability


class FakeRedis(InMemoryCacheBackend):
pass
Expand Down Expand Up @@ -31,6 +33,10 @@ def test_budget_window_in_flight_dedupe_uses_existing_batch_without_live_db(

fake_cache = BudgetWindowCache(client=FakeRedis())
simulation_entrypoint = MagicMock()
# A certified bundle asks the selected worker to certify the measurement
# before the service reaches its cache, so the gateway double answers the
# capability its own bundle binds instead of an unread attribute.
simulation_entrypoint.get_spm_capability.return_value = worker_spm_capability()
simulation_entrypoint.resolve_app_name.side_effect = (
lambda country, version, **kwargs: ("test-budget-window-worker", version)
)
Expand Down
Loading
Loading