Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,11 @@

All notable changes to this project are documented in this file.

## Unreleased

### Added
- **`validate-kgx` can repair a graph in place with `--prune` (`-p`).** The flag rewrites the edges NDJSON in place, keeping only the edges that validate strictly: real defects (non-coercible values, extras the model will never declare, rejected predicates), the deliberate pending carryovers (`synonym`/`xref`/`relation`/`provided_by`, ...), and lines that are not JSON at all are all removed, so the final graph validates with zero failures (on a real graph the carryovers are most of the edges, so the flag trades them for compliance). The rewrite is atomic and kept lines stay byte-identical, a clean file is not rewritten at all, a missing edges file is reported without ever being pruned, nodes are never touched, and the exit code goes green as soon as the nodes are clean too. The core lives in the new public `biolink.prune_kgx_edges`, which reuses `validate_kgx`'s own strict bar, so a prune can never disagree with what a plain validation run reports.

## 19.5.1 - 2026-10-01

### Changed
Expand Down
30 changes: 30 additions & 0 deletions docs/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -400,6 +400,7 @@ tablassert validate-kgx --nodes MY_KG_1.0.0.nodes.ndjson --edges MY_KG_1.0.0.edg
| `--nodes`, `-n` | Path | Yes | n/a | Built `*.nodes.ndjson` file to validate |
| `--edges`, `-e` | Path | Yes | n/a | Built `*.edges.ndjson` file to validate |
| `--limit` | int | No | `20` | Maximum example failures to retain per file |
| `--prune`, `-p` | Flag | No | `False` | Rewrite the edges file in place, keeping only edges that validate strictly (pending carryovers are dropped too); nodes are never touched, and a clean file is not rewritten |

Failures are grouped by field and error type, so a systematic modelling problem shows up as one line
rather than a million:
Expand Down Expand Up @@ -429,6 +430,35 @@ edges: 1200000/2000085 valid (800085 failures; 800085 pending biolink-model supp
The strict count is what `ok` and the exit code use; the pending count is what
[`tablassert agent`](#agent) optimizes against, so a deliberate gap never reads as a modelling error.

### Repairing a graph with `--prune`

`--prune` (`-p`) repairs an already-emitted graph instead of just scoring it: the edges file is
rewritten in place and every edge that does not validate strictly is removed. That includes real
defects (non-coercible values, extras the model will never declare, rejected predicates), the
deliberate pending carryovers (`synonym`, `xref`, `relation`, `provided_by`, ...), and lines that
are not JSON at all (counted separately as malformed). On a real graph the carryovers are most of
the edges -- a DAKP-style build prunes roughly 800085 of 2000085 -- so `--prune` trades them for a
strictly-compliant graph. Nodes are validated and reported, never pruned: dropping one would
strand its references.

```bash
tablassert validate-kgx -n MY_KG_1.0.0.nodes.ndjson -e MY_KG_1.0.0.edges.ndjson --prune
```

The rewrite is atomic (written to a sibling temp file, then swapped in) and kept lines stay
byte-identical, so a pruned graph differs from the original only by the removed records. A clean
file is not rewritten at all, and a missing edges file is reported without ever being pruned.
After pruning, every kept edge is strictly valid, so the exit code goes green as soon as the
nodes are clean too:

```text
edges: pruned 12/424159 non-compliant records
9 p_value: float_parsing
e.g. DISEASE_EDGE_000017: p_value: float_parsing
edges: 424147/424159 valid (0 failures)
KGX output is Biolink-compliant.
```

Retrieval-source entries are the mirror case. Tablassert emits `resource_id` as each `sources` entry's
sole identifier, while the pinned model still requires the inherited `Entity.id` on `RetrievalSource`.
The validator supplies that `id` to its own in-memory copy of the record whenever the installed model
Expand Down
228 changes: 193 additions & 35 deletions src/tablassert/biolink.py
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@

import inspect
import json
import os
import re
from collections import Counter
from enum import Enum
Expand Down Expand Up @@ -86,6 +87,7 @@
"legal_predicates",
"node_class",
"numeric_slot_kind",
"prune_kgx_edges",
"resolve_association_class",
"resolve_node_category",
"resolve_node_class",
Expand Down Expand Up @@ -944,45 +946,201 @@ def validate_kgx(nodes_path: Path, edges_path: Path, limit: int = 20) -> dict[st
"""
report: dict[str, Any] = {"biolink_version": BIOLINK_VERSION, "ok": True, "ok_excluding_pending": True}
for label, path, edge in (("nodes", nodes_path, False), ("edges", edges_path, True)):
total: int = 0
valid: int = 0
valid_excluding_pending: int = 0
problems: Counter[str] = Counter()
examples: list[dict[str, Any]] = []
# A missing path must never read as a clean bill of health: counting zero records
# out of zero would otherwise exit 0 on a typo'd filename and hide a broken build.
missing: bool = not path.is_file()
if not missing:
with path.open(encoding="utf-8") as handle:
for line in handle:
section: dict[str, Any] = _validate_file(path, edge=edge, limit=limit)
report[label] = section
if section["missing"] or section["total"] != section["valid"]:
report["ok"] = False
if section["missing"] or section["total"] != section["valid_excluding_pending"]:
report["ok_excluding_pending"] = False
return report


def _validate_file(path: Path, *, edge: bool, limit: int) -> dict[str, Any]:
"""Validate one KGX NDJSON file record by record (the per-file half of ``validate_kgx``).

Args:
path: Path to the ``.ndjson`` file; a missing path reports ``missing`` instead of
counting zero records out of zero as a pass.
edge: ``True`` to dispatch on the association family, ``False`` for nodes.
limit: Maximum number of example failures to retain.

Returns:
The same per-file section ``validate_kgx`` nests under ``"nodes"`` / ``"edges"``.
"""
total: int = 0
valid: int = 0
valid_excluding_pending: int = 0
problems: Counter[str] = Counter()
examples: list[dict[str, Any]] = []
# A missing path must never read as a clean bill of health: counting zero records
# out of zero would otherwise exit 0 on a typo'd filename and hide a broken build.
missing: bool = not path.is_file()
if not missing:
with path.open(encoding="utf-8") as handle:
for line in handle:
if not line.strip():
continue
total += 1
record: dict[str, Any] = json.loads(line)
errors: list[str] = validate_record(record, edge=edge)
if not errors:
valid += 1
valid_excluding_pending += 1
continue
if _record_has_only_pending_problems(record, errors):
valid_excluding_pending += 1
problems.update(errors)
if len(examples) < limit:
examples.append({"id": record.get("id"), "errors": errors})
return {
"total": total,
"valid": valid,
"valid_excluding_pending": valid_excluding_pending,
"failures": total - valid,
"missing": missing,
"problems": dict(problems.most_common()),
"examples": examples,
}


def prune_kgx_edges(edges_path: Path, limit: int = 20) -> dict[str, Any]:
"""Drop every edge record that is not strictly KGX/Biolink-compliant, in place.

A record survives only when it validates against the Biolink Pydantic class named
by its own ``category`` with zero errors -- the same strict bar a plain
``validate_kgx`` run uses for its ``ok`` verdict, so a prune can never disagree
with what validation reports:

* a strictly valid record is kept;
* a record whose every failure is a known, intentional model gap (the KGX
denormalized carryovers in :data:`KNOWN_PENDING_EDGE_FIELDS`, or a
:data:`PREDICATE_OVERRIDES` predicate the pinned classes reject) is dropped --
these carryovers are most of a real graph, so the caller asked for a
strictly-compliant graph over maximal data retention when invoking this;
* every other failing record is a real defect and is dropped;
* a line that is not valid JSON at all is dropped and counted separately in
``malformed``.

Nodes are never touched: dropping one would strand references, and cascading that
into the edges is a modelling decision, not a validation repair.

The rewrite is ATOMIC and lossless for kept records: kept lines are written verbatim
(byte-identical, order preserved, blank lines passed through) into a sibling
``.{name}.tmp`` and swapped in with :func:`os.replace` only once every line was
written, so a crash never leaves a torn edges file. When nothing is prunable the
destination is not rewritten at all (``rewritten`` is ``False`` and the bytes are
untouched), and a missing file is reported (``missing``) without ever being created
-- a typo'd path must never read as a prune.

Args:
edges_path: Path to ``<name>_<version>.edges.ndjson``.
limit: Maximum number of example failures to retain per report section.

Returns:
Mapping with ``before`` (the same per-file section ``validate_kgx`` builds over
the original file), ``after`` (the same shape over the kept records, i.e. the
file's post-prune state, equal to what ``validate_kgx`` reports on the rewritten
file), ``dropped`` (the ``pruned`` records' problem histogram and examples --
the records that were removed), ``pruned`` / ``kept`` / ``malformed`` /
``rewritten`` counts, and ``ok`` / ``ok_excluding_pending`` computed over the
post-prune state. For a readable file ``ok`` is therefore always true -- every
kept record is strictly valid -- and ``ok_excluding_pending`` agrees, because
nothing pending survives the prune.

Raises:
OSError: the destination's directory is unwritable; the error propagates and the
original file keeps its previous content.
"""
total: int = 0
valid: int = 0
pending: int = 0 # pending-only failures: counted for the BEFORE report, never kept
pruned: int = 0
malformed: int = 0
all_problems: Counter[str] = Counter()
all_examples: list[dict[str, Any]] = []
pruned_problems: Counter[str] = Counter()
pruned_examples: list[dict[str, Any]] = []
# A missing path must never read as a clean bill of health: reporting zero records
# out of zero would otherwise exit 0 on a typo'd filename and hide a broken build.
missing: bool = not edges_path.is_file()
if not missing:
tmp: Path = edges_path.with_name(f".{edges_path.name}.tmp")
try:
# newline="" on both ends keeps every kept line byte-identical: the default
# universal-newlines mode would silently rewrite "\r\n" as "\n".
with edges_path.open(encoding="utf-8", newline="") as src, tmp.open("w", encoding="utf-8", newline="") as dst:
for line in src:
if not line.strip():
dst.write(line) # blank lines are not records; pass through verbatim
continue
total += 1
record: dict[str, Any] = json.loads(line)
errors: list[str] = validate_record(record, edge=edge)
if not errors:
valid += 1
valid_excluding_pending += 1
try:
record: dict[str, Any] = json.loads(line)
except json.JSONDecodeError:
malformed += 1
pruned += 1
pruned_problems["record: json_parse_error"] += 1
if len(pruned_examples) < limit:
pruned_examples.append({"id": None, "errors": ["record: json_parse_error"]})
continue
if _record_has_only_pending_problems(record, errors):
valid_excluding_pending += 1
problems.update(errors)
if len(examples) < limit:
examples.append({"id": record.get("id"), "errors": errors})
report[label] = {
"total": total,
"valid": valid,
"valid_excluding_pending": valid_excluding_pending,
"failures": total - valid,
"missing": missing,
"problems": dict(problems.most_common()),
"examples": examples,
}
if missing or total != valid:
report["ok"] = False
if missing or total != valid_excluding_pending:
report["ok_excluding_pending"] = False
return report
errors: list[str] = validate_record(record, edge=True)
all_problems.update(errors)
if len(all_examples) < limit:
all_examples.append({"id": record.get("id"), "errors": errors})
if errors:
# The before-section keeps validate_kgx's pending score so the
# pre-report matches a plain validation run; the STRICT bar still
# drops the record -- real defects AND deliberate pending gaps.
if _record_has_only_pending_problems(record, errors):
pending += 1
pruned += 1
pruned_problems.update(errors)
if len(pruned_examples) < limit:
pruned_examples.append({"id": record.get("id"), "errors": errors})
continue
valid += 1
dst.write(line)
# ATOMIC: the swap happens only after every kept line was written, so a crash
# mid-write leaves the original file untouched and the temp never read back.
rewritten: bool = pruned > 0
if rewritten:
os.replace(tmp, edges_path)
finally:
tmp.unlink(missing_ok=True) # no-op after a successful replace; cleanup after a raise
else:
rewritten = False
kept: int = total - pruned
before: dict[str, Any] = {
"total": total,
"valid": valid,
"valid_excluding_pending": valid + pending,
"failures": total - valid,
"missing": missing,
"problems": dict(all_problems.most_common()),
"examples": all_examples,
}
dropped: dict[str, Any] = {"total": pruned, "problems": dict(pruned_problems.most_common()), "examples": pruned_examples}
after: dict[str, Any] = {
"total": kept,
"valid": valid,
"valid_excluding_pending": valid, # nothing pending survives, so the scores agree
"failures": kept - valid,
"missing": missing,
"problems": {},
"examples": [],
}
return {
"biolink_version": BIOLINK_VERSION,
"before": before,
"after": after,
"dropped": dropped,
"pruned": pruned,
"kept": kept,
"malformed": malformed,
"rewritten": rewritten,
"ok": not missing and kept == valid,
"ok_excluding_pending": not missing and kept == valid,
}


# Node slots Tablassert can always populate from a fullmap hit. A class requiring
Expand Down
Loading
Loading