Skip to content

fix: honor vLLM logits processor transforms in reconstruction - #1853

Open
emerardd wants to merge 1 commit into
TransformerLensOrg:devfrom
emerardd:fix/vllm-logits-output-transform
Open

emerardd wants to merge 1 commit into
TransformerLensOrg:devfrom
emerardd:fix/vllm-logits-output-transform

Conversation

@emerardd

@emerardd emerardd commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Description

Fix vLLM full-sequence logit reconstruction for architectures whose logits processor applies an output scale. The current driver reconstructs hidden @ lm_head.weight.T + bias and a config-derived Gemma softcap, but omits the processor's scale. For Cohere/Command-R this drops config.logit_scale; for Granite it drops the effective logit_scale / logits_scaling. Probabilities and causal loss are therefore wrong even when a positive scale leaves greedy argmax unchanged.

The worker now exposes the actual loaded model's logits processor scale and softcap over a small, plain-data RPC. The reconstruction probe caches those values alongside the unembedding, tolerates absent processors on non-final pipeline stages, verifies agreement across responding ranks, and applies softcap followed by scale after the matmul and bias. This follows the vLLM 0.20.2 processor contract, without architecture-name special cases or assuming local compatibility-fold state applies to remote weights. The architecture setup is visible in Command-R and Granite.

No additional RPC runs during cached reconstruction. When no rank exposes a processor, reconstruction is disabled with an explicit warning instead of guessing an identity transform; the existing last-token fallback cannot claim full-sequence loss support. Unknown/logits-as-input processors and inconsistent or invalid metadata fail loudly. Cache cleanup also releases the transform metadata.

Regression evidence

Offline integration tests construct real tiny HF Cohere and Granite models, capture the real final norm, and exercise actual worker parameter reads plus the actual reconstruction method through a synchronous RPC scaffold. Processor attributes in this scaffold mirror the documented vLLM contract; it is not a live vLLM engine.

On the unchanged base, the non-unit Cohere and Granite scaling cases failed while identity-scale controls passed. The Cohere maximum error was 0.3270536959; the Granite maximum error was 0.0767858624. After the fix, reconstructed logits, probabilities, and shifted causal losses match the HF controls. Unit tests additionally cover bias, scale/softcap ordering, cached reads, batched shapes, missing PP stages, replica disagreement, malformed metadata, and scaled logits/loss reaching RemoteBridge.forward(return_type="both") through both single and batched dispatch.

Type of change

  • Bug fix (non-breaking change which fixes an issue)

Validation

  • Affected suite: 251 passed, 18 skipped, 2 warnings. Files: test_vllm_logit_scaling.py, test_vllm_driver.py, test_vllm_worker_extension.py, test_vllm_boot.py, test_vllm_plugin.py, test_vllm_internals.py, and test_inspect_vllm_provider.py.
  • The skips are existing optional-vLLM tests on this Windows CPU environment; no new skips/xfails were added.
  • Changed-file Makefile-equivalent pycln/isort/Black checks passed.
  • Full mypy . passed: 387 source files.
  • git diff --check passed.

The complete test suite was not run locally. No live vLLM/CUDA, compiled graph, tensor-parallel, or pipeline-parallel execution was performed; real engine parity remains a validation boundary for review. The two reported warnings are existing SWIG deprecations. The new missing-processor warning is intentional and tested with pytest.warns.

Checklist

  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • I have not rewritten tests relating to key interfaces which would affect backward compatibility

Documentation changes are the reconstruction/probe/RPC docstrings. The warning checkbox is unchecked because the missing-processor diagnostic is intentional. The unit-test checkbox is unchecked because only the affected surface was run, not the complete unit suite.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant