Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fix vLLM full-sequence logit reconstruction for architectures whose logits processor applies an output scale. The current driver reconstructs
hidden @ lm_head.weight.T + biasand a config-derived Gemma softcap, but omits the processor's scale. For Cohere/Command-R this dropsconfig.logit_scale; for Granite it drops the effectivelogit_scale / logits_scaling. Probabilities and causal loss are therefore wrong even when a positive scale leaves greedy argmax unchanged.The worker now exposes the actual loaded model's logits processor scale and softcap over a small, plain-data RPC. The reconstruction probe caches those values alongside the unembedding, tolerates absent processors on non-final pipeline stages, verifies agreement across responding ranks, and applies softcap followed by scale after the matmul and bias. This follows the vLLM 0.20.2 processor contract, without architecture-name special cases or assuming local compatibility-fold state applies to remote weights. The architecture setup is visible in Command-R and Granite.
No additional RPC runs during cached reconstruction. When no rank exposes a processor, reconstruction is disabled with an explicit warning instead of guessing an identity transform; the existing last-token fallback cannot claim full-sequence loss support. Unknown/logits-as-input processors and inconsistent or invalid metadata fail loudly. Cache cleanup also releases the transform metadata.
Regression evidence
Offline integration tests construct real tiny HF Cohere and Granite models, capture the real final norm, and exercise actual worker parameter reads plus the actual reconstruction method through a synchronous RPC scaffold. Processor attributes in this scaffold mirror the documented vLLM contract; it is not a live vLLM engine.
On the unchanged base, the non-unit Cohere and Granite scaling cases failed while identity-scale controls passed. The Cohere maximum error was
0.3270536959; the Granite maximum error was0.0767858624. After the fix, reconstructed logits, probabilities, and shifted causal losses match the HF controls. Unit tests additionally cover bias, scale/softcap ordering, cached reads, batched shapes, missing PP stages, replica disagreement, malformed metadata, and scaled logits/loss reachingRemoteBridge.forward(return_type="both")through both single and batched dispatch.Type of change
Validation
test_vllm_logit_scaling.py,test_vllm_driver.py,test_vllm_worker_extension.py,test_vllm_boot.py,test_vllm_plugin.py,test_vllm_internals.py, andtest_inspect_vllm_provider.py.mypy .passed: 387 source files.git diff --checkpassed.The complete test suite was not run locally. No live vLLM/CUDA, compiled graph, tensor-parallel, or pipeline-parallel execution was performed; real engine parity remains a validation boundary for review. The two reported warnings are existing SWIG deprecations. The new missing-processor warning is intentional and tested with
pytest.warns.Checklist
Documentation changes are the reconstruction/probe/RPC docstrings. The warning checkbox is unchecked because the missing-processor diagnostic is intentional. The unit-test checkbox is unchecked because only the affected surface was run, not the complete unit suite.