Conversation
jlarson4
reviewed
Oct 5, 2026
| return bool(identity_ok and causal_ok) | ||
|
|
||
|
|
||
| def _accepts_mask_positions(model: torch.nn.Module) -> bool: |
Collaborator
There was a problem hiding this comment.
This mirrors TransformerBridge._accepts_derived_position_ids, but it has already drifted. The local gate accepts a forward(**kwargs) model and this one doesn't, so such a model would get different positions remotely vs locally. Can both call sites use a shared helper in order to stay in sync and reuse the existing gate tests?
| profile = profiles.TLBridgeProfile( | ||
| supported_kinds=kinds, | ||
| provides_sequence_logits=psl, | ||
| supports_attention_mask=provider == "tl_bridge", |
Collaborator
There was a problem hiding this comment.
The other capabilities in this block are read off the provider API, but mask support is keyed on the provider name here and again in profiles.for_provider. Would it be possible for the provider declare it the way it declares provides_sequence_logits?
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes #1851.
Inspect RemoteBridge accepted
attention_mask, but the Inspect driver omitted it from the request and the HF provider ran an unmasked forward. Masking only the final loss could therefore score logits computed from the wrong context, and cached activations were affected as well.This change validates and serializes single-sequence binary padding masks, forwards them to the HF-backed
tl_bridgeprovider, and derives padding-aware position IDs only where the model supports ordinary 2-D positions. Model-owned position derivation (including OPT's mask-consuming embeddings and mRoPE) is left intact. No-mask and all-ones-mask behavior is preserved. The provider's completion and logprobs use the last attended token for right-padded inputs.The Inspect
tl_bridge_vllmandvllm-lensprofiles reject supplied masks explicitly instead of silently ignoring them; the directtl_bridge_vllmcapture entry point also rejects masks. This does not change the separate native vLLM driver's padding support. Theboot_inspectAPI documentation now states the mask contract and provider limits.The offline integration regression boots the real Inspect model envelope, driver, and HF provider from locally saved tiny GPT-2, Llama, and OPT models. It covers left/right padding, raw HF parity, unpadded-control logits and residual caches, masked implicit/explicit-label loss, attention patterns and interventions, all-ones masks, binary mask dtypes, malformed masks at both public/provider boundaries, and last-attended-token completion/logprobs. No model downloads or mocked model loads are needed.
Type of change
Validation
tests/integration/model_bridge/test_inspect_attention_mask.py,tests/unit/model_bridge/test_inspect_driver.py,tests/unit/model_bridge/test_inspect_vllm_provider.py, andtests/unit/model_bridge/sources/test_inspect_provider_model_class.py.mypy .passed: 388 source files.git diff --checkpassed.Validation used the frozen lockfile with the
inspectextra. The complete unit/integration/acceptance suite was not run locally; CI will cover the broader surface. The selected run emitted existing SWIG deprecation warnings and the expected GPT-2 fused-QKV capability warning. No live vLLM/GPU run was performed; the new unsupported-mask rejection is tested before a vLLM import or engine call.Checklist
The unit-test checkbox is left unchecked because only the affected unit-test surface was run, not the complete unit suite.