Skip to content

Fix stack_neuron_results for an integer neuron_slice - #1844

Merged
jlarson4 merged 2 commits into
TransformerLensOrg:devfrom
Zhuoxi2000:fix-stack-neuron-results-int-slice
Oct 2, 2026
Merged

jlarson4 merged 2 commits into
TransformerLensOrg:devfrom
Zhuoxi2000:fix-stack-neuron-results-int-slice

Conversation

@Zhuoxi2000

@Zhuoxi2000 Zhuoxi2000 commented Oct 1, 2026 •

Copy link
Copy Markdown

Base branch: dev

Description

ActivationCache.stack_neuron_results(layer, neuron_slice=n) with an int n (a valid SliceInput) wrapped the input with Slice(n), which collapses the neuron axis. The visible symptom was TypeError: iteration over a 0-d tensor while building the labels. That crash hid two more bugs further down:

  • non-LN path: the per-layer results lose their neuron axis, so the layer concat runs along positions and the stack comes out as (8, 3, 16) instead of (2, 3, 4, 16). With apply_ln or incl_remainder it then fails with a broadcast RuntimeError.
  • LN-folded projection path (apply_ln=True with project_output_onto): IndexError on LN models and a broadcast RuntimeError on RMS models.

Fix: normalize with Slice.unwrap(neuron_slice), which turns an int into [n]. The neuron axis is then kept on every path (raw, apply_ln, project_output_onto, the folded LN+projection path, and incl_remainder), and the result matches neuron_slice=[n]. This also makes the isinstance(neuron_labels, int) guard dead, so it is removed together with the now-unused numpy import. Slice.unwrap returns an existing Slice unchanged, so an integer-mode Slice(n) is also rebuilt as a new Slice([n]) inside stack_neuron_results. The caller's object is not modified, and Slice / get_neuron_results keep their dimension-collapsing int semantics. The neuron_slice docstring says that an int or an integer-mode Slice(n) is treated like [n].

Fixes #1841

Type of change

  • Bug fix (non-breaking change which fixes an issue)

Test evidence

New tests in tests/unit/test_activation_cache.py:

  • test_stack_neuron_results_integer_neuron_slice_matches_list runs stack_neuron_results(n_layers, ...) with neuron_slice=3 and with neuron_slice=Slice(3), and compares each against neuron_slice=[3]: the labels are equal, the shapes are equal, and torch.testing.assert_close passes on the values. It also checks that the stack has one row per layer (plus the remainder). The matrix is:

    • neuron_slice form: bare int 3 / Slice(3)
    • apply_ln: off / on
    • project_output_onto: none / [d_model] vector / [d_model, 2] matrix
    • incl_remainder: off / on
    • the existing module fixture's LN and RMS models (2 layers, d_model=16, built natively, no download)

    That gives 48 cases.

  • test_stack_neuron_results_integer_neuron_slice_without_labels does the same shape and value comparison with return_labels=False, for 3 and Slice(3) on both models (4 cases), because the labels were built even when they were not returned.

Source New tests (52) Failure
dev @ cfac4be 52 failed all 52: TypeError: iteration over a 0-d tensor
dev @ 02a7f5a + label guard patched only (0-d tensor reshaped, Slice(n) kept), bare-int matrix only 24 failed 6: shape assert, e.g. (8, 3, 16) != (2, 3, 4, 16); 4: IndexError (LN, folded projection); 14: broadcast RuntimeError (remainder subtraction, apply_ln_to_stack, RMS folded projection)
first commit of this PR (bare int only) 26 failed all 26 Slice(3) cases: TypeError: iteration over a 0-d tensor
this PR 52 passed none

The label-guard row comes from a throwaway patch and is not part of this PR. It shows that the test checks the result against [n] and does not only check that the call no longer raises.

Whole file:

$ uv run pytest tests/unit/test_activation_cache.py -q
dev @ cfac4be + new tests:  52 failed, 56 passed
this PR:                    108 passed

uv run pytest transformer_lens/ActivationCache.py transformer_lens/utilities/slice.py (doctests): 5 passed.

Lint, using the versions from uv.lock (black 23.12.1, isort 5.8.0, pycln 2.5.0, mypy 1.17.0): the make check-format commands (pycln / isort / black, repo-wide) are clean, and uv run mypy . reports no issues in 387 source files.

Checklist:

  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • I have not rewritten tests relating to key interfaces which would affect backward compatibility

AI assistance: this change was drafted with an AI coding assistant (Claude) and verified locally with the tests above.

stack_neuron_results wrapped neuron_slice with Slice(), so an int (a valid
SliceInput) collapsed the neuron axis. Building the labels then raised
"TypeError: iteration over a 0-d tensor", and that crash hid two more bugs:
the non-LN path concatenated layers along positions (e.g. (8, 3, 16) instead
of (2, 3, 4, 16)), and the LN-folded projection path raised IndexError (LN)
or a broadcast RuntimeError (RMS).

Normalise with Slice.unwrap() so an int keeps the neuron axis like [n] on
every path. This makes the isinstance(neuron_labels, int) guard dead, so it
is removed along with the now-unused numpy import.

Add a regression test comparing neuron_slice=n against neuron_slice=[n]
(labels, shape and values) for apply_ln on/off, no/vector/matrix
project_output_onto, and incl_remainder on/off, on LN and RMS models.

Fixes TransformerLensOrg#1841
neuron_slice = Slice(neuron_slice)
# unwrap turns an int into [n]: layers are concatenated along the neuron axis, so it
# must not collapse.
neuron_slice = Slice.unwrap(neuron_slice)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for addressing this follow-up to #1836/#1837 and adding coverage for the different normalization and projection paths!

I noticed one remaining edge case: the signature accepts neuron_slice=Slice(n), but Slice.unwrap leaves existing Slice instances unchanged. Since Slice(n) remains in integer mode, applying it still collapses the neuron axis. This makes neuron_labels a 0-D tensor, so the unconditional label comprehension raises the original TypeError even when return_labels=False.

Could we normalize integer-mode Slice inputs locally in stack_neuron_results and extend the regression test to verify that Slice(neuron) behaves the same as [neuron]? Keeping the normalization local would preserve the dimension-collapsing semantics of Slice/get_neuron_results elsewhere.

This looks like an existing uncovered edge case rather than a regression introduced by this PR. My review is based on static code inspection only; I haven’t run the tests locally.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, good catch. Slice.unwrap passes an existing Slice through unchanged, so Slice(n) still collapsed the neuron axis and the label comprehension failed even with return_labels=False.

Fixed in a follow-up commit. After unwrap, stack_neuron_results now rebuilds an integer-mode Slice as a new Slice([n]). The normalization stays local: the caller's object is not modified, and Slice / get_neuron_results keep their dimension-collapsing int semantics.

Tests: 3 and Slice(3) are now both compared against [3] (labels, shape, values) over the same matrix, plus a return_labels=False case; the 26 new Slice(3) cases fail on the previous commit with TypeError: iteration over a 0-d tensor and pass now, and the whole tests/unit/test_activation_cache.py passes (108).

AI assistance: drafted with an AI coding assistant (Claude).

Slice.unwrap() only converts a bare int into [n]; an existing Slice is
returned unchanged. So neuron_slice=Slice(n) stayed in integer mode,
collapsed the neuron axis, and building the labels raised
"TypeError: iteration over a 0-d tensor", also with return_labels=False.

Rebuild an integer-mode Slice as a new Slice([n]) inside
stack_neuron_results only. The caller's Slice is not modified, and Slice /
get_neuron_results keep their dimension-collapsing int semantics.

Extend the regression test so Slice(n) is compared against [n] (labels,
shape, values) over the same matrix as the bare int, and add a
return_labels=False case for both forms.

@yuanwuyuan9 yuanwuyuan9 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the update! I’ve reviewed the follow-up change, and it addresses the Slice(n) edge case I raised.
Rebuilding the slice locally preserves the caller’s object and the existing semantics elsewhere, and the added coverage includes return_labels=False.
No further concerns from my static review; I haven’t run the tests locally.

@jlarson4

jlarson4 commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Looks good to me as well! Merging as is! Thanks @Zhuoxi2000

@jlarson4
jlarson4 merged commit 0c254e6 into TransformerLensOrg:dev Oct 2, 2026
27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants