Skip to content

Laguna: bridge and management port fixes - #1692

Open
troglobit wants to merge 13 commits into
mainfrom
laguna-bug
Open

troglobit wants to merge 13 commits into
mainfrom
laguna-bug

Conversation

@troglobit

@troglobit troglobit commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Description

Fixes and debug aids for the Laguna (LAN969x) boards, found while bringing the Tactical 1000 into the new test system.

Kernel, synced from kkit-linux-6.18.y:

  • Bridge IP address unreachable (patches 0082–0084). Every bridge test on the Tactical rig failed on this. The driver stands in hardware VLAN 1 for a VLAN-unaware bridge, and confd creates bridges with vlan_default_pvid 0, which adds VLAN 1 at creation and removes it right after. That forgot the VID 1 broadcast entry, so ARP from the ports never reached the CPU, and the bridge MAC entry, so unicast to the bridge was flooded to the ports. A third fix stops the driver from forgetting a bridged port's MAC on link down, which cut the bridge off whenever its first port, whose MAC the bridge uses, was disabled. Traced with symreg on the EV23X71A.
  • Management port "lockups" (patch 0081). All standalone ports share pvid 0, so they share the MAC table entry that delivers a multicast group to the CPU, but the driver forgot it on behalf of a single port. Taking any other standalone port down, or adding it to a bridge, removed the all-nodes and mDNS entries the management port relies on. Frames were received but never forwarded to the CPU, so the unit dropped out of ixll and the test rig until the port was cycled. Fixed by reference counting the shared entries.
  • Scheduling while atomic in the same path (patch 0079): ndo_set_rx_mode runs under netif_addr_lock_bh, and the MAC table access took a mutex. Now a spinlock with an atomic poll, as lan966x does. Upstream fixed it via the new ndo_set_rx_mode_async, which 6.18 does not have.
  • symreg debugfs driver from Microchip's BSP (patch 0080), for register access by name.

Infix side:

  • PTP ports kept going FAULTY on the copper ports, which failed every PTP test on the rig. The LAN8814 PHYs ask to be the default timestamp provider and the kernel grants it, so each port timestamps with its quad PHY's clock instead of the switch core's. On the EV23X71A the PHY's TX timestamp arrives over MDIO later than the 10 ms ptp4l allows; on the Tactical 1000 it never arrives. confd now selects the PTP clock of the interface's own device, the switch core on Laguna, for every port of an instance, and allows 100 ms for hardware timestamps for boards that only have a PHY clock. One clock for all ports is also what boundary and transparent clocks and the TSN schedules need. Verified on the EV23X71A and on alpha in the rig: no faults, switch PTP interrupts delivering the timestamps.
  • symreg package, selected by the Laguna boards, with the device tree node in each board's Infix overlay. Usage notes, including how to chase a port fault, in board/aarch64/microchip-lan969x/README.md.
  • aarch64_minimal_defconfig lacked the Tactical 1000, EV23X71A and Vero W6m board packages, so PR CI images had no DTB for them and hung after Starting kernel ....
  • provision-tactical did not seed uboot.env, so RAUC could not read any boot state on provisioned units.
  • kernel-refresh.sh deleted the patch directory it had just regenerated when run for the current kernel version.
  • log and follow moved to /etc/profile.d/, so they work in BusyBox ash on minimal builds too.

Risk analysis

Bridge fixes (0082–0084): the VID 1 broadcast entry is now permanent, like the VID 0 one for standalone ports. Standalone ports never classify to VID 1 and PGID_BCAST holds only bridged ports, so it is inert without a bridge. The shared-entry count only covers entries the driver installed itself (locked); auto-learned entries behave as before. Skipping learn and forget on up/down for bridged ports leaves those entries to the bridge FDB notifications the driver already handles; join still forgets the standalone entry and leave still restores it. Upstream has the same three bugs, hidden by the default vlan_default_pvid 1.

Lock conversion (0079): nothing sleeps or takes another lock under sparx5->lock, it nests inside the mutexes and never the reverse, and every caller is process context; the atomic switchdev notifier defers FDB work and never touches the MAC table. A softirq caller would already have been taking a mutex from softirq, so the spinlock tolerates exactly the contexts the mutex had to. Same fix lan966x took in 77bdaf39f3c8. Only behaviour change: a wedged MAC table command busy-waits up to 100 ms instead of sleeping.

Multicast fix (0081): list touched only on group join and leave, held across the hardware learn/forget so two ports cannot race on the same entry. Lock order is always mc_cpu_lock then sparx5->lock. Unicast addresses learned per port have the same shape if two standalone ports share a MAC; left for a follow-up, and not a configuration Infix produces.

symreg: debugfs only, root only; the driver never probes without its device tree node, so it is inert on non-Laguna boards. The tool bypasses the driver when writing, documented in the README.

Checklist

Tick relevant boxes, this PR is-a or has-a:

  • Bugfix
    • Regression tests
    • ChangeLog updates (for next release)
  • Feature
    • YANG model change => revision updated?
    • Regression tests added?
    • ChangeLog updates (for next release)
    • Documentation added?
  • Test changes
    • Checked in changed Readme.adoc (make test-spec)
    • Added new test to group Readme.adoc and yaml file
  • Code style update (formatting, renaming)
  • Refactoring (please detail in commit messages)
  • Build related changes
  • Documentation content changes
    • ChangeLog updated (for major changes)
  • Other (please describe): debug tooling (symreg)

🤖 Generated with Claude Code

Sync patches with kkit-linux-6.18.y, adding a fix for a "scheduling
while atomic" BUG on LAN969x when a multicast address is added to or
removed from an unbridged switch port.  Seen on Tactical 1000 when
IPv6 DAD re-ran after a link flap.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Refreshing the patch series for the current kernel version removed the
patches it had just regenerated, since the script always ran the
upgrade steps:

  kernel-refresh.sh -k ~/src/linux -t v6.18.55 -o 6.18.55

With -o equal to the new version, the rm of the old patch directory
hit the new one.  Without -o it was worse, patches/linux/ itself went.

Skip the directory removal, defconfig bump and tarball checksum when
no old version is given or it matches the new one.  A plain refresh
also no longer downloads the kernel tarball.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
@troglobit troglobit changed the title linux: sparx5: fix sleep in atomic context in MAC table access Laguna: management port multicast fix, provisioning fixes and symreg debug tool Oct 8, 2026
RAUC cannot read the boot state on a Tactical 1000 laid out with
provision-tactical:

  rauc: Failed to get boot state of 'rootfs.0': uboot backend: fw_printenv failed with exit code: 1
  # fw_printenv
  Cannot open /mnt/aux/uboot.env: No such file or directory

The script creates the aux filesystem but never writes uboot.env,
which the stock image gets from image-itb-aux and a regular install
from prod/provision.  Write it the same way, with the default
BOOT_ORDER and one boot attempt per slot.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The minimal builds drop bash, and BusyBox ash does not read
/etc/bash.bashrc even when started as /bin/bash, so the log and
follow helpers were gone from the shell there.

Move them to /etc/profile.d/log.sh as plain sh functions, read by
both ash and bash login shells, including the CLI shell command.
The bash completions for them stay in bash.bashrc.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Sync patches with kkit-linux-6.18.y, adding Microchip's symreg debugfs
driver from their BSP kernel.  It exposes /sys/kernel/debug/symreg/mem,
through which the symreg tool reads and writes LAN969x switch registers
by name.  Needed to debug the management port lockups on Tactical 1000.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The management port on Tactical 1000 locks up now and then, and the
only way to see what the switch is doing is to read its registers.
Add Microchip's symreg tool, which reads and writes LAN969x registers
by symbolic name and dumps the MAC, VLAN, VCAP and stream tables, via
the symreg debugfs driver added with patch 0080.

The package enables the driver in the kernel, and each Laguna board
selects the package and carries the device tree node in its Infix
overlay.  Builds without a Laguna board never see it.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The shared Laguna directory had no README.  Describe what it holds,
what a new Laguna board needs for the symreg tool, and how to use the
tool to chase a port fault, with the management port as the example.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Sync patches with kkit-linux-6.18.y.  On LAN969x, taking any
standalone port down, or adding it to a bridge, removed the multicast
MAC table entries it shared with every other standalone port.  The
management port then dropped IPv6 all-nodes and mDNS traffic until
cycled, which is how it disappeared from ixll and the test rig.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The IP address of a bridge is unreachable from the bridge ports on the
LAN969x boards, while forwarding between the ports works.  Every bridge
test on the Tactical 1000 rig fails on this.

The sparx5 driver stands in hardware VLAN 1 for a VLAN-unaware bridge,
and confd creates bridges with vlan_default_pvid 0, which adds VLAN 1 at
creation and removes it right after.  That forgets the VLAN 1 broadcast
entry, so ARP from the ports never reaches the CPU, and the bridge MAC
entry, so unicast to the bridge is flooded to the ports.  Taking a
bridged port down also forgets its MAC on the bridge VLAN, which cuts
the bridge off whenever its first port, whose MAC the bridge uses, is
disabled.

Three kernel patches, 0082 to 0084, synced from kkit-linux-6.18.y.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
@troglobit troglobit changed the title Laguna: management port multicast fix, provisioning fixes and symreg debug tool Laguna: bridge and management port fixes, provisioning fixes and symreg debug tool Oct 8, 2026
@troglobit troglobit changed the title Laguna: bridge and management port fixes, provisioning fixes and symreg debug tool Laguna: bridge and management port fixes Oct 8, 2026
PTP ports on the EV23X71A copper ports drop to the faulty state a few
seconds after reaching master or slave:

  ptp4l: timed out while polling for tx timestamp
  ptp4l: port 1 (e20): send sync failed
  ptp4l: port 1 (e20): MASTER to FAULTY on FAULT_DETECTED (FT_UNSPECIFIED)

The LAN8814 PHYs on the board do the timestamping, and the kernel picks
them over the switch core as timestamp provider.  The PHY returns the
transmit timestamp from its interrupt handler over MDIO, a bus shared by
25 PHYs and polled by phylib, so it sometimes takes longer than the
10 ms ptp4l waits by default.

Set tx_timestamp_timeout to 100 ms in the generated ptp4l configuration
whenever hardware timestamping is used.  On the EV23X71A, 20 ms already
clears every fault, and a MAC timestamper answers within microseconds,
so the extra headroom costs nothing.  This covers any port that ends up
timestamping in a PHY.  The Laguna ports move to the switch clock in the
next commit, which the Tactical 1000 needs since its PHY timestamps
never arrive at all.
PTP ports on the Tactical 1000 drop to the faulty state within half a
second of becoming master, with the 100 ms transmit timestamp allowance
in place, so every PTP test on the rig still fails:

  ptp4l: port 1 (e5): LISTENING to MASTER on ANNOUNCE_RECEIPT_TIMEOUT_EXPIRES
  ptp4l: port 1 (e5): MASTER to FAULTY on FAULT_DETECTED (FT_UNSPECIFIED)

The LAN8814 PHYs ask to be the default time stamp provider, and the
kernel grants it, so each copper port time stamps with its quad PHY's
clock, /dev/ptp1 to /dev/ptp6, instead of the switch core's.  On the
EV23X71A those timestamps arrive late; on the Tactical 1000 they do not
arrive at all.  Beyond that, a boundary or transparent clock spanning
two quads runs on two clocks, and the TSN schedules of the switch core
follow its own clock, which PTP then never disciplines.

Select the PTP hardware clock of the interface's own device for every
port of an instance, when one exists, before writing the ptp4l
configuration.  The kernel keeps that selection across link down and up.
With the switch clock, e5 on the Tactical 1000 stays master and the
switch's PTP interrupt starts delivering the transmit timestamps.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant