Skip to content

zephyr-cp: enforce the _bleio scan timeout on legacy-scan controllers - #20

Open
tyeth wants to merge 1 commit into
zephyr-cp-pico2w-usb-btconnfrom
zephyr-cp-ble-scan-timeout
Open

tyeth wants to merge 1 commit into
zephyr-cp-pico2w-usb-btconnfrom
zephyr-cp-ble-scan-timeout

Conversation

@tyeth

@tyeth tyeth commented Sep 9, 2026 •

Copy link
Copy Markdown
Owner

Update 2026-10-03: rebased onto CircuitPython 11; the console-wake commit dropped

Rebased onto upstream main @ 35210c2e31 (after 11.0.0-alpha.1). This PR is now the single commit 0f69da0e25 (the scan-timeout fix).

  • Dropped: the second commit (console RX callback wakes the main thread). Upstream made the same change in 2b8a218385.
  • Conflict resolved in Adapter.c: upstream removed bleio_adapter_reset, so it is not re-added here. stop_scan already clears the scan deadline, which covers what it did for this fix.
  • The scan-timeout enforcement itself is unchanged. The rebase found no upstream equivalent, so this fix stays.
  • CI: green on ci/pico2w-ble-assets (Pico 2 W run 37126091753, Pico W run 37127295820). UF2s are in prerelease v-zephyr-cp-ble-ci-20261003. Not hardware-tested since the rebase. The gdb evidence and timing table below are from the pre-rebase build.

The "post-scan stall" in #18 is the scan timeout being ignored

start_scan(timeout=...) never ended on the Pico 2 W. The ScanResults loop ran until Ctrl-C, and because scanresults_next() returns None when an interrupt is pending, the loop exited cleanly and the KeyboardInterrupt was raised on the first statement after the loop — stop_scan() in one run, the print() after it in another. That is exactly the two tracebacks in #18; it is not a stall in teardown or in CDC output.

gdb on the running board (ELF verified against flash with compare-sections), halted 9 s into a timeout=3 scan:

  • main thread: common_hal_bleio_scanresults_next() at ScanResults.c:27, done = false, ring buffer used = 0 (draining faster than reports arrive)
  • bt_dev.flags = 0x35 = ENABLE | READY | HAS_PUB_KEY | SCANNING
  • bt_dev.ncmd_sem.count = 1, sent_cmd = 0 — no HCI command outstanding
  • cyw43_bus_mutex.owner = 0, lock_count = 0 — gSPI bus free
  • bt_poll_thread state 0x04 in z_tick_sleep — its normal 4 ms poll
  • every work_queue_main, usbd_thread, udc_rpi_pico_thread_0, tc_rx_handler, airoc_event_task pended (0x02)

Nothing is blocked. Explicit stop_scan() takes 5–6 ms.

Cause (Zephyr host): bt_le_scan_param.timeout is passed to the controller only on the extended-scanning path. start_le_scan_legacy() (subsys/bluetooth/host/scan.c) never reads it, and CONFIG_BT_EXT_ADV=n — required because the CYW43439 has no extended advertising — selects that path. The timeout was silently dropped.

Fix: the adapter keeps the deadline and enforces it from bleio_background() (called from port_background_task()), on the main thread where bt_le_scan_stop()'s blocking HCI round-trip is safe.

timeout loop exited after reports
3 s 3.0 s —
1 s 1.0 s —
0.5 s 0.53 s 12
2 s 2.05 s 30
before: 3 s still running at 57 s 1600

Second commit: the console RX callback now wakes the main thread (the "press any key" wait after code.py never woke on a keypress).

Size (Pico 2 W): no measurable change.

🤖 Generated with Claude Code

start_scan(timeout=...) never ended on the Pico 2 W: the ScanResults loop ran
until Ctrl-C, and the KeyboardInterrupt then surfaced on the first statement
after the loop (stop_scan() in one run, the print() after it in another),
which looked like a stall in scan teardown or in CDC output.

gdb on the running board showed nothing blocked. The main thread was in
common_hal_bleio_scanresults_next() waiting for the next entry with done
still false; bt_dev.flags had BT_DEV_SCANNING set; ncmd_sem.count was 1 and
sent_cmd 0 (no HCI command outstanding); the gSPI bus mutex was free; the
BT RX poll thread was in its 4 ms k_msleep; every work queue and USB thread
was pended idle. stop_scan() itself took 5-6 ms when called explicitly.

The cause is in Zephyr's host: bt_le_scan_param.timeout is only passed to the
controller on the extended-scanning path (LE Set Extended Scan Enable carries
a duration and the controller reports LE Scan Timeout). start_le_scan_legacy()
never reads it, and the legacy path is what CONFIG_BT_EXT_ADV=n selects --
which a controller without extended advertising, such as the CYW43439,
forces. So the timeout was silently ignored and the scan ran forever.

Keep the deadline in the adapter and enforce it from bleio_background(),
called from port_background_task() on the main thread, where bt_le_scan_stop()
is safe to call (it blocks on an HCI round-trip, which must not happen on the
system work queue that also runs the USB CDC console). The ScanResults
iterator finishes at the deadline as it does on nRF and ESP32.

Measured on hardware, Pico 2 W: timeout=3 s -> loop exited after 3.0 s;
timeout=1 -> 1.0 s; timeout=0.5 -> 0.53 s; timeout=2 -> 2.05 s (30 reports).
Before the change a timeout=3 scan was still yielding at 57 s (1600 reports).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tyeth
tyeth force-pushed the zephyr-cp-pico2w-usb-btconn branch from e6b7fd1 to 7527914 Compare October 3, 2026 13:39
@tyeth
tyeth force-pushed the zephyr-cp-ble-scan-timeout branch from 3c6020d to 0f69da0 Compare October 3, 2026 13:39
tyeth added a commit to tyeth/deepsleep_espnow_wifi_and_ble_env_collector that referenced this pull request Oct 3, 2026
PR #11's transport and Pico node, carried onto the shims. node/code.py
stays tiny and picks the node: nodemain where espnow + alarm import (every
ESP32 -- both are built-ins nodemain imports first thing anyway), else
node_lite. node_lite is imported (setup) and then run(), so a MemoryError
from the import is reported as "does not fit this heap" with the REPL left
up, and one from a running loop is not mistaken for it.

  envadv.py           one reading in an 18-byte Manufacturer Specific Data
                      structure, 25 bytes on air; identical in both trees.
                      Now keeps co2/voc/nox as ints (as over ESP-NOW) and
                      carries "sim" so bench readings stay labelled.
  node/net_bleadv.py  raw _bleio, non-connectable
  node/node_lite.py   read -> advertise ble_adv_s -> optional WiFi POST ->
                      time.sleep() the rest; an awake loop, said plainly
  collector/net_blescan.py
                      short passive scan every few seconds, de-dup by
                      address + seq, into take_node_packet(mac=None)

Corrections from the hardware results on #11:
  * start_scan(timeout=) is ignored on Zephyr's legacy scan path before
    tyeth/circuitpython#20 -- scans never ended, so #11's hub would have
    stuck in its first poll(). net_blescan now enforces its own deadline
    (scan_s + 0.5 s) and stops the scan, counting `overruns`. In a room
    with no reports at all the unfixed firmware can still block between
    reports; that needs #20.
  * CONFIG_BT_MAX_CONN=1 was Zephyr's default, now 4 on the fork: the
    docstrings no longer give it as the reason for broadcasting. The
    broadcast stays (no connection, no adafruit_ble on the node); a
    connection with ESP-NOW-style confirmation is the open alternative.
  * adafruit_ble frozen into the Pico W image (#23) costs ~10 KB, not
    ~22-30 KB; noted where the portal-on-a-Pico-W question is discussed.

The hub scans only where it has no ESP-NOW of its own unless
"ble_scan_nodes": true says otherwise, so an ESP32 hub does exactly what it
did; ble_rx/ble_err join the mesh status when the scanner runs.

Exercised under CPython with stubs: node/code.py -> node_lite runs cycles
with the sim sensor and every advertisement decodes; a Pico-2-W-shaped hub
whose stub start_scan ignores its timeout stores the node as ble-62AC and
stops the scan itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant