[REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active
Andrei Rusu de Castro
arc at empyreal.works
Sun Oct 4 08:48:46 PDT 2026
Hi Devin, Jonas,
I reproduced the stall with my September patch on an MT7925 PCIe Z13
and traced the station commands sent to firmware. The missing case was
the primary update inside secondary link addition.
mt7925_mac_link_sta_add() updates the primary WCID first, then the new
secondary WCID. The new secondary is not in msta->link[] or valid_links
until both commands succeed. My first patch included an unpublished
link only when that link was the subject of the current command. It
therefore still sent a one-link MLD description to the primary WCID,
followed by a two-link description to the secondary WCID.
These are the captured fields during secondary addition, omitting an
earlier primary-only association update:
command WCID count entries (WCID/BSS)
September patch 1 1 1/0
2 2 1/0, 2/1
September patch + early 1 2 1/0, 2/1
link publication 2 2 1/0, 2/1
v2 1 2 1/0, 2/1
2 2 1/0, 2/1
V2 passes the initialized pending secondary explicitly through both
station updates, without publishing it early. It keeps the existing
publication-after-success and host error cleanup. It also keeps stable
membership for subsequent updates. There is no new shared pending state.
On Linux 7.3-rc5 with firmware 20260813113118 and an ASUS GT-BE98:
upstream MLD path stalled at 157 seconds
September patch stalled at 289 seconds; repeated at 150
v2 813 seconds, 48 samples, no failures
v2, second clean boot 812 seconds, no sustained stall under
20 Mbit/s traffic in each direction
The AP advertises three links. The active client pair was 5+6 GHz
(active_links=0x3, valid_links=0x7). Failed runs had roughly 27-second
full-band scans; both v2 runs returned to roughly 6.3-7.1 seconds.
The second v2 run transferred about 1.9 GB in each direction over 760
seconds. It had one gateway-ping timeout after a neighbor flush, with
REACHABLE neighbor state and successful IP/HTTPS probes in the same
sample, so it did not pass the strict zero-failure gate. Traffic did not
enter the persistent associated-but-dead state.
Tests used the in-tree driver plus the selected fix, with the separate
scheduled-scan withdrawal and WM2 reset correction held constant. Other
local platform patches did not change between these arms. My production
variant additionally retains the applicable local safety guards in a
separate patch; those guards are not part of this submission. That
variant passed the full zero-failure gate on two PCIe Z13 machines
(813 and 810 seconds), and both booted it successfully as their normal
default. Five ordinary reconnects also passed on the first local variant.
The changed objects build with W=1 and -Werror on both the tested rc5
source and the current mt76 integration branch. Source-extracted fixtures
exercise the pending-link records and host cleanup at eight failing MCU
command positions. Those tests do not establish firmware rollback after
a partly successful add.
Direct debugfs link switches and chip-reset tests exposed failures or
aggregation teardown warnings that also occur with my previous working
full-revert kernel. V2 is not claimed to resolve those. USB hardware and
your exact FritzBox setup have not been tested here.
The v2 posted in reply to this message is based on mt76 commit
0dbc9c9fa9b9. Could you try it on
the setups where the September patch still stalled?
Thanks,
Andrei
More information about the Linux-mediatek
mailing list