net: airoha: RX rings below 32 descriptors let hw DMA past the ring

Vitaliy Sochnev sochnev.v.74 at gmail.com
Sun Sep 6 01:02:49 PDT 2026


Correction to the platform description, and a stronger dump.

I wrote that the images "differ only in RX_DSCP_NUM(); no other patches".
That was about the difference between my own images and is misleading as a
description of the driver. The tree is OpenWrt's, and it carries an
out-of-tree HW GRO patch that touches exactly the ring under test:

  AIROHA_RXQ_LRO_EN_MASK = GENMASK(7, 0)   -> rings 0-7, including ring 4

For an LRO ring that patch sets buf_size to 16 KiB instead of PAGE_SIZE/2
and clears RX_RING_SG_EN_MASK, which mainline always sets. Both are
plausibly relevant to a DMA overrun, so the result needed rechecking with
that removed.

To be precise about the base: it is 6.18.44 with the airoha RX path
backported from mainline, including 269389ba5398 ("Set REG_RX_CPU_IDX()
once in airoha_qdma_fill_rx_queue()") and bbfb1983944f ("Reserve RX
headroom to avoid skb reallocation"), both in net today. The GRO patch is
the only out-of-tree piece touching this path.

Rechecked with its LRO mask zeroed, which restores page order 0,
buf_size = PAGE_SIZE/2 and RX_RING_SG_EN - confirmed on the board by the
posted buffer length dropping from 0x3E80 to 0x680. Everything
reproduces:

  ring   LRO   panics                        overrun signature
  16     on    3 (one with no load at all)   present, 560 words
  16     off   2 (one with no load at all)   present, see below
  32     on    none                          0 of 768
  32     off   none, 268 018 frames          0 of 768
  128    on    none                          0 of 1024
  128    off   none, 273 921 frames          0 of 1024

The clean dump, ring 4 = 16, LRO off, taken while the ring was stalled.
Descriptor 16 does not exist in a 16-entry ring:

  desc 15 (last real)      desc 16 (past the end)
  +4  ctrl 0x80000156      +4  ctrl 0xC0000000   DONE|DROP, len 0
      DONE, len 342        +8  addr 0x00000000   no buffer posted
  +20 msg1 0x2A5E0000      +16 msg0 0x00008000
                           +20 msg1 0x2A5E0000
                           +24 msg2 0x007F000E
                           +28 msg3 0x0000FFFF

msg0-msg3 are bit-identical to the dump taken with LRO on, so it is the
same engine either way. addr = 0 explains DROP: hw ran past the posted
descriptors, found no buffer in the next slot - a slot that is not part of
the ring - and marked the completion dropped, but wrote the structure
anyway.

The panic in the ring 4 = 16, LRO off run landed in yet another place:

  nf_conntrack_hash_check_insert+0x480 [nf_conntrack]
  nf_conntrack_in / nf_hook_slow / __ip_local_out / udp_send_skb
  Comm: ntpd

Five panics so far, in four distinct places, none of them networking:
cpufreq's deferred work (__queue_work), the scheduler's load balancer
(sched_balance_rq), the scheduler's stack-end check (twice), and
conntrack's hash insert.

Nothing else in the original mail changes.



More information about the Linux-mediatek mailing list