[RFC PATCH 0/4] spi: cadence-xspi: add ACMD PIO support for NAND and NOR

Fei Xie fei.xie at horizon.auto
Sun Sep 27 22:50:05 PDT 2026


Hi Nuno, Miquel,

Thanks for asking for benchmarks. I backported the 32/64-bit STIG SDMA
transfer changes from the patch Nuno pointed to [1] to our vendor 6.1
kernel, and compared that STIG path against the ACMD SPI NAND path from
this RFC. I have not applied Nuno's separate ACMD read implementation;
the figures below are for the implementation in this RFC.

The device is a GD5F4GM8RE SPI NAND (2 KiB pages, 128 KiB eraseblocks)
on a Horizon J6B board, running at 80 MHz in quad SDR mode. Both paths
used the same board PHY configuration. STIG needed one additional dummy
clock for the EBh read-cache command on this board to return correct
data at 80 MHz. This is a local test adjustment, not part of the posted
RFC or [1]; I excluded the incorrect STIG reads from the comparison.

I ran flash_speed -c 5 -d /dev/mtd18 three times with each mode. The
numbers below are the arithmetic means, in KiB/s:

                               STIG      ACMD PIO
  eraseblock write speed      3685          3555
  eraseblock read speed      12468         14222
  page write speed            3671          3510
  page read speed            11931         13716
  erase speed                40000         37647

For a separate sequential-read test, I erased and programmed the first
64 MiB with known random data, then checked that each mode read back
the complete area byte-for-byte. Both SHA-256 hashes matched the source;
the reported ECC failure counts were zero. I then ran:

  /var/busybox/time -f 'real=%e user=%U sys=%S cpu=%P' \
    dd if=/dev/mtd18 of=/dev/null bs=1M count=64

                          STIG              ACMD PIO
  elapsed (three runs)    5.30/5.28/5.26 s  4.66/4.65/4.65 s
  mean throughput         12.12 MiB/s       13.75 MiB/s
  mean system time        3.55 s            1.06 s
  process CPU usage       67%               22%

On this NAND, ACMD improves the 64 MiB read throughput by about 13.5%
and reduces the reading process's CPU usage by 45 percentage points.
The flash_speed results show a similar gain for reads, including
single-page reads. I do not see a program or erase throughput benefit:
ACMD is slightly slower for both in this setup. These are SPI NAND
results only; I have not benchmarked the NOR path on this board.

This is hardware validation of a 6.1 backport, not a claim that the RFC
was tested on a current upstream kernel. I agree the read and CPU gains
need to be weighed against the controller/MTD layering concerns raised
in review.

Thanks,
Fei

[1] https://lore.kernel.org/linux-spi/178056886874.53724.4850286391745939707.b4-ty@b4/



More information about the linux-mtd mailing list