Random corruption on SpacemiT K1 (and K3) with RVV
Karl Mehltretter
kmehltretter at gmail.com
Wed Sep 23 23:15:11 PDT 2026
On Thu, Sep 24, 2026 at 06:52:21AM +0100, Aurelien Jarno wrote:
> Thanks for your feedback. Note that at this stage I have not been able
> to reproduce the issue with QEMU. I guess it's very timing dependent,
> also I am not sure if QEMU simulates partially executed instructions
> (outside of page faults).
>
Hello Aurelien, Andy,
I tested this with a local TCG diagnostic change. It did not reproduce
the K1 failure, but it answers the partial-instruction question.
My LLM agent helped me running these tests.
In stock TCG at 7074591d7954, vle8.v can leave partial state on a memory
fault, but the vector helper runs atomically with respect to guest
interrupts [1,2]. Stock QEMU therefore cannot take an asynchronous
interrupt partway through this load.
In a bare-metal test at VLEN=256 and LMUL=8, a page fault after element
16 left vstart=16 and a snapshot containing 16 source bytes followed by
240 poison bytes. After the missing page was mapped, QEMU resumed the
load and completed the copy correctly.
I then added a hook to QEMU vector-load which performs 16 elements, sets
vstart=16, raises a timer interrupt, and resumes at the same vle8.v. The
bare-metal test completed 262,145 such restarts and 67,108,864 copied
bytes without a mismatch.
I also booted Linux 388b607d107c with:
CONFIG_RISCV_ISA_V=y
CONFIG_RISCV_ISA_V_UCOPY_THRESHOLD=1
CONFIG_RISCV_ISA_V_PREEMPTIVE=n
With the hook restricted to S-mode loads with SUM and SIE set, a checked
pipe test completed 90,308,608 bytes across copy_from_user() and
copy_to_user(), including demand-faulting source and destination pages.
The hook logged at least 327,680 forced interruptions at vstart=16,
without a byte mismatch.
As a negative control, I made one load read 16 elements and then retire
as if all 256 had completed. That produced the 16-source/240-poison
signature in bare metal, and the Linux checker reported exactly 240
differing bytes. This is an injected symptom, but confirms that the test
detects the reported failure shape.
Correct QEMU fault recovery and forced interrupt restart therefore did
not produce the corruption. Reaching the 16/240 result required
deliberately modelling a load that completed early without a trap. That
fits the observed first load/store pair, but does not establish its
cause. Your IRQ-disabled result also makes a normal asynchronous restart
a poor fit.
The normal Linux load-fault path exits before vse8.v and falls back to
the scalar copy [3,4]. Do you have the original trap PC, cause and fault
address for the second-iteration fault? Those values could show whether
an unexpected synchronous trap is involved.
Thanks,
Karl
[1] https://gitlab.com/qemu-project/qemu/-/blob/7074591d7954876951f84c15b994a43251d5a3c1/target/riscv/tcg/vector_helper.c#L403
[2] https://gitlab.com/qemu-project/qemu/-/blob/7074591d7954876951f84c15b994a43251d5a3c1/accel/tcg/cpu-exec.c#L930
[3] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/arch/riscv/lib/uaccess_vector.S?h=v7.2#n38
[4] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/arch/riscv/lib/riscv_v_helpers.c?h=v7.2#n23
More information about the linux-riscv
mailing list