Random corruption on SpacemiT K1 (and K3) with RVV

Aurelien Jarno aurelien at aurel32.net
Sun Sep 27 21:32:40 PDT 2026


Hi,

On 2026-09-25 06:40, Aurelien Jarno wrote:
> Hi Karl,
> 
> On 2026-09-24 08:15, Karl Mehltretter wrote:
> > On Thu, Sep 24, 2026 at 06:52:21AM +0100, Aurelien Jarno wrote:
> > > Thanks for your feedback. Note that at this stage I have not been able 
> > > to reproduce the issue with QEMU. I guess it's very timing dependent, 
> > > also I am not sure if QEMU simulates partially executed instructions 
> > > (outside of page faults).
> > > 
> > 
> > Hello Aurelien, Andy,
> > 
> > I tested this with a local TCG diagnostic change. It did not reproduce
> > the K1 failure, but it answers the partial-instruction question.
> > 
> > My LLM agent helped me running these tests.
> > 
> > In stock TCG at 7074591d7954, vle8.v can leave partial state on a memory
> > fault, but the vector helper runs atomically with respect to guest
> > interrupts [1,2]. Stock QEMU therefore cannot take an asynchronous
> > interrupt partway through this load.
> > 
> > In a bare-metal test at VLEN=256 and LMUL=8, a page fault after element
> > 16 left vstart=16 and a snapshot containing 16 source bytes followed by
> > 240 poison bytes. After the missing page was mapped, QEMU resumed the
> > load and completed the copy correctly.
> > 
> > I then added a hook to QEMU vector-load which performs 16 elements, sets
> > vstart=16, raises a timer interrupt, and resumes at the same vle8.v. The
> > bare-metal test completed 262,145 such restarts and 67,108,864 copied
> > bytes without a mismatch.
> > 
> > I also booted Linux 388b607d107c with:
> > 
> >   CONFIG_RISCV_ISA_V=y
> >   CONFIG_RISCV_ISA_V_UCOPY_THRESHOLD=1
> >   CONFIG_RISCV_ISA_V_PREEMPTIVE=n
> > 
> > With the hook restricted to S-mode loads with SUM and SIE set, a checked
> > pipe test completed 90,308,608 bytes across copy_from_user() and
> > copy_to_user(), including demand-faulting source and destination pages.
> > The hook logged at least 327,680 forced interruptions at vstart=16,
> > without a byte mismatch.
> > 
> > As a negative control, I made one load read 16 elements and then retire
> > as if all 256 had completed. That produced the 16-source/240-poison
> > signature in bare metal, and the Linux checker reported exactly 240
> > differing bytes. This is an injected symptom, but confirms that the test
> > detects the reported failure shape.
> > 
> > Correct QEMU fault recovery and forced interrupt restart therefore did
> > not produce the corruption. Reaching the 16/240 result required
> > deliberately modelling a load that completed early without a trap. That
> > fits the observed first load/store pair, but does not establish its
> > cause. Your IRQ-disabled result also makes a normal asynchronous restart
> > a poor fit.
> 
> Thanks for all those extensive tests.
> 
> > The normal Linux load-fault path exits before vse8.v and falls back to
> > the scalar copy [3,4]. Do you have the original trap PC, cause and fault
> > address for the second-iteration fault? Those values could show whether
> > an unexpected synchronous trap is involved.
> 
> Unfortunately the issue is quite rare, so I have not found a way to 
> trace all that information when the issue happens.
> 
> That said Han Gao pointed me to this patch:
> https://lore.kernel.org/linux-riscv/20260807-vector_fpu_regs_status_rmw_fix-v1-1-0c16848b60db@intel.com/
> 
> I have been testing it, and so far it seems to fix my issue with both 
> CONFIG_RISCV_ISA_V_PREEMPTIVE enable and disabled. Or maybe it just 
> hides the issue by changing the timing, as at this point, I haven't 
> fully understood how the bug fixed by this patch can completely explain 
> my observations.

Unfortunately, I have been able to reproduce the issue, even with this 
patch applied. It just significantly reduces the frequency of the issue.

Regards
Aurelien

-- 
Aurelien Jarno                          GPG: 4096R/1DDD8C9B
aurelien at aurel32.net                     http://aurel32.net



More information about the linux-riscv mailing list