[RFC] RISC-V KVM: guest local sfence.vma may miss stale VS-stage TLB after vCPU migration
Anup Patel
anup at brainfault.org
Mon Aug 3 07:22:14 PDT 2026
On Sun, Aug 2, 2026 at 10:30 AM Yaxing Guo <guoyaxing at bosc.ac.cn> wrote:
>
> Hi all,
>
> I would like to ask whether RISC-V KVM has a gap around guest sfence.vma
> and VS-stage TLB invalidation when a vCPU migrates across host CPUs.
>
> On bare metal, Linux uses mm_cpumask in __flush_tlb_range() to decide
> whether to send a local or remote sfence.vma. That works because the cpumask
> reflects physical CPUs. In a guest, however, the kernel only sees vCPUs, so
> from the guest point of view a task may never "migrate" even though the
> underlying vCPU moved to a different host CPU. In other words, guest
> mm_cpumask is effectively a vCPU mask, not a host CPU mask.
>
> There is also a related switch_mm()/set_mm_asid() aspect. With the ASID
> allocator enabled, RISC-V Linux does not flush on every context switch. A
> local_flush_tlb_all() is only done when a deferred ASID rollover flush is
> pending for the current CPU. That is fine on bare metal, because previous
> shootdowns are also based on physical CPUs. In a guest, however, the earlier
> shootdown may have been only a guest-local sfence.vma and may have missed an
> old host CPU where the same vCPU ran previously.
>
> A simplified sequence is:
>
> 1. A guest task with mm A runs on vcpu0 while vcpu0 is scheduled on host
> CPU1. CPU1 fills a VS-stage TLB entry for that mm/ASID.
> 2. vcpu0 migrates to host CPU0. KVM's vCPU migration sanitization is local
> to the current host CPU, so it does not invalidate the VS-stage TLB entry
> that was left on host CPU1.
When vcpu0 migrates to host CPU0, the host CPU0 may or may not have
VS-stage TLB entry so VCPU migration sanitization does the right thing.
In the future, vcpu0 may again move back to host CPU1 containing stale
VS-stage TLB entries left in step1 above so the VCPU migration sanitization
will cleanup these stale TLB enteries.
> 3. The guest later updates mm A's page table from vcpu0, for example due to
> COW. Since the guest still sees mm A as having run only on guest CPU0,
> __flush_tlb_range() can choose a local sfence.vma. That local sfence.vma
> only affects host CPU0, where vcpu0 is currently running, so the stale
> VS-stage entry on host CPU1 remains.
No need to worry about stale VS-stage entry on host CPU1 because these
will age-out or same guest may again move to host CPU1 resulting in
VCPU migration sanitization.
> 4. Later, the task is scheduled again while vcpu0 is running on host CPU1.
> If the ASID is still valid and no deferred ASID rollover flush is pending,
> set_mm_asid() may not issue any flush. The access on host CPU1 can then
> hit the old VS-stage TLB entry.
See above comments.
>
> We hit this on a XiangShan RISC-V system. A bash task in the guest was
> scheduled on vcpu0 while that vCPU was running on host CPU1. A read-only
> page was prefetched and cached in CPU1's VS-stage TLB. Later the same vCPU
> migrated to host CPU0, the guest triggered COW and updated the mapping, and
> the guest still issued only a local sfence.vma. The stale VS-stage entry on
> CPU1 was not invalidated. After running for some time, the task later
> accessed that mapping while running on host CPU1 again and hit the stale
> translation. The data in that page was a pointer from the GOT, so it
> dereferenced to NULL and eventually caused a segmentation fault.
The above example is already taken care by TLB flushes in KVM RISC-V.
This smells like some HW errata to me.
>
> I checked current upstream RISC-V KVM (v7.0.2) and I do not see HSTATUS.VTVM
> being enabled, so guest sfence.vma is not trapped. I found
> kvm_riscv_local_tlb_sanitize() on host CPU migration, which handles G-stage
> VMID entries on the current host CPU, but I do not think it generally solves
> the stale VS-stage TLB case described above.
>
> I also noticed the vendor-specific kvm_riscv_vsstage_tlb_no_gpa path in
> current upstream. It looks like a workaround for a particular implementation
> issue on Andes AX66, where the VS-stage TLB does not cache guest physical
> address and VMID, so the normal VMID/GPA-based invalidation model is
> insufficient. In that sense, it does not seem to be a generic solution for
> guest sfence.vma coherence across vCPU migration.
>
> I am wondering what the right direction should be. Enabling HSTATUS.VTVM and
> trapping guest sfence.vma would allow KVM to translate guest local sfence.vma
> into the required host-side hfence.vvma, but that may be too heavy for the
> common case. Another possible direction may be to make the current
> Andes-specific vCPU migration flush logic more generic, so implementations
> that need VS-stage TLB sanitization on vCPU migration can opt in through a
> common mechanism. A third possibility may be to define a lightweight SBI
> interface for guest local sfence.vma, so Linux could use it in
> __flush_tlb_range() for the local case when running under KVM, and KVM could
> then decide whether a local sfence.vma is sufficient or whether host-side
> hfence.vvma is also needed.
Enabling HSTATUS.VTVM has a big impact on performance so certainly
NACK from myside.
>
> My questions are:
> 1. Is the behavior above expected on current RISC-V KVM?
Yes, it should work fine unless there is some HW bug.
> 2. If not, what would be the preferred fix direction?
> 3. Should this be handled by trapping guest sfence.vma with VTVM, by
> generalizing the existing vCPU-migration VS-stage TLB flush logic, or by
> adding a lighter SBI-mediated path for guest local sfence.vma?
> 4. Or is there another intended mechanism to keep VS-stage TLB state coherent
> across vCPU migration?
>
> If needed, I can send a reproducer and the exact hardware/kernel details.
>
Even if you share some reporducing code sequence, we still have don't
have access to your HW so it won't help much.
Regards,
Anup
More information about the kvm-riscv
mailing list