[PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown
Shuai Xue
xueshuai at linux.alibaba.com
Sun Sep 13 23:44:15 PDT 2026
On 9/12/26 6:48 PM, Marc Zyngier wrote:
> Tearing down a full S2 is a pretty involved process, resulting in a
> lot of TLB invalidation. These TLBIs are either on a per leaf basis if
> the HW doesn't support range invalidation, or by top-level range if it
> does. Amusingly, the latter occurs even when nothing has been
> unmapped.
>
> Things are made worse with NV, as we have a bucket-load of shadow S2s,
> and the need to invalidate them all on the back of an MMU notifier.
> The latter will eventually be solved by the reverse-map tracking that
> Wei-Lin is working on, but we need to be better at full-S2 teardown.
>
> This small series adds a "no TLBI" unmapping primitive, which allows
> the caller to then whack the TLBs using a VMID-wide invalidation. This
> results in far fewer TLBIs, and a better recursive virtualisation as
> we get far fewer traps as a consequence.
>
> This applies on top of my shadow-s2 lifetime fixes, and is expected to
> be a prefix to Wei-Lin's series.
>
> Marc Zyngier (4):
> KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
> KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper
> KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all()
> KVM: arm64: nv: Move TLBI VMALLS12E1* emulation over to
> kvm_stage2_unmap_all()
>
> arch/arm64/include/asm/kvm_mmu.h | 1 +
> arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
> arch/arm64/include/asm/kvm_pkvm.h | 1 +
> arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
> arch/arm64/kvm/mmu.c | 15 +++++++++++
> arch/arm64/kvm/nested.c | 4 +--
> arch/arm64/kvm/pkvm.c | 2 ++
> arch/arm64/kvm/sys_regs.c | 18 ++++++-------
> 8 files changed, 75 insertions(+), 21 deletions(-)
>
Hi Marc,
Thanks for putting this series together. I reviewed the four patches
and revisited the traces from my earlier Marc-only tests.
I have a correctness concern about child page-table reclamation.
1. Child page-table reclamation before the final TLBI
In patch 1, SKIP_S2_TLBI suppresses invalidation for both leaf and table
descriptors, while stage2_unmap_walker() still immediately releases
empty child tables.
For a child table with page_count(childp) == 1, the sequence is:
stage2_unmap_walker()
stage2_unmap_put_pte()
clear the parent table descriptor
skip its TLBI
mm_ops->put_page(childp)
kvm_s2_put_page()
put_page() /* drop the child's last reference */
... process the remaining address ranges ...
__kvm_tlb_flush_vmid() /* final invalidation in patch 2 */
This path does not use the free_unlinked_table()/call_rcu() deferred
reclamation mechanism. The existing deferred-range-TLBI path still
invalidates table descriptors immediately; the new flag skips that
invalidation too.
Patch 4 provides a caller operating on active shadow MMUs:
kvm_s2_mmu_iterate_by_vmid() holds mmu_lock for write and visits valid
matching shadow MMUs, but does not require refcnt == 0 or wait for
other vCPUs using the MMU to exit.
Another vCPU can therefore still use that shadow S2. The write lock
excludes software page-table updates, not hardware table walks.
The following interleaving is allowed:
vCPU B / hardware walker vCPU A
------------------------ ----------------------------
Holds an old reference to T
Clears parent, skips TLBI
Drops T's last reference
T is reused by the allocator
Accesses T via the old reference
Performs final VMID-wide TLBI
The final flush barriers cannot retroactively protect a table that
has already been freed and reused.
Is there a lifetime guarantee for these active-MMU callers that I
have missed? Otherwise, the corresponding invalidation needs to
complete before the child table's memory can be reused.
Retaining table-descriptor TLBI looks like a smaller correction
that would still remove empty-range TLBI amplification. Batching
those invalidations as well would require deferred child-table
reclamation, accounting for the lock drop/reacquisition boundaries
in the may_block path.
Thanks.
Shuai
More information about the linux-arm-kernel
mailing list