[PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure)
Wang Han
wanghan at linux.alibaba.com
Wed Sep 2 09:35:00 PDT 2026
Hi Wei-Lin,
I tested this series on a Yitian 710 system with an ARM Neoverse-N2 CPU
(128 CPUs, 2 NUMA nodes).
Test environment
----------------
L0 kernel: Linux v7.2-rc6
L1 guest: Ubuntu 26.04 LTS, kernel 7.0.0-27-generic (aarch64)
QEMU: 10.2.3
L0 NUMA balancing was enabled (`/proc/sys/kernel/numa_balancing=1`).
The host was booted with `kvm_arm.mode=nested`.
This series fixes a functional hang that is exposed when NUMA balancing is
enabled. The previous nested stage-2 unmap path is too slow for this
workload, making the performance problem user-visible: NUMA balancing can
leave the L1 guest unable to make progress and eventually hang during boot.
The L1 was started with 8 vCPUs and 32 GiB of RAM using:
qemu-system-aarch64 -smp 8 -m 32G \
-machine virt,accel=kvm,gic-version=3,virtualization=on \
-cpu host -nographic -enable-kvm \
-drive if=pflash,format=raw,readonly=on,file=pflash0_bak.img \
-drive if=pflash,format=raw,file=pflash1_bak.img \
-drive file=./ubuntu-vm.qcow2,format=qcow2,if=virtio,cache=none,aio=native \
-nic user,model=virtio-net-pci,hostfwd=tcp::11234-:22 \
-serial mon:stdio
With upstream v7.2-rc6 (075b74841bd0065a3bda3440873c747938e69b68),
L0 NUMA balancing enabled, and the same QEMU configuration, the L1 guest
hung during boot. The original L1 console reported:
[ 76.595764] watchdog: BUG: soft lockup - CPU#1 stuck for 45s! [k8s-dqlite:3068]
[ 76.595973] watchdog: BUG: soft lockup - CPU#0 stuck for 38s! [rs:main Q:Reg:1602]
[ 76.596181] watchdog: BUG: soft lockup - CPU#7 stuck for 45s! [kubelite:2672]
[ 76.596334] watchdog: BUG: soft lockup - CPU#5 stuck for 38s! [containerd:2375]
The corresponding L0 hung-task report was:
[Wed Sep 2 22:49:08 2026] INFO: task qemu-system-aar:14681 blocked in I/O wait for more than 120 seconds.
[Wed Sep 2 22:49:08 2026] Tainted: G E N 7.2.0-rc6-poluted_opt-7-2-numa #10.al8
[Wed Sep 2 22:49:08 2026] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
[Wed Sep 2 22:49:08 2026] task:qemu-system-aar state:D stack:0 pid:14681 tgid:14681 ppid:14680 task_flags:0x8400080 flags:0x00800000
[Wed Sep 2 22:49:08 2026] Call trace:
[Wed Sep 2 22:49:08 2026] __switch_to+0x128/0x168 (T)
[Wed Sep 2 22:49:08 2026] __schedule+0x278/0x910
[Wed Sep 2 22:49:08 2026] schedule+0x3c/0xe8
[Wed Sep 2 22:49:08 2026] io_schedule+0x44/0x68
[Wed Sep 2 22:49:08 2026] softleaf_entry_wait_on_locked+0x280/0x2d0
[Wed Sep 2 22:49:08 2026] migration_entry_wait+0xdc/0x140
[Wed Sep 2 22:49:08 2026] do_swap_page+0x834/0xd80
[Wed Sep 2 22:49:08 2026] handle_pte_fault+0x208/0x2b8
[Wed Sep 2 22:49:08 2026] __handle_mm_fault+0x228/0x528
[Wed Sep 2 22:49:08 2026] handle_mm_fault+0xdc/0x2d8
[Wed Sep 2 22:49:08 2026] do_page_fault+0x388/0x790
[Wed Sep 2 22:49:08 2026] do_translation_fault+0x4c/0x88
[Wed Sep 2 22:49:08 2026] do_mem_abort+0x4c/0xa0
[Wed Sep 2 22:49:08 2026] el0_da+0x54/0x178
[Wed Sep 2 22:49:08 2026] el0t_64_sync_handler+0xd0/0xe8
[Wed Sep 2 22:49:08 2026] el0t_64_sync+0x1ac/0x1b0
The same wait and stack were observed repeatedly; the hung-task report
recurred after 241 and 362 seconds.
I applied the v5 series to the same v7.2-rc6 baseline. With L0 NUMA
balancing still enabled and the same QEMU command line, the L1 guest
booted normally and could be operated through the serial console. QEMU
no longer hung. No new soft-lockup, hung-task, blocked-I/O,
migration-entry-wait, or softleaf-entry-wait message was observed in the
patched boot log.
As a control, the unmodified v7.2-rc6 kernel with
/proc/sys/kernel/numa_balancing=0 also booted the L1 with the same QEMU
configuration. This confirms that the patch removes the NUMA-balancing
failure mode in this setup rather than merely changing the guest setup.
Tested-by: Wang Han <wanghan at linux.alibaba.com>
Thanks,
Wang Han
More information about the linux-arm-kernel
mailing list