[PATCH v16 21/45] KVM: arm64: CCA: Handle realm enter/exit

Kohei Enju enju.kohei at fujitsu.com
Mon Aug 10 01:03:02 PDT 2026


On 08/03 14:43, Steven Price wrote:
> Entering a realm is done using a SMC call to the RMM. On exit the
> exit-codes need to be handled slightly differently to the normal KVM
> path so define our own functions for realm enter/exit and hook them
> in if the guest is a realm guest.

Hi Steven,

I found that when the host kernel is booting with pseudo-NMI enabled
(irqchip.gicv3_pseudo_nmi=1), Realm VMs fail boot successfully and the
watchdog reports RCU stalls and soft lockups.

After some investigations, I confirmed that arch_timer interrupts
are not being delivered to some CPUs. This appears to be because PMR is
000000c0 (GICV3_PRIO_IRQ), as shown below:

  [  248.769692] pmr: 000000c0

For normal VMs, KVM opens PMR before entering the guest, because having
IRQs masked via PMR when entering the guest means the GIC will not
signal the CPU of interrupts of lower priority, and in the worst case
a guest exit may never occur.
(See __kvm_vcpu_run in arch/arm64/kvm/hyp/vhe/switch.c)

[  248.756980] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
[  248.758649] rcu:     3-...0: (113 ticks this GP) idle=d2ec/1/0x4000000000000000 softirq=1631/1633 fqs=297
[  248.759413] rcu:     4-...0: (7 ticks this GP) idle=b77c/1/0x4000000000000000 softirq=1589/1589 fqs=297
[  248.760101] rcu:     (detected by 7, t=6003 jiffies, g=4553, q=86 ncpus=8)
[  248.760731] Sending NMI from CPU 7 to CPUs 3:
[  248.766814] NMI backtrace for cpu 3
[  248.768285] CPU: 3 UID: 0 PID: 387 Comm: kvm-vcpu-0 Not tainted 7.2.0-rc2-00267-g7326f0114689 #238 PREEMPT(lazy)
[  248.768681] Hardware name: QEMU QEMU Virtual Machine, BIOS unknown 02/02/2022
[  248.768936] pstate: 61402009 (nZCv daif +PAN -UAO -TCO +DIT -SSBS BTYPE=--)
[  248.769021] pc : kvm_rec_enter+0xa0/0xb8
[  248.769585] lr : kvm_rec_enter+0x24/0xb8
[  248.769655] sp : ffff800084943830
[  248.769692] pmr: 000000c0
[  248.769757] x29: ffff800084943830 x28: ffff0000079c6c00 x27: 0000000000000000
[  248.770441] x26: 0000000000000000 x25: 0000000000000000 x24: 0000000000000000
[  248.770536] x23: 0000000080000000 x22: 0000000000000001 x21: 0000000049d99000
[  248.770590] x20: 0000000049d9a000 x19: ffff00000e018000 x18: 0000000000000000
[  248.770647] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000000
[  248.770756] x14: 0000000000000000 x13: 0000000000000000 x12: 0000000000000000
[  248.770863] x11: 0000000000000000 x10: 0000000000000000 x9 : ffff8000812ab1bc
[  248.771018] x8 : 0000000000000000 x7 : 0000000000000020 x6 : 0000000000000080
[  248.771123] x5 : 0000000000000004 x4 : 0000000000000040 x3 : 0000000000000000
[  248.771172] x2 : 0000000000000000 x1 : 0000000000000000 x0 : 0000000000000000
[  248.771385] Call trace:
[  248.771612]  kvm_rec_enter+0xa0/0xb8 (P)
[  248.771757]  kvm_arm_vcpu_enter_exit+0x8c/0x218
[  248.771796]  kvm_arch_vcpu_ioctl_run+0x274/0x890
[  248.771843]  kvm_vcpu_ioctl+0x180/0xb50
[  248.771883]  __arm64_sys_ioctl+0xb4/0x118
[  248.771920]  invoke_syscall.constprop.0+0xb8/0x120
[  248.771954]  do_el0_svc+0x48/0xc8
[  248.771982]  el0_svc+0x48/0x280
[  248.772012]  el0t_64_sync_handler+0xa0/0xe8
[  248.772041]  el0t_64_sync+0x1ac/0x1b0
[...]

> [...]
> +int noinstr kvm_rec_enter(struct kvm_vcpu *vcpu)
> +{
> +	struct realm_rec *rec = &vcpu->arch.rec;
> +	int ret;
> +
> +	ret = rmi_rec_enter(rec->rec_phys, rec->run_phys);
> +	if (!ret)
> +		load_realm_timer_state(vcpu);
> +
> +	return ret;
> +}

In my testing environment, adding local_daif_mask/restore() in line with
__kvm_vcpu_run() fixes the issue, and Realm VMs successfully boot
without RCU stalls or soft lockups.

However, I'm not sure if we can safely call local_daif_mask/restore()
here, bacause they call trace_hardirqs_off/on(), which presumably cannot
be called from noinstr context.

For reference, the VHE hyp path seems to call local_daif_mask/restore()
from noinstr context, but I'm not sure why this is considered safe:
  noinstr kvm_arm_vcpu_enter_exit()
    kvm_call_hyp_ret(__kvm_vcpu_run, vcpu) <- normal function call for VHE
      local_daif_mask()
        trace_hardirqs_off()
      local_daif_restore()
        trace_hardirqs_off/on()

The following change works in my test environment.
Any thoughts on the issue and the proposed fix?

diff --git a/arch/arm64/kvm/rmi.c b/arch/arm64/kvm/rmi.c
index c242dfc2c7a6..dd7cb2d7db88 100644
--- a/arch/arm64/kvm/rmi.c
+++ b/arch/arm64/kvm/rmi.c
@@ -1307,7 +1307,13 @@ int noinstr kvm_rec_enter(struct kvm_vcpu *vcpu)
        struct realm_rec *rec = &vcpu->arch.rec;
        int ret;

+       local_daif_mask();
+       pmr_sync();
+
        ret = rmi_rec_enter(rec->rec_phys, rec->run_phys);
+
+       local_daif_restore(DAIF_PROCCTX_NOIRQ);
+
        if (!ret)
                load_realm_timer_state(vcpu);

Thanks,
Kohei



More information about the linux-arm-kernel mailing list