[PATCH v15 14/37] KVM: arm64: CCA: Handle realm enter/exit
Suzuki K Poulose
suzuki.poulose at arm.com
Thu Jul 30 07:57:56 PDT 2026
On 30/07/2026 14:58, Steven Price wrote:
> On 27/07/2026 09:14, Kohei Enju wrote:
>> Hi Steven, I have a comment about rec_exit_sys_reg() below.
>>
>> On 07/15 15:28, Steven Price wrote:
>>> Entering a realm is done using a SMC call to the RMM. On exit the
>>> exit-codes need to be handled slightly differently to the normal KVM
>>> path so define our own functions for realm enter/exit and hook them
>>> in if the guest is a realm guest.
>>>
>>> Signed-off-by: Steven Price <steven.price at arm.com>
>>> Reviewed-by: Gavin Shan <gshan at redhat.com>
>>> ---
>>> Changes since v13:
>>> * The RMM is now required to provide an ESR value with the correct
>>> information to emulate MMIO, so we no longer need to hardcode 0s in
>>> rec_exit_sys_reg().
>>> * The PSCI changes mean that there is a potential race when turning on
>>> a VCPU which can cause a RMI_ERROR_REC return. Exit to user space
>>> with -EAGAIN in this case.
>>> Changes since v12:
>>> * Call guest_state_{enter,exit}_irqoff() around rmi_rec_enter().
>>> * Add handling of the IRQ exception case where IRQs need to be briefly
>>> enabled before exiting guest timing.
>>> Changes since v8:
>>> * Introduce kvm_rec_pre_enter() called before entering an atomic
>>> section to handle operations that might require memory allocation
>>> (specifically completing a RIPAS change introduced in a later patch).
>>> * Updates to align with upstream changes to hpfar_el2 which now (ab)uses
>>> HPFAR_EL2_NS as a valid flag.
>>> * Fix exit reason when racing with PSCI shutdown to return
>>> KVM_EXIT_SHUTDOWN rather than KVM_EXIT_UNKNOWN.
>>> Changes since v7:
>>> * A return of 0 from kvm_handle_sys_reg() doesn't mean the register has
>>> been read (although that can never happen in the current code). Tidy
>>> up the condition to handle any future refactoring.
>>> Changes since v6:
>>> * Use vcpu_err() rather than pr_err/kvm_err when there is an associated
>>> vcpu to the error.
>>> * Return -EFAULT for KVM_EXIT_MEMORY_FAULT as per the documentation for
>>> this exit type.
>>> * Split code handling a RIPAS change triggered by the guest to the
>>> following patch.
>>> Changes since v5:
>>> * For a RIPAS_CHANGE request from the guest perform the actual RIPAS
>>> change on next entry rather than immediately on the exit. This allows
>>> the VMM to 'reject' a RIPAS change by refusing to continue
>>> scheduling.
>>> Changes since v4:
>>> * Rename handle_rme_exit() to handle_rec_exit()
>>> * Move the loop to copy registers into the REC enter structure from the
>>> to rec_exit_handlers callbacks to kvm_rec_enter(). This fixes a bug
>>> where the handler exits to user space and user space wants to modify
>>> the GPRS.
>>> * Some code rearrangement in rec_exit_ripas_change().
>>> Changes since v2:
>>> * realm_set_ipa_state() now provides an output parameter for the
>>> top_iap that was changed. Use this to signal the VMM with the correct
>>> range that has been transitioned.
>>> * Adapt to previous patch changes.
>>> ---
>>>
>>> [...]
>>>
>>> +static int rec_exit_sys_reg(struct kvm_vcpu *vcpu)
>>> +{
>>> + struct realm_rec *rec = &vcpu->arch.rec;
>>> + unsigned long esr = kvm_vcpu_get_esr(vcpu);
>>> + int rt = kvm_vcpu_sys_get_rt(vcpu);
>>> + bool is_write = (esr & ESR_ELx_SYS64_ISS_DIR_MASK) == ESR_ELx_SYS64_ISS_DIR_WRITE;
>>> + int ret;
>>> +
>>> + if (is_write)
>>> + vcpu_set_reg(vcpu, rt, rec->run->exit.gprs[rt]);
>>
>> When rt is 31 (XZR), does exit.gprs[rt] trigger an out-of-bounds read
>> since REC_RUN_GPRS is 31? Although the padding after the gprs means this
>> OOB may not cause any practical issue, would it make sense to skip the
>> access when rt is 31?
>>
RMM mandates that the ESR_ELx.ISS.RT == 0 on an exit due to System
register access. See "A4.3.4.4 REC exit due to System register access"
So this shouldn't be a problem, but doesn't hurt to defend the host
against a malicious RMM ?
Suzuki
>> if (is_write && rt != 31)
>> vcpu_set_reg(vcpu, rt, rec->run->exit.gprs[rt]);
>
> Very true - as you say in practise this isn't a big issue because
> vcpu_set_reg() is a no-op, so it's just a read of padding. But
> definitely worth fixing.
>
>>> +
>>> + ret = kvm_handle_sys_reg(vcpu);
>>> + if (!is_write)
>>> + rec->run->enter.gprs[rt] = vcpu_get_reg(vcpu, rt);
>>
>> The same applies here:
>>
>> if (!is_write && rt != 31)
>> rec->run->enter.gprs[rt] = vcpu_get_reg(vcpu, rt);
>
> And the same here - a zero written into the padding.
>
> Thanks for the review!
>
> Steve
>
>>> +
>>> + return ret;
>>> +}
>
More information about the linux-arm-kernel
mailing list