[PATCH v3 12/18] KVM: arm64: Prevent host PC adjustments for protected vCPUs

Marc Zyngier maz at kernel.org
Mon Sep 14 06:42:47 PDT 2026


On Mon, 14 Sep 2026 12:33:32 +0100,
Fuad Tabba <fuad.tabba at linux.dev> wrote:
> 
> __kvm_adjust_pc() lets the host advance a vCPU's PC or inject an
> exception, which for a protected vCPU would let the host redirect
> guest execution. Drop the request there: the entry handlers apply the
> host's PC_UPDATE_REQ on re-entry, where EL2 allows it.
> 
> __kvm_adjust_pc() adjusts the vCPU a get/put pair returns, the one it
> was given outside pKVM. Both host callers hold the vCPU mutex, so a
> hyp vCPU loaded for the vCPU is loaded on the calling CPU. The request
> for a loaded protected vCPU is dropped. For a loaded non-protected
> vCPU, PKVM_HOST_STATE_DIRTY selects the copy, since adjusting the hyp
> vCPU while the host copy is authoritative loses the update at the next
> flush. Adjusting the hyp vCPU copies PC_UPDATE_REQ in and back out
> again: without the copy back, INCREMENT_PC outlives the adjustment and
> the next KVM_SET_VCPU_EVENTS trips WARN_ON(INCREMENT_PC) in
> kvm_pend_exception(). With no hyp vCPU loaded, as under
> KVM_SET_VCPU_EVENTS, the host copy is adjusted as before.
> 
> Until the marshalling patch clears PC_UPDATE_REQ on the host copy at
> exit, a KVM_RUN that returns to userspace with INCREMENT_PC set leaves
> it on the host copy of a loaded protected vCPU, and a
> KVM_SET_VCPU_EVENTS before the next run then trips that WARN_ON.
> 
> Suggested-by: Marc Zyngier <maz at kernel.org>
> Signed-off-by: Fuad Tabba <fuad.tabba at linux.dev>
> ---
>  arch/arm64/kvm/hyp/exception.c             | 21 +++++++++-----
>  arch/arm64/kvm/hyp/include/hyp/adjust_pc.h | 19 +++++++++++++
>  arch/arm64/kvm/hyp/nvhe/hyp-main.c         | 32 ++++++++++++++++++++++
>  3 files changed, 65 insertions(+), 7 deletions(-)
> 
> diff --git a/arch/arm64/kvm/hyp/exception.c b/arch/arm64/kvm/hyp/exception.c
> index 6e60d890afa4a..7ea62e5304ae9 100644
> --- a/arch/arm64/kvm/hyp/exception.c
> +++ b/arch/arm64/kvm/hyp/exception.c
> @@ -353,12 +353,19 @@ static void kvm_inject_exception(struct kvm_vcpu *vcpu)
>   */
>  void __kvm_adjust_pc(struct kvm_vcpu *vcpu)
>  {
> -	if (vcpu_get_flag(vcpu, PENDING_EXCEPTION)) {
> -		kvm_inject_exception(vcpu);
> -		vcpu_clear_flag(vcpu, PENDING_EXCEPTION);
> -		vcpu_clear_flag(vcpu, EXCEPT_MASK);
> -	} else if (vcpu_get_flag(vcpu, INCREMENT_PC)) {
> -		kvm_skip_instr(vcpu);
> -		vcpu_clear_flag(vcpu, INCREMENT_PC);
> +	struct kvm_vcpu *target = pkvm_adjust_pc_get(vcpu);

nit: it is really odd to see this 'pkvm' prefix in generic code. The
point of it is not only to abstract the pvkm complexity away, but also
to have a wrapper that may be of use in other situations.

Don't respin the series just for this though, I may end-up changing
this when applying it.

> +	if (!target)
> +		return;
> +

I find this one scary, see below.

[...]

> diff --git a/arch/arm64/kvm/hyp/nvhe/hyp-main.c b/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> index da8ab636063cf..1dcc75261dc08 100644
> --- a/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> +++ b/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> @@ -642,6 +642,38 @@ static void handle___pkvm_host_mkyoung_guest(struct kvm_cpu_context *host_ctxt)
>  	cpu_reg(host_ctxt, 1) =  ret;
>  }
>  
> +/*
> + * PKVM_HOST_STATE_DIRTY names the authoritative copy, the host's when set.
> + * A loaded protected vCPU takes the request at its next entry instead.
> + */
> +struct kvm_vcpu *pkvm_adjust_pc_get(struct kvm_vcpu *vcpu)
> +{
> +	struct pkvm_hyp_vcpu *hyp_vcpu;
> +
> +	if (!is_protected_kvm_enabled())
> +		return vcpu;
> +
> +	hyp_vcpu = pkvm_get_loaded_hyp_vcpu();
> +	if (!hyp_vcpu || hyp_vcpu->host_vcpu != vcpu)

Under which circumstances do we get hyp_vcpu->host_vcpu != vcpu?

> +		return vcpu;
> +
> +	if (pkvm_hyp_vcpu_is_protected(hyp_vcpu))
> +		return NULL;

Is it always the case that a protected vcpu cannot see its PC adjusted
at all? How is PC updated after an exit for MMIO? I feel there is an
interaction with the above, but I'm not 100% certain...

> +
> +	if (vcpu_get_flag(vcpu, PKVM_HOST_STATE_DIRTY))
> +		return vcpu;
> +
> +	vcpu_copy_flag(&hyp_vcpu->vcpu, vcpu, PC_UPDATE_REQ);
> +	return &hyp_vcpu->vcpu;
> +}
> +
> +/* Reflect the consumed request back, otherwise it stays pending. */
> +void pkvm_adjust_pc_put(struct kvm_vcpu *vcpu, struct kvm_vcpu *target)
> +{
> +	if (target != vcpu)
> +		vcpu_copy_flag(vcpu, target, PC_UPDATE_REQ);
> +}
> +
>  static void handle___kvm_adjust_pc(struct kvm_cpu_context *host_ctxt)
>  {
>  	DECLARE_REG(struct kvm_vcpu *, vcpu, host_ctxt, 1);

Thanks,

	M.

-- 
Without deviation from the norm, progress is not possible.



More information about the linux-arm-kernel mailing list