[PATCH v4 20/48] KVM: arm64: gic-v5: Add GICv5 IRS IODEV and MMIO emulation

Sascha Bischoff Sascha.Bischoff at arm.com
Thu Jul 30 09:06:29 PDT 2026


Hi Fuad,

On Mon, 2026-07-27 at 19:34 +0100, Fuad Tabba wrote:
> Hi Sascha,
> 
> On Fri, 24 Jul 2026 at 11:54, Sascha Bischoff
> <Sascha.Bischoff at arm.com> wrote:
> ...
> 
> > +       case GICV5_IRS_IDR1:
> > +               value = FIELD_PREP(GICV5_IRS_IDR1_PE_CNT,
> > +                                  atomic_read(&vcpu->kvm-
> > >online_vcpus));
> > +               /*
> > +                * IRS_IDR1 encodes IAFFID_BITS as N - 1.
> > +                */
> > +               vpe_id_bits = vgic_v5_vmte_vpe_id_bits(vcpu);
> > +               value |= FIELD_PREP(GICV5_IRS_IDR1_IAFFID_BITS,
> > vpe_id_bits - 1);
> 
> I think the advertised IAFFID width might not cover the IAFFIDs KVM
> hands out.
> 
> vgic_v5_vmte_vpe_id_bits() sizes the VPE ID space from the vCPU
> count,
> but the IAFFID a guest reads from ICC_IAFFIDR_EL1 is vcpu_id, which
> userspace picks. Two vCPUs created with IDs 0 and 300 give an
> IAFFID_BITS of one implemented bit, with the second vCPU reporting
> 300. I could not find anything in the spec requiring the reported
> IAFFID to fit the advertised width, but to be sufficient to identify
> every PE in the system if my read is correct.

Broadly, I think that the rules you're effectively looking for in the
A.a spec [1] are R_JKYGN and S_WVYLC:

R_JKYGN
   When the Affinity of an interrupt is configured by the GIC <domain>AFF,
   <Xt> instruction and the IRM field configures the interrupt to be
   Targeted, all of the following are true:
      * For a Physical Interrupt Domain, the IAFFID field specifies the PE
      with the corresponding interrupt Affinity ID.
      * For the Virtual Interrupt Domain, the IAFFID field specifies the VPE
      in the VM using the corresponding VPE ID.

S_WVYLC
   Arm expects that Hypervisor software emulates a virtual GICv5
   implementation where the emulated ICC_IAFFIDR_EL1.IAFFID value is the
   VPE ID of the corresponding VPE. This means that to software running in
   a VM, configuring the Affinity of a virtual interrupt using the
   virtualized ICC_IAFFIDR_EL1.IAFFID value will have the desired effect.

There are some upcoming spec clarifications here which explicitly state
that the VPE_ID_BITS field in the VMTE is used as the limit for virtual
IAFFID bits. Hence, we need to convey the same limit here to make it
clear to the guest that upper bits might be ignored or that the IAFFID
might be treated as invalid if those are non-zero.

Coming to the issue you've mentioned; you are right. The elephant in
the room is that I've been using vcpu_idx to index into the VPET. That,
alas, assumes that userspace is nice and allocates contiguous,
incremental IDs which just so happen match KVM's internal, non-sparse
representation in vcpu_idx. Not only was this a really bad assumption
on my part, but it would also have meant that the IAFFID used by the
guest didn't always match the ID.

I've reworked the code to do the following:

1. Always use vcpu->vcpu_id as the GICv5 VPE ID. This means that the
userspace-picked ID is used to index into the table. This then also
means that the IAFFID used by the guest matches what userspace has
chosen. As an aside, the emulation of guest accesses to ICC_IAFFIDR_EL1
was already returning vcpu_id and not vcpu_idx. A former bug; now a
convenient feature.

2. Updated the VPET sizing code to allow for a sparse set of VPE IDs.
We now iterate over all vCPUs on vGICv5 init, find the max ID, and use
this (rounded up to the next power of 2) for VPET sizing. Unused
indexes within this table remain invalid, don't get a VPE Descriptor,
and cannot be made resident. The only valid VPEs in the VPET are those
indexed by vcpu->vcpu_id.

Yes, this wastes some memory, but if userspace wants to pick IDs
poorly, there's little we can do here. In the worst case, where we have
one VPE with ID 256-511 (we cap to at most 512 VPEs per GICv5 VM in
KVM), we'd consume 4KB for the VPET of which we use 8 bytes (i.e., one
entry).

3. Added in code to reject vGICv5 creation if there are already vCPUs
present and they have IDs that cannot be represented.

4. Added in code to reject vCPU creation post-vGICv5 creation if the ID
is out of range for the given hardware.

5. Added in a self test for sparse vCPU IDs with GICv5. Hence, I've got
some confidence in this rework.

With those changes, userspace is now able to create vCPUs in any order
with any (in range) IDs. As an aside, KVM already has a duplicate vCPU
ID check so should reject any duplicates, so we should never be in the
situation where we have two vCPUs that use the same ID and hence VPE.

I hope this addresses your issue! It was more an an adventure than I'd
expected, but we're definitely in a better place now.

> 
> Cheers,
> /fuad

Thanks,
Sascha

[1] https://support.arm.com/documentation/111701/latest


More information about the linux-arm-kernel mailing list