[PATCH v4 35/48] KVM: arm64: gic-v5: Implement save/restore mechanisms for ISTs

Sascha Bischoff Sascha.Bischoff at arm.com
Fri Jul 31 01:55:46 PDT 2026


Hi Fuad,

On Mon, 2026-07-27 at 19:56 +0100, Fuad Tabba wrote:
> On Fri, 24 Jul 2026 at 11:58, Sascha Bischoff
> <Sascha.Bischoff at arm.com> wrote:
> ...
> 
> > +/*
> > + * Save the LPI IST to userspace memory.
> > + *
> > + * The LPI IST may be linear or two-level, so host iteration
> > depends on the
> > + * allocated host shape.
> > + */
> > +int vgic_v5_save_lpi_ist(struct kvm *kvm, struct kvm_vgic_v5_ist
> > *ist_attr)
> > +{
> > +       struct vgic_v5_ist_desc ist;
> > +       u32 __user *uaddr;
> > +       int ret;
> > +
> > +       ret = vgic_v5_get_lpi_ist_desc(kvm, &ist);
> > +       if (ret)
> > +               return ret;
> > +
> > +       if (!ist.present)
> > +               return 0;
> > +
> > +       uaddr = (u32 __user *)(unsigned long)ist_attr-
> > >lpi_ist_addr;
> > +
> > +       if (!ist.vmi->h_lpi_ist_structure)
> > +               return vgic_v5_save_linear_ist(&ist, uaddr,
> > +                                              BIT(ist.id_bits));
> > +
> > +       return vgic_v5_save_two_level_ist(&ist, uaddr);
> > +}
> 
> This bounds the walk with the ID_BITS stamped into the VMTE when the
> host IST was allocated...
> 
> ...
> 
> > +static int vgic_v5_validate_ist_attr(struct kvm *kvm,
> > +                                    const struct kvm_vgic_v5_ist
> > *ist_attr)
> > +{
> > +       unsigned int id_bits;
> > +       int ret;
> > +
> > +       /* We always have SPIs to save */
> > +       ret = vgic_v5_validate_ist_user_buffer(ist_attr-
> > >spi_ist_addr,
> > +                                       ist_attr->spi_ist_size,
> > +                                       kvm->arch.vgic.nr_spis *
> > sizeof(__u32));
> > +       if (ret)
> > +               return ret;
> > +
> > +       /* We don't always have LPIs to save */
> > +       ret = vgic_v5_irs_lpi_ist_id_bits(kvm, &id_bits);
> > +       if (ret < 0)
> > +               return ret;
> > +
> > +       /* No LPI IST */
> > +       if (!ret) {
> > +               if (ist_attr->lpi_ist_addr || ist_attr-
> > >lpi_ist_size)
> > +                       return -EINVAL;
> > +
> > +               return 0;
> > +       }
> > +
> > +       return vgic_v5_validate_ist_user_buffer(ist_attr-
> > >lpi_ist_addr,
> > +                                              ist_attr-
> > >lpi_ist_size,
> > +                                              BIT(id_bits) *
> > sizeof(__u32));
> > +}
> 
> ... and this sizes the buffer from the emulated IRS_IST_CFGR. I think
> the two are using different values, and I could not find anything
> that
> re-syncs them.
> 
> The two writers of IRS_IST_CFGR differ:
> `vgic_v5_mmio_write_irs_ist()`
> returns early while `ist_baser.valid` is set,
> `vgic_v5_mmio_uaccess_write_irs()` does not. So userspace can restore
> the IST, stamping the VMTE from the larger LPI_ID_BITS, rewrite
> IRS_IST_CFGR smaller, and then save with the matching smaller buffer.
> `vgic_v5_ist_cfgr_valid()` accepts down to IRS_IDR2.MIN_LPI_ID_BITS,
> so on a part reporting 0 that is BIT(16) entries into a buffer
> validated at 4 bytes. `test_vgic_v5_ist_save_restore()` in patch 48
> already restores and then re-saves before any vCPU runs, it just does
> not write IRS_IST_CFGR in that window.

Yeah, that's very much not desired! I've gone and have restricted what
userspace can realistically do - we shouldn't be in a situation where
userspace can force things to go out of sync, IMO. 

I've updated the code to reject userspace writes to the IRS_IST_CFGR
once and LPI IST has been allocated - unless userspace just so happens
to write the same value in which case the NOP is allowed. The
allocation of the LPI IST happens as part of restoring the IST, so
broadly userspace is allowed to do whatever it likes, until the IST has
been created, at which point the values are frozen. I've added a
similar restriction for the IRS_IST_BASER - the value is frozen once
the IST is allocated.

> 
> arm-vgic-v5.rst says the LPI buffer size must be BIT(lpi_id_bits) *
> sizeof(__u32), "where lpi_id_bits is the LPI ID width configured in
> IRS_IST_CFGR", which is the value the validation uses and not the one
> the walk uses.

Yeah, but we should mandate that these are the same. Therefore, the
code now rejects LPI IST saves when the VMTE’s IST_ID_BITS differ from
IRS_IST_CFGR.LPI_ID_BITS.

> 
> Adding the `ist_baser.valid` check to the uaccess path would break
> restore, since the same file says the IRS registers may be restored
> in
> any order. Validating against the value the walk uses, or rejecting a
> save when the register and the VMTE disagree,
> both look workable. Bounding the walk by the register instead would
> read past the host allocation whenever the register is the larger of
> the two, so I do not think that direction works.

Agreed, I don't think that direction works. The limit needs to be based
on what the host has allocated.

In general, I've done the following:

- Reject LPI IST saves when the VMTE’s IST_ID_BITS is different from
IRS_IST_CFGR.LPI_ID_BITS.
- Make IRS_IST_CFGR and IRS_IST_BASER immutable after allocating the
LPI IST. Identical userspace writes are accepted but are NOPs.
- Added tests for the BASER and CFGR restrictions to the selftests.

> 
> `put_user()` targets the caller's own mapping and the host-side reads
> stay in bounds, so I think this is a correctness problem rather than
> a
> security one. The SPI save uses `nr_spis` on both sides and restore
> stamps the VMTE from the value it validated,
> so only the LPI save looks affected, linear and two-level alike.
> 
> Cheers,
> /fuad

Thanks,
Sascha



More information about the linux-arm-kernel mailing list