[PATCH v7 02/24] KVM: arm64: Donate MMIO to the hypervisor
Mostafa Saleh
smostafa at google.com
Wed Aug 5 06:42:10 PDT 2026
On Wed, Aug 05, 2026 at 12:47:03PM +0000, Sebastian Ene wrote:
> On Wed, Jul 15, 2026 at 11:58:43AM +0000, Mostafa Saleh wrote:
> > Add a function to donate MMIO to the hypervisor so IOMMU hypervisor
> > drivers can protect and access the MMIO of IOMMUs.
> >
> > As donating MMIO is very rare, and we don’t need to encode the full
> > state, it’s reasonable to have a separate function to do this.
> > It will init the host s2 page table with an invalid leaf with the owner ID
> > to prevent the host from mapping the page on faults.
> >
> > Also, prevent kvm_pgtable_stage2_unmap() from removing owner ID from
> > stage-2 PTEs, as this can be triggered from recycle logic under memory
> > pressure. There is no code relying on this, as all ownership changes is
> > done via kvm_pgtable_stage2_set_owner()
> >
> > For the error path in IOMMU drivers, add a function to donate MMIO
> > back from hyp to host. However, that leaks the hypervisor virtual
> > address range which should be acceptable as this is quite rare and
> > it matches the behaviour of fix_map/block.
> >
> > Signed-off-by: Mostafa Saleh <smostafa at google.com>
> > ---
> > arch/arm64/kvm/hyp/include/nvhe/mem_protect.h | 7 ++
> > arch/arm64/kvm/hyp/nvhe/mem_protect.c | 91 ++++++++++++++++++-
> > arch/arm64/kvm/hyp/pgtable.c | 11 +--
> > 3 files changed, 102 insertions(+), 7 deletions(-)
> >
> > diff --git a/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h b/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h
> > index 29935c7da1de..51b0eb3844a9 100644
> > --- a/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h
> > +++ b/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h
> > @@ -36,6 +36,13 @@ int __pkvm_guest_share_host(struct pkvm_hyp_vcpu *vcpu, u64 gfn);
> > int __pkvm_guest_unshare_host(struct pkvm_hyp_vcpu *vcpu, u64 gfn);
> > int __pkvm_host_unshare_hyp(u64 pfn);
> > int __pkvm_host_donate_hyp(u64 pfn, u64 nr_pages);
> > +/*
> > + * Donate MMIO range to the hypervisor, it will be mapped in the hypervisor's
> > + * private range and unmapped from the host stage-2.
> > + */
> > +int __pkvm_host_donate_hyp_mmio(phys_addr_t addr, size_t size, unsigned long *haddr);
> > +/* Remaps MMIO range in the host, typically used in error path. */
> > +int __pkvm_hyp_donate_host_mmio(phys_addr_t addr, size_t size);
> > int __pkvm_hyp_donate_host(u64 pfn, u64 nr_pages);
> > int __pkvm_host_share_ffa(u64 pfn, u64 nr_pages);
> > int __pkvm_host_unshare_ffa(u64 pfn, u64 nr_pages);
> > diff --git a/arch/arm64/kvm/hyp/nvhe/mem_protect.c b/arch/arm64/kvm/hyp/nvhe/mem_protect.c
> > index 4e329e39a695..d803b3dd4cb4 100644
> > --- a/arch/arm64/kvm/hyp/nvhe/mem_protect.c
> > +++ b/arch/arm64/kvm/hyp/nvhe/mem_protect.c
> > @@ -378,7 +378,11 @@ static int host_stage2_unmap_dev_all(void)
> > u64 addr = 0;
> > int i, ret;
> >
> > - /* Unmap all non-memory regions to recycle the pages */
> > + /*
> > + * Unmap all non-memory regions to recycle the pages.
> > + * That relies on kvm_pgtable_stage2_unmap() not clearing
> > + * counted PTEs which include hypervisor MMIO.
> > + */
> > for (i = 0; i < hyp_memblock_nr; i++, addr = reg->base + reg->size) {
> > reg = &hyp_memory[i];
> > ret = kvm_pgtable_stage2_unmap(pgt, addr, reg->base - addr);
> > @@ -1119,6 +1123,91 @@ int __pkvm_host_donate_hyp(u64 pfn, u64 nr_pages)
> > return ret;
> > }
> >
> > +int __pkvm_host_donate_hyp_mmio(phys_addr_t addr, size_t size, unsigned long *haddr)
> > +{
> > + kvm_pte_t pte;
> > + u64 offset;
> > + int ret;
> > +
> > + /* Only before de-privilege. */
> > + if (static_branch_unlikely(&kvm_protected_mode_initialized))
> > + return -EPERM;
> > +
> > + if (!PAGE_ALIGNED(addr | size) ||
> > + !pfn_range_is_valid(hyp_phys_to_pfn(addr), size >> PAGE_SHIFT))
> > + return -EINVAL;
> > +
> > + ret = __pkvm_create_private_mapping(addr, size, PAGE_HYP_DEVICE, haddr);
>
> > + if (ret)
> > + return ret;
> > +
> > + host_lock_component();
> > + for (offset = 0; offset < size; offset += PAGE_SIZE) {
> > + if (addr_is_memory(addr + offset)) {
> > + ret = -EINVAL;
> > + goto unlock;
>
> If this fails we are left with the mapping inside the hyp because the unlock
> doesn't destroy the private mapping. Is this intended ?
>
Yes, there is no way to remove a private mapping at the moment, all
the callers to __pkvm_create_private_mapping() will leak it on failure.
> > + }
> > + ret = kvm_pgtable_get_leaf(&host_mmu.pgt, addr + offset, &pte, NULL);
> > + if (ret)
> > + goto unlock;
> > + if (pte && !kvm_pte_valid(pte)) {
> > + ret = -EPERM;
> > + goto unlock;
> > + }
> > + }
> > + /*
> > + * We set HYP as the owner of the MMIO pages in the host stage-2, for:
> > + * - host aborts: host_stage2_adjust_range() would fail for invalid non zero PTEs.
> > + * - recycle under memory pressure: host_stage2_unmap_dev_all() would call
> > + * kvm_pgtable_stage2_unmap() which will not clear non zero invalid ptes (counted).
> > + * - other MMIO donation: Would fail as we check that the PTE is valid or empty.
> > + */
> > + ret = host_stage2_try(kvm_pgtable_stage2_annotate, &host_mmu.pgt,
> > + addr, size, &host_s2_pool,
> > + KVM_HOST_INVALID_PTE_TYPE_DONATION,
> > + FIELD_PREP(KVM_HOST_DONATION_PTE_OWNER_MASK, PKVM_ID_HYP));
> > +unlock:
> > + host_unlock_component();
> > + return ret;
> > +}
> > +
> > +int __pkvm_hyp_donate_host_mmio(phys_addr_t addr, size_t size)
>
> This function seems to only update the host stage-2 annotation but it
> doesn't destroy the hyp mapping.
Yes, as mentioned above there is no way to destroy it, and this was
not desgined for frequent use. Typicaly, __pkvm_host_donate_hyp_mmio()
is called at boot per area/device. And __pkvm_hyp_donate_host_mmio()
is only used for failures.
>
> I was looking to make use of this patch in an upcoming posting for the
> v2 ITS hardening but in my case I don't need the private VA range
> creation.
I believe if you need to map MMIO, private range is the right way to
do it as the linear map was mainly designed around system memory.
Thanks,
Mostafa
More information about the linux-arm-kernel
mailing list