[PATCH v3 1/3] KVM: riscv: Implement KVM_PRE_FAULT_MEMORY

Anup Patel anup at brainfault.org
Sat Oct 3 03:01:43 PDT 2026


On Fri, Aug 14, 2026 at 5:15 PM Jinyu Tang <jinyu.tang at linux.dev> wrote:
>
> The generic KVM_PRE_FAULT_MEMORY ioctl lets userspace populate KVM page
> tables before running a vCPU over a GPA range. x86 already supports the
> ioctl, but RISC-V does not expose the capability and has no arch hook.
>
> Add the RISC-V arch hook and reuse the existing G-stage fault mapping
> path with a read access. Report progress using the G-stage mapping
> returned by the map path, so the ioctl can advance by the actual leaf
> size that covers the requested GPA. Retry until a mapping is installed
> or a signal, VM-dead request, or real error is observed.
>
> Signed-off-by: Jinyu Tang <jinyu.tang at linux.dev>
> ---
> v2 -> v3:
> - Retry internally when the map path returns success without a visible
>   G-stage mapping, instead of exposing -EAGAIN.  (Sashiko)
>
>  arch/riscv/kvm/Kconfig  |  1 +
>  arch/riscv/kvm/gstage.c |  3 +++
>  arch/riscv/kvm/mmu.c    | 42 +++++++++++++++++++++++++++++++++++++++++
>  arch/riscv/kvm/vm.c     |  1 +
>  4 files changed, 47 insertions(+)
>
> diff --git a/arch/riscv/kvm/Kconfig b/arch/riscv/kvm/Kconfig
> index ec2cee0a39e0..8ac209e8ac87 100644
> --- a/arch/riscv/kvm/Kconfig
> +++ b/arch/riscv/kvm/Kconfig
> @@ -28,6 +28,7 @@ config KVM
>         select KVM_COMMON
>         select KVM_GENERIC_DIRTYLOG_READ_PROTECT
>         select KVM_GENERIC_HARDWARE_ENABLING
> +       select KVM_GENERIC_PRE_FAULT_MEMORY
>         select KVM_MMIO
>         select VIRT_XFER_TO_GUEST_WORK
>         select SCHED_INFO
> diff --git a/arch/riscv/kvm/gstage.c b/arch/riscv/kvm/gstage.c
> index e5002cb9cbef..6bd8b8fd6ceb 100644
> --- a/arch/riscv/kvm/gstage.c
> +++ b/arch/riscv/kvm/gstage.c
> @@ -280,6 +280,9 @@ int kvm_riscv_gstage_map_page(struct kvm_gstage *gstage,
>                                                     out_map->level, true);
>                 } else if (ALIGN_DOWN(PFN_PHYS(pte_pfn(ptep_get(ptep))), page_size) == hpa) {
>                         kvm_riscv_gstage_update_pte_prot(gstage, ptep_level, gpa, ptep, prot);
> +                       out_map->addr = ALIGN_DOWN(gpa, page_size);

This ALIGN_DOWN() must be done for all cases at the start itself
right after gstage_page_size_to_level(). I will take care of this at the
time of merging this patch.

> +                       out_map->level = ptep_level;
> +                       out_map->pte = ptep_get(ptep);
>                         return 0;
>                 }
>         }
> diff --git a/arch/riscv/kvm/mmu.c b/arch/riscv/kvm/mmu.c
> index 6035b5ec9503..33d4ba406b0d 100644
> --- a/arch/riscv/kvm/mmu.c
> +++ b/arch/riscv/kvm/mmu.c
> @@ -748,6 +748,48 @@ int kvm_riscv_mmu_map(struct kvm_vcpu *vcpu, struct kvm_memory_slot *memslot,
>         return ret;
>  }
>
> +long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
> +                                   struct kvm_pre_fault_memory *range)
> +{
> +       struct kvm_gstage_mapping out_map = { 0 };
> +       struct kvm_memory_slot *memslot;
> +       unsigned long map_size;
> +       unsigned long hva;
> +       gpa_t end;
> +       gfn_t gfn;
> +       int ret;
> +
> +       gfn = gpa_to_gfn(range->gpa);
> +       memslot = kvm_vcpu_gfn_to_memslot(vcpu, gfn);
> +       if (!memslot)
> +               return -ENOENT;
> +
> +       hva = gfn_to_hva_memslot_prot(memslot, gfn, NULL);
> +       if (kvm_is_error_hva(hva))
> +               return -ENOENT;
> +
> +       for (;;) {
> +               if (signal_pending(current))
> +                       return -EINTR;
> +
> +               if (kvm_check_request(KVM_REQ_VM_DEAD, vcpu))
> +                       return -EIO;

The kvm_check_request() does not allow KVM_REQ_VM_DEAD because
this request can't be cleared so we have to use kvm_test_request() over
here. I will update it at the time of merging this patch.

> +
> +               cond_resched();
> +               ret = kvm_riscv_mmu_map(vcpu, memslot, range->gpa, hva, false, &out_map);
> +               if (ret)
> +                       return ret;
> +
> +               if (pte_val(out_map.pte))
> +                       break;
> +       }
> +
> +       map_size = PAGE_SIZE << (out_map.level * kvm_riscv_gstage_index_bits);
> +       end = ALIGN_DOWN(range->gpa, map_size) + map_size;
> +
> +       return min_t(u64, range->size, end - range->gpa);
> +}
> +
>  int kvm_riscv_mmu_alloc_pgd(struct kvm *kvm)
>  {
>         struct page *pgd_page;
> diff --git a/arch/riscv/kvm/vm.c b/arch/riscv/kvm/vm.c
> index a9f083feeb76..58500a19b33b 100644
> --- a/arch/riscv/kvm/vm.c
> +++ b/arch/riscv/kvm/vm.c
> @@ -187,6 +187,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext)
>         case KVM_CAP_MP_STATE:
>         case KVM_CAP_IMMEDIATE_EXIT:
>         case KVM_CAP_SET_GUEST_DEBUG:
> +       case KVM_CAP_PRE_FAULT_MEMORY:
>                 r = 1;
>                 break;
>         case KVM_CAP_NR_VCPUS:
> --
> 2.43.0

Otherwise, it looks good to me.

Reviewed-by: Anup Patel <anup at brainfault.org>

Queued this patch for Linux-7.4

Thanks,
Anup



More information about the linux-riscv mailing list