[RFC PATCH 0/3] KVM: Dirty page logging for guest_memfd-only memslots

Sean Christopherson seanjc at google.com
Thu Aug 13 16:32:55 PDT 2026


On Thu, Aug 13, 2026, David Hildenbrand wrote:
> On 8/13/26 18:09, Alexandru Elisei wrote:
> > On Mon, Jul 13, 2026 at 04:11:57PM +0200, David Hildenbrand wrote:
> >>> Yeah. I agree (with the caveats you mention). Arm folk need to go figure
> >>> out if it's worthwhile to support with all those caveats.
> >>
> >> Right, disabling migration (once gmem supports it) was also what I discussed
> >> with Alexandru when that topic comes up.
> >>
> >> How to communicate to gmem that it wants these fixed mappings is a good question.
> > 
> > There's already a proposal for how to do this in the migratable guest_memfd
> > series [1] - it's a new guest_memfd creation flag that disables migration.
> 
> In the guest_memfd call I was arguing against the flag in the first version, and
> instead adding it when actually required.

Ya, right now a flag is meaningless.  Telling guest_memfd not to do something it
doesn't ever do...

> I was also raising whether KVM couldn't tell guest_memfd (e.g., at creation
> time?) that it supports a CPU feature that requires S2 to be always mapped to
> disable migration.
> 
> It would then be a contract between KVM and guest_memfd without user space
> having to be involved on that level.
> 
> > With Sean's comment that he expects swap/reclaim to be fully userspace
> > driven, I believe that would be enough to guarantee on the _kernel_ side
> > that SPE will work as intended for a guest.
> 
> That's my understanding.
> 
> > 
> > I'm a slightly concerned though that all of this will work by chance, and
> > not by design, and in the future the behaviour might change to allow
> > guest_memfd memory to be unmapped from stage 2 without the VMM or KVM
> > explicitly allowing it or initiating it.

Meh, TDX on x86 already has the same requirement.  Unmapping a page from the S-EPT
kills the VM unless the VM was expecting the page to be lost. 

> Thus my idea of the explicit contract between KVM and guest_memfd. Instead of
> being a "this doesn't support migration" it would be a "S2 always mapped"
> kind-of contract.

Who would that contract be between though?  KVM can tell a guest_memfd instance
that page migration is/isn't supported, but telling guest_memfd that the VM will
always keep the entirety of the guest_memfd mapped in stage-2 is nonsensical.
guest_memfd simply doesn't care if the page is mapped or not, it only cares if
migration is supported.

> > My understanding from the conversation so far is that the plan for the
> > future of guest_memfd is to support an option/mode where the memory is
> > effectively "pinned" at stage 2 (but which allows userspace to explicitly
> > free/unmap it, of course). Is that correct, or am I being overly optimistic
> > in my interpretation?
> 
> We could then even disallow fallocate() to punch holes if that contract is
> negotiated.

Why?  If userspace pulls a stupid and kills its guest, that's userspace's problem.

KVM would also have to block memslot changes, and probably other things in the
future that would unmap stage-2 in response to userspace syscalls/ioctls.



More information about the linux-arm-kernel mailing list