[RFC PATCH 0/3] KVM: Dirty page logging for guest_memfd-only memslots

Sean Christopherson seanjc at google.com
Thu Jul 9 12:01:31 PDT 2026


On Thu, Jul 09, 2026, Mark Rutland wrote:
> Hi Sean,
> 
> Ignoring dirty page logging for the moment, I think you've raised a much
> bigger concern.
> 
> I've dumped a bit more context below, with a couple of high-level
> questions. The important thing for Alexandru and I is whether core folk
> are willing to consider some mechanism to ensure that guest PAs are
> pinned writeable and never fault (even transiently).
> 
> On Tue, Jul 07, 2026 at 10:12:41AM -0700, Sean Christopherson wrote:
> > On Tue, Jul 07, 2026, Alexandru Elisei wrote:
> > > On Mon, Jul 06, 2026 at 05:56:12PM -0700, Sean Christopherson wrote:
> > > > On Thu, Jul 02, 2026, Alexandru Elisei wrote:
> 
> > > > Please (publicly) document *why*  you want to add dirty-logging support.  It's
> > > > all but impossible to review new uAPI without knowing the use case.
> 
> > > As to why I'm working on it now, it's because of an arm64 feature that
> > > requires that memory remains mapped at stage 2, called Statistical
> > > Profiling Extension (SPE), similar to Intel's PEBS or AMD's IBS. Exposing
> > > the feature to a guest requires that memory remains mapped at stage 2
> > > outside of userspace explicitely unmapping it, and guest_memfd, with the
> > > patch to ignore the MMU notifiers [1], has this property.  I wanted to
> > > expand the functionality of guest_memfd to support migration of virtual
> > > machines when that arm64 feature is exposed to guests.
> > 
> > I'm all for adding dirty logging support for guest_memfd, but for SPE I don't
> > think relying on guest_memfd always being mapped is a good idea.  guest_memfd
> > is "pinned" purely because adding support for page migration is (very) low
> > priority for SNP, TDX, and pKVM.  guest_memfd page migration might play nice
> > with SPE?  Probably depends on whether KVM is forced to do break-before-make?
> 
> The key thing for SPE is that any pages that the SPE HW is using must
> have a valid writeable end-to-end (VA to real PA) mapping at all times
> while the guest is running (and while the host drains buffers). If that
> requirement is violated, even transiently, then data is lost and the
> guest will see an unexpected fault reported by SPE.

Well, at least it doesn't crash the host :-D

> If there's any reason we might (transiently) unmap a leaf entry (or
> entire sub-tree) from the stage-2 tables, or might (transiently) remove
> write permission, then we can't guarantee SPE will work correctly from
> the guest's PoV.
> 
> Obviously we can't guarantee that for regular memslots backed by
> userspace memory, hence we were hoping we could rely on guest_memfd.
> 
> Am I right to understand that we expect (in future) to do things with
> guest_memfd that could violate that? If so (and if we're not happy to
> have some options to say "always keep this pinned end-to-end no matter
> what"), 

Yes?  Page migration is the main one that I think will be problematic for arm64,
*unless* CPUs that support SPE don't require break-before-make (no idea if there
are even plans to ever drop the BBM rules).  Because with page migration, unless
KVM temporarily jails all vCPUs in the host while swizzling stage-2 PTEs, there
will be a small window of time where the memory isn't mapped.

I mention page migration because I expect swap/reclaim to be fully userspace
driven, i.e. if userspace crashes/corrupts the guest, the answer will be "well
don't do that".  But (at least some forms of) page migration will likely be
kernel-driven, e.g. to compact movable memory for THP.  And I really don't want
to give arch code the ability to hard-pin specific ranges of memory.

That said, odds are good that we'll end up with per-gmem flags to communicate to
guest_memfd whether or not the gmem instance supports page migration (x86's TDX
and SNP in particular require extra consideration).  So if the anticipated use
cases are fine with all-or-nothing "pinning", or with juggling guest_memfd files
in userspace if a more dynamic setup is desired, then you should be ok?

E.g. if the anticipated use cases are all slice-of-hardware style setups where
the VM will be statically assigned a chunk of memory, then for the most part this
will all Just Work.

> then I think that means that in practice we cannot virtualize
> SPE correctly, and Alexandru and I need to go back to our architects.
> 
> > And at some point guest_memfd may support userspace-driven swap, but I
> > suppose we can cross that bridge when we come to it.
> 
> Unfortunately, I think we need to figure out now whether it would be
> acceptable to suppress that (or making it mutually exclusive with SPE),
> if only to decide whether or not we continue trying to virtualize SPE.

Eh, as above, if userspace pulls a stupid and kills the guest, that's their
problem.  Just make sure the host isn't at risk :-)

> We don't need to figure out all the details; just whether or not that
> broad approach would be acceptable, or whether we have to give up on SPE
> virtualization.
> 
> > From a uABI perspective, forcing userspace to use guest_memfd to get access to
> > something like SPE isn't ideal.  While I have aspirations of using guest_memfd
> > much more broadly, I don't know that banking on guest_memfd replacing "everything"
> > is a winning strategy.
> 
> I agree this isn't ideal, and we're certainly not expecting that this is
> suitable for all users. We're just trying to get as much functionality
> as we practically can with today's SPE hardware.
> 
> Thanks,
> Mark.



More information about the linux-arm-kernel mailing list