[RFC PATCH 1/3] KVM: guest_memfd: Use memslot id to keep track of associated memslots
David Hildenbrand
david.hildenbrand at arm.com
Mon Jul 13 07:03:18 PDT 2026
On 7/7/26 19:05, Alexandru Elisei wrote:
> Hi Sean,
>
> On Mon, Jul 06, 2026 at 02:43:23PM -0700, Sean Christopherson wrote:
>> On Thu, Jul 02, 2026, Alexandru Elisei wrote:
>>> To enable memslot operations, KVM maintains two arrays of memslots, and an
>>> RCU pointer to the active (in use) array. Changes are made first to the
>>> inactive array, and the RCU pointer is updated to point to the inactive
>>> array, which becomes active.
>>>
>>> The guest_memfd file maintains an xarray of pointers to memslots that use
>>> it as the memory provider. After the RCU pointer to the active memslots is
>>> updated and until SRCU is synchronized, readers can observe the old or the
>>> new value for the active array, and therefore the old or the new pointer
>>> for a given memslot. For memslot creation or deletion that is not an issue
>>> for guest_memfd, as readers will either read the same memslot pointer saved
>>> by the guest_memfd file, or a non-existing memslot.
>>>
>>> But when changing the flags for a memslot, readers can read two different
>>> and non-NULL memslot pointers.
>>
>> And? Why does that matter? KVM memslot updates aren't atomic. Practically
>> speaking, they _can't_ be made atomic. Userspace is required to quiesce all
>> activity that must not observe inconsistent state, i.e. userspace must pause
>> (stop running) vCPUs when performing a memslot update.
>
> Is that true when KVM_MEM_LOG_DIRTY_PAGES is toggled for a memslot?
Good point.
>
> As far as I can tell, KVM today tolerates VCPUs running while the
> KVM_MEM_LOG_DIRTY_PAGES flags is being changed for a memslot. And by
> tolarate I mean that VCPUs that are running when the flag is changed don't
> return an error from KVM_RUN. If changing a memslot while VCPUs are running
> were fatal, I would think that KVM would want to take vcpu->mutex for all
> VCPUs to keep them from running. Or is it a case of KVM allowing userspace
> to shoot themselves in the foot if they really want it?
>
> When the KVM_MEM_LOG_DIRTY_PAGES flags is being changed, VCPUs handling a
> guest fault can observe either the old memslot, with the old flags, or the
> new memslot, with the flag changed, but they still continue running without
> returning an error.
Staring at QEMU, kvm_log_start()+kvm_log_stop() do not call
accel_ioctl_inhibit_begin() etc.
So at least QEMU does not force VCPUs out of KVM when only updating flags.
One option would be to require user space to do that also when starting+stopping
dirty page logging. (I'd assume that should work, but it might be tricky
depending on in which context it is called from QEMU migration code -- whether
we hold the BQL)
--
Cheers,
David
More information about the linux-arm-kernel
mailing list