[PATCH v4 3/6] KVM: arm64: Add auto DBM support for hardware dirty tracking
Tian Zheng
zhengtian10 at huawei.com
Sun Aug 2 18:33:02 PDT 2026
On 7/29/2026 11:16 PM, Leonardo Bras wrote:
> On Wed, Jul 29, 2026 at 04:51:59PM +0800, Tian Zheng wrote:
>>
>>
>> On 7/20/2026 8:58 PM, Leonardo Bras wrote:
>>> On Fri, Jul 17, 2026 at 04:21:32PM +0100, Leonardo Bras wrote:
>>>> On Fri, Jul 17, 2026 at 11:58:06AM +0800, Tian Zheng wrote:
>>>>>
>>>>> On 7/16/2026 3:39 PM, Oliver Upton wrote:
>>>>>> Hi Tian,
>>>>>>
>>>>>> On Thu, Jul 09, 2026 at 06:40:23PM +0800, Tian Zheng wrote:
>>>>>>> - if (prot & KVM_PGTABLE_PROT_W)
>>>>>>> + if (prot & KVM_PGTABLE_PROT_W) {
>>>>>>> set |= KVM_PTE_LEAF_ATTR_LO_S2_S2AP_W;
>>>>>>>
>>>>>>> + /*
>>>>>>> + * No DEVICE filter needed here: relax_perms is only called
>>>>>>> + * on FSC_PERM faults. Device pages always get full RW from
>>>>>>> + * initial mapping and are never write-protected during
>>>>>>> + * migration, so they never trigger a permission fault.
>>>>>>> + */
>>>>>>> + if (pgt->flags & KVM_PGTABLE_S2_DBM)
>>>>>>> + set |= KVM_PTE_LEAF_ATTR_HI_S2_DBM;
>>>>>>> + } else {
>>>>>>> + /*
>>>>>>> + * Clear DBM on W→RO downgrade to prevent hardware from
>>>>>>> + * silently upgrading RO+DBM back to W+dirty, which would
>>>>>>> + * bypass KVM's write tracking and cause data corruption.
>>>>>>> + */
>>>>>>> + clr |= KVM_PTE_LEAF_ATTR_HI_S2_DBM;
>>>>>>> + }
>>>>>>> +
>>>>>> This block makes it pretty evident that the DBM bit really *is* the
>>>>>> write permission bit. I'd much rather we introduce the concept of dirty
>>>>>> state to the page table library and migrate the abstract write
>>>>>> permission to the DBM field, even if we don't have FEAT_HAFDBS.
>>>>>>
>>>>
>>>> Ohh, that's an amazing idea!
>>>
>>> Thinking about that again...
>>> If we adopt the encoding with DBM being the write-permission bit, and all
>>> PTEs have it since the start, how can we have lazy-splitting happening?
>>>
>>> Only way I think of is removing both DBM and S2_S2AP_W bit from writable
>>> PTEs during dirty-track enable, and re-adding them during the first write
>>> fault. If we don't remove the DBM bit, systems with HDBSS would just dirty
>>> it by hardware, without causing a fault.
>>>
>>> DBM=0 would need to happen only in the first write-protect (only on
>>> lazy-splitting). All other write-protecting would just clean the S2_S2AP_W
>>> bit, as everything is already split.
>>>
>>> Is that what was intended?
>>>
>>> Thanks!
>>> Leo
>>>
>> Hi Leo,
>>
>> I think the cleanest way to handle this is to simply avoid setting DBM
>> on block mappings. If we only set DBM on page-level PTEs, then block
>> mappings will naturally stay DBM=0 and trigger a write fault on first
>> access — exactly what we need for lazy splitting.
>>
>> When the fault occurs, the block gets split into page-level PTEs, and at
>> that point we can set DBM=1 on the resulting leaf entries. This way:
>>
>> 1. Lazy split works naturally (fault -> split -> set DBM=1)
>>
>> 2. No need to clear DBM globally at dirty-track enable
>>
>> 3. No special handling for block mappings
>>
>> So I think global DBM is still viable — we just need to filter out block
>> mappings when setting the DBM bit. That way the lazy split path is preserved
>> without extra complexity.
>
> Hi Tian,
>
> Humm, but would not that be contrary to what Oliver suggested: changing the
> encoding from the PTE for all entries?
>
> (Like, if the PTE is writable, it has to have DBM set)
>
> IIUC what you said, on first faulting of the page in the VM:
> - If the entry is a page (level-3 leaf) and writable, add DBM
> - If it's a block entry (leaf but not a level-3), don't add DBM
>
> So after we enable dirty-logging:
> - a level-3 entry would not fault, using HDBSS, and
> - a block entry would fault, do the splitting, and add DBM to level-3
> entries during the split.
>
> If I got that correct, that would be clean indeed.
>
> But then we would have a different encoding for block entries and page
> entries. In page entries, DBM could be used to say if the page is writable,
> but on block entries one would have to look at the 'dirty-bit'.
>
> Would that be ok?
>
> Thanks!
> Leo
>
Hi Leo,
My initial concern was that clearing all DBM bits at the start of
migration would be too expensive, so I thought distinguishing between
level-3 entries and block entries would be better.
However, I ran a quick test on a 400GB VM (4 vCPUs), and the overhead
turned out to be around 30ns — which I think is acceptable.
So I think we can go with your approach: simply clear DBM globally in
kvm_arch_commit_memory_region() when dirty logging starts, before write-
protecting the memslot.
```
void kvm_arch_commit_memory_region(...)
{
// ...
if (log_dirty_pages) {
if (change == KVM_MR_DELETE)
return;
kvm_mmu_clear_dbm_memory_region(kvm, new->id);
// ...
}
// ...
}
```
Thanks!
Tian
More information about the linux-arm-kernel
mailing list