[PATCH] Derive the mem_section mem_map mask from SIZE(page)

HAGIO KAZUHITO(萩尾 一仁) kazuhito.hagio at nec.com
Wed Sep 23 17:43:30 PDT 2026


Hi Nikunj

On 2026/09/18 1:28, Nikunj Kela wrote:
> Hi Kazu,
> 
> Yes, that's fine. Please remove the "(see the BUG_ON ...)" when
> merging. The point I wanted to capture was just that masking with
> SECTION_MAP_MASK losslessly recovers the encoded mem_map (which that
> check enforces), but the reference is more confusing than helpful,
> especially now that it's a VM_WARN_ON_ONCE in
> sparse_init_one_section(). Thanks for testing and picking this up.

thanks, applied.
https://github.com/makedumpfile/makedumpfile/commit/cecce00ab4f3bab6b24f437d7a6d19bb8738081f

Thanks,
Kazu

> 
> Regards,
> -Nikunj
> 
> 
> On Thu, Sep 17, 2026 at 1:10 AM HAGIO KAZUHITO(萩尾 一仁)
> <kazuhito.hagio at nec.com> wrote:
>>
>> Hi Nikunj,
>>
>> thank you for the patch.  sorry for my delay.
>>
>> (My email will be changed to this From address soon.)
>>
>>> -----Original Message-----
>>> Linux 6.15 added a new mem_section flag bit, SECTION_IS_VMEMMAP_PREINIT
>>> (commit "mm/sparse: allow for alternate vmemmap section init at boot",
>>> Frank van der Linden), used on x86-64 to mark sections whose vmemmap
>>> was pre-initialized at boot for HugeTLB Vmemmap Optimization (HVO).
>>> With CONFIG_ZONE_DEVICE=y this is bit 5, and SECTION_MAP_LAST_BIT in
>>> the kernel moved to bit 6:
>>>
>>>      SECTION_MARKED_PRESENT_BIT,      /* 0 */
>>>      SECTION_HAS_MEM_MAP_BIT,         /* 1 */
>>>      SECTION_IS_ONLINE_BIT,           /* 2 */
>>>      SECTION_IS_EARLY_BIT,            /* 3 */
>>>      SECTION_TAINT_ZONE_DEVICE_BIT,   /* 4 */
>>>      SECTION_IS_VMEMMAP_PREINIT_BIT,  /* 5 */
>>>      SECTION_MAP_LAST_BIT,            /* 6 */
>>>
>>> makedumpfile still defines SECTION_MAP_LAST_BIT as 1UL<<5, so
>>> SECTION_MAP_MASK is ~0x1f and bit 5 survives when section_mem_map is
>>> decoded in section_mem_map_addr().  For every section that has
>>> SECTION_IS_VMEMMAP_PREINIT set, the resulting mem_map address is off
>>> by 0x20 -- half a struct page on 64-bit.  All struct page fields read
>>> from those sections are misaligned (page.flags reads page.index,
>>> page._mapcount reads the next page's lru.prev, ...), no exclusion
>>> filter matches, and the pages are kept as "dumpable kernel data".
>>> Nothing is reported, because the bogus address still passes
>>> is_kvaddr().
>>>
>>> The kernel sets this flag on every section backing a boot-time
>>> gigantic hugetlb page when HVO is enabled (hugetlb_vmemmap_init_early()).
>>> On the hosts with lot of hugepages reserved at boot-time, the core size
>>> grows significantly. It increase the time taken to generate the core as
>>> as the disk space needed to store it.
>>>
>>> Signed-off-by: Nikunj Kela <nkela at crusoe.ai>
>>> ---
>>> Tested:
>>>    - x86-64, Linux 6.17.13 (SPARSEMEM_VMEMMAP_PREINIT=y, ZONE_DEVICE=y),
>>>      1470 x 1 GiB hugepages, HVO enabled
>>>    - arm64, Linux 6.11, 512 MiB hugepages, HVO not enabled
>>>    - No 32-bit target available; the 32-bit case is a provable no-op
>>>      since mask is unchanged
>>>
>>>   makedumpfile.c | 16 +++++++++++++++-
>>>   makedumpfile.h |  5 +++++
>>>   2 files changed, 20 insertions(+), 1 deletion(-)
>>>
>>> diff --git a/makedumpfile.c b/makedumpfile.c
>>> index ba01137..be885f3 100644
>>> --- a/makedumpfile.c
>>> +++ b/makedumpfile.c
>>> @@ -3705,7 +3705,21 @@ section_mem_map_addr(unsigned long addr, unsigned long *map_mask)
>>>                return NOT_KV_ADDR;
>>>        }
>>>        map = ULONG(mem_section + OFFSET(mem_section.section_mem_map));
>>> -     mask = SECTION_MAP_MASK;
>>> +     /*
>>> +      * The encoded mem_map is struct page pointer arithmetic, so it is
>>> +      * always SIZE(page)-aligned, and the kernel guarantees that all
>>> +      * section flag bits fit below that alignment (see the BUG_ON in
>>> +      * sparse_encode_mem_map()).  Derive the mask from SIZE(page) instead
>>
>> The whole patch looks good to me and tested ok, but does the BUG_ON says
>> about the relation to that alignment?
>>
>>    BUG_ON(coded_mem_map & ~SECTION_MAP_MASK);
>>
>> (currently this is in sparse_init_one_section() as VM_WARN_ON_ONCE though.)
>>
>> This seems that all section flag bits fit below SECTION_MAP_LAST_BIT, not the
>> alignment.  Is it ok to remove the "(see the BUG_ON ...)" to avoid confusion?
>> If it's ok, I'll remove it when merging.
>>
>> Thanks,
>> Kazu
>>
>>> +      * of hardcoding the number of flag bits, which changed in Linux 4.13,
>>> +      * 5.x and again in 6.15 (SECTION_IS_VMEMMAP_PREINIT, bit 5 with
>>> +      * CONFIG_ZONE_DEVICE), and fall back to SECTION_MAP_MASK if SIZE(page)
>>> +      * is unknown or not a power of two.
>>> +      */
>>> +     if (SIZE(page) != NOT_FOUND_STRUCTURE && SIZE(page) > 0
>>> +         && !(SIZE(page) & (SIZE(page) - 1)))
>>> +             mask = ~((unsigned long)SIZE(page) - 1);
>>> +     else
>>> +             mask = SECTION_MAP_MASK;
>>>        *map_mask = map & ~mask;
>>>        map &= mask;
>>>        free(mem_section);
>>> diff --git a/makedumpfile.h b/makedumpfile.h
>>> index c4f3614..e2ab034 100644
>>> --- a/makedumpfile.h
>>> +++ b/makedumpfile.h
>>> @@ -195,6 +195,11 @@ test_bit(int nr, unsigned long addr)
>>>    *     version 4.12,
>>>    *  2. it has been verified that (1UL<<2) was never set, so it is
>>>    *     safe to mask that bit off even in old kernels.
>>> + *  3. since Linux 6.15 (SECTION_IS_VMEMMAP_PREINIT, bit 5 with
>>> + *     CONFIG_ZONE_DEVICE) there are six flag bits, so this value is
>>> + *     stale; section_mem_map_addr() derives the real mask from
>>> + *     SIZE(page) and only falls back to SECTION_MAP_MASK when that
>>> + *     is not possible.
>>>    */
>>>   #define SECTION_MAP_LAST_BIT (1UL<<5)
>>>   #define SECTION_MAP_MASK     (~(SECTION_MAP_LAST_BIT-1))
>>> --
>>> 2.54.0


More information about the kexec mailing list