[PATCH] Derive the mem_section mem_map mask from SIZE(page)

HAGIO KAZUHITO(萩尾 一仁) kazuhito.hagio at nec.com
Thu Sep 17 01:10:11 PDT 2026


Hi Nikunj,

thank you for the patch.  sorry for my delay.

(My email will be changed to this From address soon.)

> -----Original Message-----
> Linux 6.15 added a new mem_section flag bit, SECTION_IS_VMEMMAP_PREINIT
> (commit "mm/sparse: allow for alternate vmemmap section init at boot",
> Frank van der Linden), used on x86-64 to mark sections whose vmemmap
> was pre-initialized at boot for HugeTLB Vmemmap Optimization (HVO).
> With CONFIG_ZONE_DEVICE=y this is bit 5, and SECTION_MAP_LAST_BIT in
> the kernel moved to bit 6:
> 
>     SECTION_MARKED_PRESENT_BIT,      /* 0 */
>     SECTION_HAS_MEM_MAP_BIT,         /* 1 */
>     SECTION_IS_ONLINE_BIT,           /* 2 */
>     SECTION_IS_EARLY_BIT,            /* 3 */
>     SECTION_TAINT_ZONE_DEVICE_BIT,   /* 4 */
>     SECTION_IS_VMEMMAP_PREINIT_BIT,  /* 5 */
>     SECTION_MAP_LAST_BIT,            /* 6 */
> 
> makedumpfile still defines SECTION_MAP_LAST_BIT as 1UL<<5, so
> SECTION_MAP_MASK is ~0x1f and bit 5 survives when section_mem_map is
> decoded in section_mem_map_addr().  For every section that has
> SECTION_IS_VMEMMAP_PREINIT set, the resulting mem_map address is off
> by 0x20 -- half a struct page on 64-bit.  All struct page fields read
> from those sections are misaligned (page.flags reads page.index,
> page._mapcount reads the next page's lru.prev, ...), no exclusion
> filter matches, and the pages are kept as "dumpable kernel data".
> Nothing is reported, because the bogus address still passes
> is_kvaddr().
> 
> The kernel sets this flag on every section backing a boot-time
> gigantic hugetlb page when HVO is enabled (hugetlb_vmemmap_init_early()).
> On the hosts with lot of hugepages reserved at boot-time, the core size
> grows significantly. It increase the time taken to generate the core as
> as the disk space needed to store it.
> 
> Signed-off-by: Nikunj Kela <nkela at crusoe.ai>
> ---
> Tested:
>   - x86-64, Linux 6.17.13 (SPARSEMEM_VMEMMAP_PREINIT=y, ZONE_DEVICE=y),
>     1470 x 1 GiB hugepages, HVO enabled
>   - arm64, Linux 6.11, 512 MiB hugepages, HVO not enabled
>   - No 32-bit target available; the 32-bit case is a provable no-op
>     since mask is unchanged
> 
>  makedumpfile.c | 16 +++++++++++++++-
>  makedumpfile.h |  5 +++++
>  2 files changed, 20 insertions(+), 1 deletion(-)
> 
> diff --git a/makedumpfile.c b/makedumpfile.c
> index ba01137..be885f3 100644
> --- a/makedumpfile.c
> +++ b/makedumpfile.c
> @@ -3705,7 +3705,21 @@ section_mem_map_addr(unsigned long addr, unsigned long *map_mask)
>  		return NOT_KV_ADDR;
>  	}
>  	map = ULONG(mem_section + OFFSET(mem_section.section_mem_map));
> -	mask = SECTION_MAP_MASK;
> +	/*
> +	 * The encoded mem_map is struct page pointer arithmetic, so it is
> +	 * always SIZE(page)-aligned, and the kernel guarantees that all
> +	 * section flag bits fit below that alignment (see the BUG_ON in
> +	 * sparse_encode_mem_map()).  Derive the mask from SIZE(page) instead

The whole patch looks good to me and tested ok, but does the BUG_ON says
about the relation to that alignment? 

  BUG_ON(coded_mem_map & ~SECTION_MAP_MASK);

(currently this is in sparse_init_one_section() as VM_WARN_ON_ONCE though.)

This seems that all section flag bits fit below SECTION_MAP_LAST_BIT, not the
alignment.  Is it ok to remove the "(see the BUG_ON ...)" to avoid confusion?
If it's ok, I'll remove it when merging.

Thanks,
Kazu

> +	 * of hardcoding the number of flag bits, which changed in Linux 4.13,
> +	 * 5.x and again in 6.15 (SECTION_IS_VMEMMAP_PREINIT, bit 5 with
> +	 * CONFIG_ZONE_DEVICE), and fall back to SECTION_MAP_MASK if SIZE(page)
> +	 * is unknown or not a power of two.
> +	 */
> +	if (SIZE(page) != NOT_FOUND_STRUCTURE && SIZE(page) > 0
> +	    && !(SIZE(page) & (SIZE(page) - 1)))
> +		mask = ~((unsigned long)SIZE(page) - 1);
> +	else
> +		mask = SECTION_MAP_MASK;
>  	*map_mask = map & ~mask;
>  	map &= mask;
>  	free(mem_section);
> diff --git a/makedumpfile.h b/makedumpfile.h
> index c4f3614..e2ab034 100644
> --- a/makedumpfile.h
> +++ b/makedumpfile.h
> @@ -195,6 +195,11 @@ test_bit(int nr, unsigned long addr)
>   *     version 4.12,
>   *  2. it has been verified that (1UL<<2) was never set, so it is
>   *     safe to mask that bit off even in old kernels.
> + *  3. since Linux 6.15 (SECTION_IS_VMEMMAP_PREINIT, bit 5 with
> + *     CONFIG_ZONE_DEVICE) there are six flag bits, so this value is
> + *     stale; section_mem_map_addr() derives the real mask from
> + *     SIZE(page) and only falls back to SECTION_MAP_MASK when that
> + *     is not possible.
>   */
>  #define SECTION_MAP_LAST_BIT	(1UL<<5)
>  #define SECTION_MAP_MASK	(~(SECTION_MAP_LAST_BIT-1))
> --
> 2.54.0
-------------- next part --------------
A non-text attachment was scrubbed...
Name: smime.p7s
Type: application/pkcs7-signature
Size: 7878 bytes
Desc: not available
URL: <http://lists.infradead.org/pipermail/kexec/attachments/20260917/e5f9ddba/attachment-0001.p7s>


More information about the kexec mailing list