[PATCH] Derive the mem_section mem_map mask from SIZE(page)
Nikunj Kela
nkela at crusoe.ai
Thu Sep 17 09:28:38 PDT 2026
Hi Kazu,
Yes, that's fine. Please remove the "(see the BUG_ON ...)" when
merging. The point I wanted to capture was just that masking with
SECTION_MAP_MASK losslessly recovers the encoded mem_map (which that
check enforces), but the reference is more confusing than helpful,
especially now that it's a VM_WARN_ON_ONCE in
sparse_init_one_section(). Thanks for testing and picking this up.
Regards,
-Nikunj
On Thu, Sep 17, 2026 at 1:10 AM HAGIO KAZUHITO(萩尾 一仁)
<kazuhito.hagio at nec.com> wrote:
>
> Hi Nikunj,
>
> thank you for the patch. sorry for my delay.
>
> (My email will be changed to this From address soon.)
>
> > -----Original Message-----
> > Linux 6.15 added a new mem_section flag bit, SECTION_IS_VMEMMAP_PREINIT
> > (commit "mm/sparse: allow for alternate vmemmap section init at boot",
> > Frank van der Linden), used on x86-64 to mark sections whose vmemmap
> > was pre-initialized at boot for HugeTLB Vmemmap Optimization (HVO).
> > With CONFIG_ZONE_DEVICE=y this is bit 5, and SECTION_MAP_LAST_BIT in
> > the kernel moved to bit 6:
> >
> > SECTION_MARKED_PRESENT_BIT, /* 0 */
> > SECTION_HAS_MEM_MAP_BIT, /* 1 */
> > SECTION_IS_ONLINE_BIT, /* 2 */
> > SECTION_IS_EARLY_BIT, /* 3 */
> > SECTION_TAINT_ZONE_DEVICE_BIT, /* 4 */
> > SECTION_IS_VMEMMAP_PREINIT_BIT, /* 5 */
> > SECTION_MAP_LAST_BIT, /* 6 */
> >
> > makedumpfile still defines SECTION_MAP_LAST_BIT as 1UL<<5, so
> > SECTION_MAP_MASK is ~0x1f and bit 5 survives when section_mem_map is
> > decoded in section_mem_map_addr(). For every section that has
> > SECTION_IS_VMEMMAP_PREINIT set, the resulting mem_map address is off
> > by 0x20 -- half a struct page on 64-bit. All struct page fields read
> > from those sections are misaligned (page.flags reads page.index,
> > page._mapcount reads the next page's lru.prev, ...), no exclusion
> > filter matches, and the pages are kept as "dumpable kernel data".
> > Nothing is reported, because the bogus address still passes
> > is_kvaddr().
> >
> > The kernel sets this flag on every section backing a boot-time
> > gigantic hugetlb page when HVO is enabled (hugetlb_vmemmap_init_early()).
> > On the hosts with lot of hugepages reserved at boot-time, the core size
> > grows significantly. It increase the time taken to generate the core as
> > as the disk space needed to store it.
> >
> > Signed-off-by: Nikunj Kela <nkela at crusoe.ai>
> > ---
> > Tested:
> > - x86-64, Linux 6.17.13 (SPARSEMEM_VMEMMAP_PREINIT=y, ZONE_DEVICE=y),
> > 1470 x 1 GiB hugepages, HVO enabled
> > - arm64, Linux 6.11, 512 MiB hugepages, HVO not enabled
> > - No 32-bit target available; the 32-bit case is a provable no-op
> > since mask is unchanged
> >
> > makedumpfile.c | 16 +++++++++++++++-
> > makedumpfile.h | 5 +++++
> > 2 files changed, 20 insertions(+), 1 deletion(-)
> >
> > diff --git a/makedumpfile.c b/makedumpfile.c
> > index ba01137..be885f3 100644
> > --- a/makedumpfile.c
> > +++ b/makedumpfile.c
> > @@ -3705,7 +3705,21 @@ section_mem_map_addr(unsigned long addr, unsigned long *map_mask)
> > return NOT_KV_ADDR;
> > }
> > map = ULONG(mem_section + OFFSET(mem_section.section_mem_map));
> > - mask = SECTION_MAP_MASK;
> > + /*
> > + * The encoded mem_map is struct page pointer arithmetic, so it is
> > + * always SIZE(page)-aligned, and the kernel guarantees that all
> > + * section flag bits fit below that alignment (see the BUG_ON in
> > + * sparse_encode_mem_map()). Derive the mask from SIZE(page) instead
>
> The whole patch looks good to me and tested ok, but does the BUG_ON says
> about the relation to that alignment?
>
> BUG_ON(coded_mem_map & ~SECTION_MAP_MASK);
>
> (currently this is in sparse_init_one_section() as VM_WARN_ON_ONCE though.)
>
> This seems that all section flag bits fit below SECTION_MAP_LAST_BIT, not the
> alignment. Is it ok to remove the "(see the BUG_ON ...)" to avoid confusion?
> If it's ok, I'll remove it when merging.
>
> Thanks,
> Kazu
>
> > + * of hardcoding the number of flag bits, which changed in Linux 4.13,
> > + * 5.x and again in 6.15 (SECTION_IS_VMEMMAP_PREINIT, bit 5 with
> > + * CONFIG_ZONE_DEVICE), and fall back to SECTION_MAP_MASK if SIZE(page)
> > + * is unknown or not a power of two.
> > + */
> > + if (SIZE(page) != NOT_FOUND_STRUCTURE && SIZE(page) > 0
> > + && !(SIZE(page) & (SIZE(page) - 1)))
> > + mask = ~((unsigned long)SIZE(page) - 1);
> > + else
> > + mask = SECTION_MAP_MASK;
> > *map_mask = map & ~mask;
> > map &= mask;
> > free(mem_section);
> > diff --git a/makedumpfile.h b/makedumpfile.h
> > index c4f3614..e2ab034 100644
> > --- a/makedumpfile.h
> > +++ b/makedumpfile.h
> > @@ -195,6 +195,11 @@ test_bit(int nr, unsigned long addr)
> > * version 4.12,
> > * 2. it has been verified that (1UL<<2) was never set, so it is
> > * safe to mask that bit off even in old kernels.
> > + * 3. since Linux 6.15 (SECTION_IS_VMEMMAP_PREINIT, bit 5 with
> > + * CONFIG_ZONE_DEVICE) there are six flag bits, so this value is
> > + * stale; section_mem_map_addr() derives the real mask from
> > + * SIZE(page) and only falls back to SECTION_MAP_MASK when that
> > + * is not possible.
> > */
> > #define SECTION_MAP_LAST_BIT (1UL<<5)
> > #define SECTION_MAP_MASK (~(SECTION_MAP_LAST_BIT-1))
> > --
> > 2.54.0
More information about the kexec
mailing list