[PATCH v3 1/2] x86/crash: reserve elfcorehdr for CONFIG_NR_CPUS, not CONFIG_NR_CPUS_DEFAULT

Ionut Nechita (Wind River) ionut.nechita at windriver.com
Wed Aug 26 00:35:26 PDT 2026


From: Ionut Nechita <ionut.nechita at windriver.com>

kexec_file_load(2) fails with -EINVAL when loading a crash kernel on a
machine whose number of possible CPUs exceeds CONFIG_NR_CPUS_DEFAULT.

With CONFIG_CRASH_HOTPLUG=y the elfcorehdr segment is over-allocated so
it can be updated in place on CPU/memory hotplug.  On the
!CONFIG_MEMORY_HOTPLUG path, crash_load_segments() sizes that
reservation as:

	ret = crash_prepare_headers(..., &kbuf.bufsz, &pnum);
	...
	pnum += 2 + CONFIG_NR_CPUS_DEFAULT;

The value that lands in @pnum is crash_prepare_headers()'s
@nr_mem_ranges out parameter, i.e. cmem->nr_ranges - the number of
memory ranges only, not a phdr count.  The header that
crash_prepare_elf64_headers() actually builds adds one phdr per
*possible* CPU on top of those ranges:

	nr_phdr = nr_cpus + 1;		/* + vmcoreinfo */
	nr_phdr += mem->nr_ranges;
	nr_phdr++;			/* + kernel text map */

So the reservation covers

	nr_ranges + 2 + CONFIG_NR_CPUS_DEFAULT

phdrs while the buffer holds

	nr_ranges + 2 + num_possible_cpus()

phdrs, and the buffer exceeds the reservation by

	(num_possible_cpus() - CONFIG_NR_CPUS_DEFAULT) * sizeof(Elf64_Phdr)

bytes as soon as num_possible_cpus() grows past CONFIG_NR_CPUS_DEFAULT.
num_possible_cpus() is bounded by CONFIG_NR_CPUS, not by
CONFIG_NR_CPUS_DEFAULT, so this is reachable on any config that raises
CONFIG_NR_CPUS above the arch default without CONFIG_MAXSMP.

crash_prepare_elf64_headers() rounds bufsz up to ELF_CORE_HEADER_ALIGN
and kexec_add_buffer() rounds memsz up to PAGE_SIZE (both 4096), so the
excess is invisible until it outgrows that padding.  Once it does,
sanity_check_segment_list() rejects the image:

	if (image->segment[i].bufsz > image->segment[i].memsz)
		return -EINVAL;

kexec_load(2) does not report an error on the same machine, but it is
not unaffected either.  With crash hotplug enabled, kexec-tools sizes
the elfcorehdr segment from /sys/kernel/crash_elfcorehdr_size in
load_crashdump_segments(), i.e. from arch_crash_get_elfcorehdr_size(),
which derives its answer from the very same CONFIG_NR_CPUS_DEFAULT.
add_segment_phys_virt() then clamps the buffer to that segment:

	if (bufsz > memsz) {
		bufsz = memsz;
	}

so on the kexec_load(2) path the undersized reservation silently
truncates the elfcorehdr instead of failing the load, which surfaces
later as a bad or unusable dump rather than as a load failure.
Switching arch_crash_get_elfcorehdr_size() to CONFIG_NR_CPUS below fixes
that path as well.

Observed on a single-socket Xeon 6776P (144 possible CPUs) running a
PREEMPT_RT kernel with:

	# CONFIG_MAXSMP is not set
	# CONFIG_MEMORY_HOTPLUG is not set
	CONFIG_NR_CPUS_RANGE_BEGIN=2
	CONFIG_NR_CPUS_RANGE_END=512
	CONFIG_NR_CPUS_DEFAULT=64
	CONFIG_NR_CPUS=256

At 144 possible CPUs the buffer exceeds the reservation by
(144 - 64) * 56 = 4480 bytes.  That is more than the 4096 bytes of page
padding, so the overflow is guaranteed and kexec -p -s fails with
"kexec_file_load failed: Invalid argument".  Reducing the possible CPU
count to 72 leaves an excess of (72 - 64) * 56 = 448 bytes, which the
page rounding still absorbs, and the load succeeds - confirming the
reservation is the limiting factor.

Reserve the elfcorehdr for CONFIG_NR_CPUS, the compile-time upper bound
of num_possible_cpus(), so the reservation always covers the header that
is actually generated.

The CONFIG_MEMORY_HOTPLUG=y path discards @pnum and reserves
2 + CONFIG_NR_CPUS_DEFAULT + CONFIG_CRASH_MAX_MEMORY_RANGES phdrs
instead.  With the default CONFIG_CRASH_MAX_MEMORY_RANGES=8192 the
memory range allowance dwarfs the CPU shortfall, so that path does not
fail in practice; it is switched to CONFIG_NR_CPUS as well for
consistency and to stay correct for small CONFIG_CRASH_MAX_MEMORY_RANGES
values.

This does not change the reservation for defconfig-like builds, since
CONFIG_NR_CPUS defaults to CONFIG_NR_CPUS_DEFAULT.  Only configs that
raise CONFIG_NR_CPUS reserve more, and the worst case is bounded by the
top of the range (CONFIG_NR_CPUS=8192 with CONFIG_CPUMASK_OFFSTACK=y),
which is exactly what CONFIG_MAXSMP already reserves today.

Fixes: ea53ad9cf73b ("x86/crash: add x86 crash hotplug support")
Assisted-by: LLM
Signed-off-by: Ionut Nechita <ionut.nechita at windriver.com>
Reviewed-by: Jinjie Ruan <ruanjinjie at huawei.com>
Reviewed-by: Bradley Morgan <include at grrlz.net>
Reviewed-by: Sourabh Jain <sourabhjain at linux.ibm.com>
Acked-by: Baoquan He <baoquan.he at linux.dev>
---
 arch/x86/kernel/crash.c | 6 +++---
 1 file changed, 3 insertions(+), 3 deletions(-)

diff --git a/arch/x86/kernel/crash.c b/arch/x86/kernel/crash.c
index e681ec9cf1dc..e6f23933a6df 100644
--- a/arch/x86/kernel/crash.c
+++ b/arch/x86/kernel/crash.c
@@ -369,9 +369,9 @@ int crash_load_segments(struct kimage *image)
 	 * maximum CPUs and maximum memory ranges.
 	 */
 	if (IS_ENABLED(CONFIG_MEMORY_HOTPLUG))
-		pnum = 2 + CONFIG_NR_CPUS_DEFAULT + CONFIG_CRASH_MAX_MEMORY_RANGES;
+		pnum = 2 + CONFIG_NR_CPUS + CONFIG_CRASH_MAX_MEMORY_RANGES;
 	else
-		pnum += 2 + CONFIG_NR_CPUS_DEFAULT;
+		pnum += 2 + CONFIG_NR_CPUS;
 
 	if (pnum < (unsigned long)PN_XNUM) {
 		kbuf.memsz = pnum * sizeof(Elf64_Phdr);
@@ -430,7 +430,7 @@ unsigned int arch_crash_get_elfcorehdr_size(void)
 	unsigned int sz;
 
 	/* kernel_map, VMCOREINFO and maximum CPUs */
-	sz = 2 + CONFIG_NR_CPUS_DEFAULT;
+	sz = 2 + CONFIG_NR_CPUS;
 	if (IS_ENABLED(CONFIG_MEMORY_HOTPLUG))
 		sz += CONFIG_CRASH_MAX_MEMORY_RANGES;
 	sz *= sizeof(Elf64_Phdr);

base-commit: a8406e6c0b793ce0788019683837c40855b55995
-- 
2.55.0




More information about the kexec mailing list