[PATCH] kho: fix global scratch size calculation
Sourabh Jain
sourabhjain at linux.ibm.com
Sat Sep 26 07:46:33 PDT 2026
Hello Mike,
On 26/09/26 14:36, Mike Rapoport wrote:
> Hi Sourabh,
>
> On Tue, Sep 22, 2026 at 06:42:16PM +0530, Sourabh Jain wrote:
>> KHO calculates the global scratch size based on memblock-reserved kernel
>> memory. It passes NUMA_NO_NODE to memblock_reserved_kern_size() for
>> this calculation.
>>
>> When memblock_reserved_kern_size() is called with NUMA_NO_NODE, it
>> counts both:
>>
>> - memory reserved for a specific NUMA node
>> - memory reserved with NUMA_NO_NODE
>>
>> KHO needs to distinguish between these two types of reservations.
>> When calculating the size of global scratch memory, KHO only needs to
>> account for reservations made with NUMA_NO_NODE. Reservations made for
>> a specific NUMA node must not be included in the global scratch size.
>>
>> Add memblock_reserved_size_nid() to calculate reserved memory for a
>> given reservation type and NUMA node. When NUMA_NO_NODE is passed, it
>> counts only memory reserved with NUMA_NO_NODE.
>>
>> Use the new API for lowmem, global, and per-node KHO scratch size
>> calculations. For lowmem and global scratch, count only memory
>> reservations that were made with NUMA_NO_NODE. For per-node scratch,
>> count only memory reservations that were made with the corresponding
>> NUMA node ID.
>>
>> Remove memblock_reserved_hugetlb_size() since it has the same
>> implementation as the new API and differs only in the memblock
>> reservation flag being checked. The new API handles both kernel and
>> HugeTLB reservations through its reservation type argument.
>>
>> Define the new helper as a static function in the KHO implementation,
>> since it is only used by KHO and has no users outside
>> kernel/liveupdate/kexec_handover.c.
>>
>> On powerpc, the difference can be seen in the scratch_len values
>> reported by:
>>
>> cat /sys/kernel/debug/kho/out/scratch_len
>>
>> Before this change, the global scratch allocation was 0x12000000
>> (288 MB):
>>
>> 0x1000000 (16 MB)
>> 0x12000000 (288 MB) <- global allocation
>> 0x5000000 (80 MB)
>>
>> After this change, the global scratch allocation is 0xd000000
>> (208 MB):
>>
>> 0x1000000 (16 MB)
>> 0xd000000 (208 MB) <- global allocation
>> 0x5000000 (80 MB)
>>
>> The 80 MB difference is the per-node reservation that was previously
>> being included in the global allocation.
>>
>> The same issue also affects lowmem scratch memory, but its impact is
>> limited because the lowmem scratch memory calculation is restricted to
>> the first 4G of memory. The changes also cover the lowmem scratch
>> memory case.
>>
>> Cc: Alexander Graf <graf at amazon.com>
>> Cc: Andrew Morton <akpm at linux-foundation.org>
>> Cc: George Guo <guodongtai at kylinos.cn>
>> Cc: Mike Rapoport <rppt at kernel.org>
>> Cc: Pasha Tatashin <pasha.tatashin at soleen.com>
>> Cc: Pratyush Yadav <pratyush at kernel.org>
>> Cc: Ritesh Harjani (IBM) <ritesh.list at gmail.com>
>> Cc: linux-kernel at vger.kernel.org
>> Cc: linux-mm at kvack.org
>> Signed-off-by: Sourabh Jain <sourabhjain at linux.ibm.com>
>> ---
>> include/linux/memblock.h | 1 -
>> kernel/liveupdate/kexec_handover.c | 49 ++++++++++++++++++++++--------
>> mm/memblock.c | 22 --------------
>> 3 files changed, 37 insertions(+), 35 deletions(-)
>>
>> diff --git a/include/linux/memblock.h b/include/linux/memblock.h
>> index d62db9e776cf..678fe466529a 100644
>> --- a/include/linux/memblock.h
>> +++ b/include/linux/memblock.h
>> @@ -487,7 +487,6 @@ static inline __init_memblock bool memblock_bottom_up(void)
>> phys_addr_t memblock_phys_mem_size(void);
>> phys_addr_t memblock_reserved_size(void);
>> phys_addr_t memblock_reserved_kern_size(phys_addr_t limit, int nid);
>> -phys_addr_t memblock_reserved_hugetlb_size(phys_addr_t limit, int nid);
>> unsigned long memblock_estimated_nr_free_pages(void);
>> phys_addr_t memblock_start_of_DRAM(void);
>> phys_addr_t memblock_end_of_DRAM(void);
>> diff --git a/kernel/liveupdate/kexec_handover.c b/kernel/liveupdate/kexec_handover.c
>> index 7c4d86daf86d..dc809e1e768c 100644
>> --- a/kernel/liveupdate/kexec_handover.c
>> +++ b/kernel/liveupdate/kexec_handover.c
>> @@ -752,6 +752,31 @@ static int __init kho_parse_scratch_size(char *p)
>> }
>> early_param("kho_scratch", kho_parse_scratch_size);
>>
>> +static phys_addr_t __init_memblock memblock_reserved_size_nid(phys_addr_t limit, int nid,
>> + enum memblock_flags region_type)
>> +{
> Why is this in kexec_handover.c and not in memblock.c?
I kept it here because KHO is the only user of this. However, I'm happy
to move it to
memblock.c if you think that would be more appropriate.
>> + struct memblock_region *r;
>> + phys_addr_t total = 0;
>> +
>> + for_each_reserved_mem_region(r) {
>> + phys_addr_t size = r->size;
>> +
>> + if (r->base > limit)
>> + break;
>> +
>> + if (r->base + r->size > limit)
>> + size = limit - r->base;
>> +
>> +#ifdef CONFIG_NUMA
>> + if (nid == memblock_get_region_node(r))
>> +#endif
> Do we really need the #ifdef here?
Yes, otherwise, if CONFIG_NUMA is not set, the lowmem and global sizes
would be 0.
- Sourabh Jain
>
>> + if (r->flags & region_type)
>> + total += size;
>> + }
>> +
>> + return total;
>> +}
>> +
>> static void __init scratch_size_update(void)
>> {
>> /*
More information about the kexec
mailing list