[RFC PATCH] arm64: mm: Map fixmap PTE tables r/o in the linear map
Ard Biesheuvel
ardb at kernel.org
Thu Aug 27 03:16:11 PDT 2026
On Thu, 27 Aug 2026, at 12:00, Kevin Brodsky wrote:
> On 27/08/2026 11:04, Ard Biesheuvel wrote:
>> On Thu, 27 Aug 2026, at 10:55, Kevin Brodsky wrote:
>>> On 26/08/2026 11:07, Ard Biesheuvel wrote:
>>>>>> [...]
>>>>>>
>>>>>> +static int __init fixmap_remap_ro(void)
>>>>>> +{
>>>>>> + return set_direct_map_ro((unsigned long)lm_alias(&bm_pte), NR_BM_PTE_TABLES);
>>>>> Should we not also remap bm_pmd and bm_pud?
>>>>>
>>>>> For that matter, do we need RW access via the linear map for any page
>>>>> annotated with __bss_pgtbl? I suppose that might be the case for
>>>>> kasan_early_shadow_* but I don't know enough about KASAN to tell for sure.
>>>>>
>>>> bm_pte[] is special because it is only ever written via the kernel mapping,
>>>> and never via the linear map. This is why it is being singled out in this
>>>> patch.
>>> Right I see __set_fixmap(). This is a very special case of calling
>>> __set_pte() on a pointer derived from a global (i.e. pointing to the
>>> kernel image and not the LM) so it makes sense to treat it differently.
>>>
>>> I do wonder whether there is that much value in protecting bm_pte while
>>> all page tables in init_pg_dir (also at a fixed offset in the LM) remain
>>> writeable though.
>>>
>> Yes, so those need to be taken care of as well.
>
> But init_pg_dir needs to remain writeable in the LM, right?
>
>>> Either way all this is interesting for my series protecting page tables
>>> with pkeys [1] - it seems that to protect the fixmap, the easiest option
>>> would be to change the pkey of .pgtlb in both the kernel image and LM
>>> (the former for bm_pte, and the latter for everything else).
>>>
>>> [1] https://lore.kernel.org/all/20260818-kpkeys-v9-0-743ad31b2c8f@arm.com/
>>>
>> Thanks, I've been meaning to go through those but they are a bit overwhelming :-)
>
> Yes there's a fair bit to go through, sorry for that... You'll probably
> be most interested in patch 12, 13 and 19-21 - that's where page tables
> are being mapped with a special pkey. The cover letter should give
> enough context for them to make sense on their own.
>
OK, that helps ... :-)
>>>> Whether or not bm_pmd[] can be treated as a special case depends on the page
>>>> size: with 4k pages, the whole array covers a virtual region of 1G, which is
>>>> currently guaranteed to be shared only with the PCI I/O space (but we could
>>>> move that out). With 16k pages, it covers 64G, and so it is shared with the
>>>> vmemmap and other virtual mappings in the vmalloc region, and so the current
>>>> kernel mapping code expects to be able to write those entries.
>>>>
>>>> What we might do is generalize the logic that uses the fixmap to modify
>>>> pgd level entries in swapper_pg_dir, and use it for all modifications
>>>> at PMD level or higher if those tables are in .rodata
>>>>
>>>> But this is a bit more complicated than this change, so I decided to
>>>> present this as a separate change.
>>> Definitely, this is orthogonal to this patch. I wonder what the
>>> performance impact would be if we used the fixmap for setting all kernel
>>> entries at PMD and above. It would be nice to keep that logic easily
>>> togglable, because with the kpkeys approach I mentioned above we have a
>>> (most likely) much cheaper way of protecting these page tables.
>>>
>> I'm currently experimenting with using __put_kernel_nofault() to update
>> the descriptors at all levels (including PTE) and handling the fault
>> using the fixmap approach. That way, all statically allocated page
>> tables could move to .rodata (except bm_pte[]) and there should be no
>> substantial performance impact.
>
> Doesn't that still mean modifying a fixmap PTE every time we update
> kernel page tables? That doesn't exactly sound cheap.
>
Only the ones that were allocated statically, which cover the region
around the kernel image, and the special cases (fixmap, kasan). Everything
else is allocated dynamically, using memblock_alloc() if very early
during the boot.
>>>>>> +}
>>>>>> +late_initcall(fixmap_remap_ro);
>>>>> Is mark_rodata_ro() definitely too early to remap these pages RO?
>>>>>
>>>> No, we might just call this from there, afaict.
>>> Sounds good, could we then use update_mapping_prot()? I suppose not as
>>> that would require mapping bm_pte separately in map_mem()?
>>>
>> Yeah, but perhaps we should do the latter anyway, so we can remap the
>> region even if !can_set_direct_map().
>
> Yep makes sense. Not having to add yet another set_direct_map_* function
> is definitely welcome ;)
>
Ack.
More information about the linux-arm-kernel
mailing list