[PATCH RFC v9 00/25] pkeys-based page table hardening

Kevin Brodsky kevin.brodsky at arm.com
Thu Sep 3 09:47:50 PDT 2026


On 01/09/2026 16:24, Linu Cherian wrote:
> Hi Kevin,
>
> On Tue, Aug 18, 2026 at 03:08:42PM +0100, Kevin Brodsky wrote:
>> [Sending during the merge window in case reviewers have spare
>> cycles; I'm not aiming to have this series merged in v7.3.]
>>
>> This is a proposal to leverage protection keys (pkeys) to harden
>> critical kernel data, by making it mostly read-only. The series includes
>> a simple framework called "kpkeys" to manipulate pkeys for in-kernel use,
>> as well as a page table hardening feature based on that framework,
>> "kpkeys_hardened_pgtables". Both are implemented on arm64 as a proof of
>> concept, but they are designed to be compatible with any architecture
>> that supports pkeys.
>>
>> The proposed approach is a typical use of pkeys: the data to protect is
>> mapped with a given pkey P, and the pkey register is initially
>> configured to grant read-only access to P. Where the protected data
>> needs to be written to, the pkey register is temporarily switched to
>> grant write access to P on the current CPU.
>>
>> The key fact this approach relies on is that the target data is
>> only written to via a limited and well-defined API. This makes it
>> possible to explicitly switch the pkey register where needed, without
>> introducing excessively invasive changes, and only for a small amount of
>> trusted code.
>>
>> Page tables are chosen as an initial target because of their especially
>> critical nature - a single write may result in arbitrary pages becoming
>> accessible to any context (including userspace). In order to keep the
>> series digestible for reviewers, this version focuses on functionality
>> rather than performance, making it most suitable as a debug feature. The
>> key trade-off is the requirement to PTE-map the linear map - see section
>> "Protected page table allocation" for details.
>>
>> This series has similarities with the "PKS write protected page tables"
>> series posted by Rick Edgecombe a few years ago [1] but it is not
>> specific to x86/PKS - the approach is meant to be generic.
>>
>> This proposal (as of RFC v5) was presented at Linux Security Summit
>> Europe 2025 [2].
>>
>> [Table of contents]
>>
>> * kpkeys
>>   - pkey register management
>>
>> * kpkeys_hardened_pgtables
>>   - Protected page table allocation
>>   - kpkeys context switching
>>   - Performance
>>   - Limitations
>>
>> * This series
>>   - Branches
>>
>> * Threat model
>>
>> * Further use-cases
>>
>> * Open questions
>>
>> kpkeys
>> ======
>>
>> The use of pkeys involves two separate mechanisms: assigning a pkey to
>> pages, and defining the pkeys -> permissions mapping via the pkey
>> register. This is implemented through the following interface:
>>
>> - Pages are assigned a pkey in the linear map using set_memory_pkey().
>>   This is sufficient for this series, but it is also plausible for
>>   higher-level allocators to support marking allocations with a given
>>   pkey.
>>
>> - The pkey register is configured based on a *kpkeys context*. kpkeys
>>   contexts are represented as simple integers that correspond to a given
>>   configuration, for instance:
>>
>>   KPKEYS_CTX_DEFAULT:
>>         RW access to KPKEYS_PKEY_DEFAULT
>>         RO access to any other KPKEYS_PKEY_*
>>
>>   KPKEYS_CTX_<FEAT>:
>>         RW access to KPKEYS_PKEY_DEFAULT
>>         RW access to KPKEYS_PKEY_<FEAT>
>>         RO access to any other KPKEYS_PKEY_*
>>
>>   Only pkeys that are managed by the kpkeys framework are impacted;
>>   permissions for other pkeys are left unchanged (this allows for other
>>   schemes using pkeys to be used in parallel, and arch-specific use of
>>   certain pkeys).
>
> - Adding some basic details on what a scheme and context is quite helpful.
>
>  - Giving some hints (may be an example) on how multiple schemes and multiple contexts
>   play together would be quite helpful.

"scheme" doesn't mean anything precise, it's only the notion that pkeys
that aren't reserved for kpkeys (i.e. anything but 0 or 1 in this
series) may be used for other purposes. Happy to reword if you have a
suggestion.

"kpkeys context" is what is described above this paragraph, it's really
just a set of permissions for the managed pkeys. Transitioning between
context is described below.

> Adding a documentation that covers these aspects would be much
> appreciated.

For sure, I am planning to have a documentation patch in a subsequent
version.

> My understanding is that pkeys are being partitioned across different
> contexts. But then the introduction of the term "scheme" looks bit confusing to me.

I wouldn't say pkeys are partitioned across contexts. Every context has
a set of permissions for all the pkeys managed by kpkeys. Any other pkey
is ignored (permissions left unchanged) by this framework.

>>   The current kpkeys context is changed by calling
>>   kpkeys_enter_context(), which will set the pkey register
>>   accordingly and return the original state. A
>>   subsequent call to kpkeys_leave_context() restores the original
>>   state (and thus the original kpkeys context). The numeric value of
>>   KPKEYS_CTX_* (kpkeys context) is purely symbolic and thus generic,
>>   however each architecture is free to define non-default pkeys
>>   values (KPKEYS_PKEY_*).
>>
> ..snip
>
>> Open questions
>> ==============
>>
>> A few aspects in this RFC that are debatable and/or worth discussing:
>>
>> - There is currently no restriction on how kpkeys contexts map to pkeys
>>   permissions. A typical approach is to allocate one pkey per context and
>>   make it writable in that context only. As the number of contexts
> Probably to avoid the assumption, may be we can we have something like
> below 
>
> For a pkey P, we could define
> PKEY_P_PERM_CTXT_OTHERS	 //permission for pkey p in other contexts
> PKEY_P_PERM_CTXT_SELF	 //permission for pkey p in self context
>
> With the assumption of one pkey mapped for every context,
> the permission for the default context would look something like,
>
> PKEY_DEF_PERM_CTXT_SELF << PKEY_DEF_PKEY_SHIFT |
> PKEY_CT0_PERM_CTXT_OTHERS << PKEY_CT0_PKEY_SHIFT | 
> PKEY_CT1_PERM_CTXT_OTHERS << PKEY_CT1_PKEY_SHIFT |
> ...(for all valid contexts)
>
> where,
> Permission key, PKEY_DEF is associated with context DEFAULT,
> Permission key, PKEY_CT0 is associated with context CT0,
> Permission key, PKEY_CT1 is associated with context CT1

This adds assumptions rather than avoiding them. *Typically* when adding
a context you'd allocate a pkey that's only writable by this context,
but it doesn't have to be this way.

The configuration space is more easily understood by considering the
other use-cases we've investigated (struct cred protection and eBPF
isolation, linked further down). For instance, for cred protection, we
had KPKEYS_LVL_UNRESTRICTED with write access to all pkeys, and for eBPF
isolation, we need a level that is less privileged and therefore does
*not* have write access to pkey 0.

>>   increases, we may however run out of pkeys, especially on arm64 (just
>>   8 pkeys with POE). Depending on the use-cases, it may be acceptable to
>>   use the same pkey for the data associated to multiple contexts.
> Lets say two contexts A and B, use the same pkey P as their permission matches.
> But then, when we enter context A, permission for pkey P gets
> relaxed, then that would relax permission for pages associated with
> context B as well which is unintended ?

That may be exactly what is intended, it all depends on the use-case. C1
may have a private pkey P1, and C2 P2, and then P3 that is shared by C1
and C2 (writable by both)
 
> As the hardware supports 16 pkeys, should we consider removing the limit
> of 8 pkeys so that we can have unique pkeys for each context ?

FEAT_S1POE only supports 4-bit pkeys when using 128-bit page tables.

- Kevin



More information about the linux-arm-kernel mailing list