[PATCH v19 4/7] firmware: arm_rmm: Add support for SRO
Suzuki K Poulose
suzuki.poulose at arm.com
Mon Sep 28 13:45:07 PDT 2026
On 28/09/2026 18:28, Catalin Marinas wrote:
> On Mon, Sep 28, 2026 at 11:13:39AM +0100, Suzuki K Poulose wrote:
>> On 28/09/2026 10:28, Catalin Marinas wrote:
>>> On Thu, Sep 24, 2026 at 02:51:58PM +0100, Suzuki K Poulose wrote:
>>>> +long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp)
>>>> +{
>>>> + struct arm_smccc_1_2_regs *regs = &sro->regs;
>>>> + bool cancelled = false;
>>>> + unsigned long sro_handle;
>>>> +
>>>> + rmi_smccc_invoke(regs);
>>>> +
>>>> + sro_handle = regs->a1;
>>>> + while (RMI_RESULT_STATUS(regs->a0) == RMI_INCOMPLETE) {
>>>> + bool can_cancel = RMI_RESULT_CAN_CANCEL(regs->a0) == RMI_OP_CAN_CANCEL;
>>>> + int ret = 0;
>>>> +
>>>> + switch (RMI_RESULT_MEMREQ(regs->a0)) {
>>>> + case RMI_OP_MEM_REQ_NONE:
>>>> + rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
>>>> + regs);
>>>> + break;
>>>> + case RMI_OP_MEM_REQ_DONATE:
>>>> + ret = rmi_sro_donate(sro, sro_handle, regs->a2, regs,
>>>> + gfp);
>>>> + break;
>>>> + case RMI_OP_MEM_REQ_RECLAIM:
>>>> + ret = rmi_sro_reclaim(sro, sro_handle, regs);
>>>> + break;
>>>> + default:
>>>> + WARN_ON_ONCE(1);
>>>> + ret = -ENXIO;
>>>> + break;
>>>> + }
>>>
>>> Another thing I came across while looking whether we can defer the
>>> activation. It seems that the spec (I_JVYCH) lists some SROs as
>>> PE-bound. Nothing here or in rmi_sro_execute() disables migration and
>>> the memory allocation paths can even sleep with GFP_KERNEL.
>>
>> No, this is not required. I agree this is confusing. I will get it
>> clarified.
>>
>> So, there are two different sources for the SRO contexts. One is a global
>> pool and the other an Object.
>>
>> e.g., For an RMI operation on an Object, SRO context can be the object
>> itself (e.g., REC_CREATE, REALM_ACTIVATE etc.)
>>
>> However, when there is no reliable object for the command (e.g.,
>> RMI_GRANULE_RANGE_DELEGATE), the RMM must allocate a context from
>> the global pool. Now, the "PE" in there comes from a recommendation
>> to the RMM implementations, that the global pool size must depend on
>> the number of PEs on the system. This doesn't mean that the SRO
>> handles are only bound to those PEs. I will get this clarified
>> in the RMM spec.
>
> This part of the spec needs rewriting, not clarifying. No matter how
> hard you try, there's no way you can read it as a "global pool". For
> example:
>
> D_GZLMRA SRO context is bound to one of the following:
> - A PE
> - An RMM object
>
> And take a random command:
>
> B4.5.2 RMI_DPT_L0_CREATE command
> Create a Level 0 DPT.
> The RMI_DPT_L0_CREATE command may initiate a Stateful RMI
> Operation whose context is bound to the current PE.
>
> "bound to the current PE" pretty clearly shows the intention was to
> disable preemption. It also doesn't say what happens when this pool is
> exhausted (presumably it returns RMI_BLOCKED).
Agree. If there are not contexts available for the RMM to use, it
results in RMI_BLOCKED.
>
> TBH, that's a pretty significant change for a bet3/4 release, though
> arguably it can be seen as a relaxation. Code that relies on disabling
> preemption should still work (somewhat, assuming the global pool is at
> least the number of PEs and the host plays nicely to complete or cancel
> all SROs).
The only case where the RMM mandates the execution continues on a "CPU"
(virtual PE rather) is Realm Attestation ABI exposed to the Realm
(RSI_ATTEST_TOKEN_*).
Btw, SRO contexts are nothing but a "scratch" memory that the RMM uses
to hold the progress made in the SRO. e.g., the donated granules and
what is the next operation expected etc.
>
> That said, such pool is a limited resource and we need some way to probe
> its size if we want to do something smarter in the kernel, like a
> semaphore to ensure we don't randomly fail because of an RMM limitation.
> I don't really see how the number of PEs is relevant to this global
> pool sizing, it's not that we limit the realms we can start to the
> online CPUs.
The number of PEs is only a factor which can tell the RMM, how many
parallel requests it could get. That said, due to pre-emption there
could be multiple outstanding SRO operations on a single PE. But
also remember that not all SROs require a context from global pool.
Agreed that it would good to probe the number of contexts available.
Or may be even dynamically extend the pool for SRO contexts at
runtime. Will feed this back to the RMM spec, and hopefully we can
address this in the future versions without breaking the existing
ABI.
Thanks
Suzuki
More information about the linux-arm-kernel
mailing list