RISC-V genpd race in OSI mode (genpd domain power_off vs firmware suspend call)
Jimmy Ho
jimmy.ho at sifive.com
Thu Aug 27 22:03:35 PDT 2026
Resending this follow-up to Ulf’s current address
On Fri, Aug 28, 2026 at 12:56 PM Jimmy Ho <jimmy.ho at sifive.com> wrote:
>
> Hi linux-pm,
>
> We hit a race condition involving genpd's CPU PM domain (OSI-mode
> hierarchical cpuidle, similar to the ARM PSCI OSI model) on a RISC-V
> platform, and would like to get your opinion on whether this is a
> genpd-level issue and what the preferred generic fix would look like.
>
> Background
> ----------
> Observed this issue on qemu platform
> On our SoC, drivers/irqchip/irq-riscv-aplic-main.c registers a genpd
> notifier to save/restore the S-mode APLIC context:
>
> aplic_notifier()
> GENPD_NOTIFY_PRE_OFF -> aplic_save()
> GENPD_NOTIFY_ON -> aplic_restore()
>
> CPUs are attached to this genpd via dt_idle_attach_cpu() (as in
> drivers/cpuidle/cpuidle-riscv-sbi.c, mirroring cpuidle-psci.c). When
> the last CPU in the domain goes idle, genpd_power_off() decides the
> domain can be turned off (based on the "not_suspended" count computed
> under genpd_lock). When the first CPU
> resumes, genpd_power_on() decides to turn the domain back on and fires
> GENPD_NOTIFY_ON, which triggers aplic_restore().
>
> However, the actual SBI HSM suspend call (riscv_sbi_hart_suspend(),
> similar to a PSCI CPU_SUSPEND call) is *not* issued inside
> genpd_lock. genpd's power_off() callback only records which domain
> idle state was selected; the real firmware suspend call happens later,
> asynchronously, on each CPU's own cpuidle enter path
> (__sbi_enter_domain_idle_state()).
>
> The race
> --------
> Because of this gap between "genpd's software decision that this CPU
> is last-man / first-man" and "the CPU actually issuing the firmware
> suspend/resume call", we observed the following sequence
>
> 1. CPU0 requests a system domain-level power-gate (deep sleep, tears down the
> shared S-mode APLIC context) and hands off to firmware.
> 2. Before CPU0's request is physically completed, CPU1 wakes up,
> genpd_power_on() runs, GENPD_NOTIFY_ON fires, aplic_restore() runs.
> 3. CPU1 goes idle again, but this time system domain-level genpd reject
> request, so it fallback to core level power-gate
> 4. CPU0's original system deep power-gate request eventually completes at the
> firmware level -- physically collapsing the system domain, including the
> S-mode APLIC context, saved earlier by aplic_save()
> 5. CPU0 wakes up (e.g. UART interrupt), enters Linux, but system genpd's
> software state still says the domain is "on" (from step 2), so system domain
> GENPD_NOTIFY_ON will not be fired again and aplic_restore() will not be
> called
>
> This looks structurally identical to the race TF-A's PSCI OS-initiated
> mode documentation describes:
> https://trustedfirmware-a.readthedocs.io/en/latest/design_documents/psci_osi_mode.html#races-in-os-initiated-mode
>
>
> "In OS-initiated mode, there are race windows where the OS's view
> and implementation's view of the system's state differ... the OS
> might request a powerdown state for a node from one core, while at
> the same time, the implementation observes that another core in
> that node is powering up."
>
> TF-A's answer is to require the OS to encode a "last_at_pwrlvl" field
> into the StateID passed to CPU_SUSPEND, which firmware validates
> against its own view of system state before honoring a powerdown
> request. We could not find any code in current mainline Linux
> (cpuidle-psci.c or generic cpuidle/genpd code) that actually computes
> and encodes this field dynamically per suspend call, so as far as we
> can tell this race is unhandled generically today, on both arm/psci
> and riscv/sbi backends.
>
> Options we considered
> ----------------------
> 1. Implement something like PSCI's last_at_pwrlvl generically: expose
> genpd's already-computed "am I the last CPU in this domain"
> information through the power_off() callback API, so
> backends can pass it down to firmware, and have firmware validate/
> reject stale requests.
> - Pro: matches the documented PSCI OSI approach.
> - Con: We don't know how to set this flag properly in a generic way.
> last_at_pwrlvl is only meaningful on riscv/arm, so we can't just
> add it to genpd's generic code path. genpd also doesn't keep a
> counter of how many CPUs in a domain are still alive, so there's
> no existing state in genpd's power_off() callback we could read
> to derive the flag. One idea would be to walk the genpd hierarchy
> from within the power_off() callback and check the runtime_status
> of each CPU device in the same domain (genpd domains already carry
> a CPU mask, and get_cpu_device() can be used to look up the
> corresponding CPU device)
>
> 2. Do the power-domain control directly from the power_off()/power_on()
> genpd callbacks (e.g. have them poke the SoC's power controller
> registers via an SBI call, instead of relying on each CPU's own
> idle-entry path to trigger the deepest state).
> it creates two independent driver that will touch power controller
> (opensbi vs linux)
>
> 3. Have each CPU pass a monotonically increasing sequence number
> alongside its suspend request to firmware, and have firmware
> itself detect and reject/downgrade a deep power-down request if a
> *newer* sequence number from a shallower (i.e. "some CPU became
> active again") request has been observed for the same domain since.
> This is a firmware-side staleness check, entirely below the SBI
> call boundary
> this seq number can be maintained in gendpd generic code
> or riscv power_off() callback, where do you suggest this to be
> added?
>
> Do you have any suggestion to solving this issue?
> Thanks,
> Jimmy Ho
More information about the linux-riscv
mailing list