[PATCH v10 RESEND 0/9] riscv: add SBI Supervisor Software Events support
Himanshu Chauhan
himanshu.chauhan at oss.qualcomm.com
Tue Sep 22 21:17:03 PDT 2026
On Mon, Sep 21, 2026 at 07:14:57PM +0800, Zhanpeng Zhang wrote:
> This is v10 rebased onto v7.3-rc4, as a base for the RAS work requested
> by Himanshu. No additional fixes or features are included.
>
Thanks for this! This patch series sees a lot of churn. My patches don't
directly apply. Give me sometime to review and test.
Regards
Himanshu
> Only patch 7 needed adaptation: keep the upstream counter-mask bitmap
> conversion and snapshot NULL-check ordering, and adapt the SSE stop-all
> helper and early counter-mask initialization to the bitmap representation.
> The fast-only GUP user-stack copy remains unchanged.
>
> Rebase validation: RV64 defconfig with SSE, PMU-SSE, CPU PM and kexec
> enabled builds Image, modules, the SSE test module and the user-stack
> selftest. The SSE, trap, fault and PMU objects also build for RV32.
> The exported series applies cleanly to v7.3-rc4 and reproduces the branch
> tree. On EVB247, the Debian-packaged rc4 kernel boots via kexec with its
> matching initrd and modules. The SSE framework and priority tests, all four
> stress layers in stress=1, and both user-stack selftests pass. The latter
> includes 32 concurrent samplers. No new kernel errors were observed.
> Unavailable injection events were skipped by the framework test; KVM was
> disabled in this test configuration, so virtualization was not retested.
>
> The functional results below are retained from the original v10 and are
> not claims of runtime validation on this rebased kernel.
>
> RISC-V does not architecturally define a supervisor-mode non-maskable
> interrupt (NMI). An interrupt that arrives while Linux has cleared SIE stays
> pending and is not observed until interrupts are enabled again. That is
> correct for ordinary interrupt handling, but some kernel work needs an
> NMI-like notification that can run even inside an interrupt-disabled region:
> sampling a PMU overflow at the instruction that caused it, or taking a
> high-priority RAS report promptly, cannot wait for the next unmask boundary.
>
> The SBI Supervisor Software Events (SSE) extension [1] fills this gap. It lets
> Linux register handlers for events that the SBI implementation can deliver
> ahead of ordinary traps and interrupts, giving RISC-V the NMI-like supervisor
> notification mechanism it otherwise lacks.
>
> SSE can carry several event sources: high-priority RAS reports, double traps,
> and PMU overflow, with room for further standard and platform events. This
> series focuses on PMU overflow, its first user. Delivering overflows through
> SSE lets perf sample the code that was actually running while interrupts were
> disabled, rather than the later point where execution reached an
> interrupt-unmask boundary.
>
> This series implements the Linux side of that interface: the architecture
> entry machinery, a firmware driver that exposes SSE events to in-kernel
> clients, PMU overflow delivery, and regression tests. Per-hart local events
> and system-wide global events share one client API.
>
> SSE delivery model
> ==================
>
> Linux first registers a handler and an event stack with the SBI
> implementation, then enables the event. When an event source is signalled, the
> M-mode SBI implementation preempts Linux even in an interrupt-disabled region:
> it saves the interrupted supervisor state and constructs an S-mode context
> that enters the registered handler. Linux can now run its own handler, for
> example to take a perf sample or process a RAS report, then completes the
> event with another SBI call, allowing the interrupted context to resume.
>
> The typical hardware-triggered delivery flow is (software-injected events skip
> the hardware trigger):
>
> <--------- Linux kernel -----------> <-- Firmware ---> <- Hardware ->
> interrupted context SSE handler OpenSBI Hardware
> | | | |
> [1] setup | |-register & enable--> |
> | | | |
> [2] trigger | | <----trigger------|
> | | | |
> [3] save | | +--------------+ |
> | | | context save | |
> | | +--------------+ |
> | | | |
> [4] inject | | +-----------------------+ |
> | | | handler context setup | |
> | | +-----------------------+ |
> | <---inject (mret) ---| |
> | | | |
> [5] handle | +----------------+ | |
> | | event handling | | |
> | +----------------+ | |
> | | | |
> [6] complete | |-----complete-------> |
> | | | |
> [7] restore | | +-----------------+ |
> | | | context restore | |
> | | +-----------------+ |
> | | | |
> [8] resume <------------resume (mret) ------------| |
> | | | |
>
> The context used to enter the handler exists only for this handoff; it is not
> the task context that the event interrupted. The architecture entry code joins
> the two sides: it moves execution onto the event's dedicated stack and shadow
> call stack, establishes the current task, and presents the interrupted
> registers to the callback as a normal pt_regs. Clients can therefore operate
> on the original interrupted context without depending on the firmware entry
> details.
>
> Linux implementation
> ====================
>
> An SSE handler runs in NMI-like context: it must not sleep, must not take a
> page fault, and may interrupt code that holds arbitrary locks or is partway
> through kernel entry. The implementation is shaped by those constraints.
>
> Because it is NMI-like, an SSE can arrive at any point where interrupts are
> disabled, including while Linux is midway through exception entry, a task
> switch, or a KVM guest transition, where the normal kernel entry state is only
> partially established. The SSE entry wrapper (the architecture assembly that
> runs before the client callback) copes with this: it preserves Linux-owned
> stvec, hstatus, and task stack metadata across the handler and any nested
> exception, and its earliest instructions, which run before the event stack and
> current task are set up, are kept outside kprobe instrumentation.
>
> The callback receives the interrupted registers as a pt_regs and is allowed to
> edit them. On RISC-V a6 and a7 carry SBI call arguments and results, so a
> callback that wants to influence an in-flight SBI call the event interrupted
> edits them there. The entry wrapper copies just a6 and a7 from that pt_regs
> back into the context handed to the completion SBI call, so the edit takes
> effect when the interrupted code resumes; the rest of the interrupted state is
> restored by firmware and left untouched.
>
> The firmware driver maps the SBI event state machine onto kernel resource
> ownership. A callback, stack, and attribute buffer stay alive until firmware
> has removed every registration that can refer to them. Failed partial
> operations remain tracked for later cleanup, an aborted CPU-offline operation
> restores the requested event state, and shutdown and kexec mask SSE before
> Linux stops servicing handlers.
>
> PMU overflow and perf
> =====================
>
> The RISC-V SBI PMU driver delivers overflows through ordinary interrupts by
> default. When firmware implements SSE and the local PMU-overflow event, the
> driver routes overflows through SSE instead. The choice is made once at setup
> and is not switched at runtime; an operational failure disables sampling
> rather than risking two active routes for the same overflow.
>
> This changes where perf can observe an overflow, not how applications use
> perf. A normal PMU interrupt raised while S-mode interrupts are masked is
> handled only once they are enabled again, so the resulting sample often points
> at the unmask boundary rather than at the code that consumed the cycles. SSE
> can enter Linux at the original point and remove that source of sampling bias.
> No new perf option or perf.data format is introduced.
>
> The entry code supplies the interrupted pt_regs needed for register samples
> and for kernel and user callchains. DWARF callchains additionally require a
> copy of the interrupted user stack. Since an SSE handler cannot take a normal
> page fault, this series takes a temporary reference to the resident user pages
> with fast-only GUP, copies them through their kernel mappings, and truncates
> the sample at the first page that is not immediately available. The existing
> in-atomic copy remains unchanged outside SSE context.
>
> The PMU integration retains perf's throttling and stopped-event semantics. It
> restarts only runnable counters and orders the CPU power-management callbacks
> so that counters cannot resume after a hart has failed to restore its SSE
> delivery path.
>
> Hardware results
> ================
>
> We measured this on a RISC-V server platform. The same kernel source
> and perf binary were used for both routes; one delivered PMU overflows through
> ordinary interrupts and the other through SSE. The table shows the mean of
> three runs of three million single-CPU "perf bench sched pipe" operations. The
> "ops/s" columns are workload throughput (higher is better, so they show the
> profiling overhead); the "samples/s" columns are the sampling rate perf
> actually achieved against the requested -F frequency:
>
> rate IRQ ops/s SSE ops/s delta IRQ samples/s SSE samples/s
> -F 99 337,707 339,555 +0.55% 98.0 98.6
> -F 999 338,352 338,289 -0.02% 995.7 998.0
> -F 5000 329,002 333,034 +1.23% 5001.7 5001.6
>
> There were no lost samples. Across these normal frequency settings, both
> delivery modes reached the requested sample rate and workload throughput
> differed by no more than 1.23%.
>
> The "perf bench sched pipe" workload also shows why the delivery mechanism
> matters to the resulting profile. Ordinary PMU interrupts cannot enter an
> interrupt-disabled kernel critical section. Overflows raised there remain
> pending until interrupts are enabled again. Samples consequently accumulate
> at the enable boundary rather than at the code that consumed the cycles. In
> the IRQ profile, finish_task_switch() and _raw_spin_unlock_irqrestore()
> therefore accounted for 54.99% of all samples.
>
> SSE can enter Linux while S-mode interrupts are disabled. The PMU-SSE
> profile therefore samples inside those critical sections and exposes the
> scheduler, locking, address-space switching, and wake-up paths doing the
> actual work. The leading entries from the two -F 999 reports show the
> difference.
>
> With ordinary PMU interrupt delivery:
>
> overhead symbol
> 36.63% finish_task_switch.isra.0
> 18.36% _raw_spin_unlock_irqrestore
> 7.66% __internal_syscall_cancel
> 7.55% do_trap_ecall_u
> 4.19% mutex_lock
> 3.64% mutex_unlock
> 3.06% exit_to_user_mode_loop
>
> With PMU-SSE delivery:
>
> overhead symbol
> 5.48% __kprobes_text_end
> 5.29% __schedule
> 5.10% ret_from_exception
> 4.71% do_raw_spin_lock
> 4.01% do_trap_ecall_u
> 3.99% mutex_lock
> 3.66% switch_mm
> 3.43% mutex_unlock
> 3.29% exit_to_user_mode_loop
> 3.28% psi_group_change
>
> The ordinary interrupt profile is dominated by two interrupt-enable
> boundaries. With SSE, those two entries account for only 3.37%. The samples
> are instead distributed across scheduler paths within the critical sections.
>
> At perf's configured limit of 100,000 samples per second, both routes still
> made progress without lost samples. In this deliberately saturated regime SSE
> reduced workload throughput by 2.7% to 5.8%, which exposes the additional
> firmware-entry cost and marks a practical upper boundary for sampling.
> Thirty-second perf top runs at the same rate each processed about 3.1 million
> samples with no loss, stalls, or kernel failures.
>
> The DWARF callchain path gets dedicated coverage because it was the source of
> the corruption this series fixes. On the same platform,
> "perf record -a -g --call-graph dwarf,512 -F 999" layered on a concurrent
> "hackbench -g25 -l600" -- the configuration that previously corrupted
> spinlocks and mutexes under SSE -- now completes cleanly, with no lost
> samples, lockups, RCU stalls, or faults, including a 431-iteration soak.
> Patch 9 adds a regression test that drives the non-faulting user-stack copy
> through the SSE handler with 32 concurrent samplers and checks perf's
> truncation semantics.
>
> Changes in this resend
> ======================
>
> This resend only rebases v10 onto v7.3-rc4. The only merge conflict was
> in patch 7, due to the upstream PMU counter-mask bitmap conversion.
> In v11, I will address the Sashiko review feedback and improve user-stack
> copying with an NMI-safe interface similar to x86's copy_from_user_nmi().
>
> Changes in v10
> ==============
>
> V10 turns the earlier feature series into a path suitable for sustained perf
> use. In particular, it:
>
> - reconstructs and publishes the interrupted context for perf register
> samples and kernel and user callchains;
> - preserves current, task stack metadata, stvec, hstatus, and shadow-call
> stack state across synthetic entry and nested exceptions;
> - prevents fault-disabled accesses from entering the generic RISC-V page
> fault path and provides a non-faulting SSE user-stack copy;
> - makes event lifetime and rollback explicit across partial firmware
> operations, CPU hotplug, shutdown, crash, and kexec;
> - closes PMU throttle, counter restart, CPU power-management, and cleanup
> races without adding a runtime SSE-to-IRQ transition; and
> - expands the framework stress coverage and adds a regression test for
> high-frequency DWARF user-stack sampling.
>
> Changes in v9:
> - Rebased the original series onto RISC-V for-next.
> - Preserved Linux-owned trap, virtualization, and supervisor state across
> the synthetic SSE handler.
> - Added framework stress modes and updated MAINTAINERS.
>
> Previous versions:
> v9:
> https://lore.kernel.org/r/cover.1778331862.git.zhangzhanpeng.jasper@bytedance.com
> v8:
> https://lore.kernel.org/r/20251105082639.342973-1-cleger@rivosinc.com
>
> How to test
> ===========
>
> Enable the SSE framework and SSE overflow delivery:
>
> CONFIG_RISCV_SBI_SSE=y
> CONFIG_RISCV_PMU_SBI=y
> CONFIG_RISCV_PMU_SBI_SSE=y
>
> PMU-SSE also requires two OpenSBI fixes:
>
> f30a54f3b3a0 ("lib: sbi: pmu: Remove MIP clearing from pmu_sse_enable()")
> [2], included since OpenSBI v1.7,
> which keeps an overflow pending while its SSE event is temporarily
> disabled; and
> 35511bc6ee1c ("lib: sbi: sse: clear SPV for non-virtualized events") [3],
> not yet included in a tagged release,
> which stops a stale HSTATUS.SPV from being applied to a non-virtualized
> event.
>
> Build tools/testing/selftests/riscv, then run:
>
> for stress in 0 1 2; do
> ./run_sse_test.sh stress=$stress || break
> done
> ./sse_perf_ustack
>
> Useful perf regression workloads include:
>
> perf record -e cycles -a -- sleep 1
> perf top
> perf record -g -F 999 -- hackbench
> perf record --call-graph dwarf,8192 -F 999 -- hackbench
> perf record -a -C 3 -e cycles -F 999 -- \
> taskset -c 3 perf bench sched pipe -l 3000000
>
> Limitations and follow-up work
> ==============================
>
> This series does not yet deliver SSE events into a guest or unwind a guest
> stack; a later KVM-SSE series will let the host receive an event from firmware
> and inject the corresponding event into the guest.
>
> Hibernation and crash kernels are unsupported: the current SBI interface
> cannot reconstruct firmware registrations after an image is restored, and a
> crash kernel cannot take over the registrations left by the crashed kernel, so
> it leaves SSE masked.
>
> [1] https://docs.riscv.org/reference/sbi/ext-sse.html
> [2] https://github.com/riscv-software-src/opensbi/commit/f30a54f3b3a091c225a00476f4039bf399badd1f
> [3] https://github.com/riscv-software-src/opensbi/commit/35511bc6ee1c9c17b6a89b44c52e2044bb51b979
>
> Acknowledgements
> ================
>
> The original five feature patches were developed by Clément Léger and
> Himanshu Chauhan. Thanks to Susheng Yang for reporting the perf callchain
> failure and for providing a workload that made it reproducible.
>
> Sorry for keeping you waiting. Since v9 I spent a good deal of time hardening
> the lifecycle and error paths and reproducing and analysing the bugs that only
> show up in the callchain path, until the series finally passed both functional
> and sustained stress testing on hardware. I am confident in v10, but, echoing
> Clément, SSE is a genuinely complex feature: it adds a new NMI-like entry path
> into the kernel to stand in for a hardware NMI. I would therefore welcome wider
> community testing and feedback, especially under high-frequency delivery and
> more complex handlers.
>
> ---
>
> Clément Léger (5):
> riscv: add SBI SSE extension definitions
> riscv: add support for SBI Supervisor Software Events extension
> drivers: firmware: add riscv SSE support
> perf: RISC-V: add support for SSE event
> selftests/riscv: add SSE test module
>
> Zhanpeng Zhang (4):
> riscv: sse: mask events during shutdown and kexec
> riscv: mm: avoid enabling interrupts for nofault page faults
> perf: RISC-V: support callchains with SSE delivery
> selftests/riscv: add perf user-stack SSE copy regression test
>
> Documentation/arch/riscv/index.rst | 1 +
> Documentation/arch/riscv/pmu-sse.rst | 55 +
> MAINTAINERS | 22 +
> arch/riscv/include/asm/asm.h | 14 +-
> arch/riscv/include/asm/perf_event.h | 10 +
> arch/riscv/include/asm/sbi.h | 63 +
> arch/riscv/include/asm/scs.h | 7 +
> arch/riscv/include/asm/sse.h | 82 ++
> arch/riscv/include/asm/thread_info.h | 1 +
> arch/riscv/kernel/Makefile | 1 +
> arch/riscv/kernel/asm-offsets.c | 14 +
> arch/riscv/kernel/entry.S | 14 +
> arch/riscv/kernel/machine_kexec.c | 11 +
> arch/riscv/kernel/perf_callchain.c | 142 ++
> arch/riscv/kernel/reset.c | 18 +
> arch/riscv/kernel/sbi_sse.c | 246 ++++
> arch/riscv/kernel/sbi_sse_entry.S | 226 +++
> arch/riscv/kernel/smp.c | 17 +
> arch/riscv/mm/fault.c | 11 +-
> drivers/firmware/Kconfig | 1 +
> drivers/firmware/Makefile | 1 +
> drivers/firmware/riscv/Kconfig | 18 +
> drivers/firmware/riscv/Makefile | 3 +
> drivers/firmware/riscv/riscv_sbi_sse.c | 1228 +++++++++++++++++
> drivers/perf/Kconfig | 11 +
> drivers/perf/riscv_pmu.c | 14 +-
> drivers/perf/riscv_pmu_sbi.c | 540 ++++++--
> include/linux/cpuhotplug.h | 1 +
> include/linux/perf/riscv_pmu.h | 20 +-
> include/linux/riscv_sbi_sse.h | 95 ++
> tools/testing/selftests/riscv/Makefile | 2 +-
> tools/testing/selftests/riscv/sse/Makefile | 10 +
> .../selftests/riscv/sse/module/Makefile | 22 +
> .../riscv/sse/module/riscv_sse_test.c | 1154 ++++++++++++++++
> .../selftests/riscv/sse/run_sse_test.sh | 59 +
> .../selftests/riscv/sse/sse_perf_ustack.c | 564 ++++++++
> 36 files changed, 4599 insertions(+), 99 deletions(-)
> create mode 100644 Documentation/arch/riscv/pmu-sse.rst
> create mode 100644 arch/riscv/include/asm/sse.h
> create mode 100644 arch/riscv/kernel/sbi_sse.c
> create mode 100644 arch/riscv/kernel/sbi_sse_entry.S
> create mode 100644 drivers/firmware/riscv/Kconfig
> create mode 100644 drivers/firmware/riscv/Makefile
> create mode 100644 drivers/firmware/riscv/riscv_sbi_sse.c
> create mode 100644 include/linux/riscv_sbi_sse.h
> create mode 100644 tools/testing/selftests/riscv/sse/Makefile
> create mode 100644 tools/testing/selftests/riscv/sse/module/Makefile
> create mode 100644 tools/testing/selftests/riscv/sse/module/riscv_sse_test.c
> create mode 100644 tools/testing/selftests/riscv/sse/run_sse_test.sh
> create mode 100644 tools/testing/selftests/riscv/sse/sse_perf_ustack.c
>
>
> base-commit: 93f51579e7df248780214094418f205253383cc5
> --
> 2.50.1 (Apple Git-155)
More information about the linux-arm-kernel
mailing list