[PATCH v2 00/20] arm64: Preemptible this_cpu_*() operations
Usama Anjum
usama.anjum at arm.com
Wed Sep 2 04:55:12 PDT 2026
On 04/08/2026 6:04 pm, Mark Rutland wrote:
> This series reworks arm64's this_cpu_*() operations such that they do
> not need to disable preemption, avoiding related overhead in the fast
> paths. Instead, the ops begin/end a "PCPU GPR" critical section using
> unconditional/posted stores, which should be very cheap on any
> reasonable micro-architecture. During a critical section, should a
> (preemptible) exception be taken, the entry code will apply a fixup to
> the GPRs containing the percpu offset and the generated percpu address.
>
> The scheme is described in detail in patch 13, and is similar to the
> approach Heiko Carstens applied to s390 [1], which inspired this series.
>
> Please note that the fixup IS NOT a restart. The GPR fixup in the
> exception entry code DOES NOT alter the PC, and there are no necessary
> branches within the PCPU GPR critical sections.
>
> This scheme should build and function in all kernel configurations (e.g.
> regardless of KPTI, SW PAN, CNP, VA BITS), and has no dependency on new
> architectural features. The fixup logic should "just work" with kprobes,
> etc, and I don't expect that this will need to become more complicated.
>
> The patches are organised as follows:
>
> * Patches 1 to 2 are preparatory fixes for latent issues which were
> found by inspection. These will need to be backported to stable, and
> have appropriate Fixes tags.
>
> * Patches 3 to 6 are preparatory improvements to code generation issues
> found by inspection during development. These aren't strictly related
> to the PCPU GPR scheme, and it would make sense to queue these even if
> we don't go ahead with the rest of the series.
>
> * Patches 7 to 12 are preparatory work for the PCPU GPR scheme.
>
> * Patch 13 implements the core of the PCPU GPR scheme, with all the
> necessary exception handling logic, and the addition of helpers to
> begin/end a PCPU GPR critical section.
>
> * Patches 14 to 19 convert this_cpu_*() operations over to the PCPU GPR
> scheme. These changes have been made over several patches to aid
> review and bisection (if necessary).
>
> * Patch 20 removes code made redundant by earlier patches.
>
> I've given this build-testing (with GCC and clang) and some light boot
> testing, but this hasn't seen significant functional testing or
> benchmarking. From inspection of the generated code I expect this to
> have reasonable positive impact to performance where this_cpu*() ops are
> used heavily. I would be grateful if anyone could take this for a spin.
tl;dr
Comparing this series with the page-table series [a] gives 8 improvements and 3
regressions.
Fastpath is a Linux kernel performance benchmarking service. The table
compares the both patched kernels: positive values are faster, negative
values are slower, and (I)/(R) indicate statistically significant
improvements/regressions. Unmarked differences are not significant after
accounting for confidence and noise thresholds.
+---------------------------------+--------------------+-----------------+-----------------------+
| Benchmark | percpu-pgtable [2] | gpr-fixup [3] | gpr-fixup [3] |
| | vs baseline [1] | vs baseline [1] | vs percpu-pgtable [2] |
+=================================+====================+=================+=======================+
| lmbench/lat-mem-rd | -0.26% | (I) 1.24% | (I) 1.50% |
| micromm/fork | (I) 1.10% | (I) 2.46% | (I) 1.35% |
| micromm/munmap | (I) 18.10% | (I) 19.78% | (I) 1.43% |
| micromm/vmalloc | (I) 11.92% | (I) 14.68% | (I) 2.47% |
| mmtests/hackbench | 0.27% | (I) 1.09% | 0.81% |
| mmtests/kernbench | (I) 1.19% | (I) 1.18% | -0.01% |
| mmtests/sysbench-cpu | 0.01% | 0.06% | 0.05% |
| mmtests/sysbench-mutex | -0.85% | -0.14% | 0.72% |
| mmtests/sysbench-thread | (R) -4.49% | (I) 3.02% | (I) 7.86% |
| perf/futex | 0.55% | (R) -2.80% | (R) -3.34% |
| perf/sched | (I) 1.61% | -0.39% | (R) -1.97% |
| perf/syscall | (I) 1.54% | 0.98% | -0.55% |
| pts/memtier-benchmark | (I) 2.37% | (I) 2.51% | 0.14% |
| pts/nginx | 0.42% | 0.96% | 0.54% |
| pts/perl-benchmark | (I) 1.83% | (I) 1.63% | -0.20% |
| pts/pgbench | -0.09% | 0.55% | 0.65% |
| pts/pybench | -0.05% | -0.07% | -0.02% |
| pts/redis | -0.00% | -0.04% | -0.04% |
| pts/sqlite-speedtest | 0.15% | 0.44% | 0.28% |
| repro-collection/mysql-workload | 0.73% | 0.25% | -0.48% |
| schbench/thread-contention | 0.42% | (I) 1.04% | 0.61% |
| sockperf/echo-lat-tcp | (I) 2.88% | (I) 1.69% | (R) -1.16% |
| sockperf/echo-lat-udp | (I) 3.93% | (I) 5.11% | (I) 1.14% |
| sockperf/packet-tp-tcp | -0.41% | (I) 1.37% | (I) 1.79% |
| sockperf/packet-tp-udp | (I) 1.45% | (I) 1.49% | 0.04% |
| specjbb/composite | (I) 1.39% | (I) 1.65% | 0.26% |
| speedometer/v2.0 | (R) -1.04% | (R) -1.04% | 0.00% |
| speedometer/v2.1 | -0.10% | 0.10% | 0.20% |
| syscall/getpid | -0.98% | 0.52% | (I) 1.51% |
| syscall/getppid | -0.43% | 0.14% | 0.57% |
| syscall/invalid | (I) 3.06% | (I) 3.56% | 0.48% |
+---------------------------------+--------------------+-----------------+-----------------------+
[1] v7.2
[2] v7.2-percpu-pgtable [a]
[3] v7.2-percpu-gpr (This series) (Based on email review, we fixed SDEI to restore x26
instead of corrupting x22, and adjusted this_cpu_write() helpers to
avoid GCC overflow warnings.)
The constraints/workarounds for the per-CPU page-table series were supplied as
this Kconfig fragment for all the different runs.
CONFIG_EXPERT=y
CONFIG_SMP=y
CONFIG_ARM64_4K_PAGES=y
CONFIG_ARM64_VA_BITS_48=y
CONFIG_ARM64_VA_BITS_39=n
CONFIG_ARM64_VA_BITS_52=n
CONFIG_ARM64_PA_BITS_48=y
CONFIG_ARM64_PA_BITS_52=n
CONFIG_UNMAP_KERNEL_AT_EL0=n
CONFIG_ARM64_SW_TTBR0_PAN=n
CONFIG_KASAN=n
[a] https://lore.kernel.org/all/20260715180455.515692-1-yang@os.amperecomputing.com/
Hence:
Tested-by: Muhammad Usama Anjum <usama.anjum at arm.com>
Regards,
Usama
More information about the linux-arm-kernel
mailing list