[PATCH v2 13/20] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops

Mark Rutland mark.rutland at arm.com
Thu Aug 6 05:02:56 PDT 2026


On Thu, Aug 06, 2026 at 01:32:52PM +0200, David Hildenbrand (Arm) wrote:
> On 8/6/26 13:21, Mark Rutland wrote:
> > On Wed, Aug 05, 2026 at 08:47:08AM +0200, David Hildenbrand (Arm) wrote:
> >> On 8/5/26 08:45, David Hildenbrand (Arm) wrote:
> >>>
> >>> FWIW, in a recent discussion on some prototype hacking [1] we saw some overhead
> >>> in micro-benchmarks that would really hammer on a path that would now do a
> >>> preempt_disable()+preempt_enable().
> >>>
> >>> Switching from preempt_disable() to preempt_enable_no_resched() made it turn to
> >>> noise. Of course, that has other undesirable impacts, and I am not sure if we
> >>> are in the territory of code layout changes affecting the numbers.
> >>>
> >>> Just mentioning it as some data point.
> >>
> >> [1] https://lore.kernel.org/linux-mm/20260630174852-mutt-send-email-mst@kernel.org/
> > 
> > Thanks for the pointer.
> > 
> > IIUC in those cases you're using preempt_disable() .. preempt_enable()
> > directly, not this_cpu_*(), right?
> 
> It was purely preempt_disable/preempt_enable experiments without any percpu stuff.
> 
> > If so, patches 5 and 6 of this series [2,3] might have an impact, but I
> > wouldn't expect a significant change unless you're calling
> > preempt_enable a lot.
> > 
> > Please beware that it's not safe to use preempt_enable_no_resched()
> > UNLESS it is immediately followed by a call to schedule(). That's not
> > documented today (and I couldn't find a good reference), so more folk
> > are likely to be tempted to use it...
> Yes, that's also why we abandoned that (including for various other reasons :) ).

:)

> preempt_enable_no_resched() helped to identify that the preempt_enable() was
> really causing the noticeable overhead, not the other minor stuff we added on
> some hot paths.

Understood!
	
If we seeeing particularly noticeable overhead from preempt_enable() in
some workloads, there are some options we could investigate to reduce
that impact (e.g. using __preserve_most or a trampoline like x86's
preempt_schedule_thunk to reduce necessary spills and register
pressure).

Please let me know if you see anything that stands out, as any examples
would be useful for investigation. We'd want to figure out how much of
the overhead comes from register pressure, and how much of it comes from
the conditional call itself.

If you're testing with PREEMPT_DYNAMIC=y, on some architectures
(including arm64) you might see overhead reduced by:

  https://lore.kernel.org/lkml/20260803191731.3244294-1-mark.rutland@arm.com/

... but IIUC on x86 that won't change the cost of
preempt_enable[_notrace](). Today that makes a static call to
preempt_schedule[_notrace]_thunk, and a plain call will be the same
cost.

Mark.



More information about the linux-arm-kernel mailing list