[PATCH v2 13/20] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops
David Laight
david.laight.linux at gmail.com
Thu Aug 6 06:25:27 PDT 2026
On Thu, 6 Aug 2026 13:32:52 +0200
"David Hildenbrand (Arm)" <david at kernel.org> wrote:
> On 8/6/26 13:21, Mark Rutland wrote:
> > On Wed, Aug 05, 2026 at 08:47:08AM +0200, David Hildenbrand (Arm) wrote:
> >> On 8/5/26 08:45, David Hildenbrand (Arm) wrote:
> >>>
> >>> FWIW, in a recent discussion on some prototype hacking [1] we saw some overhead
> >>> in micro-benchmarks that would really hammer on a path that would now do a
> >>> preempt_disable()+preempt_enable().
> >>>
> >>> Switching from preempt_disable() to preempt_enable_no_resched() made it turn to
> >>> noise. Of course, that has other undesirable impacts, and I am not sure if we
> >>> are in the territory of code layout changes affecting the numbers.
> >>>
> >>> Just mentioning it as some data point.
> >>
> >> [1] https://lore.kernel.org/linux-mm/20260630174852-mutt-send-email-mst@kernel.org/
> >
> > Thanks for the pointer.
> >
> > IIUC in those cases you're using preempt_disable() .. preempt_enable()
> > directly, not this_cpu_*(), right?
>
> It was purely preempt_disable/preempt_enable experiments without any percpu stuff.
Did you check that preempt_enable() isn't likely to speculatively execute the
schedule() call.
Even if you write:
if (unlikely(a == b))
function();
the compiler tends to generate a forwards branch around the function call.
Since the branch is likely to be assumed 'not taken' the cpu will
speculatively execute the function.
Adding a non-empty else clause (eg an asm() comment) should get the function
call out of line and hopefully not speculatively called.
This is likely made worse because the condition is reading the full 64bits
of a location that has just had 32bits written.
This almost certainly has to wait for the write to 'drain' from the store
buffer before the read can be done from the D-cache.
(A read of the same/smaller size might be snooped from the store buffer.)
I'm not sure of the mis-predict penalty for a typical arm cpu.
I see ~20 clocks on a Zen-5 for a simple (value in register) one, here
I suspect an extra 5-10 clocks get added because of the memory accesses.
David
>
> >
> > If so, patches 5 and 6 of this series [2,3] might have an impact, but I
> > wouldn't expect a significant change unless you're calling
> > preempt_enable a lot.
> >
> > Please beware that it's not safe to use preempt_enable_no_resched()
> > UNLESS it is immediately followed by a call to schedule(). That's not
> > documented today (and I couldn't find a good reference), so more folk
> > are likely to be tempted to use it...
> Yes, that's also why we abandoned that (including for various other reasons :) ).
>
> preempt_enable_no_resched() helped to identify that the preempt_enable() was
> really causing the noticeable overhead, not the other minor stuff we added on
> some hot paths.
>
More information about the linux-arm-kernel
mailing list