[PATCH v2 1/2] riscv: kprobes: simulate nop and c.nop instructions

袁晓峰 yuanxiaofeng at eswincomputing.com
Wed Sep 2 01:43:14 PDT 2026


From: Xiaofeng Yuan <yuanxiaofeng at eswincomputing.com>
To: Nam Cao <namcao at linutronix.de>
Cc: Paul Walmsley <pjw at kernel.org>, Palmer Dabbelt <palmer at dabbelt.com>,
Albert Ou <aou at eecs.berkeley.edu>, Nam Cao <namcaov at gmail.com>,
linux-riscv at lists.infradead.org
Subject: Re: [PATCH v2 1/2] riscv: kprobes: simulate nop and c.nop instructions
In-Reply-To: <87wltdt4md.fsf at yellow.woof>


Hi Nam,


Thanks for the detailed review. You were right on several points; let
me take them in order.


Xiaofeng Yuan <yuanxiaofeng at eswincomputing.com> writes:
> > nop and c.nop have no architectural effect, so allocating an
> [..]
> > pure overhead. [...] following the approach already
> > used on arm64.
>
> That is not the reason why arm64 simulate nop. And how is single
> stepping slower than simulating?


You are right about the first part: I should not have paraphrased the
arm64 motivation.  ac4ad5c09b34 ("arm64: insn: Simulate nop
instruction for better uprobe performance") was motivated by Andrii's
uprobe benchmarks, where a probed nop was about 2x slower than an
already-emulated instruction.


On the second question, concretely what the extra cost is on riscv:
the "single-stepping" here is done by the XOL slot -- the breakpoint
trap runs the handler and redirects execution to a copy of the
probed instruction with another ebreak appended after it, which
traps back into the kernel a second time just to complete the step.
So each hit of a probed nop costs two exception round-trips (full
pt_regs + sret each) for the slot path versus one for simulation.
For uprobes it is worse still: the slot runs in user mode, so the
re-trap is a full kernel<->userspace round-trip.  I measured
this on QEMU (RISC-V virt): a ftrace uprobe event on a
USDT-style nop site reduces the per-hit cost by roughly 80
percent compared to the XOL path; a kprobe on a nop by roughly
40 percent.  (arm64 commit ac4ad5c09b34 measured ~2x on real
hardware for the same class of change.)  v3 carries these
measurements instead of the arm64 rationale I mis-stated.


> > [...] This
> > is relevant for USDT probe sites in user-space binaries, which are
> [...]
>
> I can only find https://github.com/chrisa/libusdt, which does not have
> riscv support. Which USDT are you referring to?


Sorry for the vague shorthand -- by USDT I mean SystemTap-style SDT
probes (the .note.stapsdt format parsed by libbpf and used by tools
like bpftrace).  On riscv a probe site is an SDT note whose recorded
Location is a plain nop by construction.  In-tree evidence:


- tools/testing/selftests/bpf/sdt.h emits the location as
  "990: _SDT_NOP"; _SDT_NOP is specialized only for ia64/s390
  ("nop"/"nop 0"), so riscv takes the generic "nop" path;
- libbpf documents the model: "USDT call is actually not a function
  call, but is instead replaced by a single NOP instruction ... [the]
  NOP instruction that kernel can replace with an interrupt
  instruction" (tools/lib/bpf/usdt.c), and it already has riscv
  specific note/argument parsing (the register map in
  tools/lib/bpf/usdt.c);
- the bpf selftests USDT provider (urandom_read.c: STAP_PROBE1/3)
  has riscv-specific build support
  (tools/testing/selftests/bpf/Makefile);
- and the arm64 series ac4ad5c09b34 states it itself: "Typicall uprobe
  is installed on 'nop' for USDT".


To make it concrete I cross-compiled a provider with the in-tree
header (riscv64-linux-musl-gcc -O2 -static):


  $ readelf -n usdt_probe | grep -A4 stapsdt
  [...]
      Provider: "bench"
      Name: "nop_site"
      Location: 0x00000000000006ba, Base: 0x0000000000001052
  $ objdump -d usdt_probe | grep -A2 '<probe_site>:'
  00000000000006ba <probe_site>:
    6ba:  0001      nop          (rv64gc: compressed c.nop)
    6bc:  8082      ret


and with -march=rv64imafd (no RVC):


    6c8:  00000013  nop          (4-byte nop)


so this is the ABI-defined instruction at every riscv USDT site --
exactly the two instructions this patch simulates.


To be honest about scope: I have not verified how widely USDT is
deployed on riscv workloads today.  My claim is the narrower one:
libbpf parses riscv SDT notes, the bpf selftests' USDT provider has
riscv-specific build handling, and the sdt.h macro body emits the
note location as a plain nop/c.nop by construction -- so USDT is the
intended consumer of probed nops on riscv, the same model the arm64
series describes: "Typicall uprobe is installed on 'nop' for USDT".


> > In kernel text, nops are also found at ftrace
> > mcount call sites and disabled jump_label sites.
>
> Not sure about mcount, but isn't installing kprobe on jump labels
> forbidden?


You are right, and I dropped both claims in v3.  register_kprobe()
rejects jump_label text sites via jump_label_text_reserved()
(kernel/kprobes.c), which covers the site regardless of whether it
currently holds a nop or a JAL (and static_call sites likewise), so
those nops are not natural probe candidates at all.  The mcount claim
was misleading even where it is allowed: on riscv the function entry
instruction is the auipc of the mcount pair; the nop is only at +4.


> > Measured on QEMU [...]
> > roughly a 2x [...]


> Reading the arm64's commit, simulating nop should only improve uprobe,
> not kprobe; or am I confused somewhere?


Not confused about the motivation -- on arm64 the benchmark and the
title are uprobe-specific, and the common real-world producer of
probed nops is also uprobes (USDT sites), as above.  The cost model
just doesn't stop at kprobes on our side: riscv_probe_decode_insn()
is shared by arch_prepare_kprobe() and arch_uprobe_analyze_insn(),
and the kprobe path is likewise a slot-replay-with-re-trap (second
breakpoint round-trip), so simulating the nop removes one exception
round-trip per hit of a probed nop for kprobes too.  But your
framing is fair: the natural producer of this traffic is uprobes
on USDT sites, so v3 leads with the uprobe measurement (~50us ->
~9us per hit, QEMU) and keeps the kprobe reduction (~40 percent)
as the secondary number for the shared decode path.


v3 is out with these corrections; the diff itself is unchanged.


Best regards,
Xiaofeng Yuan


More information about the linux-riscv mailing list