[PATCH v3 0/3] lib: sbi: fix shared memory double-fetch in DBTR and SSE

Himanshu Chauhan himanshu.chauhan at oss.qualcomm.com
Mon Sep 21 22:51:29 PDT 2026


Hi Liutong,

On Mon, Sep 21, 2026 at 08:08:08AM +0000, liutong wrote:
> Hi Himanshu,
> 
> On Sun, Sep 20, 2026 at 10:05:00PM -0700, Himanshu Chauhan wrote:
> > Secondly, even if there is a CPU running rogue code, cannot bring down
> > the system. The triggers programmed may be bad. But again, its a
> > S-mode bug.
> 
> Fair question, and I wasn't sure of the answer myself, so I measured it.
> QEMU 11.1.0-rc3, -M virt -smp 2 -cpu rv64,debug=on, MTTCG, OpenSBI
> 3593a5fa. Hart 0 installs an mcontrol6 EXEC|S trigger with M clear;
> hart 1 rewrites the same tdata1 to M=1.
> 
> Read back via TRIGGER_READ, i.e. csr_read(CSR_TDATA1):
> 
>   0x6000000000000054      type=mcontrol6, M=1, S=1, EXEC=1
> 
>   3593a5fa                : hit in 3 rounds
>   3593a5fa + this series  : no hit in 20000 rounds
> 
> Submitted honestly, the same value returns SBI_ERR_INVALID_PARAM.
> 
> With tdata2 aimed at sbi_ecall_handler, the next SBI call gives:
> 
>   sbi_trap_error: hart0: trap1: trap redirect failed (error -2)
>   sbi_trap_error: hart0: trap1: mcause=0x0000000000000003 mtval=0x0000000080014382
>   sbi_trap_error: hart0: trap1: mepc=0x0000000080014382 mstatus=0x8000000a00007800
> 
> mcause=3 with MPP=3, so it fired in M-mode. An M-mode trap can't be
> redirected, so it ends in sbi_trap_error() and sbi_hart_hang(). The
> nested context below it is the DBTR ecall itself. Aimed at the trap
> vector instead, the hart goes silent with no diagnostic at all: the
> breakpoint sits on the vector, so every entry re-triggers it. An
> honestly installed M=0 trigger at the same address survives 2000 SBI
> calls untouched.
>
Fair enough scenario. But still no reason for the code churn. Ideally, the
S-mode shouldn't be able to set M=1. The patch which fails if M=1 is being set
should be the one here. If OpenSBI is not doing it (which it seems so), please
send that patch.

> You're right that it isn't a machine kill - the trigger CSRs are
> per-hart and the other hart kept running throughout. But the calling
> hart ends on OpenSBI's own unrecoverable path and S-mode can't recover
> it, which is what makes me hesitant to call it an S-mode bug.
> 
> On fixing it in S-mode: I may be misreading what you have in mind. The
> only S-mode-side fix I can see is for S-mode to keep other harts off
> the buffer while the call is in flight. That works if S-mode is merely
> buggy, but the case this series is about is a deliberate one, and an
> S-mode that is attacking M-mode will not adopt that discipline. So I
> can't see how to close it from that side, which is why I went looking
> at the M-mode side instead.

Not allowing M bit from S-mode to be set will take care of above two. I would suggest that
you merge the validation and trigger allocation in one go (by modifying dbtr_find_free_slot()
to return trigger). Rollback from failure is easier by just freeing the allocated trigger.
The installtion should go through. If in any case it fails, set the prior triggers to their
origial value and not zero then as you are doing now.

Regards
Himanshu

> 
> Which turns the whole thing into one question: is DBTR's validation
> meant to hold against an S-mode that is deliberately attacking, or only
> to catch an S-mode that is buggy? That's the one part I can't settle by
> testing, so I'd rather have your answer on it.
> 
> Separately: while testing this I found sbi_dbtr_update_trig() never
> calls dbtr_trigger_valid() at all. S-mode can install a benign trigger,
> then set the M bit on it with TRIGGER_UPDATE - one hart, no race, both
> calls returning err=0, and the bit is live in CSR_TDATA1 afterwards.
> That is independent of this series and I'll send it separately.
> 
> PoC is ~200 lines, I can post it.
> 
> Regards,
> liutong
> 



More information about the opensbi mailing list