[PATCH] nvme-rdma: reject a second completion for a command already answered
Yehyeong Lee
yhlee at isslab.korea.ac.kr
Tue Sep 29 01:34:32 PDT 2026
nvme_rdma_process_nvme_rsp() does not always complete the request it just
looked up. When the request owns a memory region and the response arrived
without a remote invalidation, it posts a local invalidation and returns,
leaving nvme_rdma_inv_rkey_done() to drop the reference later. Until that
completion runs the request is still live and its generation counter has
not advanced, so a controller that sends a second completion for the same
command_id passes nvme_find_rq() and reaches nvme_rdma_end_request() a
second time.
req->ref is initialised to 2, for the send and the recv completion. A
duplicate answer produces three decrements -- the send completion, the
duplicate itself, and the deferred local invalidation -- and the last of
them takes the reference count below zero.
The generation counter added by commit e7006de6c238 ("nvme: code
command_id with a genctr for use-after-free validation") does not help
here: it is incremented in nvme_try_complete_req(), so it catches a
completion that arrives after the command has completed, not a second
completion that arrives before the first one has finished.
Both outcomes are a denial of service. A duplicate that lands inside the
window trips the refcount_t saturation warning, which taints the kernel
and panics it when panic_on_warn is set; one that arrives after the window
instead trips "got bad command_id" and resets the controller. In the runs
below every stray decrement landed on a request that had completed but had
not yet been reissued, and no KASAN report accompanied any of them, so
this is not a use-after-free with attacker-controlled reuse.
Mark a request once its completion has been consumed and refuse a second
one before it can touch status, result or the reference count. The flag
is cleared per submission, beside the refcount it protects, so a reused
tag always starts clean. A command whose duplicate is refused still
terminates normally through the local invalidation that the first response
posted.
Reproduced against a modified nvmet-rdma that answers one command twice,
first with a plain SEND and then with SEND_WITH_INV:
refcount_t: underflow; use-after-free.
WARNING: lib/refcount.c:28 at refcount_warn_saturate+0xad/0xe0,
CPU#0: kworker/u8:0/12
Modules linked in:
CPU: 0 UID: 0 PID: 12 Comm: kworker/u8:0 Not tainted
7.3.0-rc5-nvmerdma-atk-g72d3fcf802c4-dirty #27 PREEMPT(lazy)
Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix,
1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
Workqueue: rxe_wq do_work
RIP: 0010:refcount_warn_saturate+0xad/0xe0
Call Trace:
<IRQ>
__ib_process_cq+0xe1/0x390
? __pfx__raw_spin_lock_irq+0x10/0x10
ib_poll_handler+0x6e/0x200
irq_poll_softirq+0x1df/0x480
? __pfx_irq_poll_softirq+0x10/0x10
? __pfx__raw_spin_lock+0x10/0x10
handle_softirqs+0x182/0x5a0
? __pfx_handle_softirqs+0x10/0x10
do_softirq+0x3d/0x60
</IRQ>
<TASK>
__local_bh_enable_ip+0x6a/0x70
__alloc_skb+0x5f5/0x890
? enqueue_timer+0xe6/0x3f0
? __pfx___alloc_skb+0x10/0x10
? _raw_read_unlock_irqrestore+0x16/0x50
rxe_init_packet+0x16b/0x4f0
rxe_requester+0xeab/0x51f0
? rxe_completer+0x29e5/0x38c0
? wakeup_preempt_fair+0x5dd/0xdd0
? __pfx_rxe_completer+0x10/0x10
? __pfx_rxe_requester+0x10/0x10
? wakeup_preempt+0x1d5/0x330
? ttwu_do_activate+0x124/0x590
? _raw_spin_lock_irqsave+0x86/0xe0
? __pfx__raw_spin_lock_irqsave+0x10/0x10
? pwq_dec_nr_in_flight+0x2b2/0xc60
? __pfx_rxe_sender+0x10/0x10
rxe_sender+0xe/0x30
do_work+0x144/0x470
process_one_work+0x6b4/0x10e0
? __pfx___schedule+0x10/0x10
? __pfx_process_one_work+0x10/0x10
? _raw_spin_lock_irq+0x81/0xe0
? __pfx_do_work+0x10/0x10
? assign_work+0x11d/0x370
worker_thread+0x45b/0xd10
? __pfx_worker_thread+0x10/0x10
kthread+0x2c8/0x3b0
? recalc_sigpending+0x15c/0x1e0
? __pfx_kthread+0x10/0x10
ret_from_fork+0x36e/0x5a0
? __pfx_ret_from_fork+0x10/0x10
? __switch_to+0x572/0xdd0
? __pfx_kthread+0x10/0x10
ret_from_fork_asm+0x1a/0x30
</TASK>
---[ end trace 0000000000000000 ]---
The window is narrow and load dependent. With an idle completion queue
the local invalidation completes in about 85 us and the second answer
arrives too late to matter; under concurrent I/O the invalidation
completion is delayed to milliseconds and the duplicate lands inside it.
Counting the warning is misleading because REFCOUNT_WARN is WARN_ONCE, so
the stray decrements were counted directly:
unpatched: 18 duplicates accepted inside the window, 18 stray
decrements on a zero reference count, 1 warning
patched: 5 duplicates landed inside the window, all refused,
0 stray decrements, 0 warnings
Every refused command still completed, and an unmodified target is
unaffected with register_always set either way.
Fixes: 2f122e4f5107 ("nvme-rdma: wait for local invalidation before completing a request")
Cc: stable at vger.kernel.org
Signed-off-by: Yehyeong Lee <yhlee at isslab.korea.ac.kr>
---
drivers/nvme/host/rdma.c | 11 +++++++++++
1 file changed, 11 insertions(+)
diff --git a/drivers/nvme/host/rdma.c b/drivers/nvme/host/rdma.c
index 9cb811a2ce1f7..f96d421518293 100644
--- a/drivers/nvme/host/rdma.c
+++ b/drivers/nvme/host/rdma.c
@@ -82,6 +82,7 @@ struct nvme_rdma_request {
struct nvme_rdma_sgl data_sgl;
struct nvme_rdma_sgl *metadata_sgl;
bool use_sig_mr;
+ bool rsp_seen;
};
enum nvme_rdma_queue_flags {
@@ -1568,6 +1569,7 @@ static int nvme_rdma_map_data(struct nvme_rdma_queue *queue,
req->num_sge = 1;
refcount_set(&req->ref, 2); /* send and recv completions */
+ req->rsp_seen = false;
c->common.flags |= NVME_CMD_SGL_METABUF;
@@ -1738,6 +1740,15 @@ static void nvme_rdma_process_nvme_rsp(struct nvme_rdma_queue *queue,
}
req = blk_mq_rq_to_pdu(rq);
+ if (unlikely(req->rsp_seen)) {
+ dev_err(queue->ctrl->ctrl.device,
+ "Duplicate completion for command_id %#x on QP %#x\n",
+ cqe->command_id, queue->qp->qp_num);
+ nvme_rdma_error_recovery(queue->ctrl);
+ return;
+ }
+ req->rsp_seen = true;
+
req->status = cqe->status;
req->result = cqe->result;
--
2.43.0
More information about the Linux-nvme
mailing list