[PATCH] nvme-rdma: reject a second completion for a command already answered

Yehyeong Lee yhlee at isslab.korea.ac.kr
Tue Sep 29 01:34:32 PDT 2026


nvme_rdma_process_nvme_rsp() does not always complete the request it just
looked up.  When the request owns a memory region and the response arrived
without a remote invalidation, it posts a local invalidation and returns,
leaving nvme_rdma_inv_rkey_done() to drop the reference later.  Until that
completion runs the request is still live and its generation counter has
not advanced, so a controller that sends a second completion for the same
command_id passes nvme_find_rq() and reaches nvme_rdma_end_request() a
second time.

req->ref is initialised to 2, for the send and the recv completion.  A
duplicate answer produces three decrements -- the send completion, the
duplicate itself, and the deferred local invalidation -- and the last of
them takes the reference count below zero.

The generation counter added by commit e7006de6c238 ("nvme: code
command_id with a genctr for use-after-free validation") does not help
here: it is incremented in nvme_try_complete_req(), so it catches a
completion that arrives after the command has completed, not a second
completion that arrives before the first one has finished.

Both outcomes are a denial of service.  A duplicate that lands inside the
window trips the refcount_t saturation warning, which taints the kernel
and panics it when panic_on_warn is set; one that arrives after the window
instead trips "got bad command_id" and resets the controller.  In the runs
below every stray decrement landed on a request that had completed but had
not yet been reissued, and no KASAN report accompanied any of them, so
this is not a use-after-free with attacker-controlled reuse.

Mark a request once its completion has been consumed and refuse a second
one before it can touch status, result or the reference count.  The flag
is cleared per submission, beside the refcount it protects, so a reused
tag always starts clean.  A command whose duplicate is refused still
terminates normally through the local invalidation that the first response
posted.

Reproduced against a modified nvmet-rdma that answers one command twice,
first with a plain SEND and then with SEND_WITH_INV:

  refcount_t: underflow; use-after-free.
  WARNING: lib/refcount.c:28 at refcount_warn_saturate+0xad/0xe0,
    CPU#0: kworker/u8:0/12
  Modules linked in:
  CPU: 0 UID: 0 PID: 12 Comm: kworker/u8:0 Not tainted
    7.3.0-rc5-nvmerdma-atk-g72d3fcf802c4-dirty #27 PREEMPT(lazy)
  Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix,
    1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
  Workqueue: rxe_wq do_work
  RIP: 0010:refcount_warn_saturate+0xad/0xe0
  Call Trace:
   <IRQ>
   __ib_process_cq+0xe1/0x390
   ? __pfx__raw_spin_lock_irq+0x10/0x10
   ib_poll_handler+0x6e/0x200
   irq_poll_softirq+0x1df/0x480
   ? __pfx_irq_poll_softirq+0x10/0x10
   ? __pfx__raw_spin_lock+0x10/0x10
   handle_softirqs+0x182/0x5a0
   ? __pfx_handle_softirqs+0x10/0x10
   do_softirq+0x3d/0x60
   </IRQ>
   <TASK>
   __local_bh_enable_ip+0x6a/0x70
   __alloc_skb+0x5f5/0x890
   ? enqueue_timer+0xe6/0x3f0
   ? __pfx___alloc_skb+0x10/0x10
   ? _raw_read_unlock_irqrestore+0x16/0x50
   rxe_init_packet+0x16b/0x4f0
   rxe_requester+0xeab/0x51f0
   ? rxe_completer+0x29e5/0x38c0
   ? wakeup_preempt_fair+0x5dd/0xdd0
   ? __pfx_rxe_completer+0x10/0x10
   ? __pfx_rxe_requester+0x10/0x10
   ? wakeup_preempt+0x1d5/0x330
   ? ttwu_do_activate+0x124/0x590
   ? _raw_spin_lock_irqsave+0x86/0xe0
   ? __pfx__raw_spin_lock_irqsave+0x10/0x10
   ? pwq_dec_nr_in_flight+0x2b2/0xc60
   ? __pfx_rxe_sender+0x10/0x10
   rxe_sender+0xe/0x30
   do_work+0x144/0x470
   process_one_work+0x6b4/0x10e0
   ? __pfx___schedule+0x10/0x10
   ? __pfx_process_one_work+0x10/0x10
   ? _raw_spin_lock_irq+0x81/0xe0
   ? __pfx_do_work+0x10/0x10
   ? assign_work+0x11d/0x370
   worker_thread+0x45b/0xd10
   ? __pfx_worker_thread+0x10/0x10
   kthread+0x2c8/0x3b0
   ? recalc_sigpending+0x15c/0x1e0
   ? __pfx_kthread+0x10/0x10
   ret_from_fork+0x36e/0x5a0
   ? __pfx_ret_from_fork+0x10/0x10
   ? __switch_to+0x572/0xdd0
   ? __pfx_kthread+0x10/0x10
   ret_from_fork_asm+0x1a/0x30
   </TASK>
  ---[ end trace 0000000000000000 ]---

The window is narrow and load dependent.  With an idle completion queue
the local invalidation completes in about 85 us and the second answer
arrives too late to matter; under concurrent I/O the invalidation
completion is delayed to milliseconds and the duplicate lands inside it.
Counting the warning is misleading because REFCOUNT_WARN is WARN_ONCE, so
the stray decrements were counted directly:

  unpatched: 18 duplicates accepted inside the window, 18 stray
             decrements on a zero reference count, 1 warning
  patched:   5 duplicates landed inside the window, all refused,
             0 stray decrements, 0 warnings

Every refused command still completed, and an unmodified target is
unaffected with register_always set either way.

Fixes: 2f122e4f5107 ("nvme-rdma: wait for local invalidation before completing a request")
Cc: stable at vger.kernel.org
Signed-off-by: Yehyeong Lee <yhlee at isslab.korea.ac.kr>
---
 drivers/nvme/host/rdma.c | 11 +++++++++++
 1 file changed, 11 insertions(+)

diff --git a/drivers/nvme/host/rdma.c b/drivers/nvme/host/rdma.c
index 9cb811a2ce1f7..f96d421518293 100644
--- a/drivers/nvme/host/rdma.c
+++ b/drivers/nvme/host/rdma.c
@@ -82,6 +82,7 @@ struct nvme_rdma_request {
 	struct nvme_rdma_sgl	data_sgl;
 	struct nvme_rdma_sgl	*metadata_sgl;
 	bool			use_sig_mr;
+	bool			rsp_seen;
 };
 
 enum nvme_rdma_queue_flags {
@@ -1568,6 +1569,7 @@ static int nvme_rdma_map_data(struct nvme_rdma_queue *queue,
 
 	req->num_sge = 1;
 	refcount_set(&req->ref, 2); /* send and recv completions */
+	req->rsp_seen = false;
 
 	c->common.flags |= NVME_CMD_SGL_METABUF;
 
@@ -1738,6 +1740,15 @@ static void nvme_rdma_process_nvme_rsp(struct nvme_rdma_queue *queue,
 	}
 	req = blk_mq_rq_to_pdu(rq);
 
+	if (unlikely(req->rsp_seen)) {
+		dev_err(queue->ctrl->ctrl.device,
+			"Duplicate completion for command_id %#x on QP %#x\n",
+			cqe->command_id, queue->qp->qp_num);
+		nvme_rdma_error_recovery(queue->ctrl);
+		return;
+	}
+	req->rsp_seen = true;
+
 	req->status = cqe->status;
 	req->result = cqe->result;
 
-- 
2.43.0




More information about the Linux-nvme mailing list