[PATCH] nvme-rdma: don't complete a request after a bogus remote invalidation

Yehyeong Lee yhlee at isslab.korea.ac.kr
Mon Sep 28 08:38:11 PDT 2026


When a remote invalidation names an rkey that is not the one we handed out
for this request, nvme_rdma_process_nvme_rsp() logs and schedules error
recovery, then falls through to nvme_rdma_end_request().  The request is
completed with req->mr never invalidated, and nvme_rdma_unmap_data() hands
it back with ib_mr_pool_put(), which only puts it on a list.  The rkey
stays valid and still maps the pages the completed read filled, so a
controller that sent the bogus invalidation can write into those pages
after pread() has returned.

Return instead, the way the local invalidation path below already does,
and leave the request for error recovery to cancel.

Fixes: 3ef0279bb0031 ("nvme-rdma: Check remotely invalidated rkey matches our expected rkey")
Cc: stable at vger.kernel.org
Signed-off-by: Yehyeong Lee <yhlee at isslab.korea.ac.kr>
---
Reproduced on soft-RoCE (rxe), two nodes.  The target is modified to answer
one request with the rkey of another request that is still in flight, and
then to write 0xBD through the rkey the initiator has just reported as
bogus.  The NVMETHOSTILE lines are that scaffolding and are not part of
this patch.  The namespace holds 0x5A and the reader re-checks its buffer
after pread() has returned.

initiator:
  nvme nvme0: Bogus remote invalidation for rkey 0x35f4
target:
  nvmet_rdma: NVMETHOSTILE forcewr POST LEAKED rkey=0x35f4 addr=0xffff88807ffdd000 len=4096 ret=0 n=1
  nvmet_rdma: NVMETHOSTILE forcewr DONE ACCEPTED n=1
initiator:
  BCHECK[B] CORRUPT iter=350 poll=0ms pre_clean=1 off=0 p1_bytes=4096 p0_bytes=61440
  BCHECK[B] DONE iters=392 clean_p0=392 ioerr=0 CORRUPT_HITS=1 POST_COMPLETION=1

pre_clean=1 is the buffer intact at the instant pread() returned, and
p0_bytes=61440 is the remainder of it still holding 0x5A.  0x35f4 is the
rkey the initiator reported as bogus and the rkey the target then wrote
through.  With this patch POST_COMPLETION stayed 0 over five runs.

It is a race, so it is probabilistic rather than reliable, and every bogus
invalidation also trips error recovery, so the corruption arrives together
with a controller reset.  I only have rxe here; I have not checked how a
real RoCE HCA responder treats a write through an MR the ULP has stopped
using.  What I saw is an integrity violation on a buffer that is still
mapped, not a use-after-free.

Invalidating the MR locally before completing the request, which is what
the code did before 3ef0279bb0031, also closes it and is the other way to
go here.

 drivers/nvme/host/rdma.c | 1 +
 1 file changed, 1 insertion(+)

diff --git a/drivers/nvme/host/rdma.c b/drivers/nvme/host/rdma.c
index 9cb811a2ce1f7..5b9b540693b51 100644
--- a/drivers/nvme/host/rdma.c
+++ b/drivers/nvme/host/rdma.c
@@ -1748,6 +1748,7 @@ static void nvme_rdma_process_nvme_rsp(struct nvme_rdma_queue *queue,
 				"Bogus remote invalidation for rkey %#x\n",
 				req->mr ? req->mr->rkey : 0);
 			nvme_rdma_error_recovery(queue->ctrl);
+			return;
 		}
 	} else if (req->mr) {
 		int ret;
-- 
2.43.0




More information about the Linux-nvme mailing list