[PATCH v8 04/10] nvme-multipath: pass I/O type to nvme_find_path()
Nilay Shroff
nilay at linux.ibm.com
Sat Sep 12 03:22:13 PDT 2026
On 9/12/26 4:08 AM, Keith Busch wrote:
> On Sat, Aug 15, 2026 at 11:04:26PM +0530, Nilay Shroff wrote:
>> @@ -804,12 +830,24 @@ long nvme_ns_head_chr_ioctl(struct file *file, unsigned int cmd,
>> int nvme_ns_head_chr_uring_cmd(struct io_uring_cmd *ioucmd,
>> unsigned int issue_flags)
>> {
>> + struct nvme_ns *ns;
>> + unsigned int op_type;
>> struct cdev *cdev = file_inode(ioucmd->file)->i_cdev;
>> struct nvme_ns_head *head = container_of(cdev, struct nvme_ns_head, cdev);
>> int srcu_idx = srcu_read_lock(&head->srcu);
>> - struct nvme_ns *ns = nvme_find_path(head);
>> int ret = -EINVAL;
>> + const struct nvme_uring_cmd *cmd = io_uring_sqe128_cmd(ioucmd->sqe,
>> + struct nvme_uring_cmd);
>> + __u8 opcode = READ_ONCE(cmd->opcode);
>
> Just fyi, the READ_ONCE is used because this is shared memory with user
> space. We re-read the same field later when constructing the command, so
> if the user does something tricky like changing that opcode, you can
> have the wrong group for the command that actually gets dispatched.
Yes, agreed. If userspace modifies cmd->opcode while the SQE is being consumed
by the kernel, we could end up selecting a path based on one opcode value and
dispatching a command with another. However, per the io_uring ABI/protocol,
userspace must not modify or reuse an SQE after publishing it to the kernel
until the kernel has consumed it (i.e. advanced the SQ ring head). Therefore,
under the expected protocol, the SQE contents are stable while the kernel is
consuming them.
Nevertheless, if userspace modifies an SQE before it has been consumed by the
kernel, then that violates the io_uring SQE ownership protocol, and the
resulting behavior is not something we need to account for here, IMO.
Thanks,
--Nilay
More information about the Linux-nvme
mailing list