[PATCH v3] nvme-tcp: allow setting per-queue io_cpu through sysfs
Sagi Grimberg
sagi at grimberg.me
Fri Sep 11 14:11:18 PDT 2026
On 10/09/2026 0:14, Saravanan D wrote:
> nvme_tcp_set_queue_io_cpu() picks each queue's io_cpu at connect time
> as the least loaded CPU in the queue's blk-mq map group, and all socket
> work runs there for the connection's lifetime. This decision falls
> short when the host partitions its CPUs after connect time. On a 384
> cpu multi tenant host with 128 queue controllers, blk-mq folds three
> CPUs into every map group, some groups straddle two tenants' cpusets,
> and 9% of nvme_tcp_io_work executions ran outside the submitting VM's
> cpuset, seen by the neighbor as steal time.
>
> Expose each I/O queue's io_cpu as a writable sysfs attribute
>
> /sys/class/nvme/nvmeX/tcp_queues/<qid>/io_cpu
>
> so a control plane that owns CPU placement can set it directly instead
> of relying on the driver's heuristic. A written value persists across
> reconnects, marked by NVME_TCP_Q_IO_CPU_USER. Writing -1 clears the
> mark and re-runs the connect time selection. Reading returns the CPU,
> or -1 when the queue is unbound.
No need for NVME_TCP_Q_IO_CPU_USER flag and restoring to the original
connect time selection. -1 would mean WORK_CPU_UNBOUND. In fact lets
allow this file to access "unbound" string (same as -1). I don't think
we need to restore the connect time selection, the user who touches
these settings is obviously interested to be in control.
>
> The store, the connect time selection and queue stop serialize their
> accounting of nvme_tcp_cpu_queues under the queue lock.
I think that the queue accounting is something we'd do anyways,
regardless of how the
queue is set. It just represents the io_cpu spread across cpu cores.
>
> Suggested-by: Sagi Grimberg <sagi at grimberg.me>
> Link: https://lore.kernel.org/linux-nvme/220e9da3-f756-4a16-8de1-d4b171f15009@grimberg.me/
> Assisted-by: Claude:claude-opus-4-8 [Claude Code]
> Signed-off-by: Saravanan D <saravanand at crusoe.ai>
> ---
> Changes since v2 [1]:
> - Replaced the io_cpu_adopt connect option and the submitter adoption
> heuristic with a per queue writable sysfs attribute, following Sagi's
> suggestion [2]. The control plane now sets each queue's io_cpu
> directly, the assignment is kept across reconnects and writing -1
> reverts to the connect time selection.
> - Retitled from "nvme-tcp: pin io_cpu to submitter cpu".
>
> The per queue directories follow the blk-mq mq/<hctx> sysfs pattern.
>
> Tested on a 2 socket 384 cpu host with 128 queue controllers. Writing
> a cpu number changed the queue's io_cpu to it. Writing an invalid
> value was rejected. Writing -1 re-ran the connect time selection.
> Pinned queues kept their io_cpu across a controller reset while
> unpinned queues received a fresh pick.
>
> The multiple queues per hctx RFC [3] found the same need to steer the
> socket work cpu, so this attribute may gain a second user.
>
> [1] https://lore.kernel.org/linux-nvme/20260820083634.71689-1-saravanand@crusoe.ai/
> [2] https://lore.kernel.org/linux-nvme/1d56144d-6987-40c1-ac02-b15333db121e@grimberg.me/
> [3] https://lore.kernel.org/linux-nvme/20260903152623.614951-1-kbusch@meta.com/
>
> Documentation/ABI/stable/sysfs-nvme | 12 +++
> drivers/nvme/host/tcp.c | 153 +++++++++++++++++++++++++++-
> 2 files changed, 162 insertions(+), 3 deletions(-)
>
> diff --git a/Documentation/ABI/stable/sysfs-nvme b/Documentation/ABI/stable/sysfs-nvme
> index a2f5d0710db4..2bbb5a0b7c2e 100644
> --- a/Documentation/ABI/stable/sysfs-nvme
> +++ b/Documentation/ABI/stable/sysfs-nvme
> @@ -451,3 +451,15 @@ Contact: Hannes Reinecke <hare at suse.de>
> Description:
> Shows the subsystem type. Possible values: "discovery",
> "nvm", "reserved".
> +
> +What: /sys/class/nvme/nvmeX/tcp_queues/<qid>/io_cpu
> +Date: September 2026
> +KernelVersion: 7.4
> +Contact: Saravanan D <saravanand at crusoe.ai>
> +Description:
> + (RW) The CPU that runs the socket work for I/O queue <qid>
> + of an NVMe over TCP controller, selected by the driver at
> + connect time. Writing a CPU number overrides the selection
> + and persists across reconnects. Writing -1 reverts to the
> + driver's selection. Reads show the current CPU, or -1 when
> + the queue is unbound.
> diff --git a/drivers/nvme/host/tcp.c b/drivers/nvme/host/tcp.c
> index 921934028e0b..22ad1fdf4a12 100644
> --- a/drivers/nvme/host/tcp.c
> +++ b/drivers/nvme/host/tcp.c
> @@ -93,6 +93,7 @@ enum nvme_tcp_queue_flags {
> NVME_TCP_Q_LIVE = 1,
> NVME_TCP_Q_POLLING = 2,
> NVME_TCP_Q_IO_CPU_SET = 3,
> + NVME_TCP_Q_IO_CPU_USER = 4,
> };
>
> enum nvme_tcp_recv_state {
> @@ -102,6 +103,17 @@ enum nvme_tcp_recv_state {
> };
>
> struct nvme_tcp_ctrl;
> +struct nvme_tcp_queue;
> +
> +/*
> + * Allocated per registration and freed by its kobject release, so a
> + * reconnect never reuses a kobject whose release is still pending.
> + */
> +struct nvme_tcp_queue_kobj {
> + struct kobject kobj;
> + struct nvme_tcp_queue *queue;
> +};
I am wandering if it is time to introduce the core nvme queue:
struct nvme_queue {
struct kobject kobj;
unsigned int qid;
u64 flags;
};
And have the transport queue embed it:
struct nvme_tcp_queue {
struct nvme_queue nvmeq;
....
};
I suspect it will allow for better abstractions.
Thoughts? Keith, Christoph?
More information about the Linux-nvme
mailing list