[PATCH RFC] nvme-tcp: allow multiple queues per hctx

Saravanan D saravanand at crusoe.ai
Tue Sep 8 22:21:09 PDT 2026


On Tue, 8 Sep 2026 11:10:27 -0600 Keith Busch <kbusch at kernel.org> wrote:
> Initial results show there's some promise to this suggestion, however, I
> think in addition to the wq_unbound, it looks like I also need to adjust
> the "io_cpu" to change to the submitter's CPU. I incorporated that in
> this RFC, but there's also a proposal specifically for that here:
>
>   https://lore.kernel.org/linux-nvme/20260820083634.71689-1-saravanand@crusoe.ai/
>
> I need to catch up on the discussion there, but from what I can tell,
> the cpu hint provided to the queue work appears to be important.

A short summary to save you reading the whole thread. v1 and v2 adopted
the submitting CPU as io_cpu for every command except the fabrics
Connect, opt in per controller. The motivation is a multi tenant host
with more CPUs than controller queues, where the connect time pick can
land a queue's socket work on CPUs owned by a different tenant. 

On a 384 cpu host whose controllers expose 128 io queues, blk-mq folds
three cpus into every map group, 9% of nvme_tcp_io_work executions ran
outside the submitting VM's cpuset by default, and adoption brought
99.99% back, so the cpu hint matters in our case as well.

Sagi suggested replacing the heuristic in the driver with a per queue
writable sysfs attribute so a control plane can set io_cpu exactly, and
my v3 submission will implement that. The user assignment is kept across
reconnects and writing -1 reverts to the connect time selection.
I will post it shortly.

Thanks,
Saravanan D.



More information about the Linux-nvme mailing list