[PATCH v8 00/10] nvme-multipath: introduce latency I/O policy

Nilay Shroff nilay at linux.ibm.com
Sun Aug 30 23:31:10 PDT 2026


On 8/31/26 3:23 AM, Sagi Grimberg wrote:
> 
> 
> On 25/08/2026 8:05, Nilay Shroff wrote:
>> On 8/23/26 4:17 AM, Sagi Grimberg wrote:
>>>
>>>
>>> On 15/08/2026 20:34, Nilay Shroff wrote:
>>>> Hi,
>>>>
>>>> This series introduces a new latency I/O policy for NVMe native
>>>> multipath. Existing policies such as numa, round-robin, and queue-depth
>>>> are static and do not adapt to real-time transport performance. The numa
>>>> selects the path closest to the NUMA node of the current CPU, optimizing
>>>> memory and path locality, but ignores actual path performance. The
>>>> round-robin distributes I/O evenly across all paths, providing fairness
>>>> but not performance awareness. The queue-depth reacts to instantaneous
>>>> queue occupancy, avoiding heavily loaded paths, but does not account for
>>>> actual latency, throughput, or link speed.
>>>>
>>>> The new latency policy addresses these gaps selecting paths dynamically
>>>> based on measured I/O latency for both PCIe and fabrics. Latency is
>>>> derived by passively sampling I/O completions. Each path is assigned a
>>>> weight proportional to its latency score, and I/Os are then forwarded
>>>> accordingly. As condition changes (e.g. latency spikes, bandwidth
>>>> differences), path weights are updated, automatically steering traffic
>>>> toward better-performing paths.
>>>>
>>>> Early results show reduced tail latency under mixed workloads and
>>>> improved throughput by exploiting higher-speed links more effectively.
>>>> For example, with NVMf/TCP using two paths (one throttled with ~30 ms
>>>> delay), fio results with random read/write/rw workloads (direct I/O)
>>>> showed:
>>>
>>> TBH, I do not know if this measurement represent any real-life
>>> scenario. I do think that occasional packet drops are a real-life scenario, and
>>> it would be a worthy use-case to optimize for. Can you perhaps measure
>>> how the path selectors compare in this case?
>>>
>>
>> Okay so I have now measured another workload where I simulated packet loss/drops
>> which are occasional and compared different path selectors. In this measurement
>> I have a shared NVMe namespace configured which is reachable over two tcp paths.
>> Now to simulate the occasional packet loss/drop I have configured one of the paths
>> to experience occasional packet loss/drop as shown below over the period of
>> 480 seconds:
>>
>> 0            120           240           360           480
>> |-------------|-------------|-------------|-------------|
>> |<--5% drop-->|<--no drop-->|<--3% drop-->|<--no drop-->|
>>
>> As shown above for the first 120 seconds the path experiences 5% packet drop,
>> for the next 120 seconds path sees no packet drop and again for subsequent 120
>> seconds path experiences 3% packet drop and for the rest of the duration (during
>> last 120 seconds) there's no packet drop observed by the path. With this simulation,
>> I ran fio workload for 480 seconds leveraging direct I/O, bs=4k, iodepth=64,
>> numjobs=32 and ioengine=io_uring. Shown below is bw observed running fio test
>> using different I/O policies:
>>
>>              numa     round-robin queue-depth latency
>>              (MiB/s)  (MiB/s)     (MiB/s)     (MiB/s)
>>              -------  ----------- ----------- ---------
>> randread:    1288     1202        1464        1642
>> randwrite:   1456     1493        1779        1944
>> randrw:      R:623    R:594       R:750       R:822
>>              W:623    W:594       W:750       W:822
>>
>>>>
>>>>          numa         round-robin   queue-depth  adaptive
>>>>          -----------  -----------   -----------  ---------
>>>> READ:   50.0 MiB/s   105 MiB/s     230 MiB/s    350 MiB/s
>>>> WRITE:  65.9 MiB/s   125 MiB/s     385 MiB/s    446 MiB/s
>>>> RW:     R:30.6 MiB/s R:56.5 MiB/s  R:122 MiB/s  R:175 MiB/s
>>>>          W:30.7 MiB/s W:56.5 MiB/s  W:122 MiB/s  W:175 MiB/s
>>>
>>> And I'm assuming there are zero downsides for the normal
>>> case?
>>>
>> For the normal case where all paths are symmetric I saw
>> queue-depth, latency and round-robin policies yielding
>> nearly same bandwidth. However for numa policy, it depends
>> on CPU/numa locality.
>>
>>>>
>>>> This pathcset includes totla 8 patches:
>>>> [PATCH 1/10] block: expose blk_stat_{enable,disable}_accounting()
>>>>    - Make blk_stat APIs available to block drivers.
>>>>    - Needed for per-path latency measurement.
>>>>
>>>> [PATCH 2/10] block: record I/O request start time for passthru request
>>>>    - Record I/O start time for I/O passthru requests.
>>>>    - This is prep patch which allows measuring I/O completion latency
>>>>      for passthru requests.
>>>>
>>>> [PATCH 3/10] block: support nesting for blk-mq flag QUEUE_FLAG_SAME_FORCE
>>>>    - Support nesting for QUEUE_FLAG_SAME_FORCE as multiple users
>>>>      could toggle QUEUE_FLAG_SAME_FORCE.
>>>>
>>>> [PATCH 4/10] nvme-multipath: pass I/O type to nvme_find_path()
>>>>    - This is the prep patch which updates nvme_find_path() signature
>>>> [PATCH 5/10] nvme-multipath: add latency I/O policy
>>>>    - Implement path scoring based on latency (EWMA).
>>>>    - Distribute I/O proportionally to per-path weights.
>>>>
>>>> [PATCH 6/10] nvme: add generic debugfs support
>>>>    - Introduce generic debugfs support for NVMe module
>>>>
>>>> [PATCH 7/10] nvme-multipath: add debugfs attribute latency_ewma_shift
>>>>    - Adds a debugfs attribute to control ewma shift
>>>>
>>>> [PATCH 8/10] nvme-multipath: add debugfs attribute latency_batch_timeout
>>>>    - Adds a debugfs attribute to control latency batch window interval
>>>>
>>>> [PATCH 9/10] nvme-multipath: add debugfs attribute latency_stat
>>>>    - Add “latency_stat” under per-path and head debugfs directories to
>>>>      expose latency policy state and statistics.
>>>>
>>>> [PATCH 10/10] nvme-multipath: add documentation for latency I/O policy
>>>>    - Includes documentation for latency I/O multipath policy.
>>>>
>>>> LSFMM discussion:
>>>> =================
>>>> During lsfmm 2026, it was decided to rename this I/O policy from
>>>> "adaptive" to "latency". This series reflects that rename.
>>>>
>>>> The discussion at lsfmm also focused extensively on the latency
>>>> measurement model, including whether latency should be tracked
>>>> per-CPU or per-NUMA, and whether separate I/O-size buckets should
>>>> be maintained for different request sizes.
>>>>
>>>> After detailed discussion and evaluation of throughput results, the
>>>> consensus was to initially measure I/O completion latency on a
>>>> per-CPU basis. The available performance data showed that the
>>>> per-CPU implementation already provides sufficient averaging across
>>>> CPUs while keeping the design relatively simple.
>>>>
>>>> The use of additional I/O-size buckets did not demonstrate meaningful
>>>> throughput improvement in the general case and would introduce extra
>>>> complexity into the fast path and accounting logic. As a result, the
>>>> consensus was to avoid I/O-size bucketing for now and keep the policy
>>>> focused on per-CPU latency measurement.
>>>>
>>>> If future real-world workloads demonstrate a clear benefit from
>>>> I/O-size-aware latency accounting, the policy can be extended later
>>>> to support it.
>>>>
>>>> As ususal, feedback and suggestions are most welcome!
>>>
>>> Nilay, do we have evidence that round-robin/queue-depth are better
>>> for any workload? As a user, I would be very confused with the amount
>>> of path selectors I have available and which should I choose.
>>
>> From my experiments, when the paths are symmetric, both round-robin and
>> queue-depth (and for that matter latency) exhibit similar behavior, with
>> the workload being distributed roughly equally across the active paths.
>>
>> When the paths are asymmetric, I found queue-depth to perform better than
>> round-robin. Queue-depth tries to steer I/O toward the less-loaded path
>> based on the number of in-flight I/Os on each path, whereas round-robin
>> continues to distribute I/O evenly across all active paths.
>>
>> However, queue-depth still has a limitation in this scenario. It uses
>> the number of in-flight I/Os as an indirect indication of path
>> performance, it does not have a direct signal of the actual I/O
>> completion latency. For example, if one path starts experiencing packet
>> loss, I/O completion on that path can become significantly slower.
>> Queue-depth can react to this as the path accumulates more outstanding
>> I/Os, but it can still continue sending I/O to the degraded path as long
>> as its queue depth remains comparable to the healthy path. In other
>> words, it can reduce the amount of I/O sent to the degraded path, but it
>> cannot directly account for how much slower that path has become.
>>
>> The latency policy uses I/O completion latency as the signal instead.
>> When one path becomes degraded, its observed latency increases and its
>> path score/weight decreases. Consequently, the policy shifts more I/O towards
>> the healthy path. This allows the healthy path to sustain a higher queue
>> depth while the degraded path receives substantially less I/O, rather
>> than trying to maintain a similar queue depth across both paths.
>>
>> This is also reflected in the packet-loss experiment above. Round-robin
>> continues to distribute I/O across both paths, while queue-depth does a
>> better job by reacting to the increased queue occupancy of the degraded
>> path. The latency policy goes one step further by directly using the
>> increased completion latency as a signal and therefore steers more I/O
>> toward the healthy path, resulting in higher throughput.
>>
>> So based on the results I have so far, I would characterize the existing
>> policies as follows: round-robin is useful when paths are symmetric and
>> equal distribution is desired. The queue-depth is preferable when paths are
>> asymmetric and queue occupancy provides a useful indication of path
>> load and the latency policy is intended for cases where path performance
>> can vary dynamically and we want the policy to adapt based on actual
>> observed latency.
> 
> If you have policies A, B, and C and you say:
> - In certain conditions policy C > A, B
> - In some conditions C = B > A
> - In all other cases C = B = A
> 
> This means that C should always be used, and A, B should never be used.
> 
> Hence I ask, should we really have this as an option? or should we deprecate
> round-robin/queue-depth and have only numa|latency?
> 
> I would like to avoid introducing this as a config knob if it is always behaves
> better.
> 
> What do others think?

Good point, however my view is that we may not want to immediately deprecate or
remove queue-depth/round-robin. Based on my testing so far, across a variety of
workloads, including intermittent packet loss, the latency policy has performed
better than queue-depth and round-robin. However, I would like to see it evaluated
against a broader set of workloads and scenarios, including cases that I may not
have considered/known.

Once we have broader evidence and establish that latency consistently performs at
least as well as queue-depth/round-robin without any significant downside, then I
think it would be reasonable to consider deprecating those policies and keeping
only numa|latency.

Until then, I would prefer to keep queue-depth and round-robin as baseline policies
against which we can compare the latency policy.

But let's wait and see what others think.

Thanks,
--Nilay




More information about the Linux-nvme mailing list