dma_opt_mapping_size returns way too low sizes when using IOMMU
Christoph Hellwig
hch at lst.de
Mon Aug 17 02:18:40 PDT 2026
On Mon, Aug 17, 2026 at 10:12:25AM +0100, John Garry wrote:
> On 17/08/2026 09:36, Christoph Hellwig wrote:
>> Hi all,
>>
>> I got reports that NVMe devices were arbitrarily limited to 128kiB
>> transfers in recent kernel.
>
> How recent a kernel? This NVMe and DMA mapping code has not changed in
> years as far as I know.
This was hardware QA moving from an old distro kernel to a "recent" (aka
still old) one. But I've actually reproduced it locally on an
upstream kernel. I guess most kernel developers or power users simply
do not run with IOMMU enabled.
>> Both the NVMe performance numbers and common sense suggest that this
>> is NOT the optimal DMA mapping granularity. Can we pick a saner value
>> for iommu_dma_opt_mapping_size that does not restrict common I/O sizes?
>
> Note that SCSI does not use iommu_dma_opt_mapping_size() for clamping max
> HW sectors, but instead sets opt size / max sectors from this value (so it
> is not a hard limit there). Could we consider similar for NVMe?
We could consider that, but it would still reduce performane..
> The reason for which we have iommu_dma_opt_mapping_size() is that
> performance can go through the floor we can't use the rcache for getting
> the IOVA, i.e. we need to always alloc and dealloc an IOVA from the RB tree
> for each mapping, and this can greatly reduce performance when the IOVA
> space fills.
>
> If we increase IOVA_RANGE_CACHE_MAX_SIZE, then we just get caching of
> larger IOVAs and I am not sure that is a great idea.
At least on the four different nvme devices I tested, the larger I/O
sizes made up for this. But maybe the details depend on other
factors as well.
More information about the Linux-nvme
mailing list