[PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency

Mostafa Saleh smostafa at google.com
Thu Aug 13 08:12:03 PDT 2026


On Thu, Aug 13, 2026 at 11:13:20AM -0300, Jason Gunthorpe wrote:
> 
> > Sorry I lost track of this thread and I just saw v4.
> > 
> > In the mobile space, I haven't seen an SMMUv3 that does not support
> > RIL.
> 
> Oh? That's very surprising, AFAIK none of our embedded chips support it
> yet.. Even the server chips are only just getting it. Are you sure?

Yes, SMMUv3 is getting more and more common in moblie HW (opposed to
custom SoC IOMMUs). Devices that I have seen in the market in the
last couple of years have RIL. For example Pixel-10 which is currently
getting upstreamed. Also, I have a mini desktop with QCOM X1 which
have RIL.

The only SMMUv3 I have seen without RIL, is an old morello board I
have.

> 
> I gather it wasn't even available in ARM IP until recently ish?
> 
> > However, I have seen workloads that are really sensitive to translation
> > latency (display, camera...). And I'd be concerned about those
> > regressing.
> 
> Yes, those exist, but again, they are already facing these problems if
> running without RIL.

Yes, but my point is that those typically support RIL and that change
regresses them.

> 
> And, for the common case of putting something into a carve out region
> it is not so likely even an expanded RIL will intersect with a
> reserved IOVA that has a high alignment.

Not necessarily, those devices can run with a small IOVA space to
reduce the page table walk length making IOVAs quite close.

> 
> At least the things we have built are calibrated to handle a TLB
> reload occasionally. The isochronous TLB's are not even sized to be
> never-miss for all cases because things like 4k media require such a
> large amount of IOVA the area cost is too high.
> 
> While others can do something else you are reaching into a pretty
> narrow condition to hit a problem:
>  - HW that must have a never-miss TLB to work

I am not saying that, but we shouldn't over invalidate TLBs either
when it is easy to avoid that.

>  - HW that doesn't have a carve out, or has a badly aligned carve out

I do not think a carveout will help. But it's a very strong constraint
to enforce carveout on all devices specially media which are quite
complex and composite by nature.

>  - A SMMU that has RIL (non RIL is already worse)

This regression only impacts RIL, otherwise it does not matter.

>  - A non-isochronos workload that regularly exceeds the RIL/single
>    expansion thresholds
>  - Unlucky IOVA allocation that places isochronous near other
>    workloads in the IOVA space.
> 

It is not just luck, it depends on the IOVA space and access patterns
of the device.

> > There is a clear trade-off here as you mentioned with TLBI latency,
> > would it be make sense to make that behviour configurable from a
> > module param?
> 
> I think it makes sense for a driver to indicate to the core code that
> it needs isochronous and we can do more global things like change how
> single works as well. Having an isochronous flag on the domain, for
> example, would be a good overall direction.

But according to what? It makes sense to optimize server chips,
but that should not cause over-invalidation regressions on other
hardware.

> 
> I'm inclined to leave this as is and let someone come with a specific
> problematic HW, rather that try to badly guess without much
> information if it might popssibly be a problem.
> 

It's not really a guess, I mentioned some examples above, that I
have seen problems of translation latencies on them.

And why not the other way around:
- Which uses cases can't handle few RIL commands?
- Why those drivers does not unmap memory with a granule/IOVA fitting
  to RIL?
- Why those systems does not use FQ domains in the first place?

Thanks,
Mostafa

> Then we will know the HW and can mark the driver as I suggest above.
> 
> Jason



More information about the linux-arm-kernel mailing list