[PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency
Mostafa Saleh
smostafa at google.com
Thu Aug 13 08:12:03 PDT 2026
On Thu, Aug 13, 2026 at 11:13:20AM -0300, Jason Gunthorpe wrote:
>
> > Sorry I lost track of this thread and I just saw v4.
> >
> > In the mobile space, I haven't seen an SMMUv3 that does not support
> > RIL.
>
> Oh? That's very surprising, AFAIK none of our embedded chips support it
> yet.. Even the server chips are only just getting it. Are you sure?
Yes, SMMUv3 is getting more and more common in moblie HW (opposed to
custom SoC IOMMUs). Devices that I have seen in the market in the
last couple of years have RIL. For example Pixel-10 which is currently
getting upstreamed. Also, I have a mini desktop with QCOM X1 which
have RIL.
The only SMMUv3 I have seen without RIL, is an old morello board I
have.
>
> I gather it wasn't even available in ARM IP until recently ish?
>
> > However, I have seen workloads that are really sensitive to translation
> > latency (display, camera...). And I'd be concerned about those
> > regressing.
>
> Yes, those exist, but again, they are already facing these problems if
> running without RIL.
Yes, but my point is that those typically support RIL and that change
regresses them.
>
> And, for the common case of putting something into a carve out region
> it is not so likely even an expanded RIL will intersect with a
> reserved IOVA that has a high alignment.
Not necessarily, those devices can run with a small IOVA space to
reduce the page table walk length making IOVAs quite close.
>
> At least the things we have built are calibrated to handle a TLB
> reload occasionally. The isochronous TLB's are not even sized to be
> never-miss for all cases because things like 4k media require such a
> large amount of IOVA the area cost is too high.
>
> While others can do something else you are reaching into a pretty
> narrow condition to hit a problem:
> - HW that must have a never-miss TLB to work
I am not saying that, but we shouldn't over invalidate TLBs either
when it is easy to avoid that.
> - HW that doesn't have a carve out, or has a badly aligned carve out
I do not think a carveout will help. But it's a very strong constraint
to enforce carveout on all devices specially media which are quite
complex and composite by nature.
> - A SMMU that has RIL (non RIL is already worse)
This regression only impacts RIL, otherwise it does not matter.
> - A non-isochronos workload that regularly exceeds the RIL/single
> expansion thresholds
> - Unlucky IOVA allocation that places isochronous near other
> workloads in the IOVA space.
>
It is not just luck, it depends on the IOVA space and access patterns
of the device.
> > There is a clear trade-off here as you mentioned with TLBI latency,
> > would it be make sense to make that behviour configurable from a
> > module param?
>
> I think it makes sense for a driver to indicate to the core code that
> it needs isochronous and we can do more global things like change how
> single works as well. Having an isochronous flag on the domain, for
> example, would be a good overall direction.
But according to what? It makes sense to optimize server chips,
but that should not cause over-invalidation regressions on other
hardware.
>
> I'm inclined to leave this as is and let someone come with a specific
> problematic HW, rather that try to badly guess without much
> information if it might popssibly be a problem.
>
It's not really a guess, I mentioned some examples above, that I
have seen problems of translation latencies on them.
And why not the other way around:
- Which uses cases can't handle few RIL commands?
- Why those drivers does not unmap memory with a granule/IOVA fitting
to RIL?
- Why those systems does not use FQ domains in the first place?
Thanks,
Mostafa
> Then we will know the HW and can mark the driver as I suggest above.
>
> Jason
More information about the linux-arm-kernel
mailing list