[PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency

Robin Murphy robin.murphy at arm.com
Thu Aug 13 10:25:59 PDT 2026


On 13/08/2026 4:12 pm, Mostafa Saleh wrote:
> On Thu, Aug 13, 2026 at 11:13:20AM -0300, Jason Gunthorpe wrote:
>>
>>> Sorry I lost track of this thread and I just saw v4.
>>>
>>> In the mobile space, I haven't seen an SMMUv3 that does not support
>>> RIL.
>>
>> Oh? That's very surprising, AFAIK none of our embedded chips support it
>> yet.. Even the server chips are only just getting it. Are you sure?
> 
> Yes, SMMUv3 is getting more and more common in moblie HW (opposed to
> custom SoC IOMMUs). Devices that I have seen in the market in the
> last couple of years have RIL. For example Pixel-10 which is currently
> getting upstreamed. Also, I have a mini desktop with QCOM X1 which
> have RIL.
> 
> The only SMMUv3 I have seen without RIL, is an old morello board I
> have.

Arm MMU-600 is in a few embedded parts - Rockchip RK3588, some of the TI 
Keystone 3s IIRC - but indeed it's mostly all over the first couple of 
generations of Neoverse stuff - Ampere Altra, AWS Graviton2 and so on. 
Similarly the ThunderX2 and early HiSilicon SMMUs also predate RIL, but 
those whole machines are all probably old enough now to no longer be 
widely used.

Arm implementations from MMU-700 onward all support RIL - MMU-700 itself 
was again more of a server-focused design (it's quite big), but has 
found its way into at least some embedded/mobile SoCs; MMU L1 is the 
first design targeted directly at client systems, so we should be seeing 
more of that as time goes on.

>> I gather it wasn't even available in ARM IP until recently ish?

If you'd consider 2019 recent ;)

>>> However, I have seen workloads that are really sensitive to translation
>>> latency (display, camera...). And I'd be concerned about those
>>> regressing.
>>
>> Yes, those exist, but again, they are already facing these problems if
>> running without RIL.
> 
> Yes, but my point is that those typically support RIL and that change
> regresses them.

Not even that - the main concern is over-invalidation of adjacent in-use 
buffers due to rounding up; non-RIL absolutely does not have that issue 
and never has.

>> And, for the common case of putting something into a carve out region
>> it is not so likely even an expanded RIL will intersect with a
>> reserved IOVA that has a high alignment.
> 
> Not necessarily, those devices can run with a small IOVA space to
> reduce the page table walk length making IOVAs quite close.
> 
>>
>> At least the things we have built are calibrated to handle a TLB
>> reload occasionally. The isochronous TLB's are not even sized to be
>> never-miss for all cases because things like 4k media require such a
>> large amount of IOVA the area cost is too high.
>>
>> While others can do something else you are reaching into a pretty
>> narrow condition to hit a problem:
>>   - HW that must have a never-miss TLB to work
> 
> I am not saying that, but we shouldn't over invalidate TLBs either
> when it is easy to avoid that.
> 
>>   - HW that doesn't have a carve out, or has a badly aligned carve out
> 
> I do not think a carveout will help. But it's a very strong constraint
> to enforce carveout on all devices specially media which are quite
> complex and composite by nature.
> 
>>   - A SMMU that has RIL (non RIL is already worse)
> 
> This regression only impacts RIL, otherwise it does not matter.
> 
>>   - A non-isochronos workload that regularly exceeds the RIL/single
>>     expansion thresholds
>>   - Unlucky IOVA allocation that places isochronous near other
>>     workloads in the IOVA space.
>>
> 
> It is not just luck, it depends on the IOVA space and access patterns
> of the device.
> 
>>> There is a clear trade-off here as you mentioned with TLBI latency,
>>> would it be make sense to make that behviour configurable from a
>>> module param?
>>
>> I think it makes sense for a driver to indicate to the core code that
>> it needs isochronous and we can do more global things like change how
>> single works as well. Having an isochronous flag on the domain, for
>> example, would be a good overall direction.
> 
> But according to what? It makes sense to optimize server chips,
> but that should not cause over-invalidation regressions on other
> hardware.
> 
>>
>> I'm inclined to leave this as is and let someone come with a specific
>> problematic HW, rather that try to badly guess without much
>> information if it might popssibly be a problem.
>>
> 
> It's not really a guess, I mentioned some examples above, that I
> have seen problems of translation latencies on them.
> 
> And why not the other way around:
> - Which uses cases can't handle few RIL commands?
> - Why those drivers does not unmap memory with a granule/IOVA fitting
>    to RIL?
> - Why those systems does not use FQ domains in the first place?

Note that with RIL, even precise invalidation should only actually need 
at most two commands (excepting absurd off-the-scale sizes) - the trick 
is to consider that they can overlap.

Conversely though, if we really did have a demonstrable need to minimise 
the number of commands issued then we should probably also not bother 
with the TTL hint nor splitting leaf ranges from non-leaf, such that we 
can gather pretty much any unmap into a single command.

This is probably something we'll end up wanting some kind of tunable 
behaviour for, given that SMMUv3 hardware is going to be spanning an 
increasingly wide range of use-cases with increasingly opposing 
requirements - it's certainly more than just "hypervisor or not". For 
now, though, I'm also not buying a dubious latency argument based on 
apparently no real-world data other than "I think"...

Thanks,
Robin.



More information about the linux-arm-kernel mailing list