[PATCH v2 7/8] iommu/arm-smmu-v3: Change how the tlbi describes the invalidation

Nicolin Chen nicolinc at nvidia.com
Tue Jul 7 22:29:34 PDT 2026


On Mon, Jul 06, 2026 at 01:26:44PM -0300, Jason Gunthorpe wrote:
> + * Compute the stride for non-RIL single-page invalidation. Returns the log2
> + * stride of the lowest affected level. Single invalidation removes all IOPTEs
> + * that contain the IOVA invalidated, and we can reliably assume that the
> + * architected page size and table sizes (not contiguous!) are reflected in the
> + * IOTLB. Thus if there is a 2M leaf entry we only need to issue a single IOTLB
> + * invalidation within that 2M IOVA.
> + */
> +static u8 arm_smmu_tlbi_calc_stride(struct arm_smmu_tlbi *tlbi)
> +{
> +	unsigned int tg_lg2 = tlbi->smmu_domain->tgsz_lg2;
> +	u8 combined = tlbi->table_levels_bitmap | tlbi->leaf_levels_bitmap;
> +
> +	if (!combined)
> +		return U8_MAX;
> +	return (tg_lg2 - 3) * __ffs(combined) + tg_lg2;
> +}
> +
> +/*
> + * One TLBI command per stride-sized entry. Sets use_full_inv if too many

This is raised by Claude; not sure whether it is a false positive
or not.

This changes a previous per-pte invalidation to per-pmd one. Yet,
the spec states in 4.4 TLB invalidation (last paragraph):

  To match a TLB entry, the least significant bits of the address
  are ignored as needed, given the size of the entry.

So, the following scenario would likely miss leaf entries:

VFIO_IOMMU_UNMAP_DMA / IOMMU_IOAS_UNMAP
  iommu_unmap(domain, iova=0x40000000, size=2M)
    __iommu_unmap()
      iommu_pgsize()  -> picks pgsize=2M (aligned, 2M in pgsize_bitmap;
                         irrelevant that the region was mapped as 4K pages)
      arm_lpae_unmap_pages(iova, 2M, pgcount=1, gather)
        __arm_lpae_unmap(lvl=0) -> lvl=1 -> lvl=2:    size == BLOCK_SIZE(lvl 2)
          pte = READ_ONCE(*ptep)         -> a TABLE descriptor, not a block
          __arm_lpae_clear_pte()          # L2 descriptor := 0
          io_pgtable_tlb_flush_walk(iova, 2M, granule=4K)
            arm_smmu_tlb_inv_walk()
              tlbi = { start=0x40000000, last=0x401fffff,
                       table_levels_bitmap = BIT((ilog2(2M)-12)/9) = 0b010,
                       leaf_levels_bitmap  = 0 }              <-- the false claim
              arm_smmu_domain_tlbi()
                arm_smmu_tlbi_calc_single():
                  calc_stride: __ffs(0b010) = 1 -> stride = 2M
                  num_ops = 2M >> 21 = 1
                arm_smmu_domain_tlbi_inv()
                  non-RIL entry: ONE TLBI_NH_VA @0x40000000, Leaf=0
                    -> kills the walk entry + the 4K leaf at base
                    -> 511 4K leaves SURVIVE
          __arm_lpae_free_pgtable()       # subtree freed, no invalidation
    iommu_iotlb_sync(domain, &gather)     # gather->pgsize == 0 -> returns immediately

> +	 * If leaf_levels_bitmap is 0 then this is a walk cache only
> +	 * invalidation.
[...]
> +	u8 leaf_levels_bitmap;

Or is that only to implement a walkcache-only invalidation, such
that the leaf entries will have separate invalidation call(s)?

Nicolin



More information about the linux-arm-kernel mailing list