[PATCH v5 4/9] iommu/arm-smmu-v3: Optimize range invalidation for latency
Jason Gunthorpe
jgg at nvidia.com
Mon Sep 7 08:23:05 PDT 2026
On Mon, Sep 07, 2026 at 02:26:16PM +0000, Mostafa Saleh wrote:
> > +static void arm_smmu_cmdq_batch_add_ril(struct arm_smmu_device *smmu,
> > + struct arm_smmu_cmdq_batch *cmds,
> > + struct arm_smmu_cmd *ref_cmd,
> > + bool leaf_only,
> > + const struct arm_smmu_ril_range *ril,
> > + u8 ttl, u8 tg_enc)
> > +{
> > + struct arm_smmu_cmd cmd;
> > + unsigned int tgsz_lg2 = tg_enc * 2 + 10;
> > + u64 iova = ril->start_tg << tgsz_lg2;
> > + unsigned int num = ril->num - 1;
> > +
> > + /* 16K granule TTL=1 is reserved (Section 4.4.1) */
> > + if (WARN_ON(tgsz_lg2 == 14 && ttl == 1))
> > + ttl = 0;
> > +
> > + /* Verify address alignment for the TTL hint */
> > + if (ttl && !arm_smmu_ttl_addr_aligned(iova, tgsz_lg2, ttl))
> > + ttl = 0;
>
> Is that a separate fix? Can this be a separate patch in that case.
I don't know if it is because:
> Also, in what cases can this happen? io-pgtable-arm should do the
> right thing and SVA will use ttl=0
This is more getting things setup right for iommupt in the next series
which creates every imaginable kind of gather.
> > - /* Determine how many chunks of 2^scale size we have */
> > - num = (num_pages >> scale) & CMDQ_TLBI_RANGE_NUM_MAX;
> > -
> > - /* Keep the pre-DS 5-bit truncation when scale > 31 */
> > - cmd->data[0] = orig_data0 |
> > - FIELD_PREP(CMDQ_TLBI_0_NUM, num - 1) |
> > - FIELD_PREP(CMDQ_TLBI_0_SCALE, scale & 0x1f);
>
> That just deletes the DS code which was recently merged.
> I see that is added again in the last patch, but it would make more
> sense to do it in the same patch I guess.
I missed it when rebasing, I'll adjust it
Thanks,
Jason
More information about the linux-arm-kernel
mailing list