[PATCH v5 04/15] iommu/arm-smmu-v3: Drain in-flight fault events on domain detach
Nicolin Chen
nicolinc at nvidia.com
Wed Sep 23 18:42:42 PDT 2026
On Wed, Sep 23, 2026 at 08:39:20PM -0300, Jason Gunthorpe wrote:
> On Wed, Sep 23, 2026 at 03:33:06PM -0700, Nicolin Chen wrote:
> > > I'm not sure how this all can work, the queue is running on its own
> > > with some other CPU handling interrupts.
> > >
> > > You can't do this sort of Q_DIFF math unless you've somehow guaranteed
> > > one side of the queue is stable for this logic. If both pointers are
> > > moving forward then the points pointers can progress and wrap without
> > > this noticing that happened. That will lock up.
> >
> > One side of the queue (pointers) is actually stable. EVTQ/PRIQ uses
> > the snapshot mode (until_empty=false):
>
> It isn't stable, just because this reads it once doesn't mean the
> actual values are not changing, which is the point.
>
> If one of the pointers is held stable then the HW cannot advance its
> value past it.
Oh, HW pointers cannot be stable, as EVTQ and EVTQ are shared with
other devices that could constantly add new entries onto the queue.
The stunt was to pick a snapshot of CONS/PROD, while HW CONS/PROD
are still advancing. And it would only wait for the length of that
snapshot. The device is detached, so any event after that must be
irrelevant.
> If both are advancing all bets are off and you have no idea how the
> values are related to each other since everything is modulo the ring
> size.
>
> For instance you can read cons0=10, then you read prod=15, then you
> next read prod=11. What does that mean? It means since cons was
> actually advancing prod & cons went around the whole ring and
> wrapped.
>
> You could only do tricks like this if you had full 64 bit counters,
> not truncated versions with modulo that can wrap quickly.
The Q_DIFF was calculated including the WRAP bits. So, it wouldn't
be a problem when any pointer wraps (once).
There is a problem, however, if one of them wraps twice (i.e. 2 x
queue size): then it would miss the exit at the target length even
if it is already much longer; and the penalty would be a timeout,
yet at that moment the queue is definitely drained.
Or maybe I am still missing a key point?
> > > But I wonder if the point of this has been lost? Prior to calling the
> > > driver attach functions the core code already changes the xarray:
> > >
> > > curr = xa_cmpxchg(&group->pasid_array, pasid, NULL,
> > > XA_ZERO_ENTRY, GFP_KERNEL);
> > >
> > > That immediately makes the threaded IRQ safe since it calls
> > > iommu_attach_handle_get() which now fails.
>
> Hmm, actually that's a sneaky cmpxchg that is only doing reserve..
>
> > I am not sure about that. Looking at iommufd_hwpt_replace_device(),
> > there can be a old_handle != NULL, in which case the cmpxchg() would
> > not change the xarray?
>
> I think this is wrong, there is no way it can work like this where the
> attach continues to see the to-be-detached domain across the
> flushes. No amount of flushing can fix it.
>
> Somehow we broke it :\
Hmm, I will take a deeper look.
> > > So all that is needed is to synchronize_irq() to make sure the irq
> > > thread sees the xa update
> > >
> > > Then to flush the workqueue that iommu_report_device_fault() pushes
> > > into.
> > >
> > > We don't need to do anything with the HW queue.
> >
> > FWIW, the idea of HW drain came from intel_iommu_drain_pasid_prq()..
>
> Yeah, but I think they might have over done it too..
I see.
Thanks
Nicolin
More information about the linux-arm-kernel
mailing list