[PATCH v6] mm: retry page faults once under the per-VMA lock
Matthew Wilcox
willy at infradead.org
Thu Oct 1 14:12:16 PDT 2026
On Mon, Sep 28, 2026 at 10:48:41AM +0800, Barry Song wrote:
> On Mon, Sep 28, 2026 at 6:55 AM Matthew Wilcox <willy at infradead.org> wrote:
> > So while doing my slides, I realised that what we need to avoid doing
> > is (a) sleeping while holding the mmap_lock (b) returning RETRY while
> > holding the VMA lock
> >
> > And that turns out to be as simple as this patch:
> >
> > diff --git a/include/linux/mm.h b/include/linux/mm.h
> > index dd09c438fa23..94ed2333f8d8 100644
> > --- a/include/linux/mm.h
> > +++ b/include/linux/mm.h
> > @@ -723,6 +723,8 @@ enum {
> > */
> > static inline bool fault_flag_allow_retry_first(enum fault_flag flags)
> > {
> > + if (flags & FAULT_FLAG_VMA_LOCK)
> > + return false;
> > return (flags & FAULT_FLAG_ALLOW_RETRY) &&
> > (!(flags & FAULT_FLAG_TRIED));
> > }
> >
> > OK, this is a hack. The function is spectacularly badly named, and
> > needs to be renamed before a patch can go upstream. But this should
> > fix the contention on mmap_lock.
>
> Thanks for your suggestion.
> This is exactly what we did in Android Common Kernel before we had
> Lorenzo's proposal (bypassing `fault_flag_allow_retry_first()`):
>
> https://android.googlesource.com/kernel/common/+/1b9b045a586245cc1c29b2747c6586234c7f5bad%5E%21/#F2
Looks like that one didn't cover __folio_lock_or_retry(), but that
doesn't invalidate your point.
> Note that Lorenzo's proposal avoids mmap_lock contention without
> introducing any new VMA lock contention. It also doesn't require a new
> flag that would break KMI. So this is clearly the preferred approach.
But it does retry multiple times in cases where we know the fault
will always fail (eg the fault is on a device-private VMA)
> > Could somebody try it? I've verified it boots and runs some userspace
> > fine, but I don't have the workload to test the contention.
>
> Both Nanzhe and Hongru tested it before and reported the fork issue.
So what I didn't realise is that fork() waits for page faults to finish.
I don't think that's necessary, so we can just stop doing that (whitespace
damaged):
diff --git a/mm/mmap.c b/mm/mmap.c
index 4bf26b0f1e6e..e79555247d3a 100644
--- a/mm/mmap.c
+++ b/mm/mmap.c
@@ -1739,9 +1739,6 @@ __latent_entropy int dup_mmap(struct mm_struct *mm, struct mm_struct *oldmm)
for_each_vma(vmi, mpnt) {
struct file *file;
- retval = vma_start_write_killable(mpnt);
- if (retval < 0)
- goto loop_out;
if (vma_test(mpnt, VMA_DONTCOPY_BIT)) {
retval = vma_iter_clear_gfp(&vmi, mpnt->vm_start,
mpnt->vm_end, GFP_KERNEL);
I think this is safe. I've booted a kernel with this change, and
everything seems to run fine. Of course I don't have any multithreaded
applications which call fork() because that's a stupid way to write an
application, so it's not really tested.
My argument for why it's safe is that a thread which takes a page
fault during fork() might have taken the page fault either before or
after fork(). The faults will definitely happen in the parent process.
They may or may not have happened in the child process, which can't
possibly care whether or not they've happened.
The only difference I can think of being observable is that the child
may observe some later faults to have occurred, while some earlier faults
to have not occurred. I have a hard time believing any application can
possibly depend on it.
More information about the linux-riscv
mailing list