[PATCH v6] mm: retry page faults once under the per-VMA lock

Xueyuan Chen xueyuan.chen at vivo.com
Tue Sep 15 06:05:42 PDT 2026


On Fri, Sep 11, 2026 at 10:56:13AM +0800, Hongru Zhang wrote:
>From: Hongru Zhang <zhanghongru at xiaomi.com>
>
>The per-VMA lock fault path releases the per-VMA lock when the folio
>cannot be locked. It then waits for the folio to become lockable
>without holding the VMA lock and retries the page fault. However,
>the retry always takes the mmap_lock instead of the per-VMA lock.
>This can cause serious lock contention when a writer is holding the
>mmap_lock at the same time.
>
>Add a single retry under the per-VMA lock in the architecture fault
>handler. This does not touch any page fault code in mm. The retry is
>very likely to succeed because the first page fault has already waited
>for the folio to become lockable. For example, this usually means that
>any required I/O has completed. This allows faults that can make
>progress on an immediate retry to stay on the per-VMA lock path,
>avoiding waits on the mmap_lock when it is write-contended. This
>reduces page-fault latency and mmap_lock contention. Some faults may
>retry unnecessarily, for example, those in __vmf_anon_prepare() or
>device-private fault handling, which require the mmap_lock. However,
>these cases are expected to be infrequent and only add one cheap
>per-VMA lock attempt. If the second attempt still returns
>VM_FAULT_RETRY, the fault continues through the existing mmap_lock
>path.
>

Hi Hongru,

I tested cold launch time of douyin (the China's TikTok) and Honor of
Kings on SM8975. Each number below is the average of 100 runs:

     douyin:          666.0 ms -> 563.0 ms, 103 ms faster (15.5%)
     Honor of Kings: 1169.5 ms -> 1117.0 ms, 53 ms faster (4.5%)

The data shows a clear win on real apps.

Tested-by: Xueyuan Chen <xueyuan.chen21 at gmail.com>

Thanks,
Xueyuan

>Based on the stress model from Kunwu Chan and Wang Lian in RFC v2, we
>adapted a benchmark [1] to a 20-core Intel i7-12700 desktop by reducing
>the thread count and adjusting the memcg limits. The benchmark uses
>concurrent page faults under memcg pressure with parallel munmap to
>amplify mmap_lock read-write contention.
>
>Filemap throughput (higher is better)
>
>+---------+------------+---------------------+
>| Threads |  Vanilla   |       Patched       |
>+---------+------------+---------------------+
>|   40    | 1069.34 /s | 1400.13 /s (+30.9%) |
>+---------+------------+---------------------+
>|   60    | 1038.12 /s | 1683.37 /s (+62.2%) |
>+---------+------------+---------------------+
>|   80    | 1042.62 /s | 1767.83 /s (+69.6%) |
>+---------+------------+---------------------+
>
>mmap_lock contention count (lower is better)
>
>+---------+-----------+---------+-----------+
>| Threads |  Vanilla  | Patched | Reduction |
>+---------+-----------+---------+-----------+
>|   40    | 3,187,336 |  52,086 |  -98.4%   |
>+---------+-----------+---------+-----------+
>|   60    | 4,385,154 |  65,079 |  -98.5%   |
>+---------+-----------+---------+-----------+
>|   80    | 5,337,890 |  69,708 |  -98.7%   |
>+---------+-----------+---------+-----------+
>
>These results show that retrying once under the per-VMA lock keeps more
>file-backed faults on the fast path, improving throughput and reducing
>mmap_lock contention.
>
>Using benchmark [2], we tested this on a 20-core Intel i7-12700 desktop
>with a 2GB swapfile. The benchmark uses one pressure thread under memcg
>limits to keep a 128MB non-zero anonymous mapping under swap pressure,
>12 reader threads to fault it back in, and optional mmap writer threads
>to amplify mmap_lock read-write contention. Each test ran for 60 seconds
>and reported completed reader rounds per second under swap pressure.
>
>Swap throughput (higher is better)
>
>+--------------+-------------+---------------------------+
>| mmap writers |   Vanilla   |          Patched          |
>+--------------+-------------+---------------------------+
>|      0       | 17303.09 /s | 17899.48 /s       (+3.4%) |
>+--------------+-------------+---------------------------+
>|      4       | 12596.23 /s | 16095.20 /s      (+27.8%) |
>+--------------+-------------+---------------------------+
>|      8       |     0.58 /s | 15420.57 /s (+2658619.0%) |
>+--------------+-------------+---------------------------+
>
>With increasing mmap_lock write pressure, Vanilla degrades sharply and
>drops to near zero at eight writers. Patched kernel holds up much better.
>
>Performance was evaluated on a Pixel 6 running Android 17, using Baidu
>Tieba (com.baidu.tieba) and Tencent Video (com.tencent.qqlive), both
>very popular Android apps, as the workloads. For each workload, we
>performed 100 cold app launches with each kernel variant (vanilla and
>patched). Each run recorded cold app launch time, measured as the time
>to first frame, and the main thread's mmap_lock wait events. Shorter
>launch times indicate better performance.
>
>Baidu Tieba cold app launch time
>
>+-----------+----------+----------+--------+
>| Statistic |  Vanilla |  Patched | Change |
>+-----------+----------+----------+--------+
>|    Mean   | 3,672 ms | 3,580 ms |  -2.5% |
>+-----------+----------+----------+--------+
>|  Maximum  | 4,469 ms | 4,156 ms |  -7.0% |
>+-----------+----------+----------+--------+
>
>Baidu Tieba cold app launch time distribution
>
>+-------------+------------+------------+
>|  Time (ms)  |   Vanilla  |   Patched  |
>+-------------+------------+------------+
>| 2,750-2,999 |  0  (0.0%) |  1  (1.0%) |
>+-------------+------------+------------+
>| 3,000-3,249 |  8  (8.0%) | 14 (14.0%) |
>+-------------+------------+------------+
>| 3,250-3,499 | 17 (17.0%) | 30 (30.0%) |
>+-------------+------------+------------+
>| 3,500-3,749 | 38 (38.0%) | 27 (27.0%) |
>+-------------+------------+------------+
>| 3,750-3,999 | 26 (26.0%) | 18 (18.0%) |
>+-------------+------------+------------+
>| 4,000-4,249 |  8  (8.0%) | 10 (10.0%) |
>+-------------+------------+------------+
>| 4,250-4,499 |  3  (3.0%) |  0  (0.0%) |
>+-------------+------------+------------+
>
>Baidu Tieba main-thread mmap_lock wait statistics
>
>+----------------------+----------------+----------------+--------+
>|        Metric        |     Vanilla    |     Patched    | Change |
>+----------------------+----------------+----------------+--------+
>|    Total Wait Time   |  569.9 ms/run  |  469.9 ms/run  | -17.6% |
>+----------------------+----------------+----------------+--------+
>| Read-Lock Wait Count | 32.4 waits/run | 10.9 waits/run | -66.2% |
>+----------------------+----------------+----------------+--------+
>
>Tencent Video cold app launch time
>
>+-----------+----------+----------+--------+
>| Statistic |  Vanilla |  Patched | Change |
>+-----------+----------+----------+--------+
>|    Mean   | 1,907 ms | 1,840 ms |  -3.5% |
>+-----------+----------+----------+--------+
>|  Maximum  | 3,023 ms | 2,851 ms |  -5.7% |
>+-----------+----------+----------+--------+
>
>Tencent Video cold app launch time distribution
>
>+-------------+------------+------------+
>|  Time (ms)  |   Vanilla  |   Patched  |
>+-------------+------------+------------+
>| 1,250-1,499 |  3  (3.0%) |  7  (7.0%) |
>+-------------+------------+------------+
>| 1,500-1,749 | 32 (32.0%) | 33 (33.0%) |
>+-------------+------------+------------+
>| 1,750-1,999 | 39 (39.0%) | 36 (36.0%) |
>+-------------+------------+------------+
>| 2,000-2,249 | 12 (12.0%) | 15 (15.0%) |
>+-------------+------------+------------+
>| 2,250-2,499 |  7  (7.0%) |  4  (4.0%) |
>+-------------+------------+------------+
>| 2,500-2,749 |  5  (5.0%) |  3  (3.0%) |
>+-------------+------------+------------+
>| 2,750-2,999 |  1  (1.0%) |  2  (2.0%) |
>+-------------+------------+------------+
>| 3,000-3,249 |  1  (1.0%) |  0  (0.0%) |
>+-------------+------------+------------+
>
>Tencent Video main-thread mmap_lock wait statistics
>
>+----------------------+----------------+---------------+--------+
>|        Metric        |     Vanilla    |    Patched    | Change |
>+----------------------+----------------+---------------+--------+
>|    Total Wait Time   |  139.6 ms/run  |  66.4 ms/run  | -52.4% |
>+----------------------+----------------+---------------+--------+
>| Read-Lock Wait Count | 28.4 waits/run | 4.7 waits/run | -83.5% |
>+----------------------+----------------+---------------+--------+
>
>Across both workloads, the single retry under the per-VMA lock
>substantially reduced mmap_lock read-side contention, leading to lower
>app startup times at both the mean and the tail.
>
>[1] https://gist.github.com/zhr250/c36c2c54d9351df37e12fd072d4926ef
>[2] https://gist.github.com/zhr250/218ffe693f842346b56434483127422c
>
>Signed-off-by: Hongru Zhang <zhanghongru at xiaomi.com>
>Suggested-by: Barry Song <baohua at kernel.org>
>Suggested-by: Suren Baghdasaryan <surenb at google.com>
>Suggested-by: Lorenzo Stoakes (ARM) <ljs at kernel.org>
>Tested-by: Nanzhe Zhao <zhaonanzhe at xiaomi.com>
>---
>Changes since RFC v5:
>- Added measurements from 100 cold app launches per kernel for each of
>  Baidu Tieba and Tencent Video on a Pixel 6 running Android 17
>  Tested by Nanzhe Zhao. Thanks!
>- Rebased onto mm-unstable; no code changes
>
>Changes since RFC v4:
>- Drop `VM_FAULT_MAY_USE_VMA_LOCK` and always retry once under the
>  per-VMA lock, based on feedback from Lorenzo and Barry. Thanks!
>
>Changes since RFC v3:
>- Keep VM_FAULT_RETRY unchanged and add VM_FAULT_MAY_USE_VMA_LOCK as an
>  advisory bit
>- Bound VMA-lock retries with FAULT_FLAG_TRIED
>- Opt in filemap_fault() and do_swap_page() to VM_FAULT_MAY_USE_VMA_LOCK
>- Rebased on mm-unstable
>
>Changes since RFC v2:
>- Redesigned as a single blacklist-based patch (v2 was 5 per-path patches)
>- Added retry_vma loop to all architectures (not just x86)
>- Rebased on mm-unstable
>
>Changes since RFC v1:
>- collect tags from Pedro, Kunwu and Lian, thanks!
>- handle case (2), for uptodate folios, don't retry PF
>
>Link to RFC v5:
>https://lore.kernel.org/all/20260814085300.399107-1-zhanghongru@xiaomi.com/
>
>Link to RFC v4:
>https://lore.kernel.org/all/20260804095135.45897-1-zhanghongru@xiaomi.com/
>
>Link to RFC v3:
>https://lore.kernel.org/all/20260626075019.1833065-1-zhanghongru@xiaomi.com/
>
>Link to RFC v2:
>https://lore.kernel.org/all/20260430040427.4672-1-baohua@kernel.org/
>
>Link to RFC v1:
>https://lore.kernel.org/all/20251127011438.6918-1-21cnbao@gmail.com/
>
> arch/arm/mm/fault.c       | 8 ++++++++
> arch/arm64/mm/fault.c     | 8 ++++++++
> arch/loongarch/mm/fault.c | 8 ++++++++
> arch/powerpc/mm/fault.c   | 7 +++++++
> arch/riscv/mm/fault.c     | 8 ++++++++
> arch/s390/mm/fault.c      | 6 ++++++
> arch/x86/mm/fault.c       | 8 ++++++++
> 7 files changed, 53 insertions(+)
>
>diff --git a/arch/arm/mm/fault.c b/arch/arm/mm/fault.c
>index 0a09d4ff7718..70472744f5c5 100644
>--- a/arch/arm/mm/fault.c
>+++ b/arch/arm/mm/fault.c
>@@ -344,6 +344,7 @@ do_page_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
> 	vm_fault_t fault;
> 	unsigned int flags = FAULT_FLAG_DEFAULT;
> 	vm_flags_t vm_flags = VM_ACCESS_FLAGS;
>+	bool vma_lock_retried = false;
> 
> 	if (kprobe_page_fault(regs, fsr))
> 		return 0;
>@@ -395,6 +396,7 @@ do_page_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
> 	if (!(flags & FAULT_FLAG_USER))
> 		goto lock_mmap;
> 
>+lock_vma:
> 	vma = lock_vma_under_rcu(mm, addr);
> 	if (!vma)
> 		goto lock_mmap;
>@@ -424,6 +426,12 @@ do_page_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
> 			goto no_context;
> 		return 0;
> 	}
>+
>+	if (!vma_lock_retried) {
>+		vma_lock_retried = true;
>+		goto lock_vma;
>+	}
>+
> lock_mmap:
> 
> retry:
>diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c
>index 75c3e463df2e..c7fd6f485b16 100644
>--- a/arch/arm64/mm/fault.c
>+++ b/arch/arm64/mm/fault.c
>@@ -614,6 +614,7 @@ static int __kprobes do_page_fault(unsigned long far, unsigned long esr,
> 	struct vm_area_struct *vma;
> 	int si_code;
> 	int pkey = -1;
>+	bool vma_lock_retried = false;
> 
> 	if (kprobe_page_fault(regs, esr))
> 		return 0;
>@@ -682,6 +683,7 @@ static int __kprobes do_page_fault(unsigned long far, unsigned long esr,
> 	if (!(mm_flags & FAULT_FLAG_USER))
> 		goto lock_mmap;
> 
>+lock_vma:
> 	vma = lock_vma_under_rcu(mm, addr);
> 	if (!vma)
> 		goto lock_mmap;
>@@ -728,6 +730,12 @@ static int __kprobes do_page_fault(unsigned long far, unsigned long esr,
> 			goto no_context;
> 		return 0;
> 	}
>+
>+	if (!vma_lock_retried) {
>+		vma_lock_retried = true;
>+		goto lock_vma;
>+	}
>+
> lock_mmap:
> 
> retry:
>diff --git a/arch/loongarch/mm/fault.c b/arch/loongarch/mm/fault.c
>index 2c93d33356e5..ef6ea847b1e0 100644
>--- a/arch/loongarch/mm/fault.c
>+++ b/arch/loongarch/mm/fault.c
>@@ -181,6 +181,7 @@ static void __kprobes __do_page_fault(struct pt_regs *regs,
> 	struct mm_struct *mm = tsk->mm;
> 	struct vm_area_struct *vma = NULL;
> 	vm_fault_t fault;
>+	bool vma_lock_retried = false;
> 
> 	if (kprobe_page_fault(regs, current->thread.trap_nr))
> 		return;
>@@ -219,6 +220,7 @@ static void __kprobes __do_page_fault(struct pt_regs *regs,
> 	if (!(flags & FAULT_FLAG_USER))
> 		goto lock_mmap;
> 
>+lock_vma:
> 	vma = lock_vma_under_rcu(mm, address);
> 	if (!vma)
> 		goto lock_mmap;
>@@ -265,6 +267,12 @@ static void __kprobes __do_page_fault(struct pt_regs *regs,
> 			no_context(regs, write, address);
> 		return;
> 	}
>+
>+	if (!vma_lock_retried) {
>+		vma_lock_retried = true;
>+		goto lock_vma;
>+	}
>+
> lock_mmap:
> 
> retry:
>diff --git a/arch/powerpc/mm/fault.c b/arch/powerpc/mm/fault.c
>index 806c74e0d5ab..06018b6d7086 100644
>--- a/arch/powerpc/mm/fault.c
>+++ b/arch/powerpc/mm/fault.c
>@@ -422,6 +422,7 @@ static int ___do_page_fault(struct pt_regs *regs, unsigned long address,
> 	int is_write = page_fault_is_write(error_code);
> 	vm_fault_t fault, major = 0;
> 	bool kprobe_fault = kprobe_page_fault(regs, 11);
>+	bool vma_lock_retried = false;
> 
> 	if (unlikely(debugger_fault_handler(regs) || kprobe_fault))
> 		return 0;
>@@ -487,6 +488,7 @@ static int ___do_page_fault(struct pt_regs *regs, unsigned long address,
> 	if (!(flags & FAULT_FLAG_USER))
> 		goto lock_mmap;
> 
>+lock_vma:
> 	vma = lock_vma_under_rcu(mm, address);
> 	if (!vma)
> 		goto lock_mmap;
>@@ -517,6 +519,11 @@ static int ___do_page_fault(struct pt_regs *regs, unsigned long address,
> 	if (fault_signal_pending(fault, regs))
> 		return user_mode(regs) ? 0 : SIGBUS;
> 
>+	if (!vma_lock_retried) {
>+		vma_lock_retried = true;
>+		goto lock_vma;
>+	}
>+
> lock_mmap:
> 
> 	/* When running in the kernel we expect faults to occur only to
>diff --git a/arch/riscv/mm/fault.c b/arch/riscv/mm/fault.c
>index 04ed6f8acae4..ff861793dba9 100644
>--- a/arch/riscv/mm/fault.c
>+++ b/arch/riscv/mm/fault.c
>@@ -284,6 +284,7 @@ void handle_page_fault(struct pt_regs *regs)
> 	unsigned int flags = FAULT_FLAG_DEFAULT;
> 	int code = SEGV_MAPERR;
> 	vm_fault_t fault;
>+	bool vma_lock_retried = false;
> 
> 	cause = regs->cause;
> 	addr = regs->badaddr;
>@@ -347,6 +348,7 @@ void handle_page_fault(struct pt_regs *regs)
> 	if (!(flags & FAULT_FLAG_USER))
> 		goto lock_mmap;
> 
>+lock_vma:
> 	vma = lock_vma_under_rcu(mm, addr);
> 	if (!vma)
> 		goto lock_mmap;
>@@ -376,6 +378,12 @@ void handle_page_fault(struct pt_regs *regs)
> 			no_context(regs, addr);
> 		return;
> 	}
>+
>+	if (!vma_lock_retried) {
>+		vma_lock_retried = true;
>+		goto lock_vma;
>+	}
>+
> lock_mmap:
> 
> retry:
>diff --git a/arch/s390/mm/fault.c b/arch/s390/mm/fault.c
>index 46d828926009..dcd1ba24497f 100644
>--- a/arch/s390/mm/fault.c
>+++ b/arch/s390/mm/fault.c
>@@ -271,6 +271,7 @@ static void do_exception(struct pt_regs *regs, int access)
> 	unsigned int flags;
> 	vm_fault_t fault;
> 	bool is_write;
>+	bool vma_lock_retried = false;
> 
> 	/*
> 	 * The instruction that caused the program check has
>@@ -294,6 +295,7 @@ static void do_exception(struct pt_regs *regs, int access)
> 		flags |= FAULT_FLAG_WRITE;
> 	if (!(flags & FAULT_FLAG_USER))
> 		goto lock_mmap;
>+lock_vma:
> 	vma = lock_vma_under_rcu(mm, address);
> 	if (!vma)
> 		goto lock_mmap;
>@@ -318,6 +320,10 @@ static void do_exception(struct pt_regs *regs, int access)
> 			handle_fault_error_nolock(regs, 0);
> 		return;
> 	}
>+	if (!vma_lock_retried) {
>+		vma_lock_retried = true;
>+		goto lock_vma;
>+	}
> lock_mmap:
> retry:
> 	vma = lock_mm_and_find_vma(mm, address, regs);
>diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
>index aa88370ce739..df10d5cea4ee 100644
>--- a/arch/x86/mm/fault.c
>+++ b/arch/x86/mm/fault.c
>@@ -1222,6 +1222,7 @@ void do_user_addr_fault(struct pt_regs *regs,
> 	struct mm_struct *mm;
> 	vm_fault_t fault;
> 	unsigned int flags = FAULT_FLAG_DEFAULT;
>+	bool vma_lock_retried = false;
> 
> 	tsk = current;
> 	mm = tsk->mm;
>@@ -1331,6 +1332,7 @@ void do_user_addr_fault(struct pt_regs *regs,
> 	if (!(flags & FAULT_FLAG_USER))
> 		goto lock_mmap;
> 
>+lock_vma:
> 	vma = lock_vma_under_rcu(mm, address);
> 	if (!vma)
> 		goto lock_mmap;
>@@ -1360,6 +1362,12 @@ void do_user_addr_fault(struct pt_regs *regs,
> 						 ARCH_DEFAULT_PKEY);
> 		return;
> 	}
>+
>+	if (!vma_lock_retried) {
>+		vma_lock_retried = true;
>+		goto lock_vma;
>+	}
>+
> lock_mmap:
> 
> retry:
>
>base-commit: 3628c3df6cd2797b34714d23113cd44cb30801e7
>-- 
>2.43.0
>
>




More information about the linux-riscv mailing list