[PATCH] jffs2: fix wbuf_sem leak on summary failure in jffs2_flash_writev()

Shirong Zhao shxzhaosr at 163.com
Tue Sep 29 04:05:38 PDT 2026


When summary is enabled and jffs2_sum_add_kvec() fails (an allocation
failure while recording a summary entry), jffs2_flash_writev() returns
the error directly while still holding c->wbuf_sem write-locked.  No
later path releases it, so every subsequent read, write and GC blocks
on down_read()/down_write() and the filesystem becomes completely
unresponsive until remount or reboot.

Fix this the same way the wbuf recovery path already handles a lost
summary: disable summary collecting for the eraseblock with
jffs2_sum_disable_collecting() and let the write complete on the
normal path below, which releases the semaphore.  The summary is only
a mount-time optimization; the node data itself has already reached
the flash or wbuf, and a missing summary merely forces a full scan of
that eraseblock at the next mount.

An early version of this fix returned the error after unlocking (via
the shared outerr path).  That was rejected by runtime testing: with
retlen reset to 0 the allocator makes no progress, the retried write
lands on the same offset and trips the contiguity BUG() in
jffs2_link_node_ref().  Swallowing the summary error instead matches
the established recovery behavior and was verified clean.

Test comparison, same guest and test program in both runs (QEMU,
upstream 7.3.0-rc4, JFFS2 on nandsim, failslab fault injection
window-limited to jffs2_sum_add_kvec, fault count taken from the
failslab times counter):

- Deterministic 20+1+20 protocol: 20 clean 15-byte writes, 1 write
  with injection armed at probability 100, 20 clean writes, then
  umount/remount and readback of all 41 markers.
  - With this patch: exactly 1 fault fires during the armed write
    (times 1000000 -> 999999); all three writers exit 0 with no
    errors; file size 615 bytes and all 41 markers present both
    before and after remount; no BUG, oops, hung task or hang.
  - Pristine (unpatched) kernel: the same fault surfaces -ENOMEM
    ("Write of 83 bytes ... failed. returned -12"), then writer:79
    and kworker delayed_wbuf_sync both wedge in D state in
    jffs2_flash_writev -> down_write (hung task reports at 122s and
    245s, rw-semaphore "likely owned by task writer:79" itself);
    the D-state writer is unkillable and the test never completes.
- Endurance variant on the patched kernel (2 clean + 12 armed 256KiB
  block writes at probability 15 + 2 clean): 6 faults absorbed, all
  16 block markers present after remount, no anomalies.
- The original hang was first reproduced on 6.13 and re-confirmed on
  7.3.0-rc4 before the fix (writer + kworker D-state stacks in
  jffs2_flash_writev, same self-owned semaphore signature).

Fixes: e631ddba5887 ("[JFFS2] Add erase block summary support (mount time improvement)")
Cc: stable at vger.kernel.org
Signed-off-by: Shirong Zhao <shxzhaosr at 163.com>
Tested-by: Shirong Zhao <shxzhaosr at 163.com>
---
diff --git a/fs/jffs2/wbuf.c b/fs/jffs2/wbuf.c
index 3b7803c75..7ed3fa2bf 100644
--- a/fs/jffs2/wbuf.c
+++ b/fs/jffs2/wbuf.c
@@ -904,8 +904,16 @@ int jffs2_flash_writev(struct jffs2_sb_info *c, const struct kvec *invecs,

 	if (jffs2_sum_active()) {
 		int res = jffs2_sum_add_kvec(c, invecs, count, (uint32_t) to);
-		if (res)
-			return res;
+		if (res) {
+			/*
+			 * Summary recording failed (allocation failure).
+			 * The node data already reached the flash or wbuf;
+			 * disable summary for this eraseblock so the caller
+			 * still sees a successful write.  The semaphore is
+			 * released on the normal path below.
+			 */
+			jffs2_sum_disable_collecting(c->summary);
+		}
 	}

 	if (c->wbuf_len && ino)




More information about the linux-mtd mailing list