[PATCH] afs: Don't flush dirty data when zapping an invalidated vnode

Yuanfu Xie yuanfuxie at stu.pku.edu.cn
Sat Sep 26 10:55:48 PDT 2026


afs_validate() holds vnode->validate_lock for write while it
revalidates the vnode.  If it determines that the vnode's data must be
zapped, it calls afs_zap_data(), which for a regular file calls
filemap_invalidate_inode() with flush=true.  That submits WB_SYNC_ALL
writeback, which re-enters afs_writepages() - and afs_writepages()
takes validate_lock for read before it will write anything back.  The
task thus deadlocks against itself in an uninterruptible sleep and
hangs forever; with the hung-task watchdog set to panic, the kernel
panics:

  Kernel panic - not syncing: hung_task: blocked tasks
  task:repro state:D pid:1
  Call Trace:
   <TASK>
   rwsem_down_read_slowpath+0x4d0/0xde0
   down_read+0xb6/0x220
   afs_writepages+0xb8/0xd0
   do_writepages+0x22c/0x550
   filemap_writeback+0x1dc/0x260
   afs_validate+0xd1d/0xfc0
   afs_fsync+0xcb/0x1d0
   do_fsync+0xac/0x200
   __x64_sys_fsync+0x32/0x50
   entry_SYSCALL_64_after_hwframe+0x77/0x7f
   </TASK>

Trigger: fsync() on a dirty file whose server-side data version
changed (callback break, v_break change or callback expiry) - a normal
concurrency scenario, no malicious server required.  Everything stuck
behind WB_SYNC_ALL (sync(), syncfs, umount, shutdown flush) hangs with
it, and the task cannot be killed.

The read acquisition in afs_writepages() cannot be dropped: it orders
writeback against afs_setattr() truncating the pagecache under the
write lock.  So the flush has to go: invalidate the pages without
flushing, which is what the directory and symlink branches of
afs_zap_data() have always done.  Any queued writes are based on a
stale data version at this point (that is why the zap was judged) and
cannot be stored coherently anyway.

The workload that hangs the unpatched kernel was run on a build of
7.3-rc3-135-g4982d3552a3bf with only this change applied: the fsync()
that used to hang now completes cleanly, and several hundred ordinary
file operations on the same build also completed without a crash.

Fixes: d73065e60dcc8 ("afs: Use alternative invalidation to using launder_folio")
Cc: stable at vger.kernel.org
Signed-off-by: Yuanfu Xie <yuanfuxie at stu.pku.edu.cn>
---
 fs/afs/validation.c | 14 +++++++++-----
 1 file changed, 9 insertions(+), 5 deletions(-)

diff --git a/fs/afs/validation.c b/fs/afs/validation.c
index e997563af658b..20d654bbf7e8b 100644
--- a/fs/afs/validation.c
+++ b/fs/afs/validation.c
@@ -372,11 +372,15 @@ static void afs_zap_data(struct afs_vnode *vnode)
 
 	/* nuke all the non-dirty pages that aren't locked, mapped or being
 	 * written back in a regular file and completely discard the pages in a
-	 * directory or symlink */
-	if (S_ISREG(vnode->netfs.inode.i_mode))
-		filemap_invalidate_inode(&vnode->netfs.inode, true, 0, LLONG_MAX);
-	else
-		filemap_invalidate_inode(&vnode->netfs.inode, false, 0, LLONG_MAX);
+	 * directory or symlink.
+	 *
+	 * Don't flush here: we hold validate_lock for write, and writeback
+	 * would take validate_lock for read again in afs_writepages(),
+	 * deadlocking against ourselves.  The queued writes are also based
+	 * on a stale data version at this point and cannot be stored
+	 * coherently anyway.
+	 */
+	filemap_invalidate_inode(&vnode->netfs.inode, false, 0, LLONG_MAX);
 }
 
 /*
-- 
2.43.0


More information about the linux-afs mailing list