[PATCH] afs: don't re-queue expired servers in afs_finished_fs_probe()

henrymei ljp1205831794 at gmail.com
Thu Sep 10 23:59:36 PDT 2026


From: Aohan Mei <henrymei at tencent.com>

afs_finished_fs_probe() unconditionally re-adds server->probe_link to
net->fs_probe_fast/slow when a probe round completes, while
afs_server_destroyer() removes a server from those queues at most
once: after afs_remove_server_from_cell() sets AFS_SERVER_FL_EXPIRED,
every later destroyer pass early-returns and can never dequeue again.

If the probe dispatcher dequeues a server and dispatches an async
FS.GetCapabilities probe while the destroyer is between rb_erase() and
its one-shot list_del_init() — a window the synchronous
FS.GiveUpAllCallbacks RPC stretches to a full RPC round trip when
AFS_SERVER_FL_MAY_HAVE_CB is set — the completion callback re-queues
the already-expired server *after* that one-shot dequeue.  Nothing
ever removes the entry again; the call release then drops the last
reference and the server is freed via call_rcu()/kfree() with
probe_link still linked.  The next afs_fs_probe_dispatcher() run reads
fast->probed_at and writes probe_link/ref of the freed object:

    BUG: KASAN: slab-use-after-free in afs_fs_probe_dispatcher+0x520/0x640
    Read of size 8 (fast->probed_at, fs/afs/fs_probe.c)
    refcount_t: addition on 0; use-after-free  (afs_get_server)

Skip the re-queue (and the probe-timer re-arm) when the server has
AFS_SERVER_FL_EXPIRED set.  The test sits inside the net->fs_lock
seqlock critical section: the destroyer sets EXPIRED before its own
seqlocked dequeue in program order, so any finished-probe critical
section serialised after the destroyer's is guaranteed to observe the
flag, while an earlier re-queue is still removed by the destroyer's
own list_del_init().  Live servers are unaffected: an expired server
is on the destruction path and must never be probed again.

Fixes: f6cbb368bcb0 ("afs: Actively poll fileservers to maintain NAT or firewall openings")
Reported-by: TencentOS Corvus AI <corvus at tencent.com>
Cc: stable at vger.kernel.org
Assisted-by: CodeBuddy:Kimi-K3
Signed-off-by: Aohan Mei <henrymei at tencent.com>
---
 fs/afs/fs_probe.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/fs/afs/fs_probe.c b/fs/afs/fs_probe.c
--- a/fs/afs/fs_probe.c
+++ b/fs/afs/fs_probe.c
@@ -80,6 +80,16 @@ static void afs_finished_fs_probe(struct afs_net *net, struct afs_server *server
 	bool responded = test_bit(AFS_ESTATE_RESPONDED, &estate->flags);
 
 	write_seqlock(&net->fs_lock);
+	if (test_bit(AFS_SERVER_FL_EXPIRED, &server->flags)) {
+		/* The server is being destroyed and afs_server_destroyer()
+		 * dequeues probe_link from the probe queues only once, so an
+		 * expired server must not be re-queued here: the entry would
+		 * outlive the server object and the next probe dispatcher
+		 * run would touch freed memory.
+		 */
+		write_sequnlock(&net->fs_lock);
+		return;
+	}
 	if (responded) {
 		list_add_tail(&server->probe_link, &net->fs_probe_slow);
 	} else {
-- 
2.43.7



More information about the linux-afs mailing list