[PATCH] nvme-tcp: fix usage of page_frag_cache
Daniel Wagner
dwagner at suse.de
Thu Aug 6 05:46:44 PDT 2026
On Thu, Jul 16, 2026 at 04:42:19PM +0200, Daniel Wagner wrote:
> From: Dmitry Bogdanov <d.bogdanov at yadro.com>
>
> nvme uses page_frag_cache to preallocate PDU for each preallocated request
> of block device. Block devices are created in parallel threads,
> consequently page_frag_cache is used in not thread-safe manner.
> That leads to incorrect refcounting of backstore pages and premature free.
>
> That can be catched by !sendpage_ok inside network stack:
>
> WARNING: CPU: 7 PID: 467 at ../net/core/skbuff.c:6931 skb_splice_from_iter+0xfa/0x310.
> tcp_sendmsg_locked+0x782/0xce0
> tcp_sendmsg+0x27/0x40
> sock_sendmsg+0x8b/0xa0
> nvme_tcp_try_send_cmd_pdu+0x149/0x2a0
> Then random panic may occur.
>
> Fix that by serializing the usage of page_frag_cache.
>
> Fixes: 4e893ca81170 ("nvme_core: scan namespaces asynchronously")
> Signed-off-by: Dmitry Bogdanov <d.bogdanov at yadro.com>
> Signed-off-by: Daniel Wagner <wagi at kernel.org>
> ---
> If the target exposes many namespaces (>1000) and the host has many CPUs (>80),
> it is trivial to trigger the allocation race condition in nvme_tcp_init_request
> which results in the logs below:
Any chance to move forward with this on?
More information about the Linux-nvme
mailing list