From: Geliang Tang When a BPF stream_verdict program redirects an skb back to the same socket (self-redirect with BPF_F_INGRESS), sk_psock_verdict_apply() calls tcp_eat_skb() which advances tcp_sk->copied_seq. However, the skb is then delivered to the socket's psock ingress queue and later read by tcp_bpf_recvmsg_parser(), which also advances copied_seq via the copied_from_self accounting path. This double-counting causes copied_seq to advance by 2x the actual data length, triggering: TCP recvmsg seq # bug 2: copied BF2E806, seq BF2E7FD, \ rcvnxt BF2E806, fl 0 WARNING: net/ipv4/tcp.c:2745 at tcp_recvmsg_locked+0x72b/0x2640 Call Trace: tcp_recvmsg+0x10a/0x500 sock_recvmsg+0x168/0x1d0 __sys_recvfrom+0x19a/0x2a0 __x64_sys_recvfrom+0xe4/0x1f0 do_syscall_64+0xf7/0x530 entry_SYSCALL_64_after_hwframe+0x77/0x7f cleanup rbuf bug: copied BF2E806 seq BF2E806 rcvnxt BF2E806 WARNING: net/ipv4/tcp.c:1609 at tcp_cleanup_rbuf+0xf2/0x1c0 Call Trace: tcp_recvmsg_locked+0x8d1/0x2640 tcp_recvmsg+0x10a/0x500 sock_recvmsg+0x168/0x1d0 __sys_recvfrom+0x19a/0x2a0 __x64_sys_recvfrom+0xe4/0x1f0 do_syscall_64+0xf7/0x530 entry_SYSCALL_64_after_hwframe+0x77/0x7f Fix this by converting self-redirect verdict to __SK_PASS at the beginning of sk_psock_verdict_apply(). This bypasses the __SK_REDIRECT case entirely (which calls sk_psock_eat_skb), letting the __SK_PASS path queue the skb to the psock ingress queue. The data is then read via tcp_bpf_recvmsg_parser(), which advances copied_seq exactly once through copied_from_self. Cross-socket redirects continue through __SK_REDIRECT with sk_psock_eat_skb() unchanged. Fixes: e5c6de5fa025 ("bpf, sockmap: Incorrectly handling copied_seq") Suggested-by: Jakub Sitnicki Suggested-by: Jiayuan Chen Signed-off-by: Geliang Tang --- v2: - fixup the verdict as Jakub and Jiayuan suggested. v1: - https://patchwork.kernel.org/project/netdevbpf/patch/b840c35fdfdf36e9fddedfa645b12699bc51aa34.1787968065.git.tanggeliang@kylinos.cn/ --- net/core/skmsg.c | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/net/core/skmsg.c b/net/core/skmsg.c index 2521b643fa05..df385a5a961e 100644 --- a/net/core/skmsg.c +++ b/net/core/skmsg.c @@ -1000,6 +1000,10 @@ static int sk_psock_verdict_apply(struct sk_psock *psock, struct sk_buff *skb, int err = 0; u32 len, off; + if (verdict == __SK_REDIRECT && skb_bpf_ingress(skb) && + skb_bpf_redirect_fetch(skb) == psock->sk) + verdict = __SK_PASS; + switch (verdict) { case __SK_PASS: err = -EIO; -- 2.53.0