Yue Sun reported a KASAN use-after-free in tcp_sync_mss() after an AF_SMC socket entered TCP fallback and listen failed. smc_listen() installs an address-family operations table embedded in smc_sock, but its error path leaves the table installed. TCP can outlive the SMC socket and dereference the freed table. Retrying listen also saves the SMC wrapper as the original operations and can recurse indefinitely. Restore the original operations after a failed listen and before releasing the internal TCP socket. The latter also covers a fallback listener reused as an active TCP socket. Serialize installation and restoration with the TCP socket lock, and restore only if the embedded table is still installed, preserving a concurrent address-family conversion. Enable the existing RCU-delayed SMC socket destruction before publishing the listener callbacks, including on failed listen. Compare a child's operations with the embedded table itself so that concurrent restoration of the parent's operations cannot leave the child holding that table. SMC fallback exposes the kernel TCP socket through the socket file. Reject attaching the MPTCP ULP to such a socket: otherwise userspace can bypass smc_setsockopt() and MPTCP can retain the embedded table in its subflow context. MPTCP's own creation path still attaches its ULP before the new kernel socket is associated with a file. A standalone userspace reproducer forces listen to fail with EADDRINUSE, enters TCP fallback, and closes the SMC owner with data queued behind a zero receive window. A TCP probe then reports a KASAN use-after-free in __tcp_transmit_skb() when reading net_header_len. The same binary completes without a KASAN report after this change. The complete x86_64 SMC and MPTCP code was compiled. Fixes: 8270d9c21041 ("net/smc: Limit backlog connections") Fixes: 2303f994b3e1 ("mptcp: Associate MPTCP context with TCP socket") Reported-by: Yue Sun Closes: https://lore.kernel.org/netdev/20260713085238.16780-1-samsun1006219@gmail.com/ Cc: stable@vger.kernel.org Assisted-by: Codex:gpt-6-astra Signed-off-by: Kyle Zeng --- net/mptcp/subflow.c | 7 ++++--- net/smc/af_smc.c | 14 +++++++++++--- net/smc/smc_close.c | 6 ++++++ 3 files changed, 21 insertions(+), 6 deletions(-) diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c index f0a6725d2c37..da7ac71344a9 100644 --- a/net/mptcp/subflow.c +++ b/net/mptcp/subflow.c @@ -1984,10 +1984,11 @@ static int subflow_ulp_init(struct sock *sk) struct tcp_sock *tp = tcp_sk(sk); int err = 0; - /* disallow attaching ULP to a socket unless it has been - * created with sock_create_kern() + /* Only attach to a kernel-created socket that has not been + * exposed through a file. */ - if (!sk->sk_kern_sock) { + if (!sk->sk_kern_sock || + (sk->sk_socket && READ_ONCE(sk->sk_socket->file))) { err = -EOPNOTSUPP; goto out; } diff --git a/net/smc/af_smc.c b/net/smc/af_smc.c index e9f93b3ab435..dbbe7d6574e4 100644 --- a/net/smc/af_smc.c +++ b/net/smc/af_smc.c @@ -157,7 +157,7 @@ static struct sock *smc_tcp_syn_recv_sock(const struct sock *sk, rcu_assign_sk_user_data(child, NULL); /* v4-mapped sockets don't inherit parent ops. Don't restore. */ - if (inet_csk(child)->icsk_af_ops == inet_csk(sk)->icsk_af_ops) + if (inet_csk(child)->icsk_af_ops == &smc->af_ops) inet_csk(child)->icsk_af_ops = smc->ori_af_ops; } sock_put(&smc->sk); @@ -2671,6 +2671,8 @@ int smc_listen(struct socket *sock, int backlog) if (!smc->use_fallback) tcp_sk(smc->clcsock->sk)->syn_smc = 1; + sock_set_flag(sk, SOCK_RCU_FREE); + /* save original sk_data_ready function and establish * smc-specific sk_data_ready function */ @@ -2682,18 +2684,25 @@ int smc_listen(struct socket *sock, int backlog) write_unlock_bh(&smc->clcsock->sk->sk_callback_lock); /* save original ops */ + lock_sock(smc->clcsock->sk); smc->ori_af_ops = inet_csk(smc->clcsock->sk)->icsk_af_ops; smc->af_ops = *smc->ori_af_ops; smc->af_ops.syn_recv_sock = smc_tcp_syn_recv_sock; - inet_csk(smc->clcsock->sk)->icsk_af_ops = &smc->af_ops; + WRITE_ONCE(inet_csk(smc->clcsock->sk)->icsk_af_ops, &smc->af_ops); + release_sock(smc->clcsock->sk); if (smc->limit_smc_hs) tcp_sk(smc->clcsock->sk)->smc_hs_congested = smc_hs_congested; rc = kernel_listen(smc->clcsock, backlog); if (rc) { + lock_sock(smc->clcsock->sk); + if (inet_csk(smc->clcsock->sk)->icsk_af_ops == &smc->af_ops) + WRITE_ONCE(inet_csk(smc->clcsock->sk)->icsk_af_ops, + smc->ori_af_ops); + release_sock(smc->clcsock->sk); write_lock_bh(&smc->clcsock->sk->sk_callback_lock); smc_clcsock_restore_cb(&smc->clcsock->sk->sk_data_ready, &smc->clcsk_data_ready); @@ -2701,7 +2710,6 @@ int smc_listen(struct socket *sock, int backlog) write_unlock_bh(&smc->clcsock->sk->sk_callback_lock); goto out; } - sock_set_flag(sk, SOCK_RCU_FREE); sk->sk_max_ack_backlog = backlog; sk->sk_ack_backlog = 0; sk->sk_state = SMC_LISTEN; diff --git a/net/smc/smc_close.c b/net/smc/smc_close.c index bb0313ef5f7c..c59e578f3e52 100644 --- a/net/smc/smc_close.c +++ b/net/smc/smc_close.c @@ -24,12 +24,18 @@ void smc_clcsock_release(struct smc_sock *smc) { struct socket *tcp; + struct sock *sk; if (smc->listen_smc && current_work() != &smc->smc_listen_work) cancel_work_sync(&smc->smc_listen_work); mutex_lock(&smc->clcsock_release_lock); if (smc->clcsock) { tcp = smc->clcsock; + sk = tcp->sk; + lock_sock(sk); + if (inet_csk(sk)->icsk_af_ops == &smc->af_ops) + WRITE_ONCE(inet_csk(sk)->icsk_af_ops, smc->ori_af_ops); + release_sock(sk); smc->clcsock = NULL; sock_release(tcp); } -- 2.55.0.openai.867.ga7d5542d7eda