| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/09/22 19:35 | flow | patch-triage |
9mArgs:null Results: {
"EnableConfigs": [
"CGROUP_BPF",
"BPF_SYSCALL",
"BPF_JIT",
"DEBUG_INFO_BTF",
"INET"
],
"FocusSymbols": [
"__tcp_set_rcvlowat",
"tcp_set_rcvlowat",
"bpf_tcp_ops_set_rcvlowat",
"bpf_tcp_ops_cb_flags_set",
"bpf_sock_ops_cb_flags_set",
"bpf_tcp_ops_rcvlowat_kfunc_filter",
"bpf_sol_tcp_setsockopt",
"tcp_ofo_queue",
"tcp_fastopen_add_skb"
],
"KMSANReasoning": "The patch introduces TCP receive queue callback operations (enqueue_rcvq and dequeue_rcvq) in bpf_tcp_ops and adds a kfunc (bpf_tcp_ops_set_rcvlowat) along with flags to manage SO_RCVLOWAT dynamically from BPF.\n\n1. No uninitialized memory or info-leak risks: The patch does not introduce any allocations without initialization, nor does it expose kernel structs or buffers to userspace that could contain uninitialized padding or fields.\n2. Logic and bounds: All variables (such as wakeup flags, offsets, and return values) are fully initialized before use. The modifications to `__tcp_set_rcvlowat` adjust socket wake-up conditions and epoll/recvmsg wake thresholds, which do not permit reading uninitialized buffer memory (packet data consumption remains bounded by actual skb lengths).\n3. Sanitizer applicability: Any potential issues introduced by this patch (e.g., locking/context constraints in BPF callbacks, reference handling, or concurrency) fall within the domain of KASAN, LOCKDEP, and standard kernel assertions, not KMSAN.\n\nTherefore, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false,
"Reasoning": "The patch introduces support for enqueue_rcvq and dequeue_rcvq callbacks in bpf_tcp_ops along with a new flag BPF_SOCK_OPS_RCVQ_CB_FLAG and the bpf_tcp_ops_set_rcvlowat kfunc. It also refactors tcp_set_rcvlowat into __tcp_set_rcvlowat to conditionally suppress wakeups and hooks BPF operations into core TCP receive processing paths (tcp_queue_rcv, tcp_ofo_queue, tcp_fastopen_add_skb, tcp_cleanup_rbuf). These changes modify reachable core networking and BPF code paths that can be tested in standard virtualized environments.",
"WorthFuzzing": true
} |
| 1/1 | 2026/09/22 19:35 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit 00dab524a4f917b8005f7552ada3b92672cdfda5\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Tue Sep 22 19:35:27 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/include/net/tcp.h b/include/net/tcp.h\nindex d61ee00052e33..d3cf655da9ecd 100644\n--- a/include/net/tcp.h\n+++ b/include/net/tcp.h\n@@ -512,6 +512,7 @@ void tcp_set_keepalive(struct sock *sk, int val);\n void tcp_syn_ack_timeout(const struct request_sock *req);\n int tcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,\n \t\tint flags);\n+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup);\n int tcp_set_rcvlowat(struct sock *sk, int val);\n void tcp_set_rcvbuf(struct sock *sk, int val);\n int tcp_set_window_clamp(struct sock *sk, int val);\n@@ -2931,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,\n \treturn tcp_call_bpf(sk, op, 3, args);\n }\n \n+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)\n+{\n+\tif (sk_is_mptcp(sk) \u0026\u0026\n+\t (val \u0026 BPF_SOCK_OPS_RCVQ_CB_FLAG))\n+\t\treturn -EOPNOTSUPP;\n+\n+\ttcp_sk(sk)-\u003ebpf_sock_ops_cb_flags = val;\n+\treturn 0;\n+}\n+\n static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)\n {\n \ttcp_sk(sk)-\u003ebpf_sock_ops_cb_flags = 0;\n@@ -2953,6 +2964,11 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,\n \treturn -EPERM;\n }\n \n+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)\n+{\n+\treturn -EOPNOTSUPP;\n+}\n+\n static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)\n {\n }\n@@ -3053,6 +3069,12 @@ struct bpf_tcp_ops {\n \t\t\t struct request_sock *req, struct sk_buff *syn_skb,\n \t\t\t enum tcp_synack_type synack_type,\n \t\t\t u32 opt_off);\n+\n+\t/* Called when an incoming skb is enqueued to sk-\u003esk_receive_queue. */\n+\tvoid (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);\n+\n+\t/* Called after data is dequeued from sk-\u003esk_receive_queue. */\n+\tvoid (*dequeue_rcvq)(struct sock *sk);\n };\n \n #define bpf_tcp_ops_call(op, sk, ...)\t\t\t\t\t\\\n@@ -3144,6 +3166,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)\n \tbpf_tcp_ops_call(rtt, sk, mrtt, srtt);\n }\n \n+static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)\n+{\n+\tif (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))\n+\t\tbpf_tcp_ops_call(enqueue_rcvq, sk, skb);\n+}\n+\n+static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)\n+{\n+\tif (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))\n+\t\tbpf_tcp_ops_call(dequeue_rcvq, sk);\n+}\n+\n #if IS_ENABLED(CONFIG_SMC)\n extern struct static_key_false tcp_have_smc;\n #endif\ndiff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h\nindex 6330b7d745c57..8cdf22667775f 100644\n--- a/include/uapi/linux/bpf.h\n+++ b/include/uapi/linux/bpf.h\n@@ -3264,7 +3264,7 @@ union bpf_attr {\n * \tReturn\n * \t\t0\n *\n- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)\n+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)\n * \tDescription\n * \t\tAttempt to set the value of the **bpf_sock_ops_cb_flags** field\n * \t\tfor the full TCP socket associated to *bpf_sock_ops* to\n@@ -7148,8 +7148,17 @@ enum {\n \t * options first before the BPF program does.\n \t */\n \tBPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1\u003c\u003c6),\n+\t/* Call bpf when the TCP stack enqueues/dequeues payload\n+\t * to/from sk-\u003esk_receive_queue.\n+\t *\n+\t * Only bpf_tcp_ops is supported.\n+\t *\n+\t * It can be used to adjust sk-\u003esk_rcvlowat and suppress\n+\t * unnecessary wakeups before sufficient data is available.\n+\t */\n+\tBPF_SOCK_OPS_RCVQ_CB_FLAG = (1\u003c\u003c7),\n /* Mask of all currently supported cb flags */\n-\tBPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,\n+\tBPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,\n };\n \n enum {\ndiff --git a/net/core/filter.c b/net/core/filter.c\nindex 5feb998846825..f29c061bb0664 100644\n--- a/net/core/filter.c\n+++ b/net/core/filter.c\n@@ -5588,8 +5588,7 @@ static int bpf_sol_tcp_setsockopt(struct sock *sk, int optname,\n \tcase TCP_BPF_SOCK_OPS_CB_FLAGS:\n \t\tif (val \u0026 ~(BPF_SOCK_OPS_ALL_CB_FLAGS))\n \t\t\treturn -EINVAL;\n-\t\ttp-\u003ebpf_sock_ops_cb_flags = val;\n-\t\tbreak;\n+\t\treturn tcp_set_sock_ops_cb_flags(sk, val);\n \tdefault:\n \t\treturn -EINVAL;\n \t}\n@@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {\n BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,\n \t int, argval)\n {\n-\tstruct sock *sk = bpf_sock-\u003esk;\n \tint val = argval \u0026 BPF_SOCK_OPS_ALL_CB_FLAGS;\n+\tstruct sock *sk = bpf_sock-\u003esk;\n+\tint err;\n \n \tif (!is_locked_tcp_sock_ops(bpf_sock))\n \t\treturn -EOPNOTSUPP;\n@@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,\n \tif (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))\n \t\treturn -EINVAL;\n \n-\ttcp_sk(sk)-\u003ebpf_sock_ops_cb_flags = val;\n+\terr = tcp_set_sock_ops_cb_flags(sk, val);\n+\tif (err)\n+\t\treturn err;\n \n \treturn argval \u0026 (~BPF_SOCK_OPS_ALL_CB_FLAGS);\n }\ndiff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c\nindex 681fed6429999..3768b1440eb77 100644\n--- a/net/ipv4/bpf_tcp_ops.c\n+++ b/net/ipv4/bpf_tcp_ops.c\n@@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,\n {\n }\n \n+static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)\n+{\n+}\n+\n+static void dequeue_rcvq_stub(struct sock *sk)\n+{\n+}\n+\n static struct bpf_tcp_ops __bpf_tcp_ops = {\n \t.timeout_init = timeout_init_stub,\n \t.rwnd_init = rwnd_init_stub,\n@@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {\n \t.parse_hdr = parse_hdr_stub,\n \t.hdr_opt_len = hdr_opt_len_stub,\n \t.write_hdr_opt = write_hdr_opt_stub,\n+\t.enqueue_rcvq = enqueue_rcvq_stub,\n+\t.dequeue_rcvq = dequeue_rcvq_stub,\n };\n \n BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,\n@@ -210,6 +220,27 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {\n \t.ret_type\t= RET_INTEGER,\n };\n \n+BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)\n+{\n+\tint val = argval \u0026 BPF_SOCK_OPS_ALL_CB_FLAGS;\n+\tint err;\n+\n+\terr = tcp_set_sock_ops_cb_flags(sk, val);\n+\tif (err)\n+\t\treturn err;\n+\n+\treturn argval \u0026 ~BPF_SOCK_OPS_ALL_CB_FLAGS;\n+}\n+\n+static const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {\n+\t.func\t\t= bpf_tcp_ops_cb_flags_set,\n+\t.gpl_only\t= false,\n+\t.ret_type\t= RET_INTEGER,\n+\t.arg1_type\t= ARG_PTR_TO_BTF_ID,\n+\t.arg1_btf_id\t= \u0026btf_sock_ids[BTF_SOCK_TYPE_TCP],\n+\t.arg2_type\t= ARG_ANYTHING,\n+};\n+\n static const struct bpf_func_proto *\n get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)\n {\n@@ -255,6 +286,15 @@ get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)\n \t\tif (moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))\n \t\t\treturn \u0026bpf_tcp_ops_store_hdr_opt_proto;\n \t\treturn NULL;\n+\tcase BPF_FUNC_sock_ops_cb_flags_set:\n+\t\tif (moff == offsetof(struct bpf_tcp_ops, connect) ||\n+\t\t moff == offsetof(struct bpf_tcp_ops, listen) ||\n+\t\t moff == offsetof(struct bpf_tcp_ops, active_established) ||\n+\t\t moff == offsetof(struct bpf_tcp_ops, passive_established) ||\n+\t\t moff == offsetof(struct bpf_tcp_ops, enqueue_rcvq) ||\n+\t\t moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))\n+\t\t\treturn \u0026bpf_tcp_ops_cb_flags_set_proto;\n+\t\treturn NULL;\n \tdefault:\n \t\treturn bpf_base_func_proto(func_id, prog);\n \t}\n@@ -319,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {\n \t.owner = THIS_MODULE,\n };\n \n+__bpf_kfunc_start_defs();\n+\n+__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,\n+\t\t\t\t\t const struct bpf_prog_aux *aux)\n+{\n+\tu32 moff = aux-\u003eattach_st_ops_member_off;\n+\tbool wakeup = false;\n+\n+\tif (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))\n+\t\twakeup = true;\n+\n+\tif (rcvlowat \u003c 0)\n+\t\trcvlowat = INT_MAX;\n+\n+\treturn __tcp_set_rcvlowat(sk, rcvlowat, wakeup);\n+}\n+\n+__bpf_kfunc_end_defs();\n+\n+BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)\n+BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)\n+BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)\n+\n+static int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,\n+\t\t\t\t\t u32 kfunc_id)\n+{\n+\tu32 moff;\n+\n+\tif (!btf_id_set8_contains(\u0026bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))\n+\t\treturn 0;\n+\n+\tif (prog-\u003eaux-\u003est_ops != \u0026bpf_tcp_ops)\n+\t\treturn -EACCES;\n+\n+\tmoff = prog-\u003eaux-\u003eattach_st_ops_member_off;\n+\tif (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) \u0026\u0026\n+\t moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))\n+\t\treturn -EACCES;\n+\n+\treturn 0;\n+}\n+\n+static const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {\n+\t.owner = THIS_MODULE,\n+\t.set = \u0026bpf_tcp_ops_rcvlowat_kfunc_set,\n+\t.filter = bpf_tcp_ops_rcvlowat_kfunc_filter,\n+};\n+\n static int __init __bpf_tcp_ops_init(void)\n {\n-\treturn register_bpf_struct_ops(\u0026bpf_tcp_ops, bpf_tcp_ops);\n+\tint ret;\n+\n+\tret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,\n+\t\t\t\t\t\u0026bpf_tcp_ops_rcvlowat_kfunc_id_set);\n+\tret = ret ?: register_bpf_struct_ops(\u0026bpf_tcp_ops, bpf_tcp_ops);\n+\n+\treturn ret;\n }\n late_initcall(__bpf_tcp_ops_init);\ndiff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c\nindex a4456b4194122..aa7593fc8334e 100644\n--- a/net/ipv4/tcp.c\n+++ b/net/ipv4/tcp.c\n@@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)\n \t \"cleanup rbuf bug: copied %X seq %X rcvnxt %X\\n\",\n \t tp-\u003ecopied_seq, TCP_SKB_CB(skb)-\u003eend_seq, tp-\u003ercv_nxt);\n \t__tcp_cleanup_rbuf(sk, copied);\n+\n+\tbpf_tcp_ops_dequeue_rcvq(sk);\n }\n \n static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)\n@@ -1826,8 +1828,7 @@ int tcp_peek_len(struct socket *sock)\n \treturn tcp_inq(sock-\u003esk);\n }\n \n-/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */\n-int tcp_set_rcvlowat(struct sock *sk, int val)\n+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)\n {\n \tstruct tcp_sock *tp = tcp_sk(sk);\n \tint space, cap;\n@@ -1840,7 +1841,8 @@ int tcp_set_rcvlowat(struct sock *sk, int val)\n \tWRITE_ONCE(sk-\u003esk_rcvlowat, val ? : 1);\n \n \t/* Check if we need to signal EPOLLIN right now */\n-\ttcp_data_ready(sk);\n+\tif (wakeup)\n+\t\ttcp_data_ready(sk);\n \n \tif (sk-\u003esk_userlocks \u0026 SOCK_RCVBUF_LOCK)\n \t\treturn 0;\n@@ -1855,6 +1857,12 @@ int tcp_set_rcvlowat(struct sock *sk, int val)\n \treturn 0;\n }\n \n+/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */\n+int tcp_set_rcvlowat(struct sock *sk, int val)\n+{\n+\treturn __tcp_set_rcvlowat(sk, val, true);\n+}\n+\n void tcp_set_rcvbuf(struct sock *sk, int val)\n {\n \ttcp_set_window_clamp(sk, tcp_win_from_space(sk, val));\ndiff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c\nindex 471c78be55137..4939bcbc81d17 100644\n--- a/net/ipv4/tcp_fastopen.c\n+++ b/net/ipv4/tcp_fastopen.c\n@@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)\n \tTCP_SKB_CB(skb)-\u003eseq++;\n \tTCP_SKB_CB(skb)-\u003etcp_flags \u0026= ~TCPHDR_SYN;\n \n+\tbpf_tcp_ops_enqueue_rcvq(sk, skb);\n+\n \ttp-\u003ercv_nxt = TCP_SKB_CB(skb)-\u003eend_seq;\n \ttcp_add_receive_queue(sk, skb);\n \ttp-\u003esyn_data_acked = 1;\ndiff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c\nindex 6ac6f9d5b6c38..c60c61bb0a715 100644\n--- a/net/ipv4/tcp_input.c\n+++ b/net/ipv4/tcp_input.c\n@@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)\n \t\t\tcontinue;\n \t\t}\n \n+\t\tbpf_tcp_ops_enqueue_rcvq(sk, skb);\n+\n \t\ttail = skb_peek_tail(\u0026sk-\u003esk_receive_queue);\n \t\teaten = tail \u0026\u0026 tcp_try_coalesce(sk, tail, skb, \u0026fragstolen);\n \t\ttcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)-\u003eend_seq);\n@@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,\n \tint eaten;\n \tstruct sk_buff *tail = skb_peek_tail(\u0026sk-\u003esk_receive_queue);\n \n+\tbpf_tcp_ops_enqueue_rcvq(sk, skb);\n+\n \teaten = (tail \u0026\u0026\n \t\t tcp_try_coalesce(sk, tail,\n \t\t\t\t skb, fragstolen)) ? 1 : 0;\ndiff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h\nindex 6330b7d745c57..8cdf22667775f 100644\n--- a/tools/include/uapi/linux/bpf.h\n+++ b/tools/include/uapi/linux/bpf.h\n@@ -3264,7 +3264,7 @@ union bpf_attr {\n * \tReturn\n * \t\t0\n *\n- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)\n+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)\n * \tDescription\n * \t\tAttempt to set the value of the **bpf_sock_ops_cb_flags** field\n * \t\tfor the full TCP socket associated to *bpf_sock_ops* to\n@@ -7148,8 +7148,17 @@ enum {\n \t * options first before the BPF program does.\n \t */\n \tBPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1\u003c\u003c6),\n+\t/* Call bpf when the TCP stack enqueues/dequeues payload\n+\t * to/from sk-\u003esk_receive_queue.\n+\t *\n+\t * Only bpf_tcp_ops is supported.\n+\t *\n+\t * It can be used to adjust sk-\u003esk_rcvlowat and suppress\n+\t * unnecessary wakeups before sufficient data is available.\n+\t */\n+\tBPF_SOCK_OPS_RCVQ_CB_FLAG = (1\u003c\u003c7),\n /* Mask of all currently supported cb flags */\n-\tBPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,\n+\tBPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,\n };\n \n enum {\ndiff --git a/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c\nnew file mode 100644\nindex 0000000000000..c5ead4af24f85\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c\n@@ -0,0 +1,350 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Copyright 2026 Google LLC */\n+#include \u003csys/epoll.h\u003e\n+\n+#include \"test_progs.h\"\n+#include \"cgroup_helpers.h\"\n+#include \"network_helpers.h\"\n+\n+#include \"tcp_autolowat.skel.h\"\n+\n+#define SOL_BPF\t\t\t0xdeadbeef\n+#define BPF_TCP_AUTOLOWAT\t0x8badf00d\n+\n+struct rpc_descriptor {\n+\tu32 header_len;\n+\tu32 payload_len;\n+};\n+\n+enum rpc_event_type {\n+\tRPC_EVENT_END,\n+\tRPC_EVENT_AUTOLOWAT,\n+\tRPC_EVENT_SEND,\n+\tRPC_EVENT_RECV,\n+\tRPC_EVENT_EPOLL,\n+\tRPC_EVENT_RCVLOWAT,\n+};\n+\n+struct rpc_event {\n+\tenum rpc_event_type type;\n+\tunion {\n+\t\tint len;\n+\t\tint nfds;\n+\t\tint val;\n+\t\tint rcvlowat;\n+\t};\n+};\n+\n+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))\n+\n+struct rpc_test_case {\n+\tchar data[4096];\n+\tstruct rpc_descriptor desc[32];\n+\tstruct rpc_event event[32];\n+} rpc_test_cases[] = {\n+\t{\n+\t\t.desc = {\n+\t\t\t{ .header_len = 100, .payload_len = 150 },\n+\t\t},\n+\t\t.event = {\n+\t\t\t{ .type = RPC_EVENT_AUTOLOWAT,\t.val = 1},\n+\t\t\t/* Single full RPC message in skb. */\n+\t\t\t{ .type = RPC_EVENT_SEND,\t.len = RPC_DESC_SIZE + 100 + 150},\n+\t\t\t{ .type = RPC_EVENT_EPOLL,\t.nfds = 1},\n+\t\t\t{ .type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = RPC_DESC_SIZE + 100 + 150},\n+\t\t},\n+\t},\n+\t{\n+\t\t.desc = {\n+\t\t\t{.header_len = 100, .payload_len = 150},\n+\t\t\t{.header_len = 100, .payload_len = 150},\n+\t\t\t{.header_len = 100, .payload_len = 150},\n+\t\t},\n+\t\t.event = {\n+\t\t\t{ .type = RPC_EVENT_AUTOLOWAT,\t.val = 1},\n+\t\t\t/* Two full RPC messages in skb. */\n+\t\t\t{.type = RPC_EVENT_SEND,\t.len = (RPC_DESC_SIZE + 100 + 150) * 2},\n+\t\t\t{.type = RPC_EVENT_EPOLL,\t.nfds = 1},\n+\t\t\t{.type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},\n+\t\t\t/* Single full RPC message in skb. */\n+\t\t\t{ .type = RPC_EVENT_SEND,\t.len = RPC_DESC_SIZE + 100 + 150},\n+\t\t\t{ .type = RPC_EVENT_EPOLL,\t.nfds = 1},\n+\t\t\t{ .type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 3},\n+\t\t},\n+\t},\n+\t{\n+\t\t.desc = {\n+\t\t\t{.header_len = 100, .payload_len = 150},\n+\t\t\t{.header_len = 100, .payload_len = 150},\n+\t\t\t{.header_len = 100, .payload_len = 150},\n+\t\t},\n+\t\t.event = {\n+\t\t\t{ .type = RPC_EVENT_AUTOLOWAT,\t.val = 1},\n+\t\t\t/* Two full RPC messages in skb. */\n+\t\t\t{.type = RPC_EVENT_SEND,\t.len = (RPC_DESC_SIZE + 100 + 150) * 2},\n+\t\t\t{.type = RPC_EVENT_EPOLL,\t.nfds = 1},\n+\t\t\t{.type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},\n+\t\t\t/* Single full RPC message in skb. */\n+\t\t\t{ .type = RPC_EVENT_SEND,\t.len = RPC_DESC_SIZE},\n+\t\t\t{ .type = RPC_EVENT_EPOLL,\t.nfds = 1},\n+\t\t\t{ .type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},\n+\t\t},\n+\t},\n+\t{\n+\t\t.desc = {\n+\t\t\t{.header_len = 100, .payload_len = 150},\n+\t\t\t{.header_len = 200, .payload_len = 500},\n+\t\t},\n+\t\t.event = {\n+\t\t\t{ .type = RPC_EVENT_AUTOLOWAT,\t.val = 1},\n+\t\t\t/* The first descriptor is partial. */\n+\t\t\t{.type = RPC_EVENT_SEND,\t.len = 1},\n+\t\t\t{.type = RPC_EVENT_EPOLL,\t.nfds = 0},\n+\t\t\t{.type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = RPC_DESC_SIZE},\n+\t\t\t/* The first descriptor is available. */\n+\t\t\t{.type = RPC_EVENT_SEND,\t.len = RPC_DESC_SIZE - 1},\n+\t\t\t{.type = RPC_EVENT_EPOLL,\t.nfds = 0},\n+\t\t\t{.type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = RPC_DESC_SIZE + 150 + 100},\n+\t\t\t/* The first header is ready. */\n+\t\t\t{.type = RPC_EVENT_SEND,\t.len = 100},\n+\t\t\t{.type = RPC_EVENT_EPOLL,\t.nfds = 0},\n+\t\t\t{.type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = RPC_DESC_SIZE + 150 + 100},\n+\t\t\t/* skb has the first payload and 1 byte of the next descriptor. */\n+\t\t\t{.type = RPC_EVENT_SEND,\t.len = 150 + 1},\n+\t\t\t{.type = RPC_EVENT_EPOLL,\t.nfds = 1},\n+\t\t\t{.type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = RPC_DESC_SIZE + 150 + 100},\n+\t\t\t/* After reading the first RPC message, SO_RCVLOWAT should be RPC_DESC_SIZE. */\n+\t\t\t{.type = RPC_EVENT_RECV,\t.len = RPC_DESC_SIZE + 150 + 100},\n+\t\t\t{.type = RPC_EVENT_EPOLL,\t.nfds = 0},\n+\t\t\t{.type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = RPC_DESC_SIZE},\n+\t\t\t/* The second descriptor is available. */\n+\t\t\t{.type = RPC_EVENT_SEND,\t.len = RPC_DESC_SIZE - 1},\n+\t\t\t{.type = RPC_EVENT_EPOLL,\t.nfds = 0},\n+\t\t\t{.type = RPC_EVENT_RCVLOWAT,\t.rcvlowat = RPC_DESC_SIZE + 200 + 500},\n+\t\t},\n+\t},\n+};\n+\n+struct tcp_autolowat_test_cb {\n+\tint saved_netns;\n+\tunion {\n+\t\tint fd[4];\n+\t\tstruct {\n+\t\t\tint server, client, child;\n+\t\t\tint epoll;\n+\t\t};\n+\t};\n+};\n+\n+static void tcp_autolowat_teardown_cb(struct tcp_autolowat_test_cb *cb)\n+{\n+\tint i, err;\n+\n+\tfor (i = 0; i \u003c ARRAY_SIZE(cb-\u003efd); i++) {\n+\t\tif (cb-\u003efd[i] != -1)\n+\t\t\tclose(cb-\u003efd[i]);\n+\t}\n+\n+\tif (cb-\u003esaved_netns != -1) {\n+\t\terr = setns(cb-\u003esaved_netns, CLONE_NEWNET);\n+\t\tASSERT_OK(err, \"restore netns\");\n+\n+\t\tclose(cb-\u003esaved_netns);\n+\t}\n+}\n+\n+static int tcp_autolowat_setup_cb(struct tcp_autolowat_test_cb *cb, int family)\n+{\n+\tstruct epoll_event ev = {};\n+\tint err;\n+\tint i;\n+\n+\tfor (i = 0; i \u003c ARRAY_SIZE(cb-\u003efd); i++)\n+\t\tcb-\u003efd[i] = -1;\n+\n+\tcb-\u003esaved_netns = open(\"/proc/self/ns/net\", O_RDONLY);\n+\tif (!ASSERT_OK_FD(cb-\u003esaved_netns, \"save netns\"))\n+\t\tgoto err;\n+\n+\terr = unshare(CLONE_NEWNET);\n+\tif (!ASSERT_OK(err, \"unshare\"))\n+\t\tgoto err;\n+\n+\terr = system(\"ip link set dev lo up\");\n+\tif (!ASSERT_OK(err, \"set up lo\"))\n+\t\tgoto err;\n+\n+\tcb-\u003eserver = start_server(family, SOCK_STREAM, NULL, 0, 0);\n+\tif (!ASSERT_OK_FD(cb-\u003eserver, \"start_server\"))\n+\t\tgoto err;\n+\n+\tcb-\u003eclient = connect_to_fd(cb-\u003eserver, 0);\n+\tif (!ASSERT_OK_FD(cb-\u003eclient, \"connect_to_fd\"))\n+\t\tgoto err;\n+\n+\tcb-\u003echild = accept(cb-\u003eserver, NULL, NULL);\n+\tif (!ASSERT_OK_FD(cb-\u003echild, \"accept\"))\n+\t\tgoto err;\n+\n+\tcb-\u003eepoll = epoll_create1(0);\n+\tif (!ASSERT_OK_FD(cb-\u003eepoll, \"epoll_create\"))\n+\t\tgoto err;\n+\n+\tev.events = EPOLLIN;\n+\tev.data.fd = cb-\u003echild;\n+\n+\terr = epoll_ctl(cb-\u003eepoll, EPOLL_CTL_ADD, cb-\u003echild, \u0026ev);\n+\tif (!ASSERT_OK(err, \"epoll_ctl\"))\n+\t\tgoto err;\n+\n+\treturn 0;\n+\n+err:\n+\ttcp_autolowat_teardown_cb(cb);\n+\treturn -1;\n+}\n+\n+static int tcp_autolowat_build_data(struct rpc_test_case *test_case)\n+{\n+\tstruct rpc_descriptor *desc = test_case-\u003edesc;\n+\tchar *ptr = test_case-\u003edata;\n+\tint rpc_size;\n+\n+\tmemset(ptr, 0, sizeof(test_case-\u003edata));\n+\n+\twhile (desc-\u003eheader_len + desc-\u003epayload_len) {\n+\t\trpc_size = sizeof(*desc) + desc-\u003eheader_len + desc-\u003epayload_len;\n+\n+\t\tif (!ASSERT_LE(ptr + rpc_size - test_case-\u003edata,\n+\t\t\t sizeof(test_case-\u003edata), \"data overflow\"))\n+\t\t\treturn 1;\n+\n+\t\tmemcpy(ptr, desc, sizeof(*desc));\n+\t\tptr += rpc_size;\n+\t\tdesc++;\n+\t}\n+\n+\tif (!ASSERT_GT(ptr - test_case-\u003edata, 0, \"no data\"))\n+\t\treturn 1;\n+\n+\treturn 0;\n+}\n+\n+static void tcp_autolowat_run_rpc_test(struct tcp_autolowat_test_cb *cb,\n+\t\t\t\t struct rpc_test_case *test_case)\n+{\n+\tstruct rpc_event *event = test_case-\u003eevent;\n+\tchar *ptr = test_case-\u003edata;\n+\tstruct epoll_event ev;\n+\tsocklen_t optlen;\n+\tint err, optval;\n+\tchar buf[4096];\n+\n+\tif (tcp_autolowat_build_data(test_case))\n+\t\treturn;\n+\n+\twhile (1) {\n+\t\tswitch (event-\u003etype) {\n+\t\tcase RPC_EVENT_END:\n+\t\t\treturn;\n+\t\tcase RPC_EVENT_AUTOLOWAT:\n+\t\t\terr = setsockopt(cb-\u003echild, SOL_BPF, BPF_TCP_AUTOLOWAT,\n+\t\t\t\t\t \u0026event-\u003eval, sizeof(event-\u003eval));\n+\t\t\tif (!ASSERT_OK(err, \"setsockopt\"))\n+\t\t\t\treturn;\n+\t\t\tbreak;\n+\t\tcase RPC_EVENT_SEND:\n+\t\t\terr = send(cb-\u003eclient, ptr, event-\u003elen, 0);\n+\t\t\tif (!ASSERT_EQ(err, event-\u003elen, \"send\"))\n+\t\t\t\treturn;\n+\n+\t\t\tptr += event-\u003elen;\n+\t\t\tbreak;\n+\t\tcase RPC_EVENT_RECV:\n+\t\t\terr = recv(cb-\u003echild, buf, event-\u003elen, 0);\n+\t\t\tif (!ASSERT_EQ(err, event-\u003elen, \"recv\"))\n+\t\t\t\treturn;\n+\t\t\tbreak;\n+\t\tcase RPC_EVENT_EPOLL:\n+\t\t\terr = epoll_wait(cb-\u003eepoll, \u0026ev, 1, 100);\n+\t\t\tif (!ASSERT_EQ(err, event-\u003enfds, \"epoll_wait\"))\n+\t\t\t\treturn;\n+\t\t\tbreak;\n+\t\tcase RPC_EVENT_RCVLOWAT:\n+\t\t\toptval = 0;\n+\t\t\toptlen = sizeof(optval);\n+\n+\t\t\terr = getsockopt(cb-\u003echild, SOL_SOCKET, SO_RCVLOWAT, \u0026optval, \u0026optlen);\n+\t\t\tif (!ASSERT_OK(err, \"getsockopt\") ||\n+\t\t\t !ASSERT_EQ(optval, event-\u003ercvlowat, \"rcvlowat\"))\n+\t\t\t\treturn;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\tevent++;\n+\t}\n+}\n+\n+static void tcp_autolowat_run_rpc_tests(struct tcp_autolowat *skel, int family)\n+{\n+\tstruct tcp_autolowat_test_cb cb;\n+\tint err;\n+\tint i;\n+\n+\tfor (i = 0; i \u003c ARRAY_SIZE(rpc_test_cases); i++) {\n+\t\tmemset(skel-\u003ebss-\u003etest_name, 0, sizeof(skel-\u003ebss-\u003etest_name));\n+\n+\t\tsnprintf(skel-\u003ebss-\u003etest_name, sizeof(skel-\u003ebss-\u003etest_name),\n+\t\t\t \"AF_INET%c rpc_test_cases[%d]\",\n+\t\t\t family == AF_INET ? ' ' : '6', i);\n+\n+\t\tif (!test__start_subtest(skel-\u003ebss-\u003etest_name))\n+\t\t\tcontinue;\n+\n+\t\terr = tcp_autolowat_setup_cb(\u0026cb, family);\n+\t\tif (err)\n+\t\t\tcontinue;\n+\n+\t\ttcp_autolowat_run_rpc_test(\u0026cb, \u0026rpc_test_cases[i]);\n+\t\ttcp_autolowat_teardown_cb(\u0026cb);\n+\t}\n+}\n+\n+static void tcp_autolowat_run_tests(struct tcp_autolowat *skel)\n+{\n+\ttcp_autolowat_run_rpc_tests(skel, AF_INET);\n+\ttcp_autolowat_run_rpc_tests(skel, AF_INET6);\n+}\n+\n+void test_tcp_autolowat(void)\n+{\n+\tstruct tcp_autolowat *skel;\n+\tstruct bpf_link *link[2];\n+\tint cgroup;\n+\n+\tskel = tcp_autolowat__open_and_load();\n+\tif (!ASSERT_OK_PTR(skel, \"open_and_load\"))\n+\t\treturn;\n+\n+\tcgroup = test__join_cgroup(\"/tcp_autolowat\");\n+\tif (!ASSERT_GE(cgroup, 0, \"join_cgroup\"))\n+\t\tgoto destroy_skel;\n+\n+\tlink[0] = bpf_map__attach_cgroup_opts(skel-\u003emaps.tcp_autolowat_ops, cgroup, NULL);\n+\tif (!ASSERT_OK_PTR(link[0], \"attach_cgroup(tcp_autolowat_ops)\"))\n+\t\tgoto close_cgroup;\n+\n+\tlink[1] = bpf_program__attach_cgroup(skel-\u003eprogs.tcp_autolowat_setsockopt, cgroup);\n+\tif (!ASSERT_OK_PTR(link[1], \"attach_cgroup(SETSOCKOPT)\"))\n+\t\tgoto destroy_sockops;\n+\n+\ttcp_autolowat_run_tests(skel);\n+\n+\tbpf_link__destroy(link[1]);\n+destroy_sockops:\n+\tbpf_link__destroy(link[0]);\n+close_cgroup:\n+\tclose(cgroup);\n+destroy_skel:\n+\ttcp_autolowat__destroy(skel);\n+}\ndiff --git a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c\nindex 7e8fe1bad03f5..e4849d2a2956f 100644\n--- a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c\n+++ b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c\n@@ -26,7 +26,8 @@ static void verify_result(struct tcpbpf_globals *result)\n \tASSERT_EQ(result-\u003ebytes_acked, 1002, \"bytes_acked\");\n \tASSERT_EQ(result-\u003edata_segs_in, 1, \"data_segs_in\");\n \tASSERT_EQ(result-\u003edata_segs_out, 1, \"data_segs_out\");\n-\tASSERT_EQ(result-\u003ebad_cb_test_rv, 0x80, \"bad_cb_test_rv\");\n+\tASSERT_EQ(result-\u003ebad_cb_test_rv, BPF_SOCK_OPS_ALL_CB_FLAGS + 1,\n+\t\t \"bad_cb_test_rv\");\n \tASSERT_EQ(result-\u003egood_cb_test_rv, 0, \"good_cb_test_rv\");\n \tASSERT_EQ(result-\u003enum_listen, 1, \"num_listen\");\n \ndiff --git a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h\nindex 593b38f904174..4c999d59cbbce 100644\n--- a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h\n+++ b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h\n@@ -79,6 +79,8 @@\n \n #define NEXTHDR_TCP\t\t6\n \n+#define TCPHDR_FIN\t\t0x01\n+\n #define TCPOPT_NOP\t\t1\n #define TCPOPT_EOL\t\t0\n #define TCPOPT_MSS\t\t2\ndiff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c\nnew file mode 100644\nindex 0000000000000..bb96e19e7589d\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c\n@@ -0,0 +1,312 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Copyright 2026 Google LLC */\n+#include \"vmlinux.h\"\n+\n+#include \u003cstring.h\u003e\n+#include \u003climits.h\u003e\n+#include \u003cbpf/bpf_helpers.h\u003e\n+#include \u003cbpf/bpf_tracing.h\u003e\n+#include \u003cbpf/bpf_core_read.h\u003e\n+\n+#include \"bpf_kfuncs.h\"\n+#include \"bpf_tracing_net.h\"\n+\n+#define SOL_BPF\t\t\t0xdeadbeef\n+#define BPF_TCP_AUTOLOWAT\t0x8badf00d\n+\n+//#define DEBUG /* For verbose output. */\n+\n+struct rpc_descriptor {\n+\tu32 header_len;\n+\tu32 payload_len;\n+};\n+\n+#define RPC_DESC_SIZE\t\t(sizeof(struct rpc_descriptor))\n+#define MAX_RPC_DESC_PER_SKB\t100\n+\n+struct tcp_autolowat_cb {\n+\t/* Don't put this field at the end; BPF verifier complains. */\n+\tchar rpc_desc_buf[RPC_DESC_SIZE];\n+\tu32 rpc_desc_seq;\n+\tu32 rpc_end_seq;\n+#ifdef DEBUG\n+\tu32 isn;\n+#endif\n+\tu8 rpc_desc_buff_len;\n+};\n+\n+struct {\n+\t__uint(type, BPF_MAP_TYPE_SK_STORAGE);\n+\t__uint(map_flags, BPF_F_NO_PREALLOC);\n+\t__type(key, int);\n+\t__type(value, struct tcp_autolowat_cb);\n+} tcp_autolowat_map SEC(\".maps\");\n+\n+char test_name[64];\n+\n+#ifdef DEBUG\n+#define LOG(str, ...)\t\t\t\t\t\t\t\\\n+\tbpf_printk(\"%s: \" str, test_name, ##__VA_ARGS__)\n+#else\n+#define LOG(...)\n+#endif\n+\n+#define SEQ(val)\t\t\t\t\\\n+\t(val - cb-\u003eisn)\n+#define TP_SEQ(field)\t\t\t\t\\\n+\t(tp-\u003efield - cb-\u003eisn)\n+#define CB_SEQ(field)\t\t\t\t\\\n+\t(cb-\u003efield - cb-\u003eisn)\n+\n+static int tcp_parse_descriptor(struct tcp_autolowat_cb *cb,\n+\t\t\t\tstruct bpf_dynptr *dptr,\n+\t\t\t\tu32 seq, u32 end_seq)\n+{\n+\tstruct rpc_descriptor *rpc_desc;\n+\tu32 rpc_copied_seq;\n+\tu64 copy_len; /* u32 should work, but not for no_alu32 :/ */\n+\tu64 rpc_len;\n+\tint err;\n+\n+\trpc_copied_seq = cb-\u003erpc_desc_seq + cb-\u003erpc_desc_buff_len;\n+\n+\tif (before(cb-\u003erpc_desc_seq + RPC_DESC_SIZE, end_seq))\n+\t\tcopy_len = RPC_DESC_SIZE - cb-\u003erpc_desc_buff_len;\n+\telse\n+\t\tcopy_len = end_seq - rpc_copied_seq;\n+\n+\tif (copy_len == 0)\n+\t\tgoto disable; /* FIN. */\n+\tif (copy_len \u003e RPC_DESC_SIZE)\n+\t\tgoto disable; /* always false, only for verifier. */\n+\tif (cb-\u003erpc_desc_buf + cb-\u003erpc_desc_buff_len \u003e= \u0026cb-\u003erpc_desc_buf[RPC_DESC_SIZE])\n+\t\tgoto disable; /* always false, only for verifier. */\n+\n+\terr = bpf_dynptr_read(cb-\u003erpc_desc_buf + cb-\u003erpc_desc_buff_len,\n+\t\t\t copy_len, dptr, rpc_copied_seq - seq, 0);\n+\tif (err)\n+\t\tgoto disable;\n+\n+\tcb-\u003erpc_desc_buff_len += copy_len;\n+\n+\tif (cb-\u003erpc_desc_buff_len != RPC_DESC_SIZE) {\n+\t\tLOG(\"Copied %d bytes: rpc_desc_buff_len: %u\", copy_len, cb-\u003erpc_desc_buff_len);\n+\t\tgoto partial;\n+\t}\n+\n+\trpc_desc = (struct rpc_descriptor *)cb-\u003erpc_desc_buf;\n+\trpc_len = RPC_DESC_SIZE + rpc_desc-\u003eheader_len + rpc_desc-\u003epayload_len;\n+\n+\tif (rpc_len \u003e INT_MAX)\n+\t\tgoto disable;\n+\n+\tcb-\u003erpc_end_seq = cb-\u003erpc_desc_seq + rpc_len;\n+\n+\tLOG(\"Copied full descriptor: rpc_desc_seq: %u, rpc_end_seq: %u, header_len: %u, payload_len: %u\",\n+\t CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq),\n+\t rpc_desc-\u003eheader_len, rpc_desc-\u003epayload_len);\n+\n+\treturn 0;\n+disable:\n+\treturn -1;\n+partial:\n+\treturn 1;\n+}\n+\n+static void tcp_set_autolowat(struct tcp_autolowat_cb *cb,\n+\t\t\t struct sock *sk)\n+{\n+\tstruct tcp_sock *tp = (struct tcp_sock *)sk;\n+\tu32 val; /* To handle wraparound. */\n+\n+\tLOG(\"Setting rcvlowat: tp-\u003ecopied_seq: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_desc_buff_len: %u\",\n+\t TP_SEQ(copied_seq), CB_SEQ(rpc_desc_seq),\n+\t CB_SEQ(rpc_end_seq), cb-\u003erpc_desc_buff_len);\n+\n+\tif (before(tp-\u003ecopied_seq, cb-\u003erpc_desc_seq))\n+\t\tval = cb-\u003erpc_desc_seq - tp-\u003ecopied_seq;\n+\telse if (cb-\u003erpc_desc_buff_len != RPC_DESC_SIZE)\n+\t\tval = RPC_DESC_SIZE;\n+\telse\n+\t\tval = cb-\u003erpc_end_seq - tp-\u003ecopied_seq;\n+\n+\tif (val != tp-\u003einet_conn.icsk_inet.sk.sk_rcvlowat) {\n+\t\tbpf_tcp_ops_set_rcvlowat(sk, val);\n+\n+\t\tLOG(\"Set rcvlowat: expected: %u, actual: %d\\n\",\n+\t\t val, tp-\u003einet_conn.icsk_inet.sk.sk_rcvlowat);\n+\t} else {\n+\t\tLOG(\"No need to set rcvlowat: %u\\n\", val);\n+\t}\n+}\n+\n+static void tcp_disable_autolowat(struct sock *sk)\n+{\n+\tstruct tcp_sock *tp = (struct tcp_sock *)sk;\n+\tint flags;\n+\n+\tflags = tp-\u003ebpf_sock_ops_cb_flags \u0026 ~BPF_SOCK_OPS_RCVQ_CB_FLAG;\n+\tbpf_sock_ops_cb_flags_set(sk, flags);\n+\n+\tbpf_tcp_ops_set_rcvlowat(sk, 1);\n+\n+\tLOG(\"Disabled autolowat\");\n+}\n+\n+static void tcp_do_autolowat(struct tcp_autolowat_cb *cb,\n+\t\t\t struct sock *sk, struct sk_buff *skb)\n+{\n+\tstruct bpf_dynptr dptr;\n+\tstruct tcp_skb_cb *tcb;\n+\tu32 seq, end_seq;\n+\tint ret = 0, i;\n+\n+\tif (bpf_dynptr_from_skb((struct __sk_buff *)skb, 0, \u0026dptr)) {\n+\t\tret = -1;\n+\t\tgoto update;\n+\t}\n+\n+\ttcb = bpf_core_cast(skb-\u003ecb, struct tcp_skb_cb);\n+\tseq = tcb-\u003eseq;\n+\tend_seq = tcb-\u003eend_seq - !!(tcb-\u003etcp_flags \u0026 TCPHDR_FIN);\n+\n+\tLOG(\"Start parsing skb: seq: %u, end_seq: %u, len: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_buff_len: %u\",\n+\t SEQ(seq), SEQ(end_seq), end_seq - seq,\n+\t CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq), cb-\u003erpc_desc_buff_len);\n+\n+\tif (cb-\u003erpc_desc_buff_len != RPC_DESC_SIZE) {\n+\t\tret = tcp_parse_descriptor(cb, \u0026dptr, seq, end_seq);\n+\t\tif (ret)\n+\t\t\tgoto update;\n+\t}\n+\n+\ti = 0;\n+\n+\twhile (1) {\n+\t\tif (i++ \u003e MAX_RPC_DESC_PER_SKB) {\n+\t\t\tret = -1;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\tif (after(cb-\u003erpc_end_seq, end_seq)) {\n+\t\t\tLOG(\"No more descriptor: rpc_end_seq: %u, end_seq: %u\",\n+\t\t\t CB_SEQ(rpc_end_seq), SEQ(end_seq));\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\tcb-\u003erpc_desc_seq = cb-\u003erpc_end_seq;\n+\t\tcb-\u003erpc_desc_buff_len = 0;\n+\n+\t\tif (cb-\u003erpc_end_seq == end_seq)\n+\t\t\tbreak;\n+\n+\t\tLOG(\"Found next descriptor: rpc_end_seq: %u, end_seq: %u, len: %u\",\n+\t\t CB_SEQ(rpc_end_seq), SEQ(end_seq), end_seq - cb-\u003erpc_end_seq);\n+\n+\t\tret = tcp_parse_descriptor(cb, \u0026dptr, seq, end_seq);\n+\t\tif (ret)\n+\t\t\tbreak;\n+\t}\n+\n+update:\n+\tif (ret \u003e= 0)\n+\t\ttcp_set_autolowat(cb, sk);\n+\telse\n+\t\ttcp_disable_autolowat(sk);\n+}\n+\n+SEC(\"struct_ops\")\n+void BPF_PROG(tcp_autolowat_enqueue_rcvq, struct sock *sk, struct sk_buff *skb)\n+{\n+\tstruct tcp_autolowat_cb *cb;\n+\n+\tcb = bpf_sk_storage_get(\u0026tcp_autolowat_map, sk, 0, 0);\n+\tif (!cb)\n+\t\treturn;\n+\n+\ttcp_do_autolowat(cb, sk, skb);\n+}\n+\n+SEC(\"struct_ops\")\n+void BPF_PROG(tcp_autolowat_dequeue_rcvq, struct sock *sk)\n+{\n+\tstruct tcp_autolowat_cb *cb;\n+\n+\tcb = bpf_sk_storage_get(\u0026tcp_autolowat_map, sk, 0, 0);\n+\tif (!cb)\n+\t\treturn;\n+\n+\ttcp_set_autolowat(cb, sk);\n+}\n+\n+SEC(\".struct_ops.link\")\n+struct bpf_tcp_ops tcp_autolowat_ops = {\n+\t.enqueue_rcvq = (void *)tcp_autolowat_enqueue_rcvq,\n+\t.dequeue_rcvq = (void *)tcp_autolowat_dequeue_rcvq,\n+};\n+\n+static int tcp_init_autolowat_cb(struct bpf_sockopt *sockopt,\n+\t\t\t\t struct bpf_tcp_sock *btp)\n+{\n+\tstruct tcp_autolowat_cb *cb;\n+\tstruct tcp_sock *tp;\n+\tint flags;\n+\n+\tcb = bpf_sk_storage_get(\u0026tcp_autolowat_map, btp, 0,\n+\t\t\t\tBPF_SK_STORAGE_GET_F_CREATE);\n+\tif (!cb)\n+\t\treturn -1;\n+\n+\ttp = bpf_core_cast(btp, struct tcp_sock);\n+\tif (!tp)\n+\t\treturn -1;\n+\n+\tcb-\u003erpc_desc_seq = tp-\u003ecopied_seq;\n+\tcb-\u003erpc_end_seq = tp-\u003ecopied_seq;\n+#ifdef DEBUG\n+\tcb-\u003eisn = tp-\u003ecopied_seq;\n+#endif\n+\n+\tif (bpf_getsockopt(sockopt-\u003esk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,\n+\t\t\t \u0026flags, sizeof(flags)))\n+\t\treturn -1;\n+\n+\tflags |= BPF_SOCK_OPS_RCVQ_CB_FLAG;\n+\n+\tif (bpf_setsockopt(sockopt-\u003esk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,\n+\t\t\t \u0026flags, sizeof(flags)))\n+\t\treturn -1;\n+\n+\treturn 0;\n+}\n+\n+SEC(\"cgroup/setsockopt\")\n+int tcp_autolowat_setsockopt(struct bpf_sockopt *ctx)\n+{\n+\tvoid *optval_end = ctx-\u003eoptval_end;\n+\tint *optval = ctx-\u003eoptval;\n+\tstruct bpf_tcp_sock *btp;\n+\n+\tif (ctx-\u003elevel != SOL_BPF || ctx-\u003eoptname != BPF_TCP_AUTOLOWAT)\n+\t\tgoto out;\n+\n+\tif (optval + 1 \u003e optval_end)\n+\t\treturn 0; /* -EPERM */\n+\n+\tbtp = bpf_tcp_sock(ctx-\u003esk);\n+\tif (!btp)\n+\t\tgoto out;\n+\n+\tif (*optval \u0026\u0026 tcp_init_autolowat_cb(ctx, btp))\n+\t\treturn 0; /* -EPERM */\n+\n+\t/*\n+\t * BPF has consumed this option, don't call kernel\n+\t * setsockopt handler.\n+\t */\n+\tctx-\u003eoptlen = -1;\n+out:\n+\treturn 1;\n+}\n+\n+char _license[] SEC(\"license\") = \"GPL\";\ndiff --git a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c\nindex 6935f32eeb8fd..e30cb1fab0790 100644\n--- a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c\n+++ b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c\n@@ -92,7 +92,8 @@ int bpf_testcb(struct bpf_sock_ops *skops)\n \t\tbreak;\n \tcase BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB:\n \t\t/* Test failure to set largest cb flag (assumes not defined) */\n-\t\tglobal.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops, 0x80);\n+\t\tglobal.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,\n+\t\t\t\t\t\t\t\t BPF_SOCK_OPS_ALL_CB_FLAGS + 1);\n \t\t/* Set callback */\n \t\tglobal.good_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,\n \t\t\t\t\t\t BPF_SOCK_OPS_STATE_CB_FLAG);\n"
} |
| 2/1 | 2026/09/22 19:35 | agent | patch-evaluator |
4mModel:core Args: null Results: {
"EnableConfigs": [
"CGROUP_BPF",
"BPF_SYSCALL",
"BPF_JIT",
"DEBUG_INFO_BTF",
"INET"
],
"FocusSymbols": [
"__tcp_set_rcvlowat",
"tcp_set_rcvlowat",
"bpf_tcp_ops_set_rcvlowat",
"bpf_tcp_ops_cb_flags_set",
"bpf_sock_ops_cb_flags_set",
"bpf_tcp_ops_rcvlowat_kfunc_filter",
"bpf_sol_tcp_setsockopt",
"tcp_ofo_queue",
"tcp_fastopen_add_skb"
],
"Reasoning": "The patch introduces support for enqueue_rcvq and dequeue_rcvq callbacks in bpf_tcp_ops along with a new flag BPF_SOCK_OPS_RCVQ_CB_FLAG and the bpf_tcp_ops_set_rcvlowat kfunc. It also refactors tcp_set_rcvlowat into __tcp_set_rcvlowat to conditionally suppress wakeups and hooks BPF operations into core TCP receive processing paths (tcp_queue_rcv, tcp_ofo_queue, tcp_fastopen_add_skb, tcp_cleanup_rbuf). These changes modify reachable core networking and BPF code paths that can be tested in standard virtualized environments.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 00dab524a4f917b8005f7552ada3b92672cdfda5
Author: syz-cluster <triage@syzkaller.com>
Date: Tue Sep 22 19:35:27 2026 +0000
syz-cluster: applied patch under review
diff --git a/include/net/tcp.h b/include/net/tcp.h
index d61ee00052e33..d3cf655da9ecd 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -512,6 +512,7 @@ void tcp_set_keepalive(struct sock *sk, int val);
void tcp_syn_ack_timeout(const struct request_sock *req);
int tcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,
int flags);
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup);
int tcp_set_rcvlowat(struct sock *sk, int val);
void tcp_set_rcvbuf(struct sock *sk, int val);
int tcp_set_window_clamp(struct sock *sk, int val);
@@ -2931,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
return tcp_call_bpf(sk, op, 3, args);
}
+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
+{
+ if (sk_is_mptcp(sk) &&
+ (val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ return -EOPNOTSUPP;
+
+ tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+ return 0;
+}
+
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
tcp_sk(sk)->bpf_sock_ops_cb_flags = 0;
@@ -2953,6 +2964,11 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
return -EPERM;
}
+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
+{
+ return -EOPNOTSUPP;
+}
+
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
}
@@ -3053,6 +3069,12 @@ struct bpf_tcp_ops {
struct request_sock *req, struct sk_buff *syn_skb,
enum tcp_synack_type synack_type,
u32 opt_off);
+
+ /* Called when an incoming skb is enqueued to sk->sk_receive_queue. */
+ void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
+
+ /* Called after data is dequeued from sk->sk_receive_queue. */
+ void (*dequeue_rcvq)(struct sock *sk);
};
#define bpf_tcp_ops_call(op, sk, ...) \
@@ -3144,6 +3166,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
}
+static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)
+{
+ if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ bpf_tcp_ops_call(enqueue_rcvq, sk, skb);
+}
+
+static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)
+{
+ if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ bpf_tcp_ops_call(dequeue_rcvq, sk);
+}
+
#if IS_ENABLED(CONFIG_SMC)
extern struct static_key_false tcp_have_smc;
#endif
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 6330b7d745c57..8cdf22667775f 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -3264,7 +3264,7 @@ union bpf_attr {
* Return
* 0
*
- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
* Description
* Attempt to set the value of the **bpf_sock_ops_cb_flags** field
* for the full TCP socket associated to *bpf_sock_ops* to
@@ -7148,8 +7148,17 @@ enum {
* options first before the BPF program does.
*/
BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
+ /* Call bpf when the TCP stack enqueues/dequeues payload
+ * to/from sk->sk_receive_queue.
+ *
+ * Only bpf_tcp_ops is supported.
+ *
+ * It can be used to adjust sk->sk_rcvlowat and suppress
+ * unnecessary wakeups before sufficient data is available.
+ */
+ BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
/* Mask of all currently supported cb flags */
- BPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,
+ BPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,
};
enum {
diff --git a/net/core/filter.c b/net/core/filter.c
index 5feb998846825..f29c061bb0664 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -5588,8 +5588,7 @@ static int bpf_sol_tcp_setsockopt(struct sock *sk, int optname,
case TCP_BPF_SOCK_OPS_CB_FLAGS:
if (val & ~(BPF_SOCK_OPS_ALL_CB_FLAGS))
return -EINVAL;
- tp->bpf_sock_ops_cb_flags = val;
- break;
+ return tcp_set_sock_ops_cb_flags(sk, val);
default:
return -EINVAL;
}
@@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {
BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
int, argval)
{
- struct sock *sk = bpf_sock->sk;
int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+ struct sock *sk = bpf_sock->sk;
+ int err;
if (!is_locked_tcp_sock_ops(bpf_sock))
return -EOPNOTSUPP;
@@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
if (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))
return -EINVAL;
- tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+ err = tcp_set_sock_ops_cb_flags(sk, val);
+ if (err)
+ return err;
return argval & (~BPF_SOCK_OPS_ALL_CB_FLAGS);
}
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 681fed6429999..3768b1440eb77 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,
{
}
+static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)
+{
+}
+
+static void dequeue_rcvq_stub(struct sock *sk)
+{
+}
+
static struct bpf_tcp_ops __bpf_tcp_ops = {
.timeout_init = timeout_init_stub,
.rwnd_init = rwnd_init_stub,
@@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {
.parse_hdr = parse_hdr_stub,
.hdr_opt_len = hdr_opt_len_stub,
.write_hdr_opt = write_hdr_opt_stub,
+ .enqueue_rcvq = enqueue_rcvq_stub,
+ .dequeue_rcvq = dequeue_rcvq_stub,
};
BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,
@@ -210,6 +220,27 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
.ret_type = RET_INTEGER,
};
+BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
+{
+ int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+ int err;
+
+ err = tcp_set_sock_ops_cb_flags(sk, val);
+ if (err)
+ return err;
+
+ return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
+}
+
+static const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {
+ .func = bpf_tcp_ops_cb_flags_set,
+ .gpl_only = false,
+ .ret_type = RET_INTEGER,
+ .arg1_type = ARG_PTR_TO_BTF_ID,
+ .arg1_btf_id = &btf_sock_ids[BTF_SOCK_TYPE_TCP],
+ .arg2_type = ARG_ANYTHING,
+};
+
static const struct bpf_func_proto *
get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
@@ -255,6 +286,15 @@ get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
if (moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
return &bpf_tcp_ops_store_hdr_opt_proto;
return NULL;
+ case BPF_FUNC_sock_ops_cb_flags_set:
+ if (moff == offsetof(struct bpf_tcp_ops, connect) ||
+ moff == offsetof(struct bpf_tcp_ops, listen) ||
+ moff == offsetof(struct bpf_tcp_ops, active_established) ||
+ moff == offsetof(struct bpf_tcp_ops, passive_established) ||
+ moff == offsetof(struct bpf_tcp_ops, enqueue_rcvq) ||
+ moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ return &bpf_tcp_ops_cb_flags_set_proto;
+ return NULL;
default:
return bpf_base_func_proto(func_id, prog);
}
@@ -319,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
.owner = THIS_MODULE,
};
+__bpf_kfunc_start_defs();
+
+__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
+ const struct bpf_prog_aux *aux)
+{
+ u32 moff = aux->attach_st_ops_member_off;
+ bool wakeup = false;
+
+ if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ wakeup = true;
+
+ if (rcvlowat < 0)
+ rcvlowat = INT_MAX;
+
+ return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
+}
+
+__bpf_kfunc_end_defs();
+
+BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)
+BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)
+BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)
+
+static int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,
+ u32 kfunc_id)
+{
+ u32 moff;
+
+ if (!btf_id_set8_contains(&bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))
+ return 0;
+
+ if (prog->aux->st_ops != &bpf_tcp_ops)
+ return -EACCES;
+
+ moff = prog->aux->attach_st_ops_member_off;
+ if (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) &&
+ moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ return -EACCES;
+
+ return 0;
+}
+
+static const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {
+ .owner = THIS_MODULE,
+ .set = &bpf_tcp_ops_rcvlowat_kfunc_set,
+ .filter = bpf_tcp_ops_rcvlowat_kfunc_filter,
+};
+
static int __init __bpf_tcp_ops_init(void)
{
- return register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+ int ret;
+
+ ret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,
+ &bpf_tcp_ops_rcvlowat_kfunc_id_set);
+ ret = ret ?: register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+
+ return ret;
}
late_initcall(__bpf_tcp_ops_init);
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index a4456b4194122..aa7593fc8334e 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
"cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
__tcp_cleanup_rbuf(sk, copied);
+
+ bpf_tcp_ops_dequeue_rcvq(sk);
}
static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)
@@ -1826,8 +1828,7 @@ int tcp_peek_len(struct socket *sock)
return tcp_inq(sock->sk);
}
-/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
-int tcp_set_rcvlowat(struct sock *sk, int val)
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)
{
struct tcp_sock *tp = tcp_sk(sk);
int space, cap;
@@ -1840,7 +1841,8 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
WRITE_ONCE(sk->sk_rcvlowat, val ? : 1);
/* Check if we need to signal EPOLLIN right now */
- tcp_data_ready(sk);
+ if (wakeup)
+ tcp_data_ready(sk);
if (sk->sk_userlocks & SOCK_RCVBUF_LOCK)
return 0;
@@ -1855,6 +1857,12 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
return 0;
}
+/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
+int tcp_set_rcvlowat(struct sock *sk, int val)
+{
+ return __tcp_set_rcvlowat(sk, val, true);
+}
+
void tcp_set_rcvbuf(struct sock *sk, int val)
{
tcp_set_window_clamp(sk, tcp_win_from_space(sk, val));
diff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c
index 471c78be55137..4939bcbc81d17 100644
--- a/net/ipv4/tcp_fastopen.c
+++ b/net/ipv4/tcp_fastopen.c
@@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)
TCP_SKB_CB(skb)->seq++;
TCP_SKB_CB(skb)->tcp_flags &= ~TCPHDR_SYN;
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
tp->rcv_nxt = TCP_SKB_CB(skb)->end_seq;
tcp_add_receive_queue(sk, skb);
tp->syn_data_acked = 1;
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 6ac6f9d5b6c38..c60c61bb0a715 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)
continue;
}
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
tail = skb_peek_tail(&sk->sk_receive_queue);
eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
@@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
int eaten;
struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
eaten = (tail &&
tcp_try_coalesce(sk, tail,
skb, fragstolen)) ? 1 : 0;
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 6330b7d745c57..8cdf22667775f 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -3264,7 +3264,7 @@ union bpf_attr {
* Return
* 0
*
- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
* Description
* Attempt to set the value of the **bpf_sock_ops_cb_flags** field
* for the full TCP socket associated to *bpf_sock_ops* to
@@ -7148,8 +7148,17 @@ enum {
* options first before the BPF program does.
*/
BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
+ /* Call bpf when the TCP stack enqueues/dequeues payload
+ * to/from sk->sk_receive_queue.
+ *
+ * Only bpf_tcp_ops is supported.
+ *
+ * It can be used to adjust sk->sk_rcvlowat and suppress
+ * unnecessary wakeups before sufficient data is available.
+ */
+ BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
/* Mask of all currently supported cb flags */
- BPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,
+ BPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,
};
enum {
diff --git a/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
new file mode 100644
index 0000000000000..c5ead4af24f85
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
@@ -0,0 +1,350 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include <sys/epoll.h>
+
+#include "test_progs.h"
+#include "cgroup_helpers.h"
+#include "network_helpers.h"
+
+#include "tcp_autolowat.skel.h"
+
+#define SOL_BPF 0xdeadbeef
+#define BPF_TCP_AUTOLOWAT 0x8badf00d
+
+struct rpc_descriptor {
+ u32 header_len;
+ u32 payload_len;
+};
+
+enum rpc_event_type {
+ RPC_EVENT_END,
+ RPC_EVENT_AUTOLOWAT,
+ RPC_EVENT_SEND,
+ RPC_EVENT_RECV,
+ RPC_EVENT_EPOLL,
+ RPC_EVENT_RCVLOWAT,
+};
+
+struct rpc_event {
+ enum rpc_event_type type;
+ union {
+ int len;
+ int nfds;
+ int val;
+ int rcvlowat;
+ };
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+
+struct rpc_test_case {
+ char data[4096];
+ struct rpc_descriptor desc[32];
+ struct rpc_event event[32];
+} rpc_test_cases[] = {
+ {
+ .desc = {
+ { .header_len = 100, .payload_len = 150 },
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 100 + 150},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Two full RPC messages in skb. */
+ {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 3},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Two full RPC messages in skb. */
+ {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 200, .payload_len = 500},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* The first descriptor is partial. */
+ {.type = RPC_EVENT_SEND, .len = 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE},
+ /* The first descriptor is available. */
+ {.type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE - 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* The first header is ready. */
+ {.type = RPC_EVENT_SEND, .len = 100},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* skb has the first payload and 1 byte of the next descriptor. */
+ {.type = RPC_EVENT_SEND, .len = 150 + 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* After reading the first RPC message, SO_RCVLOWAT should be RPC_DESC_SIZE. */
+ {.type = RPC_EVENT_RECV, .len = RPC_DESC_SIZE + 150 + 100},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE},
+ /* The second descriptor is available. */
+ {.type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE - 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 200 + 500},
+ },
+ },
+};
+
+struct tcp_autolowat_test_cb {
+ int saved_netns;
+ union {
+ int fd[4];
+ struct {
+ int server, client, child;
+ int epoll;
+ };
+ };
+};
+
+static void tcp_autolowat_teardown_cb(struct tcp_autolowat_test_cb *cb)
+{
+ int i, err;
+
+ for (i = 0; i < ARRAY_SIZE(cb->fd); i++) {
+ if (cb->fd[i] != -1)
+ close(cb->fd[i]);
+ }
+
+ if (cb->saved_netns != -1) {
+ err = setns(cb->saved_netns, CLONE_NEWNET);
+ ASSERT_OK(err, "restore netns");
+
+ close(cb->saved_netns);
+ }
+}
+
+static int tcp_autolowat_setup_cb(struct tcp_autolowat_test_cb *cb, int family)
+{
+ struct epoll_event ev = {};
+ int err;
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(cb->fd); i++)
+ cb->fd[i] = -1;
+
+ cb->saved_netns = open("/proc/self/ns/net", O_RDONLY);
+ if (!ASSERT_OK_FD(cb->saved_netns, "save netns"))
+ goto err;
+
+ err = unshare(CLONE_NEWNET);
+ if (!ASSERT_OK(err, "unshare"))
+ goto err;
+
+ err = system("ip link set dev lo up");
+ if (!ASSERT_OK(err, "set up lo"))
+ goto err;
+
+ cb->server = start_server(family, SOCK_STREAM, NULL, 0, 0);
+ if (!ASSERT_OK_FD(cb->server, "start_server"))
+ goto err;
+
+ cb->client = connect_to_fd(cb->server, 0);
+ if (!ASSERT_OK_FD(cb->client, "connect_to_fd"))
+ goto err;
+
+ cb->child = accept(cb->server, NULL, NULL);
+ if (!ASSERT_OK_FD(cb->child, "accept"))
+ goto err;
+
+ cb->epoll = epoll_create1(0);
+ if (!ASSERT_OK_FD(cb->epoll, "epoll_create"))
+ goto err;
+
+ ev.events = EPOLLIN;
+ ev.data.fd = cb->child;
+
+ err = epoll_ctl(cb->epoll, EPOLL_CTL_ADD, cb->child, &ev);
+ if (!ASSERT_OK(err, "epoll_ctl"))
+ goto err;
+
+ return 0;
+
+err:
+ tcp_autolowat_teardown_cb(cb);
+ return -1;
+}
+
+static int tcp_autolowat_build_data(struct rpc_test_case *test_case)
+{
+ struct rpc_descriptor *desc = test_case->desc;
+ char *ptr = test_case->data;
+ int rpc_size;
+
+ memset(ptr, 0, sizeof(test_case->data));
+
+ while (desc->header_len + desc->payload_len) {
+ rpc_size = sizeof(*desc) + desc->header_len + desc->payload_len;
+
+ if (!ASSERT_LE(ptr + rpc_size - test_case->data,
+ sizeof(test_case->data), "data overflow"))
+ return 1;
+
+ memcpy(ptr, desc, sizeof(*desc));
+ ptr += rpc_size;
+ desc++;
+ }
+
+ if (!ASSERT_GT(ptr - test_case->data, 0, "no data"))
+ return 1;
+
+ return 0;
+}
+
+static void tcp_autolowat_run_rpc_test(struct tcp_autolowat_test_cb *cb,
+ struct rpc_test_case *test_case)
+{
+ struct rpc_event *event = test_case->event;
+ char *ptr = test_case->data;
+ struct epoll_event ev;
+ socklen_t optlen;
+ int err, optval;
+ char buf[4096];
+
+ if (tcp_autolowat_build_data(test_case))
+ return;
+
+ while (1) {
+ switch (event->type) {
+ case RPC_EVENT_END:
+ return;
+ case RPC_EVENT_AUTOLOWAT:
+ err = setsockopt(cb->child, SOL_BPF, BPF_TCP_AUTOLOWAT,
+ &event->val, sizeof(event->val));
+ if (!ASSERT_OK(err, "setsockopt"))
+ return;
+ break;
+ case RPC_EVENT_SEND:
+ err = send(cb->client, ptr, event->len, 0);
+ if (!ASSERT_EQ(err, event->len, "send"))
+ return;
+
+ ptr += event->len;
+ break;
+ case RPC_EVENT_RECV:
+ err = recv(cb->child, buf, event->len, 0);
+ if (!ASSERT_EQ(err, event->len, "recv"))
+ return;
+ break;
+ case RPC_EVENT_EPOLL:
+ err = epoll_wait(cb->epoll, &ev, 1, 100);
+ if (!ASSERT_EQ(err, event->nfds, "epoll_wait"))
+ return;
+ break;
+ case RPC_EVENT_RCVLOWAT:
+ optval = 0;
+ optlen = sizeof(optval);
+
+ err = getsockopt(cb->child, SOL_SOCKET, SO_RCVLOWAT, &optval, &optlen);
+ if (!ASSERT_OK(err, "getsockopt") ||
+ !ASSERT_EQ(optval, event->rcvlowat, "rcvlowat"))
+ return;
+ break;
+ }
+
+ event++;
+ }
+}
+
+static void tcp_autolowat_run_rpc_tests(struct tcp_autolowat *skel, int family)
+{
+ struct tcp_autolowat_test_cb cb;
+ int err;
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(rpc_test_cases); i++) {
+ memset(skel->bss->test_name, 0, sizeof(skel->bss->test_name));
+
+ snprintf(skel->bss->test_name, sizeof(skel->bss->test_name),
+ "AF_INET%c rpc_test_cases[%d]",
+ family == AF_INET ? ' ' : '6', i);
+
+ if (!test__start_subtest(skel->bss->test_name))
+ continue;
+
+ err = tcp_autolowat_setup_cb(&cb, family);
+ if (err)
+ continue;
+
+ tcp_autolowat_run_rpc_test(&cb, &rpc_test_cases[i]);
+ tcp_autolowat_teardown_cb(&cb);
+ }
+}
+
+static void tcp_autolowat_run_tests(struct tcp_autolowat *skel)
+{
+ tcp_autolowat_run_rpc_tests(skel, AF_INET);
+ tcp_autolowat_run_rpc_tests(skel, AF_INET6);
+}
+
+void test_tcp_autolowat(void)
+{
+ struct tcp_autolowat *skel;
+ struct bpf_link *link[2];
+ int cgroup;
+
+ skel = tcp_autolowat__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "open_and_load"))
+ return;
+
+ cgroup = test__join_cgroup("/tcp_autolowat");
+ if (!ASSERT_GE(cgroup, 0, "join_cgroup"))
+ goto destroy_skel;
+
+ link[0] = bpf_map__attach_cgroup_opts(skel->maps.tcp_autolowat_ops, cgroup, NULL);
+ if (!ASSERT_OK_PTR(link[0], "attach_cgroup(tcp_autolowat_ops)"))
+ goto close_cgroup;
+
+ link[1] = bpf_program__attach_cgroup(skel->progs.tcp_autolowat_setsockopt, cgroup);
+ if (!ASSERT_OK_PTR(link[1], "attach_cgroup(SETSOCKOPT)"))
+ goto destroy_sockops;
+
+ tcp_autolowat_run_tests(skel);
+
+ bpf_link__destroy(link[1]);
+destroy_sockops:
+ bpf_link__destroy(link[0]);
+close_cgroup:
+ close(cgroup);
+destroy_skel:
+ tcp_autolowat__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
index 7e8fe1bad03f5..e4849d2a2956f 100644
--- a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
+++ b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
@@ -26,7 +26,8 @@ static void verify_result(struct tcpbpf_globals *result)
ASSERT_EQ(result->bytes_acked, 1002, "bytes_acked");
ASSERT_EQ(result->data_segs_in, 1, "data_segs_in");
ASSERT_EQ(result->data_segs_out, 1, "data_segs_out");
- ASSERT_EQ(result->bad_cb_test_rv, 0x80, "bad_cb_test_rv");
+ ASSERT_EQ(result->bad_cb_test_rv, BPF_SOCK_OPS_ALL_CB_FLAGS + 1,
+ "bad_cb_test_rv");
ASSERT_EQ(result->good_cb_test_rv, 0, "good_cb_test_rv");
ASSERT_EQ(result->num_listen, 1, "num_listen");
diff --git a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
index 593b38f904174..4c999d59cbbce 100644
--- a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
+++ b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
@@ -79,6 +79,8 @@
#define NEXTHDR_TCP 6
+#define TCPHDR_FIN 0x01
+
#define TCPOPT_NOP 1
#define TCPOPT_EOL 0
#define TCPOPT_MSS 2
diff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
new file mode 100644
index 0000000000000..bb96e19e7589d
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
@@ -0,0 +1,312 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include "vmlinux.h"
+
+#include <string.h>
+#include <limits.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+
+#include "bpf_kfuncs.h"
+#include "bpf_tracing_net.h"
+
+#define SOL_BPF 0xdeadbeef
+#define BPF_TCP_AUTOLOWAT 0x8badf00d
+
+//#define DEBUG /* For verbose output. */
+
+struct rpc_descriptor {
+ u32 header_len;
+ u32 payload_len;
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+#define MAX_RPC_DESC_PER_SKB 100
+
+struct tcp_autolowat_cb {
+ /* Don't put this field at the end; BPF verifier complains. */
+ char rpc_desc_buf[RPC_DESC_SIZE];
+ u32 rpc_desc_seq;
+ u32 rpc_end_seq;
+#ifdef DEBUG
+ u32 isn;
+#endif
+ u8 rpc_desc_buff_len;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_SK_STORAGE);
+ __uint(map_flags, BPF_F_NO_PREALLOC);
+ __type(key, int);
+ __type(value, struct tcp_autolowat_cb);
+} tcp_autolowat_map SEC(".maps");
+
+char test_name[64];
+
+#ifdef DEBUG
+#define LOG(str, ...) \
+ bpf_printk("%s: " str, test_name, ##__VA_ARGS__)
+#else
+#define LOG(...)
+#endif
+
+#define SEQ(val) \
+ (val - cb->isn)
+#define TP_SEQ(field) \
+ (tp->field - cb->isn)
+#define CB_SEQ(field) \
+ (cb->field - cb->isn)
+
+static int tcp_parse_descriptor(struct tcp_autolowat_cb *cb,
+ struct bpf_dynptr *dptr,
+ u32 seq, u32 end_seq)
+{
+ struct rpc_descriptor *rpc_desc;
+ u32 rpc_copied_seq;
+ u64 copy_len; /* u32 should work, but not for no_alu32 :/ */
+ u64 rpc_len;
+ int err;
+
+ rpc_copied_seq = cb->rpc_desc_seq + cb->rpc_desc_buff_len;
+
+ if (before(cb->rpc_desc_seq + RPC_DESC_SIZE, end_seq))
+ copy_len = RPC_DESC_SIZE - cb->rpc_desc_buff_len;
+ else
+ copy_len = end_seq - rpc_copied_seq;
+
+ if (copy_len == 0)
+ goto disable; /* FIN. */
+ if (copy_len > RPC_DESC_SIZE)
+ goto disable; /* always false, only for verifier. */
+ if (cb->rpc_desc_buf + cb->rpc_desc_buff_len >= &cb->rpc_desc_buf[RPC_DESC_SIZE])
+ goto disable; /* always false, only for verifier. */
+
+ err = bpf_dynptr_read(cb->rpc_desc_buf + cb->rpc_desc_buff_len,
+ copy_len, dptr, rpc_copied_seq - seq, 0);
+ if (err)
+ goto disable;
+
+ cb->rpc_desc_buff_len += copy_len;
+
+ if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+ LOG("Copied %d bytes: rpc_desc_buff_len: %u", copy_len, cb->rpc_desc_buff_len);
+ goto partial;
+ }
+
+ rpc_desc = (struct rpc_descriptor *)cb->rpc_desc_buf;
+ rpc_len = RPC_DESC_SIZE + rpc_desc->header_len + rpc_desc->payload_len;
+
+ if (rpc_len > INT_MAX)
+ goto disable;
+
+ cb->rpc_end_seq = cb->rpc_desc_seq + rpc_len;
+
+ LOG("Copied full descriptor: rpc_desc_seq: %u, rpc_end_seq: %u, header_len: %u, payload_len: %u",
+ CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq),
+ rpc_desc->header_len, rpc_desc->payload_len);
+
+ return 0;
+disable:
+ return -1;
+partial:
+ return 1;
+}
+
+static void tcp_set_autolowat(struct tcp_autolowat_cb *cb,
+ struct sock *sk)
+{
+ struct tcp_sock *tp = (struct tcp_sock *)sk;
+ u32 val; /* To handle wraparound. */
+
+ LOG("Setting rcvlowat: tp->copied_seq: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_desc_buff_len: %u",
+ TP_SEQ(copied_seq), CB_SEQ(rpc_desc_seq),
+ CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+ if (before(tp->copied_seq, cb->rpc_desc_seq))
+ val = cb->rpc_desc_seq - tp->copied_seq;
+ else if (cb->rpc_desc_buff_len != RPC_DESC_SIZE)
+ val = RPC_DESC_SIZE;
+ else
+ val = cb->rpc_end_seq - tp->copied_seq;
+
+ if (val != tp->inet_conn.icsk_inet.sk.sk_rcvlowat) {
+ bpf_tcp_ops_set_rcvlowat(sk, val);
+
+ LOG("Set rcvlowat: expected: %u, actual: %d\n",
+ val, tp->inet_conn.icsk_inet.sk.sk_rcvlowat);
+ } else {
+ LOG("No need to set rcvlowat: %u\n", val);
+ }
+}
+
+static void tcp_disable_autolowat(struct sock *sk)
+{
+ struct tcp_sock *tp = (struct tcp_sock *)sk;
+ int flags;
+
+ flags = tp->bpf_sock_ops_cb_flags & ~BPF_SOCK_OPS_RCVQ_CB_FLAG;
+ bpf_sock_ops_cb_flags_set(sk, flags);
+
+ bpf_tcp_ops_set_rcvlowat(sk, 1);
+
+ LOG("Disabled autolowat");
+}
+
+static void tcp_do_autolowat(struct tcp_autolowat_cb *cb,
+ struct sock *sk, struct sk_buff *skb)
+{
+ struct bpf_dynptr dptr;
+ struct tcp_skb_cb *tcb;
+ u32 seq, end_seq;
+ int ret = 0, i;
+
+ if (bpf_dynptr_from_skb((struct __sk_buff *)skb, 0, &dptr)) {
+ ret = -1;
+ goto update;
+ }
+
+ tcb = bpf_core_cast(skb->cb, struct tcp_skb_cb);
+ seq = tcb->seq;
+ end_seq = tcb->end_seq - !!(tcb->tcp_flags & TCPHDR_FIN);
+
+ LOG("Start parsing skb: seq: %u, end_seq: %u, len: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_buff_len: %u",
+ SEQ(seq), SEQ(end_seq), end_seq - seq,
+ CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+ if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+ ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+ if (ret)
+ goto update;
+ }
+
+ i = 0;
+
+ while (1) {
+ if (i++ > MAX_RPC_DESC_PER_SKB) {
+ ret = -1;
+ break;
+ }
+
+ if (after(cb->rpc_end_seq, end_seq)) {
+ LOG("No more descriptor: rpc_end_seq: %u, end_seq: %u",
+ CB_SEQ(rpc_end_seq), SEQ(end_seq));
+ break;
+ }
+
+ cb->rpc_desc_seq = cb->rpc_end_seq;
+ cb->rpc_desc_buff_len = 0;
+
+ if (cb->rpc_end_seq == end_seq)
+ break;
+
+ LOG("Found next descriptor: rpc_end_seq: %u, end_seq: %u, len: %u",
+ CB_SEQ(rpc_end_seq), SEQ(end_seq), end_seq - cb->rpc_end_seq);
+
+ ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+ if (ret)
+ break;
+ }
+
+update:
+ if (ret >= 0)
+ tcp_set_autolowat(cb, sk);
+ else
+ tcp_disable_autolowat(sk);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_enqueue_rcvq, struct sock *sk, struct sk_buff *skb)
+{
+ struct tcp_autolowat_cb *cb;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+ if (!cb)
+ return;
+
+ tcp_do_autolowat(cb, sk, skb);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_dequeue_rcvq, struct sock *sk)
+{
+ struct tcp_autolowat_cb *cb;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+ if (!cb)
+ return;
+
+ tcp_set_autolowat(cb, sk);
+}
+
+SEC(".struct_ops.link")
+struct bpf_tcp_ops tcp_autolowat_ops = {
+ .enqueue_rcvq = (void *)tcp_autolowat_enqueue_rcvq,
+ .dequeue_rcvq = (void *)tcp_autolowat_dequeue_rcvq,
+};
+
+static int tcp_init_autolowat_cb(struct bpf_sockopt *sockopt,
+ struct bpf_tcp_sock *btp)
+{
+ struct tcp_autolowat_cb *cb;
+ struct tcp_sock *tp;
+ int flags;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, btp, 0,
+ BPF_SK_STORAGE_GET_F_CREATE);
+ if (!cb)
+ return -1;
+
+ tp = bpf_core_cast(btp, struct tcp_sock);
+ if (!tp)
+ return -1;
+
+ cb->rpc_desc_seq = tp->copied_seq;
+ cb->rpc_end_seq = tp->copied_seq;
+#ifdef DEBUG
+ cb->isn = tp->copied_seq;
+#endif
+
+ if (bpf_getsockopt(sockopt->sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
+ &flags, sizeof(flags)))
+ return -1;
+
+ flags |= BPF_SOCK_OPS_RCVQ_CB_FLAG;
+
+ if (bpf_setsockopt(sockopt->sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
+ &flags, sizeof(flags)))
+ return -1;
+
+ return 0;
+}
+
+SEC("cgroup/setsockopt")
+int tcp_autolowat_setsockopt(struct bpf_sockopt *ctx)
+{
+ void *optval_end = ctx->optval_end;
+ int *optval = ctx->optval;
+ struct bpf_tcp_sock *btp;
+
+ if (ctx->level != SOL_BPF || ctx->optname != BPF_TCP_AUTOLOWAT)
+ goto out;
+
+ if (optval + 1 > optval_end)
+ return 0; /* -EPERM */
+
+ btp = bpf_tcp_sock(ctx->sk);
+ if (!btp)
+ goto out;
+
+ if (*optval && tcp_init_autolowat_cb(ctx, btp))
+ return 0; /* -EPERM */
+
+ /*
+ * BPF has consumed this option, don't call kernel
+ * setsockopt handler.
+ */
+ ctx->optlen = -1;
+out:
+ return 1;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
index 6935f32eeb8fd..e30cb1fab0790 100644
--- a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
+++ b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
@@ -92,7 +92,8 @@ int bpf_testcb(struct bpf_sock_ops *skops)
break;
case BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB:
/* Test failure to set largest cb flag (assumes not defined) */
- global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops, 0x80);
+ global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
+ BPF_SOCK_OPS_ALL_CB_FLAGS + 1);
/* Set callback */
global.good_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
BPF_SOCK_OPS_STATE_CB_FLAG);
|
| 3/2 | 2026/09/22 19:35 | llm | patch-evaluator |
3mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 17516 output: 31 thoughts: 1053 Thoughts: **Worth Fuzzing - TCP/BPF Integration** Okay, let's take a look at this patch. My immediate thought is, this looks interesting, particularly given the BPF and TCP involvement. I need to figure out if it's worth the syzkaller treatment. First, I need to understand what's changed. Right, it looks like this introduces a BPF-based queue mechanism within TCP, with modifications to several key files. I see additions in `include/net/tcp.h`, `include/uapi/linux/bpf.h`, `net/core/filter.c`, `net/ipv4/bpf_tcp_ops.c`, `net/ipv4/tcp.c`, and `net/ipv4/tcp_fastopen.c`, and also in `net/ipv4/tcp_input.c`. Specifically, new functions like `__tcp_set_rcvlowat`, the `bpf_tcp_ops` struct is extended, and there's interaction with the core TCP data path. I'm seeing changes in `tcp_queue_rcv`, `tcp_ofo_queue`, `tcp_cleanup_rbuf`, and `tcp_set_rcvlowat`, all of which are very high-value targets, and I see calls to `bpf_tcp_ops_enqueue_rcvq` and `bpf_tcp_ops_dequeue_rcvq`. These are definitely parts of the core TCP stack. Then there are BPF struct ops, including `bpf_tcp_ops_cb_flags_set` and `bpf_tcp_ops_set_rcvlowat`. The patch is also adding new BPF functions exposed to the user. Now, is this reachable code? Absolutely. This is directly touching the TCP input path, BPF structure operations, and core BPF filtering. It seems syzkaller can reach it, too. We can create TCP sockets, attach BPF programs, send/receive data, use setsockopt, and interact with the new functions introduced here. Even without custom BPF programs, the core TCP functions being modified are definitely reachable. Given my understanding and the criteria for fuzzing, I'd say the answer is simple: WorthFuzzing = True. This is core networking, and any logic changes in such areas are definitely worth a deep dive with fuzzing. Now, I'll quickly check the required configs. I'll need to figure out what kernel configurations are required to compile and use `bpf_tcp_ops.c`. I expect there are a few options, likely related to BPF, TCP, and perhaps some new ones. I will start by looking into the kernel's `net/ipv4/Makefile` or similar. |
| 4/2 | 2026/09/22 19:39 | tool | grepper |
0mArgs:{
"Expression": "bpf_tcp_ops",
"PathPrefix": "net/ipv4/"
}Results: {
"Output": "net/ipv4/Makefile=72=obj-$(CONFIG_BPF_SYSCALL) += bpf_tcp_ca.o\nnet/ipv4/Makefile:73:obj-$(CONFIG_CGROUP_BPF) += bpf_tcp_ops.o\nnet/ipv4/Makefile-74-endif\n--\nnet/ipv4/af_inet.c=198=int __inet_listen_sk(struct sock *sk, int backlog)\n--\nnet/ipv4/af_inet.c-229-\t\ttcp_call_bpf(sk, BPF_SOCK_OPS_TCP_LISTEN_CB, 0, NULL);\nnet/ipv4/af_inet.c:230:\t\tbpf_tcp_ops_call(listen, sk);\nnet/ipv4/af_inet.c-231-\t}\n--\nnet/ipv4/bpf_tcp_ops.c=83=static void dequeue_rcvq_stub(struct sock *sk)\n--\nnet/ipv4/bpf_tcp_ops.c-86-\nnet/ipv4/bpf_tcp_ops.c:87:static struct bpf_tcp_ops __bpf_tcp_ops = {\nnet/ipv4/bpf_tcp_ops.c-88-\t.timeout_init = timeout_init_stub,\n--\nnet/ipv4/bpf_tcp_ops.c-104-\nnet/ipv4/bpf_tcp_ops.c:105:BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,\nnet/ipv4/bpf_tcp_ops.c-106-\t u32, len, u64, flags)\n--\nnet/ipv4/bpf_tcp_ops.c-114-\t/*\nnet/ipv4/bpf_tcp_ops.c:115:\t * bpf_tcp_ops does not keep track of the end of the written TCP header\nnet/ipv4/bpf_tcp_ops.c-116-\t * options, so search for it every time the helper is called. The free\n--\nnet/ipv4/bpf_tcp_ops.c-134-\nnet/ipv4/bpf_tcp_ops.c:135:static const struct bpf_func_proto bpf_tcp_ops_store_hdr_opt_proto = {\nnet/ipv4/bpf_tcp_ops.c:136:\t.func\t\t= bpf_tcp_ops_store_hdr_opt,\nnet/ipv4/bpf_tcp_ops.c-137-\t.gpl_only\t= false,\n--\nnet/ipv4/bpf_tcp_ops.c-144-\nnet/ipv4/bpf_tcp_ops.c:145:BPF_CALL_4(bpf_tcp_ops_load_hdr_opt, void *, ctx, void *, search_res,\nnet/ipv4/bpf_tcp_ops.c-146-\t u32, len, u64, flags)\n--\nnet/ipv4/bpf_tcp_ops.c-153-\t * No flags supported. In particular BPF_LOAD_HDR_OPT_TCP_SYN, which\nnet/ipv4/bpf_tcp_ops.c:154:\t * loads from the saved SYN, is not available because bpf_tcp_ops has no\nnet/ipv4/bpf_tcp_ops.c-155-\t * carrier to track the SYN source across the hooks.\n--\nnet/ipv4/bpf_tcp_ops.c-165-\nnet/ipv4/bpf_tcp_ops.c:166:static const struct bpf_func_proto bpf_tcp_ops_load_hdr_opt_proto = {\nnet/ipv4/bpf_tcp_ops.c:167:\t.func\t\t= bpf_tcp_ops_load_hdr_opt,\nnet/ipv4/bpf_tcp_ops.c-168-\t.gpl_only\t= false,\n--\nnet/ipv4/bpf_tcp_ops.c-175-\nnet/ipv4/bpf_tcp_ops.c:176:BPF_CALL_3(bpf_tcp_ops_reserve_hdr_opt, void *, ctx, u32, len, u64, flags)\nnet/ipv4/bpf_tcp_ops.c-177-{\n--\nnet/ipv4/bpf_tcp_ops.c-190-\nnet/ipv4/bpf_tcp_ops.c:191:static const struct bpf_func_proto bpf_tcp_ops_reserve_hdr_opt_proto = {\nnet/ipv4/bpf_tcp_ops.c:192:\t.func\t\t= bpf_tcp_ops_reserve_hdr_opt,\nnet/ipv4/bpf_tcp_ops.c-193-\t.gpl_only\t= false,\n--\nnet/ipv4/bpf_tcp_ops.c-199-\nnet/ipv4/bpf_tcp_ops.c:200:BPF_CALL_0(bpf_tcp_ops_get_retval)\nnet/ipv4/bpf_tcp_ops.c-201-{\n--\nnet/ipv4/bpf_tcp_ops.c-205-\t/* bpf_get_retval() is only exposed to timeout_init/rwnd_init, which\nnet/ipv4/bpf_tcp_ops.c:206:\t * always run via bpf_tcp_ops_call_int(). Its run_ctx carries the int\nnet/ipv4/bpf_tcp_ops.c:207:\t * return value chained across the bpf_tcp_ops attached to the cgroup\nnet/ipv4/bpf_tcp_ops.c-208-\t * and is this program's saved_run_ctx.\n--\nnet/ipv4/bpf_tcp_ops.c-216-\nnet/ipv4/bpf_tcp_ops.c:217:const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {\nnet/ipv4/bpf_tcp_ops.c:218:\t.func\t\t= bpf_tcp_ops_get_retval,\nnet/ipv4/bpf_tcp_ops.c-219-\t.gpl_only\t= false,\n--\nnet/ipv4/bpf_tcp_ops.c-222-\nnet/ipv4/bpf_tcp_ops.c:223:BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)\nnet/ipv4/bpf_tcp_ops.c-224-{\n--\nnet/ipv4/bpf_tcp_ops.c-234-\nnet/ipv4/bpf_tcp_ops.c:235:static const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {\nnet/ipv4/bpf_tcp_ops.c:236:\t.func\t\t= bpf_tcp_ops_cb_flags_set,\nnet/ipv4/bpf_tcp_ops.c-237-\t.gpl_only\t= false,\n--\nnet/ipv4/bpf_tcp_ops.c=245=get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)\n--\nnet/ipv4/bpf_tcp_ops.c-257-\t\t */\nnet/ipv4/bpf_tcp_ops.c:258:\t\tif (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||\nnet/ipv4/bpf_tcp_ops.c:259:\t\t moff == offsetof(struct bpf_tcp_ops, timeout_init) ||\nnet/ipv4/bpf_tcp_ops.c:260:\t\t moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||\nnet/ipv4/bpf_tcp_ops.c:261:\t\t moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))\nnet/ipv4/bpf_tcp_ops.c-262-\t\t\treturn NULL;\n--\nnet/ipv4/bpf_tcp_ops.c-264-\tcase BPF_FUNC_getsockopt:\nnet/ipv4/bpf_tcp_ops.c:265:\t\tif (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||\nnet/ipv4/bpf_tcp_ops.c:266:\t\t moff == offsetof(struct bpf_tcp_ops, timeout_init) ||\nnet/ipv4/bpf_tcp_ops.c:267:\t\t moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||\nnet/ipv4/bpf_tcp_ops.c:268:\t\t moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))\nnet/ipv4/bpf_tcp_ops.c-269-\t\t\treturn NULL;\n--\nnet/ipv4/bpf_tcp_ops.c-271-\tcase BPF_FUNC_get_retval:\nnet/ipv4/bpf_tcp_ops.c:272:\t\tif (moff == offsetof(struct bpf_tcp_ops, timeout_init) ||\nnet/ipv4/bpf_tcp_ops.c:273:\t\t moff == offsetof(struct bpf_tcp_ops, rwnd_init))\nnet/ipv4/bpf_tcp_ops.c:274:\t\t\treturn \u0026bpf_tcp_ops_get_retval_proto;\nnet/ipv4/bpf_tcp_ops.c-275-\t\treturn NULL;\nnet/ipv4/bpf_tcp_ops.c-276-\tcase BPF_FUNC_reserve_hdr_opt:\nnet/ipv4/bpf_tcp_ops.c:277:\t\tif (moff == offsetof(struct bpf_tcp_ops, hdr_opt_len))\nnet/ipv4/bpf_tcp_ops.c:278:\t\t\treturn \u0026bpf_tcp_ops_reserve_hdr_opt_proto;\nnet/ipv4/bpf_tcp_ops.c-279-\t\treturn NULL;\nnet/ipv4/bpf_tcp_ops.c-280-\tcase BPF_FUNC_load_hdr_opt:\nnet/ipv4/bpf_tcp_ops.c:281:\t\tif (moff == offsetof(struct bpf_tcp_ops, parse_hdr) ||\nnet/ipv4/bpf_tcp_ops.c:282:\t\t moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))\nnet/ipv4/bpf_tcp_ops.c:283:\t\t\treturn \u0026bpf_tcp_ops_load_hdr_opt_proto;\nnet/ipv4/bpf_tcp_ops.c-284-\t\treturn NULL;\nnet/ipv4/bpf_tcp_ops.c-285-\tcase BPF_FUNC_store_hdr_opt:\nnet/ipv4/bpf_tcp_ops.c:286:\t\tif (moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))\nnet/ipv4/bpf_tcp_ops.c:287:\t\t\treturn \u0026bpf_tcp_ops_store_hdr_opt_proto;\nnet/ipv4/bpf_tcp_ops.c-288-\t\treturn NULL;\nnet/ipv4/bpf_tcp_ops.c-289-\tcase BPF_FUNC_sock_ops_cb_flags_set:\nnet/ipv4/bpf_tcp_ops.c:290:\t\tif (moff == offsetof(struct bpf_tcp_ops, connect) ||\nnet/ipv4/bpf_tcp_ops.c:291:\t\t moff == offsetof(struct bpf_tcp_ops, listen) ||\nnet/ipv4/bpf_tcp_ops.c:292:\t\t moff == offsetof(struct bpf_tcp_ops, active_established) ||\nnet/ipv4/bpf_tcp_ops.c:293:\t\t moff == offsetof(struct bpf_tcp_ops, passive_established) ||\nnet/ipv4/bpf_tcp_ops.c:294:\t\t moff == offsetof(struct bpf_tcp_ops, enqueue_rcvq) ||\nnet/ipv4/bpf_tcp_ops.c:295:\t\t moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))\nnet/ipv4/bpf_tcp_ops.c:296:\t\t\treturn \u0026bpf_tcp_ops_cb_flags_set_proto;\nnet/ipv4/bpf_tcp_ops.c-297-\t\treturn NULL;\n--\nnet/ipv4/bpf_tcp_ops.c=303=static bool is_valid_access(int off, int size, enum bpf_access_type type,\n--\nnet/ipv4/bpf_tcp_ops.c-317-\nnet/ipv4/bpf_tcp_ops.c:318:static int bpf_tcp_ops_init_member(const struct btf_type *t,\nnet/ipv4/bpf_tcp_ops.c-319-\t\t\t\t const struct btf_member *member,\n--\nnet/ipv4/bpf_tcp_ops.c-324-\nnet/ipv4/bpf_tcp_ops.c:325:static int bpf_tcp_ops_check_member(const struct btf_type *t,\nnet/ipv4/bpf_tcp_ops.c-326-\t\t\t\t const struct btf_member *member,\n--\nnet/ipv4/bpf_tcp_ops.c-334-\nnet/ipv4/bpf_tcp_ops.c:335:static int bpf_tcp_ops_init(struct btf *btf)\nnet/ipv4/bpf_tcp_ops.c-336-{\n--\nnet/ipv4/bpf_tcp_ops.c-339-\nnet/ipv4/bpf_tcp_ops.c:340:static int bpf_tcp_ops_validate(void *kdata)\nnet/ipv4/bpf_tcp_ops.c-341-{\n--\nnet/ipv4/bpf_tcp_ops.c-344-\nnet/ipv4/bpf_tcp_ops.c:345:static const struct bpf_verifier_ops bpf_tcp_ops_verifier = {\nnet/ipv4/bpf_tcp_ops.c-346-\t.get_func_proto\t\t= get_func_proto,\n--\nnet/ipv4/bpf_tcp_ops.c-349-\nnet/ipv4/bpf_tcp_ops.c:350:static struct bpf_struct_ops bpf_tcp_ops = {\nnet/ipv4/bpf_tcp_ops.c:351:\t.verifier_ops = \u0026bpf_tcp_ops_verifier,\nnet/ipv4/bpf_tcp_ops.c:352:\t.init_member = bpf_tcp_ops_init_member,\nnet/ipv4/bpf_tcp_ops.c:353:\t.check_member = bpf_tcp_ops_check_member,\nnet/ipv4/bpf_tcp_ops.c:354:\t.init = bpf_tcp_ops_init,\nnet/ipv4/bpf_tcp_ops.c:355:\t.validate = bpf_tcp_ops_validate,\nnet/ipv4/bpf_tcp_ops.c:356:\t.name = \"bpf_tcp_ops\",\nnet/ipv4/bpf_tcp_ops.c-357-\t.cgroup_atype = CGROUP_TCP_SOCK_OPS,\nnet/ipv4/bpf_tcp_ops.c:358:\t.cfi_stubs = \u0026__bpf_tcp_ops,\nnet/ipv4/bpf_tcp_ops.c-359-\t.owner = THIS_MODULE,\n--\nnet/ipv4/bpf_tcp_ops.c=362=__bpf_kfunc_start_defs();\nnet/ipv4/bpf_tcp_ops.c-363-\nnet/ipv4/bpf_tcp_ops.c:364:__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,\nnet/ipv4/bpf_tcp_ops.c-365-\t\t\t\t\t const struct bpf_prog_aux *aux)\n--\nnet/ipv4/bpf_tcp_ops.c-369-\nnet/ipv4/bpf_tcp_ops.c:370:\tif (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))\nnet/ipv4/bpf_tcp_ops.c-371-\t\twakeup = true;\n--\nnet/ipv4/bpf_tcp_ops.c=379=__bpf_kfunc_end_defs();\nnet/ipv4/bpf_tcp_ops.c-380-\nnet/ipv4/bpf_tcp_ops.c:381:BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)\nnet/ipv4/bpf_tcp_ops.c:382:BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)\nnet/ipv4/bpf_tcp_ops.c:383:BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)\nnet/ipv4/bpf_tcp_ops.c-384-\nnet/ipv4/bpf_tcp_ops.c:385:static int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,\nnet/ipv4/bpf_tcp_ops.c-386-\t\t\t\t\t u32 kfunc_id)\n--\nnet/ipv4/bpf_tcp_ops.c-389-\nnet/ipv4/bpf_tcp_ops.c:390:\tif (!btf_id_set8_contains(\u0026bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))\nnet/ipv4/bpf_tcp_ops.c-391-\t\treturn 0;\nnet/ipv4/bpf_tcp_ops.c-392-\nnet/ipv4/bpf_tcp_ops.c:393:\tif (prog-\u003eaux-\u003est_ops != \u0026bpf_tcp_ops)\nnet/ipv4/bpf_tcp_ops.c-394-\t\treturn -EACCES;\n--\nnet/ipv4/bpf_tcp_ops.c-396-\tmoff = prog-\u003eaux-\u003eattach_st_ops_member_off;\nnet/ipv4/bpf_tcp_ops.c:397:\tif (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) \u0026\u0026\nnet/ipv4/bpf_tcp_ops.c:398:\t moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))\nnet/ipv4/bpf_tcp_ops.c-399-\t\treturn -EACCES;\n--\nnet/ipv4/bpf_tcp_ops.c-403-\nnet/ipv4/bpf_tcp_ops.c:404:static const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {\nnet/ipv4/bpf_tcp_ops.c-405-\t.owner = THIS_MODULE,\nnet/ipv4/bpf_tcp_ops.c:406:\t.set = \u0026bpf_tcp_ops_rcvlowat_kfunc_set,\nnet/ipv4/bpf_tcp_ops.c:407:\t.filter = bpf_tcp_ops_rcvlowat_kfunc_filter,\nnet/ipv4/bpf_tcp_ops.c-408-};\nnet/ipv4/bpf_tcp_ops.c-409-\nnet/ipv4/bpf_tcp_ops.c:410:static int __init __bpf_tcp_ops_init(void)\nnet/ipv4/bpf_tcp_ops.c-411-{\n--\nnet/ipv4/bpf_tcp_ops.c-414-\tret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,\nnet/ipv4/bpf_tcp_ops.c:415:\t\t\t\t\t\u0026bpf_tcp_ops_rcvlowat_kfunc_id_set);\nnet/ipv4/bpf_tcp_ops.c:416:\tret = ret ?: register_bpf_struct_ops(\u0026bpf_tcp_ops, bpf_tcp_ops);\nnet/ipv4/bpf_tcp_ops.c-417-\n--\nnet/ipv4/bpf_tcp_ops.c-419-}\nnet/ipv4/bpf_tcp_ops.c:420:late_initcall(__bpf_tcp_ops_init);\n--\nnet/ipv4/tcp.c=1604=void tcp_cleanup_rbuf(struct sock *sk, int copied)\n--\nnet/ipv4/tcp.c-1613-\nnet/ipv4/tcp.c:1614:\tbpf_tcp_ops_dequeue_rcvq(sk);\nnet/ipv4/tcp.c-1615-}\n--\nnet/ipv4/tcp.c=2969=void tcp_set_state(struct sock *sk, int state)\n--\nnet/ipv4/tcp.c-3007-\t\ttcp_call_bpf_2arg(sk, BPF_SOCK_OPS_STATE_CB, oldstate, state);\nnet/ipv4/tcp.c:3008:\tbpf_tcp_ops_call(set_state, sk, state);\nnet/ipv4/tcp.c-3009-\n--\nnet/ipv4/tcp_fastopen.c=257=void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)\n--\nnet/ipv4/tcp_fastopen.c-283-\nnet/ipv4/tcp_fastopen.c:284:\tbpf_tcp_ops_enqueue_rcvq(sk, skb);\nnet/ipv4/tcp_fastopen.c-285-\n--\nnet/ipv4/tcp_input.c=205=static void bpf_skops_established(struct sock *sk, int bpf_op,\n--\nnet/ipv4/tcp_input.c-210-\nnet/ipv4/tcp_input.c:211:static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)\nnet/ipv4/tcp_input.c-212-{\n--\nnet/ipv4/tcp_input.c-219-\nnet/ipv4/tcp_input.c:220:\tbpf_tcp_ops_call(parse_hdr, sk, skb);\nnet/ipv4/tcp_input.c-221-}\n--\nnet/ipv4/tcp_input.c=5318=static void tcp_ofo_queue(struct sock *sk)\n--\nnet/ipv4/tcp_input.c-5346-\nnet/ipv4/tcp_input.c:5347:\t\tbpf_tcp_ops_enqueue_rcvq(sk, skb);\nnet/ipv4/tcp_input.c-5348-\n--\nnet/ipv4/tcp_input.c=5546=static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,\n--\nnet/ipv4/tcp_input.c-5551-\nnet/ipv4/tcp_input.c:5552:\tbpf_tcp_ops_enqueue_rcvq(sk, skb);\nnet/ipv4/tcp_input.c-5553-\n--\nnet/ipv4/tcp_input.c=6338=static bool tcp_validate_incoming(struct sock *sk, struct sk_buff *skb,\n--\nnet/ipv4/tcp_input.c-6479-\tbpf_skops_parse_hdr(sk, skb);\nnet/ipv4/tcp_input.c:6480:\tbpf_tcp_ops_parse_hdr(sk, skb);\nnet/ipv4/tcp_input.c-6481-\n--\nnet/ipv4/tcp_input.c=6722=void tcp_init_transfer(struct sock *sk, int bpf_op, struct sk_buff *skb)\n--\nnet/ipv4/tcp_input.c-6744-\tif (bpf_op == BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB)\nnet/ipv4/tcp_input.c:6745:\t\tbpf_tcp_ops_call(active_established, sk, skb);\nnet/ipv4/tcp_input.c-6746-\telse\nnet/ipv4/tcp_input.c:6747:\t\tbpf_tcp_ops_call(passive_established, sk, skb);\nnet/ipv4/tcp_input.c-6748-\t/* Initialize congestion control unless BPF initialized it already: */\n--\nnet/ipv4/tcp_output.c=533=static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,\n--\nnet/ipv4/tcp_output.c-579-\t/*\nnet/ipv4/tcp_output.c:580:\t * bpf_tcp_ops portion is NOP-filled (everything past the sockops\nnet/ipv4/tcp_output.c-581-\t * writer's bytes). The writer finds the append point by scanning from\n--\nnet/ipv4/tcp_output.c-583-\t */\nnet/ipv4/tcp_output.c:584:\tbpf_tcp_ops_call(write_hdr_opt, sk, skb, req, syn_skb, synack_type,\nnet/ipv4/tcp_output.c-585-\t\t\t first_opt_off + nr_written);\n--\nnet/ipv4/tcp_output.c=598=static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,\n--\nnet/ipv4/tcp_output.c-606-\nnet/ipv4/tcp_output.c:607:static u32 bpf_tcp_ops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,\nnet/ipv4/tcp_output.c-608-\t\t\t\t struct request_sock *req,\n--\nnet/ipv4/tcp_output.c-618-\nnet/ipv4/tcp_output.c:619:\t/* bpf_tcp_ops_reserve_hdr_opt() reserves space via remaining_out */\nnet/ipv4/tcp_output.c:620:\tbpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, \u0026remaining_out);\nnet/ipv4/tcp_output.c-621-\n--\nnet/ipv4/tcp_output.c=995=static unsigned int tcp_syn_options(struct sock *sk, struct sk_buff *skb,\n--\nnet/ipv4/tcp_output.c-1091-\t\t\t\t\t remaining);\nnet/ipv4/tcp_output.c:1092:\tremaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,\nnet/ipv4/tcp_output.c-1093-\t\t\t\t\t remaining);\n--\nnet/ipv4/tcp_output.c=1099=static unsigned int tcp_synack_options(const struct sock *sk,\n--\nnet/ipv4/tcp_output.c-1181-\t\t\t\t\t synack_type, opts, remaining);\nnet/ipv4/tcp_output.c:1182:\tremaining = bpf_tcp_ops_hdr_opt_len((struct sock *)sk, skb, req, syn_skb,\nnet/ipv4/tcp_output.c-1183-\t\t\t\t\t synack_type, opts, remaining);\n--\nnet/ipv4/tcp_output.c=1191=static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb,\n--\nnet/ipv4/tcp_output.c-1290-\nnet/ipv4/tcp_output.c:1291:\t\tremaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,\nnet/ipv4/tcp_output.c-1292-\t\t\t\t\t\t remaining);\n--\nnet/ipv4/tcp_output.c=3600=int __tcp_retransmit_skb(struct sock *sk, struct sk_buff *skb, int segs)\n--\nnet/ipv4/tcp_output.c-3730-\t\t\t\t TCP_SKB_CB(skb)-\u003eseq, segs, err);\nnet/ipv4/tcp_output.c:3731:\tbpf_tcp_ops_call(retrans, sk, skb, err);\nnet/ipv4/tcp_output.c-3732-\n--\nnet/ipv4/tcp_output.c=4349=int tcp_connect(struct sock *sk)\n--\nnet/ipv4/tcp_output.c-4355-\ttcp_call_bpf(sk, BPF_SOCK_OPS_TCP_CONNECT_CB, 0, NULL);\nnet/ipv4/tcp_output.c:4356:\tbpf_tcp_ops_call(connect, sk);\nnet/ipv4/tcp_output.c-4357-\n--\nnet/ipv4/tcp_timer.c=243=static int tcp_write_timeout(struct sock *sk)\n--\nnet/ipv4/tcp_timer.c-292-\t\t\t\t icsk-\u003eicsk_rto, (int)expired);\nnet/ipv4/tcp_timer.c:293:\tbpf_tcp_ops_call(rto, sk);\nnet/ipv4/tcp_timer.c-294-\n"
} |
| 5/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 26909 output: 33 thoughts: 18 Thoughts: Okay, let's dive into this. Here's how I'd approach that snippet, as though it were my train of thought: **Checking the IPv4 Makefile** Alright, I'm thinking I need to understand something about the IPv4 networking stack build process. My gut feeling is that the relevant information is likely in the `net/ipv4/Makefile`. That's where I usually start when I'm looking at how IPv4 modules are compiled and linked. I'm going to jump straight to around line 73. That's often a good starting point for conditional compilation, object file handling, or perhaps even dependency definitions within that Makefile context. Let's see what that section reveals about the issue I'm working on... I'm expecting something related to a specific module, build flag, or object file. Perhaps that's where the problem lies. |
| 6/2 | 2026/09/22 19:39 | tool | read-file |
0mArgs:{
"File": "net/ipv4/Makefile",
"FirstLine": 65,
"LineCount": 20
}Results: {
"Contents": " 65:\tobj-$(CONFIG_NETLABEL) += cipso_ipv4.o\n 66:\t\n 67:\tobj-$(CONFIG_XFRM) += xfrm4_policy.o xfrm4_state.o xfrm4_input.o \\\n 68:\t\t\t xfrm4_output.o xfrm4_protocol.o\n 69:\tobj-$(CONFIG_TCP_AO) += tcp_ao.o\n 70:\t\n 71:\tifeq ($(CONFIG_BPF_JIT),y)\n 72:\tobj-$(CONFIG_BPF_SYSCALL) += bpf_tcp_ca.o\n 73:\tobj-$(CONFIG_CGROUP_BPF) += bpf_tcp_ops.o\n 74:\tendif\n 75:\t\n 76:\tifdef CONFIG_GCOV_PROFILE_NETFILTER\n 77:\tGCOV_PROFILE_netfilter.o := y\n 78:\tendif\n"
} |
| 7/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 27400 output: 31 thoughts: 54 Thoughts: Here's my summary, as though I'm thinking through this myself: **Investigating the BPF/cgroup Integration** Okay, so I see this line: `obj-$(CONFIG_CGROUP_BPF) += bpf_tcp_ops.o` and it's conditional, existing under an `ifeq ($(CONFIG_BPF_JIT),y)`. Right away, my brain goes, "That's interesting. It looks like we're conditionally including `bpf_tcp_ops.o` based on both the `CONFIG_CGROUP_BPF` *and* the BPF JIT support." The JIT part makes sense; performance is key for BPF. But what's the deal with `CGROUP_BPF`? I need to confirm the role of `CONFIG_CGROUP_BPF` here. My intuition is that this likely enables the integration of BPF programs with cgroups, allowing for network and resource control. So, the logical next step is to immediately dive into the Kconfig to understand the dependencies and purpose of `CONFIG_CGROUP_BPF`. I'll see where this is defined, its description, and any other relevant configurations. That's the key to understanding how we're leveraging BPF here, whether it's related to monitoring, filtering, or controlling TCP operations within a cgroup context. |
| 8/2 | 2026/09/22 19:39 | tool | grepper |
0mArgs:{
"Expression": "config CGROUP_BPF",
"PathPrefix": "init/Kconfig"
}Results: {
"Output": "init/Kconfig=1353=config CGROUP_PERF\n--\ninit/Kconfig-1363-\ninit/Kconfig:1364:config CGROUP_BPF\ninit/Kconfig-1365-\tbool \"Support for eBPF programs attached to cgroups\"\n"
} |
| 9/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 27870 output: 33 thoughts: 17 Thoughts: Okay, let's dive into this. I'm going to grab `init/Kconfig` and zoom in around line 1364. My expert intuition tells me that specific line is likely where a crucial configuration option or dependency is defined. It's probably related to a major feature or a fundamental building block of the kernel, given its location. Time to decipher what's going on there. I anticipate a conditional statement, a menu entry, a default setting, or perhaps a complex dependency chain. Let's see what gems are hidden within... |
| 10/2 | 2026/09/22 19:39 | tool | read-file |
0mArgs:{
"File": "init/Kconfig",
"FirstLine": 1360,
"LineCount": 20
}Results: {
"Contents": "1360:\t\t so that it can monitor performance events among cgroups.\n1361:\t\n1362:\t\t Say N if unsure.\n1363:\t\n1364:\tconfig CGROUP_BPF\n1365:\t\tbool \"Support for eBPF programs attached to cgroups\"\n1366:\t\tdepends on BPF_SYSCALL\n1367:\t\tselect SOCK_CGROUP_DATA\n1368:\t\thelp\n1369:\t\t Allow attaching eBPF programs to a cgroup using the bpf(2)\n1370:\t\t syscall command BPF_PROG_ATTACH.\n1371:\t\n1372:\t\t In which context these programs are accessed depends on the type\n1373:\t\t of attachment. For instance, programs that are attached using\n1374:\t\t BPF_CGROUP_INET_INGRESS will be executed on the ingress path of\n1375:\t\t inet sockets.\n1376:\t\n1377:\tconfig CGROUP_MISC\n1378:\t\tbool \"Misc resource controller\"\n1379:\t\tdefault n\n"
} |
| 11/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 28336 output: 36 thoughts: 120 Thoughts: Here's my thought process, as if I were analyzing these dependencies: **Kernel Configuration Dependencies - Diving Deep** Okay, so I'm looking at `CONFIG_CGROUP_BPF` and it's clear this relies on `BPF_SYSCALL`. That makes sense; the BPF syscall is fundamental. Now, I see that the `net/ipv4/Makefile` also mandates `CONFIG_BPF_JIT=y`. Seems like they're intertwined, which isn't entirely unexpected for networking BPF. Now, I'm thinking about the bigger picture. The `net/ipv4` directory is built when `CONFIG_INET=y`, so *that* dependency is also a piece of the puzzle. I should add that to my mental map. Hmm, anything else potentially lurking? I'm going to quickly check into `BPF_PROG_TYPE_STRUCT_OPS`, because it is also related. Right, time to get more concrete. Let me quickly grep through the code to explicitly confirm these dependencies and uncover any hidden configurations. I'll search for `BPF_PROG_TYPE_STRUCT_OPS` or `BPF_JIT` and see what else comes up. That should give me a more complete picture of the build requirements. |
| 12/2 | 2026/09/22 19:39 | tool | grepper |
0mArgs:{
"Expression": "BPF_PROG_TYPE_STRUCT_OPS",
"PathPrefix": "kernel/bpf/"
}Results: {
"Output": "kernel/bpf/bpf_struct_ops.c=727=static long bpf_struct_ops_map_update_elem(struct bpf_map *map, void *key,\n--\nkernel/bpf/bpf_struct_ops.c-846-\nkernel/bpf/bpf_struct_ops.c:847:\t\tif (prog-\u003etype != BPF_PROG_TYPE_STRUCT_OPS ||\nkernel/bpf/bpf_struct_ops.c-848-\t\t prog-\u003eaux-\u003eattach_btf_id != st_ops_desc-\u003etype_id ||\n--\nkernel/bpf/bpf_struct_ops.c=1540=int bpf_prog_assoc_struct_ops(struct bpf_prog *prog, struct bpf_map *map)\n--\nkernel/bpf/bpf_struct_ops.c-1551-\tif (st_ops_assoc) {\nkernel/bpf/bpf_struct_ops.c:1552:\t\tif (prog-\u003etype != BPF_PROG_TYPE_STRUCT_OPS)\nkernel/bpf/bpf_struct_ops.c-1553-\t\t\treturn -EBUSY;\n--\nkernel/bpf/bpf_struct_ops.c-1560-\t\t */\nkernel/bpf/bpf_struct_ops.c:1561:\t\tif (prog-\u003etype != BPF_PROG_TYPE_STRUCT_OPS)\nkernel/bpf/bpf_struct_ops.c-1562-\t\t\tbpf_map_inc(map);\n--\nkernel/bpf/bpf_struct_ops.c=1570=void bpf_prog_disassoc_struct_ops(struct bpf_prog *prog)\n--\nkernel/bpf/bpf_struct_ops.c-1580-\nkernel/bpf/bpf_struct_ops.c:1581:\tif (prog-\u003etype != BPF_PROG_TYPE_STRUCT_OPS)\nkernel/bpf/bpf_struct_ops.c-1582-\t\tbpf_map_put(st_ops_assoc);\n--\nkernel/bpf/btf.c=6249=static int btf_validate_prog_ctx_type(struct bpf_verifier_log *log, const struct btf *btf,\n--\nkernel/bpf/btf.c-6341-\tcase BPF_PROG_TYPE_LSM:\nkernel/bpf/btf.c:6342:\tcase BPF_PROG_TYPE_STRUCT_OPS:\nkernel/bpf/btf.c-6343-\t\t/* allow u64* as ctx */\n--\nkernel/bpf/btf.c=6648=static bool prog_args_trusted(const struct bpf_prog *prog)\n--\nkernel/bpf/btf.c-6656-\t\treturn bpf_lsm_is_trusted(prog);\nkernel/bpf/btf.c:6657:\tcase BPF_PROG_TYPE_STRUCT_OPS:\nkernel/bpf/btf.c-6658-\t\treturn true;\n--\nkernel/bpf/btf.c=9160=static int bpf_prog_type_to_kfunc_hook(enum bpf_prog_type prog_type)\n--\nkernel/bpf/btf.c-9168-\t\treturn BTF_KFUNC_HOOK_TC;\nkernel/bpf/btf.c:9169:\tcase BPF_PROG_TYPE_STRUCT_OPS:\nkernel/bpf/btf.c-9170-\t\treturn BTF_KFUNC_HOOK_STRUCT_OPS;\n--\nkernel/bpf/core.c=2503=static bool __bpf_prog_map_compatible(struct bpf_map *map,\n--\nkernel/bpf/core.c-2555-\t\t\tcase BPF_PROG_TYPE_EXT:\nkernel/bpf/core.c:2556:\t\t\tcase BPF_PROG_TYPE_STRUCT_OPS:\nkernel/bpf/core.c-2557-\t\t\t\tret = false;\n--\nkernel/bpf/cpumask.c=515=static int __init cpumask_kfunc_init(void)\n--\nkernel/bpf/cpumask.c-526-\tret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_TRACING, \u0026cpumask_kfunc_set);\nkernel/bpf/cpumask.c:527:\tret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS, \u0026cpumask_kfunc_set);\nkernel/bpf/cpumask.c-528-\tret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_SYSCALL, \u0026cpumask_kfunc_set);\n--\nkernel/bpf/helpers.c=5113=static int __init kfunc_init(void)\n--\nkernel/bpf/helpers.c-5131-\tret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_XDP, \u0026generic_kfunc_set);\nkernel/bpf/helpers.c:5132:\tret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS, \u0026generic_kfunc_set);\nkernel/bpf/helpers.c-5133-\tret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_SYSCALL, \u0026generic_kfunc_set);\n--\nkernel/bpf/syscall.c=2731=bpf_prog_load_check_attach(enum bpf_prog_type prog_type,\n--\nkernel/bpf/syscall.c-2746-\t\tcase BPF_PROG_TYPE_LSM:\nkernel/bpf/syscall.c:2747:\t\tcase BPF_PROG_TYPE_STRUCT_OPS:\nkernel/bpf/syscall.c-2748-\t\tcase BPF_PROG_TYPE_EXT:\n--\nkernel/bpf/syscall.c=2877=static bool is_perfmon_prog_type(enum bpf_prog_type prog_type)\n--\nkernel/bpf/syscall.c-2886-\tcase BPF_PROG_TYPE_LSM:\nkernel/bpf/syscall.c:2887:\tcase BPF_PROG_TYPE_STRUCT_OPS: /* has access to struct sock */\nkernel/bpf/syscall.c-2888-\tcase BPF_PROG_TYPE_EXT: /* extends any prog */\n--\nkernel/bpf/syscall.c=6266=static int prog_assoc_struct_ops(union bpf_attr *attr)\n--\nkernel/bpf/syscall.c-6281-\nkernel/bpf/syscall.c:6282:\tif (prog-\u003etype == BPF_PROG_TYPE_STRUCT_OPS) {\nkernel/bpf/syscall.c-6283-\t\tret = -EINVAL;\n--\nkernel/bpf/verifier.c=5326=static enum priv_stack_mode bpf_enable_priv_stack(struct bpf_prog *prog)\n--\nkernel/bpf/verifier.c-5342-\tcase BPF_PROG_TYPE_LSM:\nkernel/bpf/verifier.c:5343:\tcase BPF_PROG_TYPE_STRUCT_OPS:\nkernel/bpf/verifier.c-5344-\t\tif (prog-\u003eaux-\u003epriv_stack_requested || bpf_prog_check_recur(prog))\n--\nkernel/bpf/verifier.c=11274=static int check_reference_leak(struct bpf_verifier_env *env, bool exception_exit)\n--\nkernel/bpf/verifier.c-11290-\t\t */\nkernel/bpf/verifier.c:11291:\t\tif (type == BPF_PROG_TYPE_STRUCT_OPS \u0026\u0026 !exception_exit \u0026\u0026\nkernel/bpf/verifier.c-11292-\t\t reg-\u003eid == state-\u003erefs[i].id)\n--\nkernel/bpf/verifier.c=17793=static bool return_retval_range(struct bpf_verifier_env *env, struct bpf_retval_range *range)\n--\nkernel/bpf/verifier.c-17887-\t\tbreak;\nkernel/bpf/verifier.c:17888:\tcase BPF_PROG_TYPE_STRUCT_OPS:\nkernel/bpf/verifier.c-17889-\t\t*range = retval_range(0, 0);\n--\nkernel/bpf/verifier.c=17904=static bool program_returns_void(struct bpf_verifier_env *env)\n--\nkernel/bpf/verifier.c-17915-\t\tbreak;\nkernel/bpf/verifier.c:17916:\tcase BPF_PROG_TYPE_STRUCT_OPS:\nkernel/bpf/verifier.c-17917-\t\tif (!prog-\u003eaux-\u003eattach_func_proto-\u003etype)\n--\nkernel/bpf/verifier.c=17936=static int check_return_code(struct bpf_verifier_env *env, int regno, const char *reg_name)\n--\nkernel/bpf/verifier.c-17951-\nkernel/bpf/verifier.c:17952:\tif (prog_type == BPF_PROG_TYPE_STRUCT_OPS) {\nkernel/bpf/verifier.c-17953-\t\t/* Allow a struct_ops program to return a referenced kptr if it\n--\nkernel/bpf/verifier.c-17985-\nkernel/bpf/verifier.c:17986:\tif (prog_type == BPF_PROG_TYPE_STRUCT_OPS \u0026\u0026 !ret_type)\nkernel/bpf/verifier.c-17987-\t\treturn 0;\n--\nkernel/bpf/verifier.c=19857=static int do_check_common(struct bpf_verifier_env *env, int subprog)\n--\nkernel/bpf/verifier.c-19993-\t/* Acquire references for struct_ops program arguments tagged with \"__ref\" */\nkernel/bpf/verifier.c:19994:\tif (!subprog \u0026\u0026 env-\u003eprog-\u003etype == BPF_PROG_TYPE_STRUCT_OPS) {\nkernel/bpf/verifier.c-19995-\t\tfor (i = 0; i \u003c aux-\u003ectx_arg_info_size; i++) {\n--\nkernel/bpf/verifier.c=20869=static bool can_be_sleepable(struct bpf_prog *prog)\n--\nkernel/bpf/verifier.c-20890-\treturn prog-\u003etype == BPF_PROG_TYPE_KPROBE /* only for uprobes */ ||\nkernel/bpf/verifier.c:20891:\t prog-\u003etype == BPF_PROG_TYPE_STRUCT_OPS ||\nkernel/bpf/verifier.c-20892-\t prog-\u003etype == BPF_PROG_TYPE_RAW_TRACEPOINT ||\n--\nkernel/bpf/verifier.c=20896=static int check_attach_btf_id(struct bpf_verifier_env *env)\n--\nkernel/bpf/verifier.c-20918-\nkernel/bpf/verifier.c:20919:\tif (prog-\u003etype == BPF_PROG_TYPE_STRUCT_OPS)\nkernel/bpf/verifier.c-20920-\t\treturn check_struct_ops_btf_id(env);\n"
} |
| 13/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31709 output: 33 thoughts: 26 Thoughts: Here's my summary, as though I'm thinking through this process: **Diving into the `bpf_struct_ops` Structure's Construction** Alright, time to get a handle on how this `bpf_struct_ops` structure is actually put together. The natural starting point, as always, is the `Makefile`. Specifically, I need to check out what's happening within `kernel/bpf/Makefile`. That's where the build process for this critical piece of BPF infrastructure should be orchestrated. Let's see what clues are there... |
| 14/2 | 2026/09/22 19:39 | tool | grepper |
0mArgs:{
"Expression": "bpf_struct_ops\\.o",
"PathPrefix": "kernel/bpf/"
}Results: {
"Output": "kernel/bpf/Makefile=46=ifeq ($(CONFIG_BPF_JIT),y)\nkernel/bpf/Makefile:47:obj-$(CONFIG_BPF_SYSCALL) += bpf_struct_ops.o\nkernel/bpf/Makefile-48-obj-$(CONFIG_BPF_SYSCALL) += cpumask.o\n"
} |
| 15/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31987 output: 39 thoughts: 87 Thoughts: Okay, here's my summary, considering my expertise: **Checking BPF Configuration Dependencies** Alright, let's see, I'm going through the necessary kernel configurations for this BPF project. I've already confirmed `BPF_SYSCALL`, `BPF_JIT`, `CGROUP_BPF`, and `INET` are enabled. I need to make sure I haven't missed anything crucial. Let me think... Are there any other hidden dependencies? Wait, `CONFIG_NET`? Nope, `INET` handles that, that's fine. Okay, let's specifically look at this `bpf_tcp_ops.c` file. I'll take a look at line 1 of that file. This is where I can start putting things together. |
| 16/2 | 2026/09/22 19:39 | tool | read-file |
0mArgs:{
"File": "net/ipv4/bpf_tcp_ops.c",
"FirstLine": 1,
"LineCount": 40
}Results: {
"Contents": " 1:\t// SPDX-License-Identifier: GPL-2.0\n 2:\t/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */\n 3:\t\n 4:\t#include \u003clinux/bpf.h\u003e\n 5:\t#include \u003clinux/btf_ids.h\u003e\n 6:\t#include \u003clinux/bpf_verifier.h\u003e\n 7:\t#include \u003clinux/filter.h\u003e\n 8:\t#include \u003cnet/bpf_sk_storage.h\u003e\n 9:\t#include \u003cnet/tcp.h\u003e\n 10:\t\n 11:\tstatic int timeout_init_stub(struct sock *sk, struct request_sock *req__nullable)\n 12:\t{\n 13:\t\tstruct bpf_tramp_run_ctx *ctx =\n 14:\t\t\tcontainer_of(current-\u003ebpf_ctx, struct bpf_tramp_run_ctx, run_ctx);\n 15:\t\n 16:\t\treturn ctx-\u003eretval;\n 17:\t}\n 18:\t\n 19:\tstatic int rwnd_init_stub(struct sock *sk, struct request_sock *req__nullable)\n 20:\t{\n 21:\t\tstruct bpf_tramp_run_ctx *ctx =\n 22:\t\t\tcontainer_of(current-\u003ebpf_ctx, struct bpf_tramp_run_ctx, run_ctx);\n 23:\t\n 24:\t\treturn ctx-\u003eretval;\n 25:\t}\n 26:\t\n 27:\tstatic void active_established_stub(struct sock *sk, struct sk_buff *skb__nullable)\n 28:\t{\n 29:\t}\n 30:\t\n 31:\tstatic void passive_established_stub(struct sock *sk, struct sk_buff *skb)\n 32:\t{\n 33:\t}\n 34:\t\n 35:\tstatic void rto_stub(struct sock *sk)\n 36:\t{\n 37:\t}\n 38:\t\n 39:\tstatic void rtt_stub(struct sock *sk, long mrtt, u32 srtt)\n 40:\t{\n"
} |
| 17/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32845 output: 30 thoughts: 77 Thoughts: Okay, let's see what I'm thinking here. **Investigating `bpf_tcp_ops_rcvlowat_kfunc_set` and BTF Configuration** Alright, I need to get a handle on how `bpf_tcp_ops_rcvlowat_kfunc_set` in `net/ipv4/bpf_tcp_ops.c` works. First things first, I recall that `BTF_KFUNCS_START` has dependencies. It's a standard practice, and I already suspect it is related to BTF (BPF Type Format) information. To confirm my suspicion, I'll need to double-check that `CONFIG_DEBUG_INFO_BTF` is required when using `BTF_KFUNCS_START`. Let's confirm that that is the case. |
| 18/2 | 2026/09/22 19:39 | tool | grepper |
0mArgs:{
"Expression": "BTF_KFUNCS_START",
"PathPrefix": "include/"
}Results: {
"Output": "include/linux/btf_ids.h=210=extern struct btf_id_set8 name;\ninclude/linux/btf_ids.h-211-\ninclude/linux/btf_ids.h:212:#define BTF_KFUNCS_START(name)\t\t\t\t\\\ninclude/linux/btf_ids.h-213-__BTF_SET8_START(name, local, BTF_SET8_KFUNCS)\n--\ninclude/linux/btf_ids.h=216=BTF_SET8_END(name)\n--\ninclude/linux/btf_ids.h-231-#define BTF_SET8_END(name)\ninclude/linux/btf_ids.h:232:#define BTF_KFUNCS_START(name) static struct btf_id_set8 __maybe_unused name = { .flags = BTF_SET8_KFUNCS };\ninclude/linux/btf_ids.h-233-#define BTF_KFUNCS_END(name)\n"
} |
| 19/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 33398 output: 31 thoughts: 51 Thoughts: Here's my take on the thought process, summarized as you requested: **Investigating BTF Requirements for BPF Struct Ops and Kfuncs** Okay, so I'm thinking about BPF and specifically the requirements for `struct_ops` and `kfuncs`. My experience suggests that `DEBUG_INFO_BTF` is usually a necessary ingredient in the mix for these, particularly when diving into the more intricate BPF functionality. Therefore, I want to confirm that understanding. The logical next step is to examine the relevant source code for confirmation. Let's dig into `kernel/bpf/bpf_struct_ops.c` and see how `DEBUG_INFO_BTF` is used, and what its implications are in that context. That should quickly clear up whether my initial hunch is correct, or if there's more to it than I'm currently considering. |
| 20/2 | 2026/09/22 19:39 | tool | grepper |
0mArgs:{
"Expression": "DEBUG_INFO_BTF",
"PathPrefix": "kernel/bpf/"
}Results: {
"Output": "kernel/bpf/Makefile=43=ifeq ($(CONFIG_SYSFS),y)\nkernel/bpf/Makefile:44:obj-$(CONFIG_DEBUG_INFO_BTF) += sysfs_btf.o\nkernel/bpf/Makefile-45-endif\n--\nkernel/bpf/btf.c=6417=static struct btf *btf_parse_base(struct btf_verifier_env *env, const char *name,\n--\nkernel/bpf/btf.c-6422-\nkernel/bpf/btf.c:6423:\tif (!IS_ENABLED(CONFIG_DEBUG_INFO_BTF))\nkernel/bpf/btf.c-6424-\t\treturn ERR_PTR(-ENOENT);\n--\nkernel/bpf/btf.c=6503=__u32 btf_relocate_id(const struct btf *btf, __u32 id)\n--\nkernel/bpf/btf.c-6509-\nkernel/bpf/btf.c:6510:#ifdef CONFIG_DEBUG_INFO_BTF_MODULES\nkernel/bpf/btf.c-6511-\nkernel/bpf/btf.c=6512=static struct btf *btf_parse_module(const char *module_name, const void *data,\n--\nkernel/bpf/btf.c-6607-\nkernel/bpf/btf.c:6608:#endif /* CONFIG_DEBUG_INFO_BTF_MODULES */\nkernel/bpf/btf.c-6609-\n--\nkernel/bpf/btf.c=8588=enum {\n--\nkernel/bpf/btf.c-8591-\nkernel/bpf/btf.c:8592:#ifdef CONFIG_DEBUG_INFO_BTF_MODULES\nkernel/bpf/btf.c-8593-struct btf_module {\n--\nkernel/bpf/btf.c=8730=fs_initcall(btf_module_init);\nkernel/bpf/btf.c:8731:#endif /* CONFIG_DEBUG_INFO_BTF_MODULES */\nkernel/bpf/btf.c-8732-\nkernel/bpf/btf.c=8733=struct module *btf_try_get_module(const struct btf *btf)\n--\nkernel/bpf/btf.c-8735-\tstruct module *res = NULL;\nkernel/bpf/btf.c:8736:#ifdef CONFIG_DEBUG_INFO_BTF_MODULES\nkernel/bpf/btf.c-8737-\tstruct btf_module *btf_mod, *tmp;\n--\nkernel/bpf/btf.c=8763=struct btf *btf_get_module_btf(const struct module *module)\nkernel/bpf/btf.c-8764-{\nkernel/bpf/btf.c:8765:#ifdef CONFIG_DEBUG_INFO_BTF_MODULES\nkernel/bpf/btf.c-8766-\tstruct btf_module *btf_mod, *tmp;\n--\nkernel/bpf/btf.c-8776-\nkernel/bpf/btf.c:8777:#ifdef CONFIG_DEBUG_INFO_BTF_MODULES\nkernel/bpf/btf.c-8778-\tmutex_lock(\u0026btf_module_mutex);\n--\nkernel/bpf/btf.c=8793=static int check_btf_kconfigs(const struct module *module, const char *feature)\nkernel/bpf/btf.c-8794-{\nkernel/bpf/btf.c:8795:\tif (!module \u0026\u0026 IS_ENABLED(CONFIG_DEBUG_INFO_BTF)) {\nkernel/bpf/btf.c-8796-\t\tpr_err(\"missing vmlinux BTF, cannot register %s\\n\", feature);\n--\nkernel/bpf/btf.c-8798-\t}\nkernel/bpf/btf.c:8799:\tif (module \u0026\u0026 IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES))\nkernel/bpf/btf.c-8800-\t\tpr_warn(\"missing module BTF, cannot register %s\\n\", feature);\n--\nkernel/bpf/btf.c=8941=static int btf_check_kfunc_name(struct btf *btf, const char *func_name, u32 kind)\nkernel/bpf/btf.c-8942-{\nkernel/bpf/btf.c:8943:#ifdef CONFIG_DEBUG_INFO_BTF_MODULES\nkernel/bpf/btf.c-8944-\tstruct btf_module *btf_mod, *tmp;\n--\nkernel/bpf/btf.c-8957-\nkernel/bpf/btf.c:8958:#ifdef CONFIG_DEBUG_INFO_BTF_MODULES\nkernel/bpf/btf.c-8959-\tguard(mutex)(\u0026btf_module_mutex);\n--\nkernel/bpf/btf.c=9606=static struct bpf_cand_cache *populate_cand_cache(struct bpf_cand_cache *cands,\n--\nkernel/bpf/btf.c-9633-\nkernel/bpf/btf.c:9634:#ifdef CONFIG_DEBUG_INFO_BTF_MODULES\nkernel/bpf/btf.c-9635-static void __purge_cand_cache(struct btf *btf, struct bpf_cand_cache **cache,\n--\nkernel/bpf/verifier.c=2855=int bpf_add_kfunc_call(struct bpf_verifier_env *env, u32 func_id, u16 offset)\n--\nkernel/bpf/verifier.c-2871-\t\tif (!btf_vmlinux) {\nkernel/bpf/verifier.c:2872:\t\t\tverbose(env, \"calling kernel function is not supported without CONFIG_DEBUG_INFO_BTF\\n\");\nkernel/bpf/verifier.c-2873-\t\t\treturn -ENOTSUPP;\n--\nkernel/bpf/verifier.c=6241=static int check_ptr_to_map_access(struct bpf_verifier_env *env,\n--\nkernel/bpf/verifier.c-6255-\tif (!btf_vmlinux) {\nkernel/bpf/verifier.c:6256:\t\tverbose(env, \"map_ptr access not supported without CONFIG_DEBUG_INFO_BTF\\n\");\nkernel/bpf/verifier.c-6257-\t\treturn -ENOTSUPP;\n--\nkernel/bpf/verifier.c=19142=static int check_pseudo_btf_id(struct bpf_verifier_env *env,\n--\nkernel/bpf/verifier.c-19158-\t\tif (!btf_vmlinux) {\nkernel/bpf/verifier.c:19159:\t\t\tverbose(env, \"kernel is missing BTF, make sure CONFIG_DEBUG_INFO_BTF=y is specified in Kconfig.\\n\");\nkernel/bpf/verifier.c-19160-\t\t\treturn -EINVAL;\n--\nkernel/bpf/verifier.c=21053=struct btf *bpf_get_btf_vmlinux(void)\n--\nkernel/bpf/verifier.c-21057-\nkernel/bpf/verifier.c:21058:\tif (!btf \u0026\u0026 IS_ENABLED(CONFIG_DEBUG_INFO_BTF)) {\nkernel/bpf/verifier.c-21059-\t\tmutex_lock(\u0026btf_vmlinux_lock);\n"
} |
| 21/2 | 2026/09/22 19:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 35620 output: 34 thoughts: 1680 Thoughts: **Analyzing BPF-Related TCP Changes for Focused Fuzzing** Okay, so I need to figure out the right symbols to focus on for this BPF-related TCP code. First, let's confirm the necessary features: I need `DEBUG_INFO_BTF` for BTF, kfuncs, and struct_ops, and also `CGROUP_BPF`, `BPF_SYSCALL`, `BPF_JIT`, and `INET`. Good, got that. Now, let's dig into the added and modified functions. I see changes in `net/ipv4/tcp.c`, `net/ipv4/tcp_input.c`, `net/ipv4/tcp_fastopen.c`, `net/ipv4/bpf_tcp_ops.c`, and `net/core/filter.c`. I have a list: `__tcp_set_rcvlowat`, `tcp_set_rcvlowat`, `tcp_cleanup_rbuf`, `tcp_ofo_queue`, `tcp_queue_rcv`, `tcp_fastopen_add_skb`, `bpf_tcp_ops_cb_flags_set`, `bpf_tcp_ops_set_rcvlowat`, `bpf_tcp_ops_rcvlowat_kfunc_filter`, `__bpf_tcp_ops_init`, `bpf_sol_tcp_setsockopt`, and `bpf_sock_ops_cb_flags_set`. Next, I need to deal with those pesky static inline functions in the headers, specifically `tcp_set_sock_ops_cb_flags`, `bpf_tcp_ops_enqueue_rcvq`, and `bpf_tcp_ops_dequeue_rcvq` which are in `include/net/tcp.h`. The guidance is clear: find non-static caller functions in .c files. Let's trace their callers. I find `bpf_sol_tcp_setsockopt`, `bpf_sock_ops_cb_flags_set`, and `bpf_tcp_ops_cb_flags_set` calling `tcp_set_sock_ops_cb_flags`. For the others, `tcp_fastopen_add_skb`, `tcp_ofo_queue`, and `tcp_queue_rcv` call `bpf_tcp_ops_enqueue_rcvq`, and `tcp_cleanup_rbuf` calls `bpf_tcp_ops_dequeue_rcvq`. Okay, so now I have a more complete list of changed functions, including the non-inline call sites: `__tcp_set_rcvlowat`, `tcp_set_rcvlowat`, `bpf_tcp_ops_set_rcvlowat`, `bpf_tcp_ops_cb_flags_set`, `bpf_sock_ops_cb_flags_set`, `bpf_sol_tcp_setsockopt`, `bpf_tcp_ops_rcvlowat_kfunc_filter`, `tcp_cleanup_rbuf`, `tcp_queue_rcv`, `tcp_ofo_queue`, and `tcp_fastopen_add_skb`. Now the crucial part: I have to avoid "ubiquitous" functions. `tcp_cleanup_rbuf` and `tcp_queue_rcv` look suspect, since receive paths can be high frequency. I need to focus on functions specific to this feature. I need to avoid "generic, ubiquitous functions called by almost every program in the corpus". Therefore, I'm going to steer clear of `tcp_cleanup_rbuf` and `tcp_queue_rcv`. They might cause the fuzzer to waste effort. So, I need to focus on: `bpf_tcp_ops_set_rcvlowat`, `bpf_tcp_ops_cb_flags_set`, `bpf_sock_ops_cb_flags_set`, `bpf_sol_tcp_setsockopt`, `__tcp_set_rcvlowat`, `tcp_set_rcvlowat`, and `bpf_tcp_ops_rcvlowat_kfunc_filter`. Also, `tcp_fastopen_add_skb` and `tcp_ofo_queue` look specific enough. Let's check `__tcp_set_rcvlowat` and `tcp_set_rcvlowat`. Yes, `__tcp_set_rcvlowat` is new, replacing the old `tcp_set_rcvlowat`. Excellent, and `bpf_tcp_ops_set_rcvlowat` calls `__tcp_set_rc |
| 22/2 | 2026/09/22 19:40 | tool | grepper |
0mArgs:{
"Expression": "#define BPF_CALL_",
"PathPrefix": "include/linux/filter.h"
}Results: {
"Output": "include/linux/filter.h=40=struct ctl_table_header;\n--\ninclude/linux/filter.h-85-/* unused opcode to mark call to interpreter with arguments */\ninclude/linux/filter.h:86:#define BPF_CALL_ARGS\t0xe0\ninclude/linux/filter.h-87-\n--\ninclude/linux/filter.h=424=static inline int bpf_atomic_load_reg(const struct bpf_insn *insn)\n--\ninclude/linux/filter.h-512-\ninclude/linux/filter.h:513:#define BPF_CALL_REL(TGT)\t\t\t\t\t\\\ninclude/linux/filter.h-514-\t((struct bpf_insn) {\t\t\t\t\t\\\n--\ninclude/linux/filter.h-522-\ninclude/linux/filter.h:523:#define BPF_CALL_IMM(x)\t((void *)(x) - (void *)__bpf_call_base)\ninclude/linux/filter.h-524-\n--\ninclude/linux/filter.h-534-\ninclude/linux/filter.h:535:#define BPF_CALL_KFUNC(OFF, IMM)\t\t\t\t\\\ninclude/linux/filter.h-536-\t((struct bpf_insn) {\t\t\t\t\t\\\n--\ninclude/linux/filter.h-665-\ninclude/linux/filter.h:666:#define BPF_CALL_x(x, attr, name, ...)\t\t\t\t\t \\\ninclude/linux/filter.h-667-\tstatic __always_inline\t\t\t\t\t\t \\\n--\ninclude/linux/filter.h-678-#define __NOATTR\ninclude/linux/filter.h:679:#define BPF_CALL_0(name, ...)\tBPF_CALL_x(0, __NOATTR, name, __VA_ARGS__)\ninclude/linux/filter.h:680:#define BPF_CALL_1(name, ...)\tBPF_CALL_x(1, __NOATTR, name, __VA_ARGS__)\ninclude/linux/filter.h:681:#define BPF_CALL_2(name, ...)\tBPF_CALL_x(2, __NOATTR, name, __VA_ARGS__)\ninclude/linux/filter.h:682:#define BPF_CALL_3(name, ...)\tBPF_CALL_x(3, __NOATTR, name, __VA_ARGS__)\ninclude/linux/filter.h:683:#define BPF_CALL_4(name, ...)\tBPF_CALL_x(4, __NOATTR, name, __VA_ARGS__)\ninclude/linux/filter.h:684:#define BPF_CALL_5(name, ...)\tBPF_CALL_x(5, __NOATTR, name, __VA_ARGS__)\ninclude/linux/filter.h-685-\n"
} |
| 23/2 | 2026/09/22 19:40 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 39062 output: 35 thoughts: 25 Thoughts: Okay, let's dive into this. Here's my understanding, presented as a personal thought process, given my expertise: **Examining the BPF_CALL_x Macro Expansion** Alright, I need to get a clear picture of what's going on with these `BPF_CALL_x` macros in `include/linux/filter.h`. I've been working with eBPF for long enough to know the intricacies of this, so let's cut through the noise and get to the heart of it. The kernel's probably using a sophisticated system to handle these function calls inside the BPF programs, and I need to see how they're abstracting it. Knowing how these macros are expanded is crucial for understanding how the BPF interpreter resolves these calls and passes arguments. Is it through immediate values? Are they cleverly leveraging some form of indirect addressing? Perhaps it's a combination. This level of detail is key for optimization. I'm looking for any hints on how the compiler might be handling inlining, argument passing, register usage, and potential security considerations. I want to see if they are doing bounds checking or handling function pointers in a safe way. Time to dig in and see the details of how it's actually built. |
| 24/2 | 2026/09/22 19:40 | tool | read-file |
0mArgs:{
"File": "include/linux/filter.h",
"FirstLine": 666,
"LineCount": 25
}Results: {
"Contents": " 666:\t#define BPF_CALL_x(x, attr, name, ...)\t\t\t\t\t \\\n 667:\t\tstatic __always_inline\t\t\t\t\t\t \\\n 668:\t\tu64 ____##name(__BPF_MAP(x, __BPF_DECL_ARGS, __BPF_V, __VA_ARGS__)); \\\n 669:\t\ttypedef u64 (*btf_##name)(__BPF_MAP(x, __BPF_DECL_ARGS, __BPF_V, __VA_ARGS__)); \\\n 670:\t\tattr u64 name(__BPF_REG(x, __BPF_DECL_REGS, __BPF_N, __VA_ARGS__)); \\\n 671:\t\tattr u64 name(__BPF_REG(x, __BPF_DECL_REGS, __BPF_N, __VA_ARGS__)) \\\n 672:\t\t{\t\t\t\t\t\t\t\t \\\n 673:\t\t\treturn ((btf_##name)____##name)(__BPF_MAP(x,__BPF_CAST,__BPF_N,__VA_ARGS__));\\\n 674:\t\t}\t\t\t\t\t\t\t\t \\\n 675:\t\tstatic __always_inline\t\t\t\t\t\t \\\n 676:\t\tu64 ____##name(__BPF_MAP(x, __BPF_DECL_ARGS, __BPF_V, __VA_ARGS__))\n 677:\t\n 678:\t#define __NOATTR\n 679:\t#define BPF_CALL_0(name, ...)\tBPF_CALL_x(0, __NOATTR, name, __VA_ARGS__)\n 680:\t#define BPF_CALL_1(name, ...)\tBPF_CALL_x(1, __NOATTR, name, __VA_ARGS__)\n 681:\t#define BPF_CALL_2(name, ...)\tBPF_CALL_x(2, __NOATTR, name, __VA_ARGS__)\n 682:\t#define BPF_CALL_3(name, ...)\tBPF_CALL_x(3, __NOATTR, name, __VA_ARGS__)\n 683:\t#define BPF_CALL_4(name, ...)\tBPF_CALL_x(4, __NOATTR, name, __VA_ARGS__)\n 684:\t#define BPF_CALL_5(name, ...)\tBPF_CALL_x(5, __NOATTR, name, __VA_ARGS__)\n 685:\t\n 686:\t#define NOTRACE_BPF_CALL_1(name, ...)\tBPF_CALL_x(1, notrace, name, __VA_ARGS__)\n 687:\t\n 688:\t#define bpf_ctx_range(TYPE, MEMBER)\t\t\t\t\t\t\\\n 689:\t\toffsetof(TYPE, MEMBER) ... offsetofend(TYPE, MEMBER) - 1\n 690:\t#define bpf_ctx_range_till(TYPE, MEMBER1, MEMBER2)\t\t\t\t\\\n"
} |
| 25/2 | 2026/09/22 19:40 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 40125 output: 336 thoughts: 1887 Thoughts: **Analysis of TCP BPF Integration and Fuzzing Strategy** Okay, so the initial hypothesis holds: `u64 name(...)` translates directly to `name` as the symbol. That's good; it confirms the symbol names like `bpf_tcp_ops_cb_flags_set` and `bpf_sock_ops_cb_flags_set` are actually accessible. Now, let's dissect this commit thoroughly. First, I examined the modified files: `include/net/tcp.h`, `include/uapi/linux/bpf.h`, `net/core/filter.c`, `net/ipv4/bpf_tcp_ops.c`, `net/ipv4/tcp.c`, `net/ipv4/tcp_fastopen.c`, `net/ipv4/tcp_input.c`, and the `tools/...` directory. This provides the scope of the changes. Next, I meticulously went through the functions added or modified, focusing on what they do, their context, and their likely impact on the TCP stack. Specifically, I noted: `__tcp_set_rcvlowat`, `tcp_set_rcvlowat`, `tcp_cleanup_rbuf`, `bpf_tcp_ops_cb_flags_set`, `bpf_tcp_ops_set_rcvlowat`, `bpf_tcp_ops_rcvlowat_kfunc_filter`, `__bpf_tcp_ops_init`, `bpf_sock_ops_cb_flags_set`, `bpf_sol_tcp_setsockopt`, `tcp_ofo_queue`, `tcp_queue_rcv`, and `tcp_fastopen_add_skb`. I'm paying close attention to these inline functions: `tcp_set_sock_ops_cb_flags`, `bpf_tcp_ops_enqueue_rcvq`, and `bpf_tcp_ops_dequeue_rcvq`. They are important. Now, the crucial part: selecting the right `FocusSymbols`. It's vital to avoid diluting the fuzzing effort. Ubiquitous lifecycle functions are a no-go. The goal is to identify functions that implement the *specific* feature logic or entrypoints. Based on that, I've curated a list of good candidates, which are: `__tcp_set_rcvlowat`, `tcp_set_rcvlowat`, `bpf_tcp_ops_set_rcvlowat`, `bpf_tcp_ops_cb_flags_set`, `bpf_sock_ops_cb_flags_set`, `bpf_tcp_ops_rcvlowat_kfunc_filter`, `bpf_sol_tcp_setsockopt`, `tcp_ofo_queue`, and `tcp_fastopen_add_skb`. These are much better, they target specific feature logic and entrypoints. I also considered `tcp_cleanup_rbuf` and `tcp_queue_rcv`, but those are too frequently called. While relevant, including them would spread the fuzzing effort too thinly. Next, config dependencies: I've verified that `CONFIG_CGROUP_BPF` is required, which depends on `CONFIG_BPF_SYSCALL`. I've confirmed that `CONFIG_DEBUG_INFO_BTF` is also necessary for BTF support, and that `CONFIG_INET` for TCP. `CONFIG_BPF_JIT` which depends on `HAVE_EBPF_JIT`. I'm confident in the following for the `set-results` parameters: - `WorthFuzzing`: `true` (This is definitely worth fuzzing!) - `EnableConfigs`: `["CGROUP_BPF", "BPF_SYSCALL", "BPF_JIT", "DEBUG_INFO_BTF", "INET"]` (These are the necessary configurations.) - `FocusSymbols`: `[ "__tcp_set_rcvlowat", "tcp_set_rcvlowat", "bpf_tcp_ops_set_rcvlowat", "bpf_tcp_ops_cb_flags_set", "bpf_sock_ops_cb_flags_set", "bpf_tcp_ops_rcvlowat_kfunc_filter", "bpf_sol_tcp_setsockopt", "tcp_ofo_queue", "tcp_fastopen_add_skb" ]` (This list targets the core changes). - `Reasoning`: "The patch introduces support for `enqueue_rcvq` and `dequeue_rcvq` callbacks in `bpf_tcp_ops` along with the `BPF_ |
| 26/2 | 2026/09/22 19:40 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"CGROUP_BPF",
"BPF_SYSCALL",
"BPF_JIT",
"DEBUG_INFO_BTF",
"INET"
],
"FocusSymbols": [
"__tcp_set_rcvlowat",
"tcp_set_rcvlowat",
"bpf_tcp_ops_set_rcvlowat",
"bpf_tcp_ops_cb_flags_set",
"bpf_sock_ops_cb_flags_set",
"bpf_tcp_ops_rcvlowat_kfunc_filter",
"bpf_sol_tcp_setsockopt",
"tcp_ofo_queue",
"tcp_fastopen_add_skb"
],
"Reasoning": "The patch introduces support for enqueue_rcvq and dequeue_rcvq callbacks in bpf_tcp_ops along with a new flag BPF_SOCK_OPS_RCVQ_CB_FLAG and the bpf_tcp_ops_set_rcvlowat kfunc. It also refactors tcp_set_rcvlowat into __tcp_set_rcvlowat to conditionally suppress wakeups and hooks BPF operations into core TCP receive processing paths (tcp_queue_rcv, tcp_ofo_queue, tcp_fastopen_add_skb, tcp_cleanup_rbuf). These changes modify reachable core networking and BPF code paths that can be tested in standard virtualized environments.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"CGROUP_BPF",
"BPF_SYSCALL",
"BPF_JIT",
"DEBUG_INFO_BTF",
"INET"
],
"FocusSymbols": [
"__tcp_set_rcvlowat",
"tcp_set_rcvlowat",
"bpf_tcp_ops_set_rcvlowat",
"bpf_tcp_ops_cb_flags_set",
"bpf_sock_ops_cb_flags_set",
"bpf_tcp_ops_rcvlowat_kfunc_filter",
"bpf_sol_tcp_setsockopt",
"tcp_ofo_queue",
"tcp_fastopen_add_skb"
],
"Reasoning": "The patch introduces support for enqueue_rcvq and dequeue_rcvq callbacks in bpf_tcp_ops along with a new flag BPF_SOCK_OPS_RCVQ_CB_FLAG and the bpf_tcp_ops_set_rcvlowat kfunc. It also refactors tcp_set_rcvlowat into __tcp_set_rcvlowat to conditionally suppress wakeups and hooks BPF operations into core TCP receive processing paths (tcp_queue_rcv, tcp_ofo_queue, tcp_fastopen_add_skb, tcp_cleanup_rbuf). These changes modify reachable core networking and BPF code paths that can be tested in standard virtualized environments.",
"WorthFuzzing": true
} |
| 27/1 | 2026/09/22 19:40 | action | If |
4mArgs:{
"WorthFuzzing": true
}Results: null |
| 28/2 | 2026/09/22 19:40 | agent | kmsan-evaluator |
4mModel:core Args: null Results: {
"KMSANReasoning": "The patch introduces TCP receive queue callback operations (enqueue_rcvq and dequeue_rcvq) in bpf_tcp_ops and adds a kfunc (bpf_tcp_ops_set_rcvlowat) along with flags to manage SO_RCVLOWAT dynamically from BPF.\n\n1. No uninitialized memory or info-leak risks: The patch does not introduce any allocations without initialization, nor does it expose kernel structs or buffers to userspace that could contain uninitialized padding or fields.\n2. Logic and bounds: All variables (such as wakeup flags, offsets, and return values) are fully initialized before use. The modifications to `__tcp_set_rcvlowat` adjust socket wake-up conditions and epoll/recvmsg wake thresholds, which do not permit reading uninitialized buffer memory (packet data consumption remains bounded by actual skb lengths).\n3. Sanitizer applicability: Any potential issues introduced by this patch (e.g., locking/context constraints in BPF callbacks, reference handling, or concurrency) fall within the domain of KASAN, LOCKDEP, and standard kernel assertions, not KMSAN.\n\nTherefore, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 00dab524a4f917b8005f7552ada3b92672cdfda5
Author: syz-cluster <triage@syzkaller.com>
Date: Tue Sep 22 19:35:27 2026 +0000
syz-cluster: applied patch under review
diff --git a/include/net/tcp.h b/include/net/tcp.h
index d61ee00052e33..d3cf655da9ecd 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -512,6 +512,7 @@ void tcp_set_keepalive(struct sock *sk, int val);
void tcp_syn_ack_timeout(const struct request_sock *req);
int tcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,
int flags);
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup);
int tcp_set_rcvlowat(struct sock *sk, int val);
void tcp_set_rcvbuf(struct sock *sk, int val);
int tcp_set_window_clamp(struct sock *sk, int val);
@@ -2931,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
return tcp_call_bpf(sk, op, 3, args);
}
+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
+{
+ if (sk_is_mptcp(sk) &&
+ (val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ return -EOPNOTSUPP;
+
+ tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+ return 0;
+}
+
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
tcp_sk(sk)->bpf_sock_ops_cb_flags = 0;
@@ -2953,6 +2964,11 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
return -EPERM;
}
+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
+{
+ return -EOPNOTSUPP;
+}
+
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
}
@@ -3053,6 +3069,12 @@ struct bpf_tcp_ops {
struct request_sock *req, struct sk_buff *syn_skb,
enum tcp_synack_type synack_type,
u32 opt_off);
+
+ /* Called when an incoming skb is enqueued to sk->sk_receive_queue. */
+ void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
+
+ /* Called after data is dequeued from sk->sk_receive_queue. */
+ void (*dequeue_rcvq)(struct sock *sk);
};
#define bpf_tcp_ops_call(op, sk, ...) \
@@ -3144,6 +3166,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
}
+static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)
+{
+ if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ bpf_tcp_ops_call(enqueue_rcvq, sk, skb);
+}
+
+static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)
+{
+ if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ bpf_tcp_ops_call(dequeue_rcvq, sk);
+}
+
#if IS_ENABLED(CONFIG_SMC)
extern struct static_key_false tcp_have_smc;
#endif
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 6330b7d745c57..8cdf22667775f 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -3264,7 +3264,7 @@ union bpf_attr {
* Return
* 0
*
- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
* Description
* Attempt to set the value of the **bpf_sock_ops_cb_flags** field
* for the full TCP socket associated to *bpf_sock_ops* to
@@ -7148,8 +7148,17 @@ enum {
* options first before the BPF program does.
*/
BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
+ /* Call bpf when the TCP stack enqueues/dequeues payload
+ * to/from sk->sk_receive_queue.
+ *
+ * Only bpf_tcp_ops is supported.
+ *
+ * It can be used to adjust sk->sk_rcvlowat and suppress
+ * unnecessary wakeups before sufficient data is available.
+ */
+ BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
/* Mask of all currently supported cb flags */
- BPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,
+ BPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,
};
enum {
diff --git a/net/core/filter.c b/net/core/filter.c
index 5feb998846825..f29c061bb0664 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -5588,8 +5588,7 @@ static int bpf_sol_tcp_setsockopt(struct sock *sk, int optname,
case TCP_BPF_SOCK_OPS_CB_FLAGS:
if (val & ~(BPF_SOCK_OPS_ALL_CB_FLAGS))
return -EINVAL;
- tp->bpf_sock_ops_cb_flags = val;
- break;
+ return tcp_set_sock_ops_cb_flags(sk, val);
default:
return -EINVAL;
}
@@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {
BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
int, argval)
{
- struct sock *sk = bpf_sock->sk;
int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+ struct sock *sk = bpf_sock->sk;
+ int err;
if (!is_locked_tcp_sock_ops(bpf_sock))
return -EOPNOTSUPP;
@@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
if (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))
return -EINVAL;
- tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+ err = tcp_set_sock_ops_cb_flags(sk, val);
+ if (err)
+ return err;
return argval & (~BPF_SOCK_OPS_ALL_CB_FLAGS);
}
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 681fed6429999..3768b1440eb77 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,
{
}
+static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)
+{
+}
+
+static void dequeue_rcvq_stub(struct sock *sk)
+{
+}
+
static struct bpf_tcp_ops __bpf_tcp_ops = {
.timeout_init = timeout_init_stub,
.rwnd_init = rwnd_init_stub,
@@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {
.parse_hdr = parse_hdr_stub,
.hdr_opt_len = hdr_opt_len_stub,
.write_hdr_opt = write_hdr_opt_stub,
+ .enqueue_rcvq = enqueue_rcvq_stub,
+ .dequeue_rcvq = dequeue_rcvq_stub,
};
BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,
@@ -210,6 +220,27 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
.ret_type = RET_INTEGER,
};
+BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
+{
+ int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+ int err;
+
+ err = tcp_set_sock_ops_cb_flags(sk, val);
+ if (err)
+ return err;
+
+ return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
+}
+
+static const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {
+ .func = bpf_tcp_ops_cb_flags_set,
+ .gpl_only = false,
+ .ret_type = RET_INTEGER,
+ .arg1_type = ARG_PTR_TO_BTF_ID,
+ .arg1_btf_id = &btf_sock_ids[BTF_SOCK_TYPE_TCP],
+ .arg2_type = ARG_ANYTHING,
+};
+
static const struct bpf_func_proto *
get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
@@ -255,6 +286,15 @@ get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
if (moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
return &bpf_tcp_ops_store_hdr_opt_proto;
return NULL;
+ case BPF_FUNC_sock_ops_cb_flags_set:
+ if (moff == offsetof(struct bpf_tcp_ops, connect) ||
+ moff == offsetof(struct bpf_tcp_ops, listen) ||
+ moff == offsetof(struct bpf_tcp_ops, active_established) ||
+ moff == offsetof(struct bpf_tcp_ops, passive_established) ||
+ moff == offsetof(struct bpf_tcp_ops, enqueue_rcvq) ||
+ moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ return &bpf_tcp_ops_cb_flags_set_proto;
+ return NULL;
default:
return bpf_base_func_proto(func_id, prog);
}
@@ -319,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
.owner = THIS_MODULE,
};
+__bpf_kfunc_start_defs();
+
+__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
+ const struct bpf_prog_aux *aux)
+{
+ u32 moff = aux->attach_st_ops_member_off;
+ bool wakeup = false;
+
+ if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ wakeup = true;
+
+ if (rcvlowat < 0)
+ rcvlowat = INT_MAX;
+
+ return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
+}
+
+__bpf_kfunc_end_defs();
+
+BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)
+BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)
+BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)
+
+static int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,
+ u32 kfunc_id)
+{
+ u32 moff;
+
+ if (!btf_id_set8_contains(&bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))
+ return 0;
+
+ if (prog->aux->st_ops != &bpf_tcp_ops)
+ return -EACCES;
+
+ moff = prog->aux->attach_st_ops_member_off;
+ if (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) &&
+ moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ return -EACCES;
+
+ return 0;
+}
+
+static const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {
+ .owner = THIS_MODULE,
+ .set = &bpf_tcp_ops_rcvlowat_kfunc_set,
+ .filter = bpf_tcp_ops_rcvlowat_kfunc_filter,
+};
+
static int __init __bpf_tcp_ops_init(void)
{
- return register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+ int ret;
+
+ ret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,
+ &bpf_tcp_ops_rcvlowat_kfunc_id_set);
+ ret = ret ?: register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+
+ return ret;
}
late_initcall(__bpf_tcp_ops_init);
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index a4456b4194122..aa7593fc8334e 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
"cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
__tcp_cleanup_rbuf(sk, copied);
+
+ bpf_tcp_ops_dequeue_rcvq(sk);
}
static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)
@@ -1826,8 +1828,7 @@ int tcp_peek_len(struct socket *sock)
return tcp_inq(sock->sk);
}
-/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
-int tcp_set_rcvlowat(struct sock *sk, int val)
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)
{
struct tcp_sock *tp = tcp_sk(sk);
int space, cap;
@@ -1840,7 +1841,8 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
WRITE_ONCE(sk->sk_rcvlowat, val ? : 1);
/* Check if we need to signal EPOLLIN right now */
- tcp_data_ready(sk);
+ if (wakeup)
+ tcp_data_ready(sk);
if (sk->sk_userlocks & SOCK_RCVBUF_LOCK)
return 0;
@@ -1855,6 +1857,12 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
return 0;
}
+/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
+int tcp_set_rcvlowat(struct sock *sk, int val)
+{
+ return __tcp_set_rcvlowat(sk, val, true);
+}
+
void tcp_set_rcvbuf(struct sock *sk, int val)
{
tcp_set_window_clamp(sk, tcp_win_from_space(sk, val));
diff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c
index 471c78be55137..4939bcbc81d17 100644
--- a/net/ipv4/tcp_fastopen.c
+++ b/net/ipv4/tcp_fastopen.c
@@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)
TCP_SKB_CB(skb)->seq++;
TCP_SKB_CB(skb)->tcp_flags &= ~TCPHDR_SYN;
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
tp->rcv_nxt = TCP_SKB_CB(skb)->end_seq;
tcp_add_receive_queue(sk, skb);
tp->syn_data_acked = 1;
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 6ac6f9d5b6c38..c60c61bb0a715 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)
continue;
}
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
tail = skb_peek_tail(&sk->sk_receive_queue);
eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
@@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
int eaten;
struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
eaten = (tail &&
tcp_try_coalesce(sk, tail,
skb, fragstolen)) ? 1 : 0;
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 6330b7d745c57..8cdf22667775f 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -3264,7 +3264,7 @@ union bpf_attr {
* Return
* 0
*
- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
* Description
* Attempt to set the value of the **bpf_sock_ops_cb_flags** field
* for the full TCP socket associated to *bpf_sock_ops* to
@@ -7148,8 +7148,17 @@ enum {
* options first before the BPF program does.
*/
BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
+ /* Call bpf when the TCP stack enqueues/dequeues payload
+ * to/from sk->sk_receive_queue.
+ *
+ * Only bpf_tcp_ops is supported.
+ *
+ * It can be used to adjust sk->sk_rcvlowat and suppress
+ * unnecessary wakeups before sufficient data is available.
+ */
+ BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
/* Mask of all currently supported cb flags */
- BPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,
+ BPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,
};
enum {
diff --git a/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
new file mode 100644
index 0000000000000..c5ead4af24f85
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
@@ -0,0 +1,350 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include <sys/epoll.h>
+
+#include "test_progs.h"
+#include "cgroup_helpers.h"
+#include "network_helpers.h"
+
+#include "tcp_autolowat.skel.h"
+
+#define SOL_BPF 0xdeadbeef
+#define BPF_TCP_AUTOLOWAT 0x8badf00d
+
+struct rpc_descriptor {
+ u32 header_len;
+ u32 payload_len;
+};
+
+enum rpc_event_type {
+ RPC_EVENT_END,
+ RPC_EVENT_AUTOLOWAT,
+ RPC_EVENT_SEND,
+ RPC_EVENT_RECV,
+ RPC_EVENT_EPOLL,
+ RPC_EVENT_RCVLOWAT,
+};
+
+struct rpc_event {
+ enum rpc_event_type type;
+ union {
+ int len;
+ int nfds;
+ int val;
+ int rcvlowat;
+ };
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+
+struct rpc_test_case {
+ char data[4096];
+ struct rpc_descriptor desc[32];
+ struct rpc_event event[32];
+} rpc_test_cases[] = {
+ {
+ .desc = {
+ { .header_len = 100, .payload_len = 150 },
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 100 + 150},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Two full RPC messages in skb. */
+ {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 3},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Two full RPC messages in skb. */
+ {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 200, .payload_len = 500},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* The first descriptor is partial. */
+ {.type = RPC_EVENT_SEND, .len = 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE},
+ /* The first descriptor is available. */
+ {.type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE - 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* The first header is ready. */
+ {.type = RPC_EVENT_SEND, .len = 100},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* skb has the first payload and 1 byte of the next descriptor. */
+ {.type = RPC_EVENT_SEND, .len = 150 + 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* After reading the first RPC message, SO_RCVLOWAT should be RPC_DESC_SIZE. */
+ {.type = RPC_EVENT_RECV, .len = RPC_DESC_SIZE + 150 + 100},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE},
+ /* The second descriptor is available. */
+ {.type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE - 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 200 + 500},
+ },
+ },
+};
+
+struct tcp_autolowat_test_cb {
+ int saved_netns;
+ union {
+ int fd[4];
+ struct {
+ int server, client, child;
+ int epoll;
+ };
+ };
+};
+
+static void tcp_autolowat_teardown_cb(struct tcp_autolowat_test_cb *cb)
+{
+ int i, err;
+
+ for (i = 0; i < ARRAY_SIZE(cb->fd); i++) {
+ if (cb->fd[i] != -1)
+ close(cb->fd[i]);
+ }
+
+ if (cb->saved_netns != -1) {
+ err = setns(cb->saved_netns, CLONE_NEWNET);
+ ASSERT_OK(err, "restore netns");
+
+ close(cb->saved_netns);
+ }
+}
+
+static int tcp_autolowat_setup_cb(struct tcp_autolowat_test_cb *cb, int family)
+{
+ struct epoll_event ev = {};
+ int err;
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(cb->fd); i++)
+ cb->fd[i] = -1;
+
+ cb->saved_netns = open("/proc/self/ns/net", O_RDONLY);
+ if (!ASSERT_OK_FD(cb->saved_netns, "save netns"))
+ goto err;
+
+ err = unshare(CLONE_NEWNET);
+ if (!ASSERT_OK(err, "unshare"))
+ goto err;
+
+ err = system("ip link set dev lo up");
+ if (!ASSERT_OK(err, "set up lo"))
+ goto err;
+
+ cb->server = start_server(family, SOCK_STREAM, NULL, 0, 0);
+ if (!ASSERT_OK_FD(cb->server, "start_server"))
+ goto err;
+
+ cb->client = connect_to_fd(cb->server, 0);
+ if (!ASSERT_OK_FD(cb->client, "connect_to_fd"))
+ goto err;
+
+ cb->child = accept(cb->server, NULL, NULL);
+ if (!ASSERT_OK_FD(cb->child, "accept"))
+ goto err;
+
+ cb->epoll = epoll_create1(0);
+ if (!ASSERT_OK_FD(cb->epoll, "epoll_create"))
+ goto err;
+
+ ev.events = EPOLLIN;
+ ev.data.fd = cb->child;
+
+ err = epoll_ctl(cb->epoll, EPOLL_CTL_ADD, cb->child, &ev);
+ if (!ASSERT_OK(err, "epoll_ctl"))
+ goto err;
+
+ return 0;
+
+err:
+ tcp_autolowat_teardown_cb(cb);
+ return -1;
+}
+
+static int tcp_autolowat_build_data(struct rpc_test_case *test_case)
+{
+ struct rpc_descriptor *desc = test_case->desc;
+ char *ptr = test_case->data;
+ int rpc_size;
+
+ memset(ptr, 0, sizeof(test_case->data));
+
+ while (desc->header_len + desc->payload_len) {
+ rpc_size = sizeof(*desc) + desc->header_len + desc->payload_len;
+
+ if (!ASSERT_LE(ptr + rpc_size - test_case->data,
+ sizeof(test_case->data), "data overflow"))
+ return 1;
+
+ memcpy(ptr, desc, sizeof(*desc));
+ ptr += rpc_size;
+ desc++;
+ }
+
+ if (!ASSERT_GT(ptr - test_case->data, 0, "no data"))
+ return 1;
+
+ return 0;
+}
+
+static void tcp_autolowat_run_rpc_test(struct tcp_autolowat_test_cb *cb,
+ struct rpc_test_case *test_case)
+{
+ struct rpc_event *event = test_case->event;
+ char *ptr = test_case->data;
+ struct epoll_event ev;
+ socklen_t optlen;
+ int err, optval;
+ char buf[4096];
+
+ if (tcp_autolowat_build_data(test_case))
+ return;
+
+ while (1) {
+ switch (event->type) {
+ case RPC_EVENT_END:
+ return;
+ case RPC_EVENT_AUTOLOWAT:
+ err = setsockopt(cb->child, SOL_BPF, BPF_TCP_AUTOLOWAT,
+ &event->val, sizeof(event->val));
+ if (!ASSERT_OK(err, "setsockopt"))
+ return;
+ break;
+ case RPC_EVENT_SEND:
+ err = send(cb->client, ptr, event->len, 0);
+ if (!ASSERT_EQ(err, event->len, "send"))
+ return;
+
+ ptr += event->len;
+ break;
+ case RPC_EVENT_RECV:
+ err = recv(cb->child, buf, event->len, 0);
+ if (!ASSERT_EQ(err, event->len, "recv"))
+ return;
+ break;
+ case RPC_EVENT_EPOLL:
+ err = epoll_wait(cb->epoll, &ev, 1, 100);
+ if (!ASSERT_EQ(err, event->nfds, "epoll_wait"))
+ return;
+ break;
+ case RPC_EVENT_RCVLOWAT:
+ optval = 0;
+ optlen = sizeof(optval);
+
+ err = getsockopt(cb->child, SOL_SOCKET, SO_RCVLOWAT, &optval, &optlen);
+ if (!ASSERT_OK(err, "getsockopt") ||
+ !ASSERT_EQ(optval, event->rcvlowat, "rcvlowat"))
+ return;
+ break;
+ }
+
+ event++;
+ }
+}
+
+static void tcp_autolowat_run_rpc_tests(struct tcp_autolowat *skel, int family)
+{
+ struct tcp_autolowat_test_cb cb;
+ int err;
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(rpc_test_cases); i++) {
+ memset(skel->bss->test_name, 0, sizeof(skel->bss->test_name));
+
+ snprintf(skel->bss->test_name, sizeof(skel->bss->test_name),
+ "AF_INET%c rpc_test_cases[%d]",
+ family == AF_INET ? ' ' : '6', i);
+
+ if (!test__start_subtest(skel->bss->test_name))
+ continue;
+
+ err = tcp_autolowat_setup_cb(&cb, family);
+ if (err)
+ continue;
+
+ tcp_autolowat_run_rpc_test(&cb, &rpc_test_cases[i]);
+ tcp_autolowat_teardown_cb(&cb);
+ }
+}
+
+static void tcp_autolowat_run_tests(struct tcp_autolowat *skel)
+{
+ tcp_autolowat_run_rpc_tests(skel, AF_INET);
+ tcp_autolowat_run_rpc_tests(skel, AF_INET6);
+}
+
+void test_tcp_autolowat(void)
+{
+ struct tcp_autolowat *skel;
+ struct bpf_link *link[2];
+ int cgroup;
+
+ skel = tcp_autolowat__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "open_and_load"))
+ return;
+
+ cgroup = test__join_cgroup("/tcp_autolowat");
+ if (!ASSERT_GE(cgroup, 0, "join_cgroup"))
+ goto destroy_skel;
+
+ link[0] = bpf_map__attach_cgroup_opts(skel->maps.tcp_autolowat_ops, cgroup, NULL);
+ if (!ASSERT_OK_PTR(link[0], "attach_cgroup(tcp_autolowat_ops)"))
+ goto close_cgroup;
+
+ link[1] = bpf_program__attach_cgroup(skel->progs.tcp_autolowat_setsockopt, cgroup);
+ if (!ASSERT_OK_PTR(link[1], "attach_cgroup(SETSOCKOPT)"))
+ goto destroy_sockops;
+
+ tcp_autolowat_run_tests(skel);
+
+ bpf_link__destroy(link[1]);
+destroy_sockops:
+ bpf_link__destroy(link[0]);
+close_cgroup:
+ close(cgroup);
+destroy_skel:
+ tcp_autolowat__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
index 7e8fe1bad03f5..e4849d2a2956f 100644
--- a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
+++ b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
@@ -26,7 +26,8 @@ static void verify_result(struct tcpbpf_globals *result)
ASSERT_EQ(result->bytes_acked, 1002, "bytes_acked");
ASSERT_EQ(result->data_segs_in, 1, "data_segs_in");
ASSERT_EQ(result->data_segs_out, 1, "data_segs_out");
- ASSERT_EQ(result->bad_cb_test_rv, 0x80, "bad_cb_test_rv");
+ ASSERT_EQ(result->bad_cb_test_rv, BPF_SOCK_OPS_ALL_CB_FLAGS + 1,
+ "bad_cb_test_rv");
ASSERT_EQ(result->good_cb_test_rv, 0, "good_cb_test_rv");
ASSERT_EQ(result->num_listen, 1, "num_listen");
diff --git a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
index 593b38f904174..4c999d59cbbce 100644
--- a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
+++ b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
@@ -79,6 +79,8 @@
#define NEXTHDR_TCP 6
+#define TCPHDR_FIN 0x01
+
#define TCPOPT_NOP 1
#define TCPOPT_EOL 0
#define TCPOPT_MSS 2
diff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
new file mode 100644
index 0000000000000..bb96e19e7589d
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
@@ -0,0 +1,312 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include "vmlinux.h"
+
+#include <string.h>
+#include <limits.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+
+#include "bpf_kfuncs.h"
+#include "bpf_tracing_net.h"
+
+#define SOL_BPF 0xdeadbeef
+#define BPF_TCP_AUTOLOWAT 0x8badf00d
+
+//#define DEBUG /* For verbose output. */
+
+struct rpc_descriptor {
+ u32 header_len;
+ u32 payload_len;
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+#define MAX_RPC_DESC_PER_SKB 100
+
+struct tcp_autolowat_cb {
+ /* Don't put this field at the end; BPF verifier complains. */
+ char rpc_desc_buf[RPC_DESC_SIZE];
+ u32 rpc_desc_seq;
+ u32 rpc_end_seq;
+#ifdef DEBUG
+ u32 isn;
+#endif
+ u8 rpc_desc_buff_len;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_SK_STORAGE);
+ __uint(map_flags, BPF_F_NO_PREALLOC);
+ __type(key, int);
+ __type(value, struct tcp_autolowat_cb);
+} tcp_autolowat_map SEC(".maps");
+
+char test_name[64];
+
+#ifdef DEBUG
+#define LOG(str, ...) \
+ bpf_printk("%s: " str, test_name, ##__VA_ARGS__)
+#else
+#define LOG(...)
+#endif
+
+#define SEQ(val) \
+ (val - cb->isn)
+#define TP_SEQ(field) \
+ (tp->field - cb->isn)
+#define CB_SEQ(field) \
+ (cb->field - cb->isn)
+
+static int tcp_parse_descriptor(struct tcp_autolowat_cb *cb,
+ struct bpf_dynptr *dptr,
+ u32 seq, u32 end_seq)
+{
+ struct rpc_descriptor *rpc_desc;
+ u32 rpc_copied_seq;
+ u64 copy_len; /* u32 should work, but not for no_alu32 :/ */
+ u64 rpc_len;
+ int err;
+
+ rpc_copied_seq = cb->rpc_desc_seq + cb->rpc_desc_buff_len;
+
+ if (before(cb->rpc_desc_seq + RPC_DESC_SIZE, end_seq))
+ copy_len = RPC_DESC_SIZE - cb->rpc_desc_buff_len;
+ else
+ copy_len = end_seq - rpc_copied_seq;
+
+ if (copy_len == 0)
+ goto disable; /* FIN. */
+ if (copy_len > RPC_DESC_SIZE)
+ goto disable; /* always false, only for verifier. */
+ if (cb->rpc_desc_buf + cb->rpc_desc_buff_len >= &cb->rpc_desc_buf[RPC_DESC_SIZE])
+ goto disable; /* always false, only for verifier. */
+
+ err = bpf_dynptr_read(cb->rpc_desc_buf + cb->rpc_desc_buff_len,
+ copy_len, dptr, rpc_copied_seq - seq, 0);
+ if (err)
+ goto disable;
+
+ cb->rpc_desc_buff_len += copy_len;
+
+ if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+ LOG("Copied %d bytes: rpc_desc_buff_len: %u", copy_len, cb->rpc_desc_buff_len);
+ goto partial;
+ }
+
+ rpc_desc = (struct rpc_descriptor *)cb->rpc_desc_buf;
+ rpc_len = RPC_DESC_SIZE + rpc_desc->header_len + rpc_desc->payload_len;
+
+ if (rpc_len > INT_MAX)
+ goto disable;
+
+ cb->rpc_end_seq = cb->rpc_desc_seq + rpc_len;
+
+ LOG("Copied full descriptor: rpc_desc_seq: %u, rpc_end_seq: %u, header_len: %u, payload_len: %u",
+ CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq),
+ rpc_desc->header_len, rpc_desc->payload_len);
+
+ return 0;
+disable:
+ return -1;
+partial:
+ return 1;
+}
+
+static void tcp_set_autolowat(struct tcp_autolowat_cb *cb,
+ struct sock *sk)
+{
+ struct tcp_sock *tp = (struct tcp_sock *)sk;
+ u32 val; /* To handle wraparound. */
+
+ LOG("Setting rcvlowat: tp->copied_seq: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_desc_buff_len: %u",
+ TP_SEQ(copied_seq), CB_SEQ(rpc_desc_seq),
+ CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+ if (before(tp->copied_seq, cb->rpc_desc_seq))
+ val = cb->rpc_desc_seq - tp->copied_seq;
+ else if (cb->rpc_desc_buff_len != RPC_DESC_SIZE)
+ val = RPC_DESC_SIZE;
+ else
+ val = cb->rpc_end_seq - tp->copied_seq;
+
+ if (val != tp->inet_conn.icsk_inet.sk.sk_rcvlowat) {
+ bpf_tcp_ops_set_rcvlowat(sk, val);
+
+ LOG("Set rcvlowat: expected: %u, actual: %d\n",
+ val, tp->inet_conn.icsk_inet.sk.sk_rcvlowat);
+ } else {
+ LOG("No need to set rcvlowat: %u\n", val);
+ }
+}
+
+static void tcp_disable_autolowat(struct sock *sk)
+{
+ struct tcp_sock *tp = (struct tcp_sock *)sk;
+ int flags;
+
+ flags = tp->bpf_sock_ops_cb_flags & ~BPF_SOCK_OPS_RCVQ_CB_FLAG;
+ bpf_sock_ops_cb_flags_set(sk, flags);
+
+ bpf_tcp_ops_set_rcvlowat(sk, 1);
+
+ LOG("Disabled autolowat");
+}
+
+static void tcp_do_autolowat(struct tcp_autolowat_cb *cb,
+ struct sock *sk, struct sk_buff *skb)
+{
+ struct bpf_dynptr dptr;
+ struct tcp_skb_cb *tcb;
+ u32 seq, end_seq;
+ int ret = 0, i;
+
+ if (bpf_dynptr_from_skb((struct __sk_buff *)skb, 0, &dptr)) {
+ ret = -1;
+ goto update;
+ }
+
+ tcb = bpf_core_cast(skb->cb, struct tcp_skb_cb);
+ seq = tcb->seq;
+ end_seq = tcb->end_seq - !!(tcb->tcp_flags & TCPHDR_FIN);
+
+ LOG("Start parsing skb: seq: %u, end_seq: %u, len: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_buff_len: %u",
+ SEQ(seq), SEQ(end_seq), end_seq - seq,
+ CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+ if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+ ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+ if (ret)
+ goto update;
+ }
+
+ i = 0;
+
+ while (1) {
+ if (i++ > MAX_RPC_DESC_PER_SKB) {
+ ret = -1;
+ break;
+ }
+
+ if (after(cb->rpc_end_seq, end_seq)) {
+ LOG("No more descriptor: rpc_end_seq: %u, end_seq: %u",
+ CB_SEQ(rpc_end_seq), SEQ(end_seq));
+ break;
+ }
+
+ cb->rpc_desc_seq = cb->rpc_end_seq;
+ cb->rpc_desc_buff_len = 0;
+
+ if (cb->rpc_end_seq == end_seq)
+ break;
+
+ LOG("Found next descriptor: rpc_end_seq: %u, end_seq: %u, len: %u",
+ CB_SEQ(rpc_end_seq), SEQ(end_seq), end_seq - cb->rpc_end_seq);
+
+ ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+ if (ret)
+ break;
+ }
+
+update:
+ if (ret >= 0)
+ tcp_set_autolowat(cb, sk);
+ else
+ tcp_disable_autolowat(sk);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_enqueue_rcvq, struct sock *sk, struct sk_buff *skb)
+{
+ struct tcp_autolowat_cb *cb;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+ if (!cb)
+ return;
+
+ tcp_do_autolowat(cb, sk, skb);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_dequeue_rcvq, struct sock *sk)
+{
+ struct tcp_autolowat_cb *cb;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+ if (!cb)
+ return;
+
+ tcp_set_autolowat(cb, sk);
+}
+
+SEC(".struct_ops.link")
+struct bpf_tcp_ops tcp_autolowat_ops = {
+ .enqueue_rcvq = (void *)tcp_autolowat_enqueue_rcvq,
+ .dequeue_rcvq = (void *)tcp_autolowat_dequeue_rcvq,
+};
+
+static int tcp_init_autolowat_cb(struct bpf_sockopt *sockopt,
+ struct bpf_tcp_sock *btp)
+{
+ struct tcp_autolowat_cb *cb;
+ struct tcp_sock *tp;
+ int flags;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, btp, 0,
+ BPF_SK_STORAGE_GET_F_CREATE);
+ if (!cb)
+ return -1;
+
+ tp = bpf_core_cast(btp, struct tcp_sock);
+ if (!tp)
+ return -1;
+
+ cb->rpc_desc_seq = tp->copied_seq;
+ cb->rpc_end_seq = tp->copied_seq;
+#ifdef DEBUG
+ cb->isn = tp->copied_seq;
+#endif
+
+ if (bpf_getsockopt(sockopt->sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
+ &flags, sizeof(flags)))
+ return -1;
+
+ flags |= BPF_SOCK_OPS_RCVQ_CB_FLAG;
+
+ if (bpf_setsockopt(sockopt->sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
+ &flags, sizeof(flags)))
+ return -1;
+
+ return 0;
+}
+
+SEC("cgroup/setsockopt")
+int tcp_autolowat_setsockopt(struct bpf_sockopt *ctx)
+{
+ void *optval_end = ctx->optval_end;
+ int *optval = ctx->optval;
+ struct bpf_tcp_sock *btp;
+
+ if (ctx->level != SOL_BPF || ctx->optname != BPF_TCP_AUTOLOWAT)
+ goto out;
+
+ if (optval + 1 > optval_end)
+ return 0; /* -EPERM */
+
+ btp = bpf_tcp_sock(ctx->sk);
+ if (!btp)
+ goto out;
+
+ if (*optval && tcp_init_autolowat_cb(ctx, btp))
+ return 0; /* -EPERM */
+
+ /*
+ * BPF has consumed this option, don't call kernel
+ * setsockopt handler.
+ */
+ ctx->optlen = -1;
+out:
+ return 1;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
index 6935f32eeb8fd..e30cb1fab0790 100644
--- a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
+++ b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
@@ -92,7 +92,8 @@ int bpf_testcb(struct bpf_sock_ops *skops)
break;
case BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB:
/* Test failure to set largest cb flag (assumes not defined) */
- global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops, 0x80);
+ global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
+ BPF_SOCK_OPS_ALL_CB_FLAGS + 1);
/* Set callback */
global.good_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
BPF_SOCK_OPS_STATE_CB_FLAG);
|
| 29/3 | 2026/09/22 19:40 | llm | kmsan-evaluator |
3mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 16728 output: 35 thoughts: 1877 Thoughts: **KMSAN Fuzzing Assessment for This Patch** Okay, let's break this down. My goal is to determine if this patch necessitates a dedicated KernelMemorySanitizer (KMSAN) fuzzing session. The instructions are clear: if standard KASAN, LOCKDEP, or other debug tools can catch the potential issue, we don't need KMSAN. KMSAN is reserved for scenarios involving uninitialized memory reads, particularly those that leak information to user space. I've carefully reviewed the patch. It seems to revolve around adding BPF (Berkeley Packet Filter) callback functionality for TCP sockets, specifically around the receive queue (`enqueue_rcvq`, `dequeue_rcvq`) and a new `BPF_SOCK_OPS_RCVQ_CB_FLAG`. The patch modifies structures, adds function pointers, introduces helper functions to call these function pointers, and integrates them into existing TCP data flow. I also see new functionality to configure the receive low water mark. Now, let's apply the KMSAN criteria: 1. **Kernel Structures to User Space:** Are there kernel structures being sent to user space where uninitialized fields might be present? Not directly. This patch adds function pointers to `struct bpf_tcp_ops` and associates the related code with socket operations. User space interacts with these through BPF, and there is no direct copy of kernel data to user space. It *is* possible that user-provided BPF programs could interact with kernel data structures in ways that expose uninitialized memory, but this patch doesn't *directly* introduce that vector. The use of BPF itself is a significant point of concern. 2. **Conditional Logic/Branching on Uninitialized Variables:** The patch introduces conditional logic based on a flag (`BPF_SOCK_OPS_RCVQ_CB_FLAG`). The flag determines whether or not to call these BPF callbacks. If that flag is not set, there are no calls to the functions. The use of the `BPF_SOCK_OPS_TEST_FLAG` macro should mitigate any immediate uninitialized memory issues, but the control flow *does* depend on the flag's initialization and state. 3. **Complex Data Structures and Initialization:** The key new data structure is the `struct bpf_tcp_ops`. This structure is extended with the callback function pointers. The stubs provided suggests the values can be set through BPF. Careful initialization and management of this structure are crucial for this patch. While the patch itself adds the structure members, BPF allows user-space programs to write and manage these structures. This introduces potential for uninitialized fields to be read. 4. **Bounds Checks/Logic Leading to Uninitialized Byte Access:** The patch doesn't appear to directly introduce new bounds checks. There is the kfunc `bpf_tcp_ops_set_rcvlowat`. If the values are managed incorrectly, it is possible for code to access uninitialized bytes, but not immediately apparent from the patch. **Conclusion:** Based on the review, I lean toward **requiring** a KMSAN fuzzing session, but with the caveat that it should be targeted. While the patch itself doesn't scream KMSAN, the integration with BPF, which allows user-space programs to access and manipulate data structures, introduces a much higher risk of uninitialized memory reads. The new function pointers and callbacks could create a pathway for information leaks if the BPF program is not carefully designed. So, the potential is there. Since this changes core TCP structures to support BPF, KMSAN would be wise here. I recommend a focused KMSAN session that concentrates on the interaction between BPF and the receive queue, and testing variations with the new `BPF_SOCK_OPS_RCVQ_CB_FLAG`. |
| 30/3 | 2026/09/22 19:43 | tool | grepper |
0mArgs:{
"Expression": "bpf_tcp_ops_call",
"PathPrefix": "include/net/tcp.h"
}Results: {
"Output": "include/net/tcp.h=2980=struct bpf_tcp_ops {\n--\ninclude/net/tcp.h-3079-\ninclude/net/tcp.h:3080:#define bpf_tcp_ops_call(op, sk, ...)\t\t\t\t\t\\\ninclude/net/tcp.h-3081-do {\t\t\t\t\t\t\t\t\t\\\n--\ninclude/net/tcp.h-3097-\ninclude/net/tcp.h:3098:#define bpf_tcp_ops_call_int(op, init_retval, sk, ...)\t\t\t\\\ninclude/net/tcp.h-3099-({\t\t\t\t\t\t\t\t\t\\\n--\ninclude/net/tcp.h-3130-#else\ninclude/net/tcp.h:3131:#define bpf_tcp_ops_call(op, sk, ...)\t\tdo { } while (0)\ninclude/net/tcp.h:3132:#define bpf_tcp_ops_call_int(op, init_retval, sk, ...)\t(init_retval)\ninclude/net/tcp.h-3133-#endif\n--\ninclude/net/tcp.h=3135=static inline u32 tcp_timeout_init(struct sock *sk)\n--\ninclude/net/tcp.h-3139-\ttimeout = tcp_call_bpf(sk, BPF_SOCK_OPS_TIMEOUT_INIT, 0, NULL);\ninclude/net/tcp.h:3140:\ttimeout = bpf_tcp_ops_call_int(timeout_init, timeout, sk);\ninclude/net/tcp.h-3141-\tif (timeout \u003c= 0)\n--\ninclude/net/tcp.h=3146=static inline u32 tcp_rwnd_init_bpf(struct sock *sk)\n--\ninclude/net/tcp.h-3150-\trwnd = tcp_call_bpf(sk, BPF_SOCK_OPS_RWND_INIT, 0, NULL);\ninclude/net/tcp.h:3151:\trwnd = bpf_tcp_ops_call_int(rwnd_init, rwnd, sk);\ninclude/net/tcp.h-3152-\tif (rwnd \u003c 0)\n--\ninclude/net/tcp.h=3162=static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)\n--\ninclude/net/tcp.h-3165-\t\ttcp_call_bpf_2arg(sk, BPF_SOCK_OPS_RTT_CB, mrtt, srtt);\ninclude/net/tcp.h:3166:\tbpf_tcp_ops_call(rtt, sk, mrtt, srtt);\ninclude/net/tcp.h-3167-}\n--\ninclude/net/tcp.h=3169=static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)\n--\ninclude/net/tcp.h-3171-\tif (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))\ninclude/net/tcp.h:3172:\t\tbpf_tcp_ops_call(enqueue_rcvq, sk, skb);\ninclude/net/tcp.h-3173-}\n--\ninclude/net/tcp.h=3175=static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)\n--\ninclude/net/tcp.h-3177-\tif (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))\ninclude/net/tcp.h:3178:\t\tbpf_tcp_ops_call(dequeue_rcvq, sk);\ninclude/net/tcp.h-3179-}\n"
} |
| 31/3 | 2026/09/22 19:43 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 20399 output: 36 thoughts: 20 Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience: **Diving into the TCP Header** Okay, time to get to the core of this issue. Let's see... the TCP header definition. I'm going to jump straight into `include/net/tcp.h` - I suspect the problem lies somewhere in how those structs are laid out or how the flags are handled. My gut tells me to focus around line 3080. That's usually where the really low-level stuff, the critical bit-twiddling and flag definitions, tends to cluster. This feels like it has something to do with the segment handling or window management, and that's usually right in the weeds of the header definitions. Let's see what's there... |
| 32/3 | 2026/09/22 19:43 | tool | read-file |
0mArgs:{
"File": "include/net/tcp.h",
"FirstLine": 3075,
"LineCount": 40
}Results: {
"Contents": "3075:\t\n3076:\t\t/* Called after data is dequeued from sk-\u003esk_receive_queue. */\n3077:\t\tvoid (*dequeue_rcvq)(struct sock *sk);\n3078:\t};\n3079:\t\n3080:\t#define bpf_tcp_ops_call(op, sk, ...)\t\t\t\t\t\\\n3081:\tdo {\t\t\t\t\t\t\t\t\t\\\n3082:\t\tif (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) {\t\t\t\\\n3083:\t\t\tconst struct bpf_prog_array_item *item;\t\t\t\\\n3084:\t\t\tconst struct bpf_tcp_ops *tcp_ops;\t\t\t\\\n3085:\t\t\tstruct cgroup *cgrp;\t\t\t\t\t\\\n3086:\t\t\t\t\t\t\t\t\t\t\\\n3087:\t\t\tcgrp = sock_cgroup_ptr(\u0026sk-\u003esk_cgrp_data);\t\t\\\n3088:\t\t\trcu_read_lock_dont_migrate();\t\t\t\t\\\n3089:\t\t\tbpf_cgroup_struct_ops_foreach(tcp_ops, item, cgrp,\t\\\n3090:\t\t\t\t\t\t CGROUP_TCP_SOCK_OPS) {\t\\\n3091:\t\t\t\tif (tcp_ops-\u003eop)\t\t\t\t\\\n3092:\t\t\t\t\ttcp_ops-\u003eop(sk, ##__VA_ARGS__);\t\t\\\n3093:\t\t\t}\t\t\t\t\t\t\t\\\n3094:\t\t\trcu_read_unlock_migrate();\t\t\t\t\\\n3095:\t\t}\t\t\t\t\t\t\t\t\\\n3096:\t} while (0)\n3097:\t\n3098:\t#define bpf_tcp_ops_call_int(op, init_retval, sk, ...)\t\t\t\\\n3099:\t({\t\t\t\t\t\t\t\t\t\\\n3100:\t\tint __retval = (init_retval);\t\t\t\t\t\\\n3101:\t\tif (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) {\t\t\t\\\n3102:\t\t\tconst struct bpf_prog_array_item *item;\t\t\t\\\n3103:\t\t\tconst struct bpf_tcp_ops *tcp_ops;\t\t\t\\\n3104:\t\t\tstruct bpf_tramp_run_ctx run_ctx;\t\t\t\\\n3105:\t\t\tstruct bpf_run_ctx *old_run_ctx;\t\t\t\\\n3106:\t\t\tstruct sock *__sk = sk_to_full_sk(sk); \\\n3107:\t\t\tstruct request_sock *req = NULL;\t\t\t\\\n3108:\t\t\tstruct cgroup *cgrp;\t\t\t\t\t\\\n3109:\t\t\t\t\t\t\t\t\t\t\\\n3110:\t\t\tif (__sk) {\t\t\t\t\t\t\\\n3111:\t\t\t\trun_ctx.retval = (init_retval);\t\t\t\\\n3112:\t\t\t\tcgrp = sock_cgroup_ptr(\u0026__sk-\u003esk_cgrp_data);\t\\\n3113:\t\t\t\tif (!sk_fullsock(sk))\t\t\t\t\\\n3114:\t\t\t\t\treq = (struct request_sock *)sk;\t\\\n"
} |
| 33/3 | 2026/09/22 19:43 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 21377 output: 41 thoughts: 28 Thoughts: Here's a summary of my current thought process, as requested: **Diving Deep into TCP BPF Operations** Okay, so I'm digging into `net/ipv4/bpf_tcp_ops.c` right now, specifically focusing on how BPF interacts with TCP operations. I need to refresh my memory on the intricacies of the BTF data structures and how `struct ops` are defined and utilized in this context. It's crucial to understand how the kernel exposes TCP-related data for BPF programs to analyze and potentially manipulate. I know that I need to look closely at the mappings between the data structures and the BPF program's access to them. The finer points of how these are implemented is key for debugging or extending functionality. The details are important here. |
| 34/3 | 2026/09/22 19:43 | tool | read-file |
0mArgs:{
"File": "net/ipv4/bpf_tcp_ops.c",
"FirstLine": 215,
"LineCount": 60
}Results: {
"Contents": " 215:\t}\n 216:\t\n 217:\tconst struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {\n 218:\t\t.func\t\t= bpf_tcp_ops_get_retval,\n 219:\t\t.gpl_only\t= false,\n 220:\t\t.ret_type\t= RET_INTEGER,\n 221:\t};\n 222:\t\n 223:\tBPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)\n 224:\t{\n 225:\t\tint val = argval \u0026 BPF_SOCK_OPS_ALL_CB_FLAGS;\n 226:\t\tint err;\n 227:\t\n 228:\t\terr = tcp_set_sock_ops_cb_flags(sk, val);\n 229:\t\tif (err)\n 230:\t\t\treturn err;\n 231:\t\n 232:\t\treturn argval \u0026 ~BPF_SOCK_OPS_ALL_CB_FLAGS;\n 233:\t}\n 234:\t\n 235:\tstatic const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {\n 236:\t\t.func\t\t= bpf_tcp_ops_cb_flags_set,\n 237:\t\t.gpl_only\t= false,\n 238:\t\t.ret_type\t= RET_INTEGER,\n 239:\t\t.arg1_type\t= ARG_PTR_TO_BTF_ID,\n 240:\t\t.arg1_btf_id\t= \u0026btf_sock_ids[BTF_SOCK_TYPE_TCP],\n 241:\t\t.arg2_type\t= ARG_ANYTHING,\n 242:\t};\n 243:\t\n 244:\tstatic const struct bpf_func_proto *\n 245:\tget_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)\n 246:\t{\n 247:\t\tu32 moff = prog-\u003eaux-\u003eattach_st_ops_member_off;\n 248:\t\n 249:\t\tswitch (func_id) {\n 250:\t\tcase BPF_FUNC_sk_storage_get:\n 251:\t\t\treturn \u0026bpf_sk_storage_get_proto;\n 252:\t\tcase BPF_FUNC_sk_storage_delete:\n 253:\t\t\treturn \u0026bpf_sk_storage_delete_proto;\n 254:\t\tcase BPF_FUNC_setsockopt:\n 255:\t\t\t/* The sk may be an unlocked listener (synack path) or NULL\n 256:\t\t\t * fullsock; disable for members that can run unlocked.\n 257:\t\t\t */\n 258:\t\t\tif (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||\n 259:\t\t\t moff == offsetof(struct bpf_tcp_ops, timeout_init) ||\n 260:\t\t\t moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||\n 261:\t\t\t moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))\n 262:\t\t\t\treturn NULL;\n 263:\t\t\treturn \u0026bpf_sk_setsockopt_proto;\n 264:\t\tcase BPF_FUNC_getsockopt:\n 265:\t\t\tif (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||\n 266:\t\t\t moff == offsetof(struct bpf_tcp_ops, timeout_init) ||\n 267:\t\t\t moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||\n 268:\t\t\t moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))\n 269:\t\t\t\treturn NULL;\n 270:\t\t\treturn \u0026bpf_sk_getsockopt_proto;\n 271:\t\tcase BPF_FUNC_get_retval:\n 272:\t\t\tif (moff == offsetof(struct bpf_tcp_ops, timeout_init) ||\n 273:\t\t\t moff == offsetof(struct bpf_tcp_ops, rwnd_init))\n 274:\t\t\t\treturn \u0026bpf_tcp_ops_get_retval_proto;\n"
} |
| 35/3 | 2026/09/22 19:43 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 22748 output: 41 thoughts: 31 Thoughts: Okay, here's my interpretation of what's going through my mind, as an expert, thinking about that specific code section: **Diving into `bpf_tcp_ops.c`** Alright, I'm thinking, let's hone in on lines 350 to 420 in `net/ipv4/bpf_tcp_ops.c`. My mind immediately jumps to what that section likely covers, given the naming convention and my experience with the kernel's networking stack. It's almost certainly related to BPF hooks for TCP operations within IPv4. I'm guessing this region probably deals with either attaching BPF programs to key TCP events or perhaps specific processing triggered by those BPF programs. I'm internally listing some of the possibilities: * **Filter criteria:** Is it likely that the code in that region focuses on how BPF programs can filter TCP packets or events based on source and destination addresses, ports, or flags? * **Action definition:** I need to consider whether it's related to defining the actions BPF programs can take: dropping packets, modifying packet headers, redirecting traffic, or potentially generating custom statistics or logging events. * **Function hooks:** I should be thinking about the kernel functions being hooked. This is where the real complexity comes in - which specific TCP state transitions or packet processing paths are being tapped into? Are we talking about `tcp_rcv_established`, `tcp_sendmsg`, `tcp_disconnect`, or other critical entry points? I'm also wondering about the data structures being used. Is it operating on `skb` (socket buffer) directly or, more likely, accessing some helper functions to access data within the socket buffer? I know how critical performance is here, so I'm expecting to see efficient memory access patterns and minimal overhead. I'll be looking for any indications of potential contention or lock issues, knowing the impact of those on network performance. The whole point of BPF is to run fast, so I'll be looking for any signs of suboptimal code. I'm also ready to see any tricks used to make the BPF programs secure and prevent crashes. |
| 36/3 | 2026/09/22 19:44 | tool | read-file |
0mArgs:{
"File": "net/ipv4/bpf_tcp_ops.c",
"FirstLine": 350,
"LineCount": 70
}Results: {
"Contents": " 350:\tstatic struct bpf_struct_ops bpf_tcp_ops = {\n 351:\t\t.verifier_ops = \u0026bpf_tcp_ops_verifier,\n 352:\t\t.init_member = bpf_tcp_ops_init_member,\n 353:\t\t.check_member = bpf_tcp_ops_check_member,\n 354:\t\t.init = bpf_tcp_ops_init,\n 355:\t\t.validate = bpf_tcp_ops_validate,\n 356:\t\t.name = \"bpf_tcp_ops\",\n 357:\t\t.cgroup_atype = CGROUP_TCP_SOCK_OPS,\n 358:\t\t.cfi_stubs = \u0026__bpf_tcp_ops,\n 359:\t\t.owner = THIS_MODULE,\n 360:\t};\n 361:\t\n 362:\t__bpf_kfunc_start_defs();\n 363:\t\n 364:\t__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,\n 365:\t\t\t\t\t\t const struct bpf_prog_aux *aux)\n 366:\t{\n 367:\t\tu32 moff = aux-\u003eattach_st_ops_member_off;\n 368:\t\tbool wakeup = false;\n 369:\t\n 370:\t\tif (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))\n 371:\t\t\twakeup = true;\n 372:\t\n 373:\t\tif (rcvlowat \u003c 0)\n 374:\t\t\trcvlowat = INT_MAX;\n 375:\t\n 376:\t\treturn __tcp_set_rcvlowat(sk, rcvlowat, wakeup);\n 377:\t}\n 378:\t\n 379:\t__bpf_kfunc_end_defs();\n 380:\t\n 381:\tBTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)\n 382:\tBTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)\n 383:\tBTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)\n 384:\t\n 385:\tstatic int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,\n 386:\t\t\t\t\t\t u32 kfunc_id)\n 387:\t{\n 388:\t\tu32 moff;\n 389:\t\n 390:\t\tif (!btf_id_set8_contains(\u0026bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))\n 391:\t\t\treturn 0;\n 392:\t\n 393:\t\tif (prog-\u003eaux-\u003est_ops != \u0026bpf_tcp_ops)\n 394:\t\t\treturn -EACCES;\n 395:\t\n 396:\t\tmoff = prog-\u003eaux-\u003eattach_st_ops_member_off;\n 397:\t\tif (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) \u0026\u0026\n 398:\t\t moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))\n 399:\t\t\treturn -EACCES;\n 400:\t\n 401:\t\treturn 0;\n 402:\t}\n 403:\t\n 404:\tstatic const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {\n 405:\t\t.owner = THIS_MODULE,\n 406:\t\t.set = \u0026bpf_tcp_ops_rcvlowat_kfunc_set,\n 407:\t\t.filter = bpf_tcp_ops_rcvlowat_kfunc_filter,\n 408:\t};\n 409:\t\n 410:\tstatic int __init __bpf_tcp_ops_init(void)\n 411:\t{\n 412:\t\tint ret;\n 413:\t\n 414:\t\tret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,\n 415:\t\t\t\t\t\t\u0026bpf_tcp_ops_rcvlowat_kfunc_id_set);\n 416:\t\tret = ret ?: register_bpf_struct_ops(\u0026bpf_tcp_ops, bpf_tcp_ops);\n 417:\t\n 418:\t\treturn ret;\n 419:\t}\n"
} |
| 37/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24508 output: 41 thoughts: 54 Thoughts: Okay, let's dive into this. **Examining BPF TCP Operations Verification** Alright, I want to take a deeper look at the implementation, specifically the verification process for BPF TCP operations. My focus is on the function `bpf_tcp_ops_verifier` in `net/ipv4/bpf_tcp_ops.c`. I'm particularly interested in how the arguments for `enqueue_rcvq` and `dequeue_rcvq` are typed and, crucially, how those types are verified within that context. I need to understand the constraints and checks placed upon these function calls within the BPF framework to ensure program safety and correct behavior, given that we are dealing with a potentially hostile environment. This level of scrutiny will allow me to assess the soundness and any potential vulnerabilities present in the interface. |
| 38/3 | 2026/09/22 19:44 | tool | grepper |
0mArgs:{
"Expression": "bpf_tcp_ops_verifier",
"PathPrefix": "net/ipv4/bpf_tcp_ops.c"
}Results: {
"Output": "net/ipv4/bpf_tcp_ops.c=340=static int bpf_tcp_ops_validate(void *kdata)\n--\nnet/ipv4/bpf_tcp_ops.c-344-\nnet/ipv4/bpf_tcp_ops.c:345:static const struct bpf_verifier_ops bpf_tcp_ops_verifier = {\nnet/ipv4/bpf_tcp_ops.c-346-\t.get_func_proto\t\t= get_func_proto,\n--\nnet/ipv4/bpf_tcp_ops.c=350=static struct bpf_struct_ops bpf_tcp_ops = {\nnet/ipv4/bpf_tcp_ops.c:351:\t.verifier_ops = \u0026bpf_tcp_ops_verifier,\nnet/ipv4/bpf_tcp_ops.c-352-\t.init_member = bpf_tcp_ops_init_member,\n"
} |
| 39/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 25022 output: 41 thoughts: 30 Thoughts: Okay, let's dive into this. Here's how I'd approach this task, as if I were thinking it through: **Analyzing `bpf_tcp_ops_verifier` in `net/ipv4/bpf_tcp_ops.c`** Alright, time to get my hands dirty. I need to take a look at the `bpf_tcp_ops_verifier` function, specifically within the `net/ipv4/bpf_tcp_ops.c` file. Given my expertise, I already know this likely involves a BPF (Berkeley Packet Filter) verifier, specifically tailored for TCP operations. I'm assuming this is probably where the kernel ensures that any BPF programs used to manipulate TCP behavior are safe and won't crash the system or introduce security vulnerabilities. I'll expect to see checks for things like: * **Bounded Resource Consumption:** Ensuring the BPF program doesn't consume excessive CPU cycles, memory, or cause excessive lock contention. * **Memory Safety:** Verifying that memory accesses are within bounds and don't lead to out-of-bounds reads or writes. This is paramount. * **Control Flow Analysis:** Making sure the program's control flow is well-defined and doesn't contain loops or jumps that could lead to unpredictable behavior, especially since it will be running in the kernel. * **Instruction Set Restrictions:** Checking if the program uses only allowed BPF instructions relevant to TCP manipulation, and that there's no attempt to access sensitive kernel data structures without authorization. * **Type Safety:** confirming the program correctly manipulates data types and doesn't try to perform type-unsafe operations. * **Security Context Validation:** Confirming the BPF program operates within the correct security context. My immediate focus will be to see how the verifier handles the context of TCP-specific operations, how it interacts with TCP socket structures, and what security measures are put in place to prevent misuse. This is likely a critical piece of the puzzle for maintaining TCP stability and security. Time to dig in! |
| 40/3 | 2026/09/22 19:44 | tool | read-file |
0mArgs:{
"File": "net/ipv4/bpf_tcp_ops.c",
"FirstLine": 330,
"LineCount": 25
}Results: {
"Contents": " 330:\t\t\treturn -EINVAL;\n 331:\t\n 332:\t\treturn 0;\n 333:\t}\n 334:\t\n 335:\tstatic int bpf_tcp_ops_init(struct btf *btf)\n 336:\t{\n 337:\t\treturn 0;\n 338:\t}\n 339:\t\n 340:\tstatic int bpf_tcp_ops_validate(void *kdata)\n 341:\t{\n 342:\t\treturn 0;\n 343:\t}\n 344:\t\n 345:\tstatic const struct bpf_verifier_ops bpf_tcp_ops_verifier = {\n 346:\t\t.get_func_proto\t\t= get_func_proto,\n 347:\t\t.is_valid_access\t= is_valid_access,\n 348:\t};\n 349:\t\n 350:\tstatic struct bpf_struct_ops bpf_tcp_ops = {\n 351:\t\t.verifier_ops = \u0026bpf_tcp_ops_verifier,\n 352:\t\t.init_member = bpf_tcp_ops_init_member,\n 353:\t\t.check_member = bpf_tcp_ops_check_member,\n 354:\t\t.init = bpf_tcp_ops_init,\n"
} |
| 41/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 25899 output: 41 thoughts: 27 Thoughts: Here's my summary, as though I'm thinking through this process: **Diving into `is_valid_access`** Okay, so I need to understand how access is validated within the BPF TCP operations in the IPv4 stack. Let's start by looking at `is_valid_access` in `net/ipv4/bpf_tcp_ops.c`. I expect this function is crucial for ensuring that BPF programs don't wreak havoc on TCP connections or try to access memory they shouldn't. I'm hoping to quickly identify the key checks being performed here. Knowing the logic of this function will be essential for understanding the security boundaries and capabilities of these BPF programs. I'll need to figure out what it means to be a "valid access" in this context. Is it checking for out-of-bounds reads or writes, perhaps? Or is it more about controlling what TCP information the BPF program can see or modify? Let's get to the code and see. |
| 42/3 | 2026/09/22 19:44 | tool | read-file |
0mArgs:{
"File": "net/ipv4/bpf_tcp_ops.c",
"FirstLine": 295,
"LineCount": 40
}Results: {
"Contents": " 295:\t\t\t moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))\n 296:\t\t\t\treturn \u0026bpf_tcp_ops_cb_flags_set_proto;\n 297:\t\t\treturn NULL;\n 298:\t\tdefault:\n 299:\t\t\treturn bpf_base_func_proto(func_id, prog);\n 300:\t\t}\n 301:\t}\n 302:\t\n 303:\tstatic bool is_valid_access(int off, int size, enum bpf_access_type type,\n 304:\t\t\t\t const struct bpf_prog *prog, struct bpf_insn_access_aux *info)\n 305:\t{\n 306:\t\tif (!bpf_tracing_btf_ctx_access(off, size, type, prog, info))\n 307:\t\t\treturn false;\n 308:\t\n 309:\t\tif (base_type(info-\u003ereg_type) == PTR_TO_BTF_ID \u0026\u0026\n 310:\t\t !bpf_type_has_unsafe_modifiers(info-\u003ereg_type) \u0026\u0026\n 311:\t\t info-\u003ebtf_id == btf_sock_ids[BTF_SOCK_TYPE_SOCK])\n 312:\t\t\t/* promote it to tcp_sock */\n 313:\t\t\tinfo-\u003ebtf_id = btf_sock_ids[BTF_SOCK_TYPE_TCP];\n 314:\t\n 315:\t\treturn true;\n 316:\t}\n 317:\t\n 318:\tstatic int bpf_tcp_ops_init_member(const struct btf_type *t,\n 319:\t\t\t\t\t const struct btf_member *member,\n 320:\t\t\t\t\t void *kdata, const void *udata)\n 321:\t{\n 322:\t\treturn 0;\n 323:\t}\n 324:\t\n 325:\tstatic int bpf_tcp_ops_check_member(const struct btf_type *t,\n 326:\t\t\t\t\t const struct btf_member *member,\n 327:\t\t\t\t\t const struct bpf_prog *prog)\n 328:\t{\n 329:\t\tif (prog-\u003esleepable)\n 330:\t\t\treturn -EINVAL;\n 331:\t\n 332:\t\treturn 0;\n 333:\t}\n 334:\t\n"
} |
| 43/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 26819 output: 28 thoughts: 1112 Thoughts: **Analyzing `enqueue_rcvq` and `dequeue_rcvq` Usage and the Role of `rcvlowat`** Okay, let's break this down. My immediate task is to understand how `enqueue_rcvq` and `dequeue_rcvq` are used and the impact of setting `rcvlowat` within the TCP stack. First, I need to see where those functions are called. From the provided code snippets, I've identified four key locations where `enqueue_rcvq` is called: `tcp_cleanup_rbuf`, `tcp_fastopen_add_skb`, `tcp_ofo_queue`, and `tcp_queue_rcv`. `dequeue_rcvq` is called from `tcp_cleanup_rbuf`. This gives me a concrete picture of where these BPF helpers are invoked in the receive path. It looks like enqueue is used when adding received SKBs to the receive queue, in fast open, and also for out-of-order packets. The dequeue is in clean up so I need to see if it is used elsewhere. Next, I need to investigate how `__tcp_set_rcvlowat` operates. This function, and its BPF-enabled variant, `bpf_tcp_ops_set_rcvlowat`, looks crucial. It seems to directly influence the `sk->sk_rcvlowat` value. The code confirms that this variable is responsible for determining when `EPOLLIN` or similar events are triggered and also impacts wakeup calls. I need to fully grasp how the `wakeup` flag in `__tcp_set_rcvlowat` is being used and its relationship to `tcp_data_ready`, which is a key function to trigger the events. The fact that `bpf_tcp_ops_set_rcvlowat` allows the BPF program to control `rcvlowat` is very interesting. This is where I can really utilize my knowledge of the kernel. This is an area where advanced packet filtering or even sophisticated rate limiting might be achieved. I need to analyze how it affects things like receive buffer sizing and data copy behavior, I suspect these are the critical factors. So, setting `sk_rcvlowat` influences when data-ready events happen. This prompts a critical question: Does manipulating `sk_rcvlowat` also have side effects on how much data is actually copied to user space, or the size calculations? I'll need to trace the use of `sk_rcvlowat` within the relevant code paths to verify this, especially regarding any buffer length or size considerations. I must now thoroughly investigate where `sk_rcvlowat` is used throughout the kernel's TCP implementation to determine the full scope of its impact. |
| 44/3 | 2026/09/22 19:44 | tool | grepper |
0mArgs:{
"Expression": "sk_rcvlowat",
"PathPrefix": "net/"
}Results: {
"Output": "net/core/sock.c=1194=int sk_setsockopt(struct sock *sk, int level, int optname,\n--\nnet/core/sock.c-1460-\t\telse\nnet/core/sock.c:1461:\t\t\tWRITE_ONCE(sk-\u003esk_rcvlowat, val ? : 1);\nnet/core/sock.c-1462-\t\tbreak;\n--\nnet/core/sock.c=1731=int sk_getsockopt(struct sock *sk, int level, int optname,\n--\nnet/core/sock.c-1873-\tcase SO_RCVLOWAT:\nnet/core/sock.c:1874:\t\tv.val = READ_ONCE(sk-\u003esk_rcvlowat);\nnet/core/sock.c-1875-\t\tbreak;\n--\nnet/core/sock.c=3746=void sock_init_data_uid(struct socket *sock, struct sock *sk, kuid_t uid)\n--\nnet/core/sock.c-3785-\tsk-\u003esk_write_pending\t=\t0;\nnet/core/sock.c:3786:\tsk-\u003esk_rcvlowat\t\t=\t1;\nnet/core/sock.c-3787-\tsk-\u003esk_rcvtimeo\t\t=\tMAX_SCHEDULE_TIMEOUT;\n--\nnet/core/sock.c=4552=static int __init sock_struct_check(void)\n--\nnet/core/sock.c-4567-\tCACHELINE_ASSERT_GROUP_MEMBER(struct sock, sock_read_rx, sk_rcvtimeo);\nnet/core/sock.c:4568:\tCACHELINE_ASSERT_GROUP_MEMBER(struct sock, sock_read_rx, sk_rcvlowat);\nnet/core/sock.c-4569-\n--\nnet/ipv4/tcp.c=1831=int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)\n--\nnet/ipv4/tcp.c-1840-\tval = min(val, cap);\nnet/ipv4/tcp.c:1841:\tWRITE_ONCE(sk-\u003esk_rcvlowat, val ? : 1);\nnet/ipv4/tcp.c-1842-\n--\nnet/ipv4/tcp_input.c=5617=void tcp_data_ready(struct sock *sk)\nnet/ipv4/tcp_input.c-5618-{\nnet/ipv4/tcp_input.c:5619:\tif (tcp_epollin_ready(sk, sk-\u003esk_rcvlowat) || sock_flag(sk, SOCK_DONE))\nnet/ipv4/tcp_input.c-5620-\t\tREAD_ONCE(sk-\u003esk_data_ready)(sk);\n--\nnet/ipv4/tcp_input.c=6131=static void __tcp_ack_snd_check(struct sock *sk, int ofo_possible)\n--\nnet/ipv4/tcp_input.c-6144-\t */\nnet/ipv4/tcp_input.c:6145:\t (tp-\u003ercv_nxt - tp-\u003ecopied_seq \u003c sk-\u003esk_rcvlowat ||\nnet/ipv4/tcp_input.c-6146-\t __tcp_select_window(sk) \u003e= tp-\u003ercv_wnd)) ||\n--\nnet/kcm/kcmsock.c=173=static void kcm_rfree(struct sk_buff *skb)\n--\nnet/kcm/kcmsock.c-186-\tif (!READ_ONCE(kcm-\u003erx_wait) \u0026\u0026 !READ_ONCE(kcm-\u003erx_psock) \u0026\u0026\nnet/kcm/kcmsock.c:187:\t sk_rmem_alloc_get(sk) \u003c sk-\u003esk_rcvlowat) {\nnet/kcm/kcmsock.c-188-\t\tspin_lock_bh(\u0026mux-\u003erx_lock);\n--\nnet/mptcp/protocol.h=869=static inline bool mptcp_epollin_ready(const struct sock *sk)\n--\nnet/mptcp/protocol.h-878-\t */\nnet/mptcp/protocol.h:879:\treturn (data_avail \u003e= sk-\u003esk_rcvlowat) ||\nnet/mptcp/protocol.h-880-\t\ttcp_under_memory_pressure(sk);\n--\nnet/mptcp/sockopt.c=1606=void mptcp_sockopt_sync_locked(struct mptcp_sock *msk, struct sock *ssk)\n--\nnet/mptcp/sockopt.c-1611-\nnet/mptcp/sockopt.c:1612:\tssk-\u003esk_rcvlowat = 0;\nnet/mptcp/sockopt.c-1613-\n--\nnet/mptcp/sockopt.c=1630=int mptcp_set_rcvlowat(struct sock *sk, int val)\n--\nnet/mptcp/sockopt.c-1643-\tval = min(val, cap);\nnet/mptcp/sockopt.c:1644:\tWRITE_ONCE(sk-\u003esk_rcvlowat, val ? : 1);\nnet/mptcp/sockopt.c-1645-\n--\nnet/mptcp/subflow.c=1511=static void subflow_data_ready(struct sock *sk)\n--\nnet/mptcp/subflow.c-1540-\t\t */\nnet/mptcp/subflow.c:1541:\t\tif (mptcp_data_avail(msk) \u003c parent-\u003esk_rcvlowat \u0026\u0026\nnet/mptcp/subflow.c-1542-\t\t (tcp_sk(sk)-\u003ercv_nxt - tcp_sk(sk)-\u003ercv_wup) \u003e inet_csk(sk)-\u003eicsk_ack.rcv_mss)\n--\nnet/sched/em_meta.c=541=META_COLLECTOR(int_sk_prio)\n--\nnet/sched/em_meta.c-551-\nnet/sched/em_meta.c:552:META_COLLECTOR(int_sk_rcvlowat)\nnet/sched/em_meta.c-553-{\n--\nnet/sched/em_meta.c-559-\t}\nnet/sched/em_meta.c:560:\tdst-\u003evalue = READ_ONCE(sk-\u003esk_rcvlowat);\nnet/sched/em_meta.c-561-}\n--\nnet/sched/em_meta.c=621=static struct meta_ops __meta_ops[TCF_META_TYPE_MAX + 1][TCF_META_ID_MAX + 1] = {\n--\nnet/sched/em_meta.c-665-\t\t[META_ID(SK_PRIO)]\t\t= META_FUNC(int_sk_prio),\nnet/sched/em_meta.c:666:\t\t[META_ID(SK_RCVLOWAT)]\t\t= META_FUNC(int_sk_rcvlowat),\nnet/sched/em_meta.c-667-\t\t[META_ID(SK_RCVTIMEO)]\t\t= META_FUNC(int_sk_rcvtimeo),\n--\nnet/smc/af_smc.c=511=static void smc_copy_sock_settings(struct sock *nsk, struct sock *osk,\n--\nnet/smc/af_smc.c-519-\tnsk-\u003esk_priority = READ_ONCE(osk-\u003esk_priority);\nnet/smc/af_smc.c:520:\tnsk-\u003esk_rcvlowat = osk-\u003esk_rcvlowat;\nnet/smc/af_smc.c-521-\tnsk-\u003esk_bound_dev_if = osk-\u003esk_bound_dev_if;\n--\nnet/vmw_vsock/af_vsock.c=1115=void vsock_data_ready(struct sock *sk)\n--\nnet/vmw_vsock/af_vsock.c-1118-\nnet/vmw_vsock/af_vsock.c:1119:\tif (vsock_stream_has_data(vsk) \u003e= sk-\u003esk_rcvlowat ||\nnet/vmw_vsock/af_vsock.c-1120-\t sock_flag(sk, SOCK_DONE))\n--\nnet/vmw_vsock/af_vsock.c=2630=static int vsock_set_rcvlowat(struct sock *sk, int val)\n--\nnet/vmw_vsock/af_vsock.c-2649-\nnet/vmw_vsock/af_vsock.c:2650:\tWRITE_ONCE(sk-\u003esk_rcvlowat, val ? : 1);\nnet/vmw_vsock/af_vsock.c-2651-\treturn 0;\n"
} |
| 45/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 30835 output: 32 thoughts: 49 Thoughts: Okay, let's dive into this. **Investigating `tcp_recvmsg` and `sk_rcvlowat`** My immediate focus is on how `tcp_recvmsg` handles receive low watermarks. I need to determine if `sk_rcvlowat` plays a role in its operation. The best way to do this, of course, is to go directly to the source. Let's examine `tcp_recvmsg` within `net/ipv4/tcp.c`. That's where the core logic resides, and that's where I'll find the answer to my question. Time to get my hands dirty and dissect the code. |
| 46/3 | 2026/09/22 19:44 | tool | grepper |
0mArgs:{
"Expression": "rcvlowat",
"PathPrefix": "net/ipv4/tcp.c"
}Results: {
"Output": "net/ipv4/tcp.c=532=__poll_t tcp_poll(struct file *file, struct socket *sock, poll_table *wait)\n--\nnet/ipv4/tcp.c-588-\t (state != TCP_SYN_RECV || rcu_access_pointer(tp-\u003efastopen_rsk))) {\nnet/ipv4/tcp.c:589:\t\tint target = sock_rcvlowat(sk, 0, INT_MAX);\nnet/ipv4/tcp.c-590-\t\tu16 urg_data = READ_ONCE(tp-\u003eurg_data);\n--\nnet/ipv4/tcp.c=1826=int tcp_peek_len(struct socket *sock)\n--\nnet/ipv4/tcp.c-1830-\nnet/ipv4/tcp.c:1831:int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)\nnet/ipv4/tcp.c-1832-{\n--\nnet/ipv4/tcp.c-1840-\tval = min(val, cap);\nnet/ipv4/tcp.c:1841:\tWRITE_ONCE(sk-\u003esk_rcvlowat, val ? : 1);\nnet/ipv4/tcp.c-1842-\n--\nnet/ipv4/tcp.c-1860-/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */\nnet/ipv4/tcp.c:1861:int tcp_set_rcvlowat(struct sock *sk, int val)\nnet/ipv4/tcp.c-1862-{\nnet/ipv4/tcp.c:1863:\treturn __tcp_set_rcvlowat(sk, val, true);\nnet/ipv4/tcp.c-1864-}\n--\nnet/ipv4/tcp.c=2664=static int tcp_recvmsg_locked(struct sock *sk, struct msghdr *msg, size_t len,\n--\nnet/ipv4/tcp.c-2714-\nnet/ipv4/tcp.c:2715:\ttarget = sock_rcvlowat(sk, flags \u0026 MSG_WAITALL, len);\nnet/ipv4/tcp.c-2716-\n"
} |
| 47/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31640 output: 28 thoughts: 15 Thoughts: Okay, let's dive into this. Here's my take on understanding `sock_rcvlowat`: **Deconstructing `sock_rcvlowat` - My Initial Thoughts** Alright, so the first thing I'm doing is trying to grasp the function of `sock_rcvlowat`. It's pretty clear from the name that this is probably related to the receive buffer's "low water mark" for a socket. My immediate hypothesis is that it likely deals with controlling when `recv` (or similar receive system calls) return. I'm thinking this might be about blocking behavior or signaling when a certain amount of data is available. Is it about ensuring we have *at least* this much data before returning? Or maybe signaling some kind of "ready" state for the receive operation? I know that these low water marks can be critical for avoiding unnecessary system calls and optimizing network throughput, especially in high-performance applications. I need to figure out the exact nuance here – how does this "low water mark" actually influence the receive process and its behavior. Now, let's explore it... |
| 48/3 | 2026/09/22 19:44 | tool | grepper |
0mArgs:{
"Expression": "sock_rcvlowat",
"PathPrefix": "include/"
}Results: {
"Output": "include/net/sock.h=2744=static inline long sock_sndtimeo(const struct sock *sk, bool noblock)\n--\ninclude/net/sock.h-2748-\ninclude/net/sock.h:2749:static inline int sock_rcvlowat(const struct sock *sk, int waitall, int len)\ninclude/net/sock.h-2750-{\n"
} |
| 49/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32028 output: 36 thoughts: 22 Thoughts: Here's my summary, as if I were thinking these thoughts: **Analyzing `sock_rcvlowat` in `include/net/sock.h`** Okay, time to dive into the kernel code again. Let's see what's happening with `sock_rcvlowat`. Specifically, I'm going to pull up the definition, right from `include/net/sock.h`. I need to refresh my understanding of how this impacts receive buffer management. I know it's related to the minimum amount of data a socket has to buffer before being considered readable, but I want to get the specific details clear in my head. Let's see how this variable is actually used in the broader networking stack. I'm expecting to find insights into how the kernel handles the interplay between the socket's receive buffer size, application-level reads, and potential blocking behavior. Time to dig in and remind myself of the core of this crucial variable's role. |
| 50/3 | 2026/09/22 19:44 | tool | read-file |
0mArgs:{
"File": "include/net/sock.h",
"FirstLine": 2748,
"LineCount": 20
}Results: {
"Contents": "2748:\t\n2749:\tstatic inline int sock_rcvlowat(const struct sock *sk, int waitall, int len)\n2750:\t{\n2751:\t\tint v = waitall ? len : min_t(int, READ_ONCE(sk-\u003esk_rcvlowat), len);\n2752:\t\n2753:\t\treturn v ?: 1;\n2754:\t}\n2755:\t\n2756:\t/* Alas, with timeout socket operations are not restartable.\n2757:\t * Compare this to poll().\n2758:\t */\n2759:\tstatic inline int sock_intr_errno(long timeo)\n2760:\t{\n2761:\t\treturn timeo == MAX_SCHEDULE_TIMEOUT ? -ERESTARTSYS : -EINTR;\n2762:\t}\n2763:\t\n2764:\tstruct sock_skb_cb {\n2765:\t\tu32 dropcount;\n2766:\t};\n2767:\t\n"
} |
| 51/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32574 output: 37 thoughts: 80 Thoughts: Okay, here's my understanding, tailored for an expert audience, as if it were my own thought process: **Decoding `sock_rcvlowat` and its Impact on Receive Operations** Right, so `sock_rcvlowat` isn't about *exactly* how many bytes I'll get back from `recvmsg()`. Instead, it's really the minimum amount the system will wait for before it's willing to return, my target length if you will. The actual byte count, now that's totally separate, handled by `copied` which only goes up when data is *actually* copied from the socket buffer (`skbs`) into the user buffer. Interesting. I want to see how this plays out, so I'll need to look at `tcp_recvmsg_locked`. Let's zero in around line 2715 in `net/ipv4/tcp.c`. That should give me the core logic around the receive path and how these two interact. I bet I can nail down the exact moment `copied` is incremented. I'll need to carefully analyze the control flow there, especially if there are any potential race conditions... or buffering oddities that can occur. |
| 52/3 | 2026/09/22 19:44 | tool | read-file |
0mArgs:{
"File": "net/ipv4/tcp.c",
"FirstLine": 2710,
"LineCount": 40
}Results: {
"Contents": "2710:\t\t\tpeek_offset = max(sk_peek_offset(sk, flags), 0);\n2711:\t\t\tpeek_seq = tp-\u003ecopied_seq + peek_offset;\n2712:\t\t\tseq = \u0026peek_seq;\n2713:\t\t}\n2714:\t\n2715:\t\ttarget = sock_rcvlowat(sk, flags \u0026 MSG_WAITALL, len);\n2716:\t\n2717:\t\tdo {\n2718:\t\t\tu32 offset;\n2719:\t\n2720:\t\t\t/* Are we at urgent data? Stop if we have read anything or have SIGURG pending. */\n2721:\t\t\tif (unlikely(tp-\u003eurg_data) \u0026\u0026 tp-\u003eurg_seq == *seq) {\n2722:\t\t\t\tif (copied)\n2723:\t\t\t\t\tbreak;\n2724:\t\t\t\tif (signal_pending(current)) {\n2725:\t\t\t\t\tcopied = timeo ? sock_intr_errno(timeo) : -EAGAIN;\n2726:\t\t\t\t\tbreak;\n2727:\t\t\t\t}\n2728:\t\t\t}\n2729:\t\n2730:\t\t\t/* Next get a buffer. */\n2731:\t\n2732:\t\t\tlast = skb_peek_tail(\u0026sk-\u003esk_receive_queue);\n2733:\t\t\tskb_queue_walk(\u0026sk-\u003esk_receive_queue, skb) {\n2734:\t\t\t\tlast = skb;\n2735:\t\t\t\t/* Now that we have two receive queues this\n2736:\t\t\t\t * shouldn't happen.\n2737:\t\t\t\t */\n2738:\t\t\t\tif (WARN(before(*seq, TCP_SKB_CB(skb)-\u003eseq),\n2739:\t\t\t\t\t \"TCP recvmsg seq # bug: copied %X, seq %X, rcvnxt %X, fl %X\\n\",\n2740:\t\t\t\t\t *seq, TCP_SKB_CB(skb)-\u003eseq, tp-\u003ercv_nxt,\n2741:\t\t\t\t\t flags))\n2742:\t\t\t\t\tbreak;\n2743:\t\n2744:\t\t\t\toffset = *seq - TCP_SKB_CB(skb)-\u003eseq;\n2745:\t\t\t\tif (unlikely(TCP_SKB_CB(skb)-\u003etcp_flags \u0026 TCPHDR_SYN)) {\n2746:\t\t\t\t\tpr_err_once(\"%s: found a SYN, please report !\\n\", __func__);\n2747:\t\t\t\t\toffset--;\n2748:\t\t\t\t}\n2749:\t\t\t\tif (offset \u003c skb-\u003elen)\n"
} |
| 53/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 33604 output: 36 thoughts: 293 Thoughts: **Analyzing the TCP Receive Logic and BPF Integration**
Okay, so thinking through this `target` variable in `tcp_recvmsg_locked`, it's really just a limit, a threshold. It's the number of bytes we *want* to receive. Once `copied` reaches or exceeds `target`, the loop terminates. That makes sense, no uninitialized data being read or copied; it's a stopping condition. The loop keeps going until we have what we need.
Now, I need to understand these BPF changes and how they fit in, specifically, all the related calls. I'm starting with `include/net/tcp.h` and this `tcp_set_sock_ops_cb_flags` function.
```c
static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
{
if (sk_is_mptcp(sk) &&
(val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
return -EOPNOTSUPP;
tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
return 0;
}
```
Wait a second…this `sk` might not be a standard TCP socket; there's a conditional check for `sk_is_mptcp`. What if it's locked or a fullsock? I should check all the call sites of `tcp_set_sock_ops_cb_flags`. Let's see... looking at the callers, I see...
1) `bpf_sol_tcp_setsockopt(sk, optname, ...)` is one call site.
Okay, now to investigate `bpf_sol_tcp_setsockopt`, which is in `net/core/filter.c`.
|
| 54/3 | 2026/09/22 19:44 | tool | read-file |
0mArgs:{
"File": "net/core/filter.c",
"FirstLine": 5580,
"LineCount": 25
}Results: {
"Contents": "5580:\t\t\tbreak;\n5581:\t\tcase TCP_BPF_RTO_MIN:\n5582:\t\t\ttimeout = usecs_to_jiffies(val);\n5583:\t\t\tif (timeout \u003e TCP_RTO_MIN ||\n5584:\t\t\t timeout \u003c TCP_TIMEOUT_MIN)\n5585:\t\t\t\treturn -EINVAL;\n5586:\t\t\tinet_csk(sk)-\u003eicsk_rto_min = timeout;\n5587:\t\t\tbreak;\n5588:\t\tcase TCP_BPF_SOCK_OPS_CB_FLAGS:\n5589:\t\t\tif (val \u0026 ~(BPF_SOCK_OPS_ALL_CB_FLAGS))\n5590:\t\t\t\treturn -EINVAL;\n5591:\t\t\treturn tcp_set_sock_ops_cb_flags(sk, val);\n5592:\t\tdefault:\n5593:\t\t\treturn -EINVAL;\n5594:\t\t}\n5595:\t\n5596:\t\treturn 0;\n5597:\t}\n5598:\t\n5599:\tstatic int sol_tcp_sockopt_congestion(struct sock *sk, char *optval,\n5600:\t\t\t\t\t int *optlen, bool getopt)\n5601:\t{\n5602:\t\tstruct tcp_sock *tp;\n5603:\t\tint ret;\n5604:\t\n"
} |
| 55/3 | 2026/09/22 19:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 34696 output: 32 thoughts: 1595 Thoughts: **Analyzing the Impact of BPF Callback Flags on TCP Socket Operations** Okay, let's break this down. Initially, I see a change where the code now calls `tcp_set_sock_ops_cb_flags` instead of handling the `TCP_BPF_SOCK_OPS_CB_FLAGS` directly. This makes the code cleaner, which is good. Diving deeper, the introduction of `bpf_sock_ops_cb_flags_set` and `bpf_tcp_ops_cb_flags_set` intrigues me. The `bpf_tcp_ops_cb_flags_set` function with its focus on `struct sock *sk` and the `arg1_btf_id` pointing to `BTF_SOCK_TYPE_TCP` tells me we're directly dealing with TCP sockets, and this is where the interesting logic resides. My immediate thought is, when these callbacks are invoked in `bpf_tcp_ops`, are we guaranteed to hold the socket lock, or will this run in interrupt context during tcp input? The code references `enqueue_rcvq`, which points to `tcp_ofo_queue`, `tcp_queue_rcv`, and `tcp_fastopen_add_skb`. These are crucial paths. In each of these paths, the `sk` is the specific TCP socket being processed. Now, I need to understand whether this potentially opens a security issue with uninitialized memory being passed to BPF programs. The `enqueue_rcvq` callback specifically passes a `struct sk_buff *skb`. Is this `skb` uninitialized? No, it's a network received `sk_buff` or something from fastopen. This leads me to review the 4 criteria to check this issue. I must ensure the kernel is not inadvertently passing uninitialized data to userspace through BPF. 1. **Kernel Structures Sent to Userspace:** I see no new structures added that transmit kernel data to userspace that could have an issue. The changes add helpers, kfuncs, and two new callbacks in `struct bpf_tcp_ops`, modifications of existing calls and a uapi header change (comments/flag define). No data from the kernel is directly sent to the user. 2. **Conditional Logic Dependent on Uninitialized Variables:** The code uses `val` and `sk` which are initialized, so nothing there, and `bpf_sock_ops_cb_flags` is initialized when the `tcp_sock` struct is created. I need to confirm precisely how `tcp_sock` is allocated and initialized. I need to be sure the memory backing `tcp_sock` has proper initialization, via `tcp_init_sock` or the general socket initialization, which should set the data to zero. So, overall, my initial assessment suggests this is safe. However, I need to thoroughly trace the initialization of `tcp_sock` to be 100% sure we are not inadvertently exposing uninitialized memory to BPF programs through these new callbacks, but I don't see anything explicitly problematic at this moment. |
| 56/3 | 2026/09/22 19:45 | tool | grepper |
0mArgs:{
"Expression": "bpf_sock_ops_cb_flags",
"PathPrefix": "net/"
}Results: {
"Output": "net/core/filter.c=5517=static int bpf_sol_tcp_getsockopt(struct sock *sk, int optname,\n--\nnet/core/filter.c-5524-\tcase TCP_BPF_SOCK_OPS_CB_FLAGS: {\nnet/core/filter.c:5525:\t\tint cb_flags = tcp_sk(sk)-\u003ebpf_sock_ops_cb_flags;\nnet/core/filter.c-5526-\n--\nnet/core/filter.c=6166=static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {\n--\nnet/core/filter.c-6176-\nnet/core/filter.c:6177:BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,\nnet/core/filter.c-6178-\t int, argval)\n--\nnet/core/filter.c-6196-\nnet/core/filter.c:6197:static const struct bpf_func_proto bpf_sock_ops_cb_flags_set_proto = {\nnet/core/filter.c:6198:\t.func\t\t= bpf_sock_ops_cb_flags_set,\nnet/core/filter.c-6199-\t.gpl_only\t= false,\n--\nnet/core/filter.c=8848=sock_ops_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)\n--\nnet/core/filter.c-8861-\tcase BPF_FUNC_sock_ops_cb_flags_set:\nnet/core/filter.c:8862:\t\treturn \u0026bpf_sock_ops_cb_flags_set_proto;\nnet/core/filter.c-8863-\tcase BPF_FUNC_sock_map_update:\n--\nnet/core/filter.c=10830=static u32 sock_ops_convert_ctx_access(enum bpf_access_type type,\n--\nnet/core/filter.c-11135-\nnet/core/filter.c:11136:\tcase offsetof(struct bpf_sock_ops, bpf_sock_ops_cb_flags):\nnet/core/filter.c:11137:\t\tSOCK_OPS_GET_FIELD(bpf_sock_ops_cb_flags, bpf_sock_ops_cb_flags,\nnet/core/filter.c-11138-\t\t\t\t struct tcp_sock);\n"
} |
| 57/3 | 2026/09/22 19:45 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37648 output: 277 thoughts: 2608 Thoughts: **KMSAN Analysis of bpf_tcp_ops.c Patch: No Uninitialized Memory Risks Detected** Okay, let's break this down from a KMSAN perspective. My initial thought process is to meticulously go through the code, scrutinizing all variables and data structures to ensure that no uninitialized memory is introduced or exposed. I start by examining the core functions: `bpf_tcp_ops_cb_flags_set`, `bpf_tcp_ops_set_rcvlowat`, and `__tcp_set_rcvlowat`. Each function is carefully inspected. I confirm that all variables, like `val`, `err`, `moff`, and `wakeup`, are explicitly initialized before use, a critical first step. Next, I look at the `__bpf_tcp_ops_init` function and verify that the `ret` variable is initialized immediately upon declaration. This function registers the BPF struct ops. Then I thoroughly analyze the `__tcp_set_rcvlowat` function. Crucially, I examine how `sk_rcvlowat` is set and used in relation to potential data copying. I determine that it's a threshold for waking up processes or controlling `recvmsg` behaviour, and it never directly exposes uninitialized buffer contents to user space because actual data copying relies on the `skb->len` which is initialized. Then, I focused on data structures and dynamic allocation. The `__bpf_tcp_ops` struct is statically initialized with function stubs for `enqueue_rcvq` and `dequeue_rcvq`. There's no dynamic allocation of structures with uninitialized fields. This is key. The stub functions themselves are a red flag for concurrency bugs, but this would be caught by LOCKDEP and not KMSAN. I investigate whether the patch alters buffer lengths or introduces any new bounds check vulnerabilities. The code only changes the `sk_rcvlowat` threshold which only determines when `tcp_data_ready` is called and does not introduce out-of-bounds reads. I realize the `wakeup` boolean is only true for `dequeue_rcvq` because of where the `bpf_tcp_ops_set_rcvlowat` kfunc is called. The key question then becomes: can the `sk_rcvlowat` setting allow access to uninitialized data in any way, e.g., via `tcp_recvmsg_locked`? I determine that the data copying from `skbs` (network packets) to user space is safe because of the buffer lengths checks (`skb->len`). My focus shifts to the BPF verifier and struct ops. I check `bpf_tcp_ops_set_rcvlowat`'s `KF_IMPLICIT_ARGS` to ensure that it has the correct argument setup, that `sk` is type-checked and that `aux` is coming from the kernel. I then go on to think of bugs that could arise. What if someone loads a BPF struct_ops prog that implements `enqueue_rcvq` or `dequeue_rcvq`? I dig into the `get_func_proto` function to see if those functions are blocked from being called via `bpf_sk_setsockopt` and `bpf_sk_getsockopt`. Based on this analysis, I decide that they are not. But then I have to think about what happens if those functions are called from an interrupt context, can that produce a race condition? Finally, after exhausting every angle, I can confidently assert that this patch does not introduce any risks of uninitialized memory being exposed. No data structures are copied to user space without initialization, no variables are used before initialization, and no boundary checks are compromised. Any potential bugs (e.g., concurrency, locking, or out-of-bounds) would be caught by other tools like KASAN, LOCKDEP and standard debugging facilities. I double-check my reasoning against the four KMSAN criteria to ensure that I have covered every possibility. |
| 58/3 | 2026/09/22 19:45 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch introduces TCP receive queue callback operations (enqueue_rcvq and dequeue_rcvq) in bpf_tcp_ops and adds a kfunc (bpf_tcp_ops_set_rcvlowat) along with flags to manage SO_RCVLOWAT dynamically from BPF.\n\n1. No uninitialized memory or info-leak risks: The patch does not introduce any allocations without initialization, nor does it expose kernel structs or buffers to userspace that could contain uninitialized padding or fields.\n2. Logic and bounds: All variables (such as wakeup flags, offsets, and return values) are fully initialized before use. The modifications to `__tcp_set_rcvlowat` adjust socket wake-up conditions and epoll/recvmsg wake thresholds, which do not permit reading uninitialized buffer memory (packet data consumption remains bounded by actual skb lengths).\n3. Sanitizer applicability: Any potential issues introduced by this patch (e.g., locking/context constraints in BPF callbacks, reference handling, or concurrency) fall within the domain of KASAN, LOCKDEP, and standard kernel assertions, not KMSAN.\n\nTherefore, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch introduces TCP receive queue callback operations (enqueue_rcvq and dequeue_rcvq) in bpf_tcp_ops and adds a kfunc (bpf_tcp_ops_set_rcvlowat) along with flags to manage SO_RCVLOWAT dynamically from BPF.\n\n1. No uninitialized memory or info-leak risks: The patch does not introduce any allocations without initialization, nor does it expose kernel structs or buffers to userspace that could contain uninitialized padding or fields.\n2. Logic and bounds: All variables (such as wakeup flags, offsets, and return values) are fully initialized before use. The modifications to `__tcp_set_rcvlowat` adjust socket wake-up conditions and epoll/recvmsg wake thresholds, which do not permit reading uninitialized buffer memory (packet data consumption remains bounded by actual skb lengths).\n3. Sanitizer applicability: Any potential issues introduced by this patch (e.g., locking/context constraints in BPF callbacks, reference handling, or concurrency) fall within the domain of KASAN, LOCKDEP, and standard kernel assertions, not KMSAN.\n\nTherefore, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|