| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/09/26 22:18 | flow | patch-triage |
1mArgs:null Results: {
"EnableConfigs": [
"BPF_SYSCALL"
],
"FocusSymbols": [
"bpf_prog_stream_new_fd",
"prog_stream_open",
"bpf_stream_file_read",
"bpf_stream_poll",
"bpf_stream_release",
"bpf_stream_notify",
"bpf_prog_stream_init",
"bpf_prog_stream_free"
],
"KMSANReasoning": "The patch introduces the BPF_PROG_STREAM_OPEN command to the bpf() syscall, allowing user space to open a file descriptor for BPF streams with read() and poll() capabilities. It changes struct bpf_stream from being statically embedded in bpf_prog_aux to being dynamically allocated with refcounting.\n\nAnalysis of uninitialized memory risks (KMSAN):\n1. All dynamic allocations of struct bpf_stream use kzalloc_obj(), ensuring all fields (including atomic counters, flags like notify_used and dead, and backlog pointers) are fully zero-initialized before being passed to their respective init functions (refcount_set, init_llist_head, mutex_init, init_waitqueue_head, init_irq_work).\n2. Data transmitted to user space via copy_to_user() in bpf_stream_read() comes from stream element string buffers (elem-\u003estr), which are zero-allocated (__GFP_ZERO) and fully initialized via memcpy() upon creation. The readable counter only accounts for committed bytes, preventing access to uninitialized bytes.\n3. No struct copies containing uninitialized fields or alignment padding are returned to user space (the syscall only returns a file descriptor integer).\n\nApplicability of KASAN vs KMSAN:\nThe primary risks introduced by this change relate to object lifetimes, reference counting (bpf_stream_put / refcount_inc), concurrent readers/writers, and synchronization across irq_work and anon_inode file operations. Any potential bugs in these areas (such as use-after-free, double-free, or locking issues) are effectively detected by KASAN, LOCKDEP, and standard kernel debugging facilities. No uninitialized memory usage or info-leak risks are present, so a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false,
"Reasoning": "The patch introduces the BPF_PROG_STREAM_OPEN command in the bpf() system call along with anon_inode-backed file operations (read, poll, release) for BPF stream objects. It adds waitqueue handling, irq_work notifications from BPF execution contexts, reference counting, and teardown synchronization. This is new reachable executable code in the core BPF subsystem.",
"WorthFuzzing": true
} |
| 1/1 | 2026/09/26 22:18 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit 6b74fb908de0cf6223a7c477c7bf211b1f2476d2\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Sat Sep 26 22:18:53 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/include/linux/bpf.h b/include/linux/bpf.h\nindex 1e1ce2afe2ed8..a38badb1c2f7d 100644\n--- a/include/linux/bpf.h\n+++ b/include/linux/bpf.h\n@@ -17,6 +17,7 @@\n #include \u003clinux/numa.h\u003e\n #include \u003clinux/mm_types.h\u003e\n #include \u003clinux/wait.h\u003e\n+#include \u003clinux/irq_work_types.h\u003e\n #include \u003clinux/refcount.h\u003e\n #include \u003clinux/mutex.h\u003e\n #include \u003clinux/module.h\u003e\n@@ -1770,12 +1771,18 @@ enum {\n };\n \n struct bpf_stream {\n-\tatomic_t capacity;\n+\trefcount_t refcnt;\n+\tatomic_t capacity;\t/* bytes reserved against the stream limit */\n+\tatomic_t readable;\t/* published bytes available to readers */\n \tstruct llist_head log;\t/* list of in-flight stream elements in LIFO order */\n \n \tstruct mutex lock; /* lock protecting backlog_{head,tail} */\n \tstruct llist_node *backlog_head; /* list of in-flight stream elements in FIFO order */\n \tstruct llist_node *backlog_tail; /* tail of the list above */\n+\twait_queue_head_t waitq;\n+\tstruct irq_work notify_work;\n+\tbool notify_used;\t/* notify_work was queued at least once */\n+\tbool dead;\n };\n \n struct bpf_stream_stage {\n@@ -1914,7 +1921,7 @@ struct bpf_prog_aux {\n \t\tstruct work_struct work;\n \t\tstruct rcu_head\trcu;\n \t};\n-\tstruct bpf_stream stream[2];\n+\tstruct bpf_stream *stream[2];\n \tstruct mutex st_ops_assoc_mutex;\n \tstruct bpf_map __rcu *st_ops_assoc;\n };\n@@ -4223,9 +4230,10 @@ void bpf_bprintf_cleanup(struct bpf_bprintf_data *data);\n int bpf_try_get_buffers(struct bpf_bprintf_buffers **bufs);\n void bpf_put_buffers(void);\n \n-void bpf_prog_stream_init(struct bpf_prog *prog);\n+int bpf_prog_stream_init(struct bpf_prog *prog, gfp_t gfp_extra_flags);\n void bpf_prog_stream_free(struct bpf_prog *prog);\n int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len);\n+int bpf_prog_stream_new_fd(struct bpf_prog *prog, enum bpf_stream_id stream_id, u32 flags);\n void bpf_stream_stage_init(struct bpf_stream_stage *ss);\n void bpf_stream_stage_free(struct bpf_stream_stage *ss);\n __printf(2, 3)\ndiff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h\nindex 0aaa54359aebc..4687c33109968 100644\n--- a/include/uapi/linux/bpf.h\n+++ b/include/uapi/linux/bpf.h\n@@ -936,6 +936,33 @@ union bpf_iter_link_info {\n * \t\t0 on success or -1 if an error occurred (in which case,\n * \t\t*errno* is set appropriately).\n *\n+ * BPF_PROG_STREAM_OPEN\n+ *\tDescription\n+ *\t\tOpen a file descriptor for one of the BPF streams associated\n+ *\t\twith the program identified by *prog_fd*. The stream is selected\n+ *\t\tby *stream_id*.\n+ *\n+ *\t\tThe returned file descriptor supports **read**\\ (2) and\n+ *\t\t**poll**\\ (2). Reads block while the stream is empty unless\n+ *\t\t**BPF_F_STREAM_NONBLOCK** is specified in *flags*. A non-blocking\n+ *\t\tread of an empty stream fails with **EAGAIN**.\n+ *\n+ *\t\t**poll**\\ (2) reports **POLLIN** when data is available and\n+ *\t\t**POLLHUP** once the program has been freed, that is, after every\n+ *\t\treference to it, including links and other file descriptors, has\n+ *\t\tbeen dropped. Hangup may lag the final release because program\n+ *\t\tteardown is deferred. Buffered data remains readable after\n+ *\t\t**POLLHUP** and a read returns zero after all such data has been\n+ *\t\tconsumed.\n+ *\n+ *\t\tThe file descriptor is read-only and has the close-on-exec flag\n+ *\t\tset. It is not seekable and **lseek**\\ (2) fails with **ESPIPE**.\n+ *\t\t*flags* may only contain **BPF_F_STREAM_NONBLOCK**.\n+ *\n+ *\tReturn\n+ *\t\tA new file descriptor (a nonnegative integer), or -1 if an\n+ *\t\terror occurred (in which case, *errno* is set appropriately).\n+ *\n * NOTES\n *\teBPF objects (maps and programs) can be shared between processes.\n *\n@@ -993,6 +1020,7 @@ enum bpf_cmd {\n \tBPF_TOKEN_CREATE,\n \tBPF_PROG_STREAM_READ_BY_FD,\n \tBPF_PROG_ASSOC_STRUCT_OPS,\n+\tBPF_PROG_STREAM_OPEN,\n \t__MAX_BPF_CMD,\n \tBPF_COMMON_ATTRS = 1 \u003c\u003c 16, /* Indicate carrying syscall common attrs. */\n };\n@@ -1524,6 +1552,11 @@ enum {\n \tBPF_STREAM_STDERR = 2,\n };\n \n+/* flags for BPF_PROG_STREAM_OPEN command */\n+enum {\n+\tBPF_F_STREAM_NONBLOCK = (1U \u003c\u003c 0),\n+};\n+\n union bpf_attr {\n \tstruct { /* anonymous struct used by BPF_MAP_CREATE command */\n \t\t__u32\tmap_type;\t/* one of enum bpf_map_type */\n@@ -1950,6 +1983,12 @@ union bpf_attr {\n \t\t__u32\t\tflags;\n \t} prog_assoc_struct_ops;\n \n+\tstruct {\n+\t\t__u32\t\tprog_fd;\n+\t\t__u32\t\tstream_id;\n+\t\t__u32\t\tflags;\n+\t} prog_stream_open;\n+\n } __attribute__((aligned(8)));\n \n /* The description below is an attempt at providing documentation to eBPF\ndiff --git a/kernel/bpf/core.c b/kernel/bpf/core.c\nindex 273f74068068c..134700f425319 100644\n--- a/kernel/bpf/core.c\n+++ b/kernel/bpf/core.c\n@@ -141,10 +141,6 @@ struct bpf_prog *bpf_prog_alloc_no_stats(unsigned int size, gfp_t gfp_extra_flag\n \tmutex_init(\u0026fp-\u003eaux-\u003edst_mutex);\n \tmutex_init(\u0026fp-\u003eaux-\u003est_ops_assoc_mutex);\n \n-#ifdef CONFIG_BPF_SYSCALL\n-\tbpf_prog_stream_init(fp);\n-#endif\n-\n \treturn fp;\n }\n \n@@ -288,6 +284,9 @@ struct bpf_prog *bpf_prog_realloc(struct bpf_prog *fp_old, unsigned int size,\n void __bpf_prog_free(struct bpf_prog *fp)\n {\n \tif (fp-\u003eaux) {\n+#ifdef CONFIG_BPF_SYSCALL\n+\t\tbpf_prog_stream_free(fp);\n+#endif\n \t\tmutex_destroy(\u0026fp-\u003eaux-\u003eused_maps_mutex);\n \t\tmutex_destroy(\u0026fp-\u003eaux-\u003edst_mutex);\n \t\tmutex_destroy(\u0026fp-\u003eaux-\u003est_ops_assoc_mutex);\n@@ -3073,7 +3072,6 @@ static void bpf_prog_free_deferred(struct work_struct *work)\n \taux = container_of(work, struct bpf_prog_aux, work);\n #ifdef CONFIG_BPF_SYSCALL\n \tbpf_free_kfunc_btf_tab(aux-\u003ekfunc_btf_tab);\n-\tbpf_prog_stream_free(aux-\u003eprog);\n #endif\n #ifdef CONFIG_CGROUP_BPF\n \tif (aux-\u003ecgroup_atype != CGROUP_BPF_ATTACH_TYPE_INVALID)\ndiff --git a/kernel/bpf/stream.c b/kernel/bpf/stream.c\nindex 2b80a0599865e..e0ce37fcd5578 100644\n--- a/kernel/bpf/stream.c\n+++ b/kernel/bpf/stream.c\n@@ -2,11 +2,15 @@\n /* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */\n \n #include \u003clinux/bpf.h\u003e\n+#include \u003clinux/anon_inodes.h\u003e\n #include \u003clinux/filter.h\u003e\n #include \u003clinux/bpf_mem_alloc.h\u003e\n #include \u003clinux/gfp.h\u003e\n+#include \u003clinux/irq_work.h\u003e\n #include \u003clinux/memory.h\u003e\n #include \u003clinux/mutex.h\u003e\n+#include \u003clinux/poll.h\u003e\n+#include \u003clinux/refcount.h\u003e\n \n static void bpf_stream_elem_init(struct bpf_stream_elem *elem, int len)\n {\n@@ -73,16 +77,59 @@ static void bpf_stream_release_capacity(struct bpf_stream *stream, int len)\n \tatomic_sub(len, \u0026stream-\u003ecapacity);\n }\n \n+static void bpf_stream_notify(struct irq_work *work)\n+{\n+\tstruct bpf_stream *stream = container_of(work, struct bpf_stream, notify_work);\n+\n+\t/*\n+\t * Writers run in arbitrary program contexts, including NMI and regions\n+\t * that already hold wait queue or epoll locks. Wake waiters from\n+\t * irq_work instead, where taking those locks is safe.\n+\t */\n+\twake_up_interruptible_poll(\u0026stream-\u003ewaitq, EPOLLIN | EPOLLRDNORM);\n+}\n+\n+static void bpf_stream_queue_notify(struct bpf_stream *stream)\n+{\n+\t/*\n+\t * Record that the work has been used so that teardown only pays for\n+\t * irq_work_sync(), which may wait for an RCU grace period, when a\n+\t * callback could actually be in flight.\n+\t */\n+\tif (!READ_ONCE(stream-\u003enotify_used))\n+\t\tWRITE_ONCE(stream-\u003enotify_used, true);\n+\tirq_work_queue(\u0026stream-\u003enotify_work);\n+}\n+\n+static int bpf_stream_readable_bytes(struct bpf_stream *stream)\n+{\n+\treturn atomic_read_acquire(\u0026stream-\u003ereadable);\n+}\n+\n+static void bpf_stream_publish(struct bpf_stream *stream, int len)\n+{\n+\t/* Pairs with atomic_read_acquire() in bpf_stream_readable_bytes(). */\n+\t(void)atomic_add_return_release(len, \u0026stream-\u003ereadable);\n+\tbpf_stream_queue_notify(stream);\n+}\n+\n static int bpf_stream_push_str(struct bpf_stream *stream, const char *str, int len)\n {\n-\tint ret = bpf_stream_consume_capacity(stream, len);\n+\tint ret;\n+\n+\t/* Nothing to publish; do not allocate an element for it. */\n+\tif (!len)\n+\t\treturn 0;\n \n+\tret = bpf_stream_consume_capacity(stream, len);\n \tif (ret)\n \t\treturn ret;\n \n \tret = __bpf_stream_push_str(\u0026stream-\u003elog, str, len);\n \tif (ret)\n \t\tbpf_stream_release_capacity(stream, len);\n+\telse\n+\t\tbpf_stream_publish(stream, len);\n \n \treturn ret;\n }\n@@ -91,7 +138,7 @@ static struct bpf_stream *bpf_stream_get(enum bpf_stream_id stream_id, struct bp\n {\n \tif (stream_id != BPF_STDOUT \u0026\u0026 stream_id != BPF_STDERR)\n \t\treturn NULL;\n-\treturn \u0026aux-\u003estream[stream_id - 1];\n+\treturn aux-\u003estream[stream_id - 1];\n }\n \n static void bpf_stream_free_elem(struct bpf_stream_elem *elem)\n@@ -159,14 +206,16 @@ static bool bpf_stream_consume_elem(struct bpf_stream_elem *elem, int *len)\n \n static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)\n {\n-\tint rem_len = len, cons_len, ret = 0;\n+\tint read_len, rem_len, cons_len, ret = 0;\n \tstruct bpf_stream_elem *elem = NULL;\n \tstruct llist_node *node;\n \n \tmutex_lock(\u0026stream-\u003elock);\n+\tread_len = min(len, bpf_stream_readable_bytes(stream));\n+\trem_len = read_len;\n \n \twhile (rem_len) {\n-\t\tint pos = len - rem_len;\n+\t\tint pos = read_len - rem_len;\n \t\tint chunk, n;\n \t\tbool cont;\n \n@@ -188,7 +237,7 @@ static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)\n \t\t\t/* Keep any successfully copied bytes; -EFAULT only if none. */\n \t\t\telem-\u003econsumed_len -= n;\n \t\t\trem_len += n;\n-\t\t\tret = (len == rem_len) ? -EFAULT : 0;\n+\t\t\tret = (read_len == rem_len) ? -EFAULT : 0;\n \t\t\tbreak;\n \t\t}\n \n@@ -199,8 +248,9 @@ static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)\n \t\tbpf_stream_free_elem(elem);\n \t}\n \n+\tatomic_sub(read_len - rem_len, \u0026stream-\u003ereadable);\n \tmutex_unlock(\u0026stream-\u003elock);\n-\treturn ret ? ret : len - rem_len;\n+\treturn ret ? ret : read_len - rem_len;\n }\n \n int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len)\n@@ -215,6 +265,111 @@ int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, vo\n \treturn bpf_stream_read(stream, buf, len);\n }\n \n+static bool bpf_stream_has_data(struct bpf_stream *stream)\n+{\n+\treturn bpf_stream_readable_bytes(stream) \u003e 0;\n+}\n+\n+static void bpf_stream_put(struct bpf_stream *stream)\n+{\n+\tif (refcount_dec_and_test(\u0026stream-\u003erefcnt)) {\n+\t\tstruct llist_node *list;\n+\n+\t\t/* Only a stream that ever queued its work can have a callback in flight. */\n+\t\tif (READ_ONCE(stream-\u003enotify_used))\n+\t\t\tirq_work_sync(\u0026stream-\u003enotify_work);\n+\t\tlist = llist_del_all(\u0026stream-\u003elog);\n+\t\tbpf_stream_free_list(list);\n+\t\tbpf_stream_free_list(stream-\u003ebacklog_head);\n+\t\tmutex_destroy(\u0026stream-\u003elock);\n+\t\tkfree(stream);\n+\t}\n+}\n+\n+static int bpf_stream_release(struct inode *inode, struct file *file)\n+{\n+\tbpf_stream_put(file-\u003eprivate_data);\n+\treturn 0;\n+}\n+\n+static ssize_t bpf_stream_file_read(struct file *file, char __user *buf, size_t len,\n+\t\t\t\t loff_t *ppos)\n+{\n+\tstruct bpf_stream *stream = file-\u003eprivate_data;\n+\tbool dead;\n+\tint ret;\n+\n+\tif (!len)\n+\t\treturn 0;\n+\n+\tfor (;;) {\n+\t\t/*\n+\t\t * Sample teardown state before looking for data. Nothing is\n+\t\t * published once the program is gone, so finding the stream\n+\t\t * empty after observing dead means EOF. The opposite order could\n+\t\t * report EOF while data published just before teardown is still\n+\t\t * buffered.\n+\t\t */\n+\t\tdead = smp_load_acquire(\u0026stream-\u003edead);\n+\t\tret = bpf_stream_read(stream, buf, len);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t\tif (dead)\n+\t\t\treturn 0;\n+\t\tif (file-\u003ef_flags \u0026 O_NONBLOCK)\n+\t\t\treturn -EAGAIN;\n+\n+\t\tret = wait_event_interruptible(stream-\u003ewaitq,\n+\t\t\t\t\t bpf_stream_has_data(stream) ||\n+\t\t\t\t\t READ_ONCE(stream-\u003edead));\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t}\n+}\n+\n+static __poll_t bpf_stream_poll(struct file *file, struct poll_table_struct *pts)\n+{\n+\tstruct bpf_stream *stream = file-\u003eprivate_data;\n+\t__poll_t events = 0;\n+\n+\t/*\n+\t * poll_wait() only registers the wait queue callback. Register before\n+\t * checking persistent state so a concurrent publication or teardown is\n+\t * observed either by the callback or by the checks below.\n+\t */\n+\tpoll_wait(file, \u0026stream-\u003ewaitq, pts);\n+\tif (bpf_stream_has_data(stream))\n+\t\tevents |= EPOLLIN | EPOLLRDNORM;\n+\tif (READ_ONCE(stream-\u003edead))\n+\t\tevents |= EPOLLHUP;\n+\treturn events;\n+}\n+\n+static const struct file_operations bpf_stream_fops = {\n+\t.release = bpf_stream_release,\n+\t.read = bpf_stream_file_read,\n+\t.poll = bpf_stream_poll,\n+};\n+\n+int bpf_prog_stream_new_fd(struct bpf_prog *prog, enum bpf_stream_id stream_id, u32 flags)\n+{\n+\tstruct bpf_stream *stream;\n+\tint fd_flags = O_RDONLY | O_CLOEXEC;\n+\tint fd;\n+\n+\tstream = bpf_stream_get(stream_id, prog-\u003eaux);\n+\tif (!stream)\n+\t\treturn -ENOENT;\n+\tif (flags \u0026 BPF_F_STREAM_NONBLOCK)\n+\t\tfd_flags |= O_NONBLOCK;\n+\n+\trefcount_inc(\u0026stream-\u003erefcnt);\n+\tfd = anon_inode_getfd(\"bpf-stream\", \u0026bpf_stream_fops, stream, fd_flags);\n+\tif (fd \u003c 0)\n+\t\tbpf_stream_put(stream);\n+\treturn fd;\n+}\n+\n __bpf_kfunc_start_defs();\n \n /*\n@@ -282,28 +437,47 @@ __bpf_kfunc_end_defs();\n \n /* Added kfunc to common_btf_ids */\n \n-void bpf_prog_stream_init(struct bpf_prog *prog)\n+int bpf_prog_stream_init(struct bpf_prog *prog, gfp_t gfp_extra_flags)\n {\n \tint i;\n \n \tfor (i = 0; i \u003c ARRAY_SIZE(prog-\u003eaux-\u003estream); i++) {\n-\t\tatomic_set(\u0026prog-\u003eaux-\u003estream[i].capacity, 0);\n-\t\tinit_llist_head(\u0026prog-\u003eaux-\u003estream[i].log);\n-\t\tmutex_init(\u0026prog-\u003eaux-\u003estream[i].lock);\n-\t\tprog-\u003eaux-\u003estream[i].backlog_head = NULL;\n-\t\tprog-\u003eaux-\u003estream[i].backlog_tail = NULL;\n+\t\tstruct bpf_stream *stream;\n+\n+\t\t/* On failure, bpf_prog_stream_free() releases the streams allocated so far. */\n+\t\tstream = kzalloc_obj(*stream,\n+\t\t\t\t bpf_memcg_flags(GFP_KERNEL | gfp_extra_flags));\n+\t\tif (!stream)\n+\t\t\treturn -ENOMEM;\n+\n+\t\trefcount_set(\u0026stream-\u003erefcnt, 1);\n+\t\tinit_llist_head(\u0026stream-\u003elog);\n+\t\tmutex_init(\u0026stream-\u003elock);\n+\t\tinit_waitqueue_head(\u0026stream-\u003ewaitq);\n+\t\tinit_irq_work(\u0026stream-\u003enotify_work, bpf_stream_notify);\n+\t\tprog-\u003eaux-\u003estream[i] = stream;\n \t}\n+\treturn 0;\n }\n \n void bpf_prog_stream_free(struct bpf_prog *prog)\n {\n-\tstruct llist_node *list;\n \tint i;\n \n \tfor (i = 0; i \u003c ARRAY_SIZE(prog-\u003eaux-\u003estream); i++) {\n-\t\tlist = llist_del_all(\u0026prog-\u003eaux-\u003estream[i].log);\n-\t\tbpf_stream_free_list(list);\n-\t\tbpf_stream_free_list(prog-\u003eaux-\u003estream[i].backlog_head);\n+\t\tstruct bpf_stream *stream = prog-\u003eaux-\u003estream[i];\n+\n+\t\tif (!stream)\n+\t\t\tcontinue;\n+\t\t/*\n+\t\t * Pairs with smp_load_acquire() in bpf_stream_file_read(): every\n+\t\t * publication precedes the dead flag, so a reader that observes\n+\t\t * it also observes all buffered data.\n+\t\t */\n+\t\tsmp_store_release(\u0026stream-\u003edead, true);\n+\t\twake_up_interruptible_poll(\u0026stream-\u003ewaitq, EPOLLHUP);\n+\t\tbpf_stream_put(stream);\n+\t\tprog-\u003eaux-\u003estream[i] = NULL;\n \t}\n }\n \n@@ -333,8 +507,8 @@ int bpf_stream_stage_printk(struct bpf_stream_stage *ss, const char *fmt, ...)\n \tva_start(args, fmt);\n \tlen = vscnprintf(buf-\u003ebuf, ARRAY_SIZE(buf-\u003ebuf), fmt, args);\n \tva_end(args);\n-\t/* Exclude NULL byte during push. */\n-\tret = __bpf_stream_push_str(\u0026ss-\u003elog, buf-\u003ebuf, len);\n+\t/* Exclude NULL byte during push; skip empty output entirely. */\n+\tret = len ? __bpf_stream_push_str(\u0026ss-\u003elog, buf-\u003ebuf, len) : 0;\n \tif (!ret)\n \t\tss-\u003elen += len;\n \tbpf_put_buffers();\n@@ -366,6 +540,7 @@ int bpf_stream_stage_commit(struct bpf_stream_stage *ss, struct bpf_prog *prog,\n \t\tlist = tail;\n \t}\n \tllist_add_batch(head, tail, \u0026stream-\u003elog);\n+\tbpf_stream_publish(stream, ss-\u003elen);\n \treturn 0;\n }\n \ndiff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c\nindex 74496fd716d3b..8580f38b41c16 100644\n--- a/kernel/bpf/syscall.c\n+++ b/kernel/bpf/syscall.c\n@@ -3080,6 +3080,10 @@ static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_at\n \tprog-\u003eaux-\u003euser = get_current_user();\n \tprog-\u003elen = attr-\u003einsn_cnt;\n \n+\terr = bpf_prog_stream_init(prog, GFP_USER);\n+\tif (err)\n+\t\tgoto free_prog;\n+\n \terr = -EFAULT;\n \tif (copy_from_bpfptr(prog-\u003einsns,\n \t\t\t make_bpfptr(attr-\u003einsns, uattr.is_kernel),\n@@ -6313,6 +6317,28 @@ static int prog_assoc_struct_ops(union bpf_attr *attr)\n \treturn ret;\n }\n \n+#define BPF_PROG_STREAM_OPEN_LAST_FIELD prog_stream_open.flags\n+\n+static int prog_stream_open(union bpf_attr *attr)\n+{\n+\tstruct bpf_prog *prog;\n+\tu32 flags = attr-\u003eprog_stream_open.flags;\n+\tint ret;\n+\n+\tif (CHECK_ATTR(BPF_PROG_STREAM_OPEN))\n+\t\treturn -EINVAL;\n+\tif (flags \u0026 ~BPF_F_STREAM_NONBLOCK)\n+\t\treturn -EINVAL;\n+\n+\tprog = bpf_prog_get(attr-\u003eprog_stream_open.prog_fd);\n+\tif (IS_ERR(prog))\n+\t\treturn PTR_ERR(prog);\n+\n+\tret = bpf_prog_stream_new_fd(prog, attr-\u003eprog_stream_open.stream_id, flags);\n+\tbpf_prog_put(prog);\n+\treturn ret;\n+}\n+\n static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,\n \t\t bpfptr_t uattr_common, unsigned int size_common)\n {\n@@ -6485,6 +6511,9 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,\n \tcase BPF_PROG_ASSOC_STRUCT_OPS:\n \t\terr = prog_assoc_struct_ops(\u0026attr);\n \t\tbreak;\n+\tcase BPF_PROG_STREAM_OPEN:\n+\t\terr = prog_stream_open(\u0026attr);\n+\t\tbreak;\n \tdefault:\n \t\terr = -EINVAL;\n \t\tbreak;\ndiff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst\nindex 90fe8c61bf424..ef7e00eb5e479 100644\n--- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst\n+++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst\n@@ -186,6 +186,11 @@ bpftool prog tracelog { stdout | stderr } *PROG*\n error messages to the standard error stream. This facility should be used\n only for debugging purposes.\n \n+ On kernels that support opening a stream as a file descriptor, bpftool\n+ keeps printing new output as the program produces it, until the program is\n+ unloaded or \u003cCtrl+C\u003e is hit. Older kernels only allow dumping the output\n+ buffered so far, after which bpftool exits.\n+\n bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ctx_in *FILE* [ctx_out *FILE* [ctx_size_out *M*]]] [repeat *N*]\n Run BPF program *PROG* in the kernel testing infrastructure for BPF,\n meaning that the program works on the data and context provided by the\ndiff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c\nindex 24e40dfab4690..f0241dada5485 100644\n--- a/tools/bpf/bpftool/prog.c\n+++ b/tools/bpf/bpftool/prog.c\n@@ -1119,21 +1119,69 @@ enum prog_tracelog_mode {\n \tTRACE_STDERR,\n };\n \n+static volatile sig_atomic_t stream_stop;\n+\n+static void stop_stream(int signo)\n+{\n+\tstream_stop = 1;\n+}\n+\n+/* Consumes prog_fd. */\n static int\n prog_tracelog_stream(int prog_fd, enum prog_tracelog_mode mode)\n {\n+\t/* No SA_RESTART: an interrupted read() must return EINTR to end the loop. */\n+\tconst struct sigaction act = { .sa_handler = stop_stream };\n+\tconst int signals[] = { SIGHUP, SIGINT, SIGTERM };\n+\tstruct sigaction old[ARRAY_SIZE(signals)];\n \tFILE *file = mode == TRACE_STDOUT ? stdout : stderr;\n \tint stream_id = mode == TRACE_STDOUT ? 1 : 2;\n \tchar buf[512];\n-\tint ret;\n+\tunsigned int i;\n+\tint fd, ret;\n+\n+\tfd = bpf_prog_stream_open(prog_fd, stream_id, NULL);\n+\tif (fd == -EINVAL) {\n+\t\t/* Kernel predates BPF_PROG_STREAM_OPEN: dump buffered output and exit. */\n+\t\tdo {\n+\t\t\tret = bpf_prog_stream_read(prog_fd, stream_id, buf, sizeof(buf), NULL);\n+\t\t\tif (ret \u003e 0)\n+\t\t\t\tfwrite(buf, sizeof(buf[0]), ret, file);\n+\t\t} while (ret \u003e 0);\n+\t\tif (ret \u003c 0)\n+\t\t\tp_err(\"failed to read stream: %s\", strerror(-ret));\n+\t\tclose(prog_fd);\n+\t\tgoto out;\n+\t}\n+\t/*\n+\t * The stream descriptor does not keep the program alive. Drop the\n+\t * program reference so that reads return EOF once the program is gone.\n+\t */\n+\tclose(prog_fd);\n+\tif (fd \u003c 0) {\n+\t\tp_err(\"failed to open stream: %s\", strerror(-fd));\n+\t\treturn -1;\n+\t}\n \n+\tstream_stop = 0;\n+\tfor (i = 0; i \u003c ARRAY_SIZE(signals); i++)\n+\t\tsigaction(signals[i], \u0026act, \u0026old[i]);\n \tret = 0;\n-\tdo {\n-\t\tret = bpf_prog_stream_read(prog_fd, stream_id, buf, sizeof(buf), NULL);\n-\t\tif (ret \u003e 0)\n-\t\t\tfwrite(buf, sizeof(buf[0]), ret, file);\n-\t} while (ret \u003e 0);\n-\n+\twhile (!stream_stop) {\n+\t\tret = read(fd, buf, sizeof(buf));\n+\t\tif (ret \u003c= 0)\n+\t\t\tbreak;\n+\t\tfwrite(buf, sizeof(buf[0]), ret, file);\n+\t\tfflush(file);\n+\t}\n+\tif (ret \u003c 0 \u0026\u0026 !(stream_stop \u0026\u0026 errno == EINTR))\n+\t\tp_err(\"failed to read stream: %s\", strerror(errno));\n+\telse\n+\t\tret = 0;\n+\tfor (i = 0; i \u003c ARRAY_SIZE(signals); i++)\n+\t\tsigaction(signals[i], \u0026old[i], NULL);\n+\tclose(fd);\n+out:\n \tfflush(file);\n \treturn ret ? -1 : 0;\n }\ndiff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h\nindex 0aaa54359aebc..4687c33109968 100644\n--- a/tools/include/uapi/linux/bpf.h\n+++ b/tools/include/uapi/linux/bpf.h\n@@ -936,6 +936,33 @@ union bpf_iter_link_info {\n * \t\t0 on success or -1 if an error occurred (in which case,\n * \t\t*errno* is set appropriately).\n *\n+ * BPF_PROG_STREAM_OPEN\n+ *\tDescription\n+ *\t\tOpen a file descriptor for one of the BPF streams associated\n+ *\t\twith the program identified by *prog_fd*. The stream is selected\n+ *\t\tby *stream_id*.\n+ *\n+ *\t\tThe returned file descriptor supports **read**\\ (2) and\n+ *\t\t**poll**\\ (2). Reads block while the stream is empty unless\n+ *\t\t**BPF_F_STREAM_NONBLOCK** is specified in *flags*. A non-blocking\n+ *\t\tread of an empty stream fails with **EAGAIN**.\n+ *\n+ *\t\t**poll**\\ (2) reports **POLLIN** when data is available and\n+ *\t\t**POLLHUP** once the program has been freed, that is, after every\n+ *\t\treference to it, including links and other file descriptors, has\n+ *\t\tbeen dropped. Hangup may lag the final release because program\n+ *\t\tteardown is deferred. Buffered data remains readable after\n+ *\t\t**POLLHUP** and a read returns zero after all such data has been\n+ *\t\tconsumed.\n+ *\n+ *\t\tThe file descriptor is read-only and has the close-on-exec flag\n+ *\t\tset. It is not seekable and **lseek**\\ (2) fails with **ESPIPE**.\n+ *\t\t*flags* may only contain **BPF_F_STREAM_NONBLOCK**.\n+ *\n+ *\tReturn\n+ *\t\tA new file descriptor (a nonnegative integer), or -1 if an\n+ *\t\terror occurred (in which case, *errno* is set appropriately).\n+ *\n * NOTES\n *\teBPF objects (maps and programs) can be shared between processes.\n *\n@@ -993,6 +1020,7 @@ enum bpf_cmd {\n \tBPF_TOKEN_CREATE,\n \tBPF_PROG_STREAM_READ_BY_FD,\n \tBPF_PROG_ASSOC_STRUCT_OPS,\n+\tBPF_PROG_STREAM_OPEN,\n \t__MAX_BPF_CMD,\n \tBPF_COMMON_ATTRS = 1 \u003c\u003c 16, /* Indicate carrying syscall common attrs. */\n };\n@@ -1524,6 +1552,11 @@ enum {\n \tBPF_STREAM_STDERR = 2,\n };\n \n+/* flags for BPF_PROG_STREAM_OPEN command */\n+enum {\n+\tBPF_F_STREAM_NONBLOCK = (1U \u003c\u003c 0),\n+};\n+\n union bpf_attr {\n \tstruct { /* anonymous struct used by BPF_MAP_CREATE command */\n \t\t__u32\tmap_type;\t/* one of enum bpf_map_type */\n@@ -1950,6 +1983,12 @@ union bpf_attr {\n \t\t__u32\t\tflags;\n \t} prog_assoc_struct_ops;\n \n+\tstruct {\n+\t\t__u32\t\tprog_fd;\n+\t\t__u32\t\tstream_id;\n+\t\t__u32\t\tflags;\n+\t} prog_stream_open;\n+\n } __attribute__((aligned(8)));\n \n /* The description below is an attempt at providing documentation to eBPF\ndiff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c\nindex a9de7f107cf7b..b49822d212aed 100644\n--- a/tools/lib/bpf/bpf.c\n+++ b/tools/lib/bpf/bpf.c\n@@ -1466,6 +1466,25 @@ int bpf_prog_stream_read(int prog_fd, __u32 stream_id, void *buf, __u32 buf_len,\n \treturn libbpf_err_errno(err);\n }\n \n+int bpf_prog_stream_open(int prog_fd, __u32 stream_id,\n+\t\t\t const struct bpf_prog_stream_open_opts *opts)\n+{\n+\tconst size_t attr_sz = offsetofend(union bpf_attr, prog_stream_open);\n+\tunion bpf_attr attr;\n+\tint fd;\n+\n+\tif (!OPTS_VALID(opts, bpf_prog_stream_open_opts))\n+\t\treturn libbpf_err(-EINVAL);\n+\n+\tmemset(\u0026attr, 0, attr_sz);\n+\tattr.prog_stream_open.prog_fd = prog_fd;\n+\tattr.prog_stream_open.stream_id = stream_id;\n+\tattr.prog_stream_open.flags = OPTS_GET(opts, flags, 0);\n+\n+\tfd = sys_bpf_fd(BPF_PROG_STREAM_OPEN, \u0026attr, attr_sz);\n+\treturn libbpf_err_errno(fd);\n+}\n+\n int bpf_prog_assoc_struct_ops(int prog_fd, int map_fd,\n \t\t\t struct bpf_prog_assoc_struct_ops_opts *opts)\n {\ndiff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h\nindex 490e8cb4ba537..826d9cc9ab65d 100644\n--- a/tools/lib/bpf/bpf.h\n+++ b/tools/lib/bpf/bpf.h\n@@ -759,10 +759,34 @@ struct bpf_prog_stream_read_opts {\n *\n * @return The number of bytes read, on success; negative error code, otherwise\n * (errno is also set to the error code)\n+ *\n+ * For blocking reads and polling, prefer **bpf_prog_stream_open**.\n */\n LIBBPF_API int bpf_prog_stream_read(int prog_fd, __u32 stream_id, void *buf, __u32 buf_len,\n \t\t\t\t struct bpf_prog_stream_read_opts *opts);\n \n+struct bpf_prog_stream_open_opts {\n+\tsize_t sz;\n+\t__u32 flags;\n+\tsize_t :0;\n+};\n+#define bpf_prog_stream_open_opts__last_field flags\n+\n+/**\n+ * @brief **bpf_prog_stream_open** opens a file descriptor for a BPF stream of\n+ * a given BPF program.\n+ *\n+ * @param prog_fd FD for the BPF program whose BPF stream is to be opened.\n+ * @param stream_id ID of the BPF stream to be opened.\n+ * @param opts optional options, can be NULL. BPF_F_STREAM_NONBLOCK requests a\n+ * non-blocking descriptor.\n+ *\n+ * @return A new stream FD, on success; negative error code, otherwise (errno\n+ * is also set to the error code)\n+ */\n+LIBBPF_API int bpf_prog_stream_open(int prog_fd, __u32 stream_id,\n+\t\t\t\t const struct bpf_prog_stream_open_opts *opts);\n+\n struct bpf_prog_assoc_struct_ops_opts {\n \tsize_t sz;\n \t__u32 flags;\ndiff --git a/tools/lib/bpf/libbpf.map b/tools/lib/bpf/libbpf.map\nindex 18d27d20102ec..53b7b591325f5 100644\n--- a/tools/lib/bpf/libbpf.map\n+++ b/tools/lib/bpf/libbpf.map\n@@ -459,6 +459,7 @@ LIBBPF_1.7.0 {\n LIBBPF_1.8.0 {\n \tglobal:\n \t\tbpf_map__attach_cgroup_opts;\n+\t\tbpf_prog_stream_open;\n \t\tbpf_program__add_flags;\n \t\tbpf_program__attach_tracing_multi;\n \t\tbpf_program__clear_flags;\ndiff --git a/tools/testing/selftests/bpf/prog_tests/stream.c b/tools/testing/selftests/bpf/prog_tests/stream.c\nindex 74bd15c4bfbaa..2c6b441552862 100644\n--- a/tools/testing/selftests/bpf/prog_tests/stream.c\n+++ b/tools/testing/selftests/bpf/prog_tests/stream.c\n@@ -1,11 +1,17 @@\n // SPDX-License-Identifier: GPL-2.0\n /* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */\n #include \u003ctest_progs.h\u003e\n+#include \u003clinux/perf_event.h\u003e\n+#include \u003cpoll.h\u003e\n+#include \u003csys/epoll.h\u003e\n #include \u003csys/mman.h\u003e\n+#include \u003csys/syscall.h\u003e\n \n #include \"stream.skel.h\"\n #include \"stream_fail.skel.h\"\n \n+#define NMI_TIMEOUT_NS (5ULL * 1000 * 1000 * 1000)\n+\n void test_stream_failure(void)\n {\n \tRUN_TESTS(stream_fail);\n@@ -62,6 +68,415 @@ void test_stream_syscall(void)\n \tstream__destroy(skel);\n }\n \n+static bool stream_fd_trigger(struct bpf_program *prog)\n+{\n+\tLIBBPF_OPTS(bpf_test_run_opts, opts);\n+\tint ret;\n+\n+\tret = bpf_prog_test_run_opts(bpf_program__fd(prog), \u0026opts);\n+\treturn ASSERT_OK(ret, \"test_run\") \u0026\u0026 ASSERT_OK(opts.retval, \"retval\");\n+}\n+\n+static void test_stream_fd_open(void)\n+{\n+\tLIBBPF_OPTS(bpf_prog_stream_open_opts, opts);\n+\tstruct stream *skel;\n+\tint fd, prog_fd;\n+\n+\tskel = stream__open_and_load();\n+\tif (!ASSERT_OK_PTR(skel, \"stream__open_and_load\"))\n+\t\treturn;\n+\n+\tprog_fd = bpf_program__fd(skel-\u003eprogs.stream_syscall);\n+\tfd = bpf_prog_stream_open(0, BPF_STREAM_STDOUT, NULL);\n+\tASSERT_EQ(fd, -EINVAL, \"bad_prog_fd\");\n+\n+\tfd = bpf_prog_stream_open(prog_fd, 0, NULL);\n+\tASSERT_EQ(fd, -ENOENT, \"bad_stream_id\");\n+\n+\topts.flags = BPF_F_RDONLY;\n+\tfd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDOUT, \u0026opts);\n+\tASSERT_EQ(fd, -EINVAL, \"access_flag\");\n+\n+\topts.flags = 1U \u003c\u003c 31;\n+\tfd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDOUT, \u0026opts);\n+\tASSERT_EQ(fd, -EINVAL, \"unknown_flag\");\n+\n+\tfd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDERR, NULL);\n+\tif (ASSERT_OK_FD(fd, \"stderr\"))\n+\t\tclose(fd);\n+\n+\tstream__destroy(skel);\n+}\n+\n+static void test_stream_fd_nonblock(void)\n+{\n+\tLIBBPF_OPTS(bpf_prog_stream_open_opts, opts,\n+\t\t.flags = BPF_F_STREAM_NONBLOCK,\n+\t);\n+\tstruct pollfd pfd = { .events = POLLIN | POLLHUP };\n+\tstruct stream *skel;\n+\tchar buf[4] = {};\n+\tint fd, flags, ret;\n+\n+\tskel = stream__open_and_load();\n+\tif (!ASSERT_OK_PTR(skel, \"stream__open_and_load\"))\n+\t\treturn;\n+\n+\tfd = bpf_prog_stream_open(bpf_program__fd(skel-\u003eprogs.stream_syscall),\n+\t\t\t\t BPF_STREAM_STDOUT, \u0026opts);\n+\tif (!ASSERT_OK_FD(fd, \"stream_open\"))\n+\t\tgoto out_destroy;\n+\tpfd.fd = fd;\n+\n+\tflags = fcntl(fd, F_GETFD);\n+\tASSERT_GE(flags, 0, \"getfd\");\n+\tASSERT_NEQ(flags \u0026 FD_CLOEXEC, 0, \"cloexec\");\n+\tflags = fcntl(fd, F_GETFL);\n+\tASSERT_GE(flags, 0, \"getfl\");\n+\tASSERT_EQ(flags \u0026 O_ACCMODE, O_RDONLY, \"readonly\");\n+\tASSERT_NEQ(flags \u0026 O_NONBLOCK, 0, \"nonblock\");\n+\n+\tret = write(fd, \"x\", 1);\n+\tASSERT_EQ(ret, -1, \"write\");\n+\tASSERT_EQ(errno, EBADF, \"write_errno\");\n+\tret = lseek(fd, 0, SEEK_SET);\n+\tASSERT_EQ(ret, -1, \"lseek\");\n+\tASSERT_EQ(errno, ESPIPE, \"lseek_errno\");\n+\n+\tret = poll(\u0026pfd, 1, 0);\n+\tASSERT_EQ(ret, 0, \"poll_empty\");\n+\tret = read(fd, buf, sizeof(buf));\n+\tASSERT_EQ(ret, -1, \"read_empty\");\n+\tASSERT_EQ(errno, EAGAIN, \"read_empty_errno\");\n+\n+\tif (!stream_fd_trigger(skel-\u003eprogs.stream_syscall))\n+\t\tgoto out_close;\n+\tret = poll(\u0026pfd, 1, 0);\n+\tASSERT_EQ(ret, 1, \"poll_data\");\n+\tASSERT_NEQ(pfd.revents \u0026 POLLIN, 0, \"pollin\");\n+\tASSERT_EQ(pfd.revents \u0026 POLLHUP, 0, \"no_pollhup\");\n+\n+\tret = read(fd, buf, 2);\n+\tASSERT_EQ(ret, 2, \"read_first\");\n+\tASSERT_OK(memcmp(buf, \"fo\", 2), \"read_first_data\");\n+\tpfd.revents = 0;\n+\tret = poll(\u0026pfd, 1, 0);\n+\tASSERT_EQ(ret, 1, \"poll_partial\");\n+\tASSERT_NEQ(pfd.revents \u0026 POLLIN, 0, \"pollin_partial\");\n+\n+\tret = read(fd, buf, sizeof(buf));\n+\tASSERT_EQ(ret, 1, \"read_rest\");\n+\tASSERT_EQ(buf[0], 'o', \"read_rest_data\");\n+\tret = read(fd, buf, sizeof(buf));\n+\tASSERT_EQ(ret, -1, \"read_drained\");\n+\tASSERT_EQ(errno, EAGAIN, \"read_drained_errno\");\n+\n+out_close:\n+\tclose(fd);\n+out_destroy:\n+\tstream__destroy(skel);\n+}\n+\n+static void test_stream_fd_empty(void)\n+{\n+\tLIBBPF_OPTS(bpf_prog_stream_open_opts, opts,\n+\t\t.flags = BPF_F_STREAM_NONBLOCK,\n+\t);\n+\tstruct pollfd pfd = { .events = POLLIN | POLLHUP };\n+\tstruct stream *skel;\n+\tchar buf[4];\n+\tint fd, ret;\n+\n+\tskel = stream__open_and_load();\n+\tif (!ASSERT_OK_PTR(skel, \"stream__open_and_load\"))\n+\t\treturn;\n+\n+\tfd = bpf_prog_stream_open(bpf_program__fd(skel-\u003eprogs.stream_empty),\n+\t\t\t\t BPF_STREAM_STDOUT, \u0026opts);\n+\tif (!ASSERT_OK_FD(fd, \"stream_open\"))\n+\t\tgoto out_destroy;\n+\tpfd.fd = fd;\n+\n+\t/* Empty output produces neither data nor readiness. */\n+\tif (!stream_fd_trigger(skel-\u003eprogs.stream_empty))\n+\t\tgoto out_close;\n+\tret = poll(\u0026pfd, 1, 0);\n+\tASSERT_EQ(ret, 0, \"poll_empty_write\");\n+\tret = read(fd, buf, sizeof(buf));\n+\tASSERT_EQ(ret, -1, \"read_empty_write\");\n+\tASSERT_EQ(errno, EAGAIN, \"read_empty_write_errno\");\n+\n+out_close:\n+\tclose(fd);\n+out_destroy:\n+\tstream__destroy(skel);\n+}\n+\n+struct stream_fd_read_ctx {\n+\tint fd;\n+\tssize_t ret;\n+\tchar buf[4];\n+};\n+\n+static void *stream_fd_read_thread(void *arg)\n+{\n+\tstruct stream_fd_read_ctx *ctx = arg;\n+\n+\tctx-\u003eret = read(ctx-\u003efd, ctx-\u003ebuf, sizeof(ctx-\u003ebuf));\n+\treturn NULL;\n+}\n+\n+static int stream_fd_timed_join(pthread_t thread)\n+{\n+\tstruct timespec timeout;\n+\n+\tclock_gettime(CLOCK_REALTIME, \u0026timeout);\n+\ttimeout.tv_sec += 5;\n+\treturn pthread_timedjoin_np(thread, NULL, \u0026timeout);\n+}\n+\n+static void test_stream_fd_blocking(void)\n+{\n+\tstruct stream_fd_read_ctx ctx = {};\n+\tstruct stream *skel;\n+\tpthread_t thread;\n+\tbool thread_live = false;\n+\tint fd, flags, ret;\n+\n+\tskel = stream__open_and_load();\n+\tif (!ASSERT_OK_PTR(skel, \"stream__open_and_load\"))\n+\t\treturn;\n+\n+\tfd = bpf_prog_stream_open(bpf_program__fd(skel-\u003eprogs.stream_syscall),\n+\t\t\t\t BPF_STREAM_STDOUT, NULL);\n+\tif (!ASSERT_OK_FD(fd, \"stream_open\"))\n+\t\tgoto out_destroy;\n+\tctx.fd = fd;\n+\n+\tflags = fcntl(fd, F_GETFL);\n+\tASSERT_GE(flags, 0, \"getfl\");\n+\tASSERT_EQ(flags \u0026 O_NONBLOCK, 0, \"blocking\");\n+\n+\tret = pthread_create(\u0026thread, NULL, stream_fd_read_thread, \u0026ctx);\n+\tif (!ASSERT_OK(ret, \"pthread_create\"))\n+\t\tgoto out_close;\n+\tthread_live = true;\n+\n+\tusleep(50000);\n+\tret = pthread_tryjoin_np(thread, NULL);\n+\tif (!ASSERT_EQ(ret, EBUSY, \"read_blocks\")) {\n+\t\tthread_live = ret != 0;\n+\t\tgoto out_thread;\n+\t}\n+\tif (!stream_fd_trigger(skel-\u003eprogs.stream_syscall))\n+\t\tgoto out_thread;\n+\n+\tret = stream_fd_timed_join(thread);\n+\tif (!ASSERT_OK(ret, \"pthread_join\"))\n+\t\tgoto out_thread;\n+\tthread_live = false;\n+\tASSERT_EQ(ctx.ret, 3, \"read_len\");\n+\tASSERT_OK(memcmp(ctx.buf, \"foo\", 3), \"read_data\");\n+\n+\tmemset(\u0026ctx, 0, sizeof(ctx));\n+\tctx.fd = fd;\n+\tret = pthread_create(\u0026thread, NULL, stream_fd_read_thread, \u0026ctx);\n+\tif (!ASSERT_OK(ret, \"pthread_create_eof\"))\n+\t\tgoto out_close;\n+\tthread_live = true;\n+\n+\tusleep(50000);\n+\tret = pthread_tryjoin_np(thread, NULL);\n+\tif (!ASSERT_EQ(ret, EBUSY, \"read_eof_blocks\")) {\n+\t\tthread_live = ret != 0;\n+\t\tgoto out_thread;\n+\t}\n+\n+\tstream__destroy(skel);\n+\tskel = NULL;\n+\tret = stream_fd_timed_join(thread);\n+\tif (!ASSERT_OK(ret, \"pthread_join_eof\"))\n+\t\tgoto out_thread;\n+\tthread_live = false;\n+\tASSERT_EQ(ctx.ret, 0, \"read_eof\");\n+\n+out_thread:\n+\tif (thread_live) {\n+\t\tstream__destroy(skel);\n+\t\tskel = NULL;\n+\t\tret = stream_fd_timed_join(thread);\n+\t\tif (ret) {\n+\t\t\tpthread_cancel(thread);\n+\t\t\tpthread_join(thread, NULL);\n+\t\t}\n+\t}\n+out_close:\n+\tclose(fd);\n+out_destroy:\n+\tstream__destroy(skel);\n+}\n+\n+static void test_stream_fd_hup(void)\n+{\n+\tLIBBPF_OPTS(bpf_prog_stream_open_opts, opts,\n+\t\t.flags = BPF_F_STREAM_NONBLOCK,\n+\t);\n+\tstruct epoll_event event = {\n+\t\t.events = EPOLLIN | EPOLLET,\n+\t};\n+\tstruct pollfd pfd = { .events = POLLIN | POLLHUP };\n+\tstruct stream *skel;\n+\tchar buf[4] = {};\n+\tint epfd = -1, fd, ret;\n+\n+\tskel = stream__open_and_load();\n+\tif (!ASSERT_OK_PTR(skel, \"stream__open_and_load\"))\n+\t\treturn;\n+\n+\tfd = bpf_prog_stream_open(bpf_program__fd(skel-\u003eprogs.stream_syscall),\n+\t\t\t\t BPF_STREAM_STDOUT, \u0026opts);\n+\tif (!ASSERT_OK_FD(fd, \"stream_open\"))\n+\t\tgoto out_destroy;\n+\tpfd.fd = fd;\n+\tepfd = epoll_create1(EPOLL_CLOEXEC);\n+\tif (!ASSERT_OK_FD(epfd, \"epoll_create\"))\n+\t\tgoto out_close;\n+\tevent.data.fd = fd;\n+\tret = epoll_ctl(epfd, EPOLL_CTL_ADD, fd, \u0026event);\n+\tif (!ASSERT_OK(ret, \"epoll_ctl\"))\n+\t\tgoto out_close;\n+\n+\tif (!stream_fd_trigger(skel-\u003eprogs.stream_syscall))\n+\t\tgoto out_close;\n+\tret = epoll_wait(epfd, \u0026event, 1, 5000);\n+\tif (!ASSERT_EQ(ret, 1, \"epoll_wait_data\"))\n+\t\tgoto out_close;\n+\tASSERT_NEQ(event.events \u0026 EPOLLIN, 0, \"epollin\");\n+\tASSERT_EQ(event.events \u0026 EPOLLHUP, 0, \"no_epollhup\");\n+\n+\tstream__destroy(skel);\n+\tskel = NULL;\n+\tevent.events = 0;\n+\tret = epoll_wait(epfd, \u0026event, 1, 5000);\n+\tif (!ASSERT_EQ(ret, 1, \"epoll_wait_hup\"))\n+\t\tgoto out_close;\n+\tASSERT_NEQ(event.events \u0026 EPOLLIN, 0, \"epollin_with_hup\");\n+\tASSERT_NEQ(event.events \u0026 EPOLLHUP, 0, \"epollhup\");\n+\n+\tret = read(fd, buf, sizeof(buf));\n+\tASSERT_EQ(ret, 3, \"read_buffered\");\n+\tASSERT_OK(memcmp(buf, \"foo\", 3), \"read_buffered_data\");\n+\tpfd.revents = 0;\n+\tret = poll(\u0026pfd, 1, 0);\n+\tASSERT_EQ(ret, 1, \"poll_drained_hup\");\n+\tASSERT_EQ(pfd.revents \u0026 POLLIN, 0, \"no_pollin_after_drain\");\n+\tASSERT_NEQ(pfd.revents \u0026 POLLHUP, 0, \"pollhup_after_drain\");\n+\tret = read(fd, buf, sizeof(buf));\n+\tASSERT_EQ(ret, 0, \"read_eof\");\n+\n+out_close:\n+\tif (epfd \u003e= 0)\n+\t\tclose(epfd);\n+\tclose(fd);\n+out_destroy:\n+\tstream__destroy(skel);\n+}\n+\n+static void test_stream_fd_nmi_epoll(void)\n+{\n+\tLIBBPF_OPTS(bpf_prog_stream_open_opts, opts,\n+\t\t.flags = BPF_F_STREAM_NONBLOCK,\n+\t);\n+\tstruct perf_event_attr attr = {\n+\t\t.size = sizeof(attr),\n+\t\t.type = PERF_TYPE_HARDWARE,\n+\t\t.config = PERF_COUNT_HW_CPU_CYCLES,\n+\t\t.freq = 1,\n+\t\t.sample_freq = 10,\n+\t};\n+\tstruct epoll_event event = {\n+\t\t.events = EPOLLIN,\n+\t};\n+\tstruct bpf_link *link = NULL;\n+\tstruct stream *skel;\n+\t__u64 deadline;\n+\tchar buf[4] = {};\n+\tint epfd = -1, fd = -1, pmu_fd = -1;\n+\tint ret;\n+\n+\tskel = stream__open_and_load();\n+\tif (!ASSERT_OK_PTR(skel, \"stream__open_and_load\"))\n+\t\treturn;\n+\tfd = bpf_prog_stream_open(bpf_program__fd(skel-\u003eprogs.stream_nmi),\n+\t\t\t\t BPF_STREAM_STDOUT, \u0026opts);\n+\tif (!ASSERT_OK_FD(fd, \"stream_open\"))\n+\t\tgoto out;\n+\n+\tepfd = epoll_create1(EPOLL_CLOEXEC);\n+\tif (!ASSERT_OK_FD(epfd, \"epoll_create\"))\n+\t\tgoto out;\n+\tevent.data.fd = fd;\n+\tret = epoll_ctl(epfd, EPOLL_CTL_ADD, fd, \u0026event);\n+\tif (!ASSERT_OK(ret, \"epoll_ctl\"))\n+\t\tgoto out;\n+\n+\tpmu_fd = syscall(__NR_perf_event_open, \u0026attr, 0, -1, -1,\n+\t\t\t PERF_FLAG_FD_CLOEXEC);\n+\tif (pmu_fd \u003c 0 \u0026\u0026 (errno == ENOENT || errno == EOPNOTSUPP)) {\n+\t\tprintf(\"%s:SKIP:no PERF_COUNT_HW_CPU_CYCLES\\n\", __func__);\n+\t\ttest__skip();\n+\t\tgoto out;\n+\t}\n+\tif (!ASSERT_GE(pmu_fd, 0, \"perf_event_open\"))\n+\t\tgoto out;\n+\n+\tlink = bpf_program__attach_perf_event(skel-\u003eprogs.stream_nmi, pmu_fd);\n+\tif (!ASSERT_OK_PTR(link, \"attach_perf_event\")) {\n+\t\tlink = NULL;\n+\t\tgoto out;\n+\t}\n+\tpmu_fd = -1;\n+\n+\tdeadline = get_time_ns() + NMI_TIMEOUT_NS;\n+\tdo {\n+\t\tret = epoll_wait(epfd, \u0026event, 1, 0);\n+\t} while (ret == 0 \u0026\u0026 get_time_ns() \u003c deadline);\n+\tif (!ASSERT_EQ(ret, 1, \"epoll_wait\"))\n+\t\tgoto out;\n+\tASSERT_NEQ(event.events \u0026 EPOLLIN, 0, \"epollin\");\n+\n+\tret = read(fd, buf, sizeof(buf));\n+\tASSERT_EQ(ret, 3, \"read_len\");\n+\tASSERT_OK(memcmp(buf, \"nmi\", 3), \"read_data\");\n+\n+out:\n+\tbpf_link__destroy(link);\n+\tif (pmu_fd \u003e= 0)\n+\t\tclose(pmu_fd);\n+\tif (epfd \u003e= 0)\n+\t\tclose(epfd);\n+\tif (fd \u003e= 0)\n+\t\tclose(fd);\n+\tstream__destroy(skel);\n+}\n+\n+void test_stream_fd(void)\n+{\n+\tif (test__start_subtest(\"open\"))\n+\t\ttest_stream_fd_open();\n+\tif (test__start_subtest(\"nonblock\"))\n+\t\ttest_stream_fd_nonblock();\n+\tif (test__start_subtest(\"empty\"))\n+\t\ttest_stream_fd_empty();\n+\tif (test__start_subtest(\"blocking\"))\n+\t\ttest_stream_fd_blocking();\n+\tif (test__start_subtest(\"hup\"))\n+\t\ttest_stream_fd_hup();\n+\tif (test__start_subtest(\"nmi_epoll\"))\n+\t\ttest_stream_fd_nmi_epoll();\n+}\n+\n void test_stream_oversize(void)\n {\n \tLIBBPF_OPTS(bpf_test_run_opts, opts);\ndiff --git a/tools/testing/selftests/bpf/progs/stream.c b/tools/testing/selftests/bpf/progs/stream.c\nindex 12fc29e454874..1ae015c4e6b4c 100644\n--- a/tools/testing/selftests/bpf/progs/stream.c\n+++ b/tools/testing/selftests/bpf/progs/stream.c\n@@ -44,6 +44,7 @@ struct {\n \t\t _X64 _X64 _X64 _X64 _X64 _X64 _X64 _X64)\n \n int size;\n+int nmi_stream_prints;\n u64 fault_addr;\n void *arena_ptr;\n \n@@ -124,6 +125,25 @@ int stream_syscall(void *ctx)\n \treturn 0;\n }\n \n+SEC(\"syscall\")\n+__success __retval(0)\n+int stream_empty(void *ctx)\n+{\n+\treturn bpf_stream_printk(BPF_STDOUT, \"\");\n+}\n+\n+SEC(\"perf_event\")\n+int stream_nmi(void *ctx)\n+{\n+\tif (nmi_stream_prints)\n+\t\treturn 0;\n+\t/* Retry on a later sample if the write failed, e.g. with -EBUSY. */\n+\tif (bpf_stream_printk(BPF_STDOUT, \"nmi\"))\n+\t\treturn 0;\n+\tnmi_stream_prints = 1;\n+\treturn 0;\n+}\n+\n SEC(\"syscall\")\n __success __retval(0)\n int stream_oversize(void *ctx)\n"
} |
| 2/1 | 2026/09/26 22:18 | agent | patch-evaluator |
0mModel:core Args: null Results: {
"EnableConfigs": [
"BPF_SYSCALL"
],
"FocusSymbols": [
"bpf_prog_stream_new_fd",
"prog_stream_open",
"bpf_stream_file_read",
"bpf_stream_poll",
"bpf_stream_release",
"bpf_stream_notify",
"bpf_prog_stream_init",
"bpf_prog_stream_free"
],
"Reasoning": "The patch introduces the BPF_PROG_STREAM_OPEN command in the bpf() system call along with anon_inode-backed file operations (read, poll, release) for BPF stream objects. It adds waitqueue handling, irq_work notifications from BPF execution contexts, reference counting, and teardown synchronization. This is new reachable executable code in the core BPF subsystem.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 6b74fb908de0cf6223a7c477c7bf211b1f2476d2
Author: syz-cluster <triage@syzkaller.com>
Date: Sat Sep 26 22:18:53 2026 +0000
syz-cluster: applied patch under review
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 1e1ce2afe2ed8..a38badb1c2f7d 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -17,6 +17,7 @@
#include <linux/numa.h>
#include <linux/mm_types.h>
#include <linux/wait.h>
+#include <linux/irq_work_types.h>
#include <linux/refcount.h>
#include <linux/mutex.h>
#include <linux/module.h>
@@ -1770,12 +1771,18 @@ enum {
};
struct bpf_stream {
- atomic_t capacity;
+ refcount_t refcnt;
+ atomic_t capacity; /* bytes reserved against the stream limit */
+ atomic_t readable; /* published bytes available to readers */
struct llist_head log; /* list of in-flight stream elements in LIFO order */
struct mutex lock; /* lock protecting backlog_{head,tail} */
struct llist_node *backlog_head; /* list of in-flight stream elements in FIFO order */
struct llist_node *backlog_tail; /* tail of the list above */
+ wait_queue_head_t waitq;
+ struct irq_work notify_work;
+ bool notify_used; /* notify_work was queued at least once */
+ bool dead;
};
struct bpf_stream_stage {
@@ -1914,7 +1921,7 @@ struct bpf_prog_aux {
struct work_struct work;
struct rcu_head rcu;
};
- struct bpf_stream stream[2];
+ struct bpf_stream *stream[2];
struct mutex st_ops_assoc_mutex;
struct bpf_map __rcu *st_ops_assoc;
};
@@ -4223,9 +4230,10 @@ void bpf_bprintf_cleanup(struct bpf_bprintf_data *data);
int bpf_try_get_buffers(struct bpf_bprintf_buffers **bufs);
void bpf_put_buffers(void);
-void bpf_prog_stream_init(struct bpf_prog *prog);
+int bpf_prog_stream_init(struct bpf_prog *prog, gfp_t gfp_extra_flags);
void bpf_prog_stream_free(struct bpf_prog *prog);
int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len);
+int bpf_prog_stream_new_fd(struct bpf_prog *prog, enum bpf_stream_id stream_id, u32 flags);
void bpf_stream_stage_init(struct bpf_stream_stage *ss);
void bpf_stream_stage_free(struct bpf_stream_stage *ss);
__printf(2, 3)
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 0aaa54359aebc..4687c33109968 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -936,6 +936,33 @@ union bpf_iter_link_info {
* 0 on success or -1 if an error occurred (in which case,
* *errno* is set appropriately).
*
+ * BPF_PROG_STREAM_OPEN
+ * Description
+ * Open a file descriptor for one of the BPF streams associated
+ * with the program identified by *prog_fd*. The stream is selected
+ * by *stream_id*.
+ *
+ * The returned file descriptor supports **read**\ (2) and
+ * **poll**\ (2). Reads block while the stream is empty unless
+ * **BPF_F_STREAM_NONBLOCK** is specified in *flags*. A non-blocking
+ * read of an empty stream fails with **EAGAIN**.
+ *
+ * **poll**\ (2) reports **POLLIN** when data is available and
+ * **POLLHUP** once the program has been freed, that is, after every
+ * reference to it, including links and other file descriptors, has
+ * been dropped. Hangup may lag the final release because program
+ * teardown is deferred. Buffered data remains readable after
+ * **POLLHUP** and a read returns zero after all such data has been
+ * consumed.
+ *
+ * The file descriptor is read-only and has the close-on-exec flag
+ * set. It is not seekable and **lseek**\ (2) fails with **ESPIPE**.
+ * *flags* may only contain **BPF_F_STREAM_NONBLOCK**.
+ *
+ * Return
+ * A new file descriptor (a nonnegative integer), or -1 if an
+ * error occurred (in which case, *errno* is set appropriately).
+ *
* NOTES
* eBPF objects (maps and programs) can be shared between processes.
*
@@ -993,6 +1020,7 @@ enum bpf_cmd {
BPF_TOKEN_CREATE,
BPF_PROG_STREAM_READ_BY_FD,
BPF_PROG_ASSOC_STRUCT_OPS,
+ BPF_PROG_STREAM_OPEN,
__MAX_BPF_CMD,
BPF_COMMON_ATTRS = 1 << 16, /* Indicate carrying syscall common attrs. */
};
@@ -1524,6 +1552,11 @@ enum {
BPF_STREAM_STDERR = 2,
};
+/* flags for BPF_PROG_STREAM_OPEN command */
+enum {
+ BPF_F_STREAM_NONBLOCK = (1U << 0),
+};
+
union bpf_attr {
struct { /* anonymous struct used by BPF_MAP_CREATE command */
__u32 map_type; /* one of enum bpf_map_type */
@@ -1950,6 +1983,12 @@ union bpf_attr {
__u32 flags;
} prog_assoc_struct_ops;
+ struct {
+ __u32 prog_fd;
+ __u32 stream_id;
+ __u32 flags;
+ } prog_stream_open;
+
} __attribute__((aligned(8)));
/* The description below is an attempt at providing documentation to eBPF
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 273f74068068c..134700f425319 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -141,10 +141,6 @@ struct bpf_prog *bpf_prog_alloc_no_stats(unsigned int size, gfp_t gfp_extra_flag
mutex_init(&fp->aux->dst_mutex);
mutex_init(&fp->aux->st_ops_assoc_mutex);
-#ifdef CONFIG_BPF_SYSCALL
- bpf_prog_stream_init(fp);
-#endif
-
return fp;
}
@@ -288,6 +284,9 @@ struct bpf_prog *bpf_prog_realloc(struct bpf_prog *fp_old, unsigned int size,
void __bpf_prog_free(struct bpf_prog *fp)
{
if (fp->aux) {
+#ifdef CONFIG_BPF_SYSCALL
+ bpf_prog_stream_free(fp);
+#endif
mutex_destroy(&fp->aux->used_maps_mutex);
mutex_destroy(&fp->aux->dst_mutex);
mutex_destroy(&fp->aux->st_ops_assoc_mutex);
@@ -3073,7 +3072,6 @@ static void bpf_prog_free_deferred(struct work_struct *work)
aux = container_of(work, struct bpf_prog_aux, work);
#ifdef CONFIG_BPF_SYSCALL
bpf_free_kfunc_btf_tab(aux->kfunc_btf_tab);
- bpf_prog_stream_free(aux->prog);
#endif
#ifdef CONFIG_CGROUP_BPF
if (aux->cgroup_atype != CGROUP_BPF_ATTACH_TYPE_INVALID)
diff --git a/kernel/bpf/stream.c b/kernel/bpf/stream.c
index 2b80a0599865e..e0ce37fcd5578 100644
--- a/kernel/bpf/stream.c
+++ b/kernel/bpf/stream.c
@@ -2,11 +2,15 @@
/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
#include <linux/bpf.h>
+#include <linux/anon_inodes.h>
#include <linux/filter.h>
#include <linux/bpf_mem_alloc.h>
#include <linux/gfp.h>
+#include <linux/irq_work.h>
#include <linux/memory.h>
#include <linux/mutex.h>
+#include <linux/poll.h>
+#include <linux/refcount.h>
static void bpf_stream_elem_init(struct bpf_stream_elem *elem, int len)
{
@@ -73,16 +77,59 @@ static void bpf_stream_release_capacity(struct bpf_stream *stream, int len)
atomic_sub(len, &stream->capacity);
}
+static void bpf_stream_notify(struct irq_work *work)
+{
+ struct bpf_stream *stream = container_of(work, struct bpf_stream, notify_work);
+
+ /*
+ * Writers run in arbitrary program contexts, including NMI and regions
+ * that already hold wait queue or epoll locks. Wake waiters from
+ * irq_work instead, where taking those locks is safe.
+ */
+ wake_up_interruptible_poll(&stream->waitq, EPOLLIN | EPOLLRDNORM);
+}
+
+static void bpf_stream_queue_notify(struct bpf_stream *stream)
+{
+ /*
+ * Record that the work has been used so that teardown only pays for
+ * irq_work_sync(), which may wait for an RCU grace period, when a
+ * callback could actually be in flight.
+ */
+ if (!READ_ONCE(stream->notify_used))
+ WRITE_ONCE(stream->notify_used, true);
+ irq_work_queue(&stream->notify_work);
+}
+
+static int bpf_stream_readable_bytes(struct bpf_stream *stream)
+{
+ return atomic_read_acquire(&stream->readable);
+}
+
+static void bpf_stream_publish(struct bpf_stream *stream, int len)
+{
+ /* Pairs with atomic_read_acquire() in bpf_stream_readable_bytes(). */
+ (void)atomic_add_return_release(len, &stream->readable);
+ bpf_stream_queue_notify(stream);
+}
+
static int bpf_stream_push_str(struct bpf_stream *stream, const char *str, int len)
{
- int ret = bpf_stream_consume_capacity(stream, len);
+ int ret;
+
+ /* Nothing to publish; do not allocate an element for it. */
+ if (!len)
+ return 0;
+ ret = bpf_stream_consume_capacity(stream, len);
if (ret)
return ret;
ret = __bpf_stream_push_str(&stream->log, str, len);
if (ret)
bpf_stream_release_capacity(stream, len);
+ else
+ bpf_stream_publish(stream, len);
return ret;
}
@@ -91,7 +138,7 @@ static struct bpf_stream *bpf_stream_get(enum bpf_stream_id stream_id, struct bp
{
if (stream_id != BPF_STDOUT && stream_id != BPF_STDERR)
return NULL;
- return &aux->stream[stream_id - 1];
+ return aux->stream[stream_id - 1];
}
static void bpf_stream_free_elem(struct bpf_stream_elem *elem)
@@ -159,14 +206,16 @@ static bool bpf_stream_consume_elem(struct bpf_stream_elem *elem, int *len)
static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)
{
- int rem_len = len, cons_len, ret = 0;
+ int read_len, rem_len, cons_len, ret = 0;
struct bpf_stream_elem *elem = NULL;
struct llist_node *node;
mutex_lock(&stream->lock);
+ read_len = min(len, bpf_stream_readable_bytes(stream));
+ rem_len = read_len;
while (rem_len) {
- int pos = len - rem_len;
+ int pos = read_len - rem_len;
int chunk, n;
bool cont;
@@ -188,7 +237,7 @@ static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)
/* Keep any successfully copied bytes; -EFAULT only if none. */
elem->consumed_len -= n;
rem_len += n;
- ret = (len == rem_len) ? -EFAULT : 0;
+ ret = (read_len == rem_len) ? -EFAULT : 0;
break;
}
@@ -199,8 +248,9 @@ static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)
bpf_stream_free_elem(elem);
}
+ atomic_sub(read_len - rem_len, &stream->readable);
mutex_unlock(&stream->lock);
- return ret ? ret : len - rem_len;
+ return ret ? ret : read_len - rem_len;
}
int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len)
@@ -215,6 +265,111 @@ int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, vo
return bpf_stream_read(stream, buf, len);
}
+static bool bpf_stream_has_data(struct bpf_stream *stream)
+{
+ return bpf_stream_readable_bytes(stream) > 0;
+}
+
+static void bpf_stream_put(struct bpf_stream *stream)
+{
+ if (refcount_dec_and_test(&stream->refcnt)) {
+ struct llist_node *list;
+
+ /* Only a stream that ever queued its work can have a callback in flight. */
+ if (READ_ONCE(stream->notify_used))
+ irq_work_sync(&stream->notify_work);
+ list = llist_del_all(&stream->log);
+ bpf_stream_free_list(list);
+ bpf_stream_free_list(stream->backlog_head);
+ mutex_destroy(&stream->lock);
+ kfree(stream);
+ }
+}
+
+static int bpf_stream_release(struct inode *inode, struct file *file)
+{
+ bpf_stream_put(file->private_data);
+ return 0;
+}
+
+static ssize_t bpf_stream_file_read(struct file *file, char __user *buf, size_t len,
+ loff_t *ppos)
+{
+ struct bpf_stream *stream = file->private_data;
+ bool dead;
+ int ret;
+
+ if (!len)
+ return 0;
+
+ for (;;) {
+ /*
+ * Sample teardown state before looking for data. Nothing is
+ * published once the program is gone, so finding the stream
+ * empty after observing dead means EOF. The opposite order could
+ * report EOF while data published just before teardown is still
+ * buffered.
+ */
+ dead = smp_load_acquire(&stream->dead);
+ ret = bpf_stream_read(stream, buf, len);
+ if (ret)
+ return ret;
+ if (dead)
+ return 0;
+ if (file->f_flags & O_NONBLOCK)
+ return -EAGAIN;
+
+ ret = wait_event_interruptible(stream->waitq,
+ bpf_stream_has_data(stream) ||
+ READ_ONCE(stream->dead));
+ if (ret)
+ return ret;
+ }
+}
+
+static __poll_t bpf_stream_poll(struct file *file, struct poll_table_struct *pts)
+{
+ struct bpf_stream *stream = file->private_data;
+ __poll_t events = 0;
+
+ /*
+ * poll_wait() only registers the wait queue callback. Register before
+ * checking persistent state so a concurrent publication or teardown is
+ * observed either by the callback or by the checks below.
+ */
+ poll_wait(file, &stream->waitq, pts);
+ if (bpf_stream_has_data(stream))
+ events |= EPOLLIN | EPOLLRDNORM;
+ if (READ_ONCE(stream->dead))
+ events |= EPOLLHUP;
+ return events;
+}
+
+static const struct file_operations bpf_stream_fops = {
+ .release = bpf_stream_release,
+ .read = bpf_stream_file_read,
+ .poll = bpf_stream_poll,
+};
+
+int bpf_prog_stream_new_fd(struct bpf_prog *prog, enum bpf_stream_id stream_id, u32 flags)
+{
+ struct bpf_stream *stream;
+ int fd_flags = O_RDONLY | O_CLOEXEC;
+ int fd;
+
+ stream = bpf_stream_get(stream_id, prog->aux);
+ if (!stream)
+ return -ENOENT;
+ if (flags & BPF_F_STREAM_NONBLOCK)
+ fd_flags |= O_NONBLOCK;
+
+ refcount_inc(&stream->refcnt);
+ fd = anon_inode_getfd("bpf-stream", &bpf_stream_fops, stream, fd_flags);
+ if (fd < 0)
+ bpf_stream_put(stream);
+ return fd;
+}
+
__bpf_kfunc_start_defs();
/*
@@ -282,28 +437,47 @@ __bpf_kfunc_end_defs();
/* Added kfunc to common_btf_ids */
-void bpf_prog_stream_init(struct bpf_prog *prog)
+int bpf_prog_stream_init(struct bpf_prog *prog, gfp_t gfp_extra_flags)
{
int i;
for (i = 0; i < ARRAY_SIZE(prog->aux->stream); i++) {
- atomic_set(&prog->aux->stream[i].capacity, 0);
- init_llist_head(&prog->aux->stream[i].log);
- mutex_init(&prog->aux->stream[i].lock);
- prog->aux->stream[i].backlog_head = NULL;
- prog->aux->stream[i].backlog_tail = NULL;
+ struct bpf_stream *stream;
+
+ /* On failure, bpf_prog_stream_free() releases the streams allocated so far. */
+ stream = kzalloc_obj(*stream,
+ bpf_memcg_flags(GFP_KERNEL | gfp_extra_flags));
+ if (!stream)
+ return -ENOMEM;
+
+ refcount_set(&stream->refcnt, 1);
+ init_llist_head(&stream->log);
+ mutex_init(&stream->lock);
+ init_waitqueue_head(&stream->waitq);
+ init_irq_work(&stream->notify_work, bpf_stream_notify);
+ prog->aux->stream[i] = stream;
}
+ return 0;
}
void bpf_prog_stream_free(struct bpf_prog *prog)
{
- struct llist_node *list;
int i;
for (i = 0; i < ARRAY_SIZE(prog->aux->stream); i++) {
- list = llist_del_all(&prog->aux->stream[i].log);
- bpf_stream_free_list(list);
- bpf_stream_free_list(prog->aux->stream[i].backlog_head);
+ struct bpf_stream *stream = prog->aux->stream[i];
+
+ if (!stream)
+ continue;
+ /*
+ * Pairs with smp_load_acquire() in bpf_stream_file_read(): every
+ * publication precedes the dead flag, so a reader that observes
+ * it also observes all buffered data.
+ */
+ smp_store_release(&stream->dead, true);
+ wake_up_interruptible_poll(&stream->waitq, EPOLLHUP);
+ bpf_stream_put(stream);
+ prog->aux->stream[i] = NULL;
}
}
@@ -333,8 +507,8 @@ int bpf_stream_stage_printk(struct bpf_stream_stage *ss, const char *fmt, ...)
va_start(args, fmt);
len = vscnprintf(buf->buf, ARRAY_SIZE(buf->buf), fmt, args);
va_end(args);
- /* Exclude NULL byte during push. */
- ret = __bpf_stream_push_str(&ss->log, buf->buf, len);
+ /* Exclude NULL byte during push; skip empty output entirely. */
+ ret = len ? __bpf_stream_push_str(&ss->log, buf->buf, len) : 0;
if (!ret)
ss->len += len;
bpf_put_buffers();
@@ -366,6 +540,7 @@ int bpf_stream_stage_commit(struct bpf_stream_stage *ss, struct bpf_prog *prog,
list = tail;
}
llist_add_batch(head, tail, &stream->log);
+ bpf_stream_publish(stream, ss->len);
return 0;
}
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index 74496fd716d3b..8580f38b41c16 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -3080,6 +3080,10 @@ static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_at
prog->aux->user = get_current_user();
prog->len = attr->insn_cnt;
+ err = bpf_prog_stream_init(prog, GFP_USER);
+ if (err)
+ goto free_prog;
+
err = -EFAULT;
if (copy_from_bpfptr(prog->insns,
make_bpfptr(attr->insns, uattr.is_kernel),
@@ -6313,6 +6317,28 @@ static int prog_assoc_struct_ops(union bpf_attr *attr)
return ret;
}
+#define BPF_PROG_STREAM_OPEN_LAST_FIELD prog_stream_open.flags
+
+static int prog_stream_open(union bpf_attr *attr)
+{
+ struct bpf_prog *prog;
+ u32 flags = attr->prog_stream_open.flags;
+ int ret;
+
+ if (CHECK_ATTR(BPF_PROG_STREAM_OPEN))
+ return -EINVAL;
+ if (flags & ~BPF_F_STREAM_NONBLOCK)
+ return -EINVAL;
+
+ prog = bpf_prog_get(attr->prog_stream_open.prog_fd);
+ if (IS_ERR(prog))
+ return PTR_ERR(prog);
+
+ ret = bpf_prog_stream_new_fd(prog, attr->prog_stream_open.stream_id, flags);
+ bpf_prog_put(prog);
+ return ret;
+}
+
static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,
bpfptr_t uattr_common, unsigned int size_common)
{
@@ -6485,6 +6511,9 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,
case BPF_PROG_ASSOC_STRUCT_OPS:
err = prog_assoc_struct_ops(&attr);
break;
+ case BPF_PROG_STREAM_OPEN:
+ err = prog_stream_open(&attr);
+ break;
default:
err = -EINVAL;
break;
diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
index 90fe8c61bf424..ef7e00eb5e479 100644
--- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
+++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
@@ -186,6 +186,11 @@ bpftool prog tracelog { stdout | stderr } *PROG*
error messages to the standard error stream. This facility should be used
only for debugging purposes.
+ On kernels that support opening a stream as a file descriptor, bpftool
+ keeps printing new output as the program produces it, until the program is
+ unloaded or <Ctrl+C> is hit. Older kernels only allow dumping the output
+ buffered so far, after which bpftool exits.
+
bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ctx_in *FILE* [ctx_out *FILE* [ctx_size_out *M*]]] [repeat *N*]
Run BPF program *PROG* in the kernel testing infrastructure for BPF,
meaning that the program works on the data and context provided by the
diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
index 24e40dfab4690..f0241dada5485 100644
--- a/tools/bpf/bpftool/prog.c
+++ b/tools/bpf/bpftool/prog.c
@@ -1119,21 +1119,69 @@ enum prog_tracelog_mode {
TRACE_STDERR,
};
+static volatile sig_atomic_t stream_stop;
+
+static void stop_stream(int signo)
+{
+ stream_stop = 1;
+}
+
+/* Consumes prog_fd. */
static int
prog_tracelog_stream(int prog_fd, enum prog_tracelog_mode mode)
{
+ /* No SA_RESTART: an interrupted read() must return EINTR to end the loop. */
+ const struct sigaction act = { .sa_handler = stop_stream };
+ const int signals[] = { SIGHUP, SIGINT, SIGTERM };
+ struct sigaction old[ARRAY_SIZE(signals)];
FILE *file = mode == TRACE_STDOUT ? stdout : stderr;
int stream_id = mode == TRACE_STDOUT ? 1 : 2;
char buf[512];
- int ret;
+ unsigned int i;
+ int fd, ret;
+
+ fd = bpf_prog_stream_open(prog_fd, stream_id, NULL);
+ if (fd == -EINVAL) {
+ /* Kernel predates BPF_PROG_STREAM_OPEN: dump buffered output and exit. */
+ do {
+ ret = bpf_prog_stream_read(prog_fd, stream_id, buf, sizeof(buf), NULL);
+ if (ret > 0)
+ fwrite(buf, sizeof(buf[0]), ret, file);
+ } while (ret > 0);
+ if (ret < 0)
+ p_err("failed to read stream: %s", strerror(-ret));
+ close(prog_fd);
+ goto out;
+ }
+ /*
+ * The stream descriptor does not keep the program alive. Drop the
+ * program reference so that reads return EOF once the program is gone.
+ */
+ close(prog_fd);
+ if (fd < 0) {
+ p_err("failed to open stream: %s", strerror(-fd));
+ return -1;
+ }
+ stream_stop = 0;
+ for (i = 0; i < ARRAY_SIZE(signals); i++)
+ sigaction(signals[i], &act, &old[i]);
ret = 0;
- do {
- ret = bpf_prog_stream_read(prog_fd, stream_id, buf, sizeof(buf), NULL);
- if (ret > 0)
- fwrite(buf, sizeof(buf[0]), ret, file);
- } while (ret > 0);
-
+ while (!stream_stop) {
+ ret = read(fd, buf, sizeof(buf));
+ if (ret <= 0)
+ break;
+ fwrite(buf, sizeof(buf[0]), ret, file);
+ fflush(file);
+ }
+ if (ret < 0 && !(stream_stop && errno == EINTR))
+ p_err("failed to read stream: %s", strerror(errno));
+ else
+ ret = 0;
+ for (i = 0; i < ARRAY_SIZE(signals); i++)
+ sigaction(signals[i], &old[i], NULL);
+ close(fd);
+out:
fflush(file);
return ret ? -1 : 0;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 0aaa54359aebc..4687c33109968 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -936,6 +936,33 @@ union bpf_iter_link_info {
* 0 on success or -1 if an error occurred (in which case,
* *errno* is set appropriately).
*
+ * BPF_PROG_STREAM_OPEN
+ * Description
+ * Open a file descriptor for one of the BPF streams associated
+ * with the program identified by *prog_fd*. The stream is selected
+ * by *stream_id*.
+ *
+ * The returned file descriptor supports **read**\ (2) and
+ * **poll**\ (2). Reads block while the stream is empty unless
+ * **BPF_F_STREAM_NONBLOCK** is specified in *flags*. A non-blocking
+ * read of an empty stream fails with **EAGAIN**.
+ *
+ * **poll**\ (2) reports **POLLIN** when data is available and
+ * **POLLHUP** once the program has been freed, that is, after every
+ * reference to it, including links and other file descriptors, has
+ * been dropped. Hangup may lag the final release because program
+ * teardown is deferred. Buffered data remains readable after
+ * **POLLHUP** and a read returns zero after all such data has been
+ * consumed.
+ *
+ * The file descriptor is read-only and has the close-on-exec flag
+ * set. It is not seekable and **lseek**\ (2) fails with **ESPIPE**.
+ * *flags* may only contain **BPF_F_STREAM_NONBLOCK**.
+ *
+ * Return
+ * A new file descriptor (a nonnegative integer), or -1 if an
+ * error occurred (in which case, *errno* is set appropriately).
+ *
* NOTES
* eBPF objects (maps and programs) can be shared between processes.
*
@@ -993,6 +1020,7 @@ enum bpf_cmd {
BPF_TOKEN_CREATE,
BPF_PROG_STREAM_READ_BY_FD,
BPF_PROG_ASSOC_STRUCT_OPS,
+ BPF_PROG_STREAM_OPEN,
__MAX_BPF_CMD,
BPF_COMMON_ATTRS = 1 << 16, /* Indicate carrying syscall common attrs. */
};
@@ -1524,6 +1552,11 @@ enum {
BPF_STREAM_STDERR = 2,
};
+/* flags for BPF_PROG_STREAM_OPEN command */
+enum {
+ BPF_F_STREAM_NONBLOCK = (1U << 0),
+};
+
union bpf_attr {
struct { /* anonymous struct used by BPF_MAP_CREATE command */
__u32 map_type; /* one of enum bpf_map_type */
@@ -1950,6 +1983,12 @@ union bpf_attr {
__u32 flags;
} prog_assoc_struct_ops;
+ struct {
+ __u32 prog_fd;
+ __u32 stream_id;
+ __u32 flags;
+ } prog_stream_open;
+
} __attribute__((aligned(8)));
/* The description below is an attempt at providing documentation to eBPF
diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
index a9de7f107cf7b..b49822d212aed 100644
--- a/tools/lib/bpf/bpf.c
+++ b/tools/lib/bpf/bpf.c
@@ -1466,6 +1466,25 @@ int bpf_prog_stream_read(int prog_fd, __u32 stream_id, void *buf, __u32 buf_len,
return libbpf_err_errno(err);
}
+int bpf_prog_stream_open(int prog_fd, __u32 stream_id,
+ const struct bpf_prog_stream_open_opts *opts)
+{
+ const size_t attr_sz = offsetofend(union bpf_attr, prog_stream_open);
+ union bpf_attr attr;
+ int fd;
+
+ if (!OPTS_VALID(opts, bpf_prog_stream_open_opts))
+ return libbpf_err(-EINVAL);
+
+ memset(&attr, 0, attr_sz);
+ attr.prog_stream_open.prog_fd = prog_fd;
+ attr.prog_stream_open.stream_id = stream_id;
+ attr.prog_stream_open.flags = OPTS_GET(opts, flags, 0);
+
+ fd = sys_bpf_fd(BPF_PROG_STREAM_OPEN, &attr, attr_sz);
+ return libbpf_err_errno(fd);
+}
+
int bpf_prog_assoc_struct_ops(int prog_fd, int map_fd,
struct bpf_prog_assoc_struct_ops_opts *opts)
{
diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
index 490e8cb4ba537..826d9cc9ab65d 100644
--- a/tools/lib/bpf/bpf.h
+++ b/tools/lib/bpf/bpf.h
@@ -759,10 +759,34 @@ struct bpf_prog_stream_read_opts {
*
* @return The number of bytes read, on success; negative error code, otherwise
* (errno is also set to the error code)
+ *
+ * For blocking reads and polling, prefer **bpf_prog_stream_open**.
*/
LIBBPF_API int bpf_prog_stream_read(int prog_fd, __u32 stream_id, void *buf, __u32 buf_len,
struct bpf_prog_stream_read_opts *opts);
+struct bpf_prog_stream_open_opts {
+ size_t sz;
+ __u32 flags;
+ size_t :0;
+};
+#define bpf_prog_stream_open_opts__last_field flags
+
+/**
+ * @brief **bpf_prog_stream_open** opens a file descriptor for a BPF stream of
+ * a given BPF program.
+ *
+ * @param prog_fd FD for the BPF program whose BPF stream is to be opened.
+ * @param stream_id ID of the BPF stream to be opened.
+ * @param opts optional options, can be NULL. BPF_F_STREAM_NONBLOCK requests a
+ * non-blocking descriptor.
+ *
+ * @return A new stream FD, on success; negative error code, otherwise (errno
+ * is also set to the error code)
+ */
+LIBBPF_API int bpf_prog_stream_open(int prog_fd, __u32 stream_id,
+ const struct bpf_prog_stream_open_opts *opts);
+
struct bpf_prog_assoc_struct_ops_opts {
size_t sz;
__u32 flags;
diff --git a/tools/lib/bpf/libbpf.map b/tools/lib/bpf/libbpf.map
index 18d27d20102ec..53b7b591325f5 100644
--- a/tools/lib/bpf/libbpf.map
+++ b/tools/lib/bpf/libbpf.map
@@ -459,6 +459,7 @@ LIBBPF_1.7.0 {
LIBBPF_1.8.0 {
global:
bpf_map__attach_cgroup_opts;
+ bpf_prog_stream_open;
bpf_program__add_flags;
bpf_program__attach_tracing_multi;
bpf_program__clear_flags;
diff --git a/tools/testing/selftests/bpf/prog_tests/stream.c b/tools/testing/selftests/bpf/prog_tests/stream.c
index 74bd15c4bfbaa..2c6b441552862 100644
--- a/tools/testing/selftests/bpf/prog_tests/stream.c
+++ b/tools/testing/selftests/bpf/prog_tests/stream.c
@@ -1,11 +1,17 @@
// SPDX-License-Identifier: GPL-2.0
/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
#include <test_progs.h>
+#include <linux/perf_event.h>
+#include <poll.h>
+#include <sys/epoll.h>
#include <sys/mman.h>
+#include <sys/syscall.h>
#include "stream.skel.h"
#include "stream_fail.skel.h"
+#define NMI_TIMEOUT_NS (5ULL * 1000 * 1000 * 1000)
+
void test_stream_failure(void)
{
RUN_TESTS(stream_fail);
@@ -62,6 +68,415 @@ void test_stream_syscall(void)
stream__destroy(skel);
}
+static bool stream_fd_trigger(struct bpf_program *prog)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ int ret;
+
+ ret = bpf_prog_test_run_opts(bpf_program__fd(prog), &opts);
+ return ASSERT_OK(ret, "test_run") && ASSERT_OK(opts.retval, "retval");
+}
+
+static void test_stream_fd_open(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts);
+ struct stream *skel;
+ int fd, prog_fd;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ prog_fd = bpf_program__fd(skel->progs.stream_syscall);
+ fd = bpf_prog_stream_open(0, BPF_STREAM_STDOUT, NULL);
+ ASSERT_EQ(fd, -EINVAL, "bad_prog_fd");
+
+ fd = bpf_prog_stream_open(prog_fd, 0, NULL);
+ ASSERT_EQ(fd, -ENOENT, "bad_stream_id");
+
+ opts.flags = BPF_F_RDONLY;
+ fd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDOUT, &opts);
+ ASSERT_EQ(fd, -EINVAL, "access_flag");
+
+ opts.flags = 1U << 31;
+ fd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDOUT, &opts);
+ ASSERT_EQ(fd, -EINVAL, "unknown_flag");
+
+ fd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDERR, NULL);
+ if (ASSERT_OK_FD(fd, "stderr"))
+ close(fd);
+
+ stream__destroy(skel);
+}
+
+static void test_stream_fd_nonblock(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts,
+ .flags = BPF_F_STREAM_NONBLOCK,
+ );
+ struct pollfd pfd = { .events = POLLIN | POLLHUP };
+ struct stream *skel;
+ char buf[4] = {};
+ int fd, flags, ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_syscall),
+ BPF_STREAM_STDOUT, &opts);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out_destroy;
+ pfd.fd = fd;
+
+ flags = fcntl(fd, F_GETFD);
+ ASSERT_GE(flags, 0, "getfd");
+ ASSERT_NEQ(flags & FD_CLOEXEC, 0, "cloexec");
+ flags = fcntl(fd, F_GETFL);
+ ASSERT_GE(flags, 0, "getfl");
+ ASSERT_EQ(flags & O_ACCMODE, O_RDONLY, "readonly");
+ ASSERT_NEQ(flags & O_NONBLOCK, 0, "nonblock");
+
+ ret = write(fd, "x", 1);
+ ASSERT_EQ(ret, -1, "write");
+ ASSERT_EQ(errno, EBADF, "write_errno");
+ ret = lseek(fd, 0, SEEK_SET);
+ ASSERT_EQ(ret, -1, "lseek");
+ ASSERT_EQ(errno, ESPIPE, "lseek_errno");
+
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 0, "poll_empty");
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, -1, "read_empty");
+ ASSERT_EQ(errno, EAGAIN, "read_empty_errno");
+
+ if (!stream_fd_trigger(skel->progs.stream_syscall))
+ goto out_close;
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 1, "poll_data");
+ ASSERT_NEQ(pfd.revents & POLLIN, 0, "pollin");
+ ASSERT_EQ(pfd.revents & POLLHUP, 0, "no_pollhup");
+
+ ret = read(fd, buf, 2);
+ ASSERT_EQ(ret, 2, "read_first");
+ ASSERT_OK(memcmp(buf, "fo", 2), "read_first_data");
+ pfd.revents = 0;
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 1, "poll_partial");
+ ASSERT_NEQ(pfd.revents & POLLIN, 0, "pollin_partial");
+
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, 1, "read_rest");
+ ASSERT_EQ(buf[0], 'o', "read_rest_data");
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, -1, "read_drained");
+ ASSERT_EQ(errno, EAGAIN, "read_drained_errno");
+
+out_close:
+ close(fd);
+out_destroy:
+ stream__destroy(skel);
+}
+
+static void test_stream_fd_empty(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts,
+ .flags = BPF_F_STREAM_NONBLOCK,
+ );
+ struct pollfd pfd = { .events = POLLIN | POLLHUP };
+ struct stream *skel;
+ char buf[4];
+ int fd, ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_empty),
+ BPF_STREAM_STDOUT, &opts);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out_destroy;
+ pfd.fd = fd;
+
+ /* Empty output produces neither data nor readiness. */
+ if (!stream_fd_trigger(skel->progs.stream_empty))
+ goto out_close;
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 0, "poll_empty_write");
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, -1, "read_empty_write");
+ ASSERT_EQ(errno, EAGAIN, "read_empty_write_errno");
+
+out_close:
+ close(fd);
+out_destroy:
+ stream__destroy(skel);
+}
+
+struct stream_fd_read_ctx {
+ int fd;
+ ssize_t ret;
+ char buf[4];
+};
+
+static void *stream_fd_read_thread(void *arg)
+{
+ struct stream_fd_read_ctx *ctx = arg;
+
+ ctx->ret = read(ctx->fd, ctx->buf, sizeof(ctx->buf));
+ return NULL;
+}
+
+static int stream_fd_timed_join(pthread_t thread)
+{
+ struct timespec timeout;
+
+ clock_gettime(CLOCK_REALTIME, &timeout);
+ timeout.tv_sec += 5;
+ return pthread_timedjoin_np(thread, NULL, &timeout);
+}
+
+static void test_stream_fd_blocking(void)
+{
+ struct stream_fd_read_ctx ctx = {};
+ struct stream *skel;
+ pthread_t thread;
+ bool thread_live = false;
+ int fd, flags, ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_syscall),
+ BPF_STREAM_STDOUT, NULL);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out_destroy;
+ ctx.fd = fd;
+
+ flags = fcntl(fd, F_GETFL);
+ ASSERT_GE(flags, 0, "getfl");
+ ASSERT_EQ(flags & O_NONBLOCK, 0, "blocking");
+
+ ret = pthread_create(&thread, NULL, stream_fd_read_thread, &ctx);
+ if (!ASSERT_OK(ret, "pthread_create"))
+ goto out_close;
+ thread_live = true;
+
+ usleep(50000);
+ ret = pthread_tryjoin_np(thread, NULL);
+ if (!ASSERT_EQ(ret, EBUSY, "read_blocks")) {
+ thread_live = ret != 0;
+ goto out_thread;
+ }
+ if (!stream_fd_trigger(skel->progs.stream_syscall))
+ goto out_thread;
+
+ ret = stream_fd_timed_join(thread);
+ if (!ASSERT_OK(ret, "pthread_join"))
+ goto out_thread;
+ thread_live = false;
+ ASSERT_EQ(ctx.ret, 3, "read_len");
+ ASSERT_OK(memcmp(ctx.buf, "foo", 3), "read_data");
+
+ memset(&ctx, 0, sizeof(ctx));
+ ctx.fd = fd;
+ ret = pthread_create(&thread, NULL, stream_fd_read_thread, &ctx);
+ if (!ASSERT_OK(ret, "pthread_create_eof"))
+ goto out_close;
+ thread_live = true;
+
+ usleep(50000);
+ ret = pthread_tryjoin_np(thread, NULL);
+ if (!ASSERT_EQ(ret, EBUSY, "read_eof_blocks")) {
+ thread_live = ret != 0;
+ goto out_thread;
+ }
+
+ stream__destroy(skel);
+ skel = NULL;
+ ret = stream_fd_timed_join(thread);
+ if (!ASSERT_OK(ret, "pthread_join_eof"))
+ goto out_thread;
+ thread_live = false;
+ ASSERT_EQ(ctx.ret, 0, "read_eof");
+
+out_thread:
+ if (thread_live) {
+ stream__destroy(skel);
+ skel = NULL;
+ ret = stream_fd_timed_join(thread);
+ if (ret) {
+ pthread_cancel(thread);
+ pthread_join(thread, NULL);
+ }
+ }
+out_close:
+ close(fd);
+out_destroy:
+ stream__destroy(skel);
+}
+
+static void test_stream_fd_hup(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts,
+ .flags = BPF_F_STREAM_NONBLOCK,
+ );
+ struct epoll_event event = {
+ .events = EPOLLIN | EPOLLET,
+ };
+ struct pollfd pfd = { .events = POLLIN | POLLHUP };
+ struct stream *skel;
+ char buf[4] = {};
+ int epfd = -1, fd, ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_syscall),
+ BPF_STREAM_STDOUT, &opts);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out_destroy;
+ pfd.fd = fd;
+ epfd = epoll_create1(EPOLL_CLOEXEC);
+ if (!ASSERT_OK_FD(epfd, "epoll_create"))
+ goto out_close;
+ event.data.fd = fd;
+ ret = epoll_ctl(epfd, EPOLL_CTL_ADD, fd, &event);
+ if (!ASSERT_OK(ret, "epoll_ctl"))
+ goto out_close;
+
+ if (!stream_fd_trigger(skel->progs.stream_syscall))
+ goto out_close;
+ ret = epoll_wait(epfd, &event, 1, 5000);
+ if (!ASSERT_EQ(ret, 1, "epoll_wait_data"))
+ goto out_close;
+ ASSERT_NEQ(event.events & EPOLLIN, 0, "epollin");
+ ASSERT_EQ(event.events & EPOLLHUP, 0, "no_epollhup");
+
+ stream__destroy(skel);
+ skel = NULL;
+ event.events = 0;
+ ret = epoll_wait(epfd, &event, 1, 5000);
+ if (!ASSERT_EQ(ret, 1, "epoll_wait_hup"))
+ goto out_close;
+ ASSERT_NEQ(event.events & EPOLLIN, 0, "epollin_with_hup");
+ ASSERT_NEQ(event.events & EPOLLHUP, 0, "epollhup");
+
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, 3, "read_buffered");
+ ASSERT_OK(memcmp(buf, "foo", 3), "read_buffered_data");
+ pfd.revents = 0;
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 1, "poll_drained_hup");
+ ASSERT_EQ(pfd.revents & POLLIN, 0, "no_pollin_after_drain");
+ ASSERT_NEQ(pfd.revents & POLLHUP, 0, "pollhup_after_drain");
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, 0, "read_eof");
+
+out_close:
+ if (epfd >= 0)
+ close(epfd);
+ close(fd);
+out_destroy:
+ stream__destroy(skel);
+}
+
+static void test_stream_fd_nmi_epoll(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts,
+ .flags = BPF_F_STREAM_NONBLOCK,
+ );
+ struct perf_event_attr attr = {
+ .size = sizeof(attr),
+ .type = PERF_TYPE_HARDWARE,
+ .config = PERF_COUNT_HW_CPU_CYCLES,
+ .freq = 1,
+ .sample_freq = 10,
+ };
+ struct epoll_event event = {
+ .events = EPOLLIN,
+ };
+ struct bpf_link *link = NULL;
+ struct stream *skel;
+ __u64 deadline;
+ char buf[4] = {};
+ int epfd = -1, fd = -1, pmu_fd = -1;
+ int ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_nmi),
+ BPF_STREAM_STDOUT, &opts);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out;
+
+ epfd = epoll_create1(EPOLL_CLOEXEC);
+ if (!ASSERT_OK_FD(epfd, "epoll_create"))
+ goto out;
+ event.data.fd = fd;
+ ret = epoll_ctl(epfd, EPOLL_CTL_ADD, fd, &event);
+ if (!ASSERT_OK(ret, "epoll_ctl"))
+ goto out;
+
+ pmu_fd = syscall(__NR_perf_event_open, &attr, 0, -1, -1,
+ PERF_FLAG_FD_CLOEXEC);
+ if (pmu_fd < 0 && (errno == ENOENT || errno == EOPNOTSUPP)) {
+ printf("%s:SKIP:no PERF_COUNT_HW_CPU_CYCLES\n", __func__);
+ test__skip();
+ goto out;
+ }
+ if (!ASSERT_GE(pmu_fd, 0, "perf_event_open"))
+ goto out;
+
+ link = bpf_program__attach_perf_event(skel->progs.stream_nmi, pmu_fd);
+ if (!ASSERT_OK_PTR(link, "attach_perf_event")) {
+ link = NULL;
+ goto out;
+ }
+ pmu_fd = -1;
+
+ deadline = get_time_ns() + NMI_TIMEOUT_NS;
+ do {
+ ret = epoll_wait(epfd, &event, 1, 0);
+ } while (ret == 0 && get_time_ns() < deadline);
+ if (!ASSERT_EQ(ret, 1, "epoll_wait"))
+ goto out;
+ ASSERT_NEQ(event.events & EPOLLIN, 0, "epollin");
+
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, 3, "read_len");
+ ASSERT_OK(memcmp(buf, "nmi", 3), "read_data");
+
+out:
+ bpf_link__destroy(link);
+ if (pmu_fd >= 0)
+ close(pmu_fd);
+ if (epfd >= 0)
+ close(epfd);
+ if (fd >= 0)
+ close(fd);
+ stream__destroy(skel);
+}
+
+void test_stream_fd(void)
+{
+ if (test__start_subtest("open"))
+ test_stream_fd_open();
+ if (test__start_subtest("nonblock"))
+ test_stream_fd_nonblock();
+ if (test__start_subtest("empty"))
+ test_stream_fd_empty();
+ if (test__start_subtest("blocking"))
+ test_stream_fd_blocking();
+ if (test__start_subtest("hup"))
+ test_stream_fd_hup();
+ if (test__start_subtest("nmi_epoll"))
+ test_stream_fd_nmi_epoll();
+}
+
void test_stream_oversize(void)
{
LIBBPF_OPTS(bpf_test_run_opts, opts);
diff --git a/tools/testing/selftests/bpf/progs/stream.c b/tools/testing/selftests/bpf/progs/stream.c
index 12fc29e454874..1ae015c4e6b4c 100644
--- a/tools/testing/selftests/bpf/progs/stream.c
+++ b/tools/testing/selftests/bpf/progs/stream.c
@@ -44,6 +44,7 @@ struct {
_X64 _X64 _X64 _X64 _X64 _X64 _X64 _X64)
int size;
+int nmi_stream_prints;
u64 fault_addr;
void *arena_ptr;
@@ -124,6 +125,25 @@ int stream_syscall(void *ctx)
return 0;
}
+SEC("syscall")
+__success __retval(0)
+int stream_empty(void *ctx)
+{
+ return bpf_stream_printk(BPF_STDOUT, "");
+}
+
+SEC("perf_event")
+int stream_nmi(void *ctx)
+{
+ if (nmi_stream_prints)
+ return 0;
+ /* Retry on a later sample if the write failed, e.g. with -EBUSY. */
+ if (bpf_stream_printk(BPF_STDOUT, "nmi"))
+ return 0;
+ nmi_stream_prints = 1;
+ return 0;
+}
+
SEC("syscall")
__success __retval(0)
int stream_oversize(void *ctx)
|
| 3/2 | 2026/09/26 22:18 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 18622 output: 27 thoughts: 457 Thoughts: **My Initial Assessment of the BPF Stream Patch** Okay, so the user wants me to evaluate this patch series that's been committed as HEAD. My immediate thought is, let's dive into the changes. First, I need to understand *what* is being changed. The diff is touching a few key files: `include/linux/bpf.h`, `include/uapi/linux/bpf.h`, then some core BPF code in `kernel/bpf/core.c`, `kernel/bpf/stream.c`, and `kernel/bpf/syscall.c`. There's also some stuff being touched in `tools/...`, which is interesting. Right, so this is introducing a new BPF command, `BPF_PROG_STREAM_OPEN`. It looks like it's for opening file descriptors for BPF streams, specifically for stdout/stderr associated with BPF programs. Okay, so we're talking about a stream FD with standard operations like `read()`, `poll()`, and `close()`. Sounds pretty low-level. This immediately flags a bunch of potential issues. I'm seeing wait queues, `irq_work` (`notify_work`), refcounting on `struct bpf_stream`, `anon_inode` fd creation with `anon_inode_getfd`. The patch is dealing with concurrency between BPF program writers, and user-space readers/pollers, and there is teardown logic with `bpf_stream_release`, `bpf_prog_stream_free`, `bpf_stream_put`. This is all prime real estate for race conditions, use-after-free bugs, refcounting errors, and potential lockdep issues. Subsystem-wise, we're talking about BPF, the syscall handling, and VFS, with the anon_inode subsystem. The important question now: Is this code actually reachable? Absolutely! Anyone with the necessary bpf syscall capability (or, frankly, anyone with BPF enabled) can call `sys_bpf(BPF_PROG_STREAM_OPEN, ...)` and then use the stream. They can read from the fd, poll the fd, triggering all those potential teardown race conditions, use-after-free problems, refcounting problems, and lockdep issues. This is touching core kernel subsystems: BPF, VFS/anon_inode, waitqueues, and irq_work. This is *definitely* `WorthFuzzing = true`. I need to put this on the list for a deep dive and aggressive fuzzing. Now, let's figure out the Kconfig requirements. I'll need to check what configs are needed for `kernel/bpf/syscall.c` and `kernel/bpf/stream.c`. I'll run `grepper` or check the Makefile in `kernel/bpf/` to find those dependencies. |
| 4/2 | 2026/09/26 22:19 | tool | grepper |
0mArgs:{
"Expression": "stream\\.o",
"PathPrefix": "kernel/bpf"
}Results: {
"Output": "kernel/bpf/Makefile=17=obj-$(CONFIG_BPF_JIT) += trampoline.o\nkernel/bpf/Makefile:18:obj-$(CONFIG_BPF_SYSCALL) += btf.o memalloc.o rqspinlock.o stream.o\nkernel/bpf/Makefile-19-ifeq ($(CONFIG_MMU)$(CONFIG_64BIT),yy)\n"
} |
| 5/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 19833 output: 31 thoughts: 37 Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person: **Analyzing the Dependency of stream.o and BPF_SYSCALL** Right, so `stream.o` is compiled, and my immediate thought is that this is conditional on `CONFIG_BPF_SYSCALL` being enabled. Seems logical, given my understanding of the kernel's architecture and the likely role of `stream.o` in handling BPF-related system calls. The code flow is probably set up such that the BPF syscall functionality, and by extension any object files that specifically facilitate it, will not be compiled or linked if that option is disabled. It's a fairly standard approach for modularity and build-time configuration in the kernel. Therefore, the next logical step, before even digging into the code itself, is to confirm that dependency explicitly in the Kconfig files. Time to check the Kconfig for `BPF_SYSCALL` – that's where the compilation logic, the conditional build parameters, and thus the actual relationship between these two pieces of the puzzle, will be defined. It's a quick way to confirm my initial hypothesis, and understand the compile-time dependencies. |
| 6/2 | 2026/09/26 22:19 | tool | grepper |
0mError:no matches Args: {
"Expression": "config BPF_SYSCALL",
"PathPrefix": "init/Kconfig"
}Results: {
"error": "no matches"
} |
| 7/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 20175 output: 21 thoughts: 17 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Investigating the Configuration for BPF_SYSCALL** Alright, let's get down to brass tacks. I need to pinpoint the definition of `config BPF_SYSCALL` in the kernel configuration. I'm thinking I'll need to trace its origins, understand its dependencies, and see exactly *where* this crucial setting is being managed. Given my experience, I'm expecting this will involve spelunking through the Kconfig files, likely diving through the various architecture-specific or feature-specific directories. I’ll be looking for cross-references, conditional statements, and any hints about what other options control its visibility or state. Finding the definitive source of truth is key to understanding its implications within the BPF ecosystem and the kernel as a whole. Time to get digging. |
| 8/2 | 2026/09/26 22:19 | tool | grepper |
0mArgs:{
"Expression": "config BPF_SYSCALL"
}Results: {
"Output": "kernel/bpf/Kconfig=29=menu \"BPF subsystem\"\nkernel/bpf/Kconfig-30-\nkernel/bpf/Kconfig:31:config BPF_SYSCALL\nkernel/bpf/Kconfig-32-\tbool \"Enable bpf() system call\"\n"
} |
| 9/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 20481 output: 34 thoughts: 19 Thoughts: Okay, let's dive into this. Here's my perspective on understanding that `Kconfig` file, geared for an expert like myself: **Diving into the BPF Configuration** Alright, I need to get a handle on how BPF is configured within this kernel. Let's zero in on `kernel/bpf/Kconfig`, specifically around line 31. I'm assuming that's where the core BPF options are defined. My experience tells me that's the likely location for the main `config` statements that control whether BPF features are compiled into the kernel. I'm hoping to quickly identify the crucial options affecting its core functionality. It's likely these options are grouped logically, maybe separated into categories like "core BPF," "BPF helpers," or "BPF tracing." I'm also watching for any dependencies on other kernel subsystems. I want to spot those interconnections early, as that's often where the real complexity (and potential conflicts) lie. Understanding those dependencies is key to effective configuration. Let's see what we've got! |
| 10/2 | 2026/09/26 22:19 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/Kconfig",
"FirstLine": 30,
"LineCount": 40
}Results: {
"Contents": " 30:\t\n 31:\tconfig BPF_SYSCALL\n 32:\t\tbool \"Enable bpf() system call\"\n 33:\t\tselect BPF\n 34:\t\tselect IRQ_WORK\n 35:\t\tselect NEED_TASKS_RCU\n 36:\t\tselect TASKS_TRACE_RCU\n 37:\t\tselect BINARY_PRINTF\n 38:\t\tselect NET_SOCK_MSG if NET\n 39:\t\tselect NET_XGRESS if NET\n 40:\t\tselect PAGE_POOL if NET\n 41:\t\tdefault n\n 42:\t\thelp\n 43:\t\t Enable the bpf() system call that allows to manipulate BPF programs\n 44:\t\t and maps via file descriptors.\n 45:\t\n 46:\tconfig BPF_JIT\n 47:\t\tbool \"Enable BPF Just In Time compiler\"\n 48:\t\tdepends on BPF\n 49:\t\tdepends on HAVE_CBPF_JIT || HAVE_EBPF_JIT\n 50:\t\tselect EXECMEM\n 51:\t\thelp\n 52:\t\t BPF programs are normally handled by a BPF interpreter. This option\n 53:\t\t allows the kernel to generate native code when a program is loaded\n 54:\t\t into the kernel. This will significantly speed-up processing of BPF\n 55:\t\t programs.\n 56:\t\n 57:\t\t Note, an admin should enable this feature changing:\n 58:\t\t /proc/sys/net/core/bpf_jit_enable\n 59:\t\t /proc/sys/net/core/bpf_jit_harden (optional)\n 60:\t\t /proc/sys/net/core/bpf_jit_kallsyms (optional)\n 61:\t\n 62:\tconfig BPF_JIT_ALWAYS_ON\n 63:\t\tbool \"Permanently enable BPF JIT and remove BPF interpreter\"\n 64:\t\tdepends on BPF_SYSCALL \u0026\u0026 HAVE_EBPF_JIT \u0026\u0026 BPF_JIT\n 65:\t\thelp\n 66:\t\t Enables BPF JIT and removes BPF interpreter to avoid speculative\n 67:\t\t execution of BPF instructions by the interpreter.\n 68:\t\n 69:\t\t When CONFIG_BPF_JIT_ALWAYS_ON is enabled, /proc/sys/net/core/bpf_jit_enable\n"
} |
| 11/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 21373 output: 32 thoughts: 1326 Thoughts: **BPF Stream Analysis: Symbol Visibility and Key Functions**
Okay, so `BPF_SYSCALL` seems straightforward, nothing unexpected there. Time to dig deeper and see what we're really dealing with. First thing, let's confirm the obvious: the code is guarded by `#ifdef CONFIG_BPF_SYSCALL`, as expected, and we have the cleanup call `bpf_prog_stream_free(fp)`.
Now, the symbols themselves. The key here, obviously, is understanding which functions are visible in `vmlinux`/kallsyms. I need to pinpoint the entry points and the feature-specific logic, avoiding anything ubiquitous. Compiler inlining... that's always the tricky part, especially with static functions. Let's see what we have in `kernel/bpf/`.
In `syscall.c`, we have `prog_stream_open`, which is called from `__sys_bpf`, and `bpf_prog_stream_init` called from `bpf_prog_load`. Then in `stream.c`, there's a bunch more: `bpf_prog_stream_new_fd`, `bpf_stream_file_read`, `bpf_stream_poll`, `bpf_stream_release`, `bpf_stream_put`, `bpf_stream_notify`, `bpf_stream_queue_notify`, `bpf_stream_publish`, `bpf_prog_stream_init`, `bpf_prog_stream_free`, `bpf_stream_stage_commit`, `bpf_stream_stage_printk`, `bpf_stream_push_str`, `bpf_stream_read`.
Okay, let's sort this. Static vs. non-static, inlined vs. not. I need to identify those "FocusSymbols."
Looking into `kernel/bpf/stream.c`:
* `bpf_prog_stream_new_fd`, `bpf_prog_stream_init`, `bpf_prog_stream_free`, `bpf_stream_stage_commit`, `bpf_stream_stage_printk`, `bpf_prog_stream_read` are *global*, so definitely visible.
* `bpf_stream_file_read`, `bpf_stream_poll`, `bpf_stream_release` are in `bpf_stream_fops`, so no inlining. Definitely symbols.
* `bpf_stream_put`, `bpf_stream_notify` and `bpf_stream_read` are static. The first two are not inlined due to address being taken so they will be visible.
In `kernel/bpf/syscall.c`:
* `prog_stream_open` is static. Potential for inlining into `__sys_bpf`. We need to verify that; it's a `static` function, and sometimes they get inlined in these large switch cases.
Let's test the hypothesis:
* `bpf_prog_stream_new_fd` is defined globally and called from `syscall.c`, so it will definitely be a visible symbol!
* `bpf_stream_file_read`, `bpf_stream_poll`, and `bpf_stream_release` are defined in `bpf_stream_fops`. Their addresses are taken, so they *cannot* be inlined. They'll be there!
* `bpf_stream_notify`'s address is used in `init_irq_work`, so also cannot be inlined.
* `bpf_prog_stream_init` and `bpf_prog_stream_free` are declared in a header and called from both `syscall.c` and `core.c`, so these are non-static and visible, again.
Now, let's see if there are any other direct callers or functions related to streams in `syscall.c`. Need to be sure I haven't missed something.
|
| 12/2 | 2026/09/26 22:19 | tool | grepper |
0mArgs:{
"Expression": "BPF_PROG_STREAM_",
"PathPrefix": "kernel/bpf"
}Results: {
"Output": "kernel/bpf/syscall.c=6243=static int token_create(union bpf_attr *attr)\n--\nkernel/bpf/syscall.c-6254-\nkernel/bpf/syscall.c:6255:#define BPF_PROG_STREAM_READ_BY_FD_LAST_FIELD prog_stream_read.prog_fd\nkernel/bpf/syscall.c-6256-\nkernel/bpf/syscall.c=6257=static int prog_stream_read(union bpf_attr *attr)\n--\nkernel/bpf/syscall.c-6263-\nkernel/bpf/syscall.c:6264:\tif (CHECK_ATTR(BPF_PROG_STREAM_READ_BY_FD))\nkernel/bpf/syscall.c-6265-\t\treturn -EINVAL;\n--\nkernel/bpf/syscall.c=6279=static int prog_assoc_struct_ops(union bpf_attr *attr)\n--\nkernel/bpf/syscall.c-6319-\nkernel/bpf/syscall.c:6320:#define BPF_PROG_STREAM_OPEN_LAST_FIELD prog_stream_open.flags\nkernel/bpf/syscall.c-6321-\nkernel/bpf/syscall.c=6322=static int prog_stream_open(union bpf_attr *attr)\n--\nkernel/bpf/syscall.c-6327-\nkernel/bpf/syscall.c:6328:\tif (CHECK_ATTR(BPF_PROG_STREAM_OPEN))\nkernel/bpf/syscall.c-6329-\t\treturn -EINVAL;\n--\nkernel/bpf/syscall.c=6342=static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,\n--\nkernel/bpf/syscall.c-6507-\t\tbreak;\nkernel/bpf/syscall.c:6508:\tcase BPF_PROG_STREAM_READ_BY_FD:\nkernel/bpf/syscall.c-6509-\t\terr = prog_stream_read(\u0026attr);\n--\nkernel/bpf/syscall.c-6513-\t\tbreak;\nkernel/bpf/syscall.c:6514:\tcase BPF_PROG_STREAM_OPEN:\nkernel/bpf/syscall.c-6515-\t\terr = prog_stream_open(\u0026attr);\n"
} |
| 13/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24223 output: 29 thoughts: 194 Thoughts: **Analyzing the Stream of BPF Functions** Okay, let's dive into this patch. My immediate focus is on these core BPF stream functions: `bpf_prog_stream_new_fd`, `prog_stream_open`, `bpf_stream_file_read`, `bpf_stream_poll`, `bpf_stream_release`, `bpf_stream_notify`, `bpf_prog_stream_init`, and `bpf_prog_stream_free`. These are all very specialized. They aren't the kind of functions you see everywhere, so I need to understand their specific roles within this new stream implementation. The degree of the changes, too, need to be considered. Next, I need to verify what configurations this patch depends on. `CONFIG_BPF_SYSCALL` is an obvious one; however, there might be other, less apparent dependencies. Wait a minute... I see `CONFIG_ANON_INODES` starting to materialize as a potential key. Let's look at `kernel/bpf/stream.c`. Yep, I see the include for `<linux/anon_inodes.h>` and a call to `anon_inode_getfd(...)`. That's a strong indicator. I'll need to check the `Kconfig` files to fully understand this connection and confirm `ANON_INODES` is, in fact, relevant. |
| 14/2 | 2026/09/26 22:19 | tool | grepper |
0mError:no matches Args: {
"Expression": "config ANON_INODES",
"PathPrefix": "fs/Kconfig"
}Results: {
"error": "no matches"
} |
| 15/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24759 output: 19 thoughts: 16 Thoughts: Okay, here's my attempt at summarizing that thought process, keeping in mind the expert user and a first-person perspective: **Hunting for `ANON_INODES` Across the Kconfig Landscape** Right, so I need to track down the configuration option `ANON_INODES`. This isn't just a casual search; it's a targeted query. I want to know where this is defined, potentially enabling or disabling some feature, in every Kconfig file I can get my hands on. This implies a need to potentially understand its impact in different contexts. A brute-force search is in order. Time to dive into the codebase and see where this knob is being turned. Let's get cracking! |
| 16/2 | 2026/09/26 22:19 | tool | grepper |
0mError:no matches Args: {
"Expression": "config ANON_INODES"
}Results: {
"error": "no matches"
} |
| 17/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24958 output: 27 thoughts: 15 Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience: **Investigating `ANON_INODES` Definition** Okay, so I need to understand this `ANON_INODES` concept better. My first instinct is to drill down to its source, the definition. I want to see where it's being declared and, more importantly, *how* it's being used. Is it a constant, a macro, a variable, what? And what files hold this critical piece of the puzzle? Time to fire up the search tools, ideally something that handles source code indexing and cross-referencing effectively. Let's start the hunt for the definition. This will be the foundation to contextualize it within the kernel code or related libraries. Understanding the origin point is essential for determining its role. |
| 18/2 | 2026/09/26 22:19 | tool | grepper |
0mError:no matches Args: {
"Expression": "ANON_INODES",
"PathPrefix": "fs/"
}Results: {
"error": "no matches"
} |
| 19/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 25186 output: 25 thoughts: 17 Thoughts: Here's my thought process, summarized as if I were considering the provided instruction: **Investigating the Location of `anon_inodes.o`** Okay, so I need to find the definition or build instructions for `anon_inodes.o`. Given its name and typical kernel organization, it's highly likely to be found within the filesystem-related code. My first instinct, and the most logical place to start, is to check the `fs/` directory's `Makefile`. That's where the primary build logic for filesystems usually resides. Searching within that `Makefile` would give me a direct line to understanding how `anon_inodes.o` is created, its dependencies, and its place in the broader filesystem compilation. That's the most efficient way to get a quick overview of its context. |
| 20/2 | 2026/09/26 22:19 | tool | grepper |
0mArgs:{
"Expression": "anon_inodes",
"PathPrefix": "fs/"
}Results: {
"Output": "fs/Makefile=25=obj-$(CONFIG_EPOLL)\t\t+= eventpoll.o\nfs/Makefile:26:obj-y\t\t\t\t+= anon_inodes.o\nfs/Makefile-27-obj-$(CONFIG_SIGNALFD)\t\t+= signalfd.o\n--\nfs/anon_inodes.c-2-/*\nfs/anon_inodes.c:3: * fs/anon_inodes.c\nfs/anon_inodes.c-4- *\n--\nfs/anon_inodes.c-21-#include \u003clinux/magic.h\u003e\nfs/anon_inodes.c:22:#include \u003clinux/anon_inodes.h\u003e\nfs/anon_inodes.c-23-#include \u003clinux/pseudo_fs.h\u003e\n--\nfs/eventfd.c-17-#include \u003clinux/spinlock.h\u003e\nfs/eventfd.c:18:#include \u003clinux/anon_inodes.h\u003e\nfs/eventfd.c-19-#include \u003clinux/syscalls.h\u003e\n--\nfs/eventpoll.c-29-#include \u003clinux/mutex.h\u003e\nfs/eventpoll.c:30:#include \u003clinux/anon_inodes.h\u003e\nfs/eventpoll.c-31-#include \u003clinux/device.h\u003e\n--\nfs/fsopen.c-13-#include \u003clinux/security.h\u003e\nfs/fsopen.c:14:#include \u003clinux/anon_inodes.h\u003e\nfs/fsopen.c-15-#include \u003clinux/namei.h\u003e\n--\nfs/notify/fanotify/fanotify_user.c-5-#include \u003clinux/fs.h\u003e\nfs/notify/fanotify/fanotify_user.c:6:#include \u003clinux/anon_inodes.h\u003e\nfs/notify/fanotify/fanotify_user.c-7-#include \u003clinux/fsnotify_backend.h\u003e\n--\nfs/notify/inotify/inotify_user.c-27-#include \u003clinux/types.h\u003e\nfs/notify/inotify/inotify_user.c:28:#include \u003clinux/anon_inodes.h\u003e\nfs/notify/inotify/inotify_user.c-29-#include \u003clinux/uaccess.h\u003e\n--\nfs/pidfs.c-1-// SPDX-License-Identifier: GPL-2.0\nfs/pidfs.c:2:#include \u003clinux/anon_inodes.h\u003e\nfs/pidfs.c-3-#include \u003clinux/compat.h\u003e\n--\nfs/signalfd.c-29-#include \u003clinux/list.h\u003e\nfs/signalfd.c:30:#include \u003clinux/anon_inodes.h\u003e\nfs/signalfd.c-31-#include \u003clinux/signalfd.h\u003e\n--\nfs/timerfd.c-23-#include \u003clinux/hrtimer.h\u003e\nfs/timerfd.c:24:#include \u003clinux/anon_inodes.h\u003e\nfs/timerfd.c-25-#include \u003clinux/timerfd.h\u003e\n--\nfs/xfs/xfs_healthmon.c-27-\nfs/xfs/xfs_healthmon.c:28:#include \u003clinux/anon_inodes.h\u003e\nfs/xfs/xfs_healthmon.c-29-#include \u003clinux/eventpoll.h\u003e\n"
} |
| 21/2 | 2026/09/26 22:19 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 26202 output: 200 thoughts: 752 Thoughts: **BPF Stream File Descriptor Analysis** Okay, let's break this down systematically. `anon_inodes.o` is unconditionally built into `vmlinux` as part of the `fs` build, that much is clear. But my interest is the `BPF_SYSCALL` feature. Diving into `kernel/bpf/Makefile`, the `stream.o` component is only included if `CONFIG_BPF_SYSCALL` is enabled. Further confirming this, the `core.c` file uses `#ifdef CONFIG_BPF_SYSCALL` to guard the call to `bpf_prog_stream_free(fp)`. Now, what else does `CONFIG_BPF_SYSCALL` require? Examining the dependency chain, it selects several other config options: `BPF`, `IRQ_WORK`, `NEED_TASKS_RCU`, `TASKS_TRACE_RCU`, and `BINARY_PRINTF`. Crucially, in `kernel/bpf/Kconfig`, `BPF_SYSCALL` *selects* `BPF`, meaning `BPF` is a dependency, not the other way around. Thus, `BPF_SYSCALL` is the direct config. I've got a list of focus symbols to consider: `prog_stream_open`, `bpf_prog_stream_new_fd`, `bpf_stream_file_read`, `bpf_stream_poll`, `bpf_stream_release`, `bpf_stream_notify`, `bpf_prog_stream_free`, and `bpf_prog_stream_init`. These are the key functions. Are they ubiquitous? No, they're not. They are all specific to the *new* BPF stream file descriptor functionality, involving `BPF_PROG_STREAM_OPEN`, stream reading, polling, notification, and teardown. So, let's review the implications: This new functionality is reachable through the `bpf()` syscall using the `BPF_PROG_STREAM_OPEN` command, followed by `read()`, `poll()`, `close()`, and `bpf()` operations. Based on this analysis: * **Reachability:** Yes, it is reachable. * **WorthFuzzing:** Absolutely, yes. * **FocusSymbols:** Confirmed, the list provided. * **EnableConfigs:** `BPF_SYSCALL`. **Verification:** This patch introduces `BPF_PROG_STREAM_OPEN` in the `bpf()` system call. It provides file operations for BPF stream objects and integrates with anon_inode. These objects introduce new logic in BPF, and VFS, including waitqueue handling, irq\_work notifications, reference counting, and teardown. This constitutes new, reachable, user-facing kernel logic, making it a good target for fuzzing. |
| 22/2 | 2026/09/26 22:19 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"BPF_SYSCALL"
],
"FocusSymbols": [
"bpf_prog_stream_new_fd",
"prog_stream_open",
"bpf_stream_file_read",
"bpf_stream_poll",
"bpf_stream_release",
"bpf_stream_notify",
"bpf_prog_stream_init",
"bpf_prog_stream_free"
],
"Reasoning": "The patch introduces the BPF_PROG_STREAM_OPEN command in the bpf() system call along with anon_inode-backed file operations (read, poll, release) for BPF stream objects. It adds waitqueue handling, irq_work notifications from BPF execution contexts, reference counting, and teardown synchronization. This is new reachable executable code in the core BPF subsystem.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"BPF_SYSCALL"
],
"FocusSymbols": [
"bpf_prog_stream_new_fd",
"prog_stream_open",
"bpf_stream_file_read",
"bpf_stream_poll",
"bpf_stream_release",
"bpf_stream_notify",
"bpf_prog_stream_init",
"bpf_prog_stream_free"
],
"Reasoning": "The patch introduces the BPF_PROG_STREAM_OPEN command in the bpf() system call along with anon_inode-backed file operations (read, poll, release) for BPF stream objects. It adds waitqueue handling, irq_work notifications from BPF execution contexts, reference counting, and teardown synchronization. This is new reachable executable code in the core BPF subsystem.",
"WorthFuzzing": true
} |
| 23/1 | 2026/09/26 22:19 | action | If |
1mArgs:{
"WorthFuzzing": true
}Results: null |
| 24/2 | 2026/09/26 22:19 | agent | kmsan-evaluator |
1mModel:core Args: null Results: {
"KMSANReasoning": "The patch introduces the BPF_PROG_STREAM_OPEN command to the bpf() syscall, allowing user space to open a file descriptor for BPF streams with read() and poll() capabilities. It changes struct bpf_stream from being statically embedded in bpf_prog_aux to being dynamically allocated with refcounting.\n\nAnalysis of uninitialized memory risks (KMSAN):\n1. All dynamic allocations of struct bpf_stream use kzalloc_obj(), ensuring all fields (including atomic counters, flags like notify_used and dead, and backlog pointers) are fully zero-initialized before being passed to their respective init functions (refcount_set, init_llist_head, mutex_init, init_waitqueue_head, init_irq_work).\n2. Data transmitted to user space via copy_to_user() in bpf_stream_read() comes from stream element string buffers (elem-\u003estr), which are zero-allocated (__GFP_ZERO) and fully initialized via memcpy() upon creation. The readable counter only accounts for committed bytes, preventing access to uninitialized bytes.\n3. No struct copies containing uninitialized fields or alignment padding are returned to user space (the syscall only returns a file descriptor integer).\n\nApplicability of KASAN vs KMSAN:\nThe primary risks introduced by this change relate to object lifetimes, reference counting (bpf_stream_put / refcount_inc), concurrent readers/writers, and synchronization across irq_work and anon_inode file operations. Any potential bugs in these areas (such as use-after-free, double-free, or locking issues) are effectively detected by KASAN, LOCKDEP, and standard kernel debugging facilities. No uninitialized memory usage or info-leak risks are present, so a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 6b74fb908de0cf6223a7c477c7bf211b1f2476d2
Author: syz-cluster <triage@syzkaller.com>
Date: Sat Sep 26 22:18:53 2026 +0000
syz-cluster: applied patch under review
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 1e1ce2afe2ed8..a38badb1c2f7d 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -17,6 +17,7 @@
#include <linux/numa.h>
#include <linux/mm_types.h>
#include <linux/wait.h>
+#include <linux/irq_work_types.h>
#include <linux/refcount.h>
#include <linux/mutex.h>
#include <linux/module.h>
@@ -1770,12 +1771,18 @@ enum {
};
struct bpf_stream {
- atomic_t capacity;
+ refcount_t refcnt;
+ atomic_t capacity; /* bytes reserved against the stream limit */
+ atomic_t readable; /* published bytes available to readers */
struct llist_head log; /* list of in-flight stream elements in LIFO order */
struct mutex lock; /* lock protecting backlog_{head,tail} */
struct llist_node *backlog_head; /* list of in-flight stream elements in FIFO order */
struct llist_node *backlog_tail; /* tail of the list above */
+ wait_queue_head_t waitq;
+ struct irq_work notify_work;
+ bool notify_used; /* notify_work was queued at least once */
+ bool dead;
};
struct bpf_stream_stage {
@@ -1914,7 +1921,7 @@ struct bpf_prog_aux {
struct work_struct work;
struct rcu_head rcu;
};
- struct bpf_stream stream[2];
+ struct bpf_stream *stream[2];
struct mutex st_ops_assoc_mutex;
struct bpf_map __rcu *st_ops_assoc;
};
@@ -4223,9 +4230,10 @@ void bpf_bprintf_cleanup(struct bpf_bprintf_data *data);
int bpf_try_get_buffers(struct bpf_bprintf_buffers **bufs);
void bpf_put_buffers(void);
-void bpf_prog_stream_init(struct bpf_prog *prog);
+int bpf_prog_stream_init(struct bpf_prog *prog, gfp_t gfp_extra_flags);
void bpf_prog_stream_free(struct bpf_prog *prog);
int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len);
+int bpf_prog_stream_new_fd(struct bpf_prog *prog, enum bpf_stream_id stream_id, u32 flags);
void bpf_stream_stage_init(struct bpf_stream_stage *ss);
void bpf_stream_stage_free(struct bpf_stream_stage *ss);
__printf(2, 3)
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 0aaa54359aebc..4687c33109968 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -936,6 +936,33 @@ union bpf_iter_link_info {
* 0 on success or -1 if an error occurred (in which case,
* *errno* is set appropriately).
*
+ * BPF_PROG_STREAM_OPEN
+ * Description
+ * Open a file descriptor for one of the BPF streams associated
+ * with the program identified by *prog_fd*. The stream is selected
+ * by *stream_id*.
+ *
+ * The returned file descriptor supports **read**\ (2) and
+ * **poll**\ (2). Reads block while the stream is empty unless
+ * **BPF_F_STREAM_NONBLOCK** is specified in *flags*. A non-blocking
+ * read of an empty stream fails with **EAGAIN**.
+ *
+ * **poll**\ (2) reports **POLLIN** when data is available and
+ * **POLLHUP** once the program has been freed, that is, after every
+ * reference to it, including links and other file descriptors, has
+ * been dropped. Hangup may lag the final release because program
+ * teardown is deferred. Buffered data remains readable after
+ * **POLLHUP** and a read returns zero after all such data has been
+ * consumed.
+ *
+ * The file descriptor is read-only and has the close-on-exec flag
+ * set. It is not seekable and **lseek**\ (2) fails with **ESPIPE**.
+ * *flags* may only contain **BPF_F_STREAM_NONBLOCK**.
+ *
+ * Return
+ * A new file descriptor (a nonnegative integer), or -1 if an
+ * error occurred (in which case, *errno* is set appropriately).
+ *
* NOTES
* eBPF objects (maps and programs) can be shared between processes.
*
@@ -993,6 +1020,7 @@ enum bpf_cmd {
BPF_TOKEN_CREATE,
BPF_PROG_STREAM_READ_BY_FD,
BPF_PROG_ASSOC_STRUCT_OPS,
+ BPF_PROG_STREAM_OPEN,
__MAX_BPF_CMD,
BPF_COMMON_ATTRS = 1 << 16, /* Indicate carrying syscall common attrs. */
};
@@ -1524,6 +1552,11 @@ enum {
BPF_STREAM_STDERR = 2,
};
+/* flags for BPF_PROG_STREAM_OPEN command */
+enum {
+ BPF_F_STREAM_NONBLOCK = (1U << 0),
+};
+
union bpf_attr {
struct { /* anonymous struct used by BPF_MAP_CREATE command */
__u32 map_type; /* one of enum bpf_map_type */
@@ -1950,6 +1983,12 @@ union bpf_attr {
__u32 flags;
} prog_assoc_struct_ops;
+ struct {
+ __u32 prog_fd;
+ __u32 stream_id;
+ __u32 flags;
+ } prog_stream_open;
+
} __attribute__((aligned(8)));
/* The description below is an attempt at providing documentation to eBPF
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 273f74068068c..134700f425319 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -141,10 +141,6 @@ struct bpf_prog *bpf_prog_alloc_no_stats(unsigned int size, gfp_t gfp_extra_flag
mutex_init(&fp->aux->dst_mutex);
mutex_init(&fp->aux->st_ops_assoc_mutex);
-#ifdef CONFIG_BPF_SYSCALL
- bpf_prog_stream_init(fp);
-#endif
-
return fp;
}
@@ -288,6 +284,9 @@ struct bpf_prog *bpf_prog_realloc(struct bpf_prog *fp_old, unsigned int size,
void __bpf_prog_free(struct bpf_prog *fp)
{
if (fp->aux) {
+#ifdef CONFIG_BPF_SYSCALL
+ bpf_prog_stream_free(fp);
+#endif
mutex_destroy(&fp->aux->used_maps_mutex);
mutex_destroy(&fp->aux->dst_mutex);
mutex_destroy(&fp->aux->st_ops_assoc_mutex);
@@ -3073,7 +3072,6 @@ static void bpf_prog_free_deferred(struct work_struct *work)
aux = container_of(work, struct bpf_prog_aux, work);
#ifdef CONFIG_BPF_SYSCALL
bpf_free_kfunc_btf_tab(aux->kfunc_btf_tab);
- bpf_prog_stream_free(aux->prog);
#endif
#ifdef CONFIG_CGROUP_BPF
if (aux->cgroup_atype != CGROUP_BPF_ATTACH_TYPE_INVALID)
diff --git a/kernel/bpf/stream.c b/kernel/bpf/stream.c
index 2b80a0599865e..e0ce37fcd5578 100644
--- a/kernel/bpf/stream.c
+++ b/kernel/bpf/stream.c
@@ -2,11 +2,15 @@
/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
#include <linux/bpf.h>
+#include <linux/anon_inodes.h>
#include <linux/filter.h>
#include <linux/bpf_mem_alloc.h>
#include <linux/gfp.h>
+#include <linux/irq_work.h>
#include <linux/memory.h>
#include <linux/mutex.h>
+#include <linux/poll.h>
+#include <linux/refcount.h>
static void bpf_stream_elem_init(struct bpf_stream_elem *elem, int len)
{
@@ -73,16 +77,59 @@ static void bpf_stream_release_capacity(struct bpf_stream *stream, int len)
atomic_sub(len, &stream->capacity);
}
+static void bpf_stream_notify(struct irq_work *work)
+{
+ struct bpf_stream *stream = container_of(work, struct bpf_stream, notify_work);
+
+ /*
+ * Writers run in arbitrary program contexts, including NMI and regions
+ * that already hold wait queue or epoll locks. Wake waiters from
+ * irq_work instead, where taking those locks is safe.
+ */
+ wake_up_interruptible_poll(&stream->waitq, EPOLLIN | EPOLLRDNORM);
+}
+
+static void bpf_stream_queue_notify(struct bpf_stream *stream)
+{
+ /*
+ * Record that the work has been used so that teardown only pays for
+ * irq_work_sync(), which may wait for an RCU grace period, when a
+ * callback could actually be in flight.
+ */
+ if (!READ_ONCE(stream->notify_used))
+ WRITE_ONCE(stream->notify_used, true);
+ irq_work_queue(&stream->notify_work);
+}
+
+static int bpf_stream_readable_bytes(struct bpf_stream *stream)
+{
+ return atomic_read_acquire(&stream->readable);
+}
+
+static void bpf_stream_publish(struct bpf_stream *stream, int len)
+{
+ /* Pairs with atomic_read_acquire() in bpf_stream_readable_bytes(). */
+ (void)atomic_add_return_release(len, &stream->readable);
+ bpf_stream_queue_notify(stream);
+}
+
static int bpf_stream_push_str(struct bpf_stream *stream, const char *str, int len)
{
- int ret = bpf_stream_consume_capacity(stream, len);
+ int ret;
+
+ /* Nothing to publish; do not allocate an element for it. */
+ if (!len)
+ return 0;
+ ret = bpf_stream_consume_capacity(stream, len);
if (ret)
return ret;
ret = __bpf_stream_push_str(&stream->log, str, len);
if (ret)
bpf_stream_release_capacity(stream, len);
+ else
+ bpf_stream_publish(stream, len);
return ret;
}
@@ -91,7 +138,7 @@ static struct bpf_stream *bpf_stream_get(enum bpf_stream_id stream_id, struct bp
{
if (stream_id != BPF_STDOUT && stream_id != BPF_STDERR)
return NULL;
- return &aux->stream[stream_id - 1];
+ return aux->stream[stream_id - 1];
}
static void bpf_stream_free_elem(struct bpf_stream_elem *elem)
@@ -159,14 +206,16 @@ static bool bpf_stream_consume_elem(struct bpf_stream_elem *elem, int *len)
static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)
{
- int rem_len = len, cons_len, ret = 0;
+ int read_len, rem_len, cons_len, ret = 0;
struct bpf_stream_elem *elem = NULL;
struct llist_node *node;
mutex_lock(&stream->lock);
+ read_len = min(len, bpf_stream_readable_bytes(stream));
+ rem_len = read_len;
while (rem_len) {
- int pos = len - rem_len;
+ int pos = read_len - rem_len;
int chunk, n;
bool cont;
@@ -188,7 +237,7 @@ static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)
/* Keep any successfully copied bytes; -EFAULT only if none. */
elem->consumed_len -= n;
rem_len += n;
- ret = (len == rem_len) ? -EFAULT : 0;
+ ret = (read_len == rem_len) ? -EFAULT : 0;
break;
}
@@ -199,8 +248,9 @@ static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)
bpf_stream_free_elem(elem);
}
+ atomic_sub(read_len - rem_len, &stream->readable);
mutex_unlock(&stream->lock);
- return ret ? ret : len - rem_len;
+ return ret ? ret : read_len - rem_len;
}
int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len)
@@ -215,6 +265,111 @@ int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, vo
return bpf_stream_read(stream, buf, len);
}
+static bool bpf_stream_has_data(struct bpf_stream *stream)
+{
+ return bpf_stream_readable_bytes(stream) > 0;
+}
+
+static void bpf_stream_put(struct bpf_stream *stream)
+{
+ if (refcount_dec_and_test(&stream->refcnt)) {
+ struct llist_node *list;
+
+ /* Only a stream that ever queued its work can have a callback in flight. */
+ if (READ_ONCE(stream->notify_used))
+ irq_work_sync(&stream->notify_work);
+ list = llist_del_all(&stream->log);
+ bpf_stream_free_list(list);
+ bpf_stream_free_list(stream->backlog_head);
+ mutex_destroy(&stream->lock);
+ kfree(stream);
+ }
+}
+
+static int bpf_stream_release(struct inode *inode, struct file *file)
+{
+ bpf_stream_put(file->private_data);
+ return 0;
+}
+
+static ssize_t bpf_stream_file_read(struct file *file, char __user *buf, size_t len,
+ loff_t *ppos)
+{
+ struct bpf_stream *stream = file->private_data;
+ bool dead;
+ int ret;
+
+ if (!len)
+ return 0;
+
+ for (;;) {
+ /*
+ * Sample teardown state before looking for data. Nothing is
+ * published once the program is gone, so finding the stream
+ * empty after observing dead means EOF. The opposite order could
+ * report EOF while data published just before teardown is still
+ * buffered.
+ */
+ dead = smp_load_acquire(&stream->dead);
+ ret = bpf_stream_read(stream, buf, len);
+ if (ret)
+ return ret;
+ if (dead)
+ return 0;
+ if (file->f_flags & O_NONBLOCK)
+ return -EAGAIN;
+
+ ret = wait_event_interruptible(stream->waitq,
+ bpf_stream_has_data(stream) ||
+ READ_ONCE(stream->dead));
+ if (ret)
+ return ret;
+ }
+}
+
+static __poll_t bpf_stream_poll(struct file *file, struct poll_table_struct *pts)
+{
+ struct bpf_stream *stream = file->private_data;
+ __poll_t events = 0;
+
+ /*
+ * poll_wait() only registers the wait queue callback. Register before
+ * checking persistent state so a concurrent publication or teardown is
+ * observed either by the callback or by the checks below.
+ */
+ poll_wait(file, &stream->waitq, pts);
+ if (bpf_stream_has_data(stream))
+ events |= EPOLLIN | EPOLLRDNORM;
+ if (READ_ONCE(stream->dead))
+ events |= EPOLLHUP;
+ return events;
+}
+
+static const struct file_operations bpf_stream_fops = {
+ .release = bpf_stream_release,
+ .read = bpf_stream_file_read,
+ .poll = bpf_stream_poll,
+};
+
+int bpf_prog_stream_new_fd(struct bpf_prog *prog, enum bpf_stream_id stream_id, u32 flags)
+{
+ struct bpf_stream *stream;
+ int fd_flags = O_RDONLY | O_CLOEXEC;
+ int fd;
+
+ stream = bpf_stream_get(stream_id, prog->aux);
+ if (!stream)
+ return -ENOENT;
+ if (flags & BPF_F_STREAM_NONBLOCK)
+ fd_flags |= O_NONBLOCK;
+
+ refcount_inc(&stream->refcnt);
+ fd = anon_inode_getfd("bpf-stream", &bpf_stream_fops, stream, fd_flags);
+ if (fd < 0)
+ bpf_stream_put(stream);
+ return fd;
+}
+
__bpf_kfunc_start_defs();
/*
@@ -282,28 +437,47 @@ __bpf_kfunc_end_defs();
/* Added kfunc to common_btf_ids */
-void bpf_prog_stream_init(struct bpf_prog *prog)
+int bpf_prog_stream_init(struct bpf_prog *prog, gfp_t gfp_extra_flags)
{
int i;
for (i = 0; i < ARRAY_SIZE(prog->aux->stream); i++) {
- atomic_set(&prog->aux->stream[i].capacity, 0);
- init_llist_head(&prog->aux->stream[i].log);
- mutex_init(&prog->aux->stream[i].lock);
- prog->aux->stream[i].backlog_head = NULL;
- prog->aux->stream[i].backlog_tail = NULL;
+ struct bpf_stream *stream;
+
+ /* On failure, bpf_prog_stream_free() releases the streams allocated so far. */
+ stream = kzalloc_obj(*stream,
+ bpf_memcg_flags(GFP_KERNEL | gfp_extra_flags));
+ if (!stream)
+ return -ENOMEM;
+
+ refcount_set(&stream->refcnt, 1);
+ init_llist_head(&stream->log);
+ mutex_init(&stream->lock);
+ init_waitqueue_head(&stream->waitq);
+ init_irq_work(&stream->notify_work, bpf_stream_notify);
+ prog->aux->stream[i] = stream;
}
+ return 0;
}
void bpf_prog_stream_free(struct bpf_prog *prog)
{
- struct llist_node *list;
int i;
for (i = 0; i < ARRAY_SIZE(prog->aux->stream); i++) {
- list = llist_del_all(&prog->aux->stream[i].log);
- bpf_stream_free_list(list);
- bpf_stream_free_list(prog->aux->stream[i].backlog_head);
+ struct bpf_stream *stream = prog->aux->stream[i];
+
+ if (!stream)
+ continue;
+ /*
+ * Pairs with smp_load_acquire() in bpf_stream_file_read(): every
+ * publication precedes the dead flag, so a reader that observes
+ * it also observes all buffered data.
+ */
+ smp_store_release(&stream->dead, true);
+ wake_up_interruptible_poll(&stream->waitq, EPOLLHUP);
+ bpf_stream_put(stream);
+ prog->aux->stream[i] = NULL;
}
}
@@ -333,8 +507,8 @@ int bpf_stream_stage_printk(struct bpf_stream_stage *ss, const char *fmt, ...)
va_start(args, fmt);
len = vscnprintf(buf->buf, ARRAY_SIZE(buf->buf), fmt, args);
va_end(args);
- /* Exclude NULL byte during push. */
- ret = __bpf_stream_push_str(&ss->log, buf->buf, len);
+ /* Exclude NULL byte during push; skip empty output entirely. */
+ ret = len ? __bpf_stream_push_str(&ss->log, buf->buf, len) : 0;
if (!ret)
ss->len += len;
bpf_put_buffers();
@@ -366,6 +540,7 @@ int bpf_stream_stage_commit(struct bpf_stream_stage *ss, struct bpf_prog *prog,
list = tail;
}
llist_add_batch(head, tail, &stream->log);
+ bpf_stream_publish(stream, ss->len);
return 0;
}
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index 74496fd716d3b..8580f38b41c16 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -3080,6 +3080,10 @@ static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_at
prog->aux->user = get_current_user();
prog->len = attr->insn_cnt;
+ err = bpf_prog_stream_init(prog, GFP_USER);
+ if (err)
+ goto free_prog;
+
err = -EFAULT;
if (copy_from_bpfptr(prog->insns,
make_bpfptr(attr->insns, uattr.is_kernel),
@@ -6313,6 +6317,28 @@ static int prog_assoc_struct_ops(union bpf_attr *attr)
return ret;
}
+#define BPF_PROG_STREAM_OPEN_LAST_FIELD prog_stream_open.flags
+
+static int prog_stream_open(union bpf_attr *attr)
+{
+ struct bpf_prog *prog;
+ u32 flags = attr->prog_stream_open.flags;
+ int ret;
+
+ if (CHECK_ATTR(BPF_PROG_STREAM_OPEN))
+ return -EINVAL;
+ if (flags & ~BPF_F_STREAM_NONBLOCK)
+ return -EINVAL;
+
+ prog = bpf_prog_get(attr->prog_stream_open.prog_fd);
+ if (IS_ERR(prog))
+ return PTR_ERR(prog);
+
+ ret = bpf_prog_stream_new_fd(prog, attr->prog_stream_open.stream_id, flags);
+ bpf_prog_put(prog);
+ return ret;
+}
+
static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,
bpfptr_t uattr_common, unsigned int size_common)
{
@@ -6485,6 +6511,9 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,
case BPF_PROG_ASSOC_STRUCT_OPS:
err = prog_assoc_struct_ops(&attr);
break;
+ case BPF_PROG_STREAM_OPEN:
+ err = prog_stream_open(&attr);
+ break;
default:
err = -EINVAL;
break;
diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
index 90fe8c61bf424..ef7e00eb5e479 100644
--- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
+++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
@@ -186,6 +186,11 @@ bpftool prog tracelog { stdout | stderr } *PROG*
error messages to the standard error stream. This facility should be used
only for debugging purposes.
+ On kernels that support opening a stream as a file descriptor, bpftool
+ keeps printing new output as the program produces it, until the program is
+ unloaded or <Ctrl+C> is hit. Older kernels only allow dumping the output
+ buffered so far, after which bpftool exits.
+
bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ctx_in *FILE* [ctx_out *FILE* [ctx_size_out *M*]]] [repeat *N*]
Run BPF program *PROG* in the kernel testing infrastructure for BPF,
meaning that the program works on the data and context provided by the
diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
index 24e40dfab4690..f0241dada5485 100644
--- a/tools/bpf/bpftool/prog.c
+++ b/tools/bpf/bpftool/prog.c
@@ -1119,21 +1119,69 @@ enum prog_tracelog_mode {
TRACE_STDERR,
};
+static volatile sig_atomic_t stream_stop;
+
+static void stop_stream(int signo)
+{
+ stream_stop = 1;
+}
+
+/* Consumes prog_fd. */
static int
prog_tracelog_stream(int prog_fd, enum prog_tracelog_mode mode)
{
+ /* No SA_RESTART: an interrupted read() must return EINTR to end the loop. */
+ const struct sigaction act = { .sa_handler = stop_stream };
+ const int signals[] = { SIGHUP, SIGINT, SIGTERM };
+ struct sigaction old[ARRAY_SIZE(signals)];
FILE *file = mode == TRACE_STDOUT ? stdout : stderr;
int stream_id = mode == TRACE_STDOUT ? 1 : 2;
char buf[512];
- int ret;
+ unsigned int i;
+ int fd, ret;
+
+ fd = bpf_prog_stream_open(prog_fd, stream_id, NULL);
+ if (fd == -EINVAL) {
+ /* Kernel predates BPF_PROG_STREAM_OPEN: dump buffered output and exit. */
+ do {
+ ret = bpf_prog_stream_read(prog_fd, stream_id, buf, sizeof(buf), NULL);
+ if (ret > 0)
+ fwrite(buf, sizeof(buf[0]), ret, file);
+ } while (ret > 0);
+ if (ret < 0)
+ p_err("failed to read stream: %s", strerror(-ret));
+ close(prog_fd);
+ goto out;
+ }
+ /*
+ * The stream descriptor does not keep the program alive. Drop the
+ * program reference so that reads return EOF once the program is gone.
+ */
+ close(prog_fd);
+ if (fd < 0) {
+ p_err("failed to open stream: %s", strerror(-fd));
+ return -1;
+ }
+ stream_stop = 0;
+ for (i = 0; i < ARRAY_SIZE(signals); i++)
+ sigaction(signals[i], &act, &old[i]);
ret = 0;
- do {
- ret = bpf_prog_stream_read(prog_fd, stream_id, buf, sizeof(buf), NULL);
- if (ret > 0)
- fwrite(buf, sizeof(buf[0]), ret, file);
- } while (ret > 0);
-
+ while (!stream_stop) {
+ ret = read(fd, buf, sizeof(buf));
+ if (ret <= 0)
+ break;
+ fwrite(buf, sizeof(buf[0]), ret, file);
+ fflush(file);
+ }
+ if (ret < 0 && !(stream_stop && errno == EINTR))
+ p_err("failed to read stream: %s", strerror(errno));
+ else
+ ret = 0;
+ for (i = 0; i < ARRAY_SIZE(signals); i++)
+ sigaction(signals[i], &old[i], NULL);
+ close(fd);
+out:
fflush(file);
return ret ? -1 : 0;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 0aaa54359aebc..4687c33109968 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -936,6 +936,33 @@ union bpf_iter_link_info {
* 0 on success or -1 if an error occurred (in which case,
* *errno* is set appropriately).
*
+ * BPF_PROG_STREAM_OPEN
+ * Description
+ * Open a file descriptor for one of the BPF streams associated
+ * with the program identified by *prog_fd*. The stream is selected
+ * by *stream_id*.
+ *
+ * The returned file descriptor supports **read**\ (2) and
+ * **poll**\ (2). Reads block while the stream is empty unless
+ * **BPF_F_STREAM_NONBLOCK** is specified in *flags*. A non-blocking
+ * read of an empty stream fails with **EAGAIN**.
+ *
+ * **poll**\ (2) reports **POLLIN** when data is available and
+ * **POLLHUP** once the program has been freed, that is, after every
+ * reference to it, including links and other file descriptors, has
+ * been dropped. Hangup may lag the final release because program
+ * teardown is deferred. Buffered data remains readable after
+ * **POLLHUP** and a read returns zero after all such data has been
+ * consumed.
+ *
+ * The file descriptor is read-only and has the close-on-exec flag
+ * set. It is not seekable and **lseek**\ (2) fails with **ESPIPE**.
+ * *flags* may only contain **BPF_F_STREAM_NONBLOCK**.
+ *
+ * Return
+ * A new file descriptor (a nonnegative integer), or -1 if an
+ * error occurred (in which case, *errno* is set appropriately).
+ *
* NOTES
* eBPF objects (maps and programs) can be shared between processes.
*
@@ -993,6 +1020,7 @@ enum bpf_cmd {
BPF_TOKEN_CREATE,
BPF_PROG_STREAM_READ_BY_FD,
BPF_PROG_ASSOC_STRUCT_OPS,
+ BPF_PROG_STREAM_OPEN,
__MAX_BPF_CMD,
BPF_COMMON_ATTRS = 1 << 16, /* Indicate carrying syscall common attrs. */
};
@@ -1524,6 +1552,11 @@ enum {
BPF_STREAM_STDERR = 2,
};
+/* flags for BPF_PROG_STREAM_OPEN command */
+enum {
+ BPF_F_STREAM_NONBLOCK = (1U << 0),
+};
+
union bpf_attr {
struct { /* anonymous struct used by BPF_MAP_CREATE command */
__u32 map_type; /* one of enum bpf_map_type */
@@ -1950,6 +1983,12 @@ union bpf_attr {
__u32 flags;
} prog_assoc_struct_ops;
+ struct {
+ __u32 prog_fd;
+ __u32 stream_id;
+ __u32 flags;
+ } prog_stream_open;
+
} __attribute__((aligned(8)));
/* The description below is an attempt at providing documentation to eBPF
diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
index a9de7f107cf7b..b49822d212aed 100644
--- a/tools/lib/bpf/bpf.c
+++ b/tools/lib/bpf/bpf.c
@@ -1466,6 +1466,25 @@ int bpf_prog_stream_read(int prog_fd, __u32 stream_id, void *buf, __u32 buf_len,
return libbpf_err_errno(err);
}
+int bpf_prog_stream_open(int prog_fd, __u32 stream_id,
+ const struct bpf_prog_stream_open_opts *opts)
+{
+ const size_t attr_sz = offsetofend(union bpf_attr, prog_stream_open);
+ union bpf_attr attr;
+ int fd;
+
+ if (!OPTS_VALID(opts, bpf_prog_stream_open_opts))
+ return libbpf_err(-EINVAL);
+
+ memset(&attr, 0, attr_sz);
+ attr.prog_stream_open.prog_fd = prog_fd;
+ attr.prog_stream_open.stream_id = stream_id;
+ attr.prog_stream_open.flags = OPTS_GET(opts, flags, 0);
+
+ fd = sys_bpf_fd(BPF_PROG_STREAM_OPEN, &attr, attr_sz);
+ return libbpf_err_errno(fd);
+}
+
int bpf_prog_assoc_struct_ops(int prog_fd, int map_fd,
struct bpf_prog_assoc_struct_ops_opts *opts)
{
diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
index 490e8cb4ba537..826d9cc9ab65d 100644
--- a/tools/lib/bpf/bpf.h
+++ b/tools/lib/bpf/bpf.h
@@ -759,10 +759,34 @@ struct bpf_prog_stream_read_opts {
*
* @return The number of bytes read, on success; negative error code, otherwise
* (errno is also set to the error code)
+ *
+ * For blocking reads and polling, prefer **bpf_prog_stream_open**.
*/
LIBBPF_API int bpf_prog_stream_read(int prog_fd, __u32 stream_id, void *buf, __u32 buf_len,
struct bpf_prog_stream_read_opts *opts);
+struct bpf_prog_stream_open_opts {
+ size_t sz;
+ __u32 flags;
+ size_t :0;
+};
+#define bpf_prog_stream_open_opts__last_field flags
+
+/**
+ * @brief **bpf_prog_stream_open** opens a file descriptor for a BPF stream of
+ * a given BPF program.
+ *
+ * @param prog_fd FD for the BPF program whose BPF stream is to be opened.
+ * @param stream_id ID of the BPF stream to be opened.
+ * @param opts optional options, can be NULL. BPF_F_STREAM_NONBLOCK requests a
+ * non-blocking descriptor.
+ *
+ * @return A new stream FD, on success; negative error code, otherwise (errno
+ * is also set to the error code)
+ */
+LIBBPF_API int bpf_prog_stream_open(int prog_fd, __u32 stream_id,
+ const struct bpf_prog_stream_open_opts *opts);
+
struct bpf_prog_assoc_struct_ops_opts {
size_t sz;
__u32 flags;
diff --git a/tools/lib/bpf/libbpf.map b/tools/lib/bpf/libbpf.map
index 18d27d20102ec..53b7b591325f5 100644
--- a/tools/lib/bpf/libbpf.map
+++ b/tools/lib/bpf/libbpf.map
@@ -459,6 +459,7 @@ LIBBPF_1.7.0 {
LIBBPF_1.8.0 {
global:
bpf_map__attach_cgroup_opts;
+ bpf_prog_stream_open;
bpf_program__add_flags;
bpf_program__attach_tracing_multi;
bpf_program__clear_flags;
diff --git a/tools/testing/selftests/bpf/prog_tests/stream.c b/tools/testing/selftests/bpf/prog_tests/stream.c
index 74bd15c4bfbaa..2c6b441552862 100644
--- a/tools/testing/selftests/bpf/prog_tests/stream.c
+++ b/tools/testing/selftests/bpf/prog_tests/stream.c
@@ -1,11 +1,17 @@
// SPDX-License-Identifier: GPL-2.0
/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */
#include <test_progs.h>
+#include <linux/perf_event.h>
+#include <poll.h>
+#include <sys/epoll.h>
#include <sys/mman.h>
+#include <sys/syscall.h>
#include "stream.skel.h"
#include "stream_fail.skel.h"
+#define NMI_TIMEOUT_NS (5ULL * 1000 * 1000 * 1000)
+
void test_stream_failure(void)
{
RUN_TESTS(stream_fail);
@@ -62,6 +68,415 @@ void test_stream_syscall(void)
stream__destroy(skel);
}
+static bool stream_fd_trigger(struct bpf_program *prog)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ int ret;
+
+ ret = bpf_prog_test_run_opts(bpf_program__fd(prog), &opts);
+ return ASSERT_OK(ret, "test_run") && ASSERT_OK(opts.retval, "retval");
+}
+
+static void test_stream_fd_open(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts);
+ struct stream *skel;
+ int fd, prog_fd;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ prog_fd = bpf_program__fd(skel->progs.stream_syscall);
+ fd = bpf_prog_stream_open(0, BPF_STREAM_STDOUT, NULL);
+ ASSERT_EQ(fd, -EINVAL, "bad_prog_fd");
+
+ fd = bpf_prog_stream_open(prog_fd, 0, NULL);
+ ASSERT_EQ(fd, -ENOENT, "bad_stream_id");
+
+ opts.flags = BPF_F_RDONLY;
+ fd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDOUT, &opts);
+ ASSERT_EQ(fd, -EINVAL, "access_flag");
+
+ opts.flags = 1U << 31;
+ fd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDOUT, &opts);
+ ASSERT_EQ(fd, -EINVAL, "unknown_flag");
+
+ fd = bpf_prog_stream_open(prog_fd, BPF_STREAM_STDERR, NULL);
+ if (ASSERT_OK_FD(fd, "stderr"))
+ close(fd);
+
+ stream__destroy(skel);
+}
+
+static void test_stream_fd_nonblock(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts,
+ .flags = BPF_F_STREAM_NONBLOCK,
+ );
+ struct pollfd pfd = { .events = POLLIN | POLLHUP };
+ struct stream *skel;
+ char buf[4] = {};
+ int fd, flags, ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_syscall),
+ BPF_STREAM_STDOUT, &opts);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out_destroy;
+ pfd.fd = fd;
+
+ flags = fcntl(fd, F_GETFD);
+ ASSERT_GE(flags, 0, "getfd");
+ ASSERT_NEQ(flags & FD_CLOEXEC, 0, "cloexec");
+ flags = fcntl(fd, F_GETFL);
+ ASSERT_GE(flags, 0, "getfl");
+ ASSERT_EQ(flags & O_ACCMODE, O_RDONLY, "readonly");
+ ASSERT_NEQ(flags & O_NONBLOCK, 0, "nonblock");
+
+ ret = write(fd, "x", 1);
+ ASSERT_EQ(ret, -1, "write");
+ ASSERT_EQ(errno, EBADF, "write_errno");
+ ret = lseek(fd, 0, SEEK_SET);
+ ASSERT_EQ(ret, -1, "lseek");
+ ASSERT_EQ(errno, ESPIPE, "lseek_errno");
+
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 0, "poll_empty");
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, -1, "read_empty");
+ ASSERT_EQ(errno, EAGAIN, "read_empty_errno");
+
+ if (!stream_fd_trigger(skel->progs.stream_syscall))
+ goto out_close;
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 1, "poll_data");
+ ASSERT_NEQ(pfd.revents & POLLIN, 0, "pollin");
+ ASSERT_EQ(pfd.revents & POLLHUP, 0, "no_pollhup");
+
+ ret = read(fd, buf, 2);
+ ASSERT_EQ(ret, 2, "read_first");
+ ASSERT_OK(memcmp(buf, "fo", 2), "read_first_data");
+ pfd.revents = 0;
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 1, "poll_partial");
+ ASSERT_NEQ(pfd.revents & POLLIN, 0, "pollin_partial");
+
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, 1, "read_rest");
+ ASSERT_EQ(buf[0], 'o', "read_rest_data");
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, -1, "read_drained");
+ ASSERT_EQ(errno, EAGAIN, "read_drained_errno");
+
+out_close:
+ close(fd);
+out_destroy:
+ stream__destroy(skel);
+}
+
+static void test_stream_fd_empty(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts,
+ .flags = BPF_F_STREAM_NONBLOCK,
+ );
+ struct pollfd pfd = { .events = POLLIN | POLLHUP };
+ struct stream *skel;
+ char buf[4];
+ int fd, ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_empty),
+ BPF_STREAM_STDOUT, &opts);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out_destroy;
+ pfd.fd = fd;
+
+ /* Empty output produces neither data nor readiness. */
+ if (!stream_fd_trigger(skel->progs.stream_empty))
+ goto out_close;
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 0, "poll_empty_write");
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, -1, "read_empty_write");
+ ASSERT_EQ(errno, EAGAIN, "read_empty_write_errno");
+
+out_close:
+ close(fd);
+out_destroy:
+ stream__destroy(skel);
+}
+
+struct stream_fd_read_ctx {
+ int fd;
+ ssize_t ret;
+ char buf[4];
+};
+
+static void *stream_fd_read_thread(void *arg)
+{
+ struct stream_fd_read_ctx *ctx = arg;
+
+ ctx->ret = read(ctx->fd, ctx->buf, sizeof(ctx->buf));
+ return NULL;
+}
+
+static int stream_fd_timed_join(pthread_t thread)
+{
+ struct timespec timeout;
+
+ clock_gettime(CLOCK_REALTIME, &timeout);
+ timeout.tv_sec += 5;
+ return pthread_timedjoin_np(thread, NULL, &timeout);
+}
+
+static void test_stream_fd_blocking(void)
+{
+ struct stream_fd_read_ctx ctx = {};
+ struct stream *skel;
+ pthread_t thread;
+ bool thread_live = false;
+ int fd, flags, ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_syscall),
+ BPF_STREAM_STDOUT, NULL);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out_destroy;
+ ctx.fd = fd;
+
+ flags = fcntl(fd, F_GETFL);
+ ASSERT_GE(flags, 0, "getfl");
+ ASSERT_EQ(flags & O_NONBLOCK, 0, "blocking");
+
+ ret = pthread_create(&thread, NULL, stream_fd_read_thread, &ctx);
+ if (!ASSERT_OK(ret, "pthread_create"))
+ goto out_close;
+ thread_live = true;
+
+ usleep(50000);
+ ret = pthread_tryjoin_np(thread, NULL);
+ if (!ASSERT_EQ(ret, EBUSY, "read_blocks")) {
+ thread_live = ret != 0;
+ goto out_thread;
+ }
+ if (!stream_fd_trigger(skel->progs.stream_syscall))
+ goto out_thread;
+
+ ret = stream_fd_timed_join(thread);
+ if (!ASSERT_OK(ret, "pthread_join"))
+ goto out_thread;
+ thread_live = false;
+ ASSERT_EQ(ctx.ret, 3, "read_len");
+ ASSERT_OK(memcmp(ctx.buf, "foo", 3), "read_data");
+
+ memset(&ctx, 0, sizeof(ctx));
+ ctx.fd = fd;
+ ret = pthread_create(&thread, NULL, stream_fd_read_thread, &ctx);
+ if (!ASSERT_OK(ret, "pthread_create_eof"))
+ goto out_close;
+ thread_live = true;
+
+ usleep(50000);
+ ret = pthread_tryjoin_np(thread, NULL);
+ if (!ASSERT_EQ(ret, EBUSY, "read_eof_blocks")) {
+ thread_live = ret != 0;
+ goto out_thread;
+ }
+
+ stream__destroy(skel);
+ skel = NULL;
+ ret = stream_fd_timed_join(thread);
+ if (!ASSERT_OK(ret, "pthread_join_eof"))
+ goto out_thread;
+ thread_live = false;
+ ASSERT_EQ(ctx.ret, 0, "read_eof");
+
+out_thread:
+ if (thread_live) {
+ stream__destroy(skel);
+ skel = NULL;
+ ret = stream_fd_timed_join(thread);
+ if (ret) {
+ pthread_cancel(thread);
+ pthread_join(thread, NULL);
+ }
+ }
+out_close:
+ close(fd);
+out_destroy:
+ stream__destroy(skel);
+}
+
+static void test_stream_fd_hup(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts,
+ .flags = BPF_F_STREAM_NONBLOCK,
+ );
+ struct epoll_event event = {
+ .events = EPOLLIN | EPOLLET,
+ };
+ struct pollfd pfd = { .events = POLLIN | POLLHUP };
+ struct stream *skel;
+ char buf[4] = {};
+ int epfd = -1, fd, ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_syscall),
+ BPF_STREAM_STDOUT, &opts);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out_destroy;
+ pfd.fd = fd;
+ epfd = epoll_create1(EPOLL_CLOEXEC);
+ if (!ASSERT_OK_FD(epfd, "epoll_create"))
+ goto out_close;
+ event.data.fd = fd;
+ ret = epoll_ctl(epfd, EPOLL_CTL_ADD, fd, &event);
+ if (!ASSERT_OK(ret, "epoll_ctl"))
+ goto out_close;
+
+ if (!stream_fd_trigger(skel->progs.stream_syscall))
+ goto out_close;
+ ret = epoll_wait(epfd, &event, 1, 5000);
+ if (!ASSERT_EQ(ret, 1, "epoll_wait_data"))
+ goto out_close;
+ ASSERT_NEQ(event.events & EPOLLIN, 0, "epollin");
+ ASSERT_EQ(event.events & EPOLLHUP, 0, "no_epollhup");
+
+ stream__destroy(skel);
+ skel = NULL;
+ event.events = 0;
+ ret = epoll_wait(epfd, &event, 1, 5000);
+ if (!ASSERT_EQ(ret, 1, "epoll_wait_hup"))
+ goto out_close;
+ ASSERT_NEQ(event.events & EPOLLIN, 0, "epollin_with_hup");
+ ASSERT_NEQ(event.events & EPOLLHUP, 0, "epollhup");
+
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, 3, "read_buffered");
+ ASSERT_OK(memcmp(buf, "foo", 3), "read_buffered_data");
+ pfd.revents = 0;
+ ret = poll(&pfd, 1, 0);
+ ASSERT_EQ(ret, 1, "poll_drained_hup");
+ ASSERT_EQ(pfd.revents & POLLIN, 0, "no_pollin_after_drain");
+ ASSERT_NEQ(pfd.revents & POLLHUP, 0, "pollhup_after_drain");
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, 0, "read_eof");
+
+out_close:
+ if (epfd >= 0)
+ close(epfd);
+ close(fd);
+out_destroy:
+ stream__destroy(skel);
+}
+
+static void test_stream_fd_nmi_epoll(void)
+{
+ LIBBPF_OPTS(bpf_prog_stream_open_opts, opts,
+ .flags = BPF_F_STREAM_NONBLOCK,
+ );
+ struct perf_event_attr attr = {
+ .size = sizeof(attr),
+ .type = PERF_TYPE_HARDWARE,
+ .config = PERF_COUNT_HW_CPU_CYCLES,
+ .freq = 1,
+ .sample_freq = 10,
+ };
+ struct epoll_event event = {
+ .events = EPOLLIN,
+ };
+ struct bpf_link *link = NULL;
+ struct stream *skel;
+ __u64 deadline;
+ char buf[4] = {};
+ int epfd = -1, fd = -1, pmu_fd = -1;
+ int ret;
+
+ skel = stream__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "stream__open_and_load"))
+ return;
+ fd = bpf_prog_stream_open(bpf_program__fd(skel->progs.stream_nmi),
+ BPF_STREAM_STDOUT, &opts);
+ if (!ASSERT_OK_FD(fd, "stream_open"))
+ goto out;
+
+ epfd = epoll_create1(EPOLL_CLOEXEC);
+ if (!ASSERT_OK_FD(epfd, "epoll_create"))
+ goto out;
+ event.data.fd = fd;
+ ret = epoll_ctl(epfd, EPOLL_CTL_ADD, fd, &event);
+ if (!ASSERT_OK(ret, "epoll_ctl"))
+ goto out;
+
+ pmu_fd = syscall(__NR_perf_event_open, &attr, 0, -1, -1,
+ PERF_FLAG_FD_CLOEXEC);
+ if (pmu_fd < 0 && (errno == ENOENT || errno == EOPNOTSUPP)) {
+ printf("%s:SKIP:no PERF_COUNT_HW_CPU_CYCLES\n", __func__);
+ test__skip();
+ goto out;
+ }
+ if (!ASSERT_GE(pmu_fd, 0, "perf_event_open"))
+ goto out;
+
+ link = bpf_program__attach_perf_event(skel->progs.stream_nmi, pmu_fd);
+ if (!ASSERT_OK_PTR(link, "attach_perf_event")) {
+ link = NULL;
+ goto out;
+ }
+ pmu_fd = -1;
+
+ deadline = get_time_ns() + NMI_TIMEOUT_NS;
+ do {
+ ret = epoll_wait(epfd, &event, 1, 0);
+ } while (ret == 0 && get_time_ns() < deadline);
+ if (!ASSERT_EQ(ret, 1, "epoll_wait"))
+ goto out;
+ ASSERT_NEQ(event.events & EPOLLIN, 0, "epollin");
+
+ ret = read(fd, buf, sizeof(buf));
+ ASSERT_EQ(ret, 3, "read_len");
+ ASSERT_OK(memcmp(buf, "nmi", 3), "read_data");
+
+out:
+ bpf_link__destroy(link);
+ if (pmu_fd >= 0)
+ close(pmu_fd);
+ if (epfd >= 0)
+ close(epfd);
+ if (fd >= 0)
+ close(fd);
+ stream__destroy(skel);
+}
+
+void test_stream_fd(void)
+{
+ if (test__start_subtest("open"))
+ test_stream_fd_open();
+ if (test__start_subtest("nonblock"))
+ test_stream_fd_nonblock();
+ if (test__start_subtest("empty"))
+ test_stream_fd_empty();
+ if (test__start_subtest("blocking"))
+ test_stream_fd_blocking();
+ if (test__start_subtest("hup"))
+ test_stream_fd_hup();
+ if (test__start_subtest("nmi_epoll"))
+ test_stream_fd_nmi_epoll();
+}
+
void test_stream_oversize(void)
{
LIBBPF_OPTS(bpf_test_run_opts, opts);
diff --git a/tools/testing/selftests/bpf/progs/stream.c b/tools/testing/selftests/bpf/progs/stream.c
index 12fc29e454874..1ae015c4e6b4c 100644
--- a/tools/testing/selftests/bpf/progs/stream.c
+++ b/tools/testing/selftests/bpf/progs/stream.c
@@ -44,6 +44,7 @@ struct {
_X64 _X64 _X64 _X64 _X64 _X64 _X64 _X64)
int size;
+int nmi_stream_prints;
u64 fault_addr;
void *arena_ptr;
@@ -124,6 +125,25 @@ int stream_syscall(void *ctx)
return 0;
}
+SEC("syscall")
+__success __retval(0)
+int stream_empty(void *ctx)
+{
+ return bpf_stream_printk(BPF_STDOUT, "");
+}
+
+SEC("perf_event")
+int stream_nmi(void *ctx)
+{
+ if (nmi_stream_prints)
+ return 0;
+ /* Retry on a later sample if the write failed, e.g. with -EBUSY. */
+ if (bpf_stream_printk(BPF_STDOUT, "nmi"))
+ return 0;
+ nmi_stream_prints = 1;
+ return 0;
+}
+
SEC("syscall")
__success __retval(0)
int stream_oversize(void *ctx)
|
| 25/3 | 2026/09/26 22:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 17834 output: 33 thoughts: 1569 Thoughts: **KMSAN Assessment: BPF Stream Patch** Okay, let's break this down. The task is to decide if these BPF changes warrant a dedicated KMSAN fuzzing session, beyond the standard KASAN efforts. The instructions are clear: KMSAN is for detecting *uninitialized* memory reads and info-leaks, especially in scenarios KASAN, LOCKDEP, and the other standard detectors might miss. It's about efficiency; we don't want to waste resources. First, I need to understand the patch's impact. It introduces the `BPF_PROG_STREAM_OPEN` command to the bpf syscall. Critically, it changes `struct bpf_stream` from being embedded in the `bpf_prog_aux` structure to being dynamically allocated as a pointer array `struct bpf_stream *stream[2]`. This is the core change. Now, let's go through the checklist. Rule out the KASAN stuff. This is *not* about out-of-bounds reads, overflows, or use-after-free conditions. Those are already covered. This is about *uninitialized* memory. I examine the changes to `struct bpf_stream` in `include/linux/bpf.h`. The structure now includes `refcnt`, `capacity`, `readable`, a linked list for log entries, a mutex, a FIFO backlog, a wait queue, an irq work, and a `dead` flag. Important: in `bpf_prog_stream_init` is a call to `kzalloc_obj`, which zero-initializes the stream struct. Excellent! This *should* prevent uninitialized memory issues during initialization. All fields of the struct are explicitly zeroed out, including the new fields. Next, I need to check the uAPI headers. `BPF_PROG_STREAM_OPEN` is added to the bpf command list. `struct bpf_attr` is updated, and it looks like a new input is being introduced via the `bpf_attr->prog_stream_open`. Nothing is copied to userspace here, just values provided by userspace. However, `bpf_prog_stream_open` returns an fd, created by `anon_inode_getfd`. The return value of `__sys_bpf` is `ret` which is the fd, an integer, meaning nothing is being copied to userspace here. The critical test is whether the patch *introduces* new avenues for uninitialized memory usage *later* in the program flow. Let's look for user-space copies that could lead to info leaks. I'll need to check the code that deals with reading streams, especially `bpf_stream_read` and associated functions, to see if anything gets copied to userspace. I'll use `grepper` or `read-file` to explore how the `stream` structure is used to verify that no information leaks. Given that the allocation uses `kzalloc_obj` and the API itself does not directly copy memory to user space, I'm leaning towards *not* needing a dedicated KMSAN session. The crucial thing is the initial zeroing; absent that, I would be much more concerned. However, I need to be 100% sure that `bpf_stream_read`, `bpf_stream_file_read` and `__bpf_stream_push_str` (or anything else related to data being read from the stream) don't have issues. |
| 26/3 | 2026/09/26 22:19 | tool | grepper |
0mArgs:{
"Expression": "bpf_stream_",
"PathPrefix": "kernel/bpf/stream.c"
}Results: {
"Output": "kernel/bpf/stream.c-14-\nkernel/bpf/stream.c:15:static void bpf_stream_elem_init(struct bpf_stream_elem *elem, int len)\nkernel/bpf/stream.c-16-{\n--\nkernel/bpf/stream.c-21-\nkernel/bpf/stream.c:22:static struct bpf_stream_elem *bpf_stream_elem_alloc(int len)\nkernel/bpf/stream.c-23-{\nkernel/bpf/stream.c-24-\tconst int max_len = ARRAY_SIZE((struct bpf_bprintf_buffers){}.buf);\nkernel/bpf/stream.c:25:\tstruct bpf_stream_elem *elem;\nkernel/bpf/stream.c-26-\tsize_t alloc_size;\n--\nkernel/bpf/stream.c-35-\nkernel/bpf/stream.c:36:\talloc_size = offsetof(struct bpf_stream_elem, str[len]);\nkernel/bpf/stream.c-37-\telem = kmalloc_nolock(alloc_size, __GFP_ZERO, -1);\n--\nkernel/bpf/stream.c-40-\nkernel/bpf/stream.c:41:\tbpf_stream_elem_init(elem, len);\nkernel/bpf/stream.c-42-\n--\nkernel/bpf/stream.c-45-\nkernel/bpf/stream.c:46:static int __bpf_stream_push_str(struct llist_head *log, const char *str, int len)\nkernel/bpf/stream.c-47-{\nkernel/bpf/stream.c:48:\tstruct bpf_stream_elem *elem = NULL;\nkernel/bpf/stream.c-49-\n--\nkernel/bpf/stream.c-53-\t */\nkernel/bpf/stream.c:54:\telem = bpf_stream_elem_alloc(len);\nkernel/bpf/stream.c-55-\tif (!elem)\n--\nkernel/bpf/stream.c-63-\nkernel/bpf/stream.c:64:static int bpf_stream_consume_capacity(struct bpf_stream *stream, int len)\nkernel/bpf/stream.c-65-{\n--\nkernel/bpf/stream.c-74-\nkernel/bpf/stream.c:75:static void bpf_stream_release_capacity(struct bpf_stream *stream, int len)\nkernel/bpf/stream.c-76-{\n--\nkernel/bpf/stream.c-79-\nkernel/bpf/stream.c:80:static void bpf_stream_notify(struct irq_work *work)\nkernel/bpf/stream.c-81-{\n--\nkernel/bpf/stream.c-91-\nkernel/bpf/stream.c:92:static void bpf_stream_queue_notify(struct bpf_stream *stream)\nkernel/bpf/stream.c-93-{\n--\nkernel/bpf/stream.c-103-\nkernel/bpf/stream.c:104:static int bpf_stream_readable_bytes(struct bpf_stream *stream)\nkernel/bpf/stream.c-105-{\n--\nkernel/bpf/stream.c-108-\nkernel/bpf/stream.c:109:static void bpf_stream_publish(struct bpf_stream *stream, int len)\nkernel/bpf/stream.c-110-{\nkernel/bpf/stream.c:111:\t/* Pairs with atomic_read_acquire() in bpf_stream_readable_bytes(). */\nkernel/bpf/stream.c-112-\t(void)atomic_add_return_release(len, \u0026stream-\u003ereadable);\nkernel/bpf/stream.c:113:\tbpf_stream_queue_notify(stream);\nkernel/bpf/stream.c-114-}\nkernel/bpf/stream.c-115-\nkernel/bpf/stream.c:116:static int bpf_stream_push_str(struct bpf_stream *stream, const char *str, int len)\nkernel/bpf/stream.c-117-{\n--\nkernel/bpf/stream.c-123-\nkernel/bpf/stream.c:124:\tret = bpf_stream_consume_capacity(stream, len);\nkernel/bpf/stream.c-125-\tif (ret)\n--\nkernel/bpf/stream.c-127-\nkernel/bpf/stream.c:128:\tret = __bpf_stream_push_str(\u0026stream-\u003elog, str, len);\nkernel/bpf/stream.c-129-\tif (ret)\nkernel/bpf/stream.c:130:\t\tbpf_stream_release_capacity(stream, len);\nkernel/bpf/stream.c-131-\telse\nkernel/bpf/stream.c:132:\t\tbpf_stream_publish(stream, len);\nkernel/bpf/stream.c-133-\n--\nkernel/bpf/stream.c-136-\nkernel/bpf/stream.c:137:static struct bpf_stream *bpf_stream_get(enum bpf_stream_id stream_id, struct bpf_prog_aux *aux)\nkernel/bpf/stream.c-138-{\n--\nkernel/bpf/stream.c-143-\nkernel/bpf/stream.c:144:static void bpf_stream_free_elem(struct bpf_stream_elem *elem)\nkernel/bpf/stream.c-145-{\n--\nkernel/bpf/stream.c-148-\nkernel/bpf/stream.c:149:static void bpf_stream_free_list(struct llist_node *list)\nkernel/bpf/stream.c-150-{\nkernel/bpf/stream.c:151:\tstruct bpf_stream_elem *elem, *tmp;\nkernel/bpf/stream.c-152-\nkernel/bpf/stream.c-153-\tllist_for_each_entry_safe(elem, tmp, list, node)\nkernel/bpf/stream.c:154:\t\tbpf_stream_free_elem(elem);\nkernel/bpf/stream.c-155-}\nkernel/bpf/stream.c-156-\nkernel/bpf/stream.c:157:static struct llist_node *bpf_stream_backlog_peek(struct bpf_stream *stream)\nkernel/bpf/stream.c-158-{\n--\nkernel/bpf/stream.c-161-\nkernel/bpf/stream.c:162:static struct llist_node *bpf_stream_backlog_pop(struct bpf_stream *stream)\nkernel/bpf/stream.c-163-{\n--\nkernel/bpf/stream.c-173-\nkernel/bpf/stream.c:174:static void bpf_stream_backlog_fill(struct bpf_stream *stream)\nkernel/bpf/stream.c-175-{\n--\nkernel/bpf/stream.c-195-\nkernel/bpf/stream.c:196:static bool bpf_stream_consume_elem(struct bpf_stream_elem *elem, int *len)\nkernel/bpf/stream.c-197-{\n--\nkernel/bpf/stream.c-206-\nkernel/bpf/stream.c:207:static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)\nkernel/bpf/stream.c-208-{\nkernel/bpf/stream.c-209-\tint read_len, rem_len, cons_len, ret = 0;\nkernel/bpf/stream.c:210:\tstruct bpf_stream_elem *elem = NULL;\nkernel/bpf/stream.c-211-\tstruct llist_node *node;\n--\nkernel/bpf/stream.c-213-\tmutex_lock(\u0026stream-\u003elock);\nkernel/bpf/stream.c:214:\tread_len = min(len, bpf_stream_readable_bytes(stream));\nkernel/bpf/stream.c-215-\trem_len = read_len;\n--\nkernel/bpf/stream.c-221-\nkernel/bpf/stream.c:222:\t\tnode = bpf_stream_backlog_peek(stream);\nkernel/bpf/stream.c-223-\t\tif (!node) {\nkernel/bpf/stream.c:224:\t\t\tbpf_stream_backlog_fill(stream);\nkernel/bpf/stream.c:225:\t\t\tnode = bpf_stream_backlog_peek(stream);\nkernel/bpf/stream.c-226-\t\t}\n--\nkernel/bpf/stream.c-231-\t\tcons_len = elem-\u003econsumed_len;\nkernel/bpf/stream.c:232:\t\tcont = bpf_stream_consume_elem(elem, \u0026rem_len) == false;\nkernel/bpf/stream.c-233-\t\tchunk = elem-\u003econsumed_len - cons_len;\n--\nkernel/bpf/stream.c-245-\t\t\tcontinue;\nkernel/bpf/stream.c:246:\t\tbpf_stream_backlog_pop(stream);\nkernel/bpf/stream.c:247:\t\tbpf_stream_release_capacity(stream, elem-\u003etotal_len);\nkernel/bpf/stream.c:248:\t\tbpf_stream_free_elem(elem);\nkernel/bpf/stream.c-249-\t}\n--\nkernel/bpf/stream.c-255-\nkernel/bpf/stream.c:256:int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len)\nkernel/bpf/stream.c-257-{\n--\nkernel/bpf/stream.c-259-\nkernel/bpf/stream.c:260:\tstream = bpf_stream_get(stream_id, prog-\u003eaux);\nkernel/bpf/stream.c-261-\tif (!stream)\n--\nkernel/bpf/stream.c-264-\t\treturn -EINVAL;\nkernel/bpf/stream.c:265:\treturn bpf_stream_read(stream, buf, len);\nkernel/bpf/stream.c-266-}\nkernel/bpf/stream.c-267-\nkernel/bpf/stream.c:268:static bool bpf_stream_has_data(struct bpf_stream *stream)\nkernel/bpf/stream.c-269-{\nkernel/bpf/stream.c:270:\treturn bpf_stream_readable_bytes(stream) \u003e 0;\nkernel/bpf/stream.c-271-}\nkernel/bpf/stream.c-272-\nkernel/bpf/stream.c:273:static void bpf_stream_put(struct bpf_stream *stream)\nkernel/bpf/stream.c-274-{\n--\nkernel/bpf/stream.c-281-\t\tlist = llist_del_all(\u0026stream-\u003elog);\nkernel/bpf/stream.c:282:\t\tbpf_stream_free_list(list);\nkernel/bpf/stream.c:283:\t\tbpf_stream_free_list(stream-\u003ebacklog_head);\nkernel/bpf/stream.c-284-\t\tmutex_destroy(\u0026stream-\u003elock);\n--\nkernel/bpf/stream.c-288-\nkernel/bpf/stream.c:289:static int bpf_stream_release(struct inode *inode, struct file *file)\nkernel/bpf/stream.c-290-{\nkernel/bpf/stream.c:291:\tbpf_stream_put(file-\u003eprivate_data);\nkernel/bpf/stream.c-292-\treturn 0;\n--\nkernel/bpf/stream.c-294-\nkernel/bpf/stream.c:295:static ssize_t bpf_stream_file_read(struct file *file, char __user *buf, size_t len,\nkernel/bpf/stream.c-296-\t\t\t\t loff_t *ppos)\n--\nkernel/bpf/stream.c-313-\t\tdead = smp_load_acquire(\u0026stream-\u003edead);\nkernel/bpf/stream.c:314:\t\tret = bpf_stream_read(stream, buf, len);\nkernel/bpf/stream.c-315-\t\tif (ret)\n--\nkernel/bpf/stream.c-322-\t\tret = wait_event_interruptible(stream-\u003ewaitq,\nkernel/bpf/stream.c:323:\t\t\t\t\t bpf_stream_has_data(stream) ||\nkernel/bpf/stream.c-324-\t\t\t\t\t READ_ONCE(stream-\u003edead));\n--\nkernel/bpf/stream.c-329-\nkernel/bpf/stream.c:330:static __poll_t bpf_stream_poll(struct file *file, struct poll_table_struct *pts)\nkernel/bpf/stream.c-331-{\n--\nkernel/bpf/stream.c-340-\tpoll_wait(file, \u0026stream-\u003ewaitq, pts);\nkernel/bpf/stream.c:341:\tif (bpf_stream_has_data(stream))\nkernel/bpf/stream.c-342-\t\tevents |= EPOLLIN | EPOLLRDNORM;\n--\nkernel/bpf/stream.c-347-\nkernel/bpf/stream.c:348:static const struct file_operations bpf_stream_fops = {\nkernel/bpf/stream.c:349:\t.release = bpf_stream_release,\nkernel/bpf/stream.c:350:\t.read = bpf_stream_file_read,\nkernel/bpf/stream.c:351:\t.poll = bpf_stream_poll,\nkernel/bpf/stream.c-352-};\nkernel/bpf/stream.c-353-\nkernel/bpf/stream.c:354:int bpf_prog_stream_new_fd(struct bpf_prog *prog, enum bpf_stream_id stream_id, u32 flags)\nkernel/bpf/stream.c-355-{\n--\nkernel/bpf/stream.c-359-\nkernel/bpf/stream.c:360:\tstream = bpf_stream_get(stream_id, prog-\u003eaux);\nkernel/bpf/stream.c-361-\tif (!stream)\n--\nkernel/bpf/stream.c-366-\trefcount_inc(\u0026stream-\u003erefcnt);\nkernel/bpf/stream.c:367:\tfd = anon_inode_getfd(\"bpf-stream\", \u0026bpf_stream_fops, stream, fd_flags);\nkernel/bpf/stream.c-368-\tif (fd \u003c 0)\nkernel/bpf/stream.c:369:\t\tbpf_stream_put(stream);\nkernel/bpf/stream.c-370-\treturn fd;\n--\nkernel/bpf/stream.c=373=__bpf_kfunc_start_defs();\n--\nkernel/bpf/stream.c-375-/*\nkernel/bpf/stream.c:376: * Avoid using enum bpf_stream_id so that kfunc users don't have to pull in the\nkernel/bpf/stream.c-377- * enum in headers.\nkernel/bpf/stream.c-378- */\nkernel/bpf/stream.c:379:__bpf_kfunc int bpf_stream_vprintk(int stream_id, const char *fmt__str, const void *args,\nkernel/bpf/stream.c-380-\t\t\t\t u32 len__sz, struct bpf_prog_aux *aux)\n--\nkernel/bpf/stream.c-390-\nkernel/bpf/stream.c:391:\tstream = bpf_stream_get(stream_id, aux);\nkernel/bpf/stream.c-392-\tif (!stream)\n--\nkernel/bpf/stream.c-410-\t/* Exclude NULL byte during push. */\nkernel/bpf/stream.c:411:\tret = bpf_stream_push_str(stream, data.buf, ret);\nkernel/bpf/stream.c-412-\tbpf_bprintf_cleanup(\u0026data);\n--\nkernel/bpf/stream.c-417-/* Directly trigger a stack dump from the program. */\nkernel/bpf/stream.c:418:__bpf_kfunc int bpf_stream_print_stack(int stream_id, struct bpf_prog_aux *aux)\nkernel/bpf/stream.c-419-{\nkernel/bpf/stream.c:420:\tstruct bpf_stream_stage ss;\nkernel/bpf/stream.c-421-\tstruct bpf_prog *prog;\n--\nkernel/bpf/stream.c-423-\t/* Make sure the stream ID is valid. */\nkernel/bpf/stream.c:424:\tif (!bpf_stream_get(stream_id, aux))\nkernel/bpf/stream.c-425-\t\treturn -ENOENT;\n--\nkernel/bpf/stream.c-428-\nkernel/bpf/stream.c:429:\tbpf_stream_stage(ss, prog, stream_id, ({\nkernel/bpf/stream.c:430:\t\tbpf_stream_dump_stack(ss);\nkernel/bpf/stream.c-431-\t}));\n--\nkernel/bpf/stream.c=440=int bpf_prog_stream_init(struct bpf_prog *prog, gfp_t gfp_extra_flags)\n--\nkernel/bpf/stream.c-456-\t\tinit_waitqueue_head(\u0026stream-\u003ewaitq);\nkernel/bpf/stream.c:457:\t\tinit_irq_work(\u0026stream-\u003enotify_work, bpf_stream_notify);\nkernel/bpf/stream.c-458-\t\tprog-\u003eaux-\u003estream[i] = stream;\n--\nkernel/bpf/stream.c=463=void bpf_prog_stream_free(struct bpf_prog *prog)\n--\nkernel/bpf/stream.c-472-\t\t/*\nkernel/bpf/stream.c:473:\t\t * Pairs with smp_load_acquire() in bpf_stream_file_read(): every\nkernel/bpf/stream.c-474-\t\t * publication precedes the dead flag, so a reader that observes\n--\nkernel/bpf/stream.c-478-\t\twake_up_interruptible_poll(\u0026stream-\u003ewaitq, EPOLLHUP);\nkernel/bpf/stream.c:479:\t\tbpf_stream_put(stream);\nkernel/bpf/stream.c-480-\t\tprog-\u003eaux-\u003estream[i] = NULL;\n--\nkernel/bpf/stream.c-483-\nkernel/bpf/stream.c:484:void bpf_stream_stage_init(struct bpf_stream_stage *ss)\nkernel/bpf/stream.c-485-{\n--\nkernel/bpf/stream.c-489-\nkernel/bpf/stream.c:490:void bpf_stream_stage_free(struct bpf_stream_stage *ss)\nkernel/bpf/stream.c-491-{\n--\nkernel/bpf/stream.c-494-\tnode = llist_del_all(\u0026ss-\u003elog);\nkernel/bpf/stream.c:495:\tbpf_stream_free_list(node);\nkernel/bpf/stream.c-496-}\nkernel/bpf/stream.c-497-\nkernel/bpf/stream.c:498:int bpf_stream_stage_printk(struct bpf_stream_stage *ss, const char *fmt, ...)\nkernel/bpf/stream.c-499-{\n--\nkernel/bpf/stream.c-510-\t/* Exclude NULL byte during push; skip empty output entirely. */\nkernel/bpf/stream.c:511:\tret = len ? __bpf_stream_push_str(\u0026ss-\u003elog, buf-\u003ebuf, len) : 0;\nkernel/bpf/stream.c-512-\tif (!ret)\n--\nkernel/bpf/stream.c-517-\nkernel/bpf/stream.c:518:int bpf_stream_stage_commit(struct bpf_stream_stage *ss, struct bpf_prog *prog,\nkernel/bpf/stream.c:519:\t\t\t enum bpf_stream_id stream_id)\nkernel/bpf/stream.c-520-{\n--\nkernel/bpf/stream.c-524-\nkernel/bpf/stream.c:525:\tstream = bpf_stream_get(stream_id, prog-\u003eaux);\nkernel/bpf/stream.c-526-\tif (!stream)\n--\nkernel/bpf/stream.c-528-\nkernel/bpf/stream.c:529:\tret = bpf_stream_consume_capacity(stream, ss-\u003elen);\nkernel/bpf/stream.c-530-\tif (ret)\n--\nkernel/bpf/stream.c-542-\tllist_add_batch(head, tail, \u0026stream-\u003elog);\nkernel/bpf/stream.c:543:\tbpf_stream_publish(stream, ss-\u003elen);\nkernel/bpf/stream.c-544-\treturn 0;\n--\nkernel/bpf/stream.c=547=struct dump_stack_ctx {\nkernel/bpf/stream.c:548:\tstruct bpf_stream_stage *ss;\nkernel/bpf/stream.c-549-\tint err;\n--\nkernel/bpf/stream.c=552=static bool dump_stack_cb(void *cookie, u64 ip, u64 sp, u64 bp)\n--\nkernel/bpf/stream.c-565-\t\t\tgoto end;\nkernel/bpf/stream.c:566:\t\tctxp-\u003eerr = bpf_stream_stage_printk(ctxp-\u003ess, \"%pS\\n %s @ %s:%d\\n\",\nkernel/bpf/stream.c-567-\t\t\t\t\t\t (void *)(long)ip, line, file, num);\n--\nkernel/bpf/stream.c-570-end:\nkernel/bpf/stream.c:571:\tctxp-\u003eerr = bpf_stream_stage_printk(ctxp-\u003ess, \"%pS\\n\", (void *)(long)ip);\nkernel/bpf/stream.c-572-\treturn !ctxp-\u003eerr;\n--\nkernel/bpf/stream.c-574-\nkernel/bpf/stream.c:575:int bpf_stream_stage_dump_stack(struct bpf_stream_stage *ss)\nkernel/bpf/stream.c-576-{\n--\nkernel/bpf/stream.c-579-\nkernel/bpf/stream.c:580:\tret = bpf_stream_stage_printk(ss, \"CPU: %d UID: %d PID: %d Comm: %s\\n\",\nkernel/bpf/stream.c-581-\t\t\t\t raw_smp_processor_id(), __kuid_val(current_real_cred()-\u003eeuid),\n--\nkernel/bpf/stream.c-584-\t\treturn ret;\nkernel/bpf/stream.c:585:\tret = bpf_stream_stage_printk(ss, \"Call trace:\\n\");\nkernel/bpf/stream.c-586-\tif (ret)\n--\nkernel/bpf/stream.c-590-\t\treturn ctx.err;\nkernel/bpf/stream.c:591:\treturn bpf_stream_stage_printk(ss, \"\\n\");\nkernel/bpf/stream.c-592-}\n"
} |
| 27/3 | 2026/09/26 22:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 26323 output: 35 thoughts: 66 Thoughts: Here's my summary, as though I'm thinking through this process: **Diving into BPF Stream Data Handling** Alright, let's get down to brass tacks. I need to understand exactly how data is managed within these BPF streams. Specifically, I'm going to focus on the read and write operations, so I'm heading straight into `kernel/bpf/stream.c`. My targets are `bpf_stream_read`, naturally, which should give me the retrieval side of things. Then, I'll examine the initialization and allocation with `bpf_stream_elem_init` and `bpf_stream_elem_alloc` to see how elements are set up. Finally, I'm especially interested in seeing how strings are handled, so I'll be sure to understand `__bpf_stream_push_str`. I need to grasp how the structures work together to push data into, and read data from, these BPF streams. These functions should give me a good picture of the underlying data flow. |
| 28/3 | 2026/09/26 22:19 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/stream.c",
"FirstLine": 15,
"LineCount": 60
}Results: {
"Contents": " 15:\tstatic void bpf_stream_elem_init(struct bpf_stream_elem *elem, int len)\n 16:\t{\n 17:\t\tinit_llist_node(\u0026elem-\u003enode);\n 18:\t\telem-\u003etotal_len = len;\n 19:\t\telem-\u003econsumed_len = 0;\n 20:\t}\n 21:\t\n 22:\tstatic struct bpf_stream_elem *bpf_stream_elem_alloc(int len)\n 23:\t{\n 24:\t\tconst int max_len = ARRAY_SIZE((struct bpf_bprintf_buffers){}.buf);\n 25:\t\tstruct bpf_stream_elem *elem;\n 26:\t\tsize_t alloc_size;\n 27:\t\n 28:\t\t/*\n 29:\t\t * Length is the payload pushed into the stream, excluding the\n 30:\t\t * trailing NUL of the bprintf buffer. Reject anything that cannot\n 31:\t\t * fit without copying that NUL into the stream element.\n 32:\t\t */\n 33:\t\tif (len \u003c 0 || len \u003e= max_len)\n 34:\t\t\treturn NULL;\n 35:\t\n 36:\t\talloc_size = offsetof(struct bpf_stream_elem, str[len]);\n 37:\t\telem = kmalloc_nolock(alloc_size, __GFP_ZERO, -1);\n 38:\t\tif (!elem)\n 39:\t\t\treturn NULL;\n 40:\t\n 41:\t\tbpf_stream_elem_init(elem, len);\n 42:\t\n 43:\t\treturn elem;\n 44:\t}\n 45:\t\n 46:\tstatic int __bpf_stream_push_str(struct llist_head *log, const char *str, int len)\n 47:\t{\n 48:\t\tstruct bpf_stream_elem *elem = NULL;\n 49:\t\n 50:\t\t/*\n 51:\t\t * Allocate a bpf_prog_stream_elem and push it to the bpf_prog_stream\n 52:\t\t * log, elements will be popped at once and reversed to print the log.\n 53:\t\t */\n 54:\t\telem = bpf_stream_elem_alloc(len);\n 55:\t\tif (!elem)\n 56:\t\t\treturn -ENOMEM;\n 57:\t\n 58:\t\tmemcpy(elem-\u003estr, str, len);\n 59:\t\tllist_add(\u0026elem-\u003enode, log);\n 60:\t\n 61:\t\treturn 0;\n 62:\t}\n 63:\t\n 64:\tstatic int bpf_stream_consume_capacity(struct bpf_stream *stream, int len)\n 65:\t{\n 66:\t\tif (atomic_read(\u0026stream-\u003ecapacity) \u003e= BPF_STREAM_MAX_CAPACITY)\n 67:\t\t\treturn -ENOSPC;\n 68:\t\tif (atomic_add_return(len, \u0026stream-\u003ecapacity) \u003e= BPF_STREAM_MAX_CAPACITY) {\n 69:\t\t\tatomic_sub(len, \u0026stream-\u003ecapacity);\n 70:\t\t\treturn -ENOSPC;\n 71:\t\t}\n 72:\t\treturn 0;\n 73:\t}\n 74:\t\n"
} |
| 29/3 | 2026/09/26 22:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 27483 output: 36 thoughts: 27 Thoughts: Okay, let's dive into this `bpf_stream_read` function, specifically lines 200 to 260. I need to understand this core logic thoroughly. I'm focusing on these specific lines because they're likely the heart of the streaming read operation, which is critical for performance and resource utilization. I expect there to be clever handling of data buffers, potential optimizations for different read sizes, and error checking to ensure the stream is handled robustly. I'm looking for how it interacts with the underlying BPF infrastructure and how it manages potentially asynchronous data availability. Time to get down to it. |
| 30/3 | 2026/09/26 22:19 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/stream.c",
"FirstLine": 200,
"LineCount": 60
}Results: {
"Contents": " 200:\t\n 201:\t\telem-\u003econsumed_len += used;\n 202:\t\t*len -= used;\n 203:\t\n 204:\t\treturn elem-\u003econsumed_len == elem-\u003etotal_len;\n 205:\t}\n 206:\t\n 207:\tstatic int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)\n 208:\t{\n 209:\t\tint read_len, rem_len, cons_len, ret = 0;\n 210:\t\tstruct bpf_stream_elem *elem = NULL;\n 211:\t\tstruct llist_node *node;\n 212:\t\n 213:\t\tmutex_lock(\u0026stream-\u003elock);\n 214:\t\tread_len = min(len, bpf_stream_readable_bytes(stream));\n 215:\t\trem_len = read_len;\n 216:\t\n 217:\t\twhile (rem_len) {\n 218:\t\t\tint pos = read_len - rem_len;\n 219:\t\t\tint chunk, n;\n 220:\t\t\tbool cont;\n 221:\t\n 222:\t\t\tnode = bpf_stream_backlog_peek(stream);\n 223:\t\t\tif (!node) {\n 224:\t\t\t\tbpf_stream_backlog_fill(stream);\n 225:\t\t\t\tnode = bpf_stream_backlog_peek(stream);\n 226:\t\t\t}\n 227:\t\t\tif (!node)\n 228:\t\t\t\tbreak;\n 229:\t\t\telem = container_of(node, typeof(*elem), node);\n 230:\t\n 231:\t\t\tcons_len = elem-\u003econsumed_len;\n 232:\t\t\tcont = bpf_stream_consume_elem(elem, \u0026rem_len) == false;\n 233:\t\t\tchunk = elem-\u003econsumed_len - cons_len;\n 234:\t\n 235:\t\t\tn = copy_to_user(buf + pos, elem-\u003estr + cons_len, chunk);\n 236:\t\t\tif (n) {\n 237:\t\t\t\t/* Keep any successfully copied bytes; -EFAULT only if none. */\n 238:\t\t\t\telem-\u003econsumed_len -= n;\n 239:\t\t\t\trem_len += n;\n 240:\t\t\t\tret = (read_len == rem_len) ? -EFAULT : 0;\n 241:\t\t\t\tbreak;\n 242:\t\t\t}\n 243:\t\n 244:\t\t\tif (cont)\n 245:\t\t\t\tcontinue;\n 246:\t\t\tbpf_stream_backlog_pop(stream);\n 247:\t\t\tbpf_stream_release_capacity(stream, elem-\u003etotal_len);\n 248:\t\t\tbpf_stream_free_elem(elem);\n 249:\t\t}\n 250:\t\n 251:\t\tatomic_sub(read_len - rem_len, \u0026stream-\u003ereadable);\n 252:\t\tmutex_unlock(\u0026stream-\u003elock);\n 253:\t\treturn ret ? ret : read_len - rem_len;\n 254:\t}\n 255:\t\n 256:\tint bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len)\n 257:\t{\n 258:\t\tstruct bpf_stream *stream;\n 259:\t\n"
} |
| 31/3 | 2026/09/26 22:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 28584 output: 34 thoughts: 1005 Thoughts: **Analyzing the BPF Stream Read Mechanism** Okay, let's break this down. My initial focus is on `bpf_stream_read` and how a recent patch altered its behavior. I see they've introduced a `read_len` variable: before, they were using just `len`, but now they're calculating `read_len` as the minimum of the requested `len` and the available readable bytes from the stream. This `read_len` is then used consistently throughout the function. It looks like the patch consistently changed the usage of the variable `len` to `read_len` in several places: calculating the position within the buffer, error handling, and in the final return value. Now, I need to understand how the `readable` counter is managed. Let's trace through the publishing of data: `bpf_stream_push_str` calls `bpf_stream_publish`, and `bpf_stream_stage_commit` calls `bpf_stream_publish` as well. The function `bpf_stream_publish` uses `atomic_add_return_release` to increment `stream->readable` with the length of data added to the stream. This indicates a potential race condition in how the readable variable is updated. Then, within `bpf_stream_read` itself, I'm checking the actual data being read, and how this affects readable bytes. It consumes up to the `read_len`. The data is copied from elements (`elem->str`) within a backlog to the provided `buf`. This copying happens after an allocation with `kmalloc_nolock` and a `memcpy`. I'm confirming that the memory pointed to by `elem->str` should always be initialized because of `__GFP_ZERO` used during the allocation and a `memcpy` filling the buffer, so there shouldn't be any uninitialized memory issues. Wait a second, there's a new `prog_stream_open` structure in the `bpf_attr` union. Let me check the kernel code to understand what's going on. This new `prog_stream_open` structure looks like it's used for opening a BPF stream associated with a program. The code validates some flags, gets a program by its file descriptor, and then calls `bpf_prog_stream_new_fd`. I need to figure out how the system call `__sys_bpf` handles `attr`, to understand how the structure `prog_stream_open` is passed and used. |
| 32/3 | 2026/09/26 22:19 | tool | grepper |
0mArgs:{
"Expression": "int __sys_bpf",
"PathPrefix": "kernel/bpf/syscall.c"
}Results: {
"Output": "kernel/bpf/syscall.c=6322=static int prog_stream_open(union bpf_attr *attr)\n--\nkernel/bpf/syscall.c-6341-\nkernel/bpf/syscall.c:6342:static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,\nkernel/bpf/syscall.c-6343-\t\t bpfptr_t uattr_common, unsigned int size_common)\n"
} |
| 33/3 | 2026/09/26 22:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 30294 output: 37 thoughts: 27 Thoughts: Okay, here's my interpretation of what's going through my mind as I approach that section of the kernel: **Dissecting the BPF Syscall Implementation** Alright, time to dive into `kernel/bpf/syscall.c` and specifically, lines 6340 to 6380. I need to understand how this code handles the BPF syscalls, particularly the nuances of this specific block. Given the context of the larger BPF ecosystem, I'm guessing this section probably deals with something critical - maybe function argument validation, resource allocation, or potentially even some crucial interaction with the verifier. I'll need to pay close attention to the data structures used here, like the BPF program itself, any associated maps, and any potential security checks or error handling. My experience tells me that subtle vulnerabilities often hide in these areas, so I need to be methodical. Let's see how they've handled the potential for resource exhaustion or any possible integer overflows. I'm prepared to trace calls, check for race conditions, and be ready for some intricate pointer manipulation. Let's get to it. |
| 34/3 | 2026/09/26 22:19 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/syscall.c",
"FirstLine": 6340,
"LineCount": 40
}Results: {
"Contents": "6340:\t}\n6341:\t\n6342:\tstatic int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,\n6343:\t\t\t bpfptr_t uattr_common, unsigned int size_common)\n6344:\t{\n6345:\t\tstruct bpf_common_attr attr_common;\n6346:\t\tu32 offsetof_log_true_size = 0;\n6347:\t\tstruct bpf_log_attr attr_log;\n6348:\t\tunion bpf_attr attr;\n6349:\t\tint err;\n6350:\t\n6351:\t\terr = bpf_check_uarg_tail_zero(uattr, sizeof(attr), size);\n6352:\t\tif (err)\n6353:\t\t\treturn err;\n6354:\t\tsize = min_t(u32, size, sizeof(attr));\n6355:\t\n6356:\t\t/* copy attributes from user space, may be less than sizeof(bpf_attr) */\n6357:\t\tmemset(\u0026attr, 0, sizeof(attr));\n6358:\t\tif (copy_from_bpfptr(\u0026attr, uattr, size) != 0)\n6359:\t\t\treturn -EFAULT;\n6360:\t\n6361:\t\tmemset(\u0026attr_common, 0, sizeof(attr_common));\n6362:\t\tif (cmd \u0026 BPF_COMMON_ATTRS) {\n6363:\t\t\terr = bpf_check_uarg_tail_zero(uattr_common,\n6364:\t\t\t\t\t\t offsetofend(struct bpf_common_attr, log_true_size),\n6365:\t\t\t\t\t\t size_common);\n6366:\t\t\tif (err)\n6367:\t\t\t\treturn err;\n6368:\t\n6369:\t\t\tcmd \u0026= ~BPF_COMMON_ATTRS;\n6370:\t\t\tsize_common = min_t(u32, size_common, sizeof(attr_common));\n6371:\t\t\tif (copy_from_bpfptr(\u0026attr_common, uattr_common, size_common) != 0)\n6372:\t\t\t\treturn -EFAULT;\n6373:\t\t} else {\n6374:\t\t\tsize_common = 0;\n6375:\t\t}\n6376:\t\n6377:\t\terr = security_bpf(cmd, \u0026attr, size, uattr.is_kernel);\n6378:\t\tif (err \u003c 0)\n6379:\t\t\treturn err;\n"
} |
| 35/3 | 2026/09/26 22:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31267 output: 33 thoughts: 99 Thoughts: **Analyzing the BPF Program Stream Open in `__sys_bpf`** Okay, so I'm looking at the `__sys_bpf` function, specifically how it handles the `BPF_PROG_STREAM_OPEN` command. I see they're initializing the `attr` struct with `memset` to zero it out completely. Then, they use `copy_from_bpfptr` to populate it from `uattr` with a given `size`. Seems standard so far, prevents some memory issues. Now, the interesting bit: this `CHECK_ATTR(BPF_PROG_STREAM_OPEN)` macro. It's the key to validation here. I need to understand what that macro expands to in this specific context. The `attr->NAME.pad || attr->...` part is the giveaway. I'll need to dig into the definition of `CHECK_ATTR` to figure out which members of the `attr` struct it's actually checking and under what conditions it fails. I'll need to know which of `attr->NAME.pad` or `attr->...` are being tested, and what "pad" signifies here. I need to see what fields are being explicitly checked during open and how this translates into security considerations. Let's see how that macro is defined and what it is really doing. |
| 36/3 | 2026/09/26 22:19 | tool | grepper |
0mArgs:{
"Expression": "#define CHECK_ATTR",
"PathPrefix": "kernel/bpf/syscall.c"
}Results: {
"Output": "kernel/bpf/syscall.c=1209=int bpf_get_file_flag(int flags)\n--\nkernel/bpf/syscall.c-1220-/* helper macro to check that unused fields 'union bpf_attr' are zero */\nkernel/bpf/syscall.c:1221:#define CHECK_ATTR(CMD) \\\nkernel/bpf/syscall.c-1222-\tmemchr_inv((void *) \u0026attr-\u003eCMD##_LAST_FIELD + \\\n"
} |
| 37/3 | 2026/09/26 22:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31816 output: 28 thoughts: 773 Thoughts: **Analyzing BPF Program Initialization** Okay, so I'm digging into this `bpf_prog_stream_init` function. The first thing that jumps out is the `CHECK_ATTR` macro. It looks like it validates that unused fields within the `bpf_attr` structure are zeroed out, specifically for the `BPF_PROG_STREAM_OPEN` case. That means anything past `prog_stream_open.flags` needs to be zero. Good to know. Now, let's break down what `bpf_prog_stream_init` actually *does*. It iterates through an array of `bpf_stream` structures, allocating each one with `kzalloc_obj`, initializing a refcount, linked list for logging, mutex, wait queue, and irq work. Crucially, it seems this is for programs with streams. It handles the allocation of these stream objects and their associated data structures. On failure it correctly cleans up by releasing any previously allocated streams. The critical question here is *when* is `bpf_prog_stream_init` called? The provided code snippets give the answer. Looking at the diff, it looks like it was called in `bpf_prog_alloc_no_stats`, but it's now conditionally compiled out via a `#ifdef CONFIG_BPF_SYSCALL` directive. Then, looking at `bpf_prog_load` in `syscall.c`, it looks like `bpf_prog_stream_init` *is* called as part of the loading process. My concern is whether all BPF programs call `bpf_prog_stream_init`. Specifically, I need to know where `bpf_prog_alloc` or `bpf_prog_alloc_no_stats` are called. This is important to determine if there's any pathway for creating a BPF program without calling the `bpf_prog_stream_init` routine. If this is the case, it could lead to uninitialized memory. I need to be sure this function is called for *all* BPF programs that may need the features initialized within it. I'll need to follow the call paths and see if there are any scenarios where `bpf_prog_alloc` is called without subsequently calling `bpf_prog_stream_init`. |
| 38/3 | 2026/09/26 22:20 | tool | grepper |
0mArgs:{
"Expression": "bpf_prog_alloc",
"PathPrefix": "kernel/"
}Results: {
"Output": "kernel/bpf/core.c=96=enum page_size_enum {\n--\nkernel/bpf/core.c-99-\nkernel/bpf/core.c:100:struct bpf_prog *bpf_prog_alloc_no_stats(unsigned int size, gfp_t gfp_extra_flags)\nkernel/bpf/core.c-101-{\n--\nkernel/bpf/core.c-146-\nkernel/bpf/core.c:147:struct bpf_prog *bpf_prog_alloc(unsigned int size, gfp_t gfp_extra_flags)\nkernel/bpf/core.c-148-{\n--\nkernel/bpf/core.c-152-\nkernel/bpf/core.c:153:\tprog = bpf_prog_alloc_no_stats(size, gfp_extra_flags);\nkernel/bpf/core.c-154-\tif (!prog)\n--\nkernel/bpf/core.c-172-}\nkernel/bpf/core.c:173:EXPORT_SYMBOL_GPL(bpf_prog_alloc);\nkernel/bpf/core.c-174-\nkernel/bpf/core.c:175:int bpf_prog_alloc_jited_linfo(struct bpf_prog *prog)\nkernel/bpf/core.c-176-{\n--\nkernel/bpf/core.c=2659=struct bpf_prog *__bpf_prog_select_runtime(struct bpf_verifier_env *env, struct bpf_prog *fp,\n--\nkernel/bpf/core.c-2679-\tif (!bpf_prog_is_offloaded(fp-\u003eaux)) {\nkernel/bpf/core.c:2680:\t\t*err = bpf_prog_alloc_jited_linfo(fp);\nkernel/bpf/core.c-2681-\t\tif (*err)\n--\nkernel/bpf/fixups.c=1184=static int jit_subprogs(struct bpf_verifier_env *env)\n--\nkernel/bpf/fixups.c-1229-\nkernel/bpf/fixups.c:1230:\terr = bpf_prog_alloc_jited_linfo(prog);\nkernel/bpf/fixups.c-1231-\tif (err)\n--\nkernel/bpf/fixups.c-1248-\t\t */\nkernel/bpf/fixups.c:1249:\t\tfunc[i] = bpf_prog_alloc_no_stats(bpf_prog_size(len), GFP_USER);\nkernel/bpf/fixups.c-1250-\t\tif (!func[i])\n--\nkernel/bpf/syscall.c=2396=static void bpf_audit_prog(const struct bpf_prog *prog, unsigned int op)\n--\nkernel/bpf/syscall.c-2414-\nkernel/bpf/syscall.c:2415:static int bpf_prog_alloc_id(struct bpf_prog *prog)\nkernel/bpf/syscall.c-2416-{\n--\nkernel/bpf/syscall.c=2929=static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_attr *attr_log)\n--\nkernel/bpf/syscall.c-3053-\t/* plain bpf_prog allocation */\nkernel/bpf/syscall.c:3054:\tprog = bpf_prog_alloc(bpf_prog_size(attr-\u003einsn_cnt), GFP_USER);\nkernel/bpf/syscall.c-3055-\tif (!prog) {\n--\nkernel/bpf/syscall.c-3156-\nkernel/bpf/syscall.c:3157:\terr = bpf_prog_alloc_id(prog);\nkernel/bpf/syscall.c-3158-\tif (err)\n--\nkernel/bpf/syscall.c-3160-\nkernel/bpf/syscall.c:3161:\t/* Upon success of bpf_prog_alloc_id(), the BPF prog is\nkernel/bpf/syscall.c-3162-\t * effectively publicly exposed. However, retrieving via\n--\nkernel/bpf/trampoline.c=1069=static struct bpf_shim_tramp_link *cgroup_shim_alloc(const struct bpf_prog *prog,\n--\nkernel/bpf/trampoline.c-1080-\nkernel/bpf/trampoline.c:1081:\tp = bpf_prog_alloc(1, 0);\nkernel/bpf/trampoline.c-1082-\tif (!p) {\n"
} |
| 39/3 | 2026/09/26 22:20 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 34317 output: 36 thoughts: 59 Thoughts: Here's my understanding of the situation, presented as if it were my thought process: **Investigating `prog->aux->stream` After `bpf_prog_alloc_no_stats`** Alright, I need to understand what happens to `prog->aux->stream` during the allocation process, specifically when we use `bpf_prog_alloc_no_stats` to allocate the `prog` structure. This seems like a critical piece of the puzzle, likely related to how BPF programs are initialized or potentially how their output streams are handled, which I need to understand in detail. Therefore, I need to carefully examine the `bpf_prog_alloc_no_stats` function definition. Let's head over to `kernel/bpf/core.c` and dig in. I expect to find the relevant memory allocation and initialization of the `prog` and related auxiliary structures. I am looking for whether the `aux` member is initialized within the allocation or if additional logic is needed. I'll focus on what gets set for `aux` and whether it has a direct effect on the `stream` member of the aux structure. |
| 40/3 | 2026/09/26 22:20 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/core.c",
"FirstLine": 100,
"LineCount": 50
}Results: {
"Contents": " 100:\tstruct bpf_prog *bpf_prog_alloc_no_stats(unsigned int size, gfp_t gfp_extra_flags)\n 101:\t{\n 102:\t\tgfp_t gfp_flags = bpf_memcg_flags(GFP_KERNEL | __GFP_ZERO | gfp_extra_flags);\n 103:\t\tstruct bpf_prog_aux *aux;\n 104:\t\tstruct bpf_prog *fp;\n 105:\t\n 106:\t\tsize = round_up(size, __PAGE_SIZE);\n 107:\t\tfp = __vmalloc(size, gfp_flags);\n 108:\t\tif (fp == NULL)\n 109:\t\t\treturn NULL;\n 110:\t\n 111:\t\taux = kzalloc_obj(*aux, bpf_memcg_flags(GFP_KERNEL | gfp_extra_flags));\n 112:\t\tif (aux == NULL) {\n 113:\t\t\tvfree(fp);\n 114:\t\t\treturn NULL;\n 115:\t\t}\n 116:\t\tfp-\u003eactive = __alloc_percpu_gfp(sizeof(u8[BPF_NR_CONTEXTS]), 4,\n 117:\t\t\t\t\t\tbpf_memcg_flags(GFP_KERNEL | gfp_extra_flags));\n 118:\t\tif (!fp-\u003eactive) {\n 119:\t\t\tvfree(fp);\n 120:\t\t\tkfree(aux);\n 121:\t\t\treturn NULL;\n 122:\t\t}\n 123:\t\n 124:\t\tfp-\u003epages = size / PAGE_SIZE;\n 125:\t\tfp-\u003eaux = aux;\n 126:\t\tfp-\u003eaux-\u003emain_prog_aux = aux;\n 127:\t\tfp-\u003eaux-\u003eprog = fp;\n 128:\t\tfp-\u003ejit_requested = ebpf_jit_enabled();\n 129:\t\tfp-\u003ejit_required = IS_ENABLED(CONFIG_BPF_JIT_ALWAYS_ON);\n 130:\t\tfp-\u003eblinding_requested = bpf_jit_blinding_enabled(fp);\n 131:\t#ifdef CONFIG_CGROUP_BPF\n 132:\t\taux-\u003ecgroup_atype = CGROUP_BPF_ATTACH_TYPE_INVALID;\n 133:\t#endif\n 134:\t\n 135:\t\tINIT_LIST_HEAD_RCU(\u0026fp-\u003eaux-\u003eksym.lnode);\n 136:\t#ifdef CONFIG_FINEIBT\n 137:\t\tINIT_LIST_HEAD_RCU(\u0026fp-\u003eaux-\u003eksym_prefix.lnode);\n 138:\t#endif\n 139:\t\tmutex_init(\u0026fp-\u003eaux-\u003eused_maps_mutex);\n 140:\t\tmutex_init(\u0026fp-\u003eaux-\u003eext_mutex);\n 141:\t\tmutex_init(\u0026fp-\u003eaux-\u003edst_mutex);\n 142:\t\tmutex_init(\u0026fp-\u003eaux-\u003est_ops_assoc_mutex);\n 143:\t\n 144:\t\treturn fp;\n 145:\t}\n 146:\t\n 147:\tstruct bpf_prog *bpf_prog_alloc(unsigned int size, gfp_t gfp_extra_flags)\n 148:\t{\n 149:\t\tgfp_t gfp_flags = bpf_memcg_flags(GFP_KERNEL | __GFP_ZERO | gfp_extra_flags);\n"
} |
| 41/3 | 2026/09/26 22:20 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 35540 output: 29 thoughts: 484 Thoughts: **Analyzing BPF Program Stream Cleanup** Okay, so I'm looking at how `bpf_prog` streams are handled during cleanup. It starts with `aux` being allocated, and importantly, it's initialized with `kzalloc_obj`, which means `aux->stream[0]` and `aux->stream[1]` are initially NULL. That's key. The cleanup process itself, within `__bpf_prog_free`, calls `bpf_prog_stream_free`. Inside `bpf_prog_stream_free`, there's a loop that iterates through the `stream` array. The critical part is that if a `stream` pointer is NULL (as it is initially), the code simply continues to the next iteration. This avoids any double-free or invalid access since the function does a NULL check. This is important to ensure correct cleanup of the BPF stream array. Now, when a `bpf_stream` is accessed with `bpf_stream_get`, the code checks `aux->stream[stream_id - 1]`. The function `bpf_stream_get` returns NULL if it's the stream ID is not one of `BPF_STDOUT` and `BPF_STDERR`. If that entry is NULL, `bpf_stream_get` correctly returns NULL. Crucially, the calling code *should* be checking the return value and handling the NULL case which will return -ENOENT or -EINVAL, as seen in the code. I need to verify that all the callers are doing this check and responding appropriately. Therefore, I need to check `bpf_prog_stream_read_by_fd` and other potential stream callers to make sure there are no obvious flaws in stream access or memory management. A quick grep of the kernel sources for anything starting with `bpf_prog_stream_` should give me a good overview of the potential problem spots. |
| 42/3 | 2026/09/26 22:20 | tool | grepper |
0mArgs:{
"Expression": "bpf_prog_stream_",
"PathPrefix": "kernel/"
}Results: {
"Output": "kernel/bpf/core.c=284=void __bpf_prog_free(struct bpf_prog *fp)\n--\nkernel/bpf/core.c-287-#ifdef CONFIG_BPF_SYSCALL\nkernel/bpf/core.c:288:\t\tbpf_prog_stream_free(fp);\nkernel/bpf/core.c-289-#endif\n--\nkernel/bpf/stream.c=46=static int __bpf_stream_push_str(struct llist_head *log, const char *str, int len)\n--\nkernel/bpf/stream.c-50-\t/*\nkernel/bpf/stream.c:51:\t * Allocate a bpf_prog_stream_elem and push it to the bpf_prog_stream\nkernel/bpf/stream.c-52-\t * log, elements will be popped at once and reversed to print the log.\n--\nkernel/bpf/stream.c=207=static int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)\n--\nkernel/bpf/stream.c-255-\nkernel/bpf/stream.c:256:int bpf_prog_stream_read(struct bpf_prog *prog, enum bpf_stream_id stream_id, void __user *buf, u32 len)\nkernel/bpf/stream.c-257-{\n--\nkernel/bpf/stream.c=348=static const struct file_operations bpf_stream_fops = {\n--\nkernel/bpf/stream.c-353-\nkernel/bpf/stream.c:354:int bpf_prog_stream_new_fd(struct bpf_prog *prog, enum bpf_stream_id stream_id, u32 flags)\nkernel/bpf/stream.c-355-{\n--\nkernel/bpf/stream.c=436=__bpf_kfunc_end_defs();\n--\nkernel/bpf/stream.c-439-\nkernel/bpf/stream.c:440:int bpf_prog_stream_init(struct bpf_prog *prog, gfp_t gfp_extra_flags)\nkernel/bpf/stream.c-441-{\n--\nkernel/bpf/stream.c-446-\nkernel/bpf/stream.c:447:\t\t/* On failure, bpf_prog_stream_free() releases the streams allocated so far. */\nkernel/bpf/stream.c-448-\t\tstream = kzalloc_obj(*stream,\n--\nkernel/bpf/stream.c-462-\nkernel/bpf/stream.c:463:void bpf_prog_stream_free(struct bpf_prog *prog)\nkernel/bpf/stream.c-464-{\n--\nkernel/bpf/syscall.c=2929=static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_attr *attr_log)\n--\nkernel/bpf/syscall.c-3082-\nkernel/bpf/syscall.c:3083:\terr = bpf_prog_stream_init(prog, GFP_USER);\nkernel/bpf/syscall.c-3084-\tif (err)\n--\nkernel/bpf/syscall.c=6257=static int prog_stream_read(union bpf_attr *attr)\n--\nkernel/bpf/syscall.c-6270-\nkernel/bpf/syscall.c:6271:\tret = bpf_prog_stream_read(prog, attr-\u003eprog_stream_read.stream_id, buf, len);\nkernel/bpf/syscall.c-6272-\tbpf_prog_put(prog);\n--\nkernel/bpf/syscall.c=6322=static int prog_stream_open(union bpf_attr *attr)\n--\nkernel/bpf/syscall.c-6336-\nkernel/bpf/syscall.c:6337:\tret = bpf_prog_stream_new_fd(prog, attr-\u003eprog_stream_open.stream_id, flags);\nkernel/bpf/syscall.c-6338-\tbpf_prog_put(prog);\n"
} |
| 43/3 | 2026/09/26 22:20 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37537 output: 37 thoughts: 32 Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue: **Diving into `prog_stream_read`** Alright, time to get a bit deeper into this. I'm going to take a look at `prog_stream_read`, specifically in `kernel/bpf/syscall.c`. Need to understand exactly how it's functioning, what it's doing in terms of data copying, and potential areas for optimization or vulnerabilities. This is core functionality, so a solid understanding of its mechanics is critical for overall system stability and performance in any bpf related area. Let's see how things are really implemented under the hood. I suspect there may be interesting implications given the context. I'll focus on the actual data movement involved. Let's see what we're dealing with here. |
| 44/3 | 2026/09/26 22:20 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/syscall.c",
"FirstLine": 6255,
"LineCount": 30
}Results: {
"Contents": "6255:\t#define BPF_PROG_STREAM_READ_BY_FD_LAST_FIELD prog_stream_read.prog_fd\n6256:\t\n6257:\tstatic int prog_stream_read(union bpf_attr *attr)\n6258:\t{\n6259:\t\tchar __user *buf = u64_to_user_ptr(attr-\u003eprog_stream_read.stream_buf);\n6260:\t\tu32 len = attr-\u003eprog_stream_read.stream_buf_len;\n6261:\t\tstruct bpf_prog *prog;\n6262:\t\tint ret;\n6263:\t\n6264:\t\tif (CHECK_ATTR(BPF_PROG_STREAM_READ_BY_FD))\n6265:\t\t\treturn -EINVAL;\n6266:\t\n6267:\t\tprog = bpf_prog_get(attr-\u003eprog_stream_read.prog_fd);\n6268:\t\tif (IS_ERR(prog))\n6269:\t\t\treturn PTR_ERR(prog);\n6270:\t\n6271:\t\tret = bpf_prog_stream_read(prog, attr-\u003eprog_stream_read.stream_id, buf, len);\n6272:\t\tbpf_prog_put(prog);\n6273:\t\n6274:\t\treturn ret;\n6275:\t}\n6276:\t\n6277:\t#define BPF_PROG_ASSOC_STRUCT_OPS_LAST_FIELD prog_assoc_struct_ops.flags\n6278:\t\n6279:\tstatic int prog_assoc_struct_ops(union bpf_attr *attr)\n6280:\t{\n6281:\t\tstruct bpf_prog *prog;\n6282:\t\tstruct bpf_map *map;\n6283:\t\tint ret;\n6284:\t\n"
} |
| 45/3 | 2026/09/26 22:20 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 38276 output: 31 thoughts: 529 Thoughts: **Analyzing `bpf_stream_file_read` for Potential Integer Overflow** Okay, let's dive into this `bpf_stream_file_read` function in `kernel/bpf/stream.c`. I need to make sure this is robust. First, I see the basic structure: It's handling a file read operation, which is good. I need to make sure there's no way to cause an issue by providing a large value for the read length. I spot something immediately: The function takes a `size_t len`, which is the requested read length from the user space. But hang on, the crucial part is the call to `bpf_stream_read`. `bpf_stream_read` takes an *`int len`*! This is a potential vulnerability. Here's the problem: `bpf_stream_file_read` casts the `size_t len` to an `int` when calling `bpf_stream_read`. So, if `len` is greater than `INT_MAX`, what happens? It's an integer truncation, possibly wrapping around. Looking at `bpf_stream_read`, the code does a `min(len, bpf_stream_readable_bytes(stream))`. The `bpf_stream_readable_bytes` clearly returns an int value. Also, `atomic_read_acquire(&stream->readable)` indicates the use of `atomic_t readable`, which is bounded by `BPF_STREAM_MAX_CAPACITY`. This suggests the maximum read size is constrained. But is this enough? If we can supply a `len` larger than `INT_MAX`, then the `len` value passed to `bpf_stream_read` will be a truncated, small integer. However, there's `min(len, bpf_stream_readable_bytes(stream))`, where `bpf_stream_readable_bytes(stream)` is an int. This seems to mean we're safe, since any len that is larger than the int will only ever use the size of the readable bytes. Let's make sure, just in case, that `BPF_STREAM_MAX_CAPACITY` is sufficiently bounded. Let me grep for that constant. |
| 46/3 | 2026/09/26 22:20 | tool | grepper |
0mArgs:{
"Expression": "BPF_STREAM_MAX_CAPACITY",
"PathPrefix": "include/"
}Results: {
"Output": "include/linux/bpf.h=1768=enum {\ninclude/linux/bpf.h-1769-\t/* 100k bytes */\ninclude/linux/bpf.h:1770:\tBPF_STREAM_MAX_CAPACITY = 100000ULL,\ninclude/linux/bpf.h-1771-};\n"
} |
| 47/3 | 2026/09/26 22:20 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 39421 output: 36 thoughts: 676 Thoughts: **Analyzing Potential Uninitialized Memory in BPF Stream Read** Okay, so I'm looking at this code and trying to figure out if there's any way uninitialized memory could be read. The initial thought is, based on the `readable` limit of 100,000 bytes, `min(len, bpf_stream_readable_bytes(stream))` should cap the read to a reasonable size, right? Even if `len` is somehow messed up – say, truncated to a negative value when cast from `size_t` to `int` – the `min` operation shouldn't lead to reading uninitialized memory. If `len` is greater than `INT_MAX`, and cast to an `int`, it becomes a negative number. This means `min(negative number, positive number)` becomes a negative number, which would then be assigned to `rem_len`. However, the code then skips the `while` loop entirely since `rem_len` is negative, returning zero. So that also seems safe. But, I need to be *sure*. Now I'm focused on the `copy_to_user` call within `bpf_stream_read`. Specifically, the part where the data is copied: `copy_to_user(buf + pos, elem->str + cons_len, chunk);`. My critical question is: where does `elem->str` get its data, and is it always initialized before this copy? Digging into the `bpf_stream_elem_alloc` function reveals the allocation details. It determines the allocation size based on the provided length `len`. The crucial line is: `alloc_size = offsetof(struct bpf_stream_elem, str[len]);`. And, it allocates with `kmalloc_nolock(__GFP_ZERO, -1);`, which initializes the memory to zero. This seems promising. Looking further, `bpf_stream_push_str` is the function that actually *populates* the `elem->str` buffer. Here, we see `memcpy(elem->str, str, len);`. So, it copies *exactly* `len` bytes from `str` into `elem->str`. It's a precise copy operation. And, crucially, `elem->total_len = len; elem->consumed_len = 0;`, meaning the intended range `[0..len-1]` is fully populated. So, the memory in `elem->str` *should* be initialized. Plus, remember that `__GFP_ZERO` guarantees zero initialization during allocation. I still need to fully check `bpf_stream_consume_elem` just to be completely certain! Time to review the code. |
| 48/3 | 2026/09/26 22:20 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/stream.c",
"FirstLine": 180,
"LineCount": 35
}Results: {
"Contents": " 180:\t\ttail = llist_del_all(\u0026stream-\u003elog);\n 181:\t\tif (!tail)\n 182:\t\t\treturn;\n 183:\t\thead = llist_reverse_order(tail);\n 184:\t\n 185:\t\tif (!stream-\u003ebacklog_head) {\n 186:\t\t\tstream-\u003ebacklog_head = head;\n 187:\t\t\tstream-\u003ebacklog_tail = tail;\n 188:\t\t} else {\n 189:\t\t\tstream-\u003ebacklog_tail-\u003enext = head;\n 190:\t\t\tstream-\u003ebacklog_tail = tail;\n 191:\t\t}\n 192:\t\n 193:\t\treturn;\n 194:\t}\n 195:\t\n 196:\tstatic bool bpf_stream_consume_elem(struct bpf_stream_elem *elem, int *len)\n 197:\t{\n 198:\t\tint rem = elem-\u003etotal_len - elem-\u003econsumed_len;\n 199:\t\tint used = min(rem, *len);\n 200:\t\n 201:\t\telem-\u003econsumed_len += used;\n 202:\t\t*len -= used;\n 203:\t\n 204:\t\treturn elem-\u003econsumed_len == elem-\u003etotal_len;\n 205:\t}\n 206:\t\n 207:\tstatic int bpf_stream_read(struct bpf_stream *stream, void __user *buf, int len)\n 208:\t{\n 209:\t\tint read_len, rem_len, cons_len, ret = 0;\n 210:\t\tstruct bpf_stream_elem *elem = NULL;\n 211:\t\tstruct llist_node *node;\n 212:\t\n 213:\t\tmutex_lock(\u0026stream-\u003elock);\n 214:\t\tread_len = min(len, bpf_stream_readable_bytes(stream));\n"
} |
| 49/3 | 2026/09/26 22:20 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 41223 output: 407 thoughts: 4140 Thoughts: **Analysis of BPF Stream Log Implementation for Uninitialized Memory Risks** Okay, so I've gone through this code with a fine-tooth comb, focusing intensely on potential uninitialized memory issues. My primary concern, of course, is preventing any kernel memory leaks to user space. I've broken down each function and data structure to see where the risk lies. First, I looked at the data copy in `bpf_stream_file_read` and ensured only initialized bytes are copied by making sure that the relevant variable `used` precisely tracks the data written and consumed. No surprises there. Then, I dug deeper. The `bpf_stream_file_read`, `bpf_stream_poll`, `bpf_prog_stream_new_fd`, and `bpf_prog_stream_init` functions: each was analyzed meticulously. The important data structures `struct bpf_stream` and associated internal buffers are now allocated with `kzalloc_obj`, so everything starts zeroed out. This is *critical* because it gives us a baseline. The reference counts, mutexes, wait queues, everything. Initialized at creation. Next, I looked at the resource cleanup in `bpf_prog_stream_free` and `bpf_stream_put`. The refcounting logic is solid, with proper increments and decrements ensuring that the streams are freed properly, and there doesn't seem to be a race condition. It seems the code prevents use-after-free and double frees correctly. The same applies to stream stages. Now, for `bpf_stream_stage_printk`, it won't push empty outputs, which is good. And I double-checked the error handling of `bpf_stream_get` to be sure that calls to the stream with null streams return errors. I then looked at `bpf_prog_stream_new_fd`, `bpf_prog_stream_read`, and `bpf_stream_vprintk`. All of these are well-behaved if there's no stream. They handle the case of accessing a NULL stream gracefully. Then, the really important part: I focused on the KMSAN checklist. Am I exposing uninitialized kernel data to user space? No. There's nothing going on there. No conditional logic that relies on uninitialized data. Allocation and initialization of complex data structures is carefully handled, with `kzalloc_obj` ensuring zero-initialization. And, crucially, there's nothing in how I'm handling bounds or lengths that could lead to uninitialized bytes being read. The lengths are properly tracked. I had to be certain. I rechecked every data flow to ensure there are no subtle info leaks. The code is structured correctly, using `kzalloc_obj` for zero-initialization and that the data copied to userspace is carefully managed and initialized. After that extensive assessment, I asked myself if I needed KMSAN. No. Everything here can be caught by KASAN, LOCKDEP, and standard kernel debugging tools. The refcounting issues will trigger standard KASAN checks, any locking problems will be flagged by LOCKDEP and UBsan checks, and memory corruption will be caught by the other tools. This patch doesn't introduce *new* uninitialized memory problems that *only* KMSAN could find. Therefore, `NeedsKMSAN` is false. This code is safe from uninitialized memory leaks. I am confident. |
| 50/3 | 2026/09/26 22:20 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch introduces the BPF_PROG_STREAM_OPEN command to the bpf() syscall, allowing user space to open a file descriptor for BPF streams with read() and poll() capabilities. It changes struct bpf_stream from being statically embedded in bpf_prog_aux to being dynamically allocated with refcounting.\n\nAnalysis of uninitialized memory risks (KMSAN):\n1. All dynamic allocations of struct bpf_stream use kzalloc_obj(), ensuring all fields (including atomic counters, flags like notify_used and dead, and backlog pointers) are fully zero-initialized before being passed to their respective init functions (refcount_set, init_llist_head, mutex_init, init_waitqueue_head, init_irq_work).\n2. Data transmitted to user space via copy_to_user() in bpf_stream_read() comes from stream element string buffers (elem-\u003estr), which are zero-allocated (__GFP_ZERO) and fully initialized via memcpy() upon creation. The readable counter only accounts for committed bytes, preventing access to uninitialized bytes.\n3. No struct copies containing uninitialized fields or alignment padding are returned to user space (the syscall only returns a file descriptor integer).\n\nApplicability of KASAN vs KMSAN:\nThe primary risks introduced by this change relate to object lifetimes, reference counting (bpf_stream_put / refcount_inc), concurrent readers/writers, and synchronization across irq_work and anon_inode file operations. Any potential bugs in these areas (such as use-after-free, double-free, or locking issues) are effectively detected by KASAN, LOCKDEP, and standard kernel debugging facilities. No uninitialized memory usage or info-leak risks are present, so a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch introduces the BPF_PROG_STREAM_OPEN command to the bpf() syscall, allowing user space to open a file descriptor for BPF streams with read() and poll() capabilities. It changes struct bpf_stream from being statically embedded in bpf_prog_aux to being dynamically allocated with refcounting.\n\nAnalysis of uninitialized memory risks (KMSAN):\n1. All dynamic allocations of struct bpf_stream use kzalloc_obj(), ensuring all fields (including atomic counters, flags like notify_used and dead, and backlog pointers) are fully zero-initialized before being passed to their respective init functions (refcount_set, init_llist_head, mutex_init, init_waitqueue_head, init_irq_work).\n2. Data transmitted to user space via copy_to_user() in bpf_stream_read() comes from stream element string buffers (elem-\u003estr), which are zero-allocated (__GFP_ZERO) and fully initialized via memcpy() upon creation. The readable counter only accounts for committed bytes, preventing access to uninitialized bytes.\n3. No struct copies containing uninitialized fields or alignment padding are returned to user space (the syscall only returns a file descriptor integer).\n\nApplicability of KASAN vs KMSAN:\nThe primary risks introduced by this change relate to object lifetimes, reference counting (bpf_stream_put / refcount_inc), concurrent readers/writers, and synchronization across irq_work and anon_inode file operations. Any potential bugs in these areas (such as use-after-free, double-free, or locking issues) are effectively detected by KASAN, LOCKDEP, and standard kernel debugging facilities. No uninitialized memory usage or info-leak risks are present, so a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|