| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/10/10 23:15 | flow | patch-triage |
2mArgs:null Results: {
"EnableConfigs": [
"BPF_SYSCALL"
],
"FocusSymbols": [
"arena_map_alloc",
"arena_vm_fault",
"arena_get_unmapped_area",
"arena_map_mmap",
"zap_pages",
"bpf_arena_free_pages",
"bpf_arena_free_pages_non_sleepable",
"bpf_map_setattr",
"bpffs_arena_open",
"bpffs_arena_mmap",
"bpffs_arena_llseek",
"bpffs_arena_get_unmapped_area",
"bpf_mkmap"
],
"KMSANReasoning": "The patch series introduces BPF_F_ARENA_NO_FREE and BPF_F_ARENA_EXPORT flags for BPF arenas, allowing pinned arenas in bpffs to be opened and mapped as sized files with exported slices.\n\n1. Uninitialized Memory Analysis (KMSAN):\n- No kernel stack or heap structures are copied to userspace (no copy_to_user, put_user, netlink attributes, ioctl out-arguments, etc.).\n- File operations on the exported bpffs inode (bpffs_arena_fops) do not implement .read or .write handlers; memory access occurs solely via memory-mapping (mmap/fault).\n- Backing pages allocated during fault handling (arena_vm_fault) continue to use GFP_KERNEL | __GFP_ZERO.\n- All new stack variables in modified functions (e.g. start, end, vma_start, vma_end in zap_pages, and err in bpffs_arena_open) are explicitly initialized prior to evaluation.\n- No uninitialized memory or info-leak risk is introduced or exposed.\n\n2. Applicable Sanitizers (KASAN, LOCKDEP, DEBUG_VM):\n- Potential risks in this change involve offset and address calculations (e.g. kaddr calculation in arena_vm_fault and slice offsets in zap_pages), which could risk out-of-bounds accesses in the vmalloc area (caught by KASAN).\n- Reference counting and lifecycle management of the map across open/mmap/release/unlink (bpf_map_inc_with_uref, UAF or leak) are covered by KASAN and standard refcount tracking.\n- Locking and concurrency interactions (arena-\u003elock, arena-\u003espinlock) are monitored by LOCKDEP.\n\nTherefore, standard KASAN and debugging builds fully cover the risk profile of these changes, and a dedicated KMSAN fuzzing session is not warranted (NeedsKMSAN=false).",
"NeedsKMSAN": false,
"Reasoning": "The patch introduces export and retention support for BPF arena maps (BPF_F_ARENA_EXPORT and BPF_F_ARENA_NO_FREE), allowing pinned arenas in bpffs to be opened and mmap'd as shared-memory files with slice offsets. It modifies core BPF arena memory management and bpffs inode operations, altering VMA tracking, page fault resolution (accounting for slice offsets), unmapped area discovery, range zapping, and file operations (open, mmap, setattr, llseek). These paths are reachable from userspace in standard virtualized environments via standard syscalls.",
"WorthFuzzing": true
} |
| 1/1 | 2026/10/10 23:15 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit a034d16eea30a35c1529eb1fb108e6ad8ca18920\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Sat Oct 10 23:15:04 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/Documentation/bpf/map_arena.rst b/Documentation/bpf/map_arena.rst\nnew file mode 100644\nindex 0000000000000..1b9eadd5fb3a1\n--- /dev/null\n+++ b/Documentation/bpf/map_arena.rst\n@@ -0,0 +1,131 @@\n+.. SPDX-License-Identifier: GPL-2.0-only\n+\n+==================\n+BPF_MAP_TYPE_ARENA\n+==================\n+\n+A BPF arena provides shared memory that BPF programs and userspace can\n+access directly. Create the map with zero key and value sizes and\n+``BPF_F_MMAPABLE``. ``max_entries`` specifies its capacity in pages, up to\n+4 GiB. The canonical BPF address range must not cross a 4 GiB boundary.\n+\n+Pinned arena files\n+==================\n+\n+An arena created with ``BPF_F_ARENA_EXPORT`` can be pinned in bpffs and\n+opened as a sized file. Its size is ``max_entries * PAGE_SIZE``. Export\n+requires ``BPF_F_ARENA_NO_FREE`` so BPF programs cannot release backing\n+pages while external consumers may still use them. Creating an arena with\n+``BPF_F_ARENA_EXPORT`` alone fails with ``EINVAL``. Arenas without the\n+export flag retain their existing bpffs behavior and cannot be opened\n+this way, including arenas created with ``BPF_F_ARENA_NO_FREE`` alone.\n+\n+Establish the canonical BPF address range before mapping the pinned file:\n+\n+* Set ``map_extra`` to a nonzero, page-aligned address at map creation.\n+ The canonical range then covers the full map capacity, even if the map\n+ FD has not been mmaped.\n+* Alternatively, leave ``map_extra`` zero and mmap the map FD first. This\n+ mapping establishes the canonical address and must cover the full map\n+ capacity. Shorter canonical mappings fail with ``EINVAL``.\n+\n+A canonical map FD mapping may start at address zero if the system's\n+low-address mapping policy permits it.\n+\n+Mapping the pinned file before either step fails with ``EINVAL``. Exported\n+mappings do not establish or change the canonical BPF address range.\n+\n+Once the canonical range is established, it covers the entire capacity\n+reported by ``fstat()``. Consumers can map the full file or smaller slices.\n+Establishing the full canonical mapping reserves virtual address space;\n+it does not populate every arena page. Arenas without\n+``BPF_F_ARENA_EXPORT`` can still establish shorter canonical mappings.\n+\n+Open the pin with ``O_RDONLY`` for read-only access or ``O_RDWR`` for\n+writable access, and mmap it with ``MAP_SHARED``. An exported\n+mapping may use a different virtual address from the canonical mapping.\n+Its offset must be page-aligned, and its offset and length must fit within\n+both the map capacity and the established canonical range. Consumers can\n+map disjoint slices of one arena. Oversized or out-of-range mappings fail\n+with ``EINVAL``. ``MAP_PRIVATE`` mappings are unsupported.\n+\n+Exported mappings share the bytes of the arena without relocating pointers\n+stored in those bytes. Arena pointers produced by a BPF program use that\n+arena's canonical address representation, based on ``user_vm_start``.\n+A consumer mapping a slice at another address must translate such pointers\n+before dereferencing them. A guest BPF arena has its own canonical range;\n+sharing the backing pages does not make host arena pointers valid in the\n+guest arena, or vice versa. Communication structures can use offsets\n+relative to the shared slice, with each side validating the offsets and\n+translating them to addresses in its own mapping or arena.\n+\n+A pin opened with ``O_RDONLY`` supports read-only shared mappings, which\n+observe updates made by BPF programs and other writable mappings. Requests\n+for ``PROT_WRITE`` through that FD fail with ``EACCES``. A mapping made\n+without ``PROT_WRITE`` cannot subsequently acquire write permission with\n+``mprotect()``, even when the pin was opened with ``O_RDWR``.\n+\n+Pin permissions can grant readers access independently of writers. For\n+example, mode ``0640`` permits the owner to open the pin for writable\n+access and the group to open it for read-only access, subject to the\n+applicable LSM checks. Changing permissions does not revoke access through\n+existing FDs or mappings.\n+\n+The file has a fixed size. ``truncate()``, ``ftruncate()``, and ``O_TRUNC``\n+cannot change it. A truncate request that leaves the size unchanged is\n+allowed. Permission and ownership changes remain available.\n+Consumers can discover the capacity with ``fstat()`` or\n+``lseek(fd, 0, SEEK_END)``. Seeking supports ``SEEK_SET``, ``SEEK_CUR``, and\n+``SEEK_END`` within the file bounds; it does not enable ``read()`` or\n+``write()`` access to the memory.\n+\n+``fsync()``, ``fdatasync()``, and ``msync(..., MS_SYNC)`` succeed without\n+performing writeback: arena memory is volatile and has no persistent\n+backing. These calls do not synchronize access between consumers or BPF\n+programs; shared-memory protocols still need appropriate memory ordering.\n+\n+Exported mappings inherit the arena's mapping lifecycle restrictions.\n+They are not inherited across ``fork()``, and ``MADV_DOFORK`` cannot enable\n+inheritance. Consumers must create their own mappings in a child process.\n+Mappings cannot be split, expanded, or relocated with ``mremap()``.\n+Partial unmapping and protection changes that require splitting a mapping\n+fail with ``EINVAL``. Unmapping an entire mapping remains supported, as do\n+protection changes that do not require a split and satisfy the access\n+restrictions above. Consumers needing independently managed regions\n+should create separate slice mappings.\n+\n+Access control and lifetime\n+===========================\n+\n+Opening the pin uses ordinary filesystem permissions and file LSM checks,\n+as well as ``security_bpf_map()`` for the requested access mode. Mapping it\n+also passes the ordinary mmap LSM checks. Writable mappings remain subject\n+to the BPF map's frozen state and program-read-only restrictions.\n+\n+Consumers use ``open()`` and ``mmap()`` rather than ``bpf(BPF_OBJ_GET)``.\n+Consequently, seccomp rules and LSM policies specific to the ``bpf()``\n+syscall do not govern this file interface. Grant access through the pin's\n+permissions and the applicable file, mmap, and BPF map LSM policies.\n+A process prohibited from calling ``bpf()`` can still access the arena\n+through this interface if those permissions and checks allow it.\n+``unprivileged_bpf_disabled`` restricts unprivileged map and program\n+creation; it does not generally prohibit access to existing maps.\n+\n+``BPF_F_ARENA_NO_FREE`` makes ``bpf_arena_free_pages()`` a no-op in both\n+sleepable and non-sleepable BPF programs. The free operation returns no\n+indication that pages were retained. Programs using this flag must not\n+rely on freeing pages to restore arena allocation capacity. Arena pages\n+are retained until the map is destroyed. Open map and pin FDs, mappings,\n+and other map references retain the arena, including after the pin is\n+unlinked. Removing the pin therefore does not revoke existing mappings or\n+free their pages.\n+\n+The retention flag does not itself expose the pin as a sized file.\n+``BPF_F_ARENA_EXPORT`` opts into that interface and must be combined with\n+``BPF_F_ARENA_NO_FREE``. Retention-only arenas can still share memory\n+through ordinary map FD mappings and manage reusable objects within\n+their retained pages.\n+\n+For example, a consumer such as QEMU can use an exported slice as a shared\n+file-backed guest RAM backend. The file interface provides shared memory;\n+consumers must supply their own allocation and communication protocol.\ndiff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h\nindex e0ed44b1bbcb7..40a5489ac9fea 100644\n--- a/include/uapi/linux/bpf.h\n+++ b/include/uapi/linux/bpf.h\n@@ -1500,6 +1500,12 @@ enum {\n \n /* Enable BPF ringbuf overwrite mode */\n \tBPF_F_RB_OVERWRITE\t= (1U \u003c\u003c 19),\n+\n+\t/* Keep arena pages allocated until the map is destroyed. */\n+\tBPF_F_ARENA_NO_FREE\t= (1U \u003c\u003c 20),\n+\n+\t/* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */\n+\tBPF_F_ARENA_EXPORT\t= (1U \u003c\u003c 21),\n };\n \n /* Flags for BPF_PROG_QUERY. */\ndiff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c\nindex 0ff707707da04..98f6dffe0138f 100644\n--- a/kernel/bpf/arena.c\n+++ b/kernel/bpf/arena.c\n@@ -283,7 +283,13 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)\n \t /* BPF_F_MMAPABLE must be set */\n \t !(attr-\u003emap_flags \u0026 BPF_F_MMAPABLE) ||\n \t /* No unsupported flags present */\n-\t (attr-\u003emap_flags \u0026 ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE | BPF_F_NO_USER_CONV)))\n+\t (attr-\u003emap_flags \u0026 ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE |\n+\t\t\t\t BPF_F_NO_USER_CONV | BPF_F_ARENA_NO_FREE |\n+\t\t\t\t BPF_F_ARENA_EXPORT)))\n+\t\treturn ERR_PTR(-EINVAL);\n+\n+\t/* Exported memory must remain allocated while consumers use it. */\n+\tif ((attr-\u003emap_flags \u0026 BPF_F_ARENA_EXPORT) \u0026\u0026 !(attr-\u003emap_flags \u0026 BPF_F_ARENA_NO_FREE))\n \t\treturn ERR_PTR(-EINVAL);\n \n \tif (attr-\u003emap_extra \u0026 ~PAGE_MASK)\n@@ -495,7 +501,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)\n \tint ret;\n \n \tkbase = bpf_arena_get_kern_vm_start(arena);\n-\tkaddr = kbase + (u32)(vmf-\u003eaddress);\n+\t/* vmf-\u003epgoff includes the file offset of a bpffs-backed slice. */\n+\tkaddr = kbase + (u32)(arena-\u003euser_vm_start + ((u64)vmf-\u003epgoff \u003c\u003c PAGE_SHIFT));\n \n \tpage = vmalloc_to_page((void *)kaddr);\n \tif (!page \u0026\u0026 !(arena-\u003emap.map_flags \u0026 BPF_F_SEGV_ON_FAULT)) {\n@@ -612,13 +619,23 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad\n \tstruct bpf_arena *arena = container_of(map, struct bpf_arena, map);\n \tlong ret;\n \n+\tif (filp-\u003ef_op != \u0026bpf_map_fops) {\n+\t\tif (!len || pgoff \u003e= map-\u003emax_entries ||\n+\t\t len \u003e ((u64)map-\u003emax_entries - pgoff) \u003c\u003c PAGE_SHIFT)\n+\t\t\treturn -EINVAL;\n+\t\treturn mm_get_unmapped_area(filp, addr, len, pgoff, flags);\n+\t}\n+\n \tif (pgoff)\n \t\treturn -EINVAL;\n \tif (len \u003e SZ_4G)\n \t\treturn -E2BIG;\n+\tif ((map-\u003emap_flags \u0026 BPF_F_ARENA_EXPORT) \u0026\u0026\n+\t len != (u64)map-\u003emax_entries \u003c\u003c PAGE_SHIFT)\n+\t\treturn -EINVAL;\n \n-\t/* if user_vm_start was specified at arena creation time */\n-\tif (arena-\u003euser_vm_start) {\n+\t/* Once established, the canonical range cannot change. */\n+\tif (arena-\u003euser_vm_end) {\n \t\tif (len \u003e arena-\u003euser_vm_end - arena-\u003euser_vm_start)\n \t\t\treturn -E2BIG;\n \t\tif (len != arena-\u003euser_vm_end - arena-\u003euser_vm_start)\n@@ -632,7 +649,7 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad\n \t\treturn ret;\n \tif ((ret \u003e\u003e 32) == ((ret + len - 1) \u003e\u003e 32))\n \t\treturn ret;\n-\tif (WARN_ON_ONCE(arena-\u003euser_vm_start))\n+\tif (WARN_ON_ONCE(arena-\u003euser_vm_end))\n \t\t/* checks at map creation time should prevent this */\n \t\treturn -EFAULT;\n \treturn round_up(ret, SZ_4G);\n@@ -641,9 +658,13 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad\n static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)\n {\n \tstruct bpf_arena *arena = container_of(map, struct bpf_arena, map);\n+\tbool exported = vma-\u003evm_file-\u003ef_op != \u0026bpf_map_fops;\n \n \tguard(mutex)(\u0026arena-\u003elock);\n-\tif (arena-\u003euser_vm_start \u0026\u0026 arena-\u003euser_vm_start != vma-\u003evm_start)\n+\tif (!exported \u0026\u0026 (map-\u003emap_flags \u0026 BPF_F_ARENA_EXPORT) \u0026\u0026\n+\t vma-\u003evm_end - vma-\u003evm_start != (u64)map-\u003emax_entries \u003c\u003c PAGE_SHIFT)\n+\t\treturn -EINVAL;\n+\tif (!exported \u0026\u0026 arena-\u003euser_vm_end \u0026\u0026 arena-\u003euser_vm_start != vma-\u003evm_start)\n \t\t/*\n \t\t * If map_extra was not specified at arena creation time then\n \t\t * 1st user process can do mmap(NULL, ...) to pick user_vm_start\n@@ -654,19 +675,31 @@ static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)\n \t\t */\n \t\treturn -EBUSY;\n \n-\tif (arena-\u003euser_vm_end \u0026\u0026 arena-\u003euser_vm_end != vma-\u003evm_end)\n+\tif (exported \u0026\u0026 !arena-\u003euser_vm_end)\n+\t\treturn -EINVAL;\n+\tif (!exported \u0026\u0026 arena-\u003euser_vm_end \u0026\u0026 arena-\u003euser_vm_end != vma-\u003evm_end)\n \t\t/* all user processes must have the same size of mmap-ed region */\n \t\treturn -EBUSY;\n \n-\t/* Earlier checks should prevent this */\n-\tif (WARN_ON_ONCE(vma-\u003evm_end - vma-\u003evm_start \u003e SZ_4G || vma-\u003evm_pgoff))\n+\tif (exported) {\n+\t\tu64 page_cnt = (arena-\u003euser_vm_end - arena-\u003euser_vm_start) \u003e\u003e PAGE_SHIFT;\n+\n+\t\tif (vma-\u003evm_pgoff \u003e= page_cnt ||\n+\t\t (vma-\u003evm_end - vma-\u003evm_start) \u003e\u003e PAGE_SHIFT \u003e\n+\t\t page_cnt - vma-\u003evm_pgoff)\n+\t\t\treturn -EINVAL;\n+\t} else if (WARN_ON_ONCE(vma-\u003evm_end - vma-\u003evm_start \u003e SZ_4G || vma-\u003evm_pgoff)) {\n+\t\t/* Earlier checks should prevent this for map FD mappings. */\n \t\treturn -EFAULT;\n+\t}\n \n \tif (remember_vma(arena, vma))\n \t\treturn -ENOMEM;\n \n-\tarena-\u003euser_vm_start = vma-\u003evm_start;\n-\tarena-\u003euser_vm_end = vma-\u003evm_end;\n+\tif (!exported) {\n+\t\tarena-\u003euser_vm_start = vma-\u003evm_start;\n+\t\tarena-\u003euser_vm_end = vma-\u003evm_end;\n+\t}\n \t/*\n \t * bpf_map_mmap() checks that it's being mmaped as VM_SHARED and\n \t * clears VM_MAYEXEC. Set VM_DONTEXPAND to avoid potential change\n@@ -840,8 +873,12 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)\n \tstruct mm_struct *mm;\n \tstruct vma_list *vml;\n \tunsigned long vm_start;\n+\tu64 start, end, vma_start, vma_end;\n \tu64 my_gen;\n \n+\tstart = uaddr - arena-\u003euser_vm_start;\n+\tend = start + size;\n+\n \t/*\n \t * Taking mmap_read_lock() under arena-\u003elock would deadlock against\n \t * arena_vm_close(), which runs with mmap_write_lock held and then\n@@ -880,8 +917,14 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)\n \t\t */\n \t\tvma = find_vma(mm, vm_start);\n \t\tif (vma \u0026\u0026 vma-\u003evm_start == vm_start \u0026\u0026\n-\t\t vma-\u003evm_file \u0026\u0026 vma-\u003evm_file-\u003eprivate_data == \u0026arena-\u003emap)\n-\t\t\tzap_vma_range(vma, uaddr, size);\n+\t\t vma-\u003evm_file \u0026\u0026 vma-\u003evm_file-\u003eprivate_data == \u0026arena-\u003emap) {\n+\t\t\tvma_start = (u64)vma-\u003evm_pgoff \u003c\u003c PAGE_SHIFT;\n+\t\t\tvma_end = vma_start + vma-\u003evm_end - vma-\u003evm_start;\n+\t\t\tif (start \u003c vma_end \u0026\u0026 end \u003e vma_start)\n+\t\t\t\tzap_vma_range(vma, vma-\u003evm_start +\n+\t\t\t\t\t (max(start, vma_start) - vma_start),\n+\t\t\t\t\t min(end, vma_end) - max(start, vma_start));\n+\t\t}\n \t\tmmap_read_unlock(mm);\n \t\tmmput(mm);\n \n@@ -1134,7 +1177,8 @@ __bpf_kfunc void bpf_arena_free_pages(void *p__map, void *ptr__ign, u32 page_cnt\n \tstruct bpf_map *map = p__map;\n \tstruct bpf_arena *arena = container_of(map, struct bpf_arena, map);\n \n-\tif (map-\u003emap_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign)\n+\tif (map-\u003emap_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign ||\n+\t (map-\u003emap_flags \u0026 BPF_F_ARENA_NO_FREE))\n \t\treturn;\n \tarena_free_pages(arena, (long)ptr__ign, page_cnt, true);\n }\n@@ -1144,7 +1188,8 @@ void bpf_arena_free_pages_non_sleepable(void *p__map, void *ptr__ign, u32 page_c\n \tstruct bpf_map *map = p__map;\n \tstruct bpf_arena *arena = container_of(map, struct bpf_arena, map);\n \n-\tif (map-\u003emap_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign)\n+\tif (map-\u003emap_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign ||\n+\t (map-\u003emap_flags \u0026 BPF_F_ARENA_NO_FREE))\n \t\treturn;\n \tarena_free_pages(arena, (long)ptr__ign, page_cnt, false);\n }\ndiff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c\nindex 7837968c0842c..d3dc6f70e32ae 100644\n--- a/kernel/bpf/inode.c\n+++ b/kernel/bpf/inode.c\n@@ -119,8 +119,24 @@ static const struct inode_operations bpf_symlink_iops;\n static const struct inode_operations bpf_prog_iops = {\n \t.listxattr\t= bpf_fs_listxattr,\n };\n+\n+static int bpf_map_setattr(struct mnt_idmap *idmap, struct dentry *dentry,\n+\t\t\t struct iattr *attr)\n+{\n+\tstruct inode *inode = d_inode(dentry);\n+\tstruct bpf_map *map = inode-\u003ei_private;\n+\n+\tif (map-\u003emap_type == BPF_MAP_TYPE_ARENA \u0026\u0026\n+\t (map-\u003emap_flags \u0026 BPF_F_ARENA_EXPORT) \u0026\u0026\n+\t (attr-\u003eia_valid \u0026 ATTR_SIZE) \u0026\u0026 attr-\u003eia_size != i_size_read(inode))\n+\t\treturn -EINVAL;\n+\n+\treturn simple_setattr(idmap, dentry, attr);\n+}\n+\n static const struct inode_operations bpf_map_iops = {\n \t.listxattr\t= bpf_fs_listxattr,\n+\t.setattr\t= bpf_map_setattr,\n };\n static const struct inode_operations bpf_link_iops = {\n \t.listxattr\t= bpf_fs_listxattr,\n@@ -351,6 +367,68 @@ static const struct file_operations bpffs_map_fops = {\n \t.release\t= bpffs_map_release,\n };\n \n+/*\n+ * An arena with BPF_F_ARENA_EXPORT pinned in bpffs can serve as a\n+ * shared-memory file. The ordinary map FD is an anonymous inode with no\n+ * size and cannot be opened by pathname, so keep the pin's inode and\n+ * forward mmap to the map FD implementation. The pinned file uses ordinary\n+ * address selection so QEMU can map the pages at another virtual address.\n+ */\n+static int bpffs_arena_open(struct inode *inode, struct file *file)\n+{\n+\tstruct bpf_map *map = inode-\u003ei_private;\n+\tint err;\n+\n+\terr = security_bpf_map(map, file-\u003ef_mode);\n+\tif (err)\n+\t\treturn err;\n+\tbpf_map_inc_with_uref(map);\n+\tfile-\u003eprivate_data = map;\n+\n+\treturn 0;\n+}\n+\n+static int bpffs_arena_mmap(struct file *file, struct vm_area_struct *vma)\n+{\n+\t/*\n+\t * Generic mmap clears VM_SHARED for O_RDONLY files, but VM_MAYSHARE\n+\t * still distinguishes shared mappings from private ones. Restore\n+\t * VM_SHARED for the map mmap path without granting VM_MAYWRITE.\n+\t */\n+\tif (!(file-\u003ef_mode \u0026 FMODE_WRITE) \u0026\u0026 (vma-\u003evm_flags \u0026 VM_MAYSHARE))\n+\t\tvm_flags_set(vma, VM_SHARED);\n+\n+\treturn bpf_map_fops.mmap(file, vma);\n+}\n+\n+static int bpffs_arena_release(struct inode *inode, struct file *file)\n+{\n+\treturn bpf_map_fops.release(inode, file);\n+}\n+\n+static unsigned long bpffs_arena_get_unmapped_area(struct file *file,\n+\t\t\t\t\t\t unsigned long addr,\n+\t\t\t\t\t\t unsigned long len,\n+\t\t\t\t\t\t unsigned long pgoff,\n+\t\t\t\t\t\t unsigned long flags)\n+{\n+\treturn bpf_map_fops.get_unmapped_area(file, addr, len, pgoff, flags);\n+}\n+\n+static loff_t bpffs_arena_llseek(struct file *file, loff_t offset, int whence)\n+{\n+\treturn fixed_size_llseek(file, offset, whence, i_size_read(file_inode(file)));\n+}\n+\n+static const struct file_operations bpffs_arena_fops = {\n+\t.open\t\t= bpffs_arena_open,\n+\t.llseek\t\t= bpffs_arena_llseek,\n+\t.fsync\t\t= noop_fsync,\n+\t.release\t= bpffs_arena_release,\n+\t.mmap\t\t= bpffs_arena_mmap,\n+\t.get_unmapped_area = bpffs_arena_get_unmapped_area,\n+};\n+\n static int bpffs_obj_open(struct inode *inode, struct file *file)\n {\n \treturn -EIO;\n@@ -362,7 +440,7 @@ static const struct file_operations bpffs_obj_fops = {\n \n static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,\n \t\t\t const struct inode_operations *iops,\n-\t\t\t const struct file_operations *fops)\n+\t\t\t const struct file_operations *fops, loff_t size)\n {\n \tstruct inode *dir = dentry-\u003ed_parent-\u003ed_inode;\n \tstruct inode *inode;\n@@ -382,6 +460,7 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,\n \tinode-\u003ei_op = iops;\n \tinode-\u003ei_fop = fops;\n \tinode-\u003ei_private = raw;\n+\ti_size_write(inode, size);\n \n \tbpf_dentry_finalize(dentry, inode, dir);\n \treturn 0;\n@@ -390,16 +469,20 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,\n static int bpf_mkprog(struct dentry *dentry, umode_t mode, void *arg)\n {\n \treturn bpf_mkobj_ops(dentry, mode, arg, \u0026bpf_prog_iops,\n-\t\t\t \u0026bpffs_obj_fops);\n+\t\t\t \u0026bpffs_obj_fops, 0);\n }\n \n static int bpf_mkmap(struct dentry *dentry, umode_t mode, void *arg)\n {\n \tstruct bpf_map *map = arg;\n+\tbool shared_arena = map-\u003emap_type == BPF_MAP_TYPE_ARENA \u0026\u0026\n+\t\t\t (map-\u003emap_flags \u0026 BPF_F_ARENA_EXPORT);\n \n \treturn bpf_mkobj_ops(dentry, mode, arg, \u0026bpf_map_iops,\n-\t\t\t bpf_map_support_seq_show(map) ?\n-\t\t\t \u0026bpffs_map_fops : \u0026bpffs_obj_fops);\n+\t\t\t shared_arena ? \u0026bpffs_arena_fops :\n+\t\t\t bpf_map_support_seq_show(map) ?\n+\t\t\t \u0026bpffs_map_fops : \u0026bpffs_obj_fops,\n+\t\t\t shared_arena ? (loff_t)map-\u003emax_entries * PAGE_SIZE : 0);\n }\n \n static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)\n@@ -408,7 +491,7 @@ static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)\n \n \treturn bpf_mkobj_ops(dentry, mode, arg, \u0026bpf_link_iops,\n \t\t\t bpf_link_is_iter(link) ?\n-\t\t\t \u0026bpf_iter_fops : \u0026bpffs_obj_fops);\n+\t\t\t \u0026bpf_iter_fops : \u0026bpffs_obj_fops, 0);\n }\n \n static struct dentry *\n@@ -483,7 +566,7 @@ static int bpf_iter_link_pin_kernel(struct dentry *parent,\n \tif (IS_ERR(dentry))\n \t\treturn PTR_ERR(dentry);\n \tret = bpf_mkobj_ops(dentry, mode, link, \u0026bpf_link_iops,\n-\t\t\t \u0026bpf_iter_fops);\n+\t\t\t \u0026bpf_iter_fops, 0);\n \tsimple_done_creating(dentry);\n \treturn ret;\n }\ndiff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h\nindex e0ed44b1bbcb7..40a5489ac9fea 100644\n--- a/tools/include/uapi/linux/bpf.h\n+++ b/tools/include/uapi/linux/bpf.h\n@@ -1500,6 +1500,12 @@ enum {\n \n /* Enable BPF ringbuf overwrite mode */\n \tBPF_F_RB_OVERWRITE\t= (1U \u003c\u003c 19),\n+\n+\t/* Keep arena pages allocated until the map is destroyed. */\n+\tBPF_F_ARENA_NO_FREE\t= (1U \u003c\u003c 20),\n+\n+\t/* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */\n+\tBPF_F_ARENA_EXPORT\t= (1U \u003c\u003c 21),\n };\n \n /* Flags for BPF_PROG_QUERY. */\ndiff --git a/tools/testing/selftests/bpf/.gitignore b/tools/testing/selftests/bpf/.gitignore\nindex b815bf0d88774..9747ad37d698e 100644\n--- a/tools/testing/selftests/bpf/.gitignore\n+++ b/tools/testing/selftests/bpf/.gitignore\n@@ -47,3 +47,5 @@ verification_cert.h\n *.BTF.base\n usdt_1\n usdt_2\n+/arena_kvm_guest-init\n+/arena_kvm_host-runner\ndiff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile\nindex a22be7efd1fa2..29fffa7857c19 100644\n--- a/tools/testing/selftests/bpf/Makefile\n+++ b/tools/testing/selftests/bpf/Makefile\n@@ -42,7 +42,13 @@ TEST_PROGS := test_kmod.sh \\\n \ttest_bpftool_build.sh \\\n \ttest_doc_build.sh \\\n \ttest_xsk.sh \\\n-\ttest_xdp_features.sh\n+\ttest_xdp_features.sh \\\n+\tarena_kvm.py\n+\n+TEST_GEN_FILES += arena_kvm_guest-init arena_kvm_host-runner\n+ifneq ($(CLANG_CPUV4),)\n+TEST_GEN_FILES += arena_kvm_guest.bpf.o arena_kvm_host.bpf.o\n+endif\n \n TEST_PROGS_EXTENDED := \\\n \tima_setup.sh verify_sig_setup.sh\n@@ -94,6 +100,25 @@ BPFTOOLDIR := $(TOOLSDIR)/bpf/bpftool\n HOST_BPFOBJ := $(HOST_BUILD_DIR)/libbpf/libbpf.a\n BPF_TARGET_ENDIAN:=$(if $(IS_LITTLE_ENDIAN),--target=bpfel,--target=bpfeb)\n \n+# These helpers are installed beside arena_kvm.py, including for OUTPUT=.\n+$(OUTPUT)/arena_kvm_%.bpf.o: arena_kvm_%.bpf.c arena_kvm_shared.h \\\n+\t\tlibarena/include/bpf_atomic.h \\\n+\t\tlibarena/include/bpf_arena_common.h \\\n+\t\t$(INCLUDE_DIR)/vmlinux.h $(BPFOBJ) | $(OUTPUT)\n+\t$(call msg,BPF,,$@)\n+\t$(Q)$(CLANG) $(BPF_CFLAGS) $(CLANG_CFLAGS) -O2 \\\n+\t\t$(BPF_TARGET_ENDIAN) -mcpu=v4 -c $\u003c -o $@ $(call skip_on_fail,BPF)\n+\n+$(OUTPUT)/arena_kvm_guest-init: arena_kvm_guest.c arena_kvm_shared.h \\\n+\t\t$(BPFOBJ) | $(OUTPUT)\n+\t$(call msg,BINARY,,$@)\n+\t$(Q)$(CC) $(CFLAGS) $(LDFLAGS) $\u003c $(BPFOBJ) $(LDLIBS) -lzstd -o $@\n+\n+$(OUTPUT)/arena_kvm_host-runner: arena_kvm_host.c arena_kvm_shared.h \\\n+\t\t$(BPFOBJ) | $(OUTPUT)\n+\t$(call msg,BINARY,,$@)\n+\t$(Q)$(CC) $(CFLAGS) $(LDFLAGS) $\u003c $(BPFOBJ) $(LDLIBS) -lzstd -o $@\n+\n NON_CHECK_FEAT_TARGETS := clean docs-clean emit_tests\n CHECK_FEAT := $(filter-out $(NON_CHECK_FEAT_TARGETS),$(or $(MAKECMDGOALS), \"none\"))\n ifneq ($(CHECK_FEAT),)\ndiff --git a/tools/testing/selftests/bpf/arena_kvm.py b/tools/testing/selftests/bpf/arena_kvm.py\nnew file mode 100755\nindex 0000000000000..552bfcfb80ac8\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/arena_kvm.py\n@@ -0,0 +1,448 @@\n+#!/usr/bin/env python3\n+# SPDX-License-Identifier: GPL-2.0\n+# Copyright (c) 2026 NVIDIA CORPORATION \u0026 AFFILIATES\n+\"\"\"Exchange values between host and guest BPF through NUMA-backed RAM.\n+\n+Build with the BPF selftests Makefile. Helpers are found in the current\n+directory or beside this script, so OUTPUT builds and installed tests work.\n+Use --build-dir to select another helper directory and --kernel (or\n+ARENA_KVM_KERNEL) to select the kernel booted by virtme-ng. By default, use\n+the source tree's kernel when available, otherwise the running release.\n+\"\"\"\n+\n+# Dependencies: Python 3.9+, virtme-ng, busybox-static, qemu, util-linux (script).\n+\n+import argparse\n+import array\n+import ctypes\n+import mmap\n+import os\n+from pathlib import Path\n+import queue\n+import re\n+import shutil\n+import shlex\n+import signal\n+import socket\n+import struct\n+import subprocess\n+import sys\n+import threading\n+import time\n+\n+\n+PAGE_SIZE = os.sysconf(\"SC_PAGE_SIZE\")\n+KSFT_SKIP = 4\n+HELPER_TIMEOUT = 20\n+GUEST_MEMORY = 512 * 1024 * 1024\n+# RAM starts for virtme-ng's PC/q35, virt (arm64/riscv64), and pseries\n+# machines. With 512 MiB of RAM, the last NUMA node follows node 0\n+# without crossing a machine's RAM hole. Skip other layouts.\n+GUEST_RAM_BASES = {\n+ \"x86_64\": 0,\n+ \"aarch64\": 0x40000000,\n+ \"arm64\": 0x40000000,\n+ \"riscv64\": 0x80000000,\n+ \"ppc64\": 0,\n+ \"ppc64le\": 0,\n+}\n+# Leave room for guest boot allocations and allocator watermarks.\n+# QEMU's pseries machine requires each NUMA node to be 256 MiB aligned.\n+SLICE_SIZE = (256 if os.uname().machine in (\"ppc64\", \"ppc64le\") else 16) \\\n+ * 1024 * 1024\n+ARENA_SIZE = 2 * SLICE_SIZE\n+NUM_PAGES = ARENA_SIZE // PAGE_SIZE\n+GUEST_READY = 0x4152454E414B564D\n+HOST_FIRST = 0x123456789ABCDEF0\n+GUEST_FIRST = 0xFEDCBA9876543210\n+HOST_SECOND = 0x1020304050607080\n+GUEST_SECOND = 0x8070605040302010\n+\n+\n+def progress(message):\n+ print(f\"# arena_kvm: {message}\", flush=True)\n+\n+\n+class TestSkipped(Exception):\n+ pass\n+\n+\n+class TestTerminated(BaseException):\n+ def __init__(self, signum):\n+ self.signum = signum\n+\n+\n+def terminate(signum, frame):\n+ # A second signal must not interrupt cleanup of pins and descendants.\n+ signal.signal(signal.SIGTERM, signal.SIG_IGN)\n+ signal.signal(signal.SIGINT, signal.SIG_IGN)\n+ raise TestTerminated(signum)\n+\n+\n+def enable_subreaper():\n+ # Adopt orphaned descendants even when script/vng changes session.\n+ libc = ctypes.CDLL(None, use_errno=True)\n+ libc.prctl.argtypes = (ctypes.c_int, *([ctypes.c_ulong] * 4))\n+ if libc.prctl(36, 1, 0, 0, 0): # PR_SET_CHILD_SUBREAPER\n+ err = ctypes.get_errno()\n+ raise OSError(err, os.strerror(err))\n+\n+\n+def children(pid):\n+ # /proc/PID/task/TID/children requires CONFIG_CHECKPOINT_RESTORE.\n+ # PPid is always available and covers children of every parent thread.\n+ result = set()\n+ for path in Path(\"/proc\").iterdir():\n+ if not path.name.isdigit():\n+ continue\n+ try:\n+ status = (path / \"status\").read_text()\n+ except (FileNotFoundError, ProcessLookupError, PermissionError):\n+ continue\n+ for line in status.splitlines():\n+ if line.startswith(\"PPid:\") and int(line.split()[1]) == pid:\n+ result.add(int(path.name))\n+ break\n+ return result\n+\n+\n+def stop_children(processes=()):\n+ # Kill direct children, then repeat for descendants adopted by the\n+ # subreaper. Sessions and process groups do not affect adoption.\n+ # Reaping until ECHILD verifies that the entire owned tree has exited.\n+ known = {process.pid: process for process in processes}\n+ deadline = time.monotonic() + 10\n+ while True:\n+ for pid in children(os.getpid()):\n+ try:\n+ fd = os.pidfd_open(pid)\n+ except ProcessLookupError:\n+ continue\n+ try:\n+ signal.pidfd_send_signal(fd, signal.SIGKILL)\n+ except ProcessLookupError:\n+ pass\n+ finally:\n+ os.close(fd)\n+ while True:\n+ try:\n+ pid, status = os.waitpid(-1, os.WNOHANG)\n+ except ChildProcessError:\n+ return\n+ if not pid:\n+ break\n+ if pid in known:\n+ known[pid].returncode = os.waitstatus_to_exitcode(status)\n+ if time.monotonic() \u003e= deadline:\n+ raise TimeoutError(\"guest descendants did not exit\")\n+ time.sleep(0.01)\n+\n+\n+def getq(arena, offset):\n+ return struct.unpack_from(\"=Q\", arena, offset)[0]\n+\n+\n+def wait_reply(runner, obj, fd, offset, guest, lines):\n+ deadline = time.monotonic() + 20\n+ while time.monotonic() \u003c deadline:\n+ result = run_host_bpf(runner, obj, fd, offset, \"check_reply\",\n+ timeout=max(0.001, deadline - time.monotonic()))\n+ if result == 0:\n+ return\n+ if result != 1:\n+ raise RuntimeError(f\"guest BPF reply is wrong: result={result}\")\n+ if guest.poll() is not None:\n+ raise RuntimeError(\"guest exited early:\\n\" + \"\".join(lines))\n+ time.sleep(0.001)\n+ raise TimeoutError(f\"timed out waiting for BPF reply at offset {offset}\")\n+\n+\n+def read_lines(guest, messages, lines):\n+ for line in guest.stdout:\n+ lines.append(line)\n+ # Serial startup may prefix a message with NULs echoed as ^@.\n+ messages.put(re.sub(r\"^(?:\\x00|\\^@)+\", \"\", line.strip()))\n+\n+\n+def wait_message(messages, lines, expected, guest):\n+ deadline = time.monotonic() + 45\n+ while time.monotonic() \u003c deadline:\n+ try:\n+ line = messages.get(timeout=0.5)\n+ except queue.Empty:\n+ if guest.poll() is not None:\n+ break\n+ continue\n+ if expected in line:\n+ return line\n+ if \"GUEST_BPF_FAILED\" in line:\n+ break\n+ if \"GUEST_BPF_SKIPPED\" in line:\n+ raise TestSkipped(line)\n+ raise RuntimeError(f\"guest did not print {expected} \"\n+ f\"(launcher exit status: {guest.poll()}):\\n\" +\n+ \"\".join(lines))\n+\n+\n+def run_host_bpf(runner, obj, fd, offset, program=\"exchange\",\n+ timeout=HELPER_TIMEOUT):\n+ output = subprocess.check_output(\n+ [str(runner), str(fd), str(obj), str(offset), program],\n+ pass_fds=(fd,), text=True, timeout=timeout)\n+ return int(output)\n+\n+\n+def resolve_vng():\n+ executable = shutil.which(\"vng\")\n+ if not executable or not os.access(executable, os.X_OK):\n+ print(\"SKIP: virtme-ng (vng) is unavailable\")\n+ return None, None\n+ return executable, os.environ.copy()\n+\n+\n+def start_guest(pin, build, kernel, index, vng, env, guests):\n+ node0_size = GUEST_MEMORY - SLICE_SIZE\n+ base = GUEST_RAM_BASES[os.uname().machine] + node0_size\n+ progress(f\"Starting guest {index}: {SLICE_SIZE // (1024 * 1024)} MiB \"\n+ f\"slice at file offset {index * SLICE_SIZE:#x}, \"\n+ f\"NUMA node 1 physical base {base:#x}\")\n+ qemu_opts = (\n+ f\"-object memory-backend-file,id=signal,size={SLICE_SIZE},share=on,\"\n+ f\"offset={index * SLICE_SIZE},mem-path={pin} \"\n+ \"-numa node,nodeid=1,memdev=signal\"\n+ )\n+ guest_command = [str(build / \"arena_kvm_guest-init\"),\n+ str(build / \"arena_kvm_guest.bpf.o\")]\n+ command = [\n+ vng, \"--verbose\", \"--run\", str(kernel), \"--cpus\", \"2\",\n+ \"--memory\", f\"{GUEST_MEMORY // (1024 * 1024)}M\",\n+ \"--numa\", str(node0_size),\n+ \"--exec\",\n+ shlex.join(guest_command),\n+ \"--append\", f\"arena_kvm.base={base:#x} arena_kvm.size={SLICE_SIZE}\",\n+ f\"--qemu-opts={qemu_opts}\",\n+ ]\n+ lines = []\n+ messages = queue.Queue()\n+ guest = subprocess.Popen([\"script\", \"-e\", \"-q\", \"-c\", shlex.join(command),\n+ \"/dev/null\"], stdout=subprocess.PIPE,\n+ stderr=subprocess.STDOUT, text=True, env=env,\n+ start_new_session=True)\n+ # Register ownership before thread construction or startup can fail.\n+ guests.append((guest, messages, lines, None))\n+ reader = threading.Thread(target=read_lines, args=(guest, messages, lines),\n+ daemon=True)\n+ guests[-1] = guest, messages, lines, reader\n+ reader.start()\n+\n+\n+def stop_guests(guests):\n+ stop_children(process for process, _, _, _ in guests)\n+ for process, _, _, reader in guests:\n+ process.wait(timeout=10)\n+ if reader is not None and reader.ident is not None:\n+ reader.join(timeout=2)\n+ if reader.is_alive():\n+ raise RuntimeError(\"guest output reader did not exit\")\n+ process.stdout.close()\n+\n+\n+def exchange(pin, fd, arena, build, kernel, vng, env):\n+ guests = []\n+ try:\n+ for index in range(2):\n+ start_guest(pin, build, kernel, index, vng, env, guests)\n+ offsets = []\n+ for index, (process, messages, lines, _) in enumerate(guests):\n+ progress(f\"Waiting for guest {index} to allocate its BPF arena page \"\n+ \"and check that BPF cannot free it\")\n+ ready = wait_message(messages, lines, \"GUEST_BPF_READY\", process)\n+ match = re.fullmatch(r\"GUEST_BPF_READY (\\d+) (\\d+)\", ready)\n+ if not match:\n+ raise RuntimeError(f\"malformed guest page offset: {ready}\")\n+ relative = int(match.group(1))\n+ if int(match.group(2)) != PAGE_SIZE:\n+ raise RuntimeError(\"host and guest page sizes differ\")\n+ offset = index * SLICE_SIZE + relative\n+ if relative \u003e= SLICE_SIZE or relative % PAGE_SIZE or \\\n+ getq(arena, offset) != GUEST_READY:\n+ raise RuntimeError(f\"invalid guest {index} page offset {relative}\")\n+ offsets.append(offset)\n+ progress(f\"Guest {index} ready: page offset {relative:#x} in slice, \"\n+ f\"{offset:#x} in host arena\")\n+ # The QEMU VMA and the open map FD retain the arena after unlink.\n+ progress(\"Unpinning the host arena while both guests retain mappings\")\n+ os.unlink(pin)\n+ runner = build / \"arena_kvm_host-runner\"\n+ obj = build / \"arena_kvm_host.bpf.o\"\n+ for offset in offsets:\n+ if run_host_bpf(runner, obj, fd, offset, \"probe_free\") != 0 or \\\n+ getq(arena, offset) != GUEST_READY:\n+ raise RuntimeError(\"host BPF freed a shared RAM page\")\n+ progress(\"Host BPF page-free checks passed for both shared pages\")\n+\n+ for index, offset in enumerate(offsets):\n+ progress(f\"Round 1: host -\u003e guest {index}, value {HOST_FIRST:#018x}\")\n+ result = run_host_bpf(runner, obj, fd, offset)\n+ if result != 0:\n+ raise RuntimeError(f\"guest {index} first result={result}\")\n+ other = offsets[1 - index]\n+ if index == 0 and getq(arena, other + 16) != 0:\n+ raise RuntimeError(\"guest 0 write modified guest 1 signal\")\n+ if index == 0:\n+ progress(\"Guest 1 signal unchanged after host write to guest 0\")\n+\n+ for index, (process, messages, lines, _) in enumerate(guests):\n+ offset = offsets[index]\n+ wait_reply(runner, obj, fd, offset, process, lines)\n+ progress(f\"Round 1: guest {index} -\u003e host, \"\n+ f\"verified value {GUEST_FIRST:#018x}\")\n+\n+ # Guest 0 exits while guest 1 keeps its slice mapped and active.\n+ for index, (process, messages, lines, _) in enumerate(guests):\n+ offset = offsets[index]\n+ progress(f\"Round 2: host -\u003e guest {index}, value {HOST_SECOND:#018x}\")\n+ if run_host_bpf(runner, obj, fd, offset) != 0:\n+ raise RuntimeError(f\"host BPF failed guest {index} second value\")\n+ wait_reply(runner, obj, fd, offset, process, lines)\n+ progress(f\"Round 2: guest {index} -\u003e host, \"\n+ f\"verified value {GUEST_SECOND:#018x}\")\n+ if run_host_bpf(runner, obj, fd, offset) != 2:\n+ raise RuntimeError(f\"host BPF failed guest {index} reply\")\n+ wait_message(messages, lines, \"GUEST_BPF_EXCHANGED\", process)\n+ process.wait(timeout=15)\n+ if process.returncode:\n+ raise RuntimeError(f\"guest {index} failed:\\n\" + \"\".join(lines))\n+ progress(f\"Guest {index} completed and exited successfully\")\n+ if index == 0:\n+ progress(\"Continuing with guest 1 after guest 0 exits\")\n+ print(f\"OK: two guests exchanged BPF values through arena slices {offsets}\")\n+ finally:\n+ progress(\"Cleaning up guest processes\")\n+ stop_guests(guests)\n+\n+\n+def skip(reason):\n+ print(f\"SKIP: {reason}\")\n+ return KSFT_SKIP\n+\n+\n+def create_arena(pin, build):\n+ parent, child = socket.socketpair()\n+ with parent, child:\n+ result = subprocess.run(\n+ [str(build / \"arena_kvm_host-runner\"), \"--create\", pin,\n+ str(child.fileno()), str(NUM_PAGES)], pass_fds=(child.fileno(),),\n+ timeout=HELPER_TIMEOUT)\n+ child.close()\n+ if result.returncode == KSFT_SKIP:\n+ return None\n+ result.check_returncode()\n+ fds = array.array(\"i\")\n+ _, ancillary, flags, _ = parent.recvmsg(\n+ 1, socket.CMSG_SPACE(fds.itemsize))\n+ for level, kind, data in ancillary:\n+ if level == socket.SOL_SOCKET and kind == socket.SCM_RIGHTS:\n+ fds.frombytes(data[:len(data) - len(data) % fds.itemsize])\n+ if flags \u0026 socket.MSG_CTRUNC or len(fds) != 1:\n+ for fd in fds:\n+ os.close(fd)\n+ raise RuntimeError(\"host helper did not send one arena FD\")\n+ return fds[0]\n+\n+\n+def main():\n+ parser = argparse.ArgumentParser(description=__doc__)\n+ parser.add_argument(\"--build-dir\", type=Path,\n+ help=\"directory containing the built helpers\")\n+ parser.add_argument(\"--kernel\", default=os.environ.get(\"ARENA_KVM_KERNEL\"),\n+ help=\"virtme-ng kernel path or release (default: source \"\n+ \"tree kernel, otherwise running kernel)\")\n+ args = parser.parse_args()\n+ if not hasattr(os, \"pidfd_open\") or not hasattr(signal, \"pidfd_send_signal\"):\n+ return skip(\"Python 3.9+ with Linux pidfd support is required\")\n+ machine = os.uname().machine\n+ if machine not in GUEST_RAM_BASES:\n+ return skip(f\"unknown guest RAM layout for architecture {machine}\")\n+ if struct.calcsize(\"P\") != 8:\n+ return skip(\"BPF arenas require a supported 64-bit architecture\")\n+ if os.geteuid() != 0:\n+ return skip(\"this test requires root\")\n+ if SLICE_SIZE % PAGE_SIZE:\n+ return skip(\"arena slices must contain a whole number of pages\")\n+ if not os.path.ismount(\"/sys/fs/bpf\"):\n+ return skip(\"mount bpffs first\")\n+ directory = Path(__file__).resolve().parent\n+ build = args.build_dir\n+ if build is None:\n+ build = Path.cwd() if (Path.cwd() / \"arena_kvm_host-runner\").is_file() \\\n+ else directory\n+ build = build.resolve()\n+ kernel = args.kernel\n+ if kernel is None:\n+ kernel = next((str(p) for p in directory.parents\n+ if (p / \"vmlinux\").is_file()), os.uname().release)\n+ vng, env = resolve_vng()\n+ if not vng:\n+ return KSFT_SKIP\n+ qemu_arch = {\"arm64\": \"aarch64\", \"ppc64le\": \"ppc64\"}.get(machine, machine)\n+ for tool in (\"script\", f\"qemu-system-{qemu_arch}\"):\n+ if not shutil.which(tool, path=env.get(\"PATH\")):\n+ return skip(f\"{tool} is unavailable\")\n+ for name in (\"arena_kvm_guest-init\", \"arena_kvm_guest.bpf.o\",\n+ \"arena_kvm_host-runner\", \"arena_kvm_host.bpf.o\"):\n+ if not (build / name).exists():\n+ return skip(f\"missing {build / name}; build the BPF selftests\")\n+ for name in (\"arena_kvm_host.bpf.o\", \"arena_kvm_guest.bpf.o\"):\n+ result = subprocess.run([str(build / \"arena_kvm_host-runner\"),\n+ \"--check-bpf\", str(build / name)],\n+ timeout=HELPER_TIMEOUT)\n+ if result.returncode == KSFT_SKIP:\n+ return skip(\"BPF compiler lacks arena address-space conversions\")\n+ result.check_returncode()\n+ if subprocess.run([str(build / \"arena_kvm_host-runner\"),\n+ \"--check-kvm\"], timeout=HELPER_TIMEOUT).returncode:\n+ return skip(\"KVM is unavailable or its API version is unsupported\")\n+ pin = f\"/sys/fs/bpf/arena_kvm_{os.getpid()}\"\n+ progress(f\"Architecture: {machine}; \"\n+ f\"page size: {PAGE_SIZE} bytes\")\n+ progress(f\"Creating and pinning a {ARENA_SIZE // (1024 * 1024)} MiB \"\n+ f\"host arena at {pin}\")\n+ try:\n+ fd = create_arena(pin, build)\n+ if fd is None:\n+ return skip(\"pinned BPF arenas are unsupported\")\n+ try:\n+ if os.stat(pin).st_size != ARENA_SIZE:\n+ raise RuntimeError(\"pinned arena has the wrong size\")\n+ with mmap.mmap(fd, ARENA_SIZE, flags=mmap.MAP_SHARED,\n+ prot=mmap.PROT_READ | mmap.PROT_WRITE) as arena:\n+ progress(f\"Populating {NUM_PAGES} host arena pages before \"\n+ \"starting the guests\")\n+ # Populate the file pages before QEMU uses them as RAM.\n+ for i in range(NUM_PAGES):\n+ arena[i * PAGE_SIZE]\n+ exchange(pin, fd, arena, build, kernel, vng, env)\n+ finally:\n+ os.close(fd)\n+ finally:\n+ if os.path.exists(pin):\n+ os.unlink(pin)\n+\n+\n+if __name__ == \"__main__\":\n+ enable_subreaper()\n+ signal.signal(signal.SIGTERM, terminate)\n+ signal.signal(signal.SIGINT, terminate)\n+ try:\n+ status = main()\n+ except TestSkipped as err:\n+ status = skip(str(err))\n+ except TestTerminated as err:\n+ progress(f\"Terminated by signal {err.signum}\")\n+ status = 128 + err.signum\n+ finally:\n+ signal.signal(signal.SIGTERM, signal.SIG_IGN)\n+ signal.signal(signal.SIGINT, signal.SIG_IGN)\n+ stop_children()\n+ sys.exit(status)\ndiff --git a/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c b/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c\nnew file mode 100644\nindex 0000000000000..461562149c57f\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c\n@@ -0,0 +1,89 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/*\n+ * Copyright (c) 2026 NVIDIA CORPORATION \u0026 AFFILIATES\n+ */\n+#include \u003cvmlinux.h\u003e\n+#include \u003cbpf/bpf_helpers.h\u003e\n+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST\n+#include \"bpf_arena_common.h\"\n+#ifdef __TARGET_ARCH_powerpc\n+/* PowerPC does not support arena load-acquire/store-release instructions. */\n+#undef __BPF_FEATURE_LOAD_ACQ_STORE_REL\n+#endif\n+#include \"bpf_atomic.h\"\n+#endif\n+#include \"arena_kvm_shared.h\"\n+\n+struct {\n+\t__uint(type, BPF_MAP_TYPE_ARENA);\n+\t__uint(map_flags, BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE);\n+\t__uint(max_entries, 1);\n+} arena SEC(\".maps\");\n+\n+volatile __u32 signal_offset;\n+\n+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST\n+\n+SEC(\"syscall\")\n+int allocate(void *ctx)\n+{\n+\tstruct signal_page __arena *page;\n+\n+\tpage = bpf_arena_alloc_pages(\u0026arena, NULL, 1, ARENA_KVM_NODE, 0);\n+\tif (!page)\n+\t\treturn 1;\n+\tsignal_offset = (__u32)((__u64)page - (__u64)arena_base(\u0026arena));\n+\tsmp_store_release(\u0026page-\u003eready, GUEST_READY);\n+\treturn 0;\n+}\n+\n+SEC(\"syscall\")\n+int probe_free(void *ctx)\n+{\n+\tstruct signal_page __arena *page =\n+\t\t(struct signal_page __arena *)((char __arena *)arena_base(\u0026arena) +\n+\t\t\t\t\t signal_offset);\n+\n+\tbpf_arena_free_pages(\u0026arena, page, 1);\n+\treturn page-\u003eready == GUEST_READY ? 0 : 1;\n+}\n+\n+SEC(\"syscall\")\n+int exchange(void *ctx)\n+{\n+\tstruct signal_page __arena *page =\n+\t\t(struct signal_page __arena *)((char __arena *)arena_base(\u0026arena) +\n+\t\t\t\t\t signal_offset);\n+\t__u64 seq = smp_load_acquire(\u0026page-\u003eh2g_seq);\n+\t__u64 payload;\n+\n+\tif ((seq != 1 \u0026\u0026 seq != 2) || page-\u003eg2h_seq == seq)\n+\t\treturn 1;\n+\tpayload = page-\u003eh2g_payload;\n+\tif (seq == 1 \u0026\u0026 payload != HOST_FIRST)\n+\t\treturn 2;\n+\tif (seq == 2 \u0026\u0026 payload != HOST_SECOND)\n+\t\treturn 3;\n+\tpage-\u003eg2h_payload = seq == 1 ? GUEST_FIRST : GUEST_SECOND;\n+\tsmp_store_release(\u0026page-\u003eg2h_seq, seq);\n+\treturn 0;\n+}\n+\n+SEC(\"syscall\")\n+int complete(void *ctx)\n+{\n+\tstruct signal_page __arena *page =\n+\t\t(struct signal_page __arena *)((char __arena *)arena_base(\u0026arena) +\n+\t\t\t\t\t signal_offset);\n+\n+\treturn smp_load_acquire(\u0026page-\u003eh2g_seq) == 3 ? 0 : 1;\n+}\n+\n+#else\n+SEC(\"syscall\") int allocate(void *ctx) { return 1; }\n+SEC(\"syscall\") int probe_free(void *ctx) { return 1; }\n+SEC(\"syscall\") int exchange(void *ctx) { return 1; }\n+SEC(\"syscall\") int complete(void *ctx) { return 1; }\n+#endif\n+\n+char _license[] SEC(\"license\") = \"GPL\";\ndiff --git a/tools/testing/selftests/bpf/arena_kvm_guest.c b/tools/testing/selftests/bpf/arena_kvm_guest.c\nnew file mode 100644\nindex 0000000000000..da6586ae660d1\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/arena_kvm_guest.c\n@@ -0,0 +1,224 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/*\n+ * Copyright (c) 2026 NVIDIA CORPORATION \u0026 AFFILIATES\n+ *\n+ * Guest helper for a BPF arena allocated from shared NUMA RAM.\n+ */\n+#define __EXPORTED_HEADERS__\n+\n+#include \u003cbpf/bpf.h\u003e\n+#include \u003cbpf/libbpf.h\u003e\n+#include \u003cerrno.h\u003e\n+#include \u003cfcntl.h\u003e\n+#include \u003clinux/bpf.h\u003e\n+#include \u003clinux/mempolicy.h\u003e\n+#include \u003cstdint.h\u003e\n+#include \u003cstdio.h\u003e\n+#include \u003cstdlib.h\u003e\n+#include \u003cstring.h\u003e\n+#include \u003csys/mman.h\u003e\n+#include \u003csys/mount.h\u003e\n+#include \u003csys/syscall.h\u003e\n+#include \u003cunistd.h\u003e\n+\n+#include \"arena_kvm_shared.h\"\n+\n+static int signal_file_offset(void *arena, uint64_t *offset)\n+{\n+\tchar cmdline[8192], *arg, *end;\n+\tuint64_t entry, pfn, gpa, base = 0, size = 0, *value;\n+\tssize_t n;\n+\tint fd, node, found = 0;\n+\n+\t/* The launcher supplies QEMU's guest physical base for this slice. */\n+\tfd = open(\"/proc/cmdline\", O_RDONLY);\n+\tif (fd \u003c 0)\n+\t\treturn -errno;\n+\tn = read(fd, cmdline, sizeof(cmdline) - 1);\n+\tclose(fd);\n+\tif (n \u003c= 0 || n == sizeof(cmdline) - 1)\n+\t\treturn -EIO;\n+\tcmdline[n] = '\\0';\n+\tfor (arg = strtok(cmdline, \"\\n \"); arg; arg = strtok(NULL, \"\\n \")) {\n+\t\tif (!strncmp(arg, \"arena_kvm.base=\", 15)) {\n+\t\t\tvalue = \u0026base;\n+\t\t\tfound |= 1;\n+\t\t} else if (!strncmp(arg, \"arena_kvm.size=\", 15)) {\n+\t\t\tvalue = \u0026size;\n+\t\t\tfound |= 2;\n+\t\t} else {\n+\t\t\tcontinue;\n+\t\t}\n+\t\terrno = 0;\n+\t\t*value = strtoull(arg + 15, \u0026end, 0);\n+\t\tif (errno || *end || end == arg + 15 || *value % getpagesize())\n+\t\t\treturn -EINVAL;\n+\t}\n+\tif (found != 3 || size \u003c getpagesize())\n+\t\treturn -EINVAL;\n+\n+\t/* Fault the BPF-allocated page into the guest user VMA. */\n+\tif (!*(volatile uint64_t *)arena)\n+\t\treturn -EINVAL;\n+\tif (syscall(SYS_get_mempolicy, \u0026node, NULL, 0, arena,\n+\t\t MPOL_F_NODE | MPOL_F_ADDR))\n+\t\treturn -errno;\n+\tif (node != ARENA_KVM_NODE)\n+\t\treturn -EXDEV;\n+\tfd = open(\"/proc/self/pagemap\", O_RDONLY);\n+\tif (fd \u003c 0)\n+\t\treturn -errno;\n+\tn = pread(fd, \u0026entry, sizeof(entry),\n+\t\t (uintptr_t)arena / getpagesize() * sizeof(entry));\n+\tclose(fd);\n+\tif (n != sizeof(entry) || !(entry \u0026 (1ULL \u003c\u003c 63)))\n+\t\treturn -EIO;\n+\tpfn = entry \u0026 ((1ULL \u003c\u003c 55) - 1);\n+\tif (!pfn)\n+\t\treturn -EPERM;\n+\tgpa = pfn * getpagesize();\n+\tif (gpa \u003c base || gpa - base \u003e= size || size - (gpa - base) \u003c getpagesize())\n+\t\treturn -ERANGE;\n+\t*offset = gpa - base;\n+\treturn 0;\n+}\n+\n+static int create_arena(void)\n+{\n+\tunion bpf_attr attr = {\n+\t\t.map_type = BPF_MAP_TYPE_ARENA,\n+\t\t.max_entries = 1,\n+\t\t.map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE,\n+\t};\n+\n+\treturn syscall(__NR_bpf, BPF_MAP_CREATE, \u0026attr, sizeof(attr));\n+}\n+\n+static int run_prog(int fd, unsigned int *retval)\n+{\n+\tLIBBPF_OPTS(bpf_test_run_opts, opts);\n+\tint err = bpf_prog_test_run_opts(fd, \u0026opts);\n+\n+\tif (err)\n+\t\treturn err;\n+\t*retval = opts.retval;\n+\treturn 0;\n+}\n+\n+static int exchange_once(int fd)\n+{\n+\tunsigned int result = 0;\n+\tint err;\n+\n+\t/* The host bounds each protocol stage and stops the guest on timeout. */\n+\tfor (;;) {\n+\t\terr = run_prog(fd, \u0026result);\n+\t\tif (err || result \u003e 1)\n+\t\t\treturn err ? err : -EINVAL;\n+\t\tif (!result)\n+\t\t\treturn 0;\n+\t\tusleep(1000);\n+\t}\n+}\n+\n+int main(int argc, char **argv)\n+{\n+\tstruct bpf_object *obj = NULL;\n+\tstruct bpf_program *alloc, *probe_free, *exchange, *complete;\n+\tstruct bpf_map *map;\n+\tvoid *arena = MAP_FAILED;\n+\tuint64_t file_offset;\n+\tunsigned int result = 0;\n+\tint fd = -1;\n+\tint err;\n+\n+\tsetbuf(stdout, NULL);\n+\tif (argc != 2 ||\n+\t (mount(\"sysfs\", \"/sys\", \"sysfs\", 0, NULL) \u0026\u0026 errno != EBUSY) ||\n+\t (mount(\"proc\", \"/proc\", \"proc\", 0, NULL) \u0026\u0026 errno != EBUSY))\n+\t\tgoto fail;\n+\tif (access(\"/sys/devices/system/node/node1\", F_OK)) {\n+\t\tfputs(\"guest NUMA node 1 is unavailable\\n\", stderr);\n+\t\tgoto fail;\n+\t}\n+\tfd = create_arena();\n+\tif (fd \u003c 0) {\n+\t\tperror(\"create guest arena\");\n+\t\tgoto fail;\n+\t}\n+\tarena = mmap(NULL, getpagesize(), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);\n+\tif (arena == MAP_FAILED) {\n+\t\tperror(\"mmap guest arena\");\n+\t\tgoto fail;\n+\t}\n+\tobj = bpf_object__open_file(argv[1], NULL);\n+\tif (!obj)\n+\t\tgoto fail;\n+\terr = arena_kvm_check_features(obj);\n+\tif (err == 4) {\n+\t\tputs(\"GUEST_BPF_SKIPPED: compiler lacks arena address-space casts\");\n+\t\tbpf_object__close(obj);\n+\t\tmunmap(arena, getpagesize());\n+\t\tclose(fd);\n+\t\treturn 4;\n+\t}\n+\tif (err)\n+\t\tgoto fail;\n+\tmap = bpf_object__find_map_by_name(obj, \"arena\");\n+\talloc = bpf_object__find_program_by_name(obj, \"allocate\");\n+\tprobe_free = bpf_object__find_program_by_name(obj, \"probe_free\");\n+\texchange = bpf_object__find_program_by_name(obj, \"exchange\");\n+\tcomplete = bpf_object__find_program_by_name(obj, \"complete\");\n+\tif (!map || !alloc || !probe_free || !exchange || !complete ||\n+\t bpf_map__reuse_fd(map, fd))\n+\t\tgoto fail;\n+\tclose(fd);\n+\tfd = -1;\n+\terr = bpf_object__load(obj);\n+\tif (err) {\n+\t\tfprintf(stderr, \"load guest BPF: %d\\n\", err);\n+\t\tgoto fail;\n+\t}\n+\terr = run_prog(bpf_program__fd(alloc), \u0026result);\n+\tif (err || result) {\n+\t\tfprintf(stderr, \"allocate guest page: err=%d result=%u\\n\",\n+\t\t\terr, result);\n+\t\tgoto fail;\n+\t}\n+\terr = run_prog(bpf_program__fd(probe_free), \u0026result);\n+\tif (err || result) {\n+\t\tfprintf(stderr, \"guest page lifetime check: err=%d result=%u\\n\",\n+\t\t\terr, result);\n+\t\tgoto fail;\n+\t}\n+\terr = signal_file_offset(arena, \u0026file_offset);\n+\tif (err == -EXDEV) {\n+\t\tputs(\"GUEST_BPF_SKIPPED: allocation fell back from shared NUMA node\");\n+\t\tbpf_object__close(obj);\n+\t\tmunmap(arena, getpagesize());\n+\t\treturn 4;\n+\t}\n+\tif (err) {\n+\t\tfprintf(stderr, \"resolve shared page offset: %d\\n\", err);\n+\t\tgoto fail;\n+\t}\n+\tprintf(\"GUEST_BPF_READY %llu %d\\n\",\n+\t (unsigned long long)file_offset, getpagesize());\n+\tif (exchange_once(bpf_program__fd(exchange)) ||\n+\t exchange_once(bpf_program__fd(exchange)) ||\n+\t exchange_once(bpf_program__fd(complete)))\n+\t\tgoto fail;\n+\tputs(\"GUEST_BPF_EXCHANGED\");\n+\tbpf_object__close(obj);\n+\tmunmap(arena, getpagesize());\n+\treturn 0;\n+\n+fail:\n+\tputs(\"GUEST_BPF_FAILED\");\n+\tbpf_object__close(obj);\n+\tif (arena != MAP_FAILED)\n+\t\tmunmap(arena, getpagesize());\n+\tif (fd \u003e= 0)\n+\t\tclose(fd);\n+\treturn 1;\n+}\ndiff --git a/tools/testing/selftests/bpf/arena_kvm_host.bpf.c b/tools/testing/selftests/bpf/arena_kvm_host.bpf.c\nnew file mode 100644\nindex 0000000000000..9c1c535fe429b\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/arena_kvm_host.bpf.c\n@@ -0,0 +1,89 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/*\n+ * Copyright (c) 2026 NVIDIA CORPORATION \u0026 AFFILIATES\n+ */\n+#include \u003cvmlinux.h\u003e\n+#include \u003cbpf/bpf_helpers.h\u003e\n+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST\n+#include \"bpf_arena_common.h\"\n+#ifdef __TARGET_ARCH_powerpc\n+/* PowerPC does not support arena load-acquire/store-release instructions. */\n+#undef __BPF_FEATURE_LOAD_ACQ_STORE_REL\n+#endif\n+#include \"bpf_atomic.h\"\n+#endif\n+#include \"arena_kvm_shared.h\"\n+\n+struct {\n+\t__uint(type, BPF_MAP_TYPE_ARENA);\n+\t__uint(map_flags, BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE | BPF_F_ARENA_EXPORT);\n+\t/* The host runner reuses the map created with the selected capacity. */\n+\t__uint(max_entries, 1);\n+} arena SEC(\".maps\");\n+\n+volatile __u32 signal_offset;\n+\n+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST\n+\n+SEC(\"syscall\")\n+int exchange(void *ctx)\n+{\n+\tstruct signal_page __arena *page =\n+\t\t(struct signal_page __arena *)((char __arena *)arena_base(\u0026arena) +\n+\t\t\t\t\t signal_offset);\n+\t__u64 seq = page-\u003eh2g_seq;\n+\n+\tif (smp_load_acquire(\u0026page-\u003eready) != GUEST_READY)\n+\t\treturn 4;\n+\tif (!seq) {\n+\t\tpage-\u003eh2g_payload = HOST_FIRST;\n+\t\tsmp_store_release(\u0026page-\u003eh2g_seq, 1);\n+\t\treturn 0;\n+\t}\n+\tif (seq == 1 \u0026\u0026 smp_load_acquire(\u0026page-\u003eg2h_seq) == 1) {\n+\t\tif (page-\u003eg2h_payload != GUEST_FIRST)\n+\t\t\treturn 3;\n+\t\tpage-\u003eh2g_payload = HOST_SECOND;\n+\t\tsmp_store_release(\u0026page-\u003eh2g_seq, 2);\n+\t\treturn 0;\n+\t}\n+\tif (seq == 2 \u0026\u0026 smp_load_acquire(\u0026page-\u003eg2h_seq) == 2) {\n+\t\tif (page-\u003eg2h_payload != GUEST_SECOND)\n+\t\t\treturn 3;\n+\t\tsmp_store_release(\u0026page-\u003eh2g_seq, 3);\n+\t\treturn 2;\n+\t}\n+\treturn 1;\n+}\n+\n+SEC(\"syscall\")\n+int check_reply(void *ctx)\n+{\n+\tstruct signal_page __arena *page =\n+\t\t(struct signal_page __arena *)((char __arena *)arena_base(\u0026arena) +\n+\t\t\t\t\t signal_offset);\n+\t__u64 seq = page-\u003eh2g_seq;\n+\n+\tif ((seq != 1 \u0026\u0026 seq != 2) || smp_load_acquire(\u0026page-\u003eg2h_seq) != seq)\n+\t\treturn 1;\n+\treturn page-\u003eg2h_payload == (seq == 1 ? GUEST_FIRST : GUEST_SECOND) ? 0 : 3;\n+}\n+\n+SEC(\"syscall\")\n+int probe_free(void *ctx)\n+{\n+\tstruct signal_page __arena *page =\n+\t\t(struct signal_page __arena *)((char __arena *)arena_base(\u0026arena) +\n+\t\t\t\t\t signal_offset);\n+\n+\tbpf_arena_free_pages(\u0026arena, page, 1);\n+\treturn page-\u003eready == GUEST_READY ? 0 : 1;\n+}\n+\n+#else\n+SEC(\"syscall\") int exchange(void *ctx) { return 1; }\n+SEC(\"syscall\") int check_reply(void *ctx) { return 1; }\n+SEC(\"syscall\") int probe_free(void *ctx) { return 1; }\n+#endif\n+\n+char _license[] SEC(\"license\") = \"GPL\";\ndiff --git a/tools/testing/selftests/bpf/arena_kvm_host.c b/tools/testing/selftests/bpf/arena_kvm_host.c\nnew file mode 100644\nindex 0000000000000..90a77a54bd6f4\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/arena_kvm_host.c\n@@ -0,0 +1,147 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/*\n+ * Copyright (c) 2026 NVIDIA CORPORATION \u0026 AFFILIATES\n+ */\n+#include \u003cbpf/bpf.h\u003e\n+#include \u003cbpf/libbpf.h\u003e\n+#include \u003cerrno.h\u003e\n+#include \u003cfcntl.h\u003e\n+#include \u003clinux/kvm.h\u003e\n+#include \u003climits.h\u003e\n+#include \u003cstdint.h\u003e\n+#include \u003cstdio.h\u003e\n+#include \u003cstdlib.h\u003e\n+#include \u003cstring.h\u003e\n+#include \u003csys/ioctl.h\u003e\n+#include \u003csys/socket.h\u003e\n+#include \u003cunistd.h\u003e\n+\n+#include \"arena_kvm_shared.h\"\n+\n+static int create_arena(const char *pin, int socket_fd, unsigned int pages)\n+{\n+\tLIBBPF_OPTS(bpf_map_create_opts, opts,\n+\t\t.map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE |\n+\t\t\t BPF_F_ARENA_EXPORT);\n+\tchar control[CMSG_SPACE(sizeof(int))] = {};\n+\tchar byte = 0;\n+\tstruct iovec iov = { .iov_base = \u0026byte, .iov_len = 1 };\n+\tstruct msghdr msg = {\n+\t\t.msg_iov = \u0026iov,\n+\t\t.msg_iovlen = 1,\n+\t\t.msg_control = control,\n+\t\t.msg_controllen = sizeof(control),\n+\t};\n+\tstruct cmsghdr *cmsg = CMSG_FIRSTHDR(\u0026msg);\n+\tint fd, err = 1;\n+\n+\tfd = bpf_map_create(BPF_MAP_TYPE_ARENA, \"arena_kvm\", 0, 0,\n+\t\t\t pages, \u0026opts);\n+\tif (fd \u003c 0) {\n+\t\tif (errno == EOPNOTSUPP || errno == EINVAL)\n+\t\t\treturn 4; /* KSFT_SKIP: arena type or flags unsupported. */\n+\t\tperror(\"create host arena\");\n+\t\treturn 1;\n+\t}\n+\tif (bpf_obj_pin(fd, pin)) {\n+\t\tperror(\"pin host arena\");\n+\t\tgoto out;\n+\t}\n+\tcmsg-\u003ecmsg_level = SOL_SOCKET;\n+\tcmsg-\u003ecmsg_type = SCM_RIGHTS;\n+\tcmsg-\u003ecmsg_len = CMSG_LEN(sizeof(fd));\n+\tmemcpy(CMSG_DATA(cmsg), \u0026fd, sizeof(fd));\n+\tif (sendmsg(socket_fd, \u0026msg, 0) == 1)\n+\t\terr = 0;\n+\telse\n+\t\tperror(\"send arena FD\");\n+out:\n+\tclose(fd);\n+\treturn err;\n+}\n+\n+int main(int argc, char **argv)\n+{\n+\tLIBBPF_OPTS(bpf_test_run_opts, opts);\n+\tstruct bpf_object *obj = NULL;\n+\tstruct bpf_program *prog;\n+\tstruct bpf_map *arena, *bss;\n+\tuint32_t key = 0, offset;\n+\tunsigned long parsed;\n+\tint fd, err;\n+\tchar *end;\n+\n+\tif (argc == 3 \u0026\u0026 !strcmp(argv[1], \"--check-bpf\")) {\n+\t\tobj = bpf_object__open_file(argv[2], NULL);\n+\t\tif (!obj)\n+\t\t\treturn 1;\n+\t\terr = arena_kvm_check_features(obj);\n+\t\tbpf_object__close(obj);\n+\t\treturn err;\n+\t}\n+\n+\tif (argc == 2 \u0026\u0026 !strcmp(argv[1], \"--check-kvm\")) {\n+\t\tfd = open(\"/dev/kvm\", O_RDWR);\n+\t\tif (fd \u003c 0)\n+\t\t\treturn 4;\n+\t\terr = ioctl(fd, KVM_GET_API_VERSION, 0);\n+\t\tclose(fd);\n+\t\treturn err == KVM_API_VERSION ? 0 : 4;\n+\t}\n+\tif (argc == 5 \u0026\u0026 !strcmp(argv[1], \"--create\")) {\n+\t\terrno = 0;\n+\t\tparsed = strtoul(argv[3], \u0026end, 0);\n+\t\tif (errno || *end || end == argv[3] || parsed \u003e INT_MAX)\n+\t\t\treturn 1;\n+\t\tfd = parsed;\n+\t\terrno = 0;\n+\t\tparsed = strtoul(argv[4], \u0026end, 0);\n+\t\tif (errno || *end || end == argv[4] || !parsed || parsed \u003e UINT_MAX)\n+\t\t\treturn 1;\n+\t\treturn create_arena(argv[2], fd, parsed);\n+\t}\n+\tif (argc != 5)\n+\t\treturn 1;\n+\terrno = 0;\n+\tparsed = strtoul(argv[1], \u0026end, 0);\n+\tif (errno || *end || parsed \u003e INT_MAX)\n+\t\treturn 1;\n+\tfd = dup(parsed);\n+\tif (fd \u003c 0)\n+\t\treturn 1;\n+\terrno = 0;\n+\tparsed = strtoul(argv[3], \u0026end, 0);\n+\tif (errno || *end || parsed \u003e UINT_MAX || parsed % getpagesize())\n+\t\tgoto fail;\n+\toffset = parsed;\n+\tobj = bpf_object__open_file(argv[2], NULL);\n+\tif (!obj)\n+\t\tgoto fail;\n+\tarena = bpf_object__find_map_by_name(obj, \"arena\");\n+\tbss = bpf_object__find_map_by_name(obj, \".bss\");\n+\tprog = bpf_object__find_program_by_name(obj, argv[4]);\n+\tif (!arena || !bss || !prog || bpf_map__reuse_fd(arena, fd))\n+\t\tgoto fail;\n+\tif (offset \u003e= (uint64_t)bpf_map__max_entries(arena) * getpagesize())\n+\t\tgoto fail;\n+\tif (bpf_object__load(obj)) {\n+\t\tfputs(\"load host BPF failed\\n\", stderr);\n+\t\tgoto fail;\n+\t}\n+\tif (bpf_map_update_elem(bpf_map__fd(bss), \u0026key, \u0026offset, BPF_ANY)) {\n+\t\tperror(\"set signal offset\");\n+\t\tgoto fail;\n+\t}\n+\terr = bpf_prog_test_run_opts(bpf_program__fd(prog), \u0026opts);\n+\tif (err)\n+\t\tgoto fail;\n+\tprintf(\"%u\\n\", opts.retval);\n+\tbpf_object__close(obj);\n+\tclose(fd);\n+\treturn 0;\n+\n+fail:\n+\tbpf_object__close(obj);\n+\tclose(fd);\n+\treturn 1;\n+}\ndiff --git a/tools/testing/selftests/bpf/arena_kvm_shared.h b/tools/testing/selftests/bpf/arena_kvm_shared.h\nnew file mode 100644\nindex 0000000000000..63aafac25efc3\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/arena_kvm_shared.h\n@@ -0,0 +1,60 @@\n+/* SPDX-License-Identifier: GPL-2.0 */\n+/*\n+ * Copyright (c) 2026 NVIDIA CORPORATION \u0026 AFFILIATES\n+ */\n+#ifndef ARENA_KVM_SHARED_H\n+#define ARENA_KVM_SHARED_H\n+\n+/* These flags may be absent from a pre-series kernel's vmlinux.h. */\n+#ifndef BPF_F_ARENA_NO_FREE\n+#define BPF_F_ARENA_NO_FREE (1U \u003c\u003c 20)\n+#endif\n+#ifndef BPF_F_ARENA_EXPORT\n+#define BPF_F_ARENA_EXPORT (1U \u003c\u003c 21)\n+#endif\n+\n+struct arena_kvm_features {\n+\tunsigned int addr_space_cast;\n+};\n+\n+#ifdef __BPF__\n+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST\n+const volatile struct arena_kvm_features features = { .addr_space_cast = 1 };\n+#else\n+const volatile struct arena_kvm_features features = {};\n+#endif\n+#else\n+/* Inspect compiler support before loading maps or booting either guest. */\n+static inline int arena_kvm_check_features(struct bpf_object *obj)\n+{\n+\tconst struct arena_kvm_features *features;\n+\tstruct bpf_map *map;\n+\tsize_t size;\n+\n+\tmap = bpf_object__find_map_by_name(obj, \".rodata\");\n+\tif (!map)\n+\t\treturn 1;\n+\tfeatures = bpf_map__initial_value(map, \u0026size);\n+\tif (!features || size != sizeof(*features))\n+\t\treturn 1;\n+\treturn features-\u003eaddr_space_cast ? 0 : 4; /* KSFT_SKIP */\n+}\n+#endif\n+\n+#define ARENA_KVM_NODE 1\n+\n+#define GUEST_READY 0x4152454e414b564dULL\n+#define HOST_FIRST 0x123456789abcdef0ULL\n+#define GUEST_FIRST 0xfedcba9876543210ULL\n+#define HOST_SECOND 0x1020304050607080ULL\n+#define GUEST_SECOND 0x8070605040302010ULL\n+\n+struct signal_page {\n+\tvolatile unsigned long long ready;\n+\tvolatile unsigned long long h2g_payload;\n+\tvolatile unsigned long long h2g_seq;\n+\tvolatile unsigned long long g2h_payload;\n+\tvolatile unsigned long long g2h_seq;\n+};\n+\n+#endif\ndiff --git a/tools/testing/selftests/bpf/prog_tests/arena_pinned.c b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c\nnew file mode 100644\nindex 0000000000000..8b7ca90211296\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c\n@@ -0,0 +1,442 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/*\n+ * Copyright (c) 2026 NVIDIA CORPORATION \u0026 AFFILIATES\n+ */\n+#include \u003ctest_progs.h\u003e\n+#include \u003cfcntl.h\u003e\n+#include \u003csys/mman.h\u003e\n+#include \u003csys/stat.h\u003e\n+#include \u003csys/wait.h\u003e\n+\n+#define ARENA_PAGES 4\n+\n+struct arena_file {\n+\tint map_fd;\n+\tint file_fd;\n+\tchar pin[128];\n+\tvoid *canonical;\n+\tsize_t len;\n+\tbool pinned;\n+};\n+\n+static void cleanup(struct arena_file *file)\n+{\n+\tif (file-\u003ecanonical != MAP_FAILED)\n+\t\tmunmap(file-\u003ecanonical, file-\u003elen);\n+\tif (file-\u003efile_fd \u003e= 0)\n+\t\tclose(file-\u003efile_fd);\n+\tif (file-\u003epinned)\n+\t\tunlink(file-\u003epin);\n+\tif (file-\u003emap_fd \u003e= 0)\n+\t\tclose(file-\u003emap_fd);\n+}\n+\n+static void reject_mapping(int fd, size_t len, int flags, off_t offset,\n+\t\t\t const char *name)\n+{\n+\tvoid *addr;\n+\tint err;\n+\n+\taddr = mmap(NULL, len, PROT_READ | PROT_WRITE, flags, fd, offset);\n+\terr = errno;\n+\tif (ASSERT_EQ(addr, MAP_FAILED, name))\n+\t\tASSERT_EQ(err, EINVAL, \"mmap_errno\");\n+\telse\n+\t\tmunmap(addr, len);\n+}\n+\n+static bool setup(struct arena_file *file, __u64 map_extra, bool zero)\n+{\n+\tLIBBPF_OPTS(bpf_map_create_opts, opts,\n+\t\t.map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE |\n+\t\t\t BPF_F_ARENA_EXPORT,\n+\t\t.map_extra = map_extra);\n+\n+\tmemset(file, 0, sizeof(*file));\n+\tfile-\u003emap_fd = -1;\n+\tfile-\u003efile_fd = -1;\n+\tfile-\u003ecanonical = MAP_FAILED;\n+\tfile-\u003elen = ARENA_PAGES * getpagesize();\n+\tsnprintf(file-\u003epin, sizeof(file-\u003epin), \"/sys/fs/bpf/arena_pinned_%d\", getpid());\n+\tfile-\u003emap_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, \"arena_pinned\",\n+\t\t\t\t 0, 0, ARENA_PAGES, \u0026opts);\n+\tif (file-\u003emap_fd \u003c 0 \u0026\u0026 errno == EOPNOTSUPP) {\n+\t\ttest__skip();\n+\t\treturn false;\n+\t}\n+\tif (!ASSERT_GE(file-\u003emap_fd, 0, \"map_create\"))\n+\t\treturn false;\n+\tif (!ASSERT_OK(bpf_obj_pin(file-\u003emap_fd, file-\u003epin), \"obj_pin\"))\n+\t\treturn false;\n+\tfile-\u003epinned = true;\n+\tfile-\u003efile_fd = open(file-\u003epin, O_RDWR);\n+\tif (!ASSERT_GE(file-\u003efile_fd, 0, \"open_pin\"))\n+\t\treturn false;\n+\treject_mapping(file-\u003emap_fd, getpagesize(), MAP_SHARED, 0, \"short_canonical\");\n+\treject_mapping(file-\u003emap_fd, file-\u003elen + getpagesize(), MAP_SHARED,\n+\t\t 0, \"large_canonical\");\n+\tif (map_extra)\n+\t\treturn true;\n+\n+\treject_mapping(file-\u003efile_fd, getpagesize(), MAP_SHARED, 0, \"uninitialized\");\n+\tfile-\u003ecanonical = mmap(NULL, file-\u003elen, PROT_READ | PROT_WRITE,\n+\t\t\t MAP_SHARED | (zero ? MAP_FIXED_NOREPLACE : 0),\n+\t\t\t file-\u003emap_fd, 0);\n+\tif (zero \u0026\u0026 file-\u003ecanonical == MAP_FAILED \u0026\u0026\n+\t (errno == EPERM || errno == EACCES)) {\n+\t\ttest__skip();\n+\t\treturn false;\n+\t}\n+\treturn ASSERT_NEQ(file-\u003ecanonical, MAP_FAILED, \"canonical_mmap\");\n+}\n+\n+static void test_mappings(__u64 map_extra)\n+{\n+\tstruct arena_file file;\n+\tsize_t ps = getpagesize();\n+\tstruct stat st;\n+\tchar *alias = MAP_FAILED;\n+\tchar *full = MAP_FAILED;\n+\tint expected = 91;\n+\n+\tif (!setup(\u0026file, map_extra, false))\n+\t\tgoto out;\n+\tif (!ASSERT_OK(fstat(file.file_fd, \u0026st), \"stat_pin\"))\n+\t\tgoto out;\n+\tASSERT_EQ(st.st_size, ARENA_PAGES * ps, \"pin_size\");\n+\treject_mapping(file.file_fd, 8ULL \u003c\u003c 30, MAP_SHARED, 0, \"oversized\");\n+\treject_mapping(file.file_fd, ps, MAP_SHARED, ARENA_PAGES * ps, \"past_capacity\");\n+\treject_mapping(file.file_fd, ps * 2, MAP_SHARED,\n+\t\t (ARENA_PAGES - 1) * ps, \"past_capacity_end\");\n+\treject_mapping(file.file_fd, ps, MAP_PRIVATE, 0, \"private_mapping\");\n+\tfull = mmap(NULL, st.st_size, PROT_READ | PROT_WRITE, MAP_SHARED,\n+\t\t file.file_fd, 0);\n+\tif (!ASSERT_NEQ(full, MAP_FAILED, \"full_file_mmap\"))\n+\t\tgoto out;\n+\tfull[0] = 17;\n+\tfull[file.len - 1] = 62;\n+\n+\talias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED,\n+\t\t file.file_fd, (ARENA_PAGES - 1) * ps);\n+\tif (!ASSERT_NEQ(alias, MAP_FAILED, \"last_page_slice\"))\n+\t\tgoto out;\n+\talias[0] = 91;\n+\tASSERT_OK(msync(alias, ps, MS_SYNC), \"msync_slice\");\n+\tASSERT_OK(fsync(file.file_fd), \"fsync_pin\");\n+\tASSERT_OK(fdatasync(file.file_fd), \"fdatasync_pin\");\n+\tASSERT_EQ(alias[ps - 1], 62, \"full_file_last_page\");\n+\tif (file.canonical != MAP_FAILED) {\n+\t\tchar *canonical = file.canonical;\n+\n+\t\tASSERT_EQ(canonical[0], 17, \"full_file_first_page\");\n+\t\tASSERT_EQ(canonical[(ARENA_PAGES - 1) * ps], 91, \"alias_write\");\n+\t\tcanonical[(ARENA_PAGES - 1) * ps] = 42;\n+\t\tASSERT_EQ(alias[0], 42, \"canonical_write\");\n+\t\texpected = 42;\n+\t\tif (!ASSERT_OK(munmap(file.canonical, file.len), \"unmap_canonical\"))\n+\t\t\tgoto out;\n+\t\tfile.canonical = MAP_FAILED;\n+\t}\n+\tif (!ASSERT_OK(munmap(full, file.len), \"unmap_full_file\"))\n+\t\tgoto out;\n+\tfull = MAP_FAILED;\n+\tif (!ASSERT_OK(unlink(file.pin), \"unlink_pin\"))\n+\t\tgoto out;\n+\tfile.pinned = false;\n+\tclose(file.file_fd);\n+\tfile.file_fd = -1;\n+\tclose(file.map_fd);\n+\tfile.map_fd = -1;\n+\tASSERT_EQ(alias[0], expected, \"unpin_lifetime\");\n+\talias[0] = 73;\n+out:\n+\tif (full != MAP_FAILED)\n+\t\tmunmap(full, file.len);\n+\tif (alias != MAP_FAILED)\n+\t\tmunmap(alias, ps);\n+\tcleanup(\u0026file);\n+}\n+\n+static void test_unexported_pin(__u32 flags)\n+{\n+\tLIBBPF_OPTS(bpf_map_create_opts, opts,\n+\t\t.map_flags = BPF_F_MMAPABLE | flags);\n+\tchar pin[128];\n+\tstruct stat st;\n+\tchar *canonical = MAP_FAILED;\n+\tsize_t ps = getpagesize();\n+\tint map_fd, fd, err;\n+\n+\tmap_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, \"arena_unexported\",\n+\t\t\t 0, 0, ARENA_PAGES, \u0026opts);\n+\tif (map_fd \u003c 0 \u0026\u0026 errno == EOPNOTSUPP) {\n+\t\ttest__skip();\n+\t\treturn;\n+\t}\n+\tif (!ASSERT_GE(map_fd, 0, \"unexported_create\"))\n+\t\treturn;\n+\tcanonical = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, map_fd, 0);\n+\tif (!ASSERT_NEQ(canonical, MAP_FAILED, \"unexported_short_mmap\"))\n+\t\tgoto out;\n+\tcanonical[0] = 42;\n+\tASSERT_EQ(canonical[0], 42, \"unexported_short_access\");\n+\tsnprintf(pin, sizeof(pin), \"/sys/fs/bpf/arena_unexported_%d\", getpid());\n+\tif (!ASSERT_OK(bpf_obj_pin(map_fd, pin), \"unexported_pin\"))\n+\t\tgoto out;\n+\tif (ASSERT_OK(stat(pin, \u0026st), \"unexported_stat\"))\n+\t\tASSERT_EQ(st.st_size, 0, \"unexported_size\");\n+\tfd = open(pin, O_RDONLY);\n+\terr = errno;\n+\tif (ASSERT_EQ(fd, -1, \"unexported_open\"))\n+\t\tASSERT_EQ(err, EIO, \"unexported_open_errno\");\n+\telse\n+\t\tclose(fd);\n+\tfd = bpf_obj_get(pin);\n+\tif (ASSERT_GE(fd, 0, \"unexported_obj_get\"))\n+\t\tclose(fd);\n+\tASSERT_OK(unlink(pin), \"unexported_unlink\");\n+out:\n+\tif (canonical != MAP_FAILED)\n+\t\tmunmap(canonical, ps);\n+\tclose(map_fd);\n+}\n+\n+static void test_export_requires_no_free(void)\n+{\n+\tLIBBPF_OPTS(bpf_map_create_opts, opts,\n+\t\t.map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_EXPORT);\n+\tint fd;\n+\n+\tfd = bpf_map_create(BPF_MAP_TYPE_ARENA, \"arena_export\",\n+\t\t\t 0, 0, ARENA_PAGES, \u0026opts);\n+\tif (fd == -EOPNOTSUPP) {\n+\t\ttest__skip();\n+\t\treturn;\n+\t}\n+\tASSERT_EQ(fd, -EINVAL, \"export_requires_no_free\");\n+\tif (fd \u003e= 0)\n+\t\tclose(fd);\n+}\n+\n+static void test_read_only(int flags)\n+{\n+\tstruct arena_file file;\n+\tsize_t ps = getpagesize();\n+\tchar *alias = MAP_FAILED;\n+\tvoid *addr;\n+\tint fd = -1, err, ret;\n+\n+\tif (!setup(\u0026file, 0, false))\n+\t\tgoto out;\n+\tif (!ASSERT_OK(fchmod(file.file_fd, 0400), \"chmod_read_only\"))\n+\t\tgoto out;\n+\tfd = open(file.pin, O_RDONLY);\n+\tif (!ASSERT_GE(fd, 0, \"open_read_only\"))\n+\t\tgoto out;\n+\talias = mmap(NULL, ps, PROT_READ, flags, fd, ps);\n+\tif (!ASSERT_NEQ(alias, MAP_FAILED, \"read_only_mmap\"))\n+\t\tgoto out;\n+\tASSERT_EQ(alias[0], 0, \"read_only_fault\");\n+\t((char *)file.canonical)[ps] = 42;\n+\tASSERT_EQ(alias[0], 42, \"read_only_shared_update\");\n+\n+\taddr = mmap(NULL, ps, PROT_READ | PROT_WRITE, flags, fd, ps);\n+\terr = errno;\n+\tif (ASSERT_EQ(addr, MAP_FAILED, \"read_only_writable_mmap\"))\n+\t\tASSERT_EQ(err, EACCES, \"read_only_writable_errno\");\n+\telse\n+\t\tmunmap(addr, ps);\n+\tret = mprotect(alias, ps, PROT_READ | PROT_WRITE);\n+\terr = errno;\n+\tASSERT_EQ(ret, -1, \"read_only_mprotect\");\n+\tASSERT_EQ(err, EACCES, \"read_only_mprotect_errno\");\n+\taddr = mmap(NULL, ps, PROT_READ, MAP_PRIVATE, fd, ps);\n+\terr = errno;\n+\tif (ASSERT_EQ(addr, MAP_FAILED, \"read_only_private_mmap\"))\n+\t\tASSERT_EQ(err, EINVAL, \"read_only_private_errno\");\n+\telse\n+\t\tmunmap(addr, ps);\n+\n+\tclose(fd);\n+\tfd = -1;\n+\t((char *)file.canonical)[ps] = 73;\n+\tASSERT_EQ(alias[0], 73, \"read_only_after_close\");\n+out:\n+\tif (alias != MAP_FAILED)\n+\t\tmunmap(alias, ps);\n+\tif (fd \u003e= 0)\n+\t\tclose(fd);\n+\tcleanup(\u0026file);\n+}\n+\n+static void test_fixed_size(void)\n+{\n+\tstruct arena_file file;\n+\tsize_t ps = getpagesize();\n+\tstruct stat st;\n+\tunsigned char resident;\n+\tchar *alias = MAP_FAILED;\n+\tint err, fd, ret;\n+\n+\tif (!setup(\u0026file, 0, false))\n+\t\tgoto out;\n+\talias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, file.file_fd, 0);\n+\tif (!ASSERT_NEQ(alias, MAP_FAILED, \"alias_mmap\"))\n+\t\tgoto out;\n+\talias[0] = 42;\n+\tret = ftruncate(file.file_fd, 0);\n+\terr = errno;\n+\tASSERT_EQ(ret, -1, \"ftruncate_shrink\");\n+\tASSERT_EQ(err, EINVAL, \"ftruncate_shrink_errno\");\n+\tret = ftruncate(file.file_fd, file.len + ps);\n+\terr = errno;\n+\tASSERT_EQ(ret, -1, \"ftruncate_grow\");\n+\tASSERT_EQ(err, EINVAL, \"ftruncate_grow_errno\");\n+\tret = truncate(file.pin, 0);\n+\terr = errno;\n+\tASSERT_EQ(ret, -1, \"truncate_pin\");\n+\tASSERT_EQ(err, EINVAL, \"truncate_pin_errno\");\n+\tfd = open(file.pin, O_RDWR | O_TRUNC);\n+\terr = errno;\n+\tif (ASSERT_EQ(fd, -1, \"open_trunc\"))\n+\t\tASSERT_EQ(err, EINVAL, \"open_trunc_errno\");\n+\telse\n+\t\tclose(fd);\n+\tASSERT_OK(ftruncate(file.file_fd, file.len), \"ftruncate_same_size\");\n+\tif (!ASSERT_OK(mincore(alias, ps, \u0026resident), \"mincore\"))\n+\t\tgoto out;\n+\tASSERT_EQ(resident \u0026 1, 1, \"mapping_not_zapped\");\n+\tASSERT_EQ(alias[0], 42, \"mapping_value\");\n+\tASSERT_OK(fchmod(file.file_fd, 0640), \"chmod_pin\");\n+\tif (ASSERT_OK(fstat(file.file_fd, \u0026st), \"stat_pin\")) {\n+\t\tASSERT_EQ(st.st_size, file.len, \"fixed_size\");\n+\t\tASSERT_EQ(st.st_mode \u0026 0777, 0640, \"pin_mode\");\n+\t}\n+\tASSERT_EQ(lseek(file.file_fd, 0, SEEK_END), file.len, \"seek_capacity\");\n+\tASSERT_EQ(lseek(file.file_fd, -(off_t)ps, SEEK_CUR), file.len - ps, \"seek_back\");\n+\tASSERT_EQ(lseek(file.file_fd, ps, SEEK_SET), ps, \"seek_offset\");\n+\terrno = 0;\n+\tASSERT_EQ(lseek(file.file_fd, file.len + 1, SEEK_SET), -1, \"seek_past_end\");\n+\tASSERT_EQ(errno, EINVAL, \"seek_past_end_errno\");\n+\terrno = 0;\n+\tASSERT_EQ(lseek(file.file_fd, -1, SEEK_SET), -1, \"seek_before_start\");\n+\tASSERT_EQ(errno, EINVAL, \"seek_before_start_errno\");\n+\tfd = bpf_obj_get(file.pin);\n+\tif (ASSERT_GE(fd, 0, \"obj_get_after_chmod\"))\n+\t\tclose(fd);\n+out:\n+\tif (alias != MAP_FAILED)\n+\t\tmunmap(alias, ps);\n+\tcleanup(\u0026file);\n+}\n+\n+static void test_zero_canonical(void)\n+{\n+\tstruct arena_file file;\n+\tsize_t ps = getpagesize();\n+\tchar *alias = MAP_FAILED;\n+\tvoid *addr;\n+\n+\tif (!setup(\u0026file, 0, true))\n+\t\tgoto out;\n+\tif (!ASSERT_EQ((unsigned long)file.canonical, 0, \"zero_canonical\"))\n+\t\tgoto out;\n+\talias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, file.file_fd, ps);\n+\tif (!ASSERT_NEQ(alias, MAP_FAILED, \"zero_canonical_export\"))\n+\t\tgoto out;\n+\talias[0] = 42;\n+\tASSERT_EQ(alias[0], 42, \"zero_canonical_fault\");\n+\taddr = mmap((void *)ps, file.len, PROT_READ | PROT_WRITE,\n+\t\t MAP_SHARED | MAP_FIXED_NOREPLACE, file.map_fd, 0);\n+\tif (ASSERT_EQ(addr, MAP_FAILED, \"canonical_cannot_move\"))\n+\t\tASSERT_EQ(errno, EINVAL, \"canonical_cannot_move_errno\");\n+\telse\n+\t\tmunmap(addr, file.len);\n+out:\n+\tif (alias != MAP_FAILED)\n+\t\tmunmap(alias, ps);\n+\tcleanup(\u0026file);\n+}\n+\n+static void test_lifecycle(void)\n+{\n+\tstruct arena_file file;\n+\tsize_t ps = getpagesize();\n+\tchar *alias = MAP_FAILED;\n+\tvoid *target = MAP_FAILED, *addr;\n+\tpid_t child;\n+\tint status, ret, err;\n+\n+\tif (!setup(\u0026file, 0, false))\n+\t\tgoto out;\n+\talias = mmap(NULL, ps * 2, PROT_READ | PROT_WRITE, MAP_SHARED,\n+\t\t file.file_fd, 0);\n+\tif (!ASSERT_NEQ(alias, MAP_FAILED, \"lifecycle_mmap\"))\n+\t\tgoto out;\n+\talias[0] = 42;\n+\tret = munmap(alias, ps);\n+\terr = errno;\n+\tASSERT_EQ(ret, -1, \"partial_unmap\");\n+\tASSERT_EQ(err, EINVAL, \"partial_unmap_errno\");\n+\tret = mprotect(alias, ps, PROT_READ);\n+\terr = errno;\n+\tASSERT_EQ(ret, -1, \"partial_mprotect\");\n+\tASSERT_EQ(err, EINVAL, \"partial_mprotect_errno\");\n+\tret = madvise(alias, ps * 2, MADV_DOFORK);\n+\terr = errno;\n+\tASSERT_EQ(ret, -1, \"enable_inheritance\");\n+\tASSERT_EQ(err, EINVAL, \"enable_inheritance_errno\");\n+\ttarget = mmap(NULL, ps * 2, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);\n+\tif (!ASSERT_NEQ(target, MAP_FAILED, \"reserve_remap_target\"))\n+\t\tgoto out;\n+\taddr = mremap(alias, ps * 2, ps * 2, MREMAP_MAYMOVE | MREMAP_FIXED, target);\n+\terr = errno;\n+\tif (!ASSERT_EQ(addr, MAP_FAILED, \"relocate_mapping\")) {\n+\t\talias = addr;\n+\t\ttarget = MAP_FAILED;\n+\t}\n+\tASSERT_EQ(err, EINVAL, \"relocate_mapping_errno\");\n+\tchild = fork();\n+\tif (!ASSERT_GE(child, 0, \"fork\"))\n+\t\tgoto out;\n+\tif (!child) {\n+\t\tunsigned char resident[2];\n+\n+\t\tret = mincore(alias, ps * 2, resident);\n+\t\t_exit(ret == -1 \u0026\u0026 errno == ENOMEM ? 0 : 1);\n+\t}\n+\tif (ASSERT_EQ(waitpid(child, \u0026status, 0), child, \"wait_child\") \u0026\u0026\n+\t ASSERT_TRUE(WIFEXITED(status), \"child_exited\"))\n+\t\tASSERT_EQ(WEXITSTATUS(status), 0, \"mapping_not_inherited\");\n+\tASSERT_EQ(alias[0], 42, \"parent_mapping_retained\");\n+out:\n+\tif (target != MAP_FAILED)\n+\t\tmunmap(target, ps * 2);\n+\tif (alias != MAP_FAILED)\n+\t\tmunmap(alias, ps * 2);\n+\tcleanup(\u0026file);\n+}\n+\n+void serial_test_arena_pinned(void)\n+{\n+\tif (test__start_subtest(\"ordinary_pin\"))\n+\t\ttest_unexported_pin(0);\n+\tif (test__start_subtest(\"retained_pin\"))\n+\t\ttest_unexported_pin(BPF_F_ARENA_NO_FREE);\n+\tif (test__start_subtest(\"export_requires_no_free\"))\n+\t\ttest_export_requires_no_free();\n+\tif (test__start_subtest(\"mappings\"))\n+\t\ttest_mappings(0);\n+\tif (test__start_subtest(\"map_extra\"))\n+\t\ttest_mappings(1ULL \u003c\u003c 32);\n+\tif (test__start_subtest(\"fixed_size\"))\n+\t\ttest_fixed_size();\n+\tif (test__start_subtest(\"read_only\"))\n+\t\ttest_read_only(MAP_SHARED);\n+\tif (test__start_subtest(\"read_only_validate\"))\n+\t\ttest_read_only(MAP_SHARED_VALIDATE);\n+\tif (test__start_subtest(\"zero_canonical\"))\n+\t\ttest_zero_canonical();\n+\tif (test__start_subtest(\"lifecycle\"))\n+\t\ttest_lifecycle();\n+}\n"
} |
| 2/1 | 2026/10/10 23:15 | agent | patch-evaluator |
1mModel:core Args: null Results: {
"EnableConfigs": [
"BPF_SYSCALL"
],
"FocusSymbols": [
"arena_map_alloc",
"arena_vm_fault",
"arena_get_unmapped_area",
"arena_map_mmap",
"zap_pages",
"bpf_arena_free_pages",
"bpf_arena_free_pages_non_sleepable",
"bpf_map_setattr",
"bpffs_arena_open",
"bpffs_arena_mmap",
"bpffs_arena_llseek",
"bpffs_arena_get_unmapped_area",
"bpf_mkmap"
],
"Reasoning": "The patch introduces export and retention support for BPF arena maps (BPF_F_ARENA_EXPORT and BPF_F_ARENA_NO_FREE), allowing pinned arenas in bpffs to be opened and mmap'd as shared-memory files with slice offsets. It modifies core BPF arena memory management and bpffs inode operations, altering VMA tracking, page fault resolution (accounting for slice offsets), unmapped area discovery, range zapping, and file operations (open, mmap, setattr, llseek). These paths are reachable from userspace in standard virtualized environments via standard syscalls.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit a034d16eea30a35c1529eb1fb108e6ad8ca18920
Author: syz-cluster <triage@syzkaller.com>
Date: Sat Oct 10 23:15:04 2026 +0000
syz-cluster: applied patch under review
diff --git a/Documentation/bpf/map_arena.rst b/Documentation/bpf/map_arena.rst
new file mode 100644
index 0000000000000..1b9eadd5fb3a1
--- /dev/null
+++ b/Documentation/bpf/map_arena.rst
@@ -0,0 +1,131 @@
+.. SPDX-License-Identifier: GPL-2.0-only
+
+==================
+BPF_MAP_TYPE_ARENA
+==================
+
+A BPF arena provides shared memory that BPF programs and userspace can
+access directly. Create the map with zero key and value sizes and
+``BPF_F_MMAPABLE``. ``max_entries`` specifies its capacity in pages, up to
+4 GiB. The canonical BPF address range must not cross a 4 GiB boundary.
+
+Pinned arena files
+==================
+
+An arena created with ``BPF_F_ARENA_EXPORT`` can be pinned in bpffs and
+opened as a sized file. Its size is ``max_entries * PAGE_SIZE``. Export
+requires ``BPF_F_ARENA_NO_FREE`` so BPF programs cannot release backing
+pages while external consumers may still use them. Creating an arena with
+``BPF_F_ARENA_EXPORT`` alone fails with ``EINVAL``. Arenas without the
+export flag retain their existing bpffs behavior and cannot be opened
+this way, including arenas created with ``BPF_F_ARENA_NO_FREE`` alone.
+
+Establish the canonical BPF address range before mapping the pinned file:
+
+* Set ``map_extra`` to a nonzero, page-aligned address at map creation.
+ The canonical range then covers the full map capacity, even if the map
+ FD has not been mmaped.
+* Alternatively, leave ``map_extra`` zero and mmap the map FD first. This
+ mapping establishes the canonical address and must cover the full map
+ capacity. Shorter canonical mappings fail with ``EINVAL``.
+
+A canonical map FD mapping may start at address zero if the system's
+low-address mapping policy permits it.
+
+Mapping the pinned file before either step fails with ``EINVAL``. Exported
+mappings do not establish or change the canonical BPF address range.
+
+Once the canonical range is established, it covers the entire capacity
+reported by ``fstat()``. Consumers can map the full file or smaller slices.
+Establishing the full canonical mapping reserves virtual address space;
+it does not populate every arena page. Arenas without
+``BPF_F_ARENA_EXPORT`` can still establish shorter canonical mappings.
+
+Open the pin with ``O_RDONLY`` for read-only access or ``O_RDWR`` for
+writable access, and mmap it with ``MAP_SHARED``. An exported
+mapping may use a different virtual address from the canonical mapping.
+Its offset must be page-aligned, and its offset and length must fit within
+both the map capacity and the established canonical range. Consumers can
+map disjoint slices of one arena. Oversized or out-of-range mappings fail
+with ``EINVAL``. ``MAP_PRIVATE`` mappings are unsupported.
+
+Exported mappings share the bytes of the arena without relocating pointers
+stored in those bytes. Arena pointers produced by a BPF program use that
+arena's canonical address representation, based on ``user_vm_start``.
+A consumer mapping a slice at another address must translate such pointers
+before dereferencing them. A guest BPF arena has its own canonical range;
+sharing the backing pages does not make host arena pointers valid in the
+guest arena, or vice versa. Communication structures can use offsets
+relative to the shared slice, with each side validating the offsets and
+translating them to addresses in its own mapping or arena.
+
+A pin opened with ``O_RDONLY`` supports read-only shared mappings, which
+observe updates made by BPF programs and other writable mappings. Requests
+for ``PROT_WRITE`` through that FD fail with ``EACCES``. A mapping made
+without ``PROT_WRITE`` cannot subsequently acquire write permission with
+``mprotect()``, even when the pin was opened with ``O_RDWR``.
+
+Pin permissions can grant readers access independently of writers. For
+example, mode ``0640`` permits the owner to open the pin for writable
+access and the group to open it for read-only access, subject to the
+applicable LSM checks. Changing permissions does not revoke access through
+existing FDs or mappings.
+
+The file has a fixed size. ``truncate()``, ``ftruncate()``, and ``O_TRUNC``
+cannot change it. A truncate request that leaves the size unchanged is
+allowed. Permission and ownership changes remain available.
+Consumers can discover the capacity with ``fstat()`` or
+``lseek(fd, 0, SEEK_END)``. Seeking supports ``SEEK_SET``, ``SEEK_CUR``, and
+``SEEK_END`` within the file bounds; it does not enable ``read()`` or
+``write()`` access to the memory.
+
+``fsync()``, ``fdatasync()``, and ``msync(..., MS_SYNC)`` succeed without
+performing writeback: arena memory is volatile and has no persistent
+backing. These calls do not synchronize access between consumers or BPF
+programs; shared-memory protocols still need appropriate memory ordering.
+
+Exported mappings inherit the arena's mapping lifecycle restrictions.
+They are not inherited across ``fork()``, and ``MADV_DOFORK`` cannot enable
+inheritance. Consumers must create their own mappings in a child process.
+Mappings cannot be split, expanded, or relocated with ``mremap()``.
+Partial unmapping and protection changes that require splitting a mapping
+fail with ``EINVAL``. Unmapping an entire mapping remains supported, as do
+protection changes that do not require a split and satisfy the access
+restrictions above. Consumers needing independently managed regions
+should create separate slice mappings.
+
+Access control and lifetime
+===========================
+
+Opening the pin uses ordinary filesystem permissions and file LSM checks,
+as well as ``security_bpf_map()`` for the requested access mode. Mapping it
+also passes the ordinary mmap LSM checks. Writable mappings remain subject
+to the BPF map's frozen state and program-read-only restrictions.
+
+Consumers use ``open()`` and ``mmap()`` rather than ``bpf(BPF_OBJ_GET)``.
+Consequently, seccomp rules and LSM policies specific to the ``bpf()``
+syscall do not govern this file interface. Grant access through the pin's
+permissions and the applicable file, mmap, and BPF map LSM policies.
+A process prohibited from calling ``bpf()`` can still access the arena
+through this interface if those permissions and checks allow it.
+``unprivileged_bpf_disabled`` restricts unprivileged map and program
+creation; it does not generally prohibit access to existing maps.
+
+``BPF_F_ARENA_NO_FREE`` makes ``bpf_arena_free_pages()`` a no-op in both
+sleepable and non-sleepable BPF programs. The free operation returns no
+indication that pages were retained. Programs using this flag must not
+rely on freeing pages to restore arena allocation capacity. Arena pages
+are retained until the map is destroyed. Open map and pin FDs, mappings,
+and other map references retain the arena, including after the pin is
+unlinked. Removing the pin therefore does not revoke existing mappings or
+free their pages.
+
+The retention flag does not itself expose the pin as a sized file.
+``BPF_F_ARENA_EXPORT`` opts into that interface and must be combined with
+``BPF_F_ARENA_NO_FREE``. Retention-only arenas can still share memory
+through ordinary map FD mappings and manage reusable objects within
+their retained pages.
+
+For example, a consumer such as QEMU can use an exported slice as a shared
+file-backed guest RAM backend. The file interface provides shared memory;
+consumers must supply their own allocation and communication protocol.
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index e0ed44b1bbcb7..40a5489ac9fea 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1500,6 +1500,12 @@ enum {
/* Enable BPF ringbuf overwrite mode */
BPF_F_RB_OVERWRITE = (1U << 19),
+
+ /* Keep arena pages allocated until the map is destroyed. */
+ BPF_F_ARENA_NO_FREE = (1U << 20),
+
+ /* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */
+ BPF_F_ARENA_EXPORT = (1U << 21),
};
/* Flags for BPF_PROG_QUERY. */
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 0ff707707da04..98f6dffe0138f 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -283,7 +283,13 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
/* BPF_F_MMAPABLE must be set */
!(attr->map_flags & BPF_F_MMAPABLE) ||
/* No unsupported flags present */
- (attr->map_flags & ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE | BPF_F_NO_USER_CONV)))
+ (attr->map_flags & ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE |
+ BPF_F_NO_USER_CONV | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT)))
+ return ERR_PTR(-EINVAL);
+
+ /* Exported memory must remain allocated while consumers use it. */
+ if ((attr->map_flags & BPF_F_ARENA_EXPORT) && !(attr->map_flags & BPF_F_ARENA_NO_FREE))
return ERR_PTR(-EINVAL);
if (attr->map_extra & ~PAGE_MASK)
@@ -495,7 +501,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
int ret;
kbase = bpf_arena_get_kern_vm_start(arena);
- kaddr = kbase + (u32)(vmf->address);
+ /* vmf->pgoff includes the file offset of a bpffs-backed slice. */
+ kaddr = kbase + (u32)(arena->user_vm_start + ((u64)vmf->pgoff << PAGE_SHIFT));
page = vmalloc_to_page((void *)kaddr);
if (!page && !(arena->map.map_flags & BPF_F_SEGV_ON_FAULT)) {
@@ -612,13 +619,23 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
long ret;
+ if (filp->f_op != &bpf_map_fops) {
+ if (!len || pgoff >= map->max_entries ||
+ len > ((u64)map->max_entries - pgoff) << PAGE_SHIFT)
+ return -EINVAL;
+ return mm_get_unmapped_area(filp, addr, len, pgoff, flags);
+ }
+
if (pgoff)
return -EINVAL;
if (len > SZ_4G)
return -E2BIG;
+ if ((map->map_flags & BPF_F_ARENA_EXPORT) &&
+ len != (u64)map->max_entries << PAGE_SHIFT)
+ return -EINVAL;
- /* if user_vm_start was specified at arena creation time */
- if (arena->user_vm_start) {
+ /* Once established, the canonical range cannot change. */
+ if (arena->user_vm_end) {
if (len > arena->user_vm_end - arena->user_vm_start)
return -E2BIG;
if (len != arena->user_vm_end - arena->user_vm_start)
@@ -632,7 +649,7 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
return ret;
if ((ret >> 32) == ((ret + len - 1) >> 32))
return ret;
- if (WARN_ON_ONCE(arena->user_vm_start))
+ if (WARN_ON_ONCE(arena->user_vm_end))
/* checks at map creation time should prevent this */
return -EFAULT;
return round_up(ret, SZ_4G);
@@ -641,9 +658,13 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
{
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
+ bool exported = vma->vm_file->f_op != &bpf_map_fops;
guard(mutex)(&arena->lock);
- if (arena->user_vm_start && arena->user_vm_start != vma->vm_start)
+ if (!exported && (map->map_flags & BPF_F_ARENA_EXPORT) &&
+ vma->vm_end - vma->vm_start != (u64)map->max_entries << PAGE_SHIFT)
+ return -EINVAL;
+ if (!exported && arena->user_vm_end && arena->user_vm_start != vma->vm_start)
/*
* If map_extra was not specified at arena creation time then
* 1st user process can do mmap(NULL, ...) to pick user_vm_start
@@ -654,19 +675,31 @@ static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
*/
return -EBUSY;
- if (arena->user_vm_end && arena->user_vm_end != vma->vm_end)
+ if (exported && !arena->user_vm_end)
+ return -EINVAL;
+ if (!exported && arena->user_vm_end && arena->user_vm_end != vma->vm_end)
/* all user processes must have the same size of mmap-ed region */
return -EBUSY;
- /* Earlier checks should prevent this */
- if (WARN_ON_ONCE(vma->vm_end - vma->vm_start > SZ_4G || vma->vm_pgoff))
+ if (exported) {
+ u64 page_cnt = (arena->user_vm_end - arena->user_vm_start) >> PAGE_SHIFT;
+
+ if (vma->vm_pgoff >= page_cnt ||
+ (vma->vm_end - vma->vm_start) >> PAGE_SHIFT >
+ page_cnt - vma->vm_pgoff)
+ return -EINVAL;
+ } else if (WARN_ON_ONCE(vma->vm_end - vma->vm_start > SZ_4G || vma->vm_pgoff)) {
+ /* Earlier checks should prevent this for map FD mappings. */
return -EFAULT;
+ }
if (remember_vma(arena, vma))
return -ENOMEM;
- arena->user_vm_start = vma->vm_start;
- arena->user_vm_end = vma->vm_end;
+ if (!exported) {
+ arena->user_vm_start = vma->vm_start;
+ arena->user_vm_end = vma->vm_end;
+ }
/*
* bpf_map_mmap() checks that it's being mmaped as VM_SHARED and
* clears VM_MAYEXEC. Set VM_DONTEXPAND to avoid potential change
@@ -840,8 +873,12 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
struct mm_struct *mm;
struct vma_list *vml;
unsigned long vm_start;
+ u64 start, end, vma_start, vma_end;
u64 my_gen;
+ start = uaddr - arena->user_vm_start;
+ end = start + size;
+
/*
* Taking mmap_read_lock() under arena->lock would deadlock against
* arena_vm_close(), which runs with mmap_write_lock held and then
@@ -880,8 +917,14 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
*/
vma = find_vma(mm, vm_start);
if (vma && vma->vm_start == vm_start &&
- vma->vm_file && vma->vm_file->private_data == &arena->map)
- zap_vma_range(vma, uaddr, size);
+ vma->vm_file && vma->vm_file->private_data == &arena->map) {
+ vma_start = (u64)vma->vm_pgoff << PAGE_SHIFT;
+ vma_end = vma_start + vma->vm_end - vma->vm_start;
+ if (start < vma_end && end > vma_start)
+ zap_vma_range(vma, vma->vm_start +
+ (max(start, vma_start) - vma_start),
+ min(end, vma_end) - max(start, vma_start));
+ }
mmap_read_unlock(mm);
mmput(mm);
@@ -1134,7 +1177,8 @@ __bpf_kfunc void bpf_arena_free_pages(void *p__map, void *ptr__ign, u32 page_cnt
struct bpf_map *map = p__map;
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
- if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign)
+ if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign ||
+ (map->map_flags & BPF_F_ARENA_NO_FREE))
return;
arena_free_pages(arena, (long)ptr__ign, page_cnt, true);
}
@@ -1144,7 +1188,8 @@ void bpf_arena_free_pages_non_sleepable(void *p__map, void *ptr__ign, u32 page_c
struct bpf_map *map = p__map;
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
- if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign)
+ if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign ||
+ (map->map_flags & BPF_F_ARENA_NO_FREE))
return;
arena_free_pages(arena, (long)ptr__ign, page_cnt, false);
}
diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c
index 7837968c0842c..d3dc6f70e32ae 100644
--- a/kernel/bpf/inode.c
+++ b/kernel/bpf/inode.c
@@ -119,8 +119,24 @@ static const struct inode_operations bpf_symlink_iops;
static const struct inode_operations bpf_prog_iops = {
.listxattr = bpf_fs_listxattr,
};
+
+static int bpf_map_setattr(struct mnt_idmap *idmap, struct dentry *dentry,
+ struct iattr *attr)
+{
+ struct inode *inode = d_inode(dentry);
+ struct bpf_map *map = inode->i_private;
+
+ if (map->map_type == BPF_MAP_TYPE_ARENA &&
+ (map->map_flags & BPF_F_ARENA_EXPORT) &&
+ (attr->ia_valid & ATTR_SIZE) && attr->ia_size != i_size_read(inode))
+ return -EINVAL;
+
+ return simple_setattr(idmap, dentry, attr);
+}
+
static const struct inode_operations bpf_map_iops = {
.listxattr = bpf_fs_listxattr,
+ .setattr = bpf_map_setattr,
};
static const struct inode_operations bpf_link_iops = {
.listxattr = bpf_fs_listxattr,
@@ -351,6 +367,68 @@ static const struct file_operations bpffs_map_fops = {
.release = bpffs_map_release,
};
+/*
+ * An arena with BPF_F_ARENA_EXPORT pinned in bpffs can serve as a
+ * shared-memory file. The ordinary map FD is an anonymous inode with no
+ * size and cannot be opened by pathname, so keep the pin's inode and
+ * forward mmap to the map FD implementation. The pinned file uses ordinary
+ * address selection so QEMU can map the pages at another virtual address.
+ */
+static int bpffs_arena_open(struct inode *inode, struct file *file)
+{
+ struct bpf_map *map = inode->i_private;
+ int err;
+
+ err = security_bpf_map(map, file->f_mode);
+ if (err)
+ return err;
+ bpf_map_inc_with_uref(map);
+ file->private_data = map;
+
+ return 0;
+}
+
+static int bpffs_arena_mmap(struct file *file, struct vm_area_struct *vma)
+{
+ /*
+ * Generic mmap clears VM_SHARED for O_RDONLY files, but VM_MAYSHARE
+ * still distinguishes shared mappings from private ones. Restore
+ * VM_SHARED for the map mmap path without granting VM_MAYWRITE.
+ */
+ if (!(file->f_mode & FMODE_WRITE) && (vma->vm_flags & VM_MAYSHARE))
+ vm_flags_set(vma, VM_SHARED);
+
+ return bpf_map_fops.mmap(file, vma);
+}
+
+static int bpffs_arena_release(struct inode *inode, struct file *file)
+{
+ return bpf_map_fops.release(inode, file);
+}
+
+static unsigned long bpffs_arena_get_unmapped_area(struct file *file,
+ unsigned long addr,
+ unsigned long len,
+ unsigned long pgoff,
+ unsigned long flags)
+{
+ return bpf_map_fops.get_unmapped_area(file, addr, len, pgoff, flags);
+}
+
+static loff_t bpffs_arena_llseek(struct file *file, loff_t offset, int whence)
+{
+ return fixed_size_llseek(file, offset, whence, i_size_read(file_inode(file)));
+}
+
+static const struct file_operations bpffs_arena_fops = {
+ .open = bpffs_arena_open,
+ .llseek = bpffs_arena_llseek,
+ .fsync = noop_fsync,
+ .release = bpffs_arena_release,
+ .mmap = bpffs_arena_mmap,
+ .get_unmapped_area = bpffs_arena_get_unmapped_area,
+};
+
static int bpffs_obj_open(struct inode *inode, struct file *file)
{
return -EIO;
@@ -362,7 +440,7 @@ static const struct file_operations bpffs_obj_fops = {
static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
const struct inode_operations *iops,
- const struct file_operations *fops)
+ const struct file_operations *fops, loff_t size)
{
struct inode *dir = dentry->d_parent->d_inode;
struct inode *inode;
@@ -382,6 +460,7 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
inode->i_op = iops;
inode->i_fop = fops;
inode->i_private = raw;
+ i_size_write(inode, size);
bpf_dentry_finalize(dentry, inode, dir);
return 0;
@@ -390,16 +469,20 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
static int bpf_mkprog(struct dentry *dentry, umode_t mode, void *arg)
{
return bpf_mkobj_ops(dentry, mode, arg, &bpf_prog_iops,
- &bpffs_obj_fops);
+ &bpffs_obj_fops, 0);
}
static int bpf_mkmap(struct dentry *dentry, umode_t mode, void *arg)
{
struct bpf_map *map = arg;
+ bool shared_arena = map->map_type == BPF_MAP_TYPE_ARENA &&
+ (map->map_flags & BPF_F_ARENA_EXPORT);
return bpf_mkobj_ops(dentry, mode, arg, &bpf_map_iops,
- bpf_map_support_seq_show(map) ?
- &bpffs_map_fops : &bpffs_obj_fops);
+ shared_arena ? &bpffs_arena_fops :
+ bpf_map_support_seq_show(map) ?
+ &bpffs_map_fops : &bpffs_obj_fops,
+ shared_arena ? (loff_t)map->max_entries * PAGE_SIZE : 0);
}
static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)
@@ -408,7 +491,7 @@ static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)
return bpf_mkobj_ops(dentry, mode, arg, &bpf_link_iops,
bpf_link_is_iter(link) ?
- &bpf_iter_fops : &bpffs_obj_fops);
+ &bpf_iter_fops : &bpffs_obj_fops, 0);
}
static struct dentry *
@@ -483,7 +566,7 @@ static int bpf_iter_link_pin_kernel(struct dentry *parent,
if (IS_ERR(dentry))
return PTR_ERR(dentry);
ret = bpf_mkobj_ops(dentry, mode, link, &bpf_link_iops,
- &bpf_iter_fops);
+ &bpf_iter_fops, 0);
simple_done_creating(dentry);
return ret;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index e0ed44b1bbcb7..40a5489ac9fea 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1500,6 +1500,12 @@ enum {
/* Enable BPF ringbuf overwrite mode */
BPF_F_RB_OVERWRITE = (1U << 19),
+
+ /* Keep arena pages allocated until the map is destroyed. */
+ BPF_F_ARENA_NO_FREE = (1U << 20),
+
+ /* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */
+ BPF_F_ARENA_EXPORT = (1U << 21),
};
/* Flags for BPF_PROG_QUERY. */
diff --git a/tools/testing/selftests/bpf/.gitignore b/tools/testing/selftests/bpf/.gitignore
index b815bf0d88774..9747ad37d698e 100644
--- a/tools/testing/selftests/bpf/.gitignore
+++ b/tools/testing/selftests/bpf/.gitignore
@@ -47,3 +47,5 @@ verification_cert.h
*.BTF.base
usdt_1
usdt_2
+/arena_kvm_guest-init
+/arena_kvm_host-runner
diff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile
index a22be7efd1fa2..29fffa7857c19 100644
--- a/tools/testing/selftests/bpf/Makefile
+++ b/tools/testing/selftests/bpf/Makefile
@@ -42,7 +42,13 @@ TEST_PROGS := test_kmod.sh \
test_bpftool_build.sh \
test_doc_build.sh \
test_xsk.sh \
- test_xdp_features.sh
+ test_xdp_features.sh \
+ arena_kvm.py
+
+TEST_GEN_FILES += arena_kvm_guest-init arena_kvm_host-runner
+ifneq ($(CLANG_CPUV4),)
+TEST_GEN_FILES += arena_kvm_guest.bpf.o arena_kvm_host.bpf.o
+endif
TEST_PROGS_EXTENDED := \
ima_setup.sh verify_sig_setup.sh
@@ -94,6 +100,25 @@ BPFTOOLDIR := $(TOOLSDIR)/bpf/bpftool
HOST_BPFOBJ := $(HOST_BUILD_DIR)/libbpf/libbpf.a
BPF_TARGET_ENDIAN:=$(if $(IS_LITTLE_ENDIAN),--target=bpfel,--target=bpfeb)
+# These helpers are installed beside arena_kvm.py, including for OUTPUT=.
+$(OUTPUT)/arena_kvm_%.bpf.o: arena_kvm_%.bpf.c arena_kvm_shared.h \
+ libarena/include/bpf_atomic.h \
+ libarena/include/bpf_arena_common.h \
+ $(INCLUDE_DIR)/vmlinux.h $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BPF,,$@)
+ $(Q)$(CLANG) $(BPF_CFLAGS) $(CLANG_CFLAGS) -O2 \
+ $(BPF_TARGET_ENDIAN) -mcpu=v4 -c $< -o $@ $(call skip_on_fail,BPF)
+
+$(OUTPUT)/arena_kvm_guest-init: arena_kvm_guest.c arena_kvm_shared.h \
+ $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BINARY,,$@)
+ $(Q)$(CC) $(CFLAGS) $(LDFLAGS) $< $(BPFOBJ) $(LDLIBS) -lzstd -o $@
+
+$(OUTPUT)/arena_kvm_host-runner: arena_kvm_host.c arena_kvm_shared.h \
+ $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BINARY,,$@)
+ $(Q)$(CC) $(CFLAGS) $(LDFLAGS) $< $(BPFOBJ) $(LDLIBS) -lzstd -o $@
+
NON_CHECK_FEAT_TARGETS := clean docs-clean emit_tests
CHECK_FEAT := $(filter-out $(NON_CHECK_FEAT_TARGETS),$(or $(MAKECMDGOALS), "none"))
ifneq ($(CHECK_FEAT),)
diff --git a/tools/testing/selftests/bpf/arena_kvm.py b/tools/testing/selftests/bpf/arena_kvm.py
new file mode 100755
index 0000000000000..552bfcfb80ac8
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm.py
@@ -0,0 +1,448 @@
+#!/usr/bin/env python3
+# SPDX-License-Identifier: GPL-2.0
+# Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+"""Exchange values between host and guest BPF through NUMA-backed RAM.
+
+Build with the BPF selftests Makefile. Helpers are found in the current
+directory or beside this script, so OUTPUT builds and installed tests work.
+Use --build-dir to select another helper directory and --kernel (or
+ARENA_KVM_KERNEL) to select the kernel booted by virtme-ng. By default, use
+the source tree's kernel when available, otherwise the running release.
+"""
+
+# Dependencies: Python 3.9+, virtme-ng, busybox-static, qemu, util-linux (script).
+
+import argparse
+import array
+import ctypes
+import mmap
+import os
+from pathlib import Path
+import queue
+import re
+import shutil
+import shlex
+import signal
+import socket
+import struct
+import subprocess
+import sys
+import threading
+import time
+
+
+PAGE_SIZE = os.sysconf("SC_PAGE_SIZE")
+KSFT_SKIP = 4
+HELPER_TIMEOUT = 20
+GUEST_MEMORY = 512 * 1024 * 1024
+# RAM starts for virtme-ng's PC/q35, virt (arm64/riscv64), and pseries
+# machines. With 512 MiB of RAM, the last NUMA node follows node 0
+# without crossing a machine's RAM hole. Skip other layouts.
+GUEST_RAM_BASES = {
+ "x86_64": 0,
+ "aarch64": 0x40000000,
+ "arm64": 0x40000000,
+ "riscv64": 0x80000000,
+ "ppc64": 0,
+ "ppc64le": 0,
+}
+# Leave room for guest boot allocations and allocator watermarks.
+# QEMU's pseries machine requires each NUMA node to be 256 MiB aligned.
+SLICE_SIZE = (256 if os.uname().machine in ("ppc64", "ppc64le") else 16) \
+ * 1024 * 1024
+ARENA_SIZE = 2 * SLICE_SIZE
+NUM_PAGES = ARENA_SIZE // PAGE_SIZE
+GUEST_READY = 0x4152454E414B564D
+HOST_FIRST = 0x123456789ABCDEF0
+GUEST_FIRST = 0xFEDCBA9876543210
+HOST_SECOND = 0x1020304050607080
+GUEST_SECOND = 0x8070605040302010
+
+
+def progress(message):
+ print(f"# arena_kvm: {message}", flush=True)
+
+
+class TestSkipped(Exception):
+ pass
+
+
+class TestTerminated(BaseException):
+ def __init__(self, signum):
+ self.signum = signum
+
+
+def terminate(signum, frame):
+ # A second signal must not interrupt cleanup of pins and descendants.
+ signal.signal(signal.SIGTERM, signal.SIG_IGN)
+ signal.signal(signal.SIGINT, signal.SIG_IGN)
+ raise TestTerminated(signum)
+
+
+def enable_subreaper():
+ # Adopt orphaned descendants even when script/vng changes session.
+ libc = ctypes.CDLL(None, use_errno=True)
+ libc.prctl.argtypes = (ctypes.c_int, *([ctypes.c_ulong] * 4))
+ if libc.prctl(36, 1, 0, 0, 0): # PR_SET_CHILD_SUBREAPER
+ err = ctypes.get_errno()
+ raise OSError(err, os.strerror(err))
+
+
+def children(pid):
+ # /proc/PID/task/TID/children requires CONFIG_CHECKPOINT_RESTORE.
+ # PPid is always available and covers children of every parent thread.
+ result = set()
+ for path in Path("/proc").iterdir():
+ if not path.name.isdigit():
+ continue
+ try:
+ status = (path / "status").read_text()
+ except (FileNotFoundError, ProcessLookupError, PermissionError):
+ continue
+ for line in status.splitlines():
+ if line.startswith("PPid:") and int(line.split()[1]) == pid:
+ result.add(int(path.name))
+ break
+ return result
+
+
+def stop_children(processes=()):
+ # Kill direct children, then repeat for descendants adopted by the
+ # subreaper. Sessions and process groups do not affect adoption.
+ # Reaping until ECHILD verifies that the entire owned tree has exited.
+ known = {process.pid: process for process in processes}
+ deadline = time.monotonic() + 10
+ while True:
+ for pid in children(os.getpid()):
+ try:
+ fd = os.pidfd_open(pid)
+ except ProcessLookupError:
+ continue
+ try:
+ signal.pidfd_send_signal(fd, signal.SIGKILL)
+ except ProcessLookupError:
+ pass
+ finally:
+ os.close(fd)
+ while True:
+ try:
+ pid, status = os.waitpid(-1, os.WNOHANG)
+ except ChildProcessError:
+ return
+ if not pid:
+ break
+ if pid in known:
+ known[pid].returncode = os.waitstatus_to_exitcode(status)
+ if time.monotonic() >= deadline:
+ raise TimeoutError("guest descendants did not exit")
+ time.sleep(0.01)
+
+
+def getq(arena, offset):
+ return struct.unpack_from("=Q", arena, offset)[0]
+
+
+def wait_reply(runner, obj, fd, offset, guest, lines):
+ deadline = time.monotonic() + 20
+ while time.monotonic() < deadline:
+ result = run_host_bpf(runner, obj, fd, offset, "check_reply",
+ timeout=max(0.001, deadline - time.monotonic()))
+ if result == 0:
+ return
+ if result != 1:
+ raise RuntimeError(f"guest BPF reply is wrong: result={result}")
+ if guest.poll() is not None:
+ raise RuntimeError("guest exited early:\n" + "".join(lines))
+ time.sleep(0.001)
+ raise TimeoutError(f"timed out waiting for BPF reply at offset {offset}")
+
+
+def read_lines(guest, messages, lines):
+ for line in guest.stdout:
+ lines.append(line)
+ # Serial startup may prefix a message with NULs echoed as ^@.
+ messages.put(re.sub(r"^(?:\x00|\^@)+", "", line.strip()))
+
+
+def wait_message(messages, lines, expected, guest):
+ deadline = time.monotonic() + 45
+ while time.monotonic() < deadline:
+ try:
+ line = messages.get(timeout=0.5)
+ except queue.Empty:
+ if guest.poll() is not None:
+ break
+ continue
+ if expected in line:
+ return line
+ if "GUEST_BPF_FAILED" in line:
+ break
+ if "GUEST_BPF_SKIPPED" in line:
+ raise TestSkipped(line)
+ raise RuntimeError(f"guest did not print {expected} "
+ f"(launcher exit status: {guest.poll()}):\n" +
+ "".join(lines))
+
+
+def run_host_bpf(runner, obj, fd, offset, program="exchange",
+ timeout=HELPER_TIMEOUT):
+ output = subprocess.check_output(
+ [str(runner), str(fd), str(obj), str(offset), program],
+ pass_fds=(fd,), text=True, timeout=timeout)
+ return int(output)
+
+
+def resolve_vng():
+ executable = shutil.which("vng")
+ if not executable or not os.access(executable, os.X_OK):
+ print("SKIP: virtme-ng (vng) is unavailable")
+ return None, None
+ return executable, os.environ.copy()
+
+
+def start_guest(pin, build, kernel, index, vng, env, guests):
+ node0_size = GUEST_MEMORY - SLICE_SIZE
+ base = GUEST_RAM_BASES[os.uname().machine] + node0_size
+ progress(f"Starting guest {index}: {SLICE_SIZE // (1024 * 1024)} MiB "
+ f"slice at file offset {index * SLICE_SIZE:#x}, "
+ f"NUMA node 1 physical base {base:#x}")
+ qemu_opts = (
+ f"-object memory-backend-file,id=signal,size={SLICE_SIZE},share=on,"
+ f"offset={index * SLICE_SIZE},mem-path={pin} "
+ "-numa node,nodeid=1,memdev=signal"
+ )
+ guest_command = [str(build / "arena_kvm_guest-init"),
+ str(build / "arena_kvm_guest.bpf.o")]
+ command = [
+ vng, "--verbose", "--run", str(kernel), "--cpus", "2",
+ "--memory", f"{GUEST_MEMORY // (1024 * 1024)}M",
+ "--numa", str(node0_size),
+ "--exec",
+ shlex.join(guest_command),
+ "--append", f"arena_kvm.base={base:#x} arena_kvm.size={SLICE_SIZE}",
+ f"--qemu-opts={qemu_opts}",
+ ]
+ lines = []
+ messages = queue.Queue()
+ guest = subprocess.Popen(["script", "-e", "-q", "-c", shlex.join(command),
+ "/dev/null"], stdout=subprocess.PIPE,
+ stderr=subprocess.STDOUT, text=True, env=env,
+ start_new_session=True)
+ # Register ownership before thread construction or startup can fail.
+ guests.append((guest, messages, lines, None))
+ reader = threading.Thread(target=read_lines, args=(guest, messages, lines),
+ daemon=True)
+ guests[-1] = guest, messages, lines, reader
+ reader.start()
+
+
+def stop_guests(guests):
+ stop_children(process for process, _, _, _ in guests)
+ for process, _, _, reader in guests:
+ process.wait(timeout=10)
+ if reader is not None and reader.ident is not None:
+ reader.join(timeout=2)
+ if reader.is_alive():
+ raise RuntimeError("guest output reader did not exit")
+ process.stdout.close()
+
+
+def exchange(pin, fd, arena, build, kernel, vng, env):
+ guests = []
+ try:
+ for index in range(2):
+ start_guest(pin, build, kernel, index, vng, env, guests)
+ offsets = []
+ for index, (process, messages, lines, _) in enumerate(guests):
+ progress(f"Waiting for guest {index} to allocate its BPF arena page "
+ "and check that BPF cannot free it")
+ ready = wait_message(messages, lines, "GUEST_BPF_READY", process)
+ match = re.fullmatch(r"GUEST_BPF_READY (\d+) (\d+)", ready)
+ if not match:
+ raise RuntimeError(f"malformed guest page offset: {ready}")
+ relative = int(match.group(1))
+ if int(match.group(2)) != PAGE_SIZE:
+ raise RuntimeError("host and guest page sizes differ")
+ offset = index * SLICE_SIZE + relative
+ if relative >= SLICE_SIZE or relative % PAGE_SIZE or \
+ getq(arena, offset) != GUEST_READY:
+ raise RuntimeError(f"invalid guest {index} page offset {relative}")
+ offsets.append(offset)
+ progress(f"Guest {index} ready: page offset {relative:#x} in slice, "
+ f"{offset:#x} in host arena")
+ # The QEMU VMA and the open map FD retain the arena after unlink.
+ progress("Unpinning the host arena while both guests retain mappings")
+ os.unlink(pin)
+ runner = build / "arena_kvm_host-runner"
+ obj = build / "arena_kvm_host.bpf.o"
+ for offset in offsets:
+ if run_host_bpf(runner, obj, fd, offset, "probe_free") != 0 or \
+ getq(arena, offset) != GUEST_READY:
+ raise RuntimeError("host BPF freed a shared RAM page")
+ progress("Host BPF page-free checks passed for both shared pages")
+
+ for index, offset in enumerate(offsets):
+ progress(f"Round 1: host -> guest {index}, value {HOST_FIRST:#018x}")
+ result = run_host_bpf(runner, obj, fd, offset)
+ if result != 0:
+ raise RuntimeError(f"guest {index} first result={result}")
+ other = offsets[1 - index]
+ if index == 0 and getq(arena, other + 16) != 0:
+ raise RuntimeError("guest 0 write modified guest 1 signal")
+ if index == 0:
+ progress("Guest 1 signal unchanged after host write to guest 0")
+
+ for index, (process, messages, lines, _) in enumerate(guests):
+ offset = offsets[index]
+ wait_reply(runner, obj, fd, offset, process, lines)
+ progress(f"Round 1: guest {index} -> host, "
+ f"verified value {GUEST_FIRST:#018x}")
+
+ # Guest 0 exits while guest 1 keeps its slice mapped and active.
+ for index, (process, messages, lines, _) in enumerate(guests):
+ offset = offsets[index]
+ progress(f"Round 2: host -> guest {index}, value {HOST_SECOND:#018x}")
+ if run_host_bpf(runner, obj, fd, offset) != 0:
+ raise RuntimeError(f"host BPF failed guest {index} second value")
+ wait_reply(runner, obj, fd, offset, process, lines)
+ progress(f"Round 2: guest {index} -> host, "
+ f"verified value {GUEST_SECOND:#018x}")
+ if run_host_bpf(runner, obj, fd, offset) != 2:
+ raise RuntimeError(f"host BPF failed guest {index} reply")
+ wait_message(messages, lines, "GUEST_BPF_EXCHANGED", process)
+ process.wait(timeout=15)
+ if process.returncode:
+ raise RuntimeError(f"guest {index} failed:\n" + "".join(lines))
+ progress(f"Guest {index} completed and exited successfully")
+ if index == 0:
+ progress("Continuing with guest 1 after guest 0 exits")
+ print(f"OK: two guests exchanged BPF values through arena slices {offsets}")
+ finally:
+ progress("Cleaning up guest processes")
+ stop_guests(guests)
+
+
+def skip(reason):
+ print(f"SKIP: {reason}")
+ return KSFT_SKIP
+
+
+def create_arena(pin, build):
+ parent, child = socket.socketpair()
+ with parent, child:
+ result = subprocess.run(
+ [str(build / "arena_kvm_host-runner"), "--create", pin,
+ str(child.fileno()), str(NUM_PAGES)], pass_fds=(child.fileno(),),
+ timeout=HELPER_TIMEOUT)
+ child.close()
+ if result.returncode == KSFT_SKIP:
+ return None
+ result.check_returncode()
+ fds = array.array("i")
+ _, ancillary, flags, _ = parent.recvmsg(
+ 1, socket.CMSG_SPACE(fds.itemsize))
+ for level, kind, data in ancillary:
+ if level == socket.SOL_SOCKET and kind == socket.SCM_RIGHTS:
+ fds.frombytes(data[:len(data) - len(data) % fds.itemsize])
+ if flags & socket.MSG_CTRUNC or len(fds) != 1:
+ for fd in fds:
+ os.close(fd)
+ raise RuntimeError("host helper did not send one arena FD")
+ return fds[0]
+
+
+def main():
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--build-dir", type=Path,
+ help="directory containing the built helpers")
+ parser.add_argument("--kernel", default=os.environ.get("ARENA_KVM_KERNEL"),
+ help="virtme-ng kernel path or release (default: source "
+ "tree kernel, otherwise running kernel)")
+ args = parser.parse_args()
+ if not hasattr(os, "pidfd_open") or not hasattr(signal, "pidfd_send_signal"):
+ return skip("Python 3.9+ with Linux pidfd support is required")
+ machine = os.uname().machine
+ if machine not in GUEST_RAM_BASES:
+ return skip(f"unknown guest RAM layout for architecture {machine}")
+ if struct.calcsize("P") != 8:
+ return skip("BPF arenas require a supported 64-bit architecture")
+ if os.geteuid() != 0:
+ return skip("this test requires root")
+ if SLICE_SIZE % PAGE_SIZE:
+ return skip("arena slices must contain a whole number of pages")
+ if not os.path.ismount("/sys/fs/bpf"):
+ return skip("mount bpffs first")
+ directory = Path(__file__).resolve().parent
+ build = args.build_dir
+ if build is None:
+ build = Path.cwd() if (Path.cwd() / "arena_kvm_host-runner").is_file() \
+ else directory
+ build = build.resolve()
+ kernel = args.kernel
+ if kernel is None:
+ kernel = next((str(p) for p in directory.parents
+ if (p / "vmlinux").is_file()), os.uname().release)
+ vng, env = resolve_vng()
+ if not vng:
+ return KSFT_SKIP
+ qemu_arch = {"arm64": "aarch64", "ppc64le": "ppc64"}.get(machine, machine)
+ for tool in ("script", f"qemu-system-{qemu_arch}"):
+ if not shutil.which(tool, path=env.get("PATH")):
+ return skip(f"{tool} is unavailable")
+ for name in ("arena_kvm_guest-init", "arena_kvm_guest.bpf.o",
+ "arena_kvm_host-runner", "arena_kvm_host.bpf.o"):
+ if not (build / name).exists():
+ return skip(f"missing {build / name}; build the BPF selftests")
+ for name in ("arena_kvm_host.bpf.o", "arena_kvm_guest.bpf.o"):
+ result = subprocess.run([str(build / "arena_kvm_host-runner"),
+ "--check-bpf", str(build / name)],
+ timeout=HELPER_TIMEOUT)
+ if result.returncode == KSFT_SKIP:
+ return skip("BPF compiler lacks arena address-space conversions")
+ result.check_returncode()
+ if subprocess.run([str(build / "arena_kvm_host-runner"),
+ "--check-kvm"], timeout=HELPER_TIMEOUT).returncode:
+ return skip("KVM is unavailable or its API version is unsupported")
+ pin = f"/sys/fs/bpf/arena_kvm_{os.getpid()}"
+ progress(f"Architecture: {machine}; "
+ f"page size: {PAGE_SIZE} bytes")
+ progress(f"Creating and pinning a {ARENA_SIZE // (1024 * 1024)} MiB "
+ f"host arena at {pin}")
+ try:
+ fd = create_arena(pin, build)
+ if fd is None:
+ return skip("pinned BPF arenas are unsupported")
+ try:
+ if os.stat(pin).st_size != ARENA_SIZE:
+ raise RuntimeError("pinned arena has the wrong size")
+ with mmap.mmap(fd, ARENA_SIZE, flags=mmap.MAP_SHARED,
+ prot=mmap.PROT_READ | mmap.PROT_WRITE) as arena:
+ progress(f"Populating {NUM_PAGES} host arena pages before "
+ "starting the guests")
+ # Populate the file pages before QEMU uses them as RAM.
+ for i in range(NUM_PAGES):
+ arena[i * PAGE_SIZE]
+ exchange(pin, fd, arena, build, kernel, vng, env)
+ finally:
+ os.close(fd)
+ finally:
+ if os.path.exists(pin):
+ os.unlink(pin)
+
+
+if __name__ == "__main__":
+ enable_subreaper()
+ signal.signal(signal.SIGTERM, terminate)
+ signal.signal(signal.SIGINT, terminate)
+ try:
+ status = main()
+ except TestSkipped as err:
+ status = skip(str(err))
+ except TestTerminated as err:
+ progress(f"Terminated by signal {err.signum}")
+ status = 128 + err.signum
+ finally:
+ signal.signal(signal.SIGTERM, signal.SIG_IGN)
+ signal.signal(signal.SIGINT, signal.SIG_IGN)
+ stop_children()
+ sys.exit(status)
diff --git a/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c b/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c
new file mode 100644
index 0000000000000..461562149c57f
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c
@@ -0,0 +1,89 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+#include "bpf_arena_common.h"
+#ifdef __TARGET_ARCH_powerpc
+/* PowerPC does not support arena load-acquire/store-release instructions. */
+#undef __BPF_FEATURE_LOAD_ACQ_STORE_REL
+#endif
+#include "bpf_atomic.h"
+#endif
+#include "arena_kvm_shared.h"
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARENA);
+ __uint(map_flags, BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE);
+ __uint(max_entries, 1);
+} arena SEC(".maps");
+
+volatile __u32 signal_offset;
+
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+
+SEC("syscall")
+int allocate(void *ctx)
+{
+ struct signal_page __arena *page;
+
+ page = bpf_arena_alloc_pages(&arena, NULL, 1, ARENA_KVM_NODE, 0);
+ if (!page)
+ return 1;
+ signal_offset = (__u32)((__u64)page - (__u64)arena_base(&arena));
+ smp_store_release(&page->ready, GUEST_READY);
+ return 0;
+}
+
+SEC("syscall")
+int probe_free(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ bpf_arena_free_pages(&arena, page, 1);
+ return page->ready == GUEST_READY ? 0 : 1;
+}
+
+SEC("syscall")
+int exchange(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = smp_load_acquire(&page->h2g_seq);
+ __u64 payload;
+
+ if ((seq != 1 && seq != 2) || page->g2h_seq == seq)
+ return 1;
+ payload = page->h2g_payload;
+ if (seq == 1 && payload != HOST_FIRST)
+ return 2;
+ if (seq == 2 && payload != HOST_SECOND)
+ return 3;
+ page->g2h_payload = seq == 1 ? GUEST_FIRST : GUEST_SECOND;
+ smp_store_release(&page->g2h_seq, seq);
+ return 0;
+}
+
+SEC("syscall")
+int complete(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ return smp_load_acquire(&page->h2g_seq) == 3 ? 0 : 1;
+}
+
+#else
+SEC("syscall") int allocate(void *ctx) { return 1; }
+SEC("syscall") int probe_free(void *ctx) { return 1; }
+SEC("syscall") int exchange(void *ctx) { return 1; }
+SEC("syscall") int complete(void *ctx) { return 1; }
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/arena_kvm_guest.c b/tools/testing/selftests/bpf/arena_kvm_guest.c
new file mode 100644
index 0000000000000..da6586ae660d1
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_guest.c
@@ -0,0 +1,224 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ *
+ * Guest helper for a BPF arena allocated from shared NUMA RAM.
+ */
+#define __EXPORTED_HEADERS__
+
+#include <bpf/bpf.h>
+#include <bpf/libbpf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <linux/bpf.h>
+#include <linux/mempolicy.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/mount.h>
+#include <sys/syscall.h>
+#include <unistd.h>
+
+#include "arena_kvm_shared.h"
+
+static int signal_file_offset(void *arena, uint64_t *offset)
+{
+ char cmdline[8192], *arg, *end;
+ uint64_t entry, pfn, gpa, base = 0, size = 0, *value;
+ ssize_t n;
+ int fd, node, found = 0;
+
+ /* The launcher supplies QEMU's guest physical base for this slice. */
+ fd = open("/proc/cmdline", O_RDONLY);
+ if (fd < 0)
+ return -errno;
+ n = read(fd, cmdline, sizeof(cmdline) - 1);
+ close(fd);
+ if (n <= 0 || n == sizeof(cmdline) - 1)
+ return -EIO;
+ cmdline[n] = '\0';
+ for (arg = strtok(cmdline, "\n "); arg; arg = strtok(NULL, "\n ")) {
+ if (!strncmp(arg, "arena_kvm.base=", 15)) {
+ value = &base;
+ found |= 1;
+ } else if (!strncmp(arg, "arena_kvm.size=", 15)) {
+ value = &size;
+ found |= 2;
+ } else {
+ continue;
+ }
+ errno = 0;
+ *value = strtoull(arg + 15, &end, 0);
+ if (errno || *end || end == arg + 15 || *value % getpagesize())
+ return -EINVAL;
+ }
+ if (found != 3 || size < getpagesize())
+ return -EINVAL;
+
+ /* Fault the BPF-allocated page into the guest user VMA. */
+ if (!*(volatile uint64_t *)arena)
+ return -EINVAL;
+ if (syscall(SYS_get_mempolicy, &node, NULL, 0, arena,
+ MPOL_F_NODE | MPOL_F_ADDR))
+ return -errno;
+ if (node != ARENA_KVM_NODE)
+ return -EXDEV;
+ fd = open("/proc/self/pagemap", O_RDONLY);
+ if (fd < 0)
+ return -errno;
+ n = pread(fd, &entry, sizeof(entry),
+ (uintptr_t)arena / getpagesize() * sizeof(entry));
+ close(fd);
+ if (n != sizeof(entry) || !(entry & (1ULL << 63)))
+ return -EIO;
+ pfn = entry & ((1ULL << 55) - 1);
+ if (!pfn)
+ return -EPERM;
+ gpa = pfn * getpagesize();
+ if (gpa < base || gpa - base >= size || size - (gpa - base) < getpagesize())
+ return -ERANGE;
+ *offset = gpa - base;
+ return 0;
+}
+
+static int create_arena(void)
+{
+ union bpf_attr attr = {
+ .map_type = BPF_MAP_TYPE_ARENA,
+ .max_entries = 1,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE,
+ };
+
+ return syscall(__NR_bpf, BPF_MAP_CREATE, &attr, sizeof(attr));
+}
+
+static int run_prog(int fd, unsigned int *retval)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ int err = bpf_prog_test_run_opts(fd, &opts);
+
+ if (err)
+ return err;
+ *retval = opts.retval;
+ return 0;
+}
+
+static int exchange_once(int fd)
+{
+ unsigned int result = 0;
+ int err;
+
+ /* The host bounds each protocol stage and stops the guest on timeout. */
+ for (;;) {
+ err = run_prog(fd, &result);
+ if (err || result > 1)
+ return err ? err : -EINVAL;
+ if (!result)
+ return 0;
+ usleep(1000);
+ }
+}
+
+int main(int argc, char **argv)
+{
+ struct bpf_object *obj = NULL;
+ struct bpf_program *alloc, *probe_free, *exchange, *complete;
+ struct bpf_map *map;
+ void *arena = MAP_FAILED;
+ uint64_t file_offset;
+ unsigned int result = 0;
+ int fd = -1;
+ int err;
+
+ setbuf(stdout, NULL);
+ if (argc != 2 ||
+ (mount("sysfs", "/sys", "sysfs", 0, NULL) && errno != EBUSY) ||
+ (mount("proc", "/proc", "proc", 0, NULL) && errno != EBUSY))
+ goto fail;
+ if (access("/sys/devices/system/node/node1", F_OK)) {
+ fputs("guest NUMA node 1 is unavailable\n", stderr);
+ goto fail;
+ }
+ fd = create_arena();
+ if (fd < 0) {
+ perror("create guest arena");
+ goto fail;
+ }
+ arena = mmap(NULL, getpagesize(), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
+ if (arena == MAP_FAILED) {
+ perror("mmap guest arena");
+ goto fail;
+ }
+ obj = bpf_object__open_file(argv[1], NULL);
+ if (!obj)
+ goto fail;
+ err = arena_kvm_check_features(obj);
+ if (err == 4) {
+ puts("GUEST_BPF_SKIPPED: compiler lacks arena address-space casts");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ close(fd);
+ return 4;
+ }
+ if (err)
+ goto fail;
+ map = bpf_object__find_map_by_name(obj, "arena");
+ alloc = bpf_object__find_program_by_name(obj, "allocate");
+ probe_free = bpf_object__find_program_by_name(obj, "probe_free");
+ exchange = bpf_object__find_program_by_name(obj, "exchange");
+ complete = bpf_object__find_program_by_name(obj, "complete");
+ if (!map || !alloc || !probe_free || !exchange || !complete ||
+ bpf_map__reuse_fd(map, fd))
+ goto fail;
+ close(fd);
+ fd = -1;
+ err = bpf_object__load(obj);
+ if (err) {
+ fprintf(stderr, "load guest BPF: %d\n", err);
+ goto fail;
+ }
+ err = run_prog(bpf_program__fd(alloc), &result);
+ if (err || result) {
+ fprintf(stderr, "allocate guest page: err=%d result=%u\n",
+ err, result);
+ goto fail;
+ }
+ err = run_prog(bpf_program__fd(probe_free), &result);
+ if (err || result) {
+ fprintf(stderr, "guest page lifetime check: err=%d result=%u\n",
+ err, result);
+ goto fail;
+ }
+ err = signal_file_offset(arena, &file_offset);
+ if (err == -EXDEV) {
+ puts("GUEST_BPF_SKIPPED: allocation fell back from shared NUMA node");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ return 4;
+ }
+ if (err) {
+ fprintf(stderr, "resolve shared page offset: %d\n", err);
+ goto fail;
+ }
+ printf("GUEST_BPF_READY %llu %d\n",
+ (unsigned long long)file_offset, getpagesize());
+ if (exchange_once(bpf_program__fd(exchange)) ||
+ exchange_once(bpf_program__fd(exchange)) ||
+ exchange_once(bpf_program__fd(complete)))
+ goto fail;
+ puts("GUEST_BPF_EXCHANGED");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ return 0;
+
+fail:
+ puts("GUEST_BPF_FAILED");
+ bpf_object__close(obj);
+ if (arena != MAP_FAILED)
+ munmap(arena, getpagesize());
+ if (fd >= 0)
+ close(fd);
+ return 1;
+}
diff --git a/tools/testing/selftests/bpf/arena_kvm_host.bpf.c b/tools/testing/selftests/bpf/arena_kvm_host.bpf.c
new file mode 100644
index 0000000000000..9c1c535fe429b
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_host.bpf.c
@@ -0,0 +1,89 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+#include "bpf_arena_common.h"
+#ifdef __TARGET_ARCH_powerpc
+/* PowerPC does not support arena load-acquire/store-release instructions. */
+#undef __BPF_FEATURE_LOAD_ACQ_STORE_REL
+#endif
+#include "bpf_atomic.h"
+#endif
+#include "arena_kvm_shared.h"
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARENA);
+ __uint(map_flags, BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE | BPF_F_ARENA_EXPORT);
+ /* The host runner reuses the map created with the selected capacity. */
+ __uint(max_entries, 1);
+} arena SEC(".maps");
+
+volatile __u32 signal_offset;
+
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+
+SEC("syscall")
+int exchange(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = page->h2g_seq;
+
+ if (smp_load_acquire(&page->ready) != GUEST_READY)
+ return 4;
+ if (!seq) {
+ page->h2g_payload = HOST_FIRST;
+ smp_store_release(&page->h2g_seq, 1);
+ return 0;
+ }
+ if (seq == 1 && smp_load_acquire(&page->g2h_seq) == 1) {
+ if (page->g2h_payload != GUEST_FIRST)
+ return 3;
+ page->h2g_payload = HOST_SECOND;
+ smp_store_release(&page->h2g_seq, 2);
+ return 0;
+ }
+ if (seq == 2 && smp_load_acquire(&page->g2h_seq) == 2) {
+ if (page->g2h_payload != GUEST_SECOND)
+ return 3;
+ smp_store_release(&page->h2g_seq, 3);
+ return 2;
+ }
+ return 1;
+}
+
+SEC("syscall")
+int check_reply(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = page->h2g_seq;
+
+ if ((seq != 1 && seq != 2) || smp_load_acquire(&page->g2h_seq) != seq)
+ return 1;
+ return page->g2h_payload == (seq == 1 ? GUEST_FIRST : GUEST_SECOND) ? 0 : 3;
+}
+
+SEC("syscall")
+int probe_free(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ bpf_arena_free_pages(&arena, page, 1);
+ return page->ready == GUEST_READY ? 0 : 1;
+}
+
+#else
+SEC("syscall") int exchange(void *ctx) { return 1; }
+SEC("syscall") int check_reply(void *ctx) { return 1; }
+SEC("syscall") int probe_free(void *ctx) { return 1; }
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/arena_kvm_host.c b/tools/testing/selftests/bpf/arena_kvm_host.c
new file mode 100644
index 0000000000000..90a77a54bd6f4
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_host.c
@@ -0,0 +1,147 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <bpf/bpf.h>
+#include <bpf/libbpf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <linux/kvm.h>
+#include <limits.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/ioctl.h>
+#include <sys/socket.h>
+#include <unistd.h>
+
+#include "arena_kvm_shared.h"
+
+static int create_arena(const char *pin, int socket_fd, unsigned int pages)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT);
+ char control[CMSG_SPACE(sizeof(int))] = {};
+ char byte = 0;
+ struct iovec iov = { .iov_base = &byte, .iov_len = 1 };
+ struct msghdr msg = {
+ .msg_iov = &iov,
+ .msg_iovlen = 1,
+ .msg_control = control,
+ .msg_controllen = sizeof(control),
+ };
+ struct cmsghdr *cmsg = CMSG_FIRSTHDR(&msg);
+ int fd, err = 1;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_kvm", 0, 0,
+ pages, &opts);
+ if (fd < 0) {
+ if (errno == EOPNOTSUPP || errno == EINVAL)
+ return 4; /* KSFT_SKIP: arena type or flags unsupported. */
+ perror("create host arena");
+ return 1;
+ }
+ if (bpf_obj_pin(fd, pin)) {
+ perror("pin host arena");
+ goto out;
+ }
+ cmsg->cmsg_level = SOL_SOCKET;
+ cmsg->cmsg_type = SCM_RIGHTS;
+ cmsg->cmsg_len = CMSG_LEN(sizeof(fd));
+ memcpy(CMSG_DATA(cmsg), &fd, sizeof(fd));
+ if (sendmsg(socket_fd, &msg, 0) == 1)
+ err = 0;
+ else
+ perror("send arena FD");
+out:
+ close(fd);
+ return err;
+}
+
+int main(int argc, char **argv)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ struct bpf_object *obj = NULL;
+ struct bpf_program *prog;
+ struct bpf_map *arena, *bss;
+ uint32_t key = 0, offset;
+ unsigned long parsed;
+ int fd, err;
+ char *end;
+
+ if (argc == 3 && !strcmp(argv[1], "--check-bpf")) {
+ obj = bpf_object__open_file(argv[2], NULL);
+ if (!obj)
+ return 1;
+ err = arena_kvm_check_features(obj);
+ bpf_object__close(obj);
+ return err;
+ }
+
+ if (argc == 2 && !strcmp(argv[1], "--check-kvm")) {
+ fd = open("/dev/kvm", O_RDWR);
+ if (fd < 0)
+ return 4;
+ err = ioctl(fd, KVM_GET_API_VERSION, 0);
+ close(fd);
+ return err == KVM_API_VERSION ? 0 : 4;
+ }
+ if (argc == 5 && !strcmp(argv[1], "--create")) {
+ errno = 0;
+ parsed = strtoul(argv[3], &end, 0);
+ if (errno || *end || end == argv[3] || parsed > INT_MAX)
+ return 1;
+ fd = parsed;
+ errno = 0;
+ parsed = strtoul(argv[4], &end, 0);
+ if (errno || *end || end == argv[4] || !parsed || parsed > UINT_MAX)
+ return 1;
+ return create_arena(argv[2], fd, parsed);
+ }
+ if (argc != 5)
+ return 1;
+ errno = 0;
+ parsed = strtoul(argv[1], &end, 0);
+ if (errno || *end || parsed > INT_MAX)
+ return 1;
+ fd = dup(parsed);
+ if (fd < 0)
+ return 1;
+ errno = 0;
+ parsed = strtoul(argv[3], &end, 0);
+ if (errno || *end || parsed > UINT_MAX || parsed % getpagesize())
+ goto fail;
+ offset = parsed;
+ obj = bpf_object__open_file(argv[2], NULL);
+ if (!obj)
+ goto fail;
+ arena = bpf_object__find_map_by_name(obj, "arena");
+ bss = bpf_object__find_map_by_name(obj, ".bss");
+ prog = bpf_object__find_program_by_name(obj, argv[4]);
+ if (!arena || !bss || !prog || bpf_map__reuse_fd(arena, fd))
+ goto fail;
+ if (offset >= (uint64_t)bpf_map__max_entries(arena) * getpagesize())
+ goto fail;
+ if (bpf_object__load(obj)) {
+ fputs("load host BPF failed\n", stderr);
+ goto fail;
+ }
+ if (bpf_map_update_elem(bpf_map__fd(bss), &key, &offset, BPF_ANY)) {
+ perror("set signal offset");
+ goto fail;
+ }
+ err = bpf_prog_test_run_opts(bpf_program__fd(prog), &opts);
+ if (err)
+ goto fail;
+ printf("%u\n", opts.retval);
+ bpf_object__close(obj);
+ close(fd);
+ return 0;
+
+fail:
+ bpf_object__close(obj);
+ close(fd);
+ return 1;
+}
diff --git a/tools/testing/selftests/bpf/arena_kvm_shared.h b/tools/testing/selftests/bpf/arena_kvm_shared.h
new file mode 100644
index 0000000000000..63aafac25efc3
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_shared.h
@@ -0,0 +1,60 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#ifndef ARENA_KVM_SHARED_H
+#define ARENA_KVM_SHARED_H
+
+/* These flags may be absent from a pre-series kernel's vmlinux.h. */
+#ifndef BPF_F_ARENA_NO_FREE
+#define BPF_F_ARENA_NO_FREE (1U << 20)
+#endif
+#ifndef BPF_F_ARENA_EXPORT
+#define BPF_F_ARENA_EXPORT (1U << 21)
+#endif
+
+struct arena_kvm_features {
+ unsigned int addr_space_cast;
+};
+
+#ifdef __BPF__
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+const volatile struct arena_kvm_features features = { .addr_space_cast = 1 };
+#else
+const volatile struct arena_kvm_features features = {};
+#endif
+#else
+/* Inspect compiler support before loading maps or booting either guest. */
+static inline int arena_kvm_check_features(struct bpf_object *obj)
+{
+ const struct arena_kvm_features *features;
+ struct bpf_map *map;
+ size_t size;
+
+ map = bpf_object__find_map_by_name(obj, ".rodata");
+ if (!map)
+ return 1;
+ features = bpf_map__initial_value(map, &size);
+ if (!features || size != sizeof(*features))
+ return 1;
+ return features->addr_space_cast ? 0 : 4; /* KSFT_SKIP */
+}
+#endif
+
+#define ARENA_KVM_NODE 1
+
+#define GUEST_READY 0x4152454e414b564dULL
+#define HOST_FIRST 0x123456789abcdef0ULL
+#define GUEST_FIRST 0xfedcba9876543210ULL
+#define HOST_SECOND 0x1020304050607080ULL
+#define GUEST_SECOND 0x8070605040302010ULL
+
+struct signal_page {
+ volatile unsigned long long ready;
+ volatile unsigned long long h2g_payload;
+ volatile unsigned long long h2g_seq;
+ volatile unsigned long long g2h_payload;
+ volatile unsigned long long g2h_seq;
+};
+
+#endif
diff --git a/tools/testing/selftests/bpf/prog_tests/arena_pinned.c b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c
new file mode 100644
index 0000000000000..8b7ca90211296
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c
@@ -0,0 +1,442 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <test_progs.h>
+#include <fcntl.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <sys/wait.h>
+
+#define ARENA_PAGES 4
+
+struct arena_file {
+ int map_fd;
+ int file_fd;
+ char pin[128];
+ void *canonical;
+ size_t len;
+ bool pinned;
+};
+
+static void cleanup(struct arena_file *file)
+{
+ if (file->canonical != MAP_FAILED)
+ munmap(file->canonical, file->len);
+ if (file->file_fd >= 0)
+ close(file->file_fd);
+ if (file->pinned)
+ unlink(file->pin);
+ if (file->map_fd >= 0)
+ close(file->map_fd);
+}
+
+static void reject_mapping(int fd, size_t len, int flags, off_t offset,
+ const char *name)
+{
+ void *addr;
+ int err;
+
+ addr = mmap(NULL, len, PROT_READ | PROT_WRITE, flags, fd, offset);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, name))
+ ASSERT_EQ(err, EINVAL, "mmap_errno");
+ else
+ munmap(addr, len);
+}
+
+static bool setup(struct arena_file *file, __u64 map_extra, bool zero)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT,
+ .map_extra = map_extra);
+
+ memset(file, 0, sizeof(*file));
+ file->map_fd = -1;
+ file->file_fd = -1;
+ file->canonical = MAP_FAILED;
+ file->len = ARENA_PAGES * getpagesize();
+ snprintf(file->pin, sizeof(file->pin), "/sys/fs/bpf/arena_pinned_%d", getpid());
+ file->map_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_pinned",
+ 0, 0, ARENA_PAGES, &opts);
+ if (file->map_fd < 0 && errno == EOPNOTSUPP) {
+ test__skip();
+ return false;
+ }
+ if (!ASSERT_GE(file->map_fd, 0, "map_create"))
+ return false;
+ if (!ASSERT_OK(bpf_obj_pin(file->map_fd, file->pin), "obj_pin"))
+ return false;
+ file->pinned = true;
+ file->file_fd = open(file->pin, O_RDWR);
+ if (!ASSERT_GE(file->file_fd, 0, "open_pin"))
+ return false;
+ reject_mapping(file->map_fd, getpagesize(), MAP_SHARED, 0, "short_canonical");
+ reject_mapping(file->map_fd, file->len + getpagesize(), MAP_SHARED,
+ 0, "large_canonical");
+ if (map_extra)
+ return true;
+
+ reject_mapping(file->file_fd, getpagesize(), MAP_SHARED, 0, "uninitialized");
+ file->canonical = mmap(NULL, file->len, PROT_READ | PROT_WRITE,
+ MAP_SHARED | (zero ? MAP_FIXED_NOREPLACE : 0),
+ file->map_fd, 0);
+ if (zero && file->canonical == MAP_FAILED &&
+ (errno == EPERM || errno == EACCES)) {
+ test__skip();
+ return false;
+ }
+ return ASSERT_NEQ(file->canonical, MAP_FAILED, "canonical_mmap");
+}
+
+static void test_mappings(__u64 map_extra)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ struct stat st;
+ char *alias = MAP_FAILED;
+ char *full = MAP_FAILED;
+ int expected = 91;
+
+ if (!setup(&file, map_extra, false))
+ goto out;
+ if (!ASSERT_OK(fstat(file.file_fd, &st), "stat_pin"))
+ goto out;
+ ASSERT_EQ(st.st_size, ARENA_PAGES * ps, "pin_size");
+ reject_mapping(file.file_fd, 8ULL << 30, MAP_SHARED, 0, "oversized");
+ reject_mapping(file.file_fd, ps, MAP_SHARED, ARENA_PAGES * ps, "past_capacity");
+ reject_mapping(file.file_fd, ps * 2, MAP_SHARED,
+ (ARENA_PAGES - 1) * ps, "past_capacity_end");
+ reject_mapping(file.file_fd, ps, MAP_PRIVATE, 0, "private_mapping");
+ full = mmap(NULL, st.st_size, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, 0);
+ if (!ASSERT_NEQ(full, MAP_FAILED, "full_file_mmap"))
+ goto out;
+ full[0] = 17;
+ full[file.len - 1] = 62;
+
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, (ARENA_PAGES - 1) * ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "last_page_slice"))
+ goto out;
+ alias[0] = 91;
+ ASSERT_OK(msync(alias, ps, MS_SYNC), "msync_slice");
+ ASSERT_OK(fsync(file.file_fd), "fsync_pin");
+ ASSERT_OK(fdatasync(file.file_fd), "fdatasync_pin");
+ ASSERT_EQ(alias[ps - 1], 62, "full_file_last_page");
+ if (file.canonical != MAP_FAILED) {
+ char *canonical = file.canonical;
+
+ ASSERT_EQ(canonical[0], 17, "full_file_first_page");
+ ASSERT_EQ(canonical[(ARENA_PAGES - 1) * ps], 91, "alias_write");
+ canonical[(ARENA_PAGES - 1) * ps] = 42;
+ ASSERT_EQ(alias[0], 42, "canonical_write");
+ expected = 42;
+ if (!ASSERT_OK(munmap(file.canonical, file.len), "unmap_canonical"))
+ goto out;
+ file.canonical = MAP_FAILED;
+ }
+ if (!ASSERT_OK(munmap(full, file.len), "unmap_full_file"))
+ goto out;
+ full = MAP_FAILED;
+ if (!ASSERT_OK(unlink(file.pin), "unlink_pin"))
+ goto out;
+ file.pinned = false;
+ close(file.file_fd);
+ file.file_fd = -1;
+ close(file.map_fd);
+ file.map_fd = -1;
+ ASSERT_EQ(alias[0], expected, "unpin_lifetime");
+ alias[0] = 73;
+out:
+ if (full != MAP_FAILED)
+ munmap(full, file.len);
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_unexported_pin(__u32 flags)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | flags);
+ char pin[128];
+ struct stat st;
+ char *canonical = MAP_FAILED;
+ size_t ps = getpagesize();
+ int map_fd, fd, err;
+
+ map_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_unexported",
+ 0, 0, ARENA_PAGES, &opts);
+ if (map_fd < 0 && errno == EOPNOTSUPP) {
+ test__skip();
+ return;
+ }
+ if (!ASSERT_GE(map_fd, 0, "unexported_create"))
+ return;
+ canonical = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, map_fd, 0);
+ if (!ASSERT_NEQ(canonical, MAP_FAILED, "unexported_short_mmap"))
+ goto out;
+ canonical[0] = 42;
+ ASSERT_EQ(canonical[0], 42, "unexported_short_access");
+ snprintf(pin, sizeof(pin), "/sys/fs/bpf/arena_unexported_%d", getpid());
+ if (!ASSERT_OK(bpf_obj_pin(map_fd, pin), "unexported_pin"))
+ goto out;
+ if (ASSERT_OK(stat(pin, &st), "unexported_stat"))
+ ASSERT_EQ(st.st_size, 0, "unexported_size");
+ fd = open(pin, O_RDONLY);
+ err = errno;
+ if (ASSERT_EQ(fd, -1, "unexported_open"))
+ ASSERT_EQ(err, EIO, "unexported_open_errno");
+ else
+ close(fd);
+ fd = bpf_obj_get(pin);
+ if (ASSERT_GE(fd, 0, "unexported_obj_get"))
+ close(fd);
+ ASSERT_OK(unlink(pin), "unexported_unlink");
+out:
+ if (canonical != MAP_FAILED)
+ munmap(canonical, ps);
+ close(map_fd);
+}
+
+static void test_export_requires_no_free(void)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_EXPORT);
+ int fd;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_export",
+ 0, 0, ARENA_PAGES, &opts);
+ if (fd == -EOPNOTSUPP) {
+ test__skip();
+ return;
+ }
+ ASSERT_EQ(fd, -EINVAL, "export_requires_no_free");
+ if (fd >= 0)
+ close(fd);
+}
+
+static void test_read_only(int flags)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *addr;
+ int fd = -1, err, ret;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ if (!ASSERT_OK(fchmod(file.file_fd, 0400), "chmod_read_only"))
+ goto out;
+ fd = open(file.pin, O_RDONLY);
+ if (!ASSERT_GE(fd, 0, "open_read_only"))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ, flags, fd, ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "read_only_mmap"))
+ goto out;
+ ASSERT_EQ(alias[0], 0, "read_only_fault");
+ ((char *)file.canonical)[ps] = 42;
+ ASSERT_EQ(alias[0], 42, "read_only_shared_update");
+
+ addr = mmap(NULL, ps, PROT_READ | PROT_WRITE, flags, fd, ps);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, "read_only_writable_mmap"))
+ ASSERT_EQ(err, EACCES, "read_only_writable_errno");
+ else
+ munmap(addr, ps);
+ ret = mprotect(alias, ps, PROT_READ | PROT_WRITE);
+ err = errno;
+ ASSERT_EQ(ret, -1, "read_only_mprotect");
+ ASSERT_EQ(err, EACCES, "read_only_mprotect_errno");
+ addr = mmap(NULL, ps, PROT_READ, MAP_PRIVATE, fd, ps);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, "read_only_private_mmap"))
+ ASSERT_EQ(err, EINVAL, "read_only_private_errno");
+ else
+ munmap(addr, ps);
+
+ close(fd);
+ fd = -1;
+ ((char *)file.canonical)[ps] = 73;
+ ASSERT_EQ(alias[0], 73, "read_only_after_close");
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ if (fd >= 0)
+ close(fd);
+ cleanup(&file);
+}
+
+static void test_fixed_size(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ struct stat st;
+ unsigned char resident;
+ char *alias = MAP_FAILED;
+ int err, fd, ret;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, file.file_fd, 0);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "alias_mmap"))
+ goto out;
+ alias[0] = 42;
+ ret = ftruncate(file.file_fd, 0);
+ err = errno;
+ ASSERT_EQ(ret, -1, "ftruncate_shrink");
+ ASSERT_EQ(err, EINVAL, "ftruncate_shrink_errno");
+ ret = ftruncate(file.file_fd, file.len + ps);
+ err = errno;
+ ASSERT_EQ(ret, -1, "ftruncate_grow");
+ ASSERT_EQ(err, EINVAL, "ftruncate_grow_errno");
+ ret = truncate(file.pin, 0);
+ err = errno;
+ ASSERT_EQ(ret, -1, "truncate_pin");
+ ASSERT_EQ(err, EINVAL, "truncate_pin_errno");
+ fd = open(file.pin, O_RDWR | O_TRUNC);
+ err = errno;
+ if (ASSERT_EQ(fd, -1, "open_trunc"))
+ ASSERT_EQ(err, EINVAL, "open_trunc_errno");
+ else
+ close(fd);
+ ASSERT_OK(ftruncate(file.file_fd, file.len), "ftruncate_same_size");
+ if (!ASSERT_OK(mincore(alias, ps, &resident), "mincore"))
+ goto out;
+ ASSERT_EQ(resident & 1, 1, "mapping_not_zapped");
+ ASSERT_EQ(alias[0], 42, "mapping_value");
+ ASSERT_OK(fchmod(file.file_fd, 0640), "chmod_pin");
+ if (ASSERT_OK(fstat(file.file_fd, &st), "stat_pin")) {
+ ASSERT_EQ(st.st_size, file.len, "fixed_size");
+ ASSERT_EQ(st.st_mode & 0777, 0640, "pin_mode");
+ }
+ ASSERT_EQ(lseek(file.file_fd, 0, SEEK_END), file.len, "seek_capacity");
+ ASSERT_EQ(lseek(file.file_fd, -(off_t)ps, SEEK_CUR), file.len - ps, "seek_back");
+ ASSERT_EQ(lseek(file.file_fd, ps, SEEK_SET), ps, "seek_offset");
+ errno = 0;
+ ASSERT_EQ(lseek(file.file_fd, file.len + 1, SEEK_SET), -1, "seek_past_end");
+ ASSERT_EQ(errno, EINVAL, "seek_past_end_errno");
+ errno = 0;
+ ASSERT_EQ(lseek(file.file_fd, -1, SEEK_SET), -1, "seek_before_start");
+ ASSERT_EQ(errno, EINVAL, "seek_before_start_errno");
+ fd = bpf_obj_get(file.pin);
+ if (ASSERT_GE(fd, 0, "obj_get_after_chmod"))
+ close(fd);
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_zero_canonical(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *addr;
+
+ if (!setup(&file, 0, true))
+ goto out;
+ if (!ASSERT_EQ((unsigned long)file.canonical, 0, "zero_canonical"))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, file.file_fd, ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "zero_canonical_export"))
+ goto out;
+ alias[0] = 42;
+ ASSERT_EQ(alias[0], 42, "zero_canonical_fault");
+ addr = mmap((void *)ps, file.len, PROT_READ | PROT_WRITE,
+ MAP_SHARED | MAP_FIXED_NOREPLACE, file.map_fd, 0);
+ if (ASSERT_EQ(addr, MAP_FAILED, "canonical_cannot_move"))
+ ASSERT_EQ(errno, EINVAL, "canonical_cannot_move_errno");
+ else
+ munmap(addr, file.len);
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_lifecycle(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *target = MAP_FAILED, *addr;
+ pid_t child;
+ int status, ret, err;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ alias = mmap(NULL, ps * 2, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, 0);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "lifecycle_mmap"))
+ goto out;
+ alias[0] = 42;
+ ret = munmap(alias, ps);
+ err = errno;
+ ASSERT_EQ(ret, -1, "partial_unmap");
+ ASSERT_EQ(err, EINVAL, "partial_unmap_errno");
+ ret = mprotect(alias, ps, PROT_READ);
+ err = errno;
+ ASSERT_EQ(ret, -1, "partial_mprotect");
+ ASSERT_EQ(err, EINVAL, "partial_mprotect_errno");
+ ret = madvise(alias, ps * 2, MADV_DOFORK);
+ err = errno;
+ ASSERT_EQ(ret, -1, "enable_inheritance");
+ ASSERT_EQ(err, EINVAL, "enable_inheritance_errno");
+ target = mmap(NULL, ps * 2, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+ if (!ASSERT_NEQ(target, MAP_FAILED, "reserve_remap_target"))
+ goto out;
+ addr = mremap(alias, ps * 2, ps * 2, MREMAP_MAYMOVE | MREMAP_FIXED, target);
+ err = errno;
+ if (!ASSERT_EQ(addr, MAP_FAILED, "relocate_mapping")) {
+ alias = addr;
+ target = MAP_FAILED;
+ }
+ ASSERT_EQ(err, EINVAL, "relocate_mapping_errno");
+ child = fork();
+ if (!ASSERT_GE(child, 0, "fork"))
+ goto out;
+ if (!child) {
+ unsigned char resident[2];
+
+ ret = mincore(alias, ps * 2, resident);
+ _exit(ret == -1 && errno == ENOMEM ? 0 : 1);
+ }
+ if (ASSERT_EQ(waitpid(child, &status, 0), child, "wait_child") &&
+ ASSERT_TRUE(WIFEXITED(status), "child_exited"))
+ ASSERT_EQ(WEXITSTATUS(status), 0, "mapping_not_inherited");
+ ASSERT_EQ(alias[0], 42, "parent_mapping_retained");
+out:
+ if (target != MAP_FAILED)
+ munmap(target, ps * 2);
+ if (alias != MAP_FAILED)
+ munmap(alias, ps * 2);
+ cleanup(&file);
+}
+
+void serial_test_arena_pinned(void)
+{
+ if (test__start_subtest("ordinary_pin"))
+ test_unexported_pin(0);
+ if (test__start_subtest("retained_pin"))
+ test_unexported_pin(BPF_F_ARENA_NO_FREE);
+ if (test__start_subtest("export_requires_no_free"))
+ test_export_requires_no_free();
+ if (test__start_subtest("mappings"))
+ test_mappings(0);
+ if (test__start_subtest("map_extra"))
+ test_mappings(1ULL << 32);
+ if (test__start_subtest("fixed_size"))
+ test_fixed_size();
+ if (test__start_subtest("read_only"))
+ test_read_only(MAP_SHARED);
+ if (test__start_subtest("read_only_validate"))
+ test_read_only(MAP_SHARED_VALIDATE);
+ if (test__start_subtest("zero_canonical"))
+ test_zero_canonical();
+ if (test__start_subtest("lifecycle"))
+ test_lifecycle();
+}
|
| 3/2 | 2026/10/10 23:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 30641 output: 34 thoughts: 936 Thoughts: **Fuzzing Evaluation: BPF Arena Patch Series** Okay, let's break this down. My initial assessment is that this patch series is likely a good candidate for syzkaller fuzzing. The goal is to determine if it warrants fuzzing, right? Let's see. First, I see documentation changes, header updates in `bpf.h`, and, most importantly, core code modifications within the `kernel/bpf/` and `kernel/` directories. Specifically, I'm looking at `arena.c` and `inode.c`. That gets my attention immediately. This patch introduces BPF arena maps, which are potentially shareable/pinable BPF maps. The code deals with arena map allocation, mmaping, memory fault handling (`arena_vm_fault`), and bpffs interaction – all in the BPF core subsystem and filesystem interactions, which are critical areas for security. The patch also adds functionality for exporting arenas via bpffs and incorporates checks related to memory management, which, of course, is a high-value target. I like the use of flags like `BPF_F_ARENA_NO_FREE` and `BPF_F_ARENA_EXPORT`, and the associated logic. The code seems reachable; it involves userspace interaction via BPF system calls, creating and pinning the maps, opening files in the bpffs filesystem and memory mapping (mmaping), and triggering faults. These operations are standard and available in standard virtualized environments, so there aren't any hardware dependencies or special configurations. The code changes are not merely configuration options or trivial changes. The code changes seem to warrant some investigation. I'm thinking about the criteria. "WorthFuzzing=true" if the patch touches reachable executable code. And this definitely does: BPF, VFS (bpffs), and core memory management. This patch is definitely touching core subsystems, specifically BPF and, by extension, VFS, because of the introduction of bpffs and the memory mapping code, as well as modifications to how the memory is allocated. Therefore, based on that criteria, this should be fuzzed. Now, a quick configuration check. I should look at the kernel source's `Makefile` (I'll use grepper for this) to identify the specific configurations needed to enable the affected code paths (bpffs, arena maps). That will assist in setting up a reproducible fuzzing environment. |
| 4/2 | 2026/10/10 23:15 | tool | grepper |
0mArgs:{
"Expression": "arena\\.o|inode\\.o",
"PathPrefixes": [
"kernel/bpf/Makefile"
]
}Results: {
"Output": "kernel/bpf/Makefile=7=CFLAGS_core.o += -Wno-override-init $(cflags-nogcse-yy)\nkernel/bpf/Makefile-8-\nkernel/bpf/Makefile:9:obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o tnum.o cnum.o log.o token.o liveness.o const_fold.o diagnostics.o loops.o scev.o\nkernel/bpf/Makefile-10-obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_iter.o\n--\nkernel/bpf/Makefile=19=ifeq ($(CONFIG_MMU)$(CONFIG_64BIT),yy)\nkernel/bpf/Makefile:20:obj-$(CONFIG_BPF_SYSCALL) += arena.o range_tree.o\nkernel/bpf/Makefile-21-endif\n"
} |
| 5/2 | 2026/10/10 23:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32372 output: 42 thoughts: 90 Thoughts: Okay, here's my summary, considering my expertise: **Checking BPF Arena Requirements** Right, so `CONFIG_BPF_SYSCALL` is definitely a prerequisite, that's clear. Now, I need to dig deeper into the arena component. The key question is: does `arena` have any other, potentially subtle, configuration dependencies beyond `BPF_SYSCALL`? That `ifeq` statement in the Makefile is immediately helpful: `ifeq ($(CONFIG_MMU)$(CONFIG_64BIT),yy) obj-$(CONFIG_BPF_SYSCALL) += arena.o range_tree.o`. This tells me that MMU and 64-bit support are important for building *and* that those are *required* for the arena to be built alongside the BPF syscall when BPF is enabled! I need to ensure those are taken into account. I should also cross-reference this with the Kconfig files, and anywhere else `arena` is specified. I need to make sure I don't miss anything that would break the arena feature. I'll meticulously check for other dependencies, and any additional requirements that relate to the arena component. |
| 6/2 | 2026/10/10 23:15 | tool | grepper |
0mArgs:{
"Expression": "config BPF_SYSCALL",
"PathPrefixes": [
"init/Kconfig",
"kernel/bpf/Kconfig"
]
}Results: {
"Output": "kernel/bpf/Kconfig=29=menu \"BPF subsystem\"\nkernel/bpf/Kconfig-30-\nkernel/bpf/Kconfig:31:config BPF_SYSCALL\nkernel/bpf/Kconfig-32-\tbool \"Enable bpf() system call\"\n"
} |
| 7/2 | 2026/10/10 23:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32838 output: 34 thoughts: 30 Thoughts: Here's my thought process, as if I were considering this myself: **Investigating BPF_SYSCALL Dependency** Okay, so I need to figure out the dependencies for `BPF_SYSCALL`. I'm pretty sure that's controlled via a Kconfig option, so the logical place to start is `kernel/bpf/Kconfig`. Let me go ahead and open that file up... Right, now, I seem to remember that being defined around line 31, give or take. Let's head over there and take a look. I'm expecting to see what other options this relies on, or any conditions that have to be met before it is selectable. Time to find out. |
| 8/2 | 2026/10/10 23:15 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/Kconfig",
"FirstLine": 30,
"LineCount": 25
}Results: {
"Contents": " 30:\t\n 31:\tconfig BPF_SYSCALL\n 32:\t\tbool \"Enable bpf() system call\"\n 33:\t\tselect BPF\n 34:\t\tselect IRQ_WORK\n 35:\t\tselect NEED_TASKS_RCU\n 36:\t\tselect TASKS_TRACE_RCU\n 37:\t\tselect BINARY_PRINTF\n 38:\t\tselect NET_SOCK_MSG if NET\n 39:\t\tselect NET_XGRESS if NET\n 40:\t\tselect PAGE_POOL if NET\n 41:\t\tdefault n\n 42:\t\thelp\n 43:\t\t Enable the bpf() system call that allows to manipulate BPF programs\n 44:\t\t and maps via file descriptors.\n 45:\t\n 46:\tconfig BPF_JIT\n 47:\t\tbool \"Enable BPF Just In Time compiler\"\n 48:\t\tdepends on BPF\n 49:\t\tdepends on HAVE_CBPF_JIT || HAVE_EBPF_JIT\n 50:\t\tselect EXECMEM\n 51:\t\thelp\n 52:\t\t BPF programs are normally handled by a BPF interpreter. This option\n 53:\t\t allows the kernel to generate native code when a program is loaded\n 54:\t\t into the kernel. This will significantly speed-up processing of BPF\n"
} |
| 9/2 | 2026/10/10 23:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 33395 output: 44 thoughts: 83 Thoughts: Here's my thought process, as requested: **Checking BPF System Call Dependencies** Okay, so I'm thinking about BPF and specifically the `BPF_SYSCALL` functionality. I know that's enabled by the `CONFIG_BPF_SYSCALL` kernel configuration option. That's pretty straightforward. And, as a matter of course, `BPF_JIT` is almost always enabled by default, but `BPF_SYSCALL` is the critical piece here for this to function. Now, I need to check something else. What's the relationship between the BPF system call and `bpffs`? Does `bpffs` depend on `BPF_SYSCALL`, or is it independent and managed by another config option? I can't quite recall. To nail this down and confirm the dependency relationship, I'll need to look at the source code. Let's start by checking the `kernel/bpf/inode.c` file. That seems like the logical place where `bpffs` interactions with BPF would be handled. From there, I should be able to clarify the situation quickly. |
| 10/2 | 2026/10/10 23:15 | tool | grepper |
0mArgs:{
"Expression": "bpffs|bpf_fs|bpf_init_fs",
"PathPrefixes": [
"kernel/bpf/inode.c"
]
}Results: {
"Output": "\nFull output is too long, showing 200 out of 330 lines.\nUse more precise expression or PathPrefixes if possible.\n\nNumber of matching lines per file (1 files in total):\nkernel/bpf/inode.c:89\n\nkernel/bpf/inode.c=29=enum bpf_type {\n--\nkernel/bpf/inode.c-35-\nkernel/bpf/inode.c:36:struct bpf_fs_inode {\nkernel/bpf/inode.c-37-\tstruct list_head\t\txattrs;\n--\nkernel/bpf/inode.c-41-\nkernel/bpf/inode.c:42:static inline struct bpf_fs_inode *BPF_FS_I(struct inode *inode)\nkernel/bpf/inode.c-43-{\nkernel/bpf/inode.c:44:\treturn container_of(inode, struct bpf_fs_inode, vfs_inode);\nkernel/bpf/inode.c-45-}\nkernel/bpf/inode.c-46-\nkernel/bpf/inode.c:47:static struct kmem_cache *bpf_fs_inode_cachep __ro_after_init;\nkernel/bpf/inode.c-48-\nkernel/bpf/inode.c:49:static int bpf_fs_initxattrs(struct inode *inode,\nkernel/bpf/inode.c-50-\t\t\t const struct xattr *xattr_array, void *fs_info);\nkernel/bpf/inode.c:51:static ssize_t bpf_fs_listxattr(struct dentry *dentry, char *buf, size_t size);\nkernel/bpf/inode.c-52-\n--\nkernel/bpf/inode.c=119=static const struct inode_operations bpf_prog_iops = {\nkernel/bpf/inode.c:120:\t.listxattr\t= bpf_fs_listxattr,\nkernel/bpf/inode.c-121-};\n--\nkernel/bpf/inode.c=137=static const struct inode_operations bpf_map_iops = {\nkernel/bpf/inode.c:138:\t.listxattr\t= bpf_fs_listxattr,\nkernel/bpf/inode.c-139-\t.setattr\t= bpf_map_setattr,\n--\nkernel/bpf/inode.c=141=static const struct inode_operations bpf_link_iops = {\nkernel/bpf/inode.c:142:\t.listxattr\t= bpf_fs_listxattr,\nkernel/bpf/inode.c-143-};\n--\nkernel/bpf/inode.c=195=static struct dentry *bpf_mkdir(struct mnt_idmap *idmap, struct inode *dir,\n--\nkernel/bpf/inode.c-205-\tret = security_inode_init_security(inode, dir, \u0026dentry-\u003ed_name,\nkernel/bpf/inode.c:206:\t\t\t\t\t bpf_fs_initxattrs, NULL);\nkernel/bpf/inode.c-207-\tif (ret \u0026\u0026 ret != -EOPNOTSUPP) {\n--\nkernel/bpf/inode.c=300=static int map_seq_show(struct seq_file *m, void *v)\n--\nkernel/bpf/inode.c-314-\nkernel/bpf/inode.c:315:static const struct seq_operations bpffs_map_seq_ops = {\nkernel/bpf/inode.c-316-\t.start\t= map_seq_start,\n--\nkernel/bpf/inode.c-321-\nkernel/bpf/inode.c:322:static int bpffs_map_open(struct inode *inode, struct file *file)\nkernel/bpf/inode.c-323-{\n--\nkernel/bpf/inode.c-332-\nkernel/bpf/inode.c:333:\terr = seq_open(file, \u0026bpffs_map_seq_ops);\nkernel/bpf/inode.c-334-\tif (err) {\n--\nkernel/bpf/inode.c-344-\nkernel/bpf/inode.c:345:static int bpffs_map_release(struct inode *inode, struct file *file)\nkernel/bpf/inode.c-346-{\n--\nkernel/bpf/inode.c-353-\nkernel/bpf/inode.c:354:/* bpffs_map_fops should only implement the basic\nkernel/bpf/inode.c-355- * read operation for a BPF map. The purpose is to\nkernel/bpf/inode.c-356- * provide a simple user intuitive way to do\nkernel/bpf/inode.c:357: * \"cat bpffs/pathto/a-pinned-map\".\nkernel/bpf/inode.c-358- *\n--\nkernel/bpf/inode.c-363- */\nkernel/bpf/inode.c:364:static const struct file_operations bpffs_map_fops = {\nkernel/bpf/inode.c:365:\t.open\t\t= bpffs_map_open,\nkernel/bpf/inode.c-366-\t.read\t\t= seq_read,\nkernel/bpf/inode.c:367:\t.release\t= bpffs_map_release,\nkernel/bpf/inode.c-368-};\n--\nkernel/bpf/inode.c-370-/*\nkernel/bpf/inode.c:371: * An arena with BPF_F_ARENA_EXPORT pinned in bpffs can serve as a\nkernel/bpf/inode.c-372- * shared-memory file. The ordinary map FD is an anonymous inode with no\n--\nkernel/bpf/inode.c-376- */\nkernel/bpf/inode.c:377:static int bpffs_arena_open(struct inode *inode, struct file *file)\nkernel/bpf/inode.c-378-{\n--\nkernel/bpf/inode.c-390-\nkernel/bpf/inode.c:391:static int bpffs_arena_mmap(struct file *file, struct vm_area_struct *vma)\nkernel/bpf/inode.c-392-{\n--\nkernel/bpf/inode.c-403-\nkernel/bpf/inode.c:404:static int bpffs_arena_release(struct inode *inode, struct file *file)\nkernel/bpf/inode.c-405-{\n--\nkernel/bpf/inode.c-408-\nkernel/bpf/inode.c:409:static unsigned long bpffs_arena_get_unmapped_area(struct file *file,\nkernel/bpf/inode.c-410-\t\t\t\t\t\t unsigned long addr,\n--\nkernel/bpf/inode.c-417-\nkernel/bpf/inode.c:418:static loff_t bpffs_arena_llseek(struct file *file, loff_t offset, int whence)\nkernel/bpf/inode.c-419-{\n--\nkernel/bpf/inode.c-422-\nkernel/bpf/inode.c:423:static const struct file_operations bpffs_arena_fops = {\nkernel/bpf/inode.c:424:\t.open\t\t= bpffs_arena_open,\nkernel/bpf/inode.c:425:\t.llseek\t\t= bpffs_arena_llseek,\nkernel/bpf/inode.c-426-\t.fsync\t\t= noop_fsync,\nkernel/bpf/inode.c:427:\t.release\t= bpffs_arena_release,\nkernel/bpf/inode.c:428:\t.mmap\t\t= bpffs_arena_mmap,\nkernel/bpf/inode.c:429:\t.get_unmapped_area = bpffs_arena_get_unmapped_area,\nkernel/bpf/inode.c-430-};\nkernel/bpf/inode.c-431-\nkernel/bpf/inode.c:432:static int bpffs_obj_open(struct inode *inode, struct file *file)\nkernel/bpf/inode.c-433-{\n--\nkernel/bpf/inode.c-436-\nkernel/bpf/inode.c:437:static const struct file_operations bpffs_obj_fops = {\nkernel/bpf/inode.c:438:\t.open\t\t= bpffs_obj_open,\nkernel/bpf/inode.c-439-};\n--\nkernel/bpf/inode.c=441=static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,\n--\nkernel/bpf/inode.c-453-\tret = security_inode_init_security(inode, dir, \u0026dentry-\u003ed_name,\nkernel/bpf/inode.c:454:\t\t\t\t\t bpf_fs_initxattrs, NULL);\nkernel/bpf/inode.c-455-\tif (ret \u0026\u0026 ret != -EOPNOTSUPP) {\n--\nkernel/bpf/inode.c=469=static int bpf_mkprog(struct dentry *dentry, umode_t mode, void *arg)\n--\nkernel/bpf/inode.c-471-\treturn bpf_mkobj_ops(dentry, mode, arg, \u0026bpf_prog_iops,\nkernel/bpf/inode.c:472:\t\t\t \u0026bpffs_obj_fops, 0);\nkernel/bpf/inode.c-473-}\n--\nkernel/bpf/inode.c=475=static int bpf_mkmap(struct dentry *dentry, umode_t mode, void *arg)\n--\nkernel/bpf/inode.c-481-\treturn bpf_mkobj_ops(dentry, mode, arg, \u0026bpf_map_iops,\nkernel/bpf/inode.c:482:\t\t\t shared_arena ? \u0026bpffs_arena_fops :\nkernel/bpf/inode.c-483-\t\t\t bpf_map_support_seq_show(map) ?\nkernel/bpf/inode.c:484:\t\t\t \u0026bpffs_map_fops : \u0026bpffs_obj_fops,\nkernel/bpf/inode.c-485-\t\t\t shared_arena ? (loff_t)map-\u003emax_entries * PAGE_SIZE : 0);\n--\nkernel/bpf/inode.c=488=static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)\n--\nkernel/bpf/inode.c-493-\t\t\t bpf_link_is_iter(link) ?\nkernel/bpf/inode.c:494:\t\t\t \u0026bpf_iter_fops : \u0026bpffs_obj_fops, 0);\nkernel/bpf/inode.c-495-}\n--\nkernel/bpf/inode.c=498=bpf_lookup(struct inode *dir, struct dentry *dentry, unsigned flags)\n--\nkernel/bpf/inode.c-500-\t/* Dots in names (e.g. \"/sys/fs/bpf/foo.bar\") are reserved for future\nkernel/bpf/inode.c:501:\t * extensions. That allows popoulate_bpffs() create special files.\nkernel/bpf/inode.c-502-\t */\n--\nkernel/bpf/inode.c=510=static int bpf_symlink(struct mnt_idmap *idmap, struct inode *dir,\n--\nkernel/bpf/inode.c-530-\tret = security_inode_init_security(inode, dir, \u0026dentry-\u003ed_name,\nkernel/bpf/inode.c:531:\t\t\t\t\t bpf_fs_initxattrs, NULL);\nkernel/bpf/inode.c-532-\tif (ret \u0026\u0026 ret != -EOPNOTSUPP) {\n--\nkernel/bpf/inode.c=541=static const struct inode_operations bpf_symlink_iops = {\nkernel/bpf/inode.c-542-\t.get_link\t= simple_get_link,\nkernel/bpf/inode.c:543:\t.listxattr\t= bpf_fs_listxattr,\nkernel/bpf/inode.c-544-};\n--\nkernel/bpf/inode.c=546=static const struct inode_operations bpf_dir_iops = {\n--\nkernel/bpf/inode.c-553-\t.unlink\t\t= simple_unlink,\nkernel/bpf/inode.c:554:\t.listxattr\t= bpf_fs_listxattr,\nkernel/bpf/inode.c-555-};\nkernel/bpf/inode.c-556-\nkernel/bpf/inode.c:557:/* pin iterator link into bpffs */\nkernel/bpf/inode.c-558-static int bpf_iter_link_pin_kernel(struct dentry *parent,\n--\nkernel/bpf/inode.c=734=EXPORT_SYMBOL(bpf_prog_get_type_path);\nkernel/bpf/inode.c-735-\nkernel/bpf/inode.c:736:struct bpffs_btf_enums {\nkernel/bpf/inode.c-737-\tconst struct btf *btf;\n--\nkernel/bpf/inode.c-743-\nkernel/bpf/inode.c:744:static int find_bpffs_btf_enums(struct bpffs_btf_enums *info)\nkernel/bpf/inode.c-745-{\n--\nkernel/bpf/inode.c=861=static int bpf_show_options(struct seq_file *m, struct dentry *root)\n--\nkernel/bpf/inode.c-878-\t opts-\u003edelegate_progs || opts-\u003edelegate_attachs) {\nkernel/bpf/inode.c:879:\t\tstruct bpffs_btf_enums info;\nkernel/bpf/inode.c-880-\nkernel/bpf/inode.c-881-\t\t/* ignore errors, fallback to hex */\nkernel/bpf/inode.c:882:\t\t(void)find_bpffs_btf_enums(\u0026info);\nkernel/bpf/inode.c-883-\n--\nkernel/bpf/inode.c-907-\nkernel/bpf/inode.c:908:static struct inode *bpf_fs_alloc_inode(struct super_block *sb)\nkernel/bpf/inode.c-909-{\nkernel/bpf/inode.c:910:\tstruct bpf_fs_inode *bi;\nkernel/bpf/inode.c-911-\nkernel/bpf/inode.c:912:\tbi = alloc_inode_sb(sb, bpf_fs_inode_cachep, GFP_KERNEL);\nkernel/bpf/inode.c-913-\tif (!bi)\n--\nkernel/bpf/inode.c=920=static void bpf_destroy_inode(struct inode *inode)\n--\nkernel/bpf/inode.c-922-\tstruct bpf_mount_opts *opts = inode-\u003ei_sb-\u003es_fs_info;\nkernel/bpf/inode.c:923:\tstruct bpf_fs_inode *bi = BPF_FS_I(inode);\nkernel/bpf/inode.c-924-\tenum bpf_type type;\n--\nkernel/bpf/inode.c=935=static void bpf_free_inode(struct inode *inode)\n--\nkernel/bpf/inode.c-938-\t\tkfree(inode-\u003ei_link);\nkernel/bpf/inode.c:939:\tkmem_cache_free(bpf_fs_inode_cachep, BPF_FS_I(inode));\nkernel/bpf/inode.c-940-}\n\n"
} |
| 11/2 | 2026/10/10 23:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37853 output: 37 thoughts: 195 Thoughts: Here's my analysis of the situation: **Investigating the Configuration Landscape of `bpffs` and the BPF Arena** Okay, so `bpffs` is clearly baked into the `kernel/bpf/inode.c` file. Makes sense. And that's gated by `CONFIG_BPF_SYSCALL`, straightforward enough. I see the object file inclusion in the `Makefile` there: `obj-$(CONFIG_BPF_SYSCALL) += syscall.o verifier.o inode.o helpers.o ...`. Now, the `kernel/bpf/arena.c` file is a little more interesting. It's also linked under `CONFIG_BPF_SYSCALL`, but with the added requirement of `CONFIG_MMU` and `CONFIG_64BIT`. The `ifeq` condition in the `Makefile` clearly shows this: `ifeq ($(CONFIG_MMU)$(CONFIG_64BIT),yy) obj-$(CONFIG_BPF_SYSCALL) += arena.o range_tree.o`. My immediate next step is to ensure I've accounted for all the dependencies. I need to make *certain* there aren't any other configuration options involved, specific to `bpffs` or the BPF arena functionality. Perhaps there's a more granular filesystem-specific configuration option that isn't immediately obvious. The logical place to start looking for those would be the `fs/Kconfig` file, or perhaps the `kernel/bpf/Kconfig` file. I'll check those next. |
| 12/2 | 2026/10/10 23:15 | tool | grepper |
0mArgs:{
"Expression": "bpf_fs_type",
"PathPrefixes": [
"kernel/bpf/",
"fs/"
]
}Results: {
"Output": "kernel/bpf/inode.c=1351=static void bpf_kill_super(struct super_block *sb)\n--\nkernel/bpf/inode.c-1359-\nkernel/bpf/inode.c:1360:static struct file_system_type bpf_fs_type = {\nkernel/bpf/inode.c-1361-\t.owner\t\t= THIS_MODULE,\n--\nkernel/bpf/inode.c=1376=static int __init bpf_init(void)\n--\nkernel/bpf/inode.c-1390-\nkernel/bpf/inode.c:1391:\tret = register_filesystem(\u0026bpf_fs_type);\nkernel/bpf/inode.c-1392-\tif (ret) {\n"
} |
| 13/2 | 2026/10/10 23:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 38633 output: 36 thoughts: 959 Thoughts: **Analyzing the bpf Arena Implementation** Okay, let's break this down. First off, it's clear that the `bpf_fs_type` is unconditionally registered during initialization, so there's no separate configuration needed for the BPF filesystem beyond requiring the BPF syscalls. Now, let's dive into the core changes and figure out how they relate to selecting focus symbols. It looks like the `kernel/bpf/arena.c` file is the primary focus. We're dealing with arena-based memory management here. Several functions have been modified to handle new flags: `BPF_F_ARENA_NO_FREE` and `BPF_F_ARENA_EXPORT`. Specifically, `arena_map_alloc` now accounts for these flags, which is the first sign of an interesting extension. Looking at the diff, `arena_vm_fault`'s calculation of the kernel address from user space input seems standard but we should check it doesn't leave security holes, while `arena_get_unmapped_area` now incorporates checks related to the file operation (`filp->f_op`), maximum entries, and importantly, the export flag. This is critical as it looks like it's imposing size limits based on the bpf arena map, which is what we want. `arena_map_mmap` seems to be the key function for mmapping arena maps. We have checks for `exported` based on the file operations, size verifications, and assignment of `user_vm_start` and `user_vm_end` within the `arena` structure. This likely controls the region of user space we're allowing to be mapped. `zap_pages` is relevant as it handles cleaning the vmas based on the memory region, which has to be correct. The `bpf_arena_free_pages` family has checks for `BPF_F_ARENA_NO_FREE`, which affects how the memory is freed, or *isn't* freed. This is interesting as it may relate to how we expose the arena as a resource. Now, shifting to `kernel/bpf/inode.c`, we see some critical additions and modifications. `bpf_map_setattr` is a new inode operation to handle setting attributes. It rejects size changes for `BPF_MAP_TYPE_ARENA` maps that are also exported, which seems very important for the arena's consistency. This would ensure the arena's size can't be modified after its creation, which is a key part for its role. The `bpffs_arena_*` functions are all standard file system operations. `bpffs_arena_open` calls `security_bpf_map` (always a good thing), increments the map's reference count, and assigns the map to the file's private data. `bpffs_arena_mmap` sets `VM_SHARED` under certain conditions and then calls the `bpf_map_fops.mmap` function, which in turn calls `arena_map_mmap`, connecting the dots. `bpffs_arena_release` calls the appropriate release operation, while `bpffs_arena_get_unmapped_area` calls `arena_get_unmapped_area`. And we have llseek. Crucially, in `bpf_mkmap`, the file operations are chosen based on `shared_arena`, pointing either to `bpffs_arena_fops` or something else, and the size passed is based on the map's max entries and page size. This is how we are creating and setting the bpf arenas. Finally, I need to check the symbol visibility, which should answer whether they are standard functions. We need to verify if the key `arena_vm_fault`, `arena_map_mmap`, and `arena_get_unmapped_area`, `zap_pages`, `bpf_arena_free_pages` functions are exported, or if they are just normal symbols. Let's look at lines 500-660 of `kernel/bpf/arena.c`. |
| 14/2 | 2026/10/10 23:15 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/arena.c",
"FirstLine": 490,
"LineCount": 50
}Results: {
"Contents": " 490:\tstatic vm_fault_t arena_vm_fault(struct vm_fault *vmf)\n 491:\t{\n 492:\t\tstruct bpf_map *map = vmf-\u003evma-\u003evm_file-\u003eprivate_data;\n 493:\t\tstruct bpf_arena *arena = container_of(map, struct bpf_arena, map);\n 494:\t\tstruct mem_cgroup *new_memcg, *old_memcg;\n 495:\t\tLLIST_HEAD(pages);\n 496:\t\tstruct page *page, *new_page = NULL;\n 497:\t\tstruct apply_range_data data;\n 498:\t\tvm_fault_t fault_ret;\n 499:\t\tlong kbase, kaddr;\n 500:\t\tunsigned long flags;\n 501:\t\tint ret;\n 502:\t\n 503:\t\tkbase = bpf_arena_get_kern_vm_start(arena);\n 504:\t\t/* vmf-\u003epgoff includes the file offset of a bpffs-backed slice. */\n 505:\t\tkaddr = kbase + (u32)(arena-\u003euser_vm_start + ((u64)vmf-\u003epgoff \u003c\u003c PAGE_SHIFT));\n 506:\t\n 507:\t\tpage = vmalloc_to_page((void *)kaddr);\n 508:\t\tif (!page \u0026\u0026 !(arena-\u003emap.map_flags \u0026 BPF_F_SEGV_ON_FAULT)) {\n 509:\t\t\t/*\n 510:\t\t\t * Preallocate outside the lock so the allocation can reclaim;\n 511:\t\t\t * __GFP_RETRY_MAYFAIL keeps the OOM killer out of it.\n 512:\t\t\t */\n 513:\t\t\tbpf_map_memcg_enter(\u0026arena-\u003emap, \u0026old_memcg, \u0026new_memcg);\n 514:\t\t\tnew_page = alloc_pages_node(map-\u003enuma_node,\n 515:\t\t\t\t\t\t GFP_KERNEL | __GFP_ZERO |\n 516:\t\t\t\t\t\t __GFP_ACCOUNT | __GFP_NOWARN |\n 517:\t\t\t\t\t\t __GFP_RETRY_MAYFAIL, 0);\n 518:\t\t\tbpf_map_memcg_exit(old_memcg, new_memcg);\n 519:\t\t}\n 520:\t\n 521:\t\tif (raw_res_spin_lock_irqsave(\u0026arena-\u003espinlock, flags)) {\n 522:\t\t\t/*\n 523:\t\t\t * A failed lock means a possible deadlock was detected. Don't\n 524:\t\t\t * return VM_FAULT_RETRY: this handler never took mmap_lock, but\n 525:\t\t\t * the fault path would re-take it on retry and deadlock. Fail.\n 526:\t\t\t */\n 527:\t\t\tif (new_page)\n 528:\t\t\t\tfree_pages_nolock(new_page, 0);\n 529:\t\t\treturn VM_FAULT_SIGBUS;\n 530:\t\t}\n 531:\t\n 532:\t\tpage = vmalloc_to_page((void *)kaddr);\n 533:\t\tif (page) {\n 534:\t\t\tif (page == arena-\u003escratch_page) {\n 535:\t\t\t\t/* BPF triggered scratch here; don't lazy-alloc over it */\n 536:\t\t\t\tfault_ret = (arena-\u003emap.map_flags \u0026 BPF_F_SEGV_ON_FAULT)\n 537:\t\t\t\t\t ? VM_FAULT_SIGSEGV : VM_FAULT_SIGBUS;\n 538:\t\t\t\tgoto out_err_locked;\n 539:\t\t\t}\n"
} |
| 15/2 | 2026/10/10 23:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 41454 output: 31 thoughts: 658 Thoughts: **Analyzing Arena Operations Hooking** Okay, let's dive into how these arena operations are actually linked up within the system. I need to understand this mechanism thoroughly. It seems like we're looking at a set of functions that are declared as `static` within their respective `.c` files, which initially suggests they might be candidates for inlining. First, I've noticed the following key functions are tied into various operation structures: `arena_vm_fault` which is connected to `arena_vm_ops.fault`, `arena_map_mmap` which is used by `arena_map_ops.map_mmap`, `arena_get_unmapped_area` with `arena_map_ops.map_get_unmapped_area`, `arena_map_alloc` used within `arena_map_ops.map_alloc`, `bpf_map_setattr` that's part of `bpf_map_iops.setattr` and a whole host of `bpffs_arena_*` functions (`bpffs_arena_open`, `bpffs_arena_mmap`, `bpffs_arena_llseek`, etc.) that are grouped into `bpffs_arena_fops`. I'm also seeing `zap_pages`, used in `arena_free_pages` and other places, and the kfunc `bpf_arena_free_pages`. Ah, here's the crucial point: despite being declared `static` *within* the source files, these aren't actually `static inline` functions. Since their addresses are taken because they're being used as function pointers in structures like `vm_operations_struct`, `bpf_map_ops`, `inode_operations`, or `file_operations` (the standard fops/ops), the compiler *cannot* inline them. The `-fno-inline-functions-called-once` compiler flag, or the fact that they're function pointers, prevents inlining. This means they remain as discrete symbols within the kernel's symbol table, kallsyms, and vmlinux, exactly as you'd expect with standard C functions. They're regular symbols, not inlined. Now I need to verify this assumption. Let me re-confirm that the critical functions listed earlier are indeed used as function pointers within those structures: `arena_vm_fault`, `arena_map_mmap`, `arena_get_unmapped_area`, `arena_map_alloc`, `bpf_map_setattr`, `bpffs_arena_open`, `bpffs_arena_mmap`, `bpffs_arena_llseek`, `zap_pages`, and `bpf_arena_free_pages`. I am also asking myself if `bpf_map_mmap` or something similar also exists in `kernel/bpf/syscall.c`. Let me just use `grep` to search for `arena_map_ops` to get more context and double-check my understanding of how this is all connected. |
| 16/2 | 2026/10/10 23:15 | tool | grepper |
0mArgs:{
"Expression": "arena_map_ops",
"PathPrefixes": [
"kernel/bpf/"
]
}Results: {
"Output": "kernel/bpf/arena.c=724=BTF_ID_LIST_SINGLE(bpf_arena_map_btf_ids, struct, bpf_arena)\nkernel/bpf/arena.c:725:const struct bpf_map_ops arena_map_ops = {\nkernel/bpf/arena.c-726-\t.map_meta_equal = bpf_map_meta_equal,\n--\nkernel/bpf/core.c=3557=u64 bpf_check_timed_may_goto(struct bpf_timed_may_goto *p)\n--\nkernel/bpf/core.c-3575-/* for configs without MMU or 32-bit */\nkernel/bpf/core.c:3576:__weak const struct bpf_map_ops arena_map_ops;\nkernel/bpf/core.c-3577-__weak u64 bpf_arena_get_user_vm_start(struct bpf_arena *arena)\n"
} |
| 17/2 | 2026/10/10 23:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 43057 output: 33 thoughts: 38 Thoughts: Here's a summary of my thought process, as requested: **Diving into the BPF Map MMAP Mechanism** Okay, so I need to refresh my memory on how `bpf_map_fops` interacts with the `mmap` syscall for BPF maps. Let's start by zeroing in on the implementation itself. I should check the source for `bpf_map_mmap`, which, I believe, resides in `kernel/bpf/syscall.c`. That should provide the core logic for memory mapping these BPF structures. I'm especially interested in how it handles the file operations (`fops`) aspect and how permissions/memory protection are managed under the hood. I expect to see the typical mmap-related operations like `mmap`, `mprotect`, and potential checks for user-space access violations. I also expect to see the interplay with page allocation and management within the context of the BPF map's backing memory. This exploration should clarify the performance implications and potential security considerations of allowing user-space direct access to these BPF data structures. I'm keen to see how they've handled the typical challenges of memory mapping kernel-managed memory and what safeguards are in place to prevent misuse. |
| 18/2 | 2026/10/10 23:16 | tool | grepper |
0mArgs:{
"Expression": "bpf_map_fops",
"PathPrefixes": [
"kernel/bpf/"
]
}Results: {
"Output": "kernel/bpf/arena.c=614=static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long addr,\n--\nkernel/bpf/arena.c-621-\nkernel/bpf/arena.c:622:\tif (filp-\u003ef_op != \u0026bpf_map_fops) {\nkernel/bpf/arena.c-623-\t\tif (!len || pgoff \u003e= map-\u003emax_entries ||\n--\nkernel/bpf/arena.c=658=static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)\n--\nkernel/bpf/arena.c-660-\tstruct bpf_arena *arena = container_of(map, struct bpf_arena, map);\nkernel/bpf/arena.c:661:\tbool exported = vma-\u003evm_file-\u003ef_op != \u0026bpf_map_fops;\nkernel/bpf/arena.c-662-\n--\nkernel/bpf/inode.c=391=static int bpffs_arena_mmap(struct file *file, struct vm_area_struct *vma)\n--\nkernel/bpf/inode.c-400-\nkernel/bpf/inode.c:401:\treturn bpf_map_fops.mmap(file, vma);\nkernel/bpf/inode.c-402-}\n--\nkernel/bpf/inode.c=404=static int bpffs_arena_release(struct inode *inode, struct file *file)\nkernel/bpf/inode.c-405-{\nkernel/bpf/inode.c:406:\treturn bpf_map_fops.release(inode, file);\nkernel/bpf/inode.c-407-}\n--\nkernel/bpf/inode.c=409=static unsigned long bpffs_arena_get_unmapped_area(struct file *file,\n--\nkernel/bpf/inode.c-414-{\nkernel/bpf/inode.c:415:\treturn bpf_map_fops.get_unmapped_area(file, addr, len, pgoff, flags);\nkernel/bpf/inode.c-416-}\n--\nkernel/bpf/syscall.c=1177=static unsigned long bpf_get_unmapped_area(struct file *filp, unsigned long addr,\n--\nkernel/bpf/syscall.c-1191-\nkernel/bpf/syscall.c:1192:const struct file_operations bpf_map_fops = {\nkernel/bpf/syscall.c-1193-#ifdef CONFIG_PROC_FS\n--\nkernel/bpf/syscall.c=1204=int bpf_map_new_fd(struct bpf_map *map, int flags)\n--\nkernel/bpf/syscall.c-1211-\nkernel/bpf/syscall.c:1212:\treturn anon_inode_getfd(\"bpf-map\", \u0026bpf_map_fops, map,\nkernel/bpf/syscall.c-1213-\t\t\t\tflags | O_CLOEXEC);\n--\nkernel/bpf/syscall.c=5580=static int bpf_obj_get_info_by_fd(const union bpf_attr *attr,\n--\nkernel/bpf/syscall.c-5592-\t\t\t\t\t uattr);\nkernel/bpf/syscall.c:5593:\telse if (fd_file(f)-\u003ef_op == \u0026bpf_map_fops)\nkernel/bpf/syscall.c-5594-\t\treturn bpf_map_get_info_by_fd(fd_file(f), fd_file(f)-\u003eprivate_data, attr,\n"
} |
| 19/2 | 2026/10/10 23:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44332 output: 31 thoughts: 1603 Thoughts: **My Analysis of BPF Arena Maps for Fuzzing** Alright, let's break down this BPF arena map implementation and figure out its fuzzing potential. I'm seeing the creation of these arena maps via `bpf(BPF_MAP_CREATE, ...)` with the `BPF_MAP_TYPE_ARENA` type and specific flags. This eventually calls into `arena_map_alloc()` after passing through the map creation and allocation routines. The map gets pinned to bpffs using `bpf_obj_pin_path()`, `bpf_obj_do_pin()`, and sets file operations. Then we have userspace interacting with it: opening the bpffs file via `bpffs_arena_open()`, followed by `mmap()` using `bpffs_arena_mmap()` and finally, `arena_map_mmap()` to set up the mapping. Reads and writes trigger `arena_vm_fault()` due to page faults, and other operations like `truncate()`, `lseek()` use specific functions. Notably, there's a mechanism for freeing pages through `bpf_arena_free_pages()`, which has a check for `BPF_F_ARENA_NO_FREE`. `zap_pages()` is also important, which covers the arena pages across slices. Now, let's look for potential vulnerabilities, especially things syzkaller can reach: I'm concerned about how `vmf->pgoff` and `zap_pages` are used within the address calculations, and the bounds checks, especially around `arena_map_mmap`. The `bpf_map_setattr` function also checks for size attributes related to BPF arenas, which may have unexpected behavior, and the mmap of bpffs file, which can allow VM_SHARED to be set. Crucially, this entire process is readily accessible from unprivileged userspace, depending on the kernel's `kernel.unprivileged_bpf_disabled` sysctl. Standard system calls like `bpf`, `open`, `mmap`, and `truncate` are the entry points here – perfect for syzkaller to target. For syzkaller's FocusSymbols, we need to identify the relevant code that was added or modified. The key functions in `kernel/bpf/arena.c` are: `arena_map_alloc`, `arena_vm_fault`, `arena_get_unmapped_area`, `arena_map_mmap`, `zap_pages`, `bpf_arena_free_pages`, and `bpf_arena_free_pages_non_sleepable`. In `kernel/bpf/inode.c`: `bpf_map_setattr`, `bpffs_arena_open`, `bpffs_arena_mmap`, `bpffs_arena_release`, `bpffs_arena_get_unmapped_area`, `bpffs_arena_llseek`, and `bpf_mkmap` are essential. These are very specific to the arena feature. Finally, configuration is important. `CONFIG_BPF_SYSCALL` is essential, since it enables the BPF subsystem which creates all the necessary BPF and bpffs infrastructure. On my target architecture (amd64, where CONFIG_64BIT and CONFIG_MMU are y), this configuration appears to be sufficient. The `bpf_fs_type` registration in `bpf_init()` confirms that. I don't see any other required configurations in my analysis of `kernel/bpf/arena.c`. |
| 20/2 | 2026/10/10 23:16 | tool | grepper |
0mError:no matches Args: {
"Expression": "CONFIG_",
"PathPrefixes": [
"kernel/bpf/arena.c"
]
}Results: {
"error": "no matches"
} |
| 21/2 | 2026/10/10 23:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 46751 output: 31 thoughts: 33 Thoughts: Okay, here's my take on those thoughts, framed as a first-person expert summary: **BPF Code Review: CONFIG_ Awareness** Alright, let's see, I'm reviewing this BPF code and I noticed something interesting. I was digging around in `kernel/bpf/arena.c` and found, to my surprise, absolutely no conditional checks tied to `CONFIG_` preprocessor directives. That's a bit odd, usually, in a core subsystem like this, you'd expect to see some feature toggling based on configuration options. Now, that alone isn't necessarily a *problem*, but it does make me start thinking about the overall configuration strategy of this component. The absence suggests either: 1. This particular file is intentionally designed to be unconditionally present and functional regardless of how BPF is configured. That is a likely, and good, design decision. 2. The relevant `CONFIG_` options are handled elsewhere, maybe at a higher level, or in the build system itself, which is also valid. My next step, before assuming anything, is to check `kernel/bpf/inode.c`. That's an adjacent file, part of the same subsystem, and it is a good place to look for any missing puzzle pieces in the code. I am curious if they are present there, as that is a more likely place to find those kinds of directives. |
| 22/2 | 2026/10/10 23:16 | tool | grepper |
0mError:no matches Args: {
"Expression": "CONFIG_",
"PathPrefixes": [
"kernel/bpf/inode.c"
]
}Results: {
"error": "no matches"
} |
| 23/2 | 2026/10/10 23:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 47127 output: 300 thoughts: 426 Thoughts: **Fuzzing Configuration Thoughts** Okay, so it looks like we only need `CONFIG_BPF_SYSCALL` enabled for this fuzzing setup, that simplifies things nicely. Let's double-check the `WorthFuzzing` verdict and the reasoning behind it to be absolutely certain. Yep, `WorthFuzzing` is confirmed as `true`, and the reasoning is spot on: the patch introduces significant changes to BPF arena map memory management, supporting export and retention, effectively allowing shared memory files within bpffs. These modifications touch core areas, including mmap calculations, page fault handling, unmapping, and inode operations. These paths are definitely reachable through standard userspace operations via the bpf syscall, bpffs operations, and memory access which makes it a great target. The `EnableConfigs` list confirms `BPF_SYSCALL` is the only requirement, as expected. The `FocusSymbols` list looks comprehensive and covers all the critical functions involved: `arena_map_alloc`, `arena_vm_fault`, `arena_get_unmapped_area`, `arena_map_mmap`, `zap_pages`, and the arena freeing functions - `bpf_arena_free_pages`, `bpf_arena_free_pages_non_sleepable`, as well as file operation functions like `bpf_map_setattr`, `bpffs_arena_open`, `bpffs_arena_mmap`, `bpffs_arena_llseek`, `bpffs_arena_get_unmapped_area`, and `bpf_mkmap`. They are all important areas to fuzz. Before proceeding, let's just make sure everything is in order and that we have all the required components to call `set-results`. We have `WorthFuzzing` as a boolean, a list of `FocusSymbols` as strings, the `EnableConfigs` as a list of strings, and finally, `Reasoning` as a string. Check, check, and check. Everything seems to be accurate and complete. Okay, let's go ahead and call `set-results`. |
| 24/2 | 2026/10/10 23:16 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"BPF_SYSCALL"
],
"FocusSymbols": [
"arena_map_alloc",
"arena_vm_fault",
"arena_get_unmapped_area",
"arena_map_mmap",
"zap_pages",
"bpf_arena_free_pages",
"bpf_arena_free_pages_non_sleepable",
"bpf_map_setattr",
"bpffs_arena_open",
"bpffs_arena_mmap",
"bpffs_arena_llseek",
"bpffs_arena_get_unmapped_area",
"bpf_mkmap"
],
"Reasoning": "The patch introduces export and retention support for BPF arena maps (BPF_F_ARENA_EXPORT and BPF_F_ARENA_NO_FREE), allowing pinned arenas in bpffs to be opened and mmap'd as shared-memory files with slice offsets. It modifies core BPF arena memory management and bpffs inode operations, altering VMA tracking, page fault resolution (accounting for slice offsets), unmapped area discovery, range zapping, and file operations (open, mmap, setattr, llseek). These paths are reachable from userspace in standard virtualized environments via standard syscalls.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"BPF_SYSCALL"
],
"FocusSymbols": [
"arena_map_alloc",
"arena_vm_fault",
"arena_get_unmapped_area",
"arena_map_mmap",
"zap_pages",
"bpf_arena_free_pages",
"bpf_arena_free_pages_non_sleepable",
"bpf_map_setattr",
"bpffs_arena_open",
"bpffs_arena_mmap",
"bpffs_arena_llseek",
"bpffs_arena_get_unmapped_area",
"bpf_mkmap"
],
"Reasoning": "The patch introduces export and retention support for BPF arena maps (BPF_F_ARENA_EXPORT and BPF_F_ARENA_NO_FREE), allowing pinned arenas in bpffs to be opened and mmap'd as shared-memory files with slice offsets. It modifies core BPF arena memory management and bpffs inode operations, altering VMA tracking, page fault resolution (accounting for slice offsets), unmapped area discovery, range zapping, and file operations (open, mmap, setattr, llseek). These paths are reachable from userspace in standard virtualized environments via standard syscalls.",
"WorthFuzzing": true
} |
| 25/1 | 2026/10/10 23:16 | action | If |
0mArgs:{
"WorthFuzzing": true
}Results: null |
| 26/2 | 2026/10/10 23:16 | agent | kmsan-evaluator |
0mModel:core Args: null Results: {
"KMSANReasoning": "The patch series introduces BPF_F_ARENA_NO_FREE and BPF_F_ARENA_EXPORT flags for BPF arenas, allowing pinned arenas in bpffs to be opened and mapped as sized files with exported slices.\n\n1. Uninitialized Memory Analysis (KMSAN):\n- No kernel stack or heap structures are copied to userspace (no copy_to_user, put_user, netlink attributes, ioctl out-arguments, etc.).\n- File operations on the exported bpffs inode (bpffs_arena_fops) do not implement .read or .write handlers; memory access occurs solely via memory-mapping (mmap/fault).\n- Backing pages allocated during fault handling (arena_vm_fault) continue to use GFP_KERNEL | __GFP_ZERO.\n- All new stack variables in modified functions (e.g. start, end, vma_start, vma_end in zap_pages, and err in bpffs_arena_open) are explicitly initialized prior to evaluation.\n- No uninitialized memory or info-leak risk is introduced or exposed.\n\n2. Applicable Sanitizers (KASAN, LOCKDEP, DEBUG_VM):\n- Potential risks in this change involve offset and address calculations (e.g. kaddr calculation in arena_vm_fault and slice offsets in zap_pages), which could risk out-of-bounds accesses in the vmalloc area (caught by KASAN).\n- Reference counting and lifecycle management of the map across open/mmap/release/unlink (bpf_map_inc_with_uref, UAF or leak) are covered by KASAN and standard refcount tracking.\n- Locking and concurrency interactions (arena-\u003elock, arena-\u003espinlock) are monitored by LOCKDEP.\n\nTherefore, standard KASAN and debugging builds fully cover the risk profile of these changes, and a dedicated KMSAN fuzzing session is not warranted (NeedsKMSAN=false).",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit a034d16eea30a35c1529eb1fb108e6ad8ca18920
Author: syz-cluster <triage@syzkaller.com>
Date: Sat Oct 10 23:15:04 2026 +0000
syz-cluster: applied patch under review
diff --git a/Documentation/bpf/map_arena.rst b/Documentation/bpf/map_arena.rst
new file mode 100644
index 0000000000000..1b9eadd5fb3a1
--- /dev/null
+++ b/Documentation/bpf/map_arena.rst
@@ -0,0 +1,131 @@
+.. SPDX-License-Identifier: GPL-2.0-only
+
+==================
+BPF_MAP_TYPE_ARENA
+==================
+
+A BPF arena provides shared memory that BPF programs and userspace can
+access directly. Create the map with zero key and value sizes and
+``BPF_F_MMAPABLE``. ``max_entries`` specifies its capacity in pages, up to
+4 GiB. The canonical BPF address range must not cross a 4 GiB boundary.
+
+Pinned arena files
+==================
+
+An arena created with ``BPF_F_ARENA_EXPORT`` can be pinned in bpffs and
+opened as a sized file. Its size is ``max_entries * PAGE_SIZE``. Export
+requires ``BPF_F_ARENA_NO_FREE`` so BPF programs cannot release backing
+pages while external consumers may still use them. Creating an arena with
+``BPF_F_ARENA_EXPORT`` alone fails with ``EINVAL``. Arenas without the
+export flag retain their existing bpffs behavior and cannot be opened
+this way, including arenas created with ``BPF_F_ARENA_NO_FREE`` alone.
+
+Establish the canonical BPF address range before mapping the pinned file:
+
+* Set ``map_extra`` to a nonzero, page-aligned address at map creation.
+ The canonical range then covers the full map capacity, even if the map
+ FD has not been mmaped.
+* Alternatively, leave ``map_extra`` zero and mmap the map FD first. This
+ mapping establishes the canonical address and must cover the full map
+ capacity. Shorter canonical mappings fail with ``EINVAL``.
+
+A canonical map FD mapping may start at address zero if the system's
+low-address mapping policy permits it.
+
+Mapping the pinned file before either step fails with ``EINVAL``. Exported
+mappings do not establish or change the canonical BPF address range.
+
+Once the canonical range is established, it covers the entire capacity
+reported by ``fstat()``. Consumers can map the full file or smaller slices.
+Establishing the full canonical mapping reserves virtual address space;
+it does not populate every arena page. Arenas without
+``BPF_F_ARENA_EXPORT`` can still establish shorter canonical mappings.
+
+Open the pin with ``O_RDONLY`` for read-only access or ``O_RDWR`` for
+writable access, and mmap it with ``MAP_SHARED``. An exported
+mapping may use a different virtual address from the canonical mapping.
+Its offset must be page-aligned, and its offset and length must fit within
+both the map capacity and the established canonical range. Consumers can
+map disjoint slices of one arena. Oversized or out-of-range mappings fail
+with ``EINVAL``. ``MAP_PRIVATE`` mappings are unsupported.
+
+Exported mappings share the bytes of the arena without relocating pointers
+stored in those bytes. Arena pointers produced by a BPF program use that
+arena's canonical address representation, based on ``user_vm_start``.
+A consumer mapping a slice at another address must translate such pointers
+before dereferencing them. A guest BPF arena has its own canonical range;
+sharing the backing pages does not make host arena pointers valid in the
+guest arena, or vice versa. Communication structures can use offsets
+relative to the shared slice, with each side validating the offsets and
+translating them to addresses in its own mapping or arena.
+
+A pin opened with ``O_RDONLY`` supports read-only shared mappings, which
+observe updates made by BPF programs and other writable mappings. Requests
+for ``PROT_WRITE`` through that FD fail with ``EACCES``. A mapping made
+without ``PROT_WRITE`` cannot subsequently acquire write permission with
+``mprotect()``, even when the pin was opened with ``O_RDWR``.
+
+Pin permissions can grant readers access independently of writers. For
+example, mode ``0640`` permits the owner to open the pin for writable
+access and the group to open it for read-only access, subject to the
+applicable LSM checks. Changing permissions does not revoke access through
+existing FDs or mappings.
+
+The file has a fixed size. ``truncate()``, ``ftruncate()``, and ``O_TRUNC``
+cannot change it. A truncate request that leaves the size unchanged is
+allowed. Permission and ownership changes remain available.
+Consumers can discover the capacity with ``fstat()`` or
+``lseek(fd, 0, SEEK_END)``. Seeking supports ``SEEK_SET``, ``SEEK_CUR``, and
+``SEEK_END`` within the file bounds; it does not enable ``read()`` or
+``write()`` access to the memory.
+
+``fsync()``, ``fdatasync()``, and ``msync(..., MS_SYNC)`` succeed without
+performing writeback: arena memory is volatile and has no persistent
+backing. These calls do not synchronize access between consumers or BPF
+programs; shared-memory protocols still need appropriate memory ordering.
+
+Exported mappings inherit the arena's mapping lifecycle restrictions.
+They are not inherited across ``fork()``, and ``MADV_DOFORK`` cannot enable
+inheritance. Consumers must create their own mappings in a child process.
+Mappings cannot be split, expanded, or relocated with ``mremap()``.
+Partial unmapping and protection changes that require splitting a mapping
+fail with ``EINVAL``. Unmapping an entire mapping remains supported, as do
+protection changes that do not require a split and satisfy the access
+restrictions above. Consumers needing independently managed regions
+should create separate slice mappings.
+
+Access control and lifetime
+===========================
+
+Opening the pin uses ordinary filesystem permissions and file LSM checks,
+as well as ``security_bpf_map()`` for the requested access mode. Mapping it
+also passes the ordinary mmap LSM checks. Writable mappings remain subject
+to the BPF map's frozen state and program-read-only restrictions.
+
+Consumers use ``open()`` and ``mmap()`` rather than ``bpf(BPF_OBJ_GET)``.
+Consequently, seccomp rules and LSM policies specific to the ``bpf()``
+syscall do not govern this file interface. Grant access through the pin's
+permissions and the applicable file, mmap, and BPF map LSM policies.
+A process prohibited from calling ``bpf()`` can still access the arena
+through this interface if those permissions and checks allow it.
+``unprivileged_bpf_disabled`` restricts unprivileged map and program
+creation; it does not generally prohibit access to existing maps.
+
+``BPF_F_ARENA_NO_FREE`` makes ``bpf_arena_free_pages()`` a no-op in both
+sleepable and non-sleepable BPF programs. The free operation returns no
+indication that pages were retained. Programs using this flag must not
+rely on freeing pages to restore arena allocation capacity. Arena pages
+are retained until the map is destroyed. Open map and pin FDs, mappings,
+and other map references retain the arena, including after the pin is
+unlinked. Removing the pin therefore does not revoke existing mappings or
+free their pages.
+
+The retention flag does not itself expose the pin as a sized file.
+``BPF_F_ARENA_EXPORT`` opts into that interface and must be combined with
+``BPF_F_ARENA_NO_FREE``. Retention-only arenas can still share memory
+through ordinary map FD mappings and manage reusable objects within
+their retained pages.
+
+For example, a consumer such as QEMU can use an exported slice as a shared
+file-backed guest RAM backend. The file interface provides shared memory;
+consumers must supply their own allocation and communication protocol.
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index e0ed44b1bbcb7..40a5489ac9fea 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1500,6 +1500,12 @@ enum {
/* Enable BPF ringbuf overwrite mode */
BPF_F_RB_OVERWRITE = (1U << 19),
+
+ /* Keep arena pages allocated until the map is destroyed. */
+ BPF_F_ARENA_NO_FREE = (1U << 20),
+
+ /* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */
+ BPF_F_ARENA_EXPORT = (1U << 21),
};
/* Flags for BPF_PROG_QUERY. */
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 0ff707707da04..98f6dffe0138f 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -283,7 +283,13 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
/* BPF_F_MMAPABLE must be set */
!(attr->map_flags & BPF_F_MMAPABLE) ||
/* No unsupported flags present */
- (attr->map_flags & ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE | BPF_F_NO_USER_CONV)))
+ (attr->map_flags & ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE |
+ BPF_F_NO_USER_CONV | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT)))
+ return ERR_PTR(-EINVAL);
+
+ /* Exported memory must remain allocated while consumers use it. */
+ if ((attr->map_flags & BPF_F_ARENA_EXPORT) && !(attr->map_flags & BPF_F_ARENA_NO_FREE))
return ERR_PTR(-EINVAL);
if (attr->map_extra & ~PAGE_MASK)
@@ -495,7 +501,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
int ret;
kbase = bpf_arena_get_kern_vm_start(arena);
- kaddr = kbase + (u32)(vmf->address);
+ /* vmf->pgoff includes the file offset of a bpffs-backed slice. */
+ kaddr = kbase + (u32)(arena->user_vm_start + ((u64)vmf->pgoff << PAGE_SHIFT));
page = vmalloc_to_page((void *)kaddr);
if (!page && !(arena->map.map_flags & BPF_F_SEGV_ON_FAULT)) {
@@ -612,13 +619,23 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
long ret;
+ if (filp->f_op != &bpf_map_fops) {
+ if (!len || pgoff >= map->max_entries ||
+ len > ((u64)map->max_entries - pgoff) << PAGE_SHIFT)
+ return -EINVAL;
+ return mm_get_unmapped_area(filp, addr, len, pgoff, flags);
+ }
+
if (pgoff)
return -EINVAL;
if (len > SZ_4G)
return -E2BIG;
+ if ((map->map_flags & BPF_F_ARENA_EXPORT) &&
+ len != (u64)map->max_entries << PAGE_SHIFT)
+ return -EINVAL;
- /* if user_vm_start was specified at arena creation time */
- if (arena->user_vm_start) {
+ /* Once established, the canonical range cannot change. */
+ if (arena->user_vm_end) {
if (len > arena->user_vm_end - arena->user_vm_start)
return -E2BIG;
if (len != arena->user_vm_end - arena->user_vm_start)
@@ -632,7 +649,7 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
return ret;
if ((ret >> 32) == ((ret + len - 1) >> 32))
return ret;
- if (WARN_ON_ONCE(arena->user_vm_start))
+ if (WARN_ON_ONCE(arena->user_vm_end))
/* checks at map creation time should prevent this */
return -EFAULT;
return round_up(ret, SZ_4G);
@@ -641,9 +658,13 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
{
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
+ bool exported = vma->vm_file->f_op != &bpf_map_fops;
guard(mutex)(&arena->lock);
- if (arena->user_vm_start && arena->user_vm_start != vma->vm_start)
+ if (!exported && (map->map_flags & BPF_F_ARENA_EXPORT) &&
+ vma->vm_end - vma->vm_start != (u64)map->max_entries << PAGE_SHIFT)
+ return -EINVAL;
+ if (!exported && arena->user_vm_end && arena->user_vm_start != vma->vm_start)
/*
* If map_extra was not specified at arena creation time then
* 1st user process can do mmap(NULL, ...) to pick user_vm_start
@@ -654,19 +675,31 @@ static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
*/
return -EBUSY;
- if (arena->user_vm_end && arena->user_vm_end != vma->vm_end)
+ if (exported && !arena->user_vm_end)
+ return -EINVAL;
+ if (!exported && arena->user_vm_end && arena->user_vm_end != vma->vm_end)
/* all user processes must have the same size of mmap-ed region */
return -EBUSY;
- /* Earlier checks should prevent this */
- if (WARN_ON_ONCE(vma->vm_end - vma->vm_start > SZ_4G || vma->vm_pgoff))
+ if (exported) {
+ u64 page_cnt = (arena->user_vm_end - arena->user_vm_start) >> PAGE_SHIFT;
+
+ if (vma->vm_pgoff >= page_cnt ||
+ (vma->vm_end - vma->vm_start) >> PAGE_SHIFT >
+ page_cnt - vma->vm_pgoff)
+ return -EINVAL;
+ } else if (WARN_ON_ONCE(vma->vm_end - vma->vm_start > SZ_4G || vma->vm_pgoff)) {
+ /* Earlier checks should prevent this for map FD mappings. */
return -EFAULT;
+ }
if (remember_vma(arena, vma))
return -ENOMEM;
- arena->user_vm_start = vma->vm_start;
- arena->user_vm_end = vma->vm_end;
+ if (!exported) {
+ arena->user_vm_start = vma->vm_start;
+ arena->user_vm_end = vma->vm_end;
+ }
/*
* bpf_map_mmap() checks that it's being mmaped as VM_SHARED and
* clears VM_MAYEXEC. Set VM_DONTEXPAND to avoid potential change
@@ -840,8 +873,12 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
struct mm_struct *mm;
struct vma_list *vml;
unsigned long vm_start;
+ u64 start, end, vma_start, vma_end;
u64 my_gen;
+ start = uaddr - arena->user_vm_start;
+ end = start + size;
+
/*
* Taking mmap_read_lock() under arena->lock would deadlock against
* arena_vm_close(), which runs with mmap_write_lock held and then
@@ -880,8 +917,14 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
*/
vma = find_vma(mm, vm_start);
if (vma && vma->vm_start == vm_start &&
- vma->vm_file && vma->vm_file->private_data == &arena->map)
- zap_vma_range(vma, uaddr, size);
+ vma->vm_file && vma->vm_file->private_data == &arena->map) {
+ vma_start = (u64)vma->vm_pgoff << PAGE_SHIFT;
+ vma_end = vma_start + vma->vm_end - vma->vm_start;
+ if (start < vma_end && end > vma_start)
+ zap_vma_range(vma, vma->vm_start +
+ (max(start, vma_start) - vma_start),
+ min(end, vma_end) - max(start, vma_start));
+ }
mmap_read_unlock(mm);
mmput(mm);
@@ -1134,7 +1177,8 @@ __bpf_kfunc void bpf_arena_free_pages(void *p__map, void *ptr__ign, u32 page_cnt
struct bpf_map *map = p__map;
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
- if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign)
+ if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign ||
+ (map->map_flags & BPF_F_ARENA_NO_FREE))
return;
arena_free_pages(arena, (long)ptr__ign, page_cnt, true);
}
@@ -1144,7 +1188,8 @@ void bpf_arena_free_pages_non_sleepable(void *p__map, void *ptr__ign, u32 page_c
struct bpf_map *map = p__map;
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
- if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign)
+ if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign ||
+ (map->map_flags & BPF_F_ARENA_NO_FREE))
return;
arena_free_pages(arena, (long)ptr__ign, page_cnt, false);
}
diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c
index 7837968c0842c..d3dc6f70e32ae 100644
--- a/kernel/bpf/inode.c
+++ b/kernel/bpf/inode.c
@@ -119,8 +119,24 @@ static const struct inode_operations bpf_symlink_iops;
static const struct inode_operations bpf_prog_iops = {
.listxattr = bpf_fs_listxattr,
};
+
+static int bpf_map_setattr(struct mnt_idmap *idmap, struct dentry *dentry,
+ struct iattr *attr)
+{
+ struct inode *inode = d_inode(dentry);
+ struct bpf_map *map = inode->i_private;
+
+ if (map->map_type == BPF_MAP_TYPE_ARENA &&
+ (map->map_flags & BPF_F_ARENA_EXPORT) &&
+ (attr->ia_valid & ATTR_SIZE) && attr->ia_size != i_size_read(inode))
+ return -EINVAL;
+
+ return simple_setattr(idmap, dentry, attr);
+}
+
static const struct inode_operations bpf_map_iops = {
.listxattr = bpf_fs_listxattr,
+ .setattr = bpf_map_setattr,
};
static const struct inode_operations bpf_link_iops = {
.listxattr = bpf_fs_listxattr,
@@ -351,6 +367,68 @@ static const struct file_operations bpffs_map_fops = {
.release = bpffs_map_release,
};
+/*
+ * An arena with BPF_F_ARENA_EXPORT pinned in bpffs can serve as a
+ * shared-memory file. The ordinary map FD is an anonymous inode with no
+ * size and cannot be opened by pathname, so keep the pin's inode and
+ * forward mmap to the map FD implementation. The pinned file uses ordinary
+ * address selection so QEMU can map the pages at another virtual address.
+ */
+static int bpffs_arena_open(struct inode *inode, struct file *file)
+{
+ struct bpf_map *map = inode->i_private;
+ int err;
+
+ err = security_bpf_map(map, file->f_mode);
+ if (err)
+ return err;
+ bpf_map_inc_with_uref(map);
+ file->private_data = map;
+
+ return 0;
+}
+
+static int bpffs_arena_mmap(struct file *file, struct vm_area_struct *vma)
+{
+ /*
+ * Generic mmap clears VM_SHARED for O_RDONLY files, but VM_MAYSHARE
+ * still distinguishes shared mappings from private ones. Restore
+ * VM_SHARED for the map mmap path without granting VM_MAYWRITE.
+ */
+ if (!(file->f_mode & FMODE_WRITE) && (vma->vm_flags & VM_MAYSHARE))
+ vm_flags_set(vma, VM_SHARED);
+
+ return bpf_map_fops.mmap(file, vma);
+}
+
+static int bpffs_arena_release(struct inode *inode, struct file *file)
+{
+ return bpf_map_fops.release(inode, file);
+}
+
+static unsigned long bpffs_arena_get_unmapped_area(struct file *file,
+ unsigned long addr,
+ unsigned long len,
+ unsigned long pgoff,
+ unsigned long flags)
+{
+ return bpf_map_fops.get_unmapped_area(file, addr, len, pgoff, flags);
+}
+
+static loff_t bpffs_arena_llseek(struct file *file, loff_t offset, int whence)
+{
+ return fixed_size_llseek(file, offset, whence, i_size_read(file_inode(file)));
+}
+
+static const struct file_operations bpffs_arena_fops = {
+ .open = bpffs_arena_open,
+ .llseek = bpffs_arena_llseek,
+ .fsync = noop_fsync,
+ .release = bpffs_arena_release,
+ .mmap = bpffs_arena_mmap,
+ .get_unmapped_area = bpffs_arena_get_unmapped_area,
+};
+
static int bpffs_obj_open(struct inode *inode, struct file *file)
{
return -EIO;
@@ -362,7 +440,7 @@ static const struct file_operations bpffs_obj_fops = {
static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
const struct inode_operations *iops,
- const struct file_operations *fops)
+ const struct file_operations *fops, loff_t size)
{
struct inode *dir = dentry->d_parent->d_inode;
struct inode *inode;
@@ -382,6 +460,7 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
inode->i_op = iops;
inode->i_fop = fops;
inode->i_private = raw;
+ i_size_write(inode, size);
bpf_dentry_finalize(dentry, inode, dir);
return 0;
@@ -390,16 +469,20 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
static int bpf_mkprog(struct dentry *dentry, umode_t mode, void *arg)
{
return bpf_mkobj_ops(dentry, mode, arg, &bpf_prog_iops,
- &bpffs_obj_fops);
+ &bpffs_obj_fops, 0);
}
static int bpf_mkmap(struct dentry *dentry, umode_t mode, void *arg)
{
struct bpf_map *map = arg;
+ bool shared_arena = map->map_type == BPF_MAP_TYPE_ARENA &&
+ (map->map_flags & BPF_F_ARENA_EXPORT);
return bpf_mkobj_ops(dentry, mode, arg, &bpf_map_iops,
- bpf_map_support_seq_show(map) ?
- &bpffs_map_fops : &bpffs_obj_fops);
+ shared_arena ? &bpffs_arena_fops :
+ bpf_map_support_seq_show(map) ?
+ &bpffs_map_fops : &bpffs_obj_fops,
+ shared_arena ? (loff_t)map->max_entries * PAGE_SIZE : 0);
}
static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)
@@ -408,7 +491,7 @@ static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)
return bpf_mkobj_ops(dentry, mode, arg, &bpf_link_iops,
bpf_link_is_iter(link) ?
- &bpf_iter_fops : &bpffs_obj_fops);
+ &bpf_iter_fops : &bpffs_obj_fops, 0);
}
static struct dentry *
@@ -483,7 +566,7 @@ static int bpf_iter_link_pin_kernel(struct dentry *parent,
if (IS_ERR(dentry))
return PTR_ERR(dentry);
ret = bpf_mkobj_ops(dentry, mode, link, &bpf_link_iops,
- &bpf_iter_fops);
+ &bpf_iter_fops, 0);
simple_done_creating(dentry);
return ret;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index e0ed44b1bbcb7..40a5489ac9fea 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1500,6 +1500,12 @@ enum {
/* Enable BPF ringbuf overwrite mode */
BPF_F_RB_OVERWRITE = (1U << 19),
+
+ /* Keep arena pages allocated until the map is destroyed. */
+ BPF_F_ARENA_NO_FREE = (1U << 20),
+
+ /* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */
+ BPF_F_ARENA_EXPORT = (1U << 21),
};
/* Flags for BPF_PROG_QUERY. */
diff --git a/tools/testing/selftests/bpf/.gitignore b/tools/testing/selftests/bpf/.gitignore
index b815bf0d88774..9747ad37d698e 100644
--- a/tools/testing/selftests/bpf/.gitignore
+++ b/tools/testing/selftests/bpf/.gitignore
@@ -47,3 +47,5 @@ verification_cert.h
*.BTF.base
usdt_1
usdt_2
+/arena_kvm_guest-init
+/arena_kvm_host-runner
diff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile
index a22be7efd1fa2..29fffa7857c19 100644
--- a/tools/testing/selftests/bpf/Makefile
+++ b/tools/testing/selftests/bpf/Makefile
@@ -42,7 +42,13 @@ TEST_PROGS := test_kmod.sh \
test_bpftool_build.sh \
test_doc_build.sh \
test_xsk.sh \
- test_xdp_features.sh
+ test_xdp_features.sh \
+ arena_kvm.py
+
+TEST_GEN_FILES += arena_kvm_guest-init arena_kvm_host-runner
+ifneq ($(CLANG_CPUV4),)
+TEST_GEN_FILES += arena_kvm_guest.bpf.o arena_kvm_host.bpf.o
+endif
TEST_PROGS_EXTENDED := \
ima_setup.sh verify_sig_setup.sh
@@ -94,6 +100,25 @@ BPFTOOLDIR := $(TOOLSDIR)/bpf/bpftool
HOST_BPFOBJ := $(HOST_BUILD_DIR)/libbpf/libbpf.a
BPF_TARGET_ENDIAN:=$(if $(IS_LITTLE_ENDIAN),--target=bpfel,--target=bpfeb)
+# These helpers are installed beside arena_kvm.py, including for OUTPUT=.
+$(OUTPUT)/arena_kvm_%.bpf.o: arena_kvm_%.bpf.c arena_kvm_shared.h \
+ libarena/include/bpf_atomic.h \
+ libarena/include/bpf_arena_common.h \
+ $(INCLUDE_DIR)/vmlinux.h $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BPF,,$@)
+ $(Q)$(CLANG) $(BPF_CFLAGS) $(CLANG_CFLAGS) -O2 \
+ $(BPF_TARGET_ENDIAN) -mcpu=v4 -c $< -o $@ $(call skip_on_fail,BPF)
+
+$(OUTPUT)/arena_kvm_guest-init: arena_kvm_guest.c arena_kvm_shared.h \
+ $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BINARY,,$@)
+ $(Q)$(CC) $(CFLAGS) $(LDFLAGS) $< $(BPFOBJ) $(LDLIBS) -lzstd -o $@
+
+$(OUTPUT)/arena_kvm_host-runner: arena_kvm_host.c arena_kvm_shared.h \
+ $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BINARY,,$@)
+ $(Q)$(CC) $(CFLAGS) $(LDFLAGS) $< $(BPFOBJ) $(LDLIBS) -lzstd -o $@
+
NON_CHECK_FEAT_TARGETS := clean docs-clean emit_tests
CHECK_FEAT := $(filter-out $(NON_CHECK_FEAT_TARGETS),$(or $(MAKECMDGOALS), "none"))
ifneq ($(CHECK_FEAT),)
diff --git a/tools/testing/selftests/bpf/arena_kvm.py b/tools/testing/selftests/bpf/arena_kvm.py
new file mode 100755
index 0000000000000..552bfcfb80ac8
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm.py
@@ -0,0 +1,448 @@
+#!/usr/bin/env python3
+# SPDX-License-Identifier: GPL-2.0
+# Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+"""Exchange values between host and guest BPF through NUMA-backed RAM.
+
+Build with the BPF selftests Makefile. Helpers are found in the current
+directory or beside this script, so OUTPUT builds and installed tests work.
+Use --build-dir to select another helper directory and --kernel (or
+ARENA_KVM_KERNEL) to select the kernel booted by virtme-ng. By default, use
+the source tree's kernel when available, otherwise the running release.
+"""
+
+# Dependencies: Python 3.9+, virtme-ng, busybox-static, qemu, util-linux (script).
+
+import argparse
+import array
+import ctypes
+import mmap
+import os
+from pathlib import Path
+import queue
+import re
+import shutil
+import shlex
+import signal
+import socket
+import struct
+import subprocess
+import sys
+import threading
+import time
+
+
+PAGE_SIZE = os.sysconf("SC_PAGE_SIZE")
+KSFT_SKIP = 4
+HELPER_TIMEOUT = 20
+GUEST_MEMORY = 512 * 1024 * 1024
+# RAM starts for virtme-ng's PC/q35, virt (arm64/riscv64), and pseries
+# machines. With 512 MiB of RAM, the last NUMA node follows node 0
+# without crossing a machine's RAM hole. Skip other layouts.
+GUEST_RAM_BASES = {
+ "x86_64": 0,
+ "aarch64": 0x40000000,
+ "arm64": 0x40000000,
+ "riscv64": 0x80000000,
+ "ppc64": 0,
+ "ppc64le": 0,
+}
+# Leave room for guest boot allocations and allocator watermarks.
+# QEMU's pseries machine requires each NUMA node to be 256 MiB aligned.
+SLICE_SIZE = (256 if os.uname().machine in ("ppc64", "ppc64le") else 16) \
+ * 1024 * 1024
+ARENA_SIZE = 2 * SLICE_SIZE
+NUM_PAGES = ARENA_SIZE // PAGE_SIZE
+GUEST_READY = 0x4152454E414B564D
+HOST_FIRST = 0x123456789ABCDEF0
+GUEST_FIRST = 0xFEDCBA9876543210
+HOST_SECOND = 0x1020304050607080
+GUEST_SECOND = 0x8070605040302010
+
+
+def progress(message):
+ print(f"# arena_kvm: {message}", flush=True)
+
+
+class TestSkipped(Exception):
+ pass
+
+
+class TestTerminated(BaseException):
+ def __init__(self, signum):
+ self.signum = signum
+
+
+def terminate(signum, frame):
+ # A second signal must not interrupt cleanup of pins and descendants.
+ signal.signal(signal.SIGTERM, signal.SIG_IGN)
+ signal.signal(signal.SIGINT, signal.SIG_IGN)
+ raise TestTerminated(signum)
+
+
+def enable_subreaper():
+ # Adopt orphaned descendants even when script/vng changes session.
+ libc = ctypes.CDLL(None, use_errno=True)
+ libc.prctl.argtypes = (ctypes.c_int, *([ctypes.c_ulong] * 4))
+ if libc.prctl(36, 1, 0, 0, 0): # PR_SET_CHILD_SUBREAPER
+ err = ctypes.get_errno()
+ raise OSError(err, os.strerror(err))
+
+
+def children(pid):
+ # /proc/PID/task/TID/children requires CONFIG_CHECKPOINT_RESTORE.
+ # PPid is always available and covers children of every parent thread.
+ result = set()
+ for path in Path("/proc").iterdir():
+ if not path.name.isdigit():
+ continue
+ try:
+ status = (path / "status").read_text()
+ except (FileNotFoundError, ProcessLookupError, PermissionError):
+ continue
+ for line in status.splitlines():
+ if line.startswith("PPid:") and int(line.split()[1]) == pid:
+ result.add(int(path.name))
+ break
+ return result
+
+
+def stop_children(processes=()):
+ # Kill direct children, then repeat for descendants adopted by the
+ # subreaper. Sessions and process groups do not affect adoption.
+ # Reaping until ECHILD verifies that the entire owned tree has exited.
+ known = {process.pid: process for process in processes}
+ deadline = time.monotonic() + 10
+ while True:
+ for pid in children(os.getpid()):
+ try:
+ fd = os.pidfd_open(pid)
+ except ProcessLookupError:
+ continue
+ try:
+ signal.pidfd_send_signal(fd, signal.SIGKILL)
+ except ProcessLookupError:
+ pass
+ finally:
+ os.close(fd)
+ while True:
+ try:
+ pid, status = os.waitpid(-1, os.WNOHANG)
+ except ChildProcessError:
+ return
+ if not pid:
+ break
+ if pid in known:
+ known[pid].returncode = os.waitstatus_to_exitcode(status)
+ if time.monotonic() >= deadline:
+ raise TimeoutError("guest descendants did not exit")
+ time.sleep(0.01)
+
+
+def getq(arena, offset):
+ return struct.unpack_from("=Q", arena, offset)[0]
+
+
+def wait_reply(runner, obj, fd, offset, guest, lines):
+ deadline = time.monotonic() + 20
+ while time.monotonic() < deadline:
+ result = run_host_bpf(runner, obj, fd, offset, "check_reply",
+ timeout=max(0.001, deadline - time.monotonic()))
+ if result == 0:
+ return
+ if result != 1:
+ raise RuntimeError(f"guest BPF reply is wrong: result={result}")
+ if guest.poll() is not None:
+ raise RuntimeError("guest exited early:\n" + "".join(lines))
+ time.sleep(0.001)
+ raise TimeoutError(f"timed out waiting for BPF reply at offset {offset}")
+
+
+def read_lines(guest, messages, lines):
+ for line in guest.stdout:
+ lines.append(line)
+ # Serial startup may prefix a message with NULs echoed as ^@.
+ messages.put(re.sub(r"^(?:\x00|\^@)+", "", line.strip()))
+
+
+def wait_message(messages, lines, expected, guest):
+ deadline = time.monotonic() + 45
+ while time.monotonic() < deadline:
+ try:
+ line = messages.get(timeout=0.5)
+ except queue.Empty:
+ if guest.poll() is not None:
+ break
+ continue
+ if expected in line:
+ return line
+ if "GUEST_BPF_FAILED" in line:
+ break
+ if "GUEST_BPF_SKIPPED" in line:
+ raise TestSkipped(line)
+ raise RuntimeError(f"guest did not print {expected} "
+ f"(launcher exit status: {guest.poll()}):\n" +
+ "".join(lines))
+
+
+def run_host_bpf(runner, obj, fd, offset, program="exchange",
+ timeout=HELPER_TIMEOUT):
+ output = subprocess.check_output(
+ [str(runner), str(fd), str(obj), str(offset), program],
+ pass_fds=(fd,), text=True, timeout=timeout)
+ return int(output)
+
+
+def resolve_vng():
+ executable = shutil.which("vng")
+ if not executable or not os.access(executable, os.X_OK):
+ print("SKIP: virtme-ng (vng) is unavailable")
+ return None, None
+ return executable, os.environ.copy()
+
+
+def start_guest(pin, build, kernel, index, vng, env, guests):
+ node0_size = GUEST_MEMORY - SLICE_SIZE
+ base = GUEST_RAM_BASES[os.uname().machine] + node0_size
+ progress(f"Starting guest {index}: {SLICE_SIZE // (1024 * 1024)} MiB "
+ f"slice at file offset {index * SLICE_SIZE:#x}, "
+ f"NUMA node 1 physical base {base:#x}")
+ qemu_opts = (
+ f"-object memory-backend-file,id=signal,size={SLICE_SIZE},share=on,"
+ f"offset={index * SLICE_SIZE},mem-path={pin} "
+ "-numa node,nodeid=1,memdev=signal"
+ )
+ guest_command = [str(build / "arena_kvm_guest-init"),
+ str(build / "arena_kvm_guest.bpf.o")]
+ command = [
+ vng, "--verbose", "--run", str(kernel), "--cpus", "2",
+ "--memory", f"{GUEST_MEMORY // (1024 * 1024)}M",
+ "--numa", str(node0_size),
+ "--exec",
+ shlex.join(guest_command),
+ "--append", f"arena_kvm.base={base:#x} arena_kvm.size={SLICE_SIZE}",
+ f"--qemu-opts={qemu_opts}",
+ ]
+ lines = []
+ messages = queue.Queue()
+ guest = subprocess.Popen(["script", "-e", "-q", "-c", shlex.join(command),
+ "/dev/null"], stdout=subprocess.PIPE,
+ stderr=subprocess.STDOUT, text=True, env=env,
+ start_new_session=True)
+ # Register ownership before thread construction or startup can fail.
+ guests.append((guest, messages, lines, None))
+ reader = threading.Thread(target=read_lines, args=(guest, messages, lines),
+ daemon=True)
+ guests[-1] = guest, messages, lines, reader
+ reader.start()
+
+
+def stop_guests(guests):
+ stop_children(process for process, _, _, _ in guests)
+ for process, _, _, reader in guests:
+ process.wait(timeout=10)
+ if reader is not None and reader.ident is not None:
+ reader.join(timeout=2)
+ if reader.is_alive():
+ raise RuntimeError("guest output reader did not exit")
+ process.stdout.close()
+
+
+def exchange(pin, fd, arena, build, kernel, vng, env):
+ guests = []
+ try:
+ for index in range(2):
+ start_guest(pin, build, kernel, index, vng, env, guests)
+ offsets = []
+ for index, (process, messages, lines, _) in enumerate(guests):
+ progress(f"Waiting for guest {index} to allocate its BPF arena page "
+ "and check that BPF cannot free it")
+ ready = wait_message(messages, lines, "GUEST_BPF_READY", process)
+ match = re.fullmatch(r"GUEST_BPF_READY (\d+) (\d+)", ready)
+ if not match:
+ raise RuntimeError(f"malformed guest page offset: {ready}")
+ relative = int(match.group(1))
+ if int(match.group(2)) != PAGE_SIZE:
+ raise RuntimeError("host and guest page sizes differ")
+ offset = index * SLICE_SIZE + relative
+ if relative >= SLICE_SIZE or relative % PAGE_SIZE or \
+ getq(arena, offset) != GUEST_READY:
+ raise RuntimeError(f"invalid guest {index} page offset {relative}")
+ offsets.append(offset)
+ progress(f"Guest {index} ready: page offset {relative:#x} in slice, "
+ f"{offset:#x} in host arena")
+ # The QEMU VMA and the open map FD retain the arena after unlink.
+ progress("Unpinning the host arena while both guests retain mappings")
+ os.unlink(pin)
+ runner = build / "arena_kvm_host-runner"
+ obj = build / "arena_kvm_host.bpf.o"
+ for offset in offsets:
+ if run_host_bpf(runner, obj, fd, offset, "probe_free") != 0 or \
+ getq(arena, offset) != GUEST_READY:
+ raise RuntimeError("host BPF freed a shared RAM page")
+ progress("Host BPF page-free checks passed for both shared pages")
+
+ for index, offset in enumerate(offsets):
+ progress(f"Round 1: host -> guest {index}, value {HOST_FIRST:#018x}")
+ result = run_host_bpf(runner, obj, fd, offset)
+ if result != 0:
+ raise RuntimeError(f"guest {index} first result={result}")
+ other = offsets[1 - index]
+ if index == 0 and getq(arena, other + 16) != 0:
+ raise RuntimeError("guest 0 write modified guest 1 signal")
+ if index == 0:
+ progress("Guest 1 signal unchanged after host write to guest 0")
+
+ for index, (process, messages, lines, _) in enumerate(guests):
+ offset = offsets[index]
+ wait_reply(runner, obj, fd, offset, process, lines)
+ progress(f"Round 1: guest {index} -> host, "
+ f"verified value {GUEST_FIRST:#018x}")
+
+ # Guest 0 exits while guest 1 keeps its slice mapped and active.
+ for index, (process, messages, lines, _) in enumerate(guests):
+ offset = offsets[index]
+ progress(f"Round 2: host -> guest {index}, value {HOST_SECOND:#018x}")
+ if run_host_bpf(runner, obj, fd, offset) != 0:
+ raise RuntimeError(f"host BPF failed guest {index} second value")
+ wait_reply(runner, obj, fd, offset, process, lines)
+ progress(f"Round 2: guest {index} -> host, "
+ f"verified value {GUEST_SECOND:#018x}")
+ if run_host_bpf(runner, obj, fd, offset) != 2:
+ raise RuntimeError(f"host BPF failed guest {index} reply")
+ wait_message(messages, lines, "GUEST_BPF_EXCHANGED", process)
+ process.wait(timeout=15)
+ if process.returncode:
+ raise RuntimeError(f"guest {index} failed:\n" + "".join(lines))
+ progress(f"Guest {index} completed and exited successfully")
+ if index == 0:
+ progress("Continuing with guest 1 after guest 0 exits")
+ print(f"OK: two guests exchanged BPF values through arena slices {offsets}")
+ finally:
+ progress("Cleaning up guest processes")
+ stop_guests(guests)
+
+
+def skip(reason):
+ print(f"SKIP: {reason}")
+ return KSFT_SKIP
+
+
+def create_arena(pin, build):
+ parent, child = socket.socketpair()
+ with parent, child:
+ result = subprocess.run(
+ [str(build / "arena_kvm_host-runner"), "--create", pin,
+ str(child.fileno()), str(NUM_PAGES)], pass_fds=(child.fileno(),),
+ timeout=HELPER_TIMEOUT)
+ child.close()
+ if result.returncode == KSFT_SKIP:
+ return None
+ result.check_returncode()
+ fds = array.array("i")
+ _, ancillary, flags, _ = parent.recvmsg(
+ 1, socket.CMSG_SPACE(fds.itemsize))
+ for level, kind, data in ancillary:
+ if level == socket.SOL_SOCKET and kind == socket.SCM_RIGHTS:
+ fds.frombytes(data[:len(data) - len(data) % fds.itemsize])
+ if flags & socket.MSG_CTRUNC or len(fds) != 1:
+ for fd in fds:
+ os.close(fd)
+ raise RuntimeError("host helper did not send one arena FD")
+ return fds[0]
+
+
+def main():
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--build-dir", type=Path,
+ help="directory containing the built helpers")
+ parser.add_argument("--kernel", default=os.environ.get("ARENA_KVM_KERNEL"),
+ help="virtme-ng kernel path or release (default: source "
+ "tree kernel, otherwise running kernel)")
+ args = parser.parse_args()
+ if not hasattr(os, "pidfd_open") or not hasattr(signal, "pidfd_send_signal"):
+ return skip("Python 3.9+ with Linux pidfd support is required")
+ machine = os.uname().machine
+ if machine not in GUEST_RAM_BASES:
+ return skip(f"unknown guest RAM layout for architecture {machine}")
+ if struct.calcsize("P") != 8:
+ return skip("BPF arenas require a supported 64-bit architecture")
+ if os.geteuid() != 0:
+ return skip("this test requires root")
+ if SLICE_SIZE % PAGE_SIZE:
+ return skip("arena slices must contain a whole number of pages")
+ if not os.path.ismount("/sys/fs/bpf"):
+ return skip("mount bpffs first")
+ directory = Path(__file__).resolve().parent
+ build = args.build_dir
+ if build is None:
+ build = Path.cwd() if (Path.cwd() / "arena_kvm_host-runner").is_file() \
+ else directory
+ build = build.resolve()
+ kernel = args.kernel
+ if kernel is None:
+ kernel = next((str(p) for p in directory.parents
+ if (p / "vmlinux").is_file()), os.uname().release)
+ vng, env = resolve_vng()
+ if not vng:
+ return KSFT_SKIP
+ qemu_arch = {"arm64": "aarch64", "ppc64le": "ppc64"}.get(machine, machine)
+ for tool in ("script", f"qemu-system-{qemu_arch}"):
+ if not shutil.which(tool, path=env.get("PATH")):
+ return skip(f"{tool} is unavailable")
+ for name in ("arena_kvm_guest-init", "arena_kvm_guest.bpf.o",
+ "arena_kvm_host-runner", "arena_kvm_host.bpf.o"):
+ if not (build / name).exists():
+ return skip(f"missing {build / name}; build the BPF selftests")
+ for name in ("arena_kvm_host.bpf.o", "arena_kvm_guest.bpf.o"):
+ result = subprocess.run([str(build / "arena_kvm_host-runner"),
+ "--check-bpf", str(build / name)],
+ timeout=HELPER_TIMEOUT)
+ if result.returncode == KSFT_SKIP:
+ return skip("BPF compiler lacks arena address-space conversions")
+ result.check_returncode()
+ if subprocess.run([str(build / "arena_kvm_host-runner"),
+ "--check-kvm"], timeout=HELPER_TIMEOUT).returncode:
+ return skip("KVM is unavailable or its API version is unsupported")
+ pin = f"/sys/fs/bpf/arena_kvm_{os.getpid()}"
+ progress(f"Architecture: {machine}; "
+ f"page size: {PAGE_SIZE} bytes")
+ progress(f"Creating and pinning a {ARENA_SIZE // (1024 * 1024)} MiB "
+ f"host arena at {pin}")
+ try:
+ fd = create_arena(pin, build)
+ if fd is None:
+ return skip("pinned BPF arenas are unsupported")
+ try:
+ if os.stat(pin).st_size != ARENA_SIZE:
+ raise RuntimeError("pinned arena has the wrong size")
+ with mmap.mmap(fd, ARENA_SIZE, flags=mmap.MAP_SHARED,
+ prot=mmap.PROT_READ | mmap.PROT_WRITE) as arena:
+ progress(f"Populating {NUM_PAGES} host arena pages before "
+ "starting the guests")
+ # Populate the file pages before QEMU uses them as RAM.
+ for i in range(NUM_PAGES):
+ arena[i * PAGE_SIZE]
+ exchange(pin, fd, arena, build, kernel, vng, env)
+ finally:
+ os.close(fd)
+ finally:
+ if os.path.exists(pin):
+ os.unlink(pin)
+
+
+if __name__ == "__main__":
+ enable_subreaper()
+ signal.signal(signal.SIGTERM, terminate)
+ signal.signal(signal.SIGINT, terminate)
+ try:
+ status = main()
+ except TestSkipped as err:
+ status = skip(str(err))
+ except TestTerminated as err:
+ progress(f"Terminated by signal {err.signum}")
+ status = 128 + err.signum
+ finally:
+ signal.signal(signal.SIGTERM, signal.SIG_IGN)
+ signal.signal(signal.SIGINT, signal.SIG_IGN)
+ stop_children()
+ sys.exit(status)
diff --git a/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c b/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c
new file mode 100644
index 0000000000000..461562149c57f
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c
@@ -0,0 +1,89 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+#include "bpf_arena_common.h"
+#ifdef __TARGET_ARCH_powerpc
+/* PowerPC does not support arena load-acquire/store-release instructions. */
+#undef __BPF_FEATURE_LOAD_ACQ_STORE_REL
+#endif
+#include "bpf_atomic.h"
+#endif
+#include "arena_kvm_shared.h"
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARENA);
+ __uint(map_flags, BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE);
+ __uint(max_entries, 1);
+} arena SEC(".maps");
+
+volatile __u32 signal_offset;
+
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+
+SEC("syscall")
+int allocate(void *ctx)
+{
+ struct signal_page __arena *page;
+
+ page = bpf_arena_alloc_pages(&arena, NULL, 1, ARENA_KVM_NODE, 0);
+ if (!page)
+ return 1;
+ signal_offset = (__u32)((__u64)page - (__u64)arena_base(&arena));
+ smp_store_release(&page->ready, GUEST_READY);
+ return 0;
+}
+
+SEC("syscall")
+int probe_free(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ bpf_arena_free_pages(&arena, page, 1);
+ return page->ready == GUEST_READY ? 0 : 1;
+}
+
+SEC("syscall")
+int exchange(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = smp_load_acquire(&page->h2g_seq);
+ __u64 payload;
+
+ if ((seq != 1 && seq != 2) || page->g2h_seq == seq)
+ return 1;
+ payload = page->h2g_payload;
+ if (seq == 1 && payload != HOST_FIRST)
+ return 2;
+ if (seq == 2 && payload != HOST_SECOND)
+ return 3;
+ page->g2h_payload = seq == 1 ? GUEST_FIRST : GUEST_SECOND;
+ smp_store_release(&page->g2h_seq, seq);
+ return 0;
+}
+
+SEC("syscall")
+int complete(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ return smp_load_acquire(&page->h2g_seq) == 3 ? 0 : 1;
+}
+
+#else
+SEC("syscall") int allocate(void *ctx) { return 1; }
+SEC("syscall") int probe_free(void *ctx) { return 1; }
+SEC("syscall") int exchange(void *ctx) { return 1; }
+SEC("syscall") int complete(void *ctx) { return 1; }
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/arena_kvm_guest.c b/tools/testing/selftests/bpf/arena_kvm_guest.c
new file mode 100644
index 0000000000000..da6586ae660d1
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_guest.c
@@ -0,0 +1,224 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ *
+ * Guest helper for a BPF arena allocated from shared NUMA RAM.
+ */
+#define __EXPORTED_HEADERS__
+
+#include <bpf/bpf.h>
+#include <bpf/libbpf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <linux/bpf.h>
+#include <linux/mempolicy.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/mount.h>
+#include <sys/syscall.h>
+#include <unistd.h>
+
+#include "arena_kvm_shared.h"
+
+static int signal_file_offset(void *arena, uint64_t *offset)
+{
+ char cmdline[8192], *arg, *end;
+ uint64_t entry, pfn, gpa, base = 0, size = 0, *value;
+ ssize_t n;
+ int fd, node, found = 0;
+
+ /* The launcher supplies QEMU's guest physical base for this slice. */
+ fd = open("/proc/cmdline", O_RDONLY);
+ if (fd < 0)
+ return -errno;
+ n = read(fd, cmdline, sizeof(cmdline) - 1);
+ close(fd);
+ if (n <= 0 || n == sizeof(cmdline) - 1)
+ return -EIO;
+ cmdline[n] = '\0';
+ for (arg = strtok(cmdline, "\n "); arg; arg = strtok(NULL, "\n ")) {
+ if (!strncmp(arg, "arena_kvm.base=", 15)) {
+ value = &base;
+ found |= 1;
+ } else if (!strncmp(arg, "arena_kvm.size=", 15)) {
+ value = &size;
+ found |= 2;
+ } else {
+ continue;
+ }
+ errno = 0;
+ *value = strtoull(arg + 15, &end, 0);
+ if (errno || *end || end == arg + 15 || *value % getpagesize())
+ return -EINVAL;
+ }
+ if (found != 3 || size < getpagesize())
+ return -EINVAL;
+
+ /* Fault the BPF-allocated page into the guest user VMA. */
+ if (!*(volatile uint64_t *)arena)
+ return -EINVAL;
+ if (syscall(SYS_get_mempolicy, &node, NULL, 0, arena,
+ MPOL_F_NODE | MPOL_F_ADDR))
+ return -errno;
+ if (node != ARENA_KVM_NODE)
+ return -EXDEV;
+ fd = open("/proc/self/pagemap", O_RDONLY);
+ if (fd < 0)
+ return -errno;
+ n = pread(fd, &entry, sizeof(entry),
+ (uintptr_t)arena / getpagesize() * sizeof(entry));
+ close(fd);
+ if (n != sizeof(entry) || !(entry & (1ULL << 63)))
+ return -EIO;
+ pfn = entry & ((1ULL << 55) - 1);
+ if (!pfn)
+ return -EPERM;
+ gpa = pfn * getpagesize();
+ if (gpa < base || gpa - base >= size || size - (gpa - base) < getpagesize())
+ return -ERANGE;
+ *offset = gpa - base;
+ return 0;
+}
+
+static int create_arena(void)
+{
+ union bpf_attr attr = {
+ .map_type = BPF_MAP_TYPE_ARENA,
+ .max_entries = 1,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE,
+ };
+
+ return syscall(__NR_bpf, BPF_MAP_CREATE, &attr, sizeof(attr));
+}
+
+static int run_prog(int fd, unsigned int *retval)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ int err = bpf_prog_test_run_opts(fd, &opts);
+
+ if (err)
+ return err;
+ *retval = opts.retval;
+ return 0;
+}
+
+static int exchange_once(int fd)
+{
+ unsigned int result = 0;
+ int err;
+
+ /* The host bounds each protocol stage and stops the guest on timeout. */
+ for (;;) {
+ err = run_prog(fd, &result);
+ if (err || result > 1)
+ return err ? err : -EINVAL;
+ if (!result)
+ return 0;
+ usleep(1000);
+ }
+}
+
+int main(int argc, char **argv)
+{
+ struct bpf_object *obj = NULL;
+ struct bpf_program *alloc, *probe_free, *exchange, *complete;
+ struct bpf_map *map;
+ void *arena = MAP_FAILED;
+ uint64_t file_offset;
+ unsigned int result = 0;
+ int fd = -1;
+ int err;
+
+ setbuf(stdout, NULL);
+ if (argc != 2 ||
+ (mount("sysfs", "/sys", "sysfs", 0, NULL) && errno != EBUSY) ||
+ (mount("proc", "/proc", "proc", 0, NULL) && errno != EBUSY))
+ goto fail;
+ if (access("/sys/devices/system/node/node1", F_OK)) {
+ fputs("guest NUMA node 1 is unavailable\n", stderr);
+ goto fail;
+ }
+ fd = create_arena();
+ if (fd < 0) {
+ perror("create guest arena");
+ goto fail;
+ }
+ arena = mmap(NULL, getpagesize(), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
+ if (arena == MAP_FAILED) {
+ perror("mmap guest arena");
+ goto fail;
+ }
+ obj = bpf_object__open_file(argv[1], NULL);
+ if (!obj)
+ goto fail;
+ err = arena_kvm_check_features(obj);
+ if (err == 4) {
+ puts("GUEST_BPF_SKIPPED: compiler lacks arena address-space casts");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ close(fd);
+ return 4;
+ }
+ if (err)
+ goto fail;
+ map = bpf_object__find_map_by_name(obj, "arena");
+ alloc = bpf_object__find_program_by_name(obj, "allocate");
+ probe_free = bpf_object__find_program_by_name(obj, "probe_free");
+ exchange = bpf_object__find_program_by_name(obj, "exchange");
+ complete = bpf_object__find_program_by_name(obj, "complete");
+ if (!map || !alloc || !probe_free || !exchange || !complete ||
+ bpf_map__reuse_fd(map, fd))
+ goto fail;
+ close(fd);
+ fd = -1;
+ err = bpf_object__load(obj);
+ if (err) {
+ fprintf(stderr, "load guest BPF: %d\n", err);
+ goto fail;
+ }
+ err = run_prog(bpf_program__fd(alloc), &result);
+ if (err || result) {
+ fprintf(stderr, "allocate guest page: err=%d result=%u\n",
+ err, result);
+ goto fail;
+ }
+ err = run_prog(bpf_program__fd(probe_free), &result);
+ if (err || result) {
+ fprintf(stderr, "guest page lifetime check: err=%d result=%u\n",
+ err, result);
+ goto fail;
+ }
+ err = signal_file_offset(arena, &file_offset);
+ if (err == -EXDEV) {
+ puts("GUEST_BPF_SKIPPED: allocation fell back from shared NUMA node");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ return 4;
+ }
+ if (err) {
+ fprintf(stderr, "resolve shared page offset: %d\n", err);
+ goto fail;
+ }
+ printf("GUEST_BPF_READY %llu %d\n",
+ (unsigned long long)file_offset, getpagesize());
+ if (exchange_once(bpf_program__fd(exchange)) ||
+ exchange_once(bpf_program__fd(exchange)) ||
+ exchange_once(bpf_program__fd(complete)))
+ goto fail;
+ puts("GUEST_BPF_EXCHANGED");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ return 0;
+
+fail:
+ puts("GUEST_BPF_FAILED");
+ bpf_object__close(obj);
+ if (arena != MAP_FAILED)
+ munmap(arena, getpagesize());
+ if (fd >= 0)
+ close(fd);
+ return 1;
+}
diff --git a/tools/testing/selftests/bpf/arena_kvm_host.bpf.c b/tools/testing/selftests/bpf/arena_kvm_host.bpf.c
new file mode 100644
index 0000000000000..9c1c535fe429b
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_host.bpf.c
@@ -0,0 +1,89 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+#include "bpf_arena_common.h"
+#ifdef __TARGET_ARCH_powerpc
+/* PowerPC does not support arena load-acquire/store-release instructions. */
+#undef __BPF_FEATURE_LOAD_ACQ_STORE_REL
+#endif
+#include "bpf_atomic.h"
+#endif
+#include "arena_kvm_shared.h"
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARENA);
+ __uint(map_flags, BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE | BPF_F_ARENA_EXPORT);
+ /* The host runner reuses the map created with the selected capacity. */
+ __uint(max_entries, 1);
+} arena SEC(".maps");
+
+volatile __u32 signal_offset;
+
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+
+SEC("syscall")
+int exchange(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = page->h2g_seq;
+
+ if (smp_load_acquire(&page->ready) != GUEST_READY)
+ return 4;
+ if (!seq) {
+ page->h2g_payload = HOST_FIRST;
+ smp_store_release(&page->h2g_seq, 1);
+ return 0;
+ }
+ if (seq == 1 && smp_load_acquire(&page->g2h_seq) == 1) {
+ if (page->g2h_payload != GUEST_FIRST)
+ return 3;
+ page->h2g_payload = HOST_SECOND;
+ smp_store_release(&page->h2g_seq, 2);
+ return 0;
+ }
+ if (seq == 2 && smp_load_acquire(&page->g2h_seq) == 2) {
+ if (page->g2h_payload != GUEST_SECOND)
+ return 3;
+ smp_store_release(&page->h2g_seq, 3);
+ return 2;
+ }
+ return 1;
+}
+
+SEC("syscall")
+int check_reply(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = page->h2g_seq;
+
+ if ((seq != 1 && seq != 2) || smp_load_acquire(&page->g2h_seq) != seq)
+ return 1;
+ return page->g2h_payload == (seq == 1 ? GUEST_FIRST : GUEST_SECOND) ? 0 : 3;
+}
+
+SEC("syscall")
+int probe_free(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ bpf_arena_free_pages(&arena, page, 1);
+ return page->ready == GUEST_READY ? 0 : 1;
+}
+
+#else
+SEC("syscall") int exchange(void *ctx) { return 1; }
+SEC("syscall") int check_reply(void *ctx) { return 1; }
+SEC("syscall") int probe_free(void *ctx) { return 1; }
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/arena_kvm_host.c b/tools/testing/selftests/bpf/arena_kvm_host.c
new file mode 100644
index 0000000000000..90a77a54bd6f4
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_host.c
@@ -0,0 +1,147 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <bpf/bpf.h>
+#include <bpf/libbpf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <linux/kvm.h>
+#include <limits.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/ioctl.h>
+#include <sys/socket.h>
+#include <unistd.h>
+
+#include "arena_kvm_shared.h"
+
+static int create_arena(const char *pin, int socket_fd, unsigned int pages)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT);
+ char control[CMSG_SPACE(sizeof(int))] = {};
+ char byte = 0;
+ struct iovec iov = { .iov_base = &byte, .iov_len = 1 };
+ struct msghdr msg = {
+ .msg_iov = &iov,
+ .msg_iovlen = 1,
+ .msg_control = control,
+ .msg_controllen = sizeof(control),
+ };
+ struct cmsghdr *cmsg = CMSG_FIRSTHDR(&msg);
+ int fd, err = 1;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_kvm", 0, 0,
+ pages, &opts);
+ if (fd < 0) {
+ if (errno == EOPNOTSUPP || errno == EINVAL)
+ return 4; /* KSFT_SKIP: arena type or flags unsupported. */
+ perror("create host arena");
+ return 1;
+ }
+ if (bpf_obj_pin(fd, pin)) {
+ perror("pin host arena");
+ goto out;
+ }
+ cmsg->cmsg_level = SOL_SOCKET;
+ cmsg->cmsg_type = SCM_RIGHTS;
+ cmsg->cmsg_len = CMSG_LEN(sizeof(fd));
+ memcpy(CMSG_DATA(cmsg), &fd, sizeof(fd));
+ if (sendmsg(socket_fd, &msg, 0) == 1)
+ err = 0;
+ else
+ perror("send arena FD");
+out:
+ close(fd);
+ return err;
+}
+
+int main(int argc, char **argv)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ struct bpf_object *obj = NULL;
+ struct bpf_program *prog;
+ struct bpf_map *arena, *bss;
+ uint32_t key = 0, offset;
+ unsigned long parsed;
+ int fd, err;
+ char *end;
+
+ if (argc == 3 && !strcmp(argv[1], "--check-bpf")) {
+ obj = bpf_object__open_file(argv[2], NULL);
+ if (!obj)
+ return 1;
+ err = arena_kvm_check_features(obj);
+ bpf_object__close(obj);
+ return err;
+ }
+
+ if (argc == 2 && !strcmp(argv[1], "--check-kvm")) {
+ fd = open("/dev/kvm", O_RDWR);
+ if (fd < 0)
+ return 4;
+ err = ioctl(fd, KVM_GET_API_VERSION, 0);
+ close(fd);
+ return err == KVM_API_VERSION ? 0 : 4;
+ }
+ if (argc == 5 && !strcmp(argv[1], "--create")) {
+ errno = 0;
+ parsed = strtoul(argv[3], &end, 0);
+ if (errno || *end || end == argv[3] || parsed > INT_MAX)
+ return 1;
+ fd = parsed;
+ errno = 0;
+ parsed = strtoul(argv[4], &end, 0);
+ if (errno || *end || end == argv[4] || !parsed || parsed > UINT_MAX)
+ return 1;
+ return create_arena(argv[2], fd, parsed);
+ }
+ if (argc != 5)
+ return 1;
+ errno = 0;
+ parsed = strtoul(argv[1], &end, 0);
+ if (errno || *end || parsed > INT_MAX)
+ return 1;
+ fd = dup(parsed);
+ if (fd < 0)
+ return 1;
+ errno = 0;
+ parsed = strtoul(argv[3], &end, 0);
+ if (errno || *end || parsed > UINT_MAX || parsed % getpagesize())
+ goto fail;
+ offset = parsed;
+ obj = bpf_object__open_file(argv[2], NULL);
+ if (!obj)
+ goto fail;
+ arena = bpf_object__find_map_by_name(obj, "arena");
+ bss = bpf_object__find_map_by_name(obj, ".bss");
+ prog = bpf_object__find_program_by_name(obj, argv[4]);
+ if (!arena || !bss || !prog || bpf_map__reuse_fd(arena, fd))
+ goto fail;
+ if (offset >= (uint64_t)bpf_map__max_entries(arena) * getpagesize())
+ goto fail;
+ if (bpf_object__load(obj)) {
+ fputs("load host BPF failed\n", stderr);
+ goto fail;
+ }
+ if (bpf_map_update_elem(bpf_map__fd(bss), &key, &offset, BPF_ANY)) {
+ perror("set signal offset");
+ goto fail;
+ }
+ err = bpf_prog_test_run_opts(bpf_program__fd(prog), &opts);
+ if (err)
+ goto fail;
+ printf("%u\n", opts.retval);
+ bpf_object__close(obj);
+ close(fd);
+ return 0;
+
+fail:
+ bpf_object__close(obj);
+ close(fd);
+ return 1;
+}
diff --git a/tools/testing/selftests/bpf/arena_kvm_shared.h b/tools/testing/selftests/bpf/arena_kvm_shared.h
new file mode 100644
index 0000000000000..63aafac25efc3
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_shared.h
@@ -0,0 +1,60 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#ifndef ARENA_KVM_SHARED_H
+#define ARENA_KVM_SHARED_H
+
+/* These flags may be absent from a pre-series kernel's vmlinux.h. */
+#ifndef BPF_F_ARENA_NO_FREE
+#define BPF_F_ARENA_NO_FREE (1U << 20)
+#endif
+#ifndef BPF_F_ARENA_EXPORT
+#define BPF_F_ARENA_EXPORT (1U << 21)
+#endif
+
+struct arena_kvm_features {
+ unsigned int addr_space_cast;
+};
+
+#ifdef __BPF__
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+const volatile struct arena_kvm_features features = { .addr_space_cast = 1 };
+#else
+const volatile struct arena_kvm_features features = {};
+#endif
+#else
+/* Inspect compiler support before loading maps or booting either guest. */
+static inline int arena_kvm_check_features(struct bpf_object *obj)
+{
+ const struct arena_kvm_features *features;
+ struct bpf_map *map;
+ size_t size;
+
+ map = bpf_object__find_map_by_name(obj, ".rodata");
+ if (!map)
+ return 1;
+ features = bpf_map__initial_value(map, &size);
+ if (!features || size != sizeof(*features))
+ return 1;
+ return features->addr_space_cast ? 0 : 4; /* KSFT_SKIP */
+}
+#endif
+
+#define ARENA_KVM_NODE 1
+
+#define GUEST_READY 0x4152454e414b564dULL
+#define HOST_FIRST 0x123456789abcdef0ULL
+#define GUEST_FIRST 0xfedcba9876543210ULL
+#define HOST_SECOND 0x1020304050607080ULL
+#define GUEST_SECOND 0x8070605040302010ULL
+
+struct signal_page {
+ volatile unsigned long long ready;
+ volatile unsigned long long h2g_payload;
+ volatile unsigned long long h2g_seq;
+ volatile unsigned long long g2h_payload;
+ volatile unsigned long long g2h_seq;
+};
+
+#endif
diff --git a/tools/testing/selftests/bpf/prog_tests/arena_pinned.c b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c
new file mode 100644
index 0000000000000..8b7ca90211296
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c
@@ -0,0 +1,442 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <test_progs.h>
+#include <fcntl.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <sys/wait.h>
+
+#define ARENA_PAGES 4
+
+struct arena_file {
+ int map_fd;
+ int file_fd;
+ char pin[128];
+ void *canonical;
+ size_t len;
+ bool pinned;
+};
+
+static void cleanup(struct arena_file *file)
+{
+ if (file->canonical != MAP_FAILED)
+ munmap(file->canonical, file->len);
+ if (file->file_fd >= 0)
+ close(file->file_fd);
+ if (file->pinned)
+ unlink(file->pin);
+ if (file->map_fd >= 0)
+ close(file->map_fd);
+}
+
+static void reject_mapping(int fd, size_t len, int flags, off_t offset,
+ const char *name)
+{
+ void *addr;
+ int err;
+
+ addr = mmap(NULL, len, PROT_READ | PROT_WRITE, flags, fd, offset);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, name))
+ ASSERT_EQ(err, EINVAL, "mmap_errno");
+ else
+ munmap(addr, len);
+}
+
+static bool setup(struct arena_file *file, __u64 map_extra, bool zero)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT,
+ .map_extra = map_extra);
+
+ memset(file, 0, sizeof(*file));
+ file->map_fd = -1;
+ file->file_fd = -1;
+ file->canonical = MAP_FAILED;
+ file->len = ARENA_PAGES * getpagesize();
+ snprintf(file->pin, sizeof(file->pin), "/sys/fs/bpf/arena_pinned_%d", getpid());
+ file->map_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_pinned",
+ 0, 0, ARENA_PAGES, &opts);
+ if (file->map_fd < 0 && errno == EOPNOTSUPP) {
+ test__skip();
+ return false;
+ }
+ if (!ASSERT_GE(file->map_fd, 0, "map_create"))
+ return false;
+ if (!ASSERT_OK(bpf_obj_pin(file->map_fd, file->pin), "obj_pin"))
+ return false;
+ file->pinned = true;
+ file->file_fd = open(file->pin, O_RDWR);
+ if (!ASSERT_GE(file->file_fd, 0, "open_pin"))
+ return false;
+ reject_mapping(file->map_fd, getpagesize(), MAP_SHARED, 0, "short_canonical");
+ reject_mapping(file->map_fd, file->len + getpagesize(), MAP_SHARED,
+ 0, "large_canonical");
+ if (map_extra)
+ return true;
+
+ reject_mapping(file->file_fd, getpagesize(), MAP_SHARED, 0, "uninitialized");
+ file->canonical = mmap(NULL, file->len, PROT_READ | PROT_WRITE,
+ MAP_SHARED | (zero ? MAP_FIXED_NOREPLACE : 0),
+ file->map_fd, 0);
+ if (zero && file->canonical == MAP_FAILED &&
+ (errno == EPERM || errno == EACCES)) {
+ test__skip();
+ return false;
+ }
+ return ASSERT_NEQ(file->canonical, MAP_FAILED, "canonical_mmap");
+}
+
+static void test_mappings(__u64 map_extra)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ struct stat st;
+ char *alias = MAP_FAILED;
+ char *full = MAP_FAILED;
+ int expected = 91;
+
+ if (!setup(&file, map_extra, false))
+ goto out;
+ if (!ASSERT_OK(fstat(file.file_fd, &st), "stat_pin"))
+ goto out;
+ ASSERT_EQ(st.st_size, ARENA_PAGES * ps, "pin_size");
+ reject_mapping(file.file_fd, 8ULL << 30, MAP_SHARED, 0, "oversized");
+ reject_mapping(file.file_fd, ps, MAP_SHARED, ARENA_PAGES * ps, "past_capacity");
+ reject_mapping(file.file_fd, ps * 2, MAP_SHARED,
+ (ARENA_PAGES - 1) * ps, "past_capacity_end");
+ reject_mapping(file.file_fd, ps, MAP_PRIVATE, 0, "private_mapping");
+ full = mmap(NULL, st.st_size, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, 0);
+ if (!ASSERT_NEQ(full, MAP_FAILED, "full_file_mmap"))
+ goto out;
+ full[0] = 17;
+ full[file.len - 1] = 62;
+
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, (ARENA_PAGES - 1) * ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "last_page_slice"))
+ goto out;
+ alias[0] = 91;
+ ASSERT_OK(msync(alias, ps, MS_SYNC), "msync_slice");
+ ASSERT_OK(fsync(file.file_fd), "fsync_pin");
+ ASSERT_OK(fdatasync(file.file_fd), "fdatasync_pin");
+ ASSERT_EQ(alias[ps - 1], 62, "full_file_last_page");
+ if (file.canonical != MAP_FAILED) {
+ char *canonical = file.canonical;
+
+ ASSERT_EQ(canonical[0], 17, "full_file_first_page");
+ ASSERT_EQ(canonical[(ARENA_PAGES - 1) * ps], 91, "alias_write");
+ canonical[(ARENA_PAGES - 1) * ps] = 42;
+ ASSERT_EQ(alias[0], 42, "canonical_write");
+ expected = 42;
+ if (!ASSERT_OK(munmap(file.canonical, file.len), "unmap_canonical"))
+ goto out;
+ file.canonical = MAP_FAILED;
+ }
+ if (!ASSERT_OK(munmap(full, file.len), "unmap_full_file"))
+ goto out;
+ full = MAP_FAILED;
+ if (!ASSERT_OK(unlink(file.pin), "unlink_pin"))
+ goto out;
+ file.pinned = false;
+ close(file.file_fd);
+ file.file_fd = -1;
+ close(file.map_fd);
+ file.map_fd = -1;
+ ASSERT_EQ(alias[0], expected, "unpin_lifetime");
+ alias[0] = 73;
+out:
+ if (full != MAP_FAILED)
+ munmap(full, file.len);
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_unexported_pin(__u32 flags)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | flags);
+ char pin[128];
+ struct stat st;
+ char *canonical = MAP_FAILED;
+ size_t ps = getpagesize();
+ int map_fd, fd, err;
+
+ map_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_unexported",
+ 0, 0, ARENA_PAGES, &opts);
+ if (map_fd < 0 && errno == EOPNOTSUPP) {
+ test__skip();
+ return;
+ }
+ if (!ASSERT_GE(map_fd, 0, "unexported_create"))
+ return;
+ canonical = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, map_fd, 0);
+ if (!ASSERT_NEQ(canonical, MAP_FAILED, "unexported_short_mmap"))
+ goto out;
+ canonical[0] = 42;
+ ASSERT_EQ(canonical[0], 42, "unexported_short_access");
+ snprintf(pin, sizeof(pin), "/sys/fs/bpf/arena_unexported_%d", getpid());
+ if (!ASSERT_OK(bpf_obj_pin(map_fd, pin), "unexported_pin"))
+ goto out;
+ if (ASSERT_OK(stat(pin, &st), "unexported_stat"))
+ ASSERT_EQ(st.st_size, 0, "unexported_size");
+ fd = open(pin, O_RDONLY);
+ err = errno;
+ if (ASSERT_EQ(fd, -1, "unexported_open"))
+ ASSERT_EQ(err, EIO, "unexported_open_errno");
+ else
+ close(fd);
+ fd = bpf_obj_get(pin);
+ if (ASSERT_GE(fd, 0, "unexported_obj_get"))
+ close(fd);
+ ASSERT_OK(unlink(pin), "unexported_unlink");
+out:
+ if (canonical != MAP_FAILED)
+ munmap(canonical, ps);
+ close(map_fd);
+}
+
+static void test_export_requires_no_free(void)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_EXPORT);
+ int fd;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_export",
+ 0, 0, ARENA_PAGES, &opts);
+ if (fd == -EOPNOTSUPP) {
+ test__skip();
+ return;
+ }
+ ASSERT_EQ(fd, -EINVAL, "export_requires_no_free");
+ if (fd >= 0)
+ close(fd);
+}
+
+static void test_read_only(int flags)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *addr;
+ int fd = -1, err, ret;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ if (!ASSERT_OK(fchmod(file.file_fd, 0400), "chmod_read_only"))
+ goto out;
+ fd = open(file.pin, O_RDONLY);
+ if (!ASSERT_GE(fd, 0, "open_read_only"))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ, flags, fd, ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "read_only_mmap"))
+ goto out;
+ ASSERT_EQ(alias[0], 0, "read_only_fault");
+ ((char *)file.canonical)[ps] = 42;
+ ASSERT_EQ(alias[0], 42, "read_only_shared_update");
+
+ addr = mmap(NULL, ps, PROT_READ | PROT_WRITE, flags, fd, ps);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, "read_only_writable_mmap"))
+ ASSERT_EQ(err, EACCES, "read_only_writable_errno");
+ else
+ munmap(addr, ps);
+ ret = mprotect(alias, ps, PROT_READ | PROT_WRITE);
+ err = errno;
+ ASSERT_EQ(ret, -1, "read_only_mprotect");
+ ASSERT_EQ(err, EACCES, "read_only_mprotect_errno");
+ addr = mmap(NULL, ps, PROT_READ, MAP_PRIVATE, fd, ps);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, "read_only_private_mmap"))
+ ASSERT_EQ(err, EINVAL, "read_only_private_errno");
+ else
+ munmap(addr, ps);
+
+ close(fd);
+ fd = -1;
+ ((char *)file.canonical)[ps] = 73;
+ ASSERT_EQ(alias[0], 73, "read_only_after_close");
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ if (fd >= 0)
+ close(fd);
+ cleanup(&file);
+}
+
+static void test_fixed_size(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ struct stat st;
+ unsigned char resident;
+ char *alias = MAP_FAILED;
+ int err, fd, ret;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, file.file_fd, 0);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "alias_mmap"))
+ goto out;
+ alias[0] = 42;
+ ret = ftruncate(file.file_fd, 0);
+ err = errno;
+ ASSERT_EQ(ret, -1, "ftruncate_shrink");
+ ASSERT_EQ(err, EINVAL, "ftruncate_shrink_errno");
+ ret = ftruncate(file.file_fd, file.len + ps);
+ err = errno;
+ ASSERT_EQ(ret, -1, "ftruncate_grow");
+ ASSERT_EQ(err, EINVAL, "ftruncate_grow_errno");
+ ret = truncate(file.pin, 0);
+ err = errno;
+ ASSERT_EQ(ret, -1, "truncate_pin");
+ ASSERT_EQ(err, EINVAL, "truncate_pin_errno");
+ fd = open(file.pin, O_RDWR | O_TRUNC);
+ err = errno;
+ if (ASSERT_EQ(fd, -1, "open_trunc"))
+ ASSERT_EQ(err, EINVAL, "open_trunc_errno");
+ else
+ close(fd);
+ ASSERT_OK(ftruncate(file.file_fd, file.len), "ftruncate_same_size");
+ if (!ASSERT_OK(mincore(alias, ps, &resident), "mincore"))
+ goto out;
+ ASSERT_EQ(resident & 1, 1, "mapping_not_zapped");
+ ASSERT_EQ(alias[0], 42, "mapping_value");
+ ASSERT_OK(fchmod(file.file_fd, 0640), "chmod_pin");
+ if (ASSERT_OK(fstat(file.file_fd, &st), "stat_pin")) {
+ ASSERT_EQ(st.st_size, file.len, "fixed_size");
+ ASSERT_EQ(st.st_mode & 0777, 0640, "pin_mode");
+ }
+ ASSERT_EQ(lseek(file.file_fd, 0, SEEK_END), file.len, "seek_capacity");
+ ASSERT_EQ(lseek(file.file_fd, -(off_t)ps, SEEK_CUR), file.len - ps, "seek_back");
+ ASSERT_EQ(lseek(file.file_fd, ps, SEEK_SET), ps, "seek_offset");
+ errno = 0;
+ ASSERT_EQ(lseek(file.file_fd, file.len + 1, SEEK_SET), -1, "seek_past_end");
+ ASSERT_EQ(errno, EINVAL, "seek_past_end_errno");
+ errno = 0;
+ ASSERT_EQ(lseek(file.file_fd, -1, SEEK_SET), -1, "seek_before_start");
+ ASSERT_EQ(errno, EINVAL, "seek_before_start_errno");
+ fd = bpf_obj_get(file.pin);
+ if (ASSERT_GE(fd, 0, "obj_get_after_chmod"))
+ close(fd);
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_zero_canonical(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *addr;
+
+ if (!setup(&file, 0, true))
+ goto out;
+ if (!ASSERT_EQ((unsigned long)file.canonical, 0, "zero_canonical"))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, file.file_fd, ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "zero_canonical_export"))
+ goto out;
+ alias[0] = 42;
+ ASSERT_EQ(alias[0], 42, "zero_canonical_fault");
+ addr = mmap((void *)ps, file.len, PROT_READ | PROT_WRITE,
+ MAP_SHARED | MAP_FIXED_NOREPLACE, file.map_fd, 0);
+ if (ASSERT_EQ(addr, MAP_FAILED, "canonical_cannot_move"))
+ ASSERT_EQ(errno, EINVAL, "canonical_cannot_move_errno");
+ else
+ munmap(addr, file.len);
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_lifecycle(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *target = MAP_FAILED, *addr;
+ pid_t child;
+ int status, ret, err;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ alias = mmap(NULL, ps * 2, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, 0);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "lifecycle_mmap"))
+ goto out;
+ alias[0] = 42;
+ ret = munmap(alias, ps);
+ err = errno;
+ ASSERT_EQ(ret, -1, "partial_unmap");
+ ASSERT_EQ(err, EINVAL, "partial_unmap_errno");
+ ret = mprotect(alias, ps, PROT_READ);
+ err = errno;
+ ASSERT_EQ(ret, -1, "partial_mprotect");
+ ASSERT_EQ(err, EINVAL, "partial_mprotect_errno");
+ ret = madvise(alias, ps * 2, MADV_DOFORK);
+ err = errno;
+ ASSERT_EQ(ret, -1, "enable_inheritance");
+ ASSERT_EQ(err, EINVAL, "enable_inheritance_errno");
+ target = mmap(NULL, ps * 2, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+ if (!ASSERT_NEQ(target, MAP_FAILED, "reserve_remap_target"))
+ goto out;
+ addr = mremap(alias, ps * 2, ps * 2, MREMAP_MAYMOVE | MREMAP_FIXED, target);
+ err = errno;
+ if (!ASSERT_EQ(addr, MAP_FAILED, "relocate_mapping")) {
+ alias = addr;
+ target = MAP_FAILED;
+ }
+ ASSERT_EQ(err, EINVAL, "relocate_mapping_errno");
+ child = fork();
+ if (!ASSERT_GE(child, 0, "fork"))
+ goto out;
+ if (!child) {
+ unsigned char resident[2];
+
+ ret = mincore(alias, ps * 2, resident);
+ _exit(ret == -1 && errno == ENOMEM ? 0 : 1);
+ }
+ if (ASSERT_EQ(waitpid(child, &status, 0), child, "wait_child") &&
+ ASSERT_TRUE(WIFEXITED(status), "child_exited"))
+ ASSERT_EQ(WEXITSTATUS(status), 0, "mapping_not_inherited");
+ ASSERT_EQ(alias[0], 42, "parent_mapping_retained");
+out:
+ if (target != MAP_FAILED)
+ munmap(target, ps * 2);
+ if (alias != MAP_FAILED)
+ munmap(alias, ps * 2);
+ cleanup(&file);
+}
+
+void serial_test_arena_pinned(void)
+{
+ if (test__start_subtest("ordinary_pin"))
+ test_unexported_pin(0);
+ if (test__start_subtest("retained_pin"))
+ test_unexported_pin(BPF_F_ARENA_NO_FREE);
+ if (test__start_subtest("export_requires_no_free"))
+ test_export_requires_no_free();
+ if (test__start_subtest("mappings"))
+ test_mappings(0);
+ if (test__start_subtest("map_extra"))
+ test_mappings(1ULL << 32);
+ if (test__start_subtest("fixed_size"))
+ test_fixed_size();
+ if (test__start_subtest("read_only"))
+ test_read_only(MAP_SHARED);
+ if (test__start_subtest("read_only_validate"))
+ test_read_only(MAP_SHARED_VALIDATE);
+ if (test__start_subtest("zero_canonical"))
+ test_zero_canonical();
+ if (test__start_subtest("lifecycle"))
+ test_lifecycle();
+}
|
| 27/3 | 2026/10/10 23:16 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 29853 output: 34 thoughts: 1318 Thoughts: **KMSAN Fuzzing Session Justification: BPF Arena Flags Patch** Alright, let's break down this patch series and determine if it warrants a dedicated KMSAN fuzzing session. I need to be precise here; these resources are valuable. The guideline is clear: KMSAN is for catching *uninitialized memory* risks that are unique to KMSAN. So, let's analyze carefully. This patch introduces two new flags for BPF arenas: `BPF_F_ARENA_NO_FREE` and `BPF_F_ARENA_EXPORT`. The goal is to allow a pinned arena in bpffs to be opened and mmap'ed as a sized file, which could be useful for something like shared guest RAM in QEMU. Let's dig in. The changes are in documentation, the BPF header (`linux/bpf.h`), and the kernel code (`arena.c` and `inode.c`), along with updates to the tools and tests. Specifically, I see these changes: * **Header File:** The flags `BPF_F_ARENA_NO_FREE = (1U << 20)` and `BPF_F_ARENA_EXPORT = (1U << 21)` are added to the BPF definitions in `linux/bpf.h`. * **Arena Allocations and Checks:** In `arena.c`, the `arena_map_alloc()` function now checks for the new flags, specifically `BPF_F_ARENA_EXPORT` and `BPF_F_ARENA_NO_FREE`. It looks like the flags have a significant impact on how the arena is managed. * **Virtual Memory Fault Handling:** The `arena_vm_fault()` function is crucial here. It uses offset calculations to determine the kernel virtual address (`kaddr`) based on the arena's base address and the page offset (`vmf->pgoff`). This is where uninitialized memory *could* become a problem. I will have to investigate this further. * **Unmapped Area and mmap:** The `arena_get_unmapped_area()` and `arena_map_mmap()` functions have checks related to the new flags and file operations. They validate the memory region, ensuring that we don't end up with some kind of illegal mmap operations and also validate `user_vm_end` for memory bounds. * **Memory Management:** The patch includes the `zap_pages()` function, which now has new offset calculations related to the VMA. The functions `bpf_arena_free_pages` and `bpf_arena_free_pages_non_sleepable` are now NOOP if `BPF_F_ARENA_NO_FREE` is set. * **File System Integration:** The `inode.c` file is also updated, with modifications to the following functions: `bpf_map_setattr`, `bpffs_arena_open`, `bpffs_arena_mmap`, `bpffs_arena_release`, `bpffs_arena_get_unmapped_area`, `bpffs_arena_llseek`, `bpf_mkobj_ops`, and `bpf_mkmap`. These functions deal with the file system operations. * **Tools:** Lastly, tools and tests are updated to take advantage of new functionalities. Now, my primary concern is the `arena_vm_fault` function and other memory access pathways. Does it expose uninitialized memory? The critical question: Does it create a situation where a BPF program can read data that hasn't been properly initialized? That's what KMSAN is designed to catch. It's the offset calculations using `vmf->pgoff` that I need to focus on. To definitively decide, I'll need to scrutinize `arena_vm_fault` and all call sites in `arena.c` that involve memory reads or copies, paying attention to how the offset calculations are performed and whether the memory is zeroed or initialized beforehand. If there is a risk of uninitialized memory being exposed to user space through any of the new flags, then I'll set `NeedsKMSAN=true`. Otherwise, this should be fine without dedicated KMSAN coverage. |
| 28/3 | 2026/10/10 23:16 | tool | grepper |
0mArgs:{
"Expression": "arena_vm_fault",
"PathPrefixes": [
"kernel/bpf/arena.c"
]
}Results: {
"Output": "kernel/bpf/arena.c=475=static void arena_vm_close(struct vm_area_struct *vma)\n--\nkernel/bpf/arena.c-489-\nkernel/bpf/arena.c:490:static vm_fault_t arena_vm_fault(struct vm_fault *vmf)\nkernel/bpf/arena.c-491-{\n--\nkernel/bpf/arena.c=606=static const struct vm_operations_struct arena_vm_ops = {\n--\nkernel/bpf/arena.c-610-\t.close\t\t= arena_vm_close,\nkernel/bpf/arena.c:611:\t.fault = arena_vm_fault,\nkernel/bpf/arena.c-612-};\n"
} |
| 29/3 | 2026/10/10 23:16 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32326 output: 36 thoughts: 30 Thoughts: Okay, here's my interpretation of reading those lines, formatted as requested: **Deconstructing `arena.c`'s Allocation Logic** Alright, let's dive into this `kernel/bpf/arena.c` code. Specifically, I'm focusing on lines 490 through 550, which I've just pulled up using `read-file`. This block seems crucial for understanding the arena allocation strategy in BPF. My initial impression is that this section is probably handling the core logic for either allocating new arena blocks or managing the existing ones within the overall arena structure. Given the surrounding code, I'm expecting to find details on how the allocator determines *where* to place these blocks, the mechanisms for tracking free space, or maybe even some error handling related to fragmentation or resource exhaustion. I'm betting there's a loop or a recursive structure in play here, perhaps iterating through available regions within the arena to find a suitable fit for a requested allocation size. I'll need to look for any usage of `struct arena_block` or related data structures to confirm that. The presence of bitmask operations or other clever data packing is highly likely, given the performance-critical nature of the BPF subsystem. I'm also going to be keenly interested in any memory barriers or atomic operations, as we're likely dealing with concurrent access. This is going to be my key area of focus to identify any potential race conditions and the strategies implemented to avoid them. Ultimately, I'm looking to understand the algorithm's time complexity and any potential bottlenecks. I'll be keeping an eye out for any comments that provide insights into optimization choices or the reasoning behind specific design decisions. From a performance standpoint, I suspect this section is optimized for the common case, so I will be analyzing to discover the performance tradeoff that was made in this code to optimize it for its intended use case. |
| 30/3 | 2026/10/10 23:16 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/arena.c",
"FirstLine": 490,
"LineCount": 60
}Results: {
"Contents": " 490:\tstatic vm_fault_t arena_vm_fault(struct vm_fault *vmf)\n 491:\t{\n 492:\t\tstruct bpf_map *map = vmf-\u003evma-\u003evm_file-\u003eprivate_data;\n 493:\t\tstruct bpf_arena *arena = container_of(map, struct bpf_arena, map);\n 494:\t\tstruct mem_cgroup *new_memcg, *old_memcg;\n 495:\t\tLLIST_HEAD(pages);\n 496:\t\tstruct page *page, *new_page = NULL;\n 497:\t\tstruct apply_range_data data;\n 498:\t\tvm_fault_t fault_ret;\n 499:\t\tlong kbase, kaddr;\n 500:\t\tunsigned long flags;\n 501:\t\tint ret;\n 502:\t\n 503:\t\tkbase = bpf_arena_get_kern_vm_start(arena);\n 504:\t\t/* vmf-\u003epgoff includes the file offset of a bpffs-backed slice. */\n 505:\t\tkaddr = kbase + (u32)(arena-\u003euser_vm_start + ((u64)vmf-\u003epgoff \u003c\u003c PAGE_SHIFT));\n 506:\t\n 507:\t\tpage = vmalloc_to_page((void *)kaddr);\n 508:\t\tif (!page \u0026\u0026 !(arena-\u003emap.map_flags \u0026 BPF_F_SEGV_ON_FAULT)) {\n 509:\t\t\t/*\n 510:\t\t\t * Preallocate outside the lock so the allocation can reclaim;\n 511:\t\t\t * __GFP_RETRY_MAYFAIL keeps the OOM killer out of it.\n 512:\t\t\t */\n 513:\t\t\tbpf_map_memcg_enter(\u0026arena-\u003emap, \u0026old_memcg, \u0026new_memcg);\n 514:\t\t\tnew_page = alloc_pages_node(map-\u003enuma_node,\n 515:\t\t\t\t\t\t GFP_KERNEL | __GFP_ZERO |\n 516:\t\t\t\t\t\t __GFP_ACCOUNT | __GFP_NOWARN |\n 517:\t\t\t\t\t\t __GFP_RETRY_MAYFAIL, 0);\n 518:\t\t\tbpf_map_memcg_exit(old_memcg, new_memcg);\n 519:\t\t}\n 520:\t\n 521:\t\tif (raw_res_spin_lock_irqsave(\u0026arena-\u003espinlock, flags)) {\n 522:\t\t\t/*\n 523:\t\t\t * A failed lock means a possible deadlock was detected. Don't\n 524:\t\t\t * return VM_FAULT_RETRY: this handler never took mmap_lock, but\n 525:\t\t\t * the fault path would re-take it on retry and deadlock. Fail.\n 526:\t\t\t */\n 527:\t\t\tif (new_page)\n 528:\t\t\t\tfree_pages_nolock(new_page, 0);\n 529:\t\t\treturn VM_FAULT_SIGBUS;\n 530:\t\t}\n 531:\t\n 532:\t\tpage = vmalloc_to_page((void *)kaddr);\n 533:\t\tif (page) {\n 534:\t\t\tif (page == arena-\u003escratch_page) {\n 535:\t\t\t\t/* BPF triggered scratch here; don't lazy-alloc over it */\n 536:\t\t\t\tfault_ret = (arena-\u003emap.map_flags \u0026 BPF_F_SEGV_ON_FAULT)\n 537:\t\t\t\t\t ? VM_FAULT_SIGSEGV : VM_FAULT_SIGBUS;\n 538:\t\t\t\tgoto out_err_locked;\n 539:\t\t\t}\n 540:\t\t\t/* already have a page vmap-ed */\n 541:\t\t\tgoto out;\n 542:\t\t}\n 543:\t\n 544:\t\tif (arena-\u003emap.map_flags \u0026 BPF_F_SEGV_ON_FAULT) {\n 545:\t\t\t/*\n 546:\t\t\t * User space requested to segfault when page is not allocated\n 547:\t\t\t * by bpf prog\n 548:\t\t\t */\n 549:\t\t\tfault_ret = VM_FAULT_SIGSEGV;\n"
} |
| 31/3 | 2026/10/10 23:16 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 33863 output: 38 thoughts: 70 Thoughts: Here's my summary, as if I were thinking it: **Digging into Memory Allocation for BPF Arenas** Okay, I'm taking a look at this code and something jumps out at me on line 515. Specifically, I see `new_page = alloc_pages_node(map->numa_node, GFP_KERNEL | __GFP_ZERO | ...)`. The key here is that the allocated page is *zero-initialized* thanks to `__GFP_ZERO`. This is crucial for security and predictability, and it's a good practice to verify. So, my next logical step is to quickly cross-reference this with `bpf_arena_alloc_pages`. I need to confirm that this zero-initialization strategy is consistently applied throughout the memory management in this area. Making sure there are no leaks or unintentional exposures of data is paramount in BPF, and I want to verify that this best practice is being followed here as well. |
| 32/3 | 2026/10/10 23:16 | tool | grepper |
0mArgs:{
"Expression": "alloc_pages|kzalloc|vmalloc",
"PathPrefixes": [
"kernel/bpf/arena.c"
]
}Results: {
"Output": "kernel/bpf/arena.c-10-#include \u003clinux/btf_ids.h\u003e\nkernel/bpf/arena.c:11:#include \u003clinux/vmalloc.h\u003e\nkernel/bpf/arena.c-12-#include \u003clinux/pagemap.h\u003e\n--\nkernel/bpf/arena.c-39- * into kernel and user vma.\nkernel/bpf/arena.c:40: * bpf program can allocate a page via bpf_arena_alloc_pages() kfunc\nkernel/bpf/arena.c-41- * which will insert it into kernel vm_area.\n--\nkernel/bpf/arena.c=381=static void arena_map_free(struct bpf_map *map)\n--\nkernel/bpf/arena.c-399-\t * free_vm_area() calls remove_vm_area() that calls free_unmap_vmap_area().\nkernel/bpf/arena.c:400:\t * It unmaps everything from vmalloc area and clears pgtables.\nkernel/bpf/arena.c-401-\t * Call apply_to_existing_page_range() first to find populated ptes and\n--\nkernel/bpf/arena.c=490=static vm_fault_t arena_vm_fault(struct vm_fault *vmf)\n--\nkernel/bpf/arena.c-506-\nkernel/bpf/arena.c:507:\tpage = vmalloc_to_page((void *)kaddr);\nkernel/bpf/arena.c-508-\tif (!page \u0026\u0026 !(arena-\u003emap.map_flags \u0026 BPF_F_SEGV_ON_FAULT)) {\n--\nkernel/bpf/arena.c-513-\t\tbpf_map_memcg_enter(\u0026arena-\u003emap, \u0026old_memcg, \u0026new_memcg);\nkernel/bpf/arena.c:514:\t\tnew_page = alloc_pages_node(map-\u003enuma_node,\nkernel/bpf/arena.c-515-\t\t\t\t\t GFP_KERNEL | __GFP_ZERO |\n--\nkernel/bpf/arena.c-531-\nkernel/bpf/arena.c:532:\tpage = vmalloc_to_page((void *)kaddr);\nkernel/bpf/arena.c-533-\tif (page) {\n--\nkernel/bpf/arena.c=749=static int arena_adjust_tree(struct bpf_arena *arena, long uaddr, long page_cnt, long *pgoff)\n--\nkernel/bpf/arena.c-770-/*\nkernel/bpf/arena.c:771: * Allocate pages and vmap them into kernel vmalloc area.\nkernel/bpf/arena.c-772- * Later the pages will be mmaped into user space vma.\nkernel/bpf/arena.c-773- */\nkernel/bpf/arena.c:774:static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt, int node_id,\nkernel/bpf/arena.c-775-\t\t\t bool sleepable)\n--\nkernel/bpf/arena.c-806-\nkernel/bpf/arena.c:807:\tret = bpf_alloc_pages(node_id, page_cnt, \u0026pages, sleepable);\nkernel/bpf/arena.c-808-\tif (ret)\n--\nkernel/bpf/arena.c-864-/*\nkernel/bpf/arena.c:865: * If page is present in vmalloc area, unmap it from vmalloc area,\nkernel/bpf/arena.c-866- * unmap it from all user space vma-s,\n--\nkernel/bpf/arena.c=937=static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt, bool sleepable)\n--\nkernel/bpf/arena.c-1026- * Reserve an arena virtual address range without populating it. This call stops\nkernel/bpf/arena.c:1027: * bpf_arena_alloc_pages from adding pages to this range.\nkernel/bpf/arena.c-1028- */\n--\nkernel/bpf/arena.c=1137=__bpf_kfunc_start_defs();\nkernel/bpf/arena.c-1138-\nkernel/bpf/arena.c:1139:__bpf_kfunc void *bpf_arena_alloc_pages(void *p__map, void *addr__ign, u32 page_cnt,\nkernel/bpf/arena.c-1140-\t\t\t\t\tint node_id, u64 flags)\n--\nkernel/bpf/arena.c-1147-\nkernel/bpf/arena.c:1148:\treturn (void *)arena_alloc_pages(arena, (long)addr__ign, page_cnt, node_id, true);\nkernel/bpf/arena.c-1149-}\nkernel/bpf/arena.c-1150-\nkernel/bpf/arena.c:1151:void *bpf_arena_alloc_pages_non_sleepable(void *p__map, void *addr__ign, u32 page_cnt,\nkernel/bpf/arena.c-1152-\t\t\t\t\t int node_id, u64 flags)\n--\nkernel/bpf/arena.c-1159-\nkernel/bpf/arena.c:1160:\treturn (void *)arena_alloc_pages(arena, (long)addr__ign, page_cnt, node_id, false);\nkernel/bpf/arena.c-1161-}\nkernel/bpf/arena.c-1162-\nkernel/bpf/arena.c:1163:void *bpf_arena_alloc_pages_sleepable(void *p__map, void *addr__ign, u32 page_cnt,\nkernel/bpf/arena.c-1164-\t\t\t\t int node_id, u64 flags)\n--\nkernel/bpf/arena.c-1171-\nkernel/bpf/arena.c:1172:\treturn (void *)arena_alloc_pages(arena, (long)addr__ign, page_cnt, node_id, true);\nkernel/bpf/arena.c-1173-}\n--\nkernel/bpf/arena.c=1212=BTF_KFUNCS_START(arena_kfuncs)\nkernel/bpf/arena.c:1213:BTF_ID_FLAGS(func, bpf_arena_alloc_pages, KF_ARENA_RET | KF_ARENA_ARG2 | KF_SPINLOCK_SAFE)\nkernel/bpf/arena.c-1214-BTF_ID_FLAGS(func, bpf_arena_free_pages, KF_ARENA_ARG2 | KF_SPINLOCK_SAFE)\n"
} |
| 33/3 | 2026/10/10 23:16 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 35903 output: 32 thoughts: 20 Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience: **Investigating `bpf_alloc_pages` Zeroing Behavior** Okay, so I need to quickly determine whether `bpf_alloc_pages` guarantees zeroed pages upon allocation. I'm assuming that proper initialization is critical in this context to prevent information leakage or unexpected behaviors within the BPF context. The potential implications of uninitialized memory in a BPF program are significant, especially concerning security and determinism. Therefore, before I proceed with any implementation leveraging this function, a deep dive into its implementation details to confirm the zeroing behavior is paramount. This will influence how I structure subsequent data handling within the BPF program, and whether I need to add any manual zeroing operations. This is important to verify before I go any further. |
| 34/3 | 2026/10/10 23:16 | tool | grepper |
0mArgs:{
"Expression": "bpf_alloc_pages",
"PathPrefixes": [
"kernel/bpf/"
]
}Results: {
"Output": "kernel/bpf/arena.c=774=static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt, int node_id,\n--\nkernel/bpf/arena.c-806-\nkernel/bpf/arena.c:807:\tret = bpf_alloc_pages(node_id, page_cnt, \u0026pages, sleepable);\nkernel/bpf/arena.c-808-\tif (ret)\n--\nkernel/bpf/syscall.c=616=void bpf_free_pages(struct llist_head *pages)\n--\nkernel/bpf/syscall.c-625-\nkernel/bpf/syscall.c:626:int bpf_alloc_pages(int nid, unsigned long nr_pages,\nkernel/bpf/syscall.c-627-\t\t struct llist_head *pages, bool sleepable)\n"
} |
| 35/3 | 2026/10/10 23:16 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 36363 output: 36 thoughts: 23 Thoughts: Okay, let's dive into this. Here's my thought process on the matter: **Analyzing `bpf_alloc_pages` in `kernel/bpf/syscall.c`** Alright, time to crack open `kernel/bpf/syscall.c` and zero in on `bpf_alloc_pages`. I need to understand exactly how the BPF subsystem allocates physical pages for various purposes. Knowing this is critical, as it likely forms the foundation for BPF program memory management. Specifically, I need to understand: Are they using standard kernel allocation functions, and if so, which ones? Are there any BPF-specific memory allocators in play? What considerations are made regarding NUMA nodes, page size, alignment, and security? I should pay particular attention to potential vulnerabilities that might stem from memory allocation practices, like double frees or off-by-one errors. It is also important to consider potential race conditions. I should be looking for any hints of resource limits or checks for allocation failures, and how the code handles those situations. Finally, I'll be particularly interested in comparing and contrasting the behavior of `bpf_alloc_pages` with other memory allocation schemes within the kernel. It's likely that it will be using some sort of shared memory arrangement with userspace, or be doing something exotic to allow BPF programs to manage memory. |
| 36/3 | 2026/10/10 23:16 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/syscall.c",
"FirstLine": 626,
"LineCount": 25
}Results: {
"Contents": " 626:\tint bpf_alloc_pages(int nid, unsigned long nr_pages,\n 627:\t\t\t struct llist_head *pages, bool sleepable)\n 628:\t{\n 629:\t\tunsigned long i;\n 630:\t\tstruct page *pg;\n 631:\t\n 632:\t\tfor (i = 0; i \u003c nr_pages; i++) {\n 633:\t\t\tpg = bpf_alloc_page(nid, sleepable);\n 634:\t\t\tif (!pg)\n 635:\t\t\t\tgoto free_pages;\n 636:\t\t\tllist_add(\u0026pg-\u003epcp_llist, pages);\n 637:\t\t}\n 638:\t\n 639:\t\treturn 0;\n 640:\t\n 641:\tfree_pages:\n 642:\t\tbpf_free_pages(pages);\n 643:\t\treturn -ENOMEM;\n 644:\t}\n 645:\t\n 646:\tstatic int btf_field_cmp(const void *a, const void *b)\n 647:\t{\n 648:\t\tconst struct btf_field *f1 = a, *f2 = b;\n 649:\t\n 650:\t\tif (f1-\u003eoffset \u003c f2-\u003eoffset)\n"
} |
| 37/3 | 2026/10/10 23:16 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37064 output: 33 thoughts: 12 Thoughts: Okay, let's dive into this `bpf_alloc_page` function. **Analyzing `bpf_alloc_page`** Alright, so I'm thinking about `bpf_alloc_page`. My initial instinct is to consider its purpose within the broader BPF framework. It's clearly involved in memory allocation, specifically pages. That implies a need for larger, contiguous blocks of memory, which is standard for BPF programs and the data they operate on. I'm immediately picturing the kernel's page allocator at play here – that's the obvious mechanism. I need to understand how this function interfaces with that. Is it a direct wrapper, or does it add some BPF-specific layer on top? Are there any BPF-specific memory management considerations that are being handled here, like potential alignment requirements or page-related security features. I'm also curious about its implications for resource limits – is there any inherent limit on how many pages a BPF program can request? Are there any specific error handling strategies or checks against potential denial-of-service? Time to dig deeper and see how this fits into the overall BPF memory landscape. |
| 38/3 | 2026/10/10 23:16 | tool | grepper |
0mArgs:{
"Expression": "bpf_alloc_page\\(",
"PathPrefixes": [
"kernel/bpf/"
]
}Results: {
"Output": "kernel/bpf/arena.c=271=static struct bpf_map *arena_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/arena.c-325-\nkernel/bpf/arena.c:326:\tarena-\u003escratch_page = bpf_alloc_page(NUMA_NO_NODE, true);\nkernel/bpf/arena.c-327-\tif (!arena-\u003escratch_page)\n--\nkernel/bpf/arena.c=490=static vm_fault_t arena_vm_fault(struct vm_fault *vmf)\n--\nkernel/bpf/arena.c-559-\t\t */\nkernel/bpf/arena.c:560:\t\tnew_page = bpf_alloc_page(map-\u003enuma_node, false);\nkernel/bpf/arena.c-561-\t\tif (!new_page) {\n--\nkernel/bpf/syscall.c=599=static bool can_alloc_pages(void)\n--\nkernel/bpf/syscall.c-604-\nkernel/bpf/syscall.c:605:struct page *bpf_alloc_page(int nid, bool sleepable)\nkernel/bpf/syscall.c-606-{\n--\nkernel/bpf/syscall.c=626=int bpf_alloc_pages(int nid, unsigned long nr_pages,\n--\nkernel/bpf/syscall.c-632-\tfor (i = 0; i \u003c nr_pages; i++) {\nkernel/bpf/syscall.c:633:\t\tpg = bpf_alloc_page(nid, sleepable);\nkernel/bpf/syscall.c-634-\t\tif (!pg)\n"
} |
| 39/3 | 2026/10/10 23:16 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37787 output: 36 thoughts: 24 Thoughts: Here's a breakdown of my thought process, focusing on `bpf_alloc_page`: **Diving into `bpf_alloc_page`** Okay, so I'm looking at the `bpf_alloc_page` implementation, and naturally, my first thought is, "Where's the actual memory allocation happening?" It's in `kernel/bpf/syscall.c`, as indicated. Knowing that, I'm expecting to find the low-level calls to `kmalloc` or `__get_free_pages`, or something similar, likely wrapped in some BPF-specific safety checks or context management. I'm keen to see if this utilizes any specific BPF-related memory pools, potentially for performance or security reasons. I'm also curious about the potential integration with page tables if applicable, and if they're relevant for any considerations about memory access and protection within the BPF context. I'm assuming there's some sort of accounting involved, to avoid memory leaks or excessive allocation in BPF programs. I'm also prepared for a layer of abstraction or indirection, given this is the BPF subsystem, which tends to encapsulate a lot of low-level complexity. I'm interested in how it's integrated with the larger BPF memory management system, and how the kernel handles page faults or memory exhaustion conditions, considering this is likely used for dynamically allocating memory used by the BPF program. |
| 40/3 | 2026/10/10 23:16 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/syscall.c",
"FirstLine": 604,
"LineCount": 20
}Results: {
"Contents": " 604:\t\n 605:\tstruct page *bpf_alloc_page(int nid, bool sleepable)\n 606:\t{\n 607:\t\tif (!sleepable || !can_alloc_pages())\n 608:\t\t\treturn alloc_pages_nolock(__GFP_ACCOUNT, nid, 0);\n 609:\t\n 610:\t\treturn alloc_pages_node(nid,\n 611:\t\t\t\t\tGFP_KERNEL | __GFP_ZERO | __GFP_ACCOUNT\n 612:\t\t\t\t\t| __GFP_NOWARN | __GFP_RETRY_MAYFAIL,\n 613:\t\t\t\t\t0);\n 614:\t}\n 615:\t\n 616:\tvoid bpf_free_pages(struct llist_head *pages)\n 617:\t{\n 618:\t\tstruct llist_node *node;\n 619:\t\tstruct page *page, *tmp;\n 620:\t\n 621:\t\tnode = llist_del_all(pages);\n 622:\t\tllist_for_each_entry_safe(page, tmp, node, pcp_llist)\n 623:\t\t\tfree_pages_nolock(page, 0);\n"
} |
| 41/3 | 2026/10/10 23:16 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 38469 output: 445 thoughts: 4193 Thoughts: **KMSAN Assessment: No Need for KMSAN in This Patch** Okay, let's break down this patch and determine if it necessitates KMSAN. My initial thought process is to meticulously examine each modification and assess potential risks. The primary focus is on identifying any instances where uninitialized memory could be exposed, especially to userspace. First, I notice the changes involve several functions within the BPF arena implementation, including `arena_map_alloc`, `arena_vm_fault`, `arena_get_unmapped_area`, `arena_map_mmap`, `zap_pages`, and `bpf_arena_free_pages`. Also, functions in `kernel/bpf/inode.c`, such as `bpf_map_setattr`, `bpffs_arena_open`, `bpffs_arena_mmap`, and others. These functions deal with memory mapping, fault handling, and arena management, especially in the context of file-backed BPF arenas. The core change appears to enable userspace mmap of these arenas at non-canonical addresses (exported slices), along with associated file operations. My primary concern is the potential for information leaks or the unintended exposure of uninitialized memory. I immediately check for any data transfers to userspace using functions like `copy_to_user`, `put_user`, or via netlink/ioctl. I also look for conditional logic that might depend on uninitialized variables. Reviewing the code, I don't see any direct copying of kernel structures to user space. All new variables within the added functions seem to be properly initialized before use. The arena memory itself is shared memory, backed by physical pages. The critical interaction point with userspace occurs during `arena_vm_fault` when page faults happen on mapped arena files. If the page is not in the arena it's allocated with `__GFP_ZERO` which is critical for safety. The patch *doesn't* change how the pages are allocated or initialized; it's just adding the file-backed arena operations. The `alloc_pages_nolock` function is untouched, so it still works like it always has. My concern about uninitialized memory in allocation is already dealt with in the existing code. The patch's changes focus on how memory is *accessed* (address calculations in `arena_vm_fault`, `zap_pages`, and how to deal with pinned, exported arenas). My analysis indicates the primary risks are potential address calculation errors, VMA accounting issues, or locking bugs. These are all risks of out-of-bounds access, use-after-free, reference count problems, and locking deadlocks. Crucially, these issues would be caught by KASAN, LOCKDEP, or standard bug detectors (e.g. refcount tracking, DEBUG_VM). The patch *does not* introduce any new or modified code that directly exposes uninitialized kernel memory to user space. It primarily modifies address calculations, file operations, and VMA accounting. There are file operations and file size management logic added, but these don't lead to unsafe `read()` or `write()` syscall code paths. There are no changes, additions, or code removals of kernel buffers or arena memory that might cause an information leak. User space access is strictly mediated by `mmap()` (page faults) and BPF instructions, with the kernel only using the `vmalloc` mapping. The file size handling is a file *attribute* and there are no file access operations implemented! Therefore, based on my analysis of the changes and the KMSAN guidelines, this patch does *not* introduce a risk that is *uniquely* detectable by KMSAN. The risks are primarily address calculation errors, VMA accounting/UAF, and locking, which are all well-covered by KASAN, standard debuggers, and LOCKDEP. There's no risk related to uninitialized memory. Therefore, I'm confident in setting `NeedsKMSAN=false` for this patch. I should also set `KMSANReasoning` with the above analysis. |
| 42/3 | 2026/10/10 23:17 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch series introduces BPF_F_ARENA_NO_FREE and BPF_F_ARENA_EXPORT flags for BPF arenas, allowing pinned arenas in bpffs to be opened and mapped as sized files with exported slices.\n\n1. Uninitialized Memory Analysis (KMSAN):\n- No kernel stack or heap structures are copied to userspace (no copy_to_user, put_user, netlink attributes, ioctl out-arguments, etc.).\n- File operations on the exported bpffs inode (bpffs_arena_fops) do not implement .read or .write handlers; memory access occurs solely via memory-mapping (mmap/fault).\n- Backing pages allocated during fault handling (arena_vm_fault) continue to use GFP_KERNEL | __GFP_ZERO.\n- All new stack variables in modified functions (e.g. start, end, vma_start, vma_end in zap_pages, and err in bpffs_arena_open) are explicitly initialized prior to evaluation.\n- No uninitialized memory or info-leak risk is introduced or exposed.\n\n2. Applicable Sanitizers (KASAN, LOCKDEP, DEBUG_VM):\n- Potential risks in this change involve offset and address calculations (e.g. kaddr calculation in arena_vm_fault and slice offsets in zap_pages), which could risk out-of-bounds accesses in the vmalloc area (caught by KASAN).\n- Reference counting and lifecycle management of the map across open/mmap/release/unlink (bpf_map_inc_with_uref, UAF or leak) are covered by KASAN and standard refcount tracking.\n- Locking and concurrency interactions (arena-\u003elock, arena-\u003espinlock) are monitored by LOCKDEP.\n\nTherefore, standard KASAN and debugging builds fully cover the risk profile of these changes, and a dedicated KMSAN fuzzing session is not warranted (NeedsKMSAN=false).",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch series introduces BPF_F_ARENA_NO_FREE and BPF_F_ARENA_EXPORT flags for BPF arenas, allowing pinned arenas in bpffs to be opened and mapped as sized files with exported slices.\n\n1. Uninitialized Memory Analysis (KMSAN):\n- No kernel stack or heap structures are copied to userspace (no copy_to_user, put_user, netlink attributes, ioctl out-arguments, etc.).\n- File operations on the exported bpffs inode (bpffs_arena_fops) do not implement .read or .write handlers; memory access occurs solely via memory-mapping (mmap/fault).\n- Backing pages allocated during fault handling (arena_vm_fault) continue to use GFP_KERNEL | __GFP_ZERO.\n- All new stack variables in modified functions (e.g. start, end, vma_start, vma_end in zap_pages, and err in bpffs_arena_open) are explicitly initialized prior to evaluation.\n- No uninitialized memory or info-leak risk is introduced or exposed.\n\n2. Applicable Sanitizers (KASAN, LOCKDEP, DEBUG_VM):\n- Potential risks in this change involve offset and address calculations (e.g. kaddr calculation in arena_vm_fault and slice offsets in zap_pages), which could risk out-of-bounds accesses in the vmalloc area (caught by KASAN).\n- Reference counting and lifecycle management of the map across open/mmap/release/unlink (bpf_map_inc_with_uref, UAF or leak) are covered by KASAN and standard refcount tracking.\n- Locking and concurrency interactions (arena-\u003elock, arena-\u003espinlock) are monitored by LOCKDEP.\n\nTherefore, standard KASAN and debugging builds fully cover the risk profile of these changes, and a dedicated KMSAN fuzzing session is not warranted (NeedsKMSAN=false).",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|