| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/09/24 01:45 | flow | patch-triage |
2mArgs:null Results: {
"EnableConfigs": [
"BPF_SYSCALL",
"BPF_JIT"
],
"FocusSymbols": [
"check_func_callx",
"resolve_func_ptrs",
"check_func_ptr_read",
"bpf_map_claim",
"bpf_map_range_func_ptrs",
"bpf_adjust_func_ptrs",
"bpf_adjust_func_ptrs_after_remove",
"emit_indirect_call",
"bpf_jit_supports_callx",
"jit_subprogs"
],
"KMSANReasoning": "The patch series implements support for indirect calls (`callx` instruction: `BPF_JMP | BPF_CALL | BPF_X`) in the BPF verifier, x86/arm64 JIT compilers, and user-space libbpf tooling.\n\nAnalysis of memory initialization and safety:\n1. Kernel structures and memory allocations:\n - `struct bpf_verifier_env` fields (`func_ptrs`, `func_ptr_cnt`, `has_callx`, `callx_edges`) are zero-initialized via `kzalloc`.\n - `callx_edges` is allocated using `kvcalloc()`, ensuring all bits are zeroed.\n - `func_ptrs` elements are allocated via `kvrealloc()` and each populated element up to `func_ptr_cnt` is initialized with a compound literal (`struct bpf_func_ptr`), zeroing any omitted fields (such as `.used`).\n - `bpf_subprog_call_depth_info` array (`dinfo`) is allocated using `kvzalloc_objs()`.\n - The new field `user` in `struct bpf_map` is initialized to 0 during map allocation via `__bpf_map_area_alloc()` (`__GFP_ZERO`).\n2. No kernel-to-user info-leak of uninitialized memory:\n - No new `copy_to_user`, `put_user`, socket option, netlink, or ioctl copy-out paths are introduced or modified.\n3. No conditional logic or branching depends on uninitialized state.\n4. The risks associated with this patch primarily involve verifier correctness, call graph cycle detection, and JIT assembly generation, all of which fall under functional verification, UBSAN, and KASAN (e.g., out-of-bounds or use-after-free).\n\nSince there are no uninitialized memory usage risks or uninitialized memory info-leaks, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false,
"Reasoning": "The patch series implements support for BPF indirect calls ('callx' / BPF_CALL | BPF_X) of static subprograms and function pointer resolution in frozen read-only maps. This introduces substantial modifications across the BPF verifier (CFG traversal, state pruning/liveness, stack depth and recursion checks, func pointer bounds checks) and the x86 BPF JIT compiler (indirect call emission and retpoline handling), along with several new runtime defensive assertions (verifier_bug_if). These changes modify reachable core subsystem code easily exercised via the bpf() syscall.",
"WorthFuzzing": true
} |
| 1/1 | 2026/09/24 01:45 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit 47a35dcdf09a938b11851b9703bb51549ad4f665\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Thu Sep 24 01:45:15 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/Documentation/bpf/clang-notes.rst b/Documentation/bpf/clang-notes.rst\nindex 2c872a1ee08e9..3ccc7b09d19e6 100644\n--- a/Documentation/bpf/clang-notes.rst\n+++ b/Documentation/bpf/clang-notes.rst\n@@ -23,8 +23,11 @@ For CPU versions prior to 3, Clang v7.0 and later can enable ``BPF_ALU`` support\n Jump instructions\n =================\n \n-If ``-O0`` is used, Clang will generate the ``BPF_CALL | BPF_X | BPF_JMP`` (0x8d)\n-instruction, which is not supported by the Linux kernel verifier.\n+Clang generates the ``BPF_CALL | BPF_X | BPF_JMP`` (0x8d) instruction for calls\n+through a function pointer. The Linux kernel verifier accepts it only when it\n+can prove that the register holds the address of a static BPF function, see\n+Documentation/bpf/linux-notes.rst. If ``-O0`` is used, Clang will generate this\n+instruction for helper calls as well, which is not supported.\n \n Atomic operations\n =================\ndiff --git a/Documentation/bpf/linux-notes.rst b/Documentation/bpf/linux-notes.rst\nindex 00d2693de025a..6c036b54a29f2 100644\n--- a/Documentation/bpf/linux-notes.rst\n+++ b/Documentation/bpf/linux-notes.rst\n@@ -15,10 +15,33 @@ Byte swap instructions\n Jump instructions\n =================\n \n-``BPF_CALL | BPF_X | BPF_JMP`` (0x8d), where the helper function\n-integer would be read from a specified register, is not currently supported\n-by the verifier. Any programs with this instruction will fail to load\n-until such support is added.\n+``BPF_CALL | BPF_X | BPF_JMP`` (0x8d), ``callx dst``, performs an indirect\n+call of a BPF function whose address is held in the ``dst`` register. The\n+``src``, ``offset`` and ``imm`` fields are reserved and must be zero.\n+\n+The address of a BPF function gets into a register in one of two ways:\n+\n+* it is loaded by a 64-bit immediate instruction with ``src`` =\n+ ``BPF_PSEUDO_FUNC``;\n+* it is read, with a 64-bit load, from a frozen read-only array map, that no\n+ other program uses, that holds its read-only data: tables of functions,\n+ structures of operations, vtables, where pointers to functions may be mixed\n+ with other data. In the map a pointer to a function is the offset in bytes\n+ of its first instruction in the program, and that is how the verifier\n+ recognizes it. It is replaced with the address of the function when\n+ the program is loaded. The program reads it from there, which requires\n+ ``CAP_PERFMON``.\n+\n+In both cases only static functions can be referenced. Therefore all functions\n+that can be called indirectly are known to the verifier before it starts to\n+analyze the program, and ``callx`` is verified as a direct call of every\n+function that ``dst`` may point to at that instruction. The same rules apply:\n+the calls can not be recursive, and the depth of the call chain and its\n+combined stack size are limited.\n+\n+Calling helper or kernel functions through a register, indirect calls of global\n+functions, and tail calls in functions that are called via ``callx`` are not\n+supported. ``callx`` requires the BPF JIT.\n \n Maps\n ====\ndiff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c\nindex 6c04fee468766..9544b2f483e59 100644\n--- a/arch/arm64/net/bpf_jit_comp.c\n+++ b/arch/arm64/net/bpf_jit_comp.c\n@@ -1789,6 +1789,17 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn\n \t\t\temit(A64_MOV(1, r0, A64_R(0)), ctx);\n \t\tbreak;\n \t}\n+\t/* indirect call of a bpf subprog, dst holds its address */\n+\tcase BPF_JMP | BPF_CALL | BPF_X:\n+\t\t/*\n+\t\t * It's the same as a direct call of a subprog that is out of\n+\t\t * range of BL: the subprog starts with BTI JC, the arguments\n+\t\t * are in place, and the registers that hold the tail call\n+\t\t * counter and the private stack are callee saved.\n+\t\t */\n+\t\temit(A64_BLR(dst), ctx);\n+\t\temit(A64_MOV(1, bpf2a64[BPF_REG_0], A64_R(0)), ctx);\n+\t\tbreak;\n \t/* tail call */\n \tcase BPF_JMP | BPF_TAIL_CALL:\n \t\tif (emit_bpf_tail_call(ctx))\n@@ -2485,6 +2496,11 @@ bool bpf_jit_supports_subprog_tailcalls(void)\n \treturn true;\n }\n \n+bool bpf_jit_supports_callx(void)\n+{\n+\treturn true;\n+}\n+\n static void invoke_bpf_prog(struct jit_ctx *ctx, struct bpf_tramp_node *node,\n \t\t\t int bargs_off, int retval_off, int run_ctx_off,\n \t\t\t bool save_ret)\ndiff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c\nindex d4a980140b48d..9fbef7504e51a 100644\n--- a/arch/x86/net/bpf_jit_comp.c\n+++ b/arch/x86/net/bpf_jit_comp.c\n@@ -749,6 +749,46 @@ static void emit_indirect_jump(u8 **pprog, int bpf_reg, u8 *ip)\n \t*pprog = prog;\n }\n \n+static void __emit_indirect_call(u8 **pprog, int reg, bool ereg)\n+{\n+\tu8 *prog = *pprog;\n+\n+\tif (ereg)\n+\t\tEMIT1(0x41);\n+\n+\tEMIT2(0xFF, 0xD0 + reg);\n+\n+\t*pprog = prog;\n+}\n+\n+/* call *bpf_reg */\n+static int emit_indirect_call(u8 **pprog, int bpf_reg, u8 *ip)\n+{\n+\tu8 *prog = *pprog;\n+\tint reg = reg2hex[bpf_reg];\n+\tbool ereg = is_ereg(bpf_reg);\n+\tint err = 0;\n+\n+\tif (cpu_feature_enabled(X86_FEATURE_INDIRECT_THUNK_ITS)) {\n+\t\tOPTIMIZER_HIDE_VAR(reg);\n+\t\terr = emit_call(\u0026prog, its_static_thunk(reg + 8*ereg), ip);\n+\t} else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE_LFENCE)) {\n+\t\tEMIT_LFENCE();\n+\t\t__emit_indirect_call(\u0026prog, reg, ereg);\n+\t} else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE)) {\n+\t\tOPTIMIZER_HIDE_VAR(reg);\n+\t\tif (cpu_feature_enabled(X86_FEATURE_CALL_DEPTH))\n+\t\t\terr = emit_call(\u0026prog, \u0026__x86_indirect_call_thunk_array[reg + 8*ereg], ip);\n+\t\telse\n+\t\t\terr = emit_call(\u0026prog, \u0026__x86_indirect_thunk_array[reg + 8*ereg], ip);\n+\t} else {\n+\t\t__emit_indirect_call(\u0026prog, reg, ereg);\n+\t}\n+\n+\t*pprog = prog;\n+\treturn err;\n+}\n+\n static void emit_return(u8 **pprog, u8 *ip)\n {\n \tu8 *prog = *pprog;\n@@ -2941,6 +2981,24 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *\n \t\t\tbreak;\n \t\t}\n \n+\t\t\t/* callx: call of a bpf subprog whose address is in dst_reg */\n+\t\tcase BPF_JMP | BPF_CALL | BPF_X:\n+\t\t\t/*\n+\t\t\t * The verifier makes sure that callees of callx are\n+\t\t\t * not tail call reachable, hence unlike a direct call\n+\t\t\t * of a subprog there is no need to pass\n+\t\t\t * tail_call_cnt_ptr in rax.\n+\t\t\t */\n+\t\t\tif (priv_frame_ptr) {\n+\t\t\t\tpush_r9(\u0026prog);\n+\t\t\t\tip += 2;\n+\t\t\t}\n+\t\t\tif (emit_indirect_call(\u0026prog, insn-\u003edst_reg, ip))\n+\t\t\t\treturn -EINVAL;\n+\t\t\tif (priv_frame_ptr)\n+\t\t\t\tpop_r9(\u0026prog);\n+\t\t\tbreak;\n+\n \t\tcase BPF_JMP | BPF_TAIL_CALL:\n \t\t\tif (imm32)\n \t\t\t\temit_bpf_tail_call_direct(bpf_prog,\n@@ -4414,6 +4472,16 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr\n \treturn prog;\n }\n \n+bool bpf_jit_supports_callx(void)\n+{\n+\t/*\n+\t * FineIBT poisons ENDBR at the entry of a JITed function and expects\n+\t * indirect callers to go through the CFI preamble instead.\n+\t * callx doesn't do that yet.\n+\t */\n+\treturn cfi_mode != CFI_FINEIBT;\n+}\n+\n bool bpf_jit_supports_kfunc_call(void)\n {\n \treturn true;\ndiff --git a/include/linux/bpf.h b/include/linux/bpf.h\nindex fd22db8bc6c50..7747c5fc290fc 100644\n--- a/include/linux/bpf.h\n+++ b/include/linux/bpf.h\n@@ -342,8 +342,18 @@ struct bpf_map {\n \ts64 __percpu *elem_count;\n \tu64 cookie; /* write-once */\n \tchar *excl_prog_sha;\n+\t/*\n+\t * Which programs use the map, see bpf_map_claim(): 0 - none so far,\n+\t * aux of the program - only that one, the same with BPF_MAP_USER_PATCHED\n+\t * set - only that one and it stored the addresses of its functions into\n+\t * the map, BPF_MAP_USER_MANY - more than one.\n+\t */\n+\tunsigned long user;\n };\n \n+#define BPF_MAP_USER_MANY\t1UL\n+#define BPF_MAP_USER_PATCHED\t1UL\n+\n static inline const char *btf_field_type_name(enum btf_field_type type)\n {\n \tswitch (type) {\ndiff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h\nindex 92f528c456052..38a4ba50669a2 100644\n--- a/include/linux/bpf_verifier.h\n+++ b/include/linux/bpf_verifier.h\n@@ -786,6 +786,22 @@ int bpf_log_attr_finalize(struct bpf_log_attr *attr, struct bpf_verifier_log *lo\n \n #define BPF_MAX_SUBPROGS 256\n \n+/*\n+ * A pointer to a static subprog in the value of a frozen read-only array map:\n+ * a 64-bit value that is the offset in bytes of the first instruction of\n+ * the subprog in the program.\n+ */\n+struct bpf_func_ptr {\n+\tstruct bpf_map *map;\n+\tu32 map_off;\t\t/* offset of the pointer in the value of the map */\n+\tu32 orig_off;\t\t/* what the map has: the first instruction of the subprog */\n+\tu32 xlated_off;\t\t/* the same after instructions were patched and removed */\n+\tbool used;\t\t/* the program reads the pointer */\n+};\n+\n+/* the subprog that a bpf_func_ptr pointed to was removed as dead code */\n+#define BPF_FUNC_PTR_DELETED ((u32)-1)\n+\n struct bpf_subprog_arg_info {\n \tenum bpf_arg_type arg_type;\n \tunion {\n@@ -965,6 +981,20 @@ struct bpf_verifier_env {\n \tstruct bpf_subprog_info subprog_info[BPF_MAX_SUBPROGS + 2]; /* max + 2 for the fake and exception subprogs */\n \t/* subprog indices sorted in topological order: leaves first, callers last */\n \tint subprog_topo_order[BPF_MAX_SUBPROGS + 2];\n+\t/*\n+\t * Pointers to static subprogs found in frozen read-only maps of the\n+\t * program, see resolve_func_ptrs(). Sorted by map and map_off.\n+\t */\n+\tstruct bpf_func_ptr *func_ptrs;\n+\tu32 func_ptr_cnt;\n+\tbool has_callx;\n+\t/*\n+\t * Call graph edges created by callx instructions. A bitmap of\n+\t * subprog_cnt * subprog_cnt bits, where bit (caller * subprog_cnt + callee)\n+\t * is set when the main verification pass sees 'caller' calling 'callee'\n+\t * via callx. Allocated when the first such edge is recorded.\n+\t */\n+\tunsigned long *callx_edges;\n \tunion {\n \t\tstruct bpf_idmap idmap_scratch;\n \t\tstruct bpf_idset idset_scratch;\n@@ -1089,6 +1119,12 @@ static inline bool bpf_pseudo_kfunc_call(const struct bpf_insn *insn)\n \t insn-\u003esrc_reg == BPF_PSEUDO_KFUNC_CALL;\n }\n \n+/* callx: indirect call of a bpf subprog whose address is in insn-\u003edst_reg */\n+static inline bool bpf_is_callx(const struct bpf_insn *insn)\n+{\n+\treturn insn-\u003ecode == (BPF_JMP | BPF_CALL | BPF_X);\n+}\n+\n __printf(2, 0) void bpf_verifier_vlog(struct bpf_verifier_log *log,\n \t\t\t\t const char *fmt, va_list args);\n __printf(2, 3) void bpf_verifier_log_write(struct bpf_verifier_env *env,\n@@ -1311,6 +1347,13 @@ static inline bool bt_is_frame_slot_set(struct backtrack_state *bt, u32 frame, u\n }\n \n bool bpf_map_is_rdonly(const struct bpf_map *map);\n+struct bpf_func_ptr *bpf_map_func_ptrs(struct bpf_verifier_env *env,\n+\t\t\t\t const struct bpf_map *map, u32 *cnt);\n+struct bpf_func_ptr *bpf_map_range_func_ptrs(struct bpf_verifier_env *env,\n+\t\t\t\t\t const struct bpf_map *map,\n+\t\t\t\t\t u64 off, u64 size, u32 *cnt);\n+void bpf_adjust_func_ptrs(struct bpf_verifier_env *env, u32 off, u32 len);\n+void bpf_adjust_func_ptrs_after_remove(struct bpf_verifier_env *env, u32 off, u32 len);\n int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,\n \t\t\tbool is_ldsx);\n \ndiff --git a/include/linux/filter.h b/include/linux/filter.h\nindex 422284b4fa96f..4f0662e428970 100644\n--- a/include/linux/filter.h\n+++ b/include/linux/filter.h\n@@ -1237,6 +1237,7 @@ bool bpf_jit_inlines_helper_call(s32 imm);\n bool bpf_jit_supports_subprog_tailcalls(void);\n bool bpf_jit_supports_percpu_insn(void);\n bool bpf_jit_supports_kfunc_call(void);\n+bool bpf_jit_supports_callx(void);\n bool bpf_jit_supports_kfunc_ret_reg_pair(void);\n bool bpf_jit_supports_stack_args(void);\n bool bpf_jit_supports_arena_args(void);\ndiff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c\nindex 507a366dffa47..4da99dec08184 100644\n--- a/kernel/bpf/backtrack.c\n+++ b/kernel/bpf/backtrack.c\n@@ -406,15 +406,18 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,\n \t\tif (class == BPF_STX)\n \t\t\tbt_set_reg(bt, sreg);\n \t} else if (class == BPF_JMP || class == BPF_JMP32) {\n-\t\tif (bpf_pseudo_call(insn)) {\n-\t\t\tint subprog_insn_idx, subprog;\n+\t\tif (bpf_pseudo_call(insn) || bpf_is_callx(insn)) {\n+\t\t\tint subprog_insn_idx, subprog = -1;\n \n-\t\t\tsubprog_insn_idx = idx + insn-\u003eimm + 1;\n-\t\t\tsubprog = bpf_find_subprog(env, subprog_insn_idx);\n-\t\t\tif (subprog \u003c 0)\n-\t\t\t\treturn -EFAULT;\n+\t\t\tif (bpf_pseudo_call(insn)) {\n+\t\t\t\tsubprog_insn_idx = idx + insn-\u003eimm + 1;\n+\t\t\t\tsubprog = bpf_find_subprog(env, subprog_insn_idx);\n+\t\t\t\tif (subprog \u003c 0)\n+\t\t\t\t\treturn -EFAULT;\n+\t\t\t}\n \n-\t\t\tif (bpf_subprog_is_global(env, subprog)) {\n+\t\t\t/* callx calls static subprogs only */\n+\t\t\tif (subprog \u003e= 0 \u0026\u0026 bpf_subprog_is_global(env, subprog)) {\n \t\t\t\t/* check that jump history doesn't have any\n \t\t\t\t * extra instructions from subprog; the next\n \t\t\t\t * instruction after call to global subprog\n@@ -536,7 +539,8 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,\n \t\t\t * never do that.\n \t\t\t */\n \t\t\tfrom_subprog_call = subseq_idx - 1 \u003e= 0 \u0026\u0026\n-\t\t\t\t\t bpf_pseudo_call(\u0026env-\u003eprog-\u003einsnsi[subseq_idx - 1]);\n+\t\t\t\t\t (bpf_pseudo_call(\u0026env-\u003eprog-\u003einsnsi[subseq_idx - 1]) ||\n+\t\t\t\t\t bpf_is_callx(\u0026env-\u003eprog-\u003einsnsi[subseq_idx - 1]));\n \n \t\t\tr0_precise = from_subprog_call \u0026\u0026 bt_is_reg_set(bt, BPF_REG_0);\n \t\t\tr2_precise = from_subprog_call \u0026\u0026 bt_is_reg_set(bt, BPF_REG_2);\ndiff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c\nindex 842c7d1eabccc..a068c191409a4 100644\n--- a/kernel/bpf/cfg.c\n+++ b/kernel/bpf/cfg.c\n@@ -421,6 +421,75 @@ static int visit_gotox_insn(int t, struct bpf_verifier_env *env)\n \treturn keep_exploring ? KEEP_EXPLORING : DONE_EXPLORING;\n }\n \n+/*\n+ * Return pointers to functions in the read-only map that ld_imm64 instruction\n+ * 't' loads the address of, or of its value, if there are any.\n+ */\n+static struct bpf_func_ptr *insn_func_ptrs(struct bpf_verifier_env *env, int t, u32 *cnt)\n+{\n+\tstruct bpf_insn *insn = \u0026env-\u003eprog-\u003einsnsi[t];\n+\n+\t*cnt = 0;\n+\tif (!env-\u003efunc_ptr_cnt || !bpf_is_ldimm64(insn))\n+\t\treturn NULL;\n+\tif (insn-\u003esrc_reg != BPF_PSEUDO_MAP_VALUE \u0026\u0026\n+\t insn-\u003esrc_reg != BPF_PSEUDO_MAP_IDX_VALUE \u0026\u0026\n+\t insn-\u003esrc_reg != BPF_PSEUDO_MAP_FD \u0026\u0026\n+\t insn-\u003esrc_reg != BPF_PSEUDO_MAP_IDX)\n+\t\treturn NULL;\n+\n+\treturn bpf_map_func_ptrs(env, env-\u003eused_maps[env-\u003einsn_aux_data[t].map_index], cnt);\n+}\n+\n+/*\n+ * ld_imm64 that loads the address of a map that has pointers to functions\n+ * is similar to ld_imm64 with BPF_PSEUDO_FUNC that loads the address of one\n+ * function: any of them may be read from the map and called via callx later.\n+ * Treat it as a call of all of them.\n+ */\n+static int visit_func_ptrs_insn(int t, struct bpf_verifier_env *env,\n+\t\t\t\tstruct bpf_func_ptr *ptrs, u32 cnt)\n+{\n+\tint *insn_stack = env-\u003ecfg.insn_stack;\n+\tint *insn_state = env-\u003ecfg.insn_state;\n+\tbool keep_exploring = false;\n+\tint ret, w;\n+\tu32 i;\n+\n+\tret = push_insn(t, t + 2, FALLTHROUGH, env);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\tmark_prune_point(env, t);\n+\tfor (i = 0; i \u003c cnt; i++) {\n+\t\tw = ptrs[i].xlated_off;\n+\n+\t\t/*\n+\t\t * This function is called until all functions are explored,\n+\t\t * so the effects are complete in the end.\n+\t\t */\n+\t\tmerge_callee_effects(env, t, w);\n+\n+\t\t/* the same marks as push_insn() leaves on a branch target */\n+\t\tmark_prune_point(env, w);\n+\t\tmark_jmp_point(env, w);\n+\t\tmark_jump_target(env, w);\n+\n+\t\t/* EXPLORED || DISCOVERED */\n+\t\tif (insn_state[w])\n+\t\t\tcontinue;\n+\n+\t\tif (env-\u003ecfg.cur_stack \u003e= env-\u003eprog-\u003elen)\n+\t\t\treturn -E2BIG;\n+\n+\t\tinsn_stack[env-\u003ecfg.cur_stack++] = w;\n+\t\tinsn_state[w] |= DISCOVERED;\n+\t\tkeep_exploring = true;\n+\t}\n+\n+\treturn keep_exploring ? KEEP_EXPLORING : DONE_EXPLORING;\n+}\n+\n /*\n * Instructions that can abnormally return from a subprog (tail_call\n * upon success, ld_{abs,ind} upon load failure) have a hidden exit\n@@ -453,11 +522,17 @@ static int visit_abnormal_return_insn(struct bpf_verifier_env *env, int t)\n static int visit_insn(int t, struct bpf_verifier_env *env)\n {\n \tstruct bpf_insn *insns = env-\u003eprog-\u003einsnsi, *insn = \u0026insns[t];\n+\tstruct bpf_func_ptr *ptrs;\n \tint ret, off, insn_sz;\n+\tu32 cnt;\n \n \tif (bpf_pseudo_func(insn))\n \t\treturn visit_func_call_insn(t, insns, env, true);\n \n+\tptrs = insn_func_ptrs(env, t, \u0026cnt);\n+\tif (ptrs)\n+\t\treturn visit_func_ptrs_insn(t, env, ptrs, cnt);\n+\n \t/* All non-branch instructions have a single fall-through edge. */\n \tif (BPF_CLASS(insn-\u003ecode) != BPF_JMP \u0026\u0026\n \t BPF_CLASS(insn-\u003ecode) != BPF_JMP32) {\ndiff --git a/kernel/bpf/const_fold.c b/kernel/bpf/const_fold.c\nindex 7f1b30059cc87..fea639f62b3f2 100644\n--- a/kernel/bpf/const_fold.c\n+++ b/kernel/bpf/const_fold.c\n@@ -180,9 +180,17 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *\n \t\tbool is_ldsx = mode == BPF_MEMSX;\n \t\tint off = src-\u003eval + insn-\u003eoff;\n \t\tu64 val = 0;\n+\t\tu32 cnt;\n \n+\t\t/*\n+\t\t * Values of insn_array map are addresses of jitted instructions,\n+\t\t * which are not known until the program is jitted.\n+\t\t */\n \t\tif (!bpf_map_is_rdonly(map) || !map-\u003eops-\u003emap_direct_value_addr ||\n+\t\t map-\u003emap_type == BPF_MAP_TYPE_INSN_ARRAY ||\n \t\t off \u003c 0 || off + size \u003e map-\u003evalue_size ||\n+\t\t /* so are the addresses of functions that the map points to */\n+\t\t bpf_map_range_func_ptrs(env, map, off, size, \u0026cnt) ||\n \t\t bpf_map_direct_read(map, off, size, \u0026val, is_ldsx)) {\n \t\t\t*dst = unknown;\n \t\t\tbreak;\n@@ -191,7 +199,8 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *\n \t\tdst-\u003eval = val;\n \t\tbreak;\n \tcase BPF_JMP:\n-\t\tif (opcode != BPF_CALL)\n+\t\t/* both 'call imm' and 'callx reg' clobber caller saved registers */\n+\t\tif (BPF_OP(insn-\u003ecode) != BPF_CALL)\n \t\t\tbreak;\n process_call:\n \t\tfor (r = BPF_REG_0; r \u003c= BPF_REG_5; r++)\ndiff --git a/kernel/bpf/core.c b/kernel/bpf/core.c\nindex 227211166dccf..273f74068068c 100644\n--- a/kernel/bpf/core.c\n+++ b/kernel/bpf/core.c\n@@ -1831,6 +1831,7 @@ bool bpf_opcode_in_insntable(u8 code)\n \t\t[BPF_LD | BPF_IND | BPF_H] = true,\n \t\t[BPF_LD | BPF_IND | BPF_W] = true,\n \t\t[BPF_JMP | BPF_JA | BPF_X] = true,\n+\t\t[BPF_JMP | BPF_CALL | BPF_X] = true,\n \t\t[BPF_JMP | BPF_JCOND] = true,\n \t};\n #undef BPF_INSN_3_TBL\n@@ -3028,6 +3029,11 @@ void __bpf_free_used_maps(struct bpf_prog_aux *aux,\n \t\t\tmap-\u003eops-\u003emap_poke_untrack(map, aux);\n \t\tif (sleepable)\n \t\t\tatomic64_dec(\u0026map-\u003esleepable_refcnt);\n+\t\t/*\n+\t\t * The program that didn't load is not a user of the map. libbpf\n+\t\t * loads the program again to get the log of the verifier.\n+\t\t */\n+\t\tcmpxchg(\u0026map-\u003euser, (unsigned long)aux, 0);\n \t\tbpf_map_put(map);\n \t}\n }\n@@ -3287,6 +3293,12 @@ bool __weak bpf_jit_supports_kfunc_call(void)\n \treturn false;\n }\n \n+/* Return TRUE if the JIT backend supports callx (indirect call) instruction. */\n+bool __weak bpf_jit_supports_callx(void)\n+{\n+\treturn false;\n+}\n+\n bool __weak bpf_jit_supports_kfunc_ret_reg_pair(void)\n {\n \treturn false;\ndiff --git a/kernel/bpf/disasm.c b/kernel/bpf/disasm.c\nindex 3ce8d74b0e400..36d3228d77454 100644\n--- a/kernel/bpf/disasm.c\n+++ b/kernel/bpf/disasm.c\n@@ -350,7 +350,10 @@ void print_bpf_insn(const struct bpf_insn_cbs *cbs,\n \t\tif (opcode == BPF_CALL) {\n \t\t\tchar tmp[64];\n \n-\t\t\tif (insn-\u003esrc_reg == BPF_PSEUDO_CALL) {\n+\t\t\tif (BPF_SRC(insn-\u003ecode) == BPF_X) {\n+\t\t\t\tverbose(cbs-\u003eprivate_data, \"(%02x) callx r%d\",\n+\t\t\t\t\tinsn-\u003ecode, insn-\u003edst_reg);\n+\t\t\t} else if (insn-\u003esrc_reg == BPF_PSEUDO_CALL) {\n \t\t\t\tverbose(cbs-\u003eprivate_data, \"(%02x) call pc%s\",\n \t\t\t\t\tinsn-\u003ecode,\n \t\t\t\t\t__func_get_name(cbs, insn,\ndiff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c\nindex 2add8001c3ec3..e568b9b790b5c 100644\n--- a/kernel/bpf/fixups.c\n+++ b/kernel/bpf/fixups.c\n@@ -361,6 +361,7 @@ struct bpf_prog *bpf_patch_insn_data(struct bpf_verifier_env *env, u32 off,\n \tadjust_insn_aux_data(env, new_prog, off, len, \u0026original_insn);\n \tadjust_subprog_starts(env, off, len);\n \tadjust_insn_arrays(env, off, len);\n+\tbpf_adjust_func_ptrs(env, off, len);\n \tadjust_poke_descs(new_prog, off, len);\n \treturn new_prog;\n }\n@@ -559,6 +560,9 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)\n \tif (err)\n \t\treturn err;\n \n+\t/* before subprogs are adjusted, since it looks at them */\n+\tbpf_adjust_func_ptrs_after_remove(env, off, cnt);\n+\n \terr = adjust_subprog_starts_after_remove(env, off, cnt);\n \tif (err)\n \t\treturn err;\n@@ -1285,6 +1289,57 @@ static int jit_subprogs(struct bpf_verifier_env *env)\n \t\tcond_resched();\n \t}\n \n+\t/*\n+\t * The addresses of all functions are final. Replace the offsets of\n+\t * functions with them in the maps of the program, see\n+\t * resolve_func_ptrs(). The program must be the only user of such map.\n+\t * From now on no other program can use it, see bpf_map_claim().\n+\t */\n+\tfor (i = 0; i \u003c env-\u003efunc_ptr_cnt; i++) {\n+\t\tstruct bpf_func_ptr *ptr = \u0026env-\u003efunc_ptrs[i];\n+\t\tunsigned long me = (unsigned long)prog-\u003eaux;\n+\t\tu64 addr, old, new = 0;\n+\n+\t\t/* pointers are sorted by map */\n+\t\tif ((!i || ptr-\u003emap != ptr[-1].map) \u0026\u0026\n+\t\t cmpxchg(\u0026ptr-\u003emap-\u003euser, me, me | BPF_MAP_USER_PATCHED) != me) {\n+\t\t\tverbose(env, \"map '%s' is used by another program\\n\", ptr-\u003emap-\u003ename);\n+\t\t\terr = -EBUSY;\n+\t\t\tgoto out_free;\n+\t\t}\n+\n+\t\t/* it's the address of the value of the map whatever the offset is */\n+\t\terr = ptr-\u003emap-\u003eops-\u003emap_direct_value_addr(ptr-\u003emap, \u0026addr, 0);\n+\t\tif (verifier_bug_if(err, env, \"no value of map '%s'\", ptr-\u003emap-\u003ename)) {\n+\t\t\terr = -EFAULT;\n+\t\t\tgoto out_free;\n+\t\t}\n+\t\taddr += ptr-\u003emap_off;\n+\n+\t\tif (ptr-\u003exlated_off != BPF_FUNC_PTR_DELETED) {\n+\t\t\tsubprog = bpf_find_subprog(env, ptr-\u003exlated_off);\n+\t\t\tif (verifier_bug_if(subprog \u003c= 0, env, \"no function at insn %u\",\n+\t\t\t\t\t ptr-\u003exlated_off)) {\n+\t\t\t\terr = -EFAULT;\n+\t\t\t\tgoto out_free;\n+\t\t\t}\n+\t\t\tnew = (unsigned long)func[subprog]-\u003ebpf_func;\n+\t\t} else if (verifier_bug_if(ptr-\u003eused, env, \"function of map '%s' offset %u is removed\",\n+\t\t\t\t\t ptr-\u003emap-\u003ename, ptr-\u003emap_off)) {\n+\t\t\t/* the program that reads the pointer might call the function */\n+\t\t\terr = -EFAULT;\n+\t\t\tgoto out_free;\n+\t\t}\n+\t\t/* else the function is dead code, nothing calls it, the pointer is NULL */\n+\n+\t\told = (u64)ptr-\u003eorig_off * sizeof(struct bpf_insn);\n+\t\tif (verifier_bug_if(cmpxchg64((u64 *)(unsigned long)addr, old, new) != old, env,\n+\t\t\t\t \"map '%s' offset %u changed\", ptr-\u003emap-\u003ename, ptr-\u003emap_off)) {\n+\t\t\terr = -EFAULT;\n+\t\t\tgoto out_free;\n+\t\t}\n+\t}\n+\n \t/*\n \t * Cleanup func[i]-\u003eaux fields which aren't required\n \t * or can become invalid in future\ndiff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c\nindex 44ecdc5b4ec2d..5aa2f68d92b37 100644\n--- a/kernel/bpf/liveness.c\n+++ b/kernel/bpf/liveness.c\n@@ -356,12 +356,25 @@ int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_\n \treturn 0;\n }\n \n+/*\n+ * Stack accesses of callbacks and of callx callees are not tracked by\n+ * func instances keyed by the @callsite. Callbacks might be called several\n+ * times and the callee of callx is not known when stack liveness is computed.\n+ * In both cases stack slots of the outer frames that might be read by the\n+ * callee are accounted as read by the @callsite instruction itself.\n+ */\n+static bool callee_stack_access_at_callsite(struct bpf_verifier_env *env, u32 callsite)\n+{\n+\treturn bpf_calls_callback(env, callsite) ||\n+\t bpf_is_callx(\u0026env-\u003eprog-\u003einsnsi[callsite]);\n+}\n+\n bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_spi)\n {\n \t/*\n \t * Slot is alive if it is read before q-\u003einsn_idx in current func instance,\n \t * or if for some outer func instance:\n-\t * - alive before callsite if callsite calls callback, otherwise\n+\t * - alive before callsite if callsite calls callback or is callx, otherwise\n \t * - alive after callsite\n \t */\n \tstruct live_stack_query *q = \u0026env-\u003eliveness-\u003elive_stack_query;\n@@ -394,7 +407,7 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp\n \t\t/* Get callsite from verifier state, not from instance callchain */\n \t\tcallsite = q-\u003ecallsites[i];\n \n-\t\talive = bpf_calls_callback(env, callsite)\n+\t\talive = callee_stack_access_at_callsite(env, callsite)\n \t\t\t? is_live_before(instance, callsite, rel, half_spi)\n \t\t\t: is_live_before(instance, callsite + 1, rel, half_spi);\n \t\tif (alive)\n@@ -1439,7 +1452,15 @@ static int record_call_access(struct bpf_verifier_env *env,\n \tif (bpf_pseudo_call(insn))\n \t\treturn 0;\n \n-\tif (bpf_get_call_summary(env, insn, \u0026cs))\n+\tif (bpf_is_callx(insn))\n+\t\t/*\n+\t\t * The callee is not known statically. Assume that all arg\n+\t\t * slots are passed and let record_arg_access() conservatively\n+\t\t * mark the stack of all frames as read if any of them is\n+\t\t * derived from a frame pointer.\n+\t\t */\n+\t\targ_slot_cnt = MAX_BPF_FUNC_REG_ARGS + MAX_STACK_ARG_SLOTS;\n+\telse if (bpf_get_call_summary(env, insn, \u0026cs))\n \t\targ_slot_cnt = cs.arg_slot_cnt;\n \n \tfor (r = BPF_REG_1; r \u003c BPF_REG_1 + min(arg_slot_cnt, MAX_BPF_FUNC_REG_ARGS); r++) {\n@@ -1533,7 +1554,8 @@ static void print_subprog_arg_access(struct bpf_verifier_env *env,\n \t\tbool has_extra = false;\n \t\tu8 cls = BPF_CLASS(insns[idx].code);\n \t\tbool is_ldx_stx_call = cls == BPF_LDX || cls == BPF_STX ||\n-\t\t\t\t insns[idx].code == (BPF_JMP | BPF_CALL);\n+\t\t\t\t insns[idx].code == (BPF_JMP | BPF_CALL) ||\n+\t\t\t\t bpf_is_callx(\u0026insns[idx]);\n \n \t\tverbose(env, \"%3d: \", idx);\n \t\tbpf_verbose_insn(env, \u0026insns[idx]);\n@@ -1722,7 +1744,7 @@ static int compute_subprog_args(struct bpf_verifier_env *env,\n \t\tif (err)\n \t\t\tgoto err_free;\n \n-\t\tif (insn-\u003ecode == (BPF_JMP | BPF_CALL)) {\n+\t\tif (insn-\u003ecode == (BPF_JMP | BPF_CALL) || bpf_is_callx(insn)) {\n \t\t\terr = record_call_access(env, instance, at_in[i], idx);\n \t\t\tif (err)\n \t\t\t\tgoto err_free;\n@@ -2202,6 +2224,9 @@ static void compute_insn_live_regs(struct bpf_verifier_env *env,\n \t\t\t\tuse = GENMASK(min_t(u8, cs.arg_slot_cnt, MAX_BPF_FUNC_REG_ARGS), 1);\n \t\t\tdef = mask_widen(def);\n \t\t\tuse = mask_widen(use);\n+\t\t\t/* callx reads the address of the callee from dst_reg */\n+\t\t\tif (bpf_is_callx(insn))\n+\t\t\t\tuse |= dst;\n \t\t\tbreak;\n \t\tdefault:\n \t\t\tdef = 0;\ndiff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c\nindex d62c0f74cff5e..8507114cf690a 100644\n--- a/kernel/bpf/verifier.c\n+++ b/kernel/bpf/verifier.c\n@@ -3018,6 +3018,22 @@ static int add_subprogs(struct bpf_verifier_env *env)\n \t\t\treturn ret;\n \t}\n \n+\t/*\n+\t * func_info describes all functions of the program. Those that are\n+\t * referenced only from data, e.g. from a table of functions to be\n+\t * called via callx, are not seen by the loop above. They are possible\n+\t * callees that have to be known upfront as well.\n+\t */\n+\tif (env-\u003ebpf_capable) {\n+\t\tstruct bpf_prog_aux *aux = env-\u003eprog-\u003eaux;\n+\n+\t\tfor (i = 1; i \u003c aux-\u003efunc_info_cnt; i++) {\n+\t\t\tret = add_subprog(env, aux-\u003efunc_info[i].insn_off);\n+\t\t\tif (ret \u003c 0)\n+\t\t\t\treturn ret;\n+\t\t}\n+\t}\n+\n \tret = bpf_find_exception_callback_insn_off(env);\n \tif (ret \u003c 0)\n \t\treturn ret;\n@@ -3099,6 +3115,8 @@ static int check_subprogs(struct bpf_verifier_env *env)\n \t\tif (BPF_CLASS(code) == BPF_LD \u0026\u0026\n \t\t (BPF_MODE(code) == BPF_ABS || BPF_MODE(code) == BPF_IND))\n \t\t\tsubprog[cur_subprog].has_ld_abs = true;\n+\t\tif (bpf_is_callx(\u0026insn[i]))\n+\t\t\tenv-\u003ehas_callx = true;\n \t\tif (BPF_CLASS(code) != BPF_JMP \u0026\u0026 BPF_CLASS(code) != BPF_JMP32)\n \t\t\tgoto next;\n \t\tif (BPF_OP(code) == BPF_CALL)\n@@ -3144,12 +3162,50 @@ static int check_subprogs(struct bpf_verifier_env *env)\n \treturn 0;\n }\n \n+/*\n+ * The callee of callx is known to the main verification pass only, which\n+ * records the 'caller' -\u003e 'callee' edge of the call graph for the checks\n+ * that follow it: absence of recursion and the maximum stack depth.\n+ */\n+static int record_callx_edge(struct bpf_verifier_env *env, int caller, int callee)\n+{\n+\tu32 cnt = env-\u003esubprog_cnt;\n+\n+\tif (!env-\u003ecallx_edges) {\n+\t\tenv-\u003ecallx_edges = kvcalloc(BITS_TO_LONGS(cnt * cnt), sizeof(long),\n+\t\t\t\t\t GFP_KERNEL_ACCOUNT);\n+\t\tif (!env-\u003ecallx_edges)\n+\t\t\treturn -ENOMEM;\n+\t}\n+\t__set_bit(caller * cnt + callee, env-\u003ecallx_edges);\n+\treturn 0;\n+}\n+\n+/*\n+ * Return the first subprog with the number \u003e= 'from' that 'caller' calls\n+ * via callx, or -1 when there is none.\n+ */\n+static int next_callx_callee(struct bpf_verifier_env *env, int caller, int from)\n+{\n+\tu32 cnt = env-\u003esubprog_cnt;\n+\tunsigned long bit, end = (caller + 1) * cnt;\n+\n+\tif (!env-\u003ecallx_edges || from \u003e= cnt)\n+\t\treturn -1;\n+\tbit = find_next_bit(env-\u003ecallx_edges, end, caller * cnt + from);\n+\treturn bit \u003c end ? bit - caller * cnt : -1;\n+}\n+\n /*\n * Sort subprogs in topological order so that leaf subprogs come first and\n * their callers come later. This is a DFS post-order traversal of the call\n * graph. Scan only reachable instructions (those in the computed postorder) of\n * the current subprog to discover callees (direct subprogs and sync\n * callbacks).\n+ *\n+ * The callees of callx are not known before the main verification pass.\n+ * When callx is used the sort is repeated after it with the recorded callx\n+ * edges added to the call graph to reject recursion through indirect calls.\n */\n static int sort_subprogs_topo(struct bpf_verifier_env *env)\n {\n@@ -3190,12 +3246,22 @@ static int sort_subprogs_topo(struct bpf_verifier_env *env)\n \t\t\t\tint idx = insn_postorder[j];\n \t\t\t\tint callee;\n \n-\t\t\t\tif (!bpf_pseudo_call(\u0026insn[idx]) \u0026\u0026 !bpf_pseudo_func(\u0026insn[idx]))\n+\t\t\t\tif (bpf_is_callx(\u0026insn[idx])) {\n+\t\t\t\t\t/* find a callee that is not explored yet */\n+\t\t\t\t\tcallee = -1;\n+\t\t\t\t\tdo {\n+\t\t\t\t\t\tcallee = next_callx_callee(env, cur, callee + 1);\n+\t\t\t\t\t} while (callee \u003e= 0 \u0026\u0026 color[callee] == 2);\n+\t\t\t\t\tif (callee \u003c 0)\n+\t\t\t\t\t\tcontinue;\n+\t\t\t\t} else if (bpf_pseudo_call(\u0026insn[idx]) || bpf_pseudo_func(\u0026insn[idx])) {\n+\t\t\t\t\tcallee = bpf_find_subprog(env, idx + insn[idx].imm + 1);\n+\t\t\t\t\tif (callee \u003c 0) {\n+\t\t\t\t\t\tret = -EFAULT;\n+\t\t\t\t\t\tgoto out;\n+\t\t\t\t\t}\n+\t\t\t\t} else {\n \t\t\t\t\tcontinue;\n-\t\t\t\tcallee = bpf_find_subprog(env, idx + insn[idx].imm + 1);\n-\t\t\t\tif (callee \u003c 0) {\n-\t\t\t\t\tret = -EFAULT;\n-\t\t\t\t\tgoto out;\n \t\t\t\t}\n \t\t\t\tif (color[callee] == 2)\n \t\t\t\t\tcontinue;\n@@ -5367,6 +5433,9 @@ struct bpf_subprog_call_depth_info {\n \tint ret_insn; /* caller instruction where we return to. */\n \tint caller; /* caller subprogram idx */\n \tint frame; /* # of consecutive static call stack frames on top of stack */\n+\tint callx_insn; /* callx instruction whose callees are being walked */\n+\tint callx_next; /* next callee of callx_insn to walk */\n+\tbool via_callx; /* the subprogram is entered via callx */\n };\n \n /* starting from main bpf function walk all instructions of the function\n@@ -5386,11 +5455,13 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,\n \n \t/* no caller idx */\n \tdinfo[idx].caller = -1;\n+\tdinfo[idx].via_callx = false;\n \n \ti = subprog[idx].start;\n \tif (!priv_stack_supported)\n \t\tsubprog[idx].priv_stack_mode = NO_PRIV_STACK;\n process_func:\n+\tdinfo[idx].callx_insn = -1;\n \tif (subprog[idx].has_ld_abs) {\n \t\tfor (tmp = idx; tmp \u003e= 0; tmp = dinfo[tmp].caller) {\n \t\t\tif (subprog[tmp].is_cb) {\n@@ -5486,31 +5557,64 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,\n \t\t\treturn -EINVAL;\n \t\t}\n \n-\t\tif (!bpf_pseudo_call(insn + i) \u0026\u0026 !bpf_pseudo_func(insn + i))\n+\t\tif (bpf_is_callx(insn + i)) {\n+\t\t\t/*\n+\t\t\t * Walk the callees recorded by the main verification\n+\t\t\t * pass one by one, returning to this insn after each.\n+\t\t\t */\n+\t\t\tif (dinfo[idx].callx_insn != i) {\n+\t\t\t\tdinfo[idx].callx_insn = i;\n+\t\t\t\tdinfo[idx].callx_next = 0;\n+\t\t\t}\n+\t\t\tsidx = next_callx_callee(env, idx, dinfo[idx].callx_next);\n+\t\t\tif (sidx \u003c 0)\n+\t\t\t\tcontinue;\n+\t\t\tdinfo[idx].callx_next = sidx + 1;\n+\t\t\tdinfo[idx].ret_insn = i;\n+\t\t\tnext_insn = subprog[sidx].start;\n+\t\t} else if (bpf_pseudo_call(insn + i) || bpf_pseudo_func(insn + i)) {\n+\t\t\t/* find the callee */\n+\t\t\tnext_insn = i + insn[i].imm + 1;\n+\t\t\tsidx = bpf_find_subprog(env, next_insn);\n+\t\t\tif (verifier_bug_if(sidx \u003c 0, env, \"callee not found at insn %d\", next_insn))\n+\t\t\t\treturn -EFAULT;\n+\t\t\tif (subprog[sidx].is_async_cb) {\n+\t\t\t\t/* async callbacks don't increase bpf prog stack size unless called directly */\n+\t\t\t\tif (!bpf_pseudo_call(insn + i))\n+\t\t\t\t\tcontinue;\n+\t\t\t\tif (subprog[sidx].is_exception_cb) {\n+\t\t\t\t\tverbose(env, \"insn %d cannot call exception cb directly\", i);\n+\t\t\t\t\treturn -EINVAL;\n+\t\t\t\t}\n+\t\t\t}\n+\t\t\t/* remember insn to return to */\n+\t\t\tdinfo[idx].ret_insn = i + 1;\n+\t\t} else {\n \t\t\tcontinue;\n-\t\t/* remember insn and function to return to */\n+\t\t}\n \n-\t\t/* find the callee */\n-\t\tnext_insn = i + insn[i].imm + 1;\n-\t\tsidx = bpf_find_subprog(env, next_insn);\n-\t\tif (verifier_bug_if(sidx \u003c 0, env, \"callee not found at insn %d\", next_insn))\n-\t\t\treturn -EFAULT;\n-\t\tif (subprog[sidx].is_async_cb) {\n-\t\t\t/* async callbacks don't increase bpf prog stack size unless called directly */\n-\t\t\tif (!bpf_pseudo_call(insn + i))\n+\t\t/*\n+\t\t * sort_subprogs_topo() tolerates cycles in the call graph that\n+\t\t * go through the address of a function being taken, since it\n+\t\t * doesn't know what it is taken for. Such cycle is a recursion\n+\t\t * unless it's an async callback, which are skipped above.\n+\t\t * The main verification pass limits the depth of the recursion,\n+\t\t * but it doesn't follow calls of global functions.\n+\t\t */\n+\t\tfor (tmp = idx; tmp \u003e= 0; tmp = dinfo[tmp].caller) {\n+\t\t\tif (tmp != sidx)\n \t\t\t\tcontinue;\n-\t\t\tif (subprog[sidx].is_exception_cb) {\n-\t\t\t\tverbose(env, \"insn %d cannot call exception cb directly\", i);\n-\t\t\t\treturn -EINVAL;\n-\t\t\t}\n+\t\t\tverbose(env, \"recursive call from %s() to %s()\\n\",\n+\t\t\t\tbpf_subprog_name(env, idx), bpf_subprog_name(env, sidx));\n+\t\t\treturn -EINVAL;\n \t\t}\n \n \t\t/* store caller info for after we return from callee */\n \t\tdinfo[idx].frame = frame;\n-\t\tdinfo[idx].ret_insn = i + 1;\n \n \t\t/* push caller idx into callee's dinfo */\n \t\tdinfo[sidx].caller = idx;\n+\t\tdinfo[sidx].via_callx = bpf_is_callx(insn + i);\n \n \t\ti = next_insn;\n \n@@ -5544,6 +5648,14 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,\n \t\t\t\tverbose(env, \"tail_calls are not allowed in programs with stack args\\n\");\n \t\t\t\treturn -EINVAL;\n \t\t\t}\n+\t\t\t/*\n+\t\t\t * JITs pass tail call counter in a register that is\n+\t\t\t * not available when the callee is called via callx.\n+\t\t\t */\n+\t\t\tif (dinfo[tmp].via_callx) {\n+\t\t\t\tverbose(env, \"tail_calls are not allowed in functions called via callx\\n\");\n+\t\t\t\treturn -EINVAL;\n+\t\t\t}\n \t\t\tsubprog[tmp].tail_call_reachable = true;\n \t\t}\n \t} else if (!idx \u0026\u0026 subprog[0].has_tail_call \u0026\u0026 subprog[0].stack_arg_cnt) {\n@@ -5908,6 +6020,110 @@ int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,\n \treturn 0;\n }\n \n+static int cmp_func_ptrs(const void *_a, const void *_b)\n+{\n+\tconst struct bpf_func_ptr *a = _a, *b = _b;\n+\n+\tif (a-\u003emap != b-\u003emap)\n+\t\treturn a-\u003emap \u003c b-\u003emap ? -1 : 1;\n+\tif (a-\u003emap_off != b-\u003emap_off)\n+\t\treturn a-\u003emap_off \u003c b-\u003emap_off ? -1 : 1;\n+\treturn 0;\n+}\n+\n+/* Find the first pointer to a function at or after 'off' in the value of 'map' */\n+static u32 func_ptr_lower_bound(struct bpf_verifier_env *env, const struct bpf_map *map, u64 off)\n+{\n+\tu32 l = 0, r = env-\u003efunc_ptr_cnt, m;\n+\tstruct bpf_func_ptr *p;\n+\n+\twhile (l \u003c r) {\n+\t\tm = l + (r - l) / 2;\n+\t\tp = \u0026env-\u003efunc_ptrs[m];\n+\t\tif (p-\u003emap \u003c map || (p-\u003emap == map \u0026\u0026 p-\u003emap_off \u003c off))\n+\t\t\tl = m + 1;\n+\t\telse\n+\t\t\tr = m;\n+\t}\n+\treturn l;\n+}\n+\n+/*\n+ * Return pointers to functions that overlap with 'size' bytes at offset 'off'\n+ * of the value of 'map' and their number in 'cnt'.\n+ */\n+struct bpf_func_ptr *bpf_map_range_func_ptrs(struct bpf_verifier_env *env,\n+\t\t\t\t\t const struct bpf_map *map,\n+\t\t\t\t\t u64 off, u64 size, u32 *cnt)\n+{\n+\tu32 first, last;\n+\n+\t*cnt = 0;\n+\tif (!env-\u003efunc_ptr_cnt || !size)\n+\t\treturn NULL;\n+\n+\t/* a pointer that starts up to 7 bytes before 'off' overlaps too */\n+\tfirst = func_ptr_lower_bound(env, map, off \u003e= sizeof(u64) ? off - sizeof(u64) + 1 : 0);\n+\tlast = func_ptr_lower_bound(env, map, off + size);\n+\tif (first \u003e= last)\n+\t\treturn NULL;\n+\n+\t*cnt = last - first;\n+\treturn \u0026env-\u003efunc_ptrs[first];\n+}\n+\n+/* Return all pointers to functions in the value of 'map' */\n+struct bpf_func_ptr *bpf_map_func_ptrs(struct bpf_verifier_env *env,\n+\t\t\t\t const struct bpf_map *map, u32 *cnt)\n+{\n+\treturn bpf_map_range_func_ptrs(env, map, 0, (u64)map-\u003evalue_size, cnt);\n+}\n+\n+/* instructions [off, off + len) replaced the instruction at 'off' */\n+void bpf_adjust_func_ptrs(struct bpf_verifier_env *env, u32 off, u32 len)\n+{\n+\tstruct bpf_func_ptr *p;\n+\tu32 i;\n+\n+\tif (len \u003c= 1)\n+\t\treturn;\n+\n+\tfor (i = 0; i \u003c env-\u003efunc_ptr_cnt; i++) {\n+\t\tp = \u0026env-\u003efunc_ptrs[i];\n+\t\tif (p-\u003exlated_off \u003c= off || p-\u003exlated_off == BPF_FUNC_PTR_DELETED)\n+\t\t\tcontinue;\n+\t\tp-\u003exlated_off += len - 1;\n+\t}\n+}\n+\n+/*\n+ * Instructions [off, off + len) are about to be removed. It's called before\n+ * the starts of subprogs are adjusted. A subprog is gone when all of its\n+ * instructions are. Otherwise, e.g. when its first instruction is a nop,\n+ * it starts where the removed instructions did.\n+ */\n+void bpf_adjust_func_ptrs_after_remove(struct bpf_verifier_env *env, u32 off, u32 len)\n+{\n+\tstruct bpf_func_ptr *p;\n+\tint subprog;\n+\tu32 i;\n+\n+\tfor (i = 0; i \u003c env-\u003efunc_ptr_cnt; i++) {\n+\t\tp = \u0026env-\u003efunc_ptrs[i];\n+\t\tif (p-\u003exlated_off \u003c off || p-\u003exlated_off == BPF_FUNC_PTR_DELETED)\n+\t\t\tcontinue;\n+\t\tif (p-\u003exlated_off \u003e= off + len) {\n+\t\t\tp-\u003exlated_off -= len;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tsubprog = bpf_find_subprog(env, p-\u003exlated_off);\n+\t\tif (subprog \u003e 0 \u0026\u0026 env-\u003esubprog_info[subprog + 1].start \u003e off + len)\n+\t\t\tp-\u003exlated_off = off;\n+\t\telse\n+\t\t\tp-\u003exlated_off = BPF_FUNC_PTR_DELETED;\n+\t}\n+}\n+\n #define BTF_TYPE_SAFE_RCU(__type) __PASTE(__type, __safe_rcu)\n #define BTF_TYPE_SAFE_RCU_OR_NULL(__type) __PASTE(__type, __safe_rcu_or_null)\n #define BTF_TYPE_SAFE_TRUSTED(__type) __PASTE(__type, __safe_trusted)\n@@ -6417,12 +6633,100 @@ static void add_scalar_to_reg(struct bpf_reg_state *dst_reg, s64 val)\n \treg_bounds_sync(dst_reg);\n }\n \n+static void mark_reg_func_ptr(struct bpf_verifier_env *env, struct bpf_reg_state *regs,\n+\t\t\t int regno, int subprog)\n+{\n+\tmark_reg_known_zero(env, regs, regno);\n+\tregs[regno].type = PTR_TO_FUNC;\n+\tregs[regno].subprogno = subprog;\n+}\n+\n+/* a read from a table of functions branches into that many states at most */\n+#define BPF_MAX_FUNC_PTR_TARGETS 64\n+/* and the table, which might have other data in it, is that many pointers long at most */\n+#define BPF_MAX_FUNC_PTR_RANGE 4096\n+\n+/*\n+ * A read from a frozen read-only map that has pointers to functions, see\n+ * resolve_func_ptrs(). A read of exactly one pointer yields PTR_TO_FUNC.\n+ * When the offset is variable and only pointers can be read, which is how\n+ * an element of a table of functions is loaded, the verification continues\n+ * with each of them. Other reads that overlap with a pointer are rejected,\n+ * because their result is not known until the program is jitted.\n+ *\n+ * Return -ENOENT if there are no pointers to functions in the bytes that are read.\n+ */\n+static int check_func_ptr_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,\n+\t\t\t int size, int value_regno)\n+{\n+\tu64 min_off = reg_umin(reg) + off, max_off = reg_umax(reg) + off;\n+\tstruct tnum offs = tnum_add(reg-\u003evar_off, tnum_const(off));\n+\tstruct bpf_reg_state *regs = cur_regs(env);\n+\tstruct bpf_map *map = reg-\u003emap_ptr;\n+\tstruct bpf_verifier_state *branch;\n+\tint subprog, targets[BPF_MAX_FUNC_PTR_TARGETS];\n+\tstruct bpf_func_ptr *ptrs;\n+\tu32 i, cnt, n = 0;\n+\tu64 o;\n+\n+\tptrs = bpf_map_range_func_ptrs(env, map, min_off, max_off - min_off + size, \u0026cnt);\n+\tif (!ptrs)\n+\t\treturn -ENOENT;\n+\n+\tif (size != sizeof(u64) || value_regno \u003c 0 || !tnum_is_aligned(offs, sizeof(u64)) ||\n+\t (max_off - min_off) / sizeof(u64) \u003e BPF_MAX_FUNC_PTR_RANGE)\n+\t\tgoto overlap;\n+\n+\t/*\n+\t * Every offset that the read is possible at has to be the offset of\n+\t * a pointer. var_off tells the stride of the elements of an array.\n+\t */\n+\tfor (o = round_up(min_off, sizeof(u64)), i = 0; o \u003c= max_off; o += sizeof(u64)) {\n+\t\tif ((o ^ offs.value) \u0026 ~offs.mask)\n+\t\t\tcontinue;\n+\t\twhile (i \u003c cnt \u0026\u0026 ptrs[i].map_off \u003c o)\n+\t\t\ti++;\n+\t\tif (i == cnt || ptrs[i].map_off != o)\n+\t\t\tgoto overlap;\n+\t\tif (n == BPF_MAX_FUNC_PTR_TARGETS) {\n+\t\t\tverbose(env, \"read from map '%s' may yield more than %d pointers to functions\\n\",\n+\t\t\t\tmap-\u003ename, BPF_MAX_FUNC_PTR_TARGETS);\n+\t\t\treturn -E2BIG;\n+\t\t}\n+\t\tsubprog = bpf_find_subprog(env, ptrs[i].xlated_off);\n+\t\tif (verifier_bug_if(subprog \u003c= 0, env, \"no function at insn %u for map '%s' offset %u\",\n+\t\t\t\t ptrs[i].xlated_off, map-\u003ename, ptrs[i].map_off))\n+\t\t\treturn -EFAULT;\n+\t\tptrs[i].used = true;\n+\t\ttargets[n++] = subprog;\n+\t}\n+\tif (verifier_bug_if(!n, env, \"no offsets to read map '%s' at\", map-\u003ename))\n+\t\treturn -EFAULT;\n+\n+\tfor (i = 0; i \u003c n - 1; i++) {\n+\t\tbranch = push_stack(env, env-\u003einsn_idx + 1, env-\u003einsn_idx,\n+\t\t\t\t env-\u003ecur_state-\u003especulative);\n+\t\tif (IS_ERR(branch))\n+\t\t\treturn PTR_ERR(branch);\n+\t\tmark_reg_func_ptr(env, branch-\u003eframe[branch-\u003ecurframe]-\u003eregs, value_regno,\n+\t\t\t\t targets[i]);\n+\t}\n+\tmark_reg_func_ptr(env, regs, value_regno, targets[n - 1]);\n+\treturn 0;\n+\n+overlap:\n+\tverbose(env, \"read of %d bytes at offset [%llu,%llu] of map '%s' overlaps with a pointer to a function\\n\",\n+\t\tsize, min_off, max_off, map-\u003ename);\n+\treturn -EACCES;\n+}\n+\n static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,\n \t\t\t int bpf_size, int value_regno, bool is_ldsx)\n {\n \tstruct bpf_reg_state *regs = cur_regs(env);\n \tint size = bpf_size_to_bytes(bpf_size);\n \tstruct bpf_map *map = reg-\u003emap_ptr;\n+\tint err;\n \n \tswitch (map-\u003emap_type) {\n \tcase BPF_MAP_TYPE_INSN_ARRAY:\n@@ -6440,13 +6744,18 @@ static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state\n \t\tbreak;\n \t}\n \n+\tif (env-\u003efunc_ptr_cnt) {\n+\t\terr = check_func_ptr_read(env, reg, off, size, value_regno);\n+\t\tif (err != -ENOENT)\n+\t\t\treturn err;\n+\t}\n+\n \t/* If map is read-only, track its contents as scalars. */\n \tif (tnum_is_const(reg-\u003evar_off) \u0026\u0026\n \t bpf_map_is_rdonly(map) \u0026\u0026\n \t map-\u003eops-\u003emap_direct_value_addr) {\n \t\tint map_off = off + reg-\u003evar_off.value;\n \t\tu64 val = 0;\n-\t\tint err;\n \n \t\terr = bpf_map_direct_read(map, map_off, size, \u0026val, is_ldsx);\n \t\tif (err)\n@@ -8754,6 +9063,7 @@ static int check_arg_const_str(struct bpf_verifier_env *env,\n \tint map_off;\n \tu64 map_addr;\n \tchar *str_ptr;\n+\tu32 cnt;\n \n \tif (reg-\u003etype != PTR_TO_MAP_VALUE)\n \t\treturn -EINVAL;\n@@ -8803,6 +9113,11 @@ static int check_arg_const_str(struct bpf_verifier_env *env,\n \t\tverbose(env, \"string is not zero-terminated\\n\");\n \t\treturn -EINVAL;\n \t}\n+\t/* the bytes of a pointer to a function are not known until the program is jitted */\n+\tif (bpf_map_range_func_ptrs(env, map, map_off, strlen(str_ptr + map_off) + 1, \u0026cnt)) {\n+\t\tverbose(env, \"string overlaps with a pointer to a function\\n\");\n+\t\treturn -EACCES;\n+\t}\n \treturn 0;\n }\n \n@@ -10584,13 +10899,61 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins\n static int process_bpf_exit_full(struct bpf_verifier_env *env,\n \t\t\t\t bool *do_print_state, bool exception_exit);\n \n-static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,\n-\t\t\t int *insn_idx)\n+/*\n+ * Call of a static subprog. The callee is verified in the context of\n+ * the caller, hence set up a new frame and continue from the first\n+ * instruction of the callee.\n+ */\n+static int check_static_func_call(struct bpf_verifier_env *env, int subprog,\n+\t\t\t\t int *insn_idx)\n {\n \tstruct bpf_verifier_state *state = env-\u003ecur_state;\n \tstruct bpf_subprog_info *caller_info;\n \tu16 callee_incoming, stack_arg_cnt;\n \tstruct bpf_func_state *caller;\n+\tint err;\n+\n+\tcaller = state-\u003eframe[state-\u003ecurframe];\n+\n+\t/*\n+\t * Track caller's total stack arg count (incoming + max outgoing).\n+\t * This is needed so the JIT knows how much stack arg space to allocate.\n+\t */\n+\tcaller_info = \u0026env-\u003esubprog_info[caller-\u003esubprogno];\n+\tcallee_incoming = bpf_in_stack_arg_cnt(\u0026env-\u003esubprog_info[subprog]);\n+\tstack_arg_cnt = bpf_in_stack_arg_cnt(caller_info) + callee_incoming;\n+\tif (stack_arg_cnt \u003e caller_info-\u003estack_arg_cnt)\n+\t\tcaller_info-\u003estack_arg_cnt = stack_arg_cnt;\n+\n+\t/*\n+\t * For regular function entry setup new frame and continue\n+\t * from that frame.\n+\t */\n+\terr = setup_func_entry(env, subprog, *insn_idx, set_callee_state, state);\n+\tif (err)\n+\t\treturn err;\n+\n+\tbpf_diag_record_scrub(env, \u0026caller-\u003eregs[BPF_REG_0], BPF_DIAG_MOD_CALLER_SAVED);\n+\tclear_caller_saved_regs(env, caller-\u003eregs);\n+\n+\t/* and go analyze first insn of the callee */\n+\t*insn_idx = env-\u003esubprog_info[subprog].start - 1;\n+\n+\tif (env-\u003elog.level \u0026 BPF_LOG_LEVEL) {\n+\t\tverbose(env, \"caller:\\n\");\n+\t\tprint_verifier_state(env, state, caller-\u003eframeno, true);\n+\t\tverbose(env, \"callee:\\n\");\n+\t\tprint_verifier_state(env, state, state-\u003ecurframe, true);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,\n+\t\t\t int *insn_idx)\n+{\n+\tstruct bpf_verifier_state *state = env-\u003ecur_state;\n+\tstruct bpf_func_state *caller;\n \tint err, subprog, target_insn;\n \tu32 i, nregs;\n \n@@ -10679,37 +11042,75 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,\n \t\treturn 0;\n \t}\n \n-\t/*\n-\t * Track caller's total stack arg count (incoming + max outgoing).\n-\t * This is needed so the JIT knows how much stack arg space to allocate.\n-\t */\n-\tcaller_info = \u0026env-\u003esubprog_info[caller-\u003esubprogno];\n-\tcallee_incoming = bpf_in_stack_arg_cnt(\u0026env-\u003esubprog_info[subprog]);\n-\tstack_arg_cnt = bpf_in_stack_arg_cnt(caller_info) + callee_incoming;\n-\tif (stack_arg_cnt \u003e caller_info-\u003estack_arg_cnt)\n-\t\tcaller_info-\u003estack_arg_cnt = stack_arg_cnt;\n+\treturn check_static_func_call(env, subprog, insn_idx);\n+}\n \n-\t/* for regular function entry setup new frame and continue\n-\t * from that frame.\n-\t */\n-\terr = setup_func_entry(env, subprog, *insn_idx, set_callee_state, state);\n+/*\n+ * callx dst_reg: call a bpf subprog whose address is in dst_reg.\n+ *\n+ * The address of a subprog is either loaded into a register by ld_imm64 with\n+ * src_reg == BPF_PSEUDO_FUNC, or it is read from a frozen read-only map, see\n+ * resolve_func_ptrs(). Both are possible for static subprogs only. Hence all\n+ * possible callees of callx are discovered by add_subprogs() and are reachable\n+ * in the control flow graph before the main verification pass begins.\n+ * PTR_TO_FUNC register identifies the callee, so from here on callx is verified\n+ * as a direct call of that static subprog.\n+ */\n+static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,\n+\t\t\t int *insn_idx)\n+{\n+\tstruct bpf_func_state *caller = cur_func(env);\n+\tstruct bpf_reg_state *reg;\n+\tconst char *reason;\n+\tint err, subprog;\n+\n+\terr = check_reg_arg(env, insn-\u003edst_reg, SRC_OP);\n \tif (err)\n \t\treturn err;\n \n-\tbpf_diag_record_scrub(env, \u0026caller-\u003eregs[BPF_REG_0], BPF_DIAG_MOD_CALLER_SAVED);\n-\tclear_caller_saved_regs(env, caller-\u003eregs);\n+\treg = reg_state(env, insn-\u003edst_reg);\n+\tif (reg-\u003etype != PTR_TO_FUNC) {\n+\t\tverbose(env, \"R%d has type %s, expected func\\n\", insn-\u003edst_reg,\n+\t\t\treg_type_str(env, reg-\u003etype));\n+\t\treason = bpf_diag_fmt(\n+\t\t\tenv, \"R%d holds %s, but callx can only call through the address of a static BPF function.\",\n+\t\t\tinsn-\u003edst_reg, bpf_diag_reg_type_plain(env, reg-\u003etype));\n+\t\tbpf_diag_register_type(\n+\t\t\tenv, *insn_idx, insn-\u003edst_reg, \"indirect call through a non-function pointer\", reason,\n+\t\t\t\"Load the address of a static BPF function into the register before callx.\");\n+\t\treturn -EACCES;\n+\t}\n \n-\t/* and go analyze first insn of the callee */\n-\t*insn_idx = env-\u003esubprog_info[subprog].start - 1;\n+\t/*\n+\t * Arithmetic on PTR_TO_FUNC is allowed, but only unmodified address\n+\t * of a subprog can be called.\n+\t */\n+\terr = check_ptr_off_reg(env, reg, insn-\u003edst_reg);\n+\tif (err)\n+\t\treturn err;\n \n-\tif (env-\u003elog.level \u0026 BPF_LOG_LEVEL) {\n-\t\tverbose(env, \"caller:\\n\");\n-\t\tprint_verifier_state(env, state, caller-\u003eframeno, true);\n-\t\tverbose(env, \"callee:\\n\");\n-\t\tprint_verifier_state(env, state, state-\u003ecurframe, true);\n+\t/* there is no support for callx in the interpreter */\n+\tif (!env-\u003eprog-\u003ejit_requested) {\n+\t\tverbose(env, \"JIT is required to use callx\\n\");\n+\t\treturn -EOPNOTSUPP;\n+\t}\n+\tif (!bpf_jit_supports_callx()) {\n+\t\tverbose(env, \"JIT doesn't support callx\\n\");\n+\t\treturn -EOPNOTSUPP;\n \t}\n+\tenv-\u003eprog-\u003ejit_required = true;\n \n-\treturn 0;\n+\t/* PTR_TO_FUNC is a pointer to a static subprog */\n+\tsubprog = reg-\u003esubprogno;\n+\terr = btf_check_subprog_call(env, subprog, caller-\u003eregs);\n+\tif (err == -EFAULT)\n+\t\treturn err;\n+\n+\terr = record_callx_edge(env, caller-\u003esubprogno, subprog);\n+\tif (err)\n+\t\treturn err;\n+\n+\treturn check_static_func_call(env, subprog, insn_idx);\n }\n \n int map_set_for_each_callback_args(struct bpf_verifier_env *env,\n@@ -18709,7 +19110,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)\n \n \t\tenv-\u003ejmps_processed++;\n \t\tif (opcode == BPF_CALL) {\n-\t\t\tif (env-\u003ecur_state-\u003eactive_locks) {\n+\t\t\t/* similar to static subprog calls callx is allowed under a lock */\n+\t\t\tif (env-\u003ecur_state-\u003eactive_locks \u0026\u0026 !bpf_is_callx(insn)) {\n \t\t\t\tif ((insn-\u003esrc_reg == BPF_REG_0 \u0026\u0026\n \t\t\t\t insn-\u003eimm != BPF_FUNC_spin_unlock \u0026\u0026\n \t\t\t\t insn-\u003eimm != BPF_FUNC_kptr_xchg) ||\n@@ -18727,6 +19129,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)\n \t\t\tmark_reg_scratched(env, BPF_REG_0);\n \t\t\tif (bpf_in_stack_arg_cnt(\u0026env-\u003esubprog_info[cur_func(env)-\u003esubprogno]))\n \t\t\t\tcur_func(env)-\u003eno_stack_arg_load = true;\n+\t\t\tif (bpf_is_callx(insn))\n+\t\t\t\treturn check_func_callx(env, insn, \u0026env-\u003einsn_idx);\n \t\t\tif (insn-\u003esrc_reg == BPF_PSEUDO_CALL)\n \t\t\t\treturn check_func_call(env, insn, \u0026env-\u003einsn_idx);\n \t\t\tif (insn-\u003esrc_reg == BPF_PSEUDO_KFUNC_CALL)\n@@ -19302,6 +19706,30 @@ static int check_map_prog_compatibility(struct bpf_verifier_env *env,\n \treturn 0;\n }\n \n+/*\n+ * Keep track of whether the map is used by one program only. Such program may\n+ * store the addresses of its functions into the map when it's frozen, see\n+ * resolve_func_ptrs(), since nothing else relies on what the map has. After\n+ * that the map is not available to other programs.\n+ */\n+static int bpf_map_claim(struct bpf_verifier_env *env, struct bpf_map *map)\n+{\n+\tunsigned long me = (unsigned long)env-\u003eprog-\u003eaux, old;\n+\n+\tfor (;;) {\n+\t\told = READ_ONCE(map-\u003euser);\n+\t\tif (old == me || old == BPF_MAP_USER_MANY)\n+\t\t\treturn 0;\n+\t\tif (old \u0026 BPF_MAP_USER_PATCHED) {\n+\t\t\tverbose(env, \"map '%s' has addresses of functions of another program\\n\",\n+\t\t\t\tmap-\u003ename);\n+\t\t\treturn -EBUSY;\n+\t\t}\n+\t\tif (cmpxchg(\u0026map-\u003euser, old, old ? BPF_MAP_USER_MANY : me) == old)\n+\t\t\treturn 0;\n+\t}\n+}\n+\n static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)\n {\n \tint i, err;\n@@ -19339,6 +19767,10 @@ static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)\n \n \tenv-\u003eused_maps[env-\u003eused_map_cnt++] = map;\n \n+\terr = bpf_map_claim(env, map);\n+\tif (err)\n+\t\treturn err;\n+\n \tif (map-\u003emap_type == BPF_MAP_TYPE_INSN_ARRAY) {\n \t\terr = bpf_insn_array_init(map, env-\u003eprog);\n \t\tif (err) {\n@@ -19493,6 +19925,14 @@ static int check_jmp_fields(struct bpf_verifier_env *env, struct bpf_insn *insn)\n \n \tswitch (opcode) {\n \tcase BPF_CALL:\n+\t\tif (bpf_is_callx(insn)) {\n+\t\t\t/* callx dst_reg */\n+\t\t\tif (insn-\u003esrc_reg != BPF_REG_0 || insn-\u003eimm != 0 || insn-\u003eoff != 0) {\n+\t\t\t\tverbose(env, \"BPF_CALL|BPF_X uses reserved fields\\n\");\n+\t\t\t\treturn -EINVAL;\n+\t\t\t}\n+\t\t\treturn 0;\n+\t\t}\n \t\tif (BPF_SRC(insn-\u003ecode) != BPF_K ||\n \t\t (insn-\u003esrc_reg != BPF_PSEUDO_KFUNC_CALL \u0026\u0026 insn-\u003eoff != 0) ||\n \t\t (insn-\u003esrc_reg != BPF_REG_0 \u0026\u0026 insn-\u003esrc_reg != BPF_PSEUDO_CALL \u0026\u0026\n@@ -19747,6 +20187,93 @@ static int check_and_resolve_insns(struct bpf_verifier_env *env)\n \treturn 0;\n }\n \n+static int add_func_ptr(struct bpf_verifier_env *env, struct bpf_map *map, u32 map_off,\n+\t\t\tu32 xlated_off)\n+{\n+\tstruct bpf_func_ptr *ptrs;\n+\n+\t/* grow by doubling, the array is sorted and searched later */\n+\tif (!(env-\u003efunc_ptr_cnt \u0026 (env-\u003efunc_ptr_cnt - 1))) {\n+\t\tptrs = kvrealloc(env-\u003efunc_ptrs,\n+\t\t\t\t array_size(max(2 * env-\u003efunc_ptr_cnt, 16U), sizeof(*ptrs)),\n+\t\t\t\t GFP_KERNEL_ACCOUNT);\n+\t\tif (!ptrs)\n+\t\t\treturn -ENOMEM;\n+\t\tenv-\u003efunc_ptrs = ptrs;\n+\t}\n+\tenv-\u003efunc_ptrs[env-\u003efunc_ptr_cnt++] = (struct bpf_func_ptr){\n+\t\t.map = map,\n+\t\t.map_off = map_off,\n+\t\t.orig_off = xlated_off,\n+\t\t.xlated_off = xlated_off,\n+\t};\n+\treturn 0;\n+}\n+\n+/*\n+ * Compilers put pointers to functions into read-only data: tables of functions,\n+ * structures of operations, vtables, where they are mixed with other data.\n+ * The loader stores such data in a frozen read-only array map and resolves\n+ * a pointer to a static function to the offset in bytes of its first\n+ * instruction in the program: the address of the function in the program.\n+ *\n+ * Find 64-bit values that look like that in the maps of a program that uses\n+ * callx. It's a guess. When the value is not a pointer, the program either\n+ * fails to load, because it does with a pointer what can be done with\n+ * a number only, or it sees the address of a function instead of the number.\n+ * It's known before the control flow graph of the program is built and the main\n+ * verification pass begins which functions may be called via callx.\n+ *\n+ * When the program is jitted the offsets are replaced with the addresses of\n+ * the functions in the map itself, see jit_subprogs(). Hence the program has to\n+ * be the only user of the map, see bpf_map_claim(): nothing else may rely on\n+ * what the map had.\n+ *\n+ * The program reads the addresses of its functions from there like any other\n+ * data, so it has to be allowed to leak pointers.\n+ */\n+static int resolve_func_ptrs(struct bpf_verifier_env *env)\n+{\n+\tint insn_cnt = env-\u003eprog-\u003elen;\n+\tint i, err, subprog;\n+\tstruct bpf_map *map;\n+\tu64 addr, val;\n+\tu32 off;\n+\n+\tif (!env-\u003ehas_callx || !env-\u003eallow_ptr_leaks)\n+\t\treturn 0;\n+\n+\tfor (i = 0; i \u003c env-\u003eused_map_cnt; i++) {\n+\t\tmap = env-\u003eused_maps[i];\n+\t\t/* coincidences in maps that are shared with other programs don't matter */\n+\t\tif (READ_ONCE(map-\u003euser) != (unsigned long)env-\u003eprog-\u003eaux)\n+\t\t\tcontinue;\n+\t\tif (map-\u003emap_type != BPF_MAP_TYPE_ARRAY || map-\u003emax_entries != 1 ||\n+\t\t !bpf_map_is_rdonly(map) || !map-\u003eops-\u003emap_direct_value_addr ||\n+\t\t !IS_ERR_OR_NULL(map-\u003erecord))\n+\t\t\tcontinue;\n+\t\tif (map-\u003eops-\u003emap_direct_value_addr(map, \u0026addr, 0))\n+\t\t\tcontinue;\n+\n+\t\tfor (off = 0; off + sizeof(u64) \u003c= map-\u003evalue_size; off += sizeof(u64)) {\n+\t\t\tval = *(u64 *)(unsigned long)(addr + off);\n+\t\t\tif (!val || val % sizeof(struct bpf_insn) ||\n+\t\t\t val / sizeof(struct bpf_insn) \u003e= insn_cnt)\n+\t\t\t\tcontinue;\n+\t\t\tsubprog = bpf_find_subprog(env, val / sizeof(struct bpf_insn));\n+\t\t\tif (subprog \u003c= 0 || bpf_subprog_is_global(env, subprog))\n+\t\t\t\tcontinue;\n+\t\t\terr = add_func_ptr(env, map, off, val / sizeof(struct bpf_insn));\n+\t\t\tif (err)\n+\t\t\t\treturn err;\n+\t\t}\n+\t}\n+\tif (env-\u003efunc_ptr_cnt)\n+\t\tsort(env-\u003efunc_ptrs, env-\u003efunc_ptr_cnt, sizeof(*env-\u003efunc_ptrs),\n+\t\t cmp_func_ptrs, NULL);\n+\treturn 0;\n+}\n+\n /* drop refcnt of maps used by the rejected program */\n static void release_maps(struct bpf_verifier_env *env)\n {\n@@ -21717,6 +22244,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,\n \tif (ret \u003c 0)\n \t\tgoto skip_full_check;\n \n+\t/* Find pointers to functions in the read-only maps of the program. */\n+\tret = resolve_func_ptrs(env);\n+\tif (ret \u003c 0)\n+\t\tgoto skip_full_check;\n+\n \t/* Build kfunc prototypes after resolving program resources. */\n \tret = add_kfuncs(env);\n \tif (ret \u003c 0)\n@@ -21776,6 +22308,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,\n \tret = do_check_main(env);\n \tret = ret ?: do_check_subprogs(env);\n \n+\t/* reject recursion through the callx edges found by the main pass */\n+\tif (ret == 0 \u0026\u0026 env-\u003ecallx_edges)\n+\t\tret = sort_subprogs_topo(env);\n+\n \tif (ret == 0 \u0026\u0026 bpf_prog_is_offloaded(env-\u003eprog-\u003eaux))\n \t\tret = bpf_prog_offload_finalize(env);\n \n@@ -21918,6 +22454,8 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,\n \tkvfree(env-\u003escc_info);\n \tkvfree(env-\u003esucc);\n \tkvfree(env-\u003egotox_tmp_buf);\n+\tkvfree(env-\u003ecallx_edges);\n+\tkvfree(env-\u003efunc_ptrs);\n \tbpf_diag_free(env);\n \tkvfree(env);\n \treturn ret;\ndiff --git a/tools/lib/bpf/bpf_gen_internal.h b/tools/lib/bpf/bpf_gen_internal.h\nindex 6c5ad6c55e8a6..206adf28793d7 100644\n--- a/tools/lib/bpf/bpf_gen_internal.h\n+++ b/tools/lib/bpf/bpf_gen_internal.h\n@@ -51,9 +51,17 @@ struct bpf_gen {\n \t__u32 nr_ksyms;\n \tint fd_array;\n \tint nr_fd_array;\n+\t/*\n+\t * Maps with pointers to functions, that programs get their own copies\n+\t * of, take slots in fd_array after nr_obj_maps maps of the object.\n+\t */\n+\t__u32 nr_obj_maps;\n+\t__u32 nr_func_ptr_maps;\n+\t__u32 max_func_ptr_maps;\n };\n \n-void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps);\n+void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps,\n+\t\t int max_func_ptr_maps);\n int bpf_gen__finish(struct bpf_gen *gen, int nr_progs, int nr_maps);\n void bpf_gen__free(struct bpf_gen *gen);\n void bpf_gen__load_btf(struct bpf_gen *gen, const void *raw_data, __u32 raw_size);\n@@ -68,6 +76,9 @@ void bpf_gen__prog_load(struct bpf_gen *gen,\n void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *value, __u32 value_size,\n \t\t\t __u64 flags);\n void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx);\n+int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,\n+\t\t\t\t void *value, __u32 value_size, const __u32 *ptr_offs,\n+\t\t\t\t const __u64 *ptr_vals, int ptr_cnt);\n void bpf_gen__record_attach_target(struct bpf_gen *gen, const char *name, enum bpf_attach_type type);\n void bpf_gen__record_extern(struct bpf_gen *gen, const char *name, bool is_weak,\n \t\t\t bool is_typeless, bool is_ld64, int kind, int insn_idx);\ndiff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c\nindex af3a04f161ac1..251392aa8b41f 100644\n--- a/tools/lib/bpf/gen_loader.c\n+++ b/tools/lib/bpf/gen_loader.c\n@@ -112,13 +112,23 @@ static void emit2(struct bpf_gen *gen, struct bpf_insn insn1, struct bpf_insn in\n static int add_data(struct bpf_gen *gen, const void *data, __u32 size);\n static void emit_sys_close_blob(struct bpf_gen *gen, int blob_off);\n \n-void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps)\n+void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps,\n+\t\t int max_func_ptr_maps)\n {\n \tsize_t stack_sz = sizeof(struct loader_stack), nr_progs_sz;\n \tint i;\n \n \tgen-\u003efd_array = add_data(gen, NULL, MAX_FD_ARRAY_SZ * sizeof(int));\n \tgen-\u003elog_level = log_level;\n+\tgen-\u003enr_obj_maps = nr_maps;\n+\tgen-\u003emax_func_ptr_maps = max_func_ptr_maps;\n+\tif (nr_maps + max_func_ptr_maps \u003e MAX_USED_MAPS) {\n+\t\tpr_warn(\"Total maps exceeds %d\\n\", MAX_USED_MAPS);\n+\t\tgen-\u003eerror = -E2BIG;\n+\t\treturn;\n+\t}\n+\t/* their fds are closed like the fds of the maps of the object when loading fails */\n+\tnr_maps += max_func_ptr_maps;\n \t/* save ctx pointer into R6 */\n \temit(gen, BPF_MOV64_REG(BPF_REG_6, BPF_REG_1));\n \n@@ -385,6 +395,9 @@ int bpf_gen__finish(struct bpf_gen *gen, int nr_progs, int nr_maps)\n \t\treturn gen-\u003eerror;\n \t}\n \temit_sys_close_stack(gen, stack_off(btf_fd));\n+\t/* programs hold their maps with pointers to functions, nothing else needs them */\n+\tfor (i = 0; i \u003c gen-\u003enr_func_ptr_maps; i++)\n+\t\temit_sys_close_blob(gen, blob_fd_array_off(gen, gen-\u003enr_obj_maps + i));\n \tfor (i = 0; i \u003c gen-\u003enr_progs; i++)\n \t\tmove_stack2ctx(gen,\n \t\t\t sizeof(struct bpf_loader_ctx) +\n@@ -1127,12 +1140,59 @@ void bpf_gen__prog_load(struct bpf_gen *gen,\n \tgen-\u003enr_progs++;\n }\n \n+/*\n+ * if (map_desc[map_idx].initial_value) {\n+ * if (ctx-\u003eflags \u0026 BPF_SKEL_KERNEL)\n+ * bpf_probe_read_kernel(value, value_size, initial_value);\n+ * else\n+ * bpf_copy_from_user(value, value_size, initial_value);\n+ * nr_more_insns that the caller emits\n+ * }\n+ */\n+static void emit_copy_initial_value(struct bpf_gen *gen, int map_idx, int value,\n+\t\t\t\t __u32 value_size, int nr_more_insns)\n+{\n+\temit(gen, BPF_LDX_MEM(BPF_DW, BPF_REG_3, BPF_REG_6,\n+\t\t\t sizeof(struct bpf_loader_ctx) +\n+\t\t\t sizeof(struct bpf_map_desc) * map_idx +\n+\t\t\t offsetof(struct bpf_map_desc, initial_value)));\n+\temit(gen, BPF_JMP_IMM(BPF_JEQ, BPF_REG_3, 0, 8 + nr_more_insns));\n+\temit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,\n+\t\t\t\t\t 0, 0, 0, value));\n+\temit(gen, BPF_MOV64_IMM(BPF_REG_2, value_size));\n+\temit(gen, BPF_LDX_MEM(BPF_W, BPF_REG_0, BPF_REG_6,\n+\t\t\t offsetof(struct bpf_loader_ctx, flags)));\n+\temit(gen, BPF_JMP_IMM(BPF_JSET, BPF_REG_0, BPF_SKEL_KERNEL, 2));\n+\temit(gen, BPF_EMIT_CALL(BPF_FUNC_copy_from_user));\n+\temit(gen, BPF_JMP_IMM(BPF_JA, 0, 0, 1));\n+\temit(gen, BPF_EMIT_CALL(BPF_FUNC_probe_read_kernel));\n+}\n+\n+/* Update the element of the map whose fd is in the slot map_idx of fd_array */\n+static void emit_map_update_elem(struct bpf_gen *gen, int map_idx, union bpf_attr *attr,\n+\t\t\t\t int attr_size, int key, int value, __u32 value_size)\n+{\n+\tint map_update_attr;\n+\n+\tmap_update_attr = add_data(gen, attr, attr_size);\n+\tpr_debug(\"gen: map_update_elem: idx %d, value: off %d size %u, attr: off %d size %d\\n\",\n+\t\t map_idx, value, value_size, map_update_attr, attr_size);\n+\tmove_blob2blob(gen, attr_field(map_update_attr, map_fd), 4,\n+\t\t blob_fd_array_off(gen, map_idx));\n+\temit_rel_store(gen, attr_field(map_update_attr, key), key);\n+\temit_rel_store(gen, attr_field(map_update_attr, value), value);\n+\t/* emit MAP_UPDATE_ELEM command */\n+\temit_sys_bpf(gen, BPF_MAP_UPDATE_ELEM, map_update_attr, attr_size);\n+\tdebug_ret(gen, \"update_elem idx %d value_size %d\", map_idx, value_size);\n+\temit_check_err(gen);\n+}\n+\n void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,\n \t\t\t __u32 value_size, __u64 flags)\n {\n \tint attr_size = offsetofend(union bpf_attr, flags);\n-\tint map_update_attr, value, key;\n \tunion bpf_attr attr;\n+\tint value, key;\n \tint zero = 0;\n \n \tmemset(\u0026attr, 0, attr_size);\n@@ -1142,47 +1202,16 @@ void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,\n \tkey = add_data(gen, \u0026zero, sizeof(zero));\n \n \t/*\n-\t * if (map_desc[map_idx].initial_value) {\n-\t * if (ctx-\u003eflags \u0026 BPF_SKEL_KERNEL)\n-\t * bpf_probe_read_kernel(value, value_size, initial_value);\n-\t * else\n-\t * bpf_copy_from_user(value, value_size, initial_value);\n-\t * }\n-\t *\n \t * The runtime initial_value comes from the host-supplied loader\n \t * ctx and would overwrite the blob value that the program signature\n \t * covers and the kernel verifies at load time. For a signed loader\n \t * (gen_hash) the attested blob value must be authoritative, so skip\n \t * the override and leave the signed value in place.\n \t */\n-\tif (!OPTS_GET(gen-\u003eopts, gen_hash, false)) {\n-\t\temit(gen, BPF_LDX_MEM(BPF_DW, BPF_REG_3, BPF_REG_6,\n-\t\t\t\t sizeof(struct bpf_loader_ctx) +\n-\t\t\t\t sizeof(struct bpf_map_desc) * map_idx +\n-\t\t\t\t offsetof(struct bpf_map_desc, initial_value)));\n-\t\temit(gen, BPF_JMP_IMM(BPF_JEQ, BPF_REG_3, 0, 8));\n-\t\temit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,\n-\t\t\t\t\t\t 0, 0, 0, value));\n-\t\temit(gen, BPF_MOV64_IMM(BPF_REG_2, value_size));\n-\t\temit(gen, BPF_LDX_MEM(BPF_W, BPF_REG_0, BPF_REG_6,\n-\t\t\t\t offsetof(struct bpf_loader_ctx, flags)));\n-\t\temit(gen, BPF_JMP_IMM(BPF_JSET, BPF_REG_0, BPF_SKEL_KERNEL, 2));\n-\t\temit(gen, BPF_EMIT_CALL(BPF_FUNC_copy_from_user));\n-\t\temit(gen, BPF_JMP_IMM(BPF_JA, 0, 0, 1));\n-\t\temit(gen, BPF_EMIT_CALL(BPF_FUNC_probe_read_kernel));\n-\t}\n+\tif (!OPTS_GET(gen-\u003eopts, gen_hash, false))\n+\t\temit_copy_initial_value(gen, map_idx, value, value_size, 0);\n \n-\tmap_update_attr = add_data(gen, \u0026attr, attr_size);\n-\tpr_debug(\"gen: map_update_elem: idx %d, value: off %d size %u, attr: off %d size %d\\n\",\n-\t\t map_idx, value, value_size, map_update_attr, attr_size);\n-\tmove_blob2blob(gen, attr_field(map_update_attr, map_fd), 4,\n-\t\t blob_fd_array_off(gen, map_idx));\n-\temit_rel_store(gen, attr_field(map_update_attr, key), key);\n-\temit_rel_store(gen, attr_field(map_update_attr, value), value);\n-\t/* emit MAP_UPDATE_ELEM command */\n-\temit_sys_bpf(gen, BPF_MAP_UPDATE_ELEM, map_update_attr, attr_size);\n-\tdebug_ret(gen, \"update_elem idx %d value_size %d\", map_idx, value_size);\n-\temit_check_err(gen);\n+\temit_map_update_elem(gen, map_idx, \u0026attr, attr_size, key, value, value_size);\n }\n \n void bpf_gen__populate_outer_map(struct bpf_gen *gen, int outer_map_idx, int slot,\n@@ -1214,6 +1243,77 @@ void bpf_gen__populate_outer_map(struct bpf_gen *gen, int outer_map_idx, int slo\n \temit_check_err(gen);\n }\n \n+/*\n+ * A copy of the read-only data map obj_map_idx for a program, with the offsets\n+ * of its functions in it, see create_func_ptr_map() in libbpf.c. It's not\n+ * a map of the object: it's not in the loader ctx. Its content is the content\n+ * of obj_map_idx, that the host may supply when the skeleton is loaded, with\n+ * ptr_cnt 64-bit ptr_vals at ptr_offs.\n+ * Return the index of the map in fd_array for instructions to refer to.\n+ */\n+int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,\n+\t\t\t\t void *pvalue, __u32 value_size, const __u32 *ptr_offs,\n+\t\t\t\t const __u64 *ptr_vals, int ptr_cnt)\n+{\n+\tint attr_size = offsetofend(union bpf_attr, map_extra);\n+\tint map_create_attr, map_idx, key, value, zero = 0, i;\n+\tunion bpf_attr attr;\n+\n+\tif (gen-\u003enr_func_ptr_maps == gen-\u003emax_func_ptr_maps) {\n+\t\tgen-\u003eerror = -EDOM; /* internal bug */\n+\t\treturn 0;\n+\t}\n+\tmap_idx = gen-\u003enr_obj_maps + gen-\u003enr_func_ptr_maps++;\n+\n+\tmemset(\u0026attr, 0, attr_size);\n+\tattr.map_type = tgt_endian(BPF_MAP_TYPE_ARRAY);\n+\tattr.key_size = tgt_endian((__u32)sizeof(int));\n+\tattr.value_size = tgt_endian(value_size);\n+\tattr.max_entries = tgt_endian((__u32)1);\n+\tattr.map_flags = tgt_endian((__u32)BPF_F_RDONLY_PROG);\n+\tif (map_name)\n+\t\tlibbpf_strlcpy(attr.map_name, map_name, sizeof(attr.map_name));\n+\n+\tmap_create_attr = add_data(gen, \u0026attr, attr_size);\n+\tpr_debug(\"gen: func_ptr_map_create: %s idx %d value_size %u, attr: off %d size %d\\n\",\n+\t\t map_name, map_idx, value_size, map_create_attr, attr_size);\n+\temit_sys_bpf(gen, BPF_MAP_CREATE, map_create_attr, attr_size);\n+\tdebug_ret(gen, \"func_ptr_map_create %s idx %d value_size %d\", map_name, map_idx,\n+\t\t value_size);\n+\temit_check_err(gen);\n+\t/* remember map_fd in fd_array */\n+\temit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,\n+\t\t\t\t\t 0, 0, 0, blob_fd_array_off(gen, map_idx)));\n+\temit(gen, BPF_STX_MEM(BPF_W, BPF_REG_1, BPF_REG_7, 0));\n+\n+\t/* pvalue has the pointers already */\n+\tvalue = add_data(gen, pvalue, value_size);\n+\tkey = add_data(gen, \u0026zero, sizeof(zero));\n+\n+\t/* see bpf_gen__map_update_elem() */\n+\tif (!OPTS_GET(gen-\u003eopts, gen_hash, false)) {\n+\t\t/* the jump over these instructions has 16-bit offset */\n+\t\tif (ptr_cnt \u003e 10000) {\n+\t\t\tgen-\u003eerror = -E2BIG;\n+\t\t\treturn 0;\n+\t\t}\n+\t\temit_copy_initial_value(gen, obj_map_idx, value, value_size, 3 * ptr_cnt);\n+\t\t/* the content that the host supplied doesn't have them */\n+\t\tfor (i = 0; i \u003c ptr_cnt; i++) {\n+\t\t\temit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,\n+\t\t\t\t\t\t\t 0, 0, 0, value + ptr_offs[i]));\n+\t\t\temit(gen, BPF_ST_MEM(BPF_DW, BPF_REG_1, 0, ptr_vals[i]));\n+\t\t}\n+\t}\n+\n+\tattr_size = offsetofend(union bpf_attr, flags);\n+\tmemset(\u0026attr, 0, attr_size);\n+\temit_map_update_elem(gen, map_idx, \u0026attr, attr_size, key, value, value_size);\n+\n+\tbpf_gen__map_freeze(gen, map_idx);\n+\treturn map_idx;\n+}\n+\n void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx)\n {\n \tint attr_size = offsetofend(union bpf_attr, map_fd);\ndiff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c\nindex cd1ea1bb53cbf..2e11808f7508d 100644\n--- a/tools/lib/bpf/libbpf.c\n+++ b/tools/lib/bpf/libbpf.c\n@@ -544,6 +544,7 @@ struct bpf_struct_ops {\n #define PERCPU_SEC \".percpu\"\n #define BSS_SEC \".bss\"\n #define RODATA_SEC \".rodata\"\n+#define DATA_REL_RO_SEC \".data.rel.ro\"\n #define KCONFIG_SEC \".kconfig\"\n #define KSYMS_SEC \".ksyms\"\n #define STRUCT_OPS_SEC \".struct_ops\"\n@@ -601,6 +602,9 @@ struct bpf_map {\n \tbool autoattach;\n \t__u64 map_extra;\n \tstruct bpf_program *excl_prog;\n+\t/* pointers to functions in the data of an internal map, see obj-\u003efunc_ptrs */\n+\tstruct func_ptr *func_ptrs;\n+\tsize_t func_ptr_cnt;\n };\n \n enum extern_type {\n@@ -780,6 +784,29 @@ struct bpf_object {\n \t} *jumptable_maps;\n \tsize_t jumptable_map_cnt;\n \n+\t/*\n+\t * Pointers to functions found in read-only data sections: tables of\n+\t * functions, structures of operations, vtables. Sorted by section\n+\t * and offset.\n+\t */\n+\tstruct func_ptr {\n+\t\tint sec_idx;\t\t/* ELF section that contains the pointer */\n+\t\tsize_t sec_off;\t\t/* offset of the pointer in the section */\n+\t\tsize_t text_off;\t/* offset of the function in .text section */\n+\t} *func_ptrs;\n+\tsize_t func_ptr_cnt;\n+\n+\t/*\n+\t * Read-only data with pointers to functions is different for every\n+\t * program that uses it, because so are the offsets of the functions.\n+\t */\n+\tstruct {\n+\t\tstruct bpf_program *prog;\n+\t\tint map_idx;\n+\t\tint fd;\n+\t} *func_ptr_maps;\n+\tsize_t func_ptr_map_cnt;\n+\n \tstruct kern_feature_cache *feat_cache;\n \tchar *token_path;\n \tint token_fd;\n@@ -845,6 +872,17 @@ static bool insn_is_pseudo_func(struct bpf_insn *insn)\n \treturn is_ldimm64_insn(insn) \u0026\u0026 insn-\u003esrc_reg == BPF_PSEUDO_FUNC;\n }\n \n+/*\n+ * ld_imm64 that loads the address of read-only data with pointers to functions\n+ * is marked by bpf_object__relocate() before the code is relocated. Compilers\n+ * leave src_reg of other ld_imm64 zero and it's set by\n+ * bpf_object__relocate_data() later.\n+ */\n+static bool insn_is_func_ptrs_addr(struct bpf_insn *insn)\n+{\n+\treturn is_ldimm64_insn(insn) \u0026\u0026 insn-\u003esrc_reg == BPF_PSEUDO_MAP_VALUE;\n+}\n+\n static int\n bpf_object__init_prog(struct bpf_object *obj, struct bpf_program *prog,\n \t\t const char *name, size_t sec_idx, const char *sec_name,\n@@ -4024,6 +4062,17 @@ static int bpf_object__elf_collect(struct bpf_object *obj)\n \t\t\t\terr = bpf_object__add_programs(obj, data, name, idx);\n \t\t\t\tif (err)\n \t\t\t\t\treturn err;\n+\t\t\t} else if (strcmp(name, DATA_REL_RO_SEC) == 0 ||\n+\t\t\t\t str_has_pfx(name, DATA_REL_RO_SEC \".\")) {\n+\t\t\t\t/*\n+\t\t\t\t * Constants with pointers in them, e.g. vtables,\n+\t\t\t\t * that position independent code keeps here to\n+\t\t\t\t * have them relocated. There is nothing that\n+\t\t\t\t * writes to it after that.\n+\t\t\t\t */\n+\t\t\t\tsec_desc-\u003esec_type = SEC_RODATA;\n+\t\t\t\tsec_desc-\u003eshdr = sh;\n+\t\t\t\tsec_desc-\u003edata = data;\n \t\t\t} else if (strcmp(name, DATA_SEC) == 0 ||\n \t\t\t\t str_has_pfx(name, DATA_SEC \".\")) {\n \t\t\t\tsec_desc-\u003esec_type = SEC_DATA;\n@@ -4068,8 +4117,16 @@ static int bpf_object__elf_collect(struct bpf_object *obj)\n \t\t\t targ_sec_idx \u003e= obj-\u003eefile.sec_cnt)\n \t\t\t\treturn -LIBBPF_ERRNO__FORMAT;\n \n-\t\t\t/* Only do relo for section with exec instructions */\n+\t\t\t/*\n+\t\t\t * Only do relo for section with exec instructions,\n+\t\t\t * struct_ops, maps, and read-only data that might\n+\t\t\t * have pointers to functions.\n+\t\t\t */\n \t\t\tif (!section_have_execinstr(obj, targ_sec_idx) \u0026\u0026\n+\t\t\t strcmp(name, \".rel\" RODATA_SEC) \u0026\u0026\n+\t\t\t !str_has_pfx(name, \".rel\" RODATA_SEC \".\") \u0026\u0026\n+\t\t\t strcmp(name, \".rel\" DATA_REL_RO_SEC) \u0026\u0026\n+\t\t\t !str_has_pfx(name, \".rel\" DATA_REL_RO_SEC \".\") \u0026\u0026\n \t\t\t strcmp(name, \".rel\" STRUCT_OPS_SEC) \u0026\u0026\n \t\t\t strcmp(name, \".rel\" STRUCT_OPS_LINK_SEC) \u0026\u0026\n \t\t\t strcmp(name, \".rel?\" STRUCT_OPS_SEC) \u0026\u0026\n@@ -6471,6 +6528,168 @@ static int create_jt_map(struct bpf_object *obj, struct bpf_program *prog, struc\n \treturn err;\n }\n \n+/*\n+ * The kernel recognizes a pointer to a function in a frozen read-only map by\n+ * its value: the offset in bytes of the function in the program. It makes\n+ * callx work for tables of functions, structures of operations and vtables,\n+ * where pointers are mixed with other data. Functions have different offsets\n+ * in different programs, so create a copy of the map for the program.\n+ * The kernel replaces the offsets with the addresses of the functions when it\n+ * loads the program, which has to be the only user of the map.\n+ */\n+static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog, int map_idx)\n+{\n+\tLIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);\n+\tstruct bpf_map *map = \u0026obj-\u003emaps[map_idx];\n+\t__u32 value_size = map-\u003edef.value_size;\n+\tsize_t i, j, cnt, sec_insn_off;\n+\tstruct func_ptr *ptrs;\n+\tint map_fd, err, zero = 0;\n+\t__u64 val;\n+\tvoid *data, *tmp;\n+\n+\tfor (i = 0; i \u003c obj-\u003efunc_ptr_map_cnt; i++)\n+\t\tif (obj-\u003efunc_ptr_maps[i].prog == prog \u0026\u0026\n+\t\t obj-\u003efunc_ptr_maps[i].map_idx == map_idx)\n+\t\t\treturn obj-\u003efunc_ptr_maps[i].fd;\n+\n+\tdata = malloc(value_size);\n+\tif (!data)\n+\t\treturn -ENOMEM;\n+\n+\t/*\n+\t * The content of the map is final, it's frozen already. There is no map\n+\t * when light skeleton is generated, but there is what it's created with.\n+\t */\n+\tif (map-\u003emmaped) {\n+\t\tmemcpy(data, map-\u003emmaped, value_size);\n+\t} else if (obj-\u003egen_loader) {\n+\t\terr = -EINVAL;\n+\t\tgoto err_free;\n+\t} else if (bpf_map_lookup_elem(map-\u003efd, \u0026zero, data)) {\n+\t\terr = -errno;\n+\t\tpr_warn(\"prog '%s': map '%s': failed to read the content: %s\\n\",\n+\t\t\tprog-\u003ename, map-\u003ename, errstr(err));\n+\t\tgoto err_free;\n+\t}\n+\n+\tptrs = map-\u003efunc_ptrs;\n+\tcnt = map-\u003efunc_ptr_cnt;\n+\tfor (i = 0; i \u003c cnt; i++) {\n+\t\tif (ptrs[i].sec_off + sizeof(val) \u003e value_size) {\n+\t\t\terr = -LIBBPF_ERRNO__FORMAT;\n+\t\t\tgoto err_free;\n+\t\t}\n+\t\t/*\n+\t\t * Static functions were appended by bpf_object__append_func_ptrs_code().\n+\t\t * A global function is in the program only if the code refers to it.\n+\t\t */\n+\t\tsec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;\n+\t\tfor (j = 0; j \u003c prog-\u003esubprog_cnt; j++)\n+\t\t\tif (prog-\u003esubprogs[j].sec_insn_off == sec_insn_off)\n+\t\t\t\tbreak;\n+\t\tif (j == prog-\u003esubprog_cnt) {\n+\t\t\tpr_debug(\"prog '%s': map '%s': no function for the pointer at offset %zu, it's NULL\\n\",\n+\t\t\t\t prog-\u003ename, map-\u003ename, ptrs[i].sec_off);\n+\t\t\tval = 0;\n+\t\t} else {\n+\t\t\tval = (__u64)prog-\u003esubprogs[j].sub_insn_off * BPF_INSN_SZ;\n+\t\t}\n+\t\t/* light skeleton can be generated for a target of another endianness */\n+\t\tif (!is_native_endianness(obj))\n+\t\t\tval = bswap_64(val);\n+\t\tmemcpy(data + ptrs[i].sec_off, \u0026val, sizeof(val));\n+\t}\n+\n+\t/*\n+\t * The kernel takes any aligned 64-bit value that is equal to the offset\n+\t * of a function for a pointer. Tell when it's going to get it wrong.\n+\t */\n+\tfor (i = 0, j = 0; i + sizeof(val) \u003c= value_size; i += sizeof(val)) {\n+\t\t__u32 k;\n+\n+\t\twhile (j \u003c cnt \u0026\u0026 ptrs[j].sec_off \u003c i)\n+\t\t\tj++;\n+\t\tif (j \u003c cnt \u0026\u0026 ptrs[j].sec_off == i)\n+\t\t\tcontinue;\n+\t\tmemcpy(\u0026val, data + i, sizeof(val));\n+\t\tif (!is_native_endianness(obj))\n+\t\t\tval = bswap_64(val);\n+\t\tif (!val || val % BPF_INSN_SZ)\n+\t\t\tcontinue;\n+\t\tfor (k = 0; k \u003c prog-\u003esubprog_cnt; k++) {\n+\t\t\tif ((__u64)prog-\u003esubprogs[k].sub_insn_off * BPF_INSN_SZ != val)\n+\t\t\t\tcontinue;\n+\t\t\tpr_warn(\"prog '%s': map '%s': value %llu at offset %zu is the offset of a function, the kernel will treat it as a pointer to it\\n\",\n+\t\t\t\tprog-\u003ename, map-\u003ename, (unsigned long long)val, i);\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+\n+\tif (obj-\u003egen_loader) {\n+\t\t__u32 *ptr_offs = calloc(cnt, sizeof(*ptr_offs));\n+\t\t__u64 *ptr_vals = calloc(cnt, sizeof(*ptr_vals));\n+\n+\t\tif (!ptr_offs || !ptr_vals) {\n+\t\t\tfree(ptr_offs);\n+\t\t\tfree(ptr_vals);\n+\t\t\terr = -ENOMEM;\n+\t\t\tgoto err_free;\n+\t\t}\n+\t\tfor (i = 0; i \u003c cnt; i++) {\n+\t\t\tptr_offs[i] = ptrs[i].sec_off;\n+\t\t\tmemcpy(\u0026ptr_vals[i], data + ptrs[i].sec_off, sizeof(val));\n+\t\t\tif (!is_native_endianness(obj))\n+\t\t\t\tptr_vals[i] = bswap_64(ptr_vals[i]);\n+\t\t}\n+\t\t/* it's an index in fd_array of the loader, not an fd */\n+\t\tmap_fd = bpf_gen__func_ptr_map_create(obj-\u003egen_loader, map-\u003ename, map_idx, data,\n+\t\t\t\t\t\t value_size, ptr_offs, ptr_vals, cnt);\n+\t\tfree(ptr_offs);\n+\t\tfree(ptr_vals);\n+\t\tgoto done;\n+\t}\n+\n+\tmap_fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, map-\u003ename, sizeof(int), value_size, 1, \u0026opts);\n+\tif (map_fd \u003c 0) {\n+\t\terr = map_fd;\n+\t\tgoto err_free;\n+\t}\n+\n+\terr = bpf_map_update_elem(map_fd, \u0026zero, data, 0);\n+\tif (!err)\n+\t\terr = bpf_map_freeze(map_fd);\n+\tif (err) {\n+\t\terr = -errno;\n+\t\tgoto err_close;\n+\t}\n+done:\n+\n+\ttmp = libbpf_reallocarray(obj-\u003efunc_ptr_maps, obj-\u003efunc_ptr_map_cnt + 1,\n+\t\t\t\t sizeof(*obj-\u003efunc_ptr_maps));\n+\tif (!tmp) {\n+\t\terr = -ENOMEM;\n+\t\tgoto err_close;\n+\t}\n+\tobj-\u003efunc_ptr_maps = tmp;\n+\tobj-\u003efunc_ptr_maps[obj-\u003efunc_ptr_map_cnt].prog = prog;\n+\tobj-\u003efunc_ptr_maps[obj-\u003efunc_ptr_map_cnt].map_idx = map_idx;\n+\tobj-\u003efunc_ptr_maps[obj-\u003efunc_ptr_map_cnt].fd = map_fd;\n+\tobj-\u003efunc_ptr_map_cnt++;\n+\n+\tpr_debug(\"prog '%s': created a copy of map '%s' with %zu pointers to functions\\n\",\n+\t\t prog-\u003ename, map-\u003ename, cnt);\n+\tfree(data);\n+\treturn map_fd;\n+\n+err_close:\n+\tif (!obj-\u003egen_loader)\n+\t\tclose(map_fd);\n+err_free:\n+\tfree(data);\n+\treturn err;\n+}\n+\n /* Relocate data references within program code:\n * - map references;\n * - global variable references;\n@@ -6508,7 +6727,20 @@ bpf_object__relocate_data(struct bpf_object *obj, struct bpf_program *prog)\n \t\t\tif (relo-\u003emap_idx == obj-\u003earena_map_idx)\n \t\t\t\tinsn[1].imm += obj-\u003earena_data_off;\n \n-\t\t\tif (obj-\u003egen_loader) {\n+\t\t\tif (map-\u003eautocreate \u0026\u0026 map-\u003efunc_ptr_cnt) {\n+\t\t\t\tint map_fd;\n+\n+\t\t\t\t/* the program gets its own map with pointers to its functions */\n+\t\t\t\tmap_fd = create_func_ptr_map(obj, prog, relo-\u003emap_idx);\n+\t\t\t\tif (map_fd \u003c 0) {\n+\t\t\t\t\tpr_warn(\"prog '%s': relo #%d: can't create a copy of map '%s' with pointers to functions\\n\",\n+\t\t\t\t\t\tprog-\u003ename, i, map-\u003ename);\n+\t\t\t\t\treturn map_fd;\n+\t\t\t\t}\n+\t\t\t\tinsn[0].src_reg = obj-\u003egen_loader ? BPF_PSEUDO_MAP_IDX_VALUE :\n+\t\t\t\t\t\t\t\t BPF_PSEUDO_MAP_VALUE;\n+\t\t\t\tinsn[0].imm = map_fd;\n+\t\t\t} else if (obj-\u003egen_loader) {\n \t\t\t\tinsn[0].src_reg = BPF_PSEUDO_MAP_IDX_VALUE;\n \t\t\t\tinsn[0].imm = relo-\u003emap_idx;\n \t\t\t} else if (map-\u003eautocreate) {\n@@ -6835,6 +7067,51 @@ bpf_object__append_subprog_code(struct bpf_object *obj, struct bpf_program *main\n \treturn 0;\n }\n \n+static int\n+bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,\n+\t\t struct bpf_program *prog);\n+\n+/* Append to the main program all functions that the data of the map points to */\n+static int\n+bpf_object__append_func_ptrs_code(struct bpf_object *obj, struct bpf_program *main_prog,\n+\t\t\t\t const struct bpf_map *map)\n+{\n+\tstruct bpf_program *subprog;\n+\tsize_t i, cnt, sec_insn_off;\n+\tstruct func_ptr *ptrs;\n+\tint err;\n+\n+\tptrs = map-\u003efunc_ptrs;\n+\tcnt = map-\u003efunc_ptr_cnt;\n+\tfor (i = 0; i \u003c cnt; i++) {\n+\t\tsec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;\n+\t\tsubprog = find_prog_by_sec_insn(obj, obj-\u003eefile.text_shndx, sec_insn_off);\n+\t\tif (!subprog || subprog-\u003esec_insn_off != sec_insn_off) {\n+\t\t\tpr_warn(\"prog '%s': map '%s': no function at .text+%zu for the pointer at offset %zu\\n\",\n+\t\t\t\tmain_prog-\u003ename, map-\u003ename, ptrs[i].text_off, ptrs[i].sec_off);\n+\t\t\treturn -LIBBPF_ERRNO__RELOC;\n+\t\t}\n+\n+\t\t/*\n+\t\t * callx can't call global functions. Don't add one to the\n+\t\t * program only because the data points to it.\n+\t\t */\n+\t\tif (subprog-\u003esym_global)\n+\t\t\tcontinue;\n+\n+\t\t/* see the comment in bpf_object__reloc_code() */\n+\t\tif (subprog-\u003esub_insn_off == 0) {\n+\t\t\terr = bpf_object__append_subprog_code(obj, main_prog, subprog);\n+\t\t\tif (err)\n+\t\t\t\treturn err;\n+\t\t\terr = bpf_object__reloc_code(obj, main_prog, subprog);\n+\t\t\tif (err)\n+\t\t\t\treturn err;\n+\t\t}\n+\t}\n+\treturn 0;\n+}\n+\n static int\n bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,\n \t\t struct bpf_program *prog)\n@@ -6851,6 +7128,20 @@ bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,\n \n \tfor (insn_idx = 0; insn_idx \u003c prog-\u003esec_insn_cnt; insn_idx++) {\n \t\tinsn = \u0026main_prog-\u003einsns[prog-\u003esub_insn_off + insn_idx];\n+\t\tif (insn_is_func_ptrs_addr(insn)) {\n+\t\t\t/*\n+\t\t\t * The code that loads the address of the data might\n+\t\t\t * call any function that the data points to.\n+\t\t\t */\n+\t\t\trelo = find_prog_insn_relo(prog, insn_idx);\n+\t\t\tif (relo \u0026\u0026 relo-\u003etype == RELO_DATA) {\n+\t\t\t\terr = bpf_object__append_func_ptrs_code(obj, main_prog,\n+\t\t\t\t\t\t\t\t\t\u0026obj-\u003emaps[relo-\u003emap_idx]);\n+\t\t\t\tif (err)\n+\t\t\t\t\treturn err;\n+\t\t\t}\n+\t\t\tcontinue;\n+\t\t}\n \t\tif (!insn_is_subprog_call(insn) \u0026\u0026 !insn_is_pseudo_func(insn))\n \t\t\tcontinue;\n \n@@ -7541,6 +7832,9 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat\n \t\t\t/* mark the insn, so it's recognized by insn_is_pseudo_func() */\n \t\t\tif (relo-\u003etype == RELO_SUBPROG_ADDR)\n \t\t\t\tinsn[0].src_reg = BPF_PSEUDO_FUNC;\n+\t\t\t/* and by insn_is_func_ptrs_addr() */\n+\t\t\tif (relo-\u003etype == RELO_DATA \u0026\u0026 obj-\u003emaps[relo-\u003emap_idx].func_ptr_cnt)\n+\t\t\t\tinsn[0].src_reg = BPF_PSEUDO_MAP_VALUE;\n \t\t}\n \t}\n \n@@ -7757,6 +8051,98 @@ static int bpf_object__collect_map_relos(struct bpf_object *obj,\n \treturn 0;\n }\n \n+/*\n+ * Collect pointers to functions in a read-only data section. They are\n+ * R_BPF_64_ABS64 relocations against .text section, where the offset of\n+ * a static function in the section is stored in place. Relocations in data\n+ * sections were ignored before pointers to functions were supported. Those\n+ * that are something else, e.g. pointers to data, still are.\n+ */\n+static int bpf_object__collect_rodata_relos(struct bpf_object *obj,\n+\t\t\t\t\t Elf64_Shdr *shdr, Elf_Data *data)\n+{\n+\tsize_t sec_idx = shdr-\u003esh_info, sym_idx;\n+\tint i, nrels = shdr-\u003esh_size / shdr-\u003esh_entsize;\n+\tconst char *relo_sec_name;\n+\tstruct func_ptr *ptrs;\n+\tElf_Data *scn_data;\n+\tElf64_Sym *sym;\n+\tElf64_Rel *rel;\n+\t__u64 addend;\n+\n+\trelo_sec_name = elf_sec_str(obj, shdr-\u003esh_name) ?: \"\u003c?\u003e\";\n+\tscn_data = obj-\u003eefile.secs[sec_idx].data;\n+\tif (!scn_data)\n+\t\treturn -LIBBPF_ERRNO__FORMAT;\n+\n+\tfor (i = 0; i \u003c nrels; i++) {\n+\t\trel = elf_rel_by_idx(data, i);\n+\t\tif (!rel) {\n+\t\t\tpr_warn(\"sec '%s': failed to get relo #%d\\n\", relo_sec_name, i);\n+\t\t\treturn -LIBBPF_ERRNO__FORMAT;\n+\t\t}\n+\n+\t\tsym_idx = ELF64_R_SYM(rel-\u003er_info);\n+\t\tsym = elf_sym_by_idx(obj, sym_idx);\n+\t\tif (!sym) {\n+\t\t\tpr_warn(\"sec '%s': symbol #%zu not found for relo #%d\\n\",\n+\t\t\t\trelo_sec_name, sym_idx, i);\n+\t\t\treturn -LIBBPF_ERRNO__FORMAT;\n+\t\t}\n+\n+\t\tif (ELF64_R_TYPE(rel-\u003er_info) != R_BPF_64_ABS64 ||\n+\t\t !sym_is_subprog(sym, obj-\u003eefile.text_shndx)) {\n+\t\t\tpr_debug(\"sec '%s': relo #%d: not a pointer to a function, skipping...\\n\",\n+\t\t\t\t relo_sec_name, i);\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\t/* the kernel finds aligned pointers only */\n+\t\tif (rel-\u003er_offset % sizeof(__u64) || rel-\u003er_offset \u003e= scn_data-\u003ed_size ||\n+\t\t scn_data-\u003ed_size - rel-\u003er_offset \u003c sizeof(__u64)) {\n+\t\t\tpr_debug(\"sec '%s': relo #%d: unsupported offset 0x%zx, skipping...\\n\",\n+\t\t\t\t relo_sec_name, i, (size_t)rel-\u003er_offset);\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tmemcpy(\u0026addend, scn_data-\u003ed_buf + rel-\u003er_offset, sizeof(addend));\n+\t\tif (!is_native_endianness(obj))\n+\t\t\taddend = bswap_64(addend);\n+\t\tif ((sym-\u003est_value + addend) % BPF_INSN_SZ) {\n+\t\t\tpr_debug(\"sec '%s': relo #%d: bad pointer to a function at offset %zu+%llu, skipping...\\n\",\n+\t\t\t\t relo_sec_name, i, (size_t)sym-\u003est_value,\n+\t\t\t\t (unsigned long long)addend);\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tptrs = libbpf_reallocarray(obj-\u003efunc_ptrs, obj-\u003efunc_ptr_cnt + 1, sizeof(*ptrs));\n+\t\tif (!ptrs)\n+\t\t\treturn -ENOMEM;\n+\t\tobj-\u003efunc_ptrs = ptrs;\n+\n+\t\tptrs[obj-\u003efunc_ptr_cnt].sec_idx = sec_idx;\n+\t\tptrs[obj-\u003efunc_ptr_cnt].sec_off = rel-\u003er_offset;\n+\t\tptrs[obj-\u003efunc_ptr_cnt].text_off = sym-\u003est_value + addend;\n+\t\tobj-\u003efunc_ptr_cnt++;\n+\n+\t\tpr_debug(\"sec '%s': relo #%d: pointer at offset %zu to a function at .text+%zu\\n\",\n+\t\t\t relo_sec_name, i, (size_t)rel-\u003er_offset, (size_t)(sym-\u003est_value + addend));\n+\t}\n+\treturn 0;\n+}\n+\n+static int cmp_func_ptrs(const void *_a, const void *_b)\n+{\n+\tconst struct func_ptr *a = _a;\n+\tconst struct func_ptr *b = _b;\n+\n+\tif (a-\u003esec_idx != b-\u003esec_idx)\n+\t\treturn a-\u003esec_idx \u003c b-\u003esec_idx ? -1 : 1;\n+\tif (a-\u003esec_off != b-\u003esec_off)\n+\t\treturn a-\u003esec_off \u003c b-\u003esec_off ? -1 : 1;\n+\treturn 0;\n+}\n+\n static int bpf_object__collect_relos(struct bpf_object *obj)\n {\n \tint i, err;\n@@ -7779,7 +8165,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)\n \t\t\treturn -LIBBPF_ERRNO__INTERNAL;\n \t\t}\n \n-\t\tif (obj-\u003eefile.secs[idx].sec_type == SEC_ST_OPS)\n+\t\tif (obj-\u003eefile.secs[idx].sec_type == SEC_RODATA)\n+\t\t\terr = bpf_object__collect_rodata_relos(obj, shdr, data);\n+\t\telse if (obj-\u003eefile.secs[idx].sec_type == SEC_ST_OPS)\n \t\t\terr = bpf_object__collect_st_ops_relos(obj, shdr, data);\n \t\telse if (idx == obj-\u003eefile.btf_maps_shndx)\n \t\t\terr = bpf_object__collect_map_relos(obj, shdr, data);\n@@ -7789,6 +8177,25 @@ static int bpf_object__collect_relos(struct bpf_object *obj)\n \t\t\treturn err;\n \t}\n \n+\t/* sort by section, so that pointers in the data of a map are next to each other */\n+\tif (obj-\u003efunc_ptr_cnt)\n+\t\tqsort(obj-\u003efunc_ptrs, obj-\u003efunc_ptr_cnt, sizeof(*obj-\u003efunc_ptrs), cmp_func_ptrs);\n+\n+\tfor (i = 0; i \u003c obj-\u003enr_maps; i++) {\n+\t\tstruct bpf_map *map = \u0026obj-\u003emaps[i];\n+\t\tsize_t j;\n+\n+\t\tif (map-\u003elibbpf_type != LIBBPF_MAP_RODATA)\n+\t\t\tcontinue;\n+\t\tfor (j = 0; j \u003c obj-\u003efunc_ptr_cnt; j++) {\n+\t\t\tif (obj-\u003efunc_ptrs[j].sec_idx != map-\u003esec_idx)\n+\t\t\t\tcontinue;\n+\t\t\tif (!map-\u003efunc_ptr_cnt)\n+\t\t\t\tmap-\u003efunc_ptrs = \u0026obj-\u003efunc_ptrs[j];\n+\t\t\tmap-\u003efunc_ptr_cnt++;\n+\t\t}\n+\t}\n+\n \tbpf_object__sort_relos(obj);\n \treturn 0;\n }\n@@ -9208,8 +9615,17 @@ static int bpf_object_load(struct bpf_object *obj, int extra_log_level, const ch\n \t * permit cross-endian creation of \"light skeleton\".\n \t */\n \tif (obj-\u003egen_loader) {\n+\t\tint nr_func_ptr_maps = 0, nr_progs = 0, i;\n+\n+\t\t/* every program may get a copy of every map with pointers to functions */\n+\t\tfor (i = 0; i \u003c obj-\u003enr_maps; i++)\n+\t\t\tif (obj-\u003emaps[i].autocreate \u0026\u0026 obj-\u003emaps[i].func_ptr_cnt)\n+\t\t\t\tnr_func_ptr_maps++;\n+\t\tfor (i = 0; i \u003c obj-\u003enr_programs; i++)\n+\t\t\tif (obj-\u003eprograms[i].autoload \u0026\u0026 !prog_is_subprog(obj, \u0026obj-\u003eprograms[i]))\n+\t\t\t\tnr_progs++;\n \t\tbpf_gen__init(obj-\u003egen_loader, obj-\u003elog_level | extra_log_level,\n-\t\t\t obj-\u003enr_programs, obj-\u003enr_maps);\n+\t\t\t obj-\u003enr_programs, obj-\u003enr_maps, nr_func_ptr_maps * nr_progs);\n \t} else if (!is_native_endianness(obj)) {\n \t\tpr_warn(\"object '%s': loading non-native endianness is unsupported\\n\", obj-\u003ename);\n \t\treturn libbpf_err(-LIBBPF_ERRNO__ENDIAN);\n@@ -9749,6 +10165,12 @@ void bpf_object__close(struct bpf_object *obj)\n \t\tclose(obj-\u003ejumptable_maps[i].fd);\n \tzfree(\u0026obj-\u003ejumptable_maps);\n \n+\tfor (i = 0; i \u003c obj-\u003efunc_ptr_map_cnt; i++)\n+\t\tif (!obj-\u003egen_loader)\n+\t\t\tclose(obj-\u003efunc_ptr_maps[i].fd);\n+\tzfree(\u0026obj-\u003efunc_ptr_maps);\n+\tzfree(\u0026obj-\u003efunc_ptrs);\n+\n \tif (obj-\u003ebtf_module_allowlist) {\n \t\tfor (i = 0; i \u003c obj-\u003ebtf_module_allowlist_cnt; i++)\n \t\t\tzfree(\u0026obj-\u003ebtf_module_allowlist[i]);\ndiff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c\nindex 78f92c39290af..53f64a1a1f25a 100644\n--- a/tools/lib/bpf/linker.c\n+++ b/tools/lib/bpf/linker.c\n@@ -2274,6 +2274,24 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob\n \t\t\t\t\t\tinsn-\u003eimm += sec-\u003edst_off / sizeof(struct bpf_insn);\n \t\t\t\t\telse\n \t\t\t\t\t\tinsn-\u003eimm += sec-\u003edst_off;\n+\t\t\t\t} else if (sym_type == R_BPF_64_ABS64 \u0026\u0026\n+\t\t\t\t\t (sec-\u003eshdr-\u003esh_flags \u0026 SHF_EXECINSTR)) {\n+\t\t\t\t\t/*\n+\t\t\t\t\t * A pointer to a static function in a data section,\n+\t\t\t\t\t * which is stored in place as an offset of the\n+\t\t\t\t\t * function in its section. Data sections are kept\n+\t\t\t\t\t * in the byte order of the object.\n+\t\t\t\t\t */\n+\t\t\t\t\tvoid *ptr = dst_linked_sec-\u003eraw_data + dst_rel-\u003er_offset;\n+\t\t\t\t\t__u64 off;\n+\n+\t\t\t\t\tmemcpy(\u0026off, ptr, sizeof(off));\n+\t\t\t\t\tif (linker-\u003eswapped_endian)\n+\t\t\t\t\t\toff = bswap_64(off);\n+\t\t\t\t\toff += sec-\u003edst_off;\n+\t\t\t\t\tif (linker-\u003eswapped_endian)\n+\t\t\t\t\t\toff = bswap_64(off);\n+\t\t\t\t\tmemcpy(ptr, \u0026off, sizeof(off));\n \t\t\t\t} else {\n \t\t\t\t\tpr_warn(\"relocation against STT_SECTION in non-exec section is not supported!\\n\");\n \t\t\t\t\treturn -EINVAL;\ndiff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel\nindex 580d1d82c1867..3d92cdca62ed8 100644\n--- a/tools/testing/selftests/bpf/Makefile.skel\n+++ b/tools/testing/selftests/bpf/Makefile.skel\n@@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h\t\t\\\n LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \t\\\n \tcore_kern.c core_kern_overflow.c test_ringbuf.c\t\t\t\\\n \ttest_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \\\n-\ttest_ringbuf_overwrite.c\n+\ttest_ringbuf_overwrite.c callx_rodata.c\n \n LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c\n \ndiff --git a/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c\nnew file mode 100644\nindex 0000000000000..19f96792df5db\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c\n@@ -0,0 +1,192 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/*\n+ * The kernel replaces the offsets of functions in a frozen read-only map with\n+ * their addresses when the program is loaded. The program has to be the only\n+ * user of the map, and no other program can use the map after that.\n+ */\n+#include \u003ctest_progs.h\u003e\n+#include \u003clinux/filter.h\u003e\n+#include \u003cbpf/btf.h\u003e\n+\n+#if defined(__x86_64__) || defined(__aarch64__)\n+\n+#define CALLEE_INSN\t6\n+#define DATA\t\t0x1234\n+\n+/*\n+ * main: r2 = \u0026value; r2 = *(u64 *)(r2 + 8); r1 = 10; callx r2; exit\n+ * add1: r0 = r1; r0 += 1; exit\n+ *\n+ * where value is { DATA, offset of add1 in the program }.\n+ */\n+static const struct bpf_insn callx_insns[] = {\n+\tBPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),\n+\tBPF_LDX_MEM(BPF_DW, BPF_REG_2, BPF_REG_2, 8),\n+\tBPF_MOV64_IMM(BPF_REG_1, 10),\n+\tBPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, BPF_REG_2, 0, 0, 0),\n+\tBPF_EXIT_INSN(),\n+\tBPF_MOV64_REG(BPF_REG_0, BPF_REG_1),\n+\tBPF_ALU64_IMM(BPF_ADD, BPF_REG_0, 1),\n+\tBPF_EXIT_INSN(),\n+};\n+\n+/* reads the data of the map */\n+static const struct bpf_insn reader_insns[] = {\n+\tBPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),\n+\tBPF_LDX_MEM(BPF_DW, BPF_REG_0, BPF_REG_2, 0),\n+\tBPF_EXIT_INSN(),\n+};\n+\n+/* refers to the map and is rejected: r0 is not set */\n+static const struct bpf_insn bad_insns[] = {\n+\tBPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),\n+\tBPF_EXIT_INSN(),\n+};\n+\n+static char log_buf[16 * 1024];\n+static struct bpf_func_info func_info[2];\n+static int btf_fd;\n+\n+static int create_map(void)\n+{\n+\tLIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);\n+\t__u64 value[2] = { DATA, CALLEE_INSN * sizeof(struct bpf_insn) };\n+\tint fd, zero = 0;\n+\n+\tfd = bpf_map_create(BPF_MAP_TYPE_ARRAY, \"callx_rodata\", sizeof(int), sizeof(value), 1,\n+\t\t\t \u0026opts);\n+\tif (!ASSERT_OK_FD(fd, \"map_create\"))\n+\t\treturn -1;\n+\tif (!ASSERT_OK(bpf_map_update_elem(fd, \u0026zero, value, 0), \"map_update\") ||\n+\t !ASSERT_OK(bpf_map_freeze(fd), \"map_freeze\")) {\n+\t\tclose(fd);\n+\t\treturn -1;\n+\t}\n+\treturn fd;\n+}\n+\n+static int load(const struct bpf_insn *prog_insns, int cnt, int map_fd)\n+{\n+\tLIBBPF_OPTS(bpf_prog_load_opts, opts,\n+\t\t.log_buf = log_buf,\n+\t\t.log_size = sizeof(log_buf),\n+\t\t.log_level = 1,\n+\t);\n+\tstruct bpf_insn insns[ARRAY_SIZE(callx_insns)];\n+\n+\tif (prog_insns == callx_insns) {\n+\t\t/* add1() is referred to by the data only and is found through func_info */\n+\t\topts.prog_btf_fd = btf_fd;\n+\t\topts.func_info = func_info;\n+\t\topts.func_info_cnt = 2;\n+\t\topts.func_info_rec_size = sizeof(func_info[0]);\n+\t}\n+\tmemcpy(insns, prog_insns, cnt * sizeof(insns[0]));\n+\tinsns[0].imm = map_fd;\n+\tlog_buf[0] = 0;\n+\treturn bpf_prog_load(BPF_PROG_TYPE_SOCKET_FILTER, \"callx_map\", \"GPL\", insns, cnt, \u0026opts);\n+}\n+\n+#define LOAD(insns, map_fd) load(insns, ARRAY_SIZE(insns), map_fd)\n+\n+static void run(int prog_fd, int expected, const char *name)\n+{\n+\tLIBBPF_OPTS(bpf_test_run_opts, topts);\n+\tchar pkt[64] = {};\n+\n+\ttopts.data_in = pkt;\n+\ttopts.data_size_in = sizeof(pkt);\n+\tif (ASSERT_OK(bpf_prog_test_run_opts(prog_fd, \u0026topts), name))\n+\t\tASSERT_EQ(topts.retval, expected, name);\n+}\n+\n+void test_callx_func_ptr_map(void)\n+{\n+\tint int_id, proto_id, map_fd = -1, prog_fd = -1, reader_fd = -1, fd, zero = 0;\n+\tstruct btf *btf;\n+\t__u64 value[2];\n+\n+\tbtf = btf__new_empty();\n+\tif (!ASSERT_OK_PTR(btf, \"btf_new\"))\n+\t\treturn;\n+\tint_id = btf__add_int(btf, \"int\", 4, BTF_INT_SIGNED);\n+\tproto_id = btf__add_func_proto(btf, int_id);\n+\tfunc_info[0].insn_off = 0;\n+\tfunc_info[0].type_id = btf__add_func(btf, \"main_prog\", BTF_FUNC_GLOBAL, proto_id);\n+\tfunc_info[1].insn_off = CALLEE_INSN;\n+\tfunc_info[1].type_id = btf__add_func(btf, \"add1\", BTF_FUNC_STATIC, proto_id);\n+\tif (!ASSERT_GT(func_info[1].type_id, 0, \"btf_add_func\") ||\n+\t !ASSERT_OK(btf__load_into_kernel(btf), \"btf_load\"))\n+\t\tgoto out;\n+\tbtf_fd = btf__fd(btf);\n+\n+\t/*\n+\t * Another program relies on what the map has: it's verified with\n+\t * the data folded into constants. Pointers to functions are not looked\n+\t * for in such map.\n+\t */\n+\tmap_fd = create_map();\n+\tif (map_fd \u003c 0)\n+\t\tgoto out;\n+\treader_fd = LOAD(reader_insns, map_fd);\n+\tif (!ASSERT_OK_FD(reader_fd, \"load_reader\"))\n+\t\tgoto out;\n+\tfd = LOAD(callx_insns, map_fd);\n+\tif (!ASSERT_LT(fd, 0, \"load_shared\"))\n+\t\tclose(fd);\n+\tASSERT_HAS_SUBSTR(log_buf, \"unreachable insn 6\", \"log_shared\");\n+\trun(reader_fd, DATA, \"run_reader\");\n+\tclose(reader_fd);\n+\treader_fd = -1;\n+\tclose(map_fd);\n+\n+\tmap_fd = create_map();\n+\tif (map_fd \u003c 0)\n+\t\tgoto out;\n+\n+\t/* a program that is rejected is not a user, libbpf loads it again to get the log */\n+\tfd = LOAD(bad_insns, map_fd);\n+\tif (!ASSERT_LT(fd, 0, \"load_bad\"))\n+\t\tclose(fd);\n+\n+\tprog_fd = LOAD(callx_insns, map_fd);\n+\tif (!ASSERT_OK_FD(prog_fd, \"load_callx\")) {\n+\t\tprintf(\"%s\\n\", log_buf);\n+\t\tgoto out;\n+\t}\n+\trun(prog_fd, 11, \"run_callx\");\n+\n+\t/* the offset of the function is gone from the map, the data is intact */\n+\tif (ASSERT_OK(bpf_map_lookup_elem(map_fd, \u0026zero, value), \"map_lookup\")) {\n+\t\tASSERT_EQ(value[0], DATA, \"data\");\n+\t\tASSERT_NEQ(value[1], CALLEE_INSN * sizeof(struct bpf_insn), \"pointer\");\n+\t}\n+\n+\t/* no other program can use the map now, another instance of the same one too */\n+\tfd = LOAD(reader_insns, map_fd);\n+\tif (!ASSERT_EQ(fd, -EBUSY, \"load_reader_after\"))\n+\t\tclose(fd);\n+\tASSERT_HAS_SUBSTR(log_buf, \"has addresses of functions of another program\", \"log_reader\");\n+\tfd = LOAD(callx_insns, map_fd);\n+\tif (!ASSERT_EQ(fd, -EBUSY, \"load_second_instance\"))\n+\t\tclose(fd);\n+\n+\trun(prog_fd, 11, \"run_callx_again\");\n+out:\n+\tif (reader_fd \u003e= 0)\n+\t\tclose(reader_fd);\n+\tif (prog_fd \u003e= 0)\n+\t\tclose(prog_fd);\n+\tif (map_fd \u003e= 0)\n+\t\tclose(map_fd);\n+\tbtf__free(btf);\n+}\n+\n+#else\n+\n+void test_callx_func_ptr_map(void)\n+{\n+\ttest__skip();\n+}\n+\n+#endif\ndiff --git a/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c\nnew file mode 100644\nindex 0000000000000..5c60f352ff446\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c\n@@ -0,0 +1,55 @@\n+// SPDX-License-Identifier: GPL-2.0\n+#include \u003ctest_progs.h\u003e\n+\n+#include \"callx_rodata.lskel.h\"\n+\n+#if defined(__x86_64__) || defined(__aarch64__)\n+\n+static void run(int prog_fd, int expected, const char *name)\n+{\n+\tLIBBPF_OPTS(bpf_test_run_opts, topts);\n+\tchar pkt[64] = {};\n+\n+\ttopts.data_in = pkt;\n+\ttopts.data_size_in = sizeof(pkt);\n+\tif (ASSERT_OK(bpf_prog_test_run_opts(prog_fd, \u0026topts), name))\n+\t\tASSERT_EQ(topts.retval, expected, name);\n+}\n+\n+/*\n+ * Every program gets its own copy of .rodata with the offsets of its functions.\n+ * The copies have what user space puts into .rodata before the load.\n+ */\n+void test_callx_rodata_lskel(void)\n+{\n+\tstruct callx_rodata_lskel *skel;\n+\n+\tskel = callx_rodata_lskel__open();\n+\tif (!ASSERT_OK_PTR(skel, \"open\"))\n+\t\treturn;\n+\n+\tskel-\u003erodata-\u003ebias = 7;\n+\n+\tif (!ASSERT_OK(callx_rodata_lskel__load(skel), \"load\"))\n+\t\tgoto out;\n+\n+\tskel-\u003ebss-\u003eop_idx = 0;\n+\trun(skel-\u003eprogs.select_op.prog_fd, 17, \"add_bias\");\n+\tskel-\u003ebss-\u003eop_idx = 1;\n+\trun(skel-\u003eprogs.select_op.prog_fd, 30, \"mul3\");\n+\tskel-\u003ebss-\u003eop_idx = 2;\n+\trun(skel-\u003eprogs.select_op.prog_fd, -1, \"out_of_range\");\n+\t/* mul3(add_bias(4)) */\n+\trun(skel-\u003eprogs.both_ops.prog_fd, 33, \"both_ops\");\n+out:\n+\tcallx_rodata_lskel__destroy(skel);\n+}\n+\n+#else\n+\n+void test_callx_rodata_lskel(void)\n+{\n+\ttest__skip();\n+}\n+\n+#endif\ndiff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c\nindex 4f1e1c1cd5ab3..3e07b957ee80d 100644\n--- a/tools/testing/selftests/bpf/prog_tests/verifier.c\n+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c\n@@ -27,6 +27,8 @@\n #include \"verifier_btf_ctx_access.skel.h\"\n #include \"verifier_btf_unreliable_prog.skel.h\"\n #include \"verifier_call_large_imm.skel.h\"\n+#include \"verifier_callx.skel.h\"\n+#include \"verifier_callx_rodata.skel.h\"\n #include \"verifier_cfg.skel.h\"\n #include \"verifier_cgroup_inv_retcode.skel.h\"\n #include \"verifier_cgroup_skb.skel.h\"\n@@ -196,6 +198,8 @@ void test_verifier_bswap(void) { RUN(verifier_bswap); }\n void test_verifier_btf_ctx_access(void) { RUN(verifier_btf_ctx_access); }\n void test_verifier_btf_unreliable_prog(void) { RUN(verifier_btf_unreliable_prog); }\n void test_verifier_call_large_imm(void) { RUN(verifier_call_large_imm); }\n+void test_verifier_callx(void) { RUN(verifier_callx); }\n+void test_verifier_callx_rodata(void) { RUN(verifier_callx_rodata); }\n void test_verifier_cfg(void) { RUN(verifier_cfg); }\n void test_verifier_cgroup_inv_retcode(void) { RUN(verifier_cgroup_inv_retcode); }\n void test_verifier_cgroup_skb(void) { RUN(verifier_cgroup_skb); }\ndiff --git a/tools/testing/selftests/bpf/progs/callx_rodata.c b/tools/testing/selftests/bpf/progs/callx_rodata.c\nnew file mode 100644\nindex 0000000000000..7475c87cdaed4\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/progs/callx_rodata.c\n@@ -0,0 +1,43 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* callx through pointers to functions in .rodata, loaded by light skeleton */\n+\n+#include \u003clinux/bpf.h\u003e\n+#include \u003cbpf/bpf_helpers.h\u003e\n+\n+typedef int (*op_fn)(int);\n+\n+/* set by user space before the programs are loaded, it's in .rodata too */\n+const volatile int bias = 1;\n+\n+int op_idx;\n+\n+static __noinline int add_bias(int x)\n+{\n+\treturn x + bias;\n+}\n+\n+static __noinline int mul3(int x)\n+{\n+\treturn x * 3;\n+}\n+\n+static op_fn const ops[] = { add_bias, mul3 };\n+\n+SEC(\"socket\")\n+int select_op(void *ctx)\n+{\n+\tunsigned int i = op_idx;\n+\n+\tif (i \u003e= sizeof(ops) / sizeof(ops[0]))\n+\t\treturn -1;\n+\treturn ops[i](10);\n+}\n+\n+/* functions have other offsets in this program */\n+SEC(\"socket\")\n+int both_ops(void *ctx)\n+{\n+\treturn ops[1](ops[0](4));\n+}\n+\n+char _license[] SEC(\"license\") = \"GPL\";\ndiff --git a/tools/testing/selftests/bpf/progs/verifier_callx.c b/tools/testing/selftests/bpf/progs/verifier_callx.c\nnew file mode 100644\nindex 0000000000000..ec63d2107e6c9\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/progs/verifier_callx.c\n@@ -0,0 +1,984 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Tests for callx: indirect calls of bpf subprogs */\n+\n+#include \u003clinux/bpf.h\u003e\n+#include \u003cbpf/bpf_helpers.h\u003e\n+#include \"bpf_misc.h\"\n+#include \"../../../include/linux/filter.h\"\n+\n+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)\n+\n+#define CALLX_INSN(DST, SRC, OFF, IMM) \\\n+\tBPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, DST, SRC, OFF, IMM)\n+\n+struct {\n+\t__uint(type, BPF_MAP_TYPE_ARRAY);\n+\t__uint(max_entries, 1);\n+\t__type(key, int);\n+\t__type(value, long long);\n+} map_array SEC(\".maps\");\n+\n+struct val_with_lock {\n+\tstruct bpf_spin_lock lock;\n+\tint cnt;\n+};\n+\n+struct {\n+\t__uint(type, BPF_MAP_TYPE_ARRAY);\n+\t__uint(max_entries, 1);\n+\t__type(key, int);\n+\t__type(value, struct val_with_lock);\n+} map_lock SEC(\".maps\");\n+\n+struct {\n+\t__uint(type, BPF_MAP_TYPE_PROG_ARRAY);\n+\t__uint(max_entries, 1);\n+\t__uint(key_size, sizeof(int));\n+\t__uint(value_size, sizeof(int));\n+} map_prog SEC(\".maps\");\n+\n+__naked __noinline __used\n+static unsigned long add1(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = r1;\"\n+\t\t\"r0 += 1;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+__naked __noinline __used\n+static unsigned long add2(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = r1;\"\n+\t\t\"r0 += 2;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* apply(fn, x) { return fn(x); } */\n+__naked __noinline __used\n+static unsigned long apply(void)\n+{\n+\tasm volatile (\n+\t\t\"r3 = r1;\"\n+\t\t\"r1 = r2;\"\n+\t\t\"callx r3;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+SEC(\"socket\")\n+__success __retval(6)\n+__naked void callx_basic(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 5;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+/* callx is printed by the verifier log and xlated dump */\n+SEC(\"socket\")\n+__success __log_level(2)\n+__msg(\"(8d) callx r2\")\n+__xlated(\"callx r2\")\n+__naked void callx_disasm(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 5;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+/* different callees are called by the same callx on different paths */\n+SEC(\"socket\")\n+__success __retval(11)\n+__naked void callx_two_callees(void)\n+{\n+\tasm volatile (\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r6 = r0;\"\n+\t\t\"r6 \u0026= 1;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"if r6 == 0 goto +2;\"\n+\t\t\"r2 = %[add2] ll;\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t/* add1(10) - 0 or add2(10) - 1 */\n+\t\t\"r0 -= r6;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32),\n+\t\t __imm_addr(add1),\n+\t\t __imm_addr(add2)\n+\t\t: __clobber_all);\n+}\n+\n+/* pointer to a function is passed as an argument */\n+SEC(\"socket\")\n+__success __retval(45)\n+__naked void callx_fn_as_arg(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = %[add2] ll;\"\n+\t\t\"r2 = 40;\"\n+\t\t\"call apply;\"\n+\t\t\"r6 = r0;\"\n+\t\t\"r1 = %[add1] ll;\"\n+\t\t\"r2 = 2;\"\n+\t\t\"call apply;\"\n+\t\t\"r0 += r6;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1),\n+\t\t __imm_addr(add2)\n+\t\t: __clobber_all);\n+}\n+\n+__naked __noinline __used\n+static unsigned long get_add1(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = %[add1] ll;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+/* pointer to a function is returned from a subprog and called via r0 */\n+SEC(\"socket\")\n+__success __retval(2)\n+__naked void callx_r0(void)\n+{\n+\tasm volatile (\n+\t\t\"call get_add1;\"\n+\t\t\"r1 = 1;\"\n+\t\t\"callx r0;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* pointer to a function survives spill/fill */\n+SEC(\"socket\")\n+__success __retval(9)\n+__naked void callx_spill_fill(void)\n+{\n+\tasm volatile (\n+\t\t\"r2 = %[add2] ll;\"\n+\t\t\"*(u64 *)(r10 - 8) = r2;\"\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r1 = 7;\"\n+\t\t\"r9 = *(u64 *)(r10 - 8);\"\n+\t\t\"callx r9;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32),\n+\t\t __imm_addr(add2)\n+\t\t: __clobber_all);\n+}\n+\n+__naked __noinline __used\n+static unsigned long clobber_callee_saved(void)\n+{\n+\tasm volatile (\n+\t\t\"r6 = 100;\"\n+\t\t\"r7 = 100;\"\n+\t\t\"r8 = 100;\"\n+\t\t\"r9 = 100;\"\n+\t\t\"r0 = r1;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* r6-r9 are preserved across callx */\n+SEC(\"socket\")\n+__success __retval(11)\n+__naked void callx_callee_saved_regs(void)\n+{\n+\tasm volatile (\n+\t\t\"r6 = 1;\"\n+\t\t\"r7 = 2;\"\n+\t\t\"r8 = 3;\"\n+\t\t\"r9 = 4;\"\n+\t\t\"r1 = 1;\"\n+\t\t\"r2 = %[clobber_callee_saved] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"r0 += r6;\"\n+\t\t\"r0 += r7;\"\n+\t\t\"r0 += r8;\"\n+\t\t\"r0 += r9;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(clobber_callee_saved)\n+\t\t: __clobber_all);\n+}\n+\n+/* r1-r5 are scratched by callx */\n+SEC(\"socket\")\n+__failure __msg(\"R1 !read_ok\")\n+__naked void callx_scratches_args(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 5;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"r0 = r1;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+/* return value of the callee is tracked */\n+SEC(\"socket\")\n+__success __log_level(2)\n+__msg(\"R0=6\")\n+__naked void callx_retval_is_tracked(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 5;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+__naked __noinline __used\n+static unsigned long write42(void)\n+{\n+\tasm volatile (\n+\t\t\"r2 = 42;\"\n+\t\t\"*(u64 *)(r1 + 0) = r2;\"\n+\t\t\"r0 = 0;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* callee writes into the stack of the caller */\n+SEC(\"socket\")\n+__success __retval(42)\n+__naked void callx_callee_writes_caller_stack(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 0;\"\n+\t\t\"*(u64 *)(r10 - 8) = r1;\"\n+\t\t\"r1 = r10;\"\n+\t\t\"r1 += -8;\"\n+\t\t\"r2 = %[write42] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"r0 = *(u64 *)(r10 - 8);\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(write42)\n+\t\t: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"R1 has type scalar, expected func\")\n+__naked void callx_scalar(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 0;\"\n+\t\t\"callx r1;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"R2 !read_ok\")\n+__naked void callx_uninit_reg(void)\n+{\n+\tasm volatile (\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"R10 has type fp, expected func\")\n+__naked void callx_fp(void)\n+{\n+\tasm volatile (\n+\t\t\"callx r10;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"R1 has type map_value, expected func\")\n+__naked void callx_map_value(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = %[map_array] ll;\"\n+\t\t\"r2 = r10;\"\n+\t\t\"r2 += -4;\"\n+\t\t\"r3 = 0;\"\n+\t\t\"*(u32 *)(r2 + 0) = r3;\"\n+\t\t\"call %[bpf_map_lookup_elem];\"\n+\t\t\"if r0 == 0 goto 1f;\"\n+\t\t\"r1 = r0;\"\n+\t\t\"callx r1;\"\n+\t\"1:\"\n+\t\t\"r0 = 0;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_map_lookup_elem),\n+\t\t __imm_addr(map_array)\n+\t\t: __clobber_all);\n+}\n+\n+/* the address of a subprog can't be modified before the call */\n+SEC(\"socket\")\n+__failure __msg(\"dereference of modified func ptr R2 off=8 disallowed\")\n+__naked void callx_modified_ptr(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 5;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"r2 += 8;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"variable func access var_off=\")\n+__naked void callx_variable_ptr(void)\n+{\n+\tasm volatile (\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r0 \u0026= 8;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"r2 += r0;\"\n+\t\t\"r1 = 5;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32),\n+\t\t __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+__noinline __used\n+int global_add3(int x)\n+{\n+\treturn x + 3;\n+}\n+\n+/* only static subprogs can be called via callx */\n+SEC(\"socket\")\n+__failure __msg(\"callback function not static\")\n+__naked void callx_global_func(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 5;\"\n+\t\t\"r2 = %[global_add3] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(global_add3)\n+\t\t: __clobber_all);\n+}\n+\n+#define DEFINE_CALLX_RESERVED_FIELDS_PROG(NAME, SRC_REG, OFF, IMM)\t\t\\\n+\tSEC(\"socket\")\t\t\t\t\t\t\t\t\\\n+\t__failure __msg(\"BPF_CALL|BPF_X uses reserved fields\")\t\t\t\\\n+\t__naked void callx_reserved_field_ ## NAME(void)\t\t\t\\\n+\t{\t\t\t\t\t\t\t\t\t\\\n+\t\tasm volatile (\t\t\t\t\t\t\t\\\n+\t\t\t\"r1 = 5;\"\t\t\t\t\t\t\\\n+\t\t\t\"r2 = %[add1] ll;\"\t\t\t\t\t\\\n+\t\t\t\".8byte %[callx_r2];\"\t\t\t\t\t\\\n+\t\t\t\"exit;\"\t\t\t\t\t\t\t\\\n+\t\t\t:\t\t\t\t\t\t\t\\\n+\t\t\t: __imm_addr(add1),\t\t\t\t\t\\\n+\t\t\t __imm_insn(callx_r2, CALLX_INSN(BPF_REG_2, (SRC_REG), (OFF), (IMM))) \\\n+\t\t\t: __clobber_all);\t\t\t\t\t\\\n+\t}\n+\n+DEFINE_CALLX_RESERVED_FIELDS_PROG(src_reg, BPF_REG_1, 0, 0)\n+DEFINE_CALLX_RESERVED_FIELDS_PROG(off, BPF_REG_0, 1, 0)\n+DEFINE_CALLX_RESERVED_FIELDS_PROG(imm, BPF_REG_0, 0, 1)\n+\n+SEC(\"socket\")\n+__failure __msg(\"unknown opcode 8e\")\n+__naked void callx_jmp32(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 5;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\".8byte %[callx32_r2];\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1),\n+\t\t __imm_insn(callx32_r2,\n+\t\t\t BPF_RAW_INSN(BPF_JMP32 | BPF_CALL | BPF_X, BPF_REG_2, 0, 0, 0))\n+\t\t: __clobber_all);\n+}\n+\n+/* similar to calls of static subprogs callx is allowed under a lock */\n+SEC(\"tc\")\n+__success __retval(3)\n+__naked void callx_under_lock(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 0;\"\n+\t\t\"*(u32 *)(r10 - 4) = r1;\"\n+\t\t\"r2 = r10;\"\n+\t\t\"r2 += -4;\"\n+\t\t\"r1 = %[map_lock] ll;\"\n+\t\t\"call %[bpf_map_lookup_elem];\"\n+\t\t\"if r0 != 0 goto 1f;\"\n+\t\t\"exit;\"\n+\t\"1:\"\n+\t\t\"r6 = r0;\"\n+\t\t\"r1 = r6;\"\n+\t\t\"call %[bpf_spin_lock];\"\n+\t\t\"r1 = 1;\"\n+\t\t\"r2 = %[add2] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"r7 = r0;\"\n+\t\t\"r1 = r6;\"\n+\t\t\"call %[bpf_spin_unlock];\"\n+\t\t\"r0 = r7;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_map_lookup_elem),\n+\t\t __imm(bpf_spin_lock),\n+\t\t __imm(bpf_spin_unlock),\n+\t\t __imm_addr(map_lock),\n+\t\t __imm_addr(add2)\n+\t\t: __clobber_all);\n+}\n+\n+/* helpers are still not allowed under a lock in the callee */\n+__naked __noinline __used\n+static unsigned long call_helper(void)\n+{\n+\tasm volatile (\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32)\n+\t\t: __clobber_all);\n+}\n+\n+SEC(\"tc\")\n+__failure __msg(\"function calls are not allowed while holding a lock\")\n+__naked void callx_helper_under_lock(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 0;\"\n+\t\t\"*(u32 *)(r10 - 4) = r1;\"\n+\t\t\"r2 = r10;\"\n+\t\t\"r2 += -4;\"\n+\t\t\"r1 = %[map_lock] ll;\"\n+\t\t\"call %[bpf_map_lookup_elem];\"\n+\t\t\"if r0 != 0 goto 1f;\"\n+\t\t\"exit;\"\n+\t\"1:\"\n+\t\t\"r6 = r0;\"\n+\t\t\"r1 = r6;\"\n+\t\t\"call %[bpf_spin_lock];\"\n+\t\t\"r2 = %[call_helper] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"r1 = r6;\"\n+\t\t\"call %[bpf_spin_unlock];\"\n+\t\t\"r0 = 0;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_map_lookup_elem),\n+\t\t __imm(bpf_spin_lock),\n+\t\t __imm(bpf_spin_unlock),\n+\t\t __imm_addr(map_lock),\n+\t\t __imm_addr(call_helper)\n+\t\t: __clobber_all);\n+}\n+\n+/* self(fn) { return fn(fn); } */\n+__naked __noinline __used\n+static unsigned long self(void)\n+{\n+\tasm volatile (\n+\t\t\"callx r1;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* unbounded recursion is caught by the main verification pass */\n+SEC(\"socket\")\n+__failure __msg(\"frames is too deep\")\n+__naked void callx_unbounded_recursion(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = %[self] ll;\"\n+\t\t\"call self;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(self)\n+\t\t: __clobber_all);\n+}\n+\n+/* countdown(fn, n) { return n ? fn(fn, n - 1) : 0; } */\n+__naked __noinline __used\n+static unsigned long countdown(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"if r2 == 0 goto 1f;\"\n+\t\t\"r2 += -1;\"\n+\t\t\"callx r1;\"\n+\t\"1:\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/*\n+ * The depth of the recursion is known to the main verification pass,\n+ * but recursive calls are not allowed regardless.\n+ */\n+SEC(\"socket\")\n+__failure __msg(\"recursive call from countdown() to countdown()\")\n+__naked void callx_bounded_recursion(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = %[countdown] ll;\"\n+\t\t\"r2 = 2;\"\n+\t\t\"call countdown;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(countdown)\n+\t\t: __clobber_all);\n+}\n+\n+/* ping(fn1, fn2, n) { return n ? fn2(fn2, fn1, n - 1) : 0; } */\n+__naked __noinline __used\n+static unsigned long ping(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"if r3 == 0 goto 1f;\"\n+\t\t\"r3 += -1;\"\n+\t\t\"r4 = r1;\"\n+\t\t\"r1 = r2;\"\n+\t\t\"r2 = r4;\"\n+\t\t\"callx r1;\"\n+\t\"1:\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+__naked __noinline __used\n+static unsigned long pong(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"if r3 == 0 goto 1f;\"\n+\t\t\"r3 += -1;\"\n+\t\t\"r4 = r1;\"\n+\t\t\"r1 = r2;\"\n+\t\t\"r2 = r4;\"\n+\t\t\"callx r1;\"\n+\t\"1:\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"recursive call from\")\n+__naked void callx_mutual_recursion(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = %[ping] ll;\"\n+\t\t\"r2 = %[pong] ll;\"\n+\t\t\"r3 = 3;\"\n+\t\t\"call ping;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(ping),\n+\t\t __imm_addr(pong)\n+\t\t: __clobber_all);\n+}\n+\n+__naked __noinline __used\n+static unsigned long use_stack_304(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"*(u64 *)(r10 - 304) = r0;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* stack of the callee of callx is accounted */\n+SEC(\"socket\")\n+__failure __msg(\"combined stack size of 2 calls is\")\n+__naked void callx_stack_depth(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"*(u64 *)(r10 - 304) = r0;\"\n+\t\t\"r2 = %[use_stack_304] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(use_stack_304)\n+\t\t: __clobber_all);\n+}\n+\n+/* apply_stack_304(fn) { char buf[304]; return fn(); } */\n+__naked __noinline __used\n+static unsigned long apply_stack_304(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"*(u64 *)(r10 - 304) = r0;\"\n+\t\t\"callx r1;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/*\n+ * The address of use_stack_304() is taken by the main prog that doesn't\n+ * use stack, but it is called from apply_stack_304().\n+ */\n+SEC(\"socket\")\n+__failure __msg(\"combined stack size of 3 calls is\")\n+__naked void callx_stack_depth_nested(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = %[use_stack_304] ll;\"\n+\t\t\"call apply_stack_304;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(use_stack_304)\n+\t\t: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__success __retval(0)\n+__naked void callx_stack_depth_ok(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"*(u64 *)(r10 - 200) = r0;\"\n+\t\t\"r2 = %[use_stack_304] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(use_stack_304)\n+\t\t: __clobber_all);\n+}\n+\n+__naked __noinline __used\n+static unsigned long do_tail_call(void)\n+{\n+\tasm volatile (\n+\t\t\"r2 = %[map_prog] ll;\"\n+\t\t\"r3 = 0;\"\n+\t\t\"call %[bpf_tail_call];\"\n+\t\t\"r0 = 0;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_tail_call),\n+\t\t __imm_addr(map_prog)\n+\t\t: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"tail_calls are not allowed in functions called via callx\")\n+__naked void callx_tail_call_in_callee(void)\n+{\n+\tasm volatile (\n+\t\t\"r2 = %[do_tail_call] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(do_tail_call)\n+\t\t: __clobber_all);\n+}\n+\n+/* tail call in the caller of callx is fine */\n+SEC(\"socket\")\n+__success __retval(3)\n+__naked void callx_tail_call_in_caller(void)\n+{\n+\tasm volatile (\n+\t\t\"r6 = r1;\"\n+\t\t\"r1 = 2;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"r7 = r0;\"\n+\t\t\"r1 = r6;\"\n+\t\t\"r2 = %[map_prog] ll;\"\n+\t\t\"r3 = 0;\"\n+\t\t\"call %[bpf_tail_call];\"\n+\t\t\"r0 = r7;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_tail_call),\n+\t\t __imm_addr(map_prog),\n+\t\t __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+/* read_idx(p) { return ((char *)map_value)[*p]; } */\n+__naked __noinline __used\n+static unsigned long read_idx(void)\n+{\n+\tasm volatile (\n+\t\t\"r6 = *(u64 *)(r1 + 0);\"\n+\t\t\"r1 = 0;\"\n+\t\t\"*(u32 *)(r10 - 4) = r1;\"\n+\t\t\"r2 = r10;\"\n+\t\t\"r2 += -4;\"\n+\t\t\"r1 = %[map_array] ll;\"\n+\t\t\"call %[bpf_map_lookup_elem];\"\n+\t\t\"if r0 == 0 goto 1f;\"\n+\t\t\"r0 += r6;\"\n+\t\t\"r0 = *(u8 *)(r0 + 0);\"\n+\t\"1:\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_map_lookup_elem),\n+\t\t __imm_addr(map_array)\n+\t\t: __clobber_all);\n+}\n+\n+/*\n+ * Stack slots of the caller that might be read by the callee of callx\n+ * have to be considered alive at the checkpoints before callx and\n+ * inside of the callee. Otherwise the state with fp-8 == 1000 is pruned\n+ * and out of bounds access in read_idx() goes unnoticed.\n+ */\n+SEC(\"socket\")\n+__failure __msg(\"invalid access to map value, value_size=8 off=1000 size=1\")\n+__flag(BPF_F_TEST_STATE_FREQ)\n+__naked void callx_callee_reads_caller_stack(void)\n+{\n+\tasm volatile (\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r1 = 1000;\"\n+\t\t\"*(u64 *)(r10 - 8) = r1;\"\n+\t\t\"if r0 == 0 goto 1f;\"\n+\t\t\"r1 = 0;\"\n+\t\t\"*(u64 *)(r10 - 8) = r1;\"\n+\t\"1:\"\n+\t\t\"r1 = r10;\"\n+\t\t\"r1 += -8;\"\n+\t\t\"r2 = %[read_idx] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"r0 = 0;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32),\n+\t\t __imm_addr(read_idx)\n+\t\t: __clobber_all);\n+}\n+\n+/* in bounds access in read_idx() is fine */\n+SEC(\"socket\")\n+__success __retval(0)\n+__flag(BPF_F_TEST_STATE_FREQ)\n+__naked void callx_callee_reads_caller_stack_ok(void)\n+{\n+\tasm volatile (\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r1 = 7;\"\n+\t\t\"*(u64 *)(r10 - 8) = r1;\"\n+\t\t\"if r0 == 0 goto 1f;\"\n+\t\t\"r1 = 0;\"\n+\t\t\"*(u64 *)(r10 - 8) = r1;\"\n+\t\"1:\"\n+\t\t\"r1 = r10;\"\n+\t\t\"r1 += -8;\"\n+\t\t\"r2 = %[read_idx] ll;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32),\n+\t\t __imm_addr(read_idx)\n+\t\t: __clobber_all);\n+}\n+\n+/* same as above, but the pointer to the stack is passed through one more frame */\n+SEC(\"socket\")\n+__failure __msg(\"invalid access to map value, value_size=8 off=1000 size=1\")\n+__flag(BPF_F_TEST_STATE_FREQ)\n+__naked void callx_callee_reads_caller_stack_nested(void)\n+{\n+\tasm volatile (\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r1 = 1000;\"\n+\t\t\"*(u64 *)(r10 - 8) = r1;\"\n+\t\t\"if r0 == 0 goto 1f;\"\n+\t\t\"r1 = 0;\"\n+\t\t\"*(u64 *)(r10 - 8) = r1;\"\n+\t\"1:\"\n+\t\t\"r1 = %[read_idx] ll;\"\n+\t\t\"r2 = r10;\"\n+\t\t\"r2 += -8;\"\n+\t\t\"call apply;\"\n+\t\t\"r0 = 0;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32),\n+\t\t __imm_addr(read_idx)\n+\t\t: __clobber_all);\n+}\n+\n+/* registers that are constant before callx are not constant after it */\n+SEC(\"socket\")\n+__success __retval(1)\n+__naked void callx_clobbers_const_regs(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"r1 = 0;\"\n+\t\t\"r2 = %[add1] ll;\"\n+\t\t\"callx r2;\"\n+\t\t/* dead branch pruning must not assume that r0 is still 0 */\n+\t\t\"if r0 == 0 goto 1f;\"\n+\t\t\"r0 = 1;\"\n+\t\t\"exit;\"\n+\t\"1:\"\n+\t\t\"r0 = 2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(add1)\n+\t\t: __clobber_all);\n+}\n+\n+/* scalar argument passed through callx is tracked precisely */\n+__naked __noinline __used\n+static unsigned long identity(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = r1;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+long long vals[] SEC(\".data.vals\") = {1, 2, 3, 4};\n+\n+SEC(\"socket\")\n+__success __log_level(2)\n+__msg(\"mark_precise: frame0: regs=r0 stack= before 12: (95) exit\")\n+__msg(\"mark_precise: frame1: regs=r0 stack= before 11: (bf) r0 = r1\")\n+__msg(\"mark_precise: frame1: regs=r1 stack= before 4: (8d) callx r2\")\n+__msg(\"mark_precise: frame0: regs=r1 stack= before 3: (bf) r1 = r6\")\n+__msg(\"mark_precise: frame0: regs=r6 stack= before 2: (b7) r6 = 3\")\n+__retval(4)\n+__naked void callx_precision(void)\n+{\n+\tasm volatile (\n+\t\t\"r2 = %[identity] ll;\"\n+\t\t\"r6 = 3;\"\n+\t\t\"r1 = r6;\"\n+\t\t\"callx r2;\"\n+\t\t\"r0 *= 8;\"\n+\t\t\"r1 = %[vals] ll;\"\n+\t\t\"r1 += r0;\"\n+\t\t\"r0 = *(u64 *)(r1 + 0);\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm_addr(identity),\n+\t\t __imm_addr(vals)\n+\t\t: __clobber_all);\n+}\n+\n+/* function pointers in C */\n+\n+typedef int (*op_fn)(int);\n+\n+static __noinline int mul3(int x)\n+{\n+\treturn x * 3;\n+}\n+\n+static __noinline int sub7(int x)\n+{\n+\treturn x - 7;\n+}\n+\n+static __noinline int apply_op(op_fn op, int x)\n+{\n+\treturn op(x);\n+}\n+\n+SEC(\"socket\")\n+__success __retval(36)\n+int callx_c_fn_as_arg(void *ctx)\n+{\n+\t/* (5 * 3) + (28 - 7) */\n+\treturn apply_op(mul3, 5) + apply_op(sub7, 28);\n+}\n+\n+SEC(\"socket\")\n+__success __retval(30)\n+int callx_c_select(void *ctx)\n+{\n+\t__u32 rnd = bpf_get_prandom_u32() \u0026 1;\n+\top_fn op = rnd ? mul3 : sub7;\n+\tint x = rnd ? 10 : 37;\n+\n+\t/* 10 * 3 or 37 - 7 */\n+\treturn op(x);\n+}\n+\n+struct ops {\n+\top_fn first;\n+\top_fn second;\n+\tint bias;\n+};\n+\n+static __noinline int run_ops(const struct ops *ops, int x)\n+{\n+\treturn ops-\u003esecond(ops-\u003efirst(x)) + ops-\u003ebias;\n+}\n+\n+SEC(\"socket\")\n+__success __retval(100)\n+int callx_c_ops_on_stack(void *ctx)\n+{\n+\tstruct ops a, b;\n+\n+\t/* avoid an initializer with function pointers in .rodata */\n+\ta.first = mul3;\n+\ta.second = sub7;\n+\ta.bias = 10;\n+\tb.first = sub7;\n+\tb.second = mul3;\n+\tb.bias = 1;\n+\n+\t/* ((7 * 3) - 7 + 10) + ((32 - 7) * 3 + 1) */\n+\treturn run_ops(\u0026a, 7) + run_ops(\u0026b, 32);\n+}\n+\n+#else\n+\n+SEC(\"socket\")\n+__success\n+int dummy(void *ctx)\n+{\n+\treturn 0;\n+}\n+\n+#endif\n+\n+char _license[] SEC(\"license\") = \"GPL\";\ndiff --git a/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c\nnew file mode 100644\nindex 0000000000000..419a90b1f894e\n--- /dev/null\n+++ b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c\n@@ -0,0 +1,651 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Tests for callx through pointers to functions in read-only data */\n+\n+#include \u003clinux/bpf.h\u003e\n+#include \u003cbpf/bpf_helpers.h\u003e\n+#include \"bpf_misc.h\"\n+#include \"../../../include/linux/filter.h\"\n+\n+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)\n+\n+/*\n+ * Read-only data with pointers to functions, where the compiler puts tables\n+ * of functions, structures of operations and vtables. libbpf resolves\n+ * a pointer to the offset of the function in the program and the kernel\n+ * recognizes it by that value.\n+ */\n+#define DATA(SECTION, NAME, ...)\t\t\t\t\\\n+\t\".pushsection \" SECTION \",@progbits;\"\t\t\t\\\n+\t\".balign 8;\"\t\t\t\t\t\t\\\n+\t#NAME \"_%=:\"\t\t\t\t\t\t\\\n+\t__VA_ARGS__\t\t\t\t\t\t\\\n+\t\".type \" #NAME \"_%=, @object;\"\t\t\t\t\\\n+\t\".size \" #NAME \"_%=, .-\" #NAME \"_%=;\"\t\t\t\\\n+\t\".popsection;\"\n+\n+#define RODATA(NAME, ...) DATA(\".rodata,\\\"a\\\"\", NAME, __VA_ARGS__)\n+\n+#define FUNC_TABLE2(NAME, F0, F1) RODATA(NAME, \".quad \" #F0 \"; .quad \" #F1 \";\")\n+\n+__naked __noinline __used\n+static unsigned long add1(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = r1;\"\n+\t\t\"r0 += 1;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+__naked __noinline __used\n+static unsigned long add2(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = r1;\"\n+\t\t\"r0 += 2;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* the second element of the table is called */\n+SEC(\"socket\")\n+__success __retval(12)\n+__naked void callx_rodata_const_index(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, add1, add2)\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 = *(u64 *)(r2 + 8);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/*\n+ * The program reads the address of a function from the data, so pointers to\n+ * functions in data are recognized only for programs that may leak pointers.\n+ * Otherwise nothing refers to the functions.\n+ */\n+SEC(\"socket\")\n+__success __retval(12)\n+__failure_unpriv __msg_unpriv(\"unreachable insn\")\n+__caps_unpriv(CAP_BPF)\n+__naked void callx_rodata_needs_perfmon(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, add1, add2)\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 = *(u64 *)(r2 + 8);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* the address of an element is used instead of the address of the table */\n+SEC(\"socket\")\n+__success __retval(12)\n+__naked void callx_rodata_elem_addr(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, add1, add2)\n+\t\t\"r2 = tbl_%= + 8 ll;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* pointers to functions are mixed with other data like in a vtable */\n+SEC(\"socket\")\n+__success __retval(23)\n+__log_level(2)\n+__msg(\"r1 = *(u64 *)(r6 +0) ; R1=7\")\n+__msg(\"r2 = *(u64 *)(r6 +8) ; R2=func()\")\n+__naked void callx_rodata_mixed_with_data(void)\n+{\n+\tasm volatile (\n+\t\tRODATA(vt, \".quad 7; .quad add1; .quad 13; .quad add2;\")\n+\t\t\"r6 = vt_%= ll;\"\n+\t\t\"r1 = *(u64 *)(r6 + 0);\"\n+\t\t\"r2 = *(u64 *)(r6 + 8);\"\n+\t\t/* add1(7) */\n+\t\t\"callx r2;\"\n+\t\t\"r7 = r0;\"\n+\t\t\"r1 = *(u64 *)(r6 + 16);\"\n+\t\t\"r2 = *(u64 *)(r6 + 24);\"\n+\t\t/* add2(13) */\n+\t\t\"callx r2;\"\n+\t\t\"r0 += r7;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/*\n+ * Position independent code keeps constants with pointers in .data.rel.ro.\n+ * libbpf treats it as read-only data.\n+ */\n+SEC(\"socket\")\n+__success __retval(23)\n+__naked void callx_data_rel_ro(void)\n+{\n+\tasm volatile (\n+\t\tDATA(\".data.rel.ro,\\\"aw\\\"\", vt, \".quad 7; .quad add1; .quad 13; .quad add2;\")\n+\t\t\"r6 = vt_%= ll;\"\n+\t\t\"r1 = *(u64 *)(r6 + 0);\"\n+\t\t\"r2 = *(u64 *)(r6 + 8);\"\n+\t\t\"callx r2;\"\n+\t\t\"r7 = r0;\"\n+\t\t\"r1 = *(u64 *)(r6 + 16);\"\n+\t\t\"r2 = *(u64 *)(r6 + 24);\"\n+\t\t\"callx r2;\"\n+\t\t\"r0 += r7;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* an array of structures: the index selects one of the functions, not the data */\n+SEC(\"socket\")\n+__success __retval(11)\n+__naked void callx_rodata_array_of_structs(void)\n+{\n+\tasm volatile (\n+\t\tRODATA(arr, \".quad add1; .quad 0x1111; .quad add2; .quad 0x2222;\")\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r6 = r0;\"\n+\t\t\"r6 \u0026= 1;\"\n+\t\t\"r3 = r6;\"\n+\t\t\"r3 \u003c\u003c= 4;\"\n+\t\t\"r2 = arr_%= ll;\"\n+\t\t\"r2 += r3;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t/* add1(10) - 0 or add2(10) - 1 */\n+\t\t\"r0 -= r6;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32)\n+\t\t: __clobber_all);\n+}\n+\n+/* dynamic dispatch: which vtable is used is not known until run time */\n+SEC(\"socket\")\n+__success __retval(2)\n+__naked void callx_rodata_two_vtables(void)\n+{\n+\tasm volatile (\n+\t\tRODATA(vt_a, \".quad 1; .quad add1;\")\n+\t\tRODATA(vt_b, \".quad 0; .quad add2;\")\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r6 = vt_a_%= ll;\"\n+\t\t\"r0 \u0026= 1;\"\n+\t\t\"if r0 == 0 goto +2;\"\n+\t\t\"r6 = vt_b_%= ll;\"\n+\t\t\"r1 = *(u64 *)(r6 + 0);\"\n+\t\t\"r2 = *(u64 *)(r6 + 8);\"\n+\t\t/* add1(1) or add2(0) */\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32)\n+\t\t: __clobber_all);\n+}\n+\n+/* both functions are verified, either of them is called */\n+SEC(\"socket\")\n+__success __retval(11)\n+__naked void callx_rodata_var_index(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, add1, add2)\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r6 = r0;\"\n+\t\t\"r6 \u0026= 1;\"\n+\t\t\"r3 = r6;\"\n+\t\t\"r3 \u003c\u003c= 3;\"\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 += r3;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t/* add1(10) - 0 or add2(10) - 1 */\n+\t\t\"r0 -= r6;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32)\n+\t\t: __clobber_all);\n+}\n+\n+/* a table that has no symbol, the compiler generates such for a switch statement */\n+SEC(\"socket\")\n+__success __retval(11)\n+__naked void callx_rodata_no_symbol(void)\n+{\n+\tasm volatile (\n+\t\t\".pushsection .rodata,\\\"a\\\",@progbits;\"\n+\t\t\".balign 8;\"\n+\t\".Lanon_%=:\"\n+\t\t\".quad add1;\"\n+\t\t\".quad add2;\"\n+\t\t\".popsection;\"\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r6 = r0;\"\n+\t\t\"r6 \u0026= 1;\"\n+\t\t\"r3 = r6;\"\n+\t\t\"r3 \u003c\u003c= 3;\"\n+\t\t\"r2 = .Lanon_%= ll;\"\n+\t\t\"r2 += r3;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t\"r0 -= r6;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32)\n+\t\t: __clobber_all);\n+}\n+\n+/* the first instruction of the function is removed by the verifier */\n+__naked __noinline __used\n+static unsigned long nop_add1(void)\n+{\n+\tasm volatile (\n+\t\t\"goto +0;\"\n+\t\t\"r0 = r1;\"\n+\t\t\"r0 += 1;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* the pointer follows the function when instructions are removed */\n+SEC(\"socket\")\n+__success __retval(11)\n+__naked void callx_rodata_func_starts_with_nop(void)\n+{\n+\tasm volatile (\n+\t\tRODATA(tbl, \".quad nop_add1; .quad 0;\")\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* can't be called with a scalar in r1 */\n+__naked __noinline __used\n+static unsigned long deref_r1(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = *(u64 *)(r1 + 0);\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+__naked __noinline __used\n+static unsigned long ret0(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* every function that might be called is verified */\n+SEC(\"socket\")\n+__failure __msg(\"R1 invalid mem access 'scalar'\")\n+__naked void callx_rodata_all_callees_verified(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, ret0, deref_r1)\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r0 \u0026= 1;\"\n+\t\t\"r0 \u003c\u003c= 3;\"\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 += r0;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r1 = 0;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32)\n+\t\t: __clobber_all);\n+}\n+\n+/* a function that is never called is not verified, it's dead code */\n+SEC(\"socket\")\n+__success __retval(0)\n+__naked void callx_rodata_unused_callee(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, ret0, deref_r1)\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r1 = 0;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* the index may select what is not a pointer to a function */\n+SEC(\"socket\")\n+__failure __msg(\"overlaps with a pointer to a function\")\n+__naked void callx_rodata_index_beyond_table(void)\n+{\n+\tasm volatile (\n+\t\tRODATA(tbl, \".quad add1; .quad add2; .quad 0x1234; .quad 0x5678;\")\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r0 \u0026= 3;\"\n+\t\t\"r0 \u003c\u003c= 3;\"\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 += r0;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32)\n+\t\t: __clobber_all);\n+}\n+\n+/* the address of a function is not known until the program is jitted */\n+SEC(\"socket\")\n+__failure __msg(\"read of 4 bytes at offset\")\n+__msg(\"overlaps with a pointer to a function\")\n+__naked void callx_rodata_partial_read(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, add1, add2)\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r0 = *(u32 *)(r2 + 0);\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"read of 8 bytes at offset\")\n+__msg(\"overlaps with a pointer to a function\")\n+__naked void callx_rodata_misaligned_read(void)\n+{\n+\tasm volatile (\n+\t\tRODATA(tbl, \".quad add1; .quad add2; .quad 0;\")\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r0 = *(u64 *)(r2 + 4);\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* the data next to a pointer is still known to the verifier */\n+SEC(\"socket\")\n+__success __retval(0x1234)\n+__log_level(2)\n+__msg(\"R0=4660\")\n+__naked void callx_rodata_data_is_const(void)\n+{\n+\tasm volatile (\n+\t\tRODATA(tbl, \".quad add1; .quad 0x1234;\")\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r0 = *(u64 *)(r2 + 8);\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* the pointer read from the data can't be modified */\n+SEC(\"socket\")\n+__failure __msg(\"dereference of modified func ptr R2 off=8 disallowed\")\n+__naked void callx_rodata_modified_ptr(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, add1, add2)\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"r2 += 8;\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+__noinline __used\n+int global_add3(int x)\n+{\n+\treturn x + 3;\n+}\n+\n+/* a pointer to a global function is not recognized, it's a number */\n+SEC(\"socket\")\n+__failure __msg(\"R2 has type scalar, expected func\")\n+__naked void callx_rodata_global_func(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, add1, global_add3)\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 = *(u64 *)(r2 + 8);\"\n+\t\t\"r1 = 10;\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+/* selfcall(n) { return n ? tbl[0](n - 1) : 0; }, where tbl[0] == selfcall */\n+__naked __noinline __used\n+static unsigned long selfcall(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, selfcall, ret0)\n+\t\t\"r0 = 0;\"\n+\t\t\"if r1 == 0 goto 1f;\"\n+\t\t\"r1 += -1;\"\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"callx r2;\"\n+\t\"1:\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+SEC(\"socket\")\n+__failure __msg(\"recursive call from selfcall() to selfcall()\")\n+__naked void callx_rodata_recursion(void)\n+{\n+\tasm volatile (\n+\t\t\"r1 = 2;\"\n+\t\t\"call selfcall;\"\n+\t\t\"exit;\"\n+\t\t::: __clobber_all);\n+}\n+\n+__naked __noinline __used\n+static unsigned long use_stack_304(void)\n+{\n+\tasm volatile (\n+\t\t\"r0 = 0;\"\n+\t\t\"*(u64 *)(r10 - 304) = r0;\"\n+\t\t\"exit;\"\n+\t);\n+}\n+\n+/* stack of all possible callees is accounted */\n+SEC(\"socket\")\n+__failure __msg(\"combined stack size of 2 calls is\")\n+__naked void callx_rodata_stack_depth(void)\n+{\n+\tasm volatile (\n+\t\tFUNC_TABLE2(tbl, ret0, use_stack_304)\n+\t\t\"r0 = 0;\"\n+\t\t\"*(u64 *)(r10 - 304) = r0;\"\n+\t\t\"call %[bpf_get_prandom_u32];\"\n+\t\t\"r0 \u0026= 1;\"\n+\t\t\"r0 \u003c\u003c= 3;\"\n+\t\t\"r2 = tbl_%= ll;\"\n+\t\t\"r2 += r0;\"\n+\t\t\"r2 = *(u64 *)(r2 + 0);\"\n+\t\t\"callx r2;\"\n+\t\t\"exit;\"\n+\t\t:\n+\t\t: __imm(bpf_get_prandom_u32)\n+\t\t: __clobber_all);\n+}\n+\n+/* pointers to functions in C */\n+\n+typedef int (*op_fn)(int);\n+\n+#define DEFINE_OP(N) static __noinline int op##N(int x) { return x * (N + 2) + N; }\n+\n+DEFINE_OP(0) DEFINE_OP(1) DEFINE_OP(2) DEFINE_OP(3)\n+DEFINE_OP(4) DEFINE_OP(5) DEFINE_OP(6) DEFINE_OP(7)\n+DEFINE_OP(8) DEFINE_OP(9) DEFINE_OP(10) DEFINE_OP(11)\n+DEFINE_OP(12) DEFINE_OP(13) DEFINE_OP(14) DEFINE_OP(15)\n+\n+static op_fn const ops[] = {\n+\top0, op1, op2, op3, op4, op5, op6, op7,\n+\top8, op9, op10, op11, op12, op13, op14, op15,\n+};\n+\n+int op_idx = 11;\n+\n+SEC(\"socket\")\n+__success __retval(50)\n+int callx_c_table(void *ctx)\n+{\n+\tunsigned int i = op_idx;\n+\n+\tif (i \u003e= sizeof(ops) / sizeof(ops[0]))\n+\t\treturn -1;\n+\t/* op11(3) = 3 * 13 + 11 */\n+\treturn ops[i](3);\n+}\n+\n+SEC(\"socket\")\n+__success __retval(50)\n+int callx_c_table_null_check(void *ctx)\n+{\n+\tunsigned int i = op_idx;\n+\top_fn op;\n+\n+\tif (i \u003e= sizeof(ops) / sizeof(ops[0]))\n+\t\treturn -1;\n+\t/* the address of a function is not known until the program is jitted */\n+\top = ops[i];\n+\tif (!op)\n+\t\treturn -2;\n+\treturn op(3);\n+}\n+\n+/* a structure of operations, where pointers to functions are mixed with data */\n+struct shape_ops {\n+\tint id;\n+\top_fn area;\n+\tlong scale;\n+\top_fn perimeter;\n+};\n+\n+static const struct shape_ops square_ops = { 1, op1, 10, op2 };\n+static const struct shape_ops circle_ops = { 2, op3, 20, op1 };\n+\n+static __noinline int use_shape(const struct shape_ops *ops, int x)\n+{\n+\treturn ops-\u003earea(x) * ops-\u003escale + ops-\u003eperimeter(ops-\u003eid);\n+}\n+\n+SEC(\"socket\")\n+__success __retval(433)\n+int callx_c_ops_mixed_with_data(void *ctx)\n+{\n+\t/*\n+\t * square: op1(2) * 10 + op2(1) = 7 * 10 + 6 = 76\n+\t * circle: op3(3) * 20 + op1(2) = 18 * 20 + 7 = 367\n+\t * minus 10 when op_idx is not what it is set to\n+\t */\n+\treturn use_shape(\u0026square_ops, 2) + use_shape(\u0026circle_ops, 3) - (op_idx == 11 ? 10 : 0);\n+}\n+\n+/* the ops are selected at run time */\n+SEC(\"socket\")\n+__success __retval(367)\n+int callx_c_ops_selected(void *ctx)\n+{\n+\tconst struct shape_ops *ops = op_idx == 11 ? \u0026circle_ops : \u0026square_ops;\n+\n+\treturn use_shape(ops, 3);\n+}\n+\n+/* the compiler might turn the switch into a table that has no symbol */\n+static __noinline int call_by_switch(unsigned int idx, int x)\n+{\n+\top_fn op;\n+\n+\tswitch (idx) {\n+\tcase 0:\n+\t\top = op8;\n+\t\tbreak;\n+\tcase 1:\n+\t\top = op9;\n+\t\tbreak;\n+\tcase 2:\n+\t\top = op10;\n+\t\tbreak;\n+\tcase 3:\n+\t\top = op11;\n+\t\tbreak;\n+\tcase 4:\n+\t\top = op12;\n+\t\tbreak;\n+\tcase 5:\n+\t\top = op13;\n+\t\tbreak;\n+\tcase 6:\n+\t\top = op14;\n+\t\tbreak;\n+\tcase 7:\n+\t\top = op15;\n+\t\tbreak;\n+\tdefault:\n+\t\treturn -1;\n+\t}\n+\treturn op(x);\n+}\n+\n+SEC(\"socket\")\n+__success __retval(50)\n+int callx_c_switch_table(void *ctx)\n+{\n+\t/* op11(3) = 3 * 13 + 11 */\n+\treturn call_by_switch(op_idx - 8, 3);\n+}\n+\n+/*\n+ * Misaligned pointers to functions are ignored, the rest of the data is\n+ * accessible as before.\n+ */\n+static const struct {\n+\tchar tag;\n+\top_fn op;\n+} __attribute__((packed)) packed_ops[] = {\n+\t{ 5, op0 },\n+\t{ 7, op1 },\n+};\n+\n+SEC(\"socket\")\n+__success __retval(7)\n+int callx_c_packed_struct(void *ctx)\n+{\n+\treturn packed_ops[op_idx \u0026 1].tag;\n+}\n+\n+#else\n+\n+SEC(\"socket\")\n+__success\n+int dummy(void *ctx)\n+{\n+\treturn 0;\n+}\n+\n+#endif\n+\n+char _license[] SEC(\"license\") = \"GPL\";\ndiff --git a/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c b/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c\nindex 966f493487874..27fbe54e8795a 100644\n--- a/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c\n+++ b/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c\n@@ -536,4 +536,35 @@ int return_from_void_global(struct __sk_buff *skb)\n \treturn 0;\n }\n \n+int global_calls_loop(int x);\n+\n+static __noinline int static_calls_global(int x)\n+{\n+\treturn global_calls_loop(x);\n+}\n+\n+static __noinline int loop_cb_calls_static(u32 i, void *ctx)\n+{\n+\treturn static_calls_global(i);\n+}\n+\n+__noinline int global_calls_loop(int x)\n+{\n+\tbpf_loop(1, loop_cb_calls_static, NULL, 0);\n+\treturn 0;\n+}\n+\n+/*\n+ * loop_cb_calls_static() -\u003e static_calls_global() -\u003e global_calls_loop() -\u003e\n+ * bpf_loop() -\u003e loop_cb_calls_static() is an unbounded recursion that the\n+ * main verification pass can't see, because it doesn't follow calls of global\n+ * functions. None of the functions use stack.\n+ */\n+SEC(\"?raw_tp\")\n+__failure __msg(\"recursive call from global_calls_loop() to loop_cb_calls_static()\")\n+int recursion_via_global_func_and_callback(const void *ctx)\n+{\n+\treturn loop_cb_calls_static(0, NULL);\n+}\n+\n char _license[] SEC(\"license\") = \"GPL\";\ndiff --git a/tools/testing/selftests/bpf/verifier/basic_call.c b/tools/testing/selftests/bpf/verifier/basic_call.c\nindex a8c6ab4c1622f..0f93c4551f235 100644\n--- a/tools/testing/selftests/bpf/verifier/basic_call.c\n+++ b/tools/testing/selftests/bpf/verifier/basic_call.c\n@@ -4,7 +4,7 @@\n \tBPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, 0, 0, 0, 0),\n \tBPF_EXIT_INSN(),\n \t},\n-\t.errstr = \"unknown opcode 8d\",\n+\t.errstr = \"R0 !read_ok\",\n \t.result = REJECT,\n },\n {\n"
} |
| 2/1 | 2026/09/24 01:45 | agent | patch-evaluator |
1mModel:core Args: null Results: {
"EnableConfigs": [
"BPF_SYSCALL",
"BPF_JIT"
],
"FocusSymbols": [
"check_func_callx",
"resolve_func_ptrs",
"check_func_ptr_read",
"bpf_map_claim",
"bpf_map_range_func_ptrs",
"bpf_adjust_func_ptrs",
"bpf_adjust_func_ptrs_after_remove",
"emit_indirect_call",
"bpf_jit_supports_callx",
"jit_subprogs"
],
"Reasoning": "The patch series implements support for BPF indirect calls ('callx' / BPF_CALL | BPF_X) of static subprograms and function pointer resolution in frozen read-only maps. This introduces substantial modifications across the BPF verifier (CFG traversal, state pruning/liveness, stack depth and recursion checks, func pointer bounds checks) and the x86 BPF JIT compiler (indirect call emission and retpoline handling), along with several new runtime defensive assertions (verifier_bug_if). These changes modify reachable core subsystem code easily exercised via the bpf() syscall.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 47a35dcdf09a938b11851b9703bb51549ad4f665
Author: syz-cluster <triage@syzkaller.com>
Date: Thu Sep 24 01:45:15 2026 +0000
syz-cluster: applied patch under review
diff --git a/Documentation/bpf/clang-notes.rst b/Documentation/bpf/clang-notes.rst
index 2c872a1ee08e9..3ccc7b09d19e6 100644
--- a/Documentation/bpf/clang-notes.rst
+++ b/Documentation/bpf/clang-notes.rst
@@ -23,8 +23,11 @@ For CPU versions prior to 3, Clang v7.0 and later can enable ``BPF_ALU`` support
Jump instructions
=================
-If ``-O0`` is used, Clang will generate the ``BPF_CALL | BPF_X | BPF_JMP`` (0x8d)
-instruction, which is not supported by the Linux kernel verifier.
+Clang generates the ``BPF_CALL | BPF_X | BPF_JMP`` (0x8d) instruction for calls
+through a function pointer. The Linux kernel verifier accepts it only when it
+can prove that the register holds the address of a static BPF function, see
+Documentation/bpf/linux-notes.rst. If ``-O0`` is used, Clang will generate this
+instruction for helper calls as well, which is not supported.
Atomic operations
=================
diff --git a/Documentation/bpf/linux-notes.rst b/Documentation/bpf/linux-notes.rst
index 00d2693de025a..6c036b54a29f2 100644
--- a/Documentation/bpf/linux-notes.rst
+++ b/Documentation/bpf/linux-notes.rst
@@ -15,10 +15,33 @@ Byte swap instructions
Jump instructions
=================
-``BPF_CALL | BPF_X | BPF_JMP`` (0x8d), where the helper function
-integer would be read from a specified register, is not currently supported
-by the verifier. Any programs with this instruction will fail to load
-until such support is added.
+``BPF_CALL | BPF_X | BPF_JMP`` (0x8d), ``callx dst``, performs an indirect
+call of a BPF function whose address is held in the ``dst`` register. The
+``src``, ``offset`` and ``imm`` fields are reserved and must be zero.
+
+The address of a BPF function gets into a register in one of two ways:
+
+* it is loaded by a 64-bit immediate instruction with ``src`` =
+ ``BPF_PSEUDO_FUNC``;
+* it is read, with a 64-bit load, from a frozen read-only array map, that no
+ other program uses, that holds its read-only data: tables of functions,
+ structures of operations, vtables, where pointers to functions may be mixed
+ with other data. In the map a pointer to a function is the offset in bytes
+ of its first instruction in the program, and that is how the verifier
+ recognizes it. It is replaced with the address of the function when
+ the program is loaded. The program reads it from there, which requires
+ ``CAP_PERFMON``.
+
+In both cases only static functions can be referenced. Therefore all functions
+that can be called indirectly are known to the verifier before it starts to
+analyze the program, and ``callx`` is verified as a direct call of every
+function that ``dst`` may point to at that instruction. The same rules apply:
+the calls can not be recursive, and the depth of the call chain and its
+combined stack size are limited.
+
+Calling helper or kernel functions through a register, indirect calls of global
+functions, and tail calls in functions that are called via ``callx`` are not
+supported. ``callx`` requires the BPF JIT.
Maps
====
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 6c04fee468766..9544b2f483e59 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -1789,6 +1789,17 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
emit(A64_MOV(1, r0, A64_R(0)), ctx);
break;
}
+ /* indirect call of a bpf subprog, dst holds its address */
+ case BPF_JMP | BPF_CALL | BPF_X:
+ /*
+ * It's the same as a direct call of a subprog that is out of
+ * range of BL: the subprog starts with BTI JC, the arguments
+ * are in place, and the registers that hold the tail call
+ * counter and the private stack are callee saved.
+ */
+ emit(A64_BLR(dst), ctx);
+ emit(A64_MOV(1, bpf2a64[BPF_REG_0], A64_R(0)), ctx);
+ break;
/* tail call */
case BPF_JMP | BPF_TAIL_CALL:
if (emit_bpf_tail_call(ctx))
@@ -2485,6 +2496,11 @@ bool bpf_jit_supports_subprog_tailcalls(void)
return true;
}
+bool bpf_jit_supports_callx(void)
+{
+ return true;
+}
+
static void invoke_bpf_prog(struct jit_ctx *ctx, struct bpf_tramp_node *node,
int bargs_off, int retval_off, int run_ctx_off,
bool save_ret)
diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index d4a980140b48d..9fbef7504e51a 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -749,6 +749,46 @@ static void emit_indirect_jump(u8 **pprog, int bpf_reg, u8 *ip)
*pprog = prog;
}
+static void __emit_indirect_call(u8 **pprog, int reg, bool ereg)
+{
+ u8 *prog = *pprog;
+
+ if (ereg)
+ EMIT1(0x41);
+
+ EMIT2(0xFF, 0xD0 + reg);
+
+ *pprog = prog;
+}
+
+/* call *bpf_reg */
+static int emit_indirect_call(u8 **pprog, int bpf_reg, u8 *ip)
+{
+ u8 *prog = *pprog;
+ int reg = reg2hex[bpf_reg];
+ bool ereg = is_ereg(bpf_reg);
+ int err = 0;
+
+ if (cpu_feature_enabled(X86_FEATURE_INDIRECT_THUNK_ITS)) {
+ OPTIMIZER_HIDE_VAR(reg);
+ err = emit_call(&prog, its_static_thunk(reg + 8*ereg), ip);
+ } else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE_LFENCE)) {
+ EMIT_LFENCE();
+ __emit_indirect_call(&prog, reg, ereg);
+ } else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE)) {
+ OPTIMIZER_HIDE_VAR(reg);
+ if (cpu_feature_enabled(X86_FEATURE_CALL_DEPTH))
+ err = emit_call(&prog, &__x86_indirect_call_thunk_array[reg + 8*ereg], ip);
+ else
+ err = emit_call(&prog, &__x86_indirect_thunk_array[reg + 8*ereg], ip);
+ } else {
+ __emit_indirect_call(&prog, reg, ereg);
+ }
+
+ *pprog = prog;
+ return err;
+}
+
static void emit_return(u8 **pprog, u8 *ip)
{
u8 *prog = *pprog;
@@ -2941,6 +2981,24 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
break;
}
+ /* callx: call of a bpf subprog whose address is in dst_reg */
+ case BPF_JMP | BPF_CALL | BPF_X:
+ /*
+ * The verifier makes sure that callees of callx are
+ * not tail call reachable, hence unlike a direct call
+ * of a subprog there is no need to pass
+ * tail_call_cnt_ptr in rax.
+ */
+ if (priv_frame_ptr) {
+ push_r9(&prog);
+ ip += 2;
+ }
+ if (emit_indirect_call(&prog, insn->dst_reg, ip))
+ return -EINVAL;
+ if (priv_frame_ptr)
+ pop_r9(&prog);
+ break;
+
case BPF_JMP | BPF_TAIL_CALL:
if (imm32)
emit_bpf_tail_call_direct(bpf_prog,
@@ -4414,6 +4472,16 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
return prog;
}
+bool bpf_jit_supports_callx(void)
+{
+ /*
+ * FineIBT poisons ENDBR at the entry of a JITed function and expects
+ * indirect callers to go through the CFI preamble instead.
+ * callx doesn't do that yet.
+ */
+ return cfi_mode != CFI_FINEIBT;
+}
+
bool bpf_jit_supports_kfunc_call(void)
{
return true;
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index fd22db8bc6c50..7747c5fc290fc 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -342,8 +342,18 @@ struct bpf_map {
s64 __percpu *elem_count;
u64 cookie; /* write-once */
char *excl_prog_sha;
+ /*
+ * Which programs use the map, see bpf_map_claim(): 0 - none so far,
+ * aux of the program - only that one, the same with BPF_MAP_USER_PATCHED
+ * set - only that one and it stored the addresses of its functions into
+ * the map, BPF_MAP_USER_MANY - more than one.
+ */
+ unsigned long user;
};
+#define BPF_MAP_USER_MANY 1UL
+#define BPF_MAP_USER_PATCHED 1UL
+
static inline const char *btf_field_type_name(enum btf_field_type type)
{
switch (type) {
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 92f528c456052..38a4ba50669a2 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -786,6 +786,22 @@ int bpf_log_attr_finalize(struct bpf_log_attr *attr, struct bpf_verifier_log *lo
#define BPF_MAX_SUBPROGS 256
+/*
+ * A pointer to a static subprog in the value of a frozen read-only array map:
+ * a 64-bit value that is the offset in bytes of the first instruction of
+ * the subprog in the program.
+ */
+struct bpf_func_ptr {
+ struct bpf_map *map;
+ u32 map_off; /* offset of the pointer in the value of the map */
+ u32 orig_off; /* what the map has: the first instruction of the subprog */
+ u32 xlated_off; /* the same after instructions were patched and removed */
+ bool used; /* the program reads the pointer */
+};
+
+/* the subprog that a bpf_func_ptr pointed to was removed as dead code */
+#define BPF_FUNC_PTR_DELETED ((u32)-1)
+
struct bpf_subprog_arg_info {
enum bpf_arg_type arg_type;
union {
@@ -965,6 +981,20 @@ struct bpf_verifier_env {
struct bpf_subprog_info subprog_info[BPF_MAX_SUBPROGS + 2]; /* max + 2 for the fake and exception subprogs */
/* subprog indices sorted in topological order: leaves first, callers last */
int subprog_topo_order[BPF_MAX_SUBPROGS + 2];
+ /*
+ * Pointers to static subprogs found in frozen read-only maps of the
+ * program, see resolve_func_ptrs(). Sorted by map and map_off.
+ */
+ struct bpf_func_ptr *func_ptrs;
+ u32 func_ptr_cnt;
+ bool has_callx;
+ /*
+ * Call graph edges created by callx instructions. A bitmap of
+ * subprog_cnt * subprog_cnt bits, where bit (caller * subprog_cnt + callee)
+ * is set when the main verification pass sees 'caller' calling 'callee'
+ * via callx. Allocated when the first such edge is recorded.
+ */
+ unsigned long *callx_edges;
union {
struct bpf_idmap idmap_scratch;
struct bpf_idset idset_scratch;
@@ -1089,6 +1119,12 @@ static inline bool bpf_pseudo_kfunc_call(const struct bpf_insn *insn)
insn->src_reg == BPF_PSEUDO_KFUNC_CALL;
}
+/* callx: indirect call of a bpf subprog whose address is in insn->dst_reg */
+static inline bool bpf_is_callx(const struct bpf_insn *insn)
+{
+ return insn->code == (BPF_JMP | BPF_CALL | BPF_X);
+}
+
__printf(2, 0) void bpf_verifier_vlog(struct bpf_verifier_log *log,
const char *fmt, va_list args);
__printf(2, 3) void bpf_verifier_log_write(struct bpf_verifier_env *env,
@@ -1311,6 +1347,13 @@ static inline bool bt_is_frame_slot_set(struct backtrack_state *bt, u32 frame, u
}
bool bpf_map_is_rdonly(const struct bpf_map *map);
+struct bpf_func_ptr *bpf_map_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map, u32 *cnt);
+struct bpf_func_ptr *bpf_map_range_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map,
+ u64 off, u64 size, u32 *cnt);
+void bpf_adjust_func_ptrs(struct bpf_verifier_env *env, u32 off, u32 len);
+void bpf_adjust_func_ptrs_after_remove(struct bpf_verifier_env *env, u32 off, u32 len);
int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,
bool is_ldsx);
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 422284b4fa96f..4f0662e428970 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1237,6 +1237,7 @@ bool bpf_jit_inlines_helper_call(s32 imm);
bool bpf_jit_supports_subprog_tailcalls(void);
bool bpf_jit_supports_percpu_insn(void);
bool bpf_jit_supports_kfunc_call(void);
+bool bpf_jit_supports_callx(void);
bool bpf_jit_supports_kfunc_ret_reg_pair(void);
bool bpf_jit_supports_stack_args(void);
bool bpf_jit_supports_arena_args(void);
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 507a366dffa47..4da99dec08184 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -406,15 +406,18 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
if (class == BPF_STX)
bt_set_reg(bt, sreg);
} else if (class == BPF_JMP || class == BPF_JMP32) {
- if (bpf_pseudo_call(insn)) {
- int subprog_insn_idx, subprog;
+ if (bpf_pseudo_call(insn) || bpf_is_callx(insn)) {
+ int subprog_insn_idx, subprog = -1;
- subprog_insn_idx = idx + insn->imm + 1;
- subprog = bpf_find_subprog(env, subprog_insn_idx);
- if (subprog < 0)
- return -EFAULT;
+ if (bpf_pseudo_call(insn)) {
+ subprog_insn_idx = idx + insn->imm + 1;
+ subprog = bpf_find_subprog(env, subprog_insn_idx);
+ if (subprog < 0)
+ return -EFAULT;
+ }
- if (bpf_subprog_is_global(env, subprog)) {
+ /* callx calls static subprogs only */
+ if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
/* check that jump history doesn't have any
* extra instructions from subprog; the next
* instruction after call to global subprog
@@ -536,7 +539,8 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
* never do that.
*/
from_subprog_call = subseq_idx - 1 >= 0 &&
- bpf_pseudo_call(&env->prog->insnsi[subseq_idx - 1]);
+ (bpf_pseudo_call(&env->prog->insnsi[subseq_idx - 1]) ||
+ bpf_is_callx(&env->prog->insnsi[subseq_idx - 1]));
r0_precise = from_subprog_call && bt_is_reg_set(bt, BPF_REG_0);
r2_precise = from_subprog_call && bt_is_reg_set(bt, BPF_REG_2);
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 842c7d1eabccc..a068c191409a4 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -421,6 +421,75 @@ static int visit_gotox_insn(int t, struct bpf_verifier_env *env)
return keep_exploring ? KEEP_EXPLORING : DONE_EXPLORING;
}
+/*
+ * Return pointers to functions in the read-only map that ld_imm64 instruction
+ * 't' loads the address of, or of its value, if there are any.
+ */
+static struct bpf_func_ptr *insn_func_ptrs(struct bpf_verifier_env *env, int t, u32 *cnt)
+{
+ struct bpf_insn *insn = &env->prog->insnsi[t];
+
+ *cnt = 0;
+ if (!env->func_ptr_cnt || !bpf_is_ldimm64(insn))
+ return NULL;
+ if (insn->src_reg != BPF_PSEUDO_MAP_VALUE &&
+ insn->src_reg != BPF_PSEUDO_MAP_IDX_VALUE &&
+ insn->src_reg != BPF_PSEUDO_MAP_FD &&
+ insn->src_reg != BPF_PSEUDO_MAP_IDX)
+ return NULL;
+
+ return bpf_map_func_ptrs(env, env->used_maps[env->insn_aux_data[t].map_index], cnt);
+}
+
+/*
+ * ld_imm64 that loads the address of a map that has pointers to functions
+ * is similar to ld_imm64 with BPF_PSEUDO_FUNC that loads the address of one
+ * function: any of them may be read from the map and called via callx later.
+ * Treat it as a call of all of them.
+ */
+static int visit_func_ptrs_insn(int t, struct bpf_verifier_env *env,
+ struct bpf_func_ptr *ptrs, u32 cnt)
+{
+ int *insn_stack = env->cfg.insn_stack;
+ int *insn_state = env->cfg.insn_state;
+ bool keep_exploring = false;
+ int ret, w;
+ u32 i;
+
+ ret = push_insn(t, t + 2, FALLTHROUGH, env);
+ if (ret)
+ return ret;
+
+ mark_prune_point(env, t);
+ for (i = 0; i < cnt; i++) {
+ w = ptrs[i].xlated_off;
+
+ /*
+ * This function is called until all functions are explored,
+ * so the effects are complete in the end.
+ */
+ merge_callee_effects(env, t, w);
+
+ /* the same marks as push_insn() leaves on a branch target */
+ mark_prune_point(env, w);
+ mark_jmp_point(env, w);
+ mark_jump_target(env, w);
+
+ /* EXPLORED || DISCOVERED */
+ if (insn_state[w])
+ continue;
+
+ if (env->cfg.cur_stack >= env->prog->len)
+ return -E2BIG;
+
+ insn_stack[env->cfg.cur_stack++] = w;
+ insn_state[w] |= DISCOVERED;
+ keep_exploring = true;
+ }
+
+ return keep_exploring ? KEEP_EXPLORING : DONE_EXPLORING;
+}
+
/*
* Instructions that can abnormally return from a subprog (tail_call
* upon success, ld_{abs,ind} upon load failure) have a hidden exit
@@ -453,11 +522,17 @@ static int visit_abnormal_return_insn(struct bpf_verifier_env *env, int t)
static int visit_insn(int t, struct bpf_verifier_env *env)
{
struct bpf_insn *insns = env->prog->insnsi, *insn = &insns[t];
+ struct bpf_func_ptr *ptrs;
int ret, off, insn_sz;
+ u32 cnt;
if (bpf_pseudo_func(insn))
return visit_func_call_insn(t, insns, env, true);
+ ptrs = insn_func_ptrs(env, t, &cnt);
+ if (ptrs)
+ return visit_func_ptrs_insn(t, env, ptrs, cnt);
+
/* All non-branch instructions have a single fall-through edge. */
if (BPF_CLASS(insn->code) != BPF_JMP &&
BPF_CLASS(insn->code) != BPF_JMP32) {
diff --git a/kernel/bpf/const_fold.c b/kernel/bpf/const_fold.c
index 7f1b30059cc87..fea639f62b3f2 100644
--- a/kernel/bpf/const_fold.c
+++ b/kernel/bpf/const_fold.c
@@ -180,9 +180,17 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
bool is_ldsx = mode == BPF_MEMSX;
int off = src->val + insn->off;
u64 val = 0;
+ u32 cnt;
+ /*
+ * Values of insn_array map are addresses of jitted instructions,
+ * which are not known until the program is jitted.
+ */
if (!bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
+ map->map_type == BPF_MAP_TYPE_INSN_ARRAY ||
off < 0 || off + size > map->value_size ||
+ /* so are the addresses of functions that the map points to */
+ bpf_map_range_func_ptrs(env, map, off, size, &cnt) ||
bpf_map_direct_read(map, off, size, &val, is_ldsx)) {
*dst = unknown;
break;
@@ -191,7 +199,8 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
dst->val = val;
break;
case BPF_JMP:
- if (opcode != BPF_CALL)
+ /* both 'call imm' and 'callx reg' clobber caller saved registers */
+ if (BPF_OP(insn->code) != BPF_CALL)
break;
process_call:
for (r = BPF_REG_0; r <= BPF_REG_5; r++)
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 227211166dccf..273f74068068c 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -1831,6 +1831,7 @@ bool bpf_opcode_in_insntable(u8 code)
[BPF_LD | BPF_IND | BPF_H] = true,
[BPF_LD | BPF_IND | BPF_W] = true,
[BPF_JMP | BPF_JA | BPF_X] = true,
+ [BPF_JMP | BPF_CALL | BPF_X] = true,
[BPF_JMP | BPF_JCOND] = true,
};
#undef BPF_INSN_3_TBL
@@ -3028,6 +3029,11 @@ void __bpf_free_used_maps(struct bpf_prog_aux *aux,
map->ops->map_poke_untrack(map, aux);
if (sleepable)
atomic64_dec(&map->sleepable_refcnt);
+ /*
+ * The program that didn't load is not a user of the map. libbpf
+ * loads the program again to get the log of the verifier.
+ */
+ cmpxchg(&map->user, (unsigned long)aux, 0);
bpf_map_put(map);
}
}
@@ -3287,6 +3293,12 @@ bool __weak bpf_jit_supports_kfunc_call(void)
return false;
}
+/* Return TRUE if the JIT backend supports callx (indirect call) instruction. */
+bool __weak bpf_jit_supports_callx(void)
+{
+ return false;
+}
+
bool __weak bpf_jit_supports_kfunc_ret_reg_pair(void)
{
return false;
diff --git a/kernel/bpf/disasm.c b/kernel/bpf/disasm.c
index 3ce8d74b0e400..36d3228d77454 100644
--- a/kernel/bpf/disasm.c
+++ b/kernel/bpf/disasm.c
@@ -350,7 +350,10 @@ void print_bpf_insn(const struct bpf_insn_cbs *cbs,
if (opcode == BPF_CALL) {
char tmp[64];
- if (insn->src_reg == BPF_PSEUDO_CALL) {
+ if (BPF_SRC(insn->code) == BPF_X) {
+ verbose(cbs->private_data, "(%02x) callx r%d",
+ insn->code, insn->dst_reg);
+ } else if (insn->src_reg == BPF_PSEUDO_CALL) {
verbose(cbs->private_data, "(%02x) call pc%s",
insn->code,
__func_get_name(cbs, insn,
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 2add8001c3ec3..e568b9b790b5c 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -361,6 +361,7 @@ struct bpf_prog *bpf_patch_insn_data(struct bpf_verifier_env *env, u32 off,
adjust_insn_aux_data(env, new_prog, off, len, &original_insn);
adjust_subprog_starts(env, off, len);
adjust_insn_arrays(env, off, len);
+ bpf_adjust_func_ptrs(env, off, len);
adjust_poke_descs(new_prog, off, len);
return new_prog;
}
@@ -559,6 +560,9 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
if (err)
return err;
+ /* before subprogs are adjusted, since it looks at them */
+ bpf_adjust_func_ptrs_after_remove(env, off, cnt);
+
err = adjust_subprog_starts_after_remove(env, off, cnt);
if (err)
return err;
@@ -1285,6 +1289,57 @@ static int jit_subprogs(struct bpf_verifier_env *env)
cond_resched();
}
+ /*
+ * The addresses of all functions are final. Replace the offsets of
+ * functions with them in the maps of the program, see
+ * resolve_func_ptrs(). The program must be the only user of such map.
+ * From now on no other program can use it, see bpf_map_claim().
+ */
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ struct bpf_func_ptr *ptr = &env->func_ptrs[i];
+ unsigned long me = (unsigned long)prog->aux;
+ u64 addr, old, new = 0;
+
+ /* pointers are sorted by map */
+ if ((!i || ptr->map != ptr[-1].map) &&
+ cmpxchg(&ptr->map->user, me, me | BPF_MAP_USER_PATCHED) != me) {
+ verbose(env, "map '%s' is used by another program\n", ptr->map->name);
+ err = -EBUSY;
+ goto out_free;
+ }
+
+ /* it's the address of the value of the map whatever the offset is */
+ err = ptr->map->ops->map_direct_value_addr(ptr->map, &addr, 0);
+ if (verifier_bug_if(err, env, "no value of map '%s'", ptr->map->name)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ addr += ptr->map_off;
+
+ if (ptr->xlated_off != BPF_FUNC_PTR_DELETED) {
+ subprog = bpf_find_subprog(env, ptr->xlated_off);
+ if (verifier_bug_if(subprog <= 0, env, "no function at insn %u",
+ ptr->xlated_off)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ new = (unsigned long)func[subprog]->bpf_func;
+ } else if (verifier_bug_if(ptr->used, env, "function of map '%s' offset %u is removed",
+ ptr->map->name, ptr->map_off)) {
+ /* the program that reads the pointer might call the function */
+ err = -EFAULT;
+ goto out_free;
+ }
+ /* else the function is dead code, nothing calls it, the pointer is NULL */
+
+ old = (u64)ptr->orig_off * sizeof(struct bpf_insn);
+ if (verifier_bug_if(cmpxchg64((u64 *)(unsigned long)addr, old, new) != old, env,
+ "map '%s' offset %u changed", ptr->map->name, ptr->map_off)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ }
+
/*
* Cleanup func[i]->aux fields which aren't required
* or can become invalid in future
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 44ecdc5b4ec2d..5aa2f68d92b37 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -356,12 +356,25 @@ int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_
return 0;
}
+/*
+ * Stack accesses of callbacks and of callx callees are not tracked by
+ * func instances keyed by the @callsite. Callbacks might be called several
+ * times and the callee of callx is not known when stack liveness is computed.
+ * In both cases stack slots of the outer frames that might be read by the
+ * callee are accounted as read by the @callsite instruction itself.
+ */
+static bool callee_stack_access_at_callsite(struct bpf_verifier_env *env, u32 callsite)
+{
+ return bpf_calls_callback(env, callsite) ||
+ bpf_is_callx(&env->prog->insnsi[callsite]);
+}
+
bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_spi)
{
/*
* Slot is alive if it is read before q->insn_idx in current func instance,
* or if for some outer func instance:
- * - alive before callsite if callsite calls callback, otherwise
+ * - alive before callsite if callsite calls callback or is callx, otherwise
* - alive after callsite
*/
struct live_stack_query *q = &env->liveness->live_stack_query;
@@ -394,7 +407,7 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
/* Get callsite from verifier state, not from instance callchain */
callsite = q->callsites[i];
- alive = bpf_calls_callback(env, callsite)
+ alive = callee_stack_access_at_callsite(env, callsite)
? is_live_before(instance, callsite, rel, half_spi)
: is_live_before(instance, callsite + 1, rel, half_spi);
if (alive)
@@ -1439,7 +1452,15 @@ static int record_call_access(struct bpf_verifier_env *env,
if (bpf_pseudo_call(insn))
return 0;
- if (bpf_get_call_summary(env, insn, &cs))
+ if (bpf_is_callx(insn))
+ /*
+ * The callee is not known statically. Assume that all arg
+ * slots are passed and let record_arg_access() conservatively
+ * mark the stack of all frames as read if any of them is
+ * derived from a frame pointer.
+ */
+ arg_slot_cnt = MAX_BPF_FUNC_REG_ARGS + MAX_STACK_ARG_SLOTS;
+ else if (bpf_get_call_summary(env, insn, &cs))
arg_slot_cnt = cs.arg_slot_cnt;
for (r = BPF_REG_1; r < BPF_REG_1 + min(arg_slot_cnt, MAX_BPF_FUNC_REG_ARGS); r++) {
@@ -1533,7 +1554,8 @@ static void print_subprog_arg_access(struct bpf_verifier_env *env,
bool has_extra = false;
u8 cls = BPF_CLASS(insns[idx].code);
bool is_ldx_stx_call = cls == BPF_LDX || cls == BPF_STX ||
- insns[idx].code == (BPF_JMP | BPF_CALL);
+ insns[idx].code == (BPF_JMP | BPF_CALL) ||
+ bpf_is_callx(&insns[idx]);
verbose(env, "%3d: ", idx);
bpf_verbose_insn(env, &insns[idx]);
@@ -1722,7 +1744,7 @@ static int compute_subprog_args(struct bpf_verifier_env *env,
if (err)
goto err_free;
- if (insn->code == (BPF_JMP | BPF_CALL)) {
+ if (insn->code == (BPF_JMP | BPF_CALL) || bpf_is_callx(insn)) {
err = record_call_access(env, instance, at_in[i], idx);
if (err)
goto err_free;
@@ -2202,6 +2224,9 @@ static void compute_insn_live_regs(struct bpf_verifier_env *env,
use = GENMASK(min_t(u8, cs.arg_slot_cnt, MAX_BPF_FUNC_REG_ARGS), 1);
def = mask_widen(def);
use = mask_widen(use);
+ /* callx reads the address of the callee from dst_reg */
+ if (bpf_is_callx(insn))
+ use |= dst;
break;
default:
def = 0;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index d62c0f74cff5e..8507114cf690a 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -3018,6 +3018,22 @@ static int add_subprogs(struct bpf_verifier_env *env)
return ret;
}
+ /*
+ * func_info describes all functions of the program. Those that are
+ * referenced only from data, e.g. from a table of functions to be
+ * called via callx, are not seen by the loop above. They are possible
+ * callees that have to be known upfront as well.
+ */
+ if (env->bpf_capable) {
+ struct bpf_prog_aux *aux = env->prog->aux;
+
+ for (i = 1; i < aux->func_info_cnt; i++) {
+ ret = add_subprog(env, aux->func_info[i].insn_off);
+ if (ret < 0)
+ return ret;
+ }
+ }
+
ret = bpf_find_exception_callback_insn_off(env);
if (ret < 0)
return ret;
@@ -3099,6 +3115,8 @@ static int check_subprogs(struct bpf_verifier_env *env)
if (BPF_CLASS(code) == BPF_LD &&
(BPF_MODE(code) == BPF_ABS || BPF_MODE(code) == BPF_IND))
subprog[cur_subprog].has_ld_abs = true;
+ if (bpf_is_callx(&insn[i]))
+ env->has_callx = true;
if (BPF_CLASS(code) != BPF_JMP && BPF_CLASS(code) != BPF_JMP32)
goto next;
if (BPF_OP(code) == BPF_CALL)
@@ -3144,12 +3162,50 @@ static int check_subprogs(struct bpf_verifier_env *env)
return 0;
}
+/*
+ * The callee of callx is known to the main verification pass only, which
+ * records the 'caller' -> 'callee' edge of the call graph for the checks
+ * that follow it: absence of recursion and the maximum stack depth.
+ */
+static int record_callx_edge(struct bpf_verifier_env *env, int caller, int callee)
+{
+ u32 cnt = env->subprog_cnt;
+
+ if (!env->callx_edges) {
+ env->callx_edges = kvcalloc(BITS_TO_LONGS(cnt * cnt), sizeof(long),
+ GFP_KERNEL_ACCOUNT);
+ if (!env->callx_edges)
+ return -ENOMEM;
+ }
+ __set_bit(caller * cnt + callee, env->callx_edges);
+ return 0;
+}
+
+/*
+ * Return the first subprog with the number >= 'from' that 'caller' calls
+ * via callx, or -1 when there is none.
+ */
+static int next_callx_callee(struct bpf_verifier_env *env, int caller, int from)
+{
+ u32 cnt = env->subprog_cnt;
+ unsigned long bit, end = (caller + 1) * cnt;
+
+ if (!env->callx_edges || from >= cnt)
+ return -1;
+ bit = find_next_bit(env->callx_edges, end, caller * cnt + from);
+ return bit < end ? bit - caller * cnt : -1;
+}
+
/*
* Sort subprogs in topological order so that leaf subprogs come first and
* their callers come later. This is a DFS post-order traversal of the call
* graph. Scan only reachable instructions (those in the computed postorder) of
* the current subprog to discover callees (direct subprogs and sync
* callbacks).
+ *
+ * The callees of callx are not known before the main verification pass.
+ * When callx is used the sort is repeated after it with the recorded callx
+ * edges added to the call graph to reject recursion through indirect calls.
*/
static int sort_subprogs_topo(struct bpf_verifier_env *env)
{
@@ -3190,12 +3246,22 @@ static int sort_subprogs_topo(struct bpf_verifier_env *env)
int idx = insn_postorder[j];
int callee;
- if (!bpf_pseudo_call(&insn[idx]) && !bpf_pseudo_func(&insn[idx]))
+ if (bpf_is_callx(&insn[idx])) {
+ /* find a callee that is not explored yet */
+ callee = -1;
+ do {
+ callee = next_callx_callee(env, cur, callee + 1);
+ } while (callee >= 0 && color[callee] == 2);
+ if (callee < 0)
+ continue;
+ } else if (bpf_pseudo_call(&insn[idx]) || bpf_pseudo_func(&insn[idx])) {
+ callee = bpf_find_subprog(env, idx + insn[idx].imm + 1);
+ if (callee < 0) {
+ ret = -EFAULT;
+ goto out;
+ }
+ } else {
continue;
- callee = bpf_find_subprog(env, idx + insn[idx].imm + 1);
- if (callee < 0) {
- ret = -EFAULT;
- goto out;
}
if (color[callee] == 2)
continue;
@@ -5367,6 +5433,9 @@ struct bpf_subprog_call_depth_info {
int ret_insn; /* caller instruction where we return to. */
int caller; /* caller subprogram idx */
int frame; /* # of consecutive static call stack frames on top of stack */
+ int callx_insn; /* callx instruction whose callees are being walked */
+ int callx_next; /* next callee of callx_insn to walk */
+ bool via_callx; /* the subprogram is entered via callx */
};
/* starting from main bpf function walk all instructions of the function
@@ -5386,11 +5455,13 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
/* no caller idx */
dinfo[idx].caller = -1;
+ dinfo[idx].via_callx = false;
i = subprog[idx].start;
if (!priv_stack_supported)
subprog[idx].priv_stack_mode = NO_PRIV_STACK;
process_func:
+ dinfo[idx].callx_insn = -1;
if (subprog[idx].has_ld_abs) {
for (tmp = idx; tmp >= 0; tmp = dinfo[tmp].caller) {
if (subprog[tmp].is_cb) {
@@ -5486,31 +5557,64 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
return -EINVAL;
}
- if (!bpf_pseudo_call(insn + i) && !bpf_pseudo_func(insn + i))
+ if (bpf_is_callx(insn + i)) {
+ /*
+ * Walk the callees recorded by the main verification
+ * pass one by one, returning to this insn after each.
+ */
+ if (dinfo[idx].callx_insn != i) {
+ dinfo[idx].callx_insn = i;
+ dinfo[idx].callx_next = 0;
+ }
+ sidx = next_callx_callee(env, idx, dinfo[idx].callx_next);
+ if (sidx < 0)
+ continue;
+ dinfo[idx].callx_next = sidx + 1;
+ dinfo[idx].ret_insn = i;
+ next_insn = subprog[sidx].start;
+ } else if (bpf_pseudo_call(insn + i) || bpf_pseudo_func(insn + i)) {
+ /* find the callee */
+ next_insn = i + insn[i].imm + 1;
+ sidx = bpf_find_subprog(env, next_insn);
+ if (verifier_bug_if(sidx < 0, env, "callee not found at insn %d", next_insn))
+ return -EFAULT;
+ if (subprog[sidx].is_async_cb) {
+ /* async callbacks don't increase bpf prog stack size unless called directly */
+ if (!bpf_pseudo_call(insn + i))
+ continue;
+ if (subprog[sidx].is_exception_cb) {
+ verbose(env, "insn %d cannot call exception cb directly", i);
+ return -EINVAL;
+ }
+ }
+ /* remember insn to return to */
+ dinfo[idx].ret_insn = i + 1;
+ } else {
continue;
- /* remember insn and function to return to */
+ }
- /* find the callee */
- next_insn = i + insn[i].imm + 1;
- sidx = bpf_find_subprog(env, next_insn);
- if (verifier_bug_if(sidx < 0, env, "callee not found at insn %d", next_insn))
- return -EFAULT;
- if (subprog[sidx].is_async_cb) {
- /* async callbacks don't increase bpf prog stack size unless called directly */
- if (!bpf_pseudo_call(insn + i))
+ /*
+ * sort_subprogs_topo() tolerates cycles in the call graph that
+ * go through the address of a function being taken, since it
+ * doesn't know what it is taken for. Such cycle is a recursion
+ * unless it's an async callback, which are skipped above.
+ * The main verification pass limits the depth of the recursion,
+ * but it doesn't follow calls of global functions.
+ */
+ for (tmp = idx; tmp >= 0; tmp = dinfo[tmp].caller) {
+ if (tmp != sidx)
continue;
- if (subprog[sidx].is_exception_cb) {
- verbose(env, "insn %d cannot call exception cb directly", i);
- return -EINVAL;
- }
+ verbose(env, "recursive call from %s() to %s()\n",
+ bpf_subprog_name(env, idx), bpf_subprog_name(env, sidx));
+ return -EINVAL;
}
/* store caller info for after we return from callee */
dinfo[idx].frame = frame;
- dinfo[idx].ret_insn = i + 1;
/* push caller idx into callee's dinfo */
dinfo[sidx].caller = idx;
+ dinfo[sidx].via_callx = bpf_is_callx(insn + i);
i = next_insn;
@@ -5544,6 +5648,14 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
verbose(env, "tail_calls are not allowed in programs with stack args\n");
return -EINVAL;
}
+ /*
+ * JITs pass tail call counter in a register that is
+ * not available when the callee is called via callx.
+ */
+ if (dinfo[tmp].via_callx) {
+ verbose(env, "tail_calls are not allowed in functions called via callx\n");
+ return -EINVAL;
+ }
subprog[tmp].tail_call_reachable = true;
}
} else if (!idx && subprog[0].has_tail_call && subprog[0].stack_arg_cnt) {
@@ -5908,6 +6020,110 @@ int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,
return 0;
}
+static int cmp_func_ptrs(const void *_a, const void *_b)
+{
+ const struct bpf_func_ptr *a = _a, *b = _b;
+
+ if (a->map != b->map)
+ return a->map < b->map ? -1 : 1;
+ if (a->map_off != b->map_off)
+ return a->map_off < b->map_off ? -1 : 1;
+ return 0;
+}
+
+/* Find the first pointer to a function at or after 'off' in the value of 'map' */
+static u32 func_ptr_lower_bound(struct bpf_verifier_env *env, const struct bpf_map *map, u64 off)
+{
+ u32 l = 0, r = env->func_ptr_cnt, m;
+ struct bpf_func_ptr *p;
+
+ while (l < r) {
+ m = l + (r - l) / 2;
+ p = &env->func_ptrs[m];
+ if (p->map < map || (p->map == map && p->map_off < off))
+ l = m + 1;
+ else
+ r = m;
+ }
+ return l;
+}
+
+/*
+ * Return pointers to functions that overlap with 'size' bytes at offset 'off'
+ * of the value of 'map' and their number in 'cnt'.
+ */
+struct bpf_func_ptr *bpf_map_range_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map,
+ u64 off, u64 size, u32 *cnt)
+{
+ u32 first, last;
+
+ *cnt = 0;
+ if (!env->func_ptr_cnt || !size)
+ return NULL;
+
+ /* a pointer that starts up to 7 bytes before 'off' overlaps too */
+ first = func_ptr_lower_bound(env, map, off >= sizeof(u64) ? off - sizeof(u64) + 1 : 0);
+ last = func_ptr_lower_bound(env, map, off + size);
+ if (first >= last)
+ return NULL;
+
+ *cnt = last - first;
+ return &env->func_ptrs[first];
+}
+
+/* Return all pointers to functions in the value of 'map' */
+struct bpf_func_ptr *bpf_map_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map, u32 *cnt)
+{
+ return bpf_map_range_func_ptrs(env, map, 0, (u64)map->value_size, cnt);
+}
+
+/* instructions [off, off + len) replaced the instruction at 'off' */
+void bpf_adjust_func_ptrs(struct bpf_verifier_env *env, u32 off, u32 len)
+{
+ struct bpf_func_ptr *p;
+ u32 i;
+
+ if (len <= 1)
+ return;
+
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ p = &env->func_ptrs[i];
+ if (p->xlated_off <= off || p->xlated_off == BPF_FUNC_PTR_DELETED)
+ continue;
+ p->xlated_off += len - 1;
+ }
+}
+
+/*
+ * Instructions [off, off + len) are about to be removed. It's called before
+ * the starts of subprogs are adjusted. A subprog is gone when all of its
+ * instructions are. Otherwise, e.g. when its first instruction is a nop,
+ * it starts where the removed instructions did.
+ */
+void bpf_adjust_func_ptrs_after_remove(struct bpf_verifier_env *env, u32 off, u32 len)
+{
+ struct bpf_func_ptr *p;
+ int subprog;
+ u32 i;
+
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ p = &env->func_ptrs[i];
+ if (p->xlated_off < off || p->xlated_off == BPF_FUNC_PTR_DELETED)
+ continue;
+ if (p->xlated_off >= off + len) {
+ p->xlated_off -= len;
+ continue;
+ }
+ subprog = bpf_find_subprog(env, p->xlated_off);
+ if (subprog > 0 && env->subprog_info[subprog + 1].start > off + len)
+ p->xlated_off = off;
+ else
+ p->xlated_off = BPF_FUNC_PTR_DELETED;
+ }
+}
+
#define BTF_TYPE_SAFE_RCU(__type) __PASTE(__type, __safe_rcu)
#define BTF_TYPE_SAFE_RCU_OR_NULL(__type) __PASTE(__type, __safe_rcu_or_null)
#define BTF_TYPE_SAFE_TRUSTED(__type) __PASTE(__type, __safe_trusted)
@@ -6417,12 +6633,100 @@ static void add_scalar_to_reg(struct bpf_reg_state *dst_reg, s64 val)
reg_bounds_sync(dst_reg);
}
+static void mark_reg_func_ptr(struct bpf_verifier_env *env, struct bpf_reg_state *regs,
+ int regno, int subprog)
+{
+ mark_reg_known_zero(env, regs, regno);
+ regs[regno].type = PTR_TO_FUNC;
+ regs[regno].subprogno = subprog;
+}
+
+/* a read from a table of functions branches into that many states at most */
+#define BPF_MAX_FUNC_PTR_TARGETS 64
+/* and the table, which might have other data in it, is that many pointers long at most */
+#define BPF_MAX_FUNC_PTR_RANGE 4096
+
+/*
+ * A read from a frozen read-only map that has pointers to functions, see
+ * resolve_func_ptrs(). A read of exactly one pointer yields PTR_TO_FUNC.
+ * When the offset is variable and only pointers can be read, which is how
+ * an element of a table of functions is loaded, the verification continues
+ * with each of them. Other reads that overlap with a pointer are rejected,
+ * because their result is not known until the program is jitted.
+ *
+ * Return -ENOENT if there are no pointers to functions in the bytes that are read.
+ */
+static int check_func_ptr_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
+ int size, int value_regno)
+{
+ u64 min_off = reg_umin(reg) + off, max_off = reg_umax(reg) + off;
+ struct tnum offs = tnum_add(reg->var_off, tnum_const(off));
+ struct bpf_reg_state *regs = cur_regs(env);
+ struct bpf_map *map = reg->map_ptr;
+ struct bpf_verifier_state *branch;
+ int subprog, targets[BPF_MAX_FUNC_PTR_TARGETS];
+ struct bpf_func_ptr *ptrs;
+ u32 i, cnt, n = 0;
+ u64 o;
+
+ ptrs = bpf_map_range_func_ptrs(env, map, min_off, max_off - min_off + size, &cnt);
+ if (!ptrs)
+ return -ENOENT;
+
+ if (size != sizeof(u64) || value_regno < 0 || !tnum_is_aligned(offs, sizeof(u64)) ||
+ (max_off - min_off) / sizeof(u64) > BPF_MAX_FUNC_PTR_RANGE)
+ goto overlap;
+
+ /*
+ * Every offset that the read is possible at has to be the offset of
+ * a pointer. var_off tells the stride of the elements of an array.
+ */
+ for (o = round_up(min_off, sizeof(u64)), i = 0; o <= max_off; o += sizeof(u64)) {
+ if ((o ^ offs.value) & ~offs.mask)
+ continue;
+ while (i < cnt && ptrs[i].map_off < o)
+ i++;
+ if (i == cnt || ptrs[i].map_off != o)
+ goto overlap;
+ if (n == BPF_MAX_FUNC_PTR_TARGETS) {
+ verbose(env, "read from map '%s' may yield more than %d pointers to functions\n",
+ map->name, BPF_MAX_FUNC_PTR_TARGETS);
+ return -E2BIG;
+ }
+ subprog = bpf_find_subprog(env, ptrs[i].xlated_off);
+ if (verifier_bug_if(subprog <= 0, env, "no function at insn %u for map '%s' offset %u",
+ ptrs[i].xlated_off, map->name, ptrs[i].map_off))
+ return -EFAULT;
+ ptrs[i].used = true;
+ targets[n++] = subprog;
+ }
+ if (verifier_bug_if(!n, env, "no offsets to read map '%s' at", map->name))
+ return -EFAULT;
+
+ for (i = 0; i < n - 1; i++) {
+ branch = push_stack(env, env->insn_idx + 1, env->insn_idx,
+ env->cur_state->speculative);
+ if (IS_ERR(branch))
+ return PTR_ERR(branch);
+ mark_reg_func_ptr(env, branch->frame[branch->curframe]->regs, value_regno,
+ targets[i]);
+ }
+ mark_reg_func_ptr(env, regs, value_regno, targets[n - 1]);
+ return 0;
+
+overlap:
+ verbose(env, "read of %d bytes at offset [%llu,%llu] of map '%s' overlaps with a pointer to a function\n",
+ size, min_off, max_off, map->name);
+ return -EACCES;
+}
+
static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
int bpf_size, int value_regno, bool is_ldsx)
{
struct bpf_reg_state *regs = cur_regs(env);
int size = bpf_size_to_bytes(bpf_size);
struct bpf_map *map = reg->map_ptr;
+ int err;
switch (map->map_type) {
case BPF_MAP_TYPE_INSN_ARRAY:
@@ -6440,13 +6744,18 @@ static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state
break;
}
+ if (env->func_ptr_cnt) {
+ err = check_func_ptr_read(env, reg, off, size, value_regno);
+ if (err != -ENOENT)
+ return err;
+ }
+
/* If map is read-only, track its contents as scalars. */
if (tnum_is_const(reg->var_off) &&
bpf_map_is_rdonly(map) &&
map->ops->map_direct_value_addr) {
int map_off = off + reg->var_off.value;
u64 val = 0;
- int err;
err = bpf_map_direct_read(map, map_off, size, &val, is_ldsx);
if (err)
@@ -8754,6 +9063,7 @@ static int check_arg_const_str(struct bpf_verifier_env *env,
int map_off;
u64 map_addr;
char *str_ptr;
+ u32 cnt;
if (reg->type != PTR_TO_MAP_VALUE)
return -EINVAL;
@@ -8803,6 +9113,11 @@ static int check_arg_const_str(struct bpf_verifier_env *env,
verbose(env, "string is not zero-terminated\n");
return -EINVAL;
}
+ /* the bytes of a pointer to a function are not known until the program is jitted */
+ if (bpf_map_range_func_ptrs(env, map, map_off, strlen(str_ptr + map_off) + 1, &cnt)) {
+ verbose(env, "string overlaps with a pointer to a function\n");
+ return -EACCES;
+ }
return 0;
}
@@ -10584,13 +10899,61 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
static int process_bpf_exit_full(struct bpf_verifier_env *env,
bool *do_print_state, bool exception_exit);
-static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
- int *insn_idx)
+/*
+ * Call of a static subprog. The callee is verified in the context of
+ * the caller, hence set up a new frame and continue from the first
+ * instruction of the callee.
+ */
+static int check_static_func_call(struct bpf_verifier_env *env, int subprog,
+ int *insn_idx)
{
struct bpf_verifier_state *state = env->cur_state;
struct bpf_subprog_info *caller_info;
u16 callee_incoming, stack_arg_cnt;
struct bpf_func_state *caller;
+ int err;
+
+ caller = state->frame[state->curframe];
+
+ /*
+ * Track caller's total stack arg count (incoming + max outgoing).
+ * This is needed so the JIT knows how much stack arg space to allocate.
+ */
+ caller_info = &env->subprog_info[caller->subprogno];
+ callee_incoming = bpf_in_stack_arg_cnt(&env->subprog_info[subprog]);
+ stack_arg_cnt = bpf_in_stack_arg_cnt(caller_info) + callee_incoming;
+ if (stack_arg_cnt > caller_info->stack_arg_cnt)
+ caller_info->stack_arg_cnt = stack_arg_cnt;
+
+ /*
+ * For regular function entry setup new frame and continue
+ * from that frame.
+ */
+ err = setup_func_entry(env, subprog, *insn_idx, set_callee_state, state);
+ if (err)
+ return err;
+
+ bpf_diag_record_scrub(env, &caller->regs[BPF_REG_0], BPF_DIAG_MOD_CALLER_SAVED);
+ clear_caller_saved_regs(env, caller->regs);
+
+ /* and go analyze first insn of the callee */
+ *insn_idx = env->subprog_info[subprog].start - 1;
+
+ if (env->log.level & BPF_LOG_LEVEL) {
+ verbose(env, "caller:\n");
+ print_verifier_state(env, state, caller->frameno, true);
+ verbose(env, "callee:\n");
+ print_verifier_state(env, state, state->curframe, true);
+ }
+
+ return 0;
+}
+
+static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
+ int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *caller;
int err, subprog, target_insn;
u32 i, nregs;
@@ -10679,37 +11042,75 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
return 0;
}
- /*
- * Track caller's total stack arg count (incoming + max outgoing).
- * This is needed so the JIT knows how much stack arg space to allocate.
- */
- caller_info = &env->subprog_info[caller->subprogno];
- callee_incoming = bpf_in_stack_arg_cnt(&env->subprog_info[subprog]);
- stack_arg_cnt = bpf_in_stack_arg_cnt(caller_info) + callee_incoming;
- if (stack_arg_cnt > caller_info->stack_arg_cnt)
- caller_info->stack_arg_cnt = stack_arg_cnt;
+ return check_static_func_call(env, subprog, insn_idx);
+}
- /* for regular function entry setup new frame and continue
- * from that frame.
- */
- err = setup_func_entry(env, subprog, *insn_idx, set_callee_state, state);
+/*
+ * callx dst_reg: call a bpf subprog whose address is in dst_reg.
+ *
+ * The address of a subprog is either loaded into a register by ld_imm64 with
+ * src_reg == BPF_PSEUDO_FUNC, or it is read from a frozen read-only map, see
+ * resolve_func_ptrs(). Both are possible for static subprogs only. Hence all
+ * possible callees of callx are discovered by add_subprogs() and are reachable
+ * in the control flow graph before the main verification pass begins.
+ * PTR_TO_FUNC register identifies the callee, so from here on callx is verified
+ * as a direct call of that static subprog.
+ */
+static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,
+ int *insn_idx)
+{
+ struct bpf_func_state *caller = cur_func(env);
+ struct bpf_reg_state *reg;
+ const char *reason;
+ int err, subprog;
+
+ err = check_reg_arg(env, insn->dst_reg, SRC_OP);
if (err)
return err;
- bpf_diag_record_scrub(env, &caller->regs[BPF_REG_0], BPF_DIAG_MOD_CALLER_SAVED);
- clear_caller_saved_regs(env, caller->regs);
+ reg = reg_state(env, insn->dst_reg);
+ if (reg->type != PTR_TO_FUNC) {
+ verbose(env, "R%d has type %s, expected func\n", insn->dst_reg,
+ reg_type_str(env, reg->type));
+ reason = bpf_diag_fmt(
+ env, "R%d holds %s, but callx can only call through the address of a static BPF function.",
+ insn->dst_reg, bpf_diag_reg_type_plain(env, reg->type));
+ bpf_diag_register_type(
+ env, *insn_idx, insn->dst_reg, "indirect call through a non-function pointer", reason,
+ "Load the address of a static BPF function into the register before callx.");
+ return -EACCES;
+ }
- /* and go analyze first insn of the callee */
- *insn_idx = env->subprog_info[subprog].start - 1;
+ /*
+ * Arithmetic on PTR_TO_FUNC is allowed, but only unmodified address
+ * of a subprog can be called.
+ */
+ err = check_ptr_off_reg(env, reg, insn->dst_reg);
+ if (err)
+ return err;
- if (env->log.level & BPF_LOG_LEVEL) {
- verbose(env, "caller:\n");
- print_verifier_state(env, state, caller->frameno, true);
- verbose(env, "callee:\n");
- print_verifier_state(env, state, state->curframe, true);
+ /* there is no support for callx in the interpreter */
+ if (!env->prog->jit_requested) {
+ verbose(env, "JIT is required to use callx\n");
+ return -EOPNOTSUPP;
+ }
+ if (!bpf_jit_supports_callx()) {
+ verbose(env, "JIT doesn't support callx\n");
+ return -EOPNOTSUPP;
}
+ env->prog->jit_required = true;
- return 0;
+ /* PTR_TO_FUNC is a pointer to a static subprog */
+ subprog = reg->subprogno;
+ err = btf_check_subprog_call(env, subprog, caller->regs);
+ if (err == -EFAULT)
+ return err;
+
+ err = record_callx_edge(env, caller->subprogno, subprog);
+ if (err)
+ return err;
+
+ return check_static_func_call(env, subprog, insn_idx);
}
int map_set_for_each_callback_args(struct bpf_verifier_env *env,
@@ -18709,7 +19110,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->jmps_processed++;
if (opcode == BPF_CALL) {
- if (env->cur_state->active_locks) {
+ /* similar to static subprog calls callx is allowed under a lock */
+ if (env->cur_state->active_locks && !bpf_is_callx(insn)) {
if ((insn->src_reg == BPF_REG_0 &&
insn->imm != BPF_FUNC_spin_unlock &&
insn->imm != BPF_FUNC_kptr_xchg) ||
@@ -18727,6 +19129,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
mark_reg_scratched(env, BPF_REG_0);
if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
cur_func(env)->no_stack_arg_load = true;
+ if (bpf_is_callx(insn))
+ return check_func_callx(env, insn, &env->insn_idx);
if (insn->src_reg == BPF_PSEUDO_CALL)
return check_func_call(env, insn, &env->insn_idx);
if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL)
@@ -19302,6 +19706,30 @@ static int check_map_prog_compatibility(struct bpf_verifier_env *env,
return 0;
}
+/*
+ * Keep track of whether the map is used by one program only. Such program may
+ * store the addresses of its functions into the map when it's frozen, see
+ * resolve_func_ptrs(), since nothing else relies on what the map has. After
+ * that the map is not available to other programs.
+ */
+static int bpf_map_claim(struct bpf_verifier_env *env, struct bpf_map *map)
+{
+ unsigned long me = (unsigned long)env->prog->aux, old;
+
+ for (;;) {
+ old = READ_ONCE(map->user);
+ if (old == me || old == BPF_MAP_USER_MANY)
+ return 0;
+ if (old & BPF_MAP_USER_PATCHED) {
+ verbose(env, "map '%s' has addresses of functions of another program\n",
+ map->name);
+ return -EBUSY;
+ }
+ if (cmpxchg(&map->user, old, old ? BPF_MAP_USER_MANY : me) == old)
+ return 0;
+ }
+}
+
static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)
{
int i, err;
@@ -19339,6 +19767,10 @@ static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)
env->used_maps[env->used_map_cnt++] = map;
+ err = bpf_map_claim(env, map);
+ if (err)
+ return err;
+
if (map->map_type == BPF_MAP_TYPE_INSN_ARRAY) {
err = bpf_insn_array_init(map, env->prog);
if (err) {
@@ -19493,6 +19925,14 @@ static int check_jmp_fields(struct bpf_verifier_env *env, struct bpf_insn *insn)
switch (opcode) {
case BPF_CALL:
+ if (bpf_is_callx(insn)) {
+ /* callx dst_reg */
+ if (insn->src_reg != BPF_REG_0 || insn->imm != 0 || insn->off != 0) {
+ verbose(env, "BPF_CALL|BPF_X uses reserved fields\n");
+ return -EINVAL;
+ }
+ return 0;
+ }
if (BPF_SRC(insn->code) != BPF_K ||
(insn->src_reg != BPF_PSEUDO_KFUNC_CALL && insn->off != 0) ||
(insn->src_reg != BPF_REG_0 && insn->src_reg != BPF_PSEUDO_CALL &&
@@ -19747,6 +20187,93 @@ static int check_and_resolve_insns(struct bpf_verifier_env *env)
return 0;
}
+static int add_func_ptr(struct bpf_verifier_env *env, struct bpf_map *map, u32 map_off,
+ u32 xlated_off)
+{
+ struct bpf_func_ptr *ptrs;
+
+ /* grow by doubling, the array is sorted and searched later */
+ if (!(env->func_ptr_cnt & (env->func_ptr_cnt - 1))) {
+ ptrs = kvrealloc(env->func_ptrs,
+ array_size(max(2 * env->func_ptr_cnt, 16U), sizeof(*ptrs)),
+ GFP_KERNEL_ACCOUNT);
+ if (!ptrs)
+ return -ENOMEM;
+ env->func_ptrs = ptrs;
+ }
+ env->func_ptrs[env->func_ptr_cnt++] = (struct bpf_func_ptr){
+ .map = map,
+ .map_off = map_off,
+ .orig_off = xlated_off,
+ .xlated_off = xlated_off,
+ };
+ return 0;
+}
+
+/*
+ * Compilers put pointers to functions into read-only data: tables of functions,
+ * structures of operations, vtables, where they are mixed with other data.
+ * The loader stores such data in a frozen read-only array map and resolves
+ * a pointer to a static function to the offset in bytes of its first
+ * instruction in the program: the address of the function in the program.
+ *
+ * Find 64-bit values that look like that in the maps of a program that uses
+ * callx. It's a guess. When the value is not a pointer, the program either
+ * fails to load, because it does with a pointer what can be done with
+ * a number only, or it sees the address of a function instead of the number.
+ * It's known before the control flow graph of the program is built and the main
+ * verification pass begins which functions may be called via callx.
+ *
+ * When the program is jitted the offsets are replaced with the addresses of
+ * the functions in the map itself, see jit_subprogs(). Hence the program has to
+ * be the only user of the map, see bpf_map_claim(): nothing else may rely on
+ * what the map had.
+ *
+ * The program reads the addresses of its functions from there like any other
+ * data, so it has to be allowed to leak pointers.
+ */
+static int resolve_func_ptrs(struct bpf_verifier_env *env)
+{
+ int insn_cnt = env->prog->len;
+ int i, err, subprog;
+ struct bpf_map *map;
+ u64 addr, val;
+ u32 off;
+
+ if (!env->has_callx || !env->allow_ptr_leaks)
+ return 0;
+
+ for (i = 0; i < env->used_map_cnt; i++) {
+ map = env->used_maps[i];
+ /* coincidences in maps that are shared with other programs don't matter */
+ if (READ_ONCE(map->user) != (unsigned long)env->prog->aux)
+ continue;
+ if (map->map_type != BPF_MAP_TYPE_ARRAY || map->max_entries != 1 ||
+ !bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
+ !IS_ERR_OR_NULL(map->record))
+ continue;
+ if (map->ops->map_direct_value_addr(map, &addr, 0))
+ continue;
+
+ for (off = 0; off + sizeof(u64) <= map->value_size; off += sizeof(u64)) {
+ val = *(u64 *)(unsigned long)(addr + off);
+ if (!val || val % sizeof(struct bpf_insn) ||
+ val / sizeof(struct bpf_insn) >= insn_cnt)
+ continue;
+ subprog = bpf_find_subprog(env, val / sizeof(struct bpf_insn));
+ if (subprog <= 0 || bpf_subprog_is_global(env, subprog))
+ continue;
+ err = add_func_ptr(env, map, off, val / sizeof(struct bpf_insn));
+ if (err)
+ return err;
+ }
+ }
+ if (env->func_ptr_cnt)
+ sort(env->func_ptrs, env->func_ptr_cnt, sizeof(*env->func_ptrs),
+ cmp_func_ptrs, NULL);
+ return 0;
+}
+
/* drop refcnt of maps used by the rejected program */
static void release_maps(struct bpf_verifier_env *env)
{
@@ -21717,6 +22244,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ /* Find pointers to functions in the read-only maps of the program. */
+ ret = resolve_func_ptrs(env);
+ if (ret < 0)
+ goto skip_full_check;
+
/* Build kfunc prototypes after resolving program resources. */
ret = add_kfuncs(env);
if (ret < 0)
@@ -21776,6 +22308,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
ret = do_check_main(env);
ret = ret ?: do_check_subprogs(env);
+ /* reject recursion through the callx edges found by the main pass */
+ if (ret == 0 && env->callx_edges)
+ ret = sort_subprogs_topo(env);
+
if (ret == 0 && bpf_prog_is_offloaded(env->prog->aux))
ret = bpf_prog_offload_finalize(env);
@@ -21918,6 +22454,8 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
kvfree(env->scc_info);
kvfree(env->succ);
kvfree(env->gotox_tmp_buf);
+ kvfree(env->callx_edges);
+ kvfree(env->func_ptrs);
bpf_diag_free(env);
kvfree(env);
return ret;
diff --git a/tools/lib/bpf/bpf_gen_internal.h b/tools/lib/bpf/bpf_gen_internal.h
index 6c5ad6c55e8a6..206adf28793d7 100644
--- a/tools/lib/bpf/bpf_gen_internal.h
+++ b/tools/lib/bpf/bpf_gen_internal.h
@@ -51,9 +51,17 @@ struct bpf_gen {
__u32 nr_ksyms;
int fd_array;
int nr_fd_array;
+ /*
+ * Maps with pointers to functions, that programs get their own copies
+ * of, take slots in fd_array after nr_obj_maps maps of the object.
+ */
+ __u32 nr_obj_maps;
+ __u32 nr_func_ptr_maps;
+ __u32 max_func_ptr_maps;
};
-void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps);
+void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps,
+ int max_func_ptr_maps);
int bpf_gen__finish(struct bpf_gen *gen, int nr_progs, int nr_maps);
void bpf_gen__free(struct bpf_gen *gen);
void bpf_gen__load_btf(struct bpf_gen *gen, const void *raw_data, __u32 raw_size);
@@ -68,6 +76,9 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *value, __u32 value_size,
__u64 flags);
void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx);
+int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,
+ void *value, __u32 value_size, const __u32 *ptr_offs,
+ const __u64 *ptr_vals, int ptr_cnt);
void bpf_gen__record_attach_target(struct bpf_gen *gen, const char *name, enum bpf_attach_type type);
void bpf_gen__record_extern(struct bpf_gen *gen, const char *name, bool is_weak,
bool is_typeless, bool is_ld64, int kind, int insn_idx);
diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index af3a04f161ac1..251392aa8b41f 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -112,13 +112,23 @@ static void emit2(struct bpf_gen *gen, struct bpf_insn insn1, struct bpf_insn in
static int add_data(struct bpf_gen *gen, const void *data, __u32 size);
static void emit_sys_close_blob(struct bpf_gen *gen, int blob_off);
-void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps)
+void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps,
+ int max_func_ptr_maps)
{
size_t stack_sz = sizeof(struct loader_stack), nr_progs_sz;
int i;
gen->fd_array = add_data(gen, NULL, MAX_FD_ARRAY_SZ * sizeof(int));
gen->log_level = log_level;
+ gen->nr_obj_maps = nr_maps;
+ gen->max_func_ptr_maps = max_func_ptr_maps;
+ if (nr_maps + max_func_ptr_maps > MAX_USED_MAPS) {
+ pr_warn("Total maps exceeds %d\n", MAX_USED_MAPS);
+ gen->error = -E2BIG;
+ return;
+ }
+ /* their fds are closed like the fds of the maps of the object when loading fails */
+ nr_maps += max_func_ptr_maps;
/* save ctx pointer into R6 */
emit(gen, BPF_MOV64_REG(BPF_REG_6, BPF_REG_1));
@@ -385,6 +395,9 @@ int bpf_gen__finish(struct bpf_gen *gen, int nr_progs, int nr_maps)
return gen->error;
}
emit_sys_close_stack(gen, stack_off(btf_fd));
+ /* programs hold their maps with pointers to functions, nothing else needs them */
+ for (i = 0; i < gen->nr_func_ptr_maps; i++)
+ emit_sys_close_blob(gen, blob_fd_array_off(gen, gen->nr_obj_maps + i));
for (i = 0; i < gen->nr_progs; i++)
move_stack2ctx(gen,
sizeof(struct bpf_loader_ctx) +
@@ -1127,12 +1140,59 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
gen->nr_progs++;
}
+/*
+ * if (map_desc[map_idx].initial_value) {
+ * if (ctx->flags & BPF_SKEL_KERNEL)
+ * bpf_probe_read_kernel(value, value_size, initial_value);
+ * else
+ * bpf_copy_from_user(value, value_size, initial_value);
+ * nr_more_insns that the caller emits
+ * }
+ */
+static void emit_copy_initial_value(struct bpf_gen *gen, int map_idx, int value,
+ __u32 value_size, int nr_more_insns)
+{
+ emit(gen, BPF_LDX_MEM(BPF_DW, BPF_REG_3, BPF_REG_6,
+ sizeof(struct bpf_loader_ctx) +
+ sizeof(struct bpf_map_desc) * map_idx +
+ offsetof(struct bpf_map_desc, initial_value)));
+ emit(gen, BPF_JMP_IMM(BPF_JEQ, BPF_REG_3, 0, 8 + nr_more_insns));
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, value));
+ emit(gen, BPF_MOV64_IMM(BPF_REG_2, value_size));
+ emit(gen, BPF_LDX_MEM(BPF_W, BPF_REG_0, BPF_REG_6,
+ offsetof(struct bpf_loader_ctx, flags)));
+ emit(gen, BPF_JMP_IMM(BPF_JSET, BPF_REG_0, BPF_SKEL_KERNEL, 2));
+ emit(gen, BPF_EMIT_CALL(BPF_FUNC_copy_from_user));
+ emit(gen, BPF_JMP_IMM(BPF_JA, 0, 0, 1));
+ emit(gen, BPF_EMIT_CALL(BPF_FUNC_probe_read_kernel));
+}
+
+/* Update the element of the map whose fd is in the slot map_idx of fd_array */
+static void emit_map_update_elem(struct bpf_gen *gen, int map_idx, union bpf_attr *attr,
+ int attr_size, int key, int value, __u32 value_size)
+{
+ int map_update_attr;
+
+ map_update_attr = add_data(gen, attr, attr_size);
+ pr_debug("gen: map_update_elem: idx %d, value: off %d size %u, attr: off %d size %d\n",
+ map_idx, value, value_size, map_update_attr, attr_size);
+ move_blob2blob(gen, attr_field(map_update_attr, map_fd), 4,
+ blob_fd_array_off(gen, map_idx));
+ emit_rel_store(gen, attr_field(map_update_attr, key), key);
+ emit_rel_store(gen, attr_field(map_update_attr, value), value);
+ /* emit MAP_UPDATE_ELEM command */
+ emit_sys_bpf(gen, BPF_MAP_UPDATE_ELEM, map_update_attr, attr_size);
+ debug_ret(gen, "update_elem idx %d value_size %d", map_idx, value_size);
+ emit_check_err(gen);
+}
+
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,
__u32 value_size, __u64 flags)
{
int attr_size = offsetofend(union bpf_attr, flags);
- int map_update_attr, value, key;
union bpf_attr attr;
+ int value, key;
int zero = 0;
memset(&attr, 0, attr_size);
@@ -1142,47 +1202,16 @@ void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,
key = add_data(gen, &zero, sizeof(zero));
/*
- * if (map_desc[map_idx].initial_value) {
- * if (ctx->flags & BPF_SKEL_KERNEL)
- * bpf_probe_read_kernel(value, value_size, initial_value);
- * else
- * bpf_copy_from_user(value, value_size, initial_value);
- * }
- *
* The runtime initial_value comes from the host-supplied loader
* ctx and would overwrite the blob value that the program signature
* covers and the kernel verifies at load time. For a signed loader
* (gen_hash) the attested blob value must be authoritative, so skip
* the override and leave the signed value in place.
*/
- if (!OPTS_GET(gen->opts, gen_hash, false)) {
- emit(gen, BPF_LDX_MEM(BPF_DW, BPF_REG_3, BPF_REG_6,
- sizeof(struct bpf_loader_ctx) +
- sizeof(struct bpf_map_desc) * map_idx +
- offsetof(struct bpf_map_desc, initial_value)));
- emit(gen, BPF_JMP_IMM(BPF_JEQ, BPF_REG_3, 0, 8));
- emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
- 0, 0, 0, value));
- emit(gen, BPF_MOV64_IMM(BPF_REG_2, value_size));
- emit(gen, BPF_LDX_MEM(BPF_W, BPF_REG_0, BPF_REG_6,
- offsetof(struct bpf_loader_ctx, flags)));
- emit(gen, BPF_JMP_IMM(BPF_JSET, BPF_REG_0, BPF_SKEL_KERNEL, 2));
- emit(gen, BPF_EMIT_CALL(BPF_FUNC_copy_from_user));
- emit(gen, BPF_JMP_IMM(BPF_JA, 0, 0, 1));
- emit(gen, BPF_EMIT_CALL(BPF_FUNC_probe_read_kernel));
- }
+ if (!OPTS_GET(gen->opts, gen_hash, false))
+ emit_copy_initial_value(gen, map_idx, value, value_size, 0);
- map_update_attr = add_data(gen, &attr, attr_size);
- pr_debug("gen: map_update_elem: idx %d, value: off %d size %u, attr: off %d size %d\n",
- map_idx, value, value_size, map_update_attr, attr_size);
- move_blob2blob(gen, attr_field(map_update_attr, map_fd), 4,
- blob_fd_array_off(gen, map_idx));
- emit_rel_store(gen, attr_field(map_update_attr, key), key);
- emit_rel_store(gen, attr_field(map_update_attr, value), value);
- /* emit MAP_UPDATE_ELEM command */
- emit_sys_bpf(gen, BPF_MAP_UPDATE_ELEM, map_update_attr, attr_size);
- debug_ret(gen, "update_elem idx %d value_size %d", map_idx, value_size);
- emit_check_err(gen);
+ emit_map_update_elem(gen, map_idx, &attr, attr_size, key, value, value_size);
}
void bpf_gen__populate_outer_map(struct bpf_gen *gen, int outer_map_idx, int slot,
@@ -1214,6 +1243,77 @@ void bpf_gen__populate_outer_map(struct bpf_gen *gen, int outer_map_idx, int slo
emit_check_err(gen);
}
+/*
+ * A copy of the read-only data map obj_map_idx for a program, with the offsets
+ * of its functions in it, see create_func_ptr_map() in libbpf.c. It's not
+ * a map of the object: it's not in the loader ctx. Its content is the content
+ * of obj_map_idx, that the host may supply when the skeleton is loaded, with
+ * ptr_cnt 64-bit ptr_vals at ptr_offs.
+ * Return the index of the map in fd_array for instructions to refer to.
+ */
+int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,
+ void *pvalue, __u32 value_size, const __u32 *ptr_offs,
+ const __u64 *ptr_vals, int ptr_cnt)
+{
+ int attr_size = offsetofend(union bpf_attr, map_extra);
+ int map_create_attr, map_idx, key, value, zero = 0, i;
+ union bpf_attr attr;
+
+ if (gen->nr_func_ptr_maps == gen->max_func_ptr_maps) {
+ gen->error = -EDOM; /* internal bug */
+ return 0;
+ }
+ map_idx = gen->nr_obj_maps + gen->nr_func_ptr_maps++;
+
+ memset(&attr, 0, attr_size);
+ attr.map_type = tgt_endian(BPF_MAP_TYPE_ARRAY);
+ attr.key_size = tgt_endian((__u32)sizeof(int));
+ attr.value_size = tgt_endian(value_size);
+ attr.max_entries = tgt_endian((__u32)1);
+ attr.map_flags = tgt_endian((__u32)BPF_F_RDONLY_PROG);
+ if (map_name)
+ libbpf_strlcpy(attr.map_name, map_name, sizeof(attr.map_name));
+
+ map_create_attr = add_data(gen, &attr, attr_size);
+ pr_debug("gen: func_ptr_map_create: %s idx %d value_size %u, attr: off %d size %d\n",
+ map_name, map_idx, value_size, map_create_attr, attr_size);
+ emit_sys_bpf(gen, BPF_MAP_CREATE, map_create_attr, attr_size);
+ debug_ret(gen, "func_ptr_map_create %s idx %d value_size %d", map_name, map_idx,
+ value_size);
+ emit_check_err(gen);
+ /* remember map_fd in fd_array */
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, blob_fd_array_off(gen, map_idx)));
+ emit(gen, BPF_STX_MEM(BPF_W, BPF_REG_1, BPF_REG_7, 0));
+
+ /* pvalue has the pointers already */
+ value = add_data(gen, pvalue, value_size);
+ key = add_data(gen, &zero, sizeof(zero));
+
+ /* see bpf_gen__map_update_elem() */
+ if (!OPTS_GET(gen->opts, gen_hash, false)) {
+ /* the jump over these instructions has 16-bit offset */
+ if (ptr_cnt > 10000) {
+ gen->error = -E2BIG;
+ return 0;
+ }
+ emit_copy_initial_value(gen, obj_map_idx, value, value_size, 3 * ptr_cnt);
+ /* the content that the host supplied doesn't have them */
+ for (i = 0; i < ptr_cnt; i++) {
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, value + ptr_offs[i]));
+ emit(gen, BPF_ST_MEM(BPF_DW, BPF_REG_1, 0, ptr_vals[i]));
+ }
+ }
+
+ attr_size = offsetofend(union bpf_attr, flags);
+ memset(&attr, 0, attr_size);
+ emit_map_update_elem(gen, map_idx, &attr, attr_size, key, value, value_size);
+
+ bpf_gen__map_freeze(gen, map_idx);
+ return map_idx;
+}
+
void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx)
{
int attr_size = offsetofend(union bpf_attr, map_fd);
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index cd1ea1bb53cbf..2e11808f7508d 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -544,6 +544,7 @@ struct bpf_struct_ops {
#define PERCPU_SEC ".percpu"
#define BSS_SEC ".bss"
#define RODATA_SEC ".rodata"
+#define DATA_REL_RO_SEC ".data.rel.ro"
#define KCONFIG_SEC ".kconfig"
#define KSYMS_SEC ".ksyms"
#define STRUCT_OPS_SEC ".struct_ops"
@@ -601,6 +602,9 @@ struct bpf_map {
bool autoattach;
__u64 map_extra;
struct bpf_program *excl_prog;
+ /* pointers to functions in the data of an internal map, see obj->func_ptrs */
+ struct func_ptr *func_ptrs;
+ size_t func_ptr_cnt;
};
enum extern_type {
@@ -780,6 +784,29 @@ struct bpf_object {
} *jumptable_maps;
size_t jumptable_map_cnt;
+ /*
+ * Pointers to functions found in read-only data sections: tables of
+ * functions, structures of operations, vtables. Sorted by section
+ * and offset.
+ */
+ struct func_ptr {
+ int sec_idx; /* ELF section that contains the pointer */
+ size_t sec_off; /* offset of the pointer in the section */
+ size_t text_off; /* offset of the function in .text section */
+ } *func_ptrs;
+ size_t func_ptr_cnt;
+
+ /*
+ * Read-only data with pointers to functions is different for every
+ * program that uses it, because so are the offsets of the functions.
+ */
+ struct {
+ struct bpf_program *prog;
+ int map_idx;
+ int fd;
+ } *func_ptr_maps;
+ size_t func_ptr_map_cnt;
+
struct kern_feature_cache *feat_cache;
char *token_path;
int token_fd;
@@ -845,6 +872,17 @@ static bool insn_is_pseudo_func(struct bpf_insn *insn)
return is_ldimm64_insn(insn) && insn->src_reg == BPF_PSEUDO_FUNC;
}
+/*
+ * ld_imm64 that loads the address of read-only data with pointers to functions
+ * is marked by bpf_object__relocate() before the code is relocated. Compilers
+ * leave src_reg of other ld_imm64 zero and it's set by
+ * bpf_object__relocate_data() later.
+ */
+static bool insn_is_func_ptrs_addr(struct bpf_insn *insn)
+{
+ return is_ldimm64_insn(insn) && insn->src_reg == BPF_PSEUDO_MAP_VALUE;
+}
+
static int
bpf_object__init_prog(struct bpf_object *obj, struct bpf_program *prog,
const char *name, size_t sec_idx, const char *sec_name,
@@ -4024,6 +4062,17 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
err = bpf_object__add_programs(obj, data, name, idx);
if (err)
return err;
+ } else if (strcmp(name, DATA_REL_RO_SEC) == 0 ||
+ str_has_pfx(name, DATA_REL_RO_SEC ".")) {
+ /*
+ * Constants with pointers in them, e.g. vtables,
+ * that position independent code keeps here to
+ * have them relocated. There is nothing that
+ * writes to it after that.
+ */
+ sec_desc->sec_type = SEC_RODATA;
+ sec_desc->shdr = sh;
+ sec_desc->data = data;
} else if (strcmp(name, DATA_SEC) == 0 ||
str_has_pfx(name, DATA_SEC ".")) {
sec_desc->sec_type = SEC_DATA;
@@ -4068,8 +4117,16 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
targ_sec_idx >= obj->efile.sec_cnt)
return -LIBBPF_ERRNO__FORMAT;
- /* Only do relo for section with exec instructions */
+ /*
+ * Only do relo for section with exec instructions,
+ * struct_ops, maps, and read-only data that might
+ * have pointers to functions.
+ */
if (!section_have_execinstr(obj, targ_sec_idx) &&
+ strcmp(name, ".rel" RODATA_SEC) &&
+ !str_has_pfx(name, ".rel" RODATA_SEC ".") &&
+ strcmp(name, ".rel" DATA_REL_RO_SEC) &&
+ !str_has_pfx(name, ".rel" DATA_REL_RO_SEC ".") &&
strcmp(name, ".rel" STRUCT_OPS_SEC) &&
strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
@@ -6471,6 +6528,168 @@ static int create_jt_map(struct bpf_object *obj, struct bpf_program *prog, struc
return err;
}
+/*
+ * The kernel recognizes a pointer to a function in a frozen read-only map by
+ * its value: the offset in bytes of the function in the program. It makes
+ * callx work for tables of functions, structures of operations and vtables,
+ * where pointers are mixed with other data. Functions have different offsets
+ * in different programs, so create a copy of the map for the program.
+ * The kernel replaces the offsets with the addresses of the functions when it
+ * loads the program, which has to be the only user of the map.
+ */
+static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog, int map_idx)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);
+ struct bpf_map *map = &obj->maps[map_idx];
+ __u32 value_size = map->def.value_size;
+ size_t i, j, cnt, sec_insn_off;
+ struct func_ptr *ptrs;
+ int map_fd, err, zero = 0;
+ __u64 val;
+ void *data, *tmp;
+
+ for (i = 0; i < obj->func_ptr_map_cnt; i++)
+ if (obj->func_ptr_maps[i].prog == prog &&
+ obj->func_ptr_maps[i].map_idx == map_idx)
+ return obj->func_ptr_maps[i].fd;
+
+ data = malloc(value_size);
+ if (!data)
+ return -ENOMEM;
+
+ /*
+ * The content of the map is final, it's frozen already. There is no map
+ * when light skeleton is generated, but there is what it's created with.
+ */
+ if (map->mmaped) {
+ memcpy(data, map->mmaped, value_size);
+ } else if (obj->gen_loader) {
+ err = -EINVAL;
+ goto err_free;
+ } else if (bpf_map_lookup_elem(map->fd, &zero, data)) {
+ err = -errno;
+ pr_warn("prog '%s': map '%s': failed to read the content: %s\n",
+ prog->name, map->name, errstr(err));
+ goto err_free;
+ }
+
+ ptrs = map->func_ptrs;
+ cnt = map->func_ptr_cnt;
+ for (i = 0; i < cnt; i++) {
+ if (ptrs[i].sec_off + sizeof(val) > value_size) {
+ err = -LIBBPF_ERRNO__FORMAT;
+ goto err_free;
+ }
+ /*
+ * Static functions were appended by bpf_object__append_func_ptrs_code().
+ * A global function is in the program only if the code refers to it.
+ */
+ sec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;
+ for (j = 0; j < prog->subprog_cnt; j++)
+ if (prog->subprogs[j].sec_insn_off == sec_insn_off)
+ break;
+ if (j == prog->subprog_cnt) {
+ pr_debug("prog '%s': map '%s': no function for the pointer at offset %zu, it's NULL\n",
+ prog->name, map->name, ptrs[i].sec_off);
+ val = 0;
+ } else {
+ val = (__u64)prog->subprogs[j].sub_insn_off * BPF_INSN_SZ;
+ }
+ /* light skeleton can be generated for a target of another endianness */
+ if (!is_native_endianness(obj))
+ val = bswap_64(val);
+ memcpy(data + ptrs[i].sec_off, &val, sizeof(val));
+ }
+
+ /*
+ * The kernel takes any aligned 64-bit value that is equal to the offset
+ * of a function for a pointer. Tell when it's going to get it wrong.
+ */
+ for (i = 0, j = 0; i + sizeof(val) <= value_size; i += sizeof(val)) {
+ __u32 k;
+
+ while (j < cnt && ptrs[j].sec_off < i)
+ j++;
+ if (j < cnt && ptrs[j].sec_off == i)
+ continue;
+ memcpy(&val, data + i, sizeof(val));
+ if (!is_native_endianness(obj))
+ val = bswap_64(val);
+ if (!val || val % BPF_INSN_SZ)
+ continue;
+ for (k = 0; k < prog->subprog_cnt; k++) {
+ if ((__u64)prog->subprogs[k].sub_insn_off * BPF_INSN_SZ != val)
+ continue;
+ pr_warn("prog '%s': map '%s': value %llu at offset %zu is the offset of a function, the kernel will treat it as a pointer to it\n",
+ prog->name, map->name, (unsigned long long)val, i);
+ break;
+ }
+ }
+
+ if (obj->gen_loader) {
+ __u32 *ptr_offs = calloc(cnt, sizeof(*ptr_offs));
+ __u64 *ptr_vals = calloc(cnt, sizeof(*ptr_vals));
+
+ if (!ptr_offs || !ptr_vals) {
+ free(ptr_offs);
+ free(ptr_vals);
+ err = -ENOMEM;
+ goto err_free;
+ }
+ for (i = 0; i < cnt; i++) {
+ ptr_offs[i] = ptrs[i].sec_off;
+ memcpy(&ptr_vals[i], data + ptrs[i].sec_off, sizeof(val));
+ if (!is_native_endianness(obj))
+ ptr_vals[i] = bswap_64(ptr_vals[i]);
+ }
+ /* it's an index in fd_array of the loader, not an fd */
+ map_fd = bpf_gen__func_ptr_map_create(obj->gen_loader, map->name, map_idx, data,
+ value_size, ptr_offs, ptr_vals, cnt);
+ free(ptr_offs);
+ free(ptr_vals);
+ goto done;
+ }
+
+ map_fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, map->name, sizeof(int), value_size, 1, &opts);
+ if (map_fd < 0) {
+ err = map_fd;
+ goto err_free;
+ }
+
+ err = bpf_map_update_elem(map_fd, &zero, data, 0);
+ if (!err)
+ err = bpf_map_freeze(map_fd);
+ if (err) {
+ err = -errno;
+ goto err_close;
+ }
+done:
+
+ tmp = libbpf_reallocarray(obj->func_ptr_maps, obj->func_ptr_map_cnt + 1,
+ sizeof(*obj->func_ptr_maps));
+ if (!tmp) {
+ err = -ENOMEM;
+ goto err_close;
+ }
+ obj->func_ptr_maps = tmp;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].prog = prog;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].map_idx = map_idx;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].fd = map_fd;
+ obj->func_ptr_map_cnt++;
+
+ pr_debug("prog '%s': created a copy of map '%s' with %zu pointers to functions\n",
+ prog->name, map->name, cnt);
+ free(data);
+ return map_fd;
+
+err_close:
+ if (!obj->gen_loader)
+ close(map_fd);
+err_free:
+ free(data);
+ return err;
+}
+
/* Relocate data references within program code:
* - map references;
* - global variable references;
@@ -6508,7 +6727,20 @@ bpf_object__relocate_data(struct bpf_object *obj, struct bpf_program *prog)
if (relo->map_idx == obj->arena_map_idx)
insn[1].imm += obj->arena_data_off;
- if (obj->gen_loader) {
+ if (map->autocreate && map->func_ptr_cnt) {
+ int map_fd;
+
+ /* the program gets its own map with pointers to its functions */
+ map_fd = create_func_ptr_map(obj, prog, relo->map_idx);
+ if (map_fd < 0) {
+ pr_warn("prog '%s': relo #%d: can't create a copy of map '%s' with pointers to functions\n",
+ prog->name, i, map->name);
+ return map_fd;
+ }
+ insn[0].src_reg = obj->gen_loader ? BPF_PSEUDO_MAP_IDX_VALUE :
+ BPF_PSEUDO_MAP_VALUE;
+ insn[0].imm = map_fd;
+ } else if (obj->gen_loader) {
insn[0].src_reg = BPF_PSEUDO_MAP_IDX_VALUE;
insn[0].imm = relo->map_idx;
} else if (map->autocreate) {
@@ -6835,6 +7067,51 @@ bpf_object__append_subprog_code(struct bpf_object *obj, struct bpf_program *main
return 0;
}
+static int
+bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
+ struct bpf_program *prog);
+
+/* Append to the main program all functions that the data of the map points to */
+static int
+bpf_object__append_func_ptrs_code(struct bpf_object *obj, struct bpf_program *main_prog,
+ const struct bpf_map *map)
+{
+ struct bpf_program *subprog;
+ size_t i, cnt, sec_insn_off;
+ struct func_ptr *ptrs;
+ int err;
+
+ ptrs = map->func_ptrs;
+ cnt = map->func_ptr_cnt;
+ for (i = 0; i < cnt; i++) {
+ sec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;
+ subprog = find_prog_by_sec_insn(obj, obj->efile.text_shndx, sec_insn_off);
+ if (!subprog || subprog->sec_insn_off != sec_insn_off) {
+ pr_warn("prog '%s': map '%s': no function at .text+%zu for the pointer at offset %zu\n",
+ main_prog->name, map->name, ptrs[i].text_off, ptrs[i].sec_off);
+ return -LIBBPF_ERRNO__RELOC;
+ }
+
+ /*
+ * callx can't call global functions. Don't add one to the
+ * program only because the data points to it.
+ */
+ if (subprog->sym_global)
+ continue;
+
+ /* see the comment in bpf_object__reloc_code() */
+ if (subprog->sub_insn_off == 0) {
+ err = bpf_object__append_subprog_code(obj, main_prog, subprog);
+ if (err)
+ return err;
+ err = bpf_object__reloc_code(obj, main_prog, subprog);
+ if (err)
+ return err;
+ }
+ }
+ return 0;
+}
+
static int
bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
struct bpf_program *prog)
@@ -6851,6 +7128,20 @@ bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
for (insn_idx = 0; insn_idx < prog->sec_insn_cnt; insn_idx++) {
insn = &main_prog->insns[prog->sub_insn_off + insn_idx];
+ if (insn_is_func_ptrs_addr(insn)) {
+ /*
+ * The code that loads the address of the data might
+ * call any function that the data points to.
+ */
+ relo = find_prog_insn_relo(prog, insn_idx);
+ if (relo && relo->type == RELO_DATA) {
+ err = bpf_object__append_func_ptrs_code(obj, main_prog,
+ &obj->maps[relo->map_idx]);
+ if (err)
+ return err;
+ }
+ continue;
+ }
if (!insn_is_subprog_call(insn) && !insn_is_pseudo_func(insn))
continue;
@@ -7541,6 +7832,9 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
/* mark the insn, so it's recognized by insn_is_pseudo_func() */
if (relo->type == RELO_SUBPROG_ADDR)
insn[0].src_reg = BPF_PSEUDO_FUNC;
+ /* and by insn_is_func_ptrs_addr() */
+ if (relo->type == RELO_DATA && obj->maps[relo->map_idx].func_ptr_cnt)
+ insn[0].src_reg = BPF_PSEUDO_MAP_VALUE;
}
}
@@ -7757,6 +8051,98 @@ static int bpf_object__collect_map_relos(struct bpf_object *obj,
return 0;
}
+/*
+ * Collect pointers to functions in a read-only data section. They are
+ * R_BPF_64_ABS64 relocations against .text section, where the offset of
+ * a static function in the section is stored in place. Relocations in data
+ * sections were ignored before pointers to functions were supported. Those
+ * that are something else, e.g. pointers to data, still are.
+ */
+static int bpf_object__collect_rodata_relos(struct bpf_object *obj,
+ Elf64_Shdr *shdr, Elf_Data *data)
+{
+ size_t sec_idx = shdr->sh_info, sym_idx;
+ int i, nrels = shdr->sh_size / shdr->sh_entsize;
+ const char *relo_sec_name;
+ struct func_ptr *ptrs;
+ Elf_Data *scn_data;
+ Elf64_Sym *sym;
+ Elf64_Rel *rel;
+ __u64 addend;
+
+ relo_sec_name = elf_sec_str(obj, shdr->sh_name) ?: "<?>";
+ scn_data = obj->efile.secs[sec_idx].data;
+ if (!scn_data)
+ return -LIBBPF_ERRNO__FORMAT;
+
+ for (i = 0; i < nrels; i++) {
+ rel = elf_rel_by_idx(data, i);
+ if (!rel) {
+ pr_warn("sec '%s': failed to get relo #%d\n", relo_sec_name, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ sym_idx = ELF64_R_SYM(rel->r_info);
+ sym = elf_sym_by_idx(obj, sym_idx);
+ if (!sym) {
+ pr_warn("sec '%s': symbol #%zu not found for relo #%d\n",
+ relo_sec_name, sym_idx, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ if (ELF64_R_TYPE(rel->r_info) != R_BPF_64_ABS64 ||
+ !sym_is_subprog(sym, obj->efile.text_shndx)) {
+ pr_debug("sec '%s': relo #%d: not a pointer to a function, skipping...\n",
+ relo_sec_name, i);
+ continue;
+ }
+
+ /* the kernel finds aligned pointers only */
+ if (rel->r_offset % sizeof(__u64) || rel->r_offset >= scn_data->d_size ||
+ scn_data->d_size - rel->r_offset < sizeof(__u64)) {
+ pr_debug("sec '%s': relo #%d: unsupported offset 0x%zx, skipping...\n",
+ relo_sec_name, i, (size_t)rel->r_offset);
+ continue;
+ }
+
+ memcpy(&addend, scn_data->d_buf + rel->r_offset, sizeof(addend));
+ if (!is_native_endianness(obj))
+ addend = bswap_64(addend);
+ if ((sym->st_value + addend) % BPF_INSN_SZ) {
+ pr_debug("sec '%s': relo #%d: bad pointer to a function at offset %zu+%llu, skipping...\n",
+ relo_sec_name, i, (size_t)sym->st_value,
+ (unsigned long long)addend);
+ continue;
+ }
+
+ ptrs = libbpf_reallocarray(obj->func_ptrs, obj->func_ptr_cnt + 1, sizeof(*ptrs));
+ if (!ptrs)
+ return -ENOMEM;
+ obj->func_ptrs = ptrs;
+
+ ptrs[obj->func_ptr_cnt].sec_idx = sec_idx;
+ ptrs[obj->func_ptr_cnt].sec_off = rel->r_offset;
+ ptrs[obj->func_ptr_cnt].text_off = sym->st_value + addend;
+ obj->func_ptr_cnt++;
+
+ pr_debug("sec '%s': relo #%d: pointer at offset %zu to a function at .text+%zu\n",
+ relo_sec_name, i, (size_t)rel->r_offset, (size_t)(sym->st_value + addend));
+ }
+ return 0;
+}
+
+static int cmp_func_ptrs(const void *_a, const void *_b)
+{
+ const struct func_ptr *a = _a;
+ const struct func_ptr *b = _b;
+
+ if (a->sec_idx != b->sec_idx)
+ return a->sec_idx < b->sec_idx ? -1 : 1;
+ if (a->sec_off != b->sec_off)
+ return a->sec_off < b->sec_off ? -1 : 1;
+ return 0;
+}
+
static int bpf_object__collect_relos(struct bpf_object *obj)
{
int i, err;
@@ -7779,7 +8165,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return -LIBBPF_ERRNO__INTERNAL;
}
- if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
+ if (obj->efile.secs[idx].sec_type == SEC_RODATA)
+ err = bpf_object__collect_rodata_relos(obj, shdr, data);
+ else if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
err = bpf_object__collect_st_ops_relos(obj, shdr, data);
else if (idx == obj->efile.btf_maps_shndx)
err = bpf_object__collect_map_relos(obj, shdr, data);
@@ -7789,6 +8177,25 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return err;
}
+ /* sort by section, so that pointers in the data of a map are next to each other */
+ if (obj->func_ptr_cnt)
+ qsort(obj->func_ptrs, obj->func_ptr_cnt, sizeof(*obj->func_ptrs), cmp_func_ptrs);
+
+ for (i = 0; i < obj->nr_maps; i++) {
+ struct bpf_map *map = &obj->maps[i];
+ size_t j;
+
+ if (map->libbpf_type != LIBBPF_MAP_RODATA)
+ continue;
+ for (j = 0; j < obj->func_ptr_cnt; j++) {
+ if (obj->func_ptrs[j].sec_idx != map->sec_idx)
+ continue;
+ if (!map->func_ptr_cnt)
+ map->func_ptrs = &obj->func_ptrs[j];
+ map->func_ptr_cnt++;
+ }
+ }
+
bpf_object__sort_relos(obj);
return 0;
}
@@ -9208,8 +9615,17 @@ static int bpf_object_load(struct bpf_object *obj, int extra_log_level, const ch
* permit cross-endian creation of "light skeleton".
*/
if (obj->gen_loader) {
+ int nr_func_ptr_maps = 0, nr_progs = 0, i;
+
+ /* every program may get a copy of every map with pointers to functions */
+ for (i = 0; i < obj->nr_maps; i++)
+ if (obj->maps[i].autocreate && obj->maps[i].func_ptr_cnt)
+ nr_func_ptr_maps++;
+ for (i = 0; i < obj->nr_programs; i++)
+ if (obj->programs[i].autoload && !prog_is_subprog(obj, &obj->programs[i]))
+ nr_progs++;
bpf_gen__init(obj->gen_loader, obj->log_level | extra_log_level,
- obj->nr_programs, obj->nr_maps);
+ obj->nr_programs, obj->nr_maps, nr_func_ptr_maps * nr_progs);
} else if (!is_native_endianness(obj)) {
pr_warn("object '%s': loading non-native endianness is unsupported\n", obj->name);
return libbpf_err(-LIBBPF_ERRNO__ENDIAN);
@@ -9749,6 +10165,12 @@ void bpf_object__close(struct bpf_object *obj)
close(obj->jumptable_maps[i].fd);
zfree(&obj->jumptable_maps);
+ for (i = 0; i < obj->func_ptr_map_cnt; i++)
+ if (!obj->gen_loader)
+ close(obj->func_ptr_maps[i].fd);
+ zfree(&obj->func_ptr_maps);
+ zfree(&obj->func_ptrs);
+
if (obj->btf_module_allowlist) {
for (i = 0; i < obj->btf_module_allowlist_cnt; i++)
zfree(&obj->btf_module_allowlist[i]);
diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index 78f92c39290af..53f64a1a1f25a 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -2274,6 +2274,24 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
insn->imm += sec->dst_off / sizeof(struct bpf_insn);
else
insn->imm += sec->dst_off;
+ } else if (sym_type == R_BPF_64_ABS64 &&
+ (sec->shdr->sh_flags & SHF_EXECINSTR)) {
+ /*
+ * A pointer to a static function in a data section,
+ * which is stored in place as an offset of the
+ * function in its section. Data sections are kept
+ * in the byte order of the object.
+ */
+ void *ptr = dst_linked_sec->raw_data + dst_rel->r_offset;
+ __u64 off;
+
+ memcpy(&off, ptr, sizeof(off));
+ if (linker->swapped_endian)
+ off = bswap_64(off);
+ off += sec->dst_off;
+ if (linker->swapped_endian)
+ off = bswap_64(off);
+ memcpy(ptr, &off, sizeof(off));
} else {
pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
return -EINVAL;
diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
index 580d1d82c1867..3d92cdca62ed8 100644
--- a/tools/testing/selftests/bpf/Makefile.skel
+++ b/tools/testing/selftests/bpf/Makefile.skel
@@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
core_kern.c core_kern_overflow.c test_ringbuf.c \
test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
- test_ringbuf_overwrite.c
+ test_ringbuf_overwrite.c callx_rodata.c
LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
diff --git a/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
new file mode 100644
index 0000000000000..19f96792df5db
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
@@ -0,0 +1,192 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * The kernel replaces the offsets of functions in a frozen read-only map with
+ * their addresses when the program is loaded. The program has to be the only
+ * user of the map, and no other program can use the map after that.
+ */
+#include <test_progs.h>
+#include <linux/filter.h>
+#include <bpf/btf.h>
+
+#if defined(__x86_64__) || defined(__aarch64__)
+
+#define CALLEE_INSN 6
+#define DATA 0x1234
+
+/*
+ * main: r2 = &value; r2 = *(u64 *)(r2 + 8); r1 = 10; callx r2; exit
+ * add1: r0 = r1; r0 += 1; exit
+ *
+ * where value is { DATA, offset of add1 in the program }.
+ */
+static const struct bpf_insn callx_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_LDX_MEM(BPF_DW, BPF_REG_2, BPF_REG_2, 8),
+ BPF_MOV64_IMM(BPF_REG_1, 10),
+ BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, BPF_REG_2, 0, 0, 0),
+ BPF_EXIT_INSN(),
+ BPF_MOV64_REG(BPF_REG_0, BPF_REG_1),
+ BPF_ALU64_IMM(BPF_ADD, BPF_REG_0, 1),
+ BPF_EXIT_INSN(),
+};
+
+/* reads the data of the map */
+static const struct bpf_insn reader_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_LDX_MEM(BPF_DW, BPF_REG_0, BPF_REG_2, 0),
+ BPF_EXIT_INSN(),
+};
+
+/* refers to the map and is rejected: r0 is not set */
+static const struct bpf_insn bad_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_EXIT_INSN(),
+};
+
+static char log_buf[16 * 1024];
+static struct bpf_func_info func_info[2];
+static int btf_fd;
+
+static int create_map(void)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);
+ __u64 value[2] = { DATA, CALLEE_INSN * sizeof(struct bpf_insn) };
+ int fd, zero = 0;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, "callx_rodata", sizeof(int), sizeof(value), 1,
+ &opts);
+ if (!ASSERT_OK_FD(fd, "map_create"))
+ return -1;
+ if (!ASSERT_OK(bpf_map_update_elem(fd, &zero, value, 0), "map_update") ||
+ !ASSERT_OK(bpf_map_freeze(fd), "map_freeze")) {
+ close(fd);
+ return -1;
+ }
+ return fd;
+}
+
+static int load(const struct bpf_insn *prog_insns, int cnt, int map_fd)
+{
+ LIBBPF_OPTS(bpf_prog_load_opts, opts,
+ .log_buf = log_buf,
+ .log_size = sizeof(log_buf),
+ .log_level = 1,
+ );
+ struct bpf_insn insns[ARRAY_SIZE(callx_insns)];
+
+ if (prog_insns == callx_insns) {
+ /* add1() is referred to by the data only and is found through func_info */
+ opts.prog_btf_fd = btf_fd;
+ opts.func_info = func_info;
+ opts.func_info_cnt = 2;
+ opts.func_info_rec_size = sizeof(func_info[0]);
+ }
+ memcpy(insns, prog_insns, cnt * sizeof(insns[0]));
+ insns[0].imm = map_fd;
+ log_buf[0] = 0;
+ return bpf_prog_load(BPF_PROG_TYPE_SOCKET_FILTER, "callx_map", "GPL", insns, cnt, &opts);
+}
+
+#define LOAD(insns, map_fd) load(insns, ARRAY_SIZE(insns), map_fd)
+
+static void run(int prog_fd, int expected, const char *name)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, topts);
+ char pkt[64] = {};
+
+ topts.data_in = pkt;
+ topts.data_size_in = sizeof(pkt);
+ if (ASSERT_OK(bpf_prog_test_run_opts(prog_fd, &topts), name))
+ ASSERT_EQ(topts.retval, expected, name);
+}
+
+void test_callx_func_ptr_map(void)
+{
+ int int_id, proto_id, map_fd = -1, prog_fd = -1, reader_fd = -1, fd, zero = 0;
+ struct btf *btf;
+ __u64 value[2];
+
+ btf = btf__new_empty();
+ if (!ASSERT_OK_PTR(btf, "btf_new"))
+ return;
+ int_id = btf__add_int(btf, "int", 4, BTF_INT_SIGNED);
+ proto_id = btf__add_func_proto(btf, int_id);
+ func_info[0].insn_off = 0;
+ func_info[0].type_id = btf__add_func(btf, "main_prog", BTF_FUNC_GLOBAL, proto_id);
+ func_info[1].insn_off = CALLEE_INSN;
+ func_info[1].type_id = btf__add_func(btf, "add1", BTF_FUNC_STATIC, proto_id);
+ if (!ASSERT_GT(func_info[1].type_id, 0, "btf_add_func") ||
+ !ASSERT_OK(btf__load_into_kernel(btf), "btf_load"))
+ goto out;
+ btf_fd = btf__fd(btf);
+
+ /*
+ * Another program relies on what the map has: it's verified with
+ * the data folded into constants. Pointers to functions are not looked
+ * for in such map.
+ */
+ map_fd = create_map();
+ if (map_fd < 0)
+ goto out;
+ reader_fd = LOAD(reader_insns, map_fd);
+ if (!ASSERT_OK_FD(reader_fd, "load_reader"))
+ goto out;
+ fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_LT(fd, 0, "load_shared"))
+ close(fd);
+ ASSERT_HAS_SUBSTR(log_buf, "unreachable insn 6", "log_shared");
+ run(reader_fd, DATA, "run_reader");
+ close(reader_fd);
+ reader_fd = -1;
+ close(map_fd);
+
+ map_fd = create_map();
+ if (map_fd < 0)
+ goto out;
+
+ /* a program that is rejected is not a user, libbpf loads it again to get the log */
+ fd = LOAD(bad_insns, map_fd);
+ if (!ASSERT_LT(fd, 0, "load_bad"))
+ close(fd);
+
+ prog_fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_OK_FD(prog_fd, "load_callx")) {
+ printf("%s\n", log_buf);
+ goto out;
+ }
+ run(prog_fd, 11, "run_callx");
+
+ /* the offset of the function is gone from the map, the data is intact */
+ if (ASSERT_OK(bpf_map_lookup_elem(map_fd, &zero, value), "map_lookup")) {
+ ASSERT_EQ(value[0], DATA, "data");
+ ASSERT_NEQ(value[1], CALLEE_INSN * sizeof(struct bpf_insn), "pointer");
+ }
+
+ /* no other program can use the map now, another instance of the same one too */
+ fd = LOAD(reader_insns, map_fd);
+ if (!ASSERT_EQ(fd, -EBUSY, "load_reader_after"))
+ close(fd);
+ ASSERT_HAS_SUBSTR(log_buf, "has addresses of functions of another program", "log_reader");
+ fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_EQ(fd, -EBUSY, "load_second_instance"))
+ close(fd);
+
+ run(prog_fd, 11, "run_callx_again");
+out:
+ if (reader_fd >= 0)
+ close(reader_fd);
+ if (prog_fd >= 0)
+ close(prog_fd);
+ if (map_fd >= 0)
+ close(map_fd);
+ btf__free(btf);
+}
+
+#else
+
+void test_callx_func_ptr_map(void)
+{
+ test__skip();
+}
+
+#endif
diff --git a/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
new file mode 100644
index 0000000000000..5c60f352ff446
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
@@ -0,0 +1,55 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <test_progs.h>
+
+#include "callx_rodata.lskel.h"
+
+#if defined(__x86_64__) || defined(__aarch64__)
+
+static void run(int prog_fd, int expected, const char *name)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, topts);
+ char pkt[64] = {};
+
+ topts.data_in = pkt;
+ topts.data_size_in = sizeof(pkt);
+ if (ASSERT_OK(bpf_prog_test_run_opts(prog_fd, &topts), name))
+ ASSERT_EQ(topts.retval, expected, name);
+}
+
+/*
+ * Every program gets its own copy of .rodata with the offsets of its functions.
+ * The copies have what user space puts into .rodata before the load.
+ */
+void test_callx_rodata_lskel(void)
+{
+ struct callx_rodata_lskel *skel;
+
+ skel = callx_rodata_lskel__open();
+ if (!ASSERT_OK_PTR(skel, "open"))
+ return;
+
+ skel->rodata->bias = 7;
+
+ if (!ASSERT_OK(callx_rodata_lskel__load(skel), "load"))
+ goto out;
+
+ skel->bss->op_idx = 0;
+ run(skel->progs.select_op.prog_fd, 17, "add_bias");
+ skel->bss->op_idx = 1;
+ run(skel->progs.select_op.prog_fd, 30, "mul3");
+ skel->bss->op_idx = 2;
+ run(skel->progs.select_op.prog_fd, -1, "out_of_range");
+ /* mul3(add_bias(4)) */
+ run(skel->progs.both_ops.prog_fd, 33, "both_ops");
+out:
+ callx_rodata_lskel__destroy(skel);
+}
+
+#else
+
+void test_callx_rodata_lskel(void)
+{
+ test__skip();
+}
+
+#endif
diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 4f1e1c1cd5ab3..3e07b957ee80d 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -27,6 +27,8 @@
#include "verifier_btf_ctx_access.skel.h"
#include "verifier_btf_unreliable_prog.skel.h"
#include "verifier_call_large_imm.skel.h"
+#include "verifier_callx.skel.h"
+#include "verifier_callx_rodata.skel.h"
#include "verifier_cfg.skel.h"
#include "verifier_cgroup_inv_retcode.skel.h"
#include "verifier_cgroup_skb.skel.h"
@@ -196,6 +198,8 @@ void test_verifier_bswap(void) { RUN(verifier_bswap); }
void test_verifier_btf_ctx_access(void) { RUN(verifier_btf_ctx_access); }
void test_verifier_btf_unreliable_prog(void) { RUN(verifier_btf_unreliable_prog); }
void test_verifier_call_large_imm(void) { RUN(verifier_call_large_imm); }
+void test_verifier_callx(void) { RUN(verifier_callx); }
+void test_verifier_callx_rodata(void) { RUN(verifier_callx_rodata); }
void test_verifier_cfg(void) { RUN(verifier_cfg); }
void test_verifier_cgroup_inv_retcode(void) { RUN(verifier_cgroup_inv_retcode); }
void test_verifier_cgroup_skb(void) { RUN(verifier_cgroup_skb); }
diff --git a/tools/testing/selftests/bpf/progs/callx_rodata.c b/tools/testing/selftests/bpf/progs/callx_rodata.c
new file mode 100644
index 0000000000000..7475c87cdaed4
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/callx_rodata.c
@@ -0,0 +1,43 @@
+// SPDX-License-Identifier: GPL-2.0
+/* callx through pointers to functions in .rodata, loaded by light skeleton */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+
+typedef int (*op_fn)(int);
+
+/* set by user space before the programs are loaded, it's in .rodata too */
+const volatile int bias = 1;
+
+int op_idx;
+
+static __noinline int add_bias(int x)
+{
+ return x + bias;
+}
+
+static __noinline int mul3(int x)
+{
+ return x * 3;
+}
+
+static op_fn const ops[] = { add_bias, mul3 };
+
+SEC("socket")
+int select_op(void *ctx)
+{
+ unsigned int i = op_idx;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ return ops[i](10);
+}
+
+/* functions have other offsets in this program */
+SEC("socket")
+int both_ops(void *ctx)
+{
+ return ops[1](ops[0](4));
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_callx.c b/tools/testing/selftests/bpf/progs/verifier_callx.c
new file mode 100644
index 0000000000000..ec63d2107e6c9
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_callx.c
@@ -0,0 +1,984 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Tests for callx: indirect calls of bpf subprogs */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../../../include/linux/filter.h"
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+#define CALLX_INSN(DST, SRC, OFF, IMM) \
+ BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, DST, SRC, OFF, IMM)
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, int);
+ __type(value, long long);
+} map_array SEC(".maps");
+
+struct val_with_lock {
+ struct bpf_spin_lock lock;
+ int cnt;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, int);
+ __type(value, struct val_with_lock);
+} map_lock SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(int));
+ __uint(value_size, sizeof(int));
+} map_prog SEC(".maps");
+
+__naked __noinline __used
+static unsigned long add1(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long add2(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 2;"
+ "exit;"
+ );
+}
+
+/* apply(fn, x) { return fn(x); } */
+__naked __noinline __used
+static unsigned long apply(void)
+{
+ asm volatile (
+ "r3 = r1;"
+ "r1 = r2;"
+ "callx r3;"
+ "exit;"
+ );
+}
+
+SEC("socket")
+__success __retval(6)
+__naked void callx_basic(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* callx is printed by the verifier log and xlated dump */
+SEC("socket")
+__success __log_level(2)
+__msg("(8d) callx r2")
+__xlated("callx r2")
+__naked void callx_disasm(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* different callees are called by the same callx on different paths */
+SEC("socket")
+__success __retval(11)
+__naked void callx_two_callees(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r2 = %[add1] ll;"
+ "if r6 == 0 goto +2;"
+ "r2 = %[add2] ll;"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add1),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+/* pointer to a function is passed as an argument */
+SEC("socket")
+__success __retval(45)
+__naked void callx_fn_as_arg(void)
+{
+ asm volatile (
+ "r1 = %[add2] ll;"
+ "r2 = 40;"
+ "call apply;"
+ "r6 = r0;"
+ "r1 = %[add1] ll;"
+ "r2 = 2;"
+ "call apply;"
+ "r0 += r6;"
+ "exit;"
+ :
+ : __imm_addr(add1),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long get_add1(void)
+{
+ asm volatile (
+ "r0 = %[add1] ll;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* pointer to a function is returned from a subprog and called via r0 */
+SEC("socket")
+__success __retval(2)
+__naked void callx_r0(void)
+{
+ asm volatile (
+ "call get_add1;"
+ "r1 = 1;"
+ "callx r0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* pointer to a function survives spill/fill */
+SEC("socket")
+__success __retval(9)
+__naked void callx_spill_fill(void)
+{
+ asm volatile (
+ "r2 = %[add2] ll;"
+ "*(u64 *)(r10 - 8) = r2;"
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 7;"
+ "r9 = *(u64 *)(r10 - 8);"
+ "callx r9;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long clobber_callee_saved(void)
+{
+ asm volatile (
+ "r6 = 100;"
+ "r7 = 100;"
+ "r8 = 100;"
+ "r9 = 100;"
+ "r0 = r1;"
+ "exit;"
+ );
+}
+
+/* r6-r9 are preserved across callx */
+SEC("socket")
+__success __retval(11)
+__naked void callx_callee_saved_regs(void)
+{
+ asm volatile (
+ "r6 = 1;"
+ "r7 = 2;"
+ "r8 = 3;"
+ "r9 = 4;"
+ "r1 = 1;"
+ "r2 = %[clobber_callee_saved] ll;"
+ "callx r2;"
+ "r0 += r6;"
+ "r0 += r7;"
+ "r0 += r8;"
+ "r0 += r9;"
+ "exit;"
+ :
+ : __imm_addr(clobber_callee_saved)
+ : __clobber_all);
+}
+
+/* r1-r5 are scratched by callx */
+SEC("socket")
+__failure __msg("R1 !read_ok")
+__naked void callx_scratches_args(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "r0 = r1;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* return value of the callee is tracked */
+SEC("socket")
+__success __log_level(2)
+__msg("R0=6")
+__naked void callx_retval_is_tracked(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long write42(void)
+{
+ asm volatile (
+ "r2 = 42;"
+ "*(u64 *)(r1 + 0) = r2;"
+ "r0 = 0;"
+ "exit;"
+ );
+}
+
+/* callee writes into the stack of the caller */
+SEC("socket")
+__success __retval(42)
+__naked void callx_callee_writes_caller_stack(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[write42] ll;"
+ "callx r2;"
+ "r0 = *(u64 *)(r10 - 8);"
+ "exit;"
+ :
+ : __imm_addr(write42)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R1 has type scalar, expected func")
+__naked void callx_scalar(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "callx r1;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R2 !read_ok")
+__naked void callx_uninit_reg(void)
+{
+ asm volatile (
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R10 has type fp, expected func")
+__naked void callx_fp(void)
+{
+ asm volatile (
+ "callx r10;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R1 has type map_value, expected func")
+__naked void callx_map_value(void)
+{
+ asm volatile (
+ "r1 = %[map_array] ll;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r3 = 0;"
+ "*(u32 *)(r2 + 0) = r3;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 == 0 goto 1f;"
+ "r1 = r0;"
+ "callx r1;"
+ "1:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm_addr(map_array)
+ : __clobber_all);
+}
+
+/* the address of a subprog can't be modified before the call */
+SEC("socket")
+__failure __msg("dereference of modified func ptr R2 off=8 disallowed")
+__naked void callx_modified_ptr(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "r2 += 8;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("variable func access var_off=")
+__naked void callx_variable_ptr(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 8;"
+ "r2 = %[add1] ll;"
+ "r2 += r0;"
+ "r1 = 5;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add1)
+ : __clobber_all);
+}
+
+__noinline __used
+int global_add3(int x)
+{
+ return x + 3;
+}
+
+/* only static subprogs can be called via callx */
+SEC("socket")
+__failure __msg("callback function not static")
+__naked void callx_global_func(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[global_add3] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(global_add3)
+ : __clobber_all);
+}
+
+#define DEFINE_CALLX_RESERVED_FIELDS_PROG(NAME, SRC_REG, OFF, IMM) \
+ SEC("socket") \
+ __failure __msg("BPF_CALL|BPF_X uses reserved fields") \
+ __naked void callx_reserved_field_ ## NAME(void) \
+ { \
+ asm volatile ( \
+ "r1 = 5;" \
+ "r2 = %[add1] ll;" \
+ ".8byte %[callx_r2];" \
+ "exit;" \
+ : \
+ : __imm_addr(add1), \
+ __imm_insn(callx_r2, CALLX_INSN(BPF_REG_2, (SRC_REG), (OFF), (IMM))) \
+ : __clobber_all); \
+ }
+
+DEFINE_CALLX_RESERVED_FIELDS_PROG(src_reg, BPF_REG_1, 0, 0)
+DEFINE_CALLX_RESERVED_FIELDS_PROG(off, BPF_REG_0, 1, 0)
+DEFINE_CALLX_RESERVED_FIELDS_PROG(imm, BPF_REG_0, 0, 1)
+
+SEC("socket")
+__failure __msg("unknown opcode 8e")
+__naked void callx_jmp32(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ ".8byte %[callx32_r2];"
+ "exit;"
+ :
+ : __imm_addr(add1),
+ __imm_insn(callx32_r2,
+ BPF_RAW_INSN(BPF_JMP32 | BPF_CALL | BPF_X, BPF_REG_2, 0, 0, 0))
+ : __clobber_all);
+}
+
+/* similar to calls of static subprogs callx is allowed under a lock */
+SEC("tc")
+__success __retval(3)
+__naked void callx_under_lock(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_lock] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 != 0 goto 1f;"
+ "exit;"
+ "1:"
+ "r6 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_lock];"
+ "r1 = 1;"
+ "r2 = %[add2] ll;"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_unlock];"
+ "r0 = r7;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm(bpf_spin_lock),
+ __imm(bpf_spin_unlock),
+ __imm_addr(map_lock),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+/* helpers are still not allowed under a lock in the callee */
+__naked __noinline __used
+static unsigned long call_helper(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+SEC("tc")
+__failure __msg("function calls are not allowed while holding a lock")
+__naked void callx_helper_under_lock(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_lock] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 != 0 goto 1f;"
+ "exit;"
+ "1:"
+ "r6 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_lock];"
+ "r2 = %[call_helper] ll;"
+ "callx r2;"
+ "r1 = r6;"
+ "call %[bpf_spin_unlock];"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm(bpf_spin_lock),
+ __imm(bpf_spin_unlock),
+ __imm_addr(map_lock),
+ __imm_addr(call_helper)
+ : __clobber_all);
+}
+
+/* self(fn) { return fn(fn); } */
+__naked __noinline __used
+static unsigned long self(void)
+{
+ asm volatile (
+ "callx r1;"
+ "exit;"
+ );
+}
+
+/* unbounded recursion is caught by the main verification pass */
+SEC("socket")
+__failure __msg("frames is too deep")
+__naked void callx_unbounded_recursion(void)
+{
+ asm volatile (
+ "r1 = %[self] ll;"
+ "call self;"
+ "exit;"
+ :
+ : __imm_addr(self)
+ : __clobber_all);
+}
+
+/* countdown(fn, n) { return n ? fn(fn, n - 1) : 0; } */
+__naked __noinline __used
+static unsigned long countdown(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r2 == 0 goto 1f;"
+ "r2 += -1;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+/*
+ * The depth of the recursion is known to the main verification pass,
+ * but recursive calls are not allowed regardless.
+ */
+SEC("socket")
+__failure __msg("recursive call from countdown() to countdown()")
+__naked void callx_bounded_recursion(void)
+{
+ asm volatile (
+ "r1 = %[countdown] ll;"
+ "r2 = 2;"
+ "call countdown;"
+ "exit;"
+ :
+ : __imm_addr(countdown)
+ : __clobber_all);
+}
+
+/* ping(fn1, fn2, n) { return n ? fn2(fn2, fn1, n - 1) : 0; } */
+__naked __noinline __used
+static unsigned long ping(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r3 == 0 goto 1f;"
+ "r3 += -1;"
+ "r4 = r1;"
+ "r1 = r2;"
+ "r2 = r4;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long pong(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r3 == 0 goto 1f;"
+ "r3 += -1;"
+ "r4 = r1;"
+ "r1 = r2;"
+ "r2 = r4;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+SEC("socket")
+__failure __msg("recursive call from")
+__naked void callx_mutual_recursion(void)
+{
+ asm volatile (
+ "r1 = %[ping] ll;"
+ "r2 = %[pong] ll;"
+ "r3 = 3;"
+ "call ping;"
+ "exit;"
+ :
+ : __imm_addr(ping),
+ __imm_addr(pong)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long use_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "exit;"
+ );
+}
+
+/* stack of the callee of callx is accounted */
+SEC("socket")
+__failure __msg("combined stack size of 2 calls is")
+__naked void callx_stack_depth(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "r2 = %[use_stack_304] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+/* apply_stack_304(fn) { char buf[304]; return fn(); } */
+__naked __noinline __used
+static unsigned long apply_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "callx r1;"
+ "exit;"
+ );
+}
+
+/*
+ * The address of use_stack_304() is taken by the main prog that doesn't
+ * use stack, but it is called from apply_stack_304().
+ */
+SEC("socket")
+__failure __msg("combined stack size of 3 calls is")
+__naked void callx_stack_depth_nested(void)
+{
+ asm volatile (
+ "r1 = %[use_stack_304] ll;"
+ "call apply_stack_304;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+SEC("socket")
+__success __retval(0)
+__naked void callx_stack_depth_ok(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 200) = r0;"
+ "r2 = %[use_stack_304] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long do_tail_call(void)
+{
+ asm volatile (
+ "r2 = %[map_prog] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_tail_call),
+ __imm_addr(map_prog)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("tail_calls are not allowed in functions called via callx")
+__naked void callx_tail_call_in_callee(void)
+{
+ asm volatile (
+ "r2 = %[do_tail_call] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(do_tail_call)
+ : __clobber_all);
+}
+
+/* tail call in the caller of callx is fine */
+SEC("socket")
+__success __retval(3)
+__naked void callx_tail_call_in_caller(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "r1 = 2;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = r6;"
+ "r2 = %[map_prog] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "r0 = r7;"
+ "exit;"
+ :
+ : __imm(bpf_tail_call),
+ __imm_addr(map_prog),
+ __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* read_idx(p) { return ((char *)map_value)[*p]; } */
+__naked __noinline __used
+static unsigned long read_idx(void)
+{
+ asm volatile (
+ "r6 = *(u64 *)(r1 + 0);"
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_array] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 == 0 goto 1f;"
+ "r0 += r6;"
+ "r0 = *(u8 *)(r0 + 0);"
+ "1:"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm_addr(map_array)
+ : __clobber_all);
+}
+
+/*
+ * Stack slots of the caller that might be read by the callee of callx
+ * have to be considered alive at the checkpoints before callx and
+ * inside of the callee. Otherwise the state with fp-8 == 1000 is pruned
+ * and out of bounds access in read_idx() goes unnoticed.
+ */
+SEC("socket")
+__failure __msg("invalid access to map value, value_size=8 off=1000 size=1")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 1000;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[read_idx] ll;"
+ "callx r2;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* in bounds access in read_idx() is fine */
+SEC("socket")
+__success __retval(0)
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack_ok(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 7;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[read_idx] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* same as above, but the pointer to the stack is passed through one more frame */
+SEC("socket")
+__failure __msg("invalid access to map value, value_size=8 off=1000 size=1")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack_nested(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 1000;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = %[read_idx] ll;"
+ "r2 = r10;"
+ "r2 += -8;"
+ "call apply;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* registers that are constant before callx are not constant after it */
+SEC("socket")
+__success __retval(1)
+__naked void callx_clobbers_const_regs(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "r1 = 0;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ /* dead branch pruning must not assume that r0 is still 0 */
+ "if r0 == 0 goto 1f;"
+ "r0 = 1;"
+ "exit;"
+ "1:"
+ "r0 = 2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* scalar argument passed through callx is tracked precisely */
+__naked __noinline __used
+static unsigned long identity(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "exit;"
+ );
+}
+
+long long vals[] SEC(".data.vals") = {1, 2, 3, 4};
+
+SEC("socket")
+__success __log_level(2)
+__msg("mark_precise: frame0: regs=r0 stack= before 12: (95) exit")
+__msg("mark_precise: frame1: regs=r0 stack= before 11: (bf) r0 = r1")
+__msg("mark_precise: frame1: regs=r1 stack= before 4: (8d) callx r2")
+__msg("mark_precise: frame0: regs=r1 stack= before 3: (bf) r1 = r6")
+__msg("mark_precise: frame0: regs=r6 stack= before 2: (b7) r6 = 3")
+__retval(4)
+__naked void callx_precision(void)
+{
+ asm volatile (
+ "r2 = %[identity] ll;"
+ "r6 = 3;"
+ "r1 = r6;"
+ "callx r2;"
+ "r0 *= 8;"
+ "r1 = %[vals] ll;"
+ "r1 += r0;"
+ "r0 = *(u64 *)(r1 + 0);"
+ "exit;"
+ :
+ : __imm_addr(identity),
+ __imm_addr(vals)
+ : __clobber_all);
+}
+
+/* function pointers in C */
+
+typedef int (*op_fn)(int);
+
+static __noinline int mul3(int x)
+{
+ return x * 3;
+}
+
+static __noinline int sub7(int x)
+{
+ return x - 7;
+}
+
+static __noinline int apply_op(op_fn op, int x)
+{
+ return op(x);
+}
+
+SEC("socket")
+__success __retval(36)
+int callx_c_fn_as_arg(void *ctx)
+{
+ /* (5 * 3) + (28 - 7) */
+ return apply_op(mul3, 5) + apply_op(sub7, 28);
+}
+
+SEC("socket")
+__success __retval(30)
+int callx_c_select(void *ctx)
+{
+ __u32 rnd = bpf_get_prandom_u32() & 1;
+ op_fn op = rnd ? mul3 : sub7;
+ int x = rnd ? 10 : 37;
+
+ /* 10 * 3 or 37 - 7 */
+ return op(x);
+}
+
+struct ops {
+ op_fn first;
+ op_fn second;
+ int bias;
+};
+
+static __noinline int run_ops(const struct ops *ops, int x)
+{
+ return ops->second(ops->first(x)) + ops->bias;
+}
+
+SEC("socket")
+__success __retval(100)
+int callx_c_ops_on_stack(void *ctx)
+{
+ struct ops a, b;
+
+ /* avoid an initializer with function pointers in .rodata */
+ a.first = mul3;
+ a.second = sub7;
+ a.bias = 10;
+ b.first = sub7;
+ b.second = mul3;
+ b.bias = 1;
+
+ /* ((7 * 3) - 7 + 10) + ((32 - 7) * 3 + 1) */
+ return run_ops(&a, 7) + run_ops(&b, 32);
+}
+
+#else
+
+SEC("socket")
+__success
+int dummy(void *ctx)
+{
+ return 0;
+}
+
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
new file mode 100644
index 0000000000000..419a90b1f894e
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
@@ -0,0 +1,651 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Tests for callx through pointers to functions in read-only data */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../../../include/linux/filter.h"
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+/*
+ * Read-only data with pointers to functions, where the compiler puts tables
+ * of functions, structures of operations and vtables. libbpf resolves
+ * a pointer to the offset of the function in the program and the kernel
+ * recognizes it by that value.
+ */
+#define DATA(SECTION, NAME, ...) \
+ ".pushsection " SECTION ",@progbits;" \
+ ".balign 8;" \
+ #NAME "_%=:" \
+ __VA_ARGS__ \
+ ".type " #NAME "_%=, @object;" \
+ ".size " #NAME "_%=, .-" #NAME "_%=;" \
+ ".popsection;"
+
+#define RODATA(NAME, ...) DATA(".rodata,\"a\"", NAME, __VA_ARGS__)
+
+#define FUNC_TABLE2(NAME, F0, F1) RODATA(NAME, ".quad " #F0 "; .quad " #F1 ";")
+
+__naked __noinline __used
+static unsigned long add1(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long add2(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 2;"
+ "exit;"
+ );
+}
+
+/* the second element of the table is called */
+SEC("socket")
+__success __retval(12)
+__naked void callx_rodata_const_index(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * The program reads the address of a function from the data, so pointers to
+ * functions in data are recognized only for programs that may leak pointers.
+ * Otherwise nothing refers to the functions.
+ */
+SEC("socket")
+__success __retval(12)
+__failure_unpriv __msg_unpriv("unreachable insn")
+__caps_unpriv(CAP_BPF)
+__naked void callx_rodata_needs_perfmon(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the address of an element is used instead of the address of the table */
+SEC("socket")
+__success __retval(12)
+__naked void callx_rodata_elem_addr(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= + 8 ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* pointers to functions are mixed with other data like in a vtable */
+SEC("socket")
+__success __retval(23)
+__log_level(2)
+__msg("r1 = *(u64 *)(r6 +0) ; R1=7")
+__msg("r2 = *(u64 *)(r6 +8) ; R2=func()")
+__naked void callx_rodata_mixed_with_data(void)
+{
+ asm volatile (
+ RODATA(vt, ".quad 7; .quad add1; .quad 13; .quad add2;")
+ "r6 = vt_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ /* add1(7) */
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = *(u64 *)(r6 + 16);"
+ "r2 = *(u64 *)(r6 + 24);"
+ /* add2(13) */
+ "callx r2;"
+ "r0 += r7;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * Position independent code keeps constants with pointers in .data.rel.ro.
+ * libbpf treats it as read-only data.
+ */
+SEC("socket")
+__success __retval(23)
+__naked void callx_data_rel_ro(void)
+{
+ asm volatile (
+ DATA(".data.rel.ro,\"aw\"", vt, ".quad 7; .quad add1; .quad 13; .quad add2;")
+ "r6 = vt_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = *(u64 *)(r6 + 16);"
+ "r2 = *(u64 *)(r6 + 24);"
+ "callx r2;"
+ "r0 += r7;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* an array of structures: the index selects one of the functions, not the data */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_array_of_structs(void)
+{
+ asm volatile (
+ RODATA(arr, ".quad add1; .quad 0x1111; .quad add2; .quad 0x2222;")
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 4;"
+ "r2 = arr_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* dynamic dispatch: which vtable is used is not known until run time */
+SEC("socket")
+__success __retval(2)
+__naked void callx_rodata_two_vtables(void)
+{
+ asm volatile (
+ RODATA(vt_a, ".quad 1; .quad add1;")
+ RODATA(vt_b, ".quad 0; .quad add2;")
+ "call %[bpf_get_prandom_u32];"
+ "r6 = vt_a_%= ll;"
+ "r0 &= 1;"
+ "if r0 == 0 goto +2;"
+ "r6 = vt_b_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ /* add1(1) or add2(0) */
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* both functions are verified, either of them is called */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_var_index(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* a table that has no symbol, the compiler generates such for a switch statement */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_no_symbol(void)
+{
+ asm volatile (
+ ".pushsection .rodata,\"a\",@progbits;"
+ ".balign 8;"
+ ".Lanon_%=:"
+ ".quad add1;"
+ ".quad add2;"
+ ".popsection;"
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 3;"
+ "r2 = .Lanon_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* the first instruction of the function is removed by the verifier */
+__naked __noinline __used
+static unsigned long nop_add1(void)
+{
+ asm volatile (
+ "goto +0;"
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+/* the pointer follows the function when instructions are removed */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_func_starts_with_nop(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad nop_add1; .quad 0;")
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* can't be called with a scalar in r1 */
+__naked __noinline __used
+static unsigned long deref_r1(void)
+{
+ asm volatile (
+ "r0 = *(u64 *)(r1 + 0);"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long ret0(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "exit;"
+ );
+}
+
+/* every function that might be called is verified */
+SEC("socket")
+__failure __msg("R1 invalid mem access 'scalar'")
+__naked void callx_rodata_all_callees_verified(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, deref_r1)
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 1;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 0;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* a function that is never called is not verified, it's dead code */
+SEC("socket")
+__success __retval(0)
+__naked void callx_rodata_unused_callee(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, deref_r1)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 0;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the index may select what is not a pointer to a function */
+SEC("socket")
+__failure __msg("overlaps with a pointer to a function")
+__naked void callx_rodata_index_beyond_table(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad add2; .quad 0x1234; .quad 0x5678;")
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 3;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* the address of a function is not known until the program is jitted */
+SEC("socket")
+__failure __msg("read of 4 bytes at offset")
+__msg("overlaps with a pointer to a function")
+__naked void callx_rodata_partial_read(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r0 = *(u32 *)(r2 + 0);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("read of 8 bytes at offset")
+__msg("overlaps with a pointer to a function")
+__naked void callx_rodata_misaligned_read(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad add2; .quad 0;")
+ "r2 = tbl_%= ll;"
+ "r0 = *(u64 *)(r2 + 4);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the data next to a pointer is still known to the verifier */
+SEC("socket")
+__success __retval(0x1234)
+__log_level(2)
+__msg("R0=4660")
+__naked void callx_rodata_data_is_const(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad 0x1234;")
+ "r2 = tbl_%= ll;"
+ "r0 = *(u64 *)(r2 + 8);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the pointer read from the data can't be modified */
+SEC("socket")
+__failure __msg("dereference of modified func ptr R2 off=8 disallowed")
+__naked void callx_rodata_modified_ptr(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r2 += 8;"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+__noinline __used
+int global_add3(int x)
+{
+ return x + 3;
+}
+
+/* a pointer to a global function is not recognized, it's a number */
+SEC("socket")
+__failure __msg("R2 has type scalar, expected func")
+__naked void callx_rodata_global_func(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, global_add3)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* selfcall(n) { return n ? tbl[0](n - 1) : 0; }, where tbl[0] == selfcall */
+__naked __noinline __used
+static unsigned long selfcall(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, selfcall, ret0)
+ "r0 = 0;"
+ "if r1 == 0 goto 1f;"
+ "r1 += -1;"
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "callx r2;"
+ "1:"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("recursive call from selfcall() to selfcall()")
+__naked void callx_rodata_recursion(void)
+{
+ asm volatile (
+ "r1 = 2;"
+ "call selfcall;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long use_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "exit;"
+ );
+}
+
+/* stack of all possible callees is accounted */
+SEC("socket")
+__failure __msg("combined stack size of 2 calls is")
+__naked void callx_rodata_stack_depth(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, use_stack_304)
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 1;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* pointers to functions in C */
+
+typedef int (*op_fn)(int);
+
+#define DEFINE_OP(N) static __noinline int op##N(int x) { return x * (N + 2) + N; }
+
+DEFINE_OP(0) DEFINE_OP(1) DEFINE_OP(2) DEFINE_OP(3)
+DEFINE_OP(4) DEFINE_OP(5) DEFINE_OP(6) DEFINE_OP(7)
+DEFINE_OP(8) DEFINE_OP(9) DEFINE_OP(10) DEFINE_OP(11)
+DEFINE_OP(12) DEFINE_OP(13) DEFINE_OP(14) DEFINE_OP(15)
+
+static op_fn const ops[] = {
+ op0, op1, op2, op3, op4, op5, op6, op7,
+ op8, op9, op10, op11, op12, op13, op14, op15,
+};
+
+int op_idx = 11;
+
+SEC("socket")
+__success __retval(50)
+int callx_c_table(void *ctx)
+{
+ unsigned int i = op_idx;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ /* op11(3) = 3 * 13 + 11 */
+ return ops[i](3);
+}
+
+SEC("socket")
+__success __retval(50)
+int callx_c_table_null_check(void *ctx)
+{
+ unsigned int i = op_idx;
+ op_fn op;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ /* the address of a function is not known until the program is jitted */
+ op = ops[i];
+ if (!op)
+ return -2;
+ return op(3);
+}
+
+/* a structure of operations, where pointers to functions are mixed with data */
+struct shape_ops {
+ int id;
+ op_fn area;
+ long scale;
+ op_fn perimeter;
+};
+
+static const struct shape_ops square_ops = { 1, op1, 10, op2 };
+static const struct shape_ops circle_ops = { 2, op3, 20, op1 };
+
+static __noinline int use_shape(const struct shape_ops *ops, int x)
+{
+ return ops->area(x) * ops->scale + ops->perimeter(ops->id);
+}
+
+SEC("socket")
+__success __retval(433)
+int callx_c_ops_mixed_with_data(void *ctx)
+{
+ /*
+ * square: op1(2) * 10 + op2(1) = 7 * 10 + 6 = 76
+ * circle: op3(3) * 20 + op1(2) = 18 * 20 + 7 = 367
+ * minus 10 when op_idx is not what it is set to
+ */
+ return use_shape(&square_ops, 2) + use_shape(&circle_ops, 3) - (op_idx == 11 ? 10 : 0);
+}
+
+/* the ops are selected at run time */
+SEC("socket")
+__success __retval(367)
+int callx_c_ops_selected(void *ctx)
+{
+ const struct shape_ops *ops = op_idx == 11 ? &circle_ops : &square_ops;
+
+ return use_shape(ops, 3);
+}
+
+/* the compiler might turn the switch into a table that has no symbol */
+static __noinline int call_by_switch(unsigned int idx, int x)
+{
+ op_fn op;
+
+ switch (idx) {
+ case 0:
+ op = op8;
+ break;
+ case 1:
+ op = op9;
+ break;
+ case 2:
+ op = op10;
+ break;
+ case 3:
+ op = op11;
+ break;
+ case 4:
+ op = op12;
+ break;
+ case 5:
+ op = op13;
+ break;
+ case 6:
+ op = op14;
+ break;
+ case 7:
+ op = op15;
+ break;
+ default:
+ return -1;
+ }
+ return op(x);
+}
+
+SEC("socket")
+__success __retval(50)
+int callx_c_switch_table(void *ctx)
+{
+ /* op11(3) = 3 * 13 + 11 */
+ return call_by_switch(op_idx - 8, 3);
+}
+
+/*
+ * Misaligned pointers to functions are ignored, the rest of the data is
+ * accessible as before.
+ */
+static const struct {
+ char tag;
+ op_fn op;
+} __attribute__((packed)) packed_ops[] = {
+ { 5, op0 },
+ { 7, op1 },
+};
+
+SEC("socket")
+__success __retval(7)
+int callx_c_packed_struct(void *ctx)
+{
+ return packed_ops[op_idx & 1].tag;
+}
+
+#else
+
+SEC("socket")
+__success
+int dummy(void *ctx)
+{
+ return 0;
+}
+
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c b/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
index 966f493487874..27fbe54e8795a 100644
--- a/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
+++ b/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
@@ -536,4 +536,35 @@ int return_from_void_global(struct __sk_buff *skb)
return 0;
}
+int global_calls_loop(int x);
+
+static __noinline int static_calls_global(int x)
+{
+ return global_calls_loop(x);
+}
+
+static __noinline int loop_cb_calls_static(u32 i, void *ctx)
+{
+ return static_calls_global(i);
+}
+
+__noinline int global_calls_loop(int x)
+{
+ bpf_loop(1, loop_cb_calls_static, NULL, 0);
+ return 0;
+}
+
+/*
+ * loop_cb_calls_static() -> static_calls_global() -> global_calls_loop() ->
+ * bpf_loop() -> loop_cb_calls_static() is an unbounded recursion that the
+ * main verification pass can't see, because it doesn't follow calls of global
+ * functions. None of the functions use stack.
+ */
+SEC("?raw_tp")
+__failure __msg("recursive call from global_calls_loop() to loop_cb_calls_static()")
+int recursion_via_global_func_and_callback(const void *ctx)
+{
+ return loop_cb_calls_static(0, NULL);
+}
+
char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/verifier/basic_call.c b/tools/testing/selftests/bpf/verifier/basic_call.c
index a8c6ab4c1622f..0f93c4551f235 100644
--- a/tools/testing/selftests/bpf/verifier/basic_call.c
+++ b/tools/testing/selftests/bpf/verifier/basic_call.c
@@ -4,7 +4,7 @@
BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, 0, 0, 0, 0),
BPF_EXIT_INSN(),
},
- .errstr = "unknown opcode 8d",
+ .errstr = "R0 !read_ok",
.result = REJECT,
},
{
|
| 3/2 | 2026/09/24 01:45 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 59711 output: 23 thoughts: 635 Thoughts: **Evaluating the Syzkaller Fuzzing Patch Series** Okay, so I need to assess this patch series targeting syzkaller fuzzing, specifically for the amd64 architecture. My first thought is to dive into what's actually being implemented. It looks like this patch introduces support for `BPF_CALL | BPF_X | BPF_JMP` (0x8d), or the `callx dst` instruction. This is a big deal; we're talking about indirect calls of static BPF subprograms, with the address located in a destination register. The scope of this change is pretty extensive. I see modifications across the board: documentation, both the arm64 and x86 BPF JIT compilers (the x86 changes are particularly interesting, especially for indirect calls!), a bunch of include files, and a slew of files within the kernel's BPF subsystem. We're looking at code touching the verifier, control flow graph manipulation, constant folding, disassembly, fixups, liveness analysis, and so on. There's also some work in the `tools/lib/bpf` and `tools/testing/selftests` directories. This is clearly a significant new feature for BPF. I'm seeing substantial updates to the BPF verifier (which is crucial, of course) and the x86 BPF JIT compiler – `emit_indirect_call`, `do_jit`, and `bpf_jit_supports_callx`. The arm64 JIT compiler is also being modified. There are changes in BPF map handling, specifically around `bpf_map_claim`, with patching of function pointers in rodata maps at JIT time. Verifier checks are being added to detect recursion, track stack depth, perform pointer resolution, and even patch maps with `cmpxchg64` into map values. Crucially, this is core functionality: it's in the kernel's BPF subsystem and the x86 JIT compiler. The amd64 architecture is supported, so that's good! This new functionality is highly accessible via the standard `bpf()` syscall (`BPF_PROG_LOAD`, `BPF_MAP_CREATE`, etc.). Given all that, the verifier, in particular, is one of the most critical security boundaries, a prime target for syzkaller. Any bugs here can lead to kernel panics or arbitrary code execution. This is *definitely* worth fuzzing. The `WorthFuzzing` flag is set to `true`. Now, I need to figure out the specifics for syzkaller. I need to nail down the `EnableConfigs` to determine what kernel configurations are required to activate this code and then the `FocusSymbols` to find the non-ubiquitous kernel functions to target. Time to dig into the code and Kconfig files. |
| 4/2 | 2026/09/24 01:45 | tool | grepper |
0mArgs:{
"Expression": "bpf_jit_supports_callx"
}Results: {
"Output": "arch/arm64/net/bpf_jit_comp.c=2494=bool bpf_jit_supports_subprog_tailcalls(void)\n--\narch/arm64/net/bpf_jit_comp.c-2498-\narch/arm64/net/bpf_jit_comp.c:2499:bool bpf_jit_supports_callx(void)\narch/arm64/net/bpf_jit_comp.c-2500-{\n--\narch/x86/net/bpf_jit_comp.c=4283=struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_prog *prog)\n--\narch/x86/net/bpf_jit_comp.c-4474-\narch/x86/net/bpf_jit_comp.c:4475:bool bpf_jit_supports_callx(void)\narch/x86/net/bpf_jit_comp.c-4476-{\n--\ninclude/linux/filter.h=1239=bool bpf_jit_supports_kfunc_call(void);\ninclude/linux/filter.h:1240:bool bpf_jit_supports_callx(void);\ninclude/linux/filter.h-1241-bool bpf_jit_supports_kfunc_ret_reg_pair(void);\n--\nkernel/bpf/core.c=3291=bool __weak bpf_jit_supports_kfunc_call(void)\n--\nkernel/bpf/core.c-3296-/* Return TRUE if the JIT backend supports callx (indirect call) instruction. */\nkernel/bpf/core.c:3297:bool __weak bpf_jit_supports_callx(void)\nkernel/bpf/core.c-3298-{\n--\nkernel/bpf/verifier.c=11059=static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,\n--\nkernel/bpf/verifier.c-11096-\t}\nkernel/bpf/verifier.c:11097:\tif (!bpf_jit_supports_callx()) {\nkernel/bpf/verifier.c-11098-\t\tverbose(env, \"JIT doesn't support callx\\n\");\n"
} |
| 5/2 | 2026/09/24 01:45 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 61588 output: 30 thoughts: 147 Thoughts: **BPF JIT Dependency Check** Okay, here's what's running through my head. I'm looking at `kernel/bpf/verifier.c`, specifically the part dealing with `callx`. The code's pretty clear: first, it checks if JIT compilation is even *requested* for the program. If not, it bails with an "unsupported" error because the interpreter doesn't handle `callx`. But even if JIT *is* requested, there's a second, crucial check: `bpf_jit_supports_callx()`. This implies that the presence of the JIT is a *necessary*, but not *sufficient* condition. We need to confirm that this particular JIT implementation actually has `callx` support. If the JIT doesn't have the `callx` implementation, we're still going to fail. That's the real takeaway here. Looks like I need to dig into the BPF JIT configuration and make sure it's enabled and configured correctly for `callx` support. |
| 6/2 | 2026/09/24 01:45 | tool | grepper |
0mArgs:{
"Expression": "BPF_JIT",
"PathPrefix": "kernel/bpf/"
}Results: {
"Output": "kernel/bpf/Kconfig=4=config BPF\n--\nkernel/bpf/Kconfig-12-# Classic BPF JIT (cBPF)\nkernel/bpf/Kconfig:13:config HAVE_CBPF_JIT\nkernel/bpf/Kconfig-14-\tbool\n--\nkernel/bpf/Kconfig-16-# Extended BPF JIT (eBPF)\nkernel/bpf/Kconfig:17:config HAVE_EBPF_JIT\nkernel/bpf/Kconfig-18-\tbool\n--\nkernel/bpf/Kconfig-20-# KASAN support for JIT compiler\nkernel/bpf/Kconfig:21:config HAVE_EBPF_JIT_KASAN\nkernel/bpf/Kconfig-22-\tbool\n--\nkernel/bpf/Kconfig-25-# default for kernels that were compiled with BPF JIT support.\nkernel/bpf/Kconfig:26:config ARCH_WANT_DEFAULT_BPF_JIT\nkernel/bpf/Kconfig-27-\tbool\n--\nkernel/bpf/Kconfig=31=config BPF_SYSCALL\n--\nkernel/bpf/Kconfig-45-\nkernel/bpf/Kconfig:46:config BPF_JIT\nkernel/bpf/Kconfig-47-\tbool \"Enable BPF Just In Time compiler\"\nkernel/bpf/Kconfig-48-\tdepends on BPF\nkernel/bpf/Kconfig:49:\tdepends on HAVE_CBPF_JIT || HAVE_EBPF_JIT\nkernel/bpf/Kconfig-50-\tselect EXECMEM\n--\nkernel/bpf/Kconfig-61-\nkernel/bpf/Kconfig:62:config BPF_JIT_ALWAYS_ON\nkernel/bpf/Kconfig-63-\tbool \"Permanently enable BPF JIT and remove BPF interpreter\"\nkernel/bpf/Kconfig:64:\tdepends on BPF_SYSCALL \u0026\u0026 HAVE_EBPF_JIT \u0026\u0026 BPF_JIT\nkernel/bpf/Kconfig-65-\thelp\n--\nkernel/bpf/Kconfig-68-\nkernel/bpf/Kconfig:69:\t When CONFIG_BPF_JIT_ALWAYS_ON is enabled, /proc/sys/net/core/bpf_jit_enable\nkernel/bpf/Kconfig-70-\t is permanently set to 1 and setting any other value than that will\n--\nkernel/bpf/Kconfig-72-\nkernel/bpf/Kconfig:73:config BPF_JIT_DEFAULT_ON\nkernel/bpf/Kconfig:74:\tdef_bool ARCH_WANT_DEFAULT_BPF_JIT || BPF_JIT_ALWAYS_ON\nkernel/bpf/Kconfig:75:\tdepends on HAVE_EBPF_JIT \u0026\u0026 BPF_JIT\nkernel/bpf/Kconfig-76-\n--\nkernel/bpf/Kconfig=96=config BPF_LSM\n--\nkernel/bpf/Kconfig-100-\tdepends on SECURITY\nkernel/bpf/Kconfig:101:\tdepends on BPF_JIT\nkernel/bpf/Kconfig-102-\thelp\n--\nkernel/bpf/Kconfig-107-\nkernel/bpf/Kconfig:108:config BPF_JIT_KASAN\nkernel/bpf/Kconfig-109-\tbool\nkernel/bpf/Kconfig:110:\tdepends on HAVE_EBPF_JIT_KASAN\nkernel/bpf/Kconfig-111-\tdepends on KASAN_GENERIC\nkernel/bpf/Kconfig-112-\tdepends on KASAN_VMALLOC\nkernel/bpf/Kconfig:113:\tdepends on BPF_JIT\nkernel/bpf/Kconfig-114-\tdefault y\n--\nkernel/bpf/Makefile=2=obj-y := core.o\nkernel/bpf/Makefile:3:ifneq ($(CONFIG_BPF_JIT_ALWAYS_ON),y)\nkernel/bpf/Makefile-4-# ___bpf_prog_run() needs GCSE disabled on x86; see 3193c0836f203 for details\n--\nkernel/bpf/Makefile=16=obj-$(CONFIG_BPF_SYSCALL) += disasm.o mprog.o\nkernel/bpf/Makefile:17:obj-$(CONFIG_BPF_JIT) += trampoline.o\nkernel/bpf/Makefile-18-obj-$(CONFIG_BPF_SYSCALL) += btf.o memalloc.o rqspinlock.o stream.o\n--\nkernel/bpf/Makefile=21=endif\nkernel/bpf/Makefile:22:obj-$(CONFIG_BPF_JIT) += dispatcher.o\nkernel/bpf/Makefile-23-ifeq ($(CONFIG_NET),y)\n--\nkernel/bpf/Makefile=45=endif\nkernel/bpf/Makefile:46:ifeq ($(CONFIG_BPF_JIT),y)\nkernel/bpf/Makefile-47-obj-$(CONFIG_BPF_SYSCALL) += bpf_struct_ops.o\n--\nkernel/bpf/btf.c=9938=bool btf_type_ids_nocast_alias(struct bpf_verifier_log *log,\n--\nkernel/bpf/btf.c-9992-\nkernel/bpf/btf.c:9993:#ifdef CONFIG_BPF_JIT\nkernel/bpf/btf.c-9994-static int\n--\nkernel/bpf/core.c=100=struct bpf_prog *bpf_prog_alloc_no_stats(unsigned int size, gfp_t gfp_extra_flags)\n--\nkernel/bpf/core.c-128-\tfp-\u003ejit_requested = ebpf_jit_enabled();\nkernel/bpf/core.c:129:\tfp-\u003ejit_required = IS_ENABLED(CONFIG_BPF_JIT_ALWAYS_ON);\nkernel/bpf/core.c-130-\tfp-\u003eblinding_requested = bpf_jit_blinding_enabled(fp);\n--\nkernel/bpf/core.c=541=void bpf_prog_kallsyms_del_all(struct bpf_prog *fp)\n--\nkernel/bpf/core.c-546-\nkernel/bpf/core.c:547:#ifdef CONFIG_BPF_JIT\nkernel/bpf/core.c-548-/* All BPF JIT sysctl knobs here. */\nkernel/bpf/core.c:549:int bpf_jit_enable __read_mostly = IS_BUILTIN(CONFIG_BPF_JIT_DEFAULT_ON);\nkernel/bpf/core.c:550:int bpf_jit_kallsyms __read_mostly = IS_BUILTIN(CONFIG_BPF_JIT_DEFAULT_ON);\nkernel/bpf/core.c-551-int bpf_jit_harden __read_mostly;\n--\nkernel/bpf/core.c=1357=static int bpf_jit_blind_insn(const struct bpf_insn *from,\n--\nkernel/bpf/core.c-1365-\nkernel/bpf/core.c:1366:\tBUILD_BUG_ON(BPF_REG_PARAMS + 2 != MAX_BPF_JIT_REG);\nkernel/bpf/core.c:1367:\tBUILD_BUG_ON(BPF_REG_AX + 1 != MAX_BPF_JIT_REG);\nkernel/bpf/core.c-1368-\n--\nkernel/bpf/core.c=1650=u16 bpf_out_stack_arg_cnt(const struct bpf_verifier_env *env, const struct bpf_prog *prog)\n--\nkernel/bpf/core.c-1658-}\nkernel/bpf/core.c:1659:#endif /* CONFIG_BPF_JIT */\nkernel/bpf/core.c-1660-\n--\nkernel/bpf/core.c=1818=bool bpf_opcode_in_insntable(u8 code)\n--\nkernel/bpf/core.c-1841-\nkernel/bpf/core.c:1842:#ifndef CONFIG_BPF_JIT_ALWAYS_ON\nkernel/bpf/core.c-1843-/* Absolute value of s32 without undefined behavior for S32_MIN */\n--\nkernel/bpf/core.c=2494=static unsigned int __bpf_prog_ret0_warn(const void *ctx,\n--\nkernel/bpf/core.c-2496-{\nkernel/bpf/core.c:2497:\t/* If this handler ever gets executed, then BPF_JIT_ALWAYS_ON\nkernel/bpf/core.c-2498-\t * is not working properly, so warn about it!\n--\nkernel/bpf/core.c=2606=static bool bpf_prog_select_interpreter(struct bpf_prog *fp)\n--\nkernel/bpf/core.c-2608-\tbool select_interpreter = false;\nkernel/bpf/core.c:2609:#ifndef CONFIG_BPF_JIT_ALWAYS_ON\nkernel/bpf/core.c-2610-\tu32 stack_depth = max_t(u32, fp-\u003eaux-\u003estack_depth, 1);\n--\nkernel/bpf/core.c=2629=static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struct bpf_prog *prog)\nkernel/bpf/core.c-2630-{\nkernel/bpf/core.c:2631:#ifdef CONFIG_BPF_JIT\nkernel/bpf/core.c-2632-\tstruct bpf_prog *orig_prog;\n--\nkernel/bpf/core.c=3360=u32 bpf_jit_plan_arg_moves(const struct bpf_jit_arg_abi *abi,\n--\nkernel/bpf/core.c-3382-\t\tif (pos \u003c slot) {\nkernel/bpf/core.c:3383:\t\t\tmoves[n].dst = BPF_JIT_ARG_TMP;\nkernel/bpf/core.c-3384-\t\t\tback = slot;\n--\nkernel/bpf/core.c-3393-\t\tmoves[n].dst = pos_of_slot[back];\nkernel/bpf/core.c:3394:\t\tmoves[n].src = BPF_JIT_ARG_TMP;\nkernel/bpf/core.c-3395-\t\tn++;\n--\nkernel/bpf/fixups.c=188=static int add_kfunc_in_insns(struct bpf_verifier_env *env,\n--\nkernel/bpf/fixups.c-202-\nkernel/bpf/fixups.c:203:#ifndef CONFIG_BPF_JIT_ALWAYS_ON\nkernel/bpf/fixups.c-204-static int get_callee_stack_depth(struct bpf_verifier_env *env,\n--\nkernel/bpf/fixups.c=1492=int bpf_fixup_call_args(struct bpf_verifier_env *env)\nkernel/bpf/fixups.c-1493-{\nkernel/bpf/fixups.c:1494:#ifndef CONFIG_BPF_JIT_ALWAYS_ON\nkernel/bpf/fixups.c-1495-\tstruct bpf_prog *prog = env-\u003eprog;\n--\nkernel/bpf/fixups.c-1520-\t}\nkernel/bpf/fixups.c:1521:#ifndef CONFIG_BPF_JIT_ALWAYS_ON\nkernel/bpf/fixups.c-1522-\tif (prog-\u003ejit_required) {\n--\nkernel/bpf/syscall.c=6530=int kern_sys_bpf(int cmd, union bpf_attr *attr, unsigned int size)\n--\nkernel/bpf/syscall.c-6535-\tswitch (cmd) {\nkernel/bpf/syscall.c:6536:#ifdef CONFIG_BPF_JIT /* __bpf_prog_enter_sleepable used by trampoline and JIT */\nkernel/bpf/syscall.c-6537-\tcase BPF_PROG_TEST_RUN:\n--\nkernel/bpf/verifier.c=4727=BTF_ID(struct, cgroup)\nkernel/bpf/verifier.c-4728-#endif\nkernel/bpf/verifier.c:4729:#ifdef CONFIG_BPF_JIT\nkernel/bpf/verifier.c-4730-BTF_ID(struct, bpf_cpumask)\n"
} |
| 7/2 | 2026/09/24 01:45 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 65211 output: 32 thoughts: 67 Thoughts: Here's my thought process, as requested: **Digging into the BPF JIT Compilation Path** Okay, so I'm looking into how the BPF JIT compiler is implemented, specifically on x86. My first instinct is to check how it all hangs together. I need to zero in on how the syscall interface interacts with the JIT engine. So, let's start by looking into the kernel's definitions related to `BPF_SYSCALL` and `BPF_JIT`. These will give me the high-level entry points. Next, I need to see the actual JIT compilation code itself. The obvious place to begin is with the x86-specific implementation. I’m going to examine `arch/x86/net/bpf_jit_comp.c`. That file should contain the nitty-gritty of the code generation for x86. Finally, I need to understand the build process for this. How is `bpf_jit_comp.c` compiled and linked into the kernel? I’ll quickly check the `arch/x86/net/Makefile` to see how the build system handles this module. That will give me the compilation flags and dependencies. This approach will allow me to piece together the entire flow, from syscall entry point to generated machine code. |
| 8/2 | 2026/09/24 01:45 | tool | grepper |
0mArgs:{
"Expression": "bpf_jit_comp",
"PathPrefix": "arch/x86/"
}Results: {
"Output": "arch/x86/net/Makefile=6=ifeq ($(CONFIG_X86_32),y)\narch/x86/net/Makefile:7: obj-$(CONFIG_BPF_JIT) += bpf_jit_comp32.o\narch/x86/net/Makefile-8-else\narch/x86/net/Makefile:9: obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o\narch/x86/net/Makefile-10-endif\n"
} |
| 9/2 | 2026/09/24 01:45 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 65742 output: 31 thoughts: 1238 Thoughts: **Analyzing BPF Callx and Function Pointer Support** Okay, so I'm looking at this new BPF feature, `BPF_CALLX`, which clearly depends on both `BPF_SYSCALL` and `BPF_JIT`. On x86, we've got `HAVE_EBPF_JIT` enabled by default, which is convenient. The initial check I made confirmed that `BPF_JIT_ALWAYS_ON` or just `BPF_JIT` is needed. The verifier in `verifier.c` explicitly mandates the JIT if `callx` is being used, through `env->prog->jit_requested` which in turn depends on the result of `ebpf_jit_enabled()`, which boils down to whether `bpf_jit_enable` is true. Now I'm diving into the actual patch, checking the added or modified functions. `check_func_callx` handles the `BPF_CALL | BPF_X` instruction, and there is a lot of new functionality here to get the new behaviour operational. It appears that significant changes have been made to the verifier, particularly around handling indirect calls and static function pointers. We also have `emit_indirect_call` in the x86 JIT, which is crucial for actually generating the code. The patch also introduces new functions for managing function pointers in maps like `resolve_func_ptrs`, `check_func_ptr_read`, and several `bpf_map_*` functions. We also have changes in the CFG and liveness analysis, which indicates the feature touches quite a few key parts of the BPF subsystem. Next, I need to figure out if these new functions are good candidates for focus symbols for the fuzzer. Based on the guidelines, I need to ensure that the candidate functions are feature-specific and are not ubiquitous lifecyle hot-paths. Functions like `check_func_callx`, `resolve_func_ptrs`, `emit_indirect_call`, and especially `jit_subprogs` seem like solid candidates since they implement the core logic for enabling `callx` and handling function pointers. Functions like `bpf_map_claim` may also be worth considering because we're tying this feature to map management. I need to be careful, however, about functions that might be too generic. I'll need to check if these symbols actually exist in the vmlinux or object files. Now I'll start checking the `verifier.c` and `bpf_jit_comp.c` files for symbols. |
| 10/2 | 2026/09/24 01:45 | tool | grepper |
0mArgs:{
"Expression": "check_func_callx",
"PathPrefix": "kernel/bpf/"
}Results: {
"Output": "kernel/bpf/verifier.c=10952=static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,\n--\nkernel/bpf/verifier.c-11058- */\nkernel/bpf/verifier.c:11059:static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,\nkernel/bpf/verifier.c-11060-\t\t\t int *insn_idx)\n--\nkernel/bpf/verifier.c=19061=static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)\n--\nkernel/bpf/verifier.c-19132-\t\t\tif (bpf_is_callx(insn))\nkernel/bpf/verifier.c:19133:\t\t\t\treturn check_func_callx(env, insn, \u0026env-\u003einsn_idx);\nkernel/bpf/verifier.c-19134-\t\t\tif (insn-\u003esrc_reg == BPF_PSEUDO_CALL)\n"
} |
| 11/2 | 2026/09/24 01:45 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 67858 output: 33 thoughts: 952 Thoughts: **Identifying Key BPF Functionality for Focused Fuzzing** Okay, so I'm trying to figure out the right functions to focus on for fuzzing related to this BPF change. The instructions are clear: target specific logic, avoid the ubiquitous hot-paths. Let's break this down. First, `check_func_callx` in `kernel/bpf/verifier.c` immediately stands out. It's a static function, and while static functions can be inlined, this one is substantial and definitely not marked for inlining. That means it likely shows up in kallsyms and is specific. Excellent. Now, what else? I should look for *non-static* functions that are either new or have been heavily modified, as that's probably a good indicator of core logic. I'm seeing a cluster of potentially interesting functions: * **In `include/linux/bpf_verifier.h`:** `bpf_map_func_ptrs`, `bpf_map_range_func_ptrs`, `bpf_adjust_func_ptrs`, and `bpf_adjust_func_ptrs_after_remove`. * **In `include/linux/filter.h`:** `bpf_jit_supports_callx`. * **In `arch/x86/net/bpf_jit_comp.c`:** `bpf_jit_supports_callx` (again!) and `emit_indirect_call`. `emit_indirect_call` is static but potentially important because it is non-trivial. * **In `kernel/bpf/verifier.c`:** `check_func_callx`, `resolve_func_ptrs`, `check_func_ptr_read`, and `bpf_map_claim`. * **In `kernel/bpf/fixups.c`:** `bpf_patch_insn_data` (calls `bpf_adjust_func_ptrs`), and `jit_subprogs`. `bpf_check` is the main entry point, but it's called often, so it's probably too generic. I need to avoid that. So, the functions that really stand out as *specific* to this feature based on all the above analysis are the `check_func_callx` and other functions with `callx` or `func_ptr` in the name, like `resolve_func_ptrs`, `check_func_ptr_read`, and `bpf_jit_supports_callx`. Even `emit_indirect_call`, despite being static, looks critical because it's responsible for the *actual* indirect calls based on CPU features. I checked `kernel/bpf/verifier.c` and found that the map pointer functions are global, further solidifying their importance: `bpf_map_range_func_ptrs`, `bpf_map_func_ptrs`, `bpf_adjust_func_ptrs`, and `bpf_adjust_func_ptrs_after_remove`. `bpf_jit_supports_callx` is also global (defined in `arch/x86/net/bpf_jit_comp.c` and `kernel/bpf/core.c`). Now let's double-check `arch/x86/net/bpf_jit_comp.c` around `emit_indirect_call`. |
| 12/2 | 2026/09/24 01:45 | tool | grepper |
0mArgs:{
"Expression": "emit_indirect_call",
"PathPrefix": "arch/x86/net/"
}Results: {
"Output": "arch/x86/net/bpf_jit_comp.c=725=static void emit_indirect_jump(u8 **pprog, int bpf_reg, u8 *ip)\n--\narch/x86/net/bpf_jit_comp.c-751-\narch/x86/net/bpf_jit_comp.c:752:static void __emit_indirect_call(u8 **pprog, int reg, bool ereg)\narch/x86/net/bpf_jit_comp.c-753-{\n--\narch/x86/net/bpf_jit_comp.c-764-/* call *bpf_reg */\narch/x86/net/bpf_jit_comp.c:765:static int emit_indirect_call(u8 **pprog, int bpf_reg, u8 *ip)\narch/x86/net/bpf_jit_comp.c-766-{\n--\narch/x86/net/bpf_jit_comp.c-776-\t\tEMIT_LFENCE();\narch/x86/net/bpf_jit_comp.c:777:\t\t__emit_indirect_call(\u0026prog, reg, ereg);\narch/x86/net/bpf_jit_comp.c-778-\t} else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE)) {\n--\narch/x86/net/bpf_jit_comp.c-784-\t} else {\narch/x86/net/bpf_jit_comp.c:785:\t\t__emit_indirect_call(\u0026prog, reg, ereg);\narch/x86/net/bpf_jit_comp.c-786-\t}\n--\narch/x86/net/bpf_jit_comp.c=2002=static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *addrs, u8 *image,\n--\narch/x86/net/bpf_jit_comp.c-2995-\t\t\t}\narch/x86/net/bpf_jit_comp.c:2996:\t\t\tif (emit_indirect_call(\u0026prog, insn-\u003edst_reg, ip))\narch/x86/net/bpf_jit_comp.c-2997-\t\t\t\treturn -EINVAL;\n"
} |
| 13/2 | 2026/09/24 01:45 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 70203 output: 30 thoughts: 350 Thoughts: **Analyzing `jit_subprogs` and Defensive Checks in BPF JIT Compilation** Okay, let's dive into `jit_subprogs` within `kernel/bpf/fixups.c`. My understanding is that this is where the magic happens – specifically, when the JIT compiler is working on subprograms, it updates pointers within BPF maps. I need to refresh my memory on the exact mechanism, so let's zero in on lines 1289-1340. Right, here it is: the code iterates through `env->func_ptr_cnt`, and for each function pointer, it's replacing offsets with their final addresses in maps. The comment explicitly states that these are *final* addresses, and the map becomes exclusively owned by the program at this point. That makes sense, ensuring no unexpected modification later. Ah, and here we are. I see a `verifier_bug_if()` call. This is interesting – it means there's a runtime defensive check happening in a code path that's definitely reachable. `cmpxchg64` is being used to atomically swap the old offset with the new address within the map. The `verifier_bug_if` is a conditional check, and if the compare-and-swap fails unexpectedly (i.e., `cmpxchg64` *doesn't* return the expected `old` value), then a bug is indicated. This indicates a very specific type of potential race condition or modification in the map that the JIT compiler is actively trying to prevent. The program is asserting that the map entry doesn't change unexpectedly from under it after the pointers have been declared final and exclusively used by the program. That's a strong guarantee, and the defensive programming is clear. This is a very targeted check against a fairly complex scenario. |
| 14/2 | 2026/09/24 01:46 | tool | grepper |
0mArgs:{
"Expression": "verifier_bug_if",
"PathPrefix": "kernel/bpf/"
}Results: {
"Output": "kernel/bpf/backtrack.c=11=int bpf_push_jmp_history(struct bpf_verifier_env *env, struct bpf_verifier_state *cur,\n--\nkernel/bpf/backtrack.c-22-\t\t */\nkernel/bpf/backtrack.c:23:\t\tverifier_bug_if((env-\u003ecur_hist_ent-\u003eflags \u0026 insn_flags) \u0026\u0026\nkernel/bpf/backtrack.c-24-\t\t\t\t(env-\u003ecur_hist_ent-\u003eflags \u0026 insn_flags) != insn_flags,\n--\nkernel/bpf/backtrack.c-29-\t\tenv-\u003ecur_hist_ent-\u003eframe = frame;\nkernel/bpf/backtrack.c:30:\t\tverifier_bug_if(env-\u003ecur_hist_ent-\u003elinked_regs != 0, env,\nkernel/bpf/backtrack.c-31-\t\t\t\t\"insn history: insn_idx %d linked_regs: %#llx\",\n--\nkernel/bpf/backtrack.c=265=static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,\n--\nkernel/bpf/backtrack.c-426-\t\t\t\t */\nkernel/bpf/backtrack.c:427:\t\t\t\tverifier_bug_if(idx + 1 != subseq_idx, env,\nkernel/bpf/backtrack.c-428-\t\t\t\t\t\t\"extra insn from subprog\");\n--\nkernel/bpf/backtrack.c=817=int bpf_mark_chain_precision(struct bpf_verifier_env *env,\n--\nkernel/bpf/backtrack.c-952-\t\t\tfor_each_set_bit(i, mask, 64) {\nkernel/bpf/backtrack.c:953:\t\t\t\tif (verifier_bug_if(i \u003e= func-\u003eallocated_stack / BPF_REG_SIZE,\nkernel/bpf/backtrack.c-954-\t\t\t\t\t\t env, \"stack slot %d, total slots %d\",\n--\nkernel/bpf/fixups.c=204=static int get_callee_stack_depth(struct bpf_verifier_env *env,\n--\nkernel/bpf/fixups.c-209-\tsubprog = bpf_find_subprog(env, start);\nkernel/bpf/fixups.c:210:\tif (verifier_bug_if(subprog \u003c 0, env, \"get stack depth: no program at insn %d\", start))\nkernel/bpf/fixups.c-211-\t\treturn -EFAULT;\n--\nkernel/bpf/fixups.c=675=int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env,\n--\nkernel/bpf/fixups.c-749-\nkernel/bpf/fixups.c:750:\t\tif (verifier_bug_if(load_reg == -1, env,\nkernel/bpf/fixups.c-751-\t\t\t\t \"zext_dst is set, but no reg is defined\"))\n--\nkernel/bpf/fixups.c=777=int bpf_convert_ctx_accesses(struct bpf_verifier_env *env)\n--\nkernel/bpf/fixups.c-993-\t\t\t}\nkernel/bpf/fixups.c:994:\t\t\tif (verifier_bug_if(bpf_may_fault_on_deref(ptr_type), env,\nkernel/bpf/fixups.c-995-\t\t\t\t\t \"access to a fault prone pointer is not rewritten as a probed one\"))\n--\nkernel/bpf/fixups.c=1102=static int jit_subprogs(struct bpf_verifier_env *env)\n--\nkernel/bpf/fixups.c-1119-\t\tsubprog = bpf_find_subprog(env, i + insn-\u003eimm + 1);\nkernel/bpf/fixups.c:1120:\t\tif (verifier_bug_if(subprog \u003c 0, env, \"No program to jit at insn %d\",\nkernel/bpf/fixups.c-1121-\t\t\t\t i + insn-\u003eimm + 1))\n--\nkernel/bpf/fixups.c-1312-\t\terr = ptr-\u003emap-\u003eops-\u003emap_direct_value_addr(ptr-\u003emap, \u0026addr, 0);\nkernel/bpf/fixups.c:1313:\t\tif (verifier_bug_if(err, env, \"no value of map '%s'\", ptr-\u003emap-\u003ename)) {\nkernel/bpf/fixups.c-1314-\t\t\terr = -EFAULT;\n--\nkernel/bpf/fixups.c-1320-\t\t\tsubprog = bpf_find_subprog(env, ptr-\u003exlated_off);\nkernel/bpf/fixups.c:1321:\t\t\tif (verifier_bug_if(subprog \u003c= 0, env, \"no function at insn %u\",\nkernel/bpf/fixups.c-1322-\t\t\t\t\t ptr-\u003exlated_off)) {\n--\nkernel/bpf/fixups.c-1326-\t\t\tnew = (unsigned long)func[subprog]-\u003ebpf_func;\nkernel/bpf/fixups.c:1327:\t\t} else if (verifier_bug_if(ptr-\u003eused, env, \"function of map '%s' offset %u is removed\",\nkernel/bpf/fixups.c-1328-\t\t\t\t\t ptr-\u003emap-\u003ename, ptr-\u003emap_off)) {\n--\nkernel/bpf/fixups.c-1335-\t\told = (u64)ptr-\u003eorig_off * sizeof(struct bpf_insn);\nkernel/bpf/fixups.c:1336:\t\tif (verifier_bug_if(cmpxchg64((u64 *)(unsigned long)addr, old, new) != old, env,\nkernel/bpf/fixups.c-1337-\t\t\t\t \"map '%s' offset %u changed\", ptr-\u003emap-\u003ename, ptr-\u003emap_off)) {\n--\nkernel/bpf/states.c=273=int bpf_update_branch_counts(struct bpf_verifier_env *env, struct bpf_verifier_state *st)\n--\nkernel/bpf/states.c-281-\nkernel/bpf/states.c:282:\t\t/* verifier_bug_if(br \u003e 1, ...) technically makes sense here,\nkernel/bpf/states.c-283-\t\t * but see comment in push_stack(), hence:\nkernel/bpf/states.c-284-\t\t */\nkernel/bpf/states.c:285:\t\tverifier_bug_if((int)br \u003c 0, env, \"%s:branches_to_explore=%d\", __func__, br);\nkernel/bpf/states.c-286-\t\tif (br)\n--\nkernel/bpf/states.c=1236=int bpf_is_state_visited(struct bpf_verifier_env *env, int insn_idx)\n--\nkernel/bpf/states.c-1580-\tnew-\u003einsn_idx = insn_idx;\nkernel/bpf/states.c:1581:\tverifier_bug_if(new-\u003ebranches != 1, env,\nkernel/bpf/states.c-1582-\t\t\t\"%s:branches_to_explore=%d insn %d\",\n--\nkernel/bpf/verifier.c=437=static int bpf_compute_subprog_ret_regs(struct bpf_verifier_env *env)\n--\nkernel/bpf/verifier.c-456-\t\t\tcontinue;\nkernel/bpf/verifier.c:457:\t\tif (verifier_bug_if(IS_ERR(btf_resolve_size(btf, type, \u0026size)), env,\nkernel/bpf/verifier.c-458-\t\t\t\t \"cannot size return type of subprog %d\", subprog))\n--\nkernel/bpf/verifier.c=5445=static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,\n--\nkernel/bpf/verifier.c-5578-\t\t\tsidx = bpf_find_subprog(env, next_insn);\nkernel/bpf/verifier.c:5579:\t\t\tif (verifier_bug_if(sidx \u003c 0, env, \"callee not found at insn %d\", next_insn))\nkernel/bpf/verifier.c-5580-\t\t\t\treturn -EFAULT;\n--\nkernel/bpf/verifier.c=6659=static int check_func_ptr_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,\n--\nkernel/bpf/verifier.c-6696-\t\tsubprog = bpf_find_subprog(env, ptrs[i].xlated_off);\nkernel/bpf/verifier.c:6697:\t\tif (verifier_bug_if(subprog \u003c= 0, env, \"no function at insn %u for map '%s' offset %u\",\nkernel/bpf/verifier.c-6698-\t\t\t\t ptrs[i].xlated_off, map-\u003ename, ptrs[i].map_off))\n--\nkernel/bpf/verifier.c-6702-\t}\nkernel/bpf/verifier.c:6703:\tif (verifier_bug_if(!n, env, \"no offsets to read map '%s' at\", map-\u003ename))\nkernel/bpf/verifier.c-6704-\t\treturn -EFAULT;\n--\nkernel/bpf/verifier.c=10952=static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,\n--\nkernel/bpf/verifier.c-10961-\tsubprog = bpf_find_subprog(env, target_insn);\nkernel/bpf/verifier.c:10962:\tif (verifier_bug_if(subprog \u003c 0, env, \"target of func call at insn %d is not a program\",\nkernel/bpf/verifier.c-10963-\t\t\t target_insn))\n--\nkernel/bpf/verifier.c=15357=static int adjust_ptr_min_max_vals(struct bpf_verifier_env *env, struct bpf_insn *insn,\n--\nkernel/bpf/verifier.c-15615-\t\t\t\t \u0026info, true);\nkernel/bpf/verifier.c:15616:\t\tif (verifier_bug_if(!can_skip_alu_sanitation(env, insn)\nkernel/bpf/verifier.c-15617-\t\t\t\t \u0026\u0026 !env-\u003ecur_state-\u003especulative\n--\nkernel/bpf/verifier.c=19003=static int check_indirect_jump(struct bpf_verifier_env *env, struct bpf_insn *insn)\n--\nkernel/bpf/verifier.c-19020-\tmap = dst_reg-\u003emap_ptr;\nkernel/bpf/verifier.c:19021:\tif (verifier_bug_if(!map, env, \"R%d has an empty map pointer\", insn-\u003edst_reg))\nkernel/bpf/verifier.c-19022-\t\treturn -EFAULT;\nkernel/bpf/verifier.c-19023-\nkernel/bpf/verifier.c:19024:\tif (verifier_bug_if(map-\u003emap_type != BPF_MAP_TYPE_INSN_ARRAY, env,\nkernel/bpf/verifier.c-19025-\t\t\t \"R%d has incorrect map type %d\", insn-\u003edst_reg, map-\u003emap_type))\n--\nkernel/bpf/verifier.c=19174=static int do_check(struct bpf_verifier_env *env)\n--\nkernel/bpf/verifier.c-19310-\t\t\t\t\tcontinue;\nkernel/bpf/verifier.c:19311:\t\t\t\tif (verifier_bug_if((u32)regs[r].var_off.value != cval,\nkernel/bpf/verifier.c-19312-\t\t\t\t\t\t env, \"const R%d: %u != %llu\",\n--\nkernel/bpf/verifier.c-19358-\t\t\t */\nkernel/bpf/verifier.c:19359:\t\t\tif (verifier_bug_if((BPF_CLASS(insn-\u003ecode) == BPF_JMP ||\nkernel/bpf/verifier.c-19360-\t\t\t\t\t BPF_CLASS(insn-\u003ecode) == BPF_JMP32) \u0026\u0026\n"
} |
| 15/2 | 2026/09/24 01:46 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 74148 output: 33 thoughts: 60 Thoughts: Here's my summary, as the thought process unfolds: **Examining the `verifier_bug_if` Macro** Okay, let's dissect this. I see `verifier_bug_if` defined as a macro. Looking at the definition, it appears to be a conditional bug reporting mechanism within the kernel. The structure is typical for these sorts of checks: a condition (`cond`), an environment (`env`), and a format string with potential arguments (`fmt, ...`). The key thing is the inclusion of `WARN_ONCE`, which immediately flags this as a "one-time warning" or a "soft error". This indicates that it's designed to point out a potential, but not necessarily fatal, issue only once. It's a way to highlight a suspicious behavior without immediately panicking the system. That's good for debugging. Now, I need to see what `WARN_ONCE` expands to, as that will provide more context. Knowing the environment variable will be useful too. |
| 16/2 | 2026/09/24 01:46 | tool | grepper |
0mError:no matches Args: {
"Expression": "#define verifier_bug_if",
"PathPrefix": "kernel/bpf/"
}Results: {
"error": "no matches"
} |
| 17/2 | 2026/09/24 01:46 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 74464 output: 27 thoughts: 14 Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience: **Hunting Down That Pesky `verifier_bug_if`** Okay, so I need to track down where this `verifier_bug_if` macro is defined. Based on convention and my experience with the codebase, my first instinct is to look in the `include/` directory. That's usually where the core definitions and headers live. I'm assuming it's likely a macro used for some kind of verification or error handling within the kernel or related modules. My working hypothesis is that it's probably part of some preprocessor definition related to memory safety, security checks, or perhaps even debugging aids specific to a verification framework. Let's start the search there. |
| 18/2 | 2026/09/24 01:46 | tool | grepper |
0mArgs:{
"Expression": "verifier_bug_if",
"PathPrefix": "include/"
}Results: {
"Output": "include/linux/bpf_verifier.h=1139=__printf(3, 4) void verbose_linfo(struct bpf_verifier_env *env,\n--\ninclude/linux/bpf_verifier.h-1142-\ninclude/linux/bpf_verifier.h:1143:#define verifier_bug_if(cond, env, fmt, args...)\t\t\t\t\t\t\\\ninclude/linux/bpf_verifier.h-1144-\t({\t\t\t\t\t\t\t\t\t\t\t\\\n"
} |
| 19/2 | 2026/09/24 01:46 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 74792 output: 39 thoughts: 24 Thoughts: Okay, let's dive into this. Here's my thought process on the matter: **Delving into the BPF Verifier Header** Alright, I need to understand this specific section of the BPF verifier. I'm going to navigate to `include/linux/bpf_verifier.h` around line 1143. That should give me the relevant context for what's going on. I'm going to look for the key data structures and functions that are defined or used around that area. I'll likely be scanning for things that are related to instruction validation, register management, or perhaps error handling, depending on the surrounding code. I'll be focused on understanding the logic flow at this point. I'll be very interested in the specifics. |
| 20/2 | 2026/09/24 01:46 | tool | read-file |
0mArgs:{
"File": "include/linux/bpf_verifier.h",
"FirstLine": 1140,
"LineCount": 20
}Results: {
"Contents": "1140:\t\t\t\t\t u32 insn_off,\n1141:\t\t\t\t\t const char *prefix_fmt, ...);\n1142:\t\n1143:\t#define verifier_bug_if(cond, env, fmt, args...)\t\t\t\t\t\t\\\n1144:\t\t({\t\t\t\t\t\t\t\t\t\t\t\\\n1145:\t\t\tbool __cond = (cond);\t\t\t\t\t\t\t\t\\\n1146:\t\t\tif (unlikely(__cond))\t\t\t\t\t\t\t\t\\\n1147:\t\t\t\tverifier_bug(env, fmt \" (\" #cond \")\", ##args);\t\t\t\t\\\n1148:\t\t\t(__cond);\t\t\t\t\t\t\t\t\t\\\n1149:\t\t})\n1150:\t#define verifier_bug(env, fmt, args...)\t\t\t\t\t\t\t\t\\\n1151:\t\t({\t\t\t\t\t\t\t\t\t\t\t\\\n1152:\t\t\tBPF_WARN_ONCE(1, \"verifier bug: \" fmt \"\\n\", ##args);\t\t\t\t\\\n1153:\t\t\tbpf_log(\u0026env-\u003elog, \"verifier bug: \" fmt \"\\n\", ##args);\t\t\t\t\\\n1154:\t\t})\n1155:\t\n1156:\tstatic inline void mark_prune_point(struct bpf_verifier_env *env, int idx)\n1157:\t{\n1158:\t\tenv-\u003einsn_aux_data[idx].prune_point = true;\n1159:\t}\n"
} |
| 21/2 | 2026/09/24 01:46 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 75365 output: 274 thoughts: 1689 Thoughts: **Analysis of Patch Requirements for Fuzzing** Alright, let's break this down. This patch introduces `verifier_bug_if` calls, essentially kernel warnings/assertions, within `jit_subprogs` and `check_func_ptr_read`. The documentation explicitly states that new or modified assertions in reachable code *must* be fuzzed. This seems like a no-brainer given the potential for unexpected sequences to violate these new invariants. We're also dealing with modifications in BPF, which is flagged as a core subsystem requiring fuzzing attention, so that confirms the requirement. Now, let's figure out what we need configured to actually exercise this code. `BPF_SYSCALL` and `BPF_JIT` are obviously critical. Digging into the Kconfig, `BPF_JIT` depends on `BPF` and either `HAVE_CBPF_JIT` or `HAVE_EBPF_JIT`, and selects `EXECMEM`. We also need to consider if `BPF_JIT_ALWAYS_ON` is needed, looking at `bpf_jit_comp.c`, and it appears that `CONFIG_BPF_JIT=y` enables `bpf_jit_comp.c`, and by extension the `callx` support. The JIT requires `jit_requested` which is enabled by `ebpf_jit_enabled()`, which in turn, is controlled by `bpf_jit_enable`. If `CONFIG_BPF_JIT_ALWAYS_ON` is set, or if it's set via sysctl, then `bpf_jit_enable` will be set. `BPF_SYSCALL` and `BPF_JIT` are definitely required. We can now look at specific functions implementing the new logic. The patch introduces several new functions related to the function pointer and indirect call features. I need to enumerate them: * `check_func_callx` (indirect call handling) * `resolve_func_ptrs` (read-only map function pointer scan) * `check_func_ptr_read` (function pointer memory read validation) * `bpf_map_claim` (function pointer map ownership tracking) * `record_callx_edge` (callgraph edge recording) * `bpf_map_range_func_ptrs` (function pointer map range search) * `bpf_adjust_func_ptrs` (adjusting function pointers after instruction patching) * `bpf_adjust_func_ptrs_after_remove` (adjusts func pointers after instruction removal) * `jit_subprogs` (patches function pointers in frozen maps and JITs subprograms) * `emit_indirect_call` (x86 JIT indirect call emission) * `bpf_jit_supports_callx` (arch-specific support check) * `visit_func_ptrs_insn` (CFG analysis). We need to confirm none of these are ubiquitous. They are all specific to the `callx` support and function pointer tracking. The static functions are suitable candidates. The non-static are OK too, but we may want to try to constrain the fuzzer to the entrypoints. The final checklist for `set-results` looks like: - `WorthFuzzing`: True, given the introduction of new invariants and core subsystem modification. - `Reasoning`: The patch adds support for indirect function calls (`callx`), function pointer tracking in frozen maps, recursion and stack depth checks, new defensive assertions, and JIT emission of indirect calls. - `FocusSymbols`: Listed functions above. - `EnableConfigs`: `["BPF_SYSCALL", "BPF_JIT"]` Looks like we're good to go. Let's call `set-results`. |
| 22/2 | 2026/09/24 01:46 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"BPF_SYSCALL",
"BPF_JIT"
],
"FocusSymbols": [
"check_func_callx",
"resolve_func_ptrs",
"check_func_ptr_read",
"bpf_map_claim",
"bpf_map_range_func_ptrs",
"bpf_adjust_func_ptrs",
"bpf_adjust_func_ptrs_after_remove",
"emit_indirect_call",
"bpf_jit_supports_callx",
"jit_subprogs"
],
"Reasoning": "The patch series implements support for BPF indirect calls ('callx' / BPF_CALL | BPF_X) of static subprograms and function pointer resolution in frozen read-only maps. This introduces substantial modifications across the BPF verifier (CFG traversal, state pruning/liveness, stack depth and recursion checks, func pointer bounds checks) and the x86 BPF JIT compiler (indirect call emission and retpoline handling), along with several new runtime defensive assertions (verifier_bug_if). These changes modify reachable core subsystem code easily exercised via the bpf() syscall.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"BPF_SYSCALL",
"BPF_JIT"
],
"FocusSymbols": [
"check_func_callx",
"resolve_func_ptrs",
"check_func_ptr_read",
"bpf_map_claim",
"bpf_map_range_func_ptrs",
"bpf_adjust_func_ptrs",
"bpf_adjust_func_ptrs_after_remove",
"emit_indirect_call",
"bpf_jit_supports_callx",
"jit_subprogs"
],
"Reasoning": "The patch series implements support for BPF indirect calls ('callx' / BPF_CALL | BPF_X) of static subprograms and function pointer resolution in frozen read-only maps. This introduces substantial modifications across the BPF verifier (CFG traversal, state pruning/liveness, stack depth and recursion checks, func pointer bounds checks) and the x86 BPF JIT compiler (indirect call emission and retpoline handling), along with several new runtime defensive assertions (verifier_bug_if). These changes modify reachable core subsystem code easily exercised via the bpf() syscall.",
"WorthFuzzing": true
} |
| 23/1 | 2026/09/24 01:46 | action | If |
1mArgs:{
"WorthFuzzing": true
}Results: null |
| 24/2 | 2026/09/24 01:46 | agent | kmsan-evaluator |
1mModel:core Args: null Results: {
"KMSANReasoning": "The patch series implements support for indirect calls (`callx` instruction: `BPF_JMP | BPF_CALL | BPF_X`) in the BPF verifier, x86/arm64 JIT compilers, and user-space libbpf tooling.\n\nAnalysis of memory initialization and safety:\n1. Kernel structures and memory allocations:\n - `struct bpf_verifier_env` fields (`func_ptrs`, `func_ptr_cnt`, `has_callx`, `callx_edges`) are zero-initialized via `kzalloc`.\n - `callx_edges` is allocated using `kvcalloc()`, ensuring all bits are zeroed.\n - `func_ptrs` elements are allocated via `kvrealloc()` and each populated element up to `func_ptr_cnt` is initialized with a compound literal (`struct bpf_func_ptr`), zeroing any omitted fields (such as `.used`).\n - `bpf_subprog_call_depth_info` array (`dinfo`) is allocated using `kvzalloc_objs()`.\n - The new field `user` in `struct bpf_map` is initialized to 0 during map allocation via `__bpf_map_area_alloc()` (`__GFP_ZERO`).\n2. No kernel-to-user info-leak of uninitialized memory:\n - No new `copy_to_user`, `put_user`, socket option, netlink, or ioctl copy-out paths are introduced or modified.\n3. No conditional logic or branching depends on uninitialized state.\n4. The risks associated with this patch primarily involve verifier correctness, call graph cycle detection, and JIT assembly generation, all of which fall under functional verification, UBSAN, and KASAN (e.g., out-of-bounds or use-after-free).\n\nSince there are no uninitialized memory usage risks or uninitialized memory info-leaks, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 47a35dcdf09a938b11851b9703bb51549ad4f665
Author: syz-cluster <triage@syzkaller.com>
Date: Thu Sep 24 01:45:15 2026 +0000
syz-cluster: applied patch under review
diff --git a/Documentation/bpf/clang-notes.rst b/Documentation/bpf/clang-notes.rst
index 2c872a1ee08e9..3ccc7b09d19e6 100644
--- a/Documentation/bpf/clang-notes.rst
+++ b/Documentation/bpf/clang-notes.rst
@@ -23,8 +23,11 @@ For CPU versions prior to 3, Clang v7.0 and later can enable ``BPF_ALU`` support
Jump instructions
=================
-If ``-O0`` is used, Clang will generate the ``BPF_CALL | BPF_X | BPF_JMP`` (0x8d)
-instruction, which is not supported by the Linux kernel verifier.
+Clang generates the ``BPF_CALL | BPF_X | BPF_JMP`` (0x8d) instruction for calls
+through a function pointer. The Linux kernel verifier accepts it only when it
+can prove that the register holds the address of a static BPF function, see
+Documentation/bpf/linux-notes.rst. If ``-O0`` is used, Clang will generate this
+instruction for helper calls as well, which is not supported.
Atomic operations
=================
diff --git a/Documentation/bpf/linux-notes.rst b/Documentation/bpf/linux-notes.rst
index 00d2693de025a..6c036b54a29f2 100644
--- a/Documentation/bpf/linux-notes.rst
+++ b/Documentation/bpf/linux-notes.rst
@@ -15,10 +15,33 @@ Byte swap instructions
Jump instructions
=================
-``BPF_CALL | BPF_X | BPF_JMP`` (0x8d), where the helper function
-integer would be read from a specified register, is not currently supported
-by the verifier. Any programs with this instruction will fail to load
-until such support is added.
+``BPF_CALL | BPF_X | BPF_JMP`` (0x8d), ``callx dst``, performs an indirect
+call of a BPF function whose address is held in the ``dst`` register. The
+``src``, ``offset`` and ``imm`` fields are reserved and must be zero.
+
+The address of a BPF function gets into a register in one of two ways:
+
+* it is loaded by a 64-bit immediate instruction with ``src`` =
+ ``BPF_PSEUDO_FUNC``;
+* it is read, with a 64-bit load, from a frozen read-only array map, that no
+ other program uses, that holds its read-only data: tables of functions,
+ structures of operations, vtables, where pointers to functions may be mixed
+ with other data. In the map a pointer to a function is the offset in bytes
+ of its first instruction in the program, and that is how the verifier
+ recognizes it. It is replaced with the address of the function when
+ the program is loaded. The program reads it from there, which requires
+ ``CAP_PERFMON``.
+
+In both cases only static functions can be referenced. Therefore all functions
+that can be called indirectly are known to the verifier before it starts to
+analyze the program, and ``callx`` is verified as a direct call of every
+function that ``dst`` may point to at that instruction. The same rules apply:
+the calls can not be recursive, and the depth of the call chain and its
+combined stack size are limited.
+
+Calling helper or kernel functions through a register, indirect calls of global
+functions, and tail calls in functions that are called via ``callx`` are not
+supported. ``callx`` requires the BPF JIT.
Maps
====
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 6c04fee468766..9544b2f483e59 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -1789,6 +1789,17 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
emit(A64_MOV(1, r0, A64_R(0)), ctx);
break;
}
+ /* indirect call of a bpf subprog, dst holds its address */
+ case BPF_JMP | BPF_CALL | BPF_X:
+ /*
+ * It's the same as a direct call of a subprog that is out of
+ * range of BL: the subprog starts with BTI JC, the arguments
+ * are in place, and the registers that hold the tail call
+ * counter and the private stack are callee saved.
+ */
+ emit(A64_BLR(dst), ctx);
+ emit(A64_MOV(1, bpf2a64[BPF_REG_0], A64_R(0)), ctx);
+ break;
/* tail call */
case BPF_JMP | BPF_TAIL_CALL:
if (emit_bpf_tail_call(ctx))
@@ -2485,6 +2496,11 @@ bool bpf_jit_supports_subprog_tailcalls(void)
return true;
}
+bool bpf_jit_supports_callx(void)
+{
+ return true;
+}
+
static void invoke_bpf_prog(struct jit_ctx *ctx, struct bpf_tramp_node *node,
int bargs_off, int retval_off, int run_ctx_off,
bool save_ret)
diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index d4a980140b48d..9fbef7504e51a 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -749,6 +749,46 @@ static void emit_indirect_jump(u8 **pprog, int bpf_reg, u8 *ip)
*pprog = prog;
}
+static void __emit_indirect_call(u8 **pprog, int reg, bool ereg)
+{
+ u8 *prog = *pprog;
+
+ if (ereg)
+ EMIT1(0x41);
+
+ EMIT2(0xFF, 0xD0 + reg);
+
+ *pprog = prog;
+}
+
+/* call *bpf_reg */
+static int emit_indirect_call(u8 **pprog, int bpf_reg, u8 *ip)
+{
+ u8 *prog = *pprog;
+ int reg = reg2hex[bpf_reg];
+ bool ereg = is_ereg(bpf_reg);
+ int err = 0;
+
+ if (cpu_feature_enabled(X86_FEATURE_INDIRECT_THUNK_ITS)) {
+ OPTIMIZER_HIDE_VAR(reg);
+ err = emit_call(&prog, its_static_thunk(reg + 8*ereg), ip);
+ } else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE_LFENCE)) {
+ EMIT_LFENCE();
+ __emit_indirect_call(&prog, reg, ereg);
+ } else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE)) {
+ OPTIMIZER_HIDE_VAR(reg);
+ if (cpu_feature_enabled(X86_FEATURE_CALL_DEPTH))
+ err = emit_call(&prog, &__x86_indirect_call_thunk_array[reg + 8*ereg], ip);
+ else
+ err = emit_call(&prog, &__x86_indirect_thunk_array[reg + 8*ereg], ip);
+ } else {
+ __emit_indirect_call(&prog, reg, ereg);
+ }
+
+ *pprog = prog;
+ return err;
+}
+
static void emit_return(u8 **pprog, u8 *ip)
{
u8 *prog = *pprog;
@@ -2941,6 +2981,24 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
break;
}
+ /* callx: call of a bpf subprog whose address is in dst_reg */
+ case BPF_JMP | BPF_CALL | BPF_X:
+ /*
+ * The verifier makes sure that callees of callx are
+ * not tail call reachable, hence unlike a direct call
+ * of a subprog there is no need to pass
+ * tail_call_cnt_ptr in rax.
+ */
+ if (priv_frame_ptr) {
+ push_r9(&prog);
+ ip += 2;
+ }
+ if (emit_indirect_call(&prog, insn->dst_reg, ip))
+ return -EINVAL;
+ if (priv_frame_ptr)
+ pop_r9(&prog);
+ break;
+
case BPF_JMP | BPF_TAIL_CALL:
if (imm32)
emit_bpf_tail_call_direct(bpf_prog,
@@ -4414,6 +4472,16 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
return prog;
}
+bool bpf_jit_supports_callx(void)
+{
+ /*
+ * FineIBT poisons ENDBR at the entry of a JITed function and expects
+ * indirect callers to go through the CFI preamble instead.
+ * callx doesn't do that yet.
+ */
+ return cfi_mode != CFI_FINEIBT;
+}
+
bool bpf_jit_supports_kfunc_call(void)
{
return true;
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index fd22db8bc6c50..7747c5fc290fc 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -342,8 +342,18 @@ struct bpf_map {
s64 __percpu *elem_count;
u64 cookie; /* write-once */
char *excl_prog_sha;
+ /*
+ * Which programs use the map, see bpf_map_claim(): 0 - none so far,
+ * aux of the program - only that one, the same with BPF_MAP_USER_PATCHED
+ * set - only that one and it stored the addresses of its functions into
+ * the map, BPF_MAP_USER_MANY - more than one.
+ */
+ unsigned long user;
};
+#define BPF_MAP_USER_MANY 1UL
+#define BPF_MAP_USER_PATCHED 1UL
+
static inline const char *btf_field_type_name(enum btf_field_type type)
{
switch (type) {
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 92f528c456052..38a4ba50669a2 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -786,6 +786,22 @@ int bpf_log_attr_finalize(struct bpf_log_attr *attr, struct bpf_verifier_log *lo
#define BPF_MAX_SUBPROGS 256
+/*
+ * A pointer to a static subprog in the value of a frozen read-only array map:
+ * a 64-bit value that is the offset in bytes of the first instruction of
+ * the subprog in the program.
+ */
+struct bpf_func_ptr {
+ struct bpf_map *map;
+ u32 map_off; /* offset of the pointer in the value of the map */
+ u32 orig_off; /* what the map has: the first instruction of the subprog */
+ u32 xlated_off; /* the same after instructions were patched and removed */
+ bool used; /* the program reads the pointer */
+};
+
+/* the subprog that a bpf_func_ptr pointed to was removed as dead code */
+#define BPF_FUNC_PTR_DELETED ((u32)-1)
+
struct bpf_subprog_arg_info {
enum bpf_arg_type arg_type;
union {
@@ -965,6 +981,20 @@ struct bpf_verifier_env {
struct bpf_subprog_info subprog_info[BPF_MAX_SUBPROGS + 2]; /* max + 2 for the fake and exception subprogs */
/* subprog indices sorted in topological order: leaves first, callers last */
int subprog_topo_order[BPF_MAX_SUBPROGS + 2];
+ /*
+ * Pointers to static subprogs found in frozen read-only maps of the
+ * program, see resolve_func_ptrs(). Sorted by map and map_off.
+ */
+ struct bpf_func_ptr *func_ptrs;
+ u32 func_ptr_cnt;
+ bool has_callx;
+ /*
+ * Call graph edges created by callx instructions. A bitmap of
+ * subprog_cnt * subprog_cnt bits, where bit (caller * subprog_cnt + callee)
+ * is set when the main verification pass sees 'caller' calling 'callee'
+ * via callx. Allocated when the first such edge is recorded.
+ */
+ unsigned long *callx_edges;
union {
struct bpf_idmap idmap_scratch;
struct bpf_idset idset_scratch;
@@ -1089,6 +1119,12 @@ static inline bool bpf_pseudo_kfunc_call(const struct bpf_insn *insn)
insn->src_reg == BPF_PSEUDO_KFUNC_CALL;
}
+/* callx: indirect call of a bpf subprog whose address is in insn->dst_reg */
+static inline bool bpf_is_callx(const struct bpf_insn *insn)
+{
+ return insn->code == (BPF_JMP | BPF_CALL | BPF_X);
+}
+
__printf(2, 0) void bpf_verifier_vlog(struct bpf_verifier_log *log,
const char *fmt, va_list args);
__printf(2, 3) void bpf_verifier_log_write(struct bpf_verifier_env *env,
@@ -1311,6 +1347,13 @@ static inline bool bt_is_frame_slot_set(struct backtrack_state *bt, u32 frame, u
}
bool bpf_map_is_rdonly(const struct bpf_map *map);
+struct bpf_func_ptr *bpf_map_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map, u32 *cnt);
+struct bpf_func_ptr *bpf_map_range_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map,
+ u64 off, u64 size, u32 *cnt);
+void bpf_adjust_func_ptrs(struct bpf_verifier_env *env, u32 off, u32 len);
+void bpf_adjust_func_ptrs_after_remove(struct bpf_verifier_env *env, u32 off, u32 len);
int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,
bool is_ldsx);
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 422284b4fa96f..4f0662e428970 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1237,6 +1237,7 @@ bool bpf_jit_inlines_helper_call(s32 imm);
bool bpf_jit_supports_subprog_tailcalls(void);
bool bpf_jit_supports_percpu_insn(void);
bool bpf_jit_supports_kfunc_call(void);
+bool bpf_jit_supports_callx(void);
bool bpf_jit_supports_kfunc_ret_reg_pair(void);
bool bpf_jit_supports_stack_args(void);
bool bpf_jit_supports_arena_args(void);
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 507a366dffa47..4da99dec08184 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -406,15 +406,18 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
if (class == BPF_STX)
bt_set_reg(bt, sreg);
} else if (class == BPF_JMP || class == BPF_JMP32) {
- if (bpf_pseudo_call(insn)) {
- int subprog_insn_idx, subprog;
+ if (bpf_pseudo_call(insn) || bpf_is_callx(insn)) {
+ int subprog_insn_idx, subprog = -1;
- subprog_insn_idx = idx + insn->imm + 1;
- subprog = bpf_find_subprog(env, subprog_insn_idx);
- if (subprog < 0)
- return -EFAULT;
+ if (bpf_pseudo_call(insn)) {
+ subprog_insn_idx = idx + insn->imm + 1;
+ subprog = bpf_find_subprog(env, subprog_insn_idx);
+ if (subprog < 0)
+ return -EFAULT;
+ }
- if (bpf_subprog_is_global(env, subprog)) {
+ /* callx calls static subprogs only */
+ if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
/* check that jump history doesn't have any
* extra instructions from subprog; the next
* instruction after call to global subprog
@@ -536,7 +539,8 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
* never do that.
*/
from_subprog_call = subseq_idx - 1 >= 0 &&
- bpf_pseudo_call(&env->prog->insnsi[subseq_idx - 1]);
+ (bpf_pseudo_call(&env->prog->insnsi[subseq_idx - 1]) ||
+ bpf_is_callx(&env->prog->insnsi[subseq_idx - 1]));
r0_precise = from_subprog_call && bt_is_reg_set(bt, BPF_REG_0);
r2_precise = from_subprog_call && bt_is_reg_set(bt, BPF_REG_2);
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 842c7d1eabccc..a068c191409a4 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -421,6 +421,75 @@ static int visit_gotox_insn(int t, struct bpf_verifier_env *env)
return keep_exploring ? KEEP_EXPLORING : DONE_EXPLORING;
}
+/*
+ * Return pointers to functions in the read-only map that ld_imm64 instruction
+ * 't' loads the address of, or of its value, if there are any.
+ */
+static struct bpf_func_ptr *insn_func_ptrs(struct bpf_verifier_env *env, int t, u32 *cnt)
+{
+ struct bpf_insn *insn = &env->prog->insnsi[t];
+
+ *cnt = 0;
+ if (!env->func_ptr_cnt || !bpf_is_ldimm64(insn))
+ return NULL;
+ if (insn->src_reg != BPF_PSEUDO_MAP_VALUE &&
+ insn->src_reg != BPF_PSEUDO_MAP_IDX_VALUE &&
+ insn->src_reg != BPF_PSEUDO_MAP_FD &&
+ insn->src_reg != BPF_PSEUDO_MAP_IDX)
+ return NULL;
+
+ return bpf_map_func_ptrs(env, env->used_maps[env->insn_aux_data[t].map_index], cnt);
+}
+
+/*
+ * ld_imm64 that loads the address of a map that has pointers to functions
+ * is similar to ld_imm64 with BPF_PSEUDO_FUNC that loads the address of one
+ * function: any of them may be read from the map and called via callx later.
+ * Treat it as a call of all of them.
+ */
+static int visit_func_ptrs_insn(int t, struct bpf_verifier_env *env,
+ struct bpf_func_ptr *ptrs, u32 cnt)
+{
+ int *insn_stack = env->cfg.insn_stack;
+ int *insn_state = env->cfg.insn_state;
+ bool keep_exploring = false;
+ int ret, w;
+ u32 i;
+
+ ret = push_insn(t, t + 2, FALLTHROUGH, env);
+ if (ret)
+ return ret;
+
+ mark_prune_point(env, t);
+ for (i = 0; i < cnt; i++) {
+ w = ptrs[i].xlated_off;
+
+ /*
+ * This function is called until all functions are explored,
+ * so the effects are complete in the end.
+ */
+ merge_callee_effects(env, t, w);
+
+ /* the same marks as push_insn() leaves on a branch target */
+ mark_prune_point(env, w);
+ mark_jmp_point(env, w);
+ mark_jump_target(env, w);
+
+ /* EXPLORED || DISCOVERED */
+ if (insn_state[w])
+ continue;
+
+ if (env->cfg.cur_stack >= env->prog->len)
+ return -E2BIG;
+
+ insn_stack[env->cfg.cur_stack++] = w;
+ insn_state[w] |= DISCOVERED;
+ keep_exploring = true;
+ }
+
+ return keep_exploring ? KEEP_EXPLORING : DONE_EXPLORING;
+}
+
/*
* Instructions that can abnormally return from a subprog (tail_call
* upon success, ld_{abs,ind} upon load failure) have a hidden exit
@@ -453,11 +522,17 @@ static int visit_abnormal_return_insn(struct bpf_verifier_env *env, int t)
static int visit_insn(int t, struct bpf_verifier_env *env)
{
struct bpf_insn *insns = env->prog->insnsi, *insn = &insns[t];
+ struct bpf_func_ptr *ptrs;
int ret, off, insn_sz;
+ u32 cnt;
if (bpf_pseudo_func(insn))
return visit_func_call_insn(t, insns, env, true);
+ ptrs = insn_func_ptrs(env, t, &cnt);
+ if (ptrs)
+ return visit_func_ptrs_insn(t, env, ptrs, cnt);
+
/* All non-branch instructions have a single fall-through edge. */
if (BPF_CLASS(insn->code) != BPF_JMP &&
BPF_CLASS(insn->code) != BPF_JMP32) {
diff --git a/kernel/bpf/const_fold.c b/kernel/bpf/const_fold.c
index 7f1b30059cc87..fea639f62b3f2 100644
--- a/kernel/bpf/const_fold.c
+++ b/kernel/bpf/const_fold.c
@@ -180,9 +180,17 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
bool is_ldsx = mode == BPF_MEMSX;
int off = src->val + insn->off;
u64 val = 0;
+ u32 cnt;
+ /*
+ * Values of insn_array map are addresses of jitted instructions,
+ * which are not known until the program is jitted.
+ */
if (!bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
+ map->map_type == BPF_MAP_TYPE_INSN_ARRAY ||
off < 0 || off + size > map->value_size ||
+ /* so are the addresses of functions that the map points to */
+ bpf_map_range_func_ptrs(env, map, off, size, &cnt) ||
bpf_map_direct_read(map, off, size, &val, is_ldsx)) {
*dst = unknown;
break;
@@ -191,7 +199,8 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
dst->val = val;
break;
case BPF_JMP:
- if (opcode != BPF_CALL)
+ /* both 'call imm' and 'callx reg' clobber caller saved registers */
+ if (BPF_OP(insn->code) != BPF_CALL)
break;
process_call:
for (r = BPF_REG_0; r <= BPF_REG_5; r++)
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 227211166dccf..273f74068068c 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -1831,6 +1831,7 @@ bool bpf_opcode_in_insntable(u8 code)
[BPF_LD | BPF_IND | BPF_H] = true,
[BPF_LD | BPF_IND | BPF_W] = true,
[BPF_JMP | BPF_JA | BPF_X] = true,
+ [BPF_JMP | BPF_CALL | BPF_X] = true,
[BPF_JMP | BPF_JCOND] = true,
};
#undef BPF_INSN_3_TBL
@@ -3028,6 +3029,11 @@ void __bpf_free_used_maps(struct bpf_prog_aux *aux,
map->ops->map_poke_untrack(map, aux);
if (sleepable)
atomic64_dec(&map->sleepable_refcnt);
+ /*
+ * The program that didn't load is not a user of the map. libbpf
+ * loads the program again to get the log of the verifier.
+ */
+ cmpxchg(&map->user, (unsigned long)aux, 0);
bpf_map_put(map);
}
}
@@ -3287,6 +3293,12 @@ bool __weak bpf_jit_supports_kfunc_call(void)
return false;
}
+/* Return TRUE if the JIT backend supports callx (indirect call) instruction. */
+bool __weak bpf_jit_supports_callx(void)
+{
+ return false;
+}
+
bool __weak bpf_jit_supports_kfunc_ret_reg_pair(void)
{
return false;
diff --git a/kernel/bpf/disasm.c b/kernel/bpf/disasm.c
index 3ce8d74b0e400..36d3228d77454 100644
--- a/kernel/bpf/disasm.c
+++ b/kernel/bpf/disasm.c
@@ -350,7 +350,10 @@ void print_bpf_insn(const struct bpf_insn_cbs *cbs,
if (opcode == BPF_CALL) {
char tmp[64];
- if (insn->src_reg == BPF_PSEUDO_CALL) {
+ if (BPF_SRC(insn->code) == BPF_X) {
+ verbose(cbs->private_data, "(%02x) callx r%d",
+ insn->code, insn->dst_reg);
+ } else if (insn->src_reg == BPF_PSEUDO_CALL) {
verbose(cbs->private_data, "(%02x) call pc%s",
insn->code,
__func_get_name(cbs, insn,
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 2add8001c3ec3..e568b9b790b5c 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -361,6 +361,7 @@ struct bpf_prog *bpf_patch_insn_data(struct bpf_verifier_env *env, u32 off,
adjust_insn_aux_data(env, new_prog, off, len, &original_insn);
adjust_subprog_starts(env, off, len);
adjust_insn_arrays(env, off, len);
+ bpf_adjust_func_ptrs(env, off, len);
adjust_poke_descs(new_prog, off, len);
return new_prog;
}
@@ -559,6 +560,9 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
if (err)
return err;
+ /* before subprogs are adjusted, since it looks at them */
+ bpf_adjust_func_ptrs_after_remove(env, off, cnt);
+
err = adjust_subprog_starts_after_remove(env, off, cnt);
if (err)
return err;
@@ -1285,6 +1289,57 @@ static int jit_subprogs(struct bpf_verifier_env *env)
cond_resched();
}
+ /*
+ * The addresses of all functions are final. Replace the offsets of
+ * functions with them in the maps of the program, see
+ * resolve_func_ptrs(). The program must be the only user of such map.
+ * From now on no other program can use it, see bpf_map_claim().
+ */
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ struct bpf_func_ptr *ptr = &env->func_ptrs[i];
+ unsigned long me = (unsigned long)prog->aux;
+ u64 addr, old, new = 0;
+
+ /* pointers are sorted by map */
+ if ((!i || ptr->map != ptr[-1].map) &&
+ cmpxchg(&ptr->map->user, me, me | BPF_MAP_USER_PATCHED) != me) {
+ verbose(env, "map '%s' is used by another program\n", ptr->map->name);
+ err = -EBUSY;
+ goto out_free;
+ }
+
+ /* it's the address of the value of the map whatever the offset is */
+ err = ptr->map->ops->map_direct_value_addr(ptr->map, &addr, 0);
+ if (verifier_bug_if(err, env, "no value of map '%s'", ptr->map->name)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ addr += ptr->map_off;
+
+ if (ptr->xlated_off != BPF_FUNC_PTR_DELETED) {
+ subprog = bpf_find_subprog(env, ptr->xlated_off);
+ if (verifier_bug_if(subprog <= 0, env, "no function at insn %u",
+ ptr->xlated_off)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ new = (unsigned long)func[subprog]->bpf_func;
+ } else if (verifier_bug_if(ptr->used, env, "function of map '%s' offset %u is removed",
+ ptr->map->name, ptr->map_off)) {
+ /* the program that reads the pointer might call the function */
+ err = -EFAULT;
+ goto out_free;
+ }
+ /* else the function is dead code, nothing calls it, the pointer is NULL */
+
+ old = (u64)ptr->orig_off * sizeof(struct bpf_insn);
+ if (verifier_bug_if(cmpxchg64((u64 *)(unsigned long)addr, old, new) != old, env,
+ "map '%s' offset %u changed", ptr->map->name, ptr->map_off)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ }
+
/*
* Cleanup func[i]->aux fields which aren't required
* or can become invalid in future
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 44ecdc5b4ec2d..5aa2f68d92b37 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -356,12 +356,25 @@ int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_
return 0;
}
+/*
+ * Stack accesses of callbacks and of callx callees are not tracked by
+ * func instances keyed by the @callsite. Callbacks might be called several
+ * times and the callee of callx is not known when stack liveness is computed.
+ * In both cases stack slots of the outer frames that might be read by the
+ * callee are accounted as read by the @callsite instruction itself.
+ */
+static bool callee_stack_access_at_callsite(struct bpf_verifier_env *env, u32 callsite)
+{
+ return bpf_calls_callback(env, callsite) ||
+ bpf_is_callx(&env->prog->insnsi[callsite]);
+}
+
bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_spi)
{
/*
* Slot is alive if it is read before q->insn_idx in current func instance,
* or if for some outer func instance:
- * - alive before callsite if callsite calls callback, otherwise
+ * - alive before callsite if callsite calls callback or is callx, otherwise
* - alive after callsite
*/
struct live_stack_query *q = &env->liveness->live_stack_query;
@@ -394,7 +407,7 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
/* Get callsite from verifier state, not from instance callchain */
callsite = q->callsites[i];
- alive = bpf_calls_callback(env, callsite)
+ alive = callee_stack_access_at_callsite(env, callsite)
? is_live_before(instance, callsite, rel, half_spi)
: is_live_before(instance, callsite + 1, rel, half_spi);
if (alive)
@@ -1439,7 +1452,15 @@ static int record_call_access(struct bpf_verifier_env *env,
if (bpf_pseudo_call(insn))
return 0;
- if (bpf_get_call_summary(env, insn, &cs))
+ if (bpf_is_callx(insn))
+ /*
+ * The callee is not known statically. Assume that all arg
+ * slots are passed and let record_arg_access() conservatively
+ * mark the stack of all frames as read if any of them is
+ * derived from a frame pointer.
+ */
+ arg_slot_cnt = MAX_BPF_FUNC_REG_ARGS + MAX_STACK_ARG_SLOTS;
+ else if (bpf_get_call_summary(env, insn, &cs))
arg_slot_cnt = cs.arg_slot_cnt;
for (r = BPF_REG_1; r < BPF_REG_1 + min(arg_slot_cnt, MAX_BPF_FUNC_REG_ARGS); r++) {
@@ -1533,7 +1554,8 @@ static void print_subprog_arg_access(struct bpf_verifier_env *env,
bool has_extra = false;
u8 cls = BPF_CLASS(insns[idx].code);
bool is_ldx_stx_call = cls == BPF_LDX || cls == BPF_STX ||
- insns[idx].code == (BPF_JMP | BPF_CALL);
+ insns[idx].code == (BPF_JMP | BPF_CALL) ||
+ bpf_is_callx(&insns[idx]);
verbose(env, "%3d: ", idx);
bpf_verbose_insn(env, &insns[idx]);
@@ -1722,7 +1744,7 @@ static int compute_subprog_args(struct bpf_verifier_env *env,
if (err)
goto err_free;
- if (insn->code == (BPF_JMP | BPF_CALL)) {
+ if (insn->code == (BPF_JMP | BPF_CALL) || bpf_is_callx(insn)) {
err = record_call_access(env, instance, at_in[i], idx);
if (err)
goto err_free;
@@ -2202,6 +2224,9 @@ static void compute_insn_live_regs(struct bpf_verifier_env *env,
use = GENMASK(min_t(u8, cs.arg_slot_cnt, MAX_BPF_FUNC_REG_ARGS), 1);
def = mask_widen(def);
use = mask_widen(use);
+ /* callx reads the address of the callee from dst_reg */
+ if (bpf_is_callx(insn))
+ use |= dst;
break;
default:
def = 0;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index d62c0f74cff5e..8507114cf690a 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -3018,6 +3018,22 @@ static int add_subprogs(struct bpf_verifier_env *env)
return ret;
}
+ /*
+ * func_info describes all functions of the program. Those that are
+ * referenced only from data, e.g. from a table of functions to be
+ * called via callx, are not seen by the loop above. They are possible
+ * callees that have to be known upfront as well.
+ */
+ if (env->bpf_capable) {
+ struct bpf_prog_aux *aux = env->prog->aux;
+
+ for (i = 1; i < aux->func_info_cnt; i++) {
+ ret = add_subprog(env, aux->func_info[i].insn_off);
+ if (ret < 0)
+ return ret;
+ }
+ }
+
ret = bpf_find_exception_callback_insn_off(env);
if (ret < 0)
return ret;
@@ -3099,6 +3115,8 @@ static int check_subprogs(struct bpf_verifier_env *env)
if (BPF_CLASS(code) == BPF_LD &&
(BPF_MODE(code) == BPF_ABS || BPF_MODE(code) == BPF_IND))
subprog[cur_subprog].has_ld_abs = true;
+ if (bpf_is_callx(&insn[i]))
+ env->has_callx = true;
if (BPF_CLASS(code) != BPF_JMP && BPF_CLASS(code) != BPF_JMP32)
goto next;
if (BPF_OP(code) == BPF_CALL)
@@ -3144,12 +3162,50 @@ static int check_subprogs(struct bpf_verifier_env *env)
return 0;
}
+/*
+ * The callee of callx is known to the main verification pass only, which
+ * records the 'caller' -> 'callee' edge of the call graph for the checks
+ * that follow it: absence of recursion and the maximum stack depth.
+ */
+static int record_callx_edge(struct bpf_verifier_env *env, int caller, int callee)
+{
+ u32 cnt = env->subprog_cnt;
+
+ if (!env->callx_edges) {
+ env->callx_edges = kvcalloc(BITS_TO_LONGS(cnt * cnt), sizeof(long),
+ GFP_KERNEL_ACCOUNT);
+ if (!env->callx_edges)
+ return -ENOMEM;
+ }
+ __set_bit(caller * cnt + callee, env->callx_edges);
+ return 0;
+}
+
+/*
+ * Return the first subprog with the number >= 'from' that 'caller' calls
+ * via callx, or -1 when there is none.
+ */
+static int next_callx_callee(struct bpf_verifier_env *env, int caller, int from)
+{
+ u32 cnt = env->subprog_cnt;
+ unsigned long bit, end = (caller + 1) * cnt;
+
+ if (!env->callx_edges || from >= cnt)
+ return -1;
+ bit = find_next_bit(env->callx_edges, end, caller * cnt + from);
+ return bit < end ? bit - caller * cnt : -1;
+}
+
/*
* Sort subprogs in topological order so that leaf subprogs come first and
* their callers come later. This is a DFS post-order traversal of the call
* graph. Scan only reachable instructions (those in the computed postorder) of
* the current subprog to discover callees (direct subprogs and sync
* callbacks).
+ *
+ * The callees of callx are not known before the main verification pass.
+ * When callx is used the sort is repeated after it with the recorded callx
+ * edges added to the call graph to reject recursion through indirect calls.
*/
static int sort_subprogs_topo(struct bpf_verifier_env *env)
{
@@ -3190,12 +3246,22 @@ static int sort_subprogs_topo(struct bpf_verifier_env *env)
int idx = insn_postorder[j];
int callee;
- if (!bpf_pseudo_call(&insn[idx]) && !bpf_pseudo_func(&insn[idx]))
+ if (bpf_is_callx(&insn[idx])) {
+ /* find a callee that is not explored yet */
+ callee = -1;
+ do {
+ callee = next_callx_callee(env, cur, callee + 1);
+ } while (callee >= 0 && color[callee] == 2);
+ if (callee < 0)
+ continue;
+ } else if (bpf_pseudo_call(&insn[idx]) || bpf_pseudo_func(&insn[idx])) {
+ callee = bpf_find_subprog(env, idx + insn[idx].imm + 1);
+ if (callee < 0) {
+ ret = -EFAULT;
+ goto out;
+ }
+ } else {
continue;
- callee = bpf_find_subprog(env, idx + insn[idx].imm + 1);
- if (callee < 0) {
- ret = -EFAULT;
- goto out;
}
if (color[callee] == 2)
continue;
@@ -5367,6 +5433,9 @@ struct bpf_subprog_call_depth_info {
int ret_insn; /* caller instruction where we return to. */
int caller; /* caller subprogram idx */
int frame; /* # of consecutive static call stack frames on top of stack */
+ int callx_insn; /* callx instruction whose callees are being walked */
+ int callx_next; /* next callee of callx_insn to walk */
+ bool via_callx; /* the subprogram is entered via callx */
};
/* starting from main bpf function walk all instructions of the function
@@ -5386,11 +5455,13 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
/* no caller idx */
dinfo[idx].caller = -1;
+ dinfo[idx].via_callx = false;
i = subprog[idx].start;
if (!priv_stack_supported)
subprog[idx].priv_stack_mode = NO_PRIV_STACK;
process_func:
+ dinfo[idx].callx_insn = -1;
if (subprog[idx].has_ld_abs) {
for (tmp = idx; tmp >= 0; tmp = dinfo[tmp].caller) {
if (subprog[tmp].is_cb) {
@@ -5486,31 +5557,64 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
return -EINVAL;
}
- if (!bpf_pseudo_call(insn + i) && !bpf_pseudo_func(insn + i))
+ if (bpf_is_callx(insn + i)) {
+ /*
+ * Walk the callees recorded by the main verification
+ * pass one by one, returning to this insn after each.
+ */
+ if (dinfo[idx].callx_insn != i) {
+ dinfo[idx].callx_insn = i;
+ dinfo[idx].callx_next = 0;
+ }
+ sidx = next_callx_callee(env, idx, dinfo[idx].callx_next);
+ if (sidx < 0)
+ continue;
+ dinfo[idx].callx_next = sidx + 1;
+ dinfo[idx].ret_insn = i;
+ next_insn = subprog[sidx].start;
+ } else if (bpf_pseudo_call(insn + i) || bpf_pseudo_func(insn + i)) {
+ /* find the callee */
+ next_insn = i + insn[i].imm + 1;
+ sidx = bpf_find_subprog(env, next_insn);
+ if (verifier_bug_if(sidx < 0, env, "callee not found at insn %d", next_insn))
+ return -EFAULT;
+ if (subprog[sidx].is_async_cb) {
+ /* async callbacks don't increase bpf prog stack size unless called directly */
+ if (!bpf_pseudo_call(insn + i))
+ continue;
+ if (subprog[sidx].is_exception_cb) {
+ verbose(env, "insn %d cannot call exception cb directly", i);
+ return -EINVAL;
+ }
+ }
+ /* remember insn to return to */
+ dinfo[idx].ret_insn = i + 1;
+ } else {
continue;
- /* remember insn and function to return to */
+ }
- /* find the callee */
- next_insn = i + insn[i].imm + 1;
- sidx = bpf_find_subprog(env, next_insn);
- if (verifier_bug_if(sidx < 0, env, "callee not found at insn %d", next_insn))
- return -EFAULT;
- if (subprog[sidx].is_async_cb) {
- /* async callbacks don't increase bpf prog stack size unless called directly */
- if (!bpf_pseudo_call(insn + i))
+ /*
+ * sort_subprogs_topo() tolerates cycles in the call graph that
+ * go through the address of a function being taken, since it
+ * doesn't know what it is taken for. Such cycle is a recursion
+ * unless it's an async callback, which are skipped above.
+ * The main verification pass limits the depth of the recursion,
+ * but it doesn't follow calls of global functions.
+ */
+ for (tmp = idx; tmp >= 0; tmp = dinfo[tmp].caller) {
+ if (tmp != sidx)
continue;
- if (subprog[sidx].is_exception_cb) {
- verbose(env, "insn %d cannot call exception cb directly", i);
- return -EINVAL;
- }
+ verbose(env, "recursive call from %s() to %s()\n",
+ bpf_subprog_name(env, idx), bpf_subprog_name(env, sidx));
+ return -EINVAL;
}
/* store caller info for after we return from callee */
dinfo[idx].frame = frame;
- dinfo[idx].ret_insn = i + 1;
/* push caller idx into callee's dinfo */
dinfo[sidx].caller = idx;
+ dinfo[sidx].via_callx = bpf_is_callx(insn + i);
i = next_insn;
@@ -5544,6 +5648,14 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
verbose(env, "tail_calls are not allowed in programs with stack args\n");
return -EINVAL;
}
+ /*
+ * JITs pass tail call counter in a register that is
+ * not available when the callee is called via callx.
+ */
+ if (dinfo[tmp].via_callx) {
+ verbose(env, "tail_calls are not allowed in functions called via callx\n");
+ return -EINVAL;
+ }
subprog[tmp].tail_call_reachable = true;
}
} else if (!idx && subprog[0].has_tail_call && subprog[0].stack_arg_cnt) {
@@ -5908,6 +6020,110 @@ int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,
return 0;
}
+static int cmp_func_ptrs(const void *_a, const void *_b)
+{
+ const struct bpf_func_ptr *a = _a, *b = _b;
+
+ if (a->map != b->map)
+ return a->map < b->map ? -1 : 1;
+ if (a->map_off != b->map_off)
+ return a->map_off < b->map_off ? -1 : 1;
+ return 0;
+}
+
+/* Find the first pointer to a function at or after 'off' in the value of 'map' */
+static u32 func_ptr_lower_bound(struct bpf_verifier_env *env, const struct bpf_map *map, u64 off)
+{
+ u32 l = 0, r = env->func_ptr_cnt, m;
+ struct bpf_func_ptr *p;
+
+ while (l < r) {
+ m = l + (r - l) / 2;
+ p = &env->func_ptrs[m];
+ if (p->map < map || (p->map == map && p->map_off < off))
+ l = m + 1;
+ else
+ r = m;
+ }
+ return l;
+}
+
+/*
+ * Return pointers to functions that overlap with 'size' bytes at offset 'off'
+ * of the value of 'map' and their number in 'cnt'.
+ */
+struct bpf_func_ptr *bpf_map_range_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map,
+ u64 off, u64 size, u32 *cnt)
+{
+ u32 first, last;
+
+ *cnt = 0;
+ if (!env->func_ptr_cnt || !size)
+ return NULL;
+
+ /* a pointer that starts up to 7 bytes before 'off' overlaps too */
+ first = func_ptr_lower_bound(env, map, off >= sizeof(u64) ? off - sizeof(u64) + 1 : 0);
+ last = func_ptr_lower_bound(env, map, off + size);
+ if (first >= last)
+ return NULL;
+
+ *cnt = last - first;
+ return &env->func_ptrs[first];
+}
+
+/* Return all pointers to functions in the value of 'map' */
+struct bpf_func_ptr *bpf_map_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map, u32 *cnt)
+{
+ return bpf_map_range_func_ptrs(env, map, 0, (u64)map->value_size, cnt);
+}
+
+/* instructions [off, off + len) replaced the instruction at 'off' */
+void bpf_adjust_func_ptrs(struct bpf_verifier_env *env, u32 off, u32 len)
+{
+ struct bpf_func_ptr *p;
+ u32 i;
+
+ if (len <= 1)
+ return;
+
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ p = &env->func_ptrs[i];
+ if (p->xlated_off <= off || p->xlated_off == BPF_FUNC_PTR_DELETED)
+ continue;
+ p->xlated_off += len - 1;
+ }
+}
+
+/*
+ * Instructions [off, off + len) are about to be removed. It's called before
+ * the starts of subprogs are adjusted. A subprog is gone when all of its
+ * instructions are. Otherwise, e.g. when its first instruction is a nop,
+ * it starts where the removed instructions did.
+ */
+void bpf_adjust_func_ptrs_after_remove(struct bpf_verifier_env *env, u32 off, u32 len)
+{
+ struct bpf_func_ptr *p;
+ int subprog;
+ u32 i;
+
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ p = &env->func_ptrs[i];
+ if (p->xlated_off < off || p->xlated_off == BPF_FUNC_PTR_DELETED)
+ continue;
+ if (p->xlated_off >= off + len) {
+ p->xlated_off -= len;
+ continue;
+ }
+ subprog = bpf_find_subprog(env, p->xlated_off);
+ if (subprog > 0 && env->subprog_info[subprog + 1].start > off + len)
+ p->xlated_off = off;
+ else
+ p->xlated_off = BPF_FUNC_PTR_DELETED;
+ }
+}
+
#define BTF_TYPE_SAFE_RCU(__type) __PASTE(__type, __safe_rcu)
#define BTF_TYPE_SAFE_RCU_OR_NULL(__type) __PASTE(__type, __safe_rcu_or_null)
#define BTF_TYPE_SAFE_TRUSTED(__type) __PASTE(__type, __safe_trusted)
@@ -6417,12 +6633,100 @@ static void add_scalar_to_reg(struct bpf_reg_state *dst_reg, s64 val)
reg_bounds_sync(dst_reg);
}
+static void mark_reg_func_ptr(struct bpf_verifier_env *env, struct bpf_reg_state *regs,
+ int regno, int subprog)
+{
+ mark_reg_known_zero(env, regs, regno);
+ regs[regno].type = PTR_TO_FUNC;
+ regs[regno].subprogno = subprog;
+}
+
+/* a read from a table of functions branches into that many states at most */
+#define BPF_MAX_FUNC_PTR_TARGETS 64
+/* and the table, which might have other data in it, is that many pointers long at most */
+#define BPF_MAX_FUNC_PTR_RANGE 4096
+
+/*
+ * A read from a frozen read-only map that has pointers to functions, see
+ * resolve_func_ptrs(). A read of exactly one pointer yields PTR_TO_FUNC.
+ * When the offset is variable and only pointers can be read, which is how
+ * an element of a table of functions is loaded, the verification continues
+ * with each of them. Other reads that overlap with a pointer are rejected,
+ * because their result is not known until the program is jitted.
+ *
+ * Return -ENOENT if there are no pointers to functions in the bytes that are read.
+ */
+static int check_func_ptr_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
+ int size, int value_regno)
+{
+ u64 min_off = reg_umin(reg) + off, max_off = reg_umax(reg) + off;
+ struct tnum offs = tnum_add(reg->var_off, tnum_const(off));
+ struct bpf_reg_state *regs = cur_regs(env);
+ struct bpf_map *map = reg->map_ptr;
+ struct bpf_verifier_state *branch;
+ int subprog, targets[BPF_MAX_FUNC_PTR_TARGETS];
+ struct bpf_func_ptr *ptrs;
+ u32 i, cnt, n = 0;
+ u64 o;
+
+ ptrs = bpf_map_range_func_ptrs(env, map, min_off, max_off - min_off + size, &cnt);
+ if (!ptrs)
+ return -ENOENT;
+
+ if (size != sizeof(u64) || value_regno < 0 || !tnum_is_aligned(offs, sizeof(u64)) ||
+ (max_off - min_off) / sizeof(u64) > BPF_MAX_FUNC_PTR_RANGE)
+ goto overlap;
+
+ /*
+ * Every offset that the read is possible at has to be the offset of
+ * a pointer. var_off tells the stride of the elements of an array.
+ */
+ for (o = round_up(min_off, sizeof(u64)), i = 0; o <= max_off; o += sizeof(u64)) {
+ if ((o ^ offs.value) & ~offs.mask)
+ continue;
+ while (i < cnt && ptrs[i].map_off < o)
+ i++;
+ if (i == cnt || ptrs[i].map_off != o)
+ goto overlap;
+ if (n == BPF_MAX_FUNC_PTR_TARGETS) {
+ verbose(env, "read from map '%s' may yield more than %d pointers to functions\n",
+ map->name, BPF_MAX_FUNC_PTR_TARGETS);
+ return -E2BIG;
+ }
+ subprog = bpf_find_subprog(env, ptrs[i].xlated_off);
+ if (verifier_bug_if(subprog <= 0, env, "no function at insn %u for map '%s' offset %u",
+ ptrs[i].xlated_off, map->name, ptrs[i].map_off))
+ return -EFAULT;
+ ptrs[i].used = true;
+ targets[n++] = subprog;
+ }
+ if (verifier_bug_if(!n, env, "no offsets to read map '%s' at", map->name))
+ return -EFAULT;
+
+ for (i = 0; i < n - 1; i++) {
+ branch = push_stack(env, env->insn_idx + 1, env->insn_idx,
+ env->cur_state->speculative);
+ if (IS_ERR(branch))
+ return PTR_ERR(branch);
+ mark_reg_func_ptr(env, branch->frame[branch->curframe]->regs, value_regno,
+ targets[i]);
+ }
+ mark_reg_func_ptr(env, regs, value_regno, targets[n - 1]);
+ return 0;
+
+overlap:
+ verbose(env, "read of %d bytes at offset [%llu,%llu] of map '%s' overlaps with a pointer to a function\n",
+ size, min_off, max_off, map->name);
+ return -EACCES;
+}
+
static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
int bpf_size, int value_regno, bool is_ldsx)
{
struct bpf_reg_state *regs = cur_regs(env);
int size = bpf_size_to_bytes(bpf_size);
struct bpf_map *map = reg->map_ptr;
+ int err;
switch (map->map_type) {
case BPF_MAP_TYPE_INSN_ARRAY:
@@ -6440,13 +6744,18 @@ static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state
break;
}
+ if (env->func_ptr_cnt) {
+ err = check_func_ptr_read(env, reg, off, size, value_regno);
+ if (err != -ENOENT)
+ return err;
+ }
+
/* If map is read-only, track its contents as scalars. */
if (tnum_is_const(reg->var_off) &&
bpf_map_is_rdonly(map) &&
map->ops->map_direct_value_addr) {
int map_off = off + reg->var_off.value;
u64 val = 0;
- int err;
err = bpf_map_direct_read(map, map_off, size, &val, is_ldsx);
if (err)
@@ -8754,6 +9063,7 @@ static int check_arg_const_str(struct bpf_verifier_env *env,
int map_off;
u64 map_addr;
char *str_ptr;
+ u32 cnt;
if (reg->type != PTR_TO_MAP_VALUE)
return -EINVAL;
@@ -8803,6 +9113,11 @@ static int check_arg_const_str(struct bpf_verifier_env *env,
verbose(env, "string is not zero-terminated\n");
return -EINVAL;
}
+ /* the bytes of a pointer to a function are not known until the program is jitted */
+ if (bpf_map_range_func_ptrs(env, map, map_off, strlen(str_ptr + map_off) + 1, &cnt)) {
+ verbose(env, "string overlaps with a pointer to a function\n");
+ return -EACCES;
+ }
return 0;
}
@@ -10584,13 +10899,61 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
static int process_bpf_exit_full(struct bpf_verifier_env *env,
bool *do_print_state, bool exception_exit);
-static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
- int *insn_idx)
+/*
+ * Call of a static subprog. The callee is verified in the context of
+ * the caller, hence set up a new frame and continue from the first
+ * instruction of the callee.
+ */
+static int check_static_func_call(struct bpf_verifier_env *env, int subprog,
+ int *insn_idx)
{
struct bpf_verifier_state *state = env->cur_state;
struct bpf_subprog_info *caller_info;
u16 callee_incoming, stack_arg_cnt;
struct bpf_func_state *caller;
+ int err;
+
+ caller = state->frame[state->curframe];
+
+ /*
+ * Track caller's total stack arg count (incoming + max outgoing).
+ * This is needed so the JIT knows how much stack arg space to allocate.
+ */
+ caller_info = &env->subprog_info[caller->subprogno];
+ callee_incoming = bpf_in_stack_arg_cnt(&env->subprog_info[subprog]);
+ stack_arg_cnt = bpf_in_stack_arg_cnt(caller_info) + callee_incoming;
+ if (stack_arg_cnt > caller_info->stack_arg_cnt)
+ caller_info->stack_arg_cnt = stack_arg_cnt;
+
+ /*
+ * For regular function entry setup new frame and continue
+ * from that frame.
+ */
+ err = setup_func_entry(env, subprog, *insn_idx, set_callee_state, state);
+ if (err)
+ return err;
+
+ bpf_diag_record_scrub(env, &caller->regs[BPF_REG_0], BPF_DIAG_MOD_CALLER_SAVED);
+ clear_caller_saved_regs(env, caller->regs);
+
+ /* and go analyze first insn of the callee */
+ *insn_idx = env->subprog_info[subprog].start - 1;
+
+ if (env->log.level & BPF_LOG_LEVEL) {
+ verbose(env, "caller:\n");
+ print_verifier_state(env, state, caller->frameno, true);
+ verbose(env, "callee:\n");
+ print_verifier_state(env, state, state->curframe, true);
+ }
+
+ return 0;
+}
+
+static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
+ int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *caller;
int err, subprog, target_insn;
u32 i, nregs;
@@ -10679,37 +11042,75 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
return 0;
}
- /*
- * Track caller's total stack arg count (incoming + max outgoing).
- * This is needed so the JIT knows how much stack arg space to allocate.
- */
- caller_info = &env->subprog_info[caller->subprogno];
- callee_incoming = bpf_in_stack_arg_cnt(&env->subprog_info[subprog]);
- stack_arg_cnt = bpf_in_stack_arg_cnt(caller_info) + callee_incoming;
- if (stack_arg_cnt > caller_info->stack_arg_cnt)
- caller_info->stack_arg_cnt = stack_arg_cnt;
+ return check_static_func_call(env, subprog, insn_idx);
+}
- /* for regular function entry setup new frame and continue
- * from that frame.
- */
- err = setup_func_entry(env, subprog, *insn_idx, set_callee_state, state);
+/*
+ * callx dst_reg: call a bpf subprog whose address is in dst_reg.
+ *
+ * The address of a subprog is either loaded into a register by ld_imm64 with
+ * src_reg == BPF_PSEUDO_FUNC, or it is read from a frozen read-only map, see
+ * resolve_func_ptrs(). Both are possible for static subprogs only. Hence all
+ * possible callees of callx are discovered by add_subprogs() and are reachable
+ * in the control flow graph before the main verification pass begins.
+ * PTR_TO_FUNC register identifies the callee, so from here on callx is verified
+ * as a direct call of that static subprog.
+ */
+static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,
+ int *insn_idx)
+{
+ struct bpf_func_state *caller = cur_func(env);
+ struct bpf_reg_state *reg;
+ const char *reason;
+ int err, subprog;
+
+ err = check_reg_arg(env, insn->dst_reg, SRC_OP);
if (err)
return err;
- bpf_diag_record_scrub(env, &caller->regs[BPF_REG_0], BPF_DIAG_MOD_CALLER_SAVED);
- clear_caller_saved_regs(env, caller->regs);
+ reg = reg_state(env, insn->dst_reg);
+ if (reg->type != PTR_TO_FUNC) {
+ verbose(env, "R%d has type %s, expected func\n", insn->dst_reg,
+ reg_type_str(env, reg->type));
+ reason = bpf_diag_fmt(
+ env, "R%d holds %s, but callx can only call through the address of a static BPF function.",
+ insn->dst_reg, bpf_diag_reg_type_plain(env, reg->type));
+ bpf_diag_register_type(
+ env, *insn_idx, insn->dst_reg, "indirect call through a non-function pointer", reason,
+ "Load the address of a static BPF function into the register before callx.");
+ return -EACCES;
+ }
- /* and go analyze first insn of the callee */
- *insn_idx = env->subprog_info[subprog].start - 1;
+ /*
+ * Arithmetic on PTR_TO_FUNC is allowed, but only unmodified address
+ * of a subprog can be called.
+ */
+ err = check_ptr_off_reg(env, reg, insn->dst_reg);
+ if (err)
+ return err;
- if (env->log.level & BPF_LOG_LEVEL) {
- verbose(env, "caller:\n");
- print_verifier_state(env, state, caller->frameno, true);
- verbose(env, "callee:\n");
- print_verifier_state(env, state, state->curframe, true);
+ /* there is no support for callx in the interpreter */
+ if (!env->prog->jit_requested) {
+ verbose(env, "JIT is required to use callx\n");
+ return -EOPNOTSUPP;
+ }
+ if (!bpf_jit_supports_callx()) {
+ verbose(env, "JIT doesn't support callx\n");
+ return -EOPNOTSUPP;
}
+ env->prog->jit_required = true;
- return 0;
+ /* PTR_TO_FUNC is a pointer to a static subprog */
+ subprog = reg->subprogno;
+ err = btf_check_subprog_call(env, subprog, caller->regs);
+ if (err == -EFAULT)
+ return err;
+
+ err = record_callx_edge(env, caller->subprogno, subprog);
+ if (err)
+ return err;
+
+ return check_static_func_call(env, subprog, insn_idx);
}
int map_set_for_each_callback_args(struct bpf_verifier_env *env,
@@ -18709,7 +19110,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->jmps_processed++;
if (opcode == BPF_CALL) {
- if (env->cur_state->active_locks) {
+ /* similar to static subprog calls callx is allowed under a lock */
+ if (env->cur_state->active_locks && !bpf_is_callx(insn)) {
if ((insn->src_reg == BPF_REG_0 &&
insn->imm != BPF_FUNC_spin_unlock &&
insn->imm != BPF_FUNC_kptr_xchg) ||
@@ -18727,6 +19129,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
mark_reg_scratched(env, BPF_REG_0);
if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
cur_func(env)->no_stack_arg_load = true;
+ if (bpf_is_callx(insn))
+ return check_func_callx(env, insn, &env->insn_idx);
if (insn->src_reg == BPF_PSEUDO_CALL)
return check_func_call(env, insn, &env->insn_idx);
if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL)
@@ -19302,6 +19706,30 @@ static int check_map_prog_compatibility(struct bpf_verifier_env *env,
return 0;
}
+/*
+ * Keep track of whether the map is used by one program only. Such program may
+ * store the addresses of its functions into the map when it's frozen, see
+ * resolve_func_ptrs(), since nothing else relies on what the map has. After
+ * that the map is not available to other programs.
+ */
+static int bpf_map_claim(struct bpf_verifier_env *env, struct bpf_map *map)
+{
+ unsigned long me = (unsigned long)env->prog->aux, old;
+
+ for (;;) {
+ old = READ_ONCE(map->user);
+ if (old == me || old == BPF_MAP_USER_MANY)
+ return 0;
+ if (old & BPF_MAP_USER_PATCHED) {
+ verbose(env, "map '%s' has addresses of functions of another program\n",
+ map->name);
+ return -EBUSY;
+ }
+ if (cmpxchg(&map->user, old, old ? BPF_MAP_USER_MANY : me) == old)
+ return 0;
+ }
+}
+
static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)
{
int i, err;
@@ -19339,6 +19767,10 @@ static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)
env->used_maps[env->used_map_cnt++] = map;
+ err = bpf_map_claim(env, map);
+ if (err)
+ return err;
+
if (map->map_type == BPF_MAP_TYPE_INSN_ARRAY) {
err = bpf_insn_array_init(map, env->prog);
if (err) {
@@ -19493,6 +19925,14 @@ static int check_jmp_fields(struct bpf_verifier_env *env, struct bpf_insn *insn)
switch (opcode) {
case BPF_CALL:
+ if (bpf_is_callx(insn)) {
+ /* callx dst_reg */
+ if (insn->src_reg != BPF_REG_0 || insn->imm != 0 || insn->off != 0) {
+ verbose(env, "BPF_CALL|BPF_X uses reserved fields\n");
+ return -EINVAL;
+ }
+ return 0;
+ }
if (BPF_SRC(insn->code) != BPF_K ||
(insn->src_reg != BPF_PSEUDO_KFUNC_CALL && insn->off != 0) ||
(insn->src_reg != BPF_REG_0 && insn->src_reg != BPF_PSEUDO_CALL &&
@@ -19747,6 +20187,93 @@ static int check_and_resolve_insns(struct bpf_verifier_env *env)
return 0;
}
+static int add_func_ptr(struct bpf_verifier_env *env, struct bpf_map *map, u32 map_off,
+ u32 xlated_off)
+{
+ struct bpf_func_ptr *ptrs;
+
+ /* grow by doubling, the array is sorted and searched later */
+ if (!(env->func_ptr_cnt & (env->func_ptr_cnt - 1))) {
+ ptrs = kvrealloc(env->func_ptrs,
+ array_size(max(2 * env->func_ptr_cnt, 16U), sizeof(*ptrs)),
+ GFP_KERNEL_ACCOUNT);
+ if (!ptrs)
+ return -ENOMEM;
+ env->func_ptrs = ptrs;
+ }
+ env->func_ptrs[env->func_ptr_cnt++] = (struct bpf_func_ptr){
+ .map = map,
+ .map_off = map_off,
+ .orig_off = xlated_off,
+ .xlated_off = xlated_off,
+ };
+ return 0;
+}
+
+/*
+ * Compilers put pointers to functions into read-only data: tables of functions,
+ * structures of operations, vtables, where they are mixed with other data.
+ * The loader stores such data in a frozen read-only array map and resolves
+ * a pointer to a static function to the offset in bytes of its first
+ * instruction in the program: the address of the function in the program.
+ *
+ * Find 64-bit values that look like that in the maps of a program that uses
+ * callx. It's a guess. When the value is not a pointer, the program either
+ * fails to load, because it does with a pointer what can be done with
+ * a number only, or it sees the address of a function instead of the number.
+ * It's known before the control flow graph of the program is built and the main
+ * verification pass begins which functions may be called via callx.
+ *
+ * When the program is jitted the offsets are replaced with the addresses of
+ * the functions in the map itself, see jit_subprogs(). Hence the program has to
+ * be the only user of the map, see bpf_map_claim(): nothing else may rely on
+ * what the map had.
+ *
+ * The program reads the addresses of its functions from there like any other
+ * data, so it has to be allowed to leak pointers.
+ */
+static int resolve_func_ptrs(struct bpf_verifier_env *env)
+{
+ int insn_cnt = env->prog->len;
+ int i, err, subprog;
+ struct bpf_map *map;
+ u64 addr, val;
+ u32 off;
+
+ if (!env->has_callx || !env->allow_ptr_leaks)
+ return 0;
+
+ for (i = 0; i < env->used_map_cnt; i++) {
+ map = env->used_maps[i];
+ /* coincidences in maps that are shared with other programs don't matter */
+ if (READ_ONCE(map->user) != (unsigned long)env->prog->aux)
+ continue;
+ if (map->map_type != BPF_MAP_TYPE_ARRAY || map->max_entries != 1 ||
+ !bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
+ !IS_ERR_OR_NULL(map->record))
+ continue;
+ if (map->ops->map_direct_value_addr(map, &addr, 0))
+ continue;
+
+ for (off = 0; off + sizeof(u64) <= map->value_size; off += sizeof(u64)) {
+ val = *(u64 *)(unsigned long)(addr + off);
+ if (!val || val % sizeof(struct bpf_insn) ||
+ val / sizeof(struct bpf_insn) >= insn_cnt)
+ continue;
+ subprog = bpf_find_subprog(env, val / sizeof(struct bpf_insn));
+ if (subprog <= 0 || bpf_subprog_is_global(env, subprog))
+ continue;
+ err = add_func_ptr(env, map, off, val / sizeof(struct bpf_insn));
+ if (err)
+ return err;
+ }
+ }
+ if (env->func_ptr_cnt)
+ sort(env->func_ptrs, env->func_ptr_cnt, sizeof(*env->func_ptrs),
+ cmp_func_ptrs, NULL);
+ return 0;
+}
+
/* drop refcnt of maps used by the rejected program */
static void release_maps(struct bpf_verifier_env *env)
{
@@ -21717,6 +22244,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ /* Find pointers to functions in the read-only maps of the program. */
+ ret = resolve_func_ptrs(env);
+ if (ret < 0)
+ goto skip_full_check;
+
/* Build kfunc prototypes after resolving program resources. */
ret = add_kfuncs(env);
if (ret < 0)
@@ -21776,6 +22308,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
ret = do_check_main(env);
ret = ret ?: do_check_subprogs(env);
+ /* reject recursion through the callx edges found by the main pass */
+ if (ret == 0 && env->callx_edges)
+ ret = sort_subprogs_topo(env);
+
if (ret == 0 && bpf_prog_is_offloaded(env->prog->aux))
ret = bpf_prog_offload_finalize(env);
@@ -21918,6 +22454,8 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
kvfree(env->scc_info);
kvfree(env->succ);
kvfree(env->gotox_tmp_buf);
+ kvfree(env->callx_edges);
+ kvfree(env->func_ptrs);
bpf_diag_free(env);
kvfree(env);
return ret;
diff --git a/tools/lib/bpf/bpf_gen_internal.h b/tools/lib/bpf/bpf_gen_internal.h
index 6c5ad6c55e8a6..206adf28793d7 100644
--- a/tools/lib/bpf/bpf_gen_internal.h
+++ b/tools/lib/bpf/bpf_gen_internal.h
@@ -51,9 +51,17 @@ struct bpf_gen {
__u32 nr_ksyms;
int fd_array;
int nr_fd_array;
+ /*
+ * Maps with pointers to functions, that programs get their own copies
+ * of, take slots in fd_array after nr_obj_maps maps of the object.
+ */
+ __u32 nr_obj_maps;
+ __u32 nr_func_ptr_maps;
+ __u32 max_func_ptr_maps;
};
-void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps);
+void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps,
+ int max_func_ptr_maps);
int bpf_gen__finish(struct bpf_gen *gen, int nr_progs, int nr_maps);
void bpf_gen__free(struct bpf_gen *gen);
void bpf_gen__load_btf(struct bpf_gen *gen, const void *raw_data, __u32 raw_size);
@@ -68,6 +76,9 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *value, __u32 value_size,
__u64 flags);
void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx);
+int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,
+ void *value, __u32 value_size, const __u32 *ptr_offs,
+ const __u64 *ptr_vals, int ptr_cnt);
void bpf_gen__record_attach_target(struct bpf_gen *gen, const char *name, enum bpf_attach_type type);
void bpf_gen__record_extern(struct bpf_gen *gen, const char *name, bool is_weak,
bool is_typeless, bool is_ld64, int kind, int insn_idx);
diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index af3a04f161ac1..251392aa8b41f 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -112,13 +112,23 @@ static void emit2(struct bpf_gen *gen, struct bpf_insn insn1, struct bpf_insn in
static int add_data(struct bpf_gen *gen, const void *data, __u32 size);
static void emit_sys_close_blob(struct bpf_gen *gen, int blob_off);
-void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps)
+void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps,
+ int max_func_ptr_maps)
{
size_t stack_sz = sizeof(struct loader_stack), nr_progs_sz;
int i;
gen->fd_array = add_data(gen, NULL, MAX_FD_ARRAY_SZ * sizeof(int));
gen->log_level = log_level;
+ gen->nr_obj_maps = nr_maps;
+ gen->max_func_ptr_maps = max_func_ptr_maps;
+ if (nr_maps + max_func_ptr_maps > MAX_USED_MAPS) {
+ pr_warn("Total maps exceeds %d\n", MAX_USED_MAPS);
+ gen->error = -E2BIG;
+ return;
+ }
+ /* their fds are closed like the fds of the maps of the object when loading fails */
+ nr_maps += max_func_ptr_maps;
/* save ctx pointer into R6 */
emit(gen, BPF_MOV64_REG(BPF_REG_6, BPF_REG_1));
@@ -385,6 +395,9 @@ int bpf_gen__finish(struct bpf_gen *gen, int nr_progs, int nr_maps)
return gen->error;
}
emit_sys_close_stack(gen, stack_off(btf_fd));
+ /* programs hold their maps with pointers to functions, nothing else needs them */
+ for (i = 0; i < gen->nr_func_ptr_maps; i++)
+ emit_sys_close_blob(gen, blob_fd_array_off(gen, gen->nr_obj_maps + i));
for (i = 0; i < gen->nr_progs; i++)
move_stack2ctx(gen,
sizeof(struct bpf_loader_ctx) +
@@ -1127,12 +1140,59 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
gen->nr_progs++;
}
+/*
+ * if (map_desc[map_idx].initial_value) {
+ * if (ctx->flags & BPF_SKEL_KERNEL)
+ * bpf_probe_read_kernel(value, value_size, initial_value);
+ * else
+ * bpf_copy_from_user(value, value_size, initial_value);
+ * nr_more_insns that the caller emits
+ * }
+ */
+static void emit_copy_initial_value(struct bpf_gen *gen, int map_idx, int value,
+ __u32 value_size, int nr_more_insns)
+{
+ emit(gen, BPF_LDX_MEM(BPF_DW, BPF_REG_3, BPF_REG_6,
+ sizeof(struct bpf_loader_ctx) +
+ sizeof(struct bpf_map_desc) * map_idx +
+ offsetof(struct bpf_map_desc, initial_value)));
+ emit(gen, BPF_JMP_IMM(BPF_JEQ, BPF_REG_3, 0, 8 + nr_more_insns));
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, value));
+ emit(gen, BPF_MOV64_IMM(BPF_REG_2, value_size));
+ emit(gen, BPF_LDX_MEM(BPF_W, BPF_REG_0, BPF_REG_6,
+ offsetof(struct bpf_loader_ctx, flags)));
+ emit(gen, BPF_JMP_IMM(BPF_JSET, BPF_REG_0, BPF_SKEL_KERNEL, 2));
+ emit(gen, BPF_EMIT_CALL(BPF_FUNC_copy_from_user));
+ emit(gen, BPF_JMP_IMM(BPF_JA, 0, 0, 1));
+ emit(gen, BPF_EMIT_CALL(BPF_FUNC_probe_read_kernel));
+}
+
+/* Update the element of the map whose fd is in the slot map_idx of fd_array */
+static void emit_map_update_elem(struct bpf_gen *gen, int map_idx, union bpf_attr *attr,
+ int attr_size, int key, int value, __u32 value_size)
+{
+ int map_update_attr;
+
+ map_update_attr = add_data(gen, attr, attr_size);
+ pr_debug("gen: map_update_elem: idx %d, value: off %d size %u, attr: off %d size %d\n",
+ map_idx, value, value_size, map_update_attr, attr_size);
+ move_blob2blob(gen, attr_field(map_update_attr, map_fd), 4,
+ blob_fd_array_off(gen, map_idx));
+ emit_rel_store(gen, attr_field(map_update_attr, key), key);
+ emit_rel_store(gen, attr_field(map_update_attr, value), value);
+ /* emit MAP_UPDATE_ELEM command */
+ emit_sys_bpf(gen, BPF_MAP_UPDATE_ELEM, map_update_attr, attr_size);
+ debug_ret(gen, "update_elem idx %d value_size %d", map_idx, value_size);
+ emit_check_err(gen);
+}
+
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,
__u32 value_size, __u64 flags)
{
int attr_size = offsetofend(union bpf_attr, flags);
- int map_update_attr, value, key;
union bpf_attr attr;
+ int value, key;
int zero = 0;
memset(&attr, 0, attr_size);
@@ -1142,47 +1202,16 @@ void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,
key = add_data(gen, &zero, sizeof(zero));
/*
- * if (map_desc[map_idx].initial_value) {
- * if (ctx->flags & BPF_SKEL_KERNEL)
- * bpf_probe_read_kernel(value, value_size, initial_value);
- * else
- * bpf_copy_from_user(value, value_size, initial_value);
- * }
- *
* The runtime initial_value comes from the host-supplied loader
* ctx and would overwrite the blob value that the program signature
* covers and the kernel verifies at load time. For a signed loader
* (gen_hash) the attested blob value must be authoritative, so skip
* the override and leave the signed value in place.
*/
- if (!OPTS_GET(gen->opts, gen_hash, false)) {
- emit(gen, BPF_LDX_MEM(BPF_DW, BPF_REG_3, BPF_REG_6,
- sizeof(struct bpf_loader_ctx) +
- sizeof(struct bpf_map_desc) * map_idx +
- offsetof(struct bpf_map_desc, initial_value)));
- emit(gen, BPF_JMP_IMM(BPF_JEQ, BPF_REG_3, 0, 8));
- emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
- 0, 0, 0, value));
- emit(gen, BPF_MOV64_IMM(BPF_REG_2, value_size));
- emit(gen, BPF_LDX_MEM(BPF_W, BPF_REG_0, BPF_REG_6,
- offsetof(struct bpf_loader_ctx, flags)));
- emit(gen, BPF_JMP_IMM(BPF_JSET, BPF_REG_0, BPF_SKEL_KERNEL, 2));
- emit(gen, BPF_EMIT_CALL(BPF_FUNC_copy_from_user));
- emit(gen, BPF_JMP_IMM(BPF_JA, 0, 0, 1));
- emit(gen, BPF_EMIT_CALL(BPF_FUNC_probe_read_kernel));
- }
+ if (!OPTS_GET(gen->opts, gen_hash, false))
+ emit_copy_initial_value(gen, map_idx, value, value_size, 0);
- map_update_attr = add_data(gen, &attr, attr_size);
- pr_debug("gen: map_update_elem: idx %d, value: off %d size %u, attr: off %d size %d\n",
- map_idx, value, value_size, map_update_attr, attr_size);
- move_blob2blob(gen, attr_field(map_update_attr, map_fd), 4,
- blob_fd_array_off(gen, map_idx));
- emit_rel_store(gen, attr_field(map_update_attr, key), key);
- emit_rel_store(gen, attr_field(map_update_attr, value), value);
- /* emit MAP_UPDATE_ELEM command */
- emit_sys_bpf(gen, BPF_MAP_UPDATE_ELEM, map_update_attr, attr_size);
- debug_ret(gen, "update_elem idx %d value_size %d", map_idx, value_size);
- emit_check_err(gen);
+ emit_map_update_elem(gen, map_idx, &attr, attr_size, key, value, value_size);
}
void bpf_gen__populate_outer_map(struct bpf_gen *gen, int outer_map_idx, int slot,
@@ -1214,6 +1243,77 @@ void bpf_gen__populate_outer_map(struct bpf_gen *gen, int outer_map_idx, int slo
emit_check_err(gen);
}
+/*
+ * A copy of the read-only data map obj_map_idx for a program, with the offsets
+ * of its functions in it, see create_func_ptr_map() in libbpf.c. It's not
+ * a map of the object: it's not in the loader ctx. Its content is the content
+ * of obj_map_idx, that the host may supply when the skeleton is loaded, with
+ * ptr_cnt 64-bit ptr_vals at ptr_offs.
+ * Return the index of the map in fd_array for instructions to refer to.
+ */
+int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,
+ void *pvalue, __u32 value_size, const __u32 *ptr_offs,
+ const __u64 *ptr_vals, int ptr_cnt)
+{
+ int attr_size = offsetofend(union bpf_attr, map_extra);
+ int map_create_attr, map_idx, key, value, zero = 0, i;
+ union bpf_attr attr;
+
+ if (gen->nr_func_ptr_maps == gen->max_func_ptr_maps) {
+ gen->error = -EDOM; /* internal bug */
+ return 0;
+ }
+ map_idx = gen->nr_obj_maps + gen->nr_func_ptr_maps++;
+
+ memset(&attr, 0, attr_size);
+ attr.map_type = tgt_endian(BPF_MAP_TYPE_ARRAY);
+ attr.key_size = tgt_endian((__u32)sizeof(int));
+ attr.value_size = tgt_endian(value_size);
+ attr.max_entries = tgt_endian((__u32)1);
+ attr.map_flags = tgt_endian((__u32)BPF_F_RDONLY_PROG);
+ if (map_name)
+ libbpf_strlcpy(attr.map_name, map_name, sizeof(attr.map_name));
+
+ map_create_attr = add_data(gen, &attr, attr_size);
+ pr_debug("gen: func_ptr_map_create: %s idx %d value_size %u, attr: off %d size %d\n",
+ map_name, map_idx, value_size, map_create_attr, attr_size);
+ emit_sys_bpf(gen, BPF_MAP_CREATE, map_create_attr, attr_size);
+ debug_ret(gen, "func_ptr_map_create %s idx %d value_size %d", map_name, map_idx,
+ value_size);
+ emit_check_err(gen);
+ /* remember map_fd in fd_array */
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, blob_fd_array_off(gen, map_idx)));
+ emit(gen, BPF_STX_MEM(BPF_W, BPF_REG_1, BPF_REG_7, 0));
+
+ /* pvalue has the pointers already */
+ value = add_data(gen, pvalue, value_size);
+ key = add_data(gen, &zero, sizeof(zero));
+
+ /* see bpf_gen__map_update_elem() */
+ if (!OPTS_GET(gen->opts, gen_hash, false)) {
+ /* the jump over these instructions has 16-bit offset */
+ if (ptr_cnt > 10000) {
+ gen->error = -E2BIG;
+ return 0;
+ }
+ emit_copy_initial_value(gen, obj_map_idx, value, value_size, 3 * ptr_cnt);
+ /* the content that the host supplied doesn't have them */
+ for (i = 0; i < ptr_cnt; i++) {
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, value + ptr_offs[i]));
+ emit(gen, BPF_ST_MEM(BPF_DW, BPF_REG_1, 0, ptr_vals[i]));
+ }
+ }
+
+ attr_size = offsetofend(union bpf_attr, flags);
+ memset(&attr, 0, attr_size);
+ emit_map_update_elem(gen, map_idx, &attr, attr_size, key, value, value_size);
+
+ bpf_gen__map_freeze(gen, map_idx);
+ return map_idx;
+}
+
void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx)
{
int attr_size = offsetofend(union bpf_attr, map_fd);
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index cd1ea1bb53cbf..2e11808f7508d 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -544,6 +544,7 @@ struct bpf_struct_ops {
#define PERCPU_SEC ".percpu"
#define BSS_SEC ".bss"
#define RODATA_SEC ".rodata"
+#define DATA_REL_RO_SEC ".data.rel.ro"
#define KCONFIG_SEC ".kconfig"
#define KSYMS_SEC ".ksyms"
#define STRUCT_OPS_SEC ".struct_ops"
@@ -601,6 +602,9 @@ struct bpf_map {
bool autoattach;
__u64 map_extra;
struct bpf_program *excl_prog;
+ /* pointers to functions in the data of an internal map, see obj->func_ptrs */
+ struct func_ptr *func_ptrs;
+ size_t func_ptr_cnt;
};
enum extern_type {
@@ -780,6 +784,29 @@ struct bpf_object {
} *jumptable_maps;
size_t jumptable_map_cnt;
+ /*
+ * Pointers to functions found in read-only data sections: tables of
+ * functions, structures of operations, vtables. Sorted by section
+ * and offset.
+ */
+ struct func_ptr {
+ int sec_idx; /* ELF section that contains the pointer */
+ size_t sec_off; /* offset of the pointer in the section */
+ size_t text_off; /* offset of the function in .text section */
+ } *func_ptrs;
+ size_t func_ptr_cnt;
+
+ /*
+ * Read-only data with pointers to functions is different for every
+ * program that uses it, because so are the offsets of the functions.
+ */
+ struct {
+ struct bpf_program *prog;
+ int map_idx;
+ int fd;
+ } *func_ptr_maps;
+ size_t func_ptr_map_cnt;
+
struct kern_feature_cache *feat_cache;
char *token_path;
int token_fd;
@@ -845,6 +872,17 @@ static bool insn_is_pseudo_func(struct bpf_insn *insn)
return is_ldimm64_insn(insn) && insn->src_reg == BPF_PSEUDO_FUNC;
}
+/*
+ * ld_imm64 that loads the address of read-only data with pointers to functions
+ * is marked by bpf_object__relocate() before the code is relocated. Compilers
+ * leave src_reg of other ld_imm64 zero and it's set by
+ * bpf_object__relocate_data() later.
+ */
+static bool insn_is_func_ptrs_addr(struct bpf_insn *insn)
+{
+ return is_ldimm64_insn(insn) && insn->src_reg == BPF_PSEUDO_MAP_VALUE;
+}
+
static int
bpf_object__init_prog(struct bpf_object *obj, struct bpf_program *prog,
const char *name, size_t sec_idx, const char *sec_name,
@@ -4024,6 +4062,17 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
err = bpf_object__add_programs(obj, data, name, idx);
if (err)
return err;
+ } else if (strcmp(name, DATA_REL_RO_SEC) == 0 ||
+ str_has_pfx(name, DATA_REL_RO_SEC ".")) {
+ /*
+ * Constants with pointers in them, e.g. vtables,
+ * that position independent code keeps here to
+ * have them relocated. There is nothing that
+ * writes to it after that.
+ */
+ sec_desc->sec_type = SEC_RODATA;
+ sec_desc->shdr = sh;
+ sec_desc->data = data;
} else if (strcmp(name, DATA_SEC) == 0 ||
str_has_pfx(name, DATA_SEC ".")) {
sec_desc->sec_type = SEC_DATA;
@@ -4068,8 +4117,16 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
targ_sec_idx >= obj->efile.sec_cnt)
return -LIBBPF_ERRNO__FORMAT;
- /* Only do relo for section with exec instructions */
+ /*
+ * Only do relo for section with exec instructions,
+ * struct_ops, maps, and read-only data that might
+ * have pointers to functions.
+ */
if (!section_have_execinstr(obj, targ_sec_idx) &&
+ strcmp(name, ".rel" RODATA_SEC) &&
+ !str_has_pfx(name, ".rel" RODATA_SEC ".") &&
+ strcmp(name, ".rel" DATA_REL_RO_SEC) &&
+ !str_has_pfx(name, ".rel" DATA_REL_RO_SEC ".") &&
strcmp(name, ".rel" STRUCT_OPS_SEC) &&
strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
@@ -6471,6 +6528,168 @@ static int create_jt_map(struct bpf_object *obj, struct bpf_program *prog, struc
return err;
}
+/*
+ * The kernel recognizes a pointer to a function in a frozen read-only map by
+ * its value: the offset in bytes of the function in the program. It makes
+ * callx work for tables of functions, structures of operations and vtables,
+ * where pointers are mixed with other data. Functions have different offsets
+ * in different programs, so create a copy of the map for the program.
+ * The kernel replaces the offsets with the addresses of the functions when it
+ * loads the program, which has to be the only user of the map.
+ */
+static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog, int map_idx)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);
+ struct bpf_map *map = &obj->maps[map_idx];
+ __u32 value_size = map->def.value_size;
+ size_t i, j, cnt, sec_insn_off;
+ struct func_ptr *ptrs;
+ int map_fd, err, zero = 0;
+ __u64 val;
+ void *data, *tmp;
+
+ for (i = 0; i < obj->func_ptr_map_cnt; i++)
+ if (obj->func_ptr_maps[i].prog == prog &&
+ obj->func_ptr_maps[i].map_idx == map_idx)
+ return obj->func_ptr_maps[i].fd;
+
+ data = malloc(value_size);
+ if (!data)
+ return -ENOMEM;
+
+ /*
+ * The content of the map is final, it's frozen already. There is no map
+ * when light skeleton is generated, but there is what it's created with.
+ */
+ if (map->mmaped) {
+ memcpy(data, map->mmaped, value_size);
+ } else if (obj->gen_loader) {
+ err = -EINVAL;
+ goto err_free;
+ } else if (bpf_map_lookup_elem(map->fd, &zero, data)) {
+ err = -errno;
+ pr_warn("prog '%s': map '%s': failed to read the content: %s\n",
+ prog->name, map->name, errstr(err));
+ goto err_free;
+ }
+
+ ptrs = map->func_ptrs;
+ cnt = map->func_ptr_cnt;
+ for (i = 0; i < cnt; i++) {
+ if (ptrs[i].sec_off + sizeof(val) > value_size) {
+ err = -LIBBPF_ERRNO__FORMAT;
+ goto err_free;
+ }
+ /*
+ * Static functions were appended by bpf_object__append_func_ptrs_code().
+ * A global function is in the program only if the code refers to it.
+ */
+ sec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;
+ for (j = 0; j < prog->subprog_cnt; j++)
+ if (prog->subprogs[j].sec_insn_off == sec_insn_off)
+ break;
+ if (j == prog->subprog_cnt) {
+ pr_debug("prog '%s': map '%s': no function for the pointer at offset %zu, it's NULL\n",
+ prog->name, map->name, ptrs[i].sec_off);
+ val = 0;
+ } else {
+ val = (__u64)prog->subprogs[j].sub_insn_off * BPF_INSN_SZ;
+ }
+ /* light skeleton can be generated for a target of another endianness */
+ if (!is_native_endianness(obj))
+ val = bswap_64(val);
+ memcpy(data + ptrs[i].sec_off, &val, sizeof(val));
+ }
+
+ /*
+ * The kernel takes any aligned 64-bit value that is equal to the offset
+ * of a function for a pointer. Tell when it's going to get it wrong.
+ */
+ for (i = 0, j = 0; i + sizeof(val) <= value_size; i += sizeof(val)) {
+ __u32 k;
+
+ while (j < cnt && ptrs[j].sec_off < i)
+ j++;
+ if (j < cnt && ptrs[j].sec_off == i)
+ continue;
+ memcpy(&val, data + i, sizeof(val));
+ if (!is_native_endianness(obj))
+ val = bswap_64(val);
+ if (!val || val % BPF_INSN_SZ)
+ continue;
+ for (k = 0; k < prog->subprog_cnt; k++) {
+ if ((__u64)prog->subprogs[k].sub_insn_off * BPF_INSN_SZ != val)
+ continue;
+ pr_warn("prog '%s': map '%s': value %llu at offset %zu is the offset of a function, the kernel will treat it as a pointer to it\n",
+ prog->name, map->name, (unsigned long long)val, i);
+ break;
+ }
+ }
+
+ if (obj->gen_loader) {
+ __u32 *ptr_offs = calloc(cnt, sizeof(*ptr_offs));
+ __u64 *ptr_vals = calloc(cnt, sizeof(*ptr_vals));
+
+ if (!ptr_offs || !ptr_vals) {
+ free(ptr_offs);
+ free(ptr_vals);
+ err = -ENOMEM;
+ goto err_free;
+ }
+ for (i = 0; i < cnt; i++) {
+ ptr_offs[i] = ptrs[i].sec_off;
+ memcpy(&ptr_vals[i], data + ptrs[i].sec_off, sizeof(val));
+ if (!is_native_endianness(obj))
+ ptr_vals[i] = bswap_64(ptr_vals[i]);
+ }
+ /* it's an index in fd_array of the loader, not an fd */
+ map_fd = bpf_gen__func_ptr_map_create(obj->gen_loader, map->name, map_idx, data,
+ value_size, ptr_offs, ptr_vals, cnt);
+ free(ptr_offs);
+ free(ptr_vals);
+ goto done;
+ }
+
+ map_fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, map->name, sizeof(int), value_size, 1, &opts);
+ if (map_fd < 0) {
+ err = map_fd;
+ goto err_free;
+ }
+
+ err = bpf_map_update_elem(map_fd, &zero, data, 0);
+ if (!err)
+ err = bpf_map_freeze(map_fd);
+ if (err) {
+ err = -errno;
+ goto err_close;
+ }
+done:
+
+ tmp = libbpf_reallocarray(obj->func_ptr_maps, obj->func_ptr_map_cnt + 1,
+ sizeof(*obj->func_ptr_maps));
+ if (!tmp) {
+ err = -ENOMEM;
+ goto err_close;
+ }
+ obj->func_ptr_maps = tmp;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].prog = prog;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].map_idx = map_idx;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].fd = map_fd;
+ obj->func_ptr_map_cnt++;
+
+ pr_debug("prog '%s': created a copy of map '%s' with %zu pointers to functions\n",
+ prog->name, map->name, cnt);
+ free(data);
+ return map_fd;
+
+err_close:
+ if (!obj->gen_loader)
+ close(map_fd);
+err_free:
+ free(data);
+ return err;
+}
+
/* Relocate data references within program code:
* - map references;
* - global variable references;
@@ -6508,7 +6727,20 @@ bpf_object__relocate_data(struct bpf_object *obj, struct bpf_program *prog)
if (relo->map_idx == obj->arena_map_idx)
insn[1].imm += obj->arena_data_off;
- if (obj->gen_loader) {
+ if (map->autocreate && map->func_ptr_cnt) {
+ int map_fd;
+
+ /* the program gets its own map with pointers to its functions */
+ map_fd = create_func_ptr_map(obj, prog, relo->map_idx);
+ if (map_fd < 0) {
+ pr_warn("prog '%s': relo #%d: can't create a copy of map '%s' with pointers to functions\n",
+ prog->name, i, map->name);
+ return map_fd;
+ }
+ insn[0].src_reg = obj->gen_loader ? BPF_PSEUDO_MAP_IDX_VALUE :
+ BPF_PSEUDO_MAP_VALUE;
+ insn[0].imm = map_fd;
+ } else if (obj->gen_loader) {
insn[0].src_reg = BPF_PSEUDO_MAP_IDX_VALUE;
insn[0].imm = relo->map_idx;
} else if (map->autocreate) {
@@ -6835,6 +7067,51 @@ bpf_object__append_subprog_code(struct bpf_object *obj, struct bpf_program *main
return 0;
}
+static int
+bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
+ struct bpf_program *prog);
+
+/* Append to the main program all functions that the data of the map points to */
+static int
+bpf_object__append_func_ptrs_code(struct bpf_object *obj, struct bpf_program *main_prog,
+ const struct bpf_map *map)
+{
+ struct bpf_program *subprog;
+ size_t i, cnt, sec_insn_off;
+ struct func_ptr *ptrs;
+ int err;
+
+ ptrs = map->func_ptrs;
+ cnt = map->func_ptr_cnt;
+ for (i = 0; i < cnt; i++) {
+ sec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;
+ subprog = find_prog_by_sec_insn(obj, obj->efile.text_shndx, sec_insn_off);
+ if (!subprog || subprog->sec_insn_off != sec_insn_off) {
+ pr_warn("prog '%s': map '%s': no function at .text+%zu for the pointer at offset %zu\n",
+ main_prog->name, map->name, ptrs[i].text_off, ptrs[i].sec_off);
+ return -LIBBPF_ERRNO__RELOC;
+ }
+
+ /*
+ * callx can't call global functions. Don't add one to the
+ * program only because the data points to it.
+ */
+ if (subprog->sym_global)
+ continue;
+
+ /* see the comment in bpf_object__reloc_code() */
+ if (subprog->sub_insn_off == 0) {
+ err = bpf_object__append_subprog_code(obj, main_prog, subprog);
+ if (err)
+ return err;
+ err = bpf_object__reloc_code(obj, main_prog, subprog);
+ if (err)
+ return err;
+ }
+ }
+ return 0;
+}
+
static int
bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
struct bpf_program *prog)
@@ -6851,6 +7128,20 @@ bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
for (insn_idx = 0; insn_idx < prog->sec_insn_cnt; insn_idx++) {
insn = &main_prog->insns[prog->sub_insn_off + insn_idx];
+ if (insn_is_func_ptrs_addr(insn)) {
+ /*
+ * The code that loads the address of the data might
+ * call any function that the data points to.
+ */
+ relo = find_prog_insn_relo(prog, insn_idx);
+ if (relo && relo->type == RELO_DATA) {
+ err = bpf_object__append_func_ptrs_code(obj, main_prog,
+ &obj->maps[relo->map_idx]);
+ if (err)
+ return err;
+ }
+ continue;
+ }
if (!insn_is_subprog_call(insn) && !insn_is_pseudo_func(insn))
continue;
@@ -7541,6 +7832,9 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
/* mark the insn, so it's recognized by insn_is_pseudo_func() */
if (relo->type == RELO_SUBPROG_ADDR)
insn[0].src_reg = BPF_PSEUDO_FUNC;
+ /* and by insn_is_func_ptrs_addr() */
+ if (relo->type == RELO_DATA && obj->maps[relo->map_idx].func_ptr_cnt)
+ insn[0].src_reg = BPF_PSEUDO_MAP_VALUE;
}
}
@@ -7757,6 +8051,98 @@ static int bpf_object__collect_map_relos(struct bpf_object *obj,
return 0;
}
+/*
+ * Collect pointers to functions in a read-only data section. They are
+ * R_BPF_64_ABS64 relocations against .text section, where the offset of
+ * a static function in the section is stored in place. Relocations in data
+ * sections were ignored before pointers to functions were supported. Those
+ * that are something else, e.g. pointers to data, still are.
+ */
+static int bpf_object__collect_rodata_relos(struct bpf_object *obj,
+ Elf64_Shdr *shdr, Elf_Data *data)
+{
+ size_t sec_idx = shdr->sh_info, sym_idx;
+ int i, nrels = shdr->sh_size / shdr->sh_entsize;
+ const char *relo_sec_name;
+ struct func_ptr *ptrs;
+ Elf_Data *scn_data;
+ Elf64_Sym *sym;
+ Elf64_Rel *rel;
+ __u64 addend;
+
+ relo_sec_name = elf_sec_str(obj, shdr->sh_name) ?: "<?>";
+ scn_data = obj->efile.secs[sec_idx].data;
+ if (!scn_data)
+ return -LIBBPF_ERRNO__FORMAT;
+
+ for (i = 0; i < nrels; i++) {
+ rel = elf_rel_by_idx(data, i);
+ if (!rel) {
+ pr_warn("sec '%s': failed to get relo #%d\n", relo_sec_name, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ sym_idx = ELF64_R_SYM(rel->r_info);
+ sym = elf_sym_by_idx(obj, sym_idx);
+ if (!sym) {
+ pr_warn("sec '%s': symbol #%zu not found for relo #%d\n",
+ relo_sec_name, sym_idx, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ if (ELF64_R_TYPE(rel->r_info) != R_BPF_64_ABS64 ||
+ !sym_is_subprog(sym, obj->efile.text_shndx)) {
+ pr_debug("sec '%s': relo #%d: not a pointer to a function, skipping...\n",
+ relo_sec_name, i);
+ continue;
+ }
+
+ /* the kernel finds aligned pointers only */
+ if (rel->r_offset % sizeof(__u64) || rel->r_offset >= scn_data->d_size ||
+ scn_data->d_size - rel->r_offset < sizeof(__u64)) {
+ pr_debug("sec '%s': relo #%d: unsupported offset 0x%zx, skipping...\n",
+ relo_sec_name, i, (size_t)rel->r_offset);
+ continue;
+ }
+
+ memcpy(&addend, scn_data->d_buf + rel->r_offset, sizeof(addend));
+ if (!is_native_endianness(obj))
+ addend = bswap_64(addend);
+ if ((sym->st_value + addend) % BPF_INSN_SZ) {
+ pr_debug("sec '%s': relo #%d: bad pointer to a function at offset %zu+%llu, skipping...\n",
+ relo_sec_name, i, (size_t)sym->st_value,
+ (unsigned long long)addend);
+ continue;
+ }
+
+ ptrs = libbpf_reallocarray(obj->func_ptrs, obj->func_ptr_cnt + 1, sizeof(*ptrs));
+ if (!ptrs)
+ return -ENOMEM;
+ obj->func_ptrs = ptrs;
+
+ ptrs[obj->func_ptr_cnt].sec_idx = sec_idx;
+ ptrs[obj->func_ptr_cnt].sec_off = rel->r_offset;
+ ptrs[obj->func_ptr_cnt].text_off = sym->st_value + addend;
+ obj->func_ptr_cnt++;
+
+ pr_debug("sec '%s': relo #%d: pointer at offset %zu to a function at .text+%zu\n",
+ relo_sec_name, i, (size_t)rel->r_offset, (size_t)(sym->st_value + addend));
+ }
+ return 0;
+}
+
+static int cmp_func_ptrs(const void *_a, const void *_b)
+{
+ const struct func_ptr *a = _a;
+ const struct func_ptr *b = _b;
+
+ if (a->sec_idx != b->sec_idx)
+ return a->sec_idx < b->sec_idx ? -1 : 1;
+ if (a->sec_off != b->sec_off)
+ return a->sec_off < b->sec_off ? -1 : 1;
+ return 0;
+}
+
static int bpf_object__collect_relos(struct bpf_object *obj)
{
int i, err;
@@ -7779,7 +8165,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return -LIBBPF_ERRNO__INTERNAL;
}
- if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
+ if (obj->efile.secs[idx].sec_type == SEC_RODATA)
+ err = bpf_object__collect_rodata_relos(obj, shdr, data);
+ else if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
err = bpf_object__collect_st_ops_relos(obj, shdr, data);
else if (idx == obj->efile.btf_maps_shndx)
err = bpf_object__collect_map_relos(obj, shdr, data);
@@ -7789,6 +8177,25 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return err;
}
+ /* sort by section, so that pointers in the data of a map are next to each other */
+ if (obj->func_ptr_cnt)
+ qsort(obj->func_ptrs, obj->func_ptr_cnt, sizeof(*obj->func_ptrs), cmp_func_ptrs);
+
+ for (i = 0; i < obj->nr_maps; i++) {
+ struct bpf_map *map = &obj->maps[i];
+ size_t j;
+
+ if (map->libbpf_type != LIBBPF_MAP_RODATA)
+ continue;
+ for (j = 0; j < obj->func_ptr_cnt; j++) {
+ if (obj->func_ptrs[j].sec_idx != map->sec_idx)
+ continue;
+ if (!map->func_ptr_cnt)
+ map->func_ptrs = &obj->func_ptrs[j];
+ map->func_ptr_cnt++;
+ }
+ }
+
bpf_object__sort_relos(obj);
return 0;
}
@@ -9208,8 +9615,17 @@ static int bpf_object_load(struct bpf_object *obj, int extra_log_level, const ch
* permit cross-endian creation of "light skeleton".
*/
if (obj->gen_loader) {
+ int nr_func_ptr_maps = 0, nr_progs = 0, i;
+
+ /* every program may get a copy of every map with pointers to functions */
+ for (i = 0; i < obj->nr_maps; i++)
+ if (obj->maps[i].autocreate && obj->maps[i].func_ptr_cnt)
+ nr_func_ptr_maps++;
+ for (i = 0; i < obj->nr_programs; i++)
+ if (obj->programs[i].autoload && !prog_is_subprog(obj, &obj->programs[i]))
+ nr_progs++;
bpf_gen__init(obj->gen_loader, obj->log_level | extra_log_level,
- obj->nr_programs, obj->nr_maps);
+ obj->nr_programs, obj->nr_maps, nr_func_ptr_maps * nr_progs);
} else if (!is_native_endianness(obj)) {
pr_warn("object '%s': loading non-native endianness is unsupported\n", obj->name);
return libbpf_err(-LIBBPF_ERRNO__ENDIAN);
@@ -9749,6 +10165,12 @@ void bpf_object__close(struct bpf_object *obj)
close(obj->jumptable_maps[i].fd);
zfree(&obj->jumptable_maps);
+ for (i = 0; i < obj->func_ptr_map_cnt; i++)
+ if (!obj->gen_loader)
+ close(obj->func_ptr_maps[i].fd);
+ zfree(&obj->func_ptr_maps);
+ zfree(&obj->func_ptrs);
+
if (obj->btf_module_allowlist) {
for (i = 0; i < obj->btf_module_allowlist_cnt; i++)
zfree(&obj->btf_module_allowlist[i]);
diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index 78f92c39290af..53f64a1a1f25a 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -2274,6 +2274,24 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
insn->imm += sec->dst_off / sizeof(struct bpf_insn);
else
insn->imm += sec->dst_off;
+ } else if (sym_type == R_BPF_64_ABS64 &&
+ (sec->shdr->sh_flags & SHF_EXECINSTR)) {
+ /*
+ * A pointer to a static function in a data section,
+ * which is stored in place as an offset of the
+ * function in its section. Data sections are kept
+ * in the byte order of the object.
+ */
+ void *ptr = dst_linked_sec->raw_data + dst_rel->r_offset;
+ __u64 off;
+
+ memcpy(&off, ptr, sizeof(off));
+ if (linker->swapped_endian)
+ off = bswap_64(off);
+ off += sec->dst_off;
+ if (linker->swapped_endian)
+ off = bswap_64(off);
+ memcpy(ptr, &off, sizeof(off));
} else {
pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
return -EINVAL;
diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
index 580d1d82c1867..3d92cdca62ed8 100644
--- a/tools/testing/selftests/bpf/Makefile.skel
+++ b/tools/testing/selftests/bpf/Makefile.skel
@@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
core_kern.c core_kern_overflow.c test_ringbuf.c \
test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
- test_ringbuf_overwrite.c
+ test_ringbuf_overwrite.c callx_rodata.c
LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
diff --git a/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
new file mode 100644
index 0000000000000..19f96792df5db
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
@@ -0,0 +1,192 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * The kernel replaces the offsets of functions in a frozen read-only map with
+ * their addresses when the program is loaded. The program has to be the only
+ * user of the map, and no other program can use the map after that.
+ */
+#include <test_progs.h>
+#include <linux/filter.h>
+#include <bpf/btf.h>
+
+#if defined(__x86_64__) || defined(__aarch64__)
+
+#define CALLEE_INSN 6
+#define DATA 0x1234
+
+/*
+ * main: r2 = &value; r2 = *(u64 *)(r2 + 8); r1 = 10; callx r2; exit
+ * add1: r0 = r1; r0 += 1; exit
+ *
+ * where value is { DATA, offset of add1 in the program }.
+ */
+static const struct bpf_insn callx_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_LDX_MEM(BPF_DW, BPF_REG_2, BPF_REG_2, 8),
+ BPF_MOV64_IMM(BPF_REG_1, 10),
+ BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, BPF_REG_2, 0, 0, 0),
+ BPF_EXIT_INSN(),
+ BPF_MOV64_REG(BPF_REG_0, BPF_REG_1),
+ BPF_ALU64_IMM(BPF_ADD, BPF_REG_0, 1),
+ BPF_EXIT_INSN(),
+};
+
+/* reads the data of the map */
+static const struct bpf_insn reader_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_LDX_MEM(BPF_DW, BPF_REG_0, BPF_REG_2, 0),
+ BPF_EXIT_INSN(),
+};
+
+/* refers to the map and is rejected: r0 is not set */
+static const struct bpf_insn bad_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_EXIT_INSN(),
+};
+
+static char log_buf[16 * 1024];
+static struct bpf_func_info func_info[2];
+static int btf_fd;
+
+static int create_map(void)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);
+ __u64 value[2] = { DATA, CALLEE_INSN * sizeof(struct bpf_insn) };
+ int fd, zero = 0;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, "callx_rodata", sizeof(int), sizeof(value), 1,
+ &opts);
+ if (!ASSERT_OK_FD(fd, "map_create"))
+ return -1;
+ if (!ASSERT_OK(bpf_map_update_elem(fd, &zero, value, 0), "map_update") ||
+ !ASSERT_OK(bpf_map_freeze(fd), "map_freeze")) {
+ close(fd);
+ return -1;
+ }
+ return fd;
+}
+
+static int load(const struct bpf_insn *prog_insns, int cnt, int map_fd)
+{
+ LIBBPF_OPTS(bpf_prog_load_opts, opts,
+ .log_buf = log_buf,
+ .log_size = sizeof(log_buf),
+ .log_level = 1,
+ );
+ struct bpf_insn insns[ARRAY_SIZE(callx_insns)];
+
+ if (prog_insns == callx_insns) {
+ /* add1() is referred to by the data only and is found through func_info */
+ opts.prog_btf_fd = btf_fd;
+ opts.func_info = func_info;
+ opts.func_info_cnt = 2;
+ opts.func_info_rec_size = sizeof(func_info[0]);
+ }
+ memcpy(insns, prog_insns, cnt * sizeof(insns[0]));
+ insns[0].imm = map_fd;
+ log_buf[0] = 0;
+ return bpf_prog_load(BPF_PROG_TYPE_SOCKET_FILTER, "callx_map", "GPL", insns, cnt, &opts);
+}
+
+#define LOAD(insns, map_fd) load(insns, ARRAY_SIZE(insns), map_fd)
+
+static void run(int prog_fd, int expected, const char *name)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, topts);
+ char pkt[64] = {};
+
+ topts.data_in = pkt;
+ topts.data_size_in = sizeof(pkt);
+ if (ASSERT_OK(bpf_prog_test_run_opts(prog_fd, &topts), name))
+ ASSERT_EQ(topts.retval, expected, name);
+}
+
+void test_callx_func_ptr_map(void)
+{
+ int int_id, proto_id, map_fd = -1, prog_fd = -1, reader_fd = -1, fd, zero = 0;
+ struct btf *btf;
+ __u64 value[2];
+
+ btf = btf__new_empty();
+ if (!ASSERT_OK_PTR(btf, "btf_new"))
+ return;
+ int_id = btf__add_int(btf, "int", 4, BTF_INT_SIGNED);
+ proto_id = btf__add_func_proto(btf, int_id);
+ func_info[0].insn_off = 0;
+ func_info[0].type_id = btf__add_func(btf, "main_prog", BTF_FUNC_GLOBAL, proto_id);
+ func_info[1].insn_off = CALLEE_INSN;
+ func_info[1].type_id = btf__add_func(btf, "add1", BTF_FUNC_STATIC, proto_id);
+ if (!ASSERT_GT(func_info[1].type_id, 0, "btf_add_func") ||
+ !ASSERT_OK(btf__load_into_kernel(btf), "btf_load"))
+ goto out;
+ btf_fd = btf__fd(btf);
+
+ /*
+ * Another program relies on what the map has: it's verified with
+ * the data folded into constants. Pointers to functions are not looked
+ * for in such map.
+ */
+ map_fd = create_map();
+ if (map_fd < 0)
+ goto out;
+ reader_fd = LOAD(reader_insns, map_fd);
+ if (!ASSERT_OK_FD(reader_fd, "load_reader"))
+ goto out;
+ fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_LT(fd, 0, "load_shared"))
+ close(fd);
+ ASSERT_HAS_SUBSTR(log_buf, "unreachable insn 6", "log_shared");
+ run(reader_fd, DATA, "run_reader");
+ close(reader_fd);
+ reader_fd = -1;
+ close(map_fd);
+
+ map_fd = create_map();
+ if (map_fd < 0)
+ goto out;
+
+ /* a program that is rejected is not a user, libbpf loads it again to get the log */
+ fd = LOAD(bad_insns, map_fd);
+ if (!ASSERT_LT(fd, 0, "load_bad"))
+ close(fd);
+
+ prog_fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_OK_FD(prog_fd, "load_callx")) {
+ printf("%s\n", log_buf);
+ goto out;
+ }
+ run(prog_fd, 11, "run_callx");
+
+ /* the offset of the function is gone from the map, the data is intact */
+ if (ASSERT_OK(bpf_map_lookup_elem(map_fd, &zero, value), "map_lookup")) {
+ ASSERT_EQ(value[0], DATA, "data");
+ ASSERT_NEQ(value[1], CALLEE_INSN * sizeof(struct bpf_insn), "pointer");
+ }
+
+ /* no other program can use the map now, another instance of the same one too */
+ fd = LOAD(reader_insns, map_fd);
+ if (!ASSERT_EQ(fd, -EBUSY, "load_reader_after"))
+ close(fd);
+ ASSERT_HAS_SUBSTR(log_buf, "has addresses of functions of another program", "log_reader");
+ fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_EQ(fd, -EBUSY, "load_second_instance"))
+ close(fd);
+
+ run(prog_fd, 11, "run_callx_again");
+out:
+ if (reader_fd >= 0)
+ close(reader_fd);
+ if (prog_fd >= 0)
+ close(prog_fd);
+ if (map_fd >= 0)
+ close(map_fd);
+ btf__free(btf);
+}
+
+#else
+
+void test_callx_func_ptr_map(void)
+{
+ test__skip();
+}
+
+#endif
diff --git a/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
new file mode 100644
index 0000000000000..5c60f352ff446
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
@@ -0,0 +1,55 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <test_progs.h>
+
+#include "callx_rodata.lskel.h"
+
+#if defined(__x86_64__) || defined(__aarch64__)
+
+static void run(int prog_fd, int expected, const char *name)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, topts);
+ char pkt[64] = {};
+
+ topts.data_in = pkt;
+ topts.data_size_in = sizeof(pkt);
+ if (ASSERT_OK(bpf_prog_test_run_opts(prog_fd, &topts), name))
+ ASSERT_EQ(topts.retval, expected, name);
+}
+
+/*
+ * Every program gets its own copy of .rodata with the offsets of its functions.
+ * The copies have what user space puts into .rodata before the load.
+ */
+void test_callx_rodata_lskel(void)
+{
+ struct callx_rodata_lskel *skel;
+
+ skel = callx_rodata_lskel__open();
+ if (!ASSERT_OK_PTR(skel, "open"))
+ return;
+
+ skel->rodata->bias = 7;
+
+ if (!ASSERT_OK(callx_rodata_lskel__load(skel), "load"))
+ goto out;
+
+ skel->bss->op_idx = 0;
+ run(skel->progs.select_op.prog_fd, 17, "add_bias");
+ skel->bss->op_idx = 1;
+ run(skel->progs.select_op.prog_fd, 30, "mul3");
+ skel->bss->op_idx = 2;
+ run(skel->progs.select_op.prog_fd, -1, "out_of_range");
+ /* mul3(add_bias(4)) */
+ run(skel->progs.both_ops.prog_fd, 33, "both_ops");
+out:
+ callx_rodata_lskel__destroy(skel);
+}
+
+#else
+
+void test_callx_rodata_lskel(void)
+{
+ test__skip();
+}
+
+#endif
diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 4f1e1c1cd5ab3..3e07b957ee80d 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -27,6 +27,8 @@
#include "verifier_btf_ctx_access.skel.h"
#include "verifier_btf_unreliable_prog.skel.h"
#include "verifier_call_large_imm.skel.h"
+#include "verifier_callx.skel.h"
+#include "verifier_callx_rodata.skel.h"
#include "verifier_cfg.skel.h"
#include "verifier_cgroup_inv_retcode.skel.h"
#include "verifier_cgroup_skb.skel.h"
@@ -196,6 +198,8 @@ void test_verifier_bswap(void) { RUN(verifier_bswap); }
void test_verifier_btf_ctx_access(void) { RUN(verifier_btf_ctx_access); }
void test_verifier_btf_unreliable_prog(void) { RUN(verifier_btf_unreliable_prog); }
void test_verifier_call_large_imm(void) { RUN(verifier_call_large_imm); }
+void test_verifier_callx(void) { RUN(verifier_callx); }
+void test_verifier_callx_rodata(void) { RUN(verifier_callx_rodata); }
void test_verifier_cfg(void) { RUN(verifier_cfg); }
void test_verifier_cgroup_inv_retcode(void) { RUN(verifier_cgroup_inv_retcode); }
void test_verifier_cgroup_skb(void) { RUN(verifier_cgroup_skb); }
diff --git a/tools/testing/selftests/bpf/progs/callx_rodata.c b/tools/testing/selftests/bpf/progs/callx_rodata.c
new file mode 100644
index 0000000000000..7475c87cdaed4
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/callx_rodata.c
@@ -0,0 +1,43 @@
+// SPDX-License-Identifier: GPL-2.0
+/* callx through pointers to functions in .rodata, loaded by light skeleton */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+
+typedef int (*op_fn)(int);
+
+/* set by user space before the programs are loaded, it's in .rodata too */
+const volatile int bias = 1;
+
+int op_idx;
+
+static __noinline int add_bias(int x)
+{
+ return x + bias;
+}
+
+static __noinline int mul3(int x)
+{
+ return x * 3;
+}
+
+static op_fn const ops[] = { add_bias, mul3 };
+
+SEC("socket")
+int select_op(void *ctx)
+{
+ unsigned int i = op_idx;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ return ops[i](10);
+}
+
+/* functions have other offsets in this program */
+SEC("socket")
+int both_ops(void *ctx)
+{
+ return ops[1](ops[0](4));
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_callx.c b/tools/testing/selftests/bpf/progs/verifier_callx.c
new file mode 100644
index 0000000000000..ec63d2107e6c9
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_callx.c
@@ -0,0 +1,984 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Tests for callx: indirect calls of bpf subprogs */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../../../include/linux/filter.h"
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+#define CALLX_INSN(DST, SRC, OFF, IMM) \
+ BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, DST, SRC, OFF, IMM)
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, int);
+ __type(value, long long);
+} map_array SEC(".maps");
+
+struct val_with_lock {
+ struct bpf_spin_lock lock;
+ int cnt;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, int);
+ __type(value, struct val_with_lock);
+} map_lock SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(int));
+ __uint(value_size, sizeof(int));
+} map_prog SEC(".maps");
+
+__naked __noinline __used
+static unsigned long add1(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long add2(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 2;"
+ "exit;"
+ );
+}
+
+/* apply(fn, x) { return fn(x); } */
+__naked __noinline __used
+static unsigned long apply(void)
+{
+ asm volatile (
+ "r3 = r1;"
+ "r1 = r2;"
+ "callx r3;"
+ "exit;"
+ );
+}
+
+SEC("socket")
+__success __retval(6)
+__naked void callx_basic(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* callx is printed by the verifier log and xlated dump */
+SEC("socket")
+__success __log_level(2)
+__msg("(8d) callx r2")
+__xlated("callx r2")
+__naked void callx_disasm(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* different callees are called by the same callx on different paths */
+SEC("socket")
+__success __retval(11)
+__naked void callx_two_callees(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r2 = %[add1] ll;"
+ "if r6 == 0 goto +2;"
+ "r2 = %[add2] ll;"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add1),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+/* pointer to a function is passed as an argument */
+SEC("socket")
+__success __retval(45)
+__naked void callx_fn_as_arg(void)
+{
+ asm volatile (
+ "r1 = %[add2] ll;"
+ "r2 = 40;"
+ "call apply;"
+ "r6 = r0;"
+ "r1 = %[add1] ll;"
+ "r2 = 2;"
+ "call apply;"
+ "r0 += r6;"
+ "exit;"
+ :
+ : __imm_addr(add1),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long get_add1(void)
+{
+ asm volatile (
+ "r0 = %[add1] ll;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* pointer to a function is returned from a subprog and called via r0 */
+SEC("socket")
+__success __retval(2)
+__naked void callx_r0(void)
+{
+ asm volatile (
+ "call get_add1;"
+ "r1 = 1;"
+ "callx r0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* pointer to a function survives spill/fill */
+SEC("socket")
+__success __retval(9)
+__naked void callx_spill_fill(void)
+{
+ asm volatile (
+ "r2 = %[add2] ll;"
+ "*(u64 *)(r10 - 8) = r2;"
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 7;"
+ "r9 = *(u64 *)(r10 - 8);"
+ "callx r9;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long clobber_callee_saved(void)
+{
+ asm volatile (
+ "r6 = 100;"
+ "r7 = 100;"
+ "r8 = 100;"
+ "r9 = 100;"
+ "r0 = r1;"
+ "exit;"
+ );
+}
+
+/* r6-r9 are preserved across callx */
+SEC("socket")
+__success __retval(11)
+__naked void callx_callee_saved_regs(void)
+{
+ asm volatile (
+ "r6 = 1;"
+ "r7 = 2;"
+ "r8 = 3;"
+ "r9 = 4;"
+ "r1 = 1;"
+ "r2 = %[clobber_callee_saved] ll;"
+ "callx r2;"
+ "r0 += r6;"
+ "r0 += r7;"
+ "r0 += r8;"
+ "r0 += r9;"
+ "exit;"
+ :
+ : __imm_addr(clobber_callee_saved)
+ : __clobber_all);
+}
+
+/* r1-r5 are scratched by callx */
+SEC("socket")
+__failure __msg("R1 !read_ok")
+__naked void callx_scratches_args(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "r0 = r1;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* return value of the callee is tracked */
+SEC("socket")
+__success __log_level(2)
+__msg("R0=6")
+__naked void callx_retval_is_tracked(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long write42(void)
+{
+ asm volatile (
+ "r2 = 42;"
+ "*(u64 *)(r1 + 0) = r2;"
+ "r0 = 0;"
+ "exit;"
+ );
+}
+
+/* callee writes into the stack of the caller */
+SEC("socket")
+__success __retval(42)
+__naked void callx_callee_writes_caller_stack(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[write42] ll;"
+ "callx r2;"
+ "r0 = *(u64 *)(r10 - 8);"
+ "exit;"
+ :
+ : __imm_addr(write42)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R1 has type scalar, expected func")
+__naked void callx_scalar(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "callx r1;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R2 !read_ok")
+__naked void callx_uninit_reg(void)
+{
+ asm volatile (
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R10 has type fp, expected func")
+__naked void callx_fp(void)
+{
+ asm volatile (
+ "callx r10;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R1 has type map_value, expected func")
+__naked void callx_map_value(void)
+{
+ asm volatile (
+ "r1 = %[map_array] ll;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r3 = 0;"
+ "*(u32 *)(r2 + 0) = r3;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 == 0 goto 1f;"
+ "r1 = r0;"
+ "callx r1;"
+ "1:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm_addr(map_array)
+ : __clobber_all);
+}
+
+/* the address of a subprog can't be modified before the call */
+SEC("socket")
+__failure __msg("dereference of modified func ptr R2 off=8 disallowed")
+__naked void callx_modified_ptr(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "r2 += 8;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("variable func access var_off=")
+__naked void callx_variable_ptr(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 8;"
+ "r2 = %[add1] ll;"
+ "r2 += r0;"
+ "r1 = 5;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add1)
+ : __clobber_all);
+}
+
+__noinline __used
+int global_add3(int x)
+{
+ return x + 3;
+}
+
+/* only static subprogs can be called via callx */
+SEC("socket")
+__failure __msg("callback function not static")
+__naked void callx_global_func(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[global_add3] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(global_add3)
+ : __clobber_all);
+}
+
+#define DEFINE_CALLX_RESERVED_FIELDS_PROG(NAME, SRC_REG, OFF, IMM) \
+ SEC("socket") \
+ __failure __msg("BPF_CALL|BPF_X uses reserved fields") \
+ __naked void callx_reserved_field_ ## NAME(void) \
+ { \
+ asm volatile ( \
+ "r1 = 5;" \
+ "r2 = %[add1] ll;" \
+ ".8byte %[callx_r2];" \
+ "exit;" \
+ : \
+ : __imm_addr(add1), \
+ __imm_insn(callx_r2, CALLX_INSN(BPF_REG_2, (SRC_REG), (OFF), (IMM))) \
+ : __clobber_all); \
+ }
+
+DEFINE_CALLX_RESERVED_FIELDS_PROG(src_reg, BPF_REG_1, 0, 0)
+DEFINE_CALLX_RESERVED_FIELDS_PROG(off, BPF_REG_0, 1, 0)
+DEFINE_CALLX_RESERVED_FIELDS_PROG(imm, BPF_REG_0, 0, 1)
+
+SEC("socket")
+__failure __msg("unknown opcode 8e")
+__naked void callx_jmp32(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ ".8byte %[callx32_r2];"
+ "exit;"
+ :
+ : __imm_addr(add1),
+ __imm_insn(callx32_r2,
+ BPF_RAW_INSN(BPF_JMP32 | BPF_CALL | BPF_X, BPF_REG_2, 0, 0, 0))
+ : __clobber_all);
+}
+
+/* similar to calls of static subprogs callx is allowed under a lock */
+SEC("tc")
+__success __retval(3)
+__naked void callx_under_lock(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_lock] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 != 0 goto 1f;"
+ "exit;"
+ "1:"
+ "r6 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_lock];"
+ "r1 = 1;"
+ "r2 = %[add2] ll;"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_unlock];"
+ "r0 = r7;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm(bpf_spin_lock),
+ __imm(bpf_spin_unlock),
+ __imm_addr(map_lock),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+/* helpers are still not allowed under a lock in the callee */
+__naked __noinline __used
+static unsigned long call_helper(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+SEC("tc")
+__failure __msg("function calls are not allowed while holding a lock")
+__naked void callx_helper_under_lock(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_lock] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 != 0 goto 1f;"
+ "exit;"
+ "1:"
+ "r6 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_lock];"
+ "r2 = %[call_helper] ll;"
+ "callx r2;"
+ "r1 = r6;"
+ "call %[bpf_spin_unlock];"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm(bpf_spin_lock),
+ __imm(bpf_spin_unlock),
+ __imm_addr(map_lock),
+ __imm_addr(call_helper)
+ : __clobber_all);
+}
+
+/* self(fn) { return fn(fn); } */
+__naked __noinline __used
+static unsigned long self(void)
+{
+ asm volatile (
+ "callx r1;"
+ "exit;"
+ );
+}
+
+/* unbounded recursion is caught by the main verification pass */
+SEC("socket")
+__failure __msg("frames is too deep")
+__naked void callx_unbounded_recursion(void)
+{
+ asm volatile (
+ "r1 = %[self] ll;"
+ "call self;"
+ "exit;"
+ :
+ : __imm_addr(self)
+ : __clobber_all);
+}
+
+/* countdown(fn, n) { return n ? fn(fn, n - 1) : 0; } */
+__naked __noinline __used
+static unsigned long countdown(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r2 == 0 goto 1f;"
+ "r2 += -1;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+/*
+ * The depth of the recursion is known to the main verification pass,
+ * but recursive calls are not allowed regardless.
+ */
+SEC("socket")
+__failure __msg("recursive call from countdown() to countdown()")
+__naked void callx_bounded_recursion(void)
+{
+ asm volatile (
+ "r1 = %[countdown] ll;"
+ "r2 = 2;"
+ "call countdown;"
+ "exit;"
+ :
+ : __imm_addr(countdown)
+ : __clobber_all);
+}
+
+/* ping(fn1, fn2, n) { return n ? fn2(fn2, fn1, n - 1) : 0; } */
+__naked __noinline __used
+static unsigned long ping(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r3 == 0 goto 1f;"
+ "r3 += -1;"
+ "r4 = r1;"
+ "r1 = r2;"
+ "r2 = r4;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long pong(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r3 == 0 goto 1f;"
+ "r3 += -1;"
+ "r4 = r1;"
+ "r1 = r2;"
+ "r2 = r4;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+SEC("socket")
+__failure __msg("recursive call from")
+__naked void callx_mutual_recursion(void)
+{
+ asm volatile (
+ "r1 = %[ping] ll;"
+ "r2 = %[pong] ll;"
+ "r3 = 3;"
+ "call ping;"
+ "exit;"
+ :
+ : __imm_addr(ping),
+ __imm_addr(pong)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long use_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "exit;"
+ );
+}
+
+/* stack of the callee of callx is accounted */
+SEC("socket")
+__failure __msg("combined stack size of 2 calls is")
+__naked void callx_stack_depth(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "r2 = %[use_stack_304] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+/* apply_stack_304(fn) { char buf[304]; return fn(); } */
+__naked __noinline __used
+static unsigned long apply_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "callx r1;"
+ "exit;"
+ );
+}
+
+/*
+ * The address of use_stack_304() is taken by the main prog that doesn't
+ * use stack, but it is called from apply_stack_304().
+ */
+SEC("socket")
+__failure __msg("combined stack size of 3 calls is")
+__naked void callx_stack_depth_nested(void)
+{
+ asm volatile (
+ "r1 = %[use_stack_304] ll;"
+ "call apply_stack_304;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+SEC("socket")
+__success __retval(0)
+__naked void callx_stack_depth_ok(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 200) = r0;"
+ "r2 = %[use_stack_304] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long do_tail_call(void)
+{
+ asm volatile (
+ "r2 = %[map_prog] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_tail_call),
+ __imm_addr(map_prog)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("tail_calls are not allowed in functions called via callx")
+__naked void callx_tail_call_in_callee(void)
+{
+ asm volatile (
+ "r2 = %[do_tail_call] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(do_tail_call)
+ : __clobber_all);
+}
+
+/* tail call in the caller of callx is fine */
+SEC("socket")
+__success __retval(3)
+__naked void callx_tail_call_in_caller(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "r1 = 2;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = r6;"
+ "r2 = %[map_prog] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "r0 = r7;"
+ "exit;"
+ :
+ : __imm(bpf_tail_call),
+ __imm_addr(map_prog),
+ __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* read_idx(p) { return ((char *)map_value)[*p]; } */
+__naked __noinline __used
+static unsigned long read_idx(void)
+{
+ asm volatile (
+ "r6 = *(u64 *)(r1 + 0);"
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_array] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 == 0 goto 1f;"
+ "r0 += r6;"
+ "r0 = *(u8 *)(r0 + 0);"
+ "1:"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm_addr(map_array)
+ : __clobber_all);
+}
+
+/*
+ * Stack slots of the caller that might be read by the callee of callx
+ * have to be considered alive at the checkpoints before callx and
+ * inside of the callee. Otherwise the state with fp-8 == 1000 is pruned
+ * and out of bounds access in read_idx() goes unnoticed.
+ */
+SEC("socket")
+__failure __msg("invalid access to map value, value_size=8 off=1000 size=1")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 1000;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[read_idx] ll;"
+ "callx r2;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* in bounds access in read_idx() is fine */
+SEC("socket")
+__success __retval(0)
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack_ok(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 7;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[read_idx] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* same as above, but the pointer to the stack is passed through one more frame */
+SEC("socket")
+__failure __msg("invalid access to map value, value_size=8 off=1000 size=1")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack_nested(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 1000;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = %[read_idx] ll;"
+ "r2 = r10;"
+ "r2 += -8;"
+ "call apply;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* registers that are constant before callx are not constant after it */
+SEC("socket")
+__success __retval(1)
+__naked void callx_clobbers_const_regs(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "r1 = 0;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ /* dead branch pruning must not assume that r0 is still 0 */
+ "if r0 == 0 goto 1f;"
+ "r0 = 1;"
+ "exit;"
+ "1:"
+ "r0 = 2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* scalar argument passed through callx is tracked precisely */
+__naked __noinline __used
+static unsigned long identity(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "exit;"
+ );
+}
+
+long long vals[] SEC(".data.vals") = {1, 2, 3, 4};
+
+SEC("socket")
+__success __log_level(2)
+__msg("mark_precise: frame0: regs=r0 stack= before 12: (95) exit")
+__msg("mark_precise: frame1: regs=r0 stack= before 11: (bf) r0 = r1")
+__msg("mark_precise: frame1: regs=r1 stack= before 4: (8d) callx r2")
+__msg("mark_precise: frame0: regs=r1 stack= before 3: (bf) r1 = r6")
+__msg("mark_precise: frame0: regs=r6 stack= before 2: (b7) r6 = 3")
+__retval(4)
+__naked void callx_precision(void)
+{
+ asm volatile (
+ "r2 = %[identity] ll;"
+ "r6 = 3;"
+ "r1 = r6;"
+ "callx r2;"
+ "r0 *= 8;"
+ "r1 = %[vals] ll;"
+ "r1 += r0;"
+ "r0 = *(u64 *)(r1 + 0);"
+ "exit;"
+ :
+ : __imm_addr(identity),
+ __imm_addr(vals)
+ : __clobber_all);
+}
+
+/* function pointers in C */
+
+typedef int (*op_fn)(int);
+
+static __noinline int mul3(int x)
+{
+ return x * 3;
+}
+
+static __noinline int sub7(int x)
+{
+ return x - 7;
+}
+
+static __noinline int apply_op(op_fn op, int x)
+{
+ return op(x);
+}
+
+SEC("socket")
+__success __retval(36)
+int callx_c_fn_as_arg(void *ctx)
+{
+ /* (5 * 3) + (28 - 7) */
+ return apply_op(mul3, 5) + apply_op(sub7, 28);
+}
+
+SEC("socket")
+__success __retval(30)
+int callx_c_select(void *ctx)
+{
+ __u32 rnd = bpf_get_prandom_u32() & 1;
+ op_fn op = rnd ? mul3 : sub7;
+ int x = rnd ? 10 : 37;
+
+ /* 10 * 3 or 37 - 7 */
+ return op(x);
+}
+
+struct ops {
+ op_fn first;
+ op_fn second;
+ int bias;
+};
+
+static __noinline int run_ops(const struct ops *ops, int x)
+{
+ return ops->second(ops->first(x)) + ops->bias;
+}
+
+SEC("socket")
+__success __retval(100)
+int callx_c_ops_on_stack(void *ctx)
+{
+ struct ops a, b;
+
+ /* avoid an initializer with function pointers in .rodata */
+ a.first = mul3;
+ a.second = sub7;
+ a.bias = 10;
+ b.first = sub7;
+ b.second = mul3;
+ b.bias = 1;
+
+ /* ((7 * 3) - 7 + 10) + ((32 - 7) * 3 + 1) */
+ return run_ops(&a, 7) + run_ops(&b, 32);
+}
+
+#else
+
+SEC("socket")
+__success
+int dummy(void *ctx)
+{
+ return 0;
+}
+
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
new file mode 100644
index 0000000000000..419a90b1f894e
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
@@ -0,0 +1,651 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Tests for callx through pointers to functions in read-only data */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../../../include/linux/filter.h"
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+/*
+ * Read-only data with pointers to functions, where the compiler puts tables
+ * of functions, structures of operations and vtables. libbpf resolves
+ * a pointer to the offset of the function in the program and the kernel
+ * recognizes it by that value.
+ */
+#define DATA(SECTION, NAME, ...) \
+ ".pushsection " SECTION ",@progbits;" \
+ ".balign 8;" \
+ #NAME "_%=:" \
+ __VA_ARGS__ \
+ ".type " #NAME "_%=, @object;" \
+ ".size " #NAME "_%=, .-" #NAME "_%=;" \
+ ".popsection;"
+
+#define RODATA(NAME, ...) DATA(".rodata,\"a\"", NAME, __VA_ARGS__)
+
+#define FUNC_TABLE2(NAME, F0, F1) RODATA(NAME, ".quad " #F0 "; .quad " #F1 ";")
+
+__naked __noinline __used
+static unsigned long add1(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long add2(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 2;"
+ "exit;"
+ );
+}
+
+/* the second element of the table is called */
+SEC("socket")
+__success __retval(12)
+__naked void callx_rodata_const_index(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * The program reads the address of a function from the data, so pointers to
+ * functions in data are recognized only for programs that may leak pointers.
+ * Otherwise nothing refers to the functions.
+ */
+SEC("socket")
+__success __retval(12)
+__failure_unpriv __msg_unpriv("unreachable insn")
+__caps_unpriv(CAP_BPF)
+__naked void callx_rodata_needs_perfmon(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the address of an element is used instead of the address of the table */
+SEC("socket")
+__success __retval(12)
+__naked void callx_rodata_elem_addr(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= + 8 ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* pointers to functions are mixed with other data like in a vtable */
+SEC("socket")
+__success __retval(23)
+__log_level(2)
+__msg("r1 = *(u64 *)(r6 +0) ; R1=7")
+__msg("r2 = *(u64 *)(r6 +8) ; R2=func()")
+__naked void callx_rodata_mixed_with_data(void)
+{
+ asm volatile (
+ RODATA(vt, ".quad 7; .quad add1; .quad 13; .quad add2;")
+ "r6 = vt_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ /* add1(7) */
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = *(u64 *)(r6 + 16);"
+ "r2 = *(u64 *)(r6 + 24);"
+ /* add2(13) */
+ "callx r2;"
+ "r0 += r7;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * Position independent code keeps constants with pointers in .data.rel.ro.
+ * libbpf treats it as read-only data.
+ */
+SEC("socket")
+__success __retval(23)
+__naked void callx_data_rel_ro(void)
+{
+ asm volatile (
+ DATA(".data.rel.ro,\"aw\"", vt, ".quad 7; .quad add1; .quad 13; .quad add2;")
+ "r6 = vt_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = *(u64 *)(r6 + 16);"
+ "r2 = *(u64 *)(r6 + 24);"
+ "callx r2;"
+ "r0 += r7;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* an array of structures: the index selects one of the functions, not the data */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_array_of_structs(void)
+{
+ asm volatile (
+ RODATA(arr, ".quad add1; .quad 0x1111; .quad add2; .quad 0x2222;")
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 4;"
+ "r2 = arr_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* dynamic dispatch: which vtable is used is not known until run time */
+SEC("socket")
+__success __retval(2)
+__naked void callx_rodata_two_vtables(void)
+{
+ asm volatile (
+ RODATA(vt_a, ".quad 1; .quad add1;")
+ RODATA(vt_b, ".quad 0; .quad add2;")
+ "call %[bpf_get_prandom_u32];"
+ "r6 = vt_a_%= ll;"
+ "r0 &= 1;"
+ "if r0 == 0 goto +2;"
+ "r6 = vt_b_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ /* add1(1) or add2(0) */
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* both functions are verified, either of them is called */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_var_index(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* a table that has no symbol, the compiler generates such for a switch statement */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_no_symbol(void)
+{
+ asm volatile (
+ ".pushsection .rodata,\"a\",@progbits;"
+ ".balign 8;"
+ ".Lanon_%=:"
+ ".quad add1;"
+ ".quad add2;"
+ ".popsection;"
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 3;"
+ "r2 = .Lanon_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* the first instruction of the function is removed by the verifier */
+__naked __noinline __used
+static unsigned long nop_add1(void)
+{
+ asm volatile (
+ "goto +0;"
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+/* the pointer follows the function when instructions are removed */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_func_starts_with_nop(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad nop_add1; .quad 0;")
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* can't be called with a scalar in r1 */
+__naked __noinline __used
+static unsigned long deref_r1(void)
+{
+ asm volatile (
+ "r0 = *(u64 *)(r1 + 0);"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long ret0(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "exit;"
+ );
+}
+
+/* every function that might be called is verified */
+SEC("socket")
+__failure __msg("R1 invalid mem access 'scalar'")
+__naked void callx_rodata_all_callees_verified(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, deref_r1)
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 1;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 0;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* a function that is never called is not verified, it's dead code */
+SEC("socket")
+__success __retval(0)
+__naked void callx_rodata_unused_callee(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, deref_r1)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 0;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the index may select what is not a pointer to a function */
+SEC("socket")
+__failure __msg("overlaps with a pointer to a function")
+__naked void callx_rodata_index_beyond_table(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad add2; .quad 0x1234; .quad 0x5678;")
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 3;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* the address of a function is not known until the program is jitted */
+SEC("socket")
+__failure __msg("read of 4 bytes at offset")
+__msg("overlaps with a pointer to a function")
+__naked void callx_rodata_partial_read(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r0 = *(u32 *)(r2 + 0);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("read of 8 bytes at offset")
+__msg("overlaps with a pointer to a function")
+__naked void callx_rodata_misaligned_read(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad add2; .quad 0;")
+ "r2 = tbl_%= ll;"
+ "r0 = *(u64 *)(r2 + 4);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the data next to a pointer is still known to the verifier */
+SEC("socket")
+__success __retval(0x1234)
+__log_level(2)
+__msg("R0=4660")
+__naked void callx_rodata_data_is_const(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad 0x1234;")
+ "r2 = tbl_%= ll;"
+ "r0 = *(u64 *)(r2 + 8);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the pointer read from the data can't be modified */
+SEC("socket")
+__failure __msg("dereference of modified func ptr R2 off=8 disallowed")
+__naked void callx_rodata_modified_ptr(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r2 += 8;"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+__noinline __used
+int global_add3(int x)
+{
+ return x + 3;
+}
+
+/* a pointer to a global function is not recognized, it's a number */
+SEC("socket")
+__failure __msg("R2 has type scalar, expected func")
+__naked void callx_rodata_global_func(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, global_add3)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* selfcall(n) { return n ? tbl[0](n - 1) : 0; }, where tbl[0] == selfcall */
+__naked __noinline __used
+static unsigned long selfcall(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, selfcall, ret0)
+ "r0 = 0;"
+ "if r1 == 0 goto 1f;"
+ "r1 += -1;"
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "callx r2;"
+ "1:"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("recursive call from selfcall() to selfcall()")
+__naked void callx_rodata_recursion(void)
+{
+ asm volatile (
+ "r1 = 2;"
+ "call selfcall;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long use_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "exit;"
+ );
+}
+
+/* stack of all possible callees is accounted */
+SEC("socket")
+__failure __msg("combined stack size of 2 calls is")
+__naked void callx_rodata_stack_depth(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, use_stack_304)
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 1;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* pointers to functions in C */
+
+typedef int (*op_fn)(int);
+
+#define DEFINE_OP(N) static __noinline int op##N(int x) { return x * (N + 2) + N; }
+
+DEFINE_OP(0) DEFINE_OP(1) DEFINE_OP(2) DEFINE_OP(3)
+DEFINE_OP(4) DEFINE_OP(5) DEFINE_OP(6) DEFINE_OP(7)
+DEFINE_OP(8) DEFINE_OP(9) DEFINE_OP(10) DEFINE_OP(11)
+DEFINE_OP(12) DEFINE_OP(13) DEFINE_OP(14) DEFINE_OP(15)
+
+static op_fn const ops[] = {
+ op0, op1, op2, op3, op4, op5, op6, op7,
+ op8, op9, op10, op11, op12, op13, op14, op15,
+};
+
+int op_idx = 11;
+
+SEC("socket")
+__success __retval(50)
+int callx_c_table(void *ctx)
+{
+ unsigned int i = op_idx;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ /* op11(3) = 3 * 13 + 11 */
+ return ops[i](3);
+}
+
+SEC("socket")
+__success __retval(50)
+int callx_c_table_null_check(void *ctx)
+{
+ unsigned int i = op_idx;
+ op_fn op;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ /* the address of a function is not known until the program is jitted */
+ op = ops[i];
+ if (!op)
+ return -2;
+ return op(3);
+}
+
+/* a structure of operations, where pointers to functions are mixed with data */
+struct shape_ops {
+ int id;
+ op_fn area;
+ long scale;
+ op_fn perimeter;
+};
+
+static const struct shape_ops square_ops = { 1, op1, 10, op2 };
+static const struct shape_ops circle_ops = { 2, op3, 20, op1 };
+
+static __noinline int use_shape(const struct shape_ops *ops, int x)
+{
+ return ops->area(x) * ops->scale + ops->perimeter(ops->id);
+}
+
+SEC("socket")
+__success __retval(433)
+int callx_c_ops_mixed_with_data(void *ctx)
+{
+ /*
+ * square: op1(2) * 10 + op2(1) = 7 * 10 + 6 = 76
+ * circle: op3(3) * 20 + op1(2) = 18 * 20 + 7 = 367
+ * minus 10 when op_idx is not what it is set to
+ */
+ return use_shape(&square_ops, 2) + use_shape(&circle_ops, 3) - (op_idx == 11 ? 10 : 0);
+}
+
+/* the ops are selected at run time */
+SEC("socket")
+__success __retval(367)
+int callx_c_ops_selected(void *ctx)
+{
+ const struct shape_ops *ops = op_idx == 11 ? &circle_ops : &square_ops;
+
+ return use_shape(ops, 3);
+}
+
+/* the compiler might turn the switch into a table that has no symbol */
+static __noinline int call_by_switch(unsigned int idx, int x)
+{
+ op_fn op;
+
+ switch (idx) {
+ case 0:
+ op = op8;
+ break;
+ case 1:
+ op = op9;
+ break;
+ case 2:
+ op = op10;
+ break;
+ case 3:
+ op = op11;
+ break;
+ case 4:
+ op = op12;
+ break;
+ case 5:
+ op = op13;
+ break;
+ case 6:
+ op = op14;
+ break;
+ case 7:
+ op = op15;
+ break;
+ default:
+ return -1;
+ }
+ return op(x);
+}
+
+SEC("socket")
+__success __retval(50)
+int callx_c_switch_table(void *ctx)
+{
+ /* op11(3) = 3 * 13 + 11 */
+ return call_by_switch(op_idx - 8, 3);
+}
+
+/*
+ * Misaligned pointers to functions are ignored, the rest of the data is
+ * accessible as before.
+ */
+static const struct {
+ char tag;
+ op_fn op;
+} __attribute__((packed)) packed_ops[] = {
+ { 5, op0 },
+ { 7, op1 },
+};
+
+SEC("socket")
+__success __retval(7)
+int callx_c_packed_struct(void *ctx)
+{
+ return packed_ops[op_idx & 1].tag;
+}
+
+#else
+
+SEC("socket")
+__success
+int dummy(void *ctx)
+{
+ return 0;
+}
+
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c b/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
index 966f493487874..27fbe54e8795a 100644
--- a/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
+++ b/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
@@ -536,4 +536,35 @@ int return_from_void_global(struct __sk_buff *skb)
return 0;
}
+int global_calls_loop(int x);
+
+static __noinline int static_calls_global(int x)
+{
+ return global_calls_loop(x);
+}
+
+static __noinline int loop_cb_calls_static(u32 i, void *ctx)
+{
+ return static_calls_global(i);
+}
+
+__noinline int global_calls_loop(int x)
+{
+ bpf_loop(1, loop_cb_calls_static, NULL, 0);
+ return 0;
+}
+
+/*
+ * loop_cb_calls_static() -> static_calls_global() -> global_calls_loop() ->
+ * bpf_loop() -> loop_cb_calls_static() is an unbounded recursion that the
+ * main verification pass can't see, because it doesn't follow calls of global
+ * functions. None of the functions use stack.
+ */
+SEC("?raw_tp")
+__failure __msg("recursive call from global_calls_loop() to loop_cb_calls_static()")
+int recursion_via_global_func_and_callback(const void *ctx)
+{
+ return loop_cb_calls_static(0, NULL);
+}
+
char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/verifier/basic_call.c b/tools/testing/selftests/bpf/verifier/basic_call.c
index a8c6ab4c1622f..0f93c4551f235 100644
--- a/tools/testing/selftests/bpf/verifier/basic_call.c
+++ b/tools/testing/selftests/bpf/verifier/basic_call.c
@@ -4,7 +4,7 @@
BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, 0, 0, 0, 0),
BPF_EXIT_INSN(),
},
- .errstr = "unknown opcode 8d",
+ .errstr = "R0 !read_ok",
.result = REJECT,
},
{
|
| 25/3 | 2026/09/24 01:46 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 58923 output: 20 thoughts: 2502 Thoughts: **KMSAN Fuzzing Session Assessment for BPF callx Support** Okay, let's analyze this patch series and determine if a dedicated KMSAN fuzzing session is warranted. My understanding is that KMSAN specifically flags reads of uninitialized memory in *kernel C code*, not in BPF programs themselves, or memory not managed by the kernel. The goal is to identify if these changes introduce any new risk of uninitialized memory usage within the kernel code *itself*, or new ways for kernel to user memory leaks. The primary focus should be on the four specific scenarios where KMSAN is relevant: 1) Kernel data sent to userspace, potentially with uninitialized fields, 2) Conditional logic depending on uninitialized variables, 3) Allocation of complex data structures with uninitialized fields that might be read later, and 4) Bounds checks or lengths that expose uninitialized bytes. Examining the diff, this patch implements `BPF_CALL | BPF_X | BPF_JMP` which means indirect function calls (callx) in BPF programs. I'm seeing changes in: * Documentation. * Architecture-specific JIT compilers (x86 and arm64). This emits the actual indirect call instruction. * `include/linux/bpf.h`, including the addition of `unsigned long user;` to `struct bpf_map`. Crucially, I need to verify how `struct bpf_map` is allocated. I need to be sure any new fields are appropriately initialized at allocation. * `include/linux/bpf_verifier.h` which introduces `struct bpf_func_ptr` and adds to `struct bpf_verifier_env`: and the fields: `struct bpf_func_ptr *func_ptrs;`, `u32 func_ptr_cnt;`, `bool has_callx;`, and `unsigned long *callx_edges;`. These are initialized in their respective functions. * Core BPF infrastructure files, including the verifier, backtracking, configuration, constant folding, disassembly, fixups, and liveness analysis. The verifier has been updated to handle the `callx` instruction, track function pointers, and handle edges, but I'll need to confirm initialization and potential information leaks. Notably, the use of `kvrealloc` and `kvcalloc` are good because they handle the zero-initialization of memory. The `add_func_ptr` uses a compound literal (where missing values are set to zero). The `record_callx_edge` uses `kvcalloc`, which is good. `resolve_func_ptrs` reads 64-bit values from frozen, read-only maps. * Userspace libbpf, gen\_loader, linker, and selftest changes. My immediate concerns center on these points: 1. **`struct bpf_map` and its new `user` field:** I need to find all allocations of `struct bpf_map` using `grepper` (or similar) to confirm that the new `user` field is *always* initialized. Is it allocated with `kzalloc`, `bpf_map_alloc_percpu` or similar? If so, this is likely safe, but I need definitive confirmation. 2. **Verifier changes:** While I see the verifier has been updated, I need to double-check that the code additions, especially when handling `bpf_func_ptr` and `callx_edges`, never read uninitialized memory or leak data to user space. I need to ensure that the logic within the verifier correctly handles conditional logic and branching with zero-initialized values. The stack liveness analysis requires close scrutiny, particularly with the new `callee_stack_access_at_callsite` function and the way it handles uninitialized stack slots. 3. **JIT Compiler:** The JIT code itself is very low-level. However, it looks like local variables are initialized and no uninitialized variables or data are leaked to user space. 4. **No user space interaction (bpf programs memory, or stack)** KMSAN does *not* instrument JITed BPF code. I must remember that. The verifier ensures memory (registers, stack slots) is initialized *within* the BPF program. My conclusion at this point is that I *don't* see any immediate red flags. The code appears to be very careful about initialization. However, I need to confirm: 1. How and where `struct bpf_map` is allocated and the `user` field initialized. 2. Thoroughly review the verifier changes with a focus on potential data leaks or reads of uninitialized memory. Based on this review, I'm leaning toward |
| 26/3 | 2026/09/24 01:46 | tool | grepper |
0mArgs:{
"Expression": "bpf_map_claim"
}Results: {
"Output": "include/linux/bpf.h=302=struct bpf_map {\n--\ninclude/linux/bpf.h-345-\t/*\ninclude/linux/bpf.h:346:\t * Which programs use the map, see bpf_map_claim(): 0 - none so far,\ninclude/linux/bpf.h-347-\t * aux of the program - only that one, the same with BPF_MAP_USER_PATCHED\n--\nkernel/bpf/fixups.c=1102=static int jit_subprogs(struct bpf_verifier_env *env)\n--\nkernel/bpf/fixups.c-1295-\t * resolve_func_ptrs(). The program must be the only user of such map.\nkernel/bpf/fixups.c:1296:\t * From now on no other program can use it, see bpf_map_claim().\nkernel/bpf/fixups.c-1297-\t */\n--\nkernel/bpf/verifier.c=19599=static int check_map_prog_compatibility(struct bpf_verifier_env *env,\n--\nkernel/bpf/verifier.c-19714- */\nkernel/bpf/verifier.c:19715:static int bpf_map_claim(struct bpf_verifier_env *env, struct bpf_map *map)\nkernel/bpf/verifier.c-19716-{\n--\nkernel/bpf/verifier.c=19733=static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)\n--\nkernel/bpf/verifier.c-19769-\nkernel/bpf/verifier.c:19770:\terr = bpf_map_claim(env, map);\nkernel/bpf/verifier.c-19771-\tif (err)\n--\nkernel/bpf/verifier.c=20190=static int add_func_ptr(struct bpf_verifier_env *env, struct bpf_map *map, u32 map_off,\n--\nkernel/bpf/verifier.c-20228- * the functions in the map itself, see jit_subprogs(). Hence the program has to\nkernel/bpf/verifier.c:20229: * be the only user of the map, see bpf_map_claim(): nothing else may rely on\nkernel/bpf/verifier.c-20230- * what the map had.\n"
} |
| 27/3 | 2026/09/24 01:46 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 63101 output: 19 thoughts: 18 Thoughts: Here's your summary: **Digging into Map Creation and Allocation** Alright, let's get down to brass tacks. I need to understand the nuts and bolts of how maps are being constructed and how their memory is managed. Specifically, I need to zero in on the process. It's time to examine the `array_map_alloc` function. This should give me a solid grasp on the underlying mechanisms. Let's see what's happening under the hood. |
| 28/3 | 2026/09/24 01:47 | tool | grepper |
0mArgs:{
"Expression": "array_map_alloc"
}Results: {
"Output": "include/linux/bpf.h=3201=struct bpf_prog *bpf_prog_get_type_path(const char *name, enum bpf_prog_type type);\ninclude/linux/bpf.h:3202:int array_map_alloc_check(union bpf_attr *attr);\ninclude/linux/bpf.h-3203-\n--\nkernel/bpf/arraymap.c=33=static int bpf_array_alloc_percpu(struct bpf_array *array)\n--\nkernel/bpf/arraymap.c-52-/* Called from syscall */\nkernel/bpf/arraymap.c:53:int array_map_alloc_check(union bpf_attr *attr)\nkernel/bpf/arraymap.c-54-{\n--\nkernel/bpf/arraymap.c-83-\nkernel/bpf/arraymap.c:84:static struct bpf_map *array_map_alloc(union bpf_attr *attr)\nkernel/bpf/arraymap.c-85-{\n--\nkernel/bpf/arraymap.c=859=const struct bpf_map_ops array_map_ops = {\nkernel/bpf/arraymap.c-860-\t.map_meta_equal = array_map_meta_equal,\nkernel/bpf/arraymap.c:861:\t.map_alloc_check = array_map_alloc_check,\nkernel/bpf/arraymap.c:862:\t.map_alloc = array_map_alloc,\nkernel/bpf/arraymap.c-863-\t.map_free = array_map_free,\n--\nkernel/bpf/arraymap.c=886=const struct bpf_map_ops percpu_array_map_ops = {\nkernel/bpf/arraymap.c-887-\t.map_meta_equal = array_map_meta_equal,\nkernel/bpf/arraymap.c:888:\t.map_alloc_check = array_map_alloc_check,\nkernel/bpf/arraymap.c:889:\t.map_alloc = array_map_alloc,\nkernel/bpf/arraymap.c-890-\t.map_free = array_map_free,\n--\nkernel/bpf/arraymap.c-909-\nkernel/bpf/arraymap.c:910:static int fd_array_map_alloc_check(union bpf_attr *attr)\nkernel/bpf/arraymap.c-911-{\n--\nkernel/bpf/arraymap.c-917-\t\treturn -EINVAL;\nkernel/bpf/arraymap.c:918:\treturn array_map_alloc_check(attr);\nkernel/bpf/arraymap.c-919-}\n--\nkernel/bpf/arraymap.c=1225=static void prog_array_map_clear(struct bpf_map *map)\n--\nkernel/bpf/arraymap.c-1232-\nkernel/bpf/arraymap.c:1233:static struct bpf_map *prog_array_map_alloc(union bpf_attr *attr)\nkernel/bpf/arraymap.c-1234-{\n--\nkernel/bpf/arraymap.c-1245-\nkernel/bpf/arraymap.c:1246:\tmap = array_map_alloc(attr);\nkernel/bpf/arraymap.c-1247-\tif (IS_ERR(map)) {\n--\nkernel/bpf/arraymap.c=1277=const struct bpf_map_ops prog_array_map_ops = {\nkernel/bpf/arraymap.c:1278:\t.map_alloc_check = fd_array_map_alloc_check,\nkernel/bpf/arraymap.c:1279:\t.map_alloc = prog_array_map_alloc,\nkernel/bpf/arraymap.c-1280-\t.map_free = prog_array_map_free,\n--\nkernel/bpf/arraymap.c=1383=const struct bpf_map_ops perf_event_array_map_ops = {\nkernel/bpf/arraymap.c-1384-\t.map_meta_equal = bpf_map_meta_equal,\nkernel/bpf/arraymap.c:1385:\t.map_alloc_check = fd_array_map_alloc_check,\nkernel/bpf/arraymap.c:1386:\t.map_alloc = array_map_alloc,\nkernel/bpf/arraymap.c-1387-\t.map_free = perf_event_fd_array_map_free,\n--\nkernel/bpf/arraymap.c=1419=const struct bpf_map_ops cgroup_array_map_ops = {\nkernel/bpf/arraymap.c-1420-\t.map_meta_equal = bpf_map_meta_equal,\nkernel/bpf/arraymap.c:1421:\t.map_alloc_check = fd_array_map_alloc_check,\nkernel/bpf/arraymap.c:1422:\t.map_alloc = array_map_alloc,\nkernel/bpf/arraymap.c-1423-\t.map_free = cgroup_fd_array_free,\n--\nkernel/bpf/arraymap.c=1435=static struct bpf_map *array_of_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/arraymap.c-1442-\nkernel/bpf/arraymap.c:1443:\tmap = array_map_alloc(attr);\nkernel/bpf/arraymap.c-1444-\tif (IS_ERR(map)) {\n--\nkernel/bpf/arraymap.c=1505=const struct bpf_map_ops array_of_maps_map_ops = {\nkernel/bpf/arraymap.c:1506:\t.map_alloc_check = fd_array_map_alloc_check,\nkernel/bpf/arraymap.c-1507-\t.map_alloc = array_of_map_alloc,\n--\nkernel/bpf/reuseport_array.c=41=static int reuseport_array_alloc_check(union bpf_attr *attr)\n--\nkernel/bpf/reuseport_array.c-46-\nkernel/bpf/reuseport_array.c:47:\treturn array_map_alloc_check(attr);\nkernel/bpf/reuseport_array.c-48-}\n"
} |
| 29/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 64892 output: 36 thoughts: 29 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Diving into `array_map_alloc` for a closer look** Right, let's get down to brass tacks. I'm going to take a peek at the `array_map_alloc` function. Specifically, I'm heading straight into `kernel/bpf/arraymap.c` to see how it's implemented. Time to get the `read-file` tool ready. I need to refresh my memory and scrutinize its allocation strategy within the BPF context. I'm looking for anything that might leap out—potential performance bottlenecks, memory management quirks, or perhaps vulnerabilities if I'm feeling ambitious. |
| 30/3 | 2026/09/24 01:47 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/arraymap.c",
"FirstLine": 84,
"LineCount": 40
}Results: {
"Contents": " 84:\tstatic struct bpf_map *array_map_alloc(union bpf_attr *attr)\n 85:\t{\n 86:\t\tbool percpu = attr-\u003emap_type == BPF_MAP_TYPE_PERCPU_ARRAY;\n 87:\t\tint numa_node = bpf_map_attr_numa_node(attr);\n 88:\t\tu32 elem_size, index_mask, max_entries;\n 89:\t\tbool bypass_spec_v1 = bpf_bypass_spec_v1(NULL);\n 90:\t\tu64 array_size, mask64;\n 91:\t\tstruct bpf_array *array;\n 92:\t\n 93:\t\telem_size = round_up(attr-\u003evalue_size, 8);\n 94:\t\n 95:\t\tmax_entries = attr-\u003emax_entries;\n 96:\t\n 97:\t\t/* On 32 bit archs roundup_pow_of_two() with max_entries that has\n 98:\t\t * upper most bit set in u32 space is undefined behavior due to\n 99:\t\t * resulting 1U \u003c\u003c 32, so do it manually here in u64 space.\n 100:\t\t */\n 101:\t\tmask64 = fls_long(max_entries - 1);\n 102:\t\tmask64 = 1ULL \u003c\u003c mask64;\n 103:\t\tmask64 -= 1;\n 104:\t\n 105:\t\tindex_mask = mask64;\n 106:\t\tif (!bypass_spec_v1) {\n 107:\t\t\t/* round up array size to nearest power of 2,\n 108:\t\t\t * since cpu will speculate within index_mask limits\n 109:\t\t\t */\n 110:\t\t\tmax_entries = index_mask + 1;\n 111:\t\t\t/* Check for overflows. */\n 112:\t\t\tif (max_entries \u003c attr-\u003emax_entries)\n 113:\t\t\t\treturn ERR_PTR(-E2BIG);\n 114:\t\t}\n 115:\t\n 116:\t\tarray_size = sizeof(*array);\n 117:\t\tif (percpu) {\n 118:\t\t\tarray_size += (u64) max_entries * sizeof(void *);\n 119:\t\t} else {\n 120:\t\t\t/* rely on vmalloc() to return page-aligned memory and\n 121:\t\t\t * ensure array-\u003evalue is exactly page-aligned\n 122:\t\t\t */\n 123:\t\t\tif (attr-\u003emap_flags \u0026 BPF_F_MMAPABLE) {\n"
} |
| 31/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 65781 output: 37 thoughts: 14 Thoughts: Okay, here's my attempt at summarizing the thought process, speaking as if the described thoughts were my own, and assuming the reader is an expert: **Diving Deeper into `array_map_alloc`** Alright, let's take a closer look. The problem's probably lurking a bit further down in `array_map_alloc`. There's likely some nuance there that's eluding me at first glance. Need to dig in and see exactly what's going on with memory allocation or iteration within that function. My initial assumptions might be correct, but let's confirm them by getting into the implementation details. Time to get my hands dirty. |
| 32/3 | 2026/09/24 01:47 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/arraymap.c",
"FirstLine": 124,
"LineCount": 40
}Results: {
"Contents": " 124:\t\t\t\tarray_size = PAGE_ALIGN(array_size);\n 125:\t\t\t\tarray_size += PAGE_ALIGN((u64) max_entries * elem_size);\n 126:\t\t\t} else {\n 127:\t\t\t\tarray_size += (u64) max_entries * elem_size;\n 128:\t\t\t}\n 129:\t\t}\n 130:\t\n 131:\t\t/* allocate all map elements and zero-initialize them */\n 132:\t\tif (attr-\u003emap_flags \u0026 BPF_F_MMAPABLE) {\n 133:\t\t\tvoid *data;\n 134:\t\n 135:\t\t\t/* kmalloc'ed memory can't be mmap'ed, use explicit vmalloc */\n 136:\t\t\tdata = bpf_map_area_mmapable_alloc(array_size, numa_node);\n 137:\t\t\tif (!data)\n 138:\t\t\t\treturn ERR_PTR(-ENOMEM);\n 139:\t\t\tarray = data + PAGE_ALIGN(sizeof(struct bpf_array))\n 140:\t\t\t\t- offsetof(struct bpf_array, value);\n 141:\t\t} else {\n 142:\t\t\tarray = bpf_map_area_alloc(array_size, numa_node);\n 143:\t\t}\n 144:\t\tif (!array)\n 145:\t\t\treturn ERR_PTR(-ENOMEM);\n 146:\t\tarray-\u003eindex_mask = index_mask;\n 147:\t\tarray-\u003emap.bypass_spec_v1 = bypass_spec_v1;\n 148:\t\n 149:\t\t/* copy mandatory map attributes */\n 150:\t\tbpf_map_init_from_attr(\u0026array-\u003emap, attr);\n 151:\t\tarray-\u003eelem_size = elem_size;\n 152:\t\n 153:\t\tif (percpu \u0026\u0026 bpf_array_alloc_percpu(array)) {\n 154:\t\t\tbpf_map_area_free(array);\n 155:\t\t\treturn ERR_PTR(-ENOMEM);\n 156:\t\t}\n 157:\t\n 158:\t\treturn \u0026array-\u003emap;\n 159:\t}\n 160:\t\n 161:\tstatic void *array_map_elem_ptr(struct bpf_array* array, u32 index)\n 162:\t{\n 163:\t\treturn array-\u003evalue + (u64)array-\u003eelem_size * index;\n"
} |
| 33/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 66623 output: 22 thoughts: 37 Thoughts: Here's my summary, as though I were thinking through the process: **Diving into `bpf_map_area_alloc`** Okay, so I'm looking at `bpf_map_area_alloc`, and my immediate thought is that this function is responsible for allocating zeroed memory. From the name, I'd guess this is likely used for BPF map areas, which makes sense. I see this is either going to use `kzalloc` or `vzalloc`. Now, let's take a closer look at *how* exactly this allocation is happening and, more importantly, what considerations are being made for the different allocation methods. We've got work to do, but this is the right place to be. |
| 34/3 | 2026/09/24 01:47 | tool | grepper |
0mArgs:{
"Expression": "bpf_map_area_alloc"
}Results: {
"Output": "include/linux/bpf.h=2868=void bpf_map_put(struct bpf_map *map);\ninclude/linux/bpf.h:2869:void *bpf_map_area_alloc(u64 size, int numa_node);\ninclude/linux/bpf.h-2870-void *bpf_map_area_mmapable_alloc(u64 size, int numa_node);\n--\nkernel/bpf/arena.c=266=static struct bpf_map *arena_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/arena.c-299-\nkernel/bpf/arena.c:300:\tarena = bpf_map_area_alloc(sizeof(*arena), numa_node);\nkernel/bpf/arena.c-301-\tif (!arena)\n--\nkernel/bpf/arraymap.c=84=static struct bpf_map *array_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/arraymap.c-141-\t} else {\nkernel/bpf/arraymap.c:142:\t\tarray = bpf_map_area_alloc(array_size, numa_node);\nkernel/bpf/arraymap.c-143-\t}\n--\nkernel/bpf/bloom_filter.c=99=static struct bpf_map *bloom_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/bloom_filter.c-144-\tbitset_bytes = BITS_TO_LONGS((u64)bitset_mask + 1) * sizeof(unsigned long);\nkernel/bpf/bloom_filter.c:145:\tbloom = bpf_map_area_alloc(sizeof(*bloom) + bitset_bytes, numa_node);\nkernel/bpf/bloom_filter.c-146-\n--\nkernel/bpf/bpf_insn_array.c=44=static struct bpf_map *insn_array_alloc(union bpf_attr *attr)\n--\nkernel/bpf/bpf_insn_array.c-48-\nkernel/bpf/bpf_insn_array.c:49:\tinsn_array = bpf_map_area_alloc(size, NUMA_NO_NODE);\nkernel/bpf/bpf_insn_array.c-50-\tif (!insn_array)\n--\nkernel/bpf/bpf_local_storage.c=776=bpf_local_storage_map_alloc(union bpf_attr *attr,\n--\nkernel/bpf/bpf_local_storage.c-783-\nkernel/bpf/bpf_local_storage.c:784:\tsmap = bpf_map_area_alloc(sizeof(*smap), NUMA_NO_NODE);\nkernel/bpf/bpf_local_storage.c-785-\tif (!smap)\n--\nkernel/bpf/bpf_struct_ops.c=1078=static struct bpf_map *bpf_struct_ops_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/bpf_struct_ops.c-1138-\nkernel/bpf/bpf_struct_ops.c:1139:\tst_map = bpf_map_area_alloc(st_map_size, NUMA_NO_NODE);\nkernel/bpf/bpf_struct_ops.c-1140-\tif (!st_map) {\n--\nkernel/bpf/bpf_struct_ops.c-1147-\nkernel/bpf/bpf_struct_ops.c:1148:\tst_map-\u003euvalue = bpf_map_area_alloc(vt-\u003esize, NUMA_NO_NODE);\nkernel/bpf/bpf_struct_ops.c-1149-\tst_map-\u003efuncs_cnt = count_func_ptrs(btf, t);\nkernel/bpf/bpf_struct_ops.c-1150-\tst_map-\u003elinks =\nkernel/bpf/bpf_struct_ops.c:1151:\t\tbpf_map_area_alloc(st_map-\u003efuncs_cnt * sizeof(struct bpf_link *),\nkernel/bpf/bpf_struct_ops.c-1152-\t\t\t\t NUMA_NO_NODE);\n--\nkernel/bpf/bpf_struct_ops.c-1154-\tst_map-\u003eksyms =\nkernel/bpf/bpf_struct_ops.c:1155:\t\tbpf_map_area_alloc(st_map-\u003efuncs_cnt * sizeof(struct bpf_ksym *),\nkernel/bpf/bpf_struct_ops.c-1156-\t\t\t\t NUMA_NO_NODE);\n--\nkernel/bpf/cpumap.c=85=static struct bpf_map *cpu_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/cpumap.c-100-\nkernel/bpf/cpumap.c:101:\tcmap = bpf_map_area_alloc(sizeof(*cmap), NUMA_NO_NODE);\nkernel/bpf/cpumap.c-102-\tif (!cmap)\n--\nkernel/bpf/cpumap.c-107-\t/* Alloc array for possible remote \"destination\" CPUs */\nkernel/bpf/cpumap.c:108:\tcmap-\u003ecpu_map = bpf_map_area_alloc(cmap-\u003emap.max_entries *\nkernel/bpf/cpumap.c-109-\t\t\t\t\t sizeof(struct bpf_cpu_map_entry *),\n--\nkernel/bpf/devmap.c=91=static struct hlist_head *dev_map_create_hash(unsigned int entries,\n--\nkernel/bpf/devmap.c-96-\nkernel/bpf/devmap.c:97:\thash = bpf_map_area_alloc((u64) entries * sizeof(*hash), numa_node);\nkernel/bpf/devmap.c-98-\tif (hash != NULL)\n--\nkernel/bpf/devmap.c=136=static int dev_map_init_map(struct bpf_dtab *dtab, union bpf_attr *attr)\n--\nkernel/bpf/devmap.c-153-\t} else {\nkernel/bpf/devmap.c:154:\t\tdtab-\u003enetdev_map = bpf_map_area_alloc((u64) dtab-\u003emap.max_entries *\nkernel/bpf/devmap.c-155-\t\t\t\t\t\t sizeof(struct bpf_dtab_netdev *),\n--\nkernel/bpf/devmap.c=164=static struct bpf_map *dev_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/devmap.c-168-\nkernel/bpf/devmap.c:169:\tdtab = bpf_map_area_alloc(sizeof(*dtab), NUMA_NO_NODE);\nkernel/bpf/devmap.c-170-\tif (!dtab)\n--\nkernel/bpf/hashtab.c=319=static int prealloc_init(struct bpf_htab *htab)\n--\nkernel/bpf/hashtab.c-326-\nkernel/bpf/hashtab.c:327:\thtab-\u003eelems = bpf_map_area_alloc((u64)htab-\u003eelem_size * num_entries,\nkernel/bpf/hashtab.c-328-\t\t\t\t\t htab-\u003emap.numa_node);\n--\nkernel/bpf/hashtab.c=548=static struct bpf_map *htab_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/hashtab.c-561-\nkernel/bpf/hashtab.c:562:\thtab = bpf_map_area_alloc(sizeof(*htab), NUMA_NO_NODE);\nkernel/bpf/hashtab.c-563-\tif (!htab)\n--\nkernel/bpf/hashtab.c-604-\terr = -ENOMEM;\nkernel/bpf/hashtab.c:605:\thtab-\u003ebuckets = bpf_map_area_alloc(htab-\u003en_buckets *\nkernel/bpf/hashtab.c-606-\t\t\t\t\t sizeof(struct bucket),\n--\nkernel/bpf/hashtab.c=2798=static struct bpf_map *rhtab_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/hashtab.c-2803-\nkernel/bpf/hashtab.c:2804:\trhtab = bpf_map_area_alloc(sizeof(*rhtab), NUMA_NO_NODE);\nkernel/bpf/hashtab.c-2805-\tif (!rhtab)\n--\nkernel/bpf/local_storage.c=297=static struct bpf_map *cgroup_storage_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/local_storage.c-327-\nkernel/bpf/local_storage.c:328:\tmap = bpf_map_area_alloc(sizeof(struct bpf_cgroup_storage_map), numa_node);\nkernel/bpf/local_storage.c-329-\tif (!map)\n--\nkernel/bpf/lpm_trie.c=571=static struct bpf_map *trie_alloc(union bpf_attr *attr)\n--\nkernel/bpf/lpm_trie.c-587-\nkernel/bpf/lpm_trie.c:588:\ttrie = bpf_map_area_alloc(sizeof(*trie), NUMA_NO_NODE);\nkernel/bpf/lpm_trie.c-589-\tif (!trie)\n--\nkernel/bpf/offload.c=503=struct bpf_map *bpf_map_offload_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/offload.c-515-\nkernel/bpf/offload.c:516:\toffmap = bpf_map_area_alloc(sizeof(*offmap), NUMA_NO_NODE);\nkernel/bpf/offload.c-517-\tif (!offmap)\n--\nkernel/bpf/queue_stack_maps.c=68=static struct bpf_map *queue_stack_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/queue_stack_maps.c-76-\nkernel/bpf/queue_stack_maps.c:77:\tqs = bpf_map_area_alloc(queue_size, numa_node);\nkernel/bpf/queue_stack_maps.c-78-\tif (!qs)\n--\nkernel/bpf/reuseport_array.c=149=static struct bpf_map *reuseport_array_alloc(union bpf_attr *attr)\n--\nkernel/bpf/reuseport_array.c-154-\t/* allocate all map elements and zero-initialize them */\nkernel/bpf/reuseport_array.c:155:\tarray = bpf_map_area_alloc(struct_size(array, ptrs, attr-\u003emax_entries), numa_node);\nkernel/bpf/reuseport_array.c-156-\tif (!array)\n--\nkernel/bpf/ringbuf.c=93=static struct bpf_ringbuf *bpf_ringbuf_area_alloc(size_t data_sz, int numa_node)\n--\nkernel/bpf/ringbuf.c-122-\tarray_size = (nr_meta_pages + 2 * nr_data_pages) * sizeof(*pages);\nkernel/bpf/ringbuf.c:123:\tpages = bpf_map_area_alloc(array_size, numa_node);\nkernel/bpf/ringbuf.c-124-\tif (!pages)\n--\nkernel/bpf/ringbuf.c=194=static struct bpf_map *ringbuf_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/ringbuf.c-212-\nkernel/bpf/ringbuf.c:213:\trb_map = bpf_map_area_alloc(sizeof(*rb_map), NUMA_NO_NODE);\nkernel/bpf/ringbuf.c-214-\tif (!rb_map)\n--\nkernel/bpf/stackmap.c=68=static int prealloc_elems_and_freelist(struct bpf_stack_map *smap)\n--\nkernel/bpf/stackmap.c-73-\nkernel/bpf/stackmap.c:74:\tsmap-\u003eelems = bpf_map_area_alloc(elem_size * smap-\u003emap.max_entries,\nkernel/bpf/stackmap.c-75-\t\t\t\t\t smap-\u003emap.numa_node);\n--\nkernel/bpf/stackmap.c=93=static struct bpf_map *stack_map_alloc(union bpf_attr *attr)\n--\nkernel/bpf/stackmap.c-125-\tcost = n_buckets * sizeof(struct stack_map_bucket *) + sizeof(*smap);\nkernel/bpf/stackmap.c:126:\tsmap = bpf_map_area_alloc(cost, bpf_map_attr_numa_node(attr));\nkernel/bpf/stackmap.c-127-\tif (!smap)\n--\nkernel/bpf/syscall.c=308=static int bpf_map_copy_value(struct bpf_map *map, void *key, void *value,\n--\nkernel/bpf/syscall.c-371- */\nkernel/bpf/syscall.c:372:static void *__bpf_map_area_alloc(u64 size, int numa_node, bool mmapable)\nkernel/bpf/syscall.c-373-{\n--\nkernel/bpf/syscall.c-408-\nkernel/bpf/syscall.c:409:void *bpf_map_area_alloc(u64 size, int numa_node)\nkernel/bpf/syscall.c-410-{\nkernel/bpf/syscall.c:411:\treturn __bpf_map_area_alloc(size, numa_node, false);\nkernel/bpf/syscall.c-412-}\n--\nkernel/bpf/syscall.c=414=void *bpf_map_area_mmapable_alloc(u64 size, int numa_node)\nkernel/bpf/syscall.c-415-{\nkernel/bpf/syscall.c:416:\treturn __bpf_map_area_alloc(size, numa_node, true);\nkernel/bpf/syscall.c-417-}\n--\nnet/core/sock_map.c=39=static struct bpf_map *sock_map_alloc(union bpf_attr *attr)\n--\nnet/core/sock_map.c-49-\nnet/core/sock_map.c:50:\tstab = bpf_map_area_alloc(sizeof(*stab), NUMA_NO_NODE);\nnet/core/sock_map.c-51-\tif (!stab)\n--\nnet/core/sock_map.c-56-\nnet/core/sock_map.c:57:\tstab-\u003esks = bpf_map_area_alloc((u64) stab-\u003emap.max_entries *\nnet/core/sock_map.c-58-\t\t\t\t sizeof(struct sock *),\n--\nnet/core/sock_map.c=1098=static struct bpf_map *sock_hash_alloc(union bpf_attr *attr)\n--\nnet/core/sock_map.c-1111-\nnet/core/sock_map.c:1112:\thtab = bpf_map_area_alloc(sizeof(*htab), NUMA_NO_NODE);\nnet/core/sock_map.c-1113-\tif (!htab)\n--\nnet/core/sock_map.c-1126-\nnet/core/sock_map.c:1127:\thtab-\u003ebuckets = bpf_map_area_alloc(htab-\u003ebuckets_num *\nnet/core/sock_map.c-1128-\t\t\t\t\t sizeof(struct bpf_shtab_bucket),\n--\nnet/xdp/xskmap.c=64=static struct bpf_map *xsk_map_alloc(union bpf_attr *attr)\n--\nnet/xdp/xskmap.c-77-\nnet/xdp/xskmap.c:78:\tm = bpf_map_area_alloc(size, numa_node);\nnet/xdp/xskmap.c-79-\tif (!m)\n"
} |
| 35/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 71076 output: 36 thoughts: 37 Thoughts: Here's a breakdown of my thought process: **Diving into `__bpf_map_area_alloc`** Okay, so I need to understand what's happening with the allocation process in `__bpf_map_area_alloc`. The likely place to start would be in `kernel/bpf/syscall.c`, as this is where the system call interface resides. Specifically, lines 372 through 400 seem like a reasonable range to get an initial grasp on the allocation logic, focusing on how memory is being set aside for the BPF maps. Let's see what we can glean from that section. |
| 36/3 | 2026/09/24 01:47 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/syscall.c",
"FirstLine": 372,
"LineCount": 35
}Results: {
"Contents": " 372:\tstatic void *__bpf_map_area_alloc(u64 size, int numa_node, bool mmapable)\n 373:\t{\n 374:\t\t/* We really just want to fail instead of triggering OOM killer\n 375:\t\t * under memory pressure, therefore we set __GFP_NORETRY to kmalloc,\n 376:\t\t * which is used for lower order allocation requests.\n 377:\t\t *\n 378:\t\t * It has been observed that higher order allocation requests done by\n 379:\t\t * vmalloc with __GFP_NORETRY being set might fail due to not trying\n 380:\t\t * to reclaim memory from the page cache, thus we set\n 381:\t\t * __GFP_RETRY_MAYFAIL to avoid such situations.\n 382:\t\t */\n 383:\t\n 384:\t\tgfp_t gfp = bpf_memcg_flags(__GFP_NOWARN | __GFP_ZERO);\n 385:\t\tunsigned int flags = 0;\n 386:\t\tunsigned long align = 1;\n 387:\t\tvoid *area;\n 388:\t\n 389:\t\tif (size \u003e= SIZE_MAX)\n 390:\t\t\treturn NULL;\n 391:\t\n 392:\t\t/* kmalloc()'ed memory can't be mmap()'ed */\n 393:\t\tif (mmapable) {\n 394:\t\t\tBUG_ON(!PAGE_ALIGNED(size));\n 395:\t\t\talign = SHMLBA;\n 396:\t\t\tflags = VM_USERMAP;\n 397:\t\t} else if (size \u003c= (PAGE_SIZE \u003c\u003c PAGE_ALLOC_COSTLY_ORDER)) {\n 398:\t\t\tarea = kmalloc_node(size, gfp | GFP_USER | __GFP_NORETRY,\n 399:\t\t\t\t\t numa_node);\n 400:\t\t\tif (area != NULL)\n 401:\t\t\t\treturn area;\n 402:\t\t}\n 403:\t\n 404:\t\treturn __vmalloc_node_range(size, align, VMALLOC_START, VMALLOC_END,\n 405:\t\t\t\tgfp | GFP_KERNEL | __GFP_RETRY_MAYFAIL, PAGE_KERNEL,\n 406:\t\t\t\tflags, numa_node, __builtin_return_address(0));\n"
} |
| 37/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 71888 output: 21 thoughts: 648 Thoughts: **Analyzing Function Pointer Resolution in BPF Verification** Okay, so I'm looking at how function pointers are handled during BPF program verification, specifically within `resolve_func_ptrs`. I see that `__bpf_map_area_alloc` uses `__GFP_ZERO`, meaning the initial value of `map->user` is zero, which is important for understanding how maps are tracked. The core of the problem lies in `resolve_func_ptrs`. I need to understand what this code is *actually* doing. It iterates through the used maps in the environment (`env`). The first critical check: `if (READ_ONCE(map->user) != (unsigned long)env->prog->aux)`. This suggests it's trying to filter for maps that *belong* to the current program or are somehow associated. That's a good first step, avoiding interference between unrelated programs using the same maps. Now, a series of conditions narrow down the target maps. They must be `BPF_MAP_TYPE_ARRAY` with `max_entries` of 1, read-only, have `map_direct_value_addr` set, and not have `map->record`. This reeks of specialized usage. The most interesting part is what follows, `map->ops->map_direct_value_addr(map, &addr, 0)`. This call fetches a memory address, into `addr`, from the map. It looks like `addr` is a direct pointer to some data *within* the map. What is this `map_direct_value_addr`? I'm going to have to check what function is used for this `array_map_direct_value_addr`. The function pointer must be stored in the map's value. I need to be able to access it for verification. The subsequent loop then iterates through the map's value, assuming it's a series of 64-bit values. Each value is interpreted as a function pointer address. There's a set of checks: is `val` non-zero, aligned to `sizeof(struct bpf_insn)`, and within the program's instruction count? If so, the code tries to find the subprogram associated with that address using `bpf_find_subprog`, and then it uses `add_func_ptr` to add the function pointer to the environment. Finally, the code sorts the found function pointers. The key takeaway: the intent is to safely resolve function pointers embedded in BPF maps and make them usable, but only if they are properly "associated" with the running program in some way, and I need to carefully follow what `map_direct_value_addr` is doing because it seems to be the place where we retrieve the function pointer from the BPF map. |
| 38/3 | 2026/09/24 01:47 | tool | grepper |
0mArgs:{
"Expression": "map_direct_value_addr"
}Results: {
"Output": "Documentation/bpf/signing.rst=128=reuses the existing ``fd_array``:\n--\nDocumentation/bpf/signing.rst-141- ``excl_prog_sha``) and a plain array map (``BPF_MAP_TYPE_ARRAY``); only an\nDocumentation/bpf/signing.rst:142: array map exposes its value buffer through ``map_direct_value_addr()`` as a\nDocumentation/bpf/signing.rst-143- kernel address spanning ``value_size`` bytes. A map that is not exclusive, not\n--\nDocumentation/bpf/signing.rst=697=Notes and limitations\n--\nDocumentation/bpf/signing.rst-703-- The metadata container is a single-element array map, accessed through\nDocumentation/bpf/signing.rst:704: ``map_direct_value_addr``.\nDocumentation/bpf/signing.rst-705-- The bpf keyring needs ``CONFIG_KEYS``; without it there is no bpf keyring\n--\ninclude/linux/bpf.h=87=struct bpf_map_ops {\n--\ninclude/linux/bpf.h-143-\t/* Direct value access helpers. */\ninclude/linux/bpf.h:144:\tint (*map_direct_value_addr)(const struct bpf_map *map,\ninclude/linux/bpf.h-145-\t\t\t\t u64 *imm, u32 off);\n--\nkernel/bpf/arena.c=589=static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)\n--\nkernel/bpf/arena.c-628-\nkernel/bpf/arena.c:629:static int arena_map_direct_value_addr(const struct bpf_map *map, u64 *imm, u32 off)\nkernel/bpf/arena.c-630-{\n--\nkernel/bpf/arena.c=640=const struct bpf_map_ops arena_map_ops = {\n--\nkernel/bpf/arena.c-643-\t.map_free = arena_map_free,\nkernel/bpf/arena.c:644:\t.map_direct_value_addr = arena_map_direct_value_addr,\nkernel/bpf/arena.c-645-\t.map_mmap = arena_map_mmap,\n--\nkernel/bpf/arraymap.c=178=static int array_map_get_hash(struct bpf_map *map)\n--\nkernel/bpf/arraymap.c-186-\nkernel/bpf/arraymap.c:187:static int array_map_direct_value_addr(const struct bpf_map *map, u64 *imm,\nkernel/bpf/arraymap.c-188-\t\t\t\t u32 off)\n--\nkernel/bpf/arraymap.c=251=static void *percpu_array_map_lookup_elem(struct bpf_map *map, void *key)\n--\nkernel/bpf/arraymap.c-261-\nkernel/bpf/arraymap.c:262:static int percpu_array_map_direct_value_addr(const struct bpf_map *map, u64 *imm, u32 off)\nkernel/bpf/arraymap.c-263-{\n--\nkernel/bpf/arraymap.c=859=const struct bpf_map_ops array_map_ops = {\n--\nkernel/bpf/arraymap.c-869-\t.map_gen_lookup = array_map_gen_lookup,\nkernel/bpf/arraymap.c:870:\t.map_direct_value_addr = array_map_direct_value_addr,\nkernel/bpf/arraymap.c-871-\t.map_direct_value_meta = array_map_direct_value_meta,\n--\nkernel/bpf/arraymap.c=886=const struct bpf_map_ops percpu_array_map_ops = {\n--\nkernel/bpf/arraymap.c-893-\t.map_gen_lookup = percpu_array_map_gen_lookup,\nkernel/bpf/arraymap.c:894:\t.map_direct_value_addr = percpu_array_map_direct_value_addr,\nkernel/bpf/arraymap.c-895-\t.map_direct_value_meta = percpu_array_map_direct_value_meta,\n--\nkernel/bpf/bpf_insn_array.c=115=static u64 insn_array_mem_usage(const struct bpf_map *map)\n--\nkernel/bpf/bpf_insn_array.c-119-\nkernel/bpf/bpf_insn_array.c:120:static int insn_array_map_direct_value_addr(const struct bpf_map *map, u64 *imm, u32 off)\nkernel/bpf/bpf_insn_array.c-121-{\n--\nkernel/bpf/bpf_insn_array.c=136=const struct bpf_map_ops insn_array_map_ops = {\n--\nkernel/bpf/bpf_insn_array.c-145-\t.map_mem_usage = insn_array_mem_usage,\nkernel/bpf/bpf_insn_array.c:146:\t.map_direct_value_addr = insn_array_map_direct_value_addr,\nkernel/bpf/bpf_insn_array.c-147-\t.map_btf_id = \u0026insn_array_btf_ids[0],\n--\nkernel/bpf/const_fold.c=50=static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *ci_out,\n--\nkernel/bpf/const_fold.c-188-\t\t */\nkernel/bpf/const_fold.c:189:\t\tif (!bpf_map_is_rdonly(map) || !map-\u003eops-\u003emap_direct_value_addr ||\nkernel/bpf/const_fold.c-190-\t\t map-\u003emap_type == BPF_MAP_TYPE_INSN_ARRAY ||\n--\nkernel/bpf/fixups.c=1102=static int jit_subprogs(struct bpf_verifier_env *env)\n--\nkernel/bpf/fixups.c-1311-\t\t/* it's the address of the value of the map whatever the offset is */\nkernel/bpf/fixups.c:1312:\t\terr = ptr-\u003emap-\u003eops-\u003emap_direct_value_addr(ptr-\u003emap, \u0026addr, 0);\nkernel/bpf/fixups.c-1313-\t\tif (verifier_bug_if(err, env, \"no value of map '%s'\", ptr-\u003emap-\u003ename)) {\n--\nkernel/bpf/verifier.c=5990=int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,\n--\nkernel/bpf/verifier.c-5998-\t\treturn -EINVAL;\nkernel/bpf/verifier.c:5999:\terr = map-\u003eops-\u003emap_direct_value_addr(map, \u0026addr, off);\nkernel/bpf/verifier.c-6000-\tif (err)\n--\nkernel/bpf/verifier.c=6723=static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,\n--\nkernel/bpf/verifier.c-6755-\t bpf_map_is_rdonly(map) \u0026\u0026\nkernel/bpf/verifier.c:6756:\t map-\u003eops-\u003emap_direct_value_addr) {\nkernel/bpf/verifier.c-6757-\t\tint map_off = off + reg-\u003evar_off.value;\n--\nkernel/bpf/verifier.c=9058=static int check_arg_const_str(struct bpf_verifier_env *env,\n--\nkernel/bpf/verifier.c-9092-\nkernel/bpf/verifier.c:9093:\tif (!map-\u003eops-\u003emap_direct_value_addr) {\nkernel/bpf/verifier.c-9094-\t\tverbose(env, \"no direct value access support for this map type\\n\");\n--\nkernel/bpf/verifier.c-9104-\tmap_off = reg-\u003evar_off.value;\nkernel/bpf/verifier.c:9105:\terr = map-\u003eops-\u003emap_direct_value_addr(map, \u0026map_addr, map_off);\nkernel/bpf/verifier.c-9106-\tif (err) {\n--\nkernel/bpf/verifier.c=11745=static int check_bpf_snprintf_call(struct bpf_verifier_env *env,\n--\nkernel/bpf/verifier.c-11761-\t/* fmt being ARG_PTR_TO_CONST_STR guarantees that var_off is const\nkernel/bpf/verifier.c:11762:\t * and map_direct_value_addr is set.\nkernel/bpf/verifier.c-11763-\t */\nkernel/bpf/verifier.c-11764-\tfmt_map_off = fmt_reg-\u003evar_off.value;\nkernel/bpf/verifier.c:11765:\terr = fmt_map-\u003eops-\u003emap_direct_value_addr(fmt_map, \u0026fmt_addr,\nkernel/bpf/verifier.c-11766-\t\t\t\t\t\t fmt_map_off);\n--\nkernel/bpf/verifier.c=20047=static int check_and_resolve_insns(struct bpf_verifier_env *env)\n--\nkernel/bpf/verifier.c-20146-\nkernel/bpf/verifier.c:20147:\t\t\t\tif (!map-\u003eops-\u003emap_direct_value_addr) {\nkernel/bpf/verifier.c-20148-\t\t\t\t\tverbose(env, \"no direct value access support for this map type\\n\");\n--\nkernel/bpf/verifier.c-20151-\nkernel/bpf/verifier.c:20152:\t\t\t\terr = map-\u003eops-\u003emap_direct_value_addr(map, \u0026addr, off);\nkernel/bpf/verifier.c-20153-\t\t\t\tif (err) {\n--\nkernel/bpf/verifier.c=20235=static int resolve_func_ptrs(struct bpf_verifier_env *env)\n--\nkernel/bpf/verifier.c-20251-\t\tif (map-\u003emap_type != BPF_MAP_TYPE_ARRAY || map-\u003emax_entries != 1 ||\nkernel/bpf/verifier.c:20252:\t\t !bpf_map_is_rdonly(map) || !map-\u003eops-\u003emap_direct_value_addr ||\nkernel/bpf/verifier.c-20253-\t\t !IS_ERR_OR_NULL(map-\u003erecord))\nkernel/bpf/verifier.c-20254-\t\t\tcontinue;\nkernel/bpf/verifier.c:20255:\t\tif (map-\u003eops-\u003emap_direct_value_addr(map, \u0026addr, 0))\nkernel/bpf/verifier.c-20256-\t\t\tcontinue;\n--\nkernel/bpf/verifier.c=21988=static int bpf_prog_verify_signature(struct bpf_verifier_env *env,\n--\nkernel/bpf/verifier.c-22037-\t\tif (map-\u003emap_type != BPF_MAP_TYPE_ARRAY ||\nkernel/bpf/verifier.c:22038:\t\t !map-\u003eops-\u003emap_direct_value_addr) {\nkernel/bpf/verifier.c-22039-\t\t\tverbose(env, \"signed program metadata map '%s' must be an array\\n\",\n--\nkernel/bpf/verifier.c-22079-\nkernel/bpf/verifier.c:22080:\t\terr = map-\u003eops-\u003emap_direct_value_addr(map, \u0026addr, 0);\nkernel/bpf/verifier.c-22081-\t\tif (err) {\n--\ntools/testing/selftests/bpf/prog_tests/signed_loader.c=451=static void signed_nonarray_fd_array_rejected(void)\n--\ntools/testing/selftests/bpf/prog_tests/signed_loader.c-462-\t\t * (-EINVAL) rather than folded - this is the type gate that\ntools/testing/selftests/bpf/prog_tests/signed_loader.c:463:\t\t * keeps arena maps (map_direct_value_addr() returns a user\ntools/testing/selftests/bpf/prog_tests/signed_loader.c-464-\t\t * address) and insn-array maps (buffer smaller than value_size)\n"
} |
| 39/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 76432 output: 37 thoughts: 27 Thoughts: Here's a summary of my thought process, as if I were considering the code you mentioned: **Examining `array_map_direct_value_addr` in `kernel/bpf/arraymap.c`** Okay, let's dive into this. I'm focusing on `array_map_direct_value_addr` within `kernel/bpf/arraymap.c`. This function name immediately suggests it's involved in direct access to the values within a BPF array map. Given my experience, I'm expecting this function to be crucial for performance-sensitive BPF applications where we bypass the usual indirection. I know a direct value address is a key optimization in this context. My first thought is to understand the context of this function. How does it fit into the overall array map access scheme? I'll probably be looking for its role in the `bpf_map_ops` structure, and how it interacts with the map's internal data structures (the `array_map` itself, of course). I want to see how this function handles things like bounds checking and any potential alignment considerations. Does it handle any locking? Does it assume any specific memory layout? Is it dealing with userspace vs. kernel space addresses? I'm particularly interested in any potential vulnerabilities or performance bottlenecks here. My expertise tells me that direct memory access can be a double-edged sword: highly performant but also prone to issues if not implemented correctly. I'll be scrutinizing the code for any potential race conditions, incorrect pointer arithmetic, or integer overflows, particularly when dealing with index calculations or sizes. I'll be paying close attention to any error handling. What kind of error codes are returned, and under what conditions? Are there any potential for denial-of-service (DoS) scenarios if invalid inputs are provided? I'm anticipating that this function, if it's indeed meant for high-performance direct value access, will have to be incredibly robust. I also want to see the error cases in the return types. |
| 40/3 | 2026/09/24 01:47 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/arraymap.c",
"FirstLine": 186,
"LineCount": 25
}Results: {
"Contents": " 186:\t\n 187:\tstatic int array_map_direct_value_addr(const struct bpf_map *map, u64 *imm,\n 188:\t\t\t\t\t u32 off)\n 189:\t{\n 190:\t\tstruct bpf_array *array = container_of(map, struct bpf_array, map);\n 191:\t\n 192:\t\tif (map-\u003emax_entries != 1)\n 193:\t\t\treturn -ENOTSUPP;\n 194:\t\tif (off \u003e= map-\u003evalue_size)\n 195:\t\t\treturn -EINVAL;\n 196:\t\n 197:\t\t*imm = (unsigned long)array-\u003evalue;\n 198:\t\treturn 0;\n 199:\t}\n 200:\t\n 201:\tstatic int array_map_direct_value_meta(const struct bpf_map *map, u64 imm,\n 202:\t\t\t\t\t u32 *off)\n 203:\t{\n 204:\t\tstruct bpf_array *array = container_of(map, struct bpf_array, map);\n 205:\t\tu64 base = (unsigned long)array-\u003evalue;\n 206:\t\tu64 range = array-\u003eelem_size;\n 207:\t\n 208:\t\tif (map-\u003emax_entries != 1)\n 209:\t\t\treturn -ENOTSUPP;\n 210:\t\tif (imm \u003c base || imm \u003e= base + range)\n"
} |
| 41/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 77327 output: 25 thoughts: 5359 Thoughts: **KMSAN Analysis of BPF Function Pointer Patch** Alright, after meticulously reviewing the code changes, my analysis focuses on potential KMSAN violations. The core of this patch introduces mechanisms for handling function pointers within BPF maps. My primary concern is if the patch introduces any risk of uninitialized memory usage or information leaks to user space. Firstly, the crucial point is that `array->value` is zero-initialized and updated through `bpf_map_update_elem`. The `resolve_func_ptrs` only acts on frozen (read-only) maps. The `jit_subprogs()` function iterates through function pointers, checks and updates the map values using `cmpxchg64`. All variables like `addr`, `old`, `new` are initialized. Importantly, the `new` value here, which represents the function pointer itself, is the kernel address of the JITed function. The `check_func_ptr_read()` verifies reads from function pointer maps. The `targets` array stores subprogram indices. The array `targets` is initialized before use; all the variables in this function are properly initialized, including the `targets` array. `mark_reg_func_ptr` initializes the register state. `record_callx_edge` allocates memory via `kvcalloc`, zeroing all memory. Similarly, `next_callx_callee` is clean and bounds-checked. The `check_max_stack_depth_subprog` initializes all fields of the `dinfo` structure. Function pointers are searched on `env->func_ptrs` using binary search. The range is always set and initialized. `bpf_map_range_func_ptrs` and `insn_func_ptrs` both set their respective count to 0 right away. The `check_arg_const_str` checks the size from `bpf_map_range_func_ptrs`. The `add_func_ptr` function warrants a deeper look. While `kvrealloc` itself does not zero new memory, the code immediately initializes each element of `env->func_ptrs` when added. Crucially, the code never reads beyond the valid elements of the array, ensuring no uninitialized memory access. Every added element is fully initialized, and other methods loop strictly up to `env->func_ptr_cnt`. Regarding potential kernel-to-user information leaks, it is crucial to understand that while a kernel pointer is written to the BPF map value (in `jit_subprogs`), this is an intended feature. The BPF program, with `CAP_PERFMON` capability, reads this function pointer directly from the map. This is designed behavior, not an information leak violation. KMSAN detects the copying of uninitialized memory to user space. In this case, `(unsigned long)func[subprog]->bpf_func` (the function pointer) is properly initialized, and KMSAN's shadow for the 64-bit value is initialized. Rule 1 is not violated because the kernel structures are not sent to the userspace. Rule 2 is also not violated because the conditional logic and branching does not depend on uninitialized variables. Rule 3 is not violated because any data structures introduced or modified are either fully initialized or zeroed before use. Rule 4 is not violated because of bounds checks in the surrounding code. Consequently, there are no risks of uninitialized memory usage or information leaks, thus, the `NeedsKMSAN=false`. |
| 42/3 | 2026/09/24 01:47 | tool | grepper |
0mArgs:{
"Expression": "bpf_subprog_call_depth_info"
}Results: {
"Output": "kernel/bpf/verifier.c=5420=static int round_up_stack_depth(struct bpf_verifier_env *env, int stack_depth)\n--\nkernel/bpf/verifier.c-5431-/* temporary state used for call frame depth calculation */\nkernel/bpf/verifier.c:5432:struct bpf_subprog_call_depth_info {\nkernel/bpf/verifier.c-5433-\tint ret_insn; /* caller instruction where we return to. */\n--\nkernel/bpf/verifier.c=5445=static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,\nkernel/bpf/verifier.c:5446:\t\t\t\t\t struct bpf_subprog_call_depth_info *dinfo,\nkernel/bpf/verifier.c-5447-\t\t\t\t\t bool priv_stack_supported)\n--\nkernel/bpf/verifier.c=5690=static int check_max_stack_depth(struct bpf_verifier_env *env)\n--\nkernel/bpf/verifier.c-5692-\tenum priv_stack_mode priv_stack_mode = PRIV_STACK_UNKNOWN;\nkernel/bpf/verifier.c:5693:\tstruct bpf_subprog_call_depth_info *dinfo;\nkernel/bpf/verifier.c-5694-\tstruct bpf_subprog_info *si = env-\u003esubprog_info;\n"
} |
| 43/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 83832 output: 37 thoughts: 24 Thoughts: Here's my summary, as if I were thinking it: **Digging into the BPF Verifier's Stack Depth** Okay, time to get to the heart of this. I need to understand how the BPF verifier handles stack depth, and where the relevant checks are performed. That means I need to jump into the `kernel/bpf/verifier.c` file. Specifically, I'm going to focus my attention on the `check_max_stack_depth` function. That's the logical place to start looking for the core logic around stack depth limitations and how the verifier enforces them. This should give me a solid foundation for tracing the control flow and understanding the verifier's approach. Let's see what that function reveals. |
| 44/3 | 2026/09/24 01:47 | tool | read-file |
0mArgs:{
"File": "kernel/bpf/verifier.c",
"FirstLine": 5690,
"LineCount": 40
}Results: {
"Contents": "5690:\tstatic int check_max_stack_depth(struct bpf_verifier_env *env)\n5691:\t{\n5692:\t\tenum priv_stack_mode priv_stack_mode = PRIV_STACK_UNKNOWN;\n5693:\t\tstruct bpf_subprog_call_depth_info *dinfo;\n5694:\t\tstruct bpf_subprog_info *si = env-\u003esubprog_info;\n5695:\t\tbool priv_stack_supported;\n5696:\t\tint ret;\n5697:\t\n5698:\t\tdinfo = kvzalloc_objs(*dinfo, env-\u003esubprog_cnt, GFP_KERNEL_ACCOUNT);\n5699:\t\tif (!dinfo)\n5700:\t\t\treturn -ENOMEM;\n5701:\t\n5702:\t\tfor (int i = 0; i \u003c env-\u003esubprog_cnt; i++) {\n5703:\t\t\tif (si[i].has_tail_call) {\n5704:\t\t\t\tpriv_stack_mode = NO_PRIV_STACK;\n5705:\t\t\t\tbreak;\n5706:\t\t\t}\n5707:\t\t}\n5708:\t\n5709:\t\tif (priv_stack_mode == PRIV_STACK_UNKNOWN)\n5710:\t\t\tpriv_stack_mode = bpf_enable_priv_stack(env-\u003eprog);\n5711:\t\n5712:\t\t/* All async_cb subprogs use normal kernel stack. If a particular\n5713:\t\t * subprog appears in both main prog and async_cb subtree, that\n5714:\t\t * subprog will use normal kernel stack to avoid potential nesting.\n5715:\t\t * The reverse subprog traversal ensures when main prog subtree is\n5716:\t\t * checked, the subprogs appearing in async_cb subtrees are already\n5717:\t\t * marked as using normal kernel stack, so stack size checking can\n5718:\t\t * be done properly.\n5719:\t\t */\n5720:\t\tfor (int i = env-\u003esubprog_cnt - 1; i \u003e= 0; i--) {\n5721:\t\t\tif (!i || si[i].is_async_cb) {\n5722:\t\t\t\tpriv_stack_supported = !i \u0026\u0026 priv_stack_mode == PRIV_STACK_ADAPTIVE;\n5723:\t\t\t\tret = check_max_stack_depth_subprog(env, i, dinfo,\n5724:\t\t\t\t\t\tpriv_stack_supported);\n5725:\t\t\t\tif (ret \u003c 0) {\n5726:\t\t\t\t\tkvfree(dinfo);\n5727:\t\t\t\t\treturn ret;\n5728:\t\t\t\t}\n5729:\t\t\t}\n"
} |
| 45/3 | 2026/09/24 01:47 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 84745 output: 459 thoughts: 3304 Thoughts: **Analysis of the BPF Verifier Patch for Uninitialized Memory Risks** Okay, let's break this down. My primary concern here is ensuring this BPF verifier patch doesn't introduce or expose any uninitialized memory vulnerabilities in the kernel. My approach is systematic: scrutinize all memory allocations, initializations, and data flows introduced or modified by the patch, using my deep understanding of the kernel's memory management and BPF verifier architecture. First, the allocation of `dinfo` with `kvzalloc_objs` is good. It's completely zeroed, which is exactly what we want. Then, I checked the `backtrack_insn` function and confirmed that `subprog` is initialized to -1 before potentially being set by `bpf_find_subprog`. The conditional logic using the result of these function calls looks safe, and there are no instances where the code directly relies on uninitialized memory within these areas. I made sure to note how we cover `bpf_is_callx`, and ensured that all potential values used are covered. Next, I went through `cfg.c` and specifically `visit_func_ptrs_insn` and `check_func_callx`. All variables are properly initialized and used within these functions, which confirms their safety. The same is true for `check_static_func_call`. Now, the addition and modifications to `bpf_map_claim` and `__bpf_free_used_maps`, are safe; all the variables are initialized properly, and the map claim logic looks robust. The cleanup phase in `bpf_check` is important, and I confirmed that allocated structures like `env->callx_edges` and `env->func_ptrs` are freed appropriately. The allocation of `env` itself is done with `kzalloc`, meaning the `func_ptrs`, and `callx_edges` were all set to NULL/0 when created. Then, there's the critical question. Does this require KMSAN testing? I carefully considered the KMSAN and KASAN differences. KASAN is for out-of-bounds access, UAF, double-frees, etc. KMSAN is specifically for uninitialized memory reads and information leaks. Given the context of this patch, I asked four key questions: 1. Are we sending kernel structures to userspace without zeroing? No. 2. Do we have conditional logic with uninitialized variables? No. All new variables are allocated and initialized. 3. Are we allocating complex data structures and potentially reading uninitialized fields? No. Everything is zeroed or initialized. 4. Are we messing with buffer boundaries in a way that allows access to uninitialized data? No. Checks are added to prevent that. Given all of this, the types of potential bugs are logic errors in CFG traversal, OOB array access, JIT errors, and verifier correctness. These would be caught by standard debugging tools (KASAN, UBSAN, KCSAN, etc.). No specific KMSAN fuzzing is required. KASAN is sufficient for memory safety in this patch. |
| 46/3 | 2026/09/24 01:48 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch series implements support for indirect calls (`callx` instruction: `BPF_JMP | BPF_CALL | BPF_X`) in the BPF verifier, x86/arm64 JIT compilers, and user-space libbpf tooling.\n\nAnalysis of memory initialization and safety:\n1. Kernel structures and memory allocations:\n - `struct bpf_verifier_env` fields (`func_ptrs`, `func_ptr_cnt`, `has_callx`, `callx_edges`) are zero-initialized via `kzalloc`.\n - `callx_edges` is allocated using `kvcalloc()`, ensuring all bits are zeroed.\n - `func_ptrs` elements are allocated via `kvrealloc()` and each populated element up to `func_ptr_cnt` is initialized with a compound literal (`struct bpf_func_ptr`), zeroing any omitted fields (such as `.used`).\n - `bpf_subprog_call_depth_info` array (`dinfo`) is allocated using `kvzalloc_objs()`.\n - The new field `user` in `struct bpf_map` is initialized to 0 during map allocation via `__bpf_map_area_alloc()` (`__GFP_ZERO`).\n2. No kernel-to-user info-leak of uninitialized memory:\n - No new `copy_to_user`, `put_user`, socket option, netlink, or ioctl copy-out paths are introduced or modified.\n3. No conditional logic or branching depends on uninitialized state.\n4. The risks associated with this patch primarily involve verifier correctness, call graph cycle detection, and JIT assembly generation, all of which fall under functional verification, UBSAN, and KASAN (e.g., out-of-bounds or use-after-free).\n\nSince there are no uninitialized memory usage risks or uninitialized memory info-leaks, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch series implements support for indirect calls (`callx` instruction: `BPF_JMP | BPF_CALL | BPF_X`) in the BPF verifier, x86/arm64 JIT compilers, and user-space libbpf tooling.\n\nAnalysis of memory initialization and safety:\n1. Kernel structures and memory allocations:\n - `struct bpf_verifier_env` fields (`func_ptrs`, `func_ptr_cnt`, `has_callx`, `callx_edges`) are zero-initialized via `kzalloc`.\n - `callx_edges` is allocated using `kvcalloc()`, ensuring all bits are zeroed.\n - `func_ptrs` elements are allocated via `kvrealloc()` and each populated element up to `func_ptr_cnt` is initialized with a compound literal (`struct bpf_func_ptr`), zeroing any omitted fields (such as `.used`).\n - `bpf_subprog_call_depth_info` array (`dinfo`) is allocated using `kvzalloc_objs()`.\n - The new field `user` in `struct bpf_map` is initialized to 0 during map allocation via `__bpf_map_area_alloc()` (`__GFP_ZERO`).\n2. No kernel-to-user info-leak of uninitialized memory:\n - No new `copy_to_user`, `put_user`, socket option, netlink, or ioctl copy-out paths are introduced or modified.\n3. No conditional logic or branching depends on uninitialized state.\n4. The risks associated with this patch primarily involve verifier correctness, call graph cycle detection, and JIT assembly generation, all of which fall under functional verification, UBSAN, and KASAN (e.g., out-of-bounds or use-after-free).\n\nSince there are no uninitialized memory usage risks or uninitialized memory info-leaks, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|