With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a user address with EFLAGS.AC clear as a kernel bug: it does not consult the exception table and oopses right away. That is correct for ordinary kernel code, where get_kernel_nofault() and the other nofault accessors never let a user address reach a faulting instruction. JITed BPF programs are different. A privileged program may dereference a pointer the verifier cannot prove valid, and the verifier marks such loads PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM load, so that a fault on an unmapped kernel address zeroes the destination register and the program continues. Since a user address would oops instead, the x86 JIT also emits an address range check in front of every PROBE_MEM load, which keeps user addresses, the guard page above TASK_SIZE_MAX and the vsyscall page away from the load. The check duplicates the fault handler's knowledge of the address space layout, got the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM runtime load check"), and costs nine instructions and 32 to 39 bytes of code per load. When the faulting instruction belongs to a BPF program, resolve its exception table entry instead of oopsing, exactly as is done for faults on kernel addresses. The is_bpf_text_address() lookup sits inside the unlikely() SMAP branch that currently ends in page_fault_oops(), so no path that does not oops today executes any additional code, and the oops itself is unchanged when no entry matches. Non-BPF code keeps the existing behaviour: a kernel-mode user access without STAC still oopses, extable entry or not. The next patch uses this to drop the range check from the JIT when SMAP is enabled. Nothing changes about which addresses a BPF program may read: a user address never becomes readable, since SMAP forbids the access, and a PROBE_MEM load of a kernel address is handled as before. Without SMAP the JIT keeps its range check and this path is never reached. Acked-by: Puranjay Mohan Signed-off-by: Kumar Kartikeya Dwivedi --- arch/x86/mm/fault.c | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c index aa88370ce739..2060e5f35d77 100644 --- a/arch/x86/mm/fault.c +++ b/arch/x86/mm/fault.c @@ -20,6 +20,7 @@ #include #include /* find_and_lock_vma() */ #include +#include /* is_bpf_text_address() */ #include /* boot_cpu_has, ... */ #include /* dotraplinkage, ... */ @@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs, if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) && !(error_code & X86_PF_USER) && !(regs->flags & X86_EFLAGS_AC))) { + /* + * JITed BPF programs dereference untrusted pointers with loads + * that carry an exception table entry (PROBE_MEM). SMAP makes + * sure such a load cannot read user memory, so resolve the + * fault through the exception table, as for an unmapped kernel + * address, instead of oopsing. + */ + if (is_bpf_text_address(regs->ip) && + fixup_exception(regs, X86_TRAP_PF, error_code, address)) + return; /* * No extable entry here. This was a kernel access to an * invalid pointer. get_kernel_nofault() will not get here. -- 2.53.0-Meta The x86 JIT guards every PROBE_MEM load with a range check that keeps user addresses, the guard page above TASK_SIZE_MAX and the vsyscall page away from the load, because a kernel-mode fault on those addresses would oops under SMAP instead of reaching the load's exception table entry. The check is nine instructions and 39 bytes (32 when the offset is zero) in front of a load of a few bytes. For "r7 = *(u64 *)(r0 + 2200)" on a 5-level paging kernel, with VSYSCALL_ADDR and TASK_SIZE_MAX + PAGE_SIZE - VSYSCALL_ADDR as the two constants: movq $-10485760, %r10 movq %rax, %r11 addq $2200, %r11 subq %r10, %r11 movabsq $72057594048413696, %r10 cmpq %r10, %r11 ja load xorl %edi, %edi jmp done load: movq 2200(%rax), %rdi done: The previous patch made do_user_addr_fault() resolve the exception table entries of BPF programs for faults on user addresses when SMAP is enabled. On such kernels, emit the bare load with its exception table entry, as the arm64, riscv, s390 and loongarch JITs already do. All BPF programs run with SMAP active once the CPU feature is enabled, and the feature cannot change after boot, so checking it at JIT time is sufficient. Kernels without SMAP, including those booted with nosmap, keep the range check. Measured with veristat over every object of the BPF selftests and over 466 production objects from Meta's fleet, on an x86-64 guest with SMAP, with and without this series on top of bpf-next: programs with PROBE_MEM reduction per program mean median max BPF selftests 3324 91 35.0% 37.5% 74.4% Meta production programs 1710 228 7.0% 1.8% 69.2% No program grows. The socket and task iterators of the selftests lose about half of their code, dump_tcp6 goes from 4386 to 2124 bytes, and the smallest production programs lose 60% to 69%. Loads through trusted pointers and the probe_read helpers do not use PROBE_MEM, which is why most programs are unaffected. Signed-off-by: Kumar Kartikeya Dwivedi --- arch/x86/net/bpf_jit_comp.c | 21 ++++++++++++++++----- 1 file changed, 16 insertions(+), 5 deletions(-) diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c index 083fcd6cf15b..793e7cd5a5c4 100644 --- a/arch/x86/net/bpf_jit_comp.c +++ b/arch/x86/net/bpf_jit_comp.c @@ -2133,6 +2133,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * const s32 imm32 = insn->imm; u32 dst_reg = insn->dst_reg; u32 src_reg = insn->src_reg; + bool probe_mem, bounds_check; bool accesses_stack_only; u8 b2 = 0, b3 = 0; u8 *start_of_ldx; @@ -2709,6 +2710,15 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * case BPF_LDX | BPF_PROBE_MEMSX | BPF_B: case BPF_LDX | BPF_PROBE_MEMSX | BPF_H: case BPF_LDX | BPF_PROBE_MEMSX | BPF_W: + probe_mem = BPF_MODE(insn->code) == BPF_PROBE_MEM || + BPF_MODE(insn->code) == BPF_PROBE_MEMSX; + /* + * With SMAP enabled, a load from a user address faults and + * do_user_addr_fault() resolves the exception table entry of the + * program, as for an unmapped kernel address, so the address range + * check is only needed without SMAP. + */ + bounds_check = probe_mem && !cpu_feature_enabled(X86_FEATURE_SMAP); insn_off = insn->off; if (src_reg == BPF_REG_PARAMS) { if (insn_off == 8) { @@ -2724,8 +2734,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * */ } - if (BPF_MODE(insn->code) == BPF_PROBE_MEM || - BPF_MODE(insn->code) == BPF_PROBE_MEMSX) { + if (bounds_check) { /* Conservatively check that src_reg + insn->off is a kernel address: * src_reg + insn->off > TASK_SIZE_MAX + PAGE_SIZE * and @@ -2772,6 +2781,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * /* populate jmp_offset for JAE above to jump to start_of_ldx */ start_of_ldx = prog; end_of_jmp[-1] = start_of_ldx - end_of_jmp; + } else if (probe_mem) { + start_of_ldx = prog; } else if (!accesses_stack_only) { err = emit_kasan_check(env, &prog, src_reg, insn_off, @@ -2785,14 +2796,14 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * emit_ldsx(&prog, BPF_SIZE(insn->code), dst_reg, src_reg, insn_off); else emit_ldx(&prog, BPF_SIZE(insn->code), dst_reg, src_reg, insn_off); - if (BPF_MODE(insn->code) == BPF_PROBE_MEM || - BPF_MODE(insn->code) == BPF_PROBE_MEMSX) { + if (probe_mem) { struct exception_table_entry *ex; u8 *_insn = image + proglen + (start_of_ldx - temp); s64 delta; /* populate jmp_offset for JMP above */ - start_of_ldx[-1] = prog - start_of_ldx; + if (bounds_check) + start_of_ldx[-1] = prog - start_of_ldx; if (!bpf_prog->aux->extable) break; -- 2.53.0-Meta Add a test that performs PROBE_MEM loads of three sizes through a bpf_core_cast() pointer whose value is chosen by userspace: NULL, a low user address, the last user page, a non-canonical address and, on x86-64, the vsyscall page and an offset into it. Each load must read zero and the kernel must survive. A load through the current task pointer checks that the same loads read real values when the address is valid. The test passes on kernels that reject the addresses with the JIT's range check and on kernels that rely on the fault handler; with only the JIT change of the previous patch applied, the first subtest oopses, which is what the fault handler change prevents. Signed-off-by: Kumar Kartikeya Dwivedi --- .../bpf/prog_tests/probe_mem_fault.c | 69 +++++++++++++++++++ .../selftests/bpf/progs/probe_mem_fault.c | 41 +++++++++++ 2 files changed, 110 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c create mode 100644 tools/testing/selftests/bpf/progs/probe_mem_fault.c diff --git a/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c b/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c new file mode 100644 index 000000000000..57a313e35eb6 --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c @@ -0,0 +1,69 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */ +#include +#include "probe_mem_fault.skel.h" + +#if defined(__x86_64__) +#include +#endif + +/* + * Addresses a PROBE_MEM load has to survive. Either the JIT's address check + * or the fault handler must turn each load into a zero result. + */ +static const struct { + const char *name; + unsigned long addr; +} bad_addrs[] = { + { "null", 0 }, + { "low_user", 4096 }, + { "last_user_page", (1UL << 47) - 4096 }, + { "non_canonical", 1UL << 63 }, +#if defined(__x86_64__) + { "vsyscall", VSYSCALL_ADDR }, + { "vsyscall_tail", VSYSCALL_ADDR + 0x800 }, +#endif +}; + +static void trigger(struct probe_mem_fault *skel, int *runs) +{ + skel->bss->val_dw = ~0ULL; + skel->bss->val_w = ~0U; + skel->bss->val_b = ~0; + usleep(1); + ASSERT_EQ(skel->bss->runs, ++*runs, "runs"); +} + +void test_probe_mem_fault(void) +{ + struct probe_mem_fault *skel; + int runs = 0, i; + + skel = probe_mem_fault__open_and_load(); + if (!ASSERT_OK_PTR(skel, "open_and_load")) + return; + + skel->bss->target_pid = getpid(); + if (!ASSERT_OK(probe_mem_fault__attach(skel), "attach")) + goto out; + + /* A valid kernel address is read for real. */ + skel->bss->use_current_task = true; + trigger(skel, &runs); + ASSERT_EQ(skel->bss->val_w, getpid(), "pid"); + ASSERT_NEQ(skel->bss->val_dw, 0, "start_time"); + ASSERT_NEQ(skel->bss->val_b, 0, "comm"); + + skel->bss->use_current_task = false; + for (i = 0; i < ARRAY_SIZE(bad_addrs); i++) { + if (!test__start_subtest(bad_addrs[i].name)) + continue; + skel->bss->addr = bad_addrs[i].addr; + trigger(skel, &runs); + ASSERT_EQ(skel->bss->val_dw, 0, "start_time"); + ASSERT_EQ(skel->bss->val_w, 0, "pid"); + ASSERT_EQ(skel->bss->val_b, 0, "comm"); + } +out: + probe_mem_fault__destroy(skel); +} diff --git a/tools/testing/selftests/bpf/progs/probe_mem_fault.c b/tools/testing/selftests/bpf/progs/probe_mem_fault.c new file mode 100644 index 000000000000..748be45418f5 --- /dev/null +++ b/tools/testing/selftests/bpf/progs/probe_mem_fault.c @@ -0,0 +1,41 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */ +#include +#include +#include +#include +#include "bpf_misc.h" + +char _license[] SEC("license") = "GPL"; + +int target_pid; +bool use_current_task; +unsigned long addr; +int runs; +__u64 val_dw; +__u32 val_w; +__u8 val_b; + +SEC("fentry/" SYS_PREFIX "sys_nanosleep") +int probe_mem_fault(void *ctx) +{ + struct task_struct *task; + unsigned long p = addr; + + if ((bpf_get_current_pid_tgid() >> 32) != target_pid) + return 0; + + if (use_current_task) + p = (unsigned long)bpf_get_current_task_btf(); + /* + * bpf_core_cast() yields an untrusted pointer, so every load through it + * is a PROBE_MEM load. Whatever the address is, the load must either + * read the field or produce zero; the kernel must not oops. + */ + task = bpf_core_cast((void *)p, struct task_struct); + val_dw = task->start_time; + val_w = task->pid; + val_b = task->comm[0]; + runs++; + return 0; +} -- 2.53.0-Meta