The x86-64 JIT encodes frame sizes as 32-bit immediates in its prologue, epilogue and tail call sequences, a tail call pops the frame of the program making it and lands in the target's prologue before the target allocates its own frame, and private stacks are allocated from each program's depth, so nothing in it depends on frames staying within 512 bytes. Report bpf_jit_supports_large_stack(), which raises the budget of JITed programs to MAX_BPF_STACK_JIT: 2 KiB combined over a call chain, or per frame on a private stack, with no separate limit on a single frame. Interpreted programs keep 512 bytes. The worst case kernel stack use of a chain of tail calls grows accordingly: the callers of a tail call may still leave at most 256 bytes on the stack, so 33 programs can accumulate 8 KiB of dead frames below the last one, which may now use 2 KiB instead of 512 bytes, for a little over 10 KiB in total on a 16 KiB kernel stack. The per-cpu memory behind a private stack grows in proportion to the frames a program asks for, up to 2 KiB plus guards per frame. The budget does not depend on the privileges of the loader: an unprivileged program cannot call other BPF functions, so a tail call from one leaves no frame behind and its worst case is a single 2 KiB frame. Programs nested through helpers or attach points, such as a tracing program entered from a helper of a networking program, are not accounted against each other before or after this change; each level of nesting may now add up to 1.5 KiB more. The one nesting a program can force on itself, bpf_clone_redirect() to its own device, keeps that program at 512 bytes. The verifier state of a frame grows with the stack it uses, up to four times as many stack slots as before; the allocation is on demand, so only programs using deep frames pay for them. Signed-off-by: Kumar Kartikeya Dwivedi --- Documentation/bpf/bpf_design_QA.rst | 10 +++++++--- arch/x86/net/bpf_jit_comp.c | 12 ++++++++++++ 2 files changed, 19 insertions(+), 3 deletions(-) diff --git a/Documentation/bpf/bpf_design_QA.rst b/Documentation/bpf/bpf_design_QA.rst index eb19c945f4d5..be5fc4ac00d6 100644 --- a/Documentation/bpf/bpf_design_QA.rst +++ b/Documentation/bpf/bpf_design_QA.rst @@ -221,9 +221,13 @@ newer kernels. BPF programs need to change accordingly when this happens. Q: How much stack space a BPF program uses? ------------------------------------------- -A: Currently all program types are limited to 512 bytes of stack -space, but the verifier computes the actual amount of stack used -and both interpreter and most JITed code consume necessary amount. +A: A program may use up to 2 KiB of stack, combined over its call +chain, when the JIT of the architecture reports support for large +stacks (currently x86-64); a single function may use all of it, and +every frame of a program running on a private stack gets the whole +amount. Elsewhere, and whenever the interpreter is used, the limit is +512 bytes. The verifier computes the actual amount of stack used and +both interpreter and most JITed code consume necessary amount. Q: Can BPF be offloaded to HW? ------------------------------ diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c index 9fbef7504e51..e2e531dd1e0b 100644 --- a/arch/x86/net/bpf_jit_comp.c +++ b/arch/x86/net/bpf_jit_comp.c @@ -4520,6 +4520,18 @@ bool bpf_jit_supports_subprog_tailcalls(void) return true; } +/* + * Frame sizes are 32-bit immediates in the prologue, epilogue and tail call + * sequences, a tail call pops the caller's frame and lands in the target's + * prologue before the target allocates its own, and private stacks are + * allocated from the program's own depth, so MAX_BPF_STACK_JIT frames need + * nothing special. + */ +bool bpf_jit_supports_large_stack(void) +{ + return true; +} + bool bpf_jit_supports_percpu_insn(void) { return true; -- 2.53.0