From: Tao Cui cgv_node_less() compares cvtimes with a plain <, which breaks once cvtime wraps. A weight-1 cgroup in a hierarchy summing to 10000 advances cvtime at up to 10000x wall time, so 2^64 ns of cvtime is weeks of continuous saturation away -- unlikely but reachable on a long-running host. At the wrap instant the plain comparison puts the wrapped node behind everything else permanently. Compare with (s64)(a - b) < 0 instead, as CFS does for vruntime. A cyclic comparison is valid as an rbtree comparator only because cgrp_cap_budget() clamps every node to within max_budget behind cvtime_now, so any two nodes are far less than 2^63 apart and the cyclic order agrees with the true order. Compile-tested and smoke-tested in a VM: weight distribution and dispatch unaffected. Fixes: 7b742aa2c2c9 ("sched_ext: Add a cgroup scheduler which uses flattened hierarchy") Reported-by: Sashiko Link: https://lore.kernel.org/r/3f1ce004-e259-4e72-a5f7-14a5050053bd@linux.dev Signed-off-by: Tao Cui --- tools/sched_ext/scx_flatcg.bpf.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/tools/sched_ext/scx_flatcg.bpf.c b/tools/sched_ext/scx_flatcg.bpf.c index 454ebb820c5e..be03b409db5e 100644 --- a/tools/sched_ext/scx_flatcg.bpf.c +++ b/tools/sched_ext/scx_flatcg.bpf.c @@ -144,7 +144,8 @@ static bool cgv_node_less(struct bpf_rb_node *a, const struct bpf_rb_node *b) cgc_a = container_of(a, struct cgv_node, rb_node); cgc_b = container_of(b, struct cgv_node, rb_node); - return cgc_a->cvtime < cgc_b->cvtime; + /* wrap-safe: cap_budget keeps nodes within 2^63 of each other */ + return (s64)(cgc_a->cvtime - cgc_b->cvtime) < 0; } static struct fcg_cpu_ctx *find_cpu_ctx(void) -- 2.43.0