A watermark boost is not confined to one watermark: wmark_pages() adds it to min, low and high alike, so every watermark check sees it, including should_reclaim_retry() and the last ditch ALLOC_WMARK_HIGH attempt in __alloc_pages_may_oom(). Once the boost exceeds the memory still free the allocator gives up and invokes the OOM killer, and a capture kernel that is still booting has nothing to kill: the boot panics and the vmcore is lost. Seen on an arm64 machine with 64K pages, CONFIG_PAGE_BLOCK_MAX_ORDER=10 (pageblock = 64M) and crashkernel=512M, running a distribution kernel based on 7.0.14. A high order UNMOVABLE allocation fell back to a MOVABLE pageblock while the capture kernel was still in do_initcalls(): Node 0 DMA free:68096kB boost:65536kB min:68160kB low:68800kB high:69440kB managed:479168kB Out of memory and no killable processes... Kernel panic - not syncing: System is deadlocked on memory The zone was not short of memory. Subtracting the boost gives min:2624kB low:3264kB high:3904kB, so the 68096kB still free sat 17 times above the high watermark and the allocator would not even have entered its slow path. The boost supplied 65536kB of the 68160kB min and by itself put the zone 64kB under water. It is that large because boost_watermark() clamps it with max(pageblock_nr_pages, max_boost); watermark_boost_factor alone would have allowed 5824kB. Commit 14f69140ff9c ("mm: limit boost_watermark on small zones") already tried to protect capture kernels, but it infers them from the zone size and skips the boost only below four pageblocks. arm64 64K pageblocks were 512M then, so the guard reached zones up to 2G; CONFIG_PAGE_BLOCK_MAX_ORDER can cap them at 64M, which shrinks the guard to zones under 256M and lets this 468M zone through. kdump is a property of the kernel, not of the zone, so test for it directly. A capture kernel exits within seconds and never uses the fragmentation avoidance the boost buys. Normal kernels are unaffected: the size based check still covers their genuinely tiny zones. Passing sysctl.vm.watermark_boost_factor=0 to the capture kernel does not cover this window: sysctl.* parameters are written through procfs by do_sysctl_args(), which runs after do_initcalls() where the panic above happened, and watermark_boost_factor has no early_param of its own. Fixes: 1c30844d2dfe ("mm: reclaim small amounts of memory when an external fragmentation event occurs") Cc: stable@vger.kernel.org # v5.0 Cc: Henry Willard Cc: David Hildenbrand Signed-off-by: Yuanhe Shu --- Build tested with CONFIG_CRASH_DUMP=y and =n. Tested on the affected machine with the original crashkernel=512M and capture kernel command line. Without this patch both capture boots panicked at 7.9s inside do_initcalls() with a 64M boost in place. With it, and with boost_watermark() instrumented, every call was suppressed - including those inside do_initcalls(), where the panics happened - free memory came down to within 64kB of the min watermark, where a single boost would have put it 65472kB under, and the dump completed. --- mm/page_alloc.c | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/mm/page_alloc.c b/mm/page_alloc.c index 12fac9084c48..276fa7169b99 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -37,6 +37,7 @@ #include #include #include +#include #include #include #include @@ -2178,6 +2179,14 @@ static inline bool boost_watermark(struct zone *zone) if (!watermark_boost_factor) return false; + + /* + * A kdump capture kernel exits before a boost can pay off, while + * the raised watermark can exceed the memory left for the dump. + */ + if (is_kdump_kernel()) + return false; + /* * Don't bother in zones that are unlikely to produce results. * On small machines, including kdump capture kernels running base-commit: 704340f1cd0dcef829eb62f5b48ae95a2ce17bdf -- 2.43.7