Deferring every exit would churn small processes for no benefit and let unreaped memory grow without bound. Split the defer decision into a pure predicate and a budget reservation: - async_mm_teardown_eligible() (side-effect free): defer only when RSS >= async_mm_teardown_thresh_pages and both MMF_OOM_SKIP and MMF_OOM_TARGETED are clear: * MMF_OOM_SKIP -- the mm is hidden from the OOM killer and reaper. Usually that means it has already been reaped or run through exit_mmap() and has no RSS left worth deferring; in the mm-pinned-by-init case (__oom_kill_process()) the flag is set at kill time with RSS intact, and keeping such an mm synchronous is equally what we want; * MMF_OOM_TARGETED -- the mm is the target of an in-flight OOM kill (added in the preceding patch). Such a victim is torn down synchronously so its memory is released promptly rather than queued behind the low-priority, housekeeping-confined mm_reaper, which would work against the oom_reaper during the exact pressure that triggered the kill. The flag is set while some task still holds a live reference to the mm, and this check is reached only after mm_users has dropped to 0 -- which requires that task's own exit_mm() (or exec_mmap(), the other path that sheds ->mm) to have already cleared ->mm under that same task_lock() -- so marking and checking can never overlap in time. Whichever thread performs the final decrement is therefore guaranteed to observe the flag already set, per the preceding patch. - async_mm_teardown_reserve(): charge RSS against a pending-pages budget with atomic_long_add_return(); if the result exceeds async_mm_teardown_max_pending_pages it subtracts back and returns false, so the caller falls back to synchronous __mmput(). mmput_exit() reads RSS once and checks eligibility before reservation, so the budget is only touched for processes that pass the cheap checks. On success it stores the charged amount in mm->async_reap_node's companion field async_reap_rss before enqueueing, and the reaper releases exactly that stored charge after __mmput(). The charge cannot be re-derived at reap time: RSS keeps shrinking after mm_users reaches zero -- rmap-based reclaim can still swap out anon pages and truncation can still unmap file pages, neither of which requires mm_users -- and get_mm_rss() is an approximate per-CPU counter read besides. Two independent reads at enqueue and reap would systematically under-release, drifting the pending counter upward until the cap permanently forces synchronous fallback. An eligible mm would otherwise have its entire __mmput() -- including exit_aio() -- deferred onto the single mm_reaper kthread. exit_aio() can block indefinitely waiting for in-flight AIO to drain, so mmput_exit() calls it eagerly, on the exiting task's own context, before reservation: a stuck AIO context then only stalls this task, same as a synchronous mmput() would, instead of blocking mm_reaper and stranding every other mm already queued behind it. exit_aio() is idempotent -- it is a no-op once mm->ioctx_table has been torn down -- so __mmput() calling it again unconditionally is harmless, whether that second call lands on the reaper or -- when the reservation fails -- back on the exiting task itself. Because reservation is add-then-check rather than read-then-add, committed pending never stays above the cap: an exit that would exceed it tears down synchronously instead. Under concurrent exits a failing reservation can transiently inflate the counter between its add and roll-back, which may cause another exiter to fall back to sync conservatively -- it can over-reject but never over-admit. The cap is a coarse safety bound on unreaped memory, not a reclaim mechanism: pages queued here are invisible to reclaim until the reaper frees them. Integrating the queue with reclaim (unmap-first draining, a shrinker) is future work; see the cover letter. The RSS threshold defaults to a fixed 64MB at init; the pending cap defaults to totalram_pages() / 4 as a placeholder pending tuning. Both are static module variables here and not yet user-tunable. The static key that gates the whole path, and sysctl registration, follow in the next patch. This patch is dormant: nothing calls mmput_exit() yet (exit_mm() is routed to it later in the series), so no teardown is deferred. Signed-off-by: Aditya Sharma --- include/linux/mm_types.h | 1 + kernel/fork.c | 79 +++++++++++++++++++++++++++++++++++++++- 2 files changed, 79 insertions(+), 1 deletion(-) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 3240c029f..ee6c0ac84 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -1370,6 +1370,7 @@ struct mm_struct { struct work_struct async_put_work; #ifdef CONFIG_ASYNC_MM_TEARDOWN struct llist_node async_reap_node; + unsigned long async_reap_rss; #endif /* CONFIG_ASYNC_MM_TEARDOWN */ #ifdef CONFIG_IOMMU_MM_DATA diff --git a/kernel/fork.c b/kernel/fork.c index 884e4e767..d45418e66 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -113,6 +113,7 @@ #include #include #include +#include #include #include @@ -3412,6 +3413,48 @@ subsys_initcall(init_fork_sysctl); #ifdef CONFIG_ASYNC_MM_TEARDOWN static LLIST_HEAD(mm_reaper_list); static DECLARE_WAIT_QUEUE_HEAD(mm_reaper_wait); +static atomic_long_t mm_reaper_pending_pages; +static unsigned long sysctl_async_mm_teardown_thresh_pages; +static unsigned long sysctl_async_mm_teardown_max_pending_pages; + +static bool async_mm_teardown_eligible(struct mm_struct *mm, unsigned long rss) +{ + if (rss < READ_ONCE(sysctl_async_mm_teardown_thresh_pages)) + return false; + if (mm_flags_test(MMF_OOM_SKIP, mm)) /* reaped, or hidden from the reaper */ + return false; + /* + * MMF_OOM_TARGETED is set by the OOM killer while some task still + * holds a live reference to this mm (see mark_oom_victim() and + * __oom_kill_process()), and this eligibility check is reached only + * from mmput_exit(), after mm_users has dropped to 0 -- so marking + * and this check can never overlap in time. Whichever thread + * performs the final decrement is therefore guaranteed to observe + * the flag already set, regardless of whether that thread was + * itself ever passed to mark_oom_victim() (a CLONE_VM-sharing + * process in another thread group never is, but still observes + * this flag on the shared mm). + */ + if (mm_flags_test(MMF_OOM_TARGETED, mm)) + return false; + return true; +} + +/* + * Charge rss against the pending-teardown budget. Returns true if it fits + * under the cap, in which case the caller enqueues and the reaper releases + * the same rss after __mmput(). On overshoot nothing stays charged and the + * caller must tear down synchronously. + */ +static bool async_mm_teardown_reserve(unsigned long rss) +{ + if (atomic_long_add_return(rss, &mm_reaper_pending_pages) > + READ_ONCE(sysctl_async_mm_teardown_max_pending_pages)) { + atomic_long_sub(rss, &mm_reaper_pending_pages); + return false; + } + return true; +} static void async_mm_teardown_queue(struct mm_struct *mm) { @@ -3429,7 +3472,10 @@ static int mm_reaper(void *unused) wait_event_freezable(mm_reaper_wait, !llist_empty(&mm_reaper_list)); batch = llist_del_all(&mm_reaper_list); llist_for_each_entry_safe(mm, n, batch, async_reap_node) { + unsigned long pages = mm->async_reap_rss; + __mmput(mm); /* may free mm via mmdrop */ + atomic_long_sub(pages, &mm_reaper_pending_pages); cond_resched(); try_to_freeze(); } @@ -3439,10 +3485,38 @@ static int mm_reaper(void *unused) void mmput_exit(struct mm_struct *mm) { + unsigned long rss; + might_sleep(); + /* + * Fully ordered RMW: combined with exit_mm() (or exec_mmap(), the + * other path that sheds ->mm) clearing ->mm under task_lock() before + * this call, it guarantees MMF_OOM_TARGETED set at any mark site is + * visible here -- see mark_oom_victim(). Any new caller of + * mmput_exit() outside exit_mm() must re-verify that argument. + */ if (!atomic_dec_and_test(&mm->mm_users)) return; - async_mm_teardown_queue(mm); + + rss = get_mm_rss(mm); + + if (async_mm_teardown_eligible(mm, rss)) { + /* + * exit_aio() can block indefinitely. Run it here so a stuck + * AIO only hangs this task (same as how it would happen in + * case of synchronous mmput()) instead of stranding every + * teardown queued behind this mm + */ + exit_aio(mm); + + if (async_mm_teardown_reserve(rss)) { + mm->async_reap_rss = rss; + async_mm_teardown_queue(mm); + return; + } + } + + __mmput(mm); } static int __init mm_reaper_init(void) @@ -3456,7 +3530,10 @@ static int __init mm_reaper_init(void) } set_user_nice(th, 19); /* minimize competition with other fair class tasks */ kthread_affine_preferred(th, housekeeping_cpumask(HK_TYPE_KTHREAD)); + sysctl_async_mm_teardown_thresh_pages = SZ_64M >> PAGE_SHIFT; + sysctl_async_mm_teardown_max_pending_pages = totalram_pages() / 4; /* TODO: placeholder */ wake_up_process(th); + return 0; } subsys_initcall(mm_reaper_init); -- 2.34.1