From: Ackerley Tng When mem_cgroup_charge_hugetlb() fails with -ENOMEM, alloc_hugetlb_folio() currently propagates this error. This results in the page fault handler returning VM_FAULT_OOM. Because HugeTLB allocations are high-order and use __GFP_RETRY_MAYFAIL, they bypass the OOM killer. Returning VM_FAULT_OOM to the #PF handler without triggering the OOM killer (or having it make progress) leads to an infinite loop of retrying the fault. Avoid this loop by returning -ENOSPC when charging fails, which maps to VM_FAULT_SIGBUS, terminating the process cleanly. Make mem_cgroup_charge_hugetlb() fault handling use a common error handling path, the same handling used for hugetlb_cgroup_uncharge_cgroup{,_rsvd}(), which also don't trigger the OOM killer and hence opt to terminate the process with a SIGBUS. Fixes: 991135774c0e0 ("memcg/hugetlb: introduce mem_cgroup_charge_hugetlb") Cc: stable@vger.kernel.org Signed-off-by: Ackerley Tng --- mm/hugetlb.c | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 7857728457952..01b57f6d3b804 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -2823,7 +2823,6 @@ void wait_for_freed_hugetlb_folios(void) * * Return: A pointer to the allocated folio, or an ERR_PTR on failure. * -ENOSPC if cgroup charging fails or no folio is available. - * -ENOMEM if mem cgroup charging fails. */ struct folio *hugetlb_alloc_folio(struct hstate *h, struct mempolicy_interpreted *mpoli, u8 alloc_flags) @@ -2896,7 +2895,18 @@ struct folio *hugetlb_alloc_folio(struct hstate *h, * were committed to the folio and freeing the folio * would have cleared those up. */ - return ERR_PTR(ret); + /* + * Return -ENOSPC when this function fails to allocate + * or charge a huge page. If a standard (PAGE_SIZE) + * page allocation fails, the OOM killer is given a + * chance to run, which may resolve the failure on + * retry. However, for HugeTLB allocations, the OOM + * killer is not triggered. Returning -ENOMEM (or + * anything resulting in VM_FAULT_OOM) would leak to + * the #PF handler, causing it to loop indefinitely + * retrying the fault. + */ + return ERR_PTR(-ENOSPC); } return folio; -- 2.55.0.970.g62bdec98f9-goog