A netadmin in a user+net namespace can create many IPv4 and IPv6 multicast routing tables with MRT_TABLE and MRT6_TABLE. Each unseen id allocates an mr_table via the shared mr_table_alloc(), links it into the per-net list, and leaves it until netns teardown. Those objects were not charged to memcg, so the host unreclaimable slab grows with the table count. Account mr_table allocations with GFP_KERNEL_ACCOUNT and mark the IPv4/IPv6 MFC caches SLAB_ACCOUNT. This matches the established handling of IP addresses, routes and alternate interface names. Unresolved MFC entries are still allocated from softIRQ with GFP_ATOMIC and are not charged. They expire after 10 seconds and are bounded by the socket receive queue; see commit 0079ad8e8dc3 ("ipmr: remove hard code cache_resolve_queue_len limit"). Fixes: f0ad0860d01e ("ipv4: ipmr: support multiple tables") Fixes: d1db275dd3f6 ("ipv6: ip6mr: support multiple tables") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Signed-off-by: Zihan Xi --- changes in v4: - Reword the opening paragraph for IPv4 and IPv6, since mr_table_alloc() is shared by both families, as suggested by Ido Schimmel. - Note that unresolved MFC entries stay GFP_ATOMIC / unaccounted; they expire after 10 seconds and are bounded by the socket receive queue (commit 0079ad8e8dc3). - Point Fixes: at f0ad0860d01e and d1db275dd3f6, the commits that let user space create these tables. - Drop the 4e16880cb422 archaeology paragraph from the commit message and cover letter. - v3 Link: https://lore.kernel.org/all/cover.1788765185.git.zihanx@nebusec.ai/ changes in v3: - Drop the table lifetime / unpublished-alloc / empty-reclaim approach from v2. - Charge mr_table allocations with GFP_KERNEL_ACCOUNT and mark the IPv4/IPv6 MFC caches SLAB_ACCOUNT, as suggested by Ido Schimmel. - Point Fixes: at 4e16880cb422, the first GFP_KERNEL IPv6 heap vif6_table. 6bd521433942 only wrapped that already-heap state; d1db275dd3f6 only expanded the table count from 1 to N. - Cover: unfixed evidence is Slab/SUnreclaim growth with RET:0, not a host OOM. The memcg OOM log is from the patched kernel under 64M memory.max using poc.static. - v2 Link: https://lore.kernel.org/all/cover.1788622674.git.zihanx@nebusec.ai/ changes in v2: - Drop the shared mr_table refcount / list_del_rcu path that broke the ipmr forwarding selftest. - Limit the v2 approach to net/ipv6/ip6mr.c. - v1 Link: https://lore.kernel.org/all/cover.1784795838.git.zihanx@nebusec.ai/ net/ipv4/ipmr.c | 3 ++- net/ipv4/ipmr_base.c | 2 +- net/ipv6/ip6mr.c | 2 +- 3 files changed, 4 insertions(+), 3 deletions(-) diff --git a/net/ipv4/ipmr.c b/net/ipv4/ipmr.c index e5f2b1c6150d2..b9c544d48c452 100644 --- a/net/ipv4/ipmr.c +++ b/net/ipv4/ipmr.c @@ -3376,7 +3376,8 @@ int __init ip_mr_init(void) { int err; - mrt_cachep = KMEM_CACHE(mfc_cache, SLAB_HWCACHE_ALIGN | SLAB_PANIC); + mrt_cachep = KMEM_CACHE(mfc_cache, + SLAB_HWCACHE_ALIGN | SLAB_PANIC | SLAB_ACCOUNT); err = register_pernet_subsys(&ipmr_net_ops); if (err) diff --git a/net/ipv4/ipmr_base.c b/net/ipv4/ipmr_base.c index 867b24beded11..a0ec6d19a237f 100644 --- a/net/ipv4/ipmr_base.c +++ b/net/ipv4/ipmr_base.c @@ -52,7 +52,7 @@ mr_table_alloc(struct net *net, u32 id, struct mr_table *mrt; int err; - mrt = kzalloc_obj(*mrt); + mrt = kzalloc_obj(*mrt, GFP_KERNEL_ACCOUNT); if (!mrt) return ERR_PTR(-ENOMEM); mrt->id = id; diff --git a/net/ipv6/ip6mr.c b/net/ipv6/ip6mr.c index 3f2ed9b77deb5..9d8116b5edb17 100644 --- a/net/ipv6/ip6mr.c +++ b/net/ipv6/ip6mr.c @@ -1427,7 +1427,7 @@ int __init ip6_mr_init(void) { int err; - mrt_cachep = KMEM_CACHE(mfc6_cache, SLAB_HWCACHE_ALIGN); + mrt_cachep = KMEM_CACHE(mfc6_cache, SLAB_HWCACHE_ALIGN | SLAB_ACCOUNT); if (!mrt_cachep) return -ENOMEM; base-commit: 641d03105cc0d2437e32fdeec164f91a4ccef6c4 -- 2.43.0