AI Trajectory


Trajectory: (Expand/Collapse All)
Seq Timestamp Type Name Duration
0/0 2026/09/14 12:10 flow patch-triage
34m Args:
null

Results:
{
  "EnableConfigs": [
    "SWAP",
    "MEMCG",
    "ZSWAP",
    "VM_EVENT_COUNTERS"
  ],
  "FocusSymbols": [
    "folio_alloc_swap",
    "swap_writeout",
    "get_nr_swap_pages_eligible",
    "mem_cgroup_get_nr_swap_pages_eligible",
    "folio_swap_full",
    "user_proactive_reclaim",
    "__x64_sys_swapon"
  ],
  "KMSANReasoning": "The patch series introduces offload-only swap areas (SWAP_FLAG_OFFLOAD_ONLY / SWP_OFFLOAD_ONLY), allowing swap devices to be designated solely for proactive reclaim (e.g. cgroup v2 memory.reclaim).\n\nKMSAN vs KASAN Applicability Analysis:\n- Uninitialized Memory Risks (KMSAN):\n  1. No new structures or buffers are copied to user space; user-space interfaces modified are the input flag to the `swapon` syscall and an integer event counter reported via `/proc/vmstat`.\n  2. The new fields added to kernel structures (`allow_offload_swap` in `struct reclaim_state`, `struct scan_control`, and `struct percpu_swap_cluster`) are always initialized before use: `scan_control` instances across vmscan are zero-initialized via C99 designated initializers, `percpu_swap_cluster` is defined in static per-CPU storage, and `current-\u003ereclaim_state` is strictly scoped within `set_task_reclaim_state()`.\n  3. No branching or conditional logic depends on uninitialized data.\n\n- Memory Corruption / Concurrency Risks (KASAN / LOCKDEP):\n  1. Any potential risks introduced by the patch involve swap slot accounting, cluster selection logic, priority routing, or device-level locking/synchronization.\n  2. These failure modes (use-after-free, double frees, invalid dereferences, locking deadlocks) are already covered by standard KASAN and LOCKDEP configurations.\n\nBecause there are no uninitialized memory use or kernel info-leak risks introduced by this patch series, spawning a dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false,
  "Reasoning": "The patch series modifies core memory management code (mm/swapfile.c, mm/vmscan.c, mm/page_io.c, mm/memcontrol.c) to introduce offload-only swap areas via the new SWAP_FLAG_OFFLOAD_ONLY flag in swapon(). It alters swap slot allocation (both fast and slow paths), eligible swap accounting, swapcache reclaim heuristics, swap writeout refusal paths for retained entries, and zswap integration. Furthermore, it adds new runtime assertions (VM_WARN_ON_FOLIO) in folio_swap_full(). All modified paths are fully reachable in virtualized environments via swapon(), cgroup v2 memory.reclaim, and standard memory reclaim under pressure.",
  "WorthFuzzing": true
}

1/1 2026/09/14 12:10 action read-patch-diff
0m Args:
null

Results:
{
  "PatchDiff": "commit b18205544615b8a7fdfb59a7d24849375e130428\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate:   Mon Sep 14 12:10:57 2026 +0000\n\n    syz-cluster: applied patch under review\n\ndiff --git a/Documentation/mm/swap.rst b/Documentation/mm/swap.rst\nindex 78819bd4d7452..7d5bbd65e6160 100644\n--- a/Documentation/mm/swap.rst\n+++ b/Documentation/mm/swap.rst\n@@ -3,3 +3,90 @@\n ====\n Swap\n ====\n+\n+Offload-only swap areas\n+-----------------------\n+\n+``SWAP_FLAG_OFFLOAD_ONLY`` marks a swap area as a destination for explicit\n+userspace proactive reclaim.  The stable qualifying interfaces are cgroup v2\n+``memory.reclaim`` and ``/sys/devices/system/node/nodeX/reclaim``.  The swap\n+allocator excludes such an area from kswapd, direct reclaim, and other\n+pressure-driven swap allocation.  Normal swap priority ordering still applies\n+among the areas eligible for the current reclaim context.\n+\n+The MGLRU debugfs eviction interface currently establishes the same internal\n+proactive-reclaim provenance and can therefore use an offload-only area.\n+Debugfs is not a stable userspace ABI, however, so that behaviour is not part\n+of this interface's permanent contract.\n+\n+This permits a system to combine a small conventional swap area, which is\n+engineered for forward progress in emergency reclaim, with a larger or more\n+complex area used for ordinary cold-page offload.  For example, the latter may\n+be RAM-compressed or may use a filesystem with compression, checksums, or\n+redundancy.  Making every such write path safe in direct reclaim can require\n+backend-specific reserves, preallocation, non-blocking allocation, and\n+recursion rules.  Excluding new pressure-reclaim writes can reduce that\n+requirement and its complexity.\n+\n+For a RAM-compressed area such as zram, unused logical slots also do not imply\n+that enough physical memory remains to store their future contents.  Static\n+swap priority cannot express that distinction or provide late fallback after a\n+selected area's write fails.\n+\n+The policy is attached to an activated swap area, not to its underlying\n+physical storage.  A raw swap partition and a filesystem swapfile on the same\n+device are separate areas and may use different policies.  The kernel does not\n+infer this policy from the block driver, filesystem, or swap priority.\n+\n+This is a reclaim-provenance policy, not a measurement of current memory\n+headroom.  Userspace should only request proactive offload while its own\n+watermark or PSI policy considers memory pressure low.\n+\n+DAMON reclaim and ``MADV_PAGEOUT`` do not currently establish the proactive\n+reclaim context, so they cannot allocate slots from an offload-only area.\n+Offload-only areas are also ineligible for hibernation image allocation.\n+\n+Offload-only areas bypass zswap stores.  Zswap writeback may run after the\n+proactive context which selected the slot has ended, so admitting the folio to\n+zswap would otherwise defer the backend write beyond that context.  The\n+hierarchical cgroup v2 ``memory.zswap.writeback=0`` policy remains\n+authoritative: when zswap is enabled, it also refuses direct proactive writes\n+to an offload-only area.  Marking an area offload-only does not override a\n+cgroup policy which disables all swapping attempts to devices.\n+\n+The flag controls both allocation of new swap slots and newly initiated\n+non-zero backend writes.  A folio can retain its swap entry after swapin.  If\n+ordinary reclaim later tries to rewrite such an offload-only entry, the VM\n+redirties and activates the folio instead; proactive reclaim may retry the\n+write.  Zero-filled folios may still update the in-memory swap zeromap without\n+backend I/O.\n+\n+The flag does not prevent reads, swapoff, or writes which are already queued or\n+in flight.  It therefore does not by itself provide a forward progress\n+guarantee for an I/O path which allocates memory: earlier writes must still be\n+able to complete, and reads must remain reclaim-safe.  Repeatedly refusing\n+retained-entry writes can also reduce reclaim efficiency and lead to OOM while\n+the dirty folios remain resident.\n+\n+Page-cluster discard is incompatible with an offload-only area because its\n+work item can run after the context which freed the entries has ended.  Swapon\n+therefore rejects a resolved page-cluster discard policy combined with\n+``SWAP_FLAG_OFFLOAD_ONLY``.  Swapon-time discard is permitted because it\n+completes synchronously during activation.  The existing discard precedence\n+still applies: requesting both discard-once and discard-pages selects\n+discard-once.  Discard requests which the swap area does not support remain\n+ignored.\n+\n+Architecture-specific swap metadata preparation still runs before a retained\n+write is refused, so that metadata remains coherent with the dirty resident\n+folio.  This policy controls swap-backend I/O; it does not promise that core VM\n+or architecture preparation performs no allocation.\n+\n+``/proc/vmstat`` reports ``swpout_offload_refused`` in base pages.  The counter\n+advances when ordinary reclaim refuses a newly initiated write through a\n+retained offload-only entry.  It can reveal repeated activation of such folios,\n+but it is not a count of skipped areas during new-slot allocation.\n+\n+An offload-only area should therefore be configured with a reclaim-safe swap\n+area as fallback.  If no eligible swap space remains, swap allocation fails\n+and the existing reclaim and OOM policy applies.\ndiff --git a/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c b/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c\nindex e0d1f369a1632..b58e61f15ab10 100644\n--- a/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c\n+++ b/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c\n@@ -21,7 +21,7 @@\n \n static bool swap_available(void)\n {\n-\treturn get_nr_swap_pages() \u003e 0;\n+\treturn get_nr_swap_pages_eligible() \u003e 0;\n }\n \n static bool can_release_pages(struct drm_i915_gem_object *obj)\ndiff --git a/drivers/gpu/drm/i915/gem/selftests/huge_pages.c b/drivers/gpu/drm/i915/gem/selftests/huge_pages.c\nindex 44718e7282914..73b62b0655103 100644\n--- a/drivers/gpu/drm/i915/gem/selftests/huge_pages.c\n+++ b/drivers/gpu/drm/i915/gem/selftests/huge_pages.c\n@@ -1895,8 +1895,8 @@ static int igt_shrink_thp(void *arg)\n \ti915_gem_context_unlock_engines(ctx);\n \t/*\n \t * Nuke everything *before* we unpin the pages so we can be reasonably\n-\t * sure that when later checking get_nr_swap_pages() that some random\n-\t * leftover object doesn't steal the remaining swap space.\n+\t * sure that when later checking get_nr_swap_pages_eligible() that some\n+\t * random leftover object doesn't steal the remaining swap space.\n \t */\n \ti915_gem_shrink(NULL, i915, -1UL, NULL,\n \t\t\tI915_SHRINK_BOUND |\n@@ -1910,7 +1910,7 @@ static int igt_shrink_thp(void *arg)\n \t * Now that the pages are *unpinned* shrinking should invoke\n \t * shmem to truncate our pages, if we have available swap.\n \t */\n-\tshould_swap = get_nr_swap_pages() \u003e 0;\n+\tshould_swap = get_nr_swap_pages_eligible() \u003e 0;\n \ti915_gem_shrink(NULL, i915, -1UL, NULL,\n \t\t\tI915_SHRINK_BOUND |\n \t\t\tI915_SHRINK_UNBOUND |\ndiff --git a/drivers/gpu/drm/msm/msm_gem_shrinker.c b/drivers/gpu/drm/msm/msm_gem_shrinker.c\nindex 9d2788f79acee..da6a66b747b02 100644\n--- a/drivers/gpu/drm/msm/msm_gem_shrinker.c\n+++ b/drivers/gpu/drm/msm/msm_gem_shrinker.c\n@@ -21,7 +21,7 @@ module_param(enable_eviction, bool, 0600);\n \n static bool can_swap(void)\n {\n-\treturn enable_eviction \u0026\u0026 get_nr_swap_pages() \u003e 0;\n+\treturn enable_eviction \u0026\u0026 get_nr_swap_pages_eligible() \u003e 0;\n }\n \n static bool can_block(struct shrink_control *sc)\ndiff --git a/drivers/gpu/drm/panthor/panthor_gem.c b/drivers/gpu/drm/panthor/panthor_gem.c\nindex 72908be5e144f..23488f417cf7a 100644\n--- a/drivers/gpu/drm/panthor/panthor_gem.c\n+++ b/drivers/gpu/drm/panthor/panthor_gem.c\n@@ -1371,7 +1371,7 @@ panthor_dummy_bo_create(struct panthor_device *ptdev)\n \n static bool can_swap(void)\n {\n-\treturn get_nr_swap_pages() \u003e 0;\n+\treturn get_nr_swap_pages_eligible() \u003e 0;\n }\n \n static bool can_block(struct shrink_control *sc)\ndiff --git a/drivers/gpu/drm/ttm/ttm_backup.c b/drivers/gpu/drm/ttm/ttm_backup.c\nindex 0c2d53a13b2a6..bf9d41bdf652c 100644\n--- a/drivers/gpu/drm/ttm/ttm_backup.c\n+++ b/drivers/gpu/drm/ttm/ttm_backup.c\n@@ -206,7 +206,7 @@ u64 ttm_backup_bytes_avail(void)\n \t * number also depends on shmem actually swapping out backed-up\n \t * shmem objects without too much buffering.\n \t */\n-\treturn (u64)get_nr_swap_pages() \u003c\u003c PAGE_SHIFT;\n+\treturn (u64)get_nr_swap_pages_eligible() \u003c\u003c PAGE_SHIFT;\n }\n EXPORT_SYMBOL_GPL(ttm_backup_bytes_avail);\n \ndiff --git a/drivers/gpu/drm/xe/tests/xe_bo.c b/drivers/gpu/drm/xe/tests/xe_bo.c\nindex 6a17e13d58cf1..14d6bb8e41c9c 100644\n--- a/drivers/gpu/drm/xe/tests/xe_bo.c\n+++ b/drivers/gpu/drm/xe/tests/xe_bo.c\n@@ -695,7 +695,7 @@ static int shrink_test_run_device(struct xe_device *xe)\n \t}\n \tto_alloc = ram * 2;\n \n-\tram_and_swap = ram + get_nr_swap_pages() * PAGE_SIZE;\n+\tram_and_swap = ram + get_nr_swap_pages_eligible() * PAGE_SIZE;\n \tif (to_alloc \u003e ram_and_swap)\n \t\tpurgeable = to_alloc - ram_and_swap;\n \tpurgeable += div64_u64(purgeable, 5);\ndiff --git a/include/linux/swap.h b/include/linux/swap.h\nindex 43155e122b5c3..19a0bb4ba6c93 100644\n--- a/include/linux/swap.h\n+++ b/include/linux/swap.h\n@@ -21,10 +21,11 @@\n #define SWAP_FLAG_DISCARD\t0x10000 /* enable discard for swap */\n #define SWAP_FLAG_DISCARD_ONCE\t0x20000 /* discard swap area at swapon-time */\n #define SWAP_FLAG_DISCARD_PAGES 0x40000 /* discard page-clusters after use */\n+#define SWAP_FLAG_OFFLOAD_ONLY  0x80000 /* only use for proactive reclaim */\n \n #define SWAP_FLAGS_VALID\t(SWAP_FLAG_PRIO_MASK | SWAP_FLAG_PREFER | \\\n \t\t\t\t SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE | \\\n-\t\t\t\t SWAP_FLAG_DISCARD_PAGES)\n+\t\t\t\t SWAP_FLAG_DISCARD_PAGES | SWAP_FLAG_OFFLOAD_ONLY)\n /*\n  * MAX_SWAPFILES defines the maximum number of swaptypes: things which can\n  * be swapped to.  The swap type and the offset into that swap type are\n@@ -140,12 +141,20 @@ union swap_header {\n struct reclaim_state {\n \t/* pages reclaimed outside of LRU-based reclaim */\n \tunsigned long reclaimed;\n+\t/* this reclaim context may use offload-only swap */\n+\tbool allow_offload_swap;\n #ifdef CONFIG_LRU_GEN\n \t/* per-thread mm walk data */\n \tstruct lru_gen_mm_walk *mm_walk;\n #endif\n };\n \n+static inline bool current_reclaim_allows_offload_swap(void)\n+{\n+\treturn current-\u003ereclaim_state \u0026\u0026\n+\t       current-\u003ereclaim_state-\u003eallow_offload_swap;\n+}\n+\n /*\n  * mm_account_reclaimed_pages(): account reclaimed pages outside of LRU-based\n  * reclaim\n@@ -201,6 +210,7 @@ enum {\n \tSWP_STABLE_WRITES = (1 \u003c\u003c 11),\t/* no overwrite PG_writeback pages */\n \tSWP_SYNCHRONOUS_IO = (1 \u003c\u003c 12),\t/* synchronous IO is efficient */\n \tSWP_HIBERNATION = (1 \u003c\u003c 13),\t/* pinned for hibernation */\n+\tSWP_OFFLOAD_ONLY = (1 \u003c\u003c 14),\t/* proactive-reclaim swap only */\n \t\t\t\t\t/* add others here before... */\n };\n \n@@ -389,6 +399,9 @@ static inline long get_nr_swap_pages(void)\n \treturn atomic_long_read(\u0026nr_swap_pages);\n }\n \n+long get_nr_swap_pages_eligible(void);\n+bool folio_swap_full(struct folio *folio);\n+\n extern void si_swapinfo(struct sysinfo *);\n extern int pin_hibernation_swap_type(dev_t device, sector_t offset);\n extern void unpin_hibernation_swap_type(int type);\n@@ -443,10 +456,16 @@ static inline void put_swap_device(struct swap_info_struct *si)\n }\n \n #define get_nr_swap_pages()\t\t\t0L\n+#define get_nr_swap_pages_eligible()\t\t0L\n #define total_swap_pages\t\t\t0L\n #define total_swapcache_pages()\t\t\t0UL\n #define vm_swap_full()\t\t\t\t0\n \n+static inline bool folio_swap_full(struct folio *folio)\n+{\n+\treturn false;\n+}\n+\n #define si_swapinfo(val) \\\n \tdo { (val)-\u003efreeswap = (val)-\u003etotalswap = 0; } while (0)\n #define free_folio_and_swap_cache(folio) \\\n@@ -531,6 +550,7 @@ static inline void mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_p\n \n long mem_cgroup_get_folio_swap_margin(struct folio *folio);\n extern long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg);\n+long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg);\n extern bool mem_cgroup_swap_full(struct folio *folio);\n #else\n static inline int mem_cgroup_try_charge_swap(struct folio *folio)\n@@ -553,9 +573,14 @@ static inline long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)\n \treturn get_nr_swap_pages();\n }\n \n+static inline long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg)\n+{\n+\treturn get_nr_swap_pages_eligible();\n+}\n+\n static inline bool mem_cgroup_swap_full(struct folio *folio)\n {\n-\treturn vm_swap_full();\n+\treturn folio_swap_full(folio);\n }\n #endif\n \ndiff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h\nindex 2628ccda076a0..fc8458b314a50 100644\n--- a/include/linux/vm_event_item.h\n+++ b/include/linux/vm_event_item.h\n@@ -32,6 +32,7 @@\n \tHIGHMEM_ZONE(xx) xx##_MOVABLE, DEVICE_ZONE(xx)\n \n enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT,\n+\t\tSWPOUT_OFFLOAD_REFUSED,\n \t\tFOR_ALL_ZONES(PGALLOC)\n \t\tFOR_ALL_ZONES(ALLOCSTALL)\n \t\tFOR_ALL_ZONES(PGSCAN_SKIP)\ndiff --git a/mm/memcontrol.c b/mm/memcontrol.c\nindex 1460cba53588e..1412084d2f430 100644\n--- a/mm/memcontrol.c\n+++ b/mm/memcontrol.c\n@@ -6002,16 +6002,28 @@ void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages)\n \trcu_read_unlock();\n }\n \n-long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)\n+static long\n+mem_cgroup_get_nr_swap_pages_with_limit(struct mem_cgroup *memcg,\n+\t\t\t\t\tlong nr_swap_pages)\n {\n-\tlong nr_swap_pages = get_nr_swap_pages();\n-\n \tif (!mem_cgroup_disabled() \u0026\u0026 !do_memsw_account())\n \t\tnr_swap_pages = min(nr_swap_pages, page_counter_margin(\u0026memcg-\u003eswap));\n \n \treturn nr_swap_pages;\n }\n \n+long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)\n+{\n+\treturn mem_cgroup_get_nr_swap_pages_with_limit(memcg,\n+\t\t\tget_nr_swap_pages());\n+}\n+\n+long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg)\n+{\n+\treturn mem_cgroup_get_nr_swap_pages_with_limit(memcg,\n+\t\t\tget_nr_swap_pages_eligible());\n+}\n+\n /**\n  * mem_cgroup_get_folio_swap_margin - get a folio's memcg swap margin\n  * @folio: folio whose memcg margin is queried\n@@ -6042,7 +6054,7 @@ bool mem_cgroup_swap_full(struct folio *folio)\n \n \tVM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);\n \n-\tif (vm_swap_full())\n+\tif (folio_swap_full(folio))\n \t\treturn true;\n \tif (do_memsw_account() || !folio_memcg_charged(folio))\n \t\treturn ret;\ndiff --git a/mm/page_io.c b/mm/page_io.c\nindex 1da4ff484f097..1b281a1d6df0a 100644\n--- a/mm/page_io.c\n+++ b/mm/page_io.c\n@@ -203,6 +203,7 @@ static void swap_zeromap_folio_clear(struct folio *folio)\n  */\n int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)\n {\n+\tstruct swap_info_struct *sis = __swap_entry_to_info(folio-\u003eswap);\n \tint ret = 0;\n \n \tif (folio_free_swap(folio))\n@@ -210,7 +211,9 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)\n \n \t/*\n \t * Arch code may have to preserve more data than just the folio\n-\t * contents, e.g. memory tags.\n+\t * contents, e.g. memory tags.  Do this before refusing a retained\n+\t * offload-only entry below: a later sibling swap-PTE fault can restore\n+\t * swap-indexed metadata into this resident folio.\n \t */\n \tret = arch_prepare_to_swap(folio);\n \tif (ret) {\n@@ -228,6 +231,18 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)\n \t\tgoto out_unlock;\n \t}\n \n+\t/*\n+\t * A folio can retain an existing swap entry after swapin.  Do not let\n+\t * ordinary reclaim use an offload-only entry through that path.\n+\t */\n+\tif ((READ_ONCE(sis-\u003eflags) \u0026 SWP_OFFLOAD_ONLY) \u0026\u0026\n+\t    !current_reclaim_allows_offload_swap()) {\n+\t\tcount_vm_events(SWPOUT_OFFLOAD_REFUSED,\n+\t\t\t\tfolio_nr_pages(folio));\n+\t\tfolio_mark_dirty(folio);\n+\t\treturn AOP_WRITEPAGE_ACTIVATE;\n+\t}\n+\n \t/*\n \t * Clear bits this folio occupies in the zeromap to prevent zero data\n \t * being read in from any previous zero writes that occupied the same\n@@ -235,7 +250,12 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)\n \t */\n \tswap_zeromap_folio_clear(folio);\n \n-\tif (zswap_store(folio)) {\n+\t/*\n+\t * Zswap writeback can happen much later from pressure reclaim or its\n+\t * shrinker workqueue.  Do not let it defer an offload-only backend write\n+\t * beyond the proactive reclaim context which admitted the swap slot.\n+\t */\n+\tif (!(READ_ONCE(sis-\u003eflags) \u0026 SWP_OFFLOAD_ONLY) \u0026\u0026 zswap_store(folio)) {\n \t\tcount_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT);\n \t\tgoto out_unlock;\n \t}\ndiff --git a/mm/swapfile.c b/mm/swapfile.c\nindex 280dd906eb187..256617059f302 100644\n--- a/mm/swapfile.c\n+++ b/mm/swapfile.c\n@@ -65,13 +65,17 @@ static void move_cluster(struct swap_info_struct *si,\n  */\n static DEFINE_SPINLOCK(swap_lock);\n static unsigned int nr_swapfiles;\n-atomic_long_t nr_swap_pages;\n /*\n  * Some modules use swappable objects and may try to swap them out under\n  * memory pressure (via the shrinker). Before doing so, they may wish to\n  * check to see if any swap space is available.\n+ *\n+ * This remains the raw free-space counter for accounting users. Reclaim\n+ * decisions subtract nr_swap_pages_offload_only below when necessary.\n  */\n+atomic_long_t nr_swap_pages;\n EXPORT_SYMBOL_GPL(nr_swap_pages);\n+static atomic_long_t nr_swap_pages_offload_only;\n /* protected with swap_lock. reading in vm_swap_full() doesn't need lock */\n long total_swap_pages;\n #define DEF_SWAP_PRIO  -1\n@@ -120,6 +124,7 @@ atomic_t nr_rotate_swap = ATOMIC_INIT(0);\n struct percpu_swap_cluster {\n \tstruct swap_info_struct *si[SWAP_NR_ORDERS];\n \tunsigned long offset[SWAP_NR_ORDERS];\n+\tbool allow_offload_swap[SWAP_NR_ORDERS];\n \tlocal_lock_t lock;\n };\n \n@@ -163,6 +168,31 @@ static long swap_usage_in_pages(struct swap_info_struct *si)\n \treturn atomic_long_read(\u0026si-\u003einuse_pages) \u0026 SWAP_USAGE_COUNTER_MASK;\n }\n \n+static bool swap_area_needs_reclaim(struct swap_info_struct *si)\n+{\n+\tif (vm_swap_full())\n+\t\treturn true;\n+\n+\t/*\n+\t * Free offload-only slots must not keep a full conventional area\n+\t * pinned in swapcache.  Recover that area's unused cache entries\n+\t * even when the raw pool is not full.  This is independent of the\n+\t * current task: proactive reclaim can fill the conventional area too.\n+\t */\n+\treturn !(READ_ONCE(si-\u003eflags) \u0026 SWP_OFFLOAD_ONLY) \u0026\u0026\n+\t\tatomic_long_read(\u0026nr_swap_pages_offload_only) \u003e 0 \u0026\u0026\n+\t\tswap_usage_in_pages(si) == si-\u003epages;\n+}\n+\n+/* The caller must hold the lock on a folio in swapcache. */\n+bool folio_swap_full(struct folio *folio)\n+{\n+\tVM_WARN_ON_FOLIO(!folio_test_locked(folio), folio);\n+\tVM_WARN_ON_FOLIO(!folio_test_swapcache(folio), folio);\n+\n+\treturn swap_area_needs_reclaim(__swap_entry_to_info(folio-\u003eswap));\n+}\n+\n /* Reclaim the swap entry anyway if possible */\n #define TTRS_ANYWAY\t\t0x1\n /*\n@@ -903,7 +933,7 @@ static bool cluster_scan_range(struct swap_info_struct *si,\n \t\tif (swp_tb_is_null(swp_tb))\n \t\t\tcontinue;\n \t\tif (swp_tb_is_folio(swp_tb) \u0026\u0026 !__swp_tb_get_count(swp_tb)) {\n-\t\t\tif (!vm_swap_full())\n+\t\t\tif (!swap_area_needs_reclaim(si))\n \t\t\t\treturn false;\n \t\t\t*need_reclaim = true;\n \t\t\tcontinue;\n@@ -1014,6 +1044,8 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,\n \tif (si-\u003eflags \u0026 SWP_SOLIDSTATE) {\n \t\tthis_cpu_write(percpu_swap_cluster.offset[order], next);\n \t\tthis_cpu_write(percpu_swap_cluster.si[order], si);\n+\t\tthis_cpu_write(percpu_swap_cluster.allow_offload_swap[order],\n+\t\t\t       current_reclaim_allows_offload_swap());\n \t} else {\n \t\tsi-\u003eglobal_cluster-\u003enext[order] = next;\n \t}\n@@ -1161,7 +1193,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,\n \t}\n \n \t/* Try reclaim full clusters if free and nonfull lists are drained */\n-\tif (vm_swap_full())\n+\tif (swap_area_needs_reclaim(si))\n \t\tswap_reclaim_full_clusters(si, false);\n \n \tif (order \u003c PMD_ORDER) {\n@@ -1315,9 +1347,11 @@ static void swap_range_alloc(struct swap_info_struct *si,\n \t\t\t     unsigned int nr_entries)\n {\n \tif (swap_usage_add(si, nr_entries)) {\n-\t\tif (vm_swap_full())\n+\t\tif (swap_area_needs_reclaim(si))\n \t\t\tschedule_work(\u0026si-\u003ereclaim_work);\n \t}\n+\tif (si-\u003eflags \u0026 SWP_OFFLOAD_ONLY)\n+\t\tatomic_long_sub(nr_entries, \u0026nr_swap_pages_offload_only);\n \tatomic_long_sub(nr_entries, \u0026nr_swap_pages);\n }\n \n@@ -1346,6 +1380,8 @@ static void swap_range_free(struct swap_info_struct *si, unsigned long offset,\n \t * only after the above cleanups are done.\n \t */\n \tsmp_wmb();\n+\tif (si-\u003eflags \u0026 SWP_OFFLOAD_ONLY)\n+\t\tatomic_long_add(nr_entries, \u0026nr_swap_pages_offload_only);\n \tatomic_long_add(nr_entries, \u0026nr_swap_pages);\n \tswap_usage_sub(si, nr_entries);\n }\n@@ -1366,6 +1402,31 @@ static bool get_swap_device_info(struct swap_info_struct *si)\n \treturn true;\n }\n \n+static bool swap_area_eligible(struct swap_info_struct *si)\n+{\n+\tif (!(READ_ONCE(si-\u003eflags) \u0026 SWP_OFFLOAD_ONLY))\n+\t\treturn true;\n+\n+\treturn current_reclaim_allows_offload_swap();\n+}\n+\n+long get_nr_swap_pages_eligible(void)\n+{\n+\tlong nr_pages;\n+\n+\tif (current_reclaim_allows_offload_swap())\n+\t\treturn get_nr_swap_pages();\n+\n+\t/*\n+\t * The reads are intentionally unpaired.  This is a capacity hint; the\n+\t * allocator enforces eligibility.  Clamp a transient negative result.\n+\t */\n+\tnr_pages = get_nr_swap_pages() -\n+\t\t   atomic_long_read(\u0026nr_swap_pages_offload_only);\n+\treturn max(nr_pages, 0L);\n+}\n+EXPORT_SYMBOL_GPL(get_nr_swap_pages_eligible);\n+\n /*\n  * Fast path try to get swap entries with specified order from current\n  * CPU's swap entry pool (a cluster).\n@@ -1385,6 +1446,20 @@ static bool swap_alloc_fast(struct folio *folio)\n \toffset = this_cpu_read(percpu_swap_cluster.offset[order]);\n \tif (!si || !offset || !get_swap_device_info(si))\n \t\treturn false;\n+\tif (!swap_area_eligible(si)) {\n+\t\tput_swap_device(si);\n+\t\treturn false;\n+\t}\n+\t/*\n+\t * Pressure reclaim may cache a lower-priority conventional area while\n+\t * an offload-only area is ineligible.  Drop that cache on a context\n+\t * change so proactive reclaim returns to the normal priority search.\n+\t */\n+\tif (this_cpu_read(percpu_swap_cluster.allow_offload_swap[order]) !=\n+\t    current_reclaim_allows_offload_swap()) {\n+\t\tput_swap_device(si);\n+\t\treturn false;\n+\t}\n \n \tci = swap_cluster_lock(si, offset);\n \tif (cluster_is_usable(ci, order)) {\n@@ -1407,6 +1482,9 @@ static void swap_alloc_slow(struct folio *folio)\n \tspin_lock(\u0026swap_avail_lock);\n start_over:\n \tplist_for_each_entry_safe(si, next, \u0026swap_avail_head, avail_list) {\n+\t\tif (!swap_area_eligible(si))\n+\t\t\tcontinue;\n+\n \t\t/* Rotate the device and switch to a new cluster */\n \t\tplist_requeue(\u0026si-\u003eavail_list, \u0026swap_avail_head);\n \t\tspin_unlock(\u0026swap_avail_lock);\n@@ -1450,7 +1528,8 @@ static bool swap_sync_discard(void)\n \tplist_for_each_entry_safe(si, next, \u0026swap_active_head, list) {\n \t\tspin_unlock(\u0026swap_lock);\n \t\tif (get_swap_device_info(si)) {\n-\t\t\tif (si-\u003eflags \u0026 SWP_PAGE_DISCARD)\n+\t\t\tif (swap_area_eligible(si) \u0026\u0026\n+\t\t\t    (si-\u003eflags \u0026 SWP_PAGE_DISCARD))\n \t\t\t\tret = swap_do_scheduled_discard(si);\n \t\t\tput_swap_device(si);\n \t\t}\n@@ -1738,8 +1817,8 @@ static int swap_dup_entries_cluster(struct swap_info_struct *si,\n  *\n  * Context: Caller needs to hold the folio lock.\n  * Return: %0 on success, %-E2BIG if splitting the folio might allow swapout,\n- * %-ENOSPC if no global swap space is available, or %-ENOMEM if splitting\n- * would not help.\n+ * %-ENOSPC if no global swap space is eligible for the caller, or %-ENOMEM\n+ * if splitting would not help.\n  */\n int folio_alloc_swap(struct folio *folio)\n {\n@@ -1790,7 +1869,7 @@ int folio_alloc_swap(struct folio *folio)\n \treturn 0;\n \n failed:\n-\tif (get_nr_swap_pages() \u003c= 0)\n+\tif (get_nr_swap_pages_eligible() \u003c= 0)\n \t\treturn -ENOSPC;\n \tif (mem_cgroup_get_folio_swap_margin(folio) \u003c= 0)\n \t\treturn -ENOMEM;\n@@ -2180,7 +2259,7 @@ swp_entry_t swap_alloc_hibernation_slot(int type)\n \tstruct swap_cluster_info *ci;\n \tswp_entry_t entry = {0};\n \n-\tif (!si)\n+\tif (!si || (si-\u003eflags \u0026 SWP_OFFLOAD_ONLY))\n \t\tgoto fail;\n \n \t/*\n@@ -2247,7 +2326,8 @@ static int __find_hibernation_swap_type(dev_t device, sector_t offset)\n \tfor (type = 0; type \u003c nr_swapfiles; type++) {\n \t\tstruct swap_info_struct *sis = swap_info[type];\n \n-\t\tif (!(sis-\u003eflags \u0026 SWP_WRITEOK))\n+\t\tif (!(sis-\u003eflags \u0026 SWP_WRITEOK) ||\n+\t\t    (sis-\u003eflags \u0026 SWP_OFFLOAD_ONLY))\n \t\t\tcontinue;\n \n \t\tif (device == sis-\u003ebdev-\u003ebd_dev) {\n@@ -2434,7 +2514,8 @@ int find_first_swap(dev_t *device)\n \tfor (type = 0; type \u003c nr_swapfiles; type++) {\n \t\tstruct swap_info_struct *sis = swap_info[type];\n \n-\t\tif (!(sis-\u003eflags \u0026 SWP_WRITEOK))\n+\t\tif (!(sis-\u003eflags \u0026 SWP_WRITEOK) ||\n+\t\t    (sis-\u003eflags \u0026 SWP_OFFLOAD_ONLY))\n \t\t\tcontinue;\n \t\t*device = sis-\u003ebdev-\u003ebd_dev;\n \t\tspin_unlock(\u0026swap_lock);\n@@ -2474,7 +2555,8 @@ unsigned int count_swap_pages(int type, int free)\n \t\tstruct swap_info_struct *sis = swap_info[type];\n \n \t\tspin_lock(\u0026sis-\u003elock);\n-\t\tif (sis-\u003eflags \u0026 SWP_WRITEOK) {\n+\t\tif ((sis-\u003eflags \u0026 SWP_WRITEOK) \u0026\u0026\n+\t\t    !(sis-\u003eflags \u0026 SWP_OFFLOAD_ONLY)) {\n \t\t\tn = sis-\u003epages;\n \t\t\tif (free)\n \t\t\t\tn -= swap_usage_in_pages(sis);\n@@ -3083,6 +3165,8 @@ static int setup_swap_extents(struct swap_info_struct *sis,\n \n static void _enable_swap_info(struct swap_info_struct *si)\n {\n+\tif (si-\u003eflags \u0026 SWP_OFFLOAD_ONLY)\n+\t\tatomic_long_add(si-\u003epages, \u0026nr_swap_pages_offload_only);\n \tatomic_long_add(si-\u003epages, \u0026nr_swap_pages);\n \ttotal_swap_pages += si-\u003epages;\n \n@@ -3231,6 +3315,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)\n \tspin_lock(\u0026p-\u003elock);\n \tdel_from_avail_list(p, true);\n \tplist_del(\u0026p-\u003elist, \u0026swap_active_head);\n+\tif (p-\u003eflags \u0026 SWP_OFFLOAD_ONLY)\n+\t\tatomic_long_sub(p-\u003epages, \u0026nr_swap_pages_offload_only);\n \tatomic_long_sub(p-\u003epages, \u0026nr_swap_pages);\n \ttotal_swap_pages -= p-\u003epages;\n \tspin_unlock(\u0026p-\u003elock);\n@@ -3727,7 +3813,6 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)\n \n \tif (swap_flags \u0026 ~SWAP_FLAGS_VALID)\n \t\treturn -EINVAL;\n-\n \tif (!capable(CAP_SYS_ADMIN))\n \t\treturn -EPERM;\n \n@@ -3841,6 +3926,9 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)\n \tif (error)\n \t\tgoto bad_swap_unlock_inode;\n \n+\tif (swap_flags \u0026 SWAP_FLAG_OFFLOAD_ONLY)\n+\t\tsi-\u003eflags |= SWP_OFFLOAD_ONLY;\n+\n \tif ((swap_flags \u0026 SWAP_FLAG_DISCARD) \u0026\u0026\n \t    si-\u003ebdev \u0026\u0026 bdev_max_discard_sectors(si-\u003ebdev)) {\n \t\t/*\n@@ -3863,6 +3951,18 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)\n \t\telse if (swap_flags \u0026 SWAP_FLAG_DISCARD_PAGES)\n \t\t\tsi-\u003eflags \u0026= ~SWP_AREA_DISCARD;\n \n+\t\t/*\n+\t\t * Cluster discard can run later from discard_work, after the\n+\t\t * context which freed the entries has ended.  Swapon-time discard\n+\t\t * is explicit and synchronous, but page discard cannot honour\n+\t\t * offload provenance.\n+\t\t */\n+\t\tif ((si-\u003eflags \u0026 SWP_OFFLOAD_ONLY) \u0026\u0026\n+\t\t    (si-\u003eflags \u0026 SWP_PAGE_DISCARD)) {\n+\t\t\terror = -EINVAL;\n+\t\t\tgoto bad_swap_unlock_inode;\n+\t\t}\n+\n \t\t/* issue a swapon-time discard if it's still required */\n \t\tif (si-\u003eflags \u0026 SWP_AREA_DISCARD) {\n \t\t\tint err = discard_swap(si);\ndiff --git a/mm/vmscan.c b/mm/vmscan.c\nindex c2eb8fa9d5e50..201b71b760492 100644\n--- a/mm/vmscan.c\n+++ b/mm/vmscan.c\n@@ -123,6 +123,9 @@ struct scan_control {\n \t/* Proactive reclaim invoked by userspace */\n \tunsigned int proactive:1;\n \n+\t/* This reclaim context may use offload-only swap */\n+\tunsigned int allow_offload_swap:1;\n+\n \t/*\n \t * Cgroup memory below memory.low is protected as long as we\n \t * don't threaten to OOM. If any cgroup is reclaimed at\n@@ -291,14 +294,18 @@ static inline bool is_exec_file_folio(const struct folio *folio,\n }\n \n static void set_task_reclaim_state(struct task_struct *task,\n-\t\t\t\t   struct reclaim_state *rs)\n+\t\t\t\t   struct scan_control *sc)\n {\n+\tstruct reclaim_state *rs = sc ? \u0026sc-\u003ereclaim_state : NULL;\n+\n \t/* Check for an overwrite */\n \tWARN_ON_ONCE(rs \u0026\u0026 task-\u003ereclaim_state);\n \n \t/* Check for the nulling of an already-nulled member */\n \tWARN_ON_ONCE(!rs \u0026\u0026 !task-\u003ereclaim_state);\n \n+\tif (rs)\n+\t\trs-\u003eallow_offload_swap = sc-\u003eallow_offload_swap;\n \ttask-\u003ereclaim_state = rs;\n }\n \n@@ -418,7 +425,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,\n \t\t * And under GFP_NOIO, is there enough swapcached anon to make\n \t\t * scanning anon worthwhile?\n \t\t */\n-\t\tif (get_nr_swap_pages() \u003e 0 \u0026\u0026\n+\t\tif (get_nr_swap_pages_eligible() \u003e 0 \u0026\u0026\n \t\t    !reclaimable_anon_is_low(memcg, nid, sc))\n \t\t\treturn true;\n \t} else {\n@@ -426,7 +433,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,\n \t\t * Is the memcg below its swap limit, and under GFP_NOIO does\n \t\t * it have enough swapcached anon to make scanning worthwhile?\n \t\t */\n-\t\tif (mem_cgroup_get_nr_swap_pages(memcg) \u003e 0 \u0026\u0026\n+\t\tif (mem_cgroup_get_nr_swap_pages_eligible(memcg) \u003e 0 \u0026\u0026\n \t\t    !reclaimable_anon_is_low(memcg, nid, sc))\n \t\t\treturn true;\n \t}\n@@ -2853,7 +2860,7 @@ static int get_swappiness(struct lruvec *lruvec, struct scan_control *sc)\n \t\treturn 0;\n \n \tif (!can_demote(pgdat-\u003enode_id, sc, memcg) \u0026\u0026\n-\t    mem_cgroup_get_nr_swap_pages(memcg) \u003c MIN_LRU_BATCH)\n+\t    mem_cgroup_get_nr_swap_pages_eligible(memcg) \u003c MIN_LRU_BATCH)\n \t\treturn 0;\n \n \treturn swappiness;\n@@ -5926,6 +5933,7 @@ static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,\n \t\t.reclaim_idx = MAX_NR_ZONES - 1,\n \t\t.gfp_mask = GFP_KERNEL,\n \t\t.proactive = true,\n+\t\t.allow_offload_swap = true,\n \t};\n \n \tbuf = kvmalloc(len + 1, GFP_KERNEL);\n@@ -5937,7 +5945,7 @@ static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,\n \t\treturn -EFAULT;\n \t}\n \n-\tset_task_reclaim_state(current, \u0026sc.reclaim_state);\n+\tset_task_reclaim_state(current, \u0026sc);\n \tflags = memalloc_noreclaim_save();\n \tblk_start_plug(\u0026plug);\n \tif (!set_mm_walk(NULL, true)) {\n@@ -6941,7 +6949,7 @@ unsigned long try_to_free_pages(struct zonelist *zonelist, int order,\n \tif (throttle_direct_reclaim(sc.gfp_mask, zonelist, nodemask))\n \t\treturn 1;\n \n-\tset_task_reclaim_state(current, \u0026sc.reclaim_state);\n+\tset_task_reclaim_state(current, \u0026sc);\n \ttrace_mm_vmscan_direct_reclaim_begin(sc.gfp_mask, order, NULL);\n \n \tnr_reclaimed = do_try_to_free_pages(zonelist, \u0026sc);\n@@ -6974,6 +6982,8 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,\n \t\t.may_unmap = 1,\n \t\t.may_swap = !!(reclaim_options \u0026 MEMCG_RECLAIM_MAY_SWAP),\n \t\t.proactive = !!(reclaim_options \u0026 MEMCG_RECLAIM_PROACTIVE),\n+\t\t.allow_offload_swap =\n+\t\t\t!!(reclaim_options \u0026 MEMCG_RECLAIM_PROACTIVE),\n \t};\n \t/*\n \t * Traverse the ZONELIST_FALLBACK zonelist of the current node to put\n@@ -6982,7 +6992,7 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,\n \t */\n \tstruct zonelist *zonelist = node_zonelist(numa_node_id(), sc.gfp_mask);\n \n-\tset_task_reclaim_state(current, \u0026sc.reclaim_state);\n+\tset_task_reclaim_state(current, \u0026sc);\n \ttrace_mm_vmscan_memcg_reclaim_begin(sc.gfp_mask, 0, memcg);\n \tnoreclaim_flag = memalloc_noreclaim_save();\n \n@@ -7272,7 +7282,7 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx)\n \n \ttrace_mm_vmscan_balance_pgdat_begin(pgdat-\u003enode_id, order,\n \t\t\t\t\t    highest_zoneidx);\n-\tset_task_reclaim_state(current, \u0026sc.reclaim_state);\n+\tset_task_reclaim_state(current, \u0026sc);\n \tpsi_memstall_enter(\u0026pflags);\n \t__fs_reclaim_acquire(_THIS_IP_);\n \n@@ -7769,7 +7779,7 @@ unsigned long shrink_all_memory(unsigned long nr_to_reclaim)\n \n \tfs_reclaim_acquire(sc.gfp_mask);\n \tnoreclaim_flag = memalloc_noreclaim_save();\n-\tset_task_reclaim_state(current, \u0026sc.reclaim_state);\n+\tset_task_reclaim_state(current, \u0026sc);\n \n \tnr_reclaimed = do_try_to_free_pages(zonelist, \u0026sc);\n \n@@ -7950,7 +7960,7 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,\n \t * We need to be able to allocate from the reserves for RECLAIM_UNMAP\n \t */\n \tnoreclaim_flag = memalloc_noreclaim_save();\n-\tset_task_reclaim_state(p, \u0026sc-\u003ereclaim_state);\n+\tset_task_reclaim_state(p, sc);\n \n \tdo {\n \t\tshrink_node(pgdat, sc);\n@@ -8136,6 +8146,7 @@ int user_proactive_reclaim(char *buf,\n \t\t\t\t.may_unmap = 1,\n \t\t\t\t.may_swap = 1,\n \t\t\t\t.proactive = 1,\n+\t\t\t\t.allow_offload_swap = 1,\n \t\t\t};\n \n \t\t\tif (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED,\ndiff --git a/mm/vmstat.c b/mm/vmstat.c\nindex a3e809c57f295..b924c715886e9 100644\n--- a/mm/vmstat.c\n+++ b/mm/vmstat.c\n@@ -1331,6 +1331,7 @@ const char * const vmstat_text[] = {\n \t[I(PGPGOUT)]\t\t\t\t= \"pgpgout\",\n \t[I(PSWPIN)]\t\t\t\t= \"pswpin\",\n \t[I(PSWPOUT)]\t\t\t\t= \"pswpout\",\n+\t[I(SWPOUT_OFFLOAD_REFUSED)]\t\t= \"swpout_offload_refused\",\n \n #define OFF (NR_VM_ZONE_STAT_ITEMS + NR_VM_NUMA_EVENT_ITEMS + \\\n \t     NR_VM_NODE_STAT_ITEMS + NR_VM_STAT_ITEMS)\ndiff --git a/tools/testing/selftests/zram/.gitignore b/tools/testing/selftests/zram/.gitignore\nindex 088cd9bad87a7..74b0217c60f04 100644\n--- a/tools/testing/selftests/zram/.gitignore\n+++ b/tools/testing/selftests/zram/.gitignore\n@@ -1,2 +1,4 @@\n # SPDX-License-Identifier: GPL-2.0-only\n err.log\n+swap_offload\n+workingset_offload\ndiff --git a/tools/testing/selftests/zram/Makefile b/tools/testing/selftests/zram/Makefile\nindex 7f78eb1b59cb4..b83e6a1563b20 100644\n--- a/tools/testing/selftests/zram/Makefile\n+++ b/tools/testing/selftests/zram/Makefile\n@@ -1,9 +1,9 @@\n # SPDX-License-Identifier: GPL-2.0\n all:\n \n-TEST_PROGS := zram.sh\n+TEST_GEN_FILES := swap_offload workingset_offload\n+TEST_PROGS := zram.sh zram03.sh zram04.sh zram05.sh\n TEST_FILES := zram01.sh zram02.sh zram_lib.sh\n EXTRA_CLEAN := err.log\n \n include ../lib.mk\n-\ndiff --git a/tools/testing/selftests/zram/config b/tools/testing/selftests/zram/config\nindex e0cc47e2c7e22..018429d8f5ca8 100644\n--- a/tools/testing/selftests/zram/config\n+++ b/tools/testing/selftests/zram/config\n@@ -1,2 +1,10 @@\n+CONFIG_CGROUPS=y\n+CONFIG_BLK_DEV_DM=y\n+CONFIG_DM_DELAY=y\n+CONFIG_MEMCG=y\n+CONFIG_SWAP=y\n+CONFIG_TRANSPARENT_HUGEPAGE=y\n+CONFIG_VM_EVENT_COUNTERS=y\n CONFIG_ZSMALLOC=y\n-CONFIG_ZRAM=m\n+CONFIG_ZRAM=y\n+CONFIG_ZSWAP=y\ndiff --git a/tools/testing/selftests/zram/swap_offload.c b/tools/testing/selftests/zram/swap_offload.c\nnew file mode 100644\nindex 0000000000000..1850b3b4809f0\n--- /dev/null\n+++ b/tools/testing/selftests/zram/swap_offload.c\n@@ -0,0 +1,402 @@\n+// SPDX-License-Identifier: GPL-2.0\n+#define _GNU_SOURCE\n+\n+#include \u003cerrno.h\u003e\n+#include \u003cfcntl.h\u003e\n+#include \u003climits.h\u003e\n+#include \u003csched.h\u003e\n+#include \u003csignal.h\u003e\n+#include \u003cstdio.h\u003e\n+#include \u003cstdlib.h\u003e\n+#include \u003cstring.h\u003e\n+#include \u003csys/mman.h\u003e\n+#include \u003csys/syscall.h\u003e\n+#include \u003cunistd.h\u003e\n+\n+#define SWAP_FLAG_PREFER\t0x8000\n+#define SWAP_FLAG_DISCARD\t0x10000\n+#define SWAP_FLAG_DISCARD_ONCE\t0x20000\n+#define SWAP_FLAG_DISCARD_PAGES\t0x40000\n+#define SWAP_FLAG_OFFLOAD_ONLY\t0x80000\n+#define KSFT_SKIP\t4\n+\n+static int activate(const char *path, int priority, int discard_flags)\n+{\n+\tint flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |\n+\t\tpriority | discard_flags;\n+\n+\tif (syscall(SYS_swapon, path, flags)) {\n+\t\tperror(\"swapon\");\n+\t\treturn 1;\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int reject_page_discard(const char *path, int priority)\n+{\n+\tint flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |\n+\t\tSWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_PAGES | priority;\n+\tint ret;\n+\n+\terrno = 0;\n+\tret = syscall(SYS_swapon, path, flags);\n+\tif (ret == -1 \u0026\u0026 errno == EINVAL)\n+\t\treturn 0;\n+\tif (!ret) {\n+\t\tsyscall(SYS_swapoff, path);\n+\t\tfprintf(stderr, \"offload-only page discard was accepted\\n\");\n+\t} else {\n+\t\tfprintf(stderr, \"swapon returned unexpected error: %s\\n\",\n+\t\t\tstrerror(errno));\n+\t}\n+\treturn 1;\n+}\n+\n+static int accept_discard_once_pages(const char *path, int priority)\n+{\n+\tint flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |\n+\t\tSWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE |\n+\t\tSWAP_FLAG_DISCARD_PAGES | priority;\n+\n+\tif (syscall(SYS_swapon, path, flags)) {\n+\t\tperror(\"swapon discard-once+discard-pages\");\n+\t\treturn 1;\n+\t}\n+\tif (syscall(SYS_swapoff, path)) {\n+\t\tperror(\"swapoff discard-once+discard-pages\");\n+\t\treturn 1;\n+\t}\n+\treturn 0;\n+}\n+\n+static int join_cgroup(const char *procs)\n+{\n+\tchar pid[32];\n+\tint fd, len;\n+\n+\tfd = open(procs, O_WRONLY);\n+\tif (fd \u003c 0) {\n+\t\tperror(\"open cgroup.procs\");\n+\t\treturn 1;\n+\t}\n+\n+\tlen = snprintf(pid, sizeof(pid), \"%d\\n\", getpid());\n+\tif (write(fd, pid, len) != len) {\n+\t\tperror(\"write cgroup.procs\");\n+\t\tclose(fd);\n+\t\treturn 1;\n+\t}\n+\tclose(fd);\n+\treturn 0;\n+}\n+\n+static int pin_to_one_cpu(const char *pid_arg)\n+{\n+\tcpu_set_t allowed, selected;\n+\tchar *end;\n+\tlong pid;\n+\tint cpu;\n+\n+\terrno = 0;\n+\tpid = strtol(pid_arg, \u0026end, 10);\n+\tif (errno || *end || pid \u003c= 0) {\n+\t\tfprintf(stderr, \"invalid pid: %s\\n\", pid_arg);\n+\t\treturn 1;\n+\t}\n+\n+\tif (sched_getaffinity(pid, sizeof(allowed), \u0026allowed)) {\n+\t\tperror(\"sched_getaffinity\");\n+\t\treturn 1;\n+\t}\n+\tfor (cpu = 0; cpu \u003c CPU_SETSIZE; cpu++)\n+\t\tif (CPU_ISSET(cpu, \u0026allowed))\n+\t\t\tbreak;\n+\tif (cpu == CPU_SETSIZE) {\n+\t\tfprintf(stderr, \"pid %ld has no allowed CPU\\n\", pid);\n+\t\treturn 1;\n+\t}\n+\n+\tCPU_ZERO(\u0026selected);\n+\tCPU_SET(cpu, \u0026selected);\n+\tif (sched_setaffinity(pid, sizeof(selected), \u0026selected)) {\n+\t\tperror(\"sched_setaffinity\");\n+\t\treturn 1;\n+\t}\n+\treturn 0;\n+}\n+\n+static int allocate(const char *size_arg, const char *procs,\n+\t\t    const char *ready, const char *verified)\n+{\n+\tunsigned char *memory;\n+\tunsigned long size;\n+\tunsigned long page_size;\n+\tchar *end;\n+\tsigset_t signals;\n+\tint signal;\n+\tint fd;\n+\n+\terrno = 0;\n+\tsize = strtoul(size_arg, \u0026end, 0);\n+\tif (errno || *end || !size) {\n+\t\tfprintf(stderr, \"invalid allocation size: %s\\n\", size_arg);\n+\t\treturn 1;\n+\t}\n+\tif (join_cgroup(procs))\n+\t\treturn 1;\n+\n+\tmemory = mmap(NULL, size, PROT_READ | PROT_WRITE,\n+\t\t      MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);\n+\tif (memory == MAP_FAILED) {\n+\t\tperror(\"mmap\");\n+\t\treturn 1;\n+\t}\n+\n+\tpage_size = getpagesize();\n+\tfor (unsigned long i = 0; i \u003c size; i += page_size)\n+\t\tmemset(memory + i, i / page_size % 251 + 1,\n+\t\t       page_size \u003c size - i ? page_size : size - i);\n+\tif (mprotect(memory, size, PROT_READ)) {\n+\t\tperror(\"mprotect\");\n+\t\treturn 1;\n+\t}\n+\n+\tsigemptyset(\u0026signals);\n+\tsigaddset(\u0026signals, SIGUSR1);\n+\tif (sigprocmask(SIG_BLOCK, \u0026signals, NULL)) {\n+\t\tperror(\"sigprocmask\");\n+\t\treturn 1;\n+\t}\n+\n+\tfd = open(ready, O_WRONLY | O_CREAT | O_EXCL, 0600);\n+\tif (fd \u003c 0) {\n+\t\tperror(\"create ready file\");\n+\t\treturn 1;\n+\t}\n+\tclose(fd);\n+\n+\terrno = sigwait(\u0026signals, \u0026signal);\n+\tif (errno) {\n+\t\tperror(\"sigwait\");\n+\t\treturn 1;\n+\t}\n+\tfor (unsigned long i = 0; i \u003c size; i++) {\n+\t\tunsigned char expected = i / page_size % 251 + 1;\n+\n+\t\tif (memory[i] != expected) {\n+\t\t\tfprintf(stderr,\n+\t\t\t\t\"data mismatch at %lu: got %u, expected %u\\n\",\n+\t\t\t\ti, memory[i], expected);\n+\t\t\treturn 1;\n+\t\t}\n+\t}\n+\n+\tfd = open(verified, O_WRONLY | O_CREAT | O_EXCL, 0600);\n+\tif (fd \u003c 0) {\n+\t\tperror(\"create verified file\");\n+\t\treturn 1;\n+\t}\n+\tclose(fd);\n+\n+\tfor (;;)\n+\t\tpause();\n+}\n+\n+static int create_marker(const char *path, int write_pid)\n+{\n+\tchar *temporary;\n+\tint fd, ret = 1;\n+\n+\tif (asprintf(\u0026temporary, \"%s.XXXXXX\", path) \u003c 0) {\n+\t\tperror(\"asprintf marker path\");\n+\t\treturn 1;\n+\t}\n+\tfd = mkstemp(temporary);\n+\tif (fd \u003c 0) {\n+\t\tperror(path);\n+\t\tgoto out_free;\n+\t}\n+\tif (write_pid \u0026\u0026 dprintf(fd, \"%d\\n\", getpid()) \u003c 0) {\n+\t\tperror(\"write marker\");\n+\t\tclose(fd);\n+\t\tgoto out_unlink;\n+\t}\n+\tif (close(fd)) {\n+\t\tperror(\"close marker\");\n+\t\tgoto out_unlink;\n+\t}\n+\t/* Publish only after the payload is complete for the shell reader. */\n+\tif (rename(temporary, path)) {\n+\t\tperror(\"publish marker\");\n+\t\tgoto out_unlink;\n+\t}\n+\tret = 0;\n+out_unlink:\n+\tunlink(temporary);\n+out_free:\n+\tfree(temporary);\n+\treturn ret;\n+}\n+\n+static unsigned char retained_byte(unsigned long offset, long page_size,\n+\t\t\t\t   int touched)\n+{\n+\tunsigned char value = offset / page_size % 251 + 1;\n+\n+\tif (touched \u0026\u0026 !offset)\n+\t\tvalue ^= 0x5a;\n+\treturn value;\n+}\n+\n+static int verify_retained(unsigned char *memory, unsigned long size,\n+\t\t\t   long page_size, int full)\n+{\n+\tunsigned long limit = full ? size : 1;\n+\n+\tfor (unsigned long offset = 0; offset \u003c limit; offset++) {\n+\t\tunsigned char expected = retained_byte(offset, page_size, 1);\n+\n+\t\tif (memory[offset] != expected) {\n+\t\t\tfprintf(stderr,\n+\t\t\t\t\"retained data mismatch at %lu: got %u, expected %u\\n\",\n+\t\t\t\toffset, memory[offset], expected);\n+\t\t\treturn 1;\n+\t\t}\n+\t}\n+\treturn 0;\n+}\n+\n+static int retained(const char *size_arg, const char *procs,\n+\t\t    const char *ready, const char *touched,\n+\t\t    const char *verified)\n+{\n+\tunsigned long allocation_size, mapping_start, size;\n+\tunsigned char *mapping, *memory;\n+\tchar *end;\n+\tsigset_t signals;\n+\tlong page_size;\n+\tint signal;\n+\n+\terrno = 0;\n+\tsize = strtoul(size_arg, \u0026end, 0);\n+\tif (errno || *end || !size || (size \u0026 (size - 1))) {\n+\t\tfprintf(stderr, \"allocation size must be a power of two\\n\");\n+\t\treturn 1;\n+\t}\n+\tif (join_cgroup(procs))\n+\t\treturn 1;\n+\n+\tpage_size = sysconf(_SC_PAGESIZE);\n+\tif (page_size \u003c= 0) {\n+\t\tperror(\"sysconf _SC_PAGESIZE\");\n+\t\treturn 1;\n+\t}\n+\tif (size \u003c (unsigned long)page_size || size % page_size) {\n+\t\tfprintf(stderr, \"allocation size must contain whole pages\\n\");\n+\t\treturn 1;\n+\t}\n+\tif (size \u003e ULONG_MAX / 2) {\n+\t\tfprintf(stderr, \"allocation size is too large\\n\");\n+\t\treturn 1;\n+\t}\n+\tallocation_size = size * 2;\n+\tmapping = mmap(NULL, allocation_size, PROT_READ | PROT_WRITE,\n+\t\t       MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);\n+\tif (mapping == MAP_FAILED) {\n+\t\tperror(\"mmap\");\n+\t\treturn 1;\n+\t}\n+\tmapping_start = (unsigned long)mapping;\n+\tmemory = (unsigned char *)((mapping_start + size - 1) \u0026 ~(size - 1));\n+\tif ((unsigned long)memory != mapping_start)\n+\t\tmunmap((void *)mapping_start,\n+\t\t       (unsigned long)memory - mapping_start);\n+\tmunmap(memory + size, allocation_size - size -\n+\t       ((unsigned long)memory - mapping_start));\n+\tif (madvise(memory, size, MADV_HUGEPAGE)) {\n+\t\tperror(\"madvise MADV_HUGEPAGE\");\n+\t\treturn 1;\n+\t}\n+\tfor (unsigned long offset = 0; offset \u003c size; offset += page_size)\n+\t\tmemset(memory + offset, retained_byte(offset, page_size, 0),\n+\t\t       page_size);\n+#ifdef MADV_COLLAPSE\n+\tif (madvise(memory, size, MADV_COLLAPSE)) {\n+\t\tint error = errno;\n+\n+\t\tperror(\"madvise MADV_COLLAPSE\");\n+\t\tif (error == EAGAIN || error == EINVAL || error == ENOMEM)\n+\t\t\treturn KSFT_SKIP;\n+\t\treturn 1;\n+\t}\n+#else\n+\tfprintf(stderr, \"MADV_COLLAPSE is unavailable\\n\");\n+\treturn KSFT_SKIP;\n+#endif\n+\n+\tsigemptyset(\u0026signals);\n+\tsigaddset(\u0026signals, SIGUSR1);\n+\tsigaddset(\u0026signals, SIGUSR2);\n+\tsigaddset(\u0026signals, SIGALRM);\n+\tif (sigprocmask(SIG_BLOCK, \u0026signals, NULL)) {\n+\t\tperror(\"sigprocmask\");\n+\t\treturn 1;\n+\t}\n+\tif (create_marker(ready, 1))\n+\t\treturn 1;\n+\n+\tfor (;;) {\n+\t\terrno = sigwait(\u0026signals, \u0026signal);\n+\t\tif (errno) {\n+\t\t\tperror(\"sigwait\");\n+\t\t\treturn 1;\n+\t\t}\n+\t\tif (signal == SIGUSR1) {\n+\t\t\tunsigned char value;\n+\n+\t\t\tif (mprotect(memory, page_size, PROT_READ)) {\n+\t\t\t\tperror(\"mprotect read\");\n+\t\t\t\treturn 1;\n+\t\t\t}\n+\t\t\tvalue = memory[0];\n+\t\t\tif (mprotect(memory, page_size, PROT_READ | PROT_WRITE)) {\n+\t\t\t\tperror(\"mprotect write\");\n+\t\t\t\treturn 1;\n+\t\t\t}\n+\t\t\tmemory[0] = value ^ 0x5a;\n+\t\t\tif (create_marker(touched, 0))\n+\t\t\t\treturn 1;\n+\t\t} else {\n+\t\t\tif (verify_retained(memory, size, page_size,\n+\t\t\t\t\t    signal == SIGALRM))\n+\t\t\t\treturn 1;\n+\t\t\tif (create_marker(verified, 0))\n+\t\t\t\treturn 1;\n+\t\t}\n+\t}\n+}\n+\n+int main(int argc, char **argv)\n+{\n+\tif (argc == 4 \u0026\u0026 !strcmp(argv[1], \"activate\"))\n+\t\treturn activate(argv[2], atoi(argv[3]),\n+\t\t\t\tSWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE);\n+\tif (argc == 4 \u0026\u0026 !strcmp(argv[1], \"activate-no-discard\"))\n+\t\treturn activate(argv[2], atoi(argv[3]), 0);\n+\tif (argc == 4 \u0026\u0026 !strcmp(argv[1], \"reject-page-discard\"))\n+\t\treturn reject_page_discard(argv[2], atoi(argv[3]));\n+\tif (argc == 4 \u0026\u0026 !strcmp(argv[1], \"accept-discard-once-pages\"))\n+\t\treturn accept_discard_once_pages(argv[2], atoi(argv[3]));\n+\tif (argc == 3 \u0026\u0026 !strcmp(argv[1], \"pin\"))\n+\t\treturn pin_to_one_cpu(argv[2]);\n+\tif (argc == 6 \u0026\u0026 !strcmp(argv[1], \"allocate\"))\n+\t\treturn allocate(argv[2], argv[3], argv[4], argv[5]);\n+\tif (argc == 7 \u0026\u0026 !strcmp(argv[1], \"retained\"))\n+\t\treturn retained(argv[2], argv[3], argv[4], argv[5], argv[6]);\n+\n+\tfprintf(stderr,\n+\t\t\"usage: %s activate|activate-no-discard DEVICE PRIORITY | reject-page-discard DEVICE PRIORITY | accept-discard-once-pages DEVICE PRIORITY | pin PID | allocate BYTES CGROUP.PROCS READY VERIFIED | retained BYTES CGROUP.PROCS READY TOUCHED VERIFIED\\n\",\n+\t\targv[0]);\n+\treturn 1;\n+}\ndiff --git a/tools/testing/selftests/zram/workingset_offload.c b/tools/testing/selftests/zram/workingset_offload.c\nnew file mode 100644\nindex 0000000000000..6d41badb8d0be\n--- /dev/null\n+++ b/tools/testing/selftests/zram/workingset_offload.c\n@@ -0,0 +1,206 @@\n+// SPDX-License-Identifier: GPL-2.0\n+#define _GNU_SOURCE\n+\n+#include \u003cerrno.h\u003e\n+#include \u003cfcntl.h\u003e\n+#include \u003csignal.h\u003e\n+#include \u003cstdbool.h\u003e\n+#include \u003cstdio.h\u003e\n+#include \u003cstdlib.h\u003e\n+#include \u003cstring.h\u003e\n+#include \u003csys/mman.h\u003e\n+#include \u003csys/stat.h\u003e\n+#include \u003cunistd.h\u003e\n+\n+#define ANON_SIZE (64UL \u003c\u003c 20)\n+#define TARGET_SIZE (16UL \u003c\u003c 20)\n+#define FILLER_SIZE (32UL \u003c\u003c 20)\n+\n+static int join_cgroup(const char *path)\n+{\n+\tchar pid[32];\n+\tint fd, len;\n+\n+\tfd = open(path, O_WRONLY);\n+\tif (fd \u003c 0)\n+\t\treturn -1;\n+\tlen = snprintf(pid, sizeof(pid), \"%d\\n\", getpid());\n+\tif (write(fd, pid, len) != len) {\n+\t\tclose(fd);\n+\t\treturn -1;\n+\t}\n+\treturn close(fd);\n+}\n+\n+static int create_file(const char *path, size_t size)\n+{\n+\tunsigned char page[4096];\n+\tsize_t offset;\n+\tint fd;\n+\n+\tfd = open(path, O_CREAT | O_TRUNC | O_RDWR, 0600);\n+\tif (fd \u003c 0)\n+\t\treturn -1;\n+\tmemset(page, 0xa5, sizeof(page));\n+\tfor (offset = 0; offset \u003c size; offset += sizeof(page)) {\n+\t\tif (pwrite(fd, page, sizeof(page), offset) != sizeof(page)) {\n+\t\t\tclose(fd);\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\tif (fsync(fd) || posix_fadvise(fd, 0, size, POSIX_FADV_DONTNEED)) {\n+\t\tclose(fd);\n+\t\treturn -1;\n+\t}\n+\treturn fd;\n+}\n+\n+static int cache_file(int fd, size_t size)\n+{\n+\tunsigned char byte;\n+\tsize_t offset;\n+\n+\tif (posix_fadvise(fd, 0, size, POSIX_FADV_RANDOM))\n+\t\treturn -1;\n+\tfor (offset = 0; offset \u003c size; offset += getpagesize())\n+\t\tif (pread(fd, \u0026byte, 1, offset) != 1)\n+\t\t\treturn -1;\n+\treturn 0;\n+}\n+\n+static unsigned long memory_stat(const char *path, const char *key)\n+{\n+\tunsigned long value;\n+\tchar name[64];\n+\tFILE *file;\n+\n+\tfile = fopen(path, \"r\");\n+\tif (!file)\n+\t\treturn 0;\n+\twhile (fscanf(file, \"%63s %lu\", name, \u0026value) == 2) {\n+\t\tif (!strcmp(name, key)) {\n+\t\t\tfclose(file);\n+\t\t\treturn value;\n+\t\t}\n+\t}\n+\tfclose(file);\n+\treturn 0;\n+}\n+\n+static int touch_ready(const char *path)\n+{\n+\tint fd = open(path, O_WRONLY | O_CREAT | O_EXCL, 0600);\n+\n+\tif (fd \u003c 0)\n+\t\treturn -1;\n+\treturn close(fd);\n+}\n+\n+static bool calibration_sufficient(size_t pages, unsigned long selected,\n+\t\t\t\t   unsigned long refaulted,\n+\t\t\t\t   unsigned long anon)\n+{\n+\treturn selected \u003e= pages / 2 \u0026\u0026 refaulted \u003e= selected * 3 / 4 \u0026\u0026\n+\t       anon \u003e= TARGET_SIZE;\n+}\n+\n+int main(int argc, char **argv)\n+{\n+\tunsigned long refault_before, refault_after;\n+\tunsigned long activate_before, activate_after;\n+\tunsigned long anon, selected = 0;\n+\tunsigned char *resident, *anon_memory;\n+\tunsigned char byte, checksum = 0;\n+\tchar stat_path[4096];\n+\tsigset_t signals;\n+\tsize_t pages, i;\n+\tvoid *mapping;\n+\tint target_fd, filler_fd, signal;\n+\n+\tif (argc != 6) {\n+\t\tfprintf(stderr, \"usage: %s CGROUP.PROCS TARGET FILLER READY GO\\n\",\n+\t\t\targv[0]);\n+\t\treturn 1;\n+\t}\n+\tif (join_cgroup(argv[1])) {\n+\t\tperror(\"join cgroup\");\n+\t\treturn 1;\n+\t}\n+\n+\tanon_memory = mmap(NULL, ANON_SIZE, PROT_READ | PROT_WRITE,\n+\t\t\t   MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);\n+\tif (anon_memory == MAP_FAILED) {\n+\t\tperror(\"mmap anonymous\");\n+\t\treturn 1;\n+\t}\n+\tfor (i = 0; i \u003c ANON_SIZE; i += getpagesize())\n+\t\tanon_memory[i] = i / getpagesize() % 251 + 1;\n+\n+\ttarget_fd = create_file(argv[2], TARGET_SIZE);\n+\tfiller_fd = create_file(argv[3], FILLER_SIZE);\n+\tif (target_fd \u003c 0 || filler_fd \u003c 0) {\n+\t\tperror(\"create files\");\n+\t\treturn 1;\n+\t}\n+\tif (cache_file(target_fd, TARGET_SIZE) ||\n+\t    cache_file(filler_fd, FILLER_SIZE)) {\n+\t\tperror(\"populate page cache\");\n+\t\treturn 1;\n+\t}\n+\n+\tsigemptyset(\u0026signals);\n+\tsigaddset(\u0026signals, SIGUSR1);\n+\tif (sigprocmask(SIG_BLOCK, \u0026signals, NULL) || touch_ready(argv[4])) {\n+\t\tperror(\"prepare signal\");\n+\t\treturn 1;\n+\t}\n+\tif (sigwait(\u0026signals, \u0026signal)) {\n+\t\tperror(\"sigwait\");\n+\t\treturn 1;\n+\t}\n+\n+\tmapping = mmap(NULL, TARGET_SIZE, PROT_READ, MAP_SHARED, target_fd, 0);\n+\tif (mapping == MAP_FAILED) {\n+\t\tperror(\"mmap target\");\n+\t\treturn 1;\n+\t}\n+\tpages = TARGET_SIZE / getpagesize();\n+\tresident = calloc(pages, 1);\n+\tif (!resident || mincore(mapping, TARGET_SIZE, resident)) {\n+\t\tperror(\"mincore\");\n+\t\treturn 1;\n+\t}\n+\tmunmap(mapping, TARGET_SIZE);\n+\n+\tsnprintf(stat_path, sizeof(stat_path), \"%.*s/memory.stat\",\n+\t\t (int)(strlen(argv[1]) - strlen(\"/cgroup.procs\")), argv[1]);\n+\trefault_before = memory_stat(stat_path, \"workingset_refault_file\");\n+\tactivate_before = memory_stat(stat_path, \"workingset_activate_file\");\n+\tfor (i = 0; i \u003c pages; i++) {\n+\t\tif (resident[i] \u0026 1)\n+\t\t\tcontinue;\n+\t\tif (pread(target_fd, \u0026byte, 1, i * getpagesize()) != 1) {\n+\t\t\tperror(\"refault target\");\n+\t\t\treturn 1;\n+\t\t}\n+\t\tselected++;\n+\t}\n+\trefault_after = memory_stat(stat_path, \"workingset_refault_file\");\n+\tactivate_after = memory_stat(stat_path, \"workingset_activate_file\");\n+\tanon = memory_stat(stat_path, \"active_anon\") +\n+\t       memory_stat(stat_path, \"inactive_anon\");\n+\n+\tfor (i = 0; i \u003c ANON_SIZE; i += getpagesize())\n+\t\tchecksum ^= anon_memory[i];\n+\tprintf(\"selected=%lu refault=%lu activate=%lu anon=%lu checksum=%u\\n\",\n+\t       selected, refault_after - refault_before,\n+\t       activate_after - activate_before, anon, checksum);\n+\n+\tfree(resident);\n+\tclose(filler_fd);\n+\tclose(target_fd);\n+\tif (!calibration_sufficient(pages, selected,\n+\t\t\t\t    refault_after - refault_before, anon))\n+\t\treturn 2;\n+\treturn (activate_after - activate_before) * 100 \u003c selected * 80;\n+}\ndiff --git a/tools/testing/selftests/zram/zram03.sh b/tools/testing/selftests/zram/zram03.sh\nnew file mode 100755\nindex 0000000000000..3ccc9a135f997\n--- /dev/null\n+++ b/tools/testing/selftests/zram/zram03.sh\n@@ -0,0 +1,166 @@\n+#!/bin/sh\n+# SPDX-License-Identifier: GPL-2.0\n+# Test proactive-only swap allocation and pressure fallback.\n+\n+set -eu\n+\n+# shellcheck source=zram_lib.sh\n+. ./zram_lib.sh\n+\n+TCID=\"zram03\"\n+cg=\"/sys/fs/cgroup/zram-offload-$$\"\n+cgroup_root=\"/sys/fs/cgroup\"\n+ready=\"/tmp/zram-offload-ready-$$\"\n+verified=\"/tmp/zram-offload-verified-$$\"\n+allocator_pid=\"\"\n+offload=\"\"\n+zswap_enabled=\"\"\n+zswap_writeback=\"\"\n+\n+fail()\n+{\n+\techo \"$TCID: [FAIL] $*\" \u003e\u00262\n+\texit 1\n+}\n+\n+skip()\n+{\n+\techo \"$TCID: [SKIP] $*\" \u003e\u00262\n+\texit \"$ksft_skip\"\n+}\n+\n+cleanup()\n+{\n+\tset +e\n+\tif [ -n \"$allocator_pid\" ]; then\n+\t\tkill \"$allocator_pid\"\n+\t\twait \"$allocator_pid\"\n+\tfi\n+\trm -f \"$ready\" \"$verified\"\n+\trmdir \"$cg\"\n+\tcgroup_disable_memory_controller \"$cgroup_root\"\n+\tif [ -n \"$offload\" ] \u0026\u0026 [ \"$dev_makeswap\" -lt \"$dev_end\" ]; then\n+\t\tswapoff \"$offload\" \u003e/dev/null 2\u003e\u00261\n+\tfi\n+\tif [ \"$dev_end\" -ge \"$dev_start\" ]; then\n+\t\tzram_cleanup\n+\tfi\n+\tif [ -n \"$zswap_enabled\" ]; then\n+\t\techo \"$zswap_enabled\" \u003e /sys/module/zswap/parameters/enabled\n+\tfi\n+}\n+\n+swap_used_kb()\n+{\n+\tawk -v device=\"$1\" '$1 == device { print $4 }' /proc/swaps\n+}\n+\n+zram_orig_data_size()\n+{\n+\tawk '{ print $1 }' \"/sys/block/${1##*/}/mm_stat\"\n+}\n+\n+wait_file()\n+{\n+\tfor _ in $(seq 1 400); do\n+\t\t[ -e \"$1\" ] \u0026\u0026 return 0\n+\t\tsleep 0.05\n+\tdone\n+\treturn 1\n+}\n+\n+check_prereqs\n+# The feature marker also requires CONFIG_VM_EVENT_COUNTERS.\n+grep -q '^swpout_offload_refused ' /proc/vmstat ||\n+\tskip \"offload refusal counters are unavailable\"\n+[ -x ./swap_offload ] || skip \"swap_offload helper is unavailable\"\n+[ -e /sys/fs/cgroup/cgroup.controllers ] || skip \"cgroup v2 controllers are unavailable\"\n+grep -qw memory /sys/fs/cgroup/cgroup.controllers || skip \"memory controller is unavailable\"\n+[ -e /sys/module/zswap/parameters/enabled ] || skip \"zswap is unavailable\"\n+zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)\n+\n+# Keep every reclaim transition on one CPU.  This makes the test exercise the\n+# cached-cluster provenance mismatch rather than relying on scheduler placement.\n+./swap_offload pin \"$$\" || skip \"cannot pin test to one allowed CPU\"\n+\n+# Swap priorities are global. Put both test areas ahead of any existing swap\n+# so distro-managed zram or another high-priority area cannot absorb the test\n+# allocations and make the routing assertions fail spuriously.\n+max_prio=$(awk 'BEGIN { max = -1 } NR \u003e 1 \u0026\u0026 $5 \u003e max { max = $5 } END { print max }' /proc/swaps)\n+[ \"$max_prio\" -le 32765 ] || skip \"cannot outrank existing swap priority $max_prio\"\n+safe_prio=$((max_prio + 1))\n+offload_prio=$((max_prio + 2))\n+\n+dev_num=2\n+zram_sizes=\"67108864 67108864\"\n+trap cleanup EXIT\n+trap 'exit 129' HUP\n+trap 'exit 130' INT\n+trap 'exit 143' TERM\n+zram_load\n+echo Y \u003e /sys/module/zswap/parameters/enabled || skip \"cannot enable zswap\"\n+zram_set_disksizes\n+\n+safe=\"/dev/zram${dev_start}\"\n+offload=\"/dev/zram$((dev_start + 1))\"\n+mkswap \"$safe\" \u003e/dev/null\n+mkswap \"$offload\" \u003e/dev/null\n+swapon -p \"$safe_prio\" \"$safe\"\n+dev_makeswap=$dev_start\n+./swap_offload reject-page-discard \"$offload\" \"$offload_prio\" ||\n+\tfail \"offload-only page discard was accepted\"\n+./swap_offload accept-discard-once-pages \"$offload\" \"$offload_prio\" ||\n+\tfail \"resolved offload-only discard-once was rejected\"\n+./swap_offload activate \"$offload\" \"$offload_prio\"\n+dev_makeswap=$dev_end\n+\n+cgroup_enable_memory_controller \"$cgroup_root\" ||\n+\tskip \"cannot enable the cgroup v2 memory controller\"\n+mkdir \"$cg\" || skip \"cannot create test cgroup\"\n+[ -e \"$cg/memory.max\" ] || skip \"cgroup v2 memory controller is unavailable\"\n+echo max \u003e \"$cg/memory.swap.max\"\n+echo 1 \u003e \"$cg/memory.zswap.writeback\" ||\n+\tskip \"cannot enable zswap writeback for test cgroup\"\n+read -r zswap_writeback \u003c \"$cg/memory.zswap.writeback\"\n+[ \"$zswap_writeback\" -eq 1 ] ||\n+\tskip \"an ancestor disables zswap writeback\"\n+\n+./swap_offload allocate 67108864 \"$cg/cgroup.procs\" \"$ready\" \"$verified\" \u0026\n+allocator_pid=$!\n+wait_file \"$ready\" || fail \"allocator did not become ready\"\n+\n+echo \"16M swappiness=max\" \u003e \"$cg/memory.reclaim\" ||\n+\tfail \"proactive reclaim failed\"\n+offload_before=$(swap_used_kb \"$offload\")\n+safe_before=$(swap_used_kb \"$safe\")\n+offload_orig=$(zram_orig_data_size \"$offload\")\n+[ \"${offload_before:-0}\" -gt 0 ] || fail \"proactive reclaim missed offload area\"\n+[ \"${safe_before:-0}\" -eq 0 ] || fail \"proactive reclaim used lower-priority safe area\"\n+[ \"$((offload_orig + 1048576))\" -ge \"$((offload_before * 1024))\" ] ||\n+\tfail \"zswap deferred the offload-only backend write\"\n+\n+echo 24M \u003e \"$cg/memory.max\"\n+sleep 1\n+offload_pressure=$(swap_used_kb \"$offload\")\n+safe_pressure=$(swap_used_kb \"$safe\")\n+safe_orig=$(zram_orig_data_size \"$safe\")\n+[ \"${safe_pressure:-0}\" -gt 0 ] || fail \"pressure reclaim missed safe fallback\"\n+[ \"$offload_pressure\" -le \"$offload_before\" ] ||\n+\tfail \"pressure reclaim allocated offload-only slots\"\n+[ \"$safe_orig\" -lt \"$((safe_pressure * 1024))\" ] ||\n+\tfail \"zswap did not retain any unmarked safe-area pages\"\n+\n+echo max \u003e \"$cg/memory.max\"\n+echo \"4M swappiness=max\" \u003e \"$cg/memory.reclaim\" ||\n+\tfail \"second proactive reclaim failed\"\n+offload_after=$(swap_used_kb \"$offload\")\n+[ \"$offload_after\" -gt \"$offload_pressure\" ] ||\n+\tfail \"proactive priority was not restored after pressure reclaim\"\n+\n+swapoff \"$offload\" || fail \"swapoff could not recover offloaded pages\"\n+kill -USR1 \"$allocator_pid\" || fail \"cannot request data verification\"\n+wait_file \"$verified\" || fail \"allocator did not verify recovered data\"\n+kill -0 \"$allocator_pid\" || fail \"allocator died during swapoff\"\n+dev_makeswap=$dev_start\n+\n+echo \"$TCID: [PASS]\"\ndiff --git a/tools/testing/selftests/zram/zram04.sh b/tools/testing/selftests/zram/zram04.sh\nnew file mode 100755\nindex 0000000000000..a8f52a82e774b\n--- /dev/null\n+++ b/tools/testing/selftests/zram/zram04.sh\n@@ -0,0 +1,149 @@\n+#!/bin/sh\n+# SPDX-License-Identifier: GPL-2.0\n+# Reproduce offload-only swap leaking into the workingset refault heuristic.\n+\n+set -eu\n+\n+# shellcheck source=zram_lib.sh\n+. ./zram_lib.sh\n+\n+TCID=\"zram04\"\n+cg=\"/sys/fs/cgroup/zram-workingset-$$\"\n+cgroup_root=\"/sys/fs/cgroup\"\n+tmp=\"${TMPDIR:-/var/tmp}/zram-workingset-$$\"\n+ready=\"$tmp/ready\"\n+worker=\"\"\n+tmp_fs=\"\"\n+mglru=\"\"\n+zswap_enabled=\"\"\n+\n+fail()\n+{\n+\techo \"$TCID: [FAIL] $*\" \u003e\u00262\n+\texit 1\n+}\n+\n+skip()\n+{\n+\techo \"$TCID: [SKIP] $*\" \u003e\u00262\n+\texit \"$ksft_skip\"\n+}\n+\n+cleanup()\n+{\n+\tset +e\n+\t[ -n \"$worker\" ] \u0026\u0026 kill \"$worker\"\n+\t[ -n \"$worker\" ] \u0026\u0026 wait \"$worker\"\n+\trm -rf \"$tmp\"\n+\trmdir \"$cg\"\n+\tcgroup_disable_memory_controller \"$cgroup_root\"\n+\t[ \"$dev_end\" -ge \"$dev_start\" ] \u0026\u0026 zram_cleanup\n+\t[ -n \"$mglru\" ] \u0026\u0026 echo \"$mglru\" \u003e /sys/kernel/mm/lru_gen/enabled\n+\t[ -n \"$zswap_enabled\" ] \u0026\u0026\n+\t\techo \"$zswap_enabled\" \u003e /sys/module/zswap/parameters/enabled\n+}\n+\n+wait_helper_ready()\n+{\n+\tfor _ in $(seq 1 400); do\n+\t\t[ -e \"$ready\" ] \u0026\u0026 return 0\n+\t\tworker_state=$(awk '{ print $3 }' \"/proc/$worker/stat\" \\\n+\t\t\t2\u003e/dev/null || :)\n+\t\tif ! kill -0 \"$worker\" 2\u003e/dev/null ||\n+\t\t   [ \"$worker_state\" = Z ]; then\n+\t\t\tif wait \"$worker\"; then\n+\t\t\t\tstatus=0\n+\t\t\telse\n+\t\t\t\tstatus=$?\n+\t\t\tfi\n+\t\t\tworker=\"\"\n+\t\t\t[ \"$status\" -eq 2 ] \u0026\u0026\n+\t\t\t\tskip \"workingset calibration was insufficient\"\n+\t\t\tfail \"workingset helper exited with status $status before readiness\"\n+\t\tfi\n+\t\tsleep 0.05\n+\tdone\n+\tfail \"workingset helper timed out before readiness\"\n+}\n+\n+check_prereqs\n+# The feature marker also requires CONFIG_VM_EVENT_COUNTERS.\n+grep -q '^swpout_offload_refused ' /proc/vmstat ||\n+\tskip \"offload refusal counters are unavailable\"\n+[ -x ./swap_offload ] || skip \"swap_offload helper is unavailable\"\n+[ -x ./workingset_offload ] || skip \"workingset helper is unavailable\"\n+[ -e /sys/fs/cgroup/cgroup.controllers ] || skip \"cgroup v2 is unavailable\"\n+grep -qw memory /sys/fs/cgroup/cgroup.controllers ||\n+\tskip \"cgroup v2 memory controller is unavailable\"\n+[ \"$(awk 'END { print NR }' /proc/swaps)\" -eq 1 ] ||\n+\tskip \"test requires no pre-existing swap\"\n+tmp_fs=$(stat -f -c %T \"${TMPDIR:-/var/tmp}\") ||\n+\tskip \"cannot identify the test filesystem\"\n+[ \"$tmp_fs\" != tmpfs ] ||\n+\tskip \"test files require a disk-backed filesystem\"\n+\n+[ -e /sys/kernel/mm/lru_gen/enabled ] \u0026\u0026\n+\tmglru=$(cat /sys/kernel/mm/lru_gen/enabled)\n+[ -e /sys/module/zswap/parameters/enabled ] \u0026\u0026\n+\tzswap_enabled=$(cat /sys/module/zswap/parameters/enabled)\n+trap cleanup EXIT\n+trap 'exit 129' HUP\n+trap 'exit 130' INT\n+trap 'exit 143' TERM\n+[ -n \"$mglru\" ] \u0026\u0026 echo 0 \u003e /sys/kernel/mm/lru_gen/enabled\n+[ -n \"$zswap_enabled\" ] \u0026\u0026 echo N \u003e /sys/module/zswap/parameters/enabled\n+\n+dev_num=1\n+zram_sizes=\"134217728\"\n+zram_load\n+zram_set_disksizes\n+offload=\"/dev/zram${dev_start}\"\n+mkswap \"$offload\" \u003e/dev/null\n+./swap_offload activate \"$offload\" 1\n+dev_makeswap=$dev_end\n+\n+cgroup_enable_memory_controller \"$cgroup_root\" ||\n+\tskip \"cannot enable the cgroup v2 memory controller\"\n+mkdir \"$cg\" || skip \"cannot create test cgroup\"\n+[ -e \"$cg/memory.max\" ] || skip \"cgroup v2 memory controller is unavailable\"\n+echo max \u003e \"$cg/memory.max\"\n+echo max \u003e \"$cg/memory.swap.max\"\n+mkdir \"$tmp\"\n+\n+./workingset_offload \"$cg/cgroup.procs\" \"$tmp/target\" \"$tmp/filler\" \\\n+\t\"$ready\" unused \u0026\n+worker=$!\n+wait_helper_ready\n+\n+# Shadow retention uses local LRU/slab statistics updated by memcg flushing.\n+# Check that the 64 MiB anonymous and 48 MiB file setup is visible before\n+# evicting file pages.\n+stats_ready=0\n+for _ in $(seq 1 50); do\n+\tif awk '\n+\t\t$1 == \"active_anon\" || $1 == \"inactive_anon\" { anon += $2 }\n+\t\t$1 == \"file\" { file = $2 }\n+\t\tEND { exit !(anon \u003e= 67108864 \u0026\u0026 file \u003e= 50331648) }\n+\t' \"$cg/memory.stat\"; then\n+\t\tstats_ready=1\n+\t\tbreak\n+\tfi\n+\tsleep 0.1\n+done\n+[ \"$stats_ready\" -eq 1 ] ||\n+\tskip \"initial working-set statistics did not become visible\"\n+\n+echo \"48M swappiness=0\" \u003e \"$cg/memory.reclaim\" ||\n+\tskip \"file-only proactive reclaim failed\"\n+kill -USR1 \"$worker\"\n+if wait \"$worker\"; then\n+\tworker=\"\"\n+\techo \"$TCID: [PASS]\"\n+\texit 0\n+else\n+\tstatus=$?\n+fi\n+worker=\"\"\n+[ \"$status\" -eq 2 ] \u0026\u0026 skip \"workingset calibration was insufficient\"\n+echo \"$TCID: [FAIL] refaulted file pages were not activated\" \u003e\u00262\n+exit 1\ndiff --git a/tools/testing/selftests/zram/zram05.sh b/tools/testing/selftests/zram/zram05.sh\nnew file mode 100755\nindex 0000000000000..e4bbf6bfa0e92\n--- /dev/null\n+++ b/tools/testing/selftests/zram/zram05.sh\n@@ -0,0 +1,360 @@\n+#!/bin/sh\n+# SPDX-License-Identifier: GPL-2.0\n+# Test retained offload-only entries under ordinary and proactive reclaim.\n+\n+set -eu\n+\n+# shellcheck source=zram_lib.sh\n+. ./zram_lib.sh\n+\n+TCID=\"zram05\"\n+cg=\"/sys/fs/cgroup/zram-retained-$$\"\n+cgroup_root=\"/sys/fs/cgroup\"\n+tmp=\"${TMPDIR:-/var/tmp}/zram-retained-$$\"\n+ready=\"$tmp/ready\"\n+touched=\"$tmp/touched\"\n+verified=\"$tmp/verified\"\n+holder_pid=\"\"\n+safe=\"\"\n+safe_active=0\n+offload_backing=\"\"\n+offload=\"\"\n+offload_active=0\n+dm_name=\"zram-retained-$$\"\n+dm_active=0\n+dm_node_created=0\n+zswap_enabled=\"\"\n+thp_size=0\n+thp_kib=0\n+thp_sectors=0\n+page_kib=0\n+\n+fail()\n+{\n+\techo \"$TCID: [FAIL] $*\" \u003e\u00262\n+\texit 1\n+}\n+\n+skip()\n+{\n+\techo \"$TCID: [SKIP] $*\" \u003e\u00262\n+\texit \"$ksft_skip\"\n+}\n+\n+cleanup()\n+{\n+\tset +e\n+\tif [ -n \"$holder_pid\" ]; then\n+\t\tkill \"$holder_pid\"\n+\t\twait \"$holder_pid\"\n+\tfi\n+\tif [ \"$offload_active\" -eq 1 ]; then\n+\t\tswapoff \"$offload\" \u003e/dev/null 2\u003e\u00261\n+\tfi\n+\tif [ \"$safe_active\" -eq 1 ]; then\n+\t\tswapoff \"$safe\" \u003e/dev/null 2\u003e\u00261\n+\tfi\n+\tif [ \"$dm_active\" -eq 1 ]; then\n+\t\tdmsetup --noudevsync --noudevrules remove \"$dm_name\" \\\n+\t\t\t\u003e/dev/null 2\u003e\u00261\n+\tfi\n+\tif [ \"$dm_node_created\" -eq 1 ]; then\n+\t\trm -f \"$offload\"\n+\tfi\n+\trm -rf \"$tmp\"\n+\trmdir \"$cg\"\n+\tcgroup_disable_memory_controller \"$cgroup_root\"\n+\tif [ \"$dev_end\" -ge \"$dev_start\" ]; then\n+\t\tzram_cleanup\n+\tfi\n+\tif [ -n \"$zswap_enabled\" ]; then\n+\t\techo \"$zswap_enabled\" \u003e /sys/module/zswap/parameters/enabled\n+\tfi\n+}\n+\n+wait_file()\n+{\n+\tfor _ in $(seq 1 400); do\n+\t\t[ -e \"$1\" ] \u0026\u0026 return 0\n+\t\tif [ -n \"$holder_pid\" ]; then\n+\t\t\tholder_state=$(awk '{ print $3 }' \\\n+\t\t\t\t\"/proc/$holder_pid/stat\" 2\u003e/dev/null || :)\n+\t\telse\n+\t\t\tholder_state=\"\"\n+\t\tfi\n+\t\tif [ -n \"$holder_pid\" ] \u0026\u0026\n+\t\t   { ! kill -0 \"$holder_pid\" 2\u003e/dev/null ||\n+\t\t     [ \"$holder_state\" = Z ]; }; then\n+\t\t\tif wait \"$holder_pid\"; then\n+\t\t\t\thelper_status=0\n+\t\t\telse\n+\t\t\t\thelper_status=$?\n+\t\t\tfi\n+\t\t\tholder_pid=\"\"\n+\t\t\treturn 2\n+\t\tfi\n+\t\tsleep 0.05\n+\tdone\n+\treturn 1\n+}\n+\n+require_helper_file()\n+{\n+\tif wait_file \"$1\"; then\n+\t\treturn 0\n+\telse\n+\t\tstatus=$?\n+\tfi\n+\tif [ \"$status\" -eq 2 ]; then\n+\t\t[ \"$helper_status\" -eq \"$ksft_skip\" ] \u0026\u0026 skip \"$2 is unavailable\"\n+\t\tfail \"retained helper exited with status $helper_status while $2\"\n+\tfi\n+\tfail \"retained helper timed out while $2\"\n+}\n+\n+written_sectors()\n+{\n+\tawk '{ print $7 }' \"/sys/block/${offload_backing##*/}/stat\"\n+}\n+\n+vmstat_value()\n+{\n+\tawk -v name=\"$1\" '$1 == name { print $2 }' /proc/vmstat\n+}\n+\n+wait_offload_quiet()\n+{\n+\t# A bio queued inside dm-delay is not yet visible in either the backing\n+\t# device statistics or the mapped device's inflight counters.\n+\tsleep 4\n+\tprevious=-1\n+\tstable=0\n+\tfor _ in $(seq 1 240); do\n+\t\tcurrent=$(written_sectors)\n+\t\tread -r reads writes \u003c \"$dm_inflight\"\n+\t\tif [ \"$current\" -eq \"$previous\" ] \u0026\u0026 \\\n+\t\t   [ \"$reads\" -eq 0 ] \u0026\u0026 [ \"$writes\" -eq 0 ]; then\n+\t\t\tstable=$((stable + 1))\n+\t\t\t[ \"$stable\" -ge 20 ] \u0026\u0026 return 0\n+\t\telse\n+\t\t\tstable=0\n+\t\tfi\n+\t\tprevious=$current\n+\t\tsleep 0.05\n+\tdone\n+\treturn 1\n+}\n+\n+check_prereqs\n+[ -x ./swap_offload ] || skip \"swap_offload helper is unavailable\"\n+command -v dmsetup \u003e/dev/null 2\u003e\u00261 || skip \"dmsetup is unavailable\"\n+command -v blockdev \u003e/dev/null 2\u003e\u00261 || skip \"blockdev is unavailable\"\n+[ -e /sys/fs/cgroup/cgroup.controllers ] ||\n+\tskip \"cgroup v2 controllers are unavailable\"\n+grep -qw memory /sys/fs/cgroup/cgroup.controllers ||\n+\tskip \"memory controller is unavailable\"\n+[ -d /sys/kernel/mm/transparent_hugepage ] ||\n+\tskip \"transparent huge pages are unavailable\"\n+[ -r /sys/kernel/mm/transparent_hugepage/hpage_pmd_size ] ||\n+\tskip \"PMD huge-page size is unavailable\"\n+grep -q '^swpout_offload_refused ' /proc/vmstat ||\n+\tskip \"offload refusal counters are unavailable\"\n+\n+thp_size=$(cat /sys/kernel/mm/transparent_hugepage/hpage_pmd_size)\n+case \"$thp_size\" in\n+\t''|*[!0-9]*) skip \"invalid PMD huge-page size: $thp_size\" ;;\n+esac\n+[ \"$thp_size\" -gt 0 ] || skip \"PMD huge-page size is zero\"\n+page_kib=$(awk '/KernelPageSize:/ { print $2; exit }' /proc/self/smaps)\n+[ \"${page_kib:-0}\" -gt 0 ] || skip \"cannot determine the base page size\"\n+page_size=$((page_kib * 1024))\n+[ \"$page_size\" -gt 0 ] || skip \"base page size is zero\"\n+[ $((thp_size % page_size)) -eq 0 ] ||\n+\tskip \"PMD huge-page size is not page aligned\"\n+thp_kib=$((thp_size / 1024))\n+thp_sectors=$((thp_size / 512))\n+expected_retained_kib=$((thp_kib - page_kib))\n+expected_refused=$((thp_size / page_size))\n+\n+mkdir \"$tmp\"\n+trap cleanup EXIT\n+trap 'exit 129' HUP\n+trap 'exit 130' INT\n+trap 'exit 143' TERM\n+dmsetup targets \u003e \"$tmp/dm-targets\" 2\u003e/dev/null ||\n+\tskip \"cannot query device-mapper targets\"\n+grep -q '^delay[[:space:]]' \"$tmp/dm-targets\" ||\n+\tskip \"device-mapper delay target is unavailable\"\n+\n+if [ -e /sys/module/zswap/parameters/enabled ]; then\n+\tzswap_enabled=$(cat /sys/module/zswap/parameters/enabled)\n+\techo N \u003e /sys/module/zswap/parameters/enabled ||\n+\t\tskip \"cannot disable zswap\"\n+fi\n+\n+# Swap priorities are global.  The ordinary zram device must be preferred to\n+# any pre-existing swap, while the delayed offload device remains first for\n+# eligible proactive reclaim.\n+max_prio=$(awk 'BEGIN { max = -1 } NR \u003e 1 \u0026\u0026 $5 \u003e max { max = $5 } END { print max }' /proc/swaps)\n+[ \"$max_prio\" -le 32765 ] ||\n+\tskip \"cannot outrank existing swap priority $max_prio\"\n+safe_prio=$((max_prio + 1))\n+offload_prio=$((max_prio + 2))\n+\n+dev_num=2\n+zram_size=$((thp_size * 4))\n+[ \"$zram_size\" -ge 67108864 ] || zram_size=67108864\n+zram_sizes=\"$zram_size $zram_size\"\n+zram_load\n+zram_set_disksizes\n+safe=\"/dev/zram${dev_start}\"\n+offload_backing=\"/dev/zram$((dev_start + 1))\"\n+\n+sectors=$(blockdev --getsz \"$offload_backing\") ||\n+\tskip \"cannot read offload backing size\"\n+[ \"$sectors\" -gt 0 ] || skip \"offload backing has zero size\"\n+dm_table=\"0 $sectors delay $offload_backing 0 0 $offload_backing 0 3000\"\n+dmsetup --noudevsync --noudevrules create \"$dm_name\" --table \"$dm_table\" ||\n+\tskip \"cannot create delayed offload device\"\n+dm_active=1\n+dmsetup --noudevsync --noudevrules info --columns --noheadings \\\n+\t--separator ' ' -o major,minor \"$dm_name\" \u003e \"$tmp/dm-devno\" ||\n+\tskip \"cannot identify delayed offload device\"\n+read -r dm_major dm_minor \u003c \"$tmp/dm-devno\"\n+offload=\"/dev/mapper/$dm_name\"\n+if [ ! -e \"$offload\" ]; then\n+\tmkdir -p /dev/mapper\n+\tmknod \"$offload\" b \"$dm_major\" \"$dm_minor\" ||\n+\t\tskip \"cannot create delayed offload device node\"\n+\tdm_node_created=1\n+fi\n+dm_inflight=\"/sys/dev/block/$dm_major:$dm_minor/inflight\"\n+[ -r \"$dm_inflight\" ] || skip \"cannot observe delayed offload I/O\"\n+\n+mkswap \"$safe\" \u003e/dev/null || fail \"cannot initialise safe swap\"\n+mkswap \"$offload\" \u003e/dev/null || fail \"cannot initialise offload swap\"\n+swapon -p \"$safe_prio\" \"$safe\" || fail \"cannot activate safe swap\"\n+dev_makeswap=$dev_start\n+safe_active=1\n+./swap_offload activate-no-discard \"$offload\" \"$offload_prio\" ||\n+\tfail \"cannot activate offload-only swap\"\n+offload_active=1\n+\n+cgroup_enable_memory_controller \"$cgroup_root\" ||\n+\tskip \"cannot enable the cgroup v2 memory controller\"\n+mkdir \"$cg\" || skip \"cannot create test cgroup\"\n+[ -e \"$cg/memory.max\" ] ||\n+\tskip \"cgroup v2 memory controller is unavailable\"\n+echo max \u003e \"$cg/memory.swap.max\"\n+\n+./swap_offload retained \"$thp_size\" \"$cg/cgroup.procs\" \"$ready\" \\\n+\t\"$touched\" \"$verified\" \u0026\n+holder_pid=$!\n+require_helper_file \"$ready\" \"creating a PMD-sized anonymous huge folio\"\n+read -r reported_pid \u003c \"$ready\"\n+[ \"$reported_pid\" -eq \"$holder_pid\" ] ||\n+\tfail \"retained helper reported the wrong pid\"\n+anon_huge_kib=$(awk '/AnonHugePages:/ { print $2 }' \\\n+\t\"/proc/$holder_pid/smaps_rollup\")\n+[ \"${anon_huge_kib:-0}\" -ge \"$thp_kib\" ] ||\n+\tskip \"anonymous huge folio was not created\"\n+echo \"$TCID: anonymous_huge_kib=$anon_huge_kib\"\n+\n+# This is the only reclaim before the retained entry is made dirty.  It puts\n+# the whole folio on offload-only swap so the later single-page fault leaves\n+# sibling swap PTEs referring to the existing slot.\n+echo \"$thp_size swappiness=max\" \u003e \"$cg/memory.reclaim\" 2\u003e/dev/null || :\n+wait_offload_quiet || fail \"offload device did not quiesce after setup\"\n+kill -0 \"$holder_pid\" || fail \"retained helper died during setup reclaim\"\n+kill -USR1 \"$holder_pid\" || fail \"cannot request the dirty-page transition\"\n+require_helper_file \"$touched\" \"faulting and dirtying the retained entry\"\n+retained_kib=$(awk '/VmSwap:/ { print $2 }' \"/proc/$holder_pid/status\")\n+[ \"${retained_kib:-0}\" -eq \"$expected_retained_kib\" ] ||\n+\tfail \"retained $retained_kib KiB, expected $expected_retained_kib KiB\"\n+echo \"$TCID: retained_kib=$retained_kib\"\n+\n+wait_offload_quiet || fail \"offload device did not quiesce before pressure\"\n+ordinary_before=$(written_sectors)\n+refused_before=$(vmstat_value swpout_offload_refused)\n+echo $((thp_size / 2)) \u003e \"$cg/memory.high\" ||\n+\tfail \"ordinary pressure reclaim failed\"\n+wait_offload_quiet || fail \"offload device did not quiesce after pressure\"\n+ordinary_after=$(written_sectors)\n+refused_after=$(vmstat_value swpout_offload_refused)\n+ordinary_writes=$((ordinary_after - ordinary_before))\n+refused_pages=$((refused_after - refused_before))\n+echo \"$TCID: ordinary_sectors=$ordinary_writes refused_pages=$refused_pages\"\n+[ \"$ordinary_writes\" -eq 0 ] ||\n+\tfail \"ordinary reclaim wrote $ordinary_writes offload sectors\"\n+[ \"$refused_pages\" -ge \"$expected_refused\" ] ||\n+\tfail \"ordinary reclaim refused $refused_pages pages, expected at least $expected_refused\"\n+kill -0 \"$holder_pid\" || fail \"retained helper died after refused write\"\n+rm -f \"$verified\"\n+kill -USR2 \"$holder_pid\" || fail \"cannot request dirty-byte verification\"\n+require_helper_file \"$verified\" \"verifying the dirty retained byte\"\n+\n+# Proactive reclaim must still be able to rewrite the retained PMD-sized slot.\n+# Repeated requests make the test insensitive to a short-lived writeback\n+# collision; the backing-sector delta still requires exactly one folio write.\n+wait_offload_quiet || fail \"offload device did not quiesce before recovery\"\n+recovery_before=$(written_sectors)\n+echo max \u003e \"$cg/memory.high\"\n+for _ in $(seq 1 8); do\n+\techo \"$thp_size swappiness=max\" \u003e \"$cg/memory.reclaim\" 2\u003e/dev/null || :\n+done\n+wait_offload_quiet || fail \"offload device did not quiesce after recovery\"\n+recovery_after=$(written_sectors)\n+recovery_writes=$((recovery_after - recovery_before))\n+echo \"$TCID: recovery_sectors=$recovery_writes\"\n+[ \"$recovery_writes\" -eq \"$thp_sectors\" ] ||\n+\tfail \"proactive recovery wrote $recovery_writes sectors, expected $thp_sectors\"\n+kill -0 \"$holder_pid\" || fail \"retained helper died during recovery\"\n+rm -f \"$verified\"\n+kill -ALRM \"$holder_pid\" || fail \"cannot request full data verification\"\n+require_helper_file \"$verified\" \"verifying all retained data\"\n+kill -0 \"$holder_pid\" || fail \"retained helper failed full data verification\"\n+\n+kill \"$holder_pid\" || fail \"cannot stop retained helper\"\n+if wait \"$holder_pid\"; then\n+\t:\n+else\n+\tstatus=$?\n+\t[ \"$status\" -eq 143 ] || fail \"retained helper exited with status $status\"\n+fi\n+holder_pid=\"\"\n+swapoff \"$offload\" || fail \"cannot deactivate offload-only swap\"\n+offload_active=0\n+swapoff \"$safe\" || fail \"cannot deactivate safe swap\"\n+safe_active=0\n+dev_makeswap=-1\n+dmsetup --noudevsync --noudevrules remove \"$dm_name\" ||\n+\tfail \"cannot remove delayed offload device\"\n+dm_active=0\n+if [ \"$dm_node_created\" -eq 1 ]; then\n+\trm -f \"$offload\" || fail \"cannot remove delayed offload device node\"\n+\tdm_node_created=0\n+fi\n+rmdir \"$cg\" || fail \"cannot remove test cgroup\"\n+cg=\"\"\n+cgroup_disable_memory_controller \"$cgroup_root\" ||\n+\tfail \"cannot restore the cgroup memory controller\"\n+for i in $(seq \"$dev_start\" \"$dev_end\"); do\n+\techo 1 \u003e \"/sys/block/zram${i}/reset\" || fail \"cannot reset zram$i\"\n+\tif [ \"$sys_control\" -eq 1 ]; then\n+\t\techo \"$i\" \u003e /sys/class/zram-control/hot_remove ||\n+\t\t\tfail \"cannot remove zram$i\"\n+\tfi\n+done\n+if [ \"$module_load\" -eq 1 ]; then\n+\trmmod zram || fail \"cannot unload zram\"\n+fi\n+dev_end=-1\n+if [ -n \"$zswap_enabled\" ]; then\n+\techo \"$zswap_enabled\" \u003e /sys/module/zswap/parameters/enabled ||\n+\t\tfail \"cannot restore zswap\"\n+\tzswap_enabled=\"\"\n+fi\n+rm -rf \"$tmp\" || fail \"cannot remove temporary files\"\n+tmp=\"\"\n+\n+echo \"$TCID: [PASS]\"\ndiff --git a/tools/testing/selftests/zram/zram_lib.sh b/tools/testing/selftests/zram/zram_lib.sh\nindex 0d44d83888f9d..f14966d3820d8 100755\n--- a/tools/testing/selftests/zram/zram_lib.sh\n+++ b/tools/testing/selftests/zram/zram_lib.sh\n@@ -17,6 +17,10 @@ kernel_version=`uname -r | cut -d'.' -f1,2`\n kernel_major=${kernel_version%.*}\n kernel_minor=${kernel_version#*.}\n \n+# Whether this test enabled the memory controller on its cgroup parent.  Tests\n+# must leave a delegation which was already present alone.\n+cgroup_memory_controller_enabled=0\n+\n trap INT\n \n check_prereqs()\n@@ -30,6 +34,30 @@ check_prereqs()\n \tfi\n }\n \n+cgroup_enable_memory_controller()\n+{\n+\tlocal cgroup_root=$1\n+\n+\tif grep -qw memory \"$cgroup_root/cgroup.subtree_control\"; then\n+\t\treturn 0\n+\tfi\n+\n+\tif ! echo +memory \u003e \"$cgroup_root/cgroup.subtree_control\"; then\n+\t\treturn 1\n+\tfi\n+\n+\tcgroup_memory_controller_enabled=1\n+}\n+\n+cgroup_disable_memory_controller()\n+{\n+\tlocal cgroup_root=$1\n+\n+\t[ \"$cgroup_memory_controller_enabled\" -eq 1 ] || return 0\n+\techo -memory \u003e \"$cgroup_root/cgroup.subtree_control\" || return 1\n+\tcgroup_memory_controller_enabled=0\n+}\n+\n kernel_gte()\n {\n \tmajor=${1%.*}\n"
}

2/1 2026/09/14 12:10 agent patch-evaluator
32m Model:
core

Args:
null

Results:
{
  "EnableConfigs": [
    "SWAP",
    "MEMCG",
    "ZSWAP",
    "VM_EVENT_COUNTERS"
  ],
  "FocusSymbols": [
    "folio_alloc_swap",
    "swap_writeout",
    "get_nr_swap_pages_eligible",
    "mem_cgroup_get_nr_swap_pages_eligible",
    "folio_swap_full",
    "user_proactive_reclaim",
    "__x64_sys_swapon"
  ],
  "Reasoning": "The patch series modifies core memory management code (mm/swapfile.c, mm/vmscan.c, mm/page_io.c, mm/memcontrol.c) to introduce offload-only swap areas via the new SWAP_FLAG_OFFLOAD_ONLY flag in swapon(). It alters swap slot allocation (both fast and slow paths), eligible swap accounting, swapcache reclaim heuristics, swap writeout refusal paths for retained entries, and zswap integration. Furthermore, it adds new runtime assertions (VM_WARN_ON_FOLIO) in folio_swap_full(). All modified paths are fully reachable in virtualized environments via swapon(), cgroup v2 memory.reclaim, and standard memory reclaim under pressure.",
  "WorthFuzzing": true
}

Instruction:
You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.

IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.

================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.

- REACHABILITY IS THE PRIMARY GATE:
  Fuzzing can only discover bugs in code that can actually execute in standard virtualized
  environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
  If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
  regardless of whether it adds assertions or complex logic.

- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
  Patch authors routinely label changes as "cleanups", "refactorings", or state
  "No functional change intended". Do NOT take these claims at face value.
  Code refactorings that rearrange logic, introduce helper functions, or alter state management
  in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
  If reachable executable code is modified or refactored, it MUST be fuzzed.

- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
  When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
  BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
  Even if the author believes the invariant always holds, fuzzing is essential to verify whether
  an unusual sequence of operations can violate it.

================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:

- Non-kernel and non-executable changes:
  * Modifications to Documentation/, comments, or spelling fixes.
  * User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
    that do not affect the compiled kernel image (vmlinux) or kernel modules.
  * Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
    that do not alter control flow or data structures.
  * Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
  * Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
    ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
  * Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
    physical PCI hot-unplug or manual sysfs driver unbinding.
  * Code paths exclusive to architectures other than the target architecture.

================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
  * Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
    BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
  * Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
  * Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
  * Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).

================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:

- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
  Do NOT list generic, ubiquitous functions called by almost every program in the corpus
  (including, but not limited to: general memory allocators and deallocators, page fault
  and trap handlers, or core synchronization primitives; this is not an exhaustive list).
  Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
  which severely dilutes fuzzing effort away from the actual changes.

- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
  List functions that specifically implement the logic being added or altered, or direct API entrypoints
  for the subsystem feature under review.

- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
  Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
  distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
  If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
  functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).

================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.


Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.

Prompt:
Target architecture: amd64

For your convenience, here is the diff of the changes:
commit b18205544615b8a7fdfb59a7d24849375e130428
Author: syz-cluster <triage@syzkaller.com>
Date:   Mon Sep 14 12:10:57 2026 +0000

    syz-cluster: applied patch under review

diff --git a/Documentation/mm/swap.rst b/Documentation/mm/swap.rst
index 78819bd4d7452..7d5bbd65e6160 100644
--- a/Documentation/mm/swap.rst
+++ b/Documentation/mm/swap.rst
@@ -3,3 +3,90 @@
 ====
 Swap
 ====
+
+Offload-only swap areas
+-----------------------
+
+``SWAP_FLAG_OFFLOAD_ONLY`` marks a swap area as a destination for explicit
+userspace proactive reclaim.  The stable qualifying interfaces are cgroup v2
+``memory.reclaim`` and ``/sys/devices/system/node/nodeX/reclaim``.  The swap
+allocator excludes such an area from kswapd, direct reclaim, and other
+pressure-driven swap allocation.  Normal swap priority ordering still applies
+among the areas eligible for the current reclaim context.
+
+The MGLRU debugfs eviction interface currently establishes the same internal
+proactive-reclaim provenance and can therefore use an offload-only area.
+Debugfs is not a stable userspace ABI, however, so that behaviour is not part
+of this interface's permanent contract.
+
+This permits a system to combine a small conventional swap area, which is
+engineered for forward progress in emergency reclaim, with a larger or more
+complex area used for ordinary cold-page offload.  For example, the latter may
+be RAM-compressed or may use a filesystem with compression, checksums, or
+redundancy.  Making every such write path safe in direct reclaim can require
+backend-specific reserves, preallocation, non-blocking allocation, and
+recursion rules.  Excluding new pressure-reclaim writes can reduce that
+requirement and its complexity.
+
+For a RAM-compressed area such as zram, unused logical slots also do not imply
+that enough physical memory remains to store their future contents.  Static
+swap priority cannot express that distinction or provide late fallback after a
+selected area's write fails.
+
+The policy is attached to an activated swap area, not to its underlying
+physical storage.  A raw swap partition and a filesystem swapfile on the same
+device are separate areas and may use different policies.  The kernel does not
+infer this policy from the block driver, filesystem, or swap priority.
+
+This is a reclaim-provenance policy, not a measurement of current memory
+headroom.  Userspace should only request proactive offload while its own
+watermark or PSI policy considers memory pressure low.
+
+DAMON reclaim and ``MADV_PAGEOUT`` do not currently establish the proactive
+reclaim context, so they cannot allocate slots from an offload-only area.
+Offload-only areas are also ineligible for hibernation image allocation.
+
+Offload-only areas bypass zswap stores.  Zswap writeback may run after the
+proactive context which selected the slot has ended, so admitting the folio to
+zswap would otherwise defer the backend write beyond that context.  The
+hierarchical cgroup v2 ``memory.zswap.writeback=0`` policy remains
+authoritative: when zswap is enabled, it also refuses direct proactive writes
+to an offload-only area.  Marking an area offload-only does not override a
+cgroup policy which disables all swapping attempts to devices.
+
+The flag controls both allocation of new swap slots and newly initiated
+non-zero backend writes.  A folio can retain its swap entry after swapin.  If
+ordinary reclaim later tries to rewrite such an offload-only entry, the VM
+redirties and activates the folio instead; proactive reclaim may retry the
+write.  Zero-filled folios may still update the in-memory swap zeromap without
+backend I/O.
+
+The flag does not prevent reads, swapoff, or writes which are already queued or
+in flight.  It therefore does not by itself provide a forward progress
+guarantee for an I/O path which allocates memory: earlier writes must still be
+able to complete, and reads must remain reclaim-safe.  Repeatedly refusing
+retained-entry writes can also reduce reclaim efficiency and lead to OOM while
+the dirty folios remain resident.
+
+Page-cluster discard is incompatible with an offload-only area because its
+work item can run after the context which freed the entries has ended.  Swapon
+therefore rejects a resolved page-cluster discard policy combined with
+``SWAP_FLAG_OFFLOAD_ONLY``.  Swapon-time discard is permitted because it
+completes synchronously during activation.  The existing discard precedence
+still applies: requesting both discard-once and discard-pages selects
+discard-once.  Discard requests which the swap area does not support remain
+ignored.
+
+Architecture-specific swap metadata preparation still runs before a retained
+write is refused, so that metadata remains coherent with the dirty resident
+folio.  This policy controls swap-backend I/O; it does not promise that core VM
+or architecture preparation performs no allocation.
+
+``/proc/vmstat`` reports ``swpout_offload_refused`` in base pages.  The counter
+advances when ordinary reclaim refuses a newly initiated write through a
+retained offload-only entry.  It can reveal repeated activation of such folios,
+but it is not a count of skipped areas during new-slot allocation.
+
+An offload-only area should therefore be configured with a reclaim-safe swap
+area as fallback.  If no eligible swap space remains, swap allocation fails
+and the existing reclaim and OOM policy applies.
diff --git a/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c b/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
index e0d1f369a1632..b58e61f15ab10 100644
--- a/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
+++ b/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
@@ -21,7 +21,7 @@
 
 static bool swap_available(void)
 {
-	return get_nr_swap_pages() > 0;
+	return get_nr_swap_pages_eligible() > 0;
 }
 
 static bool can_release_pages(struct drm_i915_gem_object *obj)
diff --git a/drivers/gpu/drm/i915/gem/selftests/huge_pages.c b/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
index 44718e7282914..73b62b0655103 100644
--- a/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
+++ b/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
@@ -1895,8 +1895,8 @@ static int igt_shrink_thp(void *arg)
 	i915_gem_context_unlock_engines(ctx);
 	/*
 	 * Nuke everything *before* we unpin the pages so we can be reasonably
-	 * sure that when later checking get_nr_swap_pages() that some random
-	 * leftover object doesn't steal the remaining swap space.
+	 * sure that when later checking get_nr_swap_pages_eligible() that some
+	 * random leftover object doesn't steal the remaining swap space.
 	 */
 	i915_gem_shrink(NULL, i915, -1UL, NULL,
 			I915_SHRINK_BOUND |
@@ -1910,7 +1910,7 @@ static int igt_shrink_thp(void *arg)
 	 * Now that the pages are *unpinned* shrinking should invoke
 	 * shmem to truncate our pages, if we have available swap.
 	 */
-	should_swap = get_nr_swap_pages() > 0;
+	should_swap = get_nr_swap_pages_eligible() > 0;
 	i915_gem_shrink(NULL, i915, -1UL, NULL,
 			I915_SHRINK_BOUND |
 			I915_SHRINK_UNBOUND |
diff --git a/drivers/gpu/drm/msm/msm_gem_shrinker.c b/drivers/gpu/drm/msm/msm_gem_shrinker.c
index 9d2788f79acee..da6a66b747b02 100644
--- a/drivers/gpu/drm/msm/msm_gem_shrinker.c
+++ b/drivers/gpu/drm/msm/msm_gem_shrinker.c
@@ -21,7 +21,7 @@ module_param(enable_eviction, bool, 0600);
 
 static bool can_swap(void)
 {
-	return enable_eviction && get_nr_swap_pages() > 0;
+	return enable_eviction && get_nr_swap_pages_eligible() > 0;
 }
 
 static bool can_block(struct shrink_control *sc)
diff --git a/drivers/gpu/drm/panthor/panthor_gem.c b/drivers/gpu/drm/panthor/panthor_gem.c
index 72908be5e144f..23488f417cf7a 100644
--- a/drivers/gpu/drm/panthor/panthor_gem.c
+++ b/drivers/gpu/drm/panthor/panthor_gem.c
@@ -1371,7 +1371,7 @@ panthor_dummy_bo_create(struct panthor_device *ptdev)
 
 static bool can_swap(void)
 {
-	return get_nr_swap_pages() > 0;
+	return get_nr_swap_pages_eligible() > 0;
 }
 
 static bool can_block(struct shrink_control *sc)
diff --git a/drivers/gpu/drm/ttm/ttm_backup.c b/drivers/gpu/drm/ttm/ttm_backup.c
index 0c2d53a13b2a6..bf9d41bdf652c 100644
--- a/drivers/gpu/drm/ttm/ttm_backup.c
+++ b/drivers/gpu/drm/ttm/ttm_backup.c
@@ -206,7 +206,7 @@ u64 ttm_backup_bytes_avail(void)
 	 * number also depends on shmem actually swapping out backed-up
 	 * shmem objects without too much buffering.
 	 */
-	return (u64)get_nr_swap_pages() << PAGE_SHIFT;
+	return (u64)get_nr_swap_pages_eligible() << PAGE_SHIFT;
 }
 EXPORT_SYMBOL_GPL(ttm_backup_bytes_avail);
 
diff --git a/drivers/gpu/drm/xe/tests/xe_bo.c b/drivers/gpu/drm/xe/tests/xe_bo.c
index 6a17e13d58cf1..14d6bb8e41c9c 100644
--- a/drivers/gpu/drm/xe/tests/xe_bo.c
+++ b/drivers/gpu/drm/xe/tests/xe_bo.c
@@ -695,7 +695,7 @@ static int shrink_test_run_device(struct xe_device *xe)
 	}
 	to_alloc = ram * 2;
 
-	ram_and_swap = ram + get_nr_swap_pages() * PAGE_SIZE;
+	ram_and_swap = ram + get_nr_swap_pages_eligible() * PAGE_SIZE;
 	if (to_alloc > ram_and_swap)
 		purgeable = to_alloc - ram_and_swap;
 	purgeable += div64_u64(purgeable, 5);
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 43155e122b5c3..19a0bb4ba6c93 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -21,10 +21,11 @@
 #define SWAP_FLAG_DISCARD	0x10000 /* enable discard for swap */
 #define SWAP_FLAG_DISCARD_ONCE	0x20000 /* discard swap area at swapon-time */
 #define SWAP_FLAG_DISCARD_PAGES 0x40000 /* discard page-clusters after use */
+#define SWAP_FLAG_OFFLOAD_ONLY  0x80000 /* only use for proactive reclaim */
 
 #define SWAP_FLAGS_VALID	(SWAP_FLAG_PRIO_MASK | SWAP_FLAG_PREFER | \
 				 SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE | \
-				 SWAP_FLAG_DISCARD_PAGES)
+				 SWAP_FLAG_DISCARD_PAGES | SWAP_FLAG_OFFLOAD_ONLY)
 /*
  * MAX_SWAPFILES defines the maximum number of swaptypes: things which can
  * be swapped to.  The swap type and the offset into that swap type are
@@ -140,12 +141,20 @@ union swap_header {
 struct reclaim_state {
 	/* pages reclaimed outside of LRU-based reclaim */
 	unsigned long reclaimed;
+	/* this reclaim context may use offload-only swap */
+	bool allow_offload_swap;
 #ifdef CONFIG_LRU_GEN
 	/* per-thread mm walk data */
 	struct lru_gen_mm_walk *mm_walk;
 #endif
 };
 
+static inline bool current_reclaim_allows_offload_swap(void)
+{
+	return current->reclaim_state &&
+	       current->reclaim_state->allow_offload_swap;
+}
+
 /*
  * mm_account_reclaimed_pages(): account reclaimed pages outside of LRU-based
  * reclaim
@@ -201,6 +210,7 @@ enum {
 	SWP_STABLE_WRITES = (1 << 11),	/* no overwrite PG_writeback pages */
 	SWP_SYNCHRONOUS_IO = (1 << 12),	/* synchronous IO is efficient */
 	SWP_HIBERNATION = (1 << 13),	/* pinned for hibernation */
+	SWP_OFFLOAD_ONLY = (1 << 14),	/* proactive-reclaim swap only */
 					/* add others here before... */
 };
 
@@ -389,6 +399,9 @@ static inline long get_nr_swap_pages(void)
 	return atomic_long_read(&nr_swap_pages);
 }
 
+long get_nr_swap_pages_eligible(void);
+bool folio_swap_full(struct folio *folio);
+
 extern void si_swapinfo(struct sysinfo *);
 extern int pin_hibernation_swap_type(dev_t device, sector_t offset);
 extern void unpin_hibernation_swap_type(int type);
@@ -443,10 +456,16 @@ static inline void put_swap_device(struct swap_info_struct *si)
 }
 
 #define get_nr_swap_pages()			0L
+#define get_nr_swap_pages_eligible()		0L
 #define total_swap_pages			0L
 #define total_swapcache_pages()			0UL
 #define vm_swap_full()				0
 
+static inline bool folio_swap_full(struct folio *folio)
+{
+	return false;
+}
+
 #define si_swapinfo(val) \
 	do { (val)->freeswap = (val)->totalswap = 0; } while (0)
 #define free_folio_and_swap_cache(folio) \
@@ -531,6 +550,7 @@ static inline void mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_p
 
 long mem_cgroup_get_folio_swap_margin(struct folio *folio);
 extern long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg);
+long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg);
 extern bool mem_cgroup_swap_full(struct folio *folio);
 #else
 static inline int mem_cgroup_try_charge_swap(struct folio *folio)
@@ -553,9 +573,14 @@ static inline long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
 	return get_nr_swap_pages();
 }
 
+static inline long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg)
+{
+	return get_nr_swap_pages_eligible();
+}
+
 static inline bool mem_cgroup_swap_full(struct folio *folio)
 {
-	return vm_swap_full();
+	return folio_swap_full(folio);
 }
 #endif
 
diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h
index 2628ccda076a0..fc8458b314a50 100644
--- a/include/linux/vm_event_item.h
+++ b/include/linux/vm_event_item.h
@@ -32,6 +32,7 @@
 	HIGHMEM_ZONE(xx) xx##_MOVABLE, DEVICE_ZONE(xx)
 
 enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT,
+		SWPOUT_OFFLOAD_REFUSED,
 		FOR_ALL_ZONES(PGALLOC)
 		FOR_ALL_ZONES(ALLOCSTALL)
 		FOR_ALL_ZONES(PGSCAN_SKIP)
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1460cba53588e..1412084d2f430 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -6002,16 +6002,28 @@ void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages)
 	rcu_read_unlock();
 }
 
-long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
+static long
+mem_cgroup_get_nr_swap_pages_with_limit(struct mem_cgroup *memcg,
+					long nr_swap_pages)
 {
-	long nr_swap_pages = get_nr_swap_pages();
-
 	if (!mem_cgroup_disabled() && !do_memsw_account())
 		nr_swap_pages = min(nr_swap_pages, page_counter_margin(&memcg->swap));
 
 	return nr_swap_pages;
 }
 
+long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
+{
+	return mem_cgroup_get_nr_swap_pages_with_limit(memcg,
+			get_nr_swap_pages());
+}
+
+long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg)
+{
+	return mem_cgroup_get_nr_swap_pages_with_limit(memcg,
+			get_nr_swap_pages_eligible());
+}
+
 /**
  * mem_cgroup_get_folio_swap_margin - get a folio's memcg swap margin
  * @folio: folio whose memcg margin is queried
@@ -6042,7 +6054,7 @@ bool mem_cgroup_swap_full(struct folio *folio)
 
 	VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
 
-	if (vm_swap_full())
+	if (folio_swap_full(folio))
 		return true;
 	if (do_memsw_account() || !folio_memcg_charged(folio))
 		return ret;
diff --git a/mm/page_io.c b/mm/page_io.c
index 1da4ff484f097..1b281a1d6df0a 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -203,6 +203,7 @@ static void swap_zeromap_folio_clear(struct folio *folio)
  */
 int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
 {
+	struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
 	int ret = 0;
 
 	if (folio_free_swap(folio))
@@ -210,7 +211,9 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
 
 	/*
 	 * Arch code may have to preserve more data than just the folio
-	 * contents, e.g. memory tags.
+	 * contents, e.g. memory tags.  Do this before refusing a retained
+	 * offload-only entry below: a later sibling swap-PTE fault can restore
+	 * swap-indexed metadata into this resident folio.
 	 */
 	ret = arch_prepare_to_swap(folio);
 	if (ret) {
@@ -228,6 +231,18 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
 		goto out_unlock;
 	}
 
+	/*
+	 * A folio can retain an existing swap entry after swapin.  Do not let
+	 * ordinary reclaim use an offload-only entry through that path.
+	 */
+	if ((READ_ONCE(sis->flags) & SWP_OFFLOAD_ONLY) &&
+	    !current_reclaim_allows_offload_swap()) {
+		count_vm_events(SWPOUT_OFFLOAD_REFUSED,
+				folio_nr_pages(folio));
+		folio_mark_dirty(folio);
+		return AOP_WRITEPAGE_ACTIVATE;
+	}
+
 	/*
 	 * Clear bits this folio occupies in the zeromap to prevent zero data
 	 * being read in from any previous zero writes that occupied the same
@@ -235,7 +250,12 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
 	 */
 	swap_zeromap_folio_clear(folio);
 
-	if (zswap_store(folio)) {
+	/*
+	 * Zswap writeback can happen much later from pressure reclaim or its
+	 * shrinker workqueue.  Do not let it defer an offload-only backend write
+	 * beyond the proactive reclaim context which admitted the swap slot.
+	 */
+	if (!(READ_ONCE(sis->flags) & SWP_OFFLOAD_ONLY) && zswap_store(folio)) {
 		count_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT);
 		goto out_unlock;
 	}
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 280dd906eb187..256617059f302 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -65,13 +65,17 @@ static void move_cluster(struct swap_info_struct *si,
  */
 static DEFINE_SPINLOCK(swap_lock);
 static unsigned int nr_swapfiles;
-atomic_long_t nr_swap_pages;
 /*
  * Some modules use swappable objects and may try to swap them out under
  * memory pressure (via the shrinker). Before doing so, they may wish to
  * check to see if any swap space is available.
+ *
+ * This remains the raw free-space counter for accounting users. Reclaim
+ * decisions subtract nr_swap_pages_offload_only below when necessary.
  */
+atomic_long_t nr_swap_pages;
 EXPORT_SYMBOL_GPL(nr_swap_pages);
+static atomic_long_t nr_swap_pages_offload_only;
 /* protected with swap_lock. reading in vm_swap_full() doesn't need lock */
 long total_swap_pages;
 #define DEF_SWAP_PRIO  -1
@@ -120,6 +124,7 @@ atomic_t nr_rotate_swap = ATOMIC_INIT(0);
 struct percpu_swap_cluster {
 	struct swap_info_struct *si[SWAP_NR_ORDERS];
 	unsigned long offset[SWAP_NR_ORDERS];
+	bool allow_offload_swap[SWAP_NR_ORDERS];
 	local_lock_t lock;
 };
 
@@ -163,6 +168,31 @@ static long swap_usage_in_pages(struct swap_info_struct *si)
 	return atomic_long_read(&si->inuse_pages) & SWAP_USAGE_COUNTER_MASK;
 }
 
+static bool swap_area_needs_reclaim(struct swap_info_struct *si)
+{
+	if (vm_swap_full())
+		return true;
+
+	/*
+	 * Free offload-only slots must not keep a full conventional area
+	 * pinned in swapcache.  Recover that area's unused cache entries
+	 * even when the raw pool is not full.  This is independent of the
+	 * current task: proactive reclaim can fill the conventional area too.
+	 */
+	return !(READ_ONCE(si->flags) & SWP_OFFLOAD_ONLY) &&
+		atomic_long_read(&nr_swap_pages_offload_only) > 0 &&
+		swap_usage_in_pages(si) == si->pages;
+}
+
+/* The caller must hold the lock on a folio in swapcache. */
+bool folio_swap_full(struct folio *folio)
+{
+	VM_WARN_ON_FOLIO(!folio_test_locked(folio), folio);
+	VM_WARN_ON_FOLIO(!folio_test_swapcache(folio), folio);
+
+	return swap_area_needs_reclaim(__swap_entry_to_info(folio->swap));
+}
+
 /* Reclaim the swap entry anyway if possible */
 #define TTRS_ANYWAY		0x1
 /*
@@ -903,7 +933,7 @@ static bool cluster_scan_range(struct swap_info_struct *si,
 		if (swp_tb_is_null(swp_tb))
 			continue;
 		if (swp_tb_is_folio(swp_tb) && !__swp_tb_get_count(swp_tb)) {
-			if (!vm_swap_full())
+			if (!swap_area_needs_reclaim(si))
 				return false;
 			*need_reclaim = true;
 			continue;
@@ -1014,6 +1044,8 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
 	if (si->flags & SWP_SOLIDSTATE) {
 		this_cpu_write(percpu_swap_cluster.offset[order], next);
 		this_cpu_write(percpu_swap_cluster.si[order], si);
+		this_cpu_write(percpu_swap_cluster.allow_offload_swap[order],
+			       current_reclaim_allows_offload_swap());
 	} else {
 		si->global_cluster->next[order] = next;
 	}
@@ -1161,7 +1193,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
 	}
 
 	/* Try reclaim full clusters if free and nonfull lists are drained */
-	if (vm_swap_full())
+	if (swap_area_needs_reclaim(si))
 		swap_reclaim_full_clusters(si, false);
 
 	if (order < PMD_ORDER) {
@@ -1315,9 +1347,11 @@ static void swap_range_alloc(struct swap_info_struct *si,
 			     unsigned int nr_entries)
 {
 	if (swap_usage_add(si, nr_entries)) {
-		if (vm_swap_full())
+		if (swap_area_needs_reclaim(si))
 			schedule_work(&si->reclaim_work);
 	}
+	if (si->flags & SWP_OFFLOAD_ONLY)
+		atomic_long_sub(nr_entries, &nr_swap_pages_offload_only);
 	atomic_long_sub(nr_entries, &nr_swap_pages);
 }
 
@@ -1346,6 +1380,8 @@ static void swap_range_free(struct swap_info_struct *si, unsigned long offset,
 	 * only after the above cleanups are done.
 	 */
 	smp_wmb();
+	if (si->flags & SWP_OFFLOAD_ONLY)
+		atomic_long_add(nr_entries, &nr_swap_pages_offload_only);
 	atomic_long_add(nr_entries, &nr_swap_pages);
 	swap_usage_sub(si, nr_entries);
 }
@@ -1366,6 +1402,31 @@ static bool get_swap_device_info(struct swap_info_struct *si)
 	return true;
 }
 
+static bool swap_area_eligible(struct swap_info_struct *si)
+{
+	if (!(READ_ONCE(si->flags) & SWP_OFFLOAD_ONLY))
+		return true;
+
+	return current_reclaim_allows_offload_swap();
+}
+
+long get_nr_swap_pages_eligible(void)
+{
+	long nr_pages;
+
+	if (current_reclaim_allows_offload_swap())
+		return get_nr_swap_pages();
+
+	/*
+	 * The reads are intentionally unpaired.  This is a capacity hint; the
+	 * allocator enforces eligibility.  Clamp a transient negative result.
+	 */
+	nr_pages = get_nr_swap_pages() -
+		   atomic_long_read(&nr_swap_pages_offload_only);
+	return max(nr_pages, 0L);
+}
+EXPORT_SYMBOL_GPL(get_nr_swap_pages_eligible);
+
 /*
  * Fast path try to get swap entries with specified order from current
  * CPU's swap entry pool (a cluster).
@@ -1385,6 +1446,20 @@ static bool swap_alloc_fast(struct folio *folio)
 	offset = this_cpu_read(percpu_swap_cluster.offset[order]);
 	if (!si || !offset || !get_swap_device_info(si))
 		return false;
+	if (!swap_area_eligible(si)) {
+		put_swap_device(si);
+		return false;
+	}
+	/*
+	 * Pressure reclaim may cache a lower-priority conventional area while
+	 * an offload-only area is ineligible.  Drop that cache on a context
+	 * change so proactive reclaim returns to the normal priority search.
+	 */
+	if (this_cpu_read(percpu_swap_cluster.allow_offload_swap[order]) !=
+	    current_reclaim_allows_offload_swap()) {
+		put_swap_device(si);
+		return false;
+	}
 
 	ci = swap_cluster_lock(si, offset);
 	if (cluster_is_usable(ci, order)) {
@@ -1407,6 +1482,9 @@ static void swap_alloc_slow(struct folio *folio)
 	spin_lock(&swap_avail_lock);
 start_over:
 	plist_for_each_entry_safe(si, next, &swap_avail_head, avail_list) {
+		if (!swap_area_eligible(si))
+			continue;
+
 		/* Rotate the device and switch to a new cluster */
 		plist_requeue(&si->avail_list, &swap_avail_head);
 		spin_unlock(&swap_avail_lock);
@@ -1450,7 +1528,8 @@ static bool swap_sync_discard(void)
 	plist_for_each_entry_safe(si, next, &swap_active_head, list) {
 		spin_unlock(&swap_lock);
 		if (get_swap_device_info(si)) {
-			if (si->flags & SWP_PAGE_DISCARD)
+			if (swap_area_eligible(si) &&
+			    (si->flags & SWP_PAGE_DISCARD))
 				ret = swap_do_scheduled_discard(si);
 			put_swap_device(si);
 		}
@@ -1738,8 +1817,8 @@ static int swap_dup_entries_cluster(struct swap_info_struct *si,
  *
  * Context: Caller needs to hold the folio lock.
  * Return: %0 on success, %-E2BIG if splitting the folio might allow swapout,
- * %-ENOSPC if no global swap space is available, or %-ENOMEM if splitting
- * would not help.
+ * %-ENOSPC if no global swap space is eligible for the caller, or %-ENOMEM
+ * if splitting would not help.
  */
 int folio_alloc_swap(struct folio *folio)
 {
@@ -1790,7 +1869,7 @@ int folio_alloc_swap(struct folio *folio)
 	return 0;
 
 failed:
-	if (get_nr_swap_pages() <= 0)
+	if (get_nr_swap_pages_eligible() <= 0)
 		return -ENOSPC;
 	if (mem_cgroup_get_folio_swap_margin(folio) <= 0)
 		return -ENOMEM;
@@ -2180,7 +2259,7 @@ swp_entry_t swap_alloc_hibernation_slot(int type)
 	struct swap_cluster_info *ci;
 	swp_entry_t entry = {0};
 
-	if (!si)
+	if (!si || (si->flags & SWP_OFFLOAD_ONLY))
 		goto fail;
 
 	/*
@@ -2247,7 +2326,8 @@ static int __find_hibernation_swap_type(dev_t device, sector_t offset)
 	for (type = 0; type < nr_swapfiles; type++) {
 		struct swap_info_struct *sis = swap_info[type];
 
-		if (!(sis->flags & SWP_WRITEOK))
+		if (!(sis->flags & SWP_WRITEOK) ||
+		    (sis->flags & SWP_OFFLOAD_ONLY))
 			continue;
 
 		if (device == sis->bdev->bd_dev) {
@@ -2434,7 +2514,8 @@ int find_first_swap(dev_t *device)
 	for (type = 0; type < nr_swapfiles; type++) {
 		struct swap_info_struct *sis = swap_info[type];
 
-		if (!(sis->flags & SWP_WRITEOK))
+		if (!(sis->flags & SWP_WRITEOK) ||
+		    (sis->flags & SWP_OFFLOAD_ONLY))
 			continue;
 		*device = sis->bdev->bd_dev;
 		spin_unlock(&swap_lock);
@@ -2474,7 +2555,8 @@ unsigned int count_swap_pages(int type, int free)
 		struct swap_info_struct *sis = swap_info[type];
 
 		spin_lock(&sis->lock);
-		if (sis->flags & SWP_WRITEOK) {
+		if ((sis->flags & SWP_WRITEOK) &&
+		    !(sis->flags & SWP_OFFLOAD_ONLY)) {
 			n = sis->pages;
 			if (free)
 				n -= swap_usage_in_pages(sis);
@@ -3083,6 +3165,8 @@ static int setup_swap_extents(struct swap_info_struct *sis,
 
 static void _enable_swap_info(struct swap_info_struct *si)
 {
+	if (si->flags & SWP_OFFLOAD_ONLY)
+		atomic_long_add(si->pages, &nr_swap_pages_offload_only);
 	atomic_long_add(si->pages, &nr_swap_pages);
 	total_swap_pages += si->pages;
 
@@ -3231,6 +3315,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
 	spin_lock(&p->lock);
 	del_from_avail_list(p, true);
 	plist_del(&p->list, &swap_active_head);
+	if (p->flags & SWP_OFFLOAD_ONLY)
+		atomic_long_sub(p->pages, &nr_swap_pages_offload_only);
 	atomic_long_sub(p->pages, &nr_swap_pages);
 	total_swap_pages -= p->pages;
 	spin_unlock(&p->lock);
@@ -3727,7 +3813,6 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
 
 	if (swap_flags & ~SWAP_FLAGS_VALID)
 		return -EINVAL;
-
 	if (!capable(CAP_SYS_ADMIN))
 		return -EPERM;
 
@@ -3841,6 +3926,9 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
 	if (error)
 		goto bad_swap_unlock_inode;
 
+	if (swap_flags & SWAP_FLAG_OFFLOAD_ONLY)
+		si->flags |= SWP_OFFLOAD_ONLY;
+
 	if ((swap_flags & SWAP_FLAG_DISCARD) &&
 	    si->bdev && bdev_max_discard_sectors(si->bdev)) {
 		/*
@@ -3863,6 +3951,18 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
 		else if (swap_flags & SWAP_FLAG_DISCARD_PAGES)
 			si->flags &= ~SWP_AREA_DISCARD;
 
+		/*
+		 * Cluster discard can run later from discard_work, after the
+		 * context which freed the entries has ended.  Swapon-time discard
+		 * is explicit and synchronous, but page discard cannot honour
+		 * offload provenance.
+		 */
+		if ((si->flags & SWP_OFFLOAD_ONLY) &&
+		    (si->flags & SWP_PAGE_DISCARD)) {
+			error = -EINVAL;
+			goto bad_swap_unlock_inode;
+		}
+
 		/* issue a swapon-time discard if it's still required */
 		if (si->flags & SWP_AREA_DISCARD) {
 			int err = discard_swap(si);
diff --git a/mm/vmscan.c b/mm/vmscan.c
index c2eb8fa9d5e50..201b71b760492 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -123,6 +123,9 @@ struct scan_control {
 	/* Proactive reclaim invoked by userspace */
 	unsigned int proactive:1;
 
+	/* This reclaim context may use offload-only swap */
+	unsigned int allow_offload_swap:1;
+
 	/*
 	 * Cgroup memory below memory.low is protected as long as we
 	 * don't threaten to OOM. If any cgroup is reclaimed at
@@ -291,14 +294,18 @@ static inline bool is_exec_file_folio(const struct folio *folio,
 }
 
 static void set_task_reclaim_state(struct task_struct *task,
-				   struct reclaim_state *rs)
+				   struct scan_control *sc)
 {
+	struct reclaim_state *rs = sc ? &sc->reclaim_state : NULL;
+
 	/* Check for an overwrite */
 	WARN_ON_ONCE(rs && task->reclaim_state);
 
 	/* Check for the nulling of an already-nulled member */
 	WARN_ON_ONCE(!rs && !task->reclaim_state);
 
+	if (rs)
+		rs->allow_offload_swap = sc->allow_offload_swap;
 	task->reclaim_state = rs;
 }
 
@@ -418,7 +425,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,
 		 * And under GFP_NOIO, is there enough swapcached anon to make
 		 * scanning anon worthwhile?
 		 */
-		if (get_nr_swap_pages() > 0 &&
+		if (get_nr_swap_pages_eligible() > 0 &&
 		    !reclaimable_anon_is_low(memcg, nid, sc))
 			return true;
 	} else {
@@ -426,7 +433,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,
 		 * Is the memcg below its swap limit, and under GFP_NOIO does
 		 * it have enough swapcached anon to make scanning worthwhile?
 		 */
-		if (mem_cgroup_get_nr_swap_pages(memcg) > 0 &&
+		if (mem_cgroup_get_nr_swap_pages_eligible(memcg) > 0 &&
 		    !reclaimable_anon_is_low(memcg, nid, sc))
 			return true;
 	}
@@ -2853,7 +2860,7 @@ static int get_swappiness(struct lruvec *lruvec, struct scan_control *sc)
 		return 0;
 
 	if (!can_demote(pgdat->node_id, sc, memcg) &&
-	    mem_cgroup_get_nr_swap_pages(memcg) < MIN_LRU_BATCH)
+	    mem_cgroup_get_nr_swap_pages_eligible(memcg) < MIN_LRU_BATCH)
 		return 0;
 
 	return swappiness;
@@ -5926,6 +5933,7 @@ static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,
 		.reclaim_idx = MAX_NR_ZONES - 1,
 		.gfp_mask = GFP_KERNEL,
 		.proactive = true,
+		.allow_offload_swap = true,
 	};
 
 	buf = kvmalloc(len + 1, GFP_KERNEL);
@@ -5937,7 +5945,7 @@ static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,
 		return -EFAULT;
 	}
 
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 	flags = memalloc_noreclaim_save();
 	blk_start_plug(&plug);
 	if (!set_mm_walk(NULL, true)) {
@@ -6941,7 +6949,7 @@ unsigned long try_to_free_pages(struct zonelist *zonelist, int order,
 	if (throttle_direct_reclaim(sc.gfp_mask, zonelist, nodemask))
 		return 1;
 
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 	trace_mm_vmscan_direct_reclaim_begin(sc.gfp_mask, order, NULL);
 
 	nr_reclaimed = do_try_to_free_pages(zonelist, &sc);
@@ -6974,6 +6982,8 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,
 		.may_unmap = 1,
 		.may_swap = !!(reclaim_options & MEMCG_RECLAIM_MAY_SWAP),
 		.proactive = !!(reclaim_options & MEMCG_RECLAIM_PROACTIVE),
+		.allow_offload_swap =
+			!!(reclaim_options & MEMCG_RECLAIM_PROACTIVE),
 	};
 	/*
 	 * Traverse the ZONELIST_FALLBACK zonelist of the current node to put
@@ -6982,7 +6992,7 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,
 	 */
 	struct zonelist *zonelist = node_zonelist(numa_node_id(), sc.gfp_mask);
 
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 	trace_mm_vmscan_memcg_reclaim_begin(sc.gfp_mask, 0, memcg);
 	noreclaim_flag = memalloc_noreclaim_save();
 
@@ -7272,7 +7282,7 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx)
 
 	trace_mm_vmscan_balance_pgdat_begin(pgdat->node_id, order,
 					    highest_zoneidx);
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 	psi_memstall_enter(&pflags);
 	__fs_reclaim_acquire(_THIS_IP_);
 
@@ -7769,7 +7779,7 @@ unsigned long shrink_all_memory(unsigned long nr_to_reclaim)
 
 	fs_reclaim_acquire(sc.gfp_mask);
 	noreclaim_flag = memalloc_noreclaim_save();
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 
 	nr_reclaimed = do_try_to_free_pages(zonelist, &sc);
 
@@ -7950,7 +7960,7 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
 	 * We need to be able to allocate from the reserves for RECLAIM_UNMAP
 	 */
 	noreclaim_flag = memalloc_noreclaim_save();
-	set_task_reclaim_state(p, &sc->reclaim_state);
+	set_task_reclaim_state(p, sc);
 
 	do {
 		shrink_node(pgdat, sc);
@@ -8136,6 +8146,7 @@ int user_proactive_reclaim(char *buf,
 				.may_unmap = 1,
 				.may_swap = 1,
 				.proactive = 1,
+				.allow_offload_swap = 1,
 			};
 
 			if (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED,
diff --git a/mm/vmstat.c b/mm/vmstat.c
index a3e809c57f295..b924c715886e9 100644
--- a/mm/vmstat.c
+++ b/mm/vmstat.c
@@ -1331,6 +1331,7 @@ const char * const vmstat_text[] = {
 	[I(PGPGOUT)]				= "pgpgout",
 	[I(PSWPIN)]				= "pswpin",
 	[I(PSWPOUT)]				= "pswpout",
+	[I(SWPOUT_OFFLOAD_REFUSED)]		= "swpout_offload_refused",
 
 #define OFF (NR_VM_ZONE_STAT_ITEMS + NR_VM_NUMA_EVENT_ITEMS + \
 	     NR_VM_NODE_STAT_ITEMS + NR_VM_STAT_ITEMS)
diff --git a/tools/testing/selftests/zram/.gitignore b/tools/testing/selftests/zram/.gitignore
index 088cd9bad87a7..74b0217c60f04 100644
--- a/tools/testing/selftests/zram/.gitignore
+++ b/tools/testing/selftests/zram/.gitignore
@@ -1,2 +1,4 @@
 # SPDX-License-Identifier: GPL-2.0-only
 err.log
+swap_offload
+workingset_offload
diff --git a/tools/testing/selftests/zram/Makefile b/tools/testing/selftests/zram/Makefile
index 7f78eb1b59cb4..b83e6a1563b20 100644
--- a/tools/testing/selftests/zram/Makefile
+++ b/tools/testing/selftests/zram/Makefile
@@ -1,9 +1,9 @@
 # SPDX-License-Identifier: GPL-2.0
 all:
 
-TEST_PROGS := zram.sh
+TEST_GEN_FILES := swap_offload workingset_offload
+TEST_PROGS := zram.sh zram03.sh zram04.sh zram05.sh
 TEST_FILES := zram01.sh zram02.sh zram_lib.sh
 EXTRA_CLEAN := err.log
 
 include ../lib.mk
-
diff --git a/tools/testing/selftests/zram/config b/tools/testing/selftests/zram/config
index e0cc47e2c7e22..018429d8f5ca8 100644
--- a/tools/testing/selftests/zram/config
+++ b/tools/testing/selftests/zram/config
@@ -1,2 +1,10 @@
+CONFIG_CGROUPS=y
+CONFIG_BLK_DEV_DM=y
+CONFIG_DM_DELAY=y
+CONFIG_MEMCG=y
+CONFIG_SWAP=y
+CONFIG_TRANSPARENT_HUGEPAGE=y
+CONFIG_VM_EVENT_COUNTERS=y
 CONFIG_ZSMALLOC=y
-CONFIG_ZRAM=m
+CONFIG_ZRAM=y
+CONFIG_ZSWAP=y
diff --git a/tools/testing/selftests/zram/swap_offload.c b/tools/testing/selftests/zram/swap_offload.c
new file mode 100644
index 0000000000000..1850b3b4809f0
--- /dev/null
+++ b/tools/testing/selftests/zram/swap_offload.c
@@ -0,0 +1,402 @@
+// SPDX-License-Identifier: GPL-2.0
+#define _GNU_SOURCE
+
+#include <errno.h>
+#include <fcntl.h>
+#include <limits.h>
+#include <sched.h>
+#include <signal.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/syscall.h>
+#include <unistd.h>
+
+#define SWAP_FLAG_PREFER	0x8000
+#define SWAP_FLAG_DISCARD	0x10000
+#define SWAP_FLAG_DISCARD_ONCE	0x20000
+#define SWAP_FLAG_DISCARD_PAGES	0x40000
+#define SWAP_FLAG_OFFLOAD_ONLY	0x80000
+#define KSFT_SKIP	4
+
+static int activate(const char *path, int priority, int discard_flags)
+{
+	int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+		priority | discard_flags;
+
+	if (syscall(SYS_swapon, path, flags)) {
+		perror("swapon");
+		return 1;
+	}
+
+	return 0;
+}
+
+static int reject_page_discard(const char *path, int priority)
+{
+	int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+		SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_PAGES | priority;
+	int ret;
+
+	errno = 0;
+	ret = syscall(SYS_swapon, path, flags);
+	if (ret == -1 && errno == EINVAL)
+		return 0;
+	if (!ret) {
+		syscall(SYS_swapoff, path);
+		fprintf(stderr, "offload-only page discard was accepted\n");
+	} else {
+		fprintf(stderr, "swapon returned unexpected error: %s\n",
+			strerror(errno));
+	}
+	return 1;
+}
+
+static int accept_discard_once_pages(const char *path, int priority)
+{
+	int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+		SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE |
+		SWAP_FLAG_DISCARD_PAGES | priority;
+
+	if (syscall(SYS_swapon, path, flags)) {
+		perror("swapon discard-once+discard-pages");
+		return 1;
+	}
+	if (syscall(SYS_swapoff, path)) {
+		perror("swapoff discard-once+discard-pages");
+		return 1;
+	}
+	return 0;
+}
+
+static int join_cgroup(const char *procs)
+{
+	char pid[32];
+	int fd, len;
+
+	fd = open(procs, O_WRONLY);
+	if (fd < 0) {
+		perror("open cgroup.procs");
+		return 1;
+	}
+
+	len = snprintf(pid, sizeof(pid), "%d\n", getpid());
+	if (write(fd, pid, len) != len) {
+		perror("write cgroup.procs");
+		close(fd);
+		return 1;
+	}
+	close(fd);
+	return 0;
+}
+
+static int pin_to_one_cpu(const char *pid_arg)
+{
+	cpu_set_t allowed, selected;
+	char *end;
+	long pid;
+	int cpu;
+
+	errno = 0;
+	pid = strtol(pid_arg, &end, 10);
+	if (errno || *end || pid <= 0) {
+		fprintf(stderr, "invalid pid: %s\n", pid_arg);
+		return 1;
+	}
+
+	if (sched_getaffinity(pid, sizeof(allowed), &allowed)) {
+		perror("sched_getaffinity");
+		return 1;
+	}
+	for (cpu = 0; cpu < CPU_SETSIZE; cpu++)
+		if (CPU_ISSET(cpu, &allowed))
+			break;
+	if (cpu == CPU_SETSIZE) {
+		fprintf(stderr, "pid %ld has no allowed CPU\n", pid);
+		return 1;
+	}
+
+	CPU_ZERO(&selected);
+	CPU_SET(cpu, &selected);
+	if (sched_setaffinity(pid, sizeof(selected), &selected)) {
+		perror("sched_setaffinity");
+		return 1;
+	}
+	return 0;
+}
+
+static int allocate(const char *size_arg, const char *procs,
+		    const char *ready, const char *verified)
+{
+	unsigned char *memory;
+	unsigned long size;
+	unsigned long page_size;
+	char *end;
+	sigset_t signals;
+	int signal;
+	int fd;
+
+	errno = 0;
+	size = strtoul(size_arg, &end, 0);
+	if (errno || *end || !size) {
+		fprintf(stderr, "invalid allocation size: %s\n", size_arg);
+		return 1;
+	}
+	if (join_cgroup(procs))
+		return 1;
+
+	memory = mmap(NULL, size, PROT_READ | PROT_WRITE,
+		      MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+	if (memory == MAP_FAILED) {
+		perror("mmap");
+		return 1;
+	}
+
+	page_size = getpagesize();
+	for (unsigned long i = 0; i < size; i += page_size)
+		memset(memory + i, i / page_size % 251 + 1,
+		       page_size < size - i ? page_size : size - i);
+	if (mprotect(memory, size, PROT_READ)) {
+		perror("mprotect");
+		return 1;
+	}
+
+	sigemptyset(&signals);
+	sigaddset(&signals, SIGUSR1);
+	if (sigprocmask(SIG_BLOCK, &signals, NULL)) {
+		perror("sigprocmask");
+		return 1;
+	}
+
+	fd = open(ready, O_WRONLY | O_CREAT | O_EXCL, 0600);
+	if (fd < 0) {
+		perror("create ready file");
+		return 1;
+	}
+	close(fd);
+
+	errno = sigwait(&signals, &signal);
+	if (errno) {
+		perror("sigwait");
+		return 1;
+	}
+	for (unsigned long i = 0; i < size; i++) {
+		unsigned char expected = i / page_size % 251 + 1;
+
+		if (memory[i] != expected) {
+			fprintf(stderr,
+				"data mismatch at %lu: got %u, expected %u\n",
+				i, memory[i], expected);
+			return 1;
+		}
+	}
+
+	fd = open(verified, O_WRONLY | O_CREAT | O_EXCL, 0600);
+	if (fd < 0) {
+		perror("create verified file");
+		return 1;
+	}
+	close(fd);
+
+	for (;;)
+		pause();
+}
+
+static int create_marker(const char *path, int write_pid)
+{
+	char *temporary;
+	int fd, ret = 1;
+
+	if (asprintf(&temporary, "%s.XXXXXX", path) < 0) {
+		perror("asprintf marker path");
+		return 1;
+	}
+	fd = mkstemp(temporary);
+	if (fd < 0) {
+		perror(path);
+		goto out_free;
+	}
+	if (write_pid && dprintf(fd, "%d\n", getpid()) < 0) {
+		perror("write marker");
+		close(fd);
+		goto out_unlink;
+	}
+	if (close(fd)) {
+		perror("close marker");
+		goto out_unlink;
+	}
+	/* Publish only after the payload is complete for the shell reader. */
+	if (rename(temporary, path)) {
+		perror("publish marker");
+		goto out_unlink;
+	}
+	ret = 0;
+out_unlink:
+	unlink(temporary);
+out_free:
+	free(temporary);
+	return ret;
+}
+
+static unsigned char retained_byte(unsigned long offset, long page_size,
+				   int touched)
+{
+	unsigned char value = offset / page_size % 251 + 1;
+
+	if (touched && !offset)
+		value ^= 0x5a;
+	return value;
+}
+
+static int verify_retained(unsigned char *memory, unsigned long size,
+			   long page_size, int full)
+{
+	unsigned long limit = full ? size : 1;
+
+	for (unsigned long offset = 0; offset < limit; offset++) {
+		unsigned char expected = retained_byte(offset, page_size, 1);
+
+		if (memory[offset] != expected) {
+			fprintf(stderr,
+				"retained data mismatch at %lu: got %u, expected %u\n",
+				offset, memory[offset], expected);
+			return 1;
+		}
+	}
+	return 0;
+}
+
+static int retained(const char *size_arg, const char *procs,
+		    const char *ready, const char *touched,
+		    const char *verified)
+{
+	unsigned long allocation_size, mapping_start, size;
+	unsigned char *mapping, *memory;
+	char *end;
+	sigset_t signals;
+	long page_size;
+	int signal;
+
+	errno = 0;
+	size = strtoul(size_arg, &end, 0);
+	if (errno || *end || !size || (size & (size - 1))) {
+		fprintf(stderr, "allocation size must be a power of two\n");
+		return 1;
+	}
+	if (join_cgroup(procs))
+		return 1;
+
+	page_size = sysconf(_SC_PAGESIZE);
+	if (page_size <= 0) {
+		perror("sysconf _SC_PAGESIZE");
+		return 1;
+	}
+	if (size < (unsigned long)page_size || size % page_size) {
+		fprintf(stderr, "allocation size must contain whole pages\n");
+		return 1;
+	}
+	if (size > ULONG_MAX / 2) {
+		fprintf(stderr, "allocation size is too large\n");
+		return 1;
+	}
+	allocation_size = size * 2;
+	mapping = mmap(NULL, allocation_size, PROT_READ | PROT_WRITE,
+		       MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+	if (mapping == MAP_FAILED) {
+		perror("mmap");
+		return 1;
+	}
+	mapping_start = (unsigned long)mapping;
+	memory = (unsigned char *)((mapping_start + size - 1) & ~(size - 1));
+	if ((unsigned long)memory != mapping_start)
+		munmap((void *)mapping_start,
+		       (unsigned long)memory - mapping_start);
+	munmap(memory + size, allocation_size - size -
+	       ((unsigned long)memory - mapping_start));
+	if (madvise(memory, size, MADV_HUGEPAGE)) {
+		perror("madvise MADV_HUGEPAGE");
+		return 1;
+	}
+	for (unsigned long offset = 0; offset < size; offset += page_size)
+		memset(memory + offset, retained_byte(offset, page_size, 0),
+		       page_size);
+#ifdef MADV_COLLAPSE
+	if (madvise(memory, size, MADV_COLLAPSE)) {
+		int error = errno;
+
+		perror("madvise MADV_COLLAPSE");
+		if (error == EAGAIN || error == EINVAL || error == ENOMEM)
+			return KSFT_SKIP;
+		return 1;
+	}
+#else
+	fprintf(stderr, "MADV_COLLAPSE is unavailable\n");
+	return KSFT_SKIP;
+#endif
+
+	sigemptyset(&signals);
+	sigaddset(&signals, SIGUSR1);
+	sigaddset(&signals, SIGUSR2);
+	sigaddset(&signals, SIGALRM);
+	if (sigprocmask(SIG_BLOCK, &signals, NULL)) {
+		perror("sigprocmask");
+		return 1;
+	}
+	if (create_marker(ready, 1))
+		return 1;
+
+	for (;;) {
+		errno = sigwait(&signals, &signal);
+		if (errno) {
+			perror("sigwait");
+			return 1;
+		}
+		if (signal == SIGUSR1) {
+			unsigned char value;
+
+			if (mprotect(memory, page_size, PROT_READ)) {
+				perror("mprotect read");
+				return 1;
+			}
+			value = memory[0];
+			if (mprotect(memory, page_size, PROT_READ | PROT_WRITE)) {
+				perror("mprotect write");
+				return 1;
+			}
+			memory[0] = value ^ 0x5a;
+			if (create_marker(touched, 0))
+				return 1;
+		} else {
+			if (verify_retained(memory, size, page_size,
+					    signal == SIGALRM))
+				return 1;
+			if (create_marker(verified, 0))
+				return 1;
+		}
+	}
+}
+
+int main(int argc, char **argv)
+{
+	if (argc == 4 && !strcmp(argv[1], "activate"))
+		return activate(argv[2], atoi(argv[3]),
+				SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE);
+	if (argc == 4 && !strcmp(argv[1], "activate-no-discard"))
+		return activate(argv[2], atoi(argv[3]), 0);
+	if (argc == 4 && !strcmp(argv[1], "reject-page-discard"))
+		return reject_page_discard(argv[2], atoi(argv[3]));
+	if (argc == 4 && !strcmp(argv[1], "accept-discard-once-pages"))
+		return accept_discard_once_pages(argv[2], atoi(argv[3]));
+	if (argc == 3 && !strcmp(argv[1], "pin"))
+		return pin_to_one_cpu(argv[2]);
+	if (argc == 6 && !strcmp(argv[1], "allocate"))
+		return allocate(argv[2], argv[3], argv[4], argv[5]);
+	if (argc == 7 && !strcmp(argv[1], "retained"))
+		return retained(argv[2], argv[3], argv[4], argv[5], argv[6]);
+
+	fprintf(stderr,
+		"usage: %s activate|activate-no-discard DEVICE PRIORITY | reject-page-discard DEVICE PRIORITY | accept-discard-once-pages DEVICE PRIORITY | pin PID | allocate BYTES CGROUP.PROCS READY VERIFIED | retained BYTES CGROUP.PROCS READY TOUCHED VERIFIED\n",
+		argv[0]);
+	return 1;
+}
diff --git a/tools/testing/selftests/zram/workingset_offload.c b/tools/testing/selftests/zram/workingset_offload.c
new file mode 100644
index 0000000000000..6d41badb8d0be
--- /dev/null
+++ b/tools/testing/selftests/zram/workingset_offload.c
@@ -0,0 +1,206 @@
+// SPDX-License-Identifier: GPL-2.0
+#define _GNU_SOURCE
+
+#include <errno.h>
+#include <fcntl.h>
+#include <signal.h>
+#include <stdbool.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <unistd.h>
+
+#define ANON_SIZE (64UL << 20)
+#define TARGET_SIZE (16UL << 20)
+#define FILLER_SIZE (32UL << 20)
+
+static int join_cgroup(const char *path)
+{
+	char pid[32];
+	int fd, len;
+
+	fd = open(path, O_WRONLY);
+	if (fd < 0)
+		return -1;
+	len = snprintf(pid, sizeof(pid), "%d\n", getpid());
+	if (write(fd, pid, len) != len) {
+		close(fd);
+		return -1;
+	}
+	return close(fd);
+}
+
+static int create_file(const char *path, size_t size)
+{
+	unsigned char page[4096];
+	size_t offset;
+	int fd;
+
+	fd = open(path, O_CREAT | O_TRUNC | O_RDWR, 0600);
+	if (fd < 0)
+		return -1;
+	memset(page, 0xa5, sizeof(page));
+	for (offset = 0; offset < size; offset += sizeof(page)) {
+		if (pwrite(fd, page, sizeof(page), offset) != sizeof(page)) {
+			close(fd);
+			return -1;
+		}
+	}
+	if (fsync(fd) || posix_fadvise(fd, 0, size, POSIX_FADV_DONTNEED)) {
+		close(fd);
+		return -1;
+	}
+	return fd;
+}
+
+static int cache_file(int fd, size_t size)
+{
+	unsigned char byte;
+	size_t offset;
+
+	if (posix_fadvise(fd, 0, size, POSIX_FADV_RANDOM))
+		return -1;
+	for (offset = 0; offset < size; offset += getpagesize())
+		if (pread(fd, &byte, 1, offset) != 1)
+			return -1;
+	return 0;
+}
+
+static unsigned long memory_stat(const char *path, const char *key)
+{
+	unsigned long value;
+	char name[64];
+	FILE *file;
+
+	file = fopen(path, "r");
+	if (!file)
+		return 0;
+	while (fscanf(file, "%63s %lu", name, &value) == 2) {
+		if (!strcmp(name, key)) {
+			fclose(file);
+			return value;
+		}
+	}
+	fclose(file);
+	return 0;
+}
+
+static int touch_ready(const char *path)
+{
+	int fd = open(path, O_WRONLY | O_CREAT | O_EXCL, 0600);
+
+	if (fd < 0)
+		return -1;
+	return close(fd);
+}
+
+static bool calibration_sufficient(size_t pages, unsigned long selected,
+				   unsigned long refaulted,
+				   unsigned long anon)
+{
+	return selected >= pages / 2 && refaulted >= selected * 3 / 4 &&
+	       anon >= TARGET_SIZE;
+}
+
+int main(int argc, char **argv)
+{
+	unsigned long refault_before, refault_after;
+	unsigned long activate_before, activate_after;
+	unsigned long anon, selected = 0;
+	unsigned char *resident, *anon_memory;
+	unsigned char byte, checksum = 0;
+	char stat_path[4096];
+	sigset_t signals;
+	size_t pages, i;
+	void *mapping;
+	int target_fd, filler_fd, signal;
+
+	if (argc != 6) {
+		fprintf(stderr, "usage: %s CGROUP.PROCS TARGET FILLER READY GO\n",
+			argv[0]);
+		return 1;
+	}
+	if (join_cgroup(argv[1])) {
+		perror("join cgroup");
+		return 1;
+	}
+
+	anon_memory = mmap(NULL, ANON_SIZE, PROT_READ | PROT_WRITE,
+			   MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+	if (anon_memory == MAP_FAILED) {
+		perror("mmap anonymous");
+		return 1;
+	}
+	for (i = 0; i < ANON_SIZE; i += getpagesize())
+		anon_memory[i] = i / getpagesize() % 251 + 1;
+
+	target_fd = create_file(argv[2], TARGET_SIZE);
+	filler_fd = create_file(argv[3], FILLER_SIZE);
+	if (target_fd < 0 || filler_fd < 0) {
+		perror("create files");
+		return 1;
+	}
+	if (cache_file(target_fd, TARGET_SIZE) ||
+	    cache_file(filler_fd, FILLER_SIZE)) {
+		perror("populate page cache");
+		return 1;
+	}
+
+	sigemptyset(&signals);
+	sigaddset(&signals, SIGUSR1);
+	if (sigprocmask(SIG_BLOCK, &signals, NULL) || touch_ready(argv[4])) {
+		perror("prepare signal");
+		return 1;
+	}
+	if (sigwait(&signals, &signal)) {
+		perror("sigwait");
+		return 1;
+	}
+
+	mapping = mmap(NULL, TARGET_SIZE, PROT_READ, MAP_SHARED, target_fd, 0);
+	if (mapping == MAP_FAILED) {
+		perror("mmap target");
+		return 1;
+	}
+	pages = TARGET_SIZE / getpagesize();
+	resident = calloc(pages, 1);
+	if (!resident || mincore(mapping, TARGET_SIZE, resident)) {
+		perror("mincore");
+		return 1;
+	}
+	munmap(mapping, TARGET_SIZE);
+
+	snprintf(stat_path, sizeof(stat_path), "%.*s/memory.stat",
+		 (int)(strlen(argv[1]) - strlen("/cgroup.procs")), argv[1]);
+	refault_before = memory_stat(stat_path, "workingset_refault_file");
+	activate_before = memory_stat(stat_path, "workingset_activate_file");
+	for (i = 0; i < pages; i++) {
+		if (resident[i] & 1)
+			continue;
+		if (pread(target_fd, &byte, 1, i * getpagesize()) != 1) {
+			perror("refault target");
+			return 1;
+		}
+		selected++;
+	}
+	refault_after = memory_stat(stat_path, "workingset_refault_file");
+	activate_after = memory_stat(stat_path, "workingset_activate_file");
+	anon = memory_stat(stat_path, "active_anon") +
+	       memory_stat(stat_path, "inactive_anon");
+
+	for (i = 0; i < ANON_SIZE; i += getpagesize())
+		checksum ^= anon_memory[i];
+	printf("selected=%lu refault=%lu activate=%lu anon=%lu checksum=%u\n",
+	       selected, refault_after - refault_before,
+	       activate_after - activate_before, anon, checksum);
+
+	free(resident);
+	close(filler_fd);
+	close(target_fd);
+	if (!calibration_sufficient(pages, selected,
+				    refault_after - refault_before, anon))
+		return 2;
+	return (activate_after - activate_before) * 100 < selected * 80;
+}
diff --git a/tools/testing/selftests/zram/zram03.sh b/tools/testing/selftests/zram/zram03.sh
new file mode 100755
index 0000000000000..3ccc9a135f997
--- /dev/null
+++ b/tools/testing/selftests/zram/zram03.sh
@@ -0,0 +1,166 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Test proactive-only swap allocation and pressure fallback.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram03"
+cg="/sys/fs/cgroup/zram-offload-$$"
+cgroup_root="/sys/fs/cgroup"
+ready="/tmp/zram-offload-ready-$$"
+verified="/tmp/zram-offload-verified-$$"
+allocator_pid=""
+offload=""
+zswap_enabled=""
+zswap_writeback=""
+
+fail()
+{
+	echo "$TCID: [FAIL] $*" >&2
+	exit 1
+}
+
+skip()
+{
+	echo "$TCID: [SKIP] $*" >&2
+	exit "$ksft_skip"
+}
+
+cleanup()
+{
+	set +e
+	if [ -n "$allocator_pid" ]; then
+		kill "$allocator_pid"
+		wait "$allocator_pid"
+	fi
+	rm -f "$ready" "$verified"
+	rmdir "$cg"
+	cgroup_disable_memory_controller "$cgroup_root"
+	if [ -n "$offload" ] && [ "$dev_makeswap" -lt "$dev_end" ]; then
+		swapoff "$offload" >/dev/null 2>&1
+	fi
+	if [ "$dev_end" -ge "$dev_start" ]; then
+		zram_cleanup
+	fi
+	if [ -n "$zswap_enabled" ]; then
+		echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled
+	fi
+}
+
+swap_used_kb()
+{
+	awk -v device="$1" '$1 == device { print $4 }' /proc/swaps
+}
+
+zram_orig_data_size()
+{
+	awk '{ print $1 }' "/sys/block/${1##*/}/mm_stat"
+}
+
+wait_file()
+{
+	for _ in $(seq 1 400); do
+		[ -e "$1" ] && return 0
+		sleep 0.05
+	done
+	return 1
+}
+
+check_prereqs
+# The feature marker also requires CONFIG_VM_EVENT_COUNTERS.
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+	skip "offload refusal counters are unavailable"
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] || skip "cgroup v2 controllers are unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers || skip "memory controller is unavailable"
+[ -e /sys/module/zswap/parameters/enabled ] || skip "zswap is unavailable"
+zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+
+# Keep every reclaim transition on one CPU.  This makes the test exercise the
+# cached-cluster provenance mismatch rather than relying on scheduler placement.
+./swap_offload pin "$$" || skip "cannot pin test to one allowed CPU"
+
+# Swap priorities are global. Put both test areas ahead of any existing swap
+# so distro-managed zram or another high-priority area cannot absorb the test
+# allocations and make the routing assertions fail spuriously.
+max_prio=$(awk 'BEGIN { max = -1 } NR > 1 && $5 > max { max = $5 } END { print max }' /proc/swaps)
+[ "$max_prio" -le 32765 ] || skip "cannot outrank existing swap priority $max_prio"
+safe_prio=$((max_prio + 1))
+offload_prio=$((max_prio + 2))
+
+dev_num=2
+zram_sizes="67108864 67108864"
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+zram_load
+echo Y > /sys/module/zswap/parameters/enabled || skip "cannot enable zswap"
+zram_set_disksizes
+
+safe="/dev/zram${dev_start}"
+offload="/dev/zram$((dev_start + 1))"
+mkswap "$safe" >/dev/null
+mkswap "$offload" >/dev/null
+swapon -p "$safe_prio" "$safe"
+dev_makeswap=$dev_start
+./swap_offload reject-page-discard "$offload" "$offload_prio" ||
+	fail "offload-only page discard was accepted"
+./swap_offload accept-discard-once-pages "$offload" "$offload_prio" ||
+	fail "resolved offload-only discard-once was rejected"
+./swap_offload activate "$offload" "$offload_prio"
+dev_makeswap=$dev_end
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+	skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+[ -e "$cg/memory.max" ] || skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.swap.max"
+echo 1 > "$cg/memory.zswap.writeback" ||
+	skip "cannot enable zswap writeback for test cgroup"
+read -r zswap_writeback < "$cg/memory.zswap.writeback"
+[ "$zswap_writeback" -eq 1 ] ||
+	skip "an ancestor disables zswap writeback"
+
+./swap_offload allocate 67108864 "$cg/cgroup.procs" "$ready" "$verified" &
+allocator_pid=$!
+wait_file "$ready" || fail "allocator did not become ready"
+
+echo "16M swappiness=max" > "$cg/memory.reclaim" ||
+	fail "proactive reclaim failed"
+offload_before=$(swap_used_kb "$offload")
+safe_before=$(swap_used_kb "$safe")
+offload_orig=$(zram_orig_data_size "$offload")
+[ "${offload_before:-0}" -gt 0 ] || fail "proactive reclaim missed offload area"
+[ "${safe_before:-0}" -eq 0 ] || fail "proactive reclaim used lower-priority safe area"
+[ "$((offload_orig + 1048576))" -ge "$((offload_before * 1024))" ] ||
+	fail "zswap deferred the offload-only backend write"
+
+echo 24M > "$cg/memory.max"
+sleep 1
+offload_pressure=$(swap_used_kb "$offload")
+safe_pressure=$(swap_used_kb "$safe")
+safe_orig=$(zram_orig_data_size "$safe")
+[ "${safe_pressure:-0}" -gt 0 ] || fail "pressure reclaim missed safe fallback"
+[ "$offload_pressure" -le "$offload_before" ] ||
+	fail "pressure reclaim allocated offload-only slots"
+[ "$safe_orig" -lt "$((safe_pressure * 1024))" ] ||
+	fail "zswap did not retain any unmarked safe-area pages"
+
+echo max > "$cg/memory.max"
+echo "4M swappiness=max" > "$cg/memory.reclaim" ||
+	fail "second proactive reclaim failed"
+offload_after=$(swap_used_kb "$offload")
+[ "$offload_after" -gt "$offload_pressure" ] ||
+	fail "proactive priority was not restored after pressure reclaim"
+
+swapoff "$offload" || fail "swapoff could not recover offloaded pages"
+kill -USR1 "$allocator_pid" || fail "cannot request data verification"
+wait_file "$verified" || fail "allocator did not verify recovered data"
+kill -0 "$allocator_pid" || fail "allocator died during swapoff"
+dev_makeswap=$dev_start
+
+echo "$TCID: [PASS]"
diff --git a/tools/testing/selftests/zram/zram04.sh b/tools/testing/selftests/zram/zram04.sh
new file mode 100755
index 0000000000000..a8f52a82e774b
--- /dev/null
+++ b/tools/testing/selftests/zram/zram04.sh
@@ -0,0 +1,149 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Reproduce offload-only swap leaking into the workingset refault heuristic.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram04"
+cg="/sys/fs/cgroup/zram-workingset-$$"
+cgroup_root="/sys/fs/cgroup"
+tmp="${TMPDIR:-/var/tmp}/zram-workingset-$$"
+ready="$tmp/ready"
+worker=""
+tmp_fs=""
+mglru=""
+zswap_enabled=""
+
+fail()
+{
+	echo "$TCID: [FAIL] $*" >&2
+	exit 1
+}
+
+skip()
+{
+	echo "$TCID: [SKIP] $*" >&2
+	exit "$ksft_skip"
+}
+
+cleanup()
+{
+	set +e
+	[ -n "$worker" ] && kill "$worker"
+	[ -n "$worker" ] && wait "$worker"
+	rm -rf "$tmp"
+	rmdir "$cg"
+	cgroup_disable_memory_controller "$cgroup_root"
+	[ "$dev_end" -ge "$dev_start" ] && zram_cleanup
+	[ -n "$mglru" ] && echo "$mglru" > /sys/kernel/mm/lru_gen/enabled
+	[ -n "$zswap_enabled" ] &&
+		echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled
+}
+
+wait_helper_ready()
+{
+	for _ in $(seq 1 400); do
+		[ -e "$ready" ] && return 0
+		worker_state=$(awk '{ print $3 }' "/proc/$worker/stat" \
+			2>/dev/null || :)
+		if ! kill -0 "$worker" 2>/dev/null ||
+		   [ "$worker_state" = Z ]; then
+			if wait "$worker"; then
+				status=0
+			else
+				status=$?
+			fi
+			worker=""
+			[ "$status" -eq 2 ] &&
+				skip "workingset calibration was insufficient"
+			fail "workingset helper exited with status $status before readiness"
+		fi
+		sleep 0.05
+	done
+	fail "workingset helper timed out before readiness"
+}
+
+check_prereqs
+# The feature marker also requires CONFIG_VM_EVENT_COUNTERS.
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+	skip "offload refusal counters are unavailable"
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+[ -x ./workingset_offload ] || skip "workingset helper is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] || skip "cgroup v2 is unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers ||
+	skip "cgroup v2 memory controller is unavailable"
+[ "$(awk 'END { print NR }' /proc/swaps)" -eq 1 ] ||
+	skip "test requires no pre-existing swap"
+tmp_fs=$(stat -f -c %T "${TMPDIR:-/var/tmp}") ||
+	skip "cannot identify the test filesystem"
+[ "$tmp_fs" != tmpfs ] ||
+	skip "test files require a disk-backed filesystem"
+
+[ -e /sys/kernel/mm/lru_gen/enabled ] &&
+	mglru=$(cat /sys/kernel/mm/lru_gen/enabled)
+[ -e /sys/module/zswap/parameters/enabled ] &&
+	zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+[ -n "$mglru" ] && echo 0 > /sys/kernel/mm/lru_gen/enabled
+[ -n "$zswap_enabled" ] && echo N > /sys/module/zswap/parameters/enabled
+
+dev_num=1
+zram_sizes="134217728"
+zram_load
+zram_set_disksizes
+offload="/dev/zram${dev_start}"
+mkswap "$offload" >/dev/null
+./swap_offload activate "$offload" 1
+dev_makeswap=$dev_end
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+	skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+[ -e "$cg/memory.max" ] || skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.max"
+echo max > "$cg/memory.swap.max"
+mkdir "$tmp"
+
+./workingset_offload "$cg/cgroup.procs" "$tmp/target" "$tmp/filler" \
+	"$ready" unused &
+worker=$!
+wait_helper_ready
+
+# Shadow retention uses local LRU/slab statistics updated by memcg flushing.
+# Check that the 64 MiB anonymous and 48 MiB file setup is visible before
+# evicting file pages.
+stats_ready=0
+for _ in $(seq 1 50); do
+	if awk '
+		$1 == "active_anon" || $1 == "inactive_anon" { anon += $2 }
+		$1 == "file" { file = $2 }
+		END { exit !(anon >= 67108864 && file >= 50331648) }
+	' "$cg/memory.stat"; then
+		stats_ready=1
+		break
+	fi
+	sleep 0.1
+done
+[ "$stats_ready" -eq 1 ] ||
+	skip "initial working-set statistics did not become visible"
+
+echo "48M swappiness=0" > "$cg/memory.reclaim" ||
+	skip "file-only proactive reclaim failed"
+kill -USR1 "$worker"
+if wait "$worker"; then
+	worker=""
+	echo "$TCID: [PASS]"
+	exit 0
+else
+	status=$?
+fi
+worker=""
+[ "$status" -eq 2 ] && skip "workingset calibration was insufficient"
+echo "$TCID: [FAIL] refaulted file pages were not activated" >&2
+exit 1
diff --git a/tools/testing/selftests/zram/zram05.sh b/tools/testing/selftests/zram/zram05.sh
new file mode 100755
index 0000000000000..e4bbf6bfa0e92
--- /dev/null
+++ b/tools/testing/selftests/zram/zram05.sh
@@ -0,0 +1,360 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Test retained offload-only entries under ordinary and proactive reclaim.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram05"
+cg="/sys/fs/cgroup/zram-retained-$$"
+cgroup_root="/sys/fs/cgroup"
+tmp="${TMPDIR:-/var/tmp}/zram-retained-$$"
+ready="$tmp/ready"
+touched="$tmp/touched"
+verified="$tmp/verified"
+holder_pid=""
+safe=""
+safe_active=0
+offload_backing=""
+offload=""
+offload_active=0
+dm_name="zram-retained-$$"
+dm_active=0
+dm_node_created=0
+zswap_enabled=""
+thp_size=0
+thp_kib=0
+thp_sectors=0
+page_kib=0
+
+fail()
+{
+	echo "$TCID: [FAIL] $*" >&2
+	exit 1
+}
+
+skip()
+{
+	echo "$TCID: [SKIP] $*" >&2
+	exit "$ksft_skip"
+}
+
+cleanup()
+{
+	set +e
+	if [ -n "$holder_pid" ]; then
+		kill "$holder_pid"
+		wait "$holder_pid"
+	fi
+	if [ "$offload_active" -eq 1 ]; then
+		swapoff "$offload" >/dev/null 2>&1
+	fi
+	if [ "$safe_active" -eq 1 ]; then
+		swapoff "$safe" >/dev/null 2>&1
+	fi
+	if [ "$dm_active" -eq 1 ]; then
+		dmsetup --noudevsync --noudevrules remove "$dm_name" \
+			>/dev/null 2>&1
+	fi
+	if [ "$dm_node_created" -eq 1 ]; then
+		rm -f "$offload"
+	fi
+	rm -rf "$tmp"
+	rmdir "$cg"
+	cgroup_disable_memory_controller "$cgroup_root"
+	if [ "$dev_end" -ge "$dev_start" ]; then
+		zram_cleanup
+	fi
+	if [ -n "$zswap_enabled" ]; then
+		echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled
+	fi
+}
+
+wait_file()
+{
+	for _ in $(seq 1 400); do
+		[ -e "$1" ] && return 0
+		if [ -n "$holder_pid" ]; then
+			holder_state=$(awk '{ print $3 }' \
+				"/proc/$holder_pid/stat" 2>/dev/null || :)
+		else
+			holder_state=""
+		fi
+		if [ -n "$holder_pid" ] &&
+		   { ! kill -0 "$holder_pid" 2>/dev/null ||
+		     [ "$holder_state" = Z ]; }; then
+			if wait "$holder_pid"; then
+				helper_status=0
+			else
+				helper_status=$?
+			fi
+			holder_pid=""
+			return 2
+		fi
+		sleep 0.05
+	done
+	return 1
+}
+
+require_helper_file()
+{
+	if wait_file "$1"; then
+		return 0
+	else
+		status=$?
+	fi
+	if [ "$status" -eq 2 ]; then
+		[ "$helper_status" -eq "$ksft_skip" ] && skip "$2 is unavailable"
+		fail "retained helper exited with status $helper_status while $2"
+	fi
+	fail "retained helper timed out while $2"
+}
+
+written_sectors()
+{
+	awk '{ print $7 }' "/sys/block/${offload_backing##*/}/stat"
+}
+
+vmstat_value()
+{
+	awk -v name="$1" '$1 == name { print $2 }' /proc/vmstat
+}
+
+wait_offload_quiet()
+{
+	# A bio queued inside dm-delay is not yet visible in either the backing
+	# device statistics or the mapped device's inflight counters.
+	sleep 4
+	previous=-1
+	stable=0
+	for _ in $(seq 1 240); do
+		current=$(written_sectors)
+		read -r reads writes < "$dm_inflight"
+		if [ "$current" -eq "$previous" ] && \
+		   [ "$reads" -eq 0 ] && [ "$writes" -eq 0 ]; then
+			stable=$((stable + 1))
+			[ "$stable" -ge 20 ] && return 0
+		else
+			stable=0
+		fi
+		previous=$current
+		sleep 0.05
+	done
+	return 1
+}
+
+check_prereqs
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+command -v dmsetup >/dev/null 2>&1 || skip "dmsetup is unavailable"
+command -v blockdev >/dev/null 2>&1 || skip "blockdev is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] ||
+	skip "cgroup v2 controllers are unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers ||
+	skip "memory controller is unavailable"
+[ -d /sys/kernel/mm/transparent_hugepage ] ||
+	skip "transparent huge pages are unavailable"
+[ -r /sys/kernel/mm/transparent_hugepage/hpage_pmd_size ] ||
+	skip "PMD huge-page size is unavailable"
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+	skip "offload refusal counters are unavailable"
+
+thp_size=$(cat /sys/kernel/mm/transparent_hugepage/hpage_pmd_size)
+case "$thp_size" in
+	''|*[!0-9]*) skip "invalid PMD huge-page size: $thp_size" ;;
+esac
+[ "$thp_size" -gt 0 ] || skip "PMD huge-page size is zero"
+page_kib=$(awk '/KernelPageSize:/ { print $2; exit }' /proc/self/smaps)
+[ "${page_kib:-0}" -gt 0 ] || skip "cannot determine the base page size"
+page_size=$((page_kib * 1024))
+[ "$page_size" -gt 0 ] || skip "base page size is zero"
+[ $((thp_size % page_size)) -eq 0 ] ||
+	skip "PMD huge-page size is not page aligned"
+thp_kib=$((thp_size / 1024))
+thp_sectors=$((thp_size / 512))
+expected_retained_kib=$((thp_kib - page_kib))
+expected_refused=$((thp_size / page_size))
+
+mkdir "$tmp"
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+dmsetup targets > "$tmp/dm-targets" 2>/dev/null ||
+	skip "cannot query device-mapper targets"
+grep -q '^delay[[:space:]]' "$tmp/dm-targets" ||
+	skip "device-mapper delay target is unavailable"
+
+if [ -e /sys/module/zswap/parameters/enabled ]; then
+	zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+	echo N > /sys/module/zswap/parameters/enabled ||
+		skip "cannot disable zswap"
+fi
+
+# Swap priorities are global.  The ordinary zram device must be preferred to
+# any pre-existing swap, while the delayed offload device remains first for
+# eligible proactive reclaim.
+max_prio=$(awk 'BEGIN { max = -1 } NR > 1 && $5 > max { max = $5 } END { print max }' /proc/swaps)
+[ "$max_prio" -le 32765 ] ||
+	skip "cannot outrank existing swap priority $max_prio"
+safe_prio=$((max_prio + 1))
+offload_prio=$((max_prio + 2))
+
+dev_num=2
+zram_size=$((thp_size * 4))
+[ "$zram_size" -ge 67108864 ] || zram_size=67108864
+zram_sizes="$zram_size $zram_size"
+zram_load
+zram_set_disksizes
+safe="/dev/zram${dev_start}"
+offload_backing="/dev/zram$((dev_start + 1))"
+
+sectors=$(blockdev --getsz "$offload_backing") ||
+	skip "cannot read offload backing size"
+[ "$sectors" -gt 0 ] || skip "offload backing has zero size"
+dm_table="0 $sectors delay $offload_backing 0 0 $offload_backing 0 3000"
+dmsetup --noudevsync --noudevrules create "$dm_name" --table "$dm_table" ||
+	skip "cannot create delayed offload device"
+dm_active=1
+dmsetup --noudevsync --noudevrules info --columns --noheadings \
+	--separator ' ' -o major,minor "$dm_name" > "$tmp/dm-devno" ||
+	skip "cannot identify delayed offload device"
+read -r dm_major dm_minor < "$tmp/dm-devno"
+offload="/dev/mapper/$dm_name"
+if [ ! -e "$offload" ]; then
+	mkdir -p /dev/mapper
+	mknod "$offload" b "$dm_major" "$dm_minor" ||
+		skip "cannot create delayed offload device node"
+	dm_node_created=1
+fi
+dm_inflight="/sys/dev/block/$dm_major:$dm_minor/inflight"
+[ -r "$dm_inflight" ] || skip "cannot observe delayed offload I/O"
+
+mkswap "$safe" >/dev/null || fail "cannot initialise safe swap"
+mkswap "$offload" >/dev/null || fail "cannot initialise offload swap"
+swapon -p "$safe_prio" "$safe" || fail "cannot activate safe swap"
+dev_makeswap=$dev_start
+safe_active=1
+./swap_offload activate-no-discard "$offload" "$offload_prio" ||
+	fail "cannot activate offload-only swap"
+offload_active=1
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+	skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+[ -e "$cg/memory.max" ] ||
+	skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.swap.max"
+
+./swap_offload retained "$thp_size" "$cg/cgroup.procs" "$ready" \
+	"$touched" "$verified" &
+holder_pid=$!
+require_helper_file "$ready" "creating a PMD-sized anonymous huge folio"
+read -r reported_pid < "$ready"
+[ "$reported_pid" -eq "$holder_pid" ] ||
+	fail "retained helper reported the wrong pid"
+anon_huge_kib=$(awk '/AnonHugePages:/ { print $2 }' \
+	"/proc/$holder_pid/smaps_rollup")
+[ "${anon_huge_kib:-0}" -ge "$thp_kib" ] ||
+	skip "anonymous huge folio was not created"
+echo "$TCID: anonymous_huge_kib=$anon_huge_kib"
+
+# This is the only reclaim before the retained entry is made dirty.  It puts
+# the whole folio on offload-only swap so the later single-page fault leaves
+# sibling swap PTEs referring to the existing slot.
+echo "$thp_size swappiness=max" > "$cg/memory.reclaim" 2>/dev/null || :
+wait_offload_quiet || fail "offload device did not quiesce after setup"
+kill -0 "$holder_pid" || fail "retained helper died during setup reclaim"
+kill -USR1 "$holder_pid" || fail "cannot request the dirty-page transition"
+require_helper_file "$touched" "faulting and dirtying the retained entry"
+retained_kib=$(awk '/VmSwap:/ { print $2 }' "/proc/$holder_pid/status")
+[ "${retained_kib:-0}" -eq "$expected_retained_kib" ] ||
+	fail "retained $retained_kib KiB, expected $expected_retained_kib KiB"
+echo "$TCID: retained_kib=$retained_kib"
+
+wait_offload_quiet || fail "offload device did not quiesce before pressure"
+ordinary_before=$(written_sectors)
+refused_before=$(vmstat_value swpout_offload_refused)
+echo $((thp_size / 2)) > "$cg/memory.high" ||
+	fail "ordinary pressure reclaim failed"
+wait_offload_quiet || fail "offload device did not quiesce after pressure"
+ordinary_after=$(written_sectors)
+refused_after=$(vmstat_value swpout_offload_refused)
+ordinary_writes=$((ordinary_after - ordinary_before))
+refused_pages=$((refused_after - refused_before))
+echo "$TCID: ordinary_sectors=$ordinary_writes refused_pages=$refused_pages"
+[ "$ordinary_writes" -eq 0 ] ||
+	fail "ordinary reclaim wrote $ordinary_writes offload sectors"
+[ "$refused_pages" -ge "$expected_refused" ] ||
+	fail "ordinary reclaim refused $refused_pages pages, expected at least $expected_refused"
+kill -0 "$holder_pid" || fail "retained helper died after refused write"
+rm -f "$verified"
+kill -USR2 "$holder_pid" || fail "cannot request dirty-byte verification"
+require_helper_file "$verified" "verifying the dirty retained byte"
+
+# Proactive reclaim must still be able to rewrite the retained PMD-sized slot.
+# Repeated requests make the test insensitive to a short-lived writeback
+# collision; the backing-sector delta still requires exactly one folio write.
+wait_offload_quiet || fail "offload device did not quiesce before recovery"
+recovery_before=$(written_sectors)
+echo max > "$cg/memory.high"
+for _ in $(seq 1 8); do
+	echo "$thp_size swappiness=max" > "$cg/memory.reclaim" 2>/dev/null || :
+done
+wait_offload_quiet || fail "offload device did not quiesce after recovery"
+recovery_after=$(written_sectors)
+recovery_writes=$((recovery_after - recovery_before))
+echo "$TCID: recovery_sectors=$recovery_writes"
+[ "$recovery_writes" -eq "$thp_sectors" ] ||
+	fail "proactive recovery wrote $recovery_writes sectors, expected $thp_sectors"
+kill -0 "$holder_pid" || fail "retained helper died during recovery"
+rm -f "$verified"
+kill -ALRM "$holder_pid" || fail "cannot request full data verification"
+require_helper_file "$verified" "verifying all retained data"
+kill -0 "$holder_pid" || fail "retained helper failed full data verification"
+
+kill "$holder_pid" || fail "cannot stop retained helper"
+if wait "$holder_pid"; then
+	:
+else
+	status=$?
+	[ "$status" -eq 143 ] || fail "retained helper exited with status $status"
+fi
+holder_pid=""
+swapoff "$offload" || fail "cannot deactivate offload-only swap"
+offload_active=0
+swapoff "$safe" || fail "cannot deactivate safe swap"
+safe_active=0
+dev_makeswap=-1
+dmsetup --noudevsync --noudevrules remove "$dm_name" ||
+	fail "cannot remove delayed offload device"
+dm_active=0
+if [ "$dm_node_created" -eq 1 ]; then
+	rm -f "$offload" || fail "cannot remove delayed offload device node"
+	dm_node_created=0
+fi
+rmdir "$cg" || fail "cannot remove test cgroup"
+cg=""
+cgroup_disable_memory_controller "$cgroup_root" ||
+	fail "cannot restore the cgroup memory controller"
+for i in $(seq "$dev_start" "$dev_end"); do
+	echo 1 > "/sys/block/zram${i}/reset" || fail "cannot reset zram$i"
+	if [ "$sys_control" -eq 1 ]; then
+		echo "$i" > /sys/class/zram-control/hot_remove ||
+			fail "cannot remove zram$i"
+	fi
+done
+if [ "$module_load" -eq 1 ]; then
+	rmmod zram || fail "cannot unload zram"
+fi
+dev_end=-1
+if [ -n "$zswap_enabled" ]; then
+	echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled ||
+		fail "cannot restore zswap"
+	zswap_enabled=""
+fi
+rm -rf "$tmp" || fail "cannot remove temporary files"
+tmp=""
+
+echo "$TCID: [PASS]"
diff --git a/tools/testing/selftests/zram/zram_lib.sh b/tools/testing/selftests/zram/zram_lib.sh
index 0d44d83888f9d..f14966d3820d8 100755
--- a/tools/testing/selftests/zram/zram_lib.sh
+++ b/tools/testing/selftests/zram/zram_lib.sh
@@ -17,6 +17,10 @@ kernel_version=`uname -r | cut -d'.' -f1,2`
 kernel_major=${kernel_version%.*}
 kernel_minor=${kernel_version#*.}
 
+# Whether this test enabled the memory controller on its cgroup parent.  Tests
+# must leave a delegation which was already present alone.
+cgroup_memory_controller_enabled=0
+
 trap INT
 
 check_prereqs()
@@ -30,6 +34,30 @@ check_prereqs()
 	fi
 }
 
+cgroup_enable_memory_controller()
+{
+	local cgroup_root=$1
+
+	if grep -qw memory "$cgroup_root/cgroup.subtree_control"; then
+		return 0
+	fi
+
+	if ! echo +memory > "$cgroup_root/cgroup.subtree_control"; then
+		return 1
+	fi
+
+	cgroup_memory_controller_enabled=1
+}
+
+cgroup_disable_memory_controller()
+{
+	local cgroup_root=$1
+
+	[ "$cgroup_memory_controller_enabled" -eq 1 ] || return 0
+	echo -memory > "$cgroup_root/cgroup.subtree_control" || return 1
+	cgroup_memory_controller_enabled=0
+}
+
 kernel_gte()
 {
 	major=${1%.*}

3/2 2026/09/14 12:10 llm patch-evaluator
5m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 31065
							output: 17
							thoughts: 1200
						

Thoughts:
**Evaluation of the Patch Series for Syzkaller Fuzzing**

Alright, let's break this down. My initial assessment is that this patch series introduces changes that significantly impact core memory management, making it a prime candidate for fuzzing with syzkaller.

First, I need to look at the scope. The touched files are spread across critical areas: swap, memory management, cgroups, and even some DRM drivers that might interact with these core subsystems. This immediately raises my interest.

The core of the patch introduces a new `SWAP_FLAG_OFFLOAD_ONLY` for `swapon`, enabling a swap area to be used *exclusively* for proactive reclaim. This is where things get interesting. We're talking about new control paths, conditional logic, and potentially complex interactions with the existing swap and reclaim mechanisms. It is adding a new flag for `swapon`, which is a syscall.

The changes touch many files. We have new flags and functions added to `include/linux/swap.h`, as well as updates to `mm/memcontrol.c`, `mm/page_io.c`, `mm/swapfile.c`, `mm/vmscan.c`, and `mm/vmstat.c`. This means that we have modifications in swap allocation, page reclaim, and even accounting.

The critical thing is that this patch directly interacts with executable code. The `swapon` syscall is accessible from user space (syzkaller can test it), and proactive reclaim via cgroups or sysfs nodes introduces new reclaim paths. The changes are intertwined with swap allocation and writeout, making them very relevant for fuzzing.

My analysis aligns with Section 3, regarding Core Subsystems, which states that any modifications in memory management, like this patch series, are WorthFuzzing.

My conclusion is: **WorthFuzzing = true.**

I also need to check the configurations, `EnableConfigs`, to make sure that the required configs (e.g., swap, cgroups) are enabled to trigger the specific codepaths added by this patch series. I will search through the Kconfig files or read the relevant files for the minimal set of dependencies.



4/2 2026/09/14 12:16 tool grepper
0m Args:
{
  "Expression": "config SWAP"
}

Results:
{
  "Output": "arch/mips/Kconfig=1261=config MIPS_SPRAM\n--\narch/mips/Kconfig-1263-\narch/mips/Kconfig:1264:config SWAP_IO_SPACE\narch/mips/Kconfig-1265-\tbool\n--\narch/sh/Kconfig=127=config IO_TRAPPED\n--\narch/sh/Kconfig-129-\narch/sh/Kconfig:130:config SWAP_IO_SPACE\narch/sh/Kconfig-131-\tbool\n--\nmm/Kconfig=9=config ARCH_NO_SWAP\n--\nmm/Kconfig-11-\nmm/Kconfig:12:menuconfig SWAP\nmm/Kconfig-13-\tbool \"Support for paging of anonymous memory (swap)\"\n"
}

5/2 2026/09/14 12:16 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 32951
							output: 31
							thoughts: 16
						

Thoughts:
Here's my attempt at summarizing the thought process, tailored for an expert audience:

**Diving into the Configuration Landscape**

Okay, so I'm thinking, let's zero in on `mm/Kconfig` around line 12. My mind's immediately going to be running through the usual suspects in this area – what kind of memory management options are probably being defined there? Given it's a `Kconfig` file, I'm expecting to see dependencies, help text, maybe some default values related to memory management features. I need to get a clear picture of what the maintainers have decided to expose to users, and what’s hidden. Specifically, I’ll be looking for any hints about potential configuration interactions within the memory subsystem, perhaps even some obscure options that I’ve overlooked in the past. This is where the real nuance lies, in the details of the configuration options. The goal is a deeper understanding, not just a surface-level overview.



6/2 2026/09/14 12:16 tool read-file
0m Args:
{
  "File": "mm/Kconfig",
  "FirstLine": 12,
  "LineCount": 40
}

Results:
{
  "Contents": "  12:\tmenuconfig SWAP\n  13:\t\tbool \"Support for paging of anonymous memory (swap)\"\n  14:\t\tdepends on MMU \u0026\u0026 BLOCK \u0026\u0026 !ARCH_NO_SWAP\n  15:\t\tdefault y\n  16:\t\thelp\n  17:\t\t  This option allows you to choose whether you want to have support\n  18:\t\t  for so called swap devices or swap files in your kernel that are\n  19:\t\t  used to provide more virtual memory than the actual RAM present\n  20:\t\t  in your computer.  If unsure say Y.\n  21:\t\n  22:\tconfig ZSWAP\n  23:\t\tbool \"Compressed cache for swap pages\"\n  24:\t\tdepends on SWAP\n  25:\t\tselect CRYPTO\n  26:\t\tselect ZSMALLOC\n  27:\t\thelp\n  28:\t\t  A lightweight compressed cache for swap pages.  It takes\n  29:\t\t  pages that are in the process of being swapped out and attempts to\n  30:\t\t  compress them into a dynamically allocated RAM-based memory pool.\n  31:\t\t  This can result in a significant I/O reduction on swap device and,\n  32:\t\t  in the case where decompressing from RAM is faster than swap device\n  33:\t\t  reads, can also improve workload performance.\n  34:\t\n  35:\tconfig ZSWAP_DEFAULT_ON\n  36:\t\tbool \"Enable the compressed cache for swap pages by default\"\n  37:\t\tdepends on ZSWAP\n  38:\t\thelp\n  39:\t\t  If selected, the compressed cache for swap pages will be enabled\n  40:\t\t  at boot, otherwise it will be disabled.\n  41:\t\n  42:\t\t  The selection made here can be overridden by using the kernel\n  43:\t\t  command line 'zswap.enabled=' option.\n  44:\t\n  45:\tconfig ZSWAP_SHRINKER_DEFAULT_ON\n  46:\t\tbool \"Shrink the zswap pool on memory pressure\"\n  47:\t\tdepends on ZSWAP\n  48:\t\tdefault n\n  49:\t\thelp\n  50:\t\t  If selected, the zswap shrinker will be enabled, and the pages\n  51:\t\t  stored in the zswap pool will become available for reclaim (i.e\n"
}

7/2 2026/09/14 12:16 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 33777
							output: 26
							thoughts: 63
						

Thoughts:
**Examining the Memory Management Configuration**

Okay, let's dive in. I need to understand the configuration related to swap and memory control. First, it looks like `CONFIG_SWAP` is essential – no surprises there, given its direct link to `mm/swapfile.c` and overall swap functionality. Now, I need to see what's connected to memcontrol as it involves the cgroup swap tracking in `memcontrol.c`. The next step is clearly to determine whether `memcontrol.c` utilizes `MEMCG_SWAP` or just `MEMCG` itself. I need this to create the full picture of the configuration.



8/2 2026/09/14 12:16 tool grepper
0m Args:
{
  "Expression": "CONFIG_MEMCG",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "\nFull output is too long, showing 500 out of 582 lines.\nUse more precise expression if possible.\n\nmm/Makefile=104=obj-$(CONFIG_LIVEUPDATE_MEMFD) += memfd_luo.o\nmm/Makefile:105:obj-$(CONFIG_MEMCG_V1) += memcontrol-v1.o\nmm/Makefile:106:obj-$(CONFIG_MEMCG) += memcontrol.o vmpressure.o\nmm/Makefile-107-ifdef CONFIG_BPF_SYSCALL\nmm/Makefile:108:obj-$(CONFIG_MEMCG) += bpf_memcontrol.o\nmm/Makefile-109-endif\n--\nmm/damon/sysfs-common.c=106=static bool damon_sysfs_memcg_path_eq(struct mem_cgroup *memcg,\n--\nmm/damon/sysfs-common.c-108-{\nmm/damon/sysfs-common.c:109:#ifdef CONFIG_MEMCG\nmm/damon/sysfs-common.c-110-\tcgroup_path(memcg-\u003ecss.cgroup, memcg_path_buf, PATH_MAX);\n--\nmm/damon/sysfs-common.c-112-\t\treturn true;\nmm/damon/sysfs-common.c:113:#endif /* CONFIG_MEMCG */\nmm/damon/sysfs-common.c-114-\treturn false;\n--\nmm/debug.c=70=static void __dump_folio(const struct folio *folio, const struct page *page,\n--\nmm/debug.c-96-\nmm/debug.c:97:#ifdef CONFIG_MEMCG\nmm/debug.c-98-\tif (folio-\u003ememcg_data)\n--\nmm/debug.c=171=void dump_mm(const struct mm_struct *mm)\n--\nmm/debug.c-184-#endif\nmm/debug.c:185:#ifdef CONFIG_MEMCG\nmm/debug.c-186-\t\t\"owner %px \"\n--\nmm/debug.c-213-#endif\nmm/debug.c:214:#ifdef CONFIG_MEMCG\nmm/debug.c-215-\t\tmm-\u003eowner,\n--\nmm/folio.c=1112=void folio_batch_remove_exceptionals(struct folio_batch *fbatch)\n--\nmm/folio.c-1123-\nmm/folio.c:1124:#ifdef CONFIG_MEMCG\nmm/folio.c-1125-static void lruvec_reparent_lru(struct lruvec *child_lruvec,\n--\nmm/huge_memory.c=3693=static void __split_folio_to_order(struct folio *folio, int old_order,\n--\nmm/huge_memory.c-3799-\t\t\tfolio_set_idle(new_folio);\nmm/huge_memory.c:3800:#ifdef CONFIG_MEMCG\nmm/huge_memory.c-3801-\t\tnew_folio-\u003ememcg_data = folio-\u003ememcg_data;\n--\nmm/hwpoison-inject.c=45=static int hwpoison_filter_flags(struct page *p)\n--\nmm/hwpoison-inject.c-66- */\nmm/hwpoison-inject.c:67:#ifdef CONFIG_MEMCG\nmm/hwpoison-inject.c-68-static u64 hwpoison_filter_memcg;\n--\nmm/hwpoison-inject.c=162=static int __init pfn_inject_init(void)\n--\nmm/hwpoison-inject.c-191-\nmm/hwpoison-inject.c:192:#ifdef CONFIG_MEMCG\nmm/hwpoison-inject.c-193-\tdebugfs_create_u64(\"corrupt-filter-memcg\", 0600, hwpoison_dir,\n--\nmm/list_lru.c=29=static inline void unlock_list_lru(struct list_lru_one *l, bool irq_off,\n--\nmm/list_lru.c-39-\nmm/list_lru.c:40:#ifdef CONFIG_MEMCG\nmm/list_lru.c-41-static LIST_HEAD(memcg_list_lrus);\n--\nmm/list_lru.c=141=lock_list_lru_of_memcg(struct list_lru *lru, int nid,\n--\nmm/list_lru.c-150-}\nmm/list_lru.c:151:#endif /* CONFIG_MEMCG */\nmm/list_lru.c-152-\n--\nmm/list_lru.c=413=unsigned long list_lru_walk_node(struct list_lru *lru, int nid,\n--\nmm/list_lru.c-421-\nmm/list_lru.c:422:#ifdef CONFIG_MEMCG\nmm/list_lru.c-423-\tif (*nr_to_walk \u003e 0 \u0026\u0026 list_lru_memcg_aware(lru)) {\n--\nmm/list_lru.c=451=static void init_one_lru(struct list_lru *lru, struct list_lru_one *l)\n--\nmm/list_lru.c-461-\nmm/list_lru.c:462:#ifdef CONFIG_MEMCG\nmm/list_lru.c-463-static struct list_lru_memcg *memcg_init_list_lru_one(struct list_lru *lru, gfp_t gfp)\n--\nmm/list_lru.c=659=static void memcg_destroy_list_lru(struct list_lru *lru)\n--\nmm/list_lru.c-661-}\nmm/list_lru.c:662:#endif /* CONFIG_MEMCG */\nmm/list_lru.c-663-\nmm/list_lru.c=664=int __list_lru_init(struct list_lru *lru, bool memcg_aware, struct shrinker *shrinker)\n--\nmm/list_lru.c-667-\nmm/list_lru.c:668:#ifdef CONFIG_MEMCG\nmm/list_lru.c-669-\t/*\n--\nmm/list_lru.c=697=void list_lru_destroy(struct list_lru *lru)\n--\nmm/list_lru.c-708-\nmm/list_lru.c:709:#ifdef CONFIG_MEMCG\nmm/list_lru.c-710-\tlru-\u003eshrinker_id = -1;\n--\nmm/memcontrol-v1.c=1123=void memcg1_oom_finish(struct mem_cgroup *memcg, bool locked)\n--\nmm/memcontrol-v1.c-1130- * cgroup v1 userspace vmpressure interface (memory.pressure_level /\nmm/memcontrol-v1.c:1131: * cgroup.event_control). Kept here so v2-only kernels (CONFIG_MEMCG_V1=n)\nmm/memcontrol-v1.c-1132- * drop the whole eventfd accumulator, its work item, and the per-memcg\n--\nmm/memcontrol-v1.h=25=struct mem_cgroup *mem_cgroup_private_id_get_online(struct mem_cgroup *memcg,\n--\nmm/memcontrol-v1.h-28-/* Cgroup v1-specific declarations */\nmm/memcontrol-v1.h:29:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol-v1.h-30-\n--\nmm/memcontrol-v1.h=86=extern struct cftype mem_cgroup_legacy_files[];\nmm/memcontrol-v1.h-87-\nmm/memcontrol-v1.h:88:#else\t/* CONFIG_MEMCG_V1 */\nmm/memcontrol-v1.h-89-\n--\nmm/memcontrol-v1.h=118=static inline void memcg1_uncharge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages) {}\nmm/memcontrol-v1.h-119-\nmm/memcontrol-v1.h:120:#endif\t/* CONFIG_MEMCG_V1 */\nmm/memcontrol-v1.h-121-\n--\nmm/memcontrol.c=216=static inline struct obj_cgroup *__memcg_reparent_objcgs(struct mem_cgroup *memcg,\n--\nmm/memcontrol.c-235-\nmm/memcontrol.c:236:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-237-static void __mem_cgroup_flush_stats(struct mem_cgroup *memcg, bool force);\n--\nmm/memcontrol.c=551=unsigned long lruvec_page_state_local(struct lruvec *lruvec,\n--\nmm/memcontrol.c-573-\nmm/memcontrol.c:574:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-575-static void __mod_memcg_lruvec_state(struct mem_cgroup_per_node *pn,\n--\nmm/memcontrol.c=599=static const unsigned int memcg_vm_event_stat[] = {\nmm/memcontrol.c:600:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-601-\tPGPGIN,\n--\nmm/memcontrol.c=674=struct memcg_vmstats {\n--\nmm/memcontrol.c-678-\nmm/memcontrol.c:679:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-680-\t/* Non-hierarchical (CPU aggregated) page state \u0026 events */\n--\nmm/memcontrol.c-695- * memory.stat and memory.numa_stat, and by reparenting on offline, all of which\nmm/memcontrol.c:696: * are v1-only. If the kernel is built without CONFIG_MEMCG_V1, or if the boot\nmm/memcontrol.c-697- * param cgroup_no_v1= has blocked the memory controller from v1 hierarchies,\n--\nmm/memcontrol.c=700=static long *memcg_state_local_array(struct mem_cgroup *memcg)\nmm/memcontrol.c-701-{\nmm/memcontrol.c:702:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-703-\tif (!cgroup1_ssid_disabled(memory_cgrp_id))\n--\nmm/memcontrol.c=709=static unsigned long *memcg_events_local_array(struct mem_cgroup *memcg)\nmm/memcontrol.c-710-{\nmm/memcontrol.c:711:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-712-\tif (!cgroup1_ssid_disabled(memory_cgrp_id))\n--\nmm/memcontrol.c=938=void mod_memcg_state(struct mem_cgroup *memcg, enum memcg_stat_item idx,\n--\nmm/memcontrol.c-950-\nmm/memcontrol.c:951:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-952-/* idx can be of type enum memcg_stat_item or node_stat_item. */\n--\nmm/memcontrol.c=1126=bool memcg_vm_event_item_valid(enum vm_event_item idx)\n--\nmm/memcontrol.c-1133-\nmm/memcontrol.c:1134:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-1135-unsigned long memcg_events_local(struct mem_cgroup *memcg, int event)\n--\nmm/memcontrol.c=1718=unsigned long memcg_page_state_output(struct mem_cgroup *memcg, int item)\n--\nmm/memcontrol.c-1723-\nmm/memcontrol.c:1724:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-1725-unsigned long memcg_page_state_local_output(struct mem_cgroup *memcg, int item)\n--\nmm/memcontrol.c=1744=static void memcg_stat_format(struct mem_cgroup *memcg, struct seq_buf *s)\n--\nmm/memcontrol.c-1790-\tfor (i = 0; i \u003c ARRAY_SIZE(memcg_vm_event_stat); i++) {\nmm/memcontrol.c:1791:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-1792-\t\tif (memcg_vm_event_stat[i] == PGPGIN ||\n--\nmm/memcontrol.c=1842=void mem_cgroup_print_oom_meminfo(struct mem_cgroup *memcg)\n--\nmm/memcontrol.c-1863-\t\t\tatomic_long_read(\u0026memcg-\u003ememory_events[MEMCG_SWAP_MAX]));\nmm/memcontrol.c:1864:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-1865-\telse {\n--\nmm/memcontrol.c=2910=static void commit_charge(struct folio *folio, struct obj_cgroup *objcg)\n--\nmm/memcontrol.c-2922-\nmm/memcontrol.c:2923:#ifdef CONFIG_MEMCG_NMI_SAFETY_REQUIRES_ATOMIC\nmm/memcontrol.c-2924-static inline void account_slab_nmi_safe(struct mem_cgroup *memcg,\n--\nmm/memcontrol.c=3097=__always_inline struct obj_cgroup *current_obj_cgroup(void)\n--\nmm/memcontrol.c-3102-\nmm/memcontrol.c:3103:\tif (IS_ENABLED(CONFIG_MEMCG_NMI_UNSAFE) \u0026\u0026 in_nmi())\nmm/memcontrol.c-3104-\t\treturn NULL;\n--\nmm/memcontrol.c=3143=struct obj_cgroup *get_obj_cgroup_from_folio(struct folio *folio)\n--\nmm/memcontrol.c-3153-\nmm/memcontrol.c:3154:#ifdef CONFIG_MEMCG_NMI_SAFETY_REQUIRES_ATOMIC\nmm/memcontrol.c-3155-static inline void account_kmem_nmi_safe(struct mem_cgroup *memcg, int val)\n--\nmm/memcontrol.c=4266=mem_cgroup_css_alloc(struct cgroup_subsys_state *parent_css)\n--\nmm/memcontrol.c-4286-\t\tpage_counter_init(\u0026memcg-\u003eswap, \u0026parent-\u003eswap, false);\nmm/memcontrol.c:4287:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-4288-\t\tWRITE_ONCE(memcg-\u003eswappiness, mem_cgroup_swappiness(parent));\n--\nmm/memcontrol.c-4300-\t\tpage_counter_init(\u0026memcg-\u003eswap, NULL, false);\nmm/memcontrol.c:4301:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-4302-\t\tpage_counter_init(\u0026memcg-\u003ekmem, NULL, false);\n--\nmm/memcontrol.c=4468=static void mem_cgroup_css_reset(struct cgroup_subsys_state *css)\n--\nmm/memcontrol.c-4478-#endif\nmm/memcontrol.c:4479:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-4480-\tpage_counter_set_max(\u0026memcg-\u003ekmem, PAGE_COUNTER_MAX);\n--\nmm/memcontrol.c=4510=static void mem_cgroup_stat_aggregate(struct aggregate_control *ac)\n--\nmm/memcontrol.c-4545-\nmm/memcontrol.c:4546:#ifdef CONFIG_MEMCG_NMI_SAFETY_REQUIRES_ATOMIC\nmm/memcontrol.c-4547-static void flush_nmi_stats(struct mem_cgroup *memcg, struct mem_cgroup *parent)\n--\nmm/memcontrol.c=5215=struct cgroup_subsys memory_cgrp_subsys = {\n--\nmm/memcontrol.c-5226-\t.dfl_cftypes = memory_files,\nmm/memcontrol.c:5227:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-5228-\t.legacy_cftypes = mem_cgroup_legacy_files,\n--\nmm/memcontrol.c=5791=static void __init memcg_struct_check(void)\n--\nmm/memcontrol.c-5806-\t\t\t\t      private_id_ref);\nmm/memcontrol.c:5807:#ifdef CONFIG_MEMCG_NMI_SAFETY_REQUIRES_ATOMIC\nmm/memcontrol.c-5808-\tCACHELINE_ASSERT_GROUP_MEMBER(struct mem_cgroup, memcg_write_hot,\n--\nmm/memcontrol.c-5875-\t\t\t\t      memcg_pn_write_hot, iter);\nmm/memcontrol.c:5876:#ifdef CONFIG_MEMCG_NMI_SAFETY_REQUIRES_ATOMIC\nmm/memcontrol.c-5877-\tCACHELINE_ASSERT_GROUP_MEMBER(struct mem_cgroup_per_node,\n--\nmm/memcontrol.c=6409=static int __init mem_cgroup_swap_init(void)\n--\nmm/memcontrol.c-6414-\tWARN_ON(cgroup_add_dfl_cftypes(\u0026memory_cgrp_subsys, swap_files));\nmm/memcontrol.c:6415:#ifdef CONFIG_MEMCG_V1\nmm/memcontrol.c-6416-\tWARN_ON(cgroup_add_legacy_cftypes(\u0026memory_cgrp_subsys, memsw_files));\n--\nmm/page_alloc.c=1015=static inline bool page_expected_state(struct page *page,\n--\nmm/page_alloc.c-1022-\t\t\tpage_ref_count(page) |\nmm/page_alloc.c:1023:#ifdef CONFIG_MEMCG\nmm/page_alloc.c-1024-\t\t\tpage-\u003ememcg_data |\n--\nmm/page_alloc.c=1033=static const char *page_bad_reason(struct page *page, unsigned long flags)\n--\nmm/page_alloc.c-1048-\t}\nmm/page_alloc.c:1049:#ifdef CONFIG_MEMCG\nmm/page_alloc.c-1050-\tif (unlikely(page-\u003ememcg_data))\n--\nmm/page_counter.c=293=int page_counter_memparse(const char *buf, const char *max,\n--\nmm/page_counter.c-313-\nmm/page_counter.c:314:#if IS_ENABLED(CONFIG_MEMCG) || IS_ENABLED(CONFIG_CGROUP_DMEM)\nmm/page_counter.c-315-/*\n--\nmm/page_counter.c=453=void page_counter_calculate_protection(struct page_counter *root,\n--\nmm/page_counter.c-493-}\nmm/page_counter.c:494:#endif /* CONFIG_MEMCG || CONFIG_CGROUP_DMEM */\n--\nmm/page_io.c=204=int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)\n--\nmm/page_io.c-277-\nmm/page_io.c:278:#if defined(CONFIG_MEMCG) \u0026\u0026 defined(CONFIG_BLK_CGROUP)\nmm/page_io.c-279-static struct cgroup_subsys_state *folio_memcg_blkg_css(struct folio *folio)\n--\nmm/page_io.c=321=static void bio_associate_blkg_from_folio(struct bio *bio, struct folio *folio)\n--\nmm/page_io.c-323-}\nmm/page_io.c:324:#endif /* CONFIG_MEMCG \u0026\u0026 CONFIG_BLK_CGROUP */\nmm/page_io.c-325-\n--\nmm/page_owner.c=476=void pagetypeinfo_showmixedcount_print(struct seq_file *m,\n--\nmm/page_owner.c-552-\nmm/page_owner.c:553:#ifdef CONFIG_MEMCG\nmm/page_owner.c-554-/*\n--\nmm/percpu-internal.h=35=struct pcpuobj_ext {\nmm/percpu-internal.h:36:#ifdef CONFIG_MEMCG\nmm/percpu-internal.h-37-\tstruct obj_cgroup\t*cgroup;\n--\nmm/percpu-internal.h-43-\nmm/percpu-internal.h:44:#if defined(CONFIG_MEMCG) || defined(CONFIG_MEM_ALLOC_PROFILING)\nmm/percpu-internal.h-45-#define NEED_PCPUOBJ_EXT\n--\nmm/percpu-internal.h=153=static inline size_t pcpu_obj_full_size(size_t size)\n--\nmm/percpu-internal.h-156-\nmm/percpu-internal.h:157:#ifdef CONFIG_MEMCG\nmm/percpu-internal.h-158-\tif (!mem_cgroup_kmem_disabled())\n--\nmm/percpu.c=1592=static struct pcpu_chunk *pcpu_chunk_addr_search(void *addr)\n--\nmm/percpu.c-1612-\nmm/percpu.c:1613:#ifdef CONFIG_MEMCG\nmm/percpu.c-1614-static bool pcpu_memcg_pre_alloc_hook(size_t size, gfp_t gfp,\n--\nmm/percpu.c=1653=static void pcpu_memcg_free_hook(struct pcpu_chunk *chunk, int off, size_t size)\n--\nmm/percpu.c-1674-\nmm/percpu.c:1675:#else /* CONFIG_MEMCG */\nmm/percpu.c-1676-static bool\n--\nmm/percpu.c=1688=static void pcpu_memcg_free_hook(struct pcpu_chunk *chunk, int off, size_t size)\n--\nmm/percpu.c-1690-}\nmm/percpu.c:1691:#endif /* CONFIG_MEMCG */\nmm/percpu.c-1692-\n--\nmm/shmem.c=739=static struct obj_cgroup *shmem_get_and_clear_objcg(struct shmem_inode_info *info)\n--\nmm/shmem.c-747-\nmm/shmem.c:748:#ifdef CONFIG_MEMCG\nmm/shmem.c-749-static struct obj_cgroup *\n--\nmm/shrinker.c=11=DEFINE_MUTEX(shrinker_mutex);\nmm/shrinker.c-12-\nmm/shrinker.c:13:#ifdef CONFIG_MEMCG\nmm/shrinker.c-14-static int shrinker_nr_max;\n--\nmm/shrinker.c=336=static long add_nr_deferred_memcg(long nr, int nid, struct shrinker *shrinker,\n--\nmm/shrinker.c-340-}\nmm/shrinker.c:341:#endif /* CONFIG_MEMCG */\nmm/shrinker.c-342-\n--\nmm/shrinker.c=378=static unsigned long do_shrink_slab(struct shrink_control *shrinkctl,\n--\nmm/shrinker.c-476-\nmm/shrinker.c:477:#ifdef CONFIG_MEMCG\nmm/shrinker.c-478-static unsigned long shrink_slab_memcg(gfp_t gfp_mask, int nid,\n--\nmm/shrinker.c-597-}\nmm/shrinker.c:598:#else /* !CONFIG_MEMCG */\nmm/shrinker.c-599-static unsigned long shrink_slab_memcg(gfp_t gfp_mask, int nid,\n--\nmm/shrinker.c-603-}\nmm/shrinker.c:604:#endif /* CONFIG_MEMCG */\nmm/shrinker.c-605-\n--\nmm/shrinker.c=690=struct shrinker *shrinker_alloc(unsigned int flags, const char *fmt, ...)\n--\nmm/shrinker.c-728-\t *  - non-memcg-aware shrinkers\nmm/shrinker.c:729:\t *  - !CONFIG_MEMCG\nmm/shrinker.c-730-\t *  - memcg is disabled by kernel command line\n--\nmm/slab.h=140=SLAB_MATCH(_refcount, __page_refcount);\nmm/slab.h:141:#ifdef CONFIG_MEMCG\nmm/slab.h-142-SLAB_MATCH(memcg_data, obj_exts);\n--\nmm/slab.h=604=struct slabobj_ext {\n--\nmm/slab.h-609-\tunion {\nmm/slab.h:610:#ifdef CONFIG_MEMCG\nmm/slab.h-611-\t\tstruct obj_cgroup *_objcg;\n--\nmm/slab.h=629=static inline bool slab_obj_ext_has_codetag(void)\n--\nmm/slab.h-634-\nmm/slab.h:635:#ifdef CONFIG_MEMCG\nmm/slab.h-636-static inline bool cache_needs_objcg(struct kmem_cache *cache)\n--\nmm/slab.h=713=static inline unsigned long slab_obj_exts(struct slab *slab)\n--\nmm/slab.h-716-\nmm/slab.h:717:#ifdef CONFIG_MEMCG\nmm/slab.h-718-\t/*\n--\nmm/slab.h=771=slab_obj_ext(struct kmem_cache *s, struct slab *slab, unsigned long obj_exts,\n--\nmm/slab.h-795-\nmm/slab.h:796:#ifdef CONFIG_MEMCG\nmm/slab.h-797-static inline struct obj_cgroup *\n--\nmm/slab.h=854=static inline enum node_stat_item cache_vmstat_idx(struct kmem_cache *s)\n--\nmm/slab.h-859-\nmm/slab.h:860:#ifdef CONFIG_MEMCG\nmm/slab.h-861-bool __memcg_slab_post_alloc_hook(struct kmem_cache *s, struct list_lru *lru,\n--\nmm/slab_common.c=815=EXPORT_SYMBOL(kmalloc_size_roundup);\n--\nmm/slab_common.c-822-\nmm/slab_common.c:823:#ifdef CONFIG_MEMCG\nmm/slab_common.c-824-#define KMALLOC_CGROUP_NAME(sz)\t.name[KMALLOC_CGROUP] = \"kmalloc-cg-\" #sz,\n--\nmm/slab_common.c=963=new_kmalloc_cache(int idx, enum kmalloc_cache_type type)\n--\nmm/slab_common.c-971-\t\tflags |= SLAB_RECLAIM_ACCOUNT;\nmm/slab_common.c:972:\t} else if (IS_ENABLED(CONFIG_MEMCG) \u0026\u0026 (type == KMALLOC_CGROUP)) {\nmm/slab_common.c-973-\t\tif (mem_cgroup_kmem_disabled()) {\n--\nmm/slab_common.c=1027=void __init create_kmalloc_caches(void)\n--\nmm/slab_common.c-1032-\t/*\nmm/slab_common.c:1033:\t * Including KMALLOC_CGROUP if CONFIG_MEMCG defined\nmm/slab_common.c-1034-\t */\n--\nmm/slub.c=2164=int alloc_slab_obj_exts(struct slab *slab, struct kmem_cache *s,\n--\nmm/slub.c-2217-\tnew_exts = (unsigned long)vec;\nmm/slub.c:2218:#ifdef CONFIG_MEMCG\nmm/slub.c-2219-\tnew_exts |= MEMCG_DATA_OBJEXTS;\n--\nmm/slub.c=2285=static void alloc_slab_obj_exts_early(struct kmem_cache *s, struct slab *slab)\n--\nmm/slub.c-2301-\nmm/slub.c:2302:#ifdef CONFIG_MEMCG\nmm/slub.c-2303-\t\tobj_exts |= MEMCG_DATA_OBJEXTS;\n--\nmm/slub.c-2317-\nmm/slub.c:2318:#ifdef CONFIG_MEMCG\nmm/slub.c-2319-\t\tobj_exts |= MEMCG_DATA_OBJEXTS;\n--\nmm/slub.c=2495=static inline void slab_obj_ext_has_codetag_init(void)\n--\nmm/slub.c-2501-\nmm/slub.c:2502:#ifdef CONFIG_MEMCG\nmm/slub.c-2503-\n--\nmm/slub.c=2553=bool memcg_slab_post_charge(void *p, gfp_t flags)\n--\nmm/slub.c-2610-\nmm/slub.c:2611:#else /* CONFIG_MEMCG */\nmm/slub.c-2612-static inline bool memcg_slab_post_alloc_hook(struct kmem_cache *s,\n--\nmm/slub.c=2625=static inline bool memcg_slab_post_charge(void *p, gfp_t flags)\n--\nmm/slub.c-2628-}\nmm/slub.c:2629:#endif /* CONFIG_MEMCG */\nmm/slub.c-2630-\n--\nmm/slub.c=6493=void slab_free(struct kmem_cache *s, struct slab *slab, void *object,\n--\nmm/slub.c-6508-\nmm/slub.c:6509:#ifdef CONFIG_MEMCG\nmm/slub.c-6510-/* Do not inline the rare memcg charging failed path into the allocation path */\n--\nmm/swap.h=50=struct swap_cluster_info {\n--\nmm/swap.h-60-\tunsigned int *extend_table;\t/* For large swap count, protected by ci-\u003elock */\nmm/swap.h:61:#ifdef CONFIG_MEMCG\nmm/swap.h-62-\tstruct swap_memcg_table *memcg_table;\t/* Swap table entries' cgroup record */\n--\nmm/swap.h=85=static inline int mem_cgroup_swappiness(struct mem_cgroup *memcg)\nmm/swap.h-86-{\nmm/swap.h:87:#ifdef CONFIG_MEMCG_V1\nmm/swap.h-88-\tif (!cgroup_subsys_on_dfl(memory_cgrp_subsys) \u0026\u0026\n--\nmm/swap_table.h=296=static inline void __swap_table_clear_zero(struct swap_cluster_info *ci,\n--\nmm/swap_table.h-310-\nmm/swap_table.h:311:#ifdef CONFIG_MEMCG\nmm/swap_table.h-312-static inline void __swap_cgroup_set(struct swap_cluster_info *ci,\n--\nmm/swapfile.c=460=static void swap_cluster_free_table(struct swap_cluster_info *ci)\n--\nmm/swapfile.c-463-\nmm/swapfile.c:464:#ifdef CONFIG_MEMCG\nmm/swapfile.c-465-\tkfree(ci-\u003ememcg_table);\n--\nmm/swapfile.c=482=static int swap_cluster_alloc_table(struct swap_cluster_info *ci, gfp_t gfp)\n--\nmm/swapfile.c-501-\nmm/swapfile.c:502:#ifdef CONFIG_MEMCG\nmm/swapfile.c-503-\tif (!mem_cgroup_disabled()) {\n--\nmm/swapfile.c=4082=int swap_dup_entry_direct(swp_entry_t entry)\n--\nmm/swapfile.c-4101-\nmm/swapfile.c:4102:#if defined(CONFIG_MEMCG) \u0026\u0026 defined(CONFIG_BLK_CGROUP)\nmm/swapfile.c-4103-static bool __has_usable_swap(void)\n--\nmm/vmscan.c=226=static int sc_swappiness(struct scan_control *sc, struct mem_cgroup *memcg)\n--\nmm/vmscan.c-232-\nmm/vmscan.c:233:#ifdef CONFIG_MEMCG\nmm/vmscan.c-234-\n--\nmm/vmscan.c=2830=static struct lruvec *get_lruvec(struct mem_cgroup *memcg, int nid)\n--\nmm/vmscan.c-2833-\nmm/vmscan.c:2834:#ifdef CONFIG_MEMCG\nmm/vmscan.c-2835-\tif (memcg) {\n--\nmm/vmscan.c=2988=static struct lru_gen_mm_list *get_mm_list(struct mem_cgroup *memcg)\n--\nmm/vmscan.c-2994-\n\n"
}

9/2 2026/09/14 12:16 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 42611
							output: 26
							thoughts: 540
						

Thoughts:
**My Analysis of Swap and Memory Cgroup Interactions**

Okay, let's dive into this code. First, I need to confirm something. I'm checking `include/linux/swap.h` to see if `mem_cgroup_get_nr_swap_pages_eligible` has a dependency on `CONFIG_MEMCG`. Looking at the header file, it seems like it does. The code is structured with an `#ifdef CONFIG_MEMCG` block. When `CONFIG_MEMCG` is enabled, there's a specific implementation with `mem_cgroup_get_nr_swap_pages_eligible` taking a `struct mem_cgroup *memcg` argument. Otherwise, it defaults to calling `get_nr_swap_pages_eligible()`. Got it.

Next, I'm considering `include/linux/vm_event_item.h`. I'm curious about the `SWPOUT_OFFLOAD_REFUSED` event and if it's tied to `CONFIG_VM_EVENT_COUNTERS` for the `proc vmstat` interface. My initial thought is it likely is, so let's check `mm/vmstat.c`. And yes, in `mm/vmstat.c`, I see a conditional inclusion: `#ifdef CONFIG_VM_EVENT_COUNTERS`, and within that block, `SWPOUT_OFFLOAD_REFUSED` is included in `vmstat_text[]`. That makes sense. Also, I see its usage in `count_vm_events(SWPOUT_OFFLOAD_REFUSED, ...)` within `mm/page_io.c`. This also means that both `mm/swapfile.c` and `mm/page_io.c` require `CONFIG_SWAP`.

Now, I'm shifting focus to `CONFIG_MEMCG`. Proactive reclaim via cgroup v2's `memory.reclaim` interface *definitely* relies on `CONFIG_MEMCG`. Now I'm mentally tracing where `memory.reclaim` is defined, and I'll trace to it...



10/2 2026/09/14 12:16 tool grepper
0m Args:
{
  "Expression": "memory\\.reclaim",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "mm/memcontrol-v1.c=1968=static int mem_cgroup_swappiness_write(struct cgroup_subsys_state *css,\n--\nmm/memcontrol-v1.c-1977-\t\tpr_info_once(\"Per memcg swappiness does not exist in cgroup v2. \"\nmm/memcontrol-v1.c:1978:\t\t\t     \"See memory.reclaim or memory.swap.max there\\n \");\nmm/memcontrol-v1.c-1979-\t\tWRITE_ONCE(memcg-\u003eswappiness, val);\n"
}

11/2 2026/09/14 12:16 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 43781
							output: 29
							thoughts: 15
						

Thoughts:
Here's my thought process, as if I were considering this myself:

**Tracing the `memory.reclaim` Handler**

Okay, I need to figure out where the `memory.reclaim` handler is actually registered. This typically involves digging into the kernel code or the relevant memory management subsystem. My immediate approach will be to leverage my knowledge of the system and start with the core modules. I'll probably look at the places where memory management events are handled. I expect there to be a central dispatcher or registration mechanism within the kernel's memory management code. So, I'll be looking for something that seems to manage events. I know where to look. I'll search for things like registration calls, handler tables, or event dispatchers related to memory or resource management. I'll be looking for keywords like "register_handler," "event_table," or even something specific to the reclaim process. I'll also think about where this handler might be _used_ -- that might give clues on where it's defined. My next step will be to dive into the codebase, ideally equipped with a powerful search tool and a good understanding of the system's architecture.



12/2 2026/09/14 12:16 tool grepper
0m Args:
{
  "Expression": "\\.name = \"reclaim\"",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "mm/memcontrol.c=5141=static struct cftype memory_files[] = {\n--\nmm/memcontrol.c-5207-\t{\nmm/memcontrol.c:5208:\t\t.name = \"reclaim\",\nmm/memcontrol.c-5209-\t\t.flags = CFTYPE_NS_DELEGATABLE,\n"
}

13/2 2026/09/14 12:16 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 44171
							output: 21
							thoughts: 60
						

Thoughts:
**Investigating Memory Reclaim Mechanisms**

Okay, let's dive in. I need to understand the relationship between general memory reclaim and node-specific reclaim in this system. My immediate starting point is `mm/memcontrol.c`; I know that defines `memory.reclaim`, so that's the likely source for the core reclaim function. But node reclaim feels like a separate, possibly more granular, control. Let me double-check - I suspect the `/sys/devices/system/node/nodeX/reclaim` entries are relevant there. I need to be sure though, before I proceed with any experiments. To clarify the link, I'll quickly grep the codebase to identify all locations where the function `user_proactive_reclaim` is actually used. This should give me a good picture of how the different reclaim mechanisms are orchestrated, so I can continue with my assessment.



14/2 2026/09/14 12:16 tool grepper
0m Args:
{
  "Expression": "user_proactive_reclaim"
}

Results:
{
  "Output": "mm/internal.h=686=extern void reclaim_throttle(pg_data_t *pgdat, enum vmscan_throttle_state reason);\nmm/internal.h:687:int user_proactive_reclaim(char *buf,\nmm/internal.h-688-\t\t\t   struct mem_cgroup *memcg, pg_data_t *pgdat);\n--\nmm/memcontrol.c=5128=static ssize_t memory_reclaim(struct kernfs_open_file *of, char *buf,\n--\nmm/memcontrol.c-5133-\nmm/memcontrol.c:5134:\tret = user_proactive_reclaim(buf, memcg, NULL);\nmm/memcontrol.c-5135-\tif (ret)\n--\nmm/vmscan.c=8056=static const match_table_t tokens = {\n--\nmm/vmscan.c-8061-\nmm/vmscan.c:8062:int user_proactive_reclaim(char *buf,\nmm/vmscan.c-8063-\t\t\t   struct mem_cgroup *memcg, pg_data_t *pgdat)\n--\nmm/vmscan.c=8216=static ssize_t reclaim_store(struct device *dev,\n--\nmm/vmscan.c-8221-\nmm/vmscan.c:8222:\tret = user_proactive_reclaim((char *)buf, NULL, NODE_DATA(nid));\nmm/vmscan.c-8223-\treturn ret ? ret : count;\n"
}

15/2 2026/09/14 12:17 llm patch-evaluator
12m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 44841
							output: 27
							thoughts: 1587
						

Thoughts:
**Fuzzing Focus: Memory Management and Swap Subsystem**

Okay, let's break this down. My initial assessment is that `user_proactive_reclaim` is a key area of interest, as it's used in both cgroup memory reclaim (`memory_reclaim`) and the `/sys/devices/system/node/nodeX/reclaim` interface. Additionally, its use in `lru_gen_seq_write` (within MGLRU and debugfs, requiring `CONFIG_LRU_GEN`) suggests a potential interaction with generational LRU mechanisms, which is interesting.

Now, let's address the questions at hand.

1.  **WorthFuzzing?** Absolutely, the answer is a resounding "yes." This patch touches *core* memory management code. We're looking at significant changes to `mm/swapfile.c` with modifications to functions like `swapon`, `swapoff`, folio-related allocation, and swap synchronization, allocation, and deallocation. Changes to `mm/page_io.c` involving `swap_writeout`'s checks for `SWP_OFFLOAD_ONLY` and the use of `current_reclaim_allows_offload_swap()`, indicate a shift in how swap is handled during memory pressure. There's also modification of `mm/vmscan.c` with changes to `scan_control`, `reclaim_state`, and the integration of the aforementioned `user_proactive_reclaim`, as well as modifications to reclaim logic. And, let's not forget the alterations in `mm/memcontrol.c` and DRM shrinker callers to `get_nr_swap_pages_eligible()`. The new `SWAP_FLAG_OFFLOAD_ONLY` flag is a big deal, along with the swapout refusal paths. This is a core subsystem and a lot of reachable executable code is affected by this.

2.  **FocusSymbols?** We need to identify specific, non-ubiquitous kernel functions to target in fuzzing. Let's see what stands out:

    *   `get_nr_swap_pages_eligible` (New, non-static, *exported!* This is a good target.)
    *   `folio_swap_full` (New, non-static)
    *   `mem_cgroup_get_nr_swap_pages_eligible` (New, non-static)
    *   `swap_writeout` (Non-static, directly modified, so definitely a candidate)
    *   `user_proactive_reclaim` (Non-static, and sets allow\_offload\_swap, another strong candidate)
    *   `folio_alloc_swap` (Non-static, potentially impacted)
    *   `__do_sys_swapon` (or `sys_swapon`) and `__do_sys_swapoff` (or `sys_swapoff`): are good syscall entrypoints for this area
    *   And as a side note, let's note that `swap_alloc_slow` and `swap_sync_discard` are static functions in `swapfile.c`, and we must check if they are inlined or not.

    We need to avoid the "ubiquitous lifecycle hot-paths." No general memory allocators or deallocators, or page fault handlers.

    Okay, so based on the rules, we must now make a final list. I'll need to check the exact symbol names in the kernel's vmlinux and source code. Syzkaller tends to target the kernel function names, like `__do_sys_swapon`, and those function calls that are not inlined are the targets. Looking at `mm/swapfile.c`, the `SYSCALL_DEFINE2` macro for `swapon` generates the symbols like `__do_sys_swapon` and `__se_sys_swapon`. I'll focus on the functions that are either new or directly modified and non-static, because these functions are not inlined.

    So, my refined list of FocusSymbols (for fuzzing) is:

    *   `get_nr_swap_pages_eligible`
    *   `folio_swap_full`
    *   `mem_cgroup_get_nr_swap_pages_eligible`
    *   `swap_writeout`
    *   `user_proactive_reclaim`
    *   `__do_sys_swapon` / `__se_sys_swapon` (or `sys_swapon`)
    *   `__do_sys_swapoff` / `__se_sys_swapoff` (or `swapoff`)



16/2 2026/09/14 12:29 tool grepper
0m Args:
{
  "Expression": "folio_alloc_swap",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "mm/shmem.c=1808=int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,\n--\nmm/shmem.c-1896-\nmm/shmem.c:1897:\tret = folio_alloc_swap(folio);\nmm/shmem.c-1898-\tif (!ret) {\n--\nmm/swap.h=225=extern int swap_retry_table_alloc(swp_entry_t entry, gfp_t gfp);\n--\nmm/swap.h-233- *\nmm/swap.h:234: * folio_alloc_swap(): the entry point for a folio to be swapped\nmm/swap.h-235- * out. It allocates swap slots and pins the slots with swap cache.\n--\nmm/swap.h-245- */\nmm/swap.h:246:int folio_alloc_swap(struct folio *folio);\nmm/swap.h-247-int folio_dup_swap(struct folio *folio, struct page *page);\n--\nmm/swap.h=361=static inline struct swap_info_struct *__swap_entry_to_info(swp_entry_t entry)\n--\nmm/swap.h-365-\nmm/swap.h:366:static inline int folio_alloc_swap(struct folio *folio)\nmm/swap.h-367-{\n--\nmm/swapfile.c=94=static PLIST_HEAD(swap_active_head);\n--\nmm/swapfile.c-98- * protected with swap_avail_lock, ordered by priority.\nmm/swapfile.c:99: * This is used by folio_alloc_swap() instead of swap_active_head\nmm/swapfile.c-100- * because swap_active_head includes all swap_info_structs,\nmm/swapfile.c:101: * but folio_alloc_swap() doesn't need to look at full ones.\nmm/swapfile.c-102- * This uses its own lock instead of swap_lock because when a\n--\nmm/swapfile.c=949=static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,\n--\nmm/swapfile.c-962-\t/*\nmm/swapfile.c:963:\t * All mm swap allocation starts with a folio (folio_alloc_swap),\nmm/swapfile.c-964-\t * it's also the only allocation path for large orders allocation.\n--\nmm/swapfile.c=1777=static int swap_dup_entries_cluster(struct swap_info_struct *si,\n--\nmm/swapfile.c-1811-/**\nmm/swapfile.c:1812: * folio_alloc_swap - allocate swap space for a folio\nmm/swapfile.c-1813- * @folio: folio we want to move to swap\n--\nmm/swapfile.c-1822- */\nmm/swapfile.c:1823:int folio_alloc_swap(struct folio *folio)\nmm/swapfile.c-1824-{\n--\nmm/swapfile.c=2917=static int try_to_unuse(unsigned int type)\n--\nmm/swapfile.c-3000-\t * and even shmem_writeout() could have been preempted after\nmm/swapfile.c:3001:\t * folio_alloc_swap(), temporarily hiding that swap.  It's easy\nmm/swapfile.c-3002-\t * and robust (though cpu-intensive) just to keep retrying.\n--\nmm/vmscan.c=1139=static unsigned int shrink_folio_list(struct list_head *folio_list,\n--\nmm/vmscan.c-1362-\t\t\t}\nmm/vmscan.c:1363:\t\t\tret = folio_alloc_swap(folio);\nmm/vmscan.c-1364-\t\t\tif (ret) {\n--\nmm/vmscan.c-1381-\t\t\t\tcount_mthp_stat(order, MTHP_STAT_SWPOUT_FALLBACK);\nmm/vmscan.c:1382:\t\t\t\tif (folio_alloc_swap(folio))\nmm/vmscan.c-1383-\t\t\t\t\tgoto activate_locked_split;\n"
}

17/2 2026/09/14 12:29 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 48555
							output: 26
							thoughts: 28
						

Thoughts:
Here's my analysis of the provided information, as if these were my internal thoughts:

**Dissecting the Swap Allocation Process**

Okay, so `folio_alloc_swap` is the key entry point for handling swap allocation. That's good to know – it gives me a central point to start digging into the details of the allocation flow. Now, the next logical step in my investigation is to identify all the places where `swap_writeout` is called. Understanding where this function is invoked will illuminate the broader context surrounding the swap-out process. I need to trace this dependency to fully grasp the execution path of swap allocation.



18/2 2026/09/14 12:29 tool grepper
0m Args:
{
  "Expression": "swap_writeout",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "mm/page_io.c=183=static void swap_zeromap_folio_clear(struct folio *folio)\n--\nmm/page_io.c-203- */\nmm/page_io.c:204:int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)\nmm/page_io.c-205-{\n--\nmm/page_io.c-270-\nmm/page_io.c:271:\t__swap_writeout(ctx, folio);\nmm/page_io.c-272-\treturn 0;\n--\nmm/page_io.c=354=static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)\n--\nmm/page_io.c-391-\nmm/page_io.c:392:void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)\nmm/page_io.c-393-{\n--\nmm/shmem.c=1808=int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,\n--\nmm/shmem.c-1920-\t\tBUG_ON(folio_mapped(folio));\nmm/shmem.c:1921:\t\terror = swap_writeout(ctx, folio);\nmm/shmem.c-1922-\t\tif (error != AOP_WRITEPAGE_ACTIVATE) {\n--\nmm/swap.h=259=void swap_write_submit(struct swap_io_ctx *ctx);\nmm/swap.h:260:int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio);\nmm/swap.h:261:void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio);\nmm/swap.h-262-\n--\nmm/swap.h=421=static inline void swap_update_readahead(struct folio *folio,\n--\nmm/swap.h-425-\nmm/swap.h:426:static inline int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)\nmm/swap.h-427-{\n--\nmm/swapfile.c=3107=EXPORT_SYMBOL_GPL(add_swap_extent);\n--\nmm/swapfile.c-3111- * onto a contiguous range of disk blocks.  A rbtree of swap extents is\nmm/swapfile.c:3112: * built at swapon time and is then used at swap_writeout/swap_read_folio\nmm/swapfile.c-3113- * time for locating where on disk a page belongs.\n--\nmm/vmscan.c=700=static pageout_t pageout(struct swap_io_ctx *ctx, struct address_space *mapping,\n--\nmm/vmscan.c-738-\telse\nmm/vmscan.c:739:\t\tres = swap_writeout(ctx, folio);\nmm/vmscan.c-740-\n--\nmm/zswap.c=856=static bool zswap_compress(struct folio *folio, long index,\n--\nmm/zswap.c-896-\t * of writebacks.  If writeback is disabled, reject the page since it\nmm/zswap.c:897:\t * only adds metadata overhead.  swap_writeout() will put the page back\nmm/zswap.c-898-\t * to the active LRU list in the case.\n--\nmm/zswap.c=1005=static int zswap_writeback_entry(struct zswap_entry *entry,\n--\nmm/zswap.c-1070-\t/* start writeback */\nmm/zswap.c:1071:\t__swap_writeout(\u0026ctx, folio);\nmm/zswap.c-1072-\tswap_write_submit(\u0026ctx);\n"
}

19/2 2026/09/14 12:29 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 49734
							output: 31
							thoughts: 282
						

Thoughts:
**My Mental Checklist for Swap Space Analysis**

Okay, let's break this down. First, I'm thinking about the connections between different swap-related functions in the kernel. `folio_alloc_swap` is my starting point, it's called during page reclaim when an anonymous folio is reclaimed via `shrink_folio_list`, that seems logical enough. Then, `swap_writeout` pops up during the actual writeback process, either from `pageout` or `shmem_writeout`.

Next, I'm noting the proactive reclaim path from the user through `/sys/devices/system/node/nodeX/reclaim` or a cgroup's `memory.reclaim`. That's a key control point.

I'm making a mental map of where `get_nr_swap_pages_eligible` is used, and it's popping up in `folio_alloc_swap`, `can_reclaim_anon_pages`, `mem_cgroup_get_nr_swap_pages_eligible`, and also some driver-specific contexts - i915, panthor, msm, and ttm. I'm keeping an eye on that list as it is a good indicator of what is using this information. The same thought goes for `folio_swap_full` which is triggered by `mem_cgroup_swap_full`. Also related is `mem_cgroup_get_nr_swap_pages_eligible`, called by `can_reclaim_anon_pages` and `get_swappiness`.

Now, I need to zoom out a bit. What about the core operations of *swapping* a device in or out? That brings me to `swapon` and `swapoff`. I should check on the syscall implementations with `__do_sys_swapon`, `__x64_sys_swapon`, and `sys_swapon`. I'll need to double check the syzkaller or kernel source to see how they define the swapon symbol.



20/2 2026/09/14 12:29 tool grepper
0m Args:
{
  "Expression": "SYSCALL_DEFINE2\\(swapon",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "mm/swapfile.c=3718=static int setup_swap_clusters_info(struct swap_info_struct *si,\n--\nmm/swapfile.c-3797-\nmm/swapfile.c:3798:SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)\nmm/swapfile.c-3799-{\n"
}

21/2 2026/09/14 12:29 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 50587
							output: 19
							thoughts: 285
						

Thoughts:
**My Thought Process on `swapon` Implementation in Linux**

Okay, so I'm digging into how `swapon` is implemented on x86_64 Linux. Right, the `SYSCALL_DEFINE2(swapon, ...)` macro is used, which means we're going to have `__do_sys_swapon` (the real guts of the function) and `__x64_sys_swapon` (the system call stub). Now, the first thing I need to do is quickly check whether `__do_sys_swapon` is a standard naming convention. I have a hunch it is. Before I do that, though, I should take a look at the important symbols. Can I quickly list some relevant symbols related to swap management to orient myself? Let's check `folio_alloc_swap`, `swap_writeout`, `get_nr_swap_pages_eligible`, and things like `mem_cgroup_get_nr_swap_pages_eligible`, and `folio_swap_full`, as well as `user_proactive_reclaim`.

Hold on, before I get too far, let's also remember to quickly check `__do_sys_swapoff` as well. Now, the main question is, how standard is this `__do_sys_` naming? Let me check `kallsyms` or search the source directory for something like `__do_sys_swapon` to see if it even exists. Aha, I remember. In `include/linux/syscalls.h`, the macro definition for `SYSCALL_DEFINEx` creates `__do_sys_##sname` for us. That seems to be the convention, then. Just to be completely sure, I'll grep the entire repository for `__do_sys_swapon` to see if there are other instances of this naming convention.



22/2 2026/09/14 12:29 tool grepper
0m Args:
{
  "Expression": "__do_sys_"
}

Results:
{
  "Output": "Documentation/dev-tools/kasan.rst=169=A typical KASAN report looks like this::\n--\nDocumentation/dev-tools/kasan.rst-186-     load_module+0x75df/0x8070\nDocumentation/dev-tools/kasan.rst:187:     __do_sys_init_module+0x1c6/0x200\nDocumentation/dev-tools/kasan.rst-188-     __x64_sys_init_module+0x6e/0xb0\n--\nDocumentation/dev-tools/kasan.rst-207-     load_module+0x75df/0x8070\nDocumentation/dev-tools/kasan.rst:208:     __do_sys_init_module+0x1c6/0x200\nDocumentation/dev-tools/kasan.rst-209-     __x64_sys_init_module+0x6e/0xb0\n--\nDocumentation/dev-tools/kmemleak.rst=250=memory leaks``. Then read the file to see then::\n--\nDocumentation/dev-tools/kmemleak.rst-262-            [\u003c00000000542b9814\u003e] load_module+0x203c/0x2480\nDocumentation/dev-tools/kmemleak.rst:263:            [\u003c00000000c2850256\u003e] __do_sys_finit_module+0xba/0xe0\nDocumentation/dev-tools/kmemleak.rst-264-            [\u003c000000006564e7ef\u003e] do_syscall_64+0x43/0x110\n--\nDocumentation/mm/page_owner.rst=65=Usage\n--\nDocumentation/mm/page_owner.rst-127-\t kernel_read_file+0x309/0x4d0\nDocumentation/mm/page_owner.rst:128:\t __do_sys_finit_module+0x381/0x730\nDocumentation/mm/page_owner.rst-129-\t do_syscall_64+0x8d/0x150\n--\nDocumentation/mm/page_owner.rst-154-\t kernel_read_file+0x309/0x4d0\nDocumentation/mm/page_owner.rst:155:\t __do_sys_finit_module+0x381/0x730\nDocumentation/mm/page_owner.rst-156-\t do_syscall_64+0x8d/0x150\n--\nDocumentation/trace/timerlat-tracer.rst=142=irq disabled. For example::\n--\nDocumentation/trace/timerlat-tracer.rst-165-  =\u003e do_init_module\nDocumentation/trace/timerlat-tracer.rst:166:  =\u003e __do_sys_finit_module\nDocumentation/trace/timerlat-tracer.rst-167-  =\u003e do_syscall_64\n--\nDocumentation/translations/zh_CN/dev-tools/kasan.rst=96=KASAN受到通用 ``panic_on_warn`` 命令行参数的影响。当它被启用时,KASAN\n--\nDocumentation/translations/zh_CN/dev-tools/kasan.rst-154-     load_module+0x75df/0x8070\nDocumentation/translations/zh_CN/dev-tools/kasan.rst:155:     __do_sys_init_module+0x1c6/0x200\nDocumentation/translations/zh_CN/dev-tools/kasan.rst-156-     __x64_sys_init_module+0x6e/0xb0\n--\nDocumentation/translations/zh_CN/dev-tools/kasan.rst-175-     load_module+0x75df/0x8070\nDocumentation/translations/zh_CN/dev-tools/kasan.rst:176:     __do_sys_init_module+0x1c6/0x200\nDocumentation/translations/zh_CN/dev-tools/kasan.rst-177-     __x64_sys_init_module+0x6e/0xb0\n--\nDocumentation/translations/zh_CN/dev-tools/kmemleak.rst=202=kmemleak-test。设置 CONFIG_SAMPLE_KMEMLEAK 为模块(不能作为内建模块使用)\n--\nDocumentation/translations/zh_CN/dev-tools/kmemleak.rst-222-       [\u003c00000000542b9814\u003e] load_module+0x203c/0x2480\nDocumentation/translations/zh_CN/dev-tools/kmemleak.rst:223:       [\u003c00000000c2850256\u003e] __do_sys_finit_module+0xba/0xe0\nDocumentation/translations/zh_CN/dev-tools/kmemleak.rst-224-       [\u003c000000006564e7ef\u003e] do_syscall_64+0x43/0x110\n--\nDocumentation/translations/zh_CN/mm/page_owner.rst=32=page owner在默认情况下是禁用的。所以,如果你想使用它,你需要在你的启动cmdline\n--\nDocumentation/translations/zh_CN/mm/page_owner.rst-102-\t kernel_read_file+0x309/0x4d0\nDocumentation/translations/zh_CN/mm/page_owner.rst:103:\t __do_sys_finit_module+0x381/0x730\nDocumentation/translations/zh_CN/mm/page_owner.rst-104-\t do_syscall_64+0x8d/0x150\n--\nDocumentation/translations/zh_TW/dev-tools/kasan.rst=96=KASAN受到通用 ``panic_on_warn`` 命令行參數的影響。當它被啓用時,KASAN\n--\nDocumentation/translations/zh_TW/dev-tools/kasan.rst-154-     load_module+0x75df/0x8070\nDocumentation/translations/zh_TW/dev-tools/kasan.rst:155:     __do_sys_init_module+0x1c6/0x200\nDocumentation/translations/zh_TW/dev-tools/kasan.rst-156-     __x64_sys_init_module+0x6e/0xb0\n--\nDocumentation/translations/zh_TW/dev-tools/kasan.rst-175-     load_module+0x75df/0x8070\nDocumentation/translations/zh_TW/dev-tools/kasan.rst:176:     __do_sys_init_module+0x1c6/0x200\nDocumentation/translations/zh_TW/dev-tools/kasan.rst-177-     __x64_sys_init_module+0x6e/0xb0\n--\narch/s390/include/asm/syscall_wrapper.h-19-\tALLOW_ERROR_INJECTION(__s390x_sys_##sname, ERRNO);\t\t\\\narch/s390/include/asm/syscall_wrapper.h:20:\tstatic inline long __do_sys_##sname(void);\t\t\t\\\narch/s390/include/asm/syscall_wrapper.h-21-\tlong __s390x_sys_##sname(struct pt_regs *__unused)\t\t\\\narch/s390/include/asm/syscall_wrapper.h-22-\t{\t\t\t\t\t\t\t\t\\\narch/s390/include/asm/syscall_wrapper.h:23:\t\treturn __do_sys_##sname();\t\t\t\t\\\narch/s390/include/asm/syscall_wrapper.h-24-\t}\t\t\t\t\t\t\t\t\\\narch/s390/include/asm/syscall_wrapper.h:25:\tstatic inline long __do_sys_##sname(void)\narch/s390/include/asm/syscall_wrapper.h-26-\n--\narch/x86/include/asm/syscall_wrapper.h=12=extern long __ia32_sys_ni_syscall(const struct pt_regs *regs);\n--\narch/x86/include/asm/syscall_wrapper.h-27- * perform sign-extension (omitted for zero-argument syscalls).  Finally the\narch/x86/include/asm/syscall_wrapper.h:28: * arguments are passed to the __do_sys_*() function which is the actual\narch/x86/include/asm/syscall_wrapper.h-29- * syscall.  These wrappers are marked as inline so the compiler can optimize\n--\narch/x86/include/asm/syscall_wrapper.h-245-\tSYSCALL_METADATA(_##sname, 0);\t\t\t\t\t\\\narch/x86/include/asm/syscall_wrapper.h:246:\tstatic long __do_sys_##sname(const struct pt_regs *__unused);\t\\\narch/x86/include/asm/syscall_wrapper.h-247-\t__X64_SYS_STUB0(sname)\t\t\t\t\t\t\\\narch/x86/include/asm/syscall_wrapper.h-248-\t__IA32_SYS_STUB0(sname)\t\t\t\t\t\t\\\narch/x86/include/asm/syscall_wrapper.h:249:\tstatic long __do_sys_##sname(const struct pt_regs *__unused)\narch/x86/include/asm/syscall_wrapper.h-250-\n--\ninclude/linux/syscalls.h=211=static inline int is_syscall_trace_event(struct trace_event_call *tp_event)\n--\ninclude/linux/syscalls.h-242- * sign-extends 32-bit ints to longs whenever needed. The actual work is\ninclude/linux/syscalls.h:243: * done within __do_sys_*().\ninclude/linux/syscalls.h-244- */\n"
}

23/2 2026/09/14 12:29 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 53836
							output: 44
							thoughts: 137
						

Thoughts:
**Examining the Swap System Calls**

Okay, here we go, looking at `__do_sys_swapon` and `__do_sys_swapoff`. These are the key functions I'm interested in. Now, I need to check how these are handled specifically for the x86_64 architecture. Let's see... looking at `arch/x86/include/asm/syscall_wrapper.h`, I see that `__do_sys_##sname` – which, in our case, becomes `__do_sys_swapon` or `__do_sys_swapoff` – is declared as `static`. Right, so it's likely to be inlined based on compiler optimisations. This means it may or may not be the direct callee of `__x64_sys_##sname`. Best to review the surrounding lines in the header. Let's start around line 245 of that file and see what's what.



24/2 2026/09/14 12:29 tool grepper
0m Args:
{
  "Expression": "__X64_SYS_STUB",
  "PathPrefix": "arch/x86/include/asm/syscall_wrapper.h"
}

Results:
{
  "Output": "arch/x86/include/asm/syscall_wrapper.h=12=extern long __ia32_sys_ni_syscall(const struct pt_regs *regs);\n--\narch/x86/include/asm/syscall_wrapper.h-106-#ifdef CONFIG_X86_64\narch/x86/include/asm/syscall_wrapper.h:107:#define __X64_SYS_STUB0(name)\t\t\t\t\t\t\\\narch/x86/include/asm/syscall_wrapper.h-108-\t__SYS_STUB0(x64, sys_##name)\narch/x86/include/asm/syscall_wrapper.h-109-\narch/x86/include/asm/syscall_wrapper.h:110:#define __X64_SYS_STUBx(x, name, ...)\t\t\t\t\t\\\narch/x86/include/asm/syscall_wrapper.h-111-\t__SYS_STUBx(x64, sys##name,\t\t\t\t\t\\\n--\narch/x86/include/asm/syscall_wrapper.h-117-#else /* CONFIG_X86_64 */\narch/x86/include/asm/syscall_wrapper.h:118:#define __X64_SYS_STUB0(name)\narch/x86/include/asm/syscall_wrapper.h:119:#define __X64_SYS_STUBx(x, name, ...)\narch/x86/include/asm/syscall_wrapper.h-120-#define __X64_COND_SYSCALL(name)\n--\narch/x86/include/asm/syscall_wrapper.h-225-\tstatic inline long __do_sys##name(__MAP(x,__SC_DECL,__VA_ARGS__));\\\narch/x86/include/asm/syscall_wrapper.h:226:\t__X64_SYS_STUBx(x, name, __VA_ARGS__)\t\t\t\t\\\narch/x86/include/asm/syscall_wrapper.h-227-\t__IA32_SYS_STUBx(x, name, __VA_ARGS__)\t\t\t\t\\\n--\narch/x86/include/asm/syscall_wrapper.h-246-\tstatic long __do_sys_##sname(const struct pt_regs *__unused);\t\\\narch/x86/include/asm/syscall_wrapper.h:247:\t__X64_SYS_STUB0(sname)\t\t\t\t\t\t\\\narch/x86/include/asm/syscall_wrapper.h-248-\t__IA32_SYS_STUB0(sname)\t\t\t\t\t\t\\\n"
}

25/2 2026/09/14 12:29 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 54889
							output: 43
							thoughts: 88
						

Thoughts:
Here's my thinking process:

**Deconstructing the System Call Implementation**

Okay, so I'm looking at how these system calls are defined. It seems pretty straightforward on the surface, but let's dig in. I'm seeing that the `__do_sys_##name` macro is a static inline function definition. That means it's likely optimized for performance, and the compiler will probably just insert the code directly where it's called. The `__X64_SYS_STUBx` macro is where things are getting interesting though because it's defining either `__x64_sys_##name` or `__se_sys_##name`. I need to figure out which it's doing and why. I should take a look at the actual code in the `__X64_SYS_STUBx` macro to see the specific details of what is happening. Lines 110-116 should have what I need. Let me jump there and see...



26/2 2026/09/14 12:29 tool read-file
0m Args:
{
  "File": "arch/x86/include/asm/syscall_wrapper.h",
  "FirstLine": 100,
  "LineCount": 25
}

Results:
{
  "Contents": " 100:\t\t__weak long __##abi##_##name(const struct pt_regs *__unused);\t\\\n 101:\t\t__weak long __##abi##_##name(const struct pt_regs *__unused)\t\\\n 102:\t\t{\t\t\t\t\t\t\t\t\\\n 103:\t\t\treturn sys_ni_syscall();\t\t\t\t\\\n 104:\t\t}\n 105:\t\n 106:\t#ifdef CONFIG_X86_64\n 107:\t#define __X64_SYS_STUB0(name)\t\t\t\t\t\t\\\n 108:\t\t__SYS_STUB0(x64, sys_##name)\n 109:\t\n 110:\t#define __X64_SYS_STUBx(x, name, ...)\t\t\t\t\t\\\n 111:\t\t__SYS_STUBx(x64, sys##name,\t\t\t\t\t\\\n 112:\t\t\t    SC_X86_64_REGS_TO_ARGS(x, __VA_ARGS__))\n 113:\t\n 114:\t#define __X64_COND_SYSCALL(name)\t\t\t\t\t\\\n 115:\t\t__COND_SYSCALL(x64, sys_##name)\n 116:\t\n 117:\t#else /* CONFIG_X86_64 */\n 118:\t#define __X64_SYS_STUB0(name)\n 119:\t#define __X64_SYS_STUBx(x, name, ...)\n 120:\t#define __X64_COND_SYSCALL(name)\n 121:\t#endif /* CONFIG_X86_64 */\n 122:\t\n 123:\t#if defined(CONFIG_X86_32) || defined(CONFIG_IA32_EMULATION)\n 124:\t#define __IA32_SYS_STUB0(name)\t\t\t\t\t\t\\\n"
}

27/2 2026/09/14 12:29 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 55730
							output: 26
							thoughts: 1107
						

Thoughts:
**Analyzing the Impact of the Swap Patch**

Okay, let's break this down. First, the patch introduces these `__x64_sys_##name` system calls: `__x64_sys_swapon` and `__x64_sys_swapoff`. But the real question is, what other functions are directly modified or added by the patch and how should I approach it in terms of symbol coverage?

I've got a list. In `mm/swapfile.c`, we've got `get_nr_swap_pages_eligible`, `folio_swap_full`, `folio_alloc_swap`, and the new system calls `__x64_sys_swapon` (or `__do_sys_swapon`) and `__x64_sys_swapoff` (or `__do_sys_swapoff`). Then, `swap_writeout` in `mm/page_io.c`. In `mm/memcontrol.c`, there's `mem_cgroup_get_nr_swap_pages_eligible` and `mem_cgroup_swap_full`. And lastly, in `mm/vmscan.c`, we have `user_proactive_reclaim`.

Now, the important question: are any of these ubiquitous? They would create a large volume of hits. Let's see... `get_nr_swap_pages_eligible`? Nope, it's pretty specific to swap and reclaim paths. `folio_swap_full`? Used with memcg, also not general. `folio_alloc_swap`? Only during pageout of anonymous pages to swap. `swap_writeout`? Only for writing swap pages. `mem_cgroup_get_nr_swap_pages_eligible`? Memcg reclaim/swappiness, not ubiquitous. `mem_cgroup_swap_full`?  Folio full checks, specific. `user_proactive_reclaim`?  This one's triggered by `memory.reclaim` and `/sys/devices/system/node/nodeX/reclaim` calls -- a proactive reclaim entrypoint, not a general thing. And finally, the syscalls themselves, `__x64_sys_swapon` and `__x64_sys_swapoff` are entry points! Not general usage either.

Hmmm, what about `try_to_free_mem_cgroup_pages`? Too broad, reclaim is general. And `lru_gen_seq_write`? Debugfs. Not. And `swap_sync_discard` is static.

Wait, the instructions say something important about static inline functions in headers. They lack unique symbol addresses and can't be targeted, so targeting non-static functions in .c files that exercise them is the right approach. Bingo.

So, given those instructions, I've got a refined list of non-static functions:
`get_nr_swap_pages_eligible`, `mem_cgroup_get_nr_swap_pages_eligible`, `folio_swap_full`, `folio_alloc_swap`, `swap_writeout`, `user_proactive_reclaim`, `__x64_sys_swapon`, and `__x64_sys_swapoff`.

Wait, let's check for any *other* non-static functions in `mm/swapfile.c` that are altered: `find_first_swap` (modified with a check for `SWP_OFFLOAD_ONLY`), `count_swap_pages` (ditto), and `swap_alloc_hibernation_slot` (ditto).

Finally, I need to know what config options are *required* to compile and exercise the modified code. Definitely `CONFIG_SWAP`. If swap is disabled, `swapfile.c` isn't even compiled. Let's make sure by checking the `mm/Makefile`. That's it, that's the list of functions affected and the important config dependencies.



28/2 2026/09/14 12:30 tool grepper
0m Args:
{
  "Expression": "swapfile\\.o",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "mm/Makefile=77=endif\nmm/Makefile-78-\nmm/Makefile:79:obj-$(CONFIG_SWAP)\t+= page_io.o swap_state.o swapfile.o\nmm/Makefile-80-obj-$(CONFIG_ZSWAP)\t+= zswap.o\n"
}

29/2 2026/09/14 12:30 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 57811
							output: 26
							thoughts: 101
						

Thoughts:
Okay, here's my interpretation of those build instructions, from my perspective, as if they're my mental flow:

**Analyzing Kernel Build Dependencies**

Alright, so the build system is pretty clear here. I see that if `CONFIG_SWAP` is defined, we're including `page_io.o`, `swap_state.o`, and `swapfile.o` in the `obj-$(CONFIG_SWAP)` list. That's straightforward; that makes perfect sense for swap functionality.

Now, for `memcontrol.c`, the requirement is even more explicit: `CONFIG_MEMCG`.  Right, `memcontrol.c` deals with memory control groups, so that's a direct dependency.  No surprises there.

Finally, I'm thinking about `SWPOUT_OFFLOAD_REFUSED`. I vaguely remember needing to check a configuration option for this. Let me quickly double-check that `CONFIG_VM_EVENT_COUNTERS` is defined, probably in `mm/Makefile` or maybe `mm/vmstat.c`.  I'll need to verify that setting. Just a quick check to make sure the relevant event counters are enabled before going further with the analysis of `SWPOUT_OFFLOAD_REFUSED`.



30/2 2026/09/14 12:30 tool grepper
0m Args:
{
  "Expression": "vmstat\\.o",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "mm/Makefile=30=KCOV_INSTRUMENT_mmzone.o := n\nmm/Makefile:31:KCOV_INSTRUMENT_vmstat.o := n\nmm/Makefile-32-KCOV_INSTRUMENT_failslab.o := n\n--\nmm/Makefile=51=obj-y\t\t\t:= filemap.o mempool.o oom_kill.o fadvise.o \\\n--\nmm/Makefile-53-\t\t\t   readahead.o folio.o truncate.o vmscan.o shrinker.o \\\nmm/Makefile:54:\t\t\t   shmem.o util.o mmzone.o vmstat.o backing-dev.o \\\nmm/Makefile-55-\t\t\t   mm_init.o percpu.o slab_common.o \\\n"
}

31/2 2026/09/14 12:30 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 58404
							output: 29
							thoughts: 20
						

Thoughts:
Here's my take on those thoughts:

**Checking Kernel Configuration for Virtual Machine Event Counters**

Okay, so I need to quickly ascertain the kernel's stance on virtual machine event counters. My first move is to dive straight into `mm/Kconfig`. That's where I expect to find the configuration option, specifically `CONFIG_VM_EVENT_COUNTERS`, which should define whether or not these VM-specific performance metrics are enabled. Time to get to the heart of the matter and confirm if this feature is compiled in!



32/2 2026/09/14 12:30 tool grepper
0m Args:
{
  "Expression": "config VM_EVENT_COUNTERS",
  "PathPrefix": "mm/"
}

Results:
{
  "Output": "mm/Kconfig=1274=config ARCH_USES_PG_ARCH_3\n--\nmm/Kconfig-1276-\nmm/Kconfig:1277:config VM_EVENT_COUNTERS\nmm/Kconfig-1278-\tdefault y\n"
}

33/2 2026/09/14 12:30 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 58642
							output: 33
							thoughts: 225
						

Thoughts:
**My Analysis of the Swap System Dependencies**

Okay, let's break this down. First, the core `SWAP` functionality clearly hinges on the presence of an `MMU`, the `BLOCK` layer, but crucially, *not* the `ARCH_NO_SWAP` configuration. That's the baseline. Now, what about `ZSWAP`? The interaction with `zswap_store` in `swap_writeout` seems particularly interesting here: specifically, the condition `if (!(READ_ONCE(sis->flags) & SWP_OFFLOAD_ONLY) && zswap_store(folio)) {` tells me that `ZSWAP` is critical if we're evaluating zswap's integration with swap that's configured for offload-only mode. We need to focus on how this is behaving.

Next up, `MEMCG`. It looks like `mem_cgroup_get_nr_swap_pages_eligible` is conditionally compiled based on `CONFIG_MEMCG` within `mm/memcontrol.c`. This suggests that the memory cgroup feature has impact on what pages are determined to be swap eligible. This means that memory accounting and swap are linked in a way we need to investigate.

Finally, `LRU_GEN`. The modifications in `mm/vmscan.c` to `lru_gen_seq_write` and `mm_walk` are controlled by `CONFIG_LRU_GEN`. I should probably check that configuration's implications in the `mm/vmscan.c` file to understand the logic.



34/2 2026/09/14 12:30 tool grepper
0m Args:
{
  "Expression": "CONFIG_LRU_GEN",
  "PathPrefix": "include/linux/swap.h"
}

Results:
{
  "Output": "include/linux/swap.h=141=struct reclaim_state {\n--\ninclude/linux/swap.h-145-\tbool allow_offload_swap;\ninclude/linux/swap.h:146:#ifdef CONFIG_LRU_GEN\ninclude/linux/swap.h-147-\t/* per-thread mm walk data */\n"
}

35/2 2026/09/14 12:30 llm patch-evaluator
1m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 59331
							output: 34
							thoughts: 22
						

Thoughts:
Here's my summary, as if I were thinking these thoughts:

**Diving into `lru_gen_seq_write`**

Okay, time to get my hands dirty and jump into the code. I need to take a look at the `lru_gen_seq_write` function. Specifically, I'm heading over to `mm/vmscan.c`. This is where I'll find the core of the LRU generation and writeback sequence logic, crucial for understanding how the kernel manages memory and pages. I'm expecting to see how this function interacts with the page replacement algorithm and the writeback mechanisms to optimize memory usage and minimize I/O overhead. This will likely involve examining its interaction with various data structures and flags within the page cache and the swap subsystem.



36/2 2026/09/14 12:31 tool grepper
0m Args:
{
  "Expression": "lru_gen_seq_write",
  "PathPrefix": "mm/vmscan.c"
}

Results:
{
  "Output": "mm/vmscan.c=5879=static int run_cmd(char cmd, u64 memcg_id, int nid, unsigned long seq,\n--\nmm/vmscan.c-5920-/* see Documentation/admin-guide/mm/multigen_lru.rst for details */\nmm/vmscan.c:5921:static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,\nmm/vmscan.c-5922-\t\t\t\t size_t len, loff_t *pos)\n--\nmm/vmscan.c=6011=static const struct file_operations lru_gen_rw_fops = {\n--\nmm/vmscan.c-6013-\t.read = seq_read,\nmm/vmscan.c:6014:\t.write = lru_gen_seq_write,\nmm/vmscan.c-6015-\t.llseek = seq_lseek,\n"
}

37/2 2026/09/14 12:31 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 59805
							output: 35
							thoughts: 32
						

Thoughts:
Here's a summary of my thought process, as requested:

**Deconstructing `lru_gen_seq_write` Protection**

Okay, so I need to understand the safeguards surrounding `lru_gen_seq_write`. Where does this beast live? Ah, right, it's in the memory management section, `mm/vmscan.c`. Specifically, line 5921. Now, let's drill down into that code and figure out the exact mechanisms in place to keep this thing safe and sound. What are the key elements protecting it? I need to analyze its use within the broader context of the virtual memory subsystem. Let's dig in and see how the kernel handles this particular write. This might involve looking at locks, atomic operations, or whatever synchronization primitives are employed. Let's see what's what.



38/2 2026/09/14 12:32 tool read-file
0m Args:
{
  "File": "mm/vmscan.c",
  "FirstLine": 5915,
  "LineCount": 25
}

Results:
{
  "Contents": "5915:\t\tmem_cgroup_put(memcg);\n5916:\t\n5917:\t\treturn err;\n5918:\t}\n5919:\t\n5920:\t/* see Documentation/admin-guide/mm/multigen_lru.rst for details */\n5921:\tstatic ssize_t lru_gen_seq_write(struct file *file, const char __user *src,\n5922:\t\t\t\t\t size_t len, loff_t *pos)\n5923:\t{\n5924:\t\tvoid *buf;\n5925:\t\tchar *cur, *next;\n5926:\t\tunsigned int flags;\n5927:\t\tstruct blk_plug plug;\n5928:\t\tint err = -EINVAL;\n5929:\t\tstruct scan_control sc = {\n5930:\t\t\t.may_writepage = true,\n5931:\t\t\t.may_unmap = true,\n5932:\t\t\t.may_swap = true,\n5933:\t\t\t.reclaim_idx = MAX_NR_ZONES - 1,\n5934:\t\t\t.gfp_mask = GFP_KERNEL,\n5935:\t\t\t.proactive = true,\n5936:\t\t\t.allow_offload_swap = true,\n5937:\t\t};\n5938:\t\n5939:\t\tbuf = kvmalloc(len + 1, GFP_KERNEL);\n"
}

39/2 2026/09/14 12:32 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 60426
							output: 34
							thoughts: 27
						

Thoughts:
Here's my take on those thoughts, tailored for an expert audience:

**Diving into vmscan.c: Examining the lru_gen_seq_write Guard**

Okay, so I'm thinking, let's zero in on the `#ifdef` block surrounding `lru_gen_seq_write` within `mm/vmscan.c`. Immediately, my brain is starting to parse what's likely the implications here. This macro suggests conditional compilation, which means this code path is probably enabled or disabled based on a specific kernel configuration option. Given its name, `lru_gen_seq_write` clearly hints at something related to the Least Recently Used (LRU) cache, specifically how the write side of sequential access is handled. I need to figure out what triggers this conditional compilation: what feature is controlled by this flag. Is it an optimization, a debugging tool, or a particular memory management mode? The conditional nature suggests that this code segment is likely an advanced feature or a performance tweak that may not be universally desirable, or, alternatively, it is a feature requiring a certain kernel setting to be enabled. My goal is to understand precisely what effect it has on the page reclaim behavior and how it relates to the wider system memory management strategy. Looking at the `#ifdef` can reveal the configuration option and then I can delve into the related documentation to grasp its role, and how it impacts the performance and behavior of the virtual memory system.



40/2 2026/09/14 12:32 tool grepper
0m Args:
{
  "Expression": "#ifdef CONFIG_LRU_GEN",
  "PathPrefix": "mm/vmscan.c"
}

Results:
{
  "Output": "mm/vmscan.c=905=enum folio_references {\n--\nmm/vmscan.c-910-\nmm/vmscan.c:911:#ifdef CONFIG_LRU_GEN\nmm/vmscan.c-912-/*\n--\nmm/vmscan.c=2766=static bool can_age_anon_pages(struct lruvec *lruvec,\n--\nmm/vmscan.c-2777-\nmm/vmscan.c:2778:#ifdef CONFIG_LRU_GEN\nmm/vmscan.c-2779-\nmm/vmscan.c=2780=DEFINE_STATIC_KEY_FALSE(lru_switch);\nmm/vmscan.c:2781:#ifdef CONFIG_LRU_GEN_ENABLED\nmm/vmscan.c-2782-DEFINE_STATIC_KEY_ARRAY_TRUE(lru_gen_caps, NR_LRU_GEN_CAPS);\n--\nmm/vmscan.c=2966=static void reset_bloom_filter(struct lru_gen_mm_state *mm_state, unsigned long seq)\n--\nmm/vmscan.c-2985-\nmm/vmscan.c:2986:#ifdef CONFIG_LRU_GEN_WALKS_MMU\nmm/vmscan.c-2987-\n"
}

41/2 2026/09/14 12:32 llm patch-evaluator
8m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 61128
							output: 34
							thoughts: 66
						

Thoughts:
Here's my summary, as the expert:

**Navigating the Codebase: Configuration Dependencies**

Okay, so I'm diving into this code, specifically looking at how the Least Recently Used (LRU) generation sequence write function, `lru_gen_seq_write`, is configured. Right, it's under `CONFIG_LRU_GEN` – that makes sense. Now, the `user_proactive_reclaim` functionality… that's the one I'm less sure about. I seem to recall it being tied to either `CONFIG_PROC_FS` or `CONFIG_SYSFS`, possibly both. To be certain, I should quickly verify its location to understand the exact configuration dependencies. Time to drill down on `user_proactive_reclaim` and see where it lives in the build configuration.



42/2 2026/09/14 12:41 tool grepper
0m Args:
{
  "Expression": "int user_proactive_reclaim",
  "PathPrefix": "mm/vmscan.c"
}

Results:
{
  "Output": "mm/vmscan.c=8056=static const match_table_t tokens = {\n--\nmm/vmscan.c-8061-\nmm/vmscan.c:8062:int user_proactive_reclaim(char *buf,\nmm/vmscan.c-8063-\t\t\t   struct mem_cgroup *memcg, pg_data_t *pgdat)\n"
}

43/2 2026/09/14 12:41 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 61513
							output: 35
							thoughts: 28
						

Thoughts:
Okay, let's dive into this. Here's my take on those lines in `mm/vmscan.c`, assuming I'm an expert:

**Investigating Memory Management in `mm/vmscan.c`**

Alright, I need to get my head around what's happening specifically around lines 8050 to 8070 of `mm/vmscan.c`. It's pretty clear this is going to be some core memory scanning logic, given the file context. I expect this section to be deeply involved in page reclamation, likely dealing with how the kernel decides what pages to evict to free up memory. It could be related to various aspects of the scanning process: maybe the choice of which page lists to scan, the criteria used for selecting candidate pages, or perhaps even the scoring mechanism that determines which pages are most suitable for eviction.

I'm mentally preparing for complex conditional statements, likely nested loops, and potentially some intricate interaction with zone reclaim or NUMA-related logic. I wouldn't be surprised to see some interplay with the LRU lists, probably looking at age, accessed flags, dirty bits, and maybe even some accounting of page faults. I am also betting there are several calls to helper functions, especially if this section is part of a larger scan.

Given the line numbers, it's likely part of a function or a block of code within a function, but I am going to have to see what function this block sits in to give it true context. I am betting the code in this section is fairly critical. It might be where the kernel makes some of its more important decisions about memory usage.



44/2 2026/09/14 12:41 tool read-file
0m Args:
{
  "File": "mm/vmscan.c",
  "FirstLine": 8050,
  "LineCount": 30
}

Results:
{
  "Contents": "8050:\t\n8051:\tenum {\n8052:\t\tMEMORY_RECLAIM_SWAPPINESS = 0,\n8053:\t\tMEMORY_RECLAIM_SWAPPINESS_MAX,\n8054:\t\tMEMORY_RECLAIM_NULL,\n8055:\t};\n8056:\tstatic const match_table_t tokens = {\n8057:\t\t{ MEMORY_RECLAIM_SWAPPINESS, \"swappiness=%d\"},\n8058:\t\t{ MEMORY_RECLAIM_SWAPPINESS_MAX, \"swappiness=max\"},\n8059:\t\t{ MEMORY_RECLAIM_NULL, NULL },\n8060:\t};\n8061:\t\n8062:\tint user_proactive_reclaim(char *buf,\n8063:\t\t\t\t   struct mem_cgroup *memcg, pg_data_t *pgdat)\n8064:\t{\n8065:\t\tunsigned int nr_retries = MAX_RECLAIM_RETRIES;\n8066:\t\tunsigned long nr_to_reclaim, nr_reclaimed = 0;\n8067:\t\tint swappiness = -1;\n8068:\t\tchar *old_buf, *start;\n8069:\t\tsubstring_t args[MAX_OPT_ARGS];\n8070:\t\tgfp_t gfp_mask = GFP_KERNEL;\n8071:\t\n8072:\t\tif (!buf || (!memcg \u0026\u0026 !pgdat) || (memcg \u0026\u0026 pgdat))\n8073:\t\t\treturn -EINVAL;\n8074:\t\n8075:\t\tbuf = strstrip(buf);\n8076:\t\n8077:\t\told_buf = buf;\n8078:\t\tnr_to_reclaim = memparse(buf, \u0026buf) / PAGE_SIZE;\n8079:\t\tif (buf == old_buf)\n"
}

45/2 2026/09/14 12:41 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 62408
							output: 30
							thoughts: 76
						

Thoughts:
Here's my take on the situation:

**Reclaiming Memory: A Deep Dive**

Okay, so I'm thinking about memory reclaim. I noticed `user_proactive_reclaim` isn't conditional on an `#ifdef` within `mm/vmscan.c`. It's always compiled along with `vmscan.o`, that makes sense. However, when it comes to controlling the reclaim functionality via `memory.reclaim`, the system configuration *absolutely* needs `CONFIG_MEMCG` enabled. Got it. Now, for the finer-grained node reclaim via the `/sys/devices/system/node/nodeX/reclaim` interface… I need to confirm the interactions there. What's the dependency tree like at that level? Let me dig a bit deeper.



46/2 2026/09/14 12:42 tool grepper
0m Args:
{
  "Expression": "reclaim_store",
  "PathPrefix": "mm/vmscan.c"
}

Results:
{
  "Output": "mm/vmscan.c=8213=EXPORT_SYMBOL_GPL(check_move_unevictable_folios);\n--\nmm/vmscan.c-8215-#if defined(CONFIG_SYSFS) \u0026\u0026 defined(CONFIG_NUMA)\nmm/vmscan.c:8216:static ssize_t reclaim_store(struct device *dev,\nmm/vmscan.c-8217-\t\t\t     struct device_attribute *attr,\n"
}

47/2 2026/09/14 12:42 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 62802
							output: 280
							thoughts: 1720
						

Thoughts:
**Fuzzing Target Identification: Offload-Only Swap Feature**

Okay, so I've just gone through the checklist to determine if this new offload-only swap feature warrants fuzzing. The goal is to reach this new functionality, so let's start by looking at what the configuration requirements are. We need `SWAP` for swap support, `MEMCG` for the cgroup v2 memory controller and proactive reclaim, `ZSWAP` is recommended to check the zswap bypass logic, and `VM_EVENT_COUNTERS` for related counters in `/proc/vmstat`.

Now, let's nail down what the key symbols are. The most relevant and non-ubiquitous functions include `folio_alloc_swap`, which is the core swap allocation function, now updated to check eligible swap areas; `swap_writeout`, the direct writeout, now with offload refusal and zswap bypass; `get_nr_swap_pages_eligible` and `mem_cgroup_get_nr_swap_pages_eligible`, the new reporting functions; `folio_swap_full`, a new function checking if a swap area needs reclaim; `user_proactive_reclaim`, the entry point for proactive reclaim; and `__x64_sys_swapon`, the system call entry point, and finally, `__x64_sys_swapoff`. I verified the correct syscall name by checking the x86_64 kallsyms.

Importantly, none of these are *ubiquitous*. `user_proactive_reclaim` is only called during explicit reclaim; `swap_writeout` only when swapping out; `folio_alloc_swap` only during anonymous page swap allocation; `get_nr_swap_pages_eligible` is called by swap/reclaim paths; `mem_cgroup_get_nr_swap_pages_eligible` during memcg scan/swappiness check; and `folio_swap_full` only during memcg_swap_full/swapcache check.

Critical observation: this patch introduces two new assertions: `VM_WARN_ON_FOLIO(!folio_test_locked(folio), folio);` and `VM_WARN_ON_FOLIO(!folio_test_swapcache(folio), folio);` within `folio_swap_full`. These assertions are strong signals that the patch introduces new invariants and that fuzzing is necessary to verify them. The `folio_swap_full` call is also now the replacement for `vm_swap_full()`, so the assertions are reachable code.

Therefore, this patch is absolutely worth fuzzing. I'm going to set `WorthFuzzing` to `true` and provide the following reasoning:

1.  This patch modifies core memory management subsystems like `mm/swapfile.c`, `mm/page_io.c`, `mm/vmscan.c`, and `mm/memcontrol.c`.
2.  It introduces a new `SWAP_FLAG_OFFLOAD_ONLY` flag to the `swapon` syscall and enforces new validation rules against page discard.
3.  The swap slot allocation logic is changed (both fast and slow paths), and now tracks eligible swap space separately from the raw space.
4.  It introduces refusal paths for retained swap entries during non-proactive reclaim in `swap_writeout`, with a zswap bypass for offload-only areas.
5.  It adds two new runtime assertions (`VM_WARN_ON_FOLIO`) in `folio_swap_full()`.

I will set the `FocusSymbols` to:

*   `folio_alloc_swap`
*   `swap_writeout`
*   `get_nr_swap_pages_eligible`
*   `mem_cgroup_get_nr_swap_pages_eligible`
*   `folio_swap_full`
*   `user_proactive_reclaim`
*   `__x64_sys_swapon`

And the `EnableConfigs` to:

*   `SWAP`
*   `MEMCG`
*   `ZSWAP`
*   `VM_EVENT_COUNTERS`

With this information I am ready to call `set-results`.



48/2 2026/09/14 12:43 tool set-results
0m Args:
{
  "EnableConfigs": [
    "SWAP",
    "MEMCG",
    "ZSWAP",
    "VM_EVENT_COUNTERS"
  ],
  "FocusSymbols": [
    "folio_alloc_swap",
    "swap_writeout",
    "get_nr_swap_pages_eligible",
    "mem_cgroup_get_nr_swap_pages_eligible",
    "folio_swap_full",
    "user_proactive_reclaim",
    "__x64_sys_swapon"
  ],
  "Reasoning": "The patch series modifies core memory management code (mm/swapfile.c, mm/vmscan.c, mm/page_io.c, mm/memcontrol.c) to introduce offload-only swap areas via the new SWAP_FLAG_OFFLOAD_ONLY flag in swapon(). It alters swap slot allocation (both fast and slow paths), eligible swap accounting, swapcache reclaim heuristics, swap writeout refusal paths for retained entries, and zswap integration. Furthermore, it adds new runtime assertions (VM_WARN_ON_FOLIO) in folio_swap_full(). All modified paths are fully reachable in virtualized environments via swapon(), cgroup v2 memory.reclaim, and standard memory reclaim under pressure.",
  "WorthFuzzing": true
}

Results:
{
  "EnableConfigs": [
    "SWAP",
    "MEMCG",
    "ZSWAP",
    "VM_EVENT_COUNTERS"
  ],
  "FocusSymbols": [
    "folio_alloc_swap",
    "swap_writeout",
    "get_nr_swap_pages_eligible",
    "mem_cgroup_get_nr_swap_pages_eligible",
    "folio_swap_full",
    "user_proactive_reclaim",
    "__x64_sys_swapon"
  ],
  "Reasoning": "The patch series modifies core memory management code (mm/swapfile.c, mm/vmscan.c, mm/page_io.c, mm/memcontrol.c) to introduce offload-only swap areas via the new SWAP_FLAG_OFFLOAD_ONLY flag in swapon(). It alters swap slot allocation (both fast and slow paths), eligible swap accounting, swapcache reclaim heuristics, swap writeout refusal paths for retained entries, and zswap integration. Furthermore, it adds new runtime assertions (VM_WARN_ON_FOLIO) in folio_swap_full(). All modified paths are fully reachable in virtualized environments via swapon(), cgroup v2 memory.reclaim, and standard memory reclaim under pressure.",
  "WorthFuzzing": true
}

49/1 2026/09/14 12:43 action If
2m Args:
{
  "WorthFuzzing": true
}

Results:
null

50/2 2026/09/14 12:43 agent kmsan-evaluator
2m Model:
core

Args:
null

Results:
{
  "KMSANReasoning": "The patch series introduces offload-only swap areas (SWAP_FLAG_OFFLOAD_ONLY / SWP_OFFLOAD_ONLY), allowing swap devices to be designated solely for proactive reclaim (e.g. cgroup v2 memory.reclaim).\n\nKMSAN vs KASAN Applicability Analysis:\n- Uninitialized Memory Risks (KMSAN):\n  1. No new structures or buffers are copied to user space; user-space interfaces modified are the input flag to the `swapon` syscall and an integer event counter reported via `/proc/vmstat`.\n  2. The new fields added to kernel structures (`allow_offload_swap` in `struct reclaim_state`, `struct scan_control`, and `struct percpu_swap_cluster`) are always initialized before use: `scan_control` instances across vmscan are zero-initialized via C99 designated initializers, `percpu_swap_cluster` is defined in static per-CPU storage, and `current-\u003ereclaim_state` is strictly scoped within `set_task_reclaim_state()`.\n  3. No branching or conditional logic depends on uninitialized data.\n\n- Memory Corruption / Concurrency Risks (KASAN / LOCKDEP):\n  1. Any potential risks introduced by the patch involve swap slot accounting, cluster selection logic, priority routing, or device-level locking/synchronization.\n  2. These failure modes (use-after-free, double frees, invalid dereferences, locking deadlocks) are already covered by standard KASAN and LOCKDEP configurations.\n\nBecause there are no uninitialized memory use or kernel info-leak risks introduced by this patch series, spawning a dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

Instruction:
You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.

CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
  a comprehensive suite of debugging tools and sanitizers, including KASAN
  (out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
  (locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
  or page allocations) and kernel-to-user memory info-leaks.

Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.

Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.

Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
   attributes, ioctl output arguments, socket options, or BPF buffers) where fields
   or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
   or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
   could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
   uninitialized bytes of existing buffers.

Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
  dereferences, locking deadlocks, or use-after-free bugs (these are already caught
  by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
  or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.

Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.


Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.

Prompt:
Target architecture: amd64

For your convenience, here is the diff of the changes:
commit b18205544615b8a7fdfb59a7d24849375e130428
Author: syz-cluster <triage@syzkaller.com>
Date:   Mon Sep 14 12:10:57 2026 +0000

    syz-cluster: applied patch under review

diff --git a/Documentation/mm/swap.rst b/Documentation/mm/swap.rst
index 78819bd4d7452..7d5bbd65e6160 100644
--- a/Documentation/mm/swap.rst
+++ b/Documentation/mm/swap.rst
@@ -3,3 +3,90 @@
 ====
 Swap
 ====
+
+Offload-only swap areas
+-----------------------
+
+``SWAP_FLAG_OFFLOAD_ONLY`` marks a swap area as a destination for explicit
+userspace proactive reclaim.  The stable qualifying interfaces are cgroup v2
+``memory.reclaim`` and ``/sys/devices/system/node/nodeX/reclaim``.  The swap
+allocator excludes such an area from kswapd, direct reclaim, and other
+pressure-driven swap allocation.  Normal swap priority ordering still applies
+among the areas eligible for the current reclaim context.
+
+The MGLRU debugfs eviction interface currently establishes the same internal
+proactive-reclaim provenance and can therefore use an offload-only area.
+Debugfs is not a stable userspace ABI, however, so that behaviour is not part
+of this interface's permanent contract.
+
+This permits a system to combine a small conventional swap area, which is
+engineered for forward progress in emergency reclaim, with a larger or more
+complex area used for ordinary cold-page offload.  For example, the latter may
+be RAM-compressed or may use a filesystem with compression, checksums, or
+redundancy.  Making every such write path safe in direct reclaim can require
+backend-specific reserves, preallocation, non-blocking allocation, and
+recursion rules.  Excluding new pressure-reclaim writes can reduce that
+requirement and its complexity.
+
+For a RAM-compressed area such as zram, unused logical slots also do not imply
+that enough physical memory remains to store their future contents.  Static
+swap priority cannot express that distinction or provide late fallback after a
+selected area's write fails.
+
+The policy is attached to an activated swap area, not to its underlying
+physical storage.  A raw swap partition and a filesystem swapfile on the same
+device are separate areas and may use different policies.  The kernel does not
+infer this policy from the block driver, filesystem, or swap priority.
+
+This is a reclaim-provenance policy, not a measurement of current memory
+headroom.  Userspace should only request proactive offload while its own
+watermark or PSI policy considers memory pressure low.
+
+DAMON reclaim and ``MADV_PAGEOUT`` do not currently establish the proactive
+reclaim context, so they cannot allocate slots from an offload-only area.
+Offload-only areas are also ineligible for hibernation image allocation.
+
+Offload-only areas bypass zswap stores.  Zswap writeback may run after the
+proactive context which selected the slot has ended, so admitting the folio to
+zswap would otherwise defer the backend write beyond that context.  The
+hierarchical cgroup v2 ``memory.zswap.writeback=0`` policy remains
+authoritative: when zswap is enabled, it also refuses direct proactive writes
+to an offload-only area.  Marking an area offload-only does not override a
+cgroup policy which disables all swapping attempts to devices.
+
+The flag controls both allocation of new swap slots and newly initiated
+non-zero backend writes.  A folio can retain its swap entry after swapin.  If
+ordinary reclaim later tries to rewrite such an offload-only entry, the VM
+redirties and activates the folio instead; proactive reclaim may retry the
+write.  Zero-filled folios may still update the in-memory swap zeromap without
+backend I/O.
+
+The flag does not prevent reads, swapoff, or writes which are already queued or
+in flight.  It therefore does not by itself provide a forward progress
+guarantee for an I/O path which allocates memory: earlier writes must still be
+able to complete, and reads must remain reclaim-safe.  Repeatedly refusing
+retained-entry writes can also reduce reclaim efficiency and lead to OOM while
+the dirty folios remain resident.
+
+Page-cluster discard is incompatible with an offload-only area because its
+work item can run after the context which freed the entries has ended.  Swapon
+therefore rejects a resolved page-cluster discard policy combined with
+``SWAP_FLAG_OFFLOAD_ONLY``.  Swapon-time discard is permitted because it
+completes synchronously during activation.  The existing discard precedence
+still applies: requesting both discard-once and discard-pages selects
+discard-once.  Discard requests which the swap area does not support remain
+ignored.
+
+Architecture-specific swap metadata preparation still runs before a retained
+write is refused, so that metadata remains coherent with the dirty resident
+folio.  This policy controls swap-backend I/O; it does not promise that core VM
+or architecture preparation performs no allocation.
+
+``/proc/vmstat`` reports ``swpout_offload_refused`` in base pages.  The counter
+advances when ordinary reclaim refuses a newly initiated write through a
+retained offload-only entry.  It can reveal repeated activation of such folios,
+but it is not a count of skipped areas during new-slot allocation.
+
+An offload-only area should therefore be configured with a reclaim-safe swap
+area as fallback.  If no eligible swap space remains, swap allocation fails
+and the existing reclaim and OOM policy applies.
diff --git a/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c b/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
index e0d1f369a1632..b58e61f15ab10 100644
--- a/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
+++ b/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
@@ -21,7 +21,7 @@
 
 static bool swap_available(void)
 {
-	return get_nr_swap_pages() > 0;
+	return get_nr_swap_pages_eligible() > 0;
 }
 
 static bool can_release_pages(struct drm_i915_gem_object *obj)
diff --git a/drivers/gpu/drm/i915/gem/selftests/huge_pages.c b/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
index 44718e7282914..73b62b0655103 100644
--- a/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
+++ b/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
@@ -1895,8 +1895,8 @@ static int igt_shrink_thp(void *arg)
 	i915_gem_context_unlock_engines(ctx);
 	/*
 	 * Nuke everything *before* we unpin the pages so we can be reasonably
-	 * sure that when later checking get_nr_swap_pages() that some random
-	 * leftover object doesn't steal the remaining swap space.
+	 * sure that when later checking get_nr_swap_pages_eligible() that some
+	 * random leftover object doesn't steal the remaining swap space.
 	 */
 	i915_gem_shrink(NULL, i915, -1UL, NULL,
 			I915_SHRINK_BOUND |
@@ -1910,7 +1910,7 @@ static int igt_shrink_thp(void *arg)
 	 * Now that the pages are *unpinned* shrinking should invoke
 	 * shmem to truncate our pages, if we have available swap.
 	 */
-	should_swap = get_nr_swap_pages() > 0;
+	should_swap = get_nr_swap_pages_eligible() > 0;
 	i915_gem_shrink(NULL, i915, -1UL, NULL,
 			I915_SHRINK_BOUND |
 			I915_SHRINK_UNBOUND |
diff --git a/drivers/gpu/drm/msm/msm_gem_shrinker.c b/drivers/gpu/drm/msm/msm_gem_shrinker.c
index 9d2788f79acee..da6a66b747b02 100644
--- a/drivers/gpu/drm/msm/msm_gem_shrinker.c
+++ b/drivers/gpu/drm/msm/msm_gem_shrinker.c
@@ -21,7 +21,7 @@ module_param(enable_eviction, bool, 0600);
 
 static bool can_swap(void)
 {
-	return enable_eviction && get_nr_swap_pages() > 0;
+	return enable_eviction && get_nr_swap_pages_eligible() > 0;
 }
 
 static bool can_block(struct shrink_control *sc)
diff --git a/drivers/gpu/drm/panthor/panthor_gem.c b/drivers/gpu/drm/panthor/panthor_gem.c
index 72908be5e144f..23488f417cf7a 100644
--- a/drivers/gpu/drm/panthor/panthor_gem.c
+++ b/drivers/gpu/drm/panthor/panthor_gem.c
@@ -1371,7 +1371,7 @@ panthor_dummy_bo_create(struct panthor_device *ptdev)
 
 static bool can_swap(void)
 {
-	return get_nr_swap_pages() > 0;
+	return get_nr_swap_pages_eligible() > 0;
 }
 
 static bool can_block(struct shrink_control *sc)
diff --git a/drivers/gpu/drm/ttm/ttm_backup.c b/drivers/gpu/drm/ttm/ttm_backup.c
index 0c2d53a13b2a6..bf9d41bdf652c 100644
--- a/drivers/gpu/drm/ttm/ttm_backup.c
+++ b/drivers/gpu/drm/ttm/ttm_backup.c
@@ -206,7 +206,7 @@ u64 ttm_backup_bytes_avail(void)
 	 * number also depends on shmem actually swapping out backed-up
 	 * shmem objects without too much buffering.
 	 */
-	return (u64)get_nr_swap_pages() << PAGE_SHIFT;
+	return (u64)get_nr_swap_pages_eligible() << PAGE_SHIFT;
 }
 EXPORT_SYMBOL_GPL(ttm_backup_bytes_avail);
 
diff --git a/drivers/gpu/drm/xe/tests/xe_bo.c b/drivers/gpu/drm/xe/tests/xe_bo.c
index 6a17e13d58cf1..14d6bb8e41c9c 100644
--- a/drivers/gpu/drm/xe/tests/xe_bo.c
+++ b/drivers/gpu/drm/xe/tests/xe_bo.c
@@ -695,7 +695,7 @@ static int shrink_test_run_device(struct xe_device *xe)
 	}
 	to_alloc = ram * 2;
 
-	ram_and_swap = ram + get_nr_swap_pages() * PAGE_SIZE;
+	ram_and_swap = ram + get_nr_swap_pages_eligible() * PAGE_SIZE;
 	if (to_alloc > ram_and_swap)
 		purgeable = to_alloc - ram_and_swap;
 	purgeable += div64_u64(purgeable, 5);
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 43155e122b5c3..19a0bb4ba6c93 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -21,10 +21,11 @@
 #define SWAP_FLAG_DISCARD	0x10000 /* enable discard for swap */
 #define SWAP_FLAG_DISCARD_ONCE	0x20000 /* discard swap area at swapon-time */
 #define SWAP_FLAG_DISCARD_PAGES 0x40000 /* discard page-clusters after use */
+#define SWAP_FLAG_OFFLOAD_ONLY  0x80000 /* only use for proactive reclaim */
 
 #define SWAP_FLAGS_VALID	(SWAP_FLAG_PRIO_MASK | SWAP_FLAG_PREFER | \
 				 SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE | \
-				 SWAP_FLAG_DISCARD_PAGES)
+				 SWAP_FLAG_DISCARD_PAGES | SWAP_FLAG_OFFLOAD_ONLY)
 /*
  * MAX_SWAPFILES defines the maximum number of swaptypes: things which can
  * be swapped to.  The swap type and the offset into that swap type are
@@ -140,12 +141,20 @@ union swap_header {
 struct reclaim_state {
 	/* pages reclaimed outside of LRU-based reclaim */
 	unsigned long reclaimed;
+	/* this reclaim context may use offload-only swap */
+	bool allow_offload_swap;
 #ifdef CONFIG_LRU_GEN
 	/* per-thread mm walk data */
 	struct lru_gen_mm_walk *mm_walk;
 #endif
 };
 
+static inline bool current_reclaim_allows_offload_swap(void)
+{
+	return current->reclaim_state &&
+	       current->reclaim_state->allow_offload_swap;
+}
+
 /*
  * mm_account_reclaimed_pages(): account reclaimed pages outside of LRU-based
  * reclaim
@@ -201,6 +210,7 @@ enum {
 	SWP_STABLE_WRITES = (1 << 11),	/* no overwrite PG_writeback pages */
 	SWP_SYNCHRONOUS_IO = (1 << 12),	/* synchronous IO is efficient */
 	SWP_HIBERNATION = (1 << 13),	/* pinned for hibernation */
+	SWP_OFFLOAD_ONLY = (1 << 14),	/* proactive-reclaim swap only */
 					/* add others here before... */
 };
 
@@ -389,6 +399,9 @@ static inline long get_nr_swap_pages(void)
 	return atomic_long_read(&nr_swap_pages);
 }
 
+long get_nr_swap_pages_eligible(void);
+bool folio_swap_full(struct folio *folio);
+
 extern void si_swapinfo(struct sysinfo *);
 extern int pin_hibernation_swap_type(dev_t device, sector_t offset);
 extern void unpin_hibernation_swap_type(int type);
@@ -443,10 +456,16 @@ static inline void put_swap_device(struct swap_info_struct *si)
 }
 
 #define get_nr_swap_pages()			0L
+#define get_nr_swap_pages_eligible()		0L
 #define total_swap_pages			0L
 #define total_swapcache_pages()			0UL
 #define vm_swap_full()				0
 
+static inline bool folio_swap_full(struct folio *folio)
+{
+	return false;
+}
+
 #define si_swapinfo(val) \
 	do { (val)->freeswap = (val)->totalswap = 0; } while (0)
 #define free_folio_and_swap_cache(folio) \
@@ -531,6 +550,7 @@ static inline void mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_p
 
 long mem_cgroup_get_folio_swap_margin(struct folio *folio);
 extern long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg);
+long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg);
 extern bool mem_cgroup_swap_full(struct folio *folio);
 #else
 static inline int mem_cgroup_try_charge_swap(struct folio *folio)
@@ -553,9 +573,14 @@ static inline long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
 	return get_nr_swap_pages();
 }
 
+static inline long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg)
+{
+	return get_nr_swap_pages_eligible();
+}
+
 static inline bool mem_cgroup_swap_full(struct folio *folio)
 {
-	return vm_swap_full();
+	return folio_swap_full(folio);
 }
 #endif
 
diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h
index 2628ccda076a0..fc8458b314a50 100644
--- a/include/linux/vm_event_item.h
+++ b/include/linux/vm_event_item.h
@@ -32,6 +32,7 @@
 	HIGHMEM_ZONE(xx) xx##_MOVABLE, DEVICE_ZONE(xx)
 
 enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT,
+		SWPOUT_OFFLOAD_REFUSED,
 		FOR_ALL_ZONES(PGALLOC)
 		FOR_ALL_ZONES(ALLOCSTALL)
 		FOR_ALL_ZONES(PGSCAN_SKIP)
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1460cba53588e..1412084d2f430 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -6002,16 +6002,28 @@ void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages)
 	rcu_read_unlock();
 }
 
-long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
+static long
+mem_cgroup_get_nr_swap_pages_with_limit(struct mem_cgroup *memcg,
+					long nr_swap_pages)
 {
-	long nr_swap_pages = get_nr_swap_pages();
-
 	if (!mem_cgroup_disabled() && !do_memsw_account())
 		nr_swap_pages = min(nr_swap_pages, page_counter_margin(&memcg->swap));
 
 	return nr_swap_pages;
 }
 
+long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
+{
+	return mem_cgroup_get_nr_swap_pages_with_limit(memcg,
+			get_nr_swap_pages());
+}
+
+long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg)
+{
+	return mem_cgroup_get_nr_swap_pages_with_limit(memcg,
+			get_nr_swap_pages_eligible());
+}
+
 /**
  * mem_cgroup_get_folio_swap_margin - get a folio's memcg swap margin
  * @folio: folio whose memcg margin is queried
@@ -6042,7 +6054,7 @@ bool mem_cgroup_swap_full(struct folio *folio)
 
 	VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
 
-	if (vm_swap_full())
+	if (folio_swap_full(folio))
 		return true;
 	if (do_memsw_account() || !folio_memcg_charged(folio))
 		return ret;
diff --git a/mm/page_io.c b/mm/page_io.c
index 1da4ff484f097..1b281a1d6df0a 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -203,6 +203,7 @@ static void swap_zeromap_folio_clear(struct folio *folio)
  */
 int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
 {
+	struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
 	int ret = 0;
 
 	if (folio_free_swap(folio))
@@ -210,7 +211,9 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
 
 	/*
 	 * Arch code may have to preserve more data than just the folio
-	 * contents, e.g. memory tags.
+	 * contents, e.g. memory tags.  Do this before refusing a retained
+	 * offload-only entry below: a later sibling swap-PTE fault can restore
+	 * swap-indexed metadata into this resident folio.
 	 */
 	ret = arch_prepare_to_swap(folio);
 	if (ret) {
@@ -228,6 +231,18 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
 		goto out_unlock;
 	}
 
+	/*
+	 * A folio can retain an existing swap entry after swapin.  Do not let
+	 * ordinary reclaim use an offload-only entry through that path.
+	 */
+	if ((READ_ONCE(sis->flags) & SWP_OFFLOAD_ONLY) &&
+	    !current_reclaim_allows_offload_swap()) {
+		count_vm_events(SWPOUT_OFFLOAD_REFUSED,
+				folio_nr_pages(folio));
+		folio_mark_dirty(folio);
+		return AOP_WRITEPAGE_ACTIVATE;
+	}
+
 	/*
 	 * Clear bits this folio occupies in the zeromap to prevent zero data
 	 * being read in from any previous zero writes that occupied the same
@@ -235,7 +250,12 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
 	 */
 	swap_zeromap_folio_clear(folio);
 
-	if (zswap_store(folio)) {
+	/*
+	 * Zswap writeback can happen much later from pressure reclaim or its
+	 * shrinker workqueue.  Do not let it defer an offload-only backend write
+	 * beyond the proactive reclaim context which admitted the swap slot.
+	 */
+	if (!(READ_ONCE(sis->flags) & SWP_OFFLOAD_ONLY) && zswap_store(folio)) {
 		count_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT);
 		goto out_unlock;
 	}
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 280dd906eb187..256617059f302 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -65,13 +65,17 @@ static void move_cluster(struct swap_info_struct *si,
  */
 static DEFINE_SPINLOCK(swap_lock);
 static unsigned int nr_swapfiles;
-atomic_long_t nr_swap_pages;
 /*
  * Some modules use swappable objects and may try to swap them out under
  * memory pressure (via the shrinker). Before doing so, they may wish to
  * check to see if any swap space is available.
+ *
+ * This remains the raw free-space counter for accounting users. Reclaim
+ * decisions subtract nr_swap_pages_offload_only below when necessary.
  */
+atomic_long_t nr_swap_pages;
 EXPORT_SYMBOL_GPL(nr_swap_pages);
+static atomic_long_t nr_swap_pages_offload_only;
 /* protected with swap_lock. reading in vm_swap_full() doesn't need lock */
 long total_swap_pages;
 #define DEF_SWAP_PRIO  -1
@@ -120,6 +124,7 @@ atomic_t nr_rotate_swap = ATOMIC_INIT(0);
 struct percpu_swap_cluster {
 	struct swap_info_struct *si[SWAP_NR_ORDERS];
 	unsigned long offset[SWAP_NR_ORDERS];
+	bool allow_offload_swap[SWAP_NR_ORDERS];
 	local_lock_t lock;
 };
 
@@ -163,6 +168,31 @@ static long swap_usage_in_pages(struct swap_info_struct *si)
 	return atomic_long_read(&si->inuse_pages) & SWAP_USAGE_COUNTER_MASK;
 }
 
+static bool swap_area_needs_reclaim(struct swap_info_struct *si)
+{
+	if (vm_swap_full())
+		return true;
+
+	/*
+	 * Free offload-only slots must not keep a full conventional area
+	 * pinned in swapcache.  Recover that area's unused cache entries
+	 * even when the raw pool is not full.  This is independent of the
+	 * current task: proactive reclaim can fill the conventional area too.
+	 */
+	return !(READ_ONCE(si->flags) & SWP_OFFLOAD_ONLY) &&
+		atomic_long_read(&nr_swap_pages_offload_only) > 0 &&
+		swap_usage_in_pages(si) == si->pages;
+}
+
+/* The caller must hold the lock on a folio in swapcache. */
+bool folio_swap_full(struct folio *folio)
+{
+	VM_WARN_ON_FOLIO(!folio_test_locked(folio), folio);
+	VM_WARN_ON_FOLIO(!folio_test_swapcache(folio), folio);
+
+	return swap_area_needs_reclaim(__swap_entry_to_info(folio->swap));
+}
+
 /* Reclaim the swap entry anyway if possible */
 #define TTRS_ANYWAY		0x1
 /*
@@ -903,7 +933,7 @@ static bool cluster_scan_range(struct swap_info_struct *si,
 		if (swp_tb_is_null(swp_tb))
 			continue;
 		if (swp_tb_is_folio(swp_tb) && !__swp_tb_get_count(swp_tb)) {
-			if (!vm_swap_full())
+			if (!swap_area_needs_reclaim(si))
 				return false;
 			*need_reclaim = true;
 			continue;
@@ -1014,6 +1044,8 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
 	if (si->flags & SWP_SOLIDSTATE) {
 		this_cpu_write(percpu_swap_cluster.offset[order], next);
 		this_cpu_write(percpu_swap_cluster.si[order], si);
+		this_cpu_write(percpu_swap_cluster.allow_offload_swap[order],
+			       current_reclaim_allows_offload_swap());
 	} else {
 		si->global_cluster->next[order] = next;
 	}
@@ -1161,7 +1193,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
 	}
 
 	/* Try reclaim full clusters if free and nonfull lists are drained */
-	if (vm_swap_full())
+	if (swap_area_needs_reclaim(si))
 		swap_reclaim_full_clusters(si, false);
 
 	if (order < PMD_ORDER) {
@@ -1315,9 +1347,11 @@ static void swap_range_alloc(struct swap_info_struct *si,
 			     unsigned int nr_entries)
 {
 	if (swap_usage_add(si, nr_entries)) {
-		if (vm_swap_full())
+		if (swap_area_needs_reclaim(si))
 			schedule_work(&si->reclaim_work);
 	}
+	if (si->flags & SWP_OFFLOAD_ONLY)
+		atomic_long_sub(nr_entries, &nr_swap_pages_offload_only);
 	atomic_long_sub(nr_entries, &nr_swap_pages);
 }
 
@@ -1346,6 +1380,8 @@ static void swap_range_free(struct swap_info_struct *si, unsigned long offset,
 	 * only after the above cleanups are done.
 	 */
 	smp_wmb();
+	if (si->flags & SWP_OFFLOAD_ONLY)
+		atomic_long_add(nr_entries, &nr_swap_pages_offload_only);
 	atomic_long_add(nr_entries, &nr_swap_pages);
 	swap_usage_sub(si, nr_entries);
 }
@@ -1366,6 +1402,31 @@ static bool get_swap_device_info(struct swap_info_struct *si)
 	return true;
 }
 
+static bool swap_area_eligible(struct swap_info_struct *si)
+{
+	if (!(READ_ONCE(si->flags) & SWP_OFFLOAD_ONLY))
+		return true;
+
+	return current_reclaim_allows_offload_swap();
+}
+
+long get_nr_swap_pages_eligible(void)
+{
+	long nr_pages;
+
+	if (current_reclaim_allows_offload_swap())
+		return get_nr_swap_pages();
+
+	/*
+	 * The reads are intentionally unpaired.  This is a capacity hint; the
+	 * allocator enforces eligibility.  Clamp a transient negative result.
+	 */
+	nr_pages = get_nr_swap_pages() -
+		   atomic_long_read(&nr_swap_pages_offload_only);
+	return max(nr_pages, 0L);
+}
+EXPORT_SYMBOL_GPL(get_nr_swap_pages_eligible);
+
 /*
  * Fast path try to get swap entries with specified order from current
  * CPU's swap entry pool (a cluster).
@@ -1385,6 +1446,20 @@ static bool swap_alloc_fast(struct folio *folio)
 	offset = this_cpu_read(percpu_swap_cluster.offset[order]);
 	if (!si || !offset || !get_swap_device_info(si))
 		return false;
+	if (!swap_area_eligible(si)) {
+		put_swap_device(si);
+		return false;
+	}
+	/*
+	 * Pressure reclaim may cache a lower-priority conventional area while
+	 * an offload-only area is ineligible.  Drop that cache on a context
+	 * change so proactive reclaim returns to the normal priority search.
+	 */
+	if (this_cpu_read(percpu_swap_cluster.allow_offload_swap[order]) !=
+	    current_reclaim_allows_offload_swap()) {
+		put_swap_device(si);
+		return false;
+	}
 
 	ci = swap_cluster_lock(si, offset);
 	if (cluster_is_usable(ci, order)) {
@@ -1407,6 +1482,9 @@ static void swap_alloc_slow(struct folio *folio)
 	spin_lock(&swap_avail_lock);
 start_over:
 	plist_for_each_entry_safe(si, next, &swap_avail_head, avail_list) {
+		if (!swap_area_eligible(si))
+			continue;
+
 		/* Rotate the device and switch to a new cluster */
 		plist_requeue(&si->avail_list, &swap_avail_head);
 		spin_unlock(&swap_avail_lock);
@@ -1450,7 +1528,8 @@ static bool swap_sync_discard(void)
 	plist_for_each_entry_safe(si, next, &swap_active_head, list) {
 		spin_unlock(&swap_lock);
 		if (get_swap_device_info(si)) {
-			if (si->flags & SWP_PAGE_DISCARD)
+			if (swap_area_eligible(si) &&
+			    (si->flags & SWP_PAGE_DISCARD))
 				ret = swap_do_scheduled_discard(si);
 			put_swap_device(si);
 		}
@@ -1738,8 +1817,8 @@ static int swap_dup_entries_cluster(struct swap_info_struct *si,
  *
  * Context: Caller needs to hold the folio lock.
  * Return: %0 on success, %-E2BIG if splitting the folio might allow swapout,
- * %-ENOSPC if no global swap space is available, or %-ENOMEM if splitting
- * would not help.
+ * %-ENOSPC if no global swap space is eligible for the caller, or %-ENOMEM
+ * if splitting would not help.
  */
 int folio_alloc_swap(struct folio *folio)
 {
@@ -1790,7 +1869,7 @@ int folio_alloc_swap(struct folio *folio)
 	return 0;
 
 failed:
-	if (get_nr_swap_pages() <= 0)
+	if (get_nr_swap_pages_eligible() <= 0)
 		return -ENOSPC;
 	if (mem_cgroup_get_folio_swap_margin(folio) <= 0)
 		return -ENOMEM;
@@ -2180,7 +2259,7 @@ swp_entry_t swap_alloc_hibernation_slot(int type)
 	struct swap_cluster_info *ci;
 	swp_entry_t entry = {0};
 
-	if (!si)
+	if (!si || (si->flags & SWP_OFFLOAD_ONLY))
 		goto fail;
 
 	/*
@@ -2247,7 +2326,8 @@ static int __find_hibernation_swap_type(dev_t device, sector_t offset)
 	for (type = 0; type < nr_swapfiles; type++) {
 		struct swap_info_struct *sis = swap_info[type];
 
-		if (!(sis->flags & SWP_WRITEOK))
+		if (!(sis->flags & SWP_WRITEOK) ||
+		    (sis->flags & SWP_OFFLOAD_ONLY))
 			continue;
 
 		if (device == sis->bdev->bd_dev) {
@@ -2434,7 +2514,8 @@ int find_first_swap(dev_t *device)
 	for (type = 0; type < nr_swapfiles; type++) {
 		struct swap_info_struct *sis = swap_info[type];
 
-		if (!(sis->flags & SWP_WRITEOK))
+		if (!(sis->flags & SWP_WRITEOK) ||
+		    (sis->flags & SWP_OFFLOAD_ONLY))
 			continue;
 		*device = sis->bdev->bd_dev;
 		spin_unlock(&swap_lock);
@@ -2474,7 +2555,8 @@ unsigned int count_swap_pages(int type, int free)
 		struct swap_info_struct *sis = swap_info[type];
 
 		spin_lock(&sis->lock);
-		if (sis->flags & SWP_WRITEOK) {
+		if ((sis->flags & SWP_WRITEOK) &&
+		    !(sis->flags & SWP_OFFLOAD_ONLY)) {
 			n = sis->pages;
 			if (free)
 				n -= swap_usage_in_pages(sis);
@@ -3083,6 +3165,8 @@ static int setup_swap_extents(struct swap_info_struct *sis,
 
 static void _enable_swap_info(struct swap_info_struct *si)
 {
+	if (si->flags & SWP_OFFLOAD_ONLY)
+		atomic_long_add(si->pages, &nr_swap_pages_offload_only);
 	atomic_long_add(si->pages, &nr_swap_pages);
 	total_swap_pages += si->pages;
 
@@ -3231,6 +3315,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
 	spin_lock(&p->lock);
 	del_from_avail_list(p, true);
 	plist_del(&p->list, &swap_active_head);
+	if (p->flags & SWP_OFFLOAD_ONLY)
+		atomic_long_sub(p->pages, &nr_swap_pages_offload_only);
 	atomic_long_sub(p->pages, &nr_swap_pages);
 	total_swap_pages -= p->pages;
 	spin_unlock(&p->lock);
@@ -3727,7 +3813,6 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
 
 	if (swap_flags & ~SWAP_FLAGS_VALID)
 		return -EINVAL;
-
 	if (!capable(CAP_SYS_ADMIN))
 		return -EPERM;
 
@@ -3841,6 +3926,9 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
 	if (error)
 		goto bad_swap_unlock_inode;
 
+	if (swap_flags & SWAP_FLAG_OFFLOAD_ONLY)
+		si->flags |= SWP_OFFLOAD_ONLY;
+
 	if ((swap_flags & SWAP_FLAG_DISCARD) &&
 	    si->bdev && bdev_max_discard_sectors(si->bdev)) {
 		/*
@@ -3863,6 +3951,18 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
 		else if (swap_flags & SWAP_FLAG_DISCARD_PAGES)
 			si->flags &= ~SWP_AREA_DISCARD;
 
+		/*
+		 * Cluster discard can run later from discard_work, after the
+		 * context which freed the entries has ended.  Swapon-time discard
+		 * is explicit and synchronous, but page discard cannot honour
+		 * offload provenance.
+		 */
+		if ((si->flags & SWP_OFFLOAD_ONLY) &&
+		    (si->flags & SWP_PAGE_DISCARD)) {
+			error = -EINVAL;
+			goto bad_swap_unlock_inode;
+		}
+
 		/* issue a swapon-time discard if it's still required */
 		if (si->flags & SWP_AREA_DISCARD) {
 			int err = discard_swap(si);
diff --git a/mm/vmscan.c b/mm/vmscan.c
index c2eb8fa9d5e50..201b71b760492 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -123,6 +123,9 @@ struct scan_control {
 	/* Proactive reclaim invoked by userspace */
 	unsigned int proactive:1;
 
+	/* This reclaim context may use offload-only swap */
+	unsigned int allow_offload_swap:1;
+
 	/*
 	 * Cgroup memory below memory.low is protected as long as we
 	 * don't threaten to OOM. If any cgroup is reclaimed at
@@ -291,14 +294,18 @@ static inline bool is_exec_file_folio(const struct folio *folio,
 }
 
 static void set_task_reclaim_state(struct task_struct *task,
-				   struct reclaim_state *rs)
+				   struct scan_control *sc)
 {
+	struct reclaim_state *rs = sc ? &sc->reclaim_state : NULL;
+
 	/* Check for an overwrite */
 	WARN_ON_ONCE(rs && task->reclaim_state);
 
 	/* Check for the nulling of an already-nulled member */
 	WARN_ON_ONCE(!rs && !task->reclaim_state);
 
+	if (rs)
+		rs->allow_offload_swap = sc->allow_offload_swap;
 	task->reclaim_state = rs;
 }
 
@@ -418,7 +425,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,
 		 * And under GFP_NOIO, is there enough swapcached anon to make
 		 * scanning anon worthwhile?
 		 */
-		if (get_nr_swap_pages() > 0 &&
+		if (get_nr_swap_pages_eligible() > 0 &&
 		    !reclaimable_anon_is_low(memcg, nid, sc))
 			return true;
 	} else {
@@ -426,7 +433,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,
 		 * Is the memcg below its swap limit, and under GFP_NOIO does
 		 * it have enough swapcached anon to make scanning worthwhile?
 		 */
-		if (mem_cgroup_get_nr_swap_pages(memcg) > 0 &&
+		if (mem_cgroup_get_nr_swap_pages_eligible(memcg) > 0 &&
 		    !reclaimable_anon_is_low(memcg, nid, sc))
 			return true;
 	}
@@ -2853,7 +2860,7 @@ static int get_swappiness(struct lruvec *lruvec, struct scan_control *sc)
 		return 0;
 
 	if (!can_demote(pgdat->node_id, sc, memcg) &&
-	    mem_cgroup_get_nr_swap_pages(memcg) < MIN_LRU_BATCH)
+	    mem_cgroup_get_nr_swap_pages_eligible(memcg) < MIN_LRU_BATCH)
 		return 0;
 
 	return swappiness;
@@ -5926,6 +5933,7 @@ static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,
 		.reclaim_idx = MAX_NR_ZONES - 1,
 		.gfp_mask = GFP_KERNEL,
 		.proactive = true,
+		.allow_offload_swap = true,
 	};
 
 	buf = kvmalloc(len + 1, GFP_KERNEL);
@@ -5937,7 +5945,7 @@ static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,
 		return -EFAULT;
 	}
 
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 	flags = memalloc_noreclaim_save();
 	blk_start_plug(&plug);
 	if (!set_mm_walk(NULL, true)) {
@@ -6941,7 +6949,7 @@ unsigned long try_to_free_pages(struct zonelist *zonelist, int order,
 	if (throttle_direct_reclaim(sc.gfp_mask, zonelist, nodemask))
 		return 1;
 
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 	trace_mm_vmscan_direct_reclaim_begin(sc.gfp_mask, order, NULL);
 
 	nr_reclaimed = do_try_to_free_pages(zonelist, &sc);
@@ -6974,6 +6982,8 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,
 		.may_unmap = 1,
 		.may_swap = !!(reclaim_options & MEMCG_RECLAIM_MAY_SWAP),
 		.proactive = !!(reclaim_options & MEMCG_RECLAIM_PROACTIVE),
+		.allow_offload_swap =
+			!!(reclaim_options & MEMCG_RECLAIM_PROACTIVE),
 	};
 	/*
 	 * Traverse the ZONELIST_FALLBACK zonelist of the current node to put
@@ -6982,7 +6992,7 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,
 	 */
 	struct zonelist *zonelist = node_zonelist(numa_node_id(), sc.gfp_mask);
 
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 	trace_mm_vmscan_memcg_reclaim_begin(sc.gfp_mask, 0, memcg);
 	noreclaim_flag = memalloc_noreclaim_save();
 
@@ -7272,7 +7282,7 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx)
 
 	trace_mm_vmscan_balance_pgdat_begin(pgdat->node_id, order,
 					    highest_zoneidx);
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 	psi_memstall_enter(&pflags);
 	__fs_reclaim_acquire(_THIS_IP_);
 
@@ -7769,7 +7779,7 @@ unsigned long shrink_all_memory(unsigned long nr_to_reclaim)
 
 	fs_reclaim_acquire(sc.gfp_mask);
 	noreclaim_flag = memalloc_noreclaim_save();
-	set_task_reclaim_state(current, &sc.reclaim_state);
+	set_task_reclaim_state(current, &sc);
 
 	nr_reclaimed = do_try_to_free_pages(zonelist, &sc);
 
@@ -7950,7 +7960,7 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
 	 * We need to be able to allocate from the reserves for RECLAIM_UNMAP
 	 */
 	noreclaim_flag = memalloc_noreclaim_save();
-	set_task_reclaim_state(p, &sc->reclaim_state);
+	set_task_reclaim_state(p, sc);
 
 	do {
 		shrink_node(pgdat, sc);
@@ -8136,6 +8146,7 @@ int user_proactive_reclaim(char *buf,
 				.may_unmap = 1,
 				.may_swap = 1,
 				.proactive = 1,
+				.allow_offload_swap = 1,
 			};
 
 			if (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED,
diff --git a/mm/vmstat.c b/mm/vmstat.c
index a3e809c57f295..b924c715886e9 100644
--- a/mm/vmstat.c
+++ b/mm/vmstat.c
@@ -1331,6 +1331,7 @@ const char * const vmstat_text[] = {
 	[I(PGPGOUT)]				= "pgpgout",
 	[I(PSWPIN)]				= "pswpin",
 	[I(PSWPOUT)]				= "pswpout",
+	[I(SWPOUT_OFFLOAD_REFUSED)]		= "swpout_offload_refused",
 
 #define OFF (NR_VM_ZONE_STAT_ITEMS + NR_VM_NUMA_EVENT_ITEMS + \
 	     NR_VM_NODE_STAT_ITEMS + NR_VM_STAT_ITEMS)
diff --git a/tools/testing/selftests/zram/.gitignore b/tools/testing/selftests/zram/.gitignore
index 088cd9bad87a7..74b0217c60f04 100644
--- a/tools/testing/selftests/zram/.gitignore
+++ b/tools/testing/selftests/zram/.gitignore
@@ -1,2 +1,4 @@
 # SPDX-License-Identifier: GPL-2.0-only
 err.log
+swap_offload
+workingset_offload
diff --git a/tools/testing/selftests/zram/Makefile b/tools/testing/selftests/zram/Makefile
index 7f78eb1b59cb4..b83e6a1563b20 100644
--- a/tools/testing/selftests/zram/Makefile
+++ b/tools/testing/selftests/zram/Makefile
@@ -1,9 +1,9 @@
 # SPDX-License-Identifier: GPL-2.0
 all:
 
-TEST_PROGS := zram.sh
+TEST_GEN_FILES := swap_offload workingset_offload
+TEST_PROGS := zram.sh zram03.sh zram04.sh zram05.sh
 TEST_FILES := zram01.sh zram02.sh zram_lib.sh
 EXTRA_CLEAN := err.log
 
 include ../lib.mk
-
diff --git a/tools/testing/selftests/zram/config b/tools/testing/selftests/zram/config
index e0cc47e2c7e22..018429d8f5ca8 100644
--- a/tools/testing/selftests/zram/config
+++ b/tools/testing/selftests/zram/config
@@ -1,2 +1,10 @@
+CONFIG_CGROUPS=y
+CONFIG_BLK_DEV_DM=y
+CONFIG_DM_DELAY=y
+CONFIG_MEMCG=y
+CONFIG_SWAP=y
+CONFIG_TRANSPARENT_HUGEPAGE=y
+CONFIG_VM_EVENT_COUNTERS=y
 CONFIG_ZSMALLOC=y
-CONFIG_ZRAM=m
+CONFIG_ZRAM=y
+CONFIG_ZSWAP=y
diff --git a/tools/testing/selftests/zram/swap_offload.c b/tools/testing/selftests/zram/swap_offload.c
new file mode 100644
index 0000000000000..1850b3b4809f0
--- /dev/null
+++ b/tools/testing/selftests/zram/swap_offload.c
@@ -0,0 +1,402 @@
+// SPDX-License-Identifier: GPL-2.0
+#define _GNU_SOURCE
+
+#include <errno.h>
+#include <fcntl.h>
+#include <limits.h>
+#include <sched.h>
+#include <signal.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/syscall.h>
+#include <unistd.h>
+
+#define SWAP_FLAG_PREFER	0x8000
+#define SWAP_FLAG_DISCARD	0x10000
+#define SWAP_FLAG_DISCARD_ONCE	0x20000
+#define SWAP_FLAG_DISCARD_PAGES	0x40000
+#define SWAP_FLAG_OFFLOAD_ONLY	0x80000
+#define KSFT_SKIP	4
+
+static int activate(const char *path, int priority, int discard_flags)
+{
+	int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+		priority | discard_flags;
+
+	if (syscall(SYS_swapon, path, flags)) {
+		perror("swapon");
+		return 1;
+	}
+
+	return 0;
+}
+
+static int reject_page_discard(const char *path, int priority)
+{
+	int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+		SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_PAGES | priority;
+	int ret;
+
+	errno = 0;
+	ret = syscall(SYS_swapon, path, flags);
+	if (ret == -1 && errno == EINVAL)
+		return 0;
+	if (!ret) {
+		syscall(SYS_swapoff, path);
+		fprintf(stderr, "offload-only page discard was accepted\n");
+	} else {
+		fprintf(stderr, "swapon returned unexpected error: %s\n",
+			strerror(errno));
+	}
+	return 1;
+}
+
+static int accept_discard_once_pages(const char *path, int priority)
+{
+	int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+		SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE |
+		SWAP_FLAG_DISCARD_PAGES | priority;
+
+	if (syscall(SYS_swapon, path, flags)) {
+		perror("swapon discard-once+discard-pages");
+		return 1;
+	}
+	if (syscall(SYS_swapoff, path)) {
+		perror("swapoff discard-once+discard-pages");
+		return 1;
+	}
+	return 0;
+}
+
+static int join_cgroup(const char *procs)
+{
+	char pid[32];
+	int fd, len;
+
+	fd = open(procs, O_WRONLY);
+	if (fd < 0) {
+		perror("open cgroup.procs");
+		return 1;
+	}
+
+	len = snprintf(pid, sizeof(pid), "%d\n", getpid());
+	if (write(fd, pid, len) != len) {
+		perror("write cgroup.procs");
+		close(fd);
+		return 1;
+	}
+	close(fd);
+	return 0;
+}
+
+static int pin_to_one_cpu(const char *pid_arg)
+{
+	cpu_set_t allowed, selected;
+	char *end;
+	long pid;
+	int cpu;
+
+	errno = 0;
+	pid = strtol(pid_arg, &end, 10);
+	if (errno || *end || pid <= 0) {
+		fprintf(stderr, "invalid pid: %s\n", pid_arg);
+		return 1;
+	}
+
+	if (sched_getaffinity(pid, sizeof(allowed), &allowed)) {
+		perror("sched_getaffinity");
+		return 1;
+	}
+	for (cpu = 0; cpu < CPU_SETSIZE; cpu++)
+		if (CPU_ISSET(cpu, &allowed))
+			break;
+	if (cpu == CPU_SETSIZE) {
+		fprintf(stderr, "pid %ld has no allowed CPU\n", pid);
+		return 1;
+	}
+
+	CPU_ZERO(&selected);
+	CPU_SET(cpu, &selected);
+	if (sched_setaffinity(pid, sizeof(selected), &selected)) {
+		perror("sched_setaffinity");
+		return 1;
+	}
+	return 0;
+}
+
+static int allocate(const char *size_arg, const char *procs,
+		    const char *ready, const char *verified)
+{
+	unsigned char *memory;
+	unsigned long size;
+	unsigned long page_size;
+	char *end;
+	sigset_t signals;
+	int signal;
+	int fd;
+
+	errno = 0;
+	size = strtoul(size_arg, &end, 0);
+	if (errno || *end || !size) {
+		fprintf(stderr, "invalid allocation size: %s\n", size_arg);
+		return 1;
+	}
+	if (join_cgroup(procs))
+		return 1;
+
+	memory = mmap(NULL, size, PROT_READ | PROT_WRITE,
+		      MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+	if (memory == MAP_FAILED) {
+		perror("mmap");
+		return 1;
+	}
+
+	page_size = getpagesize();
+	for (unsigned long i = 0; i < size; i += page_size)
+		memset(memory + i, i / page_size % 251 + 1,
+		       page_size < size - i ? page_size : size - i);
+	if (mprotect(memory, size, PROT_READ)) {
+		perror("mprotect");
+		return 1;
+	}
+
+	sigemptyset(&signals);
+	sigaddset(&signals, SIGUSR1);
+	if (sigprocmask(SIG_BLOCK, &signals, NULL)) {
+		perror("sigprocmask");
+		return 1;
+	}
+
+	fd = open(ready, O_WRONLY | O_CREAT | O_EXCL, 0600);
+	if (fd < 0) {
+		perror("create ready file");
+		return 1;
+	}
+	close(fd);
+
+	errno = sigwait(&signals, &signal);
+	if (errno) {
+		perror("sigwait");
+		return 1;
+	}
+	for (unsigned long i = 0; i < size; i++) {
+		unsigned char expected = i / page_size % 251 + 1;
+
+		if (memory[i] != expected) {
+			fprintf(stderr,
+				"data mismatch at %lu: got %u, expected %u\n",
+				i, memory[i], expected);
+			return 1;
+		}
+	}
+
+	fd = open(verified, O_WRONLY | O_CREAT | O_EXCL, 0600);
+	if (fd < 0) {
+		perror("create verified file");
+		return 1;
+	}
+	close(fd);
+
+	for (;;)
+		pause();
+}
+
+static int create_marker(const char *path, int write_pid)
+{
+	char *temporary;
+	int fd, ret = 1;
+
+	if (asprintf(&temporary, "%s.XXXXXX", path) < 0) {
+		perror("asprintf marker path");
+		return 1;
+	}
+	fd = mkstemp(temporary);
+	if (fd < 0) {
+		perror(path);
+		goto out_free;
+	}
+	if (write_pid && dprintf(fd, "%d\n", getpid()) < 0) {
+		perror("write marker");
+		close(fd);
+		goto out_unlink;
+	}
+	if (close(fd)) {
+		perror("close marker");
+		goto out_unlink;
+	}
+	/* Publish only after the payload is complete for the shell reader. */
+	if (rename(temporary, path)) {
+		perror("publish marker");
+		goto out_unlink;
+	}
+	ret = 0;
+out_unlink:
+	unlink(temporary);
+out_free:
+	free(temporary);
+	return ret;
+}
+
+static unsigned char retained_byte(unsigned long offset, long page_size,
+				   int touched)
+{
+	unsigned char value = offset / page_size % 251 + 1;
+
+	if (touched && !offset)
+		value ^= 0x5a;
+	return value;
+}
+
+static int verify_retained(unsigned char *memory, unsigned long size,
+			   long page_size, int full)
+{
+	unsigned long limit = full ? size : 1;
+
+	for (unsigned long offset = 0; offset < limit; offset++) {
+		unsigned char expected = retained_byte(offset, page_size, 1);
+
+		if (memory[offset] != expected) {
+			fprintf(stderr,
+				"retained data mismatch at %lu: got %u, expected %u\n",
+				offset, memory[offset], expected);
+			return 1;
+		}
+	}
+	return 0;
+}
+
+static int retained(const char *size_arg, const char *procs,
+		    const char *ready, const char *touched,
+		    const char *verified)
+{
+	unsigned long allocation_size, mapping_start, size;
+	unsigned char *mapping, *memory;
+	char *end;
+	sigset_t signals;
+	long page_size;
+	int signal;
+
+	errno = 0;
+	size = strtoul(size_arg, &end, 0);
+	if (errno || *end || !size || (size & (size - 1))) {
+		fprintf(stderr, "allocation size must be a power of two\n");
+		return 1;
+	}
+	if (join_cgroup(procs))
+		return 1;
+
+	page_size = sysconf(_SC_PAGESIZE);
+	if (page_size <= 0) {
+		perror("sysconf _SC_PAGESIZE");
+		return 1;
+	}
+	if (size < (unsigned long)page_size || size % page_size) {
+		fprintf(stderr, "allocation size must contain whole pages\n");
+		return 1;
+	}
+	if (size > ULONG_MAX / 2) {
+		fprintf(stderr, "allocation size is too large\n");
+		return 1;
+	}
+	allocation_size = size * 2;
+	mapping = mmap(NULL, allocation_size, PROT_READ | PROT_WRITE,
+		       MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+	if (mapping == MAP_FAILED) {
+		perror("mmap");
+		return 1;
+	}
+	mapping_start = (unsigned long)mapping;
+	memory = (unsigned char *)((mapping_start + size - 1) & ~(size - 1));
+	if ((unsigned long)memory != mapping_start)
+		munmap((void *)mapping_start,
+		       (unsigned long)memory - mapping_start);
+	munmap(memory + size, allocation_size - size -
+	       ((unsigned long)memory - mapping_start));
+	if (madvise(memory, size, MADV_HUGEPAGE)) {
+		perror("madvise MADV_HUGEPAGE");
+		return 1;
+	}
+	for (unsigned long offset = 0; offset < size; offset += page_size)
+		memset(memory + offset, retained_byte(offset, page_size, 0),
+		       page_size);
+#ifdef MADV_COLLAPSE
+	if (madvise(memory, size, MADV_COLLAPSE)) {
+		int error = errno;
+
+		perror("madvise MADV_COLLAPSE");
+		if (error == EAGAIN || error == EINVAL || error == ENOMEM)
+			return KSFT_SKIP;
+		return 1;
+	}
+#else
+	fprintf(stderr, "MADV_COLLAPSE is unavailable\n");
+	return KSFT_SKIP;
+#endif
+
+	sigemptyset(&signals);
+	sigaddset(&signals, SIGUSR1);
+	sigaddset(&signals, SIGUSR2);
+	sigaddset(&signals, SIGALRM);
+	if (sigprocmask(SIG_BLOCK, &signals, NULL)) {
+		perror("sigprocmask");
+		return 1;
+	}
+	if (create_marker(ready, 1))
+		return 1;
+
+	for (;;) {
+		errno = sigwait(&signals, &signal);
+		if (errno) {
+			perror("sigwait");
+			return 1;
+		}
+		if (signal == SIGUSR1) {
+			unsigned char value;
+
+			if (mprotect(memory, page_size, PROT_READ)) {
+				perror("mprotect read");
+				return 1;
+			}
+			value = memory[0];
+			if (mprotect(memory, page_size, PROT_READ | PROT_WRITE)) {
+				perror("mprotect write");
+				return 1;
+			}
+			memory[0] = value ^ 0x5a;
+			if (create_marker(touched, 0))
+				return 1;
+		} else {
+			if (verify_retained(memory, size, page_size,
+					    signal == SIGALRM))
+				return 1;
+			if (create_marker(verified, 0))
+				return 1;
+		}
+	}
+}
+
+int main(int argc, char **argv)
+{
+	if (argc == 4 && !strcmp(argv[1], "activate"))
+		return activate(argv[2], atoi(argv[3]),
+				SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE);
+	if (argc == 4 && !strcmp(argv[1], "activate-no-discard"))
+		return activate(argv[2], atoi(argv[3]), 0);
+	if (argc == 4 && !strcmp(argv[1], "reject-page-discard"))
+		return reject_page_discard(argv[2], atoi(argv[3]));
+	if (argc == 4 && !strcmp(argv[1], "accept-discard-once-pages"))
+		return accept_discard_once_pages(argv[2], atoi(argv[3]));
+	if (argc == 3 && !strcmp(argv[1], "pin"))
+		return pin_to_one_cpu(argv[2]);
+	if (argc == 6 && !strcmp(argv[1], "allocate"))
+		return allocate(argv[2], argv[3], argv[4], argv[5]);
+	if (argc == 7 && !strcmp(argv[1], "retained"))
+		return retained(argv[2], argv[3], argv[4], argv[5], argv[6]);
+
+	fprintf(stderr,
+		"usage: %s activate|activate-no-discard DEVICE PRIORITY | reject-page-discard DEVICE PRIORITY | accept-discard-once-pages DEVICE PRIORITY | pin PID | allocate BYTES CGROUP.PROCS READY VERIFIED | retained BYTES CGROUP.PROCS READY TOUCHED VERIFIED\n",
+		argv[0]);
+	return 1;
+}
diff --git a/tools/testing/selftests/zram/workingset_offload.c b/tools/testing/selftests/zram/workingset_offload.c
new file mode 100644
index 0000000000000..6d41badb8d0be
--- /dev/null
+++ b/tools/testing/selftests/zram/workingset_offload.c
@@ -0,0 +1,206 @@
+// SPDX-License-Identifier: GPL-2.0
+#define _GNU_SOURCE
+
+#include <errno.h>
+#include <fcntl.h>
+#include <signal.h>
+#include <stdbool.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <unistd.h>
+
+#define ANON_SIZE (64UL << 20)
+#define TARGET_SIZE (16UL << 20)
+#define FILLER_SIZE (32UL << 20)
+
+static int join_cgroup(const char *path)
+{
+	char pid[32];
+	int fd, len;
+
+	fd = open(path, O_WRONLY);
+	if (fd < 0)
+		return -1;
+	len = snprintf(pid, sizeof(pid), "%d\n", getpid());
+	if (write(fd, pid, len) != len) {
+		close(fd);
+		return -1;
+	}
+	return close(fd);
+}
+
+static int create_file(const char *path, size_t size)
+{
+	unsigned char page[4096];
+	size_t offset;
+	int fd;
+
+	fd = open(path, O_CREAT | O_TRUNC | O_RDWR, 0600);
+	if (fd < 0)
+		return -1;
+	memset(page, 0xa5, sizeof(page));
+	for (offset = 0; offset < size; offset += sizeof(page)) {
+		if (pwrite(fd, page, sizeof(page), offset) != sizeof(page)) {
+			close(fd);
+			return -1;
+		}
+	}
+	if (fsync(fd) || posix_fadvise(fd, 0, size, POSIX_FADV_DONTNEED)) {
+		close(fd);
+		return -1;
+	}
+	return fd;
+}
+
+static int cache_file(int fd, size_t size)
+{
+	unsigned char byte;
+	size_t offset;
+
+	if (posix_fadvise(fd, 0, size, POSIX_FADV_RANDOM))
+		return -1;
+	for (offset = 0; offset < size; offset += getpagesize())
+		if (pread(fd, &byte, 1, offset) != 1)
+			return -1;
+	return 0;
+}
+
+static unsigned long memory_stat(const char *path, const char *key)
+{
+	unsigned long value;
+	char name[64];
+	FILE *file;
+
+	file = fopen(path, "r");
+	if (!file)
+		return 0;
+	while (fscanf(file, "%63s %lu", name, &value) == 2) {
+		if (!strcmp(name, key)) {
+			fclose(file);
+			return value;
+		}
+	}
+	fclose(file);
+	return 0;
+}
+
+static int touch_ready(const char *path)
+{
+	int fd = open(path, O_WRONLY | O_CREAT | O_EXCL, 0600);
+
+	if (fd < 0)
+		return -1;
+	return close(fd);
+}
+
+static bool calibration_sufficient(size_t pages, unsigned long selected,
+				   unsigned long refaulted,
+				   unsigned long anon)
+{
+	return selected >= pages / 2 && refaulted >= selected * 3 / 4 &&
+	       anon >= TARGET_SIZE;
+}
+
+int main(int argc, char **argv)
+{
+	unsigned long refault_before, refault_after;
+	unsigned long activate_before, activate_after;
+	unsigned long anon, selected = 0;
+	unsigned char *resident, *anon_memory;
+	unsigned char byte, checksum = 0;
+	char stat_path[4096];
+	sigset_t signals;
+	size_t pages, i;
+	void *mapping;
+	int target_fd, filler_fd, signal;
+
+	if (argc != 6) {
+		fprintf(stderr, "usage: %s CGROUP.PROCS TARGET FILLER READY GO\n",
+			argv[0]);
+		return 1;
+	}
+	if (join_cgroup(argv[1])) {
+		perror("join cgroup");
+		return 1;
+	}
+
+	anon_memory = mmap(NULL, ANON_SIZE, PROT_READ | PROT_WRITE,
+			   MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+	if (anon_memory == MAP_FAILED) {
+		perror("mmap anonymous");
+		return 1;
+	}
+	for (i = 0; i < ANON_SIZE; i += getpagesize())
+		anon_memory[i] = i / getpagesize() % 251 + 1;
+
+	target_fd = create_file(argv[2], TARGET_SIZE);
+	filler_fd = create_file(argv[3], FILLER_SIZE);
+	if (target_fd < 0 || filler_fd < 0) {
+		perror("create files");
+		return 1;
+	}
+	if (cache_file(target_fd, TARGET_SIZE) ||
+	    cache_file(filler_fd, FILLER_SIZE)) {
+		perror("populate page cache");
+		return 1;
+	}
+
+	sigemptyset(&signals);
+	sigaddset(&signals, SIGUSR1);
+	if (sigprocmask(SIG_BLOCK, &signals, NULL) || touch_ready(argv[4])) {
+		perror("prepare signal");
+		return 1;
+	}
+	if (sigwait(&signals, &signal)) {
+		perror("sigwait");
+		return 1;
+	}
+
+	mapping = mmap(NULL, TARGET_SIZE, PROT_READ, MAP_SHARED, target_fd, 0);
+	if (mapping == MAP_FAILED) {
+		perror("mmap target");
+		return 1;
+	}
+	pages = TARGET_SIZE / getpagesize();
+	resident = calloc(pages, 1);
+	if (!resident || mincore(mapping, TARGET_SIZE, resident)) {
+		perror("mincore");
+		return 1;
+	}
+	munmap(mapping, TARGET_SIZE);
+
+	snprintf(stat_path, sizeof(stat_path), "%.*s/memory.stat",
+		 (int)(strlen(argv[1]) - strlen("/cgroup.procs")), argv[1]);
+	refault_before = memory_stat(stat_path, "workingset_refault_file");
+	activate_before = memory_stat(stat_path, "workingset_activate_file");
+	for (i = 0; i < pages; i++) {
+		if (resident[i] & 1)
+			continue;
+		if (pread(target_fd, &byte, 1, i * getpagesize()) != 1) {
+			perror("refault target");
+			return 1;
+		}
+		selected++;
+	}
+	refault_after = memory_stat(stat_path, "workingset_refault_file");
+	activate_after = memory_stat(stat_path, "workingset_activate_file");
+	anon = memory_stat(stat_path, "active_anon") +
+	       memory_stat(stat_path, "inactive_anon");
+
+	for (i = 0; i < ANON_SIZE; i += getpagesize())
+		checksum ^= anon_memory[i];
+	printf("selected=%lu refault=%lu activate=%lu anon=%lu checksum=%u\n",
+	       selected, refault_after - refault_before,
+	       activate_after - activate_before, anon, checksum);
+
+	free(resident);
+	close(filler_fd);
+	close(target_fd);
+	if (!calibration_sufficient(pages, selected,
+				    refault_after - refault_before, anon))
+		return 2;
+	return (activate_after - activate_before) * 100 < selected * 80;
+}
diff --git a/tools/testing/selftests/zram/zram03.sh b/tools/testing/selftests/zram/zram03.sh
new file mode 100755
index 0000000000000..3ccc9a135f997
--- /dev/null
+++ b/tools/testing/selftests/zram/zram03.sh
@@ -0,0 +1,166 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Test proactive-only swap allocation and pressure fallback.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram03"
+cg="/sys/fs/cgroup/zram-offload-$$"
+cgroup_root="/sys/fs/cgroup"
+ready="/tmp/zram-offload-ready-$$"
+verified="/tmp/zram-offload-verified-$$"
+allocator_pid=""
+offload=""
+zswap_enabled=""
+zswap_writeback=""
+
+fail()
+{
+	echo "$TCID: [FAIL] $*" >&2
+	exit 1
+}
+
+skip()
+{
+	echo "$TCID: [SKIP] $*" >&2
+	exit "$ksft_skip"
+}
+
+cleanup()
+{
+	set +e
+	if [ -n "$allocator_pid" ]; then
+		kill "$allocator_pid"
+		wait "$allocator_pid"
+	fi
+	rm -f "$ready" "$verified"
+	rmdir "$cg"
+	cgroup_disable_memory_controller "$cgroup_root"
+	if [ -n "$offload" ] && [ "$dev_makeswap" -lt "$dev_end" ]; then
+		swapoff "$offload" >/dev/null 2>&1
+	fi
+	if [ "$dev_end" -ge "$dev_start" ]; then
+		zram_cleanup
+	fi
+	if [ -n "$zswap_enabled" ]; then
+		echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled
+	fi
+}
+
+swap_used_kb()
+{
+	awk -v device="$1" '$1 == device { print $4 }' /proc/swaps
+}
+
+zram_orig_data_size()
+{
+	awk '{ print $1 }' "/sys/block/${1##*/}/mm_stat"
+}
+
+wait_file()
+{
+	for _ in $(seq 1 400); do
+		[ -e "$1" ] && return 0
+		sleep 0.05
+	done
+	return 1
+}
+
+check_prereqs
+# The feature marker also requires CONFIG_VM_EVENT_COUNTERS.
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+	skip "offload refusal counters are unavailable"
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] || skip "cgroup v2 controllers are unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers || skip "memory controller is unavailable"
+[ -e /sys/module/zswap/parameters/enabled ] || skip "zswap is unavailable"
+zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+
+# Keep every reclaim transition on one CPU.  This makes the test exercise the
+# cached-cluster provenance mismatch rather than relying on scheduler placement.
+./swap_offload pin "$$" || skip "cannot pin test to one allowed CPU"
+
+# Swap priorities are global. Put both test areas ahead of any existing swap
+# so distro-managed zram or another high-priority area cannot absorb the test
+# allocations and make the routing assertions fail spuriously.
+max_prio=$(awk 'BEGIN { max = -1 } NR > 1 && $5 > max { max = $5 } END { print max }' /proc/swaps)
+[ "$max_prio" -le 32765 ] || skip "cannot outrank existing swap priority $max_prio"
+safe_prio=$((max_prio + 1))
+offload_prio=$((max_prio + 2))
+
+dev_num=2
+zram_sizes="67108864 67108864"
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+zram_load
+echo Y > /sys/module/zswap/parameters/enabled || skip "cannot enable zswap"
+zram_set_disksizes
+
+safe="/dev/zram${dev_start}"
+offload="/dev/zram$((dev_start + 1))"
+mkswap "$safe" >/dev/null
+mkswap "$offload" >/dev/null
+swapon -p "$safe_prio" "$safe"
+dev_makeswap=$dev_start
+./swap_offload reject-page-discard "$offload" "$offload_prio" ||
+	fail "offload-only page discard was accepted"
+./swap_offload accept-discard-once-pages "$offload" "$offload_prio" ||
+	fail "resolved offload-only discard-once was rejected"
+./swap_offload activate "$offload" "$offload_prio"
+dev_makeswap=$dev_end
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+	skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+[ -e "$cg/memory.max" ] || skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.swap.max"
+echo 1 > "$cg/memory.zswap.writeback" ||
+	skip "cannot enable zswap writeback for test cgroup"
+read -r zswap_writeback < "$cg/memory.zswap.writeback"
+[ "$zswap_writeback" -eq 1 ] ||
+	skip "an ancestor disables zswap writeback"
+
+./swap_offload allocate 67108864 "$cg/cgroup.procs" "$ready" "$verified" &
+allocator_pid=$!
+wait_file "$ready" || fail "allocator did not become ready"
+
+echo "16M swappiness=max" > "$cg/memory.reclaim" ||
+	fail "proactive reclaim failed"
+offload_before=$(swap_used_kb "$offload")
+safe_before=$(swap_used_kb "$safe")
+offload_orig=$(zram_orig_data_size "$offload")
+[ "${offload_before:-0}" -gt 0 ] || fail "proactive reclaim missed offload area"
+[ "${safe_before:-0}" -eq 0 ] || fail "proactive reclaim used lower-priority safe area"
+[ "$((offload_orig + 1048576))" -ge "$((offload_before * 1024))" ] ||
+	fail "zswap deferred the offload-only backend write"
+
+echo 24M > "$cg/memory.max"
+sleep 1
+offload_pressure=$(swap_used_kb "$offload")
+safe_pressure=$(swap_used_kb "$safe")
+safe_orig=$(zram_orig_data_size "$safe")
+[ "${safe_pressure:-0}" -gt 0 ] || fail "pressure reclaim missed safe fallback"
+[ "$offload_pressure" -le "$offload_before" ] ||
+	fail "pressure reclaim allocated offload-only slots"
+[ "$safe_orig" -lt "$((safe_pressure * 1024))" ] ||
+	fail "zswap did not retain any unmarked safe-area pages"
+
+echo max > "$cg/memory.max"
+echo "4M swappiness=max" > "$cg/memory.reclaim" ||
+	fail "second proactive reclaim failed"
+offload_after=$(swap_used_kb "$offload")
+[ "$offload_after" -gt "$offload_pressure" ] ||
+	fail "proactive priority was not restored after pressure reclaim"
+
+swapoff "$offload" || fail "swapoff could not recover offloaded pages"
+kill -USR1 "$allocator_pid" || fail "cannot request data verification"
+wait_file "$verified" || fail "allocator did not verify recovered data"
+kill -0 "$allocator_pid" || fail "allocator died during swapoff"
+dev_makeswap=$dev_start
+
+echo "$TCID: [PASS]"
diff --git a/tools/testing/selftests/zram/zram04.sh b/tools/testing/selftests/zram/zram04.sh
new file mode 100755
index 0000000000000..a8f52a82e774b
--- /dev/null
+++ b/tools/testing/selftests/zram/zram04.sh
@@ -0,0 +1,149 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Reproduce offload-only swap leaking into the workingset refault heuristic.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram04"
+cg="/sys/fs/cgroup/zram-workingset-$$"
+cgroup_root="/sys/fs/cgroup"
+tmp="${TMPDIR:-/var/tmp}/zram-workingset-$$"
+ready="$tmp/ready"
+worker=""
+tmp_fs=""
+mglru=""
+zswap_enabled=""
+
+fail()
+{
+	echo "$TCID: [FAIL] $*" >&2
+	exit 1
+}
+
+skip()
+{
+	echo "$TCID: [SKIP] $*" >&2
+	exit "$ksft_skip"
+}
+
+cleanup()
+{
+	set +e
+	[ -n "$worker" ] && kill "$worker"
+	[ -n "$worker" ] && wait "$worker"
+	rm -rf "$tmp"
+	rmdir "$cg"
+	cgroup_disable_memory_controller "$cgroup_root"
+	[ "$dev_end" -ge "$dev_start" ] && zram_cleanup
+	[ -n "$mglru" ] && echo "$mglru" > /sys/kernel/mm/lru_gen/enabled
+	[ -n "$zswap_enabled" ] &&
+		echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled
+}
+
+wait_helper_ready()
+{
+	for _ in $(seq 1 400); do
+		[ -e "$ready" ] && return 0
+		worker_state=$(awk '{ print $3 }' "/proc/$worker/stat" \
+			2>/dev/null || :)
+		if ! kill -0 "$worker" 2>/dev/null ||
+		   [ "$worker_state" = Z ]; then
+			if wait "$worker"; then
+				status=0
+			else
+				status=$?
+			fi
+			worker=""
+			[ "$status" -eq 2 ] &&
+				skip "workingset calibration was insufficient"
+			fail "workingset helper exited with status $status before readiness"
+		fi
+		sleep 0.05
+	done
+	fail "workingset helper timed out before readiness"
+}
+
+check_prereqs
+# The feature marker also requires CONFIG_VM_EVENT_COUNTERS.
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+	skip "offload refusal counters are unavailable"
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+[ -x ./workingset_offload ] || skip "workingset helper is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] || skip "cgroup v2 is unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers ||
+	skip "cgroup v2 memory controller is unavailable"
+[ "$(awk 'END { print NR }' /proc/swaps)" -eq 1 ] ||
+	skip "test requires no pre-existing swap"
+tmp_fs=$(stat -f -c %T "${TMPDIR:-/var/tmp}") ||
+	skip "cannot identify the test filesystem"
+[ "$tmp_fs" != tmpfs ] ||
+	skip "test files require a disk-backed filesystem"
+
+[ -e /sys/kernel/mm/lru_gen/enabled ] &&
+	mglru=$(cat /sys/kernel/mm/lru_gen/enabled)
+[ -e /sys/module/zswap/parameters/enabled ] &&
+	zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+[ -n "$mglru" ] && echo 0 > /sys/kernel/mm/lru_gen/enabled
+[ -n "$zswap_enabled" ] && echo N > /sys/module/zswap/parameters/enabled
+
+dev_num=1
+zram_sizes="134217728"
+zram_load
+zram_set_disksizes
+offload="/dev/zram${dev_start}"
+mkswap "$offload" >/dev/null
+./swap_offload activate "$offload" 1
+dev_makeswap=$dev_end
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+	skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+[ -e "$cg/memory.max" ] || skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.max"
+echo max > "$cg/memory.swap.max"
+mkdir "$tmp"
+
+./workingset_offload "$cg/cgroup.procs" "$tmp/target" "$tmp/filler" \
+	"$ready" unused &
+worker=$!
+wait_helper_ready
+
+# Shadow retention uses local LRU/slab statistics updated by memcg flushing.
+# Check that the 64 MiB anonymous and 48 MiB file setup is visible before
+# evicting file pages.
+stats_ready=0
+for _ in $(seq 1 50); do
+	if awk '
+		$1 == "active_anon" || $1 == "inactive_anon" { anon += $2 }
+		$1 == "file" { file = $2 }
+		END { exit !(anon >= 67108864 && file >= 50331648) }
+	' "$cg/memory.stat"; then
+		stats_ready=1
+		break
+	fi
+	sleep 0.1
+done
+[ "$stats_ready" -eq 1 ] ||
+	skip "initial working-set statistics did not become visible"
+
+echo "48M swappiness=0" > "$cg/memory.reclaim" ||
+	skip "file-only proactive reclaim failed"
+kill -USR1 "$worker"
+if wait "$worker"; then
+	worker=""
+	echo "$TCID: [PASS]"
+	exit 0
+else
+	status=$?
+fi
+worker=""
+[ "$status" -eq 2 ] && skip "workingset calibration was insufficient"
+echo "$TCID: [FAIL] refaulted file pages were not activated" >&2
+exit 1
diff --git a/tools/testing/selftests/zram/zram05.sh b/tools/testing/selftests/zram/zram05.sh
new file mode 100755
index 0000000000000..e4bbf6bfa0e92
--- /dev/null
+++ b/tools/testing/selftests/zram/zram05.sh
@@ -0,0 +1,360 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Test retained offload-only entries under ordinary and proactive reclaim.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram05"
+cg="/sys/fs/cgroup/zram-retained-$$"
+cgroup_root="/sys/fs/cgroup"
+tmp="${TMPDIR:-/var/tmp}/zram-retained-$$"
+ready="$tmp/ready"
+touched="$tmp/touched"
+verified="$tmp/verified"
+holder_pid=""
+safe=""
+safe_active=0
+offload_backing=""
+offload=""
+offload_active=0
+dm_name="zram-retained-$$"
+dm_active=0
+dm_node_created=0
+zswap_enabled=""
+thp_size=0
+thp_kib=0
+thp_sectors=0
+page_kib=0
+
+fail()
+{
+	echo "$TCID: [FAIL] $*" >&2
+	exit 1
+}
+
+skip()
+{
+	echo "$TCID: [SKIP] $*" >&2
+	exit "$ksft_skip"
+}
+
+cleanup()
+{
+	set +e
+	if [ -n "$holder_pid" ]; then
+		kill "$holder_pid"
+		wait "$holder_pid"
+	fi
+	if [ "$offload_active" -eq 1 ]; then
+		swapoff "$offload" >/dev/null 2>&1
+	fi
+	if [ "$safe_active" -eq 1 ]; then
+		swapoff "$safe" >/dev/null 2>&1
+	fi
+	if [ "$dm_active" -eq 1 ]; then
+		dmsetup --noudevsync --noudevrules remove "$dm_name" \
+			>/dev/null 2>&1
+	fi
+	if [ "$dm_node_created" -eq 1 ]; then
+		rm -f "$offload"
+	fi
+	rm -rf "$tmp"
+	rmdir "$cg"
+	cgroup_disable_memory_controller "$cgroup_root"
+	if [ "$dev_end" -ge "$dev_start" ]; then
+		zram_cleanup
+	fi
+	if [ -n "$zswap_enabled" ]; then
+		echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled
+	fi
+}
+
+wait_file()
+{
+	for _ in $(seq 1 400); do
+		[ -e "$1" ] && return 0
+		if [ -n "$holder_pid" ]; then
+			holder_state=$(awk '{ print $3 }' \
+				"/proc/$holder_pid/stat" 2>/dev/null || :)
+		else
+			holder_state=""
+		fi
+		if [ -n "$holder_pid" ] &&
+		   { ! kill -0 "$holder_pid" 2>/dev/null ||
+		     [ "$holder_state" = Z ]; }; then
+			if wait "$holder_pid"; then
+				helper_status=0
+			else
+				helper_status=$?
+			fi
+			holder_pid=""
+			return 2
+		fi
+		sleep 0.05
+	done
+	return 1
+}
+
+require_helper_file()
+{
+	if wait_file "$1"; then
+		return 0
+	else
+		status=$?
+	fi
+	if [ "$status" -eq 2 ]; then
+		[ "$helper_status" -eq "$ksft_skip" ] && skip "$2 is unavailable"
+		fail "retained helper exited with status $helper_status while $2"
+	fi
+	fail "retained helper timed out while $2"
+}
+
+written_sectors()
+{
+	awk '{ print $7 }' "/sys/block/${offload_backing##*/}/stat"
+}
+
+vmstat_value()
+{
+	awk -v name="$1" '$1 == name { print $2 }' /proc/vmstat
+}
+
+wait_offload_quiet()
+{
+	# A bio queued inside dm-delay is not yet visible in either the backing
+	# device statistics or the mapped device's inflight counters.
+	sleep 4
+	previous=-1
+	stable=0
+	for _ in $(seq 1 240); do
+		current=$(written_sectors)
+		read -r reads writes < "$dm_inflight"
+		if [ "$current" -eq "$previous" ] && \
+		   [ "$reads" -eq 0 ] && [ "$writes" -eq 0 ]; then
+			stable=$((stable + 1))
+			[ "$stable" -ge 20 ] && return 0
+		else
+			stable=0
+		fi
+		previous=$current
+		sleep 0.05
+	done
+	return 1
+}
+
+check_prereqs
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+command -v dmsetup >/dev/null 2>&1 || skip "dmsetup is unavailable"
+command -v blockdev >/dev/null 2>&1 || skip "blockdev is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] ||
+	skip "cgroup v2 controllers are unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers ||
+	skip "memory controller is unavailable"
+[ -d /sys/kernel/mm/transparent_hugepage ] ||
+	skip "transparent huge pages are unavailable"
+[ -r /sys/kernel/mm/transparent_hugepage/hpage_pmd_size ] ||
+	skip "PMD huge-page size is unavailable"
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+	skip "offload refusal counters are unavailable"
+
+thp_size=$(cat /sys/kernel/mm/transparent_hugepage/hpage_pmd_size)
+case "$thp_size" in
+	''|*[!0-9]*) skip "invalid PMD huge-page size: $thp_size" ;;
+esac
+[ "$thp_size" -gt 0 ] || skip "PMD huge-page size is zero"
+page_kib=$(awk '/KernelPageSize:/ { print $2; exit }' /proc/self/smaps)
+[ "${page_kib:-0}" -gt 0 ] || skip "cannot determine the base page size"
+page_size=$((page_kib * 1024))
+[ "$page_size" -gt 0 ] || skip "base page size is zero"
+[ $((thp_size % page_size)) -eq 0 ] ||
+	skip "PMD huge-page size is not page aligned"
+thp_kib=$((thp_size / 1024))
+thp_sectors=$((thp_size / 512))
+expected_retained_kib=$((thp_kib - page_kib))
+expected_refused=$((thp_size / page_size))
+
+mkdir "$tmp"
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+dmsetup targets > "$tmp/dm-targets" 2>/dev/null ||
+	skip "cannot query device-mapper targets"
+grep -q '^delay[[:space:]]' "$tmp/dm-targets" ||
+	skip "device-mapper delay target is unavailable"
+
+if [ -e /sys/module/zswap/parameters/enabled ]; then
+	zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+	echo N > /sys/module/zswap/parameters/enabled ||
+		skip "cannot disable zswap"
+fi
+
+# Swap priorities are global.  The ordinary zram device must be preferred to
+# any pre-existing swap, while the delayed offload device remains first for
+# eligible proactive reclaim.
+max_prio=$(awk 'BEGIN { max = -1 } NR > 1 && $5 > max { max = $5 } END { print max }' /proc/swaps)
+[ "$max_prio" -le 32765 ] ||
+	skip "cannot outrank existing swap priority $max_prio"
+safe_prio=$((max_prio + 1))
+offload_prio=$((max_prio + 2))
+
+dev_num=2
+zram_size=$((thp_size * 4))
+[ "$zram_size" -ge 67108864 ] || zram_size=67108864
+zram_sizes="$zram_size $zram_size"
+zram_load
+zram_set_disksizes
+safe="/dev/zram${dev_start}"
+offload_backing="/dev/zram$((dev_start + 1))"
+
+sectors=$(blockdev --getsz "$offload_backing") ||
+	skip "cannot read offload backing size"
+[ "$sectors" -gt 0 ] || skip "offload backing has zero size"
+dm_table="0 $sectors delay $offload_backing 0 0 $offload_backing 0 3000"
+dmsetup --noudevsync --noudevrules create "$dm_name" --table "$dm_table" ||
+	skip "cannot create delayed offload device"
+dm_active=1
+dmsetup --noudevsync --noudevrules info --columns --noheadings \
+	--separator ' ' -o major,minor "$dm_name" > "$tmp/dm-devno" ||
+	skip "cannot identify delayed offload device"
+read -r dm_major dm_minor < "$tmp/dm-devno"
+offload="/dev/mapper/$dm_name"
+if [ ! -e "$offload" ]; then
+	mkdir -p /dev/mapper
+	mknod "$offload" b "$dm_major" "$dm_minor" ||
+		skip "cannot create delayed offload device node"
+	dm_node_created=1
+fi
+dm_inflight="/sys/dev/block/$dm_major:$dm_minor/inflight"
+[ -r "$dm_inflight" ] || skip "cannot observe delayed offload I/O"
+
+mkswap "$safe" >/dev/null || fail "cannot initialise safe swap"
+mkswap "$offload" >/dev/null || fail "cannot initialise offload swap"
+swapon -p "$safe_prio" "$safe" || fail "cannot activate safe swap"
+dev_makeswap=$dev_start
+safe_active=1
+./swap_offload activate-no-discard "$offload" "$offload_prio" ||
+	fail "cannot activate offload-only swap"
+offload_active=1
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+	skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+[ -e "$cg/memory.max" ] ||
+	skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.swap.max"
+
+./swap_offload retained "$thp_size" "$cg/cgroup.procs" "$ready" \
+	"$touched" "$verified" &
+holder_pid=$!
+require_helper_file "$ready" "creating a PMD-sized anonymous huge folio"
+read -r reported_pid < "$ready"
+[ "$reported_pid" -eq "$holder_pid" ] ||
+	fail "retained helper reported the wrong pid"
+anon_huge_kib=$(awk '/AnonHugePages:/ { print $2 }' \
+	"/proc/$holder_pid/smaps_rollup")
+[ "${anon_huge_kib:-0}" -ge "$thp_kib" ] ||
+	skip "anonymous huge folio was not created"
+echo "$TCID: anonymous_huge_kib=$anon_huge_kib"
+
+# This is the only reclaim before the retained entry is made dirty.  It puts
+# the whole folio on offload-only swap so the later single-page fault leaves
+# sibling swap PTEs referring to the existing slot.
+echo "$thp_size swappiness=max" > "$cg/memory.reclaim" 2>/dev/null || :
+wait_offload_quiet || fail "offload device did not quiesce after setup"
+kill -0 "$holder_pid" || fail "retained helper died during setup reclaim"
+kill -USR1 "$holder_pid" || fail "cannot request the dirty-page transition"
+require_helper_file "$touched" "faulting and dirtying the retained entry"
+retained_kib=$(awk '/VmSwap:/ { print $2 }' "/proc/$holder_pid/status")
+[ "${retained_kib:-0}" -eq "$expected_retained_kib" ] ||
+	fail "retained $retained_kib KiB, expected $expected_retained_kib KiB"
+echo "$TCID: retained_kib=$retained_kib"
+
+wait_offload_quiet || fail "offload device did not quiesce before pressure"
+ordinary_before=$(written_sectors)
+refused_before=$(vmstat_value swpout_offload_refused)
+echo $((thp_size / 2)) > "$cg/memory.high" ||
+	fail "ordinary pressure reclaim failed"
+wait_offload_quiet || fail "offload device did not quiesce after pressure"
+ordinary_after=$(written_sectors)
+refused_after=$(vmstat_value swpout_offload_refused)
+ordinary_writes=$((ordinary_after - ordinary_before))
+refused_pages=$((refused_after - refused_before))
+echo "$TCID: ordinary_sectors=$ordinary_writes refused_pages=$refused_pages"
+[ "$ordinary_writes" -eq 0 ] ||
+	fail "ordinary reclaim wrote $ordinary_writes offload sectors"
+[ "$refused_pages" -ge "$expected_refused" ] ||
+	fail "ordinary reclaim refused $refused_pages pages, expected at least $expected_refused"
+kill -0 "$holder_pid" || fail "retained helper died after refused write"
+rm -f "$verified"
+kill -USR2 "$holder_pid" || fail "cannot request dirty-byte verification"
+require_helper_file "$verified" "verifying the dirty retained byte"
+
+# Proactive reclaim must still be able to rewrite the retained PMD-sized slot.
+# Repeated requests make the test insensitive to a short-lived writeback
+# collision; the backing-sector delta still requires exactly one folio write.
+wait_offload_quiet || fail "offload device did not quiesce before recovery"
+recovery_before=$(written_sectors)
+echo max > "$cg/memory.high"
+for _ in $(seq 1 8); do
+	echo "$thp_size swappiness=max" > "$cg/memory.reclaim" 2>/dev/null || :
+done
+wait_offload_quiet || fail "offload device did not quiesce after recovery"
+recovery_after=$(written_sectors)
+recovery_writes=$((recovery_after - recovery_before))
+echo "$TCID: recovery_sectors=$recovery_writes"
+[ "$recovery_writes" -eq "$thp_sectors" ] ||
+	fail "proactive recovery wrote $recovery_writes sectors, expected $thp_sectors"
+kill -0 "$holder_pid" || fail "retained helper died during recovery"
+rm -f "$verified"
+kill -ALRM "$holder_pid" || fail "cannot request full data verification"
+require_helper_file "$verified" "verifying all retained data"
+kill -0 "$holder_pid" || fail "retained helper failed full data verification"
+
+kill "$holder_pid" || fail "cannot stop retained helper"
+if wait "$holder_pid"; then
+	:
+else
+	status=$?
+	[ "$status" -eq 143 ] || fail "retained helper exited with status $status"
+fi
+holder_pid=""
+swapoff "$offload" || fail "cannot deactivate offload-only swap"
+offload_active=0
+swapoff "$safe" || fail "cannot deactivate safe swap"
+safe_active=0
+dev_makeswap=-1
+dmsetup --noudevsync --noudevrules remove "$dm_name" ||
+	fail "cannot remove delayed offload device"
+dm_active=0
+if [ "$dm_node_created" -eq 1 ]; then
+	rm -f "$offload" || fail "cannot remove delayed offload device node"
+	dm_node_created=0
+fi
+rmdir "$cg" || fail "cannot remove test cgroup"
+cg=""
+cgroup_disable_memory_controller "$cgroup_root" ||
+	fail "cannot restore the cgroup memory controller"
+for i in $(seq "$dev_start" "$dev_end"); do
+	echo 1 > "/sys/block/zram${i}/reset" || fail "cannot reset zram$i"
+	if [ "$sys_control" -eq 1 ]; then
+		echo "$i" > /sys/class/zram-control/hot_remove ||
+			fail "cannot remove zram$i"
+	fi
+done
+if [ "$module_load" -eq 1 ]; then
+	rmmod zram || fail "cannot unload zram"
+fi
+dev_end=-1
+if [ -n "$zswap_enabled" ]; then
+	echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled ||
+		fail "cannot restore zswap"
+	zswap_enabled=""
+fi
+rm -rf "$tmp" || fail "cannot remove temporary files"
+tmp=""
+
+echo "$TCID: [PASS]"
diff --git a/tools/testing/selftests/zram/zram_lib.sh b/tools/testing/selftests/zram/zram_lib.sh
index 0d44d83888f9d..f14966d3820d8 100755
--- a/tools/testing/selftests/zram/zram_lib.sh
+++ b/tools/testing/selftests/zram/zram_lib.sh
@@ -17,6 +17,10 @@ kernel_version=`uname -r | cut -d'.' -f1,2`
 kernel_major=${kernel_version%.*}
 kernel_minor=${kernel_version#*.}
 
+# Whether this test enabled the memory controller on its cgroup parent.  Tests
+# must leave a delegation which was already present alone.
+cgroup_memory_controller_enabled=0
+
 trap INT
 
 check_prereqs()
@@ -30,6 +34,30 @@ check_prereqs()
 	fi
 }
 
+cgroup_enable_memory_controller()
+{
+	local cgroup_root=$1
+
+	if grep -qw memory "$cgroup_root/cgroup.subtree_control"; then
+		return 0
+	fi
+
+	if ! echo +memory > "$cgroup_root/cgroup.subtree_control"; then
+		return 1
+	fi
+
+	cgroup_memory_controller_enabled=1
+}
+
+cgroup_disable_memory_controller()
+{
+	local cgroup_root=$1
+
+	[ "$cgroup_memory_controller_enabled" -eq 1 ] || return 0
+	echo -memory > "$cgroup_root/cgroup.subtree_control" || return 1
+	cgroup_memory_controller_enabled=0
+}
+
 kernel_gte()
 {
 	major=${1%.*}

51/3 2026/09/14 12:43 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 30277
							output: 31
							thoughts: 1691
						

Thoughts:
**KMSAN Assessment: Offload-Only Swap Areas Patch**

Alright, let's break this down. The core of this patch introduces a new concept: "offload-only" swap areas, marked by `SWAP_FLAG_OFFLOAD_ONLY` or `SWP_OFFLOAD_ONLY`. The idea is to make certain swap areas *exclusively* for proactive reclaim, like the ones triggered by `memory.reclaim` or sysfs, effectively excluding them from the emergency reclaim path. My job is to determine whether this warrants a dedicated KMSAN session.

First, I need to remind myself of the rules. KMSAN is expensive, so I only invoke it if the patch introduces risks of uninitialized memory *that standard detectors (KASAN, LOCKDEP, etc.) would miss*. KMSAN is on the lookout for reads of uninitialized data, especially in kernel-to-user space interactions, conditional logic, and complex data structure initialization. We avoid it if the risk is related to OOB access, null pointer dereferences, locking issues, or use-after-free conditions. Those are already well-covered.

Now, let's analyze the patch changes.

*   **Documentation and Driver Updates:** These are unlikely to introduce uninitialized memory issues directly, but I'll glance at them briefly to make sure.

*   **Header Files:** The new `SWAP_FLAG_OFFLOAD_ONLY` flag and related structures (`struct reclaim_state`, `SWP_OFFLOAD_ONLY` enum) need careful scrutiny. The key concern is whether these new flags are handled in a way that could lead to uninitialized fields being read.

*   **`page_io.c`:** The core logic now checks the `SWP_OFFLOAD_ONLY` flag in the `swap_writeout` function. If the flag is set and `current_reclaim_allows_offload_swap()` is false, it marks the folio dirty, otherwise, zswap may be used. I will confirm that all relevant variables are initialized prior to the conditional checks.

*   **`swapfile.c`:** This is the heart of the changes. The addition of `atomic_long_t nr_swap_pages_offload_only` and the `percpu_swap_cluster` struct are potential areas of concern. Specifically, `percpu_swap_cluster.allow_offload_swap[SWAP_NR_ORDERS]` should be analyzed for how it's initialized (e.g., zeroed, assigned default values, etc.). Calls to `alloc_swap_scan_cluster`, `swap_range_alloc`, `swap_range_free`, and others involving the manipulation of these new variables must be meticulously examined. I will ensure that the logic around `swap_area_eligible(si)` and the handling of the new `nr_swap_pages_offload_only` counter don't expose any uninitialized memory.

*   **`vmscan.c`:** The addition of `scan_control.allow_offload_swap` and `reclaim_state.allow_offload_swap` raises the most significant questions. The `set_task_reclaim_state` function is critical. The initialization of `sc` needs to be investigated in each caller. Since `sc.reclaim_state` is conditionally initialized, the initialization of the `struct scan_control` in all of the callers to `set_task_reclaim_state` is of critical importance. I noticed the use of designated initializers in one caller (`lru_gen_seq_write`), which *should* zero the unset fields, and so that *should* be safe. But the other callers must be verified as well, especially `try_to_free_pages`, `try_to_free_mem_cgroup_pages`, `balance_pgdat`, `shrink_all_memory`, `__node_reclaim`, and `user_proactive_reclaim`. I must grep through these usages to confirm that initialization of `sc.reclaim_state` and any related structures is consistently handled correctly in all cases.

Given the potential for subtle uninitialized memory reads, particularly in the interaction between `vmscan.c`, `swapfile.c`, and the new reclaim state, I'm leaning towards the cautious side. The changes introduce new conditional logic and data structures, and it's imperative to confirm that all code paths are safe. I will need to use a combination of code review and grep to verify the initialization of the added structures and variables. Unless I find explicit evidence that *all* potentially uninitialized memory issues are covered by KASAN or other existing sanitizers, I'll recommend the KMSAN session.



52/3 2026/09/14 12:43 tool grepper
0m Args:
{
  "Expression": "struct scan_control sc",
  "PathPrefix": "mm/vmscan.c"
}

Results:
{
  "Output": "mm/vmscan.c=1675=unsigned int reclaim_clean_pages_from_list(struct zone *zone,\n--\nmm/vmscan.c-1677-{\nmm/vmscan.c:1678:\tstruct scan_control sc = {\nmm/vmscan.c-1679-\t\t.gfp_mask = GFP_KERNEL,\n--\nmm/vmscan.c=2240=static unsigned int reclaim_folio_list(struct list_head *folio_list,\n--\nmm/vmscan.c-2245-\tstruct folio *folio;\nmm/vmscan.c:2246:\tstruct scan_control sc = {\nmm/vmscan.c-2247-\t\t.gfp_mask = GFP_KERNEL,\n--\nmm/vmscan.c=5921=static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,\n--\nmm/vmscan.c-5928-\tint err = -EINVAL;\nmm/vmscan.c:5929:\tstruct scan_control sc = {\nmm/vmscan.c-5930-\t\t.may_writepage = true,\n--\nmm/vmscan.c=6920=unsigned long try_to_free_pages(struct zonelist *zonelist, int order,\n--\nmm/vmscan.c-6923-\tunsigned long nr_reclaimed;\nmm/vmscan.c:6924:\tstruct scan_control sc = {\nmm/vmscan.c-6925-\t\t.nr_to_reclaim = SWAP_CLUSTER_MAX,\n--\nmm/vmscan.c=6965=unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,\n--\nmm/vmscan.c-6972-\tunsigned int noreclaim_flag;\nmm/vmscan.c:6973:\tstruct scan_control sc = {\nmm/vmscan.c-6974-\t\t.nr_to_reclaim = max(nr_pages, SWAP_CLUSTER_MAX),\n--\nmm/vmscan.c=7269=static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx)\n--\nmm/vmscan.c-7276-\tstruct zone *zone;\nmm/vmscan.c:7277:\tstruct scan_control sc = {\nmm/vmscan.c-7278-\t\t.gfp_mask = GFP_KERNEL,\n--\nmm/vmscan.c=7764=unsigned long shrink_all_memory(unsigned long nr_to_reclaim)\nmm/vmscan.c-7765-{\nmm/vmscan.c:7766:\tstruct scan_control sc = {\nmm/vmscan.c-7767-\t\t.nr_to_reclaim = nr_to_reclaim,\n--\nmm/vmscan.c=7980=unsigned long node_reclaim(struct pglist_data *pgdat, gfp_t gfp_mask, unsigned int order)\n--\nmm/vmscan.c-7984-\tconst unsigned long nr_pages = 1 \u003c\u003c order;\nmm/vmscan.c:7985:\tstruct scan_control sc = {\nmm/vmscan.c-7986-\t\t.nr_to_reclaim = max(nr_pages, SWAP_CLUSTER_MAX),\n--\nmm/vmscan.c=8062=int user_proactive_reclaim(char *buf,\n--\nmm/vmscan.c-8138-\t\t} else {\nmm/vmscan.c:8139:\t\t\tstruct scan_control sc = {\nmm/vmscan.c-8140-\t\t\t\t.gfp_mask = current_gfp_context(gfp_mask),\n"
}

53/3 2026/09/14 12:43 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 33962
							output: 32
							thoughts: 25
						

Thoughts:
Here's my take on the `struct percpu_swap_cluster` initialization and allocation within `mm/swapfile.c`.

**Diving into percpu_swap_cluster**

Alright, time to refresh my memory on how `struct percpu_swap_cluster` is handled. Specifically, I need to focus on its initialization and allocation within the `mm/swapfile.c` file. I'm assuming there's some interesting interplay between per-CPU data and the swap subsystem happening here, and I want to understand the lifecycle of these clusters. Are they allocated at boot? On demand? What's the impact of NUMA architecture in this context? I'm hoping to see how the system ensures that swap cluster management is efficient and takes advantage of potential locality. Let's see how the kernel engineers designed this.



54/3 2026/09/14 12:43 tool grepper
0m Args:
{
  "Expression": "percpu_swap_cluster",
  "PathPrefix": "mm/swapfile.c"
}

Results:
{
  "Output": "mm/swapfile.c=122=atomic_t nr_rotate_swap = ATOMIC_INIT(0);\nmm/swapfile.c-123-\nmm/swapfile.c:124:struct percpu_swap_cluster {\nmm/swapfile.c-125-\tstruct swap_info_struct *si[SWAP_NR_ORDERS];\n--\nmm/swapfile.c-130-\nmm/swapfile.c:131:static DEFINE_PER_CPU(struct percpu_swap_cluster, percpu_swap_cluster) = {\nmm/swapfile.c-132-\t.si = { NULL },\n--\nmm/swapfile.c=565=swap_cluster_populate(struct swap_info_struct *si,\n--\nmm/swapfile.c-573-\t */\nmm/swapfile.c:574:\tlockdep_assert_held(\u0026this_cpu_ptr(\u0026percpu_swap_cluster)-\u003elock);\nmm/swapfile.c-575-\tif (!(si-\u003eflags \u0026 SWP_SOLIDSTATE))\n--\nmm/swapfile.c-590-\t\tspin_unlock(\u0026si-\u003eglobal_cluster_lock);\nmm/swapfile.c:591:\tlocal_unlock(\u0026percpu_swap_cluster.lock);\nmm/swapfile.c-592-\n--\nmm/swapfile.c-603-\t */\nmm/swapfile.c:604:\tlocal_lock(\u0026percpu_swap_cluster.lock);\nmm/swapfile.c-605-\tif (!(si-\u003eflags \u0026 SWP_SOLIDSTATE))\n--\nmm/swapfile.c=701=static bool swap_do_scheduled_discard(struct swap_info_struct *si)\n--\nmm/swapfile.c-711-\t\t * Delete the cluster from list to prepare for discard, but keep\nmm/swapfile.c:712:\t\t * the CLUSTER_FLAG_DISCARD flag, percpu_swap_cluster could be\nmm/swapfile.c-713-\t\t * pointing to it, or ran into by relocate_cluster.\n--\nmm/swapfile.c=1002=static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,\n--\nmm/swapfile.c-1044-\tif (si-\u003eflags \u0026 SWP_SOLIDSTATE) {\nmm/swapfile.c:1045:\t\tthis_cpu_write(percpu_swap_cluster.offset[order], next);\nmm/swapfile.c:1046:\t\tthis_cpu_write(percpu_swap_cluster.si[order], si);\nmm/swapfile.c:1047:\t\tthis_cpu_write(percpu_swap_cluster.allow_offload_swap[order],\nmm/swapfile.c-1048-\t\t\t       current_reclaim_allows_offload_swap());\n--\nmm/swapfile.c=1434=static bool swap_alloc_fast(struct folio *folio)\n--\nmm/swapfile.c-1444-\t */\nmm/swapfile.c:1445:\tsi = this_cpu_read(percpu_swap_cluster.si[order]);\nmm/swapfile.c:1446:\toffset = this_cpu_read(percpu_swap_cluster.offset[order]);\nmm/swapfile.c-1447-\tif (!si || !offset || !get_swap_device_info(si))\n--\nmm/swapfile.c-1457-\t */\nmm/swapfile.c:1458:\tif (this_cpu_read(percpu_swap_cluster.allow_offload_swap[order]) !=\nmm/swapfile.c-1459-\t    current_reclaim_allows_offload_swap()) {\n--\nmm/swapfile.c=1823=int folio_alloc_swap(struct folio *folio)\n--\nmm/swapfile.c-1849-again:\nmm/swapfile.c:1850:\tlocal_lock(\u0026percpu_swap_cluster.lock);\nmm/swapfile.c-1851-\tif (!swap_alloc_fast(folio))\nmm/swapfile.c-1852-\t\tswap_alloc_slow(folio);\nmm/swapfile.c:1853:\tlocal_unlock(\u0026percpu_swap_cluster.lock);\nmm/swapfile.c-1854-\n--\nmm/swapfile.c=2255=swp_entry_t swap_alloc_hibernation_slot(int type)\n--\nmm/swapfile.c-2268-\t */\nmm/swapfile.c:2269:\tlocal_lock(\u0026percpu_swap_cluster.lock);\nmm/swapfile.c:2270:\tpcp_si = this_cpu_read(percpu_swap_cluster.si[0]);\nmm/swapfile.c:2271:\tpcp_offset = this_cpu_read(percpu_swap_cluster.offset[0]);\nmm/swapfile.c-2272-\tif (pcp_si == si \u0026\u0026 pcp_offset) {\n--\nmm/swapfile.c-2280-\t\toffset = cluster_alloc_swap_entry(si, NULL);\nmm/swapfile.c:2281:\tlocal_unlock(\u0026percpu_swap_cluster.lock);\nmm/swapfile.c-2282-\tif (offset)\n--\nmm/swapfile.c=3222=static void free_swap_cluster_info(struct swap_cluster_info *cluster_info,\n--\nmm/swapfile.c-3246- */\nmm/swapfile.c:3247:static void flush_percpu_swap_cluster(struct swap_info_struct *si)\nmm/swapfile.c-3248-{\n--\nmm/swapfile.c-3252-\tfor_each_possible_cpu(cpu) {\nmm/swapfile.c:3253:\t\tpcp_si = per_cpu_ptr(percpu_swap_cluster.si, cpu);\nmm/swapfile.c-3254-\t\t/*\n--\nmm/swapfile.c=3265=SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)\n--\nmm/swapfile.c-3350-\tflush_work(\u0026p-\u003ereclaim_work);\nmm/swapfile.c:3351:\tflush_percpu_swap_cluster(p);\nmm/swapfile.c-3352-\n"
}

55/3 2026/09/14 12:43 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 35825
							output: 34
							thoughts: 271
						

Thoughts:
**My Analysis of Task Reclaim State**

Okay, so `percpu_swap_cluster` is defined as static per-CPU storage, initialized to zero. That makes sense; it's going to reside in the BSS/per-CPU init section. No surprises there.

Now, let's dive into `set_task_reclaim_state`. Examining the code: We've got `task->reclaim_state` being set. There are WARN_ON_ONCE checks here for potential overwrites and double-nulling, good defensive programming. The reclaim state allows for offload swapping if `rs` is non-null.

But, a critical question arises: Where is `task->reclaim_state` *cleared*? The code sets it, but I don't see an explicit nulling operation. We should be searching for the places where `set_task_reclaim_state` is called to understand the lifecycle of `task->reclaim_state`. This is crucial to preventing memory leaks or unexpected behavior. Time to search for the calls.



56/3 2026/09/14 12:44 tool grepper
0m Args:
{
  "Expression": "set_task_reclaim_state",
  "PathPrefix": "mm/vmscan.c"
}

Results:
{
  "Output": "mm/vmscan.c=290=static inline bool is_exec_file_folio(const struct folio *folio,\n--\nmm/vmscan.c-295-\nmm/vmscan.c:296:static void set_task_reclaim_state(struct task_struct *task,\nmm/vmscan.c-297-\t\t\t\t   struct scan_control *sc)\n--\nmm/vmscan.c=5921=static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,\n--\nmm/vmscan.c-5947-\nmm/vmscan.c:5948:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-5949-\tflags = memalloc_noreclaim_save();\n--\nmm/vmscan.c-5998-\tmemalloc_noreclaim_restore(flags);\nmm/vmscan.c:5999:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-6000-\n--\nmm/vmscan.c=6920=unsigned long try_to_free_pages(struct zonelist *zonelist, int order,\n--\nmm/vmscan.c-6951-\nmm/vmscan.c:6952:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-6953-\ttrace_mm_vmscan_direct_reclaim_begin(sc.gfp_mask, order, NULL);\n--\nmm/vmscan.c-6957-\ttrace_mm_vmscan_direct_reclaim_end(nr_reclaimed, NULL);\nmm/vmscan.c:6958:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-6959-\n--\nmm/vmscan.c=6965=unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,\n--\nmm/vmscan.c-6994-\nmm/vmscan.c:6995:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-6996-\ttrace_mm_vmscan_memcg_reclaim_begin(sc.gfp_mask, 0, memcg);\n--\nmm/vmscan.c-7002-\ttrace_mm_vmscan_memcg_reclaim_end(nr_reclaimed, memcg);\nmm/vmscan.c:7003:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-7004-\n--\nmm/vmscan.c=7269=static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx)\n--\nmm/vmscan.c-7284-\t\t\t\t\t    highest_zoneidx);\nmm/vmscan.c:7285:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-7286-\tpsi_memstall_enter(\u0026pflags);\n--\nmm/vmscan.c-7470-\tpsi_memstall_leave(\u0026pflags);\nmm/vmscan.c:7471:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-7472-\n--\nmm/vmscan.c=7764=unsigned long shrink_all_memory(unsigned long nr_to_reclaim)\n--\nmm/vmscan.c-7781-\tnoreclaim_flag = memalloc_noreclaim_save();\nmm/vmscan.c:7782:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-7783-\n--\nmm/vmscan.c-7785-\nmm/vmscan.c:7786:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-7787-\tmemalloc_noreclaim_restore(noreclaim_flag);\n--\nmm/vmscan.c=7944=static unsigned long __node_reclaim(struct pglist_data *pgdat,\n--\nmm/vmscan.c-7962-\tnoreclaim_flag = memalloc_noreclaim_save();\nmm/vmscan.c:7963:\tset_task_reclaim_state(p, sc);\nmm/vmscan.c-7964-\n--\nmm/vmscan.c-7968-\nmm/vmscan.c:7969:\tset_task_reclaim_state(p, NULL);\nmm/vmscan.c-7970-\tmemalloc_noreclaim_restore(noreclaim_flag);\n"
}

57/3 2026/09/14 12:44 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 37558
							output: 35
							thoughts: 25
						

Thoughts:
Okay, here's my interpretation of that thought process, assuming I'm deep in my domain of expertise:

**Checking `user_proactive_reclaim`**

Alright, that's interesting. The mention of `user_proactive_reclaim` just pinged something in my brain. It could be relevant to the current problem. My initial thought is to dive right in and dissect what that particular element entails. It might be the key to unlocking the next stage. Let's make sure that's exactly what we need, though, and go over all the related parameters to ensure it matches the broader goals.



58/3 2026/09/14 12:44 tool read-file
0m Args:
{
  "File": "mm/vmscan.c",
  "FirstLine": 8130,
  "LineCount": 40
}

Results:
{
  "Contents": "8130:\t\t\t\tunsigned int reclaim_options;\n8131:\t\n8132:\t\t\t\treclaim_options = MEMCG_RECLAIM_MAY_SWAP |\n8133:\t\t\t\t\t\t  MEMCG_RECLAIM_PROACTIVE;\n8134:\t\t\t\treclaimed = try_to_free_mem_cgroup_pages(memcg,\n8135:\t\t\t\t\t\t\t batch_size, gfp_mask,\n8136:\t\t\t\t\t\t\t reclaim_options,\n8137:\t\t\t\t\t\t\t swappiness == -1 ? NULL : \u0026swappiness);\n8138:\t\t\t} else {\n8139:\t\t\t\tstruct scan_control sc = {\n8140:\t\t\t\t\t.gfp_mask = current_gfp_context(gfp_mask),\n8141:\t\t\t\t\t.reclaim_idx = gfp_zone(gfp_mask),\n8142:\t\t\t\t\t.proactive_swappiness = swappiness == -1 ? NULL : \u0026swappiness,\n8143:\t\t\t\t\t.priority = DEF_PRIORITY,\n8144:\t\t\t\t\t.may_writepage = 1,\n8145:\t\t\t\t\t.nr_to_reclaim = max(batch_size, SWAP_CLUSTER_MAX),\n8146:\t\t\t\t\t.may_unmap = 1,\n8147:\t\t\t\t\t.may_swap = 1,\n8148:\t\t\t\t\t.proactive = 1,\n8149:\t\t\t\t\t.allow_offload_swap = 1,\n8150:\t\t\t\t};\n8151:\t\n8152:\t\t\t\tif (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED,\n8153:\t\t\t\t\t\t\t  \u0026pgdat-\u003eflags))\n8154:\t\t\t\t\treturn -EBUSY;\n8155:\t\n8156:\t\t\t\treclaimed = __node_reclaim(pgdat, batch_size, \u0026sc);\n8157:\t\t\t\tclear_bit_unlock(PGDAT_RECLAIM_LOCKED, \u0026pgdat-\u003eflags);\n8158:\t\t\t}\n8159:\t\n8160:\t\t\tif (!reclaimed \u0026\u0026 !nr_retries--)\n8161:\t\t\t\treturn -EAGAIN;\n8162:\t\n8163:\t\t\tnr_reclaimed += reclaimed;\n8164:\t\t}\n8165:\t\n8166:\t\treturn 0;\n8167:\t}\n8168:\t\n8169:\t/**\n"
}

59/3 2026/09/14 12:44 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 38366
							output: 18
							thoughts: 367
						

Thoughts:
**My Analysis of `__node_reclaim` and `reclaim_state`**

Okay, so `__node_reclaim` is calling `set_task_reclaim_state` with `sc`, and then again with `NULL`. That's the core of the issue. When `sc` is NULL, I see that the `task->reclaim_state` is set to NULL because `rs` becomes NULL. Seems straightforward.

But when `sc` *isn't* NULL, things get more interesting. `rs` gets set to point to `&sc->reclaim_state`. Importantly, from what I can tell from the callers, `sc` is defined with a designated initializer. This means every member of `sc`, including the nested `sc.reclaim_state`, gets zeroed out. That means `sc->reclaim_state.reclaimed` is initialized to zero. And `rs->allow_offload_swap` gets explicitly assigned the value of `sc->allow_offload_swap`. All of this happens, in the most common case, with designated initializers, so the reclaim state is zeroed.

Hold on a minute, let me take a quick look at `include/linux/swap.h` to refresh my memory about `struct reclaim_state`. Aha!

```c
struct reclaim_state {
	/* pages reclaimed outside of LRU-based reclaim */
	unsigned long reclaimed;
	/* this reclaim context may use offload-only swap */
	bool allow_offload_swap;
#ifdef CONFIG_LRU_GEN
	/* per-thread mm walk data */
	struct lru_gen_mm_walk *mm_walk;
#endif
};
```

Okay, so `reclaimed` is an `unsigned long`, `allow_offload_swap` is a `bool`, and there's a conditional `mm_walk` member for `CONFIG_LRU_GEN`.

Now, the critical question: *is `reclaim_state` ever created elsewhere or dynamically allocated in a way that *doesn't* guarantee zero-initialization?* That's the crucial question. I need to make sure this is being treated consistently. Let's grep for `reclaim_state` across the codebase to confirm if my assumption about initialization is correct, or if there might be other places where it is created and potentially not zeroed. I want to make sure I am not missing anything that would contradict my assumption.



60/3 2026/09/14 12:44 tool grepper
0m Args:
{
  "Expression": "reclaim_state"
}

Results:
{
  "Output": "drivers/gpu/drm/panthor/panthor_gem.c=121=static bool is_gpu_mapped(struct panthor_gem_object *bo,\ndrivers/gpu/drm/panthor/panthor_gem.c:122:\t\t\t  enum panthor_gem_reclaim_state *state)\ndrivers/gpu/drm/panthor/panthor_gem.c-123-{\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-142-\ndrivers/gpu/drm/panthor/panthor_gem.c:143:static enum panthor_gem_reclaim_state\ndrivers/gpu/drm/panthor/panthor_gem.c:144:panthor_gem_evaluate_reclaim_state_locked(struct panthor_gem_object *bo)\ndrivers/gpu/drm/panthor/panthor_gem.c-145-{\ndrivers/gpu/drm/panthor/panthor_gem.c:146:\tenum panthor_gem_reclaim_state gpu_mapped_state;\ndrivers/gpu/drm/panthor/panthor_gem.c-147-\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-167-\ndrivers/gpu/drm/panthor/panthor_gem.c:168:void panthor_gem_update_reclaim_state_locked(struct panthor_gem_object *bo,\ndrivers/gpu/drm/panthor/panthor_gem.c:169:\t\t\t\t\t     enum panthor_gem_reclaim_state *old_statep)\ndrivers/gpu/drm/panthor/panthor_gem.c-170-{\ndrivers/gpu/drm/panthor/panthor_gem.c-171-\tstruct panthor_device *ptdev = container_of(bo-\u003ebase.dev, struct panthor_device, base);\ndrivers/gpu/drm/panthor/panthor_gem.c:172:\tenum panthor_gem_reclaim_state old_state = bo-\u003ereclaim_state;\ndrivers/gpu/drm/panthor/panthor_gem.c:173:\tenum panthor_gem_reclaim_state new_state;\ndrivers/gpu/drm/panthor/panthor_gem.c-174-\tbool was_gpu_mapped, is_gpu_mapped;\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-178-\ndrivers/gpu/drm/panthor/panthor_gem.c:179:\tnew_state = panthor_gem_evaluate_reclaim_state_locked(bo);\ndrivers/gpu/drm/panthor/panthor_gem.c-180-\tif (new_state == old_state)\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-213-\ndrivers/gpu/drm/panthor/panthor_gem.c:214:\tbo-\u003ereclaim_state = new_state;\ndrivers/gpu/drm/panthor/panthor_gem.c-215-}\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=259=static int panthor_gem_backing_pin_locked(struct panthor_gem_object *bo)\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-272-\t\tmutex_lock(\u0026bo-\u003ebase.gpuva.lock);\ndrivers/gpu/drm/panthor/panthor_gem.c:273:\t\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_gem.c-274-\t\tmutex_unlock(\u0026bo-\u003ebase.gpuva.lock);\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=280=static void panthor_gem_backing_unpin_locked(struct panthor_gem_object *bo)\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-289-\t\tmutex_lock(\u0026bo-\u003ebase.gpuva.lock);\ndrivers/gpu/drm/panthor/panthor_gem.c:290:\t\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_gem.c-291-\t\tmutex_unlock(\u0026bo-\u003ebase.gpuva.lock);\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=676=static void panthor_gem_evict_locked(struct panthor_gem_object *bo)\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-693-\tpanthor_gem_backing_cleanup_locked(bo);\ndrivers/gpu/drm/panthor/panthor_gem.c:694:\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_gem.c-695-}\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=742=static int panthor_gem_mmap(struct drm_gem_object *obj, struct vm_area_struct *vma)\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-771-\t\t\tmutex_lock(\u0026bo-\u003ebase.gpuva.lock);\ndrivers/gpu/drm/panthor/panthor_gem.c:772:\t\t\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_gem.c-773-\t\t\tmutex_unlock(\u0026bo-\u003ebase.gpuva.lock);\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=842=static vm_fault_t blocking_page_setup(struct vm_fault *vmf, unsigned int order,\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-862-\t\tmutex_lock(\u0026bo-\u003ebase.gpuva.lock);\ndrivers/gpu/drm/panthor/panthor_gem.c:863:\t\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_gem.c-864-\t\tmutex_unlock(\u0026bo-\u003ebase.gpuva.lock);\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=937=static void panthor_gem_vm_close(struct vm_area_struct *vma)\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-949-\t\tmutex_lock(\u0026bo-\u003ebase.gpuva.lock);\ndrivers/gpu/drm/panthor/panthor_gem.c:950:\t\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_gem.c-951-\t\tmutex_unlock(\u0026bo-\u003ebase.gpuva.lock);\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=983=panthor_gem_alloc_object(u32 flags)\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-990-\ndrivers/gpu/drm/panthor/panthor_gem.c:991:\tbo-\u003ereclaim_state = PANTHOR_GEM_UNRECLAIMABLE;\ndrivers/gpu/drm/panthor/panthor_gem.c-992-\tbo-\u003ebase.funcs = \u0026panthor_gem_funcs;\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=1413=static bool panthor_gem_try_evict_no_resv_wait(struct drm_gem_object *obj,\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-1421-\tstruct drm_gpuvm_bo *vm_bo, *last_locked = NULL;\ndrivers/gpu/drm/panthor/panthor_gem.c:1422:\tenum panthor_gem_reclaim_state old_state;\ndrivers/gpu/drm/panthor/panthor_gem.c-1423-\tint ret = 0;\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-1466-\t */\ndrivers/gpu/drm/panthor/panthor_gem.c:1467:\tpanthor_gem_update_reclaim_state_locked(bo, \u0026old_state);\ndrivers/gpu/drm/panthor/panthor_gem.c:1468:\tif (old_state \u003c bo-\u003ereclaim_state) {\ndrivers/gpu/drm/panthor/panthor_gem.c-1469-\t\tret = -EAGAIN;\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=1642=static void panthor_gem_debugfs_bo_print(struct panthor_gem_object *bo,\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-1645-{\ndrivers/gpu/drm/panthor/panthor_gem.c:1646:\tenum panthor_gem_reclaim_state reclaim_state = bo-\u003ereclaim_state;\ndrivers/gpu/drm/panthor/panthor_gem.c-1647-\tunsigned int refcount = kref_read(\u0026bo-\u003ebase.refcount);\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-1686-\ttotals-\u003eresident += resident_size;\ndrivers/gpu/drm/panthor/panthor_gem.c:1687:\tif (reclaim_state != PANTHOR_GEM_UNRECLAIMABLE)\ndrivers/gpu/drm/panthor/panthor_gem.c-1688-\t\ttotals-\u003ereclaimable += resident_size;\n--\ndrivers/gpu/drm/panthor/panthor_gem.h=96=struct panthor_gem_dev_map {\n--\ndrivers/gpu/drm/panthor/panthor_gem.h-101-/**\ndrivers/gpu/drm/panthor/panthor_gem.h:102: * enum panthor_gem_reclaim_state - Reclaim state of a GEM object\ndrivers/gpu/drm/panthor/panthor_gem.h-103- *\n--\ndrivers/gpu/drm/panthor/panthor_gem.h-106- */\ndrivers/gpu/drm/panthor/panthor_gem.h:107:enum panthor_gem_reclaim_state {\ndrivers/gpu/drm/panthor/panthor_gem.h-108-\t/**\n--\ndrivers/gpu/drm/panthor/panthor_gem.h=163=struct panthor_gem_object {\n--\ndrivers/gpu/drm/panthor/panthor_gem.h-175-\ndrivers/gpu/drm/panthor/panthor_gem.h:176:\t/** @reclaim_state: Cached reclaim state */\ndrivers/gpu/drm/panthor/panthor_gem.h:177:\tenum panthor_gem_reclaim_state reclaim_state;\ndrivers/gpu/drm/panthor/panthor_gem.h-178-\n--\ndrivers/gpu/drm/panthor/panthor_gem.h=266=int panthor_gem_swapin_locked(struct panthor_gem_object *bo);\ndrivers/gpu/drm/panthor/panthor_gem.h:267:void panthor_gem_update_reclaim_state_locked(struct panthor_gem_object *bo,\ndrivers/gpu/drm/panthor/panthor_gem.h:268:\t\t\t\t\t     enum panthor_gem_reclaim_state *old_state);\ndrivers/gpu/drm/panthor/panthor_gem.h-269-int panthor_gem_shrinker_init(struct panthor_device *ptdev);\n--\ndrivers/gpu/drm/panthor/panthor_mmu.c=1153=static void panthor_vm_bo_free(struct drm_gpuvm_bo *vm_bo)\n--\ndrivers/gpu/drm/panthor/panthor_mmu.c-1161-\tmutex_lock(\u0026bo-\u003ebase.gpuva.lock);\ndrivers/gpu/drm/panthor/panthor_mmu.c:1162:\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_mmu.c-1163-\tmutex_unlock(\u0026bo-\u003ebase.gpuva.lock);\n--\ndrivers/gpu/drm/panthor/panthor_mmu.c=1322=static int panthor_vm_prepare_map_op_ctx(struct panthor_vm_op_ctx *op_ctx,\n--\ndrivers/gpu/drm/panthor/panthor_mmu.c-1405-\tmutex_lock(\u0026bo-\u003ebase.gpuva.lock);\ndrivers/gpu/drm/panthor/panthor_mmu.c:1406:\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_mmu.c-1407-\tmutex_unlock(\u0026bo-\u003ebase.gpuva.lock);\n--\ndrivers/gpu/drm/panthor/panthor_mmu.c=2662=static int panthor_vm_bo_validate(struct drm_gpuvm_bo *vm_bo,\n--\ndrivers/gpu/drm/panthor/panthor_mmu.c-2677-\tmutex_lock(\u0026bo-\u003ebase.gpuva.lock);\ndrivers/gpu/drm/panthor/panthor_mmu.c:2678:\tpanthor_gem_update_reclaim_state_locked(bo, NULL);\ndrivers/gpu/drm/panthor/panthor_mmu.c-2679-\tmutex_unlock(\u0026bo-\u003ebase.gpuva.lock);\n--\ninclude/linux/sched.h=77=struct rcu_node;\ninclude/linux/sched.h:78:struct reclaim_state;\ninclude/linux/sched.h-79-struct root_domain;\n--\ninclude/linux/sched.h=835=struct task_struct {\n--\ninclude/linux/sched.h-1318-\t/* VM state: */\ninclude/linux/sched.h:1319:\tstruct reclaim_state\t\t*reclaim_state;\ninclude/linux/sched.h-1320-\n--\ninclude/linux/swap.h=120=union swap_header {\n--\ninclude/linux/swap.h-137-/*\ninclude/linux/swap.h:138: * current-\u003ereclaim_state points to one of these when a task is running\ninclude/linux/swap.h-139- * memory reclaim\ninclude/linux/swap.h-140- */\ninclude/linux/swap.h:141:struct reclaim_state {\ninclude/linux/swap.h-142-\t/* pages reclaimed outside of LRU-based reclaim */\n--\ninclude/linux/swap.h=152=static inline bool current_reclaim_allows_offload_swap(void)\ninclude/linux/swap.h-153-{\ninclude/linux/swap.h:154:\treturn current-\u003ereclaim_state \u0026\u0026\ninclude/linux/swap.h:155:\t       current-\u003ereclaim_state-\u003eallow_offload_swap;\ninclude/linux/swap.h-156-}\n--\ninclude/linux/swap.h=166=static inline void mm_account_reclaimed_pages(unsigned long pages)\ninclude/linux/swap.h-167-{\ninclude/linux/swap.h:168:\tif (current-\u003ereclaim_state)\ninclude/linux/swap.h:169:\t\tcurrent-\u003ereclaim_state-\u003ereclaimed += pages;\ninclude/linux/swap.h-170-}\n--\nmm/vmscan.c=76=struct scan_control {\n--\nmm/vmscan.c-183-\t/* for recording the reclaimed slab by now */\nmm/vmscan.c:184:\tstruct reclaim_state reclaim_state;\nmm/vmscan.c-185-};\n--\nmm/vmscan.c=290=static inline bool is_exec_file_folio(const struct folio *folio,\n--\nmm/vmscan.c-295-\nmm/vmscan.c:296:static void set_task_reclaim_state(struct task_struct *task,\nmm/vmscan.c-297-\t\t\t\t   struct scan_control *sc)\nmm/vmscan.c-298-{\nmm/vmscan.c:299:\tstruct reclaim_state *rs = sc ? \u0026sc-\u003ereclaim_state : NULL;\nmm/vmscan.c-300-\nmm/vmscan.c-301-\t/* Check for an overwrite */\nmm/vmscan.c:302:\tWARN_ON_ONCE(rs \u0026\u0026 task-\u003ereclaim_state);\nmm/vmscan.c-303-\nmm/vmscan.c-304-\t/* Check for the nulling of an already-nulled member */\nmm/vmscan.c:305:\tWARN_ON_ONCE(!rs \u0026\u0026 !task-\u003ereclaim_state);\nmm/vmscan.c-306-\n--\nmm/vmscan.c-308-\t\trs-\u003eallow_offload_swap = sc-\u003eallow_offload_swap;\nmm/vmscan.c:309:\ttask-\u003ereclaim_state = rs;\nmm/vmscan.c-310-}\n--\nmm/vmscan.c-312-/*\nmm/vmscan.c:313: * flush_reclaim_state(): add pages reclaimed outside of LRU-based reclaim to\nmm/vmscan.c-314- * scan_control-\u003enr_reclaimed.\nmm/vmscan.c-315- */\nmm/vmscan.c:316:static void flush_reclaim_state(struct scan_control *sc)\nmm/vmscan.c-317-{\nmm/vmscan.c-318-\t/*\nmm/vmscan.c:319:\t * Currently, reclaim_state-\u003ereclaimed includes three types of pages\nmm/vmscan.c-320-\t * freed outside of vmscan:\n--\nmm/vmscan.c-345-\t */\nmm/vmscan.c:346:\tif (current-\u003ereclaim_state \u0026\u0026 root_reclaim(sc)) {\nmm/vmscan.c:347:\t\tsc-\u003enr_reclaimed += current-\u003ereclaim_state-\u003ereclaimed;\nmm/vmscan.c:348:\t\tcurrent-\u003ereclaim_state-\u003ereclaimed = 0;\nmm/vmscan.c-349-\t}\n--\nmm/vmscan.c=3969=static struct lru_gen_mm_walk *set_mm_walk(struct pglist_data *pgdat, bool force_alloc)\nmm/vmscan.c-3970-{\nmm/vmscan.c:3971:\tstruct lru_gen_mm_walk *walk = current-\u003ereclaim_state-\u003emm_walk;\nmm/vmscan.c-3972-\n--\nmm/vmscan.c-3983-\nmm/vmscan.c:3984:\tcurrent-\u003ereclaim_state-\u003emm_walk = walk;\nmm/vmscan.c-3985-\n--\nmm/vmscan.c=3989=static void clear_mm_walk(void)\nmm/vmscan.c-3990-{\nmm/vmscan.c:3991:\tstruct lru_gen_mm_walk *walk = current-\u003ereclaim_state-\u003emm_walk;\nmm/vmscan.c-3992-\n--\nmm/vmscan.c-3995-\nmm/vmscan.c:3996:\tcurrent-\u003ereclaim_state-\u003emm_walk = NULL;\nmm/vmscan.c-3997-\n--\nmm/vmscan.c=4400=bool lru_gen_look_around(struct page_vma_mapped_walk *pvmw, unsigned int nr)\n--\nmm/vmscan.c-4431-\t/* avoid taking the LRU lock under the PTL when possible */\nmm/vmscan.c:4432:\twalk = current-\u003ereclaim_state ? current-\u003ereclaim_state-\u003emm_walk : NULL;\nmm/vmscan.c-4433-\n--\nmm/vmscan.c=5022=static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,\n--\nmm/vmscan.c-5090-\nmm/vmscan.c:5091:\twalk = current-\u003ereclaim_state-\u003emm_walk;\nmm/vmscan.c-5092-\tif (walk \u0026\u0026 walk-\u003ebatched) {\n--\nmm/vmscan.c=5247=static int shrink_one(struct lruvec *lruvec, struct scan_control *sc)\n--\nmm/vmscan.c-5280-\nmm/vmscan.c:5281:\tflush_reclaim_state(sc);\nmm/vmscan.c-5282-\n--\nmm/vmscan.c=5921=static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,\n--\nmm/vmscan.c-5947-\nmm/vmscan.c:5948:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-5949-\tflags = memalloc_noreclaim_save();\n--\nmm/vmscan.c-5998-\tmemalloc_noreclaim_restore(flags);\nmm/vmscan.c:5999:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-6000-\n--\nmm/vmscan.c=6410=static void shrink_node(pg_data_t *pgdat, struct scan_control *sc)\n--\nmm/vmscan.c-6436-\nmm/vmscan.c:6437:\tflush_reclaim_state(sc);\nmm/vmscan.c-6438-\n--\nmm/vmscan.c=6920=unsigned long try_to_free_pages(struct zonelist *zonelist, int order,\n--\nmm/vmscan.c-6951-\nmm/vmscan.c:6952:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-6953-\ttrace_mm_vmscan_direct_reclaim_begin(sc.gfp_mask, order, NULL);\n--\nmm/vmscan.c-6957-\ttrace_mm_vmscan_direct_reclaim_end(nr_reclaimed, NULL);\nmm/vmscan.c:6958:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-6959-\n--\nmm/vmscan.c=6965=unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,\n--\nmm/vmscan.c-6994-\nmm/vmscan.c:6995:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-6996-\ttrace_mm_vmscan_memcg_reclaim_begin(sc.gfp_mask, 0, memcg);\n--\nmm/vmscan.c-7002-\ttrace_mm_vmscan_memcg_reclaim_end(nr_reclaimed, memcg);\nmm/vmscan.c:7003:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-7004-\n--\nmm/vmscan.c=7269=static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx)\n--\nmm/vmscan.c-7284-\t\t\t\t\t    highest_zoneidx);\nmm/vmscan.c:7285:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-7286-\tpsi_memstall_enter(\u0026pflags);\n--\nmm/vmscan.c-7470-\tpsi_memstall_leave(\u0026pflags);\nmm/vmscan.c:7471:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-7472-\n--\nmm/vmscan.c=7764=unsigned long shrink_all_memory(unsigned long nr_to_reclaim)\n--\nmm/vmscan.c-7781-\tnoreclaim_flag = memalloc_noreclaim_save();\nmm/vmscan.c:7782:\tset_task_reclaim_state(current, \u0026sc);\nmm/vmscan.c-7783-\n--\nmm/vmscan.c-7785-\nmm/vmscan.c:7786:\tset_task_reclaim_state(current, NULL);\nmm/vmscan.c-7787-\tmemalloc_noreclaim_restore(noreclaim_flag);\n--\nmm/vmscan.c=7944=static unsigned long __node_reclaim(struct pglist_data *pgdat,\n--\nmm/vmscan.c-7962-\tnoreclaim_flag = memalloc_noreclaim_save();\nmm/vmscan.c:7963:\tset_task_reclaim_state(p, sc);\nmm/vmscan.c-7964-\n--\nmm/vmscan.c-7968-\nmm/vmscan.c:7969:\tset_task_reclaim_state(p, NULL);\nmm/vmscan.c-7970-\tmemalloc_noreclaim_restore(noreclaim_flag);\n"
}

61/3 2026/09/14 12:44 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 45707
							output: 35
							thoughts: 156
						

Thoughts:
**Analyzing `set_task_reclaim_state` and `user_proactive_reclaim`**

Okay, so I'm looking at how `set_task_reclaim_state` is used. My initial observation is that every place that calls it seems to be passing either a pointer to an on-stack `struct scan_control` (initialized with a zeroed `reclaim_state`, which means `sc.reclaim_state` starts at zero) or `NULL`. Then, inside `set_task_reclaim_state`, it assigns `rs->allow_offload_swap` based on `sc->allow_offload_swap`.

Now, I need to understand this better, specifically in the context of `user_proactive_reclaim`. Wait a minute, what is going on here in `mm/vmscan.c`? 

```c
int user_proactive_reclaim(char *buf,
...
```

I should probably dive in and analyze everything within `user_proactive_reclaim` to get a good understanding of this relationship.



62/3 2026/09/14 12:44 tool read-file
0m Args:
{
  "File": "mm/vmscan.c",
  "FirstLine": 8115,
  "LineCount": 45
}

Results:
{
  "Contents": "8115:\t\t\tif (signal_pending(current))\n8116:\t\t\t\treturn -ERESTARTSYS;\n8117:\t\n8118:\t\t\t/* cgroup_rmdir() waits for us with cgroup_mutex held. */\n8119:\t\t\tif (memcg \u0026\u0026 memcg_is_dying(memcg))\n8120:\t\t\t\treturn -EAGAIN;\n8121:\t\n8122:\t\t\t/*\n8123:\t\t\t * This is the final attempt, drain percpu lru caches in the\n8124:\t\t\t * hope of introducing more evictable pages.\n8125:\t\t\t */\n8126:\t\t\tif (!nr_retries)\n8127:\t\t\t\tlru_add_drain_all();\n8128:\t\n8129:\t\t\tif (memcg) {\n8130:\t\t\t\tunsigned int reclaim_options;\n8131:\t\n8132:\t\t\t\treclaim_options = MEMCG_RECLAIM_MAY_SWAP |\n8133:\t\t\t\t\t\t  MEMCG_RECLAIM_PROACTIVE;\n8134:\t\t\t\treclaimed = try_to_free_mem_cgroup_pages(memcg,\n8135:\t\t\t\t\t\t\t batch_size, gfp_mask,\n8136:\t\t\t\t\t\t\t reclaim_options,\n8137:\t\t\t\t\t\t\t swappiness == -1 ? NULL : \u0026swappiness);\n8138:\t\t\t} else {\n8139:\t\t\t\tstruct scan_control sc = {\n8140:\t\t\t\t\t.gfp_mask = current_gfp_context(gfp_mask),\n8141:\t\t\t\t\t.reclaim_idx = gfp_zone(gfp_mask),\n8142:\t\t\t\t\t.proactive_swappiness = swappiness == -1 ? NULL : \u0026swappiness,\n8143:\t\t\t\t\t.priority = DEF_PRIORITY,\n8144:\t\t\t\t\t.may_writepage = 1,\n8145:\t\t\t\t\t.nr_to_reclaim = max(batch_size, SWAP_CLUSTER_MAX),\n8146:\t\t\t\t\t.may_unmap = 1,\n8147:\t\t\t\t\t.may_swap = 1,\n8148:\t\t\t\t\t.proactive = 1,\n8149:\t\t\t\t\t.allow_offload_swap = 1,\n8150:\t\t\t\t};\n8151:\t\n8152:\t\t\t\tif (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED,\n8153:\t\t\t\t\t\t\t  \u0026pgdat-\u003eflags))\n8154:\t\t\t\t\treturn -EBUSY;\n8155:\t\n8156:\t\t\t\treclaimed = __node_reclaim(pgdat, batch_size, \u0026sc);\n8157:\t\t\t\tclear_bit_unlock(PGDAT_RECLAIM_LOCKED, \u0026pgdat-\u003eflags);\n8158:\t\t\t}\n8159:\t\n"
}

63/3 2026/09/14 12:44 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 46858
							output: 35
							thoughts: 320
						

Thoughts:
**My Analysis of Memory Reclaim in the Linux Kernel**

Okay, so I'm digging into how the kernel handles memory reclaim, specifically within the context of cgroups. I see that when `user_proactive_reclaim` is triggered, the function `try_to_free_mem_cgroup_pages` is called, provided the `memcg` is not a null pointer. Makes sense – we only try to free pages if we're dealing with a cgroup.

Looking inside `try_to_free_mem_cgroup_pages`, I see the core logic wrapped in a `scan_control` structure. The initialization looks pretty standard: C99 designated initializers are used, so any unspecified fields will be zeroed out automatically, which is cleaner than older approaches.

I need to understand what this `scan_control` struct is configuring for the memory reclaim. `nr_to_reclaim` is set to the maximum between `nr_pages` and `SWAP_CLUSTER_MAX`, indicating it sets a target number of pages, or it might be clustering swap operations? The `gfp_mask` is set using `current_gfp_context` which seems logical. The `reclaim_idx` is set to `MAX_NR_ZONES - 1` indicating we start at the highest zone, which is probably the most likely place to start reclaiming from. `target_mem_cgroup` is set to the provided `memcg`, ensuring the reclaim operates within the correct cgroup.

The rest of the fields give more insight into reclaim behavior, such as:
-   `may_writepage` is controlled by a `laptop_mode` variable which is not fully clear from the code, but indicates writepage operations should be avoided on laptops.
-   `may_unmap` and `may_swap` are self-explanatory.
-   `proactive` and `allow_offload_swap` are controlled by the `reclaim_options` flag.

Now, my next step is to compare `node_reclaim` and `__node_reclaim` to see if I can contrast how global and cgroup-specific memory is managed.



64/3 2026/09/14 12:44 tool read-file
0m Args:
{
  "File": "mm/vmscan.c",
  "FirstLine": 7980,
  "LineCount": 40
}

Results:
{
  "Contents": "7980:\tunsigned long node_reclaim(struct pglist_data *pgdat, gfp_t gfp_mask, unsigned int order)\n7981:\t{\n7982:\t\tunsigned long ret;\n7983:\t\t/* Minimum pages needed in order to stay on node */\n7984:\t\tconst unsigned long nr_pages = 1 \u003c\u003c order;\n7985:\t\tstruct scan_control sc = {\n7986:\t\t\t.nr_to_reclaim = max(nr_pages, SWAP_CLUSTER_MAX),\n7987:\t\t\t.gfp_mask = current_gfp_context(gfp_mask),\n7988:\t\t\t.order = order,\n7989:\t\t\t.priority = NODE_RECLAIM_PRIORITY,\n7990:\t\t\t.may_writepage = !!(node_reclaim_mode \u0026 RECLAIM_WRITE),\n7991:\t\t\t.may_unmap = !!(node_reclaim_mode \u0026 RECLAIM_UNMAP),\n7992:\t\t\t.may_swap = 1,\n7993:\t\t\t.reclaim_idx = gfp_zone(gfp_mask),\n7994:\t\t};\n7995:\t\n7996:\t\t/*\n7997:\t\t * Node reclaim reclaims unmapped file backed pages and\n7998:\t\t * slab pages if we are over the defined limits.\n7999:\t\t *\n8000:\t\t * A small portion of unmapped file backed pages is needed for\n8001:\t\t * file I/O otherwise pages read by file I/O will be immediately\n8002:\t\t * thrown out if the node is overallocated. So we do not reclaim\n8003:\t\t * if less than a specified percentage of the node is used by\n8004:\t\t * unmapped file backed pages.\n8005:\t\t */\n8006:\t\tif (node_pagecache_reclaimable(pgdat) \u003c= pgdat-\u003emin_unmapped_pages \u0026\u0026\n8007:\t\t    node_page_state_pages(pgdat, NR_SLAB_RECLAIMABLE_B) \u003c=\n8008:\t\t    pgdat-\u003emin_slab_pages)\n8009:\t\t\treturn 0;\n8010:\t\n8011:\t\t/*\n8012:\t\t * Do not scan if the allocation should not be delayed.\n8013:\t\t */\n8014:\t\tif (!gfpflags_allow_blocking(gfp_mask) || (current-\u003eflags \u0026 PF_MEMALLOC))\n8015:\t\t\treturn 0;\n8016:\t\n8017:\t\t/*\n8018:\t\t * Only run node reclaim on the local node or on nodes that do not\n8019:\t\t * have associated processors. This will favor the local processor\n"
}

65/3 2026/09/14 12:44 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 48381
							output: 35
							thoughts: 13
						

Thoughts:
Okay, let's dive into this `node_reclaim` function. I need to understand its role within this system. I'm assuming it's handling the process of returning nodes to a pool, or perhaps freeing them for reuse. It's crucial to grasp the nuances here, especially if we're dealing with memory management or resource allocation. The context will tell me whether this is about physical nodes, virtual machines, or some other computational unit. I need to look closely at what's being reclaimed – data, resources, the node itself? Also, I will need to consider who initiated the reclaim, the reasons why, and any associated cleanup procedures. Is there any state that needs to be reset, resources that have to be deallocated, or connections that need to be severed? I'll be looking for potential bottlenecks, race conditions, or anything that could lead to resource leaks or unexpected behavior. Essentially, I need to understand the function’s purpose, its inputs, outputs, and its potential side effects within the larger architecture. Let's dig in and see the code!



66/3 2026/09/14 12:44 tool read-file
0m Args:
{
  "File": "mm/vmscan.c",
  "FirstLine": 8020,
  "LineCount": 40
}

Results:
{
  "Contents": "8020:\t\t * over remote processors and spread off node memory allocations\n8021:\t\t * as wide as possible.\n8022:\t\t */\n8023:\t\tif (node_state(pgdat-\u003enode_id, N_CPU) \u0026\u0026 pgdat-\u003enode_id != numa_node_id())\n8024:\t\t\treturn 0;\n8025:\t\n8026:\t\tif (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED, \u0026pgdat-\u003eflags))\n8027:\t\t\treturn 0;\n8028:\t\n8029:\t\tret = __node_reclaim(pgdat, nr_pages, \u0026sc);\n8030:\t\tclear_bit_unlock(PGDAT_RECLAIM_LOCKED, \u0026pgdat-\u003eflags);\n8031:\t\n8032:\t\tif (ret \u003e= nr_pages)\n8033:\t\t\tcount_vm_event(PGSCAN_ZONE_RECLAIM_SUCCESS);\n8034:\t\telse\n8035:\t\t\tcount_vm_event(PGSCAN_ZONE_RECLAIM_FAILED);\n8036:\t\n8037:\t\treturn ret;\n8038:\t}\n8039:\t\n8040:\t#else\n8041:\t\n8042:\tstatic unsigned long __node_reclaim(struct pglist_data *pgdat,\n8043:\t\t\t\t\t    unsigned long nr_pages,\n8044:\t\t\t\t\t    struct scan_control *sc)\n8045:\t{\n8046:\t\treturn 0;\n8047:\t}\n8048:\t\n8049:\t#endif\n8050:\t\n8051:\tenum {\n8052:\t\tMEMORY_RECLAIM_SWAPPINESS = 0,\n8053:\t\tMEMORY_RECLAIM_SWAPPINESS_MAX,\n8054:\t\tMEMORY_RECLAIM_NULL,\n8055:\t};\n8056:\tstatic const match_table_t tokens = {\n8057:\t\t{ MEMORY_RECLAIM_SWAPPINESS, \"swappiness=%d\"},\n8058:\t\t{ MEMORY_RECLAIM_SWAPPINESS_MAX, \"swappiness=max\"},\n8059:\t\t{ MEMORY_RECLAIM_NULL, NULL },\n"
}

67/3 2026/09/14 12:44 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 49248
							output: 19
							thoughts: 2727
						

Thoughts:
**KMSAN Analysis of the Offload Swap Patch**

Okay, let's break this down systematically. My primary concern is always around potential KMSAN issues: uninitialized memory reads or writes, and data leakage to userspace. I've gone through the patch meticulously.

First, the initialization. I see the `node_reclaim` function in `mm/vmscan.c` designates initializes `struct scan_control sc`. This is crucial. `sc.allow_offload_swap` *will* be `0` (false) here.  Then, within `__node_reclaim`, the `set_task_reclaim_state(p, sc)` call is key. It's setting the `current->reclaim_state` and, critically, `rs->allow_offload_swap` from the `sc->allow_offload_swap` field. This ensures that the reclaim state is *always* initialized, specifically to the value from the scan control structure.

Now, that `current_reclaim_allows_offload_swap()` function… I'm verifying it's safe. It reads `current->reclaim_state` and checks the `allow_offload_swap` member. Since `current->reclaim_state` is only ever set through the `set_task_reclaim_state` which always initializes it, it should always be safe to read.

Moving on to `folio_swap_full`, I had a slight moment of concern about `folio->swap` potentially being uninitialized. However, the `folio_test_swapcache(folio)` check *guarantees* that `folio->swap` is a valid swap entry, and the subsequent use of that pointer within `swap_area_needs_reclaim()` *should* be fine. The function uses `si->flags`, `si->pages` and other fields within the `swap_info_struct`, which are initialized during swap device setup (swapon). The `swap_area_needs_reclaim` then makes use of the `READ_ONCE` construct, so should not have any issues in itself either.

The calculations within `get_nr_swap_pages_eligible()` and `get_nr_swap_pages()` rely on atomic operations – specifically, reads from and calculations with `atomic_long_t` variables. I've confirmed that `nr_swap_pages` and `nr_swap_pages_offload_only` are correctly initialized as atomic variables.

The `swap_alloc_fast()` function uses `percpu_swap_cluster`, and the `this_cpu_read` calls within that function and associated percpu variables, I see this is declared with `DEFINE_PER_CPU`, which means this is zero-initialized and set appropriately when allocating/deallocating. Furthermore, the `alloc_swap_scan_cluster` write to this is only executed when valid swap is allocated. So no uninitialized reads here.

Now, the user-space interface. This is *critical*. I've verified that the patch:
1.  **Adds a `SWPOUT_OFFLOAD_REFUSED` vmstat event:** This is just an incrementing per-CPU counter – simple integer, and initialized to 0. No uninitialized memory or data leakage.
2.  **Modifies the `swapon` syscall:** Specifically the `SWAP_FLAG_OFFLOAD_ONLY` flag. The syscall receives an integer parameter by *value*. No user-space buffers are involved. The code checks `swap_flags` and uses bitwise operations, but doesn't expose any uninitialized data to the user.
3.  **Documentation:** The documentation changes are just informative, and have no code.
4.  **Self-Tests:** I looked at the selftest code, which is userspace code, so I did not need to analyse it.

Now, the four KMSAN criteria, let me tick them off:
1.  **Kernel structures to user space:** I have confirmed that no kernel structures containing newly added fields are copied to user space. There's no use of `copy_to_user`, `put_user`, netlink, or ioctl output modifications that expose uninitialized data.
2.  **Conditional logic on uninitialized variables:** I've carefully reviewed all added and modified conditional logic (within the KMSAN specific section), and confirmed that the variables used are initialized.
3.  **Initialization of complex data structures:** The patch introduces `struct reclaim_state`, `struct scan_control`, and `struct percpu_swap_cluster`.  I confirmed that these structures are *always* zero-initialized or explicitly set during allocation.
4.  **Bounds checks and lengths impacting existing buffers:** The patch primarily changes swap slot management.  It doesn't introduce


68/3 2026/09/14 12:45 tool grepper
0m Args:
{
  "Expression": "-\u003ereclaim_state"
}

Results:
{
  "Output": "drivers/gpu/drm/panthor/panthor_gem.c=168=void panthor_gem_update_reclaim_state_locked(struct panthor_gem_object *bo,\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-171-\tstruct panthor_device *ptdev = container_of(bo-\u003ebase.dev, struct panthor_device, base);\ndrivers/gpu/drm/panthor/panthor_gem.c:172:\tenum panthor_gem_reclaim_state old_state = bo-\u003ereclaim_state;\ndrivers/gpu/drm/panthor/panthor_gem.c-173-\tenum panthor_gem_reclaim_state new_state;\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-213-\ndrivers/gpu/drm/panthor/panthor_gem.c:214:\tbo-\u003ereclaim_state = new_state;\ndrivers/gpu/drm/panthor/panthor_gem.c-215-}\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=983=panthor_gem_alloc_object(u32 flags)\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-990-\ndrivers/gpu/drm/panthor/panthor_gem.c:991:\tbo-\u003ereclaim_state = PANTHOR_GEM_UNRECLAIMABLE;\ndrivers/gpu/drm/panthor/panthor_gem.c-992-\tbo-\u003ebase.funcs = \u0026panthor_gem_funcs;\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=1413=static bool panthor_gem_try_evict_no_resv_wait(struct drm_gem_object *obj,\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-1467-\tpanthor_gem_update_reclaim_state_locked(bo, \u0026old_state);\ndrivers/gpu/drm/panthor/panthor_gem.c:1468:\tif (old_state \u003c bo-\u003ereclaim_state) {\ndrivers/gpu/drm/panthor/panthor_gem.c-1469-\t\tret = -EAGAIN;\n--\ndrivers/gpu/drm/panthor/panthor_gem.c=1642=static void panthor_gem_debugfs_bo_print(struct panthor_gem_object *bo,\n--\ndrivers/gpu/drm/panthor/panthor_gem.c-1645-{\ndrivers/gpu/drm/panthor/panthor_gem.c:1646:\tenum panthor_gem_reclaim_state reclaim_state = bo-\u003ereclaim_state;\ndrivers/gpu/drm/panthor/panthor_gem.c-1647-\tunsigned int refcount = kref_read(\u0026bo-\u003ebase.refcount);\n--\ninclude/linux/swap.h=120=union swap_header {\n--\ninclude/linux/swap.h-137-/*\ninclude/linux/swap.h:138: * current-\u003ereclaim_state points to one of these when a task is running\ninclude/linux/swap.h-139- * memory reclaim\n--\ninclude/linux/swap.h=152=static inline bool current_reclaim_allows_offload_swap(void)\ninclude/linux/swap.h-153-{\ninclude/linux/swap.h:154:\treturn current-\u003ereclaim_state \u0026\u0026\ninclude/linux/swap.h:155:\t       current-\u003ereclaim_state-\u003eallow_offload_swap;\ninclude/linux/swap.h-156-}\n--\ninclude/linux/swap.h=166=static inline void mm_account_reclaimed_pages(unsigned long pages)\ninclude/linux/swap.h-167-{\ninclude/linux/swap.h:168:\tif (current-\u003ereclaim_state)\ninclude/linux/swap.h:169:\t\tcurrent-\u003ereclaim_state-\u003ereclaimed += pages;\ninclude/linux/swap.h-170-}\n--\nmm/vmscan.c=296=static void set_task_reclaim_state(struct task_struct *task,\n--\nmm/vmscan.c-298-{\nmm/vmscan.c:299:\tstruct reclaim_state *rs = sc ? \u0026sc-\u003ereclaim_state : NULL;\nmm/vmscan.c-300-\nmm/vmscan.c-301-\t/* Check for an overwrite */\nmm/vmscan.c:302:\tWARN_ON_ONCE(rs \u0026\u0026 task-\u003ereclaim_state);\nmm/vmscan.c-303-\nmm/vmscan.c-304-\t/* Check for the nulling of an already-nulled member */\nmm/vmscan.c:305:\tWARN_ON_ONCE(!rs \u0026\u0026 !task-\u003ereclaim_state);\nmm/vmscan.c-306-\n--\nmm/vmscan.c-308-\t\trs-\u003eallow_offload_swap = sc-\u003eallow_offload_swap;\nmm/vmscan.c:309:\ttask-\u003ereclaim_state = rs;\nmm/vmscan.c-310-}\n--\nmm/vmscan.c=316=static void flush_reclaim_state(struct scan_control *sc)\n--\nmm/vmscan.c-345-\t */\nmm/vmscan.c:346:\tif (current-\u003ereclaim_state \u0026\u0026 root_reclaim(sc)) {\nmm/vmscan.c:347:\t\tsc-\u003enr_reclaimed += current-\u003ereclaim_state-\u003ereclaimed;\nmm/vmscan.c:348:\t\tcurrent-\u003ereclaim_state-\u003ereclaimed = 0;\nmm/vmscan.c-349-\t}\n--\nmm/vmscan.c=3969=static struct lru_gen_mm_walk *set_mm_walk(struct pglist_data *pgdat, bool force_alloc)\nmm/vmscan.c-3970-{\nmm/vmscan.c:3971:\tstruct lru_gen_mm_walk *walk = current-\u003ereclaim_state-\u003emm_walk;\nmm/vmscan.c-3972-\n--\nmm/vmscan.c-3983-\nmm/vmscan.c:3984:\tcurrent-\u003ereclaim_state-\u003emm_walk = walk;\nmm/vmscan.c-3985-\n--\nmm/vmscan.c=3989=static void clear_mm_walk(void)\nmm/vmscan.c-3990-{\nmm/vmscan.c:3991:\tstruct lru_gen_mm_walk *walk = current-\u003ereclaim_state-\u003emm_walk;\nmm/vmscan.c-3992-\n--\nmm/vmscan.c-3995-\nmm/vmscan.c:3996:\tcurrent-\u003ereclaim_state-\u003emm_walk = NULL;\nmm/vmscan.c-3997-\n--\nmm/vmscan.c=4400=bool lru_gen_look_around(struct page_vma_mapped_walk *pvmw, unsigned int nr)\n--\nmm/vmscan.c-4431-\t/* avoid taking the LRU lock under the PTL when possible */\nmm/vmscan.c:4432:\twalk = current-\u003ereclaim_state ? current-\u003ereclaim_state-\u003emm_walk : NULL;\nmm/vmscan.c-4433-\n--\nmm/vmscan.c=5022=static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,\n--\nmm/vmscan.c-5090-\nmm/vmscan.c:5091:\twalk = current-\u003ereclaim_state-\u003emm_walk;\nmm/vmscan.c-5092-\tif (walk \u0026\u0026 walk-\u003ebatched) {\n"
}

69/3 2026/09/14 12:45 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 55004
							output: 390
							thoughts: 708
						

Thoughts:
**KMSAN Applicability Assessment: No Risk of Uninitialized Memory Reads or Leaks**

Okay, let's break this down. My analysis confirms that `set_task_reclaim_state` is the single point in the kernel that modifies `task->reclaim_state`, and it's always called with a zeroed `scan_control` structure. The reclaim state is also guaranteed to be reset to NULL on exit, maintaining strict scope.

Now, concerning KMSAN (Kernel Memory Sanitizer) and potential vulnerabilities, this patch series introduces `SWAP_FLAG_OFFLOAD_ONLY` to segregate swap devices for proactive reclaim, essentially changing swap allocation/reclaim logic and adding some tracking. We're talking about modifications to reclaim routines, priority calculations, and the introduction of new atomic counters. But does this open any doors for KMSAN to flag uninitialized memory reads or information leaks to userspace?

The answer, after careful consideration, is no. We're not exposing any data structures or padding bytes directly to userspace. All new struct members, namely `allow_offload_swap` across `reclaim_state`, `scan_control`, and `percpu_swap_cluster`, are guaranteed to be initialized, either through designated initializers (hence zeroed before use) or BSS. The new atomic variable is zero-initialized too. The swapon flag is parsed from syscall arguments. The code is clean.

Any problems introduced by this patch series would be standard functional bugs, like incorrect reclaim routing, priority calculation, accounting, or locking/concurrency problems. These are the kinds of issues that standard functional testing, KASAN (Kernel Address Sanitizer), LOCKDEP (Lock Dependency Detector), and existing debuggers would pick up, no doubt. These checks are already in place and well-suited to the task.

In conclusion, there's no risk of uninitialized memory reads or information leaks to userspace due to this patch. KMSAN fuzzing isn't needed here. The changes are logically sound, and any potential issues are well-covered by existing tools. Therefore, `NeedsKMSAN` is definitely `false`. Time to set the results.



70/3 2026/09/14 12:45 tool set-results
0m Args:
{
  "KMSANReasoning": "The patch series introduces offload-only swap areas (SWAP_FLAG_OFFLOAD_ONLY / SWP_OFFLOAD_ONLY), allowing swap devices to be designated solely for proactive reclaim (e.g. cgroup v2 memory.reclaim).\n\nKMSAN vs KASAN Applicability Analysis:\n- Uninitialized Memory Risks (KMSAN):\n  1. No new structures or buffers are copied to user space; user-space interfaces modified are the input flag to the `swapon` syscall and an integer event counter reported via `/proc/vmstat`.\n  2. The new fields added to kernel structures (`allow_offload_swap` in `struct reclaim_state`, `struct scan_control`, and `struct percpu_swap_cluster`) are always initialized before use: `scan_control` instances across vmscan are zero-initialized via C99 designated initializers, `percpu_swap_cluster` is defined in static per-CPU storage, and `current-\u003ereclaim_state` is strictly scoped within `set_task_reclaim_state()`.\n  3. No branching or conditional logic depends on uninitialized data.\n\n- Memory Corruption / Concurrency Risks (KASAN / LOCKDEP):\n  1. Any potential risks introduced by the patch involve swap slot accounting, cluster selection logic, priority routing, or device-level locking/synchronization.\n  2. These failure modes (use-after-free, double frees, invalid dereferences, locking deadlocks) are already covered by standard KASAN and LOCKDEP configurations.\n\nBecause there are no uninitialized memory use or kernel info-leak risks introduced by this patch series, spawning a dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

Results:
{
  "KMSANReasoning": "The patch series introduces offload-only swap areas (SWAP_FLAG_OFFLOAD_ONLY / SWP_OFFLOAD_ONLY), allowing swap devices to be designated solely for proactive reclaim (e.g. cgroup v2 memory.reclaim).\n\nKMSAN vs KASAN Applicability Analysis:\n- Uninitialized Memory Risks (KMSAN):\n  1. No new structures or buffers are copied to user space; user-space interfaces modified are the input flag to the `swapon` syscall and an integer event counter reported via `/proc/vmstat`.\n  2. The new fields added to kernel structures (`allow_offload_swap` in `struct reclaim_state`, `struct scan_control`, and `struct percpu_swap_cluster`) are always initialized before use: `scan_control` instances across vmscan are zero-initialized via C99 designated initializers, `percpu_swap_cluster` is defined in static per-CPU storage, and `current-\u003ereclaim_state` is strictly scoped within `set_task_reclaim_state()`.\n  3. No branching or conditional logic depends on uninitialized data.\n\n- Memory Corruption / Concurrency Risks (KASAN / LOCKDEP):\n  1. Any potential risks introduced by the patch involve swap slot accounting, cluster selection logic, priority routing, or device-level locking/synchronization.\n  2. These failure modes (use-after-free, double frees, invalid dereferences, locking deadlocks) are already covered by standard KASAN and LOCKDEP configurations.\n\nBecause there are no uninitialized memory use or kernel info-leak risks introduced by this patch series, spawning a dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

LLM Calls Summary:
Total Calls Total Tokens Avg Tokens Total Duration (Seconds) Avg Duration (Seconds)
Tool Calls Summary:
Total Calls Total Duration (Seconds) Avg Duration (Seconds)