| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/10/09 22:39 | flow | patch-triage |
4mArgs:null Results: {
"EnableConfigs": [
"NUMA",
"LRU_GEN",
"MEMORY_HOTPLUG"
],
"FocusSymbols": [
"max_kswapds_per_node_sysctl_handler",
"update_max_kswapds_per_node",
"wakeup_kswapd",
"kswapd_idle_avg_sample_workfn",
"kswapd_run",
"kswapd_stop"
],
"KMSANReasoning": "The patch implements support for multiple kswapd threads per NUMA node, an idle CPU rolling average mechanism to scale thread wakeup, and a sysctl interface (`max_kswapds_per_node`). \n\nA dedicated KMSAN session is not needed for the following reasons:\n1. No uninitialized memory exposure to userspace: The sysctl handler uses standard `proc_dointvec_minmax` with a fully initialized integer, and there are no info-leak vectors or uninitialized structures copied to user space.\n2. Safe data structure allocations: The NUMA idle-tracking arrays (`kswapd_idle_avg`, `kswapd_prev_idle`, `kswapd_prev_idle_valid`) are allocated using `kcalloc`, ensuring all fields and elements are zero-initialized. The `kswapd` thread array in `pglist_data` is part of node data zeroed at boot/hotplug.\n3. Deterministic initialization of local variables: All local variables (including `running` in `kswapd_is_running` which is unconditionally assigned in the first iteration of the fixed 16-iteration loop) are assigned before being read.\n4. Bug profile: The risks introduced by this change relate to concurrency (multiple kswapd threads contending on reclaim paths), task lifecycle/locking (`kthread_run`/`kthread_stop`), and bounds checks on sysctl parameters—issues that are covered by KASAN, LOCKDEP, and standard kernel debug facilities, rather than KMSAN uninitialized memory tracking.",
"NeedsKMSAN": false,
"Reasoning": "The patch adds multi-kswapd thread support per NUMA node, dynamic scaling of kswapd threads via the /proc/sys/vm/max_kswapds_per_node sysctl, per-node rolling average idle CPU sampling via delayed work and CPU hotplug hooks, altered wait queue wakeups (prepare_to_wait_exclusive and wake_up_nr), and changes to Multi-Gen LRU scanning. These modifications are in core memory management (mm/vmscan.c, mm/page_alloc.c, mm/compaction.c) and are reachable in virtualized environments via memory allocation/pressure and sysctl writes. Concurrent multi-thread reclaim and dynamic thread stopping/starting introduce potential race conditions and locking subtleties that warrant fuzzing.",
"WorthFuzzing": true
} |
| 1/1 | 2026/10/09 22:39 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit 3135ebe1de131a892c48dfa2c20fa9c02d6f74fc\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Fri Oct 9 22:39:58 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/include/linux/mmzone.h b/include/linux/mmzone.h\nindex 94f9c3ff54160..3f0ae2b08835c 100644\n--- a/include/linux/mmzone.h\n+++ b/include/linux/mmzone.h\n@@ -62,6 +62,8 @@\n */\n #define PAGE_ALLOC_COSTLY_ORDER 3\n \n+#define MAX_KSWAPD_THREADS 16\n+\n #if !defined(CONFIG_HAVE_GIGANTIC_FOLIOS)\n /*\n * We don't expect any folios that exceed buddy sizes (and consequently\n@@ -1529,7 +1531,8 @@ typedef struct pglist_data {\n #ifdef CONFIG_MEMORY_HOTPLUG\n \tstruct mutex kswapd_lock;\n #endif\n-\tstruct task_struct *kswapd;\t/* Protected by kswapd_lock */\n+\t/* Protected by kswapd_lock */\n+\tstruct task_struct *kswapd[MAX_KSWAPD_THREADS];\n \tint kswapd_order;\n \tenum zone_type kswapd_highest_zoneidx;\n \ndiff --git a/include/trace/events/vmscan.h b/include/trace/events/vmscan.h\nindex b4bf7b8def1f5..6fd87d811e984 100644\n--- a/include/trace/events/vmscan.h\n+++ b/include/trace/events/vmscan.h\n@@ -174,6 +174,38 @@ TRACE_EVENT(mm_vmscan_wakeup_kswapd,\n \t\tshow_gfp_flags(__entry-\u003egfp_flags))\n );\n \n+TRACE_EVENT(mm_vmscan_kswapd_threads_to_wake,\n+\n+\tTP_PROTO(int nid, int threads_to_wake, int threads_woken,\n+\t\t int node_cpus, int avg_idle_cores),\n+\n+\tTP_ARGS(nid, threads_to_wake, threads_woken, node_cpus,\n+\t\tavg_idle_cores),\n+\n+\tTP_STRUCT__entry(\n+\t\t__field(int,\t nid)\n+\t\t__field(int,\t threads_to_wake)\n+\t\t__field(int,\t threads_woken)\n+\t\t__field(int,\t node_cpus)\n+\t\t__field(int,\t avg_idle_cores)\n+\t),\n+\n+\tTP_fast_assign(\n+\t\t__entry-\u003enid\t\t\t= nid;\n+\t\t__entry-\u003ethreads_to_wake\t= threads_to_wake;\n+\t\t__entry-\u003ethreads_woken\t\t= threads_woken;\n+\t\t__entry-\u003enode_cpus\t\t= node_cpus;\n+\t\t__entry-\u003eavg_idle_cores\t\t= avg_idle_cores;\n+\t),\n+\n+\tTP_printk(\"nid=%d threads_to_wake=%d threads_woken=%d node_cpus=%d avg_idle_cores=%d\",\n+\t\t __entry-\u003enid,\n+\t\t __entry-\u003ethreads_to_wake,\n+\t\t __entry-\u003ethreads_woken,\n+\t\t __entry-\u003enode_cpus,\n+\t\t __entry-\u003eavg_idle_cores)\n+);\n+\n DECLARE_EVENT_CLASS(mm_vmscan_direct_reclaim_begin_template,\n \n \tTP_PROTO(gfp_t gfp_flags, int order, struct mem_cgroup *memcg),\ndiff --git a/mm/compaction.c b/mm/compaction.c\nindex a049415512c67..5db7d94aefc23 100644\n--- a/mm/compaction.c\n+++ b/mm/compaction.c\n@@ -2198,9 +2198,13 @@ static isolate_migrate_t isolate_migratepages(struct compact_control *cc)\n static bool kswapd_is_running(pg_data_t *pgdat)\n {\n \tbool running;\n-\n+\tint i;\n \tpgdat_kswapd_lock(pgdat);\n-\trunning = pgdat-\u003ekswapd \u0026\u0026 task_is_running(pgdat-\u003ekswapd);\n+\tfor (i = 0; i \u003c MAX_KSWAPD_THREADS; ++i) {\n+\t\trunning = pgdat-\u003ekswapd[i] \u0026\u0026 task_is_running(pgdat-\u003ekswapd[i]);\n+\t\tif (running)\n+\t\t\tbreak;\n+\t}\n \tpgdat_kswapd_unlock(pgdat);\n \n \treturn running;\ndiff --git a/mm/internal.h b/mm/internal.h\nindex 38b1165212c94..1417c5b718ff2 100644\n--- a/mm/internal.h\n+++ b/mm/internal.h\n@@ -717,10 +717,13 @@ void set_recommended_min_free_kbytes(void);\n \n extern char * const zone_names[MAX_NR_ZONES];\n \n+extern int max_kswapds_per_node;\n+extern int max_kswapd_threads;\n extern int min_free_kbytes;\n extern int defrag_mode;\n \n void setup_per_zone_wmarks(void);\n+void update_max_kswapds_per_node(void);\n void calculate_min_free_kbytes(void);\n int __meminit init_per_zone_wmark_min(void);\n \ndiff --git a/mm/page_alloc.c b/mm/page_alloc.c\nindex 12fac9084c483..d56273f8c7de5 100644\n--- a/mm/page_alloc.c\n+++ b/mm/page_alloc.c\n@@ -277,6 +277,7 @@ int user_min_free_kbytes = -1;\n static int watermark_boost_factor __read_mostly = 15000;\n static int watermark_scale_factor = 10;\n int defrag_mode;\n+int max_kswapd_threads = MAX_KSWAPD_THREADS;\n \n /* movable_zone is the \"real\" zone pages in ZONE_MOVABLE are taken from */\n int movable_zone;\n@@ -6777,6 +6778,22 @@ static int min_free_kbytes_sysctl_handler(const struct ctl_table *table, int wri\n \treturn 0;\n }\n \n+static int max_kswapds_per_node_sysctl_handler(const struct ctl_table *table,\n+\t\t\t\t\t int write, void __user *buffer,\n+\t\t\t\t\t size_t *length, loff_t *ppos)\n+{\n+\tint rc;\n+\n+\trc = proc_dointvec_minmax(table, write, buffer, length, ppos);\n+\tif (rc)\n+\t\treturn rc;\n+\n+\tif (write)\n+\t\tupdate_max_kswapds_per_node();\n+\n+\treturn 0;\n+}\n+\n static int watermark_scale_factor_sysctl_handler(const struct ctl_table *table, int write,\n \t\tvoid *buffer, size_t *length, loff_t *ppos)\n {\n@@ -6950,6 +6967,15 @@ static const struct ctl_table page_alloc_sysctl_table[] = {\n \t\t.proc_handler\t= proc_dointvec_minmax,\n \t\t.extra1\t\t= SYSCTL_ZERO,\n \t},\n+\t{\n+\t\t.procname = \"max_kswapds_per_node\",\n+\t\t.data = \u0026max_kswapds_per_node,\n+\t\t.maxlen = sizeof(max_kswapds_per_node),\n+\t\t.mode = 0644,\n+\t\t.proc_handler = max_kswapds_per_node_sysctl_handler,\n+\t\t.extra1 = SYSCTL_ONE,\n+\t\t.extra2 = \u0026max_kswapd_threads,\n+\t},\n \t{\n \t\t.procname\t= \"watermark_scale_factor\",\n \t\t.data\t\t= \u0026watermark_scale_factor,\ndiff --git a/mm/vmscan.c b/mm/vmscan.c\nindex f11491ee9ed5c..3d7a1a56b16a1 100644\n--- a/mm/vmscan.c\n+++ b/mm/vmscan.c\n@@ -32,6 +32,7 @@\n #include \u003clinux/rmap.h\u003e\n #include \u003clinux/topology.h\u003e\n #include \u003clinux/cpu.h\u003e\n+#include \u003clinux/cpuhotplug.h\u003e\n #include \u003clinux/cpuset.h\u003e\n #include \u003clinux/compaction.h\u003e\n #include \u003clinux/notifier.h\u003e\n@@ -59,12 +60,14 @@\n #include \u003clinux/mmu_notifier.h\u003e\n #include \u003clinux/parser.h\u003e\n #include \u003clinux/swap_ops.h\u003e\n+#include \u003clinux/smp.h\u003e\n \n #include \u003casm/tlbflush.h\u003e\n #include \u003casm/div64.h\u003e\n \n #include \u003clinux/swapops.h\u003e\n #include \u003clinux/sched/sysctl.h\u003e\n+#include \u003clinux/slab.h\u003e\n \n #include \"internal.h\"\n #include \"page_alloc.h\"\n@@ -181,6 +184,285 @@ struct scan_control {\n \tstruct reclaim_state reclaim_state;\n };\n \n+/*\n+ * Max and current number of kswapd threads per node.\n+ */\n+#define DEF_MAX_KSWAPDS_PER_NODE 1\n+int max_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;\n+int current_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;\n+\n+#ifdef CONFIG_NUMA\n+/*\n+ * Per-node 10-second rolling average of idle CPU cores.\n+ * Sampled once per second; used by wakeup_kswapd() to decide how many\n+ * kswapd threads to wake without adding contention on a loaded node.\n+ */\n+#define KSWAPD_NUMA_AVG_WINDOW\t10\n+\n+struct kswapd_node_idle_avg {\n+\tu32\tsamples[KSWAPD_NUMA_AVG_WINDOW];\n+\tu32\tsum;\n+\tu8\tidx;\n+\tu8\tcount;\n+\tu32\tavg_idle_cores;\t/* rolling average, written only by the work item */\n+\tint\tnid;\n+\tint\tnode_cpus;\t/* updated by cpuhp callbacks */\n+\tbool\tstarted;\n+\tu64\tlast_sample_ns;\n+\tstruct delayed_work work;\n+};\n+\n+static struct kswapd_node_idle_avg *kswapd_idle_avg;\n+static u64 *kswapd_prev_idle;\n+static bool *kswapd_prev_idle_valid;\n+\n+static int kswapd_node_first_online_cpu(int nid)\n+{\n+\treturn cpumask_first_and(cpumask_of_node(nid), cpu_online_mask);\n+}\n+\n+static u64 kswapd_read_idle_cpu(int cpu)\n+{\n+\tif (!cpu_online(cpu))\n+\t\treturn 0;\n+\n+\treturn kcpustat_field(CPUTIME_IDLE, cpu);\n+}\n+\n+static void kswapd_idle_avg_queue_work(struct kswapd_node_idle_avg *avg)\n+{\n+\tint cpu;\n+\n+\tcpu = kswapd_node_first_online_cpu(avg-\u003enid);\n+\tif (cpu \u003c nr_cpu_ids)\n+\t\tqueue_delayed_work_on(cpu, system_wq, \u0026avg-\u003ework, HZ);\n+}\n+\n+static void kswapd_idle_avg_sample_workfn(struct work_struct *work)\n+{\n+\tstruct kswapd_node_idle_avg *avg;\n+\tu64 now_ns, elapsed_ns;\n+\tu64 idle_cores_milli = 0;\n+\tu32 idle_cores;\n+\tint cpu, node_cpus;\n+\n+\tif (!kswapd_idle_avg || !kswapd_prev_idle || !kswapd_prev_idle_valid)\n+\t\treturn;\n+\n+\tavg = container_of(to_delayed_work(work), struct kswapd_node_idle_avg,\n+\t\t\t work);\n+\tnode_cpus = 0;\n+\tnow_ns = ktime_get_ns();\n+\n+\tif (!avg-\u003elast_sample_ns) {\n+\t\tavg-\u003elast_sample_ns = now_ns;\n+\t\tfor_each_cpu(cpu, cpumask_of_node(avg-\u003enid)) {\n+\t\t\tif (!cpu_online(cpu))\n+\t\t\t\tcontinue;\n+\t\t\tkswapd_prev_idle[cpu] = kswapd_read_idle_cpu(cpu);\n+\t\t\tkswapd_prev_idle_valid[cpu] = true;\n+\t\t}\n+\t\tkswapd_idle_avg_queue_work(avg);\n+\t\treturn;\n+\t}\n+\n+\telapsed_ns = now_ns - avg-\u003elast_sample_ns;\n+\tavg-\u003elast_sample_ns = now_ns;\n+\n+\tif (!elapsed_ns) {\n+\t\tkswapd_idle_avg_queue_work(avg);\n+\t\treturn;\n+\t}\n+\n+\tfor_each_cpu(cpu, cpumask_of_node(avg-\u003enid)) {\n+\t\tu64 idle_now, idle_delta;\n+\n+\t\tif (!cpu_online(cpu))\n+\t\t\tcontinue;\n+\n+\t\tnode_cpus++;\n+\t\tidle_now = kswapd_read_idle_cpu(cpu);\n+\n+\t\tif (!kswapd_prev_idle_valid[cpu]) {\n+\t\t\tkswapd_prev_idle_valid[cpu] = true;\n+\t\t\tkswapd_prev_idle[cpu] = idle_now;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tidle_delta = idle_now - kswapd_prev_idle[cpu];\n+\t\tkswapd_prev_idle[cpu] = idle_now;\n+\n+\t\tidle_cores_milli += div_u64(min(idle_delta, elapsed_ns) * 1000,\n+\t\t\t\t\t\t elapsed_ns);\n+\t}\n+\n+\tidle_cores = min_t(u32, DIV_ROUND_CLOSEST_ULL(idle_cores_milli, 1000),\n+\t\t\t node_cpus);\n+\n+\tif (avg-\u003ecount == KSWAPD_NUMA_AVG_WINDOW)\n+\t\tavg-\u003esum -= avg-\u003esamples[avg-\u003eidx];\n+\telse\n+\t\tavg-\u003ecount++;\n+\n+\tavg-\u003esamples[avg-\u003eidx] = idle_cores;\n+\tavg-\u003esum += idle_cores;\n+\tavg-\u003eidx = (avg-\u003eidx + 1) % KSWAPD_NUMA_AVG_WINDOW;\n+\tWRITE_ONCE(avg-\u003eavg_idle_cores, DIV_ROUND_CLOSEST(avg-\u003esum, avg-\u003ecount));\n+\n+\tif (READ_ONCE(avg-\u003estarted))\n+\t\tkswapd_idle_avg_queue_work(avg);\n+}\n+\n+static void kswapd_idle_avg_start_node(int nid)\n+{\n+\tstruct kswapd_node_idle_avg *avg;\n+\n+\tif (!kswapd_idle_avg || nid \u003c 0 || nid \u003e= nr_node_ids)\n+\t\treturn;\n+\n+\tavg = \u0026kswapd_idle_avg[nid];\n+\tif (READ_ONCE(avg-\u003estarted))\n+\t\treturn;\n+\n+\tWRITE_ONCE(avg-\u003estarted, true);\n+\tavg-\u003elast_sample_ns = 0;\n+\tkswapd_idle_avg_queue_work(avg);\n+}\n+\n+static void kswapd_idle_avg_stop_node(int nid)\n+{\n+\tstruct kswapd_node_idle_avg *avg;\n+\tint cpu;\n+\n+\tif (!kswapd_idle_avg || nid \u003c 0 || nid \u003e= nr_node_ids)\n+\t\treturn;\n+\n+\tavg = \u0026kswapd_idle_avg[nid];\n+\tif (!READ_ONCE(avg-\u003estarted))\n+\t\treturn;\n+\n+\tWRITE_ONCE(avg-\u003estarted, false);\n+\tcancel_delayed_work_sync(\u0026avg-\u003ework);\n+\n+\tfor_each_cpu(cpu, cpumask_of_node(nid))\n+\t\tkswapd_prev_idle_valid[cpu] = false;\n+}\n+\n+static int kswapd_idle_avg_cpu_online(unsigned int cpu)\n+{\n+\tif (kswapd_prev_idle_valid)\n+\t\tkswapd_prev_idle_valid[cpu] = false;\n+\n+\tif (kswapd_idle_avg) {\n+\t\tint nid = cpu_to_node(cpu);\n+\n+\t\tWRITE_ONCE(kswapd_idle_avg[nid].node_cpus,\n+\t\t\t cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask));\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int kswapd_idle_avg_cpu_offline(unsigned int cpu)\n+{\n+\tif (kswapd_prev_idle_valid)\n+\t\tkswapd_prev_idle_valid[cpu] = false;\n+\n+\tif (kswapd_idle_avg) {\n+\t\tint nid = cpu_to_node(cpu);\n+\n+\t\tWRITE_ONCE(kswapd_idle_avg[nid].node_cpus,\n+\t\t\t cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask));\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static u32 kswapd_node_avg_idle_cores(int nid)\n+{\n+\tif (!kswapd_idle_avg || nid \u003c 0 || nid \u003e= nr_node_ids)\n+\t\treturn 0;\n+\n+\treturn READ_ONCE(kswapd_idle_avg[nid].avg_idle_cores);\n+}\n+\n+static int kswapd_node_cpu_count(int nid)\n+{\n+\tif (!kswapd_idle_avg || nid \u003c 0 || nid \u003e= nr_node_ids)\n+\t\treturn cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask);\n+\n+\treturn READ_ONCE(kswapd_idle_avg[nid].node_cpus);\n+}\n+\n+static void kswapd_idle_avg_init(void)\n+{\n+\tint cpu;\n+\tint nid;\n+\tint ret;\n+\n+\tkswapd_idle_avg = kcalloc(nr_node_ids, sizeof(*kswapd_idle_avg),\n+\t\t\t\t GFP_KERNEL);\n+\tif (!kswapd_idle_avg)\n+\t\treturn;\n+\n+\tkswapd_prev_idle = kcalloc(nr_cpu_ids, sizeof(*kswapd_prev_idle),\n+\t\t\t\t GFP_KERNEL);\n+\tif (!kswapd_prev_idle)\n+\t\tgoto free_idle_avg;\n+\n+\tkswapd_prev_idle_valid = kcalloc(nr_cpu_ids,\n+\t\t\t\t\t sizeof(*kswapd_prev_idle_valid),\n+\t\t\t\t\t GFP_KERNEL);\n+\tif (!kswapd_prev_idle_valid)\n+\t\tgoto free_prev_idle;\n+\n+\tfor (nid = 0; nid \u003c nr_node_ids; nid++) {\n+\t\tstruct kswapd_node_idle_avg *avg = \u0026kswapd_idle_avg[nid];\n+\n+\t\tavg-\u003enid = nid;\n+\t\tINIT_DELAYED_WORK(\u0026avg-\u003ework, kswapd_idle_avg_sample_workfn);\n+\t}\n+\n+\tret = cpuhp_setup_state_nocalls(CPUHP_AP_ONLINE_DYN,\n+\t\t\t\t\t\"mm/vmscan:online\",\n+\t\t\t\t\tkswapd_idle_avg_cpu_online,\n+\t\t\t\t\tkswapd_idle_avg_cpu_offline);\n+\tif (ret \u003c 0)\n+\t\tgoto free_prev_idle_valid;\n+\n+\tcpus_read_lock();\n+\tfor_each_online_cpu(cpu)\n+\t\tkswapd_idle_avg_cpu_online(cpu);\n+\tcpus_read_unlock();\n+\n+\treturn;\n+\n+free_prev_idle_valid:\n+\tkfree(kswapd_prev_idle_valid);\n+\tkswapd_prev_idle_valid = NULL;\n+free_prev_idle:\n+\tkfree(kswapd_prev_idle);\n+\tkswapd_prev_idle = NULL;\n+free_idle_avg:\n+\tkfree(kswapd_idle_avg);\n+\tkswapd_idle_avg = NULL;\n+}\n+#else\n+static inline u32 kswapd_node_avg_idle_cores(int nid)\n+{\n+\treturn 0;\n+}\n+\n+static inline int kswapd_node_cpu_count(int nid)\n+{\n+\treturn cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask);\n+}\n+\n+static inline void kswapd_idle_avg_init(void) {}\n+static inline void kswapd_idle_avg_start_node(int nid) {}\n+static inline void kswapd_idle_avg_stop_node(int nid) {}\n+#endif /* CONFIG_NUMA */\n+\n #ifdef ARCH_HAS_PREFETCHW\n #define prefetchw_prev_lru_folio(_folio, _base, _field)\t\t\t\\\n \tdo {\t\t\t\t\t\t\t\t\\\n@@ -4113,10 +4395,8 @@ static bool try_to_inc_max_seq(struct lruvec *lruvec, unsigned long seq,\n \t\t\twalk_mm(mm, walk);\n \t} while (mm);\n done:\n-\tif (success) {\n+\tif (success)\n \t\tsuccess = inc_max_seq(lruvec, seq, swappiness);\n-\t\tWARN_ON_ONCE(!success);\n-\t}\n \n \treturn success;\n }\n@@ -4741,7 +5021,7 @@ static int scan_folios(unsigned long nr_to_scan, struct lruvec *lruvec,\n \tVM_WARN_ON_ONCE(nr_to_scan \u003e MAX_LRU_BATCH);\n \tVM_WARN_ON_ONCE(!list_empty(list));\n \n-\tif (get_nr_gens(lruvec, type) == MIN_NR_GENS)\n+\tif (!current_is_kswapd() \u0026\u0026 get_nr_gens(lruvec, type) == MIN_NR_GENS)\n \t\treturn 0;\n \n \tgen = lru_gen_from_seq(lrugen-\u003emin_seq[type]);\n@@ -6671,7 +6951,7 @@ static bool allow_direct_reclaim(pg_data_t *pgdat)\n \t\tif (READ_ONCE(pgdat-\u003ekswapd_highest_zoneidx) \u003e ZONE_NORMAL)\n \t\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, ZONE_NORMAL);\n \n-\t\twake_up_interruptible(\u0026pgdat-\u003ekswapd_wait);\n+\t\twake_up_interruptible_all(\u0026pgdat-\u003ekswapd_wait);\n \t}\n \n \treturn wmark_ok;\n@@ -7401,7 +7681,8 @@ static void kswapd_try_to_sleep(pg_data_t *pgdat, int alloc_order, int reclaim_o\n \tif (freezing(current) || kthread_should_stop())\n \t\treturn;\n \n-\tprepare_to_wait(\u0026pgdat-\u003ekswapd_wait, \u0026wait, TASK_INTERRUPTIBLE);\n+\tprepare_to_wait_exclusive(\u0026pgdat-\u003ekswapd_wait, \u0026wait,\n+\t\t\t\t TASK_INTERRUPTIBLE);\n \n \t/*\n \t * Try to sleep for a short interval. Note that kcompactd will only be\n@@ -7442,7 +7723,8 @@ static void kswapd_try_to_sleep(pg_data_t *pgdat, int alloc_order, int reclaim_o\n \t\t}\n \n \t\tfinish_wait(\u0026pgdat-\u003ekswapd_wait, \u0026wait);\n-\t\tprepare_to_wait(\u0026pgdat-\u003ekswapd_wait, \u0026wait, TASK_INTERRUPTIBLE);\n+\t\tprepare_to_wait_exclusive(\u0026pgdat-\u003ekswapd_wait, \u0026wait,\n+\t\t\t\t\t TASK_INTERRUPTIBLE);\n \t}\n \n \t/*\n@@ -7575,6 +7857,8 @@ void wakeup_kswapd(struct zone *zone, gfp_t gfp_flags, int order,\n {\n \tpg_data_t *pgdat;\n \tenum zone_type curr_idx;\n+\tint nid, threads_to_wake, threads_woken;\n+\tint node_cpus, avg_idle_cores;\n \n \tif (!managed_zone(zone))\n \t\treturn;\n@@ -7612,7 +7896,29 @@ void wakeup_kswapd(struct zone *zone, gfp_t gfp_flags, int order,\n \n \ttrace_mm_vmscan_wakeup_kswapd(pgdat-\u003enode_id, highest_zoneidx, order,\n \t\t\t\t gfp_flags);\n-\twake_up_interruptible(\u0026pgdat-\u003ekswapd_wait);\n+\n+\t/* Determine how many kswapd threads to wake based on the 10-second\n+\t * rolling average of idle cores on this NUMA node. More idle cores\n+\t * means more threads can run without adding contention; fewer idle\n+\t * cores means we keep the wakeup conservative.\n+\t */\n+\tnid = pgdat-\u003enode_id;\n+\tavg_idle_cores = kswapd_node_avg_idle_cores(nid);\n+\tnode_cpus = kswapd_node_cpu_count(nid);\n+\n+\t/*\n+\t * If no average is available yet (early boot), fall back to waking\n+\t * a single thread to avoid stalling reclaim.\n+\t */\n+\tthreads_to_wake = max(1,\n+\t\t\t min3(avg_idle_cores, node_cpus,\n+\t\t\t\t READ_ONCE(current_kswapds_per_node)));\n+\n+\tthreads_woken = wake_up_nr(\u0026pgdat-\u003ekswapd_wait, threads_to_wake);\n+\n+\ttrace_mm_vmscan_kswapd_threads_to_wake(nid, threads_to_wake,\n+\t\t\t\t\t threads_woken, node_cpus,\n+\t\t\t\t\t avg_idle_cores);\n }\n \n void kswapd_clear_hopeless(pg_data_t *pgdat, enum kswapd_clear_hopeless_reason reason)\n@@ -7638,7 +7944,8 @@ void kswapd_try_clear_hopeless(struct pglist_data *pgdat,\n \n bool kswapd_test_hopeless(pg_data_t *pgdat)\n {\n-\treturn atomic_read(\u0026pgdat-\u003ekswapd_failures) \u003e= MAX_RECLAIM_RETRIES;\n+\treturn atomic_read(\u0026pgdat-\u003ekswapd_failures) \u003e=\n+\t\tMAX_RECLAIM_RETRIES * READ_ONCE(current_kswapds_per_node);\n }\n \n #ifdef CONFIG_HIBERNATION\n@@ -7680,27 +7987,93 @@ unsigned long shrink_all_memory(unsigned long nr_to_reclaim)\n }\n #endif /* CONFIG_HIBERNATION */\n \n+static void update_kswapds_per_node_node(int nid)\n+{\n+\tpg_data_t *pgdat;\n+\tint drop, increase;\n+\tint last_idx, start_idx, hid;\n+\tint nr_threads = current_kswapds_per_node;\n+\n+\tpgdat = NODE_DATA(nid);\n+\tpgdat_kswapd_lock(pgdat);\n+\tlast_idx = nr_threads - 1;\n+\tif (max_kswapds_per_node \u003c nr_threads) {\n+\t\tdrop = nr_threads - max_kswapds_per_node;\n+\t\tfor (hid = last_idx; hid \u003e (last_idx - drop); hid--) {\n+\t\t\tif (pgdat-\u003ekswapd[hid]) {\n+\t\t\t\tkthread_stop(pgdat-\u003ekswapd[hid]);\n+\t\t\t\tpgdat-\u003ekswapd[hid] = NULL;\n+\t\t\t}\n+\t\t}\n+\t} else {\n+\t\tincrease = max_kswapds_per_node - nr_threads;\n+\t\tstart_idx = last_idx + 1;\n+\t\tfor (hid = start_idx; hid \u003c (start_idx + increase); hid++) {\n+\t\t\tpgdat-\u003ekswapd[hid] = kthread_run(kswapd, pgdat, \"kswapd%d:%d\",\n+\t\t\t\t\t\t\t nid, hid);\n+\t\t\tif (IS_ERR(pgdat-\u003ekswapd[hid])) {\n+\t\t\t\tpr_err(\"Failed to start kswapd%d on node %d\\n\", hid, nid);\n+\t\t\t\tpgdat-\u003ekswapd[hid] = NULL;\n+\t\t\t\t/*\n+\t\t\t\t * We are out of resources. Do not start any\n+\t\t\t\t * more threads.\n+\t\t\t\t */\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t}\n+\t}\n+\tpgdat_kswapd_unlock(pgdat);\n+}\n+\n+void update_max_kswapds_per_node(void)\n+{\n+\tint nid;\n+\n+\tif (current_kswapds_per_node == max_kswapds_per_node)\n+\t\treturn;\n+\n+\t/*\n+\t * Hold the memory hotplug lock to avoid racing with memory\n+\t * hotplug initiated updates\n+\t */\n+\tmem_hotplug_begin();\n+\tfor_each_node_state(nid, N_MEMORY)\n+\t\tupdate_kswapds_per_node_node(nid);\n+\n+\tpr_info(\"max_kswapds_per_node changed, old:%d new:%d\\n\",\n+\t\tcurrent_kswapds_per_node, max_kswapds_per_node);\n+\tcurrent_kswapds_per_node = max_kswapds_per_node;\n+\tmem_hotplug_done();\n+}\n+\n /*\n * This kswapd start function will be called by init and node-hot-add.\n */\n void __meminit kswapd_run(int nid)\n {\n \tpg_data_t *pgdat = NODE_DATA(nid);\n+\tint hid, nr_threads;\n \n \tpgdat_kswapd_lock(pgdat);\n-\tif (!pgdat-\u003ekswapd) {\n-\t\tpgdat-\u003ekswapd = kthread_create_on_node(kswapd, pgdat, nid, \"kswapd%d\", nid);\n-\t\tif (IS_ERR(pgdat-\u003ekswapd)) {\n-\t\t\t/* failure at boot is fatal */\n-\t\t\tpr_err(\"Failed to start kswapd on node %d, ret=%pe\\n\",\n-\t\t\t\t nid, pgdat-\u003ekswapd);\n-\t\t\tBUG_ON(system_state \u003c SYSTEM_RUNNING);\n-\t\t\tpgdat-\u003ekswapd = NULL;\n-\t\t} else {\n-\t\t\twake_up_process(pgdat-\u003ekswapd);\n+\tnr_threads = max_kswapds_per_node;\n+\tfor (hid = 0; hid \u003c nr_threads; hid++) {\n+\t\tif (!pgdat-\u003ekswapd[hid]) {\n+\t\t\tpgdat-\u003ekswapd[hid] =\n+\t\t\t\tkthread_create_on_node(kswapd, pgdat, nid,\n+\t\t\t\t\t\t \"kswapd%d:%d\", nid, hid);\n+\t\t\tif (IS_ERR(pgdat-\u003ekswapd[hid])) {\n+\t\t\t\t/* failure at boot is fatal */\n+\t\t\t\tpr_err(\"Failed to start kswapd%d on node %d, ret=%pe\\n\", hid,\n+\t\t\t\t nid, pgdat-\u003ekswapd[hid]);\n+\t\t\t\tpgdat-\u003ekswapd[hid] = NULL;\n+\t\t\t} else {\n+\t\t\t\twake_up_process(pgdat-\u003ekswapd[hid]);\n+\t\t\t}\n \t\t}\n \t}\n+\tcurrent_kswapds_per_node = nr_threads;\n \tpgdat_kswapd_unlock(pgdat);\n+\tkswapd_idle_avg_start_node(nid);\n }\n \n /*\n@@ -7710,13 +8083,17 @@ void __meminit kswapd_run(int nid)\n void __meminit kswapd_stop(int nid)\n {\n \tpg_data_t *pgdat = NODE_DATA(nid);\n-\tstruct task_struct *kswapd;\n+\tint hid;\n+\tint nr_threads = current_kswapds_per_node;\n+\n+\tkswapd_idle_avg_stop_node(nid);\n \n \tpgdat_kswapd_lock(pgdat);\n-\tkswapd = pgdat-\u003ekswapd;\n-\tif (kswapd) {\n-\t\tkthread_stop(kswapd);\n-\t\tpgdat-\u003ekswapd = NULL;\n+\tfor (hid = 0; hid \u003c nr_threads; hid++) {\n+\t\tif (pgdat-\u003ekswapd[hid]) {\n+\t\t\tkthread_stop(pgdat-\u003ekswapd[hid]);\n+\t\t\tpgdat-\u003ekswapd[hid] = NULL;\n+\t\t}\n \t}\n \tpgdat_kswapd_unlock(pgdat);\n }\n@@ -7747,6 +8124,7 @@ static int __init kswapd_init(void)\n {\n \tint nid;\n \n+\tkswapd_idle_avg_init();\n \tfor_each_node_state(nid, N_MEMORY)\n \t\tkswapd_run(nid);\n \tregister_sysctl_init(\"vm\", vmscan_sysctl_table);\n"
} |
| 2/1 | 2026/10/09 22:39 | agent | patch-evaluator |
3mModel:core Args: null Results: {
"EnableConfigs": [
"NUMA",
"LRU_GEN",
"MEMORY_HOTPLUG"
],
"FocusSymbols": [
"max_kswapds_per_node_sysctl_handler",
"update_max_kswapds_per_node",
"wakeup_kswapd",
"kswapd_idle_avg_sample_workfn",
"kswapd_run",
"kswapd_stop"
],
"Reasoning": "The patch adds multi-kswapd thread support per NUMA node, dynamic scaling of kswapd threads via the /proc/sys/vm/max_kswapds_per_node sysctl, per-node rolling average idle CPU sampling via delayed work and CPU hotplug hooks, altered wait queue wakeups (prepare_to_wait_exclusive and wake_up_nr), and changes to Multi-Gen LRU scanning. These modifications are in core memory management (mm/vmscan.c, mm/page_alloc.c, mm/compaction.c) and are reachable in virtualized environments via memory allocation/pressure and sysctl writes. Concurrent multi-thread reclaim and dynamic thread stopping/starting introduce potential race conditions and locking subtleties that warrant fuzzing.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 3135ebe1de131a892c48dfa2c20fa9c02d6f74fc
Author: syz-cluster <triage@syzkaller.com>
Date: Fri Oct 9 22:39:58 2026 +0000
syz-cluster: applied patch under review
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index 94f9c3ff54160..3f0ae2b08835c 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -62,6 +62,8 @@
*/
#define PAGE_ALLOC_COSTLY_ORDER 3
+#define MAX_KSWAPD_THREADS 16
+
#if !defined(CONFIG_HAVE_GIGANTIC_FOLIOS)
/*
* We don't expect any folios that exceed buddy sizes (and consequently
@@ -1529,7 +1531,8 @@ typedef struct pglist_data {
#ifdef CONFIG_MEMORY_HOTPLUG
struct mutex kswapd_lock;
#endif
- struct task_struct *kswapd; /* Protected by kswapd_lock */
+ /* Protected by kswapd_lock */
+ struct task_struct *kswapd[MAX_KSWAPD_THREADS];
int kswapd_order;
enum zone_type kswapd_highest_zoneidx;
diff --git a/include/trace/events/vmscan.h b/include/trace/events/vmscan.h
index b4bf7b8def1f5..6fd87d811e984 100644
--- a/include/trace/events/vmscan.h
+++ b/include/trace/events/vmscan.h
@@ -174,6 +174,38 @@ TRACE_EVENT(mm_vmscan_wakeup_kswapd,
show_gfp_flags(__entry->gfp_flags))
);
+TRACE_EVENT(mm_vmscan_kswapd_threads_to_wake,
+
+ TP_PROTO(int nid, int threads_to_wake, int threads_woken,
+ int node_cpus, int avg_idle_cores),
+
+ TP_ARGS(nid, threads_to_wake, threads_woken, node_cpus,
+ avg_idle_cores),
+
+ TP_STRUCT__entry(
+ __field(int, nid)
+ __field(int, threads_to_wake)
+ __field(int, threads_woken)
+ __field(int, node_cpus)
+ __field(int, avg_idle_cores)
+ ),
+
+ TP_fast_assign(
+ __entry->nid = nid;
+ __entry->threads_to_wake = threads_to_wake;
+ __entry->threads_woken = threads_woken;
+ __entry->node_cpus = node_cpus;
+ __entry->avg_idle_cores = avg_idle_cores;
+ ),
+
+ TP_printk("nid=%d threads_to_wake=%d threads_woken=%d node_cpus=%d avg_idle_cores=%d",
+ __entry->nid,
+ __entry->threads_to_wake,
+ __entry->threads_woken,
+ __entry->node_cpus,
+ __entry->avg_idle_cores)
+);
+
DECLARE_EVENT_CLASS(mm_vmscan_direct_reclaim_begin_template,
TP_PROTO(gfp_t gfp_flags, int order, struct mem_cgroup *memcg),
diff --git a/mm/compaction.c b/mm/compaction.c
index a049415512c67..5db7d94aefc23 100644
--- a/mm/compaction.c
+++ b/mm/compaction.c
@@ -2198,9 +2198,13 @@ static isolate_migrate_t isolate_migratepages(struct compact_control *cc)
static bool kswapd_is_running(pg_data_t *pgdat)
{
bool running;
-
+ int i;
pgdat_kswapd_lock(pgdat);
- running = pgdat->kswapd && task_is_running(pgdat->kswapd);
+ for (i = 0; i < MAX_KSWAPD_THREADS; ++i) {
+ running = pgdat->kswapd[i] && task_is_running(pgdat->kswapd[i]);
+ if (running)
+ break;
+ }
pgdat_kswapd_unlock(pgdat);
return running;
diff --git a/mm/internal.h b/mm/internal.h
index 38b1165212c94..1417c5b718ff2 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -717,10 +717,13 @@ void set_recommended_min_free_kbytes(void);
extern char * const zone_names[MAX_NR_ZONES];
+extern int max_kswapds_per_node;
+extern int max_kswapd_threads;
extern int min_free_kbytes;
extern int defrag_mode;
void setup_per_zone_wmarks(void);
+void update_max_kswapds_per_node(void);
void calculate_min_free_kbytes(void);
int __meminit init_per_zone_wmark_min(void);
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 12fac9084c483..d56273f8c7de5 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -277,6 +277,7 @@ int user_min_free_kbytes = -1;
static int watermark_boost_factor __read_mostly = 15000;
static int watermark_scale_factor = 10;
int defrag_mode;
+int max_kswapd_threads = MAX_KSWAPD_THREADS;
/* movable_zone is the "real" zone pages in ZONE_MOVABLE are taken from */
int movable_zone;
@@ -6777,6 +6778,22 @@ static int min_free_kbytes_sysctl_handler(const struct ctl_table *table, int wri
return 0;
}
+static int max_kswapds_per_node_sysctl_handler(const struct ctl_table *table,
+ int write, void __user *buffer,
+ size_t *length, loff_t *ppos)
+{
+ int rc;
+
+ rc = proc_dointvec_minmax(table, write, buffer, length, ppos);
+ if (rc)
+ return rc;
+
+ if (write)
+ update_max_kswapds_per_node();
+
+ return 0;
+}
+
static int watermark_scale_factor_sysctl_handler(const struct ctl_table *table, int write,
void *buffer, size_t *length, loff_t *ppos)
{
@@ -6950,6 +6967,15 @@ static const struct ctl_table page_alloc_sysctl_table[] = {
.proc_handler = proc_dointvec_minmax,
.extra1 = SYSCTL_ZERO,
},
+ {
+ .procname = "max_kswapds_per_node",
+ .data = &max_kswapds_per_node,
+ .maxlen = sizeof(max_kswapds_per_node),
+ .mode = 0644,
+ .proc_handler = max_kswapds_per_node_sysctl_handler,
+ .extra1 = SYSCTL_ONE,
+ .extra2 = &max_kswapd_threads,
+ },
{
.procname = "watermark_scale_factor",
.data = &watermark_scale_factor,
diff --git a/mm/vmscan.c b/mm/vmscan.c
index f11491ee9ed5c..3d7a1a56b16a1 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -32,6 +32,7 @@
#include <linux/rmap.h>
#include <linux/topology.h>
#include <linux/cpu.h>
+#include <linux/cpuhotplug.h>
#include <linux/cpuset.h>
#include <linux/compaction.h>
#include <linux/notifier.h>
@@ -59,12 +60,14 @@
#include <linux/mmu_notifier.h>
#include <linux/parser.h>
#include <linux/swap_ops.h>
+#include <linux/smp.h>
#include <asm/tlbflush.h>
#include <asm/div64.h>
#include <linux/swapops.h>
#include <linux/sched/sysctl.h>
+#include <linux/slab.h>
#include "internal.h"
#include "page_alloc.h"
@@ -181,6 +184,285 @@ struct scan_control {
struct reclaim_state reclaim_state;
};
+/*
+ * Max and current number of kswapd threads per node.
+ */
+#define DEF_MAX_KSWAPDS_PER_NODE 1
+int max_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;
+int current_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;
+
+#ifdef CONFIG_NUMA
+/*
+ * Per-node 10-second rolling average of idle CPU cores.
+ * Sampled once per second; used by wakeup_kswapd() to decide how many
+ * kswapd threads to wake without adding contention on a loaded node.
+ */
+#define KSWAPD_NUMA_AVG_WINDOW 10
+
+struct kswapd_node_idle_avg {
+ u32 samples[KSWAPD_NUMA_AVG_WINDOW];
+ u32 sum;
+ u8 idx;
+ u8 count;
+ u32 avg_idle_cores; /* rolling average, written only by the work item */
+ int nid;
+ int node_cpus; /* updated by cpuhp callbacks */
+ bool started;
+ u64 last_sample_ns;
+ struct delayed_work work;
+};
+
+static struct kswapd_node_idle_avg *kswapd_idle_avg;
+static u64 *kswapd_prev_idle;
+static bool *kswapd_prev_idle_valid;
+
+static int kswapd_node_first_online_cpu(int nid)
+{
+ return cpumask_first_and(cpumask_of_node(nid), cpu_online_mask);
+}
+
+static u64 kswapd_read_idle_cpu(int cpu)
+{
+ if (!cpu_online(cpu))
+ return 0;
+
+ return kcpustat_field(CPUTIME_IDLE, cpu);
+}
+
+static void kswapd_idle_avg_queue_work(struct kswapd_node_idle_avg *avg)
+{
+ int cpu;
+
+ cpu = kswapd_node_first_online_cpu(avg->nid);
+ if (cpu < nr_cpu_ids)
+ queue_delayed_work_on(cpu, system_wq, &avg->work, HZ);
+}
+
+static void kswapd_idle_avg_sample_workfn(struct work_struct *work)
+{
+ struct kswapd_node_idle_avg *avg;
+ u64 now_ns, elapsed_ns;
+ u64 idle_cores_milli = 0;
+ u32 idle_cores;
+ int cpu, node_cpus;
+
+ if (!kswapd_idle_avg || !kswapd_prev_idle || !kswapd_prev_idle_valid)
+ return;
+
+ avg = container_of(to_delayed_work(work), struct kswapd_node_idle_avg,
+ work);
+ node_cpus = 0;
+ now_ns = ktime_get_ns();
+
+ if (!avg->last_sample_ns) {
+ avg->last_sample_ns = now_ns;
+ for_each_cpu(cpu, cpumask_of_node(avg->nid)) {
+ if (!cpu_online(cpu))
+ continue;
+ kswapd_prev_idle[cpu] = kswapd_read_idle_cpu(cpu);
+ kswapd_prev_idle_valid[cpu] = true;
+ }
+ kswapd_idle_avg_queue_work(avg);
+ return;
+ }
+
+ elapsed_ns = now_ns - avg->last_sample_ns;
+ avg->last_sample_ns = now_ns;
+
+ if (!elapsed_ns) {
+ kswapd_idle_avg_queue_work(avg);
+ return;
+ }
+
+ for_each_cpu(cpu, cpumask_of_node(avg->nid)) {
+ u64 idle_now, idle_delta;
+
+ if (!cpu_online(cpu))
+ continue;
+
+ node_cpus++;
+ idle_now = kswapd_read_idle_cpu(cpu);
+
+ if (!kswapd_prev_idle_valid[cpu]) {
+ kswapd_prev_idle_valid[cpu] = true;
+ kswapd_prev_idle[cpu] = idle_now;
+ continue;
+ }
+
+ idle_delta = idle_now - kswapd_prev_idle[cpu];
+ kswapd_prev_idle[cpu] = idle_now;
+
+ idle_cores_milli += div_u64(min(idle_delta, elapsed_ns) * 1000,
+ elapsed_ns);
+ }
+
+ idle_cores = min_t(u32, DIV_ROUND_CLOSEST_ULL(idle_cores_milli, 1000),
+ node_cpus);
+
+ if (avg->count == KSWAPD_NUMA_AVG_WINDOW)
+ avg->sum -= avg->samples[avg->idx];
+ else
+ avg->count++;
+
+ avg->samples[avg->idx] = idle_cores;
+ avg->sum += idle_cores;
+ avg->idx = (avg->idx + 1) % KSWAPD_NUMA_AVG_WINDOW;
+ WRITE_ONCE(avg->avg_idle_cores, DIV_ROUND_CLOSEST(avg->sum, avg->count));
+
+ if (READ_ONCE(avg->started))
+ kswapd_idle_avg_queue_work(avg);
+}
+
+static void kswapd_idle_avg_start_node(int nid)
+{
+ struct kswapd_node_idle_avg *avg;
+
+ if (!kswapd_idle_avg || nid < 0 || nid >= nr_node_ids)
+ return;
+
+ avg = &kswapd_idle_avg[nid];
+ if (READ_ONCE(avg->started))
+ return;
+
+ WRITE_ONCE(avg->started, true);
+ avg->last_sample_ns = 0;
+ kswapd_idle_avg_queue_work(avg);
+}
+
+static void kswapd_idle_avg_stop_node(int nid)
+{
+ struct kswapd_node_idle_avg *avg;
+ int cpu;
+
+ if (!kswapd_idle_avg || nid < 0 || nid >= nr_node_ids)
+ return;
+
+ avg = &kswapd_idle_avg[nid];
+ if (!READ_ONCE(avg->started))
+ return;
+
+ WRITE_ONCE(avg->started, false);
+ cancel_delayed_work_sync(&avg->work);
+
+ for_each_cpu(cpu, cpumask_of_node(nid))
+ kswapd_prev_idle_valid[cpu] = false;
+}
+
+static int kswapd_idle_avg_cpu_online(unsigned int cpu)
+{
+ if (kswapd_prev_idle_valid)
+ kswapd_prev_idle_valid[cpu] = false;
+
+ if (kswapd_idle_avg) {
+ int nid = cpu_to_node(cpu);
+
+ WRITE_ONCE(kswapd_idle_avg[nid].node_cpus,
+ cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask));
+ }
+
+ return 0;
+}
+
+static int kswapd_idle_avg_cpu_offline(unsigned int cpu)
+{
+ if (kswapd_prev_idle_valid)
+ kswapd_prev_idle_valid[cpu] = false;
+
+ if (kswapd_idle_avg) {
+ int nid = cpu_to_node(cpu);
+
+ WRITE_ONCE(kswapd_idle_avg[nid].node_cpus,
+ cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask));
+ }
+
+ return 0;
+}
+
+static u32 kswapd_node_avg_idle_cores(int nid)
+{
+ if (!kswapd_idle_avg || nid < 0 || nid >= nr_node_ids)
+ return 0;
+
+ return READ_ONCE(kswapd_idle_avg[nid].avg_idle_cores);
+}
+
+static int kswapd_node_cpu_count(int nid)
+{
+ if (!kswapd_idle_avg || nid < 0 || nid >= nr_node_ids)
+ return cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask);
+
+ return READ_ONCE(kswapd_idle_avg[nid].node_cpus);
+}
+
+static void kswapd_idle_avg_init(void)
+{
+ int cpu;
+ int nid;
+ int ret;
+
+ kswapd_idle_avg = kcalloc(nr_node_ids, sizeof(*kswapd_idle_avg),
+ GFP_KERNEL);
+ if (!kswapd_idle_avg)
+ return;
+
+ kswapd_prev_idle = kcalloc(nr_cpu_ids, sizeof(*kswapd_prev_idle),
+ GFP_KERNEL);
+ if (!kswapd_prev_idle)
+ goto free_idle_avg;
+
+ kswapd_prev_idle_valid = kcalloc(nr_cpu_ids,
+ sizeof(*kswapd_prev_idle_valid),
+ GFP_KERNEL);
+ if (!kswapd_prev_idle_valid)
+ goto free_prev_idle;
+
+ for (nid = 0; nid < nr_node_ids; nid++) {
+ struct kswapd_node_idle_avg *avg = &kswapd_idle_avg[nid];
+
+ avg->nid = nid;
+ INIT_DELAYED_WORK(&avg->work, kswapd_idle_avg_sample_workfn);
+ }
+
+ ret = cpuhp_setup_state_nocalls(CPUHP_AP_ONLINE_DYN,
+ "mm/vmscan:online",
+ kswapd_idle_avg_cpu_online,
+ kswapd_idle_avg_cpu_offline);
+ if (ret < 0)
+ goto free_prev_idle_valid;
+
+ cpus_read_lock();
+ for_each_online_cpu(cpu)
+ kswapd_idle_avg_cpu_online(cpu);
+ cpus_read_unlock();
+
+ return;
+
+free_prev_idle_valid:
+ kfree(kswapd_prev_idle_valid);
+ kswapd_prev_idle_valid = NULL;
+free_prev_idle:
+ kfree(kswapd_prev_idle);
+ kswapd_prev_idle = NULL;
+free_idle_avg:
+ kfree(kswapd_idle_avg);
+ kswapd_idle_avg = NULL;
+}
+#else
+static inline u32 kswapd_node_avg_idle_cores(int nid)
+{
+ return 0;
+}
+
+static inline int kswapd_node_cpu_count(int nid)
+{
+ return cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask);
+}
+
+static inline void kswapd_idle_avg_init(void) {}
+static inline void kswapd_idle_avg_start_node(int nid) {}
+static inline void kswapd_idle_avg_stop_node(int nid) {}
+#endif /* CONFIG_NUMA */
+
#ifdef ARCH_HAS_PREFETCHW
#define prefetchw_prev_lru_folio(_folio, _base, _field) \
do { \
@@ -4113,10 +4395,8 @@ static bool try_to_inc_max_seq(struct lruvec *lruvec, unsigned long seq,
walk_mm(mm, walk);
} while (mm);
done:
- if (success) {
+ if (success)
success = inc_max_seq(lruvec, seq, swappiness);
- WARN_ON_ONCE(!success);
- }
return success;
}
@@ -4741,7 +5021,7 @@ static int scan_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
VM_WARN_ON_ONCE(nr_to_scan > MAX_LRU_BATCH);
VM_WARN_ON_ONCE(!list_empty(list));
- if (get_nr_gens(lruvec, type) == MIN_NR_GENS)
+ if (!current_is_kswapd() && get_nr_gens(lruvec, type) == MIN_NR_GENS)
return 0;
gen = lru_gen_from_seq(lrugen->min_seq[type]);
@@ -6671,7 +6951,7 @@ static bool allow_direct_reclaim(pg_data_t *pgdat)
if (READ_ONCE(pgdat->kswapd_highest_zoneidx) > ZONE_NORMAL)
WRITE_ONCE(pgdat->kswapd_highest_zoneidx, ZONE_NORMAL);
- wake_up_interruptible(&pgdat->kswapd_wait);
+ wake_up_interruptible_all(&pgdat->kswapd_wait);
}
return wmark_ok;
@@ -7401,7 +7681,8 @@ static void kswapd_try_to_sleep(pg_data_t *pgdat, int alloc_order, int reclaim_o
if (freezing(current) || kthread_should_stop())
return;
- prepare_to_wait(&pgdat->kswapd_wait, &wait, TASK_INTERRUPTIBLE);
+ prepare_to_wait_exclusive(&pgdat->kswapd_wait, &wait,
+ TASK_INTERRUPTIBLE);
/*
* Try to sleep for a short interval. Note that kcompactd will only be
@@ -7442,7 +7723,8 @@ static void kswapd_try_to_sleep(pg_data_t *pgdat, int alloc_order, int reclaim_o
}
finish_wait(&pgdat->kswapd_wait, &wait);
- prepare_to_wait(&pgdat->kswapd_wait, &wait, TASK_INTERRUPTIBLE);
+ prepare_to_wait_exclusive(&pgdat->kswapd_wait, &wait,
+ TASK_INTERRUPTIBLE);
}
/*
@@ -7575,6 +7857,8 @@ void wakeup_kswapd(struct zone *zone, gfp_t gfp_flags, int order,
{
pg_data_t *pgdat;
enum zone_type curr_idx;
+ int nid, threads_to_wake, threads_woken;
+ int node_cpus, avg_idle_cores;
if (!managed_zone(zone))
return;
@@ -7612,7 +7896,29 @@ void wakeup_kswapd(struct zone *zone, gfp_t gfp_flags, int order,
trace_mm_vmscan_wakeup_kswapd(pgdat->node_id, highest_zoneidx, order,
gfp_flags);
- wake_up_interruptible(&pgdat->kswapd_wait);
+
+ /* Determine how many kswapd threads to wake based on the 10-second
+ * rolling average of idle cores on this NUMA node. More idle cores
+ * means more threads can run without adding contention; fewer idle
+ * cores means we keep the wakeup conservative.
+ */
+ nid = pgdat->node_id;
+ avg_idle_cores = kswapd_node_avg_idle_cores(nid);
+ node_cpus = kswapd_node_cpu_count(nid);
+
+ /*
+ * If no average is available yet (early boot), fall back to waking
+ * a single thread to avoid stalling reclaim.
+ */
+ threads_to_wake = max(1,
+ min3(avg_idle_cores, node_cpus,
+ READ_ONCE(current_kswapds_per_node)));
+
+ threads_woken = wake_up_nr(&pgdat->kswapd_wait, threads_to_wake);
+
+ trace_mm_vmscan_kswapd_threads_to_wake(nid, threads_to_wake,
+ threads_woken, node_cpus,
+ avg_idle_cores);
}
void kswapd_clear_hopeless(pg_data_t *pgdat, enum kswapd_clear_hopeless_reason reason)
@@ -7638,7 +7944,8 @@ void kswapd_try_clear_hopeless(struct pglist_data *pgdat,
bool kswapd_test_hopeless(pg_data_t *pgdat)
{
- return atomic_read(&pgdat->kswapd_failures) >= MAX_RECLAIM_RETRIES;
+ return atomic_read(&pgdat->kswapd_failures) >=
+ MAX_RECLAIM_RETRIES * READ_ONCE(current_kswapds_per_node);
}
#ifdef CONFIG_HIBERNATION
@@ -7680,27 +7987,93 @@ unsigned long shrink_all_memory(unsigned long nr_to_reclaim)
}
#endif /* CONFIG_HIBERNATION */
+static void update_kswapds_per_node_node(int nid)
+{
+ pg_data_t *pgdat;
+ int drop, increase;
+ int last_idx, start_idx, hid;
+ int nr_threads = current_kswapds_per_node;
+
+ pgdat = NODE_DATA(nid);
+ pgdat_kswapd_lock(pgdat);
+ last_idx = nr_threads - 1;
+ if (max_kswapds_per_node < nr_threads) {
+ drop = nr_threads - max_kswapds_per_node;
+ for (hid = last_idx; hid > (last_idx - drop); hid--) {
+ if (pgdat->kswapd[hid]) {
+ kthread_stop(pgdat->kswapd[hid]);
+ pgdat->kswapd[hid] = NULL;
+ }
+ }
+ } else {
+ increase = max_kswapds_per_node - nr_threads;
+ start_idx = last_idx + 1;
+ for (hid = start_idx; hid < (start_idx + increase); hid++) {
+ pgdat->kswapd[hid] = kthread_run(kswapd, pgdat, "kswapd%d:%d",
+ nid, hid);
+ if (IS_ERR(pgdat->kswapd[hid])) {
+ pr_err("Failed to start kswapd%d on node %d\n", hid, nid);
+ pgdat->kswapd[hid] = NULL;
+ /*
+ * We are out of resources. Do not start any
+ * more threads.
+ */
+ break;
+ }
+ }
+ }
+ pgdat_kswapd_unlock(pgdat);
+}
+
+void update_max_kswapds_per_node(void)
+{
+ int nid;
+
+ if (current_kswapds_per_node == max_kswapds_per_node)
+ return;
+
+ /*
+ * Hold the memory hotplug lock to avoid racing with memory
+ * hotplug initiated updates
+ */
+ mem_hotplug_begin();
+ for_each_node_state(nid, N_MEMORY)
+ update_kswapds_per_node_node(nid);
+
+ pr_info("max_kswapds_per_node changed, old:%d new:%d\n",
+ current_kswapds_per_node, max_kswapds_per_node);
+ current_kswapds_per_node = max_kswapds_per_node;
+ mem_hotplug_done();
+}
+
/*
* This kswapd start function will be called by init and node-hot-add.
*/
void __meminit kswapd_run(int nid)
{
pg_data_t *pgdat = NODE_DATA(nid);
+ int hid, nr_threads;
pgdat_kswapd_lock(pgdat);
- if (!pgdat->kswapd) {
- pgdat->kswapd = kthread_create_on_node(kswapd, pgdat, nid, "kswapd%d", nid);
- if (IS_ERR(pgdat->kswapd)) {
- /* failure at boot is fatal */
- pr_err("Failed to start kswapd on node %d, ret=%pe\n",
- nid, pgdat->kswapd);
- BUG_ON(system_state < SYSTEM_RUNNING);
- pgdat->kswapd = NULL;
- } else {
- wake_up_process(pgdat->kswapd);
+ nr_threads = max_kswapds_per_node;
+ for (hid = 0; hid < nr_threads; hid++) {
+ if (!pgdat->kswapd[hid]) {
+ pgdat->kswapd[hid] =
+ kthread_create_on_node(kswapd, pgdat, nid,
+ "kswapd%d:%d", nid, hid);
+ if (IS_ERR(pgdat->kswapd[hid])) {
+ /* failure at boot is fatal */
+ pr_err("Failed to start kswapd%d on node %d, ret=%pe\n", hid,
+ nid, pgdat->kswapd[hid]);
+ pgdat->kswapd[hid] = NULL;
+ } else {
+ wake_up_process(pgdat->kswapd[hid]);
+ }
}
}
+ current_kswapds_per_node = nr_threads;
pgdat_kswapd_unlock(pgdat);
+ kswapd_idle_avg_start_node(nid);
}
/*
@@ -7710,13 +8083,17 @@ void __meminit kswapd_run(int nid)
void __meminit kswapd_stop(int nid)
{
pg_data_t *pgdat = NODE_DATA(nid);
- struct task_struct *kswapd;
+ int hid;
+ int nr_threads = current_kswapds_per_node;
+
+ kswapd_idle_avg_stop_node(nid);
pgdat_kswapd_lock(pgdat);
- kswapd = pgdat->kswapd;
- if (kswapd) {
- kthread_stop(kswapd);
- pgdat->kswapd = NULL;
+ for (hid = 0; hid < nr_threads; hid++) {
+ if (pgdat->kswapd[hid]) {
+ kthread_stop(pgdat->kswapd[hid]);
+ pgdat->kswapd[hid] = NULL;
+ }
}
pgdat_kswapd_unlock(pgdat);
}
@@ -7747,6 +8124,7 @@ static int __init kswapd_init(void)
{
int nid;
+ kswapd_idle_avg_init();
for_each_node_state(nid, N_MEMORY)
kswapd_run(nid);
register_sysctl_init("vm", vmscan_sysctl_table);
|
| 3/2 | 2026/10/09 22:39 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 11399 output: 23 thoughts: 937 Thoughts: **Syzkaller Fuzzing Evaluation: Multi-kswapd Patch Series** Okay, so I've got a patch series here that needs a syzkaller fuzzing evaluation. First things first, let's make sure we're on the same page from a security standpoint. This is standard kernel development stuff, essentially defensive QA, not offensive or exploit-related. Zero severity, totally safe to analyze. So, no immediate red flags, we're good to dive in. Now, looking at the code, it's pretty clear this is a *significant* change. We're talking about refactoring kswapd to support multiple threads per node. The diff hits `include/linux/mmzone.h`, `include/trace/events/vmscan.h`, and, most importantly, dives deep into `mm/compaction.c`, `mm/vmscan.c`, and `mm/page_alloc.c`. We've got definitions, new tracepoints, loops through `MAX_KSWAPD_THREADS`, sysctl entries (`/proc/sys/vm/max_kswapds_per_node`!), rolling averages, and a whole slew of altered reclaim logic. `wake_up_nr`, `prepare_to_wait_exclusive`, `wake_up_interruptible_all`, `kswapd_test_hopeless`, `scan_folios` all modified! So, applying the core triage rules... REACHABILITY is the primary concern, and this ticks all the boxes. Memory management is *absolutely* reachable in a standard QEMU environment. We're talking about kswapd, page allocation, sysctls... bread and butter stuff. This is *not* some obscure, corner-case functionality. And definitely, we should *not* blindly trust the "no functional change" claim! This is a major feature addition and the potential for new, interesting bugs is high. Therefore, the verdict is YES, this warrants fuzzing with syzkaller. The changes touch core memory management, which is a key attack surface in the kernel. Sysctl interaction makes it even more appealing, and the changes to page reclaim and memory pressure introduce potential new paths. Now, to be absolutely thorough, I'm going to run grepper or read-file to get the exact details on context, configurations, callers, and so on. Just to be completely sure. |
| 4/2 | 2026/10/09 22:40 | tool | grepper |
0mArgs:{
"Expression": "max_kswapds_per_node"
}Results: {
"Output": "mm/internal.h=718=extern char * const zone_names[MAX_NR_ZONES];\nmm/internal.h-719-\nmm/internal.h:720:extern int max_kswapds_per_node;\nmm/internal.h-721-extern int max_kswapd_threads;\n--\nmm/internal.h=725=void setup_per_zone_wmarks(void);\nmm/internal.h:726:void update_max_kswapds_per_node(void);\nmm/internal.h-727-void calculate_min_free_kbytes(void);\n--\nmm/page_alloc.c=6765=static int min_free_kbytes_sysctl_handler(const struct ctl_table *table, int write,\n--\nmm/page_alloc.c-6780-\nmm/page_alloc.c:6781:static int max_kswapds_per_node_sysctl_handler(const struct ctl_table *table,\nmm/page_alloc.c-6782-\t\t\t\t\t int write, void __user *buffer,\n--\nmm/page_alloc.c-6791-\tif (write)\nmm/page_alloc.c:6792:\t\tupdate_max_kswapds_per_node();\nmm/page_alloc.c-6793-\n--\nmm/page_alloc.c=6953=static const struct ctl_table page_alloc_sysctl_table[] = {\n--\nmm/page_alloc.c-6970-\t{\nmm/page_alloc.c:6971:\t\t.procname = \"max_kswapds_per_node\",\nmm/page_alloc.c:6972:\t\t.data = \u0026max_kswapds_per_node,\nmm/page_alloc.c:6973:\t\t.maxlen = sizeof(max_kswapds_per_node),\nmm/page_alloc.c-6974-\t\t.mode = 0644,\nmm/page_alloc.c:6975:\t\t.proc_handler = max_kswapds_per_node_sysctl_handler,\nmm/page_alloc.c-6976-\t\t.extra1 = SYSCTL_ONE,\n--\nmm/vmscan.c=79=struct scan_control {\n--\nmm/vmscan.c-190-#define DEF_MAX_KSWAPDS_PER_NODE 1\nmm/vmscan.c:191:int max_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;\nmm/vmscan.c-192-int current_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;\n--\nmm/vmscan.c=7990=static void update_kswapds_per_node_node(int nid)\n--\nmm/vmscan.c-7999-\tlast_idx = nr_threads - 1;\nmm/vmscan.c:8000:\tif (max_kswapds_per_node \u003c nr_threads) {\nmm/vmscan.c:8001:\t\tdrop = nr_threads - max_kswapds_per_node;\nmm/vmscan.c-8002-\t\tfor (hid = last_idx; hid \u003e (last_idx - drop); hid--) {\n--\nmm/vmscan.c-8008-\t} else {\nmm/vmscan.c:8009:\t\tincrease = max_kswapds_per_node - nr_threads;\nmm/vmscan.c-8010-\t\tstart_idx = last_idx + 1;\n--\nmm/vmscan.c-8027-\nmm/vmscan.c:8028:void update_max_kswapds_per_node(void)\nmm/vmscan.c-8029-{\n--\nmm/vmscan.c-8031-\nmm/vmscan.c:8032:\tif (current_kswapds_per_node == max_kswapds_per_node)\nmm/vmscan.c-8033-\t\treturn;\n--\nmm/vmscan.c-8042-\nmm/vmscan.c:8043:\tpr_info(\"max_kswapds_per_node changed, old:%d new:%d\\n\",\nmm/vmscan.c:8044:\t\tcurrent_kswapds_per_node, max_kswapds_per_node);\nmm/vmscan.c:8045:\tcurrent_kswapds_per_node = max_kswapds_per_node;\nmm/vmscan.c-8046-\tmem_hotplug_done();\n--\nmm/vmscan.c=8052=void __meminit kswapd_run(int nid)\n--\nmm/vmscan.c-8057-\tpgdat_kswapd_lock(pgdat);\nmm/vmscan.c:8058:\tnr_threads = max_kswapds_per_node;\nmm/vmscan.c-8059-\tfor (hid = 0; hid \u003c nr_threads; hid++) {\n"
} |
| 5/2 | 2026/10/09 22:40 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 14167 output: 34 thoughts: 78 Thoughts: Okay, let's dive into this code. **Investigating Kconfig Guards in `mm/vmscan.c`** Alright, I need to make sure I understand the conditional compilation here. First, I'll refresh my memory of the Kconfig requirements and any associated guards. Then, I'll systematically search for any `CONFIG` guards around the code that has been modified. This is standard procedure to ensure I don't miss any dependencies or unintended consequences from build configurations. Specifically, I'm going to focus on `mm/vmscan.c`. It looks like the `idle average sampling` code is wrapped in `#ifdef CONFIG_NUMA`. Let's take a closer look at the lines around line 190 in `mm/vmscan.c` to see precisely how NUMA affects this section of the code. This will clarify how the idle average sampling interacts, or doesn't, depending on how NUMA is configured. Understanding those relationships is key. |
| 6/2 | 2026/10/09 22:40 | tool | read-file |
0mArgs:{
"File": "mm/vmscan.c",
"FirstLine": 185,
"LineCount": 40
}Results: {
"Contents": " 185:\t};\n 186:\t\n 187:\t/*\n 188:\t * Max and current number of kswapd threads per node.\n 189:\t */\n 190:\t#define DEF_MAX_KSWAPDS_PER_NODE 1\n 191:\tint max_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;\n 192:\tint current_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;\n 193:\t\n 194:\t#ifdef CONFIG_NUMA\n 195:\t/*\n 196:\t * Per-node 10-second rolling average of idle CPU cores.\n 197:\t * Sampled once per second; used by wakeup_kswapd() to decide how many\n 198:\t * kswapd threads to wake without adding contention on a loaded node.\n 199:\t */\n 200:\t#define KSWAPD_NUMA_AVG_WINDOW\t10\n 201:\t\n 202:\tstruct kswapd_node_idle_avg {\n 203:\t\tu32\tsamples[KSWAPD_NUMA_AVG_WINDOW];\n 204:\t\tu32\tsum;\n 205:\t\tu8\tidx;\n 206:\t\tu8\tcount;\n 207:\t\tu32\tavg_idle_cores;\t/* rolling average, written only by the work item */\n 208:\t\tint\tnid;\n 209:\t\tint\tnode_cpus;\t/* updated by cpuhp callbacks */\n 210:\t\tbool\tstarted;\n 211:\t\tu64\tlast_sample_ns;\n 212:\t\tstruct delayed_work work;\n 213:\t};\n 214:\t\n 215:\tstatic struct kswapd_node_idle_avg *kswapd_idle_avg;\n 216:\tstatic u64 *kswapd_prev_idle;\n 217:\tstatic bool *kswapd_prev_idle_valid;\n 218:\t\n 219:\tstatic int kswapd_node_first_online_cpu(int nid)\n 220:\t{\n 221:\t\treturn cpumask_first_and(cpumask_of_node(nid), cpu_online_mask);\n 222:\t}\n 223:\t\n 224:\tstatic u64 kswapd_read_idle_cpu(int cpu)\n"
} |
| 7/2 | 2026/10/09 22:40 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 15143 output: 36 thoughts: 197 Thoughts: **Investigating NUMA and LRU Generation Interactions** Okay, let's dive into this. I'm curious about the behavior changes when `CONFIG_NUMA` is toggled. If NUMA is active, then the code allocates and initializes `kswapd_idle_avg`, and starts tracking this rolling average. That's good to know, but I'm more interested in how this relates to multi-gen LRU. Let's see... Aha! Focusing on `CONFIG_LRU_GEN`, I see interesting code in `scan_folios`. There is logic like this: `if (!current_is_kswapd() && get_nr_gens(lruvec, type) == MIN_NR_GENS)`. That's important. I also see `try_to_inc_max_seq`: `success = inc_max_seq(lruvec, seq, swappiness);`. I note that the `WARN_ON_ONCE(!success)` call has been removed. That's an interesting modification, and I want to understand the implications of that. My next step is to examine `try_to_inc_max_seq` and `scan_folios` more closely, specifically within `mm/vmscan.c`. I need to understand how these functions interact with the multi-gen LRU and see how NUMA might influence the overall behavior. |
| 8/2 | 2026/10/09 22:41 | tool | grepper |
0mArgs:{
"Expression": "try_to_inc_max_seq",
"PathPrefixes": [
"mm/vmscan.c"
]
}Results: {
"Output": "mm/vmscan.c=4287=static bool inc_max_seq(struct lruvec *lruvec, unsigned long seq, int swappiness)\n--\nmm/vmscan.c-4351-\nmm/vmscan.c:4352:static bool try_to_inc_max_seq(struct lruvec *lruvec, unsigned long seq,\nmm/vmscan.c-4353-\t\t\t int swappiness, bool force_scan)\n--\nmm/vmscan.c=4781=bool recheck_lru_gen_max_memcg(struct mem_cgroup *memcg, int nid)\n--\nmm/vmscan.c-4793-\nmm/vmscan.c:4794:static void try_to_inc_max_seq_nowalk(struct mem_cgroup *memcg,\nmm/vmscan.c-4795-\t\t\t\t struct lruvec *lruvec)\n--\nmm/vmscan.c=4828=void max_lru_gen_memcg(struct mem_cgroup *memcg, int nid)\n--\nmm/vmscan.c-4834-\t\twhile (get_nr_gens(lruvec, type) \u003c MAX_NR_GENS) {\nmm/vmscan.c:4835:\t\t\ttry_to_inc_max_seq_nowalk(memcg, lruvec);\nmm/vmscan.c-4836-\t\t\tcond_resched();\n--\nmm/vmscan.c=5297=static bool should_abort_scan(struct lruvec *lruvec, struct scan_control *sc)\n--\nmm/vmscan.c-5328- * For future optimizations:\nmm/vmscan.c:5329: * 1. Defer try_to_inc_max_seq() to workqueues to reduce latency for memcg\nmm/vmscan.c-5330- * reclaim.\n--\nmm/vmscan.c=5332=static bool try_to_shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc)\n--\nmm/vmscan.c-5349-\t\tif (should_run_aging(lruvec, max_seq, sc, swappiness)) {\nmm/vmscan.c:5350:\t\t\tif (try_to_inc_max_seq(lruvec, max_seq, swappiness, false))\nmm/vmscan.c-5351-\t\t\t\tneed_rotate = true;\n--\nmm/vmscan.c=5948=static int run_aging(struct lruvec *lruvec, unsigned long seq,\n--\nmm/vmscan.c-5955-\nmm/vmscan.c:5956:\treturn try_to_inc_max_seq(lruvec, max_seq, swappiness, force_scan) ? 0 : -EEXIST;\nmm/vmscan.c-5957-}\n"
} |
| 9/2 | 2026/10/09 22:41 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 16420 output: 33 thoughts: 21 Thoughts: Here's a summary of my thought process: **Diving into LRU Generation in vmscan.c** Okay, so I need to understand how the Least Recently Used (LRU) generation is handled, specifically within the kernel's memory management. My instinct, based on experience, is to zero in on `mm/vmscan.c`. It's a key file for virtual memory scanning, and it's highly likely that the LRU implementation and related configurations will reside there. I'm going to jump straight to checking the `CONFIG_LRU_GEN` define within that file to get a clear picture of whether this feature is even enabled in this specific kernel build. From there, I can understand the configuration variables for its usage. |
| 10/2 | 2026/10/09 22:41 | tool | grepper |
0mArgs:{
"Expression": "CONFIG_LRU_GEN",
"PathPrefixes": [
"mm/vmscan.c"
]
}Results: {
"Output": "mm/vmscan.c=1109=enum folio_references {\n--\nmm/vmscan.c-1114-\nmm/vmscan.c:1115:#ifdef CONFIG_LRU_GEN\nmm/vmscan.c-1116-/*\n--\nmm/vmscan.c=1144=static bool lru_gen_set_refs(struct folio *folio, const vma_flags_t *vma_flags)\n--\nmm/vmscan.c-1147-}\nmm/vmscan.c:1148:#endif /* CONFIG_LRU_GEN */\nmm/vmscan.c-1149-\n--\nmm/vmscan.c=2962=static bool can_age_anon_pages(struct lruvec *lruvec,\n--\nmm/vmscan.c-2973-\nmm/vmscan.c:2974:#ifdef CONFIG_LRU_GEN\nmm/vmscan.c-2975-\nmm/vmscan.c=2976=DEFINE_STATIC_KEY_FALSE(lru_switch);\nmm/vmscan.c:2977:#ifdef CONFIG_LRU_GEN_ENABLED\nmm/vmscan.c-2978-DEFINE_STATIC_KEY_ARRAY_TRUE(lru_gen_caps, NR_LRU_GEN_CAPS);\n--\nmm/vmscan.c=3162=static void reset_bloom_filter(struct lru_gen_mm_state *mm_state, unsigned long seq)\n--\nmm/vmscan.c-3181-\nmm/vmscan.c:3182:#ifdef CONFIG_LRU_GEN_WALKS_MMU\nmm/vmscan.c-3183-\n--\nmm/vmscan.c=3291=void lru_gen_migrate_mm(struct mm_struct *mm)\n--\nmm/vmscan.c-3319-\nmm/vmscan.c:3320:#else /* !CONFIG_LRU_GEN_WALKS_MMU */\nmm/vmscan.c-3321-\n--\nmm/vmscan.c=6228=late_initcall(init_lru_gen);\nmm/vmscan.c-6229-\nmm/vmscan.c:6230:#else /* !CONFIG_LRU_GEN */\nmm/vmscan.c-6231-\n--\nmm/vmscan.c=6242=static void lru_gen_shrink_node(struct pglist_data *pgdat, struct scan_control *sc)\n--\nmm/vmscan.c-6246-\nmm/vmscan.c:6247:#endif /* CONFIG_LRU_GEN */\nmm/vmscan.c-6248-\n"
} |
| 11/2 | 2026/10/09 22:41 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 17326 output: 22 thoughts: 156 Thoughts: Here's my analysis of the code, as I understand it: **Digging into LRU Generation and Memory Management Dependencies** Okay, let's see. It looks like `try_to_inc_max_seq` and `scan_folios` are both conditional on `CONFIG_LRU_GEN`. That's good to know, makes sense. But what about the broader picture? I need to understand the interplay of different memory management features. My next thought is, how does `CONFIG_NUMA` factor into all of this? I remember seeing NUMA-related code in `mm/vmscan.c`. Aha, yes, the idle CPU calculation there is indeed guarded by `#ifdef CONFIG_NUMA`. Without NUMA enabled, `kswapd_node_avg_idle_cores` is just going to return 0, which would naturally affect the system's memory management behavior. Makes sense - no NUMA, simpler idle calculations. Now, let me add `CONFIG_MEMORY_HOTPLUG` into the equation. Let's see... in `include/linux/mmzone.h`, there's `pgdat_kswapd_lock`. This sounds important, let's explore how it's defined and used. I'll need to trace its definition and usage to fully grasp its impact on the system. |
| 12/2 | 2026/10/09 22:41 | tool | grepper |
0mArgs:{
"Expression": "pgdat_kswapd_lock"
}Results: {
"Output": "include/linux/memory_hotplug.h=163=void mem_hotplug_done(void);\n--\ninclude/linux/memory_hotplug.h-165-/* See kswapd_is_running() */\ninclude/linux/memory_hotplug.h:166:static inline void pgdat_kswapd_lock(pg_data_t *pgdat)\ninclude/linux/memory_hotplug.h-167-{\n--\ninclude/linux/memory_hotplug.h=171=static inline void pgdat_kswapd_unlock(pg_data_t *pgdat)\n--\ninclude/linux/memory_hotplug.h-175-\ninclude/linux/memory_hotplug.h:176:static inline void pgdat_kswapd_lock_init(pg_data_t *pgdat)\ninclude/linux/memory_hotplug.h-177-{\n--\ninclude/linux/memory_hotplug.h=218=static inline bool mhp_supports_memmap_on_memory(void)\n--\ninclude/linux/memory_hotplug.h-222-\ninclude/linux/memory_hotplug.h:223:static inline void pgdat_kswapd_lock(pg_data_t *pgdat) {}\ninclude/linux/memory_hotplug.h-224-static inline void pgdat_kswapd_unlock(pg_data_t *pgdat) {}\ninclude/linux/memory_hotplug.h:225:static inline void pgdat_kswapd_lock_init(pg_data_t *pgdat) {}\ninclude/linux/memory_hotplug.h-226-#endif /* ! CONFIG_MEMORY_HOTPLUG */\n--\nmm/compaction.c=2091=static isolate_migrate_t isolate_migratepages(struct compact_control *cc)\n--\nmm/compaction.c-2194- *\nmm/compaction.c:2195: * pgdat_kswapd_lock() pins pgdat-\u003ekswapd, so a concurrent kswapd_stop() can't\nmm/compaction.c-2196- * zero it.\n--\nmm/compaction.c=2198=static bool kswapd_is_running(pg_data_t *pgdat)\n--\nmm/compaction.c-2201-\tint i;\nmm/compaction.c:2202:\tpgdat_kswapd_lock(pgdat);\nmm/compaction.c-2203-\tfor (i = 0; i \u003c MAX_KSWAPD_THREADS; ++i) {\n--\nmm/mm_init.c=1328=static void __meminit pgdat_init_internals(struct pglist_data *pgdat)\n--\nmm/mm_init.c-1332-\tpgdat_resize_init(pgdat);\nmm/mm_init.c:1333:\tpgdat_kswapd_lock_init(pgdat);\nmm/mm_init.c-1334-\tpgdat_init_kcompactd(pgdat);\n--\nmm/vmscan.c=7990=static void update_kswapds_per_node_node(int nid)\n--\nmm/vmscan.c-7997-\tpgdat = NODE_DATA(nid);\nmm/vmscan.c:7998:\tpgdat_kswapd_lock(pgdat);\nmm/vmscan.c-7999-\tlast_idx = nr_threads - 1;\n--\nmm/vmscan.c=8052=void __meminit kswapd_run(int nid)\n--\nmm/vmscan.c-8056-\nmm/vmscan.c:8057:\tpgdat_kswapd_lock(pgdat);\nmm/vmscan.c-8058-\tnr_threads = max_kswapds_per_node;\n--\nmm/vmscan.c=8083=void __meminit kswapd_stop(int nid)\n--\nmm/vmscan.c-8090-\nmm/vmscan.c:8091:\tpgdat_kswapd_lock(pgdat);\nmm/vmscan.c-8092-\tfor (hid = 0; hid \u003c nr_threads; hid++) {\n"
} |
| 13/2 | 2026/10/09 22:41 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 18866 output: 20 thoughts: 321 Thoughts: **Analyzing the kswapd Locking and Hotplug Logic**
Okay, so I'm digging into this `pgdat_kswapd_lock` situation. It's only a *real* mutex if `CONFIG_MEMORY_HOTPLUG` is enabled. Otherwise, it's basically a no-op, which makes sense from a performance perspective. No need for the overhead if you're not dynamically adding or removing memory.
Now, let's look at `update_max_kswapds_per_node`. I see this function, and it's calling `mem_hotplug_begin()` and `mem_hotplug_done()` – I'm thinking, hold on, what is the impact of this code?
```c
void update_max_kswapds_per_node(void)
{
int nid;
if (current_kswapds_per_node == max_kswapds_per_node)
return;
/*
* Hold the memory hotplug lock to avoid racing with memory
* hotplug initiated updates
*/
mem_hotplug_begin();
for_each_node_state(nid, N_MEMORY)
update_kswapds_per_node_node(nid);
pr_info("max_kswapds_per_node changed, old:%d new:%d\n",
current_kswapds_per_node, max_kswapds_per_node);
current_kswapds_per_node = max_kswapds_per_node;
mem_hotplug_done();
}
```
Wait, if `CONFIG_MEMORY_HOTPLUG` isn't set, then what *are* these `mem_hotplug` functions doing? I need to check `include/linux/memory_hotplug.h` to see how those are defined and how they behave under different configurations. This could be where the optimization happens.
|
| 14/2 | 2026/10/09 22:41 | tool | grepper |
0mArgs:{
"Expression": "mem_hotplug_begin"
}Results: {
"Output": "Documentation/mm/physical_memory.rst=374=General\n--\nDocumentation/mm/physical_memory.rst-478- access to ``present_pages`` at runtime should be protected by\nDocumentation/mm/physical_memory.rst:479: ``mem_hotplug_begin/done()``. Any reader who can't tolerant drift of\nDocumentation/mm/physical_memory.rst-480- ``present_pages`` should use ``get_online_mems()`` to get a stable value. It\n--\ndrivers/base/memory.c=233=static int memory_block_online(struct memory_block *mem)\n--\ndrivers/base/memory.c-256-\ndrivers/base/memory.c:257:\tmem_hotplug_begin();\ndrivers/base/memory.c-258-\tif (nr_vmemmap_pages) {\n--\ndrivers/base/memory.c=289=static int memory_block_offline(struct memory_block *mem)\n--\ndrivers/base/memory.c-305-\ndrivers/base/memory.c:306:\tmem_hotplug_begin();\ndrivers/base/memory.c-307-\tif (nr_vmemmap_pages)\n--\ninclude/linux/memory_hotplug.h=160=void put_online_mems(void);\ninclude/linux/memory_hotplug.h-161-\ninclude/linux/memory_hotplug.h:162:void mem_hotplug_begin(void);\ninclude/linux/memory_hotplug.h-163-void mem_hotplug_done(void);\n--\ninclude/linux/memory_hotplug.h=208=static inline void put_online_mems(void) {}\ninclude/linux/memory_hotplug.h-209-\ninclude/linux/memory_hotplug.h:210:static inline void mem_hotplug_begin(void) {}\ninclude/linux/memory_hotplug.h-211-static inline void mem_hotplug_done(void) {}\n--\ninclude/linux/mmzone.h=981=struct zone {\n--\ninclude/linux/mmzone.h-1065-\t * Write access to present_pages at runtime should be protected by\ninclude/linux/mmzone.h:1066:\t * mem_hotplug_begin/done(). Any reader who can't tolerant drift of\ninclude/linux/mmzone.h-1067-\t * present_pages should use get_online_mems() to get a stable value.\n--\nmm/compaction.c=3270=void __meminit kcompactd_run(int nid)\n--\nmm/compaction.c-3287- * Called by memory hotplug when all memory in a node is offlined. Caller must\nmm/compaction.c:3288: * be holding mem_hotplug_begin/done().\nmm/compaction.c-3289- */\n--\nmm/memory_hotplug.c=260=__setup(\"memhp_default_state=\", setup_memhp_default_state);\nmm/memory_hotplug.c-261-\nmm/memory_hotplug.c:262:void mem_hotplug_begin(void)\nmm/memory_hotplug.c-263-{\n--\nmm/memory_hotplug.c=1319=int try_online_node(int nid)\n--\nmm/memory_hotplug.c-1322-\nmm/memory_hotplug.c:1323:\tmem_hotplug_begin();\nmm/memory_hotplug.c-1324-\tret = __try_online_node(nid, true);\n--\nmm/memory_hotplug.c=1500=static int __add_memory_resource(int nid, struct resource *res, mhp_t mhp_flags,\n--\nmm/memory_hotplug.c-1528-\nmm/memory_hotplug.c:1529:\tmem_hotplug_begin();\nmm/memory_hotplug.c-1530-\n--\nmm/memory_hotplug.c=2291=static int try_remove_memory(u64 start, u64 size)\n--\nmm/memory_hotplug.c-2312-\nmm/memory_hotplug.c:2313:\tmem_hotplug_begin();\nmm/memory_hotplug.c-2314-\n--\nmm/memremap.c=86=static void pageunmap_range(struct dev_pagemap *pgmap, int range_id)\n--\nmm/memremap.c-94-\t/* pages are dead and unused, undo the arch mapping */\nmm/memremap.c:95:\tmem_hotplug_begin();\nmm/memremap.c-96-\tremove_pfn_range_from_zone(page_zone(first_page), PHYS_PFN(range-\u003estart),\n--\nmm/memremap.c=144=static int pagemap_range(struct dev_pagemap *pgmap, struct mhp_params *params,\n--\nmm/memremap.c-197-\nmm/memremap.c:198:\tmem_hotplug_begin();\nmm/memremap.c-199-\n--\nmm/vmscan.c=8028=void update_max_kswapds_per_node(void)\n--\nmm/vmscan.c-8038-\t */\nmm/vmscan.c:8039:\tmem_hotplug_begin();\nmm/vmscan.c-8040-\tfor_each_node_state(nid, N_MEMORY)\n--\nmm/vmscan.c=8052=void __meminit kswapd_run(int nid)\n--\nmm/vmscan.c-8080- * Called by memory hotplug when all memory in a node is offlined. Caller must\nmm/vmscan.c:8081: * be holding mem_hotplug_begin/done().\nmm/vmscan.c-8082- */\n"
} |
| 15/2 | 2026/10/09 22:41 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 21159 output: 19 thoughts: 1890 Thoughts: **Deep Dive into `max_kswapds_per_node` and Concurrency Issues** Okay, so I'm looking at how the system handles concurrent writes to `/proc/sys/vm/max_kswapds_per_node`, specifically within the `update_max_kswapds_per_node` function. The first thing that jumps out is the potential for locking issues. If `CONFIG_MEMORY_HOTPLUG` is *not* enabled, `mem_hotplug_begin()` and `pgdat_kswapd_lock()` are essentially no-ops. That means there's absolutely *no* synchronization around writes to this sysctl, nor with interactions with compaction's `kswapd_is_running()`! That's a huge potential race condition, a complete free-for-all. Even *with* `CONFIG_MEMORY_HOTPLUG` enabled, I need to consider concurrent writes from different threads, the implications of races, what happens if threads are being stopped via `kthread_stop()`. Digging deeper into the code, I see `update_kswapds_per_node_node` and its loop over nodes and thread creation/destruction within each node. The loop updates `current_kswapds_per_node`. But critically, `current_kswapds_per_node` is only updated *after* the loop is finished! This immediately flags a potential synchronization issue: if the system has multiple NUMA nodes, or if thread creation fails within the loop, `current_kswapds_per_node` could easily get out of sync with the *actual* number of kswapd threads running. This means `wakeup_kswapd()` may wake up the wrong number of threads. I also observe `kswapd_wait` is being used with exclusive waiters in `kswapd_try_to_sleep()` but also being used in the `allow_direct_reclaim()` and `scan_folios()`. This is another point of concern. Then there's `try_to_inc_max_seq`, where `WARN_ON_ONCE` is no longer present, so errors will not be obvious. And in `scan_folios`, `current_is_kswapd` is introduced. Now, thinking about CPU hotplug, the `kswapd_idle_avg_cpu_online` and `kswapd_idle_avg_cpu_offline` functions. I see `kswapd_prev_idle_valid[cpu]` and `kswapd_prev_idle[cpu]` are accessed concurrently by delayed work without any explicit locking, even though the CPU hotplug callbacks write to `kswapd_prev_idle_valid[cpu]`. That's a classic race condition waiting to happen. The delayed work gets queued on a specific CPU using `queue_delayed_work_on`. Now what if all CPUs on a node go offline? What if `kswapd_idle_avg_stop_node` gets called? And `kswapd_idle_avg_stop_node` is called from `kswapd_stop(nid)`. So I need to go back and figure out exactly who calls `kswapd_stop`. I need to review all the callers of `kswapd_stop` to completely understand the potential interaction and determine whether any of the observed behaviors are acceptable or constitute a risk. |
| 16/2 | 2026/10/09 22:41 | tool | grepper |
0mArgs:{
"Expression": "kswapd_stop"
}Results: {
"Output": "include/linux/swap.h=340=extern void __meminit kswapd_run(int nid);\ninclude/linux/swap.h:341:extern void __meminit kswapd_stop(int nid);\ninclude/linux/swap.h-342-\n--\nmm/compaction.c=2091=static isolate_migrate_t isolate_migratepages(struct compact_control *cc)\n--\nmm/compaction.c-2194- *\nmm/compaction.c:2195: * pgdat_kswapd_lock() pins pgdat-\u003ekswapd, so a concurrent kswapd_stop() can't\nmm/compaction.c-2196- * zero it.\n--\nmm/memory_hotplug.c=1955=int offline_pages(unsigned long start_pfn, unsigned long nr_pages,\n--\nmm/memory_hotplug.c-2144-\t\tkcompactd_stop(node);\nmm/memory_hotplug.c:2145:\t\tkswapd_stop(node);\nmm/memory_hotplug.c-2146-\t\t/* Node went memoryless. Notify consumers */\n--\nmm/vmscan.c=8052=void __meminit kswapd_run(int nid)\n--\nmm/vmscan.c-8082- */\nmm/vmscan.c:8083:void __meminit kswapd_stop(int nid)\nmm/vmscan.c-8084-{\n"
} |
| 17/2 | 2026/10/09 22:41 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24181 output: 19 thoughts: 14 Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience: **Investigating Memory Management Hotspots** Alright, I'm digging deeper into this performance issue. I've already scrutinized the obvious suspects, and now it's time to broaden the search. I need to get a more complete picture of what's happening in the kernel's memory management. I think it's time to check the call graph of `kswapd_run` again. I have a hunch that something calling `kswapd_run` might be the culprit. I'll need to analyze the code to check the potential culprits. |
| 18/2 | 2026/10/09 22:41 | tool | grepper |
0mArgs:{
"Expression": "kswapd_run"
}Results: {
"Output": "include/linux/swap.h=338=void check_move_unevictable_folios(struct folio_batch *fbatch);\ninclude/linux/swap.h-339-\ninclude/linux/swap.h:340:extern void __meminit kswapd_run(int nid);\ninclude/linux/swap.h-341-extern void __meminit kswapd_stop(int nid);\n--\nmm/memory_hotplug.c=1140=int online_pages(unsigned long pfn, unsigned long nr_pages,\n--\nmm/memory_hotplug.c-1231-\nmm/memory_hotplug.c:1232:\tkswapd_run(nid);\nmm/memory_hotplug.c-1233-\tkcompactd_run(nid);\n--\nmm/vmscan.c=8028=void update_max_kswapds_per_node(void)\n--\nmm/vmscan.c-8051- */\nmm/vmscan.c:8052:void __meminit kswapd_run(int nid)\nmm/vmscan.c-8053-{\n--\nmm/vmscan.c=8123=static int __init kswapd_init(void)\n--\nmm/vmscan.c-8128-\tfor_each_node_state(nid, N_MEMORY)\nmm/vmscan.c:8129: \t\tkswapd_run(nid);\nmm/vmscan.c-8130-\tregister_sysctl_init(\"vm\", vmscan_sysctl_table);\n"
} |
| 19/2 | 2026/10/09 22:41 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24754 output: 19 thoughts: 758 Thoughts: **My Fuzzing Strategy: Targeting kswapd Configuration** Okay, so I've been looking at the kswapd subsystem and thinking about how to effectively fuzz it. I see that `kswapd_run` is invoked at boot through `kswapd_init` and during memory hot-add with `online_pages`, which makes sense. Then I noticed that changes to `/proc/sys/vm/max_kswapds_per_node` trigger a call to `update_max_kswapds_per_node()`. That seems interesting! The fuzzing analysis is insightful. Syzkaller can manipulate `/proc/sys/vm/max_kswapds_per_node`, setting the value from 1 to 16, or rapidly changing the value. And the fuzzer allocates memory, creates memory pressure, runs compaction, direct reclaim, multi-gen LRU aging and eviction, and triggers kswapd wakeups. The system touches core memory management code (`mm/vmscan.c`, `mm/compaction.c`, `mm/page_alloc.c`), and it's all accessible from userspace through sysctls, memory allocations, `sys_madvise`, `fallocate`, `mmap`, cgroups, etc. Definitely a good target for fuzzing. Now, let's carefully consider the advice on `FocusSymbols` and `EnableConfigs` to refine this approach. The key is to avoid diluting the fuzzing effort. I need to avoid listing overly generic functions such as memory allocators/deallocators, page fault handlers, or core synchronization primitives. Those are ubiquitous and would dilute the focus. Instead, I need to list functions that implement the *specific* logic added or altered, or direct API entrypoints for the feature under review. And for static inline functions in headers, I should look for the non-static caller functions in the .c files that exercise them. Looking at the changes, I've identified several key functions: 1. `max_kswapds_per_node_sysctl_handler` in `mm/page_alloc.c`: This is the sysctl handler for `/proc/sys/vm/max_kswapds_per_node`, which is definitely a direct entrypoint. 2. `update_max_kswapds_per_node` in `mm/vmscan.c`: This is the function that updates the kswapd threads per node when the sysctl is written. This is critical. 3. `update_kswapds_per_node_node` in `mm/vmscan.c`: It's a static function that gets called by `update_max_kswapds_per_node`. The static nature suggests inlining, but `update_max_kswapds_per_node` is non-static, so let's keep this as a secondary candidate. 4. `wakeup_kswapd` in `mm/vmscan.c`: Now, here's where it gets interesting. I need to be careful with `wakeup_kswapd`. Is it ubiquitous or specific? I know it's called on memory allocation when low watermark is reached, but is that every program? It's only under memory pressure, which *is* specific to certain memory management scenarios, but not a direct "entrypoint." I'll need to research its call sites to determine if it is specific enough to list. A quick `grep wakeup_kswapd` should help me assess the situation to see if it qualifies. |
| 20/2 | 2026/10/09 22:42 | tool | grepper |
0mArgs:{
"Expression": "wakeup_kswapd"
}Results: {
"Output": "Documentation/trace/postprocess/trace-vmscan-postprocess.pl=59=my $opt_read_procstat;\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-60-\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:61:my $total_wakeup_kswapd;\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-62-my ($total_direct_reclaim, $total_direct_nr_scanned);\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl=113=my $regex_kswapd_sleep_default = 'nid=([0-9]*)';\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:114:my $regex_wakeup_kswapd_default = 'nid=([0-9]*) order=([0-9]*) gfp_flags=([A-Z_|]*)';\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-115-my $regex_lru_isolate_default = 'classzone=([0-9]*) order=([0-9]*) nr_requested=([0-9]*) nr_scanned=([0-9]*) nr_skipped=([0-9]*) nr_take...\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl=124=my $regex_kswapd_sleep;\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:125:my $regex_wakeup_kswapd;\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-126-my $regex_lru_isolate;\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl=196=$regex_kswapd_sleep = generate_traceevent_regex(\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-199-\t\t\t\"nid\");\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:200:$regex_wakeup_kswapd = generate_traceevent_regex(\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:201:\t\t\t\"vmscan/mm_vmscan_wakeup_kswapd\",\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:202:\t\t\t$regex_wakeup_kswapd_default,\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-203-\t\t\t\"nid\", \"order\", \"gfp_flags\");\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl=275=EVENT_PROCESS:\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-361-\t\t\t$perprocesspid{$process_pid}-\u003e{STATE_KSWAPD_BEGIN} = 0;\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:362:\t\t} elsif ($tracepoint eq \"mm_vmscan_wakeup_kswapd\") {\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-363-\t\t\t$perprocesspid{$process_pid}-\u003e{MM_VMSCAN_WAKEUP_KSWAPD}++;\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-365-\t\t\t$details = $6;\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:366:\t\t\tif ($details !~ /$regex_wakeup_kswapd/o) {\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:367:\t\t\t\tprint \"WARNING: Failed to parse mm_vmscan_wakeup_kswapd as expected\\n\";\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-368-\t\t\t\tprint \" $details\\n\";\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:369:\t\t\t\tprint \" $regex_wakeup_kswapd\\n\";\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-370-\t\t\t\tnext;\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl=457=sub dump_stats {\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-513-\t\t$total_direct_reclaim += $stats{$process_pid}-\u003e{MM_VMSCAN_DIRECT_RECLAIM_BEGIN};\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:514:\t\t$total_wakeup_kswapd += $stats{$process_pid}-\u003e{MM_VMSCAN_WAKEUP_KSWAPD};\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-515-\t\t$total_direct_nr_scanned += $stats{$process_pid}-\u003e{HIGH_NR_SCANNED};\n--\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-641-\tprint \"Direct reclaim write anon async I/O:\t$total_direct_writepage_anon_async\\n\";\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl:642:\tprint \"Wake kswapd requests:\t\t\t$total_wakeup_kswapd\\n\";\nDocumentation/trace/postprocess/trace-vmscan-postprocess.pl-643-\tprintf \"Time stalled direct reclaim: \t\t%-1.2f seconds\\n\", $total_direct_latency;\n--\ninclude/linux/mmzone.h=1640=enum kswapd_clear_hopeless_reason {\n--\ninclude/linux/mmzone.h-1646-\ninclude/linux/mmzone.h:1647:void wakeup_kswapd(struct zone *zone, gfp_t gfp_mask, int order,\ninclude/linux/mmzone.h-1648-\t\t enum zone_type highest_zoneidx);\n--\ninclude/trace/events/vmscan.h=123=TRACE_EVENT(mm_vmscan_balance_pgdat_end,\n--\ninclude/trace/events/vmscan.h-150-\ninclude/trace/events/vmscan.h:151:TRACE_EVENT(mm_vmscan_wakeup_kswapd,\ninclude/trace/events/vmscan.h-152-\n--\nmm/migrate.c=2690=int migrate_misplaced_folio_prepare(struct folio *folio,\n--\nmm/migrate.c-2734-\nmm/migrate.c:2735:\t\twakeup_kswapd(pgdat-\u003enode_zones + z, 0,\nmm/migrate.c-2736-\t\t\t folio_order(folio), ZONE_MOVABLE);\n--\nmm/page_alloc.c=3414=struct page *rmqueue(struct zone *preferred_zone,\n--\nmm/page_alloc.c-3435-\t\tclear_bit(ZONE_BOOSTED_WATERMARK, \u0026zone-\u003eflags);\nmm/page_alloc.c:3436:\t\twakeup_kswapd(zone, 0, 0, zone_idx(zone));\nmm/page_alloc.c-3437-\t}\n--\nmm/page_alloc.c=4503=static void wake_all_kswapds(unsigned int order, gfp_t gfp_mask,\n--\nmm/page_alloc.c-4522-\t\t\tcontinue;\nmm/page_alloc.c:4523:\t\twakeup_kswapd(zone, gfp_mask, reclaim_order, highest_zoneidx);\nmm/page_alloc.c-4524-\t\tlast_pgdat = zone-\u003ezone_pgdat;\n--\nmm/vmscan.c=192=int current_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;\n--\nmm/vmscan.c-196- * Per-node 10-second rolling average of idle CPU cores.\nmm/vmscan.c:197: * Sampled once per second; used by wakeup_kswapd() to decide how many\nmm/vmscan.c-198- * kswapd threads to wake without adding contention on a loaded node.\n--\nmm/vmscan.c=7774=static int kswapd(void *p)\n--\nmm/vmscan.c-7854- */\nmm/vmscan.c:7855:void wakeup_kswapd(struct zone *zone, gfp_t gfp_flags, int order,\nmm/vmscan.c-7856-\t\t enum zone_type highest_zoneidx)\n--\nmm/vmscan.c-7896-\nmm/vmscan.c:7897:\ttrace_mm_vmscan_wakeup_kswapd(pgdat-\u003enode_id, highest_zoneidx, order,\nmm/vmscan.c-7898-\t\t\t\t gfp_flags);\n"
} |
| 21/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 28479 output: 33 thoughts: 506 Thoughts: **Analyzing the Impact of the Patch on the Kernel's Kswapd Implementation** Okay, let's break this down. First, I'm tracing where `wakeup_kswapd` is called. It seems to be triggered from several key areas: the slow path of page allocation (`__alloc_pages_slowpath`), during migration (`migrate_misplaced_folio_prepare`), and within `rmqueue` when `ZONE_BOOSTED_WATERMARK` is active. Since this patch modifies `wakeup_kswapd`, those call sites are definitely worth a closer look. Now, I'm checking the helper functions. I see `kswapd_idle_avg_sample_workfn` is static and called through a workqueue, so I'll keep that in mind. `kswapd_idle_avg_cpu_online` and `kswapd_idle_avg_cpu_offline` are also static, but hooked into cpuhp, which is interesting. The patch also directly modifies `kswapd_run`, `kswapd_stop`, `kswapd_test_hopeless`, `kswapd_try_to_sleep`, and `allow_direct_reclaim`. That's a lot of core functionality being touched. Next, I'm looking at other related symbols. `kswapd_is_running` in `mm/compaction.c` is static, but it's called from `isolate_migratepages`, so it's still relevant. `update_max_kswapds_per_node` in `mm/vmscan.c` is non-static, which is good to know, and `max_kswapds_per_node_sysctl_handler` in `mm/page_alloc.c` is static, though it's a sysctl handler for `max_kswapds_per_node`. I need to ensure the patch doesn't break our sysctl access to that value. I'm also paying attention to the exposed symbols. `wakeup_kswapd`, `update_max_kswapds_per_node`, `kswapd_test_hopeless`, `kswapd_run`, and `kswapd_stop` are visible outside their respective files. This means external modules or parts of the kernel could potentially depend on the behavior of these functions, so I must examine the patch's changes carefully. Wait, what about `kswapd` itself? Oh, right, `static int kswapd(void *p)` is the main kthread function. I nearly overlooked it! I need to ensure that whatever this patch does doesn't introduce any regressions within the core kswapd thread. Now, let's examine the definition of `kswapd` in `mm/vmscan.c`. |
| 22/2 | 2026/10/09 22:42 | tool | grepper |
0mArgs:{
"Expression": "^static int kswapd",
"PathPrefixes": [
"mm/vmscan.c"
]
}Results: {
"Output": "mm/vmscan.c=217=static bool *kswapd_prev_idle_valid;\nmm/vmscan.c-218-\nmm/vmscan.c:219:static int kswapd_node_first_online_cpu(int nid)\nmm/vmscan.c-220-{\n--\nmm/vmscan.c=332=static void kswapd_idle_avg_stop_node(int nid)\n--\nmm/vmscan.c-350-\nmm/vmscan.c:351:static int kswapd_idle_avg_cpu_online(unsigned int cpu)\nmm/vmscan.c-352-{\n--\nmm/vmscan.c-365-\nmm/vmscan.c:366:static int kswapd_idle_avg_cpu_offline(unsigned int cpu)\nmm/vmscan.c-367-{\n--\nmm/vmscan.c=381=static u32 kswapd_node_avg_idle_cores(int nid)\n--\nmm/vmscan.c-388-\nmm/vmscan.c:389:static int kswapd_node_cpu_count(int nid)\nmm/vmscan.c-390-{\n--\nmm/vmscan.c=7675=static void kswapd_try_to_sleep(pg_data_t *pgdat, int alloc_order, int reclaim_order,\n--\nmm/vmscan.c-7773- */\nmm/vmscan.c:7774:static int kswapd(void *p)\nmm/vmscan.c-7775-{\n"
} |
| 23/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 30037 output: 35 thoughts: 21 Thoughts: Here's a summary of my thought process, tailored for an expert like yourself: **Diving into `kswapd` Context** Alright, time to get to the heart of the matter. I need to understand what's triggering `kswapd` and the surrounding context, so I'm heading straight for the source code. My immediate focus is going to be `mm/vmscan.c`. That's where all the magic related to the page scanner, and by extension, `kswapd`, resides. I want to trace the calls *to* `kswapd` and, crucially, understand the situations that cause it to be invoked. I’m thinking about looking for function calls to `kswapd_run`, or similar, and following those execution paths to map out the triggering conditions. I know `vmscan.c` is the nerve center, but I also expect to see relevant interactions with other memory management components, like the page allocator and the reclaim path. Let's see what we can dig up. |
| 24/2 | 2026/10/09 22:42 | tool | read-file |
0mArgs:{
"File": "mm/vmscan.c",
"FirstLine": 7770,
"LineCount": 80
}Results: {
"Contents": "7770:\t *\n7771:\t * If there are applications that are active memory-allocators\n7772:\t * (most normal use), this basically shouldn't matter.\n7773:\t */\n7774:\tstatic int kswapd(void *p)\n7775:\t{\n7776:\t\tunsigned int alloc_order, reclaim_order;\n7777:\t\tunsigned int highest_zoneidx = MAX_NR_ZONES - 1;\n7778:\t\tpg_data_t *pgdat = (pg_data_t *)p;\n7779:\t\tstruct task_struct *tsk = current;\n7780:\t\n7781:\t\t/*\n7782:\t\t * Tell the memory management that we're a \"memory allocator\",\n7783:\t\t * and that if we need more memory we should get access to it\n7784:\t\t * regardless (see \"__alloc_pages()\"). \"kswapd\" should\n7785:\t\t * never get caught in the normal page freeing logic.\n7786:\t\t *\n7787:\t\t * (Kswapd normally doesn't need memory anyway, but sometimes\n7788:\t\t * you need a small amount of memory in order to be able to\n7789:\t\t * page out something else, and this flag essentially protects\n7790:\t\t * us from recursively trying to free more memory as we're\n7791:\t\t * trying to free the first piece of memory in the first place).\n7792:\t\t */\n7793:\t\ttsk-\u003eflags |= PF_MEMALLOC | PF_KSWAPD;\n7794:\t\tset_freezable();\n7795:\t\n7796:\t\tWRITE_ONCE(pgdat-\u003ekswapd_order, 0);\n7797:\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, MAX_NR_ZONES);\n7798:\t\tatomic_set(\u0026pgdat-\u003enr_writeback_throttled, 0);\n7799:\t\tfor ( ; ; ) {\n7800:\t\t\tbool was_frozen;\n7801:\t\n7802:\t\t\talloc_order = reclaim_order = READ_ONCE(pgdat-\u003ekswapd_order);\n7803:\t\t\thighest_zoneidx = kswapd_highest_zoneidx(pgdat,\n7804:\t\t\t\t\t\t\t\thighest_zoneidx);\n7805:\t\n7806:\tkswapd_try_sleep:\n7807:\t\t\tkswapd_try_to_sleep(pgdat, alloc_order, reclaim_order,\n7808:\t\t\t\t\t\thighest_zoneidx);\n7809:\t\n7810:\t\t\t/* Read the new order and highest_zoneidx */\n7811:\t\t\talloc_order = READ_ONCE(pgdat-\u003ekswapd_order);\n7812:\t\t\thighest_zoneidx = kswapd_highest_zoneidx(pgdat,\n7813:\t\t\t\t\t\t\t\thighest_zoneidx);\n7814:\t\t\tWRITE_ONCE(pgdat-\u003ekswapd_order, 0);\n7815:\t\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, MAX_NR_ZONES);\n7816:\t\n7817:\t\t\tif (kthread_freezable_should_stop(\u0026was_frozen))\n7818:\t\t\t\tbreak;\n7819:\t\n7820:\t\t\t/*\n7821:\t\t\t * We can speed up thawing tasks if we don't call balance_pgdat\n7822:\t\t\t * after returning from the refrigerator\n7823:\t\t\t */\n7824:\t\t\tif (was_frozen)\n7825:\t\t\t\tcontinue;\n7826:\t\n7827:\t\t\t/*\n7828:\t\t\t * Reclaim begins at the requested order but if a high-order\n7829:\t\t\t * reclaim fails then kswapd falls back to reclaiming for\n7830:\t\t\t * order-0. If that happens, kswapd will consider sleeping\n7831:\t\t\t * for the order it finished reclaiming at (reclaim_order)\n7832:\t\t\t * but kcompactd is woken to compact for the original\n7833:\t\t\t * request (alloc_order).\n7834:\t\t\t */\n7835:\t\t\ttrace_mm_vmscan_kswapd_wake(pgdat-\u003enode_id, highest_zoneidx,\n7836:\t\t\t\t\t\t\talloc_order);\n7837:\t\t\treclaim_order = balance_pgdat(pgdat, alloc_order,\n7838:\t\t\t\t\t\t\thighest_zoneidx);\n7839:\t\t\tif (reclaim_order \u003c alloc_order)\n7840:\t\t\t\tgoto kswapd_try_sleep;\n7841:\t\t}\n7842:\t\n7843:\t\ttsk-\u003eflags \u0026= ~(PF_MEMALLOC | PF_KSWAPD);\n7844:\t\n7845:\t\treturn 0;\n7846:\t}\n7847:\t\n7848:\t/*\n7849:\t * A zone is low on free memory or too fragmented for high-order memory. If\n"
} |
| 25/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31603 output: 36 thoughts: 378 Thoughts: **Analyzing Potential Concurrency Issues in kswapd()** Okay, let's dive into this `kswapd()` code snippet. I'm immediately noticing a potential race condition here. I'm seeing that when one kswapd thread wakes up, it reads `pgdat->kswapd_order` and then, crucially, clears it with `WRITE_ONCE(pgdat->kswapd_order, 0)`. The subsequent kswapd thread wakes up and reads this cleared value! This looks like a definite point for concurrent read/writes to potentially wreak havoc. Further, consider that the code also uses `atomic_set(&pgdat->nr_writeback_throttled, 0);` alongside `balance_pgdat(pgdat, alloc_order, highest_zoneidx)`. Multiple kswapd threads executing concurrently on the same `pgdat` could race within `balance_pgdat`, potentially leading to issues with wait queues and `kswapd_try_to_sleep`. This is definitely something that needs thorough investigation, and potentially fuzzing, to uncover potential edge cases. I can see this being heavily impactful. Now, let's explore the configuration space. What options influence this area? I need to understand the configuration space to see if any toggles will let me tune the system into exposing these race conditions. I'm thinking about the `max_kswapds_per_node` sysctl setting. Okay, that looks promising. Specifically, `page_alloc_sysctl_table` from `mm/page_alloc.c` contains the relevant variables. I need to see if `page_alloc_sysctl_table` itself is guarded by any conditional compilation. I'll need to check the definition and registration of this table within `mm/page_alloc.c` for any clues. This feels like a good starting point to control the number of kswapd threads and potentially influence the probability of hitting this race. |
| 26/2 | 2026/10/09 22:42 | tool | grepper |
0mArgs:{
"Expression": "page_alloc_sysctl_table",
"PathPrefixes": [
"mm/page_alloc.c"
]
}Results: {
"Output": "mm/page_alloc.c=6912=static int percpu_pagelist_high_fraction_sysctl_handler(const struct ctl_table *table,\n--\nmm/page_alloc.c-6952-\nmm/page_alloc.c:6953:static const struct ctl_table page_alloc_sysctl_table[] = {\nmm/page_alloc.c-6954-\t{\n--\nmm/page_alloc.c=7042=void __init page_alloc_sysctl_init(void)\nmm/page_alloc.c-7043-{\nmm/page_alloc.c:7044:\tregister_sysctl_init(\"vm\", page_alloc_sysctl_table);\nmm/page_alloc.c-7045-}\n"
} |
| 27/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32641 output: 36 thoughts: 31 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Investigating a Conditional Compilation Issue in the Kernel** Alright, so I need to quickly determine if there's a conditional compilation directive, likely an `#ifdef`, surrounding the `page_alloc_sysctl_table`. My initial hunch is that it might be conditionally included based on some configuration option. Line 6953 seems like a reasonable starting point, given the context. I'll need to rapidly examine the surrounding lines to understand the dependency, if one exists. This will likely provide insight into whether the symbol is available based on the current build configuration. Knowing this will immediately inform subsequent analysis related to memory management and sysctl parameters. |
| 28/2 | 2026/10/09 22:42 | tool | read-file |
0mArgs:{
"File": "mm/page_alloc.c",
"FirstLine": 6945,
"LineCount": 40
}Results: {
"Contents": "6945:\t\n6946:\t\tfor_each_populated_zone(zone)\n6947:\t\t\tzone_set_pageset_high_and_batch(zone, 0);\n6948:\tout:\n6949:\t\tmutex_unlock(\u0026pcp_batch_high_lock);\n6950:\t\treturn ret;\n6951:\t}\n6952:\t\n6953:\tstatic const struct ctl_table page_alloc_sysctl_table[] = {\n6954:\t\t{\n6955:\t\t\t.procname\t= \"min_free_kbytes\",\n6956:\t\t\t.data\t\t= \u0026min_free_kbytes,\n6957:\t\t\t.maxlen\t\t= sizeof(min_free_kbytes),\n6958:\t\t\t.mode\t\t= 0644,\n6959:\t\t\t.proc_handler\t= min_free_kbytes_sysctl_handler,\n6960:\t\t\t.extra1\t\t= SYSCTL_ZERO,\n6961:\t\t},\n6962:\t\t{\n6963:\t\t\t.procname\t= \"watermark_boost_factor\",\n6964:\t\t\t.data\t\t= \u0026watermark_boost_factor,\n6965:\t\t\t.maxlen\t\t= sizeof(watermark_boost_factor),\n6966:\t\t\t.mode\t\t= 0644,\n6967:\t\t\t.proc_handler\t= proc_dointvec_minmax,\n6968:\t\t\t.extra1\t\t= SYSCTL_ZERO,\n6969:\t\t},\n6970:\t\t{\n6971:\t\t\t.procname = \"max_kswapds_per_node\",\n6972:\t\t\t.data = \u0026max_kswapds_per_node,\n6973:\t\t\t.maxlen = sizeof(max_kswapds_per_node),\n6974:\t\t\t.mode = 0644,\n6975:\t\t\t.proc_handler = max_kswapds_per_node_sysctl_handler,\n6976:\t\t\t.extra1 = SYSCTL_ONE,\n6977:\t\t\t.extra2 = \u0026max_kswapd_threads,\n6978:\t\t},\n6979:\t\t{\n6980:\t\t\t.procname\t= \"watermark_scale_factor\",\n6981:\t\t\t.data\t\t= \u0026watermark_scale_factor,\n6982:\t\t\t.maxlen\t\t= sizeof(watermark_scale_factor),\n6983:\t\t\t.mode\t\t= 0644,\n6984:\t\t\t.proc_handler\t= watermark_scale_factor_sysctl_handler,\n"
} |
| 29/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 33540 output: 28 thoughts: 975 Thoughts: **Investigating kswapd and NUMA Configuration Dependencies** Okay, so I was looking into `page_alloc_sysctl_table` and noticed there aren't any ifdefs there. But, that's not the whole story because the general sysctl registration relies on `CONFIG_SYSCTL`, which is, of course, pretty standard. The more interesting question is, what about `CONFIG_NUMA`? I started digging into `mm/vmscan.c` and found something crucial. The idle average sampling logic is *inside* an `#ifdef CONFIG_NUMA` block. Taking a look at what happens without `CONFIG_NUMA`, `kswapd_node_avg_idle_cores(nid)` simply returns 0. Consequently, in `wakeup_kswapd()`, when `avg_idle_cores` is 0, the `min3` function returns 0. The subsequent `max(1, 0)` then always results in `threads_to_wake` being 1! That's the key. This implies that without `CONFIG_NUMA`, you'll *always* have only one kswapd thread being woken up, effectively bypassing the logic for multiple kswapd threads. In addition, the code for `kswapd_idle_avg_init()`, `kswapd_idle_avg_sample_workfn`, the delayed work, and the CPU hotplug hooks are completely compiled out if `CONFIG_NUMA` is not enabled. So it is absolutely *critical* to enable `CONFIG_NUMA` if you want to test the kswapd thread wake-up logic with multiple threads and all of the `kswapd_idle_avg_*` code. This is very important. Now, let's see if there are other relevant configurations. How about `CONFIG_LRU_GEN`? I've found it being used in the functions `try_to_inc_max_seq` and `scan_folios`. These are Multi-Gen LRU functions. Time to look into `mm/Kconfig` for the configuration. |
| 30/2 | 2026/10/09 22:42 | tool | grepper |
0mArgs:{
"Expression": "config LRU_GEN",
"PathPrefixes": [
"mm/"
]
}Results: {
"Output": "mm/Kconfig=1393=config PTE_MARKER_UFFD_WP\n--\nmm/Kconfig-1403-# multi-gen LRU {\nmm/Kconfig:1404:config LRU_GEN\nmm/Kconfig-1405-\tbool \"Multi-Gen LRU\"\n--\nmm/Kconfig-1412-\nmm/Kconfig:1413:config LRU_GEN_ENABLED\nmm/Kconfig-1414-\tbool \"Enable by default\"\n--\nmm/Kconfig-1418-\nmm/Kconfig:1419:config LRU_GEN_STATS\nmm/Kconfig-1420-\tbool \"Full stats for debugging\"\n--\nmm/Kconfig-1427-\nmm/Kconfig:1428:config LRU_GEN_WALKS_MMU\nmm/Kconfig-1429-\tdef_bool y\n"
} |
| 31/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 35228 output: 26 thoughts: 312 Thoughts: **Analyzing Memory Hotplug and Compaction Dependencies** Okay, let's see about `CONFIG_MEMORY_HOTPLUG`. I need to figure out if it's actually *required* here. I'll start by checking `mm/vmscan.c` since that's a key area for memory management. Hmm, I see `mem_hotplug_begin()` and `mem_hotplug_done()` being used, and also `pgdat_kswapd_lock()`—that's interesting. Now, I should double-check the definition in `include/linux/mmzone.h`. Yes, there it is: the `kswapd_lock` mutex is conditionally compiled! It's *only* defined if `CONFIG_MEMORY_HOTPLUG` is enabled. That's a crucial clue. And let's not forget the other definition of `pgdat_kswapd_lock` in `include/linux/memory_hotplug.h`. Okay, `pgdat_kswapd_lock` is also a no-op if hotplug is disabled! So if `CONFIG_MEMORY_HOTPLUG` is not defined, then that function is basically empty. Now I should look at `CONFIG_COMPACTION`. Over in `mm/compaction.c`, I see `kswapd_is_running(pg_data_t *pgdat)`. I'll go check the Kconfig file to look at the interactions between these two config options. |
| 32/2 | 2026/10/09 22:42 | tool | grepper |
0mArgs:{
"Expression": "config COMPACTION",
"PathPrefixes": [
"mm/"
]
}Results: {
"Output": "mm/Kconfig=630=config BALLOON_MIGRATION\n--\nmm/Kconfig-641-# support for memory compaction\nmm/Kconfig:642:config COMPACTION\nmm/Kconfig-643-\tbool \"Allow for memory compaction\"\n"
} |
| 33/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 35960 output: 19 thoughts: 15 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Analyzing `CONFIG_NUMA` Dependencies in Kconfig** Right, let's dive into the Kconfig for `CONFIG_NUMA`. I need to ensure a clean picture of its dependencies and implications. This isn't just about the immediate `depends on` clauses; I want to understand the *entire* dependency tree, upstream and downstream. Specifically, I'm thinking: are there any subtle conflicts or unexpected interactions hiding within the nested relationships? Any potential for accidentally enabling NUMA when it's not truly supported by the hardware or desired in a particular configuration? We know Kconfig can be tricky, so let's check for any obscure dependencies which might trip us up, or prevent the correct settings from being applied in the first place. This requires a meticulous walkthrough, paying close attention to any conditional logic or environment variables that might influence the ultimate outcome of `CONFIG_NUMA`'s setting. Gotta ensure the kernel behaves *exactly* as expected. Now, where to begin... |
| 34/2 | 2026/10/09 22:42 | tool | grepper |
0mArgs:{
"Expression": "config NUMA\\b"
}Results: {
"Output": "arch/arm64/Kconfig=1656=config HOTPLUG_CPU\n--\narch/arm64/Kconfig-1663-# Common NUMA Features\narch/arm64/Kconfig:1664:config NUMA\narch/arm64/Kconfig-1665-\tbool \"NUMA Memory Allocation and Scheduler Support\"\n--\narch/loongarch/Kconfig=544=config NR_CPUS\n--\narch/loongarch/Kconfig-552-\narch/loongarch/Kconfig:553:config NUMA\narch/loongarch/Kconfig-554-\tbool \"NUMA Support\"\n--\narch/mips/Kconfig=2596=config ARCH_SPARSEMEM_ENABLE\n--\narch/mips/Kconfig-2598-\narch/mips/Kconfig:2599:config NUMA\narch/mips/Kconfig-2600-\tbool \"NUMA Support\"\n--\narch/powerpc/Kconfig=758=config IRQ_ALL_CPUS\n--\narch/powerpc/Kconfig-766-\narch/powerpc/Kconfig:767:config NUMA\narch/powerpc/Kconfig-768-\tbool \"NUMA Memory Allocation and Scheduler Support\"\n--\narch/riscv/Kconfig=481=endchoice\n--\narch/riscv/Kconfig-483-# Common NUMA Features\narch/riscv/Kconfig:484:config NUMA\narch/riscv/Kconfig-485-\tbool \"NUMA Memory Allocation and Scheduler Support\"\n--\narch/s390/Kconfig=547=config HOTPLUG_CPU\n--\narch/s390/Kconfig-549-\narch/s390/Kconfig:550:config NUMA\narch/s390/Kconfig-551-\tbool \"NUMA support\"\n--\narch/sh/mm/Kconfig=102=config VSYSCALL\n--\narch/sh/mm/Kconfig-114-\narch/sh/mm/Kconfig:115:config NUMA\narch/sh/mm/Kconfig-116-\tbool \"Non-Uniform Memory Access (NUMA) Support\"\n--\narch/sparc/Kconfig=254=config GENERIC_LOCKBREAK\n--\narch/sparc/Kconfig-258-\narch/sparc/Kconfig:259:config NUMA\narch/sparc/Kconfig-260-\tbool \"NUMA support\"\n--\narch/x86/Kconfig=1492=config AMD_MEM_ENCRYPT\n--\narch/x86/Kconfig-1508-# Common NUMA Features\narch/x86/Kconfig:1509:config NUMA\narch/x86/Kconfig-1510-\tbool \"NUMA Memory Allocation and Scheduler Support\"\n"
} |
| 35/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 36946 output: 37 thoughts: 21 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Diving into the x86 Kernel Configuration** Alright, let's get down to brass tacks. My initial instinct, given this context, is to zero in on `arch/x86/Kconfig` around line 1509. I'm immediately thinking there's something specific I'm trying to configure or troubleshoot within the x86 architecture's kernel configuration. It's not a generic issue; I've narrowed the scope. Something specific is triggering this precise location within the configuration files, and likely, it's something relatively intricate. I know that particular region of the Kconfig file holds clues, and that's the starting point. I could be looking for a particular feature flag, a hardware-related option, or even a dependency definition. It could be related to a specific CPU model, a security feature, or a device driver configuration. I have a hunch about what I am looking for; let's get to it. |
| 36/2 | 2026/10/09 22:42 | tool | read-file |
0mArgs:{
"File": "arch/x86/Kconfig",
"FirstLine": 1505,
"LineCount": 20
}Results: {
"Contents": "1505:\t\t This requires an AMD processor that supports Secure Memory\n1506:\t\t Encryption (SME).\n1507:\t\n1508:\t# Common NUMA Features\n1509:\tconfig NUMA\n1510:\t\tbool \"NUMA Memory Allocation and Scheduler Support\"\n1511:\t\tdepends on SMP\n1512:\t\tdepends on X86_64\n1513:\t\tselect USE_PERCPU_NUMA_NODE_ID\n1514:\t\tselect OF_NUMA if OF\n1515:\t\thelp\n1516:\t\t Enable NUMA (Non-Uniform Memory Access) support.\n1517:\t\n1518:\t\t The kernel will try to allocate memory used by a CPU on the\n1519:\t\t local memory controller of the CPU and add some more\n1520:\t\t NUMA awareness to the kernel.\n1521:\t\n1522:\t\t For 64-bit this is recommended if the system is Intel Core i7\n1523:\t\t (or later), AMD Opteron, or EM64T NUMA.\n1524:\t\n"
} |
| 37/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37525 output: 32 thoughts: 48 Thoughts: Okay, here's my interpretation of those thoughts, tailored for an expert audience, presented in the first person: **Kernel Configuration Dependencies: A Quick Mental Walkthrough** Alright, so I'm thinking about x86_64 system configuration, particularly the interplay of features. It's pretty straightforward: I know that `NUMA` on x86_64 systems has a fundamental dependency on both `SMP` (Symmetric Multi-Processing) and the `X86_64` architecture being enabled. That's just a given, the x86_64 architecture and multi-processor support are prerequisites. Now, I'm trying to recall the dependencies or implications of `MEMORY_HOTPLUG`. My immediate thought is, "Where's the definitive source?" Right, that's defined in the memory management section's Kconfig file (`mm/Kconfig`). I'll need to check the Kconfig file for any interdependencies or conditions. Let's take a look. |
| 38/2 | 2026/10/09 22:42 | tool | grepper |
0mArgs:{
"Expression": "config MEMORY_HOTPLUG\\b",
"PathPrefixes": [
"mm/"
]
}Results: {
"Output": "mm/Kconfig=510=config ARCH_ENABLE_MEMORY_HOTPLUG\n--\nmm/Kconfig-513-# eventually, we can have this option just 'select SPARSEMEM'\nmm/Kconfig:514:menuconfig MEMORY_HOTPLUG\nmm/Kconfig-515-\tbool \"Memory hotplug\"\n"
} |
| 39/2 | 2026/10/09 22:42 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37912 output: 32 thoughts: 17 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Analyzing `mm/Kconfig` - Line 514 and Surrounding Context** Alright, let's dive in. I need to zero in on line 514 within the `mm/Kconfig` file. I'm assuming, given the context, that a configuration option or dependency is likely defined or referenced around that point. My goal is to quickly grasp the implications of whatever's happening there. Is it a new memory management feature being introduced? A dependency check on an existing feature? Perhaps a conditional compilation statement affecting how a memory management subsystem behaves? My immediate thoughts are, where does that Kconfig fragment fit in the broader scheme? What is the *overall* context of the module, and what is being configured? Line 514 is the *starting point*. I am anticipating that it's just the tip of the iceberg, and it is surrounded by a bunch of `config` blocks, `depends on` statements, `select` statements, or some other crucial Kconfig directives. The goal is a quick high-level mental map of the configuration space in that area. I’ll be looking for potential conflicts, dependencies that might not be immediately obvious, or areas that could benefit from refactoring to improve readability or reduce duplication. My immediate concern is that a change here could cause downstream integration or implementation pain. I need to quickly assess the potential impact of any changes at or around that line. That's the priority now. |
| 40/2 | 2026/10/09 22:43 | tool | read-file |
0mArgs:{
"File": "mm/Kconfig",
"FirstLine": 510,
"LineCount": 25
}Results: {
"Contents": " 510:\tconfig ARCH_ENABLE_MEMORY_HOTPLUG\n 511:\t\tbool\n 512:\t\n 513:\t# eventually, we can have this option just 'select SPARSEMEM'\n 514:\tmenuconfig MEMORY_HOTPLUG\n 515:\t\tbool \"Memory hotplug\"\n 516:\t\tselect MEMORY_ISOLATION\n 517:\t\tdepends on SPARSEMEM_VMEMMAP\n 518:\t\tdepends on ARCH_ENABLE_MEMORY_HOTPLUG\n 519:\t\tdepends on 64BIT\n 520:\t\tselect NUMA_KEEP_MEMINFO if NUMA\n 521:\t\n 522:\tif MEMORY_HOTPLUG\n 523:\t\n 524:\tchoice\n 525:\t\tprompt \"Memory Hotplug Default Online Type\"\n 526:\t\tdefault MHP_DEFAULT_ONLINE_TYPE_OFFLINE\n 527:\t\thelp\n 528:\t\t Default memory type for hotplugged memory.\n 529:\t\n 530:\t\t This option sets the default policy setting for memory hotplug\n 531:\t\t onlining policy (/sys/devices/system/memory/auto_online_blocks) which\n 532:\t\t determines what happens to newly added memory regions. Policy setting\n 533:\t\t can always be changed at runtime.\n 534:\t\n"
} |
| 41/2 | 2026/10/09 22:43 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 38647 output: 279 thoughts: 2388 Thoughts: **Focus Symbol Selection and Reasoning for Memory Management Patch** Alright, let's break down this patch and pinpoint the optimal focus symbols for fuzzing. We're dealing with enhancements to the memory management subsystem, specifically related to NUMA, multi-Gen LRU, and memory hotplug. Key configurations involved are NUMA, LRU_GEN, and MEMORY_HOTPLUG, and the goal is to identify functions that specifically implement these new features or act as entry points to them, while avoiding the overly-ubiquitous. Firstly, the dependencies. `NUMA` is critical, enabling idle core calculations, cpuhp callbacks for kswapd, and multi-thread wakeup calculations. `LRU_GEN` is essential for changes in `try_to_inc_max_seq` and `scan_folios`. Finally, `MEMORY_HOTPLUG` enables `kswapd_lock` protection and the hotplugging routines `kswapd_run` and `kswapd_stop`. Now, the function analysis. We have several candidates: `max_kswapds_per_node_sysctl_handler` is a sysctl handler for writing `max_kswapds_per_node`. `update_max_kswapds_per_node` reconfigures kswapd threads. `wakeup_kswapd` wakes up kswapd threads based on idle core averages. `kswapd_run` starts new kswapd threads and `kswapd_stop` stops them. `kswapd_test_hopeless` checks for kswapd failures. `kswapd_try_to_sleep` and `kswapd` are part of the main thread logic, along with `kswapd_idle_avg_sample_workfn`. Finally, `scan_folios` and `try_to_inc_max_seq` are also modified. My thought is that the non-static functions are good targets, and while static functions can be inlined, the ones here can't be due to their use as work functions, and their use in `kthread_run` and sysctl entries. Let's carefully verify the choices for FocusSymbols: `update_max_kswapds_per_node`, `max_kswapds_per_node_sysctl_handler`, `wakeup_kswapd`, `kswapd_idle_avg_sample_workfn`, `kswapd`, `kswapd_run`, and `kswapd_stop`. `wakeup_kswapd` is only triggered during specific memory pressure conditions, like low watermarks or misplaced folios, therefore it is not ubiquitous. This is critical. We must verify that these symbols are not ubiquitous. They all look focused. `kswapd` is a background thread but is not called for every system call - it's specifically for the page reclaim loop. When a user process runs out of memory or needs to enter the slow path for page allocation, the calls in question are executed in the context of the calling user task, or in the context of a kernel thread. The EnableConfigs are clear: `NUMA` is essential for per-node idle CPU averages, hotplug callbacks, and calculations. `LRU_GEN` deals with the Multi-Gen LRU logic. `MEMORY_HOTPLUG` enables `kswapd_lock` and the hotplug routines. The reasoning is straightforward: this patch modifies core memory management functionality through sysctl, memory pressure, and allocation which could expose concurrency issues, race conditions, or deadlocks. These are all reachable code paths. We can conclude that it's worth fuzzing these features. The selected FocusSymbols are targeted and not ubiquitous, and we have the necessary supporting configurations. Let's move forward! |
| 42/2 | 2026/10/09 22:43 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"NUMA",
"LRU_GEN",
"MEMORY_HOTPLUG"
],
"FocusSymbols": [
"max_kswapds_per_node_sysctl_handler",
"update_max_kswapds_per_node",
"wakeup_kswapd",
"kswapd_idle_avg_sample_workfn",
"kswapd_run",
"kswapd_stop"
],
"Reasoning": "The patch adds multi-kswapd thread support per NUMA node, dynamic scaling of kswapd threads via the /proc/sys/vm/max_kswapds_per_node sysctl, per-node rolling average idle CPU sampling via delayed work and CPU hotplug hooks, altered wait queue wakeups (prepare_to_wait_exclusive and wake_up_nr), and changes to Multi-Gen LRU scanning. These modifications are in core memory management (mm/vmscan.c, mm/page_alloc.c, mm/compaction.c) and are reachable in virtualized environments via memory allocation/pressure and sysctl writes. Concurrent multi-thread reclaim and dynamic thread stopping/starting introduce potential race conditions and locking subtleties that warrant fuzzing.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"NUMA",
"LRU_GEN",
"MEMORY_HOTPLUG"
],
"FocusSymbols": [
"max_kswapds_per_node_sysctl_handler",
"update_max_kswapds_per_node",
"wakeup_kswapd",
"kswapd_idle_avg_sample_workfn",
"kswapd_run",
"kswapd_stop"
],
"Reasoning": "The patch adds multi-kswapd thread support per NUMA node, dynamic scaling of kswapd threads via the /proc/sys/vm/max_kswapds_per_node sysctl, per-node rolling average idle CPU sampling via delayed work and CPU hotplug hooks, altered wait queue wakeups (prepare_to_wait_exclusive and wake_up_nr), and changes to Multi-Gen LRU scanning. These modifications are in core memory management (mm/vmscan.c, mm/page_alloc.c, mm/compaction.c) and are reachable in virtualized environments via memory allocation/pressure and sysctl writes. Concurrent multi-thread reclaim and dynamic thread stopping/starting introduce potential race conditions and locking subtleties that warrant fuzzing.",
"WorthFuzzing": true
} |
| 43/1 | 2026/10/09 22:43 | action | If |
1mArgs:{
"WorthFuzzing": true
}Results: null |
| 44/2 | 2026/10/09 22:43 | agent | kmsan-evaluator |
1mModel:core Args: null Results: {
"KMSANReasoning": "The patch implements support for multiple kswapd threads per NUMA node, an idle CPU rolling average mechanism to scale thread wakeup, and a sysctl interface (`max_kswapds_per_node`). \n\nA dedicated KMSAN session is not needed for the following reasons:\n1. No uninitialized memory exposure to userspace: The sysctl handler uses standard `proc_dointvec_minmax` with a fully initialized integer, and there are no info-leak vectors or uninitialized structures copied to user space.\n2. Safe data structure allocations: The NUMA idle-tracking arrays (`kswapd_idle_avg`, `kswapd_prev_idle`, `kswapd_prev_idle_valid`) are allocated using `kcalloc`, ensuring all fields and elements are zero-initialized. The `kswapd` thread array in `pglist_data` is part of node data zeroed at boot/hotplug.\n3. Deterministic initialization of local variables: All local variables (including `running` in `kswapd_is_running` which is unconditionally assigned in the first iteration of the fixed 16-iteration loop) are assigned before being read.\n4. Bug profile: The risks introduced by this change relate to concurrency (multiple kswapd threads contending on reclaim paths), task lifecycle/locking (`kthread_run`/`kthread_stop`), and bounds checks on sysctl parameters—issues that are covered by KASAN, LOCKDEP, and standard kernel debug facilities, rather than KMSAN uninitialized memory tracking.",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 3135ebe1de131a892c48dfa2c20fa9c02d6f74fc
Author: syz-cluster <triage@syzkaller.com>
Date: Fri Oct 9 22:39:58 2026 +0000
syz-cluster: applied patch under review
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index 94f9c3ff54160..3f0ae2b08835c 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -62,6 +62,8 @@
*/
#define PAGE_ALLOC_COSTLY_ORDER 3
+#define MAX_KSWAPD_THREADS 16
+
#if !defined(CONFIG_HAVE_GIGANTIC_FOLIOS)
/*
* We don't expect any folios that exceed buddy sizes (and consequently
@@ -1529,7 +1531,8 @@ typedef struct pglist_data {
#ifdef CONFIG_MEMORY_HOTPLUG
struct mutex kswapd_lock;
#endif
- struct task_struct *kswapd; /* Protected by kswapd_lock */
+ /* Protected by kswapd_lock */
+ struct task_struct *kswapd[MAX_KSWAPD_THREADS];
int kswapd_order;
enum zone_type kswapd_highest_zoneidx;
diff --git a/include/trace/events/vmscan.h b/include/trace/events/vmscan.h
index b4bf7b8def1f5..6fd87d811e984 100644
--- a/include/trace/events/vmscan.h
+++ b/include/trace/events/vmscan.h
@@ -174,6 +174,38 @@ TRACE_EVENT(mm_vmscan_wakeup_kswapd,
show_gfp_flags(__entry->gfp_flags))
);
+TRACE_EVENT(mm_vmscan_kswapd_threads_to_wake,
+
+ TP_PROTO(int nid, int threads_to_wake, int threads_woken,
+ int node_cpus, int avg_idle_cores),
+
+ TP_ARGS(nid, threads_to_wake, threads_woken, node_cpus,
+ avg_idle_cores),
+
+ TP_STRUCT__entry(
+ __field(int, nid)
+ __field(int, threads_to_wake)
+ __field(int, threads_woken)
+ __field(int, node_cpus)
+ __field(int, avg_idle_cores)
+ ),
+
+ TP_fast_assign(
+ __entry->nid = nid;
+ __entry->threads_to_wake = threads_to_wake;
+ __entry->threads_woken = threads_woken;
+ __entry->node_cpus = node_cpus;
+ __entry->avg_idle_cores = avg_idle_cores;
+ ),
+
+ TP_printk("nid=%d threads_to_wake=%d threads_woken=%d node_cpus=%d avg_idle_cores=%d",
+ __entry->nid,
+ __entry->threads_to_wake,
+ __entry->threads_woken,
+ __entry->node_cpus,
+ __entry->avg_idle_cores)
+);
+
DECLARE_EVENT_CLASS(mm_vmscan_direct_reclaim_begin_template,
TP_PROTO(gfp_t gfp_flags, int order, struct mem_cgroup *memcg),
diff --git a/mm/compaction.c b/mm/compaction.c
index a049415512c67..5db7d94aefc23 100644
--- a/mm/compaction.c
+++ b/mm/compaction.c
@@ -2198,9 +2198,13 @@ static isolate_migrate_t isolate_migratepages(struct compact_control *cc)
static bool kswapd_is_running(pg_data_t *pgdat)
{
bool running;
-
+ int i;
pgdat_kswapd_lock(pgdat);
- running = pgdat->kswapd && task_is_running(pgdat->kswapd);
+ for (i = 0; i < MAX_KSWAPD_THREADS; ++i) {
+ running = pgdat->kswapd[i] && task_is_running(pgdat->kswapd[i]);
+ if (running)
+ break;
+ }
pgdat_kswapd_unlock(pgdat);
return running;
diff --git a/mm/internal.h b/mm/internal.h
index 38b1165212c94..1417c5b718ff2 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -717,10 +717,13 @@ void set_recommended_min_free_kbytes(void);
extern char * const zone_names[MAX_NR_ZONES];
+extern int max_kswapds_per_node;
+extern int max_kswapd_threads;
extern int min_free_kbytes;
extern int defrag_mode;
void setup_per_zone_wmarks(void);
+void update_max_kswapds_per_node(void);
void calculate_min_free_kbytes(void);
int __meminit init_per_zone_wmark_min(void);
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 12fac9084c483..d56273f8c7de5 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -277,6 +277,7 @@ int user_min_free_kbytes = -1;
static int watermark_boost_factor __read_mostly = 15000;
static int watermark_scale_factor = 10;
int defrag_mode;
+int max_kswapd_threads = MAX_KSWAPD_THREADS;
/* movable_zone is the "real" zone pages in ZONE_MOVABLE are taken from */
int movable_zone;
@@ -6777,6 +6778,22 @@ static int min_free_kbytes_sysctl_handler(const struct ctl_table *table, int wri
return 0;
}
+static int max_kswapds_per_node_sysctl_handler(const struct ctl_table *table,
+ int write, void __user *buffer,
+ size_t *length, loff_t *ppos)
+{
+ int rc;
+
+ rc = proc_dointvec_minmax(table, write, buffer, length, ppos);
+ if (rc)
+ return rc;
+
+ if (write)
+ update_max_kswapds_per_node();
+
+ return 0;
+}
+
static int watermark_scale_factor_sysctl_handler(const struct ctl_table *table, int write,
void *buffer, size_t *length, loff_t *ppos)
{
@@ -6950,6 +6967,15 @@ static const struct ctl_table page_alloc_sysctl_table[] = {
.proc_handler = proc_dointvec_minmax,
.extra1 = SYSCTL_ZERO,
},
+ {
+ .procname = "max_kswapds_per_node",
+ .data = &max_kswapds_per_node,
+ .maxlen = sizeof(max_kswapds_per_node),
+ .mode = 0644,
+ .proc_handler = max_kswapds_per_node_sysctl_handler,
+ .extra1 = SYSCTL_ONE,
+ .extra2 = &max_kswapd_threads,
+ },
{
.procname = "watermark_scale_factor",
.data = &watermark_scale_factor,
diff --git a/mm/vmscan.c b/mm/vmscan.c
index f11491ee9ed5c..3d7a1a56b16a1 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -32,6 +32,7 @@
#include <linux/rmap.h>
#include <linux/topology.h>
#include <linux/cpu.h>
+#include <linux/cpuhotplug.h>
#include <linux/cpuset.h>
#include <linux/compaction.h>
#include <linux/notifier.h>
@@ -59,12 +60,14 @@
#include <linux/mmu_notifier.h>
#include <linux/parser.h>
#include <linux/swap_ops.h>
+#include <linux/smp.h>
#include <asm/tlbflush.h>
#include <asm/div64.h>
#include <linux/swapops.h>
#include <linux/sched/sysctl.h>
+#include <linux/slab.h>
#include "internal.h"
#include "page_alloc.h"
@@ -181,6 +184,285 @@ struct scan_control {
struct reclaim_state reclaim_state;
};
+/*
+ * Max and current number of kswapd threads per node.
+ */
+#define DEF_MAX_KSWAPDS_PER_NODE 1
+int max_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;
+int current_kswapds_per_node = DEF_MAX_KSWAPDS_PER_NODE;
+
+#ifdef CONFIG_NUMA
+/*
+ * Per-node 10-second rolling average of idle CPU cores.
+ * Sampled once per second; used by wakeup_kswapd() to decide how many
+ * kswapd threads to wake without adding contention on a loaded node.
+ */
+#define KSWAPD_NUMA_AVG_WINDOW 10
+
+struct kswapd_node_idle_avg {
+ u32 samples[KSWAPD_NUMA_AVG_WINDOW];
+ u32 sum;
+ u8 idx;
+ u8 count;
+ u32 avg_idle_cores; /* rolling average, written only by the work item */
+ int nid;
+ int node_cpus; /* updated by cpuhp callbacks */
+ bool started;
+ u64 last_sample_ns;
+ struct delayed_work work;
+};
+
+static struct kswapd_node_idle_avg *kswapd_idle_avg;
+static u64 *kswapd_prev_idle;
+static bool *kswapd_prev_idle_valid;
+
+static int kswapd_node_first_online_cpu(int nid)
+{
+ return cpumask_first_and(cpumask_of_node(nid), cpu_online_mask);
+}
+
+static u64 kswapd_read_idle_cpu(int cpu)
+{
+ if (!cpu_online(cpu))
+ return 0;
+
+ return kcpustat_field(CPUTIME_IDLE, cpu);
+}
+
+static void kswapd_idle_avg_queue_work(struct kswapd_node_idle_avg *avg)
+{
+ int cpu;
+
+ cpu = kswapd_node_first_online_cpu(avg->nid);
+ if (cpu < nr_cpu_ids)
+ queue_delayed_work_on(cpu, system_wq, &avg->work, HZ);
+}
+
+static void kswapd_idle_avg_sample_workfn(struct work_struct *work)
+{
+ struct kswapd_node_idle_avg *avg;
+ u64 now_ns, elapsed_ns;
+ u64 idle_cores_milli = 0;
+ u32 idle_cores;
+ int cpu, node_cpus;
+
+ if (!kswapd_idle_avg || !kswapd_prev_idle || !kswapd_prev_idle_valid)
+ return;
+
+ avg = container_of(to_delayed_work(work), struct kswapd_node_idle_avg,
+ work);
+ node_cpus = 0;
+ now_ns = ktime_get_ns();
+
+ if (!avg->last_sample_ns) {
+ avg->last_sample_ns = now_ns;
+ for_each_cpu(cpu, cpumask_of_node(avg->nid)) {
+ if (!cpu_online(cpu))
+ continue;
+ kswapd_prev_idle[cpu] = kswapd_read_idle_cpu(cpu);
+ kswapd_prev_idle_valid[cpu] = true;
+ }
+ kswapd_idle_avg_queue_work(avg);
+ return;
+ }
+
+ elapsed_ns = now_ns - avg->last_sample_ns;
+ avg->last_sample_ns = now_ns;
+
+ if (!elapsed_ns) {
+ kswapd_idle_avg_queue_work(avg);
+ return;
+ }
+
+ for_each_cpu(cpu, cpumask_of_node(avg->nid)) {
+ u64 idle_now, idle_delta;
+
+ if (!cpu_online(cpu))
+ continue;
+
+ node_cpus++;
+ idle_now = kswapd_read_idle_cpu(cpu);
+
+ if (!kswapd_prev_idle_valid[cpu]) {
+ kswapd_prev_idle_valid[cpu] = true;
+ kswapd_prev_idle[cpu] = idle_now;
+ continue;
+ }
+
+ idle_delta = idle_now - kswapd_prev_idle[cpu];
+ kswapd_prev_idle[cpu] = idle_now;
+
+ idle_cores_milli += div_u64(min(idle_delta, elapsed_ns) * 1000,
+ elapsed_ns);
+ }
+
+ idle_cores = min_t(u32, DIV_ROUND_CLOSEST_ULL(idle_cores_milli, 1000),
+ node_cpus);
+
+ if (avg->count == KSWAPD_NUMA_AVG_WINDOW)
+ avg->sum -= avg->samples[avg->idx];
+ else
+ avg->count++;
+
+ avg->samples[avg->idx] = idle_cores;
+ avg->sum += idle_cores;
+ avg->idx = (avg->idx + 1) % KSWAPD_NUMA_AVG_WINDOW;
+ WRITE_ONCE(avg->avg_idle_cores, DIV_ROUND_CLOSEST(avg->sum, avg->count));
+
+ if (READ_ONCE(avg->started))
+ kswapd_idle_avg_queue_work(avg);
+}
+
+static void kswapd_idle_avg_start_node(int nid)
+{
+ struct kswapd_node_idle_avg *avg;
+
+ if (!kswapd_idle_avg || nid < 0 || nid >= nr_node_ids)
+ return;
+
+ avg = &kswapd_idle_avg[nid];
+ if (READ_ONCE(avg->started))
+ return;
+
+ WRITE_ONCE(avg->started, true);
+ avg->last_sample_ns = 0;
+ kswapd_idle_avg_queue_work(avg);
+}
+
+static void kswapd_idle_avg_stop_node(int nid)
+{
+ struct kswapd_node_idle_avg *avg;
+ int cpu;
+
+ if (!kswapd_idle_avg || nid < 0 || nid >= nr_node_ids)
+ return;
+
+ avg = &kswapd_idle_avg[nid];
+ if (!READ_ONCE(avg->started))
+ return;
+
+ WRITE_ONCE(avg->started, false);
+ cancel_delayed_work_sync(&avg->work);
+
+ for_each_cpu(cpu, cpumask_of_node(nid))
+ kswapd_prev_idle_valid[cpu] = false;
+}
+
+static int kswapd_idle_avg_cpu_online(unsigned int cpu)
+{
+ if (kswapd_prev_idle_valid)
+ kswapd_prev_idle_valid[cpu] = false;
+
+ if (kswapd_idle_avg) {
+ int nid = cpu_to_node(cpu);
+
+ WRITE_ONCE(kswapd_idle_avg[nid].node_cpus,
+ cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask));
+ }
+
+ return 0;
+}
+
+static int kswapd_idle_avg_cpu_offline(unsigned int cpu)
+{
+ if (kswapd_prev_idle_valid)
+ kswapd_prev_idle_valid[cpu] = false;
+
+ if (kswapd_idle_avg) {
+ int nid = cpu_to_node(cpu);
+
+ WRITE_ONCE(kswapd_idle_avg[nid].node_cpus,
+ cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask));
+ }
+
+ return 0;
+}
+
+static u32 kswapd_node_avg_idle_cores(int nid)
+{
+ if (!kswapd_idle_avg || nid < 0 || nid >= nr_node_ids)
+ return 0;
+
+ return READ_ONCE(kswapd_idle_avg[nid].avg_idle_cores);
+}
+
+static int kswapd_node_cpu_count(int nid)
+{
+ if (!kswapd_idle_avg || nid < 0 || nid >= nr_node_ids)
+ return cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask);
+
+ return READ_ONCE(kswapd_idle_avg[nid].node_cpus);
+}
+
+static void kswapd_idle_avg_init(void)
+{
+ int cpu;
+ int nid;
+ int ret;
+
+ kswapd_idle_avg = kcalloc(nr_node_ids, sizeof(*kswapd_idle_avg),
+ GFP_KERNEL);
+ if (!kswapd_idle_avg)
+ return;
+
+ kswapd_prev_idle = kcalloc(nr_cpu_ids, sizeof(*kswapd_prev_idle),
+ GFP_KERNEL);
+ if (!kswapd_prev_idle)
+ goto free_idle_avg;
+
+ kswapd_prev_idle_valid = kcalloc(nr_cpu_ids,
+ sizeof(*kswapd_prev_idle_valid),
+ GFP_KERNEL);
+ if (!kswapd_prev_idle_valid)
+ goto free_prev_idle;
+
+ for (nid = 0; nid < nr_node_ids; nid++) {
+ struct kswapd_node_idle_avg *avg = &kswapd_idle_avg[nid];
+
+ avg->nid = nid;
+ INIT_DELAYED_WORK(&avg->work, kswapd_idle_avg_sample_workfn);
+ }
+
+ ret = cpuhp_setup_state_nocalls(CPUHP_AP_ONLINE_DYN,
+ "mm/vmscan:online",
+ kswapd_idle_avg_cpu_online,
+ kswapd_idle_avg_cpu_offline);
+ if (ret < 0)
+ goto free_prev_idle_valid;
+
+ cpus_read_lock();
+ for_each_online_cpu(cpu)
+ kswapd_idle_avg_cpu_online(cpu);
+ cpus_read_unlock();
+
+ return;
+
+free_prev_idle_valid:
+ kfree(kswapd_prev_idle_valid);
+ kswapd_prev_idle_valid = NULL;
+free_prev_idle:
+ kfree(kswapd_prev_idle);
+ kswapd_prev_idle = NULL;
+free_idle_avg:
+ kfree(kswapd_idle_avg);
+ kswapd_idle_avg = NULL;
+}
+#else
+static inline u32 kswapd_node_avg_idle_cores(int nid)
+{
+ return 0;
+}
+
+static inline int kswapd_node_cpu_count(int nid)
+{
+ return cpumask_weight_and(cpumask_of_node(nid), cpu_online_mask);
+}
+
+static inline void kswapd_idle_avg_init(void) {}
+static inline void kswapd_idle_avg_start_node(int nid) {}
+static inline void kswapd_idle_avg_stop_node(int nid) {}
+#endif /* CONFIG_NUMA */
+
#ifdef ARCH_HAS_PREFETCHW
#define prefetchw_prev_lru_folio(_folio, _base, _field) \
do { \
@@ -4113,10 +4395,8 @@ static bool try_to_inc_max_seq(struct lruvec *lruvec, unsigned long seq,
walk_mm(mm, walk);
} while (mm);
done:
- if (success) {
+ if (success)
success = inc_max_seq(lruvec, seq, swappiness);
- WARN_ON_ONCE(!success);
- }
return success;
}
@@ -4741,7 +5021,7 @@ static int scan_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
VM_WARN_ON_ONCE(nr_to_scan > MAX_LRU_BATCH);
VM_WARN_ON_ONCE(!list_empty(list));
- if (get_nr_gens(lruvec, type) == MIN_NR_GENS)
+ if (!current_is_kswapd() && get_nr_gens(lruvec, type) == MIN_NR_GENS)
return 0;
gen = lru_gen_from_seq(lrugen->min_seq[type]);
@@ -6671,7 +6951,7 @@ static bool allow_direct_reclaim(pg_data_t *pgdat)
if (READ_ONCE(pgdat->kswapd_highest_zoneidx) > ZONE_NORMAL)
WRITE_ONCE(pgdat->kswapd_highest_zoneidx, ZONE_NORMAL);
- wake_up_interruptible(&pgdat->kswapd_wait);
+ wake_up_interruptible_all(&pgdat->kswapd_wait);
}
return wmark_ok;
@@ -7401,7 +7681,8 @@ static void kswapd_try_to_sleep(pg_data_t *pgdat, int alloc_order, int reclaim_o
if (freezing(current) || kthread_should_stop())
return;
- prepare_to_wait(&pgdat->kswapd_wait, &wait, TASK_INTERRUPTIBLE);
+ prepare_to_wait_exclusive(&pgdat->kswapd_wait, &wait,
+ TASK_INTERRUPTIBLE);
/*
* Try to sleep for a short interval. Note that kcompactd will only be
@@ -7442,7 +7723,8 @@ static void kswapd_try_to_sleep(pg_data_t *pgdat, int alloc_order, int reclaim_o
}
finish_wait(&pgdat->kswapd_wait, &wait);
- prepare_to_wait(&pgdat->kswapd_wait, &wait, TASK_INTERRUPTIBLE);
+ prepare_to_wait_exclusive(&pgdat->kswapd_wait, &wait,
+ TASK_INTERRUPTIBLE);
}
/*
@@ -7575,6 +7857,8 @@ void wakeup_kswapd(struct zone *zone, gfp_t gfp_flags, int order,
{
pg_data_t *pgdat;
enum zone_type curr_idx;
+ int nid, threads_to_wake, threads_woken;
+ int node_cpus, avg_idle_cores;
if (!managed_zone(zone))
return;
@@ -7612,7 +7896,29 @@ void wakeup_kswapd(struct zone *zone, gfp_t gfp_flags, int order,
trace_mm_vmscan_wakeup_kswapd(pgdat->node_id, highest_zoneidx, order,
gfp_flags);
- wake_up_interruptible(&pgdat->kswapd_wait);
+
+ /* Determine how many kswapd threads to wake based on the 10-second
+ * rolling average of idle cores on this NUMA node. More idle cores
+ * means more threads can run without adding contention; fewer idle
+ * cores means we keep the wakeup conservative.
+ */
+ nid = pgdat->node_id;
+ avg_idle_cores = kswapd_node_avg_idle_cores(nid);
+ node_cpus = kswapd_node_cpu_count(nid);
+
+ /*
+ * If no average is available yet (early boot), fall back to waking
+ * a single thread to avoid stalling reclaim.
+ */
+ threads_to_wake = max(1,
+ min3(avg_idle_cores, node_cpus,
+ READ_ONCE(current_kswapds_per_node)));
+
+ threads_woken = wake_up_nr(&pgdat->kswapd_wait, threads_to_wake);
+
+ trace_mm_vmscan_kswapd_threads_to_wake(nid, threads_to_wake,
+ threads_woken, node_cpus,
+ avg_idle_cores);
}
void kswapd_clear_hopeless(pg_data_t *pgdat, enum kswapd_clear_hopeless_reason reason)
@@ -7638,7 +7944,8 @@ void kswapd_try_clear_hopeless(struct pglist_data *pgdat,
bool kswapd_test_hopeless(pg_data_t *pgdat)
{
- return atomic_read(&pgdat->kswapd_failures) >= MAX_RECLAIM_RETRIES;
+ return atomic_read(&pgdat->kswapd_failures) >=
+ MAX_RECLAIM_RETRIES * READ_ONCE(current_kswapds_per_node);
}
#ifdef CONFIG_HIBERNATION
@@ -7680,27 +7987,93 @@ unsigned long shrink_all_memory(unsigned long nr_to_reclaim)
}
#endif /* CONFIG_HIBERNATION */
+static void update_kswapds_per_node_node(int nid)
+{
+ pg_data_t *pgdat;
+ int drop, increase;
+ int last_idx, start_idx, hid;
+ int nr_threads = current_kswapds_per_node;
+
+ pgdat = NODE_DATA(nid);
+ pgdat_kswapd_lock(pgdat);
+ last_idx = nr_threads - 1;
+ if (max_kswapds_per_node < nr_threads) {
+ drop = nr_threads - max_kswapds_per_node;
+ for (hid = last_idx; hid > (last_idx - drop); hid--) {
+ if (pgdat->kswapd[hid]) {
+ kthread_stop(pgdat->kswapd[hid]);
+ pgdat->kswapd[hid] = NULL;
+ }
+ }
+ } else {
+ increase = max_kswapds_per_node - nr_threads;
+ start_idx = last_idx + 1;
+ for (hid = start_idx; hid < (start_idx + increase); hid++) {
+ pgdat->kswapd[hid] = kthread_run(kswapd, pgdat, "kswapd%d:%d",
+ nid, hid);
+ if (IS_ERR(pgdat->kswapd[hid])) {
+ pr_err("Failed to start kswapd%d on node %d\n", hid, nid);
+ pgdat->kswapd[hid] = NULL;
+ /*
+ * We are out of resources. Do not start any
+ * more threads.
+ */
+ break;
+ }
+ }
+ }
+ pgdat_kswapd_unlock(pgdat);
+}
+
+void update_max_kswapds_per_node(void)
+{
+ int nid;
+
+ if (current_kswapds_per_node == max_kswapds_per_node)
+ return;
+
+ /*
+ * Hold the memory hotplug lock to avoid racing with memory
+ * hotplug initiated updates
+ */
+ mem_hotplug_begin();
+ for_each_node_state(nid, N_MEMORY)
+ update_kswapds_per_node_node(nid);
+
+ pr_info("max_kswapds_per_node changed, old:%d new:%d\n",
+ current_kswapds_per_node, max_kswapds_per_node);
+ current_kswapds_per_node = max_kswapds_per_node;
+ mem_hotplug_done();
+}
+
/*
* This kswapd start function will be called by init and node-hot-add.
*/
void __meminit kswapd_run(int nid)
{
pg_data_t *pgdat = NODE_DATA(nid);
+ int hid, nr_threads;
pgdat_kswapd_lock(pgdat);
- if (!pgdat->kswapd) {
- pgdat->kswapd = kthread_create_on_node(kswapd, pgdat, nid, "kswapd%d", nid);
- if (IS_ERR(pgdat->kswapd)) {
- /* failure at boot is fatal */
- pr_err("Failed to start kswapd on node %d, ret=%pe\n",
- nid, pgdat->kswapd);
- BUG_ON(system_state < SYSTEM_RUNNING);
- pgdat->kswapd = NULL;
- } else {
- wake_up_process(pgdat->kswapd);
+ nr_threads = max_kswapds_per_node;
+ for (hid = 0; hid < nr_threads; hid++) {
+ if (!pgdat->kswapd[hid]) {
+ pgdat->kswapd[hid] =
+ kthread_create_on_node(kswapd, pgdat, nid,
+ "kswapd%d:%d", nid, hid);
+ if (IS_ERR(pgdat->kswapd[hid])) {
+ /* failure at boot is fatal */
+ pr_err("Failed to start kswapd%d on node %d, ret=%pe\n", hid,
+ nid, pgdat->kswapd[hid]);
+ pgdat->kswapd[hid] = NULL;
+ } else {
+ wake_up_process(pgdat->kswapd[hid]);
+ }
}
}
+ current_kswapds_per_node = nr_threads;
pgdat_kswapd_unlock(pgdat);
+ kswapd_idle_avg_start_node(nid);
}
/*
@@ -7710,13 +8083,17 @@ void __meminit kswapd_run(int nid)
void __meminit kswapd_stop(int nid)
{
pg_data_t *pgdat = NODE_DATA(nid);
- struct task_struct *kswapd;
+ int hid;
+ int nr_threads = current_kswapds_per_node;
+
+ kswapd_idle_avg_stop_node(nid);
pgdat_kswapd_lock(pgdat);
- kswapd = pgdat->kswapd;
- if (kswapd) {
- kthread_stop(kswapd);
- pgdat->kswapd = NULL;
+ for (hid = 0; hid < nr_threads; hid++) {
+ if (pgdat->kswapd[hid]) {
+ kthread_stop(pgdat->kswapd[hid]);
+ pgdat->kswapd[hid] = NULL;
+ }
}
pgdat_kswapd_unlock(pgdat);
}
@@ -7747,6 +8124,7 @@ static int __init kswapd_init(void)
{
int nid;
+ kswapd_idle_avg_init();
for_each_node_state(nid, N_MEMORY)
kswapd_run(nid);
register_sysctl_init("vm", vmscan_sysctl_table);
|
| 45/3 | 2026/10/09 22:43 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 10611 output: 41 thoughts: 1046 Thoughts: **KMSAN Fuzzing Session Justification** Okay, so I'm evaluating whether this kernel patch series warrants a dedicated KMSAN fuzzing run. Safety check complete – we're good to go. Let's dig in. This patch introduces multiple kswapd threads per node, adds a sysctl for their configuration, and modifies a number of core memory management files. Specifically, it touches `include/linux/mmzone.h`, `include/trace/events/vmscan.h`, `mm/compaction.c`, `mm/internal.h`, `mm/page_alloc.c`, and `mm/vmscan.c`. It looks like we're dealing with per-node idle CPU core averages, delayed work, and thread management related to kswapd. Now, the crucial question: KMSAN applicability. The rules are clear: KMSAN is for catching *uninitialized memory* issues *not* detected by KASAN, LOCKDEP, or standard bug detectors. We need to look for kernel structures sent to userspace with potential initialization issues, conditional logic dependent on uninitialized data, allocation of complex structures with potentially uninitialized fields, or ways that uninitialized data might be accessed through bounds checks. So, am I seeing anything that screams KMSAN? First, the `pgdat->kswapd[MAX_KSWAPD_THREADS]` array is part of `pglist_data`. The question is, how is `pglist_data` allocated? If it's zero-initialized or allocated in BSS or via `kzalloc`, we may be ok. Let me check with `grepper`... |
| 46/3 | 2026/10/09 22:43 | tool | grepper |
0mArgs:{
"Expression": "(alloc_node_data|bootmem_init|free_area_init)",
"PathPrefixes": [
"mm/"
]
}Results: {
"Output": "mm/arch_numa.c=207=static void __init setup_node_data(int nid, u64 start_pfn, u64 end_pfn)\n--\nmm/arch_numa.c-211-\nmm/arch_numa.c:212:\talloc_node_data(nid);\nmm/arch_numa.c-213-\n--\nmm/hugetlb.c=3213=static bool __init hugetlb_bootmem_page_earlycma(struct huge_bootmem_page *m)\n--\nmm/hugetlb.c-3227- */\nmm/hugetlb.c:3228:static void __init hugetlb_bootmem_init_migratetype(struct folio *folio,\nmm/hugetlb.c-3229-\t\t\t\t\t\t\t struct hstate *h)\n--\nmm/hugetlb.c=3244=static void __init prep_and_add_bootmem_folios(struct hstate *h,\n--\nmm/hugetlb.c-3264-\t\t}\nmm/hugetlb.c:3265:\t\thugetlb_bootmem_init_migratetype(folio, h);\nmm/hugetlb.c-3266-\t\t/* Subdivide locks to achieve better parallel performance */\n--\nmm/hugetlb.c=3278=static void __init gather_bootmem_prealloc_node(unsigned long nid)\n--\nmm/hugetlb.c-3322-\t\t * For CMA pages, this is done in init_cma_pageblock\nmm/hugetlb.c:3323:\t\t * (via hugetlb_bootmem_init_migratetype), so skip it here.\nmm/hugetlb.c-3324-\t\t */\n--\nmm/memblock.c=2280=static void __init free_unused_memmap(void)\n--\nmm/memblock.c-2290-\t * This relies on each bank being in address order.\nmm/memblock.c:2291:\t * The banks are sorted previously in bootmem_init().\nmm/memblock.c-2292-\t */\n--\nmm/memory_hotplug.c=1256=static pg_data_t *hotadd_init_pgdat(int nid)\n--\nmm/memory_hotplug.c-1260-\t/*\nmm/memory_hotplug.c:1261:\t * NODE_DATA is preallocated (free_area_init) but its internal\nmm/memory_hotplug.c-1262-\t * state is not allocated completely. Add missing pieces.\n--\nmm/memory_hotplug.c-1268-\t/* init node's zones as empty zones, we don't have any present pages.*/\nmm/memory_hotplug.c:1269:\tif (free_area_init_core_hotplug(pgdat))\nmm/memory_hotplug.c-1270-\t\treturn NULL;\n--\nmm/mm_init.c=1463=static inline void __init set_pageblock_order(void)\n--\nmm/mm_init.c-1476-#ifdef CONFIG_MEMORY_HOTPLUG\nmm/mm_init.c:1477:int __ref free_area_init_core_hotplug(struct pglist_data *pgdat)\nmm/mm_init.c-1478-{\n--\nmm/mm_init.c-1534-\nmm/mm_init.c:1535:static void __init free_area_init_core(struct pglist_data *pgdat)\nmm/mm_init.c-1536-{\n--\nmm/mm_init.c=1638=void __init get_pfn_range_for_nid(unsigned int nid,\n--\nmm/mm_init.c-1655-\nmm/mm_init.c:1656:static void __init free_area_init_node(int nid)\nmm/mm_init.c-1657-{\n--\nmm/mm_init.c-1684-\nmm/mm_init.c:1685:\tfree_area_init_core(pgdat);\nmm/mm_init.c-1686-\tlru_gen_init_pgdat(pgdat);\n--\nmm/mm_init.c=1728=static void __init set_high_memory(void)\n--\nmm/mm_init.c-1749-/**\nmm/mm_init.c:1750: * free_area_init - Initialise all pg_data_t and zone data\nmm/mm_init.c-1751- *\nmm/mm_init.c:1752: * This will call free_area_init_node() for each active node in the system.\nmm/mm_init.c-1753- * Using the page ranges provided by memblock_set_node(), the size of each\n--\nmm/mm_init.c-1760- */\nmm/mm_init.c:1761:static void __init free_area_init(void)\nmm/mm_init.c-1762-{\n--\nmm/mm_init.c-1839-\t\tpgdat = NODE_DATA(nid);\nmm/mm_init.c:1840:\t\tfree_area_init_node(nid);\nmm/mm_init.c-1841-\n--\nmm/mm_init.c=2627=void __init mm_core_init_early(void)\n--\nmm/mm_init.c-2630-\nmm/mm_init.c:2631:\tfree_area_init();\nmm/mm_init.c-2632-\n--\nmm/numa.c=9=EXPORT_SYMBOL(node_data);\n--\nmm/numa.c-11-/* Allocate NODE_DATA for a node on the local memory */\nmm/numa.c:12:void __init alloc_node_data(int nid)\nmm/numa.c-13-{\n--\nmm/page_alloc.c=5996=static void __build_all_zonelists(void *data)\n--\nmm/page_alloc.c-6028-\t\t * All possible nodes have pgdat preallocated\nmm/page_alloc.c:6029:\t\t * in free_area_init\nmm/page_alloc.c-6030-\t\t */\n"
} |
| 47/3 | 2026/10/09 22:43 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 13666 output: 33 thoughts: 59 Thoughts: Here's my summary, as though I'm thinking through this process: **Diving into NUMA Initialization** Okay, so I'm looking into how `NODE_DATA(nid)` is handled, particularly in the context of NUMA. Right, it's preallocated, which is good to know – it's already there when the system boots. It's either allocated with `memblock_alloc_node` which should zero the memory, or it's part of the `static contig_page_data` structure in UMA which is in the BSS segment. I want to confirm exactly how the data is being handled. I should take a look at `alloc_node_data` in `mm/numa.c`. That's where the heart of this initialization should be. Let me check the kernel code to confirm the initial state and allocation strategy. |
| 48/3 | 2026/10/09 22:43 | tool | read-file |
0mArgs:{
"File": "mm/numa.c",
"FirstLine": 10,
"LineCount": 30
}Results: {
"Contents": " 10:\t\n 11:\t/* Allocate NODE_DATA for a node on the local memory */\n 12:\tvoid __init alloc_node_data(int nid)\n 13:\t{\n 14:\t\tconst size_t nd_size = roundup(sizeof(pg_data_t), SMP_CACHE_BYTES);\n 15:\t\tu64 nd_pa;\n 16:\t\tint tnid;\n 17:\t\n 18:\t\t/* Allocate node data. Try node-local memory and then any node. */\n 19:\t\tnd_pa = memblock_phys_alloc_try_nid(nd_size, SMP_CACHE_BYTES, nid);\n 20:\t\tif (!nd_pa)\n 21:\t\t\tpanic(\"Cannot allocate %zu bytes for node %d data\\n\",\n 22:\t\t\t nd_size, nid);\n 23:\t\n 24:\t\t/* report and initialize */\n 25:\t\tpr_info(\"NODE_DATA(%d) allocated [mem %#010Lx-%#010Lx]\\n\", nid,\n 26:\t\t\tnd_pa, nd_pa + nd_size - 1);\n 27:\t\ttnid = early_pfn_to_nid(nd_pa \u003e\u003e PAGE_SHIFT);\n 28:\t\tif (tnid != nid)\n 29:\t\t\tpr_info(\" NODE_DATA(%d) on node %d\\n\", nid, tnid);\n 30:\t\n 31:\t\tnode_data[nid] = __va(nd_pa);\n 32:\t\tmemset(NODE_DATA(nid), 0, sizeof(pg_data_t));\n 33:\t}\n 34:\t\n 35:\tvoid __init alloc_offline_node_data(int nid)\n 36:\t{\n 37:\t\tpg_data_t *pgdat;\n 38:\t\tnode_data[nid] = memblock_alloc_or_panic(sizeof(*pgdat), SMP_CACHE_BYTES);\n 39:\t}\n"
} |
| 49/3 | 2026/10/09 22:43 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 14456 output: 20 thoughts: 3765 Thoughts: **Analysis of `kswapd` Initialization and Usage** Alright, let's break this down systematically. My initial impression is that the code looks reasonably safe regarding uninitialized memory reads, but we need to be thorough. First, I see consistent initialization with `memset` and `kcalloc` ensuring zeroed memory for key structures like `pg_data_t` and the idle average arrays (`kswapd_idle_avg`, `kswapd_prev_idle`, `kswapd_prev_idle_valid`). This is a good starting point. Now, I'm checking the flow of data. Nothing in the sysctl handler (`max_kswapds_per_node_sysctl_handler`) exposes uninitialized memory, just an integer (`max_kswapds_per_node`). Similarly, the tracepoint arguments in `mm_vmscan_kswapd_threads_to_wake` use basic integer values passed by value. And local variables used in `wakeup_kswapd` are explicitly initialized. The `kswapd_node_avg_idle_cores` function uses `READ_ONCE`, which seems appropriate for shared data. The `kswapd_node_cpu_count` function is equally safe. It's safe to assume `nr_node_ids` and `nr_cpu_ids` are valid. In `kswapd_idle_avg_sample_workfn`, I focus on memory access patterns. The `kswapd_prev_idle` array is indexed by CPU ID within a `for_each_cpu` loop, and the array size is determined by `nr_cpu_ids`, so this check is valid. The `kswapd_prev_idle_valid` boolean array is allocated with `kcalloc` so initialised to false. It only reads the memory if `kswapd_prev_idle_valid` is `true`. The `avg` structure, initialized with `kcalloc`, ensures initial values and avoids division-by-zero errors. `avg->count` is at least 1 after it has been sampled at least once, so the division is always safe. This looks good. Now, my attention is drawn to the `kswapd_is_running` function, particularly the loop through `pgdat->kswapd`. I need to ensure the loop accesses are safe and consider the possibility of zero threads. Even if `pgdat->kswapd[i]` is null, the code checks before trying to use that memory, and the loop is guaranteed to run, so the `running` variable is always assigned a value before being returned. Finally, I'm analyzing `update_kswapds_per_node_node` and `kswapd_stop` to examine potential out-of-bounds array access. It uses `max_kswapds_per_node` which is derived from `DEF_MAX_KSWAPDS_PER_NODE` (1) and `max_kswapd_threads` (16) via sysctl. It accesses `pgdat->kswapd` up to `MAX_KSWAPD_THREADS` which is 16. The sysctl bounds `max_kswapds_per_node` between 1 and `max_kswapd_threads`. The bounds check on the loop indexes are correct and the access bounds are within 0..15, as `pgdat->kswapd` array size is `MAX_KSWAPD_THREADS` (16). This means it's safe. While there might be concurrency issues or logic bugs with thread management, these are not directly related to uninitialized memory access. I still need to look for any other access to `pgdat->kswapd`. So, in summary, based on this review, I *think* the risk of reading uninitialized memory is low, but a complete review of all `kswapd` accesses is needed. |
| 50/3 | 2026/10/09 22:44 | tool | grepper |
0mArgs:{
"Expression": "-\u003ekswapd\\b"
}Results: {
"Output": "fs/xfs/libxfs/xfs_btree.c=3006=xfs_btree_split_worker(\n--\nfs/xfs/libxfs/xfs_btree.c-3019-\t */\nfs/xfs/libxfs/xfs_btree.c:3020:\tif (args-\u003ekswapd)\nfs/xfs/libxfs/xfs_btree.c-3021-\t\tnew_pflags |= PF_MEMALLOC | PF_KSWAPD;\n--\nmm/compaction.c=2091=static isolate_migrate_t isolate_migratepages(struct compact_control *cc)\n--\nmm/compaction.c-2194- *\nmm/compaction.c:2195: * pgdat_kswapd_lock() pins pgdat-\u003ekswapd, so a concurrent kswapd_stop() can't\nmm/compaction.c-2196- * zero it.\n--\nmm/compaction.c=2198=static bool kswapd_is_running(pg_data_t *pgdat)\n--\nmm/compaction.c-2203-\tfor (i = 0; i \u003c MAX_KSWAPD_THREADS; ++i) {\nmm/compaction.c:2204:\t\trunning = pgdat-\u003ekswapd[i] \u0026\u0026 task_is_running(pgdat-\u003ekswapd[i]);\nmm/compaction.c-2205-\t\tif (running)\n--\nmm/vmscan.c=7990=static void update_kswapds_per_node_node(int nid)\n--\nmm/vmscan.c-8002-\t\tfor (hid = last_idx; hid \u003e (last_idx - drop); hid--) {\nmm/vmscan.c:8003:\t\t\tif (pgdat-\u003ekswapd[hid]) {\nmm/vmscan.c:8004:\t\t\t\tkthread_stop(pgdat-\u003ekswapd[hid]);\nmm/vmscan.c:8005:\t\t\t\tpgdat-\u003ekswapd[hid] = NULL;\nmm/vmscan.c-8006-\t\t\t}\n--\nmm/vmscan.c-8011-\t\tfor (hid = start_idx; hid \u003c (start_idx + increase); hid++) {\nmm/vmscan.c:8012:\t\t\tpgdat-\u003ekswapd[hid] = kthread_run(kswapd, pgdat, \"kswapd%d:%d\",\nmm/vmscan.c-8013-\t\t\t\t\t\t\t nid, hid);\nmm/vmscan.c:8014:\t\t\tif (IS_ERR(pgdat-\u003ekswapd[hid])) {\nmm/vmscan.c-8015-\t\t\t\tpr_err(\"Failed to start kswapd%d on node %d\\n\", hid, nid);\nmm/vmscan.c:8016:\t\t\t\tpgdat-\u003ekswapd[hid] = NULL;\nmm/vmscan.c-8017-\t\t\t\t/*\n--\nmm/vmscan.c=8052=void __meminit kswapd_run(int nid)\n--\nmm/vmscan.c-8059-\tfor (hid = 0; hid \u003c nr_threads; hid++) {\nmm/vmscan.c:8060:\t\tif (!pgdat-\u003ekswapd[hid]) {\nmm/vmscan.c:8061:\t\t\tpgdat-\u003ekswapd[hid] =\nmm/vmscan.c-8062-\t\t\t\tkthread_create_on_node(kswapd, pgdat, nid,\nmm/vmscan.c-8063-\t\t\t\t\t\t \"kswapd%d:%d\", nid, hid);\nmm/vmscan.c:8064:\t\t\tif (IS_ERR(pgdat-\u003ekswapd[hid])) {\nmm/vmscan.c-8065-\t\t\t\t/* failure at boot is fatal */\nmm/vmscan.c-8066-\t\t\t\tpr_err(\"Failed to start kswapd%d on node %d, ret=%pe\\n\", hid,\nmm/vmscan.c:8067:\t\t\t\t nid, pgdat-\u003ekswapd[hid]);\nmm/vmscan.c:8068:\t\t\t\tpgdat-\u003ekswapd[hid] = NULL;\nmm/vmscan.c-8069-\t\t\t} else {\nmm/vmscan.c:8070:\t\t\t\twake_up_process(pgdat-\u003ekswapd[hid]);\nmm/vmscan.c-8071-\t\t\t}\n--\nmm/vmscan.c=8083=void __meminit kswapd_stop(int nid)\n--\nmm/vmscan.c-8092-\tfor (hid = 0; hid \u003c nr_threads; hid++) {\nmm/vmscan.c:8093:\t\tif (pgdat-\u003ekswapd[hid]) {\nmm/vmscan.c:8094:\t\t\tkthread_stop(pgdat-\u003ekswapd[hid]);\nmm/vmscan.c:8095:\t\t\tpgdat-\u003ekswapd[hid] = NULL;\nmm/vmscan.c-8096-\t\t}\n"
} |
| 51/3 | 2026/10/09 22:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 20361 output: 28 thoughts: 85 Thoughts: Okay, here's my thought process as I digest this: **Investigating Potential `pgdat->kswapd` Usage Issues** Right, so compaction and vmscan are the two main culprits using `pgdat->kswapd`, and they've both seen recent updates, which is a good starting point. But I need to be thorough. I can't just assume those are the *only* places. My expert intuition tells me to dig deeper. I need to make sure I'm not missing something subtle, some indirect dependency. Let me broaden my search scope. I'll need to check the broader `pglist_data` structure and any related functions. Let's see... a good place to start would be to look for any instances of `kswapd` specifically. I'll search for things like `kswapd_wait`, `kswapd_highest_zoneidx`, `kswapd_order`, and `kswapd_failures`. Those keywords should flag any areas where `kswapd` is being used in a broader context or in less obvious ways. That will give me a better understanding of the overall system's use of `kswapd` and if there are any other areas of potential concern. |
| 52/3 | 2026/10/09 22:44 | tool | grepper |
0mArgs:{
"Expression": "kswapd_order|kswapd_highest_zoneidx"
}Results: {
"Output": "Documentation/mm/physical_memory.rst=273=See also Documentation/mm/page_reclaim.rst.\n--\nDocumentation/mm/physical_memory.rst-286-\nDocumentation/mm/physical_memory.rst:287:``kswapd_order``\nDocumentation/mm/physical_memory.rst-288- Controls the order kswapd tries to reclaim\nDocumentation/mm/physical_memory.rst-289-\nDocumentation/mm/physical_memory.rst:290:``kswapd_highest_zoneidx``\nDocumentation/mm/physical_memory.rst-291- The highest zone index to be reclaimed by kswapd\n--\nDocumentation/translations/zh_CN/mm/physical_memory.rst=205=node_zonelists引用。\n--\nDocumentation/translations/zh_CN/mm/physical_memory.rst-273-\nDocumentation/translations/zh_CN/mm/physical_memory.rst:274:``kswapd_order``\nDocumentation/translations/zh_CN/mm/physical_memory.rst-275-控制kswapd尝试回收的order。\nDocumentation/translations/zh_CN/mm/physical_memory.rst-276-\nDocumentation/translations/zh_CN/mm/physical_memory.rst:277:``kswapd_highest_zoneidx``\nDocumentation/translations/zh_CN/mm/physical_memory.rst-278-kswapd线程可以回收的最高区域索引。\n--\ninclude/linux/mmzone.h=1480=typedef struct pglist_data {\n--\ninclude/linux/mmzone.h-1535-\tstruct task_struct *kswapd[MAX_KSWAPD_THREADS];\ninclude/linux/mmzone.h:1536:\tint kswapd_order;\ninclude/linux/mmzone.h:1537:\tenum zone_type kswapd_highest_zoneidx;\ninclude/linux/mmzone.h-1538-\n--\nmm/mm_init.c=1477=int __ref free_area_init_core_hotplug(struct pglist_data *pgdat)\n--\nmm/mm_init.c-1495-\t * Reset the nr_zones, order and highest_zoneidx before reuse.\nmm/mm_init.c:1496:\t * Note that kswapd will init kswapd_highest_zoneidx properly\nmm/mm_init.c-1497-\t * when it starts in the near future.\n--\nmm/mm_init.c-1499-\tpgdat-\u003enr_zones = 0;\nmm/mm_init.c:1500:\tpgdat-\u003ekswapd_order = 0;\nmm/mm_init.c:1501:\tpgdat-\u003ekswapd_highest_zoneidx = 0;\nmm/mm_init.c-1502-\tpgdat-\u003enode_start_pfn = 0;\n--\nmm/mm_init.c=1656=static void __init free_area_init_node(int nid)\n--\nmm/mm_init.c-1662-\t/* pg_data_t should be reset to zero when it's allocated */\nmm/mm_init.c:1663:\tWARN_ON(pgdat-\u003enr_zones || pgdat-\u003ekswapd_highest_zoneidx);\nmm/mm_init.c-1664-\n--\nmm/vmscan.c=6924=static bool allow_direct_reclaim(pg_data_t *pgdat)\n--\nmm/vmscan.c-6950-\tif (!wmark_ok \u0026\u0026 waitqueue_active(\u0026pgdat-\u003ekswapd_wait)) {\nmm/vmscan.c:6951:\t\tif (READ_ONCE(pgdat-\u003ekswapd_highest_zoneidx) \u003e ZONE_NORMAL)\nmm/vmscan.c:6952:\t\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, ZONE_NORMAL);\nmm/vmscan.c-6953-\n--\nmm/vmscan.c=7437=static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx)\n--\nmm/vmscan.c-7660-/*\nmm/vmscan.c:7661: * The pgdat-\u003ekswapd_highest_zoneidx is used to pass the highest zone index to\nmm/vmscan.c-7662- * be reclaimed by kswapd from the waker. If the value is MAX_NR_ZONES which is\n--\nmm/vmscan.c-7666- */\nmm/vmscan.c:7667:static enum zone_type kswapd_highest_zoneidx(pg_data_t *pgdat,\nmm/vmscan.c-7668-\t\t\t\t\t enum zone_type prev_highest_zoneidx)\nmm/vmscan.c-7669-{\nmm/vmscan.c:7670:\tenum zone_type curr_idx = READ_ONCE(pgdat-\u003ekswapd_highest_zoneidx);\nmm/vmscan.c-7671-\n--\nmm/vmscan.c=7675=static void kswapd_try_to_sleep(pg_data_t *pgdat, int alloc_order, int reclaim_order,\n--\nmm/vmscan.c-7711-\t\t/*\nmm/vmscan.c:7712:\t\t * If woken prematurely then reset kswapd_highest_zoneidx and\nmm/vmscan.c-7713-\t\t * order. The values will either be from a wakeup request or\n--\nmm/vmscan.c-7716-\t\tif (remaining) {\nmm/vmscan.c:7717:\t\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx,\nmm/vmscan.c:7718:\t\t\t\t\tkswapd_highest_zoneidx(pgdat,\nmm/vmscan.c-7719-\t\t\t\t\t\t\thighest_zoneidx));\nmm/vmscan.c-7720-\nmm/vmscan.c:7721:\t\t\tif (READ_ONCE(pgdat-\u003ekswapd_order) \u003c reclaim_order)\nmm/vmscan.c:7722:\t\t\t\tWRITE_ONCE(pgdat-\u003ekswapd_order, reclaim_order);\nmm/vmscan.c-7723-\t\t}\n--\nmm/vmscan.c=7774=static int kswapd(void *p)\n--\nmm/vmscan.c-7795-\nmm/vmscan.c:7796:\tWRITE_ONCE(pgdat-\u003ekswapd_order, 0);\nmm/vmscan.c:7797:\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, MAX_NR_ZONES);\nmm/vmscan.c-7798-\tatomic_set(\u0026pgdat-\u003enr_writeback_throttled, 0);\n--\nmm/vmscan.c-7801-\nmm/vmscan.c:7802:\t\talloc_order = reclaim_order = READ_ONCE(pgdat-\u003ekswapd_order);\nmm/vmscan.c:7803:\t\thighest_zoneidx = kswapd_highest_zoneidx(pgdat,\nmm/vmscan.c-7804-\t\t\t\t\t\t\thighest_zoneidx);\n--\nmm/vmscan.c-7810-\t\t/* Read the new order and highest_zoneidx */\nmm/vmscan.c:7811:\t\talloc_order = READ_ONCE(pgdat-\u003ekswapd_order);\nmm/vmscan.c:7812:\t\thighest_zoneidx = kswapd_highest_zoneidx(pgdat,\nmm/vmscan.c-7813-\t\t\t\t\t\t\thighest_zoneidx);\nmm/vmscan.c:7814:\t\tWRITE_ONCE(pgdat-\u003ekswapd_order, 0);\nmm/vmscan.c:7815:\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, MAX_NR_ZONES);\nmm/vmscan.c-7816-\n--\nmm/vmscan.c=7855=void wakeup_kswapd(struct zone *zone, gfp_t gfp_flags, int order,\n--\nmm/vmscan.c-7869-\tpgdat = zone-\u003ezone_pgdat;\nmm/vmscan.c:7870:\tcurr_idx = READ_ONCE(pgdat-\u003ekswapd_highest_zoneidx);\nmm/vmscan.c-7871-\nmm/vmscan.c-7872-\tif (curr_idx == MAX_NR_ZONES || curr_idx \u003c highest_zoneidx)\nmm/vmscan.c:7873:\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, highest_zoneidx);\nmm/vmscan.c-7874-\nmm/vmscan.c:7875:\tif (READ_ONCE(pgdat-\u003ekswapd_order) \u003c order)\nmm/vmscan.c:7876:\t\tWRITE_ONCE(pgdat-\u003ekswapd_order, order);\nmm/vmscan.c-7877-\n"
} |
| 53/3 | 2026/10/09 22:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 23033 output: 35 thoughts: 47 Thoughts: Okay, let's dive into this. **Examining kswapd's Multithreaded Behavior** Alright, I need to get a better handle on how `kswapd` handles concurrency when multiple threads are involved. My focus is going to be on understanding the mechanics of `kswapd(void *p)` – I'll start by taking a look at `mm/vmscan.c` to find the relevant code, specifically around line 7770. That should give me a good vantage point to observe its behavior in a multithreaded context. I expect to be able to follow the logic and identify how locking, synchronization, and potentially other mechanisms are used to ensure the proper functioning of this critical memory management process. I'm keen to see if there are any subtle inefficiencies or potential bottlenecks that jump out as I review this section. |
| 54/3 | 2026/10/09 22:44 | tool | read-file |
0mArgs:{
"File": "mm/vmscan.c",
"FirstLine": 7770,
"LineCount": 60
}Results: {
"Contents": "7770:\t *\n7771:\t * If there are applications that are active memory-allocators\n7772:\t * (most normal use), this basically shouldn't matter.\n7773:\t */\n7774:\tstatic int kswapd(void *p)\n7775:\t{\n7776:\t\tunsigned int alloc_order, reclaim_order;\n7777:\t\tunsigned int highest_zoneidx = MAX_NR_ZONES - 1;\n7778:\t\tpg_data_t *pgdat = (pg_data_t *)p;\n7779:\t\tstruct task_struct *tsk = current;\n7780:\t\n7781:\t\t/*\n7782:\t\t * Tell the memory management that we're a \"memory allocator\",\n7783:\t\t * and that if we need more memory we should get access to it\n7784:\t\t * regardless (see \"__alloc_pages()\"). \"kswapd\" should\n7785:\t\t * never get caught in the normal page freeing logic.\n7786:\t\t *\n7787:\t\t * (Kswapd normally doesn't need memory anyway, but sometimes\n7788:\t\t * you need a small amount of memory in order to be able to\n7789:\t\t * page out something else, and this flag essentially protects\n7790:\t\t * us from recursively trying to free more memory as we're\n7791:\t\t * trying to free the first piece of memory in the first place).\n7792:\t\t */\n7793:\t\ttsk-\u003eflags |= PF_MEMALLOC | PF_KSWAPD;\n7794:\t\tset_freezable();\n7795:\t\n7796:\t\tWRITE_ONCE(pgdat-\u003ekswapd_order, 0);\n7797:\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, MAX_NR_ZONES);\n7798:\t\tatomic_set(\u0026pgdat-\u003enr_writeback_throttled, 0);\n7799:\t\tfor ( ; ; ) {\n7800:\t\t\tbool was_frozen;\n7801:\t\n7802:\t\t\talloc_order = reclaim_order = READ_ONCE(pgdat-\u003ekswapd_order);\n7803:\t\t\thighest_zoneidx = kswapd_highest_zoneidx(pgdat,\n7804:\t\t\t\t\t\t\t\thighest_zoneidx);\n7805:\t\n7806:\tkswapd_try_sleep:\n7807:\t\t\tkswapd_try_to_sleep(pgdat, alloc_order, reclaim_order,\n7808:\t\t\t\t\t\thighest_zoneidx);\n7809:\t\n7810:\t\t\t/* Read the new order and highest_zoneidx */\n7811:\t\t\talloc_order = READ_ONCE(pgdat-\u003ekswapd_order);\n7812:\t\t\thighest_zoneidx = kswapd_highest_zoneidx(pgdat,\n7813:\t\t\t\t\t\t\t\thighest_zoneidx);\n7814:\t\t\tWRITE_ONCE(pgdat-\u003ekswapd_order, 0);\n7815:\t\t\tWRITE_ONCE(pgdat-\u003ekswapd_highest_zoneidx, MAX_NR_ZONES);\n7816:\t\n7817:\t\t\tif (kthread_freezable_should_stop(\u0026was_frozen))\n7818:\t\t\t\tbreak;\n7819:\t\n7820:\t\t\t/*\n7821:\t\t\t * We can speed up thawing tasks if we don't call balance_pgdat\n7822:\t\t\t * after returning from the refrigerator\n7823:\t\t\t */\n7824:\t\t\tif (was_frozen)\n7825:\t\t\t\tcontinue;\n7826:\t\n7827:\t\t\t/*\n7828:\t\t\t * Reclaim begins at the requested order but if a high-order\n7829:\t\t\t * reclaim fails then kswapd falls back to reclaiming for\n"
} |
| 55/3 | 2026/10/09 22:44 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24287 output: 353 thoughts: 1429 Thoughts: **Analysis of the vmscan.c Patch for Uninitialized Memory Risks** Okay, let's break this down. My primary concern here is always uninitialized memory, which can lead to nasty vulnerabilities. First, I went through the code changes methodically, focusing on how data is handled and whether there's any potential for reading uninitialized values. I examined the diff, specifically lines 4400, 5024, 6954, 7684, and 7947, looking for anything that could introduce or expose uninitialized memory. I paid close attention to where kernel structures are passed to user space, conditional logic involving potentially uninitialized variables, the allocation and initialization of data structures, and the possibility of accessing uninitialized bytes of existing buffers due to bounds-check issues. I specifically looked for: 1. **User-Kernel Interaction:** There are no `copy_to_user`, `put_user`, netlink calls, or ioctls, that would allow uninitialized data to leak to userspace. The only user-kernel interaction is the sysctl handler `max_kswapds_per_node_sysctl_handler` operating on `max_kswapds_per_node`, an `int` which is globally initialized. 2. **Conditional Logic:** All the variables used in any conditional branches, like `avg_idle_cores`, `node_cpus`, `current_kswapds_per_node`, etc., are explicitly initialized, zero-initialized with `kcalloc`, or deterministically calculated. No concerns there. 3. **Data Structure Initialization:** I checked how complex data structures like `kswapd_idle_avg`, `kswapd_prev_idle`, `kswapd_prev_idle_valid`, and the `pgdat->kswapd[MAX_KSWAPD_THREADS]` array within `pglist_data` are initialized. These are allocated with `kcalloc` or zero-initialized at boot or hotplug. They are good to go. 4. **Bounds Checks:** There are no issues with bounds checks or length calculations that could lead to reading uninitialized bytes. I carefully considered the types of bugs that *could* be introduced: * Concurrency bugs like races and lock contention: These are the usual suspects. LOCKDEP, KCSAN, and functional testing will catch these. * Resource leaks, deadlocks, and task stopping races: Standard bug analysis techniques would catch these too. * Division by zero: I validated that the divisor isn't zero. * Memory leaks: Error handling in `kswapd_idle_avg_init` could cause leaks, but not uninitialized memory bugs. * Out of bounds: The sysctl min/max values protect `max_kswapds_per_node`. Even if that failed, KASAN would catch out-of-bounds accesses. The conclusion: this patch does *not* introduce or expose any uninitialized memory risks. KASAN and standard tools are fully adequate to identify bugs that might come up with this change (use-after-free, out of bounds, locking, etc.), KMSAN would add no value in this case. I took one more look to double-check: `kswapd_is_running` is well defined, with `running` getting initialized, and the fields in `trace_mm_vmscan_kswapd_threads_to_wake` are derived from the function parameters. Finally, the Sysctl handling via `proc_dointvec_minmax` on a globally initialized `int`. All checks out. Thus, based on my analysis, the `NeedsKMSAN` flag is `false`. |
| 56/3 | 2026/10/09 22:44 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch implements support for multiple kswapd threads per NUMA node, an idle CPU rolling average mechanism to scale thread wakeup, and a sysctl interface (`max_kswapds_per_node`). \n\nA dedicated KMSAN session is not needed for the following reasons:\n1. No uninitialized memory exposure to userspace: The sysctl handler uses standard `proc_dointvec_minmax` with a fully initialized integer, and there are no info-leak vectors or uninitialized structures copied to user space.\n2. Safe data structure allocations: The NUMA idle-tracking arrays (`kswapd_idle_avg`, `kswapd_prev_idle`, `kswapd_prev_idle_valid`) are allocated using `kcalloc`, ensuring all fields and elements are zero-initialized. The `kswapd` thread array in `pglist_data` is part of node data zeroed at boot/hotplug.\n3. Deterministic initialization of local variables: All local variables (including `running` in `kswapd_is_running` which is unconditionally assigned in the first iteration of the fixed 16-iteration loop) are assigned before being read.\n4. Bug profile: The risks introduced by this change relate to concurrency (multiple kswapd threads contending on reclaim paths), task lifecycle/locking (`kthread_run`/`kthread_stop`), and bounds checks on sysctl parameters—issues that are covered by KASAN, LOCKDEP, and standard kernel debug facilities, rather than KMSAN uninitialized memory tracking.",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch implements support for multiple kswapd threads per NUMA node, an idle CPU rolling average mechanism to scale thread wakeup, and a sysctl interface (`max_kswapds_per_node`). \n\nA dedicated KMSAN session is not needed for the following reasons:\n1. No uninitialized memory exposure to userspace: The sysctl handler uses standard `proc_dointvec_minmax` with a fully initialized integer, and there are no info-leak vectors or uninitialized structures copied to user space.\n2. Safe data structure allocations: The NUMA idle-tracking arrays (`kswapd_idle_avg`, `kswapd_prev_idle`, `kswapd_prev_idle_valid`) are allocated using `kcalloc`, ensuring all fields and elements are zero-initialized. The `kswapd` thread array in `pglist_data` is part of node data zeroed at boot/hotplug.\n3. Deterministic initialization of local variables: All local variables (including `running` in `kswapd_is_running` which is unconditionally assigned in the first iteration of the fixed 16-iteration loop) are assigned before being read.\n4. Bug profile: The risks introduced by this change relate to concurrency (multiple kswapd threads contending on reclaim paths), task lifecycle/locking (`kthread_run`/`kthread_stop`), and bounds checks on sysctl parameters—issues that are covered by KASAN, LOCKDEP, and standard kernel debug facilities, rather than KMSAN uninitialized memory tracking.",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|