The extents_status shrinker's count_objects() callback, ext4_es_count(), reports the number of shrinkable extent_status objects via a percpu counter. do_shrink_slab() reads this count exactly once per invocation and derives a one-shot scan budget (total_scan) from it, then repeatedly calls scan_objects() in fixed-size batches until that budget is exhausted. ext4_es_scan() never updates sc->nr_scanned, even though include/linux/shrinker.h documents that "the callee should track its actual progress" in that field. Because sc->nr_scanned defaults to sc->nr_to_scan before every call, do_shrink_slab() always believes a full batch was examined, regardless of what __es_shrink() actually did. When __es_shrink() hits an empty sbi->s_es_list and returns immediately having examined nothing, do_shrink_slab() has no way to tell "there was nothing to scan" from "a full batch was scanned and none of it was freeable". It keeps calling scan_objects() until the original, one-shot total_scan budget is drained, even though the shrinkable list emptied out long before that budget was used up. Make __es_shrink() report the number of extent_status objects it actually examined through a new nr_scanned output parameter, derived from the existing per-extent nr_to_scan counter that es_reclaim_extents() already decrements as it walks the tree. Have ext4_es_scan() copy this value into sc->nr_scanned, and return SHRINK_STOP once nr_scanned comes back as zero, since that only happens when sbi->s_es_list was already empty and no further calls in this reclaim pass can make progress. Tested by fallocate(2)-ing 10000 4K files (to populate the shrinker with reclaimable unwritten extents without also exercising the extent_status "referenced" second-chance path, which needs a separate two-pass accounting of its own) and triggering "echo 2 > /proc/sys/vm/drop_caches", while tracing the ext4_es_shrink* tracepoints: total scan_objects() calls with calls nr_shrunk == 0 before this patch 429 189 (44%) after this patch 165 1 (0.6%) after (rerun) 242 1 (0.4%) nr_skipped stayed at 0 throughout every run, confirming the wasted calls came from the stale one-shot budget racing ahead of the real list state, not from the existing precached/trylock skip paths. Signed-off-by: Qiliang Yuan --- fs/ext4/extents_status.c | 26 ++++++++++++++++++++++---- 1 file changed, 22 insertions(+), 4 deletions(-) diff --git a/fs/ext4/extents_status.c b/fs/ext4/extents_status.c index 6e4a191e82191..05ded0c018e08 100644 --- a/fs/ext4/extents_status.c +++ b/fs/ext4/extents_status.c @@ -184,7 +184,7 @@ static int __es_remove_extent(struct inode *inode, ext4_lblk_t lblk, struct extent_status *prealloc); static int es_reclaim_extents(struct ext4_inode_info *ei, int *nr_to_scan); static int __es_shrink(struct ext4_sb_info *sbi, int nr_to_scan, - struct ext4_inode_info *locked_ei); + struct ext4_inode_info *locked_ei, int *nr_scanned); static int __revise_pending(struct inode *inode, ext4_lblk_t lblk, ext4_lblk_t len, struct pending_reservation **prealloc); @@ -1670,7 +1670,7 @@ void ext4_es_remove_extent(struct inode *inode, ext4_lblk_t lblk, } static int __es_shrink(struct ext4_sb_info *sbi, int nr_to_scan, - struct ext4_inode_info *locked_ei) + struct ext4_inode_info *locked_ei, int *nr_scanned) { struct ext4_inode_info *ei; struct ext4_es_stats *es_stats; @@ -1679,6 +1679,7 @@ static int __es_shrink(struct ext4_sb_info *sbi, int nr_to_scan, int nr_to_walk; int nr_shrunk = 0; int retried = 0, nr_skipped = 0; + int orig_nr_to_scan = nr_to_scan; es_stats = &sbi->s_es_stats; start_time = ktime_get(); @@ -1738,6 +1739,8 @@ static int __es_shrink(struct ext4_sb_info *sbi, int nr_to_scan, nr_shrunk = es_reclaim_extents(locked_ei, &nr_to_scan); out: + *nr_scanned = orig_nr_to_scan - nr_to_scan; + scan_time = ktime_to_ns(ktime_sub(ktime_get(), start_time)); if (likely(es_stats->es_stats_scan_time)) es_stats->es_stats_scan_time = (scan_time + @@ -1774,15 +1777,30 @@ static unsigned long ext4_es_scan(struct shrinker *shrink, { struct ext4_sb_info *sbi = shrink->private_data; int nr_to_scan = sc->nr_to_scan; - int ret, nr_shrunk; + int ret, nr_shrunk, nr_scanned; ret = percpu_counter_read_positive(&sbi->s_es_stats.es_stats_shk_cnt); trace_ext4_es_shrink_scan_enter(sbi->s_sb, nr_to_scan, ret); - nr_shrunk = __es_shrink(sbi, nr_to_scan, NULL); + nr_shrunk = __es_shrink(sbi, nr_to_scan, NULL, &nr_scanned); + sc->nr_scanned = nr_scanned; ret = percpu_counter_read_positive(&sbi->s_es_stats.es_stats_shk_cnt); trace_ext4_es_shrink_scan_exit(sbi->s_sb, nr_shrunk, ret); + + /* + * es_stats_shk_cnt is a percpu counter, so count_objects() can report + * a stale/approximate value that is still positive even though + * sbi->s_es_list is actually empty by the time we get here. When that + * happens __es_shrink() returns immediately having examined nothing + * (nr_scanned == 0). Without SHRINK_STOP, do_shrink_slab() has no way + * to tell "nothing was there" from "nothing was scanned yet" and will + * keep calling us with the same stale freeable count until its scan + * budget for this priority level is exhausted one batch at a time. + */ + if (nr_scanned == 0) + return SHRINK_STOP; + return nr_shrunk; } --- base-commit: 502d801f0ab03e4f32f9a33d203154ce84887921 change-id: 20260928-fix-ext4-es-scan-nr-scanned-5af70744f6cd Best regards, -- Qiliang Yuan