__swap_cache_add_check() turns away folio entries and slots with no count and lets everything else in. That is safe only when the caller owns the slot. Cluster readahead owns nothing, it walks a raw page_cluster sized window of offsets around the faulting entry. A hibernation slot is not a folio, and the count test does not stop it either, because the type the previous patch added has the count bits set. So readahead allocates a folio and reads the offset off the device into it, for a slot that nothing will ever swap in. A bad slot listed in the swap header gets in the same way, its count bits are set too, and the folio entry that replaces it drops the bad marker. The first patch keeps that folio from doing harm, but the folio and the read still happen. Require a shadow entry instead. A slot dropped from the swap cache always gets one, empty if there is no workingset value. The check runs before the folio allocation in __swap_cache_alloc(), so readahead now skips the offset without allocating or reading. Signed-off-by: Youngjun Park --- mm/swap_state.c | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-) diff --git a/mm/swap_state.c b/mm/swap_state.c index 5be825911e64..9f2cc5918713 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -180,9 +180,14 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci, old_tb = __swap_table_get(ci, ci_off); if (swp_tb_is_folio(old_tb)) return -EEXIST; - if (!__swp_tb_get_count(old_tb)) + /* + * Only a swapped-out slot may be brought into the swap cache. + * Cluster readahead walks raw offset ranges, so it can land on + * slots that are free, bad, or owned by hibernation. + */ + if (!swp_tb_is_shadow(old_tb) || !__swp_tb_get_count(old_tb)) return -ENOENT; - if (shadowp && swp_tb_is_shadow(old_tb)) + if (shadowp) *shadowp = swp_tb_to_shadow(old_tb); if (memcg_id) *memcg_id = __swap_cgroup_get(ci, ci_off); -- 2.48.1