From: Tao Cui blk_rq_stat_sum() folds src into dst but only advances dst->mean and dst->nr_samples, leaving dst->batch untouched. The mean is computed from src->batch (the raw per-cpu sum), so a stat that has already been through one sum carries batch=0; if that aggregated stat is then used as the src of another sum, its samples add nothing to the new mean. iolatency hits exactly that: iolatency_check_latencies() first sums the per-cpu stats into a local stat, then sums that local stat into iolat->cur_stat. After the first sum the local stat has batch=0, so every later window drives cur_stat->mean toward zero. On non-SSD devices it stays at 0, making the latency_sum_ok(&cur_stat) check that gates scaling up always true -- the scale-up hysteresis is effectively defeated. SSD devices use the percentile path and are unaffected. Keep dst->batch in sync across sums so an aggregated stat can be reused as a src. blk_rq_stat.batch is internal to blk_rq_stat_init/_add/_sum (no other reader in the tree), so the wbt and blk-mq consumers, which only read ->mean/->min/->nr_samples, behave as before. Fixes: 34dbad5d26e2 ("blk-stat: convert to callback-based statistics reporting") Signed-off-by: Tao Cui --- block/blk-stat.c | 1 + 1 file changed, 1 insertion(+) diff --git a/block/blk-stat.c b/block/blk-stat.c index de126e1ea5ac..4d4781350083 100644 --- a/block/blk-stat.c +++ b/block/blk-stat.c @@ -36,6 +36,7 @@ void blk_rq_stat_sum(struct blk_rq_stat *dst, struct blk_rq_stat *src) dst->mean = div_u64(src->batch + dst->mean * dst->nr_samples, dst->nr_samples + src->nr_samples); + dst->batch += src->batch; dst->nr_samples += src->nr_samples; } -- 2.43.0