From: Tao Cui Standalone flushes issued by blkdev_issue_flush() are represented as dataless REQ_OP_WRITE | REQ_PREFLUSH bios, which calc_vtime_cost_builtin() prices at zero. The flush component of flush-heavy workloads such as database commits, journal flushes, and metadata sync is thus neither charged nor throttled: a cgroup at 1% weight can issue ~510k flushes per 12s, monopolizing the device while iocost reports zero usage. The same is true for the flush component of data-bearing REQ_OP_WRITE | REQ_PREFLUSH bios, e.g. journal commit writes: they are charged for their data only, and the cache flush the flush machine runs ahead of it is free. Charge the flush component of any REQ_PREFLUSH bio on top of its data cost, priced as a pageless random write (LCOEF_WRANDIO), which provides an approximation of the device time consumed by a flush. For profiles where WRANDIO clamps to zero (ssd_dfl / ssd_fast), use a one-page floor (LCOEF_WPAGE). A dataless flush bio falls out of the switch with zero data cost and picks up the same surcharge, so standalone and pre-flush forms are priced the same way. After this patch, the same 1%-weight cgroup is limited to 24 flushes per 12s; on ext4, write+fsync workloads are correctly accounted through the journal layer (~2.2us per flush on the ssd_fast profile). A standalone flush must also not update iocg->cursor: its bi_sector (usually 0) is not a data position, so setting the cursor from it would misclassify the following READ/WRITE bios, and a zero cursor defeats the !iocg->cursor sentinel in calc_vtime_cost_builtin(). Skip the cursor update for dataless bios. Fixes: 7caa47151ab2 ("blkcg: implement blk-iocost") Signed-off-by: Tao Cui --- Changes in v2: - Skip the iocg->cursor update for dataless flush bios, which would otherwise corrupt the seq/rand classification of the following IOs (reported in review of v1). - Charge the flush component of data-bearing REQ_PREFLUSH bios too; v1 only priced standalone flushes (reported in review of v1). --- block/blk-iocost.c | 20 +++++++++++++++++--- 1 file changed, 17 insertions(+), 3 deletions(-) diff --git a/block/blk-iocost.c b/block/blk-iocost.c index 2745bffcd5eef..082f26d6e27b6 100644 --- a/block/blk-iocost.c +++ b/block/blk-iocost.c @@ -2532,8 +2532,20 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg, u64 pages = max_t(u64, bio_sectors(bio) >> IOC_SECT_TO_PAGE_SHIFT, 1); u64 seek_pages = 0; u64 cost = 0; + u64 flush_cost = 0; - /* Can't calculate cost for empty bio */ + /* + * A WRITE|REQ_PREFLUSH bio carries a flush component: the flush + * machine runs a cache flush for it, either standalone (dataless) + * or ahead of the data. Charge the flush on top of the data cost, + * priced as a pageless random write with a one-page floor so fast + * profiles still charge something. Flush bios are never merged. + */ + if (!is_merge && (bio->bi_opf & REQ_PREFLUSH)) + flush_cost = max(ioc->params.lcoefs[LCOEF_WRANDIO], + ioc->params.lcoefs[LCOEF_WPAGE]); + + /* Can't calculate data cost for empty bio */ if (!bio->bi_iter.bi_size) goto out; @@ -2566,7 +2578,7 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg, } cost += pages * coef_page; out: - *costp = cost; + *costp = cost + flush_cost; } static u64 calc_vtime_cost(struct bio *bio, struct ioc_gq *iocg, bool is_merge) @@ -2708,7 +2720,9 @@ static void ioc_rqos_throttle(struct rq_qos *rqos, struct bio *bio) if (!iocg_activate(iocg, &now)) return; - iocg->cursor = bio_end_sector(bio); + /* dataless bios have no meaningful position for seq/rand detection */ + if (bio->bi_iter.bi_size) + iocg->cursor = bio_end_sector(bio); vtime = atomic64_read(&iocg->vtime); cost = adjust_inuse_and_calc_cost(iocg, vtime, abs_cost, &now); -- 2.43.0