From: Tao Cui Standalone flushes issued by blkdev_issue_flush() are represented as dataless REQ_OP_WRITE | REQ_PREFLUSH bios, which calc_vtime_cost_builtin() prices at zero. The flush component of flush-heavy workloads such as database commits, journal flushes, and metadata sync is thus neither charged nor throttled: a cgroup at 1% weight can issue ~510k flushes per 12s, monopolizing the device while iocost reports zero usage. Price them as pageless random writes (LCOEF_WRANDIO), which provides an approximation of the device time consumed by a flush. For profiles where WRANDIO clamps to zero (ssd_dfl / ssd_fast), use a one-page floor (LCOEF_WPAGE). After this patch, the same 1%-weight cgroup is limited to 24 flushes per 12s; on ext4, write+fsync workloads are correctly accounted through the journal layer (~2.2us per flush on the ssd_fast profile). Fixes: 7caa47151ab2 ("blkcg: implement blk-iocost") Signed-off-by: Tao Cui --- block/blk-iocost.c | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/block/blk-iocost.c b/block/blk-iocost.c index 2745bffcd5ee..abc512532ed9 100644 --- a/block/blk-iocost.c +++ b/block/blk-iocost.c @@ -2533,8 +2533,8 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg, u64 seek_pages = 0; u64 cost = 0; - /* Can't calculate cost for empty bio */ - if (!bio->bi_iter.bi_size) + /* Dataless WRITE|REQ_PREFLUSH (standalone flush) is priced below */ + if (!bio->bi_iter.bi_size && !(bio->bi_opf & REQ_PREFLUSH)) goto out; switch (bio_op(bio)) { @@ -2544,6 +2544,15 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg, coef_page = ioc->params.lcoefs[LCOEF_RPAGE]; break; case REQ_OP_WRITE: + if (!bio->bi_iter.bi_size) { + /* + * dataless WRITE|REQ_PREFLUSH: standalone flush; + * at least one page so fast profiles still charge + */ + cost = max(ioc->params.lcoefs[LCOEF_WRANDIO], + ioc->params.lcoefs[LCOEF_WPAGE]); + goto out; + } coef_seqio = ioc->params.lcoefs[LCOEF_WSEQIO]; coef_randio = ioc->params.lcoefs[LCOEF_WRANDIO]; coef_page = ioc->params.lcoefs[LCOEF_WPAGE]; -- 2.43.0 From: Tao Cui Zone append is a primary write operation for zoned devices; zoned btrfs and f2fs use it for data writes. It is priced at zero, so the zone append portion of zoned workloads runs outside the controller: a 1%-weight cgroup issued 16000 appends at zero cost on a zoned null_blk. A zone append advances the zone write pointer and is therefore sequential from the device's perspective; the actual sector is only returned after completion, so the cursor-based seq/rand classification doesn't apply. Price it as a page-counted sequential write. After this patch, 16000 appends from the same cgroup are charged 533264 usec (33us per append). Also skip the cursor update for ZA bios: bi_sector is the zone start, not the actual write position (which is only returned after completion). Setting the cursor from ZA would misclassify subsequent READ/WRITE bios. Fixes: 7caa47151ab2 ("blkcg: implement blk-iocost") Signed-off-by: Tao Cui --- block/blk-iocost.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-) diff --git a/block/blk-iocost.c b/block/blk-iocost.c index abc512532ed9..d7860ec1ad49 100644 --- a/block/blk-iocost.c +++ b/block/blk-iocost.c @@ -2557,6 +2557,17 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg, coef_randio = ioc->params.lcoefs[LCOEF_WRANDIO]; coef_page = ioc->params.lcoefs[LCOEF_WPAGE]; break; + case REQ_OP_ZONE_APPEND: + /* + * A zone append advances the zone write pointer and is + * therefore sequential from the device's perspective, so + * the cursor-based classification below doesn't apply. + * Compute the full cost here. + */ + if (!is_merge) + cost += ioc->params.lcoefs[LCOEF_WSEQIO]; + cost += pages * ioc->params.lcoefs[LCOEF_WPAGE]; + goto out; default: goto out; } @@ -2717,7 +2728,9 @@ static void ioc_rqos_throttle(struct rq_qos *rqos, struct bio *bio) if (!iocg_activate(iocg, &now)) return; - iocg->cursor = bio_end_sector(bio); + /* ZA bi_sector is zone start, not the write position */ + if (bio_op(bio) != REQ_OP_ZONE_APPEND) + iocg->cursor = bio_end_sector(bio); vtime = atomic64_read(&iocg->vtime); cost = adjust_inuse_and_calc_cost(iocg, vtime, abs_cost, &now); -- 2.43.0 From: Tao Cui ioc_rqos_done() only accounts READ and WRITE completions, so zone append completions are excluded from the latency window: the vrate feedback loop cannot see ZA-induced latency, leaving it unable to respond to device saturation caused by zone appends. Similarly, calc_size_vtime_cost_builtin() does not classify zone append as a write operation. Charging (the cost model) and feedback (the latency window) are separate mechanisms, so this is not covered by the previous patch that prices zone appends. Treat zone append as WRITE for both latency accounting and cost classification. Fixes: 7caa47151ab2 ("blkcg: implement blk-iocost") Signed-off-by: Tao Cui --- block/blk-iocost.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/block/blk-iocost.c b/block/blk-iocost.c index d7860ec1ad49..1bf622e39ce4 100644 --- a/block/blk-iocost.c +++ b/block/blk-iocost.c @@ -2606,6 +2606,7 @@ static void calc_size_vtime_cost_builtin(struct request *rq, struct ioc *ioc, case REQ_OP_READ: *costp = pages * ioc->params.lcoefs[LCOEF_RPAGE]; break; + case REQ_OP_ZONE_APPEND: case REQ_OP_WRITE: *costp = pages * ioc->params.lcoefs[LCOEF_WPAGE]; break; @@ -2876,6 +2877,7 @@ static void ioc_rqos_done(struct rq_qos *rqos, struct request *rq) pidx = QOS_RLAT; rw = READ; break; + case REQ_OP_ZONE_APPEND: case REQ_OP_WRITE: pidx = QOS_WLAT; rw = WRITE; -- 2.43.0 From: Tao Cui The comment says that priority-inversion IOs are "punted to @ioc->aux_iocg", but no aux_iocg field ever existed in struct ioc. The comment was introduced already stale by commit da437b95db83 ("blk-iocost: grab ioc->lock for debt handling"). Update it to describe the current use_debt / iocg->abs_vdebt mechanism. Signed-off-by: Tao Cui --- block/blk-iocost.c | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/block/blk-iocost.c b/block/blk-iocost.c index 1bf622e39ce4..f99c4f5c83fd 100644 --- a/block/blk-iocost.c +++ b/block/blk-iocost.c @@ -2748,10 +2748,11 @@ static void ioc_rqos_throttle(struct rq_qos *rqos, struct bio *bio) /* * We're over budget. This can be handled in two ways. IOs which may - * cause priority inversions are punted to @ioc->aux_iocg and charged as - * debt. Otherwise, the issuer is blocked on @iocg->waitq. Debt handling - * requires @ioc->lock, waitq handling @iocg->waitq.lock. Determine - * whether debt handling is needed and acquire locks accordingly. + * cause priority inversions are issued regardless and charged against + * @iocg->abs_vdebt as debt. Otherwise, the issuer is blocked on + * @iocg->waitq. Debt handling requires @ioc->lock, waitq handling + * @iocg->waitq.lock. Determine whether debt handling is needed and + * acquire locks accordingly. */ use_debt = bio_issue_as_root_blkg(bio) || fatal_signal_pending(current); ioc_locked = use_debt || READ_ONCE(iocg->abs_vdebt); -- 2.43.0