AI Trajectory


Trajectory: (Expand/Collapse All)
Seq Timestamp Type Name Duration
0/0 2026/10/02 08:23 flow patch-triage
1m Args:
null

Results:
{
  "EnableConfigs": [
    "BLK_DEV_UBLK"
  ],
  "FocusSymbols": [
    "ublk_stop_dev",
    "ublk_ctrl_start_dev",
    "ublk_ch_uring_cmd_local",
    "ublk_uring_cmd_cancel_fn",
    "ublk_ch_release_work_fn",
    "ublk_batch_attach",
    "ublk_reset_ch_dev",
    "ublk_ch_open"
  ],
  "KMSANReasoning": "The patch series modifies the ublk driver (drivers/block/ublk_drv.c) and adds selftests to address race conditions and synchronization issues during device cancellation and command completion (such as STOP_DEV racing with START_DEV, FETCH, and process exits). Specifically, it manages internal state flags (UB_STATE_STOPPING, ub-\u003ecanceling), file references (ub-\u003ech_file), memory barriers, and the timing of io_uring_cmd_mark_cancelable() and io_uring_cmd_done().\n\nAll modified kernel structures (struct ublk_device and struct ublk_batch_fetch_cmd) are allocated using zeroing allocators (kzalloc_flex and kzalloc_obj). Furthermore, INIT_LIST_HEAD(\u0026fcmd-\u003enode) is added to ensure proper initialization before cancellation checks. The patch does not introduce or modify any data copied to user space, does not alter buffer length or bounds calculations, and introduces no uninitialized memory reads or information leaks.\n\nAny memory safety or concurrency risks associated with these changes involve use-after-free, NULL pointer dereferences, or lock/barrier ordering, which are fully covered by standard KASAN and LOCKDEP. Therefore, a dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false,
  "Reasoning": "The patch modifies drivers/block/ublk_drv.c to fix races and state management between STOP_DEV, START_DEV, FETCH_REQ, and device release/cancellation in the ublk driver. It introduces a new device state (UB_STATE_STOPPING), adjusts file reference counting and memory ordering, alters cancelation routines, and modifies assertion checks. The modified code is directly reachable from userspace via io_uring commands on /dev/ublk-control and /dev/ublkc*, making it highly relevant for fuzzing concurrency and lifecycle edge cases.",
  "WorthFuzzing": true
}

1/1 2026/10/02 08:23 action read-patch-diff
0m Args:
null

Results:
{
  "PatchDiff": "commit 30c2ad39819c25e745543e6ae191d461a2c8251e\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate:   Fri Oct 2 08:23:10 2026 +0000\n\n    syz-cluster: applied patch under review\n\ndiff --git a/drivers/block/ublk_drv.c b/drivers/block/ublk_drv.c\nindex 66eb55e7162e5..f57d544c1da23 100644\n--- a/drivers/block/ublk_drv.c\n+++ b/drivers/block/ublk_drv.c\n@@ -321,6 +321,8 @@ struct ublk_device {\n #define UB_STATE_OPEN\t\t0\n #define UB_STATE_USED\t\t1\n #define UB_STATE_DELETED\t2\n+/* STOP_DEV canceled the server's commands, until its release */\n+#define UB_STATE_STOPPING\t3\n \tunsigned long\t\tstate;\n \tint\t\t\tub_number;\n \n@@ -334,6 +336,13 @@ struct ublk_device {\n \tu16\t\t\tnr_queue_ready;\n \tbool \t\t\tunprivileged_daemons;\n \tstruct mutex cancel_mutex;\n+\t/* the open /dev/ublkcN, protected by cancel_mutex */\n+\tstruct file *ch_file;\n+\t/*\n+\t * A cancel started in this FETCH round. Set by ublk_set_canceling(),\n+\t * cleared only by ublk_reset_ch_dev() when a new round starts. While\n+\t * it is set, no queue clears its -\u003ecanceling.\n+\t */\n \tbool canceling;\n \tpid_t \tublksrv_tgid;\n \tstruct delayed_work\texit_work;\n@@ -810,6 +819,8 @@ ublk_batch_alloc_fcmd(struct io_uring_cmd *cmd)\n \tif (fcmd) {\n \t\tfcmd-\u003ecmd = cmd;\n \t\tfcmd-\u003ebuf_group = READ_ONCE(cmd-\u003esqe-\u003ebuf_index);\n+\t\t/* a cancel may look at it before it is linked */\n+\t\tINIT_LIST_HEAD(\u0026fcmd-\u003enode);\n \t}\n \treturn fcmd;\n }\n@@ -2399,6 +2410,9 @@ static int ublk_ch_open(struct inode *inode, struct file *filp)\n \t\treturn -EBUSY;\n \tfilp-\u003eprivate_data = ub;\n \tub-\u003eublksrv_tgid = current-\u003etgid;\n+\tmutex_lock(\u0026ub-\u003ecancel_mutex);\n+\tub-\u003ech_file = filp;\n+\tmutex_unlock(\u0026ub-\u003ecancel_mutex);\n \treturn 0;\n }\n \n@@ -2415,11 +2429,17 @@ static void ublk_reset_ch_dev(struct ublk_device *ub)\n \t\tspin_unlock(\u0026ubq-\u003ecancel_lock);\n \t}\n \n+\t/* a new FETCH round starts, the queues stay canceling until ready */\n+\tmutex_lock(\u0026ub-\u003ecancel_mutex);\n+\tub-\u003ecanceling = false;\n+\tmutex_unlock(\u0026ub-\u003ecancel_mutex);\n+\n \t/* set to NULL, otherwise new tasks cannot mmap io_cmd_buf */\n \tub-\u003emm = NULL;\n \tub-\u003enr_queue_ready = 0;\n \tub-\u003eunprivileged_daemons = false;\n \tub-\u003eublksrv_tgid = -1;\n+\tclear_bit(UB_STATE_STOPPING, \u0026ub-\u003estate);\n }\n \n static struct gendisk *ublk_get_disk(struct ublk_device *ub)\n@@ -2539,12 +2559,15 @@ static void ublk_ch_release_work_fn(struct work_struct *work)\n \t}\n \n \t/*\n-\t * disk isn't attached yet, either device isn't live, or it has\n-\t * been removed already, so we needn't to do anything\n+\t * No disk: the device isn't live, or it has been removed already.\n+\t * There are no requests to abort, but the round still has to be\n+\t * reset, so that a new server can fetch and start the device.\n \t */\n \tdisk = ublk_get_disk(ub);\n-\tif (!disk)\n-\t\tgoto out;\n+\tif (!disk) {\n+\t\tmutex_lock(\u0026ub-\u003emutex);\n+\t\tgoto reset;\n+\t}\n \n \t/*\n \t * All uring_cmd are done now, so abort any request outstanding to\n@@ -2576,7 +2599,7 @@ static void ublk_ch_release_work_fn(struct work_struct *work)\n \n \t/* double check after grabbing lock */\n \tif (!ub-\u003eub_disk)\n-\t\tgoto unlock;\n+\t\tgoto reset;\n \n \t/*\n \t * Transition the device to the nosrv state. What exactly this\n@@ -2602,13 +2625,14 @@ static void ublk_ch_release_work_fn(struct work_struct *work)\n \t\t\t\tWRITE_ONCE(ublk_get_queue(ub, i)-\u003efail_io, true);\n \t\t}\n \t}\n-unlock:\n+reset:\n+\t/*\n+\t * All uring_cmd has been done now, reset device \u0026 ubq. Under\n+\t * ub-\u003emutex, so START_DEV sees the round either ready or reset.\n+\t */\n+\tublk_reset_ch_dev(ub);\n \tmutex_unlock(\u0026ub-\u003emutex);\n \tublk_put_disk(disk);\n-\n-\t/* all uring_cmd has been done now, reset device \u0026 ubq */\n-\tublk_reset_ch_dev(ub);\n-out:\n \tclear_bit(UB_STATE_OPEN, \u0026ub-\u003estate);\n \n \t/* put the reference grabbed in ublk_ch_release() */\n@@ -2619,6 +2643,9 @@ static int ublk_ch_release(struct inode *inode, struct file *filp)\n {\n \tstruct ublk_device *ub = filp-\u003eprivate_data;\n \n+\tmutex_lock(\u0026ub-\u003ecancel_mutex);\n+\tub-\u003ech_file = NULL;\n+\tmutex_unlock(\u0026ub-\u003ecancel_mutex);\n \t/*\n \t * Grab ublk device reference, so it won't be gone until we are\n \t * really released from work function.\n@@ -2790,7 +2817,8 @@ static void ublk_cancel_cmd(struct ublk_queue *ubq, u16 tag,\n \tdone = !!(io-\u003eflags \u0026 UBLK_IO_FLAG_CANCELED);\n \tif (!done) {\n \t\tio-\u003eflags |= UBLK_IO_FLAG_CANCELED;\n-\t\tcmd = io-\u003ecmd;\n+\t\t/* dependency ordered against smp_wmb() in ublk_prep_cancel() */\n+\t\tcmd = READ_ONCE(io-\u003ecmd);\n \t\tio-\u003ecmd = NULL;\n \t}\n \tspin_unlock(\u0026ubq-\u003ecancel_lock);\n@@ -2879,6 +2907,7 @@ static void ublk_uring_cmd_cancel_fn(struct io_uring_cmd *cmd,\n {\n \tstruct ublk_uring_cmd_pdu *pdu = ublk_get_uring_cmd_pdu(cmd);\n \tstruct ublk_queue *ubq = pdu-\u003eubq;\n+\tstruct io_uring_cmd *cur;\n \tstruct task_struct *task;\n \tstruct ublk_io *io;\n \n@@ -2895,7 +2924,9 @@ static void ublk_uring_cmd_cancel_fn(struct io_uring_cmd *cmd,\n \n \tublk_start_cancel(ubq-\u003edev);\n \n-\tWARN_ON_ONCE(io-\u003ecmd != cmd);\n+\t/* NULL if STOP_DEV's cancel took it meanwhile */\n+\tcur = READ_ONCE(io-\u003ecmd);\n+\tWARN_ON_ONCE(cur \u0026\u0026 cur != cmd);\n \tublk_cancel_cmd(ubq, pdu-\u003etag, issue_flags);\n }\n \n@@ -3006,13 +3037,54 @@ static void ublk_stop_dev_unlocked(struct ublk_device *ub)\n \tput_disk(disk);\n }\n \n+static struct file *ublk_get_ch_file(struct ublk_device *ub)\n+{\n+\tstruct file *file;\n+\n+\tmutex_lock(\u0026ub-\u003ecancel_mutex);\n+\tfile = ub-\u003ech_file;\n+\tif (file \u0026\u0026 !file_ref_get(\u0026file-\u003ef_ref))\n+\t\tfile = NULL;\n+\tmutex_unlock(\u0026ub-\u003ecancel_mutex);\n+\treturn file;\n+}\n+\n static void ublk_stop_dev(struct ublk_device *ub)\n {\n+\tstruct file *file;\n+\n+\t/*\n+\t * FETCH, PREP and START_DEV take ub-\u003emutex. If a server has\n+\t * /dev/ublkcN open, set STOPPING, which turns it away, and hold a\n+\t * reference on the file: the server's release, whose reset clears\n+\t * STOPPING and lets a new server attach, can't run before we drop it.\n+\t * So the cancel can run after the unlock and only meets this server's\n+\t * commands; it has to: io_uring_cmd_done() may take uring_lock, under\n+\t * which FETCH takes ub-\u003emutex.\n+\t */\n \tmutex_lock(\u0026ub-\u003emutex);\n \tublk_stop_dev_unlocked(ub);\n-\tmutex_unlock(\u0026ub-\u003emutex);\n \tcancel_work_sync(\u0026ub-\u003epartition_scan_work);\n+\tfile = ublk_get_ch_file(ub);\n+\tif (file) {\n+\t\t/*\n+\t\t * The server has to close /dev/ublkcN before this device can\n+\t\t * be started again: the reset in its release clears STOPPING.\n+\t\t */\n+\t\tset_bit(UB_STATE_STOPPING, \u0026ub-\u003estate);\n+\t\t/* for wake_up_var() below, see wake_up_bit() */\n+\t\tsmp_mb__after_atomic();\n+\t}\n+\tmutex_unlock(\u0026ub-\u003emutex);\n+\n+\t/* no server: nothing to cancel */\n+\tif (!file)\n+\t\treturn;\n+\n+\t/* wake a START_DEV waiting for the device to get ready */\n+\twake_up_var(\u0026ub-\u003enr_queue_ready);\n \tublk_cancel_dev(ub);\n+\t__fput_sync(file);\n }\n \n static void ublk_reset_io_flags(struct ublk_queue *ubq, struct ublk_io *io)\n@@ -3024,11 +3096,19 @@ static void ublk_reset_io_flags(struct ublk_queue *ubq, struct ublk_io *io)\n }\n \n /* reset per-queue io flags */\n-static void ublk_queue_reset_io_flags(struct ublk_queue *ubq)\n+static void ublk_queue_reset_io_flags(struct ublk_device *ub,\n+\t\t\t\t      struct ublk_queue *ubq)\n {\n-\tspin_lock(\u0026ubq-\u003ecancel_lock);\n-\tubq-\u003ecanceling = false;\n-\tspin_unlock(\u0026ubq-\u003ecancel_lock);\n+\t/*\n+\t * A cancel in this FETCH round took a command which still counts as\n+\t * ready, so the queue has to stay canceling. ub-\u003ecanceling is set\n+\t * under cancel_mutex before any command is taken: either we see it\n+\t * here, or the cancel marks this queue again later.\n+\t */\n+\tmutex_lock(\u0026ub-\u003ecancel_mutex);\n+\tif (!ub-\u003ecanceling)\n+\t\tubq-\u003ecanceling = false;\n+\tmutex_unlock(\u0026ub-\u003ecancel_mutex);\n \tubq-\u003efail_io = false;\n \tubq-\u003eforce_abort = false;\n }\n@@ -3051,24 +3131,17 @@ static void ublk_mark_io_ready(struct ublk_device *ub, u16 q_id,\n \t\tub-\u003enr_queue_ready++;\n \n \t\t/*\n-\t\t * Reset queue flags as soon as this queue is ready.\n-\t\t * This clears the canceling flag, allowing batch FETCH commands\n-\t\t * to succeed during recovery without waiting for all queues.\n+\t\t * Reset queue flags as soon as this queue is ready. Unless\n+\t\t * this round saw a cancel, this clears the canceling flag,\n+\t\t * allowing batch FETCH commands to succeed during recovery\n+\t\t * without waiting for all queues.\n \t\t */\n-\t\tublk_queue_reset_io_flags(ubq);\n+\t\tublk_queue_reset_io_flags(ub, ubq);\n \t}\n \n-\t/* Check if all queues are ready */\n-\tif (ublk_dev_ready(ub)) {\n-\t\t/*\n-\t\t * All queues ready - clear device-level canceling flag\n-\t\t * and wake ublk_dev_ready() waiters.\n-\t\t */\n-\t\tmutex_lock(\u0026ub-\u003ecancel_mutex);\n-\t\tub-\u003ecanceling = false;\n-\t\tmutex_unlock(\u0026ub-\u003ecancel_mutex);\n+\t/* All queues ready - wake ublk_dev_ready() waiters */\n+\tif (ublk_dev_ready(ub))\n \t\twake_up_var(\u0026ub-\u003enr_queue_ready);\n-\t}\n }\n \n static inline int ublk_check_cmd_op(u32 cmd_op)\n@@ -3152,6 +3225,12 @@ ublk_fill_io_cmd(struct ublk_io *io, struct io_uring_cmd *cmd)\n \treturn req;\n }\n \n+/*\n+ * Call before ublk_fill_io_cmd() publishes @cmd in io-\u003ecmd: a control-path\n+ * cancel may complete any command found there, and io_uring_cmd_done() only\n+ * takes it off the cancelable list if it is marked already. The handlers\n+ * hold uring_lock, so marking takes no lock.\n+ */\n static inline void ublk_prep_cancel(struct io_uring_cmd *cmd,\n \t\t\t\t    unsigned int issue_flags,\n \t\t\t\t    struct ublk_queue *ubq, u16 tag)\n@@ -3165,6 +3244,8 @@ static inline void ublk_prep_cancel(struct io_uring_cmd *cmd,\n \tpdu-\u003eubq = ubq;\n \tpdu-\u003etag = tag;\n \tio_uring_cmd_mark_cancelable(cmd, issue_flags);\n+\t/* pairs with the cancel loading cmd from io-\u003ecmd, then cmd-\u003eflags */\n+\tsmp_wmb();\n }\n \n static void ublk_io_release(void *priv)\n@@ -3269,6 +3350,9 @@ static int ublk_check_fetch_buf(const struct ublk_device *ub, __u64 buf_addr)\n static int __ublk_fetch(struct io_uring_cmd *cmd, struct ublk_device *ub,\n \t\t\tstruct ublk_io *io, u16 q_id)\n {\n+\tif (test_bit(UB_STATE_STOPPING, \u0026ub-\u003estate))\n+\t\treturn UBLK_IO_RES_ABORT;\n+\n \t/* UBLK_IO_FETCH_REQ is only allowed before dev is setup */\n \tif (ublk_dev_ready(ub))\n \t\treturn -EBUSY;\n@@ -3412,11 +3496,11 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,\n \t\tret = ublk_check_fetch_buf(ub, addr);\n \t\tif (ret)\n \t\t\tgoto out;\n+\t\t/* before ublk_fetch() publishes io-\u003ecmd, see ublk_prep_cancel() */\n+\t\tublk_prep_cancel(cmd, issue_flags, ubq, tag);\n \t\tret = ublk_fetch(cmd, ub, io, addr, q_id);\n \t\tif (ret)\n-\t\t\tgoto out;\n-\n-\t\tublk_prep_cancel(cmd, issue_flags, ubq, tag);\n+\t\t\tgoto out_done;\n \t\treturn -EIOCBQUEUED;\n \t}\n \n@@ -3460,6 +3544,7 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,\n \t\tif (ret)\n \t\t\tgoto out;\n \t\tio-\u003eres = result;\n+\t\tublk_prep_cancel(cmd, issue_flags, ubq, tag);\n \t\treq = ublk_fill_io_cmd(io, cmd);\n \t\tublk_apply_io_buf(ub, io, cmd, addr, \u0026auto_buf, \u0026buf_idx);\n \t\tif (buf_idx != UBLK_INVALID_BUF_IDX)\n@@ -3478,19 +3563,24 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,\n \t\t * uring_cmd active first and prepare for handling new requeued\n \t\t * request\n \t\t */\n+\t\tublk_prep_cancel(cmd, issue_flags, ubq, tag);\n \t\treq = ublk_fill_io_cmd(io, cmd);\n \t\tio-\u003ebuf.addr = addr;\n \t\tif (likely(ublk_get_data(ubq, io, req))) {\n \t\t\t__ublk_prep_compl_io_cmd(io, req);\n-\t\t\treturn UBLK_IO_RES_OK;\n+\t\t\tret = UBLK_IO_RES_OK;\n+\t\t\tgoto out_done;\n \t\t}\n \t\tbreak;\n \tdefault:\n \t\tgoto out;\n \t}\n-\tublk_prep_cancel(cmd, issue_flags, ubq, tag);\n \treturn -EIOCBQUEUED;\n \n+ out_done:\n+\t/* marked cancelable: complete through io_uring_cmd_done() */\n+\tio_uring_cmd_done(cmd, ret, issue_flags);\n+\treturn -EIOCBQUEUED;\n  out:\n \tpr_devel(\"%s: complete: cmd op %d, tag %d ret %x io_flags %x\\n\",\n \t\t\t__func__, cmd_op, tag, ret, io ? io-\u003eflags : 0);\n@@ -3889,6 +3979,15 @@ static int ublk_batch_attach(struct ublk_queue *ubq,\n \tbool free = false;\n \tstruct ublk_uring_cmd_pdu *pdu = ublk_get_uring_cmd_pdu(data-\u003ecmd);\n \n+\t/*\n+\t * Mark it cancelable before linking it into fcmd_head, where a cancel\n+\t * from the control path can take and complete it: see\n+\t * ublk_prep_cancel(). evts_lock orders the mark before the link.\n+\t */\n+\tpdu-\u003eubq = ubq;\n+\tpdu-\u003efcmd = fcmd;\n+\tio_uring_cmd_mark_cancelable(fcmd-\u003ecmd, data-\u003eissue_flags);\n+\n \tspin_lock(\u0026ubq-\u003eevts_lock);\n \tif (unlikely(ubq-\u003eforce_abort || ubq-\u003ecanceling)) {\n \t\tfree = true;\n@@ -3899,14 +3998,12 @@ static int ublk_batch_attach(struct ublk_queue *ubq,\n \tspin_unlock(\u0026ubq-\u003eevts_lock);\n \n \tif (unlikely(free)) {\n+\t\t/* off the cancelable list first, then nothing can see fcmd */\n+\t\tio_uring_cmd_done(data-\u003ecmd, -ENODEV, data-\u003eissue_flags);\n \t\tublk_batch_free_fcmd(fcmd);\n-\t\treturn -ENODEV;\n+\t\treturn -EIOCBQUEUED;\n \t}\n \n-\tpdu-\u003eubq = ubq;\n-\tpdu-\u003efcmd = fcmd;\n-\tio_uring_cmd_mark_cancelable(fcmd-\u003ecmd, data-\u003eissue_flags);\n-\n \tif (!new_fcmd)\n \t\tgoto out;\n \n@@ -3914,9 +4011,12 @@ static int ublk_batch_attach(struct ublk_queue *ubq,\n \t * If the two fetch commands are originated from same io_ring_ctx,\n \t * run batch dispatch directly. Otherwise, schedule task work for\n \t * doing it.\n+\t *\n+\t * Use data-\u003ecmd, not fcmd-\u003ecmd: once fcmd is linked and not active,\n+\t * a cancel from the control path may complete and free it.\n \t */\n \tif (io_uring_cmd_ctx_handle(new_fcmd-\u003ecmd) ==\n-\t\t\tio_uring_cmd_ctx_handle(fcmd-\u003ecmd)) {\n+\t\t\tio_uring_cmd_ctx_handle(data-\u003ecmd)) {\n \t\tdata-\u003ecmd = new_fcmd-\u003ecmd;\n \t\tublk_batch_dispatch(ubq, data, new_fcmd);\n \t} else {\n@@ -4431,21 +4531,27 @@ static bool ublk_validate_user_pid(struct ublk_device *ub, pid_t ublksrv_pid)\n \treturn ub-\u003eublksrv_tgid == ublksrv_pid;\n }\n \n+static bool ublk_dev_ready_or_stopping(const struct ublk_device *ub)\n+{\n+\treturn ublk_dev_ready(ub) || test_bit(UB_STATE_STOPPING, \u0026ub-\u003estate);\n+}\n+\n /*\n- * Wait until all queues have fetched their I/O commands, and return with\n- * ub-\u003emutex held and readiness guaranteed: then every queue's -\u003ecanceling\n- * is cleared. Ready may regress between wakeup and mutex_lock() (F_BATCH\n- * UNPREP, daemon death), so re-check it under the mutex and wait again.\n+ * Wait until all queues have fetched their I/O commands, or STOP_DEV set\n+ * UB_STATE_STOPPING, and return with ub-\u003emutex held. The queues stay\n+ * canceling if this round saw a cancel, see ublk_queue_reset_io_flags().\n+ * Ready may regress between wakeup and mutex_lock() (F_BATCH UNPREP,\n+ * daemon death), so re-check it under the mutex and wait again.\n  */\n static int ublk_wait_dev_ready_and_lock(struct ublk_device *ub)\n {\n \twhile (true) {\n \t\tif (wait_var_event_interruptible(\u0026ub-\u003enr_queue_ready,\n-\t\t\t\t\t\t ublk_dev_ready(ub)))\n+\t\t\t\t\t\t ublk_dev_ready_or_stopping(ub)))\n \t\t\treturn -EINTR;\n \n \t\tmutex_lock(\u0026ub-\u003emutex);\n-\t\tif (ublk_dev_ready(ub))\n+\t\tif (ublk_dev_ready_or_stopping(ub))\n \t\t\treturn 0;\n \t\tmutex_unlock(\u0026ub-\u003emutex);\n \t}\n@@ -4545,6 +4651,10 @@ static int ublk_ctrl_start_dev(struct ublk_device *ub,\n \t\tret = -EEXIST;\n \t\tgoto out_unlock;\n \t}\n+\tif (test_bit(UB_STATE_STOPPING, \u0026ub-\u003estate)) {\n+\t\tret = -EBUSY;\n+\t\tgoto out_unlock;\n+\t}\n \n \tdisk = blk_mq_alloc_disk(\u0026ub-\u003etag_set, \u0026lim, NULL);\n \tif (IS_ERR(disk)) {\ndiff --git a/tools/testing/selftests/ublk/.gitignore b/tools/testing/selftests/ublk/.gitignore\nindex e17bd28f27e04..d8e93fef7fcd3 100644\n--- a/tools/testing/selftests/ublk/.gitignore\n+++ b/tools/testing/selftests/ublk/.gitignore\n@@ -3,3 +3,4 @@\n /tools\n kublk\n metadata_size\n+ublk_cancel_ready\ndiff --git a/tools/testing/selftests/ublk/Makefile b/tools/testing/selftests/ublk/Makefile\nindex 37883e9d50ec3..fc64b8f028332 100644\n--- a/tools/testing/selftests/ublk/Makefile\n+++ b/tools/testing/selftests/ublk/Makefile\n@@ -19,6 +19,7 @@ TEST_PROGS += test_generic_12.sh\n TEST_PROGS += test_generic_13.sh\n TEST_PROGS += test_generic_16.sh\n TEST_PROGS += test_generic_17.sh\n+TEST_PROGS += test_generic_18.sh\n \n TEST_PROGS += test_batch_01.sh\n TEST_PROGS += test_batch_02.sh\n@@ -76,13 +77,14 @@ TEST_FILES := settings\n TEST_FILES += test_common.sh\n TEST_FILES += trace\n \n-TEST_GEN_PROGS_EXTENDED = kublk metadata_size\n-STANDALONE_UTILS := metadata_size.c\n+TEST_GEN_PROGS_EXTENDED = kublk metadata_size ublk_cancel_ready\n+STANDALONE_UTILS := metadata_size.c ublk_cancel_ready.c\n \n LOCAL_HDRS += $(wildcard *.h)\n include ../lib.mk\n \n $(OUTPUT)/kublk: $(filter-out $(STANDALONE_UTILS),$(wildcard *.c))\n+$(OUTPUT)/ublk_cancel_ready: ublk_cancel_ready.c ctrl.c\n \n check:\n \tshellcheck -x -f gcc *.sh\ndiff --git a/tools/testing/selftests/ublk/ctrl.c b/tools/testing/selftests/ublk/ctrl.c\nnew file mode 100644\nindex 0000000000000..53d54799f9097\n--- /dev/null\n+++ b/tools/testing/selftests/ublk/ctrl.c\n@@ -0,0 +1,268 @@\n+// SPDX-License-Identifier: GPL-2.0\n+\n+/* ublk control commands, shared by kublk and ublk_cancel_ready */\n+\n+#include \"kublk.h\"\n+\n+static void ublk_ctrl_init_cmd(struct ublk_dev *dev,\n+\t\tstruct io_uring_sqe *sqe,\n+\t\tstruct ublk_ctrl_cmd_data *data)\n+{\n+\tstruct ublksrv_ctrl_dev_info *info = \u0026dev-\u003edev_info;\n+\tstruct ublksrv_ctrl_cmd *cmd = (struct ublksrv_ctrl_cmd *)ublk_get_sqe_cmd(sqe);\n+\n+\tsqe-\u003efd = dev-\u003ectrl_fd;\n+\tsqe-\u003eopcode = IORING_OP_URING_CMD;\n+\tsqe-\u003eioprio = 0;\n+\n+\tif (data-\u003eflags \u0026 CTRL_CMD_HAS_BUF) {\n+\t\tcmd-\u003eaddr = data-\u003eaddr;\n+\t\tcmd-\u003elen = data-\u003elen;\n+\t}\n+\n+\tif (data-\u003eflags \u0026 CTRL_CMD_HAS_DATA)\n+\t\tcmd-\u003edata[0] = data-\u003edata[0];\n+\n+\tcmd-\u003edev_id = info-\u003edev_id;\n+\tcmd-\u003equeue_id = -1;\n+\n+\tublk_set_sqe_cmd_op(sqe, data-\u003ecmd_op);\n+\n+\tio_uring_sqe_set_data(sqe, cmd);\n+}\n+\n+int __ublk_ctrl_cmd(struct ublk_dev *dev,\n+\t\tstruct ublk_ctrl_cmd_data *data)\n+{\n+\tstruct io_uring_sqe *sqe;\n+\tstruct io_uring_cqe *cqe;\n+\tint ret = -EINVAL;\n+\n+\tsqe = io_uring_get_sqe(\u0026dev-\u003ering);\n+\tif (!sqe) {\n+\t\tublk_err(\"%s: can't get sqe ret %d\\n\", __func__, ret);\n+\t\treturn ret;\n+\t}\n+\n+\tublk_ctrl_init_cmd(dev, sqe, data);\n+\n+\tret = io_uring_submit(\u0026dev-\u003ering);\n+\tif (ret \u003c 0) {\n+\t\tublk_err(\"uring submit ret %d\\n\", ret);\n+\t\treturn ret;\n+\t}\n+\n+\tret = io_uring_wait_cqe(\u0026dev-\u003ering, \u0026cqe);\n+\tif (ret \u003c 0) {\n+\t\tublk_err(\"wait cqe: %s\\n\", strerror(-ret));\n+\t\treturn ret;\n+\t}\n+\tio_uring_cqe_seen(\u0026dev-\u003ering, cqe);\n+\n+\treturn cqe-\u003eres;\n+}\n+\n+void ublk_ctrl_deinit(struct ublk_dev *dev)\n+{\n+\tio_uring_queue_exit(\u0026dev-\u003ering);\n+\tclose(dev-\u003ectrl_fd);\n+\tfree(dev);\n+}\n+\n+struct ublk_dev *ublk_ctrl_init(void)\n+{\n+\tstruct ublk_dev *dev = (struct ublk_dev *)calloc(1, sizeof(*dev));\n+\tstruct ublksrv_ctrl_dev_info *info;\n+\tint ret;\n+\n+\tif (!dev)\n+\t\treturn NULL;\n+\tinfo = \u0026dev-\u003edev_info;\n+\tdev-\u003ectrl_fd = open(CTRL_DEV, O_RDWR);\n+\tif (dev-\u003ectrl_fd \u003c 0) {\n+\t\tfree(dev);\n+\t\treturn NULL;\n+\t}\n+\n+\tinfo-\u003emax_io_buf_bytes = UBLK_IO_MAX_BYTES;\n+\n+\tret = ublk_setup_ring(\u0026dev-\u003ering, UBLK_CTRL_RING_DEPTH,\n+\t\t\tUBLK_CTRL_RING_DEPTH, IORING_SETUP_SQE128);\n+\tif (ret \u003c 0) {\n+\t\tublk_err(\"queue_init: %s\\n\", strerror(-ret));\n+\t\tclose(dev-\u003ectrl_fd);\n+\t\tfree(dev);\n+\t\treturn NULL;\n+\t}\n+\tdev-\u003enr_fds = 1;\n+\n+\treturn dev;\n+}\n+\n+int ublk_ctrl_stop_dev(struct ublk_dev *dev)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_STOP_DEV,\n+\t};\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_try_stop_dev(struct ublk_dev *dev)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_TRY_STOP_DEV,\n+\t};\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_start_dev(struct ublk_dev *dev,\n+\t\tint daemon_pid)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_START_DEV,\n+\t\t.flags\t= CTRL_CMD_HAS_DATA,\n+\t};\n+\n+\tdev-\u003edev_info.ublksrv_pid = data.data[0] = daemon_pid;\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_start_user_recovery(struct ublk_dev *dev)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_START_USER_RECOVERY,\n+\t};\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_END_USER_RECOVERY,\n+\t\t.flags\t= CTRL_CMD_HAS_DATA,\n+\t};\n+\n+\tdev-\u003edev_info.ublksrv_pid = data.data[0] = daemon_pid;\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_add_dev(struct ublk_dev *dev)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_ADD_DEV,\n+\t\t.flags\t= CTRL_CMD_HAS_BUF,\n+\t\t.addr = (__u64) (uintptr_t) \u0026dev-\u003edev_info,\n+\t\t.len = sizeof(struct ublksrv_ctrl_dev_info),\n+\t};\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_del_dev(struct ublk_dev *dev)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op = UBLK_U_CMD_DEL_DEV,\n+\t\t.flags = 0,\n+\t};\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_get_info(struct ublk_dev *dev)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_GET_DEV_INFO,\n+\t\t.flags\t= CTRL_CMD_HAS_BUF,\n+\t\t.addr = (__u64) (uintptr_t) \u0026dev-\u003edev_info,\n+\t\t.len = sizeof(struct ublksrv_ctrl_dev_info),\n+\t};\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_set_params(struct ublk_dev *dev,\n+\t\tstruct ublk_params *params)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_SET_PARAMS,\n+\t\t.flags\t= CTRL_CMD_HAS_BUF,\n+\t\t.addr = (__u64) (uintptr_t) params,\n+\t\t.len = sizeof(*params),\n+\t};\n+\tparams-\u003elen = sizeof(*params);\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_get_params(struct ublk_dev *dev,\n+\t\tstruct ublk_params *params)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_GET_PARAMS,\n+\t\t.flags\t= CTRL_CMD_HAS_BUF,\n+\t\t.addr = (__u64)params,\n+\t\t.len = sizeof(*params),\n+\t};\n+\n+\tparams-\u003elen = sizeof(*params);\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_get_features(struct ublk_dev *dev,\n+\t\t__u64 *features)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_GET_FEATURES,\n+\t\t.flags\t= CTRL_CMD_HAS_BUF,\n+\t\t.addr = (__u64) (uintptr_t) features,\n+\t\t.len = sizeof(*features),\n+\t};\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_update_size(struct ublk_dev *dev,\n+\t\t__u64 nr_sects)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_UPDATE_SIZE,\n+\t\t.flags\t= CTRL_CMD_HAS_DATA,\n+\t};\n+\n+\tdata.data[0] = nr_sects;\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_quiesce_dev(struct ublk_dev *dev, unsigned int timeout_ms)\n+{\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op\t= UBLK_U_CMD_QUIESCE_DEV,\n+\t\t.flags\t= CTRL_CMD_HAS_DATA,\n+\t};\n+\n+\tdata.data[0] = timeout_ms;\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\n+\n+int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,\n+\t\t      __u32 flags)\n+{\n+\tstruct ublk_shmem_buf_reg buf_reg = {\n+\t\t.addr = (unsigned long)addr,\n+\t\t.len = size,\n+\t\t.flags = flags,\n+\t};\n+\tstruct ublk_ctrl_cmd_data data = {\n+\t\t.cmd_op = UBLK_U_CMD_REG_BUF,\n+\t\t.flags = CTRL_CMD_HAS_BUF,\n+\t\t.addr = (unsigned long)\u0026buf_reg,\n+\t\t.len = sizeof(buf_reg),\n+\t};\n+\n+\treturn __ublk_ctrl_cmd(dev, \u0026data);\n+}\ndiff --git a/tools/testing/selftests/ublk/kublk.c b/tools/testing/selftests/ublk/kublk.c\nindex 2400b46157664..15ba060ce7a3c 100644\n--- a/tools/testing/selftests/ublk/kublk.c\n+++ b/tools/testing/selftests/ublk/kublk.c\n@@ -37,203 +37,6 @@ static const struct ublk_tgt_ops *ublk_find_tgt(const char *name)\n \treturn NULL;\n }\n \n-static inline int ublk_setup_ring(struct io_uring *r, int depth,\n-\t\tint cq_depth, unsigned flags)\n-{\n-\tstruct io_uring_params p;\n-\n-\tmemset(\u0026p, 0, sizeof(p));\n-\tp.flags = flags | IORING_SETUP_CQSIZE;\n-\tp.cq_entries = cq_depth;\n-\n-\treturn io_uring_queue_init_params(depth, r, \u0026p);\n-}\n-\n-static void ublk_ctrl_init_cmd(struct ublk_dev *dev,\n-\t\tstruct io_uring_sqe *sqe,\n-\t\tstruct ublk_ctrl_cmd_data *data)\n-{\n-\tstruct ublksrv_ctrl_dev_info *info = \u0026dev-\u003edev_info;\n-\tstruct ublksrv_ctrl_cmd *cmd = (struct ublksrv_ctrl_cmd *)ublk_get_sqe_cmd(sqe);\n-\n-\tsqe-\u003efd = dev-\u003ectrl_fd;\n-\tsqe-\u003eopcode = IORING_OP_URING_CMD;\n-\tsqe-\u003eioprio = 0;\n-\n-\tif (data-\u003eflags \u0026 CTRL_CMD_HAS_BUF) {\n-\t\tcmd-\u003eaddr = data-\u003eaddr;\n-\t\tcmd-\u003elen = data-\u003elen;\n-\t}\n-\n-\tif (data-\u003eflags \u0026 CTRL_CMD_HAS_DATA)\n-\t\tcmd-\u003edata[0] = data-\u003edata[0];\n-\n-\tcmd-\u003edev_id = info-\u003edev_id;\n-\tcmd-\u003equeue_id = -1;\n-\n-\tublk_set_sqe_cmd_op(sqe, data-\u003ecmd_op);\n-\n-\tio_uring_sqe_set_data(sqe, cmd);\n-}\n-\n-static int __ublk_ctrl_cmd(struct ublk_dev *dev,\n-\t\tstruct ublk_ctrl_cmd_data *data)\n-{\n-\tstruct io_uring_sqe *sqe;\n-\tstruct io_uring_cqe *cqe;\n-\tint ret = -EINVAL;\n-\n-\tsqe = io_uring_get_sqe(\u0026dev-\u003ering);\n-\tif (!sqe) {\n-\t\tublk_err(\"%s: can't get sqe ret %d\\n\", __func__, ret);\n-\t\treturn ret;\n-\t}\n-\n-\tublk_ctrl_init_cmd(dev, sqe, data);\n-\n-\tret = io_uring_submit(\u0026dev-\u003ering);\n-\tif (ret \u003c 0) {\n-\t\tublk_err(\"uring submit ret %d\\n\", ret);\n-\t\treturn ret;\n-\t}\n-\n-\tret = io_uring_wait_cqe(\u0026dev-\u003ering, \u0026cqe);\n-\tif (ret \u003c 0) {\n-\t\tublk_err(\"wait cqe: %s\\n\", strerror(-ret));\n-\t\treturn ret;\n-\t}\n-\tio_uring_cqe_seen(\u0026dev-\u003ering, cqe);\n-\n-\treturn cqe-\u003eres;\n-}\n-\n-static int ublk_ctrl_stop_dev(struct ublk_dev *dev)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_STOP_DEV,\n-\t};\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_try_stop_dev(struct ublk_dev *dev)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_TRY_STOP_DEV,\n-\t};\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_start_dev(struct ublk_dev *dev,\n-\t\tint daemon_pid)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_START_DEV,\n-\t\t.flags\t= CTRL_CMD_HAS_DATA,\n-\t};\n-\n-\tdev-\u003edev_info.ublksrv_pid = data.data[0] = daemon_pid;\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_start_user_recovery(struct ublk_dev *dev)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_START_USER_RECOVERY,\n-\t};\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_END_USER_RECOVERY,\n-\t\t.flags\t= CTRL_CMD_HAS_DATA,\n-\t};\n-\n-\tdev-\u003edev_info.ublksrv_pid = data.data[0] = daemon_pid;\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_add_dev(struct ublk_dev *dev)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_ADD_DEV,\n-\t\t.flags\t= CTRL_CMD_HAS_BUF,\n-\t\t.addr = (__u64) (uintptr_t) \u0026dev-\u003edev_info,\n-\t\t.len = sizeof(struct ublksrv_ctrl_dev_info),\n-\t};\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_del_dev(struct ublk_dev *dev)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op = UBLK_U_CMD_DEL_DEV,\n-\t\t.flags = 0,\n-\t};\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_get_info(struct ublk_dev *dev)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_GET_DEV_INFO,\n-\t\t.flags\t= CTRL_CMD_HAS_BUF,\n-\t\t.addr = (__u64) (uintptr_t) \u0026dev-\u003edev_info,\n-\t\t.len = sizeof(struct ublksrv_ctrl_dev_info),\n-\t};\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_set_params(struct ublk_dev *dev,\n-\t\tstruct ublk_params *params)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_SET_PARAMS,\n-\t\t.flags\t= CTRL_CMD_HAS_BUF,\n-\t\t.addr = (__u64) (uintptr_t) params,\n-\t\t.len = sizeof(*params),\n-\t};\n-\tparams-\u003elen = sizeof(*params);\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_get_params(struct ublk_dev *dev,\n-\t\tstruct ublk_params *params)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_GET_PARAMS,\n-\t\t.flags\t= CTRL_CMD_HAS_BUF,\n-\t\t.addr = (__u64)params,\n-\t\t.len = sizeof(*params),\n-\t};\n-\n-\tparams-\u003elen = sizeof(*params);\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_get_features(struct ublk_dev *dev,\n-\t\t__u64 *features)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_GET_FEATURES,\n-\t\t.flags\t= CTRL_CMD_HAS_BUF,\n-\t\t.addr = (__u64) (uintptr_t) features,\n-\t\t.len = sizeof(*features),\n-\t};\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n static int parse_param_types(const char *arg, __u32 *types)\n {\n \tchar buf[128], *save = NULL, *tok;\n@@ -283,30 +86,6 @@ static void ublk_init_params_from_ctx(const struct dev_ctx *ctx,\n \t};\n }\n \n-static int ublk_ctrl_update_size(struct ublk_dev *dev,\n-\t\t__u64 nr_sects)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_UPDATE_SIZE,\n-\t\t.flags\t= CTRL_CMD_HAS_DATA,\n-\t};\n-\n-\tdata.data[0] = nr_sects;\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n-static int ublk_ctrl_quiesce_dev(struct ublk_dev *dev,\n-\t\t\t\t unsigned int timeout_ms)\n-{\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op\t= UBLK_U_CMD_QUIESCE_DEV,\n-\t\t.flags\t= CTRL_CMD_HAS_DATA,\n-\t};\n-\n-\tdata.data[0] = timeout_ms;\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n static const char *ublk_dev_state_desc(struct ublk_dev *dev)\n {\n \tswitch (dev-\u003edev_info.state) {\n@@ -426,38 +205,6 @@ static void ublk_ctrl_dump(struct ublk_dev *dev)\n \tfflush(stdout);\n }\n \n-static void ublk_ctrl_deinit(struct ublk_dev *dev)\n-{\n-\tclose(dev-\u003ectrl_fd);\n-\tfree(dev);\n-}\n-\n-static struct ublk_dev *ublk_ctrl_init(void)\n-{\n-\tstruct ublk_dev *dev = (struct ublk_dev *)calloc(1, sizeof(*dev));\n-\tstruct ublksrv_ctrl_dev_info *info = \u0026dev-\u003edev_info;\n-\tint ret;\n-\n-\tdev-\u003ectrl_fd = open(CTRL_DEV, O_RDWR);\n-\tif (dev-\u003ectrl_fd \u003c 0) {\n-\t\tfree(dev);\n-\t\treturn NULL;\n-\t}\n-\n-\tinfo-\u003emax_io_buf_bytes = UBLK_IO_MAX_BYTES;\n-\n-\tret = ublk_setup_ring(\u0026dev-\u003ering, UBLK_CTRL_RING_DEPTH,\n-\t\t\tUBLK_CTRL_RING_DEPTH, IORING_SETUP_SQE128);\n-\tif (ret \u003c 0) {\n-\t\tublk_err(\"queue_init: %s\\n\", strerror(-ret));\n-\t\tfree(dev);\n-\t\treturn NULL;\n-\t}\n-\tdev-\u003enr_fds = 1;\n-\n-\treturn dev;\n-}\n-\n static size_t __ublk_queue_cmd_buf_sz(const struct ublk_queue *q, __u16 depth)\n {\n \tsize_t size = depth * (size_t)q-\u003eio_desc_size;\n@@ -1283,24 +1030,6 @@ static void ublk_shmem_unregister_all(void)\n \tshmem_count = 0;\n }\n \n-static int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,\n-\t\t\t     __u32 flags)\n-{\n-\tstruct ublk_shmem_buf_reg buf_reg = {\n-\t\t.addr = (unsigned long)addr,\n-\t\t.len = size,\n-\t\t.flags = flags,\n-\t};\n-\tstruct ublk_ctrl_cmd_data data = {\n-\t\t.cmd_op = UBLK_U_CMD_REG_BUF,\n-\t\t.flags = CTRL_CMD_HAS_BUF,\n-\t\t.addr = (unsigned long)\u0026buf_reg,\n-\t\t.len = sizeof(buf_reg),\n-\t};\n-\n-\treturn __ublk_ctrl_cmd(dev, \u0026data);\n-}\n-\n /*\n  * Handle one client connection: receive memfd, mmap it, register\n  * the VA range with kernel, send back the assigned index.\ndiff --git a/tools/testing/selftests/ublk/kublk.h b/tools/testing/selftests/ublk/kublk.h\nindex d98f3d612d888..99b8ceff853cb 100644\n--- a/tools/testing/selftests/ublk/kublk.h\n+++ b/tools/testing/selftests/ublk/kublk.h\n@@ -294,6 +294,38 @@ struct ublk_dev {\n \n extern int ublk_queue_io_cmd(struct ublk_thread *t, struct ublk_io *io);\n \n+static inline int ublk_setup_ring(struct io_uring *r, int depth,\n+\t\tint cq_depth, unsigned int flags)\n+{\n+\tstruct io_uring_params p;\n+\n+\tmemset(\u0026p, 0, sizeof(p));\n+\tp.flags = flags | IORING_SETUP_CQSIZE;\n+\tp.cq_entries = cq_depth;\n+\n+\treturn io_uring_queue_init_params(depth, r, \u0026p);\n+}\n+\n+/* ctrl.c: control commands */\n+struct ublk_dev *ublk_ctrl_init(void);\n+void ublk_ctrl_deinit(struct ublk_dev *dev);\n+int __ublk_ctrl_cmd(struct ublk_dev *dev, struct ublk_ctrl_cmd_data *data);\n+int ublk_ctrl_add_dev(struct ublk_dev *dev);\n+int ublk_ctrl_del_dev(struct ublk_dev *dev);\n+int ublk_ctrl_get_info(struct ublk_dev *dev);\n+int ublk_ctrl_set_params(struct ublk_dev *dev, struct ublk_params *params);\n+int ublk_ctrl_get_params(struct ublk_dev *dev, struct ublk_params *params);\n+int ublk_ctrl_get_features(struct ublk_dev *dev, __u64 *features);\n+int ublk_ctrl_start_dev(struct ublk_dev *dev, int daemon_pid);\n+int ublk_ctrl_stop_dev(struct ublk_dev *dev);\n+int ublk_ctrl_try_stop_dev(struct ublk_dev *dev);\n+int ublk_ctrl_start_user_recovery(struct ublk_dev *dev);\n+int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid);\n+int ublk_ctrl_update_size(struct ublk_dev *dev, __u64 nr_sects);\n+int ublk_ctrl_quiesce_dev(struct ublk_dev *dev, unsigned int timeout_ms);\n+int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,\n+\t\t      __u32 flags);\n+\n static inline int __ublk_use_batch_io(__u64 flags)\n {\n \treturn flags \u0026 UBLK_F_BATCH_IO;\ndiff --git a/tools/testing/selftests/ublk/test_generic_18.sh b/tools/testing/selftests/ublk/test_generic_18.sh\nnew file mode 100755\nindex 0000000000000..223944dba149a\n--- /dev/null\n+++ b/tools/testing/selftests/ublk/test_generic_18.sh\n@@ -0,0 +1,37 @@\n+#!/bin/bash\n+# SPDX-License-Identifier: GPL-2.0\n+\n+. \"$(cd \"$(dirname \"$0\")\" \u0026\u0026 pwd)\"/test_common.sh\n+\n+ERR_CODE=0\n+CANCEL_PROG=\"$(_ublk_test_top_dir)/ublk_cancel_ready\"\n+\n+_prep_test \"generic\" \"start device over canceled io commands\"\n+\n+# the modes are described in ublk_cancel_ready.c\n+for mode in stop_start partial_fetch recovery stop_restart stop_attached \\\n+\t\tstop_live_restart race_start race_fetch race_async_fetch; do\n+\tdmesg_before=$(dmesg | wc -l)\n+\ttimeout 60 \"$CANCEL_PROG\" \"$mode\" \u003e \"$UBLK_TMP\" 2\u003e\u00261\n+\tres=$?\n+\tmsg=\"\"\n+\n+\tif dmesg | tail -n +\"$((dmesg_before + 1))\" | \\\n+\t\t\tgrep -q -e \"BUG:\" -e \"Oops\" -e \"WARNING:\"; then\n+\t\tmsg=\"$mode: kernel oops/warning\"\n+\t\tERR_CODE=255\n+\telif [ \"$res\" -eq \"$UBLK_SKIP_CODE\" ]; then\n+\t\t[ \"$ERR_CODE\" -eq 0 ] \u0026\u0026 ERR_CODE=$UBLK_SKIP_CODE\n+\telif [ \"$res\" -ne 0 ]; then\n+\t\tmsg=\"$mode: failed ($res)\"\n+\t\tERR_CODE=255\n+\tfi\n+\t# the output once: on failure, or always when not quiet\n+\t[ -n \"$msg\" ] \u0026\u0026 echo \"$msg\"\n+\tif [ -n \"$msg\" ] || [ \"$UBLK_TEST_QUIET\" -eq 0 ]; then\n+\t\tcat \"$UBLK_TMP\"\n+\tfi\n+done\n+\n+_cleanup_test\n+_show_result $TID $ERR_CODE\ndiff --git a/tools/testing/selftests/ublk/ublk_cancel_ready.c b/tools/testing/selftests/ublk/ublk_cancel_ready.c\nnew file mode 100644\nindex 0000000000000..0ed0d01fa74ad\n--- /dev/null\n+++ b/tools/testing/selftests/ublk/ublk_cancel_ready.c\n@@ -0,0 +1,984 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/*\n+ * Bring a ublk device live while some of its fetched io commands are\n+ * canceled.\n+ *\n+ * A cancel completes a fetched command and clears io-\u003ecmd, but the io\n+ * still counts as ready. Only ubq-\u003ecanceling keeps ublk_queue_rq() away\n+ * from the NULL io-\u003ecmd. Each mode below loses -\u003ecanceling in another way:\n+ *\n+ * stop_start:    fetch every tag, STOP_DEV before START_DEV, START_DEV.\n+ *                STOP_DEV cancels the commands of the attached server,\n+ *                so START_DEV must get -EBUSY until that server is gone.\n+ *                An unfixed kernel starts the device over them.\n+ *\n+ * partial_fetch: task A fetches tags 0..depth-2 and dies, so its\n+ *                commands are canceled. Task B fetches the last tag, and\n+ *                an unfixed kernel clears -\u003ecanceling when the queue gets\n+ *                ready. /dev/ublkcN stays open all the time, so\n+ *                ublk_ch_release() never resets the queue.\n+ *\n+ * recovery:      a UBLK_F_USER_RECOVERY device with two queues loses its\n+ *                server. During recovery task Q0 fetches queue 0, which\n+ *                clears q0-\u003ecanceling, then dies. An unfixed kernel still\n+ *                has ub-\u003ecanceling set because queue 1 is not ready, so\n+ *                ublk_start_cancel() does not mark queue 0 again. Reads\n+ *                are issued on a CPU mapped to queue 0.\n+ *\n+ * In partial_fetch and recovery, START_DEV / END_USER_RECOVERY may refuse\n+ * with -ENODEV, or bring the device live with its queue still canceling:\n+ * then every read has to complete, with -EIO. An unfixed kernel oopses\n+ * in ublk_queue_cmd() on a NULL io-\u003ecmd.\n+ *\n+ * stop_restart:  control mode. STOP_DEV on a new device, before any\n+ *                server opened it, takes no command. A server started\n+ *                afterwards has to start the device and serve I/O.\n+ *\n+ * stop_attached: STOP_DEV on a device whose server opened it but fetched\n+ *                nothing stops that server too: FETCH gets ABORT and\n+ *                START_DEV -EBUSY, until a new server opens the device.\n+ *\n+ * stop_live_restart: STOP_DEV on a live device; once its server is gone,\n+ *                a new server has to fetch and start it again. An unfixed\n+ *                ublk_ch_release() skips the reset once the disk is gone,\n+ *                so the new FETCH gets -EBUSY.\n+ *\n+ * race_start:    STOP_DEV and START_DEV at the same time on a device whose\n+ *                server fetched every tag. START_DEV either wins, and\n+ *                reads complete (they fail once STOP_DEV removes the\n+ *                disk), or gets -EBUSY. An unfixed kernel cancels the\n+ *                commands of the live disk.\n+ *\n+ * race_fetch:    STOP_DEV while a server opens the device and fetches,\n+ *                then START_DEV. If STOP_DEV came before the open, it must\n+ *                not take any command, so START_DEV works and every read\n+ *                succeeds; otherwise START_DEV gets -EBUSY. An unfixed\n+ *                kernel takes commands fetched after its unlock and goes\n+ *                live over them.\n+ *\n+ * race_async_fetch: FETCH with IOSQE_ASYNC while STOP_DEV cancels,\n+ *                then close the ring. An unfixed FETCH marks its command\n+ *                cancelable only after publishing it and dropping\n+ *                ub-\u003emutex; a cancel in between completes a command which\n+ *                is then put on io_uring's cancelable list. Closing the\n+ *                ring walks that list: KASAN reports a use after free.\n+ *                Needs a KASAN kernel to see the bug; otherwise it must\n+ *                just not crash.\n+ *\n+ * A dying task is a child process which fetches and then calls exec().\n+ * exec() cancels the task's uring_cmds before it returns, so the cancel\n+ * is done once the child is reaped. A thread exit does not cancel them,\n+ * and closing the ring cancels them later, from io_ring_exit_work().\n+ */\n+#include \u003csched.h\u003e\n+\n+#include \"kublk.h\"\n+#include \"../kselftest.h\"\n+\n+#define NR_QUEUES\t2\n+#define DEPTH\t\t4\n+#define BUF_SIZE\t(64 \u003c\u003c 10)\n+#define DEV_SECTORS\t(64 \u003c\u003c 11)\t/* 64MB */\n+#define NR_READS\t(DEPTH * 2)\n+#define SERVE_DELAY_US\t200000\n+#define RACE_LOOPS\t50\n+#define ASYNC_LOOPS\t300\n+\n+/* one control handle per thread, the race modes send commands from two */\n+static __thread struct ublk_dev *ctrl_dev;\n+static int cdev_fd = -1;\n+static int dev_id = -1;\n+static int nr_queues;\n+static void *bufs[NR_QUEUES][DEPTH];\n+static const struct ublksrv_io_desc *iods[NR_QUEUES];\n+static size_t iods_len;\n+\n+static struct io_uring_sqe *get_sqe(struct io_uring *ring)\n+{\n+\tstruct io_uring_sqe *sqe = io_uring_get_sqe(ring);\n+\n+\tif (!sqe) {\n+\t\tfprintf(stderr, \"out of sqes\\n\");\n+\t\texit(KSFT_FAIL);\n+\t}\n+\treturn sqe;\n+}\n+\n+static struct ublk_dev *ctrl(void)\n+{\n+\tif (!ctrl_dev) {\n+\t\tctrl_dev = ublk_ctrl_init();\n+\t\tif (!ctrl_dev) {\n+\t\t\tfprintf(stderr, \"ublk_ctrl_init failed\\n\");\n+\t\t\texit(KSFT_FAIL);\n+\t\t}\n+\t}\n+\tctrl_dev-\u003edev_info.dev_id = dev_id;\n+\treturn ctrl_dev;\n+}\n+\n+static void ctrl_put(void)\n+{\n+\tif (ctrl_dev)\n+\t\tublk_ctrl_deinit(ctrl_dev);\n+\tctrl_dev = NULL;\n+}\n+\n+static int dev_state(void)\n+{\n+\tstruct ublk_dev *dev = ctrl();\n+\tint ret = ublk_ctrl_get_info(dev);\n+\n+\treturn ret ? ret : dev-\u003edev_info.state;\n+}\n+\n+static int open_cdev(void)\n+{\n+\tsize_t max_len = UBLK_MAX_QUEUE_DEPTH * sizeof(struct ublksrv_io_desc);\n+\tint pg = getpagesize();\n+\tchar path[64];\n+\n+\tsnprintf(path, sizeof(path), \"/dev/ublkc%d\", dev_id);\n+\tfor (int i = 0; i \u003c 100 \u0026\u0026 cdev_fd \u003c 0; i++) {\n+\t\tcdev_fd = open(path, O_RDWR);\n+\t\tif (cdev_fd \u003c 0)\n+\t\t\tusleep(50000);\n+\t}\n+\tif (cdev_fd \u003c 0)\n+\t\treturn -errno;\n+\n+\t/* queue q's descriptors start at q * the size for the max depth */\n+\tmax_len = (max_len + pg - 1) \u0026 ~(size_t)(pg - 1);\n+\tiods_len = (DEPTH * sizeof(struct ublksrv_io_desc) + pg - 1) \u0026\n+\t\t~(size_t)(pg - 1);\n+\tfor (int q = 0; q \u003c nr_queues; q++) {\n+\t\tvoid *p = mmap(NULL, iods_len, PROT_READ,\n+\t\t\t       MAP_SHARED | MAP_POPULATE, cdev_fd,\n+\t\t\t       UBLKSRV_CMD_BUF_OFFSET + q * max_len);\n+\n+\t\tif (p == MAP_FAILED)\n+\t\t\treturn -errno;\n+\t\tiods[q] = p;\n+\t}\n+\treturn 0;\n+}\n+\n+/* the last reference to /dev/ublkcN runs ublk_ch_release() */\n+static void close_cdev(void)\n+{\n+\tfor (int q = 0; q \u003c nr_queues; q++) {\n+\t\tif (iods[q])\n+\t\t\tmunmap((void *)iods[q], iods_len);\n+\t\tiods[q] = NULL;\n+\t}\n+\tif (cdev_fd \u003e= 0)\n+\t\tclose(cdev_fd);\n+\tcdev_fd = -1;\n+}\n+\n+/* ADD_DEV and SET_PARAMS, without opening /dev/ublkcN */\n+static int add_dev_noopen(int queues, __u64 flags)\n+{\n+\tstruct ublk_dev *dev = ctrl();\n+\tstruct ublksrv_ctrl_dev_info info = {\n+\t\t.nr_hw_queues\t= queues,\n+\t\t.queue_depth\t= DEPTH,\n+\t\t.max_io_buf_bytes = BUF_SIZE,\n+\t\t.dev_id\t\t= -1,\n+\t\t.flags\t\t= UBLK_F_NO_AUTO_PART_SCAN | flags,\n+\t};\n+\tstruct ublk_params p = {\n+\t\t.types\t= UBLK_PARAM_TYPE_BASIC,\n+\t\t.basic\t= {\n+\t\t\t.logical_bs_shift\t= 9,\n+\t\t\t.physical_bs_shift\t= 12,\n+\t\t\t.io_opt_shift\t\t= 12,\n+\t\t\t.io_min_shift\t\t= 9,\n+\t\t\t.max_sectors\t\t= BUF_SIZE \u003e\u003e 9,\n+\t\t\t.dev_sectors\t\t= DEV_SECTORS,\n+\t\t},\n+\t};\n+\tint ret;\n+\n+\tnr_queues = queues;\n+\tdev-\u003edev_info = info;\n+\tret = ublk_ctrl_add_dev(dev);\n+\tif (ret)\n+\t\treturn ret;\n+\tdev_id = dev-\u003edev_info.dev_id;\n+\n+\tret = ublk_ctrl_set_params(ctrl(), \u0026p);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\tfor (int q = 0; q \u003c queues; q++)\n+\t\tfor (int i = 0; i \u003c DEPTH; i++)\n+\t\t\tif (!bufs[q][i] \u0026\u0026\n+\t\t\t    posix_memalign(\u0026bufs[q][i], getpagesize(), BUF_SIZE))\n+\t\t\t\treturn -ENOMEM;\n+\treturn 0;\n+}\n+\n+static int add_dev(int queues, __u64 flags)\n+{\n+\treturn add_dev_noopen(queues, flags) ?: open_cdev();\n+}\n+\n+static void cleanup(void)\n+{\n+\tif (dev_id \u003c 0)\n+\t\treturn;\n+\tublk_ctrl_stop_dev(ctrl());\n+\tclose_cdev();\n+\tublk_ctrl_del_dev(ctrl());\n+\tdev_id = -1;\n+}\n+\n+/* race_async_fetch sets IOSQE_ASYNC on the io commands */\n+static int io_cmd_sqe_flags;\n+\n+static void queue_io_cmd(struct io_uring *ring, __u32 op, int q, int tag,\n+\t\t\t int res)\n+{\n+\tstruct io_uring_sqe *sqe = get_sqe(ring);\n+\tstruct ublksrv_io_cmd *cmd = (struct ublksrv_io_cmd *)sqe-\u003ecmd;\n+\n+\tmemset(sqe, 0, sizeof(*sqe));\n+\tsqe-\u003efd = cdev_fd;\n+\tsqe-\u003eopcode = IORING_OP_URING_CMD;\n+\tsqe-\u003eflags = io_cmd_sqe_flags;\n+\tublk_set_sqe_cmd_op(sqe, op);\n+\tcmd-\u003eq_id = q;\n+\tcmd-\u003etag = tag;\n+\tcmd-\u003eresult = res;\n+\tcmd-\u003eaddr = (__u64)(uintptr_t)bufs[q][tag];\n+\tio_uring_sqe_set_data64(sqe, (q \u003c\u003c 16) | tag);\n+}\n+\n+struct async_arg {\n+\tpthread_barrier_t go, stopped;\n+};\n+\n+/*\n+ * FETCH every tag with IOSQE_ASYNC, racing STOP_DEV, then close the ring once\n+ * STOP_DEV returned: that walks io_uring's list of cancelable commands.\n+ */\n+static void *race_async_fn(void *data)\n+{\n+\tstruct async_arg *a = data;\n+\tstruct io_uring ring;\n+\tint ok = !io_uring_queue_init(DEPTH, \u0026ring, 0);\n+\n+\tif (ok)\n+\t\tfor (int tag = 0; tag \u003c DEPTH; tag++)\n+\t\t\tqueue_io_cmd(\u0026ring, UBLK_U_IO_FETCH_REQ, 0, tag, 0);\n+\tpthread_barrier_wait(\u0026a-\u003ego);\n+\tif (ok)\n+\t\tio_uring_submit(\u0026ring);\n+\tpthread_barrier_wait(\u0026a-\u003estopped);\n+\tif (ok)\n+\t\tio_uring_queue_exit(\u0026ring);\n+\treturn NULL;\n+}\n+\n+/* reap @nr completions of canceled fetch commands, return how many */\n+static int reap_aborts(struct io_uring *ring, int nr)\n+{\n+\tstruct __kernel_timespec ts = { .tv_sec = 2 };\n+\tstruct io_uring_cqe *cqe;\n+\tint aborted = 0;\n+\n+\twhile (nr--) {\n+\t\tif (io_uring_wait_cqe_timeout(ring, \u0026cqe, \u0026ts))\n+\t\t\tbreak;\n+\t\tif (cqe-\u003eres == UBLK_IO_RES_ABORT)\n+\t\t\taborted++;\n+\t\tio_uring_cqe_seen(ring, cqe);\n+\t}\n+\treturn aborted;\n+}\n+\n+/*\n+ * A server thread fetches @nr_tags tags of queue @q, then completes each\n+ * request after @delay_us, until its commands are aborted.\n+ */\n+struct server {\n+\tint q, first_tag, nr_tags, delay_us;\n+\tint failed;\t/* result which ended the serve loop, 0 if none */\n+\tpthread_t thread;\n+\tpthread_barrier_t fetched;\n+\t/* race_fetch: wait for @go and @race_delay_us, open, fetch one by one */\n+\tpthread_barrier_t go;\n+\tint race, race_delay_us;\n+};\n+\n+static void *server_fn(void *data)\n+{\n+\tstruct server *s = data;\n+\tstruct io_uring ring;\n+\tstruct io_uring_cqe *cqe;\n+\n+\tif (io_uring_queue_init(DEPTH, \u0026ring, 0)) {\n+\t\tpthread_barrier_wait(\u0026s-\u003efetched);\n+\t\treturn NULL;\n+\t}\n+\tif (s-\u003erace) {\n+\t\tpthread_barrier_wait(\u0026s-\u003ego);\n+\t\tusleep(s-\u003erace_delay_us);\n+\t\tif (open_cdev()) {\n+\t\t\tpthread_barrier_wait(\u0026s-\u003efetched);\n+\t\t\tio_uring_queue_exit(\u0026ring);\n+\t\t\treturn NULL;\n+\t\t}\n+\t}\n+\tfor (int tag = s-\u003efirst_tag; tag \u003c s-\u003efirst_tag + s-\u003enr_tags; tag++) {\n+\t\tqueue_io_cmd(\u0026ring, UBLK_U_IO_FETCH_REQ, s-\u003eq, tag, 0);\n+\t\tif (s-\u003erace)\n+\t\t\tio_uring_submit(\u0026ring);\n+\t}\n+\tio_uring_submit(\u0026ring);\n+\tpthread_barrier_wait(\u0026s-\u003efetched);\n+\n+\twhile (!io_uring_wait_cqe(\u0026ring, \u0026cqe)) {\n+\t\tint tag = cqe-\u003euser_data \u0026 0xffff;\n+\t\tconst struct ublksrv_io_desc *iod = \u0026iods[s-\u003eq][tag];\n+\t\tint res = cqe-\u003eres;\n+\n+\t\tio_uring_cqe_seen(\u0026ring, cqe);\n+\t\tif (res != UBLK_IO_RES_OK) {\n+\t\t\t__atomic_store_n(\u0026s-\u003efailed, res, __ATOMIC_RELEASE);\n+\t\t\tbreak;\n+\t\t}\n+\t\tusleep(s-\u003edelay_us);\n+\t\tres = ublksrv_get_op(iod) \u003c= UBLK_IO_OP_WRITE ?\n+\t\t\tiod-\u003enr_sectors \u003c\u003c 9 : 0;\n+\t\tqueue_io_cmd(\u0026ring, UBLK_U_IO_COMMIT_AND_FETCH_REQ, s-\u003eq, tag,\n+\t\t\t     res);\n+\t\tio_uring_submit(\u0026ring);\n+\t}\n+\tio_uring_queue_exit(\u0026ring);\n+\treturn NULL;\n+}\n+\n+/* returns once the fetch commands are issued; with @race, call race_go() */\n+static void server_start(struct server *s)\n+{\n+\tpthread_barrier_init(\u0026s-\u003efetched, NULL, 2);\n+\tpthread_barrier_init(\u0026s-\u003ego, NULL, 2);\n+\tpthread_create(\u0026s-\u003ethread, NULL, server_fn, s);\n+\tif (!s-\u003erace)\n+\t\tpthread_barrier_wait(\u0026s-\u003efetched);\n+}\n+\n+/* wait up to 5s for the serve loop of @s to end, return its result */\n+static int server_wait_failed(struct server *s)\n+{\n+\tint res = 0;\n+\n+\tfor (int i = 0; i \u003c 500 \u0026\u0026 !res; i++) {\n+\t\tres = __atomic_load_n(\u0026s-\u003efailed, __ATOMIC_ACQUIRE);\n+\t\tif (!res)\n+\t\t\tusleep(10000);\n+\t}\n+\treturn res;\n+}\n+\n+/*\n+ * A task fetches @nr_tags tags of queue @q and dies. Returns once its\n+ * commands are canceled, see the top of this file.\n+ */\n+static int fetch_and_die(int q, int first_tag, int nr_tags)\n+{\n+\tint status;\n+\tpid_t pid = fork();\n+\n+\tif (pid \u003c 0)\n+\t\treturn -1;\n+\tif (!pid) {\n+\t\tstruct io_uring ring;\n+\n+\t\tif (io_uring_queue_init(DEPTH, \u0026ring, 0))\n+\t\t\t_exit(1);\n+\t\tfor (int tag = first_tag; tag \u003c first_tag + nr_tags; tag++)\n+\t\t\tqueue_io_cmd(\u0026ring, UBLK_U_IO_FETCH_REQ, q, tag, 0);\n+\t\tif (io_uring_submit(\u0026ring) != nr_tags)\n+\t\t\t_exit(1);\n+\t\texeclp(\"true\", \"true\", NULL);\n+\t\t_exit(1);\n+\t}\n+\tif (waitpid(pid, \u0026status, 0) != pid || !WIFEXITED(status) ||\n+\t    WEXITSTATUS(status))\n+\t\treturn -1;\n+\tprintf(\"q%d: task died with %d fetch cmds in flight\\n\", q, nr_tags);\n+\treturn 0;\n+}\n+\n+struct reads {\n+\tstruct io_uring ring;\n+\tvoid *buf;\n+\tint fd, done, ok, eio, other;\n+};\n+\n+/*\n+ * Issue NR_READS reads at once, from @cpu if it is not negative: blk-mq\n+ * maps the submitting CPU to the hw queue. A server holding its live\n+ * tags for a while makes the other reads take the canceled tags.\n+ */\n+static int open_tries = 100;\t/* 50ms each */\n+\n+static int reads_submit(struct reads *r, int cpu)\n+{\n+\tchar path[64];\n+\tcpu_set_t set, old;\n+\n+\tsnprintf(path, sizeof(path), \"/dev/ublkb%d\", dev_id);\n+\tr-\u003efd = -1;\n+\tfor (int i = 0; i \u003c open_tries \u0026\u0026 r-\u003efd \u003c 0; i++) {\n+\t\tr-\u003efd = open(path, O_RDONLY | O_DIRECT);\n+\t\tif (r-\u003efd \u003c 0)\n+\t\t\tusleep(50000);\n+\t}\n+\tif (r-\u003efd \u003c 0) {\n+\t\tfprintf(stderr, \"open %s: %s\\n\", path, strerror(errno));\n+\t\treturn -1;\n+\t}\n+\tif (posix_memalign(\u0026r-\u003ebuf, 4096, NR_READS * 4096) ||\n+\t    io_uring_queue_init(NR_READS, \u0026r-\u003ering, 0))\n+\t\treturn -1;\n+\n+\tfor (int i = 0; i \u003c NR_READS; i++)\n+\t\tio_uring_prep_read(get_sqe(\u0026r-\u003ering), r-\u003efd,\n+\t\t\t\t   r-\u003ebuf + i * 4096, 4096, i * 4096);\n+\n+\tif (cpu \u003e= 0) {\n+\t\tsched_getaffinity(0, sizeof(old), \u0026old);\n+\t\tCPU_ZERO(\u0026set);\n+\t\tCPU_SET(cpu, \u0026set);\n+\t\tsched_setaffinity(0, sizeof(set), \u0026set);\n+\t}\n+\tio_uring_submit(\u0026r-\u003ering);\n+\tif (cpu \u003e= 0)\n+\t\tsched_setaffinity(0, sizeof(old), \u0026old);\n+\treturn 0;\n+}\n+\n+static void reads_reap(struct reads *r, int timeout_s)\n+{\n+\tstruct __kernel_timespec ts = { .tv_sec = timeout_s };\n+\tstruct io_uring_cqe *cqe;\n+\n+\twhile (r-\u003edone \u003c NR_READS \u0026\u0026\n+\t       !io_uring_wait_cqe_timeout(\u0026r-\u003ering, \u0026cqe, \u0026ts)) {\n+\t\tif (cqe-\u003eres == 4096)\n+\t\t\tr-\u003eok++;\n+\t\telse if (cqe-\u003eres == -EIO)\n+\t\t\tr-\u003eeio++;\n+\t\telse\n+\t\t\tr-\u003eother++;\n+\t\tr-\u003edone++;\n+\t\tio_uring_cqe_seen(\u0026r-\u003ering, cqe);\n+\t}\n+}\n+\n+static void reads_put(struct reads *r)\n+{\n+\tio_uring_queue_exit(\u0026r-\u003ering);\n+\tclose(r-\u003efd);\n+\tfree(r-\u003ebuf);\n+}\n+\n+static int reads_result(struct reads *r)\n+{\n+\tprintf(\"reads: %d ok, %d -EIO, %d other, %d not completed\\n\",\n+\t       r-\u003eok, r-\u003eeio, r-\u003eother, NR_READS - r-\u003edone);\n+\treads_put(r);\n+\treturn r-\u003eother || r-\u003edone \u003c NR_READS ? KSFT_FAIL : KSFT_PASS;\n+}\n+\n+static int start_dev(void)\n+{\n+\tint ret = ublk_ctrl_start_dev(ctrl(), getpid());\n+\n+\tprintf(\"START_DEV: %d\\n\", ret);\n+\treturn ret;\n+}\n+\n+static int expect_err(const char *what, int ret, int want)\n+{\n+\tif (ret == want)\n+\t\treturn KSFT_PASS;\n+\tfprintf(stderr, \"%s: %d, expected %d\\n\", what, ret, want);\n+\treturn KSFT_FAIL;\n+}\n+\n+/* START_DEV must fail with @want; if it went live, show what a read does */\n+static int start_dev_expect(int want)\n+{\n+\tstruct reads r = {};\n+\tint ret = start_dev();\n+\n+\tif (!ret \u0026\u0026 !reads_submit(\u0026r, -1)) {\n+\t\treads_reap(\u0026r, 10);\n+\t\treads_result(\u0026r);\n+\t}\n+\treturn expect_err(\"START_DEV\", ret, want);\n+}\n+\n+/* -ENODEV, or live over a canceling queue: then reads must complete */\n+static int expect_enodev_or_reads(const char *what, int ret, struct reads *r)\n+{\n+\tif (ret == -ENODEV)\n+\t\treturn KSFT_PASS;\n+\tif (ret) {\n+\t\tfprintf(stderr, \"%s: %d, expected 0 or %d\\n\", what, ret,\n+\t\t\t-ENODEV);\n+\t\treturn KSFT_FAIL;\n+\t}\n+\treads_reap(r, 10);\n+\treturn reads_result(r);\n+}\n+\n+static int test_stop_start(void)\n+{\n+\tstruct io_uring ring;\n+\tint ret;\n+\n+\tif (add_dev(1, 0))\n+\t\treturn KSFT_FAIL;\n+\tif (io_uring_queue_init(DEPTH, \u0026ring, 0))\n+\t\treturn KSFT_FAIL;\n+\tfor (int tag = 0; tag \u003c DEPTH; tag++)\n+\t\tqueue_io_cmd(\u0026ring, UBLK_U_IO_FETCH_REQ, 0, tag, 0);\n+\tio_uring_submit(\u0026ring);\n+\n+\t/* device is ready but not started: state is UBLK_S_DEV_DEAD */\n+\tret = ublk_ctrl_stop_dev(ctrl());\n+\tprintf(\"STOP_DEV: %d, canceled fetch cmds: %d/%d\\n\", ret,\n+\t       reap_aborts(\u0026ring, DEPTH), DEPTH);\n+\n+\t/* STOP_DEV canceled the attached server: no start until it exits */\n+\tret = start_dev_expect(-EBUSY);\n+\tio_uring_queue_exit(\u0026ring);\n+\treturn ret;\n+}\n+\n+static int test_partial_fetch(void)\n+{\n+\tstruct server b = { .q = 0, .first_tag = DEPTH - 1, .nr_tags = 1,\n+\t\t\t    .delay_us = SERVE_DELAY_US };\n+\tstruct reads r = {};\n+\tint ret;\n+\n+\tif (add_dev(1, 0) || fetch_and_die(0, 0, DEPTH - 1))\n+\t\treturn KSFT_FAIL;\n+\n+\t/* the last FETCH makes the queue ready and clears -\u003ecanceling */\n+\tserver_start(\u0026b);\n+\n+\tret = start_dev();\n+\tif (!ret \u0026\u0026 reads_submit(\u0026r, -1))\n+\t\tret = KSFT_FAIL;\n+\telse\n+\t\tret = expect_enodev_or_reads(\"START_DEV\", ret, \u0026r);\n+\n+\tcleanup();\n+\tpthread_join(b.thread, NULL);\n+\treturn ret;\n+}\n+\n+static int test_stop_restart(void)\n+{\n+\tstruct server s = { .q = 0, .nr_tags = DEPTH };\n+\tstruct reads r = {};\n+\tint ret;\n+\n+\tif (add_dev_noopen(1, 0))\n+\t\treturn KSFT_FAIL;\n+\n+\t/* no server is attached, so there is nothing to cancel */\n+\tret = ublk_ctrl_stop_dev(ctrl());\n+\tprintf(\"STOP_DEV: %d\\n\", ret);\n+\tif (open_cdev())\n+\t\treturn KSFT_FAIL;\n+\n+\tserver_start(\u0026s);\n+\tif (start_dev()) {\n+\t\tret = KSFT_FAIL;\n+\t} else if (reads_submit(\u0026r, -1)) {\n+\t\tret = KSFT_FAIL;\n+\t} else {\n+\t\treads_reap(\u0026r, 10);\n+\t\tret = reads_result(\u0026r);\n+\t\tif (r.ok != NR_READS)\n+\t\t\tret = KSFT_FAIL;\n+\t}\n+\n+\tcleanup();\n+\tpthread_join(s.thread, NULL);\n+\treturn ret;\n+}\n+\n+struct start_arg {\n+\tpthread_barrier_t go;\n+\tint ret, reads_ok;\n+};\n+\n+static void *race_start_fn(void *data)\n+{\n+\tstruct start_arg *a = data;\n+\tstruct reads r = {};\n+\n+\tpthread_barrier_wait(\u0026a-\u003ego);\n+\ta-\u003eret = ublk_ctrl_start_dev(ctrl(), getpid());\n+\ta-\u003ereads_ok = 1;\n+\t/*\n+\t * The disk may be gone already if STOP_DEV came right after, and\n+\t * reads may fail then: they only have to complete.\n+\t */\n+\tif (!a-\u003eret \u0026\u0026 !reads_submit(\u0026r, -1)) {\n+\t\treads_reap(\u0026r, 5);\n+\t\ta-\u003ereads_ok = r.done == NR_READS;\n+\t\tif (a-\u003ereads_ok)\n+\t\t\treads_put(\u0026r);\n+\t\telse\n+\t\t\treads_result(\u0026r);\n+\t}\n+\tctrl_put();\n+\treturn NULL;\n+}\n+\n+static int test_race_start(void)\n+{\n+\tint live = 0, ebusy = 0;\n+\n+\topen_tries = 4;\n+\tfor (int i = 0; i \u003c RACE_LOOPS; i++) {\n+\t\tstruct server s = { .q = 0, .nr_tags = DEPTH };\n+\t\tstruct start_arg a = {};\n+\t\tpthread_t t;\n+\n+\t\tif (add_dev(1, 0))\n+\t\t\treturn KSFT_FAIL;\n+\t\tserver_start(\u0026s);\n+\t\tpthread_barrier_init(\u0026a.go, NULL, 2);\n+\t\tpthread_create(\u0026t, NULL, race_start_fn, \u0026a);\n+\t\tpthread_barrier_wait(\u0026a.go);\n+\t\tublk_ctrl_stop_dev(ctrl());\n+\t\tpthread_join(t, NULL);\n+\t\tcleanup();\n+\t\tpthread_join(s.thread, NULL);\n+\n+\t\tif (a.ret == 0)\n+\t\t\tlive++;\n+\t\telse if (a.ret == -EBUSY)\n+\t\t\tebusy++;\n+\t\tif ((a.ret \u0026\u0026 a.ret != -EBUSY) || !a.reads_ok) {\n+\t\t\tfprintf(stderr, \"loop %d: START_DEV %d, reads %s\\n\",\n+\t\t\t\ti, a.ret, a.reads_ok ? \"ok\" : \"failed\");\n+\t\t\treturn KSFT_FAIL;\n+\t\t}\n+\t}\n+\tprintf(\"%d loops: START_DEV won %d, got -EBUSY %d\\n\", RACE_LOOPS,\n+\t       live, ebusy);\n+\treturn KSFT_PASS;\n+}\n+\n+static int test_race_fetch(void)\n+{\n+\tint live = 0, ebusy = 0;\n+\n+\tfor (int i = 0; i \u003c RACE_LOOPS; i++) {\n+\t\t/* vary who goes first: STOP_DEV, or the server's open */\n+\t\tstruct server s = { .q = 0, .nr_tags = DEPTH, .race = 1,\n+\t\t\t\t    .race_delay_us = (i % 10) * 50 };\n+\t\tstruct reads r = {};\n+\t\tint ret;\n+\n+\t\tif (add_dev_noopen(1, 0))\n+\t\t\treturn KSFT_FAIL;\n+\t\tserver_start(\u0026s);\n+\t\tpthread_barrier_wait(\u0026s.go);\n+\t\tublk_ctrl_stop_dev(ctrl());\n+\t\tpthread_barrier_wait(\u0026s.fetched);\n+\n+\t\tret = ublk_ctrl_start_dev(ctrl(), getpid());\n+\t\tif (!ret) {\n+\t\t\t/* nothing was taken, so every read has to succeed */\n+\t\t\tlive++;\n+\t\t\tif (reads_submit(\u0026r, -1))\n+\t\t\t\treturn KSFT_FAIL;\n+\t\t\treads_reap(\u0026r, 5);\n+\t\t\tif (r.ok != NR_READS) {\n+\t\t\t\tfprintf(stderr, \"loop %d: live, but \", i);\n+\t\t\t\treads_result(\u0026r);\n+\t\t\t\treturn KSFT_FAIL;\n+\t\t\t}\n+\t\t\treads_put(\u0026r);\n+\t\t} else if (ret == -EBUSY) {\n+\t\t\tebusy++;\n+\t\t} else {\n+\t\t\tfprintf(stderr, \"loop %d: START_DEV %d\\n\", i, ret);\n+\t\t\treturn KSFT_FAIL;\n+\t\t}\n+\t\tcleanup();\n+\t\tpthread_join(s.thread, NULL);\n+\t}\n+\tprintf(\"%d loops: START_DEV worked %d, got -EBUSY %d\\n\", RACE_LOOPS,\n+\t       live, ebusy);\n+\treturn KSFT_PASS;\n+}\n+\n+static int test_race_async_fetch(void)\n+{\n+\tio_cmd_sqe_flags = IOSQE_ASYNC;\n+\tfor (int i = 0; i \u003c ASYNC_LOOPS; i++) {\n+\t\tstruct async_arg a;\n+\t\tpthread_t t;\n+\n+\t\tif (add_dev(1, 0))\n+\t\t\treturn KSFT_FAIL;\n+\t\tpthread_barrier_init(\u0026a.go, NULL, 2);\n+\t\tpthread_barrier_init(\u0026a.stopped, NULL, 2);\n+\t\tpthread_create(\u0026t, NULL, race_async_fn, \u0026a);\n+\t\tpthread_barrier_wait(\u0026a.go);\n+\t\t/* vary where STOP_DEV lands relative to the FETCHes */\n+\t\tusleep((i % 5) * 1000);\n+\t\tublk_ctrl_stop_dev(ctrl());\n+\t\tpthread_barrier_wait(\u0026a.stopped);\n+\t\tpthread_join(t, NULL);\n+\t\tcleanup();\n+\t}\n+\tprintf(\"%d loops done\\n\", ASYNC_LOOPS);\n+\treturn KSFT_PASS;\n+}\n+\n+/*\n+ * The stopped server is gone: reopen /dev/ublkcN, start a new server and\n+ * the device, and read from it.\n+ */\n+static int restart_server(void)\n+{\n+\tstruct server s = { .q = 0, .nr_tags = DEPTH };\n+\tstruct reads r = {};\n+\tint ret;\n+\n+\tclose_cdev();\n+\tif (open_cdev())\n+\t\treturn KSFT_FAIL;\n+\tserver_start(\u0026s);\n+\tusleep(200000);\n+\tif (s.failed) {\n+\t\tfprintf(stderr, \"new server: FETCH failed %d\\n\", s.failed);\n+\t\tcleanup();\n+\t\tpthread_join(s.thread, NULL);\n+\t\treturn KSFT_FAIL;\n+\t}\n+\t/* -EEXIST until the old disk is freed, e.g. after a udev probe */\n+\tfor (int i = 0; i \u003c 100; i++) {\n+\t\tret = start_dev();\n+\t\tif (ret != -EEXIST)\n+\t\t\tbreak;\n+\t\tusleep(50000);\n+\t}\n+\tif (ret || reads_submit(\u0026r, -1)) {\n+\t\tfprintf(stderr, \"new server: START_DEV %d\\n\", ret);\n+\t\tret = KSFT_FAIL;\n+\t} else {\n+\t\treads_reap(\u0026r, 10);\n+\t\tret = reads_result(\u0026r);\n+\t\tif (r.ok != NR_READS)\n+\t\t\tret = KSFT_FAIL;\n+\t}\n+\tcleanup();\n+\tpthread_join(s.thread, NULL);\n+\treturn ret;\n+}\n+\n+/*\n+ * stop_attached: STOP_DEV while a server has the device open but has\n+ * fetched nothing stops that server too: its FETCH gets ABORT and\n+ * START_DEV -EBUSY. Once it is gone, a new server can start the device.\n+ */\n+static int test_stop_attached(void)\n+{\n+\tstruct server s = { .q = 0, .nr_tags = DEPTH };\n+\tint ret;\n+\n+\tif (add_dev(1, 0))\n+\t\treturn KSFT_FAIL;\n+\tret = ublk_ctrl_stop_dev(ctrl());\n+\tprintf(\"STOP_DEV: %d\\n\", ret);\n+\n+\tserver_start(\u0026s);\n+\tret = server_wait_failed(\u0026s);\n+\tprintf(\"FETCH after STOP_DEV: %d\\n\", ret);\n+\tif (ret != UBLK_IO_RES_ABORT) {\n+\t\tcleanup();\n+\t\tpthread_join(s.thread, NULL);\n+\t\treturn KSFT_FAIL;\n+\t}\n+\tpthread_join(s.thread, NULL);\n+\tif (expect_err(\"START_DEV\", start_dev(), -EBUSY))\n+\t\treturn KSFT_FAIL;\n+\treturn restart_server();\n+}\n+\n+/*\n+ * stop_live_restart: STOP_DEV on a live device; once its server is gone,\n+ * a new server has to fetch and start it again.\n+ */\n+static int test_stop_live_restart(void)\n+{\n+\tstruct server s = { .q = 0, .nr_tags = DEPTH };\n+\tint ret;\n+\n+\tif (add_dev(1, 0))\n+\t\treturn KSFT_FAIL;\n+\tserver_start(\u0026s);\n+\tif (expect_err(\"START_DEV\", start_dev(), 0))\n+\t\treturn KSFT_FAIL;\n+\tret = ublk_ctrl_stop_dev(ctrl());\n+\tprintf(\"STOP_DEV: %d\\n\", ret);\n+\tpthread_join(s.thread, NULL);\n+\treturn restart_server();\n+}\n+\n+/* first CPU which blk-mq maps to hw queue @q */\n+static int queue_cpu(int q)\n+{\n+\tchar path[96];\n+\tFILE *f;\n+\tint cpu = -1;\n+\n+\tsnprintf(path, sizeof(path), \"/sys/block/ublkb%d/mq/%d/cpu_list\",\n+\t\t dev_id, q);\n+\tfor (int i = 0; i \u003c 100 \u0026\u0026 !(f = fopen(path, \"r\")); i++)\n+\t\tusleep(50000);\n+\tif (!f)\n+\t\treturn -1;\n+\tif (fscanf(f, \"%d\", \u0026cpu) != 1)\n+\t\tcpu = -1;\n+\tfclose(f);\n+\treturn cpu;\n+}\n+\n+static int test_recovery(void)\n+{\n+\tstruct server q1 = { .q = 1, .nr_tags = DEPTH };\n+\tstruct io_uring ring;\n+\tstruct reads r = {};\n+\tint ret, cpu;\n+\n+\tif (add_dev(NR_QUEUES, UBLK_F_USER_RECOVERY))\n+\t\treturn KSFT_FAIL;\n+\n+\t/* the first server: fetch everything, start, then die */\n+\tif (io_uring_queue_init(NR_QUEUES * DEPTH, \u0026ring, 0))\n+\t\treturn KSFT_FAIL;\n+\tfor (int q = 0; q \u003c NR_QUEUES; q++)\n+\t\tfor (int tag = 0; tag \u003c DEPTH; tag++)\n+\t\t\tqueue_io_cmd(\u0026ring, UBLK_U_IO_FETCH_REQ, q, tag, 0);\n+\tio_uring_submit(\u0026ring);\n+\tif (start_dev())\n+\t\treturn KSFT_FAIL;\n+\tcpu = queue_cpu(0);\n+\tif (cpu \u003c 0) {\n+\t\tprintf(\"no CPU maps to queue 0, can't aim the reads\\n\");\n+\t\treturn KSFT_SKIP;\n+\t}\n+\n+\tio_uring_queue_exit(\u0026ring);\n+\tclose_cdev();\n+\tfor (int i = 0; i \u003c 100 \u0026\u0026 dev_state() != UBLK_S_DEV_QUIESCED; i++)\n+\t\tusleep(50000);\n+\tprintf(\"server exited, state %d (QUIESCED is %d)\\n\", dev_state(),\n+\t       UBLK_S_DEV_QUIESCED);\n+\n+\tfor (int i = 0; i \u003c 100; i++) {\n+\t\tret = ublk_ctrl_start_user_recovery(ctrl());\n+\t\tif (ret != -EBUSY)\n+\t\t\tbreak;\n+\t\tusleep(50000);\n+\t}\n+\tprintf(\"START_USER_RECOVERY: %d\\n\", ret);\n+\tif (ret || open_cdev())\n+\t\treturn KSFT_FAIL;\n+\n+\t/* queue 0 gets ready, then its task dies before queue 1 is ready */\n+\tif (fetch_and_die(0, 0, DEPTH))\n+\t\treturn KSFT_FAIL;\n+\n+\tprintf(\"reads on cpu %d, which maps to queue 0\\n\", cpu);\n+\tif (reads_submit(\u0026r, cpu))\n+\t\treturn KSFT_FAIL;\n+\n+\tserver_start(\u0026q1);\n+\tret = ublk_ctrl_end_user_recovery(ctrl(), getpid());\n+\tprintf(\"END_USER_RECOVERY: %d\\n\", ret);\n+\tif (ret \u0026\u0026 ret != -ENODEV) {\n+\t\tfprintf(stderr, \"END_USER_RECOVERY: %d, expected 0 or %d\\n\",\n+\t\t\tret, -ENODEV);\n+\t\tret = KSFT_FAIL;\n+\t} else {\n+\t\tret = KSFT_PASS;\n+\t}\n+\n+\t/*\n+\t * After -ENODEV, requests held back on queue 0 wait for STOP_DEV to\n+\t * fail them. Close the disk before DEL_DEV, which waits for its last\n+\t * reference.\n+\t */\n+\treads_reap(\u0026r, 1);\n+\tublk_ctrl_stop_dev(ctrl());\n+\treads_reap(\u0026r, 5);\n+\tif (reads_result(\u0026r))\n+\t\tret = KSFT_FAIL;\n+\tcleanup();\n+\tpthread_join(q1.thread, NULL);\n+\treturn ret;\n+}\n+\n+static const struct {\n+\tconst char *name;\n+\tint (*fn)(void);\n+} modes[] = {\n+\t{ \"stop_start\",\t\ttest_stop_start },\n+\t{ \"partial_fetch\",\ttest_partial_fetch },\n+\t{ \"recovery\",\t\ttest_recovery },\n+\t{ \"stop_restart\",\ttest_stop_restart },\n+\t{ \"race_start\",\t\ttest_race_start },\n+\t{ \"race_fetch\",\t\ttest_race_fetch },\n+\t{ \"race_async_fetch\",\ttest_race_async_fetch },\n+\t{ \"stop_attached\",\ttest_stop_attached },\n+\t{ \"stop_live_restart\",\ttest_stop_live_restart },\n+};\n+\n+int main(int argc, char **argv)\n+{\n+\tint (*fn)(void) = NULL;\n+\tint ret;\n+\n+\tfor (int i = 0; argc == 2 \u0026\u0026 i \u003c ARRAY_SIZE(modes); i++)\n+\t\tif (!strcmp(argv[1], modes[i].name))\n+\t\t\tfn = modes[i].fn;\n+\tif (!fn) {\n+\t\tfprintf(stderr, \"usage: %s MODE, modes:\", argv[0]);\n+\t\tfor (int i = 0; i \u003c ARRAY_SIZE(modes); i++)\n+\t\t\tfprintf(stderr, \" %s\", modes[i].name);\n+\t\tfprintf(stderr, \"\\n\");\n+\t\treturn KSFT_FAIL;\n+\t}\n+\n+\t/* keep the output of a run that ends in an oops */\n+\tsetvbuf(stdout, NULL, _IOLBF, 0);\n+\n+\tif (access(CTRL_DEV, F_OK)) {\n+\t\tperror(CTRL_DEV);\n+\t\treturn KSFT_SKIP;\n+\t}\n+\tprintf(\"%s\\n\", argv[1]);\n+\tret = fn();\n+\tcleanup();\n+\tctrl_put();\n+\treturn ret;\n+}\n"
}

2/1 2026/10/02 08:23 agent patch-evaluator
0m Model:
core

Args:
null

Results:
{
  "EnableConfigs": [
    "BLK_DEV_UBLK"
  ],
  "FocusSymbols": [
    "ublk_stop_dev",
    "ublk_ctrl_start_dev",
    "ublk_ch_uring_cmd_local",
    "ublk_uring_cmd_cancel_fn",
    "ublk_ch_release_work_fn",
    "ublk_batch_attach",
    "ublk_reset_ch_dev",
    "ublk_ch_open"
  ],
  "Reasoning": "The patch modifies drivers/block/ublk_drv.c to fix races and state management between STOP_DEV, START_DEV, FETCH_REQ, and device release/cancellation in the ublk driver. It introduces a new device state (UB_STATE_STOPPING), adjusts file reference counting and memory ordering, alters cancelation routines, and modifies assertion checks. The modified code is directly reachable from userspace via io_uring commands on /dev/ublk-control and /dev/ublkc*, making it highly relevant for fuzzing concurrency and lifecycle edge cases.",
  "WorthFuzzing": true
}

Instruction:
You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.

IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.

================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.

- REACHABILITY IS THE PRIMARY GATE:
  Fuzzing can only discover bugs in code that can actually execute in standard virtualized
  environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
  If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
  regardless of whether it adds assertions or complex logic.

- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
  Patch authors routinely label changes as "cleanups", "refactorings", or state
  "No functional change intended". Do NOT take these claims at face value.
  Code refactorings that rearrange logic, introduce helper functions, or alter state management
  in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
  If reachable executable code is modified or refactored, it MUST be fuzzed.

- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
  When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
  BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
  Even if the author believes the invariant always holds, fuzzing is essential to verify whether
  an unusual sequence of operations can violate it.

================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:

- Non-kernel and non-executable changes:
  * Modifications to Documentation/, comments, or spelling fixes.
  * User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
    that do not affect the compiled kernel image (vmlinux) or kernel modules.
  * Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
    that do not alter control flow or data structures.
  * Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
  * Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
    ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
  * Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
    physical PCI hot-unplug or manual sysfs driver unbinding.
  * Code paths exclusive to architectures other than the target architecture.

================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
  * Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
    BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
  * Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
  * Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
  * Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).

================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:

- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
  Do NOT list generic, ubiquitous functions called by almost every program in the corpus
  (including, but not limited to: general memory allocators and deallocators, page fault
  and trap handlers, or core synchronization primitives; this is not an exhaustive list).
  Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
  which severely dilutes fuzzing effort away from the actual changes.

- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
  List functions that specifically implement the logic being added or altered, or direct API entrypoints
  for the subsystem feature under review.

- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
  Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
  distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
  If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
  functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).

================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.


Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.

Prompt:
Target architecture: amd64

For your convenience, here is the diff of the changes:
commit 30c2ad39819c25e745543e6ae191d461a2c8251e
Author: syz-cluster <triage@syzkaller.com>
Date:   Fri Oct 2 08:23:10 2026 +0000

    syz-cluster: applied patch under review

diff --git a/drivers/block/ublk_drv.c b/drivers/block/ublk_drv.c
index 66eb55e7162e5..f57d544c1da23 100644
--- a/drivers/block/ublk_drv.c
+++ b/drivers/block/ublk_drv.c
@@ -321,6 +321,8 @@ struct ublk_device {
 #define UB_STATE_OPEN		0
 #define UB_STATE_USED		1
 #define UB_STATE_DELETED	2
+/* STOP_DEV canceled the server's commands, until its release */
+#define UB_STATE_STOPPING	3
 	unsigned long		state;
 	int			ub_number;
 
@@ -334,6 +336,13 @@ struct ublk_device {
 	u16			nr_queue_ready;
 	bool 			unprivileged_daemons;
 	struct mutex cancel_mutex;
+	/* the open /dev/ublkcN, protected by cancel_mutex */
+	struct file *ch_file;
+	/*
+	 * A cancel started in this FETCH round. Set by ublk_set_canceling(),
+	 * cleared only by ublk_reset_ch_dev() when a new round starts. While
+	 * it is set, no queue clears its ->canceling.
+	 */
 	bool canceling;
 	pid_t 	ublksrv_tgid;
 	struct delayed_work	exit_work;
@@ -810,6 +819,8 @@ ublk_batch_alloc_fcmd(struct io_uring_cmd *cmd)
 	if (fcmd) {
 		fcmd->cmd = cmd;
 		fcmd->buf_group = READ_ONCE(cmd->sqe->buf_index);
+		/* a cancel may look at it before it is linked */
+		INIT_LIST_HEAD(&fcmd->node);
 	}
 	return fcmd;
 }
@@ -2399,6 +2410,9 @@ static int ublk_ch_open(struct inode *inode, struct file *filp)
 		return -EBUSY;
 	filp->private_data = ub;
 	ub->ublksrv_tgid = current->tgid;
+	mutex_lock(&ub->cancel_mutex);
+	ub->ch_file = filp;
+	mutex_unlock(&ub->cancel_mutex);
 	return 0;
 }
 
@@ -2415,11 +2429,17 @@ static void ublk_reset_ch_dev(struct ublk_device *ub)
 		spin_unlock(&ubq->cancel_lock);
 	}
 
+	/* a new FETCH round starts, the queues stay canceling until ready */
+	mutex_lock(&ub->cancel_mutex);
+	ub->canceling = false;
+	mutex_unlock(&ub->cancel_mutex);
+
 	/* set to NULL, otherwise new tasks cannot mmap io_cmd_buf */
 	ub->mm = NULL;
 	ub->nr_queue_ready = 0;
 	ub->unprivileged_daemons = false;
 	ub->ublksrv_tgid = -1;
+	clear_bit(UB_STATE_STOPPING, &ub->state);
 }
 
 static struct gendisk *ublk_get_disk(struct ublk_device *ub)
@@ -2539,12 +2559,15 @@ static void ublk_ch_release_work_fn(struct work_struct *work)
 	}
 
 	/*
-	 * disk isn't attached yet, either device isn't live, or it has
-	 * been removed already, so we needn't to do anything
+	 * No disk: the device isn't live, or it has been removed already.
+	 * There are no requests to abort, but the round still has to be
+	 * reset, so that a new server can fetch and start the device.
 	 */
 	disk = ublk_get_disk(ub);
-	if (!disk)
-		goto out;
+	if (!disk) {
+		mutex_lock(&ub->mutex);
+		goto reset;
+	}
 
 	/*
 	 * All uring_cmd are done now, so abort any request outstanding to
@@ -2576,7 +2599,7 @@ static void ublk_ch_release_work_fn(struct work_struct *work)
 
 	/* double check after grabbing lock */
 	if (!ub->ub_disk)
-		goto unlock;
+		goto reset;
 
 	/*
 	 * Transition the device to the nosrv state. What exactly this
@@ -2602,13 +2625,14 @@ static void ublk_ch_release_work_fn(struct work_struct *work)
 				WRITE_ONCE(ublk_get_queue(ub, i)->fail_io, true);
 		}
 	}
-unlock:
+reset:
+	/*
+	 * All uring_cmd has been done now, reset device & ubq. Under
+	 * ub->mutex, so START_DEV sees the round either ready or reset.
+	 */
+	ublk_reset_ch_dev(ub);
 	mutex_unlock(&ub->mutex);
 	ublk_put_disk(disk);
-
-	/* all uring_cmd has been done now, reset device & ubq */
-	ublk_reset_ch_dev(ub);
-out:
 	clear_bit(UB_STATE_OPEN, &ub->state);
 
 	/* put the reference grabbed in ublk_ch_release() */
@@ -2619,6 +2643,9 @@ static int ublk_ch_release(struct inode *inode, struct file *filp)
 {
 	struct ublk_device *ub = filp->private_data;
 
+	mutex_lock(&ub->cancel_mutex);
+	ub->ch_file = NULL;
+	mutex_unlock(&ub->cancel_mutex);
 	/*
 	 * Grab ublk device reference, so it won't be gone until we are
 	 * really released from work function.
@@ -2790,7 +2817,8 @@ static void ublk_cancel_cmd(struct ublk_queue *ubq, u16 tag,
 	done = !!(io->flags & UBLK_IO_FLAG_CANCELED);
 	if (!done) {
 		io->flags |= UBLK_IO_FLAG_CANCELED;
-		cmd = io->cmd;
+		/* dependency ordered against smp_wmb() in ublk_prep_cancel() */
+		cmd = READ_ONCE(io->cmd);
 		io->cmd = NULL;
 	}
 	spin_unlock(&ubq->cancel_lock);
@@ -2879,6 +2907,7 @@ static void ublk_uring_cmd_cancel_fn(struct io_uring_cmd *cmd,
 {
 	struct ublk_uring_cmd_pdu *pdu = ublk_get_uring_cmd_pdu(cmd);
 	struct ublk_queue *ubq = pdu->ubq;
+	struct io_uring_cmd *cur;
 	struct task_struct *task;
 	struct ublk_io *io;
 
@@ -2895,7 +2924,9 @@ static void ublk_uring_cmd_cancel_fn(struct io_uring_cmd *cmd,
 
 	ublk_start_cancel(ubq->dev);
 
-	WARN_ON_ONCE(io->cmd != cmd);
+	/* NULL if STOP_DEV's cancel took it meanwhile */
+	cur = READ_ONCE(io->cmd);
+	WARN_ON_ONCE(cur && cur != cmd);
 	ublk_cancel_cmd(ubq, pdu->tag, issue_flags);
 }
 
@@ -3006,13 +3037,54 @@ static void ublk_stop_dev_unlocked(struct ublk_device *ub)
 	put_disk(disk);
 }
 
+static struct file *ublk_get_ch_file(struct ublk_device *ub)
+{
+	struct file *file;
+
+	mutex_lock(&ub->cancel_mutex);
+	file = ub->ch_file;
+	if (file && !file_ref_get(&file->f_ref))
+		file = NULL;
+	mutex_unlock(&ub->cancel_mutex);
+	return file;
+}
+
 static void ublk_stop_dev(struct ublk_device *ub)
 {
+	struct file *file;
+
+	/*
+	 * FETCH, PREP and START_DEV take ub->mutex. If a server has
+	 * /dev/ublkcN open, set STOPPING, which turns it away, and hold a
+	 * reference on the file: the server's release, whose reset clears
+	 * STOPPING and lets a new server attach, can't run before we drop it.
+	 * So the cancel can run after the unlock and only meets this server's
+	 * commands; it has to: io_uring_cmd_done() may take uring_lock, under
+	 * which FETCH takes ub->mutex.
+	 */
 	mutex_lock(&ub->mutex);
 	ublk_stop_dev_unlocked(ub);
-	mutex_unlock(&ub->mutex);
 	cancel_work_sync(&ub->partition_scan_work);
+	file = ublk_get_ch_file(ub);
+	if (file) {
+		/*
+		 * The server has to close /dev/ublkcN before this device can
+		 * be started again: the reset in its release clears STOPPING.
+		 */
+		set_bit(UB_STATE_STOPPING, &ub->state);
+		/* for wake_up_var() below, see wake_up_bit() */
+		smp_mb__after_atomic();
+	}
+	mutex_unlock(&ub->mutex);
+
+	/* no server: nothing to cancel */
+	if (!file)
+		return;
+
+	/* wake a START_DEV waiting for the device to get ready */
+	wake_up_var(&ub->nr_queue_ready);
 	ublk_cancel_dev(ub);
+	__fput_sync(file);
 }
 
 static void ublk_reset_io_flags(struct ublk_queue *ubq, struct ublk_io *io)
@@ -3024,11 +3096,19 @@ static void ublk_reset_io_flags(struct ublk_queue *ubq, struct ublk_io *io)
 }
 
 /* reset per-queue io flags */
-static void ublk_queue_reset_io_flags(struct ublk_queue *ubq)
+static void ublk_queue_reset_io_flags(struct ublk_device *ub,
+				      struct ublk_queue *ubq)
 {
-	spin_lock(&ubq->cancel_lock);
-	ubq->canceling = false;
-	spin_unlock(&ubq->cancel_lock);
+	/*
+	 * A cancel in this FETCH round took a command which still counts as
+	 * ready, so the queue has to stay canceling. ub->canceling is set
+	 * under cancel_mutex before any command is taken: either we see it
+	 * here, or the cancel marks this queue again later.
+	 */
+	mutex_lock(&ub->cancel_mutex);
+	if (!ub->canceling)
+		ubq->canceling = false;
+	mutex_unlock(&ub->cancel_mutex);
 	ubq->fail_io = false;
 	ubq->force_abort = false;
 }
@@ -3051,24 +3131,17 @@ static void ublk_mark_io_ready(struct ublk_device *ub, u16 q_id,
 		ub->nr_queue_ready++;
 
 		/*
-		 * Reset queue flags as soon as this queue is ready.
-		 * This clears the canceling flag, allowing batch FETCH commands
-		 * to succeed during recovery without waiting for all queues.
+		 * Reset queue flags as soon as this queue is ready. Unless
+		 * this round saw a cancel, this clears the canceling flag,
+		 * allowing batch FETCH commands to succeed during recovery
+		 * without waiting for all queues.
 		 */
-		ublk_queue_reset_io_flags(ubq);
+		ublk_queue_reset_io_flags(ub, ubq);
 	}
 
-	/* Check if all queues are ready */
-	if (ublk_dev_ready(ub)) {
-		/*
-		 * All queues ready - clear device-level canceling flag
-		 * and wake ublk_dev_ready() waiters.
-		 */
-		mutex_lock(&ub->cancel_mutex);
-		ub->canceling = false;
-		mutex_unlock(&ub->cancel_mutex);
+	/* All queues ready - wake ublk_dev_ready() waiters */
+	if (ublk_dev_ready(ub))
 		wake_up_var(&ub->nr_queue_ready);
-	}
 }
 
 static inline int ublk_check_cmd_op(u32 cmd_op)
@@ -3152,6 +3225,12 @@ ublk_fill_io_cmd(struct ublk_io *io, struct io_uring_cmd *cmd)
 	return req;
 }
 
+/*
+ * Call before ublk_fill_io_cmd() publishes @cmd in io->cmd: a control-path
+ * cancel may complete any command found there, and io_uring_cmd_done() only
+ * takes it off the cancelable list if it is marked already. The handlers
+ * hold uring_lock, so marking takes no lock.
+ */
 static inline void ublk_prep_cancel(struct io_uring_cmd *cmd,
 				    unsigned int issue_flags,
 				    struct ublk_queue *ubq, u16 tag)
@@ -3165,6 +3244,8 @@ static inline void ublk_prep_cancel(struct io_uring_cmd *cmd,
 	pdu->ubq = ubq;
 	pdu->tag = tag;
 	io_uring_cmd_mark_cancelable(cmd, issue_flags);
+	/* pairs with the cancel loading cmd from io->cmd, then cmd->flags */
+	smp_wmb();
 }
 
 static void ublk_io_release(void *priv)
@@ -3269,6 +3350,9 @@ static int ublk_check_fetch_buf(const struct ublk_device *ub, __u64 buf_addr)
 static int __ublk_fetch(struct io_uring_cmd *cmd, struct ublk_device *ub,
 			struct ublk_io *io, u16 q_id)
 {
+	if (test_bit(UB_STATE_STOPPING, &ub->state))
+		return UBLK_IO_RES_ABORT;
+
 	/* UBLK_IO_FETCH_REQ is only allowed before dev is setup */
 	if (ublk_dev_ready(ub))
 		return -EBUSY;
@@ -3412,11 +3496,11 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,
 		ret = ublk_check_fetch_buf(ub, addr);
 		if (ret)
 			goto out;
+		/* before ublk_fetch() publishes io->cmd, see ublk_prep_cancel() */
+		ublk_prep_cancel(cmd, issue_flags, ubq, tag);
 		ret = ublk_fetch(cmd, ub, io, addr, q_id);
 		if (ret)
-			goto out;
-
-		ublk_prep_cancel(cmd, issue_flags, ubq, tag);
+			goto out_done;
 		return -EIOCBQUEUED;
 	}
 
@@ -3460,6 +3544,7 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,
 		if (ret)
 			goto out;
 		io->res = result;
+		ublk_prep_cancel(cmd, issue_flags, ubq, tag);
 		req = ublk_fill_io_cmd(io, cmd);
 		ublk_apply_io_buf(ub, io, cmd, addr, &auto_buf, &buf_idx);
 		if (buf_idx != UBLK_INVALID_BUF_IDX)
@@ -3478,19 +3563,24 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,
 		 * uring_cmd active first and prepare for handling new requeued
 		 * request
 		 */
+		ublk_prep_cancel(cmd, issue_flags, ubq, tag);
 		req = ublk_fill_io_cmd(io, cmd);
 		io->buf.addr = addr;
 		if (likely(ublk_get_data(ubq, io, req))) {
 			__ublk_prep_compl_io_cmd(io, req);
-			return UBLK_IO_RES_OK;
+			ret = UBLK_IO_RES_OK;
+			goto out_done;
 		}
 		break;
 	default:
 		goto out;
 	}
-	ublk_prep_cancel(cmd, issue_flags, ubq, tag);
 	return -EIOCBQUEUED;
 
+ out_done:
+	/* marked cancelable: complete through io_uring_cmd_done() */
+	io_uring_cmd_done(cmd, ret, issue_flags);
+	return -EIOCBQUEUED;
  out:
 	pr_devel("%s: complete: cmd op %d, tag %d ret %x io_flags %x\n",
 			__func__, cmd_op, tag, ret, io ? io->flags : 0);
@@ -3889,6 +3979,15 @@ static int ublk_batch_attach(struct ublk_queue *ubq,
 	bool free = false;
 	struct ublk_uring_cmd_pdu *pdu = ublk_get_uring_cmd_pdu(data->cmd);
 
+	/*
+	 * Mark it cancelable before linking it into fcmd_head, where a cancel
+	 * from the control path can take and complete it: see
+	 * ublk_prep_cancel(). evts_lock orders the mark before the link.
+	 */
+	pdu->ubq = ubq;
+	pdu->fcmd = fcmd;
+	io_uring_cmd_mark_cancelable(fcmd->cmd, data->issue_flags);
+
 	spin_lock(&ubq->evts_lock);
 	if (unlikely(ubq->force_abort || ubq->canceling)) {
 		free = true;
@@ -3899,14 +3998,12 @@ static int ublk_batch_attach(struct ublk_queue *ubq,
 	spin_unlock(&ubq->evts_lock);
 
 	if (unlikely(free)) {
+		/* off the cancelable list first, then nothing can see fcmd */
+		io_uring_cmd_done(data->cmd, -ENODEV, data->issue_flags);
 		ublk_batch_free_fcmd(fcmd);
-		return -ENODEV;
+		return -EIOCBQUEUED;
 	}
 
-	pdu->ubq = ubq;
-	pdu->fcmd = fcmd;
-	io_uring_cmd_mark_cancelable(fcmd->cmd, data->issue_flags);
-
 	if (!new_fcmd)
 		goto out;
 
@@ -3914,9 +4011,12 @@ static int ublk_batch_attach(struct ublk_queue *ubq,
 	 * If the two fetch commands are originated from same io_ring_ctx,
 	 * run batch dispatch directly. Otherwise, schedule task work for
 	 * doing it.
+	 *
+	 * Use data->cmd, not fcmd->cmd: once fcmd is linked and not active,
+	 * a cancel from the control path may complete and free it.
 	 */
 	if (io_uring_cmd_ctx_handle(new_fcmd->cmd) ==
-			io_uring_cmd_ctx_handle(fcmd->cmd)) {
+			io_uring_cmd_ctx_handle(data->cmd)) {
 		data->cmd = new_fcmd->cmd;
 		ublk_batch_dispatch(ubq, data, new_fcmd);
 	} else {
@@ -4431,21 +4531,27 @@ static bool ublk_validate_user_pid(struct ublk_device *ub, pid_t ublksrv_pid)
 	return ub->ublksrv_tgid == ublksrv_pid;
 }
 
+static bool ublk_dev_ready_or_stopping(const struct ublk_device *ub)
+{
+	return ublk_dev_ready(ub) || test_bit(UB_STATE_STOPPING, &ub->state);
+}
+
 /*
- * Wait until all queues have fetched their I/O commands, and return with
- * ub->mutex held and readiness guaranteed: then every queue's ->canceling
- * is cleared. Ready may regress between wakeup and mutex_lock() (F_BATCH
- * UNPREP, daemon death), so re-check it under the mutex and wait again.
+ * Wait until all queues have fetched their I/O commands, or STOP_DEV set
+ * UB_STATE_STOPPING, and return with ub->mutex held. The queues stay
+ * canceling if this round saw a cancel, see ublk_queue_reset_io_flags().
+ * Ready may regress between wakeup and mutex_lock() (F_BATCH UNPREP,
+ * daemon death), so re-check it under the mutex and wait again.
  */
 static int ublk_wait_dev_ready_and_lock(struct ublk_device *ub)
 {
 	while (true) {
 		if (wait_var_event_interruptible(&ub->nr_queue_ready,
-						 ublk_dev_ready(ub)))
+						 ublk_dev_ready_or_stopping(ub)))
 			return -EINTR;
 
 		mutex_lock(&ub->mutex);
-		if (ublk_dev_ready(ub))
+		if (ublk_dev_ready_or_stopping(ub))
 			return 0;
 		mutex_unlock(&ub->mutex);
 	}
@@ -4545,6 +4651,10 @@ static int ublk_ctrl_start_dev(struct ublk_device *ub,
 		ret = -EEXIST;
 		goto out_unlock;
 	}
+	if (test_bit(UB_STATE_STOPPING, &ub->state)) {
+		ret = -EBUSY;
+		goto out_unlock;
+	}
 
 	disk = blk_mq_alloc_disk(&ub->tag_set, &lim, NULL);
 	if (IS_ERR(disk)) {
diff --git a/tools/testing/selftests/ublk/.gitignore b/tools/testing/selftests/ublk/.gitignore
index e17bd28f27e04..d8e93fef7fcd3 100644
--- a/tools/testing/selftests/ublk/.gitignore
+++ b/tools/testing/selftests/ublk/.gitignore
@@ -3,3 +3,4 @@
 /tools
 kublk
 metadata_size
+ublk_cancel_ready
diff --git a/tools/testing/selftests/ublk/Makefile b/tools/testing/selftests/ublk/Makefile
index 37883e9d50ec3..fc64b8f028332 100644
--- a/tools/testing/selftests/ublk/Makefile
+++ b/tools/testing/selftests/ublk/Makefile
@@ -19,6 +19,7 @@ TEST_PROGS += test_generic_12.sh
 TEST_PROGS += test_generic_13.sh
 TEST_PROGS += test_generic_16.sh
 TEST_PROGS += test_generic_17.sh
+TEST_PROGS += test_generic_18.sh
 
 TEST_PROGS += test_batch_01.sh
 TEST_PROGS += test_batch_02.sh
@@ -76,13 +77,14 @@ TEST_FILES := settings
 TEST_FILES += test_common.sh
 TEST_FILES += trace
 
-TEST_GEN_PROGS_EXTENDED = kublk metadata_size
-STANDALONE_UTILS := metadata_size.c
+TEST_GEN_PROGS_EXTENDED = kublk metadata_size ublk_cancel_ready
+STANDALONE_UTILS := metadata_size.c ublk_cancel_ready.c
 
 LOCAL_HDRS += $(wildcard *.h)
 include ../lib.mk
 
 $(OUTPUT)/kublk: $(filter-out $(STANDALONE_UTILS),$(wildcard *.c))
+$(OUTPUT)/ublk_cancel_ready: ublk_cancel_ready.c ctrl.c
 
 check:
 	shellcheck -x -f gcc *.sh
diff --git a/tools/testing/selftests/ublk/ctrl.c b/tools/testing/selftests/ublk/ctrl.c
new file mode 100644
index 0000000000000..53d54799f9097
--- /dev/null
+++ b/tools/testing/selftests/ublk/ctrl.c
@@ -0,0 +1,268 @@
+// SPDX-License-Identifier: GPL-2.0
+
+/* ublk control commands, shared by kublk and ublk_cancel_ready */
+
+#include "kublk.h"
+
+static void ublk_ctrl_init_cmd(struct ublk_dev *dev,
+		struct io_uring_sqe *sqe,
+		struct ublk_ctrl_cmd_data *data)
+{
+	struct ublksrv_ctrl_dev_info *info = &dev->dev_info;
+	struct ublksrv_ctrl_cmd *cmd = (struct ublksrv_ctrl_cmd *)ublk_get_sqe_cmd(sqe);
+
+	sqe->fd = dev->ctrl_fd;
+	sqe->opcode = IORING_OP_URING_CMD;
+	sqe->ioprio = 0;
+
+	if (data->flags & CTRL_CMD_HAS_BUF) {
+		cmd->addr = data->addr;
+		cmd->len = data->len;
+	}
+
+	if (data->flags & CTRL_CMD_HAS_DATA)
+		cmd->data[0] = data->data[0];
+
+	cmd->dev_id = info->dev_id;
+	cmd->queue_id = -1;
+
+	ublk_set_sqe_cmd_op(sqe, data->cmd_op);
+
+	io_uring_sqe_set_data(sqe, cmd);
+}
+
+int __ublk_ctrl_cmd(struct ublk_dev *dev,
+		struct ublk_ctrl_cmd_data *data)
+{
+	struct io_uring_sqe *sqe;
+	struct io_uring_cqe *cqe;
+	int ret = -EINVAL;
+
+	sqe = io_uring_get_sqe(&dev->ring);
+	if (!sqe) {
+		ublk_err("%s: can't get sqe ret %d\n", __func__, ret);
+		return ret;
+	}
+
+	ublk_ctrl_init_cmd(dev, sqe, data);
+
+	ret = io_uring_submit(&dev->ring);
+	if (ret < 0) {
+		ublk_err("uring submit ret %d\n", ret);
+		return ret;
+	}
+
+	ret = io_uring_wait_cqe(&dev->ring, &cqe);
+	if (ret < 0) {
+		ublk_err("wait cqe: %s\n", strerror(-ret));
+		return ret;
+	}
+	io_uring_cqe_seen(&dev->ring, cqe);
+
+	return cqe->res;
+}
+
+void ublk_ctrl_deinit(struct ublk_dev *dev)
+{
+	io_uring_queue_exit(&dev->ring);
+	close(dev->ctrl_fd);
+	free(dev);
+}
+
+struct ublk_dev *ublk_ctrl_init(void)
+{
+	struct ublk_dev *dev = (struct ublk_dev *)calloc(1, sizeof(*dev));
+	struct ublksrv_ctrl_dev_info *info;
+	int ret;
+
+	if (!dev)
+		return NULL;
+	info = &dev->dev_info;
+	dev->ctrl_fd = open(CTRL_DEV, O_RDWR);
+	if (dev->ctrl_fd < 0) {
+		free(dev);
+		return NULL;
+	}
+
+	info->max_io_buf_bytes = UBLK_IO_MAX_BYTES;
+
+	ret = ublk_setup_ring(&dev->ring, UBLK_CTRL_RING_DEPTH,
+			UBLK_CTRL_RING_DEPTH, IORING_SETUP_SQE128);
+	if (ret < 0) {
+		ublk_err("queue_init: %s\n", strerror(-ret));
+		close(dev->ctrl_fd);
+		free(dev);
+		return NULL;
+	}
+	dev->nr_fds = 1;
+
+	return dev;
+}
+
+int ublk_ctrl_stop_dev(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_STOP_DEV,
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_try_stop_dev(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_TRY_STOP_DEV,
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_start_dev(struct ublk_dev *dev,
+		int daemon_pid)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_START_DEV,
+		.flags	= CTRL_CMD_HAS_DATA,
+	};
+
+	dev->dev_info.ublksrv_pid = data.data[0] = daemon_pid;
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_start_user_recovery(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_START_USER_RECOVERY,
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_END_USER_RECOVERY,
+		.flags	= CTRL_CMD_HAS_DATA,
+	};
+
+	dev->dev_info.ublksrv_pid = data.data[0] = daemon_pid;
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_add_dev(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_ADD_DEV,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64) (uintptr_t) &dev->dev_info,
+		.len = sizeof(struct ublksrv_ctrl_dev_info),
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_del_dev(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op = UBLK_U_CMD_DEL_DEV,
+		.flags = 0,
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_get_info(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_GET_DEV_INFO,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64) (uintptr_t) &dev->dev_info,
+		.len = sizeof(struct ublksrv_ctrl_dev_info),
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_set_params(struct ublk_dev *dev,
+		struct ublk_params *params)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_SET_PARAMS,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64) (uintptr_t) params,
+		.len = sizeof(*params),
+	};
+	params->len = sizeof(*params);
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_get_params(struct ublk_dev *dev,
+		struct ublk_params *params)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_GET_PARAMS,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64)params,
+		.len = sizeof(*params),
+	};
+
+	params->len = sizeof(*params);
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_get_features(struct ublk_dev *dev,
+		__u64 *features)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_GET_FEATURES,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64) (uintptr_t) features,
+		.len = sizeof(*features),
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_update_size(struct ublk_dev *dev,
+		__u64 nr_sects)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_UPDATE_SIZE,
+		.flags	= CTRL_CMD_HAS_DATA,
+	};
+
+	data.data[0] = nr_sects;
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_quiesce_dev(struct ublk_dev *dev, unsigned int timeout_ms)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_QUIESCE_DEV,
+		.flags	= CTRL_CMD_HAS_DATA,
+	};
+
+	data.data[0] = timeout_ms;
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,
+		      __u32 flags)
+{
+	struct ublk_shmem_buf_reg buf_reg = {
+		.addr = (unsigned long)addr,
+		.len = size,
+		.flags = flags,
+	};
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op = UBLK_U_CMD_REG_BUF,
+		.flags = CTRL_CMD_HAS_BUF,
+		.addr = (unsigned long)&buf_reg,
+		.len = sizeof(buf_reg),
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
diff --git a/tools/testing/selftests/ublk/kublk.c b/tools/testing/selftests/ublk/kublk.c
index 2400b46157664..15ba060ce7a3c 100644
--- a/tools/testing/selftests/ublk/kublk.c
+++ b/tools/testing/selftests/ublk/kublk.c
@@ -37,203 +37,6 @@ static const struct ublk_tgt_ops *ublk_find_tgt(const char *name)
 	return NULL;
 }
 
-static inline int ublk_setup_ring(struct io_uring *r, int depth,
-		int cq_depth, unsigned flags)
-{
-	struct io_uring_params p;
-
-	memset(&p, 0, sizeof(p));
-	p.flags = flags | IORING_SETUP_CQSIZE;
-	p.cq_entries = cq_depth;
-
-	return io_uring_queue_init_params(depth, r, &p);
-}
-
-static void ublk_ctrl_init_cmd(struct ublk_dev *dev,
-		struct io_uring_sqe *sqe,
-		struct ublk_ctrl_cmd_data *data)
-{
-	struct ublksrv_ctrl_dev_info *info = &dev->dev_info;
-	struct ublksrv_ctrl_cmd *cmd = (struct ublksrv_ctrl_cmd *)ublk_get_sqe_cmd(sqe);
-
-	sqe->fd = dev->ctrl_fd;
-	sqe->opcode = IORING_OP_URING_CMD;
-	sqe->ioprio = 0;
-
-	if (data->flags & CTRL_CMD_HAS_BUF) {
-		cmd->addr = data->addr;
-		cmd->len = data->len;
-	}
-
-	if (data->flags & CTRL_CMD_HAS_DATA)
-		cmd->data[0] = data->data[0];
-
-	cmd->dev_id = info->dev_id;
-	cmd->queue_id = -1;
-
-	ublk_set_sqe_cmd_op(sqe, data->cmd_op);
-
-	io_uring_sqe_set_data(sqe, cmd);
-}
-
-static int __ublk_ctrl_cmd(struct ublk_dev *dev,
-		struct ublk_ctrl_cmd_data *data)
-{
-	struct io_uring_sqe *sqe;
-	struct io_uring_cqe *cqe;
-	int ret = -EINVAL;
-
-	sqe = io_uring_get_sqe(&dev->ring);
-	if (!sqe) {
-		ublk_err("%s: can't get sqe ret %d\n", __func__, ret);
-		return ret;
-	}
-
-	ublk_ctrl_init_cmd(dev, sqe, data);
-
-	ret = io_uring_submit(&dev->ring);
-	if (ret < 0) {
-		ublk_err("uring submit ret %d\n", ret);
-		return ret;
-	}
-
-	ret = io_uring_wait_cqe(&dev->ring, &cqe);
-	if (ret < 0) {
-		ublk_err("wait cqe: %s\n", strerror(-ret));
-		return ret;
-	}
-	io_uring_cqe_seen(&dev->ring, cqe);
-
-	return cqe->res;
-}
-
-static int ublk_ctrl_stop_dev(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_STOP_DEV,
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_try_stop_dev(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_TRY_STOP_DEV,
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_start_dev(struct ublk_dev *dev,
-		int daemon_pid)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_START_DEV,
-		.flags	= CTRL_CMD_HAS_DATA,
-	};
-
-	dev->dev_info.ublksrv_pid = data.data[0] = daemon_pid;
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_start_user_recovery(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_START_USER_RECOVERY,
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_END_USER_RECOVERY,
-		.flags	= CTRL_CMD_HAS_DATA,
-	};
-
-	dev->dev_info.ublksrv_pid = data.data[0] = daemon_pid;
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_add_dev(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_ADD_DEV,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64) (uintptr_t) &dev->dev_info,
-		.len = sizeof(struct ublksrv_ctrl_dev_info),
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_del_dev(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op = UBLK_U_CMD_DEL_DEV,
-		.flags = 0,
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_get_info(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_GET_DEV_INFO,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64) (uintptr_t) &dev->dev_info,
-		.len = sizeof(struct ublksrv_ctrl_dev_info),
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_set_params(struct ublk_dev *dev,
-		struct ublk_params *params)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_SET_PARAMS,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64) (uintptr_t) params,
-		.len = sizeof(*params),
-	};
-	params->len = sizeof(*params);
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_get_params(struct ublk_dev *dev,
-		struct ublk_params *params)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_GET_PARAMS,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64)params,
-		.len = sizeof(*params),
-	};
-
-	params->len = sizeof(*params);
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_get_features(struct ublk_dev *dev,
-		__u64 *features)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_GET_FEATURES,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64) (uintptr_t) features,
-		.len = sizeof(*features),
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
 static int parse_param_types(const char *arg, __u32 *types)
 {
 	char buf[128], *save = NULL, *tok;
@@ -283,30 +86,6 @@ static void ublk_init_params_from_ctx(const struct dev_ctx *ctx,
 	};
 }
 
-static int ublk_ctrl_update_size(struct ublk_dev *dev,
-		__u64 nr_sects)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_UPDATE_SIZE,
-		.flags	= CTRL_CMD_HAS_DATA,
-	};
-
-	data.data[0] = nr_sects;
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_quiesce_dev(struct ublk_dev *dev,
-				 unsigned int timeout_ms)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_QUIESCE_DEV,
-		.flags	= CTRL_CMD_HAS_DATA,
-	};
-
-	data.data[0] = timeout_ms;
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
 static const char *ublk_dev_state_desc(struct ublk_dev *dev)
 {
 	switch (dev->dev_info.state) {
@@ -426,38 +205,6 @@ static void ublk_ctrl_dump(struct ublk_dev *dev)
 	fflush(stdout);
 }
 
-static void ublk_ctrl_deinit(struct ublk_dev *dev)
-{
-	close(dev->ctrl_fd);
-	free(dev);
-}
-
-static struct ublk_dev *ublk_ctrl_init(void)
-{
-	struct ublk_dev *dev = (struct ublk_dev *)calloc(1, sizeof(*dev));
-	struct ublksrv_ctrl_dev_info *info = &dev->dev_info;
-	int ret;
-
-	dev->ctrl_fd = open(CTRL_DEV, O_RDWR);
-	if (dev->ctrl_fd < 0) {
-		free(dev);
-		return NULL;
-	}
-
-	info->max_io_buf_bytes = UBLK_IO_MAX_BYTES;
-
-	ret = ublk_setup_ring(&dev->ring, UBLK_CTRL_RING_DEPTH,
-			UBLK_CTRL_RING_DEPTH, IORING_SETUP_SQE128);
-	if (ret < 0) {
-		ublk_err("queue_init: %s\n", strerror(-ret));
-		free(dev);
-		return NULL;
-	}
-	dev->nr_fds = 1;
-
-	return dev;
-}
-
 static size_t __ublk_queue_cmd_buf_sz(const struct ublk_queue *q, __u16 depth)
 {
 	size_t size = depth * (size_t)q->io_desc_size;
@@ -1283,24 +1030,6 @@ static void ublk_shmem_unregister_all(void)
 	shmem_count = 0;
 }
 
-static int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,
-			     __u32 flags)
-{
-	struct ublk_shmem_buf_reg buf_reg = {
-		.addr = (unsigned long)addr,
-		.len = size,
-		.flags = flags,
-	};
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op = UBLK_U_CMD_REG_BUF,
-		.flags = CTRL_CMD_HAS_BUF,
-		.addr = (unsigned long)&buf_reg,
-		.len = sizeof(buf_reg),
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
 /*
  * Handle one client connection: receive memfd, mmap it, register
  * the VA range with kernel, send back the assigned index.
diff --git a/tools/testing/selftests/ublk/kublk.h b/tools/testing/selftests/ublk/kublk.h
index d98f3d612d888..99b8ceff853cb 100644
--- a/tools/testing/selftests/ublk/kublk.h
+++ b/tools/testing/selftests/ublk/kublk.h
@@ -294,6 +294,38 @@ struct ublk_dev {
 
 extern int ublk_queue_io_cmd(struct ublk_thread *t, struct ublk_io *io);
 
+static inline int ublk_setup_ring(struct io_uring *r, int depth,
+		int cq_depth, unsigned int flags)
+{
+	struct io_uring_params p;
+
+	memset(&p, 0, sizeof(p));
+	p.flags = flags | IORING_SETUP_CQSIZE;
+	p.cq_entries = cq_depth;
+
+	return io_uring_queue_init_params(depth, r, &p);
+}
+
+/* ctrl.c: control commands */
+struct ublk_dev *ublk_ctrl_init(void);
+void ublk_ctrl_deinit(struct ublk_dev *dev);
+int __ublk_ctrl_cmd(struct ublk_dev *dev, struct ublk_ctrl_cmd_data *data);
+int ublk_ctrl_add_dev(struct ublk_dev *dev);
+int ublk_ctrl_del_dev(struct ublk_dev *dev);
+int ublk_ctrl_get_info(struct ublk_dev *dev);
+int ublk_ctrl_set_params(struct ublk_dev *dev, struct ublk_params *params);
+int ublk_ctrl_get_params(struct ublk_dev *dev, struct ublk_params *params);
+int ublk_ctrl_get_features(struct ublk_dev *dev, __u64 *features);
+int ublk_ctrl_start_dev(struct ublk_dev *dev, int daemon_pid);
+int ublk_ctrl_stop_dev(struct ublk_dev *dev);
+int ublk_ctrl_try_stop_dev(struct ublk_dev *dev);
+int ublk_ctrl_start_user_recovery(struct ublk_dev *dev);
+int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid);
+int ublk_ctrl_update_size(struct ublk_dev *dev, __u64 nr_sects);
+int ublk_ctrl_quiesce_dev(struct ublk_dev *dev, unsigned int timeout_ms);
+int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,
+		      __u32 flags);
+
 static inline int __ublk_use_batch_io(__u64 flags)
 {
 	return flags & UBLK_F_BATCH_IO;
diff --git a/tools/testing/selftests/ublk/test_generic_18.sh b/tools/testing/selftests/ublk/test_generic_18.sh
new file mode 100755
index 0000000000000..223944dba149a
--- /dev/null
+++ b/tools/testing/selftests/ublk/test_generic_18.sh
@@ -0,0 +1,37 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+
+. "$(cd "$(dirname "$0")" && pwd)"/test_common.sh
+
+ERR_CODE=0
+CANCEL_PROG="$(_ublk_test_top_dir)/ublk_cancel_ready"
+
+_prep_test "generic" "start device over canceled io commands"
+
+# the modes are described in ublk_cancel_ready.c
+for mode in stop_start partial_fetch recovery stop_restart stop_attached \
+		stop_live_restart race_start race_fetch race_async_fetch; do
+	dmesg_before=$(dmesg | wc -l)
+	timeout 60 "$CANCEL_PROG" "$mode" > "$UBLK_TMP" 2>&1
+	res=$?
+	msg=""
+
+	if dmesg | tail -n +"$((dmesg_before + 1))" | \
+			grep -q -e "BUG:" -e "Oops" -e "WARNING:"; then
+		msg="$mode: kernel oops/warning"
+		ERR_CODE=255
+	elif [ "$res" -eq "$UBLK_SKIP_CODE" ]; then
+		[ "$ERR_CODE" -eq 0 ] && ERR_CODE=$UBLK_SKIP_CODE
+	elif [ "$res" -ne 0 ]; then
+		msg="$mode: failed ($res)"
+		ERR_CODE=255
+	fi
+	# the output once: on failure, or always when not quiet
+	[ -n "$msg" ] && echo "$msg"
+	if [ -n "$msg" ] || [ "$UBLK_TEST_QUIET" -eq 0 ]; then
+		cat "$UBLK_TMP"
+	fi
+done
+
+_cleanup_test
+_show_result $TID $ERR_CODE
diff --git a/tools/testing/selftests/ublk/ublk_cancel_ready.c b/tools/testing/selftests/ublk/ublk_cancel_ready.c
new file mode 100644
index 0000000000000..0ed0d01fa74ad
--- /dev/null
+++ b/tools/testing/selftests/ublk/ublk_cancel_ready.c
@@ -0,0 +1,984 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Bring a ublk device live while some of its fetched io commands are
+ * canceled.
+ *
+ * A cancel completes a fetched command and clears io->cmd, but the io
+ * still counts as ready. Only ubq->canceling keeps ublk_queue_rq() away
+ * from the NULL io->cmd. Each mode below loses ->canceling in another way:
+ *
+ * stop_start:    fetch every tag, STOP_DEV before START_DEV, START_DEV.
+ *                STOP_DEV cancels the commands of the attached server,
+ *                so START_DEV must get -EBUSY until that server is gone.
+ *                An unfixed kernel starts the device over them.
+ *
+ * partial_fetch: task A fetches tags 0..depth-2 and dies, so its
+ *                commands are canceled. Task B fetches the last tag, and
+ *                an unfixed kernel clears ->canceling when the queue gets
+ *                ready. /dev/ublkcN stays open all the time, so
+ *                ublk_ch_release() never resets the queue.
+ *
+ * recovery:      a UBLK_F_USER_RECOVERY device with two queues loses its
+ *                server. During recovery task Q0 fetches queue 0, which
+ *                clears q0->canceling, then dies. An unfixed kernel still
+ *                has ub->canceling set because queue 1 is not ready, so
+ *                ublk_start_cancel() does not mark queue 0 again. Reads
+ *                are issued on a CPU mapped to queue 0.
+ *
+ * In partial_fetch and recovery, START_DEV / END_USER_RECOVERY may refuse
+ * with -ENODEV, or bring the device live with its queue still canceling:
+ * then every read has to complete, with -EIO. An unfixed kernel oopses
+ * in ublk_queue_cmd() on a NULL io->cmd.
+ *
+ * stop_restart:  control mode. STOP_DEV on a new device, before any
+ *                server opened it, takes no command. A server started
+ *                afterwards has to start the device and serve I/O.
+ *
+ * stop_attached: STOP_DEV on a device whose server opened it but fetched
+ *                nothing stops that server too: FETCH gets ABORT and
+ *                START_DEV -EBUSY, until a new server opens the device.
+ *
+ * stop_live_restart: STOP_DEV on a live device; once its server is gone,
+ *                a new server has to fetch and start it again. An unfixed
+ *                ublk_ch_release() skips the reset once the disk is gone,
+ *                so the new FETCH gets -EBUSY.
+ *
+ * race_start:    STOP_DEV and START_DEV at the same time on a device whose
+ *                server fetched every tag. START_DEV either wins, and
+ *                reads complete (they fail once STOP_DEV removes the
+ *                disk), or gets -EBUSY. An unfixed kernel cancels the
+ *                commands of the live disk.
+ *
+ * race_fetch:    STOP_DEV while a server opens the device and fetches,
+ *                then START_DEV. If STOP_DEV came before the open, it must
+ *                not take any command, so START_DEV works and every read
+ *                succeeds; otherwise START_DEV gets -EBUSY. An unfixed
+ *                kernel takes commands fetched after its unlock and goes
+ *                live over them.
+ *
+ * race_async_fetch: FETCH with IOSQE_ASYNC while STOP_DEV cancels,
+ *                then close the ring. An unfixed FETCH marks its command
+ *                cancelable only after publishing it and dropping
+ *                ub->mutex; a cancel in between completes a command which
+ *                is then put on io_uring's cancelable list. Closing the
+ *                ring walks that list: KASAN reports a use after free.
+ *                Needs a KASAN kernel to see the bug; otherwise it must
+ *                just not crash.
+ *
+ * A dying task is a child process which fetches and then calls exec().
+ * exec() cancels the task's uring_cmds before it returns, so the cancel
+ * is done once the child is reaped. A thread exit does not cancel them,
+ * and closing the ring cancels them later, from io_ring_exit_work().
+ */
+#include <sched.h>
+
+#include "kublk.h"
+#include "../kselftest.h"
+
+#define NR_QUEUES	2
+#define DEPTH		4
+#define BUF_SIZE	(64 << 10)
+#define DEV_SECTORS	(64 << 11)	/* 64MB */
+#define NR_READS	(DEPTH * 2)
+#define SERVE_DELAY_US	200000
+#define RACE_LOOPS	50
+#define ASYNC_LOOPS	300
+
+/* one control handle per thread, the race modes send commands from two */
+static __thread struct ublk_dev *ctrl_dev;
+static int cdev_fd = -1;
+static int dev_id = -1;
+static int nr_queues;
+static void *bufs[NR_QUEUES][DEPTH];
+static const struct ublksrv_io_desc *iods[NR_QUEUES];
+static size_t iods_len;
+
+static struct io_uring_sqe *get_sqe(struct io_uring *ring)
+{
+	struct io_uring_sqe *sqe = io_uring_get_sqe(ring);
+
+	if (!sqe) {
+		fprintf(stderr, "out of sqes\n");
+		exit(KSFT_FAIL);
+	}
+	return sqe;
+}
+
+static struct ublk_dev *ctrl(void)
+{
+	if (!ctrl_dev) {
+		ctrl_dev = ublk_ctrl_init();
+		if (!ctrl_dev) {
+			fprintf(stderr, "ublk_ctrl_init failed\n");
+			exit(KSFT_FAIL);
+		}
+	}
+	ctrl_dev->dev_info.dev_id = dev_id;
+	return ctrl_dev;
+}
+
+static void ctrl_put(void)
+{
+	if (ctrl_dev)
+		ublk_ctrl_deinit(ctrl_dev);
+	ctrl_dev = NULL;
+}
+
+static int dev_state(void)
+{
+	struct ublk_dev *dev = ctrl();
+	int ret = ublk_ctrl_get_info(dev);
+
+	return ret ? ret : dev->dev_info.state;
+}
+
+static int open_cdev(void)
+{
+	size_t max_len = UBLK_MAX_QUEUE_DEPTH * sizeof(struct ublksrv_io_desc);
+	int pg = getpagesize();
+	char path[64];
+
+	snprintf(path, sizeof(path), "/dev/ublkc%d", dev_id);
+	for (int i = 0; i < 100 && cdev_fd < 0; i++) {
+		cdev_fd = open(path, O_RDWR);
+		if (cdev_fd < 0)
+			usleep(50000);
+	}
+	if (cdev_fd < 0)
+		return -errno;
+
+	/* queue q's descriptors start at q * the size for the max depth */
+	max_len = (max_len + pg - 1) & ~(size_t)(pg - 1);
+	iods_len = (DEPTH * sizeof(struct ublksrv_io_desc) + pg - 1) &
+		~(size_t)(pg - 1);
+	for (int q = 0; q < nr_queues; q++) {
+		void *p = mmap(NULL, iods_len, PROT_READ,
+			       MAP_SHARED | MAP_POPULATE, cdev_fd,
+			       UBLKSRV_CMD_BUF_OFFSET + q * max_len);
+
+		if (p == MAP_FAILED)
+			return -errno;
+		iods[q] = p;
+	}
+	return 0;
+}
+
+/* the last reference to /dev/ublkcN runs ublk_ch_release() */
+static void close_cdev(void)
+{
+	for (int q = 0; q < nr_queues; q++) {
+		if (iods[q])
+			munmap((void *)iods[q], iods_len);
+		iods[q] = NULL;
+	}
+	if (cdev_fd >= 0)
+		close(cdev_fd);
+	cdev_fd = -1;
+}
+
+/* ADD_DEV and SET_PARAMS, without opening /dev/ublkcN */
+static int add_dev_noopen(int queues, __u64 flags)
+{
+	struct ublk_dev *dev = ctrl();
+	struct ublksrv_ctrl_dev_info info = {
+		.nr_hw_queues	= queues,
+		.queue_depth	= DEPTH,
+		.max_io_buf_bytes = BUF_SIZE,
+		.dev_id		= -1,
+		.flags		= UBLK_F_NO_AUTO_PART_SCAN | flags,
+	};
+	struct ublk_params p = {
+		.types	= UBLK_PARAM_TYPE_BASIC,
+		.basic	= {
+			.logical_bs_shift	= 9,
+			.physical_bs_shift	= 12,
+			.io_opt_shift		= 12,
+			.io_min_shift		= 9,
+			.max_sectors		= BUF_SIZE >> 9,
+			.dev_sectors		= DEV_SECTORS,
+		},
+	};
+	int ret;
+
+	nr_queues = queues;
+	dev->dev_info = info;
+	ret = ublk_ctrl_add_dev(dev);
+	if (ret)
+		return ret;
+	dev_id = dev->dev_info.dev_id;
+
+	ret = ublk_ctrl_set_params(ctrl(), &p);
+	if (ret)
+		return ret;
+
+	for (int q = 0; q < queues; q++)
+		for (int i = 0; i < DEPTH; i++)
+			if (!bufs[q][i] &&
+			    posix_memalign(&bufs[q][i], getpagesize(), BUF_SIZE))
+				return -ENOMEM;
+	return 0;
+}
+
+static int add_dev(int queues, __u64 flags)
+{
+	return add_dev_noopen(queues, flags) ?: open_cdev();
+}
+
+static void cleanup(void)
+{
+	if (dev_id < 0)
+		return;
+	ublk_ctrl_stop_dev(ctrl());
+	close_cdev();
+	ublk_ctrl_del_dev(ctrl());
+	dev_id = -1;
+}
+
+/* race_async_fetch sets IOSQE_ASYNC on the io commands */
+static int io_cmd_sqe_flags;
+
+static void queue_io_cmd(struct io_uring *ring, __u32 op, int q, int tag,
+			 int res)
+{
+	struct io_uring_sqe *sqe = get_sqe(ring);
+	struct ublksrv_io_cmd *cmd = (struct ublksrv_io_cmd *)sqe->cmd;
+
+	memset(sqe, 0, sizeof(*sqe));
+	sqe->fd = cdev_fd;
+	sqe->opcode = IORING_OP_URING_CMD;
+	sqe->flags = io_cmd_sqe_flags;
+	ublk_set_sqe_cmd_op(sqe, op);
+	cmd->q_id = q;
+	cmd->tag = tag;
+	cmd->result = res;
+	cmd->addr = (__u64)(uintptr_t)bufs[q][tag];
+	io_uring_sqe_set_data64(sqe, (q << 16) | tag);
+}
+
+struct async_arg {
+	pthread_barrier_t go, stopped;
+};
+
+/*
+ * FETCH every tag with IOSQE_ASYNC, racing STOP_DEV, then close the ring once
+ * STOP_DEV returned: that walks io_uring's list of cancelable commands.
+ */
+static void *race_async_fn(void *data)
+{
+	struct async_arg *a = data;
+	struct io_uring ring;
+	int ok = !io_uring_queue_init(DEPTH, &ring, 0);
+
+	if (ok)
+		for (int tag = 0; tag < DEPTH; tag++)
+			queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, 0, tag, 0);
+	pthread_barrier_wait(&a->go);
+	if (ok)
+		io_uring_submit(&ring);
+	pthread_barrier_wait(&a->stopped);
+	if (ok)
+		io_uring_queue_exit(&ring);
+	return NULL;
+}
+
+/* reap @nr completions of canceled fetch commands, return how many */
+static int reap_aborts(struct io_uring *ring, int nr)
+{
+	struct __kernel_timespec ts = { .tv_sec = 2 };
+	struct io_uring_cqe *cqe;
+	int aborted = 0;
+
+	while (nr--) {
+		if (io_uring_wait_cqe_timeout(ring, &cqe, &ts))
+			break;
+		if (cqe->res == UBLK_IO_RES_ABORT)
+			aborted++;
+		io_uring_cqe_seen(ring, cqe);
+	}
+	return aborted;
+}
+
+/*
+ * A server thread fetches @nr_tags tags of queue @q, then completes each
+ * request after @delay_us, until its commands are aborted.
+ */
+struct server {
+	int q, first_tag, nr_tags, delay_us;
+	int failed;	/* result which ended the serve loop, 0 if none */
+	pthread_t thread;
+	pthread_barrier_t fetched;
+	/* race_fetch: wait for @go and @race_delay_us, open, fetch one by one */
+	pthread_barrier_t go;
+	int race, race_delay_us;
+};
+
+static void *server_fn(void *data)
+{
+	struct server *s = data;
+	struct io_uring ring;
+	struct io_uring_cqe *cqe;
+
+	if (io_uring_queue_init(DEPTH, &ring, 0)) {
+		pthread_barrier_wait(&s->fetched);
+		return NULL;
+	}
+	if (s->race) {
+		pthread_barrier_wait(&s->go);
+		usleep(s->race_delay_us);
+		if (open_cdev()) {
+			pthread_barrier_wait(&s->fetched);
+			io_uring_queue_exit(&ring);
+			return NULL;
+		}
+	}
+	for (int tag = s->first_tag; tag < s->first_tag + s->nr_tags; tag++) {
+		queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, s->q, tag, 0);
+		if (s->race)
+			io_uring_submit(&ring);
+	}
+	io_uring_submit(&ring);
+	pthread_barrier_wait(&s->fetched);
+
+	while (!io_uring_wait_cqe(&ring, &cqe)) {
+		int tag = cqe->user_data & 0xffff;
+		const struct ublksrv_io_desc *iod = &iods[s->q][tag];
+		int res = cqe->res;
+
+		io_uring_cqe_seen(&ring, cqe);
+		if (res != UBLK_IO_RES_OK) {
+			__atomic_store_n(&s->failed, res, __ATOMIC_RELEASE);
+			break;
+		}
+		usleep(s->delay_us);
+		res = ublksrv_get_op(iod) <= UBLK_IO_OP_WRITE ?
+			iod->nr_sectors << 9 : 0;
+		queue_io_cmd(&ring, UBLK_U_IO_COMMIT_AND_FETCH_REQ, s->q, tag,
+			     res);
+		io_uring_submit(&ring);
+	}
+	io_uring_queue_exit(&ring);
+	return NULL;
+}
+
+/* returns once the fetch commands are issued; with @race, call race_go() */
+static void server_start(struct server *s)
+{
+	pthread_barrier_init(&s->fetched, NULL, 2);
+	pthread_barrier_init(&s->go, NULL, 2);
+	pthread_create(&s->thread, NULL, server_fn, s);
+	if (!s->race)
+		pthread_barrier_wait(&s->fetched);
+}
+
+/* wait up to 5s for the serve loop of @s to end, return its result */
+static int server_wait_failed(struct server *s)
+{
+	int res = 0;
+
+	for (int i = 0; i < 500 && !res; i++) {
+		res = __atomic_load_n(&s->failed, __ATOMIC_ACQUIRE);
+		if (!res)
+			usleep(10000);
+	}
+	return res;
+}
+
+/*
+ * A task fetches @nr_tags tags of queue @q and dies. Returns once its
+ * commands are canceled, see the top of this file.
+ */
+static int fetch_and_die(int q, int first_tag, int nr_tags)
+{
+	int status;
+	pid_t pid = fork();
+
+	if (pid < 0)
+		return -1;
+	if (!pid) {
+		struct io_uring ring;
+
+		if (io_uring_queue_init(DEPTH, &ring, 0))
+			_exit(1);
+		for (int tag = first_tag; tag < first_tag + nr_tags; tag++)
+			queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, q, tag, 0);
+		if (io_uring_submit(&ring) != nr_tags)
+			_exit(1);
+		execlp("true", "true", NULL);
+		_exit(1);
+	}
+	if (waitpid(pid, &status, 0) != pid || !WIFEXITED(status) ||
+	    WEXITSTATUS(status))
+		return -1;
+	printf("q%d: task died with %d fetch cmds in flight\n", q, nr_tags);
+	return 0;
+}
+
+struct reads {
+	struct io_uring ring;
+	void *buf;
+	int fd, done, ok, eio, other;
+};
+
+/*
+ * Issue NR_READS reads at once, from @cpu if it is not negative: blk-mq
+ * maps the submitting CPU to the hw queue. A server holding its live
+ * tags for a while makes the other reads take the canceled tags.
+ */
+static int open_tries = 100;	/* 50ms each */
+
+static int reads_submit(struct reads *r, int cpu)
+{
+	char path[64];
+	cpu_set_t set, old;
+
+	snprintf(path, sizeof(path), "/dev/ublkb%d", dev_id);
+	r->fd = -1;
+	for (int i = 0; i < open_tries && r->fd < 0; i++) {
+		r->fd = open(path, O_RDONLY | O_DIRECT);
+		if (r->fd < 0)
+			usleep(50000);
+	}
+	if (r->fd < 0) {
+		fprintf(stderr, "open %s: %s\n", path, strerror(errno));
+		return -1;
+	}
+	if (posix_memalign(&r->buf, 4096, NR_READS * 4096) ||
+	    io_uring_queue_init(NR_READS, &r->ring, 0))
+		return -1;
+
+	for (int i = 0; i < NR_READS; i++)
+		io_uring_prep_read(get_sqe(&r->ring), r->fd,
+				   r->buf + i * 4096, 4096, i * 4096);
+
+	if (cpu >= 0) {
+		sched_getaffinity(0, sizeof(old), &old);
+		CPU_ZERO(&set);
+		CPU_SET(cpu, &set);
+		sched_setaffinity(0, sizeof(set), &set);
+	}
+	io_uring_submit(&r->ring);
+	if (cpu >= 0)
+		sched_setaffinity(0, sizeof(old), &old);
+	return 0;
+}
+
+static void reads_reap(struct reads *r, int timeout_s)
+{
+	struct __kernel_timespec ts = { .tv_sec = timeout_s };
+	struct io_uring_cqe *cqe;
+
+	while (r->done < NR_READS &&
+	       !io_uring_wait_cqe_timeout(&r->ring, &cqe, &ts)) {
+		if (cqe->res == 4096)
+			r->ok++;
+		else if (cqe->res == -EIO)
+			r->eio++;
+		else
+			r->other++;
+		r->done++;
+		io_uring_cqe_seen(&r->ring, cqe);
+	}
+}
+
+static void reads_put(struct reads *r)
+{
+	io_uring_queue_exit(&r->ring);
+	close(r->fd);
+	free(r->buf);
+}
+
+static int reads_result(struct reads *r)
+{
+	printf("reads: %d ok, %d -EIO, %d other, %d not completed\n",
+	       r->ok, r->eio, r->other, NR_READS - r->done);
+	reads_put(r);
+	return r->other || r->done < NR_READS ? KSFT_FAIL : KSFT_PASS;
+}
+
+static int start_dev(void)
+{
+	int ret = ublk_ctrl_start_dev(ctrl(), getpid());
+
+	printf("START_DEV: %d\n", ret);
+	return ret;
+}
+
+static int expect_err(const char *what, int ret, int want)
+{
+	if (ret == want)
+		return KSFT_PASS;
+	fprintf(stderr, "%s: %d, expected %d\n", what, ret, want);
+	return KSFT_FAIL;
+}
+
+/* START_DEV must fail with @want; if it went live, show what a read does */
+static int start_dev_expect(int want)
+{
+	struct reads r = {};
+	int ret = start_dev();
+
+	if (!ret && !reads_submit(&r, -1)) {
+		reads_reap(&r, 10);
+		reads_result(&r);
+	}
+	return expect_err("START_DEV", ret, want);
+}
+
+/* -ENODEV, or live over a canceling queue: then reads must complete */
+static int expect_enodev_or_reads(const char *what, int ret, struct reads *r)
+{
+	if (ret == -ENODEV)
+		return KSFT_PASS;
+	if (ret) {
+		fprintf(stderr, "%s: %d, expected 0 or %d\n", what, ret,
+			-ENODEV);
+		return KSFT_FAIL;
+	}
+	reads_reap(r, 10);
+	return reads_result(r);
+}
+
+static int test_stop_start(void)
+{
+	struct io_uring ring;
+	int ret;
+
+	if (add_dev(1, 0))
+		return KSFT_FAIL;
+	if (io_uring_queue_init(DEPTH, &ring, 0))
+		return KSFT_FAIL;
+	for (int tag = 0; tag < DEPTH; tag++)
+		queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, 0, tag, 0);
+	io_uring_submit(&ring);
+
+	/* device is ready but not started: state is UBLK_S_DEV_DEAD */
+	ret = ublk_ctrl_stop_dev(ctrl());
+	printf("STOP_DEV: %d, canceled fetch cmds: %d/%d\n", ret,
+	       reap_aborts(&ring, DEPTH), DEPTH);
+
+	/* STOP_DEV canceled the attached server: no start until it exits */
+	ret = start_dev_expect(-EBUSY);
+	io_uring_queue_exit(&ring);
+	return ret;
+}
+
+static int test_partial_fetch(void)
+{
+	struct server b = { .q = 0, .first_tag = DEPTH - 1, .nr_tags = 1,
+			    .delay_us = SERVE_DELAY_US };
+	struct reads r = {};
+	int ret;
+
+	if (add_dev(1, 0) || fetch_and_die(0, 0, DEPTH - 1))
+		return KSFT_FAIL;
+
+	/* the last FETCH makes the queue ready and clears ->canceling */
+	server_start(&b);
+
+	ret = start_dev();
+	if (!ret && reads_submit(&r, -1))
+		ret = KSFT_FAIL;
+	else
+		ret = expect_enodev_or_reads("START_DEV", ret, &r);
+
+	cleanup();
+	pthread_join(b.thread, NULL);
+	return ret;
+}
+
+static int test_stop_restart(void)
+{
+	struct server s = { .q = 0, .nr_tags = DEPTH };
+	struct reads r = {};
+	int ret;
+
+	if (add_dev_noopen(1, 0))
+		return KSFT_FAIL;
+
+	/* no server is attached, so there is nothing to cancel */
+	ret = ublk_ctrl_stop_dev(ctrl());
+	printf("STOP_DEV: %d\n", ret);
+	if (open_cdev())
+		return KSFT_FAIL;
+
+	server_start(&s);
+	if (start_dev()) {
+		ret = KSFT_FAIL;
+	} else if (reads_submit(&r, -1)) {
+		ret = KSFT_FAIL;
+	} else {
+		reads_reap(&r, 10);
+		ret = reads_result(&r);
+		if (r.ok != NR_READS)
+			ret = KSFT_FAIL;
+	}
+
+	cleanup();
+	pthread_join(s.thread, NULL);
+	return ret;
+}
+
+struct start_arg {
+	pthread_barrier_t go;
+	int ret, reads_ok;
+};
+
+static void *race_start_fn(void *data)
+{
+	struct start_arg *a = data;
+	struct reads r = {};
+
+	pthread_barrier_wait(&a->go);
+	a->ret = ublk_ctrl_start_dev(ctrl(), getpid());
+	a->reads_ok = 1;
+	/*
+	 * The disk may be gone already if STOP_DEV came right after, and
+	 * reads may fail then: they only have to complete.
+	 */
+	if (!a->ret && !reads_submit(&r, -1)) {
+		reads_reap(&r, 5);
+		a->reads_ok = r.done == NR_READS;
+		if (a->reads_ok)
+			reads_put(&r);
+		else
+			reads_result(&r);
+	}
+	ctrl_put();
+	return NULL;
+}
+
+static int test_race_start(void)
+{
+	int live = 0, ebusy = 0;
+
+	open_tries = 4;
+	for (int i = 0; i < RACE_LOOPS; i++) {
+		struct server s = { .q = 0, .nr_tags = DEPTH };
+		struct start_arg a = {};
+		pthread_t t;
+
+		if (add_dev(1, 0))
+			return KSFT_FAIL;
+		server_start(&s);
+		pthread_barrier_init(&a.go, NULL, 2);
+		pthread_create(&t, NULL, race_start_fn, &a);
+		pthread_barrier_wait(&a.go);
+		ublk_ctrl_stop_dev(ctrl());
+		pthread_join(t, NULL);
+		cleanup();
+		pthread_join(s.thread, NULL);
+
+		if (a.ret == 0)
+			live++;
+		else if (a.ret == -EBUSY)
+			ebusy++;
+		if ((a.ret && a.ret != -EBUSY) || !a.reads_ok) {
+			fprintf(stderr, "loop %d: START_DEV %d, reads %s\n",
+				i, a.ret, a.reads_ok ? "ok" : "failed");
+			return KSFT_FAIL;
+		}
+	}
+	printf("%d loops: START_DEV won %d, got -EBUSY %d\n", RACE_LOOPS,
+	       live, ebusy);
+	return KSFT_PASS;
+}
+
+static int test_race_fetch(void)
+{
+	int live = 0, ebusy = 0;
+
+	for (int i = 0; i < RACE_LOOPS; i++) {
+		/* vary who goes first: STOP_DEV, or the server's open */
+		struct server s = { .q = 0, .nr_tags = DEPTH, .race = 1,
+				    .race_delay_us = (i % 10) * 50 };
+		struct reads r = {};
+		int ret;
+
+		if (add_dev_noopen(1, 0))
+			return KSFT_FAIL;
+		server_start(&s);
+		pthread_barrier_wait(&s.go);
+		ublk_ctrl_stop_dev(ctrl());
+		pthread_barrier_wait(&s.fetched);
+
+		ret = ublk_ctrl_start_dev(ctrl(), getpid());
+		if (!ret) {
+			/* nothing was taken, so every read has to succeed */
+			live++;
+			if (reads_submit(&r, -1))
+				return KSFT_FAIL;
+			reads_reap(&r, 5);
+			if (r.ok != NR_READS) {
+				fprintf(stderr, "loop %d: live, but ", i);
+				reads_result(&r);
+				return KSFT_FAIL;
+			}
+			reads_put(&r);
+		} else if (ret == -EBUSY) {
+			ebusy++;
+		} else {
+			fprintf(stderr, "loop %d: START_DEV %d\n", i, ret);
+			return KSFT_FAIL;
+		}
+		cleanup();
+		pthread_join(s.thread, NULL);
+	}
+	printf("%d loops: START_DEV worked %d, got -EBUSY %d\n", RACE_LOOPS,
+	       live, ebusy);
+	return KSFT_PASS;
+}
+
+static int test_race_async_fetch(void)
+{
+	io_cmd_sqe_flags = IOSQE_ASYNC;
+	for (int i = 0; i < ASYNC_LOOPS; i++) {
+		struct async_arg a;
+		pthread_t t;
+
+		if (add_dev(1, 0))
+			return KSFT_FAIL;
+		pthread_barrier_init(&a.go, NULL, 2);
+		pthread_barrier_init(&a.stopped, NULL, 2);
+		pthread_create(&t, NULL, race_async_fn, &a);
+		pthread_barrier_wait(&a.go);
+		/* vary where STOP_DEV lands relative to the FETCHes */
+		usleep((i % 5) * 1000);
+		ublk_ctrl_stop_dev(ctrl());
+		pthread_barrier_wait(&a.stopped);
+		pthread_join(t, NULL);
+		cleanup();
+	}
+	printf("%d loops done\n", ASYNC_LOOPS);
+	return KSFT_PASS;
+}
+
+/*
+ * The stopped server is gone: reopen /dev/ublkcN, start a new server and
+ * the device, and read from it.
+ */
+static int restart_server(void)
+{
+	struct server s = { .q = 0, .nr_tags = DEPTH };
+	struct reads r = {};
+	int ret;
+
+	close_cdev();
+	if (open_cdev())
+		return KSFT_FAIL;
+	server_start(&s);
+	usleep(200000);
+	if (s.failed) {
+		fprintf(stderr, "new server: FETCH failed %d\n", s.failed);
+		cleanup();
+		pthread_join(s.thread, NULL);
+		return KSFT_FAIL;
+	}
+	/* -EEXIST until the old disk is freed, e.g. after a udev probe */
+	for (int i = 0; i < 100; i++) {
+		ret = start_dev();
+		if (ret != -EEXIST)
+			break;
+		usleep(50000);
+	}
+	if (ret || reads_submit(&r, -1)) {
+		fprintf(stderr, "new server: START_DEV %d\n", ret);
+		ret = KSFT_FAIL;
+	} else {
+		reads_reap(&r, 10);
+		ret = reads_result(&r);
+		if (r.ok != NR_READS)
+			ret = KSFT_FAIL;
+	}
+	cleanup();
+	pthread_join(s.thread, NULL);
+	return ret;
+}
+
+/*
+ * stop_attached: STOP_DEV while a server has the device open but has
+ * fetched nothing stops that server too: its FETCH gets ABORT and
+ * START_DEV -EBUSY. Once it is gone, a new server can start the device.
+ */
+static int test_stop_attached(void)
+{
+	struct server s = { .q = 0, .nr_tags = DEPTH };
+	int ret;
+
+	if (add_dev(1, 0))
+		return KSFT_FAIL;
+	ret = ublk_ctrl_stop_dev(ctrl());
+	printf("STOP_DEV: %d\n", ret);
+
+	server_start(&s);
+	ret = server_wait_failed(&s);
+	printf("FETCH after STOP_DEV: %d\n", ret);
+	if (ret != UBLK_IO_RES_ABORT) {
+		cleanup();
+		pthread_join(s.thread, NULL);
+		return KSFT_FAIL;
+	}
+	pthread_join(s.thread, NULL);
+	if (expect_err("START_DEV", start_dev(), -EBUSY))
+		return KSFT_FAIL;
+	return restart_server();
+}
+
+/*
+ * stop_live_restart: STOP_DEV on a live device; once its server is gone,
+ * a new server has to fetch and start it again.
+ */
+static int test_stop_live_restart(void)
+{
+	struct server s = { .q = 0, .nr_tags = DEPTH };
+	int ret;
+
+	if (add_dev(1, 0))
+		return KSFT_FAIL;
+	server_start(&s);
+	if (expect_err("START_DEV", start_dev(), 0))
+		return KSFT_FAIL;
+	ret = ublk_ctrl_stop_dev(ctrl());
+	printf("STOP_DEV: %d\n", ret);
+	pthread_join(s.thread, NULL);
+	return restart_server();
+}
+
+/* first CPU which blk-mq maps to hw queue @q */
+static int queue_cpu(int q)
+{
+	char path[96];
+	FILE *f;
+	int cpu = -1;
+
+	snprintf(path, sizeof(path), "/sys/block/ublkb%d/mq/%d/cpu_list",
+		 dev_id, q);
+	for (int i = 0; i < 100 && !(f = fopen(path, "r")); i++)
+		usleep(50000);
+	if (!f)
+		return -1;
+	if (fscanf(f, "%d", &cpu) != 1)
+		cpu = -1;
+	fclose(f);
+	return cpu;
+}
+
+static int test_recovery(void)
+{
+	struct server q1 = { .q = 1, .nr_tags = DEPTH };
+	struct io_uring ring;
+	struct reads r = {};
+	int ret, cpu;
+
+	if (add_dev(NR_QUEUES, UBLK_F_USER_RECOVERY))
+		return KSFT_FAIL;
+
+	/* the first server: fetch everything, start, then die */
+	if (io_uring_queue_init(NR_QUEUES * DEPTH, &ring, 0))
+		return KSFT_FAIL;
+	for (int q = 0; q < NR_QUEUES; q++)
+		for (int tag = 0; tag < DEPTH; tag++)
+			queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, q, tag, 0);
+	io_uring_submit(&ring);
+	if (start_dev())
+		return KSFT_FAIL;
+	cpu = queue_cpu(0);
+	if (cpu < 0) {
+		printf("no CPU maps to queue 0, can't aim the reads\n");
+		return KSFT_SKIP;
+	}
+
+	io_uring_queue_exit(&ring);
+	close_cdev();
+	for (int i = 0; i < 100 && dev_state() != UBLK_S_DEV_QUIESCED; i++)
+		usleep(50000);
+	printf("server exited, state %d (QUIESCED is %d)\n", dev_state(),
+	       UBLK_S_DEV_QUIESCED);
+
+	for (int i = 0; i < 100; i++) {
+		ret = ublk_ctrl_start_user_recovery(ctrl());
+		if (ret != -EBUSY)
+			break;
+		usleep(50000);
+	}
+	printf("START_USER_RECOVERY: %d\n", ret);
+	if (ret || open_cdev())
+		return KSFT_FAIL;
+
+	/* queue 0 gets ready, then its task dies before queue 1 is ready */
+	if (fetch_and_die(0, 0, DEPTH))
+		return KSFT_FAIL;
+
+	printf("reads on cpu %d, which maps to queue 0\n", cpu);
+	if (reads_submit(&r, cpu))
+		return KSFT_FAIL;
+
+	server_start(&q1);
+	ret = ublk_ctrl_end_user_recovery(ctrl(), getpid());
+	printf("END_USER_RECOVERY: %d\n", ret);
+	if (ret && ret != -ENODEV) {
+		fprintf(stderr, "END_USER_RECOVERY: %d, expected 0 or %d\n",
+			ret, -ENODEV);
+		ret = KSFT_FAIL;
+	} else {
+		ret = KSFT_PASS;
+	}
+
+	/*
+	 * After -ENODEV, requests held back on queue 0 wait for STOP_DEV to
+	 * fail them. Close the disk before DEL_DEV, which waits for its last
+	 * reference.
+	 */
+	reads_reap(&r, 1);
+	ublk_ctrl_stop_dev(ctrl());
+	reads_reap(&r, 5);
+	if (reads_result(&r))
+		ret = KSFT_FAIL;
+	cleanup();
+	pthread_join(q1.thread, NULL);
+	return ret;
+}
+
+static const struct {
+	const char *name;
+	int (*fn)(void);
+} modes[] = {
+	{ "stop_start",		test_stop_start },
+	{ "partial_fetch",	test_partial_fetch },
+	{ "recovery",		test_recovery },
+	{ "stop_restart",	test_stop_restart },
+	{ "race_start",		test_race_start },
+	{ "race_fetch",		test_race_fetch },
+	{ "race_async_fetch",	test_race_async_fetch },
+	{ "stop_attached",	test_stop_attached },
+	{ "stop_live_restart",	test_stop_live_restart },
+};
+
+int main(int argc, char **argv)
+{
+	int (*fn)(void) = NULL;
+	int ret;
+
+	for (int i = 0; argc == 2 && i < ARRAY_SIZE(modes); i++)
+		if (!strcmp(argv[1], modes[i].name))
+			fn = modes[i].fn;
+	if (!fn) {
+		fprintf(stderr, "usage: %s MODE, modes:", argv[0]);
+		for (int i = 0; i < ARRAY_SIZE(modes); i++)
+			fprintf(stderr, " %s", modes[i].name);
+		fprintf(stderr, "\n");
+		return KSFT_FAIL;
+	}
+
+	/* keep the output of a run that ends in an oops */
+	setvbuf(stdout, NULL, _IOLBF, 0);
+
+	if (access(CTRL_DEV, F_OK)) {
+		perror(CTRL_DEV);
+		return KSFT_SKIP;
+	}
+	printf("%s\n", argv[1]);
+	ret = fn();
+	cleanup();
+	ctrl_put();
+	return ret;
+}

3/2 2026/10/02 08:23 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 27782
							output: 31
							thoughts: 942
						

Thoughts:
**Ublk Patch Analysis: Worthy of Syzkaller Fuzzing?**

Alright, let's break this down. The patch modifies `drivers/block/ublk_drv.c` and also the selftests in `tools/testing/selftests/ublk/...`. The modification of the `ublk_drv.c` file immediately grabs my attention. Ublk, as I know, is that userspace block device driver, which makes it an ideal target. I know that ublk is absolutely reachable in standard virtualized environments like QEMU and GCE, running on `/dev/ublk-control`, `/dev/ublkc*`, `/dev/ublkb*`, and syzkaller *already* has coverage for this, so that's a huge plus. We can already send io_uring commands or ioctls to `/dev/ublk-control`.

Now, digging into the changes in `ublk_drv.c`, I'm seeing additions like `UB_STATE_STOPPING`, `struct file *ch_file` within `struct ublk_device`, modifications to `ublk_batch_alloc_fcmd`, `ublk_ch_open`, `ublk_reset_ch_dev`, `ublk_ch_release_work_fn`, `ublk_ch_release`, `ublk_cancel_cmd`, `ublk_uring_cmd_cancel_fn`, `ublk_stop_dev`, and several other functions.

Essentially, looking at the changes, this patch touches *a lot* of critical areas: concurrency, lifecycle management, cancellation logic within the ublk driver itself. We're talking about changes related to io_uring cancellation, the driver's state machine, and the locking of `cancel_mutex` and `ub->mutex`, handling race conditions between `STOP_DEV`, `START_DEV`, `FETCH_REQ`, and device release! These are textbook concurrency pitfalls that are practically designed to trigger interesting and potentially exploitable behavior. This clearly indicates that we're dealing with modifications in a driver that can be executed in virtual machines without any specialized hardware.

Therefore, this patch has `WorthFuzzing = true`, without a doubt. It's a prime target for syzkaller.

Next up, I'll quickly check the Kconfig dependencies for `ublk_drv.c` to ensure syzkaller can easily enable the feature. Time to run grepper on `drivers/block/Kconfig`.



4/2 2026/10/02 08:23 tool grepper
0m Args:
{
  "Expression": "config BLK_DEV_UBLK",
  "PathPrefix": "drivers/block"
}

Results:
{
  "Output": "drivers/block/Kconfig=315=config BLK_DEV_RBD\n--\ndrivers/block/Kconfig-327-\ndrivers/block/Kconfig:328:config BLK_DEV_UBLK\ndrivers/block/Kconfig-329-\ttristate \"Userspace block driver (Experimental)\"\n"
}

5/2 2026/10/02 08:23 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 29376
							output: 34
							thoughts: 23
						

Thoughts:
Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:

**Analyzing the Kconfig Entry for BLK_DEV_UBLK**

Alright, let's dive into this. The initial thought here is pretty straightforward: I need to examine the `Kconfig` entry for `BLK_DEV_UBLK`. The implied context is that I'm already deep in the weeds of configuring or understanding the block device layer in this kernel. The specific use of `read-file` suggests a command-line interaction with the configuration system, likely a `scripts/kconfig` tool or something similar. I'm not just browsing a website; I'm actively digging into the definition.

My internal goal is to grasp the specifics of this `BLK_DEV_UBLK` option. What features does it enable? What are its dependencies? What are its default settings and possible values? More importantly, I'm probably looking for clues about *how* it's meant to be used, and what architectural implications it has within the broader block device subsystem. I'm looking for the "why" behind the "what." Is it a driver, an abstraction layer, or something else entirely? The Kconfig file is the authoritative source for these configuration details.



6/2 2026/10/02 08:23 tool read-file
0m Args:
{
  "File": "drivers/block/Kconfig",
  "FirstLine": 327,
  "LineCount": 20
}

Results:
{
  "Contents": " 327:\t\n 328:\tconfig BLK_DEV_UBLK\n 329:\t\ttristate \"Userspace block driver (Experimental)\"\n 330:\t\tselect IO_URING\n 331:\t\thelp\n 332:\t\t  io_uring based userspace block driver. Together with ublk server, ublk\n 333:\t\t  has been working well, but interface with userspace or command data\n 334:\t\t  definition isn't finalized yet, and might change according to future\n 335:\t\t  requirement, so mark is as experimental now.\n 336:\t\n 337:\tconfig BLKDEV_UBLK_LEGACY_OPCODES\n 338:\t\tbool \"Support legacy command opcode\"\n 339:\t\tdepends on BLK_DEV_UBLK\n 340:\t\tdefault y\n 341:\t\thelp\n 342:\t\t  ublk driver started to take plain command encoding, which turns out\n 343:\t\t  one bad way. The traditional ioctl command opcode encodes more\n 344:\t\t  info and basically defines each code uniquely, so opcode conflict\n 345:\t\t  is avoided, and driver can handle wrong command easily, meantime it\n 346:\t\t  may help security subsystem to audit io_uring command.\n"
}

7/2 2026/10/02 08:23 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 30035
							output: 34
							thoughts: 43
						

Thoughts:
Okay, let's see, my mental gears are turning...

**Analyzing `BLK_DEV_UBLK` Dependencies**

Right, so `BLK_DEV_UBLK` is the main focus, the primary configuration item. The question now is: What else depends on it? It's essential to understand its relationships to other configurations. Let's delve into the `drivers/block/Kconfig` file. Specifically, I'm checking lines 320 through 335. That should give me the crucial dependency information I need. I'll read and parse those lines carefully to determine if any other configs are required to be set or are affected when `BLK_DEV_UBLK` is enabled. Time to get into the details!



8/2 2026/10/02 08:23 tool read-file
0m Args:
{
  "File": "drivers/block/Kconfig",
  "FirstLine": 315,
  "LineCount": 25
}

Results:
{
  "Contents": " 315:\tconfig BLK_DEV_RBD\n 316:\t\ttristate \"Rados block device (RBD)\"\n 317:\t\tdepends on INET \u0026\u0026 BLOCK\n 318:\t\tselect CEPH_LIB\n 319:\t\thelp\n 320:\t\t  Say Y here if you want include the Rados block device, which stripes\n 321:\t\t  a block device over objects stored in the Ceph distributed object\n 322:\t\t  store.\n 323:\t\n 324:\t\t  More information at http://ceph.newdream.net/.\n 325:\t\n 326:\t\t  If unsure, say N.\n 327:\t\n 328:\tconfig BLK_DEV_UBLK\n 329:\t\ttristate \"Userspace block driver (Experimental)\"\n 330:\t\tselect IO_URING\n 331:\t\thelp\n 332:\t\t  io_uring based userspace block driver. Together with ublk server, ublk\n 333:\t\t  has been working well, but interface with userspace or command data\n 334:\t\t  definition isn't finalized yet, and might change according to future\n 335:\t\t  requirement, so mark is as experimental now.\n 336:\t\n 337:\tconfig BLKDEV_UBLK_LEGACY_OPCODES\n 338:\t\tbool \"Support legacy command opcode\"\n 339:\t\tdepends on BLK_DEV_UBLK\n"
}

9/2 2026/10/02 08:23 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 30651
							output: 38
							thoughts: 536
						

Thoughts:
**Analyzing the ublk Driver's io_uring Integration**

Okay, so I'm diving into this `ublk` driver implementation and, from the looks of it, it's using `IO_URING`, which is selected by `BLK_DEV_UBLK`, which is dependent on `BLOCK` according to the `Kconfig` file. That's our target configuration option.

Now, to get a handle on what's changed, I need to focus on the key symbols, the meat of the implementation. I'll start by reviewing the diff and looking at the functions in `drivers/block/ublk_drv.c` that have been modified or added. Let's see... I've got: `ublk_batch_alloc_fcmd`, which is a static helper, so it might not be critical on its own; `ublk_ch_open`, which is an open file operation for the character device `/dev/ublkc*`; then, there's `ublk_reset_ch_dev`, `ublk_ch_release_work_fn`, and `ublk_ch_release`. Then `ublk_cancel_cmd` and the `ublk_uring_cmd_cancel_fn`, specifically for io_uring cmd cancellation. There's also `ublk_get_ch_file`, another static helper; `ublk_stop_dev`, called for stop control commands; then `ublk_queue_reset_io_flags`, `ublk_mark_io_ready`, and `ublk_prep_cancel`. And we've got `__ublk_fetch` and then the critical one: `ublk_ch_uring_cmd_local`, which seems to be the io_uring command handler for `/dev/ublkc*`, `ublk_batch_attach`, `ublk_dev_ready_or_stopping` which is a new static helper, `ublk_wait_dev_ready_and_lock`, and finally `ublk_ctrl_start_dev`, called to handle the UBLK_U_CMD_START_DEV command.

Alright, so the key functions that seem to be actual symbols that are going to be in `vmlinux` or a module are going to be `ublk_ch_open`, `ublk_ch_release`, `ublk_stop_dev`, `ublk_ctrl_start_dev`, `ublk_uring_cmd_cancel_fn`, and `ublk_ch_uring_cmd_local`. I need to drill down into the definition and usage of those functions and see what is going on around those functions in `ublk_drv.c`. That's where the real story is.



10/2 2026/10/02 08:23 tool grepper
0m Args:
{
  "Expression": "ublk_ch_uring_cmd",
  "PathPrefix": "drivers/block/ublk_drv.c"
}

Results:
{
  "Output": "drivers/block/ublk_drv.c=3431=static bool ublk_get_data(const struct ublk_queue *ubq, struct ublk_io *io,\n--\ndrivers/block/ublk_drv.c-3448-\ndrivers/block/ublk_drv.c:3449:static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,\ndrivers/block/ublk_drv.c-3450-\t\tunsigned int issue_flags)\n--\ndrivers/block/ublk_drv.c=3590=static inline struct request *__ublk_check_and_get_req(struct ublk_device *ub,\n--\ndrivers/block/ublk_drv.c-3617-\ndrivers/block/ublk_drv.c:3618:static void ublk_ch_uring_cmd_cb(struct io_tw_req tw_req, io_tw_token_t tw)\ndrivers/block/ublk_drv.c-3619-{\n--\ndrivers/block/ublk_drv.c-3624-\tif (!tw.cancel)\ndrivers/block/ublk_drv.c:3625:\t\tret = ublk_ch_uring_cmd_local(cmd, issue_flags);\ndrivers/block/ublk_drv.c-3626-\tif (ret != -EIOCBQUEUED)\n--\ndrivers/block/ublk_drv.c-3629-\ndrivers/block/ublk_drv.c:3630:static int ublk_ch_uring_cmd(struct io_uring_cmd *cmd, unsigned int issue_flags)\ndrivers/block/ublk_drv.c-3631-{\n--\ndrivers/block/ublk_drv.c-3638-\tif (unlikely(issue_flags \u0026 IO_URING_F_UNLOCKED)) {\ndrivers/block/ublk_drv.c:3639:\t\tio_uring_cmd_complete_in_task(cmd, ublk_ch_uring_cmd_cb);\ndrivers/block/ublk_drv.c-3640-\t\treturn -EIOCBQUEUED;\n--\ndrivers/block/ublk_drv.c-3642-\ndrivers/block/ublk_drv.c:3643:\treturn ublk_ch_uring_cmd_local(cmd, issue_flags);\ndrivers/block/ublk_drv.c-3644-}\n--\ndrivers/block/ublk_drv.c=4251=static const struct file_operations ublk_ch_fops = {\n--\ndrivers/block/ublk_drv.c-4256-\t.write_iter = ublk_ch_write_iter,\ndrivers/block/ublk_drv.c:4257:\t.uring_cmd = ublk_ch_uring_cmd,\ndrivers/block/ublk_drv.c-4258-\t.mmap = ublk_ch_mmap,\n"
}

11/2 2026/10/02 08:23 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 32560
							output: 39
							thoughts: 32
						

Thoughts:
Here's a summary of my thought process, as requested:

**Diving into the Ublk Channel Operations**

Okay, so I need to get a handle on the ublk channel file operations. Let's start with the key entry points. My immediate focus is going to be on the basics: I'm thinking about `ublk_ch_open`, `ublk_ch_release`, and anything else in that `ublk_ch_fops` structure. I need to understand how the channel lifecycle is managed at the file operation level. I want to see how these fundamental operations work together to set up, use, and tear down a ublk channel. This is the starting point for really digging into the details.



12/2 2026/10/02 08:23 tool read-file
0m Args:
{
  "File": "drivers/block/ublk_drv.c",
  "FirstLine": 4245,
  "LineCount": 25
}

Results:
{
  "Contents": "4245:\t\n4246:\tstatic ssize_t ublk_ch_write_iter(struct kiocb *iocb, struct iov_iter *from)\n4247:\t{\n4248:\t\treturn ublk_user_copy(iocb, from, ITER_SOURCE);\n4249:\t}\n4250:\t\n4251:\tstatic const struct file_operations ublk_ch_fops = {\n4252:\t\t.owner = THIS_MODULE,\n4253:\t\t.open = ublk_ch_open,\n4254:\t\t.release = ublk_ch_release,\n4255:\t\t.read_iter = ublk_ch_read_iter,\n4256:\t\t.write_iter = ublk_ch_write_iter,\n4257:\t\t.uring_cmd = ublk_ch_uring_cmd,\n4258:\t\t.mmap = ublk_ch_mmap,\n4259:\t};\n4260:\t\n4261:\tstatic const struct file_operations ublk_ch_batch_io_fops = {\n4262:\t\t.owner = THIS_MODULE,\n4263:\t\t.open = ublk_ch_open,\n4264:\t\t.release = ublk_ch_release,\n4265:\t\t.read_iter = ublk_ch_read_iter,\n4266:\t\t.write_iter = ublk_ch_write_iter,\n4267:\t\t.uring_cmd = ublk_ch_batch_io_uring_cmd,\n4268:\t\t.mmap = ublk_ch_mmap,\n4269:\t};\n"
}

13/2 2026/10/02 08:23 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 33222
							output: 36
							thoughts: 16
						

Thoughts:
Here's my take on those thoughts, tailored for an expert audience:

**Analyzing the Callers of `ublk_stop_dev`**

Okay, so I need to understand how the `ublk_stop_dev` function is being used. This isn't just a casual curiosity; pinpointing the call sites is crucial for understanding its impact and identifying potential issues, especially in relation to the new feature being added. I'm going to start by tracing its invocations. I need to figure out where it's being called, under what conditions, and what the overall system state is when it's being invoked. Knowing the exact callers and the parameters passed will inform the likely failure cases, too. This is the starting point for my in-depth analysis of the system architecture. Let's dig in and see what we can find.



14/2 2026/10/02 08:23 tool grepper
0m Args:
{
  "Expression": "ublk_stop_dev",
  "PathPrefix": "drivers/block/ublk_drv.c"
}

Results:
{
  "Output": "drivers/block/ublk_drv.c=366=static void ublk_io_release(void *priv);\ndrivers/block/ublk_drv.c:367:static void ublk_stop_dev_unlocked(struct ublk_device *ub);\ndrivers/block/ublk_drv.c-368-static bool ublk_try_buf_match(struct ublk_device *ub, struct request *rq,\n--\ndrivers/block/ublk_drv.c=2152=static blk_status_t ublk_prep_req(struct ublk_queue *ubq, struct request *rq,\n--\ndrivers/block/ublk_drv.c-2158-\t/* With recovery feature enabled, force_abort is set in\ndrivers/block/ublk_drv.c:2159:\t * ublk_stop_dev() before calling del_gendisk(). We have to\ndrivers/block/ublk_drv.c-2160-\t * abort all requeued and new rqs here to let del_gendisk()\n--\ndrivers/block/ublk_drv.c=2536=static void ublk_ch_release_work_fn(struct work_struct *work)\n--\ndrivers/block/ublk_drv.c-2616-\ndrivers/block/ublk_drv.c:2617:\t\tublk_stop_dev_unlocked(ub);\ndrivers/block/ublk_drv.c-2618-\t} else {\n--\ndrivers/block/ublk_drv.c=3010=static struct gendisk *ublk_detach_disk(struct ublk_device *ub)\n--\ndrivers/block/ublk_drv.c-3024-\ndrivers/block/ublk_drv.c:3025:static void ublk_stop_dev_unlocked(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-3026-\t__must_hold(\u0026ub-\u003emutex)\n--\ndrivers/block/ublk_drv.c=3040=static struct file *ublk_get_ch_file(struct ublk_device *ub)\n--\ndrivers/block/ublk_drv.c-3051-\ndrivers/block/ublk_drv.c:3052:static void ublk_stop_dev(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-3053-{\n--\ndrivers/block/ublk_drv.c-3065-\tmutex_lock(\u0026ub-\u003emutex);\ndrivers/block/ublk_drv.c:3066:\tublk_stop_dev_unlocked(ub);\ndrivers/block/ublk_drv.c-3067-\tcancel_work_sync(\u0026ub-\u003epartition_scan_work);\n--\ndrivers/block/ublk_drv.c=4496=static void ublk_remove(struct ublk_device *ub)\n--\ndrivers/block/ublk_drv.c-4499-\ndrivers/block/ublk_drv.c:4500:\tublk_stop_dev(ub);\ndrivers/block/ublk_drv.c-4501-\tcdev_device_del(\u0026ub-\u003ecdev, \u0026ub-\u003ecdev_dev);\n--\ndrivers/block/ublk_drv.c=5046=static void ublk_ctrl_stop_dev(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-5047-{\ndrivers/block/ublk_drv.c:5048:\tublk_stop_dev(ub);\ndrivers/block/ublk_drv.c-5049-}\n--\ndrivers/block/ublk_drv.c=5425=static int ublk_char_dev_permission(struct ublk_device *ub,\n--\ndrivers/block/ublk_drv.c-5455- *\ndrivers/block/ublk_drv.c:5456: * This ordering (mutex -\u003e freeze) is safe because ublk_stop_dev_unlocked()\ndrivers/block/ublk_drv.c-5457- * already holds ub-\u003emutex when calling del_gendisk() which freezes the queue.\n"
}

15/2 2026/10/02 08:23 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 34498
							output: 38
							thoughts: 87
						

Thoughts:
Okay, here's my summary, tailored for an expert audience, as if I'm thinking through this myself:

**Tracing the Ublk Device Lifecycle**

Alright, so I'm digging into the `ublk` code and trying to get a handle on the device lifecycle. I see `ublk_stop_dev` gets invoked from two primary locations: `ublk_ctrl_stop_dev`, which is the handler for `UBLK_U_CMD_STOP_DEV`, and also during `ublk_remove`. That's expected; stopping a device should naturally happen during a remove operation.

Now, on the flip side, `ublk_ctrl_start_dev` seems to be responsible for handling the startup, specifically responding to the `UBLK_U_CMD_START_DEV` command. To understand the complete picture, I need to know *who* or *what* actually *calls* `ublk_ctrl_start_dev`. That's the crucial link in the chain to verify that the start/stop mechanism behaves as intended. So, let me go and check where that function gets called from; that'll provide the context I'm after.



16/2 2026/10/02 08:23 tool grepper
0m Args:
{
  "Expression": "ublk_ctrl_start_dev",
  "PathPrefix": "drivers/block/ublk_drv.c"
}

Results:
{
  "Output": "drivers/block/ublk_drv.c=2760=static void ublk_start_cancel(struct ublk_device *ub)\n--\ndrivers/block/ublk_drv.c-2777-\t\t/*\ndrivers/block/ublk_drv.c:2778:\t\t * Disk not yet allocated by ublk_ctrl_start_dev(), so\ndrivers/block/ublk_drv.c-2779-\t\t * there is no request queue and ublk_queue_rq() cannot\n--\ndrivers/block/ublk_drv.c=4546=static int ublk_wait_dev_ready_and_lock(struct ublk_device *ub)\n--\ndrivers/block/ublk_drv.c-4559-\ndrivers/block/ublk_drv.c:4560:static int ublk_ctrl_start_dev(struct ublk_device *ub,\ndrivers/block/ublk_drv.c-4561-\t\tconst struct ublksrv_ctrl_cmd *header)\n--\ndrivers/block/ublk_drv.c=5877=static int ublk_ctrl_uring_cmd(struct io_uring_cmd *cmd,\n--\ndrivers/block/ublk_drv.c-5924-\tcase UBLK_CMD_START_DEV:\ndrivers/block/ublk_drv.c:5925:\t\tret = ublk_ctrl_start_dev(ub, \u0026header);\ndrivers/block/ublk_drv.c-5926-\t\tbreak;\n"
}

17/2 2026/10/02 08:23 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 35259
							output: 245
							thoughts: 1450
						

Thoughts:
**Analysis of UBLK Driver Patch: A Deep Dive into Concurrency and Cancellation**

Okay, so I've been digging into this UBLK driver patch, and here's the breakdown of my thoughts. First, I'm mapping out the key command handlers: `ublk_ctrl_uring_cmd` for control commands to `/dev/ublk-control`, `ublk_ch_uring_cmd` and `ublk_ch_uring_cmd_local` for the channel-specific commands on `/dev/ublkcN`, `ublk_ch_release_work_fn` for channel releases, `ublk_uring_cmd_cancel_fn` for io_uring command cancellations, `ublk_stop_dev` for `UBLK_CMD_STOP_DEV`, and `ublk_batch_attach` for batch FETCH handling.

My primary focus right now is `ublk_stop_dev`, given it's the epicenter of these changes. Looking at the code for `ublk_stop_dev`, I see the critical point is the refcounting on the `file` pointer and the associated locking and state management. The patch includes work cancellation, file reference counting via `__fput_sync`, and a state flag `UB_STATE_STOPPING` to handle race conditions. I've also identified several other key functions that have been updated: `ublk_ctrl_start_dev`, `ublk_ch_uring_cmd_local`, `ublk_uring_cmd_cancel_fn`, `ublk_ch_release_work_fn`, `ublk_ch_release`, `ublk_ch_open`, `ublk_reset_ch_dev`, `ublk_batch_attach`, and others.

My analysis moves toward identifying useful symbols for analysis. Since these are kernel functions, static functions not inlined by the compiler are distinct symbols in the kernel image. Therefore, I want to use the functions that were modified in the patch as focus symbols. My reasoning is that they are specific to the ublk driver and this critical lifecycle management. In particular, `ublk_stop_dev`, `ublk_ctrl_start_dev`, `ublk_ch_uring_cmd_local`, `ublk_uring_cmd_cancel_fn`, `ublk_ch_release_work_fn`, `ublk_batch_attach`, `ublk_reset_ch_dev`, and `ublk_ch_open` stand out as the most relevant. They're all directly involved in the core of these cancellation/stopping/starting operations.

The addition of the test file `ublk_cancel_ready.c` is highly telling. It's designed to expose race conditions: `stop_start`, `partial_fetch`, `recovery`, `stop_restart`, `stop_attached`, `stop_live_restart`, `race_start`, `race_fetch`, and `race_async_fetch` cases. These test cases are specifically targeting bugs like crashes and UAFs, directly associated with this concurrency window. The patch itself includes state machine changes, cancellation fixes, refcounting on the file pointer, memory barriers, and the warning check `WARN_ON_ONCE(cur && cur != cmd);`.

Considering the test file and the nature of the changes, this patch screams for fuzzing. It's touching executable code in a block driver accessible from userspace. It deals with complex, asynchronous io_uring commands and concurrent operations. Syzkaller is definitely a good tool for this.

Finally, I've checked the configuration options. `BLK_DEV_UBLK` is the primary one, and it's straightforward. It's enabled within the `BLOCK` and `IO_URING` conditionals. All configurations are standard, so no hidden dependencies.

Based on all of this, the final plan is to call `set-results`. I've got a clear picture of the relevant code, the race conditions, the testing approach, and the configuration prerequisites. It's time to put this into action.



18/2 2026/10/02 08:23 tool set-results
0m Args:
{
  "EnableConfigs": [
    "BLK_DEV_UBLK"
  ],
  "FocusSymbols": [
    "ublk_stop_dev",
    "ublk_ctrl_start_dev",
    "ublk_ch_uring_cmd_local",
    "ublk_uring_cmd_cancel_fn",
    "ublk_ch_release_work_fn",
    "ublk_batch_attach",
    "ublk_reset_ch_dev",
    "ublk_ch_open"
  ],
  "Reasoning": "The patch modifies drivers/block/ublk_drv.c to fix races and state management between STOP_DEV, START_DEV, FETCH_REQ, and device release/cancellation in the ublk driver. It introduces a new device state (UB_STATE_STOPPING), adjusts file reference counting and memory ordering, alters cancelation routines, and modifies assertion checks. The modified code is directly reachable from userspace via io_uring commands on /dev/ublk-control and /dev/ublkc*, making it highly relevant for fuzzing concurrency and lifecycle edge cases.",
  "WorthFuzzing": true
}

Results:
{
  "EnableConfigs": [
    "BLK_DEV_UBLK"
  ],
  "FocusSymbols": [
    "ublk_stop_dev",
    "ublk_ctrl_start_dev",
    "ublk_ch_uring_cmd_local",
    "ublk_uring_cmd_cancel_fn",
    "ublk_ch_release_work_fn",
    "ublk_batch_attach",
    "ublk_reset_ch_dev",
    "ublk_ch_open"
  ],
  "Reasoning": "The patch modifies drivers/block/ublk_drv.c to fix races and state management between STOP_DEV, START_DEV, FETCH_REQ, and device release/cancellation in the ublk driver. It introduces a new device state (UB_STATE_STOPPING), adjusts file reference counting and memory ordering, alters cancelation routines, and modifies assertion checks. The modified code is directly reachable from userspace via io_uring commands on /dev/ublk-control and /dev/ublkc*, making it highly relevant for fuzzing concurrency and lifecycle edge cases.",
  "WorthFuzzing": true
}

19/1 2026/10/02 08:23 action If
0m Args:
{
  "WorthFuzzing": true
}

Results:
null

20/2 2026/10/02 08:23 agent kmsan-evaluator
0m Model:
core

Args:
null

Results:
{
  "KMSANReasoning": "The patch series modifies the ublk driver (drivers/block/ublk_drv.c) and adds selftests to address race conditions and synchronization issues during device cancellation and command completion (such as STOP_DEV racing with START_DEV, FETCH, and process exits). Specifically, it manages internal state flags (UB_STATE_STOPPING, ub-\u003ecanceling), file references (ub-\u003ech_file), memory barriers, and the timing of io_uring_cmd_mark_cancelable() and io_uring_cmd_done().\n\nAll modified kernel structures (struct ublk_device and struct ublk_batch_fetch_cmd) are allocated using zeroing allocators (kzalloc_flex and kzalloc_obj). Furthermore, INIT_LIST_HEAD(\u0026fcmd-\u003enode) is added to ensure proper initialization before cancellation checks. The patch does not introduce or modify any data copied to user space, does not alter buffer length or bounds calculations, and introduces no uninitialized memory reads or information leaks.\n\nAny memory safety or concurrency risks associated with these changes involve use-after-free, NULL pointer dereferences, or lock/barrier ordering, which are fully covered by standard KASAN and LOCKDEP. Therefore, a dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

Instruction:
You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.

CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
  a comprehensive suite of debugging tools and sanitizers, including KASAN
  (out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
  (locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
  or page allocations) and kernel-to-user memory info-leaks.

Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.

Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.

Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
   attributes, ioctl output arguments, socket options, or BPF buffers) where fields
   or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
   or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
   could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
   uninitialized bytes of existing buffers.

Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
  dereferences, locking deadlocks, or use-after-free bugs (these are already caught
  by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
  or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.

Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.


Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.

Prompt:
Target architecture: amd64

For your convenience, here is the diff of the changes:
commit 30c2ad39819c25e745543e6ae191d461a2c8251e
Author: syz-cluster <triage@syzkaller.com>
Date:   Fri Oct 2 08:23:10 2026 +0000

    syz-cluster: applied patch under review

diff --git a/drivers/block/ublk_drv.c b/drivers/block/ublk_drv.c
index 66eb55e7162e5..f57d544c1da23 100644
--- a/drivers/block/ublk_drv.c
+++ b/drivers/block/ublk_drv.c
@@ -321,6 +321,8 @@ struct ublk_device {
 #define UB_STATE_OPEN		0
 #define UB_STATE_USED		1
 #define UB_STATE_DELETED	2
+/* STOP_DEV canceled the server's commands, until its release */
+#define UB_STATE_STOPPING	3
 	unsigned long		state;
 	int			ub_number;
 
@@ -334,6 +336,13 @@ struct ublk_device {
 	u16			nr_queue_ready;
 	bool 			unprivileged_daemons;
 	struct mutex cancel_mutex;
+	/* the open /dev/ublkcN, protected by cancel_mutex */
+	struct file *ch_file;
+	/*
+	 * A cancel started in this FETCH round. Set by ublk_set_canceling(),
+	 * cleared only by ublk_reset_ch_dev() when a new round starts. While
+	 * it is set, no queue clears its ->canceling.
+	 */
 	bool canceling;
 	pid_t 	ublksrv_tgid;
 	struct delayed_work	exit_work;
@@ -810,6 +819,8 @@ ublk_batch_alloc_fcmd(struct io_uring_cmd *cmd)
 	if (fcmd) {
 		fcmd->cmd = cmd;
 		fcmd->buf_group = READ_ONCE(cmd->sqe->buf_index);
+		/* a cancel may look at it before it is linked */
+		INIT_LIST_HEAD(&fcmd->node);
 	}
 	return fcmd;
 }
@@ -2399,6 +2410,9 @@ static int ublk_ch_open(struct inode *inode, struct file *filp)
 		return -EBUSY;
 	filp->private_data = ub;
 	ub->ublksrv_tgid = current->tgid;
+	mutex_lock(&ub->cancel_mutex);
+	ub->ch_file = filp;
+	mutex_unlock(&ub->cancel_mutex);
 	return 0;
 }
 
@@ -2415,11 +2429,17 @@ static void ublk_reset_ch_dev(struct ublk_device *ub)
 		spin_unlock(&ubq->cancel_lock);
 	}
 
+	/* a new FETCH round starts, the queues stay canceling until ready */
+	mutex_lock(&ub->cancel_mutex);
+	ub->canceling = false;
+	mutex_unlock(&ub->cancel_mutex);
+
 	/* set to NULL, otherwise new tasks cannot mmap io_cmd_buf */
 	ub->mm = NULL;
 	ub->nr_queue_ready = 0;
 	ub->unprivileged_daemons = false;
 	ub->ublksrv_tgid = -1;
+	clear_bit(UB_STATE_STOPPING, &ub->state);
 }
 
 static struct gendisk *ublk_get_disk(struct ublk_device *ub)
@@ -2539,12 +2559,15 @@ static void ublk_ch_release_work_fn(struct work_struct *work)
 	}
 
 	/*
-	 * disk isn't attached yet, either device isn't live, or it has
-	 * been removed already, so we needn't to do anything
+	 * No disk: the device isn't live, or it has been removed already.
+	 * There are no requests to abort, but the round still has to be
+	 * reset, so that a new server can fetch and start the device.
 	 */
 	disk = ublk_get_disk(ub);
-	if (!disk)
-		goto out;
+	if (!disk) {
+		mutex_lock(&ub->mutex);
+		goto reset;
+	}
 
 	/*
 	 * All uring_cmd are done now, so abort any request outstanding to
@@ -2576,7 +2599,7 @@ static void ublk_ch_release_work_fn(struct work_struct *work)
 
 	/* double check after grabbing lock */
 	if (!ub->ub_disk)
-		goto unlock;
+		goto reset;
 
 	/*
 	 * Transition the device to the nosrv state. What exactly this
@@ -2602,13 +2625,14 @@ static void ublk_ch_release_work_fn(struct work_struct *work)
 				WRITE_ONCE(ublk_get_queue(ub, i)->fail_io, true);
 		}
 	}
-unlock:
+reset:
+	/*
+	 * All uring_cmd has been done now, reset device & ubq. Under
+	 * ub->mutex, so START_DEV sees the round either ready or reset.
+	 */
+	ublk_reset_ch_dev(ub);
 	mutex_unlock(&ub->mutex);
 	ublk_put_disk(disk);
-
-	/* all uring_cmd has been done now, reset device & ubq */
-	ublk_reset_ch_dev(ub);
-out:
 	clear_bit(UB_STATE_OPEN, &ub->state);
 
 	/* put the reference grabbed in ublk_ch_release() */
@@ -2619,6 +2643,9 @@ static int ublk_ch_release(struct inode *inode, struct file *filp)
 {
 	struct ublk_device *ub = filp->private_data;
 
+	mutex_lock(&ub->cancel_mutex);
+	ub->ch_file = NULL;
+	mutex_unlock(&ub->cancel_mutex);
 	/*
 	 * Grab ublk device reference, so it won't be gone until we are
 	 * really released from work function.
@@ -2790,7 +2817,8 @@ static void ublk_cancel_cmd(struct ublk_queue *ubq, u16 tag,
 	done = !!(io->flags & UBLK_IO_FLAG_CANCELED);
 	if (!done) {
 		io->flags |= UBLK_IO_FLAG_CANCELED;
-		cmd = io->cmd;
+		/* dependency ordered against smp_wmb() in ublk_prep_cancel() */
+		cmd = READ_ONCE(io->cmd);
 		io->cmd = NULL;
 	}
 	spin_unlock(&ubq->cancel_lock);
@@ -2879,6 +2907,7 @@ static void ublk_uring_cmd_cancel_fn(struct io_uring_cmd *cmd,
 {
 	struct ublk_uring_cmd_pdu *pdu = ublk_get_uring_cmd_pdu(cmd);
 	struct ublk_queue *ubq = pdu->ubq;
+	struct io_uring_cmd *cur;
 	struct task_struct *task;
 	struct ublk_io *io;
 
@@ -2895,7 +2924,9 @@ static void ublk_uring_cmd_cancel_fn(struct io_uring_cmd *cmd,
 
 	ublk_start_cancel(ubq->dev);
 
-	WARN_ON_ONCE(io->cmd != cmd);
+	/* NULL if STOP_DEV's cancel took it meanwhile */
+	cur = READ_ONCE(io->cmd);
+	WARN_ON_ONCE(cur && cur != cmd);
 	ublk_cancel_cmd(ubq, pdu->tag, issue_flags);
 }
 
@@ -3006,13 +3037,54 @@ static void ublk_stop_dev_unlocked(struct ublk_device *ub)
 	put_disk(disk);
 }
 
+static struct file *ublk_get_ch_file(struct ublk_device *ub)
+{
+	struct file *file;
+
+	mutex_lock(&ub->cancel_mutex);
+	file = ub->ch_file;
+	if (file && !file_ref_get(&file->f_ref))
+		file = NULL;
+	mutex_unlock(&ub->cancel_mutex);
+	return file;
+}
+
 static void ublk_stop_dev(struct ublk_device *ub)
 {
+	struct file *file;
+
+	/*
+	 * FETCH, PREP and START_DEV take ub->mutex. If a server has
+	 * /dev/ublkcN open, set STOPPING, which turns it away, and hold a
+	 * reference on the file: the server's release, whose reset clears
+	 * STOPPING and lets a new server attach, can't run before we drop it.
+	 * So the cancel can run after the unlock and only meets this server's
+	 * commands; it has to: io_uring_cmd_done() may take uring_lock, under
+	 * which FETCH takes ub->mutex.
+	 */
 	mutex_lock(&ub->mutex);
 	ublk_stop_dev_unlocked(ub);
-	mutex_unlock(&ub->mutex);
 	cancel_work_sync(&ub->partition_scan_work);
+	file = ublk_get_ch_file(ub);
+	if (file) {
+		/*
+		 * The server has to close /dev/ublkcN before this device can
+		 * be started again: the reset in its release clears STOPPING.
+		 */
+		set_bit(UB_STATE_STOPPING, &ub->state);
+		/* for wake_up_var() below, see wake_up_bit() */
+		smp_mb__after_atomic();
+	}
+	mutex_unlock(&ub->mutex);
+
+	/* no server: nothing to cancel */
+	if (!file)
+		return;
+
+	/* wake a START_DEV waiting for the device to get ready */
+	wake_up_var(&ub->nr_queue_ready);
 	ublk_cancel_dev(ub);
+	__fput_sync(file);
 }
 
 static void ublk_reset_io_flags(struct ublk_queue *ubq, struct ublk_io *io)
@@ -3024,11 +3096,19 @@ static void ublk_reset_io_flags(struct ublk_queue *ubq, struct ublk_io *io)
 }
 
 /* reset per-queue io flags */
-static void ublk_queue_reset_io_flags(struct ublk_queue *ubq)
+static void ublk_queue_reset_io_flags(struct ublk_device *ub,
+				      struct ublk_queue *ubq)
 {
-	spin_lock(&ubq->cancel_lock);
-	ubq->canceling = false;
-	spin_unlock(&ubq->cancel_lock);
+	/*
+	 * A cancel in this FETCH round took a command which still counts as
+	 * ready, so the queue has to stay canceling. ub->canceling is set
+	 * under cancel_mutex before any command is taken: either we see it
+	 * here, or the cancel marks this queue again later.
+	 */
+	mutex_lock(&ub->cancel_mutex);
+	if (!ub->canceling)
+		ubq->canceling = false;
+	mutex_unlock(&ub->cancel_mutex);
 	ubq->fail_io = false;
 	ubq->force_abort = false;
 }
@@ -3051,24 +3131,17 @@ static void ublk_mark_io_ready(struct ublk_device *ub, u16 q_id,
 		ub->nr_queue_ready++;
 
 		/*
-		 * Reset queue flags as soon as this queue is ready.
-		 * This clears the canceling flag, allowing batch FETCH commands
-		 * to succeed during recovery without waiting for all queues.
+		 * Reset queue flags as soon as this queue is ready. Unless
+		 * this round saw a cancel, this clears the canceling flag,
+		 * allowing batch FETCH commands to succeed during recovery
+		 * without waiting for all queues.
 		 */
-		ublk_queue_reset_io_flags(ubq);
+		ublk_queue_reset_io_flags(ub, ubq);
 	}
 
-	/* Check if all queues are ready */
-	if (ublk_dev_ready(ub)) {
-		/*
-		 * All queues ready - clear device-level canceling flag
-		 * and wake ublk_dev_ready() waiters.
-		 */
-		mutex_lock(&ub->cancel_mutex);
-		ub->canceling = false;
-		mutex_unlock(&ub->cancel_mutex);
+	/* All queues ready - wake ublk_dev_ready() waiters */
+	if (ublk_dev_ready(ub))
 		wake_up_var(&ub->nr_queue_ready);
-	}
 }
 
 static inline int ublk_check_cmd_op(u32 cmd_op)
@@ -3152,6 +3225,12 @@ ublk_fill_io_cmd(struct ublk_io *io, struct io_uring_cmd *cmd)
 	return req;
 }
 
+/*
+ * Call before ublk_fill_io_cmd() publishes @cmd in io->cmd: a control-path
+ * cancel may complete any command found there, and io_uring_cmd_done() only
+ * takes it off the cancelable list if it is marked already. The handlers
+ * hold uring_lock, so marking takes no lock.
+ */
 static inline void ublk_prep_cancel(struct io_uring_cmd *cmd,
 				    unsigned int issue_flags,
 				    struct ublk_queue *ubq, u16 tag)
@@ -3165,6 +3244,8 @@ static inline void ublk_prep_cancel(struct io_uring_cmd *cmd,
 	pdu->ubq = ubq;
 	pdu->tag = tag;
 	io_uring_cmd_mark_cancelable(cmd, issue_flags);
+	/* pairs with the cancel loading cmd from io->cmd, then cmd->flags */
+	smp_wmb();
 }
 
 static void ublk_io_release(void *priv)
@@ -3269,6 +3350,9 @@ static int ublk_check_fetch_buf(const struct ublk_device *ub, __u64 buf_addr)
 static int __ublk_fetch(struct io_uring_cmd *cmd, struct ublk_device *ub,
 			struct ublk_io *io, u16 q_id)
 {
+	if (test_bit(UB_STATE_STOPPING, &ub->state))
+		return UBLK_IO_RES_ABORT;
+
 	/* UBLK_IO_FETCH_REQ is only allowed before dev is setup */
 	if (ublk_dev_ready(ub))
 		return -EBUSY;
@@ -3412,11 +3496,11 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,
 		ret = ublk_check_fetch_buf(ub, addr);
 		if (ret)
 			goto out;
+		/* before ublk_fetch() publishes io->cmd, see ublk_prep_cancel() */
+		ublk_prep_cancel(cmd, issue_flags, ubq, tag);
 		ret = ublk_fetch(cmd, ub, io, addr, q_id);
 		if (ret)
-			goto out;
-
-		ublk_prep_cancel(cmd, issue_flags, ubq, tag);
+			goto out_done;
 		return -EIOCBQUEUED;
 	}
 
@@ -3460,6 +3544,7 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,
 		if (ret)
 			goto out;
 		io->res = result;
+		ublk_prep_cancel(cmd, issue_flags, ubq, tag);
 		req = ublk_fill_io_cmd(io, cmd);
 		ublk_apply_io_buf(ub, io, cmd, addr, &auto_buf, &buf_idx);
 		if (buf_idx != UBLK_INVALID_BUF_IDX)
@@ -3478,19 +3563,24 @@ static int ublk_ch_uring_cmd_local(struct io_uring_cmd *cmd,
 		 * uring_cmd active first and prepare for handling new requeued
 		 * request
 		 */
+		ublk_prep_cancel(cmd, issue_flags, ubq, tag);
 		req = ublk_fill_io_cmd(io, cmd);
 		io->buf.addr = addr;
 		if (likely(ublk_get_data(ubq, io, req))) {
 			__ublk_prep_compl_io_cmd(io, req);
-			return UBLK_IO_RES_OK;
+			ret = UBLK_IO_RES_OK;
+			goto out_done;
 		}
 		break;
 	default:
 		goto out;
 	}
-	ublk_prep_cancel(cmd, issue_flags, ubq, tag);
 	return -EIOCBQUEUED;
 
+ out_done:
+	/* marked cancelable: complete through io_uring_cmd_done() */
+	io_uring_cmd_done(cmd, ret, issue_flags);
+	return -EIOCBQUEUED;
  out:
 	pr_devel("%s: complete: cmd op %d, tag %d ret %x io_flags %x\n",
 			__func__, cmd_op, tag, ret, io ? io->flags : 0);
@@ -3889,6 +3979,15 @@ static int ublk_batch_attach(struct ublk_queue *ubq,
 	bool free = false;
 	struct ublk_uring_cmd_pdu *pdu = ublk_get_uring_cmd_pdu(data->cmd);
 
+	/*
+	 * Mark it cancelable before linking it into fcmd_head, where a cancel
+	 * from the control path can take and complete it: see
+	 * ublk_prep_cancel(). evts_lock orders the mark before the link.
+	 */
+	pdu->ubq = ubq;
+	pdu->fcmd = fcmd;
+	io_uring_cmd_mark_cancelable(fcmd->cmd, data->issue_flags);
+
 	spin_lock(&ubq->evts_lock);
 	if (unlikely(ubq->force_abort || ubq->canceling)) {
 		free = true;
@@ -3899,14 +3998,12 @@ static int ublk_batch_attach(struct ublk_queue *ubq,
 	spin_unlock(&ubq->evts_lock);
 
 	if (unlikely(free)) {
+		/* off the cancelable list first, then nothing can see fcmd */
+		io_uring_cmd_done(data->cmd, -ENODEV, data->issue_flags);
 		ublk_batch_free_fcmd(fcmd);
-		return -ENODEV;
+		return -EIOCBQUEUED;
 	}
 
-	pdu->ubq = ubq;
-	pdu->fcmd = fcmd;
-	io_uring_cmd_mark_cancelable(fcmd->cmd, data->issue_flags);
-
 	if (!new_fcmd)
 		goto out;
 
@@ -3914,9 +4011,12 @@ static int ublk_batch_attach(struct ublk_queue *ubq,
 	 * If the two fetch commands are originated from same io_ring_ctx,
 	 * run batch dispatch directly. Otherwise, schedule task work for
 	 * doing it.
+	 *
+	 * Use data->cmd, not fcmd->cmd: once fcmd is linked and not active,
+	 * a cancel from the control path may complete and free it.
 	 */
 	if (io_uring_cmd_ctx_handle(new_fcmd->cmd) ==
-			io_uring_cmd_ctx_handle(fcmd->cmd)) {
+			io_uring_cmd_ctx_handle(data->cmd)) {
 		data->cmd = new_fcmd->cmd;
 		ublk_batch_dispatch(ubq, data, new_fcmd);
 	} else {
@@ -4431,21 +4531,27 @@ static bool ublk_validate_user_pid(struct ublk_device *ub, pid_t ublksrv_pid)
 	return ub->ublksrv_tgid == ublksrv_pid;
 }
 
+static bool ublk_dev_ready_or_stopping(const struct ublk_device *ub)
+{
+	return ublk_dev_ready(ub) || test_bit(UB_STATE_STOPPING, &ub->state);
+}
+
 /*
- * Wait until all queues have fetched their I/O commands, and return with
- * ub->mutex held and readiness guaranteed: then every queue's ->canceling
- * is cleared. Ready may regress between wakeup and mutex_lock() (F_BATCH
- * UNPREP, daemon death), so re-check it under the mutex and wait again.
+ * Wait until all queues have fetched their I/O commands, or STOP_DEV set
+ * UB_STATE_STOPPING, and return with ub->mutex held. The queues stay
+ * canceling if this round saw a cancel, see ublk_queue_reset_io_flags().
+ * Ready may regress between wakeup and mutex_lock() (F_BATCH UNPREP,
+ * daemon death), so re-check it under the mutex and wait again.
  */
 static int ublk_wait_dev_ready_and_lock(struct ublk_device *ub)
 {
 	while (true) {
 		if (wait_var_event_interruptible(&ub->nr_queue_ready,
-						 ublk_dev_ready(ub)))
+						 ublk_dev_ready_or_stopping(ub)))
 			return -EINTR;
 
 		mutex_lock(&ub->mutex);
-		if (ublk_dev_ready(ub))
+		if (ublk_dev_ready_or_stopping(ub))
 			return 0;
 		mutex_unlock(&ub->mutex);
 	}
@@ -4545,6 +4651,10 @@ static int ublk_ctrl_start_dev(struct ublk_device *ub,
 		ret = -EEXIST;
 		goto out_unlock;
 	}
+	if (test_bit(UB_STATE_STOPPING, &ub->state)) {
+		ret = -EBUSY;
+		goto out_unlock;
+	}
 
 	disk = blk_mq_alloc_disk(&ub->tag_set, &lim, NULL);
 	if (IS_ERR(disk)) {
diff --git a/tools/testing/selftests/ublk/.gitignore b/tools/testing/selftests/ublk/.gitignore
index e17bd28f27e04..d8e93fef7fcd3 100644
--- a/tools/testing/selftests/ublk/.gitignore
+++ b/tools/testing/selftests/ublk/.gitignore
@@ -3,3 +3,4 @@
 /tools
 kublk
 metadata_size
+ublk_cancel_ready
diff --git a/tools/testing/selftests/ublk/Makefile b/tools/testing/selftests/ublk/Makefile
index 37883e9d50ec3..fc64b8f028332 100644
--- a/tools/testing/selftests/ublk/Makefile
+++ b/tools/testing/selftests/ublk/Makefile
@@ -19,6 +19,7 @@ TEST_PROGS += test_generic_12.sh
 TEST_PROGS += test_generic_13.sh
 TEST_PROGS += test_generic_16.sh
 TEST_PROGS += test_generic_17.sh
+TEST_PROGS += test_generic_18.sh
 
 TEST_PROGS += test_batch_01.sh
 TEST_PROGS += test_batch_02.sh
@@ -76,13 +77,14 @@ TEST_FILES := settings
 TEST_FILES += test_common.sh
 TEST_FILES += trace
 
-TEST_GEN_PROGS_EXTENDED = kublk metadata_size
-STANDALONE_UTILS := metadata_size.c
+TEST_GEN_PROGS_EXTENDED = kublk metadata_size ublk_cancel_ready
+STANDALONE_UTILS := metadata_size.c ublk_cancel_ready.c
 
 LOCAL_HDRS += $(wildcard *.h)
 include ../lib.mk
 
 $(OUTPUT)/kublk: $(filter-out $(STANDALONE_UTILS),$(wildcard *.c))
+$(OUTPUT)/ublk_cancel_ready: ublk_cancel_ready.c ctrl.c
 
 check:
 	shellcheck -x -f gcc *.sh
diff --git a/tools/testing/selftests/ublk/ctrl.c b/tools/testing/selftests/ublk/ctrl.c
new file mode 100644
index 0000000000000..53d54799f9097
--- /dev/null
+++ b/tools/testing/selftests/ublk/ctrl.c
@@ -0,0 +1,268 @@
+// SPDX-License-Identifier: GPL-2.0
+
+/* ublk control commands, shared by kublk and ublk_cancel_ready */
+
+#include "kublk.h"
+
+static void ublk_ctrl_init_cmd(struct ublk_dev *dev,
+		struct io_uring_sqe *sqe,
+		struct ublk_ctrl_cmd_data *data)
+{
+	struct ublksrv_ctrl_dev_info *info = &dev->dev_info;
+	struct ublksrv_ctrl_cmd *cmd = (struct ublksrv_ctrl_cmd *)ublk_get_sqe_cmd(sqe);
+
+	sqe->fd = dev->ctrl_fd;
+	sqe->opcode = IORING_OP_URING_CMD;
+	sqe->ioprio = 0;
+
+	if (data->flags & CTRL_CMD_HAS_BUF) {
+		cmd->addr = data->addr;
+		cmd->len = data->len;
+	}
+
+	if (data->flags & CTRL_CMD_HAS_DATA)
+		cmd->data[0] = data->data[0];
+
+	cmd->dev_id = info->dev_id;
+	cmd->queue_id = -1;
+
+	ublk_set_sqe_cmd_op(sqe, data->cmd_op);
+
+	io_uring_sqe_set_data(sqe, cmd);
+}
+
+int __ublk_ctrl_cmd(struct ublk_dev *dev,
+		struct ublk_ctrl_cmd_data *data)
+{
+	struct io_uring_sqe *sqe;
+	struct io_uring_cqe *cqe;
+	int ret = -EINVAL;
+
+	sqe = io_uring_get_sqe(&dev->ring);
+	if (!sqe) {
+		ublk_err("%s: can't get sqe ret %d\n", __func__, ret);
+		return ret;
+	}
+
+	ublk_ctrl_init_cmd(dev, sqe, data);
+
+	ret = io_uring_submit(&dev->ring);
+	if (ret < 0) {
+		ublk_err("uring submit ret %d\n", ret);
+		return ret;
+	}
+
+	ret = io_uring_wait_cqe(&dev->ring, &cqe);
+	if (ret < 0) {
+		ublk_err("wait cqe: %s\n", strerror(-ret));
+		return ret;
+	}
+	io_uring_cqe_seen(&dev->ring, cqe);
+
+	return cqe->res;
+}
+
+void ublk_ctrl_deinit(struct ublk_dev *dev)
+{
+	io_uring_queue_exit(&dev->ring);
+	close(dev->ctrl_fd);
+	free(dev);
+}
+
+struct ublk_dev *ublk_ctrl_init(void)
+{
+	struct ublk_dev *dev = (struct ublk_dev *)calloc(1, sizeof(*dev));
+	struct ublksrv_ctrl_dev_info *info;
+	int ret;
+
+	if (!dev)
+		return NULL;
+	info = &dev->dev_info;
+	dev->ctrl_fd = open(CTRL_DEV, O_RDWR);
+	if (dev->ctrl_fd < 0) {
+		free(dev);
+		return NULL;
+	}
+
+	info->max_io_buf_bytes = UBLK_IO_MAX_BYTES;
+
+	ret = ublk_setup_ring(&dev->ring, UBLK_CTRL_RING_DEPTH,
+			UBLK_CTRL_RING_DEPTH, IORING_SETUP_SQE128);
+	if (ret < 0) {
+		ublk_err("queue_init: %s\n", strerror(-ret));
+		close(dev->ctrl_fd);
+		free(dev);
+		return NULL;
+	}
+	dev->nr_fds = 1;
+
+	return dev;
+}
+
+int ublk_ctrl_stop_dev(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_STOP_DEV,
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_try_stop_dev(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_TRY_STOP_DEV,
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_start_dev(struct ublk_dev *dev,
+		int daemon_pid)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_START_DEV,
+		.flags	= CTRL_CMD_HAS_DATA,
+	};
+
+	dev->dev_info.ublksrv_pid = data.data[0] = daemon_pid;
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_start_user_recovery(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_START_USER_RECOVERY,
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_END_USER_RECOVERY,
+		.flags	= CTRL_CMD_HAS_DATA,
+	};
+
+	dev->dev_info.ublksrv_pid = data.data[0] = daemon_pid;
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_add_dev(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_ADD_DEV,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64) (uintptr_t) &dev->dev_info,
+		.len = sizeof(struct ublksrv_ctrl_dev_info),
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_del_dev(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op = UBLK_U_CMD_DEL_DEV,
+		.flags = 0,
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_get_info(struct ublk_dev *dev)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_GET_DEV_INFO,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64) (uintptr_t) &dev->dev_info,
+		.len = sizeof(struct ublksrv_ctrl_dev_info),
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_set_params(struct ublk_dev *dev,
+		struct ublk_params *params)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_SET_PARAMS,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64) (uintptr_t) params,
+		.len = sizeof(*params),
+	};
+	params->len = sizeof(*params);
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_get_params(struct ublk_dev *dev,
+		struct ublk_params *params)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_GET_PARAMS,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64)params,
+		.len = sizeof(*params),
+	};
+
+	params->len = sizeof(*params);
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_get_features(struct ublk_dev *dev,
+		__u64 *features)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_GET_FEATURES,
+		.flags	= CTRL_CMD_HAS_BUF,
+		.addr = (__u64) (uintptr_t) features,
+		.len = sizeof(*features),
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_update_size(struct ublk_dev *dev,
+		__u64 nr_sects)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_UPDATE_SIZE,
+		.flags	= CTRL_CMD_HAS_DATA,
+	};
+
+	data.data[0] = nr_sects;
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_quiesce_dev(struct ublk_dev *dev, unsigned int timeout_ms)
+{
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op	= UBLK_U_CMD_QUIESCE_DEV,
+		.flags	= CTRL_CMD_HAS_DATA,
+	};
+
+	data.data[0] = timeout_ms;
+	return __ublk_ctrl_cmd(dev, &data);
+}
+
+int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,
+		      __u32 flags)
+{
+	struct ublk_shmem_buf_reg buf_reg = {
+		.addr = (unsigned long)addr,
+		.len = size,
+		.flags = flags,
+	};
+	struct ublk_ctrl_cmd_data data = {
+		.cmd_op = UBLK_U_CMD_REG_BUF,
+		.flags = CTRL_CMD_HAS_BUF,
+		.addr = (unsigned long)&buf_reg,
+		.len = sizeof(buf_reg),
+	};
+
+	return __ublk_ctrl_cmd(dev, &data);
+}
diff --git a/tools/testing/selftests/ublk/kublk.c b/tools/testing/selftests/ublk/kublk.c
index 2400b46157664..15ba060ce7a3c 100644
--- a/tools/testing/selftests/ublk/kublk.c
+++ b/tools/testing/selftests/ublk/kublk.c
@@ -37,203 +37,6 @@ static const struct ublk_tgt_ops *ublk_find_tgt(const char *name)
 	return NULL;
 }
 
-static inline int ublk_setup_ring(struct io_uring *r, int depth,
-		int cq_depth, unsigned flags)
-{
-	struct io_uring_params p;
-
-	memset(&p, 0, sizeof(p));
-	p.flags = flags | IORING_SETUP_CQSIZE;
-	p.cq_entries = cq_depth;
-
-	return io_uring_queue_init_params(depth, r, &p);
-}
-
-static void ublk_ctrl_init_cmd(struct ublk_dev *dev,
-		struct io_uring_sqe *sqe,
-		struct ublk_ctrl_cmd_data *data)
-{
-	struct ublksrv_ctrl_dev_info *info = &dev->dev_info;
-	struct ublksrv_ctrl_cmd *cmd = (struct ublksrv_ctrl_cmd *)ublk_get_sqe_cmd(sqe);
-
-	sqe->fd = dev->ctrl_fd;
-	sqe->opcode = IORING_OP_URING_CMD;
-	sqe->ioprio = 0;
-
-	if (data->flags & CTRL_CMD_HAS_BUF) {
-		cmd->addr = data->addr;
-		cmd->len = data->len;
-	}
-
-	if (data->flags & CTRL_CMD_HAS_DATA)
-		cmd->data[0] = data->data[0];
-
-	cmd->dev_id = info->dev_id;
-	cmd->queue_id = -1;
-
-	ublk_set_sqe_cmd_op(sqe, data->cmd_op);
-
-	io_uring_sqe_set_data(sqe, cmd);
-}
-
-static int __ublk_ctrl_cmd(struct ublk_dev *dev,
-		struct ublk_ctrl_cmd_data *data)
-{
-	struct io_uring_sqe *sqe;
-	struct io_uring_cqe *cqe;
-	int ret = -EINVAL;
-
-	sqe = io_uring_get_sqe(&dev->ring);
-	if (!sqe) {
-		ublk_err("%s: can't get sqe ret %d\n", __func__, ret);
-		return ret;
-	}
-
-	ublk_ctrl_init_cmd(dev, sqe, data);
-
-	ret = io_uring_submit(&dev->ring);
-	if (ret < 0) {
-		ublk_err("uring submit ret %d\n", ret);
-		return ret;
-	}
-
-	ret = io_uring_wait_cqe(&dev->ring, &cqe);
-	if (ret < 0) {
-		ublk_err("wait cqe: %s\n", strerror(-ret));
-		return ret;
-	}
-	io_uring_cqe_seen(&dev->ring, cqe);
-
-	return cqe->res;
-}
-
-static int ublk_ctrl_stop_dev(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_STOP_DEV,
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_try_stop_dev(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_TRY_STOP_DEV,
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_start_dev(struct ublk_dev *dev,
-		int daemon_pid)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_START_DEV,
-		.flags	= CTRL_CMD_HAS_DATA,
-	};
-
-	dev->dev_info.ublksrv_pid = data.data[0] = daemon_pid;
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_start_user_recovery(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_START_USER_RECOVERY,
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_END_USER_RECOVERY,
-		.flags	= CTRL_CMD_HAS_DATA,
-	};
-
-	dev->dev_info.ublksrv_pid = data.data[0] = daemon_pid;
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_add_dev(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_ADD_DEV,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64) (uintptr_t) &dev->dev_info,
-		.len = sizeof(struct ublksrv_ctrl_dev_info),
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_del_dev(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op = UBLK_U_CMD_DEL_DEV,
-		.flags = 0,
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_get_info(struct ublk_dev *dev)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_GET_DEV_INFO,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64) (uintptr_t) &dev->dev_info,
-		.len = sizeof(struct ublksrv_ctrl_dev_info),
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_set_params(struct ublk_dev *dev,
-		struct ublk_params *params)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_SET_PARAMS,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64) (uintptr_t) params,
-		.len = sizeof(*params),
-	};
-	params->len = sizeof(*params);
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_get_params(struct ublk_dev *dev,
-		struct ublk_params *params)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_GET_PARAMS,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64)params,
-		.len = sizeof(*params),
-	};
-
-	params->len = sizeof(*params);
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_get_features(struct ublk_dev *dev,
-		__u64 *features)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_GET_FEATURES,
-		.flags	= CTRL_CMD_HAS_BUF,
-		.addr = (__u64) (uintptr_t) features,
-		.len = sizeof(*features),
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
 static int parse_param_types(const char *arg, __u32 *types)
 {
 	char buf[128], *save = NULL, *tok;
@@ -283,30 +86,6 @@ static void ublk_init_params_from_ctx(const struct dev_ctx *ctx,
 	};
 }
 
-static int ublk_ctrl_update_size(struct ublk_dev *dev,
-		__u64 nr_sects)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_UPDATE_SIZE,
-		.flags	= CTRL_CMD_HAS_DATA,
-	};
-
-	data.data[0] = nr_sects;
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
-static int ublk_ctrl_quiesce_dev(struct ublk_dev *dev,
-				 unsigned int timeout_ms)
-{
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op	= UBLK_U_CMD_QUIESCE_DEV,
-		.flags	= CTRL_CMD_HAS_DATA,
-	};
-
-	data.data[0] = timeout_ms;
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
 static const char *ublk_dev_state_desc(struct ublk_dev *dev)
 {
 	switch (dev->dev_info.state) {
@@ -426,38 +205,6 @@ static void ublk_ctrl_dump(struct ublk_dev *dev)
 	fflush(stdout);
 }
 
-static void ublk_ctrl_deinit(struct ublk_dev *dev)
-{
-	close(dev->ctrl_fd);
-	free(dev);
-}
-
-static struct ublk_dev *ublk_ctrl_init(void)
-{
-	struct ublk_dev *dev = (struct ublk_dev *)calloc(1, sizeof(*dev));
-	struct ublksrv_ctrl_dev_info *info = &dev->dev_info;
-	int ret;
-
-	dev->ctrl_fd = open(CTRL_DEV, O_RDWR);
-	if (dev->ctrl_fd < 0) {
-		free(dev);
-		return NULL;
-	}
-
-	info->max_io_buf_bytes = UBLK_IO_MAX_BYTES;
-
-	ret = ublk_setup_ring(&dev->ring, UBLK_CTRL_RING_DEPTH,
-			UBLK_CTRL_RING_DEPTH, IORING_SETUP_SQE128);
-	if (ret < 0) {
-		ublk_err("queue_init: %s\n", strerror(-ret));
-		free(dev);
-		return NULL;
-	}
-	dev->nr_fds = 1;
-
-	return dev;
-}
-
 static size_t __ublk_queue_cmd_buf_sz(const struct ublk_queue *q, __u16 depth)
 {
 	size_t size = depth * (size_t)q->io_desc_size;
@@ -1283,24 +1030,6 @@ static void ublk_shmem_unregister_all(void)
 	shmem_count = 0;
 }
 
-static int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,
-			     __u32 flags)
-{
-	struct ublk_shmem_buf_reg buf_reg = {
-		.addr = (unsigned long)addr,
-		.len = size,
-		.flags = flags,
-	};
-	struct ublk_ctrl_cmd_data data = {
-		.cmd_op = UBLK_U_CMD_REG_BUF,
-		.flags = CTRL_CMD_HAS_BUF,
-		.addr = (unsigned long)&buf_reg,
-		.len = sizeof(buf_reg),
-	};
-
-	return __ublk_ctrl_cmd(dev, &data);
-}
-
 /*
  * Handle one client connection: receive memfd, mmap it, register
  * the VA range with kernel, send back the assigned index.
diff --git a/tools/testing/selftests/ublk/kublk.h b/tools/testing/selftests/ublk/kublk.h
index d98f3d612d888..99b8ceff853cb 100644
--- a/tools/testing/selftests/ublk/kublk.h
+++ b/tools/testing/selftests/ublk/kublk.h
@@ -294,6 +294,38 @@ struct ublk_dev {
 
 extern int ublk_queue_io_cmd(struct ublk_thread *t, struct ublk_io *io);
 
+static inline int ublk_setup_ring(struct io_uring *r, int depth,
+		int cq_depth, unsigned int flags)
+{
+	struct io_uring_params p;
+
+	memset(&p, 0, sizeof(p));
+	p.flags = flags | IORING_SETUP_CQSIZE;
+	p.cq_entries = cq_depth;
+
+	return io_uring_queue_init_params(depth, r, &p);
+}
+
+/* ctrl.c: control commands */
+struct ublk_dev *ublk_ctrl_init(void);
+void ublk_ctrl_deinit(struct ublk_dev *dev);
+int __ublk_ctrl_cmd(struct ublk_dev *dev, struct ublk_ctrl_cmd_data *data);
+int ublk_ctrl_add_dev(struct ublk_dev *dev);
+int ublk_ctrl_del_dev(struct ublk_dev *dev);
+int ublk_ctrl_get_info(struct ublk_dev *dev);
+int ublk_ctrl_set_params(struct ublk_dev *dev, struct ublk_params *params);
+int ublk_ctrl_get_params(struct ublk_dev *dev, struct ublk_params *params);
+int ublk_ctrl_get_features(struct ublk_dev *dev, __u64 *features);
+int ublk_ctrl_start_dev(struct ublk_dev *dev, int daemon_pid);
+int ublk_ctrl_stop_dev(struct ublk_dev *dev);
+int ublk_ctrl_try_stop_dev(struct ublk_dev *dev);
+int ublk_ctrl_start_user_recovery(struct ublk_dev *dev);
+int ublk_ctrl_end_user_recovery(struct ublk_dev *dev, int daemon_pid);
+int ublk_ctrl_update_size(struct ublk_dev *dev, __u64 nr_sects);
+int ublk_ctrl_quiesce_dev(struct ublk_dev *dev, unsigned int timeout_ms);
+int ublk_ctrl_reg_buf(struct ublk_dev *dev, void *addr, size_t size,
+		      __u32 flags);
+
 static inline int __ublk_use_batch_io(__u64 flags)
 {
 	return flags & UBLK_F_BATCH_IO;
diff --git a/tools/testing/selftests/ublk/test_generic_18.sh b/tools/testing/selftests/ublk/test_generic_18.sh
new file mode 100755
index 0000000000000..223944dba149a
--- /dev/null
+++ b/tools/testing/selftests/ublk/test_generic_18.sh
@@ -0,0 +1,37 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+
+. "$(cd "$(dirname "$0")" && pwd)"/test_common.sh
+
+ERR_CODE=0
+CANCEL_PROG="$(_ublk_test_top_dir)/ublk_cancel_ready"
+
+_prep_test "generic" "start device over canceled io commands"
+
+# the modes are described in ublk_cancel_ready.c
+for mode in stop_start partial_fetch recovery stop_restart stop_attached \
+		stop_live_restart race_start race_fetch race_async_fetch; do
+	dmesg_before=$(dmesg | wc -l)
+	timeout 60 "$CANCEL_PROG" "$mode" > "$UBLK_TMP" 2>&1
+	res=$?
+	msg=""
+
+	if dmesg | tail -n +"$((dmesg_before + 1))" | \
+			grep -q -e "BUG:" -e "Oops" -e "WARNING:"; then
+		msg="$mode: kernel oops/warning"
+		ERR_CODE=255
+	elif [ "$res" -eq "$UBLK_SKIP_CODE" ]; then
+		[ "$ERR_CODE" -eq 0 ] && ERR_CODE=$UBLK_SKIP_CODE
+	elif [ "$res" -ne 0 ]; then
+		msg="$mode: failed ($res)"
+		ERR_CODE=255
+	fi
+	# the output once: on failure, or always when not quiet
+	[ -n "$msg" ] && echo "$msg"
+	if [ -n "$msg" ] || [ "$UBLK_TEST_QUIET" -eq 0 ]; then
+		cat "$UBLK_TMP"
+	fi
+done
+
+_cleanup_test
+_show_result $TID $ERR_CODE
diff --git a/tools/testing/selftests/ublk/ublk_cancel_ready.c b/tools/testing/selftests/ublk/ublk_cancel_ready.c
new file mode 100644
index 0000000000000..0ed0d01fa74ad
--- /dev/null
+++ b/tools/testing/selftests/ublk/ublk_cancel_ready.c
@@ -0,0 +1,984 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Bring a ublk device live while some of its fetched io commands are
+ * canceled.
+ *
+ * A cancel completes a fetched command and clears io->cmd, but the io
+ * still counts as ready. Only ubq->canceling keeps ublk_queue_rq() away
+ * from the NULL io->cmd. Each mode below loses ->canceling in another way:
+ *
+ * stop_start:    fetch every tag, STOP_DEV before START_DEV, START_DEV.
+ *                STOP_DEV cancels the commands of the attached server,
+ *                so START_DEV must get -EBUSY until that server is gone.
+ *                An unfixed kernel starts the device over them.
+ *
+ * partial_fetch: task A fetches tags 0..depth-2 and dies, so its
+ *                commands are canceled. Task B fetches the last tag, and
+ *                an unfixed kernel clears ->canceling when the queue gets
+ *                ready. /dev/ublkcN stays open all the time, so
+ *                ublk_ch_release() never resets the queue.
+ *
+ * recovery:      a UBLK_F_USER_RECOVERY device with two queues loses its
+ *                server. During recovery task Q0 fetches queue 0, which
+ *                clears q0->canceling, then dies. An unfixed kernel still
+ *                has ub->canceling set because queue 1 is not ready, so
+ *                ublk_start_cancel() does not mark queue 0 again. Reads
+ *                are issued on a CPU mapped to queue 0.
+ *
+ * In partial_fetch and recovery, START_DEV / END_USER_RECOVERY may refuse
+ * with -ENODEV, or bring the device live with its queue still canceling:
+ * then every read has to complete, with -EIO. An unfixed kernel oopses
+ * in ublk_queue_cmd() on a NULL io->cmd.
+ *
+ * stop_restart:  control mode. STOP_DEV on a new device, before any
+ *                server opened it, takes no command. A server started
+ *                afterwards has to start the device and serve I/O.
+ *
+ * stop_attached: STOP_DEV on a device whose server opened it but fetched
+ *                nothing stops that server too: FETCH gets ABORT and
+ *                START_DEV -EBUSY, until a new server opens the device.
+ *
+ * stop_live_restart: STOP_DEV on a live device; once its server is gone,
+ *                a new server has to fetch and start it again. An unfixed
+ *                ublk_ch_release() skips the reset once the disk is gone,
+ *                so the new FETCH gets -EBUSY.
+ *
+ * race_start:    STOP_DEV and START_DEV at the same time on a device whose
+ *                server fetched every tag. START_DEV either wins, and
+ *                reads complete (they fail once STOP_DEV removes the
+ *                disk), or gets -EBUSY. An unfixed kernel cancels the
+ *                commands of the live disk.
+ *
+ * race_fetch:    STOP_DEV while a server opens the device and fetches,
+ *                then START_DEV. If STOP_DEV came before the open, it must
+ *                not take any command, so START_DEV works and every read
+ *                succeeds; otherwise START_DEV gets -EBUSY. An unfixed
+ *                kernel takes commands fetched after its unlock and goes
+ *                live over them.
+ *
+ * race_async_fetch: FETCH with IOSQE_ASYNC while STOP_DEV cancels,
+ *                then close the ring. An unfixed FETCH marks its command
+ *                cancelable only after publishing it and dropping
+ *                ub->mutex; a cancel in between completes a command which
+ *                is then put on io_uring's cancelable list. Closing the
+ *                ring walks that list: KASAN reports a use after free.
+ *                Needs a KASAN kernel to see the bug; otherwise it must
+ *                just not crash.
+ *
+ * A dying task is a child process which fetches and then calls exec().
+ * exec() cancels the task's uring_cmds before it returns, so the cancel
+ * is done once the child is reaped. A thread exit does not cancel them,
+ * and closing the ring cancels them later, from io_ring_exit_work().
+ */
+#include <sched.h>
+
+#include "kublk.h"
+#include "../kselftest.h"
+
+#define NR_QUEUES	2
+#define DEPTH		4
+#define BUF_SIZE	(64 << 10)
+#define DEV_SECTORS	(64 << 11)	/* 64MB */
+#define NR_READS	(DEPTH * 2)
+#define SERVE_DELAY_US	200000
+#define RACE_LOOPS	50
+#define ASYNC_LOOPS	300
+
+/* one control handle per thread, the race modes send commands from two */
+static __thread struct ublk_dev *ctrl_dev;
+static int cdev_fd = -1;
+static int dev_id = -1;
+static int nr_queues;
+static void *bufs[NR_QUEUES][DEPTH];
+static const struct ublksrv_io_desc *iods[NR_QUEUES];
+static size_t iods_len;
+
+static struct io_uring_sqe *get_sqe(struct io_uring *ring)
+{
+	struct io_uring_sqe *sqe = io_uring_get_sqe(ring);
+
+	if (!sqe) {
+		fprintf(stderr, "out of sqes\n");
+		exit(KSFT_FAIL);
+	}
+	return sqe;
+}
+
+static struct ublk_dev *ctrl(void)
+{
+	if (!ctrl_dev) {
+		ctrl_dev = ublk_ctrl_init();
+		if (!ctrl_dev) {
+			fprintf(stderr, "ublk_ctrl_init failed\n");
+			exit(KSFT_FAIL);
+		}
+	}
+	ctrl_dev->dev_info.dev_id = dev_id;
+	return ctrl_dev;
+}
+
+static void ctrl_put(void)
+{
+	if (ctrl_dev)
+		ublk_ctrl_deinit(ctrl_dev);
+	ctrl_dev = NULL;
+}
+
+static int dev_state(void)
+{
+	struct ublk_dev *dev = ctrl();
+	int ret = ublk_ctrl_get_info(dev);
+
+	return ret ? ret : dev->dev_info.state;
+}
+
+static int open_cdev(void)
+{
+	size_t max_len = UBLK_MAX_QUEUE_DEPTH * sizeof(struct ublksrv_io_desc);
+	int pg = getpagesize();
+	char path[64];
+
+	snprintf(path, sizeof(path), "/dev/ublkc%d", dev_id);
+	for (int i = 0; i < 100 && cdev_fd < 0; i++) {
+		cdev_fd = open(path, O_RDWR);
+		if (cdev_fd < 0)
+			usleep(50000);
+	}
+	if (cdev_fd < 0)
+		return -errno;
+
+	/* queue q's descriptors start at q * the size for the max depth */
+	max_len = (max_len + pg - 1) & ~(size_t)(pg - 1);
+	iods_len = (DEPTH * sizeof(struct ublksrv_io_desc) + pg - 1) &
+		~(size_t)(pg - 1);
+	for (int q = 0; q < nr_queues; q++) {
+		void *p = mmap(NULL, iods_len, PROT_READ,
+			       MAP_SHARED | MAP_POPULATE, cdev_fd,
+			       UBLKSRV_CMD_BUF_OFFSET + q * max_len);
+
+		if (p == MAP_FAILED)
+			return -errno;
+		iods[q] = p;
+	}
+	return 0;
+}
+
+/* the last reference to /dev/ublkcN runs ublk_ch_release() */
+static void close_cdev(void)
+{
+	for (int q = 0; q < nr_queues; q++) {
+		if (iods[q])
+			munmap((void *)iods[q], iods_len);
+		iods[q] = NULL;
+	}
+	if (cdev_fd >= 0)
+		close(cdev_fd);
+	cdev_fd = -1;
+}
+
+/* ADD_DEV and SET_PARAMS, without opening /dev/ublkcN */
+static int add_dev_noopen(int queues, __u64 flags)
+{
+	struct ublk_dev *dev = ctrl();
+	struct ublksrv_ctrl_dev_info info = {
+		.nr_hw_queues	= queues,
+		.queue_depth	= DEPTH,
+		.max_io_buf_bytes = BUF_SIZE,
+		.dev_id		= -1,
+		.flags		= UBLK_F_NO_AUTO_PART_SCAN | flags,
+	};
+	struct ublk_params p = {
+		.types	= UBLK_PARAM_TYPE_BASIC,
+		.basic	= {
+			.logical_bs_shift	= 9,
+			.physical_bs_shift	= 12,
+			.io_opt_shift		= 12,
+			.io_min_shift		= 9,
+			.max_sectors		= BUF_SIZE >> 9,
+			.dev_sectors		= DEV_SECTORS,
+		},
+	};
+	int ret;
+
+	nr_queues = queues;
+	dev->dev_info = info;
+	ret = ublk_ctrl_add_dev(dev);
+	if (ret)
+		return ret;
+	dev_id = dev->dev_info.dev_id;
+
+	ret = ublk_ctrl_set_params(ctrl(), &p);
+	if (ret)
+		return ret;
+
+	for (int q = 0; q < queues; q++)
+		for (int i = 0; i < DEPTH; i++)
+			if (!bufs[q][i] &&
+			    posix_memalign(&bufs[q][i], getpagesize(), BUF_SIZE))
+				return -ENOMEM;
+	return 0;
+}
+
+static int add_dev(int queues, __u64 flags)
+{
+	return add_dev_noopen(queues, flags) ?: open_cdev();
+}
+
+static void cleanup(void)
+{
+	if (dev_id < 0)
+		return;
+	ublk_ctrl_stop_dev(ctrl());
+	close_cdev();
+	ublk_ctrl_del_dev(ctrl());
+	dev_id = -1;
+}
+
+/* race_async_fetch sets IOSQE_ASYNC on the io commands */
+static int io_cmd_sqe_flags;
+
+static void queue_io_cmd(struct io_uring *ring, __u32 op, int q, int tag,
+			 int res)
+{
+	struct io_uring_sqe *sqe = get_sqe(ring);
+	struct ublksrv_io_cmd *cmd = (struct ublksrv_io_cmd *)sqe->cmd;
+
+	memset(sqe, 0, sizeof(*sqe));
+	sqe->fd = cdev_fd;
+	sqe->opcode = IORING_OP_URING_CMD;
+	sqe->flags = io_cmd_sqe_flags;
+	ublk_set_sqe_cmd_op(sqe, op);
+	cmd->q_id = q;
+	cmd->tag = tag;
+	cmd->result = res;
+	cmd->addr = (__u64)(uintptr_t)bufs[q][tag];
+	io_uring_sqe_set_data64(sqe, (q << 16) | tag);
+}
+
+struct async_arg {
+	pthread_barrier_t go, stopped;
+};
+
+/*
+ * FETCH every tag with IOSQE_ASYNC, racing STOP_DEV, then close the ring once
+ * STOP_DEV returned: that walks io_uring's list of cancelable commands.
+ */
+static void *race_async_fn(void *data)
+{
+	struct async_arg *a = data;
+	struct io_uring ring;
+	int ok = !io_uring_queue_init(DEPTH, &ring, 0);
+
+	if (ok)
+		for (int tag = 0; tag < DEPTH; tag++)
+			queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, 0, tag, 0);
+	pthread_barrier_wait(&a->go);
+	if (ok)
+		io_uring_submit(&ring);
+	pthread_barrier_wait(&a->stopped);
+	if (ok)
+		io_uring_queue_exit(&ring);
+	return NULL;
+}
+
+/* reap @nr completions of canceled fetch commands, return how many */
+static int reap_aborts(struct io_uring *ring, int nr)
+{
+	struct __kernel_timespec ts = { .tv_sec = 2 };
+	struct io_uring_cqe *cqe;
+	int aborted = 0;
+
+	while (nr--) {
+		if (io_uring_wait_cqe_timeout(ring, &cqe, &ts))
+			break;
+		if (cqe->res == UBLK_IO_RES_ABORT)
+			aborted++;
+		io_uring_cqe_seen(ring, cqe);
+	}
+	return aborted;
+}
+
+/*
+ * A server thread fetches @nr_tags tags of queue @q, then completes each
+ * request after @delay_us, until its commands are aborted.
+ */
+struct server {
+	int q, first_tag, nr_tags, delay_us;
+	int failed;	/* result which ended the serve loop, 0 if none */
+	pthread_t thread;
+	pthread_barrier_t fetched;
+	/* race_fetch: wait for @go and @race_delay_us, open, fetch one by one */
+	pthread_barrier_t go;
+	int race, race_delay_us;
+};
+
+static void *server_fn(void *data)
+{
+	struct server *s = data;
+	struct io_uring ring;
+	struct io_uring_cqe *cqe;
+
+	if (io_uring_queue_init(DEPTH, &ring, 0)) {
+		pthread_barrier_wait(&s->fetched);
+		return NULL;
+	}
+	if (s->race) {
+		pthread_barrier_wait(&s->go);
+		usleep(s->race_delay_us);
+		if (open_cdev()) {
+			pthread_barrier_wait(&s->fetched);
+			io_uring_queue_exit(&ring);
+			return NULL;
+		}
+	}
+	for (int tag = s->first_tag; tag < s->first_tag + s->nr_tags; tag++) {
+		queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, s->q, tag, 0);
+		if (s->race)
+			io_uring_submit(&ring);
+	}
+	io_uring_submit(&ring);
+	pthread_barrier_wait(&s->fetched);
+
+	while (!io_uring_wait_cqe(&ring, &cqe)) {
+		int tag = cqe->user_data & 0xffff;
+		const struct ublksrv_io_desc *iod = &iods[s->q][tag];
+		int res = cqe->res;
+
+		io_uring_cqe_seen(&ring, cqe);
+		if (res != UBLK_IO_RES_OK) {
+			__atomic_store_n(&s->failed, res, __ATOMIC_RELEASE);
+			break;
+		}
+		usleep(s->delay_us);
+		res = ublksrv_get_op(iod) <= UBLK_IO_OP_WRITE ?
+			iod->nr_sectors << 9 : 0;
+		queue_io_cmd(&ring, UBLK_U_IO_COMMIT_AND_FETCH_REQ, s->q, tag,
+			     res);
+		io_uring_submit(&ring);
+	}
+	io_uring_queue_exit(&ring);
+	return NULL;
+}
+
+/* returns once the fetch commands are issued; with @race, call race_go() */
+static void server_start(struct server *s)
+{
+	pthread_barrier_init(&s->fetched, NULL, 2);
+	pthread_barrier_init(&s->go, NULL, 2);
+	pthread_create(&s->thread, NULL, server_fn, s);
+	if (!s->race)
+		pthread_barrier_wait(&s->fetched);
+}
+
+/* wait up to 5s for the serve loop of @s to end, return its result */
+static int server_wait_failed(struct server *s)
+{
+	int res = 0;
+
+	for (int i = 0; i < 500 && !res; i++) {
+		res = __atomic_load_n(&s->failed, __ATOMIC_ACQUIRE);
+		if (!res)
+			usleep(10000);
+	}
+	return res;
+}
+
+/*
+ * A task fetches @nr_tags tags of queue @q and dies. Returns once its
+ * commands are canceled, see the top of this file.
+ */
+static int fetch_and_die(int q, int first_tag, int nr_tags)
+{
+	int status;
+	pid_t pid = fork();
+
+	if (pid < 0)
+		return -1;
+	if (!pid) {
+		struct io_uring ring;
+
+		if (io_uring_queue_init(DEPTH, &ring, 0))
+			_exit(1);
+		for (int tag = first_tag; tag < first_tag + nr_tags; tag++)
+			queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, q, tag, 0);
+		if (io_uring_submit(&ring) != nr_tags)
+			_exit(1);
+		execlp("true", "true", NULL);
+		_exit(1);
+	}
+	if (waitpid(pid, &status, 0) != pid || !WIFEXITED(status) ||
+	    WEXITSTATUS(status))
+		return -1;
+	printf("q%d: task died with %d fetch cmds in flight\n", q, nr_tags);
+	return 0;
+}
+
+struct reads {
+	struct io_uring ring;
+	void *buf;
+	int fd, done, ok, eio, other;
+};
+
+/*
+ * Issue NR_READS reads at once, from @cpu if it is not negative: blk-mq
+ * maps the submitting CPU to the hw queue. A server holding its live
+ * tags for a while makes the other reads take the canceled tags.
+ */
+static int open_tries = 100;	/* 50ms each */
+
+static int reads_submit(struct reads *r, int cpu)
+{
+	char path[64];
+	cpu_set_t set, old;
+
+	snprintf(path, sizeof(path), "/dev/ublkb%d", dev_id);
+	r->fd = -1;
+	for (int i = 0; i < open_tries && r->fd < 0; i++) {
+		r->fd = open(path, O_RDONLY | O_DIRECT);
+		if (r->fd < 0)
+			usleep(50000);
+	}
+	if (r->fd < 0) {
+		fprintf(stderr, "open %s: %s\n", path, strerror(errno));
+		return -1;
+	}
+	if (posix_memalign(&r->buf, 4096, NR_READS * 4096) ||
+	    io_uring_queue_init(NR_READS, &r->ring, 0))
+		return -1;
+
+	for (int i = 0; i < NR_READS; i++)
+		io_uring_prep_read(get_sqe(&r->ring), r->fd,
+				   r->buf + i * 4096, 4096, i * 4096);
+
+	if (cpu >= 0) {
+		sched_getaffinity(0, sizeof(old), &old);
+		CPU_ZERO(&set);
+		CPU_SET(cpu, &set);
+		sched_setaffinity(0, sizeof(set), &set);
+	}
+	io_uring_submit(&r->ring);
+	if (cpu >= 0)
+		sched_setaffinity(0, sizeof(old), &old);
+	return 0;
+}
+
+static void reads_reap(struct reads *r, int timeout_s)
+{
+	struct __kernel_timespec ts = { .tv_sec = timeout_s };
+	struct io_uring_cqe *cqe;
+
+	while (r->done < NR_READS &&
+	       !io_uring_wait_cqe_timeout(&r->ring, &cqe, &ts)) {
+		if (cqe->res == 4096)
+			r->ok++;
+		else if (cqe->res == -EIO)
+			r->eio++;
+		else
+			r->other++;
+		r->done++;
+		io_uring_cqe_seen(&r->ring, cqe);
+	}
+}
+
+static void reads_put(struct reads *r)
+{
+	io_uring_queue_exit(&r->ring);
+	close(r->fd);
+	free(r->buf);
+}
+
+static int reads_result(struct reads *r)
+{
+	printf("reads: %d ok, %d -EIO, %d other, %d not completed\n",
+	       r->ok, r->eio, r->other, NR_READS - r->done);
+	reads_put(r);
+	return r->other || r->done < NR_READS ? KSFT_FAIL : KSFT_PASS;
+}
+
+static int start_dev(void)
+{
+	int ret = ublk_ctrl_start_dev(ctrl(), getpid());
+
+	printf("START_DEV: %d\n", ret);
+	return ret;
+}
+
+static int expect_err(const char *what, int ret, int want)
+{
+	if (ret == want)
+		return KSFT_PASS;
+	fprintf(stderr, "%s: %d, expected %d\n", what, ret, want);
+	return KSFT_FAIL;
+}
+
+/* START_DEV must fail with @want; if it went live, show what a read does */
+static int start_dev_expect(int want)
+{
+	struct reads r = {};
+	int ret = start_dev();
+
+	if (!ret && !reads_submit(&r, -1)) {
+		reads_reap(&r, 10);
+		reads_result(&r);
+	}
+	return expect_err("START_DEV", ret, want);
+}
+
+/* -ENODEV, or live over a canceling queue: then reads must complete */
+static int expect_enodev_or_reads(const char *what, int ret, struct reads *r)
+{
+	if (ret == -ENODEV)
+		return KSFT_PASS;
+	if (ret) {
+		fprintf(stderr, "%s: %d, expected 0 or %d\n", what, ret,
+			-ENODEV);
+		return KSFT_FAIL;
+	}
+	reads_reap(r, 10);
+	return reads_result(r);
+}
+
+static int test_stop_start(void)
+{
+	struct io_uring ring;
+	int ret;
+
+	if (add_dev(1, 0))
+		return KSFT_FAIL;
+	if (io_uring_queue_init(DEPTH, &ring, 0))
+		return KSFT_FAIL;
+	for (int tag = 0; tag < DEPTH; tag++)
+		queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, 0, tag, 0);
+	io_uring_submit(&ring);
+
+	/* device is ready but not started: state is UBLK_S_DEV_DEAD */
+	ret = ublk_ctrl_stop_dev(ctrl());
+	printf("STOP_DEV: %d, canceled fetch cmds: %d/%d\n", ret,
+	       reap_aborts(&ring, DEPTH), DEPTH);
+
+	/* STOP_DEV canceled the attached server: no start until it exits */
+	ret = start_dev_expect(-EBUSY);
+	io_uring_queue_exit(&ring);
+	return ret;
+}
+
+static int test_partial_fetch(void)
+{
+	struct server b = { .q = 0, .first_tag = DEPTH - 1, .nr_tags = 1,
+			    .delay_us = SERVE_DELAY_US };
+	struct reads r = {};
+	int ret;
+
+	if (add_dev(1, 0) || fetch_and_die(0, 0, DEPTH - 1))
+		return KSFT_FAIL;
+
+	/* the last FETCH makes the queue ready and clears ->canceling */
+	server_start(&b);
+
+	ret = start_dev();
+	if (!ret && reads_submit(&r, -1))
+		ret = KSFT_FAIL;
+	else
+		ret = expect_enodev_or_reads("START_DEV", ret, &r);
+
+	cleanup();
+	pthread_join(b.thread, NULL);
+	return ret;
+}
+
+static int test_stop_restart(void)
+{
+	struct server s = { .q = 0, .nr_tags = DEPTH };
+	struct reads r = {};
+	int ret;
+
+	if (add_dev_noopen(1, 0))
+		return KSFT_FAIL;
+
+	/* no server is attached, so there is nothing to cancel */
+	ret = ublk_ctrl_stop_dev(ctrl());
+	printf("STOP_DEV: %d\n", ret);
+	if (open_cdev())
+		return KSFT_FAIL;
+
+	server_start(&s);
+	if (start_dev()) {
+		ret = KSFT_FAIL;
+	} else if (reads_submit(&r, -1)) {
+		ret = KSFT_FAIL;
+	} else {
+		reads_reap(&r, 10);
+		ret = reads_result(&r);
+		if (r.ok != NR_READS)
+			ret = KSFT_FAIL;
+	}
+
+	cleanup();
+	pthread_join(s.thread, NULL);
+	return ret;
+}
+
+struct start_arg {
+	pthread_barrier_t go;
+	int ret, reads_ok;
+};
+
+static void *race_start_fn(void *data)
+{
+	struct start_arg *a = data;
+	struct reads r = {};
+
+	pthread_barrier_wait(&a->go);
+	a->ret = ublk_ctrl_start_dev(ctrl(), getpid());
+	a->reads_ok = 1;
+	/*
+	 * The disk may be gone already if STOP_DEV came right after, and
+	 * reads may fail then: they only have to complete.
+	 */
+	if (!a->ret && !reads_submit(&r, -1)) {
+		reads_reap(&r, 5);
+		a->reads_ok = r.done == NR_READS;
+		if (a->reads_ok)
+			reads_put(&r);
+		else
+			reads_result(&r);
+	}
+	ctrl_put();
+	return NULL;
+}
+
+static int test_race_start(void)
+{
+	int live = 0, ebusy = 0;
+
+	open_tries = 4;
+	for (int i = 0; i < RACE_LOOPS; i++) {
+		struct server s = { .q = 0, .nr_tags = DEPTH };
+		struct start_arg a = {};
+		pthread_t t;
+
+		if (add_dev(1, 0))
+			return KSFT_FAIL;
+		server_start(&s);
+		pthread_barrier_init(&a.go, NULL, 2);
+		pthread_create(&t, NULL, race_start_fn, &a);
+		pthread_barrier_wait(&a.go);
+		ublk_ctrl_stop_dev(ctrl());
+		pthread_join(t, NULL);
+		cleanup();
+		pthread_join(s.thread, NULL);
+
+		if (a.ret == 0)
+			live++;
+		else if (a.ret == -EBUSY)
+			ebusy++;
+		if ((a.ret && a.ret != -EBUSY) || !a.reads_ok) {
+			fprintf(stderr, "loop %d: START_DEV %d, reads %s\n",
+				i, a.ret, a.reads_ok ? "ok" : "failed");
+			return KSFT_FAIL;
+		}
+	}
+	printf("%d loops: START_DEV won %d, got -EBUSY %d\n", RACE_LOOPS,
+	       live, ebusy);
+	return KSFT_PASS;
+}
+
+static int test_race_fetch(void)
+{
+	int live = 0, ebusy = 0;
+
+	for (int i = 0; i < RACE_LOOPS; i++) {
+		/* vary who goes first: STOP_DEV, or the server's open */
+		struct server s = { .q = 0, .nr_tags = DEPTH, .race = 1,
+				    .race_delay_us = (i % 10) * 50 };
+		struct reads r = {};
+		int ret;
+
+		if (add_dev_noopen(1, 0))
+			return KSFT_FAIL;
+		server_start(&s);
+		pthread_barrier_wait(&s.go);
+		ublk_ctrl_stop_dev(ctrl());
+		pthread_barrier_wait(&s.fetched);
+
+		ret = ublk_ctrl_start_dev(ctrl(), getpid());
+		if (!ret) {
+			/* nothing was taken, so every read has to succeed */
+			live++;
+			if (reads_submit(&r, -1))
+				return KSFT_FAIL;
+			reads_reap(&r, 5);
+			if (r.ok != NR_READS) {
+				fprintf(stderr, "loop %d: live, but ", i);
+				reads_result(&r);
+				return KSFT_FAIL;
+			}
+			reads_put(&r);
+		} else if (ret == -EBUSY) {
+			ebusy++;
+		} else {
+			fprintf(stderr, "loop %d: START_DEV %d\n", i, ret);
+			return KSFT_FAIL;
+		}
+		cleanup();
+		pthread_join(s.thread, NULL);
+	}
+	printf("%d loops: START_DEV worked %d, got -EBUSY %d\n", RACE_LOOPS,
+	       live, ebusy);
+	return KSFT_PASS;
+}
+
+static int test_race_async_fetch(void)
+{
+	io_cmd_sqe_flags = IOSQE_ASYNC;
+	for (int i = 0; i < ASYNC_LOOPS; i++) {
+		struct async_arg a;
+		pthread_t t;
+
+		if (add_dev(1, 0))
+			return KSFT_FAIL;
+		pthread_barrier_init(&a.go, NULL, 2);
+		pthread_barrier_init(&a.stopped, NULL, 2);
+		pthread_create(&t, NULL, race_async_fn, &a);
+		pthread_barrier_wait(&a.go);
+		/* vary where STOP_DEV lands relative to the FETCHes */
+		usleep((i % 5) * 1000);
+		ublk_ctrl_stop_dev(ctrl());
+		pthread_barrier_wait(&a.stopped);
+		pthread_join(t, NULL);
+		cleanup();
+	}
+	printf("%d loops done\n", ASYNC_LOOPS);
+	return KSFT_PASS;
+}
+
+/*
+ * The stopped server is gone: reopen /dev/ublkcN, start a new server and
+ * the device, and read from it.
+ */
+static int restart_server(void)
+{
+	struct server s = { .q = 0, .nr_tags = DEPTH };
+	struct reads r = {};
+	int ret;
+
+	close_cdev();
+	if (open_cdev())
+		return KSFT_FAIL;
+	server_start(&s);
+	usleep(200000);
+	if (s.failed) {
+		fprintf(stderr, "new server: FETCH failed %d\n", s.failed);
+		cleanup();
+		pthread_join(s.thread, NULL);
+		return KSFT_FAIL;
+	}
+	/* -EEXIST until the old disk is freed, e.g. after a udev probe */
+	for (int i = 0; i < 100; i++) {
+		ret = start_dev();
+		if (ret != -EEXIST)
+			break;
+		usleep(50000);
+	}
+	if (ret || reads_submit(&r, -1)) {
+		fprintf(stderr, "new server: START_DEV %d\n", ret);
+		ret = KSFT_FAIL;
+	} else {
+		reads_reap(&r, 10);
+		ret = reads_result(&r);
+		if (r.ok != NR_READS)
+			ret = KSFT_FAIL;
+	}
+	cleanup();
+	pthread_join(s.thread, NULL);
+	return ret;
+}
+
+/*
+ * stop_attached: STOP_DEV while a server has the device open but has
+ * fetched nothing stops that server too: its FETCH gets ABORT and
+ * START_DEV -EBUSY. Once it is gone, a new server can start the device.
+ */
+static int test_stop_attached(void)
+{
+	struct server s = { .q = 0, .nr_tags = DEPTH };
+	int ret;
+
+	if (add_dev(1, 0))
+		return KSFT_FAIL;
+	ret = ublk_ctrl_stop_dev(ctrl());
+	printf("STOP_DEV: %d\n", ret);
+
+	server_start(&s);
+	ret = server_wait_failed(&s);
+	printf("FETCH after STOP_DEV: %d\n", ret);
+	if (ret != UBLK_IO_RES_ABORT) {
+		cleanup();
+		pthread_join(s.thread, NULL);
+		return KSFT_FAIL;
+	}
+	pthread_join(s.thread, NULL);
+	if (expect_err("START_DEV", start_dev(), -EBUSY))
+		return KSFT_FAIL;
+	return restart_server();
+}
+
+/*
+ * stop_live_restart: STOP_DEV on a live device; once its server is gone,
+ * a new server has to fetch and start it again.
+ */
+static int test_stop_live_restart(void)
+{
+	struct server s = { .q = 0, .nr_tags = DEPTH };
+	int ret;
+
+	if (add_dev(1, 0))
+		return KSFT_FAIL;
+	server_start(&s);
+	if (expect_err("START_DEV", start_dev(), 0))
+		return KSFT_FAIL;
+	ret = ublk_ctrl_stop_dev(ctrl());
+	printf("STOP_DEV: %d\n", ret);
+	pthread_join(s.thread, NULL);
+	return restart_server();
+}
+
+/* first CPU which blk-mq maps to hw queue @q */
+static int queue_cpu(int q)
+{
+	char path[96];
+	FILE *f;
+	int cpu = -1;
+
+	snprintf(path, sizeof(path), "/sys/block/ublkb%d/mq/%d/cpu_list",
+		 dev_id, q);
+	for (int i = 0; i < 100 && !(f = fopen(path, "r")); i++)
+		usleep(50000);
+	if (!f)
+		return -1;
+	if (fscanf(f, "%d", &cpu) != 1)
+		cpu = -1;
+	fclose(f);
+	return cpu;
+}
+
+static int test_recovery(void)
+{
+	struct server q1 = { .q = 1, .nr_tags = DEPTH };
+	struct io_uring ring;
+	struct reads r = {};
+	int ret, cpu;
+
+	if (add_dev(NR_QUEUES, UBLK_F_USER_RECOVERY))
+		return KSFT_FAIL;
+
+	/* the first server: fetch everything, start, then die */
+	if (io_uring_queue_init(NR_QUEUES * DEPTH, &ring, 0))
+		return KSFT_FAIL;
+	for (int q = 0; q < NR_QUEUES; q++)
+		for (int tag = 0; tag < DEPTH; tag++)
+			queue_io_cmd(&ring, UBLK_U_IO_FETCH_REQ, q, tag, 0);
+	io_uring_submit(&ring);
+	if (start_dev())
+		return KSFT_FAIL;
+	cpu = queue_cpu(0);
+	if (cpu < 0) {
+		printf("no CPU maps to queue 0, can't aim the reads\n");
+		return KSFT_SKIP;
+	}
+
+	io_uring_queue_exit(&ring);
+	close_cdev();
+	for (int i = 0; i < 100 && dev_state() != UBLK_S_DEV_QUIESCED; i++)
+		usleep(50000);
+	printf("server exited, state %d (QUIESCED is %d)\n", dev_state(),
+	       UBLK_S_DEV_QUIESCED);
+
+	for (int i = 0; i < 100; i++) {
+		ret = ublk_ctrl_start_user_recovery(ctrl());
+		if (ret != -EBUSY)
+			break;
+		usleep(50000);
+	}
+	printf("START_USER_RECOVERY: %d\n", ret);
+	if (ret || open_cdev())
+		return KSFT_FAIL;
+
+	/* queue 0 gets ready, then its task dies before queue 1 is ready */
+	if (fetch_and_die(0, 0, DEPTH))
+		return KSFT_FAIL;
+
+	printf("reads on cpu %d, which maps to queue 0\n", cpu);
+	if (reads_submit(&r, cpu))
+		return KSFT_FAIL;
+
+	server_start(&q1);
+	ret = ublk_ctrl_end_user_recovery(ctrl(), getpid());
+	printf("END_USER_RECOVERY: %d\n", ret);
+	if (ret && ret != -ENODEV) {
+		fprintf(stderr, "END_USER_RECOVERY: %d, expected 0 or %d\n",
+			ret, -ENODEV);
+		ret = KSFT_FAIL;
+	} else {
+		ret = KSFT_PASS;
+	}
+
+	/*
+	 * After -ENODEV, requests held back on queue 0 wait for STOP_DEV to
+	 * fail them. Close the disk before DEL_DEV, which waits for its last
+	 * reference.
+	 */
+	reads_reap(&r, 1);
+	ublk_ctrl_stop_dev(ctrl());
+	reads_reap(&r, 5);
+	if (reads_result(&r))
+		ret = KSFT_FAIL;
+	cleanup();
+	pthread_join(q1.thread, NULL);
+	return ret;
+}
+
+static const struct {
+	const char *name;
+	int (*fn)(void);
+} modes[] = {
+	{ "stop_start",		test_stop_start },
+	{ "partial_fetch",	test_partial_fetch },
+	{ "recovery",		test_recovery },
+	{ "stop_restart",	test_stop_restart },
+	{ "race_start",		test_race_start },
+	{ "race_fetch",		test_race_fetch },
+	{ "race_async_fetch",	test_race_async_fetch },
+	{ "stop_attached",	test_stop_attached },
+	{ "stop_live_restart",	test_stop_live_restart },
+};
+
+int main(int argc, char **argv)
+{
+	int (*fn)(void) = NULL;
+	int ret;
+
+	for (int i = 0; argc == 2 && i < ARRAY_SIZE(modes); i++)
+		if (!strcmp(argv[1], modes[i].name))
+			fn = modes[i].fn;
+	if (!fn) {
+		fprintf(stderr, "usage: %s MODE, modes:", argv[0]);
+		for (int i = 0; i < ARRAY_SIZE(modes); i++)
+			fprintf(stderr, " %s", modes[i].name);
+		fprintf(stderr, "\n");
+		return KSFT_FAIL;
+	}
+
+	/* keep the output of a run that ends in an oops */
+	setvbuf(stdout, NULL, _IOLBF, 0);
+
+	if (access(CTRL_DEV, F_OK)) {
+		perror(CTRL_DEV);
+		return KSFT_SKIP;
+	}
+	printf("%s\n", argv[1]);
+	ret = fn();
+	cleanup();
+	ctrl_put();
+	return ret;
+}

21/3 2026/10/02 08:23 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 26994
							output: 22
							thoughts: 1720
						

Thoughts:
**KMSAN Fuzzing Session Assessment**

Okay, let's break this down. The task is to determine if this patch series warrants a dedicated KMSAN fuzzing session for the ublk driver. The primary goal is to avoid unnecessary resource expenditure, so we'll only trigger KMSAN if the changes introduce *unique* uninitialized memory risks that KASAN wouldn't catch.

First, let's re-familiarize myself with the patch. It seems heavily focused on concurrency issues, race conditions, and handling command cancellation in the `ublk_drv` driver, touching both the core driver code (`drivers/block/ublk_drv.c`) and associated selftests. The commit context and comments point directly to fixing use-after-free bugs related to command cancellation during device bring-up and teardown, particularly around walking lists of commands and access to `io->cmd`.

I see changes that add, or modify, several structures including `ublk_device`. I want to investigate the context, specifically how the `ublk_device` is initialized. Let's see if the allocation uses `kzalloc` or similar to guarantee that the allocated memory is initialized to zero. I am comfortable with using grep or a quick read-file to confirm this.

*   `ublk_batch_alloc_fcmd`: This has `INIT_LIST_HEAD(&fcmd->node);`. This is an initialization, not an uninitialized access.

Now, let's methodically go through the KMSAN criteria.

1.  **Kernel structures sent to user space:** I don't see any explicit `copy_to_user`, `put_user`, netlink usage, ioctl outputs, socket options, or BPF buffers that could leak uninitialized data.
2.  **Conditional logic based on uninitialized variables:** There is a lot of new conditional logic in the patch dealing with the state of the device and command cancellation. But, from the code I have already reviewed it does not appear that any uninitialized variables are being used in these conditionals.
3.  **Allocation/Initialization of complex data structures:** The changes deal with command structures (`fcmd`), and the device structure (`ublk_device`). `INIT_LIST_HEAD` initializes the lists and the allocation is most likely kzalloc.
4.  **Bounds checks exposing uninitialized bytes:** The fixes address concurrency and use-after-free issues, which can indirectly lead to uninitialized memory access, but the *primary* bug is not directly introducing a new bounds check or length issue that *explicitly* leaks uninitialized bytes.

Based on this review, I believe the risk here is primarily related to concurrency, synchronization, and use-after-free bugs that KASAN should handle well. While race conditions *can* lead to uninitialized memory reads, these patches are addressing the underlying concurrency issues, not directly introducing new uninitialized memory vulnerabilities that are *uniquely* detectable by KMSAN. If I determine the ublk_device is allocated using kzalloc then this concludes the matter.

Therefore, for this patch series, a dedicated KMSAN fuzzing session is likely **not** justified. KASAN and standard bug detectors are sufficient.



22/3 2026/10/02 08:23 tool grepper
0m Error:
no matches

Args:
{
  "Expression": "kzalloc\\(.*ublk_device"
}

Results:
{
  "error": "no matches"
}

23/3 2026/10/02 08:24 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 29425
							output: 37
							thoughts: 15
						

Thoughts:
Okay, let's dive into this. Here's how I'd approach this task, as if these were my internal thoughts:

**Investigating `struct ublk_device` Allocation**

Alright, time to get my hands dirty. My first thought is: I need to pinpoint exactly where this `struct ublk_device` is being allocated. Understanding its allocation strategy is fundamental to understanding its lifecycle, potential memory issues, and how it interacts with other parts of the Ublk subsystem. Is it on the stack, the heap, or some specialized memory pool? Is it a statically defined structure? Where is this happening? Let's trace back from there.



24/3 2026/10/02 08:24 tool grepper
0m Args:
{
  "Expression": "struct ublk_device \\*ub",
  "PathPrefix": "drivers/block/ublk_drv.c"
}

Results:
{
  "Output": "\nFull output is too long, showing 200 out of 582 lines.\nUse more precise expression if possible.\n\ndrivers/block/ublk_drv.c=146=struct ublk_batch_io_data {\ndrivers/block/ublk_drv.c:147:\tstruct ublk_device *ub;\ndrivers/block/ublk_drv.c-148-\tstruct io_uring_cmd *cmd;\n--\ndrivers/block/ublk_drv.c=366=static void ublk_io_release(void *priv);\ndrivers/block/ublk_drv.c:367:static void ublk_stop_dev_unlocked(struct ublk_device *ub);\ndrivers/block/ublk_drv.c:368:static bool ublk_try_buf_match(struct ublk_device *ub, struct request *rq,\ndrivers/block/ublk_drv.c-369-\t\t\t\t  u32 *buf_idx, u32 *buf_off);\ndrivers/block/ublk_drv.c:370:static void ublk_buf_cleanup(struct ublk_device *ub);\ndrivers/block/ublk_drv.c:371:static void ublk_abort_queue(struct ublk_device *ub, struct ublk_queue *ubq);\ndrivers/block/ublk_drv.c:372:static inline struct request *__ublk_check_and_get_req(struct ublk_device *ub,\ndrivers/block/ublk_drv.c-373-\t\tu16 q_id, u16 tag, struct ublk_io *io);\ndrivers/block/ublk_drv.c=374=static void ublk_batch_dispatch(struct ublk_queue *ubq,\n--\ndrivers/block/ublk_drv.c-377-\ndrivers/block/ublk_drv.c:378:static inline bool ublk_dev_support_batch_io(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-379-{\n--\ndrivers/block/ublk_drv.c=424=static inline bool ublk_support_zero_copy(const struct ublk_queue *ubq)\n--\ndrivers/block/ublk_drv.c-428-\ndrivers/block/ublk_drv.c:429:static inline bool ublk_dev_support_zero_copy(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-430-{\n--\ndrivers/block/ublk_drv.c=439=static inline bool ublk_iod_is_shmem_zc(const struct ublk_queue *ubq, u16 tag)\n--\ndrivers/block/ublk_drv.c-443-\ndrivers/block/ublk_drv.c:444:static inline bool ublk_dev_support_shmem_zc(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-445-{\n--\ndrivers/block/ublk_drv.c=449=static inline bool ublk_support_auto_buf_reg(const struct ublk_queue *ubq)\n--\ndrivers/block/ublk_drv.c-453-\ndrivers/block/ublk_drv.c:454:static inline bool ublk_dev_support_auto_buf_reg(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-455-{\n--\ndrivers/block/ublk_drv.c=459=static inline bool ublk_support_user_copy(const struct ublk_queue *ubq)\n--\ndrivers/block/ublk_drv.c-463-\ndrivers/block/ublk_drv.c:464:static inline bool ublk_dev_support_user_copy(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-465-{\n--\ndrivers/block/ublk_drv.c-468-\ndrivers/block/ublk_drv.c:469:static inline bool ublk_dev_is_zoned(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-470-{\n--\ndrivers/block/ublk_drv.c=474=static inline bool ublk_queue_is_zoned(const struct ublk_queue *ubq)\n--\ndrivers/block/ublk_drv.c-478-\ndrivers/block/ublk_drv.c:479:static inline bool ublk_dev_support_integrity(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-480-{\n--\ndrivers/block/ublk_drv.c=562=static struct ublk_zoned_report_desc *ublk_zoned_get_report_desc(\n--\ndrivers/block/ublk_drv.c-567-\ndrivers/block/ublk_drv.c:568:static int ublk_get_nr_zones(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-569-{\n--\ndrivers/block/ublk_drv.c-575-\ndrivers/block/ublk_drv.c:576:static int ublk_revalidate_disk_zones(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-577-{\n--\ndrivers/block/ublk_drv.c-580-\ndrivers/block/ublk_drv.c:581:static int ublk_dev_param_zoned_validate(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-582-{\n--\ndrivers/block/ublk_drv.c-602-\ndrivers/block/ublk_drv.c:603:static void ublk_dev_param_zoned_apply(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-604-{\n--\ndrivers/block/ublk_drv.c-608-/* Based on virtblk_alloc_report_buffer */\ndrivers/block/ublk_drv.c:609:static void *ublk_alloc_report_buffer(struct ublk_device *ublk,\ndrivers/block/ublk_drv.c-610-\t\t\t\t      unsigned int nr_zones, size_t *buflen)\n--\ndrivers/block/ublk_drv.c=636=static int ublk_report_zones(struct gendisk *disk, sector_t sector,\n--\ndrivers/block/ublk_drv.c-638-{\ndrivers/block/ublk_drv.c:639:\tstruct ublk_device *ub = disk-\u003eprivate_data;\ndrivers/block/ublk_drv.c-640-\tunsigned int zone_size_sectors = disk-\u003equeue-\u003elimits.chunk_sectors;\n--\ndrivers/block/ublk_drv.c=733=static void ublk_setup_iod_zoned(struct ublk_queue *ubq, struct request *req)\n--\ndrivers/block/ublk_drv.c-773-\ndrivers/block/ublk_drv.c:774:static int ublk_dev_param_zoned_validate(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-775-{\n--\ndrivers/block/ublk_drv.c-778-\ndrivers/block/ublk_drv.c:779:static void ublk_dev_param_zoned_apply(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-780-{\n--\ndrivers/block/ublk_drv.c-782-\ndrivers/block/ublk_drv.c:783:static int ublk_revalidate_disk_zones(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-784-{\n--\ndrivers/block/ublk_drv.c=899=static inline u16 ublk_pos_to_tag(loff_t pos)\n--\ndrivers/block/ublk_drv.c-904-\ndrivers/block/ublk_drv.c:905:static void ublk_dev_param_basic_apply(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-906-{\n--\ndrivers/block/ublk_drv.c=945=static enum blk_integrity_checksum ublk_integrity_csum_type(u8 csum_type)\n--\ndrivers/block/ublk_drv.c-961-\ndrivers/block/ublk_drv.c:962:static int ublk_validate_params(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-963-{\n--\ndrivers/block/ublk_drv.c-1066-\ndrivers/block/ublk_drv.c:1067:static void ublk_apply_params(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1068-{\n--\ndrivers/block/ublk_drv.c=1075=static inline bool ublk_need_map_io(const struct ublk_queue *ubq)\n--\ndrivers/block/ublk_drv.c-1080-\ndrivers/block/ublk_drv.c:1081:static inline bool ublk_dev_need_map_io(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1082-{\n--\ndrivers/block/ublk_drv.c=1088=static inline bool ublk_need_req_ref(const struct ublk_queue *ubq)\n--\ndrivers/block/ublk_drv.c-1104-\ndrivers/block/ublk_drv.c:1105:static inline bool ublk_dev_need_req_ref(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1106-{\n--\ndrivers/block/ublk_drv.c=1230=static inline bool ublk_need_get_data(const struct ublk_queue *ubq)\n--\ndrivers/block/ublk_drv.c-1234-\ndrivers/block/ublk_drv.c:1235:static inline bool ublk_dev_need_get_data(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1236-{\n--\ndrivers/block/ublk_drv.c-1240-/* Called in slow path only, keep it noinline for trace purpose */\ndrivers/block/ublk_drv.c:1241:static noinline struct ublk_device *ublk_get_device(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1242-{\n--\ndrivers/block/ublk_drv.c-1248-/* Called in slow path only, keep it noinline for trace purpose */\ndrivers/block/ublk_drv.c:1249:static noinline void ublk_put_device(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1250-{\n--\ndrivers/block/ublk_drv.c=1260=static inline struct ublksrv_io_desc *\ndrivers/block/ublk_drv.c:1261:ublk_queue_cmd_buf(struct ublk_device *ub, u16 q_id)\ndrivers/block/ublk_drv.c-1262-{\n--\ndrivers/block/ublk_drv.c-1265-\ndrivers/block/ublk_drv.c:1266:static inline size_t __ublk_queue_cmd_buf_size(const struct ublk_device *ub,\ndrivers/block/ublk_drv.c-1267-\t\t\t\t\t       u16 depth)\n--\ndrivers/block/ublk_drv.c-1271-\ndrivers/block/ublk_drv.c:1272:static inline size_t ublk_queue_cmd_buf_size(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1273-{\n--\ndrivers/block/ublk_drv.c-1276-\ndrivers/block/ublk_drv.c:1277:static size_t ublk_max_cmd_buf_size(const struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1278-{\n--\ndrivers/block/ublk_drv.c-1285- */\ndrivers/block/ublk_drv.c:1286:static inline bool ublk_nosrv_should_reissue_outstanding(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1287-{\n--\ndrivers/block/ublk_drv.c-1295- */\ndrivers/block/ublk_drv.c:1296:static inline bool ublk_nosrv_dev_should_queue_io(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1297-{\n--\ndrivers/block/ublk_drv.c=1307=static inline bool ublk_nosrv_should_queue_io(struct ublk_queue *ubq)\n--\ndrivers/block/ublk_drv.c-1318- */\ndrivers/block/ublk_drv.c:1319:static inline bool ublk_nosrv_should_stop_dev(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1320-{\n--\ndrivers/block/ublk_drv.c-1323-\ndrivers/block/ublk_drv.c:1324:static inline bool ublk_dev_in_recoverable_state(struct ublk_device *ub)\ndrivers/block/ublk_drv.c-1325-{\n--\ndrivers/block/ublk_drv.c=1330=static void ublk_free_disk(struct gendisk *disk)\ndrivers/block/ublk_drv.c-1331-{\ndrivers/block/ublk_drv.c:1332:\tstruct ublk_device *ub = disk-\u003eprivate_data;\ndrivers/block/ublk_drv.c-1333-\n--\ndrivers/block/ublk_drv.c=1350=static int ublk_open(struct gendisk *disk, blk_mode_t mode)\ndrivers/block/ublk_drv.c-1351-{\ndrivers/block/ublk_drv.c:1352:\tstruct ublk_device *ub = disk-\u003eprivate_data;\ndrivers/block/ublk_drv.c-1353-\n--\ndrivers/block/ublk_drv.c=1831=static bool __ublk_batch_prep_dispatch(struct ublk_queue *ubq,\n--\ndrivers/block/ublk_drv.c-1834-{\ndrivers/block/ublk_drv.c:1835:\tstruct ublk_device *ub = data-\u003eub;\ndrivers/block/ublk_drv.c-1836-\tstruct ublk_io *io = \u0026ubq-\u003eios[tag];\n--\ndrivers/block/ublk_drv.c=2346=static int ublk_init_hctx(struct blk_mq_hw_ctx *hctx, void *driver_data,\n--\ndrivers/block/ublk_drv.c-2348-{\ndrivers/block/ublk_drv.c:2349:\tstruct ublk_device *ub = driver_data;\ndrivers/block/ublk_drv.c-2350-\tstruct ublk_queue *ubq = ublk_get_queue(ub, hctx-\u003equeue_num);\n--\ndrivers/block/ublk_drv.c=2363=static const struct blk_mq_ops ublk_batch_mq_ops = {\n--\ndrivers/block/ublk_drv.c-2370-\ndrivers/block/ublk_drv.c:2371:static void ublk_queue_reinit(struct ublk_device *ub, struct ublk_queue *ubq)\ndrivers/block/ublk_drv.c-2372-{\n--\ndrivers/block/ublk_drv.c=2404=static int ublk_ch_open(struct inode *inode, struct file *filp)\ndrivers/block/ublk_drv.c-2405-{\n\n"
}

25/3 2026/10/02 08:24 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 33726
							output: 38
							thoughts: 21
						

Thoughts:
Okay, here's my thought process laid out, as if I were thinking it:

**Locating the `ublk_device` Allocation**

Alright, time to track down how `ublk_device` instances are created. I'm going to start with the standard approach for kernel memory allocation: `kzalloc`. That's usually the go-to, zeroed memory allocation function. So, I'll leverage `grepper` – that's the fastest way to get a bird's-eye view of where `kzalloc` is used within the kernel source. By searching for `kzalloc`, I can quickly pinpoint the locations where a block of zeroed memory is allocated. Then, hopefully, I'll be able to quickly spot where a `ublk_device` is allocated.



26/3 2026/10/02 08:24 tool grepper
0m Error:
no matches

Args:
{
  "Expression": "kzalloc.*sizeof\\(.*ublk",
  "PathPrefix": "drivers/block/ublk_drv.c"
}

Results:
{
  "error": "no matches"
}

27/3 2026/10/02 08:24 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 33968
							output: 38
							thoughts: 16
						

Thoughts:
Okay, here's my summary, tailored for an expert audience, presented as if it were my thought process:

**Hunting Down the Allocation Function in ublk_drv.c**

Alright, time to get down to brass tacks. I need to understand the memory allocation strategy within the Ublk driver. My immediate goal is to locate the primary allocation function. This is fundamental, understanding how the driver manages memory is crucial to, say, analyzing potential performance bottlenecks or debugging memory-related issues. The starting point, naturally, is `ublk_drv.c`. That’s where the core driver logic resides. I'll need to use my intuition here, perhaps `kmalloc`, `kzalloc`, or some custom `ublk_alloc` style function are what I should be looking for. I'll need to quickly scan the file, paying close attention to any function or macro definitions that handle memory management.



28/3 2026/10/02 08:24 tool grepper
0m Error:
no matches

Args:
{
  "Expression": "struct ublk_device \\*.*alloc",
  "PathPrefix": "drivers/block/ublk_drv.c"
}

Results:
{
  "error": "no matches"
}

29/3 2026/10/02 08:24 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 34226
							output: 32
							thoughts: 17
						

Thoughts:
Okay, here's my attempt at summarizing that thought process, keeping in mind the expert user:

**Deep Dive: Locating `ublk_device` Allocation**

Alright, let's get down to brass tacks. I need to pinpoint the allocation of a `ublk_device` structure and, specifically, the mechanism used. My immediate thought is to look for instances where it's being allocated with `kzalloc`. This tells me a few things right off the bat. First, it implies we're dealing with kernel memory. Second, `kzalloc` indicates that the memory is being initialized to zero, which is a common practice for device structures to ensure a known starting state and avoid potential security vulnerabilities from uninitialized data. Knowing this detail, my mental search is already filtering out a lot of possible areas. I'm focusing on code paths that involve device initialization, module loading, or perhaps dynamic device registration within the Ublk subsystem. Finding this call will likely be a key piece of understanding how this device is brought into existence and how its initial state is set.



30/3 2026/10/02 08:24 tool grepper
0m Args:
{
  "Expression": "kzalloc",
  "PathPrefix": "drivers/block/ublk_drv.c"
}

Results:
{
  "Output": "drivers/block/ublk_drv.c=815=ublk_batch_alloc_fcmd(struct io_uring_cmd *cmd)\ndrivers/block/ublk_drv.c-816-{\ndrivers/block/ublk_drv.c:817:\tstruct ublk_batch_fetch_cmd *fcmd = kzalloc_obj(*fcmd, GFP_NOIO);\ndrivers/block/ublk_drv.c-818-\n--\ndrivers/block/ublk_drv.c=4774=static int ublk_ctrl_add_dev(const struct ublksrv_ctrl_cmd *header)\n--\ndrivers/block/ublk_drv.c-4882-\tret = -ENOMEM;\ndrivers/block/ublk_drv.c:4883:\tub = kzalloc_flex(*ub, queues, info.nr_hw_queues);\ndrivers/block/ublk_drv.c-4884-\tif (!ub)\n--\ndrivers/block/ublk_drv.c=5493=static int __ublk_ctrl_reg_buf(struct ublk_device *ub,\n--\ndrivers/block/ublk_drv.c-5509-\ndrivers/block/ublk_drv.c:5510:\t\trange = kzalloc_obj(*range);\ndrivers/block/ublk_drv.c-5511-\t\tif (!range) {\n"
}

31/3 2026/10/02 08:24 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 34858
							output: 290
							thoughts: 837
						

Thoughts:
**KMSAN Analysis: No Uninitialized Memory Risks**

Okay, let's break this down. My initial observation is that `ub` is initialized using `kzalloc_flex`, and `fcmd` uses `kzalloc_obj`. This means both structures are guaranteed to be zero-initialized. My concern now is whether any subsequent code modifications introduce any risks by exposing uninitialized memory or struct padding to user space.

I've examined the code changes, and I see nothing new being copied to user space – no new data structures or fields are being exposed via `copy_to_user`, `put_user`, netlink, ioctls, or anything like that. The modifications are focused on internal state management within the `ublk_drv.c` driver.

Specifically, the changes deal with `UB_STATE_STOPPING` which is an internal state bit, `ch_file`, an internal pointer and `canceling`, a boolean flag, within the `ub` struct, as well as modifications to the `fcmd` struct, and functions like `ublk_stop_dev`, `ublk_prep_cancel`, and `ublk_batch_attach`. These changes are primarily aimed at fixing race conditions and synchronization issues related to command completion, cancellation, and device lifecycle events. They are addressing bugs related to race conditions between `STOP_DEV`, `START_DEV`, `FETCH`, process termination (`exec/exit`), potential use-after-free scenarios with `fcmd` and `cmd`, NULL pointer dereferences of `io->cmd`, and deadlock issues related to lock ordering.

Let's carefully verify the KMSAN criteria.

1.  **Kernel structures sent to user space:** None of the changes introduce any new cases where kernel structures, including potentially uninitialized fields or padding, are copied to user space. The core structures are already zeroed using the `kzalloc` family of functions.

2.  **Conditional logic based on uninitialized data:** All relevant fields (`ub->state`, `ub->ch_file`, `ub->canceling`, `fcmd->node`) are *guaranteed* to be zero-initialized. Their states are explicitly managed through the code.

3.  **Allocation/Initialization of complex data:**  `fcmd->node` is now explicitly initialized with `INIT_LIST_HEAD(&fcmd->node)`. This reinforces the zero-initialization strategy and avoids potential issues here.

4.  **Bounds/Length checks:** No modifications to bounds checks, data lengths, or related logic have been identified that could lead to out-of-bounds reads of uninitialized memory.

Based on this analysis, I'm confident that this patch does *not* introduce any uninitialized memory risks. The bugs it fixes or could potentially introduce are focused on use-after-free conditions (which KASAN will catch), NULL pointer dereferences (again, KASAN would detect those), or concurrency and locking problems (detectable by LOCKDEP and standard bug detectors).

Therefore, a dedicated KMSAN fuzzing session isn't warranted in this instance. The `set-results` call is prepared with `NeedsKMSAN = false` and a description of the patch analysis, emphasizing the key points – that it touches the ublk device cancellation state machine, command completion synchronization, cancellation flags, file reference handling, how structures are allocated with zero-initialization (kzalloc_flex, kzalloc_obj), and that no kernel structures or padding are exposed/copied to userspace. Potential bugs related to use-after-free, NULL pointer dereferences, or race conditions are detectable by KASAN/LOCKDEP, thus there are no uninitialized memory risks. Let's invoke `set-results`.



32/3 2026/10/02 08:24 tool set-results
0m Args:
{
  "KMSANReasoning": "The patch series modifies the ublk driver (drivers/block/ublk_drv.c) and adds selftests to address race conditions and synchronization issues during device cancellation and command completion (such as STOP_DEV racing with START_DEV, FETCH, and process exits). Specifically, it manages internal state flags (UB_STATE_STOPPING, ub-\u003ecanceling), file references (ub-\u003ech_file), memory barriers, and the timing of io_uring_cmd_mark_cancelable() and io_uring_cmd_done().\n\nAll modified kernel structures (struct ublk_device and struct ublk_batch_fetch_cmd) are allocated using zeroing allocators (kzalloc_flex and kzalloc_obj). Furthermore, INIT_LIST_HEAD(\u0026fcmd-\u003enode) is added to ensure proper initialization before cancellation checks. The patch does not introduce or modify any data copied to user space, does not alter buffer length or bounds calculations, and introduces no uninitialized memory reads or information leaks.\n\nAny memory safety or concurrency risks associated with these changes involve use-after-free, NULL pointer dereferences, or lock/barrier ordering, which are fully covered by standard KASAN and LOCKDEP. Therefore, a dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

Results:
{
  "KMSANReasoning": "The patch series modifies the ublk driver (drivers/block/ublk_drv.c) and adds selftests to address race conditions and synchronization issues during device cancellation and command completion (such as STOP_DEV racing with START_DEV, FETCH, and process exits). Specifically, it manages internal state flags (UB_STATE_STOPPING, ub-\u003ecanceling), file references (ub-\u003ech_file), memory barriers, and the timing of io_uring_cmd_mark_cancelable() and io_uring_cmd_done().\n\nAll modified kernel structures (struct ublk_device and struct ublk_batch_fetch_cmd) are allocated using zeroing allocators (kzalloc_flex and kzalloc_obj). Furthermore, INIT_LIST_HEAD(\u0026fcmd-\u003enode) is added to ensure proper initialization before cancellation checks. The patch does not introduce or modify any data copied to user space, does not alter buffer length or bounds calculations, and introduces no uninitialized memory reads or information leaks.\n\nAny memory safety or concurrency risks associated with these changes involve use-after-free, NULL pointer dereferences, or lock/barrier ordering, which are fully covered by standard KASAN and LOCKDEP. Therefore, a dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

LLM Calls Summary:
Total Calls Total Tokens Avg Tokens Total Duration (Seconds) Avg Duration (Seconds)
Tool Calls Summary:
Total Calls Total Duration (Seconds) Avg Duration (Seconds)