| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/10/06 13:26 | flow | patch-triage |
16mArgs:null Results: {
"EnableConfigs": [
"BLK_DEV_ZONED",
"BLK_DEV_ZONED_LOOP",
"BLK_DEV_SD"
],
"FocusSymbols": [
"blkdev_zone_storage_elements_ioctl",
"bdev_report_storage_elements",
"bdev_remove_storage_element",
"bdev_restore_storage_elements",
"bdev_zone_is_offline",
"zloop_remove_element",
"zloop_restore_elements",
"zloop_report_elements",
"zloop_ctl_degrade_element",
"sd_zbc_report_storage_elements",
"sd_zbc_remove_storage_element",
"sd_zbc_restore_storage_elements"
],
"KMSANReasoning": "The patch series introduces storage element management for zoned block devices across the block core, the SCSI ZBC driver (sd_zbc), and the zoned loop driver (zloop), adding ioctls BLKGETNRSTORELEMS, BLKREPORTSTORELEMS, BLKREMOVESTORELEM, and BLKRESTORESTORELEMS.\n\nRegarding KMSAN vs KASAN applicability:\n1. Heap allocations: All dynamically allocated structures, including the storage element arrays (e.g. `zlo-\u003eelements`, `elements` in `disk_wait_for_se_mgmt_completion` and `blkdev_report_storage_elements_ioctl`) and SCSI buffers (`buf` in `sd_zbc_report_storage_elements`), are allocated using `kzalloc_objs()` or `kzalloc()`, ensuring all bytes are zero-initialized upon allocation.\n2. Kernel-to-user data transfers:\n - In `blkdev_get_nr_storage_elements_ioctl()`, `nr_elements` is a scalar unsigned int initialized to 0 before being passed to `put_user()`.\n - In `blkdev_report_storage_elements_ioctl()`, `struct blk_storage_elements_report` consists of two `__u32` fields (8 bytes total, aligned to 8 bytes, zero padding) and is completely initialized from userspace via `copy_from_user()` before the count is updated and written back.\n - `struct blk_storage_element` has explicit members summing to exactly 24 bytes (4 + 4 + 8 + 1 + 1 + 1 + 5 bytes) with natural 8-byte alignment, leaving zero internal or tail padding bytes. Moreover, in `sd_zbc_parse_storage_element()`, each element descriptor is explicitly zeroed with `memset()` before parsing.\n3. Out-of-bounds risk: In `zloop_report_elements()`, `*nr_elements` is assigned `zlo-\u003enr_elements`, which in `blkdev_report_storage_elements_ioctl()` could cause `copy_to_user()` to copy more elements than the allocated `nr_elements` buffer if the caller requested fewer elements. However, this is a slab-out-of-bounds heap read, which is directly detectible by KASAN rather than an uninitialized memory defect.\n\nBecause all touched structures and buffers are zero-initialized and no uninitialized memory usage or info-leaks are introduced, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false,
"Reasoning": "The patch introduces new storage element management functionality for zoned block devices. It adds four new user-accessible block device ioctls (BLKGETNRSTORELEMS, BLKREPORTSTORELEMS, BLKREMOVESTORELEM, BLKRESTORESTORELEMS) in the core block layer, adds support in the scsi disk driver (sd_zbc), implements storage element emulation and a new control operation in the zoned loop driver (zloop), and introduces offline zone checks during bio submission. This reachable code path involves complex state transitions, synchronization (queue freeze/unfreeze, truncate, cache flush, delayed work), and memory management, making it an excellent target for syzkaller fuzzing.",
"WorthFuzzing": true
} |
| 1/1 | 2026/10/06 13:26 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit 1145222b274a2f5837f52a04a8d52b5f3f74f8a8\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Tue Oct 6 13:26:40 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/Documentation/admin-guide/blockdev/zoned_loop.rst b/Documentation/admin-guide/blockdev/zoned_loop.rst\nindex 64277494fb36a..ec39f141af8d6 100644\n--- a/Documentation/admin-guide/blockdev/zoned_loop.rst\n+++ b/Documentation/admin-guide/blockdev/zoned_loop.rst\n@@ -61,7 +61,7 @@ The options available for the add command can be listed by reading the\n /dev/zloop-control device::\n \n \t$ cat /dev/zloop-control\n- add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache\n+ add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache,stor_elements=%u\n remove id=%d\n \n In more details, the options that can be used with the \"add\" command are as\n@@ -113,6 +113,12 @@ discard_write_cache Discard all data that was not explicitly persisted using a\n each zone file to the size recorded during the last flush\n operation. This simulates power fail events where\n uncommitted data is lost.\n+stor_elements Control storage element emulation. The default value is 0,\n+ indicating no emulation. A value of 1 indicates that all\n+ access storage elements (equivalent to read+write head of\n+ a disk) are emulated. A value of 2 enables fractional\n+ access storage element (equivalent to pairs of read and\n+ write heads of a disk) emulation .\n =================== =========================================================\n \n 3) Deleting a Zoned Device\ndiff --git a/block/blk-core.c b/block/blk-core.c\nindex 13dc70e8f55d9..d420c80d2d938 100644\n--- a/block/blk-core.c\n+++ b/block/blk-core.c\n@@ -866,6 +866,9 @@ void submit_bio_noacct(struct bio *bio)\n \n \tswitch (bio_op(bio)) {\n \tcase REQ_OP_READ:\n+\t\tif (bdev_is_zoned(bdev) \u0026\u0026\n+\t\t bdev_zone_is_offline(bdev, bio-\u003ebi_iter.bi_sector))\n+\t\t\tgoto end_io;\n \t\tbreak;\n \tcase REQ_OP_WRITE:\n \t\tif (bio-\u003ebi_opf \u0026 REQ_ATOMIC) {\ndiff --git a/block/blk-zoned.c b/block/blk-zoned.c\nindex 19268afb8752e..6ebb04e0a1593 100644\n--- a/block/blk-zoned.c\n+++ b/block/blk-zoned.c\n@@ -18,6 +18,8 @@\n #include \u003clinux/mempool.h\u003e\n #include \u003clinux/kthread.h\u003e\n #include \u003clinux/freezer.h\u003e\n+#include \u003clinux/delay.h\u003e\n+#include \u003clinux/uaccess.h\u003e\n \n #include \u003ctrace/events/block.h\u003e\n \n@@ -308,6 +310,24 @@ bool bdev_zone_is_seq(struct block_device *bdev, sector_t sector)\n }\n EXPORT_SYMBOL_GPL(bdev_zone_is_seq);\n \n+/**\n+ * bdev_zone_is_offline - check if a sector belongs to an offline zone\n+ * @bdev: block device to check\n+ * @sector: sector number\n+ *\n+ * Check if @sector on @bdev is contained in an offline zone.\n+ */\n+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector)\n+{\n+\tenum blk_zone_cond cond;\n+\n+\tif (!bdev_is_zoned(bdev))\n+\t\treturn false;\n+\n+\tcond = disk_zone_get_cond(bdev-\u003ebd_disk, sector);\n+\treturn cond == BLK_ZONE_COND_OFFLINE;\n+}\n+\n /**\n * bdev_zone_mgmt_allowed - check if management operations are allowed on a zone\n * @bdev: block device to check\n@@ -2700,5 +2720,326 @@ int queue_zone_wplugs_show(void *data, struct seq_file *m)\n \n \treturn 0;\n }\n-\n #endif\n+\n+static int disk_wait_for_se_mgmt_completion(struct gendisk *disk)\n+{\n+\tstruct blk_storage_element *elements, *e;\n+\tunsigned int i, nr_se, nr_elements = 0;\n+\tint ret;\n+\n+\tret = disk-\u003efops-\u003ese_ops-\u003ereport_elements(disk, NULL, \u0026nr_elements);\n+\tif (ret) {\n+\t\tpr_err(\"Failed to get number of storage elements\\n\");\n+\t\treturn ret;\n+\t}\n+\n+\telements = kzalloc_objs(struct blk_storage_element, nr_elements);\n+\tif (!elements)\n+\t\treturn -ENOMEM;\n+\n+\twhile (1) {\n+\t\t/*\n+\t\t * Check if we have storage elements being removed or restored.\n+\t\t */\n+\t\tnr_se = nr_elements;\n+\t\tret = disk-\u003efops-\u003ese_ops-\u003ereport_elements(disk, elements,\n+\t\t\t\t\t\t\t \u0026nr_se);\n+\t\tif (ret) {\n+\t\t\tpr_err(\"Failed to get storage elements\\n\");\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\te = elements;\n+\t\tfor (i = 0; i \u003c nr_se; i++, e++) {\n+\t\t\tif (e-\u003estatus == BLK_SE_STS_REMOVE_IN_PROGRESS ||\n+\t\t\t e-\u003estatus == BLK_SE_STS_RESTORE_IN_PROGRESS)\n+\t\t\t\tbreak;\n+\t\t}\n+\t\tif (i \u003e= nr_se)\n+\t\t\tbreak;\n+\n+\t\t/* Not done yet: wait and retry. */\n+\t\tmsleep(500);\n+\t}\n+\n+\tkfree(elements);\n+\n+\treturn ret;\n+}\n+\n+/**\n+ * bdev_report_storage_elements - report the storage elements of a block device\n+ *\n+ * Fill at most @nr_elements storage element descriptors in the array @elements.\n+ * The number of storage elements filled in the array is returned using\n+ * @nr_elements. If @elements is NULL, only @nr_elements is returned.\n+ *\n+ * Returns 0 on success and a negative error code on failure.\n+ */\n+int bdev_report_storage_elements(struct block_device *bdev,\n+\t\t\t\t struct blk_storage_element *elements,\n+\t\t\t\t unsigned int *nr_elements)\n+{\n+\tstruct gendisk *disk = bdev-\u003ebd_disk;\n+\n+\tif (!bdev_is_zoned(bdev) || !disk-\u003efops-\u003ese_ops)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tif (!nr_elements)\n+\t\treturn -EINVAL;\n+\n+\tif (*nr_elements \u0026\u0026 !elements)\n+\t\treturn -EINVAL;\n+\n+\treturn disk-\u003efops-\u003ese_ops-\u003ereport_elements(disk, elements, nr_elements);\n+}\n+EXPORT_SYMBOL_GPL(bdev_report_storage_elements);\n+\n+static int blkdev_get_nr_storage_elements_ioctl(struct block_device *bdev,\n+\t\t\t\t\t\tvoid __user *argp)\n+{\n+\tunsigned int nr_elements = 0;\n+\tint ret;\n+\n+\tret = bdev_report_storage_elements(bdev, NULL, \u0026nr_elements);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\tif (put_user(nr_elements, (unsigned int __user *)argp))\n+\t\treturn -EFAULT;\n+\n+\treturn 0;\n+}\n+\n+static int blkdev_report_storage_elements_ioctl(struct block_device *bdev,\n+\t\t\t\t\t\tvoid __user *argp)\n+{\n+\tstruct blk_storage_elements_report rep;\n+\tstruct blk_storage_element *elements;\n+\tunsigned int nr_elements = 0;\n+\tunsigned long retc;\n+\tint ret;\n+\n+\tif (!argp)\n+\t\treturn -EINVAL;\n+\n+\tif (copy_from_user(\u0026rep, argp,\n+\t\t\t sizeof(struct blk_storage_elements_report)))\n+\t\treturn -EFAULT;\n+\n+\tret = bdev_report_storage_elements(bdev, NULL, \u0026nr_elements);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\tnr_elements = min(rep.nr_elements, nr_elements);\n+\tif (!nr_elements)\n+\t\treturn -EINVAL;\n+\n+\telements = kzalloc_objs(struct blk_storage_element, nr_elements);\n+\tif (!elements)\n+\t\treturn -ENOMEM;\n+\n+\tret = bdev_report_storage_elements(bdev, elements, \u0026nr_elements);\n+\tif (ret)\n+\t\tgoto free_elements;\n+\n+\tretc = copy_to_user(argp + sizeof(struct blk_storage_elements_report),\n+\t\t\t elements,\n+\t\t\t sizeof(struct blk_storage_element) * nr_elements);\n+\tif (retc) {\n+\t\tret = -EFAULT;\n+\t\tgoto free_elements;\n+\t}\n+\n+\trep.nr_elements = nr_elements;\n+\tretc = copy_to_user(argp, \u0026rep,\n+\t\t\t sizeof(struct blk_storage_elements_report));\n+\tif (retc)\n+\t\tret = -EFAULT;\n+\n+free_elements:\n+\tkfree(elements);\n+\treturn ret;\n+}\n+\n+/**\n+ * bdev_remove_storage_element - Remove (depopulate) a storage element of a\n+ *\t\t\t\t block device\n+ *\n+ * Remove (depopulate) the storage element identified by @element_id from the\n+ * block device @bdev. The caller is responsible for taking care of any\n+ * necessary device write cache flush and invalidation of cached data for the\n+ * zones that will be offlined.\n+ *\n+ * Returns 0 on success and a negative error code on failure.\n+ */\n+int bdev_remove_storage_element(struct block_device *bdev,\n+\t\t\t\tunsigned int element_id)\n+{\n+\tstruct gendisk *disk = bdev-\u003ebd_disk;\n+\tunsigned int memflags;\n+\tint ret;\n+\n+\tif (!bdev_is_zoned(bdev) || !disk-\u003efops-\u003ese_ops)\n+\t\treturn -EOPNOTSUPP;\n+\n+\t/* Zero is not a valid storage element ID. */\n+\tif (!element_id)\n+\t\treturn -EINVAL;\n+\n+\t/*\n+\t * Freeze and unfreeze the queue to flush any outstanding command.\n+\t * The caller is responsible for not queuing up more I/Os by higher\n+\t * level means.\n+\t */\n+\tmemflags = blk_mq_freeze_queue(disk-\u003equeue);\n+\tblk_mq_unfreeze_queue(disk-\u003equeue, memflags);\n+\n+\tret = disk-\u003efops-\u003ese_ops-\u003eremove_element(disk, element_id);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\t/* Revalidate the device zones once the opration completes. */\n+\tret = disk_wait_for_se_mgmt_completion(disk);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\treturn blk_revalidate_disk_zones(disk);\n+}\n+EXPORT_SYMBOL_GPL(bdev_remove_storage_element);\n+\n+static int blkdev_remove_storage_element_ioctl(struct block_device *bdev,\n+\t\t\t\t\tblk_mode_t mode, void __user *argp)\n+{\n+\tunsigned int element_id;\n+\tint ret;\n+\n+\tif (!(mode \u0026 BLK_OPEN_WRITE))\n+\t\treturn -EBADF;\n+\tif (bdev_read_only(bdev))\n+\t\treturn -EPERM;\n+\n+\tif (get_user(element_id, (unsigned int __user *)argp))\n+\t\treturn -EFAULT;\n+\n+\t/*\n+\t * Flush the device volatile write cache and invalidate all cached data\n+\t * so that reads do not return old data for zones that went offline.\n+\t */\n+\tinode_lock(bdev-\u003ebd_mapping-\u003ehost);\n+\tfilemap_invalidate_lock(bdev-\u003ebd_mapping);\n+\n+\tret = blkdev_issue_flush(bdev-\u003ebd_disk-\u003epart0);\n+\tif (ret)\n+\t\tgoto out_unlock;\n+\n+\tret = truncate_bdev_range(bdev, mode, 0,\n+\t\t\t(get_capacity(bdev-\u003ebd_disk) \u003c\u003c SECTOR_SHIFT) - 1);\n+\tif (ret)\n+\t\tgoto out_unlock;\n+\n+\tret = bdev_remove_storage_element(bdev, element_id);\n+\n+out_unlock:\n+\tfilemap_invalidate_unlock(bdev-\u003ebd_mapping);\n+\tinode_unlock(bdev-\u003ebd_mapping-\u003ehost);\n+\n+\treturn ret;\n+}\n+\n+/**\n+ * bdev_restore_storage_elements - Restore all depopulated storage elements of a\n+ *\t\t\t\t block device\n+ *\n+ * Restore all storage elements of @bdev that have been depopulated. Not all\n+ * elements may be restored by this operation.\n+ *\n+ * Returns 0 on success and a negative error code on failure.\n+ */\n+int bdev_restore_storage_elements(struct block_device *bdev)\n+{\n+\tstruct gendisk *disk = bdev-\u003ebd_disk;\n+\tunsigned int memflags;\n+\tint ret;\n+\n+\tif (!bdev_is_zoned(bdev) || !disk-\u003efops-\u003ese_ops)\n+\t\treturn -EOPNOTSUPP;\n+\n+\t/*\n+\t * Freeze and unfreeze the queue to flush any outstanding commands.\n+\t * The caller is responsible for not queuing up more I/O by higher\n+\t * level means.\n+\t */\n+\tmemflags = blk_mq_freeze_queue(disk-\u003equeue);\n+\tblk_mq_unfreeze_queue(disk-\u003equeue, memflags);\n+\n+\tret = disk-\u003efops-\u003ese_ops-\u003erestore_elements(disk);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\tret = disk_wait_for_se_mgmt_completion(disk);\n+\tif (ret)\n+\t\treturn ret;\n+\n+\treturn blk_revalidate_disk_zones(disk);\n+}\n+EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);\n+\n+static int blkdev_restore_storage_elements_ioctl(struct block_device *bdev,\n+\t\t\t\t\t\t blk_mode_t mode)\n+{\n+\tint ret;\n+\n+\tif (!(mode \u0026 BLK_OPEN_WRITE))\n+\t\treturn -EBADF;\n+\tif (bdev_read_only(bdev))\n+\t\treturn -EPERM;\n+\n+\t/*\n+\t * Flush the device volatile write cache and invalidate all cached data\n+\t * so that reads do not return old data for zones that went offline.\n+\t */\n+\tinode_lock(bdev-\u003ebd_mapping-\u003ehost);\n+\tfilemap_invalidate_lock(bdev-\u003ebd_mapping);\n+\n+\tret = blkdev_issue_flush(bdev-\u003ebd_disk-\u003epart0);\n+\tif (ret)\n+\t\tgoto out_unlock;\n+\n+\tret = truncate_bdev_range(bdev, mode, 0,\n+\t\t\t(get_capacity(bdev-\u003ebd_disk) \u003c\u003c SECTOR_SHIFT) - 1);\n+\tif (ret)\n+\t\tgoto out_unlock;\n+\n+\tret = bdev_restore_storage_elements(bdev);\n+\n+out_unlock:\n+\tfilemap_invalidate_unlock(bdev-\u003ebd_mapping);\n+\tinode_unlock(bdev-\u003ebd_mapping-\u003ehost);\n+\n+\treturn ret;\n+}\n+\n+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,\n+\t\t\t\t blk_mode_t mode, unsigned int cmd,\n+\t\t\t\t void __user *argp)\n+{\n+\tif (!bdev_is_zoned(bdev) || !bdev-\u003ebd_disk-\u003efops-\u003ese_ops)\n+\t\treturn -ENOTTY;\n+\n+\tswitch (cmd) {\n+\tcase BLKGETNRSTORELEMS:\n+\t\treturn blkdev_get_nr_storage_elements_ioctl(bdev, argp);\n+\tcase BLKREPORTSTORELEMS:\n+\t\treturn blkdev_report_storage_elements_ioctl(bdev, argp);\n+\tcase BLKREMOVESTORELEM:\n+\t\treturn blkdev_remove_storage_element_ioctl(bdev, mode, argp);\n+\tcase BLKRESTORESTORELEMS:\n+\t\treturn blkdev_restore_storage_elements_ioctl(bdev, mode);\n+\tdefault:\n+\t\tbreak;\n+\t}\n+\n+\treturn -ENOTTY;\n+}\ndiff --git a/block/blk.h b/block/blk.h\nindex 2cc03aa54c532..274afb46a809a 100644\n--- a/block/blk.h\n+++ b/block/blk.h\n@@ -578,6 +578,10 @@ int blkdev_report_zones_ioctl(struct block_device *bdev, unsigned int cmd,\n int blkdev_zone_mgmt_ioctl(struct block_device *bdev, blk_mode_t mode,\n \t\tunsigned int cmd, unsigned long arg);\n bool bdev_zone_mgmt_allowed(struct block_device *bdev, sector_t sector);\n+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector);\n+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,\n+\t\t\t\t blk_mode_t mode, unsigned int cmd,\n+\t\t\t\t void __user *argp);\n #else /* CONFIG_BLK_DEV_ZONED */\n static inline void disk_init_zone_resources(struct gendisk *disk)\n {\n@@ -625,6 +629,17 @@ static inline bool bdev_zone_mgmt_allowed(struct block_device *bdev,\n {\n \treturn false;\n }\n+static inline bool bdev_zone_is_offline(struct block_device *bdev,\n+\t\t\t\t\tsector_t sector)\n+{\n+\treturn false;\n+}\n+static inline int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,\n+\t\t\t\t\tblk_mode_t mode, unsigned int cmd,\n+\t\t\t\t\tvoid __user *argp)\n+{\n+\treturn -ENOTTY;\n+}\n #endif /* CONFIG_BLK_DEV_ZONED */\n \n struct block_device *bdev_alloc(struct gendisk *disk, u8 partno);\ndiff --git a/block/ioctl.c b/block/ioctl.c\nindex 64b4e6c0f6965..ccf806c2e37d1 100644\n--- a/block/ioctl.c\n+++ b/block/ioctl.c\n@@ -679,6 +679,11 @@ static int blkdev_common_ioctl(struct block_device *bdev, blk_mode_t mode,\n \t\treturn put_uint(argp, bdev_zone_sectors(bdev));\n \tcase BLKGETNRZONES:\n \t\treturn put_uint(argp, bdev_nr_zones(bdev));\n+\tcase BLKGETNRSTORELEMS:\n+\tcase BLKREPORTSTORELEMS:\n+\tcase BLKREMOVESTORELEM:\n+\tcase BLKRESTORESTORELEMS:\n+\t\treturn blkdev_zone_storage_elements_ioctl(bdev, mode, cmd, argp);\n \tcase BLKROGET:\n \t\treturn put_int(argp, bdev_read_only(bdev) != 0);\n \tcase BLKSSZGET: /* get block device logical block size */\ndiff --git a/drivers/block/zloop.c b/drivers/block/zloop.c\nindex 394dc2408ed73..3913bd2299108 100644\n--- a/drivers/block/zloop.c\n+++ b/drivers/block/zloop.c\n@@ -37,6 +37,8 @@ enum {\n \tZLOOP_OPT_ORDERED_ZONE_APPEND\t= (1 \u003c\u003c 10),\n \tZLOOP_OPT_DISCARD_WRITE_CACHE\t= (1 \u003c\u003c 11),\n \tZLOOP_OPT_MAX_OPEN_ZONES\t= (1 \u003c\u003c 12),\n+\tZLOOP_OPT_STOR_ELEMENTS\t\t= (1 \u003c\u003c 13),\n+\tZLOOP_OPT_ELEMENT_ID\t\t= (1 \u003c\u003c 14),\n };\n \n static const match_table_t zloop_opt_tokens = {\n@@ -51,11 +53,23 @@ static const match_table_t zloop_opt_tokens = {\n \t{ ZLOOP_OPT_BUFFERED_IO,\t\"buffered_io\"\t\t},\n \t{ ZLOOP_OPT_ZONE_APPEND,\t\"zone_append=%u\"\t},\n \t{ ZLOOP_OPT_ORDERED_ZONE_APPEND, \"ordered_zone_append\"\t},\n-\t{ ZLOOP_OPT_DISCARD_WRITE_CACHE, \"discard_write_cache\" },\n+\t{ ZLOOP_OPT_DISCARD_WRITE_CACHE, \"discard_write_cache\"\t},\n \t{ ZLOOP_OPT_MAX_OPEN_ZONES,\t\"max_open_zones=%u\"\t},\n+\t{ ZLOOP_OPT_STOR_ELEMENTS,\t\"stor_elements=%u\"\t},\n+\t{ ZLOOP_OPT_ELEMENT_ID,\t\t\"element_id=%u\"\t\t},\n \t{ ZLOOP_OPT_ERR,\t\tNULL\t\t\t}\n };\n \n+/* Storage elements emulation types. */\n+enum zloop_stor_elements {\n+\t/* No emulation. */\n+\tZLOOP_STOR_ELEMENTS_NONE,\n+\t/* Emulate read+write storage elements. */\n+\tZLOOP_STOR_ELEMENTS_RDWR,\n+\t/* Emulate pairs of associated read and write storage elements. */\n+\tZLOOP_STOR_ELEMENTS_PAIRS,\n+};\n+\n /* Default values for the \"add\" operation. */\n #define ZLOOP_DEF_ID\t\t\t-1\n #define ZLOOP_DEF_ZONE_SIZE\t\t((256ULL * SZ_1M) \u003e\u003e SECTOR_SHIFT)\n@@ -68,6 +82,8 @@ static const match_table_t zloop_opt_tokens = {\n #define ZLOOP_DEF_BUFFERED_IO\t\tfalse\n #define ZLOOP_DEF_ZONE_APPEND\t\ttrue\n #define ZLOOP_DEF_ORDERED_ZONE_APPEND\tfalse\n+#define ZLOOP_DEF_STOR_ELEMENTS\t\tZLOOP_STOR_ELEMENTS_NONE\n+#define ZLOOP_DEF_ELEMENT_ID\t\t0\n \n /* Arbitrary limit on the zone size (16GB). */\n #define ZLOOP_MAX_ZONE_SIZE_MB\t\t16384\n@@ -87,6 +103,8 @@ struct zloop_options {\n \tbool\t\t\tzone_append;\n \tbool\t\t\tordered_zone_append;\n \tbool\t\t\tdiscard_write_cache;\n+\tenum zloop_stor_elements stor_elements;\n+\tunsigned int\t\telement_id;\n };\n \n /*\n@@ -117,6 +135,8 @@ struct zloop_zone {\n \tenum blk_zone_cond\tcond;\n \tsector_t\t\tstart;\n \tsector_t\t\twp;\n+\tunsigned int\t\twr_se_id;\n+\tunsigned int\t\trd_se_id;\n \n \tgfp_t\t\t\told_gfp_mask;\n };\n@@ -133,6 +153,7 @@ struct zloop_device {\n \tbool\t\t\tzone_append;\n \tbool\t\t\tordered_zone_append;\n \tbool\t\t\tdiscard_write_cache;\n+\tenum zloop_stor_elements stor_elements;\n \n \tconst char\t\t*base_dir;\n \tstruct file\t\t*data_dir;\n@@ -150,6 +171,17 @@ struct zloop_device {\n \tstruct list_head\topen_zones_lru_list;\n \tunsigned int\t\tnr_open_zones;\n \n+\t/* For storage elements emulation. */\n+\tstruct mutex\t\tstor_elements_lock;\n+\tunsigned int\t\tnr_elements;\n+\tunsigned int\t\tmax_nr_removed_elements;\n+\tunsigned int\t\tnr_removed_elements;\n+\tstruct delayed_work\tremove_element_work;\n+\tunsigned int\t\tremove_element_id;\n+\tstruct delayed_work\trestore_elements_work;\n+\tbool\t\t\trestore_in_progress;\n+\tstruct blk_storage_element *elements;\n+\n \tstruct zloop_zone\tzones[] __counted_by(nr_zones);\n };\n \n@@ -289,22 +321,30 @@ static bool zloop_do_open_zone(struct zloop_device *zlo,\n \t}\n }\n \n-static void zloop_mark_full(struct zloop_device *zlo, struct zloop_zone *zone)\n+static void zloop_set_zone_cond(struct zloop_device *zlo,\n+\t\t\t\tstruct zloop_zone *zone,\n+\t\t\t\tenum blk_zone_cond cond)\n {\n \tlockdep_assert_held(\u0026zone-\u003ewp_lock);\n \n \tzloop_lru_remove_open_zone(zlo, zone);\n-\tzone-\u003econd = BLK_ZONE_COND_FULL;\n-\tzone-\u003ewp = ULLONG_MAX;\n+\tzone-\u003econd = cond;\n+\tif (cond == BLK_ZONE_COND_EMPTY)\n+\t\tzone-\u003ewp = zone-\u003estart;\n+\telse\n+\t\tzone-\u003ewp = ULLONG_MAX;\n }\n \n-static void zloop_mark_empty(struct zloop_device *zlo, struct zloop_zone *zone)\n+static inline void zloop_set_zone_full(struct zloop_device *zlo,\n+\t\t\t\t struct zloop_zone *zone)\n {\n-\tlockdep_assert_held(\u0026zone-\u003ewp_lock);\n+\tzloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_FULL);\n+}\n \n-\tzloop_lru_remove_open_zone(zlo, zone);\n-\tzone-\u003econd = BLK_ZONE_COND_EMPTY;\n-\tzone-\u003ewp = zone-\u003estart;\n+static inline void zloop_set_zone_empty(struct zloop_device *zlo,\n+\t\t\t\t\tstruct zloop_zone *zone)\n+{\n+\tzloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_EMPTY);\n }\n \n static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)\n@@ -339,9 +379,9 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)\n \n \tspin_lock(\u0026zone-\u003ewp_lock);\n \tif (!file_sectors) {\n-\t\tzloop_mark_empty(zlo, zone);\n+\t\tzloop_set_zone_empty(zlo, zone);\n \t} else if (file_sectors == zlo-\u003ezone_capacity) {\n-\t\tzloop_mark_full(zlo, zone);\n+\t\tzloop_set_zone_full(zlo, zone);\n \t} else {\n \t\tif (zone-\u003econd != BLK_ZONE_COND_IMP_OPEN \u0026\u0026\n \t\t zone-\u003econd != BLK_ZONE_COND_EXP_OPEN)\n@@ -353,6 +393,19 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)\n \treturn 0;\n }\n \n+static bool zloop_zone_is_offline_or_readonly(struct zloop_device *zlo,\n+\t\t\t\t\t struct zloop_zone *zone)\n+{\n+\tbool ret;\n+\n+\tspin_lock(\u0026zone-\u003ewp_lock);\n+\tret = zone-\u003econd == BLK_ZONE_COND_OFFLINE ||\n+\t\tzone-\u003econd == BLK_ZONE_COND_READONLY;\n+\tspin_unlock(\u0026zone-\u003ewp_lock);\n+\n+\treturn ret;\n+}\n+\n static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)\n {\n \tstruct zloop_zone *zone = \u0026zlo-\u003ezones[zone_no];\n@@ -363,6 +416,11 @@ static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)\n \n \tmutex_lock(\u0026zone-\u003elock);\n \n+\tif (zloop_zone_is_offline_or_readonly(zlo, zone)) {\n+\t\tret = -EIO;\n+\t\tgoto unlock;\n+\t}\n+\n \tif (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags)) {\n \t\tret = zloop_update_seq_zone(zlo, zone_no);\n \t\tif (ret)\n@@ -388,6 +446,11 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)\n \n \tmutex_lock(\u0026zone-\u003elock);\n \n+\tif (zloop_zone_is_offline_or_readonly(zlo, zone)) {\n+\t\tret = -EIO;\n+\t\tgoto unlock;\n+\t}\n+\n \tif (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags)) {\n \t\tret = zloop_update_seq_zone(zlo, zone_no);\n \t\tif (ret)\n@@ -420,7 +483,22 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)\n \treturn ret;\n }\n \n-static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)\n+static int zloop_do_reset_zone(struct zloop_device *zlo,\n+\t\t\t struct zloop_zone *zone)\n+{\n+\tif (vfs_truncate(\u0026zone-\u003efile-\u003ef_path, 0))\n+\t\treturn -EIO;\n+\n+\tspin_lock(\u0026zone-\u003ewp_lock);\n+\tzloop_set_zone_empty(zlo, zone);\n+\tclear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\n+\tspin_unlock(\u0026zone-\u003ewp_lock);\n+\n+\treturn 0;\n+}\n+\n+static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no,\n+\t\t\t bool all_zones)\n {\n \tstruct zloop_zone *zone = \u0026zlo-\u003ezones[zone_no];\n \tint ret = 0;\n@@ -430,20 +508,19 @@ static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)\n \n \tmutex_lock(\u0026zone-\u003elock);\n \n+\tif (zloop_zone_is_offline_or_readonly(zlo, zone)) {\n+\t\tif (!all_zones)\n+\t\t\tret = -EIO;\n+\t\tgoto unlock;\n+\t}\n+\n \tif (!test_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags) \u0026\u0026\n \t zone-\u003econd == BLK_ZONE_COND_EMPTY)\n \t\tgoto unlock;\n \n-\tif (vfs_truncate(\u0026zone-\u003efile-\u003ef_path, 0)) {\n+\tret = zloop_do_reset_zone(zlo, zone);\n+\tif (ret)\n \t\tset_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\n-\t\tret = -EIO;\n-\t\tgoto unlock;\n-\t}\n-\n-\tspin_lock(\u0026zone-\u003ewp_lock);\n-\tzloop_mark_empty(zlo, zone);\n-\tclear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\n-\tspin_unlock(\u0026zone-\u003ewp_lock);\n \n unlock:\n \tmutex_unlock(\u0026zone-\u003elock);\n@@ -457,7 +534,7 @@ static int zloop_reset_all_zones(struct zloop_device *zlo)\n \tint ret;\n \n \tfor (i = zlo-\u003enr_conv_zones; i \u003c zlo-\u003enr_zones; i++) {\n-\t\tret = zloop_reset_zone(zlo, i);\n+\t\tret = zloop_reset_zone(zlo, i, true);\n \t\tif (ret)\n \t\t\treturn ret;\n \t}\n@@ -475,6 +552,11 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)\n \n \tmutex_lock(\u0026zone-\u003elock);\n \n+\tif (zloop_zone_is_offline_or_readonly(zlo, zone)) {\n+\t\tret = -EIO;\n+\t\tgoto unlock;\n+\t}\n+\n \tif (!test_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags) \u0026\u0026\n \t zone-\u003econd == BLK_ZONE_COND_FULL)\n \t\tgoto unlock;\n@@ -487,7 +569,7 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)\n \t}\n \n \tspin_lock(\u0026zone-\u003ewp_lock);\n-\tzloop_mark_full(zlo, zone);\n+\tzloop_set_zone_full(zlo, zone);\n \tclear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\n \tspin_unlock(\u0026zone-\u003ewp_lock);\n \n@@ -624,7 +706,7 @@ static int zloop_seq_write_prep(struct zloop_cmd *cmd)\n \tif (!is_append || !zlo-\u003eordered_zone_append) {\n \t\tzone-\u003ewp += nr_sectors;\n \t\tif (zone-\u003ewp == zone_end)\n-\t\t\tzloop_mark_full(zlo, zone);\n+\t\t\tzloop_set_zone_full(zlo, zone);\n \t}\n out_unlock:\n \tspin_unlock(\u0026zone-\u003ewp_lock);\n@@ -766,7 +848,7 @@ static void zloop_handle_cmd(struct zloop_cmd *cmd)\n \t\tcmd-\u003eret = zloop_flush(zlo);\n \t\tbreak;\n \tcase REQ_OP_ZONE_RESET:\n-\t\tcmd-\u003eret = zloop_reset_zone(zlo, rq_zone_no(rq));\n+\t\tcmd-\u003eret = zloop_reset_zone(zlo, rq_zone_no(rq), false);\n \t\tbreak;\n \tcase REQ_OP_ZONE_RESET_ALL:\n \t\tcmd-\u003eret = zloop_reset_all_zones(zlo);\n@@ -875,30 +957,69 @@ static void zloop_complete_rq(struct request *rq)\n \tblk_mq_end_request(rq, sts);\n }\n \n-static bool zloop_set_zone_append_sector(struct request *rq)\n+static bool zloop_set_zone_append_sector(struct zloop_device *zlo,\n+\t\t\t\t\t struct zloop_zone *zone,\n+\t\t\t\t\t struct request *rq)\n {\n-\tstruct zloop_device *zlo = rq-\u003eq-\u003equeuedata;\n-\tunsigned int zone_no = rq_zone_no(rq);\n-\tstruct zloop_zone *zone = \u0026zlo-\u003ezones[zone_no];\n \tsector_t zone_end = zone-\u003estart + zlo-\u003ezone_capacity;\n \tsector_t nr_sectors = blk_rq_sectors(rq);\n \n-\tspin_lock(\u0026zone-\u003ewp_lock);\n-\n \tif (zone-\u003econd == BLK_ZONE_COND_FULL ||\n-\t zone-\u003ewp + nr_sectors \u003e zone_end) {\n-\t\tspin_unlock(\u0026zone-\u003ewp_lock);\n+\t zone-\u003ewp + nr_sectors \u003e zone_end)\n \t\treturn false;\n-\t}\n \n \trq-\u003e__sector = zone-\u003ewp;\n \tzone-\u003ewp += blk_rq_sectors(rq);\n \tif (zone-\u003ewp \u003e= zone_end)\n-\t\tzloop_mark_full(zlo, zone);\n+\t\tzloop_set_zone_full(zlo, zone);\n+\n+\treturn true;\n+}\n+\n+\n+static bool zloop_prep_rq(struct zloop_device *zlo, struct request *rq)\n+{\n+\tstruct zloop_zone *zone = \u0026zlo-\u003ezones[rq_zone_no(rq)];\n+\tbool is_write = op_is_write(req_op(rq));\n+\tbool ret = true;\n+\n+\tspin_lock(\u0026zone-\u003ewp_lock);\n+\n+\tif (zlo-\u003enr_elements) {\n+\t\tstruct blk_storage_element *se;\n+\n+\t\tif (zone-\u003econd == BLK_ZONE_COND_OFFLINE ||\n+\t\t (zone-\u003econd == BLK_ZONE_COND_READONLY \u0026\u0026 is_write)) {\n+\t\t\tret = false;\n+\t\t\tgoto unlock;\n+\t\t}\n+\n+\t\t/*\n+\t\t * Check the health state of the storage element serving the\n+\t\t * zone.\n+\t\t */\n+\t\tif (is_write)\n+\t\t\tse = \u0026zlo-\u003eelements[zone-\u003ewr_se_id - 1];\n+\t\telse\n+\t\t\tse = \u0026zlo-\u003eelements[zone-\u003erd_se_id - 1];\n+\t\tif (READ_ONCE(se-\u003estatus) == BLK_SE_STS_DEGRADED) {\n+\t\t\tret = false;\n+\t\t\tgoto unlock;\n+\t\t}\n+\t}\n+\n+\t/*\n+\t * If we need to strongly order zone append operations, set the request\n+\t * sector to the zone write pointer location now instead of when the\n+\t * command work runs.\n+\t */\n+\tif (zlo-\u003eordered_zone_append \u0026\u0026 req_op(rq) == REQ_OP_ZONE_APPEND)\n+\t\tret = zloop_set_zone_append_sector(zlo, zone, rq);\n \n+unlock:\n \tspin_unlock(\u0026zone-\u003ewp_lock);\n \n-\treturn true;\n+\treturn ret;\n }\n \n static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,\n@@ -913,14 +1034,15 @@ static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,\n \t\treturn BLK_STS_IOERR;\n \t}\n \n-\t/*\n-\t * If we need to strongly order zone append operations, set the request\n-\t * sector to the zone write pointer location now instead of when the\n-\t * command work runs.\n-\t */\n-\tif (zlo-\u003eordered_zone_append \u0026\u0026 req_op(rq) == REQ_OP_ZONE_APPEND) {\n-\t\tif (!zloop_set_zone_append_sector(rq))\n+\tswitch (req_op(rq)) {\n+\tcase REQ_OP_READ:\n+\tcase REQ_OP_WRITE:\n+\tcase REQ_OP_ZONE_APPEND:\n+\t\tif (!zloop_prep_rq(zlo, rq))\n \t\t\treturn BLK_STS_IOERR;\n+\t\tbreak;\n+\tdefault:\n+\t\tbreak;\n \t}\n \n \tblk_mq_start_request(rq);\n@@ -1002,11 +1124,277 @@ static int zloop_report_zones(struct gendisk *disk, sector_t sector,\n \treturn nr_zones;\n }\n \n+static int zloop_report_elements(struct gendisk *disk,\n+\t\t\t\t struct blk_storage_element *elements,\n+\t\t\t\t unsigned int *nr_elements)\n+{\n+\tstruct zloop_device *zlo = disk-\u003eprivate_data;\n+\tunsigned int nr_report = *nr_elements;\n+\n+\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_NONE)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tmutex_lock(\u0026zlo-\u003estor_elements_lock);\n+\n+\t*nr_elements = zlo-\u003enr_elements;\n+\tif (elements) {\n+\t\tstruct blk_storage_element *se = zlo-\u003eelements;\n+\t\tunsigned int i;\n+\n+\t\tfor (i = 0; i \u003c min(nr_report, zlo-\u003enr_elements); i++, se++) {\n+\t\t\tswitch (READ_ONCE(se-\u003estatus)) {\n+\t\t\tcase BLK_SE_STS_REMOVED:\n+\t\t\tcase BLK_SE_STS_RESTORE_ERROR:\n+\t\t\t\tse-\u003erestore_allowed = 1;\n+\t\t\t\tbreak;\n+\t\t\tdefault:\n+\t\t\t\tse-\u003erestore_allowed = 0;\n+\t\t\t}\n+\t\t\tmemcpy(\u0026elements[i], se,\n+\t\t\t sizeof(struct blk_storage_element));\n+\t\t}\n+\t}\n+\n+\tmutex_unlock(\u0026zlo-\u003estor_elements_lock);\n+\n+\treturn 0;\n+}\n+\n+static int zloop_degrade_element(struct zloop_device *zlo,\n+\t\t\t\t unsigned int element_id)\n+{\n+\tstruct blk_storage_element *se;\n+\tint ret = 0;\n+\n+\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_NONE)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tif (!element_id || element_id \u003e zlo-\u003enr_elements)\n+\t\treturn -EINVAL;\n+\n+\tmutex_lock(\u0026zlo-\u003estor_elements_lock);\n+\n+\tse = \u0026zlo-\u003eelements[element_id - 1];\n+\tif (READ_ONCE(se-\u003estatus) == BLK_SE_STS_OK)\n+\t\tWRITE_ONCE(se-\u003estatus, BLK_SE_STS_DEGRADED);\n+\telse\n+\t\tret = -EINVAL;\n+\n+\tmutex_unlock(\u0026zlo-\u003estor_elements_lock);\n+\n+\treturn ret;\n+}\n+\n+static void zloop_remove_element_work(struct work_struct *work)\n+{\n+\tstruct zloop_device *zlo = container_of(work, struct zloop_device,\n+\t\t\t\t\t\tremove_element_work.work);\n+\tstruct blk_storage_element *se, *paired_se = NULL;\n+\tenum blk_zone_cond cond;\n+\tstruct zloop_zone *zone;\n+\tunsigned int i;\n+\n+\tmutex_lock(\u0026zlo-\u003estor_elements_lock);\n+\n+\t/*\n+\t * If the element to remove is a read element, zones must go offline and\n+\t * the associated write element is also removed.\n+\t */\n+\tse = \u0026zlo-\u003eelements[zlo-\u003eremove_element_id - 1];\n+\tswitch (se-\u003etype) {\n+\tcase BLK_SE_TYPE_RDWR:\n+\t\tcond = BLK_ZONE_COND_OFFLINE;\n+\t\tbreak;\n+\tcase BLK_SE_TYPE_READ:\n+\t\tcond = BLK_ZONE_COND_OFFLINE;\n+\t\tpaired_se = \u0026zlo-\u003eelements[se-\u003epaired_id - 1];\n+\t\tbreak;\n+\tcase BLK_SE_TYPE_WRITE:\n+\t\tcond = BLK_ZONE_COND_READONLY;\n+\t\tbreak;\n+\tdefault:\n+\t\tWARN_ON_ONCE(1);\n+\t}\n+\n+\t/*\n+\t * Change the condition of the zones owned by the (pair of) elements\n+\t * being removed and mark the elements removed.\n+\t */\n+\tfor (i = 0, zone = zlo-\u003ezones; i \u003c zlo-\u003enr_zones; i++, zone++) {\n+\t\tif (zone-\u003ewr_se_id != se-\u003eid \u0026\u0026 zone-\u003erd_se_id != se-\u003eid)\n+\t\t\tcontinue;\n+\t\tmutex_lock(\u0026zone-\u003elock);\n+\t\tspin_lock(\u0026zone-\u003ewp_lock);\n+\t\tzloop_set_zone_cond(zlo, zone, cond);\n+\t\tspin_unlock(\u0026zone-\u003ewp_lock);\n+\t\tmutex_unlock(\u0026zone-\u003elock);\n+\t}\n+\n+\tWRITE_ONCE(se-\u003estatus, BLK_SE_STS_REMOVED);\n+\n+\tzlo-\u003enr_removed_elements++;\n+\tif (paired_se \u0026\u0026 READ_ONCE(paired_se-\u003estatus) != BLK_SE_STS_REMOVED) {\n+\t\tWRITE_ONCE(paired_se-\u003estatus, BLK_SE_STS_REMOVED);\n+\t\tzlo-\u003enr_removed_elements++;\n+\t}\n+\n+\tzlo-\u003eremove_element_id = 0;\n+\n+\tmutex_unlock(\u0026zlo-\u003estor_elements_lock);\n+}\n+\n+static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)\n+{\n+\tstruct zloop_device *zlo = disk-\u003eprivate_data;\n+\tstruct blk_storage_element *se, *paired_se = NULL;\n+\tunsigned int nr_remove = 1;\n+\tint ret = 0;\n+\n+\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_NONE)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tif (element_id \u003e zlo-\u003enr_elements)\n+\t\treturn -EINVAL;\n+\n+\tmutex_lock(\u0026zlo-\u003estor_elements_lock);\n+\n+\tif (zlo-\u003eremove_element_id || zlo-\u003erestore_in_progress) {\n+\t\tret = -EBUSY;\n+\t\tgoto unlock;\n+\t}\n+\n+\t/*\n+\t * Get the element to remove. If it is already removed, we have nothing\n+\t * to do.\n+\t */\n+\tse = \u0026zlo-\u003eelements[element_id - 1];\n+\tif (READ_ONCE(se-\u003estatus) == BLK_SE_STS_REMOVED)\n+\t\tgoto unlock;\n+\n+\t/*\n+\t * If the element to remove is a read element, the associated write\n+\t * element must also be removed.\n+\t */\n+\tif (se-\u003etype == BLK_SE_TYPE_READ) {\n+\t\tpaired_se = \u0026zlo-\u003eelements[se-\u003epaired_id - 1];\n+\t\tif (READ_ONCE(paired_se-\u003estatus) != BLK_SE_STS_REMOVED)\n+\t\t\tnr_remove = 2;\n+\t}\n+\n+\tif (zlo-\u003enr_removed_elements + nr_remove \u003e\n+\t zlo-\u003emax_nr_removed_elements) {\n+\t\tret = -EBUSY;\n+\t\tgoto unlock;\n+\t}\n+\n+\t/*\n+\t * Schedule the element removal with a delay, to emulate the (generally\n+\t * short) time it takes for a real device to depopulate a head and\n+\t * modify the zones.\n+\t */\n+\tzlo-\u003eremove_element_id = element_id;\n+\tWRITE_ONCE(se-\u003estatus, BLK_SE_STS_REMOVE_IN_PROGRESS);\n+\tif (paired_se \u0026\u0026 READ_ONCE(paired_se-\u003estatus) != BLK_SE_STS_REMOVED)\n+\t\tWRITE_ONCE(paired_se-\u003estatus, BLK_SE_STS_REMOVE_IN_PROGRESS);\n+\n+\tschedule_delayed_work(\u0026zlo-\u003eremove_element_work,\n+\t\t\t msecs_to_jiffies(2000));\n+\n+unlock:\n+\tmutex_unlock(\u0026zlo-\u003estor_elements_lock);\n+\n+\treturn ret;\n+}\n+\n+static void zloop_restore_elements_work(struct work_struct *work)\n+{\n+\tstruct zloop_device *zlo = container_of(work, struct zloop_device,\n+\t\t\t\t\t\trestore_elements_work.work);\n+\tstruct zloop_zone *zone = zlo-\u003ezones;\n+\tstruct blk_storage_element *se;\n+\tunsigned int i;\n+\tint ret = 0;\n+\n+\tmutex_lock(\u0026zlo-\u003estor_elements_lock);\n+\n+\t/* Reset all zones. */\n+\tfor (i = 0; i \u003c zlo-\u003enr_zones \u0026\u0026 ret == 0; i++, zone++) {\n+\t\tmutex_lock(\u0026zone-\u003elock);\n+\t\tif (test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags)) {\n+\t\t\tspin_lock(\u0026zone-\u003ewp_lock);\n+\t\t\tzloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_NOT_WP);\n+\t\t\tspin_unlock(\u0026zone-\u003ewp_lock);\n+\t\t} else {\n+\t\t\tret = zloop_do_reset_zone(zlo, zone);\n+\t\t}\n+\t\tmutex_unlock(\u0026zone-\u003elock);\n+\t}\n+\n+\t/* Restore all removed elements. */\n+\tfor (i = 0, se = zlo-\u003eelements; i \u003c zlo-\u003enr_elements; i++, se++) {\n+\t\tif (READ_ONCE(se-\u003estatus) != BLK_SE_STS_RESTORE_IN_PROGRESS)\n+\t\t\tcontinue;\n+\t\tif (!ret)\n+\t\t\tWRITE_ONCE(se-\u003estatus, BLK_SE_STS_OK);\n+\t\telse\n+\t\t\tWRITE_ONCE(se-\u003estatus, BLK_SE_STS_RESTORE_ERROR);\n+\t}\n+\n+\tzlo-\u003enr_removed_elements = 0;\n+\tzlo-\u003erestore_in_progress = false;\n+\n+\tmutex_unlock(\u0026zlo-\u003estor_elements_lock);\n+}\n+\n+static int zloop_restore_elements(struct gendisk *disk)\n+{\n+\tstruct zloop_device *zlo = disk-\u003eprivate_data;\n+\tstruct blk_storage_element *se;\n+\tunsigned int i;\n+\tint ret = 0;\n+\n+\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_NONE)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tmutex_lock(\u0026zlo-\u003estor_elements_lock);\n+\n+\tif (zlo-\u003eremove_element_id || zlo-\u003erestore_in_progress) {\n+\t\tret = -EBUSY;\n+\t\tgoto unlock;\n+\t}\n+\n+\t/* If we have no removed elements, we have nothing to do. */\n+\tif (!zlo-\u003enr_removed_elements)\n+\t\tgoto unlock;\n+\n+\t/*\n+\t * Schedule the elements restoration with a delay, to emulate the time\n+\t * it takes for a real device to restore all removed heads and reset\n+\t * all zones.\n+\t */\n+\tzlo-\u003erestore_in_progress = true;\n+\tfor (i = 0, se = zlo-\u003eelements; i \u003c zlo-\u003enr_elements; i++, se++) {\n+\t\tif (READ_ONCE(se-\u003estatus) == BLK_SE_STS_REMOVED)\n+\t\t\tWRITE_ONCE(se-\u003estatus, BLK_SE_STS_RESTORE_IN_PROGRESS);\n+\t}\n+\n+\tschedule_delayed_work(\u0026zlo-\u003erestore_elements_work,\n+\t\t\t msecs_to_jiffies(5000));\n+\n+unlock:\n+\tmutex_unlock(\u0026zlo-\u003estor_elements_lock);\n+\n+\treturn ret;\n+}\n+\n static void zloop_free_disk(struct gendisk *disk)\n {\n \tstruct zloop_device *zlo = disk-\u003eprivate_data;\n \tunsigned int i;\n \n+\tcancel_delayed_work_sync(\u0026zlo-\u003eremove_element_work);\n+\tcancel_delayed_work_sync(\u0026zlo-\u003erestore_elements_work);\n+\n \tblk_mq_free_tag_set(\u0026zlo-\u003etag_set);\n \n \tfor (i = 0; i \u003c zlo-\u003enr_zones; i++) {\n@@ -1019,15 +1407,24 @@ static void zloop_free_disk(struct gendisk *disk)\n \n \tfput(zlo-\u003edata_dir);\n \tdestroy_workqueue(zlo-\u003eworkqueue);\n+\tkfree(zlo-\u003eelements);\n \tkfree(zlo-\u003ebase_dir);\n \tkvfree(zlo);\n }\n \n+\n+static const struct blk_storage_elements_ops zloop_se_ops = {\n+\t.report_elements\t= zloop_report_elements,\n+\t.remove_element\t\t= zloop_remove_element,\n+\t.restore_elements\t= zloop_restore_elements,\n+};\n+\n static const struct block_device_operations zloop_fops = {\n \t.owner\t\t\t= THIS_MODULE,\n \t.open\t\t\t= zloop_open,\n \t.report_zones\t\t= zloop_report_zones,\n \t.free_disk\t\t= zloop_free_disk,\n+\t.se_ops\t\t\t= \u0026zloop_se_ops,\n };\n \n __printf(3, 4)\n@@ -1112,6 +1509,24 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,\n \tif (!opts-\u003ebuffered_io)\n \t\toflags |= O_DIRECT;\n \n+\tif (zlo-\u003estor_elements != ZLOOP_STOR_ELEMENTS_NONE) {\n+\t\tunsigned int nr_elems = zlo-\u003enr_elements;\n+\t\tunsigned int se_idx;\n+\n+\t\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)\n+\t\t\tnr_elems /= 2;\n+\t\tse_idx = zone_no % nr_elems;\n+\t\tzlo-\u003eelements[se_idx].nr_zones++;\n+\n+\t\tzone-\u003ewr_se_id = se_idx + 1;\n+\t\tzone-\u003erd_se_id = zone-\u003ewr_se_id;\n+\t\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {\n+\t\t\tzone-\u003erd_se_id += nr_elems;\n+\t\t\tzlo-\u003eelements[zone-\u003erd_se_id - 1].nr_zones =\n+\t\t\t\tzlo-\u003eelements[se_idx].nr_zones;\n+\t\t}\n+\t}\n+\n \tif (zone_no \u003c zlo-\u003enr_conv_zones) {\n \t\t/* Conventional zone file. */\n \t\tset_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags);\n@@ -1182,6 +1597,79 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,\n \treturn ret;\n }\n \n+#define ZLOOP_MIN_STOR_ELEMENTS\t\t\t2\n+#define ZLOOP_MAX_STOR_ELEMENTS\t\t\t32\n+#define ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS\t32\n+\n+static int zloop_create_storage_elements(struct zloop_device *zlo)\n+{\n+\tstruct blk_storage_element *se, *paired_se;\n+\tunsigned int i, nr_elems, nr_elements;\n+\n+\t/*\n+\t * Calculate the number of storage elements we are going to emulate.\n+\t * To achieve a somewhat realistic emulation, we want at least 2 storage\n+\t * elements, and no more than 32, targeting at least 32 zones per\n+\t * element.\n+\t */\n+\tif (zlo-\u003enr_zones \u003c= 64)\n+\t\tnr_elements = 2;\n+\telse\n+\t\tnr_elements =\n+\t\t\tmin(ZLOOP_MAX_STOR_ELEMENTS,\n+\t\t\t zlo-\u003enr_zones / ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS);\n+\n+\t/*\n+\t * If we are emulating pairs of read and write storage elements, we need\n+\t * double the number of storage element descriptors.\n+\t */\n+\tnr_elems = nr_elements;\n+\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)\n+\t\tnr_elems *= 2;\n+\tzlo-\u003eelements = kzalloc_objs(struct blk_storage_element, nr_elems);\n+\tif (!zlo-\u003eelements)\n+\t\treturn -ENOMEM;\n+\n+\tfor (i = 0, se = zlo-\u003eelements; i \u003c nr_elements; i++, se++) {\n+\t\tse-\u003eid = i + 1;\n+\t\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)\n+\t\t\tse-\u003etype = BLK_SE_TYPE_WRITE;\n+\t\telse\n+\t\t\tse-\u003etype = BLK_SE_TYPE_RDWR;\n+\t\tse-\u003estatus = BLK_SE_STS_OK;\n+\t}\n+\n+\t/*\n+\t * If we are emulating pairs of read and write storage elements,\n+\t * initialize the read elements paired with the write elements we just\n+\t * initialized.\n+\t */\n+\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {\n+\t\tfor (i = 0; i \u003c nr_elements; i++, se++) {\n+\t\t\tse-\u003eid = nr_elements + i + 1;\n+\t\t\tpaired_se = \u0026zlo-\u003eelements[i];\n+\t\t\tse-\u003epaired_id = paired_se-\u003eid;\n+\t\t\tpaired_se-\u003epaired_id = se-\u003eid;\n+\t\t\tse-\u003etype = BLK_SE_TYPE_READ;\n+\t\t\tse-\u003estatus = BLK_SE_STS_OK;\n+\t\t}\n+\t}\n+\n+\t/*\n+\t * Make sure we do not allow removing all storage elements as that does\n+\t * not make any sense. This is consistent with the device advertized\n+\t * limit of SCSI and ATA devices supporting the storage element\n+\t * depopulation feature.\n+\t */\n+\tzlo-\u003enr_elements = nr_elems;\n+\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)\n+\t\tzlo-\u003emax_nr_removed_elements = zlo-\u003enr_elements - 2;\n+\telse\n+\t\tzlo-\u003emax_nr_removed_elements = zlo-\u003enr_elements - 1;\n+\n+\treturn 0;\n+}\n+\n static bool zloop_dev_exists(struct zloop_device *zlo)\n {\n \tstruct file *cnv, *seq;\n@@ -1237,6 +1725,10 @@ static int zloop_ctl_add(struct zloop_options *opts)\n \tWRITE_ONCE(zlo-\u003estate, Zlo_creating);\n \tspin_lock_init(\u0026zlo-\u003eopen_zones_lock);\n \tINIT_LIST_HEAD(\u0026zlo-\u003eopen_zones_lru_list);\n+\tmutex_init(\u0026zlo-\u003estor_elements_lock);\n+\tINIT_DELAYED_WORK(\u0026zlo-\u003eremove_element_work, zloop_remove_element_work);\n+\tINIT_DELAYED_WORK(\u0026zlo-\u003erestore_elements_work,\n+\t\t\t zloop_restore_elements_work);\n \n \tret = mutex_lock_killable(\u0026zloop_ctl_mutex);\n \tif (ret)\n@@ -1270,12 +1762,19 @@ static int zloop_ctl_add(struct zloop_options *opts)\n \tif (zlo-\u003ezone_append)\n \t\tzlo-\u003eordered_zone_append = opts-\u003eordered_zone_append;\n \tzlo-\u003ediscard_write_cache = opts-\u003ediscard_write_cache;\n+\tzlo-\u003estor_elements = opts-\u003estor_elements;\n+\n+\tif (zlo-\u003estor_elements != ZLOOP_STOR_ELEMENTS_NONE) {\n+\t\tret = zloop_create_storage_elements(zlo);\n+\t\tif (ret)\n+\t\t\tgoto out_free_idr;\n+\t}\n \n \tzlo-\u003eworkqueue = alloc_workqueue(\"zloop%d\", WQ_UNBOUND | WQ_FREEZABLE,\n \t\t\t\topts-\u003enr_queues * opts-\u003equeue_depth, zlo-\u003eid);\n \tif (!zlo-\u003eworkqueue) {\n \t\tret = -ENOMEM;\n-\t\tgoto out_free_idr;\n+\t\tgoto out_destroy_storage_elements;\n \t}\n \n \tif (opts-\u003ebase_dir)\n@@ -1361,6 +1860,10 @@ static int zloop_ctl_add(struct zloop_options *opts)\n \t\tzlo-\u003eid, zlo-\u003enr_zones,\n \t\t((sector_t)zlo-\u003ezone_size \u003c\u003c SECTOR_SHIFT) \u003e\u003e 20,\n \t\tzlo-\u003eblock_size);\n+\tif (zlo-\u003enr_elements)\n+\t\tpr_info(\"zloop%d: %d storage elements\\n\",\n+\t\t\tzlo-\u003eid, zlo-\u003enr_elements);\n+\n \tpr_info(\"zloop%d: using %s%s zone append\\n\",\n \t\tzlo-\u003eid,\n \t\tzlo-\u003eordered_zone_append ? \"ordered \" : \"\",\n@@ -1384,6 +1887,8 @@ static int zloop_ctl_add(struct zloop_options *opts)\n \tkfree(zlo-\u003ebase_dir);\n out_destroy_workqueue:\n \tdestroy_workqueue(zlo-\u003eworkqueue);\n+out_destroy_storage_elements:\n+\tkfree(zlo-\u003eelements);\n out_free_idr:\n \tmutex_lock(\u0026zloop_ctl_mutex);\n \tidr_remove(\u0026zloop_index_idr, zlo-\u003eid);\n@@ -1483,6 +1988,42 @@ static int zloop_ctl_remove(struct zloop_options *opts)\n \treturn 0;\n }\n \n+static int zloop_ctl_degrade_element(struct zloop_options *opts)\n+{\n+\tstruct zloop_device *zlo;\n+\tint ret = 0;\n+\n+\tif (!(opts-\u003emask \u0026 ZLOOP_OPT_ID)) {\n+\t\tpr_err(\"No ID specified for degrade_element\\n\");\n+\t\treturn -EINVAL;\n+\t}\n+\n+\tif (opts-\u003emask \u0026 ~(ZLOOP_OPT_ID | ZLOOP_OPT_ELEMENT_ID)) {\n+\t\tpr_err(\"Invalid option specified for degrade_element\\n\");\n+\t\treturn -EINVAL;\n+\t}\n+\n+\tmutex_lock(\u0026zloop_ctl_mutex);\n+\n+\tzlo = idr_find(\u0026zloop_index_idr, opts-\u003eid);\n+\tif (!zlo || zlo-\u003estate == Zlo_creating)\n+\t\tret = -ENODEV;\n+\telse if (zlo-\u003estate == Zlo_deleting)\n+\t\tret = -EINVAL;\n+\tif (ret)\n+\t\tgoto unlock;\n+\n+\tret = zloop_degrade_element(zlo, opts-\u003eelement_id);\n+\tif (!ret)\n+\t\tpr_info(\"Degraded element %u of device %u\\n\",\n+\t\t\topts-\u003eid, opts-\u003eelement_id);\n+\n+unlock:\n+\tmutex_unlock(\u0026zloop_ctl_mutex);\n+\n+\treturn ret;\n+}\n+\n static int zloop_parse_options(struct zloop_options *opts, const char *buf)\n {\n \tsubstring_t args[MAX_OPT_ARGS];\n@@ -1502,6 +2043,8 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)\n \topts-\u003ebuffered_io = ZLOOP_DEF_BUFFERED_IO;\n \topts-\u003ezone_append = ZLOOP_DEF_ZONE_APPEND;\n \topts-\u003eordered_zone_append = ZLOOP_DEF_ORDERED_ZONE_APPEND;\n+\topts-\u003estor_elements = ZLOOP_DEF_STOR_ELEMENTS;\n+\topts-\u003eelement_id = ZLOOP_DEF_ELEMENT_ID;\n \n \tif (!buf)\n \t\treturn 0;\n@@ -1636,6 +2179,30 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)\n \t\tcase ZLOOP_OPT_DISCARD_WRITE_CACHE:\n \t\t\topts-\u003ediscard_write_cache = true;\n \t\t\tbreak;\n+\t\tcase ZLOOP_OPT_STOR_ELEMENTS:\n+\t\t\tif (match_uint(args, \u0026token)) {\n+\t\t\t\tret = -EINVAL;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\t\t\tswitch (token) {\n+\t\t\tcase ZLOOP_STOR_ELEMENTS_NONE:\n+\t\t\tcase ZLOOP_STOR_ELEMENTS_RDWR:\n+\t\t\tcase ZLOOP_STOR_ELEMENTS_PAIRS:\n+\t\t\t\tbreak;\n+\t\t\tdefault:\n+\t\t\t\tpr_err(\"Invalid stor_elements value\\n\");\n+\t\t\t\tret = -EINVAL;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\t\t\topts-\u003estor_elements = token;\n+\t\t\tbreak;\n+\t\tcase ZLOOP_OPT_ELEMENT_ID:\n+\t\t\tif (match_uint(args, \u0026token)) {\n+\t\t\t\tret = -EINVAL;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\t\t\topts-\u003eelement_id = token;\n+\t\t\tbreak;\n \t\tcase ZLOOP_OPT_ERR:\n \t\tdefault:\n \t\t\tpr_warn(\"unknown parameter or missing value '%s'\\n\", p);\n@@ -1664,14 +2231,16 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)\n enum {\n \tZLOOP_CTL_ADD,\n \tZLOOP_CTL_REMOVE,\n+\tZLOOP_CTL_DEGRADE_ELEMENT,\n };\n \n static struct zloop_ctl_op {\n \tint\t\tcode;\n \tconst char\t*name;\n } zloop_ctl_ops[] = {\n-\t{ ZLOOP_CTL_ADD,\t\"add\" },\n-\t{ ZLOOP_CTL_REMOVE,\t\"remove\" },\n+\t{ ZLOOP_CTL_ADD,\t\t\"add\" },\n+\t{ ZLOOP_CTL_REMOVE,\t\t\"remove\" },\n+\t{ ZLOOP_CTL_DEGRADE_ELEMENT,\t\"degrade_element\" },\n \t{ -1,\tNULL },\n };\n \n@@ -1719,6 +2288,9 @@ static ssize_t zloop_ctl_write(struct file *file, const char __user *ubuf,\n \tcase ZLOOP_CTL_REMOVE:\n \t\tret = zloop_ctl_remove(\u0026opts);\n \t\tbreak;\n+\tcase ZLOOP_CTL_DEGRADE_ELEMENT:\n+\t\tret = zloop_ctl_degrade_element(\u0026opts);\n+\t\tbreak;\n \tdefault:\n \t\tpr_err(\"Invalid operation\\n\");\n \t\tret = -EINVAL;\n@@ -1742,6 +2314,8 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)\n \t\ttok = \u0026zloop_opt_tokens[i];\n \t\tif (!tok-\u003epattern)\n \t\t\tbreak;\n+\t\tif (tok-\u003etoken == ZLOOP_OPT_ELEMENT_ID)\n+\t\t\tcontinue;\n \t\tif (i)\n \t\t\tseq_putc(seq_file, ',');\n \t\tseq_puts(seq_file, tok-\u003epattern);\n@@ -1752,6 +2326,10 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)\n \tseq_puts(seq_file, zloop_ctl_ops[1].name);\n \tseq_puts(seq_file, \" id=%d\\n\");\n \n+\t/* Degrade element operation */\n+\tseq_puts(seq_file, zloop_ctl_ops[2].name);\n+\tseq_puts(seq_file, \" id=%d,element_id=%d\\n\");\n+\n \treturn 0;\n }\n \ndiff --git a/drivers/scsi/sd.c b/drivers/scsi/sd.c\nindex a1b21ea14e549..4372fa80f792d 100644\n--- a/drivers/scsi/sd.c\n+++ b/drivers/scsi/sd.c\n@@ -3937,6 +3937,9 @@ static const struct block_device_operations sd_fops = {\n \t.get_unique_id\t\t= sd_get_unique_id,\n \t.free_disk\t\t= scsi_disk_free_disk,\n \t.pr_ops\t\t\t= \u0026sd_pr_ops,\n+#ifdef CONFIG_BLK_DEV_ZONED\n+\t.se_ops\t\t\t= \u0026sd_zbc_se_ops,\n+#endif\n };\n \n /**\ndiff --git a/drivers/scsi/sd.h b/drivers/scsi/sd.h\nindex 574af82430169..b68b7ee5a5fa5 100644\n--- a/drivers/scsi/sd.h\n+++ b/drivers/scsi/sd.h\n@@ -156,6 +156,7 @@ struct scsi_disk {\n \tunsigned\tignore_medium_access_errors : 1;\n \tunsigned\trscs : 1; /* reduced stream control support */\n \tunsigned\tuse_atomic_write_boundary : 1;\n+\tunsigned\tmodify_zones_supported : 1;\n };\n #define to_scsi_disk(obj) container_of(obj, struct scsi_disk, disk_dev)\n \n@@ -242,6 +243,8 @@ unsigned int sd_zbc_complete(struct scsi_cmnd *cmd, unsigned int good_bytes,\n int sd_zbc_report_zones(struct gendisk *disk, sector_t sector,\n \t\tunsigned int nr_zones, struct blk_report_zones_args *args);\n \n+extern const struct blk_storage_elements_ops sd_zbc_se_ops;\n+\n #else /* CONFIG_BLK_DEV_ZONED */\n \n static inline int sd_zbc_read_zones(struct scsi_disk *sdkp,\ndiff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c\nindex 456beaf2e7690..6698d156a0c34 100644\n--- a/drivers/scsi/sd_zbc.c\n+++ b/drivers/scsi/sd_zbc.c\n@@ -516,6 +516,21 @@ static int sd_zbc_check_capacity(struct scsi_disk *sdkp, unsigned char *buf,\n \treturn 0;\n }\n \n+/*\n+ * sd_zbc_check_modify_zones - Check if the device supports depopulation\n+ * @sdkp: Target disk\n+ * @buf: command buffer\n+ *\n+ * Check if the device supports the REMOVE ELEMENT AND MODIFY ZONES command.\n+ */\n+static inline bool sd_zbc_check_modify_zones(struct scsi_disk *sdkp,\n+\t\t\t\t\t unsigned char *buf)\n+{\n+\treturn scsi_report_opcode(sdkp-\u003edevice, buf, SD_BUF_SIZE,\n+\t\t\t\t SERVICE_ACTION_IN_16,\n+\t\t\t\t SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES) == 1;\n+}\n+\n static void sd_zbc_print_zones(struct scsi_disk *sdkp)\n {\n \tif (sdkp-\u003edevice-\u003etype != TYPE_ZBC || !sdkp-\u003ecapacity)\n@@ -533,6 +548,239 @@ static void sd_zbc_print_zones(struct scsi_disk *sdkp)\n \t\t\t sdkp-\u003ezone_info.zone_blocks);\n }\n \n+static void sd_zbc_parse_storage_element(struct scsi_disk *sdkp, u8 *desc,\n+\t\t\t\t\t struct blk_storage_element *element)\n+{\n+\tstruct scsi_device *sdp = sdkp-\u003edevice;\n+\tsector_t zone_sectors = sd_zbc_zone_sectors(sdkp);\n+\tu64 capacity;\n+\n+\tmemset(element, 0, sizeof(*element));\n+\n+\telement-\u003eid = get_unaligned_be32(\u0026desc[4]);\n+\n+\tswitch (desc[14]) {\n+\tcase SCSI_PHYS_ELEM_TYPE_ALL_ACCESS_STORAGE:\n+\t\telement-\u003etype = BLK_SE_TYPE_RDWR;\n+\t\tcapacity = get_unaligned_be64(\u0026desc[16]);\n+\t\tif (!zone_sectors || capacity == ULLONG_MAX)\n+\t\t\telement-\u003enr_zones = 0;\n+\t\telse\n+\t\t\telement-\u003enr_zones =\n+\t\t\t\tlogical_to_sectors(sdp, capacity) \u003e\u003e\n+\t\t\t\t\tilog2(zone_sectors);\n+\t\tbreak;\n+\tcase SCSI_PHYS_ELEM_TYPE_FRAC_ACCESS_STORAGE:\n+\t\telement-\u003epaired_id = get_unaligned_be32(\u0026desc[16]);\n+\t\tif (desc[20] \u0026 0x01)\n+\t\t\telement-\u003etype = BLK_SE_TYPE_READ;\n+\t\telse\n+\t\t\telement-\u003etype = BLK_SE_TYPE_WRITE;\n+\t\telement-\u003enr_zones = get_unaligned_be64(\u0026desc[24]);\n+\t\tbreak;\n+\tdefault:\n+\t\telement-\u003etype = BLK_SE_TYPE_UNKNOWN;\n+\t}\n+\n+\tswitch (desc[15]) {\n+\tcase SCSI_PHYS_ELEM_HEALTH_WITHIN_SPEC_LIMITS:\n+\tcase SCSI_PHYS_ELEM_HEALTH_AT_SPEC_LIMITS:\n+\t\telement-\u003estatus = BLK_SE_STS_OK;\n+\t\tbreak;\n+\tcase SCSI_PHYS_ELEM_HEALTH_OUTSIDE_SPEC_LIMITS:\n+\t\telement-\u003estatus = BLK_SE_STS_DEGRADED;\n+\t\tbreak;\n+\tcase SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_ERR:\n+\t\telement-\u003estatus = BLK_SE_STS_RESTORE_ERROR;\n+\t\tbreak;\n+\tcase SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_IN_PROGRESS:\n+\t\telement-\u003estatus = BLK_SE_STS_RESTORE_IN_PROGRESS;\n+\t\tbreak;\n+\tcase SCSI_PHYS_ELEM_HEALTH_DEPOP_ERR:\n+\t\telement-\u003estatus = BLK_SE_STS_REMOVE_ERROR;\n+\t\tbreak;\n+\tcase SCSI_PHYS_ELEM_HEALTH_DEPOP_IN_PROGRESS:\n+\t\telement-\u003estatus = BLK_SE_STS_REMOVE_IN_PROGRESS;\n+\t\tbreak;\n+\tcase SCSI_PHYS_ELEM_HEALTH_DEPOP_OK:\n+\t\telement-\u003estatus = BLK_SE_STS_REMOVED;\n+\t\telement-\u003erestore_allowed = desc[13] \u0026 0x01;\n+\t\tbreak;\n+\tcase SCSI_PHYS_ELEM_HEALTH_NOT_REPORTED:\n+\tdefault:\n+\t\telement-\u003estatus = BLK_SE_STS_UNKNOWN;\n+\t\tbreak;\n+\t}\n+}\n+\n+/*\n+ * Large hard limit on the number of storage elements. This accommodates all\n+ * known devices today and likely forever :)\n+ */\n+#define SD_ZBC_MAX_STORAGE_ELEMENTS\t255\n+\n+static int sd_zbc_report_storage_elements(struct gendisk *disk,\n+\t\t\t\t\t struct blk_storage_element *elements,\n+\t\t\t\t\t unsigned int *nr_elements)\n+{\n+\tstruct scsi_disk *sdkp = scsi_disk(disk);\n+\tstruct scsi_device *sdp = sdkp-\u003edevice;\n+\tconst int timeout = sdp-\u003erequest_queue-\u003erq_timeout;\n+\tstruct scsi_sense_hdr sshdr;\n+\tconst struct scsi_exec_args exec_args = {\n+\t\t.sshdr = \u0026sshdr,\n+\t};\n+\tunsigned char cmd[16];\n+\tunsigned int nr_descs, nr_se;\n+\tunsigned int buf_size;\n+\tint i, ret = 0, result;\n+\tu8 *desc, *buf;\n+\n+\tif (!sdkp-\u003emodify_zones_supported)\n+\t\treturn -EOPNOTSUPP;\n+\n+\t/*\n+\t * We need at least 32B for the report header and 32B for each\n+\t * descriptor.\n+\t */\n+\tnr_se = min(SD_ZBC_MAX_STORAGE_ELEMENTS, *nr_elements);\n+\tbuf_size = ALIGN((nr_se + 1) * 32, SECTOR_SIZE);\n+\n+\tbuf = kzalloc(buf_size, GFP_KERNEL);\n+\tif (!buf)\n+\t\treturn -ENOMEM;\n+\n+\tmemset(cmd, 0, 16);\n+\tcmd[0] = SERVICE_ACTION_IN_16;\n+\tcmd[1] = SAI_GET_PHYSICAL_ELEMENT_STATUS;\n+\tput_unaligned_be32(buf_size, \u0026cmd[10]);\n+\n+\tresult = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, buf, buf_size,\n+\t\t\t\t timeout, 1, \u0026exec_args);\n+\tif (result) {\n+\t\tsd_printk(KERN_ERR, sdkp,\n+\t\t\t \"GET PHYSICAL ELEMENT STATUS failed\\n\");\n+\t\tsd_print_result(sdkp, \"GET PHYSICAL ELEMENT STATUS\", result);\n+\t\tif (result \u003e 0 \u0026\u0026 scsi_sense_valid(\u0026sshdr))\n+\t\t\tsd_print_sense_hdr(sdkp, \u0026sshdr);\n+\t\tret = -EIO;\n+\t\tgoto free_buf;\n+\t}\n+\n+\tnr_descs = get_unaligned_be32(\u0026buf[0]);\n+\tif (!nr_descs) {\n+\t\tsd_printk(KERN_ERR, sdkp,\n+\t\t\t \"Invalid number of phys element descriptors\\n\");\n+\t\tret = -EIO;\n+\t\tgoto free_buf;\n+\t}\n+\tif (nr_descs \u003e SD_ZBC_MAX_STORAGE_ELEMENTS) {\n+\t\tsd_printk(KERN_ERR, sdkp,\n+\t\t\t \"Unsupported number of phys element descriptors\\n\");\n+\t\tret = -EIO;\n+\t\tgoto free_buf;\n+\t}\n+\n+\tif (!elements) {\n+\t\t*nr_elements = nr_descs;\n+\t\tgoto free_buf;\n+\t}\n+\n+\tnr_descs = get_unaligned_be32(\u0026buf[4]);\n+\tif (!nr_descs) {\n+\t\tsd_printk(KERN_ERR, sdkp,\n+\t\t\t\"Invalid number of reported phys element descriptors\\n\");\n+\t\tret = -EIO;\n+\t\tgoto free_buf;\n+\t}\n+\n+\tdesc = \u0026buf[32];\n+\tfor (i = 0; i \u003c min(nr_se, nr_descs); i++, desc += 32)\n+\t\tsd_zbc_parse_storage_element(sdkp, desc, \u0026elements[i]);\n+\t*nr_elements = i;\n+\n+free_buf:\n+\tkfree(buf);\n+\n+\treturn ret;\n+}\n+\n+static int sd_zbc_remove_storage_element(struct gendisk *disk,\n+\t\t\t\t\t unsigned int element_id)\n+{\n+\tstruct scsi_disk *sdkp = scsi_disk(disk);\n+\tstruct scsi_device *sdp = sdkp-\u003edevice;\n+\tconst int timeout = sdp-\u003erequest_queue-\u003erq_timeout;\n+\tstruct scsi_sense_hdr sshdr;\n+\tconst struct scsi_exec_args exec_args = {\n+\t\t.sshdr = \u0026sshdr,\n+\t};\n+\tunsigned char cmd[16];\n+\tint result;\n+\n+\tif (!sdkp-\u003emodify_zones_supported)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tmemset(cmd, 0, 16);\n+\tcmd[0] = SERVICE_ACTION_IN_16;\n+\tcmd[1] = SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES;\n+\tput_unaligned_be32(element_id, \u0026cmd[10]);\n+\n+\tresult = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,\n+\t\t\t\t timeout, 1, \u0026exec_args);\n+\tif (result) {\n+\t\tsd_printk(KERN_ERR, sdkp,\n+\t\t\t \"REMOVE ELEMENT AND MODIFY ZONES failed\\n\");\n+\t\tsd_print_result(sdkp,\n+\t\t\t\t\"REMOVE ELEMENT AND MODIFY ZONES\", result);\n+\t\tif (result \u003e 0 \u0026\u0026 scsi_sense_valid(\u0026sshdr))\n+\t\t\tsd_print_sense_hdr(sdkp, \u0026sshdr);\n+\t\treturn -EIO;\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int sd_zbc_restore_storage_elements(struct gendisk *disk)\n+{\n+\tstruct scsi_disk *sdkp = scsi_disk(disk);\n+\tstruct scsi_device *sdp = sdkp-\u003edevice;\n+\tconst int timeout = sdp-\u003erequest_queue-\u003erq_timeout;\n+\tstruct scsi_sense_hdr sshdr;\n+\tconst struct scsi_exec_args exec_args = {\n+\t\t.sshdr = \u0026sshdr,\n+\t};\n+\tunsigned char cmd[16];\n+\tint result;\n+\n+\tif (!sdkp-\u003emodify_zones_supported)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tmemset(cmd, 0, 16);\n+\tcmd[0] = SERVICE_ACTION_IN_16;\n+\tcmd[1] = SAI_RESTORE_ELEMENTS_AND_REBUILD;\n+\n+\tresult = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,\n+\t\t\t\t timeout, 1, \u0026exec_args);\n+\tif (result) {\n+\t\tsd_printk(KERN_ERR, sdkp,\n+\t\t\t \"RESTORE ELEMENTS AND REBUILD failed\\n\");\n+\t\tsd_print_result(sdkp,\n+\t\t\t\t\"RESTORE ELEMENTS AND REBUILD\", result);\n+\t\tif (result \u003e 0 \u0026\u0026 scsi_sense_valid(\u0026sshdr))\n+\t\t\tsd_print_sense_hdr(sdkp, \u0026sshdr);\n+\t\treturn -EIO;\n+\t}\n+\n+\treturn 0;\n+}\n+\n+const struct blk_storage_elements_ops sd_zbc_se_ops = {\n+\t.report_elements\t= sd_zbc_report_storage_elements,\n+\t.remove_element\t\t= sd_zbc_remove_storage_element,\n+\t.restore_elements\t= sd_zbc_restore_storage_elements,\n+};\n+\n /*\n * Call blk_revalidate_disk_zones() if any of the zoned disk properties have\n * changed that make it necessary to call that function. Called by\n@@ -554,9 +802,15 @@ int sd_zbc_revalidate_zones(struct scsi_disk *sdkp)\n \tif (!blk_queue_is_zoned(q))\n \t\treturn 0;\n \n+\t/*\n+\t * If the zone size and number of zones has not changed, and the disk\n+\t * does not support depopulating heads, skip the rather slow call to\n+\t * blk_revalidate_disk_zones().\n+\t */\n \tif (sdkp-\u003ezone_info.zone_blocks == zone_blocks \u0026\u0026\n \t sdkp-\u003ezone_info.nr_zones == nr_zones \u0026\u0026\n-\t disk-\u003enr_zones == nr_zones)\n+\t disk-\u003enr_zones == nr_zones \u0026\u0026\n+\t !sdkp-\u003emodify_zones_supported)\n \t\treturn 0;\n \n \tsdkp-\u003ezone_info.zone_blocks = zone_blocks;\n@@ -620,6 +874,9 @@ int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim,\n \tif (ret != 0)\n \t\tgoto err;\n \n+\t/* Check if REMOVE ELEMENT AND MODIFY ZONES is supported. */\n+\tsdkp-\u003emodify_zones_supported = sd_zbc_check_modify_zones(sdkp, buf);\n+\n \tnr_zones = round_up(sdkp-\u003ecapacity, zone_blocks) \u003e\u003e ilog2(zone_blocks);\n \tif (nr_zones \u003e INT_MAX) {\n \t\tsd_printk(KERN_ERR, sdkp, \"Too many zones (%llu)\\n\",\ndiff --git a/include/linux/blkdev.h b/include/linux/blkdev.h\nindex d003a9d2d1f6c..859917b3b2568 100644\n--- a/include/linux/blkdev.h\n+++ b/include/linux/blkdev.h\n@@ -1574,6 +1574,21 @@ enum blk_unique_id {\n \tBLK_UID_NAA\t= 3,\n };\n \n+struct blk_storage_elements_ops {\n+\tint (*report_elements)(struct gendisk *disk,\n+\t\t\t struct blk_storage_element *elements,\n+\t\t\t unsigned int *nr_elements);\n+\tint (*remove_element)(struct gendisk *disk, unsigned int element_id);\n+\tint (*restore_elements)(struct gendisk *disk);\n+};\n+\n+int bdev_report_storage_elements(struct block_device *bdev,\n+\t\t\t\t struct blk_storage_element *elements,\n+\t\t\t\t unsigned int *nr_elements);\n+int bdev_remove_storage_element(struct block_device *bdev,\n+\t\t\t\tunsigned int element_id);\n+int bdev_restore_storage_elements(struct block_device *bdev);\n+\n struct block_device_operations {\n \tvoid (*submit_bio)(struct bio *bio);\n \tint (*poll_bio)(struct bio *bio, struct io_comp_batch *iob,\n@@ -1601,6 +1616,7 @@ struct block_device_operations {\n \t\t\tenum blk_unique_id id_type);\n \tstruct module *owner;\n \tconst struct pr_ops *pr_ops;\n+\tconst struct blk_storage_elements_ops\t*se_ops;\n \n \t/*\n \t * Special callback for probing GPT entry at a given sector.\ndiff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h\nindex 6638361209667..532cf69b5b114 100644\n--- a/include/uapi/linux/blkzoned.h\n+++ b/include/uapi/linux/blkzoned.h\n@@ -208,4 +208,85 @@ struct blk_zone_range {\n #define BLKFINISHZONE\t_IOW(0x12, 136, struct blk_zone_range)\n #define BLKREPORTZONEV2\t_IOWR(0x12, 142, struct blk_zone_report)\n \n+/**\n+ * enum blk_storage_element_status - Status of a zoned device storage elements.\n+ *\n+ * @BLK_SE_TYPE_RDWR: The storage element handles both reads and writes.\n+ * @BLK_SE_TYPE_READ: The storage element handles reads only.\n+ * @BLK_SE_TYPE_WRITE: The storage element handles writes only.\n+ * @BLK_SE_TYPE_UNKNOWN: The storage element type is not known.\n+ */\n+enum blk_storage_element_type {\n+\tBLK_SE_TYPE_RDWR\t\t= 0x01,\n+\tBLK_SE_TYPE_READ\t\t= 0x02,\n+\tBLK_SE_TYPE_WRITE\t\t= 0x03,\n+\tBLK_SE_TYPE_UNKNOWN\t\t= 0xFF,\n+};\n+\n+/**\n+ * enum blk_storage_element_status - Status of a zoned device storage elements.\n+ *\n+ * @BLK_SE_STS_OK: The storage element is operating normally.\n+ * @BLK_SE_STS_DEGRADED: The storage element has degraded and is not operating\n+ *\t\t\t normally.\n+ * @BLK_SE_STS_REMOVE_IN_PROGRESS: The storage element is being removed.\n+ * @BLK_SE_STS_REMOVE_ERROR: The storage element removal failed.\n+ * @BLK_SE_STS_RESTORE_IN_PROGRESS: The storage element is being restored.\n+ * @BLK_SE_STS_RESTORE_ERROR: The storage element restoration failed.\n+ * @BLK_SE_STS_REMOVED: The storage element was removed.\n+ * @BLK_SE_STS_UNKNOWN: The storage element status is unknown.\n+ */\n+enum blk_storage_element_status {\n+\tBLK_SE_STS_OK\t\t\t= 0x01,\n+\tBLK_SE_STS_DEGRADED\t\t= 0x02,\n+\tBLK_SE_STS_REMOVE_IN_PROGRESS\t= 0x03,\n+\tBLK_SE_STS_REMOVE_ERROR\t\t= 0x04,\n+\tBLK_SE_STS_RESTORE_IN_PROGRESS\t= 0x05,\n+\tBLK_SE_STS_RESTORE_ERROR\t= 0x06,\n+\tBLK_SE_STS_REMOVED\t\t= 0x07,\n+\tBLK_SE_STS_UNKNOWN\t\t= 0xFF,\n+};\n+\n+/**\n+ * struct blk_storage_element - Zoned device storage element descriptor.\n+ *\n+ * @id: The ID of the element (cannot be 0).\n+ * @paired_id: The ID of the paired element for an element that is not\n+ *\t of type BLK_SE_TYPE_RDWR.\n+ * @type: The type of the storage element (enum blk_storage_element_type).\n+ * @status: The health status of the storage element\n+ *\t (enum blk_storage_element_status).\n+ * @restore_allowed: Indicate if the storage element can be restored.\n+ * @nr_zones: The number of zones that the storage element handles.\n+ */\n+struct blk_storage_element {\n+\t__u32\tid;\n+\t__u32\tpaired_id;\n+\t__u64\tnr_zones;\n+\t__u8\ttype;\n+\t__u8\tstatus;\n+\t__u8\trestore_allowed;\n+\t__u8\treserved[5];\n+};\n+\n+struct blk_storage_elements_report {\n+\t__u32\t\t\t\tnr_elements;\n+\t__u32\t\t\t\treserved;\n+\tstruct blk_storage_element\telements[];\n+};\n+\n+/**\n+ * Zoned block device storage element management ioctl's:\n+ *\n+ * @BLKGETNRSTORELEMS: Get the number of storage elements of the device.\n+ * @BLKREPORTSTORELEMS: Get the device storage elements.\n+ * @BLKREMOVESTORELEM: Remove (depopulate) one storage element of a device.\n+ * @BLKRESTORESTORELEM: Restore (repopulate if possible) all storage elements\n+ *\t\t\tthat have been removed.\n+ */\n+#define BLKGETNRSTORELEMS _IOR(0x12, 143, __u32)\n+#define BLKREPORTSTORELEMS _IOWR(0x12, 144, struct blk_storage_elements_report)\n+#define BLKREMOVESTORELEM _IOW(0x12, 145, __u32)\n+#define BLKRESTORESTORELEMS _IO(0x12, 146)\n+\n #endif /* _UAPI_BLKZONED_H */\ndiff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h\nindex 34c6f219462a5..8a979326aa7f0 100644\n--- a/include/uapi/linux/fs.h\n+++ b/include/uapi/linux/fs.h\n@@ -309,6 +309,7 @@ struct file_attr {\n /* 130-136 and 142 are used by zoned block device ioctls (uapi/linux/blkzoned.h) */\n /* 137-141 are used by blk-crypto ioctls (uapi/linux/blk-crypto.h) */\n #define BLKTRACESETUP2 _IOWR(0x12, 142, struct blk_user_trace_setup2)\n+/* 143-146 are used by storage element management for zoned block devices. */\n \n #define BMAP_IOCTL 1\t\t/* obsolete - kept for compatibility */\n #define FIBMAP\t _IO(0x00,1)\t/* bmap access */\n"
} |
| 2/1 | 2026/10/06 13:26 | agent | patch-evaluator |
2mModel:core Args: null Results: {
"EnableConfigs": [
"BLK_DEV_ZONED",
"BLK_DEV_ZONED_LOOP",
"BLK_DEV_SD"
],
"FocusSymbols": [
"blkdev_zone_storage_elements_ioctl",
"bdev_report_storage_elements",
"bdev_remove_storage_element",
"bdev_restore_storage_elements",
"bdev_zone_is_offline",
"zloop_remove_element",
"zloop_restore_elements",
"zloop_report_elements",
"zloop_ctl_degrade_element",
"sd_zbc_report_storage_elements",
"sd_zbc_remove_storage_element",
"sd_zbc_restore_storage_elements"
],
"Reasoning": "The patch introduces new storage element management functionality for zoned block devices. It adds four new user-accessible block device ioctls (BLKGETNRSTORELEMS, BLKREPORTSTORELEMS, BLKREMOVESTORELEM, BLKRESTORESTORELEMS) in the core block layer, adds support in the scsi disk driver (sd_zbc), implements storage element emulation and a new control operation in the zoned loop driver (zloop), and introduces offline zone checks during bio submission. This reachable code path involves complex state transitions, synchronization (queue freeze/unfreeze, truncate, cache flush, delayed work), and memory management, making it an excellent target for syzkaller fuzzing.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 1145222b274a2f5837f52a04a8d52b5f3f74f8a8
Author: syz-cluster <triage@syzkaller.com>
Date: Tue Oct 6 13:26:40 2026 +0000
syz-cluster: applied patch under review
diff --git a/Documentation/admin-guide/blockdev/zoned_loop.rst b/Documentation/admin-guide/blockdev/zoned_loop.rst
index 64277494fb36a..ec39f141af8d6 100644
--- a/Documentation/admin-guide/blockdev/zoned_loop.rst
+++ b/Documentation/admin-guide/blockdev/zoned_loop.rst
@@ -61,7 +61,7 @@ The options available for the add command can be listed by reading the
/dev/zloop-control device::
$ cat /dev/zloop-control
- add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache
+ add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache,stor_elements=%u
remove id=%d
In more details, the options that can be used with the "add" command are as
@@ -113,6 +113,12 @@ discard_write_cache Discard all data that was not explicitly persisted using a
each zone file to the size recorded during the last flush
operation. This simulates power fail events where
uncommitted data is lost.
+stor_elements Control storage element emulation. The default value is 0,
+ indicating no emulation. A value of 1 indicates that all
+ access storage elements (equivalent to read+write head of
+ a disk) are emulated. A value of 2 enables fractional
+ access storage element (equivalent to pairs of read and
+ write heads of a disk) emulation .
=================== =========================================================
3) Deleting a Zoned Device
diff --git a/block/blk-core.c b/block/blk-core.c
index 13dc70e8f55d9..d420c80d2d938 100644
--- a/block/blk-core.c
+++ b/block/blk-core.c
@@ -866,6 +866,9 @@ void submit_bio_noacct(struct bio *bio)
switch (bio_op(bio)) {
case REQ_OP_READ:
+ if (bdev_is_zoned(bdev) &&
+ bdev_zone_is_offline(bdev, bio->bi_iter.bi_sector))
+ goto end_io;
break;
case REQ_OP_WRITE:
if (bio->bi_opf & REQ_ATOMIC) {
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index 19268afb8752e..6ebb04e0a1593 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -18,6 +18,8 @@
#include <linux/mempool.h>
#include <linux/kthread.h>
#include <linux/freezer.h>
+#include <linux/delay.h>
+#include <linux/uaccess.h>
#include <trace/events/block.h>
@@ -308,6 +310,24 @@ bool bdev_zone_is_seq(struct block_device *bdev, sector_t sector)
}
EXPORT_SYMBOL_GPL(bdev_zone_is_seq);
+/**
+ * bdev_zone_is_offline - check if a sector belongs to an offline zone
+ * @bdev: block device to check
+ * @sector: sector number
+ *
+ * Check if @sector on @bdev is contained in an offline zone.
+ */
+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector)
+{
+ enum blk_zone_cond cond;
+
+ if (!bdev_is_zoned(bdev))
+ return false;
+
+ cond = disk_zone_get_cond(bdev->bd_disk, sector);
+ return cond == BLK_ZONE_COND_OFFLINE;
+}
+
/**
* bdev_zone_mgmt_allowed - check if management operations are allowed on a zone
* @bdev: block device to check
@@ -2700,5 +2720,326 @@ int queue_zone_wplugs_show(void *data, struct seq_file *m)
return 0;
}
-
#endif
+
+static int disk_wait_for_se_mgmt_completion(struct gendisk *disk)
+{
+ struct blk_storage_element *elements, *e;
+ unsigned int i, nr_se, nr_elements = 0;
+ int ret;
+
+ ret = disk->fops->se_ops->report_elements(disk, NULL, &nr_elements);
+ if (ret) {
+ pr_err("Failed to get number of storage elements\n");
+ return ret;
+ }
+
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return -ENOMEM;
+
+ while (1) {
+ /*
+ * Check if we have storage elements being removed or restored.
+ */
+ nr_se = nr_elements;
+ ret = disk->fops->se_ops->report_elements(disk, elements,
+ &nr_se);
+ if (ret) {
+ pr_err("Failed to get storage elements\n");
+ break;
+ }
+
+ e = elements;
+ for (i = 0; i < nr_se; i++, e++) {
+ if (e->status == BLK_SE_STS_REMOVE_IN_PROGRESS ||
+ e->status == BLK_SE_STS_RESTORE_IN_PROGRESS)
+ break;
+ }
+ if (i >= nr_se)
+ break;
+
+ /* Not done yet: wait and retry. */
+ msleep(500);
+ }
+
+ kfree(elements);
+
+ return ret;
+}
+
+/**
+ * bdev_report_storage_elements - report the storage elements of a block device
+ *
+ * Fill at most @nr_elements storage element descriptors in the array @elements.
+ * The number of storage elements filled in the array is returned using
+ * @nr_elements. If @elements is NULL, only @nr_elements is returned.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_report_storage_elements(struct block_device *bdev,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct gendisk *disk = bdev->bd_disk;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ if (!nr_elements)
+ return -EINVAL;
+
+ if (*nr_elements && !elements)
+ return -EINVAL;
+
+ return disk->fops->se_ops->report_elements(disk, elements, nr_elements);
+}
+EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
+
+static int blkdev_get_nr_storage_elements_ioctl(struct block_device *bdev,
+ void __user *argp)
+{
+ unsigned int nr_elements = 0;
+ int ret;
+
+ ret = bdev_report_storage_elements(bdev, NULL, &nr_elements);
+ if (ret)
+ return ret;
+
+ if (put_user(nr_elements, (unsigned int __user *)argp))
+ return -EFAULT;
+
+ return 0;
+}
+
+static int blkdev_report_storage_elements_ioctl(struct block_device *bdev,
+ void __user *argp)
+{
+ struct blk_storage_elements_report rep;
+ struct blk_storage_element *elements;
+ unsigned int nr_elements = 0;
+ unsigned long retc;
+ int ret;
+
+ if (!argp)
+ return -EINVAL;
+
+ if (copy_from_user(&rep, argp,
+ sizeof(struct blk_storage_elements_report)))
+ return -EFAULT;
+
+ ret = bdev_report_storage_elements(bdev, NULL, &nr_elements);
+ if (ret)
+ return ret;
+
+ nr_elements = min(rep.nr_elements, nr_elements);
+ if (!nr_elements)
+ return -EINVAL;
+
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return -ENOMEM;
+
+ ret = bdev_report_storage_elements(bdev, elements, &nr_elements);
+ if (ret)
+ goto free_elements;
+
+ retc = copy_to_user(argp + sizeof(struct blk_storage_elements_report),
+ elements,
+ sizeof(struct blk_storage_element) * nr_elements);
+ if (retc) {
+ ret = -EFAULT;
+ goto free_elements;
+ }
+
+ rep.nr_elements = nr_elements;
+ retc = copy_to_user(argp, &rep,
+ sizeof(struct blk_storage_elements_report));
+ if (retc)
+ ret = -EFAULT;
+
+free_elements:
+ kfree(elements);
+ return ret;
+}
+
+/**
+ * bdev_remove_storage_element - Remove (depopulate) a storage element of a
+ * block device
+ *
+ * Remove (depopulate) the storage element identified by @element_id from the
+ * block device @bdev. The caller is responsible for taking care of any
+ * necessary device write cache flush and invalidation of cached data for the
+ * zones that will be offlined.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /* Zero is not a valid storage element ID. */
+ if (!element_id)
+ return -EINVAL;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding command.
+ * The caller is responsible for not queuing up more I/Os by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ ret = disk->fops->se_ops->remove_element(disk, element_id);
+ if (ret)
+ return ret;
+
+ /* Revalidate the device zones once the opration completes. */
+ ret = disk_wait_for_se_mgmt_completion(disk);
+ if (ret)
+ return ret;
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
+
+static int blkdev_remove_storage_element_ioctl(struct block_device *bdev,
+ blk_mode_t mode, void __user *argp)
+{
+ unsigned int element_id;
+ int ret;
+
+ if (!(mode & BLK_OPEN_WRITE))
+ return -EBADF;
+ if (bdev_read_only(bdev))
+ return -EPERM;
+
+ if (get_user(element_id, (unsigned int __user *)argp))
+ return -EFAULT;
+
+ /*
+ * Flush the device volatile write cache and invalidate all cached data
+ * so that reads do not return old data for zones that went offline.
+ */
+ inode_lock(bdev->bd_mapping->host);
+ filemap_invalidate_lock(bdev->bd_mapping);
+
+ ret = blkdev_issue_flush(bdev->bd_disk->part0);
+ if (ret)
+ goto out_unlock;
+
+ ret = truncate_bdev_range(bdev, mode, 0,
+ (get_capacity(bdev->bd_disk) << SECTOR_SHIFT) - 1);
+ if (ret)
+ goto out_unlock;
+
+ ret = bdev_remove_storage_element(bdev, element_id);
+
+out_unlock:
+ filemap_invalidate_unlock(bdev->bd_mapping);
+ inode_unlock(bdev->bd_mapping->host);
+
+ return ret;
+}
+
+/**
+ * bdev_restore_storage_elements - Restore all depopulated storage elements of a
+ * block device
+ *
+ * Restore all storage elements of @bdev that have been depopulated. Not all
+ * elements may be restored by this operation.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_restore_storage_elements(struct block_device *bdev)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding commands.
+ * The caller is responsible for not queuing up more I/O by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ ret = disk->fops->se_ops->restore_elements(disk);
+ if (ret)
+ return ret;
+
+ ret = disk_wait_for_se_mgmt_completion(disk);
+ if (ret)
+ return ret;
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
+
+static int blkdev_restore_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode)
+{
+ int ret;
+
+ if (!(mode & BLK_OPEN_WRITE))
+ return -EBADF;
+ if (bdev_read_only(bdev))
+ return -EPERM;
+
+ /*
+ * Flush the device volatile write cache and invalidate all cached data
+ * so that reads do not return old data for zones that went offline.
+ */
+ inode_lock(bdev->bd_mapping->host);
+ filemap_invalidate_lock(bdev->bd_mapping);
+
+ ret = blkdev_issue_flush(bdev->bd_disk->part0);
+ if (ret)
+ goto out_unlock;
+
+ ret = truncate_bdev_range(bdev, mode, 0,
+ (get_capacity(bdev->bd_disk) << SECTOR_SHIFT) - 1);
+ if (ret)
+ goto out_unlock;
+
+ ret = bdev_restore_storage_elements(bdev);
+
+out_unlock:
+ filemap_invalidate_unlock(bdev->bd_mapping);
+ inode_unlock(bdev->bd_mapping->host);
+
+ return ret;
+}
+
+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp)
+{
+ if (!bdev_is_zoned(bdev) || !bdev->bd_disk->fops->se_ops)
+ return -ENOTTY;
+
+ switch (cmd) {
+ case BLKGETNRSTORELEMS:
+ return blkdev_get_nr_storage_elements_ioctl(bdev, argp);
+ case BLKREPORTSTORELEMS:
+ return blkdev_report_storage_elements_ioctl(bdev, argp);
+ case BLKREMOVESTORELEM:
+ return blkdev_remove_storage_element_ioctl(bdev, mode, argp);
+ case BLKRESTORESTORELEMS:
+ return blkdev_restore_storage_elements_ioctl(bdev, mode);
+ default:
+ break;
+ }
+
+ return -ENOTTY;
+}
diff --git a/block/blk.h b/block/blk.h
index 2cc03aa54c532..274afb46a809a 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -578,6 +578,10 @@ int blkdev_report_zones_ioctl(struct block_device *bdev, unsigned int cmd,
int blkdev_zone_mgmt_ioctl(struct block_device *bdev, blk_mode_t mode,
unsigned int cmd, unsigned long arg);
bool bdev_zone_mgmt_allowed(struct block_device *bdev, sector_t sector);
+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector);
+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp);
#else /* CONFIG_BLK_DEV_ZONED */
static inline void disk_init_zone_resources(struct gendisk *disk)
{
@@ -625,6 +629,17 @@ static inline bool bdev_zone_mgmt_allowed(struct block_device *bdev,
{
return false;
}
+static inline bool bdev_zone_is_offline(struct block_device *bdev,
+ sector_t sector)
+{
+ return false;
+}
+static inline int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp)
+{
+ return -ENOTTY;
+}
#endif /* CONFIG_BLK_DEV_ZONED */
struct block_device *bdev_alloc(struct gendisk *disk, u8 partno);
diff --git a/block/ioctl.c b/block/ioctl.c
index 64b4e6c0f6965..ccf806c2e37d1 100644
--- a/block/ioctl.c
+++ b/block/ioctl.c
@@ -679,6 +679,11 @@ static int blkdev_common_ioctl(struct block_device *bdev, blk_mode_t mode,
return put_uint(argp, bdev_zone_sectors(bdev));
case BLKGETNRZONES:
return put_uint(argp, bdev_nr_zones(bdev));
+ case BLKGETNRSTORELEMS:
+ case BLKREPORTSTORELEMS:
+ case BLKREMOVESTORELEM:
+ case BLKRESTORESTORELEMS:
+ return blkdev_zone_storage_elements_ioctl(bdev, mode, cmd, argp);
case BLKROGET:
return put_int(argp, bdev_read_only(bdev) != 0);
case BLKSSZGET: /* get block device logical block size */
diff --git a/drivers/block/zloop.c b/drivers/block/zloop.c
index 394dc2408ed73..3913bd2299108 100644
--- a/drivers/block/zloop.c
+++ b/drivers/block/zloop.c
@@ -37,6 +37,8 @@ enum {
ZLOOP_OPT_ORDERED_ZONE_APPEND = (1 << 10),
ZLOOP_OPT_DISCARD_WRITE_CACHE = (1 << 11),
ZLOOP_OPT_MAX_OPEN_ZONES = (1 << 12),
+ ZLOOP_OPT_STOR_ELEMENTS = (1 << 13),
+ ZLOOP_OPT_ELEMENT_ID = (1 << 14),
};
static const match_table_t zloop_opt_tokens = {
@@ -51,11 +53,23 @@ static const match_table_t zloop_opt_tokens = {
{ ZLOOP_OPT_BUFFERED_IO, "buffered_io" },
{ ZLOOP_OPT_ZONE_APPEND, "zone_append=%u" },
{ ZLOOP_OPT_ORDERED_ZONE_APPEND, "ordered_zone_append" },
- { ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
+ { ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
{ ZLOOP_OPT_MAX_OPEN_ZONES, "max_open_zones=%u" },
+ { ZLOOP_OPT_STOR_ELEMENTS, "stor_elements=%u" },
+ { ZLOOP_OPT_ELEMENT_ID, "element_id=%u" },
{ ZLOOP_OPT_ERR, NULL }
};
+/* Storage elements emulation types. */
+enum zloop_stor_elements {
+ /* No emulation. */
+ ZLOOP_STOR_ELEMENTS_NONE,
+ /* Emulate read+write storage elements. */
+ ZLOOP_STOR_ELEMENTS_RDWR,
+ /* Emulate pairs of associated read and write storage elements. */
+ ZLOOP_STOR_ELEMENTS_PAIRS,
+};
+
/* Default values for the "add" operation. */
#define ZLOOP_DEF_ID -1
#define ZLOOP_DEF_ZONE_SIZE ((256ULL * SZ_1M) >> SECTOR_SHIFT)
@@ -68,6 +82,8 @@ static const match_table_t zloop_opt_tokens = {
#define ZLOOP_DEF_BUFFERED_IO false
#define ZLOOP_DEF_ZONE_APPEND true
#define ZLOOP_DEF_ORDERED_ZONE_APPEND false
+#define ZLOOP_DEF_STOR_ELEMENTS ZLOOP_STOR_ELEMENTS_NONE
+#define ZLOOP_DEF_ELEMENT_ID 0
/* Arbitrary limit on the zone size (16GB). */
#define ZLOOP_MAX_ZONE_SIZE_MB 16384
@@ -87,6 +103,8 @@ struct zloop_options {
bool zone_append;
bool ordered_zone_append;
bool discard_write_cache;
+ enum zloop_stor_elements stor_elements;
+ unsigned int element_id;
};
/*
@@ -117,6 +135,8 @@ struct zloop_zone {
enum blk_zone_cond cond;
sector_t start;
sector_t wp;
+ unsigned int wr_se_id;
+ unsigned int rd_se_id;
gfp_t old_gfp_mask;
};
@@ -133,6 +153,7 @@ struct zloop_device {
bool zone_append;
bool ordered_zone_append;
bool discard_write_cache;
+ enum zloop_stor_elements stor_elements;
const char *base_dir;
struct file *data_dir;
@@ -150,6 +171,17 @@ struct zloop_device {
struct list_head open_zones_lru_list;
unsigned int nr_open_zones;
+ /* For storage elements emulation. */
+ struct mutex stor_elements_lock;
+ unsigned int nr_elements;
+ unsigned int max_nr_removed_elements;
+ unsigned int nr_removed_elements;
+ struct delayed_work remove_element_work;
+ unsigned int remove_element_id;
+ struct delayed_work restore_elements_work;
+ bool restore_in_progress;
+ struct blk_storage_element *elements;
+
struct zloop_zone zones[] __counted_by(nr_zones);
};
@@ -289,22 +321,30 @@ static bool zloop_do_open_zone(struct zloop_device *zlo,
}
}
-static void zloop_mark_full(struct zloop_device *zlo, struct zloop_zone *zone)
+static void zloop_set_zone_cond(struct zloop_device *zlo,
+ struct zloop_zone *zone,
+ enum blk_zone_cond cond)
{
lockdep_assert_held(&zone->wp_lock);
zloop_lru_remove_open_zone(zlo, zone);
- zone->cond = BLK_ZONE_COND_FULL;
- zone->wp = ULLONG_MAX;
+ zone->cond = cond;
+ if (cond == BLK_ZONE_COND_EMPTY)
+ zone->wp = zone->start;
+ else
+ zone->wp = ULLONG_MAX;
}
-static void zloop_mark_empty(struct zloop_device *zlo, struct zloop_zone *zone)
+static inline void zloop_set_zone_full(struct zloop_device *zlo,
+ struct zloop_zone *zone)
{
- lockdep_assert_held(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_FULL);
+}
- zloop_lru_remove_open_zone(zlo, zone);
- zone->cond = BLK_ZONE_COND_EMPTY;
- zone->wp = zone->start;
+static inline void zloop_set_zone_empty(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_EMPTY);
}
static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
@@ -339,9 +379,9 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
spin_lock(&zone->wp_lock);
if (!file_sectors) {
- zloop_mark_empty(zlo, zone);
+ zloop_set_zone_empty(zlo, zone);
} else if (file_sectors == zlo->zone_capacity) {
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
} else {
if (zone->cond != BLK_ZONE_COND_IMP_OPEN &&
zone->cond != BLK_ZONE_COND_EXP_OPEN)
@@ -353,6 +393,19 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
return 0;
}
+static bool zloop_zone_is_offline_or_readonly(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ bool ret;
+
+ spin_lock(&zone->wp_lock);
+ ret = zone->cond == BLK_ZONE_COND_OFFLINE ||
+ zone->cond == BLK_ZONE_COND_READONLY;
+ spin_unlock(&zone->wp_lock);
+
+ return ret;
+}
+
static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)
{
struct zloop_zone *zone = &zlo->zones[zone_no];
@@ -363,6 +416,11 @@ static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
ret = zloop_update_seq_zone(zlo, zone_no);
if (ret)
@@ -388,6 +446,11 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
ret = zloop_update_seq_zone(zlo, zone_no);
if (ret)
@@ -420,7 +483,22 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)
return ret;
}
-static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)
+static int zloop_do_reset_zone(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ if (vfs_truncate(&zone->file->f_path, 0))
+ return -EIO;
+
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_empty(zlo, zone);
+ clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
+ spin_unlock(&zone->wp_lock);
+
+ return 0;
+}
+
+static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no,
+ bool all_zones)
{
struct zloop_zone *zone = &zlo->zones[zone_no];
int ret = 0;
@@ -430,20 +508,19 @@ static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ if (!all_zones)
+ ret = -EIO;
+ goto unlock;
+ }
+
if (!test_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags) &&
zone->cond == BLK_ZONE_COND_EMPTY)
goto unlock;
- if (vfs_truncate(&zone->file->f_path, 0)) {
+ ret = zloop_do_reset_zone(zlo, zone);
+ if (ret)
set_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
- ret = -EIO;
- goto unlock;
- }
-
- spin_lock(&zone->wp_lock);
- zloop_mark_empty(zlo, zone);
- clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
- spin_unlock(&zone->wp_lock);
unlock:
mutex_unlock(&zone->lock);
@@ -457,7 +534,7 @@ static int zloop_reset_all_zones(struct zloop_device *zlo)
int ret;
for (i = zlo->nr_conv_zones; i < zlo->nr_zones; i++) {
- ret = zloop_reset_zone(zlo, i);
+ ret = zloop_reset_zone(zlo, i, true);
if (ret)
return ret;
}
@@ -475,6 +552,11 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (!test_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags) &&
zone->cond == BLK_ZONE_COND_FULL)
goto unlock;
@@ -487,7 +569,7 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)
}
spin_lock(&zone->wp_lock);
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
spin_unlock(&zone->wp_lock);
@@ -624,7 +706,7 @@ static int zloop_seq_write_prep(struct zloop_cmd *cmd)
if (!is_append || !zlo->ordered_zone_append) {
zone->wp += nr_sectors;
if (zone->wp == zone_end)
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
}
out_unlock:
spin_unlock(&zone->wp_lock);
@@ -766,7 +848,7 @@ static void zloop_handle_cmd(struct zloop_cmd *cmd)
cmd->ret = zloop_flush(zlo);
break;
case REQ_OP_ZONE_RESET:
- cmd->ret = zloop_reset_zone(zlo, rq_zone_no(rq));
+ cmd->ret = zloop_reset_zone(zlo, rq_zone_no(rq), false);
break;
case REQ_OP_ZONE_RESET_ALL:
cmd->ret = zloop_reset_all_zones(zlo);
@@ -875,30 +957,69 @@ static void zloop_complete_rq(struct request *rq)
blk_mq_end_request(rq, sts);
}
-static bool zloop_set_zone_append_sector(struct request *rq)
+static bool zloop_set_zone_append_sector(struct zloop_device *zlo,
+ struct zloop_zone *zone,
+ struct request *rq)
{
- struct zloop_device *zlo = rq->q->queuedata;
- unsigned int zone_no = rq_zone_no(rq);
- struct zloop_zone *zone = &zlo->zones[zone_no];
sector_t zone_end = zone->start + zlo->zone_capacity;
sector_t nr_sectors = blk_rq_sectors(rq);
- spin_lock(&zone->wp_lock);
-
if (zone->cond == BLK_ZONE_COND_FULL ||
- zone->wp + nr_sectors > zone_end) {
- spin_unlock(&zone->wp_lock);
+ zone->wp + nr_sectors > zone_end)
return false;
- }
rq->__sector = zone->wp;
zone->wp += blk_rq_sectors(rq);
if (zone->wp >= zone_end)
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
+
+ return true;
+}
+
+
+static bool zloop_prep_rq(struct zloop_device *zlo, struct request *rq)
+{
+ struct zloop_zone *zone = &zlo->zones[rq_zone_no(rq)];
+ bool is_write = op_is_write(req_op(rq));
+ bool ret = true;
+
+ spin_lock(&zone->wp_lock);
+
+ if (zlo->nr_elements) {
+ struct blk_storage_element *se;
+
+ if (zone->cond == BLK_ZONE_COND_OFFLINE ||
+ (zone->cond == BLK_ZONE_COND_READONLY && is_write)) {
+ ret = false;
+ goto unlock;
+ }
+
+ /*
+ * Check the health state of the storage element serving the
+ * zone.
+ */
+ if (is_write)
+ se = &zlo->elements[zone->wr_se_id - 1];
+ else
+ se = &zlo->elements[zone->rd_se_id - 1];
+ if (READ_ONCE(se->status) == BLK_SE_STS_DEGRADED) {
+ ret = false;
+ goto unlock;
+ }
+ }
+
+ /*
+ * If we need to strongly order zone append operations, set the request
+ * sector to the zone write pointer location now instead of when the
+ * command work runs.
+ */
+ if (zlo->ordered_zone_append && req_op(rq) == REQ_OP_ZONE_APPEND)
+ ret = zloop_set_zone_append_sector(zlo, zone, rq);
+unlock:
spin_unlock(&zone->wp_lock);
- return true;
+ return ret;
}
static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,
@@ -913,14 +1034,15 @@ static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,
return BLK_STS_IOERR;
}
- /*
- * If we need to strongly order zone append operations, set the request
- * sector to the zone write pointer location now instead of when the
- * command work runs.
- */
- if (zlo->ordered_zone_append && req_op(rq) == REQ_OP_ZONE_APPEND) {
- if (!zloop_set_zone_append_sector(rq))
+ switch (req_op(rq)) {
+ case REQ_OP_READ:
+ case REQ_OP_WRITE:
+ case REQ_OP_ZONE_APPEND:
+ if (!zloop_prep_rq(zlo, rq))
return BLK_STS_IOERR;
+ break;
+ default:
+ break;
}
blk_mq_start_request(rq);
@@ -1002,11 +1124,277 @@ static int zloop_report_zones(struct gendisk *disk, sector_t sector,
return nr_zones;
}
+static int zloop_report_elements(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct zloop_device *zlo = disk->private_data;
+ unsigned int nr_report = *nr_elements;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ *nr_elements = zlo->nr_elements;
+ if (elements) {
+ struct blk_storage_element *se = zlo->elements;
+ unsigned int i;
+
+ for (i = 0; i < min(nr_report, zlo->nr_elements); i++, se++) {
+ switch (READ_ONCE(se->status)) {
+ case BLK_SE_STS_REMOVED:
+ case BLK_SE_STS_RESTORE_ERROR:
+ se->restore_allowed = 1;
+ break;
+ default:
+ se->restore_allowed = 0;
+ }
+ memcpy(&elements[i], se,
+ sizeof(struct blk_storage_element));
+ }
+ }
+
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return 0;
+}
+
+static int zloop_degrade_element(struct zloop_device *zlo,
+ unsigned int element_id)
+{
+ struct blk_storage_element *se;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ if (!element_id || element_id > zlo->nr_elements)
+ return -EINVAL;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ se = &zlo->elements[element_id - 1];
+ if (READ_ONCE(se->status) == BLK_SE_STS_OK)
+ WRITE_ONCE(se->status, BLK_SE_STS_DEGRADED);
+ else
+ ret = -EINVAL;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
+static void zloop_remove_element_work(struct work_struct *work)
+{
+ struct zloop_device *zlo = container_of(work, struct zloop_device,
+ remove_element_work.work);
+ struct blk_storage_element *se, *paired_se = NULL;
+ enum blk_zone_cond cond;
+ struct zloop_zone *zone;
+ unsigned int i;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ /*
+ * If the element to remove is a read element, zones must go offline and
+ * the associated write element is also removed.
+ */
+ se = &zlo->elements[zlo->remove_element_id - 1];
+ switch (se->type) {
+ case BLK_SE_TYPE_RDWR:
+ cond = BLK_ZONE_COND_OFFLINE;
+ break;
+ case BLK_SE_TYPE_READ:
+ cond = BLK_ZONE_COND_OFFLINE;
+ paired_se = &zlo->elements[se->paired_id - 1];
+ break;
+ case BLK_SE_TYPE_WRITE:
+ cond = BLK_ZONE_COND_READONLY;
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ }
+
+ /*
+ * Change the condition of the zones owned by the (pair of) elements
+ * being removed and mark the elements removed.
+ */
+ for (i = 0, zone = zlo->zones; i < zlo->nr_zones; i++, zone++) {
+ if (zone->wr_se_id != se->id && zone->rd_se_id != se->id)
+ continue;
+ mutex_lock(&zone->lock);
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, cond);
+ spin_unlock(&zone->wp_lock);
+ mutex_unlock(&zone->lock);
+ }
+
+ WRITE_ONCE(se->status, BLK_SE_STS_REMOVED);
+
+ zlo->nr_removed_elements++;
+ if (paired_se && READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED) {
+ WRITE_ONCE(paired_se->status, BLK_SE_STS_REMOVED);
+ zlo->nr_removed_elements++;
+ }
+
+ zlo->remove_element_id = 0;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+}
+
+static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)
+{
+ struct zloop_device *zlo = disk->private_data;
+ struct blk_storage_element *se, *paired_se = NULL;
+ unsigned int nr_remove = 1;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ if (element_id > zlo->nr_elements)
+ return -EINVAL;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ if (zlo->remove_element_id || zlo->restore_in_progress) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /*
+ * Get the element to remove. If it is already removed, we have nothing
+ * to do.
+ */
+ se = &zlo->elements[element_id - 1];
+ if (READ_ONCE(se->status) == BLK_SE_STS_REMOVED)
+ goto unlock;
+
+ /*
+ * If the element to remove is a read element, the associated write
+ * element must also be removed.
+ */
+ if (se->type == BLK_SE_TYPE_READ) {
+ paired_se = &zlo->elements[se->paired_id - 1];
+ if (READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED)
+ nr_remove = 2;
+ }
+
+ if (zlo->nr_removed_elements + nr_remove >
+ zlo->max_nr_removed_elements) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /*
+ * Schedule the element removal with a delay, to emulate the (generally
+ * short) time it takes for a real device to depopulate a head and
+ * modify the zones.
+ */
+ zlo->remove_element_id = element_id;
+ WRITE_ONCE(se->status, BLK_SE_STS_REMOVE_IN_PROGRESS);
+ if (paired_se && READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED)
+ WRITE_ONCE(paired_se->status, BLK_SE_STS_REMOVE_IN_PROGRESS);
+
+ schedule_delayed_work(&zlo->remove_element_work,
+ msecs_to_jiffies(2000));
+
+unlock:
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
+static void zloop_restore_elements_work(struct work_struct *work)
+{
+ struct zloop_device *zlo = container_of(work, struct zloop_device,
+ restore_elements_work.work);
+ struct zloop_zone *zone = zlo->zones;
+ struct blk_storage_element *se;
+ unsigned int i;
+ int ret = 0;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ /* Reset all zones. */
+ for (i = 0; i < zlo->nr_zones && ret == 0; i++, zone++) {
+ mutex_lock(&zone->lock);
+ if (test_bit(ZLOOP_ZONE_CONV, &zone->flags)) {
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_NOT_WP);
+ spin_unlock(&zone->wp_lock);
+ } else {
+ ret = zloop_do_reset_zone(zlo, zone);
+ }
+ mutex_unlock(&zone->lock);
+ }
+
+ /* Restore all removed elements. */
+ for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
+ if (READ_ONCE(se->status) != BLK_SE_STS_RESTORE_IN_PROGRESS)
+ continue;
+ if (!ret)
+ WRITE_ONCE(se->status, BLK_SE_STS_OK);
+ else
+ WRITE_ONCE(se->status, BLK_SE_STS_RESTORE_ERROR);
+ }
+
+ zlo->nr_removed_elements = 0;
+ zlo->restore_in_progress = false;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+}
+
+static int zloop_restore_elements(struct gendisk *disk)
+{
+ struct zloop_device *zlo = disk->private_data;
+ struct blk_storage_element *se;
+ unsigned int i;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ if (zlo->remove_element_id || zlo->restore_in_progress) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /* If we have no removed elements, we have nothing to do. */
+ if (!zlo->nr_removed_elements)
+ goto unlock;
+
+ /*
+ * Schedule the elements restoration with a delay, to emulate the time
+ * it takes for a real device to restore all removed heads and reset
+ * all zones.
+ */
+ zlo->restore_in_progress = true;
+ for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
+ if (READ_ONCE(se->status) == BLK_SE_STS_REMOVED)
+ WRITE_ONCE(se->status, BLK_SE_STS_RESTORE_IN_PROGRESS);
+ }
+
+ schedule_delayed_work(&zlo->restore_elements_work,
+ msecs_to_jiffies(5000));
+
+unlock:
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
static void zloop_free_disk(struct gendisk *disk)
{
struct zloop_device *zlo = disk->private_data;
unsigned int i;
+ cancel_delayed_work_sync(&zlo->remove_element_work);
+ cancel_delayed_work_sync(&zlo->restore_elements_work);
+
blk_mq_free_tag_set(&zlo->tag_set);
for (i = 0; i < zlo->nr_zones; i++) {
@@ -1019,15 +1407,24 @@ static void zloop_free_disk(struct gendisk *disk)
fput(zlo->data_dir);
destroy_workqueue(zlo->workqueue);
+ kfree(zlo->elements);
kfree(zlo->base_dir);
kvfree(zlo);
}
+
+static const struct blk_storage_elements_ops zloop_se_ops = {
+ .report_elements = zloop_report_elements,
+ .remove_element = zloop_remove_element,
+ .restore_elements = zloop_restore_elements,
+};
+
static const struct block_device_operations zloop_fops = {
.owner = THIS_MODULE,
.open = zloop_open,
.report_zones = zloop_report_zones,
.free_disk = zloop_free_disk,
+ .se_ops = &zloop_se_ops,
};
__printf(3, 4)
@@ -1112,6 +1509,24 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,
if (!opts->buffered_io)
oflags |= O_DIRECT;
+ if (zlo->stor_elements != ZLOOP_STOR_ELEMENTS_NONE) {
+ unsigned int nr_elems = zlo->nr_elements;
+ unsigned int se_idx;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ nr_elems /= 2;
+ se_idx = zone_no % nr_elems;
+ zlo->elements[se_idx].nr_zones++;
+
+ zone->wr_se_id = se_idx + 1;
+ zone->rd_se_id = zone->wr_se_id;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {
+ zone->rd_se_id += nr_elems;
+ zlo->elements[zone->rd_se_id - 1].nr_zones =
+ zlo->elements[se_idx].nr_zones;
+ }
+ }
+
if (zone_no < zlo->nr_conv_zones) {
/* Conventional zone file. */
set_bit(ZLOOP_ZONE_CONV, &zone->flags);
@@ -1182,6 +1597,79 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,
return ret;
}
+#define ZLOOP_MIN_STOR_ELEMENTS 2
+#define ZLOOP_MAX_STOR_ELEMENTS 32
+#define ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS 32
+
+static int zloop_create_storage_elements(struct zloop_device *zlo)
+{
+ struct blk_storage_element *se, *paired_se;
+ unsigned int i, nr_elems, nr_elements;
+
+ /*
+ * Calculate the number of storage elements we are going to emulate.
+ * To achieve a somewhat realistic emulation, we want at least 2 storage
+ * elements, and no more than 32, targeting at least 32 zones per
+ * element.
+ */
+ if (zlo->nr_zones <= 64)
+ nr_elements = 2;
+ else
+ nr_elements =
+ min(ZLOOP_MAX_STOR_ELEMENTS,
+ zlo->nr_zones / ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS);
+
+ /*
+ * If we are emulating pairs of read and write storage elements, we need
+ * double the number of storage element descriptors.
+ */
+ nr_elems = nr_elements;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ nr_elems *= 2;
+ zlo->elements = kzalloc_objs(struct blk_storage_element, nr_elems);
+ if (!zlo->elements)
+ return -ENOMEM;
+
+ for (i = 0, se = zlo->elements; i < nr_elements; i++, se++) {
+ se->id = i + 1;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ se->type = BLK_SE_TYPE_WRITE;
+ else
+ se->type = BLK_SE_TYPE_RDWR;
+ se->status = BLK_SE_STS_OK;
+ }
+
+ /*
+ * If we are emulating pairs of read and write storage elements,
+ * initialize the read elements paired with the write elements we just
+ * initialized.
+ */
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {
+ for (i = 0; i < nr_elements; i++, se++) {
+ se->id = nr_elements + i + 1;
+ paired_se = &zlo->elements[i];
+ se->paired_id = paired_se->id;
+ paired_se->paired_id = se->id;
+ se->type = BLK_SE_TYPE_READ;
+ se->status = BLK_SE_STS_OK;
+ }
+ }
+
+ /*
+ * Make sure we do not allow removing all storage elements as that does
+ * not make any sense. This is consistent with the device advertized
+ * limit of SCSI and ATA devices supporting the storage element
+ * depopulation feature.
+ */
+ zlo->nr_elements = nr_elems;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ zlo->max_nr_removed_elements = zlo->nr_elements - 2;
+ else
+ zlo->max_nr_removed_elements = zlo->nr_elements - 1;
+
+ return 0;
+}
+
static bool zloop_dev_exists(struct zloop_device *zlo)
{
struct file *cnv, *seq;
@@ -1237,6 +1725,10 @@ static int zloop_ctl_add(struct zloop_options *opts)
WRITE_ONCE(zlo->state, Zlo_creating);
spin_lock_init(&zlo->open_zones_lock);
INIT_LIST_HEAD(&zlo->open_zones_lru_list);
+ mutex_init(&zlo->stor_elements_lock);
+ INIT_DELAYED_WORK(&zlo->remove_element_work, zloop_remove_element_work);
+ INIT_DELAYED_WORK(&zlo->restore_elements_work,
+ zloop_restore_elements_work);
ret = mutex_lock_killable(&zloop_ctl_mutex);
if (ret)
@@ -1270,12 +1762,19 @@ static int zloop_ctl_add(struct zloop_options *opts)
if (zlo->zone_append)
zlo->ordered_zone_append = opts->ordered_zone_append;
zlo->discard_write_cache = opts->discard_write_cache;
+ zlo->stor_elements = opts->stor_elements;
+
+ if (zlo->stor_elements != ZLOOP_STOR_ELEMENTS_NONE) {
+ ret = zloop_create_storage_elements(zlo);
+ if (ret)
+ goto out_free_idr;
+ }
zlo->workqueue = alloc_workqueue("zloop%d", WQ_UNBOUND | WQ_FREEZABLE,
opts->nr_queues * opts->queue_depth, zlo->id);
if (!zlo->workqueue) {
ret = -ENOMEM;
- goto out_free_idr;
+ goto out_destroy_storage_elements;
}
if (opts->base_dir)
@@ -1361,6 +1860,10 @@ static int zloop_ctl_add(struct zloop_options *opts)
zlo->id, zlo->nr_zones,
((sector_t)zlo->zone_size << SECTOR_SHIFT) >> 20,
zlo->block_size);
+ if (zlo->nr_elements)
+ pr_info("zloop%d: %d storage elements\n",
+ zlo->id, zlo->nr_elements);
+
pr_info("zloop%d: using %s%s zone append\n",
zlo->id,
zlo->ordered_zone_append ? "ordered " : "",
@@ -1384,6 +1887,8 @@ static int zloop_ctl_add(struct zloop_options *opts)
kfree(zlo->base_dir);
out_destroy_workqueue:
destroy_workqueue(zlo->workqueue);
+out_destroy_storage_elements:
+ kfree(zlo->elements);
out_free_idr:
mutex_lock(&zloop_ctl_mutex);
idr_remove(&zloop_index_idr, zlo->id);
@@ -1483,6 +1988,42 @@ static int zloop_ctl_remove(struct zloop_options *opts)
return 0;
}
+static int zloop_ctl_degrade_element(struct zloop_options *opts)
+{
+ struct zloop_device *zlo;
+ int ret = 0;
+
+ if (!(opts->mask & ZLOOP_OPT_ID)) {
+ pr_err("No ID specified for degrade_element\n");
+ return -EINVAL;
+ }
+
+ if (opts->mask & ~(ZLOOP_OPT_ID | ZLOOP_OPT_ELEMENT_ID)) {
+ pr_err("Invalid option specified for degrade_element\n");
+ return -EINVAL;
+ }
+
+ mutex_lock(&zloop_ctl_mutex);
+
+ zlo = idr_find(&zloop_index_idr, opts->id);
+ if (!zlo || zlo->state == Zlo_creating)
+ ret = -ENODEV;
+ else if (zlo->state == Zlo_deleting)
+ ret = -EINVAL;
+ if (ret)
+ goto unlock;
+
+ ret = zloop_degrade_element(zlo, opts->element_id);
+ if (!ret)
+ pr_info("Degraded element %u of device %u\n",
+ opts->id, opts->element_id);
+
+unlock:
+ mutex_unlock(&zloop_ctl_mutex);
+
+ return ret;
+}
+
static int zloop_parse_options(struct zloop_options *opts, const char *buf)
{
substring_t args[MAX_OPT_ARGS];
@@ -1502,6 +2043,8 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
opts->buffered_io = ZLOOP_DEF_BUFFERED_IO;
opts->zone_append = ZLOOP_DEF_ZONE_APPEND;
opts->ordered_zone_append = ZLOOP_DEF_ORDERED_ZONE_APPEND;
+ opts->stor_elements = ZLOOP_DEF_STOR_ELEMENTS;
+ opts->element_id = ZLOOP_DEF_ELEMENT_ID;
if (!buf)
return 0;
@@ -1636,6 +2179,30 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
case ZLOOP_OPT_DISCARD_WRITE_CACHE:
opts->discard_write_cache = true;
break;
+ case ZLOOP_OPT_STOR_ELEMENTS:
+ if (match_uint(args, &token)) {
+ ret = -EINVAL;
+ goto out;
+ }
+ switch (token) {
+ case ZLOOP_STOR_ELEMENTS_NONE:
+ case ZLOOP_STOR_ELEMENTS_RDWR:
+ case ZLOOP_STOR_ELEMENTS_PAIRS:
+ break;
+ default:
+ pr_err("Invalid stor_elements value\n");
+ ret = -EINVAL;
+ goto out;
+ }
+ opts->stor_elements = token;
+ break;
+ case ZLOOP_OPT_ELEMENT_ID:
+ if (match_uint(args, &token)) {
+ ret = -EINVAL;
+ goto out;
+ }
+ opts->element_id = token;
+ break;
case ZLOOP_OPT_ERR:
default:
pr_warn("unknown parameter or missing value '%s'\n", p);
@@ -1664,14 +2231,16 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
enum {
ZLOOP_CTL_ADD,
ZLOOP_CTL_REMOVE,
+ ZLOOP_CTL_DEGRADE_ELEMENT,
};
static struct zloop_ctl_op {
int code;
const char *name;
} zloop_ctl_ops[] = {
- { ZLOOP_CTL_ADD, "add" },
- { ZLOOP_CTL_REMOVE, "remove" },
+ { ZLOOP_CTL_ADD, "add" },
+ { ZLOOP_CTL_REMOVE, "remove" },
+ { ZLOOP_CTL_DEGRADE_ELEMENT, "degrade_element" },
{ -1, NULL },
};
@@ -1719,6 +2288,9 @@ static ssize_t zloop_ctl_write(struct file *file, const char __user *ubuf,
case ZLOOP_CTL_REMOVE:
ret = zloop_ctl_remove(&opts);
break;
+ case ZLOOP_CTL_DEGRADE_ELEMENT:
+ ret = zloop_ctl_degrade_element(&opts);
+ break;
default:
pr_err("Invalid operation\n");
ret = -EINVAL;
@@ -1742,6 +2314,8 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)
tok = &zloop_opt_tokens[i];
if (!tok->pattern)
break;
+ if (tok->token == ZLOOP_OPT_ELEMENT_ID)
+ continue;
if (i)
seq_putc(seq_file, ',');
seq_puts(seq_file, tok->pattern);
@@ -1752,6 +2326,10 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)
seq_puts(seq_file, zloop_ctl_ops[1].name);
seq_puts(seq_file, " id=%d\n");
+ /* Degrade element operation */
+ seq_puts(seq_file, zloop_ctl_ops[2].name);
+ seq_puts(seq_file, " id=%d,element_id=%d\n");
+
return 0;
}
diff --git a/drivers/scsi/sd.c b/drivers/scsi/sd.c
index a1b21ea14e549..4372fa80f792d 100644
--- a/drivers/scsi/sd.c
+++ b/drivers/scsi/sd.c
@@ -3937,6 +3937,9 @@ static const struct block_device_operations sd_fops = {
.get_unique_id = sd_get_unique_id,
.free_disk = scsi_disk_free_disk,
.pr_ops = &sd_pr_ops,
+#ifdef CONFIG_BLK_DEV_ZONED
+ .se_ops = &sd_zbc_se_ops,
+#endif
};
/**
diff --git a/drivers/scsi/sd.h b/drivers/scsi/sd.h
index 574af82430169..b68b7ee5a5fa5 100644
--- a/drivers/scsi/sd.h
+++ b/drivers/scsi/sd.h
@@ -156,6 +156,7 @@ struct scsi_disk {
unsigned ignore_medium_access_errors : 1;
unsigned rscs : 1; /* reduced stream control support */
unsigned use_atomic_write_boundary : 1;
+ unsigned modify_zones_supported : 1;
};
#define to_scsi_disk(obj) container_of(obj, struct scsi_disk, disk_dev)
@@ -242,6 +243,8 @@ unsigned int sd_zbc_complete(struct scsi_cmnd *cmd, unsigned int good_bytes,
int sd_zbc_report_zones(struct gendisk *disk, sector_t sector,
unsigned int nr_zones, struct blk_report_zones_args *args);
+extern const struct blk_storage_elements_ops sd_zbc_se_ops;
+
#else /* CONFIG_BLK_DEV_ZONED */
static inline int sd_zbc_read_zones(struct scsi_disk *sdkp,
diff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c
index 456beaf2e7690..6698d156a0c34 100644
--- a/drivers/scsi/sd_zbc.c
+++ b/drivers/scsi/sd_zbc.c
@@ -516,6 +516,21 @@ static int sd_zbc_check_capacity(struct scsi_disk *sdkp, unsigned char *buf,
return 0;
}
+/*
+ * sd_zbc_check_modify_zones - Check if the device supports depopulation
+ * @sdkp: Target disk
+ * @buf: command buffer
+ *
+ * Check if the device supports the REMOVE ELEMENT AND MODIFY ZONES command.
+ */
+static inline bool sd_zbc_check_modify_zones(struct scsi_disk *sdkp,
+ unsigned char *buf)
+{
+ return scsi_report_opcode(sdkp->device, buf, SD_BUF_SIZE,
+ SERVICE_ACTION_IN_16,
+ SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES) == 1;
+}
+
static void sd_zbc_print_zones(struct scsi_disk *sdkp)
{
if (sdkp->device->type != TYPE_ZBC || !sdkp->capacity)
@@ -533,6 +548,239 @@ static void sd_zbc_print_zones(struct scsi_disk *sdkp)
sdkp->zone_info.zone_blocks);
}
+static void sd_zbc_parse_storage_element(struct scsi_disk *sdkp, u8 *desc,
+ struct blk_storage_element *element)
+{
+ struct scsi_device *sdp = sdkp->device;
+ sector_t zone_sectors = sd_zbc_zone_sectors(sdkp);
+ u64 capacity;
+
+ memset(element, 0, sizeof(*element));
+
+ element->id = get_unaligned_be32(&desc[4]);
+
+ switch (desc[14]) {
+ case SCSI_PHYS_ELEM_TYPE_ALL_ACCESS_STORAGE:
+ element->type = BLK_SE_TYPE_RDWR;
+ capacity = get_unaligned_be64(&desc[16]);
+ if (!zone_sectors || capacity == ULLONG_MAX)
+ element->nr_zones = 0;
+ else
+ element->nr_zones =
+ logical_to_sectors(sdp, capacity) >>
+ ilog2(zone_sectors);
+ break;
+ case SCSI_PHYS_ELEM_TYPE_FRAC_ACCESS_STORAGE:
+ element->paired_id = get_unaligned_be32(&desc[16]);
+ if (desc[20] & 0x01)
+ element->type = BLK_SE_TYPE_READ;
+ else
+ element->type = BLK_SE_TYPE_WRITE;
+ element->nr_zones = get_unaligned_be64(&desc[24]);
+ break;
+ default:
+ element->type = BLK_SE_TYPE_UNKNOWN;
+ }
+
+ switch (desc[15]) {
+ case SCSI_PHYS_ELEM_HEALTH_WITHIN_SPEC_LIMITS:
+ case SCSI_PHYS_ELEM_HEALTH_AT_SPEC_LIMITS:
+ element->status = BLK_SE_STS_OK;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_OUTSIDE_SPEC_LIMITS:
+ element->status = BLK_SE_STS_DEGRADED;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_ERR:
+ element->status = BLK_SE_STS_RESTORE_ERROR;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_IN_PROGRESS:
+ element->status = BLK_SE_STS_RESTORE_IN_PROGRESS;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_ERR:
+ element->status = BLK_SE_STS_REMOVE_ERROR;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_IN_PROGRESS:
+ element->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_OK:
+ element->status = BLK_SE_STS_REMOVED;
+ element->restore_allowed = desc[13] & 0x01;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_NOT_REPORTED:
+ default:
+ element->status = BLK_SE_STS_UNKNOWN;
+ break;
+ }
+}
+
+/*
+ * Large hard limit on the number of storage elements. This accommodates all
+ * known devices today and likely forever :)
+ */
+#define SD_ZBC_MAX_STORAGE_ELEMENTS 255
+
+static int sd_zbc_report_storage_elements(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ unsigned int nr_descs, nr_se;
+ unsigned int buf_size;
+ int i, ret = 0, result;
+ u8 *desc, *buf;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ /*
+ * We need at least 32B for the report header and 32B for each
+ * descriptor.
+ */
+ nr_se = min(SD_ZBC_MAX_STORAGE_ELEMENTS, *nr_elements);
+ buf_size = ALIGN((nr_se + 1) * 32, SECTOR_SIZE);
+
+ buf = kzalloc(buf_size, GFP_KERNEL);
+ if (!buf)
+ return -ENOMEM;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_GET_PHYSICAL_ELEMENT_STATUS;
+ put_unaligned_be32(buf_size, &cmd[10]);
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, buf, buf_size,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "GET PHYSICAL ELEMENT STATUS failed\n");
+ sd_print_result(sdkp, "GET PHYSICAL ELEMENT STATUS", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ nr_descs = get_unaligned_be32(&buf[0]);
+ if (!nr_descs) {
+ sd_printk(KERN_ERR, sdkp,
+ "Invalid number of phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+ if (nr_descs > SD_ZBC_MAX_STORAGE_ELEMENTS) {
+ sd_printk(KERN_ERR, sdkp,
+ "Unsupported number of phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ if (!elements) {
+ *nr_elements = nr_descs;
+ goto free_buf;
+ }
+
+ nr_descs = get_unaligned_be32(&buf[4]);
+ if (!nr_descs) {
+ sd_printk(KERN_ERR, sdkp,
+ "Invalid number of reported phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ desc = &buf[32];
+ for (i = 0; i < min(nr_se, nr_descs); i++, desc += 32)
+ sd_zbc_parse_storage_element(sdkp, desc, &elements[i]);
+ *nr_elements = i;
+
+free_buf:
+ kfree(buf);
+
+ return ret;
+}
+
+static int sd_zbc_remove_storage_element(struct gendisk *disk,
+ unsigned int element_id)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ int result;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES;
+ put_unaligned_be32(element_id, &cmd[10]);
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "REMOVE ELEMENT AND MODIFY ZONES failed\n");
+ sd_print_result(sdkp,
+ "REMOVE ELEMENT AND MODIFY ZONES", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+static int sd_zbc_restore_storage_elements(struct gendisk *disk)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ int result;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_RESTORE_ELEMENTS_AND_REBUILD;
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "RESTORE ELEMENTS AND REBUILD failed\n");
+ sd_print_result(sdkp,
+ "RESTORE ELEMENTS AND REBUILD", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+const struct blk_storage_elements_ops sd_zbc_se_ops = {
+ .report_elements = sd_zbc_report_storage_elements,
+ .remove_element = sd_zbc_remove_storage_element,
+ .restore_elements = sd_zbc_restore_storage_elements,
+};
+
/*
* Call blk_revalidate_disk_zones() if any of the zoned disk properties have
* changed that make it necessary to call that function. Called by
@@ -554,9 +802,15 @@ int sd_zbc_revalidate_zones(struct scsi_disk *sdkp)
if (!blk_queue_is_zoned(q))
return 0;
+ /*
+ * If the zone size and number of zones has not changed, and the disk
+ * does not support depopulating heads, skip the rather slow call to
+ * blk_revalidate_disk_zones().
+ */
if (sdkp->zone_info.zone_blocks == zone_blocks &&
sdkp->zone_info.nr_zones == nr_zones &&
- disk->nr_zones == nr_zones)
+ disk->nr_zones == nr_zones &&
+ !sdkp->modify_zones_supported)
return 0;
sdkp->zone_info.zone_blocks = zone_blocks;
@@ -620,6 +874,9 @@ int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim,
if (ret != 0)
goto err;
+ /* Check if REMOVE ELEMENT AND MODIFY ZONES is supported. */
+ sdkp->modify_zones_supported = sd_zbc_check_modify_zones(sdkp, buf);
+
nr_zones = round_up(sdkp->capacity, zone_blocks) >> ilog2(zone_blocks);
if (nr_zones > INT_MAX) {
sd_printk(KERN_ERR, sdkp, "Too many zones (%llu)\n",
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index d003a9d2d1f6c..859917b3b2568 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1574,6 +1574,21 @@ enum blk_unique_id {
BLK_UID_NAA = 3,
};
+struct blk_storage_elements_ops {
+ int (*report_elements)(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements);
+ int (*remove_element)(struct gendisk *disk, unsigned int element_id);
+ int (*restore_elements)(struct gendisk *disk);
+};
+
+int bdev_report_storage_elements(struct block_device *bdev,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements);
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id);
+int bdev_restore_storage_elements(struct block_device *bdev);
+
struct block_device_operations {
void (*submit_bio)(struct bio *bio);
int (*poll_bio)(struct bio *bio, struct io_comp_batch *iob,
@@ -1601,6 +1616,7 @@ struct block_device_operations {
enum blk_unique_id id_type);
struct module *owner;
const struct pr_ops *pr_ops;
+ const struct blk_storage_elements_ops *se_ops;
/*
* Special callback for probing GPT entry at a given sector.
diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
index 6638361209667..532cf69b5b114 100644
--- a/include/uapi/linux/blkzoned.h
+++ b/include/uapi/linux/blkzoned.h
@@ -208,4 +208,85 @@ struct blk_zone_range {
#define BLKFINISHZONE _IOW(0x12, 136, struct blk_zone_range)
#define BLKREPORTZONEV2 _IOWR(0x12, 142, struct blk_zone_report)
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_TYPE_RDWR: The storage element handles both reads and writes.
+ * @BLK_SE_TYPE_READ: The storage element handles reads only.
+ * @BLK_SE_TYPE_WRITE: The storage element handles writes only.
+ * @BLK_SE_TYPE_UNKNOWN: The storage element type is not known.
+ */
+enum blk_storage_element_type {
+ BLK_SE_TYPE_RDWR = 0x01,
+ BLK_SE_TYPE_READ = 0x02,
+ BLK_SE_TYPE_WRITE = 0x03,
+ BLK_SE_TYPE_UNKNOWN = 0xFF,
+};
+
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_STS_OK: The storage element is operating normally.
+ * @BLK_SE_STS_DEGRADED: The storage element has degraded and is not operating
+ * normally.
+ * @BLK_SE_STS_REMOVE_IN_PROGRESS: The storage element is being removed.
+ * @BLK_SE_STS_REMOVE_ERROR: The storage element removal failed.
+ * @BLK_SE_STS_RESTORE_IN_PROGRESS: The storage element is being restored.
+ * @BLK_SE_STS_RESTORE_ERROR: The storage element restoration failed.
+ * @BLK_SE_STS_REMOVED: The storage element was removed.
+ * @BLK_SE_STS_UNKNOWN: The storage element status is unknown.
+ */
+enum blk_storage_element_status {
+ BLK_SE_STS_OK = 0x01,
+ BLK_SE_STS_DEGRADED = 0x02,
+ BLK_SE_STS_REMOVE_IN_PROGRESS = 0x03,
+ BLK_SE_STS_REMOVE_ERROR = 0x04,
+ BLK_SE_STS_RESTORE_IN_PROGRESS = 0x05,
+ BLK_SE_STS_RESTORE_ERROR = 0x06,
+ BLK_SE_STS_REMOVED = 0x07,
+ BLK_SE_STS_UNKNOWN = 0xFF,
+};
+
+/**
+ * struct blk_storage_element - Zoned device storage element descriptor.
+ *
+ * @id: The ID of the element (cannot be 0).
+ * @paired_id: The ID of the paired element for an element that is not
+ * of type BLK_SE_TYPE_RDWR.
+ * @type: The type of the storage element (enum blk_storage_element_type).
+ * @status: The health status of the storage element
+ * (enum blk_storage_element_status).
+ * @restore_allowed: Indicate if the storage element can be restored.
+ * @nr_zones: The number of zones that the storage element handles.
+ */
+struct blk_storage_element {
+ __u32 id;
+ __u32 paired_id;
+ __u64 nr_zones;
+ __u8 type;
+ __u8 status;
+ __u8 restore_allowed;
+ __u8 reserved[5];
+};
+
+struct blk_storage_elements_report {
+ __u32 nr_elements;
+ __u32 reserved;
+ struct blk_storage_element elements[];
+};
+
+/**
+ * Zoned block device storage element management ioctl's:
+ *
+ * @BLKGETNRSTORELEMS: Get the number of storage elements of the device.
+ * @BLKREPORTSTORELEMS: Get the device storage elements.
+ * @BLKREMOVESTORELEM: Remove (depopulate) one storage element of a device.
+ * @BLKRESTORESTORELEM: Restore (repopulate if possible) all storage elements
+ * that have been removed.
+ */
+#define BLKGETNRSTORELEMS _IOR(0x12, 143, __u32)
+#define BLKREPORTSTORELEMS _IOWR(0x12, 144, struct blk_storage_elements_report)
+#define BLKREMOVESTORELEM _IOW(0x12, 145, __u32)
+#define BLKRESTORESTORELEMS _IO(0x12, 146)
+
#endif /* _UAPI_BLKZONED_H */
diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index 34c6f219462a5..8a979326aa7f0 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -309,6 +309,7 @@ struct file_attr {
/* 130-136 and 142 are used by zoned block device ioctls (uapi/linux/blkzoned.h) */
/* 137-141 are used by blk-crypto ioctls (uapi/linux/blk-crypto.h) */
#define BLKTRACESETUP2 _IOWR(0x12, 142, struct blk_user_trace_setup2)
+/* 143-146 are used by storage element management for zoned block devices. */
#define BMAP_IOCTL 1 /* obsolete - kept for compatibility */
#define FIBMAP _IO(0x00,1) /* bmap access */
|
| 3/2 | 2026/10/06 13:26 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 26101 output: 32 thoughts: 1259 Thoughts: **Fuzzing Evaluation: Zoned Block Device Enhancements** Okay, let's break this down. My goal is to determine if this patch series is a good candidate for syzkaller fuzzing. The instructions are clear – reachability is key. Can it run within a standard virtualized environment? Absolutely, this looks like it's designed to. We've got `zloop`, which is a software emulated block device, and potentially the `scsi_debug` driver. Both are prime targets for fuzzing within QEMU. The "WorthFuzzing" criteria boil down to whether the modified code is executable, reachable code. This series clearly hits that mark. It's touching core subsystems: the block layer (`blk-core.c`, `blk-zoned.c`, `blk.h`, `ioctl.c`, `blkdev.h`, `blkzoned.h`, `fs.h`), and the `zloop` driver is a driver we use! The addition of new ioctls (`BLKGETNRSTORELEMS`, etc.) also makes it very appealing. Looking deeper, the changes introduce new ioctls to interact with zoned block devices. `zloop` and `sd_zbc` (ZBC SCSI) support the management of storage elements, which sounds like an interesting feature to fuzz. The fact that any user or root in syzkaller can interact with these via block devices or the `/dev/zloop-control` interface is very significant. We can easily trigger these ioctls within the fuzzer. Now, I'll need to figure out which Kconfig flags are necessary to enable all of these components. `zloop.c`, `blk-zoned.c`, and `sd_zbc.c` are definitely the Kconfig considerations. Finally, I'll also need to identify any specific functions or areas of the code that would be particularly useful to focus the fuzzer on. |
| 4/2 | 2026/10/06 13:27 | tool | grepper |
0mArgs:{
"Expression": "config BLK_DEV_ZONED",
"PathPrefixes": [
"block/"
]
}Results: {
"Output": "block/Kconfig=77=config BLK_DEV_WRITE_MOUNTED\n--\nblock/Kconfig-96-\nblock/Kconfig:97:config BLK_DEV_ZONED\nblock/Kconfig-98-\tbool \"Zoned block device support\"\n"
} |
| 5/2 | 2026/10/06 13:27 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 27882 output: 27 thoughts: 10 Thoughts: Okay, here's my attempt at summarizing that thought process, tailored for an expert audience: **Examining Kconfig for zloop** Right, let's get down to brass tacks. I need to understand the current configuration options related to the "zloop" functionality. The most logical place to start digging is the Kconfig files. They're the source of truth for the build system and the primary mechanism for controlling kernel features. I need to quickly identify all relevant Kconfig entries, evaluate their dependencies, and understand any associated default values or help text. This will help me confirm my assumptions about the current state, and see what the options are to change things. Let's see what we can find. |
| 6/2 | 2026/10/06 13:27 | tool | grepper |
0mArgs:{
"Expression": "ZLOOP",
"PathPrefixes": [
"drivers/block/"
]
}Results: {
"Output": "\nFull output is too long, showing 200 out of 390 lines.\nUse more precise expression or PathPrefixes if possible.\n\nNumber of matching lines per file (1 files in total):\ndrivers/block/zloop.c:136\n\ndrivers/block/zloop.c=25=enum {\ndrivers/block/zloop.c:26:\tZLOOP_OPT_ERR\t\t\t= 0,\ndrivers/block/zloop.c:27:\tZLOOP_OPT_ID\t\t\t= (1 \u003c\u003c 0),\ndrivers/block/zloop.c:28:\tZLOOP_OPT_CAPACITY\t\t= (1 \u003c\u003c 1),\ndrivers/block/zloop.c:29:\tZLOOP_OPT_ZONE_SIZE\t\t= (1 \u003c\u003c 2),\ndrivers/block/zloop.c:30:\tZLOOP_OPT_ZONE_CAPACITY\t\t= (1 \u003c\u003c 3),\ndrivers/block/zloop.c:31:\tZLOOP_OPT_NR_CONV_ZONES\t\t= (1 \u003c\u003c 4),\ndrivers/block/zloop.c:32:\tZLOOP_OPT_BASE_DIR\t\t= (1 \u003c\u003c 5),\ndrivers/block/zloop.c:33:\tZLOOP_OPT_NR_QUEUES\t\t= (1 \u003c\u003c 6),\ndrivers/block/zloop.c:34:\tZLOOP_OPT_QUEUE_DEPTH\t\t= (1 \u003c\u003c 7),\ndrivers/block/zloop.c:35:\tZLOOP_OPT_BUFFERED_IO\t\t= (1 \u003c\u003c 8),\ndrivers/block/zloop.c:36:\tZLOOP_OPT_ZONE_APPEND\t\t= (1 \u003c\u003c 9),\ndrivers/block/zloop.c:37:\tZLOOP_OPT_ORDERED_ZONE_APPEND\t= (1 \u003c\u003c 10),\ndrivers/block/zloop.c:38:\tZLOOP_OPT_DISCARD_WRITE_CACHE\t= (1 \u003c\u003c 11),\ndrivers/block/zloop.c:39:\tZLOOP_OPT_MAX_OPEN_ZONES\t= (1 \u003c\u003c 12),\ndrivers/block/zloop.c:40:\tZLOOP_OPT_STOR_ELEMENTS\t\t= (1 \u003c\u003c 13),\ndrivers/block/zloop.c:41:\tZLOOP_OPT_ELEMENT_ID\t\t= (1 \u003c\u003c 14),\ndrivers/block/zloop.c-42-};\n--\ndrivers/block/zloop.c=44=static const match_table_t zloop_opt_tokens = {\ndrivers/block/zloop.c:45:\t{ ZLOOP_OPT_ID,\t\t\t\"id=%d\"\t},\ndrivers/block/zloop.c:46:\t{ ZLOOP_OPT_CAPACITY,\t\t\"capacity_mb=%u\"\t},\ndrivers/block/zloop.c:47:\t{ ZLOOP_OPT_ZONE_SIZE,\t\t\"zone_size_mb=%u\"\t},\ndrivers/block/zloop.c:48:\t{ ZLOOP_OPT_ZONE_CAPACITY,\t\"zone_capacity_mb=%u\"\t},\ndrivers/block/zloop.c:49:\t{ ZLOOP_OPT_NR_CONV_ZONES,\t\"conv_zones=%u\"\t\t},\ndrivers/block/zloop.c:50:\t{ ZLOOP_OPT_BASE_DIR,\t\t\"base_dir=%s\"\t\t},\ndrivers/block/zloop.c:51:\t{ ZLOOP_OPT_NR_QUEUES,\t\t\"nr_queues=%u\"\t\t},\ndrivers/block/zloop.c:52:\t{ ZLOOP_OPT_QUEUE_DEPTH,\t\"queue_depth=%u\"\t},\ndrivers/block/zloop.c:53:\t{ ZLOOP_OPT_BUFFERED_IO,\t\"buffered_io\"\t\t},\ndrivers/block/zloop.c:54:\t{ ZLOOP_OPT_ZONE_APPEND,\t\"zone_append=%u\"\t},\ndrivers/block/zloop.c:55:\t{ ZLOOP_OPT_ORDERED_ZONE_APPEND, \"ordered_zone_append\"\t},\ndrivers/block/zloop.c:56:\t{ ZLOOP_OPT_DISCARD_WRITE_CACHE, \"discard_write_cache\"\t},\ndrivers/block/zloop.c:57:\t{ ZLOOP_OPT_MAX_OPEN_ZONES,\t\"max_open_zones=%u\"\t},\ndrivers/block/zloop.c:58:\t{ ZLOOP_OPT_STOR_ELEMENTS,\t\"stor_elements=%u\"\t},\ndrivers/block/zloop.c:59:\t{ ZLOOP_OPT_ELEMENT_ID,\t\t\"element_id=%u\"\t\t},\ndrivers/block/zloop.c:60:\t{ ZLOOP_OPT_ERR,\t\tNULL\t\t\t}\ndrivers/block/zloop.c-61-};\n--\ndrivers/block/zloop.c=64=enum zloop_stor_elements {\ndrivers/block/zloop.c-65-\t/* No emulation. */\ndrivers/block/zloop.c:66:\tZLOOP_STOR_ELEMENTS_NONE,\ndrivers/block/zloop.c-67-\t/* Emulate read+write storage elements. */\ndrivers/block/zloop.c:68:\tZLOOP_STOR_ELEMENTS_RDWR,\ndrivers/block/zloop.c-69-\t/* Emulate pairs of associated read and write storage elements. */\ndrivers/block/zloop.c:70:\tZLOOP_STOR_ELEMENTS_PAIRS,\ndrivers/block/zloop.c-71-};\n--\ndrivers/block/zloop.c-73-/* Default values for the \"add\" operation. */\ndrivers/block/zloop.c:74:#define ZLOOP_DEF_ID\t\t\t-1\ndrivers/block/zloop.c:75:#define ZLOOP_DEF_ZONE_SIZE\t\t((256ULL * SZ_1M) \u003e\u003e SECTOR_SHIFT)\ndrivers/block/zloop.c:76:#define ZLOOP_DEF_NR_ZONES\t\t64\ndrivers/block/zloop.c:77:#define ZLOOP_DEF_NR_CONV_ZONES\t\t8\ndrivers/block/zloop.c:78:#define ZLOOP_DEF_MAX_OPEN_ZONES\t0\ndrivers/block/zloop.c:79:#define ZLOOP_DEF_BASE_DIR\t\t\"/var/local/zloop\"\ndrivers/block/zloop.c:80:#define ZLOOP_DEF_NR_QUEUES\t\t1\ndrivers/block/zloop.c:81:#define ZLOOP_DEF_QUEUE_DEPTH\t\t128\ndrivers/block/zloop.c:82:#define ZLOOP_DEF_BUFFERED_IO\t\tfalse\ndrivers/block/zloop.c:83:#define ZLOOP_DEF_ZONE_APPEND\t\ttrue\ndrivers/block/zloop.c:84:#define ZLOOP_DEF_ORDERED_ZONE_APPEND\tfalse\ndrivers/block/zloop.c:85:#define ZLOOP_DEF_STOR_ELEMENTS\t\tZLOOP_STOR_ELEMENTS_NONE\ndrivers/block/zloop.c:86:#define ZLOOP_DEF_ELEMENT_ID\t\t0\ndrivers/block/zloop.c-87-\ndrivers/block/zloop.c-88-/* Arbitrary limit on the zone size (16GB). */\ndrivers/block/zloop.c:89:#define ZLOOP_MAX_ZONE_SIZE_MB\t\t16384\ndrivers/block/zloop.c-90-\n--\ndrivers/block/zloop.c=119=enum zloop_zone_flags {\ndrivers/block/zloop.c:120:\tZLOOP_ZONE_CONV = 0,\ndrivers/block/zloop.c:121:\tZLOOP_ZONE_SEQ_ERROR,\ndrivers/block/zloop.c-122-};\n--\ndrivers/block/zloop.c=350=static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)\n--\ndrivers/block/zloop.c-362-\t\t zone_no, ret);\ndrivers/block/zloop.c:363:\t\tset_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\ndrivers/block/zloop.c-364-\t\treturn ret;\n--\ndrivers/block/zloop.c=409=static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)\n--\ndrivers/block/zloop.c-413-\ndrivers/block/zloop.c:414:\tif (test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags))\ndrivers/block/zloop.c-415-\t\treturn -EIO;\n--\ndrivers/block/zloop.c-423-\ndrivers/block/zloop.c:424:\tif (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags)) {\ndrivers/block/zloop.c-425-\t\tret = zloop_update_seq_zone(zlo, zone_no);\n--\ndrivers/block/zloop.c=439=static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)\n--\ndrivers/block/zloop.c-443-\ndrivers/block/zloop.c:444:\tif (test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags))\ndrivers/block/zloop.c-445-\t\treturn -EIO;\n--\ndrivers/block/zloop.c-453-\ndrivers/block/zloop.c:454:\tif (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags)) {\ndrivers/block/zloop.c-455-\t\tret = zloop_update_seq_zone(zlo, zone_no);\n--\ndrivers/block/zloop.c=486=static int zloop_do_reset_zone(struct zloop_device *zlo,\n--\ndrivers/block/zloop.c-493-\tzloop_set_zone_empty(zlo, zone);\ndrivers/block/zloop.c:494:\tclear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\ndrivers/block/zloop.c-495-\tspin_unlock(\u0026zone-\u003ewp_lock);\n--\ndrivers/block/zloop.c=500=static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no,\n--\ndrivers/block/zloop.c-505-\ndrivers/block/zloop.c:506:\tif (test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags))\ndrivers/block/zloop.c-507-\t\treturn -EIO;\n--\ndrivers/block/zloop.c-516-\ndrivers/block/zloop.c:517:\tif (!test_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags) \u0026\u0026\ndrivers/block/zloop.c-518-\t zone-\u003econd == BLK_ZONE_COND_EMPTY)\n--\ndrivers/block/zloop.c-522-\tif (ret)\ndrivers/block/zloop.c:523:\t\tset_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\ndrivers/block/zloop.c-524-\n--\ndrivers/block/zloop.c=545=static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)\n--\ndrivers/block/zloop.c-549-\ndrivers/block/zloop.c:550:\tif (test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags))\ndrivers/block/zloop.c-551-\t\treturn -EIO;\n--\ndrivers/block/zloop.c-559-\ndrivers/block/zloop.c:560:\tif (!test_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags) \u0026\u0026\ndrivers/block/zloop.c-561-\t zone-\u003econd == BLK_ZONE_COND_FULL)\n--\ndrivers/block/zloop.c-565-\t\t\t zlo-\u003ezone_capacity \u003c\u003c SECTOR_SHIFT)) {\ndrivers/block/zloop.c:566:\t\tset_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\ndrivers/block/zloop.c-567-\t\tret = -EIO;\n--\ndrivers/block/zloop.c-572-\tzloop_set_zone_full(zlo, zone);\ndrivers/block/zloop.c:573:\tclear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\ndrivers/block/zloop.c-574-\tspin_unlock(\u0026zone-\u003ewp_lock);\n--\ndrivers/block/zloop.c=716=static void zloop_rw(struct zloop_cmd *cmd)\n--\ndrivers/block/zloop.c-748-\ndrivers/block/zloop.c:749:\tif (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags)) {\ndrivers/block/zloop.c-750-\t\tmutex_lock(\u0026zone-\u003elock);\n--\ndrivers/block/zloop.c-756-\ndrivers/block/zloop.c:757:\tif (!test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags) \u0026\u0026 is_write) {\ndrivers/block/zloop.c-758-\t\tmutex_lock(\u0026zone-\u003elock);\n--\ndrivers/block/zloop.c=906=static void zloop_complete_rq(struct request *rq)\n--\ndrivers/block/zloop.c-937-\ndrivers/block/zloop.c:938:\t\tif (cmd-\u003eret \u003c 0 \u0026\u0026 !test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags)) {\ndrivers/block/zloop.c-939-\t\t\t/*\n--\ndrivers/block/zloop.c-943-\t\t\t */\ndrivers/block/zloop.c:944:\t\t\tset_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags);\ndrivers/block/zloop.c-945-\t\t\tbreak;\n--\ndrivers/block/zloop.c=1076=static int zloop_report_zones(struct gendisk *disk, sector_t sector,\n--\ndrivers/block/zloop.c-1094-\ndrivers/block/zloop.c:1095:\t\tif (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, \u0026zone-\u003eflags)) {\ndrivers/block/zloop.c-1096-\t\t\tret = zloop_update_seq_zone(zlo, zone_no);\n--\ndrivers/block/zloop.c-1108-\t\tblkz.cond = zone-\u003econd;\ndrivers/block/zloop.c:1109:\t\tif (test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags)) {\ndrivers/block/zloop.c-1110-\t\t\tblkz.type = BLK_ZONE_TYPE_CONVENTIONAL;\n--\ndrivers/block/zloop.c=1127=static int zloop_report_elements(struct gendisk *disk,\n--\ndrivers/block/zloop.c-1133-\ndrivers/block/zloop.c:1134:\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_NONE)\ndrivers/block/zloop.c-1135-\t\treturn -EOPNOTSUPP;\n--\ndrivers/block/zloop.c=1163=static int zloop_degrade_element(struct zloop_device *zlo,\n--\ndrivers/block/zloop.c-1168-\ndrivers/block/zloop.c:1169:\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_NONE)\ndrivers/block/zloop.c-1170-\t\treturn -EOPNOTSUPP;\n--\ndrivers/block/zloop.c=1246=static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)\n--\ndrivers/block/zloop.c-1252-\ndrivers/block/zloop.c:1253:\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_NONE)\ndrivers/block/zloop.c-1254-\t\treturn -EOPNOTSUPP;\n--\ndrivers/block/zloop.c=1309=static void zloop_restore_elements_work(struct work_struct *work)\n--\ndrivers/block/zloop.c-1322-\t\tmutex_lock(\u0026zone-\u003elock);\ndrivers/block/zloop.c:1323:\t\tif (test_bit(ZLOOP_ZONE_CONV, \u0026zone-\u003eflags)) {\ndrivers/block/zloop.c-1324-\t\t\tspin_lock(\u0026zone-\u003ewp_lock);\n--\ndrivers/block/zloop.c=1349=static int zloop_restore_elements(struct gendisk *disk)\n--\ndrivers/block/zloop.c-1355-\ndrivers/block/zloop.c:1356:\tif (zlo-\u003estor_elements == ZLOOP_STOR_ELEMENTS_NONE)\ndrivers/block/zloop.c-1357-\t\treturn -EOPNOTSUPP;\n--\ndrivers/block/zloop.c=1492=static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,\n--\ndrivers/block/zloop.c-1511-\ndrivers/block/zloop.c:1512:\tif (zlo-\u003estor_elements != ZLOOP_STOR_ELEMENTS_NONE) {\ndrivers/block/zloop.c-1513-\t\tunsigned int nr_elems = zlo-\u003enr_elements;\n\n"
} |
| 7/2 | 2026/10/06 13:27 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32532 output: 37 thoughts: 14 Thoughts: Okay, here's my summary, presented as a stream of thought: **Finding the Right Kconfig Option for zloop.c** Alright, let's get down to brass tacks. I need to pinpoint the Kconfig option that controls the configuration for `zloop.c`. No time to waste on guesswork; I need the precise option. It's the key to controlling that module. I'm imagining it's probably linked to a specific feature or subsystem. I need to think systematically. Is it related to networking? Security? Power management? Or perhaps it's more general purpose, likely under a more common config grouping. I should probably start by doing a quick search through the `Kconfig` files, focusing on likely areas, and using a key phrase like "zloop". I know my way around the kernel source, so I know where the usual suspects are. This should be straightforward. Let's get to it. |
| 8/2 | 2026/10/06 13:27 | tool | grepper |
0mArgs:{
"Expression": "zloop",
"PathPrefixes": [
"drivers/block/Makefile",
"drivers/block/Kconfig"
]
}Results: {
"Output": "drivers/block/Kconfig=356=config BLK_DEV_ZONED_LOOP\n--\ndrivers/block/Kconfig-367-\t echo \"add id=0,zone_size_mb=256,capacity_mb=16384,conv_zones=11\" \u003e \\\ndrivers/block/Kconfig:368:\t\t/dev/zloop-control\ndrivers/block/Kconfig-369-\n--\ndrivers/block/Makefile=40=obj-$(CONFIG_BLK_DEV_UBLK)\t\t\t+= ublk_drv.o\ndrivers/block/Makefile:41:obj-$(CONFIG_BLK_DEV_ZONED_LOOP) += zloop.o\ndrivers/block/Makefile-42-\n"
} |
| 9/2 | 2026/10/06 13:27 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32971 output: 34 thoughts: 33 Thoughts: Okay, let's dive into this. Here's how I'd approach that kernel configuration investigation, and the thought process behind it: **Investigating `BLK_DEV_ZONED_LOOP` Dependencies in `Kconfig`** Alright, I need to understand the dependencies for `BLK_DEV_ZONED_LOOP`. My first instinct is to go directly to the source: `drivers/block/Kconfig`. I know that's where the kernel configuration options, their descriptions, and their dependencies are defined for block devices, so that's where I need to start. I'm going to jump to around line 356, as a preliminary guess based on past experience; it's a decent ballpark for where a newer block device might be defined. Hopefully, `BLK_DEV_ZONED_LOOP` is defined near that point, or at least a related option is, and I can quickly spot the relationships. I'm looking for things like `depends on`, `select`, or potentially even `default` statements that would reveal the prerequisites for this option. I want to build a clear picture of what needs to be enabled or configured upstream for this zoned loop device to function correctly. This is just the first step in debugging the kernel, of course. |
| 10/2 | 2026/10/06 13:27 | tool | read-file |
0mArgs:{
"File": "drivers/block/Kconfig",
"FirstLine": 350,
"LineCount": 30
}Results: {
"Contents": " 350:\t\t Say N if you don't want to support legacy command opcode. It is\n 351:\t\t suggested to enable N if your application(ublk server) switches to\n 352:\t\t ioctl command encoding.\n 353:\t\n 354:\tsource \"drivers/block/rnbd/Kconfig\"\n 355:\t\n 356:\tconfig BLK_DEV_ZONED_LOOP\n 357:\t\ttristate \"Zoned loopback device support\"\n 358:\t\tdepends on BLK_DEV_ZONED\n 359:\t\thelp\n 360:\t\t Saying Y here will allow you to use create a zoned block device using\n 361:\t\t regular files for zones (one file per zones). This is useful to test\n 362:\t\t file systems, device mapper and applications that support zoned block\n 363:\t\t devices. To create a zoned loop device, no user utility is needed, a\n 364:\t\t zoned loop device can be created (or re-started) using a command\n 365:\t\t like:\n 366:\t\n 367:\t\t echo \"add id=0,zone_size_mb=256,capacity_mb=16384,conv_zones=11\" \u003e \\\n 368:\t\t\t/dev/zloop-control\n 369:\t\n 370:\t\t See Documentation/admin-guide/blockdev/zoned_loop.rst for usage\n 371:\t\t details.\n 372:\t\n 373:\t\t If unsure, say N.\n 374:\t\n 375:\tendif # BLK_DEV\n"
} |
| 11/2 | 2026/10/06 13:27 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 33743 output: 30 thoughts: 21 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Investigating `sd_zbc.c` and its Build Configuration** Right, before diving too deep into this, I need to make sure I understand the build process and dependencies for this `sd_zbc.c` module. My immediate thought is, "Where does this fit into the bigger picture?" So, I need to check the `drivers/scsi/sd_zbc.c` directory's `Makefile` first. That will show me how it's compiled, any flags or linked libraries, and its relationship to other SCSI components. Understanding the build process is crucial for troubleshooting and for correctly contributing changes. After that, the `Kconfig` file in the same directory, is where I'll find its configuration options, which are just as essential. I need to see how the ZBC feature is enabled or disabled, any dependencies it might have on other kernel modules or features, and if there are any specific configuration parameters that I need to be aware of. This is the place to check which default values are set. This includes knowing whether it's enabled by default, or if it needs to be explicitly compiled. And if it's modular, how it gets loaded. Those two files, working in tandem, will give me the proper contextual understanding of this module. |
| 12/2 | 2026/10/06 13:27 | tool | grepper |
0mArgs:{
"Expression": "sd_zbc",
"PathPrefixes": [
"drivers/scsi/"
]
}Results: {
"Output": "\nFull output is too long, showing 200 out of 271 lines.\nUse more precise expression or PathPrefixes if possible.\n\nNumber of matching lines per file (4 files in total):\ndrivers/scsi/Makefile:2\ndrivers/scsi/sd.c:10\ndrivers/scsi/sd.h:11\ndrivers/scsi/sd_zbc.c:49\n\ndrivers/scsi/Makefile=36=CONTEXT_ANALYSIS_sd_dif.o := y\ndrivers/scsi/Makefile:37:CONTEXT_ANALYSIS_sd_zbc.o := y\ndrivers/scsi/Makefile-38-CONTEXT_ANALYSIS_sr.o := y\n--\ndrivers/scsi/Makefile=199=sd_mod-$(CONFIG_BLK_DEV_INTEGRITY) += sd_dif.o\ndrivers/scsi/Makefile:200:sd_mod-$(CONFIG_BLK_DEV_ZONED) += sd_zbc.o\ndrivers/scsi/Makefile-201-\n--\ndrivers/scsi/sd.c=1537=static blk_status_t sd_init_command(struct scsi_cmnd *cmd)\n--\ndrivers/scsi/sd.c-1562-\tcase REQ_OP_ZONE_RESET:\ndrivers/scsi/sd.c:1563:\t\treturn sd_zbc_setup_zone_mgmt_cmnd(cmd, ZO_RESET_WRITE_POINTER,\ndrivers/scsi/sd.c-1564-\t\t\t\t\t\t false);\ndrivers/scsi/sd.c-1565-\tcase REQ_OP_ZONE_RESET_ALL:\ndrivers/scsi/sd.c:1566:\t\treturn sd_zbc_setup_zone_mgmt_cmnd(cmd, ZO_RESET_WRITE_POINTER,\ndrivers/scsi/sd.c-1567-\t\t\t\t\t\t true);\ndrivers/scsi/sd.c-1568-\tcase REQ_OP_ZONE_OPEN:\ndrivers/scsi/sd.c:1569:\t\treturn sd_zbc_setup_zone_mgmt_cmnd(cmd, ZO_OPEN_ZONE, false);\ndrivers/scsi/sd.c-1570-\tcase REQ_OP_ZONE_CLOSE:\ndrivers/scsi/sd.c:1571:\t\treturn sd_zbc_setup_zone_mgmt_cmnd(cmd, ZO_CLOSE_ZONE, false);\ndrivers/scsi/sd.c-1572-\tcase REQ_OP_ZONE_FINISH:\ndrivers/scsi/sd.c:1573:\t\treturn sd_zbc_setup_zone_mgmt_cmnd(cmd, ZO_FINISH_ZONE, false);\ndrivers/scsi/sd.c-1574-\tdefault:\n--\ndrivers/scsi/sd.c=2369=static int sd_done(struct scsi_cmnd *SCpnt)\n--\ndrivers/scsi/sd.c-2474-\tif (sdkp-\u003edevice-\u003etype == TYPE_ZBC)\ndrivers/scsi/sd.c:2475:\t\tgood_bytes = sd_zbc_complete(SCpnt, good_bytes, \u0026sshdr);\ndrivers/scsi/sd.c-2476-\n--\ndrivers/scsi/sd.c=3773=static void sd_revalidate_disk(struct gendisk *disk)\n--\ndrivers/scsi/sd.c-3830-\t\t\tsd_read_block_characteristics(sdkp, lim);\ndrivers/scsi/sd.c:3831:\t\t\tsd_zbc_read_zones(sdkp, lim, buffer);\ndrivers/scsi/sd.c-3832-\t\t}\n--\ndrivers/scsi/sd.c-3897-\t */\ndrivers/scsi/sd.c:3898:\tif (sd_zbc_revalidate_zones(sdkp))\ndrivers/scsi/sd.c-3899-\t\tset_capacity_and_notify(disk, 0);\n--\ndrivers/scsi/sd.c=3927=static const struct block_device_operations sd_fops = {\n--\ndrivers/scsi/sd.c-3935-\t.unlock_native_capacity\t= sd_unlock_native_capacity,\ndrivers/scsi/sd.c:3936:\t.report_zones\t\t= sd_zbc_report_zones,\ndrivers/scsi/sd.c-3937-\t.get_unique_id\t\t= sd_get_unique_id,\n--\ndrivers/scsi/sd.c-3940-#ifdef CONFIG_BLK_DEV_ZONED\ndrivers/scsi/sd.c:3941:\t.se_ops\t\t\t= \u0026sd_zbc_se_ops,\ndrivers/scsi/sd.c-3942-#endif\n--\ndrivers/scsi/sd.h=232=void sd_dif_config_host(struct scsi_disk *sdkp, struct queue_limits *lim);\n--\ndrivers/scsi/sd.h-235-\ndrivers/scsi/sd.h:236:int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim,\ndrivers/scsi/sd.h-237-\t\tu8 buf[SD_BUF_SIZE]);\ndrivers/scsi/sd.h:238:int sd_zbc_revalidate_zones(struct scsi_disk *sdkp);\ndrivers/scsi/sd.h:239:blk_status_t sd_zbc_setup_zone_mgmt_cmnd(struct scsi_cmnd *cmd,\ndrivers/scsi/sd.h-240-\t\t\t\t\t unsigned char op, bool all);\ndrivers/scsi/sd.h:241:unsigned int sd_zbc_complete(struct scsi_cmnd *cmd, unsigned int good_bytes,\ndrivers/scsi/sd.h-242-\t\t\t struct scsi_sense_hdr *sshdr);\ndrivers/scsi/sd.h:243:int sd_zbc_report_zones(struct gendisk *disk, sector_t sector,\ndrivers/scsi/sd.h-244-\t\tunsigned int nr_zones, struct blk_report_zones_args *args);\ndrivers/scsi/sd.h-245-\ndrivers/scsi/sd.h:246:extern const struct blk_storage_elements_ops sd_zbc_se_ops;\ndrivers/scsi/sd.h-247-\n--\ndrivers/scsi/sd.h-249-\ndrivers/scsi/sd.h:250:static inline int sd_zbc_read_zones(struct scsi_disk *sdkp,\ndrivers/scsi/sd.h-251-\t\tstruct queue_limits *lim, u8 buf[SD_BUF_SIZE])\n--\ndrivers/scsi/sd.h-255-\ndrivers/scsi/sd.h:256:static inline int sd_zbc_revalidate_zones(struct scsi_disk *sdkp)\ndrivers/scsi/sd.h-257-{\n--\ndrivers/scsi/sd.h-260-\ndrivers/scsi/sd.h:261:static inline blk_status_t sd_zbc_setup_zone_mgmt_cmnd(struct scsi_cmnd *cmd,\ndrivers/scsi/sd.h-262-\t\t\t\t\t\t unsigned char op,\n--\ndrivers/scsi/sd.h-267-\ndrivers/scsi/sd.h:268:static inline unsigned int sd_zbc_complete(struct scsi_cmnd *cmd,\ndrivers/scsi/sd.h-269-\t\t\tunsigned int good_bytes, struct scsi_sense_hdr *sshdr)\n--\ndrivers/scsi/sd.h-273-\ndrivers/scsi/sd.h:274:#define sd_zbc_report_zones NULL\ndrivers/scsi/sd.h-275-\n--\ndrivers/scsi/sd_zbc.c-26-/* Whether or not a SCSI zone descriptor describes a gap zone. */\ndrivers/scsi/sd_zbc.c:27:static bool sd_zbc_is_gap_zone(const u8 buf[64])\ndrivers/scsi/sd_zbc.c-28-{\n--\ndrivers/scsi/sd_zbc.c-32-/**\ndrivers/scsi/sd_zbc.c:33: * sd_zbc_parse_report - Parse a SCSI zone descriptor\ndrivers/scsi/sd_zbc.c-34- * @sdkp: SCSI disk pointer.\n--\ndrivers/scsi/sd_zbc.c-36- * @idx: Index of the zone relative to the first zone reported by the current\ndrivers/scsi/sd_zbc.c:37: *\tsd_zbc_report_zones() call.\ndrivers/scsi/sd_zbc.c-38- * @args: report zones arguments (callback, etc)\n--\ndrivers/scsi/sd_zbc.c-44- */\ndrivers/scsi/sd_zbc.c:45:static int sd_zbc_parse_report(struct scsi_disk *sdkp, const u8 buf[64],\ndrivers/scsi/sd_zbc.c-46-\t\t\tunsigned int idx, struct blk_report_zones_args *args)\n--\ndrivers/scsi/sd_zbc.c-51-\ndrivers/scsi/sd_zbc.c:52:\tif (WARN_ON_ONCE(sd_zbc_is_gap_zone(buf)))\ndrivers/scsi/sd_zbc.c-53-\t\treturn -EINVAL;\n--\ndrivers/scsi/sd_zbc.c-91-/**\ndrivers/scsi/sd_zbc.c:92: * sd_zbc_do_report_zones - Issue a REPORT ZONES scsi command.\ndrivers/scsi/sd_zbc.c-93- * @sdkp: The target disk\n--\ndrivers/scsi/sd_zbc.c-104- */\ndrivers/scsi/sd_zbc.c:105:static int sd_zbc_do_report_zones(struct scsi_disk *sdkp, unsigned char *buf,\ndrivers/scsi/sd_zbc.c-106-\t\t\t\t unsigned int buflen, sector_t lba,\n--\ndrivers/scsi/sd_zbc.c-149-/**\ndrivers/scsi/sd_zbc.c:150: * sd_zbc_alloc_report_buffer() - Allocate a buffer for report zones reply.\ndrivers/scsi/sd_zbc.c-151- * @sdkp: The target disk\n--\ndrivers/scsi/sd_zbc.c-161- */\ndrivers/scsi/sd_zbc.c:162:static void *sd_zbc_alloc_report_buffer(struct scsi_disk *sdkp,\ndrivers/scsi/sd_zbc.c-163-\t\t\t\t\tunsigned int nr_zones, size_t *buflen)\n--\ndrivers/scsi/sd_zbc.c-200-/**\ndrivers/scsi/sd_zbc.c:201: * sd_zbc_zone_sectors - Get the device zone size in number of 512B sectors.\ndrivers/scsi/sd_zbc.c-202- * @sdkp: The target disk\ndrivers/scsi/sd_zbc.c-203- */\ndrivers/scsi/sd_zbc.c:204:static inline sector_t sd_zbc_zone_sectors(struct scsi_disk *sdkp)\ndrivers/scsi/sd_zbc.c-205-{\n--\ndrivers/scsi/sd_zbc.c-209-/**\ndrivers/scsi/sd_zbc.c:210: * sd_zbc_report_zones - SCSI .report_zones() callback.\ndrivers/scsi/sd_zbc.c-211- * @disk: Disk to report zones for.\n--\ndrivers/scsi/sd_zbc.c-218- */\ndrivers/scsi/sd_zbc.c:219:int sd_zbc_report_zones(struct gendisk *disk, sector_t sector,\ndrivers/scsi/sd_zbc.c-220-\t\t\tunsigned int nr_zones,\n--\ndrivers/scsi/sd_zbc.c-239-\ndrivers/scsi/sd_zbc.c:240:\tbuf = sd_zbc_alloc_report_buffer(sdkp, nr_zones, \u0026buflen);\ndrivers/scsi/sd_zbc.c-241-\tif (!buf)\n--\ndrivers/scsi/sd_zbc.c-244-\twhile (zone_idx \u003c nr_zones \u0026\u0026 lba \u003c sdkp-\u003ecapacity) {\ndrivers/scsi/sd_zbc.c:245:\t\tret = sd_zbc_do_report_zones(sdkp, buf, buflen, lba, true);\ndrivers/scsi/sd_zbc.c-246-\t\tif (ret)\n--\ndrivers/scsi/sd_zbc.c-269-\t\t\tlba = start_lba + zone_length;\ndrivers/scsi/sd_zbc.c:270:\t\t\tif (sd_zbc_is_gap_zone(\u0026buf[offset])) {\ndrivers/scsi/sd_zbc.c-271-\t\t\t\tif (sdkp-\u003ezone_starting_lba_gran)\n--\ndrivers/scsi/sd_zbc.c-278-\ndrivers/scsi/sd_zbc.c:279:\t\t\tret = sd_zbc_parse_report(sdkp, buf + offset, zone_idx,\ndrivers/scsi/sd_zbc.c-280-\t\t\t\t\t\t args);\n--\ndrivers/scsi/sd_zbc.c-293-\ndrivers/scsi/sd_zbc.c:294:static blk_status_t sd_zbc_cmnd_checks(struct scsi_cmnd *cmd)\ndrivers/scsi/sd_zbc.c-295-{\n--\ndrivers/scsi/sd_zbc.c-306-\ndrivers/scsi/sd_zbc.c:307:\tif (sector \u0026 (sd_zbc_zone_sectors(sdkp) - 1))\ndrivers/scsi/sd_zbc.c-308-\t\t/* Unaligned request */\n--\ndrivers/scsi/sd_zbc.c-314-/**\ndrivers/scsi/sd_zbc.c:315: * sd_zbc_setup_zone_mgmt_cmnd - Prepare a zone ZBC_OUT command. The operations\ndrivers/scsi/sd_zbc.c-316- *\t\t\tcan be RESET WRITE POINTER, OPEN, CLOSE or FINISH.\n--\ndrivers/scsi/sd_zbc.c-323- */\ndrivers/scsi/sd_zbc.c:324:blk_status_t sd_zbc_setup_zone_mgmt_cmnd(struct scsi_cmnd *cmd,\ndrivers/scsi/sd_zbc.c-325-\t\t\t\t\t unsigned char op, bool all)\n--\ndrivers/scsi/sd_zbc.c-332-\ndrivers/scsi/sd_zbc.c:333:\tret = sd_zbc_cmnd_checks(cmd);\ndrivers/scsi/sd_zbc.c-334-\tif (ret != BLK_STS_OK)\n--\ndrivers/scsi/sd_zbc.c-354-/**\ndrivers/scsi/sd_zbc.c:355: * sd_zbc_complete - ZBC command post processing.\ndrivers/scsi/sd_zbc.c-356- * @cmd: Completed command\n--\ndrivers/scsi/sd_zbc.c-362- */\ndrivers/scsi/sd_zbc.c:363:unsigned int sd_zbc_complete(struct scsi_cmnd *cmd, unsigned int good_bytes,\ndrivers/scsi/sd_zbc.c-364-\t\t struct scsi_sense_hdr *sshdr)\n--\ndrivers/scsi/sd_zbc.c-384-/**\ndrivers/scsi/sd_zbc.c:385: * sd_zbc_check_zoned_characteristics - Check zoned block device characteristics\ndrivers/scsi/sd_zbc.c-386- * @sdkp: Target disk\n--\ndrivers/scsi/sd_zbc.c-390- */\ndrivers/scsi/sd_zbc.c:391:static int sd_zbc_check_zoned_characteristics(struct scsi_disk *sdkp,\ndrivers/scsi/sd_zbc.c-392-\t\t\t\t\t unsigned char *buf)\n--\ndrivers/scsi/sd_zbc.c-455-/**\ndrivers/scsi/sd_zbc.c:456: * sd_zbc_check_capacity - Check the device capacity\ndrivers/scsi/sd_zbc.c-457- * @sdkp: Target disk\n--\ndrivers/scsi/sd_zbc.c-466- */\ndrivers/scsi/sd_zbc.c:467:static int sd_zbc_check_capacity(struct scsi_disk *sdkp, unsigned char *buf,\ndrivers/scsi/sd_zbc.c-468-\t\t\t\t u32 *zblocks)\n--\ndrivers/scsi/sd_zbc.c-475-\t/* Do a report zone to get max_lba and the size of the first zone */\ndrivers/scsi/sd_zbc.c:476:\tret = sd_zbc_do_report_zones(sdkp, buf, SD_BUF_SIZE, 0, false);\ndrivers/scsi/sd_zbc.c-477-\tif (ret)\n\n"
} |
| 13/2 | 2026/10/06 13:27 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 38563 output: 36 thoughts: 333 Thoughts: **Zoned Block Device Implementation Details** Okay, so I'm looking at the changes related to zoned block device support in the kernel. First, it looks like `sd_mod` is being compiled when `CONFIG_BLK_DEV_ZONED` is enabled, and also `CONFIG_BLK_DEV_SD` compiles `sd_mod`. That's expected. Similarly, `CONFIG_BLK_DEV_ZONED_LOOP` brings in `zloop.o`. I should double check how these modules interact. The core zoned block device functionality is compiled into `block/blk-zoned.c` when `CONFIG_BLK_DEV_ZONED` is set. Now, let's see what else has been touched. I need to get a complete picture. Okay, I see a good list of files that have been modified: `Documentation/admin-guide/blockdev/zoned_loop.rst`, `block/blk-core.c`, `block/blk-zoned.c`, `block/blk.h`, `block/ioctl.c`, `drivers/block/zloop.c`, `drivers/scsi/sd.c`, `drivers/scsi/sd.h`, `drivers/scsi/sd_zbc.c`, `include/linux/blkdev.h`, `include/uapi/linux/blkzoned.h`, and `include/uapi/linux/fs.h`. That makes sense; documentation, core block layer, the loop device, SCSI driver updates, and header files across the user and kernel space are all involved. I'm particularly interested in the `block/blk-zoned.c` file. I need to see what sort of conditional compilation is going on in there. Let's look at the if-def guards. Aha, I see something suspicious around line 2700: a lone `#endif`. I need to figure out what that `#endif` is closing. Better go read around that line and see what it is related to. |
| 14/2 | 2026/10/06 13:28 | tool | read-file |
0mArgs:{
"File": "block/blk-zoned.c",
"FirstLine": 2695,
"LineCount": 40
}Results: {
"Contents": "2695:\t\tunsigned int zwp_zone_no, zwp_ref;\n2696:\t\tunsigned int zwp_bio_list_size;\n2697:\t\tenum blk_zone_cond zwp_cond;\n2698:\t\tunsigned long flags;\n2699:\t\n2700:\t\tspin_lock_irqsave(\u0026zwplug-\u003elock, flags);\n2701:\t\tzwp_zone_no = zwplug-\u003ezone_no;\n2702:\t\tzwp_flags = zwplug-\u003eflags;\n2703:\t\tzwp_ref = refcount_read(\u0026zwplug-\u003eref);\n2704:\t\tzwp_cond = zwplug-\u003econd;\n2705:\t\tzwp_wp_offset = zwplug-\u003ewp_offset;\n2706:\t\tzwp_bio_list_size = bio_list_size(\u0026zwplug-\u003ebio_list);\n2707:\t\tspin_unlock_irqrestore(\u0026zwplug-\u003elock, flags);\n2708:\t\n2709:\t\tseq_printf(m,\n2710:\t\t\t\"Zone no: %u, flags: 0x%x, ref: %u, cond: %s, wp ofst: %u, pending BIO: %u\\n\",\n2711:\t\t\tzwp_zone_no, zwp_flags, zwp_ref, blk_zone_cond_str(zwp_cond),\n2712:\t\t\tzwp_wp_offset, zwp_bio_list_size);\n2713:\t}\n2714:\t\n2715:\tint queue_zone_wplugs_show(void *data, struct seq_file *m)\n2716:\t{\n2717:\t\tstruct request_queue *q = data;\n2718:\t\n2719:\t\tdisk_for_all_zone_wplugs(q-\u003edisk, queue_zone_wplug_show, m);\n2720:\t\n2721:\t\treturn 0;\n2722:\t}\n2723:\t#endif\n2724:\t\n2725:\tstatic int disk_wait_for_se_mgmt_completion(struct gendisk *disk)\n2726:\t{\n2727:\t\tstruct blk_storage_element *elements, *e;\n2728:\t\tunsigned int i, nr_se, nr_elements = 0;\n2729:\t\tint ret;\n2730:\t\n2731:\t\tret = disk-\u003efops-\u003ese_ops-\u003ereport_elements(disk, NULL, \u0026nr_elements);\n2732:\t\tif (ret) {\n2733:\t\t\tpr_err(\"Failed to get number of storage elements\\n\");\n2734:\t\t\treturn ret;\n"
} |
| 15/2 | 2026/10/06 13:28 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 40070 output: 36 thoughts: 27 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Investigating a Preprocessor Directive** Alright, the `#endif` at line 2723...that's piquing my interest. We're dealing with a preprocessor directive, so whatever code it's closing out is conditionally compiled. My first instinct, like always, is to understand the scope and the condition. We've got to trace back to the corresponding `#if`, `#ifdef`, or `#ifndef`. Let's assume we are using a typical pre-processor. I immediately want to scan above line 2695. Going directly to line 2723 is premature without proper context, so let's start there. I'll be looking for the opening directive that this `#endif` is paired with, specifically the one that starts its conditional block of code. From there I'll be able to work out why that block was conditionally compiled. This should quickly tell me why this section of code is being protected and under what circumstances. |
| 16/2 | 2026/10/06 13:28 | tool | read-file |
0mArgs:{
"File": "block/blk-zoned.c",
"FirstLine": 2680,
"LineCount": 20
}Results: {
"Contents": "2680:\t\n2681:\t\t/*\n2682:\t\t * Retry without BLKDEV_ZERO_NOFALLBACK to force the fallback to a\n2683:\t\t * regular write with zero-pages.\n2684:\t\t */\n2685:\t\treturn blkdev_issue_zeroout(bdev, sector, nr_sects, gfp_mask, 0);\n2686:\t}\n2687:\tEXPORT_SYMBOL_GPL(blk_zone_issue_zeroout);\n2688:\t\n2689:\t#ifdef CONFIG_BLK_DEBUG_FS\n2690:\tstatic void queue_zone_wplug_show(struct blk_zone_wplug *zwplug,\n2691:\t\t\t\t\t void *data)\n2692:\t{\n2693:\t\tstruct seq_file *m = data;\n2694:\t\tunsigned int zwp_wp_offset, zwp_flags;\n2695:\t\tunsigned int zwp_zone_no, zwp_ref;\n2696:\t\tunsigned int zwp_bio_list_size;\n2697:\t\tenum blk_zone_cond zwp_cond;\n2698:\t\tunsigned long flags;\n2699:\t\n"
} |
| 17/2 | 2026/10/06 13:28 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 40680 output: 27 thoughts: 82 Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience: **Analyzing a Conditional Compilation Block and Code Location** Alright, so I'm looking at this code and, specifically, I see that the `#endif` on line 2723 is the closing bracket for the `#ifdef CONFIG_BLK_DEBUG_FS` conditional. So, anything *inside* that block is compiled only if `CONFIG_BLK_DEBUG_FS` is defined. That's straightforward enough. Now, the interesting bit is what *follows*. `disk_wait_for_se_mgmt_completion` and everything after that – the rest of these functions – are *outside* of that conditional block. This is critical because it means regardless of whether debugging features are enabled, these functions will *always* be included in the build. I need to keep that in mind when I'm assessing the core functionality vs. debug-specific features. Since I'm in `block/blk-zoned.c`, it’s probably worthwhile to quickly cross-reference this with the build process. I'll take a quick peek at `block/Makefile` to confirm the expected compilation dependencies and how the overall module is constructed. I want to see if this is as I would expect, based on what I have already worked through. |
| 18/2 | 2026/10/06 13:28 | tool | grepper |
0mArgs:{
"Expression": "blk-zoned",
"PathPrefixes": [
"block/Makefile"
]
}Results: {
"Output": "block/Makefile=31=obj-$(CONFIG_BLK_DEV_INTEGRITY) += bio-integrity.o blk-integrity.o t10-pi.o \\\nblock/Makefile-32-\t\t\t\t bio-integrity-auto.o bio-integrity-fs.o\nblock/Makefile:33:obj-$(CONFIG_BLK_DEV_ZONED)\t+= blk-zoned.o\nblock/Makefile-34-obj-$(CONFIG_BLK_WBT)\t\t+= blk-wbt.o\n"
} |
| 19/2 | 2026/10/06 13:28 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 41200 output: 34 thoughts: 2064 Thoughts: **Zoned Block Device Implementation Analysis and Fuzzing Strategy**
Okay, so I've been digging into this new zoned block device functionality in the kernel, specifically around storage element management. First, it's pretty clear that `blk-zoned.c` is the core, and it's compiled when `CONFIG_BLK_DEV_ZONED` is enabled. That's my entry point. I see the ioctl handlers defined in `block/blk.h` and the dispatch in `block/ioctl.c`. These ioctls are accessible through standard block device descriptors, which is good. The `zloop` device (`/dev/zloop-control`) looks particularly interesting, as it allows userspace to control emulated storage elements when `CONFIG_BLK_DEV_ZONED_LOOP` is enabled (which depends on `BLK_DEV_ZONED`), so I'll need to remember that.
I see how a zoned block device `/dev/zloop0` is registered with `se_ops`, which the user space accesses through the ioctls, specifically: `BLKGETNRSTORELEMS`, `BLKREPORTSTORELEMS`, `BLKREMOVESTORELEM`, and `BLKRESTORESTORELEMS`. Also, the `/dev/zloop-control` supports writing `degrade_element id=...,element_id=...`. The standard reads/writes to `/dev/zloop0` also go through checks for offline zones.
The `blkdev_remove_storage_element_ioctl` function calls `blkdev_issue_flush`, `truncate_bdev_range`, and ultimately `bdev_remove_storage_element`. Inside `bdev_remove_storage_element`, I see some pretty significant operations: `blk_mq_freeze_queue`, `disk->fops->se_ops->remove_element` which calls the removal function, followed by a wait for completion (`disk_wait_for_se_mgmt_completion`) and zone revalidation. The `disk_wait_for_se_mgmt_completion` is particularly concerning because it contains a hardcoded `msleep(500)` loop.
Now, `zloop_remove_element` uses `schedule_delayed_work`, scheduling work to happen after 2 seconds. The `disk_wait_for_se_mgmt_completion` polls `report_elements`, waiting for the `BLK_SE_STS_REMOVE_IN_PROGRESS` status to clear. On the SCSI side, `sd_zbc.c` provides similar SCSI commands for the same functions.
This is a **lot** of new, complex functionality, and many things can go wrong here! There's a lot of potential for subtle bugs, race conditions, memory issues, and deadlocks. We're talking about new block device ioctls, a new zloop control interface, delayed work, queue freezing, zone revalidation, inode/filemap locks while freezing and/or calling `se_ops` – it's a perfect storm. Even without considering potential bugs, this code is *brand new user-facing functionality*.
The key question is, is it worth fuzzing? Absolutely! The criteria are clearly met:
- `WorthFuzzing`: True. Given the complexity and the potential for device corruption or data loss, this is a clear "yes".
- `Reachable`: Yes. Standard ioctl calls on block devices and writes to `/dev/zloop-control` are how users interact.
- `FocusSymbols`: OK, I need to narrow down which functions are most relevant. I need to avoid ubiquitous functions like `submit_bio_noacct` or `blkdev_common_ioctl`. The functions added or modified here are important, e.g., anything in `blk-zoned.c`, `zloop.c`, and `sd_zbc.c`. The ones I'll focus on are:
- `blkdev_zone_storage_elements_ioctl`
- `bdev_remove_storage_element`
- `bdev_restore_storage_elements`
- `bdev_report_storage_elements`
- `disk_wait_for_se_mgmt_completion`
- `bdev_zone_is_offline`
- `zloop_remove_element`
- `zloop_restore_elements`
- `zloop_report_elements`
- `zloop_ctl_degrade_element`
- `sd_zbc_report_storage_elements`
- `sd_zbc_remove_storage_element`
- `sd
|
| 20/2 | 2026/10/06 13:28 | tool | grepper |
0mArgs:{
"Expression": "config BLK_DEV_ZONED",
"PathPrefixes": [
"block/Kconfig"
]
}Results: {
"Output": "block/Kconfig=77=config BLK_DEV_WRITE_MOUNTED\n--\nblock/Kconfig-96-\nblock/Kconfig:97:config BLK_DEV_ZONED\nblock/Kconfig-98-\tbool \"Zoned block device support\"\n"
} |
| 21/2 | 2026/10/06 13:28 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44402 output: 31 thoughts: 16 Thoughts: Okay, let's dive into this. Here's my take on those instructions, as if they're my internal thought process: **Diving into the Kernel Configuration Landscape** Right, so we're talking about kernel configuration, and specifically the `block/Kconfig` file. I'm assuming we're dealing with block device configuration, likely within the Linux kernel itself, given the context. My immediate focus is around line 97. That line likely holds a key piece of information, a pivotal configuration option, or perhaps the start of a new section or definition related to block device configuration. I need to understand what's *there* specifically. Is it a default setting, a conditional definition, a help text description, or something else entirely? Whatever it is, that's where I need to orient myself to understand the current configuration's scope. Line 97, here we go! |
| 22/2 | 2026/10/06 13:29 | tool | read-file |
0mArgs:{
"File": "block/Kconfig",
"FirstLine": 95,
"LineCount": 20
}Results: {
"Contents": " 95:\t\twith the bdev_allow_write_mounted boot option.\n 96:\t\n 97:\tconfig BLK_DEV_ZONED\n 98:\t\tbool \"Zoned block device support\"\n 99:\t\thelp\n 100:\t\tBlock layer zoned block device support. This option enables\n 101:\t\tsupport for ZAC/ZBC/ZNS host-managed and host-aware zoned block\n 102:\t\tdevices.\n 103:\t\n 104:\t\tSay yes here if you have a ZAC, ZBC, or ZNS storage device.\n 105:\t\n 106:\tconfig BLK_DEV_THROTTLING\n 107:\t\tbool \"Block layer bio throttling support\"\n 108:\t\tdepends on BLK_CGROUP\n 109:\t\tselect BLK_CGROUP_RWSTAT\n 110:\t\thelp\n 111:\t\tBlock layer bio throttling support. It can be used to limit\n 112:\t\tthe IO rate to a device. IO rate policies are per cgroup and\n 113:\t\tone needs to mount and use blkio cgroup controller for creating\n 114:\t\tcgroups and specifying per device IO rate policies.\n"
} |
| 23/2 | 2026/10/06 13:29 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44956 output: 338 thoughts: 629 Thoughts: **Analyzing Zoned Block Device Configuration and Symbols** Okay, so I'm looking at these zoned block device configurations and figuring out what needs to be enabled and which symbols are key. `BLK_DEV_ZONED` is completely independent, that's straightforward. Now, `BLK_DEV_ZONED_LOOP` relies on `BLK_DEV_ZONED`, so that's a dependency. And then, there's `BLK_DEV_SD`, which is just standard SCSI disk support, pretty standard stuff. Alright, let's see if there are any other dependencies. Quick check, `CONFIG_SCSI` seems to tie into `CONFIG_BLK_DEV_SD`. So, in total, we're looking at `BLK_DEV_ZONED`, `BLK_DEV_ZONED_LOOP`, and `BLK_DEV_SD`. Seems like a reasonable set. Now, the important part: the symbols. Need to make sure they're accessible. Let's dig into these FocusSymbols and check if they are statically declared, which may reduce their availability. In `block/blk-zoned.c`, I've got `bdev_zone_is_offline`, `bdev_report_storage_elements`, `bdev_remove_storage_element`, `bdev_restore_storage_elements`, and `blkdev_zone_storage_elements_ioctl` – all non-static. Then there's `disk_wait_for_se_mgmt_completion`, which is static but still callable. Looking at `drivers/block/zloop.c`, I have `zloop_ctl_degrade_element` which is static, and `zloop_report_elements`, `zloop_remove_element`, and `zloop_restore_elements` which are all static, and pointed to by `zloop_se_ops`. And finally in `drivers/scsi/sd_zbc.c`, there's `sd_zbc_report_storage_elements`, `sd_zbc_remove_storage_element`, and `sd_zbc_restore_storage_elements` – all static, but used through `sd_zbc_se_ops`. Good. These are all valid C functions and symbols, they are not static inline functions in headers, so they are available in object files or symbol tables, if not optimised away, and their entrypoints are accessible (exported or non-static). Alright, let's get the `set-results` call ready. Here's how I'm thinking of structuring the call: - `WorthFuzzing`: True - `Reasoning`: "Identified zoned block device configurations and essential symbols to enable. Verified that the symbols are valid and accessible." - `FocusSymbols`: \[ "bdev\_zone\_is\_offline", "bdev\_report\_storage\_elements", "bdev\_remove\_storage\_element", "bdev\_restore\_storage\_elements", "blkdev\_zone\_storage\_elements\_ioctl", "disk\_wait\_for\_se\_mgmt\_completion", "zloop\_ctl\_degrade\_element", "zloop\_report\_elements", "zloop\_remove\_element", "zloop\_restore\_elements", "sd\_zbc\_report\_storage\_elements", "sd\_zbc\_remove\_storage\_element", "sd\_zbc\_restore\_storage\_elements" ] - `EnableConfigs`: \[ "BLK\_DEV\_ZONED", "BLK\_DEV\_ZONED\_LOOP", "BLK\_DEV\_SD" ] Ready to go. |
| 24/2 | 2026/10/06 13:29 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"BLK_DEV_ZONED",
"BLK_DEV_ZONED_LOOP",
"BLK_DEV_SD"
],
"FocusSymbols": [
"blkdev_zone_storage_elements_ioctl",
"bdev_report_storage_elements",
"bdev_remove_storage_element",
"bdev_restore_storage_elements",
"bdev_zone_is_offline",
"zloop_remove_element",
"zloop_restore_elements",
"zloop_report_elements",
"zloop_ctl_degrade_element",
"sd_zbc_report_storage_elements",
"sd_zbc_remove_storage_element",
"sd_zbc_restore_storage_elements"
],
"Reasoning": "The patch introduces new storage element management functionality for zoned block devices. It adds four new user-accessible block device ioctls (BLKGETNRSTORELEMS, BLKREPORTSTORELEMS, BLKREMOVESTORELEM, BLKRESTORESTORELEMS) in the core block layer, adds support in the scsi disk driver (sd_zbc), implements storage element emulation and a new control operation in the zoned loop driver (zloop), and introduces offline zone checks during bio submission. This reachable code path involves complex state transitions, synchronization (queue freeze/unfreeze, truncate, cache flush, delayed work), and memory management, making it an excellent target for syzkaller fuzzing.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"BLK_DEV_ZONED",
"BLK_DEV_ZONED_LOOP",
"BLK_DEV_SD"
],
"FocusSymbols": [
"blkdev_zone_storage_elements_ioctl",
"bdev_report_storage_elements",
"bdev_remove_storage_element",
"bdev_restore_storage_elements",
"bdev_zone_is_offline",
"zloop_remove_element",
"zloop_restore_elements",
"zloop_report_elements",
"zloop_ctl_degrade_element",
"sd_zbc_report_storage_elements",
"sd_zbc_remove_storage_element",
"sd_zbc_restore_storage_elements"
],
"Reasoning": "The patch introduces new storage element management functionality for zoned block devices. It adds four new user-accessible block device ioctls (BLKGETNRSTORELEMS, BLKREPORTSTORELEMS, BLKREMOVESTORELEM, BLKRESTORESTORELEMS) in the core block layer, adds support in the scsi disk driver (sd_zbc), implements storage element emulation and a new control operation in the zoned loop driver (zloop), and introduces offline zone checks during bio submission. This reachable code path involves complex state transitions, synchronization (queue freeze/unfreeze, truncate, cache flush, delayed work), and memory management, making it an excellent target for syzkaller fuzzing.",
"WorthFuzzing": true
} |
| 25/1 | 2026/10/06 13:29 | action | If |
14mArgs:{
"WorthFuzzing": true
}Results: null |
| 26/2 | 2026/10/06 13:29 | agent | kmsan-evaluator |
14mModel:core Args: null Results: {
"KMSANReasoning": "The patch series introduces storage element management for zoned block devices across the block core, the SCSI ZBC driver (sd_zbc), and the zoned loop driver (zloop), adding ioctls BLKGETNRSTORELEMS, BLKREPORTSTORELEMS, BLKREMOVESTORELEM, and BLKRESTORESTORELEMS.\n\nRegarding KMSAN vs KASAN applicability:\n1. Heap allocations: All dynamically allocated structures, including the storage element arrays (e.g. `zlo-\u003eelements`, `elements` in `disk_wait_for_se_mgmt_completion` and `blkdev_report_storage_elements_ioctl`) and SCSI buffers (`buf` in `sd_zbc_report_storage_elements`), are allocated using `kzalloc_objs()` or `kzalloc()`, ensuring all bytes are zero-initialized upon allocation.\n2. Kernel-to-user data transfers:\n - In `blkdev_get_nr_storage_elements_ioctl()`, `nr_elements` is a scalar unsigned int initialized to 0 before being passed to `put_user()`.\n - In `blkdev_report_storage_elements_ioctl()`, `struct blk_storage_elements_report` consists of two `__u32` fields (8 bytes total, aligned to 8 bytes, zero padding) and is completely initialized from userspace via `copy_from_user()` before the count is updated and written back.\n - `struct blk_storage_element` has explicit members summing to exactly 24 bytes (4 + 4 + 8 + 1 + 1 + 1 + 5 bytes) with natural 8-byte alignment, leaving zero internal or tail padding bytes. Moreover, in `sd_zbc_parse_storage_element()`, each element descriptor is explicitly zeroed with `memset()` before parsing.\n3. Out-of-bounds risk: In `zloop_report_elements()`, `*nr_elements` is assigned `zlo-\u003enr_elements`, which in `blkdev_report_storage_elements_ioctl()` could cause `copy_to_user()` to copy more elements than the allocated `nr_elements` buffer if the caller requested fewer elements. However, this is a slab-out-of-bounds heap read, which is directly detectible by KASAN rather than an uninitialized memory defect.\n\nBecause all touched structures and buffers are zero-initialized and no uninitialized memory usage or info-leaks are introduced, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 1145222b274a2f5837f52a04a8d52b5f3f74f8a8
Author: syz-cluster <triage@syzkaller.com>
Date: Tue Oct 6 13:26:40 2026 +0000
syz-cluster: applied patch under review
diff --git a/Documentation/admin-guide/blockdev/zoned_loop.rst b/Documentation/admin-guide/blockdev/zoned_loop.rst
index 64277494fb36a..ec39f141af8d6 100644
--- a/Documentation/admin-guide/blockdev/zoned_loop.rst
+++ b/Documentation/admin-guide/blockdev/zoned_loop.rst
@@ -61,7 +61,7 @@ The options available for the add command can be listed by reading the
/dev/zloop-control device::
$ cat /dev/zloop-control
- add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache
+ add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache,stor_elements=%u
remove id=%d
In more details, the options that can be used with the "add" command are as
@@ -113,6 +113,12 @@ discard_write_cache Discard all data that was not explicitly persisted using a
each zone file to the size recorded during the last flush
operation. This simulates power fail events where
uncommitted data is lost.
+stor_elements Control storage element emulation. The default value is 0,
+ indicating no emulation. A value of 1 indicates that all
+ access storage elements (equivalent to read+write head of
+ a disk) are emulated. A value of 2 enables fractional
+ access storage element (equivalent to pairs of read and
+ write heads of a disk) emulation .
=================== =========================================================
3) Deleting a Zoned Device
diff --git a/block/blk-core.c b/block/blk-core.c
index 13dc70e8f55d9..d420c80d2d938 100644
--- a/block/blk-core.c
+++ b/block/blk-core.c
@@ -866,6 +866,9 @@ void submit_bio_noacct(struct bio *bio)
switch (bio_op(bio)) {
case REQ_OP_READ:
+ if (bdev_is_zoned(bdev) &&
+ bdev_zone_is_offline(bdev, bio->bi_iter.bi_sector))
+ goto end_io;
break;
case REQ_OP_WRITE:
if (bio->bi_opf & REQ_ATOMIC) {
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index 19268afb8752e..6ebb04e0a1593 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -18,6 +18,8 @@
#include <linux/mempool.h>
#include <linux/kthread.h>
#include <linux/freezer.h>
+#include <linux/delay.h>
+#include <linux/uaccess.h>
#include <trace/events/block.h>
@@ -308,6 +310,24 @@ bool bdev_zone_is_seq(struct block_device *bdev, sector_t sector)
}
EXPORT_SYMBOL_GPL(bdev_zone_is_seq);
+/**
+ * bdev_zone_is_offline - check if a sector belongs to an offline zone
+ * @bdev: block device to check
+ * @sector: sector number
+ *
+ * Check if @sector on @bdev is contained in an offline zone.
+ */
+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector)
+{
+ enum blk_zone_cond cond;
+
+ if (!bdev_is_zoned(bdev))
+ return false;
+
+ cond = disk_zone_get_cond(bdev->bd_disk, sector);
+ return cond == BLK_ZONE_COND_OFFLINE;
+}
+
/**
* bdev_zone_mgmt_allowed - check if management operations are allowed on a zone
* @bdev: block device to check
@@ -2700,5 +2720,326 @@ int queue_zone_wplugs_show(void *data, struct seq_file *m)
return 0;
}
-
#endif
+
+static int disk_wait_for_se_mgmt_completion(struct gendisk *disk)
+{
+ struct blk_storage_element *elements, *e;
+ unsigned int i, nr_se, nr_elements = 0;
+ int ret;
+
+ ret = disk->fops->se_ops->report_elements(disk, NULL, &nr_elements);
+ if (ret) {
+ pr_err("Failed to get number of storage elements\n");
+ return ret;
+ }
+
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return -ENOMEM;
+
+ while (1) {
+ /*
+ * Check if we have storage elements being removed or restored.
+ */
+ nr_se = nr_elements;
+ ret = disk->fops->se_ops->report_elements(disk, elements,
+ &nr_se);
+ if (ret) {
+ pr_err("Failed to get storage elements\n");
+ break;
+ }
+
+ e = elements;
+ for (i = 0; i < nr_se; i++, e++) {
+ if (e->status == BLK_SE_STS_REMOVE_IN_PROGRESS ||
+ e->status == BLK_SE_STS_RESTORE_IN_PROGRESS)
+ break;
+ }
+ if (i >= nr_se)
+ break;
+
+ /* Not done yet: wait and retry. */
+ msleep(500);
+ }
+
+ kfree(elements);
+
+ return ret;
+}
+
+/**
+ * bdev_report_storage_elements - report the storage elements of a block device
+ *
+ * Fill at most @nr_elements storage element descriptors in the array @elements.
+ * The number of storage elements filled in the array is returned using
+ * @nr_elements. If @elements is NULL, only @nr_elements is returned.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_report_storage_elements(struct block_device *bdev,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct gendisk *disk = bdev->bd_disk;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ if (!nr_elements)
+ return -EINVAL;
+
+ if (*nr_elements && !elements)
+ return -EINVAL;
+
+ return disk->fops->se_ops->report_elements(disk, elements, nr_elements);
+}
+EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
+
+static int blkdev_get_nr_storage_elements_ioctl(struct block_device *bdev,
+ void __user *argp)
+{
+ unsigned int nr_elements = 0;
+ int ret;
+
+ ret = bdev_report_storage_elements(bdev, NULL, &nr_elements);
+ if (ret)
+ return ret;
+
+ if (put_user(nr_elements, (unsigned int __user *)argp))
+ return -EFAULT;
+
+ return 0;
+}
+
+static int blkdev_report_storage_elements_ioctl(struct block_device *bdev,
+ void __user *argp)
+{
+ struct blk_storage_elements_report rep;
+ struct blk_storage_element *elements;
+ unsigned int nr_elements = 0;
+ unsigned long retc;
+ int ret;
+
+ if (!argp)
+ return -EINVAL;
+
+ if (copy_from_user(&rep, argp,
+ sizeof(struct blk_storage_elements_report)))
+ return -EFAULT;
+
+ ret = bdev_report_storage_elements(bdev, NULL, &nr_elements);
+ if (ret)
+ return ret;
+
+ nr_elements = min(rep.nr_elements, nr_elements);
+ if (!nr_elements)
+ return -EINVAL;
+
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return -ENOMEM;
+
+ ret = bdev_report_storage_elements(bdev, elements, &nr_elements);
+ if (ret)
+ goto free_elements;
+
+ retc = copy_to_user(argp + sizeof(struct blk_storage_elements_report),
+ elements,
+ sizeof(struct blk_storage_element) * nr_elements);
+ if (retc) {
+ ret = -EFAULT;
+ goto free_elements;
+ }
+
+ rep.nr_elements = nr_elements;
+ retc = copy_to_user(argp, &rep,
+ sizeof(struct blk_storage_elements_report));
+ if (retc)
+ ret = -EFAULT;
+
+free_elements:
+ kfree(elements);
+ return ret;
+}
+
+/**
+ * bdev_remove_storage_element - Remove (depopulate) a storage element of a
+ * block device
+ *
+ * Remove (depopulate) the storage element identified by @element_id from the
+ * block device @bdev. The caller is responsible for taking care of any
+ * necessary device write cache flush and invalidation of cached data for the
+ * zones that will be offlined.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /* Zero is not a valid storage element ID. */
+ if (!element_id)
+ return -EINVAL;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding command.
+ * The caller is responsible for not queuing up more I/Os by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ ret = disk->fops->se_ops->remove_element(disk, element_id);
+ if (ret)
+ return ret;
+
+ /* Revalidate the device zones once the opration completes. */
+ ret = disk_wait_for_se_mgmt_completion(disk);
+ if (ret)
+ return ret;
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
+
+static int blkdev_remove_storage_element_ioctl(struct block_device *bdev,
+ blk_mode_t mode, void __user *argp)
+{
+ unsigned int element_id;
+ int ret;
+
+ if (!(mode & BLK_OPEN_WRITE))
+ return -EBADF;
+ if (bdev_read_only(bdev))
+ return -EPERM;
+
+ if (get_user(element_id, (unsigned int __user *)argp))
+ return -EFAULT;
+
+ /*
+ * Flush the device volatile write cache and invalidate all cached data
+ * so that reads do not return old data for zones that went offline.
+ */
+ inode_lock(bdev->bd_mapping->host);
+ filemap_invalidate_lock(bdev->bd_mapping);
+
+ ret = blkdev_issue_flush(bdev->bd_disk->part0);
+ if (ret)
+ goto out_unlock;
+
+ ret = truncate_bdev_range(bdev, mode, 0,
+ (get_capacity(bdev->bd_disk) << SECTOR_SHIFT) - 1);
+ if (ret)
+ goto out_unlock;
+
+ ret = bdev_remove_storage_element(bdev, element_id);
+
+out_unlock:
+ filemap_invalidate_unlock(bdev->bd_mapping);
+ inode_unlock(bdev->bd_mapping->host);
+
+ return ret;
+}
+
+/**
+ * bdev_restore_storage_elements - Restore all depopulated storage elements of a
+ * block device
+ *
+ * Restore all storage elements of @bdev that have been depopulated. Not all
+ * elements may be restored by this operation.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_restore_storage_elements(struct block_device *bdev)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding commands.
+ * The caller is responsible for not queuing up more I/O by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ ret = disk->fops->se_ops->restore_elements(disk);
+ if (ret)
+ return ret;
+
+ ret = disk_wait_for_se_mgmt_completion(disk);
+ if (ret)
+ return ret;
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
+
+static int blkdev_restore_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode)
+{
+ int ret;
+
+ if (!(mode & BLK_OPEN_WRITE))
+ return -EBADF;
+ if (bdev_read_only(bdev))
+ return -EPERM;
+
+ /*
+ * Flush the device volatile write cache and invalidate all cached data
+ * so that reads do not return old data for zones that went offline.
+ */
+ inode_lock(bdev->bd_mapping->host);
+ filemap_invalidate_lock(bdev->bd_mapping);
+
+ ret = blkdev_issue_flush(bdev->bd_disk->part0);
+ if (ret)
+ goto out_unlock;
+
+ ret = truncate_bdev_range(bdev, mode, 0,
+ (get_capacity(bdev->bd_disk) << SECTOR_SHIFT) - 1);
+ if (ret)
+ goto out_unlock;
+
+ ret = bdev_restore_storage_elements(bdev);
+
+out_unlock:
+ filemap_invalidate_unlock(bdev->bd_mapping);
+ inode_unlock(bdev->bd_mapping->host);
+
+ return ret;
+}
+
+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp)
+{
+ if (!bdev_is_zoned(bdev) || !bdev->bd_disk->fops->se_ops)
+ return -ENOTTY;
+
+ switch (cmd) {
+ case BLKGETNRSTORELEMS:
+ return blkdev_get_nr_storage_elements_ioctl(bdev, argp);
+ case BLKREPORTSTORELEMS:
+ return blkdev_report_storage_elements_ioctl(bdev, argp);
+ case BLKREMOVESTORELEM:
+ return blkdev_remove_storage_element_ioctl(bdev, mode, argp);
+ case BLKRESTORESTORELEMS:
+ return blkdev_restore_storage_elements_ioctl(bdev, mode);
+ default:
+ break;
+ }
+
+ return -ENOTTY;
+}
diff --git a/block/blk.h b/block/blk.h
index 2cc03aa54c532..274afb46a809a 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -578,6 +578,10 @@ int blkdev_report_zones_ioctl(struct block_device *bdev, unsigned int cmd,
int blkdev_zone_mgmt_ioctl(struct block_device *bdev, blk_mode_t mode,
unsigned int cmd, unsigned long arg);
bool bdev_zone_mgmt_allowed(struct block_device *bdev, sector_t sector);
+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector);
+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp);
#else /* CONFIG_BLK_DEV_ZONED */
static inline void disk_init_zone_resources(struct gendisk *disk)
{
@@ -625,6 +629,17 @@ static inline bool bdev_zone_mgmt_allowed(struct block_device *bdev,
{
return false;
}
+static inline bool bdev_zone_is_offline(struct block_device *bdev,
+ sector_t sector)
+{
+ return false;
+}
+static inline int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp)
+{
+ return -ENOTTY;
+}
#endif /* CONFIG_BLK_DEV_ZONED */
struct block_device *bdev_alloc(struct gendisk *disk, u8 partno);
diff --git a/block/ioctl.c b/block/ioctl.c
index 64b4e6c0f6965..ccf806c2e37d1 100644
--- a/block/ioctl.c
+++ b/block/ioctl.c
@@ -679,6 +679,11 @@ static int blkdev_common_ioctl(struct block_device *bdev, blk_mode_t mode,
return put_uint(argp, bdev_zone_sectors(bdev));
case BLKGETNRZONES:
return put_uint(argp, bdev_nr_zones(bdev));
+ case BLKGETNRSTORELEMS:
+ case BLKREPORTSTORELEMS:
+ case BLKREMOVESTORELEM:
+ case BLKRESTORESTORELEMS:
+ return blkdev_zone_storage_elements_ioctl(bdev, mode, cmd, argp);
case BLKROGET:
return put_int(argp, bdev_read_only(bdev) != 0);
case BLKSSZGET: /* get block device logical block size */
diff --git a/drivers/block/zloop.c b/drivers/block/zloop.c
index 394dc2408ed73..3913bd2299108 100644
--- a/drivers/block/zloop.c
+++ b/drivers/block/zloop.c
@@ -37,6 +37,8 @@ enum {
ZLOOP_OPT_ORDERED_ZONE_APPEND = (1 << 10),
ZLOOP_OPT_DISCARD_WRITE_CACHE = (1 << 11),
ZLOOP_OPT_MAX_OPEN_ZONES = (1 << 12),
+ ZLOOP_OPT_STOR_ELEMENTS = (1 << 13),
+ ZLOOP_OPT_ELEMENT_ID = (1 << 14),
};
static const match_table_t zloop_opt_tokens = {
@@ -51,11 +53,23 @@ static const match_table_t zloop_opt_tokens = {
{ ZLOOP_OPT_BUFFERED_IO, "buffered_io" },
{ ZLOOP_OPT_ZONE_APPEND, "zone_append=%u" },
{ ZLOOP_OPT_ORDERED_ZONE_APPEND, "ordered_zone_append" },
- { ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
+ { ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
{ ZLOOP_OPT_MAX_OPEN_ZONES, "max_open_zones=%u" },
+ { ZLOOP_OPT_STOR_ELEMENTS, "stor_elements=%u" },
+ { ZLOOP_OPT_ELEMENT_ID, "element_id=%u" },
{ ZLOOP_OPT_ERR, NULL }
};
+/* Storage elements emulation types. */
+enum zloop_stor_elements {
+ /* No emulation. */
+ ZLOOP_STOR_ELEMENTS_NONE,
+ /* Emulate read+write storage elements. */
+ ZLOOP_STOR_ELEMENTS_RDWR,
+ /* Emulate pairs of associated read and write storage elements. */
+ ZLOOP_STOR_ELEMENTS_PAIRS,
+};
+
/* Default values for the "add" operation. */
#define ZLOOP_DEF_ID -1
#define ZLOOP_DEF_ZONE_SIZE ((256ULL * SZ_1M) >> SECTOR_SHIFT)
@@ -68,6 +82,8 @@ static const match_table_t zloop_opt_tokens = {
#define ZLOOP_DEF_BUFFERED_IO false
#define ZLOOP_DEF_ZONE_APPEND true
#define ZLOOP_DEF_ORDERED_ZONE_APPEND false
+#define ZLOOP_DEF_STOR_ELEMENTS ZLOOP_STOR_ELEMENTS_NONE
+#define ZLOOP_DEF_ELEMENT_ID 0
/* Arbitrary limit on the zone size (16GB). */
#define ZLOOP_MAX_ZONE_SIZE_MB 16384
@@ -87,6 +103,8 @@ struct zloop_options {
bool zone_append;
bool ordered_zone_append;
bool discard_write_cache;
+ enum zloop_stor_elements stor_elements;
+ unsigned int element_id;
};
/*
@@ -117,6 +135,8 @@ struct zloop_zone {
enum blk_zone_cond cond;
sector_t start;
sector_t wp;
+ unsigned int wr_se_id;
+ unsigned int rd_se_id;
gfp_t old_gfp_mask;
};
@@ -133,6 +153,7 @@ struct zloop_device {
bool zone_append;
bool ordered_zone_append;
bool discard_write_cache;
+ enum zloop_stor_elements stor_elements;
const char *base_dir;
struct file *data_dir;
@@ -150,6 +171,17 @@ struct zloop_device {
struct list_head open_zones_lru_list;
unsigned int nr_open_zones;
+ /* For storage elements emulation. */
+ struct mutex stor_elements_lock;
+ unsigned int nr_elements;
+ unsigned int max_nr_removed_elements;
+ unsigned int nr_removed_elements;
+ struct delayed_work remove_element_work;
+ unsigned int remove_element_id;
+ struct delayed_work restore_elements_work;
+ bool restore_in_progress;
+ struct blk_storage_element *elements;
+
struct zloop_zone zones[] __counted_by(nr_zones);
};
@@ -289,22 +321,30 @@ static bool zloop_do_open_zone(struct zloop_device *zlo,
}
}
-static void zloop_mark_full(struct zloop_device *zlo, struct zloop_zone *zone)
+static void zloop_set_zone_cond(struct zloop_device *zlo,
+ struct zloop_zone *zone,
+ enum blk_zone_cond cond)
{
lockdep_assert_held(&zone->wp_lock);
zloop_lru_remove_open_zone(zlo, zone);
- zone->cond = BLK_ZONE_COND_FULL;
- zone->wp = ULLONG_MAX;
+ zone->cond = cond;
+ if (cond == BLK_ZONE_COND_EMPTY)
+ zone->wp = zone->start;
+ else
+ zone->wp = ULLONG_MAX;
}
-static void zloop_mark_empty(struct zloop_device *zlo, struct zloop_zone *zone)
+static inline void zloop_set_zone_full(struct zloop_device *zlo,
+ struct zloop_zone *zone)
{
- lockdep_assert_held(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_FULL);
+}
- zloop_lru_remove_open_zone(zlo, zone);
- zone->cond = BLK_ZONE_COND_EMPTY;
- zone->wp = zone->start;
+static inline void zloop_set_zone_empty(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_EMPTY);
}
static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
@@ -339,9 +379,9 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
spin_lock(&zone->wp_lock);
if (!file_sectors) {
- zloop_mark_empty(zlo, zone);
+ zloop_set_zone_empty(zlo, zone);
} else if (file_sectors == zlo->zone_capacity) {
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
} else {
if (zone->cond != BLK_ZONE_COND_IMP_OPEN &&
zone->cond != BLK_ZONE_COND_EXP_OPEN)
@@ -353,6 +393,19 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
return 0;
}
+static bool zloop_zone_is_offline_or_readonly(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ bool ret;
+
+ spin_lock(&zone->wp_lock);
+ ret = zone->cond == BLK_ZONE_COND_OFFLINE ||
+ zone->cond == BLK_ZONE_COND_READONLY;
+ spin_unlock(&zone->wp_lock);
+
+ return ret;
+}
+
static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)
{
struct zloop_zone *zone = &zlo->zones[zone_no];
@@ -363,6 +416,11 @@ static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
ret = zloop_update_seq_zone(zlo, zone_no);
if (ret)
@@ -388,6 +446,11 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
ret = zloop_update_seq_zone(zlo, zone_no);
if (ret)
@@ -420,7 +483,22 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)
return ret;
}
-static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)
+static int zloop_do_reset_zone(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ if (vfs_truncate(&zone->file->f_path, 0))
+ return -EIO;
+
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_empty(zlo, zone);
+ clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
+ spin_unlock(&zone->wp_lock);
+
+ return 0;
+}
+
+static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no,
+ bool all_zones)
{
struct zloop_zone *zone = &zlo->zones[zone_no];
int ret = 0;
@@ -430,20 +508,19 @@ static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ if (!all_zones)
+ ret = -EIO;
+ goto unlock;
+ }
+
if (!test_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags) &&
zone->cond == BLK_ZONE_COND_EMPTY)
goto unlock;
- if (vfs_truncate(&zone->file->f_path, 0)) {
+ ret = zloop_do_reset_zone(zlo, zone);
+ if (ret)
set_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
- ret = -EIO;
- goto unlock;
- }
-
- spin_lock(&zone->wp_lock);
- zloop_mark_empty(zlo, zone);
- clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
- spin_unlock(&zone->wp_lock);
unlock:
mutex_unlock(&zone->lock);
@@ -457,7 +534,7 @@ static int zloop_reset_all_zones(struct zloop_device *zlo)
int ret;
for (i = zlo->nr_conv_zones; i < zlo->nr_zones; i++) {
- ret = zloop_reset_zone(zlo, i);
+ ret = zloop_reset_zone(zlo, i, true);
if (ret)
return ret;
}
@@ -475,6 +552,11 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (!test_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags) &&
zone->cond == BLK_ZONE_COND_FULL)
goto unlock;
@@ -487,7 +569,7 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)
}
spin_lock(&zone->wp_lock);
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
spin_unlock(&zone->wp_lock);
@@ -624,7 +706,7 @@ static int zloop_seq_write_prep(struct zloop_cmd *cmd)
if (!is_append || !zlo->ordered_zone_append) {
zone->wp += nr_sectors;
if (zone->wp == zone_end)
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
}
out_unlock:
spin_unlock(&zone->wp_lock);
@@ -766,7 +848,7 @@ static void zloop_handle_cmd(struct zloop_cmd *cmd)
cmd->ret = zloop_flush(zlo);
break;
case REQ_OP_ZONE_RESET:
- cmd->ret = zloop_reset_zone(zlo, rq_zone_no(rq));
+ cmd->ret = zloop_reset_zone(zlo, rq_zone_no(rq), false);
break;
case REQ_OP_ZONE_RESET_ALL:
cmd->ret = zloop_reset_all_zones(zlo);
@@ -875,30 +957,69 @@ static void zloop_complete_rq(struct request *rq)
blk_mq_end_request(rq, sts);
}
-static bool zloop_set_zone_append_sector(struct request *rq)
+static bool zloop_set_zone_append_sector(struct zloop_device *zlo,
+ struct zloop_zone *zone,
+ struct request *rq)
{
- struct zloop_device *zlo = rq->q->queuedata;
- unsigned int zone_no = rq_zone_no(rq);
- struct zloop_zone *zone = &zlo->zones[zone_no];
sector_t zone_end = zone->start + zlo->zone_capacity;
sector_t nr_sectors = blk_rq_sectors(rq);
- spin_lock(&zone->wp_lock);
-
if (zone->cond == BLK_ZONE_COND_FULL ||
- zone->wp + nr_sectors > zone_end) {
- spin_unlock(&zone->wp_lock);
+ zone->wp + nr_sectors > zone_end)
return false;
- }
rq->__sector = zone->wp;
zone->wp += blk_rq_sectors(rq);
if (zone->wp >= zone_end)
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
+
+ return true;
+}
+
+
+static bool zloop_prep_rq(struct zloop_device *zlo, struct request *rq)
+{
+ struct zloop_zone *zone = &zlo->zones[rq_zone_no(rq)];
+ bool is_write = op_is_write(req_op(rq));
+ bool ret = true;
+
+ spin_lock(&zone->wp_lock);
+
+ if (zlo->nr_elements) {
+ struct blk_storage_element *se;
+
+ if (zone->cond == BLK_ZONE_COND_OFFLINE ||
+ (zone->cond == BLK_ZONE_COND_READONLY && is_write)) {
+ ret = false;
+ goto unlock;
+ }
+
+ /*
+ * Check the health state of the storage element serving the
+ * zone.
+ */
+ if (is_write)
+ se = &zlo->elements[zone->wr_se_id - 1];
+ else
+ se = &zlo->elements[zone->rd_se_id - 1];
+ if (READ_ONCE(se->status) == BLK_SE_STS_DEGRADED) {
+ ret = false;
+ goto unlock;
+ }
+ }
+
+ /*
+ * If we need to strongly order zone append operations, set the request
+ * sector to the zone write pointer location now instead of when the
+ * command work runs.
+ */
+ if (zlo->ordered_zone_append && req_op(rq) == REQ_OP_ZONE_APPEND)
+ ret = zloop_set_zone_append_sector(zlo, zone, rq);
+unlock:
spin_unlock(&zone->wp_lock);
- return true;
+ return ret;
}
static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,
@@ -913,14 +1034,15 @@ static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,
return BLK_STS_IOERR;
}
- /*
- * If we need to strongly order zone append operations, set the request
- * sector to the zone write pointer location now instead of when the
- * command work runs.
- */
- if (zlo->ordered_zone_append && req_op(rq) == REQ_OP_ZONE_APPEND) {
- if (!zloop_set_zone_append_sector(rq))
+ switch (req_op(rq)) {
+ case REQ_OP_READ:
+ case REQ_OP_WRITE:
+ case REQ_OP_ZONE_APPEND:
+ if (!zloop_prep_rq(zlo, rq))
return BLK_STS_IOERR;
+ break;
+ default:
+ break;
}
blk_mq_start_request(rq);
@@ -1002,11 +1124,277 @@ static int zloop_report_zones(struct gendisk *disk, sector_t sector,
return nr_zones;
}
+static int zloop_report_elements(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct zloop_device *zlo = disk->private_data;
+ unsigned int nr_report = *nr_elements;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ *nr_elements = zlo->nr_elements;
+ if (elements) {
+ struct blk_storage_element *se = zlo->elements;
+ unsigned int i;
+
+ for (i = 0; i < min(nr_report, zlo->nr_elements); i++, se++) {
+ switch (READ_ONCE(se->status)) {
+ case BLK_SE_STS_REMOVED:
+ case BLK_SE_STS_RESTORE_ERROR:
+ se->restore_allowed = 1;
+ break;
+ default:
+ se->restore_allowed = 0;
+ }
+ memcpy(&elements[i], se,
+ sizeof(struct blk_storage_element));
+ }
+ }
+
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return 0;
+}
+
+static int zloop_degrade_element(struct zloop_device *zlo,
+ unsigned int element_id)
+{
+ struct blk_storage_element *se;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ if (!element_id || element_id > zlo->nr_elements)
+ return -EINVAL;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ se = &zlo->elements[element_id - 1];
+ if (READ_ONCE(se->status) == BLK_SE_STS_OK)
+ WRITE_ONCE(se->status, BLK_SE_STS_DEGRADED);
+ else
+ ret = -EINVAL;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
+static void zloop_remove_element_work(struct work_struct *work)
+{
+ struct zloop_device *zlo = container_of(work, struct zloop_device,
+ remove_element_work.work);
+ struct blk_storage_element *se, *paired_se = NULL;
+ enum blk_zone_cond cond;
+ struct zloop_zone *zone;
+ unsigned int i;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ /*
+ * If the element to remove is a read element, zones must go offline and
+ * the associated write element is also removed.
+ */
+ se = &zlo->elements[zlo->remove_element_id - 1];
+ switch (se->type) {
+ case BLK_SE_TYPE_RDWR:
+ cond = BLK_ZONE_COND_OFFLINE;
+ break;
+ case BLK_SE_TYPE_READ:
+ cond = BLK_ZONE_COND_OFFLINE;
+ paired_se = &zlo->elements[se->paired_id - 1];
+ break;
+ case BLK_SE_TYPE_WRITE:
+ cond = BLK_ZONE_COND_READONLY;
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ }
+
+ /*
+ * Change the condition of the zones owned by the (pair of) elements
+ * being removed and mark the elements removed.
+ */
+ for (i = 0, zone = zlo->zones; i < zlo->nr_zones; i++, zone++) {
+ if (zone->wr_se_id != se->id && zone->rd_se_id != se->id)
+ continue;
+ mutex_lock(&zone->lock);
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, cond);
+ spin_unlock(&zone->wp_lock);
+ mutex_unlock(&zone->lock);
+ }
+
+ WRITE_ONCE(se->status, BLK_SE_STS_REMOVED);
+
+ zlo->nr_removed_elements++;
+ if (paired_se && READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED) {
+ WRITE_ONCE(paired_se->status, BLK_SE_STS_REMOVED);
+ zlo->nr_removed_elements++;
+ }
+
+ zlo->remove_element_id = 0;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+}
+
+static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)
+{
+ struct zloop_device *zlo = disk->private_data;
+ struct blk_storage_element *se, *paired_se = NULL;
+ unsigned int nr_remove = 1;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ if (element_id > zlo->nr_elements)
+ return -EINVAL;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ if (zlo->remove_element_id || zlo->restore_in_progress) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /*
+ * Get the element to remove. If it is already removed, we have nothing
+ * to do.
+ */
+ se = &zlo->elements[element_id - 1];
+ if (READ_ONCE(se->status) == BLK_SE_STS_REMOVED)
+ goto unlock;
+
+ /*
+ * If the element to remove is a read element, the associated write
+ * element must also be removed.
+ */
+ if (se->type == BLK_SE_TYPE_READ) {
+ paired_se = &zlo->elements[se->paired_id - 1];
+ if (READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED)
+ nr_remove = 2;
+ }
+
+ if (zlo->nr_removed_elements + nr_remove >
+ zlo->max_nr_removed_elements) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /*
+ * Schedule the element removal with a delay, to emulate the (generally
+ * short) time it takes for a real device to depopulate a head and
+ * modify the zones.
+ */
+ zlo->remove_element_id = element_id;
+ WRITE_ONCE(se->status, BLK_SE_STS_REMOVE_IN_PROGRESS);
+ if (paired_se && READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED)
+ WRITE_ONCE(paired_se->status, BLK_SE_STS_REMOVE_IN_PROGRESS);
+
+ schedule_delayed_work(&zlo->remove_element_work,
+ msecs_to_jiffies(2000));
+
+unlock:
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
+static void zloop_restore_elements_work(struct work_struct *work)
+{
+ struct zloop_device *zlo = container_of(work, struct zloop_device,
+ restore_elements_work.work);
+ struct zloop_zone *zone = zlo->zones;
+ struct blk_storage_element *se;
+ unsigned int i;
+ int ret = 0;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ /* Reset all zones. */
+ for (i = 0; i < zlo->nr_zones && ret == 0; i++, zone++) {
+ mutex_lock(&zone->lock);
+ if (test_bit(ZLOOP_ZONE_CONV, &zone->flags)) {
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_NOT_WP);
+ spin_unlock(&zone->wp_lock);
+ } else {
+ ret = zloop_do_reset_zone(zlo, zone);
+ }
+ mutex_unlock(&zone->lock);
+ }
+
+ /* Restore all removed elements. */
+ for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
+ if (READ_ONCE(se->status) != BLK_SE_STS_RESTORE_IN_PROGRESS)
+ continue;
+ if (!ret)
+ WRITE_ONCE(se->status, BLK_SE_STS_OK);
+ else
+ WRITE_ONCE(se->status, BLK_SE_STS_RESTORE_ERROR);
+ }
+
+ zlo->nr_removed_elements = 0;
+ zlo->restore_in_progress = false;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+}
+
+static int zloop_restore_elements(struct gendisk *disk)
+{
+ struct zloop_device *zlo = disk->private_data;
+ struct blk_storage_element *se;
+ unsigned int i;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ if (zlo->remove_element_id || zlo->restore_in_progress) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /* If we have no removed elements, we have nothing to do. */
+ if (!zlo->nr_removed_elements)
+ goto unlock;
+
+ /*
+ * Schedule the elements restoration with a delay, to emulate the time
+ * it takes for a real device to restore all removed heads and reset
+ * all zones.
+ */
+ zlo->restore_in_progress = true;
+ for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
+ if (READ_ONCE(se->status) == BLK_SE_STS_REMOVED)
+ WRITE_ONCE(se->status, BLK_SE_STS_RESTORE_IN_PROGRESS);
+ }
+
+ schedule_delayed_work(&zlo->restore_elements_work,
+ msecs_to_jiffies(5000));
+
+unlock:
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
static void zloop_free_disk(struct gendisk *disk)
{
struct zloop_device *zlo = disk->private_data;
unsigned int i;
+ cancel_delayed_work_sync(&zlo->remove_element_work);
+ cancel_delayed_work_sync(&zlo->restore_elements_work);
+
blk_mq_free_tag_set(&zlo->tag_set);
for (i = 0; i < zlo->nr_zones; i++) {
@@ -1019,15 +1407,24 @@ static void zloop_free_disk(struct gendisk *disk)
fput(zlo->data_dir);
destroy_workqueue(zlo->workqueue);
+ kfree(zlo->elements);
kfree(zlo->base_dir);
kvfree(zlo);
}
+
+static const struct blk_storage_elements_ops zloop_se_ops = {
+ .report_elements = zloop_report_elements,
+ .remove_element = zloop_remove_element,
+ .restore_elements = zloop_restore_elements,
+};
+
static const struct block_device_operations zloop_fops = {
.owner = THIS_MODULE,
.open = zloop_open,
.report_zones = zloop_report_zones,
.free_disk = zloop_free_disk,
+ .se_ops = &zloop_se_ops,
};
__printf(3, 4)
@@ -1112,6 +1509,24 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,
if (!opts->buffered_io)
oflags |= O_DIRECT;
+ if (zlo->stor_elements != ZLOOP_STOR_ELEMENTS_NONE) {
+ unsigned int nr_elems = zlo->nr_elements;
+ unsigned int se_idx;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ nr_elems /= 2;
+ se_idx = zone_no % nr_elems;
+ zlo->elements[se_idx].nr_zones++;
+
+ zone->wr_se_id = se_idx + 1;
+ zone->rd_se_id = zone->wr_se_id;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {
+ zone->rd_se_id += nr_elems;
+ zlo->elements[zone->rd_se_id - 1].nr_zones =
+ zlo->elements[se_idx].nr_zones;
+ }
+ }
+
if (zone_no < zlo->nr_conv_zones) {
/* Conventional zone file. */
set_bit(ZLOOP_ZONE_CONV, &zone->flags);
@@ -1182,6 +1597,79 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,
return ret;
}
+#define ZLOOP_MIN_STOR_ELEMENTS 2
+#define ZLOOP_MAX_STOR_ELEMENTS 32
+#define ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS 32
+
+static int zloop_create_storage_elements(struct zloop_device *zlo)
+{
+ struct blk_storage_element *se, *paired_se;
+ unsigned int i, nr_elems, nr_elements;
+
+ /*
+ * Calculate the number of storage elements we are going to emulate.
+ * To achieve a somewhat realistic emulation, we want at least 2 storage
+ * elements, and no more than 32, targeting at least 32 zones per
+ * element.
+ */
+ if (zlo->nr_zones <= 64)
+ nr_elements = 2;
+ else
+ nr_elements =
+ min(ZLOOP_MAX_STOR_ELEMENTS,
+ zlo->nr_zones / ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS);
+
+ /*
+ * If we are emulating pairs of read and write storage elements, we need
+ * double the number of storage element descriptors.
+ */
+ nr_elems = nr_elements;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ nr_elems *= 2;
+ zlo->elements = kzalloc_objs(struct blk_storage_element, nr_elems);
+ if (!zlo->elements)
+ return -ENOMEM;
+
+ for (i = 0, se = zlo->elements; i < nr_elements; i++, se++) {
+ se->id = i + 1;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ se->type = BLK_SE_TYPE_WRITE;
+ else
+ se->type = BLK_SE_TYPE_RDWR;
+ se->status = BLK_SE_STS_OK;
+ }
+
+ /*
+ * If we are emulating pairs of read and write storage elements,
+ * initialize the read elements paired with the write elements we just
+ * initialized.
+ */
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {
+ for (i = 0; i < nr_elements; i++, se++) {
+ se->id = nr_elements + i + 1;
+ paired_se = &zlo->elements[i];
+ se->paired_id = paired_se->id;
+ paired_se->paired_id = se->id;
+ se->type = BLK_SE_TYPE_READ;
+ se->status = BLK_SE_STS_OK;
+ }
+ }
+
+ /*
+ * Make sure we do not allow removing all storage elements as that does
+ * not make any sense. This is consistent with the device advertized
+ * limit of SCSI and ATA devices supporting the storage element
+ * depopulation feature.
+ */
+ zlo->nr_elements = nr_elems;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ zlo->max_nr_removed_elements = zlo->nr_elements - 2;
+ else
+ zlo->max_nr_removed_elements = zlo->nr_elements - 1;
+
+ return 0;
+}
+
static bool zloop_dev_exists(struct zloop_device *zlo)
{
struct file *cnv, *seq;
@@ -1237,6 +1725,10 @@ static int zloop_ctl_add(struct zloop_options *opts)
WRITE_ONCE(zlo->state, Zlo_creating);
spin_lock_init(&zlo->open_zones_lock);
INIT_LIST_HEAD(&zlo->open_zones_lru_list);
+ mutex_init(&zlo->stor_elements_lock);
+ INIT_DELAYED_WORK(&zlo->remove_element_work, zloop_remove_element_work);
+ INIT_DELAYED_WORK(&zlo->restore_elements_work,
+ zloop_restore_elements_work);
ret = mutex_lock_killable(&zloop_ctl_mutex);
if (ret)
@@ -1270,12 +1762,19 @@ static int zloop_ctl_add(struct zloop_options *opts)
if (zlo->zone_append)
zlo->ordered_zone_append = opts->ordered_zone_append;
zlo->discard_write_cache = opts->discard_write_cache;
+ zlo->stor_elements = opts->stor_elements;
+
+ if (zlo->stor_elements != ZLOOP_STOR_ELEMENTS_NONE) {
+ ret = zloop_create_storage_elements(zlo);
+ if (ret)
+ goto out_free_idr;
+ }
zlo->workqueue = alloc_workqueue("zloop%d", WQ_UNBOUND | WQ_FREEZABLE,
opts->nr_queues * opts->queue_depth, zlo->id);
if (!zlo->workqueue) {
ret = -ENOMEM;
- goto out_free_idr;
+ goto out_destroy_storage_elements;
}
if (opts->base_dir)
@@ -1361,6 +1860,10 @@ static int zloop_ctl_add(struct zloop_options *opts)
zlo->id, zlo->nr_zones,
((sector_t)zlo->zone_size << SECTOR_SHIFT) >> 20,
zlo->block_size);
+ if (zlo->nr_elements)
+ pr_info("zloop%d: %d storage elements\n",
+ zlo->id, zlo->nr_elements);
+
pr_info("zloop%d: using %s%s zone append\n",
zlo->id,
zlo->ordered_zone_append ? "ordered " : "",
@@ -1384,6 +1887,8 @@ static int zloop_ctl_add(struct zloop_options *opts)
kfree(zlo->base_dir);
out_destroy_workqueue:
destroy_workqueue(zlo->workqueue);
+out_destroy_storage_elements:
+ kfree(zlo->elements);
out_free_idr:
mutex_lock(&zloop_ctl_mutex);
idr_remove(&zloop_index_idr, zlo->id);
@@ -1483,6 +1988,42 @@ static int zloop_ctl_remove(struct zloop_options *opts)
return 0;
}
+static int zloop_ctl_degrade_element(struct zloop_options *opts)
+{
+ struct zloop_device *zlo;
+ int ret = 0;
+
+ if (!(opts->mask & ZLOOP_OPT_ID)) {
+ pr_err("No ID specified for degrade_element\n");
+ return -EINVAL;
+ }
+
+ if (opts->mask & ~(ZLOOP_OPT_ID | ZLOOP_OPT_ELEMENT_ID)) {
+ pr_err("Invalid option specified for degrade_element\n");
+ return -EINVAL;
+ }
+
+ mutex_lock(&zloop_ctl_mutex);
+
+ zlo = idr_find(&zloop_index_idr, opts->id);
+ if (!zlo || zlo->state == Zlo_creating)
+ ret = -ENODEV;
+ else if (zlo->state == Zlo_deleting)
+ ret = -EINVAL;
+ if (ret)
+ goto unlock;
+
+ ret = zloop_degrade_element(zlo, opts->element_id);
+ if (!ret)
+ pr_info("Degraded element %u of device %u\n",
+ opts->id, opts->element_id);
+
+unlock:
+ mutex_unlock(&zloop_ctl_mutex);
+
+ return ret;
+}
+
static int zloop_parse_options(struct zloop_options *opts, const char *buf)
{
substring_t args[MAX_OPT_ARGS];
@@ -1502,6 +2043,8 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
opts->buffered_io = ZLOOP_DEF_BUFFERED_IO;
opts->zone_append = ZLOOP_DEF_ZONE_APPEND;
opts->ordered_zone_append = ZLOOP_DEF_ORDERED_ZONE_APPEND;
+ opts->stor_elements = ZLOOP_DEF_STOR_ELEMENTS;
+ opts->element_id = ZLOOP_DEF_ELEMENT_ID;
if (!buf)
return 0;
@@ -1636,6 +2179,30 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
case ZLOOP_OPT_DISCARD_WRITE_CACHE:
opts->discard_write_cache = true;
break;
+ case ZLOOP_OPT_STOR_ELEMENTS:
+ if (match_uint(args, &token)) {
+ ret = -EINVAL;
+ goto out;
+ }
+ switch (token) {
+ case ZLOOP_STOR_ELEMENTS_NONE:
+ case ZLOOP_STOR_ELEMENTS_RDWR:
+ case ZLOOP_STOR_ELEMENTS_PAIRS:
+ break;
+ default:
+ pr_err("Invalid stor_elements value\n");
+ ret = -EINVAL;
+ goto out;
+ }
+ opts->stor_elements = token;
+ break;
+ case ZLOOP_OPT_ELEMENT_ID:
+ if (match_uint(args, &token)) {
+ ret = -EINVAL;
+ goto out;
+ }
+ opts->element_id = token;
+ break;
case ZLOOP_OPT_ERR:
default:
pr_warn("unknown parameter or missing value '%s'\n", p);
@@ -1664,14 +2231,16 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
enum {
ZLOOP_CTL_ADD,
ZLOOP_CTL_REMOVE,
+ ZLOOP_CTL_DEGRADE_ELEMENT,
};
static struct zloop_ctl_op {
int code;
const char *name;
} zloop_ctl_ops[] = {
- { ZLOOP_CTL_ADD, "add" },
- { ZLOOP_CTL_REMOVE, "remove" },
+ { ZLOOP_CTL_ADD, "add" },
+ { ZLOOP_CTL_REMOVE, "remove" },
+ { ZLOOP_CTL_DEGRADE_ELEMENT, "degrade_element" },
{ -1, NULL },
};
@@ -1719,6 +2288,9 @@ static ssize_t zloop_ctl_write(struct file *file, const char __user *ubuf,
case ZLOOP_CTL_REMOVE:
ret = zloop_ctl_remove(&opts);
break;
+ case ZLOOP_CTL_DEGRADE_ELEMENT:
+ ret = zloop_ctl_degrade_element(&opts);
+ break;
default:
pr_err("Invalid operation\n");
ret = -EINVAL;
@@ -1742,6 +2314,8 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)
tok = &zloop_opt_tokens[i];
if (!tok->pattern)
break;
+ if (tok->token == ZLOOP_OPT_ELEMENT_ID)
+ continue;
if (i)
seq_putc(seq_file, ',');
seq_puts(seq_file, tok->pattern);
@@ -1752,6 +2326,10 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)
seq_puts(seq_file, zloop_ctl_ops[1].name);
seq_puts(seq_file, " id=%d\n");
+ /* Degrade element operation */
+ seq_puts(seq_file, zloop_ctl_ops[2].name);
+ seq_puts(seq_file, " id=%d,element_id=%d\n");
+
return 0;
}
diff --git a/drivers/scsi/sd.c b/drivers/scsi/sd.c
index a1b21ea14e549..4372fa80f792d 100644
--- a/drivers/scsi/sd.c
+++ b/drivers/scsi/sd.c
@@ -3937,6 +3937,9 @@ static const struct block_device_operations sd_fops = {
.get_unique_id = sd_get_unique_id,
.free_disk = scsi_disk_free_disk,
.pr_ops = &sd_pr_ops,
+#ifdef CONFIG_BLK_DEV_ZONED
+ .se_ops = &sd_zbc_se_ops,
+#endif
};
/**
diff --git a/drivers/scsi/sd.h b/drivers/scsi/sd.h
index 574af82430169..b68b7ee5a5fa5 100644
--- a/drivers/scsi/sd.h
+++ b/drivers/scsi/sd.h
@@ -156,6 +156,7 @@ struct scsi_disk {
unsigned ignore_medium_access_errors : 1;
unsigned rscs : 1; /* reduced stream control support */
unsigned use_atomic_write_boundary : 1;
+ unsigned modify_zones_supported : 1;
};
#define to_scsi_disk(obj) container_of(obj, struct scsi_disk, disk_dev)
@@ -242,6 +243,8 @@ unsigned int sd_zbc_complete(struct scsi_cmnd *cmd, unsigned int good_bytes,
int sd_zbc_report_zones(struct gendisk *disk, sector_t sector,
unsigned int nr_zones, struct blk_report_zones_args *args);
+extern const struct blk_storage_elements_ops sd_zbc_se_ops;
+
#else /* CONFIG_BLK_DEV_ZONED */
static inline int sd_zbc_read_zones(struct scsi_disk *sdkp,
diff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c
index 456beaf2e7690..6698d156a0c34 100644
--- a/drivers/scsi/sd_zbc.c
+++ b/drivers/scsi/sd_zbc.c
@@ -516,6 +516,21 @@ static int sd_zbc_check_capacity(struct scsi_disk *sdkp, unsigned char *buf,
return 0;
}
+/*
+ * sd_zbc_check_modify_zones - Check if the device supports depopulation
+ * @sdkp: Target disk
+ * @buf: command buffer
+ *
+ * Check if the device supports the REMOVE ELEMENT AND MODIFY ZONES command.
+ */
+static inline bool sd_zbc_check_modify_zones(struct scsi_disk *sdkp,
+ unsigned char *buf)
+{
+ return scsi_report_opcode(sdkp->device, buf, SD_BUF_SIZE,
+ SERVICE_ACTION_IN_16,
+ SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES) == 1;
+}
+
static void sd_zbc_print_zones(struct scsi_disk *sdkp)
{
if (sdkp->device->type != TYPE_ZBC || !sdkp->capacity)
@@ -533,6 +548,239 @@ static void sd_zbc_print_zones(struct scsi_disk *sdkp)
sdkp->zone_info.zone_blocks);
}
+static void sd_zbc_parse_storage_element(struct scsi_disk *sdkp, u8 *desc,
+ struct blk_storage_element *element)
+{
+ struct scsi_device *sdp = sdkp->device;
+ sector_t zone_sectors = sd_zbc_zone_sectors(sdkp);
+ u64 capacity;
+
+ memset(element, 0, sizeof(*element));
+
+ element->id = get_unaligned_be32(&desc[4]);
+
+ switch (desc[14]) {
+ case SCSI_PHYS_ELEM_TYPE_ALL_ACCESS_STORAGE:
+ element->type = BLK_SE_TYPE_RDWR;
+ capacity = get_unaligned_be64(&desc[16]);
+ if (!zone_sectors || capacity == ULLONG_MAX)
+ element->nr_zones = 0;
+ else
+ element->nr_zones =
+ logical_to_sectors(sdp, capacity) >>
+ ilog2(zone_sectors);
+ break;
+ case SCSI_PHYS_ELEM_TYPE_FRAC_ACCESS_STORAGE:
+ element->paired_id = get_unaligned_be32(&desc[16]);
+ if (desc[20] & 0x01)
+ element->type = BLK_SE_TYPE_READ;
+ else
+ element->type = BLK_SE_TYPE_WRITE;
+ element->nr_zones = get_unaligned_be64(&desc[24]);
+ break;
+ default:
+ element->type = BLK_SE_TYPE_UNKNOWN;
+ }
+
+ switch (desc[15]) {
+ case SCSI_PHYS_ELEM_HEALTH_WITHIN_SPEC_LIMITS:
+ case SCSI_PHYS_ELEM_HEALTH_AT_SPEC_LIMITS:
+ element->status = BLK_SE_STS_OK;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_OUTSIDE_SPEC_LIMITS:
+ element->status = BLK_SE_STS_DEGRADED;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_ERR:
+ element->status = BLK_SE_STS_RESTORE_ERROR;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_IN_PROGRESS:
+ element->status = BLK_SE_STS_RESTORE_IN_PROGRESS;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_ERR:
+ element->status = BLK_SE_STS_REMOVE_ERROR;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_IN_PROGRESS:
+ element->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_OK:
+ element->status = BLK_SE_STS_REMOVED;
+ element->restore_allowed = desc[13] & 0x01;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_NOT_REPORTED:
+ default:
+ element->status = BLK_SE_STS_UNKNOWN;
+ break;
+ }
+}
+
+/*
+ * Large hard limit on the number of storage elements. This accommodates all
+ * known devices today and likely forever :)
+ */
+#define SD_ZBC_MAX_STORAGE_ELEMENTS 255
+
+static int sd_zbc_report_storage_elements(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ unsigned int nr_descs, nr_se;
+ unsigned int buf_size;
+ int i, ret = 0, result;
+ u8 *desc, *buf;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ /*
+ * We need at least 32B for the report header and 32B for each
+ * descriptor.
+ */
+ nr_se = min(SD_ZBC_MAX_STORAGE_ELEMENTS, *nr_elements);
+ buf_size = ALIGN((nr_se + 1) * 32, SECTOR_SIZE);
+
+ buf = kzalloc(buf_size, GFP_KERNEL);
+ if (!buf)
+ return -ENOMEM;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_GET_PHYSICAL_ELEMENT_STATUS;
+ put_unaligned_be32(buf_size, &cmd[10]);
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, buf, buf_size,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "GET PHYSICAL ELEMENT STATUS failed\n");
+ sd_print_result(sdkp, "GET PHYSICAL ELEMENT STATUS", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ nr_descs = get_unaligned_be32(&buf[0]);
+ if (!nr_descs) {
+ sd_printk(KERN_ERR, sdkp,
+ "Invalid number of phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+ if (nr_descs > SD_ZBC_MAX_STORAGE_ELEMENTS) {
+ sd_printk(KERN_ERR, sdkp,
+ "Unsupported number of phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ if (!elements) {
+ *nr_elements = nr_descs;
+ goto free_buf;
+ }
+
+ nr_descs = get_unaligned_be32(&buf[4]);
+ if (!nr_descs) {
+ sd_printk(KERN_ERR, sdkp,
+ "Invalid number of reported phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ desc = &buf[32];
+ for (i = 0; i < min(nr_se, nr_descs); i++, desc += 32)
+ sd_zbc_parse_storage_element(sdkp, desc, &elements[i]);
+ *nr_elements = i;
+
+free_buf:
+ kfree(buf);
+
+ return ret;
+}
+
+static int sd_zbc_remove_storage_element(struct gendisk *disk,
+ unsigned int element_id)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ int result;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES;
+ put_unaligned_be32(element_id, &cmd[10]);
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "REMOVE ELEMENT AND MODIFY ZONES failed\n");
+ sd_print_result(sdkp,
+ "REMOVE ELEMENT AND MODIFY ZONES", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+static int sd_zbc_restore_storage_elements(struct gendisk *disk)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ int result;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_RESTORE_ELEMENTS_AND_REBUILD;
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "RESTORE ELEMENTS AND REBUILD failed\n");
+ sd_print_result(sdkp,
+ "RESTORE ELEMENTS AND REBUILD", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+const struct blk_storage_elements_ops sd_zbc_se_ops = {
+ .report_elements = sd_zbc_report_storage_elements,
+ .remove_element = sd_zbc_remove_storage_element,
+ .restore_elements = sd_zbc_restore_storage_elements,
+};
+
/*
* Call blk_revalidate_disk_zones() if any of the zoned disk properties have
* changed that make it necessary to call that function. Called by
@@ -554,9 +802,15 @@ int sd_zbc_revalidate_zones(struct scsi_disk *sdkp)
if (!blk_queue_is_zoned(q))
return 0;
+ /*
+ * If the zone size and number of zones has not changed, and the disk
+ * does not support depopulating heads, skip the rather slow call to
+ * blk_revalidate_disk_zones().
+ */
if (sdkp->zone_info.zone_blocks == zone_blocks &&
sdkp->zone_info.nr_zones == nr_zones &&
- disk->nr_zones == nr_zones)
+ disk->nr_zones == nr_zones &&
+ !sdkp->modify_zones_supported)
return 0;
sdkp->zone_info.zone_blocks = zone_blocks;
@@ -620,6 +874,9 @@ int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim,
if (ret != 0)
goto err;
+ /* Check if REMOVE ELEMENT AND MODIFY ZONES is supported. */
+ sdkp->modify_zones_supported = sd_zbc_check_modify_zones(sdkp, buf);
+
nr_zones = round_up(sdkp->capacity, zone_blocks) >> ilog2(zone_blocks);
if (nr_zones > INT_MAX) {
sd_printk(KERN_ERR, sdkp, "Too many zones (%llu)\n",
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index d003a9d2d1f6c..859917b3b2568 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1574,6 +1574,21 @@ enum blk_unique_id {
BLK_UID_NAA = 3,
};
+struct blk_storage_elements_ops {
+ int (*report_elements)(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements);
+ int (*remove_element)(struct gendisk *disk, unsigned int element_id);
+ int (*restore_elements)(struct gendisk *disk);
+};
+
+int bdev_report_storage_elements(struct block_device *bdev,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements);
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id);
+int bdev_restore_storage_elements(struct block_device *bdev);
+
struct block_device_operations {
void (*submit_bio)(struct bio *bio);
int (*poll_bio)(struct bio *bio, struct io_comp_batch *iob,
@@ -1601,6 +1616,7 @@ struct block_device_operations {
enum blk_unique_id id_type);
struct module *owner;
const struct pr_ops *pr_ops;
+ const struct blk_storage_elements_ops *se_ops;
/*
* Special callback for probing GPT entry at a given sector.
diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
index 6638361209667..532cf69b5b114 100644
--- a/include/uapi/linux/blkzoned.h
+++ b/include/uapi/linux/blkzoned.h
@@ -208,4 +208,85 @@ struct blk_zone_range {
#define BLKFINISHZONE _IOW(0x12, 136, struct blk_zone_range)
#define BLKREPORTZONEV2 _IOWR(0x12, 142, struct blk_zone_report)
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_TYPE_RDWR: The storage element handles both reads and writes.
+ * @BLK_SE_TYPE_READ: The storage element handles reads only.
+ * @BLK_SE_TYPE_WRITE: The storage element handles writes only.
+ * @BLK_SE_TYPE_UNKNOWN: The storage element type is not known.
+ */
+enum blk_storage_element_type {
+ BLK_SE_TYPE_RDWR = 0x01,
+ BLK_SE_TYPE_READ = 0x02,
+ BLK_SE_TYPE_WRITE = 0x03,
+ BLK_SE_TYPE_UNKNOWN = 0xFF,
+};
+
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_STS_OK: The storage element is operating normally.
+ * @BLK_SE_STS_DEGRADED: The storage element has degraded and is not operating
+ * normally.
+ * @BLK_SE_STS_REMOVE_IN_PROGRESS: The storage element is being removed.
+ * @BLK_SE_STS_REMOVE_ERROR: The storage element removal failed.
+ * @BLK_SE_STS_RESTORE_IN_PROGRESS: The storage element is being restored.
+ * @BLK_SE_STS_RESTORE_ERROR: The storage element restoration failed.
+ * @BLK_SE_STS_REMOVED: The storage element was removed.
+ * @BLK_SE_STS_UNKNOWN: The storage element status is unknown.
+ */
+enum blk_storage_element_status {
+ BLK_SE_STS_OK = 0x01,
+ BLK_SE_STS_DEGRADED = 0x02,
+ BLK_SE_STS_REMOVE_IN_PROGRESS = 0x03,
+ BLK_SE_STS_REMOVE_ERROR = 0x04,
+ BLK_SE_STS_RESTORE_IN_PROGRESS = 0x05,
+ BLK_SE_STS_RESTORE_ERROR = 0x06,
+ BLK_SE_STS_REMOVED = 0x07,
+ BLK_SE_STS_UNKNOWN = 0xFF,
+};
+
+/**
+ * struct blk_storage_element - Zoned device storage element descriptor.
+ *
+ * @id: The ID of the element (cannot be 0).
+ * @paired_id: The ID of the paired element for an element that is not
+ * of type BLK_SE_TYPE_RDWR.
+ * @type: The type of the storage element (enum blk_storage_element_type).
+ * @status: The health status of the storage element
+ * (enum blk_storage_element_status).
+ * @restore_allowed: Indicate if the storage element can be restored.
+ * @nr_zones: The number of zones that the storage element handles.
+ */
+struct blk_storage_element {
+ __u32 id;
+ __u32 paired_id;
+ __u64 nr_zones;
+ __u8 type;
+ __u8 status;
+ __u8 restore_allowed;
+ __u8 reserved[5];
+};
+
+struct blk_storage_elements_report {
+ __u32 nr_elements;
+ __u32 reserved;
+ struct blk_storage_element elements[];
+};
+
+/**
+ * Zoned block device storage element management ioctl's:
+ *
+ * @BLKGETNRSTORELEMS: Get the number of storage elements of the device.
+ * @BLKREPORTSTORELEMS: Get the device storage elements.
+ * @BLKREMOVESTORELEM: Remove (depopulate) one storage element of a device.
+ * @BLKRESTORESTORELEM: Restore (repopulate if possible) all storage elements
+ * that have been removed.
+ */
+#define BLKGETNRSTORELEMS _IOR(0x12, 143, __u32)
+#define BLKREPORTSTORELEMS _IOWR(0x12, 144, struct blk_storage_elements_report)
+#define BLKREMOVESTORELEM _IOW(0x12, 145, __u32)
+#define BLKRESTORESTORELEMS _IO(0x12, 146)
+
#endif /* _UAPI_BLKZONED_H */
diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index 34c6f219462a5..8a979326aa7f0 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -309,6 +309,7 @@ struct file_attr {
/* 130-136 and 142 are used by zoned block device ioctls (uapi/linux/blkzoned.h) */
/* 137-141 are used by blk-crypto ioctls (uapi/linux/blk-crypto.h) */
#define BLKTRACESETUP2 _IOWR(0x12, 142, struct blk_user_trace_setup2)
+/* 143-146 are used by storage element management for zoned block devices. */
#define BMAP_IOCTL 1 /* obsolete - kept for compatibility */
#define FIBMAP _IO(0x00,1) /* bmap access */
|
| 27/3 | 2026/10/06 13:29 | llm | kmsan-evaluator |
6mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 25313 output: 18 thoughts: 3752 Thoughts: **Analysis of Zoned Block Device Storage Element Management**
Okay, here's the breakdown of this patch, as I see it. This patch introduces storage element management for zoned block devices (blk-zoned, zloop, and sd_zbc). My primary concern here, as always, is whether we're introducing any new uninitialized memory vulnerabilities. I need to make sure we're not leaking kernel information or introducing ways for userspace to read uninitialized memory.
First, I examined all the places where kernel structures are being sent to user space (using `copy_to_user`, `put_user`, etc.).
* In `blkdev_get_nr_storage_elements_ioctl`, the `nr_elements` variable is safe. It's initialized to zero and then either overwritten or left at zero before the `put_user` call, which copies a single, fully initialized 4-byte value. No issues there.
* Then, `blkdev_report_storage_elements_ioctl`... Here's where it gets interesting and where my spidey-sense started tingling. The `struct blk_storage_elements_report rep` gets copied from user space *without* initialization. The size is 8 bytes, from `copy_from_user` and the fields are `nr_elements` and `reserved`. The structure itself doesn't have any padding and is fully filled by the user, initially. Then `kzalloc_objs` is used to allocate the array of `struct blk_storage_element` elements, so these are all zeroed. Good. I need to know what a `struct blk_storage_element` contains. It's a standard structure, with the sizes documented. Looks good so far.
Now, let's trace this and see what happens with the `elements` array. I'm especially concerned about potential out-of-bounds reads and writes.
* The code calls `bdev_report_storage_elements`, which then calls into the specific driver implementations. I checked out `zloop`, where `zloop_report_elements` copies storage element data. *Here's where it goes haywire*. `zloop_report_elements` *overwrites* the `*nr_elements` variable passed in, setting it to `zlo->nr_elements`, which represents the *actual* number of storage elements. *This is extremely problematic!*
In the main `ioctl` handler, the number of allocated elements (`nr_elements`) can be limited to a smaller value (say 2) based on the user-provided `rep.nr_elements`.
`kzalloc_objs` allocates space for *that* user-specified number (2 elements).
But when the driver calls `bdev_report_storage_elements`, `*nr_elements` is *overwritten* with the *actual* number of elements managed by the driver (say 10). *This means the subsequent `copy_to_user` writes a larger data than allocated.*
So, then the code calls `copy_to_user` with the *original* size allocation. Now we copy `elements` buffer *with the larger actual size*.
**HOLY COW!** This is an out-of-bounds read (and write) vulnerability! KASAN *should* be catching this slab-out-of-bounds access on that `copy_to_user` call, and I bet it does. I should check the build configurations to verify this, but I am fairly certain.
* `sd_zbc.c` seems fine, it uses `memset`, so no issues there.
* The potential memory corruption is not from uninitialized data, but from a clear out-of-bounds write.
Based on this analysis, while I didn't see explicit uninitialized memory access, the control flow and data size calculations *do* introduce an out-of-bounds write. This is a critical security issue! This is caught by KASAN. So, NeedsKMSAN should be false!
|
| 28/3 | 2026/10/06 13:36 | tool | grepper |
0mArgs:{
"Expression": "kzalloc_objs"
}Results: {
"Output": "\nFull output is too long, showing 200 out of 14599 lines.\nUse more precise expression or PathPrefixes if possible.\n\nNumber of matching lines per file (1543 files in total):\nDocumentation/process/coding-style.rst:2\nDocumentation/process/deprecated.rst:1\nDocumentation/translations/it_IT/process/coding-style.rst:2\nDocumentation/translations/it_IT/process/deprecated.rst:1\nDocumentation/translations/pt_BR/process/deprecated.rst:1\narch/alpha/kernel/module.c:1\narch/arm/kernel/vdso.c:1\narch/arm/mach-footbridge/dc21285.c:1\narch/arm/mach-omap1/mcbsp.c:1\narch/arm/mach-omap2/omap_device.c:1\narch/arm/mach-omap2/pm33xx-core.c:1\narch/arm/mach-omap2/sr_device.c:1\narch/arm/mach-versatile/spc.c:1\narch/arm/xen/enlighten.c:1\narch/arm64/kernel/vdso.c:1\narch/arm64/kvm/vgic/vgic-init.c:2\narch/arm64/kvm/vgic/vgic-irqfd.c:1\narch/arm64/kvm/vgic/vgic-v4.c:1\narch/csky/kernel/vdso.c:1\narch/loongarch/kernel/vdso.c:1\narch/loongarch/kvm/intc/pch_pic.c:1\narch/mips/alchemy/common/clock.c:2\narch/mips/alchemy/common/dbdma.c:1\narch/mips/alchemy/common/platform.c:1\narch/mips/alchemy/devboards/platform.c:2\narch/mips/bcm47xx/buttons.c:4\narch/mips/kernel/smp-cps.c:3\narch/parisc/kernel/vdso.c:1\narch/parisc/net/bpf_jit_core.c:1\narch/powerpc/kernel/smp.c:1\narch/powerpc/kernel/vdso.c:1\narch/powerpc/kvm/e500_mmu.c:4\narch/powerpc/mm/drmem.c:2\narch/powerpc/perf/imc-pmu.c:7\narch/powerpc/perf/vpa-dtl.c:1\narch/powerpc/platforms/44x/pci.c:1\narch/powerpc/platforms/powernv/idle.c:1\narch/powerpc/platforms/powernv/memtrace.c:1\narch/powerpc/platforms/powernv/opal-async.c:1\narch/powerpc/platforms/powernv/opal-imc.c:1\narch/powerpc/platforms/powernv/opal-irqchip.c:1\narch/powerpc/platforms/powernv/opal-powercap.c:3\narch/powerpc/platforms/powernv/opal-psr.c:1\narch/powerpc/platforms/powernv/opal-sensor-groups.c:3\narch/powerpc/platforms/powernv/opal-sysparam.c:1\narch/powerpc/platforms/pseries/papr_platform_attributes.c:2\narch/powerpc/sysdev/mpic.c:1\narch/powerpc/sysdev/mpic_msgr.c:1\narch/powerpc/sysdev/xive/common.c:1\narch/riscv/kernel/tests/kprobes/test-kprobes.c:1\n... and 1493 more files\n\nDocumentation/process/coding-style.rst=938=The kernel provides the following general purpose memory allocators:\nDocumentation/process/coding-style.rst:939:kmalloc(), kzalloc(), kmalloc_objs(), kzalloc_objs(), vmalloc(), and\nDocumentation/process/coding-style.rst-940-vzalloc(). Please refer to the API documentation for further information\n--\nDocumentation/process/coding-style.rst=964=The preferred form for allocating a zeroed array is the following:\n--\nDocumentation/process/coding-style.rst-967-\nDocumentation/process/coding-style.rst:968:\tp = kzalloc_objs(*p, n, ...);\nDocumentation/process/coding-style.rst-969-\n--\nDocumentation/process/deprecated.rst=398=become, respectively::\n--\nDocumentation/process/deprecated.rst-402-\tptr = kmalloc_objs(*ptr, count [, gfp] );\nDocumentation/process/deprecated.rst:403:\tptr = kzalloc_objs(*ptr, count [, gfp] );\nDocumentation/process/deprecated.rst-404-\tptr = kmalloc_flex(*ptr, flex_member, count [, gfp] );\n--\nDocumentation/translations/it_IT/process/coding-style.rst=940=Il kernel fornisce i seguenti assegnatori ad uso generico:\nDocumentation/translations/it_IT/process/coding-style.rst:941:kmalloc(), kzalloc(), kmalloc_objs(), kzalloc_objs(), vmalloc(), e vzalloc().\nDocumentation/translations/it_IT/process/coding-style.rst-942-Per maggiori informazioni, consultate la documentazione dell'API:\n--\nDocumentation/translations/it_IT/process/coding-style.rst=965=Il modo preferito per assegnare un vettore a zero è il seguente:\n--\nDocumentation/translations/it_IT/process/coding-style.rst-968-\nDocumentation/translations/it_IT/process/coding-style.rst:969:\tp = kzalloc_objs(*p, n, ...);\nDocumentation/translations/it_IT/process/coding-style.rst-970-\n--\nDocumentation/translations/it_IT/process/deprecated.rst=441=diventano, rispettivamente::\n--\nDocumentation/translations/it_IT/process/deprecated.rst-445-\tptr = kmalloc_objs(*ptr, count [, gfp] );\nDocumentation/translations/it_IT/process/deprecated.rst:446:\tptr = kzalloc_objs(*ptr, count [, gfp] );\nDocumentation/translations/it_IT/process/deprecated.rst-447-\tptr = kmalloc_flex(*ptr, flex_member, count [, gfp] );\n--\nDocumentation/translations/pt_BR/process/deprecated.rst=410=tornam-se, respectivamente::\n--\nDocumentation/translations/pt_BR/process/deprecated.rst-414- ptr = kmalloc_objs(*ptr, count [, gfp] );\nDocumentation/translations/pt_BR/process/deprecated.rst:415: ptr = kzalloc_objs(*ptr, count [, gfp] );\nDocumentation/translations/pt_BR/process/deprecated.rst-416- ptr = kmalloc_flex(*ptr, flex_member, count [, gfp] );\n--\narch/alpha/kernel/module.c=64=module_frob_arch_sections(Elf64_Ehdr *hdr, Elf64_Shdr *sechdrs,\n--\narch/alpha/kernel/module.c-95-\tnsyms = symtab-\u003esh_size / sizeof(Elf64_Sym);\narch/alpha/kernel/module.c:96:\tchains = kzalloc_objs(struct got_entry, nsyms);\narch/alpha/kernel/module.c-97-\tif (!chains) {\n--\narch/arm/kernel/vdso.c=169=static int __init vdso_init(void)\n--\narch/arm/kernel/vdso.c-181-\t/* Allocate the VDSO text pagelist */\narch/arm/kernel/vdso.c:182:\tvdso_text_pagelist = kzalloc_objs(struct page *, text_pages);\narch/arm/kernel/vdso.c-183-\tif (vdso_text_pagelist == NULL)\n--\narch/arm/mach-footbridge/dc21285.c=261=int __init dc21285_setup(int nr, struct pci_sys_data *sys)\n--\narch/arm/mach-footbridge/dc21285.c-264-\narch/arm/mach-footbridge/dc21285.c:265:\tres = kzalloc_objs(struct resource, 2);\narch/arm/mach-footbridge/dc21285.c-266-\tif (!res) {\n--\narch/arm/mach-omap1/mcbsp.c=292=static void omap_mcbsp_register_board_cfg(struct resource *res, int res_count,\n--\narch/arm/mach-omap1/mcbsp.c-296-\narch/arm/mach-omap1/mcbsp.c:297:\tomap_mcbsp_devices = kzalloc_objs(struct platform_device *, size);\narch/arm/mach-omap1/mcbsp.c-298-\tif (!omap_mcbsp_devices) {\n--\narch/arm/mach-omap2/omap_device.c=131=static int omap_device_build_from_dt(struct platform_device *pdev)\n--\narch/arm/mach-omap2/omap_device.c-158-\narch/arm/mach-omap2/omap_device.c:159:\thwmods = kzalloc_objs(struct omap_hwmod *, oh_cnt);\narch/arm/mach-omap2/omap_device.c-160-\tif (!hwmods) {\n--\narch/arm/mach-omap2/pm33xx-core.c=379=static int __init amx3_idle_init(struct device_node *cpu_node, int cpu)\n--\narch/arm/mach-omap2/pm33xx-core.c-412-\narch/arm/mach-omap2/pm33xx-core.c:413:\tidle_states = kzalloc_objs(*idle_states, state_count);\narch/arm/mach-omap2/pm33xx-core.c-414-\tif (!idle_states)\n--\narch/arm/mach-omap2/sr_device.c=30=static void __init sr_set_nvalues(struct omap_volt_data *volt_data,\n--\narch/arm/mach-omap2/sr_device.c-41-\narch/arm/mach-omap2/sr_device.c:42:\tnvalue_table = kzalloc_objs(*nvalue_table, count);\narch/arm/mach-omap2/sr_device.c-43-\tif (!nvalue_table)\n--\narch/arm/mach-versatile/spc.c=393=static int ve_spc_populate_opps(uint32_t cluster)\n--\narch/arm/mach-versatile/spc.c-397-\narch/arm/mach-versatile/spc.c:398:\topps = kzalloc_objs(*opps, MAX_OPPS);\narch/arm/mach-versatile/spc.c-399-\tif (!opps)\n--\narch/arm/xen/enlighten.c=316=int __init arch_xen_unpopulated_init(struct resource **res)\n--\narch/arm/xen/enlighten.c-343-\narch/arm/xen/enlighten.c:344:\tregs = kzalloc_objs(*regs, nr_reg);\narch/arm/xen/enlighten.c-345-\tif (!regs) {\n--\narch/arm64/kernel/vdso.c=68=static int __init __vdso_init(enum vdso_abi abi)\n--\narch/arm64/kernel/vdso.c-83-\narch/arm64/kernel/vdso.c:84:\tvdso_pagelist = kzalloc_objs(struct page *, vdso_info[abi].vdso_pages);\narch/arm64/kernel/vdso.c-85-\tif (vdso_pagelist == NULL)\n--\narch/arm64/kvm/vgic/vgic-init.c=208=static int kvm_vgic_dist_init(struct kvm *kvm, unsigned int nr_spis)\n--\narch/arm64/kvm/vgic/vgic-init.c-217-\tdist-\u003eactive_spis = (atomic_t)ATOMIC_INIT(0);\narch/arm64/kvm/vgic/vgic-init.c:218:\tdist-\u003espis = kzalloc_objs(struct vgic_irq, nr_spis, GFP_KERNEL_ACCOUNT);\narch/arm64/kvm/vgic/vgic-init.c-219-\tif (!dist-\u003espis)\n--\narch/arm64/kvm/vgic/vgic-init.c=320=static int vgic_allocate_private_irqs_locked(struct kvm_vcpu *vcpu, u32 type)\n--\narch/arm64/kvm/vgic/vgic-init.c-335-\narch/arm64/kvm/vgic/vgic-init.c:336:\tvgic_cpu-\u003eprivate_irqs = kzalloc_objs(struct vgic_irq,\narch/arm64/kvm/vgic/vgic-init.c-337-\t\t\t\t\t num_private_irqs,\n--\narch/arm64/kvm/vgic/vgic-irqfd.c=142=int kvm_vgic_setup_default_irq_routing(struct kvm *kvm)\n--\narch/arm64/kvm/vgic/vgic-irqfd.c-148-\narch/arm64/kvm/vgic/vgic-irqfd.c:149:\tentries = kzalloc_objs(*entries, nr, GFP_KERNEL_ACCOUNT);\narch/arm64/kvm/vgic/vgic-irqfd.c-150-\tif (!entries)\n--\narch/arm64/kvm/vgic/vgic-v4.c=242=int vgic_v4_init(struct kvm *kvm)\n--\narch/arm64/kvm/vgic/vgic-v4.c-258-\narch/arm64/kvm/vgic/vgic-v4.c:259:\tdist-\u003eits_vm.vpes = kzalloc_objs(*dist-\u003eits_vm.vpes, nr_vcpus,\narch/arm64/kvm/vgic/vgic-v4.c-260-\t\t\t\t\t GFP_KERNEL_ACCOUNT);\n--\narch/csky/kernel/vdso.c=17=static int __init vdso_init(void)\n--\narch/csky/kernel/vdso.c-22-\tvdso_pagelist =\narch/csky/kernel/vdso.c:23:\t\tkzalloc_objs(struct page *, vdso_pages);\narch/csky/kernel/vdso.c-24-\tif (unlikely(vdso_pagelist == NULL)) {\n--\narch/loongarch/kernel/vdso.c=45=static int __init init_vdso(void)\n--\narch/loongarch/kernel/vdso.c-55-\tvdso_info.code_mapping.pages =\narch/loongarch/kernel/vdso.c:56:\t\tkzalloc_objs(struct page *, vdso_info.size / PAGE_SIZE);\narch/loongarch/kernel/vdso.c-57-\n--\narch/loongarch/kvm/intc/pch_pic.c=413=static int kvm_setup_default_irq_routing(struct kvm *kvm)\n--\narch/loongarch/kvm/intc/pch_pic.c-418-\narch/loongarch/kvm/intc/pch_pic.c:419:\tentries = kzalloc_objs(*entries, nr);\narch/loongarch/kvm/intc/pch_pic.c-420-\tif (!entries)\n--\narch/mips/alchemy/common/clock.c=754=static int __init alchemy_clk_init_fgens(int ctype)\n--\narch/mips/alchemy/common/clock.c-777-\narch/mips/alchemy/common/clock.c:778:\ta = kzalloc_objs(*a, 6);\narch/mips/alchemy/common/clock.c-779-\tif (!a)\n--\narch/mips/alchemy/common/clock.c=960=static int __init alchemy_clk_setup_imux(int ctype)\n--\narch/mips/alchemy/common/clock.c-998-\narch/mips/alchemy/common/clock.c:999:\ta = kzalloc_objs(*a, 6);\narch/mips/alchemy/common/clock.c-1000-\tif (!a)\n--\narch/mips/alchemy/common/dbdma.c=1056=static int __init dbdma_setup(unsigned int irq, dbdev_tab_t *idtable)\n--\narch/mips/alchemy/common/dbdma.c-1059-\narch/mips/alchemy/common/dbdma.c:1060:\tdbdev_tab = kzalloc_objs(dbdev_tab_t, DBDEV_TAB_SIZE);\narch/mips/alchemy/common/dbdma.c-1061-\tif (!dbdev_tab)\n--\narch/mips/alchemy/common/platform.c=203=static int __init _new_usbres(struct resource **r, struct platform_device **d)\narch/mips/alchemy/common/platform.c-204-{\narch/mips/alchemy/common/platform.c:205:\t*r = kzalloc_objs(struct resource, 2);\narch/mips/alchemy/common/platform.c-206-\tif (!*r)\n--\narch/mips/alchemy/devboards/platform.c=70=int __init db1x_register_pcmcia_socket(phys_addr_t pcmcia_attr_start,\n--\narch/mips/alchemy/devboards/platform.c-91-\narch/mips/alchemy/devboards/platform.c:92:\tsr = kzalloc_objs(struct resource, cnt);\narch/mips/alchemy/devboards/platform.c-93-\tif (!sr)\n--\narch/mips/alchemy/devboards/platform.c=154=int __init db1x_register_norflash(unsigned long size, int width,\n--\narch/mips/alchemy/devboards/platform.c-166-\tret = -ENOMEM;\narch/mips/alchemy/devboards/platform.c:167:\tparts = kzalloc_objs(struct mtd_partition, 5);\narch/mips/alchemy/devboards/platform.c-168-\tif (!parts)\n--\narch/mips/bcm47xx/buttons.c=501=bcm47xx_buttons_add(const struct bcm47xx_gpio_key *buttons, int nbuttons)\n--\narch/mips/bcm47xx/buttons.c-525-\tconst struct software_node **node_group __free(kfree) =\narch/mips/bcm47xx/buttons.c:526:\t\tkzalloc_objs(*node_group, 1 + nbuttons + 1);\narch/mips/bcm47xx/buttons.c-527-\tif (!node_group)\n--\narch/mips/bcm47xx/buttons.c-531-\tstruct property_entry *props __free(kfree) =\narch/mips/bcm47xx/buttons.c:532:\t\tkzalloc_objs(*props, nbuttons * 3);\narch/mips/bcm47xx/buttons.c-533-\tif (!props)\n--\narch/mips/bcm47xx/buttons.c-537-\tstruct software_node *nodes __free(kfree) =\narch/mips/bcm47xx/buttons.c:538:\t\tkzalloc_objs(*nodes, 1 + nbuttons);\narch/mips/bcm47xx/buttons.c-539-\tif (!nodes)\n--\narch/mips/bcm47xx/buttons.c-542-\tstruct software_node_ref_args *ref_args __free(kfree) =\narch/mips/bcm47xx/buttons.c:543:\t\tkzalloc_objs(*ref_args, nbuttons);\narch/mips/bcm47xx/buttons.c-544-\tif (!ref_args)\n--\narch/mips/kernel/smp-cps.c=295=static void __init cps_prepare_cpus(unsigned int max_cpus)\n--\narch/mips/kernel/smp-cps.c-343-\tnclusters = mips_cps_numclusters();\narch/mips/kernel/smp-cps.c:344:\tmips_cps_cluster_bootcfg = kzalloc_objs(*mips_cps_cluster_bootcfg,\narch/mips/kernel/smp-cps.c-345-\t\t\t\t\t\tnclusters);\n--\narch/mips/kernel/smp-cps.c-354-\t\tncores = mips_cps_numcores(cl);\n\n"
} |
| 29/3 | 2026/10/06 13:36 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 35383 output: 20 thoughts: 19 Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person: **Diving into Kernel Memory Allocation with grepper** Alright, so I'm thinking I need to understand `kzalloc_objs` better. Time to get my hands dirty. My first instinct, as always, is to fire up `grepper`. I'll use it to pinpoint the exact definition of `kzalloc_objs` within the kernel source tree. Knowing the definition is the crucial first step. It'll give me the precise context, data structures involved, and hopefully some hints about its purpose and how it's used. From there, I can start forming some hypotheses about its functionality and interactions. I'm hoping `grepper` will deliver me to the goods quickly so I can begin a deeper analysis. |
| 30/3 | 2026/10/06 13:36 | tool | grepper |
0mArgs:{
"Expression": "#define kzalloc_objs"
}Results: {
"Output": "include/linux/slab.h=1073=void *kmalloc_nolock(size_t size, gfp_t gfp_flags, int node);\n--\ninclude/linux/slab.h-1156-\t__alloc_objs(kzalloc, default_gfp(__VA_ARGS__), typeof(P), 1)\ninclude/linux/slab.h:1157:#define kzalloc_objs(P, COUNT, ...) \\\ninclude/linux/slab.h-1158-\t__alloc_objs(kzalloc, default_gfp(__VA_ARGS__), typeof(P), COUNT)\n"
} |
| 31/3 | 2026/10/06 13:36 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 35757 output: 21 thoughts: 3018 Thoughts: **Analysis of Potential Slab Out-of-Bounds Read in Storage Element Reporting** Okay, let's break this down. My initial focus was on understanding how storage elements are reported and managed. The code uses `kzalloc_objs` to allocate memory for `struct blk_storage_element` structures, which is good practice for zero-initializing. I then investigated `disk_wait_for_se_mgmt_completion` and its interaction with `report_elements`, ensuring memory allocation happens correctly based on the number of elements. I then investigated `sd_zbc_report_storage_elements` to see how it handled the reporting, and looked at what might happen when the number of reported elements is zero. The critical observation came when reviewing the `zloop_report_elements` function. I considered a scenario where the number of elements requested by userspace (`rep.nr_elements`, for instance, 2) differs from the number of elements the kernel driver knows about (`zlo->nr_elements`, for instance, 10). The code correctly allocates memory for the user-specified number of elements, however, inside `zloop_report_elements` the number of elements (`*nr_elements`) is then updated to be 10. The copy operation then tries to copy 10 elements, but the buffer `elements` only has space for 2 elements. That leads to a read beyond the allocated buffer! `kzalloc_objs` is used, so the initial allocation is zeroed. That's a **slab out-of-bounds read**. The `copy_to_user` function attempts to copy data from the kernel-allocated `elements` buffer (2 elements) to the userspace buffer, however, the kernel driver writes 10 elements. This is a very clear example of a vulnerability. My initial reaction was that this is something that needs KMSAN. However, because KASAN immediately detects this issue, I can confidently state that KASAN should be used. |
| 32/3 | 2026/10/06 13:37 | tool | grepper |
0mArgs:{
"Expression": "disk_zone_get_cond"
}Results: {
"Output": "block/blk-zoned.c=251=static inline u8 disk_zone_get_state(struct gendisk *disk, sector_t sector)\n--\nblock/blk-zoned.c-266-\nblock/blk-zoned.c:267:static enum blk_zone_cond disk_zone_get_cond(struct gendisk *disk,\nblock/blk-zoned.c-268-\t\t\t\t\t sector_t sector)\n--\nblock/blk-zoned.c=282=static inline bool disk_zone_is_offline_or_readonly(struct gendisk *disk,\n--\nblock/blk-zoned.c-284-{\nblock/blk-zoned.c:285:\tenum blk_zone_cond cond = disk_zone_get_cond(disk, sector);\nblock/blk-zoned.c-286-\n--\nblock/blk-zoned.c=320=bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector)\n--\nblock/blk-zoned.c-326-\nblock/blk-zoned.c:327:\tcond = disk_zone_get_cond(bdev-\u003ebd_disk, sector);\nblock/blk-zoned.c-328-\treturn cond == BLK_ZONE_COND_OFFLINE;\n"
} |
| 33/3 | 2026/10/06 13:37 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 39546 output: 35 thoughts: 29 Thoughts: Okay, let's dive into this. Here's how I'd approach understanding `disk_zone_get_cond` in `block/blk-zoned.c`: **Dissecting `disk_zone_get_cond`** Alright, time to crack open the code and get familiar with `disk_zone_get_cond`. My initial goal is to understand precisely how this function determines the conditions of a zone on a zoned block device. I need to see how it interacts with the underlying data structures, what information it's relying on, and how it handles potential error scenarios. I'm thinking, is this function just reading some pre-existing zone metadata, or is it actively interacting with the device to determine the zone's status? Is it a simple state machine, or something more complex? I'm anticipating there will be some checks against sector numbers and zone descriptors. I also expect error conditions to be handled. I'll be looking for any clues regarding locking or concurrency control, if multiple threads could be accessing this at the same time. I'll start by making note of the return value and then trace its execution step by step. I'm also planning to cross-reference this function with other related code in the block layer, to get a holistic view of the zoned device management. Ultimately, I want to form a mental model of how this function fits into the overall zoned block device framework. |
| 34/3 | 2026/10/06 13:37 | tool | read-file |
0mArgs:{
"File": "block/blk-zoned.c",
"FirstLine": 250,
"LineCount": 40
}Results: {
"Contents": " 250:\t\n 251:\tstatic inline u8 disk_zone_get_state(struct gendisk *disk, sector_t sector)\n 252:\t{\n 253:\t\tunsigned int zno = disk_zone_no(disk, sector);\n 254:\t\tu8 *zones_state, zs;\n 255:\t\n 256:\t\trcu_read_lock();\n 257:\t\tzones_state = rcu_dereference(disk-\u003ezones_state);\n 258:\t\tif (likely(zones_state \u0026\u0026 zno \u003c disk-\u003enr_zones))\n 259:\t\t\tzs = zones_state[zno];\n 260:\t\telse\n 261:\t\t\tzs = BLK_ZFLAG_CONV;\n 262:\t\trcu_read_unlock();\n 263:\t\n 264:\t\treturn zs;\n 265:\t}\n 266:\t\n 267:\tstatic enum blk_zone_cond disk_zone_get_cond(struct gendisk *disk,\n 268:\t\t\t\t\t\t sector_t sector)\n 269:\t{\n 270:\t\tu8 zs = disk_zone_get_state(disk, sector);\n 271:\t\n 272:\t\treturn blk_zstate_to_zone_cond(zs);\n 273:\t}\n 274:\t\n 275:\tstatic inline bool\n 276:\tdisk_zone_cond_is_offline_or_readonly(enum blk_zone_cond cond)\n 277:\t{\n 278:\t\treturn cond == BLK_ZONE_COND_READONLY ||\n 279:\t\t\tcond == BLK_ZONE_COND_OFFLINE;\n 280:\t}\n 281:\t\n 282:\tstatic inline bool disk_zone_is_offline_or_readonly(struct gendisk *disk,\n 283:\t\t\t\t\t\t\t sector_t sector)\n 284:\t{\n 285:\t\tenum blk_zone_cond cond = disk_zone_get_cond(disk, sector);\n 286:\t\n 287:\t\treturn disk_zone_cond_is_offline_or_readonly(cond);\n 288:\t}\n 289:\t\n"
} |
| 35/3 | 2026/10/06 13:37 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 40500 output: 25 thoughts: 3150 Thoughts: **Analyzing Storage Element Reporting in Block Devices** Okay, here's the breakdown of my thought process as I've been dissecting this storage element reporting code: First, I needed to understand the underlying state reporting mechanism, which led me to the `disk_zone_get_cond` function. Then, I needed to understand how a specific function, `disk_wait_for_se_mgmt_completion`, was behaving with the storage elements. It was using `report_elements`, and that’s where things get interesting. The handling of `nr_elements = 0` during memory allocation with `kzalloc_objs` immediately raised a red flag. If zero elements are reported, then `kzalloc` might have allocated a small or zero-size memory location, and then that zero-sized memory will be used in subsequent operations. This then made me think if the driver had a specialized version, like `zloop_report_elements`, and if so, how it would behave when `nr_elements` is zero. Turns out, it handles that case fine since the loop will run 0 times and will not cause an issue. I also checked if the memory will be uninitialized. No, it is `kzalloc` so the memory is initialized. Then, I shifted my focus to `blkdev_report_storage_elements_ioctl`. The `copy_from_user` call caught my attention. I needed to ensure it completely succeeds, and what the size and alignment of that `struct blk_storage_elements_report` are, especially given that it includes a flexible array member (`elements[]`). I had to carefully evaluate the data structures, the `struct blk_storage_element`'s alignment and size (24 bytes, aligned to 8-bytes). The flexible array member of `struct blk_storage_elements_report` does not influence the size of the structure itself which remained 8 bytes, also with 8-byte alignment, because the flexible array member is *at the end* of the struct. Finally, I considered the implications of the user-space pointer arithmetic and alignment, particularly the `argp + sizeof(...)` used with `copy_to_user`. I kept in mind that `copy_to_user` handles unaligned pointers. Then I looked for SCSI-specific implementation, especially related to `sd_zbc_report_storage_elements` for SCSI/sd_zbc. The `buf = kzalloc(buf_size, GFP_KERNEL);` followed by `scsi_execute_cmd` made me wonder about partial buffer fills and what the SCSI device reports. In particular, I was interested in what happens if the returned SCSI data is smaller than the allocated buffer. I then examined how the number of descriptors is handled, looking at the data at offsets 0 and 4 in `buf`. I need to consult the SCSI spec to completely understand how those are filled and what it means if the SCSI command does not fill the buffer. |
| 36/3 | 2026/10/06 13:38 | tool | grepper |
0mArgs:{
"Expression": "SAI_GET_PHYSICAL_ELEMENT_STATUS"
}Results: {
"Output": "drivers/ata/libata-scsi.c=3767=static const struct ata_scsi_cmd ata_supported_cmds[] = {\n--\ndrivers/ata/libata-scsi.c-3829-\t\t.sa_valid = true,\ndrivers/ata/libata-scsi.c:3830:\t\t.sa = SAI_GET_PHYSICAL_ELEMENT_STATUS\ndrivers/ata/libata-scsi.c-3831-\t},\n--\ndrivers/ata/libata-scsi.c=3891=static bool ata_scsi_cmd_is_supported(struct ata_device *dev, u8 op, u16 sa,\n--\ndrivers/ata/libata-scsi.c-3928-\t\tswitch (sa) {\ndrivers/ata/libata-scsi.c:3929:\t\tcase SAI_GET_PHYSICAL_ELEMENT_STATUS:\ndrivers/ata/libata-scsi.c-3930-\t\tcase SAI_REMOVE_ELEMENT_AND_TRUNCATE:\n--\ndrivers/ata/libata-scsi.c=4966=static inline ata_xlat_func_t ata_get_xlat_func(struct ata_device *dev,\n--\ndrivers/ata/libata-scsi.c-5006-\t\tsa = cdb[1] \u0026 0x1f;\ndrivers/ata/libata-scsi.c:5007:\t\tif (sa == SAI_GET_PHYSICAL_ELEMENT_STATUS)\ndrivers/ata/libata-scsi.c-5008-\t\t\treturn ata_scsi_get_phys_element_status_xlat;\n--\ndrivers/scsi/sd_zbc.c=622=static int sd_zbc_report_storage_elements(struct gendisk *disk,\n--\ndrivers/scsi/sd_zbc.c-654-\tcmd[0] = SERVICE_ACTION_IN_16;\ndrivers/scsi/sd_zbc.c:655:\tcmd[1] = SAI_GET_PHYSICAL_ELEMENT_STATUS;\ndrivers/scsi/sd_zbc.c-656-\tput_unaligned_be32(buf_size, \u0026cmd[10]);\n--\ninclude/scsi/scsi_proto.h-131-#define SAI_GET_STREAM_STATUS 0x16\ninclude/scsi/scsi_proto.h:132:#define SAI_GET_PHYSICAL_ELEMENT_STATUS 0x17\ninclude/scsi/scsi_proto.h-133-#define SAI_REMOVE_ELEMENT_AND_TRUNCATE 0x18\n"
} |
| 37/3 | 2026/10/06 13:38 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44959 output: 29 thoughts: 35 Thoughts: Okay, let's dive into this. Here's how I'd approach summarizing the thought process, as if it were my own: **Analyzing `ata_scsi_get_phys_element_status_xlat` in `drivers/ata/libata-scsi.c`** Alright, I need to understand what's happening with the translation of physical element status within the ATA/SCSI interface. I'm focusing on `ata_scsi_get_phys_element_status_xlat` in `drivers/ata/libata-scsi.c`. This suggests that there's a specific conversion being done between ATA drive information and the SCSI command set, particularly related to the physical characteristics or status of the storage elements. It could be dealing with things like drive sectors, LUN mapping, or even more granular details related to the underlying ATA device. I'd expect to see a function that maps ATA-specific status codes or descriptors into corresponding SCSI interpretations. I'll need to figure out what information is being passed *into* this function, likely related to the ATA device's current state. Then I'll be interested in how that information is translated to a SCSI-compatible structure or format. It's probably mapping ATA error codes or status bits to SCSI sense keys, ASC/ASCQ codes, or other relevant SCSI-defined status indicators. I'll need to look for how this function fits into the overall command path and error handling within the libata SCSI layer. I'm imagining it's a critical component in ensuring that the higher-level SCSI layers and applications receive the correct interpretations of the ATA hardware’s status. Let's see what the code actually reveals. |
| 38/3 | 2026/10/06 13:38 | tool | grepper |
0mArgs:{
"Expression": "ata_scsi_get_phys_element_status_xlat"
}Results: {
"Output": "drivers/ata/libata-scsi.c=4786=static unsigned int\ndrivers/ata/libata-scsi.c:4787:ata_scsi_get_phys_element_status_xlat(struct ata_queued_cmd *qc)\ndrivers/ata/libata-scsi.c-4788-{\n--\ndrivers/ata/libata-scsi.c=4966=static inline ata_xlat_func_t ata_get_xlat_func(struct ata_device *dev,\n--\ndrivers/ata/libata-scsi.c-5007-\t\tif (sa == SAI_GET_PHYSICAL_ELEMENT_STATUS)\ndrivers/ata/libata-scsi.c:5008:\t\t\treturn ata_scsi_get_phys_element_status_xlat;\ndrivers/ata/libata-scsi.c-5009-\t\tif (sa == SAI_REMOVE_ELEMENT_AND_TRUNCATE)\n"
} |
| 39/3 | 2026/10/06 13:38 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 45630 output: 40 thoughts: 24 Thoughts: Okay, here's my interpretation of what's going on, thinking through this code: **Diving into `drivers/ata/libata-scsi.c` - Line 4786** Alright, I'm heading over to `drivers/ata/libata-scsi.c` and specifically focusing around line 4786. *Why*? Well, let's see. My mind's probably zeroing in on this because I suspect there's something interesting happening there related to SCSI command handling within the libata framework. Could be error handling, command queuing, or device state transitions. This line number likely points me toward a critical function, a conditional branch, or a data structure that's central to the ATA-to-SCSI translation process. I'm expecting to find something that helps to bridge the ATA command set with the SCSI subsystem - the interface in action. Gotta get the gist of what's transpiring there. Now to read the code and understand the context! |
| 40/3 | 2026/10/06 13:39 | tool | read-file |
0mArgs:{
"File": "drivers/ata/libata-scsi.c",
"FirstLine": 4786,
"LineCount": 50
}Results: {
"Contents": "4786:\tstatic unsigned int\n4787:\tata_scsi_get_phys_element_status_xlat(struct ata_queued_cmd *qc)\n4788:\t{\n4789:\t\tstruct scsi_cmnd *scmd = qc-\u003escsicmd;\n4790:\t\tconst u8 *cdb = scmd-\u003ecmnd;\n4791:\t\tstruct ata_device *dev = qc-\u003edev;\n4792:\t\tstruct ata_taskfile *tf = \u0026qc-\u003etf;\n4793:\t\tu32 starting_element, len;\n4794:\t\n4795:\t\t/* ATA_CMD_GET_PHYS_ELEMENT_STATUS is a DMA command. */\n4796:\t\tif (!(dev-\u003eflags \u0026 ATA_DFLAG_DEPOP) || !ata_dma_enabled(dev)) {\n4797:\t\t\tata_scsi_set_sense(dev, scmd, ILLEGAL_REQUEST, 0x20, 0x0);\n4798:\t\t\treturn 1;\n4799:\t\t}\n4800:\t\n4801:\t\tlen = get_unaligned_be32(\u0026cdb[10]) / ATA_SECT_SIZE;\n4802:\t\tif (!len || len \u003e U16_MAX) {\n4803:\t\t\tata_scsi_set_invalid_field(dev, scmd, 10, 0);\n4804:\t\t\treturn 1;\n4805:\t\t}\n4806:\t\n4807:\t\ttf-\u003eprotocol = ATA_PROT_DMA;\n4808:\t\ttf-\u003ecommand = ATA_CMD_GET_PHYS_ELEMENT_STATUS;\n4809:\t\ttf-\u003ehob_feature = cdb[14];\n4810:\t\ttf-\u003ehob_nsect = (len \u003e\u003e 8) \u0026 0xff;\n4811:\t\ttf-\u003ensect = len \u0026 0xff;\n4812:\t\n4813:\t\tstarting_element = get_unaligned_be32(\u0026cdb[6]);\n4814:\t\tif (starting_element) {\n4815:\t\t\ttf-\u003ehob_lbal = (starting_element \u003e\u003e 24) \u0026 0xff;\n4816:\t\t\ttf-\u003elbah = (starting_element \u003e\u003e 16) \u0026 0xff;\n4817:\t\t\ttf-\u003elbam = (starting_element \u003e\u003e 8) \u0026 0xff;\n4818:\t\t\ttf-\u003elbal = starting_element \u0026 0xff;\n4819:\t\t}\n4820:\t\ttf-\u003edevice = ATA_LBA;\n4821:\t\ttf-\u003eflags |= ATA_TFLAG_ISADDR | ATA_TFLAG_DEVICE | ATA_TFLAG_LBA48;\n4822:\t\n4823:\t\tata_qc_set_pc_nbytes(qc);\n4824:\t\n4825:\t\tqc-\u003eflags |= ATA_QCFLAG_RESULT_TF;\n4826:\t\tqc-\u003ecomplete_fn = ata_scsi_get_phys_element_status_complete;\n4827:\t\n4828:\t\treturn 0;\n4829:\t}\n4830:\t\n4831:\tstatic unsigned int\n4832:\tata_scsi_remove_element_and_truncate_xlat(struct ata_queued_cmd *qc)\n4833:\t{\n4834:\t\tstruct scsi_cmnd *scmd = qc-\u003escsicmd;\n4835:\t\tconst u8 *cdb = scmd-\u003ecmnd;\n"
} |
| 41/3 | 2026/10/06 13:39 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 46822 output: 35 thoughts: 2102 Thoughts: **Analysis of blk-storage-element Initialization and Potential Uninitialized Variable Use** Okay, let's break down these concerns about uninitialized memory access within this storage element code. First, the `sd_zbc_parse_storage_element` function seems clean; the explicit `memset` call guarantees that the `element` structure is zeroed before any fields are modified. I double-checked the structure layout, and there's no compiler-inserted padding, so this approach is safe. Moving on to `zloop`, I see `kzalloc_objs` being used to allocate and zero the `elements` array in both `zloop_create_storage_elements` and `blkdev_report_storage_elements_ioctl`. That's good practice. I'm now confident that the allocation and initialization of memory is handled correctly. Now, I scrutinized the `blkdev_report_storage_elements_ioctl` function. The `rep` struct, defined locally, seemed like a potential issue at first, since it is a struct on the stack. The `copy_from_user` call initializes `rep`. The code then correctly updates the `nr_elements` and the `copy_to_user` copies the updated values back. No risk here either. Similarly, in `blkdev_get_nr_storage_elements_ioctl`, `nr_elements` is properly initialized before being passed to `put_user`. Also fine. `element_id` is initialized through `get_user` in `blkdev_remove_storage_element_ioctl`. I scanned the `blkdev_restore_storage_elements_ioctl` as well, but it takes no user argument, so it is safe. Next, I checked `zloop_parse_options`. `opts->stor_elements` and `opts->element_id` are initialized by default and the options parsing is safe. Finally, I considered `zloop_prep_rq`. The potential issue lies in the access to `zlo->elements` within this function. The code checks `zlo->nr_elements`, which prevents a crash on unallocated elements. However, I need to ensure that `zone->wr_se_id` and `zone->rd_se_id` are *always* initialized if `zlo->nr_elements` is not 0 and the storage elements feature is enabled. Looking at `zloop_init_zone`, I see that `zone->wr_se_id` and `zone->rd_se_id` are *always* initialized when storage elements are used. The code calculates a `se_idx` based on the zone number, then sets both IDs. Also, looking at `zloop.c`, I believe that `zloop_init_zone` is called for every zone. No issue here. In summary, after my review, I am *very* confident that there are no uninitialized variable accesses within the code I reviewed, and that memory is allocated and initialized correctly in all the critical paths, with one exception that is safe (the allocation size in `zloop` might not equal the total elements requested), but that is not an uninitialized memory issue. |
| 42/3 | 2026/10/06 13:39 | tool | grepper |
0mArgs:{
"Expression": "zloop_init_zone",
"PathPrefixes": [
"drivers/block/zloop.c"
]
}Results: {
"Output": "drivers/block/zloop.c=1449=static int zloop_get_block_size(struct zloop_device *zlo,\n--\ndrivers/block/zloop.c-1491-\ndrivers/block/zloop.c:1492:static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,\ndrivers/block/zloop.c-1493-\t\t\t unsigned int zone_no, bool restore)\n--\ndrivers/block/zloop.c=1692=static int zloop_ctl_add(struct zloop_options *opts)\n--\ndrivers/block/zloop.c-1799-\t * If we already have zone files, we are restoring a device created by a\ndrivers/block/zloop.c:1800:\t * previous add operation. In this case, zloop_init_zone() will check\ndrivers/block/zloop.c-1801-\t * that the zone files are consistent with the zone configuration given.\n--\ndrivers/block/zloop.c-1804-\tfor (i = 0; i \u003c nr_zones; i++) {\ndrivers/block/zloop.c:1805:\t\tret = zloop_init_zone(zlo, opts, i, restore);\ndrivers/block/zloop.c-1806-\t\tif (ret)\n"
} |
| 43/3 | 2026/10/06 13:39 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 49990 output: 33 thoughts: 35 Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person: **Analyzing ZLO Allocation in the Context of zloop_ctl_add** Right, so `zlo` is allocated with `kvzalloc` within the `zloop_ctl_add` function. I need to refresh my memory on the exact allocation process. My focus is always on understanding the memory allocation methods and their implications. I need to quickly verify how this `zlo` allocation functions. Is there any possibility of memory leaks or potential memory corruption down the line in this specific module? Are there size constraints that need to be considered? This is a key entry point, so I need to check the exact `kvzalloc` usage to make sure that the system is safe. |
| 44/3 | 2026/10/06 13:39 | tool | grepper |
0mArgs:{
"Expression": "zlo = ",
"PathPrefixes": [
"drivers/block/zloop.c"
]
}Results: {
"Output": "drivers/block/zloop.c=201=static unsigned int rq_zone_no(struct request *rq)\ndrivers/block/zloop.c-202-{\ndrivers/block/zloop.c:203:\tstruct zloop_device *zlo = rq-\u003eq-\u003equeuedata;\ndrivers/block/zloop.c-204-\n--\ndrivers/block/zloop.c=602=static int zloop_do_rw(struct zloop_cmd *cmd)\n--\ndrivers/block/zloop.c-606-\tunsigned int nr_bvec = blk_rq_nr_bvec(rq);\ndrivers/block/zloop.c:607:\tstruct zloop_device *zlo = rq-\u003eq-\u003equeuedata;\ndrivers/block/zloop.c-608-\tstruct zloop_zone *zone = \u0026zlo-\u003ezones[rq_zone_no(rq)];\n--\ndrivers/block/zloop.c=655=static int zloop_seq_write_prep(struct zloop_cmd *cmd)\n--\ndrivers/block/zloop.c-657-\tstruct request *rq = blk_mq_rq_from_pdu(cmd);\ndrivers/block/zloop.c:658:\tstruct zloop_device *zlo = rq-\u003eq-\u003equeuedata;\ndrivers/block/zloop.c-659-\tunsigned int zone_no = rq_zone_no(rq);\n--\ndrivers/block/zloop.c=716=static void zloop_rw(struct zloop_cmd *cmd)\n--\ndrivers/block/zloop.c-718-\tstruct request *rq = blk_mq_rq_from_pdu(cmd);\ndrivers/block/zloop.c:719:\tstruct zloop_device *zlo = rq-\u003eq-\u003equeuedata;\ndrivers/block/zloop.c-720-\tunsigned int zone_no = rq_zone_no(rq);\n--\ndrivers/block/zloop.c=828=static void zloop_handle_cmd(struct zloop_cmd *cmd)\n--\ndrivers/block/zloop.c-830-\tstruct request *rq = blk_mq_rq_from_pdu(cmd);\ndrivers/block/zloop.c:831:\tstruct zloop_device *zlo = rq-\u003eq-\u003equeuedata;\ndrivers/block/zloop.c-832-\n--\ndrivers/block/zloop.c=906=static void zloop_complete_rq(struct request *rq)\n--\ndrivers/block/zloop.c-908-\tstruct zloop_cmd *cmd = blk_mq_rq_to_pdu(rq);\ndrivers/block/zloop.c:909:\tstruct zloop_device *zlo = rq-\u003eq-\u003equeuedata;\ndrivers/block/zloop.c-910-\tunsigned int zone_no = cmd-\u003esector \u003e\u003e zlo-\u003ezone_shift;\n--\ndrivers/block/zloop.c=1025=static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,\n--\ndrivers/block/zloop.c-1029-\tstruct zloop_cmd *cmd = blk_mq_rq_to_pdu(rq);\ndrivers/block/zloop.c:1030:\tstruct zloop_device *zlo = rq-\u003eq-\u003equeuedata;\ndrivers/block/zloop.c-1031-\n--\ndrivers/block/zloop.c=1061=static int zloop_open(struct gendisk *disk, blk_mode_t mode)\ndrivers/block/zloop.c-1062-{\ndrivers/block/zloop.c:1063:\tstruct zloop_device *zlo = disk-\u003eprivate_data;\ndrivers/block/zloop.c-1064-\tint ret;\n--\ndrivers/block/zloop.c=1076=static int zloop_report_zones(struct gendisk *disk, sector_t sector,\n--\ndrivers/block/zloop.c-1078-{\ndrivers/block/zloop.c:1079:\tstruct zloop_device *zlo = disk-\u003eprivate_data;\ndrivers/block/zloop.c-1080-\tstruct blk_zone blkz = {};\n--\ndrivers/block/zloop.c=1127=static int zloop_report_elements(struct gendisk *disk,\n--\ndrivers/block/zloop.c-1130-{\ndrivers/block/zloop.c:1131:\tstruct zloop_device *zlo = disk-\u003eprivate_data;\ndrivers/block/zloop.c-1132-\tunsigned int nr_report = *nr_elements;\n--\ndrivers/block/zloop.c=1188=static void zloop_remove_element_work(struct work_struct *work)\ndrivers/block/zloop.c-1189-{\ndrivers/block/zloop.c:1190:\tstruct zloop_device *zlo = container_of(work, struct zloop_device,\ndrivers/block/zloop.c-1191-\t\t\t\t\t\tremove_element_work.work);\n--\ndrivers/block/zloop.c=1246=static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)\ndrivers/block/zloop.c-1247-{\ndrivers/block/zloop.c:1248:\tstruct zloop_device *zlo = disk-\u003eprivate_data;\ndrivers/block/zloop.c-1249-\tstruct blk_storage_element *se, *paired_se = NULL;\n--\ndrivers/block/zloop.c=1309=static void zloop_restore_elements_work(struct work_struct *work)\ndrivers/block/zloop.c-1310-{\ndrivers/block/zloop.c:1311:\tstruct zloop_device *zlo = container_of(work, struct zloop_device,\ndrivers/block/zloop.c-1312-\t\t\t\t\t\trestore_elements_work.work);\n--\ndrivers/block/zloop.c=1349=static int zloop_restore_elements(struct gendisk *disk)\ndrivers/block/zloop.c-1350-{\ndrivers/block/zloop.c:1351:\tstruct zloop_device *zlo = disk-\u003eprivate_data;\ndrivers/block/zloop.c-1352-\tstruct blk_storage_element *se;\n--\ndrivers/block/zloop.c=1390=static void zloop_free_disk(struct gendisk *disk)\ndrivers/block/zloop.c-1391-{\ndrivers/block/zloop.c:1392:\tstruct zloop_device *zlo = disk-\u003eprivate_data;\ndrivers/block/zloop.c-1393-\tunsigned int i;\n--\ndrivers/block/zloop.c=1692=static int zloop_ctl_add(struct zloop_options *opts)\n--\ndrivers/block/zloop.c-1719-\ndrivers/block/zloop.c:1720:\tzlo = kvzalloc_flex(*zlo, zones, nr_zones);\ndrivers/block/zloop.c-1721-\tif (!zlo) {\n--\ndrivers/block/zloop.c=1944=static int zloop_ctl_remove(struct zloop_options *opts)\n--\ndrivers/block/zloop.c-1962-\ndrivers/block/zloop.c:1963:\tzlo = idr_find(\u0026zloop_index_idr, opts-\u003eid);\ndrivers/block/zloop.c-1964-\tif (!zlo || zlo-\u003estate == Zlo_creating) {\n--\ndrivers/block/zloop.c=1991=static int zloop_ctl_degrade_element(struct zloop_options *opts)\n--\ndrivers/block/zloop.c-2007-\ndrivers/block/zloop.c:2008:\tzlo = idr_find(\u0026zloop_index_idr, opts-\u003eid);\ndrivers/block/zloop.c-2009-\tif (!zlo || zlo-\u003estate == Zlo_creating)\n"
} |
| 45/3 | 2026/10/06 13:39 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 52303 output: 32 thoughts: 528 Thoughts: **Analyzing Zoned Block Device I/O Handling** Okay, so I'm looking at this code and trying to trace the path of I/O operations, specifically in the context of zoned block devices. First, `zlo` gets allocated, which means `kvzalloc_flex` is being used, and crucially, it's zeroed out during allocation. That's a good starting point. Now, the interesting part is how ioctls are handled. My focus is on the `block/ioctl.c` file. I see the dispatch for several ioctl commands: `BLKGETNRSTORELEMS`, `BLKREPORTSTORELEMS`, `BLKREMOVESTORELEM`, and `BLKRESTORESTORELEMS`. They all funnel into `blkdev_zone_storage_elements_ioctl`. Next I have to check `blkdev_zone_storage_elements_ioctl`. Makes sense, there's a check for zoned block devices and storage element operations being supported and a switch statement dispatching to different functions, depending on the particular ioctl. That looks standard. Hold on, I see something in `submit_bio_noacct`! There's this check: `if (bdev_is_zoned(bdev) && bdev_zone_is_offline(bdev, bio->bi_iter.bi_sector)) goto end_io;`. That's intriguing. A read operation might get short-circuited if the zone is offline. This jumps to `end_io`. I need to figure out what happens when we `goto end_io`. I'll go check `block/blk-core.c` around line 866. |
| 46/3 | 2026/10/06 13:40 | tool | grepper |
0mArgs:{
"Expression": "end_io:",
"PathPrefixes": [
"block/blk-core.c"
]
}Results: {
"Output": "block/blk-core.c=817=void submit_bio_noacct(struct bio *bio)\n--\nblock/blk-core.c-935-\tstatus = BLK_STS_NOTSUPP;\nblock/blk-core.c:936:end_io:\nblock/blk-core.c-937-\tbio_endio_status(bio, status);\n"
} |
| 47/3 | 2026/10/06 13:40 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 53329 output: 35 thoughts: 26 Thoughts: Here's my take on the kernel code in `block/blk-core.c`, specifically around those two line numbers. I'll focus on what's going through my mind as I'm analyzing it: **Diving into `blk-core.c` - Lines 866 and 936** Alright, let's zoom in on `block/blk-core.c` around line 866 and line 936. I know this file; it's the heart of the block layer, so I need to understand what's happening in those sections. * **Line 866:** Okay, I'm at `blk_rq_prep_clone()`. Hmm, this function. It looks like it handles prepping a request for cloning. From what I recall, cloning, in this context, is probably about duplicating an existing request for a multi-queue block device. It's about efficiently handling the same work across different queues or perhaps for different, but related, purposes. I have to see how it's handled; are we dealing with a shared request data structure or creating a completely independent copy? I'm checking to see how this request is initialized, how fields are copied, and whether there are any memory allocation concerns. I'm looking for potential race conditions or locking issues that could arise when copying and modifying a shared request. Is there any refcounting? What exactly is being "cloned" here, and is the original request properly tracked. What are the calling contexts and implications of failure? I must see what happens with request metadata during the clone operation. * **Line 936:** Now, let's see, that's `blk_rq_unprep_clone()`. Okay, this is the cleanup function that goes hand in hand with the earlier prep. I'm immediately concerned with ensuring that the counterpart to the earlier clone isn't leaving resources hanging. Is it properly freeing any allocated memory, decrementing reference counts, or signaling the completion of the cloned request appropriately? I need to analyze how data is handled when the work is complete. The function likely reverses what `blk_rq_prep_clone()` does. I'm checking for memory leaks, double-frees, or any errors that might happen when you undo the clone operation. Are there any assumptions made about the state of the original request? Is there any risk of corruption if a clone operation fails midway through? The symmetry of this prep/unprep pair will determine if this code is safe. |
| 48/3 | 2026/10/06 13:40 | tool | read-file |
0mArgs:{
"File": "block/blk-core.c",
"FirstLine": 850,
"LineCount": 40
}Results: {
"Contents": " 850:\t\t/*\n 851:\t\t * Filter flush bio's early so that bio based drivers without flush\n 852:\t\t * support don't have to worry about them.\n 853:\t\t */\n 854:\t\tif (op_is_flush(bio-\u003ebi_opf)) {\n 855:\t\t\tif (WARN_ON_ONCE(bio_op(bio) != REQ_OP_WRITE \u0026\u0026\n 856:\t\t\t\t\t bio_op(bio) != REQ_OP_ZONE_APPEND))\n 857:\t\t\t\tgoto end_io;\n 858:\t\t\tif (!bdev_write_cache(bdev)) {\n 859:\t\t\t\tbio-\u003ebi_opf \u0026= ~(REQ_PREFLUSH | REQ_FUA);\n 860:\t\t\t\tif (!bio_sectors(bio)) {\n 861:\t\t\t\t\tstatus = BLK_STS_OK;\n 862:\t\t\t\t\tgoto end_io;\n 863:\t\t\t\t}\n 864:\t\t\t}\n 865:\t\t}\n 866:\t\n 867:\t\tswitch (bio_op(bio)) {\n 868:\t\tcase REQ_OP_READ:\n 869:\t\t\tif (bdev_is_zoned(bdev) \u0026\u0026\n 870:\t\t\t bdev_zone_is_offline(bdev, bio-\u003ebi_iter.bi_sector))\n 871:\t\t\t\tgoto end_io;\n 872:\t\t\tbreak;\n 873:\t\tcase REQ_OP_WRITE:\n 874:\t\t\tif (bio-\u003ebi_opf \u0026 REQ_ATOMIC) {\n 875:\t\t\t\tstatus = blk_validate_atomic_write_op_size(q, bio);\n 876:\t\t\t\tif (status != BLK_STS_OK)\n 877:\t\t\t\t\tgoto end_io;\n 878:\t\t\t}\n 879:\t\t\tbreak;\n 880:\t\tcase REQ_OP_FLUSH:\n 881:\t\t\t/*\n 882:\t\t\t * REQ_OP_FLUSH can't be submitted through bios, it is only\n 883:\t\t\t * synthetized in struct request by the flush state machine.\n 884:\t\t\t */\n 885:\t\t\tgoto not_supported;\n 886:\t\tcase REQ_OP_DISCARD:\n 887:\t\t\tif (!bdev_max_discard_sectors(bdev))\n 888:\t\t\t\tgoto not_supported;\n 889:\t\t\tbreak;\n"
} |
| 49/3 | 2026/10/06 13:40 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 54556 output: 35 thoughts: 41 Thoughts: Here's my analysis of the provided thought process, tailored for an expert audience: **Investigating `status` Initialization in `submit_bio_noacct`** Okay, I'm digging into `submit_bio_noacct`, specifically the initialization of the `status` variable. My gut says there might be something subtle going on here, something that could be a source of a potential issue down the line, so I need to be precise. Where is it actually *set*? I'm betting that the initialization is what's important here. I'm heading straight to `block/blk-core.c`, lines 817-850. Let's see if there is any gotchas in there. |
| 50/3 | 2026/10/06 13:40 | tool | read-file |
0mArgs:{
"File": "block/blk-core.c",
"FirstLine": 817,
"LineCount": 35
}Results: {
"Contents": " 817:\tvoid submit_bio_noacct(struct bio *bio)\n 818:\t{\n 819:\t\tstruct block_device *bdev = bio-\u003ebi_bdev;\n 820:\t\tstruct request_queue *q = bdev_get_queue(bdev);\n 821:\t\tblk_status_t status = BLK_STS_IOERR;\n 822:\t\n 823:\t\tmight_sleep();\n 824:\t\n 825:\t\t/*\n 826:\t\t * For a REQ_NOWAIT based request, return -EOPNOTSUPP\n 827:\t\t * if queue does not support NOWAIT.\n 828:\t\t */\n 829:\t\tif ((bio-\u003ebi_opf \u0026 REQ_NOWAIT) \u0026\u0026 !bdev_nowait(bdev))\n 830:\t\t\tgoto not_supported;\n 831:\t\n 832:\t\tif (bio_has_crypt_ctx(bio)) {\n 833:\t\t\tif (WARN_ON_ONCE(!bio_has_data(bio)))\n 834:\t\t\t\tgoto end_io;\n 835:\t\t\tif (!blk_crypto_supported(bio))\n 836:\t\t\t\tgoto not_supported;\n 837:\t\t}\n 838:\t\n 839:\t\tif (should_fail_bio(bio))\n 840:\t\t\tgoto end_io;\n 841:\t\tbio_check_ro(bio);\n 842:\t\tif (!bio_flagged(bio, BIO_REMAPPED)) {\n 843:\t\t\tif (unlikely(bio_check_eod(bio)))\n 844:\t\t\t\tgoto end_io;\n 845:\t\t\tif (bdev_is_partition(bdev) \u0026\u0026\n 846:\t\t\t unlikely(blk_partition_remap(bio)))\n 847:\t\t\t\tgoto end_io;\n 848:\t\t}\n 849:\t\n 850:\t\t/*\n 851:\t\t * Filter flush bio's early so that bio based drivers without flush\n"
} |
| 51/3 | 2026/10/06 13:40 | llm | kmsan-evaluator |
2mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 55332 output: 40 thoughts: 5437 Thoughts: **Analysis of Storage Element Management and Zone Operations** Okay, let's break this down systematically. My initial assessment of `status = BLK_STS_IOERR` looks sound; the `goto end_io;` seems correct, and the initialization checks out. The focus now is to meticulously examine the code changes, especially within `sd_zbc.c` and its interaction with the Zoned Block Device (ZBD) loop driver. The `sd_zbc_parse_storage_element` function appears robust. It correctly handles the parsing of storage element descriptors, initializing all fields of the `element` structure, either through `memset` or explicit assignment based on the descriptor's content. The health status interpretation looks well-defined, mapping SCSI health codes to appropriate `BLK_SE_STS_*` values. I am concerned with the usage of unaligned accessors, but they should be alright. In `sd_zbc_report_storage_elements`, `buf` is allocated with `kzalloc`. The allocation of `elements` using `kzalloc_objs` is also correct, meaning both buffers are zeroed on creation. This should prevent uninitialized memory reads. The subsequent iteration through the descriptors and the population of the `elements` array seems logically sound. I will need to check what calls this function. Ah, the `zloop_report_elements` function is where it gets interesting. I remember the slab-out-of-bounds bug from before. The core issue is this potential inconsistency: the driver *calculates* the `nr_elements` based on what the user-supplied buffer can hold, but then *reports* `nr_elements` (potentially a larger number) in `zloop_report_elements`. In other words, if the user provides a small buffer, the driver might report a larger number of elements and then try to copy data *beyond* the allocated size. KASAN is absolutely my friend here, because any attempted read past the `elements` buffer allocated by the ioctl will trigger a slab-out-of-bounds read, which is a classic KASAN issue. The prompt is also quite clear: if KASAN can catch the bug, it is not worth running a KMSAN session. But what about the opposite scenario? Could there be any info leaks? `elements` is allocated via `kzalloc_objs`, ensuring zeroed memory. The content of `zlo->elements` and `buf` is also zeroed. `rep` is a stack variable that is filled by `copy_from_user` and is then overwritten by `rep.nr_elements = nr_elements;`. The `elements` buffer is filled, and the content is then copied to user space. I need to confirm there is no use of uninitialized memory in `zloop_report_elements`. Moving on to `disk_wait_for_se_mgmt_completion`, allocation and usage of `elements` using `kzalloc_objs` and the subsequent loop to check status seem safe, as they are fully contained within the kernel and therefore the elements are initialized. `bdev_zone_is_offline` checks a zone state. `disk->zones_state` is allocated during zone initialization. So there shouldn't be any uninitialized memory in this case. I need to check `zloop_remove_element` and `zloop_ctl_degrade_element` to make sure there are no out-of-bounds array access. The code explicitly checks to see if `element_id == 0`, and then returns an error. This prevents the out-of-bounds write. The code in `zloop_prep_rq` checks that zone write/read element ids are non-zero. The allocation of the elements is fine. I've reviewed the ioctl definitions. There doesn't appear to be an issue here. The check for `BLKREPORTSTORELEMS` is important. The code calls `copy_to_user` to copy the `elements` buffer to user space. We need to make sure that the length provided to `copy_to_user` is the correct length. `copy_to_user` is going to read from `elements` which is zeroed, and it can only write a number of elements based on the number specified by the user. Nothing uninitialized in `zloop_ctl_show`. The log messages in `zloop_ctl_degrade_element` use initialized variables. No issues there. Now let's check `sd_zbc.c`: Specifically `sd_zbc_check_modify_zones` and its use in `sd_zbc_read_zones`. The `buf` is a local variable, and is used to communicate with the disk. |
| 52/3 | 2026/10/06 13:42 | tool | grepper |
0mArgs:{
"Expression": "sd_zbc_read_zones",
"PathPrefixes": [
"drivers/scsi/sd_zbc.c"
]
}Results: {
"Output": "drivers/scsi/sd_zbc.c=789=int sd_zbc_revalidate_zones(struct scsi_disk *sdkp)\n--\ndrivers/scsi/sd_zbc.c-833-/**\ndrivers/scsi/sd_zbc.c:834: * sd_zbc_read_zones - Read zone information and update the request queue\ndrivers/scsi/sd_zbc.c-835- * @sdkp: SCSI disk pointer.\n--\ndrivers/scsi/sd_zbc.c-842- */\ndrivers/scsi/sd_zbc.c:843:int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim,\ndrivers/scsi/sd_zbc.c-844-\t\tu8 buf[SD_BUF_SIZE])\n"
} |
| 53/3 | 2026/10/06 13:42 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 62041 output: 576 thoughts: 3366 Thoughts: **Analysis of Potential KMSAN Needs** Okay, so I'm diving into this code, specifically looking at where we might need KMSAN. The core here seems to be around how we handle storage elements, particularly for these zoned block devices. We're passing a `buf` around to `sd_zbc_read_zones` and using it as scratch space in `scsi_report_opcode`. The key functions are `sd_zbc_report_storage_elements`, `sd_zbc_remove_storage_element`, and `sd_zbc_restore_storage_elements`. My thought process is to meticulously check each of the KMSAN criteria one by one. First, let's look at `sd_zbc_report_storage_elements`. This function uses `kzalloc` to allocate `buf` and then zeroes it. We also have `nr_elements`, which acts as an input/output parameter. I'm carefully examining the bounds, especially around how `desc` is incremented within the loop. The calculations for `buf_size` and the loop conditions seem sound; there's a guard to make sure we don't overrun `buf_size` or read beyond the allocated memory. Even if there's a device returning fewer bytes, `kzalloc` guarantees that the unreturned bytes are initialized to zero. I move onto `sd_zbc_remove_storage_element` and `sd_zbc_restore_storage_elements`. There's nothing uninitialized in these. Now, let's go through the KMSAN criteria one by one. * **Criterion 1:** Are there kernel structures sent to user space where fields or padding might be uninitialized? I'm checking `nr_elements`, `blk_storage_elements_report`, and `blk_storage_element` carefully. For `nr_elements`, that is fine. The `blk_storage_elements_report` struct has no padding. The array `elements` passed to user space is either `kzalloc`'ed to zero or zeroed via `memset`, so it can not have uninitialized fields. I'm also remembering that the `zloop_report_elements` routine might report an invalid number of elements. But this is an out-of-bounds read, and KASAN is built to find this kind of SLAB violation. No uninitialized data appears to leak. * **Criterion 2:** Any conditional logic dependent on uninitialized variables? I'm running through all the functions like `bdev_zone_is_offline`, `disk_wait_for_se_mgmt_completion`, `zloop_report_elements`, `zloop_prep_rq`, etc. It looks like every variable used in a conditional check is initialized before it's used. * **Criterion 3:** Any allocation of complex data structures where uninitialized fields could be read? All the potentially complex structures (`zlo->elements`, `elements`, `buf`) are allocated with `kzalloc` or `kzalloc_objs`, which guarantees they start zeroed. * **Criterion 4:** Any issues with bounds checks or logic that allows access to uninitialized bytes of buffers? I'm double-checking the bounds calculations and the loops. I still think the earlier analysis is correct. Since I don't see any of the classic KMSAN "red flags," the conclusion is that `NeedsKMSAN` should be `false`. This is about KASAN (SLAB OOB read) not KMSAN. The code uses `kzalloc` consistently and zero-initializes the fields before copying it to user space. It is a SLAB OOB read, not a KMSAN issue of uninitialized memory. |
| 54/3 | 2026/10/06 13:43 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch series introduces storage element management for zoned block devices across the block core, the SCSI ZBC driver (sd_zbc), and the zoned loop driver (zloop), adding ioctls BLKGETNRSTORELEMS, BLKREPORTSTORELEMS, BLKREMOVESTORELEM, and BLKRESTORESTORELEMS.\n\nRegarding KMSAN vs KASAN applicability:\n1. Heap allocations: All dynamically allocated structures, including the storage element arrays (e.g. `zlo-\u003eelements`, `elements` in `disk_wait_for_se_mgmt_completion` and `blkdev_report_storage_elements_ioctl`) and SCSI buffers (`buf` in `sd_zbc_report_storage_elements`), are allocated using `kzalloc_objs()` or `kzalloc()`, ensuring all bytes are zero-initialized upon allocation.\n2. Kernel-to-user data transfers:\n - In `blkdev_get_nr_storage_elements_ioctl()`, `nr_elements` is a scalar unsigned int initialized to 0 before being passed to `put_user()`.\n - In `blkdev_report_storage_elements_ioctl()`, `struct blk_storage_elements_report` consists of two `__u32` fields (8 bytes total, aligned to 8 bytes, zero padding) and is completely initialized from userspace via `copy_from_user()` before the count is updated and written back.\n - `struct blk_storage_element` has explicit members summing to exactly 24 bytes (4 + 4 + 8 + 1 + 1 + 1 + 5 bytes) with natural 8-byte alignment, leaving zero internal or tail padding bytes. Moreover, in `sd_zbc_parse_storage_element()`, each element descriptor is explicitly zeroed with `memset()` before parsing.\n3. Out-of-bounds risk: In `zloop_report_elements()`, `*nr_elements` is assigned `zlo-\u003enr_elements`, which in `blkdev_report_storage_elements_ioctl()` could cause `copy_to_user()` to copy more elements than the allocated `nr_elements` buffer if the caller requested fewer elements. However, this is a slab-out-of-bounds heap read, which is directly detectible by KASAN rather than an uninitialized memory defect.\n\nBecause all touched structures and buffers are zero-initialized and no uninitialized memory usage or info-leaks are introduced, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch series introduces storage element management for zoned block devices across the block core, the SCSI ZBC driver (sd_zbc), and the zoned loop driver (zloop), adding ioctls BLKGETNRSTORELEMS, BLKREPORTSTORELEMS, BLKREMOVESTORELEM, and BLKRESTORESTORELEMS.\n\nRegarding KMSAN vs KASAN applicability:\n1. Heap allocations: All dynamically allocated structures, including the storage element arrays (e.g. `zlo-\u003eelements`, `elements` in `disk_wait_for_se_mgmt_completion` and `blkdev_report_storage_elements_ioctl`) and SCSI buffers (`buf` in `sd_zbc_report_storage_elements`), are allocated using `kzalloc_objs()` or `kzalloc()`, ensuring all bytes are zero-initialized upon allocation.\n2. Kernel-to-user data transfers:\n - In `blkdev_get_nr_storage_elements_ioctl()`, `nr_elements` is a scalar unsigned int initialized to 0 before being passed to `put_user()`.\n - In `blkdev_report_storage_elements_ioctl()`, `struct blk_storage_elements_report` consists of two `__u32` fields (8 bytes total, aligned to 8 bytes, zero padding) and is completely initialized from userspace via `copy_from_user()` before the count is updated and written back.\n - `struct blk_storage_element` has explicit members summing to exactly 24 bytes (4 + 4 + 8 + 1 + 1 + 1 + 5 bytes) with natural 8-byte alignment, leaving zero internal or tail padding bytes. Moreover, in `sd_zbc_parse_storage_element()`, each element descriptor is explicitly zeroed with `memset()` before parsing.\n3. Out-of-bounds risk: In `zloop_report_elements()`, `*nr_elements` is assigned `zlo-\u003enr_elements`, which in `blkdev_report_storage_elements_ioctl()` could cause `copy_to_user()` to copy more elements than the allocated `nr_elements` buffer if the caller requested fewer elements. However, this is a slab-out-of-bounds heap read, which is directly detectible by KASAN rather than an uninitialized memory defect.\n\nBecause all touched structures and buffers are zero-initialized and no uninitialized memory usage or info-leaks are introduced, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|