| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/10/06 21:22 | flow | patch-triage |
11mArgs:null Results: {
"EnableConfigs": [
"NETFS_SUPPORT",
"FSCACHE",
"CACHEFILES",
"AFS_FS",
"NETFS_PGPRIV2"
],
"FocusSymbols": [
"bvecq_buffer_init",
"bvecq_buffer_append",
"bvecq_pos_advance",
"bvecq_zero",
"bvecq_slice",
"bvecq_load_from_ra",
"netfs_extract_iter",
"netfs_readahead",
"netfs_writepages",
"netfs_writeback_single",
"netfs_read_single",
"netfs_unbuffered_read_iter_locked",
"netfs_unbuffered_write_iter_locked",
"netfs_retry_reads",
"netfs_retry_writes",
"netfs_pgpriv2_unlock_copied_folios",
"afs_symlink_writepages"
],
"KMSANReasoning": "The patch series refactors netfs and related filesystem code (AFS, cachefiles) by replacing the legacy rolling_buffer abstraction with segmented bio_vec queues (bvecq) and cursor tracking (bvecq_pos: load_cursor, dispatch_cursor, collect_cursor).\n\nKMSAN vs KASAN Applicability Analysis:\n1. No uninitialized struct/stack/heap usage: All new and modified structures are cleanly initialized before use. Specifically, netfs_io_request and netfs_io_subrequest are fully zeroed via memset() in netfs_alloc_request() and netfs_alloc_subrequest(); new stack cursor variables (such as dispatch_cursor) are zero-initialized with '= {}' or bvecq_pos_set(); bvecq structures are allocated and zero-initialized via C99 designated initializers in bvecq_alloc_one(); and readahead_control._nr_folios is initialized to zero upon allocation and appropriately tracked.\n2. No kernel-to-user info-leaks: None of the modified data structures are copied or exposed to user space (no copy_to_user, put_user, netlink skb, ioctl, or sockopt changes). Buffer contents transferred to user space in DIO reads are filled by network/cache reads or explicitly zeroed (via bvecq_zero, iov_iter_zero, or folio_zero_segments for gaps/unwritten ranges).\n3. Types of bugs exposed: The refactoring involves queue manipulation, cursor advancement, slot indexing, and reference counting (bvecq_get/bvecq_put). The potential defects here include out-of-bounds accesses on bvec array slots, NULL pointer dereferences on cursor exhaustion, and use-after-free or refcount issues on bvecq nodes. All of these error classes are fully detected by standard debugging tools (KASAN, refcount_t sanity checks, and slab poisoning).\n\nConsequently, the patch series does not introduce uninitialized memory read risks or info-leak vulnerabilities that would require a dedicated KMSAN fuzzing session.",
"NeedsKMSAN": false,
"Reasoning": "The patch replaces the rolling_buffer implementation in netfs with a cursor-based bio_vec queue (bvecq_pos/bvecq) mechanism across read, write, direct I/O, readahead, retry, and writeback paths. It introduces new buffer-management primitives (bvecq_slice, bvecq_pos_advance, bvecq_zero, bvecq_load_from_ra, netfs_extract_iter), refactors I/O dispatch and folio unlock accounting, and updates network filesystems (AFS, Cachefiles) and mm/readahead folio tracking. These changes are in core reachable kernel filesystem paths and warrant fuzzing to identify potential regressions, buffer overflows, or invariant violations.",
"WorthFuzzing": true
} |
| 1/1 | 2026/10/06 21:22 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit baca3ab6cbbda96e6e631014dc876f11c374596c\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Tue Oct 6 21:22:55 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/fs/afs/dir.c b/fs/afs/dir.c\nindex 5fac02d2d2814..39cba3f37ecd4 100644\n--- a/fs/afs/dir.c\n+++ b/fs/afs/dir.c\n@@ -2229,8 +2229,9 @@ static int afs_dir_writepages(struct address_space *mapping,\n \n \tif (test_bit(AFS_VNODE_DIR_VALID, \u0026dvnode-\u003eflags)) {\n \t\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0,\n-\t\t\t\t i_size_read(\u0026dvnode-\u003enetfs.inode));\n-\t\tret = netfs_writeback_single(mapping, wbc, \u0026iter);\n+\t\t\t\t dvnode-\u003edirectory_size);\n+\t\tret = netfs_writeback_single(mapping, wbc, \u0026iter,\n+\t\t\t\t\t i_size_read(\u0026dvnode-\u003enetfs.inode));\n \t\tif (ret == 1)\n \t\t\tret = 0; /* Skipped write due to lock conflict. */\n \t}\ndiff --git a/fs/afs/symlink.c b/fs/afs/symlink.c\nindex 9a611efe6b264..ae03ceff42b85 100644\n--- a/fs/afs/symlink.c\n+++ b/fs/afs/symlink.c\n@@ -248,9 +248,9 @@ int afs_symlink_writepages(struct address_space *mapping,\n \n \tif (vnode-\u003edirectory \u0026\u0026\n \t atomic64_read(\u0026vnode-\u003ecb_expires_at) != AFS_NO_CB_PROMISE) {\n-\t\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, vnode-\u003edirectory, 0, 0,\n-\t\t\t\t i_size_read(\u0026vnode-\u003enetfs.inode));\n-\t\tret = netfs_writeback_single(mapping, wbc, \u0026iter);\n+\t\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, vnode-\u003edirectory, 0, 0, PAGE_SIZE);\n+\t\tret = netfs_writeback_single(mapping, wbc, \u0026iter,\n+\t\t\t\t\t i_size_read(\u0026vnode-\u003enetfs.inode));\n \t}\n \n \tif (ret == 0) {\ndiff --git a/fs/cachefiles/io.c b/fs/cachefiles/io.c\nindex d05059822288c..788439ea6e1c8 100644\n--- a/fs/cachefiles/io.c\n+++ b/fs/cachefiles/io.c\n@@ -546,7 +546,6 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq)\n \tstruct netfs_cache_resources *cres = \u0026wreq-\u003ecache_resources;\n \tstruct cachefiles_object *object = cachefiles_cres_object(cres);\n \tstruct cachefiles_cache *cache = object-\u003evolume-\u003ecache;\n-\tstruct netfs_io_stream *stream = \u0026wreq-\u003eio_streams[subreq-\u003estream_nr];\n \tconst struct cred *saved_cred;\n \tsize_t off, pre, post, len = subreq-\u003elen;\n \tuoff_t start = subreq-\u003estart;\n@@ -571,17 +570,6 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq)\n \t}\n \n \t/* We also need to end on the cache granularity boundary */\n-\tif (start + len == wreq-\u003ei_size) {\n-\t\tsize_t part = len \u0026 (cache-\u003ebsize - 1);\n-\t\tsize_t need = cache-\u003ebsize - part;\n-\n-\t\tif (part \u0026\u0026 stream-\u003esubmit_extendable_to \u003e= need) {\n-\t\t\tlen += need;\n-\t\t\tsubreq-\u003elen += need;\n-\t\t\tsubreq-\u003eio_iter.count += need;\n-\t\t}\n-\t}\n-\n \tpost = len \u0026 (cache-\u003ebsize - 1);\n \tif (post) {\n \t\tlen -= post;\ndiff --git a/fs/netfs/Makefile b/fs/netfs/Makefile\nindex b1ea4439c1bb4..421dd0be413b3 100644\n--- a/fs/netfs/Makefile\n+++ b/fs/netfs/Makefile\n@@ -14,7 +14,6 @@ netfs-y := \\\n \tread_collect.o \\\n \tread_retry.o \\\n \tread_single.o \\\n-\trolling_buffer.o \\\n \twrite_collect.o \\\n \twrite_issue.o \\\n \twrite_retry.o\ndiff --git a/fs/netfs/buffered_read.c b/fs/netfs/buffered_read.c\nindex 052684ce1e347..3ca75b5314eee 100644\n--- a/fs/netfs/buffered_read.c\n+++ b/fs/netfs/buffered_read.c\n@@ -114,26 +114,21 @@ static int netfs_begin_cache_read(struct netfs_io_request *rreq, struct netfs_in\n static ssize_t netfs_prepare_read_iterator(struct netfs_io_subrequest *subreq)\n {\n \tstruct netfs_io_request *rreq = subreq-\u003erreq;\n+\tstruct netfs_io_stream *stream = \u0026rreq-\u003eio_streams[0];\n+\tssize_t extracted;\n \tsize_t rsize = subreq-\u003elen;\n \n \tif (subreq-\u003esource == NETFS_DOWNLOAD_FROM_SERVER)\n-\t\trsize = umin(rsize, rreq-\u003eio_streams[0].sreq_max_len);\n-\n-\tsubreq-\u003elen = rsize;\n-\tif (unlikely(rreq-\u003eio_streams[0].sreq_max_segs)) {\n-\t\tsize_t limit = netfs_limit_iter(\u0026rreq-\u003ebuffer.iter, 0, rsize,\n-\t\t\t\t\t\trreq-\u003eio_streams[0].sreq_max_segs);\n-\n-\t\tif (limit \u003c rsize) {\n-\t\t\tsubreq-\u003elen = limit;\n-\t\t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_limited);\n-\t\t}\n+\t\trsize = umin(rsize, stream-\u003esreq_max_len);\n+\n+\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026rreq-\u003edispatch_cursor);\n+\textracted = bvecq_slice(\u0026rreq-\u003edispatch_cursor, rsize,\n+\t\t\t\tstream-\u003esreq_max_segs, \u0026subreq-\u003enr_segs);\n+\tif (extracted \u003c subreq-\u003elen) {\n+\t\tsubreq-\u003elen = extracted;\n+\t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_limited);\n \t}\n \n-\tsubreq-\u003eio_iter\t= rreq-\u003ebuffer.iter;\n-\n-\tiov_iter_truncate(\u0026subreq-\u003eio_iter, subreq-\u003elen);\n-\trolling_buffer_advance(\u0026rreq-\u003ebuffer, subreq-\u003elen);\n \treturn subreq-\u003elen;\n }\n \n@@ -192,6 +187,9 @@ void netfs_queue_read(struct netfs_io_request *rreq,\n static void netfs_issue_read(struct netfs_io_request *rreq,\n \t\t\t struct netfs_io_subrequest *subreq)\n {\n+\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\n+\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n+\n \tswitch (subreq-\u003esource) {\n \tcase NETFS_DOWNLOAD_FROM_SERVER:\n \t\trreq-\u003enetfs_ops-\u003eissue_read(subreq);\n@@ -200,10 +198,9 @@ static void netfs_issue_read(struct netfs_io_request *rreq,\n \t\tnetfs_read_cache_to_pagecache(rreq, subreq);\n \t\tbreak;\n \tdefault:\n-\t\t__set_bit(NETFS_SREQ_CLEAR_TAIL, \u0026subreq-\u003eflags);\n-\t\tsubreq-\u003eerror = 0;\n-\t\tiov_iter_zero(subreq-\u003elen, \u0026subreq-\u003eio_iter);\n+\t\tbvecq_zero(\u0026subreq-\u003eio_buffer, subreq-\u003elen);\n \t\tsubreq-\u003etransferred = subreq-\u003elen;\n+\t\tsubreq-\u003eerror = 0;\n \t\tnetfs_read_subreq_terminated(subreq);\n \t\tbreak;\n \t}\n@@ -215,31 +212,31 @@ static void netfs_issue_read(struct netfs_io_request *rreq,\n * otherwise we set the deprecated PG_private_2.\n */\n static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,\n-\t\t\t\t struct bvecq **bq,\n-\t\t\t\t unsigned int *offset,\n-\t\t\t\t int *slot,\n-\t\t\t\t size_t len,\n-\t\t\t\t bool copy)\n+\t\t\t\t struct bvecq_pos *mark, size_t len, bool copy)\n {\n+\tstruct bvecq *bq = mark-\u003ebvecq;\n+\tunsigned int offset = mark-\u003eoffset;\n+\tint slot = mark-\u003eslot;\n+\n \twhile (len \u003e 0) {\n-\t\tstruct folio *folio;\n \t\tsize_t fsize, overlap;\n \n-\t\tif (!*bq)\n+\t\tif (!bq)\n \t\t\tbreak;\n-\t\tif (!bvecq_acquire_slot(*bq, *slot)) {\n-\t\t\t*bq = bvecq_next(*bq);\n-\t\t\t*slot = 0;\n-\t\t\t*offset = 0;\n+\t\tif (!bvecq_acquire_slot(bq, slot)) {\n+\t\t\tbq = bq-\u003enext;\n+\t\t\tslot = 0;\n+\t\t\toffset = 0;\n \t\t\tcontinue;\n \t\t}\n \n \t\t/* Determine how much the subreq overlaps the folio, if at all. */\n-\t\tfsize = (*bq)-\u003ebv[*slot].bv_len;\n-\t\toverlap = min(len, fsize - *offset);\n+\t\tfsize = bq-\u003ebv[slot].bv_len;\n+\t\toverlap = min(len, fsize - offset);\n \n \t\tif (overlap \u003e 0 \u0026\u0026 copy) {\n-\t\t\tfolio = bvec_folio(\u0026(*bq)-\u003ebv[*slot]);\n+\t\t\tstruct folio *folio = bvec_folio(\u0026bq-\u003ebv[slot]);\n+\n \t\t\tif (netfs_using_pgpriv2(rreq)) {\n \t\t\t\tif (!folio_test_private_2(folio))\n \t\t\t\t\tfolio_start_private_2(folio);\n@@ -251,12 +248,20 @@ static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,\n \t\t}\n \n \t\tlen -= overlap;\n-\t\t*offset += overlap;\n-\t\tif (*offset \u003e= fsize) {\n-\t\t\t*slot += 1;\n-\t\t\t*offset = 0;\n+\t\toffset += overlap;\n+\t\tif (offset \u003e= fsize) {\n+\t\t\tslot += 1;\n+\t\t\toffset = 0;\n \t\t}\n \t}\n+\n+\tif (bq) {\n+\t\tbvecq_pos_move(mark, bq);\n+\t\tmark-\u003eoffset = offset;\n+\t\tmark-\u003eslot = slot;\n+\t} else {\n+\t\tbvecq_pos_unset(mark);\n+\t}\n }\n \n /*\n@@ -275,11 +280,14 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)\n \t\t.cached_to[1]\t= ULLONG_MAX,\n \t};\n \tstruct fscache_occupancy *occ = \u0026_occ;\n-\tstruct bvecq *bq = rreq-\u003ebuffer.tail;\n-\tunsigned int offset = 0;\n+\tstruct bvecq_pos mark_cursor;\n \tssize_t size = rreq-\u003elen;\n \tuoff_t start = rreq-\u003estart;\n-\tint ret = 0, slot = 0;\n+\tint ret = 0;\n+\n+\t_enter(\"R=%08x\", rreq-\u003edebug_id);\n+\n+\tbvecq_pos_set(\u0026mark_cursor, \u0026rreq-\u003edispatch_cursor);\n \n \tdo {\n \t\tint (*prepare_read)(struct netfs_io_subrequest *subreq) = NULL;\n@@ -408,10 +416,10 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)\n \t\tif (size \u003c= 0)\n \t\t\tnetfs_all_subreqs_queued(rreq);\n \n-\t\tif (bq) {\n+\t\tif (mark_cursor.bvecq) {\n \t\t\t/* See if the cache indicated this should be cached. */\n \t\t\tcopy = test_bit(NETFS_SREQ_COPY_TO_CACHE, \u0026subreq-\u003eflags);\n-\t\t\tnetfs_mark_copy_to_cache(rreq, \u0026bq, \u0026slot, \u0026offset, slice, copy);\n+\t\t\tnetfs_mark_copy_to_cache(rreq, \u0026mark_cursor, slice, copy);\n \t\t}\n \n \t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_submit);\n@@ -432,6 +440,9 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)\n \n \t/* Defer error return as we may need to wait for outstanding I/O. */\n \tcmpxchg(\u0026rreq-\u003eerror, 0, ret);\n+\n+\tbvecq_pos_unset(\u0026mark_cursor);\n+\tbvecq_pos_unset(\u0026rreq-\u003edispatch_cursor);\n }\n \n /**\n@@ -479,7 +490,7 @@ void netfs_readahead(struct readahead_control *ractl)\n \t * acquires a ref on each folio that we will need to release later -\n \t * but we don't want to do that until after we've started the I/O.\n \t */\n-\tadded = rolling_buffer_bulk_load_from_ra(\u0026rreq-\u003ebuffer, ractl, rreq-\u003egfp);\n+\tadded = bvecq_load_from_ra(\u0026rreq-\u003edispatch_cursor, ractl);\n \tif (added \u003c 0) {\n \t\tret = added;\n \t\tgoto cleanup_free;\n@@ -488,6 +499,7 @@ void netfs_readahead(struct readahead_control *ractl)\n \n \trreq-\u003esubmitted = rreq-\u003estart + added;\n \trreq-\u003ecleaned_to = rreq-\u003estart;\n+\tbvecq_pos_set(\u0026rreq-\u003ecollect_cursor, \u0026rreq-\u003edispatch_cursor);\n \tnetfs_read_set_unlock_at(rreq);\n \n \tnetfs_read_to_pagecache(rreq);\n@@ -500,20 +512,26 @@ void netfs_readahead(struct readahead_control *ractl)\n EXPORT_SYMBOL(netfs_readahead);\n \n /*\n- * Create a rolling buffer with a single occupying folio.\n+ * Create a buffer queue with a single occupying folio.\n */\n static int netfs_create_singular_buffer(struct netfs_io_request *rreq, struct folio *folio)\n {\n-\tssize_t added;\n+\tstruct bvecq *bq;\n+\tsize_t fsize = folio_size(folio);\n \n-\tif (rolling_buffer_init(\u0026rreq-\u003ebuffer, ITER_DEST, rreq-\u003egfp, false) \u003c 0)\n+\tbq = bvecq_alloc_one(1, rreq-\u003egfp, false);\n+\tif (!bq)\n \t\treturn -ENOMEM;\n \n-\tadded = rolling_buffer_append(\u0026rreq-\u003ebuffer, folio, rreq-\u003egfp);\n-\tif (added \u003c 0)\n-\t\treturn added;\n-\trreq-\u003esubmitted = rreq-\u003estart + added;\n-\trreq-\u003eprogress_at = added;\n+\trreq-\u003edispatch_cursor.bvecq = bq;\n+\trreq-\u003edispatch_cursor.slot = 0;\n+\trreq-\u003edispatch_cursor.offset = 0;\n+\n+\tbvec_set_folio(\u0026bq-\u003ebv[0], folio, fsize, 0);\n+\tbvecq_filled_to(bq, 1);\n+\tbvecq_pos_set(\u0026rreq-\u003ecollect_cursor, \u0026rreq-\u003edispatch_cursor);\n+\trreq-\u003esubmitted = rreq-\u003estart + fsize;\n+\trreq-\u003eprogress_at = fsize;\n \treturn 0;\n }\n \n@@ -527,14 +545,14 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)\n \tstruct netfs_group *group = netfs_folio_group(folio);\n \tstruct netfs_folio *finfo = netfs_folio_info(folio);\n \tstruct netfs_inode *ctx = netfs_inode(mapping-\u003ehost);\n-\tstruct bio_vec *bvec = NULL;\n+\tstruct bvecq *bq = NULL;\n \tunsigned int from = finfo-\u003edirty_offset;\n \tunsigned int to = from + finfo-\u003edirty_len;\n \tunsigned int off = 0;\n \tsize_t flen = folio_size(folio);\n \tsize_t nr_bvec = flen / PAGE_SIZE + 2;\n \tsize_t part;\n-\tint ret, i = 0, sink_from = -1, sink_to = -1;\n+\tint ret, i = 0;\n \n \t_enter(\"%lx\", folio-\u003eindex);\n \n@@ -555,31 +573,46 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)\n \t * end get copied to, but the middle is discarded.\n \t */\n \tret = -ENOMEM;\n-\tbvec = kmalloc_objs(*bvec, nr_bvec);\n-\tif (!bvec)\n+\tbq = bvecq_alloc_chain(nr_bvec, rreq-\u003egfp, false);\n+\tif (!bq)\n \t\tgoto discard;\n+\trreq-\u003edispatch_cursor.bvecq = bq;\n \n \ttrace_netfs_folio(folio, netfs_folio_trace_read_gaps);\n \n+\tfor (struct bvecq *p = bq; p; p = p-\u003enext)\n+\t\tp-\u003emem_type = BVECQ_MEM_PAGECACHE;\n+\n \tif (from \u003e 0) {\n-\t\tbvec_set_folio(\u0026bvec[i++], folio, from, 0);\n+\t\tfolio_get(folio);\n+\t\tbvec_set_folio(\u0026bq-\u003ebv[i++], folio, from, 0);\n \t\toff = from;\n \t}\n-\tsink_from = i;\n \twhile (off \u003c to) {\n \t\tstruct folio *sink = folio_alloc(GFP_KERNEL, 0);\n \n \t\tif (!sink)\n \t\t\tgoto discard;\n-\t\tpart = min_t(size_t, to - off, PAGE_SIZE);\n-\t\tbvec_set_folio(\u0026bvec[i], sink, part, 0);\n+\t\tif (i \u003e= bq-\u003emax_slots) {\n+\t\t\tbvecq_filled_to(bq, i);\n+\t\t\tbq = bq-\u003enext;\n+\t\t\ti = 0;\n+\t\t}\n+\t\tpart = min(to - off, PAGE_SIZE);\n+\t\tbvec_set_folio(\u0026bq-\u003ebv[i++], sink, part, 0);\n \t\toff += part;\n-\t\tsink_to = i;\n-\t\ti++;\n \t}\n-\tif (to \u003c flen)\n-\t\tbvec_set_folio(\u0026bvec[i++], folio, flen - to, to);\n-\tiov_iter_bvec(\u0026rreq-\u003ebuffer.iter, ITER_DEST, bvec, i, rreq-\u003elen);\n+\tif (to \u003c flen) {\n+\t\tif (i \u003e= bq-\u003emax_slots) {\n+\t\t\tbvecq_filled_to(bq, i);\n+\t\t\tbq = bq-\u003enext;\n+\t\t\ti = 0;\n+\t\t}\n+\t\tfolio_get(folio);\n+\t\tbvec_set_folio(\u0026bq-\u003ebv[i++], folio, flen - to, to);\n+\t}\n+\tbvecq_filled_to(bq, i);\n+\n \trreq-\u003esubmitted = rreq-\u003estart + flen;\n \n \tnetfs_read_to_pagecache(rreq);\n@@ -587,11 +620,10 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)\n \tret = netfs_wait_for_read(rreq);\n \tif (ret \u003e= 0) {\n \t\tif (ret \u003c flen) {\n-\t\t\tstruct iov_iter iter;\n-\n-\t\t\tiov_iter_bvec(\u0026iter, ITER_DEST, bvec, i, flen);\n-\t\t\tiov_iter_advance(\u0026iter, ret);\n-\t\t\tiov_iter_zero(flen - ret, \u0026iter);\n+\t\t\tif (ret \u003c from)\n+\t\t\t\tfolio_zero_segments(folio, ret, from, to, flen);\n+\t\t\telse\n+\t\t\t\tfolio_zero_segment(folio, max(to, ret), flen);\n \t\t}\n \t\tif (group)\n \t\t\tfolio_change_private(folio, group);\n@@ -603,22 +635,16 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)\n \t\tfolio_mark_uptodate(folio);\n \t}\n \n-\tif (sink_to \u003e= 0)\n-\t\tfor (; sink_from \u003c= sink_to; sink_from++)\n-\t\t\tfolio_put(bvec_folio(\u0026bvec[sink_from]));\n-\tkfree(bvec);\n+\tbvecq_pos_unset(\u0026rreq-\u003edispatch_cursor);\n \tfolio_unlock(folio);\n \tnetfs_put_request(rreq, netfs_rreq_trace_put_return);\n \treturn ret \u003c 0 ? ret : 0;\n \n discard:\n+\tbvecq_pos_unset(\u0026rreq-\u003edispatch_cursor);\n \tnetfs_put_failed_request(rreq);\n alloc_error:\n \tfolio_unlock(folio);\n-\tif (sink_to \u003e= 0)\n-\t\tfor (; sink_from \u003c= sink_to; sink_from++)\n-\t\t\tfolio_put(bvec_folio(\u0026bvec[sink_from]));\n-\tkfree(bvec);\n \treturn ret;\n }\n \ndiff --git a/fs/netfs/bvecq.c b/fs/netfs/bvecq.c\nindex 5b747f6b59382..2edc0045c2453 100644\n--- a/fs/netfs/bvecq.c\n+++ b/fs/netfs/bvecq.c\n@@ -342,3 +342,313 @@ int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size,\n \treturn 0;\n }\n EXPORT_SYMBOL(bvecq_expand_buffer);\n+\n+/**\n+ * bvecq_buffer_init - Initialise a buffer and set position\n+ * @pos: The position to point at the new buffer.\n+ * @gfp: The allocation constraints.\n+ * @for_writeback: True if allocating for writeback\n+ *\n+ * Initialise a rolling buffer. We allocate an unpopulated bvecq node to so\n+ * that the pointers can be independently driven by the producer and the\n+ * consumer.\n+ *\n+ * Return 0 if successful; -ENOMEM on allocation failure.\n+ */\n+int bvecq_buffer_init(struct bvecq_pos *pos, gfp_t gfp, bool for_writeback)\n+{\n+\tstruct bvecq *bq;\n+\n+\tbq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, for_writeback);\n+\tif (!bq)\n+\t\treturn -ENOMEM;\n+\n+\tpos-\u003ebvecq = bq; /* Comes with a ref. */\n+\tpos-\u003eslot = 0;\n+\tpos-\u003eoffset = 0;\n+\treturn 0;\n+}\n+\n+/**\n+ * bvecq_buffer_append - Append a new bvecq node to a buffer\n+ * @pos: The position of the last node.\n+ * @bq: The buffer to add.\n+ *\n+ * Add a new node on to the buffer chain at the specified position, either\n+ * because the previous one is full or because we have a discontiguity to\n+ * contend with, and update @pos to point to it.\n+ */\n+void bvecq_buffer_append(struct bvecq_pos *pos, struct bvecq *bq)\n+{\n+\tstruct bvecq *head = pos-\u003ebvecq;\n+\n+\tpos-\u003ebvecq = bvecq_get(bq);\n+\tpos-\u003eslot = 0;\n+\tpos-\u003eoffset = 0;\n+\n+\t/* [!] NOTE: After we set head-\u003enext, the consumer is at liberty to\n+\t * immediately delete the old head.\n+\t */\n+\tbvecq_append(head, bq);\n+\tbvecq_put(head);\n+}\n+\n+/**\n+ * bvecq_pos_advance - Advance a bvecq position\n+ * @pos: The position to advance.\n+ * @amount: The amount of bytes to advance by.\n+ *\n+ * Advance the specified bvecq position by @amount bytes. @pos is updated and\n+ * bvecq ref counts may have been manipulated. If the position hits the end of\n+ * the queue, then it is left pointing beyond the last slot of the last bvecq\n+ * so that it doesn't break the chain.\n+ */\n+void bvecq_pos_advance(struct bvecq_pos *pos, size_t amount)\n+{\n+\tstruct bvecq *bq = pos-\u003ebvecq, *next;\n+\tunsigned int slot = pos-\u003eslot;\n+\tsize_t offset = pos-\u003eoffset;\n+\n+\twhile (amount) {\n+\t\tsize_t part;\n+\n+\t\tif (!bvecq_acquire_slot(bq, slot)) {\n+\t\t\tnext = bvecq_next(bq);\n+\t\t\tif (!next) {\n+\t\t\t\tWARN_ON_ONCE(amount \u003e 0);\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t\tif (bvecq_acquire_slot(bq, slot))\n+\t\t\t\tcontinue; /* More slots got added. */\n+\t\t\tbq = next;\n+\t\t\tslot = 0;\n+\t\t\toffset = 0;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tpart = bq-\u003ebv[slot].bv_len - offset;\n+\n+\t\tif (part \u003e amount) {\n+\t\t\toffset += amount;\n+\t\t\tbreak;\n+\t\t}\n+\t\tamount -= part;\n+\t\toffset = 0;\n+\t\tslot++;\n+\t}\n+\n+\tpos-\u003eslot = slot;\n+\tpos-\u003eoffset = offset;\n+\tbvecq_pos_move(pos, bq);\n+}\n+\n+/*\n+ * Clear part of the memory pointed to by a bio_vec.\n+ */\n+static void bvec_zero(const struct bio_vec *bv, size_t offset, size_t len)\n+{\n+\tstruct page *page = bv-\u003ebv_page;\n+\n+\toffset += bv-\u003ebv_offset;\n+\n+\tpage += offset / PAGE_SIZE;\n+\toffset = offset % PAGE_SIZE;\n+\n+\twhile (len) {\n+\t\tsize_t part = min(len, PAGE_SIZE - offset);\n+\t\tchar *p = kmap_local_page(page);\n+\n+\t\tmemset(p + offset, 0, part);\n+\t\tkunmap_local(p);\n+\n+\t\tlen -= part;\n+\t\toffset = 0;\n+\t\tpage++;\n+\t}\n+}\n+\n+/**\n+ * bvecq_zero - Clear memory starting at the bvecq position.\n+ * @pos: The position in the bvecq chain to start clearing.\n+ * @amount: The number of bytes to clear.\n+ *\n+ * Clear memory fragments pointed to by a bvec queue. @pos is updated and\n+ * bvecq ref counts may have been manipulated. If the position hits the end of\n+ * the queue, then it is left pointing beyond the last slot of the last bvecq\n+ * so that it doesn't break the chain.\n+ *\n+ * Return: The number of bytes cleared.\n+ */\n+ssize_t bvecq_zero(struct bvecq_pos *pos, size_t amount)\n+{\n+\tstruct bvecq *bq = pos-\u003ebvecq, *next;\n+\tunsigned int slot = pos-\u003eslot;\n+\tssize_t cleared = 0;\n+\tsize_t offset = pos-\u003eoffset;\n+\n+\twhile (amount) {\n+\t\tconst struct bio_vec *bv;\n+\t\tsize_t part;\n+\n+\t\tif (!bvecq_acquire_slot(bq, slot)) {\n+\t\t\tnext = bvecq_next(bq);\n+\t\t\tif (!next) {\n+\t\t\t\tWARN_ON_ONCE(amount \u003e 0);\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t\tif (bvecq_acquire_slot(bq, slot))\n+\t\t\t\tcontinue; /* More slots got added. */\n+\t\t\tbq = next;\n+\t\t\tslot = 0;\n+\t\t\toffset = 0;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tbv = \u0026bq-\u003ebv[slot];\n+\t\tif (offset \u003e= bv-\u003ebv_len) {\n+\t\t\tslot++;\n+\t\t\toffset = 0;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tpart = min(bv-\u003ebv_len - offset, amount);\n+\t\tbvec_zero(bv, offset, part);\n+\t\tcleared += part;\n+\t\toffset += part;\n+\t\tamount -= part;\n+\t}\n+\n+\tpos-\u003eslot = slot;\n+\tpos-\u003eoffset = offset;\n+\tbvecq_pos_move(pos, bq);\n+\treturn cleared;\n+}\n+\n+/**\n+ * bvecq_slice - Find a slice of a bvecq queue\n+ * @pos: The position to start at.\n+ * @max_size: The maximum size of the slice (or ULONG_MAX).\n+ * @max_slots: The maximum number of slots in the slice (or INT_MAX).\n+ * @_nr_slots: Where to put the number of slots (updated).\n+ *\n+ * Determine the size and number of slots that can be obtained the next slice\n+ * of bvec queue up to the maximum size and slot count specified.\n+ *\n+ * @pos is updated to the end of the slice. If the position hits the end of\n+ * the queue, then it is left pointing beyond the last slot of the last bvecq\n+ * so that it doesn't break the chain.\n+ *\n+ * Return: The number of bytes in the slice.\n+ */\n+size_t bvecq_slice(struct bvecq_pos *pos, size_t max_size,\n+\t\t unsigned int max_slots, unsigned int *_nr_slots)\n+{\n+\tstruct bvecq *bq, *next;\n+\tunsigned int slot = pos-\u003eslot, nslots = 0;\n+\tsize_t size = 0, offset = pos-\u003eoffset;\n+\n+\tbq = pos-\u003ebvecq;\n+\tfor (;;) {\n+\t\tfor (; slot \u003c bvecq_nr_slots_acquire(bq); slot++) {\n+\t\t\tconst struct bio_vec *bvec = \u0026bq-\u003ebv[slot];\n+\n+\t\t\tif (offset \u003c bvec-\u003ebv_len \u0026\u0026 bvec-\u003ebv_page) {\n+\t\t\t\tsize_t part = min(bvec-\u003ebv_len - offset, max_size);\n+\n+\t\t\t\tsize += part;\n+\t\t\t\toffset += part;\n+\t\t\t\tmax_size -= part;\n+\t\t\t\tnslots++;\n+\t\t\t\tif (!max_size || nslots \u003e= max_slots)\n+\t\t\t\t\tgoto out;\n+\t\t\t}\n+\t\t\toffset = 0;\n+\t\t}\n+\n+\t\t/* pos-\u003ebvecq isn't allowed to go NULL as the queue may get\n+\t\t * extended and we would lose our place.\n+\t\t */\n+\t\tnext = bvecq_next(bq);\n+\t\tif (!next)\n+\t\t\tbreak;\n+\t\tif (bvecq_acquire_slot(bq, slot))\n+\t\t\tcontinue; /* More slots got added. */\n+\t\tslot = 0;\n+\t\tbq = next;\n+\t}\n+\n+out:\n+\t*_nr_slots = nslots;\n+\tif (slot == bvecq_nr_slots_acquire(bq)) {\n+\t\tnext = bvecq_next(bq);\n+\t\tif (next) {\n+\t\t\tbq = next;\n+\t\t\tslot = 0;\n+\t\t\toffset = 0;\n+\t\t}\n+\t}\n+\tbvecq_pos_move(pos, bq);\n+\tpos-\u003eslot = slot;\n+\tpos-\u003eoffset = offset;\n+\treturn size;\n+}\n+\n+/**\n+ * bvecq_load_from_ra - Allocate a bvecq chain and load from readahead\n+ * @pos: Blank position object to attach the new chain to.\n+ * @ractl: The readahead control context.\n+ *\n+ * Decant the set of folios to be read from the readahead context into a bvecq\n+ * chain. Each folio occupies one bio_vec element.\n+ *\n+ * Return: Amount of data loaded or -ENOMEM on allocation failure.\n+ */\n+ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl)\n+{\n+\tXA_STATE(xas, \u0026ractl-\u003emapping-\u003ei_pages, ractl-\u003e_index);\n+\tstruct folio *folio;\n+\tstruct bvecq *bq;\n+\tunsigned int slot = 0;\n+\tsize_t loaded = 0;\n+\n+\tbq = bvecq_alloc_chain(ractl-\u003e_nr_folios, GFP_KERNEL, false);\n+\tif (!bq)\n+\t\treturn -ENOMEM;\n+\n+\tpos-\u003ebvecq = bq;\n+\tpos-\u003eslot = 0;\n+\tpos-\u003eoffset = 0;\n+\n+\trcu_read_lock();\n+\n+\txas_for_each(\u0026xas, folio, ractl-\u003e_index + ractl-\u003e_nr_pages - 1) {\n+\t\tsize_t len;\n+\n+\t\tif (xas_retry(\u0026xas, folio))\n+\t\t\tcontinue;\n+\t\tVM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);\n+\n+\t\tlen = folio_size(folio);\n+\t\tbvec_set_folio(\u0026bq-\u003ebv[slot], folio, len, 0);\n+\t\tloaded += len;\n+\t\tslot++;\n+\t\ttrace_netfs_folio(folio, netfs_folio_trace_read);\n+\n+\t\tif (slot \u003e= bq-\u003emax_slots) {\n+\t\t\tbvecq_filled_to(bq, slot);\n+\t\t\tbq = bq-\u003enext;\n+\t\t\tif (!bq)\n+\t\t\t\tbreak;\n+\t\t\tslot = 0;\n+\t\t}\n+\t}\n+\n+\trcu_read_unlock();\n+\n+\tif (bq)\n+\t\tbvecq_filled_to(bq, slot);\n+\n+\tractl-\u003e_index += ractl-\u003e_nr_pages;\n+\tractl-\u003e_nr_pages = 0;\n+\treturn loaded;\n+}\ndiff --git a/fs/netfs/direct_read.c b/fs/netfs/direct_read.c\nindex 8c15f30797238..dae890e8df285 100644\n--- a/fs/netfs/direct_read.c\n+++ b/fs/netfs/direct_read.c\n@@ -16,44 +16,21 @@\n #include \u003clinux/netfs.h\u003e\n #include \"internal.h\"\n \n-static void netfs_prepare_dio_read_iterator(struct netfs_io_subrequest *subreq)\n-{\n-\tstruct netfs_io_request *rreq = subreq-\u003erreq;\n-\tsize_t rsize;\n-\n-\trsize = umin(subreq-\u003elen, rreq-\u003eio_streams[0].sreq_max_len);\n-\tsubreq-\u003elen = rsize;\n-\n-\tif (unlikely(rreq-\u003eio_streams[0].sreq_max_segs)) {\n-\t\tsize_t limit = netfs_limit_iter(\u0026rreq-\u003ebuffer.iter, 0, rsize,\n-\t\t\t\t\t\trreq-\u003eio_streams[0].sreq_max_segs);\n-\n-\t\tif (limit \u003c rsize) {\n-\t\t\tsubreq-\u003elen = limit;\n-\t\t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_limited);\n-\t\t}\n-\t}\n-\n-\ttrace_netfs_sreq(subreq, netfs_sreq_trace_prepare);\n-\n-\tsubreq-\u003eio_iter\t= rreq-\u003ebuffer.iter;\n-\tiov_iter_truncate(\u0026subreq-\u003eio_iter, subreq-\u003elen);\n-\tiov_iter_advance(\u0026rreq-\u003ebuffer.iter, subreq-\u003elen);\n-}\n-\n /*\n * Perform a read to a buffer from the server, slicing up the region to be read\n * according to the network rsize.\n */\n static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)\n {\n+\tstruct netfs_io_stream *stream = \u0026rreq-\u003eio_streams[0];\n \tssize_t size = rreq-\u003elen;\n \tuoff_t start = rreq-\u003estart;\n \tint ret;\n \n+\tbvecq_pos_set(\u0026rreq-\u003ecollect_cursor, \u0026rreq-\u003edispatch_cursor);\n+\n \tdo {\n \t\tstruct netfs_io_subrequest *subreq;\n-\t\tssize_t slice;\n \n \t\tsubreq = netfs_alloc_subrequest(rreq, NETFS_DOWNLOAD_FROM_SERVER);\n \t\tif (!subreq) {\n@@ -78,14 +55,22 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)\n \t\t\t}\n \t\t}\n \n-\t\tnetfs_prepare_dio_read_iterator(subreq);\n-\t\tslice = subreq-\u003elen;\n-\t\tsize -= slice;\n-\t\tstart += slice;\n-\t\trreq-\u003esubmitted += slice;\n+\t\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026rreq-\u003edispatch_cursor);\n+\t\tsubreq-\u003elen = bvecq_slice(\u0026rreq-\u003edispatch_cursor,\n+\t\t\t\t\t umin(size, stream-\u003esreq_max_len),\n+\t\t\t\t\t stream-\u003esreq_max_segs,\n+\t\t\t\t\t \u0026subreq-\u003enr_segs);\n+\n+\t\tsize -= subreq-\u003elen;\n+\t\tstart += subreq-\u003elen;\n+\t\trreq-\u003esubmitted += subreq-\u003elen;\n \t\tif (size \u003c= 0)\n \t\t\tnetfs_all_subreqs_queued(rreq);\n \n+\t\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\n+\t\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset,\n+\t\t\t\t subreq-\u003elen);\n+\n \t\trreq-\u003enetfs_ops-\u003eissue_read(subreq);\n \n \t\tif (test_bit(NETFS_RREQ_PAUSE, \u0026rreq-\u003eflags))\n@@ -99,6 +84,8 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)\n \t\tnetfs_all_subreqs_queued(rreq);\n \t\tnetfs_wake_collector(rreq);\n \t}\n+\n+\tbvecq_pos_unset(\u0026rreq-\u003edispatch_cursor);\n }\n \n /*\n@@ -177,25 +164,17 @@ ssize_t netfs_unbuffered_read_iter_locked(struct kiocb *iocb, struct iov_iter *i\n \t * buffer for ourselves as the caller's iterator will be trashed when\n \t * we return.\n \t *\n-\t * In such a case, extract an iterator to represent as much of the the\n-\t * output buffer as we can manage. Note that the extraction might not\n-\t * be able to allocate a sufficiently large bvec array and may shorten\n-\t * the request.\n+\t * Extract a buffer queue to represent as much of the output buffer as\n+\t * we can manage. The fragments are extracted into a bvecq which will\n+\t * have sufficient nodes allocated to hold all the data, though this\n+\t * may end up truncated if ENOMEM is encountered.\n \t */\n-\tif (user_backed_iter(iter)) {\n-\t\tret = netfs_extract_user_iter(iter, rreq-\u003elen, \u0026rreq-\u003ebuffer.iter, 0);\n-\t\tif (ret \u003c 0)\n-\t\t\tgoto error_put;\n-\t\trreq-\u003edirect_bv = (struct bio_vec *)rreq-\u003ebuffer.iter.bvec;\n-\t\trreq-\u003edirect_bv_count = ret;\n-\t\trreq-\u003edirect_bv_unpin = iov_iter_extract_will_pin(iter);\n-\t\trreq-\u003elen = iov_iter_count(\u0026rreq-\u003ebuffer.iter);\n-\t} else {\n-\t\trreq-\u003ebuffer.iter = *iter;\n-\t\trreq-\u003elen = orig_count;\n-\t\trreq-\u003edirect_bv_unpin = false;\n-\t\tiov_iter_advance(iter, orig_count);\n-\t}\n+\tret = netfs_extract_iter(iter, rreq-\u003elen, INT_MAX,\n+\t\t\t\t \u0026rreq-\u003edispatch_cursor.bvecq, 0, rreq-\u003egfp);\n+\tif (ret \u003c 0)\n+\t\tgoto error_put;\n+\n+\trreq-\u003elen = ret;\n \n \t// TODO: Set up bounce buffer if needed\n \ndiff --git a/fs/netfs/direct_write.c b/fs/netfs/direct_write.c\nindex cc46b7d9321f1..65c61fc67f9bf 100644\n--- a/fs/netfs/direct_write.c\n+++ b/fs/netfs/direct_write.c\n@@ -73,7 +73,11 @@ static void netfs_unbuffered_write_collect(struct netfs_io_request *wreq,\n \tspin_unlock(\u0026wreq-\u003elock);\n \n \twreq-\u003etransferred += subreq-\u003etransferred;\n-\tiov_iter_advance(\u0026wreq-\u003ebuffer.iter, subreq-\u003etransferred);\n+\tif (subreq-\u003etransferred \u003c subreq-\u003elen) {\n+\t\tbvecq_pos_unset(\u0026wreq-\u003edispatch_cursor);\n+\t\tbvecq_pos_transfer(\u0026wreq-\u003edispatch_cursor, \u0026subreq-\u003eio_buffer);\n+\t\tbvecq_pos_advance(\u0026wreq-\u003edispatch_cursor, subreq-\u003etransferred);\n+\t}\n \n \tstream-\u003ecollected_to = subreq-\u003estart + subreq-\u003etransferred;\n \twreq-\u003ecollected_to = stream-\u003ecollected_to;\n@@ -99,6 +103,8 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)\n \n \t_enter(\"%llx\", wreq-\u003elen);\n \n+\tbvecq_pos_set(\u0026wreq-\u003ecollect_cursor, \u0026wreq-\u003edispatch_cursor);\n+\n \tif (wreq-\u003eorigin == NETFS_DIO_WRITE)\n \t\tinode_dio_begin(wreq-\u003einode);\n \n@@ -116,6 +122,8 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)\n \t\t\t\tbreak;\n \t\t\t}\n \t\t\tstream-\u003econstruct = NULL;\n+\t\t} else {\n+\t\t\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026wreq-\u003edispatch_cursor);\n \t\t}\n \n \t\t/* Check if (re-)preparation failed. */\n@@ -125,9 +133,16 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)\n \t\t\tbreak;\n \t\t}\n \n-\t\tiov_iter_truncate(\u0026subreq-\u003eio_iter, wreq-\u003elen - wreq-\u003etransferred);\n+\t\tsubreq-\u003elen = bvecq_slice(\u0026wreq-\u003edispatch_cursor, stream-\u003esreq_max_len,\n+\t\t\t\t\t stream-\u003esreq_max_segs, \u0026subreq-\u003enr_segs);\n+\n+\t\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_SOURCE,\n+\t\t\t\t subreq-\u003eio_buffer.bvecq, subreq-\u003eio_buffer.slot,\n+\t\t\t\t subreq-\u003eio_buffer.offset,\n+\t\t\t\t subreq-\u003elen);\n+\n \t\tif (!iov_iter_count(\u0026subreq-\u003eio_iter)) {\n-\t\t\tpr_warn(\"netfs: Unexpected zero-length iterator R=%08x\\n\",\n+\t\t\tpr_warn(\"netfs: Unexpected zero-length slice R=%08x\\n\",\n \t\t\t\twreq-\u003edebug_id);\n \t\t\t__set_bit(NETFS_SREQ_FAILED, \u0026subreq-\u003eflags);\n \t\t\tnetfs_write_subrequest_terminated(subreq, -EIO);\n@@ -135,12 +150,6 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)\n \t\t\tbreak;\n \t\t}\n \n-\t\tsubreq-\u003elen = netfs_limit_iter(\u0026subreq-\u003eio_iter, 0,\n-\t\t\t\t\t stream-\u003esreq_max_len,\n-\t\t\t\t\t stream-\u003esreq_max_segs);\n-\t\tiov_iter_truncate(\u0026subreq-\u003eio_iter, subreq-\u003elen);\n-\t\tstream-\u003esubmit_extendable_to = subreq-\u003elen;\n-\n \t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_submit);\n \t\tstream-\u003eissue_write(subreq);\n \n@@ -175,9 +184,13 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)\n \t\t */\n \t\tsubreq-\u003eerror = -EAGAIN;\n \t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_retry);\n+\n+\t\tbvecq_pos_unset(\u0026wreq-\u003edispatch_cursor);\n+\t\tbvecq_pos_transfer(\u0026wreq-\u003edispatch_cursor, \u0026subreq-\u003eio_buffer);\n+\n \t\tif (subreq-\u003etransferred \u003e 0) {\n-\t\t\tiov_iter_advance(\u0026wreq-\u003ebuffer.iter, subreq-\u003etransferred);\n \t\t\twreq-\u003etransferred += subreq-\u003etransferred;\n+\t\t\tbvecq_pos_advance(\u0026wreq-\u003edispatch_cursor, subreq-\u003etransferred);\n \t\t}\n \n \t\tif (stream-\u003esource == NETFS_UPLOAD_TO_SERVER \u0026\u0026\n@@ -188,7 +201,6 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)\n \t\t__clear_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags);\n \t\t__clear_bit(NETFS_SREQ_BOUNDARY, \u0026subreq-\u003eflags);\n \t\t__clear_bit(NETFS_SREQ_FAILED, \u0026subreq-\u003eflags);\n-\t\tsubreq-\u003eio_iter\t\t= wreq-\u003ebuffer.iter;\n \t\tsubreq-\u003estart\t\t= wreq-\u003estart + wreq-\u003etransferred;\n \t\tsubreq-\u003elen\t\t= wreq-\u003elen - wreq-\u003etransferred;\n \t\tsubreq-\u003etransferred\t= 0;\n@@ -204,6 +216,7 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)\n \t\tnetfs_stat(\u0026netfs_n_wh_retry_write_subreq);\n \t}\n \n+\tbvecq_pos_unset(\u0026wreq-\u003edispatch_cursor);\n \tnetfs_unbuffered_write_done(wreq);\n \t_leave(\" = %d\", ret);\n \treturn ret;\n@@ -222,10 +235,10 @@ static void netfs_unbuffered_write_async(struct work_struct *work)\n * encrypted file. This can also be used for direct I/O writes.\n */\n ssize_t netfs_unbuffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *iter,\n-\t\t\t\t\t\t struct netfs_group *netfs_group)\n+\t\t\t\t\t struct netfs_group *netfs_group)\n {\n \tstruct netfs_io_request *wreq;\n-\tssize_t ret, n;\n+\tssize_t ret;\n \tuoff_t start = iocb-\u003eki_pos;\n \tuoff_t end = start + iov_iter_count(iter);\n \tsize_t len = iov_iter_count(iter);\n@@ -261,25 +274,17 @@ ssize_t netfs_unbuffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *\n \t\t * allocate a sufficiently large bvec array and may shorten the\n \t\t * request.\n \t\t */\n-\t\tif (user_backed_iter(iter)) {\n-\t\t\tn = netfs_extract_user_iter(iter, len, \u0026wreq-\u003ebuffer.iter, 0);\n-\t\t\tif (n \u003c 0) {\n-\t\t\t\tret = n;\n-\t\t\t\tgoto error_put;\n-\t\t\t}\n-\t\t\twreq-\u003edirect_bv = (struct bio_vec *)wreq-\u003ebuffer.iter.bvec;\n-\t\t\twreq-\u003edirect_bv_count = n;\n-\t\t\twreq-\u003edirect_bv_unpin = iov_iter_extract_will_pin(iter);\n-\t\t} else {\n-\t\t\t/* If this is a kernel-generated async DIO request,\n-\t\t\t * assume that any resources the iterator points to\n-\t\t\t * (eg. a bio_vec array) will persist till the end of\n-\t\t\t * the op.\n-\t\t\t */\n-\t\t\twreq-\u003ebuffer.iter = *iter;\n-\t\t}\n+\t\tssize_t n = netfs_extract_iter(iter, len, INT_MAX,\n+\t\t\t\t\t \u0026wreq-\u003edispatch_cursor.bvecq, 0, wreq-\u003egfp);\n \n-\t\twreq-\u003elen = iov_iter_count(\u0026wreq-\u003ebuffer.iter);\n+\t\tif (n \u003c 0) {\n+\t\t\tret = n;\n+\t\t\tgoto error_put;\n+\t\t}\n+\t\twreq-\u003elen = n;\n+\t\t_debug(\"dio-write %zx/%zx %u/%u\",\n+\t\t n, len, wreq-\u003edispatch_cursor.bvecq-\u003enr_slots,\n+\t\t wreq-\u003edispatch_cursor.bvecq-\u003emax_slots);\n \t}\n \n \t__set_bit(NETFS_RREQ_USE_IO_ITER, \u0026wreq-\u003eflags);\ndiff --git a/fs/netfs/internal.h b/fs/netfs/internal.h\nindex f2a86abae9b3e..2760bce732b83 100644\n--- a/fs/netfs/internal.h\n+++ b/fs/netfs/internal.h\n@@ -70,7 +70,6 @@ static inline void netfs_proc_del_rreq(struct netfs_io_request *rreq) {}\n /*\n * misc.c\n */\n-void netfs_reset_iter(struct netfs_io_subrequest *subreq);\n void netfs_wake_collector(struct netfs_io_request *rreq);\n void netfs_subreq_clear_in_progress(struct netfs_io_subrequest *subreq);\n void netfs_wait_for_in_progress_stream(struct netfs_io_request *rreq,\n@@ -239,8 +238,7 @@ void netfs_prepare_write(struct netfs_io_request *wreq,\n \t\t\t struct netfs_io_stream *stream,\n \t\t\t uoff_t start);\n void netfs_reissue_write(struct netfs_io_stream *stream,\n-\t\t\t struct netfs_io_subrequest *subreq,\n-\t\t\t struct iov_iter *source);\n+\t\t\t struct netfs_io_subrequest *subreq);\n void netfs_issue_write(struct netfs_io_request *wreq,\n \t\t struct netfs_io_stream *stream);\n size_t netfs_advance_write(struct netfs_io_request *wreq,\ndiff --git a/fs/netfs/iterator.c b/fs/netfs/iterator.c\nindex 31748526d5682..dc97e5b0d4495 100644\n--- a/fs/netfs/iterator.c\n+++ b/fs/netfs/iterator.c\n@@ -14,296 +14,144 @@\n #include \"internal.h\"\n \n /**\n- * netfs_extract_user_iter - Extract the pages from a user iterator into a bvec\n+ * netfs_extract_iter - Extract virtually contiguous pages from an iterator into a bvecq\n * @orig: The original iterator\n- * @orig_len: The amount of iterator to copy\n- * @new: The iterator to be set up\n+ * @max_len: Maximum number of bytes to extract\n+ * @max_pages: Maximum number of pages to extract\n+ * @_bvecq_head: Where to cache the bvec queue\n * @extraction_flags: Flags to qualify the request\n+ * @gfp: Allocation mode for bvecq structs.\n *\n- * Extract the page fragments from the given amount of the source iterator and\n- * build up a second iterator that refers to all of those bits. This allows\n- * the original iterator to be disposed of.\n+ * Extract virtually contiguous page fragments from the source iterator up to\n+ * the given maxima and build bvec queue that refers to all of those bits.\n+ * This allows the original iterator to disposed of.\n *\n- * @extraction_flags can have ITER_ALLOW_P2PDMA set to request peer-to-peer DMA be\n- * allowed on the pages extracted.\n+ * @extraction_flags can have ITER_ALLOW_P2PDMA set to request peer-to-peer DMA\n+ * be allowed on the pages extracted.\n *\n- * On success, the number of elements in the bvec is returned, the original\n- * iterator will have been advanced by the amount extracted.\n+ * On success or partial success, the amount of data in the bvec is returned,\n+ * the original iterator will have been advanced by the amount extracted.\n *\n- * The iov_iter_extract_mode() function should be used to query how cleanup\n- * should be performed.\n+ * If an error occurs and no pages are extracted, an error will be returned and\n+ * any allocated bvecq will be freed. If there is no data to be extracted (or\n+ * @max_len or @max_pages are zero), a single empty bvecq will be returned.\n+ *\n+ * The bvecq segments are marked with indications on how to get clean up the\n+ * extracted fragments.\n */\n-ssize_t netfs_extract_user_iter(struct iov_iter *orig, size_t orig_len,\n-\t\t\t\tstruct iov_iter *new,\n-\t\t\t\tiov_iter_extraction_t extraction_flags)\n+ssize_t netfs_extract_iter(struct iov_iter *orig, size_t max_len, size_t max_pages,\n+\t\t\t struct bvecq **_bvecq_head,\n+\t\t\t iov_iter_extraction_t extraction_flags, gfp_t gfp)\n {\n-\tstruct bio_vec *bv = NULL;\n-\tstruct page **pages;\n-\tunsigned int cur_npages;\n-\tunsigned int max_pages;\n-\tunsigned int npages = 0;\n-\tunsigned int i;\n+\tstruct bvecq *bq_tail = NULL, *bq;\n \tssize_t ret = 0;\n-\tsize_t count = orig_len, offset, len;\n-\tsize_t bv_size, pg_size;\n+\tsize_t extracted = 0;\n \n-\tif (WARN_ON_ONCE(!iter_is_ubuf(orig) \u0026\u0026 !iter_is_iovec(orig)))\n-\t\treturn -EIO;\n+\t_enter(\"{%u,%zx},%zx\", orig-\u003eiter_type, orig-\u003ecount, max_len);\n \n-\tmax_pages = iov_iter_npages(orig, INT_MAX);\n-\tbv_size = array_size(max_pages, sizeof(*bv));\n-\tbv = kvmalloc(bv_size, GFP_KERNEL);\n-\tif (!bv)\n-\t\treturn -ENOMEM;\n+\t*_bvecq_head = NULL;\n+\tif (max_len \u003e orig-\u003ecount)\n+\t\tmax_len = orig-\u003ecount;\n+\tif (!max_len || !max_pages)\n+\t\tgoto alloc_empty;\n+\tif (WARN_ON_ONCE(max_pages \u003e INT_MAX))\n+\t\tmax_pages = INT_MAX; /* Protect iov_iter_npages(). */\n \n-\t/* Put the page list at the end of the bvec list storage. bvec\n-\t * elements are larger than page pointers, so as long as we work\n-\t * 0-\u003elast, we should be fine.\n-\t */\n-\tpg_size = array_size(max_pages, sizeof(*pages));\n-\tpages = (void *)bv + bv_size - pg_size;\n+\tmax_pages = iov_iter_npages(orig, max_pages);\n+\tif (!max_pages)\n+\t\tgoto alloc_empty;\n \n-\twhile (count \u0026\u0026 npages \u003c max_pages) {\n-\t\tret = iov_iter_extract_pages(orig, \u0026pages, count,\n-\t\t\t\t\t max_pages - npages, extraction_flags,\n-\t\t\t\t\t \u0026offset);\n-\t\tif (unlikely(ret \u003c= 0)) {\n-\t\t\tret = ret ?: -EIO;\n+\tdo {\n+\t\tbq = bvecq_alloc_one(max_pages, gfp, false);\n+\t\tif (!bq) {\n+\t\t\tret = -ENOMEM;\n \t\t\tbreak;\n \t\t}\n+\t\tif (user_backed_iter(orig))\n+\t\t\tbq-\u003emem_type = iov_iter_extract_will_pin(orig) ?\n+\t\t\t\tBVECQ_MEM_GUP : BVECQ_MEM_PAGECACHE;\n \n-\t\tif (WARN(ret \u003e count,\n-\t\t\t \"%s: extract_pages overrun %zd \u003e %zu bytes\\n\",\n-\t\t\t __func__, ret, count)) {\n-\t\t\tret = -EIO;\n-\t\t\tbreak;\n-\t\t}\n+\t\tif (bq_tail)\n+\t\t\tbvecq_append(bq_tail, bq);\n+\t\telse\n+\t\t\t*_bvecq_head = bq;\n+\t\tbq_tail = bq;\n \n-\t\tcur_npages = DIV_ROUND_UP(offset + ret, PAGE_SIZE);\n-\t\tif (WARN(cur_npages \u003e max_pages - npages,\n-\t\t\t \"%s: extract_pages overrun %u \u003e %u pages\\n\",\n-\t\t\t __func__, npages + cur_npages, max_pages)) {\n-\t\t\tret = -EIO;\n+\t\tif (max_len == 0)\n \t\t\tbreak;\n-\t\t}\n-\n-\t\tcount -= ret;\n-\t\tret += offset;\n-\n-\t\tfor (i = 0; i \u003c cur_npages; i++) {\n-\t\t\tlen = ret \u003e PAGE_SIZE ? PAGE_SIZE : ret;\n-\t\t\tbvec_set_page(bv + npages + i, *pages++, len - offset, offset);\n-\t\t\tret -= len;\n-\t\t\toffset = 0;\n-\t\t}\n-\n-\t\tnpages += cur_npages;\n-\t}\n-\n-\t/* Note: Don't try to clean up after EIO. Either we got no pages, so\n-\t * nothing to clean up, or we got a buffer overrun, memory corruption\n-\t * and can't trust the stuff in the buffer (a WARN was emitted).\n-\t */\n-\n-\tif (ret \u003c 0 \u0026\u0026 (ret == -ENOMEM || npages == 0)) {\n-\t\tfor (i = 0; i \u003c npages; i++)\n-\t\t\tunpin_user_page(bv[i].bv_page);\n-\t\tkvfree(bv);\n-\t\treturn ret;\n-\t}\n \n-\tiov_iter_bvec(new, orig-\u003edata_source, bv, npages, orig_len - count);\n-\treturn npages;\n-}\n-EXPORT_SYMBOL_GPL(netfs_extract_user_iter);\n-\n-/*\n- * Select the span of a bvec iterator we're going to use. Limit it by both maximum\n- * size and maximum number of segments. Returns the size of the span in bytes.\n- */\n-static size_t netfs_limit_bvec(const struct iov_iter *iter, size_t start_offset,\n-\t\t\t size_t max_size, size_t max_segs)\n-{\n-\tconst struct bio_vec *bvecs = iter-\u003ebvec;\n-\tunsigned int nbv = iter-\u003enr_segs, ix = 0, nsegs = 0;\n-\tsize_t len, span = 0, n = iter-\u003ecount;\n-\tsize_t skip = iter-\u003eiov_offset + start_offset;\n-\n-\tif (WARN_ON(!iov_iter_is_bvec(iter)) ||\n-\t WARN_ON(start_offset \u003e n) ||\n-\t n == 0)\n-\t\treturn 0;\n-\n-\twhile (n \u0026\u0026 ix \u003c nbv \u0026\u0026 skip) {\n-\t\tlen = bvecs[ix].bv_len;\n-\t\tif (skip \u003c len)\n-\t\t\tbreak;\n-\t\tskip -= len;\n-\t\tn -= len;\n-\t\tix++;\n-\t}\n-\n-\twhile (n \u0026\u0026 ix \u003c nbv) {\n-\t\tlen = min3(n, bvecs[ix].bv_len - skip, max_size);\n-\t\tspan += len;\n-\t\tnsegs++;\n-\t\tix++;\n-\t\tif (span \u003e= max_size || nsegs \u003e= max_segs)\n-\t\t\tbreak;\n-\t\tskip = 0;\n-\t\tn -= len;\n-\t}\n-\n-\treturn min(span, max_size);\n-}\n-\n-/*\n- * Select the span of a kvec iterator we're going to use. Limit it by both\n- * maximum size and maximum number of segments. Returns the size of the span\n- * in bytes.\n- */\n-static size_t netfs_limit_kvec(const struct iov_iter *iter, size_t start_offset,\n-\t\t\t size_t max_size, size_t max_segs)\n-{\n-\tconst struct kvec *kvecs = iter-\u003ekvec;\n-\tunsigned int nkv = iter-\u003enr_segs, ix = 0, nsegs = 0;\n-\tsize_t len, span = 0, n = iter-\u003ecount;\n-\tsize_t skip = iter-\u003eiov_offset + start_offset;\n-\n-\tif (WARN_ON(!iov_iter_is_kvec(iter)) ||\n-\t WARN_ON(start_offset \u003e n) ||\n-\t n == 0)\n-\t\treturn 0;\n-\n-\twhile (n \u0026\u0026 ix \u003c nkv \u0026\u0026 skip) {\n-\t\tlen = kvecs[ix].iov_len;\n-\t\tif (skip \u003c len)\n-\t\t\tbreak;\n-\t\tskip -= len;\n-\t\tn -= len;\n-\t\tix++;\n-\t}\n-\n-\twhile (n \u0026\u0026 ix \u003c nkv) {\n-\t\tlen = min3(n, kvecs[ix].iov_len - skip, max_size);\n-\t\tspan += len;\n-\t\tnsegs++;\n-\t\tix++;\n-\t\tif (span \u003e= max_size || nsegs \u003e= max_segs)\n-\t\t\tbreak;\n-\t\tskip = 0;\n-\t\tn -= len;\n-\t}\n-\n-\treturn min(span, max_size);\n-}\n-\n-/*\n- * Select the span of an xarray iterator we're going to use. Limit it by both\n- * maximum size and maximum number of segments. It is assumed that segments\n- * can be larger than a page in size, provided they're physically contiguous.\n- * Returns the size of the span in bytes.\n- */\n-static size_t netfs_limit_xarray(const struct iov_iter *iter, size_t start_offset,\n-\t\t\t\t size_t max_size, size_t max_segs)\n-{\n-\tstruct folio *folio;\n-\tunsigned int nsegs = 0;\n-\tuoff_t pos = iter-\u003exarray_start + iter-\u003eiov_offset;\n-\tpgoff_t index = pos / PAGE_SIZE;\n-\tsize_t span = 0, n = iter-\u003ecount;\n-\n-\tXA_STATE(xas, iter-\u003exarray, index);\n-\n-\tif (WARN_ON(!iov_iter_is_xarray(iter)) ||\n-\t WARN_ON(start_offset \u003e n) ||\n-\t n == 0)\n-\t\treturn 0;\n-\tmax_size = min(max_size, n - start_offset);\n-\n-\trcu_read_lock();\n-\txas_for_each(\u0026xas, folio, ULONG_MAX) {\n-\t\tsize_t offset, flen, len;\n-\t\tif (xas_retry(\u0026xas, folio))\n-\t\t\tcontinue;\n-\t\tif (WARN_ON(xa_is_value(folio)))\n-\t\t\tbreak;\n-\t\tif (WARN_ON(folio_test_hugetlb(folio)))\n-\t\t\tbreak;\n-\n-\t\tflen = folio_size(folio);\n-\t\toffset = offset_in_folio(folio, pos);\n-\t\tlen = min(max_size, flen - offset);\n-\t\tspan += len;\n-\t\tnsegs++;\n-\t\tif (span \u003e= max_size || nsegs \u003e= max_segs)\n-\t\t\tbreak;\n-\t}\n-\n-\trcu_read_unlock();\n-\treturn min(span, max_size);\n-}\n-\n-/*\n- * Select the span of a bvecq iterator we're going to use. Limit it by both\n- * maximum size and maximum number of segments. Returns the size of the span\n- * in bytes.\n- */\n-static size_t netfs_limit_bvecq(const struct iov_iter *iter, size_t start_offset,\n-\t\t\t\tsize_t max_size, size_t max_segs)\n-{\n-\tconst struct bvecq *bq = iter-\u003ebvecq;\n-\tunsigned int nsegs = 0;\n-\tunsigned int slot = iter-\u003ebvecq_slot;\n-\tsize_t span = 0, n = iter-\u003ecount;\n-\n-\tif (WARN_ON(!iov_iter_is_bvecq(iter)) ||\n-\t WARN_ON(start_offset \u003e n) ||\n-\t n == 0)\n-\t\treturn 0;\n-\tmax_size = umin(max_size, n - start_offset);\n-\n-\tif (!bvecq_acquire_slot(bq, slot)) {\n-\t\tbq = bvecq_next(bq);\n-\t\tslot = 0;\n-\t}\n-\n-\tstart_offset += iter-\u003eiov_offset;\n-\tdo {\n-\t\tsize_t flen;\n-\n-\t\tflen = bq-\u003ebv[slot].bv_len;\n-\t\tif (start_offset \u003c flen) {\n-\t\t\tspan += flen - start_offset;\n-\t\t\tnsegs++;\n-\t\t\tstart_offset = 0;\n-\t\t} else {\n-\t\t\tstart_offset -= flen;\n-\t\t}\n-\t\tif (span \u003e= max_size || nsegs \u003e= max_segs)\n-\t\t\tbreak;\n-\n-\t\tslot++;\n-\t\tif (!bvecq_acquire_slot(bq, slot)) {\n-\t\t\tbq = bvecq_next(bq);\n-\t\t\tslot = 0;\n-\t\t}\n-\t} while (bq);\n-\n-\treturn umin(span, max_size);\n-}\n+\t\tstruct bio_vec *bv = bq-\u003ebv;\n+\t\tunsigned int slot = 0;\n+\t\tdo {\n+\t\t\tstruct page **pages;\n+\t\t\tssize_t got;\n+\t\t\tsize_t offset;\n+\t\t\tsize_t space = bq-\u003emax_slots - slot;\n+\t\t\tsize_t bv_size = array_size(bq-\u003emax_slots, sizeof(*bv));\n+\t\t\tsize_t pg_size = array_size(space, sizeof(*pages));\n+\n+\t\t\t/* Put the page list at the end of the bvec list\n+\t\t\t * storage. bvec elements are larger than page\n+\t\t\t * pointers, so as long as we work 0-\u003elast, we should\n+\t\t\t * be fine.\n+\t\t\t */\n+\t\t\tpages = (void *)bv + bv_size - pg_size;\n+\n+\t\t\tgot = iov_iter_extract_pages(orig, \u0026pages, max_len,\n+\t\t\t\t\t\t min(space, max_pages),\n+\t\t\t\t\t\t extraction_flags, \u0026offset);\n+\t\t\tif (got \u003c 0) {\n+\t\t\t\tret = got;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\n+\t\t\tif (got == 0) {\n+\t\t\t\tpr_err(\"extract_pages gave nothing from %zx, %zx\\n\",\n+\t\t\t\t extracted, max_len);\n+\t\t\t\tret = -EIO;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\n+\t\t\tif (WARN(got \u003e max_len,\n+\t\t\t\t \"%s: extract_pages overrun %zx \u003e %zx bytes\\n\",\n+\t\t\t\t __func__, got, max_len)) {\n+\t\t\t\tret = -EIO;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\n+\t\t\textracted += got;\n+\t\t\tmax_len -= got;\n+\n+\t\t\tdo {\n+\t\t\t\tsize_t len = umin(got, PAGE_SIZE - offset);\n+\n+\t\t\t\tBUG_ON(slot \u003e= bq-\u003emax_slots);\n+\n+\t\t\t\tbvec_set_page(\u0026bq-\u003ebv[slot], *pages++, len, offset);\n+\t\t\t\tslot++;\n+\t\t\t\tmax_pages--;\n+\t\t\t\tgot -= len;\n+\t\t\t\toffset = 0;\n+\t\t\t} while (got \u003e 0);\n+\n+\t\t\tbvecq_filled_to(bq, slot);\n+\t\t} while (max_len \u003e 0 \u0026\u0026 max_pages \u003e 0 \u0026\u0026 !bvecq_is_full(bq));\n+\n+\t} while (max_len \u003e 0 \u0026\u0026 max_pages \u003e 0);\n+\n+out:\n+\tif (extracted || ret == 0)\n+\t\treturn extracted;\n+\tbvecq_put(*_bvecq_head);\n+\t*_bvecq_head = NULL;\n+\treturn ret;\n+\n+alloc_empty:\n+\tbq = bvecq_alloc_one(1, gfp, false);\n+\tif (!bq)\n+\t\treturn -ENOMEM;\n+\t*_bvecq_head = bq;\n+\treturn 0;\n \n-size_t netfs_limit_iter(const struct iov_iter *iter, size_t start_offset,\n-\t\t\tsize_t max_size, size_t max_segs)\n-{\n-\tif (iov_iter_is_bvecq(iter))\n-\t\treturn netfs_limit_bvecq(iter, start_offset, max_size, max_segs);\n-\tif (iov_iter_is_bvec(iter))\n-\t\treturn netfs_limit_bvec(iter, start_offset, max_size, max_segs);\n-\tif (iov_iter_is_xarray(iter))\n-\t\treturn netfs_limit_xarray(iter, start_offset, max_size, max_segs);\n-\tif (iov_iter_is_kvec(iter))\n-\t\treturn netfs_limit_kvec(iter, start_offset, max_size, max_segs);\n-\tBUG();\n }\n-EXPORT_SYMBOL(netfs_limit_iter);\n+EXPORT_SYMBOL_GPL(netfs_extract_iter);\ndiff --git a/fs/netfs/misc.c b/fs/netfs/misc.c\nindex a0cc248a284d5..130bd432b1948 100644\n--- a/fs/netfs/misc.c\n+++ b/fs/netfs/misc.c\n@@ -9,24 +9,6 @@\n #include \u003clinux/rmap.h\u003e\n #include \"internal.h\"\n \n-/*\n- * Reset the subrequest iterator to refer just to the region remaining to be\n- * read. The iterator may or may not have been advanced by socket ops or\n- * extraction ops to an extent that may or may not match the amount actually\n- * read.\n- */\n-void netfs_reset_iter(struct netfs_io_subrequest *subreq)\n-{\n-\tstruct iov_iter *io_iter = \u0026subreq-\u003eio_iter;\n-\tsize_t remain = subreq-\u003elen - subreq-\u003etransferred;\n-\n-\tif (io_iter-\u003ecount \u003e remain)\n-\t\tiov_iter_advance(io_iter, io_iter-\u003ecount - remain);\n-\telse if (io_iter-\u003ecount \u003c remain)\n-\t\tiov_iter_revert(io_iter, remain - io_iter-\u003ecount);\n-\tiov_iter_truncate(\u0026subreq-\u003eio_iter, remain);\n-}\n-\n /**\n * netfs_dirty_folio - Mark folio dirty and pin a cache object for writeback\n * @mapping: The mapping the folio belongs to.\ndiff --git a/fs/netfs/objects.c b/fs/netfs/objects.c\nindex 4b8d20559b0e1..bf17dc31fd9d8 100644\n--- a/fs/netfs/objects.c\n+++ b/fs/netfs/objects.c\n@@ -133,7 +133,6 @@ static void netfs_free_request_rcu(struct rcu_head *rcu)\n static void netfs_deinit_request(struct netfs_io_request *rreq)\n {\n \tstruct netfs_inode *ictx = netfs_inode(rreq-\u003einode);\n-\tunsigned int i;\n \n \ttrace_netfs_rreq(rreq, netfs_rreq_trace_free);\n \n@@ -148,16 +147,10 @@ static void netfs_deinit_request(struct netfs_io_request *rreq)\n \t\trreq-\u003enetfs_ops-\u003efree_request(rreq);\n \tif (rreq-\u003ecache_resources.ops)\n \t\trreq-\u003ecache_resources.ops-\u003eend_operation(\u0026rreq-\u003ecache_resources);\n-\tif (rreq-\u003edirect_bv) {\n-\t\tfor (i = 0; i \u003c rreq-\u003edirect_bv_count; i++) {\n-\t\t\tif (rreq-\u003edirect_bv[i].bv_page) {\n-\t\t\t\tif (rreq-\u003edirect_bv_unpin)\n-\t\t\t\t\tunpin_user_page(rreq-\u003edirect_bv[i].bv_page);\n-\t\t\t}\n-\t\t}\n-\t\tkvfree(rreq-\u003edirect_bv);\n-\t}\n-\trolling_buffer_clear(\u0026rreq-\u003ebuffer);\n+\tbvecq_pos_unset(\u0026rreq-\u003eload_cursor);\n+\tbvecq_pos_unset(\u0026rreq-\u003edispatch_cursor);\n+\tbvecq_pos_unset(\u0026rreq-\u003ecollect_cursor);\n+\tbvecq_put(rreq-\u003espare);\n \n \tif (atomic_dec_and_test(\u0026ictx-\u003eio_count))\n \t\twake_up_var(\u0026ictx-\u003eio_count);\n@@ -251,6 +244,7 @@ static void netfs_free_subrequest(struct netfs_io_subrequest *subreq)\n \ttrace_netfs_sreq(subreq, netfs_sreq_trace_free);\n \tif (rreq-\u003enetfs_ops-\u003efree_subrequest)\n \t\trreq-\u003enetfs_ops-\u003efree_subrequest(subreq);\n+\tbvecq_pos_unset(\u0026subreq-\u003eio_buffer);\n \tmempool_free(subreq, rreq-\u003enetfs_ops-\u003esubrequest_pool ?: \u0026netfs_subrequest_pool);\n \tnetfs_stat_d(\u0026netfs_n_rh_sreq);\n \tnetfs_put_request(rreq, netfs_rreq_trace_put_subreq);\ndiff --git a/fs/netfs/read_collect.c b/fs/netfs/read_collect.c\nindex 75c8874ec4595..368f14cd2d982 100644\n--- a/fs/netfs/read_collect.c\n+++ b/fs/netfs/read_collect.c\n@@ -26,27 +26,35 @@\n */\n static void netfs_clear_unread(struct netfs_io_subrequest *subreq)\n {\n-\tnetfs_reset_iter(subreq);\n-\tWARN_ON_ONCE(subreq-\u003elen - subreq-\u003etransferred != iov_iter_count(\u0026subreq-\u003eio_iter));\n-\tiov_iter_zero(iov_iter_count(\u0026subreq-\u003eio_iter), \u0026subreq-\u003eio_iter);\n+\tstruct iov_iter iter;\n+\n+\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\n+\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n+\tiov_iter_advance(\u0026iter, subreq-\u003etransferred);\n+\tiov_iter_zero(subreq-\u003elen, \u0026iter);\n+\n \tif (subreq-\u003estart + subreq-\u003etransferred \u003e= subreq-\u003erreq-\u003ei_size)\n \t\t__set_bit(NETFS_SREQ_HIT_EOF, \u0026subreq-\u003eflags);\n }\n \n static void netfs_clear_unread_dio(struct netfs_io_subrequest *subreq)\n {\n-\tuoff_t pos = subreq-\u003estart + subreq-\u003etransferred;\n \tstruct netfs_io_request *rreq = subreq-\u003erreq;\n+\tstruct iov_iter iter;\n+\tuoff_t pos = subreq-\u003estart + subreq-\u003etransferred;\n \tsize_t fill;\n \n \tif (pos \u003e= rreq-\u003ei_size)\n \t\treturn;\n \n-\tfill = min_t(uoff_t, rreq-\u003ei_size - pos,\n-\t\t subreq-\u003elen - subreq-\u003etransferred);\n+\tfill = umin(rreq-\u003ei_size - pos, subreq-\u003elen - subreq-\u003etransferred);\n \n-\tnetfs_reset_iter(subreq);\n-\tsubreq-\u003etransferred += iov_iter_zero(fill, \u0026subreq-\u003eio_iter);\n+\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\n+\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n+\tiov_iter_advance(\u0026iter, subreq-\u003etransferred);\n+\tiov_iter_zero(fill, \u0026iter);\n+\n+\tsubreq-\u003etransferred += iov_iter_zero(fill, \u0026iter);\n }\n \n /*\n@@ -138,8 +146,8 @@ static void netfs_unlock_read_folio(struct netfs_io_request *rreq,\n */\n void netfs_read_set_unlock_at(struct netfs_io_request *rreq)\n {\n-\tconst struct bvecq *bq = rreq-\u003ebuffer.tail;\n-\tunsigned int slot = rreq-\u003ebuffer.first_tail_slot;\n+\tconst struct bvecq *bq = rreq-\u003ecollect_cursor.bvecq;\n+\tunsigned int slot = rreq-\u003ecollect_cursor.slot;\n \tsize_t cleaned_to = rreq-\u003ecleaned_to - rreq-\u003estart;\n \tsize_t progress_at = cleaned_to;\n \tsize_t minimum = 256 * 1024;\n@@ -169,8 +177,8 @@ void netfs_read_set_unlock_at(struct netfs_io_request *rreq)\n static void netfs_read_unlock_folios(struct netfs_io_request *rreq,\n \t\t\t\t unsigned int *notes)\n {\n-\tstruct bvecq *bq = rreq-\u003ebuffer.tail;\n-\tunsigned int slot = rreq-\u003ebuffer.first_tail_slot;\n+\tstruct bvecq *bq = rreq-\u003ecollect_cursor.bvecq;\n+\tunsigned int slot = rreq-\u003ecollect_cursor.slot;\n \tuoff_t collected_to = rreq-\u003ecollected_to;\n \n \tif (rreq-\u003ecleaned_to \u003e= rreq-\u003ecollected_to)\n@@ -178,15 +186,6 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,\n \n \t// TODO: Begin decryption\n \n-\twhile (!bvecq_acquire_slot(bq, slot)) {\n-\t\tbq = rolling_buffer_delete_spent(\u0026rreq-\u003ebuffer);\n-\t\tif (!bq) {\n-\t\t\tWRITE_ONCE(rreq-\u003eprogress_at, rreq-\u003elen);\n-\t\t\treturn;\n-\t\t}\n-\t\tslot = 0;\n-\t}\n-\n \t/* We have to wait for readahead refs to have been released before we\n \t * can unlock any folios as the ref-dropper walks i_pages and the only\n \t * thing preventing these folios from being removed is the folio lock.\n@@ -196,9 +195,24 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,\n \n \tfor (;;) {\n \t\tstruct folio *folio;\n-\t\tuoff_t fpos, fend;\n+\t\tuoff_t fpos = rreq-\u003ecleaned_to, fend;\n \t\tsize_t fsize;\n \n+\t\t/* Clean up the head bvecq segment. If we clear an entire\n+\t\t * segment, then we can get rid of it provided it's not also\n+\t\t * the tail segment being filled by the issuer.\n+\t\t */\n+\t\tif (!bvecq_acquire_slot(bq, slot)) {\n+\t\t\trreq-\u003ecollect_cursor.slot = slot;\n+\t\t\tif (!bvecq_delete_spent(\u0026rreq-\u003ecollect_cursor)) {\n+\t\t\t\tWRITE_ONCE(rreq-\u003eprogress_at, rreq-\u003elen);\n+\t\t\t\treturn;\n+\t\t\t}\n+\t\t\tbq = rreq-\u003ecollect_cursor.bvecq;\n+\t\t\tslot = rreq-\u003ecollect_cursor.slot;\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tfolio = bvec_folio(\u0026bq-\u003ebv[slot]);\n \t\tif (WARN_ONCE(!folio_test_locked(folio),\n \t\t\t \"R=%08x: folio %lx is not locked\\n\",\n@@ -206,7 +220,6 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,\n \t\t\ttrace_netfs_folio(folio, netfs_folio_trace_not_locked);\n \n \t\tfsize = bq-\u003ebv[slot].bv_len;\n-\t\tfpos = folio_pos(folio);\n \t\tfend = fpos + fsize;\n \n \t\ttrace_netfs_collect_folio(rreq, folio);\n@@ -216,30 +229,16 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,\n \t\t\tbreak;\n \n \t\tnetfs_unlock_read_folio(rreq, bq, slot);\n-\t\tWRITE_ONCE(rreq-\u003ecleaned_to, fpos + fsize);\n-\t\t*notes |= MADE_PROGRESS;\n-\n-\t\t/* Clean up the head bq. If we clear an entire bq, then\n-\t\t * we can get rid of it provided it's not also the tail bq\n-\t\t * being filled by the issuer.\n-\t\t */\n-\t\tbq-\u003ebv[slot].bv_page = NULL;\n \t\tslot++;\n-\t\twhile (!bvecq_acquire_slot(bq, slot)) {\n-\t\t\tbq = rolling_buffer_delete_spent(\u0026rreq-\u003ebuffer);\n-\t\t\tif (!bq)\n-\t\t\t\tgoto done;\n-\t\t\tslot = 0;\n-\t\t}\n+\t\tWRITE_ONCE(rreq-\u003ecleaned_to, fend);\n+\t\t*notes |= MADE_PROGRESS;\n \n \t\tif (fpos + fsize \u003e= collected_to)\n \t\t\tbreak;\n \t}\n \n-\trreq-\u003ebuffer.tail = bq;\n-done:\n-\trreq-\u003ebuffer.first_tail_slot = slot;\n-\n+\tbvecq_pos_move(\u0026rreq-\u003ecollect_cursor, bq);\n+\trreq-\u003ecollect_cursor.slot = slot;\n \tnetfs_read_set_unlock_at(rreq);\n }\n \n@@ -422,12 +421,15 @@ static void netfs_rreq_assess_dio(struct netfs_io_request *rreq)\n \n \tif (rreq-\u003eorigin == NETFS_UNBUFFERED_READ ||\n \t rreq-\u003eorigin == NETFS_DIO_READ) {\n-\t\tfor (i = 0; i \u003c rreq-\u003edirect_bv_count; i++) {\n-\t\t\tflush_dcache_page(rreq-\u003edirect_bv[i].bv_page);\n-\t\t\t// TODO: cifs marks pages in the destination buffer\n-\t\t\t// dirty under some circumstances after a read. Do we\n-\t\t\t// need to do that too?\n-\t\t\tset_page_dirty(rreq-\u003edirect_bv[i].bv_page);\n+\t\tfor (struct bvecq *bq = rreq-\u003ecollect_cursor.bvecq; bq; bq = bvecq_next(bq)) {\n+\t\t\tunsigned int nr_slots = bvecq_nr_slots_acquire(bq);\n+\t\t\t/* Read the slot count before the slots. */\n+\n+\t\t\t/* Mark the target buffers dirty. */\n+\t\t\tfor (i = 0; i \u003c nr_slots; i++) {\n+\t\t\t\tflush_dcache_page(bq-\u003ebv[i].bv_page);\n+\t\t\t\tset_page_dirty(bq-\u003ebv[i].bv_page);\n+\t\t\t}\n \t\t}\n \t}\n \n@@ -521,7 +523,15 @@ bool netfs_read_collection(struct netfs_io_request *rreq)\n \n \ttrace_netfs_rreq(rreq, netfs_rreq_trace_done);\n \tnetfs_clear_subrequests(rreq);\n-\tnetfs_unlock_abandoned_read_pages(rreq);\n+\tswitch (rreq-\u003eorigin) {\n+\tcase NETFS_READAHEAD:\n+\tcase NETFS_READPAGE:\n+\tcase NETFS_READ_FOR_WRITE:\n+\t\tnetfs_unlock_abandoned_read_pages(rreq);\n+\t\tbreak;\n+\tdefault:\n+\t\tbreak;\n+\t}\n \tif (unlikely(rreq-\u003ecopy_to_cache))\n \t\tnetfs_pgpriv2_end_copy_to_cache(rreq);\n \treturn true;\ndiff --git a/fs/netfs/read_pgpriv2.c b/fs/netfs/read_pgpriv2.c\nindex f8e5667e278e1..f23c4cbfed581 100644\n--- a/fs/netfs/read_pgpriv2.c\n+++ b/fs/netfs/read_pgpriv2.c\n@@ -19,6 +19,9 @@\n static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio *folio)\n {\n \tstruct netfs_io_stream *cache = \u0026creq-\u003eio_streams[1];\n+\tstruct bvecq *queue;\n+\tunsigned int slot;\n+\tsize_t dio_size = PAGE_SIZE;\n \tsize_t fsize = folio_size(folio), flen = fsize;\n \tuoff_t fpos = folio_pos(folio), i_size;\n \tbool to_eof = false;\n@@ -48,18 +51,37 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio\n \t\tto_eof = true;\n \t}\n \n+\tflen = round_up(flen, dio_size);\n+\n \t_debug(\"folio %zx %zx\", flen, fsize);\n \n \ttrace_netfs_folio(folio, netfs_folio_trace_store_copy);\n \n-\t/* Attach the folio to the rolling buffer. */\n-\tif (rolling_buffer_append(\u0026creq-\u003ebuffer, folio, creq-\u003egfp) \u003c 0) {\n-\t\tset_bit(NETFS_RREQ_CANCEL_CACHING, \u0026creq-\u003eflags);\n-\t\tfolio_end_private_2(folio);\n-\t\treturn;\n+\t/* Institute a new bvec queue segment if the current one is full or if\n+\t * we encounter a discontiguity. The discontiguity break is important\n+\t * when it comes to bulk unlocking folios by file range.\n+\t */\n+\tqueue = creq-\u003eload_cursor.bvecq;\n+\tif (bvecq_is_full(queue) ||\n+\t (fpos != creq-\u003elast_end \u0026\u0026 creq-\u003elast_end \u003e 0 \u0026\u0026 queue-\u003enr_slots \u003e 0)) {\n+\t\tbvecq_buffer_append(\u0026creq-\u003eload_cursor, creq-\u003espare);\n+\t\tcreq-\u003espare = NULL;\n+\n+\t\tqueue = creq-\u003eload_cursor.bvecq;\n \t}\n \n-\tcache-\u003esubmit_extendable_to = fsize;\n+\t/* Attach the folio to the rolling buffer. */\n+\tslot = queue-\u003enr_slots;\n+\tbvec_set_folio(\u0026queue-\u003ebv[slot], folio, fsize, 0);\n+\ttrace_netfs_bv_slot(queue, slot);\n+\tslot++;\n+\tbvecq_filled_to(queue, slot);\n+\tcreq-\u003eload_cursor.slot = slot;\n+\tcreq-\u003eload_cursor.offset = 0;\n+\tcreq-\u003elast_end = fpos + flen;\n+\n+\tbvecq_pos_nudge(\u0026creq-\u003edispatch_cursor);\n+\n \tcache-\u003esubmit_off = 0;\n \tcache-\u003esubmit_len = flen;\n \n@@ -71,10 +93,9 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio\n \tdo {\n \t\tssize_t part;\n \n-\t\tcreq-\u003ebuffer.iter.iov_offset = cache-\u003esubmit_off;\n+\t\tcreq-\u003edispatch_cursor.offset = cache-\u003esubmit_off;\n \n \t\tatomic64_set(\u0026creq-\u003eissued_to, fpos + cache-\u003esubmit_off);\n-\t\tcache-\u003esubmit_extendable_to = fsize - cache-\u003esubmit_off;\n \t\tpart = netfs_advance_write(creq, cache, fpos + cache-\u003esubmit_off,\n \t\t\t\t\t cache-\u003esubmit_len, to_eof);\n \t\tcache-\u003esubmit_off += part;\n@@ -84,8 +105,7 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio\n \t\t\tcache-\u003esubmit_len -= part;\n \t} while (cache-\u003esubmit_len \u003e 0);\n \n-\tcreq-\u003ebuffer.iter.iov_offset = 0;\n-\trolling_buffer_advance(\u0026creq-\u003ebuffer, fsize);\n+\tbvecq_pos_step(\u0026creq-\u003edispatch_cursor);\n \tatomic64_set(\u0026creq-\u003eissued_to, fpos + fsize);\n \n \tif (flen \u003c fsize)\n@@ -111,6 +131,11 @@ static struct netfs_io_request *netfs_pgpriv2_begin_copy_to_cache(\n \tif (!creq-\u003eio_streams[1].avail)\n \t\tgoto cancel_put;\n \n+\tif (bvecq_buffer_init(\u0026creq-\u003eload_cursor, creq-\u003egfp, false) \u003c 0)\n+\t\tgoto cancel_put;\n+\tbvecq_pos_set(\u0026creq-\u003edispatch_cursor, \u0026creq-\u003eload_cursor);\n+\tbvecq_pos_set(\u0026creq-\u003ecollect_cursor, \u0026creq-\u003edispatch_cursor);\n+\n \t__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, \u0026creq-\u003eflags);\n \ttrace_netfs_copy2cache(rreq, creq);\n \ttrace_netfs_write(creq, netfs_write_trace_copy_to_cache);\n@@ -143,6 +168,14 @@ void netfs_pgpriv2_copy_to_cache(struct netfs_io_request *rreq, struct folio *fo\n \t\treturn;\n \t}\n \n+\tif (!creq-\u003espare) {\n+\t\tcreq-\u003espare = bvecq_alloc_one(BVECQ_POOL_SLOTS, creq-\u003egfp, false);\n+\t\tif (!creq-\u003espare) {\n+\t\t\tset_bit(NETFS_RREQ_CANCEL_CACHING, \u0026creq-\u003eflags);\n+\t\t\treturn;\n+\t\t}\n+\t}\n+\n \ttrace_netfs_folio(folio, netfs_folio_trace_pgpriv2_copy);\n \tnetfs_pgpriv2_copy_folio(creq, folio);\n }\n@@ -173,16 +206,18 @@ void netfs_pgpriv2_end_copy_to_cache(struct netfs_io_request *rreq)\n */\n bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)\n {\n-\tstruct bvecq *bq = creq-\u003ebuffer.tail;\n-\tunsigned int slot = creq-\u003ebuffer.first_tail_slot;\n+\tstruct bvecq *bq = creq-\u003ecollect_cursor.bvecq;\n+\tunsigned int slot;\n \tuoff_t collected_to = creq-\u003ecollected_to;\n \tbool made_progress = false;\n \n+\tslot = creq-\u003ecollect_cursor.slot;\n \twhile (!bvecq_acquire_slot(bq, slot)) {\n-\t\tbq = rolling_buffer_delete_spent(\u0026creq-\u003ebuffer);\n-\t\tif (!bq)\n-\t\t\treturn false;\n-\t\tslot = 0;\n+\t\tcreq-\u003ecollect_cursor.slot = slot;\n+\t\tif (!bvecq_delete_spent(\u0026creq-\u003ecollect_cursor))\n+\t\t\tgoto out;\n+\t\tbq = creq-\u003ecollect_cursor.bvecq;\n+\t\tslot = creq-\u003ecollect_cursor.slot;\n \t}\n \n \tfor (;;) {\n@@ -213,25 +248,25 @@ bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)\n \t\tcreq-\u003ecleaned_to = fpos + fsize;\n \t\tmade_progress = true;\n \n-\t\t/* Clean up the head bq. If we clear an entire bq, then\n-\t\t * we can get rid of it provided it's not also the tail bq\n-\t\t * being filled by the issuer.\n+\t\t/* Clean up the head segment. If we clear an entire segment,\n+\t\t * then we can get rid of it provided it's not also the tail\n+\t\t * segment being filled by the issuer.\n \t\t */\n \t\tbq-\u003ebv[slot].bv_page = NULL;\n \t\tslot++;\n \t\twhile (!bvecq_acquire_slot(bq, slot)) {\n-\t\t\tbq = rolling_buffer_delete_spent(\u0026creq-\u003ebuffer);\n-\t\t\tif (!bq)\n-\t\t\t\tgoto done;\n-\t\t\tslot = 0;\n+\t\t\tcreq-\u003ecollect_cursor.slot = slot;\n+\t\t\tif (!bvecq_delete_spent(\u0026creq-\u003ecollect_cursor))\n+\t\t\t\tgoto out;\n+\t\t\tbq = creq-\u003ecollect_cursor.bvecq;\n+\t\t\tslot = creq-\u003ecollect_cursor.slot;\n \t\t}\n \n \t\tif (fpos + fsize \u003e= collected_to)\n \t\t\tbreak;\n \t}\n \n-\tcreq-\u003ebuffer.tail = bq;\n-done:\n-\tcreq-\u003ebuffer.first_tail_slot = slot;\n+\tcreq-\u003ecollect_cursor.slot = slot;\n+out:\n \treturn made_progress;\n }\ndiff --git a/fs/netfs/read_retry.c b/fs/netfs/read_retry.c\nindex 142c3fb8dab17..7490b9ee1bf7c 100644\n--- a/fs/netfs/read_retry.c\n+++ b/fs/netfs/read_retry.c\n@@ -12,6 +12,10 @@\n static void netfs_reissue_read(struct netfs_io_request *rreq,\n \t\t\t struct netfs_io_subrequest *subreq)\n {\n+\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\n+\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n+\tiov_iter_advance(\u0026subreq-\u003eio_iter, subreq-\u003etransferred);\n+\n \tsubreq-\u003eerror = 0;\n \t__clear_bit(NETFS_SREQ_MADE_PROGRESS, \u0026subreq-\u003eflags);\n \t__set_bit(NETFS_SREQ_IN_PROGRESS, \u0026subreq-\u003eflags);\n@@ -27,6 +31,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n {\n \tstruct netfs_io_subrequest *subreq;\n \tstruct netfs_io_stream *stream = \u0026rreq-\u003eio_streams[0];\n+\tstruct bvecq_pos dispatch_cursor = {};\n \tstruct list_head *next;\n \n \t_enter(\"R=%x\", rreq-\u003edebug_id);\n@@ -46,9 +51,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \t\t\tif (test_bit(NETFS_SREQ_FAILED, \u0026subreq-\u003eflags))\n \t\t\t\tbreak;\n \t\t\tif (__test_and_clear_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags)) {\n-\t\t\t\t__clear_bit(NETFS_SREQ_MADE_PROGRESS, \u0026subreq-\u003eflags);\n \t\t\t\tsubreq-\u003eretry_count++;\n-\t\t\t\tnetfs_reset_iter(subreq);\n \t\t\t\tnetfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);\n \t\t\t\tnetfs_reissue_read(rreq, subreq);\n \t\t\t}\n@@ -74,11 +77,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \n \tdo {\n \t\tstruct netfs_io_subrequest *from, *to, *tmp;\n-\t\tstruct iov_iter source;\n \t\tuoff_t start, len;\n \t\tsize_t part;\n \t\tbool boundary = false, subreq_superfluous = false;\n \n+\t\tbvecq_pos_unset(\u0026dispatch_cursor);\n+\n \t\t/* Go through the subreqs and find the next span of contiguous\n \t\t * buffer that we then rejig (cifs, for example, needs the\n \t\t * rsize renegotiating) and reissue.\n@@ -105,7 +109,8 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \t\t\t\tbreak;\n \n \t\t\tsubreq = list_entry(next, struct netfs_io_subrequest, rreq_link);\n-\t\t\tif (subreq-\u003estart + subreq-\u003etransferred != start + len ||\n+\t\t\tif (subreq-\u003estart != start + len ||\n+\t\t\t subreq-\u003etransferred \u003e 0 ||\n \t\t\t test_bit(NETFS_SREQ_BOUNDARY, \u0026subreq-\u003eflags) ||\n \t\t\t !test_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags))\n \t\t\t\tbreak;\n@@ -118,11 +123,14 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \t\t/* Determine the set of buffers we're going to use. Each\n \t\t * subreq gets a subset of a single overall contiguous buffer.\n \t\t */\n-\t\tnetfs_reset_iter(from);\n-\t\tsource = from-\u003eio_iter;\n-\t\tsource.count = len;\n+\t\tbvecq_pos_transfer(\u0026dispatch_cursor, \u0026from-\u003eio_buffer);\n+\t\tbvecq_pos_advance(\u0026dispatch_cursor, from-\u003etransferred);\n+\t\tfrom-\u003etransferred = 0;\n \n-\t\t/* Work through the sublist. */\n+\t\t/* Work through the sublist. The chain of buffers we're going\n+\t\t * to fill is attached to dispatch_cursor and we need to read\n+\t\t * 'len' amount of data from 'start'.\n+\t\t */\n \t\tsubreq = from;\n \t\tlist_for_each_entry_from(subreq, \u0026stream-\u003esubrequests, rreq_link) {\n \t\t\tif (!len) {\n@@ -130,16 +138,21 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \t\t\t\tbreak;\n \t\t\t}\n \t\t\tsubreq-\u003esource\t= NETFS_DOWNLOAD_FROM_SERVER;\n-\t\t\tsubreq-\u003estart\t= start - subreq-\u003etransferred;\n-\t\t\tsubreq-\u003elen\t= len + subreq-\u003etransferred;\n+\t\t\tsubreq-\u003estart\t= start;\n+\t\t\tsubreq-\u003elen\t= len;\n \t\t\t__clear_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags);\n \t\t\t__clear_bit(NETFS_SREQ_MADE_PROGRESS, \u0026subreq-\u003eflags);\n \t\t\tsubreq-\u003eretry_count++;\n+\t\t\tsubreq-\u003etransferred = 0;\n+\n+\t\t\tbvecq_pos_unset(\u0026subreq-\u003eio_buffer);\n+\t\t\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026dispatch_cursor);\n \n \t\t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_retry);\n \n \t\t\t/* Renegotiate max_len (rsize) */\n-\t\t\tstream-\u003esreq_max_len = subreq-\u003elen;\n+\t\t\tstream-\u003esreq_max_len = len;\n+\t\t\tstream-\u003esreq_max_segs = INT_MAX;\n \t\t\tif (rreq-\u003enetfs_ops-\u003eprepare_read \u0026\u0026\n \t\t\t rreq-\u003enetfs_ops-\u003eprepare_read(subreq) \u003c 0) {\n \t\t\t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_reprep_failed);\n@@ -147,13 +160,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \t\t\t\tgoto abandon;\n \t\t\t}\n \n-\t\t\tpart = umin(len, stream-\u003esreq_max_len);\n-\t\t\tif (unlikely(stream-\u003esreq_max_segs))\n-\t\t\t\tpart = netfs_limit_iter(\u0026source, 0, part, stream-\u003esreq_max_segs);\n-\t\t\tsubreq-\u003elen = subreq-\u003etransferred + part;\n-\t\t\tsubreq-\u003eio_iter = source;\n-\t\t\tiov_iter_truncate(\u0026subreq-\u003eio_iter, part);\n-\t\t\tiov_iter_advance(\u0026source, part);\n+\t\t\tpart = bvecq_slice(\u0026dispatch_cursor,\n+\t\t\t\t\t umin(len, stream-\u003esreq_max_len),\n+\t\t\t\t\t stream-\u003esreq_max_segs,\n+\t\t\t\t\t \u0026subreq-\u003enr_segs);\n+\t\t\tsubreq-\u003elen = part;\n+\n \t\t\tlen -= part;\n \t\t\tstart += part;\n \t\t\tif (!len) {\n@@ -216,9 +228,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \t\t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_retry);\n \n \t\t\tstream-\u003esreq_max_len\t= umin(len, rreq-\u003ersize);\n-\t\t\tstream-\u003esreq_max_segs\t= 0;\n-\t\t\tif (unlikely(stream-\u003esreq_max_segs))\n-\t\t\t\tpart = netfs_limit_iter(\u0026source, 0, part, stream-\u003esreq_max_segs);\n+\t\t\tstream-\u003esreq_max_segs\t= INT_MAX;\n \n \t\t\tnetfs_stat(\u0026netfs_n_rh_download);\n \t\t\tif (rreq-\u003enetfs_ops-\u003eprepare_read(subreq) \u003c 0) {\n@@ -227,11 +237,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \t\t\t\tgoto abandon;\n \t\t\t}\n \n-\t\t\tpart = umin(len, stream-\u003esreq_max_len);\n-\t\t\tsubreq-\u003elen = subreq-\u003etransferred + part;\n-\t\t\tsubreq-\u003eio_iter = source;\n-\t\t\tiov_iter_truncate(\u0026subreq-\u003eio_iter, part);\n-\t\t\tiov_iter_advance(\u0026source, part);\n+\t\t\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026dispatch_cursor);\n+\t\t\tpart = bvecq_slice(\u0026dispatch_cursor,\n+\t\t\t\t\t umin(len, stream-\u003esreq_max_len),\n+\t\t\t\t\t stream-\u003esreq_max_segs,\n+\t\t\t\t\t \u0026subreq-\u003enr_segs);\n+\t\t\tsubreq-\u003elen = part;\n \n \t\t\tlen -= part;\n \t\t\tstart += part;\n@@ -245,12 +256,14 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \n \t} while (!list_is_head(next, \u0026stream-\u003esubrequests));\n \n+out:\n+\tbvecq_pos_unset(\u0026dispatch_cursor);\n \treturn;\n \n \t/* If we hit an error, fail all remaining incomplete subrequests */\n abandon_after:\n \tif (list_is_last(\u0026subreq-\u003erreq_link, \u0026stream-\u003esubrequests))\n-\t\treturn;\n+\t\tgoto out;\n \tsubreq = list_next_entry(subreq, rreq_link);\n abandon:\n \tlist_for_each_entry_from(subreq, \u0026stream-\u003esubrequests, rreq_link) {\n@@ -261,6 +274,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n \t\t__set_bit(NETFS_SREQ_FAILED, \u0026subreq-\u003eflags);\n \t\t__clear_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags);\n \t}\n+\tgoto out;\n }\n \n /*\n@@ -300,26 +314,24 @@ void netfs_unlock_abandoned_read_pages(struct netfs_io_request *rreq)\n \tif (test_bit(NETFS_RREQ_NEED_PUT_RA_REFS, \u0026rreq-\u003eflags))\n \t\tnetfs_wait_for_put_ra_refs(rreq);\n \n-\tfor (p = rreq-\u003ebuffer.tail; p; p = p-\u003enext) {\n-\t\tfor (int slot = rreq-\u003ebuffer.first_tail_slot;\n-\t\t bvecq_acquire_slot(p, slot);\n-\t\t slot++) {\n-\t\t\tstruct folio *folio;\n+\tfor (p = rreq-\u003ecollect_cursor.bvecq; p; p = bvecq_next(p)) {\n+\t\tunsigned int nr_slots = bvecq_nr_slots_acquire(p);\n \n+\t\tfor (int slot = 0; slot \u003c nr_slots; slot++) {\n \t\t\tif (!p-\u003ebv[slot].bv_page)\n \t\t\t\tcontinue;\n \n-\t\t\tfolio = bvec_folio(\u0026p-\u003ebv[slot]);\n+\t\t\tstruct folio *folio = bvec_folio(\u0026p-\u003ebv[slot]);\n+\n \t\t\tnetfs_cancel_copy_to_cache(rreq, folio);\n \n \t\t\tif (folio == rreq-\u003eno_unlock_folio \u0026\u0026\n \t\t\t test_bit(NETFS_RREQ_NO_UNLOCK_FOLIO, \u0026rreq-\u003eflags)) {\n \t\t\t\t_debug(\"no unlock\");\n-\t\t\t} else {\n-\t\t\t\ttrace_netfs_folio(folio, netfs_folio_trace_abandon);\n-\t\t\t\tfolio_unlock(folio);\n+\t\t\t\tcontinue;\n \t\t\t}\n+\t\t\ttrace_netfs_folio(folio, netfs_folio_trace_abandon);\n+\t\t\tfolio_unlock(folio);\n \t\t}\n-\t\trreq-\u003ebuffer.first_tail_slot = 0;\n \t}\n }\ndiff --git a/fs/netfs/read_single.c b/fs/netfs/read_single.c\nindex b248e34bd0c86..c70941121de04 100644\n--- a/fs/netfs/read_single.c\n+++ b/fs/netfs/read_single.c\n@@ -101,7 +101,11 @@ static int netfs_single_dispatch_read(struct netfs_io_request *rreq)\n \n \tsubreq-\u003estart\t= 0;\n \tsubreq-\u003elen\t= rreq-\u003elen;\n-\tsubreq-\u003eio_iter\t= rreq-\u003ebuffer.iter;\n+\n+\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026rreq-\u003edispatch_cursor);\n+\n+\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\n+\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n \n \tnetfs_queue_read(rreq, subreq);\n \n@@ -174,6 +178,15 @@ ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_ite\n \tif (IS_ERR(rreq))\n \t\treturn PTR_ERR(rreq);\n \n+\tret = netfs_extract_iter(iter, rreq-\u003elen, INT_MAX, \u0026rreq-\u003edispatch_cursor.bvecq,\n+\t\t\t\t 0, rreq-\u003egfp);\n+\tif (ret \u003c 0)\n+\t\tgoto cleanup_free;\n+\tif (ret \u003c rreq-\u003elen) {\n+\t\tret = -EIO;\n+\t\tgoto cleanup_free;\n+\t}\n+\n \trreq-\u003eprogress_at = rreq-\u003elen;\n \n \tret = netfs_single_begin_cache_read(rreq, ictx);\n@@ -183,7 +196,6 @@ ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_ite\n \tnetfs_stat(\u0026netfs_n_rh_read_single);\n \ttrace_netfs_read(rreq, 0, rreq-\u003elen, netfs_read_trace_read_single);\n \n-\trreq-\u003ebuffer.iter = *iter;\n \tnetfs_single_dispatch_read(rreq);\n \n \tret = netfs_wait_for_read(rreq);\ndiff --git a/fs/netfs/rolling_buffer.c b/fs/netfs/rolling_buffer.c\ndeleted file mode 100644\nindex 66ce9add40122..0000000000000\n--- a/fs/netfs/rolling_buffer.c\n+++ /dev/null\n@@ -1,182 +0,0 @@\n-// SPDX-License-Identifier: GPL-2.0-or-later\n-/* Rolling buffer helpers\n- *\n- * Copyright (C) 2024 Red Hat, Inc. All Rights Reserved.\n- * Written by David Howells (dhowells@redhat.com)\n- */\n-\n-#include \u003clinux/bitops.h\u003e\n-#include \u003clinux/mempool.h\u003e\n-#include \u003clinux/pagemap.h\u003e\n-#include \u003clinux/rolling_buffer.h\u003e\n-#include \u003clinux/slab.h\u003e\n-#include \"internal.h\"\n-\n-/*\n- * Initialise a rolling buffer. We allocate an empty folio queue struct to so\n- * that the pointers can be independently driven by the producer and the\n- * consumer.\n- */\n-int rolling_buffer_init(struct rolling_buffer *roll, unsigned int direction,\n-\t\t\tgfp_t gfp, bool for_writeback)\n-{\n-\tstruct bvecq *bq;\n-\n-\troll-\u003efor_writeback = for_writeback;\n-\n-\tbq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, for_writeback);\n-\tif (!bq)\n-\t\treturn -ENOMEM;\n-\n-\troll-\u003ehead = bq;\n-\troll-\u003etail = bq;\n-\tiov_iter_bvec_queue(\u0026roll-\u003eiter, direction, bq, 0, 0, 0);\n-\treturn 0;\n-}\n-\n-/*\n- * Add another bvecq to a rolling buffer if there's no space left.\n- */\n-int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp)\n-{\n-\tstruct bvecq *bq, *head = roll-\u003ehead;\n-\n-\tif (!bvecq_is_full(head))\n-\t\treturn 0;\n-\n-\tbq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, roll-\u003efor_writeback);\n-\tif (!bq)\n-\t\treturn -ENOMEM;\n-\n-\troll-\u003ehead = bq;\n-\tif (bvecq_is_full(head)) {\n-\t\t/* Make sure we don't leave the master iterator pointing to a\n-\t\t * block that might get immediately consumed.\n-\t\t */\n-\t\tif (roll-\u003eiter.bvecq == head \u0026\u0026\n-\t\t roll-\u003eiter.bvecq_slot == head-\u003enr_slots) {\n-\t\t\troll-\u003eiter.bvecq = bq;\n-\t\t\troll-\u003eiter.bvecq_slot = 0;\n-\t\t}\n-\t}\n-\n-\t/* Make sure the initialisation is stored before the next pointer.\n-\t *\n-\t * [!] NOTE: After we set head-\u003enext, the consumer is at liberty to\n-\t * immediately delete the old head.\n-\t */\n-\tbvecq_append(head, bq);\n-\treturn 0;\n-}\n-\n-/*\n- * Decant the entire list of folios to read into a rolling buffer.\n- */\n-ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,\n-\t\t\t\t\t struct readahead_control *ractl,\n-\t\t\t\t\t gfp_t gfp)\n-{\n-\tstruct bvecq *bq;\n-\tsize_t loaded = 0;\n-\n-\twhile (ractl-\u003e_nr_pages - ractl-\u003e_batch_count \u003e 0) {\n-\t\tstruct page **pages;\n-\t\tunsigned int nr;\n-\n-\t\t/* Allocate a bvecq to put some folios into and attach it to\n-\t\t * the rolling buffer.\n-\t\t */\n-\t\tbq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, false);\n-\t\tif (!bq)\n-\t\t\tgoto nomem_unlock;\n-\t\tbq-\u003emem_type = BVECQ_MEM_EXTERNAL; /* Folio cleanup handled separately. */\n-\n-\t\tif (!roll-\u003etail)\n-\t\t\troll-\u003etail = bq;\n-\t\telse\n-\t\t\tbvecq_append(roll-\u003ehead, bq);\n-\t\troll-\u003ehead = bq;\n-\n-\t\t/* Get a bunch of folios and note their sizes. */\n-\t\tpages = (struct page **)(bq-\u003ebv + bq-\u003emax_slots);\n-\t\tpages -= bq-\u003emax_slots;\n-\t\tnr = __readahead_batch(ractl, pages, bq-\u003emax_slots);\n-\t\tif (WARN_ON_ONCE(!nr))\n-\t\t\tbreak;\n-\n-\t\tfor (int slot = 0; slot \u003c nr; slot++) {\n-\t\t\tstruct folio *folio = page_folio(pages[slot]);\n-\t\t\tsize_t len = folio_size(folio);\n-\n-\t\t\tbvec_set_folio(\u0026bq-\u003ebv[slot], folio, len, 0);\n-\t\t\tloaded += len;\n-\t\t\ttrace_netfs_folio(folio, netfs_folio_trace_read);\n-\t\t}\n-\n-\t\tbvecq_filled_to(bq, nr);\n-\t}\n-\n-\tWRITE_ONCE(roll-\u003eiter.count, loaded);\n-\tiov_iter_bvec_queue(\u0026roll-\u003eiter, ITER_DEST, roll-\u003etail, 0, 0, loaded);\n-\treturn loaded;\n-\n-nomem_unlock:\n-\tfor (bq = roll-\u003etail; bq; bq = bq-\u003enext) {\n-\t\tfor (int slot = 0; slot \u003c bq-\u003enr_slots; slot++) {\n-\t\t\tstruct folio *folio = bvec_folio(\u0026bq-\u003ebv[slot]);\n-\n-\t\t\tfolio_unlock(folio);\n-\t\t\tfolio_put(folio);\n-\t\t}\n-\t}\n-\trolling_buffer_clear(roll);\n-\troll-\u003ehead = NULL;\n-\troll-\u003etail = NULL;\n-\treturn -ENOMEM;\n-}\n-\n-/*\n- * Append a folio to the rolling buffer.\n- */\n-ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio,\n-\t\t\t gfp_t gfp)\n-{\n-\tssize_t size = folio_size(folio);\n-\tint slot;\n-\n-\tif (rolling_buffer_make_space(roll, gfp) \u003c 0)\n-\t\treturn -ENOMEM;\n-\n-\tslot = roll-\u003ehead-\u003enr_slots;\n-\tbvec_set_folio(\u0026roll-\u003ehead-\u003ebv[slot], folio, size, 0);\n-\tbvecq_filled_to(roll-\u003ehead, slot + 1);\n-\n-\tWRITE_ONCE(roll-\u003eiter.count, roll-\u003eiter.count + size);\n-\treturn size;\n-}\n-\n-/*\n- * Delete a spent buffer from a rolling queue and return the next in line. We\n- * don't return the last buffer to keep the pointers independent, but return\n- * NULL instead.\n- */\n-struct bvecq *rolling_buffer_delete_spent(struct rolling_buffer *roll)\n-{\n-\tstruct bvecq *spent = roll-\u003etail, *next = bvecq_next(spent);\n-\n-\tif (!next)\n-\t\treturn NULL;\n-\tnext-\u003eprev = NULL;\n-\troll-\u003etail = next;\n-\tspent-\u003enext = NULL;\n-\tbvecq_put(spent);\n-\treturn next;\n-}\n-\n-/*\n- * Clear out a rolling queue.\n- */\n-void rolling_buffer_clear(struct rolling_buffer *roll)\n-{\n-\tbvecq_put(roll-\u003etail);\n-}\ndiff --git a/fs/netfs/write_collect.c b/fs/netfs/write_collect.c\nindex 91b42820c8925..dcacbab254b92 100644\n--- a/fs/netfs/write_collect.c\n+++ b/fs/netfs/write_collect.c\n@@ -114,12 +114,12 @@ int netfs_folio_written_back(struct folio *folio)\n static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,\n \t\t\t\t\t unsigned int *notes)\n {\n-\tstruct bvecq *bq = wreq-\u003ebuffer.tail;\n-\tunsigned int slot = wreq-\u003ebuffer.first_tail_slot;\n+\tstruct bvecq *bq = wreq-\u003ecollect_cursor.bvecq;\n+\tunsigned int slot = wreq-\u003ecollect_cursor.slot;\n \tuoff_t collected_to = wreq-\u003ecollected_to;\n \n \tif (WARN_ON_ONCE(!bq)) {\n-\t\tpr_err(\"[!] Writeback unlock found empty rolling buffer!\\n\");\n+\t\tpr_err(\"[!] Writeback unlock found empty buffer!\\n\");\n \t\tnetfs_dump_request(wreq);\n \t\treturn;\n \t}\n@@ -130,19 +130,28 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,\n \t\treturn;\n \t}\n \n-\twhile (!bvecq_acquire_slot(bq, slot)) {\n-\t\tbq = rolling_buffer_delete_spent(\u0026wreq-\u003ebuffer);\n-\t\tif (!bq)\n-\t\t\treturn;\n-\t\tslot = 0;\n-\t}\n-\n \tfor (;;) {\n \t\tstruct folio *folio;\n \t\tstruct netfs_folio *finfo;\n \t\tuoff_t fpos, fend;\n \t\tsize_t fsize, flen;\n \n+\t\t/* Try to clean up the head of the queue if it appears to be\n+\t\t * used up, but we need to be very careful - the cleanup can\n+\t\t * catch the dispatcher, which could lead to us having nothing\n+\t\t * left in the queue, causing the front and back pointers to\n+\t\t * end up on different tracks. To avoid this, we must always\n+\t\t * keep at least one segment in the queue.\n+\t\t */\n+\t\tif (!bvecq_acquire_slot(bq, slot)) {\n+\t\t\twreq-\u003ecollect_cursor.slot = slot;\n+\t\t\tif (!bvecq_delete_spent(\u0026wreq-\u003ecollect_cursor))\n+\t\t\t\treturn;\n+\t\t\tbq = wreq-\u003ecollect_cursor.bvecq;\n+\t\t\tslot = wreq-\u003ecollect_cursor.slot;\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tfolio = bvec_folio(\u0026bq-\u003ebv[slot]);\n \t\tif (WARN_ONCE(!folio_test_writeback(folio),\n \t\t\t \"R=%08x: folio %lx is not under writeback\\n\",\n@@ -166,26 +175,13 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,\n \t\twreq-\u003ecleaned_to = fpos + fsize;\n \t\t*notes |= MADE_PROGRESS;\n \n-\t\t/* Clean up the head bq. If we clear an entire bq, then\n-\t\t * we can get rid of it provided it's not also the tail bq\n-\t\t * being filled by the issuer.\n-\t\t */\n \t\tbq-\u003ebv[slot].bv_page = NULL;\n \t\tslot++;\n-\t\twhile (!bvecq_acquire_slot(bq, slot)) {\n-\t\t\tbq = rolling_buffer_delete_spent(\u0026wreq-\u003ebuffer);\n-\t\t\tif (!bq)\n-\t\t\t\tgoto done;\n-\t\t\tslot = 0;\n-\t\t}\n-\n \t\tif (fpos + fsize \u003e= collected_to)\n \t\t\tbreak;\n \t}\n \n-\twreq-\u003ebuffer.tail = bq;\n-done:\n-\twreq-\u003ebuffer.first_tail_slot = slot;\n+\twreq-\u003ecollect_cursor.slot = slot;\n }\n \n /*\n@@ -230,7 +226,8 @@ static void netfs_collect_write_results(struct netfs_io_request *wreq)\n \ttrace_netfs_rreq(wreq, netfs_rreq_trace_collect);\n \n reassess_streams:\n-\tissued_to = atomic64_read(\u0026wreq-\u003eissued_to);\n+\t/* Order reading the issued_to point before reading the queue it refers to. */\n+\tissued_to = atomic64_read_acquire(\u0026wreq-\u003eissued_to);\n \tsmp_rmb();\n \tcollected_to = ULLONG_MAX;\n \tif (wreq-\u003eorigin == NETFS_WRITEBACK ||\n@@ -560,8 +557,12 @@ void netfs_write_subrequest_terminated(void *_op, ssize_t transferred_or_error)\n \t\t\t * data is tracked.\n \t\t\t */\n \t\t\tnetfs_stat(\u0026netfs_n_wh_write_failed);\n-\t\t\tif (test_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags))\n-\t\t\t\tbreak;\n+\t\t\tif (test_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags)) {\n+\t\t\t\t/* We don't retry failed cache writes. */\n+\t\t\t\t__clear_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags);\n+\t\t\t\tif (!subreq-\u003eerror)\n+\t\t\t\t\tsubreq-\u003eerror = -ENOBUFS;\n+\t\t\t}\n \n \t\t\ttrace_netfs_failure(wreq, subreq, transferred_or_error, netfs_fail_write);\n \t\t\t__set_bit(NETFS_SREQ_CANCELLED, \u0026subreq-\u003eflags);\ndiff --git a/fs/netfs/write_issue.c b/fs/netfs/write_issue.c\nindex f0f4786666512..8fab5cf00d7b6 100644\n--- a/fs/netfs/write_issue.c\n+++ b/fs/netfs/write_issue.c\n@@ -107,10 +107,6 @@ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,\n \tictx = netfs_inode(wreq-\u003einode);\n \tif (is_cacheable)\n \t\tfscache_begin_write_operation(\u0026wreq-\u003ecache_resources, netfs_i_cookie(ictx));\n-\tif (rolling_buffer_init(\u0026wreq-\u003ebuffer, ITER_SOURCE, wreq-\u003egfp,\n-\t\t\t\t(origin == NETFS_WRITEBACK ||\n-\t\t\t\t origin == NETFS_WRITEBACK_SINGLE)) \u003c 0)\n-\t\tgoto nomem;\n \n \twreq-\u003ecleaned_to = wreq-\u003estart;\n \tif (wreq-\u003ecache_resources.dio_size \u003e 1)\n@@ -135,9 +131,6 @@ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,\n \t}\n \n \treturn wreq;\n-nomem:\n-\tnetfs_put_failed_request(wreq);\n-\treturn ERR_PTR(-ENOMEM);\n }\n \n /**\n@@ -163,22 +156,14 @@ void netfs_prepare_write(struct netfs_io_request *wreq,\n \t\t\t uoff_t start)\n {\n \tstruct netfs_io_subrequest *subreq;\n-\tstruct iov_iter *wreq_iter = \u0026wreq-\u003ebuffer.iter;\n-\n-\t/* Make sure we don't point the iterator at a used-up bvecq struct\n-\t * being used as a placeholder to prevent the queue from collapsing.\n-\t * In such a case, extend the queue.\n-\t */\n-\tif (iov_iter_is_bvecq(wreq_iter) \u0026\u0026\n-\t !bvecq_acquire_slot(wreq_iter-\u003ebvecq, wreq_iter-\u003ebvecq_slot))\n-\t\trolling_buffer_make_space(\u0026wreq-\u003ebuffer, wreq-\u003egfp);\n \n \tsubreq = netfs_alloc_subrequest(wreq, stream-\u003esource);\n \tif (!subreq)\n \t\treturn;\n \tsubreq-\u003estart\t\t= start;\n \tsubreq-\u003estream_nr\t= stream-\u003estream_nr;\n-\tsubreq-\u003eio_iter\t\t= *wreq_iter;\n+\n+\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026wreq-\u003edispatch_cursor);\n \n \t_enter(\"R=%x[%x]\", wreq-\u003edebug_id, subreq-\u003edebug_index);\n \n@@ -259,15 +244,14 @@ static void netfs_do_issue_write(struct netfs_io_stream *stream,\n }\n \n void netfs_reissue_write(struct netfs_io_stream *stream,\n-\t\t\t struct netfs_io_subrequest *subreq,\n-\t\t\t struct iov_iter *source)\n+\t\t\t struct netfs_io_subrequest *subreq)\n {\n-\tsize_t size = subreq-\u003elen - subreq-\u003etransferred;\n-\n \t// TODO: Use encrypted buffer\n-\tsubreq-\u003eio_iter = *source;\n-\tiov_iter_advance(source, size);\n-\tiov_iter_truncate(\u0026subreq-\u003eio_iter, size);\n+\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_SOURCE,\n+\t\t\t subreq-\u003eio_buffer.bvecq, subreq-\u003eio_buffer.slot,\n+\t\t\t subreq-\u003eio_buffer.offset,\n+\t\t\t subreq-\u003elen);\n+\tiov_iter_advance(\u0026subreq-\u003eio_iter, subreq-\u003etransferred);\n \n \tsubreq-\u003eretry_count++;\n \tsubreq-\u003eerror = 0;\n@@ -285,8 +269,12 @@ void netfs_issue_write(struct netfs_io_request *wreq,\n \tif (!subreq)\n \t\treturn;\n \n+\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_SOURCE,\n+\t\t\t subreq-\u003eio_buffer.bvecq, subreq-\u003eio_buffer.slot,\n+\t\t\t subreq-\u003eio_buffer.offset,\n+\t\t\t subreq-\u003elen);\n+\n \tstream-\u003econstruct = NULL;\n-\tsubreq-\u003eio_iter.count = subreq-\u003elen;\n \tnetfs_do_issue_write(stream, subreq);\n }\n \n@@ -323,7 +311,6 @@ size_t netfs_advance_write(struct netfs_io_request *wreq,\n \t_debug(\"part %zx/%zx %zx/%zx\", subreq-\u003elen, stream-\u003esreq_max_len, part, len);\n \tsubreq-\u003elen += part;\n \tsubreq-\u003enr_segs++;\n-\tstream-\u003esubmit_extendable_to -= part;\n \n \tif (subreq-\u003elen \u003e= stream-\u003esreq_max_len ||\n \t subreq-\u003enr_segs \u003e= stream-\u003esreq_max_segs ||\n@@ -347,7 +334,8 @@ static int netfs_write_folio(struct netfs_io_request *wreq,\n \tstruct netfs_io_stream *stream;\n \tstruct netfs_group *fgroup; /* TODO: Use this with ceph */\n \tstruct netfs_folio *finfo;\n-\tsize_t iter_off = 0;\n+\tstruct bvecq *queue = wreq-\u003eload_cursor.bvecq;\n+\tunsigned int slot;\n \tsize_t fsize = folio_size(folio), flen = fsize, foff = 0;\n \tuoff_t fpos = folio_pos(folio), i_size;\n \tbool to_eof = false, streamw = false;\n@@ -355,12 +343,20 @@ static int netfs_write_folio(struct netfs_io_request *wreq,\n \n \t_enter(\"\");\n \n-\tif (rolling_buffer_make_space(\u0026wreq-\u003ebuffer, wreq-\u003egfp) \u003c 0)\n-\t\treturn -ENOMEM;\n+\tif (!wreq-\u003espare) {\n+\t\twreq-\u003espare = bvecq_alloc_one(BVECQ_POOL_SLOTS, wreq-\u003egfp, true);\n+\t\tif (!wreq-\u003espare)\n+\t\t\treturn -ENOMEM;\n+\t}\n \n-\t/* netfs_perform_write() may shift i_size around the page or from out\n-\t * of the page to beyond it, but cannot move i_size into or through the\n-\t * page since we have it locked.\n+\t/* netfs_perform_write() may shift i_size around the folio or from out\n+\t * of the folio to beyond it, but cannot move i_size into or through\n+\t * the folio since we have it locked.\n+\t *\n+\t * Truncate could in theory move i_size into or before the folio, but\n+\t * it should take steps to prevent writeback from happening\n+\t * concurrently and should wait for any in-progress writebacks before\n+\t * proceeding.\n \t */\n \ti_size = i_size_read(wreq-\u003einode);\n \n@@ -452,8 +448,29 @@ static int netfs_write_folio(struct netfs_io_request *wreq,\n \t\ttrace_netfs_folio(folio, netfs_folio_trace_store_plus);\n \t}\n \n+\t/* Institute a new bvec queue segment if the current one is full or if\n+\t * we encounter a discontiguity. The discontiguity break is important\n+\t * when it comes to bulk unlocking folios by file range.\n+\t */\n+\tif (bvecq_is_full(queue) ||\n+\t (fpos != wreq-\u003elast_end \u0026\u0026 wreq-\u003elast_end \u003e 0)) {\n+\t\tbvecq_buffer_append(\u0026wreq-\u003eload_cursor, wreq-\u003espare);\n+\t\twreq-\u003espare = NULL;\n+\n+\t\tqueue = wreq-\u003eload_cursor.bvecq;\n+\t\tbvecq_pos_move(\u0026wreq-\u003edispatch_cursor, queue);\n+\t\twreq-\u003edispatch_cursor.slot = 0;\n+\t}\n+\n \t/* Attach the folio to the rolling buffer. */\n-\trolling_buffer_append(\u0026wreq-\u003ebuffer, folio, wreq-\u003egfp);\n+\tslot = queue-\u003enr_slots;\n+\tbvec_set_folio(\u0026queue-\u003ebv[slot], folio, fsize, 0);\n+\ttrace_netfs_bv_slot(queue, slot);\n+\tslot++;\n+\tbvecq_filled_to(queue, slot);\n+\twreq-\u003eload_cursor.slot = slot;\n+\twreq-\u003eload_cursor.offset = 0;\n+\twreq-\u003elast_end = fpos + fsize;\n \n \t/* Move the submission point forward to allow for write-streaming data\n \t * not starting at the front of the page. We don't do write-streaming\n@@ -462,10 +479,19 @@ static int netfs_write_folio(struct netfs_io_request *wreq,\n \t * Also skip uploading for data that's been read and just needs copying\n \t * to the cache.\n \t */\n+\tbvecq_pos_nudge(\u0026wreq-\u003edispatch_cursor);\n+\n \tfor (int s = 0; s \u003c NR_IO_STREAMS; s++) {\n+\t\tsize_t soff = foff, slen = flen, alignment = 1;\n+\n \t\tstream = \u0026wreq-\u003eio_streams[s];\n-\t\tstream-\u003esubmit_off = foff;\n-\t\tstream-\u003esubmit_len = flen;\n+\t\tif (stream-\u003esource == NETFS_WRITE_TO_CACHE)\n+\t\t\talignment = wreq-\u003ecache_resources.dio_size;\n+\t\tstream = \u0026wreq-\u003eio_streams[s];\n+\t\tstream-\u003esubmit_off = round_down(soff, alignment);\n+\t\tslen += foff - stream-\u003esubmit_off;\n+\t\tstream-\u003esubmit_len = round_up(slen, alignment);\n+\n \t\tif (!stream-\u003eavail ||\n \t\t (stream-\u003esource == NETFS_WRITE_TO_CACHE \u0026\u0026 streamw) ||\n \t\t (stream-\u003esource == NETFS_UPLOAD_TO_SERVER \u0026\u0026\n@@ -499,14 +525,10 @@ static int netfs_write_folio(struct netfs_io_request *wreq,\n \t\t\tbreak;\n \t\tstream = \u0026wreq-\u003eio_streams[choose_s];\n \n-\t\t/* Advance the iterator(s). */\n-\t\tif (stream-\u003esubmit_off \u003e iter_off) {\n-\t\t\trolling_buffer_advance(\u0026wreq-\u003ebuffer, stream-\u003esubmit_off - iter_off);\n-\t\t\titer_off = stream-\u003esubmit_off;\n-\t\t}\n+\t\t/* Advance the cursor. */\n+\t\twreq-\u003edispatch_cursor.offset = stream-\u003esubmit_off;\n \n \t\tatomic64_set(\u0026wreq-\u003eissued_to, fpos + stream-\u003esubmit_off);\n-\t\tstream-\u003esubmit_extendable_to = fsize - stream-\u003esubmit_off;\n \t\tpart = netfs_advance_write(wreq, stream, fpos + stream-\u003esubmit_off,\n \t\t\t\t\t stream-\u003esubmit_len, to_eof);\n \t\tstream-\u003esubmit_off += part;\n@@ -518,9 +540,9 @@ static int netfs_write_folio(struct netfs_io_request *wreq,\n \t\t\tdebug = true;\n \t}\n \n-\tif (fsize \u003e iter_off)\n-\t\trolling_buffer_advance(\u0026wreq-\u003ebuffer, fsize - iter_off);\n-\tatomic64_set(\u0026wreq-\u003eissued_to, fpos + fsize);\n+\tbvecq_pos_step(\u0026wreq-\u003edispatch_cursor);\n+\t/* Order loading the queue before updating the issue_to point */\n+\tatomic64_set_release(\u0026wreq-\u003eissued_to, fpos + fsize);\n \n \tif (!debug)\n \t\tkdebug(\"R=%x: No submit\", wreq-\u003edebug_id);\n@@ -581,6 +603,11 @@ int netfs_writepages(struct address_space *mapping,\n \t\tgoto couldnt_start;\n \t}\n \n+\tif (bvecq_buffer_init(\u0026wreq-\u003eload_cursor, wreq-\u003egfp, true) \u003c 0)\n+\t\tgoto nomem;\n+\tbvecq_pos_set(\u0026wreq-\u003edispatch_cursor, \u0026wreq-\u003eload_cursor);\n+\tbvecq_pos_set(\u0026wreq-\u003ecollect_cursor, \u0026wreq-\u003edispatch_cursor);\n+\n \t__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, \u0026wreq-\u003eflags);\n \ttrace_netfs_write(wreq, netfs_write_trace_writeback);\n \tnetfs_stat(\u0026netfs_n_wh_writepages);\n@@ -605,12 +632,17 @@ int netfs_writepages(struct address_space *mapping,\n \t} while ((folio = writeback_iter(mapping, wbc, folio, \u0026error)));\n \n \tnetfs_end_issue_write(wreq);\n+\tbvecq_pos_unset(\u0026wreq-\u003eload_cursor);\n+\tbvecq_pos_unset(\u0026wreq-\u003edispatch_cursor);\n \tnetfs_wake_collector(wreq);\n \n \tnetfs_put_request(wreq, netfs_rreq_trace_put_return);\n \t_leave(\" = %d\", error);\n \treturn error;\n \n+nomem:\n+\terror = -ENOMEM;\n+\tnetfs_put_failed_request(wreq);\n couldnt_start:\n \tif (error == -ENOMEM) {\n \t\tfolio_redirty_for_writepage(wbc, folio);\n@@ -631,23 +663,28 @@ EXPORT_SYMBOL(netfs_writepages);\n * netfs_writeback_single - Write back a monolithic payload\n * @mapping: The mapping to write from\n * @wbc: Hints from the VM\n- * @iter: Data to write.\n+ * @iter: Buffer to write from\n+ * @len: Amount to write from buffer\n *\n * Write a monolithic, non-pagecache object back to the server and/or the\n- * cache. The caller must explicitly set NETFS_RREQ_UPLOAD_TO_SERVER when\n- * initialising the request if it wants the data to be written to the server\n- * (for AFS directories and symlinks, this is not possible; things like mkdir,\n- * symlink, rmdir and unlink must be used instead).\n+ * cache. There's a maximum of one subrequest per stream. The buffer should be\n+ * rounded out sufficiently that it can accommodate cache DIO rounding.\n+ *\n+ * The caller must explicitly set NETFS_RREQ_UPLOAD_TO_SERVER when initialising\n+ * the request if it wants the data to be written to the server (for AFS\n+ * directories and symlinks, this is not possible; things like mkdir, symlink,\n+ * rmdir and unlink must be used instead).\n *\n * Return: 0 if successful; 1 if skipped due to lock conflict and WB_SYNC_NONE;\n * or a negative error code.\n */\n int netfs_writeback_single(struct address_space *mapping,\n \t\t\t struct writeback_control *wbc,\n-\t\t\t struct iov_iter *iter)\n+\t\t\t struct iov_iter *iter, size_t len)\n {\n \tstruct netfs_io_request *wreq;\n \tstruct netfs_inode *ictx = netfs_inode(mapping-\u003ehost);\n+\tsize_t clen;\n \tint ret;\n \n \tif (!netfs_wb_begin(ictx, wbc-\u003esync_mode == WB_SYNC_NONE)) {\n@@ -661,10 +698,27 @@ int netfs_writeback_single(struct address_space *mapping,\n \t\tret = PTR_ERR(wreq);\n \t\tgoto couldnt_start;\n \t}\n+\twreq-\u003elen = len;\n+\tclen = len;\n+\n+\tif (wreq-\u003ecache_resources.dio_size \u003e 1) {\n+\t\tclen = round_up(len, wreq-\u003ecache_resources.dio_size);\n+\t\tif (clen \u003e iov_iter_count(iter)) {\n+\t\t\tret = -EIO;\n+\t\t\tgoto cleanup_free;\n+\t\t}\n+\t}\n \n-\twreq-\u003ebuffer.iter = *iter;\n-\twreq-\u003elen = iov_iter_count(iter);\n-\twreq-\u003esubmitted = wreq-\u003elen;\n+\tret = netfs_extract_iter(iter, clen, INT_MAX, \u0026wreq-\u003edispatch_cursor.bvecq,\n+\t\t\t\t 0, wreq-\u003egfp);\n+\tif (ret \u003c 0)\n+\t\tgoto cleanup_free;\n+\tif (ret \u003c clen) {\n+\t\tret = -EIO;\n+\t\tgoto cleanup_free;\n+\t}\n+\n+\tbvecq_pos_set(\u0026wreq-\u003ecollect_cursor, \u0026wreq-\u003edispatch_cursor);\n \n \t__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, \u0026wreq-\u003eflags);\n \ttrace_netfs_write(wreq, netfs_write_trace_writeback_single);\n@@ -685,12 +739,14 @@ int netfs_writeback_single(struct address_space *mapping,\n \n \t\tsubreq = stream-\u003econstruct;\n \t\tsubreq-\u003elen = wreq-\u003elen;\n+\t\tif (stream-\u003esource == NETFS_WRITE_TO_CACHE)\n+\t\t\tsubreq-\u003elen = clen;\n \t\tstream-\u003esubmit_len = subreq-\u003elen;\n-\t\tstream-\u003esubmit_extendable_to = round_up(wreq-\u003elen, PAGE_SIZE);\n \n \t\tnetfs_issue_write(wreq, stream);\n \t}\n \n+\twreq-\u003esubmitted = wreq-\u003elen;\n \tnetfs_all_subreqs_queued(wreq);\n \tnetfs_wake_collector(wreq);\n \n@@ -705,6 +761,8 @@ int netfs_writeback_single(struct address_space *mapping,\n \t_leave(\" = %d\", ret);\n \treturn ret;\n \n+cleanup_free:\n+\tnetfs_put_failed_request(wreq);\n couldnt_start:\n \tnetfs_wb_end(ictx);\n \t_leave(\" = %d\", ret);\ndiff --git a/fs/netfs/write_retry.c b/fs/netfs/write_retry.c\nindex 2f20577563e14..235e75eb10914 100644\n--- a/fs/netfs/write_retry.c\n+++ b/fs/netfs/write_retry.c\n@@ -17,15 +17,17 @@\n static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \t\t\t\t struct netfs_io_stream *stream)\n {\n+\tstruct bvecq_pos dispatch_cursor = {};\n \tstruct list_head *next;\n \n \t_enter(\"R=%x[%x:]\", wreq-\u003edebug_id, stream-\u003estream_nr);\n \n \tif (list_empty(\u0026stream-\u003esubrequests))\n \t\treturn;\n+\tif (WARN_ON_ONCE(stream-\u003esource != NETFS_UPLOAD_TO_SERVER))\n+\t\treturn; /* Shouldn't be retrying cache writes. */\n \n-\tif (stream-\u003esource == NETFS_UPLOAD_TO_SERVER \u0026\u0026\n-\t wreq-\u003enetfs_ops-\u003eretry_request)\n+\tif (wreq-\u003enetfs_ops-\u003eretry_request)\n \t\twreq-\u003enetfs_ops-\u003eretry_request(wreq, stream);\n \n \tif (unlikely(stream-\u003efailed))\n@@ -39,12 +41,8 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \t\t\tif (test_bit(NETFS_SREQ_FAILED, \u0026subreq-\u003eflags))\n \t\t\t\tbreak;\n \t\t\tif (__test_and_clear_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags)) {\n-\t\t\t\tstruct iov_iter source;\n-\n-\t\t\t\tnetfs_reset_iter(subreq);\n-\t\t\t\tsource = subreq-\u003eio_iter;\n \t\t\t\tnetfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);\n-\t\t\t\tnetfs_reissue_write(stream, subreq, \u0026source);\n+\t\t\t\tnetfs_reissue_write(stream, subreq);\n \t\t\t}\n \t\t}\n \t\treturn;\n@@ -54,11 +52,12 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \n \tdo {\n \t\tstruct netfs_io_subrequest *subreq = NULL, *from, *to, *tmp;\n-\t\tstruct iov_iter source;\n \t\tuoff_t start, len;\n \t\tsize_t part;\n \t\tbool boundary = false;\n \n+\t\tbvecq_pos_unset(\u0026dispatch_cursor);\n+\n \t\t/* Go through the stream and find the next span of contiguous\n \t\t * data that we then rejig (cifs, for example, needs the wsize\n \t\t * renegotiating) and reissue.\n@@ -70,7 +69,7 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \n \t\tif (test_bit(NETFS_SREQ_FAILED, \u0026from-\u003eflags) ||\n \t\t !test_bit(NETFS_SREQ_NEED_RETRY, \u0026from-\u003eflags))\n-\t\t\treturn;\n+\t\t\tgoto out;\n \n \t\tfor (;;) {\n \t\t\t/* Read pointer to subreq before reading subreq state. */\n@@ -79,7 +78,8 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \t\t\t\tbreak;\n \n \t\t\tsubreq = list_entry(next, struct netfs_io_subrequest, rreq_link);\n-\t\t\tif (subreq-\u003estart + subreq-\u003etransferred != start + len ||\n+\t\t\tif (subreq-\u003estart != start + len ||\n+\t\t\t subreq-\u003etransferred \u003e 0 ||\n \t\t\t test_bit(NETFS_SREQ_BOUNDARY, \u0026subreq-\u003eflags) ||\n \t\t\t !test_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags))\n \t\t\t\tbreak;\n@@ -90,11 +90,13 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \t\t/* Determine the set of buffers we're going to use. Each\n \t\t * subreq gets a subset of a single overall contiguous buffer.\n \t\t */\n-\t\tnetfs_reset_iter(from);\n-\t\tsource = from-\u003eio_iter;\n-\t\tsource.count = len;\n+\t\tbvecq_pos_transfer(\u0026dispatch_cursor, \u0026from-\u003eio_buffer);\n+\t\tbvecq_pos_advance(\u0026dispatch_cursor, from-\u003etransferred);\n \n-\t\t/* Work through the sublist. */\n+\t\t/* Work through the sublist. The chain of buffers we're going\n+\t\t * to fill is attached to dispatch_cursor and we need to read\n+\t\t * 'len' amount of data from 'start'.\n+\t\t */\n \t\tsubreq = from;\n \t\tlist_for_each_entry_from(subreq, \u0026stream-\u003esubrequests, rreq_link) {\n \t\t\tif (!len)\n@@ -104,16 +106,22 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \t\t\tsubreq-\u003elen\t= len;\n \t\t\t__clear_bit(NETFS_SREQ_NEED_RETRY, \u0026subreq-\u003eflags);\n \t\t\ttrace_netfs_sreq(subreq, netfs_sreq_trace_retry);\n+\t\t\tsubreq-\u003etransferred = 0;\n+\n+\t\t\tbvecq_pos_unset(\u0026subreq-\u003eio_buffer);\n \n \t\t\t/* Renegotiate max_len (wsize) */\n \t\t\tstream-\u003esreq_max_len = len;\n+\t\t\tstream-\u003esreq_max_segs = INT_MAX;\n \t\t\tstream-\u003eprepare_write(subreq);\n \n-\t\t\tpart = umin(len, stream-\u003esreq_max_len);\n-\t\t\tif (unlikely(stream-\u003esreq_max_segs))\n-\t\t\t\tpart = netfs_limit_iter(\u0026source, 0, part, stream-\u003esreq_max_segs);\n+\t\t\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026dispatch_cursor);\n+\t\t\tpart = bvecq_slice(\u0026dispatch_cursor,\n+\t\t\t\t\t umin(len, stream-\u003esreq_max_len),\n+\t\t\t\t\t stream-\u003esreq_max_segs,\n+\t\t\t\t\t \u0026subreq-\u003enr_segs);\n \t\t\tsubreq-\u003elen = part;\n-\t\t\tsubreq-\u003etransferred = 0;\n+\n \t\t\tlen -= part;\n \t\t\tstart += part;\n \t\t\tif (len \u0026\u0026 subreq == to \u0026\u0026\n@@ -121,7 +129,7 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \t\t\t\tboundary = true;\n \n \t\t\tnetfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);\n-\t\t\tnetfs_reissue_write(stream, subreq, \u0026source);\n+\t\t\tnetfs_reissue_write(stream, subreq);\n \t\t\tif (subreq == to)\n \t\t\t\tbreak;\n \t\t}\n@@ -172,17 +180,19 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \t\t\t\tnetfs_stat(\u0026netfs_n_wh_upload);\n \t\t\t\tstream-\u003esreq_max_len = umin(len, wreq-\u003ewsize);\n \t\t\t\tbreak;\n-\t\t\tcase NETFS_WRITE_TO_CACHE:\n-\t\t\t\tnetfs_stat(\u0026netfs_n_wh_write);\n-\t\t\t\tbreak;\n \t\t\tdefault:\n \t\t\t\tWARN_ON_ONCE(1);\n \t\t\t}\n \n \t\t\tstream-\u003eprepare_write(subreq);\n \n-\t\t\tpart = umin(len, stream-\u003esreq_max_len);\n+\t\t\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026dispatch_cursor);\n+\t\t\tpart = bvecq_slice(\u0026dispatch_cursor,\n+\t\t\t\t\t umin(len, stream-\u003esreq_max_len),\n+\t\t\t\t\t stream-\u003esreq_max_segs,\n+\t\t\t\t\t \u0026subreq-\u003enr_segs);\n \t\t\tsubreq-\u003elen = subreq-\u003etransferred + part;\n+\n \t\t\tlen -= part;\n \t\t\tstart += part;\n \t\t\tif (!len \u0026\u0026 boundary) {\n@@ -190,13 +200,16 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n \t\t\t\tboundary = false;\n \t\t\t}\n \n-\t\t\tnetfs_reissue_write(stream, subreq, \u0026source);\n+\t\t\tnetfs_reissue_write(stream, subreq);\n \t\t\tif (!len)\n \t\t\t\tbreak;\n \n \t\t} while (len);\n \n \t} while (!list_is_head(next, \u0026stream-\u003esubrequests));\n+\n+out:\n+\tbvecq_pos_unset(\u0026dispatch_cursor);\n }\n \n /*\ndiff --git a/include/linux/bvecq.h b/include/linux/bvecq.h\nindex b984aaa449088..389f6407cb846 100644\n--- a/include/linux/bvecq.h\n+++ b/include/linux/bvecq.h\n@@ -54,6 +54,16 @@ struct bvecq {\n /* Number of slots in a 4K bvecq. */\n #define BVECQ_4KB_SLOTS ((4096 - sizeof(struct bvecq)) / sizeof(struct bio_vec))\n \n+/*\n+ * Position in a bio_vec queue. The bvecq holds a ref on the queue segment it\n+ * points to.\n+ */\n+struct bvecq_pos {\n+\tstruct bvecq\t\t*bvecq;\t\t/* The first bvecq */\n+\tunsigned int\t\toffset;\t\t/* The offset within the starting slot */\n+\tu16\t\t\tslot;\t\t/* The starting slot */\n+};\n+\n void bvecq_dump(const struct bvecq *bq);\n struct bvecq *bvecq_alloc_one(size_t nr_slots, gfp_t gfp, bool for_writeback);\n struct bvecq *bvecq_alloc_chain(size_t nr_slots, gfp_t gfp, bool for_writeback);\n@@ -61,6 +71,13 @@ struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp\n \t\t\t\t bool for_writeback);\n void bvecq_put(struct bvecq *bq);\n int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size, gfp_t gfp);\n+int bvecq_buffer_init(struct bvecq_pos *pos, gfp_t gfp, bool for_writeback);\n+void bvecq_buffer_append(struct bvecq_pos *pos, struct bvecq *bq);\n+void bvecq_pos_advance(struct bvecq_pos *pos, size_t amount);\n+ssize_t bvecq_zero(struct bvecq_pos *pos, size_t amount);\n+size_t bvecq_slice(struct bvecq_pos *pos, size_t max_size,\n+\t\t unsigned int max_slots, unsigned int *_nr_slots);\n+ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl);\n \n /**\n * bvecq_alloc_buffer - Allocate a bvecq chain and populate with buffers\n@@ -163,4 +180,171 @@ static inline struct bvecq *bvecq_next(const struct bvecq *bq)\n \treturn smp_load_acquire(\u0026bq-\u003enext);\n }\n \n+/**\n+ * bvecq_pos_set - Set one position to be the same as another\n+ * @pos: The position object to set\n+ * @at: The source position.\n+ *\n+ * Set @pos to have the same position as @at. This may take a ref on the\n+ * bvecq pointed to.\n+ */\n+static inline void bvecq_pos_set(struct bvecq_pos *pos, const struct bvecq_pos *at)\n+{\n+\t*pos = *at;\n+\tbvecq_get(pos-\u003ebvecq);\n+}\n+\n+/**\n+ * bvecq_pos_unset - Unset a position\n+ * @pos: The position object to unset\n+ *\n+ * Unset @pos. This does any needed ref cleanup.\n+ */\n+static inline void bvecq_pos_unset(struct bvecq_pos *pos)\n+{\n+\tbvecq_put(pos-\u003ebvecq);\n+\tpos-\u003ebvecq = NULL;\n+\tpos-\u003eslot = 0;\n+\tpos-\u003eoffset = 0;\n+}\n+\n+/**\n+ * bvecq_pos_transfer - Transfer one position to another, clearing the first\n+ * @pos: The position object to set\n+ * @from: The source position to clear.\n+ *\n+ * Set @pos to have the same position as @from and then clear @from. This may\n+ * transfer a ref on the bvecq pointed to.\n+ */\n+static inline void bvecq_pos_transfer(struct bvecq_pos *pos, struct bvecq_pos *from)\n+{\n+\t*pos = *from;\n+\tfrom-\u003ebvecq = NULL;\n+\tfrom-\u003eslot = 0;\n+\tfrom-\u003eoffset = 0;\n+}\n+\n+/**\n+ * bvecq_pos_move - Update a position to a new bvecq\n+ * @pos: The position object to update.\n+ * @to: The new bvecq to point at.\n+ *\n+ * Update @pos to point to @to if it doesn't already do so. This may\n+ * manipulate refs on the bvecqs pointed to.\n+ */\n+static inline void bvecq_pos_move(struct bvecq_pos *pos, struct bvecq *to)\n+{\n+\tstruct bvecq *old = pos-\u003ebvecq;\n+\n+\tif (old != to) {\n+\t\tpos-\u003ebvecq = bvecq_get(to);\n+\t\tbvecq_put(old);\n+\t}\n+}\n+\n+/**\n+ * bvecq_pos_nudge - Nudge a position onto the next segment if current used up\n+ * @pos: The position object to nudge.\n+ *\n+ * Update @pos to point to the next segment in the chain if we've used up the\n+ * current segment. This may manipulate refs on the bvecqs pointed to.\n+ *\n+ * Return: true if found a new segment, false if hit the end.\n+ */\n+static inline bool bvecq_pos_nudge(struct bvecq_pos *pos)\n+{\n+\tstruct bvecq *bq = pos-\u003ebvecq;\n+\n+\tfor (;;) {\n+\t\tif (!bvecq_acquire_slot(bq, pos-\u003eslot)) {\n+\t\t\tbq = bvecq_next(bq);\n+\t\t\tif (!bq)\n+\t\t\t\treturn false;\n+\t\t\tif (bvecq_acquire_slot(bq, pos-\u003eslot))\n+\t\t\t\tcontinue; /* More slots got added. */\n+\t\t\tbvecq_pos_move(pos, bq);\n+\t\t\tpos-\u003eslot = 0;\n+\t\t\tpos-\u003eoffset = 0;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (pos-\u003eoffset \u003e= bq-\u003ebv[pos-\u003eslot].bv_len) {\n+\t\t\tpos-\u003eslot++;\n+\t\t\tpos-\u003eoffset = 0;\n+\t\t\tcontinue;\n+\t\t}\n+\t\treturn true;\n+\t}\n+}\n+\n+/**\n+ * bvecq_pos_step - Step a position to the next slot if possible\n+ * @pos: The position object to step.\n+ *\n+ * Update @pos to point to the next slot in the queue if not at the end. This\n+ * may manipulate refs on the bvecqs pointed to.\n+ *\n+ * Return: true if successful, false if was at the end.\n+ */\n+static inline bool bvecq_pos_step(struct bvecq_pos *pos)\n+{\n+\tstruct bvecq *bq = pos-\u003ebvecq, *next;\n+\n+\tpos-\u003eslot++;\n+\tpos-\u003eoffset = 0;\n+\tif (bvecq_acquire_slot(bq, pos-\u003eslot))\n+\t\treturn true;\n+\tnext = bvecq_next(bq);\n+\tif (!next)\n+\t\treturn false;\n+\tif (bvecq_acquire_slot(bq, pos-\u003eslot))\n+\t\treturn true;\n+\tbvecq_pos_move(pos, next);\n+\tpos-\u003eslot = 0;\n+\treturn true;\n+}\n+\n+/**\n+ * bvecq_delete_spent - Delete the bvecq at the front if possible\n+ * @pos: The position object to update.\n+ *\n+ * Delete the used up bvecq at the front of the queue that @pos points to if it\n+ * is not the last node in the queue; if it is the last node in the queue, it\n+ * is kept so that the queue doesn't become detached from the other end. This\n+ * may manipulate refs on the bvecqs pointed to. It is also possible that the\n+ * producer will fill more slots in the current bvecq.\n+ *\n+ * Also, we have to be very careful: the consumer can catch the producer, which\n+ * could lead to us having nothing left in the queue, causing the front and\n+ * back pointers to end up on different tracks. To avoid this, we must always\n+ * keep at least one segment in the queue.\n+ *\n+ * The caller must reload from @pos after calling this.\n+ *\n+ * Return: true if there's more available; false if not.\n+ */\n+static inline bool bvecq_delete_spent(struct bvecq_pos *pos)\n+{\n+\tstruct bvecq *spent = pos-\u003ebvecq;\n+\tstruct bvecq *next;\n+\tunsigned int slot = pos-\u003eslot;\n+\n+again:\n+\t/* Read the contents of the queue node after the pointer to it. */\n+\tnext = bvecq_next(spent);\n+\tif (!next)\n+\t\treturn false; /* Nothing more to consume at the moment. */\n+\tif (slot \u003c bvecq_nr_slots_acquire(spent))\n+\t\treturn true; /* The producer added more. */\n+\tnext-\u003eprev = NULL;\n+\tbvecq_pos_move(pos, next);\n+\tpos-\u003eslot = 0;\n+\tpos-\u003eoffset = 0;\n+\tif (!bvecq_acquire_slot(next, 0)) {\n+\t\tspent = next;\n+\t\tslot = 0;\n+\t\tgoto again;\n+\t}\n+\treturn true;\n+}\n+\n #endif /* _LINUX_BVECQ_H */\ndiff --git a/include/linux/netfs.h b/include/linux/netfs.h\nindex b0dd92d12a971..0340c9ee587b6 100644\n--- a/include/linux/netfs.h\n+++ b/include/linux/netfs.h\n@@ -19,10 +19,12 @@\n #include \u003clinux/pagemap.h\u003e\n #include \u003clinux/bvecq.h\u003e\n #include \u003clinux/uio.h\u003e\n-#include \u003clinux/rolling_buffer.h\u003e\n \n enum netfs_sreq_ref_trace;\n typedef struct mempool mempool_t;\n+struct readahead_control;\n+struct netfs_io_request;\n+struct netfs_io_subrequest;\n struct fscache_occupancy;\n \n /**\n@@ -144,7 +146,6 @@ struct netfs_io_stream {\n \tunsigned int\t\tsreq_max_segs;\t/* 0 or max number of segments in an iterator */\n \tunsigned int\t\tsubmit_off;\t/* Folio offset we're submitting from */\n \tunsigned int\t\tsubmit_len;\t/* Amount of data left to submit */\n-\tunsigned int\t\tsubmit_extendable_to; /* Amount I/O can be rounded up to */\n \tvoid (*prepare_write)(struct netfs_io_subrequest *subreq);\n \tvoid (*issue_write)(struct netfs_io_subrequest *subreq);\n \t/* Collection tracking */\n@@ -187,6 +188,7 @@ struct netfs_io_subrequest {\n \tstruct netfs_io_request *rreq;\t\t/* Supervising I/O request */\n \tstruct work_struct\twork;\n \tstruct list_head\trreq_link;\t/* Link in rreq-\u003esubrequests */\n+\tstruct bvecq_pos\tio_buffer;\t/* Bookmark in the combined queue of the start */\n \tstruct iov_iter\t\tio_iter;\t/* Iterator for this subrequest */\n \tuoff_t\t\t\tstart;\t\t/* Where to start the I/O */\n \tsize_t\t\t\tlen;\t\t/* Size of the I/O */\n@@ -247,11 +249,14 @@ struct netfs_io_request {\n \tstruct netfs_io_stream\tio_streams[2];\t/* Streams of parallel I/O operations */\n #define NR_IO_STREAMS 2 //wreq-\u003enr_io_streams\n \tstruct netfs_group\t*group;\t\t/* Writeback group being written back */\n-\tstruct rolling_buffer\tbuffer;\t\t/* Unencrypted buffer */\n+\tstruct bvecq\t\t*spare;\t\t/* Advance allocation of bvecq */\n+\tstruct bvecq_pos\tload_cursor;\t/* Point at which new folios are loaded in */\n+\tstruct bvecq_pos\tdispatch_cursor; /* Point from which buffers are dispatched */\n+\tstruct bvecq_pos\tcollect_cursor;\t/* Clear-up point of I/O buffer */\n \twait_queue_head_t\twaitq;\t\t/* Processor waiter */\n \tvoid\t\t\t*netfs_priv;\t/* Private data for the netfs */\n \tvoid\t\t\t*netfs_priv2;\t/* Private data for the netfs */\n-\tstruct bio_vec\t\t*direct_bv;\t/* DIO buffer list (when handling iovec-iter) */\n+\tuoff_t\t\t\tlast_end;\t/* End pos of last folio submitted */\n \tuoff_t\t\t\tsubmitted;\t/* Amount submitted for I/O so far */\n \tuoff_t\t\t\tlen;\t\t/* Length of the request */\n \tsize_t\t\t\ttransferred;\t/* Amount to be indicated as transferred */\n@@ -266,7 +271,6 @@ struct netfs_io_request {\n \tuoff_t\t\t\tabandon_to;\t/* Position to abandon folios to */\n \tconst struct folio\t*no_unlock_folio; /* Don't unlock this folio after read */\n \tgfp_t\t\t\tgfp;\t\t/* GFP flags to use */\n-\tunsigned int\t\tdirect_bv_count; /* Number of elements in direct_bv[] */\n \tunsigned int\t\tdebug_id;\n \tunsigned int\t\trsize;\t\t/* Maximum read size (0 for none) */\n \tunsigned int\t\twsize;\t\t/* Maximum write size (0 for none) */\n@@ -274,7 +278,6 @@ struct netfs_io_request {\n \tunsigned int\t\tnr_group_rel;\t/* Number of refs to release on -\u003egroup */\n \tspinlock_t\t\tlock;\t\t/* Lock for queuing subreqs */\n \tenum netfs_io_origin\torigin;\t\t/* Origin of the request */\n-\tbool\t\t\tdirect_bv_unpin; /* T if direct_bv[] must be unpinned */\n \trefcount_t\t\tref;\n \tunsigned long\t\tflags;\n #define NETFS_RREQ_IN_PROGRESS\t\t0\t/* Unlocked when the request completes (has ref) */\n@@ -427,7 +430,7 @@ void netfs_single_mark_inode_dirty(struct inode *inode);\n ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_iter *iter);\n int netfs_writeback_single(struct address_space *mapping,\n \t\t\t struct writeback_control *wbc,\n-\t\t\t struct iov_iter *iter);\n+\t\t\t struct iov_iter *iter, size_t len);\n \n /* Address operations API */\n struct readahead_control;\n@@ -454,11 +457,9 @@ void netfs_get_subrequest(struct netfs_io_subrequest *subreq,\n \t\t\t enum netfs_sreq_ref_trace what);\n void netfs_put_subrequest(struct netfs_io_subrequest *subreq,\n \t\t\t enum netfs_sreq_ref_trace what);\n-ssize_t netfs_extract_user_iter(struct iov_iter *orig, size_t orig_len,\n-\t\t\t\tstruct iov_iter *new,\n-\t\t\t\tiov_iter_extraction_t extraction_flags);\n-size_t netfs_limit_iter(const struct iov_iter *iter, size_t start_offset,\n-\t\t\tsize_t max_size, size_t max_segs);\n+ssize_t netfs_extract_iter(struct iov_iter *orig, size_t max_len, size_t max_pages,\n+\t\t\t struct bvecq **_bvecq_head,\n+\t\t\t iov_iter_extraction_t extraction_flags, gfp_t gfp);\n void netfs_prepare_write_failed(struct netfs_io_subrequest *subreq);\n void netfs_write_subrequest_terminated(void *_op, ssize_t transferred_or_error);\n \ndiff --git a/include/linux/pagemap.h b/include/linux/pagemap.h\nindex 0adfa6605653d..b7154b81d0d7c 100644\n--- a/include/linux/pagemap.h\n+++ b/include/linux/pagemap.h\n@@ -1411,6 +1411,7 @@ struct readahead_control {\n \tstruct file_ra_state *ra;\n /* private: use the readahead_* accessors instead */\n \tpgoff_t _index;\n+\tunsigned int _nr_folios;\n \tunsigned int _nr_pages;\n \tunsigned int _batch_count;\n \tbool dropbehind;\n@@ -1590,6 +1591,15 @@ static inline size_t readahead_batch_length(const struct readahead_control *rac)\n \treturn rac-\u003e_batch_count * PAGE_SIZE;\n }\n \n+/**\n+ * readahead_folio_count - Get the number of folios in this readahead request.\n+ * @rac: The readahead request.\n+ */\n+static inline unsigned int readahead_folio_count(const struct readahead_control *rac)\n+{\n+\treturn rac-\u003e_nr_folios;\n+}\n+\n static inline unsigned long dir_pages(const struct inode *inode)\n {\n \treturn (unsigned long)(inode-\u003ei_size + PAGE_SIZE - 1) \u003e\u003e\ndiff --git a/include/linux/rolling_buffer.h b/include/linux/rolling_buffer.h\ndeleted file mode 100644\nindex 5c0bc4221f010..0000000000000\n--- a/include/linux/rolling_buffer.h\n+++ /dev/null\n@@ -1,47 +0,0 @@\n-/* SPDX-License-Identifier: GPL-2.0-or-later */\n-/* Rolling buffer of folios\n- *\n- * Copyright (C) 2024 Red Hat, Inc. All Rights Reserved.\n- * Written by David Howells (dhowells@redhat.com)\n- */\n-\n-#ifndef _ROLLING_BUFFER_H\n-#define _ROLLING_BUFFER_H\n-\n-#include \u003clinux/bvecq.h\u003e\n-#include \u003clinux/uio.h\u003e\n-\n-/*\n- * Rolling buffer. Whilst the buffer is live and in use, folios and bvecq\n- * segments can be added to one end by one thread and removed from the other\n- * end by another thread. The buffer isn't allowed to be empty; it must always\n- * have at least one bvecq in it so that neither side has to modify both queue\n- * pointers.\n- *\n- * The iterator in the buffer is extended as buffers are inserted. It can be\n- * snapshotted to use a segment of the buffer.\n- */\n-struct rolling_buffer {\n-\tstruct bvecq\t\t*head;\t\t/* Producer's insertion point */\n-\tstruct bvecq\t\t*tail;\t\t/* Consumer's removal point */\n-\tstruct iov_iter\t\titer;\t\t/* Iterator tracking what's left in the buffer */\n-\tu8\t\t\tfirst_tail_slot; /* First slot in -\u003etail */\n-\tbool\t\t\tfor_writeback;\t/* T if being used for writeback */\n-};\n-\n-int rolling_buffer_init(struct rolling_buffer *roll, unsigned int direction,\n-\t\t\tgfp_t gfp, bool for_writeback);\n-int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp);\n-ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,\n-\t\t\t\t\t struct readahead_control *ractl,\n-\t\t\t\t\t gfp_t gfp);\n-ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio, gfp_t gfp);\n-struct bvecq *rolling_buffer_delete_spent(struct rolling_buffer *roll);\n-void rolling_buffer_clear(struct rolling_buffer *roll);\n-\n-static inline void rolling_buffer_advance(struct rolling_buffer *roll, size_t amount)\n-{\n-\tiov_iter_advance(\u0026roll-\u003eiter, amount);\n-}\n-\n-#endif /* _ROLLING_BUFFER_H */\ndiff --git a/include/trace/events/netfs.h b/include/trace/events/netfs.h\nindex e80c27c49ce25..81a3b7aaa187c 100644\n--- a/include/trace/events/netfs.h\n+++ b/include/trace/events/netfs.h\n@@ -229,7 +229,9 @@\n \tEM(netfs_folio_trace_sched_copy,\t\"sched-copy\")\t\\\n \tEM(netfs_folio_trace_store,\t\t\"store\")\t\\\n \tEM(netfs_folio_trace_store_copy,\t\"store-copy\")\t\\\n-\tE_(netfs_folio_trace_store_plus,\t\"store+\")\n+\tEM(netfs_folio_trace_store_plus,\t\"store+\")\t\\\n+\tEM(netfs_folio_trace_zero,\t\t\"zero\")\t\t\\\n+\tE_(netfs_folio_trace_zero_ra,\t\t\"zero-ra\")\n \n #define netfs_collect_contig_traces\t\t\t\t\\\n \tEM(netfs_contig_trace_collect,\t\t\"Collect\")\t\\\n@@ -386,10 +388,10 @@ TRACE_EVENT(netfs_sreq,\n \t\t __entry-\u003elen\t= sreq-\u003elen;\n \t\t __entry-\u003etransferred = sreq-\u003etransferred;\n \t\t __entry-\u003estart\t= sreq-\u003estart;\n-\t\t __entry-\u003eslot\t= sreq-\u003eio_iter.bvecq_slot;\n+\t\t __entry-\u003eslot\t= sreq-\u003eio_buffer.slot;\n \t\t\t ),\n \n-\t TP_printk(\"R=%08x[%x] %s %s f=%03x s=%llx %zx/%zx s=%u e=%d\",\n+\t TP_printk(\"R=%08x[%x] %s %s f=%03x s=%llx %zx/%zx bv=%u e=%d\",\n \t\t __entry-\u003erreq, __entry-\u003eindex,\n \t\t __print_symbolic(__entry-\u003esource, netfs_sreq_sources),\n \t\t __print_symbolic(__entry-\u003ewhat, netfs_sreq_traces),\n@@ -782,6 +784,30 @@ TRACE_EVENT(netfs_read_progress_at,\n \t\t __entry-\u003erreq, __entry-\u003ecleaned_to, __entry-\u003eprogress_at)\n \t );\n \n+TRACE_EVENT(netfs_bv_slot,\n+\t TP_PROTO(const struct bvecq *bq, int slot),\n+\n+\t TP_ARGS(bq, slot),\n+\n+\t TP_STRUCT__entry(\n+\t\t __field(unsigned long,\t\tpfn)\n+\t\t __field(unsigned int,\t\toffset)\n+\t\t __field(unsigned int,\t\tlen)\n+\t\t __field(unsigned int,\t\tslot)\n+\t\t\t ),\n+\n+\t TP_fast_assign(\n+\t\t __entry-\u003eslot = slot;\n+\t\t __entry-\u003epfn = page_to_pfn(bq-\u003ebv[slot].bv_page);\n+\t\t __entry-\u003eoffset = bq-\u003ebv[slot].bv_offset;\n+\t\t __entry-\u003elen = bq-\u003ebv[slot].bv_len;\n+\t\t\t ),\n+\n+\t TP_printk(\"bq[%x] p=%lx %x-%x\",\n+\t\t __entry-\u003eslot,\n+\t\t __entry-\u003epfn, __entry-\u003eoffset, __entry-\u003eoffset + __entry-\u003elen)\n+\t );\n+\n #undef EM\n #undef E_\n #endif /* _TRACE_NETFS_H */\ndiff --git a/mm/readahead.c b/mm/readahead.c\nindex 6e5563290287e..196542fb1c930 100644\n--- a/mm/readahead.c\n+++ b/mm/readahead.c\n@@ -188,6 +188,7 @@ static void read_pages(struct readahead_control *rac)\n \tif (unlikely(rac-\u003e_workingset))\n \t\tpsi_memstall_leave(\u0026rac-\u003e_pflags);\n \trac-\u003e_workingset = false;\n+\trac-\u003e_nr_folios = 0;\n \n \tBUG_ON(readahead_count(rac));\n }\n@@ -303,6 +304,7 @@ void page_cache_ra_unbounded(struct readahead_control *ractl,\n \t\tif (i == mark)\n \t\t\tfolio_set_readahead(folio);\n \t\tractl-\u003e_workingset |= folio_test_workingset(folio);\n+\t\tractl-\u003e_nr_folios++;\n \t\tractl-\u003e_nr_pages += min_nrpages;\n \t\ti += min_nrpages;\n \t}\n@@ -473,6 +475,7 @@ static inline int ra_alloc_folio(struct readahead_control *ractl, pgoff_t index,\n \t\treturn err;\n \t}\n \n+\tractl-\u003e_nr_folios++;\n \tractl-\u003e_nr_pages += 1UL \u003c\u003c order;\n \tractl-\u003e_workingset |= folio_test_workingset(folio);\n \treturn 0;\n@@ -822,6 +825,7 @@ void readahead_expand(struct readahead_control *ractl,\n \t\t\tractl-\u003e_workingset = true;\n \t\t\tpsi_memstall_enter(\u0026ractl-\u003e_pflags);\n \t\t}\n+\t\tractl-\u003e_nr_folios++;\n \t\tractl-\u003e_nr_pages += min_nrpages;\n \t\tractl-\u003e_index = folio-\u003eindex;\n \t}\n@@ -851,6 +855,7 @@ void readahead_expand(struct readahead_control *ractl,\n \t\t\tractl-\u003e_workingset = true;\n \t\t\tpsi_memstall_enter(\u0026ractl-\u003e_pflags);\n \t\t}\n+\t\tractl-\u003e_nr_folios++;\n \t\tractl-\u003e_nr_pages += min_nrpages;\n \t\tif (ra) {\n \t\t\tra-\u003esize += min_nrpages;\n"
} |
| 2/1 | 2026/10/06 21:22 | agent | patch-evaluator |
2mModel:core Args: null Results: {
"EnableConfigs": [
"NETFS_SUPPORT",
"FSCACHE",
"CACHEFILES",
"AFS_FS",
"NETFS_PGPRIV2"
],
"FocusSymbols": [
"bvecq_buffer_init",
"bvecq_buffer_append",
"bvecq_pos_advance",
"bvecq_zero",
"bvecq_slice",
"bvecq_load_from_ra",
"netfs_extract_iter",
"netfs_readahead",
"netfs_writepages",
"netfs_writeback_single",
"netfs_read_single",
"netfs_unbuffered_read_iter_locked",
"netfs_unbuffered_write_iter_locked",
"netfs_retry_reads",
"netfs_retry_writes",
"netfs_pgpriv2_unlock_copied_folios",
"afs_symlink_writepages"
],
"Reasoning": "The patch replaces the rolling_buffer implementation in netfs with a cursor-based bio_vec queue (bvecq_pos/bvecq) mechanism across read, write, direct I/O, readahead, retry, and writeback paths. It introduces new buffer-management primitives (bvecq_slice, bvecq_pos_advance, bvecq_zero, bvecq_load_from_ra, netfs_extract_iter), refactors I/O dispatch and folio unlock accounting, and updates network filesystems (AFS, Cachefiles) and mm/readahead folio tracking. These changes are in core reachable kernel filesystem paths and warrant fuzzing to identify potential regressions, buffer overflows, or invariant violations.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit baca3ab6cbbda96e6e631014dc876f11c374596c
Author: syz-cluster <triage@syzkaller.com>
Date: Tue Oct 6 21:22:55 2026 +0000
syz-cluster: applied patch under review
diff --git a/fs/afs/dir.c b/fs/afs/dir.c
index 5fac02d2d2814..39cba3f37ecd4 100644
--- a/fs/afs/dir.c
+++ b/fs/afs/dir.c
@@ -2229,8 +2229,9 @@ static int afs_dir_writepages(struct address_space *mapping,
if (test_bit(AFS_VNODE_DIR_VALID, &dvnode->flags)) {
iov_iter_bvec_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0,
- i_size_read(&dvnode->netfs.inode));
- ret = netfs_writeback_single(mapping, wbc, &iter);
+ dvnode->directory_size);
+ ret = netfs_writeback_single(mapping, wbc, &iter,
+ i_size_read(&dvnode->netfs.inode));
if (ret == 1)
ret = 0; /* Skipped write due to lock conflict. */
}
diff --git a/fs/afs/symlink.c b/fs/afs/symlink.c
index 9a611efe6b264..ae03ceff42b85 100644
--- a/fs/afs/symlink.c
+++ b/fs/afs/symlink.c
@@ -248,9 +248,9 @@ int afs_symlink_writepages(struct address_space *mapping,
if (vnode->directory &&
atomic64_read(&vnode->cb_expires_at) != AFS_NO_CB_PROMISE) {
- iov_iter_bvec_queue(&iter, ITER_SOURCE, vnode->directory, 0, 0,
- i_size_read(&vnode->netfs.inode));
- ret = netfs_writeback_single(mapping, wbc, &iter);
+ iov_iter_bvec_queue(&iter, ITER_SOURCE, vnode->directory, 0, 0, PAGE_SIZE);
+ ret = netfs_writeback_single(mapping, wbc, &iter,
+ i_size_read(&vnode->netfs.inode));
}
if (ret == 0) {
diff --git a/fs/cachefiles/io.c b/fs/cachefiles/io.c
index d05059822288c..788439ea6e1c8 100644
--- a/fs/cachefiles/io.c
+++ b/fs/cachefiles/io.c
@@ -546,7 +546,6 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq)
struct netfs_cache_resources *cres = &wreq->cache_resources;
struct cachefiles_object *object = cachefiles_cres_object(cres);
struct cachefiles_cache *cache = object->volume->cache;
- struct netfs_io_stream *stream = &wreq->io_streams[subreq->stream_nr];
const struct cred *saved_cred;
size_t off, pre, post, len = subreq->len;
uoff_t start = subreq->start;
@@ -571,17 +570,6 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq)
}
/* We also need to end on the cache granularity boundary */
- if (start + len == wreq->i_size) {
- size_t part = len & (cache->bsize - 1);
- size_t need = cache->bsize - part;
-
- if (part && stream->submit_extendable_to >= need) {
- len += need;
- subreq->len += need;
- subreq->io_iter.count += need;
- }
- }
-
post = len & (cache->bsize - 1);
if (post) {
len -= post;
diff --git a/fs/netfs/Makefile b/fs/netfs/Makefile
index b1ea4439c1bb4..421dd0be413b3 100644
--- a/fs/netfs/Makefile
+++ b/fs/netfs/Makefile
@@ -14,7 +14,6 @@ netfs-y := \
read_collect.o \
read_retry.o \
read_single.o \
- rolling_buffer.o \
write_collect.o \
write_issue.o \
write_retry.o
diff --git a/fs/netfs/buffered_read.c b/fs/netfs/buffered_read.c
index 052684ce1e347..3ca75b5314eee 100644
--- a/fs/netfs/buffered_read.c
+++ b/fs/netfs/buffered_read.c
@@ -114,26 +114,21 @@ static int netfs_begin_cache_read(struct netfs_io_request *rreq, struct netfs_in
static ssize_t netfs_prepare_read_iterator(struct netfs_io_subrequest *subreq)
{
struct netfs_io_request *rreq = subreq->rreq;
+ struct netfs_io_stream *stream = &rreq->io_streams[0];
+ ssize_t extracted;
size_t rsize = subreq->len;
if (subreq->source == NETFS_DOWNLOAD_FROM_SERVER)
- rsize = umin(rsize, rreq->io_streams[0].sreq_max_len);
-
- subreq->len = rsize;
- if (unlikely(rreq->io_streams[0].sreq_max_segs)) {
- size_t limit = netfs_limit_iter(&rreq->buffer.iter, 0, rsize,
- rreq->io_streams[0].sreq_max_segs);
-
- if (limit < rsize) {
- subreq->len = limit;
- trace_netfs_sreq(subreq, netfs_sreq_trace_limited);
- }
+ rsize = umin(rsize, stream->sreq_max_len);
+
+ bvecq_pos_set(&subreq->io_buffer, &rreq->dispatch_cursor);
+ extracted = bvecq_slice(&rreq->dispatch_cursor, rsize,
+ stream->sreq_max_segs, &subreq->nr_segs);
+ if (extracted < subreq->len) {
+ subreq->len = extracted;
+ trace_netfs_sreq(subreq, netfs_sreq_trace_limited);
}
- subreq->io_iter = rreq->buffer.iter;
-
- iov_iter_truncate(&subreq->io_iter, subreq->len);
- rolling_buffer_advance(&rreq->buffer, subreq->len);
return subreq->len;
}
@@ -192,6 +187,9 @@ void netfs_queue_read(struct netfs_io_request *rreq,
static void netfs_issue_read(struct netfs_io_request *rreq,
struct netfs_io_subrequest *subreq)
{
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
+
switch (subreq->source) {
case NETFS_DOWNLOAD_FROM_SERVER:
rreq->netfs_ops->issue_read(subreq);
@@ -200,10 +198,9 @@ static void netfs_issue_read(struct netfs_io_request *rreq,
netfs_read_cache_to_pagecache(rreq, subreq);
break;
default:
- __set_bit(NETFS_SREQ_CLEAR_TAIL, &subreq->flags);
- subreq->error = 0;
- iov_iter_zero(subreq->len, &subreq->io_iter);
+ bvecq_zero(&subreq->io_buffer, subreq->len);
subreq->transferred = subreq->len;
+ subreq->error = 0;
netfs_read_subreq_terminated(subreq);
break;
}
@@ -215,31 +212,31 @@ static void netfs_issue_read(struct netfs_io_request *rreq,
* otherwise we set the deprecated PG_private_2.
*/
static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,
- struct bvecq **bq,
- unsigned int *offset,
- int *slot,
- size_t len,
- bool copy)
+ struct bvecq_pos *mark, size_t len, bool copy)
{
+ struct bvecq *bq = mark->bvecq;
+ unsigned int offset = mark->offset;
+ int slot = mark->slot;
+
while (len > 0) {
- struct folio *folio;
size_t fsize, overlap;
- if (!*bq)
+ if (!bq)
break;
- if (!bvecq_acquire_slot(*bq, *slot)) {
- *bq = bvecq_next(*bq);
- *slot = 0;
- *offset = 0;
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = bq->next;
+ slot = 0;
+ offset = 0;
continue;
}
/* Determine how much the subreq overlaps the folio, if at all. */
- fsize = (*bq)->bv[*slot].bv_len;
- overlap = min(len, fsize - *offset);
+ fsize = bq->bv[slot].bv_len;
+ overlap = min(len, fsize - offset);
if (overlap > 0 && copy) {
- folio = bvec_folio(&(*bq)->bv[*slot]);
+ struct folio *folio = bvec_folio(&bq->bv[slot]);
+
if (netfs_using_pgpriv2(rreq)) {
if (!folio_test_private_2(folio))
folio_start_private_2(folio);
@@ -251,12 +248,20 @@ static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,
}
len -= overlap;
- *offset += overlap;
- if (*offset >= fsize) {
- *slot += 1;
- *offset = 0;
+ offset += overlap;
+ if (offset >= fsize) {
+ slot += 1;
+ offset = 0;
}
}
+
+ if (bq) {
+ bvecq_pos_move(mark, bq);
+ mark->offset = offset;
+ mark->slot = slot;
+ } else {
+ bvecq_pos_unset(mark);
+ }
}
/*
@@ -275,11 +280,14 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)
.cached_to[1] = ULLONG_MAX,
};
struct fscache_occupancy *occ = &_occ;
- struct bvecq *bq = rreq->buffer.tail;
- unsigned int offset = 0;
+ struct bvecq_pos mark_cursor;
ssize_t size = rreq->len;
uoff_t start = rreq->start;
- int ret = 0, slot = 0;
+ int ret = 0;
+
+ _enter("R=%08x", rreq->debug_id);
+
+ bvecq_pos_set(&mark_cursor, &rreq->dispatch_cursor);
do {
int (*prepare_read)(struct netfs_io_subrequest *subreq) = NULL;
@@ -408,10 +416,10 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)
if (size <= 0)
netfs_all_subreqs_queued(rreq);
- if (bq) {
+ if (mark_cursor.bvecq) {
/* See if the cache indicated this should be cached. */
copy = test_bit(NETFS_SREQ_COPY_TO_CACHE, &subreq->flags);
- netfs_mark_copy_to_cache(rreq, &bq, &slot, &offset, slice, copy);
+ netfs_mark_copy_to_cache(rreq, &mark_cursor, slice, copy);
}
trace_netfs_sreq(subreq, netfs_sreq_trace_submit);
@@ -432,6 +440,9 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)
/* Defer error return as we may need to wait for outstanding I/O. */
cmpxchg(&rreq->error, 0, ret);
+
+ bvecq_pos_unset(&mark_cursor);
+ bvecq_pos_unset(&rreq->dispatch_cursor);
}
/**
@@ -479,7 +490,7 @@ void netfs_readahead(struct readahead_control *ractl)
* acquires a ref on each folio that we will need to release later -
* but we don't want to do that until after we've started the I/O.
*/
- added = rolling_buffer_bulk_load_from_ra(&rreq->buffer, ractl, rreq->gfp);
+ added = bvecq_load_from_ra(&rreq->dispatch_cursor, ractl);
if (added < 0) {
ret = added;
goto cleanup_free;
@@ -488,6 +499,7 @@ void netfs_readahead(struct readahead_control *ractl)
rreq->submitted = rreq->start + added;
rreq->cleaned_to = rreq->start;
+ bvecq_pos_set(&rreq->collect_cursor, &rreq->dispatch_cursor);
netfs_read_set_unlock_at(rreq);
netfs_read_to_pagecache(rreq);
@@ -500,20 +512,26 @@ void netfs_readahead(struct readahead_control *ractl)
EXPORT_SYMBOL(netfs_readahead);
/*
- * Create a rolling buffer with a single occupying folio.
+ * Create a buffer queue with a single occupying folio.
*/
static int netfs_create_singular_buffer(struct netfs_io_request *rreq, struct folio *folio)
{
- ssize_t added;
+ struct bvecq *bq;
+ size_t fsize = folio_size(folio);
- if (rolling_buffer_init(&rreq->buffer, ITER_DEST, rreq->gfp, false) < 0)
+ bq = bvecq_alloc_one(1, rreq->gfp, false);
+ if (!bq)
return -ENOMEM;
- added = rolling_buffer_append(&rreq->buffer, folio, rreq->gfp);
- if (added < 0)
- return added;
- rreq->submitted = rreq->start + added;
- rreq->progress_at = added;
+ rreq->dispatch_cursor.bvecq = bq;
+ rreq->dispatch_cursor.slot = 0;
+ rreq->dispatch_cursor.offset = 0;
+
+ bvec_set_folio(&bq->bv[0], folio, fsize, 0);
+ bvecq_filled_to(bq, 1);
+ bvecq_pos_set(&rreq->collect_cursor, &rreq->dispatch_cursor);
+ rreq->submitted = rreq->start + fsize;
+ rreq->progress_at = fsize;
return 0;
}
@@ -527,14 +545,14 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
struct netfs_group *group = netfs_folio_group(folio);
struct netfs_folio *finfo = netfs_folio_info(folio);
struct netfs_inode *ctx = netfs_inode(mapping->host);
- struct bio_vec *bvec = NULL;
+ struct bvecq *bq = NULL;
unsigned int from = finfo->dirty_offset;
unsigned int to = from + finfo->dirty_len;
unsigned int off = 0;
size_t flen = folio_size(folio);
size_t nr_bvec = flen / PAGE_SIZE + 2;
size_t part;
- int ret, i = 0, sink_from = -1, sink_to = -1;
+ int ret, i = 0;
_enter("%lx", folio->index);
@@ -555,31 +573,46 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
* end get copied to, but the middle is discarded.
*/
ret = -ENOMEM;
- bvec = kmalloc_objs(*bvec, nr_bvec);
- if (!bvec)
+ bq = bvecq_alloc_chain(nr_bvec, rreq->gfp, false);
+ if (!bq)
goto discard;
+ rreq->dispatch_cursor.bvecq = bq;
trace_netfs_folio(folio, netfs_folio_trace_read_gaps);
+ for (struct bvecq *p = bq; p; p = p->next)
+ p->mem_type = BVECQ_MEM_PAGECACHE;
+
if (from > 0) {
- bvec_set_folio(&bvec[i++], folio, from, 0);
+ folio_get(folio);
+ bvec_set_folio(&bq->bv[i++], folio, from, 0);
off = from;
}
- sink_from = i;
while (off < to) {
struct folio *sink = folio_alloc(GFP_KERNEL, 0);
if (!sink)
goto discard;
- part = min_t(size_t, to - off, PAGE_SIZE);
- bvec_set_folio(&bvec[i], sink, part, 0);
+ if (i >= bq->max_slots) {
+ bvecq_filled_to(bq, i);
+ bq = bq->next;
+ i = 0;
+ }
+ part = min(to - off, PAGE_SIZE);
+ bvec_set_folio(&bq->bv[i++], sink, part, 0);
off += part;
- sink_to = i;
- i++;
}
- if (to < flen)
- bvec_set_folio(&bvec[i++], folio, flen - to, to);
- iov_iter_bvec(&rreq->buffer.iter, ITER_DEST, bvec, i, rreq->len);
+ if (to < flen) {
+ if (i >= bq->max_slots) {
+ bvecq_filled_to(bq, i);
+ bq = bq->next;
+ i = 0;
+ }
+ folio_get(folio);
+ bvec_set_folio(&bq->bv[i++], folio, flen - to, to);
+ }
+ bvecq_filled_to(bq, i);
+
rreq->submitted = rreq->start + flen;
netfs_read_to_pagecache(rreq);
@@ -587,11 +620,10 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
ret = netfs_wait_for_read(rreq);
if (ret >= 0) {
if (ret < flen) {
- struct iov_iter iter;
-
- iov_iter_bvec(&iter, ITER_DEST, bvec, i, flen);
- iov_iter_advance(&iter, ret);
- iov_iter_zero(flen - ret, &iter);
+ if (ret < from)
+ folio_zero_segments(folio, ret, from, to, flen);
+ else
+ folio_zero_segment(folio, max(to, ret), flen);
}
if (group)
folio_change_private(folio, group);
@@ -603,22 +635,16 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
folio_mark_uptodate(folio);
}
- if (sink_to >= 0)
- for (; sink_from <= sink_to; sink_from++)
- folio_put(bvec_folio(&bvec[sink_from]));
- kfree(bvec);
+ bvecq_pos_unset(&rreq->dispatch_cursor);
folio_unlock(folio);
netfs_put_request(rreq, netfs_rreq_trace_put_return);
return ret < 0 ? ret : 0;
discard:
+ bvecq_pos_unset(&rreq->dispatch_cursor);
netfs_put_failed_request(rreq);
alloc_error:
folio_unlock(folio);
- if (sink_to >= 0)
- for (; sink_from <= sink_to; sink_from++)
- folio_put(bvec_folio(&bvec[sink_from]));
- kfree(bvec);
return ret;
}
diff --git a/fs/netfs/bvecq.c b/fs/netfs/bvecq.c
index 5b747f6b59382..2edc0045c2453 100644
--- a/fs/netfs/bvecq.c
+++ b/fs/netfs/bvecq.c
@@ -342,3 +342,313 @@ int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size,
return 0;
}
EXPORT_SYMBOL(bvecq_expand_buffer);
+
+/**
+ * bvecq_buffer_init - Initialise a buffer and set position
+ * @pos: The position to point at the new buffer.
+ * @gfp: The allocation constraints.
+ * @for_writeback: True if allocating for writeback
+ *
+ * Initialise a rolling buffer. We allocate an unpopulated bvecq node to so
+ * that the pointers can be independently driven by the producer and the
+ * consumer.
+ *
+ * Return 0 if successful; -ENOMEM on allocation failure.
+ */
+int bvecq_buffer_init(struct bvecq_pos *pos, gfp_t gfp, bool for_writeback)
+{
+ struct bvecq *bq;
+
+ bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, for_writeback);
+ if (!bq)
+ return -ENOMEM;
+
+ pos->bvecq = bq; /* Comes with a ref. */
+ pos->slot = 0;
+ pos->offset = 0;
+ return 0;
+}
+
+/**
+ * bvecq_buffer_append - Append a new bvecq node to a buffer
+ * @pos: The position of the last node.
+ * @bq: The buffer to add.
+ *
+ * Add a new node on to the buffer chain at the specified position, either
+ * because the previous one is full or because we have a discontiguity to
+ * contend with, and update @pos to point to it.
+ */
+void bvecq_buffer_append(struct bvecq_pos *pos, struct bvecq *bq)
+{
+ struct bvecq *head = pos->bvecq;
+
+ pos->bvecq = bvecq_get(bq);
+ pos->slot = 0;
+ pos->offset = 0;
+
+ /* [!] NOTE: After we set head->next, the consumer is at liberty to
+ * immediately delete the old head.
+ */
+ bvecq_append(head, bq);
+ bvecq_put(head);
+}
+
+/**
+ * bvecq_pos_advance - Advance a bvecq position
+ * @pos: The position to advance.
+ * @amount: The amount of bytes to advance by.
+ *
+ * Advance the specified bvecq position by @amount bytes. @pos is updated and
+ * bvecq ref counts may have been manipulated. If the position hits the end of
+ * the queue, then it is left pointing beyond the last slot of the last bvecq
+ * so that it doesn't break the chain.
+ */
+void bvecq_pos_advance(struct bvecq_pos *pos, size_t amount)
+{
+ struct bvecq *bq = pos->bvecq, *next;
+ unsigned int slot = pos->slot;
+ size_t offset = pos->offset;
+
+ while (amount) {
+ size_t part;
+
+ if (!bvecq_acquire_slot(bq, slot)) {
+ next = bvecq_next(bq);
+ if (!next) {
+ WARN_ON_ONCE(amount > 0);
+ break;
+ }
+ if (bvecq_acquire_slot(bq, slot))
+ continue; /* More slots got added. */
+ bq = next;
+ slot = 0;
+ offset = 0;
+ continue;
+ }
+
+ part = bq->bv[slot].bv_len - offset;
+
+ if (part > amount) {
+ offset += amount;
+ break;
+ }
+ amount -= part;
+ offset = 0;
+ slot++;
+ }
+
+ pos->slot = slot;
+ pos->offset = offset;
+ bvecq_pos_move(pos, bq);
+}
+
+/*
+ * Clear part of the memory pointed to by a bio_vec.
+ */
+static void bvec_zero(const struct bio_vec *bv, size_t offset, size_t len)
+{
+ struct page *page = bv->bv_page;
+
+ offset += bv->bv_offset;
+
+ page += offset / PAGE_SIZE;
+ offset = offset % PAGE_SIZE;
+
+ while (len) {
+ size_t part = min(len, PAGE_SIZE - offset);
+ char *p = kmap_local_page(page);
+
+ memset(p + offset, 0, part);
+ kunmap_local(p);
+
+ len -= part;
+ offset = 0;
+ page++;
+ }
+}
+
+/**
+ * bvecq_zero - Clear memory starting at the bvecq position.
+ * @pos: The position in the bvecq chain to start clearing.
+ * @amount: The number of bytes to clear.
+ *
+ * Clear memory fragments pointed to by a bvec queue. @pos is updated and
+ * bvecq ref counts may have been manipulated. If the position hits the end of
+ * the queue, then it is left pointing beyond the last slot of the last bvecq
+ * so that it doesn't break the chain.
+ *
+ * Return: The number of bytes cleared.
+ */
+ssize_t bvecq_zero(struct bvecq_pos *pos, size_t amount)
+{
+ struct bvecq *bq = pos->bvecq, *next;
+ unsigned int slot = pos->slot;
+ ssize_t cleared = 0;
+ size_t offset = pos->offset;
+
+ while (amount) {
+ const struct bio_vec *bv;
+ size_t part;
+
+ if (!bvecq_acquire_slot(bq, slot)) {
+ next = bvecq_next(bq);
+ if (!next) {
+ WARN_ON_ONCE(amount > 0);
+ break;
+ }
+ if (bvecq_acquire_slot(bq, slot))
+ continue; /* More slots got added. */
+ bq = next;
+ slot = 0;
+ offset = 0;
+ continue;
+ }
+
+ bv = &bq->bv[slot];
+ if (offset >= bv->bv_len) {
+ slot++;
+ offset = 0;
+ continue;
+ }
+
+ part = min(bv->bv_len - offset, amount);
+ bvec_zero(bv, offset, part);
+ cleared += part;
+ offset += part;
+ amount -= part;
+ }
+
+ pos->slot = slot;
+ pos->offset = offset;
+ bvecq_pos_move(pos, bq);
+ return cleared;
+}
+
+/**
+ * bvecq_slice - Find a slice of a bvecq queue
+ * @pos: The position to start at.
+ * @max_size: The maximum size of the slice (or ULONG_MAX).
+ * @max_slots: The maximum number of slots in the slice (or INT_MAX).
+ * @_nr_slots: Where to put the number of slots (updated).
+ *
+ * Determine the size and number of slots that can be obtained the next slice
+ * of bvec queue up to the maximum size and slot count specified.
+ *
+ * @pos is updated to the end of the slice. If the position hits the end of
+ * the queue, then it is left pointing beyond the last slot of the last bvecq
+ * so that it doesn't break the chain.
+ *
+ * Return: The number of bytes in the slice.
+ */
+size_t bvecq_slice(struct bvecq_pos *pos, size_t max_size,
+ unsigned int max_slots, unsigned int *_nr_slots)
+{
+ struct bvecq *bq, *next;
+ unsigned int slot = pos->slot, nslots = 0;
+ size_t size = 0, offset = pos->offset;
+
+ bq = pos->bvecq;
+ for (;;) {
+ for (; slot < bvecq_nr_slots_acquire(bq); slot++) {
+ const struct bio_vec *bvec = &bq->bv[slot];
+
+ if (offset < bvec->bv_len && bvec->bv_page) {
+ size_t part = min(bvec->bv_len - offset, max_size);
+
+ size += part;
+ offset += part;
+ max_size -= part;
+ nslots++;
+ if (!max_size || nslots >= max_slots)
+ goto out;
+ }
+ offset = 0;
+ }
+
+ /* pos->bvecq isn't allowed to go NULL as the queue may get
+ * extended and we would lose our place.
+ */
+ next = bvecq_next(bq);
+ if (!next)
+ break;
+ if (bvecq_acquire_slot(bq, slot))
+ continue; /* More slots got added. */
+ slot = 0;
+ bq = next;
+ }
+
+out:
+ *_nr_slots = nslots;
+ if (slot == bvecq_nr_slots_acquire(bq)) {
+ next = bvecq_next(bq);
+ if (next) {
+ bq = next;
+ slot = 0;
+ offset = 0;
+ }
+ }
+ bvecq_pos_move(pos, bq);
+ pos->slot = slot;
+ pos->offset = offset;
+ return size;
+}
+
+/**
+ * bvecq_load_from_ra - Allocate a bvecq chain and load from readahead
+ * @pos: Blank position object to attach the new chain to.
+ * @ractl: The readahead control context.
+ *
+ * Decant the set of folios to be read from the readahead context into a bvecq
+ * chain. Each folio occupies one bio_vec element.
+ *
+ * Return: Amount of data loaded or -ENOMEM on allocation failure.
+ */
+ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl)
+{
+ XA_STATE(xas, &ractl->mapping->i_pages, ractl->_index);
+ struct folio *folio;
+ struct bvecq *bq;
+ unsigned int slot = 0;
+ size_t loaded = 0;
+
+ bq = bvecq_alloc_chain(ractl->_nr_folios, GFP_KERNEL, false);
+ if (!bq)
+ return -ENOMEM;
+
+ pos->bvecq = bq;
+ pos->slot = 0;
+ pos->offset = 0;
+
+ rcu_read_lock();
+
+ xas_for_each(&xas, folio, ractl->_index + ractl->_nr_pages - 1) {
+ size_t len;
+
+ if (xas_retry(&xas, folio))
+ continue;
+ VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
+
+ len = folio_size(folio);
+ bvec_set_folio(&bq->bv[slot], folio, len, 0);
+ loaded += len;
+ slot++;
+ trace_netfs_folio(folio, netfs_folio_trace_read);
+
+ if (slot >= bq->max_slots) {
+ bvecq_filled_to(bq, slot);
+ bq = bq->next;
+ if (!bq)
+ break;
+ slot = 0;
+ }
+ }
+
+ rcu_read_unlock();
+
+ if (bq)
+ bvecq_filled_to(bq, slot);
+
+ ractl->_index += ractl->_nr_pages;
+ ractl->_nr_pages = 0;
+ return loaded;
+}
diff --git a/fs/netfs/direct_read.c b/fs/netfs/direct_read.c
index 8c15f30797238..dae890e8df285 100644
--- a/fs/netfs/direct_read.c
+++ b/fs/netfs/direct_read.c
@@ -16,44 +16,21 @@
#include <linux/netfs.h>
#include "internal.h"
-static void netfs_prepare_dio_read_iterator(struct netfs_io_subrequest *subreq)
-{
- struct netfs_io_request *rreq = subreq->rreq;
- size_t rsize;
-
- rsize = umin(subreq->len, rreq->io_streams[0].sreq_max_len);
- subreq->len = rsize;
-
- if (unlikely(rreq->io_streams[0].sreq_max_segs)) {
- size_t limit = netfs_limit_iter(&rreq->buffer.iter, 0, rsize,
- rreq->io_streams[0].sreq_max_segs);
-
- if (limit < rsize) {
- subreq->len = limit;
- trace_netfs_sreq(subreq, netfs_sreq_trace_limited);
- }
- }
-
- trace_netfs_sreq(subreq, netfs_sreq_trace_prepare);
-
- subreq->io_iter = rreq->buffer.iter;
- iov_iter_truncate(&subreq->io_iter, subreq->len);
- iov_iter_advance(&rreq->buffer.iter, subreq->len);
-}
-
/*
* Perform a read to a buffer from the server, slicing up the region to be read
* according to the network rsize.
*/
static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)
{
+ struct netfs_io_stream *stream = &rreq->io_streams[0];
ssize_t size = rreq->len;
uoff_t start = rreq->start;
int ret;
+ bvecq_pos_set(&rreq->collect_cursor, &rreq->dispatch_cursor);
+
do {
struct netfs_io_subrequest *subreq;
- ssize_t slice;
subreq = netfs_alloc_subrequest(rreq, NETFS_DOWNLOAD_FROM_SERVER);
if (!subreq) {
@@ -78,14 +55,22 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)
}
}
- netfs_prepare_dio_read_iterator(subreq);
- slice = subreq->len;
- size -= slice;
- start += slice;
- rreq->submitted += slice;
+ bvecq_pos_set(&subreq->io_buffer, &rreq->dispatch_cursor);
+ subreq->len = bvecq_slice(&rreq->dispatch_cursor,
+ umin(size, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
+
+ size -= subreq->len;
+ start += subreq->len;
+ rreq->submitted += subreq->len;
if (size <= 0)
netfs_all_subreqs_queued(rreq);
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset,
+ subreq->len);
+
rreq->netfs_ops->issue_read(subreq);
if (test_bit(NETFS_RREQ_PAUSE, &rreq->flags))
@@ -99,6 +84,8 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)
netfs_all_subreqs_queued(rreq);
netfs_wake_collector(rreq);
}
+
+ bvecq_pos_unset(&rreq->dispatch_cursor);
}
/*
@@ -177,25 +164,17 @@ ssize_t netfs_unbuffered_read_iter_locked(struct kiocb *iocb, struct iov_iter *i
* buffer for ourselves as the caller's iterator will be trashed when
* we return.
*
- * In such a case, extract an iterator to represent as much of the the
- * output buffer as we can manage. Note that the extraction might not
- * be able to allocate a sufficiently large bvec array and may shorten
- * the request.
+ * Extract a buffer queue to represent as much of the output buffer as
+ * we can manage. The fragments are extracted into a bvecq which will
+ * have sufficient nodes allocated to hold all the data, though this
+ * may end up truncated if ENOMEM is encountered.
*/
- if (user_backed_iter(iter)) {
- ret = netfs_extract_user_iter(iter, rreq->len, &rreq->buffer.iter, 0);
- if (ret < 0)
- goto error_put;
- rreq->direct_bv = (struct bio_vec *)rreq->buffer.iter.bvec;
- rreq->direct_bv_count = ret;
- rreq->direct_bv_unpin = iov_iter_extract_will_pin(iter);
- rreq->len = iov_iter_count(&rreq->buffer.iter);
- } else {
- rreq->buffer.iter = *iter;
- rreq->len = orig_count;
- rreq->direct_bv_unpin = false;
- iov_iter_advance(iter, orig_count);
- }
+ ret = netfs_extract_iter(iter, rreq->len, INT_MAX,
+ &rreq->dispatch_cursor.bvecq, 0, rreq->gfp);
+ if (ret < 0)
+ goto error_put;
+
+ rreq->len = ret;
// TODO: Set up bounce buffer if needed
diff --git a/fs/netfs/direct_write.c b/fs/netfs/direct_write.c
index cc46b7d9321f1..65c61fc67f9bf 100644
--- a/fs/netfs/direct_write.c
+++ b/fs/netfs/direct_write.c
@@ -73,7 +73,11 @@ static void netfs_unbuffered_write_collect(struct netfs_io_request *wreq,
spin_unlock(&wreq->lock);
wreq->transferred += subreq->transferred;
- iov_iter_advance(&wreq->buffer.iter, subreq->transferred);
+ if (subreq->transferred < subreq->len) {
+ bvecq_pos_unset(&wreq->dispatch_cursor);
+ bvecq_pos_transfer(&wreq->dispatch_cursor, &subreq->io_buffer);
+ bvecq_pos_advance(&wreq->dispatch_cursor, subreq->transferred);
+ }
stream->collected_to = subreq->start + subreq->transferred;
wreq->collected_to = stream->collected_to;
@@ -99,6 +103,8 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
_enter("%llx", wreq->len);
+ bvecq_pos_set(&wreq->collect_cursor, &wreq->dispatch_cursor);
+
if (wreq->origin == NETFS_DIO_WRITE)
inode_dio_begin(wreq->inode);
@@ -116,6 +122,8 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
break;
}
stream->construct = NULL;
+ } else {
+ bvecq_pos_set(&subreq->io_buffer, &wreq->dispatch_cursor);
}
/* Check if (re-)preparation failed. */
@@ -125,9 +133,16 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
break;
}
- iov_iter_truncate(&subreq->io_iter, wreq->len - wreq->transferred);
+ subreq->len = bvecq_slice(&wreq->dispatch_cursor, stream->sreq_max_len,
+ stream->sreq_max_segs, &subreq->nr_segs);
+
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_SOURCE,
+ subreq->io_buffer.bvecq, subreq->io_buffer.slot,
+ subreq->io_buffer.offset,
+ subreq->len);
+
if (!iov_iter_count(&subreq->io_iter)) {
- pr_warn("netfs: Unexpected zero-length iterator R=%08x\n",
+ pr_warn("netfs: Unexpected zero-length slice R=%08x\n",
wreq->debug_id);
__set_bit(NETFS_SREQ_FAILED, &subreq->flags);
netfs_write_subrequest_terminated(subreq, -EIO);
@@ -135,12 +150,6 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
break;
}
- subreq->len = netfs_limit_iter(&subreq->io_iter, 0,
- stream->sreq_max_len,
- stream->sreq_max_segs);
- iov_iter_truncate(&subreq->io_iter, subreq->len);
- stream->submit_extendable_to = subreq->len;
-
trace_netfs_sreq(subreq, netfs_sreq_trace_submit);
stream->issue_write(subreq);
@@ -175,9 +184,13 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
*/
subreq->error = -EAGAIN;
trace_netfs_sreq(subreq, netfs_sreq_trace_retry);
+
+ bvecq_pos_unset(&wreq->dispatch_cursor);
+ bvecq_pos_transfer(&wreq->dispatch_cursor, &subreq->io_buffer);
+
if (subreq->transferred > 0) {
- iov_iter_advance(&wreq->buffer.iter, subreq->transferred);
wreq->transferred += subreq->transferred;
+ bvecq_pos_advance(&wreq->dispatch_cursor, subreq->transferred);
}
if (stream->source == NETFS_UPLOAD_TO_SERVER &&
@@ -188,7 +201,6 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
__clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
__clear_bit(NETFS_SREQ_BOUNDARY, &subreq->flags);
__clear_bit(NETFS_SREQ_FAILED, &subreq->flags);
- subreq->io_iter = wreq->buffer.iter;
subreq->start = wreq->start + wreq->transferred;
subreq->len = wreq->len - wreq->transferred;
subreq->transferred = 0;
@@ -204,6 +216,7 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
netfs_stat(&netfs_n_wh_retry_write_subreq);
}
+ bvecq_pos_unset(&wreq->dispatch_cursor);
netfs_unbuffered_write_done(wreq);
_leave(" = %d", ret);
return ret;
@@ -222,10 +235,10 @@ static void netfs_unbuffered_write_async(struct work_struct *work)
* encrypted file. This can also be used for direct I/O writes.
*/
ssize_t netfs_unbuffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *iter,
- struct netfs_group *netfs_group)
+ struct netfs_group *netfs_group)
{
struct netfs_io_request *wreq;
- ssize_t ret, n;
+ ssize_t ret;
uoff_t start = iocb->ki_pos;
uoff_t end = start + iov_iter_count(iter);
size_t len = iov_iter_count(iter);
@@ -261,25 +274,17 @@ ssize_t netfs_unbuffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *
* allocate a sufficiently large bvec array and may shorten the
* request.
*/
- if (user_backed_iter(iter)) {
- n = netfs_extract_user_iter(iter, len, &wreq->buffer.iter, 0);
- if (n < 0) {
- ret = n;
- goto error_put;
- }
- wreq->direct_bv = (struct bio_vec *)wreq->buffer.iter.bvec;
- wreq->direct_bv_count = n;
- wreq->direct_bv_unpin = iov_iter_extract_will_pin(iter);
- } else {
- /* If this is a kernel-generated async DIO request,
- * assume that any resources the iterator points to
- * (eg. a bio_vec array) will persist till the end of
- * the op.
- */
- wreq->buffer.iter = *iter;
- }
+ ssize_t n = netfs_extract_iter(iter, len, INT_MAX,
+ &wreq->dispatch_cursor.bvecq, 0, wreq->gfp);
- wreq->len = iov_iter_count(&wreq->buffer.iter);
+ if (n < 0) {
+ ret = n;
+ goto error_put;
+ }
+ wreq->len = n;
+ _debug("dio-write %zx/%zx %u/%u",
+ n, len, wreq->dispatch_cursor.bvecq->nr_slots,
+ wreq->dispatch_cursor.bvecq->max_slots);
}
__set_bit(NETFS_RREQ_USE_IO_ITER, &wreq->flags);
diff --git a/fs/netfs/internal.h b/fs/netfs/internal.h
index f2a86abae9b3e..2760bce732b83 100644
--- a/fs/netfs/internal.h
+++ b/fs/netfs/internal.h
@@ -70,7 +70,6 @@ static inline void netfs_proc_del_rreq(struct netfs_io_request *rreq) {}
/*
* misc.c
*/
-void netfs_reset_iter(struct netfs_io_subrequest *subreq);
void netfs_wake_collector(struct netfs_io_request *rreq);
void netfs_subreq_clear_in_progress(struct netfs_io_subrequest *subreq);
void netfs_wait_for_in_progress_stream(struct netfs_io_request *rreq,
@@ -239,8 +238,7 @@ void netfs_prepare_write(struct netfs_io_request *wreq,
struct netfs_io_stream *stream,
uoff_t start);
void netfs_reissue_write(struct netfs_io_stream *stream,
- struct netfs_io_subrequest *subreq,
- struct iov_iter *source);
+ struct netfs_io_subrequest *subreq);
void netfs_issue_write(struct netfs_io_request *wreq,
struct netfs_io_stream *stream);
size_t netfs_advance_write(struct netfs_io_request *wreq,
diff --git a/fs/netfs/iterator.c b/fs/netfs/iterator.c
index 31748526d5682..dc97e5b0d4495 100644
--- a/fs/netfs/iterator.c
+++ b/fs/netfs/iterator.c
@@ -14,296 +14,144 @@
#include "internal.h"
/**
- * netfs_extract_user_iter - Extract the pages from a user iterator into a bvec
+ * netfs_extract_iter - Extract virtually contiguous pages from an iterator into a bvecq
* @orig: The original iterator
- * @orig_len: The amount of iterator to copy
- * @new: The iterator to be set up
+ * @max_len: Maximum number of bytes to extract
+ * @max_pages: Maximum number of pages to extract
+ * @_bvecq_head: Where to cache the bvec queue
* @extraction_flags: Flags to qualify the request
+ * @gfp: Allocation mode for bvecq structs.
*
- * Extract the page fragments from the given amount of the source iterator and
- * build up a second iterator that refers to all of those bits. This allows
- * the original iterator to be disposed of.
+ * Extract virtually contiguous page fragments from the source iterator up to
+ * the given maxima and build bvec queue that refers to all of those bits.
+ * This allows the original iterator to disposed of.
*
- * @extraction_flags can have ITER_ALLOW_P2PDMA set to request peer-to-peer DMA be
- * allowed on the pages extracted.
+ * @extraction_flags can have ITER_ALLOW_P2PDMA set to request peer-to-peer DMA
+ * be allowed on the pages extracted.
*
- * On success, the number of elements in the bvec is returned, the original
- * iterator will have been advanced by the amount extracted.
+ * On success or partial success, the amount of data in the bvec is returned,
+ * the original iterator will have been advanced by the amount extracted.
*
- * The iov_iter_extract_mode() function should be used to query how cleanup
- * should be performed.
+ * If an error occurs and no pages are extracted, an error will be returned and
+ * any allocated bvecq will be freed. If there is no data to be extracted (or
+ * @max_len or @max_pages are zero), a single empty bvecq will be returned.
+ *
+ * The bvecq segments are marked with indications on how to get clean up the
+ * extracted fragments.
*/
-ssize_t netfs_extract_user_iter(struct iov_iter *orig, size_t orig_len,
- struct iov_iter *new,
- iov_iter_extraction_t extraction_flags)
+ssize_t netfs_extract_iter(struct iov_iter *orig, size_t max_len, size_t max_pages,
+ struct bvecq **_bvecq_head,
+ iov_iter_extraction_t extraction_flags, gfp_t gfp)
{
- struct bio_vec *bv = NULL;
- struct page **pages;
- unsigned int cur_npages;
- unsigned int max_pages;
- unsigned int npages = 0;
- unsigned int i;
+ struct bvecq *bq_tail = NULL, *bq;
ssize_t ret = 0;
- size_t count = orig_len, offset, len;
- size_t bv_size, pg_size;
+ size_t extracted = 0;
- if (WARN_ON_ONCE(!iter_is_ubuf(orig) && !iter_is_iovec(orig)))
- return -EIO;
+ _enter("{%u,%zx},%zx", orig->iter_type, orig->count, max_len);
- max_pages = iov_iter_npages(orig, INT_MAX);
- bv_size = array_size(max_pages, sizeof(*bv));
- bv = kvmalloc(bv_size, GFP_KERNEL);
- if (!bv)
- return -ENOMEM;
+ *_bvecq_head = NULL;
+ if (max_len > orig->count)
+ max_len = orig->count;
+ if (!max_len || !max_pages)
+ goto alloc_empty;
+ if (WARN_ON_ONCE(max_pages > INT_MAX))
+ max_pages = INT_MAX; /* Protect iov_iter_npages(). */
- /* Put the page list at the end of the bvec list storage. bvec
- * elements are larger than page pointers, so as long as we work
- * 0->last, we should be fine.
- */
- pg_size = array_size(max_pages, sizeof(*pages));
- pages = (void *)bv + bv_size - pg_size;
+ max_pages = iov_iter_npages(orig, max_pages);
+ if (!max_pages)
+ goto alloc_empty;
- while (count && npages < max_pages) {
- ret = iov_iter_extract_pages(orig, &pages, count,
- max_pages - npages, extraction_flags,
- &offset);
- if (unlikely(ret <= 0)) {
- ret = ret ?: -EIO;
+ do {
+ bq = bvecq_alloc_one(max_pages, gfp, false);
+ if (!bq) {
+ ret = -ENOMEM;
break;
}
+ if (user_backed_iter(orig))
+ bq->mem_type = iov_iter_extract_will_pin(orig) ?
+ BVECQ_MEM_GUP : BVECQ_MEM_PAGECACHE;
- if (WARN(ret > count,
- "%s: extract_pages overrun %zd > %zu bytes\n",
- __func__, ret, count)) {
- ret = -EIO;
- break;
- }
+ if (bq_tail)
+ bvecq_append(bq_tail, bq);
+ else
+ *_bvecq_head = bq;
+ bq_tail = bq;
- cur_npages = DIV_ROUND_UP(offset + ret, PAGE_SIZE);
- if (WARN(cur_npages > max_pages - npages,
- "%s: extract_pages overrun %u > %u pages\n",
- __func__, npages + cur_npages, max_pages)) {
- ret = -EIO;
+ if (max_len == 0)
break;
- }
-
- count -= ret;
- ret += offset;
-
- for (i = 0; i < cur_npages; i++) {
- len = ret > PAGE_SIZE ? PAGE_SIZE : ret;
- bvec_set_page(bv + npages + i, *pages++, len - offset, offset);
- ret -= len;
- offset = 0;
- }
-
- npages += cur_npages;
- }
-
- /* Note: Don't try to clean up after EIO. Either we got no pages, so
- * nothing to clean up, or we got a buffer overrun, memory corruption
- * and can't trust the stuff in the buffer (a WARN was emitted).
- */
-
- if (ret < 0 && (ret == -ENOMEM || npages == 0)) {
- for (i = 0; i < npages; i++)
- unpin_user_page(bv[i].bv_page);
- kvfree(bv);
- return ret;
- }
- iov_iter_bvec(new, orig->data_source, bv, npages, orig_len - count);
- return npages;
-}
-EXPORT_SYMBOL_GPL(netfs_extract_user_iter);
-
-/*
- * Select the span of a bvec iterator we're going to use. Limit it by both maximum
- * size and maximum number of segments. Returns the size of the span in bytes.
- */
-static size_t netfs_limit_bvec(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- const struct bio_vec *bvecs = iter->bvec;
- unsigned int nbv = iter->nr_segs, ix = 0, nsegs = 0;
- size_t len, span = 0, n = iter->count;
- size_t skip = iter->iov_offset + start_offset;
-
- if (WARN_ON(!iov_iter_is_bvec(iter)) ||
- WARN_ON(start_offset > n) ||
- n == 0)
- return 0;
-
- while (n && ix < nbv && skip) {
- len = bvecs[ix].bv_len;
- if (skip < len)
- break;
- skip -= len;
- n -= len;
- ix++;
- }
-
- while (n && ix < nbv) {
- len = min3(n, bvecs[ix].bv_len - skip, max_size);
- span += len;
- nsegs++;
- ix++;
- if (span >= max_size || nsegs >= max_segs)
- break;
- skip = 0;
- n -= len;
- }
-
- return min(span, max_size);
-}
-
-/*
- * Select the span of a kvec iterator we're going to use. Limit it by both
- * maximum size and maximum number of segments. Returns the size of the span
- * in bytes.
- */
-static size_t netfs_limit_kvec(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- const struct kvec *kvecs = iter->kvec;
- unsigned int nkv = iter->nr_segs, ix = 0, nsegs = 0;
- size_t len, span = 0, n = iter->count;
- size_t skip = iter->iov_offset + start_offset;
-
- if (WARN_ON(!iov_iter_is_kvec(iter)) ||
- WARN_ON(start_offset > n) ||
- n == 0)
- return 0;
-
- while (n && ix < nkv && skip) {
- len = kvecs[ix].iov_len;
- if (skip < len)
- break;
- skip -= len;
- n -= len;
- ix++;
- }
-
- while (n && ix < nkv) {
- len = min3(n, kvecs[ix].iov_len - skip, max_size);
- span += len;
- nsegs++;
- ix++;
- if (span >= max_size || nsegs >= max_segs)
- break;
- skip = 0;
- n -= len;
- }
-
- return min(span, max_size);
-}
-
-/*
- * Select the span of an xarray iterator we're going to use. Limit it by both
- * maximum size and maximum number of segments. It is assumed that segments
- * can be larger than a page in size, provided they're physically contiguous.
- * Returns the size of the span in bytes.
- */
-static size_t netfs_limit_xarray(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- struct folio *folio;
- unsigned int nsegs = 0;
- uoff_t pos = iter->xarray_start + iter->iov_offset;
- pgoff_t index = pos / PAGE_SIZE;
- size_t span = 0, n = iter->count;
-
- XA_STATE(xas, iter->xarray, index);
-
- if (WARN_ON(!iov_iter_is_xarray(iter)) ||
- WARN_ON(start_offset > n) ||
- n == 0)
- return 0;
- max_size = min(max_size, n - start_offset);
-
- rcu_read_lock();
- xas_for_each(&xas, folio, ULONG_MAX) {
- size_t offset, flen, len;
- if (xas_retry(&xas, folio))
- continue;
- if (WARN_ON(xa_is_value(folio)))
- break;
- if (WARN_ON(folio_test_hugetlb(folio)))
- break;
-
- flen = folio_size(folio);
- offset = offset_in_folio(folio, pos);
- len = min(max_size, flen - offset);
- span += len;
- nsegs++;
- if (span >= max_size || nsegs >= max_segs)
- break;
- }
-
- rcu_read_unlock();
- return min(span, max_size);
-}
-
-/*
- * Select the span of a bvecq iterator we're going to use. Limit it by both
- * maximum size and maximum number of segments. Returns the size of the span
- * in bytes.
- */
-static size_t netfs_limit_bvecq(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- const struct bvecq *bq = iter->bvecq;
- unsigned int nsegs = 0;
- unsigned int slot = iter->bvecq_slot;
- size_t span = 0, n = iter->count;
-
- if (WARN_ON(!iov_iter_is_bvecq(iter)) ||
- WARN_ON(start_offset > n) ||
- n == 0)
- return 0;
- max_size = umin(max_size, n - start_offset);
-
- if (!bvecq_acquire_slot(bq, slot)) {
- bq = bvecq_next(bq);
- slot = 0;
- }
-
- start_offset += iter->iov_offset;
- do {
- size_t flen;
-
- flen = bq->bv[slot].bv_len;
- if (start_offset < flen) {
- span += flen - start_offset;
- nsegs++;
- start_offset = 0;
- } else {
- start_offset -= flen;
- }
- if (span >= max_size || nsegs >= max_segs)
- break;
-
- slot++;
- if (!bvecq_acquire_slot(bq, slot)) {
- bq = bvecq_next(bq);
- slot = 0;
- }
- } while (bq);
-
- return umin(span, max_size);
-}
+ struct bio_vec *bv = bq->bv;
+ unsigned int slot = 0;
+ do {
+ struct page **pages;
+ ssize_t got;
+ size_t offset;
+ size_t space = bq->max_slots - slot;
+ size_t bv_size = array_size(bq->max_slots, sizeof(*bv));
+ size_t pg_size = array_size(space, sizeof(*pages));
+
+ /* Put the page list at the end of the bvec list
+ * storage. bvec elements are larger than page
+ * pointers, so as long as we work 0->last, we should
+ * be fine.
+ */
+ pages = (void *)bv + bv_size - pg_size;
+
+ got = iov_iter_extract_pages(orig, &pages, max_len,
+ min(space, max_pages),
+ extraction_flags, &offset);
+ if (got < 0) {
+ ret = got;
+ goto out;
+ }
+
+ if (got == 0) {
+ pr_err("extract_pages gave nothing from %zx, %zx\n",
+ extracted, max_len);
+ ret = -EIO;
+ goto out;
+ }
+
+ if (WARN(got > max_len,
+ "%s: extract_pages overrun %zx > %zx bytes\n",
+ __func__, got, max_len)) {
+ ret = -EIO;
+ goto out;
+ }
+
+ extracted += got;
+ max_len -= got;
+
+ do {
+ size_t len = umin(got, PAGE_SIZE - offset);
+
+ BUG_ON(slot >= bq->max_slots);
+
+ bvec_set_page(&bq->bv[slot], *pages++, len, offset);
+ slot++;
+ max_pages--;
+ got -= len;
+ offset = 0;
+ } while (got > 0);
+
+ bvecq_filled_to(bq, slot);
+ } while (max_len > 0 && max_pages > 0 && !bvecq_is_full(bq));
+
+ } while (max_len > 0 && max_pages > 0);
+
+out:
+ if (extracted || ret == 0)
+ return extracted;
+ bvecq_put(*_bvecq_head);
+ *_bvecq_head = NULL;
+ return ret;
+
+alloc_empty:
+ bq = bvecq_alloc_one(1, gfp, false);
+ if (!bq)
+ return -ENOMEM;
+ *_bvecq_head = bq;
+ return 0;
-size_t netfs_limit_iter(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- if (iov_iter_is_bvecq(iter))
- return netfs_limit_bvecq(iter, start_offset, max_size, max_segs);
- if (iov_iter_is_bvec(iter))
- return netfs_limit_bvec(iter, start_offset, max_size, max_segs);
- if (iov_iter_is_xarray(iter))
- return netfs_limit_xarray(iter, start_offset, max_size, max_segs);
- if (iov_iter_is_kvec(iter))
- return netfs_limit_kvec(iter, start_offset, max_size, max_segs);
- BUG();
}
-EXPORT_SYMBOL(netfs_limit_iter);
+EXPORT_SYMBOL_GPL(netfs_extract_iter);
diff --git a/fs/netfs/misc.c b/fs/netfs/misc.c
index a0cc248a284d5..130bd432b1948 100644
--- a/fs/netfs/misc.c
+++ b/fs/netfs/misc.c
@@ -9,24 +9,6 @@
#include <linux/rmap.h>
#include "internal.h"
-/*
- * Reset the subrequest iterator to refer just to the region remaining to be
- * read. The iterator may or may not have been advanced by socket ops or
- * extraction ops to an extent that may or may not match the amount actually
- * read.
- */
-void netfs_reset_iter(struct netfs_io_subrequest *subreq)
-{
- struct iov_iter *io_iter = &subreq->io_iter;
- size_t remain = subreq->len - subreq->transferred;
-
- if (io_iter->count > remain)
- iov_iter_advance(io_iter, io_iter->count - remain);
- else if (io_iter->count < remain)
- iov_iter_revert(io_iter, remain - io_iter->count);
- iov_iter_truncate(&subreq->io_iter, remain);
-}
-
/**
* netfs_dirty_folio - Mark folio dirty and pin a cache object for writeback
* @mapping: The mapping the folio belongs to.
diff --git a/fs/netfs/objects.c b/fs/netfs/objects.c
index 4b8d20559b0e1..bf17dc31fd9d8 100644
--- a/fs/netfs/objects.c
+++ b/fs/netfs/objects.c
@@ -133,7 +133,6 @@ static void netfs_free_request_rcu(struct rcu_head *rcu)
static void netfs_deinit_request(struct netfs_io_request *rreq)
{
struct netfs_inode *ictx = netfs_inode(rreq->inode);
- unsigned int i;
trace_netfs_rreq(rreq, netfs_rreq_trace_free);
@@ -148,16 +147,10 @@ static void netfs_deinit_request(struct netfs_io_request *rreq)
rreq->netfs_ops->free_request(rreq);
if (rreq->cache_resources.ops)
rreq->cache_resources.ops->end_operation(&rreq->cache_resources);
- if (rreq->direct_bv) {
- for (i = 0; i < rreq->direct_bv_count; i++) {
- if (rreq->direct_bv[i].bv_page) {
- if (rreq->direct_bv_unpin)
- unpin_user_page(rreq->direct_bv[i].bv_page);
- }
- }
- kvfree(rreq->direct_bv);
- }
- rolling_buffer_clear(&rreq->buffer);
+ bvecq_pos_unset(&rreq->load_cursor);
+ bvecq_pos_unset(&rreq->dispatch_cursor);
+ bvecq_pos_unset(&rreq->collect_cursor);
+ bvecq_put(rreq->spare);
if (atomic_dec_and_test(&ictx->io_count))
wake_up_var(&ictx->io_count);
@@ -251,6 +244,7 @@ static void netfs_free_subrequest(struct netfs_io_subrequest *subreq)
trace_netfs_sreq(subreq, netfs_sreq_trace_free);
if (rreq->netfs_ops->free_subrequest)
rreq->netfs_ops->free_subrequest(subreq);
+ bvecq_pos_unset(&subreq->io_buffer);
mempool_free(subreq, rreq->netfs_ops->subrequest_pool ?: &netfs_subrequest_pool);
netfs_stat_d(&netfs_n_rh_sreq);
netfs_put_request(rreq, netfs_rreq_trace_put_subreq);
diff --git a/fs/netfs/read_collect.c b/fs/netfs/read_collect.c
index 75c8874ec4595..368f14cd2d982 100644
--- a/fs/netfs/read_collect.c
+++ b/fs/netfs/read_collect.c
@@ -26,27 +26,35 @@
*/
static void netfs_clear_unread(struct netfs_io_subrequest *subreq)
{
- netfs_reset_iter(subreq);
- WARN_ON_ONCE(subreq->len - subreq->transferred != iov_iter_count(&subreq->io_iter));
- iov_iter_zero(iov_iter_count(&subreq->io_iter), &subreq->io_iter);
+ struct iov_iter iter;
+
+ iov_iter_bvec_queue(&iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
+ iov_iter_advance(&iter, subreq->transferred);
+ iov_iter_zero(subreq->len, &iter);
+
if (subreq->start + subreq->transferred >= subreq->rreq->i_size)
__set_bit(NETFS_SREQ_HIT_EOF, &subreq->flags);
}
static void netfs_clear_unread_dio(struct netfs_io_subrequest *subreq)
{
- uoff_t pos = subreq->start + subreq->transferred;
struct netfs_io_request *rreq = subreq->rreq;
+ struct iov_iter iter;
+ uoff_t pos = subreq->start + subreq->transferred;
size_t fill;
if (pos >= rreq->i_size)
return;
- fill = min_t(uoff_t, rreq->i_size - pos,
- subreq->len - subreq->transferred);
+ fill = umin(rreq->i_size - pos, subreq->len - subreq->transferred);
- netfs_reset_iter(subreq);
- subreq->transferred += iov_iter_zero(fill, &subreq->io_iter);
+ iov_iter_bvec_queue(&iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
+ iov_iter_advance(&iter, subreq->transferred);
+ iov_iter_zero(fill, &iter);
+
+ subreq->transferred += iov_iter_zero(fill, &iter);
}
/*
@@ -138,8 +146,8 @@ static void netfs_unlock_read_folio(struct netfs_io_request *rreq,
*/
void netfs_read_set_unlock_at(struct netfs_io_request *rreq)
{
- const struct bvecq *bq = rreq->buffer.tail;
- unsigned int slot = rreq->buffer.first_tail_slot;
+ const struct bvecq *bq = rreq->collect_cursor.bvecq;
+ unsigned int slot = rreq->collect_cursor.slot;
size_t cleaned_to = rreq->cleaned_to - rreq->start;
size_t progress_at = cleaned_to;
size_t minimum = 256 * 1024;
@@ -169,8 +177,8 @@ void netfs_read_set_unlock_at(struct netfs_io_request *rreq)
static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
unsigned int *notes)
{
- struct bvecq *bq = rreq->buffer.tail;
- unsigned int slot = rreq->buffer.first_tail_slot;
+ struct bvecq *bq = rreq->collect_cursor.bvecq;
+ unsigned int slot = rreq->collect_cursor.slot;
uoff_t collected_to = rreq->collected_to;
if (rreq->cleaned_to >= rreq->collected_to)
@@ -178,15 +186,6 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
// TODO: Begin decryption
- while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&rreq->buffer);
- if (!bq) {
- WRITE_ONCE(rreq->progress_at, rreq->len);
- return;
- }
- slot = 0;
- }
-
/* We have to wait for readahead refs to have been released before we
* can unlock any folios as the ref-dropper walks i_pages and the only
* thing preventing these folios from being removed is the folio lock.
@@ -196,9 +195,24 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
for (;;) {
struct folio *folio;
- uoff_t fpos, fend;
+ uoff_t fpos = rreq->cleaned_to, fend;
size_t fsize;
+ /* Clean up the head bvecq segment. If we clear an entire
+ * segment, then we can get rid of it provided it's not also
+ * the tail segment being filled by the issuer.
+ */
+ if (!bvecq_acquire_slot(bq, slot)) {
+ rreq->collect_cursor.slot = slot;
+ if (!bvecq_delete_spent(&rreq->collect_cursor)) {
+ WRITE_ONCE(rreq->progress_at, rreq->len);
+ return;
+ }
+ bq = rreq->collect_cursor.bvecq;
+ slot = rreq->collect_cursor.slot;
+ continue;
+ }
+
folio = bvec_folio(&bq->bv[slot]);
if (WARN_ONCE(!folio_test_locked(folio),
"R=%08x: folio %lx is not locked\n",
@@ -206,7 +220,6 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
trace_netfs_folio(folio, netfs_folio_trace_not_locked);
fsize = bq->bv[slot].bv_len;
- fpos = folio_pos(folio);
fend = fpos + fsize;
trace_netfs_collect_folio(rreq, folio);
@@ -216,30 +229,16 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
break;
netfs_unlock_read_folio(rreq, bq, slot);
- WRITE_ONCE(rreq->cleaned_to, fpos + fsize);
- *notes |= MADE_PROGRESS;
-
- /* Clean up the head bq. If we clear an entire bq, then
- * we can get rid of it provided it's not also the tail bq
- * being filled by the issuer.
- */
- bq->bv[slot].bv_page = NULL;
slot++;
- while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&rreq->buffer);
- if (!bq)
- goto done;
- slot = 0;
- }
+ WRITE_ONCE(rreq->cleaned_to, fend);
+ *notes |= MADE_PROGRESS;
if (fpos + fsize >= collected_to)
break;
}
- rreq->buffer.tail = bq;
-done:
- rreq->buffer.first_tail_slot = slot;
-
+ bvecq_pos_move(&rreq->collect_cursor, bq);
+ rreq->collect_cursor.slot = slot;
netfs_read_set_unlock_at(rreq);
}
@@ -422,12 +421,15 @@ static void netfs_rreq_assess_dio(struct netfs_io_request *rreq)
if (rreq->origin == NETFS_UNBUFFERED_READ ||
rreq->origin == NETFS_DIO_READ) {
- for (i = 0; i < rreq->direct_bv_count; i++) {
- flush_dcache_page(rreq->direct_bv[i].bv_page);
- // TODO: cifs marks pages in the destination buffer
- // dirty under some circumstances after a read. Do we
- // need to do that too?
- set_page_dirty(rreq->direct_bv[i].bv_page);
+ for (struct bvecq *bq = rreq->collect_cursor.bvecq; bq; bq = bvecq_next(bq)) {
+ unsigned int nr_slots = bvecq_nr_slots_acquire(bq);
+ /* Read the slot count before the slots. */
+
+ /* Mark the target buffers dirty. */
+ for (i = 0; i < nr_slots; i++) {
+ flush_dcache_page(bq->bv[i].bv_page);
+ set_page_dirty(bq->bv[i].bv_page);
+ }
}
}
@@ -521,7 +523,15 @@ bool netfs_read_collection(struct netfs_io_request *rreq)
trace_netfs_rreq(rreq, netfs_rreq_trace_done);
netfs_clear_subrequests(rreq);
- netfs_unlock_abandoned_read_pages(rreq);
+ switch (rreq->origin) {
+ case NETFS_READAHEAD:
+ case NETFS_READPAGE:
+ case NETFS_READ_FOR_WRITE:
+ netfs_unlock_abandoned_read_pages(rreq);
+ break;
+ default:
+ break;
+ }
if (unlikely(rreq->copy_to_cache))
netfs_pgpriv2_end_copy_to_cache(rreq);
return true;
diff --git a/fs/netfs/read_pgpriv2.c b/fs/netfs/read_pgpriv2.c
index f8e5667e278e1..f23c4cbfed581 100644
--- a/fs/netfs/read_pgpriv2.c
+++ b/fs/netfs/read_pgpriv2.c
@@ -19,6 +19,9 @@
static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio *folio)
{
struct netfs_io_stream *cache = &creq->io_streams[1];
+ struct bvecq *queue;
+ unsigned int slot;
+ size_t dio_size = PAGE_SIZE;
size_t fsize = folio_size(folio), flen = fsize;
uoff_t fpos = folio_pos(folio), i_size;
bool to_eof = false;
@@ -48,18 +51,37 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio
to_eof = true;
}
+ flen = round_up(flen, dio_size);
+
_debug("folio %zx %zx", flen, fsize);
trace_netfs_folio(folio, netfs_folio_trace_store_copy);
- /* Attach the folio to the rolling buffer. */
- if (rolling_buffer_append(&creq->buffer, folio, creq->gfp) < 0) {
- set_bit(NETFS_RREQ_CANCEL_CACHING, &creq->flags);
- folio_end_private_2(folio);
- return;
+ /* Institute a new bvec queue segment if the current one is full or if
+ * we encounter a discontiguity. The discontiguity break is important
+ * when it comes to bulk unlocking folios by file range.
+ */
+ queue = creq->load_cursor.bvecq;
+ if (bvecq_is_full(queue) ||
+ (fpos != creq->last_end && creq->last_end > 0 && queue->nr_slots > 0)) {
+ bvecq_buffer_append(&creq->load_cursor, creq->spare);
+ creq->spare = NULL;
+
+ queue = creq->load_cursor.bvecq;
}
- cache->submit_extendable_to = fsize;
+ /* Attach the folio to the rolling buffer. */
+ slot = queue->nr_slots;
+ bvec_set_folio(&queue->bv[slot], folio, fsize, 0);
+ trace_netfs_bv_slot(queue, slot);
+ slot++;
+ bvecq_filled_to(queue, slot);
+ creq->load_cursor.slot = slot;
+ creq->load_cursor.offset = 0;
+ creq->last_end = fpos + flen;
+
+ bvecq_pos_nudge(&creq->dispatch_cursor);
+
cache->submit_off = 0;
cache->submit_len = flen;
@@ -71,10 +93,9 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio
do {
ssize_t part;
- creq->buffer.iter.iov_offset = cache->submit_off;
+ creq->dispatch_cursor.offset = cache->submit_off;
atomic64_set(&creq->issued_to, fpos + cache->submit_off);
- cache->submit_extendable_to = fsize - cache->submit_off;
part = netfs_advance_write(creq, cache, fpos + cache->submit_off,
cache->submit_len, to_eof);
cache->submit_off += part;
@@ -84,8 +105,7 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio
cache->submit_len -= part;
} while (cache->submit_len > 0);
- creq->buffer.iter.iov_offset = 0;
- rolling_buffer_advance(&creq->buffer, fsize);
+ bvecq_pos_step(&creq->dispatch_cursor);
atomic64_set(&creq->issued_to, fpos + fsize);
if (flen < fsize)
@@ -111,6 +131,11 @@ static struct netfs_io_request *netfs_pgpriv2_begin_copy_to_cache(
if (!creq->io_streams[1].avail)
goto cancel_put;
+ if (bvecq_buffer_init(&creq->load_cursor, creq->gfp, false) < 0)
+ goto cancel_put;
+ bvecq_pos_set(&creq->dispatch_cursor, &creq->load_cursor);
+ bvecq_pos_set(&creq->collect_cursor, &creq->dispatch_cursor);
+
__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, &creq->flags);
trace_netfs_copy2cache(rreq, creq);
trace_netfs_write(creq, netfs_write_trace_copy_to_cache);
@@ -143,6 +168,14 @@ void netfs_pgpriv2_copy_to_cache(struct netfs_io_request *rreq, struct folio *fo
return;
}
+ if (!creq->spare) {
+ creq->spare = bvecq_alloc_one(BVECQ_POOL_SLOTS, creq->gfp, false);
+ if (!creq->spare) {
+ set_bit(NETFS_RREQ_CANCEL_CACHING, &creq->flags);
+ return;
+ }
+ }
+
trace_netfs_folio(folio, netfs_folio_trace_pgpriv2_copy);
netfs_pgpriv2_copy_folio(creq, folio);
}
@@ -173,16 +206,18 @@ void netfs_pgpriv2_end_copy_to_cache(struct netfs_io_request *rreq)
*/
bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)
{
- struct bvecq *bq = creq->buffer.tail;
- unsigned int slot = creq->buffer.first_tail_slot;
+ struct bvecq *bq = creq->collect_cursor.bvecq;
+ unsigned int slot;
uoff_t collected_to = creq->collected_to;
bool made_progress = false;
+ slot = creq->collect_cursor.slot;
while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&creq->buffer);
- if (!bq)
- return false;
- slot = 0;
+ creq->collect_cursor.slot = slot;
+ if (!bvecq_delete_spent(&creq->collect_cursor))
+ goto out;
+ bq = creq->collect_cursor.bvecq;
+ slot = creq->collect_cursor.slot;
}
for (;;) {
@@ -213,25 +248,25 @@ bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)
creq->cleaned_to = fpos + fsize;
made_progress = true;
- /* Clean up the head bq. If we clear an entire bq, then
- * we can get rid of it provided it's not also the tail bq
- * being filled by the issuer.
+ /* Clean up the head segment. If we clear an entire segment,
+ * then we can get rid of it provided it's not also the tail
+ * segment being filled by the issuer.
*/
bq->bv[slot].bv_page = NULL;
slot++;
while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&creq->buffer);
- if (!bq)
- goto done;
- slot = 0;
+ creq->collect_cursor.slot = slot;
+ if (!bvecq_delete_spent(&creq->collect_cursor))
+ goto out;
+ bq = creq->collect_cursor.bvecq;
+ slot = creq->collect_cursor.slot;
}
if (fpos + fsize >= collected_to)
break;
}
- creq->buffer.tail = bq;
-done:
- creq->buffer.first_tail_slot = slot;
+ creq->collect_cursor.slot = slot;
+out:
return made_progress;
}
diff --git a/fs/netfs/read_retry.c b/fs/netfs/read_retry.c
index 142c3fb8dab17..7490b9ee1bf7c 100644
--- a/fs/netfs/read_retry.c
+++ b/fs/netfs/read_retry.c
@@ -12,6 +12,10 @@
static void netfs_reissue_read(struct netfs_io_request *rreq,
struct netfs_io_subrequest *subreq)
{
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
+ iov_iter_advance(&subreq->io_iter, subreq->transferred);
+
subreq->error = 0;
__clear_bit(NETFS_SREQ_MADE_PROGRESS, &subreq->flags);
__set_bit(NETFS_SREQ_IN_PROGRESS, &subreq->flags);
@@ -27,6 +31,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
{
struct netfs_io_subrequest *subreq;
struct netfs_io_stream *stream = &rreq->io_streams[0];
+ struct bvecq_pos dispatch_cursor = {};
struct list_head *next;
_enter("R=%x", rreq->debug_id);
@@ -46,9 +51,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
if (test_bit(NETFS_SREQ_FAILED, &subreq->flags))
break;
if (__test_and_clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) {
- __clear_bit(NETFS_SREQ_MADE_PROGRESS, &subreq->flags);
subreq->retry_count++;
- netfs_reset_iter(subreq);
netfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);
netfs_reissue_read(rreq, subreq);
}
@@ -74,11 +77,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
do {
struct netfs_io_subrequest *from, *to, *tmp;
- struct iov_iter source;
uoff_t start, len;
size_t part;
bool boundary = false, subreq_superfluous = false;
+ bvecq_pos_unset(&dispatch_cursor);
+
/* Go through the subreqs and find the next span of contiguous
* buffer that we then rejig (cifs, for example, needs the
* rsize renegotiating) and reissue.
@@ -105,7 +109,8 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
break;
subreq = list_entry(next, struct netfs_io_subrequest, rreq_link);
- if (subreq->start + subreq->transferred != start + len ||
+ if (subreq->start != start + len ||
+ subreq->transferred > 0 ||
test_bit(NETFS_SREQ_BOUNDARY, &subreq->flags) ||
!test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags))
break;
@@ -118,11 +123,14 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
/* Determine the set of buffers we're going to use. Each
* subreq gets a subset of a single overall contiguous buffer.
*/
- netfs_reset_iter(from);
- source = from->io_iter;
- source.count = len;
+ bvecq_pos_transfer(&dispatch_cursor, &from->io_buffer);
+ bvecq_pos_advance(&dispatch_cursor, from->transferred);
+ from->transferred = 0;
- /* Work through the sublist. */
+ /* Work through the sublist. The chain of buffers we're going
+ * to fill is attached to dispatch_cursor and we need to read
+ * 'len' amount of data from 'start'.
+ */
subreq = from;
list_for_each_entry_from(subreq, &stream->subrequests, rreq_link) {
if (!len) {
@@ -130,16 +138,21 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
break;
}
subreq->source = NETFS_DOWNLOAD_FROM_SERVER;
- subreq->start = start - subreq->transferred;
- subreq->len = len + subreq->transferred;
+ subreq->start = start;
+ subreq->len = len;
__clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
__clear_bit(NETFS_SREQ_MADE_PROGRESS, &subreq->flags);
subreq->retry_count++;
+ subreq->transferred = 0;
+
+ bvecq_pos_unset(&subreq->io_buffer);
+ bvecq_pos_set(&subreq->io_buffer, &dispatch_cursor);
trace_netfs_sreq(subreq, netfs_sreq_trace_retry);
/* Renegotiate max_len (rsize) */
- stream->sreq_max_len = subreq->len;
+ stream->sreq_max_len = len;
+ stream->sreq_max_segs = INT_MAX;
if (rreq->netfs_ops->prepare_read &&
rreq->netfs_ops->prepare_read(subreq) < 0) {
trace_netfs_sreq(subreq, netfs_sreq_trace_reprep_failed);
@@ -147,13 +160,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
goto abandon;
}
- part = umin(len, stream->sreq_max_len);
- if (unlikely(stream->sreq_max_segs))
- part = netfs_limit_iter(&source, 0, part, stream->sreq_max_segs);
- subreq->len = subreq->transferred + part;
- subreq->io_iter = source;
- iov_iter_truncate(&subreq->io_iter, part);
- iov_iter_advance(&source, part);
+ part = bvecq_slice(&dispatch_cursor,
+ umin(len, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
+ subreq->len = part;
+
len -= part;
start += part;
if (!len) {
@@ -216,9 +228,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
trace_netfs_sreq(subreq, netfs_sreq_trace_retry);
stream->sreq_max_len = umin(len, rreq->rsize);
- stream->sreq_max_segs = 0;
- if (unlikely(stream->sreq_max_segs))
- part = netfs_limit_iter(&source, 0, part, stream->sreq_max_segs);
+ stream->sreq_max_segs = INT_MAX;
netfs_stat(&netfs_n_rh_download);
if (rreq->netfs_ops->prepare_read(subreq) < 0) {
@@ -227,11 +237,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
goto abandon;
}
- part = umin(len, stream->sreq_max_len);
- subreq->len = subreq->transferred + part;
- subreq->io_iter = source;
- iov_iter_truncate(&subreq->io_iter, part);
- iov_iter_advance(&source, part);
+ bvecq_pos_set(&subreq->io_buffer, &dispatch_cursor);
+ part = bvecq_slice(&dispatch_cursor,
+ umin(len, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
+ subreq->len = part;
len -= part;
start += part;
@@ -245,12 +256,14 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
} while (!list_is_head(next, &stream->subrequests));
+out:
+ bvecq_pos_unset(&dispatch_cursor);
return;
/* If we hit an error, fail all remaining incomplete subrequests */
abandon_after:
if (list_is_last(&subreq->rreq_link, &stream->subrequests))
- return;
+ goto out;
subreq = list_next_entry(subreq, rreq_link);
abandon:
list_for_each_entry_from(subreq, &stream->subrequests, rreq_link) {
@@ -261,6 +274,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
__set_bit(NETFS_SREQ_FAILED, &subreq->flags);
__clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
}
+ goto out;
}
/*
@@ -300,26 +314,24 @@ void netfs_unlock_abandoned_read_pages(struct netfs_io_request *rreq)
if (test_bit(NETFS_RREQ_NEED_PUT_RA_REFS, &rreq->flags))
netfs_wait_for_put_ra_refs(rreq);
- for (p = rreq->buffer.tail; p; p = p->next) {
- for (int slot = rreq->buffer.first_tail_slot;
- bvecq_acquire_slot(p, slot);
- slot++) {
- struct folio *folio;
+ for (p = rreq->collect_cursor.bvecq; p; p = bvecq_next(p)) {
+ unsigned int nr_slots = bvecq_nr_slots_acquire(p);
+ for (int slot = 0; slot < nr_slots; slot++) {
if (!p->bv[slot].bv_page)
continue;
- folio = bvec_folio(&p->bv[slot]);
+ struct folio *folio = bvec_folio(&p->bv[slot]);
+
netfs_cancel_copy_to_cache(rreq, folio);
if (folio == rreq->no_unlock_folio &&
test_bit(NETFS_RREQ_NO_UNLOCK_FOLIO, &rreq->flags)) {
_debug("no unlock");
- } else {
- trace_netfs_folio(folio, netfs_folio_trace_abandon);
- folio_unlock(folio);
+ continue;
}
+ trace_netfs_folio(folio, netfs_folio_trace_abandon);
+ folio_unlock(folio);
}
- rreq->buffer.first_tail_slot = 0;
}
}
diff --git a/fs/netfs/read_single.c b/fs/netfs/read_single.c
index b248e34bd0c86..c70941121de04 100644
--- a/fs/netfs/read_single.c
+++ b/fs/netfs/read_single.c
@@ -101,7 +101,11 @@ static int netfs_single_dispatch_read(struct netfs_io_request *rreq)
subreq->start = 0;
subreq->len = rreq->len;
- subreq->io_iter = rreq->buffer.iter;
+
+ bvecq_pos_set(&subreq->io_buffer, &rreq->dispatch_cursor);
+
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
netfs_queue_read(rreq, subreq);
@@ -174,6 +178,15 @@ ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_ite
if (IS_ERR(rreq))
return PTR_ERR(rreq);
+ ret = netfs_extract_iter(iter, rreq->len, INT_MAX, &rreq->dispatch_cursor.bvecq,
+ 0, rreq->gfp);
+ if (ret < 0)
+ goto cleanup_free;
+ if (ret < rreq->len) {
+ ret = -EIO;
+ goto cleanup_free;
+ }
+
rreq->progress_at = rreq->len;
ret = netfs_single_begin_cache_read(rreq, ictx);
@@ -183,7 +196,6 @@ ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_ite
netfs_stat(&netfs_n_rh_read_single);
trace_netfs_read(rreq, 0, rreq->len, netfs_read_trace_read_single);
- rreq->buffer.iter = *iter;
netfs_single_dispatch_read(rreq);
ret = netfs_wait_for_read(rreq);
diff --git a/fs/netfs/rolling_buffer.c b/fs/netfs/rolling_buffer.c
deleted file mode 100644
index 66ce9add40122..0000000000000
--- a/fs/netfs/rolling_buffer.c
+++ /dev/null
@@ -1,182 +0,0 @@
-// SPDX-License-Identifier: GPL-2.0-or-later
-/* Rolling buffer helpers
- *
- * Copyright (C) 2024 Red Hat, Inc. All Rights Reserved.
- * Written by David Howells (dhowells@redhat.com)
- */
-
-#include <linux/bitops.h>
-#include <linux/mempool.h>
-#include <linux/pagemap.h>
-#include <linux/rolling_buffer.h>
-#include <linux/slab.h>
-#include "internal.h"
-
-/*
- * Initialise a rolling buffer. We allocate an empty folio queue struct to so
- * that the pointers can be independently driven by the producer and the
- * consumer.
- */
-int rolling_buffer_init(struct rolling_buffer *roll, unsigned int direction,
- gfp_t gfp, bool for_writeback)
-{
- struct bvecq *bq;
-
- roll->for_writeback = for_writeback;
-
- bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, for_writeback);
- if (!bq)
- return -ENOMEM;
-
- roll->head = bq;
- roll->tail = bq;
- iov_iter_bvec_queue(&roll->iter, direction, bq, 0, 0, 0);
- return 0;
-}
-
-/*
- * Add another bvecq to a rolling buffer if there's no space left.
- */
-int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp)
-{
- struct bvecq *bq, *head = roll->head;
-
- if (!bvecq_is_full(head))
- return 0;
-
- bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, roll->for_writeback);
- if (!bq)
- return -ENOMEM;
-
- roll->head = bq;
- if (bvecq_is_full(head)) {
- /* Make sure we don't leave the master iterator pointing to a
- * block that might get immediately consumed.
- */
- if (roll->iter.bvecq == head &&
- roll->iter.bvecq_slot == head->nr_slots) {
- roll->iter.bvecq = bq;
- roll->iter.bvecq_slot = 0;
- }
- }
-
- /* Make sure the initialisation is stored before the next pointer.
- *
- * [!] NOTE: After we set head->next, the consumer is at liberty to
- * immediately delete the old head.
- */
- bvecq_append(head, bq);
- return 0;
-}
-
-/*
- * Decant the entire list of folios to read into a rolling buffer.
- */
-ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,
- struct readahead_control *ractl,
- gfp_t gfp)
-{
- struct bvecq *bq;
- size_t loaded = 0;
-
- while (ractl->_nr_pages - ractl->_batch_count > 0) {
- struct page **pages;
- unsigned int nr;
-
- /* Allocate a bvecq to put some folios into and attach it to
- * the rolling buffer.
- */
- bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, false);
- if (!bq)
- goto nomem_unlock;
- bq->mem_type = BVECQ_MEM_EXTERNAL; /* Folio cleanup handled separately. */
-
- if (!roll->tail)
- roll->tail = bq;
- else
- bvecq_append(roll->head, bq);
- roll->head = bq;
-
- /* Get a bunch of folios and note their sizes. */
- pages = (struct page **)(bq->bv + bq->max_slots);
- pages -= bq->max_slots;
- nr = __readahead_batch(ractl, pages, bq->max_slots);
- if (WARN_ON_ONCE(!nr))
- break;
-
- for (int slot = 0; slot < nr; slot++) {
- struct folio *folio = page_folio(pages[slot]);
- size_t len = folio_size(folio);
-
- bvec_set_folio(&bq->bv[slot], folio, len, 0);
- loaded += len;
- trace_netfs_folio(folio, netfs_folio_trace_read);
- }
-
- bvecq_filled_to(bq, nr);
- }
-
- WRITE_ONCE(roll->iter.count, loaded);
- iov_iter_bvec_queue(&roll->iter, ITER_DEST, roll->tail, 0, 0, loaded);
- return loaded;
-
-nomem_unlock:
- for (bq = roll->tail; bq; bq = bq->next) {
- for (int slot = 0; slot < bq->nr_slots; slot++) {
- struct folio *folio = bvec_folio(&bq->bv[slot]);
-
- folio_unlock(folio);
- folio_put(folio);
- }
- }
- rolling_buffer_clear(roll);
- roll->head = NULL;
- roll->tail = NULL;
- return -ENOMEM;
-}
-
-/*
- * Append a folio to the rolling buffer.
- */
-ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio,
- gfp_t gfp)
-{
- ssize_t size = folio_size(folio);
- int slot;
-
- if (rolling_buffer_make_space(roll, gfp) < 0)
- return -ENOMEM;
-
- slot = roll->head->nr_slots;
- bvec_set_folio(&roll->head->bv[slot], folio, size, 0);
- bvecq_filled_to(roll->head, slot + 1);
-
- WRITE_ONCE(roll->iter.count, roll->iter.count + size);
- return size;
-}
-
-/*
- * Delete a spent buffer from a rolling queue and return the next in line. We
- * don't return the last buffer to keep the pointers independent, but return
- * NULL instead.
- */
-struct bvecq *rolling_buffer_delete_spent(struct rolling_buffer *roll)
-{
- struct bvecq *spent = roll->tail, *next = bvecq_next(spent);
-
- if (!next)
- return NULL;
- next->prev = NULL;
- roll->tail = next;
- spent->next = NULL;
- bvecq_put(spent);
- return next;
-}
-
-/*
- * Clear out a rolling queue.
- */
-void rolling_buffer_clear(struct rolling_buffer *roll)
-{
- bvecq_put(roll->tail);
-}
diff --git a/fs/netfs/write_collect.c b/fs/netfs/write_collect.c
index 91b42820c8925..dcacbab254b92 100644
--- a/fs/netfs/write_collect.c
+++ b/fs/netfs/write_collect.c
@@ -114,12 +114,12 @@ int netfs_folio_written_back(struct folio *folio)
static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
unsigned int *notes)
{
- struct bvecq *bq = wreq->buffer.tail;
- unsigned int slot = wreq->buffer.first_tail_slot;
+ struct bvecq *bq = wreq->collect_cursor.bvecq;
+ unsigned int slot = wreq->collect_cursor.slot;
uoff_t collected_to = wreq->collected_to;
if (WARN_ON_ONCE(!bq)) {
- pr_err("[!] Writeback unlock found empty rolling buffer!\n");
+ pr_err("[!] Writeback unlock found empty buffer!\n");
netfs_dump_request(wreq);
return;
}
@@ -130,19 +130,28 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
return;
}
- while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&wreq->buffer);
- if (!bq)
- return;
- slot = 0;
- }
-
for (;;) {
struct folio *folio;
struct netfs_folio *finfo;
uoff_t fpos, fend;
size_t fsize, flen;
+ /* Try to clean up the head of the queue if it appears to be
+ * used up, but we need to be very careful - the cleanup can
+ * catch the dispatcher, which could lead to us having nothing
+ * left in the queue, causing the front and back pointers to
+ * end up on different tracks. To avoid this, we must always
+ * keep at least one segment in the queue.
+ */
+ if (!bvecq_acquire_slot(bq, slot)) {
+ wreq->collect_cursor.slot = slot;
+ if (!bvecq_delete_spent(&wreq->collect_cursor))
+ return;
+ bq = wreq->collect_cursor.bvecq;
+ slot = wreq->collect_cursor.slot;
+ continue;
+ }
+
folio = bvec_folio(&bq->bv[slot]);
if (WARN_ONCE(!folio_test_writeback(folio),
"R=%08x: folio %lx is not under writeback\n",
@@ -166,26 +175,13 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
wreq->cleaned_to = fpos + fsize;
*notes |= MADE_PROGRESS;
- /* Clean up the head bq. If we clear an entire bq, then
- * we can get rid of it provided it's not also the tail bq
- * being filled by the issuer.
- */
bq->bv[slot].bv_page = NULL;
slot++;
- while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&wreq->buffer);
- if (!bq)
- goto done;
- slot = 0;
- }
-
if (fpos + fsize >= collected_to)
break;
}
- wreq->buffer.tail = bq;
-done:
- wreq->buffer.first_tail_slot = slot;
+ wreq->collect_cursor.slot = slot;
}
/*
@@ -230,7 +226,8 @@ static void netfs_collect_write_results(struct netfs_io_request *wreq)
trace_netfs_rreq(wreq, netfs_rreq_trace_collect);
reassess_streams:
- issued_to = atomic64_read(&wreq->issued_to);
+ /* Order reading the issued_to point before reading the queue it refers to. */
+ issued_to = atomic64_read_acquire(&wreq->issued_to);
smp_rmb();
collected_to = ULLONG_MAX;
if (wreq->origin == NETFS_WRITEBACK ||
@@ -560,8 +557,12 @@ void netfs_write_subrequest_terminated(void *_op, ssize_t transferred_or_error)
* data is tracked.
*/
netfs_stat(&netfs_n_wh_write_failed);
- if (test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags))
- break;
+ if (test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) {
+ /* We don't retry failed cache writes. */
+ __clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
+ if (!subreq->error)
+ subreq->error = -ENOBUFS;
+ }
trace_netfs_failure(wreq, subreq, transferred_or_error, netfs_fail_write);
__set_bit(NETFS_SREQ_CANCELLED, &subreq->flags);
diff --git a/fs/netfs/write_issue.c b/fs/netfs/write_issue.c
index f0f4786666512..8fab5cf00d7b6 100644
--- a/fs/netfs/write_issue.c
+++ b/fs/netfs/write_issue.c
@@ -107,10 +107,6 @@ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,
ictx = netfs_inode(wreq->inode);
if (is_cacheable)
fscache_begin_write_operation(&wreq->cache_resources, netfs_i_cookie(ictx));
- if (rolling_buffer_init(&wreq->buffer, ITER_SOURCE, wreq->gfp,
- (origin == NETFS_WRITEBACK ||
- origin == NETFS_WRITEBACK_SINGLE)) < 0)
- goto nomem;
wreq->cleaned_to = wreq->start;
if (wreq->cache_resources.dio_size > 1)
@@ -135,9 +131,6 @@ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,
}
return wreq;
-nomem:
- netfs_put_failed_request(wreq);
- return ERR_PTR(-ENOMEM);
}
/**
@@ -163,22 +156,14 @@ void netfs_prepare_write(struct netfs_io_request *wreq,
uoff_t start)
{
struct netfs_io_subrequest *subreq;
- struct iov_iter *wreq_iter = &wreq->buffer.iter;
-
- /* Make sure we don't point the iterator at a used-up bvecq struct
- * being used as a placeholder to prevent the queue from collapsing.
- * In such a case, extend the queue.
- */
- if (iov_iter_is_bvecq(wreq_iter) &&
- !bvecq_acquire_slot(wreq_iter->bvecq, wreq_iter->bvecq_slot))
- rolling_buffer_make_space(&wreq->buffer, wreq->gfp);
subreq = netfs_alloc_subrequest(wreq, stream->source);
if (!subreq)
return;
subreq->start = start;
subreq->stream_nr = stream->stream_nr;
- subreq->io_iter = *wreq_iter;
+
+ bvecq_pos_set(&subreq->io_buffer, &wreq->dispatch_cursor);
_enter("R=%x[%x]", wreq->debug_id, subreq->debug_index);
@@ -259,15 +244,14 @@ static void netfs_do_issue_write(struct netfs_io_stream *stream,
}
void netfs_reissue_write(struct netfs_io_stream *stream,
- struct netfs_io_subrequest *subreq,
- struct iov_iter *source)
+ struct netfs_io_subrequest *subreq)
{
- size_t size = subreq->len - subreq->transferred;
-
// TODO: Use encrypted buffer
- subreq->io_iter = *source;
- iov_iter_advance(source, size);
- iov_iter_truncate(&subreq->io_iter, size);
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_SOURCE,
+ subreq->io_buffer.bvecq, subreq->io_buffer.slot,
+ subreq->io_buffer.offset,
+ subreq->len);
+ iov_iter_advance(&subreq->io_iter, subreq->transferred);
subreq->retry_count++;
subreq->error = 0;
@@ -285,8 +269,12 @@ void netfs_issue_write(struct netfs_io_request *wreq,
if (!subreq)
return;
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_SOURCE,
+ subreq->io_buffer.bvecq, subreq->io_buffer.slot,
+ subreq->io_buffer.offset,
+ subreq->len);
+
stream->construct = NULL;
- subreq->io_iter.count = subreq->len;
netfs_do_issue_write(stream, subreq);
}
@@ -323,7 +311,6 @@ size_t netfs_advance_write(struct netfs_io_request *wreq,
_debug("part %zx/%zx %zx/%zx", subreq->len, stream->sreq_max_len, part, len);
subreq->len += part;
subreq->nr_segs++;
- stream->submit_extendable_to -= part;
if (subreq->len >= stream->sreq_max_len ||
subreq->nr_segs >= stream->sreq_max_segs ||
@@ -347,7 +334,8 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
struct netfs_io_stream *stream;
struct netfs_group *fgroup; /* TODO: Use this with ceph */
struct netfs_folio *finfo;
- size_t iter_off = 0;
+ struct bvecq *queue = wreq->load_cursor.bvecq;
+ unsigned int slot;
size_t fsize = folio_size(folio), flen = fsize, foff = 0;
uoff_t fpos = folio_pos(folio), i_size;
bool to_eof = false, streamw = false;
@@ -355,12 +343,20 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
_enter("");
- if (rolling_buffer_make_space(&wreq->buffer, wreq->gfp) < 0)
- return -ENOMEM;
+ if (!wreq->spare) {
+ wreq->spare = bvecq_alloc_one(BVECQ_POOL_SLOTS, wreq->gfp, true);
+ if (!wreq->spare)
+ return -ENOMEM;
+ }
- /* netfs_perform_write() may shift i_size around the page or from out
- * of the page to beyond it, but cannot move i_size into or through the
- * page since we have it locked.
+ /* netfs_perform_write() may shift i_size around the folio or from out
+ * of the folio to beyond it, but cannot move i_size into or through
+ * the folio since we have it locked.
+ *
+ * Truncate could in theory move i_size into or before the folio, but
+ * it should take steps to prevent writeback from happening
+ * concurrently and should wait for any in-progress writebacks before
+ * proceeding.
*/
i_size = i_size_read(wreq->inode);
@@ -452,8 +448,29 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
trace_netfs_folio(folio, netfs_folio_trace_store_plus);
}
+ /* Institute a new bvec queue segment if the current one is full or if
+ * we encounter a discontiguity. The discontiguity break is important
+ * when it comes to bulk unlocking folios by file range.
+ */
+ if (bvecq_is_full(queue) ||
+ (fpos != wreq->last_end && wreq->last_end > 0)) {
+ bvecq_buffer_append(&wreq->load_cursor, wreq->spare);
+ wreq->spare = NULL;
+
+ queue = wreq->load_cursor.bvecq;
+ bvecq_pos_move(&wreq->dispatch_cursor, queue);
+ wreq->dispatch_cursor.slot = 0;
+ }
+
/* Attach the folio to the rolling buffer. */
- rolling_buffer_append(&wreq->buffer, folio, wreq->gfp);
+ slot = queue->nr_slots;
+ bvec_set_folio(&queue->bv[slot], folio, fsize, 0);
+ trace_netfs_bv_slot(queue, slot);
+ slot++;
+ bvecq_filled_to(queue, slot);
+ wreq->load_cursor.slot = slot;
+ wreq->load_cursor.offset = 0;
+ wreq->last_end = fpos + fsize;
/* Move the submission point forward to allow for write-streaming data
* not starting at the front of the page. We don't do write-streaming
@@ -462,10 +479,19 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
* Also skip uploading for data that's been read and just needs copying
* to the cache.
*/
+ bvecq_pos_nudge(&wreq->dispatch_cursor);
+
for (int s = 0; s < NR_IO_STREAMS; s++) {
+ size_t soff = foff, slen = flen, alignment = 1;
+
stream = &wreq->io_streams[s];
- stream->submit_off = foff;
- stream->submit_len = flen;
+ if (stream->source == NETFS_WRITE_TO_CACHE)
+ alignment = wreq->cache_resources.dio_size;
+ stream = &wreq->io_streams[s];
+ stream->submit_off = round_down(soff, alignment);
+ slen += foff - stream->submit_off;
+ stream->submit_len = round_up(slen, alignment);
+
if (!stream->avail ||
(stream->source == NETFS_WRITE_TO_CACHE && streamw) ||
(stream->source == NETFS_UPLOAD_TO_SERVER &&
@@ -499,14 +525,10 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
break;
stream = &wreq->io_streams[choose_s];
- /* Advance the iterator(s). */
- if (stream->submit_off > iter_off) {
- rolling_buffer_advance(&wreq->buffer, stream->submit_off - iter_off);
- iter_off = stream->submit_off;
- }
+ /* Advance the cursor. */
+ wreq->dispatch_cursor.offset = stream->submit_off;
atomic64_set(&wreq->issued_to, fpos + stream->submit_off);
- stream->submit_extendable_to = fsize - stream->submit_off;
part = netfs_advance_write(wreq, stream, fpos + stream->submit_off,
stream->submit_len, to_eof);
stream->submit_off += part;
@@ -518,9 +540,9 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
debug = true;
}
- if (fsize > iter_off)
- rolling_buffer_advance(&wreq->buffer, fsize - iter_off);
- atomic64_set(&wreq->issued_to, fpos + fsize);
+ bvecq_pos_step(&wreq->dispatch_cursor);
+ /* Order loading the queue before updating the issue_to point */
+ atomic64_set_release(&wreq->issued_to, fpos + fsize);
if (!debug)
kdebug("R=%x: No submit", wreq->debug_id);
@@ -581,6 +603,11 @@ int netfs_writepages(struct address_space *mapping,
goto couldnt_start;
}
+ if (bvecq_buffer_init(&wreq->load_cursor, wreq->gfp, true) < 0)
+ goto nomem;
+ bvecq_pos_set(&wreq->dispatch_cursor, &wreq->load_cursor);
+ bvecq_pos_set(&wreq->collect_cursor, &wreq->dispatch_cursor);
+
__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, &wreq->flags);
trace_netfs_write(wreq, netfs_write_trace_writeback);
netfs_stat(&netfs_n_wh_writepages);
@@ -605,12 +632,17 @@ int netfs_writepages(struct address_space *mapping,
} while ((folio = writeback_iter(mapping, wbc, folio, &error)));
netfs_end_issue_write(wreq);
+ bvecq_pos_unset(&wreq->load_cursor);
+ bvecq_pos_unset(&wreq->dispatch_cursor);
netfs_wake_collector(wreq);
netfs_put_request(wreq, netfs_rreq_trace_put_return);
_leave(" = %d", error);
return error;
+nomem:
+ error = -ENOMEM;
+ netfs_put_failed_request(wreq);
couldnt_start:
if (error == -ENOMEM) {
folio_redirty_for_writepage(wbc, folio);
@@ -631,23 +663,28 @@ EXPORT_SYMBOL(netfs_writepages);
* netfs_writeback_single - Write back a monolithic payload
* @mapping: The mapping to write from
* @wbc: Hints from the VM
- * @iter: Data to write.
+ * @iter: Buffer to write from
+ * @len: Amount to write from buffer
*
* Write a monolithic, non-pagecache object back to the server and/or the
- * cache. The caller must explicitly set NETFS_RREQ_UPLOAD_TO_SERVER when
- * initialising the request if it wants the data to be written to the server
- * (for AFS directories and symlinks, this is not possible; things like mkdir,
- * symlink, rmdir and unlink must be used instead).
+ * cache. There's a maximum of one subrequest per stream. The buffer should be
+ * rounded out sufficiently that it can accommodate cache DIO rounding.
+ *
+ * The caller must explicitly set NETFS_RREQ_UPLOAD_TO_SERVER when initialising
+ * the request if it wants the data to be written to the server (for AFS
+ * directories and symlinks, this is not possible; things like mkdir, symlink,
+ * rmdir and unlink must be used instead).
*
* Return: 0 if successful; 1 if skipped due to lock conflict and WB_SYNC_NONE;
* or a negative error code.
*/
int netfs_writeback_single(struct address_space *mapping,
struct writeback_control *wbc,
- struct iov_iter *iter)
+ struct iov_iter *iter, size_t len)
{
struct netfs_io_request *wreq;
struct netfs_inode *ictx = netfs_inode(mapping->host);
+ size_t clen;
int ret;
if (!netfs_wb_begin(ictx, wbc->sync_mode == WB_SYNC_NONE)) {
@@ -661,10 +698,27 @@ int netfs_writeback_single(struct address_space *mapping,
ret = PTR_ERR(wreq);
goto couldnt_start;
}
+ wreq->len = len;
+ clen = len;
+
+ if (wreq->cache_resources.dio_size > 1) {
+ clen = round_up(len, wreq->cache_resources.dio_size);
+ if (clen > iov_iter_count(iter)) {
+ ret = -EIO;
+ goto cleanup_free;
+ }
+ }
- wreq->buffer.iter = *iter;
- wreq->len = iov_iter_count(iter);
- wreq->submitted = wreq->len;
+ ret = netfs_extract_iter(iter, clen, INT_MAX, &wreq->dispatch_cursor.bvecq,
+ 0, wreq->gfp);
+ if (ret < 0)
+ goto cleanup_free;
+ if (ret < clen) {
+ ret = -EIO;
+ goto cleanup_free;
+ }
+
+ bvecq_pos_set(&wreq->collect_cursor, &wreq->dispatch_cursor);
__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, &wreq->flags);
trace_netfs_write(wreq, netfs_write_trace_writeback_single);
@@ -685,12 +739,14 @@ int netfs_writeback_single(struct address_space *mapping,
subreq = stream->construct;
subreq->len = wreq->len;
+ if (stream->source == NETFS_WRITE_TO_CACHE)
+ subreq->len = clen;
stream->submit_len = subreq->len;
- stream->submit_extendable_to = round_up(wreq->len, PAGE_SIZE);
netfs_issue_write(wreq, stream);
}
+ wreq->submitted = wreq->len;
netfs_all_subreqs_queued(wreq);
netfs_wake_collector(wreq);
@@ -705,6 +761,8 @@ int netfs_writeback_single(struct address_space *mapping,
_leave(" = %d", ret);
return ret;
+cleanup_free:
+ netfs_put_failed_request(wreq);
couldnt_start:
netfs_wb_end(ictx);
_leave(" = %d", ret);
diff --git a/fs/netfs/write_retry.c b/fs/netfs/write_retry.c
index 2f20577563e14..235e75eb10914 100644
--- a/fs/netfs/write_retry.c
+++ b/fs/netfs/write_retry.c
@@ -17,15 +17,17 @@
static void netfs_retry_write_stream(struct netfs_io_request *wreq,
struct netfs_io_stream *stream)
{
+ struct bvecq_pos dispatch_cursor = {};
struct list_head *next;
_enter("R=%x[%x:]", wreq->debug_id, stream->stream_nr);
if (list_empty(&stream->subrequests))
return;
+ if (WARN_ON_ONCE(stream->source != NETFS_UPLOAD_TO_SERVER))
+ return; /* Shouldn't be retrying cache writes. */
- if (stream->source == NETFS_UPLOAD_TO_SERVER &&
- wreq->netfs_ops->retry_request)
+ if (wreq->netfs_ops->retry_request)
wreq->netfs_ops->retry_request(wreq, stream);
if (unlikely(stream->failed))
@@ -39,12 +41,8 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
if (test_bit(NETFS_SREQ_FAILED, &subreq->flags))
break;
if (__test_and_clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) {
- struct iov_iter source;
-
- netfs_reset_iter(subreq);
- source = subreq->io_iter;
netfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);
- netfs_reissue_write(stream, subreq, &source);
+ netfs_reissue_write(stream, subreq);
}
}
return;
@@ -54,11 +52,12 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
do {
struct netfs_io_subrequest *subreq = NULL, *from, *to, *tmp;
- struct iov_iter source;
uoff_t start, len;
size_t part;
bool boundary = false;
+ bvecq_pos_unset(&dispatch_cursor);
+
/* Go through the stream and find the next span of contiguous
* data that we then rejig (cifs, for example, needs the wsize
* renegotiating) and reissue.
@@ -70,7 +69,7 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
if (test_bit(NETFS_SREQ_FAILED, &from->flags) ||
!test_bit(NETFS_SREQ_NEED_RETRY, &from->flags))
- return;
+ goto out;
for (;;) {
/* Read pointer to subreq before reading subreq state. */
@@ -79,7 +78,8 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
break;
subreq = list_entry(next, struct netfs_io_subrequest, rreq_link);
- if (subreq->start + subreq->transferred != start + len ||
+ if (subreq->start != start + len ||
+ subreq->transferred > 0 ||
test_bit(NETFS_SREQ_BOUNDARY, &subreq->flags) ||
!test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags))
break;
@@ -90,11 +90,13 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
/* Determine the set of buffers we're going to use. Each
* subreq gets a subset of a single overall contiguous buffer.
*/
- netfs_reset_iter(from);
- source = from->io_iter;
- source.count = len;
+ bvecq_pos_transfer(&dispatch_cursor, &from->io_buffer);
+ bvecq_pos_advance(&dispatch_cursor, from->transferred);
- /* Work through the sublist. */
+ /* Work through the sublist. The chain of buffers we're going
+ * to fill is attached to dispatch_cursor and we need to read
+ * 'len' amount of data from 'start'.
+ */
subreq = from;
list_for_each_entry_from(subreq, &stream->subrequests, rreq_link) {
if (!len)
@@ -104,16 +106,22 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
subreq->len = len;
__clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
trace_netfs_sreq(subreq, netfs_sreq_trace_retry);
+ subreq->transferred = 0;
+
+ bvecq_pos_unset(&subreq->io_buffer);
/* Renegotiate max_len (wsize) */
stream->sreq_max_len = len;
+ stream->sreq_max_segs = INT_MAX;
stream->prepare_write(subreq);
- part = umin(len, stream->sreq_max_len);
- if (unlikely(stream->sreq_max_segs))
- part = netfs_limit_iter(&source, 0, part, stream->sreq_max_segs);
+ bvecq_pos_set(&subreq->io_buffer, &dispatch_cursor);
+ part = bvecq_slice(&dispatch_cursor,
+ umin(len, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
subreq->len = part;
- subreq->transferred = 0;
+
len -= part;
start += part;
if (len && subreq == to &&
@@ -121,7 +129,7 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
boundary = true;
netfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);
- netfs_reissue_write(stream, subreq, &source);
+ netfs_reissue_write(stream, subreq);
if (subreq == to)
break;
}
@@ -172,17 +180,19 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
netfs_stat(&netfs_n_wh_upload);
stream->sreq_max_len = umin(len, wreq->wsize);
break;
- case NETFS_WRITE_TO_CACHE:
- netfs_stat(&netfs_n_wh_write);
- break;
default:
WARN_ON_ONCE(1);
}
stream->prepare_write(subreq);
- part = umin(len, stream->sreq_max_len);
+ bvecq_pos_set(&subreq->io_buffer, &dispatch_cursor);
+ part = bvecq_slice(&dispatch_cursor,
+ umin(len, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
subreq->len = subreq->transferred + part;
+
len -= part;
start += part;
if (!len && boundary) {
@@ -190,13 +200,16 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
boundary = false;
}
- netfs_reissue_write(stream, subreq, &source);
+ netfs_reissue_write(stream, subreq);
if (!len)
break;
} while (len);
} while (!list_is_head(next, &stream->subrequests));
+
+out:
+ bvecq_pos_unset(&dispatch_cursor);
}
/*
diff --git a/include/linux/bvecq.h b/include/linux/bvecq.h
index b984aaa449088..389f6407cb846 100644
--- a/include/linux/bvecq.h
+++ b/include/linux/bvecq.h
@@ -54,6 +54,16 @@ struct bvecq {
/* Number of slots in a 4K bvecq. */
#define BVECQ_4KB_SLOTS ((4096 - sizeof(struct bvecq)) / sizeof(struct bio_vec))
+/*
+ * Position in a bio_vec queue. The bvecq holds a ref on the queue segment it
+ * points to.
+ */
+struct bvecq_pos {
+ struct bvecq *bvecq; /* The first bvecq */
+ unsigned int offset; /* The offset within the starting slot */
+ u16 slot; /* The starting slot */
+};
+
void bvecq_dump(const struct bvecq *bq);
struct bvecq *bvecq_alloc_one(size_t nr_slots, gfp_t gfp, bool for_writeback);
struct bvecq *bvecq_alloc_chain(size_t nr_slots, gfp_t gfp, bool for_writeback);
@@ -61,6 +71,13 @@ struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp
bool for_writeback);
void bvecq_put(struct bvecq *bq);
int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size, gfp_t gfp);
+int bvecq_buffer_init(struct bvecq_pos *pos, gfp_t gfp, bool for_writeback);
+void bvecq_buffer_append(struct bvecq_pos *pos, struct bvecq *bq);
+void bvecq_pos_advance(struct bvecq_pos *pos, size_t amount);
+ssize_t bvecq_zero(struct bvecq_pos *pos, size_t amount);
+size_t bvecq_slice(struct bvecq_pos *pos, size_t max_size,
+ unsigned int max_slots, unsigned int *_nr_slots);
+ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl);
/**
* bvecq_alloc_buffer - Allocate a bvecq chain and populate with buffers
@@ -163,4 +180,171 @@ static inline struct bvecq *bvecq_next(const struct bvecq *bq)
return smp_load_acquire(&bq->next);
}
+/**
+ * bvecq_pos_set - Set one position to be the same as another
+ * @pos: The position object to set
+ * @at: The source position.
+ *
+ * Set @pos to have the same position as @at. This may take a ref on the
+ * bvecq pointed to.
+ */
+static inline void bvecq_pos_set(struct bvecq_pos *pos, const struct bvecq_pos *at)
+{
+ *pos = *at;
+ bvecq_get(pos->bvecq);
+}
+
+/**
+ * bvecq_pos_unset - Unset a position
+ * @pos: The position object to unset
+ *
+ * Unset @pos. This does any needed ref cleanup.
+ */
+static inline void bvecq_pos_unset(struct bvecq_pos *pos)
+{
+ bvecq_put(pos->bvecq);
+ pos->bvecq = NULL;
+ pos->slot = 0;
+ pos->offset = 0;
+}
+
+/**
+ * bvecq_pos_transfer - Transfer one position to another, clearing the first
+ * @pos: The position object to set
+ * @from: The source position to clear.
+ *
+ * Set @pos to have the same position as @from and then clear @from. This may
+ * transfer a ref on the bvecq pointed to.
+ */
+static inline void bvecq_pos_transfer(struct bvecq_pos *pos, struct bvecq_pos *from)
+{
+ *pos = *from;
+ from->bvecq = NULL;
+ from->slot = 0;
+ from->offset = 0;
+}
+
+/**
+ * bvecq_pos_move - Update a position to a new bvecq
+ * @pos: The position object to update.
+ * @to: The new bvecq to point at.
+ *
+ * Update @pos to point to @to if it doesn't already do so. This may
+ * manipulate refs on the bvecqs pointed to.
+ */
+static inline void bvecq_pos_move(struct bvecq_pos *pos, struct bvecq *to)
+{
+ struct bvecq *old = pos->bvecq;
+
+ if (old != to) {
+ pos->bvecq = bvecq_get(to);
+ bvecq_put(old);
+ }
+}
+
+/**
+ * bvecq_pos_nudge - Nudge a position onto the next segment if current used up
+ * @pos: The position object to nudge.
+ *
+ * Update @pos to point to the next segment in the chain if we've used up the
+ * current segment. This may manipulate refs on the bvecqs pointed to.
+ *
+ * Return: true if found a new segment, false if hit the end.
+ */
+static inline bool bvecq_pos_nudge(struct bvecq_pos *pos)
+{
+ struct bvecq *bq = pos->bvecq;
+
+ for (;;) {
+ if (!bvecq_acquire_slot(bq, pos->slot)) {
+ bq = bvecq_next(bq);
+ if (!bq)
+ return false;
+ if (bvecq_acquire_slot(bq, pos->slot))
+ continue; /* More slots got added. */
+ bvecq_pos_move(pos, bq);
+ pos->slot = 0;
+ pos->offset = 0;
+ continue;
+ }
+ if (pos->offset >= bq->bv[pos->slot].bv_len) {
+ pos->slot++;
+ pos->offset = 0;
+ continue;
+ }
+ return true;
+ }
+}
+
+/**
+ * bvecq_pos_step - Step a position to the next slot if possible
+ * @pos: The position object to step.
+ *
+ * Update @pos to point to the next slot in the queue if not at the end. This
+ * may manipulate refs on the bvecqs pointed to.
+ *
+ * Return: true if successful, false if was at the end.
+ */
+static inline bool bvecq_pos_step(struct bvecq_pos *pos)
+{
+ struct bvecq *bq = pos->bvecq, *next;
+
+ pos->slot++;
+ pos->offset = 0;
+ if (bvecq_acquire_slot(bq, pos->slot))
+ return true;
+ next = bvecq_next(bq);
+ if (!next)
+ return false;
+ if (bvecq_acquire_slot(bq, pos->slot))
+ return true;
+ bvecq_pos_move(pos, next);
+ pos->slot = 0;
+ return true;
+}
+
+/**
+ * bvecq_delete_spent - Delete the bvecq at the front if possible
+ * @pos: The position object to update.
+ *
+ * Delete the used up bvecq at the front of the queue that @pos points to if it
+ * is not the last node in the queue; if it is the last node in the queue, it
+ * is kept so that the queue doesn't become detached from the other end. This
+ * may manipulate refs on the bvecqs pointed to. It is also possible that the
+ * producer will fill more slots in the current bvecq.
+ *
+ * Also, we have to be very careful: the consumer can catch the producer, which
+ * could lead to us having nothing left in the queue, causing the front and
+ * back pointers to end up on different tracks. To avoid this, we must always
+ * keep at least one segment in the queue.
+ *
+ * The caller must reload from @pos after calling this.
+ *
+ * Return: true if there's more available; false if not.
+ */
+static inline bool bvecq_delete_spent(struct bvecq_pos *pos)
+{
+ struct bvecq *spent = pos->bvecq;
+ struct bvecq *next;
+ unsigned int slot = pos->slot;
+
+again:
+ /* Read the contents of the queue node after the pointer to it. */
+ next = bvecq_next(spent);
+ if (!next)
+ return false; /* Nothing more to consume at the moment. */
+ if (slot < bvecq_nr_slots_acquire(spent))
+ return true; /* The producer added more. */
+ next->prev = NULL;
+ bvecq_pos_move(pos, next);
+ pos->slot = 0;
+ pos->offset = 0;
+ if (!bvecq_acquire_slot(next, 0)) {
+ spent = next;
+ slot = 0;
+ goto again;
+ }
+ return true;
+}
+
#endif /* _LINUX_BVECQ_H */
diff --git a/include/linux/netfs.h b/include/linux/netfs.h
index b0dd92d12a971..0340c9ee587b6 100644
--- a/include/linux/netfs.h
+++ b/include/linux/netfs.h
@@ -19,10 +19,12 @@
#include <linux/pagemap.h>
#include <linux/bvecq.h>
#include <linux/uio.h>
-#include <linux/rolling_buffer.h>
enum netfs_sreq_ref_trace;
typedef struct mempool mempool_t;
+struct readahead_control;
+struct netfs_io_request;
+struct netfs_io_subrequest;
struct fscache_occupancy;
/**
@@ -144,7 +146,6 @@ struct netfs_io_stream {
unsigned int sreq_max_segs; /* 0 or max number of segments in an iterator */
unsigned int submit_off; /* Folio offset we're submitting from */
unsigned int submit_len; /* Amount of data left to submit */
- unsigned int submit_extendable_to; /* Amount I/O can be rounded up to */
void (*prepare_write)(struct netfs_io_subrequest *subreq);
void (*issue_write)(struct netfs_io_subrequest *subreq);
/* Collection tracking */
@@ -187,6 +188,7 @@ struct netfs_io_subrequest {
struct netfs_io_request *rreq; /* Supervising I/O request */
struct work_struct work;
struct list_head rreq_link; /* Link in rreq->subrequests */
+ struct bvecq_pos io_buffer; /* Bookmark in the combined queue of the start */
struct iov_iter io_iter; /* Iterator for this subrequest */
uoff_t start; /* Where to start the I/O */
size_t len; /* Size of the I/O */
@@ -247,11 +249,14 @@ struct netfs_io_request {
struct netfs_io_stream io_streams[2]; /* Streams of parallel I/O operations */
#define NR_IO_STREAMS 2 //wreq->nr_io_streams
struct netfs_group *group; /* Writeback group being written back */
- struct rolling_buffer buffer; /* Unencrypted buffer */
+ struct bvecq *spare; /* Advance allocation of bvecq */
+ struct bvecq_pos load_cursor; /* Point at which new folios are loaded in */
+ struct bvecq_pos dispatch_cursor; /* Point from which buffers are dispatched */
+ struct bvecq_pos collect_cursor; /* Clear-up point of I/O buffer */
wait_queue_head_t waitq; /* Processor waiter */
void *netfs_priv; /* Private data for the netfs */
void *netfs_priv2; /* Private data for the netfs */
- struct bio_vec *direct_bv; /* DIO buffer list (when handling iovec-iter) */
+ uoff_t last_end; /* End pos of last folio submitted */
uoff_t submitted; /* Amount submitted for I/O so far */
uoff_t len; /* Length of the request */
size_t transferred; /* Amount to be indicated as transferred */
@@ -266,7 +271,6 @@ struct netfs_io_request {
uoff_t abandon_to; /* Position to abandon folios to */
const struct folio *no_unlock_folio; /* Don't unlock this folio after read */
gfp_t gfp; /* GFP flags to use */
- unsigned int direct_bv_count; /* Number of elements in direct_bv[] */
unsigned int debug_id;
unsigned int rsize; /* Maximum read size (0 for none) */
unsigned int wsize; /* Maximum write size (0 for none) */
@@ -274,7 +278,6 @@ struct netfs_io_request {
unsigned int nr_group_rel; /* Number of refs to release on ->group */
spinlock_t lock; /* Lock for queuing subreqs */
enum netfs_io_origin origin; /* Origin of the request */
- bool direct_bv_unpin; /* T if direct_bv[] must be unpinned */
refcount_t ref;
unsigned long flags;
#define NETFS_RREQ_IN_PROGRESS 0 /* Unlocked when the request completes (has ref) */
@@ -427,7 +430,7 @@ void netfs_single_mark_inode_dirty(struct inode *inode);
ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_iter *iter);
int netfs_writeback_single(struct address_space *mapping,
struct writeback_control *wbc,
- struct iov_iter *iter);
+ struct iov_iter *iter, size_t len);
/* Address operations API */
struct readahead_control;
@@ -454,11 +457,9 @@ void netfs_get_subrequest(struct netfs_io_subrequest *subreq,
enum netfs_sreq_ref_trace what);
void netfs_put_subrequest(struct netfs_io_subrequest *subreq,
enum netfs_sreq_ref_trace what);
-ssize_t netfs_extract_user_iter(struct iov_iter *orig, size_t orig_len,
- struct iov_iter *new,
- iov_iter_extraction_t extraction_flags);
-size_t netfs_limit_iter(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs);
+ssize_t netfs_extract_iter(struct iov_iter *orig, size_t max_len, size_t max_pages,
+ struct bvecq **_bvecq_head,
+ iov_iter_extraction_t extraction_flags, gfp_t gfp);
void netfs_prepare_write_failed(struct netfs_io_subrequest *subreq);
void netfs_write_subrequest_terminated(void *_op, ssize_t transferred_or_error);
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 0adfa6605653d..b7154b81d0d7c 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1411,6 +1411,7 @@ struct readahead_control {
struct file_ra_state *ra;
/* private: use the readahead_* accessors instead */
pgoff_t _index;
+ unsigned int _nr_folios;
unsigned int _nr_pages;
unsigned int _batch_count;
bool dropbehind;
@@ -1590,6 +1591,15 @@ static inline size_t readahead_batch_length(const struct readahead_control *rac)
return rac->_batch_count * PAGE_SIZE;
}
+/**
+ * readahead_folio_count - Get the number of folios in this readahead request.
+ * @rac: The readahead request.
+ */
+static inline unsigned int readahead_folio_count(const struct readahead_control *rac)
+{
+ return rac->_nr_folios;
+}
+
static inline unsigned long dir_pages(const struct inode *inode)
{
return (unsigned long)(inode->i_size + PAGE_SIZE - 1) >>
diff --git a/include/linux/rolling_buffer.h b/include/linux/rolling_buffer.h
deleted file mode 100644
index 5c0bc4221f010..0000000000000
--- a/include/linux/rolling_buffer.h
+++ /dev/null
@@ -1,47 +0,0 @@
-/* SPDX-License-Identifier: GPL-2.0-or-later */
-/* Rolling buffer of folios
- *
- * Copyright (C) 2024 Red Hat, Inc. All Rights Reserved.
- * Written by David Howells (dhowells@redhat.com)
- */
-
-#ifndef _ROLLING_BUFFER_H
-#define _ROLLING_BUFFER_H
-
-#include <linux/bvecq.h>
-#include <linux/uio.h>
-
-/*
- * Rolling buffer. Whilst the buffer is live and in use, folios and bvecq
- * segments can be added to one end by one thread and removed from the other
- * end by another thread. The buffer isn't allowed to be empty; it must always
- * have at least one bvecq in it so that neither side has to modify both queue
- * pointers.
- *
- * The iterator in the buffer is extended as buffers are inserted. It can be
- * snapshotted to use a segment of the buffer.
- */
-struct rolling_buffer {
- struct bvecq *head; /* Producer's insertion point */
- struct bvecq *tail; /* Consumer's removal point */
- struct iov_iter iter; /* Iterator tracking what's left in the buffer */
- u8 first_tail_slot; /* First slot in ->tail */
- bool for_writeback; /* T if being used for writeback */
-};
-
-int rolling_buffer_init(struct rolling_buffer *roll, unsigned int direction,
- gfp_t gfp, bool for_writeback);
-int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp);
-ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,
- struct readahead_control *ractl,
- gfp_t gfp);
-ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio, gfp_t gfp);
-struct bvecq *rolling_buffer_delete_spent(struct rolling_buffer *roll);
-void rolling_buffer_clear(struct rolling_buffer *roll);
-
-static inline void rolling_buffer_advance(struct rolling_buffer *roll, size_t amount)
-{
- iov_iter_advance(&roll->iter, amount);
-}
-
-#endif /* _ROLLING_BUFFER_H */
diff --git a/include/trace/events/netfs.h b/include/trace/events/netfs.h
index e80c27c49ce25..81a3b7aaa187c 100644
--- a/include/trace/events/netfs.h
+++ b/include/trace/events/netfs.h
@@ -229,7 +229,9 @@
EM(netfs_folio_trace_sched_copy, "sched-copy") \
EM(netfs_folio_trace_store, "store") \
EM(netfs_folio_trace_store_copy, "store-copy") \
- E_(netfs_folio_trace_store_plus, "store+")
+ EM(netfs_folio_trace_store_plus, "store+") \
+ EM(netfs_folio_trace_zero, "zero") \
+ E_(netfs_folio_trace_zero_ra, "zero-ra")
#define netfs_collect_contig_traces \
EM(netfs_contig_trace_collect, "Collect") \
@@ -386,10 +388,10 @@ TRACE_EVENT(netfs_sreq,
__entry->len = sreq->len;
__entry->transferred = sreq->transferred;
__entry->start = sreq->start;
- __entry->slot = sreq->io_iter.bvecq_slot;
+ __entry->slot = sreq->io_buffer.slot;
),
- TP_printk("R=%08x[%x] %s %s f=%03x s=%llx %zx/%zx s=%u e=%d",
+ TP_printk("R=%08x[%x] %s %s f=%03x s=%llx %zx/%zx bv=%u e=%d",
__entry->rreq, __entry->index,
__print_symbolic(__entry->source, netfs_sreq_sources),
__print_symbolic(__entry->what, netfs_sreq_traces),
@@ -782,6 +784,30 @@ TRACE_EVENT(netfs_read_progress_at,
__entry->rreq, __entry->cleaned_to, __entry->progress_at)
);
+TRACE_EVENT(netfs_bv_slot,
+ TP_PROTO(const struct bvecq *bq, int slot),
+
+ TP_ARGS(bq, slot),
+
+ TP_STRUCT__entry(
+ __field(unsigned long, pfn)
+ __field(unsigned int, offset)
+ __field(unsigned int, len)
+ __field(unsigned int, slot)
+ ),
+
+ TP_fast_assign(
+ __entry->slot = slot;
+ __entry->pfn = page_to_pfn(bq->bv[slot].bv_page);
+ __entry->offset = bq->bv[slot].bv_offset;
+ __entry->len = bq->bv[slot].bv_len;
+ ),
+
+ TP_printk("bq[%x] p=%lx %x-%x",
+ __entry->slot,
+ __entry->pfn, __entry->offset, __entry->offset + __entry->len)
+ );
+
#undef EM
#undef E_
#endif /* _TRACE_NETFS_H */
diff --git a/mm/readahead.c b/mm/readahead.c
index 6e5563290287e..196542fb1c930 100644
--- a/mm/readahead.c
+++ b/mm/readahead.c
@@ -188,6 +188,7 @@ static void read_pages(struct readahead_control *rac)
if (unlikely(rac->_workingset))
psi_memstall_leave(&rac->_pflags);
rac->_workingset = false;
+ rac->_nr_folios = 0;
BUG_ON(readahead_count(rac));
}
@@ -303,6 +304,7 @@ void page_cache_ra_unbounded(struct readahead_control *ractl,
if (i == mark)
folio_set_readahead(folio);
ractl->_workingset |= folio_test_workingset(folio);
+ ractl->_nr_folios++;
ractl->_nr_pages += min_nrpages;
i += min_nrpages;
}
@@ -473,6 +475,7 @@ static inline int ra_alloc_folio(struct readahead_control *ractl, pgoff_t index,
return err;
}
+ ractl->_nr_folios++;
ractl->_nr_pages += 1UL << order;
ractl->_workingset |= folio_test_workingset(folio);
return 0;
@@ -822,6 +825,7 @@ void readahead_expand(struct readahead_control *ractl,
ractl->_workingset = true;
psi_memstall_enter(&ractl->_pflags);
}
+ ractl->_nr_folios++;
ractl->_nr_pages += min_nrpages;
ractl->_index = folio->index;
}
@@ -851,6 +855,7 @@ void readahead_expand(struct readahead_control *ractl,
ractl->_workingset = true;
psi_memstall_enter(&ractl->_pflags);
}
+ ractl->_nr_folios++;
ractl->_nr_pages += min_nrpages;
if (ra) {
ra->size += min_nrpages;
|
| 3/2 | 2026/10/06 21:22 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 47852 output: 36 thoughts: 770 Thoughts: **Analyzing a Netfs Buffer Representation Overhaul** Okay, so I need to evaluate this patch series in the kernel repo. My first step is to get a handle on what's actually changing. Let's dig into the commit diff. Alright, the list of modified files is quite extensive. It looks like we're dealing with a significant refactoring, judging by the number of files touched: `fs/afs/dir.c`, `fs/afs/symlink.c`, `fs/cachefiles/io.c`, and a whole slew of files in `fs/netfs/`. Oh, and `mm/readahead.c` is in the mix too. Seems this isn't just a localized change; it's got tentacles reaching into multiple core areas. The included and deleted header files further confirm a major change, particularly concerning `rolling_buffer` and `bvecq`. From the filenames, it appears the commit is replacing the `rolling_buffer` implementation with `bvecq` which includes the addition of the cursors. This suggests a big shift in how netfs buffers are represented. I see functions being added to `fs/netfs/bvecq.c` like `bvecq_buffer_init`, `bvecq_buffer_append`, `bvecq_pos_advance`, `bvecq_zero`, `bvecq_slice`, and `bvecq_load_from_ra`. Furthermore, it looks like quite a few critical functions are being updated to accomodate the changes, like `netfs_extract_iter`, `netfs_readahead`, `netfs_create_singular_buffer`, `netfs_read_gaps`, `netfs_unbuffered_read_iter_locked`, `netfs_unbuffered_write_iter_locked`, `netfs_writepages`, `netfs_writeback_single`. This touches the core `fs/netfs` code and `mm/readahead` code, which is concerning but not unexpected. Since it's changing the buffer representation, I need to look closely at these `mm/readahead.c` changes, particularly concerning `rac->_nr_folios`. It also modifies `afs` and `cachefiles` and touches core code, so this could have widespread implications. Netfs is a pretty key component here. As I recall, it provides a crucial layer for network filesystems like AFS, CIFS, Ceph, and 9p. It's often reached through standard filesystem operations—like mounting a 9p share, a CIFS share, or even an AFS volume. Or direct I/O, buffered I/O, syscalls, all working with netfs-backed filesystems. A thought just crossed my mind: Does syzkaller fuzz netfs? Ah, yes! It *does* fuzz 9p (using `CONFIG_NET_9P`), AFS (using `CONFIG_AFS_FS`), Ceph, CIFS (using `CONFIG_CIFS`), and probably more. 9p, in particular, is frequently used with standard QEMU. Syz_mount_image might also be involved here. And netfs itself has `CONFIG_NETFS_SUPPORT`. That makes this change even more critical because the buffer changes could uncover vulnerabilities in these commonly used network filesystems. Now, let's turn our attention to the Kconfig files. I need to understand the relevant configuration options. I'll use `grepper` or just read the Kconfig files to get a solid grasp of how these components are configured and their dependencies. |
| 4/2 | 2026/10/06 21:23 | tool | grepper |
0mArgs:{
"Expression": "config NETFS_SUPPORT",
"PathPrefixes": [
"fs/netfs/",
"fs/"
]
}Results: {
"Output": "fs/netfs/Kconfig-2-\nfs/netfs/Kconfig:3:config NETFS_SUPPORT\nfs/netfs/Kconfig-4-\ttristate\n"
} |
| 5/2 | 2026/10/06 21:23 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 49477 output: 33 thoughts: 18 Thoughts: Here's my take on those thoughts, tailored for an expert audience: **Kconfig Dive for Network Filesystem Details** Okay, so I need to get a handle on the dependencies and configuration options surrounding the network filesystems. My first instinct is to dive into `fs/netfs/Kconfig`. That's where the core declarations and relationships should reside. I expect to find the usual suspects: `depends on` statements linking to networking core features (like sockets, IP, TCP/UDP), maybe specific protocol support (NFS, SMB, etc.), and of course, the configuration variables themselves – `CONFIG_NETFS_ENABLE`, `CONFIG_NFS_SUPPORT`, potentially even granular options for buffer sizes, connection timeout, and authentication methods. I'm looking for a clear overview of how these network filesystems are integrated into the kernel build process. I'll likely also find which features they depend on to be built. |
| 6/2 | 2026/10/06 21:23 | tool | read-file |
0mArgs:{
"File": "fs/netfs/Kconfig",
"FirstLine": 1,
"LineCount": 40
}Results: {
"Contents": " 1:\t# SPDX-License-Identifier: GPL-2.0-only\n 2:\t\n 3:\tconfig NETFS_SUPPORT\n 4:\t\ttristate\n 5:\t\thelp\n 6:\t\t This option enables support for network filesystems, including\n 7:\t\t helpers for high-level buffered I/O, abstracting out read\n 8:\t\t segmentation, local caching and transparent huge page support.\n 9:\t\n 10:\tconfig NETFS_STATS\n 11:\t\tbool \"Gather statistical information on local caching\"\n 12:\t\tdepends on NETFS_SUPPORT \u0026\u0026 PROC_FS\n 13:\t\thelp\n 14:\t\t This option causes statistical information to be gathered on local\n 15:\t\t caching and exported through file:\n 16:\t\n 17:\t\t\t/proc/fs/fscache/stats\n 18:\t\n 19:\t\t The gathering of statistics adds a certain amount of overhead to\n 20:\t\t execution as there are a quite a few stats gathered, and on a\n 21:\t\t multi-CPU system these may be on cachelines that keep bouncing\n 22:\t\t between CPUs. On the other hand, the stats are very useful for\n 23:\t\t debugging purposes. Saying 'Y' here is recommended.\n 24:\t\n 25:\tconfig NETFS_PGPRIV2\n 26:\t\tbool\n 27:\t\n 28:\tconfig NETFS_DEBUG\n 29:\t\tbool \"Enable dynamic debugging netfslib and FS-Cache\"\n 30:\t\tdepends on NETFS_SUPPORT\n 31:\t\thelp\n 32:\t\t This permits debugging to be dynamically enabled in the local caching\n 33:\t\t management module. If this is set, the debugging output may be\n 34:\t\t enabled by setting bits in /sys/module/netfs/parameters/debug.\n 35:\t\n 36:\tconfig FSCACHE\n 37:\t\tbool \"General filesystem local caching manager\"\n 38:\t\tdepends on NETFS_SUPPORT\n 39:\t\thelp\n 40:\t\t This option enables a generic filesystem caching manager that can be\n"
} |
| 7/2 | 2026/10/06 21:23 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 50261 output: 19 thoughts: 11 Thoughts: Okay, here's my attempt at summarizing that thought process, tailored for an expert audience: **Analyzing the NETFS_SUPPORT Selection** Alright, let's dive in. My immediate thought is, "What precisely is influencing the selection of `NETFS_SUPPORT`?" I need to understand the causal chain. Is it a direct dependency, or is it a result of some other, more fundamental configuration option being enabled? I'm looking for the underlying rationale. Am I dealing with a specific platform, kernel version, or a particular user-space utility that's triggering this? This could be a complex web of configurations, so I need to trace the dependencies methodically. I'll need to check the relevant Kconfig files and explore any build options that might be in play. I am looking for the root cause and any knock on effects. This isn't just about identifying the setting; it's about *why* it's set. What implications does this have for the overall system configuration and the potential attack surface? Time to dig in... |
| 8/2 | 2026/10/06 21:23 | tool | grepper |
0mArgs:{
"Expression": "select NETFS_SUPPORT"
}Results: {
"Output": "fs/9p/Kconfig=2=config 9P_FS\n--\nfs/9p/Kconfig-4-\tdepends on NET_9P\nfs/9p/Kconfig:5:\tselect NETFS_SUPPORT\nfs/9p/Kconfig-6-\thelp\n--\nfs/afs/Kconfig=2=config AFS_FS\n--\nfs/afs/Kconfig-6-\tselect DNS_RESOLVER\nfs/afs/Kconfig:7:\tselect NETFS_SUPPORT\nfs/afs/Kconfig-8-\tselect CRYPTO_KRB5\n--\nfs/ceph/Kconfig=2=config CEPH_FS\n--\nfs/ceph/Kconfig-5-\tselect CEPH_LIB\nfs/ceph/Kconfig:6:\tselect NETFS_SUPPORT\nfs/ceph/Kconfig-7-\tselect NETFS_PGPRIV2\n--\nfs/nfs/Kconfig=173=config NFS_FSCACHE\n--\nfs/nfs/Kconfig-175-\tdepends on NFS_FS\nfs/nfs/Kconfig:176:\tselect NETFS_SUPPORT\nfs/nfs/Kconfig-177-\tselect NETFS_PGPRIV2\n--\nfs/smb/client/Kconfig=2=config CIFS\n--\nfs/smb/client/Kconfig-20-\tselect OID_REGISTRY\nfs/smb/client/Kconfig:21:\tselect NETFS_SUPPORT\nfs/smb/client/Kconfig-22-\thelp\n--\nnet/9p/Kconfig=6=menuconfig NET_9P\nnet/9p/Kconfig-7-\ttristate \"Plan 9 Resource Sharing Support (9P2000)\"\nnet/9p/Kconfig:8:\tselect NETFS_SUPPORT\nnet/9p/Kconfig-9-\thelp\n"
} |
| 9/2 | 2026/10/06 21:23 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 50964 output: 27 thoughts: 10 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Considering fs/cachefiles** Right, let's not get tunnel vision on this. While we've been focusing on X, it's prudent to consider other potential areas where we might be seeing the issue. My initial hunch is [insert initial hunch based on prior context/discussion]. However, it's also worth checking fs/cachefiles, particularly if [insert relevant scenario based on prior context/discussion, e.g., dealing with file caching or specific storage interactions]. This could be a blind spot. A quick investigation of its logs and configuration would be prudent. I need to make sure we've explored all the likely areas before reaching a conclusion. Let's see if the underlying cause is something else. |
| 10/2 | 2026/10/06 21:23 | tool | grepper |
0mArgs:{
"Expression": "config CACHEFILES",
"PathPrefixes": [
"fs/"
]
}Results: {
"Output": "fs/cachefiles/Kconfig-2-\nfs/cachefiles/Kconfig:3:config CACHEFILES\nfs/cachefiles/Kconfig-4-\ttristate \"Filesystem caching on files\"\n--\nfs/cachefiles/Kconfig-13-\nfs/cachefiles/Kconfig:14:config CACHEFILES_DEBUG\nfs/cachefiles/Kconfig-15-\tbool \"Debug CacheFiles\"\n--\nfs/cachefiles/Kconfig-22-\nfs/cachefiles/Kconfig:23:config CACHEFILES_ERROR_INJECTION\nfs/cachefiles/Kconfig-24-\tbool \"Provide error injection for cachefiles\"\n"
} |
| 11/2 | 2026/10/06 21:23 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 51337 output: 32 thoughts: 37 Thoughts: Okay, here's my interpretation of this Makefile snippet, framed as my own thought process: **Makefile Analysis: Network Filesystem Module** Alright, let's dive into this `fs/netfs/Makefile` entry. I see immediately that `obj-$(CONFIG_NETFS_SUPPORT) += netfs.o`. This is standard kernel module build system logic. The intention here is clear: 1. **Conditional Compilation:** The `obj-$(...)` construct is the crucial part. It's indicating a *conditional* inclusion of `netfs.o` into the object list. 2. **Configuration Dependency:** The condition itself is driven by the `CONFIG_NETFS_SUPPORT` variable. This is a typical Kconfig variable, defined elsewhere within the kernel's configuration system (probably a `Kconfig` file in the same or a parent directory). It controls whether the `netfs` module will be built and linked into the kernel image. 3. **Module Object:** If `CONFIG_NETFS_SUPPORT` is set to 'y' (or similar, indicating yes), then `netfs.o` will be included in the `obj-` list. It gets added to the list of objects that will be built for this part of the filesystem. 4. **Network Filesystem Module:** This implies that `netfs.o` is the object file containing the compiled code for the network filesystem support. It will, when linked, likely contain the implementation of the filesystem's entry points, such as `mount`, `read`, `write`, `getattr`, and all the other filesystem interface functions specific to this network filesystem. So, in essence, this line ensures that the network filesystem code is compiled and incorporated into the kernel *only* if the user has specifically enabled network filesystem support in the kernel configuration. This keeps the kernel lean and avoids including unnecessary modules. |
| 12/2 | 2026/10/06 21:23 | tool | read-file |
0mArgs:{
"File": "fs/netfs/Makefile",
"FirstLine": 1,
"LineCount": 30
}Results: {
"Contents": " 1:\t# SPDX-License-Identifier: GPL-2.0\n 2:\t\n 3:\tnetfs-y := \\\n 4:\t\tbuffered_read.o \\\n 5:\t\tbuffered_write.o \\\n 6:\t\tbvecq.o \\\n 7:\t\tdirect_read.o \\\n 8:\t\tdirect_write.o \\\n 9:\t\titerator.o \\\n 10:\t\tlocking.o \\\n 11:\t\tmain.o \\\n 12:\t\tmisc.o \\\n 13:\t\tobjects.o \\\n 14:\t\tread_collect.o \\\n 15:\t\tread_retry.o \\\n 16:\t\tread_single.o \\\n 17:\t\twrite_collect.o \\\n 18:\t\twrite_issue.o \\\n 19:\t\twrite_retry.o\n 20:\t\n 21:\tnetfs-$(CONFIG_NETFS_PGPRIV2) += read_pgpriv2.o\n 22:\tnetfs-$(CONFIG_NETFS_STATS) += stats.o\n 23:\t\n 24:\tnetfs-$(CONFIG_FSCACHE) += \\\n 25:\t\tfscache_cache.o \\\n 26:\t\tfscache_cookie.o \\\n 27:\t\tfscache_io.o \\\n 28:\t\tfscache_main.o \\\n 29:\t\tfscache_volume.o\n 30:\t\n"
} |
| 13/2 | 2026/10/06 21:23 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 52177 output: 20 thoughts: 1231 Thoughts: **WorthFuzzing and EnableConfigs Assessment** Okay, let's break this down. First, I see significant changes in `fs/netfs`, particularly around I/O buffering. The old `rolling_buffer` is gone, replaced with cursor-based `bvecq` manipulation – slicing, advancing, zeroing, all that jazz. This looks like a big refactor of I/O, touching on direct and buffered read/write paths, and all the associated locking. Then, there are core memory management modifications in `mm/readahead.c` with the addition of `_nr_folios` in readahead functions. Moreover, there have been changes to `fs/cachefiles/io.c`, `fs/afs/dir.c`, and `fs/afs/symlink.c`. Based on those changes, and considering the "WHEN TO RETURN WorthFuzzing=true" criteria, this is a clear "yes." We've got restructuring of data structures, algorithms, memory management, and assertions all being altered. Given this covers reachable executable code in core subsystems like memory management and VFS, along with assertions, this definitively warrants fuzzing. `WorthFuzzing` is `true`. Now, let's figure out what configs are required. Since `netfs` is changed, `NETFS_SUPPORT` is obviously needed. But how do we enable it? Tristate with no prompt, huh? Interesting. Looks like it can be selected directly, via other filesystems that select it (9P, AFS, CIFS, etc.), or by setting `CONFIG_NETFS_SUPPORT=y` in the config file. Then, since `fs/cachefiles/io.c` has been touched, `FSCACHE` and consequently `CACHEFILES` are probably needed. And with the changes to `fs/afs/dir.c` and `fs/afs/symlink.c`, we should enable `AFS_FS` as well. The presence of `read_pgpriv2.o` seems to require `NETFS_PGPRIV2`. So, let's add `NETFS_SUPPORT`, `FSCACHE`, `CACHEFILES`, `AFS_FS` and `NETFS_PGPRIV2`. |
| 14/2 | 2026/10/06 21:23 | tool | grepper |
0mArgs:{
"Expression": "NETFS_PGPRIV2"
}Results: {
"Output": "fs/ceph/Kconfig=2=config CEPH_FS\n--\nfs/ceph/Kconfig-6-\tselect NETFS_SUPPORT\nfs/ceph/Kconfig:7:\tselect NETFS_PGPRIV2\nfs/ceph/Kconfig-8-\tselect FS_ENCRYPTION_ALGS if FS_ENCRYPTION\n--\nfs/netfs/Kconfig=10=config NETFS_STATS\n--\nfs/netfs/Kconfig-24-\nfs/netfs/Kconfig:25:config NETFS_PGPRIV2\nfs/netfs/Kconfig-26-\tbool\n--\nfs/netfs/Makefile=3=netfs-y := \\\n--\nfs/netfs/Makefile-20-\nfs/netfs/Makefile:21:netfs-$(CONFIG_NETFS_PGPRIV2) += read_pgpriv2.o\nfs/netfs/Makefile-22-netfs-$(CONFIG_NETFS_STATS) += stats.o\n--\nfs/netfs/internal.h=120=void netfs_cache_read_terminated(void *priv, ssize_t transferred_or_error);\n--\nfs/netfs/internal.h-124- */\nfs/netfs/internal.h:125:#ifdef CONFIG_NETFS_PGPRIV2\nfs/netfs/internal.h-126-int netfs_prepare_pgpriv2_write_buffer(struct netfs_io_subrequest *subreq,\n--\nfs/netfs/main.c=37=static const char *netfs_origins[nr__netfs_io_origin] = {\n--\nfs/netfs/main.c-48-\t[NETFS_DIO_WRITE]\t\t= \"DW\",\nfs/netfs/main.c:49:\t[NETFS_PGPRIV2_COPY_TO_CACHE]\t= \"2C\",\nfs/netfs/main.c-50-};\n--\nfs/netfs/read_pgpriv2.c=118=static struct netfs_io_request *netfs_pgpriv2_begin_copy_to_cache(\n--\nfs/netfs/read_pgpriv2.c-126-\tcreq = netfs_create_write_req(rreq-\u003emapping, NULL, folio_pos(folio),\nfs/netfs/read_pgpriv2.c:127:\t\t\t\t NETFS_PGPRIV2_COPY_TO_CACHE);\nfs/netfs/read_pgpriv2.c-128-\tif (IS_ERR(creq))\n--\nfs/netfs/write_collect.c=114=static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,\n--\nfs/netfs/write_collect.c-126-\nfs/netfs/write_collect.c:127:\tif (wreq-\u003eorigin == NETFS_PGPRIV2_COPY_TO_CACHE) {\nfs/netfs/write_collect.c-128-\t\tif (netfs_pgpriv2_unlock_copied_folios(wreq))\n--\nfs/netfs/write_collect.c=216=static void netfs_collect_write_results(struct netfs_io_request *wreq)\n--\nfs/netfs/write_collect.c-233-\tif (wreq-\u003eorigin == NETFS_WRITEBACK ||\nfs/netfs/write_collect.c:234:\t wreq-\u003eorigin == NETFS_PGPRIV2_COPY_TO_CACHE)\nfs/netfs/write_collect.c-235-\t\tnotes = NEED_UNLOCK;\n--\nfs/netfs/write_issue.c=90=struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,\n--\nfs/netfs/write_issue.c-98-\t\t\t origin == NETFS_WRITEBACK_SINGLE ||\nfs/netfs/write_issue.c:99:\t\t\t origin == NETFS_PGPRIV2_COPY_TO_CACHE);\nfs/netfs/write_issue.c-100-\n--\nfs/nfs/Kconfig=173=config NFS_FSCACHE\n--\nfs/nfs/Kconfig-176-\tselect NETFS_SUPPORT\nfs/nfs/Kconfig:177:\tselect NETFS_PGPRIV2\nfs/nfs/Kconfig-178-\tselect FSCACHE\n--\nfs/nfs/fscache.c=264=static int nfs_netfs_init_request(struct netfs_io_request *rreq, struct file *file)\n--\nfs/nfs/fscache.c-266-\tif (!file) {\nfs/nfs/fscache.c:267:\t\tif (WARN_ON_ONCE(rreq-\u003eorigin != NETFS_PGPRIV2_COPY_TO_CACHE))\nfs/nfs/fscache.c-268-\t\t\treturn -EIO;\n--\ninclude/linux/netfs.h=215=enum netfs_io_origin {\n--\ninclude/linux/netfs.h-226-\tNETFS_DIO_WRITE,\t\t/* This is a direct I/O write */\ninclude/linux/netfs.h:227:\tNETFS_PGPRIV2_COPY_TO_CACHE,\t/* [DEPRECATED] This is writing read data to the cache */\ninclude/linux/netfs.h-228-\tnr__netfs_io_origin\n--\ninclude/linux/netfs.h=235=struct netfs_io_request {\n--\ninclude/linux/netfs.h-296-#define NETFS_RREQ_NEED_PUT_RA_REFS\t17\t/* Need to put the folio refs RA gave us */\ninclude/linux/netfs.h:297:#ifdef CONFIG_NETFS_PGPRIV2\ninclude/linux/netfs.h-298-#define NETFS_RREQ_USE_PGPRIV2\t\t31\t/* [DEPRECATED] Use PG_private_2 to mark\n--\ninclude/trace/events/netfs.h-46-\tEM(NETFS_DIO_WRITE,\t\t\t\"DW\")\t\t\\\ninclude/trace/events/netfs.h:47:\tE_(NETFS_PGPRIV2_COPY_TO_CACHE,\t\t\"2C\")\ninclude/trace/events/netfs.h-48-\n"
} |
| 15/2 | 2026/10/06 21:23 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 55475 output: 33 thoughts: 63 Thoughts: Okay, here's my take on those thoughts, framed as my own: **Investigating Filesystem Caching Dependencies** Alright, so `NETFS_PGPRIV2` is pulled in by both `CEPH_FS` and `NFS_FSCACHE`. That's interesting, means it's pretty important for network filesystems in at least a couple of contexts. But what's the broader picture of caching here? The key seems to be `CACHEFILES`, specifically, that this thing is *selected* or enabled under `fs/cachefiles/Kconfig`. Right, let's drill down into `fs/cachefiles/Kconfig` and see what the dependencies and relationships are in this caching ecosystem. I need to understand the role of `CACHEFILES` in this flow. |
| 16/2 | 2026/10/06 21:24 | tool | read-file |
0mArgs:{
"File": "fs/cachefiles/Kconfig",
"FirstLine": 1,
"LineCount": 20
}Results: {
"Contents": " 1:\t# SPDX-License-Identifier: GPL-2.0-only\n 2:\t\n 3:\tconfig CACHEFILES\n 4:\t\ttristate \"Filesystem caching on files\"\n 5:\t\tdepends on NETFS_SUPPORT \u0026\u0026 FSCACHE \u0026\u0026 BLOCK\n 6:\t\thelp\n 7:\t\t This permits use of a mounted filesystem as a cache for other\n 8:\t\t filesystems - primarily networking filesystems - thus allowing fast\n 9:\t\t local disk to enhance the speed of slower devices.\n 10:\t\n 11:\t\t See Documentation/filesystems/caching/cachefiles.rst for more\n 12:\t\t information.\n 13:\t\n 14:\tconfig CACHEFILES_DEBUG\n 15:\t\tbool \"Debug CacheFiles\"\n 16:\t\tdepends on CACHEFILES\n 17:\t\thelp\n 18:\t\t This permits debugging to be dynamically enabled in the filesystem\n 19:\t\t caching on files module. If this is set, the debugging output may be\n 20:\t\t enabled by setting bits in /sys/module/cachefiles/parameters/debug or\n"
} |
| 17/2 | 2026/10/06 21:24 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 56022 output: 32 thoughts: 11 Thoughts: Here's a summary of my thought process, keeping in mind my expert understanding: **Navigating the File System Landscape** Okay, so I need to dig into the file system configuration. My initial instinct is to check the relevant Kconfig files, starting with the foundational ones. The `fs/afs/Kconfig` directory is a good place to begin. This is where I'd expect to find the specific configuration options related to the AFS (Andrew File System) implementation, if it's even present or enabled. I’ll be looking for things like: * **Feature Flags:** The core options that enable or disable AFS support entirely. * **Dependencies:** Any interactions with other file systems or kernel subsystems, and their own configurations. * **Parameters:** Any configuration options that can be tuned to optimize AFS behavior, such as cache size, network settings, or security policies. * **Documentation:** Hopefully, well-written help texts associated with each option. That will be very useful. This will give me a solid base to understand how AFS is managed at compile time, and how it interplays with the rest of the file system architecture. |
| 18/2 | 2026/10/06 21:24 | tool | read-file |
0mArgs:{
"File": "fs/afs/Kconfig",
"FirstLine": 1,
"LineCount": 25
}Results: {
"Contents": " 1:\t# SPDX-License-Identifier: GPL-2.0-only\n 2:\tconfig AFS_FS\n 3:\t\ttristate \"Andrew File System support (AFS)\"\n 4:\t\tdepends on INET\n 5:\t\tselect AF_RXRPC\n 6:\t\tselect DNS_RESOLVER\n 7:\t\tselect NETFS_SUPPORT\n 8:\t\tselect CRYPTO_KRB5\n 9:\t\thelp\n 10:\t\t If you say Y here, you will get an experimental Andrew File System\n 11:\t\t driver. It currently only supports unsecured read-only AFS access.\n 12:\t\n 13:\t\t See \u003cfile:Documentation/filesystems/afs.rst\u003e for more information.\n 14:\t\n 15:\t\t If unsure, say N.\n 16:\t\n 17:\tconfig AFS_DEBUG\n 18:\t\tbool \"AFS dynamic debugging\"\n 19:\t\tdepends on AFS_FS\n 20:\t\thelp\n 21:\t\t Say Y here to make runtime controllable debugging messages appear.\n 22:\t\n 23:\t\t See \u003cfile:Documentation/filesystems/afs.rst\u003e for more information.\n 24:\t\n 25:\t\t If unsure, say N.\n"
} |
| 19/2 | 2026/10/06 21:24 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 56634 output: 63 thoughts: 2898 Thoughts: **Focus Symbol Determination for NetFS Buffer Management Rewrite** Alright, let's break this down. We're looking at a patch that fundamentally alters buffer management within NetFS, and the goal is to pinpoint the *specific* functions that represent the core changes. This isn't just a generic code review; it's about guiding fuzzing efforts to the relevant logic. So, the first step is to discard any ubiquitous functions that are called across the entire kernel, as per the FocusSymbols guidelines. Looking at the configurations, the core support revolves around `NETFS_SUPPORT`, potentially intertwined with `FSCACHE`, `CACHEFILES`, `AFS_FS` (with modifications), and maybe `NETFS_PGPRIV2` (or `CEPH_FS` / `NFS_FSCACHE`). We'll keep these in mind for later context. Now for the hard part, identifying our "focus symbols." I need to analyze the patch diff to find out which functions have been added or changed. I'm going through each file touched by the diff, making notes of any relevant functionality. Initially, I'm noting changes like `afs_dir_writepages` and `afs_symlink_writepages`, which now call `netfs_writeback_single`. `cachefiles_issue_write` is also modified, but the removal of submit_extendable_to logic is less interesting from a buffer management perspective. The changes in `fs/netfs` are where the action is: `netfs_prepare_read_iterator`, `netfs_issue_read`, `netfs_mark_copy_to_cache`, etc. are all modified, and the addition of `bvecq` functions and `netfs_extract_iter` is quite significant. The removal of `netfs_prepare_dio_read_iterator` in `fs/netfs/direct_read.c` is also important. The number of functions is high: this means that the patch introduces a significant change to the internals of netfs buffering. Hold on, I see the instructions on focus symbols are giving guidance. A key takeaway is to avoid ubiquitous lifecycle hot-paths such as memory allocators. And targeting feature-specific API entrypoints. The instruction also instructs to identify non-static, feature-specific caller functions that exercise static inlines. Right, so `mm/readahead.c` is out. Functions like `read_pages`, `page_cache_ra_unbounded`, `ra_alloc_folio`, and `readahead_expand` are called for every readahead across all page cache filesystems. Definitely too generic. They'll dilute the fuzzing. My primary focus here is the NetFS buffer management rewrite. So, `bvecq_slice`, `bvecq_load_from_ra`, `bvecq_zero`, `bvecq_pos_advance`, `bvecq_buffer_init`, `bvecq_buffer_append` from `fs/netfs/bvecq.c` and `netfs_extract_iter` from `fs/netfs/iterator.c` are going on the list. These are the core features of the rework. And also the `netfs` functions: `netfs_readahead`, `netfs_writepages`, `netfs_writeback_single`, `netfs_read_single`, `netfs_unbuffered_read_iter_locked`, `netfs_unbuffered_write_iter_locked`, `netfs_retry_read_subrequests`, `netfs_retry_write_stream`, `netfs_read_unlock_folios`, `netfs_writeback_unlock_folios`, `afs_dir_writepages`, `afs_symlink_writepages`. Now I have to double-check that these functions are actually symbols in the vmlinux image. I'll verify they are defined and exported. I've noted that `bvecq_slice`, `bvecq_load_from_ra`, `bvecq_zero`, `bvecq_pos_advance`, `bvecq_buffer_init`, `bvecq_buffer_append` are non-static in `fs/netfs/bvecq.c`. Similarly, `netfs_extract_iter` is non-static and exported with `EXPORT_SYMBOL_GPL(netfs_extract_iter)`. Okay, I see the `include/linux/bvecq.h` include file, and the static inline helpers within it: `bvecq_pos_set`, `bvecq_pos_unset`, `bvecq_pos_transfer`, `bvecq_pos_move`, `bvecq_pos_nudge`, `bvecq_pos_step`, ` |
| 20/2 | 2026/10/06 21:24 | tool | grepper |
0mArgs:{
"Expression": "netfs_read_unlock_folios|netfs_writeback_unlock_folios|netfs_retry_read_subrequests|netfs_retry_write_stream",
"PathPrefixes": [
"fs/netfs/"
]
}Results: {
"Output": "fs/netfs/read_collect.c=147=void netfs_read_set_unlock_at(struct netfs_io_request *rreq)\n--\nfs/netfs/read_collect.c-176- */\nfs/netfs/read_collect.c:177:static void netfs_read_unlock_folios(struct netfs_io_request *rreq,\nfs/netfs/read_collect.c-178-\t\t\t\t unsigned int *notes)\n--\nfs/netfs/read_collect.c=253=static void netfs_collect_read_results(struct netfs_io_request *rreq)\n--\nfs/netfs/read_collect.c-326-\t\t\t test_bit(NETFS_SREQ_HIT_EOF, \u0026front-\u003eflags))\nfs/netfs/read_collect.c:327:\t\t\t\tnetfs_read_unlock_folios(rreq, \u0026notes);\nfs/netfs/read_collect.c-328-\t\t} else {\n--\nfs/netfs/read_retry.c=12=static void netfs_reissue_read(struct netfs_io_request *rreq,\n--\nfs/netfs/read_retry.c-29- */\nfs/netfs/read_retry.c:30:static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\nfs/netfs/read_retry.c-31-{\n--\nfs/netfs/read_retry.c=283=void netfs_retry_reads(struct netfs_io_request *rreq)\n--\nfs/netfs/read_retry.c-297-\ttrace_netfs_rreq(rreq, netfs_rreq_trace_resubmit);\nfs/netfs/read_retry.c:298:\tnetfs_retry_read_subrequests(rreq);\nfs/netfs/read_retry.c-299-\ttrace_netfs_rreq(rreq, netfs_rreq_trace_retry_end);\n--\nfs/netfs/write_collect.c=57=int netfs_folio_written_back(struct folio *folio)\n--\nfs/netfs/write_collect.c-113- */\nfs/netfs/write_collect.c:114:static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,\nfs/netfs/write_collect.c-115-\t\t\t\t\t unsigned int *notes)\n--\nfs/netfs/write_collect.c=216=static void netfs_collect_write_results(struct netfs_io_request *wreq)\n--\nfs/netfs/write_collect.c-370-\t\tif (wreq-\u003ecleaned_to \u003c wreq-\u003ecollected_to)\nfs/netfs/write_collect.c:371:\t\t\tnetfs_writeback_unlock_folios(wreq, \u0026notes);\nfs/netfs/write_collect.c-372-\t} else {\n--\nfs/netfs/write_retry.c-16- */\nfs/netfs/write_retry.c:17:static void netfs_retry_write_stream(struct netfs_io_request *wreq,\nfs/netfs/write_retry.c-18-\t\t\t\t struct netfs_io_stream *stream)\n--\nfs/netfs/write_retry.c=220=void netfs_retry_writes(struct netfs_io_request *wreq)\n--\nfs/netfs/write_retry.c-247-\t\t\tstream-\u003eneed_retry = false;\nfs/netfs/write_retry.c:248:\t\t\tnetfs_retry_write_stream(wreq, stream);\nfs/netfs/write_retry.c-249-\t\t}\n"
} |
| 21/2 | 2026/10/06 21:25 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 61575 output: 387 thoughts: 1093 Thoughts: **Fuzzing Configuration for NetFS and AFS** Okay, so I need to configure this fuzzing setup correctly. First, I noted that `netfs_retry_reads`, `netfs_retry_writes`, and `netfs_read_collection` are non-static functions. I then checked the `afs_dir_writepages` function, which is static, and the `afs_symlink_writepages` which is non-static. This is good to note. I compiled a list of other non-static functions in `fs/netfs`, making sure to check `netfs_pgpriv2_unlock_copied_folios` and `netfs_unlock_abandoned_read_pages` to confirm they were also non-static. I found them to be as expected. I then re-confirmed that `afs_symlink_writepages` is indeed a non-static function. Next, I needed to determine the necessary kernel configurations. I identified `NETFS_SUPPORT`, `AFS_FS`, `FSCACHE`, `CACHEFILES`, and `NETFS_PGPRIV2` as essential. I checked `Kconfig` files in `fs/cachefiles`, `fs/afs`, `fs/ceph`, and `fs/nfs` to confirm these dependencies and select statements. Specifically, I confirmed the relationships between `CACHEFILES`, `AFS_FS`, `CEPH_FS`, `NFS_FSCACHE` and the necessary support functions. Finally, I verified the required parameters for the `set-results` call: I built the `EnableConfigs` list, which will be the list of strings without the CONFIG_ prefix. It is simply the list of the configuration options, which I constructed as `["NETFS_SUPPORT", "FSCACHE", "CACHEFILES", "AFS_FS", "NETFS_PGPRIV2"]`. I've made the `FocusSymbols` list with specific, non-ubiquitous kernel functions to focus fuzzing on. These are functions I had identified as either added or modified: "bvecq_buffer_init", "bvecq_buffer_append", "bvecq_pos_advance", "bvecq_zero", "bvecq_slice", "bvecq_load_from_ra", "netfs_extract_iter", "netfs_readahead", "netfs_writepages", "netfs_writeback_single", "netfs_read_single", "netfs_unbuffered_read_iter_locked", "netfs_unbuffered_write_iter_locked", "netfs_retry_reads", "netfs_retry_writes", "netfs_pgpriv2_unlock_copied_folios", "afs_symlink_writepages". The `Reasoning` will be a concise explanation of the fuzzing verdict. Since everything seems to be correctly set up, I'm confident in setting `WorthFuzzing` to `true`. Now, I'm ready to call `set-results`. |
| 22/2 | 2026/10/06 21:25 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"NETFS_SUPPORT",
"FSCACHE",
"CACHEFILES",
"AFS_FS",
"NETFS_PGPRIV2"
],
"FocusSymbols": [
"bvecq_buffer_init",
"bvecq_buffer_append",
"bvecq_pos_advance",
"bvecq_zero",
"bvecq_slice",
"bvecq_load_from_ra",
"netfs_extract_iter",
"netfs_readahead",
"netfs_writepages",
"netfs_writeback_single",
"netfs_read_single",
"netfs_unbuffered_read_iter_locked",
"netfs_unbuffered_write_iter_locked",
"netfs_retry_reads",
"netfs_retry_writes",
"netfs_pgpriv2_unlock_copied_folios",
"afs_symlink_writepages"
],
"Reasoning": "The patch replaces the rolling_buffer implementation in netfs with a cursor-based bio_vec queue (bvecq_pos/bvecq) mechanism across read, write, direct I/O, readahead, retry, and writeback paths. It introduces new buffer-management primitives (bvecq_slice, bvecq_pos_advance, bvecq_zero, bvecq_load_from_ra, netfs_extract_iter), refactors I/O dispatch and folio unlock accounting, and updates network filesystems (AFS, Cachefiles) and mm/readahead folio tracking. These changes are in core reachable kernel filesystem paths and warrant fuzzing to identify potential regressions, buffer overflows, or invariant violations.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"NETFS_SUPPORT",
"FSCACHE",
"CACHEFILES",
"AFS_FS",
"NETFS_PGPRIV2"
],
"FocusSymbols": [
"bvecq_buffer_init",
"bvecq_buffer_append",
"bvecq_pos_advance",
"bvecq_zero",
"bvecq_slice",
"bvecq_load_from_ra",
"netfs_extract_iter",
"netfs_readahead",
"netfs_writepages",
"netfs_writeback_single",
"netfs_read_single",
"netfs_unbuffered_read_iter_locked",
"netfs_unbuffered_write_iter_locked",
"netfs_retry_reads",
"netfs_retry_writes",
"netfs_pgpriv2_unlock_copied_folios",
"afs_symlink_writepages"
],
"Reasoning": "The patch replaces the rolling_buffer implementation in netfs with a cursor-based bio_vec queue (bvecq_pos/bvecq) mechanism across read, write, direct I/O, readahead, retry, and writeback paths. It introduces new buffer-management primitives (bvecq_slice, bvecq_pos_advance, bvecq_zero, bvecq_load_from_ra, netfs_extract_iter), refactors I/O dispatch and folio unlock accounting, and updates network filesystems (AFS, Cachefiles) and mm/readahead folio tracking. These changes are in core reachable kernel filesystem paths and warrant fuzzing to identify potential regressions, buffer overflows, or invariant violations.",
"WorthFuzzing": true
} |
| 23/1 | 2026/10/06 21:25 | action | If |
8mArgs:{
"WorthFuzzing": true
}Results: null |
| 24/2 | 2026/10/06 21:25 | agent | kmsan-evaluator |
8mModel:core Args: null Results: {
"KMSANReasoning": "The patch series refactors netfs and related filesystem code (AFS, cachefiles) by replacing the legacy rolling_buffer abstraction with segmented bio_vec queues (bvecq) and cursor tracking (bvecq_pos: load_cursor, dispatch_cursor, collect_cursor).\n\nKMSAN vs KASAN Applicability Analysis:\n1. No uninitialized struct/stack/heap usage: All new and modified structures are cleanly initialized before use. Specifically, netfs_io_request and netfs_io_subrequest are fully zeroed via memset() in netfs_alloc_request() and netfs_alloc_subrequest(); new stack cursor variables (such as dispatch_cursor) are zero-initialized with '= {}' or bvecq_pos_set(); bvecq structures are allocated and zero-initialized via C99 designated initializers in bvecq_alloc_one(); and readahead_control._nr_folios is initialized to zero upon allocation and appropriately tracked.\n2. No kernel-to-user info-leaks: None of the modified data structures are copied or exposed to user space (no copy_to_user, put_user, netlink skb, ioctl, or sockopt changes). Buffer contents transferred to user space in DIO reads are filled by network/cache reads or explicitly zeroed (via bvecq_zero, iov_iter_zero, or folio_zero_segments for gaps/unwritten ranges).\n3. Types of bugs exposed: The refactoring involves queue manipulation, cursor advancement, slot indexing, and reference counting (bvecq_get/bvecq_put). The potential defects here include out-of-bounds accesses on bvec array slots, NULL pointer dereferences on cursor exhaustion, and use-after-free or refcount issues on bvecq nodes. All of these error classes are fully detected by standard debugging tools (KASAN, refcount_t sanity checks, and slab poisoning).\n\nConsequently, the patch series does not introduce uninitialized memory read risks or info-leak vulnerabilities that would require a dedicated KMSAN fuzzing session.",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit baca3ab6cbbda96e6e631014dc876f11c374596c
Author: syz-cluster <triage@syzkaller.com>
Date: Tue Oct 6 21:22:55 2026 +0000
syz-cluster: applied patch under review
diff --git a/fs/afs/dir.c b/fs/afs/dir.c
index 5fac02d2d2814..39cba3f37ecd4 100644
--- a/fs/afs/dir.c
+++ b/fs/afs/dir.c
@@ -2229,8 +2229,9 @@ static int afs_dir_writepages(struct address_space *mapping,
if (test_bit(AFS_VNODE_DIR_VALID, &dvnode->flags)) {
iov_iter_bvec_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0,
- i_size_read(&dvnode->netfs.inode));
- ret = netfs_writeback_single(mapping, wbc, &iter);
+ dvnode->directory_size);
+ ret = netfs_writeback_single(mapping, wbc, &iter,
+ i_size_read(&dvnode->netfs.inode));
if (ret == 1)
ret = 0; /* Skipped write due to lock conflict. */
}
diff --git a/fs/afs/symlink.c b/fs/afs/symlink.c
index 9a611efe6b264..ae03ceff42b85 100644
--- a/fs/afs/symlink.c
+++ b/fs/afs/symlink.c
@@ -248,9 +248,9 @@ int afs_symlink_writepages(struct address_space *mapping,
if (vnode->directory &&
atomic64_read(&vnode->cb_expires_at) != AFS_NO_CB_PROMISE) {
- iov_iter_bvec_queue(&iter, ITER_SOURCE, vnode->directory, 0, 0,
- i_size_read(&vnode->netfs.inode));
- ret = netfs_writeback_single(mapping, wbc, &iter);
+ iov_iter_bvec_queue(&iter, ITER_SOURCE, vnode->directory, 0, 0, PAGE_SIZE);
+ ret = netfs_writeback_single(mapping, wbc, &iter,
+ i_size_read(&vnode->netfs.inode));
}
if (ret == 0) {
diff --git a/fs/cachefiles/io.c b/fs/cachefiles/io.c
index d05059822288c..788439ea6e1c8 100644
--- a/fs/cachefiles/io.c
+++ b/fs/cachefiles/io.c
@@ -546,7 +546,6 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq)
struct netfs_cache_resources *cres = &wreq->cache_resources;
struct cachefiles_object *object = cachefiles_cres_object(cres);
struct cachefiles_cache *cache = object->volume->cache;
- struct netfs_io_stream *stream = &wreq->io_streams[subreq->stream_nr];
const struct cred *saved_cred;
size_t off, pre, post, len = subreq->len;
uoff_t start = subreq->start;
@@ -571,17 +570,6 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq)
}
/* We also need to end on the cache granularity boundary */
- if (start + len == wreq->i_size) {
- size_t part = len & (cache->bsize - 1);
- size_t need = cache->bsize - part;
-
- if (part && stream->submit_extendable_to >= need) {
- len += need;
- subreq->len += need;
- subreq->io_iter.count += need;
- }
- }
-
post = len & (cache->bsize - 1);
if (post) {
len -= post;
diff --git a/fs/netfs/Makefile b/fs/netfs/Makefile
index b1ea4439c1bb4..421dd0be413b3 100644
--- a/fs/netfs/Makefile
+++ b/fs/netfs/Makefile
@@ -14,7 +14,6 @@ netfs-y := \
read_collect.o \
read_retry.o \
read_single.o \
- rolling_buffer.o \
write_collect.o \
write_issue.o \
write_retry.o
diff --git a/fs/netfs/buffered_read.c b/fs/netfs/buffered_read.c
index 052684ce1e347..3ca75b5314eee 100644
--- a/fs/netfs/buffered_read.c
+++ b/fs/netfs/buffered_read.c
@@ -114,26 +114,21 @@ static int netfs_begin_cache_read(struct netfs_io_request *rreq, struct netfs_in
static ssize_t netfs_prepare_read_iterator(struct netfs_io_subrequest *subreq)
{
struct netfs_io_request *rreq = subreq->rreq;
+ struct netfs_io_stream *stream = &rreq->io_streams[0];
+ ssize_t extracted;
size_t rsize = subreq->len;
if (subreq->source == NETFS_DOWNLOAD_FROM_SERVER)
- rsize = umin(rsize, rreq->io_streams[0].sreq_max_len);
-
- subreq->len = rsize;
- if (unlikely(rreq->io_streams[0].sreq_max_segs)) {
- size_t limit = netfs_limit_iter(&rreq->buffer.iter, 0, rsize,
- rreq->io_streams[0].sreq_max_segs);
-
- if (limit < rsize) {
- subreq->len = limit;
- trace_netfs_sreq(subreq, netfs_sreq_trace_limited);
- }
+ rsize = umin(rsize, stream->sreq_max_len);
+
+ bvecq_pos_set(&subreq->io_buffer, &rreq->dispatch_cursor);
+ extracted = bvecq_slice(&rreq->dispatch_cursor, rsize,
+ stream->sreq_max_segs, &subreq->nr_segs);
+ if (extracted < subreq->len) {
+ subreq->len = extracted;
+ trace_netfs_sreq(subreq, netfs_sreq_trace_limited);
}
- subreq->io_iter = rreq->buffer.iter;
-
- iov_iter_truncate(&subreq->io_iter, subreq->len);
- rolling_buffer_advance(&rreq->buffer, subreq->len);
return subreq->len;
}
@@ -192,6 +187,9 @@ void netfs_queue_read(struct netfs_io_request *rreq,
static void netfs_issue_read(struct netfs_io_request *rreq,
struct netfs_io_subrequest *subreq)
{
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
+
switch (subreq->source) {
case NETFS_DOWNLOAD_FROM_SERVER:
rreq->netfs_ops->issue_read(subreq);
@@ -200,10 +198,9 @@ static void netfs_issue_read(struct netfs_io_request *rreq,
netfs_read_cache_to_pagecache(rreq, subreq);
break;
default:
- __set_bit(NETFS_SREQ_CLEAR_TAIL, &subreq->flags);
- subreq->error = 0;
- iov_iter_zero(subreq->len, &subreq->io_iter);
+ bvecq_zero(&subreq->io_buffer, subreq->len);
subreq->transferred = subreq->len;
+ subreq->error = 0;
netfs_read_subreq_terminated(subreq);
break;
}
@@ -215,31 +212,31 @@ static void netfs_issue_read(struct netfs_io_request *rreq,
* otherwise we set the deprecated PG_private_2.
*/
static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,
- struct bvecq **bq,
- unsigned int *offset,
- int *slot,
- size_t len,
- bool copy)
+ struct bvecq_pos *mark, size_t len, bool copy)
{
+ struct bvecq *bq = mark->bvecq;
+ unsigned int offset = mark->offset;
+ int slot = mark->slot;
+
while (len > 0) {
- struct folio *folio;
size_t fsize, overlap;
- if (!*bq)
+ if (!bq)
break;
- if (!bvecq_acquire_slot(*bq, *slot)) {
- *bq = bvecq_next(*bq);
- *slot = 0;
- *offset = 0;
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = bq->next;
+ slot = 0;
+ offset = 0;
continue;
}
/* Determine how much the subreq overlaps the folio, if at all. */
- fsize = (*bq)->bv[*slot].bv_len;
- overlap = min(len, fsize - *offset);
+ fsize = bq->bv[slot].bv_len;
+ overlap = min(len, fsize - offset);
if (overlap > 0 && copy) {
- folio = bvec_folio(&(*bq)->bv[*slot]);
+ struct folio *folio = bvec_folio(&bq->bv[slot]);
+
if (netfs_using_pgpriv2(rreq)) {
if (!folio_test_private_2(folio))
folio_start_private_2(folio);
@@ -251,12 +248,20 @@ static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,
}
len -= overlap;
- *offset += overlap;
- if (*offset >= fsize) {
- *slot += 1;
- *offset = 0;
+ offset += overlap;
+ if (offset >= fsize) {
+ slot += 1;
+ offset = 0;
}
}
+
+ if (bq) {
+ bvecq_pos_move(mark, bq);
+ mark->offset = offset;
+ mark->slot = slot;
+ } else {
+ bvecq_pos_unset(mark);
+ }
}
/*
@@ -275,11 +280,14 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)
.cached_to[1] = ULLONG_MAX,
};
struct fscache_occupancy *occ = &_occ;
- struct bvecq *bq = rreq->buffer.tail;
- unsigned int offset = 0;
+ struct bvecq_pos mark_cursor;
ssize_t size = rreq->len;
uoff_t start = rreq->start;
- int ret = 0, slot = 0;
+ int ret = 0;
+
+ _enter("R=%08x", rreq->debug_id);
+
+ bvecq_pos_set(&mark_cursor, &rreq->dispatch_cursor);
do {
int (*prepare_read)(struct netfs_io_subrequest *subreq) = NULL;
@@ -408,10 +416,10 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)
if (size <= 0)
netfs_all_subreqs_queued(rreq);
- if (bq) {
+ if (mark_cursor.bvecq) {
/* See if the cache indicated this should be cached. */
copy = test_bit(NETFS_SREQ_COPY_TO_CACHE, &subreq->flags);
- netfs_mark_copy_to_cache(rreq, &bq, &slot, &offset, slice, copy);
+ netfs_mark_copy_to_cache(rreq, &mark_cursor, slice, copy);
}
trace_netfs_sreq(subreq, netfs_sreq_trace_submit);
@@ -432,6 +440,9 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)
/* Defer error return as we may need to wait for outstanding I/O. */
cmpxchg(&rreq->error, 0, ret);
+
+ bvecq_pos_unset(&mark_cursor);
+ bvecq_pos_unset(&rreq->dispatch_cursor);
}
/**
@@ -479,7 +490,7 @@ void netfs_readahead(struct readahead_control *ractl)
* acquires a ref on each folio that we will need to release later -
* but we don't want to do that until after we've started the I/O.
*/
- added = rolling_buffer_bulk_load_from_ra(&rreq->buffer, ractl, rreq->gfp);
+ added = bvecq_load_from_ra(&rreq->dispatch_cursor, ractl);
if (added < 0) {
ret = added;
goto cleanup_free;
@@ -488,6 +499,7 @@ void netfs_readahead(struct readahead_control *ractl)
rreq->submitted = rreq->start + added;
rreq->cleaned_to = rreq->start;
+ bvecq_pos_set(&rreq->collect_cursor, &rreq->dispatch_cursor);
netfs_read_set_unlock_at(rreq);
netfs_read_to_pagecache(rreq);
@@ -500,20 +512,26 @@ void netfs_readahead(struct readahead_control *ractl)
EXPORT_SYMBOL(netfs_readahead);
/*
- * Create a rolling buffer with a single occupying folio.
+ * Create a buffer queue with a single occupying folio.
*/
static int netfs_create_singular_buffer(struct netfs_io_request *rreq, struct folio *folio)
{
- ssize_t added;
+ struct bvecq *bq;
+ size_t fsize = folio_size(folio);
- if (rolling_buffer_init(&rreq->buffer, ITER_DEST, rreq->gfp, false) < 0)
+ bq = bvecq_alloc_one(1, rreq->gfp, false);
+ if (!bq)
return -ENOMEM;
- added = rolling_buffer_append(&rreq->buffer, folio, rreq->gfp);
- if (added < 0)
- return added;
- rreq->submitted = rreq->start + added;
- rreq->progress_at = added;
+ rreq->dispatch_cursor.bvecq = bq;
+ rreq->dispatch_cursor.slot = 0;
+ rreq->dispatch_cursor.offset = 0;
+
+ bvec_set_folio(&bq->bv[0], folio, fsize, 0);
+ bvecq_filled_to(bq, 1);
+ bvecq_pos_set(&rreq->collect_cursor, &rreq->dispatch_cursor);
+ rreq->submitted = rreq->start + fsize;
+ rreq->progress_at = fsize;
return 0;
}
@@ -527,14 +545,14 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
struct netfs_group *group = netfs_folio_group(folio);
struct netfs_folio *finfo = netfs_folio_info(folio);
struct netfs_inode *ctx = netfs_inode(mapping->host);
- struct bio_vec *bvec = NULL;
+ struct bvecq *bq = NULL;
unsigned int from = finfo->dirty_offset;
unsigned int to = from + finfo->dirty_len;
unsigned int off = 0;
size_t flen = folio_size(folio);
size_t nr_bvec = flen / PAGE_SIZE + 2;
size_t part;
- int ret, i = 0, sink_from = -1, sink_to = -1;
+ int ret, i = 0;
_enter("%lx", folio->index);
@@ -555,31 +573,46 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
* end get copied to, but the middle is discarded.
*/
ret = -ENOMEM;
- bvec = kmalloc_objs(*bvec, nr_bvec);
- if (!bvec)
+ bq = bvecq_alloc_chain(nr_bvec, rreq->gfp, false);
+ if (!bq)
goto discard;
+ rreq->dispatch_cursor.bvecq = bq;
trace_netfs_folio(folio, netfs_folio_trace_read_gaps);
+ for (struct bvecq *p = bq; p; p = p->next)
+ p->mem_type = BVECQ_MEM_PAGECACHE;
+
if (from > 0) {
- bvec_set_folio(&bvec[i++], folio, from, 0);
+ folio_get(folio);
+ bvec_set_folio(&bq->bv[i++], folio, from, 0);
off = from;
}
- sink_from = i;
while (off < to) {
struct folio *sink = folio_alloc(GFP_KERNEL, 0);
if (!sink)
goto discard;
- part = min_t(size_t, to - off, PAGE_SIZE);
- bvec_set_folio(&bvec[i], sink, part, 0);
+ if (i >= bq->max_slots) {
+ bvecq_filled_to(bq, i);
+ bq = bq->next;
+ i = 0;
+ }
+ part = min(to - off, PAGE_SIZE);
+ bvec_set_folio(&bq->bv[i++], sink, part, 0);
off += part;
- sink_to = i;
- i++;
}
- if (to < flen)
- bvec_set_folio(&bvec[i++], folio, flen - to, to);
- iov_iter_bvec(&rreq->buffer.iter, ITER_DEST, bvec, i, rreq->len);
+ if (to < flen) {
+ if (i >= bq->max_slots) {
+ bvecq_filled_to(bq, i);
+ bq = bq->next;
+ i = 0;
+ }
+ folio_get(folio);
+ bvec_set_folio(&bq->bv[i++], folio, flen - to, to);
+ }
+ bvecq_filled_to(bq, i);
+
rreq->submitted = rreq->start + flen;
netfs_read_to_pagecache(rreq);
@@ -587,11 +620,10 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
ret = netfs_wait_for_read(rreq);
if (ret >= 0) {
if (ret < flen) {
- struct iov_iter iter;
-
- iov_iter_bvec(&iter, ITER_DEST, bvec, i, flen);
- iov_iter_advance(&iter, ret);
- iov_iter_zero(flen - ret, &iter);
+ if (ret < from)
+ folio_zero_segments(folio, ret, from, to, flen);
+ else
+ folio_zero_segment(folio, max(to, ret), flen);
}
if (group)
folio_change_private(folio, group);
@@ -603,22 +635,16 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
folio_mark_uptodate(folio);
}
- if (sink_to >= 0)
- for (; sink_from <= sink_to; sink_from++)
- folio_put(bvec_folio(&bvec[sink_from]));
- kfree(bvec);
+ bvecq_pos_unset(&rreq->dispatch_cursor);
folio_unlock(folio);
netfs_put_request(rreq, netfs_rreq_trace_put_return);
return ret < 0 ? ret : 0;
discard:
+ bvecq_pos_unset(&rreq->dispatch_cursor);
netfs_put_failed_request(rreq);
alloc_error:
folio_unlock(folio);
- if (sink_to >= 0)
- for (; sink_from <= sink_to; sink_from++)
- folio_put(bvec_folio(&bvec[sink_from]));
- kfree(bvec);
return ret;
}
diff --git a/fs/netfs/bvecq.c b/fs/netfs/bvecq.c
index 5b747f6b59382..2edc0045c2453 100644
--- a/fs/netfs/bvecq.c
+++ b/fs/netfs/bvecq.c
@@ -342,3 +342,313 @@ int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size,
return 0;
}
EXPORT_SYMBOL(bvecq_expand_buffer);
+
+/**
+ * bvecq_buffer_init - Initialise a buffer and set position
+ * @pos: The position to point at the new buffer.
+ * @gfp: The allocation constraints.
+ * @for_writeback: True if allocating for writeback
+ *
+ * Initialise a rolling buffer. We allocate an unpopulated bvecq node to so
+ * that the pointers can be independently driven by the producer and the
+ * consumer.
+ *
+ * Return 0 if successful; -ENOMEM on allocation failure.
+ */
+int bvecq_buffer_init(struct bvecq_pos *pos, gfp_t gfp, bool for_writeback)
+{
+ struct bvecq *bq;
+
+ bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, for_writeback);
+ if (!bq)
+ return -ENOMEM;
+
+ pos->bvecq = bq; /* Comes with a ref. */
+ pos->slot = 0;
+ pos->offset = 0;
+ return 0;
+}
+
+/**
+ * bvecq_buffer_append - Append a new bvecq node to a buffer
+ * @pos: The position of the last node.
+ * @bq: The buffer to add.
+ *
+ * Add a new node on to the buffer chain at the specified position, either
+ * because the previous one is full or because we have a discontiguity to
+ * contend with, and update @pos to point to it.
+ */
+void bvecq_buffer_append(struct bvecq_pos *pos, struct bvecq *bq)
+{
+ struct bvecq *head = pos->bvecq;
+
+ pos->bvecq = bvecq_get(bq);
+ pos->slot = 0;
+ pos->offset = 0;
+
+ /* [!] NOTE: After we set head->next, the consumer is at liberty to
+ * immediately delete the old head.
+ */
+ bvecq_append(head, bq);
+ bvecq_put(head);
+}
+
+/**
+ * bvecq_pos_advance - Advance a bvecq position
+ * @pos: The position to advance.
+ * @amount: The amount of bytes to advance by.
+ *
+ * Advance the specified bvecq position by @amount bytes. @pos is updated and
+ * bvecq ref counts may have been manipulated. If the position hits the end of
+ * the queue, then it is left pointing beyond the last slot of the last bvecq
+ * so that it doesn't break the chain.
+ */
+void bvecq_pos_advance(struct bvecq_pos *pos, size_t amount)
+{
+ struct bvecq *bq = pos->bvecq, *next;
+ unsigned int slot = pos->slot;
+ size_t offset = pos->offset;
+
+ while (amount) {
+ size_t part;
+
+ if (!bvecq_acquire_slot(bq, slot)) {
+ next = bvecq_next(bq);
+ if (!next) {
+ WARN_ON_ONCE(amount > 0);
+ break;
+ }
+ if (bvecq_acquire_slot(bq, slot))
+ continue; /* More slots got added. */
+ bq = next;
+ slot = 0;
+ offset = 0;
+ continue;
+ }
+
+ part = bq->bv[slot].bv_len - offset;
+
+ if (part > amount) {
+ offset += amount;
+ break;
+ }
+ amount -= part;
+ offset = 0;
+ slot++;
+ }
+
+ pos->slot = slot;
+ pos->offset = offset;
+ bvecq_pos_move(pos, bq);
+}
+
+/*
+ * Clear part of the memory pointed to by a bio_vec.
+ */
+static void bvec_zero(const struct bio_vec *bv, size_t offset, size_t len)
+{
+ struct page *page = bv->bv_page;
+
+ offset += bv->bv_offset;
+
+ page += offset / PAGE_SIZE;
+ offset = offset % PAGE_SIZE;
+
+ while (len) {
+ size_t part = min(len, PAGE_SIZE - offset);
+ char *p = kmap_local_page(page);
+
+ memset(p + offset, 0, part);
+ kunmap_local(p);
+
+ len -= part;
+ offset = 0;
+ page++;
+ }
+}
+
+/**
+ * bvecq_zero - Clear memory starting at the bvecq position.
+ * @pos: The position in the bvecq chain to start clearing.
+ * @amount: The number of bytes to clear.
+ *
+ * Clear memory fragments pointed to by a bvec queue. @pos is updated and
+ * bvecq ref counts may have been manipulated. If the position hits the end of
+ * the queue, then it is left pointing beyond the last slot of the last bvecq
+ * so that it doesn't break the chain.
+ *
+ * Return: The number of bytes cleared.
+ */
+ssize_t bvecq_zero(struct bvecq_pos *pos, size_t amount)
+{
+ struct bvecq *bq = pos->bvecq, *next;
+ unsigned int slot = pos->slot;
+ ssize_t cleared = 0;
+ size_t offset = pos->offset;
+
+ while (amount) {
+ const struct bio_vec *bv;
+ size_t part;
+
+ if (!bvecq_acquire_slot(bq, slot)) {
+ next = bvecq_next(bq);
+ if (!next) {
+ WARN_ON_ONCE(amount > 0);
+ break;
+ }
+ if (bvecq_acquire_slot(bq, slot))
+ continue; /* More slots got added. */
+ bq = next;
+ slot = 0;
+ offset = 0;
+ continue;
+ }
+
+ bv = &bq->bv[slot];
+ if (offset >= bv->bv_len) {
+ slot++;
+ offset = 0;
+ continue;
+ }
+
+ part = min(bv->bv_len - offset, amount);
+ bvec_zero(bv, offset, part);
+ cleared += part;
+ offset += part;
+ amount -= part;
+ }
+
+ pos->slot = slot;
+ pos->offset = offset;
+ bvecq_pos_move(pos, bq);
+ return cleared;
+}
+
+/**
+ * bvecq_slice - Find a slice of a bvecq queue
+ * @pos: The position to start at.
+ * @max_size: The maximum size of the slice (or ULONG_MAX).
+ * @max_slots: The maximum number of slots in the slice (or INT_MAX).
+ * @_nr_slots: Where to put the number of slots (updated).
+ *
+ * Determine the size and number of slots that can be obtained the next slice
+ * of bvec queue up to the maximum size and slot count specified.
+ *
+ * @pos is updated to the end of the slice. If the position hits the end of
+ * the queue, then it is left pointing beyond the last slot of the last bvecq
+ * so that it doesn't break the chain.
+ *
+ * Return: The number of bytes in the slice.
+ */
+size_t bvecq_slice(struct bvecq_pos *pos, size_t max_size,
+ unsigned int max_slots, unsigned int *_nr_slots)
+{
+ struct bvecq *bq, *next;
+ unsigned int slot = pos->slot, nslots = 0;
+ size_t size = 0, offset = pos->offset;
+
+ bq = pos->bvecq;
+ for (;;) {
+ for (; slot < bvecq_nr_slots_acquire(bq); slot++) {
+ const struct bio_vec *bvec = &bq->bv[slot];
+
+ if (offset < bvec->bv_len && bvec->bv_page) {
+ size_t part = min(bvec->bv_len - offset, max_size);
+
+ size += part;
+ offset += part;
+ max_size -= part;
+ nslots++;
+ if (!max_size || nslots >= max_slots)
+ goto out;
+ }
+ offset = 0;
+ }
+
+ /* pos->bvecq isn't allowed to go NULL as the queue may get
+ * extended and we would lose our place.
+ */
+ next = bvecq_next(bq);
+ if (!next)
+ break;
+ if (bvecq_acquire_slot(bq, slot))
+ continue; /* More slots got added. */
+ slot = 0;
+ bq = next;
+ }
+
+out:
+ *_nr_slots = nslots;
+ if (slot == bvecq_nr_slots_acquire(bq)) {
+ next = bvecq_next(bq);
+ if (next) {
+ bq = next;
+ slot = 0;
+ offset = 0;
+ }
+ }
+ bvecq_pos_move(pos, bq);
+ pos->slot = slot;
+ pos->offset = offset;
+ return size;
+}
+
+/**
+ * bvecq_load_from_ra - Allocate a bvecq chain and load from readahead
+ * @pos: Blank position object to attach the new chain to.
+ * @ractl: The readahead control context.
+ *
+ * Decant the set of folios to be read from the readahead context into a bvecq
+ * chain. Each folio occupies one bio_vec element.
+ *
+ * Return: Amount of data loaded or -ENOMEM on allocation failure.
+ */
+ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl)
+{
+ XA_STATE(xas, &ractl->mapping->i_pages, ractl->_index);
+ struct folio *folio;
+ struct bvecq *bq;
+ unsigned int slot = 0;
+ size_t loaded = 0;
+
+ bq = bvecq_alloc_chain(ractl->_nr_folios, GFP_KERNEL, false);
+ if (!bq)
+ return -ENOMEM;
+
+ pos->bvecq = bq;
+ pos->slot = 0;
+ pos->offset = 0;
+
+ rcu_read_lock();
+
+ xas_for_each(&xas, folio, ractl->_index + ractl->_nr_pages - 1) {
+ size_t len;
+
+ if (xas_retry(&xas, folio))
+ continue;
+ VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
+
+ len = folio_size(folio);
+ bvec_set_folio(&bq->bv[slot], folio, len, 0);
+ loaded += len;
+ slot++;
+ trace_netfs_folio(folio, netfs_folio_trace_read);
+
+ if (slot >= bq->max_slots) {
+ bvecq_filled_to(bq, slot);
+ bq = bq->next;
+ if (!bq)
+ break;
+ slot = 0;
+ }
+ }
+
+ rcu_read_unlock();
+
+ if (bq)
+ bvecq_filled_to(bq, slot);
+
+ ractl->_index += ractl->_nr_pages;
+ ractl->_nr_pages = 0;
+ return loaded;
+}
diff --git a/fs/netfs/direct_read.c b/fs/netfs/direct_read.c
index 8c15f30797238..dae890e8df285 100644
--- a/fs/netfs/direct_read.c
+++ b/fs/netfs/direct_read.c
@@ -16,44 +16,21 @@
#include <linux/netfs.h>
#include "internal.h"
-static void netfs_prepare_dio_read_iterator(struct netfs_io_subrequest *subreq)
-{
- struct netfs_io_request *rreq = subreq->rreq;
- size_t rsize;
-
- rsize = umin(subreq->len, rreq->io_streams[0].sreq_max_len);
- subreq->len = rsize;
-
- if (unlikely(rreq->io_streams[0].sreq_max_segs)) {
- size_t limit = netfs_limit_iter(&rreq->buffer.iter, 0, rsize,
- rreq->io_streams[0].sreq_max_segs);
-
- if (limit < rsize) {
- subreq->len = limit;
- trace_netfs_sreq(subreq, netfs_sreq_trace_limited);
- }
- }
-
- trace_netfs_sreq(subreq, netfs_sreq_trace_prepare);
-
- subreq->io_iter = rreq->buffer.iter;
- iov_iter_truncate(&subreq->io_iter, subreq->len);
- iov_iter_advance(&rreq->buffer.iter, subreq->len);
-}
-
/*
* Perform a read to a buffer from the server, slicing up the region to be read
* according to the network rsize.
*/
static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)
{
+ struct netfs_io_stream *stream = &rreq->io_streams[0];
ssize_t size = rreq->len;
uoff_t start = rreq->start;
int ret;
+ bvecq_pos_set(&rreq->collect_cursor, &rreq->dispatch_cursor);
+
do {
struct netfs_io_subrequest *subreq;
- ssize_t slice;
subreq = netfs_alloc_subrequest(rreq, NETFS_DOWNLOAD_FROM_SERVER);
if (!subreq) {
@@ -78,14 +55,22 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)
}
}
- netfs_prepare_dio_read_iterator(subreq);
- slice = subreq->len;
- size -= slice;
- start += slice;
- rreq->submitted += slice;
+ bvecq_pos_set(&subreq->io_buffer, &rreq->dispatch_cursor);
+ subreq->len = bvecq_slice(&rreq->dispatch_cursor,
+ umin(size, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
+
+ size -= subreq->len;
+ start += subreq->len;
+ rreq->submitted += subreq->len;
if (size <= 0)
netfs_all_subreqs_queued(rreq);
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset,
+ subreq->len);
+
rreq->netfs_ops->issue_read(subreq);
if (test_bit(NETFS_RREQ_PAUSE, &rreq->flags))
@@ -99,6 +84,8 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)
netfs_all_subreqs_queued(rreq);
netfs_wake_collector(rreq);
}
+
+ bvecq_pos_unset(&rreq->dispatch_cursor);
}
/*
@@ -177,25 +164,17 @@ ssize_t netfs_unbuffered_read_iter_locked(struct kiocb *iocb, struct iov_iter *i
* buffer for ourselves as the caller's iterator will be trashed when
* we return.
*
- * In such a case, extract an iterator to represent as much of the the
- * output buffer as we can manage. Note that the extraction might not
- * be able to allocate a sufficiently large bvec array and may shorten
- * the request.
+ * Extract a buffer queue to represent as much of the output buffer as
+ * we can manage. The fragments are extracted into a bvecq which will
+ * have sufficient nodes allocated to hold all the data, though this
+ * may end up truncated if ENOMEM is encountered.
*/
- if (user_backed_iter(iter)) {
- ret = netfs_extract_user_iter(iter, rreq->len, &rreq->buffer.iter, 0);
- if (ret < 0)
- goto error_put;
- rreq->direct_bv = (struct bio_vec *)rreq->buffer.iter.bvec;
- rreq->direct_bv_count = ret;
- rreq->direct_bv_unpin = iov_iter_extract_will_pin(iter);
- rreq->len = iov_iter_count(&rreq->buffer.iter);
- } else {
- rreq->buffer.iter = *iter;
- rreq->len = orig_count;
- rreq->direct_bv_unpin = false;
- iov_iter_advance(iter, orig_count);
- }
+ ret = netfs_extract_iter(iter, rreq->len, INT_MAX,
+ &rreq->dispatch_cursor.bvecq, 0, rreq->gfp);
+ if (ret < 0)
+ goto error_put;
+
+ rreq->len = ret;
// TODO: Set up bounce buffer if needed
diff --git a/fs/netfs/direct_write.c b/fs/netfs/direct_write.c
index cc46b7d9321f1..65c61fc67f9bf 100644
--- a/fs/netfs/direct_write.c
+++ b/fs/netfs/direct_write.c
@@ -73,7 +73,11 @@ static void netfs_unbuffered_write_collect(struct netfs_io_request *wreq,
spin_unlock(&wreq->lock);
wreq->transferred += subreq->transferred;
- iov_iter_advance(&wreq->buffer.iter, subreq->transferred);
+ if (subreq->transferred < subreq->len) {
+ bvecq_pos_unset(&wreq->dispatch_cursor);
+ bvecq_pos_transfer(&wreq->dispatch_cursor, &subreq->io_buffer);
+ bvecq_pos_advance(&wreq->dispatch_cursor, subreq->transferred);
+ }
stream->collected_to = subreq->start + subreq->transferred;
wreq->collected_to = stream->collected_to;
@@ -99,6 +103,8 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
_enter("%llx", wreq->len);
+ bvecq_pos_set(&wreq->collect_cursor, &wreq->dispatch_cursor);
+
if (wreq->origin == NETFS_DIO_WRITE)
inode_dio_begin(wreq->inode);
@@ -116,6 +122,8 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
break;
}
stream->construct = NULL;
+ } else {
+ bvecq_pos_set(&subreq->io_buffer, &wreq->dispatch_cursor);
}
/* Check if (re-)preparation failed. */
@@ -125,9 +133,16 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
break;
}
- iov_iter_truncate(&subreq->io_iter, wreq->len - wreq->transferred);
+ subreq->len = bvecq_slice(&wreq->dispatch_cursor, stream->sreq_max_len,
+ stream->sreq_max_segs, &subreq->nr_segs);
+
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_SOURCE,
+ subreq->io_buffer.bvecq, subreq->io_buffer.slot,
+ subreq->io_buffer.offset,
+ subreq->len);
+
if (!iov_iter_count(&subreq->io_iter)) {
- pr_warn("netfs: Unexpected zero-length iterator R=%08x\n",
+ pr_warn("netfs: Unexpected zero-length slice R=%08x\n",
wreq->debug_id);
__set_bit(NETFS_SREQ_FAILED, &subreq->flags);
netfs_write_subrequest_terminated(subreq, -EIO);
@@ -135,12 +150,6 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
break;
}
- subreq->len = netfs_limit_iter(&subreq->io_iter, 0,
- stream->sreq_max_len,
- stream->sreq_max_segs);
- iov_iter_truncate(&subreq->io_iter, subreq->len);
- stream->submit_extendable_to = subreq->len;
-
trace_netfs_sreq(subreq, netfs_sreq_trace_submit);
stream->issue_write(subreq);
@@ -175,9 +184,13 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
*/
subreq->error = -EAGAIN;
trace_netfs_sreq(subreq, netfs_sreq_trace_retry);
+
+ bvecq_pos_unset(&wreq->dispatch_cursor);
+ bvecq_pos_transfer(&wreq->dispatch_cursor, &subreq->io_buffer);
+
if (subreq->transferred > 0) {
- iov_iter_advance(&wreq->buffer.iter, subreq->transferred);
wreq->transferred += subreq->transferred;
+ bvecq_pos_advance(&wreq->dispatch_cursor, subreq->transferred);
}
if (stream->source == NETFS_UPLOAD_TO_SERVER &&
@@ -188,7 +201,6 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
__clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
__clear_bit(NETFS_SREQ_BOUNDARY, &subreq->flags);
__clear_bit(NETFS_SREQ_FAILED, &subreq->flags);
- subreq->io_iter = wreq->buffer.iter;
subreq->start = wreq->start + wreq->transferred;
subreq->len = wreq->len - wreq->transferred;
subreq->transferred = 0;
@@ -204,6 +216,7 @@ static int netfs_unbuffered_write(struct netfs_io_request *wreq)
netfs_stat(&netfs_n_wh_retry_write_subreq);
}
+ bvecq_pos_unset(&wreq->dispatch_cursor);
netfs_unbuffered_write_done(wreq);
_leave(" = %d", ret);
return ret;
@@ -222,10 +235,10 @@ static void netfs_unbuffered_write_async(struct work_struct *work)
* encrypted file. This can also be used for direct I/O writes.
*/
ssize_t netfs_unbuffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *iter,
- struct netfs_group *netfs_group)
+ struct netfs_group *netfs_group)
{
struct netfs_io_request *wreq;
- ssize_t ret, n;
+ ssize_t ret;
uoff_t start = iocb->ki_pos;
uoff_t end = start + iov_iter_count(iter);
size_t len = iov_iter_count(iter);
@@ -261,25 +274,17 @@ ssize_t netfs_unbuffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *
* allocate a sufficiently large bvec array and may shorten the
* request.
*/
- if (user_backed_iter(iter)) {
- n = netfs_extract_user_iter(iter, len, &wreq->buffer.iter, 0);
- if (n < 0) {
- ret = n;
- goto error_put;
- }
- wreq->direct_bv = (struct bio_vec *)wreq->buffer.iter.bvec;
- wreq->direct_bv_count = n;
- wreq->direct_bv_unpin = iov_iter_extract_will_pin(iter);
- } else {
- /* If this is a kernel-generated async DIO request,
- * assume that any resources the iterator points to
- * (eg. a bio_vec array) will persist till the end of
- * the op.
- */
- wreq->buffer.iter = *iter;
- }
+ ssize_t n = netfs_extract_iter(iter, len, INT_MAX,
+ &wreq->dispatch_cursor.bvecq, 0, wreq->gfp);
- wreq->len = iov_iter_count(&wreq->buffer.iter);
+ if (n < 0) {
+ ret = n;
+ goto error_put;
+ }
+ wreq->len = n;
+ _debug("dio-write %zx/%zx %u/%u",
+ n, len, wreq->dispatch_cursor.bvecq->nr_slots,
+ wreq->dispatch_cursor.bvecq->max_slots);
}
__set_bit(NETFS_RREQ_USE_IO_ITER, &wreq->flags);
diff --git a/fs/netfs/internal.h b/fs/netfs/internal.h
index f2a86abae9b3e..2760bce732b83 100644
--- a/fs/netfs/internal.h
+++ b/fs/netfs/internal.h
@@ -70,7 +70,6 @@ static inline void netfs_proc_del_rreq(struct netfs_io_request *rreq) {}
/*
* misc.c
*/
-void netfs_reset_iter(struct netfs_io_subrequest *subreq);
void netfs_wake_collector(struct netfs_io_request *rreq);
void netfs_subreq_clear_in_progress(struct netfs_io_subrequest *subreq);
void netfs_wait_for_in_progress_stream(struct netfs_io_request *rreq,
@@ -239,8 +238,7 @@ void netfs_prepare_write(struct netfs_io_request *wreq,
struct netfs_io_stream *stream,
uoff_t start);
void netfs_reissue_write(struct netfs_io_stream *stream,
- struct netfs_io_subrequest *subreq,
- struct iov_iter *source);
+ struct netfs_io_subrequest *subreq);
void netfs_issue_write(struct netfs_io_request *wreq,
struct netfs_io_stream *stream);
size_t netfs_advance_write(struct netfs_io_request *wreq,
diff --git a/fs/netfs/iterator.c b/fs/netfs/iterator.c
index 31748526d5682..dc97e5b0d4495 100644
--- a/fs/netfs/iterator.c
+++ b/fs/netfs/iterator.c
@@ -14,296 +14,144 @@
#include "internal.h"
/**
- * netfs_extract_user_iter - Extract the pages from a user iterator into a bvec
+ * netfs_extract_iter - Extract virtually contiguous pages from an iterator into a bvecq
* @orig: The original iterator
- * @orig_len: The amount of iterator to copy
- * @new: The iterator to be set up
+ * @max_len: Maximum number of bytes to extract
+ * @max_pages: Maximum number of pages to extract
+ * @_bvecq_head: Where to cache the bvec queue
* @extraction_flags: Flags to qualify the request
+ * @gfp: Allocation mode for bvecq structs.
*
- * Extract the page fragments from the given amount of the source iterator and
- * build up a second iterator that refers to all of those bits. This allows
- * the original iterator to be disposed of.
+ * Extract virtually contiguous page fragments from the source iterator up to
+ * the given maxima and build bvec queue that refers to all of those bits.
+ * This allows the original iterator to disposed of.
*
- * @extraction_flags can have ITER_ALLOW_P2PDMA set to request peer-to-peer DMA be
- * allowed on the pages extracted.
+ * @extraction_flags can have ITER_ALLOW_P2PDMA set to request peer-to-peer DMA
+ * be allowed on the pages extracted.
*
- * On success, the number of elements in the bvec is returned, the original
- * iterator will have been advanced by the amount extracted.
+ * On success or partial success, the amount of data in the bvec is returned,
+ * the original iterator will have been advanced by the amount extracted.
*
- * The iov_iter_extract_mode() function should be used to query how cleanup
- * should be performed.
+ * If an error occurs and no pages are extracted, an error will be returned and
+ * any allocated bvecq will be freed. If there is no data to be extracted (or
+ * @max_len or @max_pages are zero), a single empty bvecq will be returned.
+ *
+ * The bvecq segments are marked with indications on how to get clean up the
+ * extracted fragments.
*/
-ssize_t netfs_extract_user_iter(struct iov_iter *orig, size_t orig_len,
- struct iov_iter *new,
- iov_iter_extraction_t extraction_flags)
+ssize_t netfs_extract_iter(struct iov_iter *orig, size_t max_len, size_t max_pages,
+ struct bvecq **_bvecq_head,
+ iov_iter_extraction_t extraction_flags, gfp_t gfp)
{
- struct bio_vec *bv = NULL;
- struct page **pages;
- unsigned int cur_npages;
- unsigned int max_pages;
- unsigned int npages = 0;
- unsigned int i;
+ struct bvecq *bq_tail = NULL, *bq;
ssize_t ret = 0;
- size_t count = orig_len, offset, len;
- size_t bv_size, pg_size;
+ size_t extracted = 0;
- if (WARN_ON_ONCE(!iter_is_ubuf(orig) && !iter_is_iovec(orig)))
- return -EIO;
+ _enter("{%u,%zx},%zx", orig->iter_type, orig->count, max_len);
- max_pages = iov_iter_npages(orig, INT_MAX);
- bv_size = array_size(max_pages, sizeof(*bv));
- bv = kvmalloc(bv_size, GFP_KERNEL);
- if (!bv)
- return -ENOMEM;
+ *_bvecq_head = NULL;
+ if (max_len > orig->count)
+ max_len = orig->count;
+ if (!max_len || !max_pages)
+ goto alloc_empty;
+ if (WARN_ON_ONCE(max_pages > INT_MAX))
+ max_pages = INT_MAX; /* Protect iov_iter_npages(). */
- /* Put the page list at the end of the bvec list storage. bvec
- * elements are larger than page pointers, so as long as we work
- * 0->last, we should be fine.
- */
- pg_size = array_size(max_pages, sizeof(*pages));
- pages = (void *)bv + bv_size - pg_size;
+ max_pages = iov_iter_npages(orig, max_pages);
+ if (!max_pages)
+ goto alloc_empty;
- while (count && npages < max_pages) {
- ret = iov_iter_extract_pages(orig, &pages, count,
- max_pages - npages, extraction_flags,
- &offset);
- if (unlikely(ret <= 0)) {
- ret = ret ?: -EIO;
+ do {
+ bq = bvecq_alloc_one(max_pages, gfp, false);
+ if (!bq) {
+ ret = -ENOMEM;
break;
}
+ if (user_backed_iter(orig))
+ bq->mem_type = iov_iter_extract_will_pin(orig) ?
+ BVECQ_MEM_GUP : BVECQ_MEM_PAGECACHE;
- if (WARN(ret > count,
- "%s: extract_pages overrun %zd > %zu bytes\n",
- __func__, ret, count)) {
- ret = -EIO;
- break;
- }
+ if (bq_tail)
+ bvecq_append(bq_tail, bq);
+ else
+ *_bvecq_head = bq;
+ bq_tail = bq;
- cur_npages = DIV_ROUND_UP(offset + ret, PAGE_SIZE);
- if (WARN(cur_npages > max_pages - npages,
- "%s: extract_pages overrun %u > %u pages\n",
- __func__, npages + cur_npages, max_pages)) {
- ret = -EIO;
+ if (max_len == 0)
break;
- }
-
- count -= ret;
- ret += offset;
-
- for (i = 0; i < cur_npages; i++) {
- len = ret > PAGE_SIZE ? PAGE_SIZE : ret;
- bvec_set_page(bv + npages + i, *pages++, len - offset, offset);
- ret -= len;
- offset = 0;
- }
-
- npages += cur_npages;
- }
-
- /* Note: Don't try to clean up after EIO. Either we got no pages, so
- * nothing to clean up, or we got a buffer overrun, memory corruption
- * and can't trust the stuff in the buffer (a WARN was emitted).
- */
-
- if (ret < 0 && (ret == -ENOMEM || npages == 0)) {
- for (i = 0; i < npages; i++)
- unpin_user_page(bv[i].bv_page);
- kvfree(bv);
- return ret;
- }
- iov_iter_bvec(new, orig->data_source, bv, npages, orig_len - count);
- return npages;
-}
-EXPORT_SYMBOL_GPL(netfs_extract_user_iter);
-
-/*
- * Select the span of a bvec iterator we're going to use. Limit it by both maximum
- * size and maximum number of segments. Returns the size of the span in bytes.
- */
-static size_t netfs_limit_bvec(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- const struct bio_vec *bvecs = iter->bvec;
- unsigned int nbv = iter->nr_segs, ix = 0, nsegs = 0;
- size_t len, span = 0, n = iter->count;
- size_t skip = iter->iov_offset + start_offset;
-
- if (WARN_ON(!iov_iter_is_bvec(iter)) ||
- WARN_ON(start_offset > n) ||
- n == 0)
- return 0;
-
- while (n && ix < nbv && skip) {
- len = bvecs[ix].bv_len;
- if (skip < len)
- break;
- skip -= len;
- n -= len;
- ix++;
- }
-
- while (n && ix < nbv) {
- len = min3(n, bvecs[ix].bv_len - skip, max_size);
- span += len;
- nsegs++;
- ix++;
- if (span >= max_size || nsegs >= max_segs)
- break;
- skip = 0;
- n -= len;
- }
-
- return min(span, max_size);
-}
-
-/*
- * Select the span of a kvec iterator we're going to use. Limit it by both
- * maximum size and maximum number of segments. Returns the size of the span
- * in bytes.
- */
-static size_t netfs_limit_kvec(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- const struct kvec *kvecs = iter->kvec;
- unsigned int nkv = iter->nr_segs, ix = 0, nsegs = 0;
- size_t len, span = 0, n = iter->count;
- size_t skip = iter->iov_offset + start_offset;
-
- if (WARN_ON(!iov_iter_is_kvec(iter)) ||
- WARN_ON(start_offset > n) ||
- n == 0)
- return 0;
-
- while (n && ix < nkv && skip) {
- len = kvecs[ix].iov_len;
- if (skip < len)
- break;
- skip -= len;
- n -= len;
- ix++;
- }
-
- while (n && ix < nkv) {
- len = min3(n, kvecs[ix].iov_len - skip, max_size);
- span += len;
- nsegs++;
- ix++;
- if (span >= max_size || nsegs >= max_segs)
- break;
- skip = 0;
- n -= len;
- }
-
- return min(span, max_size);
-}
-
-/*
- * Select the span of an xarray iterator we're going to use. Limit it by both
- * maximum size and maximum number of segments. It is assumed that segments
- * can be larger than a page in size, provided they're physically contiguous.
- * Returns the size of the span in bytes.
- */
-static size_t netfs_limit_xarray(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- struct folio *folio;
- unsigned int nsegs = 0;
- uoff_t pos = iter->xarray_start + iter->iov_offset;
- pgoff_t index = pos / PAGE_SIZE;
- size_t span = 0, n = iter->count;
-
- XA_STATE(xas, iter->xarray, index);
-
- if (WARN_ON(!iov_iter_is_xarray(iter)) ||
- WARN_ON(start_offset > n) ||
- n == 0)
- return 0;
- max_size = min(max_size, n - start_offset);
-
- rcu_read_lock();
- xas_for_each(&xas, folio, ULONG_MAX) {
- size_t offset, flen, len;
- if (xas_retry(&xas, folio))
- continue;
- if (WARN_ON(xa_is_value(folio)))
- break;
- if (WARN_ON(folio_test_hugetlb(folio)))
- break;
-
- flen = folio_size(folio);
- offset = offset_in_folio(folio, pos);
- len = min(max_size, flen - offset);
- span += len;
- nsegs++;
- if (span >= max_size || nsegs >= max_segs)
- break;
- }
-
- rcu_read_unlock();
- return min(span, max_size);
-}
-
-/*
- * Select the span of a bvecq iterator we're going to use. Limit it by both
- * maximum size and maximum number of segments. Returns the size of the span
- * in bytes.
- */
-static size_t netfs_limit_bvecq(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- const struct bvecq *bq = iter->bvecq;
- unsigned int nsegs = 0;
- unsigned int slot = iter->bvecq_slot;
- size_t span = 0, n = iter->count;
-
- if (WARN_ON(!iov_iter_is_bvecq(iter)) ||
- WARN_ON(start_offset > n) ||
- n == 0)
- return 0;
- max_size = umin(max_size, n - start_offset);
-
- if (!bvecq_acquire_slot(bq, slot)) {
- bq = bvecq_next(bq);
- slot = 0;
- }
-
- start_offset += iter->iov_offset;
- do {
- size_t flen;
-
- flen = bq->bv[slot].bv_len;
- if (start_offset < flen) {
- span += flen - start_offset;
- nsegs++;
- start_offset = 0;
- } else {
- start_offset -= flen;
- }
- if (span >= max_size || nsegs >= max_segs)
- break;
-
- slot++;
- if (!bvecq_acquire_slot(bq, slot)) {
- bq = bvecq_next(bq);
- slot = 0;
- }
- } while (bq);
-
- return umin(span, max_size);
-}
+ struct bio_vec *bv = bq->bv;
+ unsigned int slot = 0;
+ do {
+ struct page **pages;
+ ssize_t got;
+ size_t offset;
+ size_t space = bq->max_slots - slot;
+ size_t bv_size = array_size(bq->max_slots, sizeof(*bv));
+ size_t pg_size = array_size(space, sizeof(*pages));
+
+ /* Put the page list at the end of the bvec list
+ * storage. bvec elements are larger than page
+ * pointers, so as long as we work 0->last, we should
+ * be fine.
+ */
+ pages = (void *)bv + bv_size - pg_size;
+
+ got = iov_iter_extract_pages(orig, &pages, max_len,
+ min(space, max_pages),
+ extraction_flags, &offset);
+ if (got < 0) {
+ ret = got;
+ goto out;
+ }
+
+ if (got == 0) {
+ pr_err("extract_pages gave nothing from %zx, %zx\n",
+ extracted, max_len);
+ ret = -EIO;
+ goto out;
+ }
+
+ if (WARN(got > max_len,
+ "%s: extract_pages overrun %zx > %zx bytes\n",
+ __func__, got, max_len)) {
+ ret = -EIO;
+ goto out;
+ }
+
+ extracted += got;
+ max_len -= got;
+
+ do {
+ size_t len = umin(got, PAGE_SIZE - offset);
+
+ BUG_ON(slot >= bq->max_slots);
+
+ bvec_set_page(&bq->bv[slot], *pages++, len, offset);
+ slot++;
+ max_pages--;
+ got -= len;
+ offset = 0;
+ } while (got > 0);
+
+ bvecq_filled_to(bq, slot);
+ } while (max_len > 0 && max_pages > 0 && !bvecq_is_full(bq));
+
+ } while (max_len > 0 && max_pages > 0);
+
+out:
+ if (extracted || ret == 0)
+ return extracted;
+ bvecq_put(*_bvecq_head);
+ *_bvecq_head = NULL;
+ return ret;
+
+alloc_empty:
+ bq = bvecq_alloc_one(1, gfp, false);
+ if (!bq)
+ return -ENOMEM;
+ *_bvecq_head = bq;
+ return 0;
-size_t netfs_limit_iter(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
-{
- if (iov_iter_is_bvecq(iter))
- return netfs_limit_bvecq(iter, start_offset, max_size, max_segs);
- if (iov_iter_is_bvec(iter))
- return netfs_limit_bvec(iter, start_offset, max_size, max_segs);
- if (iov_iter_is_xarray(iter))
- return netfs_limit_xarray(iter, start_offset, max_size, max_segs);
- if (iov_iter_is_kvec(iter))
- return netfs_limit_kvec(iter, start_offset, max_size, max_segs);
- BUG();
}
-EXPORT_SYMBOL(netfs_limit_iter);
+EXPORT_SYMBOL_GPL(netfs_extract_iter);
diff --git a/fs/netfs/misc.c b/fs/netfs/misc.c
index a0cc248a284d5..130bd432b1948 100644
--- a/fs/netfs/misc.c
+++ b/fs/netfs/misc.c
@@ -9,24 +9,6 @@
#include <linux/rmap.h>
#include "internal.h"
-/*
- * Reset the subrequest iterator to refer just to the region remaining to be
- * read. The iterator may or may not have been advanced by socket ops or
- * extraction ops to an extent that may or may not match the amount actually
- * read.
- */
-void netfs_reset_iter(struct netfs_io_subrequest *subreq)
-{
- struct iov_iter *io_iter = &subreq->io_iter;
- size_t remain = subreq->len - subreq->transferred;
-
- if (io_iter->count > remain)
- iov_iter_advance(io_iter, io_iter->count - remain);
- else if (io_iter->count < remain)
- iov_iter_revert(io_iter, remain - io_iter->count);
- iov_iter_truncate(&subreq->io_iter, remain);
-}
-
/**
* netfs_dirty_folio - Mark folio dirty and pin a cache object for writeback
* @mapping: The mapping the folio belongs to.
diff --git a/fs/netfs/objects.c b/fs/netfs/objects.c
index 4b8d20559b0e1..bf17dc31fd9d8 100644
--- a/fs/netfs/objects.c
+++ b/fs/netfs/objects.c
@@ -133,7 +133,6 @@ static void netfs_free_request_rcu(struct rcu_head *rcu)
static void netfs_deinit_request(struct netfs_io_request *rreq)
{
struct netfs_inode *ictx = netfs_inode(rreq->inode);
- unsigned int i;
trace_netfs_rreq(rreq, netfs_rreq_trace_free);
@@ -148,16 +147,10 @@ static void netfs_deinit_request(struct netfs_io_request *rreq)
rreq->netfs_ops->free_request(rreq);
if (rreq->cache_resources.ops)
rreq->cache_resources.ops->end_operation(&rreq->cache_resources);
- if (rreq->direct_bv) {
- for (i = 0; i < rreq->direct_bv_count; i++) {
- if (rreq->direct_bv[i].bv_page) {
- if (rreq->direct_bv_unpin)
- unpin_user_page(rreq->direct_bv[i].bv_page);
- }
- }
- kvfree(rreq->direct_bv);
- }
- rolling_buffer_clear(&rreq->buffer);
+ bvecq_pos_unset(&rreq->load_cursor);
+ bvecq_pos_unset(&rreq->dispatch_cursor);
+ bvecq_pos_unset(&rreq->collect_cursor);
+ bvecq_put(rreq->spare);
if (atomic_dec_and_test(&ictx->io_count))
wake_up_var(&ictx->io_count);
@@ -251,6 +244,7 @@ static void netfs_free_subrequest(struct netfs_io_subrequest *subreq)
trace_netfs_sreq(subreq, netfs_sreq_trace_free);
if (rreq->netfs_ops->free_subrequest)
rreq->netfs_ops->free_subrequest(subreq);
+ bvecq_pos_unset(&subreq->io_buffer);
mempool_free(subreq, rreq->netfs_ops->subrequest_pool ?: &netfs_subrequest_pool);
netfs_stat_d(&netfs_n_rh_sreq);
netfs_put_request(rreq, netfs_rreq_trace_put_subreq);
diff --git a/fs/netfs/read_collect.c b/fs/netfs/read_collect.c
index 75c8874ec4595..368f14cd2d982 100644
--- a/fs/netfs/read_collect.c
+++ b/fs/netfs/read_collect.c
@@ -26,27 +26,35 @@
*/
static void netfs_clear_unread(struct netfs_io_subrequest *subreq)
{
- netfs_reset_iter(subreq);
- WARN_ON_ONCE(subreq->len - subreq->transferred != iov_iter_count(&subreq->io_iter));
- iov_iter_zero(iov_iter_count(&subreq->io_iter), &subreq->io_iter);
+ struct iov_iter iter;
+
+ iov_iter_bvec_queue(&iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
+ iov_iter_advance(&iter, subreq->transferred);
+ iov_iter_zero(subreq->len, &iter);
+
if (subreq->start + subreq->transferred >= subreq->rreq->i_size)
__set_bit(NETFS_SREQ_HIT_EOF, &subreq->flags);
}
static void netfs_clear_unread_dio(struct netfs_io_subrequest *subreq)
{
- uoff_t pos = subreq->start + subreq->transferred;
struct netfs_io_request *rreq = subreq->rreq;
+ struct iov_iter iter;
+ uoff_t pos = subreq->start + subreq->transferred;
size_t fill;
if (pos >= rreq->i_size)
return;
- fill = min_t(uoff_t, rreq->i_size - pos,
- subreq->len - subreq->transferred);
+ fill = umin(rreq->i_size - pos, subreq->len - subreq->transferred);
- netfs_reset_iter(subreq);
- subreq->transferred += iov_iter_zero(fill, &subreq->io_iter);
+ iov_iter_bvec_queue(&iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
+ iov_iter_advance(&iter, subreq->transferred);
+ iov_iter_zero(fill, &iter);
+
+ subreq->transferred += iov_iter_zero(fill, &iter);
}
/*
@@ -138,8 +146,8 @@ static void netfs_unlock_read_folio(struct netfs_io_request *rreq,
*/
void netfs_read_set_unlock_at(struct netfs_io_request *rreq)
{
- const struct bvecq *bq = rreq->buffer.tail;
- unsigned int slot = rreq->buffer.first_tail_slot;
+ const struct bvecq *bq = rreq->collect_cursor.bvecq;
+ unsigned int slot = rreq->collect_cursor.slot;
size_t cleaned_to = rreq->cleaned_to - rreq->start;
size_t progress_at = cleaned_to;
size_t minimum = 256 * 1024;
@@ -169,8 +177,8 @@ void netfs_read_set_unlock_at(struct netfs_io_request *rreq)
static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
unsigned int *notes)
{
- struct bvecq *bq = rreq->buffer.tail;
- unsigned int slot = rreq->buffer.first_tail_slot;
+ struct bvecq *bq = rreq->collect_cursor.bvecq;
+ unsigned int slot = rreq->collect_cursor.slot;
uoff_t collected_to = rreq->collected_to;
if (rreq->cleaned_to >= rreq->collected_to)
@@ -178,15 +186,6 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
// TODO: Begin decryption
- while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&rreq->buffer);
- if (!bq) {
- WRITE_ONCE(rreq->progress_at, rreq->len);
- return;
- }
- slot = 0;
- }
-
/* We have to wait for readahead refs to have been released before we
* can unlock any folios as the ref-dropper walks i_pages and the only
* thing preventing these folios from being removed is the folio lock.
@@ -196,9 +195,24 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
for (;;) {
struct folio *folio;
- uoff_t fpos, fend;
+ uoff_t fpos = rreq->cleaned_to, fend;
size_t fsize;
+ /* Clean up the head bvecq segment. If we clear an entire
+ * segment, then we can get rid of it provided it's not also
+ * the tail segment being filled by the issuer.
+ */
+ if (!bvecq_acquire_slot(bq, slot)) {
+ rreq->collect_cursor.slot = slot;
+ if (!bvecq_delete_spent(&rreq->collect_cursor)) {
+ WRITE_ONCE(rreq->progress_at, rreq->len);
+ return;
+ }
+ bq = rreq->collect_cursor.bvecq;
+ slot = rreq->collect_cursor.slot;
+ continue;
+ }
+
folio = bvec_folio(&bq->bv[slot]);
if (WARN_ONCE(!folio_test_locked(folio),
"R=%08x: folio %lx is not locked\n",
@@ -206,7 +220,6 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
trace_netfs_folio(folio, netfs_folio_trace_not_locked);
fsize = bq->bv[slot].bv_len;
- fpos = folio_pos(folio);
fend = fpos + fsize;
trace_netfs_collect_folio(rreq, folio);
@@ -216,30 +229,16 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
break;
netfs_unlock_read_folio(rreq, bq, slot);
- WRITE_ONCE(rreq->cleaned_to, fpos + fsize);
- *notes |= MADE_PROGRESS;
-
- /* Clean up the head bq. If we clear an entire bq, then
- * we can get rid of it provided it's not also the tail bq
- * being filled by the issuer.
- */
- bq->bv[slot].bv_page = NULL;
slot++;
- while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&rreq->buffer);
- if (!bq)
- goto done;
- slot = 0;
- }
+ WRITE_ONCE(rreq->cleaned_to, fend);
+ *notes |= MADE_PROGRESS;
if (fpos + fsize >= collected_to)
break;
}
- rreq->buffer.tail = bq;
-done:
- rreq->buffer.first_tail_slot = slot;
-
+ bvecq_pos_move(&rreq->collect_cursor, bq);
+ rreq->collect_cursor.slot = slot;
netfs_read_set_unlock_at(rreq);
}
@@ -422,12 +421,15 @@ static void netfs_rreq_assess_dio(struct netfs_io_request *rreq)
if (rreq->origin == NETFS_UNBUFFERED_READ ||
rreq->origin == NETFS_DIO_READ) {
- for (i = 0; i < rreq->direct_bv_count; i++) {
- flush_dcache_page(rreq->direct_bv[i].bv_page);
- // TODO: cifs marks pages in the destination buffer
- // dirty under some circumstances after a read. Do we
- // need to do that too?
- set_page_dirty(rreq->direct_bv[i].bv_page);
+ for (struct bvecq *bq = rreq->collect_cursor.bvecq; bq; bq = bvecq_next(bq)) {
+ unsigned int nr_slots = bvecq_nr_slots_acquire(bq);
+ /* Read the slot count before the slots. */
+
+ /* Mark the target buffers dirty. */
+ for (i = 0; i < nr_slots; i++) {
+ flush_dcache_page(bq->bv[i].bv_page);
+ set_page_dirty(bq->bv[i].bv_page);
+ }
}
}
@@ -521,7 +523,15 @@ bool netfs_read_collection(struct netfs_io_request *rreq)
trace_netfs_rreq(rreq, netfs_rreq_trace_done);
netfs_clear_subrequests(rreq);
- netfs_unlock_abandoned_read_pages(rreq);
+ switch (rreq->origin) {
+ case NETFS_READAHEAD:
+ case NETFS_READPAGE:
+ case NETFS_READ_FOR_WRITE:
+ netfs_unlock_abandoned_read_pages(rreq);
+ break;
+ default:
+ break;
+ }
if (unlikely(rreq->copy_to_cache))
netfs_pgpriv2_end_copy_to_cache(rreq);
return true;
diff --git a/fs/netfs/read_pgpriv2.c b/fs/netfs/read_pgpriv2.c
index f8e5667e278e1..f23c4cbfed581 100644
--- a/fs/netfs/read_pgpriv2.c
+++ b/fs/netfs/read_pgpriv2.c
@@ -19,6 +19,9 @@
static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio *folio)
{
struct netfs_io_stream *cache = &creq->io_streams[1];
+ struct bvecq *queue;
+ unsigned int slot;
+ size_t dio_size = PAGE_SIZE;
size_t fsize = folio_size(folio), flen = fsize;
uoff_t fpos = folio_pos(folio), i_size;
bool to_eof = false;
@@ -48,18 +51,37 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio
to_eof = true;
}
+ flen = round_up(flen, dio_size);
+
_debug("folio %zx %zx", flen, fsize);
trace_netfs_folio(folio, netfs_folio_trace_store_copy);
- /* Attach the folio to the rolling buffer. */
- if (rolling_buffer_append(&creq->buffer, folio, creq->gfp) < 0) {
- set_bit(NETFS_RREQ_CANCEL_CACHING, &creq->flags);
- folio_end_private_2(folio);
- return;
+ /* Institute a new bvec queue segment if the current one is full or if
+ * we encounter a discontiguity. The discontiguity break is important
+ * when it comes to bulk unlocking folios by file range.
+ */
+ queue = creq->load_cursor.bvecq;
+ if (bvecq_is_full(queue) ||
+ (fpos != creq->last_end && creq->last_end > 0 && queue->nr_slots > 0)) {
+ bvecq_buffer_append(&creq->load_cursor, creq->spare);
+ creq->spare = NULL;
+
+ queue = creq->load_cursor.bvecq;
}
- cache->submit_extendable_to = fsize;
+ /* Attach the folio to the rolling buffer. */
+ slot = queue->nr_slots;
+ bvec_set_folio(&queue->bv[slot], folio, fsize, 0);
+ trace_netfs_bv_slot(queue, slot);
+ slot++;
+ bvecq_filled_to(queue, slot);
+ creq->load_cursor.slot = slot;
+ creq->load_cursor.offset = 0;
+ creq->last_end = fpos + flen;
+
+ bvecq_pos_nudge(&creq->dispatch_cursor);
+
cache->submit_off = 0;
cache->submit_len = flen;
@@ -71,10 +93,9 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio
do {
ssize_t part;
- creq->buffer.iter.iov_offset = cache->submit_off;
+ creq->dispatch_cursor.offset = cache->submit_off;
atomic64_set(&creq->issued_to, fpos + cache->submit_off);
- cache->submit_extendable_to = fsize - cache->submit_off;
part = netfs_advance_write(creq, cache, fpos + cache->submit_off,
cache->submit_len, to_eof);
cache->submit_off += part;
@@ -84,8 +105,7 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio
cache->submit_len -= part;
} while (cache->submit_len > 0);
- creq->buffer.iter.iov_offset = 0;
- rolling_buffer_advance(&creq->buffer, fsize);
+ bvecq_pos_step(&creq->dispatch_cursor);
atomic64_set(&creq->issued_to, fpos + fsize);
if (flen < fsize)
@@ -111,6 +131,11 @@ static struct netfs_io_request *netfs_pgpriv2_begin_copy_to_cache(
if (!creq->io_streams[1].avail)
goto cancel_put;
+ if (bvecq_buffer_init(&creq->load_cursor, creq->gfp, false) < 0)
+ goto cancel_put;
+ bvecq_pos_set(&creq->dispatch_cursor, &creq->load_cursor);
+ bvecq_pos_set(&creq->collect_cursor, &creq->dispatch_cursor);
+
__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, &creq->flags);
trace_netfs_copy2cache(rreq, creq);
trace_netfs_write(creq, netfs_write_trace_copy_to_cache);
@@ -143,6 +168,14 @@ void netfs_pgpriv2_copy_to_cache(struct netfs_io_request *rreq, struct folio *fo
return;
}
+ if (!creq->spare) {
+ creq->spare = bvecq_alloc_one(BVECQ_POOL_SLOTS, creq->gfp, false);
+ if (!creq->spare) {
+ set_bit(NETFS_RREQ_CANCEL_CACHING, &creq->flags);
+ return;
+ }
+ }
+
trace_netfs_folio(folio, netfs_folio_trace_pgpriv2_copy);
netfs_pgpriv2_copy_folio(creq, folio);
}
@@ -173,16 +206,18 @@ void netfs_pgpriv2_end_copy_to_cache(struct netfs_io_request *rreq)
*/
bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)
{
- struct bvecq *bq = creq->buffer.tail;
- unsigned int slot = creq->buffer.first_tail_slot;
+ struct bvecq *bq = creq->collect_cursor.bvecq;
+ unsigned int slot;
uoff_t collected_to = creq->collected_to;
bool made_progress = false;
+ slot = creq->collect_cursor.slot;
while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&creq->buffer);
- if (!bq)
- return false;
- slot = 0;
+ creq->collect_cursor.slot = slot;
+ if (!bvecq_delete_spent(&creq->collect_cursor))
+ goto out;
+ bq = creq->collect_cursor.bvecq;
+ slot = creq->collect_cursor.slot;
}
for (;;) {
@@ -213,25 +248,25 @@ bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)
creq->cleaned_to = fpos + fsize;
made_progress = true;
- /* Clean up the head bq. If we clear an entire bq, then
- * we can get rid of it provided it's not also the tail bq
- * being filled by the issuer.
+ /* Clean up the head segment. If we clear an entire segment,
+ * then we can get rid of it provided it's not also the tail
+ * segment being filled by the issuer.
*/
bq->bv[slot].bv_page = NULL;
slot++;
while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&creq->buffer);
- if (!bq)
- goto done;
- slot = 0;
+ creq->collect_cursor.slot = slot;
+ if (!bvecq_delete_spent(&creq->collect_cursor))
+ goto out;
+ bq = creq->collect_cursor.bvecq;
+ slot = creq->collect_cursor.slot;
}
if (fpos + fsize >= collected_to)
break;
}
- creq->buffer.tail = bq;
-done:
- creq->buffer.first_tail_slot = slot;
+ creq->collect_cursor.slot = slot;
+out:
return made_progress;
}
diff --git a/fs/netfs/read_retry.c b/fs/netfs/read_retry.c
index 142c3fb8dab17..7490b9ee1bf7c 100644
--- a/fs/netfs/read_retry.c
+++ b/fs/netfs/read_retry.c
@@ -12,6 +12,10 @@
static void netfs_reissue_read(struct netfs_io_request *rreq,
struct netfs_io_subrequest *subreq)
{
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
+ iov_iter_advance(&subreq->io_iter, subreq->transferred);
+
subreq->error = 0;
__clear_bit(NETFS_SREQ_MADE_PROGRESS, &subreq->flags);
__set_bit(NETFS_SREQ_IN_PROGRESS, &subreq->flags);
@@ -27,6 +31,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
{
struct netfs_io_subrequest *subreq;
struct netfs_io_stream *stream = &rreq->io_streams[0];
+ struct bvecq_pos dispatch_cursor = {};
struct list_head *next;
_enter("R=%x", rreq->debug_id);
@@ -46,9 +51,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
if (test_bit(NETFS_SREQ_FAILED, &subreq->flags))
break;
if (__test_and_clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) {
- __clear_bit(NETFS_SREQ_MADE_PROGRESS, &subreq->flags);
subreq->retry_count++;
- netfs_reset_iter(subreq);
netfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);
netfs_reissue_read(rreq, subreq);
}
@@ -74,11 +77,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
do {
struct netfs_io_subrequest *from, *to, *tmp;
- struct iov_iter source;
uoff_t start, len;
size_t part;
bool boundary = false, subreq_superfluous = false;
+ bvecq_pos_unset(&dispatch_cursor);
+
/* Go through the subreqs and find the next span of contiguous
* buffer that we then rejig (cifs, for example, needs the
* rsize renegotiating) and reissue.
@@ -105,7 +109,8 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
break;
subreq = list_entry(next, struct netfs_io_subrequest, rreq_link);
- if (subreq->start + subreq->transferred != start + len ||
+ if (subreq->start != start + len ||
+ subreq->transferred > 0 ||
test_bit(NETFS_SREQ_BOUNDARY, &subreq->flags) ||
!test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags))
break;
@@ -118,11 +123,14 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
/* Determine the set of buffers we're going to use. Each
* subreq gets a subset of a single overall contiguous buffer.
*/
- netfs_reset_iter(from);
- source = from->io_iter;
- source.count = len;
+ bvecq_pos_transfer(&dispatch_cursor, &from->io_buffer);
+ bvecq_pos_advance(&dispatch_cursor, from->transferred);
+ from->transferred = 0;
- /* Work through the sublist. */
+ /* Work through the sublist. The chain of buffers we're going
+ * to fill is attached to dispatch_cursor and we need to read
+ * 'len' amount of data from 'start'.
+ */
subreq = from;
list_for_each_entry_from(subreq, &stream->subrequests, rreq_link) {
if (!len) {
@@ -130,16 +138,21 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
break;
}
subreq->source = NETFS_DOWNLOAD_FROM_SERVER;
- subreq->start = start - subreq->transferred;
- subreq->len = len + subreq->transferred;
+ subreq->start = start;
+ subreq->len = len;
__clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
__clear_bit(NETFS_SREQ_MADE_PROGRESS, &subreq->flags);
subreq->retry_count++;
+ subreq->transferred = 0;
+
+ bvecq_pos_unset(&subreq->io_buffer);
+ bvecq_pos_set(&subreq->io_buffer, &dispatch_cursor);
trace_netfs_sreq(subreq, netfs_sreq_trace_retry);
/* Renegotiate max_len (rsize) */
- stream->sreq_max_len = subreq->len;
+ stream->sreq_max_len = len;
+ stream->sreq_max_segs = INT_MAX;
if (rreq->netfs_ops->prepare_read &&
rreq->netfs_ops->prepare_read(subreq) < 0) {
trace_netfs_sreq(subreq, netfs_sreq_trace_reprep_failed);
@@ -147,13 +160,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
goto abandon;
}
- part = umin(len, stream->sreq_max_len);
- if (unlikely(stream->sreq_max_segs))
- part = netfs_limit_iter(&source, 0, part, stream->sreq_max_segs);
- subreq->len = subreq->transferred + part;
- subreq->io_iter = source;
- iov_iter_truncate(&subreq->io_iter, part);
- iov_iter_advance(&source, part);
+ part = bvecq_slice(&dispatch_cursor,
+ umin(len, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
+ subreq->len = part;
+
len -= part;
start += part;
if (!len) {
@@ -216,9 +228,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
trace_netfs_sreq(subreq, netfs_sreq_trace_retry);
stream->sreq_max_len = umin(len, rreq->rsize);
- stream->sreq_max_segs = 0;
- if (unlikely(stream->sreq_max_segs))
- part = netfs_limit_iter(&source, 0, part, stream->sreq_max_segs);
+ stream->sreq_max_segs = INT_MAX;
netfs_stat(&netfs_n_rh_download);
if (rreq->netfs_ops->prepare_read(subreq) < 0) {
@@ -227,11 +237,12 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
goto abandon;
}
- part = umin(len, stream->sreq_max_len);
- subreq->len = subreq->transferred + part;
- subreq->io_iter = source;
- iov_iter_truncate(&subreq->io_iter, part);
- iov_iter_advance(&source, part);
+ bvecq_pos_set(&subreq->io_buffer, &dispatch_cursor);
+ part = bvecq_slice(&dispatch_cursor,
+ umin(len, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
+ subreq->len = part;
len -= part;
start += part;
@@ -245,12 +256,14 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
} while (!list_is_head(next, &stream->subrequests));
+out:
+ bvecq_pos_unset(&dispatch_cursor);
return;
/* If we hit an error, fail all remaining incomplete subrequests */
abandon_after:
if (list_is_last(&subreq->rreq_link, &stream->subrequests))
- return;
+ goto out;
subreq = list_next_entry(subreq, rreq_link);
abandon:
list_for_each_entry_from(subreq, &stream->subrequests, rreq_link) {
@@ -261,6 +274,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)
__set_bit(NETFS_SREQ_FAILED, &subreq->flags);
__clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
}
+ goto out;
}
/*
@@ -300,26 +314,24 @@ void netfs_unlock_abandoned_read_pages(struct netfs_io_request *rreq)
if (test_bit(NETFS_RREQ_NEED_PUT_RA_REFS, &rreq->flags))
netfs_wait_for_put_ra_refs(rreq);
- for (p = rreq->buffer.tail; p; p = p->next) {
- for (int slot = rreq->buffer.first_tail_slot;
- bvecq_acquire_slot(p, slot);
- slot++) {
- struct folio *folio;
+ for (p = rreq->collect_cursor.bvecq; p; p = bvecq_next(p)) {
+ unsigned int nr_slots = bvecq_nr_slots_acquire(p);
+ for (int slot = 0; slot < nr_slots; slot++) {
if (!p->bv[slot].bv_page)
continue;
- folio = bvec_folio(&p->bv[slot]);
+ struct folio *folio = bvec_folio(&p->bv[slot]);
+
netfs_cancel_copy_to_cache(rreq, folio);
if (folio == rreq->no_unlock_folio &&
test_bit(NETFS_RREQ_NO_UNLOCK_FOLIO, &rreq->flags)) {
_debug("no unlock");
- } else {
- trace_netfs_folio(folio, netfs_folio_trace_abandon);
- folio_unlock(folio);
+ continue;
}
+ trace_netfs_folio(folio, netfs_folio_trace_abandon);
+ folio_unlock(folio);
}
- rreq->buffer.first_tail_slot = 0;
}
}
diff --git a/fs/netfs/read_single.c b/fs/netfs/read_single.c
index b248e34bd0c86..c70941121de04 100644
--- a/fs/netfs/read_single.c
+++ b/fs/netfs/read_single.c
@@ -101,7 +101,11 @@ static int netfs_single_dispatch_read(struct netfs_io_request *rreq)
subreq->start = 0;
subreq->len = rreq->len;
- subreq->io_iter = rreq->buffer.iter;
+
+ bvecq_pos_set(&subreq->io_buffer, &rreq->dispatch_cursor);
+
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_DEST, subreq->io_buffer.bvecq,
+ subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
netfs_queue_read(rreq, subreq);
@@ -174,6 +178,15 @@ ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_ite
if (IS_ERR(rreq))
return PTR_ERR(rreq);
+ ret = netfs_extract_iter(iter, rreq->len, INT_MAX, &rreq->dispatch_cursor.bvecq,
+ 0, rreq->gfp);
+ if (ret < 0)
+ goto cleanup_free;
+ if (ret < rreq->len) {
+ ret = -EIO;
+ goto cleanup_free;
+ }
+
rreq->progress_at = rreq->len;
ret = netfs_single_begin_cache_read(rreq, ictx);
@@ -183,7 +196,6 @@ ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_ite
netfs_stat(&netfs_n_rh_read_single);
trace_netfs_read(rreq, 0, rreq->len, netfs_read_trace_read_single);
- rreq->buffer.iter = *iter;
netfs_single_dispatch_read(rreq);
ret = netfs_wait_for_read(rreq);
diff --git a/fs/netfs/rolling_buffer.c b/fs/netfs/rolling_buffer.c
deleted file mode 100644
index 66ce9add40122..0000000000000
--- a/fs/netfs/rolling_buffer.c
+++ /dev/null
@@ -1,182 +0,0 @@
-// SPDX-License-Identifier: GPL-2.0-or-later
-/* Rolling buffer helpers
- *
- * Copyright (C) 2024 Red Hat, Inc. All Rights Reserved.
- * Written by David Howells (dhowells@redhat.com)
- */
-
-#include <linux/bitops.h>
-#include <linux/mempool.h>
-#include <linux/pagemap.h>
-#include <linux/rolling_buffer.h>
-#include <linux/slab.h>
-#include "internal.h"
-
-/*
- * Initialise a rolling buffer. We allocate an empty folio queue struct to so
- * that the pointers can be independently driven by the producer and the
- * consumer.
- */
-int rolling_buffer_init(struct rolling_buffer *roll, unsigned int direction,
- gfp_t gfp, bool for_writeback)
-{
- struct bvecq *bq;
-
- roll->for_writeback = for_writeback;
-
- bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, for_writeback);
- if (!bq)
- return -ENOMEM;
-
- roll->head = bq;
- roll->tail = bq;
- iov_iter_bvec_queue(&roll->iter, direction, bq, 0, 0, 0);
- return 0;
-}
-
-/*
- * Add another bvecq to a rolling buffer if there's no space left.
- */
-int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp)
-{
- struct bvecq *bq, *head = roll->head;
-
- if (!bvecq_is_full(head))
- return 0;
-
- bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, roll->for_writeback);
- if (!bq)
- return -ENOMEM;
-
- roll->head = bq;
- if (bvecq_is_full(head)) {
- /* Make sure we don't leave the master iterator pointing to a
- * block that might get immediately consumed.
- */
- if (roll->iter.bvecq == head &&
- roll->iter.bvecq_slot == head->nr_slots) {
- roll->iter.bvecq = bq;
- roll->iter.bvecq_slot = 0;
- }
- }
-
- /* Make sure the initialisation is stored before the next pointer.
- *
- * [!] NOTE: After we set head->next, the consumer is at liberty to
- * immediately delete the old head.
- */
- bvecq_append(head, bq);
- return 0;
-}
-
-/*
- * Decant the entire list of folios to read into a rolling buffer.
- */
-ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,
- struct readahead_control *ractl,
- gfp_t gfp)
-{
- struct bvecq *bq;
- size_t loaded = 0;
-
- while (ractl->_nr_pages - ractl->_batch_count > 0) {
- struct page **pages;
- unsigned int nr;
-
- /* Allocate a bvecq to put some folios into and attach it to
- * the rolling buffer.
- */
- bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, false);
- if (!bq)
- goto nomem_unlock;
- bq->mem_type = BVECQ_MEM_EXTERNAL; /* Folio cleanup handled separately. */
-
- if (!roll->tail)
- roll->tail = bq;
- else
- bvecq_append(roll->head, bq);
- roll->head = bq;
-
- /* Get a bunch of folios and note their sizes. */
- pages = (struct page **)(bq->bv + bq->max_slots);
- pages -= bq->max_slots;
- nr = __readahead_batch(ractl, pages, bq->max_slots);
- if (WARN_ON_ONCE(!nr))
- break;
-
- for (int slot = 0; slot < nr; slot++) {
- struct folio *folio = page_folio(pages[slot]);
- size_t len = folio_size(folio);
-
- bvec_set_folio(&bq->bv[slot], folio, len, 0);
- loaded += len;
- trace_netfs_folio(folio, netfs_folio_trace_read);
- }
-
- bvecq_filled_to(bq, nr);
- }
-
- WRITE_ONCE(roll->iter.count, loaded);
- iov_iter_bvec_queue(&roll->iter, ITER_DEST, roll->tail, 0, 0, loaded);
- return loaded;
-
-nomem_unlock:
- for (bq = roll->tail; bq; bq = bq->next) {
- for (int slot = 0; slot < bq->nr_slots; slot++) {
- struct folio *folio = bvec_folio(&bq->bv[slot]);
-
- folio_unlock(folio);
- folio_put(folio);
- }
- }
- rolling_buffer_clear(roll);
- roll->head = NULL;
- roll->tail = NULL;
- return -ENOMEM;
-}
-
-/*
- * Append a folio to the rolling buffer.
- */
-ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio,
- gfp_t gfp)
-{
- ssize_t size = folio_size(folio);
- int slot;
-
- if (rolling_buffer_make_space(roll, gfp) < 0)
- return -ENOMEM;
-
- slot = roll->head->nr_slots;
- bvec_set_folio(&roll->head->bv[slot], folio, size, 0);
- bvecq_filled_to(roll->head, slot + 1);
-
- WRITE_ONCE(roll->iter.count, roll->iter.count + size);
- return size;
-}
-
-/*
- * Delete a spent buffer from a rolling queue and return the next in line. We
- * don't return the last buffer to keep the pointers independent, but return
- * NULL instead.
- */
-struct bvecq *rolling_buffer_delete_spent(struct rolling_buffer *roll)
-{
- struct bvecq *spent = roll->tail, *next = bvecq_next(spent);
-
- if (!next)
- return NULL;
- next->prev = NULL;
- roll->tail = next;
- spent->next = NULL;
- bvecq_put(spent);
- return next;
-}
-
-/*
- * Clear out a rolling queue.
- */
-void rolling_buffer_clear(struct rolling_buffer *roll)
-{
- bvecq_put(roll->tail);
-}
diff --git a/fs/netfs/write_collect.c b/fs/netfs/write_collect.c
index 91b42820c8925..dcacbab254b92 100644
--- a/fs/netfs/write_collect.c
+++ b/fs/netfs/write_collect.c
@@ -114,12 +114,12 @@ int netfs_folio_written_back(struct folio *folio)
static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
unsigned int *notes)
{
- struct bvecq *bq = wreq->buffer.tail;
- unsigned int slot = wreq->buffer.first_tail_slot;
+ struct bvecq *bq = wreq->collect_cursor.bvecq;
+ unsigned int slot = wreq->collect_cursor.slot;
uoff_t collected_to = wreq->collected_to;
if (WARN_ON_ONCE(!bq)) {
- pr_err("[!] Writeback unlock found empty rolling buffer!\n");
+ pr_err("[!] Writeback unlock found empty buffer!\n");
netfs_dump_request(wreq);
return;
}
@@ -130,19 +130,28 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
return;
}
- while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&wreq->buffer);
- if (!bq)
- return;
- slot = 0;
- }
-
for (;;) {
struct folio *folio;
struct netfs_folio *finfo;
uoff_t fpos, fend;
size_t fsize, flen;
+ /* Try to clean up the head of the queue if it appears to be
+ * used up, but we need to be very careful - the cleanup can
+ * catch the dispatcher, which could lead to us having nothing
+ * left in the queue, causing the front and back pointers to
+ * end up on different tracks. To avoid this, we must always
+ * keep at least one segment in the queue.
+ */
+ if (!bvecq_acquire_slot(bq, slot)) {
+ wreq->collect_cursor.slot = slot;
+ if (!bvecq_delete_spent(&wreq->collect_cursor))
+ return;
+ bq = wreq->collect_cursor.bvecq;
+ slot = wreq->collect_cursor.slot;
+ continue;
+ }
+
folio = bvec_folio(&bq->bv[slot]);
if (WARN_ONCE(!folio_test_writeback(folio),
"R=%08x: folio %lx is not under writeback\n",
@@ -166,26 +175,13 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
wreq->cleaned_to = fpos + fsize;
*notes |= MADE_PROGRESS;
- /* Clean up the head bq. If we clear an entire bq, then
- * we can get rid of it provided it's not also the tail bq
- * being filled by the issuer.
- */
bq->bv[slot].bv_page = NULL;
slot++;
- while (!bvecq_acquire_slot(bq, slot)) {
- bq = rolling_buffer_delete_spent(&wreq->buffer);
- if (!bq)
- goto done;
- slot = 0;
- }
-
if (fpos + fsize >= collected_to)
break;
}
- wreq->buffer.tail = bq;
-done:
- wreq->buffer.first_tail_slot = slot;
+ wreq->collect_cursor.slot = slot;
}
/*
@@ -230,7 +226,8 @@ static void netfs_collect_write_results(struct netfs_io_request *wreq)
trace_netfs_rreq(wreq, netfs_rreq_trace_collect);
reassess_streams:
- issued_to = atomic64_read(&wreq->issued_to);
+ /* Order reading the issued_to point before reading the queue it refers to. */
+ issued_to = atomic64_read_acquire(&wreq->issued_to);
smp_rmb();
collected_to = ULLONG_MAX;
if (wreq->origin == NETFS_WRITEBACK ||
@@ -560,8 +557,12 @@ void netfs_write_subrequest_terminated(void *_op, ssize_t transferred_or_error)
* data is tracked.
*/
netfs_stat(&netfs_n_wh_write_failed);
- if (test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags))
- break;
+ if (test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) {
+ /* We don't retry failed cache writes. */
+ __clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
+ if (!subreq->error)
+ subreq->error = -ENOBUFS;
+ }
trace_netfs_failure(wreq, subreq, transferred_or_error, netfs_fail_write);
__set_bit(NETFS_SREQ_CANCELLED, &subreq->flags);
diff --git a/fs/netfs/write_issue.c b/fs/netfs/write_issue.c
index f0f4786666512..8fab5cf00d7b6 100644
--- a/fs/netfs/write_issue.c
+++ b/fs/netfs/write_issue.c
@@ -107,10 +107,6 @@ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,
ictx = netfs_inode(wreq->inode);
if (is_cacheable)
fscache_begin_write_operation(&wreq->cache_resources, netfs_i_cookie(ictx));
- if (rolling_buffer_init(&wreq->buffer, ITER_SOURCE, wreq->gfp,
- (origin == NETFS_WRITEBACK ||
- origin == NETFS_WRITEBACK_SINGLE)) < 0)
- goto nomem;
wreq->cleaned_to = wreq->start;
if (wreq->cache_resources.dio_size > 1)
@@ -135,9 +131,6 @@ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,
}
return wreq;
-nomem:
- netfs_put_failed_request(wreq);
- return ERR_PTR(-ENOMEM);
}
/**
@@ -163,22 +156,14 @@ void netfs_prepare_write(struct netfs_io_request *wreq,
uoff_t start)
{
struct netfs_io_subrequest *subreq;
- struct iov_iter *wreq_iter = &wreq->buffer.iter;
-
- /* Make sure we don't point the iterator at a used-up bvecq struct
- * being used as a placeholder to prevent the queue from collapsing.
- * In such a case, extend the queue.
- */
- if (iov_iter_is_bvecq(wreq_iter) &&
- !bvecq_acquire_slot(wreq_iter->bvecq, wreq_iter->bvecq_slot))
- rolling_buffer_make_space(&wreq->buffer, wreq->gfp);
subreq = netfs_alloc_subrequest(wreq, stream->source);
if (!subreq)
return;
subreq->start = start;
subreq->stream_nr = stream->stream_nr;
- subreq->io_iter = *wreq_iter;
+
+ bvecq_pos_set(&subreq->io_buffer, &wreq->dispatch_cursor);
_enter("R=%x[%x]", wreq->debug_id, subreq->debug_index);
@@ -259,15 +244,14 @@ static void netfs_do_issue_write(struct netfs_io_stream *stream,
}
void netfs_reissue_write(struct netfs_io_stream *stream,
- struct netfs_io_subrequest *subreq,
- struct iov_iter *source)
+ struct netfs_io_subrequest *subreq)
{
- size_t size = subreq->len - subreq->transferred;
-
// TODO: Use encrypted buffer
- subreq->io_iter = *source;
- iov_iter_advance(source, size);
- iov_iter_truncate(&subreq->io_iter, size);
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_SOURCE,
+ subreq->io_buffer.bvecq, subreq->io_buffer.slot,
+ subreq->io_buffer.offset,
+ subreq->len);
+ iov_iter_advance(&subreq->io_iter, subreq->transferred);
subreq->retry_count++;
subreq->error = 0;
@@ -285,8 +269,12 @@ void netfs_issue_write(struct netfs_io_request *wreq,
if (!subreq)
return;
+ iov_iter_bvec_queue(&subreq->io_iter, ITER_SOURCE,
+ subreq->io_buffer.bvecq, subreq->io_buffer.slot,
+ subreq->io_buffer.offset,
+ subreq->len);
+
stream->construct = NULL;
- subreq->io_iter.count = subreq->len;
netfs_do_issue_write(stream, subreq);
}
@@ -323,7 +311,6 @@ size_t netfs_advance_write(struct netfs_io_request *wreq,
_debug("part %zx/%zx %zx/%zx", subreq->len, stream->sreq_max_len, part, len);
subreq->len += part;
subreq->nr_segs++;
- stream->submit_extendable_to -= part;
if (subreq->len >= stream->sreq_max_len ||
subreq->nr_segs >= stream->sreq_max_segs ||
@@ -347,7 +334,8 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
struct netfs_io_stream *stream;
struct netfs_group *fgroup; /* TODO: Use this with ceph */
struct netfs_folio *finfo;
- size_t iter_off = 0;
+ struct bvecq *queue = wreq->load_cursor.bvecq;
+ unsigned int slot;
size_t fsize = folio_size(folio), flen = fsize, foff = 0;
uoff_t fpos = folio_pos(folio), i_size;
bool to_eof = false, streamw = false;
@@ -355,12 +343,20 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
_enter("");
- if (rolling_buffer_make_space(&wreq->buffer, wreq->gfp) < 0)
- return -ENOMEM;
+ if (!wreq->spare) {
+ wreq->spare = bvecq_alloc_one(BVECQ_POOL_SLOTS, wreq->gfp, true);
+ if (!wreq->spare)
+ return -ENOMEM;
+ }
- /* netfs_perform_write() may shift i_size around the page or from out
- * of the page to beyond it, but cannot move i_size into or through the
- * page since we have it locked.
+ /* netfs_perform_write() may shift i_size around the folio or from out
+ * of the folio to beyond it, but cannot move i_size into or through
+ * the folio since we have it locked.
+ *
+ * Truncate could in theory move i_size into or before the folio, but
+ * it should take steps to prevent writeback from happening
+ * concurrently and should wait for any in-progress writebacks before
+ * proceeding.
*/
i_size = i_size_read(wreq->inode);
@@ -452,8 +448,29 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
trace_netfs_folio(folio, netfs_folio_trace_store_plus);
}
+ /* Institute a new bvec queue segment if the current one is full or if
+ * we encounter a discontiguity. The discontiguity break is important
+ * when it comes to bulk unlocking folios by file range.
+ */
+ if (bvecq_is_full(queue) ||
+ (fpos != wreq->last_end && wreq->last_end > 0)) {
+ bvecq_buffer_append(&wreq->load_cursor, wreq->spare);
+ wreq->spare = NULL;
+
+ queue = wreq->load_cursor.bvecq;
+ bvecq_pos_move(&wreq->dispatch_cursor, queue);
+ wreq->dispatch_cursor.slot = 0;
+ }
+
/* Attach the folio to the rolling buffer. */
- rolling_buffer_append(&wreq->buffer, folio, wreq->gfp);
+ slot = queue->nr_slots;
+ bvec_set_folio(&queue->bv[slot], folio, fsize, 0);
+ trace_netfs_bv_slot(queue, slot);
+ slot++;
+ bvecq_filled_to(queue, slot);
+ wreq->load_cursor.slot = slot;
+ wreq->load_cursor.offset = 0;
+ wreq->last_end = fpos + fsize;
/* Move the submission point forward to allow for write-streaming data
* not starting at the front of the page. We don't do write-streaming
@@ -462,10 +479,19 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
* Also skip uploading for data that's been read and just needs copying
* to the cache.
*/
+ bvecq_pos_nudge(&wreq->dispatch_cursor);
+
for (int s = 0; s < NR_IO_STREAMS; s++) {
+ size_t soff = foff, slen = flen, alignment = 1;
+
stream = &wreq->io_streams[s];
- stream->submit_off = foff;
- stream->submit_len = flen;
+ if (stream->source == NETFS_WRITE_TO_CACHE)
+ alignment = wreq->cache_resources.dio_size;
+ stream = &wreq->io_streams[s];
+ stream->submit_off = round_down(soff, alignment);
+ slen += foff - stream->submit_off;
+ stream->submit_len = round_up(slen, alignment);
+
if (!stream->avail ||
(stream->source == NETFS_WRITE_TO_CACHE && streamw) ||
(stream->source == NETFS_UPLOAD_TO_SERVER &&
@@ -499,14 +525,10 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
break;
stream = &wreq->io_streams[choose_s];
- /* Advance the iterator(s). */
- if (stream->submit_off > iter_off) {
- rolling_buffer_advance(&wreq->buffer, stream->submit_off - iter_off);
- iter_off = stream->submit_off;
- }
+ /* Advance the cursor. */
+ wreq->dispatch_cursor.offset = stream->submit_off;
atomic64_set(&wreq->issued_to, fpos + stream->submit_off);
- stream->submit_extendable_to = fsize - stream->submit_off;
part = netfs_advance_write(wreq, stream, fpos + stream->submit_off,
stream->submit_len, to_eof);
stream->submit_off += part;
@@ -518,9 +540,9 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
debug = true;
}
- if (fsize > iter_off)
- rolling_buffer_advance(&wreq->buffer, fsize - iter_off);
- atomic64_set(&wreq->issued_to, fpos + fsize);
+ bvecq_pos_step(&wreq->dispatch_cursor);
+ /* Order loading the queue before updating the issue_to point */
+ atomic64_set_release(&wreq->issued_to, fpos + fsize);
if (!debug)
kdebug("R=%x: No submit", wreq->debug_id);
@@ -581,6 +603,11 @@ int netfs_writepages(struct address_space *mapping,
goto couldnt_start;
}
+ if (bvecq_buffer_init(&wreq->load_cursor, wreq->gfp, true) < 0)
+ goto nomem;
+ bvecq_pos_set(&wreq->dispatch_cursor, &wreq->load_cursor);
+ bvecq_pos_set(&wreq->collect_cursor, &wreq->dispatch_cursor);
+
__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, &wreq->flags);
trace_netfs_write(wreq, netfs_write_trace_writeback);
netfs_stat(&netfs_n_wh_writepages);
@@ -605,12 +632,17 @@ int netfs_writepages(struct address_space *mapping,
} while ((folio = writeback_iter(mapping, wbc, folio, &error)));
netfs_end_issue_write(wreq);
+ bvecq_pos_unset(&wreq->load_cursor);
+ bvecq_pos_unset(&wreq->dispatch_cursor);
netfs_wake_collector(wreq);
netfs_put_request(wreq, netfs_rreq_trace_put_return);
_leave(" = %d", error);
return error;
+nomem:
+ error = -ENOMEM;
+ netfs_put_failed_request(wreq);
couldnt_start:
if (error == -ENOMEM) {
folio_redirty_for_writepage(wbc, folio);
@@ -631,23 +663,28 @@ EXPORT_SYMBOL(netfs_writepages);
* netfs_writeback_single - Write back a monolithic payload
* @mapping: The mapping to write from
* @wbc: Hints from the VM
- * @iter: Data to write.
+ * @iter: Buffer to write from
+ * @len: Amount to write from buffer
*
* Write a monolithic, non-pagecache object back to the server and/or the
- * cache. The caller must explicitly set NETFS_RREQ_UPLOAD_TO_SERVER when
- * initialising the request if it wants the data to be written to the server
- * (for AFS directories and symlinks, this is not possible; things like mkdir,
- * symlink, rmdir and unlink must be used instead).
+ * cache. There's a maximum of one subrequest per stream. The buffer should be
+ * rounded out sufficiently that it can accommodate cache DIO rounding.
+ *
+ * The caller must explicitly set NETFS_RREQ_UPLOAD_TO_SERVER when initialising
+ * the request if it wants the data to be written to the server (for AFS
+ * directories and symlinks, this is not possible; things like mkdir, symlink,
+ * rmdir and unlink must be used instead).
*
* Return: 0 if successful; 1 if skipped due to lock conflict and WB_SYNC_NONE;
* or a negative error code.
*/
int netfs_writeback_single(struct address_space *mapping,
struct writeback_control *wbc,
- struct iov_iter *iter)
+ struct iov_iter *iter, size_t len)
{
struct netfs_io_request *wreq;
struct netfs_inode *ictx = netfs_inode(mapping->host);
+ size_t clen;
int ret;
if (!netfs_wb_begin(ictx, wbc->sync_mode == WB_SYNC_NONE)) {
@@ -661,10 +698,27 @@ int netfs_writeback_single(struct address_space *mapping,
ret = PTR_ERR(wreq);
goto couldnt_start;
}
+ wreq->len = len;
+ clen = len;
+
+ if (wreq->cache_resources.dio_size > 1) {
+ clen = round_up(len, wreq->cache_resources.dio_size);
+ if (clen > iov_iter_count(iter)) {
+ ret = -EIO;
+ goto cleanup_free;
+ }
+ }
- wreq->buffer.iter = *iter;
- wreq->len = iov_iter_count(iter);
- wreq->submitted = wreq->len;
+ ret = netfs_extract_iter(iter, clen, INT_MAX, &wreq->dispatch_cursor.bvecq,
+ 0, wreq->gfp);
+ if (ret < 0)
+ goto cleanup_free;
+ if (ret < clen) {
+ ret = -EIO;
+ goto cleanup_free;
+ }
+
+ bvecq_pos_set(&wreq->collect_cursor, &wreq->dispatch_cursor);
__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, &wreq->flags);
trace_netfs_write(wreq, netfs_write_trace_writeback_single);
@@ -685,12 +739,14 @@ int netfs_writeback_single(struct address_space *mapping,
subreq = stream->construct;
subreq->len = wreq->len;
+ if (stream->source == NETFS_WRITE_TO_CACHE)
+ subreq->len = clen;
stream->submit_len = subreq->len;
- stream->submit_extendable_to = round_up(wreq->len, PAGE_SIZE);
netfs_issue_write(wreq, stream);
}
+ wreq->submitted = wreq->len;
netfs_all_subreqs_queued(wreq);
netfs_wake_collector(wreq);
@@ -705,6 +761,8 @@ int netfs_writeback_single(struct address_space *mapping,
_leave(" = %d", ret);
return ret;
+cleanup_free:
+ netfs_put_failed_request(wreq);
couldnt_start:
netfs_wb_end(ictx);
_leave(" = %d", ret);
diff --git a/fs/netfs/write_retry.c b/fs/netfs/write_retry.c
index 2f20577563e14..235e75eb10914 100644
--- a/fs/netfs/write_retry.c
+++ b/fs/netfs/write_retry.c
@@ -17,15 +17,17 @@
static void netfs_retry_write_stream(struct netfs_io_request *wreq,
struct netfs_io_stream *stream)
{
+ struct bvecq_pos dispatch_cursor = {};
struct list_head *next;
_enter("R=%x[%x:]", wreq->debug_id, stream->stream_nr);
if (list_empty(&stream->subrequests))
return;
+ if (WARN_ON_ONCE(stream->source != NETFS_UPLOAD_TO_SERVER))
+ return; /* Shouldn't be retrying cache writes. */
- if (stream->source == NETFS_UPLOAD_TO_SERVER &&
- wreq->netfs_ops->retry_request)
+ if (wreq->netfs_ops->retry_request)
wreq->netfs_ops->retry_request(wreq, stream);
if (unlikely(stream->failed))
@@ -39,12 +41,8 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
if (test_bit(NETFS_SREQ_FAILED, &subreq->flags))
break;
if (__test_and_clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) {
- struct iov_iter source;
-
- netfs_reset_iter(subreq);
- source = subreq->io_iter;
netfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);
- netfs_reissue_write(stream, subreq, &source);
+ netfs_reissue_write(stream, subreq);
}
}
return;
@@ -54,11 +52,12 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
do {
struct netfs_io_subrequest *subreq = NULL, *from, *to, *tmp;
- struct iov_iter source;
uoff_t start, len;
size_t part;
bool boundary = false;
+ bvecq_pos_unset(&dispatch_cursor);
+
/* Go through the stream and find the next span of contiguous
* data that we then rejig (cifs, for example, needs the wsize
* renegotiating) and reissue.
@@ -70,7 +69,7 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
if (test_bit(NETFS_SREQ_FAILED, &from->flags) ||
!test_bit(NETFS_SREQ_NEED_RETRY, &from->flags))
- return;
+ goto out;
for (;;) {
/* Read pointer to subreq before reading subreq state. */
@@ -79,7 +78,8 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
break;
subreq = list_entry(next, struct netfs_io_subrequest, rreq_link);
- if (subreq->start + subreq->transferred != start + len ||
+ if (subreq->start != start + len ||
+ subreq->transferred > 0 ||
test_bit(NETFS_SREQ_BOUNDARY, &subreq->flags) ||
!test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags))
break;
@@ -90,11 +90,13 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
/* Determine the set of buffers we're going to use. Each
* subreq gets a subset of a single overall contiguous buffer.
*/
- netfs_reset_iter(from);
- source = from->io_iter;
- source.count = len;
+ bvecq_pos_transfer(&dispatch_cursor, &from->io_buffer);
+ bvecq_pos_advance(&dispatch_cursor, from->transferred);
- /* Work through the sublist. */
+ /* Work through the sublist. The chain of buffers we're going
+ * to fill is attached to dispatch_cursor and we need to read
+ * 'len' amount of data from 'start'.
+ */
subreq = from;
list_for_each_entry_from(subreq, &stream->subrequests, rreq_link) {
if (!len)
@@ -104,16 +106,22 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
subreq->len = len;
__clear_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags);
trace_netfs_sreq(subreq, netfs_sreq_trace_retry);
+ subreq->transferred = 0;
+
+ bvecq_pos_unset(&subreq->io_buffer);
/* Renegotiate max_len (wsize) */
stream->sreq_max_len = len;
+ stream->sreq_max_segs = INT_MAX;
stream->prepare_write(subreq);
- part = umin(len, stream->sreq_max_len);
- if (unlikely(stream->sreq_max_segs))
- part = netfs_limit_iter(&source, 0, part, stream->sreq_max_segs);
+ bvecq_pos_set(&subreq->io_buffer, &dispatch_cursor);
+ part = bvecq_slice(&dispatch_cursor,
+ umin(len, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
subreq->len = part;
- subreq->transferred = 0;
+
len -= part;
start += part;
if (len && subreq == to &&
@@ -121,7 +129,7 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
boundary = true;
netfs_get_subrequest(subreq, netfs_sreq_trace_get_resubmit);
- netfs_reissue_write(stream, subreq, &source);
+ netfs_reissue_write(stream, subreq);
if (subreq == to)
break;
}
@@ -172,17 +180,19 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
netfs_stat(&netfs_n_wh_upload);
stream->sreq_max_len = umin(len, wreq->wsize);
break;
- case NETFS_WRITE_TO_CACHE:
- netfs_stat(&netfs_n_wh_write);
- break;
default:
WARN_ON_ONCE(1);
}
stream->prepare_write(subreq);
- part = umin(len, stream->sreq_max_len);
+ bvecq_pos_set(&subreq->io_buffer, &dispatch_cursor);
+ part = bvecq_slice(&dispatch_cursor,
+ umin(len, stream->sreq_max_len),
+ stream->sreq_max_segs,
+ &subreq->nr_segs);
subreq->len = subreq->transferred + part;
+
len -= part;
start += part;
if (!len && boundary) {
@@ -190,13 +200,16 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq,
boundary = false;
}
- netfs_reissue_write(stream, subreq, &source);
+ netfs_reissue_write(stream, subreq);
if (!len)
break;
} while (len);
} while (!list_is_head(next, &stream->subrequests));
+
+out:
+ bvecq_pos_unset(&dispatch_cursor);
}
/*
diff --git a/include/linux/bvecq.h b/include/linux/bvecq.h
index b984aaa449088..389f6407cb846 100644
--- a/include/linux/bvecq.h
+++ b/include/linux/bvecq.h
@@ -54,6 +54,16 @@ struct bvecq {
/* Number of slots in a 4K bvecq. */
#define BVECQ_4KB_SLOTS ((4096 - sizeof(struct bvecq)) / sizeof(struct bio_vec))
+/*
+ * Position in a bio_vec queue. The bvecq holds a ref on the queue segment it
+ * points to.
+ */
+struct bvecq_pos {
+ struct bvecq *bvecq; /* The first bvecq */
+ unsigned int offset; /* The offset within the starting slot */
+ u16 slot; /* The starting slot */
+};
+
void bvecq_dump(const struct bvecq *bq);
struct bvecq *bvecq_alloc_one(size_t nr_slots, gfp_t gfp, bool for_writeback);
struct bvecq *bvecq_alloc_chain(size_t nr_slots, gfp_t gfp, bool for_writeback);
@@ -61,6 +71,13 @@ struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp
bool for_writeback);
void bvecq_put(struct bvecq *bq);
int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size, gfp_t gfp);
+int bvecq_buffer_init(struct bvecq_pos *pos, gfp_t gfp, bool for_writeback);
+void bvecq_buffer_append(struct bvecq_pos *pos, struct bvecq *bq);
+void bvecq_pos_advance(struct bvecq_pos *pos, size_t amount);
+ssize_t bvecq_zero(struct bvecq_pos *pos, size_t amount);
+size_t bvecq_slice(struct bvecq_pos *pos, size_t max_size,
+ unsigned int max_slots, unsigned int *_nr_slots);
+ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl);
/**
* bvecq_alloc_buffer - Allocate a bvecq chain and populate with buffers
@@ -163,4 +180,171 @@ static inline struct bvecq *bvecq_next(const struct bvecq *bq)
return smp_load_acquire(&bq->next);
}
+/**
+ * bvecq_pos_set - Set one position to be the same as another
+ * @pos: The position object to set
+ * @at: The source position.
+ *
+ * Set @pos to have the same position as @at. This may take a ref on the
+ * bvecq pointed to.
+ */
+static inline void bvecq_pos_set(struct bvecq_pos *pos, const struct bvecq_pos *at)
+{
+ *pos = *at;
+ bvecq_get(pos->bvecq);
+}
+
+/**
+ * bvecq_pos_unset - Unset a position
+ * @pos: The position object to unset
+ *
+ * Unset @pos. This does any needed ref cleanup.
+ */
+static inline void bvecq_pos_unset(struct bvecq_pos *pos)
+{
+ bvecq_put(pos->bvecq);
+ pos->bvecq = NULL;
+ pos->slot = 0;
+ pos->offset = 0;
+}
+
+/**
+ * bvecq_pos_transfer - Transfer one position to another, clearing the first
+ * @pos: The position object to set
+ * @from: The source position to clear.
+ *
+ * Set @pos to have the same position as @from and then clear @from. This may
+ * transfer a ref on the bvecq pointed to.
+ */
+static inline void bvecq_pos_transfer(struct bvecq_pos *pos, struct bvecq_pos *from)
+{
+ *pos = *from;
+ from->bvecq = NULL;
+ from->slot = 0;
+ from->offset = 0;
+}
+
+/**
+ * bvecq_pos_move - Update a position to a new bvecq
+ * @pos: The position object to update.
+ * @to: The new bvecq to point at.
+ *
+ * Update @pos to point to @to if it doesn't already do so. This may
+ * manipulate refs on the bvecqs pointed to.
+ */
+static inline void bvecq_pos_move(struct bvecq_pos *pos, struct bvecq *to)
+{
+ struct bvecq *old = pos->bvecq;
+
+ if (old != to) {
+ pos->bvecq = bvecq_get(to);
+ bvecq_put(old);
+ }
+}
+
+/**
+ * bvecq_pos_nudge - Nudge a position onto the next segment if current used up
+ * @pos: The position object to nudge.
+ *
+ * Update @pos to point to the next segment in the chain if we've used up the
+ * current segment. This may manipulate refs on the bvecqs pointed to.
+ *
+ * Return: true if found a new segment, false if hit the end.
+ */
+static inline bool bvecq_pos_nudge(struct bvecq_pos *pos)
+{
+ struct bvecq *bq = pos->bvecq;
+
+ for (;;) {
+ if (!bvecq_acquire_slot(bq, pos->slot)) {
+ bq = bvecq_next(bq);
+ if (!bq)
+ return false;
+ if (bvecq_acquire_slot(bq, pos->slot))
+ continue; /* More slots got added. */
+ bvecq_pos_move(pos, bq);
+ pos->slot = 0;
+ pos->offset = 0;
+ continue;
+ }
+ if (pos->offset >= bq->bv[pos->slot].bv_len) {
+ pos->slot++;
+ pos->offset = 0;
+ continue;
+ }
+ return true;
+ }
+}
+
+/**
+ * bvecq_pos_step - Step a position to the next slot if possible
+ * @pos: The position object to step.
+ *
+ * Update @pos to point to the next slot in the queue if not at the end. This
+ * may manipulate refs on the bvecqs pointed to.
+ *
+ * Return: true if successful, false if was at the end.
+ */
+static inline bool bvecq_pos_step(struct bvecq_pos *pos)
+{
+ struct bvecq *bq = pos->bvecq, *next;
+
+ pos->slot++;
+ pos->offset = 0;
+ if (bvecq_acquire_slot(bq, pos->slot))
+ return true;
+ next = bvecq_next(bq);
+ if (!next)
+ return false;
+ if (bvecq_acquire_slot(bq, pos->slot))
+ return true;
+ bvecq_pos_move(pos, next);
+ pos->slot = 0;
+ return true;
+}
+
+/**
+ * bvecq_delete_spent - Delete the bvecq at the front if possible
+ * @pos: The position object to update.
+ *
+ * Delete the used up bvecq at the front of the queue that @pos points to if it
+ * is not the last node in the queue; if it is the last node in the queue, it
+ * is kept so that the queue doesn't become detached from the other end. This
+ * may manipulate refs on the bvecqs pointed to. It is also possible that the
+ * producer will fill more slots in the current bvecq.
+ *
+ * Also, we have to be very careful: the consumer can catch the producer, which
+ * could lead to us having nothing left in the queue, causing the front and
+ * back pointers to end up on different tracks. To avoid this, we must always
+ * keep at least one segment in the queue.
+ *
+ * The caller must reload from @pos after calling this.
+ *
+ * Return: true if there's more available; false if not.
+ */
+static inline bool bvecq_delete_spent(struct bvecq_pos *pos)
+{
+ struct bvecq *spent = pos->bvecq;
+ struct bvecq *next;
+ unsigned int slot = pos->slot;
+
+again:
+ /* Read the contents of the queue node after the pointer to it. */
+ next = bvecq_next(spent);
+ if (!next)
+ return false; /* Nothing more to consume at the moment. */
+ if (slot < bvecq_nr_slots_acquire(spent))
+ return true; /* The producer added more. */
+ next->prev = NULL;
+ bvecq_pos_move(pos, next);
+ pos->slot = 0;
+ pos->offset = 0;
+ if (!bvecq_acquire_slot(next, 0)) {
+ spent = next;
+ slot = 0;
+ goto again;
+ }
+ return true;
+}
+
#endif /* _LINUX_BVECQ_H */
diff --git a/include/linux/netfs.h b/include/linux/netfs.h
index b0dd92d12a971..0340c9ee587b6 100644
--- a/include/linux/netfs.h
+++ b/include/linux/netfs.h
@@ -19,10 +19,12 @@
#include <linux/pagemap.h>
#include <linux/bvecq.h>
#include <linux/uio.h>
-#include <linux/rolling_buffer.h>
enum netfs_sreq_ref_trace;
typedef struct mempool mempool_t;
+struct readahead_control;
+struct netfs_io_request;
+struct netfs_io_subrequest;
struct fscache_occupancy;
/**
@@ -144,7 +146,6 @@ struct netfs_io_stream {
unsigned int sreq_max_segs; /* 0 or max number of segments in an iterator */
unsigned int submit_off; /* Folio offset we're submitting from */
unsigned int submit_len; /* Amount of data left to submit */
- unsigned int submit_extendable_to; /* Amount I/O can be rounded up to */
void (*prepare_write)(struct netfs_io_subrequest *subreq);
void (*issue_write)(struct netfs_io_subrequest *subreq);
/* Collection tracking */
@@ -187,6 +188,7 @@ struct netfs_io_subrequest {
struct netfs_io_request *rreq; /* Supervising I/O request */
struct work_struct work;
struct list_head rreq_link; /* Link in rreq->subrequests */
+ struct bvecq_pos io_buffer; /* Bookmark in the combined queue of the start */
struct iov_iter io_iter; /* Iterator for this subrequest */
uoff_t start; /* Where to start the I/O */
size_t len; /* Size of the I/O */
@@ -247,11 +249,14 @@ struct netfs_io_request {
struct netfs_io_stream io_streams[2]; /* Streams of parallel I/O operations */
#define NR_IO_STREAMS 2 //wreq->nr_io_streams
struct netfs_group *group; /* Writeback group being written back */
- struct rolling_buffer buffer; /* Unencrypted buffer */
+ struct bvecq *spare; /* Advance allocation of bvecq */
+ struct bvecq_pos load_cursor; /* Point at which new folios are loaded in */
+ struct bvecq_pos dispatch_cursor; /* Point from which buffers are dispatched */
+ struct bvecq_pos collect_cursor; /* Clear-up point of I/O buffer */
wait_queue_head_t waitq; /* Processor waiter */
void *netfs_priv; /* Private data for the netfs */
void *netfs_priv2; /* Private data for the netfs */
- struct bio_vec *direct_bv; /* DIO buffer list (when handling iovec-iter) */
+ uoff_t last_end; /* End pos of last folio submitted */
uoff_t submitted; /* Amount submitted for I/O so far */
uoff_t len; /* Length of the request */
size_t transferred; /* Amount to be indicated as transferred */
@@ -266,7 +271,6 @@ struct netfs_io_request {
uoff_t abandon_to; /* Position to abandon folios to */
const struct folio *no_unlock_folio; /* Don't unlock this folio after read */
gfp_t gfp; /* GFP flags to use */
- unsigned int direct_bv_count; /* Number of elements in direct_bv[] */
unsigned int debug_id;
unsigned int rsize; /* Maximum read size (0 for none) */
unsigned int wsize; /* Maximum write size (0 for none) */
@@ -274,7 +278,6 @@ struct netfs_io_request {
unsigned int nr_group_rel; /* Number of refs to release on ->group */
spinlock_t lock; /* Lock for queuing subreqs */
enum netfs_io_origin origin; /* Origin of the request */
- bool direct_bv_unpin; /* T if direct_bv[] must be unpinned */
refcount_t ref;
unsigned long flags;
#define NETFS_RREQ_IN_PROGRESS 0 /* Unlocked when the request completes (has ref) */
@@ -427,7 +430,7 @@ void netfs_single_mark_inode_dirty(struct inode *inode);
ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_iter *iter);
int netfs_writeback_single(struct address_space *mapping,
struct writeback_control *wbc,
- struct iov_iter *iter);
+ struct iov_iter *iter, size_t len);
/* Address operations API */
struct readahead_control;
@@ -454,11 +457,9 @@ void netfs_get_subrequest(struct netfs_io_subrequest *subreq,
enum netfs_sreq_ref_trace what);
void netfs_put_subrequest(struct netfs_io_subrequest *subreq,
enum netfs_sreq_ref_trace what);
-ssize_t netfs_extract_user_iter(struct iov_iter *orig, size_t orig_len,
- struct iov_iter *new,
- iov_iter_extraction_t extraction_flags);
-size_t netfs_limit_iter(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs);
+ssize_t netfs_extract_iter(struct iov_iter *orig, size_t max_len, size_t max_pages,
+ struct bvecq **_bvecq_head,
+ iov_iter_extraction_t extraction_flags, gfp_t gfp);
void netfs_prepare_write_failed(struct netfs_io_subrequest *subreq);
void netfs_write_subrequest_terminated(void *_op, ssize_t transferred_or_error);
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 0adfa6605653d..b7154b81d0d7c 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1411,6 +1411,7 @@ struct readahead_control {
struct file_ra_state *ra;
/* private: use the readahead_* accessors instead */
pgoff_t _index;
+ unsigned int _nr_folios;
unsigned int _nr_pages;
unsigned int _batch_count;
bool dropbehind;
@@ -1590,6 +1591,15 @@ static inline size_t readahead_batch_length(const struct readahead_control *rac)
return rac->_batch_count * PAGE_SIZE;
}
+/**
+ * readahead_folio_count - Get the number of folios in this readahead request.
+ * @rac: The readahead request.
+ */
+static inline unsigned int readahead_folio_count(const struct readahead_control *rac)
+{
+ return rac->_nr_folios;
+}
+
static inline unsigned long dir_pages(const struct inode *inode)
{
return (unsigned long)(inode->i_size + PAGE_SIZE - 1) >>
diff --git a/include/linux/rolling_buffer.h b/include/linux/rolling_buffer.h
deleted file mode 100644
index 5c0bc4221f010..0000000000000
--- a/include/linux/rolling_buffer.h
+++ /dev/null
@@ -1,47 +0,0 @@
-/* SPDX-License-Identifier: GPL-2.0-or-later */
-/* Rolling buffer of folios
- *
- * Copyright (C) 2024 Red Hat, Inc. All Rights Reserved.
- * Written by David Howells (dhowells@redhat.com)
- */
-
-#ifndef _ROLLING_BUFFER_H
-#define _ROLLING_BUFFER_H
-
-#include <linux/bvecq.h>
-#include <linux/uio.h>
-
-/*
- * Rolling buffer. Whilst the buffer is live and in use, folios and bvecq
- * segments can be added to one end by one thread and removed from the other
- * end by another thread. The buffer isn't allowed to be empty; it must always
- * have at least one bvecq in it so that neither side has to modify both queue
- * pointers.
- *
- * The iterator in the buffer is extended as buffers are inserted. It can be
- * snapshotted to use a segment of the buffer.
- */
-struct rolling_buffer {
- struct bvecq *head; /* Producer's insertion point */
- struct bvecq *tail; /* Consumer's removal point */
- struct iov_iter iter; /* Iterator tracking what's left in the buffer */
- u8 first_tail_slot; /* First slot in ->tail */
- bool for_writeback; /* T if being used for writeback */
-};
-
-int rolling_buffer_init(struct rolling_buffer *roll, unsigned int direction,
- gfp_t gfp, bool for_writeback);
-int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp);
-ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,
- struct readahead_control *ractl,
- gfp_t gfp);
-ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio, gfp_t gfp);
-struct bvecq *rolling_buffer_delete_spent(struct rolling_buffer *roll);
-void rolling_buffer_clear(struct rolling_buffer *roll);
-
-static inline void rolling_buffer_advance(struct rolling_buffer *roll, size_t amount)
-{
- iov_iter_advance(&roll->iter, amount);
-}
-
-#endif /* _ROLLING_BUFFER_H */
diff --git a/include/trace/events/netfs.h b/include/trace/events/netfs.h
index e80c27c49ce25..81a3b7aaa187c 100644
--- a/include/trace/events/netfs.h
+++ b/include/trace/events/netfs.h
@@ -229,7 +229,9 @@
EM(netfs_folio_trace_sched_copy, "sched-copy") \
EM(netfs_folio_trace_store, "store") \
EM(netfs_folio_trace_store_copy, "store-copy") \
- E_(netfs_folio_trace_store_plus, "store+")
+ EM(netfs_folio_trace_store_plus, "store+") \
+ EM(netfs_folio_trace_zero, "zero") \
+ E_(netfs_folio_trace_zero_ra, "zero-ra")
#define netfs_collect_contig_traces \
EM(netfs_contig_trace_collect, "Collect") \
@@ -386,10 +388,10 @@ TRACE_EVENT(netfs_sreq,
__entry->len = sreq->len;
__entry->transferred = sreq->transferred;
__entry->start = sreq->start;
- __entry->slot = sreq->io_iter.bvecq_slot;
+ __entry->slot = sreq->io_buffer.slot;
),
- TP_printk("R=%08x[%x] %s %s f=%03x s=%llx %zx/%zx s=%u e=%d",
+ TP_printk("R=%08x[%x] %s %s f=%03x s=%llx %zx/%zx bv=%u e=%d",
__entry->rreq, __entry->index,
__print_symbolic(__entry->source, netfs_sreq_sources),
__print_symbolic(__entry->what, netfs_sreq_traces),
@@ -782,6 +784,30 @@ TRACE_EVENT(netfs_read_progress_at,
__entry->rreq, __entry->cleaned_to, __entry->progress_at)
);
+TRACE_EVENT(netfs_bv_slot,
+ TP_PROTO(const struct bvecq *bq, int slot),
+
+ TP_ARGS(bq, slot),
+
+ TP_STRUCT__entry(
+ __field(unsigned long, pfn)
+ __field(unsigned int, offset)
+ __field(unsigned int, len)
+ __field(unsigned int, slot)
+ ),
+
+ TP_fast_assign(
+ __entry->slot = slot;
+ __entry->pfn = page_to_pfn(bq->bv[slot].bv_page);
+ __entry->offset = bq->bv[slot].bv_offset;
+ __entry->len = bq->bv[slot].bv_len;
+ ),
+
+ TP_printk("bq[%x] p=%lx %x-%x",
+ __entry->slot,
+ __entry->pfn, __entry->offset, __entry->offset + __entry->len)
+ );
+
#undef EM
#undef E_
#endif /* _TRACE_NETFS_H */
diff --git a/mm/readahead.c b/mm/readahead.c
index 6e5563290287e..196542fb1c930 100644
--- a/mm/readahead.c
+++ b/mm/readahead.c
@@ -188,6 +188,7 @@ static void read_pages(struct readahead_control *rac)
if (unlikely(rac->_workingset))
psi_memstall_leave(&rac->_pflags);
rac->_workingset = false;
+ rac->_nr_folios = 0;
BUG_ON(readahead_count(rac));
}
@@ -303,6 +304,7 @@ void page_cache_ra_unbounded(struct readahead_control *ractl,
if (i == mark)
folio_set_readahead(folio);
ractl->_workingset |= folio_test_workingset(folio);
+ ractl->_nr_folios++;
ractl->_nr_pages += min_nrpages;
i += min_nrpages;
}
@@ -473,6 +475,7 @@ static inline int ra_alloc_folio(struct readahead_control *ractl, pgoff_t index,
return err;
}
+ ractl->_nr_folios++;
ractl->_nr_pages += 1UL << order;
ractl->_workingset |= folio_test_workingset(folio);
return 0;
@@ -822,6 +825,7 @@ void readahead_expand(struct readahead_control *ractl,
ractl->_workingset = true;
psi_memstall_enter(&ractl->_pflags);
}
+ ractl->_nr_folios++;
ractl->_nr_pages += min_nrpages;
ractl->_index = folio->index;
}
@@ -851,6 +855,7 @@ void readahead_expand(struct readahead_control *ractl,
ractl->_workingset = true;
psi_memstall_enter(&ractl->_pflags);
}
+ ractl->_nr_folios++;
ractl->_nr_pages += min_nrpages;
if (ra) {
ra->size += min_nrpages;
|
| 25/3 | 2026/10/06 21:25 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 47064 output: 23 thoughts: 2257 Thoughts: **KMSAN Evaluation: Patch Series for netfs/afs/cachefiles** Okay, let's break this down systematically. The objective is to determine if this patch series warrants a dedicated KMSAN fuzzing session. My understanding is that KMSAN's focus is on detecting reads of uninitialized memory and information leaks, *not* the typical memory corruption issues KASAN handles. First, I need to make a crucial distinction: is this patch introducing or exposing uninitialized memory that *only* KMSAN would catch? Standard KASAN, LOCKDEP, and other bug detectors already handle things like out-of-bounds accesses, use-after-free, etc. I need to focus on uninitialized memory reads and kernel-to-user info leaks. The rules are pretty clear. I need to say "NeedsKMSAN=true" only if: * Kernel structures are sent to userspace with the potential for uninitialized fields. * Conditional logic relies on uninitialized variables. * Complex data structures have uninitialized fields that *could* be read. * Bounds checks allow access to uninitialized bytes of existing buffers. Otherwise, "NeedsKMSAN=false" if: * The risk is primarily KASAN-type bugs (OOB, UAF, etc.). * All touched structures are zeroed/initialized before use. * No new uninitialized memory risks are introduced. Now, let's dig into the patch diff. It appears this series is largely refactoring rolling buffers with bvecq and cursors, especially around `netfs` and `afs` file systems. * `fs/afs/dir.c` and `fs/afs/symlink.c`: The code is being modified to use `dvnode->directory_size` instead of the i-node's size, and PAGE\_SIZE in one instance. It looks like some size calculations are different. This seems fine, no immediately obvious uninitialized memory risks here. * `fs/cachefiles/io.c`: Minor code cleanup; not a concern. * `fs/netfs/Makefile`: Removing `rolling_buffer.o` - expected with a refactor. * `fs/netfs/buffered_read.c`: Replacing `rolling_buffer` with `bvecq_pos`, etc. Crucially, there's a change to how `subreq->io_iter` is handled after a read. It now uses `bvecq_zero` which is good. The code now zeros `subreq->io_buffer` (using `bvecq_zero`) or uses `folio_zero_segments`. This is good. It replaces iov iter zeros with direct bvec zeroing. All the memory is zeroed, which is a good sign. * `fs/netfs/bvecq.c`: The introduction of `bvecq_zero` and associated functions is where the action is. Looks like they zero the memory pointed to by bio\_vec slots. Okay, this is core to the bvecq functionality, and they're explicitly zeroing things, which reduces the KMSAN risk. * `fs/netfs/direct_read.c` & `fs/netfs/direct_write.c`: Using `netfs_extract_iter` and `bvecq_slice` instead of potentially problematic `netfs_extract_user_iter`. Good: the new `netfs_extract_iter` uses `bvecq`. The code appears to be allocating and managing memory within the kernel, and the `netfs_extract_iter` API appears to have been improved to cover both user and kernel iterators, extracting data into a `bvecq` chain, so far so good. * `fs/netfs/iterator.c`: More of the same: replacing `netfs_extract_user_iter` with `netfs_extract_iter` This is more refactoring. * `fs/netfs/read_collect.c`: Now, here’s where my attention spikes. Inside `netfs_clear_unread`, there's a suspicious line: `iov_iter_zero(subreq->len, &iter);` After tracing the path of iov_iter through `iov_iter_bvec_queue`, `iov_iter_advance`, and `iov_iter_zero`, it seems the count given to `iov_iter_zero` is `subreq->len`, which is correct. The `iov_iter` will contain the remaining byte count after the advance. This is NOT a KMSAN bug but a bug in calculation! The developer is trying to zero some memory, but is mis-calculating what needs to be |
| 26/3 | 2026/10/06 21:25 | tool | grepper |
0mArgs:{
"Expression": "size_t iov_iter_zero"
}Results: {
"Output": "include/linux/uio.h=282=size_t _copy_mc_to_iter(const void *addr, size_t bytes, struct iov_iter *i);\n--\ninclude/linux/uio.h-286-\ninclude/linux/uio.h:287:size_t iov_iter_zero(size_t bytes, struct iov_iter *);\ninclude/linux/uio.h-288-unsigned long iov_iter_alignment(const struct iov_iter *i);\n--\nlib/iov_iter.c=458=size_t zero_to_iter(void *iter_to, size_t progress,\n--\nlib/iov_iter.c-464-\nlib/iov_iter.c:465:size_t iov_iter_zero(size_t bytes, struct iov_iter *i)\nlib/iov_iter.c-466-{\n"
} |
| 27/3 | 2026/10/06 21:25 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 50607 output: 35 thoughts: 19 Thoughts: Here's my take on those thoughts: **Analyzing `lib/iov_iter.c` Line 465** Okay, I need to zoom in on `lib/iov_iter.c`, specifically around line 465. That's usually a good starting point for a deep dive in this part of the code. I'm probably looking for a specific issue or optimization opportunity related to the `iov_iter` structure and its associated functions. Knowing the kernel's architecture, line 465 could be related to buffer management, data copying, or perhaps even an optimization within the iteration loop itself. I'll need to refresh my memory on what's handled near that line. It might involve scatter/gather operations or direct access techniques, but whatever the case, I want to see *exactly* how the iterators are being advanced or data is being prepared or consumed at that point. I’ll be looking for any potential bottlenecks or areas for improvement in the handling of I/O vectors. |
| 28/3 | 2026/10/06 21:26 | tool | read-file |
0mArgs:{
"File": "lib/iov_iter.c",
"FirstLine": 464,
"LineCount": 20
}Results: {
"Contents": " 464:\t\n 465:\tsize_t iov_iter_zero(size_t bytes, struct iov_iter *i)\n 466:\t{\n 467:\t\treturn iterate_and_advance(i, bytes, NULL,\n 468:\t\t\t\t\t zero_to_user_iter, zero_to_iter);\n 469:\t}\n 470:\tEXPORT_SYMBOL(iov_iter_zero);\n 471:\t\n 472:\tsize_t copy_folio_from_iter_atomic(struct folio *folio, size_t offset,\n 473:\t\t\tsize_t bytes, struct iov_iter *i)\n 474:\t{\n 475:\t\tsize_t n, copied = 0;\n 476:\t\n 477:\t\tif (!page_copy_sane(\u0026folio-\u003epage, offset, bytes))\n 478:\t\t\treturn 0;\n 479:\t\tif (WARN_ON_ONCE(!i-\u003edata_source))\n 480:\t\t\treturn 0;\n 481:\t\n 482:\t\tdo {\n 483:\t\t\tchar *to = kmap_local_folio(folio, offset);\n"
} |
| 29/3 | 2026/10/06 21:26 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 51194 output: 21 thoughts: 41 Thoughts: Okay, here's my take on those thoughts, framed as a first-person expert summary: **Analyzing `iterate_and_advance` and Byte Bounds** Alright, so I'm looking at this `iterate_and_advance(i, bytes, ...)` function. Immediately, my mind flags the crucial interaction between `bytes` and `i->count`. The implication is that `bytes` is, in effect, range-checked or capped by the value of `i->count`. To fully grasp the implications of this, I need a quick refresher on how this function is defined and how `i` is being used. I need to understand the precise mechanism of this boundary. I need to make sure I am not missing any subtle interactions between these two variables. I should consult the definition and verify that my assumptions on the bounds check are accurate. This is basic, but it's important to be methodical about these checks to keep from running into a logic error further down the line. I always try to double check. |
| 30/3 | 2026/10/06 21:26 | tool | grepper |
0mError:no matches Args: {
"Expression": "#define iterate_and_advance"
}Results: {
"error": "no matches"
} |
| 31/3 | 2026/10/06 21:26 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 51482 output: 33 thoughts: 21 Thoughts: Here's my thought process, as requested: **Navigating the Codebase** Okay, so I need to get to the `iterate_and_advance` function. I vaguely remember that being a core piece related to I/O vector iteration, so my first instinct is to look within the relevant kernel module. Given that I'm working with I/O vectors, `lib/iov_iter.c` feels like the most logical place to begin. It's a fairly standard naming convention, and it's where I'd *expect* to find a function dealing with that kind of functionality. Let's start there. |
| 32/3 | 2026/10/06 21:26 | tool | grepper |
0mArgs:{
"Expression": "iterate_and_advance",
"PathPrefixes": [
"lib/iov_iter.c"
]
}Results: {
"Output": "lib/iov_iter.c=191=size_t _copy_to_iter(const void *addr, size_t bytes, struct iov_iter *i)\n--\nlib/iov_iter.c-196-\t\tmight_fault();\nlib/iov_iter.c:197:\treturn iterate_and_advance(i, bytes, (void *)addr,\nlib/iov_iter.c-198-\t\t\t\t copy_to_user_iter, memcpy_to_iter);\n--\nlib/iov_iter.c=246=size_t _copy_mc_to_iter(const void *addr, size_t bytes, struct iov_iter *i)\n--\nlib/iov_iter.c-251-\t\tmight_fault();\nlib/iov_iter.c:252:\treturn iterate_and_advance(i, bytes, (void *)addr,\nlib/iov_iter.c-253-\t\t\t\t copy_to_user_iter_mc, memcpy_to_iter_mc);\n--\nlib/iov_iter.c=259=size_t __copy_from_iter(void *addr, size_t bytes, struct iov_iter *i)\nlib/iov_iter.c-260-{\nlib/iov_iter.c:261:\treturn iterate_and_advance(i, bytes, addr,\nlib/iov_iter.c-262-\t\t\t\t copy_from_user_iter, memcpy_from_iter);\n--\nlib/iov_iter.c=283=size_t _copy_from_iter_nocache(void *addr, size_t bytes, struct iov_iter *i)\n--\nlib/iov_iter.c-287-\nlib/iov_iter.c:288:\treturn iterate_and_advance(i, bytes, addr,\nlib/iov_iter.c-289-\t\t\t\t copy_from_user_iter_nocache,\n--\nlib/iov_iter.c=326=size_t _copy_from_iter_flushcache(void *addr, size_t bytes, struct iov_iter *i)\n--\nlib/iov_iter.c-330-\nlib/iov_iter.c:331:\treturn iterate_and_advance(i, bytes, addr,\nlib/iov_iter.c-332-\t\t\t\t copy_from_user_iter_flushcache,\n--\nlib/iov_iter.c=390=size_t copy_page_to_iter_nofault(struct page *page, unsigned offset, size_t bytes,\n--\nlib/iov_iter.c-404-\nlib/iov_iter.c:405:\t\tn = iterate_and_advance(i, n, kaddr + offset,\nlib/iov_iter.c-406-\t\t\t\t\tcopy_to_user_iter_nofault,\n--\nlib/iov_iter.c=465=size_t iov_iter_zero(size_t bytes, struct iov_iter *i)\nlib/iov_iter.c-466-{\nlib/iov_iter.c:467:\treturn iterate_and_advance(i, bytes, NULL,\nlib/iov_iter.c-468-\t\t\t\t zero_to_user_iter, zero_to_iter);\n"
} |
| 33/3 | 2026/10/06 21:26 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 52487 output: 34 thoughts: 24 Thoughts: Here's my thought process, as requested: **Navigating the Codebase** Okay, so I need to understand how `iterate_and_advance` functions. Since I'm working with the kernel's I/O vector iterator implementation, my first stop has to be the source. I know that `lib/iov_iter.c` is the likely location. Let's start there. I'll need to figure out exactly how the function is used in this file and if there are calls to it from any other files. Now, let's go. |
| 34/3 | 2026/10/06 21:26 | tool | read-file |
0mArgs:{
"File": "lib/iov_iter.c",
"FirstLine": 15,
"LineCount": 50
}Results: {
"Contents": " 15:\t\n 16:\tstatic __always_inline\n 17:\tsize_t copy_to_user_iter(void __user *iter_to, size_t progress,\n 18:\t\t\t\t size_t len, void *from, void *priv2)\n 19:\t{\n 20:\t\tif (should_fail_usercopy())\n 21:\t\t\treturn len;\n 22:\t\tif (access_ok(iter_to, len)) {\n 23:\t\t\tfrom += progress;\n 24:\t\t\tinstrument_copy_to_user(iter_to, from, len);\n 25:\t\t\tlen = raw_copy_to_user(iter_to, from, len);\n 26:\t\t}\n 27:\t\treturn len;\n 28:\t}\n 29:\t\n 30:\tstatic __always_inline\n 31:\tsize_t copy_to_user_iter_nofault(void __user *iter_to, size_t progress,\n 32:\t\t\t\t\t size_t len, void *from, void *priv2)\n 33:\t{\n 34:\t\tssize_t res;\n 35:\t\n 36:\t\tif (should_fail_usercopy())\n 37:\t\t\treturn len;\n 38:\t\n 39:\t\tfrom += progress;\n 40:\t\tres = copy_to_user_nofault(iter_to, from, len);\n 41:\t\treturn res \u003c 0 ? len : res;\n 42:\t}\n 43:\t\n 44:\tstatic __always_inline\n 45:\tsize_t copy_from_user_iter(void __user *iter_from, size_t progress,\n 46:\t\t\t\t size_t len, void *to, void *priv2)\n 47:\t{\n 48:\t\tsize_t res = len;\n 49:\t\n 50:\t\tif (should_fail_usercopy())\n 51:\t\t\treturn len;\n 52:\t\tif (can_do_masked_user_access()) {\n 53:\t\t\titer_from = mask_user_address(iter_from);\n 54:\t\t} else {\n 55:\t\t\tif (!access_ok(iter_from, len))\n 56:\t\t\t\treturn res;\n 57:\t\n 58:\t\t\t/*\n 59:\t\t\t * Ensure that bad access_ok() speculation will not\n 60:\t\t\t * lead to nasty side effects *after* the copy is\n 61:\t\t\t * finished:\n 62:\t\t\t */\n 63:\t\t\tbarrier_nospec();\n 64:\t\t}\n"
} |
| 35/3 | 2026/10/06 21:26 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 53333 output: 20 thoughts: 24 Thoughts: Okay, let's break this down. **Locating the Definition** My initial thought is to pinpoint exactly where the `iterate_and_advance` function is defined. Given its name, it's highly probable that it's nestled within the kernel's I/O vector iteration code. Therefore, my immediate inclination is to check either `iov_iter.c` (likely containing the core logic) or relevant header files, potentially including declarations and supporting structures. That's the first place I'd look. I'm operating on the assumption that it's part of a standard utility for moving data through iovecs or similar structures, but I need to confirm its exact location for analysis. |
| 36/3 | 2026/10/06 21:26 | tool | grepper |
0mArgs:{
"Expression": "iterate_and_advance\\("
}Results: {
"Output": "include/linux/iov_iter.h=335=static __always_inline\ninclude/linux/iov_iter.h:336:size_t iterate_and_advance(struct iov_iter *iter, size_t len, void *priv,\ninclude/linux/iov_iter.h-337-\t\t\t iov_ustep_f ustep, iov_step_f step)\n--\nlib/iov_iter.c=191=size_t _copy_to_iter(const void *addr, size_t bytes, struct iov_iter *i)\n--\nlib/iov_iter.c-196-\t\tmight_fault();\nlib/iov_iter.c:197:\treturn iterate_and_advance(i, bytes, (void *)addr,\nlib/iov_iter.c-198-\t\t\t\t copy_to_user_iter, memcpy_to_iter);\n--\nlib/iov_iter.c=246=size_t _copy_mc_to_iter(const void *addr, size_t bytes, struct iov_iter *i)\n--\nlib/iov_iter.c-251-\t\tmight_fault();\nlib/iov_iter.c:252:\treturn iterate_and_advance(i, bytes, (void *)addr,\nlib/iov_iter.c-253-\t\t\t\t copy_to_user_iter_mc, memcpy_to_iter_mc);\n--\nlib/iov_iter.c=259=size_t __copy_from_iter(void *addr, size_t bytes, struct iov_iter *i)\nlib/iov_iter.c-260-{\nlib/iov_iter.c:261:\treturn iterate_and_advance(i, bytes, addr,\nlib/iov_iter.c-262-\t\t\t\t copy_from_user_iter, memcpy_from_iter);\n--\nlib/iov_iter.c=283=size_t _copy_from_iter_nocache(void *addr, size_t bytes, struct iov_iter *i)\n--\nlib/iov_iter.c-287-\nlib/iov_iter.c:288:\treturn iterate_and_advance(i, bytes, addr,\nlib/iov_iter.c-289-\t\t\t\t copy_from_user_iter_nocache,\n--\nlib/iov_iter.c=326=size_t _copy_from_iter_flushcache(void *addr, size_t bytes, struct iov_iter *i)\n--\nlib/iov_iter.c-330-\nlib/iov_iter.c:331:\treturn iterate_and_advance(i, bytes, addr,\nlib/iov_iter.c-332-\t\t\t\t copy_from_user_iter_flushcache,\n--\nlib/iov_iter.c=390=size_t copy_page_to_iter_nofault(struct page *page, unsigned offset, size_t bytes,\n--\nlib/iov_iter.c-404-\nlib/iov_iter.c:405:\t\tn = iterate_and_advance(i, n, kaddr + offset,\nlib/iov_iter.c-406-\t\t\t\t\tcopy_to_user_iter_nofault,\n--\nlib/iov_iter.c=465=size_t iov_iter_zero(size_t bytes, struct iov_iter *i)\nlib/iov_iter.c-466-{\nlib/iov_iter.c:467:\treturn iterate_and_advance(i, bytes, NULL,\nlib/iov_iter.c-468-\t\t\t\t zero_to_user_iter, zero_to_iter);\n"
} |
| 37/3 | 2026/10/06 21:26 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 54439 output: 37 thoughts: 19 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Diving into `iov_iter.h:335`** Alright, time to dive into that specific line in `include/linux/iov_iter.h`. My immediate thought is, "What's the context? What's the function or macro defined there?" This usually points to a data structure definition or potentially a core function related to scatter-gather I/O, considering it's in `iov_iter.h`. I'm mentally parsing the likely structure layout... Is it an `iovec`-like struct? Is it managing offsets, lengths, and flags for individual memory regions? I'm starting to anticipate common pitfalls – potential issues with boundary checks, off-by-one errors in iterating through iovecs, or complexities in handling different address spaces (user vs. kernel). This line's likely going to be key to how they're handling data transfer and buffer management. I'm already formulating some questions in my head: What are the locking implications? Are they using atomic operations? Is it thread-safe? Any potential for race conditions? I'm bracing myself for potentially intricate pointer arithmetic and clever use of bitfields to pack information. Hopefully, it's well-documented, because, without that, it might take a while to follow the logic. But, I'm ready to dig in and understand exactly what this line, 335, is doing in this critical subsystem. Let's see what gems are hidden there! |
| 38/3 | 2026/10/06 21:26 | tool | read-file |
0mArgs:{
"File": "include/linux/iov_iter.h",
"FirstLine": 335,
"LineCount": 40
}Results: {
"Contents": " 335:\tstatic __always_inline\n 336:\tsize_t iterate_and_advance(struct iov_iter *iter, size_t len, void *priv,\n 337:\t\t\t\t iov_ustep_f ustep, iov_step_f step)\n 338:\t{\n 339:\t\treturn iterate_and_advance2(iter, len, priv, NULL, ustep, step);\n 340:\t}\n 341:\t\n 342:\t/**\n 343:\t * iterate_and_advance_kernel - Iterate over a kernel-internal iterator\n 344:\t * @iter: The iterator to iterate over.\n 345:\t * @len: The amount to iterate over.\n 346:\t * @priv: Data for the step functions.\n 347:\t * @priv2: More data for the step functions.\n 348:\t * @step: Function for other iterators; given kernel addresses.\n 349:\t *\n 350:\t * Iterate over the next part of an iterator, up to the specified length. The\n 351:\t * buffer is presented in segments, which for kernel iteration are broken up by\n 352:\t * physical pages and mapped, with the mapped address being presented.\n 353:\t *\n 354:\t * [!] Note This will only handle BVEC, KVEC, BVECQ, XARRAY and DISCARD-type\n 355:\t * iterators; it will not handle UBUF or IOVEC-type iterators.\n 356:\t *\n 357:\t * A step functions, @step, must be provided, one for handling mapped kernel\n 358:\t * addresses and the other is given user addresses which have the potential to\n 359:\t * fault since no pinning is performed.\n 360:\t *\n 361:\t * The step functions are passed the address and length of the segment, @priv,\n 362:\t * @priv2 and the amount of data so far iterated over (which can, for example,\n 363:\t * be added to @priv to point to the right part of a second buffer). The step\n 364:\t * functions should return the amount of the segment they didn't process (ie. 0\n 365:\t * indicates complete processsing).\n 366:\t *\n 367:\t * This function returns the amount of data processed (ie. 0 means nothing was\n 368:\t * processed and the value of @len means processes to completion).\n 369:\t */\n 370:\tstatic __always_inline\n 371:\tsize_t iterate_and_advance_kernel(struct iov_iter *iter, size_t len, void *priv,\n 372:\t\t\t\t\t void *priv2, iov_step_f step)\n 373:\t{\n 374:\t\tif (unlikely(iter-\u003ecount \u003c len))\n"
} |
| 39/3 | 2026/10/06 21:26 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 55567 output: 28 thoughts: 215 Thoughts: **Analyzing the `iov_iter_zero` Safety and Patch Context** Okay, so looking at line 374, that `if (unlikely(iter->count < len)) len = iter->count;` is the key. It makes the `iov_iter_zero` call using `subreq->len` with the iterator safe. Even when `iter->count` is `subreq->len - subreq->transferred`, this conditional ensures `len` is clamped to `iter->count` before being used. That means the write length can't possibly exceed what's available in the iterator. I'm satisfied the integer underflow scenario is correctly handled. Now, let's zoom out and understand the *why* of this code. What's this patch all about? Aha, it seems to be David Howells, or a netfs maintainer's work. It looks like the patch replaces the `rolling_buffer` construct in netfs with a more direct approach using `bvecq` and its associated `struct bvecq_pos` cursors. This makes sense; I'm familiar with both techniques. It suggests a potential performance optimization or perhaps a simplification of the code by leveraging `bvecq`'s built-in features. I should probably check the git log or recent commits related to netfs to confirm this understanding and get a better sense of the surrounding discussion and motivation for this change. I'll search for 'rolling_buffer' or 'bvecq' in the commit history to see if there's any relevant context I might be missing. |
| 40/3 | 2026/10/06 21:27 | tool | grepper |
0mArgs:{
"Expression": "bvecq",
"PathPrefixes": [
"fs/netfs"
]
}Results: {
"Output": "\nFull output is too long, showing 200 out of 952 lines.\nUse more precise expression or PathPrefixes if possible.\n\nNumber of matching lines per file (17 files in total):\nfs/netfs/Makefile:1\nfs/netfs/buffered_read.c:31\nfs/netfs/bvecq.c:102\nfs/netfs/direct_read.c:7\nfs/netfs/direct_write.c:14\nfs/netfs/internal.h:2\nfs/netfs/iterator.c:18\nfs/netfs/main.c:7\nfs/netfs/objects.c:5\nfs/netfs/read_collect.c:14\nfs/netfs/read_pgpriv2.c:19\nfs/netfs/read_retry.c:14\nfs/netfs/read_single.c:3\nfs/netfs/stats.c:2\nfs/netfs/write_collect.c:4\nfs/netfs/write_issue.c:19\nfs/netfs/write_retry.c:10\n\nfs/netfs/Makefile=3=netfs-y := \\\n--\nfs/netfs/Makefile-5-\tbuffered_write.o \\\nfs/netfs/Makefile:6:\tbvecq.o \\\nfs/netfs/Makefile-7-\tdirect_read.o \\\n--\nfs/netfs/buffered_read.c=114=static ssize_t netfs_prepare_read_iterator(struct netfs_io_subrequest *subreq)\n--\nfs/netfs/buffered_read.c-123-\nfs/netfs/buffered_read.c:124:\tbvecq_pos_set(\u0026subreq-\u003eio_buffer, \u0026rreq-\u003edispatch_cursor);\nfs/netfs/buffered_read.c:125:\textracted = bvecq_slice(\u0026rreq-\u003edispatch_cursor, rsize,\nfs/netfs/buffered_read.c-126-\t\t\t\tstream-\u003esreq_max_segs, \u0026subreq-\u003enr_segs);\n--\nfs/netfs/buffered_read.c=187=static void netfs_issue_read(struct netfs_io_request *rreq,\n--\nfs/netfs/buffered_read.c-189-{\nfs/netfs/buffered_read.c:190:\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\nfs/netfs/buffered_read.c-191-\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n--\nfs/netfs/buffered_read.c-200-\tdefault:\nfs/netfs/buffered_read.c:201:\t\tbvecq_zero(\u0026subreq-\u003eio_buffer, subreq-\u003elen);\nfs/netfs/buffered_read.c-202-\t\tsubreq-\u003etransferred = subreq-\u003elen;\n--\nfs/netfs/buffered_read.c=214=static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,\nfs/netfs/buffered_read.c:215:\t\t\t\t struct bvecq_pos *mark, size_t len, bool copy)\nfs/netfs/buffered_read.c-216-{\nfs/netfs/buffered_read.c:217:\tstruct bvecq *bq = mark-\u003ebvecq;\nfs/netfs/buffered_read.c-218-\tunsigned int offset = mark-\u003eoffset;\n--\nfs/netfs/buffered_read.c-225-\t\t\tbreak;\nfs/netfs/buffered_read.c:226:\t\tif (!bvecq_acquire_slot(bq, slot)) {\nfs/netfs/buffered_read.c-227-\t\t\tbq = bq-\u003enext;\n--\nfs/netfs/buffered_read.c-258-\tif (bq) {\nfs/netfs/buffered_read.c:259:\t\tbvecq_pos_move(mark, bq);\nfs/netfs/buffered_read.c-260-\t\tmark-\u003eoffset = offset;\n--\nfs/netfs/buffered_read.c-262-\t} else {\nfs/netfs/buffered_read.c:263:\t\tbvecq_pos_unset(mark);\nfs/netfs/buffered_read.c-264-\t}\n--\nfs/netfs/buffered_read.c=272=static void netfs_read_to_pagecache(struct netfs_io_request *rreq)\n--\nfs/netfs/buffered_read.c-282-\tstruct fscache_occupancy *occ = \u0026_occ;\nfs/netfs/buffered_read.c:283:\tstruct bvecq_pos mark_cursor;\nfs/netfs/buffered_read.c-284-\tssize_t size = rreq-\u003elen;\n--\nfs/netfs/buffered_read.c-289-\nfs/netfs/buffered_read.c:290:\tbvecq_pos_set(\u0026mark_cursor, \u0026rreq-\u003edispatch_cursor);\nfs/netfs/buffered_read.c-291-\n--\nfs/netfs/buffered_read.c-418-\nfs/netfs/buffered_read.c:419:\t\tif (mark_cursor.bvecq) {\nfs/netfs/buffered_read.c-420-\t\t\t/* See if the cache indicated this should be cached. */\n--\nfs/netfs/buffered_read.c-443-\nfs/netfs/buffered_read.c:444:\tbvecq_pos_unset(\u0026mark_cursor);\nfs/netfs/buffered_read.c:445:\tbvecq_pos_unset(\u0026rreq-\u003edispatch_cursor);\nfs/netfs/buffered_read.c-446-}\n--\nfs/netfs/buffered_read.c=463=void netfs_readahead(struct readahead_control *ractl)\n--\nfs/netfs/buffered_read.c-488-\nfs/netfs/buffered_read.c:489:\t/* Load the folios to be read into a bvecq chain. Note that this\nfs/netfs/buffered_read.c-490-\t * acquires a ref on each folio that we will need to release later -\n--\nfs/netfs/buffered_read.c-492-\t */\nfs/netfs/buffered_read.c:493:\tadded = bvecq_load_from_ra(\u0026rreq-\u003edispatch_cursor, ractl);\nfs/netfs/buffered_read.c-494-\tif (added \u003c 0) {\n--\nfs/netfs/buffered_read.c-501-\trreq-\u003ecleaned_to = rreq-\u003estart;\nfs/netfs/buffered_read.c:502:\tbvecq_pos_set(\u0026rreq-\u003ecollect_cursor, \u0026rreq-\u003edispatch_cursor);\nfs/netfs/buffered_read.c-503-\tnetfs_read_set_unlock_at(rreq);\n--\nfs/netfs/buffered_read.c=517=static int netfs_create_singular_buffer(struct netfs_io_request *rreq, struct folio *folio)\nfs/netfs/buffered_read.c-518-{\nfs/netfs/buffered_read.c:519:\tstruct bvecq *bq;\nfs/netfs/buffered_read.c-520-\tsize_t fsize = folio_size(folio);\nfs/netfs/buffered_read.c-521-\nfs/netfs/buffered_read.c:522:\tbq = bvecq_alloc_one(1, rreq-\u003egfp, false);\nfs/netfs/buffered_read.c-523-\tif (!bq)\n--\nfs/netfs/buffered_read.c-525-\nfs/netfs/buffered_read.c:526:\trreq-\u003edispatch_cursor.bvecq = bq;\nfs/netfs/buffered_read.c-527-\trreq-\u003edispatch_cursor.slot = 0;\n--\nfs/netfs/buffered_read.c-530-\tbvec_set_folio(\u0026bq-\u003ebv[0], folio, fsize, 0);\nfs/netfs/buffered_read.c:531:\tbvecq_filled_to(bq, 1);\nfs/netfs/buffered_read.c:532:\tbvecq_pos_set(\u0026rreq-\u003ecollect_cursor, \u0026rreq-\u003edispatch_cursor);\nfs/netfs/buffered_read.c-533-\trreq-\u003esubmitted = rreq-\u003estart + fsize;\n--\nfs/netfs/buffered_read.c=541=static int netfs_read_gaps(struct file *file, struct folio *folio)\n--\nfs/netfs/buffered_read.c-547-\tstruct netfs_inode *ctx = netfs_inode(mapping-\u003ehost);\nfs/netfs/buffered_read.c:548:\tstruct bvecq *bq = NULL;\nfs/netfs/buffered_read.c-549-\tunsigned int from = finfo-\u003edirty_offset;\n--\nfs/netfs/buffered_read.c-575-\tret = -ENOMEM;\nfs/netfs/buffered_read.c:576:\tbq = bvecq_alloc_chain(nr_bvec, rreq-\u003egfp, false);\nfs/netfs/buffered_read.c-577-\tif (!bq)\nfs/netfs/buffered_read.c-578-\t\tgoto discard;\nfs/netfs/buffered_read.c:579:\trreq-\u003edispatch_cursor.bvecq = bq;\nfs/netfs/buffered_read.c-580-\n--\nfs/netfs/buffered_read.c-582-\nfs/netfs/buffered_read.c:583:\tfor (struct bvecq *p = bq; p; p = p-\u003enext)\nfs/netfs/buffered_read.c-584-\t\tp-\u003emem_type = BVECQ_MEM_PAGECACHE;\n--\nfs/netfs/buffered_read.c-596-\t\tif (i \u003e= bq-\u003emax_slots) {\nfs/netfs/buffered_read.c:597:\t\t\tbvecq_filled_to(bq, i);\nfs/netfs/buffered_read.c-598-\t\t\tbq = bq-\u003enext;\n--\nfs/netfs/buffered_read.c-606-\t\tif (i \u003e= bq-\u003emax_slots) {\nfs/netfs/buffered_read.c:607:\t\t\tbvecq_filled_to(bq, i);\nfs/netfs/buffered_read.c-608-\t\t\tbq = bq-\u003enext;\n--\nfs/netfs/buffered_read.c-613-\t}\nfs/netfs/buffered_read.c:614:\tbvecq_filled_to(bq, i);\nfs/netfs/buffered_read.c-615-\n--\nfs/netfs/buffered_read.c-637-\nfs/netfs/buffered_read.c:638:\tbvecq_pos_unset(\u0026rreq-\u003edispatch_cursor);\nfs/netfs/buffered_read.c-639-\tfolio_unlock(folio);\n--\nfs/netfs/buffered_read.c-643-discard:\nfs/netfs/buffered_read.c:644:\tbvecq_pos_unset(\u0026rreq-\u003edispatch_cursor);\nfs/netfs/buffered_read.c-645-\tnetfs_put_failed_request(rreq);\n--\nfs/netfs/bvecq.c-7-\nfs/netfs/bvecq.c:8:#include \u003clinux/bvecq.h\u003e\nfs/netfs/bvecq.c-9-#include \"internal.h\"\nfs/netfs/bvecq.c-10-\nfs/netfs/bvecq.c:11:void bvecq_dump(const struct bvecq *bq)\nfs/netfs/bvecq.c-12-{\n--\nfs/netfs/bvecq.c-14-\nfs/netfs/bvecq.c:15:\tfor (; bq; bq = bvecq_next(bq), b++) {\nfs/netfs/bvecq.c-16-\t\tint skipz = 0;\n--\nfs/netfs/bvecq.c-36-}\nfs/netfs/bvecq.c:37:EXPORT_SYMBOL(bvecq_dump);\nfs/netfs/bvecq.c-38-\nfs/netfs/bvecq.c-39-/**\nfs/netfs/bvecq.c:40: * bvecq_alloc_one - Allocate a single bvecq node with unpopulated slots\nfs/netfs/bvecq.c-41- * @nr_slots: Number of slots to allocate\n--\nfs/netfs/bvecq.c-44- *\nfs/netfs/bvecq.c:45: * Allocate a single bvecq node and initialise the header. The allocation is\nfs/netfs/bvecq.c-46- * rounded up to the size of the smallest slab granule that will accommodate it\n--\nfs/netfs/bvecq.c-57- */\nfs/netfs/bvecq.c:58:struct bvecq *bvecq_alloc_one(size_t nr_slots, gfp_t gfp, bool for_writeback)\nfs/netfs/bvecq.c-59-{\nfs/netfs/bvecq.c:60:\tstruct bvecq *bq;\nfs/netfs/bvecq.c:61:\tconst size_t pool_size = struct_size_t(struct bvecq, __bv, BVECQ_POOL_SLOTS);\nfs/netfs/bvecq.c-62-\tsize_t size;\n--\nfs/netfs/bvecq.c-78-\nfs/netfs/bvecq.c:79:\t\tbq = mempool_alloc(\u0026netfs_bvecq_pool, gfp);\nfs/netfs/bvecq.c-80-\t\tif (!bq)\n--\nfs/netfs/bvecq.c-90-success:\nfs/netfs/bvecq.c:91:\t*bq = (struct bvecq) {\nfs/netfs/bvecq.c-92-\t\t.ref\t\t= REFCOUNT_INIT(1),\n--\nfs/netfs/bvecq.c-97-\t};\nfs/netfs/bvecq.c:98:\tnetfs_stat(\u0026netfs_n_bvecq);\nfs/netfs/bvecq.c-99-\treturn bq;\nfs/netfs/bvecq.c-100-}\nfs/netfs/bvecq.c:101:EXPORT_SYMBOL(bvecq_alloc_one);\nfs/netfs/bvecq.c-102-\nfs/netfs/bvecq.c-103-/**\nfs/netfs/bvecq.c:104: * bvecq_alloc_chain - Allocate an unpopulated bvecq chain\nfs/netfs/bvecq.c-105- * @nr_slots: Number of slots to allocate\n--\nfs/netfs/bvecq.c-108- *\nfs/netfs/bvecq.c:109: * Allocate a chain of bvecq nodes providing at least the requested cumulative\nfs/netfs/bvecq.c-110- * number of slots. Each node is a maximum of 4KiB in size.\n--\nfs/netfs/bvecq.c-113- */\nfs/netfs/bvecq.c:114:struct bvecq *bvecq_alloc_chain(size_t nr_slots, gfp_t gfp, bool for_writeback)\nfs/netfs/bvecq.c-115-{\nfs/netfs/bvecq.c:116:\tstruct bvecq *head = NULL, *tail = NULL;\nfs/netfs/bvecq.c-117-\n--\nfs/netfs/bvecq.c-120-\tfor (;;) {\nfs/netfs/bvecq.c:121:\t\tstruct bvecq *bq;\nfs/netfs/bvecq.c-122-\nfs/netfs/bvecq.c:123:\t\tbq = bvecq_alloc_one(min(nr_slots, BVECQ_4KB_SLOTS), gfp, for_writeback);\nfs/netfs/bvecq.c-124-\t\tif (!bq)\n--\nfs/netfs/bvecq.c-127-\t\tif (tail)\nfs/netfs/bvecq.c:128:\t\t\tbvecq_append(tail, bq);\nfs/netfs/bvecq.c-129-\t\telse\n--\nfs/netfs/bvecq.c-138-oom:\nfs/netfs/bvecq.c:139:\tbvecq_put(head);\nfs/netfs/bvecq.c-140-\treturn NULL;\nfs/netfs/bvecq.c-141-}\nfs/netfs/bvecq.c:142:EXPORT_SYMBOL(bvecq_alloc_chain);\n\n"
} |
| 41/3 | 2026/10/06 21:27 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 60565 output: 37 thoughts: 37 Thoughts: Okay, let's dive into the core of this `bvecq` initialization. **Deep Dive into `bvecq_alloc_one`** Right, so I need to get a clear picture of how this `bvecq` is being set up, specifically within the `fs/netfs/bvecq.c` file. The function to scrutinize is `bvecq_alloc_one`. Let's assume there are no gotchas and that it is what I think it is. I'm going to look at the memory allocation, the assignment of values, and any relationships to the other structures involved, as I suspect these are critical for the correct operation of the system. I have to make sure there are no obvious flaws or performance bottlenecks in the initialization process, and I'll need to double-check that this allocation is thread-safe and properly handles error conditions like memory exhaustion. Any dependencies on external resources also need to be noted. I'll mentally map this against my existing knowledge of other networking data structures. After the allocation, I'll need to figure out which components are being initialized to avoid future issues. |
| 42/3 | 2026/10/06 21:27 | tool | read-file |
0mArgs:{
"File": "fs/netfs/bvecq.c",
"FirstLine": 58,
"LineCount": 45
}Results: {
"Contents": " 58:\tstruct bvecq *bvecq_alloc_one(size_t nr_slots, gfp_t gfp, bool for_writeback)\n 59:\t{\n 60:\t\tstruct bvecq *bq;\n 61:\t\tconst size_t pool_size = struct_size_t(struct bvecq, __bv, BVECQ_POOL_SLOTS);\n 62:\t\tsize_t size;\n 63:\t\tbool from_pool = false;\n 64:\t\n 65:\t\tsize = kmalloc_size_roundup(struct_size(bq, __bv, nr_slots));\n 66:\t\tgfp \u0026= ~(GFP_ZONEMASK | __GFP_THISNODE);\n 67:\t\n 68:\t\tif (for_writeback) {\n 69:\t\t\tif (size != pool_size) {\n 70:\t\t\t\tgfp_t gfp_temp = gfp;\n 71:\t\n 72:\t\t\t\tgfp_temp |= __GFP_NOMEMALLOC | __GFP_NORETRY | __GFP_NOWARN;\n 73:\t\t\t\tgfp_temp \u0026= ~(__GFP_DIRECT_RECLAIM | __GFP_IO);\n 74:\t\t\t\tbq = kmalloc(size, gfp_temp);\n 75:\t\t\t\tif (bq)\n 76:\t\t\t\t\tgoto success;\n 77:\t\t\t}\n 78:\t\n 79:\t\t\tbq = mempool_alloc(\u0026netfs_bvecq_pool, gfp);\n 80:\t\t\tif (!bq)\n 81:\t\t\t\treturn bq;\n 82:\t\t\tfrom_pool = true;\n 83:\t\t\tsize = pool_size;\n 84:\t\t} else {\n 85:\t\t\tbq = kmalloc(size, gfp);\n 86:\t\t\tif (!bq)\n 87:\t\t\t\treturn bq;\n 88:\t\t}\n 89:\t\n 90:\tsuccess:\n 91:\t\t*bq = (struct bvecq) {\n 92:\t\t\t.ref\t\t= REFCOUNT_INIT(1),\n 93:\t\t\t.bv\t\t= bq-\u003e__bv,\n 94:\t\t\t.inline_bv\t= true,\n 95:\t\t\t.max_slots\t= (size - sizeof(*bq)) / sizeof(bq-\u003e__bv[0]),\n 96:\t\t\t.from_pool\t= from_pool,\n 97:\t\t};\n 98:\t\tnetfs_stat(\u0026netfs_n_bvecq);\n 99:\t\treturn bq;\n 100:\t}\n 101:\tEXPORT_SYMBOL(bvecq_alloc_one);\n 102:\t\n"
} |
| 43/3 | 2026/10/06 21:27 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 61545 output: 36 thoughts: 64 Thoughts: Okay, let's break this down.
**Investigating the Zero-Initialization of `struct bvecq`**
Right, so I'm looking at this initialization: `*bq = (struct bvecq) { ... }`. My immediate thought is, "Ah, the beauty of C99 designated initializers!" That format implicitly zeroes *all* the members of the `struct bvecq` that aren't explicitly assigned values within the curly braces. That's good; it's clean and makes sure we're starting with a defined state.
But now the key question is: what *are* those members of `struct bvecq`? I can't just assume I know them off the top of my head. Time to consult the source. Let me go find the definition. It'll be in `include/linux/bvecq.h`. Right there, that's where the structure's layout lives, the truth! Now I can confirm what's being implicitly zeroed and make sure there are no surprises with my assumptions.
|
| 44/3 | 2026/10/06 21:27 | tool | read-file |
0mArgs:{
"File": "include/linux/bvecq.h",
"FirstLine": 25,
"LineCount": 40
}Results: {
"Contents": " 25:\t * Segmented bio_vec queue.\n 26:\t *\n 27:\t * These can be linked together to form messages of indefinite length and\n 28:\t * iterated over with an ITER_BVECQ iterator. The list is non-circular; next\n 29:\t * and prev are NULL at the ends.\n 30:\t *\n 31:\t * The bv pointer points to the bio_vec array; this may be __bv if allocated\n 32:\t * together. The caller is responsible for determining whether or not this is\n 33:\t * the case as the array pointed to by bv may be follow on directly from the\n 34:\t * bvecq by accident of allocation (ie. -\u003ebv == -\u003e__bv is *not* sufficient to\n 35:\t * determine this).\n 36:\t */\n 37:\tstruct bvecq {\n 38:\t\tstruct bvecq\t*next;\t\t/* Next bvec in the list or NULL */\n 39:\t\tstruct bvecq\t*prev;\t\t/* Prev bvec in the list or NULL */\n 40:\t\trefcount_t\tref;\n 41:\t\tu32\t\tpriv;\t\t/* Private data */\n 42:\t\tu16\t\tnr_slots;\t/* Number of elements in bv[] used */\n 43:\t\tu16\t\tmax_slots;\t/* Number of elements allocated in bv[] */\n 44:\t\tenum bvecq_mem\tmem_type:3;\t/* What sort of memory and how to free it */\n 45:\t\tbool\t\tinline_bv:1;\t/* T if __bv[] is being used */\n 46:\t\tbool\t\tfrom_pool:1;\t/* T if bvecq from mempool */\n 47:\t\tstruct bio_vec\t*bv;\t\t/* Pointer to array of page fragments */\n 48:\t\tstruct bio_vec\t__bv[];\t\t/* Default array (if -\u003einline_bv) */\n 49:\t};\n 50:\t\n 51:\t/* Number of slots in a 512-byte mempool-backed bvecq. */\n 52:\t#define BVECQ_POOL_SLOTS ((512 - sizeof(struct bvecq)) / sizeof(struct bio_vec))\n 53:\t\n 54:\t/* Number of slots in a 4K bvecq. */\n 55:\t#define BVECQ_4KB_SLOTS ((4096 - sizeof(struct bvecq)) / sizeof(struct bio_vec))\n 56:\t\n 57:\t/*\n 58:\t * Position in a bio_vec queue. The bvecq holds a ref on the queue segment it\n 59:\t * points to.\n 60:\t */\n 61:\tstruct bvecq_pos {\n 62:\t\tstruct bvecq\t\t*bvecq;\t\t/* The first bvecq */\n 63:\t\tunsigned int\t\toffset;\t\t/* The offset within the starting slot */\n 64:\t\tu16\t\t\tslot;\t\t/* The starting slot */\n"
} |
| 45/3 | 2026/10/06 21:27 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 62631 output: 44 thoughts: 184 Thoughts: **Analyzing the `bvecq_alloc_one` Function**
Okay, let's dive into `bvecq_alloc_one` and see what's happening. The `struct bvecq` itself is being zero-initialized using the struct initialization syntax: `= (struct bvecq){...}`. My immediate question is about `__bv[]`. Does that get zeroed out as well?
Ah, right, `kmalloc(size, ...)` doesn't actually zero-initialize the memory it allocates for `__bv[]`. That’s a potential gotcha! So, when is `__bv[]` actually accessed, and how?
Okay, digging deeper… I see `nr_slots` starts at 0, which makes sense given the zero-initialization of omitted fields in the struct.
Then, slots are added. For example, `bvec_set_page(&bq->bv[slot], ...)` or `bvec_set_folio(...)` are called. After those calls, `bvecq_filled_to(bq, slot)` is called. I had better take a look at `bvecq_filled_to` to understand exactly how the system is behaving.
|
| 46/3 | 2026/10/06 21:27 | tool | grepper |
0mArgs:{
"Expression": "bvecq_filled_to",
"PathPrefixes": [
"include/linux/bvecq.h",
"fs/netfs"
]
}Results: {
"Output": "fs/netfs/buffered_read.c=517=static int netfs_create_singular_buffer(struct netfs_io_request *rreq, struct folio *folio)\n--\nfs/netfs/buffered_read.c-530-\tbvec_set_folio(\u0026bq-\u003ebv[0], folio, fsize, 0);\nfs/netfs/buffered_read.c:531:\tbvecq_filled_to(bq, 1);\nfs/netfs/buffered_read.c-532-\tbvecq_pos_set(\u0026rreq-\u003ecollect_cursor, \u0026rreq-\u003edispatch_cursor);\n--\nfs/netfs/buffered_read.c=541=static int netfs_read_gaps(struct file *file, struct folio *folio)\n--\nfs/netfs/buffered_read.c-596-\t\tif (i \u003e= bq-\u003emax_slots) {\nfs/netfs/buffered_read.c:597:\t\t\tbvecq_filled_to(bq, i);\nfs/netfs/buffered_read.c-598-\t\t\tbq = bq-\u003enext;\n--\nfs/netfs/buffered_read.c-606-\t\tif (i \u003e= bq-\u003emax_slots) {\nfs/netfs/buffered_read.c:607:\t\t\tbvecq_filled_to(bq, i);\nfs/netfs/buffered_read.c-608-\t\t\tbq = bq-\u003enext;\n--\nfs/netfs/buffered_read.c-613-\t}\nfs/netfs/buffered_read.c:614:\tbvecq_filled_to(bq, i);\nfs/netfs/buffered_read.c-615-\n--\nfs/netfs/bvecq.c=162=struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp,\n--\nfs/netfs/bvecq.c-216-\nfs/netfs/bvecq.c:217:\t\tbvecq_filled_to(p, slot);\nfs/netfs/bvecq.c-218-\t\tcount -= got;\n--\nfs/netfs/bvecq.c=294=int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size, gfp_t gfp)\n--\nfs/netfs/bvecq.c-338-\t\t*_cur_size += PAGE_SIZE \u003c\u003c order;\nfs/netfs/bvecq.c:339:\t\tbvecq_filled_to(tail, tail-\u003enr_slots + 1);\nfs/netfs/bvecq.c-340-\t}\n--\nfs/netfs/bvecq.c=606=ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl)\n--\nfs/netfs/bvecq.c-637-\t\tif (slot \u003e= bq-\u003emax_slots) {\nfs/netfs/bvecq.c:638:\t\t\tbvecq_filled_to(bq, slot);\nfs/netfs/bvecq.c-639-\t\t\tbq = bq-\u003enext;\n--\nfs/netfs/bvecq.c-648-\tif (bq)\nfs/netfs/bvecq.c:649:\t\tbvecq_filled_to(bq, slot);\nfs/netfs/bvecq.c-650-\n--\nfs/netfs/iterator.c=42=ssize_t netfs_extract_iter(struct iov_iter *orig, size_t max_len, size_t max_pages,\n--\nfs/netfs/iterator.c-136-\nfs/netfs/iterator.c:137:\t\t\tbvecq_filled_to(bq, slot);\nfs/netfs/iterator.c-138-\t\t} while (max_len \u003e 0 \u0026\u0026 max_pages \u003e 0 \u0026\u0026 !bvecq_is_full(bq));\n--\nfs/netfs/read_pgpriv2.c=19=static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio *folio)\n--\nfs/netfs/read_pgpriv2.c-77-\tslot++;\nfs/netfs/read_pgpriv2.c:78:\tbvecq_filled_to(queue, slot);\nfs/netfs/read_pgpriv2.c-79-\tcreq-\u003eload_cursor.slot = slot;\n--\nfs/netfs/write_issue.c=328=static int netfs_write_folio(struct netfs_io_request *wreq,\n--\nfs/netfs/write_issue.c-469-\tslot++;\nfs/netfs/write_issue.c:470:\tbvecq_filled_to(queue, slot);\nfs/netfs/write_issue.c-471-\twreq-\u003eload_cursor.slot = slot;\n--\ninclude/linux/bvecq.h=111=static inline bool bvecq_is_full(const struct bvecq *bvecq)\n--\ninclude/linux/bvecq.h-116-/**\ninclude/linux/bvecq.h:117: * bvecq_filled_to - Release filled slots with release barrier\ninclude/linux/bvecq.h-118- * @bvecq: The object modified\n--\ninclude/linux/bvecq.h-120- */\ninclude/linux/bvecq.h:121:static inline void bvecq_filled_to(struct bvecq *bvecq, unsigned int to)\ninclude/linux/bvecq.h-122-{\n"
} |
| 47/3 | 2026/10/06 21:27 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 64574 output: 37 thoughts: 49 Thoughts: Okay, let's dive into this. I need to refresh my understanding of how the `bvecq` functionality in the Linux kernel is implemented. Specifically, I want to clarify the interaction between `bvecq_filled_to` and `bvecq_acquire_slot` / `bvecq_nr_slots_acquire`, as defined in `include/linux/bvecq.h`. I've worked with block devices for years, so I understand the core concepts of `bvec` queues, but details sometimes get fuzzy with time. I need to make sure I grasp the finer points of how these specific functions manage the acquisition and filling of slots within the queue. This is crucial for performance and preventing race conditions, so getting a solid mental model here is important. Let's see how these pieces fit together. |
| 48/3 | 2026/10/06 21:27 | tool | read-file |
0mArgs:{
"File": "include/linux/bvecq.h",
"FirstLine": 115,
"LineCount": 50
}Results: {
"Contents": " 115:\t\n 116:\t/**\n 117:\t * bvecq_filled_to - Release filled slots with release barrier\n 118:\t * @bvecq: The object modified\n 119:\t * @to: The latest slot filled + 1\n 120:\t */\n 121:\tstatic inline void bvecq_filled_to(struct bvecq *bvecq, unsigned int to)\n 122:\t{\n 123:\t\t/* Set the slot counter after filling the slot */\n 124:\t\tsmp_store_release(\u0026bvecq-\u003enr_slots, to);\n 125:\t}\n 126:\t\n 127:\t/**\n 128:\t * bvecq_nr_slots_acquire - Get the number of filled slots with acquire barrier\n 129:\t * @bvecq: The object to query\n 130:\t *\n 131:\t * Return: The number of filled slots\n 132:\t */\n 133:\tstatic inline unsigned int bvecq_nr_slots_acquire(const struct bvecq *bvecq)\n 134:\t{\n 135:\t\t/* Read the slot counter before looking at the slot */\n 136:\t\treturn smp_load_acquire(\u0026bvecq-\u003enr_slots);\n 137:\t}\n 138:\t\n 139:\t/**\n 140:\t * bvecq_acquire_slot - Determine if a slot is valid with acquire barrier\n 141:\t * @bvecq: The object to query\n 142:\t * @slot: The next slot\n 143:\t *\n 144:\t * Return: true if valid; false if might not be valid\n 145:\t */\n 146:\tstatic inline bool bvecq_acquire_slot(const struct bvecq *bvecq, unsigned int slot)\n 147:\t{\n 148:\t\t/* Read the slot counter before looking at the slot */\n 149:\t\treturn slot \u003c bvecq_nr_slots_acquire(bvecq);\n 150:\t}\n 151:\t\n 152:\t/**\n 153:\t * bvecq_append - Get the next bvecq with appropriate barrier\n 154:\t * @to: The bvecq to append to\n 155:\t * @add: The bvecq to append\n 156:\t *\n 157:\t * Attach a new bvecq to a chain using an appropriate barrier to protect the\n 158:\t * write.\n 159:\t *\n 160:\t * [!] Note that this function transfers the caller's ref to the chain.\n 161:\t */\n 162:\tstatic inline void bvecq_append(struct bvecq *to, struct bvecq *add)\n 163:\t{\n 164:\t\tadd-\u003eprev = to;\n"
} |
| 49/3 | 2026/10/06 21:27 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 65594 output: 26 thoughts: 1942 Thoughts: **Deep Dive into `netfs_extract_iter` and Memory Allocation in `bvecq`** Okay, let's break this down. My initial thought was centered around how `bvecq` uses producer-consumer concurrency, specifically around the `smp_store_release` and `bvecq_nr_slots_acquire` functions. I'm seeing an interesting pattern emerging, with slot initialization happening before the increment of `bvecq_filled_to`, implying a carefully managed concurrency scheme. But now, this `netfs_extract_iter` function in `fs/netfs/iterator.c` has caught my eye, and it requires a closer look. The core of my interest lies within the memory allocation and page handling inside the loop: I am particularly concerned with this chunk of code in `netfs_extract_iter`: ```c size_t space = bq->max_slots - slot; size_t bv_size = array_size(bq->max_slots, sizeof(*bv)); size_t pg_size = array_size(space, sizeof(*pages)); /* Put the page list at the end of the bvec list * storage. bvec elements are larger than page * pointers, so as long as we work 0->last, we should * be fine. */ pages = (void *)bv + bv_size - pg_size; ``` Here's my thought process: 1. **`bq->bv` is Key:** First, I establish that `bv` is simply the pointer to the `bq->bv` field, which is the start of the `bio_vec` array within the `bvecq` structure. I'm tracking this pointer and its behavior. 2. **Allocation Source:** `bvecq` memory is allocated with `bvecq_alloc_one(max_pages, gfp, false)`. Inside, the size calculation is critical, and the formula comes out to `size = kmalloc_size_roundup(struct_size(bq, __bv, nr_slots));`. From this call, the fields of `bvecq` are initialized, including `bq->max_slots`. The size of the total memory allocated is `size`, where `size` can be retrieved. 3. **`max_slots` Calculation:** The allocation sets `bq->max_slots` to `(size - sizeof(*bq)) / sizeof(bq->__bv[0])`. The size of the `bio_vec` array is set by this formula. 4. **`bio_vec` and Page Pointer Overlap:** The concern is whether the `pages` pointer array is writing beyond the end of allocated memory. I need to be sure the allocations do not overlap, which are adjacent in memory. 5. **Size of `bio_vec` and `pages`:** We know a `bio_vec` (`bv`) is 16 bytes (8 for `bv_page`, 4 for `bv_len`, and 4 for `bv_offset`) and a page pointer (`pages`) is 8 bytes. With this information I can verify that `(void *)bv + bv_size - pg_size` is safe. 6. **Pointer Arithmetic:** Calculating the offset of the `pages` array requires analyzing: `bv_size = bq->max_slots * 16`. And `pg_size = space * 8`. So `pages` starts at offset `16 * max_slots - 8 * (max_slots - slot) = 8 * max_slots + 8 * slot`. The current slot being written to is `slot`, which is at offset `16 * slot`. I have to verify that no overlapping will occur between writes. 7. **Safety Check:** Now the question is: Is `16 * slot <= 8 * max_slots + 8 * slot`? This simplifies to `8 * slot <= 8 * max_slots`, meaning that `slot <= max_slots`, which is true, because we initialized `slot = 0` and increment it in an loop that has a condition to stop at `slot < bq->max_slots`. 8. **Confirmation:** So, yes, `bq->bv[slot]` never collides with the `pages` allocation. The pointer arithmetic is sound and the allocation is safe. Now, let's explore how `iov_iter_extract_pages` works since it is responsible for populating the `pages` array. |
| 50/3 | 2026/10/06 21:27 | tool | grepper |
0mArgs:{
"Expression": "ssize_t iov_iter_extract_pages"
}Results: {
"Output": "include/linux/uio.h=371=static inline void iov_iter_ubuf(struct iov_iter *i, unsigned int direction,\n--\ninclude/linux/uio.h-386-\ninclude/linux/uio.h:387:ssize_t iov_iter_extract_pages(struct iov_iter *i, struct page ***pages,\ninclude/linux/uio.h-388-\t\t\t size_t maxsize, unsigned int maxpages,\n--\nlib/iov_iter.c=1950=static ssize_t iov_iter_extract_user_pages(struct iov_iter *i,\n--\nlib/iov_iter.c-2032- */\nlib/iov_iter.c:2033:ssize_t iov_iter_extract_pages(struct iov_iter *i,\nlib/iov_iter.c-2034-\t\t\t struct page ***pages,\n"
} |
| 51/3 | 2026/10/06 21:27 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 68796 output: 36 thoughts: 29 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Diving into `iov_iter_extract_pages`** Right, let's get into the nitty-gritty. Time to crack open the `iov_iter_extract_pages` implementation. Specifically, I'm heading straight for `lib/iov_iter.c:2033`. I need a clear understanding of how this function is handling page extraction, especially given the context of [*Insert user's field of expertise, e.g., high-performance networking, storage I/O, etc.*]. I'm expecting to see how it's dealing with scattered I/O, page faults, and potential optimization strategies. I'm keen to evaluate the memory access patterns and how efficiently it's interacting with the page cache. I'll be looking for any potential bottlenecks, race conditions, or areas ripe for improvement. I'll pay close attention to the error handling within this function and how it integrates with other parts of the I/O subsystem. This is where the rubber meets the road, so I need to thoroughly understand this critical path. |
| 52/3 | 2026/10/06 21:27 | tool | read-file |
0mArgs:{
"File": "lib/iov_iter.c",
"FirstLine": 2033,
"LineCount": 40
}Results: {
"Contents": "2033:\tssize_t iov_iter_extract_pages(struct iov_iter *i,\n2034:\t\t\t\t struct page ***pages,\n2035:\t\t\t\t size_t maxsize,\n2036:\t\t\t\t unsigned int maxpages,\n2037:\t\t\t\t iov_iter_extraction_t extraction_flags,\n2038:\t\t\t\t size_t *offset0)\n2039:\t{\n2040:\t\tmaxsize = min_t(size_t, min_t(size_t, maxsize, i-\u003ecount), MAX_RW_COUNT);\n2041:\t\tif (!maxsize)\n2042:\t\t\treturn 0;\n2043:\t\n2044:\t\tif (likely(user_backed_iter(i)))\n2045:\t\t\treturn iov_iter_extract_user_pages(i, pages, maxsize,\n2046:\t\t\t\t\t\t\t maxpages, extraction_flags,\n2047:\t\t\t\t\t\t\t offset0);\n2048:\t\tif (iov_iter_is_kvec(i))\n2049:\t\t\treturn iov_iter_extract_kvec_pages(i, pages, maxsize,\n2050:\t\t\t\t\t\t\t maxpages, extraction_flags,\n2051:\t\t\t\t\t\t\t offset0);\n2052:\t\tif (iov_iter_is_bvec(i))\n2053:\t\t\treturn iov_iter_extract_bvec_pages(i, pages, maxsize,\n2054:\t\t\t\t\t\t\t maxpages, extraction_flags,\n2055:\t\t\t\t\t\t\t offset0);\n2056:\t\tif (iov_iter_is_bvecq(i))\n2057:\t\t\treturn iov_iter_extract_bvecq_pages(i, pages, maxsize,\n2058:\t\t\t\t\t\t\t maxpages, extraction_flags,\n2059:\t\t\t\t\t\t\t offset0);\n2060:\t\tif (iov_iter_is_xarray(i))\n2061:\t\t\treturn iov_iter_extract_xarray_pages(i, pages, maxsize,\n2062:\t\t\t\t\t\t\t maxpages, extraction_flags,\n2063:\t\t\t\t\t\t\t offset0);\n2064:\t\treturn -EFAULT;\n2065:\t}\n2066:\tEXPORT_SYMBOL_GPL(iov_iter_extract_pages);\n2067:\t\n2068:\tstatic unsigned int get_contig_folio_len(struct page **pages,\n2069:\t\t\tunsigned int *num_pages, size_t left, size_t offset)\n2070:\t{\n2071:\t\tstruct folio *folio = page_folio(pages[0]);\n2072:\t\tsize_t contig_sz = min_t(size_t, PAGE_SIZE - offset, left);\n"
} |
| 53/3 | 2026/10/06 21:27 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 69817 output: 36 thoughts: 62 Thoughts: Okay, let's break this down. **Diving into `iov_iter_extract_pages` and User-Backed Iterators** Alright, so I'm looking at `iov_iter_extract_pages` and immediately my brain flags that it's taking a `struct page ***pages`. That triple pointer makes me think about memory allocation and how the pages are being handled. Now, I'm particularly interested in how user-backed iterators fit into this. How does `iov_iter_extract_user_pages` interact with that `pages` variable? My gut feeling is it's handling the translation from user-space addresses to actual page structures, but I need to confirm. Let's head over to `lib/iov_iter.c:1950` and see what's actually happening. I need the details to solidify my understanding of the memory model here. |
| 54/3 | 2026/10/06 21:27 | tool | read-file |
0mArgs:{
"File": "lib/iov_iter.c",
"FirstLine": 1950,
"LineCount": 45
}Results: {
"Contents": "1950:\tstatic ssize_t iov_iter_extract_user_pages(struct iov_iter *i,\n1951:\t\t\t\t\t\t struct page ***pages,\n1952:\t\t\t\t\t\t size_t maxsize,\n1953:\t\t\t\t\t\t unsigned int maxpages,\n1954:\t\t\t\t\t\t iov_iter_extraction_t extraction_flags,\n1955:\t\t\t\t\t\t size_t *offset0)\n1956:\t{\n1957:\t\tunsigned long addr;\n1958:\t\tunsigned int gup_flags = 0;\n1959:\t\tsize_t offset;\n1960:\t\tbool will_alloc = !*pages;\n1961:\t\tint res;\n1962:\t\n1963:\t\tif (i-\u003edata_source == ITER_DEST)\n1964:\t\t\tgup_flags |= FOLL_WRITE;\n1965:\t\tif (extraction_flags \u0026 ITER_ALLOW_P2PDMA)\n1966:\t\t\tgup_flags |= FOLL_PCI_P2PDMA;\n1967:\t\tif (i-\u003enofault)\n1968:\t\t\tgup_flags |= FOLL_NOFAULT;\n1969:\t\n1970:\t\taddr = first_iovec_segment(i, \u0026maxsize);\n1971:\t\t*offset0 = offset = addr % PAGE_SIZE;\n1972:\t\taddr \u0026= PAGE_MASK;\n1973:\t\tmaxpages = want_pages_array(pages, maxsize, offset, maxpages);\n1974:\t\tif (!maxpages)\n1975:\t\t\treturn -ENOMEM;\n1976:\t\tres = pin_user_pages_fast(addr, maxpages, gup_flags, *pages);\n1977:\t\tif (unlikely(res \u003c= 0)) {\n1978:\t\t\tif (will_alloc) {\n1979:\t\t\t\tkvfree(*pages);\n1980:\t\t\t\t*pages = NULL;\n1981:\t\t\t}\n1982:\t\t\treturn res;\n1983:\t\t}\n1984:\t\n1985:\t\tmaxsize = min_t(size_t, maxsize, res * PAGE_SIZE - offset);\n1986:\t\tiov_iter_advance(i, maxsize);\n1987:\t\treturn maxsize;\n1988:\t}\n1989:\t\n1990:\t/**\n1991:\t * iov_iter_extract_pages - Extract a list of contiguous pages from an iterator\n1992:\t * @i: The iterator to extract from\n1993:\t * @pages: Where to return the list of pages\n1994:\t * @maxsize: The maximum amount of iterator to extract\n"
} |
| 55/3 | 2026/10/06 21:27 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 70808 output: 49 thoughts: 239 Thoughts: **Analyzing Page Extraction in Network Filesystems** Okay, so I'm digging into how page extraction works, specifically with network filesystems. The code path that's grabbed my attention is within `iov_iter_extract_user_pages`. I see that when `*pages` is supplied (and the conditions align for this, especially given the `pages` pointer's initialization), and when `will_alloc` is *false*, we're relying on `pin_user_pages_fast` to populate `*pages`. This seems like the standard route. But the real question is, how does this mesh with `netfs_extract_iter`? Looking at the loop there, the code snippet uses a loop, calculating `len` and using `bvec_set_page` to move the data. The crucial part here is the `*pages++` increment. I need to make sure I understand the interaction between `pin_user_pages_fast` and this loop. This chunk seems fairly straightforward, though I need to keep in mind the potential for edge cases involving `BUG_ON(slot >= bq->max_slots)` and `max_pages--`. Hold on... I'm missing something. This assumes the `iov_iter` isn't a kvec or bvec. What if the iterator *is* a kernel vector or buffer vector? We need to determine if we are in this case, and if we are, there are likely different code paths in place. I need to quickly consult `iov_iter_extract_bvec_pages` and `iov_iter_extract_kvec_pages` within `lib/iov_iter.c`. I have to be sure I understand the full spectrum of possible iterator types and the corresponding data extraction logic. |
| 56/3 | 2026/10/06 21:28 | tool | grepper |
0mArgs:{
"Expression": "iov_iter_extract_bvec_pages|iov_iter_extract_kvec_pages",
"PathPrefixes": [
"lib/iov_iter.c"
]
}Results: {
"Output": "lib/iov_iter.c=1762=static ssize_t iov_iter_extract_xarray_pages(struct iov_iter *i,\n--\nlib/iov_iter.c-1816- */\nlib/iov_iter.c:1817:static ssize_t iov_iter_extract_bvec_pages(struct iov_iter *i,\nlib/iov_iter.c-1818-\t\t\t\t\t struct page ***pages, size_t maxsize,\n--\nlib/iov_iter.c-1885- */\nlib/iov_iter.c:1886:static ssize_t iov_iter_extract_kvec_pages(struct iov_iter *i,\nlib/iov_iter.c-1887-\t\t\t\t\t struct page ***pages, size_t maxsize,\n--\nlib/iov_iter.c=2033=ssize_t iov_iter_extract_pages(struct iov_iter *i,\n--\nlib/iov_iter.c-2048-\tif (iov_iter_is_kvec(i))\nlib/iov_iter.c:2049:\t\treturn iov_iter_extract_kvec_pages(i, pages, maxsize,\nlib/iov_iter.c-2050-\t\t\t\t\t\t maxpages, extraction_flags,\n--\nlib/iov_iter.c-2052-\tif (iov_iter_is_bvec(i))\nlib/iov_iter.c:2053:\t\treturn iov_iter_extract_bvec_pages(i, pages, maxsize,\nlib/iov_iter.c-2054-\t\t\t\t\t\t maxpages, extraction_flags,\n"
} |
| 57/3 | 2026/10/06 21:28 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 71900 output: 36 thoughts: 43 Thoughts: Here's my summary, as if I were thinking these thoughts: **Analyzing IO Vector Extraction Functions** Alright, let's dive into this code. I need to refresh my memory on how these IO vector extraction functions work. Specifically, I want to take a look at `iov_iter_extract_bvec_pages` and `iov_iter_extract_kvec_pages`. My gut tells me something might be off in the way they handle page extraction, especially given the context of the current debugging effort. I'll focus my attention around lines 1817 and 1886. I'm going to carefully review the logic in those areas to see if any corner cases are being missed or if there's any potential for unexpected behavior during buffer or kernel vector extractions. I'll be looking at how they interact with the page cache, and the overall efficiency of their data access mechanisms. |
| 58/3 | 2026/10/06 21:28 | tool | read-file |
0mArgs:{
"File": "lib/iov_iter.c",
"FirstLine": 1817,
"LineCount": 40
}Results: {
"Contents": "1817:\tstatic ssize_t iov_iter_extract_bvec_pages(struct iov_iter *i,\n1818:\t\t\t\t\t\t struct page ***pages, size_t maxsize,\n1819:\t\t\t\t\t\t unsigned int maxpages,\n1820:\t\t\t\t\t\t iov_iter_extraction_t extraction_flags,\n1821:\t\t\t\t\t\t size_t *offset0)\n1822:\t{\n1823:\t\tsize_t skip = i-\u003eiov_offset, size = 0;\n1824:\t\tstruct bvec_iter bi;\n1825:\t\tint k = 0;\n1826:\t\n1827:\t\tif (i-\u003enr_segs == 0)\n1828:\t\t\treturn 0;\n1829:\t\n1830:\t\tif (i-\u003eiov_offset == i-\u003ebvec-\u003ebv_len) {\n1831:\t\t\ti-\u003eiov_offset = 0;\n1832:\t\t\ti-\u003enr_segs--;\n1833:\t\t\ti-\u003ebvec++;\n1834:\t\t\tskip = 0;\n1835:\t\t}\n1836:\t\tbi.bi_idx = 0;\n1837:\t\tbi.bi_size = maxsize;\n1838:\t\tbi.bi_offset = skip;\n1839:\t\n1840:\t\tmaxpages = want_pages_array(pages, maxsize, skip, maxpages);\n1841:\t\tif (!maxpages)\n1842:\t\t\treturn -ENOMEM;\n1843:\t\n1844:\t\twhile (bi.bi_size \u0026\u0026 bi.bi_idx \u003c i-\u003enr_segs) {\n1845:\t\t\tstruct bio_vec bv = bvec_iter_bvec(i-\u003ebvec, bi);\n1846:\t\n1847:\t\t\t/*\n1848:\t\t\t * The iov_iter_extract_pages interface only allows an offset\n1849:\t\t\t * into the first page. Break out of the loop if we see an\n1850:\t\t\t * offset into subsequent pages, the caller will have to call\n1851:\t\t\t * iov_iter_extract_pages again for the reminder.\n1852:\t\t\t */\n1853:\t\t\tif (k) {\n1854:\t\t\t\tif (bv.bv_offset)\n1855:\t\t\t\t\tbreak;\n1856:\t\t\t} else {\n"
} |
| 59/3 | 2026/10/06 21:28 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 72793 output: 20 thoughts: 1210 Thoughts: **Diving Deep into `bvecq_zero` and Its Uses** Okay, so I'm looking at how `iov_iter_extract_pages` relates to `bvec`, `kvec`, and other similar structures. The underlying principle is clear, but I want to understand the specifics within the `fs/netfs` code. I'm focusing on `bvecq_slice`, `bvecq_zero`, and `bvecq_pos_advance`, all defined within `fs/netfs/bvecq.c`. Let's dig into lines 342-652. First, I'm examining the core of the memory zeroing functions. I want to understand how it operates on `bio_vec` structures, so I'm reading `bvec_zero` and `bvecq_zero` (lines 405 to 506). `bvec_zero` itself looks straightforward enough: it clears a portion of memory described by a `bio_vec`. It correctly handles page boundaries and uses `kmap_local_page` and `kunmap_local` for direct memory access. `bvecq_zero` is the more interesting one, as it zeroes memory fragments described by a `bvecq` queue. It iterates through the `bvecq` chain, acquiring slots and zeroing the corresponding memory regions using `bvec_zero`. The code appears robust, handling cases where the position hits the end of the queue. Now, the crucial question: *where* is `bvecq_zero` actually used? I've got to follow the call chain to understand the context. After a quick scan, I've found it in `fs/netfs/buffered_read.c` specifically in the `netfs_issue_read` function. It seems to be used when the source of a read is a 'default' source (e.g., read hole, EOF, or zero fill). In these scenarios, `bvecq_zero` zeroes the *destination* buffer. This is key. The `subreq->len` specifies how many bytes to be zeroed. I'm checking that the implementation does the zeroing correctly, specifically within `bvecq_zero`. The snippet: `part = min(bv->bv_len - offset, amount)` is the one to inspect. I'm looking for potential off-by-one errors or memory corruption. The code correctly limits the amount to be zeroed to the remaining portion of the `bio_vec`. Incrementing `offset` and decrementing `amount` looks correct. It iterates across multiple `bio_vec` entries if needed. I'm making sure that the page pointer is valid, so, it looks like `bv->bv_page` is valid in filled slots. Finally, I'm checking page handling when dealing with large pages or folios. In `bvec_zero`, it correctly calculates `page += offset / PAGE_SIZE`. This suggests it's prepared to handle multi-page folios which is great. I should probably review `bvec_set_folio` to see how those are set up. |
| 60/3 | 2026/10/06 21:28 | tool | grepper |
0mArgs:{
"Expression": "bvec_set_folio"
}Results: {
"Output": "block/bio.c=1423=static int bio_iov_iter_bounce_read(struct bio *bio, struct iov_iter *iter,\n--\nblock/bio.c-1451-\t */\nblock/bio.c:1452:\tbvec_set_folio(\u0026bio-\u003ebi_io_vec[0], folio, bio-\u003ebi_iter.bi_size, 0);\nblock/bio.c-1453-\tif (iov_iter_extract_will_pin(iter))\n--\nfs/netfs/buffered_read.c=517=static int netfs_create_singular_buffer(struct netfs_io_request *rreq, struct folio *folio)\n--\nfs/netfs/buffered_read.c-529-\nfs/netfs/buffered_read.c:530:\tbvec_set_folio(\u0026bq-\u003ebv[0], folio, fsize, 0);\nfs/netfs/buffered_read.c-531-\tbvecq_filled_to(bq, 1);\n--\nfs/netfs/buffered_read.c=541=static int netfs_read_gaps(struct file *file, struct folio *folio)\n--\nfs/netfs/buffered_read.c-587-\t\tfolio_get(folio);\nfs/netfs/buffered_read.c:588:\t\tbvec_set_folio(\u0026bq-\u003ebv[i++], folio, from, 0);\nfs/netfs/buffered_read.c-589-\t\toff = from;\n--\nfs/netfs/buffered_read.c-601-\t\tpart = min(to - off, PAGE_SIZE);\nfs/netfs/buffered_read.c:602:\t\tbvec_set_folio(\u0026bq-\u003ebv[i++], sink, part, 0);\nfs/netfs/buffered_read.c-603-\t\toff += part;\n--\nfs/netfs/buffered_read.c-611-\t\tfolio_get(folio);\nfs/netfs/buffered_read.c:612:\t\tbvec_set_folio(\u0026bq-\u003ebv[i++], folio, flen - to, to);\nfs/netfs/buffered_read.c-613-\t}\n--\nfs/netfs/bvecq.c=606=ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl)\n--\nfs/netfs/bvecq.c-631-\t\tlen = folio_size(folio);\nfs/netfs/bvecq.c:632:\t\tbvec_set_folio(\u0026bq-\u003ebv[slot], folio, len, 0);\nfs/netfs/bvecq.c-633-\t\tloaded += len;\n--\nfs/netfs/read_pgpriv2.c=19=static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio *folio)\n--\nfs/netfs/read_pgpriv2.c-74-\tslot = queue-\u003enr_slots;\nfs/netfs/read_pgpriv2.c:75:\tbvec_set_folio(\u0026queue-\u003ebv[slot], folio, fsize, 0);\nfs/netfs/read_pgpriv2.c-76-\ttrace_netfs_bv_slot(queue, slot);\n--\nfs/netfs/write_issue.c=328=static int netfs_write_folio(struct netfs_io_request *wreq,\n--\nfs/netfs/write_issue.c-466-\tslot = queue-\u003enr_slots;\nfs/netfs/write_issue.c:467:\tbvec_set_folio(\u0026queue-\u003ebv[slot], folio, fsize, 0);\nfs/netfs/write_issue.c-468-\ttrace_netfs_bv_slot(queue, slot);\n--\nfs/orangefs/inode.c=19=static int orangefs_writepage_locked(struct folio *folio,\n--\nfs/orangefs/inode.c-51-\tWARN_ON(wlen == 0);\nfs/orangefs/inode.c:52:\tbvec_set_folio(\u0026bv, folio, wlen, offset_in_folio(folio, off));\nfs/orangefs/inode.c-53-\tiov_iter_bvec(\u0026iter, ITER_SOURCE, \u0026bv, 1, wlen);\n--\nfs/orangefs/inode.c=78=static int orangefs_writepages_work(struct orangefs_writepages *ow,\n--\nfs/orangefs/inode.c-93-\t\tfolio_start_writeback(ow-\u003efolios[i]);\nfs/orangefs/inode.c:94:\t\tbvec_set_folio(\u0026ow-\u003ebv[i], ow-\u003efolios[i],\nfs/orangefs/inode.c-95-\t\t\t\tfolio_size(ow-\u003efolios[i]) - start, start);\n--\nfs/orangefs/inode.c=279=static int orangefs_read_folio(struct file *file, struct folio *folio)\n--\nfs/orangefs/inode.c-290-\toff = folio_pos(folio);\nfs/orangefs/inode.c:291:\tbvec_set_folio(\u0026bv, folio, folio_size(folio), 0);\nfs/orangefs/inode.c-292-\tiov_iter_bvec(\u0026iter, ITER_DEST, \u0026bv, 1, folio_size(folio));\n--\ninclude/linux/bvec.h=41=static inline void bvec_set_page(struct bio_vec *bv, struct page *page,\n--\ninclude/linux/bvec.h-49-/**\ninclude/linux/bvec.h:50: * bvec_set_folio - initialize a bvec based off a struct folio\ninclude/linux/bvec.h-51- * @bv:\t\tbvec to initialize\n--\ninclude/linux/bvec.h-55- */\ninclude/linux/bvec.h:56:static inline void bvec_set_folio(struct bio_vec *bv, struct folio *folio,\ninclude/linux/bvec.h-57-\t\tsize_t len, size_t offset)\n--\nmm/page_io.c=332=static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)\n--\nmm/page_io.c-350-\t}\nmm/page_io.c:351:\tbvec_set_folio(\u0026sio-\u003ebvecs[sio-\u003enr_bvecs], folio, folio_size(folio), 0);\nmm/page_io.c-352-\tsio-\u003elen += folio_size(folio);\n"
} |
| 61/3 | 2026/10/06 21:28 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 76377 output: 35 thoughts: 17 Thoughts: Okay, here's my interpretation of that request, formatted as requested: **Navigating the bvec.h Waters** Alright, let's dive in. The task at hand is to understand the context. My immediate focus is on `include/linux/bvec.h:56`. Right away, the syntax "Let's see" tells me that I'm supposed to review some documentation. My instinct is to understand the `bvec.h` header file. It is a fundamental building block of block I/O in the Linux kernel. Therefore, I'm expecting to find structures and functions related to `bio_vec`. I'll immediately start to think about how these structures are being used in more abstract areas. I need to understand what this `bio_vec` actually _represents_ and how it interacts with the underlying block device model. Line 56... I anticipate finding a structure definition, macro, or perhaps an inline function related to managing scatter-gather lists for block device I/O. I need to consider whether it describes a single entry in a scatter-gather list, or perhaps the start of the overall structure that would hold multiple entries. I'll need to figure out what data is held in that `bio_vec` structure – presumably, it's things like the page pointer, offset within the page, and the length of the data segment. I need to keep in mind possible optimizations like direct access, or the need to use DMA. |
| 62/3 | 2026/10/06 21:28 | tool | read-file |
0mArgs:{
"File": "include/linux/bvec.h",
"FirstLine": 40,
"LineCount": 25
}Results: {
"Contents": " 40:\t */\n 41:\tstatic inline void bvec_set_page(struct bio_vec *bv, struct page *page,\n 42:\t\t\tunsigned int len, unsigned int offset)\n 43:\t{\n 44:\t\tbv-\u003ebv_page = page;\n 45:\t\tbv-\u003ebv_len = len;\n 46:\t\tbv-\u003ebv_offset = offset;\n 47:\t}\n 48:\t\n 49:\t/**\n 50:\t * bvec_set_folio - initialize a bvec based off a struct folio\n 51:\t * @bv:\t\tbvec to initialize\n 52:\t * @folio:\tfolio the bvec should point to\n 53:\t * @len:\tlength of the bvec\n 54:\t * @offset:\toffset into the folio\n 55:\t */\n 56:\tstatic inline void bvec_set_folio(struct bio_vec *bv, struct folio *folio,\n 57:\t\t\tsize_t len, size_t offset)\n 58:\t{\n 59:\t\tunsigned long nr = offset / PAGE_SIZE;\n 60:\t\n 61:\t\tWARN_ON_ONCE(len \u003e UINT_MAX);\n 62:\t\tbvec_set_page(bv, folio_page(folio, nr), len, offset % PAGE_SIZE);\n 63:\t}\n 64:\t\n"
} |
| 63/3 | 2026/10/06 21:28 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 77100 output: 19 thoughts: 638 Thoughts: **Analyzing Zeroing Logic in Kernel File Operations**
Okay, so I'm looking at how zeroing is handled in the kernel, specifically focusing on `bvecq.c` and `netfs_read_gaps`. First, the `bvec_zero` function seems straightforward enough. It iterates page-by-page, calculating the appropriate offsets, and using `memset` to zero out the memory. Standard, page-aligned zeroing, which is exactly what I'd expect.
Now, let's dig into `netfs_read_gaps`. This is where it gets interesting. I'm focusing on this snippet:
```c
ret = netfs_wait_for_read(rreq);
if (ret >= 0) {
if (ret < flen) {
- struct iov_iter iter;
-
- iov_iter_bvec(&iter, ITER_DEST, bvec, i, flen);
- iov_iter_advance(&iter, ret);
- iov_iter_zero(flen - ret, &iter);
+ if (ret < from)
+ folio_zero_segments(folio, ret, from, to, flen);
+ else
+ folio_zero_segment(folio, max(to, ret), flen);
}
```
The thing that immediately jumps out at me is `from` and `to`. These are defined as `finfo->dirty_offset` and `finfo->dirty_offset + finfo->dirty_len`, respectively, which makes this about the "dirty" part of a folio. Okay, so we're dealing with a range of a folio that's been modified, and we're dealing with read gaps *around* that dirty portion.
Before this patch, it looks like `bvec` was structured with this in mind: we had an initial folio, the 'dirty' folios, and the final folio. The original code was using `iov_iter_zero` to zero memory *after* a read operation. It looks like it was zeroing portions of the folios *outside* the read data. It appears that it zeroes the first folio from read position up to `from` offset, zeroing the "dirty" folios and the final folio from `to` to `flen`.
The new code replaces this with `folio_zero_segments`. My intuition tells me that this is meant to be more efficient, especially in this context, but I need to understand what it does. So, I need to look up `folio_zero_segments` and get a handle on what segments are being targeted.
|
| 64/3 | 2026/10/06 21:28 | tool | grepper |
0mArgs:{
"Expression": "folio_zero_segments"
}Results: {
"Output": "fs/buffer.c=1995=int __block_write_begin_int(struct folio *folio, loff_t pos, unsigned len,\n--\nfs/buffer.c-2044-\t\t\t\tif (block_end \u003e to || block_start \u003c from)\nfs/buffer.c:2045:\t\t\t\t\tfolio_zero_segments(folio,\nfs/buffer.c-2046-\t\t\t\t\t\tto, block_end,\n--\nfs/ext4/inode.c=1169=int ext4_block_write_begin(handle_t *handle, struct folio *folio,\n--\nfs/ext4/inode.c-1233-\t\t\t\tif (block_end \u003e to || block_start \u003c from)\nfs/ext4/inode.c:1234:\t\t\t\t\tfolio_zero_segments(folio, to,\nfs/ext4/inode.c-1235-\t\t\t\t\t\t\t block_end,\n--\nfs/fuse/dev.c=1254=int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop,\n--\nfs/fuse/dev.c-1273-\t\t\tif (cs-\u003eskip_folio_copy)\nfs/fuse/dev.c:1274:\t\t\t\tfolio_zero_segments(folio, 0, offset,\nfs/fuse/dev.c-1275-\t\t\t\t\t\t offset + count, size);\n--\nfs/iomap/buffered-io.c=875=static int __iomap_write_begin(const struct iomap_iter *iter,\n--\nfs/iomap/buffered-io.c-922-\t\t\t\treturn -EIO;\nfs/iomap/buffered-io.c:923:\t\t\tfolio_zero_segments(folio, poff, from, to, poff + plen);\nfs/iomap/buffered-io.c-924-\t\t} else {\n--\nfs/libfs.c=943=int simple_write_begin(const struct kiocb *iocb, struct address_space *mapping,\n--\nfs/libfs.c-958-\nfs/libfs.c:959:\t\tfolio_zero_segments(folio, 0, from,\nfs/libfs.c-960-\t\t\t\tfrom + len, folio_size(folio));\n--\nfs/netfs/buffered_read.c=541=static int netfs_read_gaps(struct file *file, struct folio *folio)\n--\nfs/netfs/buffered_read.c-623-\t\t\tif (ret \u003c from)\nfs/netfs/buffered_read.c:624:\t\t\t\tfolio_zero_segments(folio, ret, from, to, flen);\nfs/netfs/buffered_read.c-625-\t\t\telse\n--\nfs/netfs/buffered_read.c=727=static bool netfs_skip_folio_read(struct folio *folio, uoff_t pos, size_t len,\n--\nfs/netfs/buffered_read.c-756-zero_out:\nfs/netfs/buffered_read.c:757:\tfolio_zero_segments(folio, 0, offset, offset + len, plen);\nfs/netfs/buffered_read.c-758-\treturn true;\n--\nfs/nfs/file.c=431=static int nfs_write_end(const struct kiocb *iocb,\n--\nfs/nfs/file.c-454-\t\tif (pglen == 0) {\nfs/nfs/file.c:455:\t\t\tfolio_zero_segments(folio, 0, offset, end, fsize);\nfs/nfs/file.c-456-\t\t\tfolio_mark_uptodate(folio);\n--\nfs/ntfs/iomap.c=15=static void ntfs_iomap_put_folio_non_resident(struct inode *inode, loff_t pos,\n--\nfs/ntfs/iomap.c-31-\t\t\t\toffset = folio_size(folio);\nfs/ntfs/iomap.c:32:\t\t\tfolio_zero_segments(folio,\nfs/ntfs/iomap.c-33-\t\t\t\t\t offset_in_folio(folio, init),\n--\nfs/ntfs/iomap.c-54-\t\t\toffset2 = folio_size(folio);\nfs/ntfs/iomap.c:55:\t\tfolio_zero_segments(folio,\nfs/ntfs/iomap.c-56-\t\t\t\t offset,\n--\ninclude/linux/highmem.h=714=static inline size_t memcpy_from_file_folio(char *to, struct folio *folio,\n--\ninclude/linux/highmem.h-732-/**\ninclude/linux/highmem.h:733: * folio_zero_segments() - Zero two byte ranges in a folio.\ninclude/linux/highmem.h-734- * @folio: The folio to write to.\n--\ninclude/linux/highmem.h-739- */\ninclude/linux/highmem.h:740:static inline void folio_zero_segments(struct folio *folio,\ninclude/linux/highmem.h-741-\t\tsize_t start1, size_t xend1, size_t start2, size_t xend2)\n--\nmm/shmem.c=3283=shmem_write_end(const struct kiocb *iocb, struct address_space *mapping,\n--\nmm/shmem.c-3294-\t\t\tsize_t from = offset_in_folio(folio, pos);\nmm/shmem.c:3295:\t\t\tfolio_zero_segments(folio, 0, from,\nmm/shmem.c-3296-\t\t\t\t\tfrom + copied, folio_size(folio));\n"
} |
| 65/3 | 2026/10/06 21:28 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 79773 output: 28 thoughts: 1215 Thoughts: **Analyzing Potential Uninitialized Memory Writes in AFS File System** Okay, so I'm digging into this code, specifically around how the AFS file system handles writeback operations, and I'm seeing something that's making me a little uneasy. The core of it seems to be in how memory is zeroed and how data is prepared for writing to the cache, with some new changes that might expose uninitialized memory. Initially, `folio_zero_segments` handles zeroing memory, and it looks like a clever approach to avoid unnecessary work by only zeroing specific parts of a folio, especially the non-dirty regions. Seems solid. But then, I jumped into `fs/afs/dir.c` and `fs/afs/symlink.c`, and I found these `afs_dir_writepages` and `afs_symlink_writepages` functions. The key point is this use of `netfs_writeback_single`. Before, they passed the inode size via `i_size_read`. Now, there's a size parameter explicitly passed to `netfs_writeback_single`, the `PAGE_SIZE` in the case of symlinks and `dvnode->directory_size` for directories. But why the change? The key appears to be in how `netfs_writeback_single` interacts with `iov_iter_bvec_queue` and the `cachefiles` subsystem. It rounds up the write length (`len`) to the cache's direct I/O (DIO) size using `round_up(len, wreq->cache_resources.dio_size)`. This becomes `clen`. So if the DIO size is larger than the data to be written (and it often is), this will increase the amount of data written to the cache. And if `clen` exceeds the `iov_iter_count(iter)`, an EIO error is returned. The question is, why was the code changed to use the `PAGE_SIZE` or `dvnode->directory_size` when passing the length to `netfs_writeback_single`? Before, the code used `i_size_read()`, meaning only the relevant part of the memory region was written. Now, it looks like `clen`, and therefore more memory, is written. That's when it clicked! For symlinks, `vnode->directory` is a full page (PAGE_SIZE), but `i_size` might be shorter (e.g., 20 bytes for the path string). This means that if the DIO size is larger than the actual symlink target size, `clen` could become larger than the valid portion of the page. This means that *uninitialized* memory might be written to the cache file. This could potentially cause a security issue if the cached data is later read. I need to confirm if `vnode->directory` in symlinks is initialized. I need to check how `vnode->directory` is allocated and initialized in `fs/afs/dir.c` and `fs/afs/symlink.c`! I must find out if the memory is zeroed or initialized to prevent leakage of uninitialized data in the cached file. |
| 66/3 | 2026/10/06 21:29 | tool | grepper |
0mArgs:{
"Expression": "vnode-\u003edirectory",
"PathPrefixes": [
"fs/afs"
]
}Results: {
"Output": "fs/afs/dir.c=135=static void afs_dir_dump(struct afs_vnode *dvnode)\n--\nfs/afs/dir.c-142-\nfs/afs/dir.c:143:\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0, i_size);\nfs/afs/dir.c-144-\titerate_bvecq(\u0026iter, iov_iter_count(\u0026iter), NULL, NULL,\n--\nfs/afs/dir.c=197=static int afs_dir_check(struct afs_vnode *dvnode)\n--\nfs/afs/dir.c-205-\nfs/afs/dir.c:206:\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0, i_size);\nfs/afs/dir.c-207-\tchecked = iterate_bvecq(\u0026iter, iov_iter_count(\u0026iter), dvnode, NULL,\n--\nfs/afs/dir.c=235=static ssize_t afs_do_read_single(struct afs_vnode *dvnode, struct file *file)\n--\nfs/afs/dir.c-249-\t/* Expand the storage. TODO: Shrink the storage too. */\nfs/afs/dir.c:250:\tif (dvnode-\u003edirectory_size \u003c i_size) {\nfs/afs/dir.c:251:\t\tsize_t cur_size = dvnode-\u003edirectory_size;\nfs/afs/dir.c-252-\nfs/afs/dir.c:253:\t\tret = bvecq_expand_buffer(\u0026dvnode-\u003edirectory, \u0026cur_size,\nfs/afs/dir.c-254-\t\t\t\t\t round_up(i_size, PAGE_SIZE),\nfs/afs/dir.c-255-\t\t\t\t\t mapping_gfp_mask(dvnode-\u003enetfs.inode.i_mapping));\nfs/afs/dir.c:256:\t\tdvnode-\u003edirectory_size = cur_size;\nfs/afs/dir.c-257-\t\tif (ret \u003c 0)\n--\nfs/afs/dir.c-260-\nfs/afs/dir.c:261:\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, dvnode-\u003edirectory, 0, 0, dvnode-\u003edirectory_size);\nfs/afs/dir.c-262-\n--\nfs/afs/dir.c=488=static int afs_dir_iterate_contents(struct inode *dir, struct dir_context *dir_ctx)\n--\nfs/afs/dir.c-500-\nfs/afs/dir.c:501:\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0, i_size);\nfs/afs/dir.c-502-\tiov_iter_advance(\u0026iter, round_down(dir_ctx-\u003epos, AFS_DIR_BLOCK_SIZE));\n--\nfs/afs/dir.c=2211=static int afs_dir_writepages(struct address_space *mapping,\n--\nfs/afs/dir.c-2230-\tif (test_bit(AFS_VNODE_DIR_VALID, \u0026dvnode-\u003eflags)) {\nfs/afs/dir.c:2231:\t\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0,\nfs/afs/dir.c:2232:\t\t\t\t dvnode-\u003edirectory_size);\nfs/afs/dir.c-2233-\t\tret = netfs_writeback_single(mapping, wbc, \u0026iter,\n--\nfs/afs/dir_edit.c=110=static union afs_xdr_dir_block *afs_dir_get_block(struct afs_dir_iter *iter, size_t block)\n--\nfs/afs/dir_edit.c-117-\nfs/afs/dir_edit.c:118:\tif (dvnode-\u003edirectory_size \u003c blend) {\nfs/afs/dir_edit.c:119:\t\tsize_t cur_size = dvnode-\u003edirectory_size;\nfs/afs/dir_edit.c-120-\nfs/afs/dir_edit.c:121:\t\tret = bvecq_expand_buffer(\u0026dvnode-\u003edirectory, \u0026cur_size, blend,\nfs/afs/dir_edit.c-122-\t\t\t\t\t mapping_gfp_mask(dvnode-\u003enetfs.inode.i_mapping));\nfs/afs/dir_edit.c:123:\t\tdvnode-\u003edirectory_size = cur_size;\nfs/afs/dir_edit.c-124-\t\tif (ret \u003c 0)\n--\nfs/afs/dir_edit.c-129-\tif (!bq)\nfs/afs/dir_edit.c:130:\t\tbq = dvnode-\u003edirectory;\nfs/afs/dir_edit.c-131-\n--\nfs/afs/dir_search.c=67=union afs_xdr_dir_block *afs_dir_find_block(struct afs_dir_iter *iter, size_t block)\n--\nfs/afs/dir_search.c-78-\nfs/afs/dir_search.c:79:\tif (dvnode-\u003edirectory_size \u003c blend)\nfs/afs/dir_search.c-80-\t\tgoto fail;\n--\nfs/afs/dir_search.c-82-\tif (!bq || blpos \u003c fpos) {\nfs/afs/dir_search.c:83:\t\tbq = dvnode-\u003edirectory;\nfs/afs/dir_search.c-84-\t\tslot = 0;\n--\nfs/afs/inode.c=655=void afs_evict_inode(struct inode *inode)\n--\nfs/afs/inode.c-685-\ttruncate_inode_pages_final(\u0026inode-\u003ei_data);\nfs/afs/inode.c:686:\tbvecq_put(vnode-\u003edirectory);\nfs/afs/inode.c-687-\tif (vnode-\u003esymlink)\n--\nfs/afs/super.c=670=static struct inode *afs_alloc_inode(struct super_block *sb)\n--\nfs/afs/super.c-687-\tvnode-\u003epermit_cache\t= NULL;\nfs/afs/super.c:688:\tvnode-\u003edirectory\t= NULL;\nfs/afs/super.c:689:\tvnode-\u003edirectory_size\t= 0;\nfs/afs/super.c-690-\n--\nfs/afs/symlink.c=56=void afs_init_new_symlink(struct afs_vnode *vnode, struct afs_operation *op)\n--\nfs/afs/symlink.c-67-\nfs/afs/symlink.c:68:\tvnode-\u003edirectory =\nfs/afs/symlink.c-69-\t\tbvecq_alloc_buffer(PAGE_SIZE,\n--\nfs/afs/symlink.c-71-\t\t\t\t false);\nfs/afs/symlink.c:72:\tif (!vnode-\u003edirectory)\nfs/afs/symlink.c-73-\t\treturn;\nfs/afs/symlink.c-74-\nfs/afs/symlink.c:75:\tvnode-\u003edirectory_size = size;\nfs/afs/symlink.c:76:\tp = bvec_kmap_partial(\u0026vnode-\u003edirectory-\u003ebv[0], 0);\nfs/afs/symlink.c-77-\tmemcpy(p, symlink-\u003econtent, size);\n--\nfs/afs/symlink.c=87=static ssize_t afs_do_read_symlink(struct afs_vnode *vnode)\n--\nfs/afs/symlink.c-99-\nfs/afs/symlink.c:100:\tif (!vnode-\u003edirectory) {\nfs/afs/symlink.c:101:\t\tvnode-\u003edirectory = bvecq_alloc_buffer(\nfs/afs/symlink.c-102-\t\t\tPAGE_SIZE, mapping_gfp_mask(vnode-\u003enetfs.inode.i_mapping), false);\nfs/afs/symlink.c:103:\t\tif (!vnode-\u003edirectory)\nfs/afs/symlink.c-104-\t\t\treturn -ENOMEM;\n--\nfs/afs/symlink.c-106-\nfs/afs/symlink.c:107:\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, vnode-\u003edirectory, 0, 0, PAGE_SIZE);\nfs/afs/symlink.c-108-\n--\nfs/afs/symlink.c-118-\t\t}\nfs/afs/symlink.c:119:\t\tvnode-\u003edirectory_size = i_size;\nfs/afs/symlink.c-120-\n--\nfs/afs/symlink.c-128-\nfs/afs/symlink.c:129:\t\tconst char *s = bvec_kmap_partial(\u0026vnode-\u003edirectory-\u003ebv[0], 0);\nfs/afs/symlink.c-130-\n--\nfs/afs/symlink.c-137-\tif (!fscache_cookie_enabled(netfs_i_cookie(\u0026vnode-\u003enetfs))) {\nfs/afs/symlink.c:138:\t\tbvecq_put(vnode-\u003edirectory);\nfs/afs/symlink.c:139:\t\tvnode-\u003edirectory = NULL;\nfs/afs/symlink.c:140:\t\tvnode-\u003edirectory_size = 0;\nfs/afs/symlink.c-141-\t}\n--\nfs/afs/symlink.c=233=int afs_symlink_writepages(struct address_space *mapping,\n--\nfs/afs/symlink.c-248-\nfs/afs/symlink.c:249:\tif (vnode-\u003edirectory \u0026\u0026\nfs/afs/symlink.c-250-\t atomic64_read(\u0026vnode-\u003ecb_expires_at) != AFS_NO_CB_PROMISE) {\nfs/afs/symlink.c:251:\t\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, vnode-\u003edirectory, 0, 0, PAGE_SIZE);\nfs/afs/symlink.c-252-\t\tret = netfs_writeback_single(mapping, wbc, \u0026iter,\n--\nfs/afs/symlink.c-257-\t\tnetfs_wb_begin(\u0026vnode-\u003enetfs, false);\nfs/afs/symlink.c:258:\t\tbvecq_put(vnode-\u003edirectory);\nfs/afs/symlink.c:259:\t\tvnode-\u003edirectory = NULL;\nfs/afs/symlink.c:260:\t\tvnode-\u003edirectory_size = 0;\nfs/afs/symlink.c-261-\t\tnetfs_wb_end(\u0026vnode-\u003enetfs);\n"
} |
| 67/3 | 2026/10/06 21:29 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 84468 output: 44 thoughts: 35 Thoughts: Here's a summary of my thought process: **Diving into `bvecq_alloc_buffer`** Okay, let's dissect this. My focus is immediately drawn to that call to `bvecq_alloc_buffer(PAGE_SIZE, ...)` . I need to understand the memory allocation strategy here. I’m thinking the first logical step is to dive into the definition of this function. I need to get a concrete understanding of what's happening at the allocation level. What's the allocation size? Is this using the standard kernel allocators, or is there a custom allocator in play? Knowing the mechanics of `bvecq_alloc_buffer` is going to be key to understanding the larger system. I’m especially interested in whether this allocates a page aligned buffer, which is likely for DMA. |
| 68/3 | 2026/10/06 21:29 | tool | grepper |
0mArgs:{
"Expression": "bvecq_alloc_buffer",
"PathPrefixes": [
"include/linux/bvecq.h",
"fs/netfs"
]
}Results: {
"Output": "fs/netfs/bvecq.c=142=EXPORT_SYMBOL(bvecq_alloc_chain);\n--\nfs/netfs/bvecq.c-144-/**\nfs/netfs/bvecq.c:145: * bvecq_alloc_buffer2 - Allocate a bvecq chain and populate with buffers\nfs/netfs/bvecq.c-146- * @size: Target size of the buffer (can be 0 for an empty buffer)\n--\nfs/netfs/bvecq.c-161- */\nfs/netfs/bvecq.c:162:struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp,\nfs/netfs/bvecq.c-163-\t\t\t\t bool for_writeback)\n--\nfs/netfs/bvecq.c-225-}\nfs/netfs/bvecq.c:226:EXPORT_SYMBOL(bvecq_alloc_buffer2);\nfs/netfs/bvecq.c-227-\n--\ninclude/linux/bvecq.h=69=struct bvecq *bvecq_alloc_chain(size_t nr_slots, gfp_t gfp, bool for_writeback);\ninclude/linux/bvecq.h:70:struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp,\ninclude/linux/bvecq.h-71-\t\t\t\t bool for_writeback);\n--\ninclude/linux/bvecq.h=80=ssize_t bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl);\n--\ninclude/linux/bvecq.h-82-/**\ninclude/linux/bvecq.h:83: * bvecq_alloc_buffer - Allocate a bvecq chain and populate with buffers\ninclude/linux/bvecq.h-84- * @size: Target size of the buffer (can be 0 for an empty buffer)\n--\ninclude/linux/bvecq.h-87- *\ninclude/linux/bvecq.h:88: * Wrapper around %bvecq_alloc_buffer2().\ninclude/linux/bvecq.h-89- */\ninclude/linux/bvecq.h:90:static inline struct bvecq *bvecq_alloc_buffer(size_t size, gfp_t gfp, bool for_writeback)\ninclude/linux/bvecq.h-91-{\ninclude/linux/bvecq.h:92:\treturn bvecq_alloc_buffer2(size, 0, gfp, for_writeback);\ninclude/linux/bvecq.h-93-}\n"
} |
| 69/3 | 2026/10/06 21:29 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 85426 output: 38 thoughts: 27 Thoughts: Okay, let's dive into this. Here's my take on examining that code: **Analyzing `bvecq_alloc_buffer2`** Alright, time to crack open `fs/netfs/bvecq.c` and take a look at `bvecq_alloc_buffer2`. From the name alone, I'm expecting this function to be involved in allocating buffers for `bvecq` structures, likely the second stage or a specialized allocation function. Given the '2' in the name, I'd anticipate it builds upon an earlier allocation step, maybe handling more complex setup or initialization beyond a simple `kmalloc`. I'll be looking to see how it interacts with the underlying memory management, any error handling, and how it integrates with the overall `bvecq` data structure. I'm keen to see if there are any specific design choices related to the network filesystem context that influence this allocation. I'll be paying close attention to potential race conditions or locking mechanisms, especially given it's likely dealing with shared resources in a networking environment. I expect this will be a pretty straightforward function, but I am keeping my eyes open for anything out of the ordinary, and how it interacts with the other functions in the module. |
| 70/3 | 2026/10/06 21:29 | tool | read-file |
0mArgs:{
"File": "fs/netfs/bvecq.c",
"FirstLine": 162,
"LineCount": 55
}Results: {
"Contents": " 162:\tstruct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp,\n 163:\t\t\t\t\t bool for_writeback)\n 164:\t{\n 165:\t\tstruct bvecq *head = NULL, *p = NULL;\n 166:\t\tsize_t nr_per_bq = BVECQ_POOL_SLOTS;\n 167:\t\tsize_t count = pre_slots + DIV_ROUND_UP(size, PAGE_SIZE);\n 168:\t\n 169:\t\t_enter(\"%zx,%zx,%u\", size, count, pre_slots);\n 170:\t\n 171:\t\tif (WARN_ON_ONCE(pre_slots \u003e nr_per_bq))\n 172:\t\t\treturn NULL;\n 173:\t\n 174:\t\thead = bvecq_alloc_chain(count, gfp, for_writeback);\n 175:\t\tif (!head)\n 176:\t\t\treturn NULL;\n 177:\t\n 178:\t\tp = head;\n 179:\t\tdo {\n 180:\t\t\tstruct page **pages;\n 181:\t\t\tsize_t unused, want, got, slot;\n 182:\t\n 183:\t\t\tif (!count)\n 184:\t\t\t\tbreak;\n 185:\t\t\tif (WARN_ON_ONCE(!p))\n 186:\t\t\t\tgoto oom;\n 187:\t\n 188:\t\t\tif (p-\u003enr_slots == 0) {\n 189:\t\t\t\t/* Need to clear pre slots and pages[], so just clear all. */\n 190:\t\t\t\tmemset(p-\u003ebv, 0, p-\u003emax_slots * sizeof(p-\u003ebv[0]));\n 191:\t\t\t\tp-\u003emem_type = BVECQ_MEM_ALLOCED;\n 192:\t\t\t\tp-\u003enr_slots = pre_slots;\n 193:\t\t\t\tcount -= pre_slots;\n 194:\t\t\t\tpre_slots = 0;\n 195:\t\t\t\tif (!count)\n 196:\t\t\t\t\tbreak;\n 197:\t\t\t}\n 198:\t\n 199:\t\t\tif (p-\u003enr_slots \u003e= p-\u003emax_slots) {\n 200:\t\t\t\tp = p-\u003enext;\n 201:\t\t\t\tcontinue;\n 202:\t\t\t}\n 203:\t\t\tunused = p-\u003emax_slots - p-\u003enr_slots;\n 204:\t\n 205:\t\t\tpages = (struct page **)\u0026p-\u003ebv[p-\u003emax_slots];\n 206:\t\t\tpages -= unused;\n 207:\t\n 208:\t\t\twant = min(count, unused);\n 209:\t\t\tgot = alloc_pages_bulk(gfp, want, pages);\n 210:\t\t\tif (!got)\n 211:\t\t\t\tgoto oom;\n 212:\t\n 213:\t\t\tslot = p-\u003enr_slots;\n 214:\t\t\tfor (int i = 0; i \u003c got; i++)\n 215:\t\t\t\tbvec_set_page(\u0026p-\u003ebv[slot++], pages[i], PAGE_SIZE, 0);\n 216:\t\n"
} |
| 71/3 | 2026/10/06 21:29 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 86586 output: 35 thoughts: 622 Thoughts: **Analyzing AFS Symlink Cache Behavior** Okay, let's break this down. My initial observation is focused on how AFS symlinks are handled, specifically the buffer allocation and writeback processes. I see that `bvecq_alloc_buffer2` is used, and it calls `alloc_pages_bulk` followed by a `memset` to zero out the buffer slots. This immediately makes me think about potential inefficiencies and how data is actually stored. Then I jump to `afs_symlink.c` and notice the allocation of a full `PAGE_SIZE` buffer for `vnode->directory` using `bvecq_alloc_buffer`. Now, the writeback process in `afs_symlink_writepages` really catches my eye. It uses an `iov_iter_bvec_queue` with `PAGE_SIZE`, but the length argument passed to `netfs_writeback_single` is the inode size, *not* `PAGE_SIZE`. My brain is now trying to reconcile the allocated size with the actual data being written back. This immediately triggers a question: What size data is really being handled? If `dio_size` is, say, 512 or 4096, it's getting rounded up, and so the write operation could be handling more data than necessary. This seems potentially wasteful in terms of both memory and I/O. So, I have to find where `vnode->directory` is written to. I look at `afs_init_new_symlink` and see that `vnode->directory_size` is set, and a `memcpy` copies the symlink's content into the `vnode->directory` buffer. This reinforces my concern: what's in the *rest* of `vnode->directory` after the copy? Is it garbage, or is it zeroed? And if it's zeroed, is it done explicitly, or implicitly by `alloc_pages_bulk`? My intuition says it must be zeroed, but I need to confirm. So, I go back to `bvecq_alloc_buffer2` and try to remember exactly how `alloc_pages_bulk` works. Does it zero the pages? Or does it need the subsequent `memset` to be safe? I need to know the allocation flags used, either `mapping_gfp_mask` or `gfp` which I'll have to investigate further. Finally, I need to go back to `fs/afs/symlink.c` and review lines 65-80, as this is the critical section for initializing the symlink's directory buffer. This is where I need to confirm the zeroing behavior and the initial data size. |
| 72/3 | 2026/10/06 21:29 | tool | read-file |
0mArgs:{
"File": "fs/afs/symlink.c",
"FirstLine": 60,
"LineCount": 25
}Results: {
"Contents": " 60:\t\tchar *p;\n 61:\t\n 62:\t\trcu_assign_pointer(vnode-\u003esymlink, symlink);\n 63:\t\top-\u003ecreate.symlink = NULL;\n 64:\t\n 65:\t\tif (!fscache_cookie_enabled(netfs_i_cookie(\u0026vnode-\u003enetfs)))\n 66:\t\t\treturn;\n 67:\t\n 68:\t\tvnode-\u003edirectory =\n 69:\t\t\tbvecq_alloc_buffer(PAGE_SIZE,\n 70:\t\t\t\t\t mapping_gfp_mask(vnode-\u003enetfs.inode.i_mapping),\n 71:\t\t\t\t\t false);\n 72:\t\tif (!vnode-\u003edirectory)\n 73:\t\t\treturn;\n 74:\t\n 75:\t\tvnode-\u003edirectory_size = size;\n 76:\t\tp = bvec_kmap_partial(\u0026vnode-\u003edirectory-\u003ebv[0], 0);\n 77:\t\tmemcpy(p, symlink-\u003econtent, size);\n 78:\t\tif (size \u003c PAGE_SIZE)\n 79:\t\t\tmemset(p + size, 0, PAGE_SIZE - size);\n 80:\t\tkunmap_local(p);\n 81:\t\tnetfs_single_mark_inode_dirty(\u0026vnode-\u003enetfs.inode);\n 82:\t}\n 83:\t\n 84:\t/*\n"
} |
| 73/3 | 2026/10/06 21:29 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 88183 output: 36 thoughts: 64 Thoughts: Okay, let's dive into this code. **Analyzing Memory Zeroing and Symlink Reads** Hmm, line 78-79... there it is: `if (size < PAGE_SIZE) memset(p + size, 0, PAGE_SIZE - size);`! *Explicitly* zeroed! I knew I had seen something in there related to memory initialization. That was the piece I was looking for. Now, I wonder if that also applies to `afs_do_read_symlink`? It would be good to verify if this behavior is consistent across different functions or, if there are some specific behaviors only applicable to some functions. Better go check... it's easy to assume consistency, but we all know that's dangerous. |
| 74/3 | 2026/10/06 21:29 | tool | read-file |
0mArgs:{
"File": "fs/afs/symlink.c",
"FirstLine": 100,
"LineCount": 35
}Results: {
"Contents": " 100:\t\tif (!vnode-\u003edirectory) {\n 101:\t\t\tvnode-\u003edirectory = bvecq_alloc_buffer(\n 102:\t\t\t\tPAGE_SIZE, mapping_gfp_mask(vnode-\u003enetfs.inode.i_mapping), false);\n 103:\t\t\tif (!vnode-\u003edirectory)\n 104:\t\t\t\treturn -ENOMEM;\n 105:\t\t}\n 106:\t\n 107:\t\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, vnode-\u003edirectory, 0, 0, PAGE_SIZE);\n 108:\t\n 109:\t\t/* AFS requires us to perform the read of a symlink as a single unit to\n 110:\t\t * avoid issues with the content being changed between reads.\n 111:\t\t */\n 112:\t\tret = netfs_read_single(\u0026vnode-\u003enetfs.inode, NULL, \u0026iter);\n 113:\t\tif (ret \u003e= 0) {\n 114:\t\t\ti_size = ret;\n 115:\t\t\tif (i_size \u003e PAGE_SIZE - 1) {\n 116:\t\t\t\ttrace_afs_file_error(vnode, -EFBIG, afs_file_error_dir_big);\n 117:\t\t\t\treturn -EFBIG;\n 118:\t\t\t}\n 119:\t\t\tvnode-\u003edirectory_size = i_size;\n 120:\t\n 121:\t\t\t/* Copy the symlink. */\n 122:\t\t\tsymlink = kmalloc_flex(struct afs_symlink, content, i_size + 1);\n 123:\t\t\tif (!symlink)\n 124:\t\t\t\treturn -ENOMEM;\n 125:\t\n 126:\t\t\trefcount_set(\u0026symlink-\u003eref, 1);\n 127:\t\t\tsymlink-\u003econtent[i_size] = 0;\n 128:\t\n 129:\t\t\tconst char *s = bvec_kmap_partial(\u0026vnode-\u003edirectory-\u003ebv[0], 0);\n 130:\t\n 131:\t\t\tmemcpy(symlink-\u003econtent, s, i_size);\n 132:\t\t\tkunmap_local(s);\n 133:\t\n 134:\t\t\tafs_replace_symlink(vnode, symlink);\n"
} |
| 75/3 | 2026/10/06 21:29 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 89042 output: 37 thoughts: 1047 Thoughts: **Analyzing NetFS Read Operations**
Okay, so I'm looking into how `afs_do_read_symlink` interacts with the network filesystem, specifically the read path. I see that `netfs_read_single` is called and reads `PAGE_SIZE` into `iter`. Hmm, I should check what `netfs_read_single` does. Let me look at `read_single.c`.
Right, it uses `netfs_extract_iter` with `rreq->len` and passes in `&rreq->dispatch_cursor.bvecq`. Now, I want to see what happens to any potential unread portion of the buffer. That must be handled by `netfs_clear_unread`, found in `fs/netfs/read_collect.c`.
Aha! `netfs_clear_unread` zeroes the unread tail of the buffer. Interesting. It uses `iov_iter_zero`. Let me also check `netfs_clear_unread_dio` as well. Seems similar.
Wait a second. What's this? There is a line `subreq->transferred += iov_iter_zero(fill, &iter);` . Let me see how it is used.
Oh, right, in `netfs_clear_unread_dio` I see the following code:
```c
iov_iter_bvec_queue(&iter, ITER_DEST, subreq->io_buffer.bvecq,
subreq->io_buffer.slot, subreq->io_buffer.offset, subreq->len);
iov_iter_advance(&iter, subreq->transferred);
iov_iter_zero(fill, &iter);
subreq->transferred += iov_iter_zero(fill, &iter);
```
Wait a minute! Did I read that right? It seems to call `iov_iter_zero` TWICE in a row! `iov_iter_zero` is called, and then *immediately* called again, with the same buffer! That's... definitely not right. I need to investigate this further. I'll need to look at the code around line 45 in `fs/netfs/read_collect.c` to understand the flow and context.
|
| 76/3 | 2026/10/06 21:30 | tool | read-file |
0mArgs:{
"File": "fs/netfs/read_collect.c",
"FirstLine": 35,
"LineCount": 30
}Results: {
"Contents": " 35:\t\n 36:\t\tif (subreq-\u003estart + subreq-\u003etransferred \u003e= subreq-\u003erreq-\u003ei_size)\n 37:\t\t\t__set_bit(NETFS_SREQ_HIT_EOF, \u0026subreq-\u003eflags);\n 38:\t}\n 39:\t\n 40:\tstatic void netfs_clear_unread_dio(struct netfs_io_subrequest *subreq)\n 41:\t{\n 42:\t\tstruct netfs_io_request *rreq = subreq-\u003erreq;\n 43:\t\tstruct iov_iter iter;\n 44:\t\tuoff_t pos = subreq-\u003estart + subreq-\u003etransferred;\n 45:\t\tsize_t fill;\n 46:\t\n 47:\t\tif (pos \u003e= rreq-\u003ei_size)\n 48:\t\t\treturn;\n 49:\t\n 50:\t\tfill = umin(rreq-\u003ei_size - pos, subreq-\u003elen - subreq-\u003etransferred);\n 51:\t\n 52:\t\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\n 53:\t\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n 54:\t\tiov_iter_advance(\u0026iter, subreq-\u003etransferred);\n 55:\t\tiov_iter_zero(fill, \u0026iter);\n 56:\t\n 57:\t\tsubreq-\u003etransferred += iov_iter_zero(fill, \u0026iter);\n 58:\t}\n 59:\t\n 60:\t/*\n 61:\t * Cancel the copy-to-cache mark on a folio.\n 62:\t */\n 63:\tvoid netfs_cancel_copy_to_cache(struct netfs_io_request *rreq, struct folio *folio)\n 64:\t{\n"
} |
| 77/3 | 2026/10/06 21:30 | llm | kmsan-evaluator |
1mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 91107 output: 38 thoughts: 1796 Thoughts: **KMSAN Considerations for `iov_iter_zero` in Netfs** Okay, so I'm looking at this `iov_iter_zero` call within the Netfs code. I'm seeing it's being used to zero out buffers, and initially, my mind went to the potential for uninitialized memory reads, especially with the way the iterator is advanced and the `fill` variable is used. It looked like it could be writing past `fill`. The second call looked especially suspicious in how it was interacting with `iter`! My immediate concerns were: is this going to lead to out-of-bounds access, or perhaps reading uninitialized memory by zeroing past some boundary? However, I quickly realized zeroing memory isn't the same as *reading* uninitialized memory. It's writing, which means zero, not some garbage. It does write up to `min(fill, (subreq->len - subreq->transferred) - fill)` bytes. Then I considered the "Beyond EOF" situation, and how `fill` is calculated. It seemed we might be zeroing memory beyond the actual file size (`rreq->i_size`) and potentially into the user's buffer for Direct I/O (DIO). That's a reasonable thing to be doing. Then I paused and thought, should I be concerned about KMSAN here? It's time to run through the check-list. Based on my understanding, KMSAN is for catching *uninitialized* reads, not out-of-bounds writes. Standard KASAN should handle OOB. So, I need to check if there is an uninitialized *read* in these changes. Here's the critical question: Does this patch introduce a risk of exposing uninitialized memory that KMSAN would catch? First, I examined whether any kernel structures are being sent to user space (Condition 1). Netfs operates primarily with page cache folios and user buffers for DIO, not kernel structs. That's a pass. Next, I looked at conditional logic or branching depending on potentially uninitialized variables (Condition 2). I went through the relevant structures introduced and modified in this patch: `dispatch_cursor` in `read_retry.c` and `write_retry.c` were zero-initialized. Then I saw a `mark_cursor` field, and I quickly realized it's initialized by `bvecq_pos_set` which copies from `rreq->dispatch_cursor`. Now I got concerned and started to trace how `rreq->dispatch_cursor` itself is initialized. It looks like it's initialized within `netfs_create_singular_buffer` to zero, and in `netfs_read_gaps`, it's partially initialized, which caused me to pause, especially in the `netfs_read_gaps` code where I had to trace `slot` and `offset` initialization again. So, in the end, it doesn't look like any of the changes introduce an uninitialized read that KMSAN would catch. Therefore, it's a pass. I should not set `NeedsKMSAN=true`. |
| 78/3 | 2026/10/06 21:31 | tool | read-file |
0mArgs:{
"File": "fs/netfs/buffered_read.c",
"FirstLine": 540,
"LineCount": 45
}Results: {
"Contents": " 540:\t */\n 541:\tstatic int netfs_read_gaps(struct file *file, struct folio *folio)\n 542:\t{\n 543:\t\tstruct netfs_io_request *rreq;\n 544:\t\tstruct address_space *mapping = folio-\u003emapping;\n 545:\t\tstruct netfs_group *group = netfs_folio_group(folio);\n 546:\t\tstruct netfs_folio *finfo = netfs_folio_info(folio);\n 547:\t\tstruct netfs_inode *ctx = netfs_inode(mapping-\u003ehost);\n 548:\t\tstruct bvecq *bq = NULL;\n 549:\t\tunsigned int from = finfo-\u003edirty_offset;\n 550:\t\tunsigned int to = from + finfo-\u003edirty_len;\n 551:\t\tunsigned int off = 0;\n 552:\t\tsize_t flen = folio_size(folio);\n 553:\t\tsize_t nr_bvec = flen / PAGE_SIZE + 2;\n 554:\t\tsize_t part;\n 555:\t\tint ret, i = 0;\n 556:\t\n 557:\t\t_enter(\"%lx\", folio-\u003eindex);\n 558:\t\n 559:\t\trreq = netfs_alloc_request(mapping, file, folio_pos(folio), flen, NETFS_READ_GAPS);\n 560:\t\tif (IS_ERR(rreq)) {\n 561:\t\t\tret = PTR_ERR(rreq);\n 562:\t\t\tgoto alloc_error;\n 563:\t\t}\n 564:\t\n 565:\t\tret = netfs_begin_cache_read(rreq, ctx);\n 566:\t\tif (ret == -ENOMEM || ret == -EINTR || ret == -ERESTARTSYS)\n 567:\t\t\tgoto discard;\n 568:\t\n 569:\t\tnetfs_stat(\u0026netfs_n_rh_read_folio);\n 570:\t\ttrace_netfs_read(rreq, rreq-\u003estart, rreq-\u003elen, netfs_read_trace_read_gaps);\n 571:\t\n 572:\t\t/* Fiddle the buffer so that a gap at the beginning and/or a gap at the\n 573:\t\t * end get copied to, but the middle is discarded.\n 574:\t\t */\n 575:\t\tret = -ENOMEM;\n 576:\t\tbq = bvecq_alloc_chain(nr_bvec, rreq-\u003egfp, false);\n 577:\t\tif (!bq)\n 578:\t\t\tgoto discard;\n 579:\t\trreq-\u003edispatch_cursor.bvecq = bq;\n 580:\t\n 581:\t\ttrace_netfs_folio(folio, netfs_folio_trace_read_gaps);\n 582:\t\n 583:\t\tfor (struct bvecq *p = bq; p; p = p-\u003enext)\n 584:\t\t\tp-\u003emem_type = BVECQ_MEM_PAGECACHE;\n"
} |
| 79/3 | 2026/10/06 21:31 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 94408 output: 31 thoughts: 42 Thoughts: Here's my thought process, as if I were considering this question: **Investigating `rreq` Allocation in `netfs_alloc_request`** Okay, so I need to figure out how the request structure `rreq` is allocated within the `netfs_alloc_request` function. That's a fundamental aspect of the network filesystem object management, and understanding the memory lifecycle of these requests is crucial for debugging and optimization. Let's dive straight into the code. Time to fire up my editor and go spelunking in `fs/netfs/objects.c`. I'll be looking specifically at the `netfs_alloc_request` function itself. I expect to see the standard memory allocation calls, probably `kmalloc` or similar, depending on the kernel's memory management strategy and the expected size of the `rreq` structure. I also need to make sure that whatever's allocated, the allocated struct is initialized correctly, and that the function handles potential allocation failures gracefully. This will probably involve error checks and returning a NULL value in case of failure. Understanding this initial allocation is key to understanding how these requests are used and how they're de-allocated later in the network filesystem lifecycle. |
| 80/3 | 2026/10/06 21:31 | tool | grepper |
0mArgs:{
"Expression": "netfs_alloc_request",
"PathPrefixes": [
"fs/netfs"
]
}Results: {
"Output": "fs/netfs/buffered_read.c=463=void netfs_readahead(struct readahead_control *ractl)\n--\nfs/netfs/buffered_read.c-471-\nfs/netfs/buffered_read.c:472:\trreq = netfs_alloc_request(ractl-\u003emapping, ractl-\u003efile, start, size,\nfs/netfs/buffered_read.c-473-\t\t\t\t NETFS_READAHEAD);\n--\nfs/netfs/buffered_read.c=541=static int netfs_read_gaps(struct file *file, struct folio *folio)\n--\nfs/netfs/buffered_read.c-558-\nfs/netfs/buffered_read.c:559:\trreq = netfs_alloc_request(mapping, file, folio_pos(folio), flen, NETFS_READ_GAPS);\nfs/netfs/buffered_read.c-560-\tif (IS_ERR(rreq)) {\n--\nfs/netfs/buffered_read.c=665=int netfs_read_folio(struct file *file, struct folio *folio)\n--\nfs/netfs/buffered_read.c-678-\nfs/netfs/buffered_read.c:679:\trreq = netfs_alloc_request(mapping, file,\nfs/netfs/buffered_read.c-680-\t\t\t\t folio_pos(folio), folio_size(folio),\n--\nfs/netfs/buffered_read.c=794=int netfs_write_begin(struct netfs_inode *ctx,\n--\nfs/netfs/buffered_read.c-833-\nfs/netfs/buffered_read.c:834:\trreq = netfs_alloc_request(mapping, file,\nfs/netfs/buffered_read.c-835-\t\t\t\t folio_pos(folio), folio_size(folio),\n--\nfs/netfs/buffered_read.c=886=int netfs_prefetch_for_write(struct file *file, struct folio *folio,\n--\nfs/netfs/buffered_read.c-899-\nfs/netfs/buffered_read.c:900:\trreq = netfs_alloc_request(mapping, file, start, flen,\nfs/netfs/buffered_read.c-901-\t\t\t\t NETFS_READ_FOR_WRITE);\n--\nfs/netfs/direct_read.c=136=ssize_t netfs_unbuffered_read_iter_locked(struct kiocb *iocb, struct iov_iter *iter)\n--\nfs/netfs/direct_read.c-152-\nfs/netfs/direct_read.c:153:\trreq = netfs_alloc_request(iocb-\u003eki_filp-\u003ef_mapping, iocb-\u003eki_filp,\nfs/netfs/direct_read.c-154-\t\t\t\t iocb-\u003eki_pos, orig_count,\n--\nfs/netfs/internal.h=82=int netfs_clear_stale_pre_isize(struct inode *inode, uoff_t from,\n--\nfs/netfs/internal.h-87- */\nfs/netfs/internal.h:88:struct netfs_io_request *netfs_alloc_request(struct address_space *mapping,\nfs/netfs/internal.h-89-\t\t\t\t\t struct file *file,\n--\nfs/netfs/objects.c=12=static void netfs_free_request(struct work_struct *work);\n--\nfs/netfs/objects.c-16- */\nfs/netfs/objects.c:17:struct netfs_io_request *netfs_alloc_request(struct address_space *mapping,\nfs/netfs/objects.c-18-\t\t\t\t\t struct file *file,\n--\nfs/netfs/objects.c=187=void netfs_put_failed_request(struct netfs_io_request *rreq)\n--\nfs/netfs/objects.c-191-\t/* new requests have two references (see\nfs/netfs/objects.c:192:\t * netfs_alloc_request(), and this function is only allowed on\nfs/netfs/objects.c-193-\t * new request objects\n--\nfs/netfs/read_single.c=170=ssize_t netfs_read_single(struct inode *inode, struct file *file, struct iov_iter *iter)\n--\nfs/netfs/read_single.c-175-\nfs/netfs/read_single.c:176:\trreq = netfs_alloc_request(inode-\u003ei_mapping, file, 0, iov_iter_count(iter),\nfs/netfs/read_single.c-177-\t\t\t\t NETFS_READ_SINGLE);\n--\nfs/netfs/write_issue.c=90=struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,\n--\nfs/netfs/write_issue.c-100-\nfs/netfs/write_issue.c:101:\twreq = netfs_alloc_request(mapping, file, start, 0, origin);\nfs/netfs/write_issue.c-102-\tif (IS_ERR(wreq))\n"
} |
| 81/3 | 2026/10/06 21:31 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 96074 output: 35 thoughts: 24 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience: **Diving into `netfs_alloc_request`** Alright, time to get my hands dirty. My initial thought is to zero in on `netfs_alloc_request`. I need to refresh my memory on the allocation strategy here, and pinpoint any potential weak spots. From the file path, `fs/netfs/objects.c`, I'm assuming we're dealing with the core request allocation within the network filesystem implementation, so I want to get a good handle on how the core object's memory is managed and how the request itself is created. I'm looking for anything that might lead to fragmentation, or contention issues under load – resource leaks, or inefficient memory handling, especially if we're seeing any performance bottlenecks. It's time to dig in and re-familiarize myself with the relevant code. |
| 82/3 | 2026/10/06 21:31 | tool | read-file |
0mArgs:{
"File": "fs/netfs/objects.c",
"FirstLine": 17,
"LineCount": 45
}Results: {
"Contents": " 17:\tstruct netfs_io_request *netfs_alloc_request(struct address_space *mapping,\n 18:\t\t\t\t\t\t struct file *file,\n 19:\t\t\t\t\t\t uoff_t start, size_t len,\n 20:\t\t\t\t\t\t enum netfs_io_origin origin)\n 21:\t{\n 22:\t\tstatic atomic_t debug_ids;\n 23:\t\tstruct inode *inode = file ? file_inode(file) : mapping-\u003ehost;\n 24:\t\tstruct netfs_inode *ctx = netfs_inode(inode);\n 25:\t\tstruct netfs_io_request *rreq;\n 26:\t\tmempool_t *mempool = ctx-\u003eops-\u003erequest_pool ?: \u0026netfs_request_pool;\n 27:\t\tstruct kmem_cache *cache = mempool-\u003epool_data;\n 28:\t\tgfp_t gfp = GFP_KERNEL;\n 29:\t\tint ret;\n 30:\t\n 31:\t\t/* Writeback is part of memory reclaim and must not fail due to ENOMEM. */\n 32:\t\tif (origin == NETFS_WRITEBACK || origin == NETFS_WRITEBACK_SINGLE) {\n 33:\t\t\tgfp = GFP_NOFS; /* Allows use of mempools. */\n 34:\t\n 35:\t\t\trreq = mempool_alloc(mempool, gfp);\n 36:\t\t} else {\n 37:\t\t\trreq = mempool_alloc_noreserve(mempool, gfp);\n 38:\t\t\tif (!rreq)\n 39:\t\t\t\treturn ERR_PTR(-ENOMEM);\n 40:\t\t}\n 41:\t\n 42:\t\tmemset(rreq, 0, kmem_cache_size(cache));\n 43:\t\tINIT_WORK(\u0026rreq-\u003ecleanup_work, netfs_free_request);\n 44:\t\trreq-\u003egfp\t\t= gfp;\n 45:\t\trreq-\u003estart\t\t= start;\n 46:\t\trreq-\u003ecollected_to\t= start;\n 47:\t\trreq-\u003ecache_coll_to\t= start;\n 48:\t\trreq-\u003ecleaned_to\t= start;\n 49:\t\trreq-\u003elen\t\t= len;\n 50:\t\trreq-\u003eprogress_at\t= 0;\n 51:\t\trreq-\u003eorigin\t\t= origin;\n 52:\t\trreq-\u003enetfs_ops\t\t= ctx-\u003eops;\n 53:\t\trreq-\u003emapping\t\t= mapping;\n 54:\t\trreq-\u003einode\t\t= inode;\n 55:\t\trreq-\u003ei_size\t\t= i_size_read(inode);\n 56:\t\trreq-\u003edebug_id\t\t= atomic_inc_return(\u0026debug_ids);\n 57:\t\trreq-\u003ewsize\t\t= INT_MAX;\n 58:\t\trreq-\u003eio_streams[0].sreq_max_len = ULONG_MAX;\n 59:\t\trreq-\u003eio_streams[0].sreq_max_segs = 0;\n 60:\t\tspin_lock_init(\u0026rreq-\u003elock);\n 61:\t\tinit_waitqueue_head(\u0026rreq-\u003ewaitq);\n"
} |
| 83/3 | 2026/10/06 21:31 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 97123 output: 32 thoughts: 102 Thoughts: **Analyzing Request Zeroing and Subrequest Allocation** Okay, let's zoom in on this `rreq` structure, specifically line 42. `memset(rreq, 0, kmem_cache_size(cache));` Ah, so the entire `rreq` is being zeroed out whenever it's allocated. That's good to know. This means that, by default, `rreq->dispatch_cursor.slot` is initialized to 0, and `offset` is also starting at 0. This initialization is crucial for ensuring the structure has a known starting state. But what about the handling of subrequests? How are *those* allocated and initialized within this system? I need to understand how these subrequests fit into the overall picture. I'll need to check the `netfs_alloc_subrequest` function within `fs/netfs/objects.c` to see how subrequests are allocated and if zeroing is applied to those. This will give me a better understanding of the overall memory management strategy and potential interactions between the main request and subrequests. |
| 84/3 | 2026/10/06 21:31 | tool | grepper |
0mArgs:{
"Expression": "netfs_alloc_subrequest",
"PathPrefixes": [
"fs/netfs"
]
}Results: {
"Output": "fs/netfs/buffered_read.c=272=static void netfs_read_to_pagecache(struct netfs_io_request *rreq)\n--\nfs/netfs/buffered_read.c-382-\nfs/netfs/buffered_read.c:383:\t\tsubreq = netfs_alloc_subrequest(rreq, source);\nfs/netfs/buffered_read.c-384-\t\tif (!subreq) {\n--\nfs/netfs/direct_read.c=23=static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)\n--\nfs/netfs/direct_read.c-34-\nfs/netfs/direct_read.c:35:\t\tsubreq = netfs_alloc_subrequest(rreq, NETFS_DOWNLOAD_FROM_SERVER);\nfs/netfs/direct_read.c-36-\t\tif (!subreq) {\n--\nfs/netfs/internal.h=95=void netfs_put_failed_request(struct netfs_io_request *rreq);\nfs/netfs/internal.h:96:struct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq,\nfs/netfs/internal.h-97-\t\t\t\t\t\t enum netfs_io_source source);\n--\nfs/netfs/objects.c=187=void netfs_put_failed_request(struct netfs_io_request *rreq)\n--\nfs/netfs/objects.c-203- */\nfs/netfs/objects.c:204:struct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq,\nfs/netfs/objects.c-205-\t\t\t\t\t\t enum netfs_io_source source)\n--\nfs/netfs/read_retry.c=30=static void netfs_retry_read_subrequests(struct netfs_io_request *rreq)\n--\nfs/netfs/read_retry.c-209-\t\tdo {\nfs/netfs/read_retry.c:210:\t\t\tsubreq = netfs_alloc_subrequest(rreq, NETFS_DOWNLOAD_FROM_SERVER);\nfs/netfs/read_retry.c-211-\t\t\tif (!subreq) {\n--\nfs/netfs/read_single.c=76=static int netfs_single_dispatch_read(struct netfs_io_request *rreq)\n--\nfs/netfs/read_single.c-97-\nfs/netfs/read_single.c:98:\tsubreq = netfs_alloc_subrequest(rreq, source);\nfs/netfs/read_single.c-99-\tif (!subreq)\n--\nfs/netfs/write_issue.c=154=void netfs_prepare_write(struct netfs_io_request *wreq,\n--\nfs/netfs/write_issue.c-159-\nfs/netfs/write_issue.c:160:\tsubreq = netfs_alloc_subrequest(wreq, stream-\u003esource);\nfs/netfs/write_issue.c-161-\tif (!subreq)\n--\nfs/netfs/write_retry.c=17=static void netfs_retry_write_stream(struct netfs_io_request *wreq,\n--\nfs/netfs/write_retry.c-159-\t\tdo {\nfs/netfs/write_retry.c:160:\t\t\tsubreq = netfs_alloc_subrequest(wreq, stream-\u003esource);\nfs/netfs/write_retry.c-161-\t\t\tsubreq-\u003estart\t\t= start;\n"
} |
| 85/3 | 2026/10/06 21:31 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 98418 output: 36 thoughts: 30 Thoughts: Okay, let's dive into this. Here's my thought process on the matter: **Analyzing `netfs_alloc_subrequest` in `fs/netfs/objects.c`** Alright, my first step is to zero in on that `netfs_alloc_subrequest` function. I need to refresh my memory and understand its purpose within the `netfs` context. It's around line 204 in `fs/netfs/objects.c`, so that's the starting point. I need to quickly ascertain: What's it allocating? Is it request memory, some sort of data structure, maybe a context object? What are the implications of the allocation? Is there a corresponding deallocation, and where is that handled? I'm already anticipating potential problems around resource exhaustion or memory leaks, so looking for the general patterns and lifecycle of what this routine is managing are paramount, especially within a filesystem module. Let's see if there are any obvious error handling paths, and what those error conditions entail. The allocation strategy itself would be interesting as well, Is it just a simple `kmalloc`, or something more involved, like a custom allocator or a slab cache? Understanding the nuances of the allocation and deallocation will provide a good foundation before doing anything else. |
| 86/3 | 2026/10/06 21:31 | tool | read-file |
0mArgs:{
"File": "fs/netfs/objects.c",
"FirstLine": 204,
"LineCount": 40
}Results: {
"Contents": " 204:\tstruct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq,\n 205:\t\t\t\t\t\t\t enum netfs_io_source source)\n 206:\t{\n 207:\t\tstruct netfs_io_subrequest *subreq;\n 208:\t\tmempool_t *mempool = rreq-\u003enetfs_ops-\u003esubrequest_pool ?: \u0026netfs_subrequest_pool;\n 209:\t\tstruct kmem_cache *cache = mempool-\u003epool_data;\n 210:\t\n 211:\t\tif (rreq-\u003egfp == GFP_KERNEL)\n 212:\t\t\tsubreq = mempool_alloc_noreserve(mempool, rreq-\u003egfp);\n 213:\t\telse\n 214:\t\t\tsubreq = mempool_alloc(mempool, rreq-\u003egfp);\n 215:\t\tif (!subreq)\n 216:\t\t\treturn NULL;\n 217:\t\n 218:\t\tmemset(subreq, 0, kmem_cache_size(cache));\n 219:\t\tINIT_WORK(\u0026subreq-\u003ework, NULL);\n 220:\t\tINIT_LIST_HEAD(\u0026subreq-\u003erreq_link);\n 221:\t\trefcount_set(\u0026subreq-\u003eref, 2);\n 222:\t\tsubreq-\u003esource = source;\n 223:\t\tsubreq-\u003erreq = rreq;\n 224:\t\tsubreq-\u003edebug_index = atomic_inc_return(\u0026rreq-\u003esubreq_counter);\n 225:\t\tnetfs_get_request(rreq, netfs_rreq_trace_get_subreq);\n 226:\t\tnetfs_stat(\u0026netfs_n_rh_sreq);\n 227:\t\treturn subreq;\n 228:\t}\n 229:\t\n 230:\tvoid netfs_get_subrequest(struct netfs_io_subrequest *subreq,\n 231:\t\t\t\t enum netfs_sreq_ref_trace what)\n 232:\t{\n 233:\t\tint r;\n 234:\t\n 235:\t\t__refcount_inc(\u0026subreq-\u003eref, \u0026r);\n 236:\t\ttrace_netfs_sreq_ref(subreq-\u003erreq-\u003edebug_id, subreq-\u003edebug_index, r + 1,\n 237:\t\t\t\t what);\n 238:\t}\n 239:\t\n 240:\tstatic void netfs_free_subrequest(struct netfs_io_subrequest *subreq)\n 241:\t{\n 242:\t\tstruct netfs_io_request *rreq = subreq-\u003erreq;\n 243:\t\n"
} |
| 87/3 | 2026/10/06 21:31 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 99469 output: 29 thoughts: 551 Thoughts: **Analyzing Memory Management and `readahead_control`** Okay, let's break this down. Line 218 immediately catches my attention: `memset(subreq, 0, kmem_cache_size(cache));`. I see that `subreq` is being zeroed out entirely. Good. Then I notice the addition of `struct bvecq_pos io_buffer;` within `subreq`. That means it's also being initialized to zero upon allocation, which is what I'd expect. Now I need to see how its lifecycle plays out. Looking at `netfs_free_subrequest`, I see `bvecq_pos_unset(&subreq->io_buffer);`. That looks like a clean deallocation of resources. Excellent. I also see similar `bvecq_pos_unset` calls in `netfs_deinit_request`, for `rreq->load_cursor`, `rreq->dispatch_cursor`, and `rreq->collect_cursor`. And then `bvecq_put(rreq->spare);`. This confirms that all references are being handled correctly from what I see. Now, let's switch gears to `readahead_control`. I'm checking the definition in `include/linux/pagemap.h`. I see `unsigned int _nr_folios;`. My immediate thought is, how is this initialized and used? Looking at `mm/readahead.c`, I see `_nr_folios` is indeed initialized to zero in `read_pages`. It looks like there is an increment happening in `page_cache_ra_unbounded`, `ra_alloc_folio` and in `readahead_expand`. Good, so the counter is going up as needed. `readahead_folio_count` looks to be a simple getter, returning `_nr_folios`. Ah, here we go! `bvecq_load_from_ra(struct bvecq_pos *pos, struct readahead_control *ractl)` uses `ractl->_nr_folios` as the parameter for the size of the bvecq chain. So `_nr_folios` directly controls the number of slots allocated! Now I need to determine where `ractl` comes from. Is it a stack structure that's local to a function, or is it passed from the VM layer? I should check where `struct readahead_control` is initialized to be sure. |
| 88/3 | 2026/10/06 21:32 | tool | grepper |
0mArgs:{
"Expression": "struct readahead_control",
"PathPrefixes": [
"mm/"
]
}Results: {
"Output": "mm/internal.h=610=int folio_unmap_invalidate(struct address_space *mapping, struct folio *folio,\n--\nmm/internal.h-612-\nmm/internal.h:613:void page_cache_ra_order(struct readahead_control *, struct file_ra_state *);\nmm/internal.h:614:void force_page_cache_ra(struct readahead_control *, unsigned long nr);\nmm/internal.h-615-static inline void force_page_cache_readahead(struct address_space *mapping,\n--\nmm/readahead.c=147=EXPORT_SYMBOL_GPL(file_ra_state_init);\n--\nmm/readahead.c-159- */\nmm/readahead.c:160:static void read_pages(struct readahead_control *rac)\nmm/readahead.c-161-{\n--\nmm/readahead.c-195-\nmm/readahead.c:196:static struct folio *ractl_alloc_folio(struct readahead_control *ractl,\nmm/readahead.c-197-\t\t\t\t gfp_t gfp_mask, unsigned int order)\n--\nmm/readahead.c-222- */\nmm/readahead.c:223:void page_cache_ra_unbounded(struct readahead_control *ractl,\nmm/readahead.c-224-\t\tunsigned long nr_to_read, unsigned long lookahead_size)\n--\nmm/readahead.c=320=EXPORT_SYMBOL_GPL(page_cache_ra_unbounded);\n--\nmm/readahead.c-327- */\nmm/readahead.c:328:static void do_page_cache_ra(struct readahead_control *ractl,\nmm/readahead.c-329-\t\tunsigned long nr_to_read, unsigned long lookahead_size)\n--\nmm/readahead.c-357- */\nmm/readahead.c:358:void force_page_cache_ra(struct readahead_control *ractl,\nmm/readahead.c-359-\t\tunsigned long nr_to_read)\n--\nmm/readahead.c=410=static unsigned long get_next_ra_size(struct file_ra_state *ra,\n--\nmm/readahead.c-460-\nmm/readahead.c:461:static inline int ra_alloc_folio(struct readahead_control *ractl, pgoff_t index,\nmm/readahead.c-462-\t\tpgoff_t mark, unsigned int order, gfp_t gfp)\n--\nmm/readahead.c-483-\nmm/readahead.c:484:void page_cache_ra_order(struct readahead_control *ractl,\nmm/readahead.c-485-\t\tstruct file_ra_state *ra)\n--\nmm/readahead.c-564-\nmm/readahead.c:565:static unsigned long ractl_max_pages(struct readahead_control *ractl,\nmm/readahead.c-566-\t\tunsigned long req_size)\n--\nmm/readahead.c-579-\nmm/readahead.c:580:void page_cache_sync_ra(struct readahead_control *ractl,\nmm/readahead.c-581-\t\tunsigned long req_count)\n--\nmm/readahead.c=654=EXPORT_SYMBOL_GPL(page_cache_sync_ra);\nmm/readahead.c-655-\nmm/readahead.c:656:void page_cache_async_ra(struct readahead_control *ractl,\nmm/readahead.c-657-\t\tstruct folio *folio, unsigned long req_count)\n--\nmm/readahead.c=765=COMPAT_SYSCALL_DEFINE4(readahead, int, fd, compat_arg_u64_dual(offset), size_t, count)\n--\nmm/readahead.c-788- */\nmm/readahead.c:789:void readahead_expand(struct readahead_control *ractl,\nmm/readahead.c-790-\t\t loff_t new_start, size_t new_len)\n"
} |
| 89/3 | 2026/10/06 21:32 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 101657 output: 19 thoughts: 33 Thoughts: Okay, let's dive into this. Here's my thought process on this specific code snippet: **Investigating `struct readahead_control` Initialization** Alright, I'm thinking about how `struct readahead_control` is initialized, specifically at the point where it's actually allocated. That's a critical area for understanding its lifecycle and how it behaves. I need to find the specific allocation and initialization patterns. My immediate next step is to use my knowledge of the codebase and quickly locate any instances where `DEFINE_READAHEAD` is used. That macro is probably involved in defining or initializing this structure, so I'll start there as my jumping-off point. It's a key to understanding how these structures are set up, especially when they're allocated. |
| 90/3 | 2026/10/06 21:32 | tool | grepper |
0mArgs:{
"Expression": "DEFINE_READAHEAD"
}Results: {
"Output": "fs/f2fs/file.c=4794=static int redirty_blocks(struct inode *inode, pgoff_t page_idx, int len)\nfs/f2fs/file.c-4795-{\nfs/f2fs/file.c:4796:\tDEFINE_READAHEAD(ractl, NULL, NULL, inode-\u003ei_mapping, page_idx);\nfs/f2fs/file.c-4797-\tstruct address_space *mapping = inode-\u003ei_mapping;\n--\nfs/verity/pagecache.c=42=void generic_readahead_merkle_tree(struct inode *inode, pgoff_t index,\n--\nfs/verity/pagecache.c-51-\t (!IS_ERR(folio) \u0026\u0026 !folio_test_uptodate(folio))) {\nfs/verity/pagecache.c:52:\t\tDEFINE_READAHEAD(ractl, NULL, NULL, inode-\u003ei_mapping, index);\nfs/verity/pagecache.c-53-\n--\ninclude/linux/pagemap.h=1408=struct readahead_control {\n--\ninclude/linux/pagemap.h-1421-\ninclude/linux/pagemap.h:1422:#define DEFINE_READAHEAD(ractl, f, r, m, i)\t\t\t\t\\\ninclude/linux/pagemap.h-1423-\tstruct readahead_control ractl = {\t\t\t\t\\\n--\ninclude/linux/pagemap.h=1454=void page_cache_sync_readahead(struct address_space *mapping,\n--\ninclude/linux/pagemap.h-1457-{\ninclude/linux/pagemap.h:1458:\tDEFINE_READAHEAD(ractl, file, ra, mapping, index);\ninclude/linux/pagemap.h-1459-\tpage_cache_sync_ra(\u0026ractl, req_count);\n--\ninclude/linux/pagemap.h=1476=void page_cache_async_readahead(struct address_space *mapping,\n--\ninclude/linux/pagemap.h-1479-{\ninclude/linux/pagemap.h:1480:\tDEFINE_READAHEAD(ractl, file, ra, mapping, folio-\u003eindex);\ninclude/linux/pagemap.h-1481-\tpage_cache_async_ra(\u0026ractl, folio, req_count);\n--\nmm/filemap.c=2672=static int filemap_readahead(struct kiocb *iocb, struct file *file,\n--\nmm/filemap.c-2675-{\nmm/filemap.c:2676:\tDEFINE_READAHEAD(ractl, file, \u0026file-\u003ef_ra, mapping, folio-\u003eindex);\nmm/filemap.c-2677-\n--\nmm/filemap.c=2686=static int filemap_get_pages(struct kiocb *iocb, size_t count,\n--\nmm/filemap.c-2705-\tif (!folio_batch_count(fbatch)) {\nmm/filemap.c:2706:\t\tDEFINE_READAHEAD(ractl, filp, \u0026filp-\u003ef_ra, mapping, index);\nmm/filemap.c-2707-\n--\nmm/filemap.c=3329=static struct file *do_sync_mmap_readahead(struct vm_fault *vmf)\n--\nmm/filemap.c-3333-\tstruct address_space *mapping = file-\u003ef_mapping;\nmm/filemap.c:3334:\tDEFINE_READAHEAD(ractl, file, ra, mapping, vmf-\u003epgoff);\nmm/filemap.c-3335-\tstruct file *fpin = NULL;\n--\nmm/filemap.c=3452=static struct file *do_async_mmap_readahead(struct vm_fault *vmf,\n--\nmm/filemap.c-3456-\tstruct file_ra_state *ra = \u0026file-\u003ef_ra;\nmm/filemap.c:3457:\tDEFINE_READAHEAD(ractl, file, ra, file-\u003ef_mapping, vmf-\u003epgoff);\nmm/filemap.c-3458-\tstruct file *fpin = NULL;\n--\nmm/internal.h=615=static inline void force_page_cache_readahead(struct address_space *mapping,\n--\nmm/internal.h-617-{\nmm/internal.h:618:\tDEFINE_READAHEAD(ractl, file, \u0026file-\u003ef_ra, mapping, index);\nmm/internal.h-619-\tforce_page_cache_ra(\u0026ractl, nr_to_read);\n"
} |
| 91/3 | 2026/10/06 21:32 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 103074 output: 22 thoughts: 1694 Thoughts: **Analyzing the bvec Queue Implementation** Okay, let's break this down. First, `DEFINE_READAHEAD` cleverly uses a designated initializer for `struct readahead_control` which ensures `_nr_folios` is initialized to 0. As folios are added during readahead, `_nr_folios` increments. Now, let's look at `bvecq_load_from_ra`. It allocates a `bvecq` (the buffer queue) based on the number of folios (`ractl->_nr_folios`). `pos` keeps track of which folio or slice within a folio is being processed and where we are in the queue. The crucial part here is the loop that iterates through folios, setting up the `bvec_set_folio` for each one, updating the `loaded` size. `bvecq_filled_to` marks slots as "filled". The `pos` structure clearly keeps track of the queue and offset, where to start and where we are in the queue. Now, `bvecq_slice`. This is the function for extracting data from the `bvecq`. Inside the loop, it calculates a slice, tracks the number of slots, and the amount of data and updates `pos` appropriately. There's a subtle point: what happens if the slice completely consumes a `bvec`? The critical check is `offset < bvec->bv_len`, and if we arrive where `offset == bvec->bv_len`, we will enter the next iteration and reset offset, move to the next slot. Now let's consider a scenario: We completely consume a `bvec` in a slice and arrive at the "out" label. The conditional `if (slot == bvecq_nr_slots_acquire(bq))` tests if we reached the end of the slot. If that is true, we should advance to the next slot and reset the offset. The `bvecq_pos_move()` updates the `pos` structure correctly. What if, on the *next* call to `bvecq_slice`, the `pos->offset` is equal to `bvec->bv_len`? The conditional `if (offset < bvec->bv_len && bvec->bv_page)` will be false and `offset` will be reset to 0, advancing the `slot++`. So, we will not re-read from the start. Finally, `bvecq_pos_nudge` looks like a helper function for moving the `pos` to a valid position, handling queue book-keeping. It ensures that the `pos` is pointing to a valid slot in a `bvecq` chain, incrementing the `slot` and resetting the `offset` when necessary, which looks like it could be called by a function like `bvecq_slice`, so we must inspect how this function and others utilize `bvecq_pos_nudge`. Now, the question is, how does this all interface with `iov_iter_bvec_queue`? Let's trace back from there. |
| 92/3 | 2026/10/06 21:32 | tool | grepper |
0mArgs:{
"Expression": "iov_iter_bvec_queue"
}Results: {
"Output": "fs/afs/dir.c=135=static void afs_dir_dump(struct afs_vnode *dvnode)\n--\nfs/afs/dir.c-142-\nfs/afs/dir.c:143:\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0, i_size);\nfs/afs/dir.c-144-\titerate_bvecq(\u0026iter, iov_iter_count(\u0026iter), NULL, NULL,\n--\nfs/afs/dir.c=197=static int afs_dir_check(struct afs_vnode *dvnode)\n--\nfs/afs/dir.c-205-\nfs/afs/dir.c:206:\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0, i_size);\nfs/afs/dir.c-207-\tchecked = iterate_bvecq(\u0026iter, iov_iter_count(\u0026iter), dvnode, NULL,\n--\nfs/afs/dir.c=235=static ssize_t afs_do_read_single(struct afs_vnode *dvnode, struct file *file)\n--\nfs/afs/dir.c-260-\nfs/afs/dir.c:261:\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, dvnode-\u003edirectory, 0, 0, dvnode-\u003edirectory_size);\nfs/afs/dir.c-262-\n--\nfs/afs/dir.c=488=static int afs_dir_iterate_contents(struct inode *dir, struct dir_context *dir_ctx)\n--\nfs/afs/dir.c-500-\nfs/afs/dir.c:501:\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0, i_size);\nfs/afs/dir.c-502-\tiov_iter_advance(\u0026iter, round_down(dir_ctx-\u003epos, AFS_DIR_BLOCK_SIZE));\n--\nfs/afs/dir.c=2211=static int afs_dir_writepages(struct address_space *mapping,\n--\nfs/afs/dir.c-2230-\tif (test_bit(AFS_VNODE_DIR_VALID, \u0026dvnode-\u003eflags)) {\nfs/afs/dir.c:2231:\t\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, dvnode-\u003edirectory, 0, 0,\nfs/afs/dir.c-2232-\t\t\t\t dvnode-\u003edirectory_size);\n--\nfs/afs/symlink.c=87=static ssize_t afs_do_read_symlink(struct afs_vnode *vnode)\n--\nfs/afs/symlink.c-106-\nfs/afs/symlink.c:107:\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, vnode-\u003edirectory, 0, 0, PAGE_SIZE);\nfs/afs/symlink.c-108-\n--\nfs/afs/symlink.c=233=int afs_symlink_writepages(struct address_space *mapping,\n--\nfs/afs/symlink.c-250-\t atomic64_read(\u0026vnode-\u003ecb_expires_at) != AFS_NO_CB_PROMISE) {\nfs/afs/symlink.c:251:\t\tiov_iter_bvec_queue(\u0026iter, ITER_SOURCE, vnode-\u003edirectory, 0, 0, PAGE_SIZE);\nfs/afs/symlink.c-252-\t\tret = netfs_writeback_single(mapping, wbc, \u0026iter,\n--\nfs/netfs/buffered_read.c=187=static void netfs_issue_read(struct netfs_io_request *rreq,\n--\nfs/netfs/buffered_read.c-189-{\nfs/netfs/buffered_read.c:190:\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\nfs/netfs/buffered_read.c-191-\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n--\nfs/netfs/direct_read.c=23=static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq)\n--\nfs/netfs/direct_read.c-69-\nfs/netfs/direct_read.c:70:\t\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\nfs/netfs/direct_read.c-71-\t\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset,\n--\nfs/netfs/direct_write.c=98=static int netfs_unbuffered_write(struct netfs_io_request *wreq)\n--\nfs/netfs/direct_write.c-138-\nfs/netfs/direct_write.c:139:\t\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_SOURCE,\nfs/netfs/direct_write.c-140-\t\t\t\t subreq-\u003eio_buffer.bvecq, subreq-\u003eio_buffer.slot,\n--\nfs/netfs/read_collect.c=27=static void netfs_clear_unread(struct netfs_io_subrequest *subreq)\n--\nfs/netfs/read_collect.c-30-\nfs/netfs/read_collect.c:31:\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\nfs/netfs/read_collect.c-32-\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n--\nfs/netfs/read_collect.c=40=static void netfs_clear_unread_dio(struct netfs_io_subrequest *subreq)\n--\nfs/netfs/read_collect.c-51-\nfs/netfs/read_collect.c:52:\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\nfs/netfs/read_collect.c-53-\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n--\nfs/netfs/read_retry.c=12=static void netfs_reissue_read(struct netfs_io_request *rreq,\n--\nfs/netfs/read_retry.c-14-{\nfs/netfs/read_retry.c:15:\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\nfs/netfs/read_retry.c-16-\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n--\nfs/netfs/read_single.c=76=static int netfs_single_dispatch_read(struct netfs_io_request *rreq)\n--\nfs/netfs/read_single.c-106-\nfs/netfs/read_single.c:107:\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_DEST, subreq-\u003eio_buffer.bvecq,\nfs/netfs/read_single.c-108-\t\t\t subreq-\u003eio_buffer.slot, subreq-\u003eio_buffer.offset, subreq-\u003elen);\n--\nfs/netfs/write_issue.c=246=void netfs_reissue_write(struct netfs_io_stream *stream,\n--\nfs/netfs/write_issue.c-249-\t// TODO: Use encrypted buffer\nfs/netfs/write_issue.c:250:\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_SOURCE,\nfs/netfs/write_issue.c-251-\t\t\t subreq-\u003eio_buffer.bvecq, subreq-\u003eio_buffer.slot,\n--\nfs/netfs/write_issue.c=264=void netfs_issue_write(struct netfs_io_request *wreq,\n--\nfs/netfs/write_issue.c-271-\nfs/netfs/write_issue.c:272:\tiov_iter_bvec_queue(\u0026subreq-\u003eio_iter, ITER_SOURCE,\nfs/netfs/write_issue.c-273-\t\t\t subreq-\u003eio_buffer.bvecq, subreq-\u003eio_buffer.slot,\n--\nfs/smb/client/smb2ops.c=4937=smb3_init_transform_rq(struct TCP_Server_Info *server, int num_rqst,\n--\nfs/smb/client/smb2ops.c-4960-\t\t\tnew-\u003erq_buffer = buffer;\nfs/smb/client/smb2ops.c:4961:\t\t\tiov_iter_bvec_queue(\u0026new-\u003erq_iter, ITER_SOURCE,\nfs/smb/client/smb2ops.c-4962-\t\t\t\t\t buffer, 0, 0, size);\n--\nfs/smb/client/smb2ops.c=5245=static void smb2_decrypt_offload(struct work_struct *work)\n--\nfs/smb/client/smb2ops.c-5252-\nfs/smb/client/smb2ops.c:5253:\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, dw-\u003ebuffer, 0, 0, dw-\u003elen);\nfs/smb/client/smb2ops.c-5254-\trc = decrypt_raw_data(dw-\u003eserver, dw-\u003ebuf, dw-\u003eserver-\u003evals-\u003eread_rsp_size,\n--\nfs/smb/client/smb2ops.c=5309=receive_encrypted_read(struct TCP_Server_Info *server, struct mid_q_entry **mid,\n--\nfs/smb/client/smb2ops.c-5352-\nfs/smb/client/smb2ops.c:5353:\tiov_iter_bvec_queue(\u0026iter, ITER_DEST, dw-\u003ebuffer, 0, 0, len);\nfs/smb/client/smb2ops.c-5354-\n--\ninclude/linux/uio.h=296=void iov_iter_discard(struct iov_iter *i, unsigned int direction, size_t count);\ninclude/linux/uio.h:297:void iov_iter_bvec_queue(struct iov_iter *i, unsigned int direction,\ninclude/linux/uio.h-298-\t\t\t const struct bvecq *bvecq,\n--\nlib/iov_iter.c=739=EXPORT_SYMBOL(iov_iter_bvec);\n--\nlib/iov_iter.c-741-/**\nlib/iov_iter.c:742: * iov_iter_bvec_queue - Initialise an I/O iterator to use a segmented bvec queue\nlib/iov_iter.c-743- * @i: The iterator to initialise.\n--\nlib/iov_iter.c-753- */\nlib/iov_iter.c:754:void iov_iter_bvec_queue(struct iov_iter *i, unsigned int direction,\nlib/iov_iter.c-755-\t\t\t const struct bvecq *bvecq, unsigned int first_slot,\n--\nlib/iov_iter.c-767-}\nlib/iov_iter.c:768:EXPORT_SYMBOL(iov_iter_bvec_queue);\nlib/iov_iter.c-769-\n--\nlib/tests/kunit_iov_iter.c=417=static void __init iov_kunit_load_bvecq(struct kunit *test,\n--\nlib/tests/kunit_iov_iter.c-434-\t}\nlib/tests/kunit_iov_iter.c:435:\tiov_iter_bvec_queue(iter, dir, bq_head, 0, 0, size);\nlib/tests/kunit_iov_iter.c-436-}\n--\nlib/tests/kunit_iov_iter.c=441=static void __init iov_kunit_copy_to_bvecq(struct kunit *test)\n--\nlib/tests/kunit_iov_iter.c-469-\nlib/tests/kunit_iov_iter.c:470:\t\tiov_iter_bvec_queue(\u0026iter, READ, bq, 0, 0, pr-\u003eto);\nlib/tests/kunit_iov_iter.c-471-\t\tiov_iter_advance(\u0026iter, pr-\u003efrom);\n--\nlib/tests/kunit_iov_iter.c=502=static void __init iov_kunit_copy_from_bvecq(struct kunit *test)\n--\nlib/tests/kunit_iov_iter.c-530-\nlib/tests/kunit_iov_iter.c:531:\t\tiov_iter_bvec_queue(\u0026iter, WRITE, bq, 0, 0, pr-\u003eto);\nlib/tests/kunit_iov_iter.c-532-\t\tiov_iter_advance(\u0026iter, pr-\u003efrom);\n--\nlib/tests/kunit_iov_iter.c=880=static void __init iov_kunit_extract_pages_bvecq(struct kunit *test)\n--\nlib/tests/kunit_iov_iter.c-902-\nlib/tests/kunit_iov_iter.c:903:\t\tiov_iter_bvec_queue(\u0026iter, WRITE, bq, 0, 0, pr-\u003eto);\nlib/tests/kunit_iov_iter.c-904-\t\tiov_iter_advance(\u0026iter, from);\n"
} |
| 93/3 | 2026/10/06 21:32 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 108965 output: 35 thoughts: 29 Thoughts: Okay, here's my interpretation of that thought process, assuming I'm already quite familiar with kernel internals and specifically I/O operations: **Diving into `iov_iter_bvec_queue`** Alright, I'm heading straight for the `iov_iter_bvec_queue` function. It's in `lib/iov_iter.c`, line 754 if I recall correctly. This likely means I'm about to dig into how block vectors (`bvecs`) are handled within the `iov_iter` framework, which is fundamental for efficient scatter-gather I/O operations in the kernel. I'm probably looking at this specific function for a reason - perhaps I'm debugging an issue related to buffered I/O, or optimizing data transfer, or maybe just refreshing my memory on the intricacies of the implementation. I'm expecting to see the logic for queuing `bvec` data into the `iov_iter`, which would involve handling things like page alignment, byte offsets, and perhaps memory mapping or DMA setup, depending on the context. I'll need to pay close attention to the interactions between the `iov_iter` structure and the underlying block device access methods. Specifically, I'll be interested in how it handles situations like short writes, partial reads, and how the iterators advance through the data segments. Gotta see how they manage the scatter-gather lists. Let's see how it's done... |
| 94/3 | 2026/10/06 21:33 | tool | read-file |
0mArgs:{
"File": "lib/iov_iter.c",
"FirstLine": 750,
"LineCount": 20
}Results: {
"Contents": " 750:\t * Set up an I/O iterator to either draw data out of the buffers attached to an\n 751:\t * inode or to inject data into those buffers. The pages *must* be prevented\n 752:\t * from evaporation, either by the caller.\n 753:\t */\n 754:\tvoid iov_iter_bvec_queue(struct iov_iter *i, unsigned int direction,\n 755:\t\t\t\t const struct bvecq *bvecq, unsigned int first_slot,\n 756:\t\t\t\t unsigned int offset, size_t count)\n 757:\t{\n 758:\t\tWARN_ON(direction \u0026 ~(READ | WRITE));\n 759:\t\t*i = (struct iov_iter) {\n 760:\t\t\t.iter_type\t= ITER_BVECQ,\n 761:\t\t\t.data_source\t= direction,\n 762:\t\t\t.bvecq\t\t= bvecq,\n 763:\t\t\t.bvecq_slot\t= first_slot,\n 764:\t\t\t.count\t\t= count,\n 765:\t\t\t.iov_offset\t= offset,\n 766:\t\t};\n 767:\t}\n 768:\tEXPORT_SYMBOL(iov_iter_bvec_queue);\n 769:\t\n"
} |
| 95/3 | 2026/10/06 21:33 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 109677 output: 42 thoughts: 1881 Thoughts: **NetFS Read/Write Optimization and Bvecq Analysis** Okay, let's break this down. First, `iov_iter_bvec_queue` is used, and it sets the `.iter_type` to `ITER_BVECQ`. Got it. Now, I'm examining `netfs_rreq_assess_dio`. The code iterates through a `bvecq` (I'm assuming this is a *bvec queue*) to flush and mark dirty pages if it's either an unbuffered read or a direct I/O read. We're iterating using `bvecq_next`, and inside the loop it reads the slot count first, then flushes and dirties each page. The comments about needing to consider dirtying pages just like CIFS do make sense given the I/O model. I then check `netfs_direct_write.c`. I see `bvecq_pos_set`, `bvecq_slice`, and `iov_iter_bvec_queue` used extensively in `netfs_unbuffered_write`. The `bvecq_pos_transfer` and `bvecq_pos_advance` functions look interesting for managing the write position within the queue during retry operations. I see it again in `netfs_unbuffered_write_iter_locked`, and `netfs_extract_iter` is called to extract data from the iter. Next, I'm analyzing `direct_read.c`. `netfs_dispatch_unbuffered_reads` is similar to the write path. We have `bvecq_pos_set`, `bvecq_slice`, `iov_iter_bvec_queue`, and also calls to `netfs_extract_iter`, suggesting that the read side has a similar structure. The use of `netfs_extract_iter` has got me thinking. I'm checking if it's modifying the `dispatch_cursor`. It seems `netfs_extract_iter` initializes `*_bvecq_head` with the current `bvecq`, leaving `slot` and `offset` untouched. Since `rreq` is zero-allocated, `dispatch_cursor.slot` and `dispatch_cursor.offset` will be zero. I note that `netfs_dispatch_unbuffered_reads` copies `rreq->dispatch_cursor` into `rreq->collect_cursor`. We also see the use of `bvecq_pos_unset` at the end which I'm assuming is the ref release. Now I'm looking at `netfs_read_single`. Here we also use `netfs_extract_iter` and similar logic for managing the read cursor. I note we copy the dispatch cursor again. The key is to figure out whether or not the collect cursor is used here. After reviewing `fs/netfs/read_single.c`, it does not appear that the `collect_cursor` is used at all in this code path. |
| 96/3 | 2026/10/06 21:33 | tool | grepper |
0mError:no matches Args: {
"Expression": "collect_cursor|read_collect|read_collection",
"PathPrefixes": [
"fs/netfs/read_single.c"
]
}Results: {
"error": "no matches"
} |
| 97/3 | 2026/10/06 21:33 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 112254 output: 38 thoughts: 19 Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience: **Analyzing the Termination of `netfs_read_single`** Okay, so I need to dig into `netfs_read_single` and really understand how it wraps things up after a read operation. Specifically, I want to drill down into the control flow related to finalization. My aim is to build a thorough mental map of the graceful shutdown path that this function takes, accounting for success, error scenarios, and edge cases. I'm going to focus on where file handles are closed, memory is freed, and any locks are released. I'm assuming there's some cleanup, and I want to verify if any resources are potentially leaked under specific error conditions. I also want to determine how the calling code should receive the final status. I need a clear picture of all possible exit points and what happens at each of those points. Ultimately, I'm trying to anticipate potential issues related to resource management and general system stability. |
| 98/3 | 2026/10/06 21:33 | tool | read-file |
0mArgs:{
"File": "fs/netfs/read_single.c",
"FirstLine": 180,
"LineCount": 40
}Results: {
"Contents": " 180:\t\n 181:\t\tret = netfs_extract_iter(iter, rreq-\u003elen, INT_MAX, \u0026rreq-\u003edispatch_cursor.bvecq,\n 182:\t\t\t\t\t 0, rreq-\u003egfp);\n 183:\t\tif (ret \u003c 0)\n 184:\t\t\tgoto cleanup_free;\n 185:\t\tif (ret \u003c rreq-\u003elen) {\n 186:\t\t\tret = -EIO;\n 187:\t\t\tgoto cleanup_free;\n 188:\t\t}\n 189:\t\n 190:\t\trreq-\u003eprogress_at = rreq-\u003elen;\n 191:\t\n 192:\t\tret = netfs_single_begin_cache_read(rreq, ictx);\n 193:\t\tif (ret == -ENOMEM || ret == -EINTR || ret == -ERESTARTSYS)\n 194:\t\t\tgoto cleanup_free;\n 195:\t\n 196:\t\tnetfs_stat(\u0026netfs_n_rh_read_single);\n 197:\t\ttrace_netfs_read(rreq, 0, rreq-\u003elen, netfs_read_trace_read_single);\n 198:\t\n 199:\t\tnetfs_single_dispatch_read(rreq);\n 200:\t\n 201:\t\tret = netfs_wait_for_read(rreq);\n 202:\t\tnetfs_put_request(rreq, netfs_rreq_trace_put_return);\n 203:\t\treturn ret;\n 204:\t\n 205:\tcleanup_free:\n 206:\t\tnetfs_put_failed_request(rreq);\n 207:\t\treturn ret;\n 208:\t}\n 209:\tEXPORT_SYMBOL(netfs_read_single);\n"
} |
| 99/3 | 2026/10/06 21:33 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 113007 output: 31 thoughts: 165 Thoughts: **De-initializing the Network Filesystem Request** Okay, so I'm looking at `netfs_deinit_request` and walking through the clean up. First, I see the calls to `bvecq_pos_unset` on the various cursors: `load_cursor`, `dispatch_cursor`, and `collect_cursor`. Right, so that's expected. We're removing those references. Let's make sure, as the expert I am, that this is the best implementation. Ah, and I see we're then calling `bvecq_put(rreq->spare)`. Got it. We're returning the `spare` buffer. Now, let's focus on those cursor unsets. `dispatch_cursor`... I'm pretty sure that was initialized, so `bvecq_pos_unset(&rreq->dispatch_cursor)` will correctly decrement that reference count. Good. Now, `collect_cursor`... this is where I need to double-check. `bvecq_pos_unset(&rreq->collect_cursor)` handles a null pointer safely, which I think is a non-issue given we are dealing with a cursor and a buffer. Excellent. I'm going to quickly verify that `bvecq_put(NULL)` is indeed a no-op, just to confirm my assumptions and make absolutely sure everything is handled correctly, but it should be fine. |
| 100/3 | 2026/10/06 21:33 | tool | grepper |
0mArgs:{
"Expression": "void bvecq_put",
"PathPrefixes": [
"fs/netfs"
]
}Results: {
"Output": "fs/netfs/bvecq.c=231=static void bvecq_free_slot(struct bvecq *bq, unsigned int slot)\n--\nfs/netfs/bvecq.c-262- */\nfs/netfs/bvecq.c:263:void bvecq_put(struct bvecq *bq)\nfs/netfs/bvecq.c-264-{\n"
} |
| 101/3 | 2026/10/06 21:33 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 113621 output: 38 thoughts: 28 Thoughts: Okay, here's my summary, taking the role of the expert: **Diving into `bvecq_put`** Right, let's take a look at `bvecq_put`. The reference points to line 263 in `fs/netfs/bvecq.c`. I'm familiar with this area; `bvecq` has always been a key data structure. I'm likely interested in a specific use-case or optimization related to how we're pushing data into it. I'll need to remember that this `put` function implies a "producer" side, the process adding the data. I expect to be able to quickly understand its responsibility for this portion of memory management: updating counters, managing locks, handling potential full queues and related error handling. I'll be keeping an eye out for potential concurrency issues or memory allocation problems, always. |
| 102/3 | 2026/10/06 21:33 | tool | read-file |
0mArgs:{
"File": "fs/netfs/bvecq.c",
"FirstLine": 263,
"LineCount": 20
}Results: {
"Contents": " 263:\tvoid bvecq_put(struct bvecq *bq)\n 264:\t{\n 265:\t\tstruct bvecq *next;\n 266:\t\n 267:\t\tfor (; bq; bq = next) {\n 268:\t\t\tif (!refcount_dec_and_test(\u0026bq-\u003eref))\n 269:\t\t\t\tbreak;\n 270:\t\t\tfor (int slot = 0; slot \u003c bq-\u003enr_slots; slot++)\n 271:\t\t\t\tbvecq_free_slot(bq, slot);\n 272:\t\t\tnext = bq-\u003enext;\n 273:\t\t\tnetfs_stat_d(\u0026netfs_n_bvecq);\n 274:\t\t\tif (bq-\u003efrom_pool)\n 275:\t\t\t\tmempool_free(bq, \u0026netfs_bvecq_pool);\n 276:\t\t\telse\n 277:\t\t\t\tkfree(bq);\n 278:\t\t}\n 279:\t}\n 280:\tEXPORT_SYMBOL(bvecq_put);\n 281:\t\n 282:\t/**\n"
} |
| 103/3 | 2026/10/06 21:33 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 114178 output: 470 thoughts: 3597 Thoughts: **KMSAN Assessment: Netfs Buffer Management Refactoring**
Alright, after meticulously reviewing this netfs refactoring, I've concluded that `NeedsKMSAN=false`. My rationale is multi-faceted, stemming from a deep dive into the code's data structures, initialization, and interactions. This patch series significantly refactors netfs's buffer management, moving away from the `rolling_buffer` approach to a segmented `bvecq` (vector of buffer descriptors) with associated cursor tracking (`bvecq_pos`). This restructuring, of course, means carefully examining potential flaws introduced.
First, I examined every structure involved: `struct netfs_io_request` (`rreq`), `struct netfs_io_subrequest` (`subreq`), `struct bvecq` and `struct bvecq_pos` and `struct readahead_control`. Crucially, none of these are ever directly exposed to userspace via `copy_to_user`, `ioctl` calls, or the like. Netfs is fundamentally an internal caching and buffering layer for network file systems. User-space interactions remain limited to the familiar `read`/`write` system calls, or direct I/O, which transfer data between user buffers and the kernel.
Next, I meticulously verified that all newly introduced variables and structure fields are *always* initialized. Every structure allocated (e.g., `rreq`, `subreq`, `bvecq`) is zeroed upon allocation using `memset`, designated initializers (`= {}`), or `kmalloc` with explicit field assignment. This eliminates the risk of conditional logic or branching that relies on uninitialized values. Furthermore, `bvecq` itself is constructed using an explicit initializer `{ ... }`, ensuring proper initialization. The code utilizes `bvec_set_page()` / `bvec_set_folio()` to ensure `bvecq->__bv[]` entries are properly populated before the memory is exposed.
A major concern was the potential for uninitialized memory usage in bounds checks or during buffer operations. I investigated the write paths in `afs_dir.c`, `afs_symlink.c`, and `cachefiles/io.c`. The code carefully calculates lengths and uses techniques such as rounding to ensure that writes do not exceed the provided lengths. I specifically reviewed how it deals with partial writes and gaps by zeroing unused portions of the destination buffers. Importantly, I examined `netfs_extract_iter` and its interaction with `iov_iter_extract_will_pin` or `BVECQ_MEM_PAGECACHE` to guarantee that reads from user space do not introduce uninitialized data. Specifically, when reading from a file hole or EOF, the destination buffer is explicitly zeroed using `bvecq_zero()` or other `folio_zero_segments()` functions.
This refactoring does introduce new risks such as use-after-free, refcount leaks, or null pointer dereferences, as well as concurrency issues within the `bvecq` structures, but these are precisely the sorts of errors that KASAN (Kernel Address Sanitizer), LOCKDEP (Lock Dependency Validator), and other standard kernel debugging tools are designed to catch. The risk of uninitialized memory issues that only KMSAN could detect is demonstrably absent.
Therefore, given the rigorous initialization and lack of kernel-to-user info leaks, a dedicated KMSAN fuzzing session isn't warranted here.
|
| 104/3 | 2026/10/06 21:34 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch series refactors netfs and related filesystem code (AFS, cachefiles) by replacing the legacy rolling_buffer abstraction with segmented bio_vec queues (bvecq) and cursor tracking (bvecq_pos: load_cursor, dispatch_cursor, collect_cursor).\n\nKMSAN vs KASAN Applicability Analysis:\n1. No uninitialized struct/stack/heap usage: All new and modified structures are cleanly initialized before use. Specifically, netfs_io_request and netfs_io_subrequest are fully zeroed via memset() in netfs_alloc_request() and netfs_alloc_subrequest(); new stack cursor variables (such as dispatch_cursor) are zero-initialized with '= {}' or bvecq_pos_set(); bvecq structures are allocated and zero-initialized via C99 designated initializers in bvecq_alloc_one(); and readahead_control._nr_folios is initialized to zero upon allocation and appropriately tracked.\n2. No kernel-to-user info-leaks: None of the modified data structures are copied or exposed to user space (no copy_to_user, put_user, netlink skb, ioctl, or sockopt changes). Buffer contents transferred to user space in DIO reads are filled by network/cache reads or explicitly zeroed (via bvecq_zero, iov_iter_zero, or folio_zero_segments for gaps/unwritten ranges).\n3. Types of bugs exposed: The refactoring involves queue manipulation, cursor advancement, slot indexing, and reference counting (bvecq_get/bvecq_put). The potential defects here include out-of-bounds accesses on bvec array slots, NULL pointer dereferences on cursor exhaustion, and use-after-free or refcount issues on bvecq nodes. All of these error classes are fully detected by standard debugging tools (KASAN, refcount_t sanity checks, and slab poisoning).\n\nConsequently, the patch series does not introduce uninitialized memory read risks or info-leak vulnerabilities that would require a dedicated KMSAN fuzzing session.",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch series refactors netfs and related filesystem code (AFS, cachefiles) by replacing the legacy rolling_buffer abstraction with segmented bio_vec queues (bvecq) and cursor tracking (bvecq_pos: load_cursor, dispatch_cursor, collect_cursor).\n\nKMSAN vs KASAN Applicability Analysis:\n1. No uninitialized struct/stack/heap usage: All new and modified structures are cleanly initialized before use. Specifically, netfs_io_request and netfs_io_subrequest are fully zeroed via memset() in netfs_alloc_request() and netfs_alloc_subrequest(); new stack cursor variables (such as dispatch_cursor) are zero-initialized with '= {}' or bvecq_pos_set(); bvecq structures are allocated and zero-initialized via C99 designated initializers in bvecq_alloc_one(); and readahead_control._nr_folios is initialized to zero upon allocation and appropriately tracked.\n2. No kernel-to-user info-leaks: None of the modified data structures are copied or exposed to user space (no copy_to_user, put_user, netlink skb, ioctl, or sockopt changes). Buffer contents transferred to user space in DIO reads are filled by network/cache reads or explicitly zeroed (via bvecq_zero, iov_iter_zero, or folio_zero_segments for gaps/unwritten ranges).\n3. Types of bugs exposed: The refactoring involves queue manipulation, cursor advancement, slot indexing, and reference counting (bvecq_get/bvecq_put). The potential defects here include out-of-bounds accesses on bvec array slots, NULL pointer dereferences on cursor exhaustion, and use-after-free or refcount issues on bvecq nodes. All of these error classes are fully detected by standard debugging tools (KASAN, refcount_t sanity checks, and slab poisoning).\n\nConsequently, the patch series does not introduce uninitialized memory read risks or info-leak vulnerabilities that would require a dedicated KMSAN fuzzing session.",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|