get_name() opens the parent directory with dentry_open() and releases it with fput(). From a kernel thread, fput() defers the final __fput() to the delayed_fput work item. NFSD threads are kernel threads, and NFSD reaches get_name() through exportfs_decode_fh() whenever a file handle's dentry is not connected to the root. On a filesystem with more directories than the dcache retains, nearly every READDIRPLUS needs a reconnect. The deferred releases arrive faster than the work item retires them, and each pending file pins the directory's readdir state. On ext4 that state is the htree fname cache for the last leaf block read. A 64-thread NFSv3 server under a parallel tree walk showed unreclaimable slab growing by about 90 MB per second, and the same workload on a 5.15 kernel ended in a global OOM. Commit 5ff318f645eb ("nfsd: use __fput_sync() to avoid delayed closing of files.") stopped NFSD's own closes from feeding the delayed_fput queue but did not cover the exportfs reconnect path. Release the file with __fput_sync(). get_name() opened the file itself, read-only, on a directory, so __fput() involves nothing that could wait on the caller. Fixes: 4a9d4b024a31 ("switch fput to task_work_add") Reported-by: Danny Pidutti Closes: https://lore.kernel.org/linux-fsdevel/DS4PR04MB9846BA3C16BC5ACED8B589228D952@DS4PR04MB9846.namprd04.prod.outlook.com/ Signed-off-by: Chuck Lever --- fs/exportfs/expfs.c | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) Mail from kernel.org to hotmail.com was rejected. kernel.org appears to be on some deny-list. Reposting. diff --git a/fs/exportfs/expfs.c b/fs/exportfs/expfs.c index eafd99507afe..b5768f94fbb5 100644 --- a/fs/exportfs/expfs.c +++ b/fs/exportfs/expfs.c @@ -336,7 +336,11 @@ static int get_name(const struct path *path, char *name, struct dentry *child) } out_close: - fput(file); + /* + * @file is the read-only directory opened above, so __fput() + * cannot block on anything the caller holds. + */ + __fput_sync(file); out: return error; } -- 2.55.0