nbd_pending_cmd_work() completes the request with BLK_STS_IOERR once its retry loop passes req->deadline, but leaves nsock->pending set. Nothing else clears it: nbd_send_cmd() only does so at its out: label after a complete send, and nbd_requeue_cmd() does not touch it. Every later request on that socket then takes this branch in nbd_handle_cmd(): if (unlikely(nsock->pending && nsock->pending != req)) { nbd_requeue_cmd(cmd); and is requeued again, forever, since the requeue path does not clear ->pending either. If the tag is recycled first, ->pending matches the new request instead and nbd_send_cmd() resumes it in place of the abandoned one, skipping its header and leaving cmd_cookie alone. Neither case leaves a usable socket. The header went out and the rest of the payload never will, so the stream no longer matches what the server expects. Mark the socket dead, which shuts it down and clears the partial send state. nbd_xmit_timeout() does the same for a timed out request and only skips it here because NBD_CMD_PARTIAL_SEND makes it defer to this work function. Without this, a write issued after the deadline fires never completes and the device has to be torn down to recover. Reproduced by making the resumed send never progress, so the retry loop runs out req->deadline. With the socket marked dead a later request is dispatched instead of requeued, and a reconnect afterwards does clean O_DIRECT I/O with no oops or warning. Fixes: 8337b029f788 ("nbd: fix partial sending") Cc: stable@vger.kernel.org Signed-off-by: Joseph Qi --- drivers/block/nbd.c | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/drivers/block/nbd.c b/drivers/block/nbd.c index ffce519bf008..c2c3dbdd631f 100644 --- a/drivers/block/nbd.c +++ b/drivers/block/nbd.c @@ -838,6 +838,14 @@ static void nbd_pending_cmd_work(struct work_struct *work) /* don't bother timeout handler for partial sending */ if (READ_ONCE(jiffies) + msecs_to_jiffies(wait_ms) >= deadline) { cmd->status = BLK_STS_IOERR; + /* + * The header is on the wire but the rest of the payload + * never will be, so the stream is out of sync with the + * server. Marking the socket dead also drops the stale + * nsock->pending, which would otherwise make + * nbd_handle_cmd() requeue every later request forever. + */ + nbd_mark_nsock_dead(nbd, nsock, 1); blk_mq_complete_request(req); break; } -- 2.39.3