flush_cached_blocks() drops CACHE_MTX before invoking the channel's write_error handler, then re-acquires it and jumps to the retry label. The retry label is above the while loop, whose first statement acquires CACHE_MTX again: retry: while (errors_found) { if ((flags & FLUSH_NOLOCK) == 0) mutex_lock(data, CACHE_MTX); <- second ... mutex_unlock(data, CACHE_MTX); (channel->write_error)(...); ... mutex_lock(data, CACHE_MTX); <- first goto retry; The mutex is created with pthread_mutex_init(..., NULL), so it is not recursive and the second acquisition blocks the thread that already holds it. The first flush that has to report a write error through a registered handler deadlocks, every time. Found by fuzzing e2fsck with corrupt images. An image whose group descriptors force bitmap relocation makes e2fsck -fn fail hundreds of writes on the read-only device; the aborted run reaches fatal_error(), whose io_channel_flush() enters this loop. e2fsck then sleeps forever at 0% CPU and does not respond to SIGTERM, so an init script waiting on fsck waits indefinitely. Read-only checking is enough to reach it. unfixed: killed after 35s, no progress fixed: exits 12 after 13ms Drop the redundant acquisition; the loop head takes the lock. Signed-off-by: Matthias Goergens --- lib/ext2fs/unix_io.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/lib/ext2fs/unix_io.c b/lib/ext2fs/unix_io.c index abd33ba..29968e3 100644 --- a/lib/ext2fs/unix_io.c +++ b/lib/ext2fs/unix_io.c @@ -738,7 +738,7 @@ retry: retval2); if (err_buf) ext2fs_free_mem(&err_buf); - mutex_lock(data, CACHE_MTX); + /* the loop head re-acquires CACHE_MTX */ goto retry; } else cache->write_err = 0; -- 2.55.0