fix(vfs): prevent shutdown hangs during lmbench ext4_copy_files_bw cleanup by using a dedicated shutdown worker - #2343
Conversation
5c4b829 to
18eaaba
Compare
|
补充说明:为什么在 VFS mount 模块中新增静态
没有外部使用者,不代表后台写回已经完成。 此前已接受的写入可能仍在内存脏页中,后台 producer 也可能仍持有 domain I/O permit。最终清理需要推进这些工作并等待它们结束,才能安全地回收缓存、inode 和文件系统状态。不能通过提前释放 permit 或跳过 drain 来绕过这个依赖。 当前 图中箭头表示“等待”。具体时序可以是:
这是执行资源上的循环等待。shutdown 即使睡眠并让出 CPU,worker 也仍停留在该任务的调用栈中,不会自动取下一项工作。队列和计数器正确加锁并不能消除这个死锁;仅调整入队顺序也不足以保证安全,因为后台工作可以继续派生任务。 专用 WQ 将“等待清理完成的执行者”和“推进被等待工作的执行者”分开: 队列放在 mount 模块内部,是因为最终 superblock 生命周期由这一层协调,具体文件系统仍负责自己的同步和回收操作。它是模块私有、全局共享的一条静态队列,不是每个 MountFS 或每个 superblock 创建一条队列;启动时预初始化,也避免在最后一个外部 pin 的释放路径上临时创建 worker。 这一选择保留了一个明确限制:不同 superblock 的最终 shutdown 仍串行执行,慢清理可能阻塞后续清理。它针对的是上述与 |
ext4_copy_files_bw cleanup by using a dedicated shutdown worker
Preserve the two 64 MiB file preparation and vary dirty, fsync and remounted state. Verify successful truncation retains inode identity and new contents survive remount. Refs: DragonOS-Community#2303, DragonOS-Community#2286
d5cdce7 to
1329a21
Compare
Final superblock shutdown can deadlock on the system workqueue: it waits for the page-cache domain's admitted I/O to drain, while an ext4 delayed-allocation progress task holding the last I/O permit is queued behind it on that same single worker. A fresh guest running the loop-backed truncation regression reached this state after successful truncation, marker fsync and close.
Move final superblock shutdown to a dedicated queue initialized during kernel startup. Filesystem producers continue to run on their existing queues while shutdown waits for their permits. Preserve the existing shutdown ordering and drain requirements.
The regression creates two 64 MiB files on a disposable sparse 1 GiB ext4 fixture, truncates the second file after dirty, fsync or remounted preparation, and verifies inode identity, zero size, and marker persistence across another remount. Each mode is its own executable to fit the runner's existing 60-second per-binary deadline; all three retain the full workload and remount checks. File descriptors use RAII so fatal assertions cannot leave the loop mount busy.
Validation on x86_64 QEMU/KVM (2 vCPUs, 2 GiB), baseline
b0dd898508479046b744e22591e10da2da9a85d0:PageCacheWritebackDomain::wait_drained, with one admitted I/O permit held by a delayed-allocation task still queued on that worker.timeout -k 2s 60sbounds. Runext4_truncate_dirty_test,ext4_truncate_synced_test, andext4_truncate_remounted_testfrom the installed dunitest tree with its fixture directory.lmdd if=/ext4/zero_file of=/ext4/test_filecopies after the two-file preparation returned 0 under the original per-run deadline.make kernel, Rust formatting andgit diff --checkpassed.Refs #2303 and #2286. This fixes the independently reproduced unmount deadlock exposed by the regression; it does not attribute the historical
open(O_TRUNC)EIO to the same cause. The complete 48-case sequence remains unverified.