Skip to content

fix(vfs): prevent shutdown hangs during lmbench ext4_copy_files_bw cleanup by using a dedicated shutdown worker - #2343

Open
mistcoversmyeyes wants to merge 3 commits into
DragonOS-Community:masterfrom
mistcoversmyeyes:fix/ext4-truncate-drain-2303
Open

mistcoversmyeyes wants to merge 3 commits into
DragonOS-Community:masterfrom
mistcoversmyeyes:fix/ext4-truncate-drain-2303

Conversation

@mistcoversmyeyes

@mistcoversmyeyes mistcoversmyeyes commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Final superblock shutdown can deadlock on the system workqueue: it waits for the page-cache domain's admitted I/O to drain, while an ext4 delayed-allocation progress task holding the last I/O permit is queued behind it on that same single worker. A fresh guest running the loop-backed truncation regression reached this state after successful truncation, marker fsync and close.

Move final superblock shutdown to a dedicated queue initialized during kernel startup. Filesystem producers continue to run on their existing queues while shutdown waits for their permits. Preserve the existing shutdown ordering and drain requirements.

The regression creates two 64 MiB files on a disposable sparse 1 GiB ext4 fixture, truncates the second file after dirty, fsync or remounted preparation, and verifies inode identity, zero size, and marker persistence across another remount. Each mode is its own executable to fit the runner's existing 60-second per-binary deadline; all three retain the full workload and remount checks. File descriptors use RAII so fatal assertions cannot leave the loop mount busy.

Validation on x86_64 QEMU/KVM (2 vCPUs, 2 GiB), baseline b0dd898508479046b744e22591e10da2da9a85d0:

  • Baseline: regression stalled in final shutdown. A stopped-state inspection found the system worker waiting in PageCacheWritebackDomain::wait_drained, with one admitted I/O permit held by a delayed-allocation task still queued on that worker.
  • Fixed: all three Make-built static musl executables passed under separate timeout -k 2s 60s bounds. Run ext4_truncate_dirty_test, ext4_truncate_synced_test, and ext4_truncate_remounted_test from the installed dunitest tree with its fixture directory.
  • Three consecutive original lmdd if=/ext4/zero_file of=/ext4/test_file copies after the two-file preparation returned 0 under the original per-run deadline.
  • make kernel, Rust formatting and git diff --check passed.

Refs #2303 and #2286. This fixes the independently reproduced unmount deadlock exposed by the regression; it does not attribute the historical open(O_TRUNC) EIO to the same cause. The complete 48-case sequence remains unverified.

@github-actions github-actions Bot added Bug fix A bug is fixed in this pull request test Unitest/User space test labels Sep 24, 2026
@mistcoversmyeyes mistcoversmyeyes changed the title test(ext4): cover truncation after large delayed writes fix(vfs): drain superblocks on a dedicated shutdown worker Sep 24, 2026
@mistcoversmyeyes
mistcoversmyeyes marked this pull request as ready for review September 24, 2026 09:37
@github-actions github-actions Bot removed the test Unitest/User space test label Sep 24, 2026
@mistcoversmyeyes
mistcoversmyeyes marked this pull request as draft September 25, 2026 09:24
@mistcoversmyeyes
mistcoversmyeyes force-pushed the fix/ext4-truncate-drain-2303 branch 2 times, most recently from 5c4b829 to 18eaaba Compare September 27, 2026 05:56
@mistcoversmyeyes

Copy link
Copy Markdown
Collaborator Author

补充说明:为什么在 VFS mount 模块中新增静态 SUPERBLOCK_SHUTDOWN_WQ。

schedule_final_shutdown(mount: Arc<MountFS>) 在调用方确认 superblock 已无活动挂载和外部 pin、将其转为 Dying 后,投递最终清理任务。它本身只安排执行;实际同步、后台工作排空和资源回收由 finish_final_shutdown() 完成。

没有外部使用者,不代表后台写回已经完成。 此前已接受的写入可能仍在内存脏页中,后台 producer 也可能仍持有 domain I/O permit。最终清理需要推进这些工作并等待它们结束,才能安全地回收缓存、inode 和文件系统状态。不能通过提前释放 permit 或跳过 drain 来绕过这个依赖。

当前 SYSTEM_WQ 只有一个 worker,按顺序执行任务;ext4 的写回推进任务通过 schedule_work() 进入该队列,并从排队时就持有 domain I/O permit。如果最终清理也在同一个 worker 上执行,而它所等待的 producer 尚在队列中,就会出现以下循环等待:

                  清理任务 shutdown
                   |             ^
    等待 producer 完成             | 等待清理任务返回,
    并释放 I/O permit |             | 才能执行下一项
                   v             |
             写回任务 --------> SYSTEM_WQ
             producer           唯一 worker
                    等待获得执行机会

图中箭头表示“等待”。具体时序可以是:

  1. shutdown 开始占用 SYSTEM_WQ 的唯一 worker。
  2. 某个已获准的 producer 持有 permit,等待该 worker 执行;它可能在 shutdown 后面排队,或在 shutdown 执行期间入队。
  3. shutdown 在同步、停止 producer 或排空 domain 的过程中,等待相关后台工作完成。
  4. producer 无法执行和释放 permit,shutdown 无法结束,worker 也就无法取出 producer。

这是执行资源上的循环等待。shutdown 即使睡眠并让出 CPU,worker 也仍停留在该任务的调用栈中,不会自动取下一项工作。队列和计数器正确加锁并不能消除这个死锁;仅调整入队顺序也不足以保证安全,因为后台工作可以继续派生任务。

专用 WQ 将“等待清理完成的执行者”和“推进被等待工作的执行者”分开: sb_shutdown worker 可以等待,而 SYSTEM_WQ worker 仍能运行 producer、释放 permit,随后唤醒清理路径。这里复用现有 WorkQueue 的排队、唤醒和内核线程机制,无须另建一套调度实现。

队列放在 mount 模块内部,是因为最终 superblock 生命周期由这一层协调,具体文件系统仍负责自己的同步和回收操作。它是模块私有、全局共享的一条静态队列,不是每个 MountFS 或每个 superblock 创建一条队列;启动时预初始化,也避免在最后一个外部 pin 的释放路径上临时创建 worker。

这一选择保留了一个明确限制:不同 superblock 的最终 shutdown 仍串行执行,慢清理可能阻塞后续清理。它针对的是上述与 SYSTEM_WQ producer 共用唯一 worker 的依赖环,并不声称消除了所有潜在等待关系。

@mistcoversmyeyes
mistcoversmyeyes marked this pull request as ready for review September 28, 2026 16:10
@mistcoversmyeyes mistcoversmyeyes changed the title fix(vfs): drain superblocks on a dedicated shutdown worker fix(vfs): prevent shutdown hangs during lmbench ext4_copy_files_bw cleanup by using a dedicated shutdown worker Sep 30, 2026
Preserve the two 64 MiB file preparation and vary dirty, fsync and remounted state. Verify successful truncation retains inode identity and new contents survive remount.

Refs: DragonOS-Community#2303, DragonOS-Community#2286
@mistcoversmyeyes
mistcoversmyeyes force-pushed the fix/ext4-truncate-drain-2303 branch from d5cdce7 to 1329a21 Compare October 2, 2026 16:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Bug fix A bug is fixed in this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug(ext4,vfs): ext4_copy_files_bw test_case blocked when opening test_file with flag OTRUNC

1 participant