Summary
With soft updates, mkdir(2), link(2) and rename(2) into a directory whose i_nlink is at UFS_LINK_MAX while i_effnlink is below it call ufs_sync_nlink1(). It syncs the file system and returns ERELOOKUP so that the syscall restarts. It busied the mount with vfs_busy(mp, 0), which sleeps for as long as an unmount is in progress. That deadlocks with the unmount:
- The caller is still between vn_start_write() and vn_finished_write(). kern_mkdirat(), kern_linkat() and kern_renameat() keep the write going across the VOP.
- The unmount sets MNTK_UNMOUNT, then goes on to VFS_UNMOUNT(). There, ffs_unmount() suspends writes and waits in vfs_write_suspend() for that write to end.
- After that, every vfs_busy() of the mount hangs too. That includes getfsstat(2) with MNT_WAIT (mount -v, mount -p, df) and lookups crossing the mount point. A reboot is the only way out, and on the test machine lookups piling up behind the covered vnode hung the whole system once.
The window is wide. After setting MNTK_UNMOUNT, dounmount() waits for the busy references to drain, which includes another thread already syncing in ufs_sync_nlink1(). It then runs vfs_periodic(). A mkdir(2) that starts in that time sleeps in vfs_busy() and the unmount never gets to suspend writes.
This is the deadlock that 29d03af1e4da fixed for VOP_RENAME() in 2014 ("Busying mp after vn_start_write() deadlocks the unmount"), and the same reasoning applies. The caller's vn_start_write() already keeps the unmount from getting past the write suspension, and ffs_unmount() tears nothing down before that point. So the mount does not need to be busied for VFS_SYNC(). The reference the callers take with vfs_ref() still keeps the mount structure around.
Found in the wild: gnop10.sh deadlocked this way on a non-debug kernel. The forced unmount there came from the deferred unmount task after injected I/O errors.
Fixes: 8db679af66b0 ("UFS: make mkdir() and link() reliable when using SU and reaching nlink limit"), bc6d0d72f4f4 ("UFS rename: make it reliable when using SU and reaching nlink limit")
Test
New stress2 test nlink6.sh:
- It fills a directory on a soft updates file system up to UFS_LINK_MAX.
- Eight workers remove and recreate subdirectories in it, so mkdir(2) keeps going through ufs_sync_nlink1().
- Meanwhile the script force-unmounts, runs fsck_ffs -fy and remounts in a loop, for 300 seconds by default.
- An unmount that does not return within 120 seconds is reported with the kernel stacks of the unmount and of the workers.
The workers chdir() into the directory and use relative names. Lookups through the mount point would wait for the unmount on the covered vnode lock and never get to race with it.
With dtrace=1 the test counts, while an unmount is in progress:
- the vfs_busy() calls, by caller
- the mkdir(2) calls that went through the sync and restarted with ERELOOKUP, which shows that the race window was hit
When it fails, the mount point stays wedged until reboot. Running it with a private mount point, such as mntpoint=/var/tmp/nlink6/mnt, keeps the hang from spreading to /mnt.
Test Plan
All testing was on AWS EC2, 4 vCPUs, 16 GB RAM.
Without the fix
nlink6.sh deadlocked during its first forced unmount:
umount mi_switch _sleep vfs_write_suspend vfs_write_suspend_umnt ffs_unmount dounmount kern_unmount nlink6 mi_switch _sleep vfs_busy ufs_sync_nlink ufs_mkdir VOP_MKDIR_APV kern_mkdirat nlink6 mi_switch _sleep vn_start_write_refed vn_start_write kern_frmdirat
Seven workers were in vfs_busy() and one in vn_start_write(), waiting for the suspension. These are the same stacks as the original gnop10.sh hang, where the unmount was in the deferred unmount task.
With the fix
- nlink6.sh, 300 seconds, run twice: 30 and 19 forced unmounts, no hang, every fsck_ffs clean.
- The race was hit: 8 mkdir(2) calls synced and restarted while an unmount was in progress, both in a 120 second run and in the 300 second run. No vfs_busy() from UFS was seen during an unmount.
- The stress2 tests below all passed on the same kernel, with no panics and only the expected kernel messages from the injected errors and forced unmounts:
- gnop10.sh, the test that originally hung, plus gnop9.sh, gnop7.sh (SU and SU+J) and fsck6.sh
- fdatasync.sh and fdatasync2.sh on UFS without soft updates, with soft updates and with SU+J
- ftruncate2.sh, fsync2.sh and fsync3.sh
- 10-minute marcus.cfg loads on -o sync mounts: soft updates, SU+J, and soft updates with snapshots
- nlink.sh through nlink5.sh were not run separately. They exercise the same ERELOOKUP path without an unmount, and nlink6.sh goes through it continuously.