Page MenuHomeFreeBSD

nvme: send the shutdown notification to all controllers before waiting
Needs ReviewPublic

Authored by wanpengqian_gmail.com on Wed, Oct 7, 1:11 AM.

Details

Reviewers
imp
Summary

At shutdown, an NVMe controller is told that the system is going down by
setting the shutdown notification bits (CC.SHN) in its configuration
register; it then flushes its caches and reports completion in its status
register (CSTS.SHST). nvme_ctrlr_shutdown() sets the bits and waits for
that report, and the kernel runs the devices' shutdown methods one after
the other, so with several controllers the waits add up: about 1.3 s for
each PM1725a, 2.6 s for two of them in series, more on a server with many
drives.

Send the notification to every initialized, non-failed controller from a
shutdown_final handler registered at SHUTDOWN_PRI_FIRST, before any
device's shutdown method runs, so the controllers do their shutdown work
at the same time; the shutdown methods then only wait for what is left.
The wait becomes the slowest drive's instead of the sum.

This is safe here: shutdown_final runs after the file systems are synced
and the buffer cache is flushed, and before any device is shut down, so no
I/O is in flight and the order of the device shutdowns (children before
parents) is unchanged; and the nvme shutdown method submits no commands
after setting CC.SHN. The latter matters because the specification lets
a controller stop processing commands once CC.SHN is set.

Linux has the same serial wait (device_shutdown() calls each driver's
.shutdown in turn). A 2022 patch split its nvme shutdown into a "set
CC.SHN" and a "wait" phase like this one; it was dropped because Linux's
nvme driver deletes its I/O queues with admin commands before it sets
CC.SHN, so the early notification is not safe there, and the maintainers
preferred a mechanism in the driver core that benefits every driver. That
replacement ("shut down devices asynchronously", v21 in September 2026)
runs each device's shutdown method in its own task with device links for
the ordering and leaves the nvme driver alone. FreeBSD's driver submits
nothing at shutdown and this change stays inside it.

Signed-off-by: Wanpeng Qian <wanpengqian@gmail.com>
Sponsored by: keelos.dev

Test Plan

A/B on a Supermicro X11SPW-TF (Xeon Silver 4110): FreeBSD main f958aa7e7 GENERIC without and with this change, both carrying a measuring patch that prints every device whose shutdown method takes more than 10 ms (in device_shutdown(), sys/kern/subr_bus.c; not part of this change). NVMe: two Samsung PM1725a (nvme1, nvme2) and one PM961 (nvme0, 14 ms, below the threshold). shutdown -r now three times with each kernel, alternating; numbers from the serial console.

nvme2 (shut down first)nvme1 (second)NVMe total
without1259–1262 ms1274–1276 ms2.54 s
with1259–1268 msbelow 10 ms, not printed1.26 s

With the change the second controller has already completed its shutdown by the time its method runs; the first one still waits its full time. The order of the device shutdowns is unchanged (nvme2, pci13, pcib13, pci12, pcib12, nvme1, pci11, ...), and so are the other devices' times (ixl0/ixl1 84 ms, the pci8/pci9 bridges 174–183 ms).

The same change has been in keelOS's stable/15 kernel since 2026-10-01 (device shutdown 2.8 s -> 1.5 s on this machine), through daily reboots on several machines.

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped
Build Status
Buildable 77785
Build 74668: arc lint + arc unit

Event Timeline

This is a very interesting notion... It might be worth doing generally in newbus. I'll have to think about that a little since it's late here.
It definitely would help the shutdown times...