Page MenuHomeFreeBSD

nvme: reset the controllers in parallel at boot
Needs ReviewPublic

Authored by wanpengqian_gmail.com on Wed, Oct 7, 5:11 AM.

Details

Reviewers
imp
Summary

nvme_ctrlr_start_config_hook() runs in the boot thread, one controller after
another, and sleeps in nvme_ctrlr_hw_reset() until the controller reports
ready: about 2.3 s for each Samsung PM1725a, in series. A machine with
several enterprise NVMe drives spends that time added up.

Let the first hook start the reset of every controller that is still waiting
for its config hook, each in that controller's own taskqueue thread, and have
each hook wait for its own controller's reset. The rest of a hook -- identify,
the I/O queues, attaching the CAM bus -- runs as before, in the hooks' order,
so the scbus and nda unit numbers do not depend on which drive becomes ready
first. A controller that is somehow not in the nvme devclass (it cannot
happen in practice) resets itself in its hook, exactly as before.

The reset wait becomes the slowest drive's instead of the sum.

With one or two drives this saves little. It is meant for servers with many
NVMe drives, where the per-drive reset times add up, and for systems that have
a boot-time budget, where shortening the device-probe phase is worth it.

Signed-off-by: Wanpeng Qian <wanpengqian@gmail.com>
Sponsored by: keelos.dev

Test Plan

A/B on a Supermicro X11SPW-TF (Xeon Silver 4110, host TSC 3.600 GHz), FreeBSD main f958aa7e7 GENERIC booted on bare metal, without and with this change, both carrying a measuring patch that prints each controller's nvme_ctrlr_hw_reset() duration and the uptime at which it completes (not part of this change). NVMe: two Samsung PM1725a (nvme1, nvme2) and one PM961 (nvme0). Booted three times with each kernel.

controllerwithout: reset / completed atwith: reset / completed at
nvme0 (PM961)17 ms / 1.287 s0 ms / 1.297 s
nvme1 (PM1725a)2333 ms / 3.727 s2302 ms / 3.598 s
nvme2 (PM1725a)2443 ms / 6.188 s2310 ms / 3.607 s

Without the change the two enterprise drives reset one after the other (completing at 3.727 s and 6.188 s, ~2.4 s apart). With it they reset at the same time (completing at 3.598 s and 3.607 s, 9 ms apart), so the probe reaches the last drive about 2.4 s sooner. Each controller's own reset takes the same ~2.3 s in both; only the overlap changes. The reset durations above are representative; the three runs agreed within ~15 ms.

scbus and nda unit numbers were identical across both kernels and all runs (nvme0->nda0, nvme1->nda1, nvme2->nda2). The guest/host behaves as before otherwise; the kernel builds with and without WITH_BHYVE_SNAPSHOT (modules unaffected).

The same change has been in keelOS's stable/15 kernel since 2026-10-01 through daily reboots on several machines (device-probe reset 4.6 s -> 1.3 s on a 2x PM1725a + 1x PM961 host).

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped
Build Status
Buildable 77792
Build 74675: arc lint + arc unit

Event Timeline

wanpengqian_gmail.com edited the test plan for this revision. (Show Details)

expand the commit message: who this helps (many-drive servers, boot-time-sensitive systems)

So why not just make the reset sequence in the startup not wait, but execute a state machine that does the init? We hold the CAM until all the busses are ready, so the nda numbering would still be stable...

Thanks. To make sure I understood: a state machine for the whole init would mean roughly

  1. nvme_ctrlr_hw_reset(): replace the pause_sbt() loops in nvme_ctrlr_wait_for_ready() with a callout that polls CSTS.RDY and advances the state (disable -> wait !RDY -> enable -> wait RDY); the same path is shared with nvme_ctrlr_reset_task(), so it would be converted too, not kept as a second copy.
  2. The admin steps after the reset: identify, set_num_qpairs, construct_io_qpairs and the rest of nvme_ctrlr_start() use nvme_completion_poll(); they would become completion callbacks that chain to the next state.
  3. The config hook only starts the machine and returns; the last state attaches the CAM bus, runs nvme_sysctl_initialize_ctrlr() and calls config_intrhook_disestablish(). Since scbus/nda numbers follow the order of xpt_bus_register() and the periph allocations (unless wired), the final attach step would still be done in nvme unit order, so the numbering stays what it is today.

Is that what you have in mind? It is a fair amount of churn in the reset path, so I want to confirm the direction before doing it.

A middle ground, if you would accept it: the config hook returns at once and the existing (synchronous) init runs in the controller's own taskqueue thread, with the CAM attach ordered by unit number and config_intrhook_disestablish() at the end of the task. Every controller then resets, identifies and builds its queues in parallel, the hooks do not block the boot thread, the numbering is unchanged, and the code stays close to what is there now. I am fine with either; tell me which you prefer and I will update the review.