When the guest resets the controller (CC.EN 1 -> 0), the I/O requests
that are in flight go on in the block layer. When one of them completed
after the reset, its completion was posted through the cqid of its
submission queue, which the reset has set to 0: it went into the admin
completion queue of the controller that the guest had enabled again in
the meantime, or, with the controller still disabled, failed the
assertion in pci_nvme_cq_update().
A guest that finds a completion it did not ask for in its new admin
queue does not recover. Windows 10, which resets the controller when a
command takes longer than its timeout, stopped with
WHEA_UNCORRECTABLE_ERROR (0x124); FreeBSD panics with "NVME polled
command failed to complete within 10s".
Count the resets, note the count in each I/O request and post the
completion only if there was no reset in between. A Dataset Management
command that has more ranges to deallocate stops there as well.
CSTS.RDY still waits for the requests in flight, as before.
Signed-off-by: Wanpeng Qian <wanpengqian@gmail.com>
Sponsored by: keelos.dev