Page MenuHomeFreeBSD

mpi3mr: Wait for pending IOCTLs to drain completely on driver unload
Needs ReviewPublic

Authored by chandrakanth.patil_broadcom.com on Sun, Oct 4, 1:22 PM.
Tags
None
Referenced Files
F174978240: D60331.diff
Wed, Oct 7, 8:15 AM
F174886015: D60331.diff
Tue, Oct 6, 6:44 PM
Unknown Object (File)
Tue, Oct 6, 4:52 AM
Unknown Object (File)
Tue, Oct 6, 3:01 AM
Unknown Object (File)
Mon, Oct 5, 3:46 AM
Subscribers
None

Details

Summary

During driver unload or module detach, the driver previously waited for
pending management IOCTLs to complete but capped the wait loop to a fixed
180-second timeout. If a long-running management command (such as firmware
flashing, diagnostic dump collection, or secure erase) took longer than 180
seconds, the detach sequence proceeded while the IOCTL was still actively
executing. This led to tearing down hardware resources, destroying mutexes,
and freeing driver memory structures while userland was still using them,
causing memory corruption.

Remove the arbitrary 180-second timeout and wait until all pending IOCTLs
have fully completed before proceeding with resource teardown. Because
all IOCTL commands have individual hardware timeouts, they are guaranteed
to complete or time out deterministically.

Test Plan
  • Clean build with WERROR=-Werror across FreeBSD 16, 15, and 14 with INVARIANTS/WITNESS enabled; git bisect verified.
  • Tested driver unload cycles while executing long-running StorCLI2 diagnostic and firmware management commands on SAS4116/SAS5116 controllers.
  • Verified that driver detach waits cleanly for active commands to finish before releasing hardware resources and driver memory.

Diff Detail

Lint
Lint Skipped
Unit
Tests Skipped

Event Timeline

sys/dev/mpi3mr/mpi3mr_pci.c
739

This strikes me as unsafe. We might hang here if there's a stuck ioctl, missed wakeup or similar bug.

sys/dev/mpi3mr/mpi3mr_pci.c
739

This strikes me as unsafe. We might hang here if there's a stuck ioctl, missed wakeup or similar bug.

Operations like firmware flash can have an IOCTL timeout of up to 500–600s. Waiting only 180s caused detach to free memory while an IOCTL was still running, might lead to a panic.

New IOCTLs are rejected once SHUTDOWN is set, and active ones always decrement pend_ioctls on completion or timeout.

Since 600s covers the maximum supported IOCTL timeout in unload, can we cap the wait at 600s in V2 patch to avoid an infinite loop?