mmu_radix_sync_icache() walks the page tables with an unlocked
pmap_extract() and passes the result straight to PHYS_TO_DMAP(), checking
only that it is non-zero. Nothing keeps the mapping - or the page table
page holding it - alive across that window: if another thread of the same
process tears a mapping down concurrently, the page table page can be
freed and reused, so pmap_extract() reads arbitrary memory and returns a
bogus physical address. __syncicache() then dereferences an unmapped
direct map address and the kernel takes a data storage interrupt:
fatal kernel trap:
exception = 0x300 (data storage interrupt)
virtual address = 0xc003317ca6022a00
dsisr = 0x40000000
srr0 = 0xc000000000f59460 (__syncicache)
lr = 0xc000000000f23588 (mmu_radix_sync_icache)
pid = 23878, comm = skyframe-evaluator-
panic: data storage interrupt trapThe faulting addresses decode to physical addresses far beyond installed
memory (~140 TB and ~900 TB on a 256 GB machine), i.e. translations that
never existed.
The hash MMU implementation of the same method, moea64_sync_icache(),
already holds PMAP_LOCK() across the loop; do the same here.
mmu_radix_extract() does not acquire the pmap lock itself, so this
introduces no recursion.
JIT workloads reach this path constantly: ppc_instr_emulate() calls
pmap_sync_icache() on the faulting address for the SIGILL "second chance"
retry, so a multithreaded JVM executing freshly written code races against
its own threads' mmap/munmap. Every panic observed here was in a JVM
thread.
MFC after: 1 week