iflib_netmap_txsync() skips zero-length fragments when it collects the
segments of a packet, so an empty packet (a single slot of length 0)
reaches the driver's txd_encap with ipi_nsegs == 0. em, igb and ix
then set the EOP bit through the pointer to the last descriptor, which
is still NULL:
Fatal trap 12: page fault while in kernel mode fault virtual address = 0x8 igb_isc_txd_encap() iflib_netmap_txsync() netmap_bwrap_notify() nm_vale_flush() netmap_vale_vp_txsync()
A VALE switch produces such packets by design: nm_vale_flush() leases
slots in the destination ring, and slots it leased and did not need (a
nearly full ring, a GSO packet cut into fewer frames than the worst
case) are filled with length 0 unless the sender holds the last lease.
A NIC attached with valectl -a or -h gets them in its TX ring.
Linux's AF_XDP does not let such a packet reach a driver either: the
core rejects a TX descriptor of length 0 (xp_aligned_validate_desc()).
Here netmap slots and NIC descriptors go in step, so the slot cannot be
skipped. A descriptor of length 0 would do for e1000 and igb, whose
data sheets allow null descriptors, but not for every NIC: ixl asserts
that no buffer has a size of 0 (the ZERO_BSIZE malicious driver
detection event). Send a minimum size frame from the interface's own
address to itself instead (ethertype 0x88b5, local experimental). A
switch that has learned the address on that port filters the frame; a
copy that is flooded before that is addressed to no other host.
Tested with a program that puts empty packets into the TX ring of
netmap:igb0 (I350) and netmap:ixl0 (X722): without the change the first
one panics the kernel on both, with it every empty packet leaves as one
frame and the ring keeps going.
Signed-off-by: Wanpeng Qian <wanpengqian@gmail.com>
Sponsored by: keelos.dev