Page MenuHomeFreeBSD

iflib: netmap: send a filler frame for an empty packet
Needs ReviewPublic

Authored by wanpengqian_gmail.com on Fri, Oct 2, 3:37 AM.

Details

Reviewers
shurd
vmaffione
kbowling
Group Reviewers
iflib
Summary

iflib_netmap_txsync() skips zero-length fragments when it collects the
segments of a packet, so an empty packet (a single slot of length 0)
reaches the driver's txd_encap with ipi_nsegs == 0. em, igb and ix
then set the EOP bit through the pointer to the last descriptor, which
is still NULL:

Fatal trap 12: page fault while in kernel mode
fault virtual address = 0x8
igb_isc_txd_encap()
iflib_netmap_txsync()
netmap_bwrap_notify()
nm_vale_flush()
netmap_vale_vp_txsync()

A VALE switch produces such packets by design: nm_vale_flush() leases
slots in the destination ring, and slots it leased and did not need (a
nearly full ring, a GSO packet cut into fewer frames than the worst
case) are filled with length 0 unless the sender holds the last lease.
A NIC attached with valectl -a or -h gets them in its TX ring.

Linux's AF_XDP does not let such a packet reach a driver either: the
core rejects a TX descriptor of length 0 (xp_aligned_validate_desc()).
Here netmap slots and NIC descriptors go in step, so the slot cannot be
skipped. A descriptor of length 0 would do for e1000 and igb, whose
data sheets allow null descriptors, but not for every NIC: ixl asserts
that no buffer has a size of 0 (the ZERO_BSIZE malicious driver
detection event). Send a minimum size frame from the interface's own
address to itself instead (ethertype 0x88b5, local experimental). A
switch that has learned the address on that port filters the frame; a
copy that is flooded before that is addressed to no other host.

Tested with a program that puts empty packets into the TX ring of
netmap:igb0 (I350) and netmap:ixl0 (X722): without the change the first
one panics the kernel on both, with it every empty packet leaves as one
frame and the ring keeps going.

Signed-off-by: Wanpeng Qian <wanpengqian@gmail.com>
Sponsored by: keelos.dev

Test Plan

Reproducer: nmempty.c (below). It opens netmap:IFNAME and sends COUNT packets on the first TX ring, by default every other one an empty packet (one slot, len 0), the rest 60-byte broadcast frames with ethertype 0x88b5.

On main (main at f958aa7e7 (2026-10-02), GENERIC amd64), in a bhyve guest with an e1000 NIC (em0, 82545 emulation):

  • With the change: nmempty em0 20. On the host, tcpdump on the guest NIC's tap shows 10 broadcast frames and 10 frames from em0's address to itself; the guest keeps running and the ring keeps going.
  • Without the change (the 16.0-CURRENT snapshot kernel main-n289650-36d3e711bc62): the first empty packet took the whole VM down, because the host's bhyve hit the assertion fixed in D60219 (e82545) at that moment. So I have no guest-side trace from main.

On hardware, with 14.5-RELEASE's kernel (keelOS, a FreeBSD 14.5 based system; its igb attachment has local SR-IOV changes in if_em.c, igb_txrx.c is unmodified; iflib_netmap_txsync() is identical in main): Supermicro X11SPW-TF with an I350-T2 (igb0) and the on-board X722 (ixl0).

Without the change:

  • nmempty igb0 4: the machine resets at the first empty packet (no console output captured for this run). The trap in the summary is from the same NIC attached to a VALE switch under load, three times; that switch had local changes in nm_vale_flush() (VLAN support) that made unused leases more frequent than in the stock code.
  • nmempty ixl0 4: panic: vm_fault_lookup: fault on nofault entry, with iflib_netmap_txsync+0x276 <- netmap_ioctl on the stack.

With the change:

  • nmempty igb0 20: dev.igb.0.mac_stats.good_pkts_txd goes up by 20, bcast_pkts_txd by 10. Another machine on the same switch sees the 10 broadcast frames and 1 of the 10 filler frames (the first one after the link came up, before the switch had learned the address).
  • nmempty ixl0 20 (hw.ixl.enable_head_writeback=1, with D60215 applied): dev.ixl.0.mac.ucast_pkts_txd and bcast_pkts_txd go up by 10 each, the TX ring is reclaimed completely (tail == head - 1), no malicious driver event.
  • A VALE switch with the NIC attached (valectl -h) and two bhyve guests: 20 minutes of mixed load without a panic.
nmempty.c
/*
 * nmempty IFNAME [COUNT [WAIT [BURST [EMPTY]]]]: sends COUNT packets on IFNAME's
 * first netmap TX ring: 60-byte broadcast frames from the interface's address with
 * ethertype 0x88b5 and, unless EMPTY is 0, every other one an empty packet (one
 * slot of length 0).  BURST (default 1) packets go into the ring before each
 * NIOCTXSYNC; 20 ms later a second NIOCTXSYNC only reclaims.
 * WAIT: seconds to wait after opening the port (link renegotiation, STP).
 * The interface leaves the host stack while this runs.
 *   cc -O2 -o nmempty nmempty.c
 */
#define NETMAP_WITH_LIBS
#include <sys/types.h>
#include <sys/ioctl.h>
#include <sys/socket.h>
#include <net/if.h>
#include <net/if_dl.h>
#include <net/netmap_user.h>
#include <ifaddrs.h>
#include <poll.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>

int
main(int argc, char **argv)
{
	unsigned char hdr[14] = { 0xff, 0xff, 0xff, 0xff, 0xff, 0xff,
	    0x02, 0x00, 0x00, 0x00, 0x00, 0x01, 0x88, 0xb5 };
	struct ifaddrs *ifa0, *ifa;
	struct nm_desc *d;
	struct netmap_ring *r;
	struct pollfd pfd;
	char name[64];
	int i, n, wait, stuck, burst, with_empty, empty = 0;

	if (argc < 2) {
		fprintf(stderr, "usage: nmempty IFNAME [COUNT [WAIT [BURST [EMPTY]]]]\n");
		return (1);
	}
	/* the frames' source: the interface's own address */
	if (getifaddrs(&ifa0) == 0) {
		for (ifa = ifa0; ifa != NULL; ifa = ifa->ifa_next)
			if (ifa->ifa_addr != NULL && ifa->ifa_addr->sa_family == AF_LINK &&
			    strcmp(ifa->ifa_name, argv[1]) == 0)
				memcpy(hdr + 6, LLADDR((struct sockaddr_dl *)ifa->ifa_addr), 6);
		freeifaddrs(ifa0);
	}
	n = argc > 2 ? atoi(argv[2]) : 10;
	wait = argc > 3 ? atoi(argv[3]) : 45;
	burst = argc > 4 && atoi(argv[4]) > 0 ? atoi(argv[4]) : 1;
	with_empty = argc > 5 ? atoi(argv[5]) : 1;
	snprintf(name, sizeof(name), "netmap:%s", argv[1]);
	if ((d = nm_open(name, NULL, 0, NULL)) == NULL) {
		perror(name);
		return (1);
	}
	sleep(wait);	/* the NIC reinitialises: the link comes back, a switch port's STP takes 30 s */
	r = NETMAP_TXRING(d->nifp, d->first_tx_ring);
	pfd.fd = d->fd;
	pfd.events = POLLOUT;
	for (i = 0; i < n; i++) {
		struct netmap_slot *slot;

		for (stuck = 0; nm_ring_empty(r); stuck++) {
			if (stuck == 5) {	/* 5 s without a free slot: the ring is not reclaimed */
				printf("%s: TX ring stuck after %d packets: head %u cur %u tail %u (of %u)\n",
				    argv[1], i, r->head, r->cur, r->tail, r->num_slots);
				nm_close(d);
				return (2);
			}
			poll(&pfd, 1, 1000);
			ioctl(d->fd, NIOCTXSYNC, NULL);
		}
		slot = &r->slot[r->head];
		if (with_empty && i % 2 == 1) {
			slot->len = 0;
			empty++;
		} else {
			char *buf = NETMAP_BUF(r, slot->buf_idx);

			memset(buf, 0, 60);
			memcpy(buf, hdr, sizeof(hdr));
			buf[14] = i >> 8;
			buf[15] = i;
			slot->len = 60;
		}
		slot->flags = 0;
		r->head = r->cur = nm_ring_next(r, r->head);
		if ((i + 1) % burst != 0 && i != n - 1)
			continue;
		ioctl(d->fd, NIOCTXSYNC, NULL);
		usleep(20000);
		ioctl(d->fd, NIOCTXSYNC, NULL);	/* nothing new: reclaims what the NIC has sent */
	}
	sleep(1);
	ioctl(d->fd, NIOCTXSYNC, NULL);
	printf("%s: %d packets sent, %d of them empty; TX ring head %u cur %u tail %u (of %u)\n",
	    argv[1], n, empty, r->head, r->cur, r->tail, r->num_slots);
	nm_close(d);
	return (0);
}

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped
Build Status
Buildable 77593
Build 74476: arc lint + arc unit

Event Timeline

Alternatives I considered, and I am happy to rework this in the direction you prefer:

  1. One zero-length segment (a null descriptor). The smallest change, and what I ran first on igb. Legal for e1000/igb by their data sheets, but ixl has MPASS(seglen != 0) for the ZERO_BSIZE MDD event, and I do not know about the other iflib drivers; it would need a per-driver flag saying that null descriptors are fine.
  2. Let every driver's txd_encap cope with ipi_nsegs == 0. Each of them would still have to put something into the descriptor, because the slot index and the descriptor index must stay in step, so this moves the same question into every driver.
  3. Keep empty packets out of NIC rings in VALE. They come from the lease scheme in nm_vale_flush(): a sender that is not the last lease holder cannot give slots back. Avoiding them needs either an exact slot count before leasing (hard for the offload mismatch path, which segments GSO packets) or holding the ring's lock across the copy. And it would not cover a netmap application that writes an empty slot itself, which should not be able to panic the kernel.
  4. Skip the slot and adjust nkr_hwofs, giving up the 1:1 mapping between slots and descriptors. The reclaim side (nr_hwtail from ift_cidx_processed) relies on it too; much more invasive.