Page MenuHomeFreeBSD

bhyve: save the state of a PCI device's INTx in a snapshot
Needs ReviewPublic

Authored by wanpengqian_gmail.com on Fri, Oct 2, 10:14 AM.

Details

Reviewers
jhb
markj
Group Reviewers
bhyve
Summary

A device model saves whether it asserts its interrupt (the e1000's
esc_irq_asserted, AHCI's lintr, virtio's ISR), and the kernel saves the
assert count of the I/O APIC pin, but pi_lintr.state, which sits between
the two, was not saved. After a restore of a device that was asserting
INTx, the state was IDLE while the pin's count was 1. When the guest
then cleared the cause, pci_lintr_deassert() did nothing, as it only
deasserts from ASSERTED, and the level triggered pin stayed asserted for
good: the guest took the same interrupt again after every EOI.

Save the state with the rest of the PCI device. This changes the format
of the snapshot.

Signed-off-by: Wanpeng Qian <wanpengqian@gmail.com>
Sponsored by: keelos.dev

Test Plan

main (f958aa7e7) with WITH_BHYVE_SNAPSHOT. FreeBSD 16.0-CURRENT guest with two vCPUs, started with bhyveload: AHCI root disk, -s 3,e1000,tap0 (the guest's em0 uses INTx), an NVMe data disk (D60234).

The test: the host floods the guest with pings (ping -f), bhyvectl --suspend two seconds into the flood, then bhyve -r; 15 seconds later the host tries an ssh command in the guest and looks at the CPU time of the two vCPU threads. Repeated in a loop (lintrtest.sh below; run5.sh starts or restores the guest).

Before (with D60241, without which the restored guest's e1000 receives nothing, a different problem): in 10 rounds, 6 restored guests came back with one vCPU thread at 72 % CPU. In such a guest the vCPUs are in em_intr(), Xapic_isr1 and lock_delay(): the interrupt of em0 is taken again after every EOI. The state survives further suspends and restores of that guest, as the pin's assert count is saved each time.

After (with D60241): 26 restores (10 + 16), the vCPU threads of every restored guest at 1 to 4 %.

How it was found (FreeBSD 14.5): Windows guests, whose e1000 and AHCI drivers use INTx, hung after a restore with one vCPU at 100 %.

The PIRQ routing of the LPC bridge (for guests that use the 8259) is not part of the snapshot, before or after this change.

lintrtest.sh
#!/bin/sh
# lintrtest.sh N: on the host, N times: suspend the guest g1 while it is flooded with pings (its
# e1000 uses INTx, so the line is asserted most of the time), restore it, and see whether it answers.
# The guest is the one run5.sh starts (bhyveload, AHCI root disk, NVMe data disk). KIND=... picks
# the bhyve of the restore and of the restarts.
export LD_LIBRARY_PATH=/root/kit/fix
K=${KIND:-fix}
ok=0
for i in $(seq 1 ${1:-5}); do
	rm -f /root/ckp/*
	(ping -f -t 6 10.9.0.2 > /dev/null 2>&1 &)
	sleep 2
	/root/kit/fix/bhyvectl --suspend=/root/ckp/s.ckp --vm=g1 2>/dev/null
	for n in $(seq 1 60); do pgrep -q bhyve || break; sleep 1; done
	pkill nc; pkill ping; sleep 2
	(/root/run5.sh $K /root/ckp/s.ckp > /dev/null 2>&1 < /dev/null &)
	sleep 15
	cpu=$(top -b -H -p $(cat /root/bhyve.pid) 2>/dev/null | grep vcpu | sed 's/.* \([0-9.]*%\) bhyve.*/\1/' | tr '\n' ' ')
	if timeout 15 /root/gs true 2>/dev/null; then
		echo "round $i: the guest answers (vCPU threads at $cpu)"; ok=$((ok + 1))
	else
		echo "round $i: the guest does NOT answer (vCPU threads at $cpu)"
		kill $(cat /root/bhyve.pid); sleep 2; pkill nc
		(/root/run5.sh $K > /dev/null 2>&1 < /dev/null &); sleep 45
		for n in $(seq 1 40); do timeout 8 /root/gs true 2>/dev/null && break; sleep 3; done
	fi
done
echo "$ok of ${1:-5} restores ok"

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped
Build Status
Buildable 77613
Build 74496: arc lint + arc unit