With direct dispatch, the receive task transmits each forwarded packet
from within if_input(), and on a lightly used ring every packet rings
the doorbell. 4a0ec469934a stopped iflib from ringing when nothing was
pending, which took three register writes per packet down to one; this
rings the remaining one once per receive burst. On an Intel I226 (igc)
a write costs about 400 cycles, a fifth of the cost of forwarding a
packet with pf disabled.
While a receive task passes a burst up, including its LRO flush, the
transmit queues it uses, up to eight, defer their doorbells, and the
task rings each of them once at the end of the burst. A queue defers
only while it is at most an eighth full, where iflib_txd_db_check()
would ring for every packet; above that iflib already batches on its
own. Both the mp_ring and the simple transmit paths take part. Other
transmitters, such as the transmit task when tx_abdicate is set, ring
as before, and nothing changes when net.iflib.min_tx_latency is set.
The deferral is off by default until it has been exercised on more
drivers and with TCP endpoints; net.iflib.tx_db_burst=1 turns it on.
The burst lives on the receive task's stack, the task stays pinned, and
the transmit path finds it on a per-CPU list by thread, so a thread
that preempts the task does not take part. A queue counts the bursts
holding its doorbell, and on the mp_ring path the transmit task leaves
such descriptors to them: taking over the ring for them would move the
draining to a task that rings per packet, and a backlog builds whenever
it falls behind. mlx5en has held its send queues' doorbells during
receive processing since 2d5e5a0d75b0; DragonFly's ifq staging and
OpenBSD's transmit mitigation batch packets at the ifq layer instead.
Sponsored by: Rubicon Communications, LLC ("Netgate")