Page MenuHomeFreeBSD

iflib: Ring transmit doorbells once per receive burst
Needs ReviewPublic

Authored by rcm on Tue, Oct 6, 6:43 PM.

Details

Reviewers
kbowling
gallatin
shurd
Group Reviewers
iflib
Summary

With direct dispatch, the receive task transmits each forwarded packet
from within if_input(), and on a lightly used ring every packet rings
the doorbell. 4a0ec469934a stopped iflib from ringing when nothing was
pending, which took three register writes per packet down to one; this
rings the remaining one once per receive burst. On an Intel I226 (igc)
a write costs about 400 cycles, a fifth of the cost of forwarding a
packet with pf disabled.

While a receive task passes a burst up, including its LRO flush, the
transmit queues it uses, up to eight, defer their doorbells, and the
task rings each of them once at the end of the burst. A queue defers
only while it is at most an eighth full, where iflib_txd_db_check()
would ring for every packet; above that iflib already batches on its
own. Both the mp_ring and the simple transmit paths take part. Other
transmitters, such as the transmit task when tx_abdicate is set, ring
as before, and nothing changes when net.iflib.min_tx_latency is set.

The deferral is off by default until it has been exercised on more
drivers and with TCP endpoints; net.iflib.tx_db_burst=1 turns it on.

The burst lives on the receive task's stack, the task stays pinned, and
the transmit path finds it on a per-CPU list by thread, so a thread
that preempts the task does not take part. A queue counts the bursts
holding its doorbell, and on the mp_ring path the transmit task leaves
such descriptors to them: taking over the ring for them would move the
draining to a task that rings per packet, and a backlog builds whenever
it falls behind. mlx5en has held its send queues' doorbells during
receive processing since 2d5e5a0d75b0; DragonFly's ifq staging and
OpenBSD's transmit mitigation batch packets at the ifq layer instead.

Sponsored by: Rubicon Communications, LLC ("Netgate")

Test Plan

On a Netgate 4200 (4 Atom cores, Intel I226) forwarding 64-byte UDP,
one core rings once per 16 packets instead of once per packet and
forwards 1.35 instead of 1.10 Mpps with pf disabled and 0.44 instead of
0.38 Mpps with pf and 198k states. With four queues per port it
forwards up to 3.5-3.6 instead of 2.6 Mpps one way and 5.4-5.6 instead
of 3.9-4.2 Mpps both ways, on either transmit path.

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped