Page MenuHomeFreeBSD

iflib: Ring transmit doorbells once per receive batch
Needs ReviewPublic

Authored by rcm on Tue, Oct 6, 6:43 PM.
Tags
None
Referenced Files
F175118217: D60416.id188906.diff
Thu, Oct 8, 9:37 AM
Unknown Object (File)
Wed, Oct 7, 1:54 PM
Unknown Object (File)
Wed, Oct 7, 10:07 AM
Unknown Object (File)
Wed, Oct 7, 10:06 AM
Unknown Object (File)
Wed, Oct 7, 10:04 AM
Unknown Object (File)
Wed, Oct 7, 4:59 AM
Unknown Object (File)
Wed, Oct 7, 4:00 AM
Unknown Object (File)
Wed, Oct 7, 12:55 AM

Details

Summary

With direct dispatch, the receive task transmits each forwarded packet
from within if_input(), and on a lightly used ring every packet rings
the doorbell. 4a0ec469934a stopped iflib from ringing when nothing was
pending, which took three register writes per packet down to one; this
rings the remaining one once per receive batch. On an Intel I226 (igc)
a write costs about 400 cycles, a fifth of the cost of forwarding a
packet with pf disabled.

For the whole of a receive task's pass over its ring, including the LRO
flush, the transmit queues it uses, up to eight, defer their doorbells,
and the task rings each of them once at the end of the pass. The whole
pass is covered, not just the if_input() call after it, because with
LRO enabled and IP forwarding on, tcp_lro_queue_mbuf() hands each
packet to if_input() from inside the receive loop; igc, em, igb, ix and
ixl enable LRO by default. A queue defers only while it is at most an
eighth full, where iflib_txd_db_check() would ring for every packet;
above that iflib already batches on its own. Both the mp_ring and the
simple transmit paths take part. Other transmitters, such as the
transmit task when tx_abdicate is set, ring as before, and nothing
changes when net.iflib.min_tx_latency is set. Batched doorbells are
off by default until they have been exercised on more drivers;
net.iflib.tx_db_batched=1 turns them on.

The batch lives on the receive task's stack, the task stays pinned, and
the transmit path finds it on a per-CPU list by thread, so a thread
that preempts the task does not take part. A queue counts the batches
holding its doorbell, and on the mp_ring path the transmit task leaves
such descriptors to them: taking over the ring for them would move the
draining to a task that rings per packet, and a backlog builds whenever
it falls behind. mlx5en has held its send queues' doorbells during
receive processing since 2d5e5a0d75b0; DragonFly's ifq staging and
OpenBSD's transmit mitigation batch packets at the ifq layer instead.

Sponsored by: Rubicon Communications, LLC ("Netgate")

Test Plan

On a Netgate 4200 (4 Atom cores, Intel I226) forwarding 64-byte UDP,
one core rings once per 16 packets instead of once per packet and
forwards 1.35 instead of 1.10 Mpps with pf disabled and 0.44 instead of
0.38 Mpps with pf and 198k states; with LRO enabled it forwards 1.31
instead of 1.08 Mpps. With four queues per port it forwards up to
3.5-3.6 instead of 2.6 Mpps one way and 5.4-5.6 instead of 3.9-4.2 Mpps
both ways, on either transmit path. As a TCP endpoint with LRO, it
receives at line rate with one doorbell per 3.5 ACKs instead of one per
ACK, and sends with one per 9 segments instead of nearly one per
segment.

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped

Event Timeline

rcm requested review of this revision.Tue, Oct 6, 6:43 PM
rcm edited the summary of this revision. (Show Details)
rcm edited the test plan for this revision. (Show Details)

Start the burst before the receive loop rather than after it, and end it on the error path too, so that packets tcp_lro_queue_mbuf() hands to if_input() from inside the loop are covered; that happens whenever LRO is enabled with IP forwarding on, and igc, em, igb, ix and ixl enable LRO by default.

rcm retitled this revision from iflib: Ring transmit doorbells once per receive burst to iflib: Ring transmit doorbells once per receive batch.
rcm edited the summary of this revision. (Show Details)

s/burst/batch/g

no functional change intended

I need to think about this more. For your workload, I see the benefit. For mine (busy CDN server), where we're running TCP, and where our TX rings are constantly > 80% full, I think this just adds overhead (due to full rings) and risk (due to TCP).

TCP is also interesting, because one RX packet can generate megabytes of TX. I'm not sure its appropriate at all for TCP. I'm going to add some folks from the transport side.

Can you make it as cheap as possible to disable? Eg, iflib_txdb_burst() checks iflib_tx_db_burst and returns immediately when its disabled, and doesn't check the pcpu pointer, etc. I can see how that would make it difficult to handle changes to iflib_tx_db_burst , but I'd rather have that as an RDTUN and have this as cheap to disable as possible.

I need to think about this more. For your workload, I see the benefit. For mine (busy CDN server), where we're running TCP, and where our TX rings are constantly > 80% full, I think this just adds overhead (due to full rings) and risk (due to TCP).

TCP is also interesting, because one RX packet can generate megabytes of TX. I'm not sure its appropriate at all for TCP. I'm going to add some folks from the transport side.

I think your TX rings a full, because they are serving a lot of TCP connections. A single incoming TCP segment belongs to a single TCP connection and only triggers the sending of TCP segments of that TCP connection. A single TCP should not send megabytes of data in response to a single incoming segment. Usually the size of such bursts is limited by a number. Sending a megabytes of data at wire speed might result in packet loss on the way to the destination if some other link is the bandwidth bottleneck.

Can you make it as cheap as possible to disable? Eg, iflib_txdb_burst() checks iflib_tx_db_burst and returns immediately when its disabled, and doesn't check the pcpu pointer, etc. I can see how that would make it difficult to handle changes to iflib_tx_db_burst , but I'd rather have that as an RDTUN and have this as cheap to disable as possible.