Page MenuHomeFreeBSD

libalias: index fully specified inbound links by remote endpoint
Needs ReviewPublic

Authored by paulo_nlink.com.br on Sun, Sep 27, 9:22 PM.
Tags
None
Referenced Files
F173756652: D60083.id.diff
Mon, Sep 28, 4:22 AM
F173738068: D60083.diff
Mon, Sep 28, 1:22 AM
F173737618: D60083.id187847.diff
Mon, Sep 28, 1:18 AM
F173735218: D60083.diff
Mon, Sep 28, 12:58 AM
F173733742: D60083.id187847.diff
Mon, Sep 28, 12:45 AM
F173729115: D60083.diff
Mon, Sep 28, 12:06 AM
F173726929: D60083.diff
Sun, Sep 27, 11:47 PM
F173726449: D60083.id187847.diff
Sun, Sep 27, 11:43 PM

Details

Reviewers
donner
Group Reviewers
network
Summary

Inbound lookups find the (alias address, alias port, link type) group
with a splay tree and then walk grp->full, a list of every fully
specified link in that group, comparing the remote address and port.
With redirect_addr in front of a busy server, every client connection
to public:443 lands in the same group, so each inbound packet that is
not near the head of the list walks all of it. TCP links live up to
24 hours unless libalias sees a clean close, so the list can grow to
hundreds of thousands of entries and saturate a core at a few hundred
packets per second.

Keep the fully specified links of a group in an RB tree ordered by
(dst_addr, dst_port). Links that share an endpoint are ordered newest
first by a per-instance insertion counter, which keeps the "most recent
link wins" behaviour of the list (tested by 3_natin:2_portoverlap).
The counter is appended at the end of struct libalias because ipfw_nat
and ng_nat access fields of that structure directly.

Lookups with an unknown remote address and a known port still scan the
group. Each link grows by 16 bytes.

With ~170,000 links in one group, LibAliasIn() took 1-2 ms per packet
and one core saturated at ~720 packets/s. With this change it takes
2-4 us per packet and the offered 1,000 packets/s are handled with the
thread nearly idle.

Sponsored by: NLINK

Test Plan

Lab reproduction on 16.0-CURRENT (main-n289577), amd64, three VNET jails:

natlab_cli (198.51.100.10-.14) --epair-- natlab_rtr --epair-- natlab_srv (10.0.0.2)

natlab_rtr runs ipfw nat with the configuration that exposed the problem
in production:

ipfw nat 1 config ip 198.51.100.1 log same_ports unreg_only redirect_addr 10.0.0.2 198.51.100.50

The client sends 1,000 SYN/s to 198.51.100.50:443 for 300 s (non-blocking
connect() then close(), so one SYN per connection). natlab_srv drops the
SYNs, so every link stays for TCP_EXPIRE_INITIAL. Link counts come from
ipfw nat show log; latency from DTrace (fbt::LibAliasIn and
fbt::FindUdpTcpIn entry/return, quantize) plus profile-997 per thread.

Experiment A: all SYNs to one port (one group).
Experiment B (control): same rate spread over ports 1000-60000.

Experiment Abefore (~170k links)after (~141k links)
LibAliasIn per packet1-2 ms2-4 us
packets handled per 10 s~7,200 (saturated)~10,000 (all offered)
epair_task CPU99%0.0-0.15%
_FindLinkIn profile samples9,911 of ~9,9702

Before the change, experiment B (~280k links spread over many groups)
already ran at 2-4 us per packet, which shows the cost came from the
length of one group's list, not from the total number of links. After
the change, the cost of A did not grow between 60k and 141k links.

Userland (tests/sys/netinet/libalias, built on the same sources):

  • 1_instance, 2_natout (including the UDP EIM cases), 3_natin: all pass. 3_natin:2_portoverlap caught an early version that broke ties between links with the same endpoint by pointer; ties are now ordered newest first, as with the list.
  • perf: no regression (existing links ~0.14 us per packet before and after; random inbound 3.2 us -> 1.7 us).
  • Additional one-group benchmark (100k links on one port): new SYN 1,252,985 ns -> ~600 ns; existing link 606,210 ns -> ~1,100 ns. I can submit it separately as tests/sys/netinet/libalias/perf_onegroup.c.

Kernel module checks:

  • Only libalias.ko was rebuilt; ipfw nat show log works with the stock ipfw_nat.ko, which confirms struct libalias offsets are unchanged (an earlier version with linkSeq in the middle broke it).
  • FullInFind()/FullInNFind() and the full_in RB functions are fully inlined at -O2 (no fbt probes, no symbols).

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped
Build Status
Buildable 77398
Build 74281: arc lint + arc unit

Event Timeline

oh, nice catch! remind me again why they're all ending up in one group here?

Thanks! It comes from how inbound links are keyed.

In StartPointIn() all inbound traffic is indexed by
(alias_addr, alias_port, link_type), and the links inside the group
were a plain list, walked comparing (dst_addr, dst_port).

For connections from outside towards a redirect, the alias port is
the port the client connected to. When no link exists yet,
FindUdpTcpIn() does:

target_addr = FindOriginalAddress(la, alias_addr);
lnk = AddLink(la, target_addr, dst_addr, alias_addr,
    alias_port, dst_port, alias_port, link_type);

The case where I hit this was a 1:1 redirect in front of a busy web
server, a common setup in a hosting datacenter:

browser  <->  ipfw nat (redirect_addr)  <->  web server (443)
client        public address                 private address

Every browser connection to <public>:443 creates one link with alias
(<public>, 443, TCP), so all clients of the server land in the same
group. Only the client address and port differ, which is exactly what
the list was scanned for.

On a heavily used HTTPS service that list gets long: browsers open
several connections each, and many clients (mobile, CGNAT, scanners)
disappear without a FIN or RST that libalias can see. Those links stay
for TCP_EXPIRE_CONNECTED (24 h) or TCP_EXPIRE_INITIAL (300 s). Every
new SYN, and every packet whose link is not near the head, walks the
whole list. In production this pinned one core in the NAT's network
thread at only a few hundred packets per second, which is what the lab
in the test plan reproduces.

Outbound-initiated flows don't have this problem, because GetNewPort()
spreads them over many alias ports (many small groups). A published
server is the opposite: one alias port for all of its clients.