Page MenuHomeFreeBSD

callout: enforce a minimum retry delay for CALLOUT_TRYLOCK
AcceptedPublic

Authored by rcm on Fri, Oct 2, 12:52 PM.
Tags
None
Referenced Files
F174445463: D60246.id188390.diff
Sat, Oct 3, 6:58 AM
F174422600: D60246.id.diff
Sat, Oct 3, 2:17 AM
F174422293: D60246.diff
Sat, Oct 3, 2:13 AM
F174420450: D60246.diff
Sat, Oct 3, 1:51 AM
F174366319: D60246.diff
Fri, Oct 2, 5:33 PM
F174353128: D60246.diff
Fri, Oct 2, 3:29 PM
Subscribers

Details

Reviewers
glebius
kib
markj
Summary

When softclock_call_cc() fails to acquire the lock of a CALLOUT_TRYLOCK
callout, it reschedules the callout half its precision after
cc_lastscan and halves the precision. Repeated failures shrink the
delay toward zero, and once the precision reaches 1 the callout is due
immediately: the timer fires again at once and softclock retries the
lock in a tight loop for as long as the lock is held.

If the lock owner runs on the callout's CPU and no other CPU is idle,
the softclock thread preempts it on every attempt, starving the thread
it is waiting on. On an 8-CPU arm64 VM, a test module holding the lock
saw 760,000 attempts per second, each with its own timer interrupt, and
progressed at 38% of its normal rate. On a 4-core amd64 system under
loopback TCP load, a netisr thread holding an inpcb lock made no
progress for 12 minutes while the TCP timer callout was retried 830,000
times per second.

Keep the half-precision retry, but never schedule it less than one
tick from now. Base the deadline on sbinuptime() rather than
cc_lastscan: cc_lastscan is the time of the last callout_process()
scan, which can be a tick or more in the past by the time softclock
runs the callout, leaving cc_lastscan + tick_sbt already expired.

Floor the precision at a tick as well. With a precision of 1 the
retries of many contended callouts cannot share a timer interrupt: the
timer is armed for the earliest retry, each interrupt collects only
those already due, and every callout_process() call walks all the
contended callouts in the callwheel bucket. On the same VM, with
10,000 callouts on one lock, the owner took 5.0 to 5.8 s to finish 1 s
of work with a precision of 1 and 1.3 s with the floor, and
callout_process() ran about 58,000 times against about 900. A single
contended callout is now retried about every two ticks.

Fixes: efcb2ec8cb81 ("callout: provide CALLOUT_TRYLOCK flag")
MFC after: 2 weeks
Sponsored by: Rubicon Communications, LLC ("Netgate")

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped

Event Timeline

rcm requested review of this revision.Fri, Oct 2, 12:52 PM
kib added inline comments.
sys/kern/kern_timeout.c
697

I suggest to create a local variable for this expression.

This revision is now accepted and ready to land.Fri, Oct 2, 9:18 PM