On Zen 6, an extra IBS NMI can arrive after later samples.
Keep the credit until the empty NMI arrives, and handle fetch and op samples when both are ready.
Details
Ran the workload 34 times on Zen 6; no unknown NMIs.
Diff Detail
- Repository
- rG FreeBSD src repository
- Lint
Lint Passed - Unit
No Test Coverage - Build Status
Buildable 77730 Build 74613: arc lint + arc unit
Event Timeline
Thanks for the note! I rebased this on top of D60004 and uploaded Diff. The updated diff is focused on the Zen 6 delayed NMI handling.
LGTM, with one question.
| sys/dev/hwpmc/hwpmc_ibs.c | ||
|---|---|---|
| 738–741 | Does it become possible that one of these intermediate NMIs could arrive with both events, thus needing an additional credit? Perhaps this would be enough to handle that case. | |
Does it become possible that one of these intermediate NMIs could arrive with both events, thus needing an additional credit?
Yes, I think you are right. If fetch and op both fire again before the delayed NMI arrives, two empty NMIs are owed, but the flag only remembers one.
Making it a counter works for me, and I'll test it on Zen 6. Maybe we could cap it with something like IBS_NMI_CREDIT_MAX, since every credit left over would hide a genuine unknown NMI later:
if (retval == 2) { if (pac->pc_nmi_credit < IBS_NMI_CREDIT_MAX) pac->pc_nmi_credit++; } else if (retval == 0 && pac->pc_nmi_credit > 0) { pac->pc_nmi_credit--; retval = 1; }
I've updated the diff with this change.
Tested the counter version on Zen 6 (1024 CPUs), hwpmc.ko only, with machdep.panic_on_nmi=255, running pmcstat with scimark4: 20 runs with -S ibs-fetch -S ibs-op, 4 with the order reversed, 5 op only and 5 fetch only. All 34 runs had 0 unknown NMIs and kern.hwpmc.stats.intr_ignored stayed at 0. For comparison, the stock module gives about 180 unclaimed NMIs per 8 runs with both events.