Diff Detail
- Lint
Lint Skipped - Unit
Tests Skipped
Event Timeline
@andrew, Is this what you were asking for? A 16 processor AWS a1.4xlarge, which is Cortex A72-based, sees a small reduction in system time using arm64_pipt_icache_sync_range(). Specifically, the reduction is about 0.75%.
Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.
| sys/arm64/arm64/cpufunc_asm.S | ||
|---|---|---|
| 162 | This limits the number of back-to-back invalidation broadcasts on a machine with a 64-byte cache line size to 512. That is the same limit that Linux places on back-to-back TLB invalidation broadcasts. | |
I think this is what you're looking for?
CPU 0: ARM Neoverse-N1 r3p1 affinity: 18 0 0
Cache Type = <IDC,64 byte CWG,64 byte ERG,64 byte D-cacheline,PIPT I-cache,64 byte I-cacheline>
Instruction Set Attributes 0 = <DP,RDM,Atomic,CRC32,SHA2,SHA1,AES+PMULL>
Instruction Set Attributes 1 = <RCPC-8.3,DCPoP>
Instruction Set Attributes 2 = <>
Processor Features 0 = <CSV3,CSV2,RAS,GIC,AdvSIMD+HP,FP+HP,EL3,EL2,EL1,EL0 32>
Processor Features 1 = <MTE_frac,PSTATE.SSBS MSR>
Processor Features 2 = <>
Trying to mount root from zfs:zroot/ROOT/bhyve []...
Memory Model Features 0 = <TGran4,TGran64,TGran16,SNSMem,BigEnd,16bit ASID,256TB PA>
Memory Model Features 1 = <XNX,PAN+ATS1E1,LO,HPD+TTPBHA,VH,16bit VMID,HAF+DS>
Memory Model Features 2 = <EVT-8.2,32bit CCIDX,48bit VA,UAO,CnP>
Memory Model Features 3 = <>
Memory Model Features 4 = <>
Debug Features 0 = <DoubleLock,SPE,2 CTX BKPTs,4 Watchpoints,6 Breakpoints,PMUv3p1,Debugv8p2>
Debug Features 1 = <>
Auxiliary Features 0 = <>
Auxiliary Features 1 = <>It's an Ampere Altra, not sure offhand which one. I can test this patch on it this weekend if you tell me what exactly you'd like to try.
Yes. I would test GENERIC-NODEBUG kernels without and with this patch. I have /etc/src.conf:
WITHOUT_LLVM_ASSERTIONS=yes # WITH_CLEAN=yes WITH_MALLOC_PRODUCTION=yes
I run:
#!/bin/csh cd /usr/src rm -fr /usr/obj/usr/src while ( 1 ) date sysctl vm.pmap vm.reserv vm.stats.vm.v_vm_faults time make -j16 buildworld > & /dev/null rm -fr /usr/obj/usr/src end
with -j adjusted for the machine and the script's output redirected to a log.
Here is Claude's summary of a log that I collected on the DevKit:
Icache synchronizations by mapping size during buildworld (arm64)
=================================================================
Source
8 Cortex-X1C/Cortex-A78C cores, three consecutive buildworld runs
Run 1 Run 2 Run 3
elapsed 2:07:05 2:06:38 2:06:29
user (s) 56,794.2 56,785.5 56,824.6
sys (s) 2,459.8 2,473.5 2,478.2
What the counters count
Synchronizations actually performed, i.e., after PGA_ICACHE_SYNCED
has let the pmap skip the ones it could.
vm.pmap.l3.icache_syncs pmap_enter() and pmap_enter_quick_locked()
vm.pmap.l3c.icache_syncs pmap_enter_l3c()
vm.pmap.l2.icache_syncs pmap_enter_l2()
Syncs per run (differences between successive counter snapshots)
Size Run 1 Run 2 Run 3 Mean Share Bytes/run
L3 (4 KB) 8,347,880 8,351,391 8,341,250 8,346,840 97.8% 31.8 GiB
L3C (64 KB) 173,976 193,834 196,820 188,210 2.2% 11.5 GiB
L2 (2 MB) 0 0 0 0 0.0% 0
From boot until the first run: 23,680 L3, 18 L3C, 0 L2.
Other pmap activity per run (means)
L3C mappings created 55.0 M L3C promotions 21.9 M
L2 mappings created 176 K L2 promotions 324 K
L3C demotions 4.23 M L2 demotions 112 K
Observations
- L3 syncs dominate the count and are very stable: the three runs are
within about 0.12% of one another. L3C varies more; run 1 was about
11% below the mean of runs 2 and 3.
- L3C syncs are only 2.2% of the calls, but each covers 64 KB, so they
are 26.5% of the bytes synchronized. With this patch's 32 KB cap, every
L3C sync takes the invalidate-all path on parts that need it, while
every L3 sync takes the ranged path.
- Only about 0.34% of L3C mappings required a sync.
- No L2 mapping required a sync, despite about 176 K L2 mappings and
324 K L2 promotions per run. Promotion never synchronizes, and this
workload never directly created a 2 MB mapping that was executable
and still needed a sync.Yes
A 16 processor AWS a1.4xlarge, which is Cortex A72-based, sees a small reduction in system time using arm64_pipt_icache_sync_range(). Specifically, the reduction is about 0.75%.
My previous attempts at implementing this have had the opposite effect.
I have an Orion-O6 with IDC & no DIC + PIPT cache, but can't test for a day or two when I'm back in the office.
Unpatched output: https://reviews.freebsd.org/P717
Patched output: https://reviews.freebsd.org/P718
Timings from an unpatched kernel:
936.90 real 45789.09 user 3280.64 sys 933.97 real 45789.59 user 3314.96 sys 935.05 real 45774.25 user 3334.41 sys 934.40 real 45794.90 user 3302.60 sys 932.76 real 45846.53 user 3315.03 sys 933.13 real 45788.78 user 3318.50 sys
and a patched kernel:
938.43 real 45847.32 user 3243.15 sys 933.22 real 45670.20 user 3268.07 sys 933.84 real 45820.26 user 3269.70 sys 933.28 real 45730.27 user 3252.91 sys 939.21 real 45792.15 user 3287.05 sys 930.78 real 45835.46 user 3260.92 sys
which indicates that the change slightly reduced system CPU time.
My previous change so reduced the number of icache syncs that there just isn't as much to gain from optimizing the sync. That said, I'd still be interested in seeing results from the Orion-O6.
One open question is whether my 32KB threshold for switching to ialluis is too low. In other words, should we do range-based syncs on 64KB page mappings? I think that is something we should look into post commit.
Thanks. This system is ZFS-based, yes? Interestingly, you're getting 25% more L3C mappings than I see on UFS-based systems, presumably from loading executable files.