Diff Detail
- Lint
Lint Skipped - Unit
Tests Skipped
Event Timeline
@andrew, Is this what you were asking for? A 16 processor AWS a1.4xlarge, which is Cortex A72-based, sees a small reduction in system time using arm64_pipt_icache_sync_range(). Specifically, the reduction is about 0.75%.
Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.
| sys/arm64/arm64/cpufunc_asm.S | ||
|---|---|---|
| 162 | This limits the number of back-to-back invalidation broadcasts on a machine with a 64-byte cache line size to 512. That is the same limit that Linux places on back-to-back TLB invalidation broadcasts. | |
I think this is what you're looking for?
CPU 0: ARM Neoverse-N1 r3p1 affinity: 18 0 0
Cache Type = <IDC,64 byte CWG,64 byte ERG,64 byte D-cacheline,PIPT I-cache,64 byte I-cacheline>
Instruction Set Attributes 0 = <DP,RDM,Atomic,CRC32,SHA2,SHA1,AES+PMULL>
Instruction Set Attributes 1 = <RCPC-8.3,DCPoP>
Instruction Set Attributes 2 = <>
Processor Features 0 = <CSV3,CSV2,RAS,GIC,AdvSIMD+HP,FP+HP,EL3,EL2,EL1,EL0 32>
Processor Features 1 = <MTE_frac,PSTATE.SSBS MSR>
Processor Features 2 = <>
Trying to mount root from zfs:zroot/ROOT/bhyve []...
Memory Model Features 0 = <TGran4,TGran64,TGran16,SNSMem,BigEnd,16bit ASID,256TB PA>
Memory Model Features 1 = <XNX,PAN+ATS1E1,LO,HPD+TTPBHA,VH,16bit VMID,HAF+DS>
Memory Model Features 2 = <EVT-8.2,32bit CCIDX,48bit VA,UAO,CnP>
Memory Model Features 3 = <>
Memory Model Features 4 = <>
Debug Features 0 = <DoubleLock,SPE,2 CTX BKPTs,4 Watchpoints,6 Breakpoints,PMUv3p1,Debugv8p2>
Debug Features 1 = <>
Auxiliary Features 0 = <>
Auxiliary Features 1 = <>It's an Ampere Altra, not sure offhand which one. I can test this patch on it this weekend if you tell me what exactly you'd like to try.
Yes. I would test GENERIC-NODEBUG kernels without and with this patch. I have /etc/src.conf:
WITHOUT_LLVM_ASSERTIONS=yes # WITH_CLEAN=yes WITH_MALLOC_PRODUCTION=yes
I run:
#!/bin/csh cd /usr/src rm -fr /usr/obj/usr/src while ( 1 ) date sysctl vm.pmap vm.reserv vm.stats.vm.v_vm_faults time make -j16 buildworld > & /dev/null rm -fr /usr/obj/usr/src end
with -j adjusted for the machine and the script's output redirected to a log.
Here is Claude's summary of a log that I collected on the DevKit:
Icache synchronizations by mapping size during buildworld (arm64)
=================================================================
Source
8 Cortex-X1C/Cortex-A78C cores, three consecutive buildworld runs
Run 1 Run 2 Run 3
elapsed 2:07:05 2:06:38 2:06:29
user (s) 56,794.2 56,785.5 56,824.6
sys (s) 2,459.8 2,473.5 2,478.2
What the counters count
Synchronizations actually performed, i.e., after PGA_ICACHE_SYNCED
has let the pmap skip the ones it could.
vm.pmap.l3.icache_syncs pmap_enter() and pmap_enter_quick_locked()
vm.pmap.l3c.icache_syncs pmap_enter_l3c()
vm.pmap.l2.icache_syncs pmap_enter_l2()
Syncs per run (differences between successive counter snapshots)
Size Run 1 Run 2 Run 3 Mean Share Bytes/run
L3 (4 KB) 8,347,880 8,351,391 8,341,250 8,346,840 97.8% 31.8 GiB
L3C (64 KB) 173,976 193,834 196,820 188,210 2.2% 11.5 GiB
L2 (2 MB) 0 0 0 0 0.0% 0
From boot until the first run: 23,680 L3, 18 L3C, 0 L2.
Other pmap activity per run (means)
L3C mappings created 55.0 M L3C promotions 21.9 M
L2 mappings created 176 K L2 promotions 324 K
L3C demotions 4.23 M L2 demotions 112 K
Observations
- L3 syncs dominate the count and are very stable: the three runs are
within about 0.12% of one another. L3C varies more; run 1 was about
11% below the mean of runs 2 and 3.
- L3C syncs are only 2.2% of the calls, but each covers 64 KB, so they
are 26.5% of the bytes synchronized. With this patch's 32 KB cap, every
L3C sync takes the invalidate-all path on parts that need it, while
every L3 sync takes the ranged path.
- Only about 0.34% of L3C mappings required a sync.
- No L2 mapping required a sync, despite about 176 K L2 mappings and
324 K L2 promotions per run. Promotion never synchronizes, and this
workload never directly created a 2 MB mapping that was executable
and still needed a sync.