Page MenuHomeFreeBSD

arm64: avoid full icache flushes on PIPT icaches
Needs ReviewPublic

Authored by alc on Fri, Oct 2, 10:22 PM.
Tags
None
Referenced Files
F174538910: D60271.diff
Sun, Oct 4, 12:55 AM
Unknown Object (File)
Sat, Oct 3, 1:38 AM
Unknown Object (File)
Sat, Oct 3, 1:06 AM
Subscribers

Details

Diff Detail

Lint
Lint Skipped
Unit
Tests Skipped

Event Timeline

alc requested review of this revision.Fri, Oct 2, 10:22 PM

@andrew, Is this what you were asking for? A 16 processor AWS a1.4xlarge, which is Cortex A72-based, sees a small reduction in system time using arm64_pipt_icache_sync_range(). Specifically, the reduction is about 0.75%.

Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.

sys/arm64/arm64/cpufunc_asm.S
162

This limits the number of back-to-back invalidation broadcasts on a machine with a 64-byte cache line size to 512. That is the same limit that Linux places on back-to-back TLB invalidation broadcasts.

In D60271#1383042, @alc wrote:

Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.

I think this is what you're looking for?

CPU  0: ARM Neoverse-N1 r3p1 affinity: 18  0  0
                   Cache Type = <IDC,64 byte CWG,64 byte ERG,64 byte D-cacheline,PIPT I-cache,64 byte I-cacheline>
 Instruction Set Attributes 0 = <DP,RDM,Atomic,CRC32,SHA2,SHA1,AES+PMULL>
 Instruction Set Attributes 1 = <RCPC-8.3,DCPoP>
 Instruction Set Attributes 2 = <>
         Processor Features 0 = <CSV3,CSV2,RAS,GIC,AdvSIMD+HP,FP+HP,EL3,EL2,EL1,EL0 32>
         Processor Features 1 = <MTE_frac,PSTATE.SSBS MSR>
         Processor Features 2 = <>
Trying to mount root from zfs:zroot/ROOT/bhyve []...
      Memory Model Features 0 = <TGran4,TGran64,TGran16,SNSMem,BigEnd,16bit ASID,256TB PA>
      Memory Model Features 1 = <XNX,PAN+ATS1E1,LO,HPD+TTPBHA,VH,16bit VMID,HAF+DS>
      Memory Model Features 2 = <EVT-8.2,32bit CCIDX,48bit VA,UAO,CnP>
      Memory Model Features 3 = <>
      Memory Model Features 4 = <>
             Debug Features 0 = <DoubleLock,SPE,2 CTX BKPTs,4 Watchpoints,6 Breakpoints,PMUv3p1,Debugv8p2>
             Debug Features 1 = <>
         Auxiliary Features 0 = <>
         Auxiliary Features 1 = <>

It's an Ampere Altra, not sure offhand which one. I can test this patch on it this weekend if you tell me what exactly you'd like to try.

In D60271#1383042, @alc wrote:

Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.

I think this is what you're looking for?

CPU  0: ARM Neoverse-N1 r3p1 affinity: 18  0  0
                   Cache Type = <IDC,64 byte CWG,64 byte ERG,64 byte D-cacheline,PIPT I-cache,64 byte I-cacheline>
 Instruction Set Attributes 0 = <DP,RDM,Atomic,CRC32,SHA2,SHA1,AES+PMULL>
 Instruction Set Attributes 1 = <RCPC-8.3,DCPoP>
 Instruction Set Attributes 2 = <>
         Processor Features 0 = <CSV3,CSV2,RAS,GIC,AdvSIMD+HP,FP+HP,EL3,EL2,EL1,EL0 32>
         Processor Features 1 = <MTE_frac,PSTATE.SSBS MSR>
         Processor Features 2 = <>
Trying to mount root from zfs:zroot/ROOT/bhyve []...
      Memory Model Features 0 = <TGran4,TGran64,TGran16,SNSMem,BigEnd,16bit ASID,256TB PA>
      Memory Model Features 1 = <XNX,PAN+ATS1E1,LO,HPD+TTPBHA,VH,16bit VMID,HAF+DS>
      Memory Model Features 2 = <EVT-8.2,32bit CCIDX,48bit VA,UAO,CnP>
      Memory Model Features 3 = <>
      Memory Model Features 4 = <>
             Debug Features 0 = <DoubleLock,SPE,2 CTX BKPTs,4 Watchpoints,6 Breakpoints,PMUv3p1,Debugv8p2>
             Debug Features 1 = <>
         Auxiliary Features 0 = <>
         Auxiliary Features 1 = <>

It's an Ampere Altra, not sure offhand which one. I can test this patch on it this weekend if you tell me what exactly you'd like to try.

Yes. I would test GENERIC-NODEBUG kernels without and with this patch. I have /etc/src.conf:

WITHOUT_LLVM_ASSERTIONS=yes
#
WITH_CLEAN=yes
WITH_MALLOC_PRODUCTION=yes

I run:

#!/bin/csh
cd /usr/src
rm -fr /usr/obj/usr/src
while ( 1 )
    date
    sysctl vm.pmap vm.reserv vm.stats.vm.v_vm_faults
    time make -j16 buildworld > & /dev/null 
    rm -fr /usr/obj/usr/src
end

with -j adjusted for the machine and the script's output redirected to a log.

Here is Claude's summary of a log that I collected on the DevKit:

Icache synchronizations by mapping size during buildworld (arm64)
=================================================================

Source
  8 Cortex-X1C/Cortex-A78C cores, three consecutive buildworld runs

                 Run 1       Run 2       Run 3
    elapsed      2:07:05     2:06:38     2:06:29
    user (s)     56,794.2    56,785.5    56,824.6
    sys (s)       2,459.8     2,473.5     2,478.2

What the counters count
  Synchronizations actually performed, i.e., after PGA_ICACHE_SYNCED
  has let the pmap skip the ones it could.

    vm.pmap.l3.icache_syncs    pmap_enter() and pmap_enter_quick_locked()
    vm.pmap.l3c.icache_syncs   pmap_enter_l3c()
    vm.pmap.l2.icache_syncs    pmap_enter_l2()

Syncs per run (differences between successive counter snapshots)

    Size           Run 1      Run 2      Run 3       Mean  Share  Bytes/run
    L3  (4 KB)  8,347,880  8,351,391  8,341,250  8,346,840  97.8%   31.8 GiB
    L3C (64 KB)   173,976    193,834    196,820    188,210   2.2%   11.5 GiB
    L2  (2 MB)          0          0          0          0   0.0%          0

    From boot until the first run: 23,680 L3, 18 L3C, 0 L2.

Other pmap activity per run (means)

    L3C mappings created   55.0 M      L3C promotions   21.9 M
    L2 mappings created    176 K       L2 promotions     324 K
    L3C demotions          4.23 M      L2 demotions      112 K

Observations
  - L3 syncs dominate the count and are very stable: the three runs are
    within about 0.12% of one another.  L3C varies more; run 1 was about
    11% below the mean of runs 2 and 3.
  - L3C syncs are only 2.2% of the calls, but each covers 64 KB, so they
    are 26.5% of the bytes synchronized.  With this patch's 32 KB cap, every
    L3C sync takes the invalidate-all path on parts that need it, while
    every L3 sync takes the ranged path.
  - Only about 0.34% of L3C mappings required a sync.
  - No L2 mapping required a sync, despite about 176 K L2 mappings and
    324 K L2 promotions per run.  Promotion never synchronizes, and this
    workload never directly created a 2 MB mapping that was executable
    and still needed a sync.