Page MenuHomeFreeBSD

arm64: avoid full icache flushes on PIPT icaches
AcceptedPublic

Authored by alc on Fri, Oct 2, 10:22 PM.
Tags
None
Referenced Files
F174890990: D60271.diff
Tue, Oct 6, 7:22 PM
F174853998: D60271.diff
Tue, Oct 6, 1:33 PM
Unknown Object (File)
Tue, Oct 6, 3:54 AM
Unknown Object (File)
Tue, Oct 6, 1:34 AM
Unknown Object (File)
Tue, Oct 6, 12:50 AM
Unknown Object (File)
Mon, Oct 5, 9:06 PM
Unknown Object (File)
Sun, Oct 4, 4:26 PM
Unknown Object (File)
Sun, Oct 4, 4:26 PM
Subscribers

Details

Diff Detail

Lint
Lint Skipped
Unit
Tests Skipped

Event Timeline

alc requested review of this revision.Fri, Oct 2, 10:22 PM

@andrew, Is this what you were asking for? A 16 processor AWS a1.4xlarge, which is Cortex A72-based, sees a small reduction in system time using arm64_pipt_icache_sync_range(). Specifically, the reduction is about 0.75%.

Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.

sys/arm64/arm64/cpufunc_asm.S
162

This limits the number of back-to-back invalidation broadcasts on a machine with a 64-byte cache line size to 512. That is the same limit that Linux places on back-to-back TLB invalidation broadcasts.

In D60271#1383042, @alc wrote:

Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.

I think this is what you're looking for?

CPU  0: ARM Neoverse-N1 r3p1 affinity: 18  0  0
                   Cache Type = <IDC,64 byte CWG,64 byte ERG,64 byte D-cacheline,PIPT I-cache,64 byte I-cacheline>
 Instruction Set Attributes 0 = <DP,RDM,Atomic,CRC32,SHA2,SHA1,AES+PMULL>
 Instruction Set Attributes 1 = <RCPC-8.3,DCPoP>
 Instruction Set Attributes 2 = <>
         Processor Features 0 = <CSV3,CSV2,RAS,GIC,AdvSIMD+HP,FP+HP,EL3,EL2,EL1,EL0 32>
         Processor Features 1 = <MTE_frac,PSTATE.SSBS MSR>
         Processor Features 2 = <>
Trying to mount root from zfs:zroot/ROOT/bhyve []...
      Memory Model Features 0 = <TGran4,TGran64,TGran16,SNSMem,BigEnd,16bit ASID,256TB PA>
      Memory Model Features 1 = <XNX,PAN+ATS1E1,LO,HPD+TTPBHA,VH,16bit VMID,HAF+DS>
      Memory Model Features 2 = <EVT-8.2,32bit CCIDX,48bit VA,UAO,CnP>
      Memory Model Features 3 = <>
      Memory Model Features 4 = <>
             Debug Features 0 = <DoubleLock,SPE,2 CTX BKPTs,4 Watchpoints,6 Breakpoints,PMUv3p1,Debugv8p2>
             Debug Features 1 = <>
         Auxiliary Features 0 = <>
         Auxiliary Features 1 = <>

It's an Ampere Altra, not sure offhand which one. I can test this patch on it this weekend if you tell me what exactly you'd like to try.

In D60271#1383042, @alc wrote:

Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.

I think this is what you're looking for?

CPU  0: ARM Neoverse-N1 r3p1 affinity: 18  0  0
                   Cache Type = <IDC,64 byte CWG,64 byte ERG,64 byte D-cacheline,PIPT I-cache,64 byte I-cacheline>
 Instruction Set Attributes 0 = <DP,RDM,Atomic,CRC32,SHA2,SHA1,AES+PMULL>
 Instruction Set Attributes 1 = <RCPC-8.3,DCPoP>
 Instruction Set Attributes 2 = <>
         Processor Features 0 = <CSV3,CSV2,RAS,GIC,AdvSIMD+HP,FP+HP,EL3,EL2,EL1,EL0 32>
         Processor Features 1 = <MTE_frac,PSTATE.SSBS MSR>
         Processor Features 2 = <>
Trying to mount root from zfs:zroot/ROOT/bhyve []...
      Memory Model Features 0 = <TGran4,TGran64,TGran16,SNSMem,BigEnd,16bit ASID,256TB PA>
      Memory Model Features 1 = <XNX,PAN+ATS1E1,LO,HPD+TTPBHA,VH,16bit VMID,HAF+DS>
      Memory Model Features 2 = <EVT-8.2,32bit CCIDX,48bit VA,UAO,CnP>
      Memory Model Features 3 = <>
      Memory Model Features 4 = <>
             Debug Features 0 = <DoubleLock,SPE,2 CTX BKPTs,4 Watchpoints,6 Breakpoints,PMUv3p1,Debugv8p2>
             Debug Features 1 = <>
         Auxiliary Features 0 = <>
         Auxiliary Features 1 = <>

It's an Ampere Altra, not sure offhand which one. I can test this patch on it this weekend if you tell me what exactly you'd like to try.

Yes. I would test GENERIC-NODEBUG kernels without and with this patch. I have /etc/src.conf:

WITHOUT_LLVM_ASSERTIONS=yes
#
WITH_CLEAN=yes
WITH_MALLOC_PRODUCTION=yes

I run:

#!/bin/csh
cd /usr/src
rm -fr /usr/obj/usr/src
while ( 1 )
    date
    sysctl vm.pmap vm.reserv vm.stats.vm.v_vm_faults
    time make -j16 buildworld > & /dev/null 
    rm -fr /usr/obj/usr/src
end

with -j adjusted for the machine and the script's output redirected to a log.

Here is Claude's summary of a log that I collected on the DevKit:

Icache synchronizations by mapping size during buildworld (arm64)
=================================================================

Source
  8 Cortex-X1C/Cortex-A78C cores, three consecutive buildworld runs

                 Run 1       Run 2       Run 3
    elapsed      2:07:05     2:06:38     2:06:29
    user (s)     56,794.2    56,785.5    56,824.6
    sys (s)       2,459.8     2,473.5     2,478.2

What the counters count
  Synchronizations actually performed, i.e., after PGA_ICACHE_SYNCED
  has let the pmap skip the ones it could.

    vm.pmap.l3.icache_syncs    pmap_enter() and pmap_enter_quick_locked()
    vm.pmap.l3c.icache_syncs   pmap_enter_l3c()
    vm.pmap.l2.icache_syncs    pmap_enter_l2()

Syncs per run (differences between successive counter snapshots)

    Size           Run 1      Run 2      Run 3       Mean  Share  Bytes/run
    L3  (4 KB)  8,347,880  8,351,391  8,341,250  8,346,840  97.8%   31.8 GiB
    L3C (64 KB)   173,976    193,834    196,820    188,210   2.2%   11.5 GiB
    L2  (2 MB)          0          0          0          0   0.0%          0

    From boot until the first run: 23,680 L3, 18 L3C, 0 L2.

Other pmap activity per run (means)

    L3C mappings created   55.0 M      L3C promotions   21.9 M
    L2 mappings created    176 K       L2 promotions     324 K
    L3C demotions          4.23 M      L2 demotions      112 K

Observations
  - L3 syncs dominate the count and are very stable: the three runs are
    within about 0.12% of one another.  L3C varies more; run 1 was about
    11% below the mean of runs 2 and 3.
  - L3C syncs are only 2.2% of the calls, but each covers 64 KB, so they
    are 26.5% of the bytes synchronized.  With this patch's 32 KB cap, every
    L3C sync takes the invalidate-all path on parts that need it, while
    every L3 sync takes the ranged path.
  - Only about 0.34% of L3C mappings required a sync.
  - No L2 mapping required a sync, despite about 176 K L2 mappings and
    324 K L2 promotions per run.  Promotion never synchronizes, and this
    workload never directly created a 2 MB mapping that was executable
    and still needed a sync.

I can test this patch on it this weekend if you tell me what exactly you'd like to try.

I had to restart this this morning, I should have some results later today.

In D60271#1383040, @alc wrote:

@andrew, Is this what you were asking for?

Yes

A 16 processor AWS a1.4xlarge, which is Cortex A72-based, sees a small reduction in system time using arm64_pipt_icache_sync_range(). Specifically, the reduction is about 0.75%.

My previous attempts at implementing this have had the opposite effect.

I have an Orion-O6 with IDC & no DIC + PIPT cache, but can't test for a day or two when I'm back in the office.

In D60271#1383119, @alc wrote:

Yes. I would test GENERIC-NODEBUG kernels without and with this patch. I have /etc/src.conf:

WITHOUT_LLVM_ASSERTIONS=yes
#
WITH_CLEAN=yes
WITH_MALLOC_PRODUCTION=yes

I run:

#!/bin/csh
cd /usr/src
rm -fr /usr/obj/usr/src
while ( 1 )
    date
    sysctl vm.pmap vm.reserv vm.stats.vm.v_vm_faults
    time make -j16 buildworld > & /dev/null 
    rm -fr /usr/obj/usr/src
end

with -j adjusted for the machine and the script's output redirected to a log.

Unpatched output: https://reviews.freebsd.org/P717
Patched output: https://reviews.freebsd.org/P718

Timings from an unpatched kernel:

936.90 real     45789.09 user      3280.64 sys
933.97 real     45789.59 user      3314.96 sys
935.05 real     45774.25 user      3334.41 sys
934.40 real     45794.90 user      3302.60 sys
932.76 real     45846.53 user      3315.03 sys
933.13 real     45788.78 user      3318.50 sys

and a patched kernel:

938.43 real     45847.32 user      3243.15 sys
933.22 real     45670.20 user      3268.07 sys
933.84 real     45820.26 user      3269.70 sys
933.28 real     45730.27 user      3252.91 sys
939.21 real     45792.15 user      3287.05 sys
930.78 real     45835.46 user      3260.92 sys

which indicates that the change slightly reduced system CPU time.

This revision is now accepted and ready to land.Tue, Oct 6, 5:08 PM
In D60271#1383040, @alc wrote:

@andrew, Is this what you were asking for?

Yes

A 16 processor AWS a1.4xlarge, which is Cortex A72-based, sees a small reduction in system time using arm64_pipt_icache_sync_range(). Specifically, the reduction is about 0.75%.

My previous attempts at implementing this have had the opposite effect.

I have an Orion-O6 with IDC & no DIC + PIPT cache, but can't test for a day or two when I'm back in the office.

My previous change so reduced the number of icache syncs that there just isn't as much to gain from optimizing the sync. That said, I'd still be interested in seeing results from the Orion-O6.

One open question is whether my 32KB threshold for switching to ialluis is too low. In other words, should we do range-based syncs on 64KB page mappings? I think that is something we should look into post commit.

In D60271#1383119, @alc wrote:

Yes. I would test GENERIC-NODEBUG kernels without and with this patch. I have /etc/src.conf:

WITHOUT_LLVM_ASSERTIONS=yes
#
WITH_CLEAN=yes
WITH_MALLOC_PRODUCTION=yes

I run:

#!/bin/csh
cd /usr/src
rm -fr /usr/obj/usr/src
while ( 1 )
    date
    sysctl vm.pmap vm.reserv vm.stats.vm.v_vm_faults
    time make -j16 buildworld > & /dev/null 
    rm -fr /usr/obj/usr/src
end

with -j adjusted for the machine and the script's output redirected to a log.

Unpatched output: https://reviews.freebsd.org/P717
Patched output: https://reviews.freebsd.org/P718

Timings from an unpatched kernel:

936.90 real     45789.09 user      3280.64 sys
933.97 real     45789.59 user      3314.96 sys
935.05 real     45774.25 user      3334.41 sys
934.40 real     45794.90 user      3302.60 sys
932.76 real     45846.53 user      3315.03 sys
933.13 real     45788.78 user      3318.50 sys

and a patched kernel:

938.43 real     45847.32 user      3243.15 sys
933.22 real     45670.20 user      3268.07 sys
933.84 real     45820.26 user      3269.70 sys
933.28 real     45730.27 user      3252.91 sys
939.21 real     45792.15 user      3287.05 sys
930.78 real     45835.46 user      3260.92 sys

which indicates that the change slightly reduced system CPU time.

Thanks. This system is ZFS-based, yes? Interestingly, you're getting 25% more L3C mappings than I see on UFS-based systems, presumably from loading executable files.