User Details
- User Since
- Dec 14 2014, 5:52 AM (616 w, 1 d)
Sun, Oct 4
Sat, Oct 3
Here is Claude's summary of a log that I collected on the DevKit:
I've been reviewing places where we already sync the icache or might need to.
Fri, Oct 2
Do any of you have access to a machine with IDC, but not DIC, e.g., an older Ampere I think? The results on a small Cortex-X1/A78 system are at best inclusive.
@andrew, Is this what you were asking for? A 16 processor AWS a1.4xlarge, which is Cortex A72-based, sees a small reduction in system time using arm64_pipt_icache_sync_range(). Specifically, the reduction is about 0.75%.
Wed, Sep 30
Sat, Sep 26
Fri, Sep 25
Here is a link to the paper I published a while back on PTE coalescing. At that time, the changes in Linux to exploit small folios were just beginning to get merged.
I've also tested the prior version of this patch on a 32-core machine (c6g.8xlarge), where icache coherence is maintained by the hardware and the sync operation is simply a dsb and an isb. I just wanted to make sure that the patch didn't negatively effect performance on such machines. I compared 18 buildworld runs on HEAD with 22 runs using this patch. The kernel was -NODEBUG, MALLOC PRODUCTION was enabled, and LLVM assertions were disabled. System time actually fell by 7.8 seconds, from 2,060 to 2,052 seconds, because we avoided a couple hundred million dsb and isb instructions at the cost of the flag maintenance. That is 0.4%, which is four times larger than the random error you would expect. User time fell by 21 seconds, or 0.05%, which is probably real but too small to assert a similar claim relative to the random error. Wall-clock time did not change measurably.
PGA_ICACHE_SYNCED and KASSERTs
Rename PGA_MACHDEP0004 to PGA_PMAP_PRIV1.
Thu, Sep 24
Wed, Sep 23
Sun, Sep 20
Wed, Sep 9
Tested. It works.
Tue, Sep 8
Yes, this will work. I will actually test it later today.
Mon, Sep 7
Introduce pmap_page_is_mapped_locked().
Sun, Sep 6
I recently updated my Microsoft Dev Kit, and because of this change I'm seeing:
bus_dmamem_alloc failed to align memory properly.
The warning stems from an allocation by the nvme driver:
KDB: stack backtrace: db_trace_self() at db_trace_self db_trace_self_wrapper() at db_trace_self_wrapper+0x44 bounce_bus_dmamem_alloc() at bounce_bus_dmamem_alloc+0x258 nvme_ctrlr_start() at nvme_ctrlr_start+0xbc8 nvme_ctrlr_start_config_hook() at nvme_ctrlr_start_config_hook+0x5ec run_interrupt_driven_config_hooks() at run_interrupt_driven_config_hooks+0x94 boot_run_interrupt_driven_config_hooks() at boot_run_interrupt_driven_config_hooks+0x30 mi_startup() at mi_startup+0x1ec virtdone() at virtdone+0x70
An added printf reports:
vaddr: 0xffffa000847d3f40, paddr: 1047d3f40, alignment: 1000
In this case, the misalignment appears to be harmless because page size alignment isn't actually needed.
Sep 4 2026
Aug 31 2026
I'm testing this patch on an EC2 a1.4xlarge (16x Cortex A72) machine to see the greatest impact. (This is the entire machine, so there isn't any variance due to other VMs on the machine.) I added a COUNTER_U64 to track icache flushes by the pmap. As expected, this patch results in an increased icache flush count, since we were not flushing on creating executable superpage mappings. During a -j16 buildworld on a -NODEBUG kernel, PRODUCTION malloc, and LLVM with assertions disabled, the number of flushes goes from ~190M to ~230M and wall clock time goes from ~2:09:10 to ~2:10:00. In particular, system time increased by ~6%.
Aug 29 2026
I uncovered this while resurrecting an old patch for reducing the number of icache flushes. If all goes well, I will post that in a week or two.
I'm curious as to why this isn't also a problem for the uses of .arch_extension in arm64/vfp.c, arm64/mte.c, and include/atomic.h?
Aug 18 2026
Revise a comment.
Aug 15 2026
Move td = curthread; out of the critical section.
Aug 14 2026
Aug 11 2026
Aug 10 2026
Aug 9 2026
I'm waiting to hear if @andrew has any comments or questions.
As an aside, as the appearance of the word "harness" in the previous comment suggests, when I had Claude reviewing the changes, it actually built a user-space test harness containing the loop to check that every page in the range would be the target of an invalidation. And, it temporarily introduced a few bugs in the loop to check that those bugs were detected. Just thought this was interesting.
I ran a dozen -j8 buildworlds, each starting from an empty /usr/obj, with MALLOC_PRODUCTION enabled and LLVM assertions disabled on a GENERIC-NODEBUG kernel. Here is Claude's analysis of the data:
The change is a win. before after delta sys 1,070.67 1,012.58 −58.1 s (−5.43%) user 23,801.46 23,816.08 +14.6 s (+0.06%) user+sys 24,872.13 24,828.66 −43.5 s (−0.17%) wall 3,221.63 (53:41.63) 3,214.40 (53:34.40) −7.2 s (−0.22%)
Aug 8 2026
While the conversion from a macro to an inline pmap_s1_invalidate_loop() had no real effect, just a couple instructions changed, pmap_s1_invalidate_strided() is not being inlined in some places due to its size. In that case, we are not benefiting from final_only always being a constant, but I think that is something for a different patch to deal with.
Convert macro to __always_inline function.
Aug 7 2026
Aug 1 2026
Jul 17 2026
I'm surprised that we didn't see this on arm64.
Jun 22 2026
May 7 2026
Apr 23 2026
Just out of curiosity, where are the pmap changes?
Apr 21 2026
Apr 20 2026
Apr 17 2026
Apr 16 2026
Mar 31 2026
"... pmap locks, than witness ..." -> "... pmap locks, then witness ..."
Mar 23 2026
Mar 5 2026
Feb 28 2026
@markj Do we know if there is actually a mix of clean and dirty pages being mapped?
Feb 27 2026
Jan 7 2026
Jan 6 2026
Jan 5 2026
Jan 1 2026
Dec 29 2025
Specifically, for now, I would commit the version in Diff 168641.
At this point, I would go back to the simpler, O(n) version, and commit that.
