Page MenuHomeFreeBSD

LinuxKPI: preserve shmem backing ownership for managed PFN mappings
Needs ReviewPublic

Authored by oleglelchuk_gmail.com on Mon, Sep 7, 2:32 PM.
Tags
None
Referenced Files
F173724585: D59481.id186141.diff
Sun, Sep 27, 11:25 PM
F173707516: D59481.id186126.diff
Sun, Sep 27, 8:56 PM
Unknown Object (File)
Sun, Sep 27, 8:35 AM
Unknown Object (File)
Sun, Sep 27, 5:38 AM
Unknown Object (File)
Sun, Sep 27, 2:43 AM
Unknown Object (File)
Sat, Sep 26, 12:43 AM
Unknown Object (File)
Thu, Sep 24, 7:11 AM
Unknown Object (File)
Thu, Sep 24, 6:45 AM

Details

Summary

Current export, September 21: complete standalone replacement against main 46d383af0dbdb3871354f344b9c1ba99da691121, pinned and rechecked before the new build. Additional review found external-page discard paths that bypassed cache-attribute restoration and an OBJ_DEAD lookup/destructor window that could outlast the driver's final page wire. The correction adds an external-page release callback and one pager-owned pin per distinct preserved physical page per VMA, with invalidation-sequence checks and cleanup before final unwiring. These are defects in the earlier candidate, distinct from the original ownership-loss hang.

Exact standalone patch:

. Application/postimage and validation record: . The revised Test Plan separates source-extracted checks and a successful combined-kernel build/boot from the live comparisons that remain outstanding. The patch remains LLM-assisted and requires independent review.

A recorded FreeBSD/i915 buffer lifetime showed a still-live graphics buffer losing its original nonzero backing contents across page retirement and reacquisition. Replacement backing was zero-filled before a later submission. Source analysis identifies a path that transfers pages out of the original shmem object. The recorded loss is consistent with that ownership transfer; exact GPU instruction fetch and a matched single-change causal comparison have not been established.

On my Intel Meteor Lake laptop, my experience with large LAN downloads has been:

  • Non-PFN kernels with my manual ZFS ARC cap removed: I have experienced a recoverable GPU hang in every attempt, using either Sway or Plasma.
  • Non-PFN kernels with the 4 GiB ARC cap: I have not experienced this hang under that usage pattern.
  • PFN kernels with my manual ARC cap removed: I have not experienced this hang under the same general usage pattern.

These are my personal observations. I have not supplied a formal trial count or performed a matched, single-change kernel comparison. The first result describes all of my attempts, not a universal 100% failure probability. The affected desktop can be recovered by terminating the compositor; these failures are not kernel panics. Here, "uncapped ARC" means automatic ARC sizing with a finite effective maximum.

The cap comparison supports memory reclamation or its timing as a possible condition exposing the ownership defect. There is a concrete source-level route: FreeBSD low-memory notifications invoke registered LinuxKPI shrinkers; the i915 shrinker can release eligible graphics-buffer pages. FreeBSD's ARC implementation also responds to low-memory notifications. A larger resident ARC can leave less immediately available memory and change the demand for reclaim. This is a plausible route from the cache-size difference to a buffer release, not a recorded call trace identifying the caller that released the affected buffer.

The capture established that a still-live buffer's nonzero backing page was retired and the same logical buffer later acquired zero-filled replacement backing before submission. It did not require an operation that overwrote the old physical page with zeros. The source-level ownership transfer is consistent with loss of recoverable contents across that transition. The working explanation is that the 4 GiB cap avoids or changes the problematic release/reacquisition sequence in my workload, while the PFN change preserves backing ownership when that sequence occurs. The cap does not repair the ownership logic.

The caller that initiated this buffer's release remains unidentified, and the GPU's subsequent instruction fetch was not observed. The later ARC/VM sample in

, including no recorded swap I/O and free pages above the target, does not reconstruct conditions at retirement or exclude earlier reclamation. I have used "memory pressure" informally; neither the downloads nor that later sample establish its precise state at the content-loss transition.

Preserve managed OBJT_SWAP/shmem pages in their backing object and supply exclusively busy pages through an external device-pager handoff. Retain references to tracked backing owners until VMA teardown, consume each handed-off page once, and invalidate mappings without removing supported pages from their original backing object. The change also balances temporary requested-page holds, delays their publication until the required operations succeed, and rejects selecting a page already held by the same population transaction.

The pager interface makes memory-locking behavior explicit. Managed LinuxKPI IO/PFN mappings opt into a user-wiring exemption before object publication. Ordinary mappings retain their locking requirements. Kernel-origin wiring retains its separate requirements, and conservative resource-limit admission and RACCT reservations still apply. The manual changes document this policy and correct an inherited statement about per-process limits.

The patch is a candidate correction for the observed managed-shmem path. Other managed backing-owner types retain legacy transfer behavior. Invalidation visits the exact currently pinned pages and can revoke their aliases beyond the requested VMA subrange. The extra pager wire survives driver unwiring until mapping revocation and cache-attribute restoration; an unmapped discarded handoff can release its pin earlier. References to previously used owners remain until VMA teardown. Review is requested on these lifetime, invalidation, pager-interface and memory-locking decisions.

The source patch contains the PFN/VM changes, their manuals, and a kernel interface version marker. Local diagnostic instrumentation and unrelated laptop changes are excluded. Kernel/module pager interfaces change and require compatible kernel and module builds.

The __FreeBSD_version bump from 1600026 to the private value 1600028 is provisional for this rebased review. The accepting committer must confirm or adjust the final number against the landing tree and add the corresponding Porter's Handbook version entry with the actual commit identity. This review does not reserve an official version number. The manual pages retain their base dates until merge.

Related work

  • D32090 already discussed the risk of shared-memory contents being lost when pages move between objects. This submission makes no claim to have first identified that hazard.
  • drm-kmod issue 481 and Bugzilla 296448 report related PFN insertion failures. The proposed mapped-page handling change in issue 481 still removes pages from their original object, and subsequent feedback reports rendering corruption and GPU hangs. Whether those reports contain the exact locally recorded lifetime failure remains unverified.
  • drm-kmod PR 484 questions the existing ownership model and proposes exporter-provided VM objects for dma-buf mmap, primarily for udmabuf. That is relevant design work with which this proposal may need coordination.

The public-source comparison has not established an equivalent implementation or the absence of one. Direction toward an existing review or preferred mapping design would be welcome.

This patch and supporting analysis were developed with AI assistance. Independent maintainer review of the implementation and design is requested.

Test Plan

Current correction, September 21, 2026, against canonical FreeBSD main 46d383af0dbdb3871354f344b9c1ba99da691121:

  • The complete standalone patch is . Strict forward application and actual application passed in a pristine full-tree Git index; reverse application restored the exact upstream tree. The twelve resulting PFN postimages match those included in the locally built combined kernel. Public application/hash records: .
  • The old vm_fault_populate_cleanup() omission was reproduced using the extracted published function. An unmapped page retained the wrong cache attribute after busy release. This is an offline cleanup reproduction, not a live kernel panic.
  • The corrected source-extracted suite passed 22 distinct cases under ASan/UBSan. It covers external-page success/discard paths, pager pin ownership and deduplication, allocation failure, OBJ_DEAD lookup, teardown ordering, and invalidation overlapping busy handoffs. Two overlap cases use real pthreads, but VM/map locking, pmap and page queues are mocked. Three deliberately broken variants (missing release callback, pager wire, or epoch checks) were rejected. Twenty further suite runs also passed; these are repetitions, not additional distinct tests.
  • Fourteen full-producer test groups were rerun against functions extracted from the actual combined build tree, covering ownership, alias refusal, busy retry, object/pin/chunk allocation failure, cache cleanup and VMA growth. Those VM/pmap operations are also mocked. Both manual pages pass lint without date changes.
  • A full amd64 GENERIC kernel and base modules were built, followed by pristine Intel drm-kmod f252a30f27d157d9c763cd408850775096a6263f and Intel firmware 2c12915d69954673c782a9a55c6ec39c19cfea93 against that kernel. INVARIANTS and WITNESS remain enabled. I have rebooted into this combined kernel; running-kernel and loaded-module identity/hash checks passed. This is not a standalone-only build or matched comparison. No new live graphics/reclamation or audio-quality test is claimed.
  • DSB remains disabled for the separate display-startup workaround. The PFN correction does not change vmap() or establish the cause of that DSB problem.

The callback interface changes require matching kernel/module builds. __FreeBSD_version=1600028 is a private provisional marker on upstream 1600026, not an official allocation; a committer must select the landing-tree value and document it appropriately.

Historical results, including the earlier 25 live API cases, 74 preservation cycles and 374 closed tracked lifetimes, remain historical and have not been rerun for this revision. A matched standalone-only graphics/reclamation comparison, repeated panic-trigger workload, live mlock/RACCT checks, whole-buffer preservation proof, exact GPU instruction-fetch trace and exhaustive real-kernel cache/TLB/locking concurrency validation remain outstanding. Passing mock tests, compiling with WITNESS and successfully booting do not establish all interleavings. The OBJ_DEAD/last-driver-wire gap is now explicitly protected by pager-owned pins and modeled in tests; independent review and live validation of that protection remain necessary.

Diff Detail

Repository
rG FreeBSD src repository
Lint
Lint Skipped
Unit
Tests Skipped

Event Timeline

oleglelchuk_gmail.com created this object with edit policy "Administrators".

Because an LLM wrote the patch, summary, and test plan, you can see things such as "The user reports", instead of "I noticed" while reading all this stuff.

oleglelchuk_gmail.com edited the summary of this revision. (Show Details)
oleglelchuk_gmail.com edited the test plan for this revision. (Show Details)

Clarification of my earlier comment: descriptions of my own desktop experience now use the first person. The patch, technical analysis, summary and test plan remain AI-assisted and are submitted for independent maintainer review.

This update adds explicit LinuxKPI and Virtual Memory review requests, a provisional kernel interface version bump, corrected manual-date handling, and a public evidence summary. It records my repeated non-PFN/PFN desktop observations without treating "memory pressure" as a measured explanation or claiming a matched causal comparison.

Historical evidence and current export summary:


Export validation and per-file hashes:

oleglelchuk_gmail.com edited the test plan for this revision. (Show Details)

It is really a lot of work to try and understand the review description. Let's please try to simplify the description of what's going on.

It sounds like the i915 GEM shrinker is reclaiming pages from a shared memory object which backs a GEM object, and this is happening while the GPU is actively working on the object. Is that right? Can you please point to the exact code paths where this happens? I would naively expect that i915 or DRM somehow pins active objects, precisely to avoid this problem, so we will have to understand why that isn't working.

If my understand is wrong, then please try to succinctly explain what is going on, in terms of actual code being executed. Writing something like, "source analysis identifies a path that transfers pages out of the original shmem object," is very unhelpful unless you actually point to the path you're talking about.

Let me explain what happened. I noticed a problem: when ZFS ARC was automatically sized and I started downloading multi-gigabyte files over my home LAN, eventually a recoverable GPU hang would happen on either Sway or Plasma. If, for example, I downloaded around 4 or 5 multi-gigabyte files in the automatically sized ZFS ARC environment, I would get a 100% guarantee a recoverable GPU hang would happen on either Sway or Plasma. I would have to remotely kill kwin_wayland_wrapper and kwin_wayland through ssh; otherwise, the computer wouldn't be accessible directly. Sway wouldn't block direct access, but I would still have to kill it due to such hangs. When I capped ZFS ARC and repeated the file downloading experiments, GPU hangs stopped happening, so I thought the problem was resolved. However, GPU hangs returned on Sway/Plasma when an LLM I was working with started doing computationally intensive things in this capped ARC environment. I asked the LLM to explain to me the cause of these GPU hangs, and it told me I was dealing with an ownership/bookkeeping bug which caused pages to lose access to their original backing objects and that's why these GPU hangs occurred. I asked it to fix the bug, and with the code that I later posted here, I stopped experiencing GPU hangs in either automatically sized ZFS ARC combined with file downloading situations or LLM-doing-computationally-intensive-work situations. I am not really knowledgeable about computer-related stuff, so I don't know what was going on behind the scenes when the GPU hangs occurred. Maybe a memory pressure thing indirectly triggered this ownership bug. All I know is that I encountered a bug and the LLM seems to have fixed it for me.

Let me explain what happened. I noticed a problem: when ZFS ARC was automatically sized and I started downloading multi-gigabyte files over my home LAN, eventually a recoverable GPU hang would happen on either Sway or Plasma. If, for example, I downloaded around 4 or 5 multi-gigabyte files in the automatically sized ZFS ARC environment, I would get a 100% guarantee a recoverable GPU hang would happen on either Sway or Plasma. I would have to remotely kill kwin_wayland_wrapper and kwin_wayland through ssh; otherwise, the computer wouldn't be accessible directly. Sway wouldn't block direct access, but I would still have to kill it due to such hangs. When I capped ZFS ARC and repeated the file downloading experiments, GPU hangs stopped happening, so I thought the problem was resolved. However, GPU hangs returned on Sway/Plasma when an LLM I was working with started doing computationally intensive things in this capped ARC environment. I asked the LLM to explain to me the cause of these GPU hangs, and it told me I was dealing with an ownership/bookkeeping bug which caused pages to lose access to their original backing objects and that's why these GPU hangs occurred. I asked it to fix the bug, and with the code that I later posted here, I stopped experiencing GPU hangs in either automatically sized ZFS ARC combined with file downloading situations or LLM-doing-computationally-intensive-work situations. I am not really knowledgeable about computer-related stuff, so I don't know what was going on behind the scenes when the GPU hangs occurred. Maybe a memory pressure thing indirectly triggered this ownership bug. All I know is that I encountered a bug and the LLM seems to have fixed it for me.

The problem is that we will not apply this patch without understanding exactly what it is doing, and it is very hard to understand what it is doing without understanding the problem it fixes. So to make progress, we will need more details about the problem, as I asked for. It may be that the underlying bug can be fixed with a two-line patch instead of the one the LLM generated.

Not quite. Since my last comment, I reproduced the hang with an unmodified
FreeBSD main GENERIC kernel at 7f8feee5ab922e6645f4061e9ca29109411aadef
and unmodified drm-kmod at 1d04976c1870bce01094b78e46c48f6eab810867.
With the LLM's assistance, I also collected a live DTrace recording. It does
not show the shrinker reclaiming pages while the GPU is actively using them.
__i915_gem_object_put_pages() refuses an object whose backing pages are
pinned. The problem is that an earlier FreeBSD PFN mmap changes the ownership
of the pages, so a later legitimate page release can destroy backing that i915
expects shmem to retain.

The mapping path is:

i915_gem_object_mmap()
  -> vm_fault_cpu()
  -> remap_io_sg()
  -> remap_sg()
  -> lkpi_vmf_insert_pfn_prot_locked()

vm_fault_cpu() pins the GEM object's shmem-backed scatter/gather pages while
calling remap_io_sg(). On Linux, remap_sg() installs a special PFN PTE and
leaves the page owned by shmem. On FreeBSD it calls
lkpi_vmf_insert_pfn_prot_locked(). For every managed page still belonging to
the shmem OBJT_SWAP object, that function calls
vm_pager_page_unswapped(page), vm_page_remove(page), and then
vm_page_iter_insert(page, vm_obj, pindex, ...). This discards the page's swap
backing, removes it from shmem, and inserts the physical page into the mmap's
OBJT_MGTDEVICE object. The exact implementation is
[linux_page.c:508-575](https://github.com/freebsd/freebsd-src/blob/7f8feee5ab922e6645f4061e9ca29109411aadef/sys/compat/linuxkpi/common/src/linux_page.c#L508-L575),
called by
[i915_mm.c:89-115](https://github.com/freebsd/drm-kmod/blob/1d04976c1870bce01094b78e46c48f6eab810867/drivers/gpu/drm/i915/i915_mm.c#L89-L115).

During reclaim, FreeBSD's
[linuxkpi_vm_lowmem()](https://github.com/freebsd/freebsd-src/blob/7f8feee5ab922e6645f4061e9ca29109411aadef/sys/compat/linuxkpi/common/src/linux_shrinker.c#L99-L142)
invokes i915's shrinker. i915 selects a releasable object and calls
__i915_gem_object_put_pages()
([i915_gem_shrinker.c:175-230](https://github.com/freebsd/drm-kmod/blob/1d04976c1870bce01094b78e46c48f6eab810867/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c#L175-L230)).
That function checks for pinned pages, revokes the mmap, unsets the pages, and
calls the shmem put-pages callback
([i915_gem_pages.c:239-267](https://github.com/freebsd/drm-kmod/blob/1d04976c1870bce01094b78e46c48f6eab810867/drivers/gpu/drm/i915/gem/i915_gem_pages.c#L239-L267)).
Once VMA invalidation or pager teardown removes a page from the temporary
OBJT_MGTDEVICE object, it is objectless because PFN insertion already removed
it from shmem. When shmem_sg_free_table() drops the last wire, an objectless
page with raw ref_count == 1 enters vm_page_free_toq(). The original shmem
slot now has neither a resident page nor swap backing. A later
linux_shmem_read_mapping_page_gfp() for that offset must create replacement
backing rather than recover the original graphics-buffer bytes. Thus the object
can be correctly inactive and unpinned when reclaimed, yet still be alive and
expected to retain contents for later use.

The DTrace run recorded 3,151 PFN insertions from a short vkcube trigger. All
3,151 entered swap_pager_unswapped, were removed from their source objects,
and were inserted into the target pager object. On release, 70 pages were
unwired; 64 were objectless with raw ref_count == 1, and exactly those 64
entered vm_page_free_toq(). The trace also observed four
__i915_gem_object_put_pages() calls in pagedaemon. The program and raw
aggregate output are attached as

(SHA-256
5199187838078a783797ea21882db1f8dc8e98344d7725c3780d29676a1e0614).

In the clean Sway reproduction, ARC grew to about 11.91 GiB, free memory reached
the VM reclaim boundary, and ARC reclaim began. Sway stopped responding, then
the render engine reported GPU HANG: ecode 12:1:85dff5fb, in sway three
seconds later. This supports ARC growth as the condition that invokes reclaim;
ARC itself is not modifying the graphics buffer.

I have not identified the exact Sway buffer from its original contents through
reclaim to the stalled GPU submission, and I have not yet performed a matched
run changing only this fix. The ownership-loss lifecycle itself was observed
directly, while the final buffer-to-hang link remains strongly supported rather
than fully proved. A smaller fix may be possible; I am not claiming that the
current twelve-file implementation is minimal. The required semantic change is
to map these managed pages without removing them from their shmem backing object
or discarding that object's backing information.

I think I understand what is going on.

Long time ago, when I first ported the GEM infrastructure from Linux to FreeBSD, I remember that the principle was that the given page can exists only in single domain (this is not same as NUMA proximity domain, but like CPU-cached or GPU-cached indicator). The GPU L1/L2 were not coherent with the CPU L1/L2, and each domain change of the page required explicit handling of the cache. This is why the pmap_invalidate_cache_pages() was added which cannot rely on the self-snoop capability of (Intel) CPUs.

Now, I believe that Intel GPUs grown the coherent L3 between CPU and GPU, and L1/L2 caches are at least invalidated when cache line crosses domain. GPU home agent might even participate in the MESI as another coherent device, but this is not important for us. What is important is that it seems to be valid for page to be both accessed by CPU and GPU, effectively making the page existing in both CPU and GPU caching domains simultaneously.

The object' ownership model for pages cannot easily model that. I suspect we have to do something weird like allowing fake pages to exists in the swap objects. So for instance when lkpi_vmf_insert_pfn_prot_locked() observes that page->object != vma->object, and page->object is the managed device pager, then instead of removing the page from the page->object, we keep it there but install a fake page into vma->object. Then something special would have to be done for pageout when it observes such page (skip), and perhaps in dozen other places which are not ready to find non-fictious pages in the swap objects.

Please add 'alc' to the reviewers.

In D59481#1370433, @kib wrote:

I think I understand what is going on.

Long time ago, when I first ported the GEM infrastructure from Linux to FreeBSD, I remember that the principle was that the given page can exists only in single domain (this is not same as NUMA proximity domain, but like CPU-cached or GPU-cached indicator). The GPU L1/L2 were not coherent with the CPU L1/L2, and each domain change of the page required explicit handling of the cache. This is why the pmap_invalidate_cache_pages() was added which cannot rely on the self-snoop capability of (Intel) CPUs.

Now, I believe that Intel GPUs grown the coherent L3 between CPU and GPU, and L1/L2 caches are at least invalidated when cache line crosses domain. GPU home agent might even participate in the MESI as another coherent device, but this is not important for us. What is important is that it seems to be valid for page to be both accessed by CPU and GPU, effectively making the page existing in both CPU and GPU caching domains simultaneously.

The object' ownership model for pages cannot easily model that. I suspect we have to do something weird like allowing fake pages to exists in the swap objects. So for instance when lkpi_vmf_insert_pfn_prot_locked() observes that page->object != vma->object, and page->object is the managed device pager, then instead of removing the page from the page->object, we keep it there but install a fake page into vma->object. Then something special would have to be done for pageout when it observes such page (skip), and perhaps in dozen other places which are not ready to find non-fictious pages in the swap objects.

My fuzzy understanding of the code is that lkpi_vmf_insert_pfn_prot_locked() is moving pages from the shmem object to the MGTDEVICE object. In that case it sounds like it should instead 1) allocate a fake page with the same paddr and insert it into the MGTDEVICE object, and 2) wire the source page.

If userspace wishes to map GPU memory into CPU page tables, then presumably it can map a MGTDEVICE directly? That is, there is no need to move a page from MGTDEVICE->SWAP.

In D59481#1370433, @kib wrote:

I think I understand what is going on.

Long time ago, when I first ported the GEM infrastructure from Linux to FreeBSD, I remember that the principle was that the given page can exists only in single domain (this is not same as NUMA proximity domain, but like CPU-cached or GPU-cached indicator). The GPU L1/L2 were not coherent with the CPU L1/L2, and each domain change of the page required explicit handling of the cache. This is why the pmap_invalidate_cache_pages() was added which cannot rely on the self-snoop capability of (Intel) CPUs.

Now, I believe that Intel GPUs grown the coherent L3 between CPU and GPU, and L1/L2 caches are at least invalidated when cache line crosses domain. GPU home agent might even participate in the MESI as another coherent device, but this is not important for us. What is important is that it seems to be valid for page to be both accessed by CPU and GPU, effectively making the page existing in both CPU and GPU caching domains simultaneously.

The object' ownership model for pages cannot easily model that. I suspect we have to do something weird like allowing fake pages to exists in the swap objects. So for instance when lkpi_vmf_insert_pfn_prot_locked() observes that page->object != vma->object, and page->object is the managed device pager, then instead of removing the page from the page->object, we keep it there but install a fake page into vma->object. Then something special would have to be done for pageout when it observes such page (skip), and perhaps in dozen other places which are not ready to find non-fictious pages in the swap objects.

My fuzzy understanding of the code is that lkpi_vmf_insert_pfn_prot_locked() is moving pages from the shmem object to the MGTDEVICE object. In that case it sounds like it should instead 1) allocate a fake page with the same paddr and insert it into the MGTDEVICE object, and 2) wire the source page.

If userspace wishes to map GPU memory into CPU page tables, then presumably it can map a MGTDEVICE directly? That is, there is no need to move a page from MGTDEVICE->SWAP.

I do not think the direction matters. In particular, wiring the source swap-backed page does not help: when would it be unwired? What prevents the source swap object from termination?

In D59481#1371356, @kib wrote:
In D59481#1370433, @kib wrote:

I think I understand what is going on.

Long time ago, when I first ported the GEM infrastructure from Linux to FreeBSD, I remember that the principle was that the given page can exists only in single domain (this is not same as NUMA proximity domain, but like CPU-cached or GPU-cached indicator). The GPU L1/L2 were not coherent with the CPU L1/L2, and each domain change of the page required explicit handling of the cache. This is why the pmap_invalidate_cache_pages() was added which cannot rely on the self-snoop capability of (Intel) CPUs.

Now, I believe that Intel GPUs grown the coherent L3 between CPU and GPU, and L1/L2 caches are at least invalidated when cache line crosses domain. GPU home agent might even participate in the MESI as another coherent device, but this is not important for us. What is important is that it seems to be valid for page to be both accessed by CPU and GPU, effectively making the page existing in both CPU and GPU caching domains simultaneously.

The object' ownership model for pages cannot easily model that. I suspect we have to do something weird like allowing fake pages to exists in the swap objects. So for instance when lkpi_vmf_insert_pfn_prot_locked() observes that page->object != vma->object, and page->object is the managed device pager, then instead of removing the page from the page->object, we keep it there but install a fake page into vma->object. Then something special would have to be done for pageout when it observes such page (skip), and perhaps in dozen other places which are not ready to find non-fictious pages in the swap objects.

My fuzzy understanding of the code is that lkpi_vmf_insert_pfn_prot_locked() is moving pages from the shmem object to the MGTDEVICE object. In that case it sounds like it should instead 1) allocate a fake page with the same paddr and insert it into the MGTDEVICE object, and 2) wire the source page.

If userspace wishes to map GPU memory into CPU page tables, then presumably it can map a MGTDEVICE directly? That is, there is no need to move a page from MGTDEVICE->SWAP.

I do not think the direction matters. In particular, wiring the source swap-backed page does not help: when would it be unwired?

When the driver knows that the GPU is no longer mapping the pages. In fact it already implements this mechanism: each GEM object type provides get_pages, which prepares the object's backing pages to be mapped into the GTT. For shmem objects, this is implemented by shmem_get_pages() -> shmem_sg_alloc_table() -> shmem_read_folio_gfp() -> vm_page_grab_valid(VM_ALLOC_WIRED).

The problem is with the implementation of lkpi_vmf_insert_pfn_prot_locked(): it moves pages out of the shmem object, so a subsequent vm_page_grab will not see moved pages, and it will allocate and zero-fill new ones. remap_pfn() and remap_sg() in i915 should be allocating fake pages instead of moving pages between objects.

What prevents the source swap object from termination?

Why does it matter? Even if the object is terminated (I would expect the driver to prevent this), wired pages will not be freed.

The technical language that you guys are using is above my head. I take it that you guys now understand the nature of the bug, so I'll just move out of your way and let you fix it. Thank you in advance for the fix! I don't understand though why no one else reported the bug before I did it myself: doesn't a regular Joe who uses Sway/Plasma often want to download files over his home LAN without changing the default "automatically sized ZFS ARC" setting first? He doesn't need to change this setting because he has no idea in the first place it needs to be changed.

In D59481#1371356, @kib wrote:
In D59481#1370433, @kib wrote:

I think I understand what is going on.

Long time ago, when I first ported the GEM infrastructure from Linux to FreeBSD, I remember that the principle was that the given page can exists only in single domain (this is not same as NUMA proximity domain, but like CPU-cached or GPU-cached indicator). The GPU L1/L2 were not coherent with the CPU L1/L2, and each domain change of the page required explicit handling of the cache. This is why the pmap_invalidate_cache_pages() was added which cannot rely on the self-snoop capability of (Intel) CPUs.

Now, I believe that Intel GPUs grown the coherent L3 between CPU and GPU, and L1/L2 caches are at least invalidated when cache line crosses domain. GPU home agent might even participate in the MESI as another coherent device, but this is not important for us. What is important is that it seems to be valid for page to be both accessed by CPU and GPU, effectively making the page existing in both CPU and GPU caching domains simultaneously.

The object' ownership model for pages cannot easily model that. I suspect we have to do something weird like allowing fake pages to exists in the swap objects. So for instance when lkpi_vmf_insert_pfn_prot_locked() observes that page->object != vma->object, and page->object is the managed device pager, then instead of removing the page from the page->object, we keep it there but install a fake page into vma->object. Then something special would have to be done for pageout when it observes such page (skip), and perhaps in dozen other places which are not ready to find non-fictious pages in the swap objects.

My fuzzy understanding of the code is that lkpi_vmf_insert_pfn_prot_locked() is moving pages from the shmem object to the MGTDEVICE object. In that case it sounds like it should instead 1) allocate a fake page with the same paddr and insert it into the MGTDEVICE object, and 2) wire the source page.

If userspace wishes to map GPU memory into CPU page tables, then presumably it can map a MGTDEVICE directly? That is, there is no need to move a page from MGTDEVICE->SWAP.

I do not think the direction matters. In particular, wiring the source swap-backed page does not help: when would it be unwired?

When the driver knows that the GPU is no longer mapping the pages. In fact it already implements this mechanism: each GEM object type provides get_pages, which prepares the object's backing pages to be mapped into the GTT. For shmem objects, this is implemented by shmem_get_pages() -> shmem_sg_alloc_table() -> shmem_read_folio_gfp() -> vm_page_grab_valid(VM_ALLOC_WIRED).

The problem is with the implementation of lkpi_vmf_insert_pfn_prot_locked(): it moves pages out of the shmem object, so a subsequent vm_page_grab will not see moved pages, and it will allocate and zero-fill new ones. remap_pfn() and remap_sg() in i915 should be allocating fake pages instead of moving pages between objects.

Might be, a fake page is probably the closest analog to the special PTE mentioned in the original linux driver sources.

What prevents the source swap object from termination?

Why does it matter? Even if the object is terminated (I would expect the driver to prevent this), wired pages will not be freed.

And the device object termination needs to unwire pages found by PHYS_TO_VM_PAGE(VM_PAGE_TO_PHYS(fake)) for each fake page owned by it?

oleglelchuk_gmail.com edited the summary of this revision. (Show Details)
oleglelchuk_gmail.com edited the test plan for this revision. (Show Details)

I have uploaded a replacement standalone patch based on FreeBSD main eca31490322f181a415c14c6bf03e656e81fb4aa, pinned on September 19. The exact patch file is

. It contains only the twelve PFN/VM/header/manual changes; my unrelated laptop patches are excluded.

This update is needed for two reasons:

  • The old diff needs rebasing onto newer main. The LLM retained the upstream PFN bounds checks and updated the sparse PFN state's page-index limit when an existing VMA grows.
  • I encountered two page ... has an unexpected memattr panics with the earlier PFN implementation. Both dumps show an unwired i915 shmem page left write-combining in an OBJT_SWAP object whose attribute is write-back/default. Wi-Fi packet-buffer allocation triggered the contiguous-reclamation check that detected this state. The LLM's analysis strongly implicates missing cache-attribute cleanup in the PFN patch; the dumps do not record the last attribute setter or establish that every possible independent cause is excluded.

In sys/compat/linuxkpi/common/src/linux_page.c, the new lkpi_vma_pfn_restore_memattr() helper restores the backing object's attribute after managed CPU mappings are revoked. It is used during invalidation, failed population and teardown. Cleanup preserves the original owner and backing contents, leaves the attribute alone while a managed CPU alias survives, and rejects a non-default shmem mapping if the driver has not wired the page. The VM assertion remains intact.

I consider it unlikely that this twelve-file approach will be merged upstream as it stands, given the alternative mapping design being discussed here. Nevertheless, people who encountered the same recoverable GPU hang during reclamation that prompted my report may want to test this version as an interim quick fix while waiting for an upstream solution. It remains an experimental workaround, not a claim that every PFN lifetime problem is solved. In particular, ordering of pager teardown against the driver's final unwire still needs review.

The LLM ran the updated offline regression and patch-application checks. I am now running a rebuilt GENERIC kernel containing these exact PFN postimages alongside my separate hardware changes, with vanilla drm-kmod. This is not a matched test of the standalone patch alone, and the previous live preservation/API results have not been rerun on this revision. The updated Test Plan distinguishes those limits from the completed checks.

oleglelchuk_gmail.com edited the summary of this revision. (Show Details)
oleglelchuk_gmail.com edited the test plan for this revision. (Show Details)

I have uploaded a new standalone patch against FreeBSD main 46d383af0dbdb3871354f344b9c1ba99da691121:

. This is a complete replacement, not an incremental patch to apply over the previous version. It contains only the twelve PFN/VM/header/manual changes.

Two additional bugs were found in the September 20 implementation. The LLM prepared a correction, and I hope it addresses both; independent review and live workload validation are still needed.

  1. Discarded external-page handoffs bypassed cache-attribute restoration. In sys/vm/vm_fault.c, vm_fault_populate_cleanup() and several retry/error branches in vm_fault_populate() released external pages directly with vm_page_xunbusy(). Once a page had been consumed from the handoff, the producer's abort cleanup could no longer find it. The correction adds populate_release_page through the VM/cdev pager interfaces. In sys/compat/linuxkpi/common/src/linux_page.c, lkpi_vma_pfn_release_page() restores the backing attribute before releasing an unmapped page; a surviving managed CPU mapping retains its attribute and pin.
  1. A dying pager could miss invalidation before its destructor restored the pages. vm_object_deallocate() marks the pager OBJ_DEAD, and vm_pager_object_lookup() skips dead objects. Consequently, lkpi_unmap_mapping_range() can miss a pager whose linux_cdev_pager_dtor() is still waiting for mmap_sem, while the driver finishes invalidation and drops its page wire. An object reference alone does not pin its resident pages. The correction keeps one pager-owned wire per distinct preserved physical page per VMA. lkpi_vma_pfn_unmap_pins() removes mappings, restores the backing attribute, releases busy ownership, then drops that wire. Epoch checks prevent stale population from adding pins after invalidation, and a pin already detached by invalidation remains owned by that invalidator.

The first omission was reproduced with the extracted old cleanup function. The second is supported by the source ordering and modeled interleaving tests; it is not a new live panic reproduction. Neither result proves which exact path caused my earlier panics. The existing VM assertions remain intact.

The LLM ran strict patch-application/reversal checks, 22 source-extracted regression cases with ASan/UBSan, three deliberately broken negative controls, and 14 full-producer test groups. The extra 20 suite runs were repetitions. VM/pmap behavior in those userspace tests is mocked. Validation and postimage hashes:

.

I have rebooted into a full GENERIC build containing these exact PFN changes alongside my separate hardware changes, with matching vanilla drm-kmod f252a30f27d157d9c763cd408850775096a6263f. Boot and module-identity checks passed. This is combined-kernel evidence; I have no new matched standalone-only graphics/reclamation result to report. DSB remains disabled, and this update makes no claim to resolve the separate display-startup issue.

The patch still uses the external-page handoff design, not the alternative fake-page design discussed here. Pins can retain pages until invalidation or VMA destruction; broad alias revocation and the user-mlock exemption still need review. The private interface marker is now 1600028, and matching kernel/modules must be rebuilt. This remains an experimental interim workaround for the original recoverable GPU hang, not an upstream-approved or comprehensively validated solution.

(I may not be following, please excuse me if my assessment is incorrect)
Please stop this.
If you don't know what is going on in the patch, and cannot explain it, and use many words, also the patch is large, it is not useful and hurts.

Please file a bug report concisely explaining, in your own words, what you are observing, and what you expect to see instead.
Here are some instructions to get you started: https://freebsdfoundation.org/our-work/journal/browser-based-edition/embedded-2/writing-effective-bug-reports/

(I may not be following, please excuse me if my assessment is incorrect)
Please stop this.
If you don't know what is going on in the patch, and cannot explain it, and use many words, also the patch is large, it is not useful and hurts.

Please file a bug report concisely explaining, in your own words, what you are observing, and what you expect to see instead.
Here are some instructions to get you started: https://freebsdfoundation.org/our-work/journal/browser-based-edition/embedded-2/writing-effective-bug-reports/

Hi. I was under the impression that the other folks who spoke in this space understood what the bug is. Pages that, from a user's point of view, contained valid data were mistakenly replaced by zero-filled pages; that's why GPU hangs occurred. The original patch and patch #2 fixed this GPU hang bug, but in turn, they introduced other bugs, so, hopefully, patch #3, which is the current patch, fixed those other bugs too, without introducing new bugs.

(I may not be following, please excuse me if my assessment is incorrect)
Please stop this.
If you don't know what is going on in the patch, and cannot explain it, and use many words, also the patch is large, it is not useful and hurts.

Please file a bug report concisely explaining, in your own words, what you are observing, and what you expect to see instead.
Here are some instructions to get you started: https://freebsdfoundation.org/our-work/journal/browser-based-edition/embedded-2/writing-effective-bug-reports/

Hi. I was under the impression that the other folks who spoke in this space understood what the bug is. Pages that, from a user's point of view contained valid data, were mistakenly replaced by zero-filled pages; that's why GPU hangs occurred. The original patch and patch #2 fixed this GPU hang bug, but in turn, they introduced other bugs, so, hopefully, patch #3, which is the current patch, fixed those other bugs too, without introducing new bugs.

The VM folks are wrapping their heads around it and likely will figure out a much smaller, easier solution. But! You're at least exploring and trying to identify the behavioural issues, so do keep tinkering. The LLM prose may be a bit much (eg we don't need to know about your local test cases passing and such) but the dives into the code paths being executed in different situations are likely helpful here.