Page MenuHomeFreeBSD

pfnmap-D59481-main-d2018cedb414-full-context.patch

Authored By
oleglelchuk_gmail.com
Tue, Sep 29, 3:38 PM
Size
467 KB
Referenced Files
None
Subscribers
None

pfnmap-D59481-main-d2018cedb414-full-context.patch

This file is larger than 256 KB, so syntax highlighting was skipped.
diff --git a/lib/libsys/mlock.2 b/lib/libsys/mlock.2
index 25346355a68ad0537b8d5aa24a968f25aa7502b8..1c624c35a25cae0d10e1504bc42ad3b013c93331 100644
--- a/lib/libsys/mlock.2
+++ b/lib/libsys/mlock.2
@@ -1,176 +1,186 @@
.\" Copyright (c) 1993
.\" The Regents of the University of California. All rights reserved.
.\"
.\" Redistribution and use in source and binary forms, with or without
.\" modification, are permitted provided that the following conditions
.\" are met:
.\" 1. Redistributions of source code must retain the above copyright
.\" notice, this list of conditions and the following disclaimer.
.\" 2. Redistributions in binary form must reproduce the above copyright
.\" notice, this list of conditions and the following disclaimer in the
.\" documentation and/or other materials provided with the distribution.
.\" 3. Neither the name of the University nor the names of its contributors
.\" may be used to endorse or promote products derived from this software
.\" without specific prior written permission.
.\"
.\" THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS ``AS IS'' AND
.\" ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
.\" IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
.\" ARE DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE
.\" FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
.\" DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS
.\" OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)
.\" HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT
.\" LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY
.\" OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
.\" SUCH DAMAGE.
.\"
.Dd May 13, 2019
.Dt MLOCK 2
.Os
.Sh NAME
.Nm mlock ,
.Nm munlock
.Nd lock (unlock) physical pages in memory
.Sh LIBRARY
.Lb libc
.Sh SYNOPSIS
.In sys/mman.h
.Ft int
.Fn mlock "const void *addr" "size_t len"
.Ft int
.Fn munlock "const void *addr" "size_t len"
.Sh DESCRIPTION
The
.Fn mlock
system call
locks into memory the physical pages associated with the virtual address
range starting at
.Fa addr
for
.Fa len
bytes.
The
.Fn munlock
system call unlocks pages previously locked by one or more
.Fn mlock
calls.
For both, the
.Fa addr
argument should be aligned to a multiple of the page size.
If the
.Fa len
argument is not a multiple of the page size, it will be rounded up
to be so.
The entire range must be allocated.
.Pp
After an
.Fn mlock
system call, the indicated pages will cause neither a non-resident page
nor address-translation fault until they are unlocked.
They may still cause protection-violation faults or TLB-miss faults on
architectures with software-managed TLBs.
The physical pages remain in memory until all locked mappings for the pages
are removed.
Multiple processes may have the same physical pages locked via their own
virtual address mappings.
A single process may likewise have pages multiply-locked via different virtual
mappings of the same physical pages.
Unlocking is performed explicitly by
.Fn munlock
or implicitly by a call to
.Fn munmap
which deallocates the unmapped address range.
Locked mappings are not inherited by the child process after a
.Xr fork 2 .
.Pp
+Special device mappings explicitly exempted by their pager are skipped by
+.Fn mlock
+and acquire no user wire reference.
+This includes LinuxKPI managed-device mappings marked as IO or PFN mappings.
+The residency and address-translation guarantees above do not apply to these
+mappings; their driver and pager control their backing lifetime.
+Ordinary mappings retain their existing locking behavior.
+Existing privilege, address-range, protection and resource-limit checks still
+apply, including conservative accounting of the requested virtual range.
+Thus a request containing an exempt mapping can still fail those checks.
+.Pp
Since physical memory is a potentially scarce resource, processes are
limited in how much they can lock down.
The amount of memory that a single process can
.Fn mlock
is limited by both the per-process
.Dv RLIMIT_MEMLOCK
resource limit and the
system-wide
.Dq wired pages
limit
.Va vm.max_user_wired .
.Va vm.max_user_wired
applies to the system as a whole, so the amount available to a single
process at any given time is the difference between
.Va vm.max_user_wired
and
.Va vm.stats.vm.v_user_wire_count .
.Pp
If
.Va security.bsd.unprivileged_mlock
is set to 0 these calls are only available to the super-user.
.Sh RETURN VALUES
.Rv -std
.Pp
-If the call succeeds, all pages in the range become locked (unlocked);
+If the call succeeds, all non-exempt pages in the range become locked
+(unlocked);
otherwise the locked status of all pages in the range remains unchanged.
.Sh ERRORS
The
.Fn mlock
system call
will fail if:
.Bl -tag -width Er
.It Bq Er EPERM
.Va security.bsd.unprivileged_mlock
is set to 0 and the caller is not the super-user.
.It Bq Er EINVAL
The address range given wraps around zero.
.It Bq Er ENOMEM
Some portion of the indicated address range is not allocated.
There was an error faulting/mapping a page.
Locking the indicated range would exceed the per-process or system-wide limits
for locked memory.
.El
The
.Fn munlock
system call
will fail if:
.Bl -tag -width Er
.It Bq Er EPERM
.Va security.bsd.unprivileged_mlock
is set to 0 and the caller is not the super-user.
.It Bq Er EINVAL
The address range given wraps around zero.
.It Bq Er ENOMEM
Some or all of the address range specified by the addr and len
arguments does not correspond to valid mapped pages in the address space
of the process.
.It Bq Er ENOMEM
Locking the pages mapped by the specified range would exceed a limit on
the amount of memory that the process may lock.
.El
.Sh "SEE ALSO"
.Xr fork 2 ,
.Xr mincore 2 ,
.Xr minherit 2 ,
.Xr mlockall 2 ,
.Xr mmap 2 ,
.Xr munlockall 2 ,
.Xr munmap 2 ,
.Xr setrlimit 2 ,
.Xr getpagesize 3
.Sh HISTORY
The
.Fn mlock
and
.Fn munlock
system calls first appeared in
.Bx 4.4 .
.Sh BUGS
Allocating too much wired memory can lead to a memory-allocation deadlock
which requires a reboot to recover from.
.Pp
The per-process and system-wide resource limits of locked memory apply
to the amount of virtual memory locked, not the amount of locked physical
pages.
Hence two distinct locked mappings of the same physical page counts as
2 pages aginst the system limit, and also against the per-process limit
if both mappings belong to the same physical map.
-.Pp
-The per-process resource limit is not currently supported.
diff --git a/lib/libsys/mlockall.2 b/lib/libsys/mlockall.2
index 58d3a9b0d78a0e99830f50b368c7c74ee61922d3..74d86859007777baaedb3e61507452ccc17ebe1d 100644
--- a/lib/libsys/mlockall.2
+++ b/lib/libsys/mlockall.2
@@ -1,144 +1,158 @@
.\" $NetBSD: mlockall.2,v 1.11 2003/04/16 13:34:54 wiz Exp $
.\"
.\" Copyright (c) 1999 The NetBSD Foundation, Inc.
.\" All rights reserved.
.\"
.\" This code is derived from software contributed to The NetBSD Foundation
.\" by Jason R. Thorpe of the Numerical Aerospace Simulation Facility,
.\" NASA Ames Research Center.
.\"
.\" Redistribution and use in source and binary forms, with or without
.\" modification, are permitted provided that the following conditions
.\" are met:
.\" 1. Redistributions of source code must retain the above copyright
.\" notice, this list of conditions and the following disclaimer.
.\" 2. Redistributions in binary form must reproduce the above copyright
.\" notice, this list of conditions and the following disclaimer in the
.\" documentation and/or other materials provided with the distribution.
.\"
.\" THIS SOFTWARE IS PROVIDED BY THE NETBSD FOUNDATION, INC. AND CONTRIBUTORS
.\" ``AS IS'' AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED
.\" TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR
.\" PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE FOUNDATION OR CONTRIBUTORS
.\" BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR
.\" CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF
.\" SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS
.\" INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN
.\" CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE)
.\" ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE
.\" POSSIBILITY OF SUCH DAMAGE.
.\"
.Dd May 13, 2019
.Dt MLOCKALL 2
.Os
.Sh NAME
.Nm mlockall ,
.Nm munlockall
.Nd lock (unlock) the address space of a process
.Sh LIBRARY
.Lb libc
.Sh SYNOPSIS
.In sys/mman.h
.Ft int
.Fn mlockall "int flags"
.Ft int
.Fn munlockall "void"
.Sh DESCRIPTION
The
.Fn mlockall
system call locks into memory the physical pages associated with the
address space of a process until the address space is unlocked, the
process exits, or execs another program image.
.Pp
The following flags affect the behavior of
.Fn mlockall :
.Bl -tag -width ".Dv MCL_CURRENT"
.It Dv MCL_CURRENT
Lock all pages currently mapped into the process's address space.
.It Dv MCL_FUTURE
Lock all pages mapped into the process's address space in the future,
at the time the mapping is established.
Note that this may cause future mappings to fail if those mappings
cause resource limits to be exceeded.
.El
.Pp
+Special device mappings explicitly exempted by their pager are skipped for
+both
+.Dv MCL_CURRENT
+and
+.Dv MCL_FUTURE .
+This includes LinuxKPI managed-device mappings marked as IO or PFN mappings.
+These mappings acquire no user wire reference and are not covered by the
+residency guarantee; their driver and pager control their backing lifetime.
+Existing resource-limit checks and conservative accounting of the requested
+virtual range still apply.
+See
+.Xr mlock 2 .
+.Pp
Since physical memory is a potentially scarce resource, processes are
limited in how much they can lock down.
A single process can lock the minimum of a system-wide
.Dq wired pages
limit
.Va vm.max_user_wired
and the per-process
.Dv RLIMIT_MEMLOCK
resource limit.
.Pp
If
.Va security.bsd.unprivileged_mlock
is set to 0 these calls are only available to the super-user.
If
.Va vm.old_mlock
is set to 1 the per-process
.Dv RLIMIT_MEMLOCK
resource limit will not be applied for
.Fn mlockall
calls.
.Pp
The
.Fn munlockall
call unlocks any locked memory regions in the process address space.
Any regions mapped after an
.Fn munlockall
call will not be locked.
.Sh RETURN VALUES
A return value of 0 indicates that the call
-succeeded and all pages in the range have either been locked or unlocked.
+succeeded and all non-exempt pages in the range have either been locked or
+unlocked.
A return value of \-1 indicates an error occurred and the locked
status of all pages in the range remains unchanged.
In this case, the global location
.Va errno
is set to indicate the error.
.Sh ERRORS
.Fn mlockall
will fail if:
.Bl -tag -width Er
.It Bq Er EINVAL
The
.Fa flags
argument is zero, or includes unimplemented flags.
.It Bq Er ENOMEM
Locking the indicated range would exceed either the system or per-process
limit for locked memory.
.It Bq Er EAGAIN
Some or all of the memory mapped into the process's address space
could not be locked when the call was made.
.It Bq Er EPERM
The calling process does not have the appropriate privilege to perform
the requested operation.
.El
.Sh SEE ALSO
.Xr mincore 2 ,
.Xr mlock 2 ,
.Xr mmap 2 ,
.Xr munmap 2 ,
.Xr setrlimit 2
.Sh STANDARDS
The
.Fn mlockall
and
.Fn munlockall
functions are believed to conform to
.St -p1003.1-2001 .
.Sh HISTORY
The
.Fn mlockall
and
.Fn munlockall
functions first appeared in
.Fx 5.1 .
.Sh BUGS
The per-process and system-wide resource limits of locked memory apply
to the amount of virtual memory locked, not the amount of locked physical
pages.
Hence two distinct locked mappings of the same physical page counts as
2 pages aginst the system limit, and also against the per-process limit
if both mappings belong to the same physical map.
diff --git a/sys/compat/linuxkpi/common/include/linux/mm.h b/sys/compat/linuxkpi/common/include/linux/mm.h
index 1732e21de5cfa810e2809e299613baf0bac8df0c..1611c4fdc396ccb9f898d6a742f418a1dd9537ca 100644
--- a/sys/compat/linuxkpi/common/include/linux/mm.h
+++ b/sys/compat/linuxkpi/common/include/linux/mm.h
@@ -1,487 +1,509 @@
/*-
* Copyright (c) 2010 Isilon Systems, Inc.
* Copyright (c) 2010 iX Systems, Inc.
* Copyright (c) 2010 Panasas, Inc.
* Copyright (c) 2013-2017 Mellanox Technologies, Ltd.
* Copyright (c) 2015 François Tigeot
* Copyright (c) 2015 Matthew Dillon <dillon@backplane.com>
* All rights reserved.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice unmodified, this list of conditions, and the following
* disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
*
* THIS SOFTWARE IS PROVIDED BY THE AUTHOR ``AS IS'' AND ANY EXPRESS OR
* IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES
* OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED.
* IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR ANY DIRECT, INDIRECT,
* INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT
* NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE,
* DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY
* THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
* (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF
* THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
*/
#ifndef _LINUXKPI_LINUX_MM_H_
#define _LINUXKPI_LINUX_MM_H_
#include <linux/spinlock.h>
#include <linux/gfp.h>
#include <linux/kernel.h>
#include <linux/mm_types.h>
#include <linux/mmzone.h>
#include <linux/pfn.h>
#include <linux/list.h>
#include <linux/mmap_lock.h>
#include <linux/overflow.h>
#include <linux/shrinker.h>
#include <linux/page.h>
#include <linux/page-flags.h>
#include <asm/pgtable.h>
#define PAGE_ALIGN(x) ALIGN(x, PAGE_SIZE)
/*
* Make sure our LinuxKPI defined virtual memory flags don't conflict
* with the ones defined by FreeBSD:
*/
CTASSERT((VM_PROT_ALL & -(1 << 8)) == 0);
#define VM_READ VM_PROT_READ
#define VM_WRITE VM_PROT_WRITE
#define VM_EXEC VM_PROT_EXECUTE
#define VM_ACCESS_FLAGS (VM_READ | VM_WRITE | VM_EXEC)
#define VM_PFNINTERNAL (1 << 8) /* FreeBSD private flag to vm_insert_pfn() */
#define VM_MIXEDMAP (1 << 9)
#define VM_NORESERVE (1 << 10)
#define VM_PFNMAP (1 << 11)
#define VM_IO (1 << 12)
#define VM_MAYWRITE (1 << 13)
#define VM_DONTCOPY (1 << 14)
#define VM_DONTEXPAND (1 << 15)
#define VM_DONTDUMP (1 << 16)
#define VM_SHARED (1 << 17)
#define VMA_MAX_PREFAULT_RECORD 1
#define FOLL_WRITE (1 << 0)
#define FOLL_FORCE (1 << 1)
#define VM_FAULT_OOM (1 << 0)
#define VM_FAULT_SIGBUS (1 << 1)
#define VM_FAULT_MAJOR (1 << 2)
#define VM_FAULT_WRITE (1 << 3)
#define VM_FAULT_HWPOISON (1 << 4)
#define VM_FAULT_HWPOISON_LARGE (1 << 5)
#define VM_FAULT_SIGSEGV (1 << 6)
#define VM_FAULT_NOPAGE (1 << 7)
#define VM_FAULT_LOCKED (1 << 8)
#define VM_FAULT_RETRY (1 << 9)
#define VM_FAULT_FALLBACK (1 << 10)
#define VM_FAULT_ERROR (VM_FAULT_OOM | VM_FAULT_SIGBUS | VM_FAULT_SIGSEGV | \
VM_FAULT_HWPOISON |VM_FAULT_HWPOISON_LARGE | VM_FAULT_FALLBACK)
#define FAULT_FLAG_WRITE (1 << 0)
#define FAULT_FLAG_MKWRITE (1 << 1)
#define FAULT_FLAG_ALLOW_RETRY (1 << 2)
#define FAULT_FLAG_RETRY_NOWAIT (1 << 3)
#define FAULT_FLAG_KILLABLE (1 << 4)
#define FAULT_FLAG_TRIED (1 << 5)
#define FAULT_FLAG_USER (1 << 6)
#define FAULT_FLAG_REMOTE (1 << 7)
#define FAULT_FLAG_INSTRUCTION (1 << 8)
#define fault_flag_allow_retry_first(flags) \
(((flags) & (FAULT_FLAG_ALLOW_RETRY | FAULT_FLAG_TRIED)) == FAULT_FLAG_ALLOW_RETRY)
typedef int (*pte_fn_t)(linux_pte_t *, unsigned long addr, void *data);
struct vm_area_struct {
vm_offset_t vm_start;
vm_offset_t vm_end;
vm_offset_t vm_pgoff;
pgprot_t vm_page_prot;
unsigned long vm_flags;
struct mm_struct *vm_mm;
void *vm_private_data;
const struct vm_operations_struct *vm_ops;
struct linux_file *vm_file;
/* internal operation */
vm_paddr_t vm_pfn; /* PFN for memory map */
vm_size_t vm_len; /* length for memory map */
vm_pindex_t vm_pfn_first;
int vm_pfn_count;
int *vm_pfn_pcount;
vm_object_t vm_obj;
vm_map_t vm_cached_map;
+ void *vm_pfn_state;
TAILQ_ENTRY(vm_area_struct) vm_entry;
};
struct vm_fault {
unsigned int flags;
pgoff_t pgoff;
union {
/* user-space address */
void *virtual_address; /* < 4.11 */
unsigned long address; /* >= 4.11 */
};
struct page *page;
struct vm_area_struct *vma;
};
+int lkpi_vma_pfn_init(struct vm_area_struct *vma);
+int lkpi_vma_pfn_begin(struct vm_area_struct *vma, uint64_t *invalidation_seq);
+int lkpi_vma_pfn_error(struct vm_area_struct *vma);
+bool lkpi_vma_pfn_unchanged(struct vm_area_struct *vma,
+ uint64_t invalidation_seq);
+bool lkpi_vma_pfn_handoff_valid(struct vm_area_struct *vma);
+bool lkpi_vma_pfn_lock(struct vm_area_struct *vma);
+void lkpi_vma_pfn_end(struct vm_area_struct *vma);
+vm_page_t lkpi_vma_pfn_take_page(struct vm_area_struct *vma,
+ vm_object_t object, vm_pindex_t pindex);
+void lkpi_vma_pfn_release_page(struct vm_area_struct *vma, vm_page_t page);
+void lkpi_vma_pfn_abort(struct vm_area_struct *vma, vm_object_t object);
+void lkpi_vma_pfn_done(struct vm_area_struct *vma, vm_object_t object);
+bool lkpi_vma_pfn_invalidate_begin(struct vm_area_struct *vma);
+bool lkpi_vma_pfn_unmap_begin(struct vm_area_struct *vma);
+void lkpi_vma_pfn_unmap(struct vm_area_struct *vma);
+void lkpi_vma_pfn_unmap_end(struct vm_area_struct *vma);
+void lkpi_vma_pfn_invalidate_end(struct vm_area_struct *vma);
+void lkpi_vma_pfn_fini(struct vm_area_struct *vma);
+void linux_cdev_pager_free_pages(vm_object_t object);
+
struct vm_operations_struct {
void (*open) (struct vm_area_struct *);
void (*close) (struct vm_area_struct *);
int (*fault) (struct vm_fault *);
int (*access) (struct vm_area_struct *, unsigned long, void *, int, int);
};
struct sysinfo {
uint64_t totalram; /* Total usable main memory size */
uint64_t freeram; /* Available memory size */
uint64_t totalhigh; /* Total high memory size */
uint64_t freehigh; /* Available high memory size */
uint32_t mem_unit; /* Memory unit size in bytes */
};
static inline struct page *
virt_to_head_page(const void *p)
{
return (virt_to_page(p));
}
static inline struct folio *
virt_to_folio(const void *p)
{
struct page *page = virt_to_page(p);
return (page_folio(page));
}
/*
* Compute log2 of the power of two rounded up count of pages
* needed for size bytes.
*/
static inline int
get_order(unsigned long size)
{
int order;
size = (size - 1) >> PAGE_SHIFT;
order = 0;
while (size) {
order++;
size >>= 1;
}
return (order);
}
/*
* Resolve a page into a virtual address:
*
* NOTE: This function only works for pages allocated by the kernel.
*/
void *linux_page_address(const struct page *);
#define page_address(page) linux_page_address(page)
static inline void *
lowmem_page_address(struct page *page)
{
return (page_address(page));
}
/*
* This only works via memory map operations.
*/
static inline int
io_remap_pfn_range(struct vm_area_struct *vma,
unsigned long addr, unsigned long pfn, unsigned long size,
vm_memattr_t prot)
{
vma->vm_page_prot = prot;
vma->vm_pfn = pfn;
vma->vm_len = size;
return (0);
}
vm_fault_t
lkpi_vmf_insert_pfn_prot_locked(struct vm_area_struct *vma, unsigned long addr,
unsigned long pfn, pgprot_t prot);
static inline vm_fault_t
vmf_insert_pfn_prot(struct vm_area_struct *vma, unsigned long addr,
unsigned long pfn, pgprot_t prot)
{
vm_fault_t ret;
VM_OBJECT_WLOCK(vma->vm_obj);
ret = lkpi_vmf_insert_pfn_prot_locked(vma, addr, pfn, prot);
VM_OBJECT_WUNLOCK(vma->vm_obj);
return (ret);
}
#define vmf_insert_pfn_prot(...) \
_Static_assert(false, \
"This function is always called in a loop. Consider using the locked version")
static inline int
apply_to_page_range(struct mm_struct *mm, unsigned long address,
unsigned long size, pte_fn_t fn, void *data)
{
return (-ENOTSUP);
}
int zap_vma_ptes(struct vm_area_struct *vma, unsigned long address,
unsigned long size);
int lkpi_remap_pfn_range(struct vm_area_struct *vma,
unsigned long start_addr, unsigned long start_pfn, unsigned long size,
pgprot_t prot);
static inline int
remap_pfn_range(struct vm_area_struct *vma, unsigned long addr,
unsigned long pfn, unsigned long size, pgprot_t prot)
{
return (lkpi_remap_pfn_range(vma, addr, pfn, size, prot));
}
static inline unsigned long
vma_pages(struct vm_area_struct *vma)
{
return ((vma->vm_end - vma->vm_start) >> PAGE_SHIFT);
}
#define offset_in_page(off) ((unsigned long)(off) & (PAGE_SIZE - 1))
#define offset_in_folio(folio, p) ((unsigned long)(p) & (folio_size(folio) - 1))
static inline void
set_page_dirty(struct page *page)
{
vm_page_dirty(page);
}
static inline void
mark_page_accessed(struct page *page)
{
vm_page_reference(page);
}
static inline void
get_page(struct page *page)
{
vm_page_wire(page);
}
static inline void
put_page(struct page *page)
{
/* `__free_page()` takes care of the refcounting (unwire). */
__free_page(page);
}
static inline void
folio_get(struct folio *folio)
{
get_page(&folio->page);
}
static inline void
folio_put(struct folio *folio)
{
put_page(&folio->page);
}
/*
* Linux uses the following "transparent" union so that `release_pages()`
* accepts both a list of `struct page` or a list of `struct folio`. This
* relies on the fact that a `struct folio` can be cast to a `struct page`.
*/
typedef union {
struct page **pages;
struct folio **folios;
} release_pages_arg __attribute__ ((__transparent_union__));
void linux_release_pages(release_pages_arg arg, int nr);
#define release_pages(arg, nr) linux_release_pages((arg), (nr))
extern long
lkpi_get_user_pages(unsigned long start, unsigned long nr_pages,
unsigned int gup_flags, struct page **);
#if defined(LINUXKPI_VERSION) && LINUXKPI_VERSION >= 60500
#define get_user_pages(start, nr_pages, gup_flags, pages) \
lkpi_get_user_pages(start, nr_pages, gup_flags, pages)
#else
#define get_user_pages(start, nr_pages, gup_flags, pages, vmas) \
lkpi_get_user_pages(start, nr_pages, gup_flags, pages)
#endif
#if defined(LINUXKPI_VERSION) && LINUXKPI_VERSION >= 60500
static inline long
pin_user_pages(unsigned long start, unsigned long nr_pages,
unsigned int gup_flags, struct page **pages)
{
return (get_user_pages(start, nr_pages, gup_flags, pages));
}
#else
static inline long
pin_user_pages(unsigned long start, unsigned long nr_pages,
unsigned int gup_flags, struct page **pages,
struct vm_area_struct **vmas)
{
return (get_user_pages(start, nr_pages, gup_flags, pages, vmas));
}
#endif
extern int
__get_user_pages_fast(unsigned long start, int nr_pages, int write,
struct page **);
static inline int
pin_user_pages_fast(unsigned long start, int nr_pages,
unsigned int gup_flags, struct page **pages)
{
return __get_user_pages_fast(
start, nr_pages, !!(gup_flags & FOLL_WRITE), pages);
}
extern long
get_user_pages_remote(struct task_struct *, struct mm_struct *,
unsigned long start, unsigned long nr_pages,
unsigned int gup_flags, struct page **,
struct vm_area_struct **);
static inline long
pin_user_pages_remote(struct task_struct *task, struct mm_struct *mm,
unsigned long start, unsigned long nr_pages,
unsigned int gup_flags, struct page **pages,
struct vm_area_struct **vmas)
{
return get_user_pages_remote(
task, mm, start, nr_pages, gup_flags, pages, vmas);
}
#define unpin_user_page(page) put_page(page)
#define unpin_user_pages(pages, npages) release_pages(pages, npages)
#define copy_highpage(to, from) pmap_copy_page(from, to)
static inline pgprot_t
vm_get_page_prot(unsigned long vm_flags)
{
return (vm_flags & VM_PROT_ALL);
}
static inline void
vm_flags_set(struct vm_area_struct *vma, unsigned long flags)
{
vma->vm_flags |= flags;
}
static inline void
vm_flags_clear(struct vm_area_struct *vma, unsigned long flags)
{
vma->vm_flags &= ~flags;
}
static inline struct page *
vmalloc_to_page(const void *addr)
{
vm_paddr_t paddr;
paddr = pmap_kextract((vm_offset_t)addr);
return (PHYS_TO_VM_PAGE(paddr));
}
static inline int
trylock_page(struct page *page)
{
return (vm_page_tryxbusy(page));
}
static inline void
unlock_page(struct page *page)
{
vm_page_xunbusy(page);
}
static inline void
split_page(struct page *page, unsigned int order)
{
pr_debug("%s: TODO\n", __func__);
}
extern int is_vmalloc_addr(const void *addr);
void si_meminfo(struct sysinfo *si);
static inline unsigned long
totalram_pages(void)
{
return ((unsigned long)physmem);
}
#define unmap_mapping_range(...) lkpi_unmap_mapping_range(__VA_ARGS__)
void lkpi_unmap_mapping_range(void *obj, loff_t const holebegin __unused,
loff_t const holelen, int even_cows __unused);
#define PAGE_ALIGNED(p) __is_aligned(p, PAGE_SIZE)
void vma_set_file(struct vm_area_struct *vma, struct linux_file *file);
static inline void
might_alloc(gfp_t gfp_mask __unused)
{
}
#define is_cow_mapping(flags) (false)
static inline bool
want_init_on_free(void)
{
return (false);
}
static inline unsigned long
folio_pfn(struct folio *folio)
{
return (page_to_pfn(&folio->page));
}
static inline long
folio_nr_pages(struct folio *folio)
{
return (1);
}
static inline size_t
folio_size(struct folio *folio)
{
return (PAGE_SIZE);
}
static inline void
folio_mark_dirty(struct folio *folio)
{
set_page_dirty(&folio->page);
}
static inline void *
folio_address(const struct folio *folio)
{
return (page_address(&folio->page));
}
#endif /* _LINUXKPI_LINUX_MM_H_ */
diff --git a/sys/compat/linuxkpi/common/src/linux_compat.c b/sys/compat/linuxkpi/common/src/linux_compat.c
index 24655eb0341d0545716f53cae38baaade2284b78..213260d2c74867e9a16a394df9000fa9d99d273e 100644
--- a/sys/compat/linuxkpi/common/src/linux_compat.c
+++ b/sys/compat/linuxkpi/common/src/linux_compat.c
@@ -1,3132 +1,3290 @@
/*-
* Copyright (c) 2010 Isilon Systems, Inc.
* Copyright (c) 2010 iX Systems, Inc.
* Copyright (c) 2010 Panasas, Inc.
* Copyright (c) 2013-2021 Mellanox Technologies, Ltd.
* All rights reserved.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice unmodified, this list of conditions, and the following
* disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
*
* THIS SOFTWARE IS PROVIDED BY THE AUTHOR ``AS IS'' AND ANY EXPRESS OR
* IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES
* OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED.
* IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR ANY DIRECT, INDIRECT,
* INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT
* NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE,
* DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY
* THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
* (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF
* THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
*/
#include <sys/cdefs.h>
#include "opt_global.h"
#include "opt_stack.h"
#include <sys/param.h>
#include <sys/systm.h>
#include <sys/malloc.h>
#include <sys/kernel.h>
#include <sys/sysctl.h>
#include <sys/proc.h>
#include <sys/sglist.h>
#include <sys/sleepqueue.h>
#include <sys/refcount.h>
#include <sys/lock.h>
#include <sys/mutex.h>
#include <sys/bus.h>
#include <sys/eventhandler.h>
#include <sys/fcntl.h>
#include <sys/file.h>
#include <sys/filio.h>
#include <sys/rwlock.h>
#include <sys/mman.h>
#include <sys/stack.h>
#include <sys/stdarg.h>
#include <sys/syscall.h>
#include <sys/sysent.h>
#include <sys/time.h>
#include <sys/user.h>
#include <vm/vm.h>
#include <vm/pmap.h>
#include <vm/vm_object.h>
#include <vm/vm_page.h>
#include <vm/vm_pager.h>
#include <vm/vm_radix.h>
#if defined(__i386__) || defined(__amd64__)
#include <machine/cputypes.h>
#include <machine/md_var.h>
#endif
#include <linux/kobject.h>
#include <linux/cpu.h>
#include <linux/device.h>
#include <linux/slab.h>
#include <linux/module.h>
#include <linux/moduleparam.h>
#include <linux/cdev.h>
#include <linux/file.h>
#include <linux/fs.h>
#include <linux/sysfs.h>
#include <linux/mm.h>
#include <linux/io.h>
#include <linux/vmalloc.h>
#include <linux/netdevice.h>
#include <linux/timer.h>
#include <linux/interrupt.h>
#include <linux/uaccess.h>
#include <linux/utsname.h>
#include <linux/list.h>
#include <linux/kthread.h>
#include <linux/kernel.h>
#include <linux/compat.h>
#include <linux/io-mapping.h>
#include <linux/poll.h>
#include <linux/smp.h>
#include <linux/wait_bit.h>
#include <linux/rcupdate.h>
#include <linux/interval_tree.h>
#include <linux/interval_tree_generic.h>
#include <linux/printk.h>
#include <linux/seq_file.h>
#include <linux/uuid.h>
#include <linux/mod_devicetable.h>
#if defined(__i386__) || defined(__amd64__)
#include <asm/cpu_device_id.h>
#include <asm/cpufeature.h>
#include <asm/smp.h>
#include <asm/processor.h>
#endif
#include <xen/xen.h>
#ifdef XENHVM
#undef xen_pv_domain
#undef xen_initial_domain
/* xen/xen-os.h redefines __must_check */
#undef __must_check
#include <xen/xen-os.h>
#endif
SYSCTL_NODE(_compat, OID_AUTO, linuxkpi, CTLFLAG_RW | CTLFLAG_MPSAFE, 0,
"LinuxKPI parameters");
int linuxkpi_debug;
SYSCTL_INT(_compat_linuxkpi, OID_AUTO, debug, CTLFLAG_RWTUN,
&linuxkpi_debug, 0, "Set to enable pr_debug() prints. Clear to disable.");
int linuxkpi_rcu_debug;
SYSCTL_INT(_compat_linuxkpi, OID_AUTO, rcu_debug, CTLFLAG_RWTUN,
&linuxkpi_rcu_debug, 0, "Set to enable RCU warning. Clear to disable.");
int linuxkpi_warn_dump_stack = 0;
SYSCTL_INT(_compat_linuxkpi, OID_AUTO, warn_dump_stack, CTLFLAG_RWTUN,
&linuxkpi_warn_dump_stack, 0,
"Set to enable stack traces from WARN_ON(). Clear to disable.");
static struct timeval lkpi_net_lastlog;
static int lkpi_net_curpps;
static int lkpi_net_maxpps = 99;
SYSCTL_INT(_compat_linuxkpi, OID_AUTO, net_ratelimit, CTLFLAG_RWTUN,
&lkpi_net_maxpps, 0, "Limit number of LinuxKPI net messages per second.");
MALLOC_DEFINE(M_KMALLOC, "lkpikmalloc", "Linux kmalloc compat");
#include <linux/rbtree.h>
/* Undo Linux compat changes. */
#undef RB_ROOT
#undef file
#undef cdev
#define RB_ROOT(head) (head)->rbh_root
static void linux_destroy_dev(struct linux_cdev *);
static void linux_cdev_deref(struct linux_cdev *ldev);
static struct vm_area_struct *linux_cdev_handle_find(void *handle);
cpumask_t cpu_online_mask;
static cpumask_t **static_single_cpu_mask;
static cpumask_t *static_single_cpu_mask_lcs;
struct kobject linux_class_root;
struct device linux_root_device;
struct class linux_class_misc;
struct list_head pci_drivers;
struct list_head pci_devices;
spinlock_t pci_lock;
struct uts_namespace init_uts_ns;
unsigned long linux_timer_hz_mask;
wait_queue_head_t linux_bit_waitq;
wait_queue_head_t linux_var_waitq;
const guid_t guid_null;
enum system_states system_state = SYSTEM_RUNNING;
struct task_struct *
__lkpi_current(void)
{
struct thread *td;
td = curthread;
linux_set_current(td);
return ((struct task_struct *)td->td_lkpi_task);
}
int
panic_cmp(struct rb_node *one, struct rb_node *two)
{
panic("no cmp");
}
RB_GENERATE(linux_root, rb_node, __entry, panic_cmp);
#define START(node) ((node)->start)
#define LAST(node) ((node)->last)
INTERVAL_TREE_DEFINE(struct interval_tree_node, rb, unsigned long,, START,
LAST,, lkpi_interval_tree)
static void
linux_device_release(struct device *dev)
{
pr_debug("linux_device_release: %s\n", dev_name(dev));
kfree(dev);
}
static ssize_t
linux_class_show(struct kobject *kobj, struct attribute *attr, char *buf)
{
struct class_attribute *dattr;
ssize_t error;
dattr = container_of(attr, struct class_attribute, attr);
error = -EIO;
if (dattr->show)
error = dattr->show(container_of(kobj, struct class, kobj),
dattr, buf);
return (error);
}
static ssize_t
linux_class_store(struct kobject *kobj, struct attribute *attr, const char *buf,
size_t count)
{
struct class_attribute *dattr;
ssize_t error;
dattr = container_of(attr, struct class_attribute, attr);
error = -EIO;
if (dattr->store)
error = dattr->store(container_of(kobj, struct class, kobj),
dattr, buf, count);
return (error);
}
static void
linux_class_release(struct kobject *kobj)
{
struct class *class;
class = container_of(kobj, struct class, kobj);
if (class->class_release)
class->class_release(class);
}
static const struct sysfs_ops linux_class_sysfs = {
.show = linux_class_show,
.store = linux_class_store,
};
const struct kobj_type linux_class_ktype = {
.release = linux_class_release,
.sysfs_ops = &linux_class_sysfs
};
static void
linux_dev_release(struct kobject *kobj)
{
struct device *dev;
dev = container_of(kobj, struct device, kobj);
/* This is the precedence defined by linux. */
if (dev->release)
dev->release(dev);
else if (dev->class && dev->class->dev_release)
dev->class->dev_release(dev);
}
static ssize_t
linux_dev_show(struct kobject *kobj, struct attribute *attr, char *buf)
{
struct device_attribute *dattr;
ssize_t error;
dattr = container_of(attr, struct device_attribute, attr);
error = -EIO;
if (dattr->show)
error = dattr->show(container_of(kobj, struct device, kobj),
dattr, buf);
return (error);
}
static ssize_t
linux_dev_store(struct kobject *kobj, struct attribute *attr, const char *buf,
size_t count)
{
struct device_attribute *dattr;
ssize_t error;
dattr = container_of(attr, struct device_attribute, attr);
error = -EIO;
if (dattr->store)
error = dattr->store(container_of(kobj, struct device, kobj),
dattr, buf, count);
return (error);
}
static const struct sysfs_ops linux_dev_sysfs = {
.show = linux_dev_show,
.store = linux_dev_store,
};
const struct kobj_type linux_dev_ktype = {
.release = linux_dev_release,
.sysfs_ops = &linux_dev_sysfs
};
struct device *
device_create(struct class *class, struct device *parent, dev_t devt,
void *drvdata, const char *fmt, ...)
{
struct device *dev;
va_list args;
dev = kzalloc(sizeof(*dev), M_WAITOK);
dev->parent = parent;
dev->class = class;
dev->devt = devt;
dev->driver_data = drvdata;
dev->release = linux_device_release;
va_start(args, fmt);
kobject_set_name_vargs(&dev->kobj, fmt, args);
va_end(args);
device_register(dev);
return (dev);
}
struct device *
device_create_groups_vargs(struct class *class, struct device *parent,
dev_t devt, void *drvdata, const struct attribute_group **groups,
const char *fmt, va_list args)
{
struct device *dev = NULL;
int retval = -ENODEV;
if (class == NULL || IS_ERR(class))
goto error;
dev = kzalloc(sizeof(*dev), GFP_KERNEL);
if (!dev) {
retval = -ENOMEM;
goto error;
}
dev->devt = devt;
dev->class = class;
dev->parent = parent;
dev->groups = groups;
dev->release = device_create_release;
/* device_initialize() needs the class and parent to be set */
device_initialize(dev);
dev_set_drvdata(dev, drvdata);
retval = kobject_set_name_vargs(&dev->kobj, fmt, args);
if (retval)
goto error;
retval = device_add(dev);
if (retval)
goto error;
return dev;
error:
put_device(dev);
return ERR_PTR(retval);
}
struct class *
lkpi_class_create(const char *name)
{
struct class *class;
int error;
class = kzalloc(sizeof(*class), M_WAITOK);
class->name = name;
class->class_release = linux_class_kfree;
error = class_register(class);
if (error) {
kfree(class);
return (NULL);
}
return (class);
}
static void
linux_kq_lock(void *arg)
{
spinlock_t *s = arg;
spin_lock(s);
}
static void
linux_kq_unlock(void *arg)
{
spinlock_t *s = arg;
spin_unlock(s);
}
static void
linux_kq_assert_lock(void *arg, int what)
{
#ifdef INVARIANTS
spinlock_t *s = arg;
if (what == LA_LOCKED)
mtx_assert(s, MA_OWNED);
else
mtx_assert(s, MA_NOTOWNED);
#endif
}
static void
linux_file_kqfilter_poll(struct linux_file *, int);
struct linux_file *
linux_file_alloc(void)
{
struct linux_file *filp;
filp = kzalloc(sizeof(*filp), GFP_KERNEL);
/* set initial refcount */
filp->f_count = 1;
/* setup fields needed by kqueue support */
spin_lock_init(&filp->f_kqlock);
knlist_init(&filp->f_selinfo.si_note, &filp->f_kqlock,
linux_kq_lock, linux_kq_unlock, linux_kq_assert_lock);
return (filp);
}
void
linux_file_free(struct linux_file *filp)
{
if (filp->_file == NULL) {
if (filp->f_op != NULL && filp->f_op->release != NULL)
filp->f_op->release(filp->f_vnode, filp);
if (filp->f_shmem != NULL)
vm_object_deallocate(filp->f_shmem);
kfree_rcu(filp, rcu);
} else {
/*
* The close method of the character device or file
* will free the linux_file structure:
*/
_fdrop(filp->_file, curthread);
}
}
struct linux_cdev *
cdev_alloc(void)
{
struct linux_cdev *cdev;
cdev = kzalloc(sizeof(struct linux_cdev), M_WAITOK);
kobject_init(&cdev->kobj, &linux_cdev_ktype);
cdev->refs = 1;
return (cdev);
}
static int
linux_cdev_pager_fault(vm_object_t vm_obj, vm_ooffset_t offset, int prot,
vm_page_t *mres)
{
struct vm_area_struct *vmap;
vmap = linux_cdev_handle_find(vm_obj->handle);
MPASS(vmap != NULL);
MPASS(vmap->vm_private_data == vm_obj->handle);
if (likely(vmap->vm_ops != NULL && offset < vmap->vm_len)) {
vm_paddr_t paddr = IDX_TO_OFF(vmap->vm_pfn) + offset;
vm_page_t page;
if (((*mres)->flags & PG_FICTITIOUS) != 0) {
/*
* If the passed in result page is a fake
* page, update it with the new physical
* address.
*/
page = *mres;
vm_page_updatefake(page, paddr, vm_obj->memattr);
} else {
/*
* Replace the passed in "mres" page with our
* own fake page and free up the all of the
* original pages.
*/
VM_OBJECT_WUNLOCK(vm_obj);
page = vm_page_getfake(paddr, vm_obj->memattr);
VM_OBJECT_WLOCK(vm_obj);
vm_page_replace(page, vm_obj, (*mres)->pindex, *mres);
*mres = page;
}
vm_page_valid(page);
return (VM_PAGER_OK);
}
return (VM_PAGER_FAIL);
}
static int
linux_cdev_pager_populate(vm_object_t vm_obj, vm_pindex_t pidx, int fault_type,
vm_prot_t max_prot, vm_pindex_t *first, vm_pindex_t *last)
{
struct vm_area_struct *vmap;
- int err;
+ bool pfn_started;
+ bool retry;
+ uint64_t invalidation_seq;
+ int err, pfn_error;
/* get VM area structure */
vmap = linux_cdev_handle_find(vm_obj->handle);
MPASS(vmap != NULL);
MPASS(vmap->vm_private_data == vm_obj->handle);
VM_OBJECT_WUNLOCK(vm_obj);
-
linux_set_current(curthread);
-
down_write(&vmap->vm_mm->mmap_sem);
- if (unlikely(vmap->vm_ops == NULL)) {
- err = VM_FAULT_SIGBUS;
+ pfn_started = false;
+ retry = false;
+ err = lkpi_vma_pfn_begin(vmap, &invalidation_seq);
+ if (err != 0) {
+ err = err == EAGAIN ? VM_FAULT_RETRY :
+ err == ENOMEM ? VM_FAULT_OOM : VM_FAULT_SIGBUS;
} else {
+ pfn_started = true;
+ /* The previous populate may still be completing its handoff. */
+ atomic_store_rel_ptr((volatile uintptr_t *)&vmap->vm_obj,
+ (uintptr_t)vm_obj);
+ vmap->vm_pfn_count = 0;
+ vmap->vm_pfn_pcount = &vmap->vm_pfn_count;
+ }
+ if (err == 0 && unlikely(vmap->vm_ops == NULL)) {
+ err = VM_FAULT_SIGBUS;
+ } else if (err == 0) {
struct vm_fault vmf;
/* fill out VM fault structure */
vmf.virtual_address = (void *)(uintptr_t)IDX_TO_OFF(pidx);
vmf.flags = (fault_type & VM_PROT_WRITE) ? FAULT_FLAG_WRITE : 0;
vmf.pgoff = 0;
vmf.page = NULL;
vmf.vma = vmap;
- vmap->vm_pfn_count = 0;
- vmap->vm_pfn_pcount = &vmap->vm_pfn_count;
- vmap->vm_obj = vm_obj;
-
err = vmap->vm_ops->fault(&vmf);
- while (vmap->vm_pfn_count == 0 && err == VM_FAULT_NOPAGE) {
- kern_yield(PRI_USER);
- err = vmap->vm_ops->fault(&vmf);
- }
+ retry = !lkpi_vma_pfn_unchanged(vmap, invalidation_seq);
+ pfn_error = lkpi_vma_pfn_error(vmap);
+ if (pfn_error != 0)
+ err = pfn_error == ENOMEM ? VM_FAULT_OOM :
+ pfn_error == EAGAIN ? VM_FAULT_RETRY : VM_FAULT_SIGBUS;
+ else if (vmap->vm_pfn_count == 0 && err == VM_FAULT_NOPAGE)
+ err = VM_FAULT_RETRY;
}
/* translate return code */
- switch (err) {
+ switch (retry ? VM_FAULT_RETRY : err) {
+ case VM_FAULT_RETRY:
+ err = VM_PAGER_RETRY;
+ break;
case VM_FAULT_OOM:
err = VM_PAGER_AGAIN;
break;
case VM_FAULT_SIGBUS:
- err = VM_PAGER_BAD;
+ err = VM_PAGER_OUT_OF_BOUNDS;
break;
case VM_FAULT_NOPAGE:
/*
- * By contract the fault handler will return having
- * busied all the pages itself. If pidx is already
- * found in the object, it will simply xbusy the first
- * page and return with vm_pfn_count set to 1.
+ * By contract, the fault handler returns with every
+ * populated page exclusively busied. A page is either
+ * installed in the pager object or handed to the VM by
+ * cdev_pg_populate_take_page(). The two representations
+ * cannot be mixed in one populated range.
*/
+ if (!lkpi_vma_pfn_handoff_valid(vmap)) {
+ err = VM_PAGER_OUT_OF_BOUNDS;
+ break;
+ }
*first = vmap->vm_pfn_first;
*last = *first + vmap->vm_pfn_count - 1;
- MPASS(pidx >= *first);
- MPASS(pidx <= *last);
err = VM_PAGER_OK;
break;
default:
err = VM_PAGER_ERROR;
break;
}
up_write(&vmap->vm_mm->mmap_sem);
VM_OBJECT_WLOCK(vm_obj);
+ if (err == VM_PAGER_OK &&
+ (*first > pidx || *last < pidx || *last >= vm_obj->size))
+ err = VM_PAGER_OUT_OF_BOUNDS;
+ if (err != VM_PAGER_OK && pfn_started) {
+ if (vmap->vm_pfn_count != 0)
+ lkpi_vma_pfn_abort(vmap, vm_obj);
+ lkpi_vma_pfn_end(vmap);
+ }
+ /* A repeatedly interrupted driver must not trap a killed process here. */
+ if (err == VM_PAGER_RETRY && P_KILLED(curproc))
+ err = VM_PAGER_ERROR;
return (err);
}
+static vm_page_t
+linux_cdev_pager_populate_take_page(vm_object_t vm_obj, vm_pindex_t pidx)
+{
+ struct vm_area_struct *vmap;
+
+ vmap = linux_cdev_handle_find(vm_obj->handle);
+ MPASS(vmap != NULL);
+ MPASS(vmap->vm_private_data == vm_obj->handle);
+ return (lkpi_vma_pfn_take_page(vmap, vm_obj, pidx));
+}
+
+static void
+linux_cdev_pager_populate_release_page(vm_object_t vm_obj, vm_page_t page)
+{
+ struct vm_area_struct *vmap;
+
+ vmap = linux_cdev_handle_find(vm_obj->handle);
+ MPASS(vmap != NULL);
+ lkpi_vma_pfn_release_page(vmap, page);
+}
+
+static void
+linux_cdev_pager_populate_done(vm_object_t vm_obj)
+{
+ struct vm_area_struct *vmap;
+
+ vmap = linux_cdev_handle_find(vm_obj->handle);
+ MPASS(vmap != NULL);
+ MPASS(vmap->vm_private_data == vm_obj->handle);
+ lkpi_vma_pfn_done(vmap, vm_obj);
+ lkpi_vma_pfn_end(vmap);
+}
+
static struct rwlock linux_vma_lock;
static TAILQ_HEAD(, vm_area_struct) linux_vma_head =
TAILQ_HEAD_INITIALIZER(linux_vma_head);
static void
linux_cdev_handle_free(struct vm_area_struct *vmap)
{
/* Drop reference on vm_file */
if (vmap->vm_file != NULL)
fput(vmap->vm_file);
/* Drop reference on mm_struct */
mmput(vmap->vm_mm);
kfree(vmap);
}
static void
linux_cdev_handle_remove(struct vm_area_struct *vmap)
{
rw_wlock(&linux_vma_lock);
TAILQ_REMOVE(&linux_vma_head, vmap, vm_entry);
rw_wunlock(&linux_vma_lock);
}
static struct vm_area_struct *
linux_cdev_handle_find(void *handle)
{
struct vm_area_struct *vmap;
rw_rlock(&linux_vma_lock);
TAILQ_FOREACH(vmap, &linux_vma_head, vm_entry) {
if (vmap->vm_private_data == handle)
break;
}
rw_runlock(&linux_vma_lock);
return (vmap);
}
static int
linux_cdev_pager_ctor(void *handle, vm_ooffset_t size, vm_prot_t prot,
vm_ooffset_t foff, struct ucred *cred, u_short *color)
{
MPASS(linux_cdev_handle_find(handle) != NULL);
*color = 0;
return (0);
}
+static int
+linux_cdev_mgtdev_pager_ctor(void *handle, vm_ooffset_t size, vm_prot_t prot,
+ vm_ooffset_t foff, struct ucred *cred, u_short *color)
+{
+ struct vm_area_struct *vmap;
+ int error;
+
+ vmap = linux_cdev_handle_find(handle);
+ MPASS(vmap != NULL);
+ error = lkpi_vma_pfn_init(vmap);
+ if (error != 0)
+ return (error);
+ *color = 0;
+ return (0);
+}
+
static void
linux_cdev_pager_dtor(void *handle)
{
const struct vm_operations_struct *vm_ops;
struct vm_area_struct *vmap;
vmap = linux_cdev_handle_find(handle);
MPASS(vmap != NULL);
/*
* Remove handle before calling close operation to prevent
* other threads from reusing the handle pointer.
*/
linux_cdev_handle_remove(vmap);
down_write(&vmap->vm_mm->mmap_sem);
+ lkpi_vma_pfn_fini(vmap);
vm_ops = vmap->vm_ops;
if (likely(vm_ops != NULL))
vm_ops->close(vmap);
up_write(&vmap->vm_mm->mmap_sem);
linux_cdev_handle_free(vmap);
}
+static bool
+linux_cdev_pager_mlock_skip(void *handle)
+{
+ struct vm_area_struct *vmap;
+
+ vmap = linux_cdev_handle_find(handle);
+ MPASS(vmap != NULL);
+ /* Linux does not mlock special IO/PFN mappings. */
+ return ((vmap->vm_flags & (VM_IO | VM_PFNMAP)) != 0);
+}
+
static struct cdev_pager_ops linux_cdev_pager_ops[2] = {
{
/* OBJT_MGTDEVICE */
.cdev_pg_populate = linux_cdev_pager_populate,
- .cdev_pg_ctor = linux_cdev_pager_ctor,
+ .cdev_pg_populate_take_page = linux_cdev_pager_populate_take_page,
+ .cdev_pg_populate_done = linux_cdev_pager_populate_done,
+ .cdev_pg_populate_release_page = linux_cdev_pager_populate_release_page,
+ .cdev_pg_mlock_skip = linux_cdev_pager_mlock_skip,
+ .cdev_pg_ctor = linux_cdev_mgtdev_pager_ctor,
.cdev_pg_dtor = linux_cdev_pager_dtor
},
{
/* OBJT_DEVICE */
.cdev_pg_fault = linux_cdev_pager_fault,
.cdev_pg_ctor = linux_cdev_pager_ctor,
.cdev_pg_dtor = linux_cdev_pager_dtor
},
};
+void
+linux_cdev_pager_free_pages(vm_object_t object)
+{
+ struct vm_area_struct *vmap;
+ bool invalidating;
+
+ vmap = linux_cdev_handle_find(object->handle);
+ if (vmap == NULL) {
+ cdev_mgtdev_pager_free_pages(object);
+ return;
+ }
+ invalidating = lkpi_vma_pfn_invalidate_begin(vmap);
+ if (!invalidating) {
+ cdev_mgtdev_pager_free_pages(object);
+ return;
+ }
+ /*
+ * Mark invalidation before removing either representation. A fault
+ * which has entered the driver but not yet selected a page will reject
+ * its result. A page selected earlier remains xbusy until
+ * vm_fault_populate() installs its PTE, so either removal pass waits and
+ * revokes that PTE before invalidation completes. Pager-owned wires
+ * protect preserved pages until that removal and cache restoration finish.
+ */
+ cdev_mgtdev_pager_free_pages(object);
+ lkpi_vma_pfn_unmap(vmap);
+ lkpi_vma_pfn_invalidate_end(vmap);
+}
+
int
zap_vma_ptes(struct vm_area_struct *vma, unsigned long address,
unsigned long size)
{
struct pctrie_iter pages;
vm_object_t obj;
vm_page_t m;
+ vm_pindex_t first, last;
+ bool unmapping;
+ bool pfn_locked;
+ int error;
- obj = vma->vm_obj;
- if (obj == NULL || (obj->flags & OBJ_UNMANAGED) != 0)
- return (-ENOTSUP);
- VM_OBJECT_RLOCK(obj);
- vm_page_iter_limit_init(&pages, obj, OFF_TO_IDX(address + size));
- VM_RADIX_FOREACH_FROM(m, &pages, OFF_TO_IDX(address))
+ pfn_locked = lkpi_vma_pfn_lock(vma);
+ obj = (void *)atomic_load_acq_ptr(
+ (volatile uintptr_t *)&vma->vm_obj);
+ if (obj == NULL || (obj->flags & OBJ_UNMANAGED) != 0) {
+ error = -ENOTSUP;
+ goto out;
+ }
+ if (size == 0) {
+ error = 0;
+ goto out;
+ }
+ if (offset_in_page(address) != 0 || offset_in_page(size) != 0 ||
+ address < vma->vm_start || address >= vma->vm_end ||
+ size > vma->vm_end - address) {
+ error = -EINVAL;
+ goto out;
+ }
+ first = OFF_TO_IDX(address - vma->vm_start);
+ last = OFF_TO_IDX(address + size - vma->vm_start);
+ VM_OBJECT_WLOCK(obj);
+ if (vma->vm_pfn_count != 0)
+ lkpi_vma_pfn_abort(vma, obj);
+ VM_OBJECT_WUNLOCK(obj);
+ unmapping = lkpi_vma_pfn_unmap_begin(vma);
+ if (unmapping)
+ lkpi_vma_pfn_unmap(vma);
+ VM_OBJECT_WLOCK(obj);
+ vm_page_iter_limit_init(&pages, obj, last);
+ VM_RADIX_FOREACH_FROM(m, &pages, first)
pmap_remove_all(m);
- VM_OBJECT_RUNLOCK(obj);
- return (0);
+ VM_OBJECT_WUNLOCK(obj);
+ if (unmapping)
+ lkpi_vma_pfn_unmap_end(vma);
+ error = 0;
+out:
+ if (pfn_locked)
+ lkpi_vma_pfn_end(vma);
+ return (error);
}
void
vma_set_file(struct vm_area_struct *vma, struct linux_file *file)
{
struct linux_file *tmp;
/* Changing an anonymous vma with this is illegal */
get_file(file);
tmp = vma->vm_file;
vma->vm_file = file;
fput(tmp);
}
static struct file_operations dummy_ldev_ops = {
/* XXXKIB */
};
static struct linux_cdev dummy_ldev = {
.ops = &dummy_ldev_ops,
};
#define LDEV_SI_DTR 0x0001
#define LDEV_SI_REF 0x0002
static void
linux_get_fop(struct linux_file *filp, const struct file_operations **fop,
struct linux_cdev **dev)
{
struct linux_cdev *ldev;
u_int siref;
ldev = filp->f_cdev;
*fop = filp->f_op;
if (ldev != NULL) {
if (ldev->kobj.ktype == &linux_cdev_static_ktype) {
refcount_acquire(&ldev->refs);
} else {
for (siref = ldev->siref;;) {
if ((siref & LDEV_SI_DTR) != 0) {
ldev = &dummy_ldev;
*fop = ldev->ops;
siref = ldev->siref;
MPASS((ldev->siref & LDEV_SI_DTR) == 0);
} else if (atomic_fcmpset_int(&ldev->siref,
&siref, siref + LDEV_SI_REF)) {
break;
}
}
}
}
*dev = ldev;
}
static void
linux_drop_fop(struct linux_cdev *ldev)
{
if (ldev == NULL)
return;
if (ldev->kobj.ktype == &linux_cdev_static_ktype) {
linux_cdev_deref(ldev);
} else {
MPASS(ldev->kobj.ktype == &linux_cdev_ktype);
MPASS((ldev->siref & ~LDEV_SI_DTR) != 0);
atomic_subtract_int(&ldev->siref, LDEV_SI_REF);
}
}
#define OPW(fp,td,code) ({ \
struct file *__fpop; \
__typeof(code) __retval; \
\
__fpop = (td)->td_fpop; \
(td)->td_fpop = (fp); \
__retval = (code); \
(td)->td_fpop = __fpop; \
__retval; \
})
static int
linux_dev_fdopen(struct cdev *dev, int fflags, struct thread *td,
struct file *file)
{
struct linux_cdev *ldev;
struct linux_file *filp;
const struct file_operations *fop;
int error;
ldev = dev->si_drv1;
filp = linux_file_alloc();
filp->f_dentry = &filp->f_dentry_store;
filp->f_op = ldev->ops;
filp->f_mode = file->f_flag;
filp->f_flags = file->f_flag;
filp->f_vnode = file->f_vnode;
filp->_file = file;
refcount_acquire(&ldev->refs);
filp->f_cdev = ldev;
linux_set_current(td);
linux_get_fop(filp, &fop, &ldev);
if (fop->open != NULL) {
error = -fop->open(file->f_vnode, filp);
if (error != 0) {
linux_drop_fop(ldev);
linux_cdev_deref(filp->f_cdev);
kfree(filp);
return (error);
}
}
/* hold on to the vnode - used for fstat() */
vref(filp->f_vnode);
/* release the file from devfs */
finit(file, filp->f_mode, DTYPE_DEV, filp, &linuxfileops);
linux_drop_fop(ldev);
return (ENXIO);
}
#define LINUX_IOCTL_MIN_PTR 0x10000UL
#define LINUX_IOCTL_MAX_PTR (LINUX_IOCTL_MIN_PTR + IOCPARM_MAX)
static inline int
linux_remap_address(void **uaddr, size_t len)
{
uintptr_t uaddr_val = (uintptr_t)(*uaddr);
if (unlikely(uaddr_val >= LINUX_IOCTL_MIN_PTR &&
uaddr_val < LINUX_IOCTL_MAX_PTR)) {
struct task_struct *pts = current;
if (pts == NULL) {
*uaddr = NULL;
return (1);
}
/* compute data offset */
uaddr_val -= LINUX_IOCTL_MIN_PTR;
/* check that length is within bounds */
if ((len > IOCPARM_MAX) ||
(uaddr_val + len) > pts->bsd_ioctl_len) {
*uaddr = NULL;
return (1);
}
/* re-add kernel buffer address */
uaddr_val += (uintptr_t)pts->bsd_ioctl_data;
/* update address location */
*uaddr = (void *)uaddr_val;
return (1);
}
return (0);
}
int
linux_copyin(const void *uaddr, void *kaddr, size_t len)
{
if (linux_remap_address(__DECONST(void **, &uaddr), len)) {
if (uaddr == NULL)
return (-EFAULT);
memcpy(kaddr, uaddr, len);
return (0);
}
return (-copyin(uaddr, kaddr, len));
}
int
linux_copyout(const void *kaddr, void *uaddr, size_t len)
{
if (linux_remap_address(&uaddr, len)) {
if (uaddr == NULL)
return (-EFAULT);
memcpy(uaddr, kaddr, len);
return (0);
}
return (-copyout(kaddr, uaddr, len));
}
size_t
linux_clear_user(void *_uaddr, size_t _len)
{
uint8_t *uaddr = _uaddr;
size_t len = _len;
/* make sure uaddr is aligned before going into the fast loop */
while (((uintptr_t)uaddr & 7) != 0 && len > 7) {
if (subyte(uaddr, 0))
return (_len);
uaddr++;
len--;
}
/* zero 8 bytes at a time */
while (len > 7) {
#ifdef __LP64__
if (suword64(uaddr, 0))
return (_len);
#else
if (suword32(uaddr, 0))
return (_len);
if (suword32(uaddr + 4, 0))
return (_len);
#endif
uaddr += 8;
len -= 8;
}
/* zero fill end, if any */
while (len > 0) {
if (subyte(uaddr, 0))
return (_len);
uaddr++;
len--;
}
return (0);
}
int
linux_access_ok(const void *uaddr, size_t len)
{
uintptr_t saddr;
uintptr_t eaddr;
/* get start and end address */
saddr = (uintptr_t)uaddr;
eaddr = (uintptr_t)uaddr + len;
/* verify addresses are valid for userspace */
return ((saddr == eaddr) ||
(eaddr > saddr && eaddr <= VM_MAXUSER_ADDRESS));
}
/*
* This function should return either EINTR or ERESTART depending on
* the signal type sent to this thread:
*/
static int
linux_get_error(struct task_struct *task, int error)
{
/* check for signal type interrupt code */
if (error == EINTR || error == ERESTARTSYS || error == ERESTART) {
error = -linux_schedule_get_interrupt_value(task);
if (error == 0)
error = EINTR;
}
return (error);
}
static int
linux_file_ioctl_sub(struct file *fp, struct linux_file *filp,
const struct file_operations *fop, u_long cmd, caddr_t data,
struct thread *td)
{
struct task_struct *task = current;
unsigned size;
int error;
bool direct;
size = IOCPARM_LEN(cmd);
/* refer to logic in sys_ioctl() */
direct = false;
if (size > 0) {
/*
* Setup hint for linux_copyin() and linux_copyout().
*
* Background: Linux kernel code expects to operate on
* userspace addresses, but FreeBSD's kern_ioctl()
* will generally provide a kernel address. For the
* native process ABI, where we know how to find the
* original address, we reach directly into the system
* call args to get it. Then, if the Linux driver
* copied out to that address, we copy the whole block
* back into the kernel buffer allocated by
* kern_ioctl() so that kern_ioctl() itself doesn't
* clobber the driver's data.
*
* Otherwise, fall back to the LINUX_IOCTL_MIN_PTR
* hack.
*/
task->bsd_ioctl_data = data;
task->bsd_ioctl_len = size;
if ((td->td_pflags & TDP_KTHREAD) == 0 &&
SV_PROC_ABI(td->td_proc) == SV_ABI_FREEBSD &&
td->td_sa.code == SYS_ioctl) {
direct = true;
data = (void *)(uintptr_t)td->td_sa.args[2];
} else {
data = (void *)LINUX_IOCTL_MIN_PTR;
}
} else {
/* fetch user-space pointer */
data = *(void **)data;
}
#ifdef COMPAT_FREEBSD32
if (SV_PROC_FLAG(td->td_proc, SV_ILP32)) {
/* try the compat IOCTL handler first */
if (fop->compat_ioctl != NULL) {
error = -OPW(fp, td, fop->compat_ioctl(filp,
cmd, (u_long)data));
} else {
error = ENOTTY;
}
/* fallback to the regular IOCTL handler, if any */
if (error == ENOTTY && fop->unlocked_ioctl != NULL) {
error = -OPW(fp, td, fop->unlocked_ioctl(filp,
cmd, (u_long)data));
}
} else
#endif
{
if (fop->unlocked_ioctl != NULL) {
error = -OPW(fp, td, fop->unlocked_ioctl(filp,
cmd, (u_long)data));
} else {
error = ENOTTY;
}
}
if (error == 0 && size > 0 && (cmd & IOC_OUT) != 0 && direct) {
void *xdata;
int error1;
/*
* Ensure that the copyout in sys_generic.c copies
* over the data which is possibly modified by the
* driver. A possible error from the copyin() is
* ignored since it is formally possible for the memory
* to become unaccessible in the meantime. Do the copying
* through the intermediate buffer instead of copying
* directly to bsd_ioctl_data, to ensure atomicity of
* the change with respect to the error.
*/
xdata = malloc(size, M_TEMP, M_WAITOK);
error1 = copyin(data, xdata, size);
if (error1 == 0)
memcpy(task->bsd_ioctl_data, xdata, size);
free(xdata, M_TEMP);
}
if (size > 0) {
task->bsd_ioctl_data = NULL;
task->bsd_ioctl_len = 0;
}
if (error == EWOULDBLOCK) {
/* update kqfilter status, if any */
linux_file_kqfilter_poll(filp,
LINUX_KQ_FLAG_HAS_READ | LINUX_KQ_FLAG_HAS_WRITE);
} else {
error = linux_get_error(task, error);
}
return (error);
}
#define LINUX_POLL_TABLE_NORMAL ((poll_table *)1)
/*
* This function atomically updates the poll wakeup state and returns
* the previous state at the time of update.
*/
static uint8_t
linux_poll_wakeup_state(atomic_t *v, const uint8_t *pstate)
{
int c, old;
c = v->counter;
while ((old = atomic_cmpxchg(v, c, pstate[c])) != c)
c = old;
return (c);
}
static int
linux_poll_wakeup_callback(wait_queue_t *wq, unsigned int wq_state, int flags, void *key)
{
static const uint8_t state[LINUX_FWQ_STATE_MAX] = {
[LINUX_FWQ_STATE_INIT] = LINUX_FWQ_STATE_INIT, /* NOP */
[LINUX_FWQ_STATE_NOT_READY] = LINUX_FWQ_STATE_NOT_READY, /* NOP */
[LINUX_FWQ_STATE_QUEUED] = LINUX_FWQ_STATE_READY,
[LINUX_FWQ_STATE_READY] = LINUX_FWQ_STATE_READY, /* NOP */
};
struct linux_file *filp = container_of(wq, struct linux_file, f_wait_queue.wq);
switch (linux_poll_wakeup_state(&filp->f_wait_queue.state, state)) {
case LINUX_FWQ_STATE_QUEUED:
linux_poll_wakeup(filp);
return (1);
default:
return (0);
}
}
void
linux_poll_wait(struct linux_file *filp, wait_queue_head_t *wqh, poll_table *p)
{
static const uint8_t state[LINUX_FWQ_STATE_MAX] = {
[LINUX_FWQ_STATE_INIT] = LINUX_FWQ_STATE_NOT_READY,
[LINUX_FWQ_STATE_NOT_READY] = LINUX_FWQ_STATE_NOT_READY, /* NOP */
[LINUX_FWQ_STATE_QUEUED] = LINUX_FWQ_STATE_QUEUED, /* NOP */
[LINUX_FWQ_STATE_READY] = LINUX_FWQ_STATE_QUEUED,
};
/* check if we are called inside the select system call */
if (p == LINUX_POLL_TABLE_NORMAL)
selrecord(curthread, &filp->f_selinfo);
switch (linux_poll_wakeup_state(&filp->f_wait_queue.state, state)) {
case LINUX_FWQ_STATE_INIT:
/* NOTE: file handles can only belong to one wait-queue */
filp->f_wait_queue.wqh = wqh;
filp->f_wait_queue.wq.func = &linux_poll_wakeup_callback;
add_wait_queue(wqh, &filp->f_wait_queue.wq);
atomic_set(&filp->f_wait_queue.state, LINUX_FWQ_STATE_QUEUED);
break;
default:
break;
}
}
static void
linux_poll_wait_dequeue(struct linux_file *filp)
{
static const uint8_t state[LINUX_FWQ_STATE_MAX] = {
[LINUX_FWQ_STATE_INIT] = LINUX_FWQ_STATE_INIT, /* NOP */
[LINUX_FWQ_STATE_NOT_READY] = LINUX_FWQ_STATE_INIT,
[LINUX_FWQ_STATE_QUEUED] = LINUX_FWQ_STATE_INIT,
[LINUX_FWQ_STATE_READY] = LINUX_FWQ_STATE_INIT,
};
seldrain(&filp->f_selinfo);
switch (linux_poll_wakeup_state(&filp->f_wait_queue.state, state)) {
case LINUX_FWQ_STATE_NOT_READY:
case LINUX_FWQ_STATE_QUEUED:
case LINUX_FWQ_STATE_READY:
remove_wait_queue(filp->f_wait_queue.wqh, &filp->f_wait_queue.wq);
break;
default:
break;
}
}
void
linux_poll_wakeup(struct linux_file *filp)
{
/* this function should be NULL-safe */
if (filp == NULL)
return;
selwakeup(&filp->f_selinfo);
spin_lock(&filp->f_kqlock);
filp->f_kqflags |= LINUX_KQ_FLAG_NEED_READ |
LINUX_KQ_FLAG_NEED_WRITE;
/* make sure the "knote" gets woken up */
KNOTE_LOCKED(&filp->f_selinfo.si_note, 1);
spin_unlock(&filp->f_kqlock);
}
static struct linux_file *
__get_file_rcu(struct linux_file **f)
{
struct linux_file *file1, *file2;
file1 = READ_ONCE(*f);
if (file1 == NULL)
return (NULL);
if (!refcount_acquire_if_not_zero(
file1->_file == NULL ? &file1->f_count : &file1->_file->f_count))
return (ERR_PTR(-EAGAIN));
file2 = READ_ONCE(*f);
if (file2 == file1)
return (file2);
fput(file1);
return (ERR_PTR(-EAGAIN));
}
struct linux_file *
linux_get_file_rcu(struct linux_file **f)
{
struct linux_file *file1;
for (;;) {
file1 = __get_file_rcu(f);
if (file1 == NULL)
return (NULL);
if (IS_ERR(file1))
continue;
return (file1);
}
}
struct linux_file *
get_file_active(struct linux_file **f)
{
struct linux_file *file1;
rcu_read_lock();
file1 = __get_file_rcu(f);
rcu_read_unlock();
if (IS_ERR(file1))
file1 = NULL;
return (file1);
}
static void
linux_file_kqfilter_detach(struct knote *kn)
{
struct linux_file *filp = kn->kn_hook;
spin_lock(&filp->f_kqlock);
knlist_remove(&filp->f_selinfo.si_note, kn, 1);
spin_unlock(&filp->f_kqlock);
}
static int
linux_file_kqfilter_read_event(struct knote *kn, long hint)
{
struct linux_file *filp = kn->kn_hook;
mtx_assert(&filp->f_kqlock, MA_OWNED);
return ((filp->f_kqflags & LINUX_KQ_FLAG_NEED_READ) ? 1 : 0);
}
static int
linux_file_kqfilter_write_event(struct knote *kn, long hint)
{
struct linux_file *filp = kn->kn_hook;
mtx_assert(&filp->f_kqlock, MA_OWNED);
return ((filp->f_kqflags & LINUX_KQ_FLAG_NEED_WRITE) ? 1 : 0);
}
static const struct filterops linux_dev_kqfiltops_read = {
.f_isfd = 1,
.f_detach = linux_file_kqfilter_detach,
.f_event = linux_file_kqfilter_read_event,
.f_copy = knote_triv_copy,
};
static const struct filterops linux_dev_kqfiltops_write = {
.f_isfd = 1,
.f_detach = linux_file_kqfilter_detach,
.f_event = linux_file_kqfilter_write_event,
.f_copy = knote_triv_copy,
};
static void
linux_file_kqfilter_poll(struct linux_file *filp, int kqflags)
{
struct thread *td;
const struct file_operations *fop;
struct linux_cdev *ldev;
int temp;
if ((filp->f_kqflags & kqflags) == 0)
return;
td = curthread;
linux_get_fop(filp, &fop, &ldev);
/* get the latest polling state */
temp = OPW(filp->_file, td, fop->poll(filp, NULL));
linux_drop_fop(ldev);
spin_lock(&filp->f_kqlock);
/* clear kqflags */
filp->f_kqflags &= ~(LINUX_KQ_FLAG_NEED_READ |
LINUX_KQ_FLAG_NEED_WRITE);
/* update kqflags */
if ((temp & (POLLIN | POLLOUT)) != 0) {
if ((temp & POLLIN) != 0)
filp->f_kqflags |= LINUX_KQ_FLAG_NEED_READ;
if ((temp & POLLOUT) != 0)
filp->f_kqflags |= LINUX_KQ_FLAG_NEED_WRITE;
/* make sure the "knote" gets woken up */
KNOTE_LOCKED(&filp->f_selinfo.si_note, 0);
}
spin_unlock(&filp->f_kqlock);
}
static int
linux_file_kqfilter(struct file *file, struct knote *kn)
{
struct linux_file *filp;
struct thread *td;
int error;
td = curthread;
filp = (struct linux_file *)file->f_data;
filp->f_flags = file->f_flag;
if (filp->f_op->poll == NULL)
return (EINVAL);
spin_lock(&filp->f_kqlock);
switch (kn->kn_filter) {
case EVFILT_READ:
filp->f_kqflags |= LINUX_KQ_FLAG_HAS_READ;
kn->kn_fop = &linux_dev_kqfiltops_read;
kn->kn_hook = filp;
knlist_add(&filp->f_selinfo.si_note, kn, 1);
error = 0;
break;
case EVFILT_WRITE:
filp->f_kqflags |= LINUX_KQ_FLAG_HAS_WRITE;
kn->kn_fop = &linux_dev_kqfiltops_write;
kn->kn_hook = filp;
knlist_add(&filp->f_selinfo.si_note, kn, 1);
error = 0;
break;
default:
error = EINVAL;
break;
}
spin_unlock(&filp->f_kqlock);
if (error == 0) {
linux_set_current(td);
/* update kqfilter status, if any */
linux_file_kqfilter_poll(filp,
LINUX_KQ_FLAG_HAS_READ | LINUX_KQ_FLAG_HAS_WRITE);
}
return (error);
}
static int
linux_file_mmap_single(struct file *fp, const struct file_operations *fop,
vm_ooffset_t *offset, vm_size_t size, struct vm_object **object,
int nprot, bool is_shared, struct thread *td)
{
struct task_struct *task;
struct vm_area_struct *vmap;
struct mm_struct *mm;
struct linux_file *filp;
vm_memattr_t attr;
int error;
filp = (struct linux_file *)fp->f_data;
filp->f_flags = fp->f_flag;
if (fop->mmap == NULL)
return (EOPNOTSUPP);
linux_set_current(td);
/*
* The same VM object might be shared by multiple processes
* and the mm_struct is usually freed when a process exits.
*
* The atomic reference below makes sure the mm_struct is
* available as long as the vmap is in the linux_vma_head.
*/
task = current;
mm = task->mm;
if (atomic_inc_not_zero(&mm->mm_users) == 0)
return (EINVAL);
vmap = kzalloc(sizeof(*vmap), GFP_KERNEL);
vmap->vm_start = 0;
vmap->vm_end = size;
vmap->vm_pgoff = *offset / PAGE_SIZE;
vmap->vm_pfn = 0;
vmap->vm_flags = vmap->vm_page_prot = (nprot & VM_PROT_ALL);
if (is_shared)
vmap->vm_flags |= VM_SHARED;
vmap->vm_ops = NULL;
vmap->vm_file = get_file(filp);
vmap->vm_mm = mm;
if (unlikely(down_write_killable(&vmap->vm_mm->mmap_sem))) {
error = linux_get_error(task, EINTR);
} else {
error = -OPW(fp, td, fop->mmap(filp, vmap));
error = linux_get_error(task, error);
up_write(&vmap->vm_mm->mmap_sem);
}
if (error != 0) {
linux_cdev_handle_free(vmap);
return (error);
}
attr = pgprot2cachemode(vmap->vm_page_prot);
if (vmap->vm_ops != NULL) {
struct vm_area_struct *ptr;
void *vm_private_data;
bool vm_no_fault;
if (vmap->vm_ops->open == NULL ||
vmap->vm_ops->close == NULL ||
vmap->vm_private_data == NULL) {
/* free allocated VM area struct */
linux_cdev_handle_free(vmap);
return (EINVAL);
}
vm_private_data = vmap->vm_private_data;
rw_wlock(&linux_vma_lock);
TAILQ_FOREACH(ptr, &linux_vma_head, vm_entry) {
if (ptr->vm_private_data == vm_private_data)
break;
}
/* check if there is an existing VM area struct */
if (ptr != NULL) {
/* check if the VM area structure is invalid */
if (ptr->vm_ops == NULL ||
ptr->vm_ops->open == NULL ||
ptr->vm_ops->close == NULL) {
error = ESTALE;
vm_no_fault = 1;
} else {
if (ptr->vm_start == vmap->vm_start &&
ptr->vm_end <= vmap->vm_end) {
/*
* Userspace wants to grow an existing
* mapping. We already have a
* `vm_object_t' for this mapping. We
* just need to update the `struct
* vm_area_struct` to have the correct
* end address.
*/
ptr->vm_end = vmap->vm_end;
}
error = EEXIST;
vm_no_fault = (ptr->vm_ops->fault == NULL);
}
} else {
/* insert VM area structure into list */
TAILQ_INSERT_TAIL(&linux_vma_head, vmap, vm_entry);
error = 0;
vm_no_fault = (vmap->vm_ops->fault == NULL);
}
rw_wunlock(&linux_vma_lock);
if (error != 0) {
/* free allocated VM area struct */
linux_cdev_handle_free(vmap);
/* check for stale VM area struct */
if (error != EEXIST)
return (error);
}
/* check if there is no fault handler */
if (vm_no_fault) {
*object = cdev_pager_allocate(vm_private_data, OBJT_DEVICE,
&linux_cdev_pager_ops[1], size, nprot, *offset,
td->td_ucred);
} else {
*object = cdev_pager_allocate(vm_private_data, OBJT_MGTDEVICE,
&linux_cdev_pager_ops[0], size, nprot, *offset,
td->td_ucred);
}
/* check if allocating the VM object failed */
if (*object == NULL) {
if (error == 0) {
/* remove VM area struct from list */
linux_cdev_handle_remove(vmap);
/* free allocated VM area struct */
linux_cdev_handle_free(vmap);
}
return (EINVAL);
}
} else {
struct sglist *sg;
sg = sglist_alloc(1, M_WAITOK);
sglist_append_phys(sg,
(vm_paddr_t)vmap->vm_pfn << PAGE_SHIFT, vmap->vm_len);
*object = vm_pager_allocate(OBJT_SG, sg, vmap->vm_len,
nprot, 0, td->td_ucred);
linux_cdev_handle_free(vmap);
if (*object == NULL) {
sglist_free(sg);
return (EINVAL);
}
}
if (attr != VM_MEMATTR_DEFAULT) {
VM_OBJECT_WLOCK(*object);
vm_object_set_memattr(*object, attr);
VM_OBJECT_WUNLOCK(*object);
}
*offset = 0;
return (0);
}
struct cdevsw linuxcdevsw = {
.d_version = D_VERSION,
.d_fdopen = linux_dev_fdopen,
.d_name = "lkpidev",
};
static int
linux_file_read(struct file *file, struct uio *uio, struct ucred *active_cred,
int flags, struct thread *td)
{
struct linux_file *filp;
const struct file_operations *fop;
struct linux_cdev *ldev;
ssize_t bytes;
int error;
error = 0;
filp = (struct linux_file *)file->f_data;
filp->f_flags = file->f_flag;
/* XXX no support for I/O vectors currently */
if (uio->uio_iovcnt != 1)
return (EOPNOTSUPP);
if (uio->uio_resid > DEVFS_IOSIZE_MAX)
return (EINVAL);
linux_set_current(td);
linux_get_fop(filp, &fop, &ldev);
if (fop->read != NULL) {
bytes = OPW(file, td, fop->read(filp,
uio->uio_iov->iov_base,
uio->uio_iov->iov_len, &uio->uio_offset));
if (bytes >= 0) {
uio->uio_iov->iov_base =
((uint8_t *)uio->uio_iov->iov_base) + bytes;
uio->uio_iov->iov_len -= bytes;
uio->uio_resid -= bytes;
} else {
error = linux_get_error(current, -bytes);
}
} else
error = ENXIO;
/* update kqfilter status, if any */
linux_file_kqfilter_poll(filp, LINUX_KQ_FLAG_HAS_READ);
linux_drop_fop(ldev);
return (error);
}
static int
linux_file_write(struct file *file, struct uio *uio, struct ucred *active_cred,
int flags, struct thread *td)
{
struct linux_file *filp;
const struct file_operations *fop;
struct linux_cdev *ldev;
ssize_t bytes;
int error;
filp = (struct linux_file *)file->f_data;
filp->f_flags = file->f_flag;
/* XXX no support for I/O vectors currently */
if (uio->uio_iovcnt != 1)
return (EOPNOTSUPP);
if (uio->uio_resid > DEVFS_IOSIZE_MAX)
return (EINVAL);
linux_set_current(td);
linux_get_fop(filp, &fop, &ldev);
if (fop->write != NULL) {
bytes = OPW(file, td, fop->write(filp,
uio->uio_iov->iov_base,
uio->uio_iov->iov_len, &uio->uio_offset));
if (bytes >= 0) {
uio->uio_iov->iov_base =
((uint8_t *)uio->uio_iov->iov_base) + bytes;
uio->uio_iov->iov_len -= bytes;
uio->uio_resid -= bytes;
error = 0;
} else {
error = linux_get_error(current, -bytes);
}
} else
error = ENXIO;
/* update kqfilter status, if any */
linux_file_kqfilter_poll(filp, LINUX_KQ_FLAG_HAS_WRITE);
linux_drop_fop(ldev);
return (error);
}
static int
linux_file_poll(struct file *file, int events, struct ucred *active_cred,
struct thread *td)
{
struct linux_file *filp;
const struct file_operations *fop;
struct linux_cdev *ldev;
int revents;
filp = (struct linux_file *)file->f_data;
filp->f_flags = file->f_flag;
linux_set_current(td);
linux_get_fop(filp, &fop, &ldev);
if (fop->poll != NULL) {
revents = OPW(file, td, fop->poll(filp,
LINUX_POLL_TABLE_NORMAL)) & events;
} else {
revents = 0;
}
linux_drop_fop(ldev);
return (revents);
}
static int
linux_file_close(struct file *file, struct thread *td)
{
struct linux_file *filp;
int (*release)(struct inode *, struct linux_file *);
const struct file_operations *fop;
struct linux_cdev *ldev;
int error;
filp = (struct linux_file *)file->f_data;
KASSERT(file_count(filp) == 0,
("File refcount(%d) is not zero", file_count(filp)));
if (td == NULL)
td = curthread;
error = 0;
filp->f_flags = file->f_flag;
linux_set_current(td);
linux_poll_wait_dequeue(filp);
linux_get_fop(filp, &fop, &ldev);
/*
* Always use the real release function, if any, to avoid
* leaking device resources:
*/
release = filp->f_op->release;
if (release != NULL)
error = -OPW(file, td, release(filp->f_vnode, filp));
funsetown(&filp->f_sigio);
if (filp->f_vnode != NULL)
vrele(filp->f_vnode);
linux_drop_fop(ldev);
ldev = filp->f_cdev;
if (ldev != NULL)
linux_cdev_deref(ldev);
linux_synchronize_rcu(RCU_TYPE_REGULAR);
kfree(filp);
return (error);
}
static int
linux_file_ioctl(struct file *fp, u_long cmd, void *data, struct ucred *cred,
struct thread *td)
{
struct linux_file *filp;
const struct file_operations *fop;
struct linux_cdev *ldev;
struct fiodgname_arg *fgn;
const char *p;
int error, i;
error = 0;
filp = (struct linux_file *)fp->f_data;
filp->f_flags = fp->f_flag;
linux_get_fop(filp, &fop, &ldev);
linux_set_current(td);
switch (cmd) {
case FIONBIO:
break;
case FIOASYNC:
if (fop->fasync == NULL)
break;
error = -OPW(fp, td, fop->fasync(0, filp, fp->f_flag & FASYNC));
break;
case FIOSETOWN:
error = fsetown(*(int *)data, &filp->f_sigio);
if (error == 0) {
if (fop->fasync == NULL)
break;
error = -OPW(fp, td, fop->fasync(0, filp,
fp->f_flag & FASYNC));
}
break;
case FIOGETOWN:
*(int *)data = fgetown(&filp->f_sigio);
break;
case FIODGNAME:
#ifdef COMPAT_FREEBSD32
case FIODGNAME_32:
#endif
if (filp->f_cdev == NULL || filp->f_cdev->cdev == NULL) {
error = ENXIO;
break;
}
fgn = data;
p = devtoname(filp->f_cdev->cdev);
i = strlen(p) + 1;
if (i > fgn->len) {
error = EINVAL;
break;
}
error = copyout(p, fiodgname_buf_get_ptr(fgn, cmd), i);
break;
default:
error = linux_file_ioctl_sub(fp, filp, fop, cmd, data, td);
break;
}
linux_drop_fop(ldev);
return (error);
}
static int
linux_file_mmap_sub(struct thread *td, vm_size_t objsize, vm_prot_t prot,
vm_prot_t maxprot, int flags, struct file *fp,
vm_ooffset_t *foff, const struct file_operations *fop, vm_object_t *objp)
{
/*
* Character devices do not provide private mappings
* of any kind:
*/
if ((maxprot & VM_PROT_WRITE) == 0 &&
(prot & VM_PROT_WRITE) != 0)
return (EACCES);
if ((flags & (MAP_PRIVATE | MAP_COPY)) != 0)
return (EINVAL);
return (linux_file_mmap_single(fp, fop, foff, objsize, objp,
(int)prot, (flags & MAP_SHARED) ? true : false, td));
}
static int
linux_file_mmap(struct file *fp, vm_map_t map, vm_offset_t *addr, vm_size_t size,
vm_prot_t prot, vm_prot_t cap_maxprot, int flags, vm_ooffset_t foff,
struct thread *td)
{
struct linux_file *filp;
const struct file_operations *fop;
struct linux_cdev *ldev;
struct mount *mp;
struct vnode *vp;
vm_object_t object;
vm_prot_t maxprot;
int error;
filp = (struct linux_file *)fp->f_data;
vp = filp->f_vnode;
if (vp == NULL)
return (EOPNOTSUPP);
/*
* Ensure that file and memory protections are
* compatible.
*/
mp = vp->v_mount;
if (mp != NULL && (mp->mnt_flag & MNT_NOEXEC) != 0) {
maxprot = VM_PROT_NONE;
if ((prot & VM_PROT_EXECUTE) != 0)
return (EACCES);
} else
maxprot = VM_PROT_EXECUTE;
if ((fp->f_flag & FREAD) != 0)
maxprot |= VM_PROT_READ;
else if ((prot & VM_PROT_READ) != 0)
return (EACCES);
/*
* If we are sharing potential changes via MAP_SHARED and we
* are trying to get write permission although we opened it
* without asking for it, bail out.
*
* Note that most character devices always share mappings.
*
* Rely on linux_file_mmap_sub() to fail invalid MAP_PRIVATE
* requests rather than doing it here.
*/
if ((flags & MAP_SHARED) != 0) {
if ((fp->f_flag & FWRITE) != 0)
maxprot |= VM_PROT_WRITE;
else if ((prot & VM_PROT_WRITE) != 0)
return (EACCES);
}
maxprot &= cap_maxprot;
linux_get_fop(filp, &fop, &ldev);
error = linux_file_mmap_sub(td, size, prot, maxprot, flags, fp,
&foff, fop, &object);
if (error != 0)
goto out;
error = vm_mmap_object(map, addr, size, prot, maxprot, flags, object,
foff, FALSE, td);
if (error != 0)
vm_object_deallocate(object);
out:
linux_drop_fop(ldev);
return (error);
}
static int
linux_file_stat(struct file *fp, struct stat *sb, struct ucred *active_cred)
{
struct linux_file *filp;
struct vnode *vp;
int error;
filp = (struct linux_file *)fp->f_data;
if (filp->f_vnode == NULL)
return (EOPNOTSUPP);
vp = filp->f_vnode;
vn_lock(vp, LK_SHARED | LK_RETRY);
error = VOP_STAT(vp, sb, curthread->td_ucred, NOCRED);
VOP_UNLOCK(vp);
return (error);
}
static int
linux_file_fill_kinfo(struct file *fp, struct kinfo_file *kif,
struct filedesc *fdp)
{
struct linux_file *filp;
struct vnode *vp;
int error;
filp = fp->f_data;
vp = filp->f_vnode;
if (vp == NULL) {
error = 0;
kif->kf_type = KF_TYPE_DEV;
} else {
vref(vp);
FILEDESC_SUNLOCK(fdp);
error = vn_fill_kinfo_vnode(vp, kif);
vrele(vp);
kif->kf_type = KF_TYPE_VNODE;
FILEDESC_SLOCK(fdp);
}
return (error);
}
unsigned int
linux_iminor(struct inode *inode)
{
struct linux_cdev *ldev;
if (inode == NULL || inode->v_rdev == NULL ||
inode->v_rdev->si_devsw != &linuxcdevsw)
return (-1U);
ldev = inode->v_rdev->si_drv1;
if (ldev == NULL)
return (-1U);
return (minor(ldev->dev));
}
static int
linux_file_kcmp(struct file *fp1, struct file *fp2, struct thread *td)
{
struct linux_file *filp1, *filp2;
if (fp2->f_type != DTYPE_DEV)
return (3);
filp1 = fp1->f_data;
filp2 = fp2->f_data;
return (kcmp_cmp((uintptr_t)filp1->f_cdev, (uintptr_t)filp2->f_cdev));
}
const struct fileops linuxfileops = {
.fo_read = linux_file_read,
.fo_write = linux_file_write,
.fo_truncate = invfo_truncate,
.fo_kqfilter = linux_file_kqfilter,
.fo_stat = linux_file_stat,
.fo_fill_kinfo = linux_file_fill_kinfo,
.fo_poll = linux_file_poll,
.fo_close = linux_file_close,
.fo_ioctl = linux_file_ioctl,
.fo_mmap = linux_file_mmap,
.fo_chmod = invfo_chmod,
.fo_chown = invfo_chown,
.fo_sendfile = invfo_sendfile,
.fo_cmp = linux_file_kcmp,
.fo_flags = DFLAG_PASSABLE,
};
static char *
devm_kvasprintf(struct device *dev, gfp_t gfp, const char *fmt, va_list ap)
{
unsigned int len;
char *p;
va_list aq;
va_copy(aq, ap);
len = vsnprintf(NULL, 0, fmt, aq);
va_end(aq);
if (dev != NULL)
p = devm_kmalloc(dev, len + 1, gfp);
else
p = kmalloc(len + 1, gfp);
if (p != NULL)
vsnprintf(p, len + 1, fmt, ap);
return (p);
}
char *
kvasprintf(gfp_t gfp, const char *fmt, va_list ap)
{
return (devm_kvasprintf(NULL, gfp, fmt, ap));
}
char *
lkpi_devm_kasprintf(struct device *dev, gfp_t gfp, const char *fmt, ...)
{
va_list ap;
char *p;
va_start(ap, fmt);
p = devm_kvasprintf(dev, gfp, fmt, ap);
va_end(ap);
return (p);
}
char *
kasprintf(gfp_t gfp, const char *fmt, ...)
{
va_list ap;
char *p;
va_start(ap, fmt);
p = kvasprintf(gfp, fmt, ap);
va_end(ap);
return (p);
}
int
__lkpi_hexdump_printf(void *arg1 __unused, const char *fmt, ...)
{
va_list ap;
int result;
va_start(ap, fmt);
result = vprintf(fmt, ap);
va_end(ap);
return (result);
}
int
__lkpi_hexdump_sbuf_printf(void *arg1, const char *fmt, ...)
{
va_list ap;
int result;
va_start(ap, fmt);
result = sbuf_vprintf(arg1, fmt, ap);
va_end(ap);
return (result);
}
void
lkpi_hex_dump(int(*_fpf)(void *, const char *, ...), void *arg1,
const char *level, const char *prefix_str,
const int prefix_type, const int rowsize, const int groupsize,
const void *buf, size_t len, const bool ascii, const bool trailing_newline)
{
typedef const struct { long long value; } __packed *print_64p_t;
typedef const struct { uint32_t value; } __packed *print_32p_t;
typedef const struct { uint16_t value; } __packed *print_16p_t;
const void *buf_old = buf;
int row, linelen, ret;
while (len > 0) {
linelen = 0;
if (level != NULL) {
ret = _fpf(arg1, "%s", level);
if (ret < 0)
break;
linelen += ret;
}
if (prefix_str != NULL) {
ret = _fpf(
arg1, "%s%s", linelen ? " " : "", prefix_str);
if (ret < 0)
break;
linelen += ret;
}
switch (prefix_type) {
case DUMP_PREFIX_ADDRESS:
ret = _fpf(
arg1, "%s[%p]", linelen ? " " : "", buf);
if (ret < 0)
return;
linelen += ret;
break;
case DUMP_PREFIX_OFFSET:
ret = _fpf(
arg1, "%s[%#tx]", linelen ? " " : "",
((const char *)buf - (const char *)buf_old));
if (ret < 0)
return;
linelen += ret;
break;
default:
break;
}
for (row = 0; row != rowsize; row++) {
if (groupsize == 8 && len > 7) {
ret = _fpf(
arg1, "%s%016llx", linelen ? " " : "",
((print_64p_t)buf)->value);
if (ret < 0)
return;
linelen += ret;
buf = (const uint8_t *)buf + 8;
len -= 8;
} else if (groupsize == 4 && len > 3) {
ret = _fpf(
arg1, "%s%08x", linelen ? " " : "",
((print_32p_t)buf)->value);
if (ret < 0)
return;
linelen += ret;
buf = (const uint8_t *)buf + 4;
len -= 4;
} else if (groupsize == 2 && len > 1) {
ret = _fpf(
arg1, "%s%04x", linelen ? " " : "",
((print_16p_t)buf)->value);
if (ret < 0)
return;
linelen += ret;
buf = (const uint8_t *)buf + 2;
len -= 2;
} else if (len > 0) {
ret = _fpf(
arg1, "%s%02x", linelen ? " " : "",
*(const uint8_t *)buf);
if (ret < 0)
return;
linelen += ret;
buf = (const uint8_t *)buf + 1;
len--;
} else {
break;
}
}
if (len > 0 && trailing_newline) {
ret = _fpf(arg1, "\n");
if (ret < 0)
break;
}
}
}
struct hdtb_context {
char *linebuf;
size_t linebuflen;
int written;
};
static int
hdtb_cb(void *arg, const char *format, ...)
{
struct hdtb_context *context;
int written;
va_list args;
context = arg;
va_start(args, format);
written = vsnprintf(
context->linebuf, context->linebuflen, format, args);
va_end(args);
if (written < 0)
return (written);
/*
* Linux' hex_dump_to_buffer() function has the same behaviour as
* snprintf() basically. Therefore, it returns the number of bytes it
* would have written if the destination buffer was large enough.
*
* If the destination buffer was exhausted, lkpi_hex_dump() will
* continue to call this callback but it will only compute the bytes it
* would have written but write nothing to that buffer.
*/
context->written += written;
if (written < context->linebuflen) {
context->linebuf += written;
context->linebuflen -= written;
} else {
context->linebuf += context->linebuflen;
context->linebuflen = 0;
}
return (written);
}
int
lkpi_hex_dump_to_buffer(const void *buf, size_t len, int rowsize,
int groupsize, char *linebuf, size_t linebuflen, bool ascii)
{
int written;
struct hdtb_context context;
context.linebuf = linebuf;
context.linebuflen = linebuflen;
context.written = 0;
if (rowsize != 16 && rowsize != 32)
rowsize = 16;
len = min(len, rowsize);
lkpi_hex_dump(
hdtb_cb, &context, NULL, NULL, DUMP_PREFIX_NONE,
rowsize, groupsize, buf, len, ascii, false);
written = context.written;
return (written);
}
static void
linux_timer_callback_wrapper(void *context)
{
struct timer_list *timer;
timer = context;
/* the timer is about to be shutdown permanently */
if (timer->function == NULL)
return;
if (linux_set_current_flags(curthread, M_NOWAIT)) {
/* try again later */
callout_reset(&timer->callout, 1,
&linux_timer_callback_wrapper, timer);
return;
}
timer->function(timer);
}
static int
linux_timer_jiffies_until(unsigned long expires)
{
unsigned long delta = expires - jiffies;
/*
* Guard against already expired values and make sure that the value can
* be used as a tick count, rather than a jiffies count.
*/
if ((long)delta < 1)
delta = 1;
else if (delta > INT_MAX)
delta = INT_MAX;
return ((int)delta);
}
int
mod_timer(struct timer_list *timer, unsigned long expires)
{
int ret;
timer->expires = expires;
ret = callout_reset(&timer->callout,
linux_timer_jiffies_until(expires),
&linux_timer_callback_wrapper, timer);
MPASS(ret == 0 || ret == 1);
return (ret == 1);
}
void
add_timer(struct timer_list *timer)
{
callout_reset(&timer->callout,
linux_timer_jiffies_until(timer->expires),
&linux_timer_callback_wrapper, timer);
}
void
add_timer_on(struct timer_list *timer, int cpu)
{
callout_reset_on(&timer->callout,
linux_timer_jiffies_until(timer->expires),
&linux_timer_callback_wrapper, timer, cpu);
}
int
timer_delete(struct timer_list *timer)
{
if (callout_stop(&(timer)->callout) == -1)
return (0);
return (1);
}
int
timer_delete_sync(struct timer_list *timer)
{
if (callout_drain(&(timer)->callout) == -1)
return (0);
return (1);
}
int
timer_shutdown_sync(struct timer_list *timer)
{
timer->function = NULL;
return (del_timer_sync(timer));
}
/* greatest common divisor, Euclid equation */
static uint64_t
lkpi_gcd_64(uint64_t a, uint64_t b)
{
uint64_t an;
uint64_t bn;
while (b != 0) {
an = b;
bn = a % b;
a = an;
b = bn;
}
return (a);
}
uint64_t lkpi_nsec2hz_rem;
uint64_t lkpi_nsec2hz_div = 1000000000ULL;
uint64_t lkpi_nsec2hz_max;
uint64_t lkpi_usec2hz_rem;
uint64_t lkpi_usec2hz_div = 1000000ULL;
uint64_t lkpi_usec2hz_max;
uint64_t lkpi_msec2hz_rem;
uint64_t lkpi_msec2hz_div = 1000ULL;
uint64_t lkpi_msec2hz_max;
static void
linux_timer_init(void *arg)
{
uint64_t gcd;
/*
* Compute an internal HZ value which can divide 2**32 to
* avoid timer rounding problems when the tick value wraps
* around 2**32:
*/
linux_timer_hz_mask = 1;
while (linux_timer_hz_mask < (unsigned long)hz)
linux_timer_hz_mask *= 2;
linux_timer_hz_mask--;
/* compute some internal constants */
lkpi_nsec2hz_rem = hz;
lkpi_usec2hz_rem = hz;
lkpi_msec2hz_rem = hz;
gcd = lkpi_gcd_64(lkpi_nsec2hz_rem, lkpi_nsec2hz_div);
lkpi_nsec2hz_rem /= gcd;
lkpi_nsec2hz_div /= gcd;
lkpi_nsec2hz_max = -1ULL / lkpi_nsec2hz_rem;
gcd = lkpi_gcd_64(lkpi_usec2hz_rem, lkpi_usec2hz_div);
lkpi_usec2hz_rem /= gcd;
lkpi_usec2hz_div /= gcd;
lkpi_usec2hz_max = -1ULL / lkpi_usec2hz_rem;
gcd = lkpi_gcd_64(lkpi_msec2hz_rem, lkpi_msec2hz_div);
lkpi_msec2hz_rem /= gcd;
lkpi_msec2hz_div /= gcd;
lkpi_msec2hz_max = -1ULL / lkpi_msec2hz_rem;
}
SYSINIT(linux_timer, SI_SUB_DRIVERS, SI_ORDER_FIRST, linux_timer_init, NULL);
void
linux_complete_common(struct completion *c, int all)
{
sleepq_lock(c);
if (all) {
c->done = UINT_MAX;
sleepq_broadcast(c, SLEEPQ_SLEEP, 0, 0);
} else {
if (c->done != UINT_MAX)
c->done++;
sleepq_signal(c, SLEEPQ_SLEEP, 0, 0);
}
sleepq_release(c);
}
/*
* Indefinite wait for done != 0 with or without signals.
*/
int
linux_wait_for_common(struct completion *c, int flags)
{
struct task_struct *task;
int error;
if (SCHEDULER_STOPPED())
return (0);
task = current;
if (flags != 0)
flags = SLEEPQ_INTERRUPTIBLE | SLEEPQ_SLEEP;
else
flags = SLEEPQ_SLEEP;
error = 0;
for (;;) {
sleepq_lock(c);
if (c->done)
break;
sleepq_add(c, NULL, "completion", flags, 0);
if (flags & SLEEPQ_INTERRUPTIBLE) {
DROP_GIANT();
error = -sleepq_wait_sig(c, 0);
PICKUP_GIANT();
if (error != 0) {
linux_schedule_save_interrupt_value(task, error);
error = -ERESTARTSYS;
goto intr;
}
} else {
DROP_GIANT();
sleepq_wait(c, 0);
PICKUP_GIANT();
}
}
if (c->done != UINT_MAX)
c->done--;
sleepq_release(c);
intr:
return (error);
}
/*
* Time limited wait for done != 0 with or without signals.
*/
unsigned long
linux_wait_for_timeout_common(struct completion *c, unsigned long timeout,
int flags)
{
struct task_struct *task;
unsigned long end = jiffies + timeout, error;
if (SCHEDULER_STOPPED())
return (0);
task = current;
if (flags != 0)
flags = SLEEPQ_INTERRUPTIBLE | SLEEPQ_SLEEP;
else
flags = SLEEPQ_SLEEP;
for (;;) {
sleepq_lock(c);
if (c->done)
break;
sleepq_add(c, NULL, "completion", flags, 0);
sleepq_set_timeout(c, linux_timer_jiffies_until(end));
DROP_GIANT();
if (flags & SLEEPQ_INTERRUPTIBLE)
error = -sleepq_timedwait_sig(c, 0);
else
error = -sleepq_timedwait(c, 0);
PICKUP_GIANT();
if (error != 0) {
/* check for timeout */
if (error == -EWOULDBLOCK) {
error = 0; /* timeout */
} else {
/* signal happened */
linux_schedule_save_interrupt_value(task, error);
error = -ERESTARTSYS;
}
goto done;
}
}
if (c->done != UINT_MAX)
c->done--;
sleepq_release(c);
/* return how many jiffies are left */
error = linux_timer_jiffies_until(end);
done:
return (error);
}
int
linux_try_wait_for_completion(struct completion *c)
{
int isdone;
sleepq_lock(c);
isdone = (c->done != 0);
if (c->done != 0 && c->done != UINT_MAX)
c->done--;
sleepq_release(c);
return (isdone);
}
int
linux_completion_done(struct completion *c)
{
int isdone;
sleepq_lock(c);
isdone = (c->done != 0);
sleepq_release(c);
return (isdone);
}
static void
linux_cdev_deref(struct linux_cdev *ldev)
{
if (refcount_release(&ldev->refs) &&
ldev->kobj.ktype == &linux_cdev_ktype)
kfree(ldev);
}
static void
linux_cdev_release(struct kobject *kobj)
{
struct linux_cdev *cdev;
struct kobject *parent;
cdev = container_of(kobj, struct linux_cdev, kobj);
parent = kobj->parent;
linux_destroy_dev(cdev);
linux_cdev_deref(cdev);
kobject_put(parent);
}
static void
linux_cdev_static_release(struct kobject *kobj)
{
struct cdev *cdev;
struct linux_cdev *ldev;
ldev = container_of(kobj, struct linux_cdev, kobj);
cdev = ldev->cdev;
if (cdev != NULL) {
destroy_dev(cdev);
ldev->cdev = NULL;
}
kobject_put(kobj->parent);
}
int
linux_cdev_device_add(struct linux_cdev *ldev, struct device *dev)
{
int ret;
if (dev->devt != 0) {
/* Set parent kernel object. */
ldev->kobj.parent = &dev->kobj;
/*
* Unlike Linux we require the kobject of the
* character device structure to have a valid name
* before calling this function:
*/
if (ldev->kobj.name == NULL)
return (-EINVAL);
ret = cdev_add(ldev, dev->devt, 1);
if (ret)
return (ret);
}
ret = device_add(dev);
if (ret != 0 && dev->devt != 0)
cdev_del(ldev);
return (ret);
}
void
linux_cdev_device_del(struct linux_cdev *ldev, struct device *dev)
{
device_del(dev);
if (dev->devt != 0)
cdev_del(ldev);
}
static void
linux_destroy_dev(struct linux_cdev *ldev)
{
if (ldev->cdev == NULL)
return;
MPASS((ldev->siref & LDEV_SI_DTR) == 0);
MPASS(ldev->kobj.ktype == &linux_cdev_ktype);
atomic_set_int(&ldev->siref, LDEV_SI_DTR);
while ((atomic_load_int(&ldev->siref) & ~LDEV_SI_DTR) != 0)
pause("ldevdtr", hz / 4);
destroy_dev(ldev->cdev);
ldev->cdev = NULL;
}
const struct kobj_type linux_cdev_ktype = {
.release = linux_cdev_release,
};
const struct kobj_type linux_cdev_static_ktype = {
.release = linux_cdev_static_release,
};
static void
linux_handle_ifnet_link_event(void *arg, struct ifnet *ifp, int linkstate)
{
struct notifier_block *nb;
struct netdev_notifier_info ni;
nb = arg;
ni.ifp = ifp;
ni.dev = (struct net_device *)ifp;
if (linkstate == LINK_STATE_UP)
nb->notifier_call(nb, NETDEV_UP, &ni);
else
nb->notifier_call(nb, NETDEV_DOWN, &ni);
}
static void
linux_handle_ifnet_arrival_event(void *arg, struct ifnet *ifp)
{
struct notifier_block *nb;
struct netdev_notifier_info ni;
nb = arg;
ni.ifp = ifp;
ni.dev = (struct net_device *)ifp;
nb->notifier_call(nb, NETDEV_REGISTER, &ni);
}
static void
linux_handle_ifnet_departure_event(void *arg, struct ifnet *ifp)
{
struct notifier_block *nb;
struct netdev_notifier_info ni;
nb = arg;
ni.ifp = ifp;
ni.dev = (struct net_device *)ifp;
nb->notifier_call(nb, NETDEV_UNREGISTER, &ni);
}
static void
linux_handle_iflladdr_event(void *arg, struct ifnet *ifp)
{
struct notifier_block *nb;
struct netdev_notifier_info ni;
nb = arg;
ni.ifp = ifp;
ni.dev = (struct net_device *)ifp;
nb->notifier_call(nb, NETDEV_CHANGEADDR, &ni);
}
static void
linux_handle_ifaddr_event(void *arg, struct ifnet *ifp)
{
struct notifier_block *nb;
struct netdev_notifier_info ni;
nb = arg;
ni.ifp = ifp;
ni.dev = (struct net_device *)ifp;
nb->notifier_call(nb, NETDEV_CHANGEIFADDR, &ni);
}
int
register_netdevice_notifier(struct notifier_block *nb)
{
nb->tags[NETDEV_UP] = EVENTHANDLER_REGISTER(
ifnet_link_event, linux_handle_ifnet_link_event, nb, 0);
nb->tags[NETDEV_REGISTER] = EVENTHANDLER_REGISTER(
ifnet_arrival_event, linux_handle_ifnet_arrival_event, nb, 0);
nb->tags[NETDEV_UNREGISTER] = EVENTHANDLER_REGISTER(
ifnet_departure_event, linux_handle_ifnet_departure_event, nb, 0);
nb->tags[NETDEV_CHANGEADDR] = EVENTHANDLER_REGISTER(
iflladdr_event, linux_handle_iflladdr_event, nb, 0);
return (0);
}
int
register_inetaddr_notifier(struct notifier_block *nb)
{
nb->tags[NETDEV_CHANGEIFADDR] = EVENTHANDLER_REGISTER(
ifaddr_event, linux_handle_ifaddr_event, nb, 0);
return (0);
}
int
unregister_netdevice_notifier(struct notifier_block *nb)
{
EVENTHANDLER_DEREGISTER(ifnet_link_event,
nb->tags[NETDEV_UP]);
EVENTHANDLER_DEREGISTER(ifnet_arrival_event,
nb->tags[NETDEV_REGISTER]);
EVENTHANDLER_DEREGISTER(ifnet_departure_event,
nb->tags[NETDEV_UNREGISTER]);
EVENTHANDLER_DEREGISTER(iflladdr_event,
nb->tags[NETDEV_CHANGEADDR]);
return (0);
}
int
unregister_inetaddr_notifier(struct notifier_block *nb)
{
EVENTHANDLER_DEREGISTER(ifaddr_event,
nb->tags[NETDEV_CHANGEIFADDR]);
return (0);
}
struct list_sort_thunk {
int (*cmp)(void *, struct list_head *, struct list_head *);
void *priv;
};
static inline int
linux_le_cmp(const void *d1, const void *d2, void *priv)
{
struct list_head *le1, *le2;
struct list_sort_thunk *thunk;
thunk = priv;
le1 = *(__DECONST(struct list_head **, d1));
le2 = *(__DECONST(struct list_head **, d2));
return ((thunk->cmp)(thunk->priv, le1, le2));
}
void
list_sort(void *priv, struct list_head *head, int (*cmp)(void *priv,
struct list_head *a, struct list_head *b))
{
struct list_sort_thunk thunk;
struct list_head **ar, *le;
size_t count, i;
count = 0;
list_for_each(le, head)
count++;
ar = malloc(sizeof(struct list_head *) * count, M_KMALLOC, M_WAITOK);
i = 0;
list_for_each(le, head)
ar[i++] = le;
thunk.cmp = cmp;
thunk.priv = priv;
qsort_r(ar, count, sizeof(struct list_head *), linux_le_cmp, &thunk);
INIT_LIST_HEAD(head);
for (i = 0; i < count; i++)
list_add_tail(ar[i], head);
free(ar, M_KMALLOC);
}
#if defined(__i386__) || defined(__amd64__)
int
linux_wbinvd_on_all_cpus(void)
{
pmap_invalidate_cache();
return (0);
}
#endif
int
linux_on_each_cpu(void callback(void *), void *data)
{
smp_rendezvous(smp_no_rendezvous_barrier, callback,
smp_no_rendezvous_barrier, data);
return (0);
}
int
linux_in_atomic(void)
{
return ((curthread->td_pflags & TDP_NOFAULTING) != 0);
}
struct linux_cdev *
linux_find_cdev(const char *name, unsigned major, unsigned minor)
{
dev_t dev = MKDEV(major, minor);
struct cdev *cdev;
dev_lock();
LIST_FOREACH(cdev, &linuxcdevsw.d_devs, si_list) {
struct linux_cdev *ldev = cdev->si_drv1;
if (ldev->dev == dev &&
strcmp(kobject_name(&ldev->kobj), name) == 0) {
break;
}
}
dev_unlock();
return (cdev != NULL ? cdev->si_drv1 : NULL);
}
int
__register_chrdev(unsigned int major, unsigned int baseminor,
unsigned int count, const char *name,
const struct file_operations *fops)
{
struct linux_cdev *cdev;
int ret = 0;
int i;
for (i = baseminor; i < baseminor + count; i++) {
cdev = cdev_alloc();
cdev->ops = fops;
kobject_set_name(&cdev->kobj, name);
ret = cdev_add(cdev, makedev(major, i), 1);
if (ret != 0)
break;
}
return (ret);
}
int
__register_chrdev_p(unsigned int major, unsigned int baseminor,
unsigned int count, const char *name,
const struct file_operations *fops, uid_t uid,
gid_t gid, int mode)
{
struct linux_cdev *cdev;
int ret = 0;
int i;
for (i = baseminor; i < baseminor + count; i++) {
cdev = cdev_alloc();
cdev->ops = fops;
kobject_set_name(&cdev->kobj, name);
ret = cdev_add_ext(cdev, makedev(major, i), uid, gid, mode);
if (ret != 0)
break;
}
return (ret);
}
void
__unregister_chrdev(unsigned int major, unsigned int baseminor,
unsigned int count, const char *name)
{
struct linux_cdev *cdevp;
int i;
for (i = baseminor; i < baseminor + count; i++) {
cdevp = linux_find_cdev(name, major, i);
if (cdevp != NULL)
cdev_del(cdevp);
}
}
void
linux_dump_stack(void)
{
#ifdef STACK
struct stack st;
stack_save(&st);
stack_print(&st);
#endif
}
int
linuxkpi_net_ratelimit(void)
{
return (ppsratecheck(&lkpi_net_lastlog, &lkpi_net_curpps,
lkpi_net_maxpps));
}
struct io_mapping *
io_mapping_create_wc(resource_size_t base, unsigned long size)
{
struct io_mapping *mapping;
mapping = kmalloc(sizeof(*mapping), GFP_KERNEL);
if (mapping == NULL)
return (NULL);
return (io_mapping_init_wc(mapping, base, size));
}
/* We likely want a linuxkpi_device.c at some point. */
bool
device_can_wakeup(struct device *dev)
{
if (dev == NULL)
return (false);
/*
* XXX-BZ iwlwifi queries it as part of enabling WoWLAN.
* Normally this would be based on a bool in dev->power.XXX.
* Check such as PCI PCIM_PCAP_*PME. We have no way to enable this yet.
* We may get away by directly calling into bsddev for as long as
* we can assume PCI only avoiding changing struct device breaking KBI.
*/
pr_debug("%s:%d: not enabled; see comment.\n", __func__, __LINE__);
return (false);
}
void
linuxkpi_device_set_wakeup_capable(struct device *dev, bool capable)
{
dev->power.can_wakeup = capable;
}
static void
devm_device_group_remove(struct device *dev, void *p)
{
const struct attribute_group **dr = p;
const struct attribute_group *group = *dr;
sysfs_remove_group(&dev->kobj, group);
}
int
lkpi_devm_device_add_group(struct device *dev,
const struct attribute_group *group)
{
const struct attribute_group **dr;
int ret;
dr = devres_alloc(devm_device_group_remove, sizeof(*dr), GFP_KERNEL);
if (dr == NULL)
return (-ENOMEM);
ret = sysfs_create_group(&dev->kobj, group);
if (ret == 0) {
*dr = group;
devres_add(dev, dr);
} else
devres_free(dr);
return (ret);
}
#if defined(__i386__) || defined(__amd64__)
bool linux_cpu_has_clflush;
struct cpuinfo_x86 boot_cpu_data;
struct cpuinfo_x86 *__cpu_data;
#endif
cpumask_t *
lkpi_get_static_single_cpu_mask(int cpuid)
{
KASSERT((cpuid >= 0 && cpuid <= mp_maxid), ("%s: invalid cpuid %d\n",
__func__, cpuid));
KASSERT(!CPU_ABSENT(cpuid), ("%s: cpu with cpuid %d is absent\n",
__func__, cpuid));
return (static_single_cpu_mask[cpuid]);
}
bool
lkpi_xen_initial_domain(void)
{
#ifdef XENHVM
return (xen_initial_domain());
#else
return (false);
#endif
}
bool
lkpi_xen_pv_domain(void)
{
#ifdef XENHVM
return (xen_pv_domain());
#else
return (false);
#endif
}
static void
linux_compat_init(void *arg)
{
struct sysctl_oid *rootoid;
int i;
#if defined(__i386__) || defined(__amd64__)
static const uint32_t x86_vendors[X86_VENDOR_NUM] = {
[X86_VENDOR_INTEL] = CPU_VENDOR_INTEL,
[X86_VENDOR_CYRIX] = CPU_VENDOR_CYRIX,
[X86_VENDOR_AMD] = CPU_VENDOR_AMD,
[X86_VENDOR_UMC] = CPU_VENDOR_UMC,
[X86_VENDOR_CENTAUR] = CPU_VENDOR_CENTAUR,
[X86_VENDOR_TRANSMETA] = CPU_VENDOR_TRANSMETA,
[X86_VENDOR_NSC] = CPU_VENDOR_NSC,
[X86_VENDOR_HYGON] = CPU_VENDOR_HYGON,
};
uint8_t x86_vendor = X86_VENDOR_UNKNOWN;
for (i = 0; i < X86_VENDOR_NUM; i++) {
if (cpu_vendor_id != 0 && cpu_vendor_id == x86_vendors[i]) {
x86_vendor = i;
break;
}
}
linux_cpu_has_clflush = (cpu_feature & CPUID_CLFSH);
boot_cpu_data.x86_clflush_size = cpu_clflush_line_size;
boot_cpu_data.x86_max_cores = mp_ncpus;
boot_cpu_data.x86 = CPUID_TO_FAMILY(cpu_id);
boot_cpu_data.x86_model = CPUID_TO_MODEL(cpu_id);
boot_cpu_data.x86_vendor = x86_vendor;
boot_cpu_data.x86_stepping = CPUID_TO_STEPPING(cpu_id);
__cpu_data = kmalloc_array(mp_maxid + 1,
sizeof(*__cpu_data), M_WAITOK | M_ZERO);
CPU_FOREACH(i) {
__cpu_data[i].x86_clflush_size = cpu_clflush_line_size;
__cpu_data[i].x86_max_cores = mp_ncpus;
__cpu_data[i].x86 = CPUID_TO_FAMILY(cpu_id);
__cpu_data[i].x86_model = CPUID_TO_MODEL(cpu_id);
__cpu_data[i].x86_vendor = x86_vendor;
}
#endif
rw_init(&linux_vma_lock, "lkpi-vma-lock");
rootoid = SYSCTL_ADD_ROOT_NODE(NULL,
OID_AUTO, "sys", CTLFLAG_RD|CTLFLAG_MPSAFE, NULL, "sys");
kobject_init(&linux_class_root, &linux_class_ktype);
kobject_set_name(&linux_class_root, "class");
linux_class_root.oidp = SYSCTL_ADD_NODE(NULL, SYSCTL_CHILDREN(rootoid),
OID_AUTO, "class", CTLFLAG_RD|CTLFLAG_MPSAFE, NULL, "class");
kobject_init(&linux_root_device.kobj, &linux_dev_ktype);
kobject_set_name(&linux_root_device.kobj, "device");
linux_root_device.kobj.oidp = SYSCTL_ADD_NODE(NULL,
SYSCTL_CHILDREN(rootoid), OID_AUTO, "device",
CTLFLAG_RD | CTLFLAG_MPSAFE, NULL, "device");
linux_root_device.bsddev = root_bus;
linux_class_misc.name = "misc";
class_register(&linux_class_misc);
INIT_LIST_HEAD(&pci_drivers);
INIT_LIST_HEAD(&pci_devices);
spin_lock_init(&pci_lock);
init_waitqueue_head(&linux_bit_waitq);
init_waitqueue_head(&linux_var_waitq);
CPU_COPY(&all_cpus, &cpu_online_mask);
/*
* Generate a single-CPU cpumask_t for each CPU (possibly) in the system.
* CPUs are indexed from 0..(mp_maxid). The entry for cpuid 0 will only
* have itself in the cpumask, cupid 1 only itself on entry 1, and so on.
* This is used by cpumask_of() (and possibly others in the future) for,
* e.g., drivers to pass hints to irq_set_affinity_hint().
*/
static_single_cpu_mask = kmalloc_array(mp_maxid + 1,
sizeof(static_single_cpu_mask), M_WAITOK | M_ZERO);
/*
* When the number of CPUs reach a threshold, we start to save memory
* given the sets are static by overlapping those having their single
* bit set at same position in a bitset word. Asymptotically, this
* regular scheme is in O(n²) whereas the overlapping one is in O(n)
* only with n being the maximum number of CPUs, so the gain will become
* huge quite quickly. The threshold for 64-bit architectures is 128
* CPUs.
*/
if (mp_ncpus < (2 * _BITSET_BITS)) {
cpumask_t *sscm_ptr;
/*
* This represents 'mp_ncpus * __bitset_words(CPU_SETSIZE) *
* (_BITSET_BITS / 8)' bytes (for comparison with the
* overlapping scheme).
*/
static_single_cpu_mask_lcs = kmalloc_array(mp_ncpus,
sizeof(*static_single_cpu_mask_lcs),
M_WAITOK | M_ZERO);
sscm_ptr = static_single_cpu_mask_lcs;
CPU_FOREACH(i) {
static_single_cpu_mask[i] = sscm_ptr++;
CPU_SET(i, static_single_cpu_mask[i]);
}
} else {
/* Pointer to a bitset word. */
__typeof(((cpuset_t *)NULL)->__bits[0]) *bwp;
/*
* Allocate memory for (static) spans of 'cpumask_t' ('cpuset_t'
* really) with a single bit set that can be reused for all
* single CPU masks by making them start at different offsets.
* We need '__bitset_words(CPU_SETSIZE) - 1' bitset words before
* the word having its single bit set, and the same amount
* after.
*/
static_single_cpu_mask_lcs = mallocarray(_BITSET_BITS,
(2 * __bitset_words(CPU_SETSIZE) - 1) * (_BITSET_BITS / 8),
M_KMALLOC, M_WAITOK | M_ZERO);
/*
* We rely below on cpuset_t and the bitset generic
* implementation assigning words in the '__bits' array in the
* same order of bits (i.e., little-endian ordering, not to be
* confused with machine endianness, which concerns bits in
* words and other integers). This is an imperfect test, but it
* will detect a change to big-endian ordering.
*/
_Static_assert(
__bitset_word(_BITSET_BITS + 1, _BITSET_BITS) == 1,
"Assumes a bitset implementation that is little-endian "
"on its words");
/* Initialize the single bit of each static span. */
bwp = (__typeof(bwp))static_single_cpu_mask_lcs +
(__bitset_words(CPU_SETSIZE) - 1);
for (i = 0; i < _BITSET_BITS; i++) {
CPU_SET(i, (cpuset_t *)bwp);
bwp += (2 * __bitset_words(CPU_SETSIZE) - 1);
}
/*
* Finally set all CPU masks to the proper word in their
* relevant span.
*/
CPU_FOREACH(i) {
bwp = (__typeof(bwp))static_single_cpu_mask_lcs;
/* Find the non-zero word of the relevant span. */
bwp += (2 * __bitset_words(CPU_SETSIZE) - 1) *
(i % _BITSET_BITS) +
__bitset_words(CPU_SETSIZE) - 1;
/* Shift to find the CPU mask start. */
bwp -= (i / _BITSET_BITS);
static_single_cpu_mask[i] = (cpuset_t *)bwp;
}
}
strlcpy(init_uts_ns.name.release, osrelease, sizeof(init_uts_ns.name.release));
}
SYSINIT(linux_compat, SI_SUB_DRIVERS, SI_ORDER_SECOND, linux_compat_init, NULL);
static void
linux_compat_uninit(void *arg)
{
linux_kobject_kfree_name(&linux_class_root);
linux_kobject_kfree_name(&linux_root_device.kobj);
linux_kobject_kfree_name(&linux_class_misc.kobj);
free(static_single_cpu_mask_lcs, M_KMALLOC);
free(static_single_cpu_mask, M_KMALLOC);
#if defined(__i386__) || defined(__amd64__)
free(__cpu_data, M_KMALLOC);
#endif
spin_lock_destroy(&pci_lock);
rw_destroy(&linux_vma_lock);
}
SYSUNINIT(linux_compat, SI_SUB_DRIVERS, SI_ORDER_SECOND, linux_compat_uninit, NULL);
#if defined(__i386__) || defined(__amd64__)
const struct x86_cpu_id *
linuxkpi_x86_match_cpu(const struct x86_cpu_id *match_array)
{
const struct x86_cpu_id *match;
for (match = match_array;
(match->flags & X86_CPU_ID_FLAG_ENTRY_VALID) != 0;
match++) {
if (match->vendor != X86_VENDOR_ANY &&
match->vendor != boot_cpu_data.x86_vendor)
continue;
if (match->family != X86_FAMILY_ANY &&
match->family != boot_cpu_data.x86)
continue;
if (match->model != X86_MODEL_ANY &&
match->model != boot_cpu_data.x86_model)
continue;
if (match->model != X86_STEPPING_ANY &&
(match->steppings & BIT(boot_cpu_data.x86_stepping)) == 0)
continue;
if (match->feature != X86_FEATURE_ANY &&
!static_cpu_has(match->feature))
continue;
return (match);
}
return (NULL);
}
#endif
/*
* NOTE: Linux frequently uses "unsigned long" for pointer to integer
* conversion and vice versa, where in FreeBSD "uintptr_t" would be
* used. Assert these types have the same size, else some parts of the
* LinuxKPI may not work like expected:
*/
CTASSERT(sizeof(unsigned long) == sizeof(uintptr_t));
diff --git a/sys/compat/linuxkpi/common/src/linux_page.c b/sys/compat/linuxkpi/common/src/linux_page.c
index a95d86753985876b4f401427ccd0547cdfe720e7..e22cbe5d1605202802b0db774d2cd68e3830042b 100644
--- a/sys/compat/linuxkpi/common/src/linux_page.c
+++ b/sys/compat/linuxkpi/common/src/linux_page.c
@@ -1,830 +1,1697 @@
/*-
* Copyright (c) 2010 Isilon Systems, Inc.
* Copyright (c) 2016 Matthew Macy (mmacy@mattmacy.io)
* Copyright (c) 2017 Mellanox Technologies, Ltd.
* All rights reserved.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice unmodified, this list of conditions, and the following
* disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
*
* THIS SOFTWARE IS PROVIDED BY THE AUTHOR ``AS IS'' AND ANY EXPRESS OR
* IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES
* OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED.
* IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR ANY DIRECT, INDIRECT,
* INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT
* NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE,
* DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY
* THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
* (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF
* THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
*/
#include <sys/param.h>
#include <sys/systm.h>
#include <sys/malloc.h>
#include <sys/kernel.h>
#include <sys/sysctl.h>
#include <sys/lock.h>
#include <sys/mutex.h>
#include <sys/rwlock.h>
+#include <sys/sx.h>
+#include <sys/tree.h>
#include <sys/proc.h>
#include <sys/sched.h>
#include <sys/memrange.h>
#include <machine/bus.h>
#include <vm/vm.h>
#include <vm/pmap.h>
#include <vm/vm_param.h>
#include <vm/vm_kern.h>
#include <vm/vm_object.h>
#include <vm/vm_map.h>
#include <vm/vm_page.h>
#include <vm/vm_pageout.h>
#include <vm/vm_pager.h>
#include <vm/vm_radix.h>
#include <vm/vm_reserv.h>
#include <vm/vm_extern.h>
#include <vm/uma.h>
#include <vm/uma_int.h>
+/*
+ * One pager-owned wire per physical page, independent of driver wiring.
+ * Generate the native tree before Linux headers redefine RB_ROOT.
+ */
+struct lkpi_vma_pfn_pin {
+ RB_ENTRY(lkpi_vma_pfn_pin) link;
+ vm_page_t m;
+};
+
+RB_HEAD(lkpi_vma_pfn_pins, lkpi_vma_pfn_pin);
+
+static int
+lkpi_vma_pfn_pin_cmp(struct lkpi_vma_pfn_pin *a,
+ struct lkpi_vma_pfn_pin *b)
+{
+
+ return ((uintptr_t)a->m < (uintptr_t)b->m ? -1 :
+ (uintptr_t)a->m > (uintptr_t)b->m);
+}
+
+RB_GENERATE_STATIC(lkpi_vma_pfn_pins, lkpi_vma_pfn_pin, link,
+ lkpi_vma_pfn_pin_cmp);
+
#include <linux/gfp.h>
#include <linux/mm.h>
#include <linux/preempt.h>
#include <linux/fs.h>
#include <linux/shmem_fs.h>
#include <linux/kernel.h>
#include <linux/idr.h>
#include <linux/io.h>
#include <linux/io-mapping.h>
#include <linux/device.h>
+#include <linux/slab.h>
#ifdef __i386__
DEFINE_IDR(mtrr_idr);
static MALLOC_DEFINE(M_LKMTRR, "idr", "Linux MTRR compat");
extern int pat_works;
#endif
void
si_meminfo(struct sysinfo *si)
{
si->totalram = physmem;
si->freeram = vm_free_count();
si->totalhigh = 0;
si->freehigh = 0;
si->mem_unit = PAGE_SIZE;
}
void *
linux_page_address(const struct page *page)
{
if (page->object != kernel_object) {
return (PMAP_HAS_DMAP ? PHYS_TO_DMAP(page_to_phys(page)) :
NULL);
}
return ((void *)(uintptr_t)(VM_MIN_KERNEL_ADDRESS +
IDX_TO_OFF(page->pindex)));
}
struct page *
linux_alloc_pages(gfp_t flags, unsigned int order)
{
struct page *page;
if (PMAP_HAS_DMAP) {
unsigned long npages = 1UL << order;
int req = VM_ALLOC_WIRED;
if ((flags & M_ZERO) != 0)
req |= VM_ALLOC_ZERO;
if (order == 0 && (flags & GFP_DMA32) == 0) {
page = vm_page_alloc_noobj(req);
if (page == NULL)
return (NULL);
} else {
vm_paddr_t pmax = (flags & GFP_DMA32) ?
BUS_SPACE_MAXADDR_32BIT : BUS_SPACE_MAXADDR;
if ((flags & __GFP_NORETRY) != 0)
req |= VM_ALLOC_NORECLAIM;
retry:
if ((flags & __GFP_THISNODE) != 0) {
int curdomain = PCPU_GET(domain);
page = vm_page_alloc_noobj_contig_domain(
curdomain, req, npages, 0, pmax,
PAGE_SIZE, 0, VM_MEMATTR_DEFAULT);
} else {
page = vm_page_alloc_noobj_contig(
req, npages, 0, pmax,
PAGE_SIZE, 0, VM_MEMATTR_DEFAULT);
}
if (page == NULL) {
if ((flags & (M_WAITOK | __GFP_NORETRY | __GFP_THISNODE)) ==
M_WAITOK) {
int err = vm_page_reclaim_contig(req,
npages, 0, pmax, PAGE_SIZE, 0);
if (err == ENOMEM)
vm_wait(NULL);
else if (err != 0)
return (NULL);
flags &= ~M_WAITOK;
goto retry;
}
return (NULL);
}
}
} else {
vm_offset_t vaddr;
vaddr = linux_alloc_kmem(flags, order);
if (vaddr == 0)
return (NULL);
page = virt_to_page((void *)vaddr);
KASSERT(vaddr == (vm_offset_t)page_address(page),
("Page address mismatch"));
}
return (page);
}
static void
_linux_free_kmem(vm_offset_t addr, unsigned int order)
{
size_t size = ((size_t)PAGE_SIZE) << order;
kmem_free((void *)addr, size);
}
void
linux_free_pages(struct page *page, unsigned int order)
{
if (PMAP_HAS_DMAP) {
unsigned long npages = 1UL << order;
unsigned long x;
for (x = 0; x != npages; x++) {
vm_page_t pgo = page + x;
/*
* The "free page" function is used in several
* contexts.
*
* Some pages are allocated by `linux_alloc_pages()`
* above, but not all of them are. For instance in the
* DRM drivers, some pages come from
* `shmem_read_mapping_page_gfp()`.
*
* That's why we need to check if the page is managed
* or not here.
*/
if ((pgo->oflags & VPO_UNMANAGED) == 0) {
vm_page_unwire(pgo, PQ_ACTIVE);
} else {
if (vm_page_unwire_noq(pgo))
vm_page_free(pgo);
}
}
} else {
vm_offset_t vaddr;
vaddr = (vm_offset_t)page_address(page);
_linux_free_kmem(vaddr, order);
}
}
void
linux_release_pages(release_pages_arg arg, int nr)
{
int i;
CTASSERT(offsetof(struct folio, page) == 0);
for (i = 0; i < nr; i++)
__free_page(arg.pages[i]);
}
vm_offset_t
linux_alloc_kmem(gfp_t flags, unsigned int order)
{
size_t size = ((size_t)PAGE_SIZE) << order;
void *addr;
addr = kmem_alloc_contig(size, flags & GFP_NATIVE_MASK, 0,
((flags & GFP_DMA32) == 0) ? -1UL : BUS_SPACE_MAXADDR_32BIT,
PAGE_SIZE, 0, VM_MEMATTR_DEFAULT);
return ((vm_offset_t)addr);
}
void
linux_free_kmem(vm_offset_t addr, unsigned int order)
{
KASSERT((addr & ~PAGE_MASK) == 0,
("%s: addr %p is not page aligned", __func__, (void *)addr));
if (addr >= VM_MIN_KERNEL_ADDRESS && addr < VM_MAX_KERNEL_ADDRESS) {
_linux_free_kmem(addr, order);
} else {
vm_page_t page;
page = DMAP_TO_VM_PAGE(addr);
linux_free_pages(page, order);
}
}
static int
linux_get_user_pages_internal(vm_map_t map, unsigned long start, int nr_pages,
int write, struct page **pages)
{
vm_prot_t prot;
size_t len;
int count;
prot = write ? (VM_PROT_READ | VM_PROT_WRITE) : VM_PROT_READ;
len = ptoa((vm_offset_t)nr_pages);
count = vm_fault_quick_hold_pages(map, start, len, prot, pages, nr_pages);
return (count == -1 ? -EFAULT : nr_pages);
}
int
__get_user_pages_fast(unsigned long start, int nr_pages, int write,
struct page **pages)
{
vm_map_t map;
vm_page_t *mp;
vm_offset_t va;
vm_offset_t end;
vm_prot_t prot;
int count;
if (nr_pages == 0 || in_interrupt())
return (0);
MPASS(pages != NULL);
map = &curthread->td_proc->p_vmspace->vm_map;
end = start + ptoa((vm_offset_t)nr_pages);
if (!vm_map_range_valid(map, start, end))
return (-EINVAL);
prot = write ? (VM_PROT_READ | VM_PROT_WRITE) : VM_PROT_READ;
for (count = 0, mp = pages, va = start; va < end;
mp++, va += PAGE_SIZE, count++) {
*mp = pmap_extract_and_hold(map->pmap, va, prot);
if (*mp == NULL)
break;
if ((prot & VM_PROT_WRITE) != 0 &&
(*mp)->dirty != VM_PAGE_BITS_ALL) {
/*
* Explicitly dirty the physical page. Otherwise, the
* caller's changes may go unnoticed because they are
* performed through an unmanaged mapping or by a DMA
* operation.
*
* The object lock is not held here.
* See vm_page_clear_dirty_mask().
*/
vm_page_dirty(*mp);
}
}
return (count);
}
long
get_user_pages_remote(struct task_struct *task, struct mm_struct *mm,
unsigned long start, unsigned long nr_pages, unsigned int gup_flags,
struct page **pages, struct vm_area_struct **vmas)
{
vm_map_t map;
map = &task->task_thread->td_proc->p_vmspace->vm_map;
return (linux_get_user_pages_internal(map, start, nr_pages,
!!(gup_flags & FOLL_WRITE), pages));
}
long
lkpi_get_user_pages(unsigned long start, unsigned long nr_pages,
unsigned int gup_flags, struct page **pages)
{
vm_map_t map;
map = &curthread->td_proc->p_vmspace->vm_map;
return (linux_get_user_pages_internal(map, start, nr_pages,
!!(gup_flags & FOLL_WRITE), pages));
}
/*
* Hash of vmmap addresses. This is infrequently accessed and does not
* need to be particularly large. This is done because we must store the
* caller's idea of the map size to properly unmap.
*/
struct vmmap {
LIST_ENTRY(vmmap) vm_next;
void *vm_addr;
unsigned long vm_size;
};
struct vmmaphd {
struct vmmap *lh_first;
};
#define VMMAP_HASH_SIZE 64
#define VMMAP_HASH_MASK (VMMAP_HASH_SIZE - 1)
#define VM_HASH(addr) ((uintptr_t)(addr) >> PAGE_SHIFT) & VMMAP_HASH_MASK
static struct vmmaphd vmmaphead[VMMAP_HASH_SIZE];
static struct mtx vmmaplock;
int
is_vmalloc_addr(const void *addr)
{
struct vmmap *vmmap;
mtx_lock(&vmmaplock);
LIST_FOREACH(vmmap, &vmmaphead[VM_HASH(addr)], vm_next)
if (addr == vmmap->vm_addr)
break;
mtx_unlock(&vmmaplock);
if (vmmap != NULL)
return (1);
return (vtoslab((vm_offset_t)addr & ~UMA_SLAB_MASK) != NULL);
}
static void
vmmap_add(void *addr, unsigned long size)
{
struct vmmap *vmmap;
vmmap = kmalloc(sizeof(*vmmap), GFP_KERNEL);
mtx_lock(&vmmaplock);
vmmap->vm_size = size;
vmmap->vm_addr = addr;
LIST_INSERT_HEAD(&vmmaphead[VM_HASH(addr)], vmmap, vm_next);
mtx_unlock(&vmmaplock);
}
static struct vmmap *
vmmap_remove(void *addr)
{
struct vmmap *vmmap;
mtx_lock(&vmmaplock);
LIST_FOREACH(vmmap, &vmmaphead[VM_HASH(addr)], vm_next)
if (vmmap->vm_addr == addr)
break;
if (vmmap)
LIST_REMOVE(vmmap, vm_next);
mtx_unlock(&vmmaplock);
return (vmmap);
}
#if defined(__i386__) || defined(__amd64__) || defined(__powerpc__) || defined(__aarch64__) || defined(__riscv)
void *
_ioremap_attr(vm_paddr_t phys_addr, unsigned long size, int attr)
{
void *addr;
addr = pmap_mapdev_attr(phys_addr, size, attr);
if (addr == NULL)
return (NULL);
vmmap_add(addr, size);
return (addr);
}
#endif
void
iounmap(void *addr)
{
struct vmmap *vmmap;
vmmap = vmmap_remove(addr);
if (vmmap == NULL)
return;
#if defined(__i386__) || defined(__amd64__) || defined(__powerpc__) || defined(__aarch64__) || defined(__riscv)
pmap_unmapdev(addr, vmmap->vm_size);
#endif
kfree(vmmap);
}
static void
lkpi_devm_memremap_unmap(struct device *dev, void *p)
{
void **dr = p;
memunmap(*dr);
}
void *
linuxkpi_devm_memremap(struct device *dev, resource_size_t offset, size_t size,
unsigned long flags)
{
void **dr, *addr;
dr = devres_alloc(lkpi_devm_memremap_unmap, sizeof(*dr), GFP_KERNEL);
if (dr == NULL)
return (ERR_PTR(-ENOMEM));
addr = memremap(offset, size, flags);
if (addr != NULL) {
*dr = addr;
devres_add(dev, dr);
} else {
addr = ERR_PTR(-ENXIO);
devres_free(dr);
}
return (addr);
}
void *
vmap(struct page **pages, unsigned int count, unsigned long flags, int prot)
{
void *off;
size_t size;
size = count * PAGE_SIZE;
off = kva_alloc(size);
if (off == NULL)
return (NULL);
vmmap_add(off, size);
pmap_qenter(off, pages, count);
return (off);
}
#define VMAP_MAX_CHUNK_SIZE (65536U / sizeof(struct vm_page)) /* KMEM_ZMAX */
void *
linuxkpi_vmap_pfn(unsigned long *pfns, unsigned int count, int prot)
{
vm_page_t m, *ma, fma;
void *off;
char *coff;
vm_paddr_t pa;
vm_memattr_t attr;
size_t size;
unsigned int i, c, chunk;
size = ptoa(count);
off = kva_alloc(size);
if (off == NULL)
return (NULL);
vmmap_add(off, size);
chunk = MIN(count, VMAP_MAX_CHUNK_SIZE);
attr = pgprot2cachemode(prot);
ma = malloc(chunk * sizeof(vm_page_t), M_TEMP, M_WAITOK | M_ZERO);
fma = NULL;
c = 0;
coff = off;
for (i = 0; i < count; i++) {
pa = IDX_TO_OFF(pfns[i]);
m = PHYS_TO_VM_PAGE(pa);
if (m == NULL) {
if (fma == NULL)
fma = malloc(chunk * sizeof(struct vm_page),
M_TEMP, M_WAITOK | M_ZERO);
m = fma + c;
vm_page_initfake(m, pa, attr);
} else {
pmap_page_set_memattr(m, attr);
}
ma[c] = m;
c++;
if (c == chunk || i == count - 1) {
pmap_qenter(coff, ma, c);
if (i == count - 1)
break;
coff += ptoa(c);
c = 0;
memset(ma, 0, chunk * sizeof(vm_page_t));
if (fma != NULL)
memset(fma, 0, chunk * sizeof(struct vm_page));
}
}
free(fma, M_TEMP);
free(ma, M_TEMP);
return (off);
}
void
vunmap(void *addr)
{
struct vmmap *vmmap;
vmmap = vmmap_remove(addr);
if (vmmap == NULL)
return;
pmap_qremove(addr, vmmap->vm_size / PAGE_SIZE);
kva_free(addr, vmmap->vm_size);
kfree(vmmap);
}
+struct lkpi_vma_pfn_object {
+ TAILQ_ENTRY(lkpi_vma_pfn_object) link;
+ vm_object_t object;
+};
+
+#define LKPI_VMA_PFN_CHUNK_PAGES 64
+
+struct lkpi_vma_pfn_chunk {
+ TAILQ_ENTRY(lkpi_vma_pfn_chunk) link;
+ vm_pindex_t first;
+ unsigned int count;
+ unsigned int remaining;
+ vm_page_t pages[LKPI_VMA_PFN_CHUNK_PAGES];
+};
+
+struct lkpi_vma_pfn_state {
+ TAILQ_HEAD(, lkpi_vma_pfn_object) objects;
+ TAILQ_HEAD(, lkpi_vma_pfn_chunk) page_chunks;
+ struct lkpi_vma_pfn_pins pins;
+ struct mtx objects_lock;
+ struct sx populate_lock;
+ struct sx unmap_lock;
+ struct lkpi_vma_pfn_chunk *last_chunk;
+ vm_pindex_t npages;
+ vm_pindex_t pending;
+ uint64_t invalidation_seq;
+ uint64_t populate_seq;
+ int error;
+};
+
+static struct lkpi_vma_pfn_state *
+lkpi_vma_pfn_get_state(struct vm_area_struct *vma)
+{
+
+ return ((void *)atomic_load_acq_ptr(
+ (volatile uintptr_t *)&vma->vm_pfn_state));
+}
+
+int
+lkpi_vma_pfn_init(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+ vm_pindex_t npages;
+
+ MPASS(lkpi_vma_pfn_get_state(vma) == NULL);
+ npages = vma_pages(vma);
+ if (npages == 0)
+ return (EINVAL);
+ state = kzalloc(sizeof(*state), GFP_KERNEL);
+ if (state == NULL)
+ return (ENOMEM);
+ TAILQ_INIT(&state->objects);
+ TAILQ_INIT(&state->page_chunks);
+ RB_INIT(&state->pins);
+ mtx_init(&state->objects_lock, "lkpi pfn objects", NULL, MTX_DEF);
+ sx_init(&state->populate_lock, "lkpi pfn populate");
+ sx_init(&state->unmap_lock, "lkpi pfn unmap");
+ state->npages = npages;
+ atomic_store_rel_ptr((volatile uintptr_t *)&vma->vm_pfn_state,
+ (uintptr_t)state);
+ return (0);
+}
+
+int
+lkpi_vma_pfn_begin(struct vm_area_struct *vma, uint64_t *invalidation_seq)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state == NULL)
+ return (EINVAL);
+ /* Never wait for a previous handoff while holding mmap_sem. */
+ if (!sx_try_xlock(&state->populate_lock))
+ return (EAGAIN);
+ if (state->pending != 0 || !TAILQ_EMPTY(&state->page_chunks)) {
+ sx_xunlock(&state->populate_lock);
+ return (EBUSY);
+ }
+ mtx_lock(&state->objects_lock);
+ if ((state->invalidation_seq & 1) != 0) {
+ mtx_unlock(&state->objects_lock);
+ sx_xunlock(&state->populate_lock);
+ return (EAGAIN);
+ }
+ state->populate_seq = state->invalidation_seq;
+ *invalidation_seq = state->populate_seq;
+ state->error = 0;
+ mtx_unlock(&state->objects_lock);
+ return (0);
+}
+
+/*
+ * Driver remappers may retry VM_FAULT_OOM internally while retaining their
+ * locks and our earlier busy pages. Stop the remapper with an error bit,
+ * retaining the real cause for linux_cdev_pager_populate() to translate after
+ * the driver has unwound. In particular, VM_FAULT_RETRY alone is not an
+ * error bit and would be mistaken for a successful insertion by remap_sg().
+ */
+static vm_fault_t
+lkpi_vma_pfn_fail(struct lkpi_vma_pfn_state *state, int error)
+{
+
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ if (state->error == 0)
+ state->error = error;
+ return (VM_FAULT_SIGBUS);
+}
+
+int
+lkpi_vma_pfn_error(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ return (state->error);
+}
+
+bool
+lkpi_vma_pfn_unchanged(struct vm_area_struct *vma,
+ uint64_t invalidation_seq)
+{
+ struct lkpi_vma_pfn_state *state;
+ bool unchanged;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state == NULL)
+ return (false);
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ mtx_lock(&state->objects_lock);
+ unchanged = state->invalidation_seq == invalidation_seq;
+ mtx_unlock(&state->objects_lock);
+ return (unchanged);
+}
+
+bool
+lkpi_vma_pfn_handoff_valid(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ return (vma->vm_pfn_count > 0 && (state->pending == 0 ||
+ state->pending == (vm_pindex_t)vma->vm_pfn_count));
+}
+
+bool
+lkpi_vma_pfn_lock(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state == NULL || sx_xlocked(&state->populate_lock))
+ return (false);
+ sx_xlock(&state->populate_lock);
+ return (true);
+}
+
+void
+lkpi_vma_pfn_end(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ MPASS(state->pending == 0);
+ sx_xunlock(&state->populate_lock);
+}
+
+static bool
+lkpi_vma_pfn_object_is_tracked_locked(struct lkpi_vma_pfn_state *state,
+ vm_object_t object)
+{
+ struct lkpi_vma_pfn_object *entry;
+
+ mtx_assert(&state->objects_lock, MA_OWNED);
+ TAILQ_FOREACH(entry, &state->objects, link) {
+ if (entry->object == object)
+ return (true);
+ }
+ return (false);
+}
+
+static bool
+lkpi_vma_pfn_object_is_tracked(struct lkpi_vma_pfn_state *state,
+ vm_object_t object)
+{
+ bool tracked;
+
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ mtx_lock(&state->objects_lock);
+ tracked = lkpi_vma_pfn_object_is_tracked_locked(state, object);
+ mtx_unlock(&state->objects_lock);
+ return (tracked);
+}
+
+static void
+lkpi_vma_pfn_drop_object_ref(vm_object_t locked_object,
+ vm_object_t referenced_object)
+{
+
+ VM_OBJECT_ASSERT_WLOCKED(locked_object);
+ VM_OBJECT_WUNLOCK(locked_object);
+ vm_object_deallocate(referenced_object);
+ VM_OBJECT_WLOCK(locked_object);
+}
+
+static int
+lkpi_vma_pfn_track_object(struct lkpi_vma_pfn_state *state,
+ vm_object_t object, bool have_reference,
+ bool *reference_consumed)
+{
+ struct lkpi_vma_pfn_object *entry;
+
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ *reference_consumed = false;
+ mtx_lock(&state->objects_lock);
+ TAILQ_FOREACH(entry, &state->objects, link) {
+ if (entry->object == object) {
+ mtx_unlock(&state->objects_lock);
+ return (0);
+ }
+ }
+ mtx_unlock(&state->objects_lock);
+ if (!have_reference)
+ return (ESTALE);
+ entry = kzalloc(sizeof(*entry), GFP_ATOMIC);
+ if (entry == NULL) {
+ return (ENOMEM);
+ }
+ entry->object = object;
+ mtx_lock(&state->objects_lock);
+ KASSERT(!lkpi_vma_pfn_object_is_tracked_locked(state, object),
+ ("%s: duplicate object %p", __func__, object));
+ TAILQ_INSERT_TAIL(&state->objects, entry, link);
+ mtx_unlock(&state->objects_lock);
+ *reference_consumed = true;
+ return (0);
+}
+
+/*
+ * Publish the pin while holding the same lock that starts invalidation.
+ * Reject a transaction spanning invalidation even when its pass has ended:
+ * no new pin may appear behind the invalidator's completed traversal.
+ */
+static int
+lkpi_vma_pfn_pin_page(struct lkpi_vma_pfn_state *state, vm_page_t page)
+{
+ struct lkpi_vma_pfn_pin key, *pin;
+
+ MPASS(vm_page_xbusied(page));
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ key.m = page;
+ mtx_lock(&state->objects_lock);
+ if (state->invalidation_seq != state->populate_seq) {
+ mtx_unlock(&state->objects_lock);
+ return (ESTALE);
+ }
+ if (RB_FIND(lkpi_vma_pfn_pins, &state->pins, &key) != NULL) {
+ mtx_unlock(&state->objects_lock);
+ return (0);
+ }
+ mtx_unlock(&state->objects_lock);
+ pin = kzalloc(sizeof(*pin), GFP_ATOMIC);
+ if (pin == NULL)
+ return (ENOMEM);
+ pin->m = page;
+ mtx_lock(&state->objects_lock);
+ if (state->invalidation_seq != state->populate_seq) {
+ mtx_unlock(&state->objects_lock);
+ kfree(pin);
+ return (ESTALE);
+ }
+ /* populate_lock serializes insertions; invalidation can only remove. */
+ vm_page_wire(page);
+ RB_INSERT(lkpi_vma_pfn_pins, &state->pins, pin);
+ mtx_unlock(&state->objects_lock);
+ return (0);
+}
+
+static int
+lkpi_vma_pfn_store_page(struct lkpi_vma_pfn_state *state,
+ vm_pindex_t pindex, vm_page_t page)
+{
+ struct lkpi_vma_pfn_chunk *chunk;
+
+ MPASS(page != NULL);
+ MPASS(vm_page_xbusied(page));
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ MPASS(pindex < state->npages);
+ MPASS(state->pending < state->npages);
+ chunk = state->last_chunk;
+ if (chunk == NULL || chunk->count == nitems(chunk->pages) ||
+ pindex != chunk->first + chunk->count) {
+ chunk = kzalloc(sizeof(*chunk), GFP_ATOMIC);
+ if (chunk == NULL)
+ return (ENOMEM);
+ chunk->first = pindex;
+ TAILQ_INSERT_TAIL(&state->page_chunks, chunk, link);
+ state->last_chunk = chunk;
+ }
+ MPASS(chunk->pages[chunk->count] == NULL);
+ chunk->pages[chunk->count++] = page;
+ chunk->remaining++;
+ state->pending++;
+ return (0);
+}
+
+static bool
+lkpi_vma_pfn_page_is_selected(struct vm_area_struct *vma, vm_page_t page)
+{
+ struct lkpi_vma_pfn_chunk *chunk;
+ struct lkpi_vma_pfn_state *state;
+ unsigned int slot;
+
+ VM_OBJECT_ASSERT_WLOCKED(vma->vm_obj);
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ TAILQ_FOREACH(chunk, &state->page_chunks, link) {
+ for (slot = 0; slot < chunk->count; slot++) {
+ if (chunk->pages[slot] == page)
+ return (true);
+ }
+ }
+ return (atomic_load_ptr(&page->object) == vma->vm_obj &&
+ page->pindex >= vma->vm_pfn_first &&
+ page->pindex - vma->vm_pfn_first <
+ (vm_pindex_t)vma->vm_pfn_count);
+}
+
+/*
+ * A preserved shmem page must regain its backing object's cache attribute
+ * before the last protecting wire is released. Removing its PTEs alone does
+ * not undo the direct-map attribute installed by the PFN fault. Leaving a
+ * WC page on a default-attribute swap object violates the reclaim contract.
+ *
+ * The caller owns xbusy and retains the backing object. Read-only fast
+ * faults can map an xbusy page, so its object must also be write locked
+ * across the unmapped check and attribute change. Drop the caller's pager
+ * lock, if different, rather than nesting VM object locks. The page's busy
+ * ownership and the populate transaction survive that lock drop.
+ *
+ * Return whether the page was unmapped under that lock. In particular, a
+ * caller must not drop its pin on the strength of an earlier unlocked check.
+ */
+static bool
+lkpi_vma_pfn_restore_memattr(vm_page_t page, vm_object_t locked_object)
+{
+ vm_object_t object;
+ vm_memattr_t memattr;
+ bool relock, unmapped;
+
+ MPASS(vm_page_xbusied(page));
+ if (locked_object != NULL)
+ VM_OBJECT_ASSERT_WLOCKED(locked_object);
+ object = page->object;
+ relock = object != NULL && object != locked_object;
+ if (relock) {
+ if (locked_object != NULL)
+ VM_OBJECT_WUNLOCK(locked_object);
+ VM_OBJECT_WLOCK(object);
+ }
+ unmapped = !pmap_page_is_mapped(page);
+ if (unmapped && (page->oflags & VPO_UNMANAGED) == 0 &&
+ (object == NULL || object->type == OBJT_SWAP)) {
+ memattr = object == NULL ? VM_MEMATTR_DEFAULT : object->memattr;
+ if (pmap_page_get_memattr(page) != memattr)
+ pmap_page_set_memattr(page, memattr);
+ }
+ if (relock) {
+ VM_OBJECT_WUNLOCK(object);
+ if (locked_object != NULL)
+ VM_OBJECT_WLOCK(locked_object);
+ }
+ return (unmapped);
+}
+
+/*
+ * Complete both successful and discarded external-page handoffs. A live
+ * mapping retains its pin. An unmapped page can be restored and released.
+ * If invalidation has detached the pin already, that thread owns its wire.
+ */
+void
+lkpi_vma_pfn_release_page(struct vm_area_struct *vma, vm_page_t page)
+{
+ struct lkpi_vma_pfn_pin key, *pin;
+ struct lkpi_vma_pfn_state *state;
+
+ MPASS(vm_page_xbusied(page));
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ pin = NULL;
+ if (lkpi_vma_pfn_restore_memattr(page, vma->vm_obj)) {
+ key.m = page;
+ mtx_lock(&state->objects_lock);
+ pin = RB_FIND(lkpi_vma_pfn_pins, &state->pins, &key);
+ if (pin != NULL)
+ RB_REMOVE(lkpi_vma_pfn_pins, &state->pins, pin);
+ mtx_unlock(&state->objects_lock);
+ }
+ vm_page_xunbusy(page);
+ if (pin != NULL) {
+ /* The wire keeps even an objectless page alive until this drop. */
+ vm_page_unwire(page, PQ_INACTIVE);
+ kfree(pin);
+ }
+}
+
+vm_page_t
+lkpi_vma_pfn_take_page(struct vm_area_struct *vma, vm_object_t object,
+ vm_pindex_t pindex)
+{
+ struct lkpi_vma_pfn_chunk *chunk;
+ struct lkpi_vma_pfn_state *state;
+ vm_page_t page;
+ unsigned int slot;
+
+ VM_OBJECT_ASSERT_WLOCKED(object);
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state == NULL || pindex >= state->npages)
+ return (NULL);
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ TAILQ_FOREACH(chunk, &state->page_chunks, link) {
+ if (pindex < chunk->first)
+ break;
+ if (pindex - chunk->first >= chunk->count)
+ continue;
+ slot = pindex - chunk->first;
+ page = chunk->pages[slot];
+ if (page == NULL)
+ return (NULL);
+ chunk->pages[slot] = NULL;
+ MPASS(chunk->remaining > 0 && state->pending > 0);
+ chunk->remaining--;
+ state->pending--;
+ if (chunk->remaining == 0) {
+ TAILQ_REMOVE(&state->page_chunks, chunk, link);
+ if (state->last_chunk == chunk)
+ state->last_chunk = NULL;
+ kfree(chunk);
+ }
+ return (page);
+ }
+ return (NULL);
+}
+
+void
+lkpi_vma_pfn_abort(struct vm_area_struct *vma, vm_object_t object)
+{
+ struct lkpi_vma_pfn_state *state;
+ vm_page_t page;
+ vm_pindex_t count, pindex;
+
+ VM_OBJECT_ASSERT_WLOCKED(object);
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state != NULL)
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ count = vma->vm_pfn_count;
+ for (pindex = vma->vm_pfn_first; count != 0;
+ count--, pindex++) {
+ page = lkpi_vma_pfn_take_page(vma, object, pindex);
+ if (page == NULL)
+ page = vm_page_lookup(object, pindex);
+ if (page != NULL) {
+ vm_page_deactivate(page);
+ lkpi_vma_pfn_release_page(vma, page);
+ }
+ }
+ vma->vm_pfn_count = 0;
+ KASSERT(state == NULL || state->pending == 0,
+ ("%s: %ju pages remain pending", __func__,
+ state == NULL ? 0 : (uintmax_t)state->pending));
+}
+
+void
+lkpi_vma_pfn_done(struct vm_area_struct *vma, vm_object_t object)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ VM_OBJECT_ASSERT_WLOCKED(object);
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ MPASS(state->pending == 0 &&
+ TAILQ_EMPTY(&state->page_chunks));
+ vma->vm_pfn_count = 0;
+}
+
+bool
+lkpi_vma_pfn_unmap_begin(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state == NULL)
+ return (false);
+ sx_xlock(&state->unmap_lock);
+ return (true);
+}
+
+void
+lkpi_vma_pfn_unmap_end(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ sx_assert(&state->unmap_lock, SA_XLOCKED);
+ sx_xunlock(&state->unmap_lock);
+}
+
+bool
+lkpi_vma_pfn_invalidate_begin(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ if (!lkpi_vma_pfn_unmap_begin(vma))
+ return (false);
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ mtx_lock(&state->objects_lock);
+ MPASS((state->invalidation_seq & 1) == 0);
+ state->invalidation_seq++;
+ mtx_unlock(&state->objects_lock);
+ return (true);
+}
+
+/*
+ * Pins, not historical owner intervals, identify the pages to revoke. In
+ * particular, the extra wire survives OBJ_DEAD making cdev_pager_lookup()
+ * miss this pager while its destructor is still waiting to run.
+ */
+static void
+lkpi_vma_pfn_unmap_pins(struct lkpi_vma_pfn_state *state)
+{
+ struct lkpi_vma_pfn_pin *pin;
+ vm_object_t object;
+ vm_page_t page;
+
+ sx_assert(&state->unmap_lock, SA_XLOCKED);
+ for (;;) {
+ mtx_lock(&state->objects_lock);
+ pin = RB_MIN(lkpi_vma_pfn_pins, &state->pins);
+ if (pin != NULL)
+ RB_REMOVE(lkpi_vma_pfn_pins, &state->pins, pin);
+ mtx_unlock(&state->objects_lock);
+ if (pin == NULL)
+ break;
+ page = pin->m;
+ /* Our wire permits waiting without owning the page's object lock. */
+ while (!vm_page_busy_acquire(page, VM_ALLOC_WAITFAIL))
+ continue;
+ /* Exclude read-only fast faults until the attribute is restored. */
+ object = page->object;
+ if (object != NULL)
+ VM_OBJECT_WLOCK(object);
+ pmap_remove_all(page);
+ lkpi_vma_pfn_restore_memattr(page, object);
+ vm_page_xunbusy(page);
+ if (object != NULL)
+ VM_OBJECT_WUNLOCK(object);
+ vm_page_unwire(page, PQ_INACTIVE);
+ kfree(pin);
+ }
+}
+
+void
+lkpi_vma_pfn_unmap(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state != NULL)
+ lkpi_vma_pfn_unmap_pins(state);
+}
+
+void
+lkpi_vma_pfn_invalidate_end(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_state *state;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ MPASS(state != NULL);
+ sx_assert(&state->unmap_lock, SA_XLOCKED);
+ mtx_lock(&state->objects_lock);
+ MPASS((state->invalidation_seq & 1) != 0);
+ state->invalidation_seq++;
+ mtx_unlock(&state->objects_lock);
+ lkpi_vma_pfn_unmap_end(vma);
+}
+
+void
+lkpi_vma_pfn_fini(struct vm_area_struct *vma)
+{
+ struct lkpi_vma_pfn_chunk *chunk;
+ struct lkpi_vma_pfn_object *entry;
+ struct lkpi_vma_pfn_state *state;
+ vm_page_t page;
+ unsigned int slot;
+
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state == NULL)
+ return;
+ sx_assert(&state->populate_lock, SA_UNLOCKED);
+ mtx_lock(&state->objects_lock);
+ MPASS((state->invalidation_seq & 1) == 0);
+ mtx_unlock(&state->objects_lock);
+ atomic_store_rel_ptr((volatile uintptr_t *)&vma->vm_pfn_state, 0);
+ while ((chunk = TAILQ_FIRST(&state->page_chunks)) != NULL) {
+ TAILQ_REMOVE(&state->page_chunks, chunk, link);
+ for (slot = 0; slot < chunk->count; slot++) {
+ page = chunk->pages[slot];
+ if (page != NULL) {
+ MPASS(state->pending > 0);
+ state->pending--;
+ lkpi_vma_pfn_restore_memattr(page, NULL);
+ vm_page_deactivate(page);
+ vm_page_xunbusy(page);
+ }
+ }
+ kfree(chunk);
+ }
+ state->last_chunk = NULL;
+ MPASS(state->pending == 0);
+ sx_xlock(&state->unmap_lock);
+ lkpi_vma_pfn_unmap_pins(state);
+ for (;;) {
+ mtx_lock(&state->objects_lock);
+ entry = TAILQ_FIRST(&state->objects);
+ if (entry != NULL)
+ TAILQ_REMOVE(&state->objects, entry, link);
+ mtx_unlock(&state->objects_lock);
+ if (entry == NULL)
+ break;
+ vm_object_deallocate(entry->object);
+ kfree(entry);
+ }
+ sx_xunlock(&state->unmap_lock);
+ sx_destroy(&state->unmap_lock);
+ sx_destroy(&state->populate_lock);
+ mtx_destroy(&state->objects_lock);
+ kfree(state);
+}
+
vm_fault_t
lkpi_vmf_insert_pfn_prot_locked(struct vm_area_struct *vma, unsigned long addr,
unsigned long pfn, pgprot_t prot)
{
+ struct lkpi_vma_pfn_state *state;
struct pctrie_iter pages;
vm_object_t vm_obj = vma->vm_obj;
vm_object_t tmp_obj;
vm_page_t page;
vm_pindex_t pindex;
+ vm_memattr_t memattr;
+ bool have_reference, reference_consumed, tracked;
+ int error;
if (addr < vma->vm_start || addr >= vma->vm_end)
return (VM_FAULT_SIGBUS);
VM_OBJECT_ASSERT_WLOCKED(vm_obj);
+ if (offset_in_page(addr) != 0 || (vm_pindex_t)pfn != pfn ||
+ OFF_TO_IDX(IDX_TO_OFF((vm_pindex_t)pfn)) !=
+ (vm_pindex_t)pfn)
+ return (VM_FAULT_SIGBUS);
+ state = lkpi_vma_pfn_get_state(vma);
+ if (state == NULL)
+ return (VM_FAULT_SIGBUS);
+ if (state->error != 0)
+ return (VM_FAULT_SIGBUS);
+ /* Reused VMAs can grow; chunks are sparse, not a fixed-size array. */
+ sx_assert(&state->populate_lock, SA_XLOCKED);
+ state->npages = vma_pages(vma);
+ if (vma->vm_pfn_count < 0 || vma->vm_pfn_count == INT_MAX)
+ return (lkpi_vma_pfn_fail(state, EINVAL));
vm_page_iter_init(&pages, vm_obj);
pindex = OFF_TO_IDX(addr - vma->vm_start);
- if (vma->vm_pfn_count == 0)
+ if (pindex >= state->npages)
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ if (vma->vm_pfn_count != 0 &&
+ pindex != vma->vm_pfn_first + vma->vm_pfn_count)
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ if (vma->vm_pfn_count == 0) {
vma->vm_pfn_first = pindex;
+ }
MPASS(pindex < OFF_TO_IDX(vma->vm_end));
+ memattr = pgprot2cachemode(prot);
retry:
- page = vm_page_grab_iter(vm_obj, pindex, VM_ALLOC_NOCREAT, &pages);
+ page = vm_page_grab_iter(vm_obj, pindex,
+ VM_ALLOC_NOCREAT | VM_ALLOC_NOWAIT, &pages);
if (page == NULL) {
+ if (vm_page_lookup(vm_obj, pindex) != NULL)
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
page = PHYS_TO_VM_PAGE(IDX_TO_OFF(pfn));
if (page == NULL)
- return (VM_FAULT_SIGBUS);
- if (!vm_page_busy_acquire(page, VM_ALLOC_WAITFAIL)) {
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ tmp_obj = atomic_load_ptr(&page->object);
+ have_reference = false;
+ tracked = tmp_obj != NULL && tmp_obj != vm_obj &&
+ lkpi_vma_pfn_object_is_tracked(state, tmp_obj);
+ if (tmp_obj != NULL && tmp_obj != vm_obj && !tracked) {
+ /*
+ * VM object locks are type-stable. Lock and revalidate the
+ * source before taking the reference that will protect this VMA.
+ */
+ VM_OBJECT_WUNLOCK(vm_obj);
+ if (!VM_OBJECT_TRYWLOCK(tmp_obj)) {
+ VM_OBJECT_WLOCK(vm_obj);
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
+ }
+ if (page->object != tmp_obj ||
+ (tmp_obj->flags & OBJ_DEAD) != 0) {
+ VM_OBJECT_WUNLOCK(tmp_obj);
+ VM_OBJECT_WLOCK(vm_obj);
+ pctrie_iter_reset(&pages);
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
+ }
+ vm_object_reference_locked(tmp_obj);
+ VM_OBJECT_WUNLOCK(tmp_obj);
+ VM_OBJECT_WLOCK(vm_obj);
+ if (vm_page_lookup(vm_obj, pindex) != NULL ||
+ atomic_load_ptr(&page->object) != tmp_obj) {
+ lkpi_vma_pfn_drop_object_ref(vm_obj, tmp_obj);
+ pctrie_iter_reset(&pages);
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
+ }
+ have_reference = true;
+ }
+ if (!vm_page_tryxbusy(page)) {
+ /*
+ * A selected page stays xbusy until this transaction is
+ * consumed or aborted. Refuse an alias rather than wait
+ * for busy ownership that this transaction must release.
+ */
+ if (lkpi_vma_pfn_page_is_selected(vma, page)) {
+ if (have_reference)
+ lkpi_vma_pfn_drop_object_ref(vm_obj, tmp_obj);
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ }
+ if (have_reference)
+ lkpi_vma_pfn_drop_object_ref(vm_obj, tmp_obj);
+ /* Drop the whole batch before waiting for another owner. */
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
+ }
+ if (page->object != tmp_obj) {
+ vm_page_xunbusy(page);
+ if (have_reference)
+ lkpi_vma_pfn_drop_object_ref(vm_obj, tmp_obj);
pctrie_iter_reset(&pages);
- goto retry;
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
+ }
+ /*
+ * Linux installs a special PFN PTE without moving a managed page
+ * out of its backing object. Preserve that ownership for shmem
+ * pages and hand the xbusy page directly to vm_fault_populate().
+ */
+ if (tmp_obj != NULL && tmp_obj != vm_obj &&
+ tmp_obj->type == OBJT_SWAP &&
+ (page->oflags & VPO_UNMANAGED) == 0) {
+ /* The driver must pin non-default mappings until invalidation. */
+ if (!vm_page_all_valid(page) ||
+ (memattr != tmp_obj->memattr && !vm_page_wired(page)) ||
+ (pmap_page_get_memattr(page) != memattr &&
+ pmap_page_is_mapped(page))) {
+ vm_page_xunbusy(page);
+ if (have_reference)
+ lkpi_vma_pfn_drop_object_ref(vm_obj, tmp_obj);
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ }
+ error = lkpi_vma_pfn_track_object(state, tmp_obj,
+ have_reference, &reference_consumed);
+ if (error != 0 || (have_reference && !reference_consumed)) {
+ vm_page_xunbusy(page);
+ if (have_reference)
+ lkpi_vma_pfn_drop_object_ref(vm_obj, tmp_obj);
+ return (lkpi_vma_pfn_fail(state,
+ error == ENOMEM ? ENOMEM : EAGAIN));
+ }
+ /*
+ * Publish our pin before changing the cache attribute. A
+ * read-only backing-object fault can create an alias even
+ * while this page is xbusy. If a later allocation fails,
+ * that alias may prevent restoration and must retain a pin.
+ * Keep xbusy across the object-lock switch so invalidation
+ * cannot consume the pin and then miss our attribute change.
+ */
+ error = lkpi_vma_pfn_pin_page(state, page);
+ if (error != 0) {
+ vm_page_xunbusy(page);
+ return (lkpi_vma_pfn_fail(state,
+ error == ESTALE ? EAGAIN : error));
+ }
+ if (pmap_page_get_memattr(page) != memattr) {
+ VM_OBJECT_WUNLOCK(vm_obj);
+ if (!VM_OBJECT_TRYWLOCK(tmp_obj)) {
+ VM_OBJECT_WLOCK(vm_obj);
+ lkpi_vma_pfn_release_page(vma, page);
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
+ }
+ MPASS(page->object == tmp_obj);
+ if (pmap_page_get_memattr(page) != memattr &&
+ pmap_page_is_mapped(page)) {
+ VM_OBJECT_WUNLOCK(tmp_obj);
+ VM_OBJECT_WLOCK(vm_obj);
+ lkpi_vma_pfn_release_page(vma, page);
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ }
+ if (pmap_page_get_memattr(page) != memattr)
+ pmap_page_set_memattr(page, memattr);
+ VM_OBJECT_WUNLOCK(tmp_obj);
+ VM_OBJECT_WLOCK(vm_obj);
+ if (vm_page_lookup(vm_obj, pindex) != NULL) {
+ lkpi_vma_pfn_release_page(vma, page);
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
+ }
+ }
+ error = lkpi_vma_pfn_store_page(state, pindex, page);
+ if (error != 0) {
+ lkpi_vma_pfn_release_page(vma, page);
+ return (lkpi_vma_pfn_fail(state, error));
+ }
+ vma->vm_pfn_count++;
+ return (VM_FAULT_NOPAGE);
}
if (page->object != NULL) {
- tmp_obj = page->object;
+ if (tracked) {
+ vm_page_xunbusy(page);
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ }
+ if (tmp_obj == vm_obj) {
+ vm_object_reference_locked(tmp_obj);
+ have_reference = true;
+ }
+ MPASS(have_reference);
vm_page_xunbusy(page);
VM_OBJECT_WUNLOCK(vm_obj);
- VM_OBJECT_WLOCK(tmp_obj);
- if (page->object == tmp_obj &&
- vm_page_busy_acquire(page, VM_ALLOC_WAITFAIL)) {
+ if (!VM_OBJECT_TRYWLOCK(tmp_obj)) {
+ vm_object_deallocate(tmp_obj);
+ VM_OBJECT_WLOCK(vm_obj);
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
+ }
+ if (page->object == tmp_obj && vm_page_tryxbusy(page)) {
KASSERT(page->object == tmp_obj,
("page has changed identity"));
- KASSERT((page->oflags & VPO_UNMANAGED) == 0,
- ("page does not belong to shmem"));
- vm_pager_page_unswapped(page);
+ if ((page->oflags & VPO_UNMANAGED) != 0 ||
+ !vm_page_wired(page)) {
+ vm_page_xunbusy(page);
+ VM_OBJECT_WUNLOCK(tmp_obj);
+ vm_object_deallocate(tmp_obj);
+ VM_OBJECT_WLOCK(vm_obj);
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ }
if (pmap_page_is_mapped(page)) {
vm_page_xunbusy(page);
VM_OBJECT_WUNLOCK(tmp_obj);
- printf("%s: page rename failed: page "
- "is mapped\n", __func__);
+ vm_object_deallocate(tmp_obj);
VM_OBJECT_WLOCK(vm_obj);
- return (VM_FAULT_NOPAGE);
+ return (lkpi_vma_pfn_fail(state, EINVAL));
}
+ vm_pager_page_unswapped(page);
vm_page_remove(page);
+ } else {
+ VM_OBJECT_WUNLOCK(tmp_obj);
+ vm_object_deallocate(tmp_obj);
+ VM_OBJECT_WLOCK(vm_obj);
+ return (lkpi_vma_pfn_fail(state, EAGAIN));
}
VM_OBJECT_WUNLOCK(tmp_obj);
+ vm_object_deallocate(tmp_obj);
pctrie_iter_reset(&pages);
VM_OBJECT_WLOCK(vm_obj);
goto retry;
}
if (vm_page_iter_insert(page, vm_obj, pindex, &pages) != 0) {
vm_page_xunbusy(page);
- return (VM_FAULT_OOM);
+ return (lkpi_vma_pfn_fail(state, ENOMEM));
}
vm_page_valid(page);
}
- pmap_page_set_memattr(page, pgprot2cachemode(prot));
+ if (!vm_page_all_valid(page) ||
+ (page->oflags & VPO_UNMANAGED) != 0 ||
+ VM_PAGE_TO_PHYS(page) != IDX_TO_OFF(pfn)) {
+ vm_page_xunbusy(page);
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ }
+ if (pmap_page_get_memattr(page) != memattr &&
+ pmap_page_is_mapped(page)) {
+ vm_page_xunbusy(page);
+ return (lkpi_vma_pfn_fail(state, EINVAL));
+ }
+ pmap_page_set_memattr(page, memattr);
vma->vm_pfn_count++;
return (VM_FAULT_NOPAGE);
}
int
lkpi_remap_pfn_range(struct vm_area_struct *vma, unsigned long start_addr,
unsigned long start_pfn, unsigned long size, pgprot_t prot)
{
vm_object_t vm_obj;
- unsigned long addr, pfn;
+ unsigned long addr, end_addr, npages, pfn;
int err = 0;
vm_obj = vma->vm_obj;
+ if (size == 0)
+ return (0);
+ if (offset_in_page(start_addr) != 0 || offset_in_page(size) != 0 ||
+ start_addr < vma->vm_start || start_addr >= vma->vm_end ||
+ size > vma->vm_end - start_addr)
+ return (-EINVAL);
+ npages = size >> PAGE_SHIFT;
+ if ((vm_pindex_t)start_pfn != start_pfn ||
+ npages - 1 > ULONG_MAX - start_pfn)
+ return (-EINVAL);
+ end_addr = start_addr + size;
VM_OBJECT_WLOCK(vm_obj);
for (addr = start_addr, pfn = start_pfn;
- addr < start_addr + size;
+ addr != end_addr;
addr += PAGE_SIZE) {
vm_fault_t ret;
-retry:
ret = lkpi_vmf_insert_pfn_prot_locked(vma, addr, pfn, prot);
if ((ret & VM_FAULT_OOM) != 0) {
- VM_OBJECT_WUNLOCK(vm_obj);
- vm_wait(NULL);
- VM_OBJECT_WLOCK(vm_obj);
- goto retry;
+ err = -ENOMEM;
+ break;
}
if ((ret & VM_FAULT_ERROR) != 0) {
err = -EFAULT;
break;
}
pfn++;
}
VM_OBJECT_WUNLOCK(vm_obj);
if (unlikely(err)) {
- zap_vma_ptes(vma, start_addr,
- (pfn - start_pfn) << PAGE_SHIFT);
+ zap_vma_ptes(vma, start_addr, addr - start_addr);
return (err);
}
return (0);
}
int
lkpi_io_mapping_map_user(struct io_mapping *iomap,
struct vm_area_struct *vma, unsigned long addr,
unsigned long pfn, unsigned long size)
{
pgprot_t prot;
int ret;
prot = cachemode2protval(iomap->attr);
ret = lkpi_remap_pfn_range(vma, addr, pfn, size, prot);
return (ret);
}
/*
* Although FreeBSD version of unmap_mapping_range has semantics and types of
* parameters compatible with Linux version, the values passed in are different
* @obj should match to vm_private_data field of vm_area_struct returned by
* mmap file operation handler, see linux_file_mmap_single() sources
* @holelen should match to size of area to be munmapped.
*/
void
lkpi_unmap_mapping_range(void *obj, loff_t const holebegin __unused,
loff_t const holelen __unused, int even_cows __unused)
{
vm_object_t devobj;
devobj = cdev_pager_lookup(obj);
if (devobj != NULL) {
- cdev_mgtdev_pager_free_pages(devobj);
+ linux_cdev_pager_free_pages(devobj);
vm_object_deallocate(devobj);
}
}
int
lkpi_arch_phys_wc_add(unsigned long base, unsigned long size)
{
#ifdef __i386__
struct mem_range_desc *mrdesc;
int error, id, act;
/* If PAT is available, do nothing */
if (pat_works)
return (0);
mrdesc = malloc(sizeof(*mrdesc), M_LKMTRR, M_WAITOK);
mrdesc->mr_base = base;
mrdesc->mr_len = size;
mrdesc->mr_flags = MDF_WRITECOMBINE;
strlcpy(mrdesc->mr_owner, "drm", sizeof(mrdesc->mr_owner));
act = MEMRANGE_SET_UPDATE;
error = mem_range_attr_set(mrdesc, &act);
if (error == 0) {
error = idr_get_new(&mtrr_idr, mrdesc, &id);
MPASS(idr_find(&mtrr_idr, id) == mrdesc);
if (error != 0) {
act = MEMRANGE_SET_REMOVE;
mem_range_attr_set(mrdesc, &act);
}
}
if (error != 0) {
free(mrdesc, M_LKMTRR);
pr_warn(
"Failed to add WC MTRR for [%p-%p]: %d; "
"performance may suffer\n",
(void *)base, (void *)(base + size - 1), error);
} else
pr_warn("Successfully added WC MTRR for [%p-%p]\n",
(void *)base, (void *)(base + size - 1));
return (error != 0 ? -error : id + __MTRR_ID_BASE);
#else
return (0);
#endif
}
void
lkpi_arch_phys_wc_del(int reg)
{
#ifdef __i386__
struct mem_range_desc *mrdesc;
int act;
/* Check if arch_phys_wc_add() failed. */
if (reg < __MTRR_ID_BASE)
return;
mrdesc = idr_find(&mtrr_idr, reg - __MTRR_ID_BASE);
MPASS(mrdesc != NULL);
idr_remove(&mtrr_idr, reg - __MTRR_ID_BASE);
act = MEMRANGE_SET_REMOVE;
mem_range_attr_set(mrdesc, &act);
free(mrdesc, M_LKMTRR);
#endif
}
int
lkpi_set_pages_attr(struct page *page, int numpages, vm_memattr_t ma)
{
while (numpages-- > 0) {
/*
* pmap_page_set_memattr() would only update the DMAP mapping
* if it's a normal page, leaving the kernel map untouched.
*/
MPASS(page->object != kernel_object);
/*
* pmap_page_set_memattr() sets page->md.pat_mode, which is
* crucial for future userspace mappings.
*/
pmap_page_set_memattr(page, ma);
page++;
}
return (0);
}
/*
* This is a highly simplified version of the Linux page_frag_cache.
* We only support up-to 1 single page as fragment size and we will
* always return a full page. This may be wasteful on small objects
* but the only known consumer (mt76) is either asking for a half-page
* or a full page. If this was to become a problem we can implement
* a more elaborate version.
*/
void *
linuxkpi_page_frag_alloc(struct page_frag_cache *pfc,
size_t fragsz, gfp_t gfp)
{
struct page *pages;
if (fragsz == 0)
return (NULL);
KASSERT(fragsz <= PAGE_SIZE, ("%s: fragsz %zu > PAGE_SIZE not yet "
"supported", __func__, fragsz));
pages = alloc_pages(gfp, flsl(howmany(fragsz, PAGE_SIZE) - 1));
if (pages == NULL)
return (NULL);
pfc->va = linux_page_address(pages);
/* Passed in as "count" to __page_frag_cache_drain(). Unused by us. */
pfc->pagecnt_bias = 0;
return (pfc->va);
}
void
linuxkpi_page_frag_free(void *addr)
{
struct page *page;
page = virt_to_page(addr);
linux_free_pages(page, 0);
}
void
linuxkpi__page_frag_cache_drain(struct page *page, size_t count __unused)
{
linux_free_pages(page, 0);
}
static void
lkpi_page_init(void *arg)
{
int i;
mtx_init(&vmmaplock, "IO Map lock", NULL, MTX_DEF);
for (i = 0; i < VMMAP_HASH_SIZE; i++)
LIST_INIT(&vmmaphead[i]);
}
SYSINIT(lkpi_page, SI_SUB_DRIVERS, SI_ORDER_SECOND, lkpi_page_init, NULL);
static void
lkpi_page_uninit(void *arg)
{
mtx_destroy(&vmmaplock);
}
SYSUNINIT(lkpi_page, SI_SUB_DRIVERS, SI_ORDER_SECOND, lkpi_page_uninit, NULL);
diff --git a/sys/sys/param.h b/sys/sys/param.h
index 5b14612a228afe11d1e7d3e110b61631ca819967..7b953293f888a6f7079b63792a8508b2cd1d44e1 100644
--- a/sys/sys/param.h
+++ b/sys/sys/param.h
@@ -1,387 +1,387 @@
/*-
* SPDX-License-Identifier: BSD-3-Clause
*
* Copyright (c) 1982, 1986, 1989, 1993
* The Regents of the University of California. All rights reserved.
* (c) UNIX System Laboratories, Inc.
* All or some portions of this file are derived from material licensed
* to the University of California by American Telephone and Telegraph
* Co. or Unix System Laboratories, Inc. and are reproduced herein with
* the permission of UNIX System Laboratories, Inc.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice, this list of conditions and the following disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
* 3. Neither the name of the University nor the names of its contributors
* may be used to endorse or promote products derived from this software
* without specific prior written permission.
*
* THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS ``AS IS'' AND
* ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
* ARE DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS
* OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)
* HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT
* LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY
* OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
* SUCH DAMAGE.
*/
#ifndef _SYS_PARAM_H_
#define _SYS_PARAM_H_
#include <sys/_null.h>
#include <sys/_param.h>
#define BSD 199506 /* System version (year & month). */
#define BSD4_3 1
#define BSD4_4 1
/*
* __FreeBSD_version numbers are documented in the Porter's Handbook.
* If you bump the version for any reason, you should update the documentation
* there.
* Currently this lives here in the doc/ repository:
*
* documentation/content/en/books/porters-handbook/versions/_index.adoc
*
* Encoding: <major><two digit minor>Rxx
* 'R' is in the range 0 to 4 if this is a release branch or
* X.0-CURRENT before releng/X.0 is created, otherwise 'R' is
* in the range 5 to 9.
* Short hand: MMmmXXX
*
* __FreeBSD_version is bumped every time there's a change in the base system
* that's noteworthy. A noteworthy change is any change which changes the
* kernel's KBI in -CURRENT, one that changes some detail about the system that
* external software (or the ports system) would want to know about, one that
* adds a system call, one that adds or deletes a shipped library, a security
* fix, or similar change not specifically noted here. Bumps should be limited
* to one per day / a couple per week except for security fixes.
*
* The approved way to obtain this from a shell script is:
* awk '/^\#define[[:space:]]*__FreeBSD_version/ {print $3}'
* Other methods to parse this file may work, but are not guaranteed against
* future changes. The above script works back to FreeBSD 3.x when this macro
* was introduced. This number is propagated to other places needing it that
* cannot include sys/param.h and should only be updated here.
*/
#undef __FreeBSD_version
-#define __FreeBSD_version 1600026
+#define __FreeBSD_version 1600028
/*
* __FreeBSD_kernel__ indicates that this system uses the kernel of FreeBSD,
* which by definition is always true on FreeBSD. This macro is also defined
* on other systems that use the kernel of FreeBSD, such as GNU/kFreeBSD.
*
* It is tempting to use this macro in userland code when we want to enable
* kernel-specific routines, and in fact it's fine to do this in code that
* is part of FreeBSD itself. However, be aware that as presence of this
* macro is still not widespread (e.g. older FreeBSD versions, 3rd party
* compilers, etc), it is STRONGLY DISCOURAGED to check for this macro in
* external applications without also checking for __FreeBSD__ as an
* alternative.
*/
#undef __FreeBSD_kernel__
#define __FreeBSD_kernel__
#if defined(_KERNEL) || defined(_WANT_P_OSREL)
#define P_OSREL_SIGWAIT 700000
#define P_OSREL_SIGSEGV 700004
#define P_OSREL_MAP_ANON 800104
#define P_OSREL_MAP_FSTRICT 1100036
#define P_OSREL_SHUTDOWN_ENOTCONN 1100077
#define P_OSREL_MAP_GUARD 1200035
#define P_OSREL_WRFSBASE 1200041
#define P_OSREL_CK_CYLGRP 1200046
#define P_OSREL_VMTOTAL64 1200054
#define P_OSREL_CK_SUPERBLOCK 1300000
#define P_OSREL_CK_INODE 1300005
#define P_OSREL_POWERPC_NEW_AUX_ARGS 1300070
#define P_OSREL_TIDPID 1400079
#define P_OSREL_ARM64_SPSR 1400084
#define P_OSREL_TLSBASE 1500044
#define P_OSREL_EXTERRCTL 1500045
#define P_OSREL_AMD64_TF_FRED 1600014
#define P_OSREL_MAJOR(x) ((x) / 100000)
#endif
#ifndef LOCORE
#include <sys/types.h>
#endif
/*
* Machine-independent constants (some used in following include files).
* Redefined constants are from POSIX 1003.1 limits file.
*
* MAXCOMLEN should be >= sizeof(ac_comm) (see <acct.h>)
*/
#include <sys/syslimits.h>
#define MAXCOMLEN 19 /* max command name remembered */
#define MAXINTERP PATH_MAX /* max interpreter file name length */
#define MAXLOGNAME 33 /* max login name length (incl. NUL) */
#define MAXUPRC CHILD_MAX /* max simultaneous processes */
#define NCARGS ARG_MAX /* max bytes for an exec function */
#define NGROUPS (NGROUPS_MAX+1) /* max number groups */
#define NOFILE OPEN_MAX /* max open files per process */
#define NOGROUP 65535 /* marker for empty group set member */
#define MAXHOSTNAMELEN 256 /* max hostname size */
#define SPECNAMELEN 255 /* max length of devicename */
/* More types and definitions used throughout the kernel. */
#ifdef _KERNEL
#include <sys/cdefs.h>
#include <sys/errno.h>
#ifndef LOCORE
#include <sys/time.h>
#include <sys/priority.h>
#endif
#ifndef FALSE
#define FALSE 0
#endif
#ifndef TRUE
#define TRUE 1
#endif
#endif
#if !defined(_KERNEL) && !defined(_STANDALONE) && !defined(LOCORE)
/* Signals. */
#include <sys/signal.h>
#endif
/* Machine type dependent parameters. */
#include <machine/param.h>
#ifndef _KERNEL
#include <sys/limits.h>
#include <sys/_maxphys.h>
#endif
#ifndef DEV_BSHIFT
#define DEV_BSHIFT 9 /* log2(DEV_BSIZE) */
#endif
#define DEV_BSIZE (1<<DEV_BSHIFT)
#ifndef BLKDEV_IOSIZE
#define BLKDEV_IOSIZE PAGE_SIZE /* default block device I/O size */
#endif
#ifndef DFLTPHYS
#define DFLTPHYS (64 * 1024) /* default max raw I/O transfer size */
#endif
#ifndef MAXDUMPPGS
#define MAXDUMPPGS (DFLTPHYS/PAGE_SIZE)
#endif
#ifdef STACKALIGNBYTES
#define STACKALIGN(p) (__align_down(p, STACKALIGNBYTES + 1))
#endif
/*
* Constants related to network buffer management.
* MCLBYTES must be no larger than PAGE_SIZE.
*/
#ifndef MSIZE
#define MSIZE 256 /* size of an mbuf */
#endif
#ifndef MCLSHIFT
#define MCLSHIFT 11 /* convert bytes to mbuf clusters */
#endif /* MCLSHIFT */
#define MCLBYTES (1 << MCLSHIFT) /* size of an mbuf cluster */
#if PAGE_SIZE <= 8192
#define MJUMPAGESIZE PAGE_SIZE
#else
#define MJUMPAGESIZE (8 * 1024)
#endif
#define MJUM9BYTES (9 * 1024) /* jumbo cluster 9k */
#define MJUM16BYTES (16 * 1024) /* jumbo cluster 16k */
/*
* Mach derived conversion macros
*/
#define round_page(x) roundup2(x, PAGE_SIZE)
#define trunc_page(x) rounddown2(x, PAGE_SIZE)
#define atop(x) ((x) >> PAGE_SHIFT)
#define ptoa(x) ((x) << PAGE_SHIFT)
#define pgtok(x) ((x) * (PAGE_SIZE / 1024))
/*
* Some macros for units conversion
*/
/* clicks to bytes */
#ifndef ctob
#define ctob(x) ((x)<<PAGE_SHIFT)
#endif
/* bytes to clicks */
#ifndef btoc
#define btoc(x) (((vm_offset_t)(x)+PAGE_MASK)>>PAGE_SHIFT)
#endif
/*
* btodb() is messy and perhaps slow because `bytes' may be an off_t. We
* want to shift an unsigned type to avoid sign extension and we don't
* want to widen `bytes' unnecessarily. Assume that the result fits in
* a daddr_t.
*/
#ifndef btodb
#define btodb(bytes) /* calculates (bytes / DEV_BSIZE) */ \
(sizeof (bytes) > sizeof(long) \
? (daddr_t)((unsigned long long)(bytes) >> DEV_BSHIFT) \
: (daddr_t)((unsigned long)(bytes) >> DEV_BSHIFT))
#endif
#ifndef dbtob
#define dbtob(db) /* calculates (db * DEV_BSIZE) */ \
((off_t)(db) << DEV_BSHIFT)
#endif
#define PRIMASK 0x0ff
#define PCATCH 0x100 /* OR'd with pri for tsleep to check signals */
#define PDROP 0x200 /* OR'd with pri to stop re-entry of interlock mutex */
#define PNOLOCK 0x400 /* OR'd with pri to allow sleeping w/o a lock */
#define PRILASTFLAG 0x400 /* Last flag defined above */
#define NZERO 0 /* default "nice" */
#define CMASK 022 /* default file mask: S_IWGRP|S_IWOTH */
#define NODEV (dev_t)(-1) /* non-existent device */
/*
* File system parameters and macros.
*
* MAXBSIZE - Filesystems are made out of blocks of at most MAXBSIZE bytes
* per block. MAXBSIZE may be made larger without effecting
* any existing filesystems as long as it does not exceed MAXPHYS,
* and may be made smaller at the risk of not being able to use
* filesystems which require a block size exceeding MAXBSIZE.
*
* MAXBCACHEBUF - Maximum size of a buffer in the buffer cache. This must
* be >= MAXBSIZE and can be set differently for different
* architectures by defining it in <machine/param.h>.
* Making this larger allows NFS to do larger reads/writes.
*
* BKVASIZE - Nominal buffer space per buffer, in bytes. BKVASIZE is the
* minimum KVM memory reservation the kernel is willing to make.
* Filesystems can of course request smaller chunks. Actual
* backing memory uses a chunk size of a page (PAGE_SIZE).
* The default value here can be overridden on a per-architecture
* basis by defining it in <machine/param.h>.
*
* If you make BKVASIZE too small you risk seriously fragmenting
* the buffer KVM map which may slow things down a bit. If you
* make it too big the kernel will not be able to optimally use
* the KVM memory reserved for the buffer cache and will wind
* up with too-few buffers.
*
* The default is 16384, roughly 2x the block size used by a
* normal UFS filesystem.
*/
#define MAXBSIZE 65536 /* must be power of 2 */
#ifndef MAXBCACHEBUF
#define MAXBCACHEBUF MAXBSIZE /* must be a power of 2 >= MAXBSIZE */
#endif
#ifndef BKVASIZE
#define BKVASIZE 16384 /* must be power of 2 */
#endif
#define BKVAMASK (BKVASIZE-1)
/*
* MAXPATHLEN defines the longest permissible path length after expanding
* symbolic links. It is used to allocate a temporary buffer from the buffer
* pool in which to do the name expansion, hence should be a power of two,
* and must be less than or equal to MAXBSIZE. MAXSYMLINKS defines the
* maximum number of symbolic links that may be expanded in a path name.
* It should be set high enough to allow all legitimate uses, but halt
* infinite loops reasonably quickly.
*/
#define MAXPATHLEN PATH_MAX
#define MAXSYMLINKS 32
/* Bit map related macros. */
#define setbit(a,i) (((unsigned char *)(a))[(i)/NBBY] |= 1<<((i)%NBBY))
#define clrbit(a,i) (((unsigned char *)(a))[(i)/NBBY] &= ~(1<<((i)%NBBY)))
#define isset(a,i) \
(((const unsigned char *)(a))[(i)/NBBY] & (1<<((i)%NBBY)))
#define isclr(a,i) \
((((const unsigned char *)(a))[(i)/NBBY] & (1<<((i)%NBBY))) == 0)
/* Macros for counting and rounding provided by <sys/_param.h>. */
/* Macros for min/max. */
#define MIN(a,b) (((a)<(b))?(a):(b))
#define MAX(a,b) (((a)>(b))?(a):(b))
#ifdef _KERNEL
/*
* Basic byte order function prototypes for non-inline functions.
*/
#ifndef LOCORE
#ifndef _BYTEORDER_PROTOTYPED
#define _BYTEORDER_PROTOTYPED
__BEGIN_DECLS
__uint32_t htonl(__uint32_t);
__uint16_t htons(__uint16_t);
__uint32_t ntohl(__uint32_t);
__uint16_t ntohs(__uint16_t);
__END_DECLS
#endif
#endif
#ifndef _BYTEORDER_FUNC_DEFINED
#define _BYTEORDER_FUNC_DEFINED
#define htonl(x) __htonl(x)
#define htons(x) __htons(x)
#define ntohl(x) __ntohl(x)
#define ntohs(x) __ntohs(x)
#endif /* !_BYTEORDER_FUNC_DEFINED */
#endif /* _KERNEL */
/*
* Scale factor for scaled integers used to count %cpu time and load avgs.
*
* The number of CPU `tick's that map to a unique `%age' can be expressed
* by the formula (1 / (2 ^ (FSHIFT - 11))). Since the intermediate
* calculation is done with 64-bit precision, the maximum load average that can
* be calculated is approximately 2^32 / FSCALE.
*
* For the scheduler to maintain a 1:1 mapping of CPU `tick' to `%age',
* FSHIFT must be at least 11. This gives a maximum load avg of 2 million.
*/
#define FSHIFT 11 /* bits to right of fixed binary point */
#define FSCALE (1<<FSHIFT)
#define dbtoc(db) /* calculates devblks to pages */ \
((db + (ctodb(1) - 1)) >> (PAGE_SHIFT - DEV_BSHIFT))
#define ctodb(db) /* calculates pages to devblks */ \
((db) << (PAGE_SHIFT - DEV_BSHIFT))
/*
* Old spelling of __containerof().
*/
#define member2struct(s, m, x) \
((struct s *)(void *)((char *)(x) - offsetof(struct s, m)))
/*
* Access a variable length array that has been declared as a fixed
* length array.
*/
#define __PAST_END(array, offset) (((__typeof__(*(array)) *)(array))[offset])
#endif /* _SYS_PARAM_H_ */
diff --git a/sys/vm/device_pager.c b/sys/vm/device_pager.c
index 9277d6255418675d86d5c77385c4549f31703397..ae23302cb4cba30072be23f65510e5ca7c54d74e 100644
--- a/sys/vm/device_pager.c
+++ b/sys/vm/device_pager.c
@@ -1,561 +1,612 @@
/*-
* SPDX-License-Identifier: BSD-3-Clause
*
* Copyright (c) 1990 University of Utah.
* Copyright (c) 1991, 1993
* The Regents of the University of California. All rights reserved.
*
* This code is derived from software contributed to Berkeley by
* the Systems Programming Group of the University of Utah Computer
* Science Department.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice, this list of conditions and the following disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
* 3. Neither the name of the University nor the names of its contributors
* may be used to endorse or promote products derived from this software
* without specific prior written permission.
*
* THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS ``AS IS'' AND
* ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
* ARE DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS
* OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)
* HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT
* LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY
* OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
* SUCH DAMAGE.
*/
#include <sys/param.h>
#include <sys/systm.h>
#include <sys/conf.h>
#include <sys/lock.h>
#include <sys/proc.h>
#include <sys/mutex.h>
#include <sys/mman.h>
#include <sys/rwlock.h>
#include <sys/sx.h>
#include <sys/user.h>
#include <sys/vmmeter.h>
#include <vm/vm.h>
#include <vm/vm_param.h>
#include <vm/vm_object.h>
#include <vm/vm_page.h>
#include <vm/vm_pager.h>
#include <vm/vm_radix.h>
#include <vm/vm_phys.h>
#include <vm/vm_radix.h>
#include <vm/uma.h>
static void dev_pager_init(void);
static vm_object_t dev_pager_alloc(void *, vm_ooffset_t, vm_prot_t,
vm_ooffset_t, struct ucred *);
static void dev_pager_dealloc(vm_object_t);
static int dev_pager_getpages(vm_object_t, vm_page_t *, int, int *, int *);
static void dev_pager_putpages(vm_object_t, vm_page_t *, int, int, int *);
static boolean_t dev_pager_haspage(vm_object_t, vm_pindex_t, int *, int *);
static void dev_pager_free_page(vm_object_t object, vm_page_t m);
static int dev_pager_populate(vm_object_t object, vm_pindex_t pidx,
int fault_type, vm_prot_t, vm_pindex_t *first, vm_pindex_t *last);
+static vm_page_t dev_pager_populate_take_page(vm_object_t object,
+ vm_pindex_t pidx);
+static void dev_pager_populate_done(vm_object_t object);
+static void dev_pager_populate_release_page(vm_object_t object, vm_page_t page);
/* list of device pager objects */
static struct pagerlst dev_pager_object_list;
/* protect list manipulation */
static struct mtx dev_pager_mtx;
const struct pagerops devicepagerops = {
.pgo_kvme_type = KVME_TYPE_DEVICE,
.pgo_init = dev_pager_init,
.pgo_alloc = dev_pager_alloc,
.pgo_dealloc = dev_pager_dealloc,
.pgo_getpages = dev_pager_getpages,
.pgo_putpages = dev_pager_putpages,
.pgo_haspage = dev_pager_haspage,
};
const struct pagerops mgtdevicepagerops = {
.pgo_kvme_type = KVME_TYPE_MGTDEVICE,
.pgo_alloc = dev_pager_alloc,
.pgo_dealloc = dev_pager_dealloc,
.pgo_getpages = dev_pager_getpages,
.pgo_putpages = dev_pager_putpages,
.pgo_haspage = dev_pager_haspage,
.pgo_populate = dev_pager_populate,
+ .pgo_populate_take_page = dev_pager_populate_take_page,
+ .pgo_populate_done = dev_pager_populate_done,
+ .pgo_populate_release_page = dev_pager_populate_release_page,
};
static int old_dev_pager_ctor(void *handle, vm_ooffset_t size, vm_prot_t prot,
vm_ooffset_t foff, struct ucred *cred, u_short *color);
static void old_dev_pager_dtor(void *handle);
static int old_dev_pager_fault(vm_object_t object, vm_ooffset_t offset,
int prot, vm_page_t *mres);
static void old_dev_pager_path(void *handle, char *path, size_t len);
static const struct cdev_pager_ops old_dev_pager_ops = {
.cdev_pg_ctor = old_dev_pager_ctor,
.cdev_pg_dtor = old_dev_pager_dtor,
.cdev_pg_fault = old_dev_pager_fault,
.cdev_pg_path = old_dev_pager_path
};
static void
dev_pager_init(void)
{
TAILQ_INIT(&dev_pager_object_list);
mtx_init(&dev_pager_mtx, "dev_pager list", NULL, MTX_DEF);
}
vm_object_t
cdev_pager_lookup(void *handle)
{
vm_object_t object;
again:
mtx_lock(&dev_pager_mtx);
object = vm_pager_object_lookup(&dev_pager_object_list, handle);
if (object != NULL && object->un_pager.devp.handle == NULL) {
msleep(&object->un_pager.devp.handle, &dev_pager_mtx,
PVM | PDROP, "cdplkp", 0);
vm_object_deallocate(object);
goto again;
}
mtx_unlock(&dev_pager_mtx);
return (object);
}
vm_object_t
cdev_pager_allocate(void *handle, enum obj_type tp,
const struct cdev_pager_ops *ops, vm_ooffset_t size, vm_prot_t prot,
vm_ooffset_t foff, struct ucred *cred)
{
vm_object_t object;
vm_pindex_t pindex;
KASSERT(handle != NULL, ("device pager with NULL handle"));
if (tp != OBJT_DEVICE && tp != OBJT_MGTDEVICE)
return (NULL);
KASSERT(tp == OBJT_MGTDEVICE || ops->cdev_pg_populate == NULL,
("populate on unmanaged device pager"));
+ KASSERT((ops->cdev_pg_populate_take_page == NULL) ==
+ (ops->cdev_pg_populate_done == NULL),
+ ("incomplete populate handoff methods"));
+ KASSERT(ops->cdev_pg_populate_take_page == NULL ||
+ ops->cdev_pg_populate != NULL,
+ ("populate handoff without populate method"));
+ KASSERT(ops->cdev_pg_populate_release_page == NULL ||
+ ops->cdev_pg_populate_take_page != NULL,
+ ("populate release without handoff method"));
/*
* Offset should be page aligned.
*/
if (foff & PAGE_MASK)
return (NULL);
/*
* Treat the mmap(2) file offset as an unsigned value for a
* device mapping. This, in effect, allows a user to pass all
* possible off_t values as the mapping cookie to the driver. At
* this point, we know that both foff and size are a multiple
* of the page size. Do a check to avoid wrap.
*/
size = round_page(size);
pindex = OFF_TO_IDX(foff) + OFF_TO_IDX(size);
if (pindex > OBJ_MAX_SIZE || pindex < OFF_TO_IDX(foff) ||
pindex < OFF_TO_IDX(size))
return (NULL);
again:
mtx_lock(&dev_pager_mtx);
/*
* Look up pager, creating as necessary.
*/
object = vm_pager_object_lookup(&dev_pager_object_list, handle);
if (object == NULL) {
vm_object_t object1;
/*
* Allocate object and associate it with the pager. Initialize
* the object's pg_color based upon the physical address of the
* device's memory.
*/
mtx_unlock(&dev_pager_mtx);
object1 = vm_object_allocate(tp, pindex);
mtx_lock(&dev_pager_mtx);
object = vm_pager_object_lookup(&dev_pager_object_list, handle);
if (object != NULL) {
object1->type = OBJT_DEAD;
vm_object_deallocate(object1);
object1 = NULL;
if (object->un_pager.devp.handle == NULL) {
msleep(&object->un_pager.devp.handle,
&dev_pager_mtx, PVM | PDROP, "cdplkp", 0);
vm_object_deallocate(object);
goto again;
}
/*
* We raced with other thread while allocating object.
*/
if (pindex > object->size)
object->size = pindex;
KASSERT(object->type == tp,
("Inconsistent device pager type %p %d",
object, tp));
KASSERT(object->un_pager.devp.ops == ops,
("Inconsistent devops %p %p", object, ops));
} else {
u_short color;
object = object1;
object1 = NULL;
object->handle = handle;
object->un_pager.devp.ops = ops;
if (object->type == OBJT_DEVICE)
vm_object_set_flag(object, OBJ_PG_DTOR);
TAILQ_INSERT_TAIL(&dev_pager_object_list, object,
pager_object_list);
mtx_unlock(&dev_pager_mtx);
if (ops->cdev_pg_populate != NULL)
vm_object_set_flag(object, OBJ_POPULATE);
if (ops->cdev_pg_ctor(handle, size, prot, foff,
cred, &color) != 0) {
mtx_lock(&dev_pager_mtx);
TAILQ_REMOVE(&dev_pager_object_list, object,
pager_object_list);
wakeup(&object->un_pager.devp.handle);
mtx_unlock(&dev_pager_mtx);
object->type = OBJT_DEAD;
vm_object_deallocate(object);
object = NULL;
mtx_lock(&dev_pager_mtx);
} else {
+ if (ops->cdev_pg_mlock_skip != NULL &&
+ ops->cdev_pg_mlock_skip(handle))
+ vm_object_set_flag(object, OBJ_NOMLOCK);
mtx_lock(&dev_pager_mtx);
object->flags |= OBJ_COLORED;
object->pg_color = color;
object->un_pager.devp.handle = handle;
wakeup(&object->un_pager.devp.handle);
}
}
MPASS(object1 == NULL);
} else {
if (object->un_pager.devp.handle == NULL) {
msleep(&object->un_pager.devp.handle,
&dev_pager_mtx, PVM | PDROP, "cdplkp", 0);
vm_object_deallocate(object);
goto again;
}
if (pindex > object->size)
object->size = pindex;
KASSERT(object->type == tp,
("Inconsistent device pager type %p %d", object, tp));
}
mtx_unlock(&dev_pager_mtx);
return (object);
}
static vm_object_t
dev_pager_alloc(void *handle, vm_ooffset_t size, vm_prot_t prot,
vm_ooffset_t foff, struct ucred *cred)
{
return (cdev_pager_allocate(handle, OBJT_DEVICE, &old_dev_pager_ops,
size, prot, foff, cred));
}
void
cdev_pager_get_path(vm_object_t object, char *path, size_t sz)
{
if (object->un_pager.devp.ops->cdev_pg_path != NULL)
object->un_pager.devp.ops->cdev_pg_path(
object->un_pager.devp.handle, path, sz);
}
void
cdev_pager_free_page(vm_object_t object, vm_page_t m)
{
if (object->type == OBJT_MGTDEVICE) {
struct pctrie_iter pages;
vm_page_iter_init(&pages, object);
vm_radix_iter_lookup(&pages, m->pindex);
cdev_mgtdev_pager_free_page(&pages, m);
} else if (object->type == OBJT_DEVICE)
dev_pager_free_page(object, m);
else
KASSERT(false,
("Invalid device type obj %p m %p", object, m));
}
void
cdev_mgtdev_pager_free_page(struct pctrie_iter *pages, vm_page_t m)
{
pmap_remove_all(m);
vm_page_iter_remove(pages, m);
}
void
cdev_mgtdev_pager_free_pages(vm_object_t object)
{
struct pctrie_iter pages;
vm_page_t m;
vm_page_iter_init(&pages, object);
VM_OBJECT_WLOCK(object);
retry:
KASSERT(pctrie_iter_is_reset(&pages),
("%s: pctrie_iter not reset for retry", __func__));
VM_RADIX_FOREACH(m, &pages) {
if (!vm_page_busy_acquire(m, VM_ALLOC_WAITFAIL)) {
pctrie_iter_reset(&pages);
goto retry;
}
cdev_mgtdev_pager_free_page(&pages, m);
}
VM_OBJECT_WUNLOCK(object);
}
static void
dev_pager_free_page(vm_object_t object, vm_page_t m)
{
VM_OBJECT_ASSERT_WLOCKED(object);
KASSERT((object->type == OBJT_DEVICE &&
(m->oflags & VPO_UNMANAGED) != 0),
("Managed device or page obj %p m %p", object, m));
vm_page_putfake(m);
}
static void
dev_pager_dealloc(vm_object_t object)
{
VM_OBJECT_WUNLOCK(object);
object->un_pager.devp.ops->cdev_pg_dtor(object->un_pager.devp.handle);
mtx_lock(&dev_pager_mtx);
TAILQ_REMOVE(&dev_pager_object_list, object, pager_object_list);
mtx_unlock(&dev_pager_mtx);
VM_OBJECT_WLOCK(object);
if (object->type == OBJT_DEVICE) {
struct pctrie_iter pages;
vm_page_t m;
vm_page_iter_init(&pages, object);
restart:
VM_RADIX_FOREACH(m, &pages) {
if (!vm_page_busy_acquire(m, VM_ALLOC_WAITFAIL)) {
pctrie_iter_reset(&pages);
goto restart;
}
if (vm_page_iter_remove(&pages, m)) {
/*
* We could end up with invalid pages installed
* by the generic page fault handler. Typically
* these are replaced by the device pager.
*/
vm_page_free(m);
} else if ((m->flags & PG_FICTITIOUS) != 0)
dev_pager_free_page(object, m);
}
}
object->handle = NULL;
object->type = OBJT_DEAD;
}
static int
dev_pager_getpages(vm_object_t object, vm_page_t *ma, int count, int *rbehind,
int *rahead)
{
int error;
/* Since our haspage reports zero after/before, the count is 1. */
KASSERT(count == 1, ("%s: count %d", __func__, count));
if (object->un_pager.devp.ops->cdev_pg_fault == NULL)
return (VM_PAGER_FAIL);
VM_OBJECT_WLOCK(object);
error = object->un_pager.devp.ops->cdev_pg_fault(object,
IDX_TO_OFF(ma[0]->pindex), PROT_READ, &ma[0]);
VM_OBJECT_ASSERT_WLOCKED(object);
if (error == VM_PAGER_OK) {
KASSERT((object->type == OBJT_DEVICE &&
(ma[0]->oflags & VPO_UNMANAGED) != 0) ||
(object->type == OBJT_MGTDEVICE &&
(ma[0]->oflags & VPO_UNMANAGED) == 0),
("Wrong page type %p %p", ma[0], object));
if (rbehind)
*rbehind = 0;
if (rahead)
*rahead = 0;
}
VM_OBJECT_WUNLOCK(object);
return (error);
}
static int
dev_pager_populate(vm_object_t object, vm_pindex_t pidx, int fault_type,
vm_prot_t max_prot, vm_pindex_t *first, vm_pindex_t *last)
{
VM_OBJECT_ASSERT_WLOCKED(object);
if (object->un_pager.devp.ops->cdev_pg_populate == NULL)
return (VM_PAGER_FAIL);
return (object->un_pager.devp.ops->cdev_pg_populate(object, pidx,
fault_type, max_prot, first, last));
}
+static vm_page_t
+dev_pager_populate_take_page(vm_object_t object, vm_pindex_t pidx)
+{
+
+ VM_OBJECT_ASSERT_WLOCKED(object);
+ if (object->un_pager.devp.ops->cdev_pg_populate_take_page == NULL)
+ return (NULL);
+ return (object->un_pager.devp.ops->cdev_pg_populate_take_page(
+ object, pidx));
+}
+
+static void
+dev_pager_populate_release_page(vm_object_t object, vm_page_t page)
+{
+
+ VM_OBJECT_ASSERT_WLOCKED(object);
+ if (object->un_pager.devp.ops->cdev_pg_populate_release_page != NULL)
+ object->un_pager.devp.ops->cdev_pg_populate_release_page(object,
+ page);
+ else
+ vm_page_xunbusy(page);
+}
+
+static void
+dev_pager_populate_done(vm_object_t object)
+{
+
+ VM_OBJECT_ASSERT_WLOCKED(object);
+ if (object->un_pager.devp.ops->cdev_pg_populate_done != NULL)
+ object->un_pager.devp.ops->cdev_pg_populate_done(object);
+}
+
static int
old_dev_pager_fault(vm_object_t object, vm_ooffset_t offset, int prot,
vm_page_t *mres)
{
vm_paddr_t paddr;
vm_page_t m_paddr, page;
struct cdev *dev;
struct cdevsw *csw;
struct file *fpop;
struct thread *td;
vm_memattr_t memattr, memattr1;
int ref, ret;
memattr = object->memattr;
VM_OBJECT_WUNLOCK(object);
dev = object->handle;
csw = dev_refthread(dev, &ref);
if (csw == NULL) {
VM_OBJECT_WLOCK(object);
return (VM_PAGER_FAIL);
}
td = curthread;
fpop = td->td_fpop;
td->td_fpop = NULL;
ret = csw->d_mmap(dev, offset, &paddr, prot, &memattr);
td->td_fpop = fpop;
dev_relthread(dev, ref);
if (ret != 0) {
printf(
"WARNING: dev_pager_getpage: map function returns error %d", ret);
VM_OBJECT_WLOCK(object);
return (VM_PAGER_FAIL);
}
/* If "paddr" is a real page, perform a sanity check on "memattr". */
if ((m_paddr = vm_phys_paddr_to_vm_page(paddr)) != NULL &&
(memattr1 = pmap_page_get_memattr(m_paddr)) != memattr) {
/*
* For the /dev/mem d_mmap routine to return the
* correct memattr, pmap_page_get_memattr() needs to
* be called, which we do there.
*/
if ((csw->d_flags & D_MEM) == 0) {
printf("WARNING: Device driver %s has set "
"\"memattr\" inconsistently (drv %u pmap %u).\n",
csw->d_name, memattr, memattr1);
}
memattr = memattr1;
}
if (((*mres)->flags & PG_FICTITIOUS) != 0) {
/*
* If the passed in result page is a fake page, update it with
* the new physical address.
*/
page = *mres;
VM_OBJECT_WLOCK(object);
vm_page_updatefake(page, paddr, memattr);
} else {
/*
* Replace the passed in reqpage page with our own fake page and
* free up the all of the original pages.
*/
page = vm_page_getfake(paddr, memattr);
VM_OBJECT_WLOCK(object);
vm_page_replace(page, object, (*mres)->pindex, *mres);
*mres = page;
}
vm_page_valid(page);
return (VM_PAGER_OK);
}
static void
dev_pager_putpages(vm_object_t object, vm_page_t *m, int count, int flags,
int *rtvals)
{
panic("dev_pager_putpage called");
}
static boolean_t
dev_pager_haspage(vm_object_t object, vm_pindex_t pindex, int *before,
int *after)
{
if (before != NULL)
*before = 0;
if (after != NULL)
*after = 0;
return (TRUE);
}
static int
old_dev_pager_ctor(void *handle, vm_ooffset_t size, vm_prot_t prot,
vm_ooffset_t foff, struct ucred *cred, u_short *color)
{
struct cdev *dev;
struct cdevsw *csw;
vm_memattr_t dummy;
vm_ooffset_t off;
vm_paddr_t paddr;
unsigned int npages;
int ref;
/*
* Make sure this device can be mapped.
*/
dev = handle;
csw = dev_refthread(dev, &ref);
if (csw == NULL)
return (ENXIO);
/*
* Check that the specified range of the device allows the desired
* protection.
*
* XXX assumes VM_PROT_* == PROT_*
*/
npages = OFF_TO_IDX(size);
paddr = 0; /* Make paddr initialized for the case of size == 0. */
for (off = foff; npages--; off += PAGE_SIZE) {
if (csw->d_mmap(dev, off, &paddr, (int)prot, &dummy) != 0) {
dev_relthread(dev, ref);
return (EINVAL);
}
}
dev_ref(dev);
dev_relthread(dev, ref);
*color = atop(paddr) - OFF_TO_IDX(off - PAGE_SIZE);
return (0);
}
static void
old_dev_pager_dtor(void *handle)
{
dev_rel(handle);
}
static void
old_dev_pager_path(void *handle, char *path, size_t len)
{
struct cdev *cdev = handle;
if (cdev != NULL)
dev_copyname(cdev, path, len);
}
diff --git a/sys/vm/vm_fault.c b/sys/vm/vm_fault.c
index 9f00e3b51a3795da9262a2fac23082c48257c505..80907348408c46c1296bce97ce7fe51412b1df8c 100644
--- a/sys/vm/vm_fault.c
+++ b/sys/vm/vm_fault.c
@@ -1,2515 +1,2627 @@
/*-
* SPDX-License-Identifier: (BSD-4-Clause AND MIT-CMU)
*
* Copyright (c) 1991, 1993
* The Regents of the University of California. All rights reserved.
* Copyright (c) 1994 John S. Dyson
* All rights reserved.
* Copyright (c) 1994 David Greenman
* All rights reserved.
*
*
* This code is derived from software contributed to Berkeley by
* The Mach Operating System project at Carnegie-Mellon University.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice, this list of conditions and the following disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
* 3. All advertising materials mentioning features or use of this software
* must display the following acknowledgement:
* This product includes software developed by the University of
* California, Berkeley and its contributors.
* 4. Neither the name of the University nor the names of its contributors
* may be used to endorse or promote products derived from this software
* without specific prior written permission.
*
* THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS ``AS IS'' AND
* ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
* ARE DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS
* OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)
* HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT
* LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY
* OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
* SUCH DAMAGE.
*
*
* Copyright (c) 1987, 1990 Carnegie-Mellon University.
* All rights reserved.
*
* Authors: Avadis Tevanian, Jr., Michael Wayne Young
*
* Permission to use, copy, modify and distribute this software and
* its documentation is hereby granted, provided that both the copyright
* notice and this permission notice appear in all copies of the
* software, derivative works or modified versions, and any portions
* thereof, and that both notices appear in supporting documentation.
*
* CARNEGIE MELLON ALLOWS FREE USE OF THIS SOFTWARE IN ITS "AS IS"
* CONDITION. CARNEGIE MELLON DISCLAIMS ANY LIABILITY OF ANY KIND
* FOR ANY DAMAGES WHATSOEVER RESULTING FROM THE USE OF THIS SOFTWARE.
*
* Carnegie Mellon requests users of this software to return to
*
* Software Distribution Coordinator or Software.Distribution@CS.CMU.EDU
* School of Computer Science
* Carnegie Mellon University
* Pittsburgh PA 15213-3890
*
* any improvements or extensions that they make and grant Carnegie the
* rights to redistribute these changes.
*/
/*
* Page fault handling module.
*/
#include "opt_ktrace.h"
#include "opt_vm.h"
#include <sys/systm.h>
#include <sys/kernel.h>
#include <sys/lock.h>
#include <sys/mman.h>
#include <sys/mutex.h>
#include <sys/pctrie.h>
#include <sys/proc.h>
#include <sys/racct.h>
#include <sys/refcount.h>
#include <sys/resourcevar.h>
#include <sys/rwlock.h>
#include <sys/sched.h>
#include <sys/sf_buf.h>
#include <sys/signalvar.h>
#include <sys/sysctl.h>
#include <sys/sysent.h>
#include <sys/vmmeter.h>
#include <sys/vnode.h>
#ifdef KTRACE
#include <sys/ktrace.h>
#endif
#include <vm/vm.h>
#include <vm/vm_param.h>
#include <vm/pmap.h>
#include <vm/vm_map.h>
#include <vm/vm_object.h>
#include <vm/vm_page.h>
#include <vm/vm_pageout.h>
#include <vm/vm_kern.h>
#include <vm/vm_pager.h>
#include <vm/vm_radix.h>
#include <vm/vm_extern.h>
#include <vm/vm_reserv.h>
#define PFBAK 4
#define PFFOR 4
#define VM_FAULT_READ_DEFAULT (1 + VM_FAULT_READ_AHEAD_INIT)
#define VM_FAULT_DONTNEED_MIN 1048576
struct faultstate {
/* Fault parameters. */
vm_offset_t vaddr;
vm_page_t *m_hold;
vm_prot_t fault_type;
vm_prot_t prot;
int fault_flags;
boolean_t wired;
/* Control state. */
struct timeval oom_start_time;
bool oom_started;
int nera;
bool can_read_lock;
/* Page reference for cow. */
vm_page_t m_cow;
/* Current object. */
vm_object_t object;
vm_pindex_t pindex;
vm_page_t m;
bool m_needs_zeroing;
/* Top-level map object. */
vm_object_t first_object;
vm_pindex_t first_pindex;
vm_page_t first_m;
/* Map state. */
vm_map_t map;
vm_map_entry_t entry;
int map_generation;
bool lookup_still_valid;
/* Vnode if locked. */
struct vnode *vp;
};
/*
* Return codes for internal fault routines.
*/
enum fault_status {
FAULT_SUCCESS = 10000, /* Return success to user. */
FAULT_FAILURE, /* Return failure to user. */
FAULT_CONTINUE, /* Continue faulting. */
FAULT_RESTART, /* Restart fault. */
+ FAULT_OOM, /* Retry after waiting for memory. */
FAULT_OUT_OF_BOUNDS, /* Invalid address for pager. */
FAULT_HARD, /* Performed I/O. */
FAULT_SOFT, /* Found valid page. */
FAULT_PROTECTION_FAILURE, /* Invalid access. */
};
enum fault_next_status {
FAULT_NEXT_GOTOBJ = 1,
FAULT_NEXT_NOOBJ,
FAULT_NEXT_RESTART,
};
static void vm_fault_dontneed(const struct faultstate *fs, vm_offset_t vaddr,
int ahead);
static void vm_fault_prefault(const struct faultstate *fs, vm_offset_t addra,
int backward, int forward, bool obj_locked);
static int vm_pfault_oom_attempts = 3;
SYSCTL_INT(_vm, OID_AUTO, pfault_oom_attempts, CTLFLAG_RWTUN,
&vm_pfault_oom_attempts, 0,
"Number of page allocation attempts in page fault handler before it "
"triggers OOM handling");
static int vm_pfault_oom_wait = 10;
SYSCTL_INT(_vm, OID_AUTO, pfault_oom_wait, CTLFLAG_RWTUN,
&vm_pfault_oom_wait, 0,
"Number of seconds to wait for free pages before retrying "
"the page fault handler");
static inline void
vm_fault_page_release(vm_page_t *mp)
{
vm_page_t m;
m = *mp;
if (m != NULL) {
/*
* We are likely to loop around again and attempt to busy
* this page. Deactivating it leaves it available for
* pageout while optimizing fault restarts.
*/
vm_page_deactivate(m);
if (vm_page_xbusied(m))
vm_page_xunbusy(m);
else
vm_page_sunbusy(m);
*mp = NULL;
}
}
static inline void
vm_fault_page_free(vm_page_t *mp)
{
vm_page_t m;
m = *mp;
if (m != NULL) {
VM_OBJECT_ASSERT_WLOCKED(m->object);
if (!vm_page_wired(m))
vm_page_free(m);
else
vm_page_xunbusy(m);
*mp = NULL;
}
}
/*
* Return true if a vm_pager_get_pages() call is needed in order to check
* whether the pager might have a particular page, false if it can be determined
* immediately that the pager can not have a copy. For swap objects, this can
* be checked quickly.
*/
static inline bool
vm_fault_object_needs_getpages(vm_object_t object)
{
VM_OBJECT_ASSERT_LOCKED(object);
return ((object->flags & OBJ_SWAP) == 0 ||
!pctrie_is_empty(&object->un_pager.swp.swp_blks));
}
static inline void
vm_fault_unlock_map(struct faultstate *fs)
{
if (fs->lookup_still_valid) {
vm_map_lookup_done(fs->map, fs->entry);
fs->lookup_still_valid = false;
}
}
static void
vm_fault_unlock_vp(struct faultstate *fs)
{
if (fs->vp != NULL) {
vput(fs->vp);
fs->vp = NULL;
}
}
static bool
vm_fault_might_be_cow(struct faultstate *fs)
{
return (fs->object != fs->first_object);
}
static void
vm_fault_deallocate(struct faultstate *fs)
{
vm_fault_page_release(&fs->m_cow);
vm_fault_page_release(&fs->m);
vm_object_pip_wakeup(fs->object);
if (vm_fault_might_be_cow(fs)) {
VM_OBJECT_WLOCK(fs->first_object);
vm_fault_page_free(&fs->first_m);
VM_OBJECT_WUNLOCK(fs->first_object);
vm_object_pip_wakeup(fs->first_object);
}
vm_object_deallocate(fs->first_object);
vm_fault_unlock_map(fs);
vm_fault_unlock_vp(fs);
}
static void
vm_fault_unlock_and_deallocate(struct faultstate *fs)
{
VM_OBJECT_UNLOCK(fs->object);
vm_fault_deallocate(fs);
}
static void
vm_fault_dirty(struct faultstate *fs, vm_page_t m)
{
bool need_dirty;
if (((fs->prot & VM_PROT_WRITE) == 0 &&
(fs->fault_flags & VM_FAULT_DIRTY) == 0) ||
(m->oflags & VPO_UNMANAGED) != 0)
return;
VM_PAGE_OBJECT_BUSY_ASSERT(m);
need_dirty = ((fs->fault_type & VM_PROT_WRITE) != 0 &&
(fs->fault_flags & VM_FAULT_WIRE) == 0) ||
(fs->fault_flags & VM_FAULT_DIRTY) != 0;
vm_object_set_writeable_dirty(m->object);
/*
* If the fault is a write, we know that this page is being
* written NOW so dirty it explicitly to save on
* pmap_is_modified() calls later.
*
* Also, since the page is now dirty, we can possibly tell
* the pager to release any swap backing the page.
*/
if (need_dirty && vm_page_set_dirty(m) == 0) {
/*
* If this is a NOSYNC mmap we do not want to set PGA_NOSYNC
* if the page is already dirty to prevent data written with
* the expectation of being synced from not being synced.
* Likewise if this entry does not request NOSYNC then make
* sure the page isn't marked NOSYNC. Applications sharing
* data should use the same flags to avoid ping ponging.
*/
if ((fs->entry->eflags & MAP_ENTRY_NOSYNC) != 0)
vm_page_aflag_set(m, PGA_NOSYNC);
else
vm_page_aflag_clear(m, PGA_NOSYNC);
}
}
static bool
vm_fault_is_read(const struct faultstate *fs)
{
return ((fs->prot & VM_PROT_WRITE) == 0 &&
(fs->fault_type & (VM_PROT_COPY | VM_PROT_WRITE)) == 0);
}
/*
* Unlocks fs.first_object and fs.map on success.
*/
static enum fault_status
vm_fault_soft_fast(struct faultstate *fs)
{
vm_page_t m, m_map;
#if VM_NRESERVLEVEL > 0
vm_page_t m_super;
int flags;
#endif
int psind;
vm_offset_t vaddr;
MPASS(fs->vp == NULL);
/*
* If we fail, vast majority of the time it is because the page is not
* there to begin with. Opportunistically perform the lookup and
* subsequent checks without the object lock, revalidate later.
*
* Note: a busy page can be mapped for read|execute access.
*/
m = vm_page_lookup_unlocked(fs->first_object, fs->first_pindex);
if (m == NULL || !vm_page_all_valid(m) ||
((fs->prot & VM_PROT_WRITE) != 0 && vm_page_busied(m))) {
VM_OBJECT_WLOCK(fs->first_object);
return (FAULT_FAILURE);
}
vaddr = fs->vaddr;
VM_OBJECT_RLOCK(fs->first_object);
/*
* Now that we stabilized the state, revalidate the page is in the shape
* we encountered above.
*/
if (m->object != fs->first_object || m->pindex != fs->first_pindex)
goto fail;
vm_object_busy(fs->first_object);
if (!vm_page_all_valid(m) ||
((fs->prot & VM_PROT_WRITE) != 0 && vm_page_busied(m)))
goto fail_busy;
m_map = m;
psind = 0;
#if VM_NRESERVLEVEL > 0
if ((m->flags & PG_FICTITIOUS) == 0 &&
(m_super = vm_reserv_to_superpage(m)) != NULL) {
psind = m_super->psind;
KASSERT(psind > 0,
("psind %d of m_super %p < 1", psind, m_super));
flags = PS_ALL_VALID;
if ((fs->prot & VM_PROT_WRITE) != 0) {
/*
* Create a superpage mapping allowing write access
* only if none of the constituent pages are busy and
* all of them are already dirty (except possibly for
* the page that was faulted on).
*/
flags |= PS_NONE_BUSY;
if ((fs->first_object->flags & OBJ_UNMANAGED) == 0)
flags |= PS_ALL_DIRTY;
}
while (rounddown2(vaddr, pagesizes[psind]) < fs->entry->start ||
roundup2(vaddr + 1, pagesizes[psind]) > fs->entry->end ||
(vaddr & (pagesizes[psind] - 1)) !=
(VM_PAGE_TO_PHYS(m) & (pagesizes[psind] - 1)) ||
!vm_page_ps_test(m_super, psind, flags, m) ||
!pmap_ps_enabled(fs->map->pmap)) {
psind--;
if (psind == 0)
break;
m_super += rounddown2(m - m_super,
atop(pagesizes[psind]));
KASSERT(m_super->psind >= psind,
("psind %d of m_super %p < %d", m_super->psind,
m_super, psind));
}
if (psind > 0) {
m_map = m_super;
vaddr = rounddown2(vaddr, pagesizes[psind]);
/* Preset the modified bit for dirty superpages. */
if ((flags & PS_ALL_DIRTY) != 0)
fs->fault_type |= VM_PROT_WRITE;
}
}
#endif
if (pmap_enter(fs->map->pmap, vaddr, m_map, fs->prot, fs->fault_type |
PMAP_ENTER_NOSLEEP | (fs->wired ? PMAP_ENTER_WIRED : 0), psind) !=
KERN_SUCCESS)
goto fail_busy;
if (fs->m_hold != NULL) {
(*fs->m_hold) = m;
vm_page_wire(m);
}
if (psind == 0 && !fs->wired)
vm_fault_prefault(fs, vaddr, PFBAK, PFFOR, true);
VM_OBJECT_RUNLOCK(fs->first_object);
vm_fault_dirty(fs, m);
vm_object_unbusy(fs->first_object);
vm_map_lookup_done(fs->map, fs->entry);
curthread->td_ru.ru_minflt++;
return (FAULT_SUCCESS);
fail_busy:
vm_object_unbusy(fs->first_object);
fail:
if (!VM_OBJECT_TRYUPGRADE(fs->first_object)) {
VM_OBJECT_RUNLOCK(fs->first_object);
VM_OBJECT_WLOCK(fs->first_object);
}
return (FAULT_FAILURE);
}
static void
vm_fault_restore_map_lock(struct faultstate *fs)
{
VM_OBJECT_ASSERT_WLOCKED(fs->first_object);
MPASS(blockcount_read(&fs->first_object->paging_in_progress) > 0);
if (!vm_map_trylock_read(fs->map)) {
VM_OBJECT_WUNLOCK(fs->first_object);
vm_map_lock_read(fs->map);
VM_OBJECT_WLOCK(fs->first_object);
}
fs->lookup_still_valid = true;
}
static void
vm_fault_populate_check_page(vm_page_t m)
{
/*
* Check each page to ensure that the pager is obeying the
- * interface: the page must be installed in the object, fully
- * valid, and exclusively busied.
+ * interface: the page must be fully valid and exclusively busied.
+ * Most populated pages are installed in the pager object, but a pager
+ * may explicitly return a page owned by another object.
*/
MPASS(m != NULL);
MPASS(vm_page_all_valid(m));
MPASS(vm_page_xbusied(m));
}
static void
vm_fault_populate_cleanup(vm_object_t object, vm_pindex_t first,
vm_pindex_t last)
{
struct pctrie_iter pages;
vm_page_t m;
+ vm_pindex_t pidx;
VM_OBJECT_ASSERT_WLOCKED(object);
MPASS(first <= last);
- vm_page_iter_limit_init(&pages, object, last + 1);
- VM_RADIX_FORALL_FROM(m, &pages, first) {
+ m = vm_pager_populate_take_page(object, first);
+ if (m == NULL) {
+ vm_page_iter_limit_init(&pages, object, last + 1);
+ VM_RADIX_FORALL_FROM(m, &pages, first) {
+ vm_fault_populate_check_page(m);
+ vm_page_deactivate(m);
+ vm_page_xunbusy(m);
+ }
+ KASSERT(pages.index == last,
+ ("%s: Object %p first %#jx last %#jx index %#jx",
+ __func__, object, (uintmax_t)first, (uintmax_t)last,
+ (uintmax_t)pages.index));
+ return;
+ }
+ for (pidx = first;; pidx++) {
vm_fault_populate_check_page(m);
vm_page_deactivate(m);
- vm_page_xunbusy(m);
+ vm_pager_populate_release_page(object, m);
+ if (pidx == last)
+ break;
+ m = vm_pager_populate_take_page(object, pidx + 1);
}
- KASSERT(pages.index == last,
- ("%s: Object %p first %#jx last %#jx index %#jx",
- __func__, object, (uintmax_t)first, (uintmax_t)last,
- (uintmax_t)pages.index));
}
static enum fault_status
vm_fault_populate(struct faultstate *fs)
{
vm_offset_t vaddr;
- vm_page_t m;
+ vm_page_t external_hold, m;
vm_pindex_t map_first, map_last, pager_first, pager_last, pidx;
+ vm_prot_t max_prot;
int bdry_idx, i, npages, psind, rv;
enum fault_status res;
MPASS(fs->object == fs->first_object);
VM_OBJECT_ASSERT_WLOCKED(fs->first_object);
MPASS(blockcount_read(&fs->first_object->paging_in_progress) > 0);
MPASS(fs->first_object->backing_object == NULL);
MPASS(fs->lookup_still_valid);
pager_first = OFF_TO_IDX(fs->entry->offset);
pager_last = pager_first + atop(fs->entry->end - fs->entry->start) - 1;
+ max_prot = fs->entry->max_protection;
+ bdry_idx = MAP_ENTRY_SPLIT_BOUNDARY_INDEX(fs->entry);
vm_fault_unlock_map(fs);
vm_fault_unlock_vp(fs);
res = FAULT_SUCCESS;
+ external_hold = NULL;
/*
* Call the pager (driver) populate() method.
*
* There is no guarantee that the method will be called again
* if the current fault is for read, and a future fault is
* for write. Report the entry's maximum allowed protection
* to the driver.
*/
rv = vm_pager_populate(fs->first_object, fs->first_pindex,
- fs->fault_type, fs->entry->max_protection, &pager_first,
+ fs->fault_type, max_prot, &pager_first,
&pager_last);
VM_OBJECT_ASSERT_WLOCKED(fs->first_object);
if (rv == VM_PAGER_BAD) {
/*
* VM_PAGER_BAD is the backdoor for a pager to request
* normal fault handling.
*/
vm_fault_restore_map_lock(fs);
if (fs->map->timestamp != fs->map_generation)
return (FAULT_RESTART);
return (FAULT_CONTINUE);
}
+ if (rv == VM_PAGER_OUT_OF_BOUNDS)
+ return (FAULT_OUT_OF_BOUNDS);
+ if (rv == VM_PAGER_AGAIN)
+ return (FAULT_OOM);
+ if (rv == VM_PAGER_RETRY)
+ return (FAULT_RESTART);
if (rv != VM_PAGER_OK)
return (FAULT_FAILURE); /* AKA SIGSEGV */
/* Ensure that the driver is obeying the interface. */
MPASS(pager_first <= pager_last);
MPASS(fs->first_pindex <= pager_last);
MPASS(fs->first_pindex >= pager_first);
MPASS(pager_last < fs->first_object->size);
- vm_fault_restore_map_lock(fs);
- bdry_idx = MAP_ENTRY_SPLIT_BOUNDARY_INDEX(fs->entry);
- if (fs->map->timestamp != fs->map_generation) {
+ /*
+ * Do not wait for a map writer while retaining the pager's busy pages
+ * and handoff lock. The writer may need one of those resources.
+ * The old entry may have been freed; use only the saved split boundary
+ * until the map lock and generation have both been revalidated.
+ */
+ fs->lookup_still_valid = vm_map_trylock_read(fs->map);
+ if (!fs->lookup_still_valid ||
+ fs->map->timestamp != fs->map_generation) {
if (bdry_idx == 0) {
vm_fault_populate_cleanup(fs->first_object, pager_first,
pager_last);
} else {
- m = vm_page_lookup(fs->first_object, pager_first);
- if (m != fs->m)
- vm_page_xunbusy(m);
+ m = vm_pager_populate_take_page(fs->first_object,
+ pager_first);
+ if (m != NULL) {
+ vm_fault_populate_check_page(m);
+ vm_page_deactivate(m);
+ vm_pager_populate_release_page(fs->first_object, m);
+ if (pager_first < pager_last)
+ vm_fault_populate_cleanup(
+ fs->first_object, pager_first + 1,
+ pager_last);
+ } else {
+ m = vm_page_lookup(fs->first_object, pager_first);
+ if (m != NULL && m != fs->m)
+ vm_page_xunbusy(m);
+ }
}
+ vm_pager_populate_done(fs->first_object);
return (FAULT_RESTART);
}
/*
* The map is unchanged after our last unlock. Process the fault.
*
* First, the special case of largepage mappings, where
* populate only busies the first page in superpage run.
*/
if (bdry_idx != 0) {
KASSERT(PMAP_HAS_LARGEPAGES,
("missing pmap support for large pages"));
+ m = vm_pager_populate_take_page(fs->first_object,
+ pager_first);
+ if (m != NULL) {
+ vm_fault_populate_check_page(m);
+ vm_page_deactivate(m);
+ vm_pager_populate_release_page(fs->first_object, m);
+ if (pager_first < pager_last)
+ vm_fault_populate_cleanup(fs->first_object,
+ pager_first + 1, pager_last);
+ res = FAULT_FAILURE;
+ goto out;
+ }
m = vm_page_lookup(fs->first_object, pager_first);
vm_fault_populate_check_page(m);
VM_OBJECT_WUNLOCK(fs->first_object);
vaddr = fs->entry->start + IDX_TO_OFF(pager_first) -
fs->entry->offset;
/* assert alignment for entry */
KASSERT((vaddr & (pagesizes[bdry_idx] - 1)) == 0,
("unaligned superpage start %#jx pager_first %#jx offset %#jx vaddr %#jx",
(uintmax_t)fs->entry->start, (uintmax_t)pager_first,
(uintmax_t)fs->entry->offset, (uintmax_t)vaddr));
KASSERT((VM_PAGE_TO_PHYS(m) & (pagesizes[bdry_idx] - 1)) == 0,
("unaligned superpage m %p %#jx", m,
(uintmax_t)VM_PAGE_TO_PHYS(m)));
rv = pmap_enter(fs->map->pmap, vaddr, m, fs->prot,
fs->fault_type | (fs->wired ? PMAP_ENTER_WIRED : 0) |
PMAP_ENTER_LARGEPAGE, bdry_idx);
VM_OBJECT_WLOCK(fs->first_object);
vm_page_xunbusy(m);
if (rv != KERN_SUCCESS) {
res = FAULT_FAILURE;
goto out;
}
if ((fs->fault_flags & VM_FAULT_WIRE) != 0) {
for (i = 0; i < atop(pagesizes[bdry_idx]); i++)
vm_page_wire(m + i);
}
if (fs->m_hold != NULL) {
*fs->m_hold = m + (fs->first_pindex - pager_first);
vm_page_wire(*fs->m_hold);
}
goto out;
}
/*
* The range [pager_first, pager_last] that is given to the
* pager is only a hint. The pager may populate any range
* within the object that includes the requested page index.
* In case the pager expanded the range, clip it to fit into
* the map entry.
*/
map_first = OFF_TO_IDX(fs->entry->offset);
if (map_first > pager_first) {
vm_fault_populate_cleanup(fs->first_object, pager_first,
map_first - 1);
pager_first = map_first;
}
map_last = map_first + atop(fs->entry->end - fs->entry->start) - 1;
if (map_last < pager_last) {
vm_fault_populate_cleanup(fs->first_object, map_last + 1,
pager_last);
pager_last = map_last;
}
for (pidx = pager_first; pidx <= pager_last; pidx += npages) {
bool writeable;
- m = vm_page_lookup(fs->first_object, pidx);
vaddr = fs->entry->start + IDX_TO_OFF(pidx) - fs->entry->offset;
+ m = vm_pager_populate_take_page(fs->first_object, pidx);
+ if (m != NULL) {
+ npages = 1;
+ vm_fault_populate_check_page(m);
+ if (fs->wired ||
+ (fs->fault_flags & VM_FAULT_WIRE) != 0) {
+ vm_page_deactivate(m);
+ vm_pager_populate_release_page(fs->first_object, m);
+ if (pidx < pager_last)
+ vm_fault_populate_cleanup(fs->first_object,
+ pidx + 1, pager_last);
+ res = FAULT_FAILURE;
+ goto out;
+ }
+ vm_fault_dirty(fs, m);
+ VM_OBJECT_WUNLOCK(fs->first_object);
+ rv = pmap_enter(fs->map->pmap, vaddr, m, fs->prot,
+ fs->fault_type | PMAP_ENTER_NOSLEEP, 0);
+ VM_OBJECT_WLOCK(fs->first_object);
+ if (rv != KERN_SUCCESS) {
+ vm_page_deactivate(m);
+ vm_pager_populate_release_page(fs->first_object, m);
+ if (pidx < pager_last)
+ vm_fault_populate_cleanup(fs->first_object,
+ pidx + 1, pager_last);
+ res = rv == KERN_RESOURCE_SHORTAGE ?
+ FAULT_OOM : FAULT_FAILURE;
+ goto out;
+ }
+ vm_page_activate(m);
+ if (fs->m_hold != NULL && pidx == fs->first_pindex) {
+ vm_page_wire(m);
+ external_hold = m;
+ }
+ vm_pager_populate_release_page(fs->first_object, m);
+ continue;
+ }
+ m = vm_page_lookup(fs->first_object, pidx);
KASSERT(m != NULL && m->pindex == pidx,
("%s: pindex mismatch", __func__));
psind = m->psind;
while (psind > 0 && ((vaddr & (pagesizes[psind] - 1)) != 0 ||
pidx + OFF_TO_IDX(pagesizes[psind]) - 1 > pager_last ||
!pmap_ps_enabled(fs->map->pmap)))
psind--;
writeable = (fs->prot & VM_PROT_WRITE) != 0;
npages = atop(pagesizes[psind]);
for (i = 0; i < npages; i++) {
vm_fault_populate_check_page(&m[i]);
vm_fault_dirty(fs, &m[i]);
/*
* If this is a writeable superpage mapping, all
* constituent pages and the new mapping should be
* dirty, otherwise the mapping should be read-only.
*/
if (writeable && psind > 0 &&
(m[i].oflags & VPO_UNMANAGED) == 0 &&
m[i].dirty != VM_PAGE_BITS_ALL)
writeable = false;
}
if (psind > 0 && writeable)
fs->fault_type |= VM_PROT_WRITE;
VM_OBJECT_WUNLOCK(fs->first_object);
rv = pmap_enter(fs->map->pmap, vaddr, m,
fs->prot & ~(writeable ? 0 : VM_PROT_WRITE),
fs->fault_type | (fs->wired ? PMAP_ENTER_WIRED : 0), psind);
/*
* pmap_enter() may fail for a superpage mapping if additional
* protection policies prevent the full mapping.
* For example, this will happen on amd64 if the entire
* address range does not share the same userspace protection
* key. Revert to single-page mappings if this happens.
*/
MPASS(rv == KERN_SUCCESS ||
(psind > 0 && rv == KERN_PROTECTION_FAILURE));
if (__predict_false(psind > 0 &&
rv == KERN_PROTECTION_FAILURE)) {
MPASS(!fs->wired);
for (i = 0; i < npages; i++) {
rv = pmap_enter(fs->map->pmap, vaddr + ptoa(i),
&m[i], fs->prot, fs->fault_type, 0);
MPASS(rv == KERN_SUCCESS);
}
}
VM_OBJECT_WLOCK(fs->first_object);
for (i = 0; i < npages; i++) {
if ((fs->fault_flags & VM_FAULT_WIRE) != 0 &&
m[i].pindex == fs->first_pindex)
vm_page_wire(&m[i]);
else
vm_page_activate(&m[i]);
if (fs->m_hold != NULL &&
m[i].pindex == fs->first_pindex) {
(*fs->m_hold) = &m[i];
vm_page_wire(&m[i]);
}
vm_page_xunbusy(&m[i]);
}
}
out:
+ /* Publish the held page only after the entire handoff succeeds. */
+ if (external_hold != NULL) {
+ if (res == FAULT_SUCCESS)
+ *fs->m_hold = external_hold;
+ else
+ vm_page_unwire(external_hold, PQ_INACTIVE);
+ }
+ vm_pager_populate_done(fs->first_object);
curthread->td_ru.ru_majflt++;
return (res);
}
static int prot_fault_translation;
SYSCTL_INT(_machdep, OID_AUTO, prot_fault_translation, CTLFLAG_RWTUN,
&prot_fault_translation, 0,
"Control signal to deliver on protection fault");
/* compat definition to keep common code for signal translation */
#define UCODE_PAGEFLT 12
#ifdef T_PAGEFLT
_Static_assert(UCODE_PAGEFLT == T_PAGEFLT, "T_PAGEFLT");
#endif
/*
* vm_fault_trap:
*
* Helper for the machine-dependent page fault trap handlers, wrapping
* vm_fault(). Issues ktrace(2) tracepoints for the faults.
*
* If the fault cannot be handled successfully by updating the
* required mapping, and the faulted instruction cannot be restarted,
* the signal number and si_code values are returned for trapsignal()
* to deliver.
*
* Returns Mach error codes, but callers should only check for
* KERN_SUCCESS.
*/
int
vm_fault_trap(vm_map_t map, vm_offset_t vaddr, vm_prot_t fault_type,
int fault_flags, int *signo, int *ucode)
{
int result;
MPASS(signo == NULL || ucode != NULL);
#ifdef KTRACE
if (map != kernel_map && KTRPOINT(curthread, KTR_FAULT))
ktrfault(vaddr, fault_type);
#endif
result = vm_fault(map, trunc_page(vaddr), fault_type, fault_flags,
NULL);
KASSERT(result == KERN_SUCCESS || result == KERN_FAILURE ||
result == KERN_INVALID_ADDRESS ||
result == KERN_RESOURCE_SHORTAGE ||
result == KERN_PROTECTION_FAILURE ||
result == KERN_OUT_OF_BOUNDS,
("Unexpected Mach error %d from vm_fault()", result));
#ifdef KTRACE
if (map != kernel_map && KTRPOINT(curthread, KTR_FAULTEND))
ktrfaultend(result);
#endif
if (result != KERN_SUCCESS && signo != NULL) {
switch (result) {
case KERN_FAILURE:
case KERN_INVALID_ADDRESS:
*signo = SIGSEGV;
*ucode = SEGV_MAPERR;
break;
case KERN_RESOURCE_SHORTAGE:
*signo = SIGBUS;
*ucode = BUS_OOMERR;
break;
case KERN_OUT_OF_BOUNDS:
*signo = SIGBUS;
*ucode = BUS_OBJERR;
break;
case KERN_PROTECTION_FAILURE:
if (prot_fault_translation == 0) {
/*
* Autodetect. This check also covers
* the images without the ABI-tag ELF
* note.
*/
if (SV_CURPROC_ABI() == SV_ABI_FREEBSD &&
curproc->p_osrel >= P_OSREL_SIGSEGV) {
*signo = SIGSEGV;
*ucode = SEGV_ACCERR;
} else {
*signo = SIGBUS;
*ucode = UCODE_PAGEFLT;
}
} else if (prot_fault_translation == 1) {
/* Always compat mode. */
*signo = SIGBUS;
*ucode = UCODE_PAGEFLT;
} else {
/* Always SIGSEGV mode. */
*signo = SIGSEGV;
*ucode = SEGV_ACCERR;
}
break;
default:
KASSERT(0, ("Unexpected Mach error %d from vm_fault()",
result));
break;
}
}
return (result);
}
static bool
vm_fault_object_ensure_wlocked(struct faultstate *fs)
{
if (fs->object == fs->first_object)
VM_OBJECT_ASSERT_WLOCKED(fs->object);
if (!fs->can_read_lock) {
VM_OBJECT_ASSERT_WLOCKED(fs->object);
return (true);
}
if (VM_OBJECT_WOWNED(fs->object))
return (true);
if (VM_OBJECT_TRYUPGRADE(fs->object))
return (true);
return (false);
}
static enum fault_status
vm_fault_lock_vnode(struct faultstate *fs, bool objlocked)
{
struct vnode *vp;
int error, locked;
if (fs->object->type != OBJT_VNODE)
return (FAULT_CONTINUE);
vp = fs->object->handle;
if (vp == fs->vp) {
ASSERT_VOP_LOCKED(vp, "saved vnode is not locked");
return (FAULT_CONTINUE);
}
/*
* Perform an unlock in case the desired vnode changed while
* the map was unlocked during a retry.
*/
vm_fault_unlock_vp(fs);
locked = VOP_ISLOCKED(vp);
if (locked != LK_EXCLUSIVE)
locked = LK_SHARED;
/*
* We must not sleep acquiring the vnode lock while we have
* the page exclusive busied or the object's
* paging-in-progress count incremented. Otherwise, we could
* deadlock.
*/
error = vget(vp, locked | LK_CANRECURSE | LK_NOWAIT);
if (error == 0) {
fs->vp = vp;
return (FAULT_CONTINUE);
}
vhold(vp);
if (objlocked)
vm_fault_unlock_and_deallocate(fs);
else
vm_fault_deallocate(fs);
error = vget(vp, locked | LK_RETRY | LK_CANRECURSE);
vdrop(vp);
fs->vp = vp;
KASSERT(error == 0, ("vm_fault: vget failed %d", error));
return (FAULT_RESTART);
}
/*
* Calculate the desired readahead. Handle drop-behind.
*
* Returns the number of readahead blocks to pass to the pager.
*/
static int
vm_fault_readahead(struct faultstate *fs)
{
int era, nera;
u_char behavior;
KASSERT(fs->lookup_still_valid, ("map unlocked"));
era = fs->entry->read_ahead;
behavior = vm_map_entry_behavior(fs->entry);
if (behavior == MAP_ENTRY_BEHAV_RANDOM) {
nera = 0;
} else if (behavior == MAP_ENTRY_BEHAV_SEQUENTIAL) {
nera = VM_FAULT_READ_AHEAD_MAX;
if (fs->vaddr == fs->entry->next_read)
vm_fault_dontneed(fs, fs->vaddr, nera);
} else if (fs->vaddr == fs->entry->next_read) {
/*
* This is a sequential fault. Arithmetically
* increase the requested number of pages in
* the read-ahead window. The requested
* number of pages is "# of sequential faults
* x (read ahead min + 1) + read ahead min"
*/
nera = VM_FAULT_READ_AHEAD_MIN;
if (era > 0) {
nera += era + 1;
if (nera > VM_FAULT_READ_AHEAD_MAX)
nera = VM_FAULT_READ_AHEAD_MAX;
}
if (era == VM_FAULT_READ_AHEAD_MAX)
vm_fault_dontneed(fs, fs->vaddr, nera);
} else {
/*
* This is a non-sequential fault.
*/
nera = 0;
}
if (era != nera) {
/*
* A read lock on the map suffices to update
* the read ahead count safely.
*/
fs->entry->read_ahead = nera;
}
return (nera);
}
static int
vm_fault_lookup(struct faultstate *fs)
{
int result;
KASSERT(!fs->lookup_still_valid,
("vm_fault_lookup: Map already locked."));
result = vm_map_lookup(&fs->map, fs->vaddr, fs->fault_type |
VM_PROT_FAULT_LOOKUP, &fs->entry, &fs->first_object,
&fs->first_pindex, &fs->prot, &fs->wired);
if (result != KERN_SUCCESS) {
vm_fault_unlock_vp(fs);
return (result);
}
fs->map_generation = fs->map->timestamp;
if (fs->entry->eflags & MAP_ENTRY_NOFAULT) {
panic("%s: fault on nofault entry, addr: %#lx",
__func__, (u_long)fs->vaddr);
}
if (fs->entry->eflags & MAP_ENTRY_IN_TRANSITION &&
fs->entry->wiring_thread != curthread) {
vm_map_unlock_read(fs->map);
vm_map_lock(fs->map);
if (vm_map_lookup_entry(fs->map, fs->vaddr, &fs->entry) &&
(fs->entry->eflags & MAP_ENTRY_IN_TRANSITION)) {
vm_fault_unlock_vp(fs);
fs->entry->eflags |= MAP_ENTRY_NEEDS_WAKEUP;
vm_map_unlock_and_wait(fs->map, 0);
} else
vm_map_unlock(fs->map);
return (KERN_RESOURCE_SHORTAGE);
}
MPASS((fs->entry->eflags & MAP_ENTRY_GUARD) == 0);
if (fs->wired)
fs->fault_type = fs->prot | (fs->fault_type & VM_PROT_COPY);
else
KASSERT((fs->fault_flags & VM_FAULT_WIRE) == 0,
("!fs->wired && VM_FAULT_WIRE"));
fs->lookup_still_valid = true;
return (KERN_SUCCESS);
}
static int
vm_fault_relookup(struct faultstate *fs)
{
vm_object_t retry_object;
vm_pindex_t retry_pindex;
vm_prot_t retry_prot;
int result;
if (!vm_map_trylock_read(fs->map))
return (KERN_RESTART);
fs->lookup_still_valid = true;
if (fs->map->timestamp == fs->map_generation)
return (KERN_SUCCESS);
result = vm_map_lookup_locked(&fs->map, fs->vaddr, fs->fault_type,
&fs->entry, &retry_object, &retry_pindex, &retry_prot,
&fs->wired);
if (result != KERN_SUCCESS) {
/*
* If retry of map lookup would have blocked then
* retry fault from start.
*/
if (result == KERN_FAILURE)
return (KERN_RESTART);
return (result);
}
if (retry_object != fs->first_object ||
retry_pindex != fs->first_pindex)
return (KERN_RESTART);
/*
* Check whether the protection has changed or the object has
* been copied while we left the map unlocked. Changing from
* read to write permission is OK - we leave the page
* write-protected, and catch the write fault. Changing from
* write to read permission means that we can't mark the page
* write-enabled after all.
*/
fs->prot &= retry_prot;
fs->fault_type &= retry_prot;
if (fs->prot == 0)
return (KERN_RESTART);
/* Reassert because wired may have changed. */
KASSERT(fs->wired || (fs->fault_flags & VM_FAULT_WIRE) == 0,
("!wired && VM_FAULT_WIRE"));
return (KERN_SUCCESS);
}
static bool
vm_fault_can_cow_rename(struct faultstate *fs)
{
return (
/* Only one shadow object and no other refs. */
fs->object->shadow_count == 1 && fs->object->ref_count == 1 &&
/* No other ways to look the object up. */
fs->object->handle == NULL && (fs->object->flags & OBJ_ANON) != 0);
}
static void
vm_fault_cow(struct faultstate *fs)
{
bool is_first_object_locked, rename_cow;
KASSERT(vm_fault_might_be_cow(fs),
("source and target COW objects are identical"));
/*
* This allows pages to be virtually copied from a backing_object
* into the first_object, where the backing object has no other
* refs to it, and cannot gain any more refs. Instead of a bcopy,
* we just move the page from the backing object to the first
* object. Note that we must mark the page dirty in the first
* object so that it will go out to swap when needed.
*/
is_first_object_locked = false;
rename_cow = false;
if (vm_fault_can_cow_rename(fs) && vm_page_xbusied(fs->m)) {
/*
* Check that we don't chase down the shadow chain and
* we can acquire locks. Recheck the conditions for
* rename after the shadow chain is stable after the
* object locking.
*/
is_first_object_locked = VM_OBJECT_TRYWLOCK(fs->first_object);
if (is_first_object_locked &&
fs->object == fs->first_object->backing_object) {
if (VM_OBJECT_TRYWLOCK(fs->object)) {
rename_cow = vm_fault_can_cow_rename(fs);
if (!rename_cow)
VM_OBJECT_WUNLOCK(fs->object);
}
}
}
if (rename_cow) {
vm_page_assert_xbusied(fs->m);
/*
* Remove but keep xbusy for replace. fs->m is moved into
* fs->first_object and left busy while fs->first_m is
* conditionally freed.
*/
vm_page_remove_xbusy(fs->m);
vm_page_replace(fs->m, fs->first_object, fs->first_pindex,
fs->first_m);
vm_page_dirty(fs->m);
#if VM_NRESERVLEVEL > 0
/*
* Rename the reservation.
*/
vm_reserv_rename(fs->m, fs->first_object, fs->object,
OFF_TO_IDX(fs->first_object->backing_object_offset));
#endif
VM_OBJECT_WUNLOCK(fs->object);
VM_OBJECT_WUNLOCK(fs->first_object);
fs->first_m = fs->m;
fs->m = NULL;
VM_CNT_INC(v_cow_optim);
} else {
if (is_first_object_locked)
VM_OBJECT_WUNLOCK(fs->first_object);
/*
* Oh, well, lets copy it.
*/
pmap_copy_page(fs->m, fs->first_m);
if (fs->wired && (fs->fault_flags & VM_FAULT_WIRE) == 0) {
vm_page_wire(fs->first_m);
vm_page_unwire(fs->m, PQ_INACTIVE);
}
/*
* Save the COW page to be released after pmap_enter is
* complete. The new copy will be marked valid when we're ready
* to map it.
*/
fs->m_cow = fs->m;
fs->m = NULL;
/*
* Typically, the shadow object is either private to this
* address space (OBJ_ONEMAPPING) or its pages are read only.
* In the highly unusual case where the pages of a shadow object
* are read/write shared between this and other address spaces,
* we need to ensure that any pmap-level mappings to the
* original, copy-on-write page from the backing object are
* removed from those other address spaces.
*
* The flag check is racy, but this is tolerable: if
* OBJ_ONEMAPPING is cleared after the check, the busy state
* ensures that new mappings of m_cow can't be created.
* pmap_enter() will replace an existing mapping in the current
* address space. If OBJ_ONEMAPPING is set after the check,
* removing mappings will at worse trigger some unnecessary page
* faults.
*
* In the fs->m shared busy case, the xbusy state of
* fs->first_m prevents new mappings of fs->m from
* being created because a parallel fault on this
* shadow chain should wait for xbusy on fs->first_m.
*/
if ((fs->first_object->flags & OBJ_ONEMAPPING) == 0)
pmap_remove_all(fs->m_cow);
}
vm_object_pip_wakeup(fs->object);
/*
* Only use the new page below...
*/
fs->object = fs->first_object;
fs->pindex = fs->first_pindex;
fs->m = fs->first_m;
VM_CNT_INC(v_cow_faults);
curthread->td_cow++;
}
static enum fault_next_status
vm_fault_next(struct faultstate *fs)
{
vm_object_t next_object;
if (fs->object == fs->first_object || !fs->can_read_lock)
VM_OBJECT_ASSERT_WLOCKED(fs->object);
else
VM_OBJECT_ASSERT_LOCKED(fs->object);
/*
* The requested page does not exist at this object/
* offset. Remove the invalid page from the object,
* waking up anyone waiting for it, and continue on to
* the next object. However, if this is the top-level
* object, we must leave the busy page in place to
* prevent another process from rushing past us, and
* inserting the page in that object at the same time
* that we are.
*/
if (fs->object == fs->first_object) {
fs->first_m = fs->m;
fs->m = NULL;
} else if (fs->m != NULL) {
if (!vm_fault_object_ensure_wlocked(fs)) {
fs->can_read_lock = false;
vm_fault_unlock_and_deallocate(fs);
return (FAULT_NEXT_RESTART);
}
vm_fault_page_free(&fs->m);
}
/*
* Move on to the next object. Lock the next object before
* unlocking the current one.
*/
next_object = fs->object->backing_object;
if (next_object == NULL)
return (FAULT_NEXT_NOOBJ);
MPASS(fs->first_m != NULL);
KASSERT(fs->object != next_object, ("object loop %p", next_object));
if (fs->can_read_lock)
VM_OBJECT_RLOCK(next_object);
else
VM_OBJECT_WLOCK(next_object);
vm_object_pip_add(next_object, 1);
if (fs->object != fs->first_object)
vm_object_pip_wakeup(fs->object);
fs->pindex += OFF_TO_IDX(fs->object->backing_object_offset);
VM_OBJECT_UNLOCK(fs->object);
fs->object = next_object;
return (FAULT_NEXT_GOTOBJ);
}
static void
vm_fault_zerofill(struct faultstate *fs)
{
/*
* If there's no object left, fill the page in the top
* object with zeros.
*/
if (vm_fault_might_be_cow(fs)) {
vm_object_pip_wakeup(fs->object);
fs->object = fs->first_object;
fs->pindex = fs->first_pindex;
}
MPASS(fs->first_m != NULL);
MPASS(fs->m == NULL);
fs->m = fs->first_m;
fs->first_m = NULL;
/*
* Zero the page if necessary and mark it valid.
*/
if (fs->m_needs_zeroing) {
pmap_zero_page(fs->m);
} else {
#ifdef INVARIANTS
if (vm_check_pg_zero) {
struct sf_buf *sf;
unsigned long *p;
int i;
sched_pin();
sf = sf_buf_alloc(fs->m, SFB_CPUPRIVATE);
p = (unsigned long *)sf_buf_kva(sf);
for (i = 0; i < PAGE_SIZE / sizeof(*p); i++, p++) {
KASSERT(*p == 0,
("zerocheck failed page %p PG_ZERO %d %jx",
fs->m, i, (uintmax_t)*p));
}
sf_buf_free(sf);
sched_unpin();
}
#endif
VM_CNT_INC(v_ozfod);
}
VM_CNT_INC(v_zfod);
vm_page_valid(fs->m);
}
/*
* Initiate page fault after timeout. Returns true if caller should
* do vm_waitpfault() after the call.
*/
static bool
vm_fault_allocate_oom(struct faultstate *fs)
{
struct timeval now;
vm_fault_unlock_and_deallocate(fs);
if (vm_pfault_oom_attempts < 0)
return (true);
if (!fs->oom_started) {
fs->oom_started = true;
getmicrotime(&fs->oom_start_time);
return (true);
}
getmicrotime(&now);
timevalsub(&now, &fs->oom_start_time);
if (now.tv_sec < vm_pfault_oom_attempts * vm_pfault_oom_wait)
return (true);
if (bootverbose)
printf(
"proc %d (%s) failed to alloc page on fault, starting OOM\n",
curproc->p_pid, curproc->p_comm);
vm_pageout_oom(VM_OOM_MEM_PF);
fs->oom_started = false;
return (false);
}
/*
* Allocate a page directly or via the object populate method.
*/
static enum fault_status
vm_fault_allocate(struct faultstate *fs, struct pctrie_iter *pages)
{
struct domainset *dset;
enum fault_status res;
if ((fs->object->flags & OBJ_SIZEVNLOCK) != 0) {
res = vm_fault_lock_vnode(fs, true);
MPASS(res == FAULT_CONTINUE || res == FAULT_RESTART);
if (res == FAULT_RESTART)
return (res);
}
if (fs->pindex >= fs->object->size) {
vm_fault_unlock_and_deallocate(fs);
return (FAULT_OUT_OF_BOUNDS);
}
if (fs->object == fs->first_object &&
(fs->first_object->flags & OBJ_POPULATE) != 0 &&
fs->first_object->shadow_count == 0) {
res = vm_fault_populate(fs);
switch (res) {
case FAULT_SUCCESS:
case FAULT_FAILURE:
+ case FAULT_OUT_OF_BOUNDS:
+ vm_fault_unlock_and_deallocate(fs);
+ return (res);
case FAULT_RESTART:
vm_fault_unlock_and_deallocate(fs);
+ kern_yield(PRI_USER);
return (res);
+ case FAULT_OOM:
+ dset = fs->object->domain.dr_policy;
+ if (dset == NULL)
+ dset = curthread->td_domain.dr_policy;
+ if (vm_fault_allocate_oom(fs))
+ vm_waitpfault(dset, vm_pfault_oom_wait * hz);
+ return (FAULT_RESTART);
case FAULT_CONTINUE:
pctrie_iter_reset(pages);
/*
* Pager's populate() method
* returned VM_PAGER_BAD.
*/
break;
default:
panic("inconsistent return codes");
}
}
/*
* Allocate a new page for this object/offset pair.
*
* If the process has a fatal signal pending, prioritize the allocation
* with the expectation that the process will exit shortly and free some
* pages. In particular, the signal may have been posted by the page
* daemon in an attempt to resolve an out-of-memory condition.
*
* The unlocked read of the p_flag is harmless. At worst, the P_KILLED
* might be not observed here, and allocation fails, causing a restart
* and new reading of the p_flag.
*/
dset = fs->object->domain.dr_policy;
if (dset == NULL)
dset = curthread->td_domain.dr_policy;
if (!vm_page_count_severe_set(&dset->ds_mask) || P_KILLED(curproc)) {
#if VM_NRESERVLEVEL > 0
vm_object_color(fs->object, atop(fs->vaddr) - fs->pindex);
#endif
if (!vm_pager_can_alloc_page(fs->object, fs->pindex)) {
vm_fault_unlock_and_deallocate(fs);
return (FAULT_FAILURE);
}
fs->m = vm_page_alloc_iter(fs->object, fs->pindex,
P_KILLED(curproc) ? VM_ALLOC_SYSTEM : 0, pages);
}
if (fs->m == NULL) {
if (vm_fault_allocate_oom(fs))
vm_waitpfault(dset, vm_pfault_oom_wait * hz);
return (FAULT_RESTART);
}
if (fs->object == fs->first_object)
fs->m_needs_zeroing = (fs->m->flags & PG_ZERO) == 0;
fs->oom_started = false;
return (FAULT_CONTINUE);
}
/*
* Call the pager to retrieve the page if there is a chance
* that the pager has it, and potentially retrieve additional
* pages at the same time.
*/
static enum fault_status
vm_fault_getpages(struct faultstate *fs, int *behindp, int *aheadp)
{
vm_offset_t e_end, e_start;
int ahead, behind, cluster_offset, rv;
enum fault_status status;
u_char behavior;
/*
* Prepare for unlocking the map. Save the map
* entry's start and end addresses, which are used to
* optimize the size of the pager operation below.
* Even if the map entry's addresses change after
* unlocking the map, using the saved addresses is
* safe.
*/
e_start = fs->entry->start;
e_end = fs->entry->end;
behavior = vm_map_entry_behavior(fs->entry);
/*
* If the pager for the current object might have
* the page, then determine the number of additional
* pages to read and potentially reprioritize
* previously read pages for earlier reclamation.
* These operations should only be performed once per
* page fault. Even if the current pager doesn't
* have the page, the number of additional pages to
* read will apply to subsequent objects in the
* shadow chain.
*/
if (fs->nera == -1 && !P_KILLED(curproc))
fs->nera = vm_fault_readahead(fs);
/*
* Release the map lock before locking the vnode or
* sleeping in the pager. (If the current object has
* a shadow, then an earlier iteration of this loop
* may have already unlocked the map.)
*/
vm_fault_unlock_map(fs);
status = vm_fault_lock_vnode(fs, false);
MPASS(status == FAULT_CONTINUE || status == FAULT_RESTART);
if (status == FAULT_RESTART)
return (status);
KASSERT(fs->vp == NULL || !vm_map_is_system(fs->map),
("vm_fault: vnode-backed object mapped by system map"));
/*
* Page in the requested page and hint the pager,
* that it may bring up surrounding pages.
*/
if (fs->nera == -1 || behavior == MAP_ENTRY_BEHAV_RANDOM ||
P_KILLED(curproc)) {
behind = 0;
ahead = 0;
} else {
/* Is this a sequential fault? */
if (fs->nera > 0) {
behind = 0;
ahead = fs->nera;
} else {
/*
* Request a cluster of pages that is
* aligned to a VM_FAULT_READ_DEFAULT
* page offset boundary within the
* object. Alignment to a page offset
* boundary is more likely to coincide
* with the underlying file system
* block than alignment to a virtual
* address boundary.
*/
cluster_offset = fs->pindex % VM_FAULT_READ_DEFAULT;
behind = ulmin(cluster_offset,
atop(fs->vaddr - e_start));
ahead = VM_FAULT_READ_DEFAULT - 1 - cluster_offset;
}
ahead = ulmin(ahead, atop(e_end - fs->vaddr) - 1);
}
*behindp = behind;
*aheadp = ahead;
rv = vm_pager_get_pages(fs->object, &fs->m, 1, behindp, aheadp);
if (rv == VM_PAGER_OK)
return (FAULT_HARD);
if (rv == VM_PAGER_ERROR)
printf("vm_fault: pager read error, pid %d (%s)\n",
curproc->p_pid, curproc->p_comm);
/*
* If an I/O error occurred or the requested page was
* outside the range of the pager, clean up and return
* an error.
*/
if (rv == VM_PAGER_ERROR || rv == VM_PAGER_BAD) {
VM_OBJECT_WLOCK(fs->object);
vm_fault_page_free(&fs->m);
vm_fault_unlock_and_deallocate(fs);
return (FAULT_OUT_OF_BOUNDS);
}
KASSERT(rv == VM_PAGER_FAIL,
("%s: unexpected pager error %d", __func__, rv));
return (FAULT_CONTINUE);
}
/*
* Wait/Retry if the page is busy. We have to do this if the page is
* either exclusive or shared busy because the vm_pager may be using
* read busy for pageouts (and even pageins if it is the vnode pager),
* and we could end up trying to pagein and pageout the same page
* simultaneously.
*
* We allow the busy case on a read fault if the page is valid. We
* cannot under any circumstances mess around with a shared busied
* page except, perhaps, to pmap it. This is controlled by the
* VM_ALLOC_SBUSY bit in the allocflags argument.
*/
static void
vm_fault_busy_sleep(struct faultstate *fs, int allocflags)
{
/*
* Reference the page before unlocking and
* sleeping so that the page daemon is less
* likely to reclaim it.
*/
vm_page_aflag_set(fs->m, PGA_REFERENCED);
if (vm_fault_might_be_cow(fs)) {
vm_fault_page_release(&fs->first_m);
vm_object_pip_wakeup(fs->first_object);
}
vm_object_pip_wakeup(fs->object);
vm_fault_unlock_map(fs);
if (!vm_page_busy_sleep(fs->m, "vmpfw", allocflags))
VM_OBJECT_UNLOCK(fs->object);
VM_CNT_INC(v_intrans);
vm_object_deallocate(fs->first_object);
}
/*
* Handle page lookup, populate, allocate, page-in for the current
* object.
*
* The object is locked on entry and will remain locked with a return
* code of FAULT_CONTINUE so that fault may follow the shadow chain.
* Otherwise, the object will be unlocked upon return.
*/
static enum fault_status
vm_fault_object(struct faultstate *fs, int *behindp, int *aheadp)
{
struct pctrie_iter pages;
enum fault_status res;
bool dead;
if (fs->object == fs->first_object || !fs->can_read_lock)
VM_OBJECT_ASSERT_WLOCKED(fs->object);
else
VM_OBJECT_ASSERT_LOCKED(fs->object);
/*
* If the object is marked for imminent termination, we retry
* here, since the collapse pass has raced with us. Otherwise,
* if we see terminally dead object, return fail.
*/
if ((fs->object->flags & OBJ_DEAD) != 0) {
dead = fs->object->type == OBJT_DEAD;
vm_fault_unlock_and_deallocate(fs);
if (dead)
return (FAULT_PROTECTION_FAILURE);
pause("vmf_de", 1);
return (FAULT_RESTART);
}
/*
* See if the page is resident.
*/
vm_page_iter_init(&pages, fs->object);
fs->m = vm_radix_iter_lookup(&pages, fs->pindex);
if (fs->m != NULL) {
/*
* If the found page is valid, will be either shadowed
* or mapped read-only, and will not be renamed for
* COW, then busy it in shared mode. This allows
* other faults needing this page to proceed in
* parallel.
*
* Unlocked check for validity, rechecked after busy
* is obtained.
*/
if (vm_page_all_valid(fs->m) &&
/*
* No write permissions for the new fs->m mapping,
* or the first object has only one mapping, so
* other writeable COW mappings of fs->m cannot
* appear under us.
*/
(vm_fault_is_read(fs) || vm_fault_might_be_cow(fs)) &&
/*
* fs->m cannot be renamed from object to
* first_object. These conditions will be
* re-checked with proper synchronization in
* vm_fault_cow().
*/
(!vm_fault_can_cow_rename(fs) ||
fs->object != fs->first_object->backing_object)) {
if (!vm_page_trysbusy(fs->m)) {
vm_fault_busy_sleep(fs, VM_ALLOC_SBUSY);
return (FAULT_RESTART);
}
/*
* Now make sure that racily checked
* conditions are still valid.
*/
if (__predict_true(vm_page_all_valid(fs->m) &&
(vm_fault_is_read(fs) ||
vm_fault_might_be_cow(fs)))) {
VM_OBJECT_UNLOCK(fs->object);
return (FAULT_SOFT);
}
vm_page_sunbusy(fs->m);
}
if (!vm_page_tryxbusy(fs->m)) {
vm_fault_busy_sleep(fs, 0);
return (FAULT_RESTART);
}
/*
* The page is marked busy for other processes and the
* pagedaemon. If it is still completely valid we are
* done.
*/
if (vm_page_all_valid(fs->m)) {
VM_OBJECT_UNLOCK(fs->object);
return (FAULT_SOFT);
}
}
/*
* Page is not resident. If the pager might contain the page
* or this is the beginning of the search, allocate a new
* page.
*/
if (fs->m == NULL && (vm_fault_object_needs_getpages(fs->object) ||
fs->object == fs->first_object)) {
if (!vm_fault_object_ensure_wlocked(fs)) {
fs->can_read_lock = false;
vm_fault_unlock_and_deallocate(fs);
return (FAULT_RESTART);
}
res = vm_fault_allocate(fs, &pages);
if (res != FAULT_CONTINUE)
return (res);
}
/*
* Check to see if the pager can possibly satisfy this fault.
* If not, skip to the next object without dropping the lock to
* preserve atomicity of shadow faults.
*/
if (vm_fault_object_needs_getpages(fs->object)) {
/*
* At this point, we have either allocated a new page
* or found an existing page that is only partially
* valid.
*
* We hold a reference on the current object and the
* page is exclusive busied. The exclusive busy
* prevents simultaneous faults and collapses while
* the object lock is dropped.
*/
VM_OBJECT_UNLOCK(fs->object);
res = vm_fault_getpages(fs, behindp, aheadp);
if (res == FAULT_CONTINUE)
VM_OBJECT_WLOCK(fs->object);
} else {
res = FAULT_CONTINUE;
}
return (res);
}
/*
* vm_fault:
*
* Handle a page fault occurring at the given address, requiring the
* given permissions, in the map specified. If successful, the page
* is inserted into the associated physical map, and optionally
* referenced and returned in *m_hold.
*
* The given address should be truncated to the proper page address.
*
* KERN_SUCCESS is returned if the page fault is handled; otherwise, a
* Mach error code explaining why the fault is fatal is returned.
*
* The map in question must be alive, either being the map for the current
* process, or the owner process hold count has been incremented to prevent
* exit().
*
* If the thread private TDP_NOFAULTING flag is set, any fault results
* in immediate protection failure. Otherwise the fault is processed,
* and caller may hold no locks.
*/
int
vm_fault(vm_map_t map, vm_offset_t vaddr, vm_prot_t fault_type,
int fault_flags, vm_page_t *m_hold)
{
struct pctrie_iter pages;
struct faultstate fs;
int ahead, behind, faultcount, rv;
enum fault_status res;
enum fault_next_status res_next;
bool hardfault;
VM_CNT_INC(v_vm_faults);
if ((curthread->td_pflags & TDP_NOFAULTING) != 0)
return (KERN_PROTECTION_FAILURE);
fs.vp = NULL;
fs.vaddr = vaddr;
fs.m_hold = m_hold;
fs.fault_flags = fault_flags;
fs.map = map;
fs.lookup_still_valid = false;
fs.oom_started = false;
fs.nera = -1;
fs.can_read_lock = true;
faultcount = 0;
hardfault = false;
RetryFault:
fs.fault_type = fault_type;
fs.m_needs_zeroing = true;
/*
* Find the backing store object and offset into it to begin the
* search.
*/
rv = vm_fault_lookup(&fs);
if (rv != KERN_SUCCESS) {
if (rv == KERN_RESOURCE_SHORTAGE)
goto RetryFault;
return (rv);
}
/*
* Try to avoid lock contention on the top-level object through
* special-case handling of some types of page faults, specifically,
* those that are mapping an existing page from the top-level object.
* Under this condition, a read lock on the object suffices, allowing
* multiple page faults of a similar type to run in parallel.
*/
if (fs.vp == NULL /* avoid locked vnode leak */ &&
(fs.entry->eflags & MAP_ENTRY_SPLIT_BOUNDARY_MASK) == 0 &&
(fs.fault_flags & (VM_FAULT_WIRE | VM_FAULT_DIRTY)) == 0) {
res = vm_fault_soft_fast(&fs);
if (res == FAULT_SUCCESS) {
VM_OBJECT_ASSERT_UNLOCKED(fs.first_object);
return (KERN_SUCCESS);
}
VM_OBJECT_ASSERT_WLOCKED(fs.first_object);
} else {
vm_page_iter_init(&pages, fs.first_object);
VM_OBJECT_WLOCK(fs.first_object);
}
/*
* Make a reference to this object to prevent its disposal while we
* are messing with it. Once we have the reference, the map is free
* to be diddled. Since objects reference their shadows (and copies),
* they will stay around as well.
*
* Bump the paging-in-progress count to prevent size changes (e.g.
* truncation operations) during I/O.
*/
vm_object_reference_locked(fs.first_object);
vm_object_pip_add(fs.first_object, 1);
fs.m_cow = fs.m = fs.first_m = NULL;
/*
* Search for the page at object/offset.
*/
fs.object = fs.first_object;
fs.pindex = fs.first_pindex;
if ((fs.entry->eflags & MAP_ENTRY_SPLIT_BOUNDARY_MASK) != 0) {
res = vm_fault_allocate(&fs, &pages);
switch (res) {
case FAULT_RESTART:
goto RetryFault;
case FAULT_SUCCESS:
return (KERN_SUCCESS);
case FAULT_FAILURE:
return (KERN_FAILURE);
case FAULT_OUT_OF_BOUNDS:
return (KERN_OUT_OF_BOUNDS);
case FAULT_CONTINUE:
break;
default:
panic("vm_fault: Unhandled status %d", res);
}
}
while (TRUE) {
KASSERT(fs.m == NULL,
("page still set %p at loop start", fs.m));
res = vm_fault_object(&fs, &behind, &ahead);
switch (res) {
case FAULT_SOFT:
goto found;
case FAULT_HARD:
faultcount = behind + 1 + ahead;
hardfault = true;
goto found;
case FAULT_RESTART:
goto RetryFault;
case FAULT_SUCCESS:
return (KERN_SUCCESS);
case FAULT_FAILURE:
return (KERN_FAILURE);
case FAULT_OUT_OF_BOUNDS:
return (KERN_OUT_OF_BOUNDS);
case FAULT_PROTECTION_FAILURE:
return (KERN_PROTECTION_FAILURE);
case FAULT_CONTINUE:
break;
default:
panic("vm_fault: Unhandled status %d", res);
}
/*
* The page was not found in the current object. Try to
* traverse into a backing object or zero fill if none is
* found.
*/
res_next = vm_fault_next(&fs);
if (res_next == FAULT_NEXT_RESTART)
goto RetryFault;
else if (res_next == FAULT_NEXT_GOTOBJ)
continue;
MPASS(res_next == FAULT_NEXT_NOOBJ);
if ((fs.fault_flags & VM_FAULT_NOFILL) != 0) {
if (fs.first_object == fs.object)
vm_fault_page_free(&fs.first_m);
vm_fault_unlock_and_deallocate(&fs);
return (KERN_OUT_OF_BOUNDS);
}
VM_OBJECT_UNLOCK(fs.object);
vm_fault_zerofill(&fs);
/* Don't try to prefault neighboring pages. */
faultcount = 1;
break;
}
found:
/*
* A valid page has been found and busied. The object lock
* must no longer be held if the page was busied.
*
* Regardless of the busy state of fs.m, fs.first_m is always
* exclusively busied after the first iteration of the loop
* calling vm_fault_object(). This is an ordering point for
* the parallel faults occuring in on the same page.
*/
vm_page_assert_busied(fs.m);
VM_OBJECT_ASSERT_UNLOCKED(fs.object);
/*
* If the page is being written, but isn't already owned by the
* top-level object, we have to copy it into a new page owned by the
* top-level object.
*/
if (vm_fault_might_be_cow(&fs)) {
/*
* We only really need to copy if we want to write it.
*/
if ((fs.fault_type & (VM_PROT_COPY | VM_PROT_WRITE)) != 0) {
vm_fault_cow(&fs);
/*
* We only try to prefault read-only mappings to the
* neighboring pages when this copy-on-write fault is
* a hard fault. In other cases, trying to prefault
* is typically wasted effort.
*/
if (faultcount == 0)
faultcount = 1;
} else {
fs.prot &= ~VM_PROT_WRITE;
}
}
/*
* We must verify that the maps have not changed since our last
* lookup.
*/
if (!fs.lookup_still_valid) {
rv = vm_fault_relookup(&fs);
if (rv != KERN_SUCCESS) {
vm_fault_deallocate(&fs);
if (rv == KERN_RESTART)
goto RetryFault;
return (rv);
}
}
VM_OBJECT_ASSERT_UNLOCKED(fs.object);
/*
* If the page was filled by a pager, save the virtual address that
* should be faulted on next under a sequential access pattern to the
* map entry. A read lock on the map suffices to update this address
* safely.
*/
if (hardfault)
fs.entry->next_read = vaddr + ptoa(ahead) + PAGE_SIZE;
/*
* If the page to be mapped was copied from a backing object, we defer
* marking it valid until here, where the fault handler is guaranteed to
* succeed. Otherwise we can end up with a shadowed, mapped page in the
* backing object, which violates an invariant of vm_object_collapse()
* that shadowed pages are not mapped.
*/
if (fs.m_cow != NULL) {
KASSERT(vm_page_none_valid(fs.m),
("vm_fault: page %p is already valid", fs.m_cow));
vm_page_valid(fs.m);
}
/*
* Page must be completely valid or it is not fit to
* map into user space. vm_pager_get_pages() ensures this.
*/
vm_page_assert_busied(fs.m);
KASSERT(vm_page_all_valid(fs.m),
("vm_fault: page %p partially invalid", fs.m));
vm_fault_dirty(&fs, fs.m);
/*
* Put this page into the physical map. We had to do the unlock above
* because pmap_enter() may sleep. We don't put the page
* back on the active queue until later so that the pageout daemon
* won't find it (yet).
*/
pmap_enter(fs.map->pmap, vaddr, fs.m, fs.prot,
fs.fault_type | (fs.wired ? PMAP_ENTER_WIRED : 0), 0);
if (faultcount != 1 && (fs.fault_flags & VM_FAULT_WIRE) == 0 &&
fs.wired == 0)
vm_fault_prefault(&fs, vaddr,
faultcount > 0 ? behind : PFBAK,
faultcount > 0 ? ahead : PFFOR, false);
/*
* If the page is not wired down, then put it where the pageout daemon
* can find it.
*/
if ((fs.fault_flags & VM_FAULT_WIRE) != 0)
vm_page_wire(fs.m);
else
vm_page_activate(fs.m);
if (fs.m_hold != NULL) {
(*fs.m_hold) = fs.m;
vm_page_wire(fs.m);
}
KASSERT(fs.first_object == fs.object || vm_page_xbusied(fs.first_m),
("first_m must be xbusy"));
if (vm_page_xbusied(fs.m))
vm_page_xunbusy(fs.m);
else
vm_page_sunbusy(fs.m);
fs.m = NULL;
/*
* Unlock everything, and return
*/
vm_fault_deallocate(&fs);
if (hardfault) {
VM_CNT_INC(v_io_faults);
curthread->td_ru.ru_majflt++;
#ifdef RACCT
if (racct_enable && fs.object->type == OBJT_VNODE) {
PROC_LOCK(curproc);
if ((fs.fault_type & (VM_PROT_COPY | VM_PROT_WRITE)) != 0) {
racct_add_force(curproc, RACCT_WRITEBPS,
PAGE_SIZE + behind * PAGE_SIZE);
racct_add_force(curproc, RACCT_WRITEIOPS, 1);
} else {
racct_add_force(curproc, RACCT_READBPS,
PAGE_SIZE + ahead * PAGE_SIZE);
racct_add_force(curproc, RACCT_READIOPS, 1);
}
PROC_UNLOCK(curproc);
}
#endif
} else
curthread->td_ru.ru_minflt++;
return (KERN_SUCCESS);
}
/*
* Speed up the reclamation of pages that precede the faulting pindex within
* the first object of the shadow chain. Essentially, perform the equivalent
* to madvise(..., MADV_DONTNEED) on a large cluster of pages that precedes
* the faulting pindex by the cluster size when the pages read by vm_fault()
* cross a cluster-size boundary. The cluster size is the greater of the
* smallest superpage size and VM_FAULT_DONTNEED_MIN.
*
* When "fs->first_object" is a shadow object, the pages in the backing object
* that precede the faulting pindex are deactivated by vm_fault(). So, this
* function must only be concerned with pages in the first object.
*/
static void
vm_fault_dontneed(const struct faultstate *fs, vm_offset_t vaddr, int ahead)
{
struct pctrie_iter pages;
vm_map_entry_t entry;
vm_object_t first_object;
vm_offset_t end, start;
vm_page_t m;
vm_size_t size;
VM_OBJECT_ASSERT_UNLOCKED(fs->object);
first_object = fs->first_object;
/* Neither fictitious nor unmanaged pages can be reclaimed. */
if ((first_object->flags & (OBJ_FICTITIOUS | OBJ_UNMANAGED)) == 0) {
VM_OBJECT_RLOCK(first_object);
size = VM_FAULT_DONTNEED_MIN;
if (MAXPAGESIZES > 1 && size < pagesizes[1])
size = pagesizes[1];
end = rounddown2(vaddr, size);
if (vaddr - end >= size - PAGE_SIZE - ptoa(ahead) &&
(entry = fs->entry)->start < end) {
if (end - entry->start < size)
start = entry->start;
else
start = end - size;
pmap_advise(fs->map->pmap, start, end, MADV_DONTNEED);
vm_page_iter_limit_init(&pages, first_object,
OFF_TO_IDX(entry->offset) +
atop(end - entry->start));
VM_RADIX_FOREACH_FROM(m, &pages,
OFF_TO_IDX(entry->offset) +
atop(start - entry->start)) {
if (!vm_page_all_valid(m) ||
vm_page_busied(m))
continue;
/*
* Don't clear PGA_REFERENCED, since it would
* likely represent a reference by a different
* process.
*
* Typically, at this point, prefetched pages
* are still in the inactive queue. Only
* pages that triggered page faults are in the
* active queue. The test for whether the page
* is in the inactive queue is racy; in the
* worst case we will requeue the page
* unnecessarily.
*/
if (!vm_page_inactive(m))
vm_page_deactivate(m);
}
}
VM_OBJECT_RUNLOCK(first_object);
}
}
/*
* vm_fault_prefault provides a quick way of clustering
* pagefaults into a processes address space. It is a "cousin"
* of vm_map_pmap_enter, except it runs at page fault time instead
* of mmap time.
*/
static void
vm_fault_prefault(const struct faultstate *fs, vm_offset_t addra,
int backward, int forward, bool obj_locked)
{
pmap_t pmap;
vm_map_entry_t entry;
vm_object_t backing_object, lobject;
vm_offset_t addr, starta;
vm_pindex_t pindex;
vm_page_t m;
vm_prot_t prot;
int i;
pmap = fs->map->pmap;
if (pmap != vmspace_pmap(curthread->td_proc->p_vmspace))
return;
entry = fs->entry;
if (addra < backward * PAGE_SIZE) {
starta = entry->start;
} else {
starta = addra - backward * PAGE_SIZE;
if (starta < entry->start)
starta = entry->start;
}
prot = entry->protection;
/*
* If pmap_enter() has enabled write access on a nearby mapping, then
* don't attempt promotion, because it will fail.
*/
if ((fs->prot & VM_PROT_WRITE) != 0)
prot |= VM_PROT_NO_PROMOTE;
/*
* Generate the sequence of virtual addresses that are candidates for
* prefaulting in an outward spiral from the faulting virtual address,
* "addra". Specifically, the sequence is "addra - PAGE_SIZE", "addra
* + PAGE_SIZE", "addra - 2 * PAGE_SIZE", "addra + 2 * PAGE_SIZE", ...
* If the candidate address doesn't have a backing physical page, then
* the loop immediately terminates.
*/
for (i = 0; i < 2 * imax(backward, forward); i++) {
addr = addra + ((i >> 1) + 1) * ((i & 1) == 0 ? -PAGE_SIZE :
PAGE_SIZE);
if (addr > addra + forward * PAGE_SIZE)
addr = 0;
if (addr < starta || addr >= entry->end)
continue;
if (!pmap_is_prefaultable(pmap, addr))
continue;
pindex = ((addr - entry->start) + entry->offset) >> PAGE_SHIFT;
lobject = entry->object.vm_object;
if (!obj_locked)
VM_OBJECT_RLOCK(lobject);
while ((m = vm_page_lookup(lobject, pindex)) == NULL &&
!vm_fault_object_needs_getpages(lobject) &&
(backing_object = lobject->backing_object) != NULL) {
KASSERT((lobject->backing_object_offset & PAGE_MASK) ==
0, ("vm_fault_prefault: unaligned object offset"));
pindex += lobject->backing_object_offset >> PAGE_SHIFT;
VM_OBJECT_RLOCK(backing_object);
if (!obj_locked || lobject != entry->object.vm_object)
VM_OBJECT_RUNLOCK(lobject);
lobject = backing_object;
}
if (m == NULL) {
if (!obj_locked || lobject != entry->object.vm_object)
VM_OBJECT_RUNLOCK(lobject);
break;
}
if (vm_page_all_valid(m) &&
(m->flags & PG_FICTITIOUS) == 0)
pmap_enter_quick(pmap, addr, m, prot);
if (!obj_locked || lobject != entry->object.vm_object)
VM_OBJECT_RUNLOCK(lobject);
}
}
/*
* Hold each of the physical pages that are mapped by the specified
* range of virtual addresses, ["addr", "addr" + "len"), if those
* mappings are valid and allow the specified types of access, "prot".
* If all of the implied pages are successfully held, then the number
* of held pages is assigned to *ppages_count, together with pointers
* to those pages in the array "ma". The returned value is zero.
*
* However, if any of the pages cannot be held, an error is returned,
* and no pages are held.
* Error values:
* ENOMEM - the range is not valid
* EINVAL - the provided vm_page array is too small to hold all pages
* EAGAIN - a page was not mapped, and the thread is in nofaulting mode
* EFAULT - a page with requested permissions cannot be mapped
* (more detailed result from vm_fault() is lost)
*/
int
vm_fault_hold_pages(vm_map_t map, vm_offset_t addr, vm_size_t len,
vm_prot_t prot, vm_page_t *ma, int max_count, int *ppages_count)
{
vm_offset_t end, va;
vm_page_t *mp;
int count, error;
boolean_t pmap_failed;
if (len == 0) {
*ppages_count = 0;
return (0);
}
end = round_page(addr + len);
addr = trunc_page(addr);
if (!vm_map_range_valid(map, addr, end))
return (ENOMEM);
if (atop(end - addr) > max_count)
return (EINVAL);
count = atop(end - addr);
/*
* Most likely, the physical pages are resident in the pmap, so it is
* faster to try pmap_extract_and_hold() first.
*/
pmap_failed = FALSE;
for (mp = ma, va = addr; va < end; mp++, va += PAGE_SIZE) {
*mp = pmap_extract_and_hold(map->pmap, va, prot);
if (*mp == NULL)
pmap_failed = TRUE;
else if ((prot & VM_PROT_WRITE) != 0 &&
(*mp)->dirty != VM_PAGE_BITS_ALL) {
/*
* Explicitly dirty the physical page. Otherwise, the
* caller's changes may go unnoticed because they are
* performed through an unmanaged mapping or by a DMA
* operation.
*
* The object lock is not held here.
* See vm_page_clear_dirty_mask().
*/
vm_page_dirty(*mp);
}
}
if (pmap_failed) {
/*
* One or more pages could not be held by the pmap. Either no
* page was mapped at the specified virtual address or that
* mapping had insufficient permissions. Attempt to fault in
* and hold these pages.
*
* If vm_fault_disable_pagefaults() was called,
* i.e., TDP_NOFAULTING is set, we must not sleep nor
* acquire MD VM locks, which means we must not call
* vm_fault(). Some (out of tree) callers mark
* too wide a code area with vm_fault_disable_pagefaults()
* already, use the VM_PROT_QUICK_NOFAULT flag to request
* the proper behaviour explicitly.
*/
if ((prot & VM_PROT_QUICK_NOFAULT) != 0 &&
(curthread->td_pflags & TDP_NOFAULTING) != 0) {
error = EAGAIN;
goto fail;
}
for (mp = ma, va = addr; va < end; mp++, va += PAGE_SIZE) {
if (*mp == NULL && vm_fault(map, va, prot,
VM_FAULT_NORMAL, mp) != KERN_SUCCESS) {
error = EFAULT;
goto fail;
}
}
}
*ppages_count = count;
return (0);
fail:
for (mp = ma; mp < ma + count; mp++)
if (*mp != NULL)
vm_page_unwire(*mp, PQ_INACTIVE);
return (error);
}
/*
* Hold each of the physical pages that are mapped by the specified range of
* virtual addresses, ["addr", "addr" + "len"), if those mappings are valid
* and allow the specified types of access, "prot". If all of the implied
* pages are successfully held, then the number of held pages is returned
* together with pointers to those pages in the array "ma". However, if any
* of the pages cannot be held, -1 is returned.
*/
int
vm_fault_quick_hold_pages(vm_map_t map, vm_offset_t addr, vm_size_t len,
vm_prot_t prot, vm_page_t *ma, int max_count)
{
int error, pages_count;
error = vm_fault_hold_pages(map, addr, len, prot, ma,
max_count, &pages_count);
if (error != 0) {
if (error == EINVAL)
panic("vm_fault_quick_hold_pages: count > max_count");
return (-1);
}
return (pages_count);
}
/*
* Routine:
* vm_fault_copy_entry
* Function:
* Create new object backing dst_entry with private copy of all
* underlying pages. When src_entry is equal to dst_entry, function
* implements COW for wired-down map entry. Otherwise, it forks
* wired entry into dst_map.
*
* In/out conditions:
* The source and destination maps must be locked for write.
* The source map entry must be wired down (or be a sharing map
* entry corresponding to a main map entry that is wired down).
*/
void
vm_fault_copy_entry(vm_map_t dst_map, vm_map_t src_map __unused,
vm_map_entry_t dst_entry, vm_map_entry_t src_entry,
vm_ooffset_t *fork_charge)
{
struct pctrie_iter pages;
vm_object_t backing_object, dst_object, object, src_object;
vm_pindex_t dst_pindex, pindex, src_pindex;
vm_prot_t access, prot;
vm_offset_t vaddr;
vm_page_t dst_m;
vm_page_t src_m;
bool upgrade;
upgrade = src_entry == dst_entry;
KASSERT(upgrade || dst_entry->object.vm_object == NULL,
("vm_fault_copy_entry: vm_object not NULL"));
/*
* If not an upgrade, then enter the mappings in the pmap as
* read and/or execute accesses. Otherwise, enter them as
* write accesses.
*
* A writeable large page mapping is only created if all of
* the constituent small page mappings are modified. Marking
* PTEs as modified on inception allows promotion to happen
* without taking potentially large number of soft faults.
*/
access = prot = dst_entry->protection;
if (!upgrade)
access &= ~VM_PROT_WRITE;
src_object = src_entry->object.vm_object;
src_pindex = OFF_TO_IDX(src_entry->offset);
if (upgrade && (dst_entry->eflags & MAP_ENTRY_NEEDS_COPY) == 0) {
dst_object = src_object;
vm_object_reference(dst_object);
} else {
/*
* Create the top-level object for the destination entry.
* Doesn't actually shadow anything - we copy the pages
* directly.
*/
dst_object = vm_object_allocate_anon(atop(dst_entry->end -
dst_entry->start), NULL, NULL);
#if VM_NRESERVLEVEL > 0
dst_object->flags |= OBJ_COLORED;
dst_object->pg_color = atop(dst_entry->start);
#endif
dst_object->domain = src_object->domain;
dst_entry->object.vm_object = dst_object;
dst_entry->offset = 0;
dst_entry->eflags &= ~MAP_ENTRY_VN_EXEC;
}
VM_OBJECT_WLOCK(dst_object);
if (fork_charge != NULL) {
KASSERT(dst_entry->cred == NULL,
("vm_fault_copy_entry: leaked swp charge"));
dst_object->cred = curthread->td_ucred;
crhold(dst_object->cred);
*fork_charge += ptoa(dst_object->size);
} else if ((dst_object->flags & OBJ_SWAP) != 0 &&
dst_object->cred == NULL) {
KASSERT(dst_entry->cred != NULL, ("no cred for entry %p",
dst_entry));
dst_object->cred = dst_entry->cred;
dst_entry->cred = NULL;
}
/*
* Loop through all of the virtual pages within the entry's
* range, copying each page from the source object to the
* destination object. Since the source is wired, those pages
* must exist. In contrast, the destination is pageable.
* Since the destination object doesn't share any backing storage
* with the source object, all of its pages must be dirtied,
* regardless of whether they can be written.
*/
vm_page_iter_init(&pages, dst_object);
for (vaddr = dst_entry->start, dst_pindex = 0;
vaddr < dst_entry->end;
vaddr += PAGE_SIZE, dst_pindex++) {
again:
/*
* Find the page in the source object, and copy it in.
* Because the source is wired down, the page will be
* in memory.
*/
if (src_object != dst_object)
VM_OBJECT_RLOCK(src_object);
object = src_object;
pindex = src_pindex + dst_pindex;
while ((src_m = vm_page_lookup(object, pindex)) == NULL &&
(backing_object = object->backing_object) != NULL) {
/*
* Unless the source mapping is read-only or
* it is presently being upgraded from
* read-only, the first object in the shadow
* chain should provide all of the pages. In
* other words, this loop body should never be
* executed when the source mapping is already
* read/write.
*/
KASSERT((src_entry->protection & VM_PROT_WRITE) == 0 ||
upgrade,
("vm_fault_copy_entry: main object missing page"));
VM_OBJECT_RLOCK(backing_object);
pindex += OFF_TO_IDX(object->backing_object_offset);
if (object != dst_object)
VM_OBJECT_RUNLOCK(object);
object = backing_object;
}
KASSERT(src_m != NULL, ("vm_fault_copy_entry: page missing"));
if (object != dst_object) {
/*
* Allocate a page in the destination object.
*/
pindex = (src_object == dst_object ? src_pindex : 0) +
dst_pindex;
dst_m = vm_page_alloc_iter(dst_object, pindex,
VM_ALLOC_NORMAL, &pages);
if (dst_m == NULL) {
VM_OBJECT_WUNLOCK(dst_object);
VM_OBJECT_RUNLOCK(object);
vm_wait(dst_object);
VM_OBJECT_WLOCK(dst_object);
pctrie_iter_reset(&pages);
goto again;
}
/*
* See the comment in vm_fault_cow().
*/
if (src_object == dst_object &&
(object->flags & OBJ_ONEMAPPING) == 0)
pmap_remove_all(src_m);
pmap_copy_page(src_m, dst_m);
/*
* The object lock does not guarantee that "src_m" will
* transition from invalid to valid, but it does ensure
* that "src_m" will not transition from valid to
* invalid.
*/
dst_m->dirty = dst_m->valid = src_m->valid;
VM_OBJECT_RUNLOCK(object);
} else {
dst_m = src_m;
if (vm_page_busy_acquire(
dst_m, VM_ALLOC_WAITFAIL) == 0) {
pctrie_iter_reset(&pages);
goto again;
}
if (dst_m->pindex >= dst_object->size) {
/*
* We are upgrading. Index can occur
* out of bounds if the object type is
* vnode and the file was truncated.
*/
vm_page_xunbusy(dst_m);
break;
}
}
/*
* Enter it in the pmap. If a wired, copy-on-write
* mapping is being replaced by a write-enabled
* mapping, then wire that new mapping.
*
* The page can be invalid if the user called
* msync(MS_INVALIDATE) or truncated the backing vnode
* or shared memory object. In this case, do not
* insert it into pmap, but still do the copy so that
* all copies of the wired map entry have similar
* backing pages.
*/
if (vm_page_all_valid(dst_m)) {
VM_OBJECT_WUNLOCK(dst_object);
pmap_enter(dst_map->pmap, vaddr, dst_m, prot,
access | (upgrade ? PMAP_ENTER_WIRED : 0), 0);
VM_OBJECT_WLOCK(dst_object);
}
/*
* Mark it no longer busy, and put it on the active list.
*/
if (upgrade) {
if (src_m != dst_m) {
vm_page_unwire(src_m, PQ_INACTIVE);
vm_page_wire(dst_m);
} else {
KASSERT(vm_page_wired(dst_m),
("dst_m %p is not wired", dst_m));
}
} else {
vm_page_activate(dst_m);
}
vm_page_xunbusy(dst_m);
}
VM_OBJECT_WUNLOCK(dst_object);
if (upgrade) {
dst_entry->eflags &= ~(MAP_ENTRY_COW | MAP_ENTRY_NEEDS_COPY);
vm_object_deallocate(src_object);
}
}
/*
* Block entry into the machine-independent layer's page fault handler by
* the calling thread. Subsequent calls to vm_fault() by that thread will
* return KERN_PROTECTION_FAILURE. Enable machine-dependent handling of
* spurious page faults.
*/
int
vm_fault_disable_pagefaults(void)
{
return (curthread_pflags_set(TDP_NOFAULTING | TDP_RESETSPUR));
}
void
vm_fault_enable_pagefaults(int save)
{
curthread_pflags_restore(save);
}
diff --git a/sys/vm/vm_map.c b/sys/vm/vm_map.c
index 94dd7d3a19bc5a0658ccf74c983aa45dad257c8a..c2d70352c1a664b318a47196cf79264bf2cfb58c 100644
--- a/sys/vm/vm_map.c
+++ b/sys/vm/vm_map.c
@@ -1,5501 +1,5512 @@
/*-
* SPDX-License-Identifier: (BSD-3-Clause AND MIT-CMU)
*
* Copyright (c) 1991, 1993
* The Regents of the University of California. All rights reserved.
*
* This code is derived from software contributed to Berkeley by
* The Mach Operating System project at Carnegie-Mellon University.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice, this list of conditions and the following disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
* 3. Neither the name of the University nor the names of its contributors
* may be used to endorse or promote products derived from this software
* without specific prior written permission.
*
* THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS ``AS IS'' AND
* ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
* ARE DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS
* OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)
* HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT
* LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY
* OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
* SUCH DAMAGE.
*
*
* Copyright (c) 1987, 1990 Carnegie-Mellon University.
* All rights reserved.
*
* Authors: Avadis Tevanian, Jr., Michael Wayne Young
*
* Permission to use, copy, modify and distribute this software and
* its documentation is hereby granted, provided that both the copyright
* notice and this permission notice appear in all copies of the
* software, derivative works or modified versions, and any portions
* thereof, and that both notices appear in supporting documentation.
*
* CARNEGIE MELLON ALLOWS FREE USE OF THIS SOFTWARE IN ITS "AS IS"
* CONDITION. CARNEGIE MELLON DISCLAIMS ANY LIABILITY OF ANY KIND
* FOR ANY DAMAGES WHATSOEVER RESULTING FROM THE USE OF THIS SOFTWARE.
*
* Carnegie Mellon requests users of this software to return to
*
* Software Distribution Coordinator or Software.Distribution@CS.CMU.EDU
* School of Computer Science
* Carnegie Mellon University
* Pittsburgh PA 15213-3890
*
* any improvements or extensions that they make and grant Carnegie the
* rights to redistribute these changes.
*/
/*
* Virtual memory mapping module.
*/
#include <sys/param.h>
#include <sys/systm.h>
#include <sys/elf.h>
#include <sys/kernel.h>
#include <sys/ktr.h>
#include <sys/lock.h>
#include <sys/mutex.h>
#include <sys/proc.h>
#include <sys/vmmeter.h>
#include <sys/mman.h>
#include <sys/vnode.h>
#include <sys/racct.h>
#include <sys/resourcevar.h>
#include <sys/rwlock.h>
#include <sys/file.h>
#include <sys/sysctl.h>
#include <sys/sysent.h>
#include <sys/shm.h>
#include <vm/vm.h>
#include <vm/vm_param.h>
#include <vm/pmap.h>
#include <vm/vm_map.h>
#include <vm/vm_page.h>
#include <vm/vm_pageout.h>
#include <vm/vm_object.h>
#include <vm/vm_pager.h>
#include <vm/vm_radix.h>
#include <vm/vm_kern.h>
#include <vm/vm_extern.h>
#include <vm/vnode_pager.h>
#include <vm/swap_pager.h>
#include <vm/uma.h>
/*
* Virtual memory maps provide for the mapping, protection,
* and sharing of virtual memory objects. In addition,
* this module provides for an efficient virtual copy of
* memory from one map to another.
*
* Synchronization is required prior to most operations.
*
* Maps consist of an ordered doubly-linked list of simple
* entries; a self-adjusting binary search tree of these
* entries is used to speed up lookups.
*
* Since portions of maps are specified by start/end addresses,
* which may not align with existing map entries, all
* routines merely "clip" entries to these start/end values.
* [That is, an entry is split into two, bordering at a
* start or end value.] Note that these clippings may not
* always be necessary (as the two resulting entries are then
* not changed); however, the clipping is done for convenience.
*
* As mentioned above, virtual copy operations are performed
* by copying VM object references from one map to
* another, and then marking both regions as copy-on-write.
*/
static struct mtx map_sleep_mtx;
static uma_zone_t mapentzone;
static uma_zone_t kmapentzone;
static uma_zone_t vmspace_zone;
static int vmspace_zinit(void *mem, int size, int flags);
static void _vm_map_init(vm_map_t map, pmap_t pmap, vm_offset_t min,
vm_offset_t max);
static void vm_map_entry_deallocate(vm_map_entry_t entry, boolean_t system_map);
static void vm_map_entry_dispose(vm_map_t map, vm_map_entry_t entry);
static void vm_map_entry_unwire(vm_map_t map, vm_map_entry_t entry);
static int vm_map_growstack(vm_map_t map, vm_offset_t addr,
vm_map_entry_t gap_entry);
static void vm_map_pmap_enter(vm_map_t map, vm_offset_t addr, vm_prot_t prot,
vm_object_t object, vm_pindex_t pindex, vm_size_t size, int flags);
#ifdef INVARIANTS
static void vmspace_zdtor(void *mem, int size, void *arg);
#endif
static int vm_map_stack_locked(vm_map_t map, vm_offset_t addrbos,
vm_size_t max_ssize, vm_size_t growsize, vm_prot_t prot, vm_prot_t max,
int cow);
static void vm_map_wire_entry_failure(vm_map_t map, vm_map_entry_t entry,
vm_offset_t failed_addr);
#define CONTAINS_BITS(set, bits) ((~(set) & (bits)) == 0)
#define ENTRY_CHARGED(e) ((e)->cred != NULL || \
((e)->object.vm_object != NULL && (e)->object.vm_object->cred != NULL && \
!((e)->eflags & MAP_ENTRY_NEEDS_COPY)))
/*
* PROC_VMSPACE_{UN,}LOCK() can be a noop as long as vmspaces are type
* stable.
*/
#define PROC_VMSPACE_LOCK(p) do { } while (0)
#define PROC_VMSPACE_UNLOCK(p) do { } while (0)
/*
* VM_MAP_RANGE_CHECK: [ internal use only ]
*
* Asserts that the starting and ending region
* addresses fall within the valid range of the map.
*/
#define VM_MAP_RANGE_CHECK(map, start, end) \
{ \
if (start < vm_map_min(map)) \
start = vm_map_min(map); \
if (end > vm_map_max(map)) \
end = vm_map_max(map); \
if (start > end) \
start = end; \
}
#ifndef UMA_USE_DMAP
/*
* Allocate a new slab for kernel map entries. The kernel map may be locked or
* unlocked, depending on whether the request is coming from the kernel map or a
* submap. This function allocates a virtual address range directly from the
* kernel map instead of the kmem_* layer to avoid recursion on the kernel map
* lock and also to avoid triggering allocator recursion in the vmem boundary
* tag allocator.
*/
static void *
kmapent_alloc(uma_zone_t zone, vm_size_t bytes, int domain, uint8_t *pflag,
int wait)
{
vm_offset_t addr;
int error, locked;
*pflag = UMA_SLAB_PRIV;
if (!(locked = vm_map_locked(kernel_map)))
vm_map_lock(kernel_map);
addr = vm_map_findspace(kernel_map, vm_map_min(kernel_map), bytes);
if (addr + bytes < addr || addr + bytes > vm_map_max(kernel_map))
panic("%s: kernel map is exhausted", __func__);
error = vm_map_insert(kernel_map, NULL, 0, addr, addr + bytes,
VM_PROT_RW, VM_PROT_RW, MAP_NOFAULT);
if (error != KERN_SUCCESS)
panic("%s: vm_map_insert() failed: %d", __func__, error);
if (!locked)
vm_map_unlock(kernel_map);
error = kmem_back_domain(domain, kernel_object, addr, bytes, M_NOWAIT |
M_USE_RESERVE | (wait & M_ZERO));
if (error == KERN_SUCCESS) {
return ((void *)addr);
} else {
if (!locked)
vm_map_lock(kernel_map);
vm_map_delete(kernel_map, addr, bytes);
if (!locked)
vm_map_unlock(kernel_map);
return (NULL);
}
}
static void
kmapent_free(void *item, vm_size_t size, uint8_t pflag)
{
vm_offset_t addr;
int error __diagused;
if ((pflag & UMA_SLAB_PRIV) == 0)
/* XXX leaked */
return;
addr = (vm_offset_t)item;
kmem_unback(kernel_object, addr, size);
error = vm_map_remove(kernel_map, addr, addr + size);
KASSERT(error == KERN_SUCCESS,
("%s: vm_map_remove failed: %d", __func__, error));
}
/*
* The worst-case upper bound on the number of kernel map entries that may be
* created before the zone must be replenished in _vm_map_unlock().
*/
#define KMAPENT_RESERVE 1
#endif /* !UMD_MD_SMALL_ALLOC */
/*
* vm_map_startup:
*
* Initialize the vm_map module. Must be called before any other vm_map
* routines.
*
* User map and entry structures are allocated from the general purpose
* memory pool. Kernel maps are statically defined. Kernel map entries
* require special handling to avoid recursion; see the comments above
* kmapent_alloc() and in vm_map_entry_create().
*/
void
vm_map_startup(void)
{
mtx_init(&map_sleep_mtx, "vm map sleep mutex", NULL, MTX_DEF);
/*
* Disable the use of per-CPU buckets: map entry allocation is
* serialized by the kernel map lock.
*/
kmapentzone = uma_zcreate("KMAP ENTRY", sizeof(struct vm_map_entry),
NULL, NULL, NULL, NULL, UMA_ALIGN_PTR,
UMA_ZONE_VM | UMA_ZONE_NOBUCKET);
#ifndef UMA_USE_DMAP
/* Reserve an extra map entry for use when replenishing the reserve. */
uma_zone_reserve(kmapentzone, KMAPENT_RESERVE + 1);
uma_prealloc(kmapentzone, KMAPENT_RESERVE + 1);
uma_zone_set_allocf(kmapentzone, kmapent_alloc);
uma_zone_set_freef(kmapentzone, kmapent_free);
#endif
mapentzone = uma_zcreate("MAP ENTRY", sizeof(struct vm_map_entry),
NULL, NULL, NULL, NULL, UMA_ALIGN_PTR, 0);
vmspace_zone = uma_zcreate("VMSPACE", sizeof(struct vmspace), NULL,
#ifdef INVARIANTS
vmspace_zdtor,
#else
NULL,
#endif
vmspace_zinit, NULL, UMA_ALIGN_PTR, UMA_ZONE_NOFREE);
}
static int
vmspace_zinit(void *mem, int size, int flags)
{
struct vmspace *vm;
vm_map_t map;
vm = (struct vmspace *)mem;
map = &vm->vm_map;
memset(map, 0, sizeof(*map)); /* set MAP_SYSTEM_MAP to false */
sx_init(&map->lock, "vm map (user)");
PMAP_LOCK_INIT(vmspace_pmap(vm));
return (0);
}
#ifdef INVARIANTS
static void
vmspace_zdtor(void *mem, int size, void *arg)
{
struct vmspace *vm;
vm = (struct vmspace *)mem;
KASSERT(vm->vm_map.nentries == 0,
("vmspace %p nentries == %d on free", vm, vm->vm_map.nentries));
KASSERT(vm->vm_map.size == 0,
("vmspace %p size == %ju on free", vm, (uintmax_t)vm->vm_map.size));
}
#endif /* INVARIANTS */
/*
* Allocate a vmspace structure, including a vm_map and pmap,
* and initialize those structures. The refcnt is set to 1.
*/
struct vmspace *
vmspace_alloc(vm_offset_t min, vm_offset_t max, pmap_pinit_t pinit)
{
struct vmspace *vm;
vm = uma_zalloc(vmspace_zone, M_WAITOK);
KASSERT(vm->vm_map.pmap == NULL, ("vm_map.pmap must be NULL"));
if (!pinit(vmspace_pmap(vm))) {
uma_zfree(vmspace_zone, vm);
return (NULL);
}
CTR1(KTR_VM, "vmspace_alloc: %p", vm);
_vm_map_init(&vm->vm_map, vmspace_pmap(vm), min, max);
refcount_init(&vm->vm_refcnt, 1);
vm->vm_shm = NULL;
vm->vm_swrss = 0;
vm->vm_tsize = 0;
vm->vm_dsize = 0;
vm->vm_ssize = 0;
vm->vm_taddr = 0;
vm->vm_daddr = 0;
vm->vm_maxsaddr = 0;
return (vm);
}
#ifdef RACCT
static void
vmspace_container_reset(struct proc *p)
{
PROC_LOCK(p);
racct_set(p, RACCT_DATA, 0);
racct_set(p, RACCT_STACK, 0);
racct_set(p, RACCT_RSS, 0);
racct_set(p, RACCT_MEMLOCK, 0);
racct_set(p, RACCT_VMEM, 0);
PROC_UNLOCK(p);
}
#endif
static inline void
vmspace_dofree(struct vmspace *vm)
{
CTR1(KTR_VM, "vmspace_free: %p", vm);
/*
* Make sure any SysV shm is freed, it might not have been in
* exit1().
*/
shmexit(vm);
/*
* Lock the map, to wait out all other references to it.
* Delete all of the mappings and pages they hold, then call
* the pmap module to reclaim anything left.
*/
(void)vm_map_remove(&vm->vm_map, vm_map_min(&vm->vm_map),
vm_map_max(&vm->vm_map));
pmap_release(vmspace_pmap(vm));
vm->vm_map.pmap = NULL;
uma_zfree(vmspace_zone, vm);
}
void
vmspace_free(struct vmspace *vm)
{
WITNESS_WARN(WARN_GIANTOK | WARN_SLEEPOK, NULL,
"vmspace_free() called");
if (refcount_release(&vm->vm_refcnt))
vmspace_dofree(vm);
}
void
vmspace_exitfree(struct proc *p)
{
struct vmspace *vm;
PROC_VMSPACE_LOCK(p);
vm = p->p_vmspace;
p->p_vmspace = NULL;
PROC_VMSPACE_UNLOCK(p);
KASSERT(vm == &vmspace0, ("vmspace_exitfree: wrong vmspace"));
vmspace_free(vm);
}
void
vmspace_exit(struct thread *td)
{
struct vmspace *vm;
struct proc *p;
bool released;
p = td->td_proc;
vm = p->p_vmspace;
/*
* Prepare to release the vmspace reference. The thread that releases
* the last reference is responsible for tearing down the vmspace.
* However, threads not releasing the final reference must switch to the
* kernel's vmspace0 before the decrement so that the subsequent pmap
* deactivation does not modify a freed vmspace.
*/
refcount_acquire(&vmspace0.vm_refcnt);
if (!(released = refcount_release_if_last(&vm->vm_refcnt))) {
if (p->p_vmspace != &vmspace0) {
PROC_VMSPACE_LOCK(p);
p->p_vmspace = &vmspace0;
PROC_VMSPACE_UNLOCK(p);
pmap_activate(td);
}
released = refcount_release(&vm->vm_refcnt);
}
if (released) {
/*
* pmap_remove_pages() expects the pmap to be active, so switch
* back first if necessary.
*/
if (p->p_vmspace != vm) {
PROC_VMSPACE_LOCK(p);
p->p_vmspace = vm;
PROC_VMSPACE_UNLOCK(p);
pmap_activate(td);
}
pmap_remove_pages(vmspace_pmap(vm));
PROC_VMSPACE_LOCK(p);
p->p_vmspace = &vmspace0;
PROC_VMSPACE_UNLOCK(p);
pmap_activate(td);
vmspace_dofree(vm);
}
#ifdef RACCT
if (racct_enable)
vmspace_container_reset(p);
#endif
}
/* Acquire reference to vmspace owned by another process. */
struct vmspace *
vmspace_acquire_ref(struct proc *p)
{
struct vmspace *vm;
PROC_VMSPACE_LOCK(p);
vm = p->p_vmspace;
if (vm == NULL || !refcount_acquire_if_not_zero(&vm->vm_refcnt)) {
PROC_VMSPACE_UNLOCK(p);
return (NULL);
}
if (vm != p->p_vmspace) {
PROC_VMSPACE_UNLOCK(p);
vmspace_free(vm);
return (NULL);
}
PROC_VMSPACE_UNLOCK(p);
return (vm);
}
/*
* Switch between vmspaces in an AIO kernel process.
*
* The new vmspace is either the vmspace of a user process obtained
* from an active AIO request or the initial vmspace of the AIO kernel
* process (when it is idling). Because user processes will block to
* drain any active AIO requests before proceeding in exit() or
* execve(), the reference count for vmspaces from AIO requests can
* never be 0. Similarly, AIO kernel processes hold an extra
* reference on their initial vmspace for the life of the process. As
* a result, the 'newvm' vmspace always has a non-zero reference
* count. This permits an additional reference on 'newvm' to be
* acquired via a simple atomic increment rather than the loop in
* vmspace_acquire_ref() above.
*/
void
vmspace_switch_aio(struct vmspace *newvm)
{
struct vmspace *oldvm;
/* XXX: Need some way to assert that this is an aio daemon. */
KASSERT(refcount_load(&newvm->vm_refcnt) > 0,
("vmspace_switch_aio: newvm unreferenced"));
oldvm = curproc->p_vmspace;
if (oldvm == newvm)
return;
/*
* Point to the new address space and refer to it.
*/
curproc->p_vmspace = newvm;
refcount_acquire(&newvm->vm_refcnt);
/* Activate the new mapping. */
pmap_activate(curthread);
vmspace_free(oldvm);
}
void
_vm_map_lock(vm_map_t map, const char *file, int line)
{
if (vm_map_is_system(map))
mtx_lock_flags_(&map->system_mtx, 0, file, line);
else
sx_xlock_(&map->lock, file, line);
map->timestamp++;
}
void
vm_map_entry_set_vnode_text(vm_map_entry_t entry, bool add)
{
vm_object_t object;
struct vnode *vp;
bool vp_held;
if ((entry->eflags & MAP_ENTRY_VN_EXEC) == 0)
return;
KASSERT((entry->eflags & MAP_ENTRY_IS_SUB_MAP) == 0,
("Submap with execs"));
object = entry->object.vm_object;
KASSERT(object != NULL, ("No object for text, entry %p", entry));
if ((object->flags & OBJ_ANON) != 0)
object = object->handle;
else
KASSERT(object->backing_object == NULL,
("non-anon object %p shadows", object));
KASSERT(object != NULL, ("No content object for text, entry %p obj %p",
entry, entry->object.vm_object));
/*
* Mostly, we do not lock the backing object. It is
* referenced by the entry we are processing, so it cannot go
* away.
*/
vm_pager_getvp(object, &vp, &vp_held);
if (vp != NULL) {
if (add) {
VOP_SET_TEXT_CHECKED(vp);
} else {
vn_lock(vp, LK_SHARED | LK_RETRY);
VOP_UNSET_TEXT_CHECKED(vp);
VOP_UNLOCK(vp);
}
if (vp_held)
vdrop(vp);
}
}
/*
* Use a different name for this vm_map_entry field when it's use
* is not consistent with its use as part of an ordered search tree.
*/
#define defer_next right
static void
vm_map_process_deferred(void)
{
struct thread *td;
vm_map_entry_t entry, next;
vm_object_t object;
td = curthread;
entry = td->td_map_def_user;
td->td_map_def_user = NULL;
while (entry != NULL) {
next = entry->defer_next;
MPASS((entry->eflags & (MAP_ENTRY_WRITECNT |
MAP_ENTRY_VN_EXEC)) != (MAP_ENTRY_WRITECNT |
MAP_ENTRY_VN_EXEC));
if ((entry->eflags & MAP_ENTRY_WRITECNT) != 0) {
/*
* Decrement the object's writemappings and
* possibly the vnode's v_writecount.
*/
KASSERT((entry->eflags & MAP_ENTRY_IS_SUB_MAP) == 0,
("Submap with writecount"));
object = entry->object.vm_object;
KASSERT(object != NULL, ("No object for writecount"));
vm_pager_release_writecount(object, entry->start,
entry->end);
}
vm_map_entry_set_vnode_text(entry, false);
vm_map_entry_deallocate(entry, FALSE);
entry = next;
}
}
#ifdef INVARIANTS
static void
_vm_map_assert_locked(vm_map_t map, const char *file, int line)
{
if (vm_map_is_system(map))
mtx_assert_(&map->system_mtx, MA_OWNED, file, line);
else
sx_assert_(&map->lock, SA_XLOCKED, file, line);
}
#define VM_MAP_ASSERT_LOCKED(map) \
_vm_map_assert_locked(map, LOCK_FILE, LOCK_LINE)
enum { VMMAP_CHECK_NONE, VMMAP_CHECK_UNLOCK, VMMAP_CHECK_ALL };
#ifdef DIAGNOSTIC
static int enable_vmmap_check = VMMAP_CHECK_UNLOCK;
#else
static int enable_vmmap_check = VMMAP_CHECK_NONE;
#endif
SYSCTL_INT(_debug, OID_AUTO, vmmap_check, CTLFLAG_RWTUN,
&enable_vmmap_check, 0, "Enable vm map consistency checking");
static void _vm_map_assert_consistent(vm_map_t map, int check);
#define VM_MAP_ASSERT_CONSISTENT(map) \
_vm_map_assert_consistent(map, VMMAP_CHECK_ALL)
#ifdef DIAGNOSTIC
#define VM_MAP_UNLOCK_CONSISTENT(map) do { \
if (map->nupdates > map->nentries) { \
_vm_map_assert_consistent(map, VMMAP_CHECK_UNLOCK); \
map->nupdates = 0; \
} \
} while (0)
#else
#define VM_MAP_UNLOCK_CONSISTENT(map)
#endif
#else
#define VM_MAP_ASSERT_LOCKED(map)
#define VM_MAP_ASSERT_CONSISTENT(map)
#define VM_MAP_UNLOCK_CONSISTENT(map)
#endif /* INVARIANTS */
void
_vm_map_unlock(vm_map_t map, const char *file, int line)
{
VM_MAP_UNLOCK_CONSISTENT(map);
if (vm_map_is_system(map)) {
#ifndef UMA_USE_DMAP
if (map == kernel_map && (map->flags & MAP_REPLENISH) != 0) {
uma_prealloc(kmapentzone, 1);
map->flags &= ~MAP_REPLENISH;
}
#endif
mtx_unlock_flags_(&map->system_mtx, 0, file, line);
} else {
sx_xunlock_(&map->lock, file, line);
vm_map_process_deferred();
}
}
void
_vm_map_lock_read(vm_map_t map, const char *file, int line)
{
if (vm_map_is_system(map))
mtx_lock_flags_(&map->system_mtx, 0, file, line);
else
sx_slock_(&map->lock, file, line);
}
void
_vm_map_unlock_read(vm_map_t map, const char *file, int line)
{
if (vm_map_is_system(map)) {
KASSERT((map->flags & MAP_REPLENISH) == 0,
("%s: MAP_REPLENISH leaked", __func__));
mtx_unlock_flags_(&map->system_mtx, 0, file, line);
} else {
sx_sunlock_(&map->lock, file, line);
vm_map_process_deferred();
}
}
int
_vm_map_trylock(vm_map_t map, const char *file, int line)
{
int error;
error = vm_map_is_system(map) ?
!mtx_trylock_flags_(&map->system_mtx, 0, file, line) :
!sx_try_xlock_(&map->lock, file, line);
if (error == 0)
map->timestamp++;
return (error == 0);
}
int
_vm_map_trylock_read(vm_map_t map, const char *file, int line)
{
int error;
error = vm_map_is_system(map) ?
!mtx_trylock_flags_(&map->system_mtx, 0, file, line) :
!sx_try_slock_(&map->lock, file, line);
return (error == 0);
}
/*
* _vm_map_lock_upgrade: [ internal use only ]
*
* Tries to upgrade a read (shared) lock on the specified map to a write
* (exclusive) lock. Returns the value "0" if the upgrade succeeds and a
* non-zero value if the upgrade fails. If the upgrade fails, the map is
* returned without a read or write lock held.
*
* Requires that the map be read locked.
*/
int
_vm_map_lock_upgrade(vm_map_t map, const char *file, int line)
{
unsigned int last_timestamp;
if (vm_map_is_system(map)) {
mtx_assert_(&map->system_mtx, MA_OWNED, file, line);
} else {
if (!sx_try_upgrade_(&map->lock, file, line)) {
last_timestamp = map->timestamp;
sx_sunlock_(&map->lock, file, line);
vm_map_process_deferred();
/*
* If the map's timestamp does not change while the
* map is unlocked, then the upgrade succeeds.
*/
sx_xlock_(&map->lock, file, line);
if (last_timestamp != map->timestamp) {
sx_xunlock_(&map->lock, file, line);
return (1);
}
}
}
map->timestamp++;
return (0);
}
void
_vm_map_lock_downgrade(vm_map_t map, const char *file, int line)
{
if (vm_map_is_system(map)) {
KASSERT((map->flags & MAP_REPLENISH) == 0,
("%s: MAP_REPLENISH leaked", __func__));
mtx_assert_(&map->system_mtx, MA_OWNED, file, line);
} else {
VM_MAP_UNLOCK_CONSISTENT(map);
sx_downgrade_(&map->lock, file, line);
}
}
/*
* vm_map_locked:
*
* Returns a non-zero value if the caller holds a write (exclusive) lock
* on the specified map and the value "0" otherwise.
*/
int
vm_map_locked(vm_map_t map)
{
if (vm_map_is_system(map))
return (mtx_owned(&map->system_mtx));
return (sx_xlocked(&map->lock));
}
/*
* _vm_map_unlock_and_wait:
*
* Atomically releases the lock on the specified map and puts the calling
* thread to sleep. The calling thread will remain asleep until either
* vm_map_wakeup() is performed on the map or the specified timeout is
* exceeded.
*
* WARNING! This function does not perform deferred deallocations of
* objects and map entries. Therefore, the calling thread is expected to
* reacquire the map lock after reawakening and later perform an ordinary
* unlock operation, such as vm_map_unlock(), before completing its
* operation on the map.
*/
int
_vm_map_unlock_and_wait(vm_map_t map, int timo, const char *file, int line)
{
VM_MAP_UNLOCK_CONSISTENT(map);
mtx_lock(&map_sleep_mtx);
if (vm_map_is_system(map)) {
KASSERT((map->flags & MAP_REPLENISH) == 0,
("%s: MAP_REPLENISH leaked", __func__));
mtx_unlock_flags_(&map->system_mtx, 0, file, line);
} else {
sx_xunlock_(&map->lock, file, line);
}
return (msleep(&map->root, &map_sleep_mtx, PDROP | PVM, "vmmaps",
timo));
}
/*
* vm_map_wakeup:
*
* Awaken any threads that have slept on the map using
* vm_map_unlock_and_wait().
*/
void
vm_map_wakeup(vm_map_t map)
{
/*
* Acquire and release map_sleep_mtx to prevent a wakeup()
* from being performed (and lost) between the map unlock
* and the msleep() in _vm_map_unlock_and_wait().
*/
mtx_lock(&map_sleep_mtx);
mtx_unlock(&map_sleep_mtx);
wakeup(&map->root);
}
void
vm_map_busy(vm_map_t map)
{
VM_MAP_ASSERT_LOCKED(map);
map->busy++;
}
void
vm_map_unbusy(vm_map_t map)
{
VM_MAP_ASSERT_LOCKED(map);
KASSERT(map->busy, ("vm_map_unbusy: not busy"));
if (--map->busy == 0 && (map->flags & MAP_BUSY_WAKEUP)) {
vm_map_modflags(map, 0, MAP_BUSY_WAKEUP);
wakeup(&map->busy);
}
}
void
vm_map_wait_busy(vm_map_t map)
{
VM_MAP_ASSERT_LOCKED(map);
while (map->busy) {
vm_map_modflags(map, MAP_BUSY_WAKEUP, 0);
if (vm_map_is_system(map))
msleep(&map->busy, &map->system_mtx, 0, "mbusy", 0);
else
sx_sleep(&map->busy, &map->lock, 0, "mbusy", 0);
}
map->timestamp++;
}
long
vmspace_resident_count(struct vmspace *vmspace)
{
return pmap_resident_count(vmspace_pmap(vmspace));
}
/*
* Initialize an existing vm_map structure
* such as that in the vmspace structure.
*/
static void
_vm_map_init(vm_map_t map, pmap_t pmap, vm_offset_t min, vm_offset_t max)
{
map->header.eflags = MAP_ENTRY_HEADER;
map->pmap = pmap;
map->header.end = min;
map->header.start = max;
map->flags = 0;
map->header.left = map->header.right = &map->header;
map->root = NULL;
map->timestamp = 0;
map->busy = 0;
map->anon_loc = 0;
#ifdef DIAGNOSTIC
map->nupdates = 0;
#endif
}
void
vm_map_init(vm_map_t map, pmap_t pmap, vm_offset_t min, vm_offset_t max)
{
_vm_map_init(map, pmap, min, max);
sx_init(&map->lock, "vm map (user)");
}
void
vm_map_init_system(vm_map_t map, pmap_t pmap, vm_offset_t min, vm_offset_t max)
{
_vm_map_init(map, pmap, min, max);
vm_map_modflags(map, MAP_SYSTEM_MAP, 0);
mtx_init(&map->system_mtx, "vm map (system)", NULL, MTX_DEF |
MTX_DUPOK);
}
/*
* vm_map_entry_dispose: [ internal use only ]
*
* Inverse of vm_map_entry_create.
*/
static void
vm_map_entry_dispose(vm_map_t map, vm_map_entry_t entry)
{
uma_zfree(vm_map_is_system(map) ? kmapentzone : mapentzone, entry);
}
/*
* vm_map_entry_create: [ internal use only ]
*
* Allocates a VM map entry for insertion.
* No entry fields are filled in.
*/
static vm_map_entry_t
vm_map_entry_create(vm_map_t map)
{
vm_map_entry_t new_entry;
#ifndef UMA_USE_DMAP
if (map == kernel_map) {
VM_MAP_ASSERT_LOCKED(map);
/*
* A new slab of kernel map entries cannot be allocated at this
* point because the kernel map has not yet been updated to
* reflect the caller's request. Therefore, we allocate a new
* map entry, dipping into the reserve if necessary, and set a
* flag indicating that the reserve must be replenished before
* the map is unlocked.
*/
new_entry = uma_zalloc(kmapentzone, M_NOWAIT | M_NOVM);
if (new_entry == NULL) {
new_entry = uma_zalloc(kmapentzone,
M_NOWAIT | M_NOVM | M_USE_RESERVE);
kernel_map->flags |= MAP_REPLENISH;
}
} else
#endif
if (vm_map_is_system(map)) {
new_entry = uma_zalloc(kmapentzone, M_NOWAIT);
} else {
new_entry = uma_zalloc(mapentzone, M_WAITOK);
}
KASSERT(new_entry != NULL,
("vm_map_entry_create: kernel resources exhausted"));
return (new_entry);
}
/*
* vm_map_entry_set_behavior:
*
* Set the expected access behavior, either normal, random, or
* sequential.
*/
static inline void
vm_map_entry_set_behavior(vm_map_entry_t entry, u_char behavior)
{
entry->eflags = (entry->eflags & ~MAP_ENTRY_BEHAV_MASK) |
(behavior & MAP_ENTRY_BEHAV_MASK);
}
/*
* vm_map_entry_max_free_{left,right}:
*
* Compute the size of the largest free gap between two entries,
* one the root of a tree and the other the ancestor of that root
* that is the least or greatest ancestor found on the search path.
*/
static inline vm_size_t
vm_map_entry_max_free_left(vm_map_entry_t root, vm_map_entry_t left_ancestor)
{
return (root->left != left_ancestor ?
root->left->max_free : root->start - left_ancestor->end);
}
static inline vm_size_t
vm_map_entry_max_free_right(vm_map_entry_t root, vm_map_entry_t right_ancestor)
{
return (root->right != right_ancestor ?
root->right->max_free : right_ancestor->start - root->end);
}
/*
* vm_map_entry_{pred,succ}:
*
* Find the {predecessor, successor} of the entry by taking one step
* in the appropriate direction and backtracking as much as necessary.
* vm_map_entry_succ is defined in vm_map.h.
*/
static inline vm_map_entry_t
vm_map_entry_pred(vm_map_entry_t entry)
{
vm_map_entry_t prior;
prior = entry->left;
if (prior->right->start < entry->start) {
do
prior = prior->right;
while (prior->right != entry);
}
return (prior);
}
static inline vm_size_t
vm_size_max(vm_size_t a, vm_size_t b)
{
return (a > b ? a : b);
}
#define SPLAY_LEFT_STEP(root, y, llist, rlist, test) do { \
vm_map_entry_t z; \
vm_size_t max_free; \
\
/* \
* Infer root->right->max_free == root->max_free when \
* y->max_free < root->max_free || root->max_free == 0. \
* Otherwise, look right to find it. \
*/ \
y = root->left; \
max_free = root->max_free; \
KASSERT(max_free == vm_size_max( \
vm_map_entry_max_free_left(root, llist), \
vm_map_entry_max_free_right(root, rlist)), \
("%s: max_free invariant fails", __func__)); \
if (max_free - 1 < vm_map_entry_max_free_left(root, llist)) \
max_free = vm_map_entry_max_free_right(root, rlist); \
if (y != llist && (test)) { \
/* Rotate right and make y root. */ \
z = y->right; \
if (z != root) { \
root->left = z; \
y->right = root; \
if (max_free < y->max_free) \
root->max_free = max_free = \
vm_size_max(max_free, z->max_free); \
} else if (max_free < y->max_free) \
root->max_free = max_free = \
vm_size_max(max_free, root->start - y->end);\
root = y; \
y = root->left; \
} \
/* Copy right->max_free. Put root on rlist. */ \
root->max_free = max_free; \
KASSERT(max_free == vm_map_entry_max_free_right(root, rlist), \
("%s: max_free not copied from right", __func__)); \
root->left = rlist; \
rlist = root; \
root = y != llist ? y : NULL; \
} while (0)
#define SPLAY_RIGHT_STEP(root, y, llist, rlist, test) do { \
vm_map_entry_t z; \
vm_size_t max_free; \
\
/* \
* Infer root->left->max_free == root->max_free when \
* y->max_free < root->max_free || root->max_free == 0. \
* Otherwise, look left to find it. \
*/ \
y = root->right; \
max_free = root->max_free; \
KASSERT(max_free == vm_size_max( \
vm_map_entry_max_free_left(root, llist), \
vm_map_entry_max_free_right(root, rlist)), \
("%s: max_free invariant fails", __func__)); \
if (max_free - 1 < vm_map_entry_max_free_right(root, rlist)) \
max_free = vm_map_entry_max_free_left(root, llist); \
if (y != rlist && (test)) { \
/* Rotate left and make y root. */ \
z = y->left; \
if (z != root) { \
root->right = z; \
y->left = root; \
if (max_free < y->max_free) \
root->max_free = max_free = \
vm_size_max(max_free, z->max_free); \
} else if (max_free < y->max_free) \
root->max_free = max_free = \
vm_size_max(max_free, y->start - root->end);\
root = y; \
y = root->right; \
} \
/* Copy left->max_free. Put root on llist. */ \
root->max_free = max_free; \
KASSERT(max_free == vm_map_entry_max_free_left(root, llist), \
("%s: max_free not copied from left", __func__)); \
root->right = llist; \
llist = root; \
root = y != rlist ? y : NULL; \
} while (0)
/*
* Walk down the tree until we find addr or a gap where addr would go, breaking
* off left and right subtrees of nodes less than, or greater than addr. Treat
* subtrees with root->max_free < length as empty trees. llist and rlist are
* the two sides in reverse order (bottom-up), with llist linked by the right
* pointer and rlist linked by the left pointer in the vm_map_entry, and both
* lists terminated by &map->header. This function, and the subsequent call to
* vm_map_splay_merge_{left,right,pred,succ}, rely on the start and end address
* values in &map->header.
*/
static __always_inline vm_map_entry_t
vm_map_splay_split(vm_map_t map, vm_offset_t addr, vm_size_t length,
vm_map_entry_t *llist, vm_map_entry_t *rlist)
{
vm_map_entry_t left, right, root, y;
left = right = &map->header;
root = map->root;
while (root != NULL && root->max_free >= length) {
KASSERT(left->end <= root->start &&
root->end <= right->start,
("%s: root not within tree bounds", __func__));
if (addr < root->start) {
SPLAY_LEFT_STEP(root, y, left, right,
y->max_free >= length && addr < y->start);
} else if (addr >= root->end) {
SPLAY_RIGHT_STEP(root, y, left, right,
y->max_free >= length && addr >= y->end);
} else
break;
}
*llist = left;
*rlist = right;
return (root);
}
static __always_inline void
vm_map_splay_findnext(vm_map_entry_t root, vm_map_entry_t *rlist)
{
vm_map_entry_t hi, right, y;
right = *rlist;
hi = root->right == right ? NULL : root->right;
if (hi == NULL)
return;
do
SPLAY_LEFT_STEP(hi, y, root, right, true);
while (hi != NULL);
*rlist = right;
}
static __always_inline void
vm_map_splay_findprev(vm_map_entry_t root, vm_map_entry_t *llist)
{
vm_map_entry_t left, lo, y;
left = *llist;
lo = root->left == left ? NULL : root->left;
if (lo == NULL)
return;
do
SPLAY_RIGHT_STEP(lo, y, left, root, true);
while (lo != NULL);
*llist = left;
}
static inline void
vm_map_entry_swap(vm_map_entry_t *a, vm_map_entry_t *b)
{
vm_map_entry_t tmp;
tmp = *b;
*b = *a;
*a = tmp;
}
/*
* Walk back up the two spines, flip the pointers and set max_free. The
* subtrees of the root go at the bottom of llist and rlist.
*/
static vm_size_t
vm_map_splay_merge_left_walk(vm_map_entry_t header, vm_map_entry_t root,
vm_map_entry_t tail, vm_size_t max_free, vm_map_entry_t llist)
{
do {
/*
* The max_free values of the children of llist are in
* llist->max_free and max_free. Update with the
* max value.
*/
llist->max_free = max_free =
vm_size_max(llist->max_free, max_free);
vm_map_entry_swap(&llist->right, &tail);
vm_map_entry_swap(&tail, &llist);
} while (llist != header);
root->left = tail;
return (max_free);
}
/*
* When llist is known to be the predecessor of root.
*/
static inline vm_size_t
vm_map_splay_merge_pred(vm_map_entry_t header, vm_map_entry_t root,
vm_map_entry_t llist)
{
vm_size_t max_free;
max_free = root->start - llist->end;
if (llist != header) {
max_free = vm_map_splay_merge_left_walk(header, root,
root, max_free, llist);
} else {
root->left = header;
header->right = root;
}
return (max_free);
}
/*
* When llist may or may not be the predecessor of root.
*/
static inline vm_size_t
vm_map_splay_merge_left(vm_map_entry_t header, vm_map_entry_t root,
vm_map_entry_t llist)
{
vm_size_t max_free;
max_free = vm_map_entry_max_free_left(root, llist);
if (llist != header) {
max_free = vm_map_splay_merge_left_walk(header, root,
root->left == llist ? root : root->left,
max_free, llist);
}
return (max_free);
}
static vm_size_t
vm_map_splay_merge_right_walk(vm_map_entry_t header, vm_map_entry_t root,
vm_map_entry_t tail, vm_size_t max_free, vm_map_entry_t rlist)
{
do {
/*
* The max_free values of the children of rlist are in
* rlist->max_free and max_free. Update with the
* max value.
*/
rlist->max_free = max_free =
vm_size_max(rlist->max_free, max_free);
vm_map_entry_swap(&rlist->left, &tail);
vm_map_entry_swap(&tail, &rlist);
} while (rlist != header);
root->right = tail;
return (max_free);
}
/*
* When rlist is known to be the succecessor of root.
*/
static inline vm_size_t
vm_map_splay_merge_succ(vm_map_entry_t header, vm_map_entry_t root,
vm_map_entry_t rlist)
{
vm_size_t max_free;
max_free = rlist->start - root->end;
if (rlist != header) {
max_free = vm_map_splay_merge_right_walk(header, root,
root, max_free, rlist);
} else {
root->right = header;
header->left = root;
}
return (max_free);
}
/*
* When rlist may or may not be the succecessor of root.
*/
static inline vm_size_t
vm_map_splay_merge_right(vm_map_entry_t header, vm_map_entry_t root,
vm_map_entry_t rlist)
{
vm_size_t max_free;
max_free = vm_map_entry_max_free_right(root, rlist);
if (rlist != header) {
max_free = vm_map_splay_merge_right_walk(header, root,
root->right == rlist ? root : root->right,
max_free, rlist);
}
return (max_free);
}
/*
* vm_map_splay:
*
* The Sleator and Tarjan top-down splay algorithm with the
* following variation. Max_free must be computed bottom-up, so
* on the downward pass, maintain the left and right spines in
* reverse order. Then, make a second pass up each side to fix
* the pointers and compute max_free. The time bound is O(log n)
* amortized.
*
* The tree is threaded, which means that there are no null pointers.
* When a node has no left child, its left pointer points to its
* predecessor, which the last ancestor on the search path from the root
* where the search branched right. Likewise, when a node has no right
* child, its right pointer points to its successor. The map header node
* is the predecessor of the first map entry, and the successor of the
* last.
*
* The new root is the vm_map_entry containing "addr", or else an
* adjacent entry (lower if possible) if addr is not in the tree.
*
* The map must be locked, and leaves it so.
*
* Returns: the new root.
*/
static vm_map_entry_t
vm_map_splay(vm_map_t map, vm_offset_t addr)
{
vm_map_entry_t header, llist, rlist, root;
vm_size_t max_free_left, max_free_right;
header = &map->header;
root = vm_map_splay_split(map, addr, 0, &llist, &rlist);
if (root != NULL) {
max_free_left = vm_map_splay_merge_left(header, root, llist);
max_free_right = vm_map_splay_merge_right(header, root, rlist);
} else if (llist != header) {
/*
* Recover the greatest node in the left
* subtree and make it the root.
*/
root = llist;
llist = root->right;
max_free_left = vm_map_splay_merge_left(header, root, llist);
max_free_right = vm_map_splay_merge_succ(header, root, rlist);
} else if (rlist != header) {
/*
* Recover the least node in the right
* subtree and make it the root.
*/
root = rlist;
rlist = root->left;
max_free_left = vm_map_splay_merge_pred(header, root, llist);
max_free_right = vm_map_splay_merge_right(header, root, rlist);
} else {
/* There is no root. */
return (NULL);
}
root->max_free = vm_size_max(max_free_left, max_free_right);
map->root = root;
VM_MAP_ASSERT_CONSISTENT(map);
return (root);
}
/*
* vm_map_entry_{un,}link:
*
* Insert/remove entries from maps. On linking, if new entry clips
* existing entry, trim existing entry to avoid overlap, and manage
* offsets. On unlinking, merge disappearing entry with neighbor, if
* called for, and manage offsets. Callers should not modify fields in
* entries already mapped.
*/
static void
vm_map_entry_link(vm_map_t map, vm_map_entry_t entry)
{
vm_map_entry_t header, llist, rlist, root;
vm_size_t max_free_left, max_free_right;
CTR3(KTR_VM,
"vm_map_entry_link: map %p, nentries %d, entry %p", map,
map->nentries, entry);
VM_MAP_ASSERT_LOCKED(map);
map->nentries++;
header = &map->header;
root = vm_map_splay_split(map, entry->start, 0, &llist, &rlist);
if (root == NULL) {
/*
* The new entry does not overlap any existing entry in the
* map, so it becomes the new root of the map tree.
*/
max_free_left = vm_map_splay_merge_pred(header, entry, llist);
max_free_right = vm_map_splay_merge_succ(header, entry, rlist);
} else if (entry->start == root->start) {
/*
* The new entry is a clone of root, with only the end field
* changed. The root entry will be shrunk to abut the new
* entry, and will be the right child of the new root entry in
* the modified map.
*/
KASSERT(entry->end < root->end,
("%s: clip_start not within entry", __func__));
vm_map_splay_findprev(root, &llist);
if ((root->eflags & MAP_ENTRY_STACK_GAP) == 0)
root->offset += entry->end - root->start;
root->start = entry->end;
max_free_left = vm_map_splay_merge_pred(header, entry, llist);
max_free_right = root->max_free = vm_size_max(
vm_map_splay_merge_pred(entry, root, entry),
vm_map_splay_merge_right(header, root, rlist));
} else {
/*
* The new entry is a clone of root, with only the start field
* changed. The root entry will be shrunk to abut the new
* entry, and will be the left child of the new root entry in
* the modified map.
*/
KASSERT(entry->end == root->end,
("%s: clip_start not within entry", __func__));
vm_map_splay_findnext(root, &rlist);
if ((entry->eflags & MAP_ENTRY_STACK_GAP) == 0)
entry->offset += entry->start - root->start;
root->end = entry->start;
max_free_left = root->max_free = vm_size_max(
vm_map_splay_merge_left(header, root, llist),
vm_map_splay_merge_succ(entry, root, entry));
max_free_right = vm_map_splay_merge_succ(header, entry, rlist);
}
entry->max_free = vm_size_max(max_free_left, max_free_right);
map->root = entry;
VM_MAP_ASSERT_CONSISTENT(map);
}
enum unlink_merge_type {
UNLINK_MERGE_NONE,
UNLINK_MERGE_NEXT
};
static void
vm_map_entry_unlink(vm_map_t map, vm_map_entry_t entry,
enum unlink_merge_type op)
{
vm_map_entry_t header, llist, rlist, root;
vm_size_t max_free_left, max_free_right;
VM_MAP_ASSERT_LOCKED(map);
header = &map->header;
root = vm_map_splay_split(map, entry->start, 0, &llist, &rlist);
KASSERT(root != NULL,
("vm_map_entry_unlink: unlink object not mapped"));
vm_map_splay_findprev(root, &llist);
vm_map_splay_findnext(root, &rlist);
if (op == UNLINK_MERGE_NEXT) {
rlist->start = root->start;
MPASS((rlist->eflags & MAP_ENTRY_STACK_GAP) == 0);
rlist->offset = root->offset;
}
if (llist != header) {
root = llist;
llist = root->right;
max_free_left = vm_map_splay_merge_left(header, root, llist);
max_free_right = vm_map_splay_merge_succ(header, root, rlist);
} else if (rlist != header) {
root = rlist;
rlist = root->left;
max_free_left = vm_map_splay_merge_pred(header, root, llist);
max_free_right = vm_map_splay_merge_right(header, root, rlist);
} else {
header->left = header->right = header;
root = NULL;
}
if (root != NULL)
root->max_free = vm_size_max(max_free_left, max_free_right);
map->root = root;
VM_MAP_ASSERT_CONSISTENT(map);
map->nentries--;
CTR3(KTR_VM, "vm_map_entry_unlink: map %p, nentries %d, entry %p", map,
map->nentries, entry);
}
/*
* vm_map_entry_resize:
*
* Resize a vm_map_entry, recompute the amount of free space that
* follows it and propagate that value up the tree.
*
* The map must be locked, and leaves it so.
*/
static void
vm_map_entry_resize(vm_map_t map, vm_map_entry_t entry, vm_size_t grow_amount)
{
vm_map_entry_t header, llist, rlist, root;
VM_MAP_ASSERT_LOCKED(map);
header = &map->header;
root = vm_map_splay_split(map, entry->start, 0, &llist, &rlist);
KASSERT(root != NULL, ("%s: resize object not mapped", __func__));
vm_map_splay_findnext(root, &rlist);
entry->end += grow_amount;
root->max_free = vm_size_max(
vm_map_splay_merge_left(header, root, llist),
vm_map_splay_merge_succ(header, root, rlist));
map->root = root;
VM_MAP_ASSERT_CONSISTENT(map);
CTR4(KTR_VM, "%s: map %p, nentries %d, entry %p",
__func__, map, map->nentries, entry);
}
/*
* vm_map_lookup_entry: [ internal use only ]
*
* Finds the map entry containing (or
* immediately preceding) the specified address
* in the given map; the entry is returned
* in the "entry" parameter. The boolean
* result indicates whether the address is
* actually contained in the map.
*/
boolean_t
vm_map_lookup_entry(
vm_map_t map,
vm_offset_t address,
vm_map_entry_t *entry) /* OUT */
{
vm_map_entry_t cur, header, lbound, ubound;
boolean_t locked;
/*
* If the map is empty, then the map entry immediately preceding
* "address" is the map's header.
*/
header = &map->header;
cur = map->root;
if (cur == NULL) {
*entry = header;
return (FALSE);
}
if (address >= cur->start && cur->end > address) {
*entry = cur;
return (TRUE);
}
if ((locked = vm_map_locked(map)) ||
sx_try_upgrade(&map->lock)) {
/*
* Splay requires a write lock on the map. However, it only
* restructures the binary search tree; it does not otherwise
* change the map. Thus, the map's timestamp need not change
* on a temporary upgrade.
*/
cur = vm_map_splay(map, address);
if (!locked) {
VM_MAP_UNLOCK_CONSISTENT(map);
sx_downgrade(&map->lock);
}
/*
* If "address" is contained within a map entry, the new root
* is that map entry. Otherwise, the new root is a map entry
* immediately before or after "address".
*/
if (address < cur->start) {
*entry = header;
return (FALSE);
}
*entry = cur;
return (address < cur->end);
}
/*
* Since the map is only locked for read access, perform a
* standard binary search tree lookup for "address".
*/
lbound = ubound = header;
for (;;) {
if (address < cur->start) {
ubound = cur;
cur = cur->left;
if (cur == lbound)
break;
} else if (cur->end <= address) {
lbound = cur;
cur = cur->right;
if (cur == ubound)
break;
} else {
*entry = cur;
return (TRUE);
}
}
*entry = lbound;
return (FALSE);
}
/*
* vm_map_insert1() is identical to vm_map_insert() except that it
* returns the newly inserted map entry in '*res'. In case the new
* entry is coalesced with a neighbor or an existing entry was
* resized, that entry is returned. In any case, the returned entry
* covers the specified address range.
*/
static int
vm_map_insert1(vm_map_t map, vm_object_t object, vm_ooffset_t offset,
vm_offset_t start, vm_offset_t end, vm_prot_t prot, vm_prot_t max, int cow,
vm_map_entry_t *res)
{
vm_map_entry_t new_entry, next_entry, prev_entry;
struct ucred *cred;
vm_eflags_t protoeflags;
vm_inherit_t inheritance;
u_long bdry;
u_int bidx;
int cflags;
VM_MAP_ASSERT_LOCKED(map);
KASSERT(object != kernel_object ||
(cow & MAP_COPY_ON_WRITE) == 0,
("vm_map_insert: kernel object and COW"));
KASSERT(object == NULL || (cow & MAP_NOFAULT) == 0 ||
(cow & MAP_SPLIT_BOUNDARY_MASK) != 0,
("vm_map_insert: paradoxical MAP_NOFAULT request, obj %p cow %#x",
object, cow));
KASSERT((prot & ~max) == 0,
("prot %#x is not subset of max_prot %#x", prot, max));
/*
* Check that the start and end points are not bogus.
*/
if (start == end || !vm_map_range_valid(map, start, end))
return (KERN_INVALID_ADDRESS);
if ((map->flags & MAP_WXORX) != 0 && (prot & (VM_PROT_WRITE |
VM_PROT_EXECUTE)) == (VM_PROT_WRITE | VM_PROT_EXECUTE))
return (KERN_PROTECTION_FAILURE);
/*
* Find the entry prior to the proposed starting address; if it's part
* of an existing entry, this range is bogus.
*/
if (vm_map_lookup_entry(map, start, &prev_entry))
return (KERN_NO_SPACE);
/*
* Assert that the next entry doesn't overlap the end point.
*/
next_entry = vm_map_entry_succ(prev_entry);
if (next_entry->start < end)
return (KERN_NO_SPACE);
if ((cow & MAP_CREATE_GUARD) != 0 && (object != NULL ||
max != VM_PROT_NONE))
return (KERN_INVALID_ARGUMENT);
protoeflags = 0;
if (cow & MAP_COPY_ON_WRITE)
protoeflags |= MAP_ENTRY_COW | MAP_ENTRY_NEEDS_COPY;
if (cow & MAP_NOFAULT)
protoeflags |= MAP_ENTRY_NOFAULT;
if (cow & MAP_DISABLE_SYNCER)
protoeflags |= MAP_ENTRY_NOSYNC;
if (cow & MAP_DISABLE_COREDUMP)
protoeflags |= MAP_ENTRY_NOCOREDUMP;
if (cow & MAP_STACK_AREA)
protoeflags |= MAP_ENTRY_GROWS_DOWN;
if (cow & MAP_WRITECOUNT)
protoeflags |= MAP_ENTRY_WRITECNT;
if (cow & MAP_VN_EXEC)
protoeflags |= MAP_ENTRY_VN_EXEC;
if ((cow & MAP_CREATE_GUARD) != 0)
protoeflags |= MAP_ENTRY_GUARD;
if ((cow & MAP_CREATE_STACK_GAP) != 0)
protoeflags |= MAP_ENTRY_STACK_GAP;
if (cow & MAP_INHERIT_SHARE)
inheritance = VM_INHERIT_SHARE;
else
inheritance = VM_INHERIT_DEFAULT;
if ((cow & MAP_SPLIT_BOUNDARY_MASK) != 0) {
/* This magically ignores index 0, for usual page size. */
bidx = (cow & MAP_SPLIT_BOUNDARY_MASK) >>
MAP_SPLIT_BOUNDARY_SHIFT;
if (bidx >= MAXPAGESIZES)
return (KERN_INVALID_ARGUMENT);
bdry = pagesizes[bidx] - 1;
if ((start & bdry) != 0 || (end & bdry) != 0)
return (KERN_INVALID_ARGUMENT);
protoeflags |= bidx << MAP_ENTRY_SPLIT_BOUNDARY_SHIFT;
}
cred = NULL;
if ((cow & (MAP_ACC_NO_CHARGE | MAP_NOFAULT | MAP_CREATE_GUARD)) != 0) {
cflags = OBJCO_NO_CHARGE;
} else {
cflags = 0;
if ((cow & MAP_ACC_CHARGED) != 0 ||
((prot & VM_PROT_WRITE) != 0 &&
((protoeflags & MAP_ENTRY_NEEDS_COPY) != 0 ||
object == NULL))) {
if ((cow & MAP_ACC_CHARGED) == 0) {
if (!swap_reserve(end - start))
return (KERN_RESOURCE_SHORTAGE);
/*
* Only inform vm_object_coalesce()
* that the object was charged if
* there is no need for CoW, so the
* swap amount reserved is applicable
* to the prev_entry->object.
*/
if ((protoeflags & MAP_ENTRY_NEEDS_COPY) == 0)
cflags |= OBJCO_CHARGED;
}
KASSERT(object == NULL ||
(protoeflags & MAP_ENTRY_NEEDS_COPY) != 0 ||
object->cred == NULL,
("overcommit: vm_map_insert o %p", object));
cred = curthread->td_ucred;
}
}
/* Expand the kernel pmap, if necessary. */
if (map == kernel_map && end > kernel_vm_end) {
int rv;
rv = pmap_growkernel(end);
if (rv != KERN_SUCCESS)
return (rv);
}
if (object != NULL) {
/*
* OBJ_ONEMAPPING must be cleared unless this mapping
* is trivially proven to be the only mapping for any
* of the object's pages. (Object granularity
* reference counting is insufficient to recognize
* aliases with precision.)
*/
if ((object->flags & OBJ_ANON) != 0) {
VM_OBJECT_WLOCK(object);
if (object->ref_count > 1 || object->shadow_count != 0)
vm_object_clear_flag(object, OBJ_ONEMAPPING);
VM_OBJECT_WUNLOCK(object);
}
} else if ((prev_entry->eflags & ~MAP_ENTRY_USER_WIRED) ==
protoeflags &&
(cow & (MAP_STACK_AREA | MAP_VN_EXEC)) == 0 &&
prev_entry->end == start && (prev_entry->cred == cred ||
(prev_entry->object.vm_object != NULL &&
prev_entry->object.vm_object->cred == cred)) &&
vm_object_coalesce(prev_entry->object.vm_object,
prev_entry->offset,
(vm_size_t)(prev_entry->end - prev_entry->start),
(vm_size_t)(end - prev_entry->end), cflags)) {
/*
* We were able to extend the object. Determine if we
* can extend the previous map entry to include the
* new range as well.
*/
if (prev_entry->inheritance == inheritance &&
prev_entry->protection == prot &&
prev_entry->max_protection == max &&
prev_entry->wired_count == 0) {
KASSERT((prev_entry->eflags & MAP_ENTRY_USER_WIRED) ==
0, ("prev_entry %p has incoherent wiring",
prev_entry));
if ((prev_entry->eflags & MAP_ENTRY_GUARD) == 0)
map->size += end - prev_entry->end;
vm_map_entry_resize(map, prev_entry,
end - prev_entry->end);
*res = vm_map_try_merge_entries(map, prev_entry,
next_entry);
return (KERN_SUCCESS);
}
/*
* If we can extend the object but cannot extend the
* map entry, we have to create a new map entry. We
* must bump the ref count on the extended object to
* account for it. object may be NULL.
*/
object = prev_entry->object.vm_object;
offset = prev_entry->offset +
(prev_entry->end - prev_entry->start);
vm_object_reference(object);
if (cred != NULL && object != NULL && object->cred != NULL &&
!(prev_entry->eflags & MAP_ENTRY_NEEDS_COPY)) {
/* Object already accounts for this uid. */
cred = NULL;
}
}
if (cred != NULL)
crhold(cred);
/*
* Create a new entry
*/
new_entry = vm_map_entry_create(map);
new_entry->start = start;
new_entry->end = end;
new_entry->cred = NULL;
new_entry->eflags = protoeflags;
new_entry->object.vm_object = object;
new_entry->offset = offset;
new_entry->inheritance = inheritance;
new_entry->protection = prot;
new_entry->max_protection = max;
new_entry->wired_count = 0;
new_entry->wiring_thread = NULL;
new_entry->read_ahead = VM_FAULT_READ_AHEAD_INIT;
new_entry->next_read = start;
KASSERT(cred == NULL || !ENTRY_CHARGED(new_entry),
("overcommit: vm_map_insert leaks vm_map %p", new_entry));
new_entry->cred = cred;
/*
* Insert the new entry into the list
*/
vm_map_entry_link(map, new_entry);
if ((new_entry->eflags & MAP_ENTRY_GUARD) == 0)
map->size += new_entry->end - new_entry->start;
/*
* Try to coalesce the new entry with both the previous and next
* entries in the list. Previously, we only attempted to coalesce
* with the previous entry when object is NULL. Here, we handle the
* other cases, which are less common.
*/
vm_map_try_merge_entries(map, prev_entry, new_entry);
*res = vm_map_try_merge_entries(map, new_entry, next_entry);
if ((cow & (MAP_PREFAULT | MAP_PREFAULT_PARTIAL)) != 0) {
vm_map_pmap_enter(map, start, prot, object, OFF_TO_IDX(offset),
end - start, cow & MAP_PREFAULT_PARTIAL);
}
return (KERN_SUCCESS);
}
/*
* vm_map_insert:
*
* Inserts the given VM object into the target map at the
* specified address range.
*
* Requires that the map be locked, and leaves it so.
*
* If object is non-NULL, ref count must be bumped by caller
* prior to making call to account for the new entry.
*/
int
vm_map_insert(vm_map_t map, vm_object_t object, vm_ooffset_t offset,
vm_offset_t start, vm_offset_t end, vm_prot_t prot, vm_prot_t max, int cow)
{
vm_map_entry_t res;
return (vm_map_insert1(map, object, offset, start, end, prot, max,
cow, &res));
}
/*
* vm_map_findspace:
*
* Find the first fit (lowest VM address) for "length" free bytes
* beginning at address >= start in the given map.
*
* In a vm_map_entry, "max_free" is the maximum amount of
* contiguous free space between an entry in its subtree and a
* neighbor of that entry. This allows finding a free region in
* one path down the tree, so O(log n) amortized with splay
* trees.
*
* The map must be locked, and leaves it so.
*
* Returns: starting address if sufficient space,
* vm_map_max(map)-length+1 if insufficient space.
*/
vm_offset_t
vm_map_findspace(vm_map_t map, vm_offset_t start, vm_size_t length)
{
vm_map_entry_t header, llist, rlist, root, y;
vm_size_t left_length, max_free_left, max_free_right;
vm_offset_t gap_end;
VM_MAP_ASSERT_LOCKED(map);
/*
* Request must fit within min/max VM address and must avoid
* address wrap.
*/
start = MAX(start, vm_map_min(map));
if (start >= vm_map_max(map) || length > vm_map_max(map) - start)
return (vm_map_max(map) - length + 1);
/* Empty tree means wide open address space. */
if (map->root == NULL)
return (start);
/*
* After splay_split, if start is within an entry, push it to the start
* of the following gap. If rlist is at the end of the gap containing
* start, save the end of that gap in gap_end to see if the gap is big
* enough; otherwise set gap_end to start skip gap-checking and move
* directly to a search of the right subtree.
*/
header = &map->header;
root = vm_map_splay_split(map, start, length, &llist, &rlist);
gap_end = rlist->start;
if (root != NULL) {
start = root->end;
if (root->right != rlist)
gap_end = start;
max_free_left = vm_map_splay_merge_left(header, root, llist);
max_free_right = vm_map_splay_merge_right(header, root, rlist);
} else if (rlist != header) {
root = rlist;
rlist = root->left;
max_free_left = vm_map_splay_merge_pred(header, root, llist);
max_free_right = vm_map_splay_merge_right(header, root, rlist);
} else {
root = llist;
llist = root->right;
max_free_left = vm_map_splay_merge_left(header, root, llist);
max_free_right = vm_map_splay_merge_succ(header, root, rlist);
}
root->max_free = vm_size_max(max_free_left, max_free_right);
map->root = root;
VM_MAP_ASSERT_CONSISTENT(map);
if (length <= gap_end - start)
return (start);
/* With max_free, can immediately tell if no solution. */
if (root->right == header || length > root->right->max_free)
return (vm_map_max(map) - length + 1);
/*
* Splay for the least large-enough gap in the right subtree.
*/
llist = rlist = header;
for (left_length = 0;;
left_length = vm_map_entry_max_free_left(root, llist)) {
if (length <= left_length)
SPLAY_LEFT_STEP(root, y, llist, rlist,
length <= vm_map_entry_max_free_left(y, llist));
else
SPLAY_RIGHT_STEP(root, y, llist, rlist,
length > vm_map_entry_max_free_left(y, root));
if (root == NULL)
break;
}
root = llist;
llist = root->right;
max_free_left = vm_map_splay_merge_left(header, root, llist);
if (rlist == header) {
root->max_free = vm_size_max(max_free_left,
vm_map_splay_merge_succ(header, root, rlist));
} else {
y = rlist;
rlist = y->left;
y->max_free = vm_size_max(
vm_map_splay_merge_pred(root, y, root),
vm_map_splay_merge_right(header, y, rlist));
root->max_free = vm_size_max(max_free_left, y->max_free);
}
map->root = root;
VM_MAP_ASSERT_CONSISTENT(map);
return (root->end);
}
int
vm_map_fixed(vm_map_t map, vm_object_t object, vm_ooffset_t offset,
vm_offset_t start, vm_size_t length, vm_prot_t prot,
vm_prot_t max, int cow)
{
vm_offset_t end;
int result;
end = start + length;
KASSERT((cow & MAP_STACK_AREA) == 0 || object == NULL,
("vm_map_fixed: non-NULL backing object for stack"));
vm_map_lock(map);
VM_MAP_RANGE_CHECK(map, start, end);
if ((cow & MAP_CHECK_EXCL) == 0) {
result = vm_map_delete(map, start, end);
if (result != KERN_SUCCESS)
goto out;
}
if ((cow & MAP_STACK_AREA) != 0) {
result = vm_map_stack_locked(map, start, length, sgrowsiz,
prot, max, cow);
} else {
result = vm_map_insert(map, object, offset, start, end,
prot, max, cow);
}
out:
vm_map_unlock(map);
return (result);
}
#if VM_NRESERVLEVEL <= 1
static const int aslr_pages_rnd_64[2] = {0x1000, 0x10};
static const int aslr_pages_rnd_32[2] = {0x100, 0x4};
#elif VM_NRESERVLEVEL == 2
static const int aslr_pages_rnd_64[3] = {0x1000, 0x1000, 0x10};
static const int aslr_pages_rnd_32[3] = {0x100, 0x100, 0x4};
#else
#error "Unsupported VM_NRESERVLEVEL"
#endif
static int cluster_anon = 1;
SYSCTL_INT(_vm, OID_AUTO, cluster_anon, CTLFLAG_RW,
&cluster_anon, 0,
"Cluster anonymous mappings: 0 = no, 1 = yes if no hint, 2 = always");
static bool
clustering_anon_allowed(vm_offset_t addr, int cow)
{
switch (cluster_anon) {
case 0:
return (false);
case 1:
return (addr == 0 || (cow & MAP_NO_HINT) != 0);
case 2:
default:
return (true);
}
}
static long aslr_restarts;
SYSCTL_LONG(_vm, OID_AUTO, aslr_restarts, CTLFLAG_RD,
&aslr_restarts, 0,
"Number of aslr failures");
/*
* Searches for the specified amount of free space in the given map with the
* specified alignment. Performs an address-ordered, first-fit search from
* the given address "*addr", with an optional upper bound "max_addr". If the
* parameter "alignment" is zero, then the alignment is computed from the
* given (object, offset) pair so as to enable the greatest possible use of
* superpage mappings. Returns KERN_SUCCESS and the address of the free space
* in "*addr" if successful. Otherwise, returns KERN_NO_SPACE.
*
* The map must be locked. Initially, there must be at least "length" bytes
* of free space at the given address.
*/
static int
vm_map_alignspace(vm_map_t map, vm_object_t object, vm_ooffset_t offset,
vm_offset_t *addr, vm_size_t length, vm_offset_t max_addr,
vm_offset_t alignment)
{
vm_offset_t aligned_addr, free_addr;
VM_MAP_ASSERT_LOCKED(map);
free_addr = *addr;
KASSERT(free_addr == vm_map_findspace(map, free_addr, length),
("caller failed to provide space %#jx at address %p",
(uintmax_t)length, (void *)free_addr));
for (;;) {
/*
* At the start of every iteration, the free space at address
* "*addr" is at least "length" bytes.
*/
if (alignment == 0)
pmap_align_superpage(object, offset, addr, length);
else
*addr = roundup2(*addr, alignment);
aligned_addr = *addr;
if (aligned_addr == free_addr) {
/*
* Alignment did not change "*addr", so "*addr" must
* still provide sufficient free space.
*/
return (KERN_SUCCESS);
}
/*
* Test for address wrap on "*addr". A wrapped "*addr" could
* be a valid address, in which case vm_map_findspace() cannot
* be relied upon to fail.
*/
if (aligned_addr < free_addr)
return (KERN_NO_SPACE);
*addr = vm_map_findspace(map, aligned_addr, length);
if (*addr + length > vm_map_max(map) ||
(max_addr != 0 && *addr + length > max_addr))
return (KERN_NO_SPACE);
free_addr = *addr;
if (free_addr == aligned_addr) {
/*
* If a successful call to vm_map_findspace() did not
* change "*addr", then "*addr" must still be aligned
* and provide sufficient free space.
*/
return (KERN_SUCCESS);
}
}
}
int
vm_map_find_aligned(vm_map_t map, vm_offset_t *addr, vm_size_t length,
vm_offset_t max_addr, vm_offset_t alignment)
{
/* XXXKIB ASLR eh ? */
*addr = vm_map_findspace(map, *addr, length);
if (*addr + length > vm_map_max(map) ||
(max_addr != 0 && *addr + length > max_addr))
return (KERN_NO_SPACE);
return (vm_map_alignspace(map, NULL, 0, addr, length, max_addr,
alignment));
}
/*
* vm_map_find finds an unallocated region in the target address
* map with the given length. The search is defined to be
* first-fit from the specified address; the region found is
* returned in the same parameter.
*
* If object is non-NULL, ref count must be bumped by caller
* prior to making call to account for the new entry.
*/
int
vm_map_find(vm_map_t map, vm_object_t object, vm_ooffset_t offset,
vm_offset_t *addr, /* IN/OUT */
vm_size_t length, vm_offset_t max_addr, int find_space,
vm_prot_t prot, vm_prot_t max, int cow)
{
int rv;
vm_map_lock(map);
rv = vm_map_find_locked(map, object, offset, addr, length, max_addr,
find_space, prot, max, cow);
vm_map_unlock(map);
return (rv);
}
int
vm_map_find_locked(vm_map_t map, vm_object_t object, vm_ooffset_t offset,
vm_offset_t *addr, /* IN/OUT */
vm_size_t length, vm_offset_t max_addr, int find_space,
vm_prot_t prot, vm_prot_t max, int cow)
{
vm_offset_t alignment, curr_min_addr, min_addr;
int gap, pidx, rv, try;
bool cluster, en_aslr, update_anon;
KASSERT((cow & MAP_STACK_AREA) == 0 || object == NULL,
("non-NULL backing object for stack"));
MPASS((cow & MAP_REMAP) == 0 || (find_space == VMFS_NO_SPACE &&
(cow & MAP_STACK_AREA) == 0));
if (find_space == VMFS_OPTIMAL_SPACE && (object == NULL ||
(object->flags & OBJ_COLORED) == 0))
find_space = VMFS_ANY_SPACE;
if (find_space >> 8 != 0) {
KASSERT((find_space & 0xff) == 0, ("bad VMFS flags"));
alignment = (vm_offset_t)1 << (find_space >> 8);
} else
alignment = 0;
en_aslr = (map->flags & MAP_ASLR) != 0;
update_anon = cluster = clustering_anon_allowed(*addr, cow) &&
(map->flags & MAP_IS_SUB_MAP) == 0 && max_addr == 0 &&
find_space != VMFS_NO_SPACE && object == NULL &&
(cow & (MAP_INHERIT_SHARE | MAP_STACK_AREA)) == 0 &&
prot != PROT_NONE;
curr_min_addr = min_addr = *addr;
if (en_aslr && min_addr == 0 && !cluster &&
find_space != VMFS_NO_SPACE &&
(map->flags & MAP_ASLR_IGNSTART) != 0)
curr_min_addr = min_addr = vm_map_min(map);
try = 0;
if (cluster) {
curr_min_addr = map->anon_loc;
if (curr_min_addr == 0)
cluster = false;
}
if (find_space != VMFS_NO_SPACE) {
KASSERT(find_space == VMFS_ANY_SPACE ||
find_space == VMFS_OPTIMAL_SPACE ||
find_space == VMFS_SUPER_SPACE ||
alignment != 0, ("unexpected VMFS flag"));
again:
/*
* When creating an anonymous mapping, try clustering
* with an existing anonymous mapping first.
*
* We make up to two attempts to find address space
* for a given find_space value. The first attempt may
* apply randomization or may cluster with an existing
* anonymous mapping. If this first attempt fails,
* perform a first-fit search of the available address
* space.
*
* If all tries failed, and find_space is
* VMFS_OPTIMAL_SPACE, fallback to VMFS_ANY_SPACE.
* Again enable clustering and randomization.
*/
try++;
MPASS(try <= 2);
if (try == 2) {
/*
* Second try: we failed either to find a
* suitable region for randomizing the
* allocation, or to cluster with an existing
* mapping. Retry with free run.
*/
curr_min_addr = (map->flags & MAP_ASLR_IGNSTART) != 0 ?
vm_map_min(map) : min_addr;
atomic_add_long(&aslr_restarts, 1);
}
if (try == 1 && en_aslr && !cluster) {
/*
* Find space for allocation, including
* gap needed for later randomization.
*/
pidx = 0;
#if VM_NRESERVLEVEL > 0
if ((find_space == VMFS_SUPER_SPACE ||
find_space == VMFS_OPTIMAL_SPACE) &&
pagesizes[VM_NRESERVLEVEL] != 0) {
/*
* Do not pointlessly increase the space that
* is requested from vm_map_findspace().
* pmap_align_superpage() will only change a
* mapping's alignment if that mapping is at
* least a superpage in size.
*/
pidx = VM_NRESERVLEVEL;
while (pidx > 0 && length < pagesizes[pidx])
pidx--;
}
#endif
gap = vm_map_max(map) > MAP_32BIT_MAX_ADDR &&
(max_addr == 0 || max_addr > MAP_32BIT_MAX_ADDR) ?
aslr_pages_rnd_64[pidx] : aslr_pages_rnd_32[pidx];
*addr = vm_map_findspace(map, curr_min_addr,
length + gap * pagesizes[pidx]);
if (*addr + length + gap * pagesizes[pidx] >
vm_map_max(map))
goto again;
/* And randomize the start address. */
*addr += (arc4random() % gap) * pagesizes[pidx];
if (max_addr != 0 && *addr + length > max_addr)
goto again;
} else {
*addr = vm_map_findspace(map, curr_min_addr, length);
if (*addr + length > vm_map_max(map) ||
(max_addr != 0 && *addr + length > max_addr)) {
if (cluster) {
cluster = false;
MPASS(try == 1);
goto again;
}
return (KERN_NO_SPACE);
}
}
if (find_space != VMFS_ANY_SPACE &&
(rv = vm_map_alignspace(map, object, offset, addr, length,
max_addr, alignment)) != KERN_SUCCESS) {
if (find_space == VMFS_OPTIMAL_SPACE) {
find_space = VMFS_ANY_SPACE;
curr_min_addr = min_addr;
cluster = update_anon;
try = 0;
goto again;
}
return (rv);
}
} else if ((cow & MAP_REMAP) != 0) {
if (!vm_map_range_valid(map, *addr, *addr + length))
return (KERN_INVALID_ADDRESS);
rv = vm_map_delete(map, *addr, *addr + length);
if (rv != KERN_SUCCESS)
return (rv);
}
if ((cow & MAP_STACK_AREA) != 0) {
rv = vm_map_stack_locked(map, *addr, length, sgrowsiz, prot,
max, cow);
} else {
rv = vm_map_insert(map, object, offset, *addr, *addr + length,
prot, max, cow);
}
/*
* Update the starting address for clustered anonymous memory mappings
* if a starting address was not previously defined or an ASLR restart
* placed an anonymous memory mapping at a lower address.
*/
if (update_anon && rv == KERN_SUCCESS && (map->anon_loc == 0 ||
*addr < map->anon_loc))
map->anon_loc = *addr;
return (rv);
}
/*
* vm_map_find_min() is a variant of vm_map_find() that takes an
* additional parameter ("default_addr") and treats the given address
* ("*addr") differently. Specifically, it treats "*addr" as a hint
* and not as the minimum address where the mapping is created.
*
* This function works in two phases. First, it tries to
* allocate above the hint. If that fails and the hint is
* greater than "default_addr", it performs a second pass, replacing
* the hint with "default_addr" as the minimum address for the
* allocation.
*/
int
vm_map_find_min(vm_map_t map, vm_object_t object, vm_ooffset_t offset,
vm_offset_t *addr, vm_size_t length, vm_offset_t default_addr,
vm_offset_t max_addr, int find_space, vm_prot_t prot, vm_prot_t max,
int cow)
{
vm_offset_t hint;
int rv;
hint = *addr;
if (hint == 0) {
cow |= MAP_NO_HINT;
*addr = hint = default_addr;
}
for (;;) {
rv = vm_map_find(map, object, offset, addr, length, max_addr,
find_space, prot, max, cow);
if (rv == KERN_SUCCESS || default_addr >= hint)
return (rv);
*addr = hint = default_addr;
}
}
/*
* A map entry with any of the following flags set must not be merged with
* another entry.
*/
#define MAP_ENTRY_NOMERGE_MASK (MAP_ENTRY_GROWS_DOWN | \
MAP_ENTRY_IN_TRANSITION | MAP_ENTRY_IS_SUB_MAP | MAP_ENTRY_VN_EXEC | \
MAP_ENTRY_STACK_GAP)
static bool
vm_map_mergeable_neighbors(vm_map_entry_t prev, vm_map_entry_t entry)
{
KASSERT((prev->eflags & MAP_ENTRY_NOMERGE_MASK) == 0 ||
(entry->eflags & MAP_ENTRY_NOMERGE_MASK) == 0,
("vm_map_mergeable_neighbors: neither %p nor %p are mergeable",
prev, entry));
return (prev->end == entry->start &&
prev->object.vm_object == entry->object.vm_object &&
(prev->object.vm_object == NULL ||
prev->offset + (prev->end - prev->start) == entry->offset) &&
prev->eflags == entry->eflags &&
prev->protection == entry->protection &&
prev->max_protection == entry->max_protection &&
prev->inheritance == entry->inheritance &&
prev->wired_count == entry->wired_count &&
prev->cred == entry->cred);
}
static void
vm_map_merged_neighbor_dispose(vm_map_t map, vm_map_entry_t entry)
{
/*
* If the backing object is a vnode object, vm_object_deallocate()
* calls vrele(). However, vrele() does not lock the vnode because
* the vnode has additional references. Thus, the map lock can be
* kept without causing a lock-order reversal with the vnode lock.
*
* Since we count the number of virtual page mappings in
* object->un_pager.vnp.writemappings, the writemappings value
* should not be adjusted when the entry is disposed of.
*/
if (entry->object.vm_object != NULL)
vm_object_deallocate(entry->object.vm_object);
if (entry->cred != NULL)
crfree(entry->cred);
vm_map_entry_dispose(map, entry);
}
/*
* vm_map_try_merge_entries:
*
* Compare two map entries that represent consecutive ranges. If
* the entries can be merged, expand the range of the second to
* cover the range of the first and delete the first. Then return
* the map entry that includes the first range.
*
* The map must be locked.
*/
vm_map_entry_t
vm_map_try_merge_entries(vm_map_t map, vm_map_entry_t prev_entry,
vm_map_entry_t entry)
{
VM_MAP_ASSERT_LOCKED(map);
if ((entry->eflags & MAP_ENTRY_NOMERGE_MASK) == 0 &&
vm_map_mergeable_neighbors(prev_entry, entry)) {
vm_map_entry_unlink(map, prev_entry, UNLINK_MERGE_NEXT);
vm_map_merged_neighbor_dispose(map, prev_entry);
return (entry);
}
return (prev_entry);
}
/*
* vm_map_entry_back:
*
* Allocate an object to back a map entry.
*/
static inline void
vm_map_entry_back(vm_map_entry_t entry)
{
vm_object_t object;
KASSERT(entry->object.vm_object == NULL,
("map entry %p has backing object", entry));
KASSERT((entry->eflags & MAP_ENTRY_IS_SUB_MAP) == 0,
("map entry %p is a submap", entry));
object = vm_object_allocate_anon(atop(entry->end - entry->start), NULL,
entry->cred);
entry->object.vm_object = object;
entry->offset = 0;
entry->cred = NULL;
}
/*
* vm_map_entry_charge_object
*
* If there is no object backing this entry, create one. Otherwise, if
* the entry has cred, give it to the backing object.
*/
static inline void
vm_map_entry_charge_object(vm_map_t map, vm_map_entry_t entry)
{
vm_object_t object;
VM_MAP_ASSERT_LOCKED(map);
KASSERT((entry->eflags & MAP_ENTRY_IS_SUB_MAP) == 0,
("map entry %p is a submap", entry));
object = entry->object.vm_object;
if (object == NULL && !vm_map_is_system(map) &&
(entry->eflags & MAP_ENTRY_GUARD) == 0)
vm_map_entry_back(entry);
else if (object != NULL &&
((entry->eflags & MAP_ENTRY_NEEDS_COPY) == 0) &&
entry->cred != NULL) {
VM_OBJECT_WLOCK(object);
KASSERT(object->cred == NULL,
("OVERCOMMIT: %s: both cred e %p", __func__, entry));
object->cred = entry->cred;
if (entry->end - entry->start < ptoa(object->size)) {
swap_reserve_force_by_cred(ptoa(object->size) -
entry->end + entry->start, object->cred);
}
VM_OBJECT_WUNLOCK(entry->object.vm_object);
entry->cred = NULL;
}
}
/*
* vm_map_entry_clone
*
* Create a duplicate map entry for clipping.
*/
static vm_map_entry_t
vm_map_entry_clone(vm_map_t map, vm_map_entry_t entry)
{
vm_map_entry_t new_entry;
VM_MAP_ASSERT_LOCKED(map);
/*
* Create a backing object now, if none exists, so that more individual
* objects won't be created after the map entry is split.
*/
vm_map_entry_charge_object(map, entry);
/* Clone the entry. */
new_entry = vm_map_entry_create(map);
*new_entry = *entry;
if (new_entry->cred != NULL)
crhold(entry->cred);
if ((entry->eflags & MAP_ENTRY_IS_SUB_MAP) == 0) {
vm_object_reference(new_entry->object.vm_object);
vm_map_entry_set_vnode_text(new_entry, true);
/*
* The object->un_pager.vnp.writemappings for the object of
* MAP_ENTRY_WRITECNT type entry shall be kept as is here. The
* virtual pages are re-distributed among the clipped entries,
* so the sum is left the same.
*/
}
return (new_entry);
}
/*
* vm_map_clip_start: [ internal use only ]
*
* Asserts that the given entry begins at or after
* the specified address; if necessary,
* it splits the entry into two.
*/
static int
vm_map_clip_start(vm_map_t map, vm_map_entry_t entry, vm_offset_t startaddr)
{
vm_map_entry_t new_entry;
int bdry_idx;
if (!vm_map_is_system(map))
WITNESS_WARN(WARN_GIANTOK | WARN_SLEEPOK, NULL,
"%s: map %p entry %p start 0x%jx", __func__, map, entry,
(uintmax_t)startaddr);
if (startaddr <= entry->start)
return (KERN_SUCCESS);
VM_MAP_ASSERT_LOCKED(map);
KASSERT(entry->end > startaddr && entry->start < startaddr,
("%s: invalid clip of entry %p", __func__, entry));
bdry_idx = MAP_ENTRY_SPLIT_BOUNDARY_INDEX(entry);
if (bdry_idx != 0) {
if ((startaddr & (pagesizes[bdry_idx] - 1)) != 0)
return (KERN_INVALID_ARGUMENT);
}
new_entry = vm_map_entry_clone(map, entry);
/*
* Split off the front portion. Insert the new entry BEFORE this one,
* so that this entry has the specified starting address.
*/
new_entry->end = startaddr;
vm_map_entry_link(map, new_entry);
return (KERN_SUCCESS);
}
/*
* vm_map_lookup_clip_start:
*
* Find the entry at or just after 'start', and clip it if 'start' is in
* the interior of the entry. Return entry after 'start', and in
* prev_entry set the entry before 'start'.
*/
static int
vm_map_lookup_clip_start(vm_map_t map, vm_offset_t start,
vm_map_entry_t *res_entry, vm_map_entry_t *prev_entry)
{
vm_map_entry_t entry;
int rv;
if (!vm_map_is_system(map))
WITNESS_WARN(WARN_GIANTOK | WARN_SLEEPOK, NULL,
"%s: map %p start 0x%jx prev %p", __func__, map,
(uintmax_t)start, prev_entry);
if (vm_map_lookup_entry(map, start, prev_entry)) {
entry = *prev_entry;
rv = vm_map_clip_start(map, entry, start);
if (rv != KERN_SUCCESS)
return (rv);
*prev_entry = vm_map_entry_pred(entry);
} else
entry = vm_map_entry_succ(*prev_entry);
*res_entry = entry;
return (KERN_SUCCESS);
}
/*
* vm_map_clip_end: [ internal use only ]
*
* Asserts that the given entry ends at or before
* the specified address; if necessary,
* it splits the entry into two.
*/
static int
vm_map_clip_end(vm_map_t map, vm_map_entry_t entry, vm_offset_t endaddr)
{
vm_map_entry_t new_entry;
int bdry_idx;
if (!vm_map_is_system(map))
WITNESS_WARN(WARN_GIANTOK | WARN_SLEEPOK, NULL,
"%s: map %p entry %p end 0x%jx", __func__, map, entry,
(uintmax_t)endaddr);
if (endaddr >= entry->end)
return (KERN_SUCCESS);
VM_MAP_ASSERT_LOCKED(map);
KASSERT(entry->start < endaddr && entry->end > endaddr,
("%s: invalid clip of entry %p", __func__, entry));
bdry_idx = MAP_ENTRY_SPLIT_BOUNDARY_INDEX(entry);
if (bdry_idx != 0) {
if ((endaddr & (pagesizes[bdry_idx] - 1)) != 0)
return (KERN_INVALID_ARGUMENT);
}
new_entry = vm_map_entry_clone(map, entry);
/*
* Split off the back portion. Insert the new entry AFTER this one,
* so that this entry has the specified ending address.
*/
new_entry->start = endaddr;
vm_map_entry_link(map, new_entry);
return (KERN_SUCCESS);
}
/*
* vm_map_submap: [ kernel use only ]
*
* Mark the given range as handled by a subordinate map.
*
* This range must have been created with vm_map_find,
* and no other operations may have been performed on this
* range prior to calling vm_map_submap.
*
* Only a limited number of operations can be performed
* within this rage after calling vm_map_submap:
* vm_fault
* [Don't try vm_map_copy!]
*
* To remove a submapping, one must first remove the
* range from the superior map, and then destroy the
* submap (if desired). [Better yet, don't try it.]
*/
int
vm_map_submap(
vm_map_t map,
vm_offset_t start,
vm_offset_t end,
vm_map_t submap)
{
vm_map_entry_t entry;
int result;
result = KERN_INVALID_ARGUMENT;
vm_map_lock(submap);
submap->flags |= MAP_IS_SUB_MAP;
vm_map_unlock(submap);
vm_map_lock(map);
VM_MAP_RANGE_CHECK(map, start, end);
if (vm_map_lookup_entry(map, start, &entry) && entry->end >= end &&
(entry->eflags & MAP_ENTRY_COW) == 0 &&
entry->object.vm_object == NULL) {
result = vm_map_clip_start(map, entry, start);
if (result != KERN_SUCCESS)
goto unlock;
result = vm_map_clip_end(map, entry, end);
if (result != KERN_SUCCESS)
goto unlock;
entry->object.sub_map = submap;
entry->eflags |= MAP_ENTRY_IS_SUB_MAP;
result = KERN_SUCCESS;
}
unlock:
vm_map_unlock(map);
if (result != KERN_SUCCESS) {
vm_map_lock(submap);
submap->flags &= ~MAP_IS_SUB_MAP;
vm_map_unlock(submap);
}
return (result);
}
/*
* The maximum number of pages to map if MAP_PREFAULT_PARTIAL is specified
*/
#define MAX_INIT_PT 96
/*
* vm_map_pmap_enter:
*
* Preload the specified map's pmap with mappings to the specified
* object's memory-resident pages. No further physical pages are
* allocated, and no further virtual pages are retrieved from secondary
* storage. If the specified flags include MAP_PREFAULT_PARTIAL, then a
* limited number of page mappings are created at the low-end of the
* specified address range. (For this purpose, a superpage mapping
* counts as one page mapping.) Otherwise, all resident pages within
* the specified address range are mapped.
*/
static void
vm_map_pmap_enter(vm_map_t map, vm_offset_t addr, vm_prot_t prot,
vm_object_t object, vm_pindex_t pindex, vm_size_t size, int flags)
{
struct pctrie_iter pages;
vm_offset_t start;
vm_page_t p, p_start;
vm_pindex_t jump, mask, psize, threshold, tmpidx;
int psind;
if ((prot & (VM_PROT_READ | VM_PROT_EXECUTE)) == 0 || object == NULL)
return;
if (object->type == OBJT_DEVICE || object->type == OBJT_SG) {
VM_OBJECT_WLOCK(object);
if (object->type == OBJT_DEVICE || object->type == OBJT_SG) {
pmap_object_init_pt(map->pmap, addr, object, pindex,
size);
VM_OBJECT_WUNLOCK(object);
return;
}
VM_OBJECT_LOCK_DOWNGRADE(object);
} else
VM_OBJECT_RLOCK(object);
psize = atop(size);
if (psize + pindex > object->size) {
if (pindex >= object->size) {
VM_OBJECT_RUNLOCK(object);
return;
}
psize = object->size - pindex;
}
start = 0;
p_start = NULL;
threshold = MAX_INIT_PT;
vm_page_iter_limit_init(&pages, object, pindex + psize);
for (p = vm_radix_iter_lookup_ge(&pages, pindex); p != NULL;
p = vm_radix_iter_jump(&pages, jump)) {
/*
* don't allow an madvise to blow away our really
* free pages allocating pv entries.
*/
tmpidx = p->pindex - pindex;
if (((flags & MAP_PREFAULT_MADVISE) != 0 &&
vm_page_count_severe()) ||
((flags & MAP_PREFAULT_PARTIAL) != 0 &&
tmpidx >= threshold)) {
psize = tmpidx;
break;
}
jump = 1;
if (vm_page_all_valid(p)) {
if (p_start == NULL) {
start = addr + ptoa(tmpidx);
p_start = p;
}
/* Jump ahead if a superpage mapping is possible. */
for (psind = p->psind; psind > 0; psind--) {
if (((addr + ptoa(tmpidx)) &
(pagesizes[psind] - 1)) == 0) {
mask = atop(pagesizes[psind]) - 1;
if (tmpidx + mask < psize &&
vm_page_ps_test(p, psind,
PS_ALL_VALID, NULL)) {
jump += mask;
threshold += mask;
break;
}
}
}
} else if (p_start != NULL) {
pmap_enter_object(map->pmap, start, addr +
ptoa(tmpidx), p_start, prot);
p_start = NULL;
}
}
if (p_start != NULL)
pmap_enter_object(map->pmap, start, addr + ptoa(psize),
p_start, prot);
VM_OBJECT_RUNLOCK(object);
}
static void
vm_map_protect_guard(vm_map_entry_t entry, vm_prot_t new_prot,
vm_prot_t new_maxprot, int flags)
{
vm_prot_t old_prot;
MPASS((entry->eflags & MAP_ENTRY_GUARD) != 0);
if ((entry->eflags & MAP_ENTRY_STACK_GAP) == 0)
return;
old_prot = PROT_EXTRACT(entry->offset);
if ((flags & VM_MAP_PROTECT_SET_MAXPROT) != 0) {
entry->offset = PROT_MAX(new_maxprot) |
(new_maxprot & old_prot);
}
if ((flags & VM_MAP_PROTECT_SET_PROT) != 0) {
entry->offset = new_prot | PROT_MAX(
PROT_MAX_EXTRACT(entry->offset));
}
}
/*
* vm_map_protect:
*
* Sets the protection and/or the maximum protection of the
* specified address region in the target map.
*/
int
vm_map_protect(vm_map_t map, vm_offset_t start, vm_offset_t end,
vm_prot_t new_prot, vm_prot_t new_maxprot, int flags)
{
vm_map_entry_t entry, first_entry, in_tran, prev_entry;
vm_object_t obj;
struct ucred *cred;
vm_offset_t orig_start;
vm_prot_t check_prot, max_prot, old_prot;
int rv;
if (start == end)
return (KERN_SUCCESS);
if (CONTAINS_BITS(flags, VM_MAP_PROTECT_SET_PROT |
VM_MAP_PROTECT_SET_MAXPROT) &&
!CONTAINS_BITS(new_maxprot, new_prot))
return (KERN_OUT_OF_BOUNDS);
orig_start = start;
again:
in_tran = NULL;
start = orig_start;
vm_map_lock(map);
if ((map->flags & MAP_WXORX) != 0 &&
(flags & VM_MAP_PROTECT_SET_PROT) != 0 &&
CONTAINS_BITS(new_prot, VM_PROT_WRITE | VM_PROT_EXECUTE)) {
vm_map_unlock(map);
return (KERN_PROTECTION_FAILURE);
}
/*
* Ensure that we are not concurrently wiring pages. vm_map_wire() may
* need to fault pages into the map and will drop the map lock while
* doing so, and the VM object may end up in an inconsistent state if we
* update the protection on the map entry in between faults.
*/
vm_map_wait_busy(map);
VM_MAP_RANGE_CHECK(map, start, end);
if (!vm_map_lookup_entry(map, start, &first_entry))
first_entry = vm_map_entry_succ(first_entry);
if ((flags & VM_MAP_PROTECT_GROWSDOWN) != 0 &&
(first_entry->eflags & MAP_ENTRY_GROWS_DOWN) != 0) {
/*
* Handle Linux's PROT_GROWSDOWN flag.
* It means that protection is applied down to the
* whole stack, including the specified range of the
* mapped region, and the grow down region (AKA
* guard).
*/
while (!CONTAINS_BITS(first_entry->eflags,
MAP_ENTRY_GUARD | MAP_ENTRY_STACK_GAP) &&
first_entry != vm_map_entry_first(map))
first_entry = vm_map_entry_pred(first_entry);
start = first_entry->start;
}
/*
* Make a first pass to check for protection violations.
*/
check_prot = 0;
if ((flags & VM_MAP_PROTECT_SET_PROT) != 0)
check_prot |= new_prot;
if ((flags & VM_MAP_PROTECT_SET_MAXPROT) != 0)
check_prot |= new_maxprot;
for (entry = first_entry; entry->start < end;
entry = vm_map_entry_succ(entry)) {
if ((entry->eflags & MAP_ENTRY_IS_SUB_MAP) != 0) {
vm_map_unlock(map);
return (KERN_INVALID_ARGUMENT);
}
if ((entry->eflags & (MAP_ENTRY_GUARD |
MAP_ENTRY_STACK_GAP)) == MAP_ENTRY_GUARD)
continue;
max_prot = (entry->eflags & MAP_ENTRY_STACK_GAP) != 0 ?
PROT_MAX_EXTRACT(entry->offset) : entry->max_protection;
if (!CONTAINS_BITS(max_prot, check_prot)) {
vm_map_unlock(map);
return (KERN_PROTECTION_FAILURE);
}
if ((entry->eflags & MAP_ENTRY_IN_TRANSITION) != 0)
in_tran = entry;
}
/*
* Postpone the operation until all in-transition map entries have
* stabilized. An in-transition entry might already have its pages
* wired and wired_count incremented, but not yet have its
* MAP_ENTRY_USER_WIRED flag set. In which case, we would fail to call
* vm_fault_copy_entry() in the final loop below.
*/
if (in_tran != NULL) {
in_tran->eflags |= MAP_ENTRY_NEEDS_WAKEUP;
vm_map_unlock_and_wait(map, 0);
goto again;
}
/*
* Before changing the protections, try to reserve swap space for any
* private (i.e., copy-on-write) mappings that are transitioning from
* read-only to read/write access. If a reservation fails, break out
* of this loop early and let the next loop simplify the entries, since
* some may now be mergeable.
*/
rv = vm_map_clip_start(map, first_entry, start);
if (rv != KERN_SUCCESS) {
vm_map_unlock(map);
return (rv);
}
for (entry = first_entry; entry->start < end;
entry = vm_map_entry_succ(entry)) {
rv = vm_map_clip_end(map, entry, end);
if (rv != KERN_SUCCESS) {
vm_map_unlock(map);
return (rv);
}
if ((flags & VM_MAP_PROTECT_SET_PROT) == 0 ||
((new_prot & ~entry->protection) & VM_PROT_WRITE) == 0 ||
ENTRY_CHARGED(entry) ||
(entry->eflags & MAP_ENTRY_GUARD) != 0)
continue;
cred = curthread->td_ucred;
obj = entry->object.vm_object;
if (obj == NULL ||
(entry->eflags & MAP_ENTRY_NEEDS_COPY) != 0) {
if (!swap_reserve(entry->end - entry->start)) {
rv = KERN_RESOURCE_SHORTAGE;
end = entry->end;
break;
}
crhold(cred);
entry->cred = cred;
continue;
}
VM_OBJECT_WLOCK(obj);
if ((obj->flags & OBJ_SWAP) == 0) {
VM_OBJECT_WUNLOCK(obj);
continue;
}
/*
* Charge for the whole object allocation now, since
* we cannot distinguish between non-charged and
* charged clipped mapping of the same object later.
*/
KASSERT(obj->cred == NULL,
("vm_map_protect: object %p overcharged (entry %p)",
obj, entry));
if (!swap_reserve(ptoa(obj->size))) {
VM_OBJECT_WUNLOCK(obj);
rv = KERN_RESOURCE_SHORTAGE;
end = entry->end;
break;
}
crhold(cred);
obj->cred = cred;
VM_OBJECT_WUNLOCK(obj);
}
/*
* If enough swap space was available, go back and fix up protections.
* Otherwise, just simplify entries, since some may have been modified.
* [Note that clipping is not necessary the second time.]
*/
for (prev_entry = vm_map_entry_pred(first_entry), entry = first_entry;
entry->start < end;
vm_map_try_merge_entries(map, prev_entry, entry),
prev_entry = entry, entry = vm_map_entry_succ(entry)) {
if (rv != KERN_SUCCESS)
continue;
if ((entry->eflags & MAP_ENTRY_GUARD) != 0) {
vm_map_protect_guard(entry, new_prot, new_maxprot,
flags);
continue;
}
old_prot = entry->protection;
if ((flags & VM_MAP_PROTECT_SET_MAXPROT) != 0) {
entry->max_protection = new_maxprot;
entry->protection = new_maxprot & old_prot;
}
if ((flags & VM_MAP_PROTECT_SET_PROT) != 0)
entry->protection = new_prot;
/*
* For user wired map entries, the normal lazy evaluation of
* write access upgrades through soft page faults is
* undesirable. Instead, immediately copy any pages that are
* copy-on-write and enable write access in the physical map.
*/
if ((entry->eflags & MAP_ENTRY_USER_WIRED) != 0 &&
(entry->protection & VM_PROT_WRITE) != 0 &&
(old_prot & VM_PROT_WRITE) == 0)
vm_fault_copy_entry(map, map, entry, entry, NULL);
/*
* When restricting access, update the physical map. Worry
* about copy-on-write here.
*/
if ((old_prot & ~entry->protection) != 0) {
#define MASK(entry) (((entry)->eflags & MAP_ENTRY_COW) ? ~VM_PROT_WRITE : \
VM_PROT_ALL)
pmap_protect(map->pmap, entry->start,
entry->end,
entry->protection & MASK(entry));
#undef MASK
}
}
vm_map_try_merge_entries(map, prev_entry, entry);
vm_map_unlock(map);
return (rv);
}
/*
* vm_map_madvise:
*
* This routine traverses a processes map handling the madvise
* system call. Advisories are classified as either those effecting
* the vm_map_entry structure, or those effecting the underlying
* objects.
*/
int
vm_map_madvise(
vm_map_t map,
vm_offset_t start,
vm_offset_t end,
int behav)
{
vm_map_entry_t entry, prev_entry;
int rv;
bool modify_map;
/*
* Some madvise calls directly modify the vm_map_entry, in which case
* we need to use an exclusive lock on the map and we need to perform
* various clipping operations. Otherwise we only need a read-lock
* on the map.
*/
switch(behav) {
case MADV_NORMAL:
case MADV_SEQUENTIAL:
case MADV_RANDOM:
case MADV_NOSYNC:
case MADV_AUTOSYNC:
case MADV_NOCORE:
case MADV_CORE:
if (start == end)
return (0);
modify_map = true;
vm_map_lock(map);
break;
case MADV_WILLNEED:
case MADV_DONTNEED:
case MADV_FREE:
if (start == end)
return (0);
modify_map = false;
vm_map_lock_read(map);
break;
default:
return (EINVAL);
}
/*
* Locate starting entry and clip if necessary.
*/
VM_MAP_RANGE_CHECK(map, start, end);
if (modify_map) {
/*
* madvise behaviors that are implemented in the vm_map_entry.
*
* We clip the vm_map_entry so that behavioral changes are
* limited to the specified address range.
*/
rv = vm_map_lookup_clip_start(map, start, &entry, &prev_entry);
if (rv != KERN_SUCCESS) {
vm_map_unlock(map);
return (vm_mmap_to_errno(rv));
}
for (; entry->start < end; prev_entry = entry,
entry = vm_map_entry_succ(entry)) {
if ((entry->eflags & MAP_ENTRY_IS_SUB_MAP) != 0)
continue;
rv = vm_map_clip_end(map, entry, end);
if (rv != KERN_SUCCESS) {
vm_map_unlock(map);
return (vm_mmap_to_errno(rv));
}
switch (behav) {
case MADV_NORMAL:
vm_map_entry_set_behavior(entry,
MAP_ENTRY_BEHAV_NORMAL);
break;
case MADV_SEQUENTIAL:
vm_map_entry_set_behavior(entry,
MAP_ENTRY_BEHAV_SEQUENTIAL);
break;
case MADV_RANDOM:
vm_map_entry_set_behavior(entry,
MAP_ENTRY_BEHAV_RANDOM);
break;
case MADV_NOSYNC:
entry->eflags |= MAP_ENTRY_NOSYNC;
break;
case MADV_AUTOSYNC:
entry->eflags &= ~MAP_ENTRY_NOSYNC;
break;
case MADV_NOCORE:
entry->eflags |= MAP_ENTRY_NOCOREDUMP;
break;
case MADV_CORE:
entry->eflags &= ~MAP_ENTRY_NOCOREDUMP;
break;
default:
break;
}
vm_map_try_merge_entries(map, prev_entry, entry);
}
vm_map_try_merge_entries(map, prev_entry, entry);
vm_map_unlock(map);
} else {
vm_pindex_t pstart, pend;
/*
* madvise behaviors that are implemented in the underlying
* vm_object.
*
* Since we don't clip the vm_map_entry, we have to clip
* the vm_object pindex and count.
*/
if (!vm_map_lookup_entry(map, start, &entry))
entry = vm_map_entry_succ(entry);
for (; entry->start < end;
entry = vm_map_entry_succ(entry)) {
vm_offset_t useEnd, useStart;
if ((entry->eflags & (MAP_ENTRY_IS_SUB_MAP |
MAP_ENTRY_GUARD)) != 0)
continue;
/*
* MADV_FREE would otherwise rewind time to
* the creation of the shadow object. Because
* we hold the VM map read-locked, neither the
* entry's object nor the presence of a
* backing object can change.
*/
if (behav == MADV_FREE &&
entry->object.vm_object != NULL &&
entry->object.vm_object->backing_object != NULL)
continue;
pstart = OFF_TO_IDX(entry->offset);
pend = pstart + atop(entry->end - entry->start);
useStart = entry->start;
useEnd = entry->end;
if (entry->start < start) {
pstart += atop(start - entry->start);
useStart = start;
}
if (entry->end > end) {
pend -= atop(entry->end - end);
useEnd = end;
}
if (pstart >= pend)
continue;
/*
* Perform the pmap_advise() before clearing
* PGA_REFERENCED in vm_page_advise(). Otherwise, a
* concurrent pmap operation, such as pmap_remove(),
* could clear a reference in the pmap and set
* PGA_REFERENCED on the page before the pmap_advise()
* had completed. Consequently, the page would appear
* referenced based upon an old reference that
* occurred before this pmap_advise() ran.
*/
if (behav == MADV_DONTNEED || behav == MADV_FREE)
pmap_advise(map->pmap, useStart, useEnd,
behav);
vm_object_madvise(entry->object.vm_object, pstart,
pend, behav);
/*
* Pre-populate paging structures in the
* WILLNEED case. For wired entries, the
* paging structures are already populated.
*/
if (behav == MADV_WILLNEED &&
entry->wired_count == 0) {
vm_map_pmap_enter(map,
useStart,
entry->protection,
entry->object.vm_object,
pstart,
ptoa(pend - pstart),
MAP_PREFAULT_MADVISE
);
}
}
vm_map_unlock_read(map);
}
return (0);
}
/*
* vm_map_inherit:
*
* Sets the inheritance of the specified address
* range in the target map. Inheritance
* affects how the map will be shared with
* child maps at the time of vmspace_fork.
*/
int
vm_map_inherit(vm_map_t map, vm_offset_t start, vm_offset_t end,
vm_inherit_t new_inheritance)
{
vm_map_entry_t entry, lentry, prev_entry, start_entry;
int rv;
switch (new_inheritance) {
case VM_INHERIT_NONE:
case VM_INHERIT_COPY:
case VM_INHERIT_SHARE:
case VM_INHERIT_ZERO:
break;
default:
return (KERN_INVALID_ARGUMENT);
}
if (start == end)
return (KERN_SUCCESS);
vm_map_lock(map);
VM_MAP_RANGE_CHECK(map, start, end);
rv = vm_map_lookup_clip_start(map, start, &start_entry, &prev_entry);
if (rv != KERN_SUCCESS)
goto unlock;
if (vm_map_lookup_entry(map, end - 1, &lentry)) {
rv = vm_map_clip_end(map, lentry, end);
if (rv != KERN_SUCCESS)
goto unlock;
}
if (new_inheritance == VM_INHERIT_COPY) {
for (entry = start_entry; entry->start < end;
prev_entry = entry, entry = vm_map_entry_succ(entry)) {
if ((entry->eflags & MAP_ENTRY_SPLIT_BOUNDARY_MASK)
!= 0) {
rv = KERN_INVALID_ARGUMENT;
goto unlock;
}
}
}
for (entry = start_entry; entry->start < end; prev_entry = entry,
entry = vm_map_entry_succ(entry)) {
KASSERT(entry->end <= end, ("non-clipped entry %p end %jx %jx",
entry, (uintmax_t)entry->end, (uintmax_t)end));
if ((entry->eflags & MAP_ENTRY_GUARD) == 0 ||
new_inheritance != VM_INHERIT_ZERO)
entry->inheritance = new_inheritance;
vm_map_try_merge_entries(map, prev_entry, entry);
}
vm_map_try_merge_entries(map, prev_entry, entry);
unlock:
vm_map_unlock(map);
return (rv);
}
/*
* vm_map_entry_in_transition:
*
* Release the map lock, and sleep until the entry is no longer in
* transition. Awake and acquire the map lock. If the map changed while
* another held the lock, lookup a possibly-changed entry at or after the
* 'start' position of the old entry.
*/
static vm_map_entry_t
vm_map_entry_in_transition(vm_map_t map, vm_offset_t in_start,
vm_offset_t *io_end, bool holes_ok, vm_map_entry_t in_entry)
{
vm_map_entry_t entry;
vm_offset_t start;
u_int last_timestamp;
VM_MAP_ASSERT_LOCKED(map);
KASSERT((in_entry->eflags & MAP_ENTRY_IN_TRANSITION) != 0,
("not in-tranition map entry %p", in_entry));
/*
* We have not yet clipped the entry.
*/
start = MAX(in_start, in_entry->start);
in_entry->eflags |= MAP_ENTRY_NEEDS_WAKEUP;
last_timestamp = map->timestamp;
if (vm_map_unlock_and_wait(map, 0)) {
/*
* Allow interruption of user wiring/unwiring?
*/
}
vm_map_lock(map);
if (last_timestamp + 1 == map->timestamp)
return (in_entry);
/*
* Look again for the entry because the map was modified while it was
* unlocked. Specifically, the entry may have been clipped, merged, or
* deleted.
*/
if (!vm_map_lookup_entry(map, start, &entry)) {
if (!holes_ok) {
*io_end = start;
return (NULL);
}
entry = vm_map_entry_succ(entry);
}
return (entry);
}
/*
* vm_map_unwire:
*
* Implements both kernel and user unwiring.
*/
int
vm_map_unwire(vm_map_t map, vm_offset_t start, vm_offset_t end,
int flags)
{
vm_map_entry_t entry, first_entry, next_entry, prev_entry;
int rv;
bool holes_ok, need_wakeup, user_unwire;
if (start == end)
return (KERN_SUCCESS);
holes_ok = (flags & VM_MAP_WIRE_HOLESOK) != 0;
user_unwire = (flags & VM_MAP_WIRE_USER) != 0;
vm_map_lock(map);
VM_MAP_RANGE_CHECK(map, start, end);
if (!vm_map_lookup_entry(map, start, &first_entry)) {
if (holes_ok)
first_entry = vm_map_entry_succ(first_entry);
else {
vm_map_unlock(map);
return (KERN_INVALID_ADDRESS);
}
}
rv = KERN_SUCCESS;
for (entry = first_entry; entry->start < end; entry = next_entry) {
if (entry->eflags & MAP_ENTRY_IN_TRANSITION) {
/*
* We have not yet clipped the entry.
*/
next_entry = vm_map_entry_in_transition(map, start,
&end, holes_ok, entry);
if (next_entry == NULL) {
if (entry == first_entry) {
vm_map_unlock(map);
return (KERN_INVALID_ADDRESS);
}
rv = KERN_INVALID_ADDRESS;
break;
}
first_entry = (entry == first_entry) ?
next_entry : NULL;
continue;
}
rv = vm_map_clip_start(map, entry, start);
if (rv != KERN_SUCCESS)
break;
rv = vm_map_clip_end(map, entry, end);
if (rv != KERN_SUCCESS)
break;
/*
* Mark the entry in case the map lock is released. (See
* above.)
*/
KASSERT((entry->eflags & MAP_ENTRY_IN_TRANSITION) == 0 &&
entry->wiring_thread == NULL,
("owned map entry %p", entry));
entry->eflags |= MAP_ENTRY_IN_TRANSITION;
entry->wiring_thread = curthread;
next_entry = vm_map_entry_succ(entry);
/*
* Check the map for holes in the specified region.
* If holes_ok, skip this check.
*/
if (!holes_ok &&
entry->end < end && next_entry->start > entry->end) {
end = entry->end;
rv = KERN_INVALID_ADDRESS;
break;
}
/*
* If system unwiring, require that the entry is system wired.
*/
if (!user_unwire &&
vm_map_entry_system_wired_count(entry) == 0) {
end = entry->end;
rv = KERN_INVALID_ARGUMENT;
break;
}
}
need_wakeup = false;
if (first_entry == NULL &&
!vm_map_lookup_entry(map, start, &first_entry)) {
KASSERT(holes_ok, ("vm_map_unwire: lookup failed"));
prev_entry = first_entry;
entry = vm_map_entry_succ(first_entry);
} else {
prev_entry = vm_map_entry_pred(first_entry);
entry = first_entry;
}
for (; entry->start < end;
prev_entry = entry, entry = vm_map_entry_succ(entry)) {
/*
* If holes_ok was specified, an empty
* space in the unwired region could have been mapped
* while the map lock was dropped for draining
* MAP_ENTRY_IN_TRANSITION. Moreover, another thread
* could be simultaneously wiring this new mapping
* entry. Detect these cases and skip any entries
* marked as in transition by us.
*/
if ((entry->eflags & MAP_ENTRY_IN_TRANSITION) == 0 ||
entry->wiring_thread != curthread) {
KASSERT(holes_ok,
("vm_map_unwire: !HOLESOK and new/changed entry"));
continue;
}
if (rv == KERN_SUCCESS && (!user_unwire ||
(entry->eflags & MAP_ENTRY_USER_WIRED))) {
if (entry->wired_count == 1)
vm_map_entry_unwire(map, entry);
else
entry->wired_count--;
if (user_unwire)
entry->eflags &= ~MAP_ENTRY_USER_WIRED;
}
KASSERT((entry->eflags & MAP_ENTRY_IN_TRANSITION) != 0,
("vm_map_unwire: in-transition flag missing %p", entry));
KASSERT(entry->wiring_thread == curthread,
("vm_map_unwire: alien wire %p", entry));
entry->eflags &= ~MAP_ENTRY_IN_TRANSITION;
entry->wiring_thread = NULL;
if (entry->eflags & MAP_ENTRY_NEEDS_WAKEUP) {
entry->eflags &= ~MAP_ENTRY_NEEDS_WAKEUP;
need_wakeup = true;
}
vm_map_try_merge_entries(map, prev_entry, entry);
}
vm_map_try_merge_entries(map, prev_entry, entry);
vm_map_unlock(map);
if (need_wakeup)
vm_map_wakeup(map);
return (rv);
}
static void
vm_map_wire_user_count_sub(u_long npages)
{
atomic_subtract_long(&vm_user_wire_count, npages);
}
static bool
vm_map_wire_user_count_add(u_long npages)
{
u_long wired;
wired = vm_user_wire_count;
do {
if (npages + wired > vm_page_max_user_wired)
return (false);
} while (!atomic_fcmpset_long(&vm_user_wire_count, &wired,
npages + wired));
return (true);
}
/*
* vm_map_wire_entry_failure:
*
* Handle a wiring failure on the given entry.
*
* The map should be locked.
*/
static void
vm_map_wire_entry_failure(vm_map_t map, vm_map_entry_t entry,
vm_offset_t failed_addr)
{
VM_MAP_ASSERT_LOCKED(map);
KASSERT((entry->eflags & MAP_ENTRY_IN_TRANSITION) != 0 &&
entry->wired_count == 1,
("vm_map_wire_entry_failure: entry %p isn't being wired", entry));
KASSERT(failed_addr < entry->end,
("vm_map_wire_entry_failure: entry %p was fully wired", entry));
/*
* If any pages at the start of this entry were successfully wired,
* then unwire them.
*/
if (failed_addr > entry->start) {
pmap_unwire(map->pmap, entry->start, failed_addr);
vm_object_unwire(entry->object.vm_object, entry->offset,
failed_addr - entry->start, PQ_ACTIVE);
}
/*
* Assign an out-of-range value to represent the failure to wire this
* entry.
*/
entry->wired_count = -1;
}
int
vm_map_wire(vm_map_t map, vm_offset_t start, vm_offset_t end, int flags)
{
int rv;
vm_map_lock(map);
rv = vm_map_wire_locked(map, start, end, flags);
vm_map_unlock(map);
return (rv);
}
/*
* vm_map_wire_locked:
*
* Implements both kernel and user wiring. Returns with the map locked,
* the map lock may be dropped.
*/
int
vm_map_wire_locked(vm_map_t map, vm_offset_t start, vm_offset_t end, int flags)
{
vm_map_entry_t entry, first_entry, next_entry, prev_entry;
vm_offset_t faddr, saved_end, saved_start;
u_long incr, npages;
u_int bidx, last_timestamp;
int rv;
bool holes_ok, need_wakeup, user_wire;
vm_prot_t prot;
VM_MAP_ASSERT_LOCKED(map);
if (start == end)
return (KERN_SUCCESS);
prot = 0;
if (flags & VM_MAP_WIRE_WRITE)
prot |= VM_PROT_WRITE;
holes_ok = (flags & VM_MAP_WIRE_HOLESOK) != 0;
user_wire = (flags & VM_MAP_WIRE_USER) != 0;
VM_MAP_RANGE_CHECK(map, start, end);
if (!vm_map_lookup_entry(map, start, &first_entry)) {
if (holes_ok)
first_entry = vm_map_entry_succ(first_entry);
else
return (KERN_INVALID_ADDRESS);
}
for (entry = first_entry; entry->start < end; entry = next_entry) {
if (entry->eflags & MAP_ENTRY_IN_TRANSITION) {
/*
* We have not yet clipped the entry.
*/
next_entry = vm_map_entry_in_transition(map, start,
&end, holes_ok, entry);
if (next_entry == NULL) {
if (entry == first_entry)
return (KERN_INVALID_ADDRESS);
rv = KERN_INVALID_ADDRESS;
goto done;
}
first_entry = (entry == first_entry) ?
next_entry : NULL;
continue;
}
rv = vm_map_clip_start(map, entry, start);
if (rv != KERN_SUCCESS)
goto done;
rv = vm_map_clip_end(map, entry, end);
if (rv != KERN_SUCCESS)
goto done;
/*
* Mark the entry in case the map lock is released. (See
* above.)
*/
KASSERT((entry->eflags & MAP_ENTRY_IN_TRANSITION) == 0 &&
entry->wiring_thread == NULL,
("owned map entry %p", entry));
entry->eflags |= MAP_ENTRY_IN_TRANSITION;
entry->wiring_thread = curthread;
if ((entry->protection & (VM_PROT_READ | VM_PROT_EXECUTE)) == 0
|| (entry->protection & prot) != prot) {
entry->eflags |= MAP_ENTRY_WIRE_SKIPPED;
if (!holes_ok) {
end = entry->end;
rv = KERN_INVALID_ADDRESS;
goto done;
}
+ } else if (user_wire &&
+ (entry->eflags & MAP_ENTRY_IS_SUB_MAP) == 0 &&
+ entry->object.vm_object != NULL &&
+ (entry->object.vm_object->flags & OBJ_NOMLOCK) != 0) {
+ /*
+ * The pager owns residency for this special device mapping.
+ * Retain range/hole checks and transition cleanup, but do not
+ * claim to pin its pages or acquire a user wire reference.
+ * Kernel wiring must still satisfy the ordinary contract.
+ */
+ entry->eflags |= MAP_ENTRY_WIRE_SKIPPED;
} else if (entry->wired_count == 0) {
entry->wired_count++;
npages = atop(entry->end - entry->start);
if (user_wire && !vm_map_wire_user_count_add(npages)) {
vm_map_wire_entry_failure(map, entry,
entry->start);
end = entry->end;
rv = KERN_RESOURCE_SHORTAGE;
goto done;
}
/*
* Release the map lock, relying on the in-transition
* mark. Mark the map busy for fork.
*/
saved_start = entry->start;
saved_end = entry->end;
last_timestamp = map->timestamp;
bidx = MAP_ENTRY_SPLIT_BOUNDARY_INDEX(entry);
incr = pagesizes[bidx];
vm_map_busy(map);
vm_map_unlock(map);
for (faddr = saved_start; faddr < saved_end;
faddr += incr) {
/*
* Simulate a fault to get the page and enter
* it into the physical map.
*/
rv = vm_fault(map, faddr, VM_PROT_NONE,
VM_FAULT_WIRE, NULL);
if (rv != KERN_SUCCESS)
break;
}
vm_map_lock(map);
vm_map_unbusy(map);
if (last_timestamp + 1 != map->timestamp) {
/*
* Look again for the entry because the map was
* modified while it was unlocked. The entry
* may have been clipped, but NOT merged or
* deleted.
*/
if (!vm_map_lookup_entry(map, saved_start,
&next_entry))
KASSERT(false,
("vm_map_wire: lookup failed"));
first_entry = (entry == first_entry) ?
next_entry : NULL;
for (entry = next_entry; entry->end < saved_end;
entry = vm_map_entry_succ(entry)) {
/*
* In case of failure, handle entries
* that were not fully wired here;
* fully wired entries are handled
* later.
*/
if (rv != KERN_SUCCESS &&
faddr < entry->end)
vm_map_wire_entry_failure(map,
entry, faddr);
}
}
if (rv != KERN_SUCCESS) {
vm_map_wire_entry_failure(map, entry, faddr);
if (user_wire)
vm_map_wire_user_count_sub(npages);
end = entry->end;
goto done;
}
} else if (!user_wire ||
(entry->eflags & MAP_ENTRY_USER_WIRED) == 0) {
entry->wired_count++;
}
/*
* Check the map for holes in the specified region.
* If holes_ok was specified, skip this check.
*/
next_entry = vm_map_entry_succ(entry);
if (!holes_ok &&
entry->end < end && next_entry->start > entry->end) {
end = entry->end;
rv = KERN_INVALID_ADDRESS;
goto done;
}
}
rv = KERN_SUCCESS;
done:
need_wakeup = false;
if (first_entry == NULL &&
!vm_map_lookup_entry(map, start, &first_entry)) {
KASSERT(holes_ok, ("vm_map_wire: lookup failed"));
prev_entry = first_entry;
entry = vm_map_entry_succ(first_entry);
} else {
prev_entry = vm_map_entry_pred(first_entry);
entry = first_entry;
}
for (; entry->start < end;
prev_entry = entry, entry = vm_map_entry_succ(entry)) {
/*
* If holes_ok was specified, an empty
* space in the unwired region could have been mapped
* while the map lock was dropped for faulting in the
* pages or draining MAP_ENTRY_IN_TRANSITION.
* Moreover, another thread could be simultaneously
* wiring this new mapping entry. Detect these cases
* and skip any entries marked as in transition not by us.
*
* Another way to get an entry not marked with
* MAP_ENTRY_IN_TRANSITION is after failed clipping,
* which set rv to KERN_INVALID_ARGUMENT.
*/
if ((entry->eflags & MAP_ENTRY_IN_TRANSITION) == 0 ||
entry->wiring_thread != curthread) {
KASSERT(holes_ok || rv == KERN_INVALID_ARGUMENT,
("vm_map_wire: !HOLESOK and new/changed entry"));
continue;
}
if ((entry->eflags & MAP_ENTRY_WIRE_SKIPPED) != 0) {
/* do nothing */
} else if (rv == KERN_SUCCESS) {
if (user_wire)
entry->eflags |= MAP_ENTRY_USER_WIRED;
} else if (entry->wired_count == -1) {
/*
* Wiring failed on this entry. Thus, unwiring is
* unnecessary.
*/
entry->wired_count = 0;
} else if (!user_wire ||
(entry->eflags & MAP_ENTRY_USER_WIRED) == 0) {
/*
* Undo the wiring. Wiring succeeded on this entry
* but failed on a later entry.
*/
if (entry->wired_count == 1) {
vm_map_entry_unwire(map, entry);
if (user_wire)
vm_map_wire_user_count_sub(
atop(entry->end - entry->start));
} else
entry->wired_count--;
}
KASSERT((entry->eflags & MAP_ENTRY_IN_TRANSITION) != 0,
("vm_map_wire: in-transition flag missing %p", entry));
KASSERT(entry->wiring_thread == curthread,
("vm_map_wire: alien wire %p", entry));
entry->eflags &= ~(MAP_ENTRY_IN_TRANSITION |
MAP_ENTRY_WIRE_SKIPPED);
entry->wiring_thread = NULL;
if (entry->eflags & MAP_ENTRY_NEEDS_WAKEUP) {
entry->eflags &= ~MAP_ENTRY_NEEDS_WAKEUP;
need_wakeup = true;
}
vm_map_try_merge_entries(map, prev_entry, entry);
}
vm_map_try_merge_entries(map, prev_entry, entry);
if (need_wakeup)
vm_map_wakeup(map);
return (rv);
}
/*
* vm_map_sync
*
* Push any dirty cached pages in the address range to their pager.
* If syncio is TRUE, dirty pages are written synchronously.
* If invalidate is TRUE, any cached pages are freed as well.
*
* If the size of the region from start to end is zero, we are
* supposed to flush all modified pages within the region containing
* start. Unfortunately, a region can be split or coalesced with
* neighboring regions, making it difficult to determine what the
* original region was. Therefore, we approximate this requirement by
* flushing the current region containing start.
*
* Returns an error if any part of the specified range is not mapped.
*/
int
vm_map_sync(
vm_map_t map,
vm_offset_t start,
vm_offset_t end,
boolean_t syncio,
boolean_t invalidate)
{
vm_map_entry_t entry, first_entry, next_entry;
vm_size_t size;
vm_object_t object;
vm_ooffset_t offset;
unsigned int last_timestamp;
int bdry_idx;
boolean_t failed;
vm_map_lock_read(map);
VM_MAP_RANGE_CHECK(map, start, end);
if (!vm_map_lookup_entry(map, start, &first_entry)) {
vm_map_unlock_read(map);
return (KERN_INVALID_ADDRESS);
} else if (start == end) {
start = first_entry->start;
end = first_entry->end;
}
/*
* Make a first pass to check for user-wired memory, holes,
* and partial invalidation of largepage mappings.
*/
for (entry = first_entry; entry->start < end; entry = next_entry) {
if (invalidate) {
if ((entry->eflags & MAP_ENTRY_USER_WIRED) != 0) {
vm_map_unlock_read(map);
return (KERN_INVALID_ARGUMENT);
}
bdry_idx = MAP_ENTRY_SPLIT_BOUNDARY_INDEX(entry);
if (bdry_idx != 0 &&
((start & (pagesizes[bdry_idx] - 1)) != 0 ||
(end & (pagesizes[bdry_idx] - 1)) != 0)) {
vm_map_unlock_read(map);
return (KERN_INVALID_ARGUMENT);
}
}
next_entry = vm_map_entry_succ(entry);
if (end > entry->end &&
entry->end != next_entry->start) {
vm_map_unlock_read(map);
return (KERN_INVALID_ADDRESS);
}
}
if (invalidate)
pmap_remove(map->pmap, start, end);
failed = FALSE;
/*
* Make a second pass, cleaning/uncaching pages from the indicated
* objects as we go.
*/
for (entry = first_entry; entry->start < end;) {
offset = entry->offset + (start - entry->start);
size = (end <= entry->end ? end : entry->end) - start;
if ((entry->eflags & MAP_ENTRY_IS_SUB_MAP) != 0) {
vm_map_t smap;
vm_map_entry_t tentry;
vm_size_t tsize;
smap = entry->object.sub_map;
vm_map_lock_read(smap);
(void) vm_map_lookup_entry(smap, offset, &tentry);
tsize = tentry->end - offset;
if (tsize < size)
size = tsize;
object = tentry->object.vm_object;
offset = tentry->offset + (offset - tentry->start);
vm_map_unlock_read(smap);
} else {
object = entry->object.vm_object;
}
vm_object_reference(object);
last_timestamp = map->timestamp;
vm_map_unlock_read(map);
if (!vm_object_sync(object, offset, size, syncio, invalidate))
failed = TRUE;
start += size;
vm_object_deallocate(object);
vm_map_lock_read(map);
if (last_timestamp == map->timestamp ||
!vm_map_lookup_entry(map, start, &entry))
entry = vm_map_entry_succ(entry);
}
vm_map_unlock_read(map);
return (failed ? KERN_FAILURE : KERN_SUCCESS);
}
/*
* vm_map_entry_unwire: [ internal use only ]
*
* Make the region specified by this entry pageable.
*
* The map in question should be locked.
* [This is the reason for this routine's existence.]
*/
static void
vm_map_entry_unwire(vm_map_t map, vm_map_entry_t entry)
{
vm_size_t size;
VM_MAP_ASSERT_LOCKED(map);
KASSERT(entry->wired_count > 0,
("vm_map_entry_unwire: entry %p isn't wired", entry));
size = entry->end - entry->start;
if ((entry->eflags & MAP_ENTRY_USER_WIRED) != 0)
vm_map_wire_user_count_sub(atop(size));
pmap_unwire(map->pmap, entry->start, entry->end);
vm_object_unwire(entry->object.vm_object, entry->offset, size,
PQ_ACTIVE);
entry->wired_count = 0;
}
static void
vm_map_entry_deallocate(vm_map_entry_t entry, boolean_t system_map)
{
if ((entry->eflags & MAP_ENTRY_IS_SUB_MAP) == 0)
vm_object_deallocate(entry->object.vm_object);
uma_zfree(system_map ? kmapentzone : mapentzone, entry);
}
/*
* vm_map_entry_delete: [ internal use only ]
*
* Deallocate the given entry from the target map.
*/
static void
vm_map_entry_delete(vm_map_t map, vm_map_entry_t entry)
{
vm_object_t object;
vm_pindex_t offidxstart, offidxend, oldsize;
vm_size_t size;
vm_map_entry_unlink(map, entry, UNLINK_MERGE_NONE);
object = entry->object.vm_object;
if ((entry->eflags & MAP_ENTRY_GUARD) != 0) {
MPASS(entry->cred == NULL);
MPASS((entry->eflags & MAP_ENTRY_IS_SUB_MAP) == 0);
MPASS(object == NULL);
vm_map_entry_deallocate(entry, vm_map_is_system(map));
return;
}
size = entry->end - entry->start;
map->size -= size;
if (entry->cred != NULL) {
swap_release_by_cred(size, entry->cred);
crfree(entry->cred);
}
if ((entry->eflags & MAP_ENTRY_IS_SUB_MAP) != 0 || object == NULL) {
entry->object.vm_object = NULL;
} else if ((object->flags & OBJ_ANON) != 0 ||
object == kernel_object) {
KASSERT(entry->cred == NULL || object->cred == NULL ||
(entry->eflags & MAP_ENTRY_NEEDS_COPY),
("OVERCOMMIT vm_map_entry_delete: both cred %p", entry));
offidxstart = OFF_TO_IDX(entry->offset);
offidxend = offidxstart + atop(size);
VM_OBJECT_WLOCK(object);
if (object->ref_count != 1 &&
((object->flags & OBJ_ONEMAPPING) != 0 ||
object == kernel_object)) {
vm_object_collapse(object);
/*
* The option OBJPR_NOTMAPPED can be passed here
* because vm_map_delete() already performed
* pmap_remove() on the only mapping to this range
* of pages.
*/
vm_object_page_remove(object, offidxstart, offidxend,
OBJPR_NOTMAPPED);
if (offidxend >= object->size &&
offidxstart < object->size) {
oldsize = object->size;
object->size = offidxstart;
if (object->cred != NULL) {
swap_release_by_cred(ptoa(oldsize -
object->size), object->cred);
}
}
}
VM_OBJECT_WUNLOCK(object);
}
if (vm_map_is_system(map))
vm_map_entry_deallocate(entry, TRUE);
else {
entry->defer_next = curthread->td_map_def_user;
curthread->td_map_def_user = entry;
}
}
/*
* vm_map_delete: [ internal use only ]
*
* Deallocates the given address range from the target
* map.
*/
int
vm_map_delete(vm_map_t map, vm_offset_t start, vm_offset_t end)
{
vm_map_entry_t entry, next_entry, scratch_entry;
int rv;
VM_MAP_ASSERT_LOCKED(map);
if (start == end)
return (KERN_SUCCESS);
/*
* Find the start of the region, and clip it.
* Step through all entries in this region.
*/
rv = vm_map_lookup_clip_start(map, start, &entry, &scratch_entry);
if (rv != KERN_SUCCESS)
return (rv);
for (; entry->start < end; entry = next_entry) {
/*
* Wait for wiring or unwiring of an entry to complete.
* Also wait for any system wirings to disappear on
* user maps.
*/
if ((entry->eflags & MAP_ENTRY_IN_TRANSITION) != 0 ||
(vm_map_pmap(map) != kernel_pmap &&
vm_map_entry_system_wired_count(entry) != 0)) {
unsigned int last_timestamp;
vm_offset_t saved_start;
saved_start = entry->start;
entry->eflags |= MAP_ENTRY_NEEDS_WAKEUP;
last_timestamp = map->timestamp;
(void) vm_map_unlock_and_wait(map, 0);
vm_map_lock(map);
if (last_timestamp + 1 != map->timestamp) {
/*
* Look again for the entry because the map was
* modified while it was unlocked.
* Specifically, the entry may have been
* clipped, merged, or deleted.
*/
rv = vm_map_lookup_clip_start(map, saved_start,
&next_entry, &scratch_entry);
if (rv != KERN_SUCCESS)
break;
} else
next_entry = entry;
continue;
}
/* XXXKIB or delete to the upper superpage boundary ? */
rv = vm_map_clip_end(map, entry, end);
if (rv != KERN_SUCCESS)
break;
next_entry = vm_map_entry_succ(entry);
/*
* Unwire before removing addresses from the pmap; otherwise,
* unwiring will put the entries back in the pmap.
*/
if (entry->wired_count != 0)
vm_map_entry_unwire(map, entry);
/*
* Remove mappings for the pages, but only if the
* mappings could exist. For instance, it does not
* make sense to call pmap_remove() for guard entries.
*/
if ((entry->eflags & MAP_ENTRY_IS_SUB_MAP) != 0 ||
entry->object.vm_object != NULL)
pmap_map_delete(map->pmap, entry->start, entry->end);
/*
* Delete the entry only after removing all pmap
* entries pointing to its pages. (Otherwise, its
* page frames may be reallocated, and any modify bits
* will be set in the wrong object!)
*/
vm_map_entry_delete(map, entry);
}
return (rv);
}
/*
* vm_map_remove:
*
* Remove the given address range from the target map.
* This is the exported form of vm_map_delete.
*/
int
vm_map_remove(vm_map_t map, vm_offset_t start, vm_offset_t end)
{
int result;
vm_map_lock(map);
VM_MAP_RANGE_CHECK(map, start, end);
result = vm_map_delete(map, start, end);
vm_map_unlock(map);
return (result);
}
/*
* vm_map_check_protection:
*
* Assert that the target map allows the specified privilege on the
* entire address region given. The entire region must be allocated.
*
* WARNING! This code does not and should not check whether the
* contents of the region is accessible. For example a smaller file
* might be mapped into a larger address space.
*
* NOTE! This code is also called by munmap().
*
* The map must be locked. A read lock is sufficient.
*/
boolean_t
vm_map_check_protection(vm_map_t map, vm_offset_t start, vm_offset_t end,
vm_prot_t protection)
{
vm_map_entry_t entry;
vm_map_entry_t tmp_entry;
if (!vm_map_lookup_entry(map, start, &tmp_entry))
return (FALSE);
entry = tmp_entry;
while (start < end) {
/*
* No holes allowed!
*/
if (start < entry->start)
return (FALSE);
/*
* Check protection associated with entry.
*/
if ((entry->protection & protection) != protection)
return (FALSE);
/* go to next entry */
start = entry->end;
entry = vm_map_entry_succ(entry);
}
return (TRUE);
}
/*
* Check whether the specified range partially overlaps a map entry with
* fixed boundaries, and return false if so.
*
* The map must be locked.
*/
bool
vm_map_check_boundary(vm_map_t map, vm_offset_t start, vm_offset_t end)
{
vm_map_entry_t entry;
int bdry_idx;
if (!vm_map_range_valid(map, start, end))
return (false);
if (start == end)
return (true);
if (vm_map_lookup_entry(map, start, &entry)) {
bdry_idx = MAP_ENTRY_SPLIT_BOUNDARY_INDEX(entry);
if (bdry_idx != 0 &&
(start & (pagesizes[bdry_idx] - 1)) != 0)
return (false);
}
if (vm_map_lookup_entry(map, end - 1, &entry)) {
bdry_idx = MAP_ENTRY_SPLIT_BOUNDARY_INDEX(entry);
if (bdry_idx != 0 &&
(end & (pagesizes[bdry_idx] - 1)) != 0)
return (false);
}
return (true);
}
/*
*
* vm_map_copy_swap_object:
*
* Copies a swap-backed object from an existing map entry to a
* new one. Carries forward the swap charge. May change the
* src object on return.
*/
static void
vm_map_copy_swap_object(vm_map_entry_t src_entry, vm_map_entry_t dst_entry,
vm_offset_t size, vm_ooffset_t *fork_charge)
{
vm_object_t src_object;
struct ucred *cred;
int charged;
src_object = src_entry->object.vm_object;
charged = ENTRY_CHARGED(src_entry);
if ((src_object->flags & OBJ_ANON) != 0) {
VM_OBJECT_WLOCK(src_object);
vm_object_collapse(src_object);
if ((src_object->flags & OBJ_ONEMAPPING) != 0) {
vm_object_split(src_entry);
src_object = src_entry->object.vm_object;
}
vm_object_reference_locked(src_object);
vm_object_clear_flag(src_object, OBJ_ONEMAPPING);
VM_OBJECT_WUNLOCK(src_object);
} else
vm_object_reference(src_object);
if (src_entry->cred != NULL &&
!(src_entry->eflags & MAP_ENTRY_NEEDS_COPY)) {
KASSERT(src_object->cred == NULL,
("OVERCOMMIT: vm_map_copy_anon_entry: cred %p",
src_object));
src_object->cred = src_entry->cred;
*fork_charge += ptoa(src_object->size) - size;
}
dst_entry->object.vm_object = src_object;
if (charged) {
cred = curthread->td_ucred;
crhold(cred);
dst_entry->cred = cred;
*fork_charge += size;
if (!(src_entry->eflags & MAP_ENTRY_NEEDS_COPY)) {
crhold(cred);
src_entry->cred = cred;
*fork_charge += size;
}
}
}
/*
* vm_map_copy_entry:
*
* Copies the contents of the source entry to the destination
* entry. The entries *must* be aligned properly.
*/
static void
vm_map_copy_entry(
vm_map_t src_map,
vm_map_t dst_map,
vm_map_entry_t src_entry,
vm_map_entry_t dst_entry,
vm_ooffset_t *fork_charge)
{
vm_object_t src_object;
vm_map_entry_t fake_entry;
vm_offset_t size;
VM_MAP_ASSERT_LOCKED(dst_map);
if ((dst_entry->eflags|src_entry->eflags) & MAP_ENTRY_IS_SUB_MAP)
return;
if (src_entry->wired_count == 0 ||
(src_entry->protection & VM_PROT_WRITE) == 0) {
/*
* If the source entry is marked needs_copy, it is already
* write-protected.
*/
if ((src_entry->eflags & MAP_ENTRY_NEEDS_COPY) == 0 &&
(src_entry->protection & VM_PROT_WRITE) != 0) {
pmap_protect(src_map->pmap,
src_entry->start,
src_entry->end,
src_entry->protection & ~VM_PROT_WRITE);
}
/*
* Make a copy of the object.
*/
size = src_entry->end - src_entry->start;
if ((src_object = src_entry->object.vm_object) != NULL) {
if ((src_object->flags & OBJ_SWAP) != 0) {
vm_map_copy_swap_object(src_entry, dst_entry,
size, fork_charge);
/* May have split/collapsed, reload obj. */
src_object = src_entry->object.vm_object;
} else {
vm_object_reference(src_object);
dst_entry->object.vm_object = src_object;
}
src_entry->eflags |= MAP_ENTRY_COW |
MAP_ENTRY_NEEDS_COPY;
dst_entry->eflags |= MAP_ENTRY_COW |
MAP_ENTRY_NEEDS_COPY;
dst_entry->offset = src_entry->offset;
if (src_entry->eflags & MAP_ENTRY_WRITECNT) {
/*
* MAP_ENTRY_WRITECNT cannot
* indicate write reference from
* src_entry, since the entry is
* marked as needs copy. Allocate a
* fake entry that is used to
* decrement object->un_pager writecount
* at the appropriate time. Attach
* fake_entry to the deferred list.
*/
fake_entry = vm_map_entry_create(dst_map);
fake_entry->eflags = MAP_ENTRY_WRITECNT;
src_entry->eflags &= ~MAP_ENTRY_WRITECNT;
vm_object_reference(src_object);
fake_entry->object.vm_object = src_object;
fake_entry->start = src_entry->start;
fake_entry->end = src_entry->end;
fake_entry->defer_next =
curthread->td_map_def_user;
curthread->td_map_def_user = fake_entry;
}
pmap_copy(dst_map->pmap, src_map->pmap,
dst_entry->start, dst_entry->end - dst_entry->start,
src_entry->start);
} else {
dst_entry->object.vm_object = NULL;
if ((dst_entry->eflags & MAP_ENTRY_GUARD) == 0)
dst_entry->offset = 0;
if (src_entry->cred != NULL) {
dst_entry->cred = curthread->td_ucred;
crhold(dst_entry->cred);
*fork_charge += size;
}
}
} else {
/*
* We don't want to make writeable wired pages copy-on-write.
* Immediately copy these pages into the new map by simulating
* page faults. The new pages are pageable.
*/
vm_fault_copy_entry(dst_map, src_map, dst_entry, src_entry,
fork_charge);
}
}
/*
* vmspace_map_entry_forked:
* Update the newly-forked vmspace each time a map entry is inherited
* or copied. The values for vm_dsize and vm_tsize are approximate
* (and mostly-obsolete ideas in the face of mmap(2) et al.)
*/
static void
vmspace_map_entry_forked(const struct vmspace *vm1, struct vmspace *vm2,
vm_map_entry_t entry)
{
vm_size_t entrysize;
vm_offset_t newend;
if ((entry->eflags & MAP_ENTRY_GUARD) != 0)
return;
entrysize = entry->end - entry->start;
vm2->vm_map.size += entrysize;
if ((entry->eflags & MAP_ENTRY_GROWS_DOWN) != 0) {
vm2->vm_ssize += btoc(entrysize);
} else if (entry->start >= (vm_offset_t)vm1->vm_daddr &&
entry->start < (vm_offset_t)vm1->vm_daddr + ctob(vm1->vm_dsize)) {
newend = MIN(entry->end,
(vm_offset_t)vm1->vm_daddr + ctob(vm1->vm_dsize));
vm2->vm_dsize += btoc(newend - entry->start);
} else if (entry->start >= (vm_offset_t)vm1->vm_taddr &&
entry->start < (vm_offset_t)vm1->vm_taddr + ctob(vm1->vm_tsize)) {
newend = MIN(entry->end,
(vm_offset_t)vm1->vm_taddr + ctob(vm1->vm_tsize));
vm2->vm_tsize += btoc(newend - entry->start);
}
}
/*
* vmspace_fork:
* Create a new process vmspace structure and vm_map
* based on those of an existing process. The new map
* is based on the old map, according to the inheritance
* values on the regions in that map.
*
* XXX It might be worth coalescing the entries added to the new vmspace.
*
* The source map must not be locked.
*/
struct vmspace *
vmspace_fork(struct vmspace *vm1, vm_ooffset_t *fork_charge)
{
struct vmspace *vm2;
vm_map_t new_map, old_map;
vm_map_entry_t new_entry, old_entry;
vm_object_t object;
int error, locked __diagused;
vm_inherit_t inh;
old_map = &vm1->vm_map;
/* Copy immutable fields of vm1 to vm2. */
vm2 = vmspace_alloc(vm_map_min(old_map), vm_map_max(old_map),
pmap_pinit);
if (vm2 == NULL)
return (NULL);
vm2->vm_taddr = vm1->vm_taddr;
vm2->vm_daddr = vm1->vm_daddr;
vm2->vm_maxsaddr = vm1->vm_maxsaddr;
vm2->vm_stacktop = vm1->vm_stacktop;
vm2->vm_shp_base = vm1->vm_shp_base;
vm_map_lock(old_map);
if (old_map->busy)
vm_map_wait_busy(old_map);
new_map = &vm2->vm_map;
locked = vm_map_trylock(new_map); /* trylock to silence WITNESS */
KASSERT(locked, ("vmspace_fork: lock failed"));
error = pmap_vmspace_copy(new_map->pmap, old_map->pmap);
if (error != 0) {
sx_xunlock(&old_map->lock);
sx_xunlock(&new_map->lock);
vm_map_process_deferred();
vmspace_free(vm2);
return (NULL);
}
new_map->anon_loc = old_map->anon_loc;
new_map->flags |= old_map->flags & (MAP_ASLR | MAP_ASLR_IGNSTART |
MAP_ASLR_STACK | MAP_WXORX);
VM_MAP_ENTRY_FOREACH(old_entry, old_map) {
if ((old_entry->eflags & MAP_ENTRY_IS_SUB_MAP) != 0)
panic("vm_map_fork: encountered a submap");
inh = old_entry->inheritance;
if ((old_entry->eflags & MAP_ENTRY_GUARD) != 0 &&
inh != VM_INHERIT_NONE)
inh = VM_INHERIT_COPY;
switch (inh) {
case VM_INHERIT_NONE:
break;
case VM_INHERIT_SHARE:
/*
* Clone the entry, creating the shared object if
* necessary.
*/
object = old_entry->object.vm_object;
if (object == NULL) {
vm_map_entry_back(old_entry);
object = old_entry->object.vm_object;
}
/*
* Add the reference before calling vm_object_shadow
* to insure that a shadow object is created.
*/
vm_object_reference(object);
if (old_entry->eflags & MAP_ENTRY_NEEDS_COPY) {
vm_object_shadow(&old_entry->object.vm_object,
&old_entry->offset,
old_entry->end - old_entry->start,
old_entry->cred,
/* Transfer the second reference too. */
true);
old_entry->eflags &= ~MAP_ENTRY_NEEDS_COPY;
old_entry->cred = NULL;
/*
* As in vm_map_merged_neighbor_dispose(),
* the vnode lock will not be acquired in
* this call to vm_object_deallocate().
*/
vm_object_deallocate(object);
object = old_entry->object.vm_object;
} else {
VM_OBJECT_WLOCK(object);
vm_object_clear_flag(object, OBJ_ONEMAPPING);
if (old_entry->cred != NULL) {
KASSERT(object->cred == NULL,
("vmspace_fork both cred"));
object->cred = old_entry->cred;
*fork_charge += old_entry->end -
old_entry->start;
old_entry->cred = NULL;
}
/*
* Assert the correct state of the vnode
* v_writecount while the object is locked, to
* not relock it later for the assertion
* correctness.
*/
if (old_entry->eflags & MAP_ENTRY_WRITECNT &&
object->type == OBJT_VNODE) {
KASSERT(((struct vnode *)object->
handle)->v_writecount > 0,
("vmspace_fork: v_writecount %p",
object));
KASSERT(object->un_pager.vnp.
writemappings > 0,
("vmspace_fork: vnp.writecount %p",
object));
}
VM_OBJECT_WUNLOCK(object);
}
/*
* Clone the entry, referencing the shared object.
*/
new_entry = vm_map_entry_create(new_map);
*new_entry = *old_entry;
new_entry->eflags &= ~(MAP_ENTRY_USER_WIRED |
MAP_ENTRY_IN_TRANSITION);
new_entry->wiring_thread = NULL;
new_entry->wired_count = 0;
if (new_entry->eflags & MAP_ENTRY_WRITECNT) {
vm_pager_update_writecount(object,
new_entry->start, new_entry->end);
}
vm_map_entry_set_vnode_text(new_entry, true);
/*
* Insert the entry into the new map -- we know we're
* inserting at the end of the new map.
*/
vm_map_entry_link(new_map, new_entry);
vmspace_map_entry_forked(vm1, vm2, new_entry);
/*
* Update the physical map
*/
pmap_copy(new_map->pmap, old_map->pmap,
new_entry->start,
(old_entry->end - old_entry->start),
old_entry->start);
break;
case VM_INHERIT_COPY:
/*
* Clone the entry and link into the map.
*/
new_entry = vm_map_entry_create(new_map);
*new_entry = *old_entry;
/*
* Copied entry is COW over the old object.
*/
new_entry->eflags &= ~(MAP_ENTRY_USER_WIRED |
MAP_ENTRY_IN_TRANSITION | MAP_ENTRY_WRITECNT);
new_entry->wiring_thread = NULL;
new_entry->wired_count = 0;
new_entry->object.vm_object = NULL;
new_entry->cred = NULL;
vm_map_entry_link(new_map, new_entry);
vmspace_map_entry_forked(vm1, vm2, new_entry);
vm_map_copy_entry(old_map, new_map, old_entry,
new_entry, fork_charge);
vm_map_entry_set_vnode_text(new_entry, true);
break;
case VM_INHERIT_ZERO:
/*
* Create a new anonymous mapping entry modelled from
* the old one.
*/
new_entry = vm_map_entry_create(new_map);
memset(new_entry, 0, sizeof(*new_entry));
new_entry->start = old_entry->start;
new_entry->end = old_entry->end;
new_entry->eflags = old_entry->eflags &
~(MAP_ENTRY_USER_WIRED | MAP_ENTRY_IN_TRANSITION |
MAP_ENTRY_WRITECNT | MAP_ENTRY_VN_EXEC |
MAP_ENTRY_SPLIT_BOUNDARY_MASK);
new_entry->protection = old_entry->protection;
new_entry->max_protection = old_entry->max_protection;
new_entry->inheritance = VM_INHERIT_ZERO;
vm_map_entry_link(new_map, new_entry);
vmspace_map_entry_forked(vm1, vm2, new_entry);
new_entry->cred = curthread->td_ucred;
crhold(new_entry->cred);
*fork_charge += (new_entry->end - new_entry->start);
break;
}
}
/*
* Use inlined vm_map_unlock() to postpone handling the deferred
* map entries, which cannot be done until both old_map and
* new_map locks are released.
*/
sx_xunlock(&old_map->lock);
sx_xunlock(&new_map->lock);
vm_map_process_deferred();
return (vm2);
}
/*
* Create a process's stack for exec_new_vmspace(). This function is never
* asked to wire the newly created stack.
*/
int
vm_map_stack(vm_map_t map, vm_offset_t addrbos, vm_size_t max_ssize,
vm_prot_t prot, vm_prot_t max, int cow)
{
vm_size_t growsize, init_ssize;
rlim_t vmemlim;
int rv;
MPASS((map->flags & MAP_WIREFUTURE) == 0);
growsize = sgrowsiz;
init_ssize = (max_ssize < growsize) ? max_ssize : growsize;
vm_map_lock(map);
vmemlim = lim_cur(curthread, RLIMIT_VMEM);
/* If we would blow our VMEM resource limit, no go */
if (map->size + init_ssize > vmemlim) {
rv = KERN_NO_SPACE;
goto out;
}
rv = vm_map_stack_locked(map, addrbos, max_ssize, growsize, prot,
max, cow);
out:
vm_map_unlock(map);
return (rv);
}
static int stack_guard_page = 1;
SYSCTL_INT(_security_bsd, OID_AUTO, stack_guard_page, CTLFLAG_RWTUN,
&stack_guard_page, 0,
"Specifies the number of guard pages for a stack that grows");
static int
vm_map_stack_locked(vm_map_t map, vm_offset_t addrbos, vm_size_t max_ssize,
vm_size_t growsize, vm_prot_t prot, vm_prot_t max, int cow)
{
vm_map_entry_t gap_entry, new_entry, prev_entry;
vm_offset_t bot, gap_bot, gap_top, top;
vm_size_t init_ssize, sgp;
int rv;
KASSERT((cow & MAP_STACK_AREA) != 0,
("New mapping is not a stack"));
if (max_ssize == 0 ||
!vm_map_range_valid(map, addrbos, addrbos + max_ssize))
return (KERN_INVALID_ADDRESS);
sgp = ((curproc->p_flag2 & P2_STKGAP_DISABLE) != 0 ||
(curproc->p_fctl0 & NT_FREEBSD_FCTL_STKGAP_DISABLE) != 0) ? 0 :
(vm_size_t)stack_guard_page * PAGE_SIZE;
if (sgp >= max_ssize)
return (KERN_INVALID_ARGUMENT);
init_ssize = growsize;
if (max_ssize < init_ssize + sgp)
init_ssize = max_ssize - sgp;
/* If addr is already mapped, no go */
if (vm_map_lookup_entry(map, addrbos, &prev_entry))
return (KERN_NO_SPACE);
/*
* If we can't accommodate max_ssize in the current mapping, no go.
*/
if (vm_map_entry_succ(prev_entry)->start < addrbos + max_ssize)
return (KERN_NO_SPACE);
/*
* We initially map a stack of only init_ssize, at the top of
* the range. We will grow as needed later.
*
* Note: we would normally expect prot and max to be VM_PROT_ALL,
* and cow to be 0. Possibly we should eliminate these as input
* parameters, and just pass these values here in the insert call.
*/
bot = addrbos + max_ssize - init_ssize;
top = bot + init_ssize;
gap_bot = addrbos;
gap_top = bot;
rv = vm_map_insert1(map, NULL, 0, bot, top, prot, max, cow,
&new_entry);
if (rv != KERN_SUCCESS)
return (rv);
KASSERT(new_entry->end == top || new_entry->start == bot,
("Bad entry start/end for new stack entry"));
KASSERT((new_entry->eflags & MAP_ENTRY_GROWS_DOWN) != 0,
("new entry lacks MAP_ENTRY_GROWS_DOWN"));
if (gap_bot == gap_top)
return (KERN_SUCCESS);
rv = vm_map_insert1(map, NULL, 0, gap_bot, gap_top, VM_PROT_NONE,
VM_PROT_NONE, MAP_CREATE_GUARD | MAP_CREATE_STACK_GAP,
&gap_entry);
if (rv == KERN_SUCCESS) {
KASSERT((gap_entry->eflags & MAP_ENTRY_GUARD) != 0,
("entry %p not gap %#x", gap_entry, gap_entry->eflags));
KASSERT((gap_entry->eflags & MAP_ENTRY_STACK_GAP) != 0,
("entry %p not stack gap %#x", gap_entry,
gap_entry->eflags));
/*
* Gap can never successfully handle a fault, so
* read-ahead logic is never used for it. Re-use
* next_read of the gap entry to store
* stack_guard_page for vm_map_growstack().
* Similarly, since a gap cannot have a backing object,
* store the original stack protections in the
* object offset.
*/
gap_entry->next_read = sgp;
gap_entry->offset = prot | PROT_MAX(max);
} else {
(void)vm_map_delete(map, bot, top);
}
return (rv);
}
static bool report_stackoverflow = true;
SYSCTL_BOOL(_vm, OID_AUTO, report_stackoverflow, CTLFLAG_RWTUN,
&report_stackoverflow, 0,
"uprintf() on stack overflow");
/*
* Attempts to grow a vm stack entry. Returns KERN_SUCCESS if we
* successfully grow the stack.
*/
static int
vm_map_growstack(vm_map_t map, vm_offset_t addr, vm_map_entry_t gap_entry)
{
vm_map_entry_t stack_entry;
struct thread *td;
struct proc *p;
struct vmspace *vm;
vm_offset_t gap_end, gap_start, grow_start;
vm_size_t grow_amount, guard, max_grow, sgp;
vm_prot_t prot, max;
rlim_t lmemlim, stacklim, vmemlim;
int rv, rv1 __diagused;
bool gap_deleted, is_procstack;
#ifdef notyet
uint64_t limit;
#endif
#ifdef RACCT
int error __diagused;
#endif
td = curthread;
p = td->td_proc;
vm = p->p_vmspace;
/*
* Disallow stack growth when the access is performed by a
* debugger or AIO daemon. The reason is that the wrong
* resource limits are applied.
*/
if (p != initproc && (map != &vm->vm_map || p->p_textvp == NULL))
return (KERN_FAILURE);
MPASS(!vm_map_is_system(map));
lmemlim = lim_cur(td, RLIMIT_MEMLOCK);
stacklim = lim_cur(td, RLIMIT_STACK);
vmemlim = lim_cur(td, RLIMIT_VMEM);
retry:
/* If addr is not in a hole for a stack grow area, no need to grow. */
if (gap_entry == NULL && !vm_map_lookup_entry(map, addr, &gap_entry))
return (KERN_FAILURE);
if ((gap_entry->eflags & MAP_ENTRY_GUARD) == 0)
return (KERN_SUCCESS);
if ((gap_entry->eflags & MAP_ENTRY_STACK_GAP) != 0) {
stack_entry = vm_map_entry_succ(gap_entry);
if ((stack_entry->eflags & MAP_ENTRY_GROWS_DOWN) == 0 ||
stack_entry->start != gap_entry->end)
return (KERN_FAILURE);
grow_amount = round_page(stack_entry->start - addr);
} else {
return (KERN_FAILURE);
}
guard = ((p->p_flag2 & P2_STKGAP_DISABLE) != 0 ||
(p->p_fctl0 & NT_FREEBSD_FCTL_STKGAP_DISABLE) != 0) ? 0 :
gap_entry->next_read;
max_grow = gap_entry->end - gap_entry->start;
if (guard > max_grow)
return (KERN_NO_SPACE);
max_grow -= guard;
if (grow_amount > max_grow) {
if (report_stackoverflow)
uprintf("pid %d comm %s tid %d stack overflow\n",
p->p_pid, p->p_comm, td->td_tid);
return (KERN_NO_SPACE);
}
/*
* If this is the main process stack, see if we're over the stack
* limit.
*/
is_procstack = addr >= (vm_offset_t)vm->vm_maxsaddr &&
addr < (vm_offset_t)vm->vm_stacktop;
if (is_procstack && (ctob(vm->vm_ssize) + grow_amount > stacklim)) {
if (report_stackoverflow)
uprintf("pid %d comm %s tid %d stack overflow\n",
p->p_pid, p->p_comm, td->td_tid);
return (KERN_NO_SPACE);
}
#ifdef RACCT
if (racct_enable) {
PROC_LOCK(p);
if (is_procstack && racct_set(p, RACCT_STACK,
ctob(vm->vm_ssize) + grow_amount)) {
PROC_UNLOCK(p);
return (KERN_NO_SPACE);
}
PROC_UNLOCK(p);
}
#endif
grow_amount = roundup(grow_amount, sgrowsiz);
if (grow_amount > max_grow)
grow_amount = max_grow;
if (is_procstack && (ctob(vm->vm_ssize) + grow_amount > stacklim)) {
grow_amount = trunc_page((vm_size_t)stacklim) -
ctob(vm->vm_ssize);
}
#ifdef notyet
PROC_LOCK(p);
limit = racct_get_available(p, RACCT_STACK);
PROC_UNLOCK(p);
if (is_procstack && (ctob(vm->vm_ssize) + grow_amount > limit))
grow_amount = limit - ctob(vm->vm_ssize);
#endif
if (!old_mlock && (map->flags & MAP_WIREFUTURE) != 0) {
if (ptoa(pmap_wired_count(map->pmap)) + grow_amount > lmemlim) {
rv = KERN_NO_SPACE;
goto out;
}
#ifdef RACCT
if (racct_enable) {
PROC_LOCK(p);
if (racct_set(p, RACCT_MEMLOCK,
ptoa(pmap_wired_count(map->pmap)) + grow_amount)) {
PROC_UNLOCK(p);
rv = KERN_NO_SPACE;
goto out;
}
PROC_UNLOCK(p);
}
#endif
}
/* If we would blow our VMEM resource limit, no go */
if (map->size + grow_amount > vmemlim) {
rv = KERN_NO_SPACE;
goto out;
}
#ifdef RACCT
if (racct_enable) {
PROC_LOCK(p);
if (racct_set(p, RACCT_VMEM, map->size + grow_amount)) {
PROC_UNLOCK(p);
rv = KERN_NO_SPACE;
goto out;
}
PROC_UNLOCK(p);
}
#endif
if (vm_map_lock_upgrade(map)) {
gap_entry = NULL;
vm_map_lock_read(map);
goto retry;
}
/*
* The gap_entry "offset" field is overloaded. See
* vm_map_stack_locked().
*/
prot = PROT_EXTRACT(gap_entry->offset);
max = PROT_MAX_EXTRACT(gap_entry->offset);
sgp = gap_entry->next_read;
grow_start = gap_entry->end - grow_amount;
if (gap_entry->start + grow_amount == gap_entry->end) {
gap_start = gap_entry->start;
gap_end = gap_entry->end;
vm_map_entry_delete(map, gap_entry);
gap_deleted = true;
} else {
MPASS(gap_entry->start < gap_entry->end - grow_amount);
vm_map_entry_resize(map, gap_entry, -grow_amount);
gap_deleted = false;
}
rv = vm_map_insert(map, NULL, 0, grow_start,
grow_start + grow_amount, prot, max, MAP_STACK_AREA);
if (rv != KERN_SUCCESS) {
if (gap_deleted) {
rv1 = vm_map_insert1(map, NULL, 0, gap_start,
gap_end, VM_PROT_NONE, VM_PROT_NONE,
MAP_CREATE_GUARD | MAP_CREATE_STACK_GAP,
&gap_entry);
MPASS(rv1 == KERN_SUCCESS);
gap_entry->next_read = sgp;
gap_entry->offset = prot | PROT_MAX(max);
} else {
vm_map_entry_resize(map, gap_entry,
grow_amount);
}
}
if (rv == KERN_SUCCESS && is_procstack)
vm->vm_ssize += btoc(grow_amount);
/*
* Heed the MAP_WIREFUTURE flag if it was set for this process.
*/
if (rv == KERN_SUCCESS && (map->flags & MAP_WIREFUTURE) != 0) {
rv = vm_map_wire_locked(map, grow_start,
grow_start + grow_amount,
VM_MAP_WIRE_USER | VM_MAP_WIRE_NOHOLES);
}
vm_map_lock_downgrade(map);
out:
#ifdef RACCT
if (racct_enable && rv != KERN_SUCCESS) {
PROC_LOCK(p);
error = racct_set(p, RACCT_VMEM, map->size);
KASSERT(error == 0, ("decreasing RACCT_VMEM failed"));
if (!old_mlock) {
error = racct_set(p, RACCT_MEMLOCK,
ptoa(pmap_wired_count(map->pmap)));
KASSERT(error == 0, ("decreasing RACCT_MEMLOCK failed"));
}
error = racct_set(p, RACCT_STACK, ctob(vm->vm_ssize));
KASSERT(error == 0, ("decreasing RACCT_STACK failed"));
PROC_UNLOCK(p);
}
#endif
return (rv);
}
/*
* Unshare the specified VM space for exec. If other processes are
* mapped to it, then create a new one. The new vmspace is null.
*/
int
vmspace_exec(struct proc *p, vm_offset_t minuser, vm_offset_t maxuser)
{
struct vmspace *oldvmspace = p->p_vmspace;
struct vmspace *newvmspace;
KASSERT((curthread->td_pflags & TDP_EXECVMSPC) == 0,
("vmspace_exec recursed"));
newvmspace = vmspace_alloc(minuser, maxuser, pmap_pinit);
if (newvmspace == NULL)
return (ENOMEM);
newvmspace->vm_swrss = oldvmspace->vm_swrss;
/*
* This code is written like this for prototype purposes. The
* goal is to avoid running down the vmspace here, but let the
* other process's that are still using the vmspace to finally
* run it down. Even though there is little or no chance of blocking
* here, it is a good idea to keep this form for future mods.
*/
PROC_VMSPACE_LOCK(p);
p->p_vmspace = newvmspace;
PROC_VMSPACE_UNLOCK(p);
if (p == curthread->td_proc)
pmap_activate(curthread);
curthread->td_pflags |= TDP_EXECVMSPC;
return (0);
}
/*
* Unshare the specified VM space for forcing COW. This
* is called by rfork, for the (RFMEM|RFPROC) == 0 case.
*/
int
vmspace_unshare(struct proc *p)
{
struct vmspace *oldvmspace = p->p_vmspace;
struct vmspace *newvmspace;
vm_ooffset_t fork_charge;
/*
* The caller is responsible for ensuring that the reference count
* cannot concurrently transition 1 -> 2.
*/
if (refcount_load(&oldvmspace->vm_refcnt) == 1)
return (0);
fork_charge = 0;
newvmspace = vmspace_fork(oldvmspace, &fork_charge);
if (newvmspace == NULL)
return (ENOMEM);
if (!swap_reserve_by_cred(fork_charge, p->p_ucred)) {
/*
* The swap reservation failed. The accounting from
* the entries of the copied newvmspace will be
* subtracted in vmspace_free(), so force the
* reservation there.
*/
swap_reserve_force_by_cred(fork_charge, p->p_ucred);
vmspace_free(newvmspace);
return (ENOMEM);
}
PROC_VMSPACE_LOCK(p);
p->p_vmspace = newvmspace;
PROC_VMSPACE_UNLOCK(p);
if (p == curthread->td_proc)
pmap_activate(curthread);
vmspace_free(oldvmspace);
return (0);
}
/*
* vm_map_lookup:
*
* Finds the VM object, offset, and
* protection for a given virtual address in the
* specified map, assuming a page fault of the
* type specified.
*
* Leaves the map in question locked for read; return
* values are guaranteed until a vm_map_lookup_done
* call is performed. Note that the map argument
* is in/out; the returned map must be used in
* the call to vm_map_lookup_done.
*
* A handle (out_entry) is returned for use in
* vm_map_lookup_done, to make that fast.
*
* If a lookup is requested with "write protection"
* specified, the map may be changed to perform virtual
* copying operations, although the data referenced will
* remain the same.
*/
int
vm_map_lookup(vm_map_t *var_map, /* IN/OUT */
vm_offset_t vaddr,
vm_prot_t fault_typea,
vm_map_entry_t *out_entry, /* OUT */
vm_object_t *object, /* OUT */
vm_pindex_t *pindex, /* OUT */
vm_prot_t *out_prot, /* OUT */
boolean_t *wired) /* OUT */
{
vm_map_entry_t entry;
vm_map_t map = *var_map;
vm_prot_t prot;
vm_prot_t fault_type;
vm_object_t eobject;
vm_size_t size;
struct ucred *cred;
RetryLookup:
vm_map_lock_read(map);
RetryLookupLocked:
/*
* Lookup the faulting address.
*/
if (!vm_map_lookup_entry(map, vaddr, out_entry)) {
vm_map_unlock_read(map);
return (KERN_INVALID_ADDRESS);
}
entry = *out_entry;
/*
* Handle submaps.
*/
if (entry->eflags & MAP_ENTRY_IS_SUB_MAP) {
vm_map_t old_map = map;
*var_map = map = entry->object.sub_map;
vm_map_unlock_read(old_map);
goto RetryLookup;
}
/*
* Check whether this task is allowed to have this page.
*/
prot = entry->protection;
if ((fault_typea & VM_PROT_FAULT_LOOKUP) != 0) {
fault_typea &= ~VM_PROT_FAULT_LOOKUP;
if (prot == VM_PROT_NONE && map != kernel_map &&
(entry->eflags & MAP_ENTRY_GUARD) != 0 &&
(entry->eflags & MAP_ENTRY_STACK_GAP) != 0 &&
vm_map_growstack(map, vaddr, entry) == KERN_SUCCESS)
goto RetryLookupLocked;
}
fault_type = fault_typea & VM_PROT_ALL;
if ((fault_type & prot) != fault_type || prot == VM_PROT_NONE) {
vm_map_unlock_read(map);
return (KERN_PROTECTION_FAILURE);
}
KASSERT((prot & VM_PROT_WRITE) == 0 || (entry->eflags &
(MAP_ENTRY_USER_WIRED | MAP_ENTRY_NEEDS_COPY)) !=
(MAP_ENTRY_USER_WIRED | MAP_ENTRY_NEEDS_COPY),
("entry %p flags %x", entry, entry->eflags));
if ((fault_typea & VM_PROT_COPY) != 0 &&
(entry->max_protection & VM_PROT_WRITE) == 0 &&
(entry->eflags & MAP_ENTRY_COW) == 0) {
vm_map_unlock_read(map);
return (KERN_PROTECTION_FAILURE);
}
/*
* If this page is not pageable, we have to get it for all possible
* accesses.
*/
*wired = (entry->wired_count != 0);
if (*wired)
fault_type = entry->protection;
size = entry->end - entry->start;
/*
* If the entry was copy-on-write, we either ...
*/
if (entry->eflags & MAP_ENTRY_NEEDS_COPY) {
/*
* If we want to write the page, we may as well handle that
* now since we've got the map locked.
*
* If we don't need to write the page, we just demote the
* permissions allowed.
*/
if ((fault_type & VM_PROT_WRITE) != 0 ||
(fault_typea & VM_PROT_COPY) != 0) {
/*
* Make a new object, and place it in the object
* chain. Note that no new references have appeared
* -- one just moved from the map to the new
* object.
*/
if (vm_map_lock_upgrade(map))
goto RetryLookup;
if (entry->cred == NULL) {
/*
* The debugger owner is charged for
* the memory.
*/
cred = curthread->td_ucred;
crhold(cred);
if (!swap_reserve_by_cred(size, cred)) {
crfree(cred);
vm_map_unlock(map);
return (KERN_RESOURCE_SHORTAGE);
}
entry->cred = cred;
}
eobject = entry->object.vm_object;
vm_object_shadow(&entry->object.vm_object,
&entry->offset, size, entry->cred, false);
if (eobject == entry->object.vm_object) {
/*
* The object was not shadowed.
*/
swap_release_by_cred(size, entry->cred);
crfree(entry->cred);
}
entry->cred = NULL;
entry->eflags &= ~MAP_ENTRY_NEEDS_COPY;
vm_map_lock_downgrade(map);
} else {
/*
* We're attempting to read a copy-on-write page --
* don't allow writes.
*/
prot &= ~VM_PROT_WRITE;
}
}
/*
* Create an object if necessary.
*/
if (entry->object.vm_object == NULL && !vm_map_is_system(map)) {
if (vm_map_lock_upgrade(map))
goto RetryLookup;
entry->object.vm_object = vm_object_allocate_anon(atop(size),
NULL, entry->cred);
entry->offset = 0;
entry->cred = NULL;
vm_map_lock_downgrade(map);
}
/*
* Return the object/offset from this entry. If the entry was
* copy-on-write or empty, it has been fixed up.
*/
*pindex = OFF_TO_IDX((vaddr - entry->start) + entry->offset);
*object = entry->object.vm_object;
*out_prot = prot;
return (KERN_SUCCESS);
}
/*
* vm_map_lookup_locked:
*
* Lookup the faulting address. A version of vm_map_lookup that returns
* KERN_FAILURE instead of blocking on map lock or memory allocation.
*/
int
vm_map_lookup_locked(vm_map_t *var_map, /* IN/OUT */
vm_offset_t vaddr,
vm_prot_t fault_typea,
vm_map_entry_t *out_entry, /* OUT */
vm_object_t *object, /* OUT */
vm_pindex_t *pindex, /* OUT */
vm_prot_t *out_prot, /* OUT */
boolean_t *wired) /* OUT */
{
vm_map_entry_t entry;
vm_map_t map = *var_map;
vm_prot_t prot;
vm_prot_t fault_type = fault_typea;
/*
* Lookup the faulting address.
*/
if (!vm_map_lookup_entry(map, vaddr, out_entry))
return (KERN_INVALID_ADDRESS);
entry = *out_entry;
/*
* Fail if the entry refers to a submap.
*/
if (entry->eflags & MAP_ENTRY_IS_SUB_MAP)
return (KERN_FAILURE);
/*
* Check whether this task is allowed to have this page.
*/
prot = entry->protection;
fault_type &= VM_PROT_READ | VM_PROT_WRITE | VM_PROT_EXECUTE;
if ((fault_type & prot) != fault_type)
return (KERN_PROTECTION_FAILURE);
/*
* If this page is not pageable, we have to get it for all possible
* accesses.
*/
*wired = (entry->wired_count != 0);
if (*wired)
fault_type = entry->protection;
if (entry->eflags & MAP_ENTRY_NEEDS_COPY) {
/*
* Fail if the entry was copy-on-write for a write fault.
*/
if (fault_type & VM_PROT_WRITE)
return (KERN_FAILURE);
/*
* We're attempting to read a copy-on-write page --
* don't allow writes.
*/
prot &= ~VM_PROT_WRITE;
}
/*
* Fail if an object should be created.
*/
if (entry->object.vm_object == NULL && !vm_map_is_system(map))
return (KERN_FAILURE);
/*
* Return the object/offset from this entry. If the entry was
* copy-on-write or empty, it has been fixed up.
*/
*pindex = OFF_TO_IDX((vaddr - entry->start) + entry->offset);
*object = entry->object.vm_object;
*out_prot = prot;
return (KERN_SUCCESS);
}
/*
* vm_map_lookup_done:
*
* Releases locks acquired by a vm_map_lookup
* (according to the handle returned by that lookup).
*/
void
vm_map_lookup_done(vm_map_t map, vm_map_entry_t entry)
{
/*
* Unlock the main-level map
*/
vm_map_unlock_read(map);
}
vm_offset_t
vm_map_max_KBI(const struct vm_map *map)
{
return (vm_map_max(map));
}
vm_offset_t
vm_map_min_KBI(const struct vm_map *map)
{
return (vm_map_min(map));
}
pmap_t
vm_map_pmap_KBI(vm_map_t map)
{
return (map->pmap);
}
bool
vm_map_range_valid_KBI(vm_map_t map, vm_offset_t start, vm_offset_t end)
{
return (vm_map_range_valid(map, start, end));
}
#ifdef INVARIANTS
static void
_vm_map_assert_consistent(vm_map_t map, int check)
{
vm_map_entry_t entry, prev;
vm_map_entry_t cur, header, lbound, ubound;
vm_size_t max_left, max_right;
#ifdef DIAGNOSTIC
++map->nupdates;
#endif
if (enable_vmmap_check != check)
return;
header = prev = &map->header;
VM_MAP_ENTRY_FOREACH(entry, map) {
KASSERT(prev->end <= entry->start,
("map %p prev->end = %jx, start = %jx", map,
(uintmax_t)prev->end, (uintmax_t)entry->start));
KASSERT(entry->start < entry->end,
("map %p start = %jx, end = %jx", map,
(uintmax_t)entry->start, (uintmax_t)entry->end));
KASSERT(entry->left == header ||
entry->left->start < entry->start,
("map %p left->start = %jx, start = %jx", map,
(uintmax_t)entry->left->start, (uintmax_t)entry->start));
KASSERT(entry->right == header ||
entry->start < entry->right->start,
("map %p start = %jx, right->start = %jx", map,
(uintmax_t)entry->start, (uintmax_t)entry->right->start));
cur = map->root;
lbound = ubound = header;
for (;;) {
if (entry->start < cur->start) {
ubound = cur;
cur = cur->left;
KASSERT(cur != lbound,
("map %p cannot find %jx",
map, (uintmax_t)entry->start));
} else if (cur->end <= entry->start) {
lbound = cur;
cur = cur->right;
KASSERT(cur != ubound,
("map %p cannot find %jx",
map, (uintmax_t)entry->start));
} else {
KASSERT(cur == entry,
("map %p cannot find %jx",
map, (uintmax_t)entry->start));
break;
}
}
max_left = vm_map_entry_max_free_left(entry, lbound);
max_right = vm_map_entry_max_free_right(entry, ubound);
KASSERT(entry->max_free == vm_size_max(max_left, max_right),
("map %p max = %jx, max_left = %jx, max_right = %jx", map,
(uintmax_t)entry->max_free,
(uintmax_t)max_left, (uintmax_t)max_right));
prev = entry;
}
KASSERT(prev->end <= entry->start,
("map %p prev->end = %jx, start = %jx", map,
(uintmax_t)prev->end, (uintmax_t)entry->start));
}
#endif
#include "opt_ddb.h"
#ifdef DDB
#include <sys/kernel.h>
#include <ddb/ddb.h>
static void
vm_map_print(vm_map_t map)
{
vm_map_entry_t entry, prev;
db_iprintf("Task map %p: pmap=%p, nentries=%d, version=%u\n",
(void *)map,
(void *)map->pmap, map->nentries, map->timestamp);
db_indent += 2;
prev = &map->header;
VM_MAP_ENTRY_FOREACH(entry, map) {
db_iprintf("map entry %p: start=%p, end=%p, eflags=%#x, \n",
(void *)entry, (void *)entry->start, (void *)entry->end,
entry->eflags);
{
static const char * const inheritance_name[4] =
{"share", "copy", "none", "donate_copy"};
db_iprintf(" prot=%x/%x/%s",
entry->protection,
entry->max_protection,
inheritance_name[(int)(unsigned char)
entry->inheritance]);
if (entry->wired_count != 0)
db_printf(", wired");
}
if (entry->eflags & MAP_ENTRY_IS_SUB_MAP) {
db_printf(", share=%p, offset=0x%jx\n",
(void *)entry->object.sub_map,
(uintmax_t)entry->offset);
if (prev == &map->header ||
prev->object.sub_map !=
entry->object.sub_map) {
db_indent += 2;
vm_map_print((vm_map_t)entry->object.sub_map);
db_indent -= 2;
}
} else {
if (entry->cred != NULL)
db_printf(", ruid %d", entry->cred->cr_ruid);
db_printf(", object=%p, offset=0x%jx",
(void *)entry->object.vm_object,
(uintmax_t)entry->offset);
if (entry->object.vm_object && entry->object.vm_object->cred)
db_printf(", obj ruid %d ",
entry->object.vm_object->cred->cr_ruid);
if (entry->eflags & MAP_ENTRY_COW)
db_printf(", copy (%s)",
(entry->eflags & MAP_ENTRY_NEEDS_COPY) ? "needed" : "done");
db_printf("\n");
if (prev == &map->header ||
prev->object.vm_object !=
entry->object.vm_object) {
db_indent += 2;
vm_object_print((db_expr_t)(intptr_t)
entry->object.vm_object,
0, 0, (char *)0);
db_indent -= 2;
}
}
prev = entry;
}
db_indent -= 2;
}
DB_SHOW_COMMAND(map, map)
{
if (!have_addr) {
db_printf("usage: show map <addr>\n");
return;
}
vm_map_print((vm_map_t)addr);
}
DB_SHOW_COMMAND(procvm, procvm)
{
struct proc *p;
if (have_addr) {
p = db_lookup_proc(addr);
} else {
p = curproc;
}
db_printf("p = %p, vmspace = %p, map = %p, pmap = %p\n",
(void *)p, (void *)p->p_vmspace, (void *)&p->p_vmspace->vm_map,
(void *)vmspace_pmap(p->p_vmspace));
vm_map_print((vm_map_t)&p->p_vmspace->vm_map);
}
#endif /* DDB */
diff --git a/sys/vm/vm_object.h b/sys/vm/vm_object.h
index 5a3a78b5b5a1d06e7dc57fb4742cbbe6102641b5..a4df62d68251cc55752ca723484042f68c0f5ed2 100644
--- a/sys/vm/vm_object.h
+++ b/sys/vm/vm_object.h
@@ -1,396 +1,397 @@
/*-
* SPDX-License-Identifier: (BSD-3-Clause AND MIT-CMU)
*
* Copyright (c) 1991, 1993
* The Regents of the University of California. All rights reserved.
*
* This code is derived from software contributed to Berkeley by
* The Mach Operating System project at Carnegie-Mellon University.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice, this list of conditions and the following disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
* 3. Neither the name of the University nor the names of its contributors
* may be used to endorse or promote products derived from this software
* without specific prior written permission.
*
* THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS ``AS IS'' AND
* ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
* ARE DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS
* OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)
* HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT
* LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY
* OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
* SUCH DAMAGE.
*
*
* Copyright (c) 1987, 1990 Carnegie-Mellon University.
* All rights reserved.
*
* Authors: Avadis Tevanian, Jr., Michael Wayne Young
*
* Permission to use, copy, modify and distribute this software and
* its documentation is hereby granted, provided that both the copyright
* notice and this permission notice appear in all copies of the
* software, derivative works or modified versions, and any portions
* thereof, and that both notices appear in supporting documentation.
*
* CARNEGIE MELLON ALLOWS FREE USE OF THIS SOFTWARE IN ITS "AS IS"
* CONDITION. CARNEGIE MELLON DISCLAIMS ANY LIABILITY OF ANY KIND
* FOR ANY DAMAGES WHATSOEVER RESULTING FROM THE USE OF THIS SOFTWARE.
*
* Carnegie Mellon requests users of this software to return to
*
* Software Distribution Coordinator or Software.Distribution@CS.CMU.EDU
* School of Computer Science
* Carnegie Mellon University
* Pittsburgh PA 15213-3890
*
* any improvements or extensions that they make and grant Carnegie the
* rights to redistribute these changes.
*/
/*
* Virtual memory object module definitions.
*/
#ifndef _VM_OBJECT_
#define _VM_OBJECT_
#include <sys/queue.h>
#include <sys/_blockcount.h>
#include <sys/_lock.h>
#include <sys/_mutex.h>
#include <sys/_pctrie.h>
#include <sys/_rwlock.h>
#include <sys/_domainset.h>
#include <vm/_vm_radix.h>
/*
* Types defined:
*
* vm_object_t Virtual memory object.
*
* List of locks
* (a) atomic
* (c) const until freed
* (o) per-object lock
* (f) free pages queue mutex
*
*/
#ifndef VM_PAGE_HAVE_PGLIST
TAILQ_HEAD(pglist, vm_page);
#define VM_PAGE_HAVE_PGLIST
#endif
struct vm_object {
struct rwlock lock;
TAILQ_ENTRY(vm_object) object_list; /* list of all objects */
LIST_HEAD(, vm_object) shadow_head; /* objects that this is a shadow for */
LIST_ENTRY(vm_object) shadow_list; /* chain of shadow objects */
struct vm_radix rtree; /* root of the resident page radix trie*/
vm_pindex_t size; /* Object size */
struct domainset_ref domain; /* NUMA policy. */
volatile int generation; /* generation ID */
int cleangeneration; /* Generation at clean time */
volatile u_int ref_count; /* How many refs?? */
int shadow_count; /* how many objects that this is a shadow for */
vm_memattr_t memattr; /* default memory attribute for pages */
objtype_t type; /* type of pager */
u_short pg_color; /* (c) color of first page in obj */
u_int flags; /* see below */
blockcount_t paging_in_progress; /* (a) Paging (in or out) so don't collapse or destroy */
blockcount_t busy; /* (a) object is busy, disallow page busy. */
int resident_page_count; /* number of resident pages */
struct vm_object *backing_object; /* object that I'm a shadow of */
vm_ooffset_t backing_object_offset;/* Offset in backing object */
TAILQ_ENTRY(vm_object) pager_object_list; /* list of all objects of this pager type */
LIST_HEAD(, vm_reserv) rvq; /* list of reservations */
void *handle;
union {
/*
* VNode pager
*
* vnp_size - current size of file
*/
struct {
off_t vnp_size;
vm_ooffset_t writemappings;
} vnp;
/*
* Device pager
*
* devp_pglist - list of allocated pages
*/
struct {
const struct cdev_pager_ops *ops;
void *handle;
} devp;
/*
* SG pager
*
* sgp_pglist - list of allocated pages
*/
struct {
TAILQ_HEAD(, vm_page) sgp_pglist;
} sgp;
/*
* Swap pager
*
* swp_priv - pager-private.
* swp_blks - pc-trie of the allocated swap blocks.
* writemappings - count of bytes mapped for write
*
*/
struct {
void *swp_priv;
struct pctrie swp_blks;
vm_ooffset_t writemappings;
} swp;
/*
* Phys pager
*/
struct {
const struct phys_pager_ops *ops;
union {
void *data_ptr;
uintptr_t data_val;
};
void *phys_priv;
} phys;
} un_pager;
struct ucred *cred;
void *umtx_data;
};
/*
* Flags
*/
#define OBJ_FICTITIOUS 0x00000001 /* (c) contains fictitious pages */
#define OBJ_UNMANAGED 0x00000002 /* (c) contains unmanaged pages */
#define OBJ_POPULATE 0x00000004 /* pager implements populate() */
#define OBJ_DEAD 0x00000008 /* dead objects (during rundown) */
#define OBJ_ANON 0x00000010 /* (c) contains anonymous memory */
#define OBJ_UMTXDEAD 0x00000020 /* umtx pshared was terminated */
#define OBJ_SIZEVNLOCK 0x00000040 /* lock vnode to check obj size */
#define OBJ_PG_DTOR 0x00000080 /* do not reset object, leave that
for dtor */
#define OBJ_SHADOWLIST 0x00000100 /* Object is on the shadow list. */
#define OBJ_SWAP 0x00000200 /* object swaps, type will be OBJT_SWAP
or dynamically registered */
#define OBJ_SPLIT 0x00000400 /* object is being split */
#define OBJ_COLLAPSING 0x00000800 /* Parent of collapse. */
#define OBJ_COLORED 0x00001000 /* pg_color is defined */
#define OBJ_ONEMAPPING 0x00002000 /* Each page has at most one managed
mapping, all in the same vm_map */
#define OBJ_PAGERPRIV1 0x00004000 /* Pager private */
#define OBJ_PAGERPRIV2 0x00008000 /* Pager private */
#define OBJ_SYSVSHM 0x00010000 /* SysV SHM */
#define OBJ_POSIXSHM 0x00020000 /* Posix SHM */
+#define OBJ_NOMLOCK 0x00040000 /* (c) skip userspace memory locking */
/*
* Helpers to perform conversion between vm_object page indexes and offsets.
* IDX_TO_OFF() converts an index into an offset.
* OFF_TO_IDX() converts an offset into an index.
* OBJ_MAX_SIZE specifies the maximum page index corresponding to the
* maximum unsigned offset.
*/
#define IDX_TO_OFF(idx) (((vm_ooffset_t)(idx)) << PAGE_SHIFT)
#define OFF_TO_IDX(off) ((vm_pindex_t)(((vm_ooffset_t)(off)) >> PAGE_SHIFT))
#define OBJ_MAX_SIZE (OFF_TO_IDX(UINT64_MAX) + 1)
#ifdef _KERNEL
#define OBJPC_SYNC 0x1 /* sync I/O */
#define OBJPC_INVAL 0x2 /* invalidate */
#define OBJPC_NOSYNC 0x4 /* skip if PGA_NOSYNC */
/*
* The following options are supported by vm_object_page_remove().
*/
#define OBJPR_CLEANONLY 0x1 /* Don't remove dirty pages. */
#define OBJPR_NOTMAPPED 0x2 /* Don't unmap pages. */
#define OBJPR_VALIDONLY 0x4 /* Ignore invalid pages. */
/*
* Options for vm_object_coalesce().
*/
#define OBJCO_CHARGED 0x1 /* The next_size was charged already */
#define OBJCO_NO_CHARGE 0x2 /* Do not do swap accounting at all */
TAILQ_HEAD(object_q, vm_object);
extern struct object_q vm_object_list; /* list of allocated objects */
extern struct mtx vm_object_list_mtx; /* lock for object list and count */
extern struct vm_object kernel_object_store;
#define kernel_object (&kernel_object_store)
#define VM_OBJECT_ASSERT_LOCKED(object) \
rw_assert(&(object)->lock, RA_LOCKED)
#define VM_OBJECT_ASSERT_RLOCKED(object) \
rw_assert(&(object)->lock, RA_RLOCKED)
#define VM_OBJECT_ASSERT_WLOCKED(object) \
rw_assert(&(object)->lock, RA_WLOCKED)
#define VM_OBJECT_ASSERT_UNLOCKED(object) \
rw_assert(&(object)->lock, RA_UNLOCKED)
#define VM_OBJECT_LOCK_DOWNGRADE(object) \
rw_downgrade(&(object)->lock)
#define VM_OBJECT_RLOCK(object) \
rw_rlock(&(object)->lock)
#define VM_OBJECT_RUNLOCK(object) \
rw_runlock(&(object)->lock)
#define VM_OBJECT_SLEEP(object, wchan, pri, wmesg, timo) \
rw_sleep((wchan), &(object)->lock, (pri), (wmesg), (timo))
#define VM_OBJECT_TRYRLOCK(object) \
rw_try_rlock(&(object)->lock)
#define VM_OBJECT_TRYWLOCK(object) \
rw_try_wlock(&(object)->lock)
#define VM_OBJECT_TRYUPGRADE(object) \
rw_try_upgrade(&(object)->lock)
#define VM_OBJECT_WLOCK(object) \
rw_wlock(&(object)->lock)
#define VM_OBJECT_WOWNED(object) \
rw_wowned(&(object)->lock)
#define VM_OBJECT_WUNLOCK(object) \
rw_wunlock(&(object)->lock)
#define VM_OBJECT_UNLOCK(object) \
rw_unlock(&(object)->lock)
#define VM_OBJECT_DROP(object) \
lock_class_rw.lc_unlock(&(object)->lock.lock_object)
#define VM_OBJECT_PICKUP(object, state) \
lock_class_rw.lc_lock(&(object)->lock.lock_object, (state))
#define VM_OBJECT_ASSERT_PAGING(object) \
KASSERT(blockcount_read(&(object)->paging_in_progress) != 0, \
("vm_object %p is not paging", object))
#define VM_OBJECT_ASSERT_REFERENCE(object) \
KASSERT((object)->reference_count != 0, \
("vm_object %p is not referenced", object))
struct vnode;
/*
* The object must be locked or thread private.
*/
static __inline void
vm_object_set_flag(vm_object_t object, u_int bits)
{
object->flags |= bits;
}
/*
* Conditionally set the object's color, which (1) enables the allocation
* of physical memory reservations for anonymous objects and larger-than-
* superpage-sized named objects and (2) determines the first page offset
* within the object at which a reservation may be allocated. In other
* words, the color determines the alignment of the object with respect
* to the largest superpage boundary. When mapping named objects, like
* files or POSIX shared memory objects, the color should be set to zero
* before a virtual address is selected for the mapping. In contrast,
* for anonymous objects, the color may be set after the virtual address
* is selected.
*
* The object must be locked.
*/
static __inline void
vm_object_color(vm_object_t object, u_short color)
{
if ((object->flags & OBJ_COLORED) == 0) {
object->pg_color = color;
vm_object_set_flag(object, OBJ_COLORED);
}
}
static __inline bool
vm_object_reserv(vm_object_t object)
{
if (object != NULL &&
(object->flags & (OBJ_COLORED | OBJ_FICTITIOUS)) == OBJ_COLORED) {
return (true);
}
return (false);
}
void vm_object_clear_flag(vm_object_t object, u_short bits);
void vm_object_pip_add(vm_object_t object, short i);
void vm_object_pip_wakeup(vm_object_t object);
void vm_object_pip_wakeupn(vm_object_t object, short i);
void vm_object_pip_wait(vm_object_t object, const char *waitid);
void vm_object_pip_wait_unlocked(vm_object_t object, const char *waitid);
void vm_object_busy(vm_object_t object);
void vm_object_unbusy(vm_object_t object);
void vm_object_busy_wait(vm_object_t object, const char *wmesg);
static inline bool
vm_object_busied(vm_object_t object)
{
return (blockcount_read(&object->busy) != 0);
}
#define VM_OBJECT_ASSERT_BUSY(object) MPASS(vm_object_busied((object)))
void umtx_shm_object_init(vm_object_t object);
void umtx_shm_object_terminated(vm_object_t object);
extern int umtx_shm_vnobj_persistent;
vm_object_t vm_object_allocate (objtype_t, vm_pindex_t);
vm_object_t vm_object_allocate_anon(vm_pindex_t, vm_object_t, struct ucred *);
vm_object_t vm_object_allocate_dyn(objtype_t, vm_pindex_t, u_short);
boolean_t vm_object_coalesce(vm_object_t, vm_ooffset_t, vm_size_t, vm_size_t,
int);
void vm_object_collapse (vm_object_t);
void vm_object_deallocate (vm_object_t);
void vm_object_destroy (vm_object_t);
void vm_object_terminate (vm_object_t);
void vm_object_set_writeable_dirty (vm_object_t);
void vm_object_set_writeable_dirty_(vm_object_t object);
bool vm_object_mightbedirty(vm_object_t object);
bool vm_object_mightbedirty_(vm_object_t object);
void vm_object_init (void);
int vm_object_kvme_type(vm_object_t object, struct vnode **vpp);
void vm_object_madvise(vm_object_t, vm_pindex_t, vm_pindex_t, int);
boolean_t vm_object_page_clean(vm_object_t object, vm_ooffset_t start,
vm_ooffset_t end, int flags);
void vm_object_page_noreuse(vm_object_t object, vm_pindex_t start,
vm_pindex_t end);
void vm_object_page_remove(vm_object_t object, vm_pindex_t start,
vm_pindex_t end, int options);
boolean_t vm_object_populate(vm_object_t, vm_pindex_t, vm_pindex_t);
void vm_object_prepare_buf_pages(vm_object_t object, vm_page_t *ma_dst,
int count, int *rbehind, int *rahead, vm_page_t *ma_src);
void vm_object_print(long addr, boolean_t have_addr, long count, char *modif);
void vm_object_reference (vm_object_t);
void vm_object_reference_locked(vm_object_t);
int vm_object_set_memattr(vm_object_t object, vm_memattr_t memattr);
void vm_object_shadow(vm_object_t *, vm_ooffset_t *, vm_size_t, struct ucred *,
bool);
void vm_object_split(vm_map_entry_t);
boolean_t vm_object_sync(vm_object_t, vm_ooffset_t, vm_size_t, boolean_t,
boolean_t);
void vm_object_unwire(vm_object_t object, vm_ooffset_t offset,
vm_size_t length, uint8_t queue);
struct vnode *vm_object_vnode(vm_object_t object);
bool vm_object_is_active(vm_object_t obj);
#endif /* _KERNEL */
#endif /* _VM_OBJECT_ */
diff --git a/sys/vm/vm_pager.c b/sys/vm/vm_pager.c
index 2c6ab1410840b8d329377257257028576188d4c7..4729b2159987a7083bf9803a03a277d34002ec63 100644
--- a/sys/vm/vm_pager.c
+++ b/sys/vm/vm_pager.c
@@ -1,626 +1,658 @@
/*-
* SPDX-License-Identifier: (BSD-3-Clause AND MIT-CMU)
*
* Copyright (c) 1991, 1993
* The Regents of the University of California. All rights reserved.
*
* This code is derived from software contributed to Berkeley by
* The Mach Operating System project at Carnegie-Mellon University.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice, this list of conditions and the following disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
* 3. Neither the name of the University nor the names of its contributors
* may be used to endorse or promote products derived from this software
* without specific prior written permission.
*
* THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS ``AS IS'' AND
* ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
* ARE DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS
* OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)
* HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT
* LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY
* OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
* SUCH DAMAGE.
*
*
* Copyright (c) 1987, 1990 Carnegie-Mellon University.
* All rights reserved.
*
* Authors: Avadis Tevanian, Jr., Michael Wayne Young
*
* Permission to use, copy, modify and distribute this software and
* its documentation is hereby granted, provided that both the copyright
* notice and this permission notice appear in all copies of the
* software, derivative works or modified versions, and any portions
* thereof, and that both notices appear in supporting documentation.
*
* CARNEGIE MELLON ALLOWS FREE USE OF THIS SOFTWARE IN ITS "AS IS"
* CONDITION. CARNEGIE MELLON DISCLAIMS ANY LIABILITY OF ANY KIND
* FOR ANY DAMAGES WHATSOEVER RESULTING FROM THE USE OF THIS SOFTWARE.
*
* Carnegie Mellon requests users of this software to return to
*
* Software Distribution Coordinator or Software.Distribution@CS.CMU.EDU
* School of Computer Science
* Carnegie Mellon University
* Pittsburgh PA 15213-3890
*
* any improvements or extensions that they make and grant Carnegie the
* rights to redistribute these changes.
*/
/*
* Paging space routine stubs. Emulates a matchmaker-like interface
* for builtin pagers.
*/
#include <sys/cdefs.h>
#include "opt_param.h"
#include <sys/param.h>
#include <sys/systm.h>
#include <sys/kernel.h>
#include <sys/vnode.h>
#include <sys/bio.h>
#include <sys/buf.h>
#include <sys/ucred.h>
#include <sys/malloc.h>
#include <sys/rwlock.h>
#include <sys/user.h>
#include <vm/vm.h>
#include <vm/vm_param.h>
#include <vm/vm_kern.h>
#include <vm/vm_object.h>
#include <vm/vm_page.h>
#include <vm/vm_pager.h>
#include <vm/vm_extern.h>
#include <vm/uma.h>
uma_zone_t pbuf_zone;
static int pbuf_init(void *, int, int);
static int pbuf_ctor(void *, int, void *, int);
static void pbuf_dtor(void *, int, void *);
static int dead_pager_getpages(vm_object_t, vm_page_t *, int, int *, int *);
static vm_object_t dead_pager_alloc(void *, vm_ooffset_t, vm_prot_t,
vm_ooffset_t, struct ucred *);
static void dead_pager_putpages(vm_object_t, vm_page_t *, int, int, int *);
static boolean_t dead_pager_haspage(vm_object_t, vm_pindex_t, int *, int *);
static void dead_pager_dealloc(vm_object_t);
static void dead_pager_getvp(vm_object_t, struct vnode **, bool *);
static int
dead_pager_getpages(vm_object_t obj, vm_page_t *ma, int count, int *rbehind,
int *rahead)
{
return (VM_PAGER_FAIL);
}
static vm_object_t
dead_pager_alloc(void *handle, vm_ooffset_t size, vm_prot_t prot,
vm_ooffset_t off, struct ucred *cred)
{
return (NULL);
}
static void
dead_pager_putpages(vm_object_t object, vm_page_t *m, int count,
int flags, int *rtvals)
{
int i;
for (i = 0; i < count; i++)
rtvals[i] = VM_PAGER_AGAIN;
}
static boolean_t
dead_pager_haspage(vm_object_t object, vm_pindex_t pindex, int *prev, int *next)
{
if (prev != NULL)
*prev = 0;
if (next != NULL)
*next = 0;
return (FALSE);
}
static void
dead_pager_dealloc(vm_object_t object)
{
}
static void
dead_pager_getvp(vm_object_t object, struct vnode **vpp, bool *vp_heldp)
{
/*
* For OBJT_DEAD objects, v_writecount was handled in
* vnode_pager_dealloc().
*/
}
static const struct pagerops deadpagerops = {
.pgo_kvme_type = KVME_TYPE_DEAD,
.pgo_alloc = dead_pager_alloc,
.pgo_dealloc = dead_pager_dealloc,
.pgo_getpages = dead_pager_getpages,
.pgo_putpages = dead_pager_putpages,
.pgo_haspage = dead_pager_haspage,
.pgo_getvp = dead_pager_getvp,
};
const struct pagerops *pagertab[16] __read_mostly = {
[OBJT_SWAP] = &swappagerops,
[OBJT_VNODE] = &vnodepagerops,
[OBJT_DEVICE] = &devicepagerops,
[OBJT_PHYS] = &physpagerops,
[OBJT_DEAD] = &deadpagerops,
[OBJT_SG] = &sgpagerops,
[OBJT_MGTDEVICE] = &mgtdevicepagerops,
};
static struct mtx pagertab_lock;
void
vm_pager_init(void)
{
const struct pagerops **pgops;
int i;
mtx_init(&pagertab_lock, "dynpag", NULL, MTX_DEF);
/*
* Initialize known pagers
*/
for (i = 0; i < OBJT_FIRST_DYN; i++) {
pgops = &pagertab[i];
if (*pgops != NULL && (*pgops)->pgo_init != NULL)
(*(*pgops)->pgo_init)();
}
}
static int nswbuf_max;
void
vm_pager_bufferinit(void)
{
/* Main zone for paging bufs. */
pbuf_zone = uma_zcreate("pbuf",
sizeof(struct buf) + PBUF_PAGES * sizeof(vm_page_t),
pbuf_ctor, pbuf_dtor, pbuf_init, NULL, UMA_ALIGN_CACHE,
UMA_ZONE_NOFREE);
/* Few systems may still use this zone directly, so it needs a limit. */
nswbuf_max += uma_zone_set_max(pbuf_zone, NSWBUF_MIN);
}
uma_zone_t
pbuf_zsecond_create(const char *name, int max)
{
uma_zone_t zone;
zone = uma_zsecond_create(name, pbuf_ctor, pbuf_dtor, NULL, NULL,
pbuf_zone);
#ifdef KMSAN
/*
* Shrink the size of the pbuf pools if KMSAN is enabled, otherwise the
* shadows of the large KVA allocations eat up too much memory.
*/
max /= 3;
#endif
/*
* uma_prealloc() rounds up to items per slab. If we would prealloc
* immediately on every pbuf_zsecond_create(), we may accumulate too
* much of difference between hard limit and prealloced items, which
* means wasted memory.
*/
if (nswbuf_max > 0)
nswbuf_max += uma_zone_set_max(zone, max);
else
uma_prealloc(pbuf_zone, uma_zone_set_max(zone, max));
return (zone);
}
static void
pbuf_prealloc(void *arg __unused)
{
uma_prealloc(pbuf_zone, nswbuf_max);
nswbuf_max = -1;
}
SYSINIT(pbuf, SI_SUB_KTHREAD_BUF, SI_ORDER_ANY, pbuf_prealloc, NULL);
/*
* Allocate an instance of a pager of the given type.
* Size, protection and offset parameters are passed in for pagers that
* need to perform page-level validation (e.g. the device pager).
*/
vm_object_t
vm_pager_allocate(objtype_t type, void *handle, vm_ooffset_t size,
vm_prot_t prot, vm_ooffset_t off, struct ucred *cred)
{
vm_object_t object;
MPASS(type < nitems(pagertab));
object = (*pagertab[type]->pgo_alloc)(handle, size, prot, off, cred);
if (object != NULL)
object->type = type;
return (object);
}
/*
* The object must be locked.
*/
void
vm_pager_deallocate(vm_object_t object)
{
VM_OBJECT_ASSERT_WLOCKED(object);
MPASS(object->type < nitems(pagertab));
(*pagertab[object->type]->pgo_dealloc) (object);
}
static void
vm_pager_assert_in(vm_object_t object, vm_page_t *m, int count)
{
#ifdef INVARIANTS
/*
* All pages must be consecutive, busied, not mapped, not fully valid,
* not dirty and belong to the proper object. Some pages may be the
* bogus page, but the first and last pages must be a real ones.
*/
VM_OBJECT_ASSERT_UNLOCKED(object);
VM_OBJECT_ASSERT_PAGING(object);
KASSERT(count > 0, ("%s: 0 count", __func__));
for (int i = 0 ; i < count; i++) {
if (m[i] == bogus_page) {
KASSERT(i != 0 && i != count - 1,
("%s: page %d is the bogus page", __func__, i));
continue;
}
vm_page_assert_xbusied(m[i]);
KASSERT(!pmap_page_is_mapped(m[i]),
("%s: page %p is mapped", __func__, m[i]));
KASSERT(m[i]->valid != VM_PAGE_BITS_ALL,
("%s: request for a valid page %p", __func__, m[i]));
KASSERT(m[i]->dirty == 0,
("%s: page %p is dirty", __func__, m[i]));
KASSERT(m[i]->object == object,
("%s: wrong object %p/%p", __func__, object, m[i]->object));
KASSERT(m[i]->pindex == m[0]->pindex + i,
("%s: page %p isn't consecutive", __func__, m[i]));
}
#endif
}
/*
* Page in the pages for the object using its associated pager.
* The requested page must be fully valid on successful return.
*/
int
vm_pager_get_pages(vm_object_t object, vm_page_t *m, int count, int *rbehind,
int *rahead)
{
#ifdef INVARIANTS
vm_pindex_t pindex = m[0]->pindex;
#endif
int r;
MPASS(object->type < nitems(pagertab));
vm_pager_assert_in(object, m, count);
r = (*pagertab[object->type]->pgo_getpages)(object, m, count, rbehind,
rahead);
if (r != VM_PAGER_OK)
return (r);
for (int i = 0; i < count; i++) {
/*
* If pager has replaced a page, assert that it had
* updated the array.
*/
#ifdef INVARIANTS
KASSERT(m[i] == vm_page_relookup(object, pindex++),
("%s: mismatch page %p pindex %ju", __func__,
m[i], (uintmax_t )pindex - 1));
#endif
/*
* Zero out partially filled data.
*/
if (m[i]->valid != VM_PAGE_BITS_ALL)
vm_page_zero_invalid(m[i], TRUE);
}
return (VM_PAGER_OK);
}
int
vm_pager_get_pages_async(vm_object_t object, vm_page_t *m, int count,
int *rbehind, int *rahead, pgo_getpages_iodone_t iodone, void *arg)
{
MPASS(object->type < nitems(pagertab));
vm_pager_assert_in(object, m, count);
return ((*pagertab[object->type]->pgo_getpages_async)(object, m,
count, rbehind, rahead, iodone, arg));
}
/*
* vm_pager_put_pages() - inline, see vm/vm_pager.h
* vm_pager_has_page() - inline, see vm/vm_pager.h
*/
/*
* Search the specified pager object list for an object with the
* specified handle. If an object with the specified handle is found,
* increase its reference count and return it. Otherwise, return NULL.
*
* The pager object list must be locked.
*/
vm_object_t
vm_pager_object_lookup(struct pagerlst *pg_list, void *handle)
{
vm_object_t object;
TAILQ_FOREACH(object, pg_list, pager_object_list) {
if (object->handle == handle) {
VM_OBJECT_WLOCK(object);
if ((object->flags & OBJ_DEAD) == 0) {
vm_object_reference_locked(object);
VM_OBJECT_WUNLOCK(object);
break;
}
VM_OBJECT_WUNLOCK(object);
}
}
return (object);
}
+void
+vm_pager_populate_release_page(vm_object_t object, vm_page_t page)
+{
+ pgo_populate_release_page_t *method;
+
+ VM_OBJECT_ASSERT_WLOCKED(object);
+ method = pagertab[object->type]->pgo_populate_release_page;
+ if (method != NULL)
+ method(object, page);
+ else
+ vm_page_xunbusy(page);
+}
+
int
vm_pager_alloc_dyn_type(struct pagerops *ops, int base_type)
{
int res;
mtx_lock(&pagertab_lock);
+ MPASS((ops->pgo_populate_take_page == NULL) ==
+ (ops->pgo_populate_done == NULL));
+ MPASS(ops->pgo_populate_take_page == NULL ||
+ ops->pgo_populate != NULL);
+ MPASS(ops->pgo_populate_release_page == NULL ||
+ ops->pgo_populate_take_page != NULL);
MPASS(base_type == -1 ||
(base_type >= OBJT_SWAP && base_type < nitems(pagertab)));
for (res = OBJT_FIRST_DYN; res < nitems(pagertab); res++) {
if (pagertab[res] == NULL)
break;
}
if (res == nitems(pagertab)) {
mtx_unlock(&pagertab_lock);
return (-1);
}
if (base_type != -1) {
MPASS(pagertab[base_type] != NULL);
#define FIX(n) \
if (ops->pgo_##n == NULL) \
ops->pgo_##n = pagertab[base_type]->pgo_##n
FIX(init);
FIX(alloc);
FIX(dealloc);
FIX(getpages);
FIX(getpages_async);
FIX(putpages);
FIX(haspage);
- FIX(populate);
FIX(pageunswapped);
FIX(update_writecount);
FIX(release_writecount);
FIX(set_writeable_dirty);
FIX(mightbedirty);
FIX(getvp);
FIX(freespace);
FIX(page_inserted);
FIX(page_removed);
FIX(can_alloc_page);
#undef FIX
+ /* The populate handoff methods form one indivisible contract. */
+ if (ops->pgo_populate == NULL) {
+ MPASS(ops->pgo_populate_take_page == NULL);
+ MPASS(ops->pgo_populate_done == NULL);
+ MPASS(ops->pgo_populate_release_page == NULL);
+ ops->pgo_populate =
+ pagertab[base_type]->pgo_populate;
+ ops->pgo_populate_take_page =
+ pagertab[base_type]->pgo_populate_take_page;
+ ops->pgo_populate_done =
+ pagertab[base_type]->pgo_populate_done;
+ ops->pgo_populate_release_page =
+ pagertab[base_type]->pgo_populate_release_page;
+ }
}
pagertab[res] = ops; /* XXXKIB should be rel, but acq is too much */
mtx_unlock(&pagertab_lock);
return (res);
}
void
vm_pager_free_dyn_type(objtype_t type)
{
MPASS(type >= OBJT_FIRST_DYN && type < nitems(pagertab));
mtx_lock(&pagertab_lock);
MPASS(pagertab[type] != NULL);
pagertab[type] = NULL;
mtx_unlock(&pagertab_lock);
}
static int
pbuf_ctor(void *mem, int size, void *arg, int flags)
{
struct buf *bp = mem;
bp->b_vp = NULL;
bp->b_bufobj = NULL;
/* copied from initpbuf() */
bp->b_rcred = NOCRED;
bp->b_wcred = NOCRED;
bp->b_qindex = 0; /* On no queue (QUEUE_NONE) */
bp->b_data = bp->b_kvabase;
bp->b_xflags = 0;
bp->b_flags = B_MAXPHYS;
bp->b_ioflags = 0;
bp->b_iodone = NULL;
bp->b_error = 0;
BUF_LOCK(bp, LK_EXCLUSIVE | LK_NOWITNESS, NULL);
return (0);
}
static void
pbuf_dtor(void *mem, int size, void *arg)
{
struct buf *bp = mem;
if (bp->b_rcred != NOCRED) {
crfree(bp->b_rcred);
bp->b_rcred = NOCRED;
}
if (bp->b_wcred != NOCRED) {
crfree(bp->b_wcred);
bp->b_wcred = NOCRED;
}
BUF_UNLOCK(bp);
}
static const char pbuf_wmesg[] = "pbufwait";
static int
pbuf_init(void *mem, int size, int flags)
{
struct buf *bp = mem;
TSENTER();
bp->b_kvabase = kva_alloc(ptoa(PBUF_PAGES));
if (bp->b_kvabase == NULL)
return (ENOMEM);
bp->b_kvasize = ptoa(PBUF_PAGES);
BUF_LOCKINIT(bp, pbuf_wmesg);
LIST_INIT(&bp->b_dep);
bp->b_rcred = bp->b_wcred = NOCRED;
bp->b_xflags = 0;
TSEXIT();
return (0);
}
/*
* Associate a p-buffer with a vnode.
*
* Also sets B_PAGING flag to indicate that vnode is not fully associated
* with the buffer. i.e. the bp has not been linked into the vnode or
* ref-counted.
*/
void
pbgetvp(struct vnode *vp, struct buf *bp)
{
KASSERT(bp->b_vp == NULL, ("pbgetvp: not free"));
KASSERT(bp->b_bufobj == NULL, ("pbgetvp: not free (bufobj)"));
bp->b_vp = vp;
bp->b_flags |= B_PAGING;
bp->b_bufobj = &vp->v_bufobj;
}
/*
* Associate a p-buffer with a vnode.
*
* Also sets B_PAGING flag to indicate that vnode is not fully associated
* with the buffer. i.e. the bp has not been linked into the vnode or
* ref-counted.
*/
void
pbgetbo(struct bufobj *bo, struct buf *bp)
{
KASSERT(bp->b_vp == NULL, ("pbgetbo: not free (vnode)"));
KASSERT(bp->b_bufobj == NULL, ("pbgetbo: not free (bufobj)"));
bp->b_flags |= B_PAGING;
bp->b_bufobj = bo;
}
/*
* Disassociate a p-buffer from a vnode.
*/
void
pbrelvp(struct buf *bp)
{
KASSERT(bp->b_vp != NULL, ("pbrelvp: NULL"));
KASSERT(bp->b_bufobj != NULL, ("pbrelvp: NULL bufobj"));
KASSERT((bp->b_xflags & (BX_VNDIRTY | BX_VNCLEAN)) == 0,
("pbrelvp: pager buf on vnode list."));
bp->b_vp = NULL;
bp->b_bufobj = NULL;
bp->b_flags &= ~B_PAGING;
}
/*
* Disassociate a p-buffer from a bufobj.
*/
void
pbrelbo(struct buf *bp)
{
KASSERT(bp->b_vp == NULL, ("pbrelbo: vnode"));
KASSERT(bp->b_bufobj != NULL, ("pbrelbo: NULL bufobj"));
KASSERT((bp->b_xflags & (BX_VNDIRTY | BX_VNCLEAN)) == 0,
("pbrelbo: pager buf on vnode list."));
bp->b_bufobj = NULL;
bp->b_flags &= ~B_PAGING;
}
void
vm_object_set_writeable_dirty(vm_object_t object)
{
pgo_set_writeable_dirty_t *method;
MPASS(object->type < nitems(pagertab));
method = pagertab[object->type]->pgo_set_writeable_dirty;
if (method != NULL)
method(object);
}
bool
vm_object_mightbedirty(vm_object_t object)
{
pgo_mightbedirty_t *method;
MPASS(object->type < nitems(pagertab));
method = pagertab[object->type]->pgo_mightbedirty;
if (method == NULL)
return (false);
return (method(object));
}
/*
* Return the kvme type of the given object.
* If vpp is not NULL, set it to the object's vm_object_vnode() or NULL.
*/
int
vm_object_kvme_type(vm_object_t object, struct vnode **vpp)
{
VM_OBJECT_ASSERT_LOCKED(object);
MPASS(object->type < nitems(pagertab));
if (vpp != NULL)
*vpp = vm_object_vnode(object);
return (pagertab[object->type]->pgo_kvme_type);
}
diff --git a/sys/vm/vm_pager.h b/sys/vm/vm_pager.h
index 7857176286854e2f296f24f2d6e03563b967bb8e..104f104d8254499bbe18cc0b9e5353f2b3270b94 100644
--- a/sys/vm/vm_pager.h
+++ b/sys/vm/vm_pager.h
@@ -1,326 +1,388 @@
/*-
* SPDX-License-Identifier: BSD-3-Clause
*
* Copyright (c) 1990 University of Utah.
* Copyright (c) 1991, 1993
* The Regents of the University of California. All rights reserved.
*
* This code is derived from software contributed to Berkeley by
* the Systems Programming Group of the University of Utah Computer
* Science Department.
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions
* are met:
* 1. Redistributions of source code must retain the above copyright
* notice, this list of conditions and the following disclaimer.
* 2. Redistributions in binary form must reproduce the above copyright
* notice, this list of conditions and the following disclaimer in the
* documentation and/or other materials provided with the distribution.
* 3. Neither the name of the University nor the names of its contributors
* may be used to endorse or promote products derived from this software
* without specific prior written permission.
*
* THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS ``AS IS'' AND
* ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
* ARE DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS
* OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)
* HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT
* LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY
* OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
* SUCH DAMAGE.
*/
/*
* Pager routine interface definition.
*/
#ifndef _VM_PAGER_
#define _VM_PAGER_
#include <sys/systm.h>
TAILQ_HEAD(pagerlst, vm_object);
struct vnode;
typedef void pgo_init_t(void);
typedef vm_object_t pgo_alloc_t(void *, vm_ooffset_t, vm_prot_t, vm_ooffset_t,
struct ucred *);
typedef void pgo_dealloc_t(vm_object_t);
typedef int pgo_getpages_t(vm_object_t, vm_page_t *, int, int *, int *);
typedef void pgo_getpages_iodone_t(void *, vm_page_t *, int, int);
typedef int pgo_getpages_async_t(vm_object_t, vm_page_t *, int, int *, int *,
pgo_getpages_iodone_t, void *);
typedef void pgo_putpages_t(vm_object_t, vm_page_t *, int, int, int *);
typedef boolean_t pgo_haspage_t(vm_object_t, vm_pindex_t, int *, int *);
typedef int pgo_populate_t(vm_object_t, vm_pindex_t, int, vm_prot_t,
vm_pindex_t *, vm_pindex_t *);
+typedef vm_page_t pgo_populate_take_page_t(vm_object_t, vm_pindex_t);
+typedef void pgo_populate_done_t(vm_object_t);
+typedef void pgo_populate_release_page_t(vm_object_t, vm_page_t);
typedef void pgo_pageunswapped_t(vm_page_t);
typedef void pgo_writecount_t(vm_object_t, vm_offset_t, vm_offset_t);
typedef void pgo_set_writeable_dirty_t(vm_object_t);
typedef bool pgo_mightbedirty_t(vm_object_t);
typedef void pgo_getvp_t(vm_object_t object, struct vnode **vpp,
bool *vp_heldp);
typedef void pgo_freespace_t(vm_object_t object, vm_pindex_t start,
vm_size_t size);
typedef void pgo_page_inserted_t(vm_object_t object, vm_page_t m);
typedef void pgo_page_removed_t(vm_object_t object, vm_page_t m);
typedef boolean_t pgo_can_alloc_page_t(vm_object_t object, vm_pindex_t pindex);
struct pagerops {
int pgo_kvme_type;
pgo_init_t *pgo_init; /* Initialize pager. */
pgo_alloc_t *pgo_alloc; /* Allocate pager. */
pgo_dealloc_t *pgo_dealloc; /* Disassociate. */
pgo_getpages_t *pgo_getpages; /* Get (read) page. */
pgo_getpages_async_t *pgo_getpages_async; /* Get page asyncly. */
pgo_putpages_t *pgo_putpages; /* Put (write) page. */
pgo_haspage_t *pgo_haspage; /* Query page. */
pgo_populate_t *pgo_populate; /* Bulk spec pagein. */
pgo_pageunswapped_t *pgo_pageunswapped;
pgo_writecount_t *pgo_update_writecount;
pgo_writecount_t *pgo_release_writecount;
pgo_set_writeable_dirty_t *pgo_set_writeable_dirty;
pgo_mightbedirty_t *pgo_mightbedirty;
pgo_getvp_t *pgo_getvp;
pgo_freespace_t *pgo_freespace;
pgo_page_inserted_t *pgo_page_inserted;
pgo_page_removed_t *pgo_page_removed;
pgo_can_alloc_page_t *pgo_can_alloc_page;
+ pgo_populate_take_page_t *pgo_populate_take_page;
+ pgo_populate_done_t *pgo_populate_done;
+ pgo_populate_release_page_t *pgo_populate_release_page;
};
extern const struct pagerops defaultpagerops;
extern const struct pagerops swappagerops;
extern const struct pagerops vnodepagerops;
extern const struct pagerops devicepagerops;
extern const struct pagerops physpagerops;
extern const struct pagerops sgpagerops;
extern const struct pagerops mgtdevicepagerops;
extern const struct pagerops swaptmpfspagerops;
/*
* get/put return values
* OK operation was successful
* BAD specified data was out of the accepted range
* FAIL specified data was in range, but doesn't exist
* PEND operations was initiated but not completed
* ERROR error while accessing data that is in range and exists
* AGAIN temporary resource shortage prevented operation from happening
+ * OUT_OF_BOUNDS pager-specific address is invalid and should raise SIGBUS
+ * RETRY restart the fault after releasing its locks and references
*/
#define VM_PAGER_OK 0
#define VM_PAGER_BAD 1
#define VM_PAGER_FAIL 2
#define VM_PAGER_PEND 3
#define VM_PAGER_ERROR 4
#define VM_PAGER_AGAIN 5
+#define VM_PAGER_OUT_OF_BOUNDS 6
+#define VM_PAGER_RETRY 7
#define VM_PAGER_PUT_SYNC 0x0001
#define VM_PAGER_PUT_INVAL 0x0002
#define VM_PAGER_PUT_NOREUSE 0x0004
#define VM_PAGER_CLUSTER_OK 0x0008
#ifdef _KERNEL
extern const struct pagerops *pagertab[] __read_mostly;
extern struct mtx_padalign pbuf_mtx;
/*
* Number of pages that pbuf buffer can store in b_pages.
* It is +1 to allow for unaligned data buffer of maxphys size.
*/
#define PBUF_PAGES (atop(maxphys) + 1)
vm_object_t vm_pager_allocate(objtype_t, void *, vm_ooffset_t, vm_prot_t,
vm_ooffset_t, struct ucred *);
void vm_pager_bufferinit(void);
void vm_pager_deallocate(vm_object_t);
int vm_pager_get_pages(vm_object_t, vm_page_t *, int, int *, int *);
int vm_pager_get_pages_async(vm_object_t, vm_page_t *, int, int *, int *,
pgo_getpages_iodone_t, void *);
void vm_pager_init(void);
vm_object_t vm_pager_object_lookup(struct pagerlst *, void *);
static __inline void
vm_pager_put_pages(vm_object_t object, vm_page_t *m, int count, int flags,
int *rtvals)
{
VM_OBJECT_ASSERT_WLOCKED(object);
(*pagertab[object->type]->pgo_putpages)
(object, m, count, flags, rtvals);
}
/*
* vm_pager_haspage
*
* Check to see if an object's pager has the requested page. The
* object's pager will also set before and after to give the caller
* some idea of the number of pages before and after the requested
* page can be I/O'd efficiently.
*
* The object must be locked.
*/
static __inline boolean_t
vm_pager_has_page(vm_object_t object, vm_pindex_t offset, int *before,
int *after)
{
boolean_t ret;
VM_OBJECT_ASSERT_LOCKED(object);
ret = (*pagertab[object->type]->pgo_haspage)
(object, offset, before, after);
return (ret);
}
static __inline int
vm_pager_populate(vm_object_t object, vm_pindex_t pidx, int fault_type,
vm_prot_t max_prot, vm_pindex_t *first, vm_pindex_t *last)
{
MPASS((object->flags & OBJ_POPULATE) != 0);
MPASS(pidx < object->size);
MPASS(blockcount_read(&object->paging_in_progress) > 0);
return ((*pagertab[object->type]->pgo_populate)(object, pidx,
fault_type, max_prot, first, last));
}
+/*
+ * Take an xbusy page populated by the pager without inserting it into the
+ * pager object. The caller must release the busy state exactly once through
+ * vm_pager_populate_release_page(), including when discarding a page. A
+ * populate operation must not mix such pages with pages installed in the
+ * pager object. Most pagers populate only object-resident pages and leave
+ * this method unset.
+ */
+static __inline vm_page_t
+vm_pager_populate_take_page(vm_object_t object, vm_pindex_t pidx)
+{
+ pgo_populate_take_page_t *method;
+
+ method = pagertab[object->type]->pgo_populate_take_page;
+ return (method == NULL ? NULL : method(object, pidx));
+}
+
+/*
+ * Release an external page after successful mapping or discard. The caller
+ * has already selected its page-queue disposition. The method must release
+ * xbusy and may drop the page's final reference. The pager object is write
+ * locked on entry and return, but the method may drop that lock to acquire
+ * the backing object's lock. The caller holds a paging-in-progress
+ * reference, and the pager must keep its handoff serialized across the
+ * lock drop. A NULL method defaults to vm_page_xunbusy(). The caller must
+ * not access the page after this call.
+ */
+void vm_pager_populate_release_page(vm_object_t object, vm_page_t page);
+
+/* Notify the pager that all pages returned by populate() were consumed. */
+static __inline void
+vm_pager_populate_done(vm_object_t object)
+{
+ pgo_populate_done_t *method;
+
+ method = pagertab[object->type]->pgo_populate_done;
+ if (method != NULL)
+ method(object);
+}
+
/*
* vm_pager_page_unswapped
*
* Destroy swap associated with the page.
*
* XXX: A much better name would be "vm_pager_page_dirtied()"
* XXX: It is not obvious if this could be profitably used by any
* XXX: pagers besides the swap_pager or if it should even be a
* XXX: generic pager_op in the first place.
*/
static __inline void
vm_pager_page_unswapped(vm_page_t m)
{
pgo_pageunswapped_t *method;
method = pagertab[m->object->type]->pgo_pageunswapped;
if (method != NULL)
method(m);
}
static __inline void
vm_pager_update_writecount(vm_object_t object, vm_offset_t start,
vm_offset_t end)
{
pgo_writecount_t *method;
method = pagertab[object->type]->pgo_update_writecount;
if (method != NULL)
method(object, start, end);
}
static __inline void
vm_pager_release_writecount(vm_object_t object, vm_offset_t start,
vm_offset_t end)
{
pgo_writecount_t *method;
method = pagertab[object->type]->pgo_release_writecount;
if (method != NULL)
method(object, start, end);
}
static __inline void
vm_pager_getvp(vm_object_t object, struct vnode **vpp, bool *vp_heldp)
{
pgo_getvp_t *method;
*vpp = NULL;
if (vp_heldp != NULL)
*vp_heldp = false;
method = pagertab[object->type]->pgo_getvp;
if (method != NULL)
method(object, vpp, vp_heldp);
}
static __inline void
vm_pager_freespace(vm_object_t object, vm_pindex_t start,
vm_size_t size)
{
pgo_freespace_t *method;
method = pagertab[object->type]->pgo_freespace;
if (method != NULL)
method(object, start, size);
}
static __inline void
vm_pager_page_inserted(vm_object_t object, vm_page_t m)
{
pgo_page_inserted_t *method;
method = pagertab[object->type]->pgo_page_inserted;
if (method != NULL)
method(object, m);
}
static __inline void
vm_pager_page_removed(vm_object_t object, vm_page_t m)
{
pgo_page_removed_t *method;
method = pagertab[object->type]->pgo_page_removed;
if (method != NULL)
method(object, m);
}
static __inline bool
vm_pager_can_alloc_page(vm_object_t object, vm_pindex_t pindex)
{
pgo_can_alloc_page_t *method;
method = pagertab[object->type]->pgo_can_alloc_page;
return (method != NULL ? method(object, pindex) : true);
}
int vm_pager_alloc_dyn_type(struct pagerops *ops, int base_type);
void vm_pager_free_dyn_type(objtype_t type);
struct cdev_pager_ops {
int (*cdev_pg_fault)(vm_object_t vm_obj, vm_ooffset_t offset,
int prot, vm_page_t *mres);
int (*cdev_pg_populate)(vm_object_t vm_obj, vm_pindex_t pidx,
int fault_type, vm_prot_t max_prot, vm_pindex_t *first,
vm_pindex_t *last);
int (*cdev_pg_ctor)(void *handle, vm_ooffset_t size, vm_prot_t prot,
vm_ooffset_t foff, struct ucred *cred, u_short *color);
void (*cdev_pg_dtor)(void *handle);
void (*cdev_pg_path)(void *handle, char *path, size_t len);
+ vm_page_t (*cdev_pg_populate_take_page)(vm_object_t vm_obj,
+ vm_pindex_t pidx);
+ void (*cdev_pg_populate_done)(vm_object_t vm_obj);
+ void (*cdev_pg_populate_release_page)(vm_object_t vm_obj, vm_page_t page);
+ /*
+ * Optional special-mapping policy, queried after successful construction
+ * and before publishing the object. True exempts this object's mappings
+ * from userspace memory locking, but does not permit kernel wiring to
+ * succeed without actually wiring pages. The result is immutable for
+ * the object's lifetime; NULL preserves ordinary memory-locking policy.
+ */
+ bool (*cdev_pg_mlock_skip)(void *handle);
};
vm_object_t cdev_pager_allocate(void *handle, enum obj_type tp,
const struct cdev_pager_ops *ops, vm_ooffset_t size, vm_prot_t prot,
vm_ooffset_t foff, struct ucred *cred);
vm_object_t cdev_pager_lookup(void *handle);
void cdev_pager_free_page(vm_object_t object, vm_page_t m);
void cdev_mgtdev_pager_free_page(struct pctrie_iter *pages, vm_page_t m);
void cdev_mgtdev_pager_free_pages(vm_object_t object);
void cdev_pager_get_path(vm_object_t object, char *path, size_t sz);
struct phys_pager_ops {
int (*phys_pg_getpages)(vm_object_t vm_obj, vm_page_t *m, int count,
int *rbehind, int *rahead);
int (*phys_pg_populate)(vm_object_t vm_obj, vm_pindex_t pidx,
int fault_type, vm_prot_t max_prot, vm_pindex_t *first,
vm_pindex_t *last);
boolean_t (*phys_pg_haspage)(vm_object_t obj, vm_pindex_t pindex,
int *before, int *after);
void (*phys_pg_ctor)(vm_object_t vm_obj, vm_prot_t prot,
vm_ooffset_t foff, struct ucred *cred);
void (*phys_pg_dtor)(vm_object_t vm_obj);
};
extern const struct phys_pager_ops default_phys_pg_ops;
vm_object_t phys_pager_allocate(void *handle, const struct phys_pager_ops *ops,
void *data, vm_ooffset_t size, vm_prot_t prot, vm_ooffset_t foff,
struct ucred *cred);
#endif /* _KERNEL */
#endif /* _VM_PAGER_ */

File Metadata

Mime Type
text/x-diff
Storage Engine
blob
Storage Format
Raw Data
Storage Handle
39953358
Default Alt Text
pfnmap-D59481-main-d2018cedb414-full-context.patch (467 KB)

Event Timeline