From: Fred Griffoul <griffoul@gmail.com>
To: Paolo Bonzini <pbonzini@redhat.com>,
Sean Christopherson <seanjc@google.com>,
Marc Zyngier <maz@kernel.org>, Oliver Upton <oupton@kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Alexander Viro <viro@zeniv.linux.org.uk>,
Christian Brauner <brauner@kernel.org>, Jan Kara <jack@suse.cz>,
Jason Gunthorpe <jgg@ziepe.ca>, Kevin Tian <kevin.tian@intel.com>,
Joerg Roedel <joro@8bytes.org>, Will Deacon <will@kernel.org>,
Robin Murphy <robin.murphy@arm.com>,
Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H . Peter Anvin" <hpa@zytor.com>,
Jonathan Corbet <corbet@lwn.net>, Shuah Khan <shuah@kernel.org>
Cc: David Woodhouse <dwmw2@infradead.org>,
Ackerley Tng <ackerleytng@google.com>,
Lorenzo Stoakes <ljs@kernel.org>,
"Liam R . Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>, Joey Gouly <joey.gouly@arm.com>,
Suzuki K Poulose <suzuki.poulose@arm.com>,
Zenghui Yu <yuzenghui@huawei.com>,
Steffen Eiden <seiden@linux.ibm.com>,
linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
kvmarm@lists.linux.dev, iommu@lists.linux.dev,
linux-fsdevel@vger.kernel.org, linux-mm@kvack.org,
linux-kselftest@vger.kernel.org
Subject: [PATCH 3/9] KVM: guest_memfd: Add a memory provider backing
Date: Tue, 6 Oct 2026 18:32:29 +0000 [thread overview]
Message-ID: <20261006183235.16576-4-griffoul@gmail.com> (raw)
In-Reply-To: <20261006183235.16576-1-griffoul@gmail.com>
From: Fred Griffoul <fgriffo@amazon.co.uk>
A driver can back a guest_memfd today only by implementing
kvm_gmem_ops itself, and must then keep any device mapping of the same
memory consistent on its own.
Add GUEST_MEMFD_FLAG_USE_PROVIDER. KVM_CREATE_GUEST_MEMFD then takes a
memory provider file in provider_fd, which uses the first reserved word
and must be zero without the flag. Guest faults ask the provider for
each frame, and a revoke removes the range from the guest and from the
VMM's mapping. A page that is not backed, or that is not RAM, exits with
KVM_EXIT_MEMORY_FAULT. Read-only pages are mapped read only, and
fallocate() is not supported.
guest_memfd maps itself into userspace from the same frames, so that
a provider needs no fault handler. Pages marked NO_USER_MAP raise
SIGBUS there.
USE_PROVIDER requires GUEST_MEMFD_FLAG_MMAP, so that the memory slot
is gmem-only. An architecture opts in; x86 does so for VMs without
private or encrypted memory, and other architectures refuse the flag
for now. Native guest_memfd files now take the invalidate lock when
a file is added and when a closing file is unbound.
Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
Documentation/virt/kvm/api.rst | 45 ++-
arch/x86/kvm/x86.c | 10 +
include/linux/kvm_host.h | 4 +
include/uapi/linux/kvm.h | 11 +-
tools/include/uapi/linux/kvm.h | 11 +-
.../testing/selftests/kvm/guest_memfd_test.c | 6 +
virt/kvm/Kconfig | 1 +
virt/kvm/guest_memfd.c | 265 +++++++++++++++++-
8 files changed, 338 insertions(+), 15 deletions(-)
diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index a5f9ee92f43e..d9aaf8a76b59 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6431,7 +6431,9 @@ and cannot be resized (guest_memfd files do however support PUNCH_HOLE).
struct kvm_create_guest_memfd {
__u64 size;
__u64 flags;
- __u64 reserved[6];
+ __s32 provider_fd;
+ __u32 pad;
+ __u64 reserved[5];
};
Conceptually, the inode backing a guest_memfd file represents physical memory,
@@ -6453,15 +6455,38 @@ a single guest_memfd file, but the bound ranges must not overlap).
The capability KVM_CAP_GUEST_MEMFD_FLAGS enumerates the `flags` that can be
specified via KVM_CREATE_GUEST_MEMFD. Currently defined flags:
- ============================ ================================================
- GUEST_MEMFD_FLAG_MMAP Enable using mmap() on the guest_memfd file
- descriptor.
- GUEST_MEMFD_FLAG_INIT_SHARED Make all memory in the file shared during
- KVM_CREATE_GUEST_MEMFD (memory files created
- without INIT_SHARED will be marked private).
- Shared memory can be faulted into host userspace
- page tables. Private memory cannot.
- ============================ ================================================
+ ============================= ================================================
+ GUEST_MEMFD_FLAG_MMAP Enable using mmap() on the guest_memfd file
+ descriptor.
+ GUEST_MEMFD_FLAG_INIT_SHARED Make all memory in the file shared during
+ KVM_CREATE_GUEST_MEMFD (memory files created
+ without INIT_SHARED will be marked private).
+ Shared memory can be faulted into host userspace
+ page tables. Private memory cannot.
+ GUEST_MEMFD_FLAG_USE_PROVIDER Take the memory of the file from the memory
+ provider behind `provider_fd`, instead of
+ allocating it. Requires
+ GUEST_MEMFD_FLAG_MMAP. Not supported for VMs
+ with private or encrypted memory.
+ ============================= ================================================
+
+`pad` must be zero. With GUEST_MEMFD_FLAG_USE_PROVIDER, `provider_fd` is a file
+handed out by a memory provider (see include/linux/mem_provider.h); without it,
+`provider_fd` must be zero.
+The provider decides which pages exist and what backs them, and can change
+this at any time. The guest and any host mapping of the guest_memfd follow
+the change. A guest access to a page that the provider does not back exits
+to userspace with KVM_EXIT_MEMORY_FAULT, and so does an access to a page that
+is not RAM. A page that the provider marks read only is mapped read only for
+the guest. fallocate() is not supported. Only x86 VMs of type
+KVM_X86_DEFAULT_VM support GUEST_MEMFD_FLAG_USE_PROVIDER.
+
+mmap() of the guest_memfd maps the provider's pages into userspace on fault.
+An access raises SIGBUS if the provider does not back the page, if the page is
+not RAM, if the provider marks it as not to be mapped into userspace, or if it
+is a write to a read-only page.
+KVM's own accesses through the memslot's userspace address fail in the same
+cases.
When the KVM MMU performs a PFN lookup to service a guest fault and the backing
guest_memfd has the GUEST_MEMFD_FLAG_MMAP set, then the fault will always be
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index afcac1042947..3669316f5712 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -14124,6 +14124,16 @@ bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm)
return !kvm_arch_has_private_mem(kvm);
}
+/*
+ * Only VMs without encrypted memory: SEV and SEV-ES guests have no private
+ * memory in KVM's sense, but their memory would need reclaiming when a
+ * provider takes it back.
+ */
+bool kvm_arch_gmem_supports_provider(struct kvm *kvm)
+{
+ return !kvm || kvm->arch.vm_type == KVM_X86_DEFAULT_VM;
+}
+
#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREPARE
int kvm_arch_gmem_prepare(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, int max_order)
{
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 7281d0e94121..3ea591642542 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -832,11 +832,15 @@ static inline bool kvm_arch_has_private_mem(struct kvm *kvm)
#ifdef CONFIG_KVM_GUEST_MEMFD
bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm);
+bool kvm_arch_gmem_supports_provider(struct kvm *kvm);
static inline u64 kvm_gmem_get_supported_flags(struct kvm *kvm)
{
u64 flags = GUEST_MEMFD_FLAG_MMAP;
+ if (kvm_arch_gmem_supports_provider(kvm))
+ flags |= GUEST_MEMFD_FLAG_USE_PROVIDER;
+
if (!kvm || kvm_arch_supports_gmem_init_shared(kvm))
flags |= GUEST_MEMFD_FLAG_INIT_SHARED;
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index 419011097fa8..820448f82e22 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -1654,11 +1654,20 @@ struct kvm_memory_attributes {
#define KVM_CREATE_GUEST_MEMFD _IOWR(KVMIO, 0xd4, struct kvm_create_guest_memfd)
#define GUEST_MEMFD_FLAG_MMAP (1ULL << 0)
#define GUEST_MEMFD_FLAG_INIT_SHARED (1ULL << 1)
+#define GUEST_MEMFD_FLAG_USE_PROVIDER (1ULL << 2)
struct kvm_create_guest_memfd {
__u64 size;
__u64 flags;
- __u64 reserved[6];
+ /*
+ * With GUEST_MEMFD_FLAG_USE_PROVIDER: a memory provider file whose
+ * pages back this guest_memfd. The provider may take
+ * any range back at any time; the guest and every host mapping of
+ * this guest_memfd follow. Must be 0 otherwise.
+ */
+ __s32 provider_fd;
+ __u32 pad;
+ __u64 reserved[5];
};
#define KVM_PRE_FAULT_MEMORY _IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)
diff --git a/tools/include/uapi/linux/kvm.h b/tools/include/uapi/linux/kvm.h
index d0c0c8605976..0f4d1dd0e931 100644
--- a/tools/include/uapi/linux/kvm.h
+++ b/tools/include/uapi/linux/kvm.h
@@ -1644,11 +1644,20 @@ struct kvm_memory_attributes {
#define KVM_CREATE_GUEST_MEMFD _IOWR(KVMIO, 0xd4, struct kvm_create_guest_memfd)
#define GUEST_MEMFD_FLAG_MMAP (1ULL << 0)
#define GUEST_MEMFD_FLAG_INIT_SHARED (1ULL << 1)
+#define GUEST_MEMFD_FLAG_USE_PROVIDER (1ULL << 2)
struct kvm_create_guest_memfd {
__u64 size;
__u64 flags;
- __u64 reserved[6];
+ /*
+ * With GUEST_MEMFD_FLAG_USE_PROVIDER: a memory provider file whose
+ * pages back this guest_memfd. The provider may take
+ * any range back at any time; the guest and every host mapping of
+ * this guest_memfd follow. Must be 0 otherwise.
+ */
+ __s32 provider_fd;
+ __u32 pad;
+ __u64 reserved[5];
};
#define KVM_PRE_FAULT_MEMORY _IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)
diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 2233d871a38f..91ff10ac6274 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -405,6 +405,12 @@ static void test_guest_memfd_flags(struct kvm_vm *vm)
for (flag = BIT(0); flag; flag <<= 1) {
fd = __vm_create_guest_memfd(vm, page_size, flag);
+ /* USE_PROVIDER also needs MMAP and a provider file. */
+ if (flag == GUEST_MEMFD_FLAG_USE_PROVIDER && (flag & valid_flags)) {
+ TEST_ASSERT(fd < 0 && errno == EINVAL,
+ "guest_memfd() with USE_PROVIDER alone should fail with EINVAL");
+ continue;
+ }
if (flag & valid_flags) {
TEST_ASSERT(fd >= 0,
"guest_memfd() with flag '0x%lx' should succeed",
diff --git a/virt/kvm/Kconfig b/virt/kvm/Kconfig
index 794976b88c6f..cfb6c4e51128 100644
--- a/virt/kvm/Kconfig
+++ b/virt/kvm/Kconfig
@@ -105,6 +105,7 @@ config KVM_GENERIC_MEMORY_ATTRIBUTES
config KVM_GUEST_MEMFD
select XARRAY_MULTI
+ select MEM_PROVIDER
bool
config HAVE_KVM_ARCH_GMEM_PREPARE
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index a509f1a96c0b..aedd8630e3ea 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -4,6 +4,7 @@
#include <linux/falloc.h>
#include <linux/fs.h>
#include <linux/kvm_host.h>
+#include <linux/mem_provider.h>
#include <linux/mempolicy.h>
#include <linux/pseudo_fs.h>
#include <linux/pagemap.h>
@@ -44,6 +45,9 @@ struct gmem_inode {
struct list_head gmem_file_list;
u64 flags;
+
+ /* The provider of the pages, with GUEST_MEMFD_FLAG_USE_PROVIDER. */
+ struct mem_provider_attachment att;
};
static __always_inline struct gmem_inode *GMEM_I(struct inode *inode)
@@ -654,7 +658,222 @@ static const struct kvm_gmem_ops kvm_gmem_native_ops = {
.fallocate = kvm_gmem_native_fallocate,
};
-static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
+/*
+ * Can @kvm use a memory provider? @kvm is NULL for the system-wide capability.
+ * An architecture opts in by overriding this. VMs whose memory is private or
+ * encrypted are not supported: each frame would need preparing before use and
+ * reclaiming when the provider takes it back.
+ */
+bool __weak kvm_arch_gmem_supports_provider(struct kvm *kvm)
+{
+ return false;
+}
+
+/*
+ * A guest_memfd backed by a memory provider (include/linux/mem_provider.h).
+ *
+ * The provider owns the frames, and guest_memfd keeps no state for them. A
+ * guest fault asks the provider for the frame each time. A fault that
+ * races with a change in the provider is retried by KVM, because the
+ * provider revokes the range after the change and the revoke opens and
+ * closes KVM's invalidation window. Host mappings are made by guest_memfd
+ * from the same frames and attributes, and removed by the revoke.
+ *
+ * Bindings and release are the same as for the native guest_memfd.
+ */
+static void kvm_gmem_provider_revoke(struct mem_provider_attachment *att,
+ loff_t offset, loff_t len)
+{
+ struct gmem_inode *gi = container_of(att, struct gmem_inode, att);
+ struct inode *inode = &gi->vfs_inode;
+ pgoff_t start, end;
+ struct gmem_file *f;
+
+ if (offset >= att->size)
+ return;
+ len = min(len, att->size - offset);
+ start = offset >> PAGE_SHIFT;
+ end = DIV_ROUND_UP(offset + len, PAGE_SIZE);
+
+ /*
+ * The bindings must be stable so that each start is matched by an
+ * end. Only VMs without private memory use a provider, but remove
+ * both kinds of mapping so that none is missed.
+ */
+ filemap_invalidate_lock(inode->i_mapping);
+ kvm_gmem_for_each_file(f, inode)
+ __kvm_gmem_invalidate_start(f, start, end,
+ KVM_FILTER_SHARED | KVM_FILTER_PRIVATE);
+ unmap_mapping_range(inode->i_mapping, (loff_t)start << PAGE_SHIFT,
+ (loff_t)(end - start) << PAGE_SHIFT, 1);
+ kvm_gmem_for_each_file(f, inode)
+ __kvm_gmem_invalidate_end(f, start, end);
+ filemap_invalidate_unlock(inode->i_mapping);
+}
+
+static int kvm_gmem_provider_get_pfn(struct file *file, struct kvm *kvm,
+ struct kvm_memory_slot *slot, gfn_t gfn,
+ kvm_pfn_t *pfn, struct page **page,
+ int *max_order, bool *writable)
+{
+ struct gmem_inode *gi = GMEM_I(file_inode(file));
+ pgoff_t index = kvm_gmem_get_index(slot, gfn);
+ unsigned long frame;
+ int order, ret;
+ u32 attrs;
+
+ if (file != READ_ONCE(slot->gmem.file))
+ return -EFAULT;
+ if (xa_load(&gmem_file_of(file)->bindings, index) != slot)
+ return -EIO;
+
+ /* A large mapping needs the block aligned in both index and gfn. */
+ order = PUD_ORDER;
+ if (index != gfn)
+ order = min_t(int, order, __ffs(index ^ gfn));
+
+ ret = mem_provider_get_page(&gi->att, index, &frame, &order, &attrs);
+ if (ret)
+ return ret;
+
+ /*
+ * KVM decides the guest's memory type itself, so accept RAM only.
+ * -EFAULT, so that the VMM gets a memory fault exit for the access.
+ */
+ if (mem_provider_type(attrs) != MEM_PROVIDER_TYPE_RAM)
+ return -EFAULT;
+
+ if (attrs & MEM_PROVIDER_ATTR_READONLY) {
+ if (!writable)
+ return -EPERM;
+ *writable = false;
+ }
+
+ *pfn = frame;
+ if (max_order)
+ *max_order = order;
+ return 0;
+}
+
+/* The protection for a host mapping of a frame with @attrs. */
+static pgprot_t kvm_gmem_provider_prot(struct vm_area_struct *vma, u32 attrs)
+{
+ vm_flags_t flags = vma->vm_flags;
+
+ if (attrs & MEM_PROVIDER_ATTR_READONLY)
+ flags &= ~VM_WRITE;
+ return vm_get_page_prot(flags);
+}
+
+/*
+ * Map page @vmf->pgoff of a provider-backed guest_memfd for the host.
+ *
+ * The invalidate lock is held shared across get_page() and the insert, and
+ * a revoke holds it exclusive while it unmaps the range, so a frame is never
+ * mapped after the revoke that removes it. A writable page is mapped
+ * writable, and a read-only page without write permission. With @mkwrite,
+ * the PTE is read only and the write must be checked: replace it under the
+ * lock, so that a change to read only cannot slip between the check and the
+ * upgrade.
+ */
+static vm_fault_t kvm_gmem_provider_map_host(struct vm_fault *vmf, bool mkwrite)
+{
+ struct vm_area_struct *vma = vmf->vma;
+ struct inode *inode = file_inode(vma->vm_file);
+ unsigned long uaddr = vmf->address & PAGE_MASK;
+ bool write = vmf->flags & FAULT_FLAG_WRITE;
+ unsigned long pfn;
+ int order = 0;
+ vm_fault_t ret;
+ u32 attrs;
+
+ if (((loff_t)vmf->pgoff << PAGE_SHIFT) >= i_size_read(inode))
+ return VM_FAULT_SIGBUS;
+
+ filemap_invalidate_lock_shared(inode->i_mapping);
+ if (mem_provider_get_page(&GMEM_I(inode)->att, vmf->pgoff, &pfn,
+ &order, &attrs)) {
+ ret = VM_FAULT_SIGBUS;
+ } else if (mem_provider_type(attrs) != MEM_PROVIDER_TYPE_RAM ||
+ (attrs & MEM_PROVIDER_ATTR_NO_USER_MAP) ||
+ (write && (attrs & MEM_PROVIDER_ATTR_READONLY))) {
+ ret = VM_FAULT_SIGBUS;
+ } else {
+ if (mkwrite)
+ zap_special_vma_range(vma, uaddr, PAGE_SIZE);
+ ret = vmf_insert_pfn_prot(vma, uaddr, pfn,
+ kvm_gmem_provider_prot(vma, attrs));
+ }
+ filemap_invalidate_unlock_shared(inode->i_mapping);
+ return ret;
+}
+
+static vm_fault_t kvm_gmem_provider_fault(struct vm_fault *vmf)
+{
+ return kvm_gmem_provider_map_host(vmf, false);
+}
+
+/* A write to a page that a read fault, or mprotect(), left read only. */
+static vm_fault_t kvm_gmem_provider_pfn_mkwrite(struct vm_fault *vmf)
+{
+ return kvm_gmem_provider_map_host(vmf, true);
+}
+
+static const struct vm_operations_struct kvm_gmem_provider_vm_ops = {
+ .fault = kvm_gmem_provider_fault,
+ .pfn_mkwrite = kvm_gmem_provider_pfn_mkwrite,
+};
+
+static int kvm_gmem_provider_mmap(struct file *file, struct vm_area_struct *vma)
+{
+ struct inode *inode = file_inode(file);
+ pgoff_t npages = i_size_read(inode) >> PAGE_SHIFT;
+
+ if (!kvm_gmem_supports_mmap(inode))
+ return -ENODEV;
+
+ if ((vma->vm_flags & (VM_SHARED | VM_MAYSHARE)) !=
+ (VM_SHARED | VM_MAYSHARE))
+ return -EINVAL;
+
+ if (vma->vm_pgoff >= npages || vma_pages(vma) > npages - vma->vm_pgoff)
+ return -EINVAL;
+
+ /*
+ * The frames may have no struct page. They are inserted on fault, so
+ * that a range that is revoked and comes back is reached again.
+ */
+ vm_flags_set(vma, VM_PFNMAP | VM_IO | VM_DONTEXPAND | VM_DONTDUMP);
+ vma->vm_ops = &kvm_gmem_provider_vm_ops;
+ return 0;
+}
+
+static const struct kvm_gmem_ops kvm_gmem_provider_ops = {
+ .bind = kvm_gmem_native_bind,
+ .unbind = kvm_gmem_native_unbind,
+ .get_pfn = kvm_gmem_provider_get_pfn,
+ .release = kvm_gmem_native_release,
+ .mmap = kvm_gmem_provider_mmap,
+ /* The provider decides which pages exist: no fallocate(). */
+};
+
+static int kvm_gmem_provider_attach(struct inode *inode, int provider_fd)
+{
+ struct file *file;
+ int ret;
+
+ file = fget(provider_fd);
+ if (!file)
+ return -EBADF;
+
+ ret = mem_provider_attach(&GMEM_I(inode)->att, file,
+ i_size_read(inode), kvm_gmem_provider_revoke);
+ fput(file);
+ return ret;
+}
+
+static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags,
+ int provider_fd)
{
static const char *name = "[kvm-gmem]";
struct gmem_file *f;
@@ -695,6 +914,12 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
GMEM_I(inode)->flags = flags;
+ if (flags & GUEST_MEMFD_FLAG_USE_PROVIDER) {
+ err = kvm_gmem_provider_attach(inode, provider_fd);
+ if (err)
+ goto err_inode;
+ }
+
file = alloc_file_pseudo(inode, kvm_gmem_mnt, name, O_RDWR, &kvm_gmem_fops);
if (IS_ERR(file)) {
err = PTR_ERR(file);
@@ -703,12 +928,18 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
file->f_flags |= O_LARGEFILE;
file->private_data = &f->backing;
- f->backing.ops = &kvm_gmem_native_ops;
+ if (flags & GUEST_MEMFD_FLAG_USE_PROVIDER)
+ f->backing.ops = &kvm_gmem_provider_ops;
+ else
+ f->backing.ops = &kvm_gmem_native_ops;
kvm_get_kvm(kvm);
f->kvm = kvm;
xa_init(&f->bindings);
+ /* A provider can revoke, and walk the list, as soon as it is attached. */
+ filemap_invalidate_lock(inode->i_mapping);
list_add(&f->entry, &GMEM_I(inode)->gmem_file_list);
+ filemap_invalidate_unlock(inode->i_mapping);
fd_install(fd, file);
return fd;
@@ -735,7 +966,20 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args)
if (size <= 0 || !PAGE_ALIGNED(size))
return -EINVAL;
- return __kvm_gmem_create(kvm, size, flags);
+ if (args->pad ||
+ (!(flags & GUEST_MEMFD_FLAG_USE_PROVIDER) && args->provider_fd))
+ return -EINVAL;
+
+ /*
+ * Without MMAP the memslot is not gmem-only, and a VM without private
+ * memory would fault through the slot's userspace address instead of
+ * the provider.
+ */
+ if ((flags & GUEST_MEMFD_FLAG_USE_PROVIDER) &&
+ !(flags & GUEST_MEMFD_FLAG_MMAP))
+ return -EINVAL;
+
+ return __kvm_gmem_create(kvm, size, flags, args->provider_fd);
}
/*
@@ -911,9 +1155,14 @@ static void kvm_gmem_native_unbind(struct file *slot_file, struct kvm *kvm,
* bindings. I.e. reaching this point means kvm_gmem_release() hasn't
* yet destroyed the bindings or freed the gmem_file, and can't do so
* until the caller drops slots_lock.
+ *
+ * A memory provider can still revoke, and walk the bindings, until
+ * the inode is evicted, so take the invalidate lock in this case too.
*/
if (!file) {
+ filemap_invalidate_lock(slot_file->f_mapping);
__kvm_gmem_unbind(slot, gmem_file_of(slot_file));
+ filemap_invalidate_unlock(slot_file->f_mapping);
return;
}
@@ -1219,6 +1468,7 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb)
mpol_shared_policy_init(&gi->policy, NULL);
gi->flags = 0;
+ gi->att.ops = NULL;
INIT_LIST_HEAD(&gi->gmem_file_list);
return &gi->vfs_inode;
}
@@ -1233,10 +1483,19 @@ static void kvm_gmem_free_inode(struct inode *inode)
kmem_cache_free(kvm_gmem_inode_cachep, GMEM_I(inode));
}
+static void kvm_gmem_evict_inode(struct inode *inode)
+{
+ /* Every file is closed, so nothing maps the provider's frames. */
+ mem_provider_detach(&GMEM_I(inode)->att);
+ truncate_inode_pages_final(&inode->i_data);
+ clear_inode(inode);
+}
+
static const struct super_operations kvm_gmem_super_operations = {
.statfs = simple_statfs,
.alloc_inode = kvm_gmem_alloc_inode,
.destroy_inode = kvm_gmem_destroy_inode,
+ .evict_inode = kvm_gmem_evict_inode,
.free_inode = kvm_gmem_free_inode,
};
next prev parent reply other threads:[~2026-10-06 18:32 UTC|newest]
Thread overview: 38+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-20 11:03 [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 02/11] KVM: selftests: sev_init2_tests: Derive SEV availability from KVM David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 03/11] KVM: SEV: Remove struct page dependency from SNP gmem paths David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 04/11] KVM: guest_memfd: Introduce guest memory ops and route native gmem through them David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 05/11] iommufd: Look up private-interconnect phys via exporter symbols David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 06/11] iommufd: Plumb dma-buf memory-type (RAM vs MMIO) through the phys map David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 07/11] KVM: guest_memfd: Add ops-driven page revocation David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 08/11] samples/kvm: Add guest_memfd backing sample David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 09/11] selftests/kvm: gmem_provider KVM-only tests David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 10/11] selftests/kvm: gmem_provider iommufd tests David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 11/11] samples/kvm, selftests/kvm: Allow the gmem_provider NVMe DMA test on arm64 David Woodhouse
2026-07-20 15:11 ` [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory Paolo Bonzini
2026-07-20 16:39 ` David Woodhouse
2026-07-23 0:24 ` Ackerley Tng
2026-07-23 9:40 ` David Woodhouse
2026-07-23 16:01 ` Ackerley Tng
2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul
2026-10-05 9:55 ` [RFC PATCH 1/6] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
2026-10-05 9:55 ` [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run Fred Griffoul
2026-10-05 10:07 ` Christian König
2026-10-05 13:20 ` Fred Griffoul
2026-10-05 14:53 ` Christian König
2026-10-05 9:55 ` [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation Fred Griffoul
2026-10-05 10:08 ` Christian König
2026-10-05 9:55 ` [RFC PATCH 4/6] dma-buf: Allow dynamic attach without a device Fred Griffoul
2026-10-05 9:55 ` [RFC PATCH 5/6] KVM: guest_memfd: Add dma-buf backing Fred Griffoul
2026-10-05 9:55 ` [RFC PATCH 6/6] samples/kvm, selftests/kvm: Exercise " Fred Griffoul
2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
2026-10-06 18:32 ` [PATCH 1/9] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
2026-10-06 18:32 ` [PATCH 2/9] mm: Add memory providers Fred Griffoul
2026-10-06 18:32 ` Fred Griffoul [this message]
2026-10-06 18:32 ` [PATCH 4/9] iommufd: Track the domains of pages that are not pinned Fred Griffoul
2026-10-06 18:32 ` [PATCH 5/9] iommufd: Map memory provider files Fred Griffoul
2026-10-06 18:32 ` [PATCH 6/9] iommufd/selftest: Add mock-domain IOVA queries Fred Griffoul
2026-10-06 18:32 ` [PATCH 7/9] iommufd/selftest: Add a mock memory provider Fred Griffoul
2026-10-06 18:32 ` [PATCH 8/9] samples/kvm: Add a memory provider sample Fred Griffoul
2026-10-06 18:32 ` [PATCH 9/9] KVM: selftests: Test a memory provider shared by KVM and iommufd Fred Griffoul
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261006183235.16576-4-griffoul@gmail.com \
--to=griffoul@gmail.com \
--cc=ackerleytng@google.com \
--cc=akpm@linux-foundation.org \
--cc=bp@alien8.de \
--cc=brauner@kernel.org \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=dwmw2@infradead.org \
--cc=hpa@zytor.com \
--cc=iommu@lists.linux.dev \
--cc=jack@suse.cz \
--cc=jgg@ziepe.ca \
--cc=joey.gouly@arm.com \
--cc=joro@8bytes.org \
--cc=kevin.tian@intel.com \
--cc=kvm@vger.kernel.org \
--cc=kvmarm@lists.linux.dev \
--cc=liam@infradead.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=maz@kernel.org \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=oupton@kernel.org \
--cc=pbonzini@redhat.com \
--cc=robin.murphy@arm.com \
--cc=rppt@kernel.org \
--cc=seanjc@google.com \
--cc=seiden@linux.ibm.com \
--cc=shuah@kernel.org \
--cc=surenb@google.com \
--cc=suzuki.poulose@arm.com \
--cc=tglx@kernel.org \
--cc=vbabka@kernel.org \
--cc=viro@zeniv.linux.org.uk \
--cc=will@kernel.org \
--cc=x86@kernel.org \
--cc=yuzenghui@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®