* [PATCH 0/2] KVM: x86: Fix lost nested APF VM-Exits
@ 2026-10-02 5:41 Loc Nguyen
2026-10-02 5:42 ` [PATCH 1/2] KVM: x86: Request event processing for exception VM-Exits Loc Nguyen
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: Loc Nguyen @ 2026-10-02 5:41 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini
Cc: Shuah Khan, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, hpa, x86, kvm, linux-kernel, linux-kselftest
A nested asynchronous page fault delivered to L1 as a synthetic #PF
VM-Exit can be lost after KVM queues the exception. Most callers of
kvm_queue_exception_vmexit() arrive through kvm_multiple_exception(),
which requests event processing, but the nested APF path calls the
helper directly.
Without KVM_REQ_EVENT, vcpu_enter_guest() can skip
kvm_check_and_inject_events() and re-enter L2. A later VM-Exit can then
clear the pending exception while reconstructing vectoring state. The
APF reason remains PAGE_NOT_PRESENT, allowing a later regular #PF to
consume the stale reason.
Patch 1 requests event processing in kvm_queue_exception_vmexit() so
that all queued exception VM-Exits are processed before re-entering the
guest.
Patch 2 adds a regression test that holds an L2 backing page missing
with userfaultfd. The test verifies that L1 receives the synthetic #PF
VM-Exit with the expected APF token and PAGE_NOT_PRESENT reason.
Testing was performed in a nested VMX environment:
- x86_64 kernel build with KVM_WERROR=y
- nested_apf_event_test fails on the base kernel with a lost nested
APF event and passes with the fix
- full KVM selftest suite passes
- kvm-unit-tests VMX suite passes with TIMEOUT=900
Loc Nguyen (2):
KVM: x86: Request event processing for exception VM-Exits
KVM: selftests: Add a test for lost nested APF VM-Exits
arch/x86/kvm/x86.c | 2 +
tools/testing/selftests/kvm/Makefile.kvm | 1 +
.../selftests/kvm/x86/nested_apf_event_test.c | 306 ++++++++++++++++++
3 files changed, 309 insertions(+)
create mode 100644 tools/testing/selftests/kvm/x86/nested_apf_event_test.c
base-commit: b378201ccd5280d0fff89bbe55e1eb00620ec0d5
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 1/2] KVM: x86: Request event processing for exception VM-Exits
2026-10-02 5:41 [PATCH 0/2] KVM: x86: Fix lost nested APF VM-Exits Loc Nguyen
@ 2026-10-02 5:42 ` Loc Nguyen
2026-10-02 5:43 ` [PATCH 2/2] KVM: selftests: Add a test for lost nested APF VM-Exits Loc Nguyen
2026-10-02 21:07 ` [PATCH 0/2] KVM: x86: Fix " Sean Christopherson
2 siblings, 0 replies; 4+ messages in thread
From: Loc Nguyen @ 2026-10-02 5:42 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini
Cc: Shuah Khan, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, hpa, x86, kvm, linux-kernel, linux-kselftest
Request event processing whenever kvm_queue_exception_vmexit() queues an
exception VM-Exit. Most callers reach the helper through
kvm_multiple_exception(), which already sets KVM_REQ_EVENT, but nested
asynchronous page faults call it directly.
Without another event request, vcpu_enter_guest() skips
kvm_check_and_inject_events() and re-enters L2. A subsequent VM-Exit can
clear the pending exception while reconstructing vectoring state, losing
the synthetic #PF VM-Exit. The APF reason remains PAGE_NOT_PRESENT and a
later regular #PF can consume the stale value.
Setting KVM_REQ_EVENT in the common helper ensures that L1 receives the
queued VM-Exit before KVM re-enters L2.
Fixes: 7709aba8f716 ("KVM: x86: Morph pending exceptions to pending VM-Exits at queue time")
Cc: stable@vger.kernel.org
Signed-off-by: Loc Nguyen <loc@ryzome.com>
---
arch/x86/kvm/x86.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index cf3fcdfd8ad2..376341a1ba04 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -450,6 +450,8 @@ static void kvm_queue_exception_vmexit(struct kvm_vcpu *vcpu, unsigned int vecto
{
struct kvm_queued_exception *ex = &vcpu->arch.exception_vmexit;
+ kvm_make_request(KVM_REQ_EVENT, vcpu);
+
ex->vector = vector;
ex->injected = false;
ex->pending = true;
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 2/2] KVM: selftests: Add a test for lost nested APF VM-Exits
2026-10-02 5:41 [PATCH 0/2] KVM: x86: Fix lost nested APF VM-Exits Loc Nguyen
2026-10-02 5:42 ` [PATCH 1/2] KVM: x86: Request event processing for exception VM-Exits Loc Nguyen
@ 2026-10-02 5:43 ` Loc Nguyen
2026-10-02 21:07 ` [PATCH 0/2] KVM: x86: Fix " Sean Christopherson
2 siblings, 0 replies; 4+ messages in thread
From: Loc Nguyen @ 2026-10-02 5:43 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini
Cc: Shuah Khan, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, hpa, x86, kvm, linux-kernel, linux-kselftest
Exercise delivery of an asynchronous page fault from L2 to L1 as a
synthetic #PF VM-Exit. Hold L2's backing page missing with userfaultfd
and bound the critical KVM_RUN with a signal-based timeout.
On an affected kernel, KVM queues the nested VM-Exit without requesting
event processing, re-enters L2, and blocks on the same missing page. The
test detects the timeout and verifies that L1's APF record still contains
KVM_PV_REASON_PAGE_NOT_PRESENT.
On a fixed kernel, verify that L1 receives an exception VM-Exit with a #PF,
error code zero, a nonzero APF token, and a matching page-not-present
reason.
Skip the test when VMX, EPT, the nested APF features, or userfaultfd fault
handling are unavailable.
Signed-off-by: Loc Nguyen <loc@ryzome.com>
---
tools/testing/selftests/kvm/Makefile.kvm | 1 +
.../selftests/kvm/x86/nested_apf_event_test.c | 306 ++++++++++++++++++
2 files changed, 307 insertions(+)
create mode 100644 tools/testing/selftests/kvm/x86/nested_apf_event_test.c
diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm
index b31ee61af284..a882c1f8a5fb 100644
--- a/tools/testing/selftests/kvm/Makefile.kvm
+++ b/tools/testing/selftests/kvm/Makefile.kvm
@@ -92,6 +92,7 @@ TEST_GEN_PROGS_x86 += x86/kvm_pv_test
TEST_GEN_PROGS_x86 += x86/kvm_buslock_test
TEST_GEN_PROGS_x86 += x86/monitor_mwait_test
TEST_GEN_PROGS_x86 += x86/msrs_test
+TEST_GEN_PROGS_x86 += x86/nested_apf_event_test
TEST_GEN_PROGS_x86 += x86/nested_close_kvm_test
TEST_GEN_PROGS_x86 += x86/nested_dirty_log_test
TEST_GEN_PROGS_x86 += x86/nested_emulation_test
diff --git a/tools/testing/selftests/kvm/x86/nested_apf_event_test.c b/tools/testing/selftests/kvm/x86/nested_apf_event_test.c
new file mode 100644
index 000000000000..97e23ef91ca6
--- /dev/null
+++ b/tools/testing/selftests/kvm/x86/nested_apf_event_test.c
@@ -0,0 +1,306 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (C) 2026, Ryzome
+ *
+ * Reproduce a lost nested asynchronous page-fault notification.
+ *
+ * L2 faults on a page held missing by userfaultfd. KVM writes
+ * PAGE_NOT_PRESENT to L1's APF record and queues a synthetic #PF VM-exit.
+ * The vulnerable path omits KVM_REQ_EVENT, re-enters L2, and loses the
+ * notification before L1 can consume it. A fixed L0 exits to L1 immediately.
+ */
+#include <errno.h>
+#include <fcntl.h>
+#include <linux/userfaultfd.h>
+#include <poll.h>
+#include <signal.h>
+#include <stdint.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/ioctl.h>
+#include <sys/mman.h>
+#include <sys/syscall.h>
+#include <unistd.h>
+
+#include <asm/kvm_para.h>
+
+#include "kvm_util.h"
+#include "processor.h"
+#include "test_util.h"
+#include "vmx.h"
+
+#define APF_VECTOR 0xF1U
+#define RUN_TIMEOUT_SECONDS 3U
+#define VMX_INTR_INFO_VECTOR_MASK 0xFFU
+#define VMX_INTR_INFO_TYPE_MASK 0x700U
+#define VMX_INTR_INFO_HARD_EXCEPTION (3U << 8)
+#define VMX_INTR_INFO_DELIVER_CODE BIT(11)
+#define VMX_INTR_INFO_VALID BIT(31)
+
+enum test_stage {
+ STAGE_APF_ARMED = 1,
+ STAGE_APF_DELIVERED,
+};
+
+static struct kvm_vcpu_pv_apf_data apf_data __aligned(64);
+static gva_t l2_target_gva;
+
+static void l2_guest_code(void)
+{
+ u64 value;
+
+ value = READ_ONCE(*(u64 *)READ_ONCE(l2_target_gva));
+ GUEST_FAIL("L2 resumed after the missing page was resolved: %#lx", value);
+}
+
+static void l1_guest_code(struct vmx_pages *vmx, gpa_t apf_data_gpa)
+{
+ u64 apf_msr;
+ u64 intr_info;
+ u64 reason;
+ u64 token;
+
+ GUEST_ASSERT(this_cpu_has(X86_FEATURE_KVM_ASYNC_PF));
+ GUEST_ASSERT(this_cpu_has(X86_FEATURE_KVM_ASYNC_PF_VMEXIT));
+ GUEST_ASSERT(this_cpu_has(X86_FEATURE_KVM_ASYNC_PF_INT));
+ prepare_for_vmx_operation(vmx);
+ load_vmcs(vmx);
+ prepare_vmcs(vmx, l2_guest_code);
+ vmwrite(GUEST_RFLAGS,
+ X86_EFLAGS_FIXED | X86_EFLAGS_IF);
+
+ wrmsr(MSR_KVM_ASYNC_PF_INT, APF_VECTOR);
+ apf_msr = apf_data_gpa | KVM_ASYNC_PF_ENABLED |
+ KVM_ASYNC_PF_SEND_ALWAYS |
+ KVM_ASYNC_PF_DELIVERY_AS_PF_VMEXIT |
+ KVM_ASYNC_PF_DELIVERY_AS_INT;
+ wrmsr(MSR_KVM_ASYNC_PF_EN, apf_msr);
+
+ /*
+ * Let userspace discard and userfaultfd-register the target only after
+ * all L1 setup exits have completed. The following KVM_RUN therefore
+ * has no pending event request to accidentally deliver the nested APF.
+ */
+ GUEST_SYNC(STAGE_APF_ARMED);
+
+ vmlaunch();
+ GUEST_ASSERT_EQ(vmread(VM_EXIT_REASON), EXIT_REASON_EXCEPTION_NMI);
+
+ intr_info = vmread(VM_EXIT_INTR_INFO);
+ GUEST_ASSERT(intr_info & VMX_INTR_INFO_VALID);
+ GUEST_ASSERT_EQ(intr_info & VMX_INTR_INFO_VECTOR_MASK, PF_VECTOR);
+ GUEST_ASSERT_EQ(intr_info & VMX_INTR_INFO_TYPE_MASK,
+ VMX_INTR_INFO_HARD_EXCEPTION);
+ GUEST_ASSERT(intr_info & VMX_INTR_INFO_DELIVER_CODE);
+ GUEST_ASSERT_EQ(vmread(VM_EXIT_INTR_ERROR_CODE), 0);
+
+ token = vmread(EXIT_QUALIFICATION);
+ reason = READ_ONCE(apf_data.flags);
+ GUEST_ASSERT(token);
+ GUEST_ASSERT_EQ(reason, KVM_PV_REASON_PAGE_NOT_PRESENT);
+
+ /* Release the APF record slot before userspace resolves the page. */
+ WRITE_ONCE(apf_data.flags, 0);
+ GUEST_SYNC_ARGS(STAGE_APF_DELIVERED, token, reason, 0, 0);
+}
+
+#ifdef __NR_userfaultfd
+static int register_missing_page(void *hva)
+{
+ struct uffdio_register uffd_register = {
+ .range = {
+ .start = (uintptr_t)hva,
+ .len = PAGE_SIZE,
+ },
+ .mode = UFFDIO_REGISTER_MODE_MISSING,
+ };
+ struct uffdio_api uffd_api = {
+ .api = UFFD_API,
+ };
+ int uffd;
+
+ uffd = syscall(__NR_userfaultfd, O_CLOEXEC | O_NONBLOCK);
+ if (uffd < 0 && (errno == EPERM || errno == EACCES))
+ ksft_exit_skip("userfaultfd is unavailable: %s\n", strerror(errno));
+ TEST_ASSERT(uffd >= 0, "userfaultfd failed: %s", strerror(errno));
+
+ TEST_ASSERT(ioctl(uffd, UFFDIO_API, &uffd_api) == 0,
+ "UFFDIO_API failed: %s", strerror(errno));
+ TEST_ASSERT(ioctl(uffd, UFFDIO_REGISTER, &uffd_register) == 0,
+ "UFFDIO_REGISTER failed: %s", strerror(errno));
+ TEST_ASSERT(uffd_register.ioctls & (1ULL << _UFFDIO_COPY),
+ "userfaultfd does not support UFFDIO_COPY");
+
+ return uffd;
+}
+
+static bool read_missing_page_event(int uffd, void *target_hva)
+{
+ struct pollfd pollfd = {
+ .fd = uffd,
+ .events = POLLIN,
+ };
+ struct uffd_msg msg;
+ ssize_t bytes;
+ int ret;
+
+ ret = poll(&pollfd, 1, 3000);
+ if (ret != 1 || !(pollfd.revents & POLLIN))
+ return false;
+
+ bytes = read(uffd, &msg, sizeof(msg));
+ if (bytes != sizeof(msg) || msg.event != UFFD_EVENT_PAGEFAULT)
+ return false;
+
+ return (msg.arg.pagefault.address & ~(PAGE_SIZE - 1ULL)) ==
+ (uintptr_t)target_hva;
+}
+
+static void resolve_missing_page(int uffd, void *target_hva)
+{
+ struct uffdio_copy uffd_copy = {
+ .dst = (uintptr_t)target_hva,
+ .len = PAGE_SIZE,
+ };
+ void *page = NULL;
+ int ret;
+
+ ret = posix_memalign(&page, PAGE_SIZE, PAGE_SIZE);
+ TEST_ASSERT(!ret, "posix_memalign failed: %s", strerror(ret));
+ memset(page, 0, PAGE_SIZE);
+ uffd_copy.src = (uintptr_t)page;
+
+ TEST_ASSERT(ioctl(uffd, UFFDIO_COPY, &uffd_copy) == 0,
+ "UFFDIO_COPY failed: %s", strerror(errno));
+ free(page);
+}
+
+static void alarm_handler(int signal)
+{
+ (void)signal;
+}
+
+static void assert_sync_stage(struct kvm_vcpu *vcpu, u64 expected_stage)
+{
+ struct ucall uc;
+
+ TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO);
+ switch (get_ucall(vcpu, &uc)) {
+ case UCALL_SYNC:
+ TEST_ASSERT_EQ(uc.args[1], expected_stage);
+ break;
+ case UCALL_ABORT:
+ REPORT_GUEST_ASSERT(uc);
+ break;
+ default:
+ TEST_FAIL("Unexpected ucall command %lu", uc.cmd);
+ }
+}
+
+static void run_test(void)
+{
+ struct kvm_vcpu_pv_apf_data *apf_hva;
+ struct sigaction action = {
+ .sa_handler = alarm_handler,
+ };
+ struct kvm_vcpu *vcpu;
+ struct kvm_vm *vm;
+ struct ucall uc = {};
+ sigset_t blocked_signals;
+ sigset_t old_signals;
+ gva_t nested_gva;
+ gva_t target_gva;
+ gpa_t apf_gpa;
+ void *target_hva;
+ u32 stale_reason;
+ bool uffd_event;
+ bool run_timed_out;
+ int run_errno = 0;
+ int run_ret;
+ int uffd;
+ u64 ucall_cmd = UCALL_NONE;
+
+ vm = vm_create_with_one_vcpu(&vcpu, l1_guest_code);
+ vm_enable_tdp(vm);
+ vcpu_alloc_vmx(vm, &nested_gva);
+ target_gva = vm_alloc_page(vm);
+ tdp_identity_map_default_memslots(vm);
+
+ l2_target_gva = target_gva;
+ sync_global_to_guest(vm, l2_target_gva);
+ apf_gpa = addr_gva2gpa(vm, (gva_t)&apf_data);
+ TEST_ASSERT(!(apf_gpa & 63U), "APF data GPA %#lx is misaligned",
+ apf_gpa);
+ apf_hva = addr_gva2hva(vm, (gva_t)&apf_data);
+ target_hva = addr_gva2hva(vm, target_gva);
+ vcpu_args_set(vcpu, 2, nested_gva, apf_gpa);
+
+ vcpu_run(vcpu);
+ assert_sync_stage(vcpu, STAGE_APF_ARMED);
+
+ TEST_ASSERT(madvise(target_hva, PAGE_SIZE, MADV_DONTNEED) == 0,
+ "madvise failed: %s", strerror(errno));
+ uffd = register_missing_page(target_hva);
+
+ TEST_ASSERT(sigaction(SIGALRM, &action, NULL) == 0,
+ "sigaction failed: %s", strerror(errno));
+ sigfillset(&blocked_signals);
+ sigdelset(&blocked_signals, SIGALRM);
+ TEST_ASSERT(sigprocmask(SIG_SETMASK, &blocked_signals, &old_signals) == 0,
+ "sigprocmask failed: %s", strerror(errno));
+
+ alarm(RUN_TIMEOUT_SECONDS);
+ run_ret = __vcpu_run(vcpu);
+ run_errno = errno;
+ alarm(0);
+ run_timed_out = run_ret < 0 && run_errno == EINTR;
+ TEST_ASSERT(sigprocmask(SIG_SETMASK, &old_signals, NULL) == 0,
+ "sigprocmask restore failed: %s", strerror(errno));
+
+ stale_reason = READ_ONCE(apf_hva->flags);
+ uffd_event = read_missing_page_event(uffd, target_hva);
+ if (uffd_event)
+ resolve_missing_page(uffd, target_hva);
+ close(uffd);
+
+ if (!run_ret)
+ ucall_cmd = get_ucall(vcpu, &uc);
+ if (ucall_cmd == UCALL_ABORT)
+ REPORT_GUEST_ASSERT(uc);
+ kvm_vm_free(vm);
+
+ TEST_ASSERT(uffd_event,
+ "L2 access did not produce the held userfaultfd fault");
+ TEST_ASSERT(!run_timed_out ||
+ stale_reason == KVM_PV_REASON_PAGE_NOT_PRESENT,
+ "KVM_RUN timed out without a stale APF reason: reason=%u",
+ stale_reason);
+ TEST_ASSERT(!run_timed_out,
+ "Lost nested APF event: KVM re-entered L2 and blocked; APF reason=%u",
+ stale_reason);
+ TEST_ASSERT(!run_ret, "KVM_RUN failed: %s", strerror(run_errno));
+ TEST_ASSERT_EQ(ucall_cmd, UCALL_SYNC);
+ TEST_ASSERT_EQ(uc.args[1], STAGE_APF_DELIVERED);
+ TEST_ASSERT_EQ(uc.args[3], KVM_PV_REASON_PAGE_NOT_PRESENT);
+ TEST_ASSERT(uc.args[2], "KVM returned a zero APF token");
+}
+#endif
+
+int main(int argc, char *argv[])
+{
+ (void)argc;
+ (void)argv;
+
+#ifndef __NR_userfaultfd
+ ksft_exit_skip("__NR_userfaultfd is unavailable\n");
+#else
+ TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_VMX));
+ TEST_REQUIRE(kvm_cpu_has_ept());
+ TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_KVM_ASYNC_PF));
+ TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_KVM_ASYNC_PF_VMEXIT));
+ TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_KVM_ASYNC_PF_INT));
+
+ run_test();
+#endif
+ return 0;
+}
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 0/2] KVM: x86: Fix lost nested APF VM-Exits
2026-10-02 5:41 [PATCH 0/2] KVM: x86: Fix lost nested APF VM-Exits Loc Nguyen
2026-10-02 5:42 ` [PATCH 1/2] KVM: x86: Request event processing for exception VM-Exits Loc Nguyen
2026-10-02 5:43 ` [PATCH 2/2] KVM: selftests: Add a test for lost nested APF VM-Exits Loc Nguyen
@ 2026-10-02 21:07 ` Sean Christopherson
2 siblings, 0 replies; 4+ messages in thread
From: Sean Christopherson @ 2026-10-02 21:07 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini, Loc Nguyen
Cc: Shuah Khan, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, hpa, x86, kvm, linux-kernel, linux-kselftest
On Fri, 02 Oct 2026 05:41:43 +0000, Loc Nguyen wrote:
> A nested asynchronous page fault delivered to L1 as a synthetic #PF
> VM-Exit can be lost after KVM queues the exception. Most callers of
> kvm_queue_exception_vmexit() arrive through kvm_multiple_exception(),
> which requests event processing, but the nested APF path calls the
> helper directly.
>
> Without KVM_REQ_EVENT, vcpu_enter_guest() can skip
> kvm_check_and_inject_events() and re-enter L2. A later VM-Exit can then
> clear the pending exception while reconstructing vectoring state. The
> APF reason remains PAGE_NOT_PRESENT, allowing a later regular #PF to
> consume the stale reason.
>
> [...]
Applied the fix to kvm-x86 misc. I'm going to hold off on the selftest for
the moment, I'd like to go straight to having it support both AMD and Intel.
Thanks!
[1/2] KVM: x86: Request event processing for exception VM-Exits
https://github.com/kvm-x86/linux/commit/88b23a2e8286
[2/2] KVM: selftests: Add a test for lost nested APF VM-Exits
[SKIP, for now]
--
https://github.com/kvm-x86/linux/tree/next
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-02 21:08 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 5:41 [PATCH 0/2] KVM: x86: Fix lost nested APF VM-Exits Loc Nguyen
2026-10-02 5:42 ` [PATCH 1/2] KVM: x86: Request event processing for exception VM-Exits Loc Nguyen
2026-10-02 5:43 ` [PATCH 2/2] KVM: selftests: Add a test for lost nested APF VM-Exits Loc Nguyen
2026-10-02 21:07 ` [PATCH 0/2] KVM: x86: Fix " Sean Christopherson
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®