mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3 0/3] KVM: PPC: Fixes for Book3S HV HPT locking and paired-single decoding
@ 2026-10-06 12:24 Amit Machhiwal
  2026-10-06 12:24 ` [PATCH v3 1/3] KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls Amit Machhiwal
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: Amit Machhiwal @ 2026-10-06 12:24 UTC (permalink / raw)
  To: Madhavan Srinivasan, linuxppc-dev
  Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
	Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
	Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
	Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
	linux-hardening, stable, Avi Kivity

This series addresses bug fixes across KVM PPC Book3S HV
locking/synchronization and paired-single instruction decoding.

Patches 1 & 2 fix synchronization and preemption issues introduced in
commit 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual mode handlers for
HPT hcalls and page faults"):

- Patch 1 adds SRCU read lock protection when walking memslots during
  virtual-mode HPT hcalls to prevent use-after-free races with concurrent
  memslot updates.

- Patch 2 adds preempt_disable() around all virtual-mode HPTE bit-lock
  holders — both the guest vCPU paths (kvmppc_hpte_hv_fault() and
  kvmppc_pseries_do_hpt_hcall()) and the host-side paths
  (kvm_unmap_rmapp(), kvm_age_rmapp(), resize_hpt_rehash_hpte()) — to
  prevent CPU stalls and deadlocks on preemption.

Patch 3 fixes paired-single D-form instruction emulation:

- Patch 3 fixes get_d_signext() to correctly extract the full 12-bit D
  displacement field and perform proper two's-complement sign extension.

Testing:
========
All test kernel builds were compiled with CONFIG_DEBUG_ATOMIC_SLEEP=y.
The following scenarios were verified:

1. Power9 PowerNV (L0) in Radix mode:
   - Booted L0 host with kernel containing all 3 patches.
   - Booted KVM guests in both Radix and Hash modes.
   - Ran kernel build workload inside guests — no errors observed.
   - Booted L1 guest with the same kernel — booted and ran cleanly.

2. Power9 PowerNV (L0) in Hash mode:
   - Booted L0 host with kernel containing all 3 patches.
   - Booted KVM guest in Hash mode.
   - Ran kernel build workload inside guest — no errors observed.
   - Booted L1 nested guest with the same kernel — booted and ran cleanly.

3. Power10 LPAR (L1):
   - Booted Power10 LPAR with the patched kernel.
   - Booted KVM guest and ran workloads — no errors observed.

Changes in v3:
==============
- v2: https://lore.kernel.org/all/20260930173750.56759-1-amachhiw@linux.ibm.com/
- Patch 2:
  - Fixed preemption window in kvm_unmap_rmapp() and kvm_age_rmapp(): moved
    preempt_disable() before lock_rmap() so both the rmap lock and the HPTE
    bit-lock are held under a single non-preemptible section, and added
    preempt_enable() on all early exits and retry paths before cpu_relax().
  - Cleaned up kvm_test_clear_dirty_npages(): removed redundant per-iteration
    preempt_disable()/preempt_enable() pairs since the caller
    kvmppc_hv_get_dirty_log_hpt() already holds preempt_disable() across the
    entire loop.
  - Reworded commit message to clearly detail the per-function locking design
    and rationale.
  - Picked up Reviewed-by tag from Shrikanth Hegde.
- Patch 3:
  - Picked up Reviewed-by tag from Shrikanth Hegde.

Changes in v2:
==============
- v1: https://lore.kernel.org/all/20260928122837.8782-1-amachhiw@linux.ibm.com/
- Patch 2: Extended preempt_disable()/preempt_enable() coverage to also wrap
  the HPTE bit-lock hold windows in four host-side virtual-mode functions:
  kvm_unmap_rmapp(), kvm_age_rmapp(), kvm_test_clear_dirty_npages(), and
  resize_hpt_rehash_hpte(). These were identified as vulnerable by Sashiko AI
  review and confirmed correct by audit.
- Dropped Reviewed-by from Ritesh as the patch was materially extended.
- Patch 2: Added warning comment above kvmppc_pseries_do_hpt_hcall()
  documenting the preemption requirement.

Amit Machhiwal (3):
  KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls
  KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE
    bit-lock users
  KVM: PPC: Fix get_d_signext() 12-bit displacement for paired-single
    D-form

 arch/powerpc/kvm/book3s_64_mmu_hv.c      | 10 +++
 arch/powerpc/kvm/book3s_hv.c             | 80 +++++++++++++-----------
 arch/powerpc/kvm/book3s_paired_singles.c |  7 +--
 3 files changed, 56 insertions(+), 41 deletions(-)


base-commit: 2c3418fffa9d037b2038a6db48be63f9e2291806
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v3 1/3] KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls
  2026-10-06 12:24 [PATCH v3 0/3] KVM: PPC: Fixes for Book3S HV HPT locking and paired-single decoding Amit Machhiwal
@ 2026-10-06 12:24 ` Amit Machhiwal
  2026-10-06 12:24 ` [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users Amit Machhiwal
  2026-10-06 12:24 ` [PATCH v3 3/3] KVM: PPC: Fix get_d_signext() 12-bit displacement for paired-single D-form Amit Machhiwal
  2 siblings, 0 replies; 5+ messages in thread
From: Amit Machhiwal @ 2026-10-06 12:24 UTC (permalink / raw)
  To: Madhavan Srinivasan, linuxppc-dev
  Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
	Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
	Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
	Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
	linux-hardening, stable, Avi Kivity

The virtual-mode HPT hcall handlers (H_ENTER, H_REMOVE, H_BULK_REMOVE,
H_CLEAR_REF, H_CLEAR_MOD) call into book3s_hv_rm_mmu.c, which accesses
memslots via kvm_memslots_raw().  kvm_memslots_raw() uses
rcu_dereference_raw_check() to bypass SRCU lockdep annotation checking.
This is safe in real mode because the entire guest entry/exit is wrapped
in srcu_read_lock/unlock inside kvmppc_run_core() and
kvmhv_run_single_vcpu().

However, since commit 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual
mode handlers for HPT hcalls and page faults"), these same handlers are
also executed in virtual mode via kvmppc_pseries_do_hcall(), which runs
after SRCU has already been released on guest exit.

A concurrent KVM_SET_USER_MEMORY_REGION deletion or move can therefore
call synchronize_srcu_expedited() — which does not wait for this thread —
and then kfree(slot) and vfree(slot->arch.rmap) while the hcall handler
still holds a raw pointer to the memslot.  lock_rmap() then writes to
freed vmalloc memory, leading to use-after-free and memory corruption.
Because kvm_memslots_raw() suppresses lockdep checks, this race is
entirely silent.

Fix this by factoring out the 7 HPT hcall handlers (H_REMOVE, H_ENTER,
H_READ, H_CLEAR_MOD, H_CLEAR_REF, H_PROTECT, H_BULK_REMOVE) into a
helper function, kvmppc_pseries_do_hpt_hcall(), and wrapping its call
site in srcu_read_lock(&kvm->srcu) / srcu_read_unlock(&kvm->srcu, idx).
Targeting only the HPT hcalls avoids wrapping non-HPT hcalls that sleep
(such as H_CONFER, H_REGISTER_VPA, H_PAGE_INIT) or handlers that already
acquire SRCU internally (such as H_RTAS).

Fixes: 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual mode handlers for HPT hcalls and page faults")
Cc: stable@vger.kernel.org # v5.14+
Suggested-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Reviewed-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
 arch/powerpc/kvm/book3s_hv.c | 70 ++++++++++++++++++------------------
 1 file changed, 35 insertions(+), 35 deletions(-)

diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index de05d721edc4..46dd550115a4 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -1159,6 +1159,38 @@ static long kvmppc_h_rpt_invalidate(struct kvm_vcpu *vcpu,
 	return H_SUCCESS;
 }
 
+static long kvmppc_pseries_do_hpt_hcall(struct kvm_vcpu *vcpu, unsigned long req)
+{
+	switch (req) {
+	case H_REMOVE:
+		return kvmppc_h_remove(vcpu, kvmppc_get_gpr(vcpu, 4),
+				       kvmppc_get_gpr(vcpu, 5),
+				       kvmppc_get_gpr(vcpu, 6));
+	case H_ENTER:
+		return kvmppc_h_enter(vcpu, kvmppc_get_gpr(vcpu, 4),
+				      kvmppc_get_gpr(vcpu, 5),
+				      kvmppc_get_gpr(vcpu, 6),
+				      kvmppc_get_gpr(vcpu, 7));
+	case H_READ:
+		return kvmppc_h_read(vcpu, kvmppc_get_gpr(vcpu, 4),
+				     kvmppc_get_gpr(vcpu, 5));
+	case H_CLEAR_MOD:
+		return kvmppc_h_clear_mod(vcpu, kvmppc_get_gpr(vcpu, 4),
+					  kvmppc_get_gpr(vcpu, 5));
+	case H_CLEAR_REF:
+		return kvmppc_h_clear_ref(vcpu, kvmppc_get_gpr(vcpu, 4),
+					  kvmppc_get_gpr(vcpu, 5));
+	case H_PROTECT:
+		return kvmppc_h_protect(vcpu, kvmppc_get_gpr(vcpu, 4),
+					kvmppc_get_gpr(vcpu, 5),
+					kvmppc_get_gpr(vcpu, 6));
+	case H_BULK_REMOVE:
+		return kvmppc_h_bulk_remove(vcpu);
+	}
+
+	return H_FUNCTION;
+}
+
 int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
 {
 	struct kvm *kvm = vcpu->kvm;
@@ -1174,47 +1206,15 @@ int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
 
 	switch (req) {
 	case H_REMOVE:
-		ret = kvmppc_h_remove(vcpu, kvmppc_get_gpr(vcpu, 4),
-					kvmppc_get_gpr(vcpu, 5),
-					kvmppc_get_gpr(vcpu, 6));
-		if (ret == H_TOO_HARD)
-			return RESUME_HOST;
-		break;
 	case H_ENTER:
-		ret = kvmppc_h_enter(vcpu, kvmppc_get_gpr(vcpu, 4),
-					kvmppc_get_gpr(vcpu, 5),
-					kvmppc_get_gpr(vcpu, 6),
-					kvmppc_get_gpr(vcpu, 7));
-		if (ret == H_TOO_HARD)
-			return RESUME_HOST;
-		break;
 	case H_READ:
-		ret = kvmppc_h_read(vcpu, kvmppc_get_gpr(vcpu, 4),
-					kvmppc_get_gpr(vcpu, 5));
-		if (ret == H_TOO_HARD)
-			return RESUME_HOST;
-		break;
 	case H_CLEAR_MOD:
-		ret = kvmppc_h_clear_mod(vcpu, kvmppc_get_gpr(vcpu, 4),
-					kvmppc_get_gpr(vcpu, 5));
-		if (ret == H_TOO_HARD)
-			return RESUME_HOST;
-		break;
 	case H_CLEAR_REF:
-		ret = kvmppc_h_clear_ref(vcpu, kvmppc_get_gpr(vcpu, 4),
-					kvmppc_get_gpr(vcpu, 5));
-		if (ret == H_TOO_HARD)
-			return RESUME_HOST;
-		break;
 	case H_PROTECT:
-		ret = kvmppc_h_protect(vcpu, kvmppc_get_gpr(vcpu, 4),
-					kvmppc_get_gpr(vcpu, 5),
-					kvmppc_get_gpr(vcpu, 6));
-		if (ret == H_TOO_HARD)
-			return RESUME_HOST;
-		break;
 	case H_BULK_REMOVE:
-		ret = kvmppc_h_bulk_remove(vcpu);
+		idx = srcu_read_lock(&kvm->srcu);
+		ret = kvmppc_pseries_do_hpt_hcall(vcpu, req);
+		srcu_read_unlock(&kvm->srcu, idx);
 		if (ret == H_TOO_HARD)
 			return RESUME_HOST;
 		break;
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users
  2026-10-06 12:24 [PATCH v3 0/3] KVM: PPC: Fixes for Book3S HV HPT locking and paired-single decoding Amit Machhiwal
  2026-10-06 12:24 ` [PATCH v3 1/3] KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls Amit Machhiwal
@ 2026-10-06 12:24 ` Amit Machhiwal
  2026-10-07  6:18   ` Ritesh Harjani
  2026-10-06 12:24 ` [PATCH v3 3/3] KVM: PPC: Fix get_d_signext() 12-bit displacement for paired-single D-form Amit Machhiwal
  2 siblings, 1 reply; 5+ messages in thread
From: Amit Machhiwal @ 2026-10-06 12:24 UTC (permalink / raw)
  To: Madhavan Srinivasan, linuxppc-dev
  Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
	Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
	Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
	Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
	linux-hardening, stable, Avi Kivity

kvmppc_hv_find_lock_hpte() requires virtual-mode callers to run with
preemption disabled, because it can return with HPTE_V_HVLOCK still held
until the caller later unlocks the HPTE.  Existing virtual-mode callers
in book3s_64_mmu_hv.c already follow that rule, but several paths do
not.

kvmppc_handle_exit_hv() calls kvmppc_hpte_hv_fault() for hash-mode
data-side and instruction-side faults after guest exit with preemption
enabled.  kvmppc_pseries_do_hcall() executes virtual-mode HPT hcall
handlers via kvmppc_pseries_do_hpt_hcall() with preemption enabled; the
handlers for H_ENTER, H_REMOVE, H_READ, H_CLEAR_MOD, H_CLEAR_REF,
H_PROTECT, and H_BULK_REMOVE all spin on try_lock_hpte() or lock_rmap().
H_ENTER also reaches kvmppc_do_h_enter(), which uses arch_spin_lock() on
kvm->mmu_lock.  That raw lock choice is intentional because
kvmppc_do_h_enter() is also called from real-mode paths, so the correct
fix is to establish the proper preemption context at the virtual-mode
caller boundary.

On the host side, kvm_unmap_rmapp(), kvm_age_rmapp(),
kvm_test_clear_dirty_npages(), and resize_hpt_rehash_hpte() also acquire
HPTE_V_HVLOCK via try_lock_hpte() in process context with preemption
enabled, serving MMU notifier callbacks, dirty-log harvesting, and HPT
resize respectively.

If any of these threads is preempted while holding HPTE_V_HVLOCK, any
other thread on the same CPU spinning on the same bit-lock can never
make progress, as the lock owner cannot be rescheduled to release it.
This is particularly acute when the spinning thread has preemption
disabled: it will never yield, causing a permanent CPU hang.

Fix this by adding preempt_disable()/preempt_enable() pairs around the
two kvmppc_hpte_hv_fault() call sites in kvmppc_handle_exit_hv() and
around the kvmppc_pseries_do_hpt_hcall() invocation in
kvmppc_pseries_do_hcall().

For kvm_unmap_rmapp() and kvm_age_rmapp(), place preempt_disable() before
lock_rmap() so that both the rmap chain lock and the subsequent
HPTE_V_HVLOCK bit-lock are held under a single non-preemptible window.
There is an ABBA ordering constraint between the two locks: the rmap chain
lock must be dropped before spinning on the HPTE bit-lock (documented in
the comment above the try_lock_hpte() call in kvm_unmap_rmapp()).  To
preserve this, preempt_enable() is called after unlock_rmap() on the
failed try_lock_hpte() retry path and on any early-exit path, before the
cpu_relax() spin, so the HPTE lock owner can be scheduled.

For kvm_test_clear_dirty_npages(), remove the per-iteration
preempt_disable()/preempt_enable() pairs: this function has a single call
site, kvmppc_hv_get_dirty_log_hpt(), which already holds preempt_disable()
across the entire loop, making the inner guards redundant.

For resize_hpt_rehash_hpte(), place preempt_disable() before the
unconditional try_lock_hpte() spin loop and preempt_enable() after
unlock_hpte() at the single exit point.  This function is called from
kvm_vm_ioctl_resize_hpt_commit(), which first quiesces all vCPUs by
clearing kvm->arch.mmu_ready and calling on_each_cpu() to flush any vCPU
currently running in guest mode back to host.  With all vCPUs out of the
guest, no vCPU thread can hold HPTE_V_HVLOCK; any remaining lock holder
(an MMU notifier callback or dirty-log walker) runs on a separate CPU
and is not preempted, so the spin always makes forward progress.

Fixes: 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual mode handlers for HPT hcalls and page faults")
Cc: stable@vger.kernel.org # v5.14+
Reviewed-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
Changes in v3:
- In kvm_unmap_rmapp() and kvm_age_rmapp(), moved preempt_disable() before
  lock_rmap() so both rmap and HPTE locks are held under a single
  non-preemptible section; added preempt_enable() on early-exit and retry paths.
- Removed redundant per-iteration preempt_disable()/preempt_enable() from
  kvm_test_clear_dirty_npages() as kvmppc_hv_get_dirty_log_hpt() already holds
  it.
- Updated commit log to explain the locking design per function.
- Picked up Reviewed-by tag from Shrikanth Hegde.

Changes in v2:
- Extended preempt_disable()/preempt_enable() to also cover four
  host-side virtual-mode HPTE bit-lock users in book3s_64_mmu_hv.c.
- Added warning comment above kvmppc_pseries_do_hpt_hcall().
- Dropped Reviewed-by as the patch was materially extended.

 arch/powerpc/kvm/book3s_64_mmu_hv.c | 10 ++++++++++
 arch/powerpc/kvm/book3s_hv.c        | 10 ++++++++++
 2 files changed, 20 insertions(+)

diff --git a/arch/powerpc/kvm/book3s_64_mmu_hv.c b/arch/powerpc/kvm/book3s_64_mmu_hv.c
index 2ccb3d138f46..908495f2b001 100644
--- a/arch/powerpc/kvm/book3s_64_mmu_hv.c
+++ b/arch/powerpc/kvm/book3s_64_mmu_hv.c
@@ -810,9 +810,11 @@ static void kvm_unmap_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
 
 	rmapp = &memslot->arch.rmap[gfn - memslot->base_gfn];
 	for (;;) {
+		preempt_disable();
 		lock_rmap(rmapp);
 		if (!(*rmapp & KVMPPC_RMAP_PRESENT)) {
 			unlock_rmap(rmapp);
+			preempt_enable();
 			break;
 		}
 
@@ -826,6 +828,7 @@ static void kvm_unmap_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
 		if (!try_lock_hpte(hptep, HPTE_V_HVLOCK)) {
 			/* unlock rmap before spinning on the HPTE lock */
 			unlock_rmap(rmapp);
+			preempt_enable();
 			while (be64_to_cpu(hptep[0]) & HPTE_V_HVLOCK)
 				cpu_relax();
 			continue;
@@ -834,6 +837,7 @@ static void kvm_unmap_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
 		kvmppc_unmap_hpte(kvm, i, memslot, rmapp, gfn);
 		unlock_rmap(rmapp);
 		__unlock_hpte(hptep, be64_to_cpu(hptep[0]));
+		preempt_enable();
 	}
 }
 
@@ -890,6 +894,7 @@ static bool kvm_age_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
 
 	rmapp = &memslot->arch.rmap[gfn - memslot->base_gfn];
  retry:
+	preempt_disable();
 	lock_rmap(rmapp);
 	if (*rmapp & KVMPPC_RMAP_REFERENCED) {
 		*rmapp &= ~KVMPPC_RMAP_REFERENCED;
@@ -897,6 +902,7 @@ static bool kvm_age_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
 	}
 	if (!(*rmapp & KVMPPC_RMAP_PRESENT)) {
 		unlock_rmap(rmapp);
+		preempt_enable();
 		return ret;
 	}
 
@@ -912,6 +918,7 @@ static bool kvm_age_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
 		if (!try_lock_hpte(hptep, HPTE_V_HVLOCK)) {
 			/* unlock rmap before spinning on the HPTE lock */
 			unlock_rmap(rmapp);
+			preempt_enable();
 			while (be64_to_cpu(hptep[0]) & HPTE_V_HVLOCK)
 				cpu_relax();
 			goto retry;
@@ -931,6 +938,7 @@ static bool kvm_age_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
 	} while ((i = j) != head);
 
 	unlock_rmap(rmapp);
+	preempt_enable();
 	return ret;
 }
 
@@ -1219,6 +1227,7 @@ static unsigned long resize_hpt_rehash_hpte(struct kvm_resize_hpt *resize,
 	if (!(vpte & HPTE_V_VALID) && !(vpte & HPTE_V_ABSENT))
 		return 0; /* nothing to do */
 
+	preempt_disable();
 	while (!try_lock_hpte(hptep, HPTE_V_HVLOCK))
 		cpu_relax();
 
@@ -1346,6 +1355,7 @@ static unsigned long resize_hpt_rehash_hpte(struct kvm_resize_hpt *resize,
 
 out:
 	unlock_hpte(hptep, vpte);
+	preempt_enable();
 	return ret;
 }
 
diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index 46dd550115a4..56083b415a29 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -1159,6 +1159,10 @@ static long kvmppc_h_rpt_invalidate(struct kvm_vcpu *vcpu,
 	return H_SUCCESS;
 }
 
+/*
+ * Must be called with preemption disabled. The HPT hcall handlers spin
+ * on HPTE bit-locks and cannot make any blocking/sleeping calls.
+ */
 static long kvmppc_pseries_do_hpt_hcall(struct kvm_vcpu *vcpu, unsigned long req)
 {
 	switch (req) {
@@ -1212,9 +1216,11 @@ int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
 	case H_CLEAR_REF:
 	case H_PROTECT:
 	case H_BULK_REMOVE:
+		preempt_disable();
 		idx = srcu_read_lock(&kvm->srcu);
 		ret = kvmppc_pseries_do_hpt_hcall(vcpu, req);
 		srcu_read_unlock(&kvm->srcu, idx);
+		preempt_enable();
 		if (ret == H_TOO_HARD)
 			return RESUME_HOST;
 		break;
@@ -1834,8 +1840,10 @@ static int kvmppc_handle_exit_hv(struct kvm_vcpu *vcpu,
 		else
 			vsid = vcpu->arch.fault_gpa;
 
+		preempt_disable();
 		err = kvmppc_hpte_hv_fault(vcpu, vcpu->arch.fault_dar,
 				vsid, vcpu->arch.fault_dsisr, true);
+		preempt_enable();
 		if (err == 0) {
 			r = RESUME_GUEST;
 		} else if (err == -1 || err == -2) {
@@ -1881,8 +1889,10 @@ static int kvmppc_handle_exit_hv(struct kvm_vcpu *vcpu,
 		else
 			vsid = vcpu->arch.fault_gpa;
 
+		preempt_disable();
 		err = kvmppc_hpte_hv_fault(vcpu, vcpu->arch.fault_dar,
 				vsid, vcpu->arch.fault_dsisr, false);
+		preempt_enable();
 		if (err == 0) {
 			r = RESUME_GUEST;
 		} else if (err == -1) {
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v3 3/3] KVM: PPC: Fix get_d_signext() 12-bit displacement for paired-single D-form
  2026-10-06 12:24 [PATCH v3 0/3] KVM: PPC: Fixes for Book3S HV HPT locking and paired-single decoding Amit Machhiwal
  2026-10-06 12:24 ` [PATCH v3 1/3] KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls Amit Machhiwal
  2026-10-06 12:24 ` [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users Amit Machhiwal
@ 2026-10-06 12:24 ` Amit Machhiwal
  2 siblings, 0 replies; 5+ messages in thread
From: Amit Machhiwal @ 2026-10-06 12:24 UTC (permalink / raw)
  To: Madhavan Srinivasan, linuxppc-dev
  Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
	Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
	Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
	Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
	linux-hardening, stable, Avi Kivity

get_d_signext() has two compounding bugs since its introduction in 2010:

1. The extraction mask 0x8FF silently drops bits 8-10 of the 12-bit D
   field (PPC ISA bits 21-23), corrupting any displacement that has any
   of those bits set.

2. The sign-magnitude idiom "-(d & 0x7ff)" is wrong for two's-complement:
   for D=0xFFC (encoding of -4) it returns -252 instead of -4.

Together these errors produce a wrong effective address for psq_l, psq_lu,
psq_st, and psq_stu whenever the displacement is negative or is a positive
value >= 0x100 with bits 8-10 set — essentially any real-world paired-
single stack-relative access.

The D field occupies the bottom 12 bits of the instruction word (confirmed
by the adjacent W and I extractions via inst_get_field(inst,16,16) and
inst_get_field(inst,17,19)). Replace the open-coded logic with the standard
sign_extend32(inst & 0xfff, 11), which correctly performs two's-complement
sign extension from 12 bits to 32 bits.

Fixes: 831317b605e7 ("KVM: PPC: Implement Paired Single emulation")
Cc: stable@vger.kernel.org
Reviewed-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
Changes in v3:
- Picked up Reviewed-by tag from Shrikanth Hegde.

 arch/powerpc/kvm/book3s_paired_singles.c | 7 +------
 1 file changed, 1 insertion(+), 6 deletions(-)

diff --git a/arch/powerpc/kvm/book3s_paired_singles.c b/arch/powerpc/kvm/book3s_paired_singles.c
index bc39c76c9d9f..532f96293de0 100644
--- a/arch/powerpc/kvm/book3s_paired_singles.c
+++ b/arch/powerpc/kvm/book3s_paired_singles.c
@@ -479,12 +479,7 @@ static bool kvmppc_inst_is_paired_single(struct kvm_vcpu *vcpu, u32 inst)
 
 static int get_d_signext(u32 inst)
 {
-	int d = inst & 0x8ff;
-
-	if (d & 0x800)
-		return -(d & 0x7ff);
-
-	return (d & 0x7ff);
+	return sign_extend32(inst & 0xfff, 11);
 }
 
 static int kvmppc_ps_three_in(struct kvm_vcpu *vcpu, bool rc,
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users
  2026-10-06 12:24 ` [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users Amit Machhiwal
@ 2026-10-07  6:18   ` Ritesh Harjani
  0 siblings, 0 replies; 5+ messages in thread
From: Ritesh Harjani @ 2026-10-07  6:18 UTC (permalink / raw)
  To: Amit Machhiwal, Madhavan Srinivasan, linuxppc-dev
  Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
	Christophe Leroy (CS GROUP),
	Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
	Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
	linux-hardening, stable, Avi Kivity

Amit Machhiwal <amachhiw@linux.ibm.com> writes:

> kvmppc_hv_find_lock_hpte() requires virtual-mode callers to run with
> preemption disabled, because it can return with HPTE_V_HVLOCK still held
> until the caller later unlocks the HPTE.  Existing virtual-mode callers
> in book3s_64_mmu_hv.c already follow that rule, but several paths do
> not.
>
> kvmppc_handle_exit_hv() calls kvmppc_hpte_hv_fault() for hash-mode
> data-side and instruction-side faults after guest exit with preemption
> enabled.  kvmppc_pseries_do_hcall() executes virtual-mode HPT hcall
> handlers via kvmppc_pseries_do_hpt_hcall() with preemption enabled; the
> handlers for H_ENTER, H_REMOVE, H_READ, H_CLEAR_MOD, H_CLEAR_REF,
> H_PROTECT, and H_BULK_REMOVE all spin on try_lock_hpte() or lock_rmap().
> H_ENTER also reaches kvmppc_do_h_enter(), which uses arch_spin_lock() on
> kvm->mmu_lock.  That raw lock choice is intentional because
> kvmppc_do_h_enter() is also called from real-mode paths, so the correct
> fix is to establish the proper preemption context at the virtual-mode
> caller boundary.
>
> On the host side, kvm_unmap_rmapp(), kvm_age_rmapp(),
> kvm_test_clear_dirty_npages(), and resize_hpt_rehash_hpte() also acquire
> HPTE_V_HVLOCK via try_lock_hpte() in process context with preemption
> enabled, serving MMU notifier callbacks, dirty-log harvesting, and HPT
> resize respectively.
>

I was going over all the callers of lock_rmap() and try_lock_hpte() on,
and I see that we might have missed kvm_htab_write() path...

... after spending sometime looks like we need this diff for
kvm_htab_write() path as well, since it calls kvmppc_do_h_remove() which
calls try_lock_hpte() and lock_rmap(), although the race window is much
narrower and maybe very hard to hit.

But still, could you kindly look into this and if needed please take it
forward too. Please note that this is not tested, so hoping that you
could take care of that too.


Thanks!
-ritesh


diff --git a/arch/powerpc/kvm/book3s_64_mmu_hv.c b/arch/powerpc/kvm/book3s_64_mmu_hv.c
index 908495f2b001b..916db7ceb51a9 100644
--- a/arch/powerpc/kvm/book3s_64_mmu_hv.c
+++ b/arch/powerpc/kvm/book3s_64_mmu_hv.c
@@ -47,6 +47,8 @@
 static long kvmppc_virtmode_do_h_enter(struct kvm *kvm, unsigned long flags,
                                long pte_index, unsigned long pteh,
                                unsigned long ptel, unsigned long *pte_idx_ret);
+static void kvmppc_virtmode_do_h_remove(struct kvm *kvm,
+                               unsigned long pte_index, unsigned long *hpret);

 struct kvm_resize_hpt {
        /* These fields read-only after init */
@@ -308,6 +310,19 @@ static long kvmppc_virtmode_do_h_enter(struct kvm *kvm, unsigned long flags,

 }

+/*
+ * Virtual-mode H_REMOVE.  kvmppc_do_h_remove() is also called from real
+ * mode, where preempt_disable() is not usable, so the guard stays here.
+ * The helper takes HPTE_V_HVLOCK and the rmap bit and does not sleep.
+ */
+static void kvmppc_virtmode_do_h_remove(struct kvm *kvm,
+                               unsigned long pte_index, unsigned long *hpret)
+{
+       preempt_disable();
+       kvmppc_do_h_remove(kvm, 0, pte_index, 0, hpret);
+       preempt_enable();
+}
+
 static struct kvmppc_slb *kvmppc_mmu_book3s_hv_find_slbe(struct kvm_vcpu *vcpu,
                                                         gva_t eaddr)
 {
@@ -1878,7 +1893,7 @@ static ssize_t kvm_htab_write(struct file *file, const char __user *buf,
                        nb += HPTE_SIZE;

                        if (be64_to_cpu(hptp[0]) & (HPTE_V_VALID | HPTE_V_ABSENT))
-                               kvmppc_do_h_remove(kvm, 0, i, 0, tmp);
+                               kvmppc_virtmode_do_h_remove(kvm, i, tmp);
                        err = -EIO;
                        ret = kvmppc_virtmode_do_h_enter(kvm, H_EXACT, i, v, r,
                                                         tmp);
@@ -1907,7 +1922,7 @@ static ssize_t kvm_htab_write(struct file *file, const char __user *buf,

                for (j = 0; j < hdr.n_invalid; ++j) {
                        if (be64_to_cpu(hptp[0]) & (HPTE_V_VALID | HPTE_V_ABSENT))
-                               kvmppc_do_h_remove(kvm, 0, i, 0, tmp);
+                               kvmppc_virtmode_do_h_remove(kvm, i, tmp);
                        ++i;
                        hptp += 2;
                }



^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-07  7:01 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 12:24 [PATCH v3 0/3] KVM: PPC: Fixes for Book3S HV HPT locking and paired-single decoding Amit Machhiwal
2026-10-06 12:24 ` [PATCH v3 1/3] KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls Amit Machhiwal
2026-10-06 12:24 ` [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users Amit Machhiwal
2026-10-07  6:18   ` Ritesh Harjani
2026-10-06 12:24 ` [PATCH v3 3/3] KVM: PPC: Fix get_d_signext() 12-bit displacement for paired-single D-form Amit Machhiwal

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®