mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
@ 2026-09-21 11:10 Gautam Menghani
  2026-09-23  3:38 ` Narayana Murty N
                   ` (2 more replies)
  0 siblings, 3 replies; 7+ messages in thread
From: Gautam Menghani @ 2026-09-21 11:10 UTC (permalink / raw)
  To: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux, amachhiw
  Cc: linuxppc-dev, kvm, linux-kernel, stable, Timothy Pearson

A huge number of spurious interrupts can be seen immediately after a KVM
on PowerNV guest boots up in XIVE mode.

$ cat /proc/interrupts  | grep SPU
SPU:     223705     192439     273526     147623   Spurious interrupts

This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
Handle pending exceptions on guest entry with MSR_EE"). The root cause
is that once LPCR_MER bit is set, it is supposed to be reset by
software. But when a vCPU starts running with LPCR_MER set, the vCPU does
not exit back to the host until the decrementer expires or there is an
hcall, etc. This is because KVM on PowerNV guests have support for
native XIVE, so they are not dependent on host for interrupt emulation.
Due to this behaviour, a huge number of spurious interrupts are seen
since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
vCPU exits to the host.

Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
available (currently in case of KVM on PowerNV only), as the XIVE hardware
can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
functionality is not required. This reduces the number of spurious
interrupts drastically.

Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
Cc: stable@vger.kernel.org # 6.8+
Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
---
v3:
1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)

v2:
1. Handle the case where xive_interrupt_pending() is true and also the
external exception bit is set. (Narayana)

 arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
 arch/powerpc/kvm/book3s_hv.c       | 2 +-
 2 files changed, 8 insertions(+), 1 deletion(-)

diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
index 169ea6a7fbad..580ad2548c2b 100644
--- a/arch/powerpc/include/asm/kvm_ppc.h
+++ b/arch/powerpc/include/asm/kvm_ppc.h
@@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
 	return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
 }
 
+static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
+{
+	return kvm->arch.xive_devices.native;
+}
+
 extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
 					   struct kvm_vcpu *vcpu, u32 cpu);
 extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
@@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
 
 static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
 	{ return 0; }
+static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
+
 static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
 			  struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
 static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index dbac3573b2c8..cb2bb29451a7 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
 			if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
 				kvmppc_inject_interrupt_hv(vcpu,
 							   BOOK3S_INTERRUPT_EXTERNAL, 0);
-			else
+			else if (!kvmppc_xive_native_enabled(vcpu->kvm))
 				lpcr |= LPCR_MER;
 		} else {
 			/*
-- 
2.55.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
  2026-09-21 11:10 [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit Gautam Menghani
@ 2026-09-23  3:38 ` Narayana Murty N
  2026-10-01  3:58   ` Gautam Menghani
  2026-09-23  7:01 ` Mukesh Kumar Chaurasiya
  2026-10-05  8:00 ` Amit Machhiwal
  2 siblings, 1 reply; 7+ messages in thread
From: Narayana Murty N @ 2026-09-23  3:38 UTC (permalink / raw)
  To: Gautam Menghani, maddy, npiggin, mpe, chleroy, ritesh.list,
	sshegde, amachhiw
  Cc: linuxppc-dev, kvm, linux-kernel, stable, Timothy Pearson

Hi Gautam,

On 21/09/26 4:40 PM, Gautam Menghani wrote:
> A huge number of spurious interrupts can be seen immediately after a KVM
> on PowerNV guest boots up in XIVE mode.
> 
> $ cat /proc/interrupts  | grep SPU
> SPU:     223705     192439     273526     147623   Spurious interrupts
> 
> This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> Handle pending exceptions on guest entry with MSR_EE"). The root cause
> is that once LPCR_MER bit is set, it is supposed to be reset by
> software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> not exit back to the host until the decrementer expires or there is an
> hcall, etc. This is because KVM on PowerNV guests have support for
> native XIVE, so they are not dependent on host for interrupt emulation.
> Due to this behaviour, a huge number of spurious interrupts are seen
> since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> vCPU exits to the host.
> 
> Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> available (currently in case of KVM on PowerNV only), as the XIVE hardware
> can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> functionality is not required. This reduces the number of spurious
> interrupts drastically.
> 
> Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> Cc: stable@vger.kernel.org # 6.8+
> Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
> Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
> ---
> v3:
> 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
> 
> v2:
> 1. Handle the case where xive_interrupt_pending() is true and also the
> external exception bit is set. (Narayana)
> 
>   arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
>   arch/powerpc/kvm/book3s_hv.c       | 2 +-
>   2 files changed, 8 insertions(+), 1 deletion(-)
> 
> diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> index 169ea6a7fbad..580ad2548c2b 100644
> --- a/arch/powerpc/include/asm/kvm_ppc.h
> +++ b/arch/powerpc/include/asm/kvm_ppc.h
> @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
>   	return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
>   }
>   
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> +{
> +	return kvm->arch.xive_devices.native;
> +}
Also, should 'kvmppc_xive_native_enabled()' check the active XIVE
device rather than just whether a native device has been created?

'kvmppc_xive_native_release()' clears 'kvm->arch.xive', but explicitly
keeps the 'kvmppc_xive' pointer under 'xive_devices' for reuse. The
comment in 'book3s_xive.c' also says that when switching between XICS
and native XIVE, the previous KVM device is released before the new one
is created.

So 'xive_devices.native != NULL' seems to mean that a native XIVE
device has been allocated/cached, rather than that it is currently
active.

Would this be more appropriate?

return kvm->arch.xive_devices.native &&
        kvm->arch.xive == kvm->arch.xive_devices.native;

> +
>   extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
>   					   struct kvm_vcpu *vcpu, u32 cpu);
>   extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
>   
>   static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
>   	{ return 0; }
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> +
>   static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
>   			  struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
>   static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> index dbac3573b2c8..cb2bb29451a7 100644
> --- a/arch/powerpc/kvm/book3s_hv.c
> +++ b/arch/powerpc/kvm/book3s_hv.c
> @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
>   			if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
>   				kvmppc_inject_interrupt_hv(vcpu,
>   							   BOOK3S_INTERRUPT_EXTERNAL, 0);
> -			else
> +			else if (!kvmppc_xive_native_enabled(vcpu->kvm))
>   				lpcr |= LPCR_MER;
Is the intent here to suppress LPCR_MER only for interrupts that
native XIVE can deliver directly?

BOOK3S_IRQPRIO_EXTERNAL is a separate software-pending exception. If
MSR_EE is clear it remains pending, so it seems this path should still
set MER even when native XIVE is active.

IOW, should the behavior be:

else if (!kvmppc_xive_native_enabled(vcpu->kvm) ||
	 test_bit(BOOK3S_IRQPRIO_EXTERNAL,
		  &vcpu->arch.pending_exceptions))
	lpcr |= LPCR_MER;

so that only the native-XIVE-pending case suppresses MER?

Thanks,
narayana Murty N
>   		} else {
>   			/*


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
  2026-09-21 11:10 [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit Gautam Menghani
  2026-09-23  3:38 ` Narayana Murty N
@ 2026-09-23  7:01 ` Mukesh Kumar Chaurasiya
  2026-09-23 12:50   ` Gautam Menghani
  2026-10-05  8:00 ` Amit Machhiwal
  2 siblings, 1 reply; 7+ messages in thread
From: Mukesh Kumar Chaurasiya @ 2026-09-23  7:01 UTC (permalink / raw)
  To: Gautam Menghani
  Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux,
	amachhiw, linuxppc-dev, kvm, linux-kernel, stable,
	Timothy Pearson

On Mon, Sep 21, 2026 at 04:40:42PM +0530, Gautam Menghani wrote:
> A huge number of spurious interrupts can be seen immediately after a KVM
> on PowerNV guest boots up in XIVE mode.
> 
> $ cat /proc/interrupts  | grep SPU
> SPU:     223705     192439     273526     147623   Spurious interrupts
> 
> This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> Handle pending exceptions on guest entry with MSR_EE"). The root cause
> is that once LPCR_MER bit is set, it is supposed to be reset by
> software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> not exit back to the host until the decrementer expires or there is an
> hcall, etc. This is because KVM on PowerNV guests have support for
> native XIVE, so they are not dependent on host for interrupt emulation.
> Due to this behaviour, a huge number of spurious interrupts are seen
> since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> vCPU exits to the host.
> 
> Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> available (currently in case of KVM on PowerNV only), as the XIVE hardware
> can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> functionality is not required. This reduces the number of spurious
> interrupts drastically.
Hey Gautam,

Thanks for this. Can you also share the numbers after the change.

Regards,
Mukesh
[...]

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
  2026-09-23  7:01 ` Mukesh Kumar Chaurasiya
@ 2026-09-23 12:50   ` Gautam Menghani
  0 siblings, 0 replies; 7+ messages in thread
From: Gautam Menghani @ 2026-09-23 12:50 UTC (permalink / raw)
  To: Mukesh Kumar Chaurasiya
  Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux,
	amachhiw, linuxppc-dev, kvm, linux-kernel, stable,
	Timothy Pearson

On Wed, Sep 23, 2026 at 12:31:05PM +0530, Mukesh Kumar Chaurasiya wrote:
> On Mon, Sep 21, 2026 at 04:40:42PM +0530, Gautam Menghani wrote:
> > A huge number of spurious interrupts can be seen immediately after a KVM
> > on PowerNV guest boots up in XIVE mode.
> > 
> > $ cat /proc/interrupts  | grep SPU
> > SPU:     223705     192439     273526     147623   Spurious interrupts
> > 
> > This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> > Handle pending exceptions on guest entry with MSR_EE"). The root cause
> > is that once LPCR_MER bit is set, it is supposed to be reset by
> > software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> > not exit back to the host until the decrementer expires or there is an
> > hcall, etc. This is because KVM on PowerNV guests have support for
> > native XIVE, so they are not dependent on host for interrupt emulation.
> > Due to this behaviour, a huge number of spurious interrupts are seen
> > since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> > vCPU exits to the host.
> > 
> > Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> > available (currently in case of KVM on PowerNV only), as the XIVE hardware
> > can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> > functionality is not required. This reduces the number of spurious
> > interrupts drastically.
> Hey Gautam,
> 
> Thanks for this. Can you also share the numbers after the change.
> 

Sure, here are the numbers seen immediately on boot with this patch
applied:

$ cat /proc/interrupts | grep SPU
SPU:        123        139         99        144   Spurious interrupts

> Regards,
> Mukesh
> [...]

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
  2026-09-23  3:38 ` Narayana Murty N
@ 2026-10-01  3:58   ` Gautam Menghani
  0 siblings, 0 replies; 7+ messages in thread
From: Gautam Menghani @ 2026-10-01  3:58 UTC (permalink / raw)
  To: Narayana Murty N
  Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, amachhiw,
	linuxppc-dev, kvm, linux-kernel, stable, Timothy Pearson

On Wed, Sep 23, 2026 at 09:08:57AM +0530, Narayana Murty N wrote:
> Hi Gautam,
> 
> On 21/09/26 4:40 PM, Gautam Menghani wrote:
> > A huge number of spurious interrupts can be seen immediately after a KVM
> > on PowerNV guest boots up in XIVE mode.
> > 
> > $ cat /proc/interrupts  | grep SPU
> > SPU:     223705     192439     273526     147623   Spurious interrupts
> > 
> > This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> > Handle pending exceptions on guest entry with MSR_EE"). The root cause
> > is that once LPCR_MER bit is set, it is supposed to be reset by
> > software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> > not exit back to the host until the decrementer expires or there is an
> > hcall, etc. This is because KVM on PowerNV guests have support for
> > native XIVE, so they are not dependent on host for interrupt emulation.
> > Due to this behaviour, a huge number of spurious interrupts are seen
> > since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> > vCPU exits to the host.
> > 
> > Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> > available (currently in case of KVM on PowerNV only), as the XIVE hardware
> > can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> > functionality is not required. This reduces the number of spurious
> > interrupts drastically.
> > 
> > Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> > Cc: stable@vger.kernel.org # 6.8+
> > Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
> > Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> > Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
> > ---
> > v3:
> > 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
> > 
> > v2:
> > 1. Handle the case where xive_interrupt_pending() is true and also the
> > external exception bit is set. (Narayana)
> > 
> >   arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
> >   arch/powerpc/kvm/book3s_hv.c       | 2 +-
> >   2 files changed, 8 insertions(+), 1 deletion(-)
> > 
> > diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> > index 169ea6a7fbad..580ad2548c2b 100644
> > --- a/arch/powerpc/include/asm/kvm_ppc.h
> > +++ b/arch/powerpc/include/asm/kvm_ppc.h
> > @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> >   	return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> >   }
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> > +{
> > +	return kvm->arch.xive_devices.native;
> > +}
> Also, should 'kvmppc_xive_native_enabled()' check the active XIVE
> device rather than just whether a native device has been created?
> 
> 'kvmppc_xive_native_release()' clears 'kvm->arch.xive', but explicitly
> keeps the 'kvmppc_xive' pointer under 'xive_devices' for reuse. The
> comment in 'book3s_xive.c' also says that when switching between XICS
> and native XIVE, the previous KVM device is released before the new one
> is created.
> 
> So 'xive_devices.native != NULL' seems to mean that a native XIVE
> device has been allocated/cached, rather than that it is currently
> active.


In a KVM guest on PowerNV, whatever interrupt mode the guest is booted
in (XIVE/ XICS), you cannot switch it with kexec. So if
'xive_devices.native' exists, we can be confident that native XIVE is
being used. So the code in the patch should be fine.

> 
> Would this be more appropriate?
> 
> return kvm->arch.xive_devices.native &&
>        kvm->arch.xive == kvm->arch.xive_devices.native;
> 
> > +
> >   extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> >   					   struct kvm_vcpu *vcpu, u32 cpu);
> >   extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> > @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
> >   static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> >   	{ return 0; }
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> > +
> >   static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> >   			  struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
> >   static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> > diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> > index dbac3573b2c8..cb2bb29451a7 100644
> > --- a/arch/powerpc/kvm/book3s_hv.c
> > +++ b/arch/powerpc/kvm/book3s_hv.c
> > @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> >   			if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> >   				kvmppc_inject_interrupt_hv(vcpu,
> >   							   BOOK3S_INTERRUPT_EXTERNAL, 0);
> > -			else
> > +			else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> >   				lpcr |= LPCR_MER;
> Is the intent here to suppress LPCR_MER only for interrupts that
> native XIVE can deliver directly?
> 
> BOOK3S_IRQPRIO_EXTERNAL is a separate software-pending exception. If
> MSR_EE is clear it remains pending, so it seems this path should still
> set MER even when native XIVE is active.
> 
> IOW, should the behavior be:
> 
> else if (!kvmppc_xive_native_enabled(vcpu->kvm) ||
> 	 test_bit(BOOK3S_IRQPRIO_EXTERNAL,
> 		  &vcpu->arch.pending_exceptions))
> 	lpcr |= LPCR_MER;
> 
> so that only the native-XIVE-pending case suppresses MER?

No this would still result in spurious flood. If native XIVE is being
used, LPCR_MER cannot be set in any case. Software interrupts will
continue to work in the same way they work in an LPAR.

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
  2026-09-21 11:10 [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit Gautam Menghani
  2026-09-23  3:38 ` Narayana Murty N
  2026-09-23  7:01 ` Mukesh Kumar Chaurasiya
@ 2026-10-05  8:00 ` Amit Machhiwal
  2026-10-05  8:31   ` Gautam Menghani
  2 siblings, 1 reply; 7+ messages in thread
From: Amit Machhiwal @ 2026-10-05  8:00 UTC (permalink / raw)
  To: Gautam Menghani
  Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux,
	amachhiw, linuxppc-dev, kvm, linux-kernel, stable,
	Timothy Pearson

Hi Gautam,

Thanks for the patch and working on this.  I have a couple of questions though.

On 2026/09/21 04:40 PM, Gautam Menghani wrote:
> A huge number of spurious interrupts can be seen immediately after a KVM
> on PowerNV guest boots up in XIVE mode.
> 
> $ cat /proc/interrupts  | grep SPU
> SPU:     223705     192439     273526     147623   Spurious interrupts
> 
> This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> Handle pending exceptions on guest entry with MSR_EE"). The root cause
> is that once LPCR_MER bit is set, it is supposed to be reset by
> software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> not exit back to the host until the decrementer expires or there is an
> hcall, etc. This is because KVM on PowerNV guests have support for
> native XIVE, so they are not dependent on host for interrupt emulation.
> Due to this behaviour, a huge number of spurious interrupts are seen
> since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> vCPU exits to the host.
> 
> Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> available (currently in case of KVM on PowerNV only), as the XIVE hardware
> can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> functionality is not required. This reduces the number of spurious
> interrupts drastically.
> 
> Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> Cc: stable@vger.kernel.org # 6.8+
> Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
> Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
> ---
> v3:
> 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
> 
> v2:
> 1. Handle the case where xive_interrupt_pending() is true and also the
> external exception bit is set. (Narayana)
> 
>  arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
>  arch/powerpc/kvm/book3s_hv.c       | 2 +-
>  2 files changed, 8 insertions(+), 1 deletion(-)
> 
> diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> index 169ea6a7fbad..580ad2548c2b 100644
> --- a/arch/powerpc/include/asm/kvm_ppc.h
> +++ b/arch/powerpc/include/asm/kvm_ppc.h
> @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
>  	return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
>  }
>  
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> +{
> +	return kvm->arch.xive_devices.native;
> +}
> +
>  extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
>  					   struct kvm_vcpu *vcpu, u32 cpu);
>  extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
>  
>  static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
>  	{ return 0; }
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> +
>  static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
>  			  struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
>  static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> index dbac3573b2c8..cb2bb29451a7 100644
> --- a/arch/powerpc/kvm/book3s_hv.c
> +++ b/arch/powerpc/kvm/book3s_hv.c
> @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
>  			if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
>  				kvmppc_inject_interrupt_hv(vcpu,
>  							   BOOK3S_INTERRUPT_EXTERNAL, 0);
> -			else
> +			else if (!kvmppc_xive_native_enabled(vcpu->kvm))
>  				lpcr |= LPCR_MER;

1. What happens when an L1 KVM guest is booted with (native) XIVE and then the
   guest is rebooted with `xive=off` i.e., with XICS?  Will we stop setting
   LPCR_MER as kvm->arch.xive_devices.native would still be set?
2. What happends when an L1 KVM guest is booted with XIVE and then we kexec into
   a new kernel with `xive=off`?  You did mention in your other reply that with
   kexec, `xive=off` is ignored currently but IMO, we would want to understand
   where this limiation lies and fix that if need be.  I do understand we are
   trying to fix spurious interrupts problem with XIVE in this patch and this
   particular problem can be taken separately but it worth investigating.

Thanks,
Amit

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
  2026-10-05  8:00 ` Amit Machhiwal
@ 2026-10-05  8:31   ` Gautam Menghani
  0 siblings, 0 replies; 7+ messages in thread
From: Gautam Menghani @ 2026-10-05  8:31 UTC (permalink / raw)
  To: Amit Machhiwal
  Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux,
	linuxppc-dev, kvm, linux-kernel, stable, Timothy Pearson

On Mon, Oct 05, 2026 at 01:30:59PM +0530, Amit Machhiwal wrote:
> Hi Gautam,
> 
> Thanks for the patch and working on this.  I have a couple of questions though.
> 
> On 2026/09/21 04:40 PM, Gautam Menghani wrote:
> > A huge number of spurious interrupts can be seen immediately after a KVM
> > on PowerNV guest boots up in XIVE mode.
> > 
> > $ cat /proc/interrupts  | grep SPU
> > SPU:     223705     192439     273526     147623   Spurious interrupts
> > 
> > This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> > Handle pending exceptions on guest entry with MSR_EE"). The root cause
> > is that once LPCR_MER bit is set, it is supposed to be reset by
> > software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> > not exit back to the host until the decrementer expires or there is an
> > hcall, etc. This is because KVM on PowerNV guests have support for
> > native XIVE, so they are not dependent on host for interrupt emulation.
> > Due to this behaviour, a huge number of spurious interrupts are seen
> > since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> > vCPU exits to the host.
> > 
> > Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> > available (currently in case of KVM on PowerNV only), as the XIVE hardware
> > can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> > functionality is not required. This reduces the number of spurious
> > interrupts drastically.
> > 
> > Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> > Cc: stable@vger.kernel.org # 6.8+
> > Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
> > Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> > Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
> > ---
> > v3:
> > 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
> > 
> > v2:
> > 1. Handle the case where xive_interrupt_pending() is true and also the
> > external exception bit is set. (Narayana)
> > 
> >  arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
> >  arch/powerpc/kvm/book3s_hv.c       | 2 +-
> >  2 files changed, 8 insertions(+), 1 deletion(-)
> > 
> > diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> > index 169ea6a7fbad..580ad2548c2b 100644
> > --- a/arch/powerpc/include/asm/kvm_ppc.h
> > +++ b/arch/powerpc/include/asm/kvm_ppc.h
> > @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> >  	return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> >  }
> >  
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> > +{
> > +	return kvm->arch.xive_devices.native;
> > +}
> > +
> >  extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> >  					   struct kvm_vcpu *vcpu, u32 cpu);
> >  extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> > @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
> >  
> >  static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> >  	{ return 0; }
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> > +
> >  static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> >  			  struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
> >  static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> > diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> > index dbac3573b2c8..cb2bb29451a7 100644
> > --- a/arch/powerpc/kvm/book3s_hv.c
> > +++ b/arch/powerpc/kvm/book3s_hv.c
> > @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> >  			if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> >  				kvmppc_inject_interrupt_hv(vcpu,
> >  							   BOOK3S_INTERRUPT_EXTERNAL, 0);
> > -			else
> > +			else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> >  				lpcr |= LPCR_MER;
> 
> 1. What happens when an L1 KVM guest is booted with (native) XIVE and then the
>    guest is rebooted with `xive=off` i.e., with XICS?  Will we stop setting
>    LPCR_MER as kvm->arch.xive_devices.native would still be set?

Yes, that's a good catch. kvm->arch.xive_devices.native is cleared only
when vm is destroyed. So this case has to be handled.

> 2. What happends when an L1 KVM guest is booted with XIVE and then we kexec into
>    a new kernel with `xive=off`?  You did mention in your other reply that with
>    kexec, `xive=off` is ignored currently but IMO, we would want to understand
>    where this limiation lies and fix that if need be.  I do understand we are
>    trying to fix spurious interrupts problem with XIVE in this patch and this
>    particular problem can be taken separately but it worth investigating.

Yes right, the kexec case is to be understood, but I'll take it up separately.
Meanwhile, I'll send a v4 to fix the reboot case.

> 
> Thanks,
> Amit

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-10-05  8:31 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-21 11:10 [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit Gautam Menghani
2026-09-23  3:38 ` Narayana Murty N
2026-10-01  3:58   ` Gautam Menghani
2026-09-23  7:01 ` Mukesh Kumar Chaurasiya
2026-09-23 12:50   ` Gautam Menghani
2026-10-05  8:00 ` Amit Machhiwal
2026-10-05  8:31   ` Gautam Menghani

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®