mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] *** cpu/hotplug, riscv: Fix CPU stuck offline after failed hotplug attempt ***
@ 2026-09-27 16:44 artemiispatov
  2026-09-27 16:44 ` [PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE artemiispatov
  2026-09-27 16:44 ` [PATCH 2/2] riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start artemiispatov
  0 siblings, 2 replies; 3+ messages in thread
From: artemiispatov @ 2026-09-27 16:44 UTC (permalink / raw)
  To: linux-kernel, linux-riscv
  Cc: tglx, peterz, anup, atishp, palmer, pjw, aou, alex, conor,
	hui.wang, artemiispatov

From: Artemii Patov <artemiispatov@gmail.com>

These two patches enable RISC-V to retry bringing an AP
online when the control CPU failed to complete the
SYNC_STATE_ALIVE -> SYNC_STATE_SHOULD_ONLINE transition and
the secondary CPU is stuck spinning in cpuhp_ap_sync_alive().

The need for this arose during debugging RISC-V SoC. Control CPU tried
to wake up a halted CPU and sent an IPI via OpenSBI. The IPI was masked
while the secondary CPU was in Debug Mode; when the CPU left Debug Mode
it saw the pending IPI bit and started processing it. Meanwhile control
CPU did not observe the ALIVE state within the timeout in
cpuhp_bp_sync_alive() and aborted the bringup. The secondary CPU
eventually entered Linux and got stuck spinning in cpuhp_ap_sync_alive().

A retry of the bringup fails due to both the current sync-state
switching in kernel/cpu.c and the incomplete SBI error handling in
arch/riscv/kernel/cpu_ops_sbi.c (see the patches for details). Both
behaviors are addressed by this patchset.

Artemii Patov (2):
  cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE
  riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start

 arch/riscv/include/asm/sbi.h    |  2 ++
 arch/riscv/kernel/cpu_ops_sbi.c |  7 ++++++-
 kernel/cpu.c                    | 17 +++++++++++++++--
 3 files changed, 23 insertions(+), 3 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE
  2026-09-27 16:44 [PATCH 0/2] *** cpu/hotplug, riscv: Fix CPU stuck offline after failed hotplug attempt *** artemiispatov
@ 2026-09-27 16:44 ` artemiispatov
  2026-09-27 16:44 ` [PATCH 2/2] riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start artemiispatov
  1 sibling, 0 replies; 3+ messages in thread
From: artemiispatov @ 2026-09-27 16:44 UTC (permalink / raw)
  To: linux-kernel, linux-riscv
  Cc: tglx, peterz, anup, atishp, palmer, pjw, aou, alex, conor,
	hui.wang, artemiispatov

From: Artemii Patov <artemiispatov@gmail.com>

There are scenarios where the AP gets stuck in cpuhp_ap_sync_alive()
during bringup. For example, when the AP starts late, the control CPU
may time out in cpuhp_wait_for_sync_state() waiting for the ALIVE state
and abort the bringup with an error.

The AP eventually reaches cpuhp_ap_sync_alive(), marks itself ALIVE and
spins, waiting for the control CPU to release it. The only way out of
this loop is the compare-exchange in cpuhp_wait_for_sync_state(), which
is invoked from cpuhp_bp_sync_alive() and transitions
SYNC_STATE_ALIVE to SYNC_STATE_SHOULD_ONLINE.

On a retry of the bringup, the ALIVE state was treated as a CPU which
did not come up: cpuhp_can_boot_ap() overwrote SYNC_STATE_ALIVE with
SYNC_STATE_KICKED and the kick stage tried to start the AP again. That
would require the AP to run through the boot stages once more until it
sets SYNC_STATE_ALIVE, which the kick mechanism cannot guarantee: on
RISC-V, sbi_hsm_hart_start() on an already running hart returns
SBI_ERR_ALREADY_AVAILABLE and does not reset it. So the AP remained
stuck and the retry failed.

Fix this by treating SYNC_STATE_ALIVE as "already alive, no kick
needed": cpuhp_can_boot_ap() leaves the state unchanged and
cpuhp_kick_ap_alive() skips the kick in that case. The retry then
proceeds to cpuhp_wait_for_sync_state(), which performs the
ALIVE -> SHOULD_ONLINE transition and releases the stuck CPU.

In cpuhp_ap_sync_alive() promote SYNC_STATE_KICKED to SYNC_STATE_ALIVE
only, so that SYNC_STATE_SHOULD_ONLINE set earlier by the control CPU
(e.g. when the AP re-enters the sync point on a retry after a failed
bringup attempt) is not overwritten.

Fixes: 6f0621238b7e ("cpu/hotplug: Add CPU state tracking and synchronization")
Signed-off-by: Artemii Patov <artemiispatov@gmail.com>
---
 kernel/cpu.c | 17 +++++++++++++++--
 1 file changed, 15 insertions(+), 2 deletions(-)

diff --git a/kernel/cpu.c b/kernel/cpu.c
index b3c8553d7bd6..eb9cb7c3e565 100644
--- a/kernel/cpu.c
+++ b/kernel/cpu.c
@@ -392,7 +392,11 @@ void cpuhp_ap_sync_alive(void)
 {
 	atomic_t *st = this_cpu_ptr(&cpuhp_state.ap_sync_state);
 
-	cpuhp_ap_update_sync_state(SYNC_STATE_ALIVE);
+	/*
+	 * Compare-exchange failure means that the state is either SYNC_STATE_ALIVE
+	 * or SYNC_STATE_SHOULD_ONLINE, both are acceptable.
+	 */
+	atomic_cmpxchg(st, SYNC_STATE_KICKED, SYNC_STATE_ALIVE);
 
 	/* Wait for the control CPU to release it. */
 	while (atomic_read(st) != SYNC_STATE_SHOULD_ONLINE)
@@ -414,7 +418,7 @@ static bool cpuhp_can_boot_ap(unsigned int cpu)
 		break;
 	case SYNC_STATE_ALIVE:
 		/* CPU is stuck cpuhp_ap_sync_alive(). */
-		break;
+		return true;
 	default:
 		/* CPU failed to report online or dead and is in limbo state. */
 		return false;
@@ -822,6 +826,15 @@ static int bringup_wait_for_ap_online(unsigned int cpu)
 #ifdef CONFIG_HOTPLUG_SPLIT_STARTUP
 static int cpuhp_kick_ap_alive(unsigned int cpu)
 {
+	struct cpuhp_cpu_state *st = per_cpu_ptr(&cpuhp_state, cpu);
+
+	/*
+	 * The AP is already alive and waiting in cpuhp_ap_sync_alive().
+	 * cpuhp_bp_sync_alive() will release it, so skip the kick.
+	 */
+	if (atomic_read(&st->ap_sync_state) == SYNC_STATE_ALIVE)
+		return 0;
+
 	if (!cpuhp_can_boot_ap(cpu))
 		return -EAGAIN;
 
-- 
2.43.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 2/2] riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start
  2026-09-27 16:44 [PATCH 0/2] *** cpu/hotplug, riscv: Fix CPU stuck offline after failed hotplug attempt *** artemiispatov
  2026-09-27 16:44 ` [PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE artemiispatov
@ 2026-09-27 16:44 ` artemiispatov
  1 sibling, 0 replies; 3+ messages in thread
From: artemiispatov @ 2026-09-27 16:44 UTC (permalink / raw)
  To: linux-kernel, linux-riscv
  Cc: tglx, peterz, anup, atishp, palmer, pjw, aou, alex, conor,
	hui.wang, artemiispatov

From: Artemii Patov <artemiispatov@gmail.com>

Commit 8ac00cddc1a6 ("cpu/hotplug: Skip kick stage if the AP is stuck in
SYNC_STATE_ALIVE") changes the sync-state switching in kernel/cpu.c so
that a second AP bringup is no longer aborted. For that to work on
RISC-V, sbi_hsm_hart_start() must treat SBI_ERR_ALREADY_AVAILABLE as
a successful start: HSM returns this error when the hart is already
running.

Add the SBI_ERR_ALREADY_AVAILABLE -> -EALREADY mapping to
sbi_err_map_linux_errno() and treat -EALREADY as success in
sbi_hsm_hart_start(). Otherwise the error propagates through
arch_cpuhp_kick_ap_alive() to cpuhp_kick_ap_alive(), which fails the
retry and aborts the bringup.

Signed-off-by: Artemii Patov <artemiispatov@gmail.com>
---
 arch/riscv/include/asm/sbi.h    | 2 ++
 arch/riscv/kernel/cpu_ops_sbi.c | 7 ++++++-
 2 files changed, 8 insertions(+), 1 deletion(-)

diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
index 5725e0ca4dda..25016f1a9ad0 100644
--- a/arch/riscv/include/asm/sbi.h
+++ b/arch/riscv/include/asm/sbi.h
@@ -664,6 +664,8 @@ static inline int sbi_err_map_linux_errno(int err)
 		return -EINVAL;
 	case SBI_ERR_BAD_RANGE:
 		return -ERANGE;
+	case SBI_ERR_ALREADY_AVAILABLE:
+		return -EALREADY;
 	case SBI_ERR_INVALID_ADDRESS:
 		return -EFAULT;
 	case SBI_ERR_NO_SHMEM:
diff --git a/arch/riscv/kernel/cpu_ops_sbi.c b/arch/riscv/kernel/cpu_ops_sbi.c
index ee6e4b5cc39e..5acca3f4f8c0 100644
--- a/arch/riscv/kernel/cpu_ops_sbi.c
+++ b/arch/riscv/kernel/cpu_ops_sbi.c
@@ -26,12 +26,17 @@ static struct sbi_hart_boot_data boot_data[NR_CPUS];
 static int sbi_hsm_hart_start(unsigned long hartid, unsigned long saddr,
 			      unsigned long priv)
 {
+	int err;
 	struct sbiret ret;
 
 	ret = sbi_ecall(SBI_EXT_HSM, SBI_EXT_HSM_HART_START,
 			hartid, saddr, priv, 0, 0, 0);
 
-	return sbi_err_map_linux_errno(ret.error);
+	err = sbi_err_map_linux_errno(ret.error);
+	if (err == -EALREADY)
+		return 0;
+
+	return err;
 }
 
 #ifdef CONFIG_HOTPLUG_CPU
-- 
2.43.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-27 16:44 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 16:44 [PATCH 0/2] *** cpu/hotplug, riscv: Fix CPU stuck offline after failed hotplug attempt *** artemiispatov
2026-09-27 16:44 ` [PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE artemiispatov
2026-09-27 16:44 ` [PATCH 2/2] riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start artemiispatov

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®