* [PATCH 0/2] *** cpu/hotplug, riscv: Fix CPU stuck offline after failed hotplug attempt ***
@ 2026-09-27 16:44 artemiispatov
2026-09-27 16:44 ` [PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE artemiispatov
2026-09-27 16:44 ` [PATCH 2/2] riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start artemiispatov
0 siblings, 2 replies; 3+ messages in thread
From: artemiispatov @ 2026-09-27 16:44 UTC (permalink / raw)
To: linux-kernel, linux-riscv
Cc: tglx, peterz, anup, atishp, palmer, pjw, aou, alex, conor,
hui.wang, artemiispatov
From: Artemii Patov <artemiispatov@gmail.com>
These two patches enable RISC-V to retry bringing an AP
online when the control CPU failed to complete the
SYNC_STATE_ALIVE -> SYNC_STATE_SHOULD_ONLINE transition and
the secondary CPU is stuck spinning in cpuhp_ap_sync_alive().
The need for this arose during debugging RISC-V SoC. Control CPU tried
to wake up a halted CPU and sent an IPI via OpenSBI. The IPI was masked
while the secondary CPU was in Debug Mode; when the CPU left Debug Mode
it saw the pending IPI bit and started processing it. Meanwhile control
CPU did not observe the ALIVE state within the timeout in
cpuhp_bp_sync_alive() and aborted the bringup. The secondary CPU
eventually entered Linux and got stuck spinning in cpuhp_ap_sync_alive().
A retry of the bringup fails due to both the current sync-state
switching in kernel/cpu.c and the incomplete SBI error handling in
arch/riscv/kernel/cpu_ops_sbi.c (see the patches for details). Both
behaviors are addressed by this patchset.
Artemii Patov (2):
cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE
riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start
arch/riscv/include/asm/sbi.h | 2 ++
arch/riscv/kernel/cpu_ops_sbi.c | 7 ++++++-
kernel/cpu.c | 17 +++++++++++++++--
3 files changed, 23 insertions(+), 3 deletions(-)
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE
2026-09-27 16:44 [PATCH 0/2] *** cpu/hotplug, riscv: Fix CPU stuck offline after failed hotplug attempt *** artemiispatov
@ 2026-09-27 16:44 ` artemiispatov
2026-09-27 16:44 ` [PATCH 2/2] riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start artemiispatov
1 sibling, 0 replies; 3+ messages in thread
From: artemiispatov @ 2026-09-27 16:44 UTC (permalink / raw)
To: linux-kernel, linux-riscv
Cc: tglx, peterz, anup, atishp, palmer, pjw, aou, alex, conor,
hui.wang, artemiispatov
From: Artemii Patov <artemiispatov@gmail.com>
There are scenarios where the AP gets stuck in cpuhp_ap_sync_alive()
during bringup. For example, when the AP starts late, the control CPU
may time out in cpuhp_wait_for_sync_state() waiting for the ALIVE state
and abort the bringup with an error.
The AP eventually reaches cpuhp_ap_sync_alive(), marks itself ALIVE and
spins, waiting for the control CPU to release it. The only way out of
this loop is the compare-exchange in cpuhp_wait_for_sync_state(), which
is invoked from cpuhp_bp_sync_alive() and transitions
SYNC_STATE_ALIVE to SYNC_STATE_SHOULD_ONLINE.
On a retry of the bringup, the ALIVE state was treated as a CPU which
did not come up: cpuhp_can_boot_ap() overwrote SYNC_STATE_ALIVE with
SYNC_STATE_KICKED and the kick stage tried to start the AP again. That
would require the AP to run through the boot stages once more until it
sets SYNC_STATE_ALIVE, which the kick mechanism cannot guarantee: on
RISC-V, sbi_hsm_hart_start() on an already running hart returns
SBI_ERR_ALREADY_AVAILABLE and does not reset it. So the AP remained
stuck and the retry failed.
Fix this by treating SYNC_STATE_ALIVE as "already alive, no kick
needed": cpuhp_can_boot_ap() leaves the state unchanged and
cpuhp_kick_ap_alive() skips the kick in that case. The retry then
proceeds to cpuhp_wait_for_sync_state(), which performs the
ALIVE -> SHOULD_ONLINE transition and releases the stuck CPU.
In cpuhp_ap_sync_alive() promote SYNC_STATE_KICKED to SYNC_STATE_ALIVE
only, so that SYNC_STATE_SHOULD_ONLINE set earlier by the control CPU
(e.g. when the AP re-enters the sync point on a retry after a failed
bringup attempt) is not overwritten.
Fixes: 6f0621238b7e ("cpu/hotplug: Add CPU state tracking and synchronization")
Signed-off-by: Artemii Patov <artemiispatov@gmail.com>
---
kernel/cpu.c | 17 +++++++++++++++--
1 file changed, 15 insertions(+), 2 deletions(-)
diff --git a/kernel/cpu.c b/kernel/cpu.c
index b3c8553d7bd6..eb9cb7c3e565 100644
--- a/kernel/cpu.c
+++ b/kernel/cpu.c
@@ -392,7 +392,11 @@ void cpuhp_ap_sync_alive(void)
{
atomic_t *st = this_cpu_ptr(&cpuhp_state.ap_sync_state);
- cpuhp_ap_update_sync_state(SYNC_STATE_ALIVE);
+ /*
+ * Compare-exchange failure means that the state is either SYNC_STATE_ALIVE
+ * or SYNC_STATE_SHOULD_ONLINE, both are acceptable.
+ */
+ atomic_cmpxchg(st, SYNC_STATE_KICKED, SYNC_STATE_ALIVE);
/* Wait for the control CPU to release it. */
while (atomic_read(st) != SYNC_STATE_SHOULD_ONLINE)
@@ -414,7 +418,7 @@ static bool cpuhp_can_boot_ap(unsigned int cpu)
break;
case SYNC_STATE_ALIVE:
/* CPU is stuck cpuhp_ap_sync_alive(). */
- break;
+ return true;
default:
/* CPU failed to report online or dead and is in limbo state. */
return false;
@@ -822,6 +826,15 @@ static int bringup_wait_for_ap_online(unsigned int cpu)
#ifdef CONFIG_HOTPLUG_SPLIT_STARTUP
static int cpuhp_kick_ap_alive(unsigned int cpu)
{
+ struct cpuhp_cpu_state *st = per_cpu_ptr(&cpuhp_state, cpu);
+
+ /*
+ * The AP is already alive and waiting in cpuhp_ap_sync_alive().
+ * cpuhp_bp_sync_alive() will release it, so skip the kick.
+ */
+ if (atomic_read(&st->ap_sync_state) == SYNC_STATE_ALIVE)
+ return 0;
+
if (!cpuhp_can_boot_ap(cpu))
return -EAGAIN;
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH 2/2] riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start
2026-09-27 16:44 [PATCH 0/2] *** cpu/hotplug, riscv: Fix CPU stuck offline after failed hotplug attempt *** artemiispatov
2026-09-27 16:44 ` [PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE artemiispatov
@ 2026-09-27 16:44 ` artemiispatov
1 sibling, 0 replies; 3+ messages in thread
From: artemiispatov @ 2026-09-27 16:44 UTC (permalink / raw)
To: linux-kernel, linux-riscv
Cc: tglx, peterz, anup, atishp, palmer, pjw, aou, alex, conor,
hui.wang, artemiispatov
From: Artemii Patov <artemiispatov@gmail.com>
Commit 8ac00cddc1a6 ("cpu/hotplug: Skip kick stage if the AP is stuck in
SYNC_STATE_ALIVE") changes the sync-state switching in kernel/cpu.c so
that a second AP bringup is no longer aborted. For that to work on
RISC-V, sbi_hsm_hart_start() must treat SBI_ERR_ALREADY_AVAILABLE as
a successful start: HSM returns this error when the hart is already
running.
Add the SBI_ERR_ALREADY_AVAILABLE -> -EALREADY mapping to
sbi_err_map_linux_errno() and treat -EALREADY as success in
sbi_hsm_hart_start(). Otherwise the error propagates through
arch_cpuhp_kick_ap_alive() to cpuhp_kick_ap_alive(), which fails the
retry and aborts the bringup.
Signed-off-by: Artemii Patov <artemiispatov@gmail.com>
---
arch/riscv/include/asm/sbi.h | 2 ++
arch/riscv/kernel/cpu_ops_sbi.c | 7 ++++++-
2 files changed, 8 insertions(+), 1 deletion(-)
diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
index 5725e0ca4dda..25016f1a9ad0 100644
--- a/arch/riscv/include/asm/sbi.h
+++ b/arch/riscv/include/asm/sbi.h
@@ -664,6 +664,8 @@ static inline int sbi_err_map_linux_errno(int err)
return -EINVAL;
case SBI_ERR_BAD_RANGE:
return -ERANGE;
+ case SBI_ERR_ALREADY_AVAILABLE:
+ return -EALREADY;
case SBI_ERR_INVALID_ADDRESS:
return -EFAULT;
case SBI_ERR_NO_SHMEM:
diff --git a/arch/riscv/kernel/cpu_ops_sbi.c b/arch/riscv/kernel/cpu_ops_sbi.c
index ee6e4b5cc39e..5acca3f4f8c0 100644
--- a/arch/riscv/kernel/cpu_ops_sbi.c
+++ b/arch/riscv/kernel/cpu_ops_sbi.c
@@ -26,12 +26,17 @@ static struct sbi_hart_boot_data boot_data[NR_CPUS];
static int sbi_hsm_hart_start(unsigned long hartid, unsigned long saddr,
unsigned long priv)
{
+ int err;
struct sbiret ret;
ret = sbi_ecall(SBI_EXT_HSM, SBI_EXT_HSM_HART_START,
hartid, saddr, priv, 0, 0, 0);
- return sbi_err_map_linux_errno(ret.error);
+ err = sbi_err_map_linux_errno(ret.error);
+ if (err == -EALREADY)
+ return 0;
+
+ return err;
}
#ifdef CONFIG_HOTPLUG_CPU
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-27 16:44 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 16:44 [PATCH 0/2] *** cpu/hotplug, riscv: Fix CPU stuck offline after failed hotplug attempt *** artemiispatov
2026-09-27 16:44 ` [PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE artemiispatov
2026-09-27 16:44 ` [PATCH 2/2] riscv: sbi: Accept SBI_ERR_ALREADY_AVAILABLE in sbi_hsm_hart_start artemiispatov
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®