[PATCH 1/2] cpu/hotplug: Skip kick stage if the AP is stuck in SYNC_STATE_ALIVE

artemiispatov at gmail.com artemiispatov at gmail.com
Sun Sep 27 09:44:17 PDT 2026


From: Artemii Patov <artemiispatov at gmail.com>

There are scenarios where the AP gets stuck in cpuhp_ap_sync_alive()
during bringup. For example, when the AP starts late, the control CPU
may time out in cpuhp_wait_for_sync_state() waiting for the ALIVE state
and abort the bringup with an error.

The AP eventually reaches cpuhp_ap_sync_alive(), marks itself ALIVE and
spins, waiting for the control CPU to release it. The only way out of
this loop is the compare-exchange in cpuhp_wait_for_sync_state(), which
is invoked from cpuhp_bp_sync_alive() and transitions
SYNC_STATE_ALIVE to SYNC_STATE_SHOULD_ONLINE.

On a retry of the bringup, the ALIVE state was treated as a CPU which
did not come up: cpuhp_can_boot_ap() overwrote SYNC_STATE_ALIVE with
SYNC_STATE_KICKED and the kick stage tried to start the AP again. That
would require the AP to run through the boot stages once more until it
sets SYNC_STATE_ALIVE, which the kick mechanism cannot guarantee: on
RISC-V, sbi_hsm_hart_start() on an already running hart returns
SBI_ERR_ALREADY_AVAILABLE and does not reset it. So the AP remained
stuck and the retry failed.

Fix this by treating SYNC_STATE_ALIVE as "already alive, no kick
needed": cpuhp_can_boot_ap() leaves the state unchanged and
cpuhp_kick_ap_alive() skips the kick in that case. The retry then
proceeds to cpuhp_wait_for_sync_state(), which performs the
ALIVE -> SHOULD_ONLINE transition and releases the stuck CPU.

In cpuhp_ap_sync_alive() promote SYNC_STATE_KICKED to SYNC_STATE_ALIVE
only, so that SYNC_STATE_SHOULD_ONLINE set earlier by the control CPU
(e.g. when the AP re-enters the sync point on a retry after a failed
bringup attempt) is not overwritten.

Fixes: 6f0621238b7e ("cpu/hotplug: Add CPU state tracking and synchronization")
Signed-off-by: Artemii Patov <artemiispatov at gmail.com>
---
 kernel/cpu.c | 17 +++++++++++++++--
 1 file changed, 15 insertions(+), 2 deletions(-)

diff --git a/kernel/cpu.c b/kernel/cpu.c
index b3c8553d7bd6..eb9cb7c3e565 100644
--- a/kernel/cpu.c
+++ b/kernel/cpu.c
@@ -392,7 +392,11 @@ void cpuhp_ap_sync_alive(void)
 {
 	atomic_t *st = this_cpu_ptr(&cpuhp_state.ap_sync_state);
 
-	cpuhp_ap_update_sync_state(SYNC_STATE_ALIVE);
+	/*
+	 * Compare-exchange failure means that the state is either SYNC_STATE_ALIVE
+	 * or SYNC_STATE_SHOULD_ONLINE, both are acceptable.
+	 */
+	atomic_cmpxchg(st, SYNC_STATE_KICKED, SYNC_STATE_ALIVE);
 
 	/* Wait for the control CPU to release it. */
 	while (atomic_read(st) != SYNC_STATE_SHOULD_ONLINE)
@@ -414,7 +418,7 @@ static bool cpuhp_can_boot_ap(unsigned int cpu)
 		break;
 	case SYNC_STATE_ALIVE:
 		/* CPU is stuck cpuhp_ap_sync_alive(). */
-		break;
+		return true;
 	default:
 		/* CPU failed to report online or dead and is in limbo state. */
 		return false;
@@ -822,6 +826,15 @@ static int bringup_wait_for_ap_online(unsigned int cpu)
 #ifdef CONFIG_HOTPLUG_SPLIT_STARTUP
 static int cpuhp_kick_ap_alive(unsigned int cpu)
 {
+	struct cpuhp_cpu_state *st = per_cpu_ptr(&cpuhp_state, cpu);
+
+	/*
+	 * The AP is already alive and waiting in cpuhp_ap_sync_alive().
+	 * cpuhp_bp_sync_alive() will release it, so skip the kick.
+	 */
+	if (atomic_read(&st->ap_sync_state) == SYNC_STATE_ALIVE)
+		return 0;
+
 	if (!cpuhp_can_boot_ap(cpu))
 		return -EAGAIN;
 
-- 
2.43.0




More information about the linux-riscv mailing list