From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0500E3FE354 for ; Thu, 4 Jun 2026 09:29:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780565381; cv=none; b=IVR4dPWrjFcM4ZUvWq8RnUtCGPUdfGBT5GXo2hPal6d6oDAbp4ac9YTkCDbrO0ZAdTKOjhgsqizyRicpR0p/ZZEoSlMhD/+fnJFRB2lcp1TDghyA7wJhMwxJp7bALnrPevLjk/eVwHwqM1EEoEkCQxBJ8kEhKxZtLxBLIuaS6gM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780565381; c=relaxed/simple; bh=LSYx6S/5jyu9KfVp/5iYPsZKG9eGFj04YvZKjTai2Y8=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=Lws1jC+BT50V0CytKBHGGkHtdQqEH1SVY7TwGPMeerwaMW/pKPHg88H3PlVQMYAQMwOpKc4xLeEGLEa4QWefVgWhzh2VQ8nTCOWuDNgyimgzQBQlq25HnDDYyGZHz9CDA7DyScAbHp3sWJc3w14LlRoYsD0wp3xXxZi7qEs/+Ms= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=GjiK6lm4; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=MrGwsVwB; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="GjiK6lm4"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="MrGwsVwB" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1780565376; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=+huBJ1nPsqC9QTJtR1rRZxAcoh6LRj4ywJbAKenhqm4=; b=GjiK6lm4DrOSLaCf6X3TvftnDIX9OO4qHHs4dMtnirzAfRlO5MP5DTAVE7Am3s1/voFyK4 IFjBCW1bQ3WCif817SybjnA3EhufLobVOGrJYVtjk8oDMXHPnMJ749Xu+SV2foxj9w95sJ AkxfpS+3OOwTaC5P1hGHJDuCSqyqisM= Received: from mail-wr1-f71.google.com (mail-wr1-f71.google.com [209.85.221.71]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-630-zJd69HULP3e0Y2M90MXGhw-1; Thu, 04 Jun 2026 05:29:35 -0400 X-MC-Unique: zJd69HULP3e0Y2M90MXGhw-1 X-Mimecast-MFC-AGG-ID: zJd69HULP3e0Y2M90MXGhw_1780565374 Received: by mail-wr1-f71.google.com with SMTP id ffacd0b85a97d-45ef6b407b4so221800f8f.1 for ; Thu, 04 Jun 2026 02:29:35 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1780565374; x=1781170174; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:content-language:from :references:cc:to:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=+huBJ1nPsqC9QTJtR1rRZxAcoh6LRj4ywJbAKenhqm4=; b=MrGwsVwBVmKBHV3LNcnxKwkeYZBZz4/9S/ZG93qcrPWEBjxvCQ/Mvet4v7imvKTi+P c3JPtQuD/YYxlw8rOLdGIDqsG36yrPhLxHAlenO4jOTeoYCJt7GDlujr+Q7EbmlBT9xF MBcEhrpr27Ao7SnmG/yRQ3Io39ZpGwaOx0He1itSGPDxPlKDZB9R5JFhXgghsW8R7imW ceopw6yzmM8ieQ9rqLdjOWZx5hXWTg4UU7ViEsLahOgKfLytUoKu360dJEwqO/+ALC/r QTzl7/RR61WMnhRPaRdMTCJCYNjtVc1dGAfSp7tU9nqJIfPGpLLo4Zqk74eH+ibMx2XI uAuQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1780565374; x=1781170174; h=content-transfer-encoding:in-reply-to:content-language:from :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=+huBJ1nPsqC9QTJtR1rRZxAcoh6LRj4ywJbAKenhqm4=; b=ChNO/BoTYFVNOu4vZjUhfUsAsNWPEYGJ9tbU1OHToUdABgQwcjmcf19kM+eerLjye5 2ux0Jlg0JbkU8ApoVLzAtdGiJwI/6q7MijbVgKGHaRdvlqGmeJC4rCAIqEry8DqfWeLV MaPt8x3SP5N5wm/HcDNYQ/emTuh3qc3g5iLfOIx6UU1AK5yN4+CKiH3a+Q1/tEA0SoQt ou5dh7Y7Z/6+TH34gUNbZRrEiVfAZel9+DbCP0YIMCcbpByH79jdLLE2ZIXiFiexr8Eu X6t1f9kN7+XGVwGbH38ff5mxPxVHY4sgethAPeEM2n5JOCXmhYLDwWiMJJ27JpiiZsFH KWmQ== X-Forwarded-Encrypted: i=1; AFNElJ/coZprxmk0NkgZrozwoVxmFb2L4MD9ZxLG+EMF/FXCOYoW5xnk/lMzvQWEI4nmOdX7R0I0hATs1eIIwzQ=@vger.kernel.org X-Gm-Message-State: AOJu0YzTOhyZgwy+Gk9ch0mI5fHheWy2beEC5gqUEpt/KlzBUcSY6iJh aGKLYPt6BnMlYeGkaYTuhXeqmCm5R8c5VULYiHI0GEJpykG8Re2FhalHsu8tg6jqsanmXPGCbRZ oRt79lxaaeHfZN6HP+SOxsP1s9XrTOTJRG9uzb29TuqZ+7r4bUZk83tBaHuSNcZ4vpQ== X-Gm-Gg: Acq92OE9RIIl+jASqk4SFPSG0mHRklwXU5WX/4W8ML7TXm4JYzMqVqUkCAujpfM5M9C NqkkBPx2+WpB8frsmgd7IQdwXztpYXYSim9VEw0ms3C69cgE53k2+JF+XYCJVp12jD0WnKvze1k rSEarq75hzxOy7pS/ELtiNqMJV3xFe6qS1PR+j4GnxrSkWz+V0sqMTBNw5pB/nkZjcV/GRao85y hSaG6BWAHvaWo/Jxs7qfkYOnXCZ5whNdyrjEAF/phcEVmVmJDLOf2dfR78i7BnsOgxehCdG/ai5 5Mb2ssDQMEo9kvDABYSy+RXjZUK2rg2n2M2ExXwjzMpTe0OSGZ720Kao4rm4p56E2O1Hr4so2+5 ilpEWgKShiFfpJOb3xebW0KwHMva5k5Hz/0HGYfDqkCSxvn725MGphN9t3U7h6oIkisA= X-Received: by 2002:a05:600c:4fc6:b0:490:4973:91a0 with SMTP id 5b1f17b1804b1-490b5e950admr112769645e9.10.1780565374438; Thu, 04 Jun 2026 02:29:34 -0700 (PDT) X-Received: by 2002:a05:600c:4fc6:b0:490:4973:91a0 with SMTP id 5b1f17b1804b1-490b5e950admr112769175e9.10.1780565373963; Thu, 04 Jun 2026 02:29:33 -0700 (PDT) Received: from [192.168.88.32] ([212.105.155.59]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-490bd670c2esm45301645e9.0.2026.06.04.02.29.32 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Thu, 04 Jun 2026 02:29:33 -0700 (PDT) Message-ID: Date: Thu, 4 Jun 2026 11:29:31 +0200 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH] net: wwan: t7xx: fix race between TX thread and system PM suspend To: Tim JH Chen , netdev@vger.kernel.org Cc: haijun.liu@mediatek.com, chandrashekar.devegowda@intel.com, ricardo.martinez@linux.intel.com, loic.poulain@oss.qualcomm.com, ryazanov.s.a@gmail.com, johannes@sipsolutions.net, andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com, kuba@kernel.org, linux-kernel@vger.kernel.org, tim.jh.chen@wnc.com.tw, Chih.Hung.Huang@wnc.com.tw References: <2f9c5f6b-1d8d-4c8b-815d-77a40aa76e23@redhat.com> <20260601015231.3211764-1-tim.jh.chen@wnc.com.tw> From: Paolo Abeni Content-Language: en-US In-Reply-To: <20260601015231.3211764-1-tim.jh.chen@wnc.com.tw> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit On 6/1/26 3:52 AM, Tim JH Chen wrote: > When system suspend is triggered while the DPMAIF TX kthread > (t7xx_dpmaif_tx_hw_push_thread) is running, a deadlock can occur > leading to a CPU soft lockup. > > The root cause is two-fold: > > 1. t7xx_dpmaif_suspend() calls t7xx_dpmaif_tx_stop() which only stops > the TX work-queue items (by clearing txq->que_started and waiting on > txq->tx_processing). It does NOT signal the kthread and does NOT > update dpmaif_ctrl->state, which stays DPMAIF_STATE_PWRON. > > 2. The kthread's state guard is only checked at the top of each loop > iteration. If the thread already passed this guard, it proceeds > unconditionally to call pm_runtime_resume_and_get() — which tries to > acquire dev->power.lock also contended by the system PM suspend path. > > The result is a spinlock deadlock observed as: > > watchdog: BUG: soft lockup - CPU#N stuck for 26s! [dpmaif_tx_hw_pu] > RIP: _raw_spin_unlock_irqrestore > Call Trace: > __pm_runtime_resume+0x5b/0x80 > t7xx_dpmaif_tx_hw_push_thread+0xc4 [mtk_t7xx] > > The condition requires ASPM L1 enabled on the endpoint (which extends > the time pm_runtime_resume_and_get() holds dev->power.lock during L1.2 > link retraining) and hundreds of repeated suspend/resume cycles to > trigger reliably. > > Fix by introducing tx_pm_lock (struct mutex) and several coordinated > changes: > > t7xx_dpmaif_suspend(): > After t7xx_dpmaif_tx_stop(), acquire tx_pm_lock. Under the lock, > snapshot dpmaif_ctrl->state into pre_suspend_state (capturing the > modem state atomically with respect to the kthread's PM section), > then set DPMAIF_STATE_PWROFF via WRITE_ONCE(). Release the lock > and call wake_up() so any sleeping kthread re-evaluates the > wait_event condition and exits. > > t7xx_dpmaif_suspend() acquires tx_pm_lock without holding any PM > lock. While it waits, the kthread may call pm_runtime_resume_and_get() > which briefly takes and releases dev->power.lock independently. > Because the suspend callback does not compete for dev->power.lock at > this point, the original spinlock deadlock cannot occur. Suspend > latency increases by at most one TX burst drain time, which is > bounded by the DRB ring depth. > > t7xx_dpmaif_resume(): > When pre_suspend_state is DPMAIF_STATE_PWRON, re-arm the HW fully > (start_txrx_qs, enable_irq, unmask_dlq_intr, start_hw) before > publishing the new state. This ensures the kthread cannot issue > ul_update_hw_drb_cnt() MMIO writes before UL_ALL_Q_EN is set by > t7xx_dpmaif_start_hw(). Publish the restored state under tx_pm_lock > to serialise with the kthread's under-lock state check. Wake up the > kthread only after HW and state are both consistent. > > When pre_suspend_state is DPMAIF_STATE_PWROFF (modem was already > stopped or in exception before suspend), skip HW re-arming entirely > to avoid leaving DMA engines running while the MD state machine > considers the modem inactive. > > t7xx_dpmaif_tx_hw_push_thread(): > Hold tx_pm_lock across the [state check -> pm_runtime_resume_and_get > -> pm_runtime_put_autosuspend] sequence. A second READ_ONCE() state > check under the lock closes the TOCTOU window between the wait_event > guard at the loop top and the pm_runtime call. READ_ONCE() is used > in all unguarded state reads in this function. > > t7xx_dpmaif_start() / t7xx_dpmaif_stop(): > Use WRITE_ONCE() for state writes to match the READ_ONCE() reads > used throughout the driver and prevent compiler optimisations from > obscuring concurrent access. > > t7xx_do_tx_hw_push(): > Use READ_ONCE() in the do/while termination condition to match the > WRITE_ONCE() annotations on the write side. > > t7xx_dpmaif_tx_thread_init(): > Initialise tx_pm_lock with mutex_init(). > > Note: t7xx_dpmaif_start() and t7xx_dpmaif_stop() (called from the > MD-FSM kthread via t7xx_dpmaif_md_state_callback()) do not hold > tx_pm_lock. A race where the FSM transitions the modem to > DPMAIF_STATE_PWROFF concurrently with the TX kthread's last burst is > pre-existing and not introduced by this patch; the do/while condition > in t7xx_do_tx_hw_push() now re-checks state with READ_ONCE() at each > iteration boundary, limiting exposure to at most one burst. > > Tested: no soft lockup observed over 500+ suspend/resume cycles with > SIM registered and ASPM L1 enabled (previously triggered in < 300). > > Fixes: 46e8f49ed7b3 ("net: wwan: t7xx: Introduce power management") > Signed-off-by: Tim JH Chen The above is way too verbose and hints that this patch should likely be split in a series. Note that process wise there are still several problems: - missing revision number in the suby prefix - mismatch between from email message and SoB - new revision MUST NOT be in reply-to of older ones. Please try to be accurate with your next resubmission, or we will have to delay processing this patch for an additional while. Sashiko has still quite a bit of concerns: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260601015231.3211764-1-tim.jh.chen%40wnc.com.tw /P