From: Paolo Abeni <pabeni@redhat.com>
To: Tim JH Chen <tim770802@gmail.com>, netdev@vger.kernel.org
Cc: Tim JH Chen <tim.jh.chen@wnc.com.tw>,
Chandrashekar Devegowda <chandrashekar.devegowda@intel.com>,
Liu Haijun <haijun.liu@mediatek.com>,
Ricardo Martinez <ricardo.martinez@linux.intel.com>,
Loic Poulain <loic.poulain@oss.qualcomm.com>,
Sergey Ryazanov <ryazanov.s.a@gmail.com>,
Johannes Berg <johannes@sipsolutions.net>,
Andrew Lunn <andrew+netdev@lunn.ch>,
"David S. Miller" <davem@davemloft.net>,
Eric Dumazet <edumazet@google.com>,
Jakub Kicinski <kuba@kernel.org>,
open list <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] net: wwan: t7xx: fix race between TX thread and system PM suspend
Date: Thu, 21 May 2026 12:33:51 +0200 [thread overview]
Message-ID: <5315a3b0-4fc8-4d74-8e27-d7b84bd794c9@redhat.com> (raw)
In-Reply-To: <20260518075033.58996-1-tim.jh.chen@wnc.com.tw>
On 5/18/26 9:50 AM, Tim JH Chen wrote:
> When system suspend is triggered while the DPMAIF TX kthread
> (t7xx_dpmaif_tx_hw_push_thread) is running, a deadlock can occur
> leading to a CPU soft lockup.
>
> The root cause is two-fold:
>
> 1. t7xx_dpmaif_suspend() calls t7xx_dpmaif_tx_stop() which only stops
> the TX work-queue items (by clearing txq->que_started and waiting on
> txq->tx_processing). It does NOT signal the kthread and does NOT
> update dpmaif_ctrl->state, which stays DPMAIF_STATE_PWRON.
>
> 2. The kthread's state guard (line: "if ... state != DPMAIF_STATE_PWRON")
> is only checked at the top of each loop iteration. If the thread
> already passed this guard, it proceeds unconditionally to call
> pm_runtime_resume_and_get() — which tries to acquire the PM spinlock
> also held (or contended) by the system PM suspend path.
>
> The result is a spinlock deadlock observed as:
>
> watchdog: BUG: soft lockup - CPU#N stuck for 26s! [dpmaif_tx_hw_pu]
> RIP: _raw_spin_unlock_irqrestore
> Call Trace:
> __pm_runtime_resume+0x5b/0x80
> t7xx_dpmaif_tx_hw_push_thread+0xc4 [mtk_t7xx]
>
> The condition requires ASPM L1 enabled on the endpoint (which extends
> the time pm_runtime_resume_and_get() holds the PM lock during L1.2
> link retraining) and hundreds of repeated suspend/resume cycles to
> trigger reliably.
>
> Fix by three coordinated changes:
>
> - In t7xx_dpmaif_suspend(): immediately set state to DPMAIF_STATE_PWROFF
> after stopping the TX queue, then call wake_up() so any sleeping thread
> re-evaluates the wait_event condition and stops.
>
> - In t7xx_dpmaif_resume(): restore state to DPMAIF_STATE_PWRON before
> re-enabling the TX queues, symmetric with the suspend change.
> Without this the kthread would never wake up after resume.
>
> - In t7xx_dpmaif_tx_hw_push_thread(): add a second state check
> immediately before pm_runtime_resume_and_get() to close the TOCTOU
> window between the wait_event guard and the pm call.
>
> Tested: no soft lockup observed over 500+ suspend/resume cycles with
> SIM registered and ASPM L1 enabled (previously triggered in < 300).
>
> Signed-off-by: Tim JH Chen <tim.jh.chen@wnc.com.tw>
This is a fix, it should target the 'net' tree including such tag into
the subj prefix and should carry a 'Fixes:' tag.
Also this is v2 of:
https://lore.kernel.org/netdev/TYZPR02MB5232A8C6A2BA56226D97CF4A90062@TYZPR02MB5232.apcprd02.prod.outlook.com/
the subj prefix should have included the relevant revision number and
you should have described what changed in the commit message after a
'---' separator.
Please have a deep read at the process documentation and specifically at:
Documentation/process/maintainer-netdev.rst
before posting the next revision.
/P
next prev parent reply other threads:[~2026-05-21 10:33 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-05-18 7:50 Tim JH Chen
2026-05-21 10:29 ` Paolo Abeni
2026-05-21 10:33 ` Paolo Abeni [this message]
2026-05-25 3:13 ` Tim JH Chen
2026-05-28 9:21 ` Paolo Abeni
2026-06-01 1:52 ` Tim JH Chen
2026-06-04 9:29 ` Paolo Abeni
-- strict thread matches above, loose matches on Subject: below --
2026-05-13 8:37 Tim JH Chen(陳仁鴻)
2026-05-15 0:19 ` Jakub Kicinski
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=5315a3b0-4fc8-4d74-8e27-d7b84bd794c9@redhat.com \
--to=pabeni@redhat.com \
--cc=andrew+netdev@lunn.ch \
--cc=chandrashekar.devegowda@intel.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=haijun.liu@mediatek.com \
--cc=johannes@sipsolutions.net \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=loic.poulain@oss.qualcomm.com \
--cc=netdev@vger.kernel.org \
--cc=ricardo.martinez@linux.intel.com \
--cc=ryazanov.s.a@gmail.com \
--cc=tim.jh.chen@wnc.com.tw \
--cc=tim770802@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®