mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Paolo Abeni <pabeni@redhat.com>
To: Tim JH Chen <tim770802@gmail.com>, netdev@vger.kernel.org
Cc: Tim JH Chen <tim.jh.chen@wnc.com.tw>,
	Chandrashekar Devegowda <chandrashekar.devegowda@intel.com>,
	Liu Haijun <haijun.liu@mediatek.com>,
	Ricardo Martinez <ricardo.martinez@linux.intel.com>,
	Loic Poulain <loic.poulain@oss.qualcomm.com>,
	Sergey Ryazanov <ryazanov.s.a@gmail.com>,
	Johannes Berg <johannes@sipsolutions.net>,
	Andrew Lunn <andrew+netdev@lunn.ch>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>,
	open list <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] net: wwan: t7xx: fix race between TX thread and system PM suspend
Date: Thu, 21 May 2026 12:33:51 +0200	[thread overview]
Message-ID: <5315a3b0-4fc8-4d74-8e27-d7b84bd794c9@redhat.com> (raw)
In-Reply-To: <20260518075033.58996-1-tim.jh.chen@wnc.com.tw>

On 5/18/26 9:50 AM, Tim JH Chen wrote:
> When system suspend is triggered while the DPMAIF TX kthread
> (t7xx_dpmaif_tx_hw_push_thread) is running, a deadlock can occur
> leading to a CPU soft lockup.
> 
> The root cause is two-fold:
> 
> 1. t7xx_dpmaif_suspend() calls t7xx_dpmaif_tx_stop() which only stops
>    the TX work-queue items (by clearing txq->que_started and waiting on
>    txq->tx_processing). It does NOT signal the kthread and does NOT
>    update dpmaif_ctrl->state, which stays DPMAIF_STATE_PWRON.
> 
> 2. The kthread's state guard (line: "if ... state != DPMAIF_STATE_PWRON")
>    is only checked at the top of each loop iteration. If the thread
>    already passed this guard, it proceeds unconditionally to call
>    pm_runtime_resume_and_get() — which tries to acquire the PM spinlock
>    also held (or contended) by the system PM suspend path.
> 
> The result is a spinlock deadlock observed as:
> 
>   watchdog: BUG: soft lockup - CPU#N stuck for 26s! [dpmaif_tx_hw_pu]
>   RIP: _raw_spin_unlock_irqrestore
>   Call Trace:
>     __pm_runtime_resume+0x5b/0x80
>     t7xx_dpmaif_tx_hw_push_thread+0xc4 [mtk_t7xx]
> 
> The condition requires ASPM L1 enabled on the endpoint (which extends
> the time pm_runtime_resume_and_get() holds the PM lock during L1.2
> link retraining) and hundreds of repeated suspend/resume cycles to
> trigger reliably.
> 
> Fix by three coordinated changes:
> 
> - In t7xx_dpmaif_suspend(): immediately set state to DPMAIF_STATE_PWROFF
>   after stopping the TX queue, then call wake_up() so any sleeping thread
>   re-evaluates the wait_event condition and stops.
> 
> - In t7xx_dpmaif_resume(): restore state to DPMAIF_STATE_PWRON before
>   re-enabling the TX queues, symmetric with the suspend change.
>   Without this the kthread would never wake up after resume.
> 
> - In t7xx_dpmaif_tx_hw_push_thread(): add a second state check
>   immediately before pm_runtime_resume_and_get() to close the TOCTOU
>   window between the wait_event guard and the pm call.
> 
> Tested: no soft lockup observed over 500+ suspend/resume cycles with
> SIM registered and ASPM L1 enabled (previously triggered in < 300).
> 
> Signed-off-by: Tim JH Chen <tim.jh.chen@wnc.com.tw>

This is a fix, it should target the 'net' tree including such tag into
the subj prefix and should carry a 'Fixes:' tag.

Also this is v2 of:

https://lore.kernel.org/netdev/TYZPR02MB5232A8C6A2BA56226D97CF4A90062@TYZPR02MB5232.apcprd02.prod.outlook.com/

the subj prefix should have included the relevant revision number and
you should have described what changed in the commit message after a
'---' separator.

Please have a deep read at the process documentation and specifically at:

Documentation/process/maintainer-netdev.rst

before posting the next revision.

/P


  parent reply	other threads:[~2026-05-21 10:33 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-05-18  7:50 Tim JH Chen
2026-05-21 10:29 ` Paolo Abeni
2026-05-21 10:33 ` Paolo Abeni [this message]
2026-05-25  3:13   ` Tim JH Chen
2026-05-28  9:21     ` Paolo Abeni
2026-06-01  1:52       ` Tim JH Chen
2026-06-04  9:29         ` Paolo Abeni
  -- strict thread matches above, loose matches on Subject: below --
2026-05-13  8:37 Tim JH Chen(陳仁鴻)
2026-05-15  0:19 ` Jakub Kicinski

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=5315a3b0-4fc8-4d74-8e27-d7b84bd794c9@redhat.com \
    --to=pabeni@redhat.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=chandrashekar.devegowda@intel.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=haijun.liu@mediatek.com \
    --cc=johannes@sipsolutions.net \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=loic.poulain@oss.qualcomm.com \
    --cc=netdev@vger.kernel.org \
    --cc=ricardo.martinez@linux.intel.com \
    --cc=ryazanov.s.a@gmail.com \
    --cc=tim.jh.chen@wnc.com.tw \
    --cc=tim770802@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®