mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Igor Paunovic <royalnet026@gmail.com>
To: Tomeu Vizoso <tomeu@tomeuvizoso.net>,
	Oded Gabbay <ogabbay@kernel.org>,
	Heiko Stuebner <heiko@sntech.de>
Cc: "Rob Herring" <robh@kernel.org>,
	"Krzysztof Kozlowski" <krzk+dt@kernel.org>,
	"Conor Dooley" <conor+dt@kernel.org>,
	"Jeff Hugo" <jeff.hugo@oss.qualcomm.com>,
	"Robert Foss" <rfoss@kernel.org>,
	"Sidong Yang" <sidong.yang@furiosa.ai>,
	"Diederik de Haas" <diederik@cknow-tech.com>,
	"Sebastian Reichel" <sebastian.reichel@collabora.com>,
	"Jiaxing Hu" <gahing@gahingwoo.com>,
	"Nicolas Dufresne" <nicolas@ndufresne.ca>,
	"Jonas Karlman" <jonas@kwiboo.se>,
	"Guangshuo Li" <lgs201920130244@gmail.com>,
	"Hüseyin BIYIK" <boogiepop@gmx.com>,
	dri-devel@lists.freedesktop.org,
	linux-rockchip@lists.infradead.org,
	linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org,
	linux-kernel@vger.kernel.org,
	"Igor Paunovic" <royalnet026@gmail.com>
Subject: [PATCH v2 02/11] accel/rocket: number the cores by devicetree position, not bind order
Date: Tue, 22 Sep 2026 10:01:05 +0200	[thread overview]
Message-ID: <20260922080114.44662-3-royalnet026@gmail.com> (raw)
In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com>

rocket_job_hw_submit() programs the S_POINTER registers of a core with an
extra bit derived from core->index, the way the vendor driver derives it
from the hardware number of the core. rocket_probe() sets core->index to
the slot the core takes in rdev->cores[], which is the order the cores
bind in.

The two agree only while the cores that bind are a prefix of the core
nodes in the devicetree, in devicetree order. Unbind them and bind them
back with a different core first, have one core's probe deferred behind a
sibling's, or disable a core other than the last one, and every task
submitted to a core whose slot is not its hardware number times out after
500 ms. The reset that follows does not help, and the inference finishes
with wrong output.

Observed on an Orange Pi 5 Plus with a KASAN build, over all six bind
orders of the three cores: only the devicetree order ran clean. The other
five produced 27 to 141 "NPU job timed out". In four of them no output
tensor changed with the input, and the harness gave up before its first
measured round; the fifth got through a six-second run with 27 timeouts
and a wrong top-1 class. All but one of the timeouts land on the cores
whose slot is not their hardware number, in proportion to the tasks the
scheduler hands them, and in both directions of the mismatch.

Number the cores by their position among the core nodes in the devicetree
instead, which is what the hardware number is.

The wrong value has been assigned since the driver was added, but it only
reached the hardware once the extra bit was introduced, hence the Fixes
tag below.

Fixes: 0810d5ad88a1 ("accel/rocket: Add job submission IOCTL")
Cc: stable@vger.kernel.org
Assisted-by: LLM sparse checkpatch
Signed-off-by: Igor Paunovic <royalnet026@gmail.com>
---
Supersedes the standalone posting:
https://lore.kernel.org/r/20260905135612.7324-1-royalnet026@gmail.com
Same diff. The message now says the numbers come from a KASAN build,
corrects the timeout range to 27-141 against the raw log (it said 140),
says what the four failed orders showed (the output did not change with
the input; it said an oracle rejected the output), notes the one timeout
that landed on a matching core, and drops the throughput figure, which
was measured under KASAN.

 drivers/accel/rocket/rocket_core.h |  5 +++++
 drivers/accel/rocket/rocket_drv.c  | 31 +++++++++++++++++++++++++++++-
 2 files changed, 35 insertions(+), 1 deletion(-)

diff --git a/drivers/accel/rocket/rocket_core.h b/drivers/accel/rocket/rocket_core.h
index f6d7382854ca9..46ed8352a79d2 100644
--- a/drivers/accel/rocket/rocket_core.h
+++ b/drivers/accel/rocket/rocket_core.h
@@ -30,6 +30,11 @@
 struct rocket_core {
 	struct device *dev;
 	struct rocket_device *rdev;
+	/*
+	 * Hardware number of the core: its position among the core nodes in
+	 * the devicetree. Not an index into rdev->cores[] - that slot is what
+	 * find_core_for_dev() returns.
+	 */
 	unsigned int index;
 
 	int irq;
diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocket_drv.c
index 2bcfe4ab3c68f..7d927bb6b322d 100644
--- a/drivers/accel/rocket/rocket_drv.c
+++ b/drivers/accel/rocket/rocket_drv.c
@@ -157,10 +157,39 @@ static const struct drm_driver rocket_drm_driver = {
 	.desc			= "rocket DRM",
 };
 
+/*
+ * The extra bit that rocket_job_hw_submit() sets in the S_POINTER registers
+ * is the hardware number of the core, which is its position among the core
+ * nodes in the devicetree: a disabled core keeps its number. The slot a core
+ * takes in rdev->cores[] is the order the cores happened to bind in, and the
+ * two only agree while the cores that bind are a prefix of those nodes, in
+ * devicetree order. Every task submitted to a core whose slot is not its
+ * hardware number then times out.
+ */
+static int rocket_core_hw_index(struct device *dev)
+{
+	struct device_node *np;
+	int index = 0;
+
+	for_each_matching_node(np, dev->driver->of_match_table) {
+		if (np == dev->of_node) {
+			of_node_put(np);
+			return index;
+		}
+		index++;
+	}
+
+	return -ENODEV;
+}
+
 static int rocket_probe(struct platform_device *pdev)
 {
+	int index = rocket_core_hw_index(&pdev->dev);
 	int ret;
 
+	if (index < 0)
+		return index;
+
 	if (rdev == NULL) {
 		/* First core probing, initialize DRM device. */
 		rdev = rocket_device_init(drm_dev, &rocket_drm_driver);
@@ -176,7 +205,7 @@ static int rocket_probe(struct platform_device *pdev)
 
 	rdev->cores[core].rdev = rdev;
 	rdev->cores[core].dev = &pdev->dev;
-	rdev->cores[core].index = core;
+	rdev->cores[core].index = index;
 
 	rdev->num_cores++;
 
-- 
2.43.0


  parent reply	other threads:[~2026-09-22  8:01 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22  8:01 [PATCH v2 00/11] accel/rocket: DVFS for the RK3588 NPU Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 01/11] accel/rocket: search every core slot when a core is removed Igor Paunovic
2026-09-22  8:01 ` Igor Paunovic [this message]
2026-09-22  8:01 ` [PATCH v2 03/11] accel/rocket: search every core slot when looking up a scheduler Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 04/11] accel/rocket: keep core slots stable across unbind and rebind Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 05/11] accel/rocket: request the core clocks by name Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 06/11] dt-bindings: npu: rockchip: allow DVFS and thermal properties Igor Paunovic
2026-09-22 16:06   ` Rob Herring
2026-09-23  8:57     ` Igor Paunovic
2026-09-23  9:15       ` Diederik de Haas
2026-09-23  9:43         ` Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 07/11] arm64: dts: rockchip: rk3588: add an OPP table for the NPU Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 08/11] accel/rocket: restore the NPU clock boot rate before powering the cores down Igor Paunovic
     [not found]   ` <20260922081326.B46651F000FF@smtp.kernel.org>
2026-09-22  8:55     ` Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 09/11] accel/rocket: add devfreq support Igor Paunovic
     [not found]   ` <20260922081855.160451F00893@smtp.kernel.org>
2026-09-22  8:56     ` Igor Paunovic
2026-09-23 13:14   ` Sidong Yang
2026-09-23 14:26     ` Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 10/11] accel/rocket: register a devfreq cooling device Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 11/11] arm64: dts: rockchip: rk3588: add passive cooling to the NPU thermal zone Igor Paunovic
2026-09-23 19:29 ` [PATCH v2 00/11] accel/rocket: DVFS for the RK3588 NPU Nicolas Dufresne
2026-09-23 19:54   ` Igor Paunovic
2026-09-24  5:37   ` Tomeu Vizoso
2026-09-24  7:36     ` Igor Paunovic

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922080114.44662-3-royalnet026@gmail.com \
    --to=royalnet026@gmail.com \
    --cc=boogiepop@gmx.com \
    --cc=conor+dt@kernel.org \
    --cc=devicetree@vger.kernel.org \
    --cc=diederik@cknow-tech.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=gahing@gahingwoo.com \
    --cc=heiko@sntech.de \
    --cc=jeff.hugo@oss.qualcomm.com \
    --cc=jonas@kwiboo.se \
    --cc=krzk+dt@kernel.org \
    --cc=lgs201920130244@gmail.com \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rockchip@lists.infradead.org \
    --cc=nicolas@ndufresne.ca \
    --cc=ogabbay@kernel.org \
    --cc=rfoss@kernel.org \
    --cc=robh@kernel.org \
    --cc=sebastian.reichel@collabora.com \
    --cc=sidong.yang@furiosa.ai \
    --cc=tomeu@tomeuvizoso.net \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®