mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Drew Fustini <fustini@kernel.org>
To: Reinette Chatre <reinette.chatre@intel.com>
Cc: "Adrien Ricciardi" <aricciardi@baylibre.com>,
	"Alexandre Ghiti" <alex@ghiti.fr>,
	"Albert Ou" <aou@eecs.berkeley.edu>,
	"Atish Kumar Patra" <atishp@rivosinc.com>,
	"Atish Patra" <atish.patra@linux.dev>,
	"Babu Moger" <babu.moger@amd.com>,
	"Ben Horgan" <ben.horgan@arm.com>,
	"Borislav Petkov" <bp@alien8.de>,
	"Chen Pei" <cp0613@linux.alibaba.com>,
	"Conor Dooley" <conor.dooley@microchip.com>,
	"Conor Dooley" <conor+dt@kernel.org>,
	"Dave Hansen" <dave.hansen@linux.intel.com>,
	"Dave Martin" <Dave.Martin@arm.com>,
	"Fenghua Yu" <fenghuay@nvidia.com>,
	"Gong Shuai" <gong.shuai@sanechips.com.cn>,
	"Gong Shuai" <gsh517@gmail.com>,
	guo.wenjia23@zte.com.cn, "James Morse" <james.morse@arm.com>,
	"Kornel Dulęba" <mindal@semihalf.com>,
	"Krzysztof Kozlowski" <krzk+dt@kernel.org>,
	liu.qingtao2@zte.com.cn,
	"Liu Zhiwei" <zhiwei_liu@linux.alibaba.com>,
	"Palmer Dabbelt" <palmer@dabbelt.com>,
	"Paul Walmsley" <pjw@kernel.org>,
	"Peter Newman" <peternewman@google.com>,
	"Radim Krčmář" <rkrcmar@ventanamicro.com>,
	"Rob Herring" <robh@kernel.org>,
	"Samuel Holland" <samuel.holland@sifive.com>,
	"Sebastian Andrzej Siewior" <bigeasy@linutronix.de>,
	"Clark Williams" <clrkwllms@kernel.org>,
	"Steven Rostedt" <rostedt@goodmis.org>,
	"Tony Luck" <tony.luck@intel.com>,
	"Vasudevan Srinivasan" <vasu@rivosinc.com>,
	"Ved Shanbhogue" <ved@rivosinc.com>,
	"Weiwei Li" <liwei1518@gmail.com>,
	"yunhui cui" <cuiyunhui@bytedance.com>,
	"Zhanpeng Zhang" <zhangzhanpeng.jasper@bytedance.com>,
	linux-kernel@vger.kernel.org, linux-riscv@lists.infradead.org,
	x86@kernel.org, devicetree@vger.kernel.org,
	linux-rt-devel@lists.linux.dev, linux-doc@vger.kernel.org
Subject: Re: [PATCH v8 2/5] riscv_cbqri: resctrl: Add cache allocation via capacity block mask
Date: Mon, 21 Sep 2026 23:57:43 -0700	[thread overview]
Message-ID: <arIm5yXDZhkH3KLm@thelio> (raw)
In-Reply-To: <fb4813cf-b21e-43ca-9bd4-4b42d2214aa6@intel.com>

On Thu, Sep 17, 2026 at 05:54:35PM -0700, Reinette Chatre wrote:
> Hi Drew,

Hi, thanks for the review.

> On 9/17/26 9:39 AM, Drew Fustini wrote:
> > diff --git a/arch/riscv/include/asm/resctrl.h b/arch/riscv/include/asm/resctrl.h
> > new file mode 100644
> > index 000000000000..b08f4e12f7aa
> > --- /dev/null
> > +++ b/arch/riscv/include/asm/resctrl.h
> 
> ...
> 
> > +/**
> > + * resctrl_arch_alloc_capable() - any CBQRI controller exposes resctrl alloc
> > + *
> > + * Returns true once at least one CBQRI controller has successfully probed for
> > + * a resctrl-exposed cache capacity allocation feature. Only meaningful after
> > + * cbqri_resctrl_setup() runs at late_initcall.
> > + */
> > +bool resctrl_arch_alloc_capable(void);
> > +
> > +/**
> > + * resctrl_arch_mon_capable() - any CBQRI controller exposes resctrl monitoring
> > + *
> > + * The CBQRI driver implements capacity allocation only and wires up no
> > + * monitoring events, so this always returns false. fs/resctrl references it
> > + * unconditionally, hence the stub.
> > + */
> > +bool resctrl_arch_mon_capable(void);
> > +
> 
> fyi ... I aim to comment more details later in this patch but for now please note that
> there are plans to remove the above two hooks since resctrl self has needed
> information via the rdt_resource::alloc_capable and rdt_resource::mon_capable flags.
> 
> For reference:
> 	https://lore.kernel.org/lkml/20260916231320.14502-7-tony.luck@intel.com/
> 
> I see this has impact on this driver that I comment more below.

Thanks for letting me know. cbqri_resctrl_control_init() already sets
rdt_resource::alloc_capable, so I will drop exposed_alloc_capable. I
will have cbqri_resctrl_teardown() clear rdt_resource::alloc_capable.

Should I wait to drop the hook until Tony's series is applied?

> > diff --git a/drivers/resctrl/cbqri_resctrl.c b/drivers/resctrl/cbqri_resctrl.c
> > new file mode 100644
> > index 000000000000..0c4bfa2a7f43
> > --- /dev/null
> > +++ b/drivers/resctrl/cbqri_resctrl.c
> > @@ -0,0 +1,785 @@
> > +// SPDX-License-Identifier: GPL-2.0-only
> > +
> > +#define pr_fmt(fmt) "%s:%s: " fmt, KBUILD_MODNAME, __func__
> > +
> > +#include <linux/bitfield.h>
> > +#include <linux/cacheinfo.h>
> > +#include <linux/cleanup.h>
> > +#include <linux/riscv_cbqri.h>
> 
> Just an observation (I have no intention to comment on the style):
> The new c files in this series seem to make an effort to have the #include
> files organized alphabetically, except when if comes to the file above?

Good catch, I'll fix that.

> > +struct cbqri_resctrl_res {
> > +	struct cbqri_controller *ctrl;
> > +	struct rdt_resource     resctrl_res;
> > +	bool                    cdp_enabled;
> > +};
> > +
> > +struct cbqri_resctrl_dom {
> > +	struct rdt_ctrl_domain  resctrl_ctrl_dom;
> > +	struct cbqri_controller *hw_ctrl;
> > +};
> 
> Is cbqri_resctrl_dom::hw_ctrl necessary? From what I can tell it is
> initialized from cbqri_resctrl_res::ctrl when a new domain is created
> and thus identical in all domains that belong to a resource. 
> 
> It looks to me as though the resource is always available when the associated
> controller information is needed so it looks like just cbqri_resctrl_res::ctrl
> could do?

cbqri_resctrl_dom::hw_ctrl is needed when a cache level has more than
one controller. Each cache instance has its own register block, so a
domain has to reach its own controller.

> The way the data is organized results in potentially confusing code since, for
> example, when initializing the control values of a single domain
> (cbqri_init_domain_ctrlval()) the RCID count is obtained from the domain's
> cbqri_resctrl_dom::hw_ctrl::rcid_count but when initializing the control values of
> all domains (resctrl_arch_reset_all_ctrls()) the RCID count is obtained from the
> resource's cbqri_resctrl_res::ctrl::rcid_count instead of each domain's domain
> cbqri_resctrl_dom::hw_ctrl::rcid_count.

Good point, I will read the rcid_count from the resource.

> > +static void cbqri_resctrl_accumulate_caps(void)
> > +{
> > +	int rid;
> > +
> > +	for (rid = 0; rid < RDT_NUM_RESOURCES; rid++) {
> > +		struct cbqri_resctrl_res *hw_res = &cbqri_resctrl_resources[rid];
> > +
> > +		if (!hw_res->ctrl)
> > +			continue;
> > +		if (hw_res->ctrl->alloc_capable)
> > +			exposed_alloc_capable = true;
> > +	}
> > +}
> 
> Since it captures whether any of the resources are alloc_capable it looks like
> exposed_alloc_capable indeed reflects the same information as what resctrl will
> use after resctrl_arch_alloc_capable() is dropped (see patch linked earlier).
> Except that cbqri_resctrl_teardown() only resets exposed_alloc_capable but not
> the rdt_resource::mon_capable and rdt_resource::alloc_capable flags that the new
> helper will use to determine if a resource is capable of allocation or monitoring.
> 
> The new helper will thus change behavior. It looks like resctrl_exit() is only
> called on error path during initialization so it seems unlikely that those flags
> will be referenced by this driver.
> 
> If there are plans to call resctrl_exit() in other paths like MPAM driver does there
> may be issues since the resctrl filesystem may still be mounted and when user space
> unmounts it the rdt_resource::mon_capable and rdt_resource::alloc_capable will be
> used and result in some arch callbacks called.
> 
> I wonder if it may simplify driver initialization to call resctrl_init() _after_
> setting up the CPU online/offline handlers (this is how the x86 driver does it).

I don't think there will be a need to call resctrl_exit() in other
paths. I will move resctrl_init() after cpuhp_setup_state().

> > +static struct rdt_ctrl_domain *cbqri_create_ctrl_domain(struct cbqri_controller *ctrl,
> > +							struct rdt_resource *res,
> > +							unsigned int cpu, int dom_id)
> > +{
> > +	struct rdt_ctrl_domain *domain;
> > +	struct list_head *pos = NULL;
> > +	int err;
> > +
> > +	domain = cbqri_new_domain(ctrl);
> > +	if (!domain)
> > +		return ERR_PTR(-ENOMEM);
> > +
> > +	cpumask_set_cpu(cpu, &domain->hdr.cpu_mask);
> > +	domain->hdr.id = dom_id;
> > +	domain->hdr.type = RESCTRL_CTRL_DOMAIN;
> > +	domain->hdr.rid = res->rid;
> > +
> > +	err = cbqri_init_domain_ctrlval(res, domain);
> > +	if (err)
> > +		goto free;
> > +
> > +	err = resctrl_online_ctrl_domain(res, domain);
> > +	if (err)
> > +		goto free;
> > +
> > +	/*
> > +	 * Publish only after the domain is fully initialized and online, so a
> > +	 * reader walking the RCU list never sees a half-built domain.
> > +	 */
> > +	resctrl_find_domain(&res->ctrl_domains, dom_id, &pos);
> 
> The caller loops over the domain list to determine whether it needs to create
> a new domain or not and then this domain create code loops over the list again
> to determine where to insert the new domain. Could this perhaps be simplified if
> the first loop  cbqri_attach_cpu_to_all_ctrls()->cbqri_find_ctrl_domain() determines
> the position at the same time as determining the presence and pass that to
> this function to just do the insert without searching the list again?

Yes, that is a good idea. I will change this.

> > +/*
> > + * Attach a CPU to the capacity controller at each cache level whose cache
> > + * the CPU shares. On failure, detach the CPU from everything attached so
> > + * far: the cpuhp core does not run this state's offline teardown when its
> > + * startup fails, so a partial attach would otherwise leak into the domain
> > + * cpu_masks. Caller holds cbqri_domain_list_lock.
> > + */
> > +static int cbqri_attach_cpu_to_all_ctrls(unsigned int cpu)
> > +{
> > +	static const u32 levels[] = { 2, 3 };
> > +	struct cbqri_controller *ctrl, *c;
> > +	struct cbqri_resctrl_res *hw_res;
> > +	struct rdt_ctrl_domain *d;
> > +	struct cacheinfo *ci;
> > +	int i, rid;
> > +
> > +	lockdep_assert_held(&cbqri_domain_list_lock);
> > +
> > +	/*
> > +	 * Hold cbqri_controllers_lock across the walk so a controller
> > +	 * registered after boot cannot corrupt it. The register path takes
> > +	 * it as a leaf and never cbqri_domain_list_lock, so this nesting
> > +	 * cannot invert.
> > +	 */
> > +	guard(mutex)(&cbqri_controllers_lock);
> > +
> > +	for (i = 0; i < ARRAY_SIZE(levels); i++) {
> > +		ci = get_cpu_cacheinfo_level(cpu, levels[i]);
> > +		if (!ci)
> > +			continue;
> > +
> > +		rid = cbqri_cache_level_to_rid(levels[i]);
> > +		hw_res = &cbqri_resctrl_resources[rid];
> > +		if (!hw_res->ctrl)
> > +			continue;
> > +
> > +		/* The controller backing this CPU's cache at this level. */
> > +		ctrl = NULL;
> > +		list_for_each_entry(c, &cbqri_controllers, list) {
> > +			if (c->type == CBQRI_CONTROLLER_TYPE_CAPACITY &&> +			    c->alloc_capable &&
> > +			    c->cache.cache_level == levels[i] &&
> > +			    c->cache.cache_id == ci->id) {
> > +				ctrl = c;
> > +				break;
> 
> Is it necessary to loop over cbqri_controllers and repeat these tests? Above seems to
> duplicate the work done during initialization (cbqri_resctrl_pick_caches()) that resulted
> in initialization of cbqri_resctrl_res::ctrl so it seems that after testing for existence
> this function could just use cbqri_resctrl_res::ctrl without again referencing cbqri_controllers?
> 
> If I understand correctly it may be that new controllers appear in cbqri_controllers 
> after this driver is initialized and the resources are initialized so the CPU online/offline
> helpers may need to take care how any controllers in cbqri_controllers not seen by
> cbqri_resctrl_setup() are handled.

The problem is that every controller that passed cbqri_cc_caps_agree()
is forgotten except the first. I will change cbqri_resctrl_pick_caches()
to keep every controller accepted for a level and have the online path
look up the cpu's cache id in that set.

> > +static int __init cbqri_arch_late_init(void)
> > +{
> > +	int err;
> > +
> > +	if (!riscv_isa_extension_available(NULL, SSQOSID))
> > +		return -ENODEV;
> > +
> > +	err = cbqri_resctrl_setup();
> > +	if (err)
> > +		return err;
> > +
> > +	err = cpuhp_setup_state(CPUHP_AP_ONLINE_DYN, "cbqri:online",
> > +				cbqri_resctrl_online_cpu,
> > +				cbqri_resctrl_offline_cpu);
> > +	if (err < 0) {
> > +		cbqri_resctrl_teardown();
> 
> cbqri_resctrl_teardown() calls resctrl_exit() that will complain
> via WARN_ON_ONCE() if any domains exist at that time. It is not clear
> to me if this can be guaranteed here.

I think moving resctrl_init() after cpuhp_setup_state() should eliminate
the possibility since a cpuhp failure then has nothing to tear down.

Thanks,
Drew

  reply	other threads:[~2026-09-22  6:57 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-17 16:39 [PATCH v8 0/5] riscv: Add initial CBQRI resctrl support Drew Fustini
2026-09-17 16:39 ` [PATCH v8 1/5] riscv_cbqri: Add capacity controller probe and allocation device ops Drew Fustini
2026-09-17 16:39 ` [PATCH v8 2/5] riscv_cbqri: resctrl: Add cache allocation via capacity block mask Drew Fustini
2026-09-17 16:57   ` sashiko-bot
2026-09-17 19:59     ` Drew Fustini
2026-09-18  0:54   ` Reinette Chatre
2026-09-22  6:57     ` Drew Fustini [this message]
2026-09-23 16:21       ` Reinette Chatre
2026-09-23  9:13   ` [External] " yunhui cui
2026-09-25  0:29     ` Drew Fustini
2026-09-17 16:39 ` [PATCH v8 3/5] riscv_cbqri: Make RESCTRL_FS available on RISC-V Drew Fustini
2026-09-17 16:39 ` [PATCH v8 4/5] dt-bindings: riscv: Add binding for CBQRI controllers Drew Fustini
2026-09-17 16:39 ` [PATCH v8 5/5] riscv_cbqri: Add CBQRI capacity allocation platform driver Drew Fustini
2026-09-17 17:11   ` sashiko-bot
2026-09-18 20:10     ` Drew Fustini

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arIm5yXDZhkH3KLm@thelio \
    --to=fustini@kernel.org \
    --cc=Dave.Martin@arm.com \
    --cc=alex@ghiti.fr \
    --cc=aou@eecs.berkeley.edu \
    --cc=aricciardi@baylibre.com \
    --cc=atish.patra@linux.dev \
    --cc=atishp@rivosinc.com \
    --cc=babu.moger@amd.com \
    --cc=ben.horgan@arm.com \
    --cc=bigeasy@linutronix.de \
    --cc=bp@alien8.de \
    --cc=clrkwllms@kernel.org \
    --cc=conor+dt@kernel.org \
    --cc=conor.dooley@microchip.com \
    --cc=cp0613@linux.alibaba.com \
    --cc=cuiyunhui@bytedance.com \
    --cc=dave.hansen@linux.intel.com \
    --cc=devicetree@vger.kernel.org \
    --cc=fenghuay@nvidia.com \
    --cc=gong.shuai@sanechips.com.cn \
    --cc=gsh517@gmail.com \
    --cc=guo.wenjia23@zte.com.cn \
    --cc=james.morse@arm.com \
    --cc=krzk+dt@kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-riscv@lists.infradead.org \
    --cc=linux-rt-devel@lists.linux.dev \
    --cc=liu.qingtao2@zte.com.cn \
    --cc=liwei1518@gmail.com \
    --cc=mindal@semihalf.com \
    --cc=palmer@dabbelt.com \
    --cc=peternewman@google.com \
    --cc=pjw@kernel.org \
    --cc=reinette.chatre@intel.com \
    --cc=rkrcmar@ventanamicro.com \
    --cc=robh@kernel.org \
    --cc=rostedt@goodmis.org \
    --cc=samuel.holland@sifive.com \
    --cc=tony.luck@intel.com \
    --cc=vasu@rivosinc.com \
    --cc=ved@rivosinc.com \
    --cc=x86@kernel.org \
    --cc=zhangzhanpeng.jasper@bytedance.com \
    --cc=zhiwei_liu@linux.alibaba.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®