From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from galois.linutronix.de (Galois.linutronix.de [193.142.43.55]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 62DB93403F7; Tue, 30 Jun 2026 09:04:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=193.142.43.55 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782810244; cv=none; b=o4kJSxYWS1hI0nI4aztiBuA1FnkKGnvuyFkT7eP7ZLKAchhZuOvMZ1aLOfy0hZ1aGMPkbwAFG6Gbtam0Y3eePHQurGfO2PvOu4ZRTewpXdzj9Veq2hivQkxWxdjUdegdYlQH6covozUc7f6tZlJZi/r2UxqkLNGxTMtVDbDn1qY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782810244; c=relaxed/simple; bh=ty1kHCbazdVJlTus6GaM2R/dae9I1dkqhb9PNH0VAuQ=; h=Date:From:To:Subject:Cc:In-Reply-To:References:MIME-Version: Message-ID:Content-Type; b=Iu0pWbGeCDG62e2N/8EwmuMmFyxbDQx8BbbSVnyGzf/eQG9X4F+fkBuDqI7RpxUkZBc07WwcUgSHAJ5vjHwhqBC7TgzuO8+1UftWDvAZCvjpwzxU3Y3eHRVFc5ifR8T1PyD5v2dmcsnNcbrNWk6kXkyTvqQJr/p72pn151FbYJU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de; spf=pass smtp.mailfrom=linutronix.de; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=glfRIcyN; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=XCUCTi6c; arc=none smtp.client-ip=193.142.43.55 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linutronix.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="glfRIcyN"; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="XCUCTi6c" Date: Tue, 30 Jun 2026 09:03:56 -0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020; t=1782810238; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=mXu6Ov5p0Q6Wwe0Z4hPg9DioBiEgGq+nlD87qeKGk9g=; b=glfRIcyNCeGa7Opb0FT3411fCzTurna9DIlXiAJnO90zxQQt3R+8PqHMBvBpEXj/ZNYAwj uIV6nIva3k67mCYkhh7dQhzh7sYpFpC5b/YxTFpfWoeW1j+1le9feVK7/OXOUm6D547uZN qKAMo9uUo8/i6k10hD7hebnPHqzqOKXsj+nQOLvfSBX7vbbTGOoO0237JGKyZR30Q/NO3/ kDU+FblwVhjN9flUaCCj2aLYBBRk/z4pX+J16yG+sRXPOToipc38pGUuxxb2LKvTYJMUzz UGI4XCqtw31d/JZ8Mmr2oSpYYrqPdHdmbj3OeOJOxRF/v/RT6F9s2NQXsqtlSw== DKIM-Signature: v=1; a=ed25519-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020e; t=1782810238; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=mXu6Ov5p0Q6Wwe0Z4hPg9DioBiEgGq+nlD87qeKGk9g=; b=XCUCTi6c0Zrw0+rjGQW2R+FqqhXemtiSnUefkH04xSwFF39hK/uHNJD3ef+KgI51YW+Jzg qyKFeGz4yEfS8xDw== From: "tip-bot2 for Peter Zijlstra" Sender: tip-bot2@linutronix.de Reply-to: linux-kernel@vger.kernel.org To: linux-tip-commits@vger.kernel.org Subject: [tip: sched/core] sched/fair: Add cgroup_mode switch Cc: "Peter Zijlstra (Intel)" , x86@kernel.org, linux-kernel@vger.kernel.org In-Reply-To: <20260605124051.338602724@infradead.org> References: <20260605124051.338602724@infradead.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Message-ID: <178281023686.3843924.13107928847097568951.tip-bot2@tip-bot2> Robot-ID: Robot-Unsubscribe: Contact to get blacklisted from these emails Precedence: bulk Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable The following commit has been merged into the sched/core branch of tip: Commit-ID: 4161cb2d867b2210b447afc4a941ab4b96e1bd13 Gitweb: https://git.kernel.org/tip/4161cb2d867b2210b447afc4a941ab4b96e= 1bd13 Author: Peter Zijlstra AuthorDate: Thu, 12 Mar 2026 14:48:44 +01:00 Committer: Peter Zijlstra CommitterDate: Tue, 30 Jun 2026 10:56:51 +02:00 sched/fair: Add cgroup_mode switch The effective task weight (W_t') for a task in cgroup g on CPU n is given by: W_t W_t' =3D W_g * F_g_n * ---------- \Sum W_t_n Where W_g is the group's weight (cpu.weight), F_g_n is the fraction of the group weight for CPU n and W_t/W is the relative weight of this task against all other tasks in the same group on the same CPU. Furthermore, this makes: \Sum W_t_n F_g_n =3D ---------- \Sum W_t The fraction of weight inside the group of CPU n against the whole group. The problem is with F_g_n, the primary goal of this fraction is to make sure that the relative weight of tasks, when distributed over CPUs is maintained. For example, consider 4 (equal weight) tasks and 2 CPUs with a 1:3 distribution, then if F_g_n would simply be 1 (no weight re-distribution) the effective relative weights (W_t') of the tasks in our group would be: CPU0 CPU1 W_g W_g/3 W_g/3 W_g/3 IOW, the lucky task on CPU0 would get an equal amount of weight as all 3 tasks on CPU1 combined. However, with the weight redistribution, this becomes: CPU0 CPU1 W_g/4 W_g/4 W_g/4 W_g/4 All tasks are equal weight (as intended). However, as is already evident from this example, the more CPUs you add, the smaller F_g_n becomes, which creates= a disparity against tasks not in our group. Specifically: avg(F_g_n) ~ 1/N This leads to a weight mismatch in the hierarchy. IOW tasks cannot compete fairly across hierarchy levels. *Notably*, what is meant by avg(F_g_n) being proportional to 1/N is that when there are at least N runnable tasks, the average of this fraction tends to 1/= N. For a hierarchy of depth d, this gets even worse, since that gets terms on the order of: avg(F_g_n)^d ~ 1/(N^d) Given fixed point arithmetic, this also leads to numerical trouble. However, the meaning of "cpu.weight" is simple and intiutive: the total weight of the cgroup. But as explored above, there is deception in this simplicity. Prepare to add a few alternative methods for distributing weight. Signed-off-by: Peter Zijlstra (Intel) Link: https://patch.msgid.link/20260605124051.338602724%40infradead.org --- kernel/sched/debug.c | 74 +++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 74 insertions(+) diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c index 40584b2..507e486 100644 --- a/kernel/sched/debug.c +++ b/kernel/sched/debug.c @@ -633,6 +633,76 @@ static void debugfs_fair_server_init(void) } } =20 +#ifdef CONFIG_FAIR_GROUP_SCHED +static int cgroup_mode =3D 0; + +static const char *cgroup_mode_str[] =3D { + "smp", +}; + +static int sched_cgroup_mode(const char *str) +{ + for (int i =3D 0; i < ARRAY_SIZE(cgroup_mode_str); i++) { + if (!strcmp(str, cgroup_mode_str[i])) + return i; + } + return -EINVAL; +} + +static ssize_t sched_cgroup_write(struct file *filp, const char __user *ubuf, + size_t cnt, loff_t *ppos) +{ + char buf[16]; + int mode; + + if (cnt > 15) + cnt =3D 15; + + if (copy_from_user(buf, ubuf, cnt)) + return -EFAULT; + + buf[cnt] =3D 0; + mode =3D sched_cgroup_mode(strstrip(buf)); + if (mode < 0) + return mode; + + WRITE_ONCE(cgroup_mode, mode); + + *ppos +=3D cnt; + return cnt; +} + +static int sched_cgroup_show(struct seq_file *m, void *v) +{ + int mode =3D READ_ONCE(cgroup_mode); + + for (int i =3D 0; i < ARRAY_SIZE(cgroup_mode_str); i++) { + if (mode =3D=3D i) + seq_puts(m, "("); + seq_puts(m, cgroup_mode_str[i]); + if (mode =3D=3D i) + seq_puts(m, ")"); + + seq_puts(m, " "); + } + seq_puts(m, "\n"); + return 0; +} + +static int sched_cgroup_open(struct inode *inode, struct file *filp) +{ + return single_open(filp, sched_cgroup_show, NULL); +} + +static const struct file_operations sched_cgroup_fops =3D { + .open =3D sched_cgroup_open, + .write =3D sched_cgroup_write, + .read =3D seq_read, + .llseek =3D seq_lseek, + .release =3D single_release, +}; +#endif + static __init int sched_init_debug(void) { struct dentry __maybe_unused *numa, *llc; @@ -686,6 +756,10 @@ static __init int sched_init_debug(void) =20 debugfs_create_file("debug", 0444, debugfs_sched, NULL, &sched_debug_fops); =20 +#ifdef CONFIG_FAIR_GROUP_SCHED + debugfs_create_file("cgroup_mode", 0644, debugfs_sched, NULL, &sched_cgroup= _fops); +#endif + debugfs_fair_server_init(); #ifdef CONFIG_SCHED_CLASS_EXT debugfs_ext_server_init();