From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 4DEA148B384 for ; Tue, 5 May 2026 17:20:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778001642; cv=none; b=I11sT2v+l6YjIW/WYZtaqLQ/Xo95HmGipRL+z2mToC3ttMvh0jvKjHuN9Amcak7B4ILqxSa/LmW/3fra62L2rF3iDeSv6VwzP5WcuPa5FnPp2Zg84PGpqLGgEqC6EQXIBQLGcABgMg4Ik9IOosKcXyUNYkkA4vSwmKaMhybEdcY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778001642; c=relaxed/simple; bh=0JW3PemZataoqIeyVS1/+Q7/158FtiDLJy5gVHboy78=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=uxYsNdJxj919cxDChjHIiAzlFR52aqrDpR3jL86jOEjamnGzNxOGSle+6p5fVwvsgAjo4ut+gau7bB8Ev+DBLzqcHVM5eX4/6H+S7MTrnYWZ2t/jPUEItpeNVNdtNu/WKMyCOmiDgEgqZ/IsRCqknCmnslfsaAvFomNmu4Mf9M0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b=YaJu5gLi; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b="YaJu5gLi" Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 39C751BB0; Tue, 5 May 2026 10:20:34 -0700 (PDT) Received: from [192.168.178.100] (usa-sjc-mx-foss1.foss.arm.com [172.31.20.19]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id D12643F836; Tue, 5 May 2026 10:20:36 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=simple/simple; d=arm.com; s=foss; t=1778001639; bh=0JW3PemZataoqIeyVS1/+Q7/158FtiDLJy5gVHboy78=; h=Date:Subject:To:Cc:References:From:In-Reply-To:From; b=YaJu5gLiXhRGKEO5uFo3A6aVCEKu3y1ccFDz8Olckkm2/bzdvQp+ed2Jgfx4HT/u3 XxZShjM27jzRei2on7UDNW2MF2ekcJfbTaZusjMJBthdCGpC4jkQt4XSp8ERg7L7ZA BbYQgzcfgjJzuPAWQEGPIL8vYQSqXBoybcDFkVcw= Message-ID: <803d8684-585e-4f41-8d9a-d9984923c3f2@arm.com> Date: Tue, 5 May 2026 19:20:35 +0200 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH 3/5] sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection To: Andrea Righi , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot Cc: Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Christian Loehle , Koba Ko , Felix Abecassis , Balbir Singh , Joel Fernandes , Shrikanth Hegde , linux-kernel@vger.kernel.org References: <20260428144352.3575863-1-arighi@nvidia.com> <20260428144352.3575863-4-arighi@nvidia.com> From: Dietmar Eggemann Content-Language: en-GB In-Reply-To: <20260428144352.3575863-4-arighi@nvidia.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 28.04.26 16:41, Andrea Righi wrote: > On systems with asymmetric CPU capacity (e.g., ACPI/CPPC reporting > different per-core frequencies), the wakeup path uses I assume those CPPC systems w/ different per-core frequencies (like your Vera) are the only real one which would make use of this. Mobile big.LITTLE/DynamIQ don't have SMT. Phil mentioned other machines (PowerPC ?) which had issues with using select_idle_capacity(): https://lore.kernel.org/r/20260325124840.GA98184@pauld.westford.csb [...] > On an SMT system with asymmetric CPU capacities, SMT-aware idle > selection has been shown to improve throughput by around 15-18% for > CPU-bound workloads, running an amount of tasks equal to the amount of > SMT cores. Just to make sure, this should be your internal NVBLAS benchmark. Is this 'ASYM (mainline) vs. ASYM + SMT' or 'NO_ASYM vs. ASYM + SMT' ? I try to match the cover letter's table numbers. [...] > @@ -7997,8 +8013,9 @@ static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool > static int > select_idle_capacity(struct task_struct *p, struct sched_domain *sd, int target) > { > + bool prefers_idle_core = sched_smt_active() && test_idle_cores(target); nit: why prefers_idle_core and not has_idle_core like in sis()? [...] > @@ -8047,12 +8102,17 @@ static inline bool asym_fits_cpu(unsigned long util, > unsigned long util_max, > int cpu) > { > - if (sched_asym_cpucap_active()) > + if (sched_asym_cpucap_active()) { > /* > * Return true only if the cpu fully fits the task requirements > * which include the utilization and the performance hints. > + * > + * When SMT is active, also require that the core has no busy > + * siblings. > */ > - return (util_fits_cpu(util, util_min, util_max, cpu) > 0); > + return (!sched_smt_active() || is_core_idle(cpu)) && > + (util_fits_cpu(util, util_min, util_max, cpu) > 0); > + } Not sure whether this has been discussed already. This makes all early bailout conditions in sis() idle core aware for 'ASYM + SMT' but it's not for 'NO_ASYM'? Otherwise, LGTM.