From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3B65E387569 for ; Tue, 19 May 2026 19:27:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779218833; cv=none; b=FbdPDNBXoIqjwK0oErJ7rFlAcusHOO8IPIRFZ/8V+/LqRZKEfyZluIzAurUpOMnCX3uAfTeU1Svgq4aWhMxx5bruEeYvsBJkuvPXoUFOAcrGowIpMOd7ZOMuwhbC4fhRSQWNCAOCtJVs9TgVarmeijKdvpX0pjYqRR8JBeFRRE8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779218833; c=relaxed/simple; bh=vKqiD5quSa784Z9Nc4Je9L+uRFDb56Kg0jbM90CRDbY=; h=Date:From:To:Cc:Subject:Message-Id:In-Reply-To:References: Mime-Version:Content-Type; b=e6l2C/mxo1MdPdEX3hxZ++XiLjpZ15roEc/1WW2Em/ZHjyOMo45xOiPulbV9GvAYoH3Rz0shJKHFYgucuJ4GEbqH9Mfd6NvIa8y/yBQPBXzr+ETLIm9kfJ0M15Aw10MGHJ2/ptkzxY0webmpiiPIjEQLndkTgVQO4wnLl1CBSFs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b=a0DQELKM; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b="a0DQELKM" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 870AF1F000E9; Tue, 19 May 2026 19:27:11 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux-foundation.org; s=korg; t=1779218831; bh=fxFwHBeP3nF1/C+zpz+/ZGN4ZMGPhZ04dCMyDEAnTzo=; h=Date:From:To:Cc:Subject:In-Reply-To:References; b=a0DQELKM2H7Bkj9ft8qMfCCtqm2YvQzoUsSSxjz1z7NGlPknAavUGATNUL8Onbzri M8v2FyaXkH4tqtXeQLySeHNhEnw+PbVad4+F8Cu0CwOLlol4S0cTOF/9ioSYEgrq+1 7kY0kpT1/VGWCQ3sBwOTAi9TD2GVFhBnF+L6t0ss= Date: Tue, 19 May 2026 12:27:11 -0700 From: Andrew Morton To: "JP Kobryn (Meta)" Cc: vbabka@kernel.org, surenb@google.com, mhocko@suse.com, jackmanb@google.com, hannes@cmpxchg.org, ziy@nvidia.com, linux-mm@kvack.org, usama.arif@linux.dev, kirill@shutemov.name, willy@infradead.org, linux-kernel@vger.kernel.org, kernel-team@meta.com Subject: Re: [PATCH] mm/page_alloc: skip high atomic reservation at or below costly order Message-Id: <20260519122711.1afe69455bcaf22a0962dce8@linux-foundation.org> In-Reply-To: <20260519012532.272770-1-jp.kobryn@linux.dev> References: <20260519012532.272770-1-jp.kobryn@linux.dev> X-Mailer: Sylpheed 3.8.0beta1 (GTK+ 2.24.33; x86_64-pc-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Mon, 18 May 2026 18:25:32 -0700 "JP Kobryn (Meta)" wrote: > We're seeing a pattern in production where 2MB THP order-9 allocations are > failing due to fragmentation and triggering reclaim on systems with plenty > of free memory. Over time, the success rate of these THP allocations do not > increase at all. > > Inspecting zone->vm_stat[NR_FREE_PAGES] via kprobe on compaction_suitable() > indicated the given zone had sufficient free pages for order-9 allocations, > yet they were going unused. Drilling down into the zone and inspecting > /proc/pagetypeinfo revealed why. Order-9 blocks were accumulating in the > zone's HighAtomic bucket (while zero were present in Movable). THP is > unable to draw blocks from HighAtomic since that bucket is not in the > fallback list. > > The heuristic for reserving pageblocks in HighAtomic is that any atomic > allocation greater than order-0 will result in the full pageblock being > captured. This means that an order-1 atomic allocation will over-reserve by > 256x, a full 512 pageblock. > > Gate the reservation on order. Skip for allocations at or below > PAGE_ALLOC_COSTLY_ORDER. This prevents smaller atomic allocations from > reserving entire pageblocks, and significantly helps when THP is in use on > a fragmented but otherwise healthy system. > > Testing was performed using an A/B instagram workload receiving prod > traffic. Each side had ~60 hosts with 64G memory. The patch resulted in > several gains: > > Unpatched > HighAtomic pageblocks per host: 309-312 (1% of zone or 620MB), > ...all order-9 blocks in HighAtomic > THP success rate: 1-6% > Compaction success rate: 0-2% > pgscan_kswapd (total across ~60 hosts, per minute): ~70.2M > Atomic order-4+ allocations: 0 > > Patched > HighAtomic pageblocks per host: 1 > THP success rate: 44-78% > Compaction success rate: 24-47% > pgscan_kswapd (total across ~60 hosts, per minute): ~29.9M > Atomic order-4+ allocations: 0 > > Note that for this workload all atomic allocations were order 0-3 > originating from the network stack, btrfs, and scheduler. > > ... > > --- a/mm/page_alloc.c > +++ b/mm/page_alloc.c > @@ -3446,6 +3446,13 @@ static void reserve_highatomic_pageblock(struct page *page, int order, > int mt; > unsigned long max_managed; > > + /* > + * Don't reserve a pageblock for lower orders. > + * Order 1-3 allocs should not capture a huge page size block. > + */ > + if (order <= PAGE_ALLOC_COSTLY_ORDER) > + return; > + > /* > * The number reserved as: minimum is 1 pageblock, maximum is > * roughly 1% of a zone. But if 1% of a zone falls below a Sashiko asked : Does skipping the HighAtomic reservation for orders 1-3 break the : anti-fragmentation guarantees for these atomic allocations? : : The MIGRATE_HIGHATOMIC reserve protects high-order atomic allocations : from failing under fragmentation by taking ownership of the entire : pageblock. : : If order-1 through order-3 atomic allocations fall back to stealing : pages, but the pageblock remains in its original migratetype, won't : order-0 non-atomic allocations consume the remaining contiguous space? : : Under memory pressure, this could leave no contiguous blocks for atomic : allocations to steal. Because these atomic allocations cannot trigger : direct reclaim or compaction, they might fail, potentially leading to : dropped packets or I/O errors in subsystems like the network stack or : BTRFS. : : Could background compaction or khugepaged be used to unreserve : HighAtomic blocks dynamically instead of disabling the reserve for : these orders?