From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f48.google.com (mail-pj1-f48.google.com [209.85.216.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 74F014EA382 for ; Wed, 13 May 2026 13:22:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.48 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778678549; cv=none; b=oxL5HUZYyS3+WTae7e0GAOpde8CsXA6yxEUV5ZwOdRbGo+7UD9g37z2lx4tvC0PwdZ3GDgn0327/jWRZ+2cdJrA6CXZsonwggdjGJzrFlqyDFT1Z+Cxptxx0uCekrBfEwGVIa/HksCQqshxpjWPjoMUpl0PCZUEpWC/WKa6w5NA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778678549; c=relaxed/simple; bh=T3b3TSRxiSzxDJ++dKuuEXK1youAm4DbF1CjY4NqZBI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=HSaNrq6L2PWuLXc+L/ZSC+k4ZRULCL/dWShls2sKgaZly2s8ICCLrsHJ48y3yQeCTGOGpQFgI996LnLnFuhTprxEJmutY6qACHO3e1LUTSXbKnh3ol7ONCwcBrm6scRK2O3xOyyAMaGe1L293VU5kUOv6ClssfucePIRoMqlW1E= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=GSYU7b2f; arc=none smtp.client-ip=209.85.216.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="GSYU7b2f" Received: by mail-pj1-f48.google.com with SMTP id 98e67ed59e1d1-3660daea6a5so3681473a91.1 for ; Wed, 13 May 2026 06:22:27 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1778678547; x=1779283347; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=5LFbqiG/RD0CCzo7+/AP1zWQqzbysWKSZNji4RwzSyY=; b=GSYU7b2fSffoFjDk8zbXvDUXYStPswBxym/b4Rw/Iva5VQvDf8q+BvUs/IC8oBeD5h zeGU9p63ybajsyXbLKyNPoFIco6prnacQSdkfhf8nyklkinFIymlWthrh7HjTHBRzBEv 05IxQ8cAncKrO0EBblLjPIj7Jd+I6QdzebMFUSxiZRG9sLfGSiYmhea4eIfKFhfnRPH+ QrqND88qg8VGuyaI3Zl3Z45RKDIzRDyuwLQSAGK8yNw4Ch+Yxie7wsc25fqpe+lqLQiR oldRljcQa4cR/5ouYogAwTB4fBoUvNQWpgUrLHbyrm1ODimR3yKlHzEczQAMib5KiXYe GIFA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1778678547; x=1779283347; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=5LFbqiG/RD0CCzo7+/AP1zWQqzbysWKSZNji4RwzSyY=; b=tVuzHxzJk+GELSWK7fhqEBtJ8gtjW1fFF24ApQie4FQtAqcNpEV/Fxv6VAeovkskfu ofaKOwugkHq0AWkJp+3JhgEIKzeLY+iLfytxcPy7Ap2LMVDDHnefKv8q3n6A88Ej+JaW Rg5B7mbYQaT3MD9QLGExvkBkR6qvwvLon1mUAqnFW6LaGbSiVb/7PvX3XrxPKh5Q6Xgg aCWZfpW2kUrcISRWolC6hUS3NODXqZ37v/dOdzgTy3y9sA7FLPA/aDQfXUp7maF+GpRg jyGwWsvyDoVcH6Tewzo+fiYIgdRFeFiStPLjuD8tPFAVgybEeswNh+HHSEIPyfyHvAE3 GPBA== X-Forwarded-Encrypted: i=1; AFNElJ9/br6bp9RaF1hruPiuJ73sakvyRq08aDdlR0tRoXp1aToOraPMwZ9om5Mi/aGV7bKxCgdP5cZmfrxGaQo=@vger.kernel.org X-Gm-Message-State: AOJu0YzYQi3oKqH+D6jz7unjvt9TcbysztiEO+4F4OVCJygNdvE3248/ qyaWOFjrdDdgq4PDuAUpu0ODvyeAcu6P5E6P5HBxFtKpVKd6rbJMKjuGvxEtoQubBcg= X-Gm-Gg: Acq92OE4/Rg50wuSYE5+ir6Ykbf0SXn+xlkDq8baU90w6YcspejONSwU6tkWe2p84dH MZP6Ryj8aS6OXFd6u97nKvSIBOYzFBgbQDeQoPzmmAHhV81kol8JwBUQKocilrrH2WGTt5T/7MD UFYli3BFUm3hVzKKbieoS80MyMEuvJq4r7WaEhDHAM+msV/Kfhg0iA5iX7EyE9xF17QrZU6MQly wUnVzuAIov/bZpPZCAKkhA/LG+gdYDH+BeGFwRH17MBy96bEJr4Q7jbsdkx71zTVsOfD/ps0mgD yLKafGpcks/7mkh6kXjYpxJSCDnmiON/FtLQkjPNGHpoK7POCqh89dD2UNOHWOT7Ifs244XBH/Y tfIMhDrE8pi7dFLpOcr46EKDMxfMIx+RFoSDTNynkYaSwRk5VziBXNKFHWED+28qC604V8llOVb bYlVmJ1ZChJOjGP91viL9gMEabF65Dq/dTSt9+rzUlJKdz4erisroTSnDgEa3y7Y7e1YIylZ0= X-Received: by 2002:a17:90a:da8b:b0:368:ddd7:abca with SMTP id 98e67ed59e1d1-368f79e8cfbmr3065574a91.26.1778678546422; Wed, 13 May 2026 06:22:26 -0700 (PDT) Received: from PXLDJ45XCM.bytedance.net ([61.213.176.10]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-368edf7cbc2sm3098406a91.14.2026.05.13.06.22.20 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Wed, 13 May 2026 06:22:25 -0700 (PDT) From: Muchun Song To: Andrew Morton , David Hildenbrand , Muchun Song , Oscar Salvador , Michael Ellerman , Madhavan Srinivasan Cc: Lorenzo Stoakes , "Liam R . Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Nicholas Piggin , Christophe Leroy , Ackerley Tng , Frank van der Linden , aneesh.kumar@linux.ibm.com, joao.m.martins@oracle.com, linux-mm@kvack.org, linuxppc-dev@lists.ozlabs.org, linux-kernel@vger.kernel.org, Muchun Song Subject: [PATCH v2 61/69] mm/hugetlb: Drop boot-time HVO handling for gigantic folios Date: Wed, 13 May 2026 21:20:26 +0800 Message-ID: <20260513132044.41690-15-songmuchun@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260513132044.41690-1-songmuchun@bytedance.com> References: <20260513130542.35604-1-songmuchun@bytedance.com> <20260513132044.41690-1-songmuchun@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit HugeTLB HVO is currently supported on x86-64, riscv64, and LoongArch. On x86-64 and riscv64, gigantic HugeTLB pages are larger than the section size, so the existing section-based vmemmap optimization infrastructure is already sufficient to cover the whole folio. On LoongArch, HugeTLB HVO is supported without gigantic HugeTLB pages. Therefore, boot-time HugeTLB HVO folios can rely on the section-based vmemmap optimization infrastructure directly, without the extra bulk optimization and fallback handling. Signed-off-by: Muchun Song --- mm/hugetlb.c | 25 ++++++------------------- mm/hugetlb_vmemmap.c | 21 ++++++--------------- mm/internal.h | 25 +++++++++++++++++++++++-- mm/sparse.c | 23 ----------------------- 4 files changed, 35 insertions(+), 59 deletions(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index bd136fc6aec0..3cb8fffb9e3e 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -3201,21 +3201,7 @@ static void __init prep_and_add_bootmem_folios(struct hstate *h, unsigned long flags; struct folio *folio, *tmp_f; - /* Send list for bulk vmemmap optimization processing */ - hugetlb_vmemmap_optimize_folios(h, folio_list); - list_for_each_entry_safe(folio, tmp_f, folio_list, lru) { - if (!folio_test_hugetlb_vmemmap_optimized(folio)) { - /* - * If HVO fails, initialize all tail struct pages - * We do not worry about potential long lock hold - * time as this is early in boot and there should - * be no contention. - */ - hugetlb_folio_init_tail_vmemmap(folio, h, - OPTIMIZED_FOLIO_VMEMMAP_NR_STRUCT_PAGES, - pages_per_huge_page(h)); - } hugetlb_bootmem_init_migratetype(folio, h); /* Subdivide locks to achieve better parallel performance */ spin_lock_irqsave(&hugetlb_lock, flags); @@ -3238,6 +3224,8 @@ static void __init gather_bootmem_prealloc_node(unsigned long nid) list_for_each_entry_safe(m, tm, &huge_boot_pages[nid], list) { struct page *page = virt_to_page(m); struct folio *folio = (void *)page; + unsigned long pfn = PHYS_PFN(__pa(m)); + unsigned long nr_pages = pages_per_huge_page(m->hstate); h = m->hstate; /* @@ -3251,13 +3239,12 @@ static void __init gather_bootmem_prealloc_node(unsigned long nid) VM_BUG_ON(!hstate_is_gigantic(h)); WARN_ON(folio_ref_count(folio) != 1); - hugetlb_folio_init_vmemmap(folio, h, - OPTIMIZED_FOLIO_VMEMMAP_NR_STRUCT_PAGES); + hugetlb_folio_init_vmemmap(folio, h, vmemmap_nr_struct_pages(pfn, nr_pages)); init_new_hugetlb_folio(folio); - if (order_vmemmap_optimizable(pfn_to_section_order(folio_pfn(folio)))) { + if (order_vmemmap_optimizable(pfn_to_section_order(pfn))) { folio_set_hugetlb_vmemmap_optimized(folio); - section_set_order_range(folio_pfn(folio), folio_nr_pages(folio), 0); + section_set_order_range(pfn, nr_pages, 0); } if (hugetlb_early_cma(h)) @@ -3274,7 +3261,7 @@ static void __init gather_bootmem_prealloc_node(unsigned long nid) * (via hugetlb_bootmem_init_migratetype), so skip it here. */ if (!folio_test_hugetlb_cma(folio)) - adjust_managed_page_count(page, pages_per_huge_page(h)); + adjust_managed_page_count(page, nr_pages); cond_resched(); } diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c index 1305bee1195a..d20d2ce13906 100644 --- a/mm/hugetlb_vmemmap.c +++ b/mm/hugetlb_vmemmap.c @@ -599,23 +599,17 @@ static int hugetlb_vmemmap_split_folio(const struct hstate *h, struct folio *fol void hugetlb_vmemmap_optimize_folios(struct hstate *h, struct list_head *folio_list) { struct folio *folio; - unsigned long nr_to_optimize = 0; LIST_HEAD(vmemmap_pages); unsigned long flags = VMEMMAP_REMAP_NO_TLB_FLUSH; - list_for_each_entry(folio, folio_list, lru) { - int ret; - - /* - * Bootmem gigantic folios may already be marked optimized when - * their vmemmap layout was prepared earlier, so skip them here. - */ - if (folio_test_hugetlb_vmemmap_optimized(folio)) - continue; + if (!vmemmap_should_optimize(h)) + return; - nr_to_optimize++; + if (list_empty(folio_list)) + return; - ret = hugetlb_vmemmap_split_folio(h, folio); + list_for_each_entry(folio, folio_list, lru) { + int ret = hugetlb_vmemmap_split_folio(h, folio); /* * Splitting the PMD requires allocating a page, thus let's fail @@ -627,9 +621,6 @@ void hugetlb_vmemmap_optimize_folios(struct hstate *h, struct list_head *folio_l break; } - if (!nr_to_optimize) - return; - flush_tlb_all(); list_for_each_entry(folio, folio_list, lru) { diff --git a/mm/internal.h b/mm/internal.h index aff7cebb1da4..416afdf7b2ec 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -949,6 +949,29 @@ void memmap_init_range(unsigned long, int, unsigned long, unsigned long, unsigned long, enum meminit_context, struct vmem_altmap *, int, bool); +static inline int vmemmap_nr_struct_pages(unsigned long pfn, unsigned long nr_pages) +{ + const unsigned int order = pfn_to_section_order(pfn); + const unsigned long pages_per_compound = 1UL << order; + + if (!order_vmemmap_optimizable(order)) + return nr_pages; + + if (order < PFN_SECTION_SHIFT) { + VM_WARN_ON_ONCE(!IS_ALIGNED(pfn | nr_pages, pages_per_compound)); + return OPTIMIZED_FOLIO_VMEMMAP_NR_STRUCT_PAGES * nr_pages / pages_per_compound; + } + + VM_WARN_ON_ONCE(!IS_ALIGNED(pfn | nr_pages, PAGES_PER_SECTION)); + /* Ensure the requested range does not cross a compound page boundary. */ + VM_WARN_ON_ONCE((pfn % pages_per_compound) + nr_pages > pages_per_compound); + + if (IS_ALIGNED(pfn, pages_per_compound)) + return OPTIMIZED_FOLIO_VMEMMAP_NR_STRUCT_PAGES; + + return 0; +} + /* * mm/sparse.c */ @@ -988,8 +1011,6 @@ static inline void __section_mark_present(struct mem_section *ms, ms->section_mem_map |= SECTION_MARKED_PRESENT; } -int vmemmap_nr_struct_pages(unsigned long pfn, unsigned long nr_pages); - static inline int section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pages) { VM_WARN_ON_ONCE(!IS_ALIGNED(pfn | nr_pages, PAGES_PER_SUBSECTION)); diff --git a/mm/sparse.c b/mm/sparse.c index 598da1651e49..21a0eb636fea 100644 --- a/mm/sparse.c +++ b/mm/sparse.c @@ -236,29 +236,6 @@ void __weak __meminit vmemmap_populate_print_last(void) { } -int __meminit vmemmap_nr_struct_pages(unsigned long pfn, unsigned long nr_pages) -{ - const unsigned int order = pfn_to_section_order(pfn); - const unsigned long pages_per_compound = 1UL << order; - - if (!order_vmemmap_optimizable(order)) - return nr_pages; - - if (order < PFN_SECTION_SHIFT) { - VM_WARN_ON_ONCE(!IS_ALIGNED(pfn | nr_pages, pages_per_compound)); - return OPTIMIZED_FOLIO_VMEMMAP_NR_STRUCT_PAGES * nr_pages / pages_per_compound; - } - - VM_WARN_ON_ONCE(!IS_ALIGNED(pfn | nr_pages, PAGES_PER_SECTION)); - /* Ensure the requested range does not cross a compound page boundary. */ - VM_WARN_ON_ONCE((pfn % pages_per_compound) + nr_pages > pages_per_compound); - - if (IS_ALIGNED(pfn, pages_per_compound)) - return OPTIMIZED_FOLIO_VMEMMAP_NR_STRUCT_PAGES; - - return 0; -} - /* * Initialize sparse on a specific node. The node spans [pnum_begin, pnum_end) * And number of present sections in this node is map_count. -- 2.54.0