From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from fout-b8-smtp.messagingengine.com (fout-b8-smtp.messagingengine.com [202.12.124.151]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E5A0E4A92F1 for ; Wed, 16 Sep 2026 09:32:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.151 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789551138; cv=none; b=YMaUPg6lRWEk9MjcRpAdZ8UDdBRG5xD1otvNOKcirtmf5vhkoTBFPiwPQu91srqKEu8EmlGCFStGj92XFPpx+JnXHgpsF/cyfhrdY5xggI+7zz+irjYqYjmJ33hFFN1LWJLezApPRUhxWQCSsaRVO6dZP4c1qqmZtrwUjcBCBTw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789551138; c=relaxed/simple; bh=0vsgegkCYa7op9rv9OivvWESbkik5p8dMrwmdH57rbs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Ok5X1mYE8aJWTSYv7AA0Kh0USz1/JR64Gfa5M+Q6Yb8/0IHorYYUYy9GEMeqVNBDXFCNZLU07i8jle8s7G/4QQDvJArsINHg6DBQ41Q2T0/q5uMrRZi3Nw9++T8CXtpoGC8gqvFxugs/BRieQoLb/84g8gcEmWAISgulv3k0SeA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=lkTwwmgm; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=IfHcTkeL; arc=none smtp.client-ip=202.12.124.151 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="lkTwwmgm"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="IfHcTkeL" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.stl.internal (Postfix) with ESMTP id 28C0F1D00098; Wed, 16 Sep 2026 05:32:02 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Wed, 16 Sep 2026 05:32:02 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm2; t=1789551122; x= 1789637522; bh=ICF9OPiue2yxmh8hP/ZUHQcKxFznzKsqrqQwWUt5ISk=; b=l kTwwmgmVxsBasNJphP2tCEw2dpkK2tdbDfRwh1Wn6LKKld4EuTVM8c1XZHRHkm+V THhHH3mh4DUAc1aBNuwdbcOzJDJqjZby1gn80iUHnDXsIm2zWN5kIYakLBKoDdci G9nzvooF83s2B0NlQG6jf62GaJ3DpFfRLOuY/cpCtlcz9Dv6ga7vax8DYx2DXHw5 UcYxqUPBFlrx0bz9Kt6d+HIK+e7R2wy+jYmcNldvYpdpQONVo3tmCiEb6kTNcxlf SBb6Ay9yKpIGFVh5X85F5XzGjrxMKESmxWVofwEWbWJGZIDso9jMhB6xua0qOQkD YzgW0Vj8BKPrWdt6oYtOQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm1; t=1789551122; x=1789637522; bh=I CF9OPiue2yxmh8hP/ZUHQcKxFznzKsqrqQwWUt5ISk=; b=IfHcTkeLQ9ubDaRHX xvvXfbj63w4k4jOZQ31eTezzSh+0fKBJPJOlgqQqeFw1cvxCggrNRs1nnRMUVJlw c/woxe7Cmtp+y77D1khAQRBQG+D8eoVmW1PkmtCiRfaO0ihzKYVMrRnrnKCB06cE R+qVSjSxl0qCWJzjOOD0YUyh98deQRE57vo4upJuuUhT5Dk7gK5f4lIjQZlBK8eI KdrujaJTAXyl283MIN1FfRyLf3OwZJEy7xxvkzbaE6PloPGmNVAtQw4DO2J4Km92 9Nh8yAawStltZzyoS91Er9RZPw+y9C5Cb7YPifrDMrpaXejNkkwhi7NxdxCsvaLY 5qItA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTE31qL+uaIHQKQPvpaICSxxs7QGWvIquuW8cKaqYQWKvt9pKCqy5bgk//TApl3wTl ad4FQSRAknjhMkzVuMAZNh0T12ywqROLvU4AnUXGVwd2kiN78Ar+1NR+fesCkATu1LALh1 4LvX/9qk4pAfO725ECjoAkI82P3z755hZxEkAV0An15e7Wr2vIbgLFM815SfkNmTU/eBmL ORNXO2IPaHkpkCtL9iOfg4dR3lpAdl3Xvq/Y35opAbs/zzbuDBIkBs44NTO3w5x5Bedurt thFl4kzzMjhv7UlAxOBz4vPRXQMYZm6skstaNA0N2ZkZEKrpjBUX5624wA8G4zEpHd89Vf 9T6n7sacdOaicBHl0a9Aiw51JNeA9NYlEYhEhIDx9KYOW0onA68maUQsM43/Hp6DjkU+Nu 72ekwCSo3DXmCWJ6wX+0Gd6ukbM49sUWHHNcWEe2Vr1f863dlft3DSoweFQ83vga9K0ojQ CYUQdYMItjRfsJ/xU+UsSJn4UiKGPcJsiAU3mczm9byTBs3kPOTyf/ZuA2Iqc3vN3ViBj6 gO9wjuvLNdFVs/TsSWTql3QMRjAHfcVzbmQP16qMqIAuCI0bGGcqhqa5juTxSnna4X1a37 vIlEH/Kvcys3ycZrZsrYEZ6/X0mz3FFd+rZMbzP7j/Xp40ENrkR1QO3B00FA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 16 Sep 2026 05:32:01 -0400 (EDT) From: Kiryl Shutsemau To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang Cc: "Kiryl Shutsemau (Meta)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Vlastimil Babka , Jann Horn Subject: [PATCH v3 05/12] mm/collapse: state what a collapse may do in the policy Date: Wed, 16 Sep 2026 10:31:32 +0100 Message-ID: <20260916093145.4022188-6-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260916093145.4022188-1-kirill@shutemov.name> References: <20260916093145.4022188-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit From: "Kiryl Shutsemau (Meta)" Tests scattered through the collapse path decide what a collapse is allowed to do by asking whether khugepaged started it. Between them they settle: - which VMAs are eligible, and how hard to try for a folio; - how many empty, swapped-out or shared PTEs a window may contain, and whether a sub-PMD window is held to a stricter rule than a PMD; - whether a range has to look used, and whether a MADV_FREE'd page is left alone; - whether the PMD is mapped as part of the request, and whether dirty pages are worth writing back and retrying. None of those is a fact about khugepaged. Each is something the caller decided before asking, and the collapse code should not have to look up who called to find out. Add struct collapse_policy for the caller to fill: khugepaged from its own settings, MADV_COLLAPSE from the fact that a user asked explicitly. Every test becomes a read of a field, and cc->is_khugepaged goes, having no reader left. khugepaged fills the policy once per scan pass, MADV_COLLAPSE once per call. That is the one change in behaviour. The max_ptes_* limits and the defrag setting behind the allocation mask are sampled once per pass rather than on every table. A table scanned early in a pass and one scanned late are then treated alike. collapse_file() also drops a NULL check on the collapse_control. It has one call site, reached only from collapse_single_pmd(), which dereferences cc unconditionally, so the check was already dead. Assisted-by: LLM Reviewed-by: Zi Yan Reviewed-by: Baolin Wang Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.h | 31 ++++++++++++- mm/khugepaged.c | 114 ++++++++++++++++++++++++++---------------------- 2 files changed, 93 insertions(+), 52 deletions(-) diff --git a/mm/collapse.h b/mm/collapse.h index b115034d9018..7044dc71c7c2 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -45,8 +45,37 @@ enum scan_result { SCAN_PAGE_DIRTY_OR_WRITEBACK, }; +/* What a collapse is allowed to do, decided by the caller that asks for it */ +struct collapse_policy { + /* Limits, stated per PMD; HPAGE_PMD_NR means "no limit" */ + unsigned int max_ptes_none; + unsigned int max_ptes_swap; + unsigned int max_ptes_shared; + + /* Take no swapped-out or shared PTE into a sub-PMD collapse */ + bool strict_sub_pmd; + + /* Leave clean lazyfree folios to reclaim rather than collapse them */ + bool skip_lazyfree; + + /* Refuse a range with no sign of use */ + bool require_referenced; + + /* Map the PMD over a file collapse instead of leaving it to a fault */ + bool install_pmd; + + /* Write dirty pages back and retry once instead of refusing them */ + bool writeback_dirty; + + /* How hard to try for a destination folio */ + gfp_t gfp; + + /* Which VMAs are eligible, as thp_vma_allowable_orders() spells it */ + enum tva_type tva_type; +}; + struct collapse_control { - bool is_khugepaged; + struct collapse_policy policy; /* Num pages scanned per node */ u32 node_load[MAX_NUMNODES]; diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 8889f75cf45f..cb08789b2d38 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -314,15 +314,12 @@ static bool pte_none_or_zero(pte_t pte) static unsigned int collapse_max_ptes_none(struct collapse_control *cc, struct vm_area_struct *vma, unsigned int order) { - const unsigned int max_ptes_none = khugepaged_max_ptes_none; + const unsigned int max_ptes_none = cc->policy.max_ptes_none; if (vma && userfaultfd_armed(vma)) return 0; - /* for MADV_COLLAPSE, allow any empty/shared zeropage PTEs */ - if (!cc->is_khugepaged) - return HPAGE_PMD_NR; - /* for PMD collapse, respect the user defined maximum */ - if (is_pmd_order(order)) + /* The limit as given, at the PMD order and wherever it is not capped */ + if (is_pmd_order(order) || !cc->policy.strict_sub_pmd) return max_ptes_none; /* * for mTHP collapse with the sysctl value set to COLLAPSE_MAX_PTES_LIMIT, @@ -354,19 +351,12 @@ static unsigned int collapse_max_ptes_shared(struct collapse_control *cc, unsigned int order) { /* - * For MADV_COLLAPSE, do not restrict the number of PTEs that map shared - * anonymous pages. + * A sub-PMD window held to the strict rule takes no shared page at all: + * an mTHP is not worth the CoW-breaking. */ - if (!cc->is_khugepaged) - return HPAGE_PMD_NR; - /* - * for mTHP collapse do not allow collapsing anonymous memory pages that - * are shared between processes. - */ - if (!is_pmd_order(order)) + if (!is_pmd_order(order) && cc->policy.strict_sub_pmd) return 0; - /* for PMD collapse, respect the user defined maximum */ - return khugepaged_max_ptes_shared; + return cc->policy.max_ptes_shared; } /** @@ -382,16 +372,12 @@ static unsigned int collapse_max_ptes_swap(struct collapse_control *cc, unsigned int order) { /* - * For MADV_COLLAPSE, do not restrict the number PTEs entries or - * pagecache entries that are non-present. + * A sub-PMD window held to the strict rule takes nothing non-present: + * reading pages back to build an mTHP is not worth the latency. */ - if (!cc->is_khugepaged) - return HPAGE_PMD_NR; - /* for mTHP collapse do not allow any non-present PTEs or pagecache entries */ - if (!is_pmd_order(order)) + if (!is_pmd_order(order) && cc->policy.strict_sub_pmd) return 0; - /* for PMD collapse, respect the user defined maximum */ - return khugepaged_max_ptes_swap; + return cc->policy.max_ptes_swap; } int hugepage_madvise(struct vm_area_struct *vma, @@ -678,7 +664,7 @@ static enum scan_result __collapse_huge_page_isolate(struct vm_area_struct *vma, * If the vma has the VM_DROPPABLE flag, the collapse will * preserve the lazyfree property without needing to skip. */ - if (cc->is_khugepaged && !(vma->vm_flags & VM_DROPPABLE) && + if (cc->policy.skip_lazyfree && !(vma->vm_flags & VM_DROPPABLE) && folio_test_lazyfree(folio) && !pte_dirty(pteval)) { result = SCAN_PAGE_LAZYFREE; goto out; @@ -767,12 +753,12 @@ static enum scan_result __collapse_huge_page_isolate(struct vm_area_struct *vma, if (folio_test_large(folio)) list_add_tail(&folio->lru, compound_pagelist); next: - if (cc->is_khugepaged && + if (cc->policy.require_referenced && folio_pte_referenced(folio, vma, addr, pteval)) referenced++; } - if (unlikely(cc->is_khugepaged && !referenced)) { + if (unlikely(cc->policy.require_referenced && !referenced)) { result = SCAN_LACK_REFERENCED_PAGE; } else { result = SCAN_SUCCEED; @@ -938,9 +924,7 @@ static void khugepaged_alloc_sleep(void) remove_wait_queue(&khugepaged_wait, &wait); } -static struct collapse_control khugepaged_collapse_control = { - .is_khugepaged = true, -}; +static struct collapse_control khugepaged_collapse_control; static bool collapse_scan_abort(int nid, struct collapse_control *cc) { @@ -976,6 +960,36 @@ static inline gfp_t alloc_hugepage_khugepaged_gfpmask(void) return khugepaged_defrag() ? GFP_TRANSHUGE : GFP_TRANSHUGE_LIGHT; } +/* khugepaged collapses on its own initiative, so it obeys its own settings */ +static void collapse_policy_khugepaged(struct collapse_policy *p) +{ + p->max_ptes_none = READ_ONCE(khugepaged_max_ptes_none); + p->max_ptes_swap = READ_ONCE(khugepaged_max_ptes_swap); + p->max_ptes_shared = READ_ONCE(khugepaged_max_ptes_shared); + p->strict_sub_pmd = true; + p->skip_lazyfree = true; + p->require_referenced = true; + p->install_pmd = false; + p->writeback_dirty = false; + p->gfp = alloc_hugepage_khugepaged_gfpmask(); + p->tva_type = TVA_KHUGEPAGED; +} + +/* MADV_COLLAPSE was asked for explicitly, so it is not held to those */ +static void collapse_policy_forced(struct collapse_policy *p) +{ + p->max_ptes_none = HPAGE_PMD_NR; + p->max_ptes_swap = HPAGE_PMD_NR; + p->max_ptes_shared = HPAGE_PMD_NR; + p->strict_sub_pmd = false; + p->skip_lazyfree = false; + p->require_referenced = false; + p->install_pmd = true; + p->writeback_dirty = true; + p->gfp = GFP_TRANSHUGE; + p->tva_type = TVA_FORCED_COLLAPSE; +} + #ifdef CONFIG_NUMA static int collapse_find_target_node(struct collapse_control *cc) { @@ -1013,8 +1027,7 @@ static enum scan_result hugepage_vma_revalidate(struct mm_struct *mm, unsigned l struct collapse_control *cc, unsigned int order) { struct vm_area_struct *vma; - enum tva_type type = cc->is_khugepaged ? TVA_KHUGEPAGED : - TVA_FORCED_COLLAPSE; + enum tva_type type = cc->policy.tva_type; if (unlikely(collapse_test_exit_or_disable(mm))) return SCAN_ANY_PROCESS; @@ -1197,8 +1210,7 @@ static enum scan_result __collapse_huge_page_swapin(struct mm_struct *mm, static enum scan_result alloc_charge_folio(struct folio **foliop, struct mm_struct *mm, struct collapse_control *cc, unsigned int order) { - gfp_t gfp = (cc->is_khugepaged ? alloc_hugepage_khugepaged_gfpmask() : - GFP_TRANSHUGE); + gfp_t gfp = cc->policy.gfp; int node = collapse_find_target_node(cc); struct folio *folio; @@ -1551,7 +1563,7 @@ static enum scan_result collapse_scan_pmd(struct mm_struct *mm, const unsigned int max_ptes_shared = collapse_max_ptes_shared(cc, HPAGE_PMD_ORDER); const unsigned int max_ptes_swap = collapse_max_ptes_swap(cc, HPAGE_PMD_ORDER); unsigned int max_ptes_none = collapse_max_ptes_none(cc, vma, HPAGE_PMD_ORDER); - enum tva_type tva_flags = cc->is_khugepaged ? TVA_KHUGEPAGED : TVA_FORCED_COLLAPSE; + enum tva_type tva_flags = cc->policy.tva_type; pmd_t *pmd; pte_t *pte, *_pte, pteval; int i; @@ -1651,7 +1663,7 @@ static enum scan_result collapse_scan_pmd(struct mm_struct *mm, * If the vma has the VM_DROPPABLE flag, the collapse will * preserve the lazyfree property without needing to skip. */ - if (cc->is_khugepaged && !(vma->vm_flags & VM_DROPPABLE) && + if (cc->policy.skip_lazyfree && !(vma->vm_flags & VM_DROPPABLE) && folio_test_lazyfree(folio) && !pte_dirty(pteval)) { result = SCAN_PAGE_LAZYFREE; failed_pfn = folio_pfn(folio); @@ -1717,13 +1729,13 @@ static enum scan_result collapse_scan_pmd(struct mm_struct *mm, goto out_unmap; } - if (cc->is_khugepaged && + if (cc->policy.require_referenced && folio_pte_referenced(folio, vma, addr, pteval)) referenced++; } - if (cc->is_khugepaged && - (!referenced || - (unmapped && referenced < HPAGE_PMD_NR / 2))) { + if (cc->policy.require_referenced && + (!referenced || + (unmapped && referenced < HPAGE_PMD_NR / 2))) { result = SCAN_LACK_REFERENCED_PAGE; } else { result = SCAN_SUCCEED; @@ -2582,11 +2594,11 @@ static enum scan_result collapse_file(struct mm_struct *mm, unsigned long addr, xas_unlock_irq(&xas); /* - * Remove pte page tables, so we can re-fault the page as huge. - * If MADV_COLLAPSE, adjust result to call try_collapse_pte_mapped_thp(). + * Remove pte page tables, so we can re-fault the page as huge. A + * caller that wants the PMD mapped now is told to go and do that. */ retract_page_tables(mapping, start); - if (cc && !cc->is_khugepaged) + if (cc->policy.install_pmd) result = SCAN_PTE_MAPPED_HUGEPAGE; folio_unlock(new_folio); @@ -2773,11 +2785,8 @@ static enum scan_result collapse_single_pmd(unsigned long addr, retry: result = collapse_scan_file(mm, addr, file, pgoff, cc); - /* - * For MADV_COLLAPSE, when encountering dirty pages, try to writeback, - * then retry the collapse one time. - */ - if (!cc->is_khugepaged && result == SCAN_PAGE_DIRTY_OR_WRITEBACK && + /* Dirty pages are worth a writeback and one more try, if asked for */ + if (cc->policy.writeback_dirty && result == SCAN_PAGE_DIRTY_OR_WRITEBACK && !triggered_wb && mapping_can_writeback(file->f_mapping)) { const loff_t lstart = (loff_t)pgoff << PAGE_SHIFT; const loff_t lend = lstart + HPAGE_PMD_SIZE - 1; @@ -2794,7 +2803,7 @@ static enum scan_result collapse_single_pmd(unsigned long addr, result = SCAN_ANY_PROCESS; else result = try_collapse_pte_mapped_thp(mm, addr, - !cc->is_khugepaged); + cc->policy.install_pmd); if (result == SCAN_PMD_MAPPED) result = SCAN_SUCCEED; mmap_read_unlock(mm); @@ -2943,6 +2952,9 @@ static void khugepaged_do_scan(struct collapse_control *cc) lru_add_drain_all(); + /* One policy for the whole pass, so every table is treated the same */ + collapse_policy_khugepaged(&cc->policy); + cc->progress = 0; while (true) { cond_resched(); @@ -3170,7 +3182,7 @@ int madvise_collapse(struct vm_area_struct *vma, unsigned long start, cc = kmalloc_obj(*cc); if (!cc) return -ENOMEM; - cc->is_khugepaged = false; + collapse_policy_forced(&cc->policy); cc->progress = 0; lru_add_drain_all(); -- 2.54.0