From: Kiryl Shutsemau <kirill@shutemov.name>
To: Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>, Zi Yan <ziy@nvidia.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: "Kiryl Shutsemau (Meta)" <kas@kernel.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
kernel-team@meta.com, "Liam R. Howlett" <liam@infradead.org>,
Nico Pache <nico.pache@linux.dev>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
Usama Arif <usama.arif@linux.dev>,
Vlastimil Babka <vbabka@kernel.org>, Jann Horn <jannh@google.com>
Subject: [PATCH v3 09/12] mm/collapse: open-code collapse_single_pmd() in its two callers
Date: Wed, 16 Sep 2026 10:31:36 +0100 [thread overview]
Message-ID: <20260916093145.4022188-10-kirill@shutemov.name> (raw)
In-Reply-To: <20260916093145.4022188-1-kirill@shutemov.name>
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
collapse_scan_pmd() and collapse_run_pmd() each have a clear locking
contract. The scan is called with mmap_lock held for reading and returns
with it still held. The collapse is called without it.
collapse_single_pmd() kept that boundary inside itself. It dropped the
lock on some paths and not others, and reported which by way of a bool its
callers had to carry along and then act on.
Open-code it in the two callers. Each scans under the lock it already
holds and, when the scan found work, gives the lock up before running the
collapse.
khugepaged's lock_dropped and madvise_collapse()'s mmap_unlocked both go:
the code dropping the lock is now the code that wanted to know.
khugepaged's walk carries on to the next table while the scan keeps
refusing, and ends once a collapse has taken the lock from under it.
madvise_collapse() re-finds its VMA after a collapse, which it did before,
and now uses a NULL vma to say that it has to. It still reports the drop
to its own caller, from the line that does it.
The lock is given up and taken again at the same points as before. No
functional change.
Assisted-by: LLM
Reviewed-by: Zi Yan <ziy@nvidia.com>
Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
---
mm/khugepaged.c | 103 +++++++++++++++++++++++-------------------------
1 file changed, 50 insertions(+), 53 deletions(-)
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index b6fc2c78e3e2..9e6b2af6519e 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -2846,28 +2846,6 @@ static enum scan_result collapse_run_pmd(struct mm_struct *mm,
return result;
}
-/*
- * Try to collapse a single PMD starting at a PMD aligned addr, and return
- * the results.
- */
-static enum scan_result collapse_single_pmd(unsigned long addr,
- struct vm_area_struct *vma, bool *lock_dropped,
- struct collapse_control *cc)
-{
- struct mm_struct *mm = vma->vm_mm;
- enum scan_result result;
-
- result = collapse_scan_pmd(vma, addr, cc);
- if (result != SCAN_SUCCEED && result != SCAN_PTE_MAPPED_HUGEPAGE)
- return result;
-
- /* The collapse takes its own locks, so give this up */
- mmap_read_unlock(mm);
- *lock_dropped = true;
-
- return collapse_run_pmd(mm, addr, result, cc);
-}
-
static void collapse_scan_mm_slot(unsigned int progress_max,
enum scan_result *result, struct collapse_control *cc)
__releases(&khugepaged_mm_lock)
@@ -2930,7 +2908,7 @@ static void collapse_scan_mm_slot(unsigned int progress_max,
VM_BUG_ON(khugepaged_scan.address & ~HPAGE_PMD_MASK);
while (khugepaged_scan.address < hend) {
- bool lock_dropped = false;
+ unsigned long addr;
cond_resched();
if (unlikely(collapse_test_exit_or_disable(mm)))
@@ -2940,23 +2918,30 @@ static void collapse_scan_mm_slot(unsigned int progress_max,
khugepaged_scan.address + HPAGE_PMD_SIZE >
hend);
- *result = collapse_single_pmd(khugepaged_scan.address,
- vma, &lock_dropped, cc);
- if (*result == SCAN_SUCCEED)
- khugepaged_pages_collapsed++;
+ addr = khugepaged_scan.address;
/* move to next address */
khugepaged_scan.address += HPAGE_PMD_SIZE;
- if (lock_dropped)
- /*
- * We released mmap_lock so break loop. Note
- * that we drop mmap_lock before all hugepage
- * allocations, so if allocation fails, we are
- * guaranteed to break here and report the
- * correct result back to caller.
- */
- goto breakouterloop_mmap_lock;
- if (cc->progress >= progress_max)
- goto breakouterloop;
+
+ *result = collapse_scan_pmd(vma, addr, cc);
+ /* Nothing to do here, and the lock is still ours */
+ if (*result != SCAN_SUCCEED &&
+ *result != SCAN_PTE_MAPPED_HUGEPAGE) {
+ if (cc->progress >= progress_max)
+ goto breakouterloop;
+ continue;
+ }
+
+ /*
+ * A collapse takes its own locks and is slow enough
+ * that a writer should not wait behind it, so give the
+ * lock up. That ends this walk: vma and the mm are
+ * whatever the collapse leaves them.
+ */
+ mmap_read_unlock(mm);
+ *result = collapse_run_pmd(mm, addr, *result, cc);
+ if (*result == SCAN_SUCCEED)
+ khugepaged_pages_collapsed++;
+ goto breakouterloop_mmap_lock;
}
}
breakouterloop:
@@ -3225,7 +3210,6 @@ int madvise_collapse(struct vm_area_struct *vma, unsigned long start,
unsigned long hstart, hend, addr;
enum scan_result last_fail = SCAN_FAIL;
int thps = 0;
- bool mmap_unlocked = false;
BUG_ON(vma->vm_start > start);
BUG_ON(vma->vm_end < end);
@@ -3248,25 +3232,40 @@ int madvise_collapse(struct vm_area_struct *vma, unsigned long start,
lru_add_drain_all();
for (addr = hstart; addr < hend; addr += HPAGE_PMD_SIZE) {
- enum scan_result result = SCAN_FAIL;
+ struct vm_area_struct *found;
+ enum scan_result result;
- if (mmap_unlocked) {
+ /*
+ * A collapse gives the lock up, so the VMA has to be found
+ * again after one: it can shrink while nothing is held. A scan
+ * that finds nothing to collapse leaves the lock alone, so a
+ * range that is already collapsed walks on without relocking.
+ */
+ if (!vma) {
cond_resched();
mmap_read_lock(mm);
- mmap_unlocked = false;
- *lock_dropped = true;
- result = hugepage_vma_revalidate(mm, addr, false, &vma,
+ result = hugepage_vma_revalidate(mm, addr, false, &found,
cc, HPAGE_PMD_ORDER);
if (result != SCAN_SUCCEED) {
last_fail = result;
- goto out_nolock;
+ goto out_locked;
}
-
+ vma = found;
hend = min(hend, vma->vm_end & HPAGE_PMD_MASK);
}
- result = collapse_single_pmd(addr, vma, &mmap_unlocked, cc);
+ result = collapse_scan_pmd(vma, addr, cc);
+ /* Nothing to do here, and the lock is still ours */
+ if (result != SCAN_SUCCEED && result != SCAN_PTE_MAPPED_HUGEPAGE)
+ goto tally;
+ /* The collapse takes its own locks, so give this up */
+ mmap_read_unlock(mm);
+ *lock_dropped = true;
+ vma = NULL;
+
+ result = collapse_run_pmd(mm, addr, result, cc);
+tally:
switch (result) {
case SCAN_SUCCEED:
case SCAN_PMD_MAPPED:
@@ -3288,17 +3287,15 @@ int madvise_collapse(struct vm_area_struct *vma, unsigned long start,
default:
last_fail = result;
/* Other error, exit */
- goto out_maybelock;
+ goto out;
}
}
-out_maybelock:
+out:
/* Caller expects us to hold mmap_lock on return */
- if (mmap_unlocked) {
- *lock_dropped = true;
+ if (!vma)
mmap_read_lock(mm);
- }
-out_nolock:
+out_locked:
mmap_assert_locked(mm);
collapse_control_release(cc);
kfree(cc);
--
2.54.0
next prev parent reply other threads:[~2026-09-16 9:32 UTC|newest]
Thread overview: 36+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 9:31 [PATCH v3 00/12] mm/collapse: separate a collapse from its callers Kiryl Shutsemau
2026-09-16 9:31 ` [PATCH v3 01/12] mm/khugepaged: drop redundant mm_struct pin in madvise_collapse() Kiryl Shutsemau
2026-09-23 11:49 ` David Hildenbrand (Arm)
2026-09-16 9:31 ` [PATCH v3 02/12] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-09-23 11:51 ` David Hildenbrand (Arm)
2026-09-16 9:31 ` [PATCH v3 03/12] mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes Kiryl Shutsemau
2026-09-23 11:52 ` David Hildenbrand (Arm)
2026-09-16 9:31 ` [PATCH v3 04/12] mm/collapse: add collapse.h for the collapse interface Kiryl Shutsemau
2026-09-23 11:56 ` David Hildenbrand (Arm)
2026-09-16 9:31 ` [PATCH v3 05/12] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-09-23 12:24 ` David Hildenbrand (Arm)
2026-09-24 13:49 ` Kiryl Shutsemau
2026-09-16 9:31 ` [PATCH v3 06/12] mm/collapse: drop the collapse_possible() wrapper Kiryl Shutsemau
2026-09-23 12:25 ` David Hildenbrand (Arm)
2026-09-16 9:31 ` [PATCH v3 07/12] mm/collapse: name the per-table scan reset for what it resets Kiryl Shutsemau
2026-09-23 12:26 ` David Hildenbrand (Arm)
2026-09-16 9:31 ` [PATCH v3 08/12] mm/collapse: separate scanning a PTE table from collapsing it Kiryl Shutsemau
2026-09-18 9:06 ` Baolin Wang
2026-09-23 13:04 ` David Hildenbrand (Arm)
2026-09-24 14:19 ` Kiryl Shutsemau
2026-09-16 9:31 ` Kiryl Shutsemau [this message]
2026-09-18 9:31 ` [PATCH v3 09/12] mm/collapse: open-code collapse_single_pmd() in its two callers Baolin Wang
2026-09-23 13:08 ` David Hildenbrand (Arm)
2026-09-24 14:56 ` Kiryl Shutsemau
2026-09-16 9:31 ` [PATCH v3 10/12] mm/collapse: work out the orders a VMA allows once per VMA Kiryl Shutsemau
2026-09-23 13:19 ` David Hildenbrand (Arm)
2026-09-24 15:04 ` Kiryl Shutsemau
2026-09-16 9:31 ` [PATCH v3 11/12] mm/collapse: declare the collapse interface in collapse.h Kiryl Shutsemau
2026-09-23 13:34 ` David Hildenbrand (Arm)
2026-09-24 15:22 ` Kiryl Shutsemau
2026-09-28 12:03 ` David Hildenbrand (Arm)
2026-09-16 9:31 ` [PATCH v3 12/12] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-09-16 22:48 ` [PATCH v3 00/12] mm/collapse: separate a collapse from its callers Andrew Morton
2026-09-17 12:26 ` Kiryl Shutsemau
2026-09-17 22:11 ` Andrew Morton
2026-09-23 13:52 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260916093145.4022188-10-kirill@shutemov.name \
--to=kirill@shutemov.name \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=jannh@google.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®