From: Zongkun Lei <leizongkun@qq.com>
To: linux-mm@kvack.org
Cc: Zongkun Lei <leizongkun@qq.com>,
Miaohe Lin <linmiaohe@huawei.com>,
Naoya Horiguchi <nao.horiguchi@gmail.com>,
Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>, Rik van Riel <riel@surriel.com>,
"Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>, Harry Yoo <harry@kernel.org>,
Jann Horn <jannh@google.com>, Lance Yang <lance.yang@linux.dev>,
linux-kernel@vger.kernel.org
Subject: [RFC PATCH 4/6] mm/memory-failure: handle swapcached hugetlb folios
Date: Tue, 29 Sep 2026 16:00:57 +0800 [thread overview]
Message-ID: <tencent_6DE2E1B379FAD189CEFE60DBA83A4C1DF005@qq.com> (raw)
In-Reply-To: <cover.1790663399.git.leizongkun@qq.com>
Anonymous hugetlb folios can now sit in the swap cache (swap-out
writeback window, cached copy after swap-in), a state memory failure
never had to handle before. Two gaps:
1. try_to_unmap() routes every hugetlb folio to
try_to_unmap_poisoned_hugetlb_one(), which requires TTU_HWPOISON.
But unmap_poisoned_folio() clears TTU_HWPOISON for a dirty
swapcache folio -- its mappings must be replaced with swap entries,
exactly like for a 4K swapcache folio, so the swap accounting stays
intact and the fault path gets to kill the owner. Routing such a
folio to the poisoned handler fires its VM_WARN_ON_ONCE and would
install hwpoison entries on top of live swap state. Route hugetlb
folios by flag instead: TTU_HWPOISON set -> the poisoned handler;
cleared -> try_to_unmap_swap_hugetlb_one().
2. me_huge_page() has no swapcache case: a clean poisoned hugetlb
folio in the swap cache falls into the truncate path, which evicts
it (the swap address space has no error_remove_folio). The folio
is never returned to the hstate pool (a silent 2M pool leak) and
the poison marker is lost, so a later fault would swap possibly
corrupt data back in. Mirror me_swapcache_dirty(): keep the
poisoned folio in the swap cache and return MF_DELAYED, so the
swap-in path sees folio_test_hwpoison() and kills the accessor.
Signed-off-by: Zongkun Lei <leizongkun@qq.com>
---
mm/memory-failure.c | 19 +++++++++++++++++++
mm/rmap.c | 32 ++++++++++++++++++++++++++------
2 files changed, 45 insertions(+), 6 deletions(-)
diff --git a/mm/memory-failure.c b/mm/memory-failure.c
index a8b03e2920ba..0c31c8ab5542 100644
--- a/mm/memory-failure.c
+++ b/mm/memory-failure.c
@@ -1151,6 +1151,25 @@ static int me_huge_page(struct page_state *ps, struct page *p)
struct address_space *mapping;
bool extra_pins = false;
+ /*
+ * A hugetlb folio reaches the swap cache only via hugetlb
+ * swap-out. Keep a poisoned one there so that the swap-in
+ * fault path intercepts folio_test_hwpoison() and kills the
+ * accessor. The truncate path below would evict a clean one
+ * (the swap address space has no error_remove_folio), losing
+ * both the folio (it is never returned to the hstate pool) and
+ * the poison marker, so a later fault would silently swap
+ * possibly-corrupt data back in. Mirror me_swapcache_dirty().
+ */
+ if (folio_test_swapcache(folio)) {
+ folio_clear_dirty(folio);
+ folio_unlock(folio);
+ /* The swap cache pin is intentionally retained. */
+ if (has_extra_refcount(ps, p, true))
+ return MF_FAILED;
+ return MF_DELAYED;
+ }
+
mapping = folio_mapping(folio);
if (mapping) {
res = truncate_error_folio(folio, page_to_pfn(p), mapping);
diff --git a/mm/rmap.c b/mm/rmap.c
index 2ce54cb5ff20..a88c911e8463 100644
--- a/mm/rmap.c
+++ b/mm/rmap.c
@@ -1991,8 +1991,9 @@ static bool try_to_unmap_poisoned_hugetlb_one(struct folio *folio,
pte_t pteval;
/*
- * The try_to_unmap() is only passed a hugetlb folio in the case
- * where the hugetlb folio is poisoned.
+ * try_to_unmap() routes a hugetlb folio here only for memory
+ * failure with TTU_HWPOISON set; a poisoned folio kept in the
+ * swap cache is routed to try_to_unmap_swap_hugetlb_one() instead.
*/
VM_WARN_ON_ONCE_FOLIO(!folio_test_hwpoison(folio), folio);
VM_WARN_ON_ONCE(!(flags & TTU_HWPOISON));
@@ -2199,8 +2200,10 @@ static bool ttu_anon_folio(struct vm_area_struct *vma, struct folio *folio,
* with a swap entry covering the whole folio; for a file-backed folio the
* PTE is simply cleared, the swap anchor lives in the hugetlbfs page
* cache instead (shmem-style). Called from the hugetlb reclaim path
- * (hugetlb_reclaim_pages()) with the folio lock held; an anonymous folio
- * is swapbacked with its swap slots allocated.
+ * (hugetlb_reclaim_pages()) and, via try_to_unmap(), from memory failure
+ * for a poisoned folio that is kept in the swap cache (see
+ * unmap_poisoned_folio()); the folio lock is held in all cases, and an
+ * anonymous folio is swapbacked with its swap slots allocated.
*
* Keep in sync with ttu_anon_swapbacked_folio().
*/
@@ -2567,13 +2570,30 @@ static int folio_not_mapped(struct folio *folio)
void try_to_unmap(struct folio *folio, enum ttu_flags flags)
{
struct rmap_walk_control rwc = {
- .rmap_one = folio_test_hugetlb(folio) ?
- try_to_unmap_poisoned_hugetlb_one : try_to_unmap_one,
+ .rmap_one = try_to_unmap_one,
.arg = (void *)flags,
.done = folio_not_mapped,
.anon_lock = folio_lock_anon_vma_read,
};
+ /*
+ * try_to_unmap() is passed a hugetlb folio only by memory failure.
+ * With TTU_HWPOISON the mappings are replaced with hwpoison
+ * entries. Without it the folio is a poisoned folio that is kept
+ * in the swap cache (unmap_poisoned_folio() cleared TTU_HWPOISON),
+ * so the mappings are replaced with swap entries, exactly like for
+ * a 4K swapcache folio: the swap accounting stays intact and the
+ * fault path gets to kill the owner.
+ */
+ if (folio_test_hugetlb(folio)) {
+ if (flags & TTU_HWPOISON)
+ rwc.rmap_one = try_to_unmap_poisoned_hugetlb_one;
+ else if (WARN_ON_ONCE(!folio_test_swapcache(folio)))
+ return;
+ else
+ rwc.rmap_one = try_to_unmap_swap_hugetlb_one;
+ }
+
if (flags & TTU_RMAP_LOCKED)
rmap_walk_locked(folio, &rwc);
else
--
2.53.0
next prev parent reply other threads:[~2026-09-29 8:01 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <cover.1790663399.git.leizongkun@qq.com>
2026-09-29 7:58 ` [RFC PATCH 1/6] mm/swap: introduce folio_swap_entry() and convert all folio->swap readers Zongkun Lei
2026-09-29 7:59 ` [RFC PATCH 2/6] mm/hugetlb: swap-in support for anonymous hugetlb folios Zongkun Lei
2026-09-29 7:59 ` [RFC PATCH 3/6] mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT Zongkun Lei
2026-09-29 8:00 ` Zongkun Lei [this message]
2026-09-29 8:01 ` [RFC PATCH 5/6] mm/hugetlb: swap support for file-backed hugetlb folios Zongkun Lei
2026-09-29 8:02 ` [RFC PATCH 6/6] selftests/mm: add hugetlb_swap test, document hugetlb swap Zongkun Lei
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=tencent_6DE2E1B379FAD189CEFE60DBA83A4C1DF005@qq.com \
--to=leizongkun@qq.com \
--cc=akpm@linux-foundation.org \
--cc=david@kernel.org \
--cc=harry@kernel.org \
--cc=jannh@google.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linmiaohe@huawei.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nao.horiguchi@gmail.com \
--cc=riel@surriel.com \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®