mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH 1/6] mm/swap: introduce folio_swap_entry() and convert all folio->swap readers
       [not found] <cover.1790663399.git.leizongkun@qq.com>
@ 2026-09-29  7:58 ` Zongkun Lei
  2026-09-29  7:59 ` [RFC PATCH 2/6] mm/hugetlb: swap-in support for anonymous hugetlb folios Zongkun Lei
                   ` (4 subsequent siblings)
  5 siblings, 0 replies; 6+ messages in thread
From: Zongkun Lei @ 2026-09-29  7:58 UTC (permalink / raw)
  To: linux-mm
  Cc: Zongkun Lei, Andrew Morton, Chris Li, Kairui Song, Kemeng Shi,
	Nhat Pham, Baoquan He, Barry Song, Youngjun Park,
	David Hildenbrand, Zi Yan, Baolin Wang, Liam R. Howlett,
	Nico Pache, Ryan Roberts, Dev Jain, Lance Yang, Usama Arif,
	Kiryl Shutsemau, Johannes Weiner, Michal Hocko, Roman Gushchin,
	Shakeel Butt, Muchun Song, Lorenzo Stoakes, Vlastimil Babka,
	Mike Rapoport, Suren Baghdasaryan, Hugh Dickins, Peter Xu,
	Qi Zheng, Axel Rasmussen, Yuanchu Xie, Wei Xu, Yosry Ahmed,
	Chengming Zhou, linux-kernel, cgroups

Hugetlb folios keep their per-folio flags in folio->private, which
shares a union with folio->swap.  While such a folio sits in the swap
cache, the low bits of folio->swap.val can therefore carry flag bits
rather than zeroes - with HVO enabled, HPG_vmemmap_optimized (bit 4)
is set on virtually every hugetlb folio.  Any naked read of
folio->swap on a swap-cached hugetlb folio would hand a corrupted
swap entry to the swap table, zswap, the memcg accounting or the I/O
paths.

Swap entries are always allocated aligned to the folio size, so
rounding the raw value down to folio_nr_pages() recovers the real
entry.  Wrap that in folio_swap_entry() and convert every reader of
folio->swap in the tree to use it.  For all non-hugetlb folios the
low bits are always zero (order-0 folios trivially, large folios
because the alignment leaves them clear), so the helper is an
identity transform and this patch changes no behaviour on any
existing path; reviewing it only requires auditing the helper itself
plus the single invariant "swap entries are folio-size aligned".

The write side keeps raw access deliberately:
__swap_cache_do_add_folio() now preserves the low flag bits when
installing the entry, and __swap_cache_do_del_folio() clears the
entry while keeping the flags, rounding the passed entry down so an
unaligned value can never leak into the swap table.

This is preparation for swap support of hugetlb folios; hugetlb
folios too small for all HPAGEFLAG bits to fit below the swap offset
will be kept out of the swap cache by the feature's size gate.

Signed-off-by: Zongkun Lei <leizongkun@qq.com>
---
 include/linux/swap.h | 22 +++++++++++++++++++++-
 mm/huge_memory.c     |  2 +-
 mm/memcontrol-v1.c   |  4 ++--
 mm/memcontrol.c      |  2 +-
 mm/memory.c          |  2 +-
 mm/page_io.c         | 18 +++++++++---------
 mm/shmem.c           |  6 +++---
 mm/swap.h            |  8 +++++---
 mm/swap_state.c      | 15 ++++++++++-----
 mm/swapfile.c        | 14 ++++++++------
 mm/userfaultfd.c     |  2 +-
 mm/util.c            |  2 +-
 mm/vmscan.c          |  4 ++--
 mm/zswap.c           |  4 ++--
 14 files changed, 67 insertions(+), 38 deletions(-)

diff --git a/include/linux/swap.h b/include/linux/swap.h
index 5658a1634b85..677081238811 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -278,10 +278,30 @@ struct swap_info_struct {
 	const struct swap_ops *ops;
 };
 
+/**
+ * folio_swap_entry - Get the swap entry recorded in a swap cache folio.
+ * @folio: The folio, must be in the swap cache.
+ *
+ * Hugetlb folios keep their per-folio flags in folio->private, which aliases
+ * folio->swap, so the low bits of folio->swap.val may carry flag bits while
+ * the folio is swap cached.  Swap entries are always aligned to the folio
+ * size, so rounding down recovers the entry.  Hugetlb folios too small for
+ * all flag bits to fit below the swap offset never enter the swap cache;
+ * see hugetlb_folio_swap_supported().
+ *
+ * For all other folios the low bits are always zero, so this is an identity
+ * transform; always use it instead of reading folio->swap directly.
+ */
+static inline swp_entry_t folio_swap_entry(const struct folio *folio)
+{
+	return (swp_entry_t) { .val = round_down(folio->swap.val,
+						 folio_nr_pages(folio)) };
+}
+
 static inline swp_entry_t page_swap_entry(struct page *page)
 {
 	struct folio *folio = page_folio(page);
-	swp_entry_t entry = folio->swap;
+	swp_entry_t entry = folio_swap_entry(folio);
 
 	entry.val += folio_page_idx(folio, page);
 	return entry;
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 54494c3fa983..18824e298cff 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -3760,7 +3760,7 @@ static void __split_folio_to_order(struct folio *folio, int old_order,
 		VM_WARN_ON_ONCE_PAGE(new_folio->private, new_head);
 
 		if (folio_test_swapcache(folio))
-			new_folio->swap.val = folio->swap.val + i;
+			new_folio->swap.val = folio_swap_entry(folio).val + i;
 
 		/* Page flags must be visible before we make the page non-compound. */
 		smp_wmb();
diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c
index bf2c7d53b01b..2186523f3f0e 100644
--- a/mm/memcontrol-v1.c
+++ b/mm/memcontrol-v1.c
@@ -299,7 +299,7 @@ void __memcg1_swapout(struct folio *folio, struct swap_cluster_info *ci)
 	swap_memcg = mem_cgroup_private_id_get_online(memcg, nr_entries);
 	mod_memcg_state(swap_memcg, MEMCG_SWAP, nr_entries);
 
-	__swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_entries,
+	__swap_cgroup_set(ci, swp_cluster_offset(folio_swap_entry(folio)), nr_entries,
 			  mem_cgroup_private_id(swap_memcg));
 
 	folio_unqueue_deferred_split(folio);
@@ -369,7 +369,7 @@ void memcg1_swapin(struct folio *folio)
 	 */
 	nr_pages = folio_nr_pages(folio);
 	ci = swap_cluster_get_and_lock(folio);
-	id = __swap_cgroup_clear(ci, swp_cluster_offset(folio->swap),
+	id = __swap_cgroup_clear(ci, swp_cluster_offset(folio_swap_entry(folio)),
 				 nr_pages);
 	swap_cluster_unlock(ci);
 	mem_cgroup_uncharge_swap(id, nr_pages);
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 256b68ffca70..0ae9e3176479 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -5781,7 +5781,7 @@ int __mem_cgroup_try_charge_swap(struct folio *folio)
 	mod_memcg_state(memcg, MEMCG_SWAP, nr_pages);
 
 	ci = swap_cluster_get_and_lock(folio);
-	__swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_pages,
+	__swap_cgroup_set(ci, swp_cluster_offset(folio_swap_entry(folio)), nr_pages,
 			  mem_cgroup_private_id(memcg));
 	swap_cluster_unlock(ci);
 
diff --git a/mm/memory.c b/mm/memory.c
index bc14cae3c49d..4dcc8d73e3ee 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -5056,7 +5056,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
 		address = folio_start;
 		ptep = folio_ptep;
 		nr_pages = nr;
-		entry = folio->swap;
+		entry = folio_swap_entry(folio);
 		page = &folio->page;
 	}
 
diff --git a/mm/page_io.c b/mm/page_io.c
index 1da4ff484f09..e1ec577d5759 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -160,7 +160,7 @@ static void swap_zeromap_folio_set(struct folio *folio)
 	struct obj_cgroup *objcg = get_obj_cgroup_from_folio(folio);
 	int nr_pages = folio_nr_pages(folio);
 	struct swap_cluster_info *ci;
-	swp_entry_t entry = folio->swap;
+	swp_entry_t entry = folio_swap_entry(folio);
 	unsigned int i;
 
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio);
@@ -183,7 +183,7 @@ static void swap_zeromap_folio_set(struct folio *folio)
 static void swap_zeromap_folio_clear(struct folio *folio)
 {
 	struct swap_cluster_info *ci;
-	swp_entry_t entry = folio->swap;
+	swp_entry_t entry = folio_swap_entry(folio);
 	unsigned int i;
 
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio);
@@ -321,7 +321,7 @@ int sio_pool_init(void)
 static bool swap_can_merge(struct swap_io_ctx *ctx, struct folio *folio,
 		int rw)
 {
-	struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
+	struct swap_info_struct *sis = __swap_entry_to_info(folio_swap_entry(folio));
 	struct bio_vec *last_bv = &ctx->sio->bvecs[ctx->sio->nr_bvecs - 1];
 	struct folio *prev_folio = bvec_folio(last_bv);
 	size_t prev_folio_size = folio_size(prev_folio);
@@ -333,7 +333,7 @@ static bool swap_can_merge(struct swap_io_ctx *ctx, struct folio *folio,
 
 static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)
 {
-	struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
+	struct swap_info_struct *sis = __swap_entry_to_info(folio_swap_entry(folio));
 	struct swap_iocb *sio = ctx->sio;
 
 	if (sio && !swap_can_merge(ctx, folio, rw)) {
@@ -432,7 +432,7 @@ static bool swap_read_folio_zeromap(struct folio *folio)
 	 * that an IO error is emitted (e.g. do_swap_page() will sigbus).
 	 * Folio lock stabilizes the cluster and map, so the check is safe.
 	 */
-	if (WARN_ON_ONCE(swap_zeromap_batch(folio->swap, nr_pages,
+	if (WARN_ON_ONCE(swap_zeromap_batch(folio_swap_entry(folio), nr_pages,
 			 &is_zeromap) != nr_pages))
 		return true;
 
@@ -453,7 +453,7 @@ static bool swap_read_folio_zeromap(struct folio *folio)
 
 void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
 {
-	struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
+	struct swap_info_struct *sis = __swap_entry_to_info(folio_swap_entry(folio));
 	bool synchronous = sis->flags & SWP_SYNCHRONOUS_IO;
 	bool workingset = folio_test_workingset(folio);
 	unsigned long pflags;
@@ -525,7 +525,7 @@ static void swap_fs_write_complete(struct kiocb *iocb, long ret)
 		 * folio_rotate_reclaimable but rate-limit the messages.
 		 */
 		pr_err_ratelimited("Write error %ld on dio swapfile (%llu)\n",
-				   ret, swap_dev_pos(folio->swap));
+				   ret, swap_dev_pos(folio_swap_entry(folio)));
 	}
 
 	swap_write_end(sio, failed);
@@ -671,8 +671,8 @@ EXPORT_SYMBOL_GPL(swap_fs_prepare_rw);
 bool swap_fs_can_merge(struct folio *folio, struct folio *prev_folio,
 		size_t prev_folio_size, int rw)
 {
-	return swap_dev_pos(folio->swap) ==
-		swap_dev_pos(prev_folio->swap) + prev_folio_size;
+	return swap_dev_pos(folio_swap_entry(folio)) ==
+		swap_dev_pos(folio_swap_entry(prev_folio)) + prev_folio_size;
 }
 EXPORT_SYMBOL_GPL(swap_fs_can_merge);
 
diff --git a/mm/shmem.c b/mm/shmem.c
index eeb9a78c125a..284d6d5d8b94 100644
--- a/mm/shmem.c
+++ b/mm/shmem.c
@@ -1914,7 +1914,7 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
 		}
 
 		folio_dup_swap(folio, NULL);
-		shmem_delete_from_page_cache(folio, swp_to_radix_entry(folio->swap));
+		shmem_delete_from_page_cache(folio, swp_to_radix_entry(folio_swap_entry(folio)));
 
 		BUG_ON(folio_mapped(folio));
 		error = swap_writeout(ctx, folio);
@@ -1929,7 +1929,7 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
 		 * it will be appropriate if other reactivate cases are added.
 		 */
 		error = shmem_add_to_page_cache(folio, mapping, index,
-				swp_to_radix_entry(folio->swap),
+				swp_to_radix_entry(folio_swap_entry(folio)),
 				__GFP_HIGH | __GFP_NOMEMALLOC | __GFP_NOWARN);
 		/* Swap entry might be erased by racing shmem_free_swap() */
 		if (!error) {
@@ -2293,7 +2293,7 @@ static int shmem_replace_folio(struct folio **foliop, gfp_t gfp,
 {
 	struct swap_cluster_info *ci;
 	struct folio *new, *old = *foliop;
-	swp_entry_t entry = old->swap;
+	swp_entry_t entry = folio_swap_entry(old);
 	int nr_pages = folio_nr_pages(old);
 	int error = 0;
 
diff --git a/mm/swap.h b/mm/swap.h
index 0b5d507739bc..19c260353e08 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -173,10 +173,12 @@ static inline struct swap_cluster_info *swap_cluster_lock(
 static inline struct swap_cluster_info *__swap_cluster_get_and_lock(
 		const struct folio *folio, bool irq)
 {
+	swp_entry_t entry = folio_swap_entry(folio);
+
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio);
-	return __swap_cluster_lock(__swap_entry_to_info(folio->swap),
-				   swp_offset(folio->swap), irq);
+	return __swap_cluster_lock(__swap_entry_to_info(entry),
+				   swp_offset(entry), irq);
 }
 
 /*
@@ -287,7 +289,7 @@ static inline loff_t swap_dev_pos(swp_entry_t entry)
 static inline bool folio_matches_swap_entry(const struct folio *folio,
 					    swp_entry_t entry)
 {
-	swp_entry_t folio_entry = folio->swap;
+	swp_entry_t folio_entry = folio_swap_entry(folio);
 	long nr_pages = folio_nr_pages(folio);
 
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 305877e1f4d7..516052efbf62 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -227,7 +227,11 @@ static void __swap_cache_do_add_folio(struct swap_cluster_info *ci,
 
 	folio_ref_add(folio, nr_pages);
 	folio_set_swapcache(folio);
-	folio->swap = entry;
+	/*
+	 * Keep the low bits, they may carry hugetlb specific flags that
+	 * share folio->private with folio->swap (see folio_swap_entry()).
+	 */
+	folio->swap.val = entry.val | (folio->swap.val & (nr_pages - 1));
 }
 
 /**
@@ -269,6 +273,7 @@ static void __swap_cache_do_del_folio(struct swap_cluster_info *ci,
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio);
 	VM_WARN_ON_ONCE_FOLIO(folio_test_writeback(folio), folio);
 
+	entry.val = round_down(entry.val, nr_pages);
 	si = __swap_entry_to_info(entry);
 	ci_start = swp_cluster_offset(entry);
 	ci_end = ci_start + nr_pages;
@@ -286,7 +291,7 @@ static void __swap_cache_do_del_folio(struct swap_cluster_info *ci,
 				 __swp_tb_get_flags(old_tb)));
 	} while (++ci_off < ci_end);
 
-	folio->swap.val = 0;
+	folio->swap.val &= nr_pages - 1;
 	folio_clear_swapcache(folio);
 
 	if (!folio_swapped) {
@@ -336,7 +341,7 @@ void __swap_cache_del_folio(struct swap_cluster_info *ci, struct folio *folio,
 void swap_cache_del_folio(struct folio *folio)
 {
 	struct swap_cluster_info *ci;
-	swp_entry_t entry = folio->swap;
+	swp_entry_t entry = folio_swap_entry(folio);
 
 	ci = swap_cluster_lock(__swap_entry_to_info(entry), swp_offset(entry));
 	__swap_cache_del_folio(ci, folio, entry, NULL);
@@ -362,7 +367,7 @@ void swap_cache_del_folio(struct folio *folio)
 void __swap_cache_replace_folio(struct swap_cluster_info *ci,
 				struct folio *old, struct folio *new)
 {
-	swp_entry_t entry = new->swap;
+	swp_entry_t entry = folio_swap_entry(new);
 	unsigned long nr_pages = folio_nr_pages(new);
 	unsigned int ci_off = swp_cluster_offset(entry);
 	unsigned int ci_end = ci_off + nr_pages;
@@ -387,7 +392,7 @@ void __swap_cache_replace_folio(struct swap_cluster_info *ci,
 	 */
 	if (IS_ENABLED(CONFIG_DEBUG_VM) &&
 	    folio_order(old) != folio_order(new)) {
-		ci_off = swp_cluster_offset(old->swap);
+		ci_off = swp_cluster_offset(folio_swap_entry(old));
 		ci_end = ci_off + folio_nr_pages(old);
 		while (ci_off++ < ci_end)
 			WARN_ON_ONCE(swp_tb_to_folio(__swap_table_get(ci, ci_off)) != old);
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 408f6c72fb5a..c01bb490f4fc 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -230,7 +230,7 @@ static int __try_to_reclaim_swap(struct swap_info_struct *si,
 		folio_put(folio);
 		goto again;
 	}
-	offset = swp_offset(folio->swap);
+	offset = swp_offset(folio_swap_entry(folio));
 
 	need_reclaim = ((flags & TTRS_ANYWAY) ||
 			((flags & TTRS_UNMAPPED) && !folio_mapped(folio)) ||
@@ -330,12 +330,14 @@ offset_to_swap_extent(struct swap_info_struct *sis, unsigned long offset)
 
 sector_t swap_folio_sector(struct folio *folio)
 {
-	struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
+	struct swap_info_struct *sis;
 	struct swap_extent *se;
 	sector_t sector;
+	swp_entry_t entry = folio_swap_entry(folio);
 	pgoff_t offset;
 
-	offset = swp_offset(folio->swap);
+	sis = __swap_entry_to_info(entry);
+	offset = swp_offset(entry);
 	se = offset_to_swap_extent(sis, offset);
 	sector = se->start_block + (offset - se->start_page);
 	return sector << (PAGE_SHIFT - 9);
@@ -1803,7 +1805,7 @@ int folio_alloc_swap(struct folio *folio)
  */
 int folio_dup_swap(struct folio *folio, struct page *page)
 {
-	swp_entry_t entry = folio->swap;
+	swp_entry_t entry = folio_swap_entry(folio);
 	unsigned long nr_pages = folio_nr_pages(folio);
 
 	VM_WARN_ON_FOLIO(!folio_test_locked(folio), folio);
@@ -1829,7 +1831,7 @@ int folio_dup_swap(struct folio *folio, struct page *page)
  */
 void folio_put_swap(struct folio *folio, struct page *page)
 {
-	swp_entry_t entry = folio->swap;
+	swp_entry_t entry = folio_swap_entry(folio);
 	unsigned long nr_pages = folio_nr_pages(folio);
 	struct swap_info_struct *si = __swap_entry_to_info(entry);
 
@@ -2033,7 +2035,7 @@ int swp_swapcount(swp_entry_t entry)
  */
 static bool folio_maybe_swapped(struct folio *folio)
 {
-	swp_entry_t entry = folio->swap;
+	swp_entry_t entry = folio_swap_entry(folio);
 	struct swap_cluster_info *ci;
 	unsigned int ci_off, ci_end;
 	bool ret = false;
diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
index 79cc7b546f13..a774877a919a 100644
--- a/mm/userfaultfd.c
+++ b/mm/userfaultfd.c
@@ -1383,7 +1383,7 @@ static int move_swap_pte(struct mm_struct *mm, struct vm_area_struct *dst_vma,
 	 * not locked.
 	 */
 	if (src_folio && unlikely(!folio_test_swapcache(src_folio) ||
-				  entry.val != src_folio->swap.val))
+				  entry.val != folio_swap_entry(src_folio).val))
 		return -EAGAIN;
 
 	double_pt_lock(dst_ptl, src_ptl);
diff --git a/mm/util.c b/mm/util.c
index bf0513d1d3d0..1209f85251a3 100644
--- a/mm/util.c
+++ b/mm/util.c
@@ -727,7 +727,7 @@ struct address_space *folio_mapping(const struct folio *folio)
 		return NULL;
 
 	if (unlikely(folio_test_swapcache(folio)))
-		return swap_address_space(folio->swap);
+		return swap_address_space(folio_swap_entry(folio));
 
 	mapping = folio->mapping;
 	if ((unsigned long)mapping & FOLIO_MAPPING_FLAGS)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index b4c9b8f3dfe9..9a0c8815fb69 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -731,7 +731,7 @@ static int __remove_mapping(struct address_space *mapping, struct folio *folio,
 	}
 
 	if (folio_test_swapcache(folio)) {
-		swp_entry_t swap = folio->swap;
+		swp_entry_t swap = folio_swap_entry(folio);
 
 		if (reclaimed && !mapping_exiting(mapping))
 			shadow = workingset_eviction(folio, target_memcg);
@@ -1049,7 +1049,7 @@ static bool may_enter_fs(struct folio *folio, gfp_t gfp_mask)
 	 * a swapfile that requires GFP_NOFS I/O.
 	 */
 	if (folio_test_swapcache(folio) && (gfp_mask & __GFP_IO) &&
-	    !(__swap_entry_to_info(folio->swap)->ops->flags &
+	    !(__swap_entry_to_info(folio_swap_entry(folio))->ops->flags &
 			SWAP_OPS_F_REQUIRE_NOFS))
 		return true;
 	return false;
diff --git a/mm/zswap.c b/mm/zswap.c
index f3ae3c81e48e..6ff2cda52b9e 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1475,7 +1475,7 @@ static bool zswap_store_page(struct page *page,
 bool zswap_store(struct folio *folio)
 {
 	long nr_pages = folio_nr_pages(folio);
-	swp_entry_t swp = folio->swap;
+	swp_entry_t swp = folio_swap_entry(folio);
 	struct obj_cgroup *objcg = NULL;
 	struct mem_cgroup *memcg = NULL;
 	struct zswap_pool *pool;
@@ -1580,7 +1580,7 @@ bool zswap_store(struct folio *folio)
  */
 int zswap_load(struct folio *folio)
 {
-	swp_entry_t swp = folio->swap;
+	swp_entry_t swp = folio_swap_entry(folio);
 	pgoff_t offset = swp_offset(swp);
 	struct xarray *tree = swap_zswap_tree(swp);
 	struct zswap_entry *entry;
-- 
2.53.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 2/6] mm/hugetlb: swap-in support for anonymous hugetlb folios
       [not found] <cover.1790663399.git.leizongkun@qq.com>
  2026-09-29  7:58 ` [RFC PATCH 1/6] mm/swap: introduce folio_swap_entry() and convert all folio->swap readers Zongkun Lei
@ 2026-09-29  7:59 ` Zongkun Lei
  2026-09-29  7:59 ` [RFC PATCH 3/6] mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT Zongkun Lei
                   ` (3 subsequent siblings)
  5 siblings, 0 replies; 6+ messages in thread
From: Zongkun Lei @ 2026-09-29  7:59 UTC (permalink / raw)
  To: linux-mm
  Cc: Zongkun Lei, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes,
	Vlastimil Babka, Jann Horn, Pedro Falcato, Muchun Song,
	Oscar Salvador, David Hildenbrand, Chris Li, Kairui Song,
	Kemeng Shi, Nhat Pham, Baoquan He, Barry Song, Youngjun Park,
	linux-kernel, linux-fsdevel

Teach the hugetlb fault path to recognize swap entries and read them
back, together with the full consumer-side lifecycle of such entries.
This is the consumer half of hugetlb swap support; the producer
(swap-out via MADV_PAGEOUT) arrives in a later patch, so every path
added here is provably unreachable until then - which is exactly why
it lands first: with swap-out first, a fault on a swapped-out hugetlb
page would hit the 'ret = 0' fall-through in hugetlb_fault() and
livelock.

* hugetlb_anon_swapin(), mirroring do_swap_page(): get_swap_device()
  pins the device, swap_cache_get_folio() finds in-flight folios, and
  on a miss the folio is allocated from the hugetlb pool and added to
  the swap cache via the new generic swap_cache_add_folio() helper
  (hugetlb folios cannot come from the buddy allocator, so
  __swap_cache_alloc() cannot be used).  Pool exhaustion retries
  briefly, then fails with SIGBUS, matching the reservation contract
  of the no-page path.  A hwpoisoned swapcache folio kills the
  faulter with VM_FAULT_HWPOISON_LARGE instead of being delivered.

* Fork: copy_hugetlb_page_range() duplicates the slot references of a
  swapped-out page with the new swap_dup_entries_direct() helper (the
  swap cache folio cannot be used: it is freed back to the pool once
  writeback completes while its slots stay pinned by the entries) and
  clears the exclusive bit on both sides, like copy_nonpresent_pte().

* Zap: __unmap_hugepage_range() returns the slot references with
  swap_put_entries_direct(); its present-folio branch additionally
  drains a leftover swapcache copy (anon, unmapped, non-poisoned) via
  folio_free_swap(), since hugetlb folios never sit on an LRU and
  vmscan cannot reclaim them later.

* hugetlb_wp() refuses to reuse a swapcache folio in place unless
  folio_free_swap() proves no slot is referenced anymore (fork may
  have duplicated the entries); hugetlb_change_protection() passes
  swap entries through, only the uffd-wp bit may change.

* MM_SWAPENTS follows the mainline contract (+N at swapout and fork,
  -N at swapin and zap, N = pages per huge page) so VmSwap stays
  meaningful.

* swapoff: hugetlb_unuse_vma() walks one VMA's swap entries of the
  dying type and swaps them in through the same worker, wired into
  unuse_mm(); try_to_unuse() must be able to drain every kind of
  swapped-out page.

* pagemap: report hugetlb swap entries with PM_SWAP, soft-dirty and
  uffd-wp bits, like small-page swap entries.

Signed-off-by: Zongkun Lei <leizongkun@qq.com>
---
 fs/proc/task_mmu.c      |  11 ++
 include/linux/hugetlb.h |   2 +
 include/linux/swap.h    |   1 +
 mm/hugetlb.c            | 414 +++++++++++++++++++++++++++++++++++++++-
 mm/swap.h               |   6 +
 mm/swap_state.c         |  32 ++++
 mm/swapfile.c           |  66 +++++++
 7 files changed, 531 insertions(+), 1 deletion(-)

diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
index e671b4fd8ded..7f4c5b385acb 100644
--- a/fs/proc/task_mmu.c
+++ b/fs/proc/task_mmu.c
@@ -2161,6 +2161,17 @@ static int pagemap_hugetlb_range(pte_t *ptep, unsigned long hmask,
 		if (pm->show_pfn)
 			frame = pte_pfn(pte) +
 				((addr & ~hmask) >> PAGE_SHIFT);
+	} else if (softleaf_is_swap(softleaf_from_pte(pte))) {
+		softleaf_t entry = softleaf_from_pte(pte);
+
+		if (pte_swp_soft_dirty(pte))
+			flags |= PM_SOFT_DIRTY;
+		if (pte_swp_uffd_any(pte))
+			flags |= PM_UFFD_WP;
+		if (pm->show_pfn)
+			frame = swp_type(entry) |
+				(swp_offset(entry) << MAX_SWAPFILES_SHIFT);
+		flags |= PM_SWAP;
 	} else if (pte_swp_uffd_any(pte)) {
 		flags |= PM_UFFD_WP;
 	}
diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
index 900c95e346b2..b751cba214be 100644
--- a/include/linux/hugetlb.h
+++ b/include/linux/hugetlb.h
@@ -1377,4 +1377,6 @@ hugetlb_walk(struct vm_area_struct *vma, unsigned long addr, unsigned long sz)
 	return huge_pte_offset(vma->vm_mm, addr, sz);
 }
 
+int hugetlb_unuse_vma(struct vm_area_struct *vma, unsigned int type);
+
 #endif /* _LINUX_HUGETLB_H */
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 677081238811..c2f52bb76af5 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -414,6 +414,7 @@ sector_t swap_folio_sector(struct folio *folio);
  * a swap count > 1. See comments of folio_*_swap helpers for more info.
  */
 int swap_dup_entry_direct(swp_entry_t entry);
+int swap_dup_entries_direct(swp_entry_t entry, int nr);
 void swap_put_entries_direct(swp_entry_t entry, int nr);
 
 /*
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index 8fa1bafa03d9..fc2077a03744 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -26,6 +26,7 @@
 #include <linux/string_choices.h>
 #include <linux/string_helpers.h>
 #include <linux/swap.h>
+#include <linux/swap_ops.h>
 #include <linux/leafops.h>
 #include <linux/jhash.h>
 #include <linux/numa.h>
@@ -48,6 +49,7 @@
 #include <linux/page_owner.h>
 #include "internal.h"
 #include "page_alloc.h"
+#include "swap.h"
 #include "hugetlb_vmemmap.h"
 #include "hugetlb_cma.h"
 #include "hugetlb_internal.h"
@@ -4968,6 +4970,50 @@ int copy_hugetlb_page_range(struct mm_struct *dst, struct mm_struct *src,
 			if (marker)
 				set_huge_pte_at(dst, addr, dst_pte,
 						make_pte_marker(marker), sz);
+		} else if (unlikely(softleaf_is_swap(softleaf))) {
+			/*
+			 * A swap entry of a swapped-out private hugetlb
+			 * page.  Share it with the child like an anonymous
+			 * swap entry: duplicate the slot references of
+			 * the whole huge page and account one more
+			 * swapped huge page.
+			 *
+			 * The swap cache folio cannot be used for this:
+			 * once pageout writeback completes, the huge page
+			 * is freed back to the hugetlb pool and leaves the
+			 * swap cache, while its slots stay pinned by this
+			 * very entry.  So duplicate the slot references
+			 * directly under the page table locks, like
+			 * copy_nonpresent_pte() does with
+			 * swap_dup_entry_direct().
+			 */
+			pte_t swp_pte = entry;
+			bool exclusive = pte_swp_exclusive(entry);
+
+			if (unlikely(swap_dup_entries_direct(softleaf,
+						pages_per_huge_page(h)) < 0)) {
+				spin_unlock(src_ptl);
+				spin_unlock(dst_ptl);
+				ret = -ENOMEM;
+				break;
+			}
+
+			mm_prepare_for_swap_entries(dst);
+			add_mm_counter(dst, MM_SWAPENTS, pages_per_huge_page(h));
+
+			/*
+			 * The entry is shared by two processes now, so
+			 * drop the exclusive bit on both sides, like
+			 * copy_nonpresent_pte() does for swap entries.
+			 */
+			if (exclusive) {
+				entry = pte_swp_clear_exclusive(entry);
+				set_huge_pte_at(src, addr, src_pte, entry, sz);
+				swp_pte = entry;
+			}
+			if (!userfaultfd_protected(dst_vma))
+				swp_pte = pte_swp_clear_uffd(swp_pte);
+			set_huge_pte_at(dst, addr, dst_pte, swp_pte, sz);
 		} else {
 			entry = huge_ptep_get(src_vma->vm_mm, addr, src_pte);
 			pte_folio = page_folio(pte_page(entry));
@@ -5236,6 +5282,8 @@ void __unmap_hugepage_range(struct mmu_gather *tlb, struct vm_area_struct *vma,
 		 * unmapped and its refcount is dropped, so just clear pte here.
 		 */
 		if (unlikely(!pte_present(pte))) {
+			const softleaf_t softleaf = softleaf_from_pte(pte);
+
 			/*
 			 * If the pte was wr-protected by uffd-wp in any of the
 			 * swap forms, meanwhile the caller does not want to
@@ -5249,6 +5297,27 @@ void __unmap_hugepage_range(struct mmu_gather *tlb, struct vm_area_struct *vma,
 						sz);
 			else
 				huge_pte_clear(mm, address, ptep, sz);
+
+			if (softleaf_is_swap(softleaf)) {
+				/*
+				 * A swapped-out private hugetlb page: this
+				 * pte held one reference on the folio-sized
+				 * swap slot range.  Return it (reclaiming
+				 * the swap cache once the last reference is
+				 * gone), like zap_pte_range() does for small
+				 * swap entries, and drop the accounting.
+				 */
+				swap_put_entries_direct(softleaf,
+						pages_per_huge_page(h));
+				/*
+				 * pages_per_huge_page() returns unsigned int,
+				 * so negating it stays unsigned and would
+				 * zero-extend to a huge positive long here;
+				 * cast to a signed type before negating.
+				 */
+				add_mm_counter(mm, MM_SWAPENTS,
+					       -(long)pages_per_huge_page(h));
+			}
 			spin_unlock(ptl);
 			continue;
 		}
@@ -5329,6 +5398,25 @@ void __unmap_hugepage_range(struct mmu_gather *tlb, struct vm_area_struct *vma,
 				vma_add_reservation(h, vma, address);
 		}
 
+		/*
+		 * An anonymous folio swapped in by a read fault keeps its
+		 * swapcache copy and slots pinned (only write faults and
+		 * swap-full free them at swapin time).  Unlike small pages,
+		 * hugetlb folios never sit on an LRU, so vmscan cannot drain
+		 * such a copy later; once the last mapping is gone this is
+		 * the only chance to free it.  Best effort: folio_free_swap()
+		 * rechecks swapcache/writeback/occupancy under the folio
+		 * lock and simply fails in any race, leaving the slots for
+		 * swapoff.  Keep poisoned copies: they exist to kill their
+		 * remaining owners via the swapin hwpoison interception.
+		 */
+		if (folio_test_anon(folio) && !folio_mapped(folio) &&
+		    folio_test_swapcache(folio) && !folio_test_hwpoison(folio) &&
+		    folio_trylock(folio)) {
+			folio_free_swap(folio);
+			folio_unlock(folio);
+		}
+
 		tlb_remove_page_size(tlb, folio_page(folio, 0),
 				     folio_size(folio));
 		/*
@@ -5508,8 +5596,16 @@ static vm_fault_t hugetlb_wp(struct vm_fault *vmf)
 	 * we run out of free hugetlb folios: we would have to kill processes
 	 * in scenarios that used to work. As a side effect, there can still
 	 * be leaks between processes, for example, with FOLL_GET users.
+	 *
+	 * A swapped-in folio that is still in the swap cache may back swap
+	 * slots referenced from other address spaces (fork() duplicates the
+	 * swap entries of a swapped-out page); reusing it in place would
+	 * corrupt the copy those slots still promise.  Try to free the swap:
+	 * it succeeds only when no slot is referenced anymore, in which case
+	 * reuse is safe; otherwise we have to copy.
 	 */
-	if (folio_mapcount(old_folio) == 1 && folio_test_anon(old_folio)) {
+	if (folio_mapcount(old_folio) == 1 && folio_test_anon(old_folio) &&
+	    (!folio_test_swapcache(old_folio) || folio_free_swap(old_folio))) {
 		if (!PageAnonExclusive(&old_folio->page)) {
 			folio_move_anon_rmap(old_folio, vma);
 			SetPageAnonExclusive(&old_folio->page);
@@ -5985,6 +6081,10 @@ u32 hugetlb_fault_mutex_hash(struct address_space *mapping, pgoff_t idx)
 }
 #endif
 
+static vm_fault_t hugetlb_anon_swapin(struct mm_struct *mm,
+		struct vm_area_struct *vma, unsigned long haddr,
+		pte_t *ptep, pte_t old_pte, unsigned int flags, u32 hash);
+
 vm_fault_t hugetlb_fault(struct mm_struct *mm, struct vm_area_struct *vma,
 			unsigned long address, unsigned int flags)
 {
@@ -6077,8 +6177,14 @@ vm_fault_t hugetlb_fault(struct mm_struct *mm, struct vm_area_struct *vma,
 		if (softleaf_is_hwpoison(softleaf)) {
 			ret = VM_FAULT_HWPOISON_LARGE |
 			    VM_FAULT_SET_HINDEX(hstate_index(h));
+			goto out_mutex;
 		}
 
+		/* Swapped out: read it back in. */
+		if (softleaf_is_swap(softleaf))
+			return hugetlb_anon_swapin(mm, vma, vmf.address,
+					vmf.pte, vmf.orig_pte, flags, hash);
+
 		goto out_mutex;
 	}
 
@@ -6588,6 +6694,23 @@ long hugetlb_change_protection(struct vm_area_struct *vma,
 			    (uffd_wp_resolve || uffd_rwp_resolve))
 				/* Safe to modify directly (non-present->none). */
 				huge_pte_clear(mm, address, ptep, psize);
+		} else if (unlikely(softleaf_is_swap(entry))) {
+			/*
+			 * A swap entry of a swapped-out private hugetlb
+			 * page: permissions do not apply to a non-present
+			 * entry, only the uffd-wp bit may change (like the
+			 * migration entry case above).  Anything else must
+			 * be left untouched or the swap entry encoding
+			 * would be corrupted.
+			 */
+			pte_t newpte = pte;
+
+			if (uffd_wp || uffd_rwp)
+				newpte = pte_swp_mkuffd(newpte);
+			else if (uffd_wp_resolve || uffd_rwp_resolve)
+				newpte = pte_swp_clear_uffd(newpte);
+			if (!pte_same(pte, newpte))
+				set_huge_pte_at(mm, address, ptep, newpte, psize);
 		} else {
 			pte_t old_pte;
 			unsigned int shift = huge_page_shift(hstate_vma(vma));
@@ -7427,3 +7550,292 @@ void fixup_hugetlb_reservations(struct vm_area_struct *vma)
 	if (is_vm_hugetlb_page(vma))
 		clear_vma_resv_huge_pages(vma);
 }
+/*
+ * Mirror of should_try_to_free_swap() in mm/memory.c for the hugetlb
+ * swapin path; keep in sync.  The exclusive test differs: hugetlb
+ * approximates it with the folio refcount (1 + nr_pages when only the
+ * swapcache pins the folio).
+ */
+static inline bool should_try_to_free_swap(struct swap_info_struct *si,
+					   struct folio *folio,
+					   struct vm_area_struct *vma,
+					   unsigned int flags)
+{
+	if (!folio_test_swapcache(folio))
+		return false;
+	/*
+	 * Always try to free swap cache for SWP_SYNCHRONOUS_IO devices. Swap
+	 * cache can help save some IO or memory overhead, but these devices
+	 * are fast, and meanwhile, swap cache pinning the slot deferring the
+	 * release of metadata or fragmentation is a more critical issue.
+	 */
+	if (data_race(si->flags & SWP_SYNCHRONOUS_IO))
+		return true;
+	if (mem_cgroup_swap_full(folio) || (vma->vm_flags & VM_LOCKED) ||
+		folio_test_mlocked(folio))
+		return true;
+
+	return (flags & FAULT_FLAG_WRITE) &&
+		folio_ref_count(folio) == (1 + folio_nr_pages(folio));
+}
+
+static vm_fault_t hugetlb_anon_swapin(struct mm_struct *mm,
+			struct vm_area_struct *vma, unsigned long haddr,
+			pte_t *ptep, pte_t old_pte,
+			unsigned int flags, u32 hash)
+{
+	struct swap_info_struct *si = NULL;
+	struct hstate *h = hstate_vma(vma);
+	int nr_pages = pages_per_huge_page(h);
+	swp_entry_t entry = softleaf_from_pte(old_pte);
+	struct folio *swapcache, *folio = NULL;
+	struct swap_io_ctx ctx = {};
+	bool new_folio = false;
+	spinlock_t *ptl;
+	pte_t new_pte;
+	vm_fault_t ret = 0;
+
+	/* Prevent swapoff from happening to us, and reject a bad entry. */
+	si = get_swap_device(entry);
+	if (IS_ERR_OR_NULL(si)) {
+		if (IS_ERR(si))
+			ret = VM_FAULT_SIGBUS;
+		goto out;
+	}
+
+	folio = swap_cache_get_folio(entry);
+	swapcache = folio;
+	if (likely(!folio)) {
+		int retries = 3;
+
+		/*
+		 * Swapping out releases the huge page back to the pool,
+		 * so swap-in has to compete for pool memory again.  If
+		 * the pool is exhausted, retry briefly (userspace may be
+		 * paging out cold pages right now) and then fail with
+		 * SIGBUS, matching the reservation-exhaustion contract
+		 * of the no-page fault path.  The fault path must not
+		 * block indefinitely.
+		 */
+		for (;;) {
+			folio = alloc_hugetlb_folio(vma, haddr, 1);
+			if (!IS_ERR(folio))
+				break;
+			if (!hugetlb_pte_stable(h, mm, haddr, ptep, old_pte)) {
+				/* Raced with swapin/unmap: drop and retry. */
+				folio = NULL;
+				goto out;
+			}
+			if (PTR_ERR(folio) != -ENOMEM) {
+				ret = vmf_error(PTR_ERR(folio));
+				goto out;
+			}
+			if (!retries--) {
+				ret = VM_FAULT_SIGBUS |
+					VM_FAULT_SET_HINDEX(hstate_index(h));
+				goto out;
+			}
+			schedule_timeout_uninterruptible(1);
+		}
+
+		new_folio = true;
+		__folio_set_locked(folio);
+		nr_pages = folio_nr_pages(folio);
+		__folio_set_swapbacked(folio);
+
+		/*
+		 * If the slot got cached or freed concurrently,
+		 * just drop the fault and let it retry.
+		 */
+		if (swap_cache_add_folio(folio, entry))
+			goto out_page;
+		swapcache = folio;
+
+		swap_read_folio(&ctx, folio);
+		swap_read_submit(&ctx);
+		folio_wait_locked(folio);
+	} else if (folio_test_hwpoison(folio)) {
+		/*
+		 * hwpoisoned dirty swapcache pages are kept for killing
+		 * owner processes (which may be unknown at hwpoison time)
+		 */
+		ret = VM_FAULT_HWPOISON_LARGE |
+				VM_FAULT_SET_HINDEX(hstate_index(h));
+		goto out_release;
+	}
+
+	if (!folio_trylock(folio))
+		goto out_release;
+
+	if (swapcache) {
+		/*
+		 * Make sure folio_free_swap() or swapoff did not release the
+		 * swapcache from under us. The page pin, and pte_same test
+		 * below, are not enough to exclude that. Even if it is still
+		 * swapcache, we need to check that the page's swap has not changed.
+		 */
+		if (unlikely(!folio_matches_swap_entry(folio, entry)))
+			goto out_page;
+	}
+
+	/*
+	 * If we are going to COW a private mapping later, we examine the
+	 * pending reservations for this page now. This will ensure that
+	 * any allocations necessary to record that reservation occur outside
+	 * the spinlock.
+	 */
+	if ((flags & FAULT_FLAG_WRITE) && !(vma->vm_flags & VM_SHARED)) {
+		if (vma_needs_reservation(h, vma, haddr) < 0) {
+			ret = VM_FAULT_OOM;
+			goto out_page;
+		}
+		/* Just decrements count, does not deallocate */
+		vma_end_reservation(h, vma, haddr);
+	}
+
+	ptl = huge_pte_lock(h, mm, ptep);
+	if (unlikely(!pte_same(huge_ptep_get(mm, haddr, ptep), old_pte)))
+		goto out_nomap;
+
+	if (unlikely(!folio_test_uptodate(folio))) {
+		ret = VM_FAULT_SIGBUS;
+		goto out_nomap;
+	}
+
+	arch_swap_restore(folio_swap(entry, folio), folio);
+
+	if (should_try_to_free_swap(si, folio, vma, flags))
+		folio_free_swap(folio);
+
+	add_mm_counter(mm, MM_SWAPENTS, -nr_pages);
+	hugetlb_count_add(nr_pages, mm);
+	/*
+	 * For a swap-cache hit on an already-anon folio (e.g. another task
+	 * faulted it in first), add a new rmap reference and carry the
+	 * anon-exclusive marker from the swap PTE. Otherwise this is a fresh
+	 * allocation that becomes anon here.  A shared swap entry (e.g.
+	 * duplicated by fork) must not make the folio anon-exclusive: the
+	 * folio still backs slots referenced from other address spaces, so
+	 * writes have to COW.
+	 */
+	if (swapcache && folio_test_anon(folio))
+		hugetlb_add_anon_rmap(folio, vma, haddr,
+				      pte_swp_exclusive(old_pte) ? RMAP_EXCLUSIVE : 0);
+	else {
+		hugetlb_add_new_anon_rmap(folio, vma, haddr);
+		if (!pte_swp_exclusive(old_pte))
+			ClearPageAnonExclusive(&folio->page);
+	}
+
+	new_pte = make_huge_pte(vma, folio, ((vma->vm_flags & VM_WRITE)
+				&& (vma->vm_flags & VM_SHARED)));
+	if (pte_swp_soft_dirty(old_pte))
+		new_pte = pte_mksoft_dirty(new_pte);
+	if (pte_swp_uffd(old_pte))
+		new_pte = huge_pte_mkuffd(new_pte);
+	set_huge_pte_at(mm, haddr, ptep, new_pte, huge_page_size(h));
+	spin_unlock(ptl);
+	/*
+	 * Drop the swap slot references after mapping, so raced page
+	 * faults will likely see the folio in swap cache and wait on
+	 * the folio lock.
+	 */
+	folio_put_swap(folio, NULL);
+	if (new_folio)
+		folio_set_hugetlb_migratable(folio);
+	folio_unlock(folio);
+
+out:
+	if (si)
+		put_swap_device(si);
+
+	hugetlb_vma_unlock_read(vma);
+	mutex_unlock(&hugetlb_fault_mutex_table[hash]);
+	return ret;
+
+out_nomap:
+	spin_unlock(ptl);
+out_page:
+	folio_unlock(folio);
+out_release:
+	if (si)
+		put_swap_device(si);
+
+	folio_put(folio);
+	hugetlb_vma_unlock_read(vma);
+	mutex_unlock(&hugetlb_fault_mutex_table[hash]);
+	return ret;
+}
+
+/*
+ * hugetlb_unuse_vma - swapin anonymous hugetlb pages from one VMA
+ * Called from hugetlb_unuse_mm() during swapoff.
+ */
+int hugetlb_unuse_vma(struct vm_area_struct *vma, unsigned int type)
+{
+	struct mm_struct *mm = vma->vm_mm;
+	struct hstate *h = hstate_vma(vma);
+	unsigned long addr = vma->vm_start;
+	unsigned long end = vma->vm_end;
+	unsigned long sz = huge_page_size(h);
+	unsigned long last_addr_mask = hugetlb_mask_last_page(h);
+	struct address_space *mapping;
+	pte_t *ptep;
+	pte_t pte;
+	spinlock_t *ptl;
+	swp_entry_t entry;
+	int ret = 0;
+	u32 hash;
+	pgoff_t idx;
+	vm_fault_t error;
+
+	for (; addr < end; addr += sz) {
+start:
+		ptep = hugetlb_walk(vma, addr, sz);
+		if (!ptep) {
+			addr |= last_addr_mask;
+			continue;
+		}
+
+		ptl = huge_pte_lock(h, mm, ptep);
+		pte = huge_ptep_get(mm, addr, ptep);
+
+		entry = softleaf_from_pte(pte);
+
+		/* Skip non-swap leaf entries: migration, hwpoison, device, markers */
+		if (!softleaf_is_swap(entry)) {
+			spin_unlock(ptl);
+			continue;
+		}
+
+		/* Check if this entry is from the swap device we're disabling */
+		if (swp_type(entry) != type) {
+			spin_unlock(ptl);
+			continue;
+		}
+
+		spin_unlock(ptl);
+
+		mapping = vma->vm_file->f_mapping;
+		idx = hugetlb_linear_page_index(vma, addr);
+		hash = hugetlb_fault_mutex_hash(mapping, idx);
+		/*
+		 * We found a swap entry matching the type. Need to swap it in.
+		 * This is similar to the swapin code in hugetlb_fault().
+		 */
+		mutex_lock(&hugetlb_fault_mutex_table[hash]);
+		hugetlb_vma_lock_read(vma);
+		error = hugetlb_anon_swapin(mm, vma, addr, ptep, pte, 0, hash);
+		/* hugetlb_anon_swapin() releases vma_lock and fault_mutex */
+		if (error) {
+			ret = vm_fault_to_errno(error, 0);
+			break;
+		}
+
+		/* Sometimes, hugetlb_anon_swapin() would return 0 for retry */
+		cond_resched();
+		goto start;
+	}
+
+	return ret;
+}
diff --git a/mm/swap.h b/mm/swap.h
index 19c260353e08..5e5a581cb030 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -314,6 +314,7 @@ bool swap_cache_has_folio(swp_entry_t entry);
 struct folio *swap_cache_get_folio(swp_entry_t entry);
 void *swap_cache_get_shadow(swp_entry_t entry);
 void swap_cache_del_folio(struct folio *folio);
+int swap_cache_add_folio(struct folio *folio, swp_entry_t entry);
 struct folio *swap_cache_alloc_folio(swp_entry_t target_entry, gfp_t gfp_mask,
 				     unsigned long orders, struct vm_fault *vmf,
 				     struct mempolicy *mpol, pgoff_t ilx);
@@ -396,6 +397,11 @@ static inline bool folio_matches_swap_entry(const struct folio *folio, swp_entry
 	return false;
 }
 
+static inline int swap_cache_add_folio(struct folio *folio, swp_entry_t entry)
+{
+	return -EINVAL;
+}
+
 static inline void show_swap_cache_info(void)
 {
 }
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 516052efbf62..43212df961bb 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -258,6 +258,38 @@ void __swap_cache_add_folio(struct swap_cluster_info *ci,
 	lruvec_stat_mod_folio(folio, NR_SWAPCACHE, nr_pages);
 }
 
+/**
+ * swap_cache_add_folio - Add an externally allocated folio to the swap cache.
+ * @folio: The folio to add, must be locked and swapbacked.
+ * @entry: The target swap entry, will be rounded down to the folio size.
+ *
+ * Mirrors __swap_cache_alloc() for callers that allocate the folio
+ * themselves, e.g. hugetlb whose folios must come from the hugetlb pool.
+ *
+ * Context: Caller must ensure @entry is valid and stabilize the swap
+ * device, e.g. via get_swap_device().
+ * Return: 0 on success, -ENOENT if the slot is not swapped out, -EEXIST
+ * if it is already cached, -EBUSY on a conflicting concurrent operation.
+ */
+int swap_cache_add_folio(struct folio *folio, swp_entry_t entry)
+{
+	struct swap_cluster_info *ci;
+	unsigned long nr_pages = folio_nr_pages(folio);
+	int err;
+
+	VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
+	VM_WARN_ON_ONCE(nr_pages > SWAPFILE_CLUSTER);
+
+	entry.val = round_down(entry.val, nr_pages);
+	ci = swap_cluster_lock(__swap_entry_to_info(entry), swp_offset(entry));
+	err = __swap_cache_add_check(ci, entry, nr_pages, NULL, NULL);
+	if (!err)
+		__swap_cache_add_folio(ci, folio, entry);
+	swap_cluster_unlock(ci);
+
+	return err;
+}
+
 static void __swap_cache_do_del_folio(struct swap_cluster_info *ci,
 				      struct folio *folio,
 				      swp_entry_t entry, void *shadow)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index c01bb490f4fc..280a31c43c81 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -2699,6 +2699,10 @@ static int unuse_mm(struct mm_struct *mm, unsigned int type)
 			ret = unuse_vma(vma, type);
 			if (ret)
 				break;
+		} else if (vma->anon_vma && is_vm_hugetlb_page(vma)) {
+			ret = hugetlb_unuse_vma(vma, type);
+			if (ret)
+				break;
 		}
 
 		cond_resched();
@@ -3904,6 +3908,68 @@ int swap_dup_entry_direct(swp_entry_t entry)
 	return swap_dup_entries_cluster(si, swp_offset(entry), 1);
 }
 
+/*
+ * swap_dup_entries_direct() - Increase reference count on a range of swap
+ *                             entries.
+ * @entry: First entry of range.
+ * @nr: Number of entries in range.
+ *
+ * For each swap entry in the contiguous range [entry.offset, entry.offset + nr),
+ * increase the swap count by one.  Used when duplicating a large (multi-page)
+ * swap entry, e.g. when forking a mm with swapped-out hugetlb pages, where
+ * the swap cache folio may already be gone and folio_dup_swap() cannot be
+ * used.  The range may span multiple swap clusters.
+ *
+ * Context: Same as swap_dup_entry_direct(): the caller must ensure there is
+ * no race condition on the reference owner, e.g. by holding the PTL of a PTE
+ * containing the entries, and every slot must have a count >= 1.
+ *
+ * Return: 0 on success.  -ENOMEM if the swap count maxed out and an extended
+ * table could not be allocated, -EINVAL if any entry is a bad entry.  On
+ * failure all references added so far are rolled back.
+ */
+int swap_dup_entries_direct(swp_entry_t entry, int nr)
+{
+	const unsigned long start_offset = swp_offset(entry);
+	const unsigned long end_offset = start_offset + nr;
+	unsigned long offset, cluster_end;
+	struct swap_info_struct *si;
+	int err;
+
+	si = swap_entry_to_info(entry);
+	if (WARN_ON_ONCE(!si)) {
+		pr_err_ratelimited("%s%08lx\n", Bad_file, entry.val);
+		return -EINVAL;
+	}
+	if (WARN_ON_ONCE(end_offset > si->max))
+		return -EINVAL;
+
+	offset = start_offset;
+	do {
+		cluster_end = min(round_up(offset + 1, SWAPFILE_CLUSTER),
+				  end_offset);
+		err = swap_dup_entries_cluster(si, offset,
+					       cluster_end - offset);
+		if (unlikely(err))
+			goto failed;
+		offset = cluster_end;
+	} while (offset < end_offset);
+	return 0;
+
+failed:
+	/* Roll back the references added in previous clusters. */
+	while (offset > start_offset) {
+		unsigned long cluster_start;
+
+		cluster_start = max(round_down(offset - 1, SWAPFILE_CLUSTER),
+				    start_offset);
+		swap_put_entries_cluster(si, cluster_start,
+					 offset - cluster_start, false);
+		offset = cluster_start;
+	}
+	return err;
+}
+
 #if defined(CONFIG_MEMCG) && defined(CONFIG_BLK_CGROUP)
 static bool __has_usable_swap(void)
 {
-- 
2.53.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 3/6] mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT
       [not found] <cover.1790663399.git.leizongkun@qq.com>
  2026-09-29  7:58 ` [RFC PATCH 1/6] mm/swap: introduce folio_swap_entry() and convert all folio->swap readers Zongkun Lei
  2026-09-29  7:59 ` [RFC PATCH 2/6] mm/hugetlb: swap-in support for anonymous hugetlb folios Zongkun Lei
@ 2026-09-29  7:59 ` Zongkun Lei
  2026-09-29  8:00 ` [RFC PATCH 4/6] mm/memory-failure: handle swapcached hugetlb folios Zongkun Lei
                   ` (2 subsequent siblings)
  5 siblings, 0 replies; 6+ messages in thread
From: Zongkun Lei @ 2026-09-29  7:59 UTC (permalink / raw)
  To: linux-mm
  Cc: Zongkun Lei, Muchun Song, Oscar Salvador, David Hildenbrand,
	Andrew Morton, Lorenzo Stoakes, Rik van Riel, Liam R. Howlett,
	Vlastimil Babka, Harry Yoo, Jann Horn, Lance Yang, Chris Li,
	Kairui Song, Kemeng Shi, Nhat Pham, Baoquan He, Barry Song,
	Youngjun Park, linux-kernel

Add the producer side of hugetlb swap: private anonymous hugetlb
folios (MAP_PRIVATE hugetlbfs/memfd mappings after COW) can now be
swapped out on explicit userspace request via madvise(2) MADV_PAGEOUT
/ process_madvise(2).  The kernel never swaps hugetlb pages on its
own: cold-page scoring and swap decisions belong to userspace.

The pieces:

- hugetlb_reclaim_pages() and hugetlb_reclaim_folio_list(): a
  deliberate mirror of reclaim_pages()/shrink_folio_list(), not a
  reuse.  The folio source (hstate activelists vs LRUs), the free
  path (free_huge_folio() vs free_unref_page_list()), the putback and
  the swap-cache removal (swap cluster lock vs i_pages) all differ,
  and hugetlb carries no scan_control/memcg-stat/demotion/workingset
  state.  All reusable leaf operations (rmap walk, slot allocation,
  writeout, referenced/pin checks) are shared; only the control flow
  is mirrored, so a future unification can mechanically fold the two
  together.  The machinery lives in hugetlb.c because it is pool
  management: folios come from folio_isolate_hugetlb() and go back
  via free_huge_folio() or the per-hstate active list, all under
  hugetlb_lock -- the same family as demote_pool_huge_page() and
  set_max_huge_pages().

- try_to_unmap_swap_hugetlb_one(): an rmap walker callback mirroring
  ttu_anon_swapbacked_folio() that replaces the huge PTE with a swap
  entry covering the whole folio (single walk, swp_pte_prepare(),
  mm_prepare_for_swap_entries(), MM_SWAPENTS accounting), plus the
  exported wrapper try_to_unmap_swap_hugetlb(), following the
  try_to_unmap_poisoned_hugetlb_one() precedent.  Cost to generic
  rmap code: one added line in rmap.h, no changes to existing code.

- hugetlb_folio_swap_supported(): the size gate.  Upper bound: swap
  space is allocated in ranges of at most SWAPFILE_CLUSTER (PMD size)
  pages, so gigantic folios cannot be swapped.  Lower bound: while
  swap cached, the HPG_* per-folio flags live in the low bits of
  folio->swap.val and the entry is recovered by rounding it down to
  the folio size (see folio_swap_entry()), which requires
  nr_pages >= 2^__NR_HPAGEFLAGS.  Relaxing the lower bound flag by
  flag breaks down at bit 5 (raw_hwp_unreliable, set on the
  memory-failure kmalloc-failure fallback path, independently of the
  folio size): extremely unlikely, but the consequence would be a
  shifted swap offset, i.e. silent data corruption, and correctness
  cannot rest on probability.  A VM_WARN_ON_ONCE_FOLIO in the
  swap-cache add path backstops the gate.  The gate scopes a new
  feature rather than removing anything: upstream supports no hugetlb
  swap at any size today; smaller sizes are left for a dedicated
  entry store follow-up.

- madvise: MADV_PAGEOUT on a private hugetlb VMA isolates the present
  hugepages and hands them to hugetlb_reclaim_pages().  Shared
  mappings (supported later in this series), userfaultfd-registered
  or locked VMAs, and hstates failing the size gate are rejected with
  EINVAL.

This patch is the producer; the consumer (fault/swapoff swap-in)
landed in the previous patch.  The order matters and must not be
flipped: without the swap-in side, hugetlb_fault() returns success
without installing anything for non-present entries that are neither
migration nor hwpoison entries, so if page-out came first, touching a
swapped-out page would fault in an endless loop.

Signed-off-by: Zongkun Lei <leizongkun@qq.com>
---
 include/linux/hugetlb.h |  36 +++++
 include/linux/rmap.h    |   1 +
 mm/hugetlb.c            | 317 ++++++++++++++++++++++++++++++++++++++++
 mm/madvise.c            |  76 ++++++++++
 mm/rmap.c               | 147 +++++++++++++++++++
 mm/swap_state.c         |   8 +
 6 files changed, 585 insertions(+)

diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
index b751cba214be..282185646fff 100644
--- a/include/linux/hugetlb.h
+++ b/include/linux/hugetlb.h
@@ -866,6 +866,30 @@ static inline struct hstate *folio_hstate(struct folio *folio)
 	return size_to_hstate(folio_size(folio));
 }
 
+/*
+ * Whether a hugetlb folio can be swapped out, i.e. enter the swap cache.
+ *
+ * Lower bound: while the folio is swap cached, the HPG_* per-folio flags
+ * are preserved in the low bits of folio->swap.val and the entry is
+ * recovered by rounding it down to the folio size (see folio_swap_entry()).
+ * That only works when the folio is large enough for all flag bits to fit
+ * below the swap offset.
+ *
+ * Upper bound: swap space is allocated in ranges of at most
+ * SWAPFILE_CLUSTER (PMD size) pages, so gigantic or larger folios cannot
+ * be swapped.
+ */
+static inline bool hugetlb_folio_swap_supported(const struct folio *folio)
+{
+	unsigned int order = folio_order(folio);
+
+	if (folio_nr_pages(folio) < (1UL << __NR_HPAGEFLAGS))
+		return false;
+	if (order_is_gigantic(order) || order > HPAGE_PMD_ORDER)
+		return false;
+	return true;
+}
+
 static inline unsigned hstate_index_to_shift(unsigned index)
 {
 	return hstates[index].order + PAGE_SHIFT;
@@ -1195,6 +1219,11 @@ static inline bool hstate_is_gigantic(struct hstate *h)
 	return false;
 }
 
+static inline bool hugetlb_folio_swap_supported(const struct folio *folio)
+{
+	return false;
+}
+
 static inline unsigned int pages_per_huge_page(struct hstate *h)
 {
 	return 1;
@@ -1377,6 +1406,13 @@ hugetlb_walk(struct vm_area_struct *vma, unsigned long addr, unsigned long sz)
 	return huge_pte_offset(vma->vm_mm, addr, sz);
 }
 
+/*
+ * Self-contained reclaim of isolated hugetlb folios.  Hugetlb folios are not
+ * on the normal LRU, so they are swapped out here instead of through the
+ * generic reclaim_pages()/shrink_folio_list() path.
+ */
+unsigned long hugetlb_reclaim_pages(struct list_head *folio_list);
+
 int hugetlb_unuse_vma(struct vm_area_struct *vma, unsigned int type);
 
 #endif /* _LINUX_HUGETLB_H */
diff --git a/include/linux/rmap.h b/include/linux/rmap.h
index 0b332770abee..cb45f6de9b79 100644
--- a/include/linux/rmap.h
+++ b/include/linux/rmap.h
@@ -966,6 +966,7 @@ struct rmap_walk_control {
 	bool (*invalid_vma)(struct vm_area_struct *vma, void *arg);
 };
 
+void try_to_unmap_swap_hugetlb(struct folio *folio);
 void rmap_walk(struct folio *folio, struct rmap_walk_control *rwc);
 void rmap_walk_locked(struct folio *folio, struct rmap_walk_control *rwc);
 struct anon_vma *folio_lock_anon_vma_read(const struct folio *folio,
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index fc2077a03744..c7af58399e5b 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -7550,6 +7550,323 @@ void fixup_hugetlb_reservations(struct vm_area_struct *vma)
 	if (is_vm_hugetlb_page(vma))
 		clear_vma_resv_huge_pages(vma);
 }
+
+/*
+ * Reclaim a list of isolated hugetlb folios by swapping them out.
+ *
+ * This is a deliberate mirror of shrink_folio_list(), not a reuse:
+ * hugetlb folios live on hstate activelists instead of LRUs, are
+ * unmapped through hugetlb rmap walks (whole-folio swap entries), and
+ * go back to the hstate pool on free instead of the buddy allocator.
+ * None of those hooks exist in the generic reclaim path, and shrinking
+ * that path's folio-size assumptions to fit hugetlb would complicate
+ * both sides, so the hugetlb-specific steps are reimplemented here next
+ * to the machinery they depend on (demotion and pool management live in
+ * this file for the same reason).
+ * Caller: proactive reclaim via MADV_PAGEOUT.
+ */
+static unsigned int hugetlb_reclaim_folio_list(struct list_head *folio_list)
+{
+	LIST_HEAD(ret_folios);
+	struct swap_io_ctx ctx = {};
+	unsigned int nr_reclaimed = 0;
+	struct folio *folio;
+
+	while (!list_empty(folio_list)) {
+		struct address_space *mapping;
+		struct swap_cluster_info *ci;
+		long nr_pages;
+		int refcount;
+
+		cond_resched();
+
+		folio = lru_to_folio(folio_list);
+		list_del(&folio->lru);
+
+		if (!folio_trylock(folio))
+			goto keep;
+
+		VM_BUG_ON_FOLIO(!folio_test_hugetlb(folio), folio);
+
+		/*
+		 * Never hand a poisoned hugepage back to the pool.  hugetlb
+		 * folios are always large, so this is the large-folio arm of
+		 * shrink_folio_list()'s hwpoison check: keep it isolated and
+		 * intact rather than unmapping or freeing it.
+		 */
+		if (folio_test_hwpoison(folio) ||
+		    folio_test_has_hwpoisoned(folio))
+			goto keep_locked;
+
+		VM_BUG_ON_FOLIO(folio_test_active(folio), folio);
+
+		nr_pages = folio_nr_pages(folio);
+
+		/*
+		 * Only hugetlb folios whose size fits both the swap
+		 * allocator and the swap-entry/flag sharing scheme can be
+		 * swapped out; see hugetlb_folio_swap_supported().
+		 */
+		if (!hugetlb_folio_swap_supported(folio))
+			goto keep_locked;
+		if (unlikely(!folio_evictable(folio)))
+			goto activate_locked;
+
+		/*
+		 * Only anonymous folios (MAP_PRIVATE mappings after COW)
+		 * are swapped out so far; file-backed hugetlbfs folios
+		 * gain swap support later in this series.
+		 */
+		if (!folio_test_anon(folio))
+			goto keep_locked;
+
+		/*
+		 * If the folio's swap write is still in flight, the bio holds a
+		 * writeback reference that end_swap_bio_write() drops
+		 * asynchronously via folio_end_writeback() -> folio_put().
+		 * Freezing and freeing the folio here (below) would race with
+		 * that completion and hand the same hugepage back to the pool
+		 * twice.  Wait for writeback to finish, then retry the folio.
+		 * Mirrors the synchronous-reclaim case in shrink_folio_list().
+		 * The write is still sitting in the deferred batch: submit it
+		 * before waiting, otherwise the wait outlives the caller.
+		 */
+		if (folio_test_writeback(folio)) {
+			folio_unlock(folio);
+			swap_write_submit(&ctx);
+			folio_wait_writeback(folio);
+			list_add_tail(&folio->lru, folio_list);
+			continue;
+		}
+
+		/*
+		 * Anonymous folios enter the swap cache here, before the
+		 * unmap below installs swap PTEs that reference the slots.
+		 * Hugetlb folios are not swap backed by default.
+		 */
+		folio_set_swapbacked(folio);
+		if (!folio_test_swapcache(folio)) {
+			if (folio_alloc_swap(folio))
+				goto activate_locked;
+
+			folio_mark_dirty(folio);
+		}
+
+		/*
+		 * Unmap from every process, installing swap PTEs that
+		 * reference the slots allocated above.
+		 */
+		if (folio_mapped(folio)) {
+			try_to_unmap_swap_hugetlb(folio);
+			if (folio_mapped(folio))
+				goto activate_locked;
+		}
+
+		/*
+		 * The folio is unmapped now, so it cannot be newly pinned.
+		 * A still-pinned folio cannot be reclaimed: the pin holds
+		 * extra references that would fail the refcount freeze below,
+		 * and writing it out to swap while a device is still DMAing
+		 * into it would lose those in-flight updates.  Mirrors
+		 * shrink_folio_list().
+		 */
+		if (folio_maybe_dma_pinned(folio))
+			goto activate_locked;
+
+		mapping = folio_mapping(folio);
+		if (folio_test_dirty(folio)) {
+			int res;
+
+			try_to_unmap_flush_dirty();
+
+			/* Could not write back, folio still locked. */
+			if (!mapping)
+				goto keep_locked;
+
+			if (folio_clear_dirty_for_io(folio)) {
+				folio_set_reclaim(folio);
+				res = swap_writeout(&ctx, folio);
+				if (res < 0) {
+					folio_lock(folio);
+					if (folio_mapping(folio) == mapping)
+						mapping_set_error(mapping, res);
+					folio_unlock(folio);
+				}
+				if (res == AOP_WRITEPAGE_ACTIVATE) {
+					folio_clear_reclaim(folio);
+					goto activate_locked;
+				}
+
+				if (!folio_test_writeback(folio)) {
+					/* synchronous write, or freed/zeromapped without IO */
+					folio_clear_reclaim(folio);
+				}
+				node_stat_add_folio(folio, NR_VMSCAN_WRITE);
+
+				/* Handed to the disk, folio unlocked. */
+				if (folio_test_writeback(folio)) {
+					/*
+					 * Asynchronous write still in flight.
+					 * Userspace-driven reclaim must not
+					 * return with the memory unaccounted:
+					 * wait for the write to finish and
+					 * retry the folio in this same pass.
+					 */
+					/* Dispatch the deferred write before waiting on it. */
+					swap_write_submit(&ctx);
+					folio_wait_writeback(folio);
+					list_add_tail(&folio->lru, folio_list);
+					continue;
+				}
+				if (folio_test_dirty(folio))
+					goto keep;
+				/* Synchronous write (e.g. ramdisk): free below. */
+				if (!folio_trylock(folio))
+					goto keep;
+				if (folio_test_dirty(folio) ||
+				    folio_test_writeback(folio))
+					goto keep_locked;
+				mapping = folio_mapping(folio);
+			}
+			/* Nothing to write otherwise, folio still locked. */
+		}
+
+		/*
+		 * Only a folio that made it into the swap cache can be freed
+		 * here; anything else (a clean folio still owned by its
+		 * mapping) gets put back onto the list.
+		 */
+		if (!folio_test_swapcache(folio))
+			goto keep_locked;
+
+		if (!mapping)
+			goto keep_locked;
+
+		/*
+		 * Remove the folio from the swap cache so it can be freed,
+		 * mirroring __remove_mapping().  With the swap table based
+		 * swap cache the entries are no longer in mapping->i_pages;
+		 * they are protected by the swap cluster lock instead.  No
+		 * workingset shadow is kept for hugetlb folios.
+		 */
+		BUG_ON(!folio_test_locked(folio));
+		BUG_ON(mapping != folio_mapping(folio));
+
+		ci = swap_cluster_get_and_lock_irq(folio);
+
+		refcount = 1 + folio_nr_pages(folio);
+		if (!folio_ref_freeze(folio, refcount))
+			goto cannot_free;
+		/* note: atomic_cmpxchg in folio_ref_freeze provides the smp_rmb */
+		if (unlikely(folio_test_dirty(folio))) {
+			folio_ref_unfreeze(folio, refcount);
+			goto cannot_free;
+		}
+
+		__memcg1_swapout(folio, ci);
+		__swap_cache_del_folio(ci, folio, folio_swap_entry(folio), NULL);
+		swap_cluster_unlock_irq(ci);
+
+		folio_unlock(folio);
+		nr_reclaimed += nr_pages;
+		INIT_LIST_HEAD(&folio->lru);
+		free_huge_folio(folio);
+		continue;
+
+cannot_free:
+		swap_cluster_unlock_irq(ci);
+		goto keep_locked;
+
+activate_locked:
+		/*
+		 * Not swapped out: drop any swap slots we reserved.  Only
+		 * anonymous folios can hold reserved-but-unused slots here.
+		 */
+		if (folio_test_anon(folio) && folio_test_swapcache(folio) &&
+		    (folio_test_mlocked(folio) || mem_cgroup_swap_full(folio)))
+			folio_free_swap(folio);
+keep_locked:
+		folio_unlock(folio);
+keep:
+		list_add(&folio->lru, &ret_folios);
+	}
+
+	/* Flush the deferred TTU_BATCH_FLUSH TLB invalidations. */
+	try_to_unmap_flush();
+	swap_write_submit(&ctx);
+
+	list_splice(&ret_folios, folio_list);
+	return nr_reclaimed;
+}
+
+static void hugetlb_putback_after_reclaim(struct folio *folio)
+{
+	struct hstate *h;
+
+	if (WARN_ON_ONCE(!folio_test_hugetlb(folio))) {
+		folio_put(folio);
+		return;
+	}
+
+	h = folio_hstate(folio);
+
+	spin_lock_irq(&hugetlb_lock);
+	folio_set_hugetlb_migratable(folio);
+	list_move(&folio->lru, &h->hugepage_activelist);
+	spin_unlock_irq(&hugetlb_lock);
+
+	folio_put(folio);
+}
+
+/*
+ * Reclaim a list of isolated hugetlb folios.  Mirrors reclaim_pages(): folios
+ * are grouped per node and run through the trimmed swap-out path above.  Any
+ * that could not be freed (lock/unmap/writeback contention, pins) are put
+ * back onto the per-hstate active list.
+ */
+unsigned long hugetlb_reclaim_pages(struct list_head *folio_list)
+{
+	unsigned int nr_reclaimed = 0;
+	LIST_HEAD(node_folio_list);
+	unsigned int noreclaim_flag;
+	struct folio *folio;
+	int nid;
+
+	if (list_empty(folio_list))
+		return 0;
+
+	noreclaim_flag = memalloc_noreclaim_save();
+
+	nid = folio_nid(lru_to_folio(folio_list));
+	do {
+		folio = lru_to_folio(folio_list);
+
+		if (nid == folio_nid(folio)) {
+			folio_clear_active(folio);
+			list_move(&folio->lru, &node_folio_list);
+			continue;
+		}
+
+		nr_reclaimed += hugetlb_reclaim_folio_list(&node_folio_list);
+		while (!list_empty(&node_folio_list)) {
+			/* putback does list_move(); do not list_del() first. */
+			folio = lru_to_folio(&node_folio_list);
+			hugetlb_putback_after_reclaim(folio);
+		}
+		nid = folio_nid(lru_to_folio(folio_list));
+	} while (!list_empty(folio_list));
+
+	nr_reclaimed += hugetlb_reclaim_folio_list(&node_folio_list);
+	while (!list_empty(&node_folio_list)) {
+		/* putback does list_move(); do not list_del() first. */
+		folio = lru_to_folio(&node_folio_list);
+		hugetlb_putback_after_reclaim(folio);
+	}
+
+	memalloc_noreclaim_restore(noreclaim_flag);
+
+	return nr_reclaimed;
+}
 /*
  * Mirror of should_try_to_free_swap() in mm/memory.c for the hugetlb
  * swapin path; keep in sync.  The exclusive test differs: hugetlb
diff --git a/mm/madvise.c b/mm/madvise.c
index 73c2901b9adb..a46306cd1d49 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -635,11 +635,87 @@ static void madvise_pageout_page_range(struct mmu_gather *tlb,
 	tlb_end_vma(tlb, vma);
 }
 
+/*
+ * Page out private anonymous hugetlb folios in the range: isolate each
+ * present hugepage, hand the batch to hugetlb_reclaim_pages() and let it
+ * unmap, write out and free the folios synchronously.
+ */
+static long madvise_pageout_hugetlb(struct madvise_behavior *madv_behavior)
+{
+	struct vm_area_struct *vma = madv_behavior->vma;
+	struct mm_struct *mm = madv_behavior->mm;
+	struct hstate *h = hstate_vma(vma);
+	unsigned long size = huge_page_size(h);
+	unsigned long addr, end = madv_behavior->range.end;
+	LIST_HEAD(folio_list);
+
+	/* uffd interaction with swap entries is not audited yet. */
+	if (userfaultfd_armed(vma))
+		return -EINVAL;
+	if (vma->vm_flags & (VM_LOCKED | VM_PFNMAP))
+		return -EINVAL;
+	/*
+	 * Swap PTEs are only installed for anonymous folios (MAP_PRIVATE
+	 * after COW); shared hugetlbfs mappings gain swap-out support
+	 * later in this series.
+	 */
+	if (vma->vm_flags & VM_MAYSHARE)
+		return -EINVAL;
+	/* A huge page must fit in a single swap cluster to be swappable. */
+	if (hstate_is_gigantic(h) || huge_page_order(h) > HPAGE_PMD_ORDER)
+		return -EINVAL;
+
+	/*
+	 * The caller's range is only PAGE_SIZE aligned and the hugetlb
+	 * end adjustment in madvise_dontneed_free_valid_vma() does not
+	 * apply to MADV_PAGEOUT, so align the range to huge pages here.
+	 * The range is clamped to this VMA, whose bounds are huge page
+	 * aligned, so the rounding cannot spill into a neighbour VMA.
+	 */
+	addr = ALIGN_DOWN(madv_behavior->range.start, size);
+	end = ALIGN(end, size);
+
+	for (; addr < end; addr += size) {
+		spinlock_t *ptl;
+		struct folio *folio;
+		pte_t *ptep, pte;
+
+		ptep = huge_pte_offset(mm, addr, size);
+		if (!ptep)
+			continue;
+		ptl = huge_pte_lock(h, mm, ptep);
+		pte = huge_ptep_get(mm, addr, ptep);
+		if (!pte_present(pte)) {
+			/* Hole or already-swapped: nothing to do. */
+			spin_unlock(ptl);
+			continue;
+		}
+		folio = page_folio(pte_page(pte));
+		folio_get(folio);
+		spin_unlock(ptl);
+
+		/*
+		 * On success folio_isolate_hugetlb() takes an additional
+		 * reference that hugetlb_reclaim_pages() consumes; drop
+		 * only the pin we grabbed under the ptl.
+		 */
+		folio_isolate_hugetlb(folio, &folio_list);
+		folio_put(folio);
+		cond_resched();
+	}
+
+	hugetlb_reclaim_pages(&folio_list);
+	return 0;
+}
+
 static long madvise_pageout(struct madvise_behavior *madv_behavior)
 {
 	struct mmu_gather tlb;
 	struct vm_area_struct *vma = madv_behavior->vma;
 
+	if (is_vm_hugetlb_page(vma))
+		return madvise_pageout_hugetlb(madv_behavior);
+
 	if (!can_madv_lru_vma(vma))
 		return -EINVAL;
 
diff --git a/mm/rmap.c b/mm/rmap.c
index fed0362e0bd0..2ce54cb5ff20 100644
--- a/mm/rmap.c
+++ b/mm/rmap.c
@@ -2193,6 +2193,132 @@ static bool ttu_anon_folio(struct vm_area_struct *vma, struct folio *folio,
 					 pteval);
 }
 
+/*
+ * Swap-out counterpart of try_to_unmap_one() for hugetlb folios.  For an
+ * anonymous folio, the single huge PTE mapping @folio in @vma is replaced
+ * with a swap entry covering the whole folio; for a file-backed folio the
+ * PTE is simply cleared, the swap anchor lives in the hugetlbfs page
+ * cache instead (shmem-style).  Called from the hugetlb reclaim path
+ * (hugetlb_reclaim_pages()) with the folio lock held; an anonymous folio
+ * is swapbacked with its swap slots allocated.
+ *
+ * Keep in sync with ttu_anon_swapbacked_folio().
+ */
+static bool try_to_unmap_swap_hugetlb_one(struct folio *folio,
+					  struct vm_area_struct *vma,
+					  unsigned long address, void *arg)
+{
+	struct mm_struct *mm = vma->vm_mm;
+	struct hstate *h = hstate_vma(vma);
+	DEFINE_FOLIO_VMA_WALK(pvmw, folio, vma, address, PVMW_SYNC);
+	const enum ttu_flags flags = (enum ttu_flags)(long)arg;
+	const unsigned long hsz = huge_page_size(h);
+	struct mmu_notifier_range range;
+	bool anon_exclusive;
+	pte_t pteval;
+
+	range.end = vma_address_end(&pvmw);
+	mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm,
+				address, range.end);
+	mmu_notifier_invalidate_range_start(&range);
+
+	/* There is only a single mapping in a VMA. */
+	if (!page_vma_mapped_walk(&pvmw))
+		goto range_end;
+
+	if ((vma->vm_flags & VM_LOCKED) && !(flags & TTU_IGNORE_MLOCK))
+		goto walk_abort;
+
+	VM_BUG_ON_FOLIO(!pvmw.pte, folio);
+
+	if (!folio_test_anon(folio)) {
+		/*
+		 * File-backed hugetlb folios swap out shmem-style: the PTE
+		 * is simply cleared and the swap anchor lives in the
+		 * hugetlbfs page cache, so a later fault re-enters through
+		 * hugetlb_no_page().  No swap PTE is installed and nothing
+		 * is charged to MM_SWAPENTS.  Shared mappings may share
+		 * PMDs, so flush the notifier-adjusted range.
+		 */
+		flush_cache_range(vma, range.start, range.end);
+		pteval = huge_ptep_clear_flush(vma, address, pvmw.pte);
+		if (huge_pte_dirty(pteval))
+			folio_mark_dirty(folio);
+		update_hiwater_rss(mm);
+		hugetlb_count_sub(folio_nr_pages(folio), mm);
+		hugetlb_remove_rmap(folio);
+		folio_put_refs(folio, 1);
+		goto walk_done;
+	}
+
+	/* Private anonymous mapping: no PMD sharing possible. */
+	flush_cache_range(vma, address, address + hsz);
+
+	/* Nuke the page table entry, TLB flush included. */
+	pteval = huge_ptep_clear_flush(vma, address, pvmw.pte);
+
+	if (huge_pte_dirty(pteval))
+		folio_mark_dirty(folio);
+
+	update_hiwater_rss(mm);
+
+	if (WARN_ON_ONCE(folio_test_swapbacked(folio) !=
+			 folio_test_swapcache(folio)))
+		goto restore;
+
+	anon_exclusive = PageAnonExclusive(&folio->page);
+	/* A writable mapping of an anon folio must be exclusive. */
+	VM_BUG_ON_PAGE(pte_write(pteval) && !anon_exclusive, &folio->page);
+
+	/*
+	 * The new swap allocator keeps the swap entries of a swap cache
+	 * folio pinned at count == 0; folio_dup_swap() bumps the whole
+	 * folio's entries to the initial count == 1 before the swap PTE
+	 * installed below starts referencing them.
+	 */
+	if (folio_dup_swap(folio, NULL) < 0)
+		goto restore;
+	if (arch_unmap_one(mm, vma, address, pteval) < 0)
+		goto restore_put_swap;
+
+	/*
+	 * See folio_try_share_anon_rmap(): the PTE has already been
+	 * cleared above.  hugetlb_try_share_anon_rmap() clears
+	 * PageAnonExclusive (exclusivity moves into the swap PTE built
+	 * below) and fails if the folio is GUP-pinned -- in which case
+	 * it must not be swapped out, so restore the PTE and abort.
+	 */
+	if (anon_exclusive && hugetlb_try_share_anon_rmap(folio))
+		goto restore_put_swap;
+
+	mm_prepare_for_swap_entries(mm);
+
+	set_huge_pte_at(mm, address, pvmw.pte,
+			swp_pte_prepare(folio_swap_entry(folio), pteval,
+					anon_exclusive), hsz);
+	add_mm_counter(mm, MM_SWAPENTS, pages_per_huge_page(h));
+	hugetlb_count_sub(folio_nr_pages(folio), mm);
+
+	hugetlb_remove_rmap(folio);
+	folio_put_refs(folio, 1);
+	goto walk_done;
+
+restore_put_swap:
+	folio_put_swap(folio, NULL);
+restore:
+	set_huge_pte_at(mm, address, pvmw.pte, pteval, hsz);
+walk_abort:
+	page_vma_mapped_walk_done(&pvmw);
+	mmu_notifier_invalidate_range_end(&range);
+	return false;
+
+walk_done:
+	page_vma_mapped_walk_done(&pvmw);
+range_end:
+	mmu_notifier_invalidate_range_end(&range);
+	return true;
+}
+
 /*
  * @arg: enum ttu_flags will be passed to this argument
  */
@@ -2779,6 +2905,27 @@ static bool try_to_migrate_one(struct folio *folio, struct vm_area_struct *vma,
 	return ret;
 }
 
+/*
+ * Try to remove all mappings of a hugetlb folio as part of swapping the
+ * folio out.  Mappings of an anonymous folio are replaced with swap
+ * entries; mappings of a file-backed folio are just cleared (the swap
+ * anchor lives in the hugetlbfs page cache).  The caller must hold the
+ * folio lock; for an anonymous folio it must also have made the folio
+ * swapbacked with its swap slots allocated.  Like try_to_unmap(), it is
+ * the caller's responsibility to check folio_mapped() to see whether the
+ * unmap succeeded.
+ */
+void try_to_unmap_swap_hugetlb(struct folio *folio)
+{
+	struct rmap_walk_control rwc = {
+		.rmap_one = try_to_unmap_swap_hugetlb_one,
+		.done = folio_not_mapped,
+		.anon_lock = folio_lock_anon_vma_read,
+	};
+
+	rmap_walk(folio, &rwc);
+}
+
 /**
  * try_to_migrate - try to replace all page table mappings with swap entries
  * @folio: the folio to replace page table entries for
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 43212df961bb..78e1e9b7821a 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -21,6 +21,7 @@
 #include <linux/migrate.h>
 #include <linux/vmalloc.h>
 #include <linux/huge_mm.h>
+#include <linux/hugetlb.h>
 #include <linux/shmem_fs.h>
 #include <linux/sysctl.h>
 #include <linux/swap_ops.h>
@@ -217,6 +218,13 @@ static void __swap_cache_do_add_folio(struct swap_cluster_info *ci,
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio);
 	VM_WARN_ON_ONCE_FOLIO(folio_test_swapcache(folio), folio);
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_swapbacked(folio), folio);
+	/*
+	 * Hugetlb folios too small for all per-folio flag bits to fit below
+	 * the swap offset must never reach the swap cache; they are kept
+	 * out by the hugetlb reclaim path (see folio_swap_entry()).
+	 */
+	VM_WARN_ON_ONCE_FOLIO(folio_test_hugetlb(folio) &&
+			      !hugetlb_folio_swap_supported(folio), folio);
 
 	ci_end = ci_off + nr_pages;
 	do {
-- 
2.53.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 4/6] mm/memory-failure: handle swapcached hugetlb folios
       [not found] <cover.1790663399.git.leizongkun@qq.com>
                   ` (2 preceding siblings ...)
  2026-09-29  7:59 ` [RFC PATCH 3/6] mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT Zongkun Lei
@ 2026-09-29  8:00 ` Zongkun Lei
  2026-09-29  8:01 ` [RFC PATCH 5/6] mm/hugetlb: swap support for file-backed " Zongkun Lei
  2026-09-29  8:02 ` [RFC PATCH 6/6] selftests/mm: add hugetlb_swap test, document hugetlb swap Zongkun Lei
  5 siblings, 0 replies; 6+ messages in thread
From: Zongkun Lei @ 2026-09-29  8:00 UTC (permalink / raw)
  To: linux-mm
  Cc: Zongkun Lei, Miaohe Lin, Naoya Horiguchi, Andrew Morton,
	David Hildenbrand, Lorenzo Stoakes, Rik van Riel,
	Liam R. Howlett, Vlastimil Babka, Harry Yoo, Jann Horn,
	Lance Yang, linux-kernel

Anonymous hugetlb folios can now sit in the swap cache (swap-out
writeback window, cached copy after swap-in), a state memory failure
never had to handle before.  Two gaps:

1. try_to_unmap() routes every hugetlb folio to
   try_to_unmap_poisoned_hugetlb_one(), which requires TTU_HWPOISON.
   But unmap_poisoned_folio() clears TTU_HWPOISON for a dirty
   swapcache folio -- its mappings must be replaced with swap entries,
   exactly like for a 4K swapcache folio, so the swap accounting stays
   intact and the fault path gets to kill the owner.  Routing such a
   folio to the poisoned handler fires its VM_WARN_ON_ONCE and would
   install hwpoison entries on top of live swap state.  Route hugetlb
   folios by flag instead: TTU_HWPOISON set -> the poisoned handler;
   cleared -> try_to_unmap_swap_hugetlb_one().

2. me_huge_page() has no swapcache case: a clean poisoned hugetlb
   folio in the swap cache falls into the truncate path, which evicts
   it (the swap address space has no error_remove_folio).  The folio
   is never returned to the hstate pool (a silent 2M pool leak) and
   the poison marker is lost, so a later fault would swap possibly
   corrupt data back in.  Mirror me_swapcache_dirty(): keep the
   poisoned folio in the swap cache and return MF_DELAYED, so the
   swap-in path sees folio_test_hwpoison() and kills the accessor.

Signed-off-by: Zongkun Lei <leizongkun@qq.com>
---
 mm/memory-failure.c | 19 +++++++++++++++++++
 mm/rmap.c           | 32 ++++++++++++++++++++++++++------
 2 files changed, 45 insertions(+), 6 deletions(-)

diff --git a/mm/memory-failure.c b/mm/memory-failure.c
index a8b03e2920ba..0c31c8ab5542 100644
--- a/mm/memory-failure.c
+++ b/mm/memory-failure.c
@@ -1151,6 +1151,25 @@ static int me_huge_page(struct page_state *ps, struct page *p)
 	struct address_space *mapping;
 	bool extra_pins = false;
 
+	/*
+	 * A hugetlb folio reaches the swap cache only via hugetlb
+	 * swap-out.  Keep a poisoned one there so that the swap-in
+	 * fault path intercepts folio_test_hwpoison() and kills the
+	 * accessor.  The truncate path below would evict a clean one
+	 * (the swap address space has no error_remove_folio), losing
+	 * both the folio (it is never returned to the hstate pool) and
+	 * the poison marker, so a later fault would silently swap
+	 * possibly-corrupt data back in.  Mirror me_swapcache_dirty().
+	 */
+	if (folio_test_swapcache(folio)) {
+		folio_clear_dirty(folio);
+		folio_unlock(folio);
+		/* The swap cache pin is intentionally retained. */
+		if (has_extra_refcount(ps, p, true))
+			return MF_FAILED;
+		return MF_DELAYED;
+	}
+
 	mapping = folio_mapping(folio);
 	if (mapping) {
 		res = truncate_error_folio(folio, page_to_pfn(p), mapping);
diff --git a/mm/rmap.c b/mm/rmap.c
index 2ce54cb5ff20..a88c911e8463 100644
--- a/mm/rmap.c
+++ b/mm/rmap.c
@@ -1991,8 +1991,9 @@ static bool try_to_unmap_poisoned_hugetlb_one(struct folio *folio,
 	pte_t pteval;
 
 	/*
-	 * The try_to_unmap() is only passed a hugetlb folio in the case
-	 * where the hugetlb folio is poisoned.
+	 * try_to_unmap() routes a hugetlb folio here only for memory
+	 * failure with TTU_HWPOISON set; a poisoned folio kept in the
+	 * swap cache is routed to try_to_unmap_swap_hugetlb_one() instead.
 	 */
 	VM_WARN_ON_ONCE_FOLIO(!folio_test_hwpoison(folio), folio);
 	VM_WARN_ON_ONCE(!(flags & TTU_HWPOISON));
@@ -2199,8 +2200,10 @@ static bool ttu_anon_folio(struct vm_area_struct *vma, struct folio *folio,
  * with a swap entry covering the whole folio; for a file-backed folio the
  * PTE is simply cleared, the swap anchor lives in the hugetlbfs page
  * cache instead (shmem-style).  Called from the hugetlb reclaim path
- * (hugetlb_reclaim_pages()) with the folio lock held; an anonymous folio
- * is swapbacked with its swap slots allocated.
+ * (hugetlb_reclaim_pages()) and, via try_to_unmap(), from memory failure
+ * for a poisoned folio that is kept in the swap cache (see
+ * unmap_poisoned_folio()); the folio lock is held in all cases, and an
+ * anonymous folio is swapbacked with its swap slots allocated.
  *
  * Keep in sync with ttu_anon_swapbacked_folio().
  */
@@ -2567,13 +2570,30 @@ static int folio_not_mapped(struct folio *folio)
 void try_to_unmap(struct folio *folio, enum ttu_flags flags)
 {
 	struct rmap_walk_control rwc = {
-		.rmap_one = folio_test_hugetlb(folio) ?
-				try_to_unmap_poisoned_hugetlb_one : try_to_unmap_one,
+		.rmap_one = try_to_unmap_one,
 		.arg = (void *)flags,
 		.done = folio_not_mapped,
 		.anon_lock = folio_lock_anon_vma_read,
 	};
 
+	/*
+	 * try_to_unmap() is passed a hugetlb folio only by memory failure.
+	 * With TTU_HWPOISON the mappings are replaced with hwpoison
+	 * entries.  Without it the folio is a poisoned folio that is kept
+	 * in the swap cache (unmap_poisoned_folio() cleared TTU_HWPOISON),
+	 * so the mappings are replaced with swap entries, exactly like for
+	 * a 4K swapcache folio: the swap accounting stays intact and the
+	 * fault path gets to kill the owner.
+	 */
+	if (folio_test_hugetlb(folio)) {
+		if (flags & TTU_HWPOISON)
+			rwc.rmap_one = try_to_unmap_poisoned_hugetlb_one;
+		else if (WARN_ON_ONCE(!folio_test_swapcache(folio)))
+			return;
+		else
+			rwc.rmap_one = try_to_unmap_swap_hugetlb_one;
+	}
+
 	if (flags & TTU_RMAP_LOCKED)
 		rmap_walk_locked(folio, &rwc);
 	else
-- 
2.53.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 5/6] mm/hugetlb: swap support for file-backed hugetlb folios
       [not found] <cover.1790663399.git.leizongkun@qq.com>
                   ` (3 preceding siblings ...)
  2026-09-29  8:00 ` [RFC PATCH 4/6] mm/memory-failure: handle swapcached hugetlb folios Zongkun Lei
@ 2026-09-29  8:01 ` Zongkun Lei
  2026-09-29  8:02 ` [RFC PATCH 6/6] selftests/mm: add hugetlb_swap test, document hugetlb swap Zongkun Lei
  5 siblings, 0 replies; 6+ messages in thread
From: Zongkun Lei @ 2026-09-29  8:01 UTC (permalink / raw)
  To: linux-mm
  Cc: Zongkun Lei, Muchun Song, Oscar Salvador, David Hildenbrand,
	Matthew Wilcox (Oracle),
	Jan Kara, Andrew Morton, Lorenzo Stoakes, Liam R. Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Jann Horn, Chris Li, Kairui Song, Kemeng Shi, Nhat Pham,
	Baoquan He, Barry Song, Youngjun Park, linux-kernel,
	linux-fsdevel

Extend hugetlb swap-out to shared file-backed mappings, using the
shmem model: swap-out only clears the PTEs -- no swap entry is ever
installed for a file folio -- and the anchor (a swap value entry,
swp_to_radix_entry()) is left in the hugetlbfs page cache, holding a
swap count reference on the slots.  Without this, instantiated shared
huge pages pin the pool absolutely until truncate/unlink: QEMU guests
backed by MAP_SHARED hugetlbfs files (-mem-path) are the canonical
victim.  It also fixes an ABI asymmetry userspace can trip over:
MADV_PAGEOUT succeeds on private hugetlb and every 4K/THP mapping
including shmem, but used to fail with EINVAL on shared hugetlb.

The pieces, each mirroring its shmem counterpart:

- Swap-out: hugetlbfs_writeout() allocates the folio-sized slot range,
  links the inode on the swapoff traversal list, moves the anchor into
  the page cache in place of the folio (holding a dup'ed swap count),
  and writes out.  The reclaim loop drops its anonymous-only gate and
  MADV_PAGEOUT drops the VM_MAYSHARE rejection.

- Refault swap-in: hugetlb_no_page() recognizes the anchor (the page
  cache lookup reports it as a hole) and swaps the folio back in via
  hugetlbfs_do_swapin(), reinstalling it in place of the anchor --
  indistinguishable from a page-cache hit afterwards.  This consumer
  must never lag the producer: the base fault path treats value
  entries as holes and would silently zero-fill over swapped data.

- read(2) swap-in: hugetlbfs_read_iter() resolves anchors through
  hugetlbfs_swapin_read() instead of zero-filling; swapin failures
  surface as -EIO/-EFAULT, never stale data.

- Lifecycle: truncate/hole-punch/evict free anchors and their slots
  via hugetlbfs_free_swap(), with UAF guards (cache entry removed
  before the anchor's swap count is dropped; a concurrent reclaim
  holding the folio is waited out).  Inode eviction drains the
  swapoff traversal list.

- swapoff: try_to_unuse() calls hugetlbfs_unuse() right after
  shmem_unuse(); since file folios install no swap PTEs, the mm walk
  cannot reach their slots -- the per-inode swaplist traversal can.

Accounting stays symmetric with free_huge_folio() on every path
(hugetlb_cgroup, memcg swap/hugetlb charge, global reserve,
NR_HUGETLB); the vma-less read/swapoff swapin charges the current
context, mirroring shmem swapoff.

Signed-off-by: Zongkun Lei <leizongkun@qq.com>
---
 fs/hugetlbfs/inode.c    | 128 ++++++-
 include/linux/hugetlb.h |  20 ++
 include/linux/pagemap.h |   7 +
 mm/hugetlb.c            | 733 ++++++++++++++++++++++++++++++++++++++--
 mm/internal.h           |  19 +-
 mm/madvise.c            |  14 +-
 mm/swapfile.c           |   7 +
 7 files changed, 882 insertions(+), 46 deletions(-)

diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c
index 7611a8470ea2..70fd4a0692eb 100644
--- a/fs/hugetlbfs/inode.c
+++ b/fs/hugetlbfs/inode.c
@@ -250,6 +250,42 @@ static ssize_t hugetlbfs_read_iter(struct kiocb *iocb, struct iov_iter *to)
 
 		/* Find the folio */
 		folio = filemap_lock_hugetlb_folio(h, mapping, index);
+		if (IS_ERR(folio)) {
+			/*
+			 * The lookup reports a swapped-out page (swap anchor
+			 * value entry) as a hole; try to swap it back in
+			 * before falling back to zero-fill.
+			 */
+			if (PTR_ERR(folio) == -ENOENT) {
+				folio = hugetlbfs_swapin_read(inode, index);
+				if (IS_ERR(folio)) {
+					int err = PTR_ERR(folio);
+
+					/* Raced with a concurrent swapin. */
+					if (err == -EAGAIN)
+						continue;
+					/*
+					 * -ENOENT is a genuine hole: fall
+					 * through and zero-fill.  Anything
+					 * else is a swapin failure that must
+					 * surface as a read error, never
+					 * stale/zeroed data.
+					 */
+					if (err != -ENOENT) {
+						/* A corrupt page or bad swap
+						 * entry reads as -EIO. */
+						if (err == -EHWPOISON ||
+						    err == -EINVAL)
+							retval = -EIO;
+						/* POSIX: only report if nothing
+						 * was copied yet. */
+						if (!retval)
+							retval = -EFAULT;
+						break;
+					}
+				}
+			}
+		}
 		if (IS_ERR(folio)) {
 			/*
 			 * We have a HOLE, zero out the user-buffer for the
@@ -494,8 +530,9 @@ hugetlb_vmdelete_list(struct address_space *mapping, pgoff_t start,
 
 /*
  * Called with hugetlb fault mutex held.
+ * Returns true if the folio was actually removed, false otherwise.
  */
-static void remove_inode_single_folio(struct hstate *h, struct inode *inode,
+static bool remove_inode_single_folio(struct hstate *h, struct inode *inode,
 		struct address_space *mapping, struct folio *folio,
 		pgoff_t index, bool truncate_op)
 {
@@ -508,6 +545,20 @@ static void remove_inode_single_folio(struct hstate *h, struct inode *inode,
 	 * lock to guarantee no concurrent migration.
 	 */
 	folio_lock(folio);
+
+	/*
+	 * Re-check the mapping under the folio lock: a concurrent
+	 * hugetlbfs_writeout() may have replaced this folio with a
+	 * swap anchor in the page cache and cleared folio->mapping
+	 * before we acquired the lock.  In that case the folio is no
+	 * longer ours to remove; the anchor will be handled when
+	 * remove_inode_hugepages() iterates the page cache again.
+	 */
+	if (unlikely(folio_mapping(folio) != mapping)) {
+		folio_unlock(folio);
+		return false;
+	}
+
 	if (unlikely(folio_mapped(folio)))
 		hugetlb_unmap_file_folio(h, mapping, folio, index);
 
@@ -527,6 +578,7 @@ static void remove_inode_single_folio(struct hstate *h, struct inode *inode,
 	}
 
 	folio_unlock(folio);
+	return true;
 }
 
 /*
@@ -556,30 +608,72 @@ static void remove_inode_hugepages(struct inode *inode, loff_t lstart,
 	struct address_space *mapping = &inode->i_data;
 	const pgoff_t end = lend >> PAGE_SHIFT;
 	struct folio_batch fbatch;
+	pgoff_t indices[FOLIO_BATCH_SIZE];
 	pgoff_t next, index;
 	int i, freed = 0;
 	bool truncate_op = (lend == LLONG_MAX);
 
 	folio_batch_init(&fbatch);
 	next = lstart >> PAGE_SHIFT;
-	while (filemap_get_folios(mapping, &next, end - 1, &fbatch)) {
+	while (find_get_entries(mapping, &next, end - 1, &fbatch, indices)) {
 		for (i = 0; i < folio_batch_count(&fbatch); ++i) {
 			struct folio *folio = fbatch.folios[i];
 			u32 hash = 0;
 
-			index = folio->index >> huge_page_order(h);
+			index = indices[i] >> huge_page_order(h);
 			hash = hugetlb_fault_mutex_hash(mapping, index);
 			mutex_lock(&hugetlb_fault_mutex_table[hash]);
 
+			if (xa_is_value(folio)) {
+				/*
+				 * Swap anchor of a swapped-out file page.
+				 * hugetlbfs page caches carry no other
+				 * value entries.
+				 */
+				if (!hugetlbfs_free_swap(mapping, indices[i],
+							 folio, h)) {
+					/* Entry changed; rescan from here. */
+					next = indices[i];
+					mutex_unlock(&hugetlb_fault_mutex_table[hash]);
+					break;
+				}
+
+				/*
+				 * Account anchors like folios: unreserve
+				 * per page for hole punch, count and
+				 * unreserve once at the end for truncate.
+				 */
+				if (!truncate_op) {
+					if (unlikely(hugetlb_unreserve_pages(inode,
+							    index, index + 1, 1)))
+						hugetlb_fix_reserve_counts(inode);
+				} else {
+					freed++;
+				}
+
+				mutex_unlock(&hugetlb_fault_mutex_table[hash]);
+				continue;
+			}
+
 			/*
 			 * Remove folio that was part of folio_batch.
 			 */
-			remove_inode_single_folio(h, inode, mapping, folio,
-						  index, truncate_op);
+			if (!remove_inode_single_folio(h, inode, mapping, folio,
+						       index, truncate_op)) {
+				/*
+				 * Raced with swapout: the folio was
+				 * replaced by a swap anchor.  Rescan
+				 * from this index to find it.
+				 */
+				next = indices[i];
+				mutex_unlock(&hugetlb_fault_mutex_table[hash]);
+				break;
+			}
 			freed++;
 
 			mutex_unlock(&hugetlb_fault_mutex_table[hash]);
 		}
+		folio_batch_remove_exceptionals(&fbatch);
 		folio_batch_release(&fbatch);
 		cond_resched();
 	}
@@ -596,6 +690,7 @@ static void hugetlbfs_evict_inode(struct inode *inode)
 
 	trace_hugetlbfs_evict_inode(inode);
 	remove_inode_hugepages(inode, 0, LLONG_MAX);
+	hugetlbfs_evict_drain_swaplist(inode);
 
 	resv_map = HUGETLBFS_I(inode)->resv_map;
 	/* Only regular and link inodes have associated reserve maps */
@@ -782,6 +877,25 @@ static long hugetlbfs_fallocate(struct file *file, int mode, loff_t offset,
 			continue;
 		}
 
+		/*
+		 * A swap anchor means the page is swapped out: it is
+		 * already backed (on swap), so there is nothing to
+		 * preallocate.  filemap_get_folio() reports anchors as
+		 * holes, so check explicitly.  The fault mutex held here
+		 * keeps the anchor stable.
+		 */
+		{
+			void *entry = filemap_get_entry(mapping,
+						index << huge_page_order(h));
+
+			if (entry) {
+				if (!xa_is_value(entry))
+					folio_put(entry);
+				mutex_unlock(&hugetlb_fault_mutex_table[hash]);
+				continue;
+			}
+		}
+
 		/*
 		 * Allocate folio without setting the avoid_reserve argument.
 		 * There certainly are no reserves associated with the
@@ -921,6 +1035,10 @@ static struct inode *hugetlbfs_get_inode(struct super_block *sb,
 		simple_inode_init_ts(inode);
 		info->resv_map = resv_map;
 		info->seals = F_SEAL_SEAL;
+		/* Initialize swapoff support fields */
+		INIT_LIST_HEAD(&info->swaplist);
+		atomic_set(&info->swapped, 0);
+		atomic_set(&info->stop_eviction, 0);
 		switch (mode & S_IFMT) {
 		default:
 			init_special_inode(inode, mode, dev);
diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
index 282185646fff..4c74323e4b01 100644
--- a/include/linux/hugetlb.h
+++ b/include/linux/hugetlb.h
@@ -508,6 +508,9 @@ struct hugetlbfs_inode_info {
 	struct inode vfs_inode;
 	struct resv_map *resv_map;
 	unsigned int seals;
+	struct list_head swaplist;	/* Link to hugetlbfs_swaplist */
+	atomic_t swapped;		/* Count of swapped pages */
+	atomic_t stop_eviction;		/* Prevent eviction during swapoff */
 };
 
 static inline struct hugetlbfs_inode_info *HUGETLBFS_I(struct inode *inode)
@@ -890,6 +893,23 @@ static inline bool hugetlb_folio_swap_supported(const struct folio *folio)
 	return true;
 }
 
+/*
+ * File-page swap anchor removal for truncate/hole-punch/evict, and the
+ * evict-time drain of the swapoff traversal list.  Defined in mm/hugetlb.c,
+ * called from fs/hugetlbfs/inode.c.
+ */
+long hugetlbfs_free_swap(struct address_space *mapping, pgoff_t index,
+			 void *radswap, struct hstate *h);
+void hugetlbfs_evict_drain_swaplist(struct inode *inode);
+
+/*
+ * Read-path swapin of a swapped-out hugetlbfs page (mm/hugetlb.c,
+ * called from fs/hugetlbfs/inode.c), and the swapoff traversal of all
+ * hugetlbfs swap anchors (called from try_to_unuse()).
+ */
+struct folio *hugetlbfs_swapin_read(struct inode *inode, pgoff_t index);
+int hugetlbfs_unuse(unsigned int type);
+
 static inline unsigned hstate_index_to_shift(unsigned index)
 {
 	return hstates[index].order + PAGE_SHIFT;
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 0adfa6605653..85b98e5ee4e3 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -987,6 +987,13 @@ static inline bool folio_contains(const struct folio *folio, pgoff_t index)
 
 unsigned filemap_get_folios(struct address_space *mapping, pgoff_t *start,
 		pgoff_t end, struct folio_batch *fbatch);
+/*
+ * find_get_entries() also returns swap/shadow value entries (unlike
+ * filemap_get_folios()); hugetlbfs needs it to enumerate the swap
+ * anchors left in its page cache by hugetlb file-page swap-out.
+ */
+unsigned find_get_entries(struct address_space *mapping, pgoff_t *start,
+		pgoff_t end, struct folio_batch *fbatch, pgoff_t *indices);
 unsigned filemap_get_folios_contig(struct address_space *mapping,
 		pgoff_t *start, pgoff_t end, struct folio_batch *fbatch);
 unsigned filemap_get_folios_tag(struct address_space *mapping, pgoff_t *start,
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index c7af58399e5b..89b066379586 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -5831,6 +5831,9 @@ static bool hugetlb_pte_stable(struct hstate *h, struct mm_struct *mm, unsigned
 	return same;
 }
 
+static struct folio *hugetlb_no_page_swapin(struct address_space *mapping,
+					    struct vm_fault *vmf);
+
 static vm_fault_t hugetlb_no_page(struct address_space *mapping,
 			struct vm_fault *vmf)
 {
@@ -5864,6 +5867,26 @@ static vm_fault_t hugetlb_no_page(struct address_space *mapping,
 	new_folio = false;
 	folio = filemap_lock_hugetlb_folio(h, mapping, vmf->pgoff);
 	if (IS_ERR(folio)) {
+		/*
+		 * A swapped-out file page leaves a swap anchor value entry
+		 * in the page cache, which filemap_lock_hugetlb_folio()
+		 * reports as -ENOENT just like a real hole.  Check for the
+		 * anchor first and swap the page back in if it is there.
+		 */
+		folio = hugetlb_no_page_swapin(mapping, vmf);
+		if (!IS_ERR(folio))
+			goto swapped_in;
+		if (PTR_ERR(folio) != -ENOENT) {
+			if (PTR_ERR(folio) == -EAGAIN)
+				ret = 0;	/* raced, retry the fault */
+			else if (PTR_ERR(folio) == -EHWPOISON)
+				ret = VM_FAULT_HWPOISON_LARGE |
+					VM_FAULT_SET_HINDEX(hstate_index(h));
+			else
+				ret = vmf_error(PTR_ERR(folio));
+			goto out;
+		}
+
 		size = i_size_read(mapping->host) >> huge_page_shift(h);
 		if (vmf->pgoff >= size)
 			goto out;
@@ -5972,6 +5995,7 @@ static vm_fault_t hugetlb_no_page(struct address_space *mapping,
 		}
 	}
 
+swapped_in:
 	/*
 	 * If we are going to COW a private mapping later, we examine the
 	 * pending reservations for this page now. This will ensure that
@@ -7551,19 +7575,667 @@ void fixup_hugetlb_reservations(struct vm_area_struct *vma)
 		clear_vma_resv_huge_pages(vma);
 }
 
+/*
+ * hugetlbfs inodes with at least one swapped-out file page; traversed by
+ * swapoff (hugetlbfs_unuse) to bring them back in.  Mirrored on
+ * shmem_swaplist.
+ */
+static LIST_HEAD(hugetlbfs_swaplist);
+static DEFINE_MUTEX(hugetlbfs_swaplist_mutex);
+
+/*
+ * Somewhat like shmem_replace_entry(), but replacing the whole huge page
+ * range in a hugetlbfs page cache.
+ */
+static int hugetlb_replace_entry(struct address_space *mapping,
+				 pgoff_t index, void *expected,
+				 void *replacement, unsigned int order)
+{
+	XA_STATE_ORDER(xas, &mapping->i_pages, index, order);
+	void *item;
+
+	VM_BUG_ON(!expected);
+	VM_BUG_ON(!replacement);
+	item = xas_load(&xas);
+	if (item != expected)
+		return -ENOENT;
+	xas_store(&xas, replacement);
+	if (WARN_ON_ONCE(xas_error(&xas)))
+		return xas_error(&xas);
+	return 0;
+}
+
+/*
+ * Somewhat like shmem_delete_from_page_cache(), but for a hugetlb folio:
+ * substitutes the swap anchor for @folio in the page cache.  After this
+ * the folio is only referenced by the swap cache.
+ */
+static void hugetlb_delete_from_page_cache(struct folio *folio, void *radswap)
+{
+	struct address_space *mapping = folio->mapping;
+	long nr = folio_nr_pages(folio);
+	int error;
+
+	xa_lock_irq(&mapping->i_pages);
+	error = hugetlb_replace_entry(mapping, folio->index, folio, radswap,
+				      folio_order(folio));
+	folio->mapping = NULL;
+	mapping->nrpages -= nr;
+	xa_unlock_irq(&mapping->i_pages);
+	folio_put_refs(folio, nr);
+	BUG_ON(error);
+}
+
+/*
+ * Swap out a file-backed hugetlb folio, shmem-style: the folio is moved
+ * into the swap cache and its place in the hugetlbfs page cache is taken
+ * by a swap anchor value entry, which holds a swap count reference on the
+ * slots and is what refault (hugetlb_no_page) and swapoff look up later.
+ * No swap PTE is ever installed for file folios.
+ *
+ * Return: the swap_writeout() result with the folio unlocked, or
+ * AOP_WRITEPAGE_ACTIVATE with the folio still locked and redirtied if it
+ * could not be swapped out.
+ */
+static int hugetlbfs_writeout(struct swap_io_ctx *ctx, struct folio *folio)
+{
+	struct address_space *mapping;
+	struct hugetlbfs_inode_info *info;
+	long nr_pages = folio_nr_pages(folio);
+
+	/* Retry of an already-anchored folio: just drive the write. */
+	if (folio_test_swapcache(folio))
+		return swap_writeout(ctx, folio);
+
+	mapping = folio->mapping;
+	info = HUGETLBFS_I(mapping->host);
+
+	/*
+	 * Move the folio into the swap cache; the slots come pinned at
+	 * count == 0.  The swap cache add requires the folio to be marked
+	 * swapbacked first.
+	 */
+	folio_set_swapbacked(folio);
+	if (folio_alloc_swap(folio))
+		goto redirty;
+
+	/*
+	 * Add the inode to the swapoff traversal list before the anchor
+	 * replaces the folio in the page cache, while the folio lock is
+	 * still serialization against a racing inode eviction.
+	 */
+	mutex_lock(&hugetlbfs_swaplist_mutex);
+	if (list_empty(&info->swaplist))
+		list_add(&info->swaplist, &hugetlbfs_swaplist);
+	atomic_add(nr_pages, &info->swapped);
+	mutex_unlock(&hugetlbfs_swaplist_mutex);
+
+	/* The anchor holds a swap count reference on the slots. */
+	folio_dup_swap(folio, NULL);
+	hugetlb_delete_from_page_cache(folio,
+			swp_to_radix_entry(folio_swap_entry(folio)));
+
+	return swap_writeout(ctx, folio);
+
+redirty:
+	folio_mark_dirty(folio);
+	return AOP_WRITEPAGE_ACTIVATE;	/* Return with folio locked */
+}
+
+/*
+ * Somewhat like shmem_add_to_page_cache(), but for a hugetlb folio:
+ * swap the folio back into the page cache in place of the swap anchor
+ * left behind by hugetlbfs_writeout().  The anchor is re-validated
+ * under the xarray lock; if it is gone (a racing swapin already
+ * installed its folio, or truncation removed it), -EEXIST is returned
+ * and the caller retries or drops the swapin.  The folio is still
+ * swap-backed here; the caller drops the swap cache membership
+ * afterwards, mirroring shmem_swapin_folio().
+ */
+static int hugetlbfs_add_to_page_cache(struct folio *folio,
+		struct address_space *mapping, pgoff_t index,
+		void *expected, gfp_t gfp)
+{
+	XA_STATE_ORDER(xas, &mapping->i_pages, index, folio_order(folio));
+	long nr = folio_nr_pages(folio);
+
+	VM_BUG_ON_FOLIO(index != round_down(index, nr), folio);
+	VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
+	VM_BUG_ON_FOLIO(!folio_test_swapbacked(folio), folio);
+
+	folio_ref_add(folio, nr);
+	folio->mapping = mapping;
+	folio->index = index;
+
+	do {
+		xas_lock_irq(&xas);
+		if (xas_load(&xas) != expected)
+			xas_set_err(&xas, -EEXIST);
+		else
+			xas_store(&xas, folio);
+		if (!xas_error(&xas))
+			mapping->nrpages += nr;
+		xas_unlock_irq(&xas);
+	} while (xas_nomem(&xas, gfp));
+
+	if (xas_error(&xas)) {
+		folio->mapping = NULL;
+		folio_put_refs(folio, nr);
+		return xas_error(&xas);
+	}
+
+	return 0;
+}
+
+/*
+ * Swap in a hugetlbfs file page whose swap anchor sits at @index in the
+ * page cache.  Shared between the fault path (hugetlb_no_page_swapin(),
+ * @vma != NULL) and the swapoff/read paths (@vma == NULL).
+ *
+ * On success returns the folio locked, with a reference, reinstalled in
+ * the page cache in place of the anchor.  On failure returns an ERR_PTR:
+ *
+ *	-EAGAIN		raced with a concurrent swapin/swapout; retry
+ *	-EEXIST		a racing swapin already installed its folio
+ *	-EHWPOISON	the swapcache copy is poisoned; kill the accessor
+ *	-ENOMEM		could not allocate a hugepage from the pool
+ *	-EIO		the swap read failed
+ */
+static struct folio *hugetlbfs_do_swapin(struct address_space *mapping,
+			pgoff_t index, swp_entry_t entry,
+			struct vm_area_struct *vma, struct hstate *h,
+			unsigned long haddr)
+{
+	struct inode *inode = mapping->host;
+	struct hugetlbfs_inode_info *info = HUGETLBFS_I(inode);
+	long nr_pages = pages_per_huge_page(h);
+	struct swap_info_struct *si;
+	struct swap_io_ctx ctx = {};
+	struct folio *folio;
+	bool new_folio = false;
+	int error;
+
+	/* Prevent swapoff from happening to us, and reject a bad entry. */
+	si = get_swap_device(entry);
+	if (IS_ERR_OR_NULL(si))
+		return ERR_PTR(-EINVAL);
+
+	folio = swap_cache_get_folio(entry);
+	if (!folio) {
+		/*
+		 * Swapping out releases the huge page back to the pool, so
+		 * swapin has to compete for pool memory again; the caller
+		 * turns an allocation failure into SIGBUS, matching the
+		 * reservation-exhaustion contract of the no-page fault path.
+		 */
+		if (vma) {
+			folio = alloc_hugetlb_folio(vma, haddr, 1);
+		} else {
+			/*
+			 * The read/swapoff path has no vma.  Route through
+			 * the same charged allocator as the fault path so
+			 * that hugetlb_cgroup/memcg charges, the global
+			 * reservation consumed by the original fault, and
+			 * NR_HUGETLB all stay paired with free_huge_folio();
+			 * the charges go to the current context, mirroring
+			 * what shmem does for swapoff-triggered swapins.
+			 */
+			struct mempolicy_interpreted mpoli = {
+				.nid = numa_node_id(),
+				.mode = MPOL_DEFAULT,
+				.nodemask = NULL,
+			};
+
+			folio = hugetlb_alloc_folio(h, &mpoli,
+					HUGETLB_ALLOC_USE_GLOBAL_RESERVATIONS |
+					HUGETLB_ALLOC_CHARG_CGROUP_RSVD);
+		}
+		if (IS_ERR(folio)) {
+			error = PTR_ERR(folio);
+			goto out_put_si;
+		}
+		new_folio = true;
+		__folio_set_locked(folio);
+		__folio_set_swapbacked(folio);
+
+		/*
+		 * If the slot got cached or freed concurrently, just drop
+		 * the fault and let it retry.
+		 */
+		if (swap_cache_add_folio(folio, entry)) {
+			error = -EAGAIN;
+			goto out_new_folio;
+		}
+
+		swap_read_folio(&ctx, folio);
+		swap_read_submit(&ctx);
+	} else if (folio_test_hwpoison(folio)) {
+		/*
+		 * hwpoisoned dirty swapcache pages are kept for killing
+		 * owner processes (which may be unknown at hwpoison time)
+		 */
+		error = -EHWPOISON;
+		goto out_put_folio;
+	}
+
+	/* The read completion (or the current writer) drops the lock. */
+	folio_lock(folio);
+	if (unlikely(!folio_matches_swap_entry(folio, entry))) {
+		/* Raced with a swapin that already finished: retry. */
+		error = -EAGAIN;
+		goto out_unlock;
+	}
+	folio_wait_writeback(folio);
+	if (!folio_test_uptodate(folio)) {
+		error = -EIO;
+		goto out_unlock;
+	}
+
+	arch_swap_restore(folio_swap(entry, folio), folio);
+
+	/*
+	 * Replace the anchor with the folio, validated against @entry
+	 * under the xarray lock; a racing swapin that already installed
+	 * its folio (or a truncation that removed the anchor) is
+	 * reported as -EEXIST.
+	 */
+	error = hugetlbfs_add_to_page_cache(folio, mapping,
+			index << huge_page_order(h),
+			swp_to_radix_entry(entry), GFP_KERNEL);
+	if (error)
+		goto out_unlock;
+
+	mutex_lock(&hugetlbfs_swaplist_mutex);
+	if (atomic_sub_return(nr_pages, &info->swapped) == 0)
+		list_del_init(&info->swaplist);
+	mutex_unlock(&hugetlbfs_swaplist_mutex);
+
+	/*
+	 * Drop the anchor's swap count reference and the swap cache
+	 * membership, mirroring shmem_swapin_folio(); the slots are freed
+	 * once both are gone.
+	 */
+	folio_put_swap(folio, NULL);
+	swap_cache_del_folio(folio);
+	folio_mark_dirty(folio);
+	put_swap_device(si);
+
+	if (new_folio)
+		folio_set_hugetlb_migratable(folio);
+
+	return folio;
+
+out_unlock:
+	if (new_folio) {
+		/*
+		 * Never leave a fresh folio orphaned in the swap cache:
+		 * its pins would pin the slots after the anchor is gone.
+		 */
+		swap_cache_del_folio(folio);
+	}
+	folio_unlock(folio);
+out_put_folio:
+	folio_put(folio);
+out_put_si:
+	put_swap_device(si);
+	return ERR_PTR(error);
+out_new_folio:
+	folio_unlock(folio);
+	folio_put(folio);
+	put_swap_device(si);
+	return ERR_PTR(error);
+}
+
+/*
+ * Swap in the hugetlbfs page at @index for the read(2) and swapoff
+ * paths.  Takes the fault mutex, which both stabilizes the anchor and
+ * serializes against fault-path swapins of the same index.
+ *
+ * On success returns the folio locked with an elevated refcount; the
+ * caller is responsible for folio_unlock() + folio_put().  On failure
+ * returns an ERR_PTR that the caller must distinguish carefully:
+ *	-ENOENT		the slot is genuinely absent (a real hole);
+ *			zero-fill is correct
+ *	-EAGAIN		raced with a concurrent swapin (slot now holds a
+ *			folio, or the swapin itself returned -EEXIST);
+ *			the caller should re-lookup
+ *	other		-ENOMEM/-EIO/-EHWPOISON/-EINVAL propagated from
+ *			the swapin; the caller must turn these into a
+ *			read error, never zero-fill
+ */
+struct folio *hugetlbfs_swapin_read(struct inode *inode, pgoff_t index)
+{
+	struct address_space *mapping = inode->i_mapping;
+	struct hstate *h = hstate_inode(inode);
+	struct folio *folio;
+	swp_entry_t entry;
+	void *xa_val;
+	u32 hash;
+
+	hash = hugetlb_fault_mutex_hash(mapping, index);
+	mutex_lock(&hugetlb_fault_mutex_table[hash]);
+
+	xa_val = filemap_get_entry(mapping, index << huge_page_order(h));
+	if (!xa_is_value(xa_val)) {
+		mutex_unlock(&hugetlb_fault_mutex_table[hash]);
+		/*
+		 * NULL means a genuine hole.  A folio means a concurrent
+		 * swapin already completed - filemap_get_entry() took a
+		 * reference we must drop, then ask the caller to re-lookup.
+		 */
+		if (xa_val) {
+			folio_put(xa_val);
+			return ERR_PTR(-EAGAIN);
+		}
+		return ERR_PTR(-ENOENT);
+	}
+
+	entry = radix_to_swp_entry(xa_val);
+	folio = hugetlbfs_do_swapin(mapping, index, entry, NULL, h, 0);
+	mutex_unlock(&hugetlb_fault_mutex_table[hash]);
+
+	if (IS_ERR(folio) && PTR_ERR(folio) == -EEXIST)
+		return ERR_PTR(-EAGAIN);	/* concurrent swapin race */
+	return folio;
+}
+
+/*
+ * Collect the xarray indices of swap anchors of swap area @type in the
+ * hugetlbfs page cache, starting from @start.  Returns the count.
+ */
+static unsigned int hugetlbfs_find_swap_entries(struct address_space *mapping,
+				pgoff_t start, pgoff_t *indices,
+				unsigned int type)
+{
+	XA_STATE(xas, &mapping->i_pages, start);
+	unsigned int nr = 0;
+	swp_entry_t entry;
+	void *xa_val;
+
+	rcu_read_lock();
+	xas_for_each(&xas, xa_val, ULONG_MAX) {
+		if (xas_retry(&xas, xa_val))
+			continue;
+		if (!xa_is_value(xa_val))
+			continue;
+		entry = radix_to_swp_entry(xa_val);
+		if (swp_type(entry) != type)
+			continue;
+		indices[nr] = xas.xa_index;
+		if (++nr == FOLIO_BATCH_SIZE)
+			break;
+	}
+	rcu_read_unlock();
+
+	return nr;
+}
+
+/*
+ * Process one inode for swapoff: swap in all its anchors belonging to
+ * swap area @type.  Similar to shmem_unuse().
+ */
+static int hugetlbfs_unuse_inode(struct inode *inode, unsigned int type)
+{
+	struct address_space *mapping = inode->i_mapping;
+	struct hstate *h = hstate_inode(inode);
+	pgoff_t indices[FOLIO_BATCH_SIZE];
+	pgoff_t start = 0;
+	unsigned int nr, i;
+	int ret = 0;
+
+	for (;;) {
+		nr = hugetlbfs_find_swap_entries(mapping, start, indices,
+						 type);
+		if (!nr)
+			break;
+
+		/* For each swap anchor found, swap it in. */
+		for (i = 0; i < nr; i++) {
+			struct folio *folio;
+
+			/*
+			 * hugetlbfs_swapin_read() takes the fault mutex,
+			 * re-validates that the slot still holds a swap
+			 * anchor and swaps the folio in.  swapoff only
+			 * needs the anchor gone, so drop the folio it
+			 * hands back.  The scan returns raw xarray
+			 * positions; shift back to hugetlb page indices.
+			 */
+			folio = hugetlbfs_swapin_read(inode,
+					indices[i] >> huge_page_order(h));
+			if (IS_ERR(folio)) {
+				ret = PTR_ERR(folio);
+				/*
+				 * -ENOENT: slot already a hole/removed.
+				 * -EAGAIN: raced with a concurrent swapin.
+				 * Either way the anchor is gone - skip it.
+				 */
+				if (ret == -ENOENT || ret == -EAGAIN) {
+					ret = 0;
+					continue;
+				}
+				break;
+			}
+
+			folio_unlock(folio);
+			folio_put(folio);
+		}
+		if (ret < 0)
+			break;
+
+		cond_resched();
+		start = indices[nr - 1] + pages_per_huge_page(h);
+	}
+
+	return ret;
+}
+
+/*
+ * Swap in all hugetlbfs pages from swap area @type.  Called from
+ * try_to_unuse() during swapoff.  Similar to shmem_unuse().
+ */
+int hugetlbfs_unuse(unsigned int type)
+{
+	struct hugetlbfs_inode_info *info, *next;
+	int error = 0;
+
+	if (list_empty(&hugetlbfs_swaplist))
+		return 0;
+
+	mutex_lock(&hugetlbfs_swaplist_mutex);
+start_over:
+	list_for_each_entry_safe(info, next, &hugetlbfs_swaplist, swaplist) {
+		if (!atomic_read(&info->swapped)) {
+			list_del_init(&info->swaplist);
+			continue;
+		}
+
+		/*
+		 * Drop the swaplist mutex while swapping in pages.
+		 * Set stop_eviction to prevent inode eviction.
+		 */
+		atomic_inc(&info->stop_eviction);
+		mutex_unlock(&hugetlbfs_swaplist_mutex);
+
+		error = hugetlbfs_unuse_inode(&info->vfs_inode, type);
+		cond_resched();
+
+		mutex_lock(&hugetlbfs_swaplist_mutex);
+		if (atomic_dec_and_test(&info->stop_eviction))
+			wake_up_var(&info->stop_eviction);
+		if (error)
+			break;
+		if (list_empty(&info->swaplist))
+			goto start_over;
+		next = list_next_entry(info, swaplist);
+		if (!atomic_read(&info->swapped))
+			list_del_init(&info->swaplist);
+	}
+	mutex_unlock(&hugetlbfs_swaplist_mutex);
+	return error;
+}
+
+/*
+ * File-backed swapin for the fault path, called from hugetlb_no_page()
+ * when the page cache lookup came back empty: a swapped-out hugetlbfs
+ * page leaves a swap anchor value entry behind, which the page cache
+ * lookup reports as a hole.
+ *
+ * Returns the swapped-in folio (locked, with a reference) on success, or
+ * an ERR_PTR: -ENOENT when there is no anchor (a genuine hole), -EAGAIN
+ * when the fault should be retried, -EHWPOISON when the swapcache copy
+ * is poisoned, or a propagated allocation/IO error.
+ */
+static struct folio *hugetlb_no_page_swapin(struct address_space *mapping,
+					    struct vm_fault *vmf)
+{
+	struct vm_area_struct *vma = vmf->vma;
+	struct hstate *h = hstate_vma(vma);
+	struct folio *folio;
+	swp_entry_t entry;
+	void *xa_val;
+
+	xa_val = xa_load(&mapping->i_pages,
+			 vmf->pgoff << huge_page_order(h));
+	if (!xa_is_value(xa_val))
+		return ERR_PTR(-ENOENT);
+	entry = radix_to_swp_entry(xa_val);
+
+	/*
+	 * The pte was examined without the page table lock; if it changed
+	 * since, this fault is stale and the swapin would be wasted work.
+	 */
+	if (!hugetlb_pte_stable(h, vma->vm_mm, vmf->address, vmf->pte,
+				vmf->orig_pte))
+		return ERR_PTR(-EAGAIN);
+
+	folio = hugetlbfs_do_swapin(mapping, vmf->pgoff, entry, vma, h,
+				    vmf->address);
+	if (IS_ERR(folio) && PTR_ERR(folio) == -EEXIST)
+		return ERR_PTR(-EAGAIN);	/* concurrent swapin race */
+	return folio;
+}
+
+/*
+ * Remove a swap anchor at @index from the hugetlbfs page cache and free
+ * the swap slots behind it, together with a cached copy if one is
+ * around.  Returns 1 if the anchor was removed, 0 if the entry at
+ * @index changed under us (the caller should rescan from @index).
+ *
+ * Called with the hugetlb fault mutex held, which serializes against
+ * the fault swapin path.
+ */
+long hugetlbfs_free_swap(struct address_space *mapping, pgoff_t index,
+			 void *radswap, struct hstate *h)
+{
+	XA_STATE_ORDER(xas, &mapping->i_pages, index, huge_page_order(h));
+	struct hugetlbfs_inode_info *info = HUGETLBFS_I(mapping->host);
+	swp_entry_t entry = radix_to_swp_entry(radswap);
+	struct swap_info_struct *si;
+	struct folio *folio;
+	void *old;
+
+	xas_lock_irq(&xas);
+	old = xas_load(&xas);
+	if (old != radswap) {
+		xas_unlock_irq(&xas);
+		return 0;
+	}
+	xas_store(&xas, NULL);
+	xas_unlock_irq(&xas);
+
+	/*
+	 * Evict a cached copy first, while the anchor's count reference
+	 * still pins the slots: this cannot race with a reallocation of
+	 * the slots, which a put-then-lookup order would allow.  A
+	 * hwpoisoned folio is safe to free here: the hugetlb free path
+	 * moves the poison marker to the raw error pages.
+	 *
+	 * Only evict when the folio is referenced by nothing but the
+	 * swap cache (and our lookup).  An extra reference means a
+	 * concurrent reclaim pass has the folio isolated; in that case
+	 * leave it cached, and keep the folio lock held across
+	 * swap_put_entries_direct() so that its __try_to_reclaim_swap()
+	 * cannot evict the folio either -- reclaim's own free path then
+	 * removes it (and frees the count-0 slots) once it resumes.
+	 */
+	si = get_swap_device(entry);
+	if (si) {
+		folio = swap_cache_get_folio(entry);
+		if (folio) {
+			folio_lock(folio);
+			folio_wait_writeback(folio);
+			if (unlikely(!folio_matches_swap_entry(folio, entry))) {
+				folio_unlock(folio);
+				folio_put(folio);
+				folio = NULL;
+			}
+		}
+		if (folio && folio_ref_count(folio) == folio_nr_pages(folio) + 1)
+			swap_cache_del_folio(folio);
+		/* Drop the anchor's swap count reference, freeing the slots. */
+		swap_put_entries_direct(entry, pages_per_huge_page(h));
+		if (folio) {
+			folio_unlock(folio);
+			folio_put(folio);
+		}
+		put_swap_device(si);
+	}
+	/*
+	 * Else the swap area is already torn down by swapoff, which has
+	 * freed the slots; there is nothing to put.
+	 */
+
+	mutex_lock(&hugetlbfs_swaplist_mutex);
+	if (atomic_sub_return(pages_per_huge_page(h), &info->swapped) == 0)
+		list_del_init(&info->swaplist);
+	mutex_unlock(&hugetlbfs_swaplist_mutex);
+
+	return 1;
+}
+
+/*
+ * Drain this inode from the swapoff traversal list and wait for any
+ * concurrent hugetlbfs_unuse() scan of it to finish.  Called from
+ * hugetlbfs_evict_inode() after remove_inode_hugepages() has removed
+ * all pages and swap anchors, so the swapped count is already 0.
+ */
+void hugetlbfs_evict_drain_swaplist(struct inode *inode)
+{
+	struct hugetlbfs_inode_info *info = HUGETLBFS_I(inode);
+
+	/*
+	 * Check list_empty under the mutex to prevent a concurrent
+	 * writeout from re-adding the inode between the check and the
+	 * wait.  Beware of the race if we peeked too early.
+	 */
+	mutex_lock(&hugetlbfs_swaplist_mutex);
+	while (!list_empty(&info->swaplist)) {
+		mutex_unlock(&hugetlbfs_swaplist_mutex);
+		wait_var_event(&info->stop_eviction,
+			       !atomic_read(&info->stop_eviction));
+		mutex_lock(&hugetlbfs_swaplist_mutex);
+		if (!atomic_read(&info->stop_eviction))
+			list_del_init(&info->swaplist);
+	}
+	mutex_unlock(&hugetlbfs_swaplist_mutex);
+}
+
 /*
  * Reclaim a list of isolated hugetlb folios by swapping them out.
  *
  * This is a deliberate mirror of shrink_folio_list(), not a reuse:
  * hugetlb folios live on hstate activelists instead of LRUs, are
- * unmapped through hugetlb rmap walks (whole-folio swap entries), and
- * go back to the hstate pool on free instead of the buddy allocator.
- * None of those hooks exist in the generic reclaim path, and shrinking
- * that path's folio-size assumptions to fit hugetlb would complicate
- * both sides, so the hugetlb-specific steps are reimplemented here next
- * to the machinery they depend on (demotion and pool management live in
- * this file for the same reason).
- * Caller: proactive reclaim via MADV_PAGEOUT.
+ * unmapped through hugetlb rmap walks (whole-folio swap entries), are
+ * written out via hugetlbfs_writeout()/the anon swap-PTE install path,
+ * and go back to the hstate pool on free instead of the buddy
+ * allocator.  None of those hooks exist in the generic reclaim path,
+ * and shrinking that path's folio-size assumptions to fit hugetlb
+ * would complicate both sides, so the hugetlb-specific steps are
+ * reimplemented here next to the machinery they depend on (demotion
+ * and pool management live in this file for the same reason).
+ * Caller: proactive reclaim via MADV_PAGEOUT; the kernel never reclaims
+ * hugetlb folios on its own.
  */
 static unsigned int hugetlb_reclaim_folio_list(struct list_head *folio_list)
 {
@@ -7612,14 +8284,6 @@ static unsigned int hugetlb_reclaim_folio_list(struct list_head *folio_list)
 		if (unlikely(!folio_evictable(folio)))
 			goto activate_locked;
 
-		/*
-		 * Only anonymous folios (MAP_PRIVATE mappings after COW)
-		 * are swapped out so far; file-backed hugetlbfs folios
-		 * gain swap support later in this series.
-		 */
-		if (!folio_test_anon(folio))
-			goto keep_locked;
-
 		/*
 		 * If the folio's swap write is still in flight, the bio holds a
 		 * writeback reference that end_swap_bio_write() drops
@@ -7640,21 +8304,28 @@ static unsigned int hugetlb_reclaim_folio_list(struct list_head *folio_list)
 		}
 
 		/*
-		 * Anonymous folios enter the swap cache here, before the
-		 * unmap below installs swap PTEs that reference the slots.
-		 * Hugetlb folios are not swap backed by default.
+		 * Both anonymous folios (MAP_PRIVATE mappings after COW) and
+		 * file-backed hugetlbfs folios can be swapped out.  Anonymous
+		 * folios enter the swap cache here, before unmap installs
+		 * swap PTEs that reference the slots.  File folios enter it
+		 * later, inside hugetlbfs_writeout(): their unmap only
+		 * clears the PTEs and installs nothing, so no slots are
+		 * needed yet.
 		 */
-		folio_set_swapbacked(folio);
-		if (!folio_test_swapcache(folio)) {
-			if (folio_alloc_swap(folio))
-				goto activate_locked;
+		if (folio_test_anon(folio)) {
+			/* Hugetlb folios are not swap backed by default. */
+			folio_set_swapbacked(folio);
+			if (!folio_test_swapcache(folio)) {
+				if (folio_alloc_swap(folio))
+					goto activate_locked;
 
-			folio_mark_dirty(folio);
+				folio_mark_dirty(folio);
+			}
 		}
 
 		/*
-		 * Unmap from every process, installing swap PTEs that
-		 * reference the slots allocated above.
+		 * Unmap from every process: installs swap PTEs for anonymous
+		 * folios, just clears the PTEs for file folios.
 		 */
 		if (folio_mapped(folio)) {
 			try_to_unmap_swap_hugetlb(folio);
@@ -7685,7 +8356,10 @@ static unsigned int hugetlb_reclaim_folio_list(struct list_head *folio_list)
 
 			if (folio_clear_dirty_for_io(folio)) {
 				folio_set_reclaim(folio);
-				res = swap_writeout(&ctx, folio);
+				if (folio_test_anon(folio))
+					res = swap_writeout(&ctx, folio);
+				else
+					res = hugetlbfs_writeout(&ctx, folio);
 				if (res < 0) {
 					folio_lock(folio);
 					if (folio_mapping(folio) == mapping)
@@ -7780,7 +8454,9 @@ static unsigned int hugetlb_reclaim_folio_list(struct list_head *folio_list)
 activate_locked:
 		/*
 		 * Not swapped out: drop any swap slots we reserved.  Only
-		 * anonymous folios can hold reserved-but-unused slots here.
+		 * anonymous folios can hold reserved-but-unused slots here;
+		 * a file folio in the swap cache is already anchored in the
+		 * page cache and must keep its slots for the anchor.
 		 */
 		if (folio_test_anon(folio) && folio_test_swapcache(folio) &&
 		    (folio_test_mlocked(folio) || mem_cgroup_swap_full(folio)))
@@ -7867,6 +8543,7 @@ unsigned long hugetlb_reclaim_pages(struct list_head *folio_list)
 
 	return nr_reclaimed;
 }
+
 /*
  * Mirror of should_try_to_free_swap() in mm/memory.c for the hugetlb
  * swapin path; keep in sync.  The exclusive test differs: hugetlb
diff --git a/mm/internal.h b/mm/internal.h
index e16f1250b25c..523472225a8b 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -605,8 +605,6 @@ static inline void force_page_cache_readahead(struct address_space *mapping,
 
 unsigned find_lock_entries(struct address_space *mapping, pgoff_t *start,
 		pgoff_t end, struct folio_batch *fbatch, pgoff_t *indices);
-unsigned find_get_entries(struct address_space *mapping, pgoff_t *start,
-		pgoff_t end, struct folio_batch *fbatch, pgoff_t *indices);
 int truncate_inode_folio(struct address_space *mapping, struct folio *folio);
 bool truncate_inode_partial_folio(struct folio *folio, loff_t start,
 		loff_t end);
@@ -1469,8 +1467,23 @@ static inline void shrinker_debugfs_remove(struct dentry *debugfs_entry,
 /* Only track the nodes of mappings with shadow entries */
 void workingset_update_node(struct xa_node *node);
 extern struct list_lru shadow_nodes;
+
+#ifdef CONFIG_HUGETLBFS
+static inline bool hugetlbfs_mapping(struct address_space *mapping)
+{
+	return mapping->host &&
+	       mapping->host->i_sb->s_magic == HUGETLBFS_MAGIC;
+}
+#else
+static inline bool hugetlbfs_mapping(struct address_space *mapping)
+{
+	return false;
+}
+#endif /* CONFIG_HUGETLBFS */
+
 #define mapping_set_update(xas, mapping) do {			\
-	if (!dax_mapping(mapping) && !shmem_mapping(mapping)) {	\
+	if (!dax_mapping(mapping) && !shmem_mapping(mapping) &&	\
+	    !hugetlbfs_mapping(mapping)) {			\
 		xas_set_update(xas, workingset_update_node);	\
 		xas_set_lru(xas, &shadow_nodes);		\
 	}							\
diff --git a/mm/madvise.c b/mm/madvise.c
index a46306cd1d49..5f7584e42c37 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -636,9 +636,10 @@ static void madvise_pageout_page_range(struct mmu_gather *tlb,
 }
 
 /*
- * Page out private anonymous hugetlb folios in the range: isolate each
- * present hugepage, hand the batch to hugetlb_reclaim_pages() and let it
- * unmap, write out and free the folios synchronously.
+ * Page out hugetlb folios in the range: isolate each present hugepage,
+ * hand the batch to hugetlb_reclaim_pages() and let it unmap, write out
+ * and free the folios synchronously.  Both private (anonymous) folios
+ * and shared (hugetlbfs page-cache) folios are supported.
  */
 static long madvise_pageout_hugetlb(struct madvise_behavior *madv_behavior)
 {
@@ -654,13 +655,6 @@ static long madvise_pageout_hugetlb(struct madvise_behavior *madv_behavior)
 		return -EINVAL;
 	if (vma->vm_flags & (VM_LOCKED | VM_PFNMAP))
 		return -EINVAL;
-	/*
-	 * Swap PTEs are only installed for anonymous folios (MAP_PRIVATE
-	 * after COW); shared hugetlbfs mappings gain swap-out support
-	 * later in this series.
-	 */
-	if (vma->vm_flags & VM_MAYSHARE)
-		return -EINVAL;
 	/* A huge page must fit in a single swap cluster to be swappable. */
 	if (hstate_is_gigantic(h) || huge_page_order(h) > HPAGE_PMD_ORDER)
 		return -EINVAL;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 280a31c43c81..97e744aa6916 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -2763,6 +2763,13 @@ static int try_to_unuse(unsigned int type)
 	if (retval)
 		return retval;
 
+#ifdef CONFIG_HUGETLB_PAGE
+	/* Swap in any swapped-out hugetlbfs file pages (swap anchors). */
+	retval = hugetlbfs_unuse(type);
+	if (retval)
+		return retval;
+#endif
+
 	prev_mm = &init_mm;
 	mmget(prev_mm);
 
-- 
2.53.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 6/6] selftests/mm: add hugetlb_swap test, document hugetlb swap
       [not found] <cover.1790663399.git.leizongkun@qq.com>
                   ` (4 preceding siblings ...)
  2026-09-29  8:01 ` [RFC PATCH 5/6] mm/hugetlb: swap support for file-backed " Zongkun Lei
@ 2026-09-29  8:02 ` Zongkun Lei
  5 siblings, 0 replies; 6+ messages in thread
From: Zongkun Lei @ 2026-09-29  8:02 UTC (permalink / raw)
  To: linux-mm
  Cc: Zongkun Lei, Muchun Song, Oscar Salvador, David Hildenbrand,
	Andrew Morton, Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka,
	Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Jonathan Corbet,
	Shuah Khan, Randy Dunlap, linux-doc, linux-kernel,
	linux-kselftest

Add tools/testing/selftests/mm/hugetlb_swap.c covering the hugetlb
swap feature end to end (37 assertions, 14 scenarios):

- anonymous pageout: PTE becomes a swap entry, the huge page returns
  to the pool, VmSwap grows by the huge page size
- fork sharing swap entries: duplicated slot references, MM_SWAPENTS
  charged to the child, COW on write, parent data unpolluted
- munmap/exit releasing swap slots; zap-time drain of the leftover
  swapcache copy after a read fault
- swapoff bringing swapped-out hugetlb pages back in
- pool exhaustion at swapin: SIGBUS after a bounded retry, matching
  the reservation contract
- hwpoison of a swapcache-resident folio: the folio stays in the swap
  cache and the accessor is killed with SIGBUS at swapin
- shared file-backed mappings: pageout clears the PTE without
  installing a swap entry, refault and read(2) both swap back in with
  data intact, repeated pageout/swapin cycles, truncate freeing the
  anchor and its slots (SIGBUS on subsequent access)
- reject paths: VM_LOCKED, userfaultfd-registered and gigantic VMAs
  fail MADV_PAGEOUT with EINVAL
- mprotect/mremap smoke over swap entries
- VmSwap accounting: +N at pageout, +N at fork, -N at swapin and zap
  (regression test for the unsigned-negation zap fix)
- soft-dirty and uffd-wp bit propagation through swapout/swapin

Verified on x86-64 (2M hugepages, dedicated swap device): 35 pass /
0 fail / 2 expected skips (mlock is a no-op on hugetlb; no free 1G
gigantic pages).

Document the feature in hugetlbpage.rst: the userspace-driven
PAGEOUT contract, the size and VMA restrictions, and the
pool-exhaustion SIGBUS semantics a manager process must plan around.

Signed-off-by: Zongkun Lei <leizongkun@qq.com>
---
 Documentation/admin-guide/mm/hugetlbpage.rst |   29 +-
 tools/testing/selftests/mm/Makefile          |    1 +
 tools/testing/selftests/mm/hugetlb_swap.c    | 1118 ++++++++++++++++++
 tools/testing/selftests/mm/run_vmtests.sh    |    2 +
 4 files changed, 1149 insertions(+), 1 deletion(-)
 create mode 100644 tools/testing/selftests/mm/hugetlb_swap.c

diff --git a/Documentation/admin-guide/mm/hugetlbpage.rst b/Documentation/admin-guide/mm/hugetlbpage.rst
index 3cc15d800be1..75b78dea5a2e 100644
--- a/Documentation/admin-guide/mm/hugetlbpage.rst
+++ b/Documentation/admin-guide/mm/hugetlbpage.rst
@@ -88,7 +88,8 @@ the user when the system is under memory pressure.  Please try again later.
 
 Pages that are used as huge pages are reserved inside the kernel and cannot
 be used for other purposes.  Huge pages cannot be swapped out under
-memory pressure.
+memory pressure; explicit userspace-driven swap-out is described in
+Swapping_ below.
 
 Once a number of huge pages have been pre-allocated to the kernel huge page
 pool, a user with appropriate privilege can use either the mmap system call
@@ -469,6 +470,32 @@ errno set to EINVAL or exclude hugetlb pages that extend beyond the length if
 not hugepage aligned.  For example, munmap(2) will fail if memory is backed by
 a hugetlb page and the length is smaller than the hugepage size.
 
+Swapping
+--------
+
+Hugetlb mappings can be swapped out on explicit userspace request via
+madvise(2) MADV_PAGEOUT / process_madvise(2).  This covers both private
+anonymous mappings (MAP_PRIVATE hugetlbfs files and MFD_HUGETLB memfds
+after COW) and shared file-backed mappings; for shared mappings the
+folios are written to swap while the PTEs are simply cleared, mirroring
+the shmem model.  The kernel never swaps hugetlb pages on its own:
+cold-page scoring and swap decisions belong to userspace.
+
+Only huge pages that fit in a single swap cluster (i.e. PMD-sized, 2M on
+x86-64 and arm64) are swappable; MADV_PAGEOUT on larger hstates,
+userfaultfd-registered or locked VMAs fails with EINVAL.  Swapped-out
+folios are swapped back in on access (for shared mappings, read(2)
+works as well).  MADV_WILLNEED has no effect on hugetlb mappings:
+hugetlbfs has no readahead support, so the advice is silently ignored
+as it has always been.
+
+Swapping out releases the huge page back to the persistent pool, so a
+later swap-in competes for pool memory.  If the pool is exhausted at
+fault time, the faulting task receives SIGBUS after a bounded retry,
+matching the reservation-exhaustion contract.  A manager process must
+therefore create headroom (page out cold pages) before touching swapped
+pages.
+
 
 Examples
 ========
diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/mm/Makefile
index 2d5366196e30..b74047ede66f 100644
--- a/tools/testing/selftests/mm/Makefile
+++ b/tools/testing/selftests/mm/Makefile
@@ -66,6 +66,7 @@ TEST_GEN_FILES += hugetlb-mremap
 TEST_GEN_FILES += hugetlb-read-hwpoison
 TEST_GEN_FILES += hugetlb-shm
 TEST_GEN_FILES += hugetlb-soft-offline
+TEST_GEN_FILES += hugetlb_swap
 TEST_GEN_FILES += khugepaged
 TEST_GEN_FILES += madv_populate
 TEST_GEN_FILES += map_fixed_noreplace
diff --git a/tools/testing/selftests/mm/hugetlb_swap.c b/tools/testing/selftests/mm/hugetlb_swap.c
new file mode 100644
index 000000000000..d8937735b44c
--- /dev/null
+++ b/tools/testing/selftests/mm/hugetlb_swap.c
@@ -0,0 +1,1118 @@
+// SPDX-License-Identifier: GPL-2.0
+#define _GNU_SOURCE
+#include <fcntl.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <stdint.h>
+#include <unistd.h>
+#include <sys/mman.h>
+#include <sys/syscall.h>
+#include <sys/wait.h>
+#include <sys/ioctl.h>
+#include <signal.h>
+#include <setjmp.h>
+#include <poll.h>
+#include <pthread.h>
+#include <linux/memfd.h>
+#include <linux/userfaultfd.h>
+#include "../kselftest.h"
+
+#ifndef MADV_PAGEOUT
+#define MADV_PAGEOUT 21
+#endif
+#ifndef MADV_HWPOISON
+#define MADV_HWPOISON 100
+#endif
+#ifndef MFD_HUGETLB
+#define MFD_HUGETLB 0x0004U
+#endif
+
+#define HUGEPAGE_SIZE (2UL * 1024 * 1024)
+#define PM_PRESENT  (1ULL << 63)
+#define PM_SWAPPED  (1ULL << 62)
+#define PM_UFFD_WP  (1ULL << 57)
+#define PM_SOFT_DIRTY (1ULL << 55)
+
+static int alloc_private_hugetlb(char **out, size_t len)
+{
+	int fd = syscall(SYS_memfd_create, "hugetlb_swap_test", MFD_HUGETLB);
+	char *p;
+
+	if (fd < 0)
+		ksft_exit_fail_msg("memfd_create(MFD_HUGETLB): %m\n");
+	if (ftruncate(fd, len))
+		ksft_exit_fail_msg("ftruncate: %m\n");
+	p = mmap(NULL, len, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);
+	if (p == MAP_FAILED)
+		ksft_exit_fail_msg("mmap private hugetlb: %m\n");
+	*out = p;
+	return fd;
+}
+
+static unsigned long pagemap_flags(void *addr)
+{
+	uint64_t v = 0;
+	int fd = open("/proc/self/pagemap", O_RDONLY);
+	off_t off = ((uintptr_t)addr / 4096) * sizeof(v);
+
+	if (fd < 0 || pread(fd, &v, sizeof(v), off) != sizeof(v))
+		ksft_exit_fail_msg("read pagemap: %m\n");
+	close(fd);
+	return v;
+}
+
+static void fill_pattern(char *p, size_t len)
+{
+	for (size_t i = 0; i < len; i += 4096)
+		p[i] = (char)(i >> 12) ^ 0x5a;
+}
+
+static int verify_pattern(char *p, size_t len)
+{
+	for (size_t i = 0; i < len; i += 4096)
+		if (p[i] != (char)((i >> 12) ^ 0x5a))
+			return -1;
+	return 0;
+}
+
+static long free_hugepages(void)
+{
+	FILE *f = fopen("/proc/meminfo", "r");
+	char line[256];
+	long v = -1;
+
+	while (f && fgets(line, sizeof(line), f))
+		if (sscanf(line, "HugePages_Free: %ld", &v) == 1)
+			break;
+	if (f)
+		fclose(f);
+	return v;
+}
+
+static long read_meminfo_kb(const char *key)
+{
+	FILE *f = fopen("/proc/meminfo", "r");
+	char line[256];
+	size_t klen = strlen(key);
+	long v = -1;
+
+	while (f && fgets(line, sizeof(line), f))
+		if (!strncmp(line, key, klen))
+			sscanf(line + klen, " %ld", &v);
+	if (f)
+		fclose(f);
+	return v;
+}
+
+/*
+ * mlock is a silent no-op on hugetlb mappings (vma_supports_mlock()
+ * excludes is_vm_hugetlb_page), so VM_LOCKED is never set; use the
+ * "lo" flag in smaps VmFlags to tell whether the vma is really locked.
+ */
+static int vma_has_locked_flag(void *addr)
+{
+	FILE *f = fopen("/proc/self/smaps", "r");
+	char line[512];
+	int in_range = 0, locked = 0;
+
+	while (f && fgets(line, sizeof(line), f)) {
+		unsigned long lo, hi;
+
+		if (sscanf(line, "%lx-%lx", &lo, &hi) == 2)
+			in_range = lo <= (unsigned long)addr && (unsigned long)addr < hi;
+		if (in_range && !strncmp(line, "VmFlags:", 8)) {
+			locked = !!strstr(line, " lo");
+			break;
+		}
+	}
+	if (f)
+		fclose(f);
+	return locked;
+}
+
+/* Prerequisite: >= 'need' free 2M hugepages, otherwise SKIP */
+static void require_pool(long need)
+{
+	long free_hp = free_hugepages();
+
+	if (free_hp < need)
+		ksft_exit_skip("need %ld free hugepages, have %ld\n", need, free_hp);
+}
+
+static void test_pageout_basic(void)
+{
+	char *p;
+	long before, after;
+	int fd;
+
+	require_pool(2);
+	before = free_hugepages();
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("MADV_PAGEOUT: %m\n");
+	else
+		ksft_test_result_pass("MADV_PAGEOUT succeeds\n");
+
+	after = free_hugepages();
+	/* allow slack: pageout should return 1 page */
+	ksft_test_result(after >= before - 1 &&
+					 (pagemap_flags(p) & PM_SWAPPED) &&
+					 !(pagemap_flags(p) & PM_PRESENT),
+					 "pageout returns memory and leaves swap entry\n");
+
+	if (verify_pattern(p, HUGEPAGE_SIZE))     /* faults back in, checks data */
+		ksft_test_result_fail("data corrupt after swapin\n");
+	else
+		ksft_test_result_pass("swapin on fault preserves data\n");
+
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+}
+
+static void test_fork_after_pageout(void)
+{
+	char *p;
+	int fd, status;
+	pid_t pid;
+
+	require_pool(3);
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+
+	errno = 0;
+	pid = fork();
+	if (pid < 0) {
+		ksft_test_result_fail("fork#1: %m\n");
+		goto out;
+	}
+	if (pid == 0) {
+		/*
+		 * The child swaps in and verifies the data, then writes a new
+		 * pattern: a shared swap entry is non-exclusive, so the write
+		 * must COW and must not pollute the parent's not-yet-swapped-in
+		 * data.
+		 */
+		if (verify_pattern(p, HUGEPAGE_SIZE))
+			_exit(1);
+		for (size_t i = 0; i < HUGEPAGE_SIZE; i += 4096)
+			p[i] = (char)0xa5;
+		_exit(0);
+	}
+	waitpid(pid, &status, 0);
+	if (!WIFEXITED(status) || WEXITSTATUS(status)) {
+		ksft_test_result_fail("child read wrong data after fork (exit=%d code=%d sig=%d)\n",
+						  WIFEXITED(status),
+						  WIFEXITED(status) ? WEXITSTATUS(status) : -1,
+						  WIFSIGNALED(status) ? WTERMSIG(status) : -1);
+		goto out;
+	}
+	/* Parent swaps in again: the child write must have COWed, data intact */
+	ksft_test_result(verify_pattern(p, HUGEPAGE_SIZE) == 0,
+					 "fork shares swap entry correctly\n");
+
+	/* Child writes (COW after swapin); parent data must not be polluted */
+	pid = fork();
+	if (pid == 0) {
+		for (size_t i = 0; i < HUGEPAGE_SIZE; i += 4096)
+			p[i] = (char)0xa5;
+		_exit(0);
+	}
+	waitpid(pid, &status, 0);
+	ksft_test_result(WIFEXITED(status) && WEXITSTATUS(status) == 0 &&
+					 verify_pattern(p, HUGEPAGE_SIZE) == 0,
+					 "child COW write does not corrupt parent data\n");
+out:
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+}
+
+static void test_munmap_releases_swap(void)
+{
+	char *p;
+	long swap_before, swap_after;
+	int fd, status;
+	pid_t pid;
+
+	require_pool(3);
+	swap_before = read_meminfo_kb("SwapFree:");
+	if (swap_before < (long)(HUGEPAGE_SIZE / 1024))
+		ksft_test_result_skip("no swap space\n");
+
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+
+	/*
+	 * Fork a child which drops the mapping without touching the
+	 * page.  Both the parent's and the child's swap PTE reference
+	 * (the fork duplicated the slot reference) must be returned
+	 * to the swap device; a leak leaves the slots allocated.
+	 */
+	errno = 0;
+	pid = fork();
+	if (pid < 0)
+		ksft_test_result_fail("munmap-test fork: %m\n");
+	if (pid == 0) {
+		munmap(p, HUGEPAGE_SIZE);
+		close(fd);
+		_exit(0);
+	}
+	waitpid(pid, &status, 0);
+
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+
+	swap_after = read_meminfo_kb("SwapFree:");
+	ksft_test_result(WIFEXITED(status) && WEXITSTATUS(status) == 0 &&
+					 swap_after >= swap_before - 64,
+					 "munmap/exit releases swapped-out huge page slots (SwapFree %ld->%ld kB)\n",
+					 swap_before, swap_after);
+}
+
+static void test_readfault_munmap_drains_swapcache(void)
+{
+	char *p;
+	long swap_before, swap_after;
+	long free_before, free_after;
+	int fd;
+
+	require_pool(2);
+	swap_before = read_meminfo_kb("SwapFree:");
+	if (swap_before < (long)(HUGEPAGE_SIZE / 1024))
+		ksft_test_result_skip("no swap space\n");
+
+	free_before = free_hugepages();
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+
+	/* Read fault: swaps the page in but may keep the swapcache copy. */
+	if (verify_pattern(p, HUGEPAGE_SIZE))
+		ksft_test_result_fail("read-fault swapin data corrupt\n");
+
+	/*
+	 * Dropping the last mapping must drain the swapcache copy and free
+	 * the slots: hugetlb folios sit on no LRU, so without the zap-time
+	 * drain nothing would ever reclaim them (pool page + slots pinned
+	 * until swapoff).
+	 */
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+
+	swap_after = read_meminfo_kb("SwapFree:");
+	free_after = free_hugepages();
+	ksft_test_result(swap_after >= swap_before - 64 &&
+					 free_after >= free_before,
+					 "munmap after read fault drains swapcache (SwapFree %ld->%ld kB, free hugepages %ld->%ld)\n",
+					 swap_before, swap_after, free_before, free_after);
+}
+
+/* Take the first active swap device/file from /proc/swaps */
+static int first_swap_device(char *buf, size_t len)
+{
+	FILE *f = fopen("/proc/swaps", "r");
+	char line[512];
+
+	if (!f)
+		return -1;
+	if (!fgets(line, sizeof(line), f) ||      /* header */
+		!fgets(line, sizeof(line), f)) {
+		fclose(f);
+		return -1;
+	}
+	fclose(f);
+	return sscanf(line, "%255s", buf) == 1 ? 0 : -1;
+}
+
+static void test_swapoff(void)
+{
+	char swapdev[256], cmd[300];
+	char *p;
+	int fd, present;
+
+	if (geteuid())
+		return ksft_test_result_skip("swapoff test needs root\n");
+	if (first_swap_device(swapdev, sizeof(swapdev)))
+		return ksft_test_result_skip("no active swap device\n");
+
+	require_pool(2);
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+
+	/*
+	 * swapoff must bring the swapped-out hugetlb page back in and
+	 * succeed.
+	 */
+	snprintf(cmd, sizeof(cmd), "swapoff %s", swapdev);
+	ksft_test_result(system(cmd) == 0,
+					 "swapoff %s succeeds with a swapped-out hugetlb page\n",
+					 swapdev);
+
+	present = (pagemap_flags(p) & PM_PRESENT) &&
+			  !(pagemap_flags(p) & PM_SWAPPED) &&
+			  verify_pattern(p, HUGEPAGE_SIZE) == 0;
+	ksft_test_result(present,
+					 "page present and data intact after swapoff\n");
+
+	/* The swap device is shared with the other test cases. */
+	snprintf(cmd, sizeof(cmd), "swapon %s", swapdev);
+	ksft_test_result(system(cmd) == 0,
+					 "swap device re-enabled for the remaining tests\n");
+
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+}
+
+static const char *nr_hugepages_path = "/proc/sys/vm/nr_hugepages";
+
+static long read_nr_hugepages(void)
+{
+	FILE *f = fopen(nr_hugepages_path, "r");
+	long v = -1;
+
+	if (f) {
+		if (fscanf(f, "%ld", &v) != 1)
+			v = -1;
+		fclose(f);
+	}
+	return v;
+}
+
+static int write_nr_hugepages(long v)
+{
+	FILE *f = fopen(nr_hugepages_path, "w");
+
+	if (!f)
+		return -1;
+	if (fprintf(f, "%ld", v) < 0) {
+		fclose(f);
+		return -1;
+	}
+	return fclose(f);
+}
+
+static sigjmp_buf sigbus_env;
+
+static void sigbus_handler(int sig, siginfo_t *si, void *uctx)
+{
+	siglongjmp(sigbus_env, 1);
+}
+
+/*
+ * Pool-exhaustion swapin contract: swap-out releases the pool page; if
+ * the pool is fully occupied when the swapped-out page is touched again,
+ * the fault must SIGBUS after a bounded retry (never swap in
+ * successfully, never OOM-kill, never fail silently).
+ */
+static void test_swapin_pool_exhausted(void)
+{
+	long saved_nr, cur;
+	int status;
+	pid_t pid;
+
+	if (geteuid())
+		return ksft_test_result_skip("pool exhaustion test needs root\n");
+
+	saved_nr = read_nr_hugepages();
+	if (saved_nr < 2 || write_nr_hugepages(2))
+		return ksft_test_result_skip("cannot set nr_hugepages\n");
+	cur = read_nr_hugepages();
+	if (cur != 2 || free_hugepages() < 2)
+		goto skip_restore;
+
+	pid = fork();
+	if (pid == 0) {
+		struct sigaction sa = { .sa_sigaction = sigbus_handler,
+								.sa_flags = SA_SIGINFO };
+		char *a, *b;
+		int fda, fdb;
+
+		/*
+		 * The child sets up its own scenario; any environment
+		 * mismatch exits with _exit(2) and the parent judges SKIP.
+		 */
+		fda = syscall(SYS_memfd_create, "hst_a", MFD_HUGETLB);
+		fdb = syscall(SYS_memfd_create, "hst_b", MFD_HUGETLB);
+		if (fda < 0 || fdb < 0)
+			_exit(2);
+		/* B must fill the whole pool (2 pages) so A's swapin has no page */
+		if (ftruncate(fda, HUGEPAGE_SIZE) ||
+			ftruncate(fdb, 2 * HUGEPAGE_SIZE))
+			_exit(3);
+		a = mmap(NULL, HUGEPAGE_SIZE, PROT_READ | PROT_WRITE,
+				 MAP_PRIVATE, fda, 0);
+		if (a == MAP_FAILED)
+			_exit(4);
+		fill_pattern(a, HUGEPAGE_SIZE);
+		if (madvise(a, HUGEPAGE_SIZE, MADV_PAGEOUT))
+			_exit(5);
+		/* Only after A's pageout frees its reservation can B fill the pool (2 pages) */
+		b = mmap(NULL, 2 * HUGEPAGE_SIZE, PROT_READ | PROT_WRITE,
+				 MAP_PRIVATE, fdb, 0);
+		if (b == MAP_FAILED)
+			_exit(4);
+		fill_pattern(b, 2 * HUGEPAGE_SIZE);     /* fill the pool */
+		if (free_hugepages() != 0)
+			_exit(6);                       /* pool not full, scenario broken */
+
+		sigaction(SIGBUS, &sa, NULL);
+		if (sigsetjmp(sigbus_env, 1) == 0) {
+			volatile char c = a[0];	/* pool exhausted: swapin must SIGBUS */
+			(void)c;
+			_exit(1);                       /* no SIGBUS: contract violated */
+		}
+		_exit(0);                           /* got SIGBUS: contract holds */
+	}
+	waitpid(pid, &status, 0);
+	if (write_nr_hugepages(saved_nr))
+		ksft_print_msg("WARN: failed to restore nr_hugepages=%ld\n", saved_nr);
+
+	if (WIFEXITED(status) && WEXITSTATUS(status) >= 2)
+		return ksft_test_result_skip("cannot set up exhausted pool (code %d)\n",
+									 WEXITSTATUS(status));
+	ksft_test_result(WIFEXITED(status) && WEXITSTATUS(status) == 0,
+					 "swapin with exhausted pool fails with SIGBUS\n");
+	return;
+
+skip_restore:
+	if (write_nr_hugepages(saved_nr))
+		ksft_print_msg("WARN: failed to restore nr_hugepages=%ld\n", saved_nr);
+	ksft_test_result_skip("cannot shrink pool to 1 page\n");
+}
+
+/*
+ * Memory-failure contract: when a hugetlb folio that still sits in the
+ * swap cache after swapin gets poisoned, me_huge_page must keep it in
+ * the swap cache and the swapin path must intercept it via
+ * folio_test_hwpoison() and SIGBUS the accessor; if the folio were
+ * evicted (the pre-fix behavior), the next access would silently swap
+ * in stale disk data and the pool page would leak permanently.
+ * A poisoned page leaves the pool when reclaimed, so this case must
+ * run last.
+ */
+static void test_hwpoison_swapcache(void)
+{
+	int status;
+	pid_t pid;
+
+	if (geteuid())
+		return ksft_test_result_skip("hwpoison test needs root\n");
+
+	require_pool(2);
+	if (read_meminfo_kb("SwapFree:") < (long)(HUGEPAGE_SIZE / 1024))
+		return ksft_test_result_skip("no swap space\n");
+
+	pid = fork();
+	if (pid == 0) {
+		struct sigaction sa = { .sa_sigaction = sigbus_handler,
+								.sa_flags = SA_SIGINFO };
+		char *a;
+		int fda;
+
+		/*
+		 * The child sets up its own scenario; any environment
+		 * mismatch exits with _exit(>=2) and the parent judges SKIP.
+		 */
+		fda = syscall(SYS_memfd_create, "hst_p", MFD_HUGETLB);
+		if (fda < 0 || ftruncate(fda, HUGEPAGE_SIZE))
+			_exit(2);
+		a = mmap(NULL, HUGEPAGE_SIZE, PROT_READ | PROT_WRITE,
+				 MAP_PRIVATE, fda, 0);
+		if (a == MAP_FAILED)
+			_exit(2);
+		fill_pattern(a, HUGEPAGE_SIZE);
+		if (madvise(a, HUGEPAGE_SIZE, MADV_PAGEOUT))
+			_exit(2);
+		/*
+		 * Read fault swaps back in: the folio is mapped again while a
+		 * copy stays in the swap cache (a read fault does not free the
+		 * swap slot).
+		 */
+		if (verify_pattern(a, HUGEPAGE_SIZE))
+			_exit(2);
+		if (!(pagemap_flags(a) & PM_PRESENT))
+			_exit(2);
+		/*
+		 * Inject poison: mf unmaps the page and reinstalls the swap
+		 * entry.  madvise holds a GUP reference for the whole call, so
+		 * me_huge_page's extra-reference check reports MF_FAILED and an
+		 * EBUSY return is expected (the folio is still kept and marked
+		 * poisoned); EINVAL would mean the kernel lacks
+		 * CONFIG_MEMORY_FAILURE.
+		 */
+		if (madvise(a, HUGEPAGE_SIZE, MADV_HWPOISON) && errno == EINVAL)
+			_exit(3);
+		if ((pagemap_flags(a) & (PM_PRESENT | PM_SWAPPED)) != PM_SWAPPED)
+			_exit(4);                   /* mf did not unmap, scenario broken */
+
+		sigaction(SIGBUS, &sa, NULL);
+		if (sigsetjmp(sigbus_env, 1) == 0) {
+			volatile char c = a[0];     /* swapin must intercept poison: SIGBUS */
+			(void)c;
+			_exit(1);			/* no SIGBUS: folio left the swap cache */
+		}
+		_exit(0);                       /* got SIGBUS: contract holds */
+	}
+	waitpid(pid, &status, 0);
+	if (WIFEXITED(status) && WEXITSTATUS(status) == 3)
+		return ksft_test_result_skip("kernel lacks CONFIG_MEMORY_FAILURE\n");
+	if (WIFEXITED(status) && WEXITSTATUS(status) >= 2)
+		return ksft_test_result_skip("cannot set up hwpoison scenario (code %d)\n",
+									 WEXITSTATUS(status));
+	ksft_test_result(WIFEXITED(status) && WEXITSTATUS(status) == 0,
+					 "poisoned swapcache hugetlb folio kills accessor with SIGBUS\n");
+}
+
+static void test_shared_file_swap(void)
+{
+	char *p;
+	long before, after;
+	int fd;
+
+	require_pool(2);
+	before = free_hugepages();
+	fd = syscall(SYS_memfd_create, "hst", MFD_HUGETLB);
+	if (fd < 0 || ftruncate(fd, HUGEPAGE_SIZE))
+		ksft_exit_fail_msg("setup shared memfd: %m\n");
+	p = mmap(NULL, HUGEPAGE_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
+	if (p == MAP_FAILED)
+		ksft_exit_fail_msg("mmap shared hugetlb: %m\n");
+	fill_pattern(p, HUGEPAGE_SIZE);
+
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("shared MADV_PAGEOUT: %m\n");
+	else
+		ksft_test_result_pass("shared MADV_PAGEOUT succeeds\n");
+
+	after = free_hugepages();
+	/* File pageout only clears the PTE: pagemap is neither PRESENT nor SWAPPED */
+	ksft_test_result(after >= before - 1 /* pageout should return 1 page */ &&
+					 !(pagemap_flags(p) & PM_PRESENT) &&
+					 !(pagemap_flags(p) & PM_SWAPPED),
+					 "shared pageout returns memory, PTE cleared (no swap entry)\n");
+
+	/* read(2) path: no fault, swap in via the anchor and check data */
+	{
+		char *buf = malloc(HUGEPAGE_SIZE);
+		ssize_t n = pread(fd, buf, HUGEPAGE_SIZE, 0);
+
+		ksft_test_result(n == (ssize_t)HUGEPAGE_SIZE &&
+						 verify_pattern(buf, HUGEPAGE_SIZE) == 0,
+						 "shared read(2) swaps in and preserves data\n");
+		free(buf);
+	}
+
+	/* Page out again after swapin: the anchor must be re-creatable */
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("second shared MADV_PAGEOUT: %m\n");
+	else
+		ksft_test_result_pass("second shared MADV_PAGEOUT succeeds\n");
+	ksft_test_result(verify_pattern(p, HUGEPAGE_SIZE) == 0,
+					 "second shared swapin on fault preserves data\n");
+
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+}
+
+static void test_truncate_after_pageout(void)
+{
+	char *p;
+	long swap_before, swap_after;
+	int fd;
+
+	require_pool(2);
+	swap_before = read_meminfo_kb("SwapFree:");
+	if (swap_before < (long)(HUGEPAGE_SIZE / 1024))
+		ksft_test_result_skip("no swap space\n");
+
+	fd = syscall(SYS_memfd_create, "hst", MFD_HUGETLB);
+	if (fd < 0 || ftruncate(fd, HUGEPAGE_SIZE))
+		ksft_exit_fail_msg("setup shared memfd: %m\n");
+	p = mmap(NULL, HUGEPAGE_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
+	if (p == MAP_FAILED)
+		ksft_exit_fail_msg("mmap shared hugetlb: %m\n");
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+
+	/* Truncate frees the anchor and the swap slots: SwapFree back to baseline */
+	if (ftruncate(fd, 0))
+		ksft_test_result_fail("truncate: %m\n");
+	swap_after = read_meminfo_kb("SwapFree:");
+	ksft_test_result(swap_after >= swap_before - 64,
+					 "truncate frees swap anchor and slots (SwapFree %ld->%ld kB)\n",
+					 swap_before, swap_after);
+
+	/* i_size is now 0: touching the mapping must SIGBUS */
+	{
+		struct sigaction sa = { .sa_sigaction = sigbus_handler,
+								.sa_flags = SA_SIGINFO };
+		struct sigaction old_sa;
+
+		sigaction(SIGBUS, &sa, &old_sa);
+		if (sigsetjmp(sigbus_env, 1)) {
+			ksft_test_result_pass("access after truncate raises SIGBUS\n");
+		} else {
+			volatile char c = p[0];
+
+			(void)c;
+			ksft_test_result_fail("access after truncate did not SIGBUS\n");
+		}
+		sigaction(SIGBUS, &old_sa, NULL);
+	}
+
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+}
+
+static void test_reject_paths(void)
+{
+	char *p;
+	int fd;
+
+	require_pool(3);
+
+	/* VM_LOCKED -> EINVAL (SKIP if mlock is a no-op on hugetlb) */
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (mlock(p, HUGEPAGE_SIZE)) {
+		ksft_test_result_skip("mlock unavailable: %m\n");
+	} else if (!vma_has_locked_flag(p)) {
+		munlock(p, HUGEPAGE_SIZE);
+		ksft_test_result_skip("mlock is a no-op on hugetlb mappings\n");
+	} else {
+		errno = 0;
+		ksft_test_result(madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT) == -1 &&
+						 errno == EINVAL,
+						 "locked mapping rejected with EINVAL\n");
+		munlock(p, HUGEPAGE_SIZE);
+	}
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+
+	/* userfaultfd-registered -> EINVAL (SKIP the sub-case if uffd unavailable) */
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	{
+		int uffd = syscall(SYS_userfaultfd, O_CLOEXEC);
+		struct uffdio_api api = { .api = UFFD_API };
+		struct uffdio_register reg;
+		int registered = 0;
+
+		if (uffd >= 0 && !ioctl(uffd, UFFDIO_API, &api)) {
+			memset(&reg, 0, sizeof(reg));
+			reg.range.start = (unsigned long)p;
+			reg.range.len = HUGEPAGE_SIZE;
+			reg.mode = UFFDIO_REGISTER_MODE_MISSING;
+			registered = !ioctl(uffd, UFFDIO_REGISTER, &reg);
+		}
+		if (!registered) {
+			ksft_test_result_skip("uffd register unsupported\n");
+		} else {
+			errno = 0;
+			ksft_test_result(madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT) == -1 &&
+							 errno == EINVAL,
+							 "uffd-registered mapping rejected with EINVAL\n");
+		}
+		if (uffd >= 0)
+			close(uffd);
+	}
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+
+	/* 1G hstate -> EINVAL (SKIP the sub-case if no free 1G page) */
+	{
+		FILE *f = fopen("/sys/kernel/mm/hugepages/hugepages-1048576kB/free_hugepages", "r");
+		long free1g = 0;
+
+		if (f) {
+			if (fscanf(f, "%ld", &free1g) != 1)
+				free1g = 0;
+			fclose(f);
+		}
+		if (free1g < 1) {
+			ksft_test_result_skip("no free 1G hugepages\n");
+		} else {
+			p = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
+					 MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
+					 (30 << MAP_HUGE_SHIFT), -1, 0);
+			if (p == MAP_FAILED) {
+				ksft_test_result_skip("cannot mmap 1G hugepage: %m\n");
+			} else {
+				p[0] = 1;
+				errno = 0;
+				ksft_test_result(madvise(p, 1UL << 30, MADV_PAGEOUT) == -1 &&
+								 errno == EINVAL,
+								 "gigantic hstate rejected with EINVAL\n");
+				munmap(p, 1UL << 30);
+			}
+		}
+	}
+}
+
+static void test_mprotect_mremap_smoke(void)
+{
+	char *p, *q;
+	int fd;
+
+	require_pool(2);
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+
+	/* mprotect must not corrupt the swap entry encoding */
+	ksft_test_result(!mprotect(p, HUGEPAGE_SIZE, PROT_READ) &&
+					 !mprotect(p, HUGEPAGE_SIZE, PROT_READ | PROT_WRITE) &&
+					 (pagemap_flags(p) & PM_SWAPPED) &&
+					 !(pagemap_flags(p) & PM_PRESENT),
+					 "mprotect preserves swap entry\n");
+
+	q = mremap(p, HUGEPAGE_SIZE, HUGEPAGE_SIZE, MREMAP_MAYMOVE);
+	if (q == MAP_FAILED) {
+		ksft_test_result_fail("mremap: %m\n");
+		munmap(p, HUGEPAGE_SIZE);
+		close(fd);
+		return;
+	}
+	ksft_test_result((pagemap_flags(q) & PM_SWAPPED) &&
+					 verify_pattern(q, HUGEPAGE_SIZE) == 0,
+					 "mremap moves swap entry, data intact\n");
+	munmap(q, HUGEPAGE_SIZE);
+	close(fd);
+}
+
+static long read_vmswap_kb(void)
+{
+	FILE *f = fopen("/proc/self/status", "r");
+	char line[256];
+	long v = -1;
+
+	while (f && fgets(line, sizeof(line), f))
+		if (sscanf(line, "VmSwap: %ld kB", &v) == 1)
+			break;
+	if (f)
+		fclose(f);
+	return v;
+}
+
+/*
+ * ksft_finished() requires plan == number of results, so skip paths
+ * must also fill in their counts; otherwise the whole binary FAILs on
+ * a plan mismatch when the environment is unsuitable.
+ */
+static void skip_results(int n, const char *msg)
+{
+	for (int i = 0; i < n; i++)
+		ksft_test_result_skip("%s", msg);
+}
+
+/*
+ * MM_SWAPENTS accounting contract (readable via VmSwap):
+ * +N at pageout, +N at fork (charged to the child), -N at swapin/zap,
+ * with N = number of small pages per huge page.  A mistake in any step
+ * is a leak or a double count.
+ */
+static void test_vmswap_accounting(void)
+{
+	char *p;
+	long v0, v1;
+	int fd, status;
+	pid_t pid;
+
+	require_pool(2);
+	if (read_meminfo_kb("SwapFree:") < (long)(HUGEPAGE_SIZE / 1024))
+		return skip_results(5, "no swap space\n");
+
+	v0 = read_vmswap_kb();
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+
+	v1 = read_vmswap_kb();
+	ksft_test_result(v1 == v0 + (long)(HUGEPAGE_SIZE / 1024),
+					 "VmSwap +%lu kB at pageout (%ld -> %ld kB)\n",
+					 HUGEPAGE_SIZE / 1024, v0, v1);
+
+	/*
+	 * fork duplicates the slot reference: the child sees the same
+	 * VmSwap; its exit does not change the parent.
+	 */
+	pid = fork();
+	if (pid < 0) {
+		ksft_test_result_fail("vmswap-test fork: %m\n");
+	} else {
+		if (pid == 0)
+			_exit(read_vmswap_kb() == v1 ? 0 : 1);  /* must not touch p */
+		if (waitpid(pid, &status, 0) != pid)
+			ksft_test_result_fail("waitpid: %m\n");
+		else
+			ksft_test_result(WIFEXITED(status) && WEXITSTATUS(status) == 0 &&
+							 read_vmswap_kb() == v1,
+							 "fork duplicates swap count, parent unchanged after child exit\n");
+	}
+
+	/*
+	 * Swapin returns the count: VmSwap back to baseline (the slots may
+	 * survive as swapcache residue, but the count is already returned).
+	 */
+	if (verify_pattern(p, HUGEPAGE_SIZE))
+		ksft_test_result_fail("data corrupt after swapin\n");
+	ksft_test_result(read_vmswap_kb() == v0,
+					 "VmSwap back to baseline after swapin (%ld kB)\n",
+					 read_vmswap_kb());
+
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+
+	/*
+	 * munmap right after pageout (no swapin): zap must return the
+	 * MM_SWAPENTS count.  Regression test: the zap path once
+	 * zero-extended -pages_per_huge_page() (unsigned) to +2^32 and
+	 * blew up VmSwap.
+	 */
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("second pageout: %m\n");
+	ksft_test_result(read_vmswap_kb() == v0 + (long)(HUGEPAGE_SIZE / 1024),
+					 "VmSwap +%lu kB after second pageout\n",
+					 HUGEPAGE_SIZE / 1024);
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+	ksft_test_result(read_vmswap_kb() == v0,
+					 "munmap of swapped-out page drops VmSwap back to baseline (%ld kB)\n",
+					 read_vmswap_kb());
+}
+
+/*
+ * Soft-dirty bit end to end: present PTE -> swap entry (carried by
+ * swp_pte_prepare at pageout) -> back to a present PTE at swapin.
+ *
+ * Note that pagemap reports PM_SOFT_DIRTY for a hugetlb present PTE
+ * from the VMA's VM_SOFTDIRTY flag rather than the PTE bit; only the
+ * swap-entry branch reports the bit in the entry itself.  So use
+ * clear_refs to drop VM_SOFTDIRTY and isolate the entry/PTE bit
+ * (clear_refs' pte walk does not touch hugetlb PTEs, leaving the entry
+ * bit alone).
+ */
+static void test_soft_dirty_swap(void)
+{
+	char *p;
+	int fd, cr;
+
+	require_pool(2);
+	if (read_meminfo_kb("SwapFree:") < (long)(HUGEPAGE_SIZE / 1024))
+		return skip_results(4, "no swap space\n");
+
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+
+	/* A fresh mapping's PTE is soft-dirty (VMA defaults to VM_SOFTDIRTY) */
+	ksft_test_result(pagemap_flags(p) & PM_SOFT_DIRTY,
+					 "fresh hugetlb page is soft-dirty\n");
+
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+	ksft_test_result((pagemap_flags(p) & PM_SWAPPED) &&
+					 (pagemap_flags(p) & PM_SOFT_DIRTY),
+					 "swap entry carries soft-dirty bit\n");
+
+	cr = open("/proc/self/clear_refs", O_WRONLY);
+	if (cr < 0 || write(cr, "4", 1) != 1) {
+		if (cr >= 0)
+			close(cr);
+		skip_results(2, "clear_refs unavailable\n");
+		goto out;
+	}
+	close(cr);
+	/* VM_SOFTDIRTY cleared: PM_SOFT_DIRTY here can only come from the swap entry itself */
+	ksft_test_result((pagemap_flags(p) & PM_SWAPPED) &&
+					 (pagemap_flags(p) & PM_SOFT_DIRTY),
+					 "swap entry soft-dirty bit survives clear_refs\n");
+
+	/*
+	 * Swapin carries the swp bit back to the present PTE (invisible to
+	 * pagemap); the next pageout's swp_pte_prepare reads the PTE bit to
+	 * rebuild the entry -- if swapin dropped the bit, PM_SOFT_DIRTY
+	 * would be absent here.
+	 */
+	if (verify_pattern(p, HUGEPAGE_SIZE))
+		ksft_test_result_fail("data corrupt after swapin\n");
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("second pageout: %m\n");
+	ksft_test_result((pagemap_flags(p) & PM_SWAPPED) &&
+					 (pagemap_flags(p) & PM_SOFT_DIRTY),
+					 "soft-dirty bit survives swapin and second pageout\n");
+
+	verify_pattern(p, HUGEPAGE_SIZE);   /* swap back to present, then clean up */
+out:
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+}
+
+struct wp_monitor_arg {
+	int uffd;
+	unsigned long start;
+	size_t len;
+	int ok;
+};
+
+/*
+ * Monitor: wait for a WP fault event, verify the WP flag, then unprotect
+ * to release the faulting thread.
+ */
+static void *wp_monitor(void *arg)
+{
+	struct wp_monitor_arg *a = arg;
+	struct pollfd pfd = { .fd = a->uffd, .events = POLLIN };
+	struct uffdio_writeprotect wp = {};
+	struct uffd_msg msg;
+	int event_ok = 0, unprotect_ok = 0;
+
+	a->ok = 0;
+	if (poll(&pfd, 1, 5000) <= 0)
+		/* no event: the write was not intercepted, main thread done */
+		return NULL;
+	if (read(a->uffd, &msg, sizeof(msg)) == sizeof(msg) &&
+		msg.event == UFFD_EVENT_PAGEFAULT &&
+		(msg.arg.pagefault.flags & UFFD_PAGEFAULT_FLAG_WP))
+		event_ok = 1;
+	/*
+	 * The event has been consumed and the main thread is blocked in the
+	 * fault waiting for unprotect; the unprotect must happen no matter
+	 * whether the event content matched, otherwise the main thread
+	 * hangs forever.
+	 */
+	wp.range.start = a->start;
+	wp.range.len = a->len;
+	wp.mode = 0;
+	unprotect_ok = !ioctl(a->uffd, UFFDIO_WRITEPROTECT, &wp);
+	a->ok = event_ok && unprotect_ok;
+	return NULL;
+}
+
+/*
+ * uffd-wp vs swap entry:
+ * register the uffd WP mode after pageout (the order cannot be flipped:
+ * an armed VMA is rejected by MADV_PAGEOUT), UFFDIO_WRITEPROTECT must
+ * be able to set the wp bit through the swap entry (the
+ * change_protection contract), swapin carries the wp bit back to the
+ * present PTE, and the following write fault must deliver a WP event
+ * to the monitor.
+ */
+static void test_uffd_wp_swap(void)
+{
+	char *p;
+	int fd, uffd;
+	pthread_t mon;
+	struct wp_monitor_arg arg;
+	struct uffdio_api api = { .api = UFFD_API };
+	struct uffdio_register reg = {};
+	struct uffdio_writeprotect wp = {};
+
+	require_pool(2);
+	if (read_meminfo_kb("SwapFree:") < (long)(HUGEPAGE_SIZE / 1024))
+		return skip_results(4, "no swap space\n");
+
+	fd = alloc_private_hugetlb(&p, HUGEPAGE_SIZE);
+	fill_pattern(p, HUGEPAGE_SIZE);
+	if (madvise(p, HUGEPAGE_SIZE, MADV_PAGEOUT))
+		ksft_test_result_fail("pageout: %m\n");
+
+	uffd = syscall(SYS_userfaultfd, O_CLOEXEC);
+	if (uffd < 0) {
+		skip_results(4, "userfaultfd unavailable\n");
+		goto out;
+	}
+	reg.range.start = (unsigned long)p;
+	reg.range.len = HUGEPAGE_SIZE;
+	reg.mode = UFFDIO_REGISTER_MODE_WP;
+	if (ioctl(uffd, UFFDIO_API, &api) || ioctl(uffd, UFFDIO_REGISTER, &reg)) {
+		skip_results(4, "uffd WP register on hugetlb unsupported\n");
+		goto out_uffd;
+	}
+
+	/*
+	 * Write-protect through the swap entry: only the uffd-wp bit may
+	 * change, the entry encoding must survive.
+	 */
+	wp.range.start = (unsigned long)p;
+	wp.range.len = HUGEPAGE_SIZE;
+	wp.mode = UFFDIO_WRITEPROTECT_MODE_WP;
+	if (ioctl(uffd, UFFDIO_WRITEPROTECT, &wp))
+		ksft_test_result_fail("UFFDIO_WRITEPROTECT: %m\n");
+	ksft_test_result((pagemap_flags(p) & PM_UFFD_WP) &&
+					 (pagemap_flags(p) & PM_SWAPPED) &&
+					 !(pagemap_flags(p) & PM_PRESENT),
+					 "uffd-wp bit set on swap entry\n");
+
+	/* Read fault swapin: the wp bit comes back with the present PTE, data intact */
+	if (verify_pattern(p, HUGEPAGE_SIZE))
+		ksft_test_result_fail("swapin data corrupt under uffd-wp\n");
+	ksft_test_result((pagemap_flags(p) & PM_PRESENT) &&
+					 (pagemap_flags(p) & PM_UFFD_WP),
+					 "swapin carries uffd-wp bit to present PTE\n");
+
+	/*
+	 * The write fault is blocked by the wp bit and raises an event;
+	 * the write completes once the monitor unprotects.
+	 */
+	arg.uffd = uffd;
+	arg.start = (unsigned long)p;
+	arg.len = HUGEPAGE_SIZE;
+	arg.ok = 0;
+	if (pthread_create(&mon, NULL, wp_monitor, &arg)) {
+		ksft_test_result_fail("pthread_create: %m\n");
+		skip_results(1, "no monitor thread\n");
+		goto out_uffd;
+	}
+	p[0] = 0x11;
+	pthread_join(mon, NULL);
+	ksft_test_result(arg.ok && p[0] == 0x11,
+					 "write fault raises uffd-wp event, write completes after unprotect\n");
+	ksft_test_result(!(pagemap_flags(p) & PM_UFFD_WP),
+					 "unprotect clears uffd-wp bit\n");
+
+out_uffd:
+	close(uffd);
+out:
+	munmap(p, HUGEPAGE_SIZE);
+	close(fd);
+}
+
+int main(void)
+{
+	ksft_print_header();
+	/*
+	 * Nearly all cases depend on swap-out; SKIP wholesale without
+	 * swap (run_vmtests friendly).
+	 */
+	if (read_meminfo_kb("SwapTotal:") <= 0)
+		ksft_exit_skip("no swap configured\n");
+	ksft_set_plan(37);
+	test_pageout_basic();
+	test_fork_after_pageout();
+	test_munmap_releases_swap();
+	test_readfault_munmap_drains_swapcache();
+	test_swapoff();
+	test_swapin_pool_exhausted();
+	test_shared_file_swap();
+	test_truncate_after_pageout();
+	test_reject_paths();
+	test_mprotect_mremap_smoke();
+	test_vmswap_accounting();
+	test_soft_dirty_swap();
+	test_uffd_wp_swap();
+	test_hwpoison_swapcache();
+	ksft_finished();
+}
diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/selftests/mm/run_vmtests.sh
index d09f9f6a384e..7e1dfa619b89 100755
--- a/tools/testing/selftests/mm/run_vmtests.sh
+++ b/tools/testing/selftests/mm/run_vmtests.sh
@@ -266,6 +266,8 @@ CATEGORY="hugetlb" run_test ./hugetlb-madvise
 CATEGORY="hugetlb" run_test ./hugetlb_dio
 CATEGORY="hugetlb" run_test ./hugetlb_fault_after_madv
 CATEGORY="hugetlb" run_test ./hugetlb_madv_vs_map
+# needs a swap device and free 2M hugepages; skips cleanly without either
+CATEGORY="hugetlb" run_test ./hugetlb_swap
 
 if test_selected "hugetlb"; then
 	echo "NOTE: These hugetlb tests provide minimal coverage.  Use"	  | tap_prefix
-- 
2.53.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-29  8:03 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
     [not found] <cover.1790663399.git.leizongkun@qq.com>
2026-09-29  7:58 ` [RFC PATCH 1/6] mm/swap: introduce folio_swap_entry() and convert all folio->swap readers Zongkun Lei
2026-09-29  7:59 ` [RFC PATCH 2/6] mm/hugetlb: swap-in support for anonymous hugetlb folios Zongkun Lei
2026-09-29  7:59 ` [RFC PATCH 3/6] mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT Zongkun Lei
2026-09-29  8:00 ` [RFC PATCH 4/6] mm/memory-failure: handle swapcached hugetlb folios Zongkun Lei
2026-09-29  8:01 ` [RFC PATCH 5/6] mm/hugetlb: swap support for file-backed " Zongkun Lei
2026-09-29  8:02 ` [RFC PATCH 6/6] selftests/mm: add hugetlb_swap test, document hugetlb swap Zongkun Lei

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®