mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly
@ 2026-09-21 15:13 Alexandre Ghiti
  2026-09-21 15:13 ` [PATCH v6 1/3] mm: swap: move LRU insertion out of the swap cache allocator Alexandre Ghiti
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Alexandre Ghiti @ 2026-09-21 15:13 UTC (permalink / raw)
  To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Andrew Morton, Chris Li,
	Kairui Song
  Cc: Kairui Song, Chengming Zhou, Matthew Wilcox (Oracle),
	Jan Kara, Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexander Viro, Christian Brauner, David Hildenbrand,
	Lorenzo Stoakes, Michal Hocko, Axel Rasmussen, Qi Zheng,
	Shakeel Butt, Wei Xu, Yuanchu Xie, Kunwu Chan, Tal Zussman,
	linux-mm, linux-kernel, linux-fsdevel, Alexandre Ghiti

When zswap writes an entry back, it allocates an order-0 swap cache folio,
decompresses into it, and issues the write. The folio is cold by
construction, yet today it is left on the LRU for page reclaim to find and
free later. That wastes a reclaim scan and keeps cold memory resident
longer than necessary.

Rather than implement this in zswap, extend the existing dropbehind
mechanism to swap cache folios and have zswap opt into it (Yosry). A
PG_dropbehind folio is already dropped from its cache once writeback
completes instead of being left for reclaim; for a swap cache folio that
"drop" is removing it from the swap cache.

  Patch 1 - move LRU insertion out of the swap cache allocator into its
            callers, so zswap writeback can allocate off the LRU.

  Patch 2 - drop dropbehind swap cache folios on writeback completion.

  Patch 3 - zswap allocates its writeback folio off the LRU and marks it
            dropbehind, opting into the mechanism above.

Note: patch 1 also appears as patch 1 of the zswap writeback refault
series [1]. It is the same patch. Both series need it and both are meant
to apply on their own, so it is posted in each; whichever lands first, the
other should drop it.

This version is based on mm-stable. v4 was based on Linus' tree because
Tal Zussman's BIO_COMPLETE_IN_TASK work had not reached the mm tree yet;
it has since, so that detour is no longer needed. Thanks to Matthew and
Barry for pointing that series out!

v1: https://lore.kernel.org/linux-mm/20260718093723.153324-1-alex@ghiti.fr/
v2: https://lore.kernel.org/linux-mm/20260727143618.1582318-1-alex@ghiti.fr/
v3: https://lore.kernel.org/linux-mm/20260818163221.589352-1-alex@ghiti.fr/
v4: https://lore.kernel.org/linux-mm/20260825135209.3135169-1-alex@ghiti.fr/
v5: https://lore.kernel.org/linux-mm/20260911121341.178028-1-alex@ghiti.fr/

[1] https://lore.kernel.org/linux-mm/20260911092012.92399-1-alex@ghiti.fr/

Changes in v6:
- No functional change: this version only updates changelogs and collects
  tags.
- Patch 1: fix the changelog. The first paragraph describes the behaviour
  before the patch, so it must name swap_cache_alloc_folio(), not the
  __swap_cache_alloc_folio() this patch introduces, and the last paragraph
  now says the helper is renamed as well as that the LRU insertion is
  deferred (Kairui).
- Patch 2: fix the changelog. The swap cluster lock is a spinlock taken
  with interrupts disabled and does not sleep; the folio lock is what the
  drop blocks on, and that is why the completion needs task context. Also
  spell out why it blocks instead of using folio_trylock() (Barry).
- Collect Reviewed-by tags. Thanks to Kairui, Barry and Nhat for the
  reviews! Kairui's tag on patch 1 was given on the posting of the zswap
  writeback refault series [1], which carries the same patch.

Changes in v5:
- Rebase from Linus' tree onto mm-stable, which now carries Christoph's
  swap_io_ctx work. __swap_writepage() takes a swap_io_ctx and the bio is
  built by swap_bdev_submit_write() from a batch of folios, so patch 2 sets
  BIO_COMPLETE_IN_TASK there instead of in the old per-folio async helper.
  The bio can hold several folios, so the flag is set if any of them is
  dropbehind; only the asynchronous branch needs it, as SWP_SYNCHRONOUS_IO
  already completes in task context. Patch 3 follows the new
  __swap_writepage() + swap_write_submit() pair.
- Patch 3: clear PG_active on the writeback folio after allocation.
  __swap_cache_alloc_folio() evaluates a refault, which can set PG_active;
  with the folio kept off the LRU nothing clears it again and the folio is
  freed with a PAGE_FLAGS_CHECK_AT_FREE flag set. This goes away once the
  refault evaluation moves out of the swap cache allocator in [1].
- Patch 2: rename remove_mapping_reclaim() to remove_mapping_set_shadow().
  Storing the workingset eviction shadow is the only thing that sets it
  apart from remove_mapping(), so name it after that rather than after the
  caller that historically did it (Nhat).
- Patch 3: the success path now returns directly, so the "if (ret)" in the
  error path is dead and the label is only reached on failure. Drop the
  check and rename the label from "out" to "err". Thanks Yosry for
  spotting this!
- Patch 1: take the version from the zswap writeback refault series [1],
  which is where it has seen the most review. The changelog now describes
  both users of the change instead of only this one, and it picks up the
  stale swap_cache_alloc_folio() comment fix in mm/swapfile.c (Kunwu). The
  two copies are now identical apart from the base they apply to.
- Collect Reviewed-by tags. Thanks to Kunwu and Nhat for the careful
  reviews, and to Usama for the ack on patch 1!

Changes in v4:
- Rebase on BIO_COMPLETE_IN_TASK: set it on dropbehind swap writeback like
  the file dropbehind paths do, and drop the folio directly from
  folio_end_writeback(). This removes the per-CPU llist, the workqueue and
  the reuse of folio->lru as the list node.
- zswap now drops its folio reference before starting writeback, so the
  swap cache holds the only one and remove_mapping() sees the refcount it
  expects. This fixes the drop on synchronous-IO devices and the race
  Sashiko reported, where the drop could run before zswap released its
  reference and fall back to the LRU. Verified on zram (the only
  SWP_SYNCHRONOUS_IO backend I have): over ~6.7M writebacks per run, 99.999%
  of the folios are dropped, and the refcount fallback fires 37-50 times.
- Use remove_mapping_reclaim() rather than adding a boolean argument to
  remove_mapping(), which keeps the calling code readable (David). This also
  leaves the existing remove_mapping() callers untouched.
- Patch 1: correct the changelog. The folio has to stay off the LRU because
  folio_add_lru() leaves a reference in the per-CPU LRU batch, not because
  of the free-time page-flag checks. Measured on zram, adding the folio to
  the LRU instead drops the freed rate from 99.999% to 2.7%.

Changes in v3:
- Drop the synchronous-IO special case in zswap writeback (Yosry, Nhat).
- Use mem_cgroup_tryget()/mem_cgroup_put(): struct mem_cgroup is only
  defined under CONFIG_MEMCG, so css_tryget()/css_put() failed to build
  with CONFIG_MEMCG=n.

Changes in v2:
- Make swap dropbehind a generic core-mm mechanism that zswap opts into,
  rather than a zswap-specific implementation (Yosry).
- Allocate off the LRU by moving folio_add_lru() out of the swap cache
  allocator into its callers; rename it to __swap_cache_alloc_folio()
  (Kairui).
- Skip the folio in the free path if it is still under writeback (Nhat).

Results
-------
Paired baseline vs series on async swap (NVMe). Each
workload runs confined to a memory cgroup (memory.max) small enough to force
zswap shrinker writeback.

Kernel build (defconfig, make -j4; memory.max = 600M):

  metric            baseline       series      delta
  pgrotated           441028         2521     -99.4%
  pgsteal_direct     3524343      2869004     -18.6%
  pgscan_direct      8393791      7765008      -7.5%
  zswpwb              705129       699019      -0.9%
  build time (s)        1155         1114      -3.6%

Of the 699019 folios written back, 698990 (99.996%) were freed promptly on
writeback completion; only 28 fell back to reclaim.

MySQL/OLTP (sysbench, 10 tables x 1M rows, 512M buffer pool, 8 threads, 300s;
memory.max = 256M):

  metric            baseline       series      delta
  transactions/s      153.87       163.30      +6.1%
  p95 latency (ms)    157.42       145.82      -7.4%
  avg latency (ms)     52.09        49.05      -5.9%
  pgrotated           743738        22460     -97.0%
  pgsteal_direct     6886490      5445278     -20.9%
  pgscan_direct     13462510     10820730     -19.6%

Future work
-----------
Barry suggested extending this to MADV_PAGEOUT and general reclaim. I
prototyped dropbehind for all reclaimed swap folios and it regressed
sysbench OLTP throughput by ~15% on NVMe swap: dropping the swap cache
immediately turns cheap in-cache refaults into disk reads and collapses
swap readahead clustering. Neither blk-wbt, mq-deadline nor a PG_workingset
gate recovered it. MADV_PAGEOUT alone may still be worth it, since there
userspace has explicitly declared the range cold, but I have not measured
that case in isolation yet.

Alexandre Ghiti (3):
  mm: swap: move LRU insertion out of the swap cache allocator
  mm: swap: drop dropbehind swap cache folios on writeback completion
  mm: zswap: drop cold writeback folios via swap dropbehind

 include/linux/swap.h |  6 +++++
 mm/filemap.c         | 19 ++++++++++++++
 mm/page_io.c         |  9 +++++++
 mm/swap.h            |  6 ++---
 mm/swap_state.c      | 62 ++++++++++++++++++++++++++++++++++++++------
 mm/swapfile.c        |  2 +-
 mm/vmscan.c          | 49 +++++++++++++++++++++++++++-------
 mm/zswap.c           | 36 +++++++++++++++++--------
 8 files changed, 156 insertions(+), 33 deletions(-)


base-commit: 0d9ff90a5422cc7509258aaaba1e7481df4d332a
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v6 1/3] mm: swap: move LRU insertion out of the swap cache allocator
  2026-09-21 15:13 [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly Alexandre Ghiti
@ 2026-09-21 15:13 ` Alexandre Ghiti
  2026-09-21 15:13 ` [PATCH v6 2/3] mm: swap: drop dropbehind swap cache folios on writeback completion Alexandre Ghiti
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 5+ messages in thread
From: Alexandre Ghiti @ 2026-09-21 15:13 UTC (permalink / raw)
  To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Andrew Morton, Chris Li,
	Kairui Song
  Cc: Kairui Song, Chengming Zhou, Matthew Wilcox (Oracle),
	Jan Kara, Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexander Viro, Christian Brauner, David Hildenbrand,
	Lorenzo Stoakes, Michal Hocko, Axel Rasmussen, Qi Zheng,
	Shakeel Butt, Wei Xu, Yuanchu Xie, Kunwu Chan, Tal Zussman,
	linux-mm, linux-kernel, linux-fsdevel, Alexandre Ghiti,
	Usama Arif

This is a preparatory patch.

swap_cache_alloc_folio() adds the new folio to the LRU itself, which
leaves its callers no way to act on the folio before it becomes visible
to reclaim.  Two users need exactly that:

 - moving the refault evaluation out of the swap cache folio allocation
   requires it to happen before folio_add_lru(): that consumes PG_active
   to file the folio on the inactive or the active list, and under MGLRU
   it also reads PG_workingset to pick the generation.  Setting either
   flag afterwards does not move the folio;

 - zswap writeback dropbehind needs the buffer folio to stay off the LRU
   entirely, as the per-CPU LRU batch would hold a reference on it and
   keep remove_mapping() from freeing it once writeback completes.

Defer the LRU insertion to the callers and rename the helper to
__swap_cache_alloc_folio(): each caller adds the folio right after the
allocation, so there is no functional change intended.

Suggested-by: Kairui Song <kasong@tencent.com>
Reviewed-by: Kairui Song <kasong@tencent.com>
Reviewed-by: Nhat Pham <nphamcs@gmail.com>
Reviewed-by: Kunwu Chan <kunwu.chan@gmail.com>
Acked-by: Usama Arif <usama.arif@linux.dev>
Reviewed-by: Barry Song <baohua@kernel.org>
Signed-off-by: Alexandre Ghiti <alex@ghiti.fr>
---
 mm/swap.h       |  6 +++---
 mm/swap_state.c | 20 ++++++++++++--------
 mm/swapfile.c   |  2 +-
 mm/zswap.c      |  5 +++--
 4 files changed, 19 insertions(+), 14 deletions(-)

diff --git a/mm/swap.h b/mm/swap.h
index 90a551a88df6..8679cb61268e 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -312,9 +312,9 @@ bool swap_cache_has_folio(swp_entry_t entry);
 struct folio *swap_cache_get_folio(swp_entry_t entry);
 void *swap_cache_get_shadow(swp_entry_t entry);
 void swap_cache_del_folio(struct folio *folio);
-struct folio *swap_cache_alloc_folio(swp_entry_t target_entry, gfp_t gfp_mask,
-				     unsigned long orders, struct vm_fault *vmf,
-				     struct mempolicy *mpol, pgoff_t ilx);
+struct folio *__swap_cache_alloc_folio(swp_entry_t target_entry, gfp_t gfp_mask,
+				       unsigned long orders, struct vm_fault *vmf,
+				       struct mempolicy *mpol, pgoff_t ilx);
 /* Below helpers require the caller to lock and pass in the swap cluster. */
 void __swap_cache_add_folio(struct swap_cluster_info *ci,
 			    struct folio *folio, swp_entry_t entry);
diff --git a/mm/swap_state.c b/mm/swap_state.c
index f3961fdd857d..e5b7fa468ade 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -489,13 +489,11 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
 	node_stat_mod_folio(folio, NR_FILE_PAGES, nr_pages);
 	lruvec_stat_mod_folio(folio, NR_SWAPCACHE, nr_pages);
 
-	/* Caller will initiate read into locked new_folio */
-	folio_add_lru(folio);
 	return folio;
 }
 
 /**
- * swap_cache_alloc_folio - Allocate folio for swapped out slot in swap cache.
+ * __swap_cache_alloc_folio - Allocate folio for swapped out slot in swap cache.
  * @targ_entry: swap entry indicating the target slot
  * @gfp: memory allocation flags
  * @orders: allocation orders, must be non zero
@@ -507,13 +505,17 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
  * doing IO (e.g. swap in or zswap writeback). The swap slot indicated by
  * @targ_entry must have a non-zero swap count (swapped out).
  *
+ * The returned folio is locked and is NOT on the LRU. The caller must either
+ * add it to the LRU with folio_add_lru() so page reclaim can find it, or free
+ * it directly once done; a folio left off the LRU is unreclaimable and leaks.
+ *
  * Context: Caller must protect the swap device with reference count or locks.
  * Return: Returns the folio if allocation succeeded and folio is in the swap
  * cache. Returns error code if failed due to race, OOM or invalid arguments.
  */
-struct folio *swap_cache_alloc_folio(swp_entry_t targ_entry, gfp_t gfp,
-				     unsigned long orders, struct vm_fault *vmf,
-				     struct mempolicy *mpol, pgoff_t ilx)
+struct folio *__swap_cache_alloc_folio(swp_entry_t targ_entry, gfp_t gfp,
+				       unsigned long orders, struct vm_fault *vmf,
+				       struct mempolicy *mpol, pgoff_t ilx)
 {
 	int order, err;
 	struct folio *ret;
@@ -649,12 +651,13 @@ static struct folio *swap_cache_read_folio(struct swap_io_ctx *ctx,
 		folio = swap_cache_get_folio(entry);
 		if (folio)
 			return folio;
-		folio = swap_cache_alloc_folio(entry, gfp, BIT(0), NULL, mpol, ilx);
+		folio = __swap_cache_alloc_folio(entry, gfp, BIT(0), NULL, mpol, ilx);
 	} while (PTR_ERR(folio) == -EEXIST);
 
 	if (IS_ERR_OR_NULL(folio))
 		return NULL;
 
+	folio_add_lru(folio);
 	swap_read_folio(ctx, folio);
 	if (readahead) {
 		folio_set_readahead(folio);
@@ -690,12 +693,13 @@ struct folio *swapin_sync(swp_entry_t entry, gfp_t gfp, unsigned long orders,
 		folio = swap_cache_get_folio(entry);
 		if (folio)
 			return folio;
-		folio = swap_cache_alloc_folio(entry, gfp, orders, vmf, mpol, ilx);
+		folio = __swap_cache_alloc_folio(entry, gfp, orders, vmf, mpol, ilx);
 	} while (PTR_ERR(folio) == -EEXIST);
 
 	if (IS_ERR(folio))
 		return folio;
 
+	folio_add_lru(folio);
 	swap_read_folio(&ctx, folio);
 	swap_read_submit(&ctx);
 	return folio;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 53bf01d5f7f1..d678a40fcaac 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1870,7 +1870,7 @@ void folio_put_swap(struct folio *folio, struct page *page)
  *   CPU1				CPU2
  *   do_swap_page()
  *     ...				swapoff+swapon
- *     swap_cache_alloc_folio()
+ *     __swap_cache_alloc_folio()
  *       // check swap_map
  *     // verify PTE not changed
  *
diff --git a/mm/zswap.c b/mm/zswap.c
index 37f34e406c8e..0d2efe21f18a 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1001,8 +1001,8 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
 		return -EEXIST;
 
 	mpol = get_task_policy(current);
-	folio = swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol,
-				       NO_INTERLEAVE_INDEX);
+	folio = __swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol,
+					 NO_INTERLEAVE_INDEX);
 	put_swap_device(si);
 
 	/*
@@ -1014,6 +1014,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
 	 */
 	if (IS_ERR(folio))
 		return PTR_ERR(folio);
+	folio_add_lru(folio);
 
 	/*
 	 * folio is locked, and the swapcache is now secured against
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v6 2/3] mm: swap: drop dropbehind swap cache folios on writeback completion
  2026-09-21 15:13 [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly Alexandre Ghiti
  2026-09-21 15:13 ` [PATCH v6 1/3] mm: swap: move LRU insertion out of the swap cache allocator Alexandre Ghiti
@ 2026-09-21 15:13 ` Alexandre Ghiti
  2026-09-21 15:13 ` [PATCH v6 3/3] mm: zswap: drop cold writeback folios via swap dropbehind Alexandre Ghiti
  2026-09-21 19:45 ` [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly Andrew Morton
  3 siblings, 0 replies; 5+ messages in thread
From: Alexandre Ghiti @ 2026-09-21 15:13 UTC (permalink / raw)
  To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Andrew Morton, Chris Li,
	Kairui Song
  Cc: Kairui Song, Chengming Zhou, Matthew Wilcox (Oracle),
	Jan Kara, Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexander Viro, Christian Brauner, David Hildenbrand,
	Lorenzo Stoakes, Michal Hocko, Axel Rasmussen, Qi Zheng,
	Shakeel Butt, Wei Xu, Yuanchu Xie, Kunwu Chan, Tal Zussman,
	linux-mm, linux-kernel, linux-fsdevel, Alexandre Ghiti

A PG_dropbehind folio is dropped from its cache once writeback completes
rather than left for reclaim to find later; this is implemented for file
folios in folio_end_dropbehind(). Extend it to swap cache folios.

The drop blocks on the folio lock, so it cannot run in interrupt
context. Set BIO_COMPLETE_IN_TASK on the write, as the file dropbehind
paths do, and drop the folio directly from folio_end_writeback().

It has to block rather than trylock: the folio is off the LRU, so
skipping it would leave it in the swap cache with nothing able to
reclaim it, and it cannot be put back while another thread holds its
lock.

Suggested-by: Yosry Ahmed <yosry@kernel.org>
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Suggested-by: Nhat Pham <nphamcs@gmail.com>
Reviewed-by: Nhat Pham <nphamcs@gmail.com>
Reviewed-by: Kunwu Chan <kunwu.chan@gmail.com>
Reviewed-by: Barry Song <baohua@kernel.org>
Signed-off-by: Alexandre Ghiti <alex@ghiti.fr>
---
 include/linux/swap.h |  6 ++++++
 mm/filemap.c         | 19 +++++++++++++++++
 mm/page_io.c         |  9 ++++++++
 mm/swap_state.c      | 42 +++++++++++++++++++++++++++++++++++++
 mm/vmscan.c          | 49 +++++++++++++++++++++++++++++++++++---------
 5 files changed, 115 insertions(+), 10 deletions(-)

diff --git a/include/linux/swap.h b/include/linux/swap.h
index 5658a1634b85..538b723a276f 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -318,6 +318,9 @@ static inline bool lru_cache_disabled(void)
 
 extern unsigned long shrink_all_memory(unsigned long nr_pages);
 long remove_mapping(struct address_space *mapping, struct folio *folio);
+long remove_mapping_set_shadow(struct address_space *mapping,
+			       struct folio *folio,
+			       struct mem_cgroup *target_memcg);
 
 #if defined(CONFIG_SYSFS) && defined(CONFIG_NUMA)
 extern int reclaim_register_node(struct node *node);
@@ -402,6 +405,8 @@ void swap_put_entries_direct(swp_entry_t entry, int nr);
  */
 bool folio_free_swap(struct folio *folio);
 
+void swap_writeback_dropbehind_folio(struct folio *folio);
+
 /* Allocate / free (hibernation) exclusive entries */
 swp_entry_t swap_alloc_hibernation_slot(int type);
 void swap_free_hibernation_slot(swp_entry_t entry);
@@ -412,6 +417,7 @@ static inline void put_swap_device(struct swap_info_struct *si)
 }
 
 #else /* CONFIG_SWAP */
+static inline void swap_writeback_dropbehind_folio(struct folio *folio) {}
 static inline struct swap_info_struct *get_swap_device(swp_entry_t entry)
 {
 	return NULL;
diff --git a/mm/filemap.c b/mm/filemap.c
index 6afec636881f..e1f1bbe943ce 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -1686,6 +1686,8 @@ EXPORT_SYMBOL_GPL(folio_end_writeback_no_dropbehind);
  */
 void folio_end_writeback(struct folio *folio)
 {
+	bool swap_dropbehind;
+
 	VM_BUG_ON_FOLIO(!folio_test_writeback(folio), folio);
 
 	/*
@@ -1695,7 +1697,24 @@ void folio_end_writeback(struct folio *folio)
 	 * reused before the folio_wake_bit().
 	 */
 	folio_get(folio);
+
+	/*
+	 * Sample this before folio_end_writeback_no_dropbehind() clears
+	 * PG_writeback: until then a racing swapin cannot remove the folio from
+	 * the swap cache. Afterwards it can, and the drop below then finds a
+	 * non-swapcache folio and puts it back on the LRU instead. The
+	 * reference taken above keeps the folio alive across that window.
+	 */
+	swap_dropbehind = folio_test_swapcache(folio) &&
+			  folio_test_dropbehind(folio);
+
 	folio_end_writeback_no_dropbehind(folio);
+
+	if (swap_dropbehind) {
+		swap_writeback_dropbehind_folio(folio);
+		return;
+	}
+
 	folio_end_dropbehind(folio);
 	folio_put(folio);
 }
diff --git a/mm/page_io.c b/mm/page_io.c
index 88962571cb93..52eae99de6e3 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -602,6 +602,15 @@ static void swap_bdev_submit_write(struct swap_io_ctx *ctx)
 		submit_bio_wait(bio);
 		end_swap_bio_write(bio);
 	} else {
+		int p;
+
+		for (p = 0; p < sio->nr_bvecs; p++) {
+			if (folio_test_dropbehind(bvec_folio(&sio->bvecs[p]))) {
+				bio_set_flag(bio, BIO_COMPLETE_IN_TASK);
+				break;
+			}
+		}
+
 		bio->bi_end_io = end_swap_bio_write;
 		submit_bio(bio);
 	}
diff --git a/mm/swap_state.c b/mm/swap_state.c
index e5b7fa468ade..b1656e2d5288 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -543,6 +543,48 @@ struct folio *__swap_cache_alloc_folio(swp_entry_t targ_entry, gfp_t gfp,
 	return ret;
 }
 
+/**
+ * swap_writeback_dropbehind_folio - drop a dropbehind swap cache folio
+ * @folio: the off-LRU folio whose writeback has completed
+ *
+ * Context: task context, with the reference taken by folio_end_writeback()
+ * donated to us.
+ */
+void swap_writeback_dropbehind_folio(struct folio *folio)
+{
+	struct mem_cgroup *memcg;
+
+	folio_lock(folio);
+
+	/* The folio was allocated off the LRU and nothing re-adds it here. */
+	VM_WARN_ON_ONCE_FOLIO(folio_test_lru(folio), folio);
+
+	rcu_read_lock();
+	memcg = folio_memcg(folio);
+	if (!mem_cgroup_tryget(memcg))
+		memcg = NULL;
+	rcu_read_unlock();
+
+	/*
+	 * Gate remove_mapping_set_shadow() on folio_test_swapcache(): a racing
+	 * swapin may have freed the swap slot (folio_free_swap()) and dropped the
+	 * folio from the cache, and it must not run on a non-swapcache folio (it
+	 * would trip __remove_mapping()'s mapping == folio_mapping() check).
+	 */
+	if (!folio_test_swapcache(folio) || folio_test_writeback(folio) ||
+	    !remove_mapping_set_shadow(swap_address_space(folio->swap), folio,
+				       memcg)) {
+		/* Raced: the folio is now owned by the swapin; put it back. */
+		folio_clear_dropbehind(folio);
+		folio_add_lru(folio);
+	}
+
+	mem_cgroup_put(memcg);
+
+	folio_unlock(folio);
+	folio_put(folio);
+}
+
 /*
  * If we are the only user, then try to free up the swap cache.
  *
diff --git a/mm/vmscan.c b/mm/vmscan.c
index f11491ee9ed5..a02f942418d3 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -783,6 +783,22 @@ static int __remove_mapping(struct address_space *mapping, struct folio *folio,
 	return 0;
 }
 
+static long __remove_mapping_unfreeze(struct address_space *mapping,
+				      struct folio *folio, bool reclaimed,
+				      struct mem_cgroup *target_memcg)
+{
+	if (__remove_mapping(mapping, folio, reclaimed, target_memcg)) {
+		/*
+		 * Unfreezing the refcount with 1 effectively
+		 * drops the pagecache ref for us without requiring another
+		 * atomic operation.
+		 */
+		folio_ref_unfreeze(folio, 1);
+		return folio_nr_pages(folio);
+	}
+	return 0;
+}
+
 /**
  * remove_mapping() - Attempt to remove a folio from its mapping.
  * @mapping: The address space.
@@ -797,16 +813,29 @@ static int __remove_mapping(struct address_space *mapping, struct folio *folio,
  */
 long remove_mapping(struct address_space *mapping, struct folio *folio)
 {
-	if (__remove_mapping(mapping, folio, false, NULL)) {
-		/*
-		 * Unfreezing the refcount with 1 effectively
-		 * drops the pagecache ref for us without requiring another
-		 * atomic operation.
-		 */
-		folio_ref_unfreeze(folio, 1);
-		return folio_nr_pages(folio);
-	}
-	return 0;
+	return __remove_mapping_unfreeze(mapping, folio, false, NULL);
+}
+
+/**
+ * remove_mapping_set_shadow() - Remove a folio and record an eviction shadow.
+ * @mapping: The address space.
+ * @folio: The folio to remove.
+ * @target_memcg: The memcg to charge the eviction shadow to; the caller must
+ *                keep it alive across the call.
+ *
+ * Like remove_mapping(), but stores a workingset eviction shadow the way page
+ * reclaim does, so that a later refault can be detected and the folio
+ * re-activated.
+ * Return: The number of pages removed from the mapping.  0 if the folio
+ * could not be removed.
+ * Context: The caller should have a single refcount on the folio and
+ * hold its lock.
+ */
+long remove_mapping_set_shadow(struct address_space *mapping,
+			       struct folio *folio,
+			       struct mem_cgroup *target_memcg)
+{
+	return __remove_mapping_unfreeze(mapping, folio, true, target_memcg);
 }
 
 /**
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v6 3/3] mm: zswap: drop cold writeback folios via swap dropbehind
  2026-09-21 15:13 [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly Alexandre Ghiti
  2026-09-21 15:13 ` [PATCH v6 1/3] mm: swap: move LRU insertion out of the swap cache allocator Alexandre Ghiti
  2026-09-21 15:13 ` [PATCH v6 2/3] mm: swap: drop dropbehind swap cache folios on writeback completion Alexandre Ghiti
@ 2026-09-21 15:13 ` Alexandre Ghiti
  2026-09-21 19:45 ` [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly Andrew Morton
  3 siblings, 0 replies; 5+ messages in thread
From: Alexandre Ghiti @ 2026-09-21 15:13 UTC (permalink / raw)
  To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Andrew Morton, Chris Li,
	Kairui Song
  Cc: Kairui Song, Chengming Zhou, Matthew Wilcox (Oracle),
	Jan Kara, Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexander Viro, Christian Brauner, David Hildenbrand,
	Lorenzo Stoakes, Michal Hocko, Axel Rasmussen, Qi Zheng,
	Shakeel Butt, Wei Xu, Yuanchu Xie, Kunwu Chan, Tal Zussman,
	linux-mm, linux-kernel, linux-fsdevel, Alexandre Ghiti

zswap writeback decompresses an entry into a fresh swap cache folio and
writes it back. The folio is cold by construction, yet it is left on the
LRU for reclaim to find and free later, wasting a reclaim scan and keeping
cold memory resident longer than necessary.

Allocate the folio off the LRU and mark it PG_dropbehind so the swap
dropbehind path frees it from the swap cache once writeback completes.

__swap_cache_alloc_folio() evaluates a refault on the new folio, and
workingset_refault() sets PG_active when it looks recent. Until now
folio_add_lru() consumed that flag and __page_cache_release() cleared it
once the folio left the LRU. This folio never reaches the LRU, so nothing
would clear PG_active and the folio would be freed with a
PAGE_FLAGS_CHECK_AT_FREE flag set, tripping bad_page() under
CONFIG_DEBUG_VM. Clear it after allocation.

That is a workaround: the refault should not be evaluated on a writeback
buffer at all. A fix for that is on the mailing list [1].

Link: https://lore.kernel.org/linux-mm/20260911092012.92399-1-alex@ghiti.fr/ [1]
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Suggested-by: Nhat Pham <nphamcs@gmail.com>
Reviewed-by: Nhat Pham <nphamcs@gmail.com>
Reviewed-by: Kunwu Chan <kunwu.chan@gmail.com>
Signed-off-by: Alexandre Ghiti <alex@ghiti.fr>
---
 mm/zswap.c | 33 +++++++++++++++++++++++----------
 1 file changed, 23 insertions(+), 10 deletions(-)

diff --git a/mm/zswap.c b/mm/zswap.c
index 0d2efe21f18a..dc8425d6b21e 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1014,7 +1014,8 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
 	 */
 	if (IS_ERR(folio))
 		return PTR_ERR(folio);
-	folio_add_lru(folio);
+
+	folio_clear_active(folio);
 
 	/*
 	 * folio is locked, and the swapcache is now secured against
@@ -1028,12 +1029,12 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
 	tree = swap_zswap_tree(swpentry);
 	if (entry != xa_load(tree, offset)) {
 		ret = -ENOMEM;
-		goto out;
+		goto err;
 	}
 
 	if (!zswap_decompress(entry, folio)) {
 		ret = -EIO;
-		goto out;
+		goto err;
 	}
 
 	xa_erase(tree, offset);
@@ -1047,18 +1048,30 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
 	/* folio is up to date */
 	folio_mark_uptodate(folio);
 
-	/* move it to the tail of the inactive list after end_writeback */
-	folio_set_reclaim(folio);
+	folio_set_dropbehind(folio);
+
+	/*
+	 * Drop our reference before starting writeback so the swap cache holds
+	 * the only one: the drop in folio_end_writeback() needs that for
+	 * remove_mapping_set_shadow() to succeed, otherwise the folio is
+	 * handed back to reclaim instead.
+	 *
+	 * Nothing can free the folio in the meantime: we hold the folio lock
+	 * until writeback starts, PG_writeback then blocks swap cache removal,
+	 * and folio_end_writeback() takes its own reference before clearing
+	 * PG_writeback and donates it to the drop.
+	 */
+	folio_put(folio);
 
 	/* start writeback */
 	__swap_writepage(&ctx, folio);
 	swap_write_submit(&ctx);
 
-out:
-	if (ret) {
-		swap_cache_del_folio(folio);
-		folio_unlock(folio);
-	}
+	return 0;
+
+err:
+	swap_cache_del_folio(folio);
+	folio_unlock(folio);
 	folio_put(folio);
 	return ret;
 }
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly
  2026-09-21 15:13 [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly Alexandre Ghiti
                   ` (2 preceding siblings ...)
  2026-09-21 15:13 ` [PATCH v6 3/3] mm: zswap: drop cold writeback folios via swap dropbehind Alexandre Ghiti
@ 2026-09-21 19:45 ` Andrew Morton
  3 siblings, 0 replies; 5+ messages in thread
From: Andrew Morton @ 2026-09-21 19:45 UTC (permalink / raw)
  To: Alexandre Ghiti
  Cc: Johannes Weiner, Yosry Ahmed, Nhat Pham, Chris Li, Kairui Song,
	Kairui Song, Chengming Zhou, Matthew Wilcox (Oracle),
	Jan Kara, Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexander Viro, Christian Brauner, David Hildenbrand,
	Lorenzo Stoakes, Michal Hocko, Axel Rasmussen, Qi Zheng,
	Shakeel Butt, Wei Xu, Yuanchu Xie, Kunwu Chan, Tal Zussman,
	linux-mm, linux-kernel, linux-fsdevel

On Mon, 21 Sep 2026 17:13:01 +0200 Alexandre Ghiti <alex@ghiti.fr> wrote:

> When zswap writes an entry back, it allocates an order-0 swap cache folio,
> decompresses into it, and issues the write. The folio is cold by
> construction, yet today it is left on the LRU for page reclaim to find and
> free later. That wastes a reclaim scan and keeps cold memory resident
> longer than necessary.
> 
> Rather than implement this in zswap, extend the existing dropbehind
> mechanism to swap cache folios and have zswap opt into it (Yosry). A
> PG_dropbehind folio is already dropped from its cache once writeback
> completes instead of being left for reclaim; for a swap cache folio that
> "drop" is removing it from the swap cache.

Thanks, I updated mm.git's mm-unstable branch to this version.

> Changes in v6:
> - No functional change: this version only updates changelogs and collects
>   tags.

Confirmed.

> - Patch 1: fix the changelog. The first paragraph describes the behaviour
>   before the patch, so it must name swap_cache_alloc_folio(), not the
>   __swap_cache_alloc_folio() this patch introduces, and the last paragraph
>   now says the helper is renamed as well as that the LRU insertion is
>   deferred (Kairui).
> - Patch 2: fix the changelog. The swap cluster lock is a spinlock taken
>   with interrupts disabled and does not sleep; the folio lock is what the
>   drop blocks on, and that is why the completion needs task context. Also
>   spell out why it blocks instead of using folio_trylock() (Barry).
> - Collect Reviewed-by tags. Thanks to Kairui, Barry and Nhat for the
>   reviews! Kairui's tag on patch 1 was given on the posting of the zswap
>   writeback refault series [1], which carries the same patch.


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-21 19:45 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-21 15:13 [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly Alexandre Ghiti
2026-09-21 15:13 ` [PATCH v6 1/3] mm: swap: move LRU insertion out of the swap cache allocator Alexandre Ghiti
2026-09-21 15:13 ` [PATCH v6 2/3] mm: swap: drop dropbehind swap cache folios on writeback completion Alexandre Ghiti
2026-09-21 15:13 ` [PATCH v6 3/3] mm: zswap: drop cold writeback folios via swap dropbehind Alexandre Ghiti
2026-09-21 19:45 ` [PATCH v6 0/3] mm: zswap: free cold writeback folios promptly Andrew Morton

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®