* [PATCH v3 0/4] PM: hibernate: flush icache for restored pages from task context
@ 2026-10-02 3:41 Xiong Xin
2026-10-02 3:41 ` [PATCH v3 1/4] PM: hibernate: reset clean_pages_on_read after load_image() Xiong Xin
` (3 more replies)
0 siblings, 4 replies; 5+ messages in thread
From: Xiong Xin @ 2026-10-02 3:41 UTC (permalink / raw)
To: rafael, lenb, pavel, james.morse, hch
Cc: will.deacon, linux-pm, linux-kernel, Xiong Xin
Hello,
With `hibernate=nocompress`, load_image() submits the read bios
asynchronously and hib_end_io() completes them from hardirq/softirq
context. The flush_icache_range() call there, added by the
commit f6cf0545ec69 ("PM / Hibernate: Call flush_icache_range() on pages
restored in-place"), ends with kick_all_cpus_sync() on ARM64 and so
triggers the WARN_ON_ONCE(!in_task()) in smp_call_function().
This series fixes it by keeping the flush in task context, and first
closes three pre-existing holes in the restore error paths that the
rework would otherwise trip over.
Patch 1 resets clean_pages_on_read when load_image() finishes, so that
a failed nocompress resume can no longer leak the flag into a later
compressed resume, where it made hib_end_io() call flush_icache_range()
from interrupt context for the reused ring pages of
load_compressed_image(). The flag is also cleared as soon as the load
is known to have failed, so that the bios still in flight do not flush
from interrupt context during the final drain.
Patch 2 stops snapshot_write_next() from releasing all snapshot pages
in place when one of its internal allocations fails: the read bios do
not take page references, so freed pages could be handed out again
while the device is still writing the image data into them. The image
is now released by the outer error paths, after swsusp_read() has
drained the batch.
Patch 3 makes the error exits of load_compressed_image() wait for the
read bios that may still be in flight before the ring pages are freed,
like save_compressed_image() and load_image() already do.
Patch 4 defers the flush on the nocompress path: hib_end_io() now only
records the restored page on a per-batch list, and load_image() flushes
the pages once hib_wait_io() has returned, preserving the asynchronous
batched I/O.
Patches 1 to 3 are new in v3; the per-patch changelogs are in the ---
sections of each commit. Many thanks to sashiko for the v2 review,
both findings are addressed.
Xiong Xin (4):
PM: hibernate: reset clean_pages_on_read after load_image()
PM: hibernate: do not free the image in place while restore I/O is in
flight
PM: hibernate: drain pending reads in compressed image load errors
PM: hibernate: flush icache for restored pages from task context
kernel/power/snapshot.c | 9 ++-----
kernel/power/swap.c | 55 +++++++++++++++++++++++++++++++++++------
2 files changed, 50 insertions(+), 14 deletions(-)
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 1/4] PM: hibernate: reset clean_pages_on_read after load_image()
2026-10-02 3:41 [PATCH v3 0/4] PM: hibernate: flush icache for restored pages from task context Xiong Xin
@ 2026-10-02 3:41 ` Xiong Xin
2026-10-02 3:41 ` [PATCH v3 2/4] PM: hibernate: do not free the image in place while restore I/O is in flight Xiong Xin
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Xiong Xin @ 2026-10-02 3:41 UTC (permalink / raw)
To: rafael, lenb, pavel, james.morse, hch
Cc: will.deacon, linux-pm, linux-kernel, Xiong Xin, Sashiko
load_image() sets clean_pages_on_read but nothing clears it. If a
nocompress resume attempt fails and a compressed resume follows in the
same boot, the stale flag makes hib_end_io() call flush_icache_range()
from interrupt context for the reused ring pages of
load_compressed_image(), triggering the WARN_ON_ONCE(!in_task()) in
smp_call_function() on ARM64.
Reset the flag before load_image() returns, so that it cannot leak
into a later compressed resume.
Also clear it once the load is known to have failed, so that the bios
still in flight do not trigger the same interrupt-context flush during
the final drain.
Fixes: f6cf0545ec69 ("PM / Hibernate: Call flush_icache_range() on pages restored in-place")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://sashiko.dev/#/patchset/20260926032416.467937-1-xiongxin%40kylinos.cn
Assisted-by: GLM:glm-5.3
Signed-off-by: Xiong Xin <xiongxin@kylinos.cn>
---
New in v3, split from the single-patch v2 per review feedback from
sashiko, as it also fixes the stale clean_pages_on_read path on its
own; it now also covers the final drain of a failing load.
kernel/power/swap.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/kernel/power/swap.c b/kernel/power/swap.c
index c78f1593600b..7f3cc2cc5cca 100644
--- a/kernel/power/swap.c
+++ b/kernel/power/swap.c
@@ -1125,6 +1125,9 @@ static int load_image(struct swap_map_handle *handle,
nr_pages / m * 10);
nr_pages++;
}
+ /* Stop flushing pages on load failure. */
+ if (ret < 0)
+ clean_pages_on_read = false;
err2 = hib_wait_io(&hb);
hib_finish_batch(&hb);
stop = ktime_get();
@@ -1137,6 +1140,7 @@ static int load_image(struct swap_map_handle *handle,
ret = -ENODATA;
}
swsusp_show_speed(start, stop, nr_to_read, "Read");
+ clean_pages_on_read = false;
return ret;
}
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 2/4] PM: hibernate: do not free the image in place while restore I/O is in flight
2026-10-02 3:41 [PATCH v3 0/4] PM: hibernate: flush icache for restored pages from task context Xiong Xin
2026-10-02 3:41 ` [PATCH v3 1/4] PM: hibernate: reset clean_pages_on_read after load_image() Xiong Xin
@ 2026-10-02 3:41 ` Xiong Xin
2026-10-02 3:41 ` [PATCH v3 3/4] PM: hibernate: drain pending reads in compressed image load errors Xiong Xin
2026-10-02 3:42 ` [PATCH v3 4/4] PM: hibernate: flush icache for restored pages from task context Xiong Xin
3 siblings, 0 replies; 5+ messages in thread
From: Xiong Xin @ 2026-10-02 3:41 UTC (permalink / raw)
To: rafael, lenb, pavel, james.morse, hch
Cc: will.deacon, linux-pm, linux-kernel, Xiong Xin
snapshot_write_next() calls swsusp_free() to release all snapshot pages
when one of its internal allocations fails, in prepare_image(),
get_buffer() or get_highmem_page_buffer(). At that point read bios
submitted by load_image() or load_compressed_image() may still be in
flight: the bios do not take page references, so nothing prevents the
freed pages from being handed out again while the device is still
writing the image data into them.
Drop the in-place releases and let the error paths of
load_image_and_restore() and hibernate() free the image, after
swsusp_read() has drained the batch.
Fixes: 343df3c79c62 ("suspend: simplify block I/O handling")
Assisted-by: GLM:glm-5.3
Signed-off-by: Xiong Xin <xiongxin@kylinos.cn>
---
New in v3: without it, patch 4 could touch snapshot pages that this
in-place release has already freed.
kernel/power/snapshot.c | 9 ++-------
1 file changed, 2 insertions(+), 7 deletions(-)
diff --git a/kernel/power/snapshot.c b/kernel/power/snapshot.c
index b209712cb2c3..244f4a19c025 100644
--- a/kernel/power/snapshot.c
+++ b/kernel/power/snapshot.c
@@ -2520,10 +2520,8 @@ static void *get_highmem_page_buffer(struct page *page,
* use a "safe" page frame to store the loaded page.
*/
pbe = chain_alloc(ca, sizeof(struct highmem_pbe));
- if (!pbe) {
- swsusp_free();
+ if (!pbe)
return ERR_PTR(-ENOMEM);
- }
pbe->orig_page = page;
if (safe_highmem_pages > 0) {
struct page *tmp;
@@ -2702,7 +2700,6 @@ static int prepare_image(struct memory_bitmap *new_bm, struct memory_bitmap *bm,
return 0;
Free:
- swsusp_free();
return error;
}
@@ -2737,10 +2734,8 @@ static void *get_buffer(struct memory_bitmap *bm, struct chain_allocator *ca)
* use a "safe" page frame to store the loaded page.
*/
pbe = chain_alloc(ca, sizeof(struct pbe));
- if (!pbe) {
- swsusp_free();
+ if (!pbe)
return ERR_PTR(-ENOMEM);
- }
pbe->orig_address = page_address(page);
pbe->address = __get_safe_page(ca->gfp_mask);
if (!pbe->address)
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 3/4] PM: hibernate: drain pending reads in compressed image load errors
2026-10-02 3:41 [PATCH v3 0/4] PM: hibernate: flush icache for restored pages from task context Xiong Xin
2026-10-02 3:41 ` [PATCH v3 1/4] PM: hibernate: reset clean_pages_on_read after load_image() Xiong Xin
2026-10-02 3:41 ` [PATCH v3 2/4] PM: hibernate: do not free the image in place while restore I/O is in flight Xiong Xin
@ 2026-10-02 3:41 ` Xiong Xin
2026-10-02 3:42 ` [PATCH v3 4/4] PM: hibernate: flush icache for restored pages from task context Xiong Xin
3 siblings, 0 replies; 5+ messages in thread
From: Xiong Xin @ 2026-10-02 3:41 UTC (permalink / raw)
To: rafael, lenb, pavel, james.morse, hch
Cc: will.deacon, linux-pm, linux-kernel, Xiong Xin
The error exits of load_compressed_image() free the ring buffer pages
without waiting for the read bios that may still be in flight: when a
read or decompression error aborts the load mid-round, the device can
still be writing into pages that are handed back to the allocator by
out_clean().
save_compressed_image() and load_image() both drain the batch on their
way out; add the same wait here, which also folds the error of any
in-flight read into the return value instead of dropping it.
Fixes: 343df3c79c62 ("suspend: simplify block I/O handling")
Assisted-by: GLM:glm-5.3
Signed-off-by: Xiong Xin <xiongxin@kylinos.cn>
---
New in v3: the compressed loader had the same missing drain on its
error exits, freeing the ring pages under in-flight reads.
kernel/power/swap.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/kernel/power/swap.c b/kernel/power/swap.c
index 7f3cc2cc5cca..571aefaf09f4 100644
--- a/kernel/power/swap.c
+++ b/kernel/power/swap.c
@@ -1204,6 +1204,7 @@ static int load_compressed_image(struct swap_map_handle *handle,
int ret = 0;
int eof = 0;
struct hib_bio_batch hb;
+ int err2;
ktime_t start;
ktime_t stop;
unsigned nr_pages;
@@ -1482,11 +1483,14 @@ static int load_compressed_image(struct swap_map_handle *handle,
}
out_finish:
+ err2 = hib_wait_io(&hb);
if (crc->run_threads) {
wait_event(crc->done, atomic_read_acquire(&crc->stop));
atomic_set(&crc->stop, 0);
}
stop = ktime_get();
+ if (!ret)
+ ret = err2;
if (!ret) {
pr_info("Image loading done\n");
ret = snapshot_write_finalize(snapshot);
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 4/4] PM: hibernate: flush icache for restored pages from task context
2026-10-02 3:41 [PATCH v3 0/4] PM: hibernate: flush icache for restored pages from task context Xiong Xin
` (2 preceding siblings ...)
2026-10-02 3:41 ` [PATCH v3 3/4] PM: hibernate: drain pending reads in compressed image load errors Xiong Xin
@ 2026-10-02 3:42 ` Xiong Xin
3 siblings, 0 replies; 5+ messages in thread
From: Xiong Xin @ 2026-10-02 3:42 UTC (permalink / raw)
To: rafael, lenb, pavel, james.morse, hch
Cc: will.deacon, linux-pm, linux-kernel, Xiong Xin, Sashiko, Riwen Lu
With `hibernate=nocompress`, load_image() submits read bios asynchronously
and hib_end_io() completes them from hardirq/softirq context. The
commit f6cf0545ec69 ("PM / Hibernate: Call flush_icache_range() on pages
restored in-place") calls flush_icache_range() there, which on ARM64 ends
kick_all_cpus_sync() and so triggers
WARN_ON_ONCE(!in_task());
in smp_call_function(), as in_task() is always false in interrupt context.
Only the nocompress path is affected: load_compressed_image() flushes from
the decompression kthread, in task context.
Defer the flush to task context. hib_end_io() now only records the
restored page on a per-batch list (page->lru is free to use, as safe
in-place pages are on no LRU). Once hib_wait_io() returns to load_image(),
hib_flush_icache_pages() walks the list and flushes each page, preserving
the asynchronous batched I/O throughput.
Fixes: f6cf0545ec69 ("PM / Hibernate: Call flush_icache_range() on pages restored in-place")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://sashiko.dev/#/patchset/20260926032416.467937-1-xiongxin%40kylinos.cn
Assisted-by: GLM:glm-5.3
Acked-by: Riwen Lu <luriwen@kylinos.cn>
Signed-off-by: Xiong Xin <xiongxin@kylinos.cn>
---
Changes in v2:
- Polish the commit message and drop the embedded call trace.
Changes in v3, addressing review feedback from sashiko:
- Take icache_lock with spin_lock_irqsave()/spin_unlock_irqrestore():
hib_end_io() can also be called from task context when submit_bio()
completes a bio inline (e.g. on error), and a plain spin_lock() held
there could deadlock against a hardirq/softirq completion of another
bio of the same batch on the same CPU.
- Drop the hib_wait_io() error check in load_image() that became
unreachable after adding the in-block break.
- Move the clean_pages_on_read reset to 1/4 and rebase on the new
2/4 and 3/4, which keep the restored pages alive until the batch
has been drained.
kernel/power/swap.c | 47 ++++++++++++++++++++++++++++++++++++++-------
1 file changed, 40 insertions(+), 7 deletions(-)
diff --git a/kernel/power/swap.c b/kernel/power/swap.c
index 571aefaf09f4..fb26f1376b92 100644
--- a/kernel/power/swap.c
+++ b/kernel/power/swap.c
@@ -223,6 +223,8 @@ struct hib_bio_batch {
wait_queue_head_t wait;
blk_status_t error;
struct blk_plug plug;
+ struct list_head icache_pages;
+ spinlock_t icache_lock;
};
static void hib_init_batch(struct hib_bio_batch *hb)
@@ -230,6 +232,8 @@ static void hib_init_batch(struct hib_bio_batch *hb)
atomic_set(&hb->count, 0);
init_waitqueue_head(&hb->wait);
hb->error = BLK_STS_OK;
+ INIT_LIST_HEAD(&hb->icache_pages);
+ spin_lock_init(&hb->icache_lock);
blk_start_plug(&hb->plug);
}
@@ -242,6 +246,7 @@ static void hib_end_io(struct bio *bio)
{
struct hib_bio_batch *hb = bio->bi_private;
struct page *page = bio_first_page_all(bio);
+ unsigned long flags;
if (bio->bi_status) {
pr_alert("Read-error on swap-device (%u:%u:%Lu)\n",
@@ -249,11 +254,19 @@ static void hib_end_io(struct bio *bio)
(unsigned long long)bio->bi_iter.bi_sector);
}
- if (bio_data_dir(bio) == WRITE)
+ if (bio_data_dir(bio) == WRITE) {
put_page(page);
- else if (clean_pages_on_read)
- flush_icache_range((unsigned long)page_address(page),
- (unsigned long)page_address(page) + PAGE_SIZE);
+ } else if (clean_pages_on_read) {
+ /*
+ * Stash the page on the batch list, load_image()
+ * will flush it once hib_wait_io() has returned to task
+ * context. The page is neither on an LRU nor in the page
+ * cache, so its lru member is free to use as link node.
+ */
+ spin_lock_irqsave(&hb->icache_lock, flags);
+ list_add(&page->lru, &hb->icache_pages);
+ spin_unlock_irqrestore(&hb->icache_lock, flags);
+ }
if (bio->bi_status && !hb->error)
hb->error = bio->bi_status;
@@ -295,6 +308,23 @@ static int hib_wait_io(struct hib_bio_batch *hb)
return blk_status_to_errno(hb->error);
}
+/*
+ * Flush the icache for pages restored since the last hib_wait_io().
+ * Must be called from task context (after hib_wait_io() returned): all
+ * bios of the batch have completed, so the icache_pages list is no longer
+ * touched by hib_end_io() and can be walked without the lock.
+ */
+static void hib_flush_icache_pages(struct hib_bio_batch *hb)
+{
+ struct page *page, *tmp;
+
+ list_for_each_entry_safe(page, tmp, &hb->icache_pages, lru) {
+ list_del_init(&page->lru);
+ flush_icache_range((unsigned long)page_address(page),
+ (unsigned long)page_address(page) + PAGE_SIZE);
+ }
+}
+
/*
* Saving part
*/
@@ -1116,10 +1146,12 @@ static int load_image(struct swap_map_handle *handle,
ret = swap_read_page(handle, data_of(*snapshot), &hb);
if (ret)
break;
- if (snapshot->sync_read)
+ if (snapshot->sync_read) {
ret = hib_wait_io(&hb);
- if (ret)
- break;
+ if (ret)
+ break;
+ hib_flush_icache_pages(&hb);
+ }
if (!(nr_pages % m))
pr_info("Image loading progress: %3d%%\n",
nr_pages / m * 10);
@@ -1134,6 +1166,7 @@ static int load_image(struct swap_map_handle *handle,
if (!ret)
ret = err2;
if (!ret) {
+ hib_flush_icache_pages(&hb);
pr_info("Image loading done\n");
ret = snapshot_write_finalize(snapshot);
if (!ret && !snapshot_image_loaded(snapshot))
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-02 3:42 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 3:41 [PATCH v3 0/4] PM: hibernate: flush icache for restored pages from task context Xiong Xin
2026-10-02 3:41 ` [PATCH v3 1/4] PM: hibernate: reset clean_pages_on_read after load_image() Xiong Xin
2026-10-02 3:41 ` [PATCH v3 2/4] PM: hibernate: do not free the image in place while restore I/O is in flight Xiong Xin
2026-10-02 3:41 ` [PATCH v3 3/4] PM: hibernate: drain pending reads in compressed image load errors Xiong Xin
2026-10-02 3:42 ` [PATCH v3 4/4] PM: hibernate: flush icache for restored pages from task context Xiong Xin
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®