From: Usama Arif <usama.arif@linux.dev>
To: Nhat Pham <nphamcs@gmail.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
chengming.zhou@linux.dev, dsterba@suse.com, hannes@cmpxchg.org,
linux-kernel@vger.kernel.org, linux-mm@kvack.org,
terrelln@fb.com, yosry@kernel.org, riel@surriel.com,
shakeel.butt@linux.dev, alex@ghiti.fr, senozhatsky@chromium.org,
kernel-team@meta.com
Subject: Re: [PATCH 0/2] mm: zswap: reduce request contention on loads
Date: Wed, 7 Oct 2026 12:56:10 +0200 [thread overview]
Message-ID: <209aaa32-7763-42d8-83e0-470066c880dd@linux.dev> (raw)
In-Reply-To: <CAKEwX=N3ks2O-dCYPRif3NnHQckdVqLiPGRn03GkCuEE+FVgtQ@mail.gmail.com>
On 06/10/2026 11:35, Nhat Pham wrote:
> On Tue, Oct 6, 2026 at 2:23 AM Usama Arif <usama.arif@linux.dev> wrote:
>>
>> Stores and loads share a per-CPU acomp request and mutex. A low-priority
>> store can be preempted right after the compressor drops its stream
>> lock, while it still holds the zswap mutex, and a higher-priority load
>> on that CPU then waits for the store to run again. This follows the work
>> from Sergey Senozhatsky's zram series which splits it for the same
>> reason [1].
>>
>> Patch 1 gives compression and decompression separate requests, waits
>> and mutexes, so loads no longer wait for stores, though they can still
>> wait for each other. Patch 2 decompresses with an on-stack request when
>> the algorithm is synchronous and needs no request context, which covers
>> all in-tree software compressors, so those loads take no zswap lock.
>> Asynchronous algorithms keep the per-CPU request and mutex. For software
>> compressors the series allocates the same number of requests as before;
>> each per-CPU context grows by 72 bytes, and the load path is about 270
>> bytes deeper on x86-64.
>>
>> The series does not fix two related cases:
>> - Stores still serialize on the compression mutex, so a high-priority
>> task that reclaims (direct reclaim, MADV_PAGEOUT) can still wait for
>> a preempted store.
>
> Any reasons why we cannot tackle this too? Or just one at a time?
Stores need more than a request. The compression mutex also protects
the per-CPU PAGE_SIZE output buffer, which has to stay ours until
zs_obj_write() copies it out, since zs_malloc() needs the compressed
length first. Loads stopped using that buffer in e2c3b6b21c77f, so an
on-stack request was enough for them, but the buffer is too big for
the stack.
>
>> - On PREEMPT_RT the codec stream locks are preemptible, so a load can
>> still wait for a preempted store inside the codec.
>
> Acked.
>
>>
>> The numbers below are the slowest read per run, as a median (min-max)
>> of 5 runs. Each run is 12 seconds in a zstd VM with lazy preemption,
>> vm.page-cluster=0 and swap on /dev/ram0. With 1 vCPU, four nice +10
>> workers page memory out and read it back while a nice 0 task spins. A
>> nice -19 reader pages out its own buffer and measures how long each
>> read of it takes. With 8 vCPUs there are 16 workers, 8 spinning tasks
>> and 8 readers.
>>
>> Before series (ms) With series (ms)
>> 1 vCPU 22.3 (21.6-22.6) 0.97 (0.72-1.4)
>> 8 vCPUs 314 (97-2542) 7.0 (5.0-98)
>>
>> Reads over 10 ms fell from 26-35 per run to none with 1 vCPU, and from
>> 3-18 per run to at most one with 8 vCPUs. The benchmark and test programs
>> were written with the help of an LLM.
>
> Great find, Usama!
Thanks for the reviews!
prev parent reply other threads:[~2026-10-07 10:56 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-06 0:22 Usama Arif
2026-10-06 0:22 ` [PATCH 1/2] mm: zswap: use separate compression and decompression requests Usama Arif
2026-10-06 9:47 ` Nhat Pham
2026-10-06 9:52 ` Sergey Senozhatsky
2026-10-07 11:02 ` Usama Arif
2026-10-06 0:22 ` [PATCH 2/2] mm: zswap: use stack requests for synchronous decompression Usama Arif
2026-10-07 5:46 ` Nhat Pham
2026-10-06 9:18 ` [PATCH 0/2] mm: zswap: reduce request contention on loads Usama Arif
2026-10-06 9:35 ` Nhat Pham
2026-10-07 10:56 ` Usama Arif [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=209aaa32-7763-42d8-83e0-470066c880dd@linux.dev \
--to=usama.arif@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=alex@ghiti.fr \
--cc=chengming.zhou@linux.dev \
--cc=dsterba@suse.com \
--cc=hannes@cmpxchg.org \
--cc=kernel-team@meta.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=nphamcs@gmail.com \
--cc=riel@surriel.com \
--cc=senozhatsky@chromium.org \
--cc=shakeel.butt@linux.dev \
--cc=terrelln@fb.com \
--cc=yosry@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®