From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A0C50576EB3 for ; Tue, 22 Sep 2026 17:16:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790097397; cv=none; b=CISrWaQ0P83O0y1cnFdBnAH5+Oeczyu24nU7bqO4mlHGiXI3a59fyBSEhrP++9K0/yeGp756CmuXcVXMO4nw7EZOQVK11eLsGS5yxdBDuYkOgHs12oqc7jDQmTtW39higdnEhQ+NtfgLtvU2nde+JDmQVmpEMq8xUxMETeJa800= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790097397; c=relaxed/simple; bh=F9wrWhleKxmtzDbWxesfxqdYHneA06HebFTlC/eppXw=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=nS4W9OZQD5aY/whiXMve1xzjzajnLJrFHc0A/nr0mN2ZaCjL6Atgw+7bqdOQo2K2L8A0dYMixSmZq/DO0tYu9qYSF1pRAYx94+zm/XhgLpC1AujUvW8nhfEoW/gLFe45Jw9nxywzMVLb0NZ93ungt4lQy/074juuDrP+N7BjKto= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=ZH4Xw0K8; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="ZH4Xw0K8" Received: by mail-qk2-f13.google.com with SMTP id af79cd13be357-939109f067fso12780285a.2 for ; Tue, 22 Sep 2026 10:16:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1790097393; x=1790702193; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=CJkQfQB3a9RKvj8bM5CBOZYh4CkWNCRPcTCANcgJj3I=; b=ZH4Xw0K8TWVxEt3LU5a0ZAvXnrqs2hWG2jpFUDzLLuhJS/3K5FQMz0vOuMG3EvZySY 7B9YNXNP9msgSJau3yceZbUWsl5YRuMfKoM44vSwLV8HiCkZfEibnCrTXQ6KAlwwsLOL hplsjEvjHkBozKr9K7rlA9BeoIRqaIUdPFxxndrSpQr10abL7O/kh57mu4vvqALKx4KA PO8gnGJ7e/9zok1lRs2Qsu8hu+L3JzuL999xZJ2NMBB5jcS+UE5kd4JaLOblFYWwPct8 wgD0PpGsvuyDyU98SvCudgBYnX8B1CtH5GX7szzLRcswZtW8b4FYcQIWAw/ClEIuZk2s 2n3g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790097393; x=1790702193; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=CJkQfQB3a9RKvj8bM5CBOZYh4CkWNCRPcTCANcgJj3I=; b=IBY6ddvNS4Vg5CSPWgsIue2/wI20eNQ6IxxxFY1wMGAiFDaiFusDXv7LOhsHmlLvfT XDMqblhwh7w9edBiOWeCBY+9wzuIHHpKhsXFshfUEOLsTuuscgdpqYi+eEvvVdXf6+1z DBK+3Cq+jH0185LPnrLCcT80JA6TA+WnEcWQhKi5ksDDQ3CNoG04TmMqCbBUk/mShpJ8 ffdHgXcEkXZ1FtxppXP77jVensdduOadrq8KnKGXbp8OE1yiX5GYFwcqCyKaa0FQx+OX hD4bvZh/5LTtTsuwPAxT8e2083OQAdifF5lqXNRnOryn2LnyxFTpc6MGIiWtt84bVDAk 2tHw== X-Forwarded-Encrypted: i=1; AKwUvBxG0cjbGDKGba3GUhXg5TvpTHfkwbEDABlmlj6O4PVnKvO2IKGnzayDw53bQGg6JhkvgwKKwh2xNxF+YU8=@vger.kernel.org X-Gm-Message-State: AFuF++nkN3v5V80MPU1akQDxcfIvbJMY/ss63TNBfSPhTn48JFvKHBFs DF/TgV3P9iUqlf0w5jJqwId/QkaLcUW1/0NbS3Phf4VHrrFXNzaU/cZOfmPtbfharhE= X-Gm-Gg: AYBFou2wrwdQlG87Tph+QDllIXw8J2uVXqoxQtwDAyd2ReiD4/oLFclpfXj4SDjJquE cFZpuCk8DIhaP20QYuug0IVmf+IRGFfxrso3Q3HNUr1r0/eT/A1lJtBzBe0Y13dJJvFSuPF7Kyh +1ZOnZCcZSbuO1ceEu/Bx7bREpQ8oQXPR2AnQtuG5fIyiICWAWK6V/2Tl6ioe10aTzO1Uob8/iy QMnzDs1/Qi3Fz7TCsXlCTWnXRve5qXWHPLYihSUzWWlIy101YUeewOIvmToxRgLjqOJYLfYCaC/ CwIkUzeGsY0ttfppRUJ9FwXr61pezsb/m0FG+oCOl+y2J+thHPohwdy4XFxROj5ae0CBCocv7Yo Eg62YLz6lEEXR+4LaTTl05Bq/46cR4jTBqPIM7AGmSmUyc7gdeFdF9u/BpCujKLvAaHqdfSgQGh JMzGtNCSJ/xtUvslAeQW1QV2ZsKw72nL0fd0A+7EKUHI19Avh6VvbG+YWvWlksfXWCTg== X-Received: by 2002:a05:620a:2993:b0:93b:c22a:e98a with SMTP id af79cd13be357-93c2509e62cmr13040785a.2.1790097393108; Tue, 22 Sep 2026 10:16:33 -0700 (PDT) Received: from localhost ([2603:7001:f100:500:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c24756838sm24592485a.1.2026.09.22.10.16.32 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 10:16:32 -0700 (PDT) Date: Tue, 22 Sep 2026 13:16:23 -0400 From: Johannes Weiner To: Chris Li Cc: Gregory Price , Baoquan He , Nhat Pham , Kairui Song , Michal Hocko , Roman Gushchin , Shakeel Butt , Yosry Ahmed , David Hildenbrand , Muchun Song , Kemeng Shi , Barry Song , YoungJun Park , Chengming Zhou , "Lorenzo Stoakes (Oracle)" , "Liam R. Howlett" , "Vlastimil Babka (SUSE)" , Mike Rapoport , Suren =?utf-8?B?QmFnaGRhc2FyeWFu77+8?= , Qi Zheng , Axel Rasmussen , Yuanchu Xie , Wei Xu , Rik van Riel , Wenchao Hao , Jonathan Corbet , Hugh Dickins , Baolin Wang , Tejun Heo , Michal =?iso-8859-1?Q?Koutn=FD?= , Shuah Khan , Kunwu Chan , Meta kernel team , Linux Memory Management List , Linux Kernel Mailing List , linux-doc@vger.kernel.org, "open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG)" , Andrew Morton , Kairui Song , Joshua Hahn Subject: Re: Path forward for Virtualized Swap? Message-ID: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: On Sat, Sep 19, 2026 at 09:02:28AM -1000, Chris Li wrote: > On Sat, Sep 19, 2026 at 6:08 AM Gregory Price wrote: > > > > On Sat, Sep 19, 2026 at 03:45:03AM -0500, Chris Li wrote: > > > > > > I found this new concept of "compression space" very confusing to me. > > > Can you explain the swap behavior and problem using only normal memory > > > usage reduction and latency without introducing a new term or new > > > metrics? > > > > > > The normal user doesn't even know what compression space is, let alone > > > what makes it transparent. > > > > > > > Sure they do - it's the amount of memory consumed by compressed data, > > including the metadata associated with it. > > > > converting Johannes statement to diagram: > > > > >> Compression space is not a separate resource. It's page tables, > > >> backing pages, and swap descriptors. It's just MEMORY. > > > > Page Data (PD) > > [page tables][ uncompressed page ] > > > > > > Compressed Data (CD) > > [ recovered space ][pte][swap meta data][compressed page] > > | | > > |---------compression space----------| > > > > > > Memory Pre-Compression > > |[ PD ][ PD ][ PD ][ PD ][ PD ][ PD ][ PD ][ PD ]| > > > > > > Memory Post-Compression > > |[CD][CD][CD][CD][CD][CD][CD][CD]-------- free space ------------| > > ^----------------------------^ > > Compression Space > > Thanks for the explanation. So the compression space is just the > actual data store backing the zswap/xswap/zram. It's actually the pre-compression side I was referring to. The address space that vswap/xswap map. If you artificially hard limit this *address space*, you are making assumptions about (1) compression ratio and (2) how much non-residency the workload can tolerate. You might well get real workloads hitting that limit while you'd still have the ability to store more. Add pressure-driven writeback in the mix, and now even compression ratio assumptions are not useful: The stuff in zswap compresses a certain way and the stuff that was written back is out of memory completely. Yet it's all mapped by that same address space. How could anyone pick an informed size limit for this space? How many setups hard-limit the process virtual address space to 2xRAM? Nobody. They limit the physical memory required to back it. That's the only thing that actually matters. > > It's actually really confusing to represent this space as a traditional > > swap device - built on the assumption of a pre-defined size limit - when > > that size limit has already been defined (the memory itself). > > First of all, the traditional swap counter has a very well-defined > meaning. It is the size of the memory that, when accessed, requires a > page fault. That's 100% wrong. First of all, it wouldn't work because you can still take page faults on the filesystem. Second, this is absolutely not their intended purpose. They are to control access to a physically limited swapfile on disk. Because it is a discrete, separate resource from memory. That's the whole reason why memory and swap controls were split in cgroup2. I designed this. What you're talking about is residency guarantees. This is what memory.min and memory.low are for.