From: David Matlack <dmatlack@google.com>
To: Pratyush Yadav <pratyush@kernel.org>
Cc: Pasha Tatashin <pasha.tatashin@soleen.com>,
Mike Rapoport <rppt@kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
Alexander Graf <graf@amazon.com>, Hugh Dickins <hughd@google.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
Samiullah Khawaja <skhawaja@google.com>,
kexec@lists.infradead.org, linux-kernel@vger.kernel.org,
linux-mm@kvack.org
Subject: Re: [RFC PATCH 0/6] luo: tmpfs preservation
Date: Thu, 24 Sep 2026 21:10:03 +0000 [thread overview]
Message-ID: <arWRq4ZRbmyWK4SI@google.com> (raw)
In-Reply-To: <20260923224408.3745689-1-pratyush@kernel.org>
On 2026-09-24 12:43 AM, Pratyush Yadav wrote:
> From: "Pratyush Yadav (Google)" <pratyush@kernel.org>
>
> I brought this idea up in this week's Hypervisor Live Update bi-weekly.
> I decided to try using an LLM to see if it can produce a
> proof-of-concept quickly. This series is the end result.
>
> The main use case is preserving in-memory files that have a filesystem
> path. We already support memfd preservation, but memfds can't be linked
> to a (user-visible) filesystem. This is needed for storing VMM packages
> for live update on hosts that don't have a disk. A cold boot fetches the
> binaries from network, but that is too slow for a live update.
>
> David tried to solve the problem by introducing
> LIVEUPDATE_SESSION_RETRIEVE_INTO_FD [0], which lets you provide a FD for
> LUO to retrieve into. This is an alternative to the idea. It uses the
> standard preservation and retrieval API that LUO already provides.
>
> The core idea is to allow userspace to preserve a tmpfs mount FD. Once
> the mount is preserved, userspace can pass in regular files in that
> mount for preservation. The files take a dependency on the mount token,
> and that is used for retrieving the files in the right mount. This saves
> us from doing a full FS preservation and makes preservation of each file
> explicit.
>
> Currently only files in the root are supported. Files in subdirectories
> will be rejected. This is mainly for simplicity. Complex mount features
> like memory policies, id mappings, or casefolding are also not
> supported. All these can be reconfigured after retrieve if really
> needed.
>
> The code re-uses a lot of the preservation and retrieval logic from
> memfd preservation. It only adds some extra file and mount metadata on
> top.
>
> As I mentioned earlier, this is heavily LLM generated. The code is not
> very polished and has some rough edges. That said, I have read all the
> code and did significant cleanups of the LLM output. This includes
> turning the 1600 or so lines it generated to a more modest 977 lines.
> So (I think) it isn't complete AI garbage. And I think it does get the
> core idea across.
Hi Pratyush,
Thanks for posting this. I wanted to compare this series against the
RETRIEVE_INTO_FD approach, since both are trying to solve the same
problem: preserving files with filesystem paths across a Live Update.
Both are RFCs, so I'm going to ignore implementation details and focus
on the design.
As we discussed at the Hypervisor Live Update bi-weekly, this will be a
topic to discuss at LPC. I'm hoping we can use this thread to align on
the pros and cons of the two approaches so I don't unintentionally
mischaracterize anything.
At the memory level the two approaches are the same: file contents are
handed over using the existing memfd folio ABI and re-inserted into a
shmem inode in the new kernel. The difference is who owns the filesystem
topology around those folios.
- RETRIEVE_INTO_FD: The kernel preserves only memory. Userspace
own the fileystem topology, recreates it after kexec with normal
syscalls, and then asks LUO to fill each empty file with its
preserved contents.
- tmpfs handlers (this series): The kernel preserves the mount and
the files as first-class LUO objects (name, mode, pos, etc.),
rebuilds them after kexec, and hands back a detached mount.
Everything else below follows from that difference.
Comparison
==========
What LUO preserves:
RETRIEVE_INTO_FD: Memory only.
tmpfs handlers: Memory plus filesystem objects and metadata.
New uAPI:
RETRIEVE_INTO_FD: One new generic ioctl, or a generic extension to
the existing RETRIEVE ioctl.
tmpfs handlers: None. Existing PRESERVE/RETRIEVE with new file
types.
New LUO ABI:
RETRIEVE_INTO_FD: None
tmpfs handlers: Mount and file structs, plus a version bump for
every new attribute added in the future.
New VFS surface:
RETRIEVE_INTO_FD: None.
tmpfs handlers: Kernel-created mounts handed to userspace as
detached mounts.
Where the filesystem topology description lives:
RETRIEVE_INTO_FD: Userspace, in a private format it can evolve
freely.
tmpfs handlers: Kernel ABI.
Mount options:
RETRIEVE_INTO_FD: Anything userspace can pass to mount.
tmpfs handlers: Only what the ABI enumerates (currently the block
limit and root mode).
Dirs, hardlinks, symlinks, ownership, xattrs:
RETRIEVE_INTO_FD: Free. Userspace recreates them with normal
syscalls.
tmpfs handlers: Each requires new kernel code and ABI.
Userspace complexity:
RETRIEVE_INTO_FD: Higher. Userspace needs a manifest, then must
mkdir, create, chown, etc. before calling
retrieve_into.
tmpfs handlers: Low. Preserve the mount and files, retrieve them,
and move_mount().
Metadata consistency:
RETRIEVE_INTO_FD: Userspace must snapshot and keep its manifest
consistent with what it preserved.
tmpfs handlers: The kernel captures metadata at freeze time,
atomically with the contents.
Trust surface for previous-kernel data:
RETRIEVE_INTO_FD: Userspace validates its own manifest. The kernel
only validates the folio list.
tmpfs handlers: The kernel must validate names, modes, etc.
Generality:
RETRIEVE_INTO_FD: Applies to any handler where "fill a
userspace-created object" makes sense.
tmpfs handlers: tmpfs only.
Arguments for RETRIEVE_INTO_FD
==============================
1. It follows the principle that LUO should only preserve what
userspace cannot recreate. Names, modes, ownership, directories,
and mount options can all be recreated by userspace. Memory
contents cannot. This approach draws the line exactly there.
2. The kernel ABI stays small. Every attribute the tmpfs handlers
learn to preserve becomes a stable KHO ABI that has to be
maintained across kernel versions. Reaching parity with what users
will eventually want (subdirectories, ownership, xattrs, ACLs,
symlinks, hardlinks, huge=, mpol, quotas, idmaps) implies a long
series of ABI revisions. RETRIEVE_INTO_FD gets all of that for
free via existing syscalls.
3. Policy stays in userspace. The target file lives on a mount that
userspace configured, in a cgroup and namespace of its choosing.
The kernel never has to guess the right configuration for a
recreated object.
4. It is a reusable LUO primitive rather than a tmpfs feature.
"Retrieve into an object that userspace already set up" could be
useful for other handlers where the object must be created in a
specific context, e.g. HugeTLBfs files come to mind. The tmpfs
series adds a one-off pair of handlers.
5. It requires no new VFS surface. There are no kernel-created
mounts being handed to userspace, so the change stays contained
to LUO and memfd, which lowers the cost of getting it upstream.
Arguments for the tmpfs handlers
================================
1. It gives a filesystem-shaped abstraction from the kernel's side.
The kernel knows that the preserved files belong to a mount,
tracks that dependency, and returns a working filesystem. This
maps directly onto the stated use case of VMM binaries on hosts
without local storage.
2. Metadata is captured atomically. Name, mode, size, and pos are
captured at freeze time alongside the contents, so they cannot
drift. With RETRIEVE_INTO_FD, userspace must keep its manifest
correct across any renames or chmods that happen after it takes
its snapshot.
3. The restored mount can be hidden until it is ready. The mount
comes back detached, and userspace chooses when and where to
attach it with move_mount(). Nobody can observe a half-built tree.
4. There is less for userspace to do. For the simple case of a
single tmpfs mount with a few files, userspace needs very little
new code.
5. It leaves the door open for whole-tree preservation. In the
future the kernel could preserve an entire mount with a single
token by walking the tree, which RETRIEVE_INTO_FD cannot express.
Where each approach hurts
=========================
RETRIEVE_INTO_FD pushes work and correctness onto userspace, likely
requiring a shared userspace library or daemon. The kernel also has
no notion that a set of tokens formed a single filesystem. It just
sees unrelated files.
The tmpfs handlers are narrow in scope today, and every extension is
expensive. There is a real risk of slowly growing a partial
filesystem serializer in the kernel. It also couples LUO to VFS
concepts such as mounts, namespaces, and names.
A possible middle ground
========================
The two approaches are not mutually exclusive. We could use
RETRIEVE_INTO_FD as the core primitive for file contents, and provide
a shared userspace library for the common "rebuild this tmpfs" case.
Kernel-side mount preservation could then be added later, if and when
a concrete need comes up that userspace cannot meet, such as an
atomic snapshot or availability before userspace runs.
That keeps the kernel ABI limited to memory, which is the part only
the kernel can preserve, while leaving room for the more integrated
model if it turns out to be needed.
Thanks,
David
prev parent reply other threads:[~2026-09-24 21:10 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-23 22:43 Pratyush Yadav
2026-09-23 22:44 ` [RFC PATCH 1/6] liveupdate: luo_file: look up outgoing tokens by id Pratyush Yadav
2026-09-23 22:55 ` sashiko-bot
2026-09-24 13:26 ` Pratyush Yadav
2026-09-23 22:44 ` [RFC PATCH 2/6] shmem: add tmpfs_create_mount() to create tmpfs mounts internally Pratyush Yadav
2026-09-23 22:54 ` sashiko-bot
2026-09-23 22:44 ` [RFC PATCH 3/6] fs/namespace: Add vfs_open_detached_mount() Pratyush Yadav
2026-09-23 23:01 ` sashiko-bot
2026-09-23 22:44 ` [RFC PATCH 4/6] mm/memfd_luo: allow preserving a tmpfs mount Pratyush Yadav
2026-09-23 23:06 ` sashiko-bot
2026-09-23 22:44 ` [RFC PATCH 5/6] mm/memfd_luo: allow preserving a tmpfs file Pratyush Yadav
2026-09-23 23:24 ` sashiko-bot
2026-09-23 22:44 ` [RFC PATCH 6/6] selftests/liveupdate: add tmpfs kexec test Pratyush Yadav
2026-09-23 23:14 ` sashiko-bot
2026-09-24 21:10 ` David Matlack [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arWRq4ZRbmyWK4SI@google.com \
--to=dmatlack@google.com \
--cc=akpm@linux-foundation.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=graf@amazon.com \
--cc=hughd@google.com \
--cc=kexec@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=pasha.tatashin@soleen.com \
--cc=pratyush@kernel.org \
--cc=rppt@kernel.org \
--cc=skhawaja@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®