From: Amir Goldstein <amir73il@gmail.com>
To: NeilBrown <neil@brown.name>, Jori Koolstra <jkoolstra@xs4all.nl>
Cc: Christian Brauner <brauner@kernel.org>,
Jeff Layton <jlayton@kernel.org>,
Al Viro <viro@zeniv.linux.org.uk>,
Aleksa Sarai <aleksa@amutable.com>, Jan Kara <jack@suse.cz>,
linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org,
Theodore Tso <tytso@mit.edu>
Subject: Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
Date: Wed, 30 Sep 2026 11:45:26 +0200 [thread overview]
Message-ID: <CAOQ4uxis4tsea6fx+9iAffVX+qwBC9KzHg5GcP7ahthKd1T9KQ@mail.gmail.com> (raw)
In-Reply-To: <179072013439.37859.465558431050704411@noble.neil.brown.name>
On Wed, Sep 30, 2026 at 12:15 AM NeilBrown <neilb@ownmail.net> wrote:
>
> On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@ownmail.net>:
> > >
> > > >
> > > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > > dentry, but it doesn't any more. The reason was because
> > > > d_splice_alias() on might return a different dentry.
> > > > In this case we want the same dentry, but we need to do a lookup on it.
> > > >
> > > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > > >
> > > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > > I'm not certain that is needed if we keep the parent locked, but we
> > > > would need to be certain.
> > > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > > to ->mkdir.
> > > >
> > > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > > probably not.
> > > >
> > > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > > comment about kernfs. But we need to double-check that this will do the
> > > > right thing with ->d_time (I think it will).
> > > > We also need to think carefully about races with
> > > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > > ->lookup.
> > >
> > > I've thought a bit more about this ... I think that doing a lookup after
> > > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > > about the fs that I would rather not assume.
> > >
> > > I would rather have the current proposed code check for a negative
> > > dentry, and fail with -EIO or similar.
> > >
> >
> > I just noticed that there's precedent for this in overlayfs in super.c:
> >
> > /* Weird filesystem returning with hashed negative (kernfs)? */
> > err = -EINVAL;
> > if (d_really_is_negative(work))
> > goto out_dput;
> >
> > Shall we just do this for current kernel release, then we can add support
> > later if wanted.
> >
> > (But let's do EOPNOTSUPP instead of EINVAL)
> >
> > What do you think?
>
> The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> might create the directory, then return -EOPNOTSUPP. This is weird and
> I'd rather it not be visible.
>
> Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> remain a -EINVAL on any filesystem which doesn't completely support
> the functionality.
>
Joining late to this party so apologies in advance if my questions
have already been addressed.
I agree with Neil's statement above, but IMO, the atomic_open() fs
match the description of "doesn't completely support the functionality."
Therefore, I think that rather than success if directory exists, they
should also return -EINVAL/-EOPNOTSUPP consistently (see below).
> To do that we need some way to detect kernfs and tracefs. I think
> the only way we can do that is to make some change to those two
> filesystems.
> Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
>
Maybe, but then I think we should also exclude all the atomic_open fs.
Food for thought:
If you agree that tracefs/kernfs and <network>fs should have consistent
"API not supported" behavior, then gating the flag combination on
sb->s_d_flags & DCACHE_OP_REVALIDATE covers all the fs that we
want to exclude, plus some fs that we shouldn't really care about.
My agent tells me that the list is:
afs, coda, ecryptfs, exfat, hfs, jfs, ocfs2, orangefs, proc, ubifs, vfat
Two exceptions that we MAY care about:
1. overlayfs DCACHE_OP_REVALIDATE is derived from underlying
layers per dentry (If any of them have the flag) but always has it
in sb default_d_op, but it could technically derive sb->s_d_flags
layers at sb fill time, without any behavior change
2. encrypted/casefolded directories with fscrypt_d_revalidate (ext4/f2fs)
with CONFIG_FS_ENCRYPTION=y
ecrypted/casefold is the non common configuration, technically
sb->s_d_flags could be set according to sb features at sb fill time
without any behavior change
Thanks,
Amir.
next prev parent reply other threads:[~2026-09-30 9:45 UTC|newest]
Thread overview: 40+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-13 18:50 [PATCH v6 00/12] " Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 01/12] fs/namei.c: use trailing_slashes() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 02/12] vfs: prepare vfs_creat|mkdir_no_perm for reuse in lookup_open() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 03/12] vfs: lookup_open(): move setting FMODE_CREATED down Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 04/12] vfs: move ->create check in lookup_open() to before try_break_deleg() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 05/12] vfs: lookup_open(): use vfs_create_no_perm() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 06/12] vfs: lookup_open(): lock the parent as I_MUTEX_PARENT Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
2026-09-18 8:17 ` Christian Brauner
2026-09-18 10:07 ` NeilBrown
2026-09-19 0:01 ` NeilBrown
2026-09-25 13:57 ` Christian Brauner
2026-09-25 21:26 ` NeilBrown
2026-09-25 22:53 ` Jori Koolstra
2026-09-29 11:54 ` Jori Koolstra
2026-09-29 22:15 ` NeilBrown
2026-09-30 9:45 ` Amir Goldstein [this message]
2026-09-30 22:16 ` Jori Koolstra
2026-10-01 9:29 ` Amir Goldstein
2026-10-01 10:08 ` NeilBrown
2026-10-01 11:29 ` Amir Goldstein
2026-10-01 15:42 ` Jori Koolstra
2026-10-01 15:59 ` Amir Goldstein
2026-10-01 16:23 ` Jori Koolstra
2026-10-01 17:53 ` Amir Goldstein
2026-09-30 22:33 ` Jori Koolstra
2026-09-30 22:56 ` NeilBrown
2026-09-30 23:25 ` Jori Koolstra
2026-10-01 1:00 ` NeilBrown
2026-10-01 14:07 ` Jori Koolstra
2026-09-25 23:13 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno Jori Koolstra
2026-09-18 8:04 ` Christian Brauner
2026-09-25 22:55 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 09/12] vfs: move O_IS_MKDIR check from lookup_open() into individual filesystems Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 10/12] vfs: refuse O_CREAT for directories through a dangling symlink Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 11/12] vfs: short-circuit MAY_WRITE access for O_DIRECTORY opens Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 12/12] selftest: add tests for open*(O_CREAT|O_DIRECTORY) Jori Koolstra
2026-09-17 10:38 ` [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Christian Brauner
2026-09-18 8:18 ` Christian Brauner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=CAOQ4uxis4tsea6fx+9iAffVX+qwBC9KzHg5GcP7ahthKd1T9KQ@mail.gmail.com \
--to=amir73il@gmail.com \
--cc=aleksa@amutable.com \
--cc=brauner@kernel.org \
--cc=jack@suse.cz \
--cc=jkoolstra@xs4all.nl \
--cc=jlayton@kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=neil@brown.name \
--cc=tytso@mit.edu \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®