mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Amir Goldstein <amir73il@gmail.com>
To: NeilBrown <neil@brown.name>, Jori Koolstra <jkoolstra@xs4all.nl>
Cc: Christian Brauner <brauner@kernel.org>,
	Jeff Layton <jlayton@kernel.org>,
	 Al Viro <viro@zeniv.linux.org.uk>,
	Aleksa Sarai <aleksa@amutable.com>, Jan Kara <jack@suse.cz>,
	 linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org,
	 Theodore Tso <tytso@mit.edu>
Subject: Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
Date: Wed, 30 Sep 2026 11:45:26 +0200	[thread overview]
Message-ID: <CAOQ4uxis4tsea6fx+9iAffVX+qwBC9KzHg5GcP7ahthKd1T9KQ@mail.gmail.com> (raw)
In-Reply-To: <179072013439.37859.465558431050704411@noble.neil.brown.name>

On Wed, Sep 30, 2026 at 12:15 AM NeilBrown <neilb@ownmail.net> wrote:
>
> On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@ownmail.net>:
> > >
> > > >
> > > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > > dentry, but it doesn't any more.  The reason was because
> > > > d_splice_alias() on might return a different dentry.
> > > > In this case we want the same dentry, but we need to do a lookup on it.
> > > >
> > > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > > >
> > > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > > I'm not certain that is needed if we keep the parent locked, but we
> > > > would need to be certain.
> > > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > > to ->mkdir.
> > > >
> > > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > > probably not.
> > > >
> > > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > > comment about kernfs.  But we need to double-check that this will do the
> > > > right thing with ->d_time (I think it will).
> > > > We also need to think carefully about races with
> > > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > > ->lookup.
> > >
> > > I've thought a bit more about this ...  I think that doing a lookup after
> > > the vfs_mkdir() results in a negative is a bit ugly.  It assumes things
> > > about the fs that I would rather not assume.
> > >
> > > I would rather have the current proposed code check for a negative
> > > dentry, and fail with -EIO or similar.
> > >
> >
> > I just noticed that there's precedent for this in overlayfs in super.c:
> >
> >               /* Weird filesystem returning with hashed negative (kernfs)? */
> >               err = -EINVAL;
> >               if (d_really_is_negative(work))
> >                       goto out_dput;
> >
> > Shall we just do this for current kernel release, then we can add support
> > later if wanted.
> >
> > (But let's do EOPNOTSUPP instead of EINVAL)
> >
> > What do you think?
>
> The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> might create the directory, then return -EOPNOTSUPP.  This is weird and
> I'd rather it not be visible.
>
> Currently O_DIRECTORY|O_CREAT results in -EINVAL.  I would rather it
> remain a -EINVAL on any filesystem which doesn't completely support
> the functionality.
>

Joining late to this party so apologies in advance if my questions
have already been addressed.

I agree with Neil's statement above, but IMO, the atomic_open() fs
match the description of "doesn't completely support the functionality."
Therefore, I think that rather than success if directory exists, they
should also return -EINVAL/-EOPNOTSUPP consistently (see below).

> To do that we need some way to detect kernfs and tracefs.  I think
> the only way we can do that is to make some change to those two
> filesystems.
> Maybe a new  SB_I_ flag in sb->s_iflags would be ok in the short term.
>

Maybe, but then I think we should also exclude all the atomic_open fs.

Food for thought:

If you agree that tracefs/kernfs and <network>fs should have consistent
"API not supported" behavior, then gating the flag combination on
sb->s_d_flags & DCACHE_OP_REVALIDATE covers all the fs that we
want to exclude, plus some fs that we shouldn't really care about.

My agent tells me that the list is:
afs, coda, ecryptfs, exfat, hfs, jfs, ocfs2, orangefs, proc, ubifs, vfat

Two exceptions that we MAY care about:
1. overlayfs DCACHE_OP_REVALIDATE is derived from underlying
    layers per dentry (If any of them have the flag) but always has it
    in sb default_d_op, but it could technically derive sb->s_d_flags
    layers at sb fill time, without any behavior change
2. encrypted/casefolded directories with fscrypt_d_revalidate (ext4/f2fs)
    with CONFIG_FS_ENCRYPTION=y
    ecrypted/casefold is the non common configuration, technically
    sb->s_d_flags could be set according to sb features at sb fill time
    without any behavior change

Thanks,
Amir.

  reply	other threads:[~2026-09-30  9:45 UTC|newest]

Thread overview: 40+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-13 18:50 [PATCH v6 00/12] " Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 01/12] fs/namei.c: use trailing_slashes() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 02/12] vfs: prepare vfs_creat|mkdir_no_perm for reuse in lookup_open() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 03/12] vfs: lookup_open(): move setting FMODE_CREATED down Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 04/12] vfs: move ->create check in lookup_open() to before try_break_deleg() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 05/12] vfs: lookup_open(): use vfs_create_no_perm() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 06/12] vfs: lookup_open(): lock the parent as I_MUTEX_PARENT Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
2026-09-18  8:17   ` Christian Brauner
2026-09-18 10:07     ` NeilBrown
2026-09-19  0:01       ` NeilBrown
2026-09-25 13:57         ` Christian Brauner
2026-09-25 21:26           ` NeilBrown
2026-09-25 22:53             ` Jori Koolstra
2026-09-29 11:54         ` Jori Koolstra
2026-09-29 22:15           ` NeilBrown
2026-09-30  9:45             ` Amir Goldstein [this message]
2026-09-30 22:16               ` Jori Koolstra
2026-10-01  9:29                 ` Amir Goldstein
2026-10-01 10:08                   ` NeilBrown
2026-10-01 11:29                     ` Amir Goldstein
2026-10-01 15:42                   ` Jori Koolstra
2026-10-01 15:59                     ` Amir Goldstein
2026-10-01 16:23                       ` Jori Koolstra
2026-10-01 17:53                         ` Amir Goldstein
2026-09-30 22:33             ` Jori Koolstra
2026-09-30 22:56               ` NeilBrown
2026-09-30 23:25                 ` Jori Koolstra
2026-10-01  1:00                   ` NeilBrown
2026-10-01 14:07                     ` Jori Koolstra
2026-09-25 23:13     ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno Jori Koolstra
2026-09-18  8:04   ` Christian Brauner
2026-09-25 22:55     ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 09/12] vfs: move O_IS_MKDIR check from lookup_open() into individual filesystems Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 10/12] vfs: refuse O_CREAT for directories through a dangling symlink Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 11/12] vfs: short-circuit MAY_WRITE access for O_DIRECTORY opens Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 12/12] selftest: add tests for open*(O_CREAT|O_DIRECTORY) Jori Koolstra
2026-09-17 10:38 ` [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Christian Brauner
2026-09-18  8:18   ` Christian Brauner

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=CAOQ4uxis4tsea6fx+9iAffVX+qwBC9KzHg5GcP7ahthKd1T9KQ@mail.gmail.com \
    --to=amir73il@gmail.com \
    --cc=aleksa@amutable.com \
    --cc=brauner@kernel.org \
    --cc=jack@suse.cz \
    --cc=jkoolstra@xs4all.nl \
    --cc=jlayton@kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=neil@brown.name \
    --cc=tytso@mit.edu \
    --cc=viro@zeniv.linux.org.uk \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®