* [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
@ 2026-09-13 18:50 Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 01/12] fs/namei.c: use trailing_slashes() Jori Koolstra
` (12 more replies)
0 siblings, 13 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
Hi Christian/Neil,
Finally got back from holiday. I know Neil is also attempting to make
changes to the same code paths as the O_CREAT|O_DIRECTORY series touch,
so I hope, now that I have a bit more time, that I can push this
forwards before I have to do more rebasing.
This series implements new semantics for the O_CREAT|O_DIRECTORY flag
combination for open*(2): perform a mkdir and open the resulting
directory; return a pinning fd (which mkdir does not).
Most of the work happens in "vfs: add O_CREAT|O_DIRECTORY to open*(2),"
as may be expected. Before that there is some clean-up work and
preparation. The "vfs: short-circuit MAY_WRITE access for O_DIRECTORY
opens" patch is also worth paying extra attention to as it
short-circuits doomed opening of directories as writable. This check (to
prevent one from opening directories as writable) is currently done very
late in do_open(), after an inode has been obtained. However, when
introducing O_CREAT|O_DIRECTORY this is unacceptable as it would create
the directory and then still fail. This patch does however change some
user visible error codes.
Moreover, "vfs: change ->create/->mkdir operations unavailable errno"
also changes the ->create and ->mkdir unavailable errnos to -EOPNOTSUPP.
A previous attempt to do this more widely stranded before.[1] However,
that regression was only for vfat and does seem specific to changing the
return code from vfs_symlink(). We can drop this patch if needed, but it
would be _really_ ugly.
Changes from vn to v(n+1):
v5: https://lore.kernel.org/linux-fsdevel/20260823160706.358293-1-jkoolstra@xs4all.nl/
- Uniformize ->create and ->mkdir unavailable errnos to -EOPNOTSUPP.
- Return -EOPNOTSUPP instead of -ENOENT when O_CREAT|O_DIRECTORY is
not supported by a ->atomic_open filesystem.
- Added test cases for dangling symlink and sticky directory
behaviour, and a test that verifies O_TMPFILE|O_CREAT is still
-EINVAL.
- Split off "vfs: lookup_open(): lock the parent as I_MUTEX_PARENT"
as a separate patch.
- Fixed the issues pointed out by Brauner in v5.
v4: https://lore.kernel.org/linux-fsdevel/20260712175539.1565444-1-jkoolstra@xs4all.nl/
- rebased on 7.2, which includes my changes to auditing in
lookup_open() as well as Neil Brown's changes to the same function
for the implementation of vfs_lookup_open().
v3: https://lore.kernel.org/linux-fsdevel/20260704164149.3480051-1-jkoolstra@xs4all.nl/
- address the inconsistency in calling audit_inode_child() in
lookup_open() versus vfs_create() in a separate series.
- fclog.c selftest header fix moved to separate patch.
- add a "with trailing slash test" success test to the included
selftests.
- pass struct qstr by pointer to trailing_slashes() and add a comment
on what it does
- changed commit message of "vfs: add O_CREAT|O_DIRECTORY to
open*(2)" to address comments of Brauner
[1]: https://lore.kernel.org/linux-fsdevel/90228c39-04e7-41ef-ad91-84a8bb866650@sirena.org.uk/
Jori Koolstra (12):
fs/namei.c: use trailing_slashes()
vfs: prepare vfs_creat|mkdir_no_perm for reuse in lookup_open()
vfs: lookup_open(): move setting FMODE_CREATED down
vfs: move ->create check in lookup_open() to before try_break_deleg()
vfs: lookup_open(): use vfs_create_no_perm()
vfs: lookup_open(): lock the parent as I_MUTEX_PARENT
vfs: add O_CREAT|O_DIRECTORY to open*(2)
vfs: change ->create/->mkdir operations unavailable errno
vfs: move O_IS_MKDIR check from lookup_open() into individual
filesystems
vfs: refuse O_CREAT for directories through a dangling symlink
vfs: short-circuit MAY_WRITE access for O_DIRECTORY opens
selftest: add tests for open*(O_CREAT|O_DIRECTORY)
fs/9p/vfs_inode.c | 5 +
fs/9p/vfs_inode_dotl.c | 5 +
fs/ceph/file.c | 5 +
fs/fuse/dir.c | 5 +
fs/gfs2/inode.c | 5 +
fs/namei.c | 247 ++++++++++-----
fs/nfs/dir.c | 10 +
fs/open.c | 32 +-
fs/smb/client/dir.c | 5 +
fs/vboxsf/dir.c | 5 +
include/linux/fcntl.h | 6 +
include/uapi/asm-generic/errno.h | 2 +-
.../testing/selftests/filesystems/.gitignore | 1 +
tools/testing/selftests/filesystems/Makefile | 2 +-
.../filesystems/open_o_creat_o_dir.c | 296 ++++++++++++++++++
.../testing/selftests/filesystems/wrappers.h | 11 +
16 files changed, 554 insertions(+), 88 deletions(-)
create mode 100644 tools/testing/selftests/filesystems/open_o_creat_o_dir.c
base-commit: 2f0c1cf72f4682178506f513bbf015e591b1aa4a
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 01/12] fs/namei.c: use trailing_slashes()
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 02/12] vfs: prepare vfs_creat|mkdir_no_perm for reuse in lookup_open() Jori Koolstra
` (11 subsequent siblings)
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
There are several places in fs/namei.c that can use the
trailing_slashes() function to improve context. To allow this broader
use its signature is changed to take a struct qstr instead of a struct
nameidata.
Reviewed-by: NeilBrown <neil@brown.name>
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 28 +++++++++++++++-------------
1 file changed, 15 insertions(+), 13 deletions(-)
diff --git a/fs/namei.c b/fs/namei.c
index 20a6534ea3ef..ab1302b38f46 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -2781,9 +2781,16 @@ static const char *path_init(struct nameidata *nd, unsigned flags)
return s;
}
+static inline bool trailing_slashes(const struct qstr *last)
+{
+ /* last->len is set by hash_name() to the length of the current
+ * component ->name, terminating with '/' or a NUL character. */
+ return (bool)last->name[last->len];
+}
+
static inline const char *lookup_last(struct nameidata *nd)
{
- if (nd->last_type == LAST_NORM && nd->last.name[nd->last.len])
+ if (nd->last_type == LAST_NORM && trailing_slashes(&nd->last))
nd->flags |= LOOKUP_FOLLOW | LOOKUP_DIRECTORY;
return walk_component(nd, WALK_TRAILING);
@@ -4695,17 +4702,12 @@ struct file *vfs_lookup_open(struct path *parent, struct qstr *last,
}
EXPORT_SYMBOL_FOR_MODULES(vfs_lookup_open, "nfsd");
-static inline bool trailing_slashes(struct nameidata *nd)
-{
- return (bool)nd->last.name[nd->last.len];
-}
-
static struct dentry *lookup_fast_for_open(struct nameidata *nd, int open_flag)
{
struct dentry *dentry;
if (open_flag & O_CREAT) {
- if (trailing_slashes(nd))
+ if (trailing_slashes(&nd->last))
return ERR_PTR(-EISDIR);
/* Don't bother on an O_EXCL create */
@@ -4713,7 +4715,7 @@ static struct dentry *lookup_fast_for_open(struct nameidata *nd, int open_flag)
return NULL;
}
- if (trailing_slashes(nd))
+ if (trailing_slashes(&nd->last))
nd->flags |= LOOKUP_FOLLOW | LOOKUP_DIRECTORY;
dentry = lookup_fast(nd);
@@ -5087,7 +5089,7 @@ static struct dentry *filename_create(int dfd, struct filename *name,
* Do the final lookup. Suppress 'create' if there is a trailing
* '/', and a directory wasn't requested.
*/
- if (last.name[last.len] && !want_dir)
+ if (trailing_slashes(&last) && !want_dir)
create_flags &= ~LOOKUP_CREATE;
dentry = start_dirop(path->dentry, &last, reval_flag | create_flags);
if (IS_ERR(dentry))
@@ -5703,7 +5705,7 @@ int filename_unlinkat(int dfd, struct filename *name)
goto exit_drop_write;
/* Why not before? Because we want correct error value */
- if (unlikely(last.name[last.len])) {
+ if (unlikely(trailing_slashes(&last))) {
if (d_is_dir(dentry))
error = -EISDIR;
else
@@ -6305,16 +6307,16 @@ int filename_renameat2(int olddfd, struct filename *from,
if (flags & RENAME_EXCHANGE) {
if (!d_is_dir(rd.new_dentry)) {
error = -ENOTDIR;
- if (new_last.name[new_last.len])
+ if (trailing_slashes(&new_last))
goto exit_unlock;
}
}
/* unless the source is a directory trailing slashes give -ENOTDIR */
if (!d_is_dir(rd.old_dentry)) {
error = -ENOTDIR;
- if (old_last.name[old_last.len])
+ if (trailing_slashes(&old_last))
goto exit_unlock;
- if (!(flags & RENAME_EXCHANGE) && new_last.name[new_last.len])
+ if (!(flags & RENAME_EXCHANGE) && trailing_slashes(&new_last))
goto exit_unlock;
}
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 02/12] vfs: prepare vfs_creat|mkdir_no_perm for reuse in lookup_open()
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 01/12] fs/namei.c: use trailing_slashes() Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 03/12] vfs: lookup_open(): move setting FMODE_CREATED down Jori Koolstra
` (10 subsequent siblings)
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
To implement O_CREAT|O_DIRECTORY we will have to repeat some of the
logic that is now in vfs_mkdir() (e.g. do error checks in the same
order). Separate this out in vfs_mkdir_no_perm(), which does all the
non-permission related work of vfs_mkdir(). Permission checking for the
lookup_open() path is timed differently because we may just be doing an
open and no create. Similar considerations give rise to
vfs_create_no_perm().
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 77 +++++++++++++++++++++++++++++++++++++-----------------
1 file changed, 53 insertions(+), 24 deletions(-)
diff --git a/fs/namei.c b/fs/namei.c
index ab1302b38f46..d8401aa3100d 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -4166,6 +4166,24 @@ static inline umode_t vfs_prepare_mode(struct mnt_idmap *idmap,
return mode;
}
+static inline
+int vfs_create_no_perm(struct mnt_idmap *idmap, struct dentry *dentry,
+ umode_t mode, struct delegated_inode *di)
+{
+ struct inode *dir = d_inode(dentry->d_parent);
+ int error;
+
+ error = try_break_deleg(dir, LEASE_BREAK_DIR_CREATE, di);
+ if (error)
+ return error;
+
+ error = dir->i_op->create(idmap, dir, dentry, mode);
+ if (!error)
+ fsnotify_create(dir, dentry);
+
+ return error;
+}
+
/**
* vfs_create - create new file
* @idmap: idmap of the mount the inode was found from
@@ -4198,13 +4216,8 @@ int vfs_create(struct mnt_idmap *idmap, struct dentry *dentry, umode_t mode,
error = security_inode_create(dir, dentry, mode);
if (error)
return error;
- error = try_break_deleg(dir, LEASE_BREAK_DIR_CREATE, di);
- if (error)
- return error;
- error = dir->i_op->create(idmap, dir, dentry, mode);
- if (!error)
- fsnotify_create(dir, dentry);
- return error;
+
+ return vfs_create_no_perm(idmap, dentry, mode, di);
}
EXPORT_SYMBOL(vfs_create);
@@ -4418,6 +4431,7 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
dput(dentry);
dentry = ERR_PTR(error);
}
+
return dentry;
}
@@ -4544,6 +4558,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
dentry = res;
}
}
+
if (dentry->d_inode || !(op->open_flag & O_CREAT)) {
/*
* No need to create a file. If lookup returned a positive
@@ -5358,6 +5373,33 @@ SYSCALL_DEFINE3(mknod, const char __user *, filename, umode_t, mode, unsigned, d
return filename_mknodat(AT_FDCWD, name, mode, dev);
}
+static inline
+struct dentry *vfs_mkdir_no_perm(struct mnt_idmap *idmap, struct inode *dir,
+ struct dentry *dentry, umode_t mode,
+ struct delegated_inode *di)
+{
+ unsigned max_links = dir->i_sb->s_max_links;
+ struct dentry *de;
+ int error;
+
+ if (max_links && dir->i_nlink >= max_links)
+ return ERR_PTR(-EMLINK);
+
+ error = try_break_deleg(dir, LEASE_BREAK_DIR_CREATE, di);
+ if (error)
+ return ERR_PTR(error);
+
+ de = dir->i_op->mkdir(idmap, dir, dentry, mode);
+ if (IS_ERR(de))
+ return de;
+ if (de) {
+ dput(dentry);
+ dentry = de;
+ }
+ fsnotify_mkdir(dir, dentry);
+ return dentry;
+}
+
/**
* vfs_mkdir - create directory returning correct dentry if possible
* @idmap: idmap of the mount the inode was found from
@@ -5385,7 +5427,6 @@ struct dentry *vfs_mkdir(struct mnt_idmap *idmap, struct inode *dir,
struct delegated_inode *delegated_inode)
{
int error;
- unsigned max_links = dir->i_sb->s_max_links;
struct dentry *de;
error = may_create_dentry(idmap, dir, dentry);
@@ -5401,24 +5442,12 @@ struct dentry *vfs_mkdir(struct mnt_idmap *idmap, struct inode *dir,
if (error)
goto err;
- error = -EMLINK;
- if (max_links && dir->i_nlink >= max_links)
- goto err;
-
- error = try_break_deleg(dir, LEASE_BREAK_DIR_CREATE, delegated_inode);
- if (error)
+ de = vfs_mkdir_no_perm(idmap, dir, dentry, mode, delegated_inode);
+ if (IS_ERR(de)) {
+ error = PTR_ERR(de);
goto err;
-
- de = dir->i_op->mkdir(idmap, dir, dentry, mode);
- error = PTR_ERR(de);
- if (IS_ERR(de))
- goto err;
- if (de) {
- dput(dentry);
- dentry = de;
}
- fsnotify_mkdir(dir, dentry);
- return dentry;
+ return de;
err:
end_creating(dentry);
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 03/12] vfs: lookup_open(): move setting FMODE_CREATED down
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 01/12] fs/namei.c: use trailing_slashes() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 02/12] vfs: prepare vfs_creat|mkdir_no_perm for reuse in lookup_open() Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 04/12] vfs: move ->create check in lookup_open() to before try_break_deleg() Jori Koolstra
` (9 subsequent siblings)
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
In preparation for using vfs_create_no_perm() in lookup_open() we need
to move setting FMODE_CREATED on the file mode to either before or after
that call, as currently it is in the middle. If try_break_deleg() fails
it is currently not set, but vfs_create_no_perm() includes a
try_break_deleg().
Going up the call chain of lookup_open() we see that it is only used in
open_last_lookups() if no error is returned from lookup_open(), so we
can safely move it to after the filesystem create() call. This also
makes more sense when reading the code as you don't have to wonder what
the implications are of setting FMODE_CREATED before the create() call.
Reviewed-by: NeilBrown <neil@brown.name>
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/fs/namei.c b/fs/namei.c
index d8401aa3100d..f1cc3ce979ef 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -4580,7 +4580,6 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
if (error)
goto out_dput;
- file->f_mode |= FMODE_CREATED;
if (!dir_inode->i_op->create) {
error = -EACCES;
goto out_dput;
@@ -4589,6 +4588,8 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
error = dir_inode->i_op->create(idmap, dir_inode, dentry, mode);
if (error)
goto out_dput;
+
+ file->f_mode |= FMODE_CREATED;
out:
if (!IS_ERR(dentry)) {
if (file->f_mode & FMODE_CREATED)
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 04/12] vfs: move ->create check in lookup_open() to before try_break_deleg()
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (2 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 03/12] vfs: lookup_open(): move setting FMODE_CREATED down Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 05/12] vfs: lookup_open(): use vfs_create_no_perm() Jori Koolstra
` (8 subsequent siblings)
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
The i_op->create check in lookup_open() takes place after the
try_break_deleg() call. This does not match the order when doing a
regular file create via mknod(2). There the call order is:
filename_mknodat()
vfs_create()
i_op->create check
try_break_deleg()
Move the i_op->create check to before try_break_deleg() in
lookup_open().
Reviewed-by: NeilBrown <neil@brown.name>
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/fs/namei.c b/fs/namei.c
index f1cc3ce979ef..e1cdece42c9b 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -4576,15 +4576,15 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
goto out_dput;
}
- error = try_break_deleg(dir_inode, LEASE_BREAK_DIR_CREATE, &delegated_inode);
- if (error)
- goto out_dput;
-
if (!dir_inode->i_op->create) {
error = -EACCES;
goto out_dput;
}
+ error = try_break_deleg(dir_inode, LEASE_BREAK_DIR_CREATE, &delegated_inode);
+ if (error)
+ goto out_dput;
+
error = dir_inode->i_op->create(idmap, dir_inode, dentry, mode);
if (error)
goto out_dput;
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 05/12] vfs: lookup_open(): use vfs_create_no_perm()
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (3 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 04/12] vfs: move ->create check in lookup_open() to before try_break_deleg() Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 06/12] vfs: lookup_open(): lock the parent as I_MUTEX_PARENT Jori Koolstra
` (7 subsequent siblings)
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
We can replace the code in the no create_error/negative dentry found
from lookup case in lookup_open() with the vfs_create_no_perm() helper.
Reviewed-by: NeilBrown <neil@brown.name>
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 25 +++++++------------------
1 file changed, 7 insertions(+), 18 deletions(-)
diff --git a/fs/namei.c b/fs/namei.c
index e1cdece42c9b..90c8be0257be 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -4430,8 +4430,14 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
}
dput(dentry);
dentry = ERR_PTR(error);
+ } else {
+ if (file->f_mode & FMODE_CREATED)
+ fsnotify_create(dir_inode, dentry);
+ if (file->f_mode & FMODE_OPENED)
+ fsnotify_open(file);
}
+
return dentry;
}
@@ -4581,22 +4587,12 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
goto out_dput;
}
- error = try_break_deleg(dir_inode, LEASE_BREAK_DIR_CREATE, &delegated_inode);
- if (error)
- goto out_dput;
-
- error = dir_inode->i_op->create(idmap, dir_inode, dentry, mode);
+ error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
if (error)
goto out_dput;
file->f_mode |= FMODE_CREATED;
out:
- if (!IS_ERR(dentry)) {
- if (file->f_mode & FMODE_CREATED)
- fsnotify_create(dir_inode, dentry);
- if (file->f_mode & FMODE_OPENED)
- fsnotify_open(file);
- }
if ((open_flag & O_CREAT) || create_error)
inode_unlock(dir_inode);
else
@@ -5216,13 +5212,6 @@ struct file *dentry_create(struct path *path, int flags, umode_t mode,
/* Drop the extra reference */
dput(orig_dentry);
- if (!error) {
- if (file->f_mode & FMODE_CREATED)
- fsnotify_create(dir->d_inode, dentry);
- if (file->f_mode & FMODE_OPENED)
- fsnotify_open(file);
- }
-
path->dentry = dentry;
} else {
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 06/12] vfs: lookup_open(): lock the parent as I_MUTEX_PARENT
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (4 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 05/12] vfs: lookup_open(): use vfs_create_no_perm() Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (6 subsequent siblings)
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
lookup_open() takes the parent directory's i_rwsem with inode_lock(),
that is, in the I_MUTEX_NORMAL class, before creating the last
component. Every other VFS path that locks a directory in order to
create something in it goes through start_dirop() and friends, which
use inode_lock_nested(dir, I_MUTEX_PARENT).
Lock the parent as I_MUTEX_PARENT here too, so that lookup_open()
matches the rest of the VFS before it starts calling ->mkdir(), which
until now was only ever reached with the parent held that way.
No functional change intended.
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/fs/namei.c b/fs/namei.c
index 90c8be0257be..0efd395a1a65 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -4482,7 +4482,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
*/
}
if (open_flag & O_CREAT)
- inode_lock(dir_inode);
+ inode_lock_nested(dir_inode, I_MUTEX_PARENT);
else
inode_lock_shared(dir_inode);
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (5 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 06/12] vfs: lookup_open(): lock the parent as I_MUTEX_PARENT Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-18 8:17 ` Christian Brauner
2026-09-13 18:50 ` [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno Jori Koolstra
` (5 subsequent siblings)
12 siblings, 1 reply; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
Currently there is no way to race-freely create and open a directory.
For regular files we have open(O_CREAT) for creating a new file inode,
and returning a pinning fd to it. The lack of such functionality for
directories means that when populating a directory tree there's always
a race involved: the inodes first need to be created, and then opened
to adjust their permissions/ownership/labels/timestamps/acls/xattrs/...,
but in the time window between the creation and the opening they might
be replaced by something else.
Addressing this race without a proper API is only partially possible:
the caller can immediately fstat() what was opened to verify that it
has the expected inode type, owner and mode. But besides being easy to
get wrong, this cannot establish who created the directory: a directory
created by another process with identical credentials is
indistinguishable from one the caller created itself, so the caller
cannot tell whether the directory is its own to manage.
Historically, the O_CREAT|O_DIRECTORY behaviour was to return ENOTDIR if
a regular file exists at the open path; EISDIR if a directory exists at
the path; and to create a regular file if no file exists at the path.
This behaviour changed accidentally with
commit 973d4b73fbaf ("do_last(): rejoin the common path even earlier in
FMODE_{OPENED,CREATED} case") causing ENOTDIR to return in the last case
while still creating the file. As this change was not detected for a
long time, Brauner proposed to adopt the more consistent NetBSD
behaviour, i.e. to return EINVAL on the O_CREAT|O_DIRECTORY combination.
This change was applied in commit 43b450632676 ("open: return EINVAL for
O_DIRECTORY | O_CREAT") in March, 2023. As the EINVAL behaviour has been
in the kernel for about 3 years now, no rollback is expected as a result
of userspace reliance on old behaviour, leaving us free to reassign the
O_CREAT|O_DIRECTORY semantics.
O_CREAT|O_DIRECTORY is made to reduce to a lookup on ->atomic_open()
filesystems. These filesystems currently cannot handle
O_CREAT|O_DIRECTORY without protocol extensions and therefore are forced
into a fallback mode by stripping the O_CREAT bit. This causes existing
directories to be successfully opened, while for targets that should
have been created, -ENOENT is returned. This -ENOENT is then converted
to -EOPNOTSUPP in later atomic_open(). The simple option of just
returning -EOPNOTSUPP directly leads to inconsistent behaviour: before
->atomic_open() is called in lookup_open(), the dcache is queried. So
returning -EOPNOTSUPP immediately would make O_CREAT|O_DIRECTORY
dependent on the cache state of the dentry.
There is no separate sysctl for directory creation implemented currently.
Therefore, for the S_ISDIR case, disabling sysctl_protected_regular is
not enough to allow creating a directory in a sticky folder, because that
may surprise users not expecting that O_CREAT|O_DIRECTORY is possible on
newer kernels.
This feature idea (and some of its description) is taken from the
UAPI group:
https://github.com/uapi-group/kernel-features?tab=readme-ov-file#race-free-creation-and-opening-of-non-file-inodes
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 116 +++++++++++++++++++++++++++++++++++-------
fs/open.c | 25 +++++----
include/linux/fcntl.h | 6 +++
3 files changed, 117 insertions(+), 30 deletions(-)
diff --git a/fs/namei.c b/fs/namei.c
index 0efd395a1a65..6ff0a3c04f02 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -1382,13 +1382,13 @@ int may_linkat(struct mnt_idmap *idmap, const struct path *link)
/**
* may_create_in_sticky - Check whether an O_CREAT open in a sticky directory
- * should be allowed, or not, on files that already
- * exist.
+ * should be allowed, or not, on files/directories that
+ * already exist.
* @idmap: idmap of the mount the inode was found from
* @nd: nameidata pathwalk data
* @inode: the inode of the file to open
*
- * Block an O_CREAT open of a FIFO (or a regular file) when:
+ * Block an O_CREAT open of a FIFO (or a regular file/directory) when:
* - sysctl_protected_fifos (or sysctl_protected_regular) is enabled
* - the file already exists
* - we are in a sticky directory
@@ -1416,6 +1416,14 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
if (likely(!(dir_mode & S_ISVTX)))
return 0;
+ /*
+ * There is no separate sysctl for directory creation in sticky
+ * folders. Therefore, for the S_ISDIR case, disabling
+ * sysctl_protected_regular is not enough to allow creating a
+ * directory in a sticky folder, because that may surprise users
+ * not expecting that O_CREAT|O_DIRECTORY is possible on newer
+ * kernels.
+ */
if (S_ISREG(inode->i_mode) && !sysctl_protected_regular)
return 0;
@@ -1447,6 +1455,12 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
"sticky_create_regular");
return -EACCES;
}
+
+ if (S_ISDIR(inode->i_mode)) {
+ audit_log_path_denied(AUDIT_ANOM_CREAT,
+ "sticky_create_dir");
+ return -EACCES;
+ }
}
return 0;
@@ -4334,21 +4348,43 @@ static inline int open_to_namei_flags(int flag)
static int may_o_create(struct mnt_idmap *idmap,
const struct path *dir, struct dentry *dentry,
- umode_t mode)
+ int open_flag, umode_t mode)
{
- int error = security_path_mknod(dir, dentry, mode, 0);
+ struct inode *dir_inode = dir->dentry->d_inode;
+ bool create_dir = O_IS_MKDIR(open_flag);
+ int error;
+
+ WARN_ON_ONCE(create_dir && !(mode & S_IFDIR));
+
+ if (create_dir)
+ error = security_path_mkdir(dir, dentry, mode);
+ else
+ error = security_path_mknod(dir, dentry, mode, 0);
if (error)
return error;
if (!fsuidgid_has_mapping(dir->dentry->d_sb, idmap))
return -EOVERFLOW;
- error = inode_permission(idmap, dir->dentry->d_inode,
- MAY_WRITE | MAY_EXEC);
+ error = inode_permission(idmap, dir_inode, MAY_WRITE | MAY_EXEC);
if (error)
return error;
- return security_inode_create(dir->dentry->d_inode, dentry, mode);
+ if (create_dir)
+ error = security_inode_mkdir(dir_inode, dentry, mode);
+ else
+ error = security_inode_create(dir_inode, dentry, mode);
+
+ return error;
+}
+
+static inline umode_t o_create_mode(struct mnt_idmap *idmap,
+ const struct inode *dir, int open_flag, umode_t mode)
+{
+ if (O_IS_MKDIR(open_flag))
+ return vfs_prepare_mode(idmap, dir, mode, S_IRWXUGO | S_ISVTX, S_IFDIR);
+ else
+ return vfs_prepare_mode(idmap, dir, mode, S_IALLUGO, S_IFREG);
}
/**
@@ -4384,8 +4420,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
file->__f_path.dentry = DENTRY_NOT_SET;
file->__f_path.mnt = path->mnt;
+
error = dir_inode->i_op->atomic_open(dir_inode, dentry, file,
- open_to_namei_flags(open_flag), mode);
+ open_to_namei_flags(open_flag), mode);
d_lookup_done(dentry);
if (!error) {
@@ -4427,12 +4464,32 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
*/
audit_inode_child(dir_inode, dentry, AUDIT_TYPE_CHILD_CREATE);
error = create_error;
+ } else if (O_IS_MKDIR(open_flag) && error == -ENOENT) {
+ /*
+ * If the underlying filesystem does not implement
+ * O_CREAT|O_DIRECTORY, it strips the O_CREAT bit and
+ * continues as a lookup. We can't simply return
+ * -EOPNOTSUPP from unsupported ->atomic_open()
+ * implementations because the dentry might be in the
+ * dcache. In that case, lookup_open() returns before
+ * reaching ->atomic_open(), and hence whether you get
+ * -EOPNOTSUPP on O_CREAT|O_DIRECTORY would not only
+ * depend on the underlying filesystem, but also on
+ * the state of the dcache. Still, we must make an
+ * effort to differentiate a regular -ENOENT from the
+ * unsupported O_CREAT|O_DIRECTORY case.
+ */
+ error = -EOPNOTSUPP;
}
dput(dentry);
dentry = ERR_PTR(error);
} else {
- if (file->f_mode & FMODE_CREATED)
- fsnotify_create(dir_inode, dentry);
+ if (file->f_mode & FMODE_CREATED) {
+ if (d_is_dir(dentry))
+ fsnotify_mkdir(dir_inode, dentry);
+ else
+ fsnotify_create(dir_inode, dentry);
+ }
if (file->f_mode & FMODE_OPENED)
fsnotify_open(file);
}
@@ -4441,6 +4498,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
return dentry;
}
+static inline
+struct dentry *vfs_mkdir_no_perm(struct mnt_idmap *, struct inode *, struct dentry *,
+ umode_t, struct delegated_inode *);
/*
* Look up and maybe create and open the last component.
*
@@ -4462,6 +4522,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
struct mnt_idmap *idmap;
struct dentry *dir = nd->path.dentry;
struct inode *dir_inode = dir->d_inode;
+ bool create_dir = O_IS_MKDIR(op->open_flag);
int open_flag;
struct dentry *dentry;
int error, create_error;
@@ -4474,6 +4535,9 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
mode = op->mode;
create_error = 0;
+ if (create_dir && dir_inode->i_op->atomic_open)
+ open_flag &= ~O_CREAT;
+
if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) {
got_write = !mnt_want_write(nd->path.mnt);
/*
@@ -4534,10 +4598,10 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
if (open_flag & O_CREAT) {
if (open_flag & O_EXCL)
open_flag &= ~O_TRUNC;
- mode = vfs_prepare_mode(idmap, dir_inode, mode, mode, mode);
+ mode = o_create_mode(idmap, dir_inode, open_flag, mode);
if (likely(got_write))
create_error = may_o_create(idmap, &nd->path,
- dentry, mode);
+ dentry, open_flag, mode);
else
create_error = -EROFS;
}
@@ -4582,12 +4646,25 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
goto out_dput;
}
- if (!dir_inode->i_op->create) {
+ /* mimic operation missing errnos of vfs_mkdir/vfs_create */
+ if (create_dir && !dir_inode->i_op->mkdir) {
+ error = -EPERM;
+ goto out_dput;
+ }
+ if (!create_dir && !dir_inode->i_op->create) {
error = -EACCES;
goto out_dput;
}
- error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
+ if (create_dir) {
+ struct dentry *res = vfs_mkdir_no_perm(idmap, dir_inode, dentry, mode,
+ &delegated_inode);
+ error = PTR_ERR_OR_ZERO(res);
+ if (!error)
+ dentry = res;
+ } else {
+ error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
+ }
if (error)
goto out_dput;
@@ -4719,7 +4796,7 @@ static struct dentry *lookup_fast_for_open(struct nameidata *nd, int open_flag)
struct dentry *dentry;
if (open_flag & O_CREAT) {
- if (trailing_slashes(&nd->last))
+ if (trailing_slashes(&nd->last) && !(open_flag & O_DIRECTORY))
return ERR_PTR(-EISDIR);
/* Don't bother on an O_EXCL create */
@@ -4820,8 +4897,9 @@ static int do_open(struct nameidata *nd,
if (open_flag & O_CREAT) {
if ((open_flag & O_EXCL) && !(file->f_mode & FMODE_CREATED))
return -EEXIST;
- if (d_is_dir(nd->path.dentry))
+ if (!(open_flag & O_DIRECTORY) && d_is_dir(nd->path.dentry))
return -EISDIR;
+
error = may_create_in_sticky(idmap, nd,
d_backing_inode(nd->path.dentry));
if (unlikely(error))
@@ -5159,7 +5237,7 @@ inline struct dentry *start_creating_user_path(
EXPORT_SYMBOL(start_creating_user_path);
/**
- * dentry_create - Create and open a file
+ * dentry_create - Create and open a regular file
* @path: path to create
* @flags: O\_ flags
* @mode: mode bits for new file
@@ -5196,7 +5274,7 @@ struct file *dentry_create(struct path *path, int flags, umode_t mode,
path->dentry = dir;
mode = vfs_prepare_mode(idmap, dir_inode, mode, S_IALLUGO, S_IFREG);
- create_error = may_o_create(idmap, path, dentry, mode);
+ create_error = may_o_create(idmap, path, dentry, flags, mode);
if (create_error)
flags &= ~O_CREAT;
diff --git a/fs/open.c b/fs/open.c
index 6b1c14e684a9..189af02a2425 100644
--- a/fs/open.c
+++ b/fs/open.c
@@ -1239,29 +1239,30 @@ inline int build_open_flags(const struct open_how *how, struct open_flags *op)
if (WILL_CREATE(flags)) {
if (how->mode & ~S_IALLUGO)
return -EINVAL;
- op->mode = how->mode | S_IFREG;
+ if (O_IS_MKDIR(flags))
+ op->mode = how->mode | S_IFDIR;
+ else
+ op->mode = how->mode | S_IFREG;
} else {
if (how->mode != 0)
return -EINVAL;
op->mode = 0;
}
- /*
- * Block bugs where O_DIRECTORY | O_CREAT created regular files.
- * Note, that blocking O_DIRECTORY | O_CREAT here also protects
- * O_TMPFILE below which requires O_DIRECTORY being raised.
- */
- if ((flags & (O_DIRECTORY | O_CREAT)) == (O_DIRECTORY | O_CREAT))
- return -EINVAL;
-
/* Now handle the creative implementation of O_TMPFILE. */
if (flags & __O_TMPFILE) {
/*
* In order to ensure programs get explicit errors when trying
* to use O_TMPFILE on old kernels we enforce that O_DIRECTORY
- * is raised alongside __O_TMPFILE.
+ * is raised alongside __O_TMPFILE, but without O_CREAT. The
+ * reason for disallowing O_CREAT|O_TMPFILE is that
+ * O_DIRECTORY|O_CREAT used to work and created a regular file
+ * if nothing existed at the open path. Hence, allowing the
+ * combination would have caused O_CREAT|O_TMPFILE to create a
+ * regular (non-temporary) file on old kernels, while the caller
+ * would believe they created an actual O_TMPFILE.
*/
- if (!(flags & O_DIRECTORY))
+ if (!(flags & O_DIRECTORY) || (flags & O_CREAT))
return -EINVAL;
if (!(acc_mode & MAY_WRITE))
return -EINVAL;
@@ -1319,6 +1320,8 @@ inline int build_open_flags(const struct open_how *how, struct open_flags *op)
op->intent = flags & O_PATH ? 0 : LOOKUP_OPEN;
if (flags & O_CREAT) {
+ if ((flags & O_DIRECTORY) && (acc_mode & MAY_WRITE))
+ return -EISDIR;
op->intent |= LOOKUP_CREATE;
if (flags & O_EXCL) {
op->intent |= LOOKUP_EXCL;
diff --git a/include/linux/fcntl.h b/include/linux/fcntl.h
index 6ad6b9e7a226..204e16bbe263 100644
--- a/include/linux/fcntl.h
+++ b/include/linux/fcntl.h
@@ -30,6 +30,12 @@
*/
#define __O_REGULAR (1 << 30)
+#define O_MKDIR_MASK (O_CREAT | O_DIRECTORY)
+static inline bool O_IS_MKDIR(unsigned int flags)
+{
+ return (flags & O_MKDIR_MASK) == O_MKDIR_MASK;
+}
+
/* List of all valid flags for the how->resolve argument: */
#define VALID_RESOLVE_FLAGS \
(RESOLVE_NO_XDEV | RESOLVE_NO_MAGICLINKS | RESOLVE_NO_SYMLINKS | \
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (6 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-18 8:04 ` Christian Brauner
2026-09-13 18:50 ` [PATCH v6 09/12] vfs: move O_IS_MKDIR check from lookup_open() into individual filesystems Jori Koolstra
` (4 subsequent siblings)
12 siblings, 1 reply; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
Currently filesystems return -EACCES for missing ->create and -EPERM for
missing ->mkdir in vfs_create/vfs_mkdir. Instead of duplicating these
dubious and inconsistent error codes to lookup_open(), change all to
-EOPNOTSUPP.
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 14 +++++---------
include/uapi/asm-generic/errno.h | 2 +-
2 files changed, 6 insertions(+), 10 deletions(-)
diff --git a/fs/namei.c b/fs/namei.c
index 6ff0a3c04f02..f31046f8b3ba 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -4224,7 +4224,7 @@ int vfs_create(struct mnt_idmap *idmap, struct dentry *dentry, umode_t mode,
return error;
if (!dir->i_op->create)
- return -EACCES; /* shouldn't it be ENOSYS? */
+ return -EOPNOTSUPP;
mode = vfs_prepare_mode(idmap, dir, mode, S_IALLUGO, S_IFREG);
error = security_inode_create(dir, dentry, mode);
@@ -4646,13 +4646,9 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
goto out_dput;
}
- /* mimic operation missing errnos of vfs_mkdir/vfs_create */
- if (create_dir && !dir_inode->i_op->mkdir) {
- error = -EPERM;
- goto out_dput;
- }
- if (!create_dir && !dir_inode->i_op->create) {
- error = -EACCES;
+ if ((create_dir && !dir_inode->i_op->mkdir)
+ || (!create_dir && !dir_inode->i_op->create)) {
+ error = -EOPNOTSUPP;
goto out_dput;
}
@@ -5501,7 +5497,7 @@ struct dentry *vfs_mkdir(struct mnt_idmap *idmap, struct inode *dir,
if (error)
goto err;
- error = -EPERM;
+ error = -EOPNOTSUPP;
if (!dir->i_op->mkdir)
goto err;
diff --git a/include/uapi/asm-generic/errno.h b/include/uapi/asm-generic/errno.h
index bd78e69e0a43..c84ebf89c8b6 100644
--- a/include/uapi/asm-generic/errno.h
+++ b/include/uapi/asm-generic/errno.h
@@ -76,7 +76,7 @@
#define ENOPROTOOPT 92 /* Protocol not available */
#define EPROTONOSUPPORT 93 /* Protocol not supported */
#define ESOCKTNOSUPPORT 94 /* Socket type not supported */
-#define EOPNOTSUPP 95 /* Operation not supported on transport endpoint */
+#define EOPNOTSUPP 95 /* Operation not supported */
#define EPFNOSUPPORT 96 /* Protocol family not supported */
#define EAFNOSUPPORT 97 /* Address family not supported by protocol */
#define EADDRINUSE 98 /* Address already in use */
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 09/12] vfs: move O_IS_MKDIR check from lookup_open() into individual filesystems
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (7 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 10/12] vfs: refuse O_CREAT for directories through a dangling symlink Jori Koolstra
` (3 subsequent siblings)
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
Individual filesystems that implement ->atomic_open() need to get the
chance to implement O_CREAT|O_DIRECTORY or not, rather than decide
this at the VFS level in lookup_open().
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/9p/vfs_inode.c | 5 +++++
fs/9p/vfs_inode_dotl.c | 5 +++++
fs/ceph/file.c | 5 +++++
fs/fuse/dir.c | 5 +++++
fs/gfs2/inode.c | 5 +++++
fs/namei.c | 3 ---
fs/nfs/dir.c | 10 ++++++++++
fs/smb/client/dir.c | 5 +++++
fs/vboxsf/dir.c | 5 +++++
9 files changed, 45 insertions(+), 3 deletions(-)
diff --git a/fs/9p/vfs_inode.c b/fs/9p/vfs_inode.c
index 3829554ca369..dd810904ff6c 100644
--- a/fs/9p/vfs_inode.c
+++ b/fs/9p/vfs_inode.c
@@ -776,6 +776,11 @@ v9fs_vfs_atomic_open(struct inode *dir, struct dentry *dentry,
struct inode *inode;
int p9_omode;
+ if (O_IS_MKDIR(flags)) {
+ flags &= ~O_CREAT;
+ mode = 0;
+ }
+
if (d_in_lookup(dentry)) {
struct dentry *res = v9fs_vfs_lookup(dir, dentry, 0);
if (res || d_really_is_positive(dentry))
diff --git a/fs/9p/vfs_inode_dotl.c b/fs/9p/vfs_inode_dotl.c
index 116b29e95f21..9308184aee61 100644
--- a/fs/9p/vfs_inode_dotl.c
+++ b/fs/9p/vfs_inode_dotl.c
@@ -238,6 +238,11 @@ v9fs_vfs_atomic_open_dotl(struct inode *dir, struct dentry *dentry,
struct v9fs_session_info *v9ses;
struct posix_acl *pacl = NULL, *dacl = NULL;
+ if (O_IS_MKDIR(flags)) {
+ flags &= ~O_CREAT;
+ omode = 0;
+ }
+
if (d_in_lookup(dentry)) {
struct dentry *res = v9fs_vfs_lookup(dir, dentry, 0);
if (res || d_really_is_positive(dentry))
diff --git a/fs/ceph/file.c b/fs/ceph/file.c
index bd3e3f5c269e..9235143edd9a 100644
--- a/fs/ceph/file.c
+++ b/fs/ceph/file.c
@@ -812,6 +812,11 @@ int ceph_atomic_open(struct inode *dir, struct dentry *dentry,
dir, ceph_vinop(dir), dentry, dentry,
d_unhashed(dentry) ? "unhashed" : "hashed", flags, mode);
+ if (O_IS_MKDIR(flags)) {
+ flags &= ~O_CREAT;
+ mode = 0;
+ }
+
if (dentry->d_name.len > NAME_MAX)
return -ENAMETOOLONG;
diff --git a/fs/fuse/dir.c b/fs/fuse/dir.c
index e49b4e874b15..a3e7daba61dc 100644
--- a/fs/fuse/dir.c
+++ b/fs/fuse/dir.c
@@ -944,6 +944,11 @@ static int fuse_atomic_open(struct inode *dir, struct dentry *entry,
struct mnt_idmap *idmap = file_mnt_idmap(file);
struct fuse_conn *fc = get_fuse_conn(dir);
+ if (O_IS_MKDIR(flags)) {
+ flags &= ~O_CREAT;
+ mode = 0;
+ }
+
if (fuse_is_bad(dir))
return -EIO;
diff --git a/fs/gfs2/inode.c b/fs/gfs2/inode.c
index f361876c5583..69e100a68e2c 100644
--- a/fs/gfs2/inode.c
+++ b/fs/gfs2/inode.c
@@ -1386,6 +1386,11 @@ static int gfs2_atomic_open(struct inode *dir, struct dentry *dentry,
{
bool excl = !!(flags & O_EXCL);
+ if (O_IS_MKDIR(flags)) {
+ flags &= ~O_CREAT;
+ mode = 0;
+ }
+
if (d_in_lookup(dentry)) {
struct dentry *d = __gfs2_lookup(dir, dentry, file);
if (file->f_mode & FMODE_OPENED) {
diff --git a/fs/namei.c b/fs/namei.c
index f31046f8b3ba..a990c9c8bddf 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -4535,9 +4535,6 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
mode = op->mode;
create_error = 0;
- if (create_dir && dir_inode->i_op->atomic_open)
- open_flag &= ~O_CREAT;
-
if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) {
got_write = !mnt_want_write(nd->path.mnt);
/*
diff --git a/fs/nfs/dir.c b/fs/nfs/dir.c
index 49394123bd09..4b1404d6859e 100644
--- a/fs/nfs/dir.c
+++ b/fs/nfs/dir.c
@@ -2121,6 +2121,11 @@ int nfs_atomic_open(struct inode *dir, struct dentry *dentry,
dfprintk(VFS, "NFS: atomic_open(%s/%llu), %pd\n",
dir->i_sb->s_id, dir->i_ino, dentry);
+ if (O_IS_MKDIR(open_flags)) {
+ open_flags &= ~O_CREAT;
+ mode = 0;
+ }
+
err = nfs_check_flags(open_flags);
if (err)
return err;
@@ -2317,6 +2322,11 @@ int nfs_atomic_open_v23(struct inode *dir, struct dentry *dentry,
*/
int error = 0;
+ if (O_IS_MKDIR(open_flags)) {
+ open_flags &= ~O_CREAT;
+ mode = 0;
+ }
+
if (dentry->d_name.len > NFS_SERVER(dir)->namelen)
return -ENAMETOOLONG;
diff --git a/fs/smb/client/dir.c b/fs/smb/client/dir.c
index 6fa6d48fdfd3..f75095a48fdb 100644
--- a/fs/smb/client/dir.c
+++ b/fs/smb/client/dir.c
@@ -534,6 +534,11 @@ int cifs_atomic_open(struct inode *dir, struct dentry *direntry,
if (unlikely(cifs_forced_shutdown(cifs_sb)))
return smb_EIO(smb_eio_trace_forced_shutdown);
+ if (O_IS_MKDIR(oflags)) {
+ oflags &= ~O_CREAT;
+ mode = 0;
+ }
+
/*
* Posix open is only called (at lookup time) for file create now. For
* opens (rather than creates), because we do not know if it is a file
diff --git a/fs/vboxsf/dir.c b/fs/vboxsf/dir.c
index 0b9eab157432..f20b61f6d8da 100644
--- a/fs/vboxsf/dir.c
+++ b/fs/vboxsf/dir.c
@@ -318,6 +318,11 @@ static int vboxsf_dir_atomic_open(struct inode *parent, struct dentry *dentry,
u64 handle;
int err;
+ if (O_IS_MKDIR(flags)) {
+ flags &= ~O_CREAT;
+ mode = 0;
+ }
+
if (d_in_lookup(dentry)) {
struct dentry *res = vboxsf_dir_lookup(parent, dentry, 0);
if (res || d_really_is_positive(dentry))
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 10/12] vfs: refuse O_CREAT for directories through a dangling symlink
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (8 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 09/12] vfs: move O_IS_MKDIR check from lookup_open() into individual filesystems Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 11/12] vfs: short-circuit MAY_WRITE access for O_DIRECTORY opens Jori Koolstra
` (2 subsequent siblings)
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
open(O_CREAT) without O_EXCL follows a trailing symlink and, when the
symlink target does not exist, creates it. Refuse to create through a
dangling symlink for directories.
In lookup_open() a negative target reached with nd->depth > 0 was
arrived at by following a trailing symlink; since the dentry is negative
the symlink is dangling. Set create_error to -EEXIST in that case
(matching the errno returned by mkdir(2).) Reusing the existing
create_error path strips O_CREAT for both the generic and
->atomic_open create paths and only reports the error when the target is
actually negative. Thus opening an existing target through a symlink,
interior symlinks, and O_EXCL (which never follows the trailing link)
are all unaffected.
Suggested-by: Christian Brauner (Amutable) <brauner@kernel.org>
Reviewed-by: NeilBrown <neil@brown.name>
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/namei.c | 5 +++++
1 file changed, 5 insertions(+)
diff --git a/fs/namei.c b/fs/namei.c
index a990c9c8bddf..fce3aaa37365 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -4601,6 +4601,11 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
dentry, open_flag, mode);
else
create_error = -EROFS;
+ /* Refuse to create a directory through a dangling (trailing)
+ * symlink. For regular files this has been allowed historically
+ * on O_CREAT without O_EXCL. */
+ if (unlikely(nd->depth) && create_dir && !create_error)
+ create_error = -EEXIST;
}
if (create_error)
open_flag &= ~O_CREAT;
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 11/12] vfs: short-circuit MAY_WRITE access for O_DIRECTORY opens
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (9 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 10/12] vfs: refuse O_CREAT for directories through a dangling symlink Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 12/12] selftest: add tests for open*(O_CREAT|O_DIRECTORY) Jori Koolstra
2026-09-17 10:38 ` [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Christian Brauner
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
Requesting write access on a directory can never succeed. Rather
than performing a path-walk to determine whether the target is
actually a directory (-EISDIR) or not (-ENOTDIR), or does not exist
(-ENOENT), etc., we short-circuit to -ENOTDIR.
Currently O_WRONLY for directories is only blocked in may_open(),
which happens after we have the inode for the target, so after any
create via O_CREAT|O_DIRECTORY.
The advantage of short-circuiting is that we don't have to add even more
logic to lookup_open() to differentiate -EISDIR/-ENOTDIR. Also, for
filesystems that define ->atomic_open, handling this cannot even be done
at the VFS level, as we can't know ahead what the result of the lookup
will be.
Suggested-by: Christian Brauner (Amutable) <brauner@kernel.org>
Reviewed-by: NeilBrown <neil@brown.name>
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
fs/open.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
diff --git a/fs/open.c b/fs/open.c
index 189af02a2425..6cb5e2ad781f 100644
--- a/fs/open.c
+++ b/fs/open.c
@@ -1319,9 +1319,16 @@ inline int build_open_flags(const struct open_how *how, struct open_flags *op)
op->intent = flags & O_PATH ? 0 : LOOKUP_OPEN;
+ /*
+ * Requesting write access on a directory can never succeed. Rather
+ * than performing a path-walk to determine whether the target is
+ * actually a directory (-EISDIR) or not (-ENOTDIR), we short-circuit
+ * to -ENOTDIR.
+ */
+ if ((flags & O_DIRECTORY) && !(flags & __O_TMPFILE) && (acc_mode & MAY_WRITE))
+ return -ENOTDIR;
+
if (flags & O_CREAT) {
- if ((flags & O_DIRECTORY) && (acc_mode & MAY_WRITE))
- return -EISDIR;
op->intent |= LOOKUP_CREATE;
if (flags & O_EXCL) {
op->intent |= LOOKUP_EXCL;
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* [PATCH v6 12/12] selftest: add tests for open*(O_CREAT|O_DIRECTORY)
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (10 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 11/12] vfs: short-circuit MAY_WRITE access for O_DIRECTORY opens Jori Koolstra
@ 2026-09-13 18:50 ` Jori Koolstra
2026-09-17 10:38 ` [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Christian Brauner
12 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-13 18:50 UTC (permalink / raw)
To: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
Cc: Jori Koolstra
Add some tests for the new valid O_CREAT|O_DIRECTORY flag combination for
open*(2) to test compliance and to showcase its behaviour.
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
---
.../testing/selftests/filesystems/.gitignore | 1 +
tools/testing/selftests/filesystems/Makefile | 2 +-
.../filesystems/open_o_creat_o_dir.c | 296 ++++++++++++++++++
.../testing/selftests/filesystems/wrappers.h | 11 +
4 files changed, 309 insertions(+), 1 deletion(-)
create mode 100644 tools/testing/selftests/filesystems/open_o_creat_o_dir.c
diff --git a/tools/testing/selftests/filesystems/.gitignore b/tools/testing/selftests/filesystems/.gitignore
index 9eb185fb2f9d..01c588d4c84f 100644
--- a/tools/testing/selftests/filesystems/.gitignore
+++ b/tools/testing/selftests/filesystems/.gitignore
@@ -1,4 +1,5 @@
# SPDX-License-Identifier: GPL-2.0-only
+open_o_creat_o_dir
dnotify_test
devpts_pts
fclog
diff --git a/tools/testing/selftests/filesystems/Makefile b/tools/testing/selftests/filesystems/Makefile
index 03be337c1f35..0959bd26875a 100644
--- a/tools/testing/selftests/filesystems/Makefile
+++ b/tools/testing/selftests/filesystems/Makefile
@@ -1,7 +1,7 @@
# SPDX-License-Identifier: GPL-2.0
CFLAGS += $(KHDR_INCLUDES)
-TEST_GEN_PROGS := devpts_pts file_stressor anon_inode_test kernfs_test fclog ustat_test
+TEST_GEN_PROGS := open_o_creat_o_dir devpts_pts file_stressor anon_inode_test kernfs_test fclog ustat_test
TEST_GEN_PROGS += idmapped_tmpfile
TEST_GEN_PROGS_EXTENDED := dnotify_test
diff --git a/tools/testing/selftests/filesystems/open_o_creat_o_dir.c b/tools/testing/selftests/filesystems/open_o_creat_o_dir.c
new file mode 100644
index 000000000000..538a7f7c5803
--- /dev/null
+++ b/tools/testing/selftests/filesystems/open_o_creat_o_dir.c
@@ -0,0 +1,296 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <sys/stat.h>
+#include <errno.h>
+#include <limits.h>
+#include <fcntl.h>
+
+#include "kselftest_harness.h"
+#include "wrappers.h"
+
+#define openat_o_mkdir_checked_flags(dfd, pathname, flags) ({ \
+ struct stat __st; \
+ int __fd = openat_o_mkdir(dfd, pathname, flags, S_IRWXU); \
+ ASSERT_GE(__fd, 0); \
+ ASSERT_EQ(fstat(__fd, &__st), 0); \
+ EXPECT_TRUE(S_ISDIR(__st.st_mode)); \
+ __fd; \
+})
+
+#define openat_o_mkdir_checked(dfd, pathname) \
+ openat_o_mkdir_checked_flags(dfd, pathname, O_RDONLY)
+
+FIXTURE(open_o_creat_o_dir) {
+ char dirpath[PATH_MAX];
+ int dfd;
+};
+
+FIXTURE_SETUP(open_o_creat_o_dir)
+{
+ strcpy(self->dirpath, "/tmp/open_o_creat_o_dir_test.XXXXXX");
+ ASSERT_NE(mkdtemp(self->dirpath), NULL);
+ self->dfd = open(self->dirpath, O_DIRECTORY);
+ ASSERT_GE(self->dfd, 0);
+}
+
+FIXTURE_TEARDOWN(open_o_creat_o_dir)
+{
+ close(self->dfd);
+ rmdir(self->dirpath);
+}
+
+/* Does open_o_creat_o_dir return a fd at all? */
+TEST_F(open_o_creat_o_dir, returns_fd)
+{
+ int fd = openat_o_mkdir_checked(self->dfd, "newdir");
+ EXPECT_EQ(close(fd), 0);
+ EXPECT_EQ(unlinkat(self->dfd, "newdir", AT_REMOVEDIR), 0);
+}
+
+/* The fd must refer to the directory that was just created. */
+TEST_F(open_o_creat_o_dir, fd_is_created_dir)
+{
+ int fd;
+ struct stat st_via_fd, st_via_path;
+ char path[PATH_MAX];
+
+ fd = openat_o_mkdir_checked(self->dfd, "checkdir");
+
+ ASSERT_EQ(fstat(fd, &st_via_fd), 0);
+
+ snprintf(path, sizeof(path), "%s/checkdir", self->dirpath);
+ ASSERT_EQ(stat(path, &st_via_path), 0);
+
+ EXPECT_EQ(st_via_fd.st_ino, st_via_path.st_ino);
+ EXPECT_EQ(st_via_fd.st_dev, st_via_path.st_dev);
+
+ EXPECT_EQ(close(fd), 0);
+ EXPECT_EQ(rmdir(path), 0);
+}
+
+/* Missing parent component must fail with ENOENT. */
+TEST_F(open_o_creat_o_dir, enoent_missing_parent)
+{
+ EXPECT_EQ(openat_o_mkdir(self->dfd, "nonexistent/child", O_RDONLY, S_IRWXU), -1);
+ EXPECT_EQ(errno, ENOENT);
+}
+
+/* An invalid dfd must fail with EBADF. */
+TEST_F(open_o_creat_o_dir, ebadf)
+{
+ EXPECT_EQ(openat_o_mkdir(FD_INVALID, "badfdir", O_RDONLY, S_IRWXU), -1);
+ EXPECT_EQ(errno, EBADF);
+}
+
+/* A dfd that points to a file (not a directory) must fail with ENOTDIR. */
+TEST_F(open_o_creat_o_dir, enotdir_dfd)
+{
+ int file_fd;
+
+ file_fd = openat(self->dfd, "file",
+ O_CREAT | O_RDONLY, S_IRWXU);
+ ASSERT_GE(file_fd, 0);
+
+ EXPECT_EQ(openat_o_mkdir(file_fd, "subdir", O_RDONLY, S_IRWXU), -1);
+ EXPECT_EQ(errno, ENOTDIR);
+
+ EXPECT_EQ(close(file_fd), 0);
+ EXPECT_EQ(unlinkat(self->dfd, "file", 0), 0);
+}
+
+/*
+ * O_EXCL together with O_CREAT|O_DIRECTORY should succeed if the target
+ * directory does not yet exist. After directory creation, repeating this
+ * call must fail with EEXIST, but should succeed if the O_EXCL is dropped.
+ */
+TEST_F(open_o_creat_o_dir, o_excl_eexist)
+{
+ int excldir_fd;
+
+ excldir_fd = openat_o_mkdir_checked_flags(self->dfd, "excldir", O_EXCL);
+
+ EXPECT_EQ(openat_o_mkdir(self->dfd, "excldir", O_EXCL, S_IRWXU), -1);
+ EXPECT_EQ(errno, EEXIST);
+
+ int excldir_reopen_fd = openat_o_mkdir_checked(self->dfd, "excldir");
+
+ EXPECT_EQ(close(excldir_reopen_fd), 0);
+ EXPECT_EQ(close(excldir_fd), 0);
+ EXPECT_EQ(unlinkat(self->dfd, "excldir", AT_REMOVEDIR), 0);
+}
+
+/*
+ * O_CREAT|O_DIRECTORY on a path that already exists as a regular file
+ * must fail with ENOTDIR.
+ */
+TEST_F(open_o_creat_o_dir, existing_file_enotdir)
+{
+ int file_fd;
+
+ file_fd = openat(self->dfd, "regfile",
+ O_CREAT | O_RDONLY, S_IRWXU);
+ ASSERT_GE(file_fd, 0);
+ EXPECT_EQ(close(file_fd), 0);
+
+ EXPECT_EQ(openat_o_mkdir(self->dfd, "regfile", O_RDONLY, S_IRWXU), -1);
+ EXPECT_EQ(errno, ENOTDIR);
+
+ EXPECT_EQ(unlinkat(self->dfd, "regfile", 0), 0);
+}
+
+/*
+ * O_CREAT|O_DIRECTORY combined with a writable access mode must be
+ * rejected: a directory cannot be opened for writing.
+ */
+TEST_F(open_o_creat_o_dir, rejects_writable_acc_mode)
+{
+ EXPECT_EQ(openat_o_mkdir(self->dfd, "rdwrdir", O_RDWR, S_IRWXU), -1);
+ EXPECT_EQ(errno, ENOTDIR);
+ /* Clean up if the kernel created the directory anyway. */
+ unlinkat(self->dfd, "rdwrdir", AT_REMOVEDIR);
+}
+
+/*
+ * openat(O_CREAT|O_DIRECTORY) with a trailing slash should work.
+ */
+TEST_F(open_o_creat_o_dir, trailing_slash)
+{
+ int fd = openat_o_mkdir_checked(self->dfd, "newdir/");
+ EXPECT_EQ(close(fd), 0);
+ EXPECT_EQ(unlinkat(self->dfd, "newdir", AT_REMOVEDIR), 0);
+}
+
+/*
+ * openat(O_CREAT) with a trailing slash but without O_DIRECTORY
+ * must fail with EISDIR and must not create anything at the path.
+ */
+TEST_F(open_o_creat_o_dir, trailing_slash_no_o_dir)
+{
+ int fd;
+ struct stat st;
+
+ fd = openat(self->dfd, "trailing/", O_CREAT | O_RDONLY, S_IRWXU);
+ EXPECT_EQ(fd, -1);
+ EXPECT_EQ(errno, EISDIR);
+
+ EXPECT_EQ(fstatat(self->dfd, "trailing", &st, 0), -1);
+ EXPECT_EQ(errno, ENOENT);
+
+ /* Best-effort cleanup in case the kernel left a file behind. */
+ if (fd >= 0)
+ close(fd);
+ unlinkat(self->dfd, "trailing", 0);
+}
+
+/*
+ * The returned fd must be usable as a dfd for further *at() calls.
+ */
+TEST_F(open_o_creat_o_dir, fd_usable_as_dfd)
+{
+ int parent_fd, child_fd;
+ char path[PATH_MAX];
+
+ parent_fd = openat_o_mkdir_checked(self->dfd, "parent");
+ child_fd = openat_o_mkdir_checked(parent_fd, "child");
+
+ EXPECT_EQ(close(child_fd), 0);
+ EXPECT_EQ(close(parent_fd), 0);
+
+ snprintf(path, sizeof(path), "%s/parent/child", self->dirpath);
+ EXPECT_EQ(rmdir(path), 0);
+ snprintf(path, sizeof(path), "%s/parent", self->dirpath);
+ EXPECT_EQ(rmdir(path), 0);
+}
+
+/*
+ * O_CREAT|O_DIRECTORY must refuse to create through a dangling trailing
+ * symlink, and must not create anything at the symlink target.
+ */
+TEST_F(open_o_creat_o_dir, dangling_symlink_eexist)
+{
+ struct stat st;
+
+ ASSERT_EQ(symlinkat("danglink_target", self->dfd, "danglink"), 0);
+
+ EXPECT_EQ(openat_o_mkdir(self->dfd, "danglink", O_RDONLY, S_IRWXU), -1);
+ EXPECT_EQ(errno, EEXIST);
+
+ /* Nothing must have been created at the target. */
+ EXPECT_EQ(fstatat(self->dfd, "danglink_target", &st, 0), -1);
+ EXPECT_EQ(errno, ENOENT);
+
+ EXPECT_EQ(unlinkat(self->dfd, "danglink", 0), 0);
+}
+
+/*
+ * A trailing symlink that resolves to an existing directory must still open.
+ */
+TEST_F(open_o_creat_o_dir, symlink_not_dangling_ok)
+{
+ int fd;
+
+ ASSERT_EQ(mkdirat(self->dfd, "realdir", 0700), 0);
+ ASSERT_EQ(symlinkat("realdir", self->dfd, "dirlink"), 0);
+
+ /* Trailing symlink resolving to an existing directory. */
+ fd = openat_o_mkdir_checked(self->dfd, "dirlink");
+ EXPECT_EQ(close(fd), 0);
+
+ EXPECT_EQ(unlinkat(self->dfd, "dirlink", 0), 0);
+ EXPECT_EQ(unlinkat(self->dfd, "realdir", AT_REMOVEDIR), 0);
+}
+
+/*
+ * An O_CREAT|O_DIRECTORY open of an existing directory owned by someone else,
+ * inside a sticky world-writable directory, must be refused.
+ */
+TEST_F(open_o_creat_o_dir, sticky_dir_eacces)
+{
+ int sticky_fd, fd;
+
+ if (geteuid() != 0)
+ SKIP(return, "needs root for fchownat");
+
+ ASSERT_EQ(mkdirat(self->dfd, "sticky", 01777), 0);
+ ASSERT_EQ(fchmodat(self->dfd, "sticky", 01777, 0), 0);
+ sticky_fd = openat(self->dfd, "sticky", O_DIRECTORY | O_RDONLY);
+ ASSERT_GE(sticky_fd, 0);
+
+ ASSERT_EQ(mkdirat(sticky_fd, "otherdir", 0700), 0);
+ if (fchownat(sticky_fd, "otherdir", 1, 1, 0)) {
+ int err = errno;
+
+ unlinkat(sticky_fd, "otherdir", AT_REMOVEDIR);
+ close(sticky_fd);
+ unlinkat(self->dfd, "sticky", AT_REMOVEDIR);
+ SKIP(return, "cannot chown to uid 1: %s", strerror(err));
+ }
+
+ EXPECT_EQ(openat_o_mkdir(sticky_fd, "otherdir", O_RDONLY, S_IRWXU), -1);
+ EXPECT_EQ(errno, EACCES);
+
+ /* Without O_CREAT the very same open must still succeed. */
+ fd = openat(sticky_fd, "otherdir", O_DIRECTORY | O_RDONLY);
+ EXPECT_GE(fd, 0);
+ if (fd >= 0) {
+ EXPECT_EQ(close(fd), 0);
+ }
+
+ EXPECT_EQ(unlinkat(sticky_fd, "otherdir", AT_REMOVEDIR), 0);
+ EXPECT_EQ(close(sticky_fd), 0);
+ EXPECT_EQ(unlinkat(self->dfd, "sticky", AT_REMOVEDIR), 0);
+}
+
+/*
+ * O_TMPFILE is encoded as __O_TMPFILE|O_DIRECTORY. Now that O_CREAT is no
+ * longer rejected alongside O_DIRECTORY, O_TMPFILE|O_CREAT must still be
+ * rejected explicitly so that it cannot create a persistent file on kernels
+ * that predate O_TMPFILE.
+ */
+TEST_F(open_o_creat_o_dir, tmpfile_with_o_creat_einval)
+{
+ EXPECT_EQ(openat(self->dfd, ".", O_TMPFILE | O_CREAT | O_RDWR, S_IRWXU),
+ -1);
+ EXPECT_EQ(errno, EINVAL);
+}
+
+TEST_HARNESS_MAIN
diff --git a/tools/testing/selftests/filesystems/wrappers.h b/tools/testing/selftests/filesystems/wrappers.h
index 420ae4f908cf..abe5b85cebdc 100644
--- a/tools/testing/selftests/filesystems/wrappers.h
+++ b/tools/testing/selftests/filesystems/wrappers.h
@@ -13,6 +13,10 @@
#define STATX_MNT_ID_UNIQUE 0x00004000U /* Want/got extended stx_mount_id */
#endif
+#ifndef FD_INVALID
+#define FD_INVALID -10009
+#endif
+
static inline int sys_fsopen(const char *fsname, unsigned int flags)
{
return syscall(__NR_fsopen, fsname, flags);
@@ -105,4 +109,11 @@ static inline int sys_open_tree(int dfd, const char *filename, unsigned int flag
return syscall(__NR_open_tree, dfd, filename, flags);
}
+static inline int openat_o_mkdir(int dfd, const char *pathname,
+ unsigned int flags, mode_t mode)
+{
+ return syscall(__NR_openat, dfd, pathname,
+ flags | O_DIRECTORY | O_CREAT, mode);
+}
+
#endif
--
2.55.0
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
` (11 preceding siblings ...)
2026-09-13 18:50 ` [PATCH v6 12/12] selftest: add tests for open*(O_CREAT|O_DIRECTORY) Jori Koolstra
@ 2026-09-17 10:38 ` Christian Brauner
2026-09-18 8:18 ` Christian Brauner
12 siblings, 1 reply; 40+ messages in thread
From: Christian Brauner @ 2026-09-17 10:38 UTC (permalink / raw)
To: Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown, Amir Goldstein,
Jan Kara, linux-fsdevel, linux-kernel, Jori Koolstra
Cc: Christian Brauner
On Sun, 13 Sep 2026 20:50:04 +0200, Jori Koolstra wrote:
> Finally got back from holiday. I know Neil is also attempting to make
> changes to the same code paths as the O_CREAT|O_DIRECTORY series touch,
> so I hope, now that I have a bit more time, that I can push this
> forwards before I have to do more rebasing.
>
> This series implements new semantics for the O_CREAT|O_DIRECTORY flag
> combination for open*(2): perform a mkdir and open the resulting
> directory; return a pinning fd (which mkdir does not).
>
> [...]
Applied to the vfs-7.4.lookup branch of the vfs/vfs.git tree.
Patches in the vfs-7.4.lookup branch should appear in linux-next soon.
Please report any outstanding bugs that were missed during review in a
new review to the original patch series allowing us to drop it.
It's encouraged to provide Acked-bys and Reviewed-bys even though the
patch has now been applied. If possible patch trailers will be updated.
Note that commit hashes shown below are subject to change due to rebase,
trailer updates or similar. If in doubt, please check the listed branch.
tree: https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
branch: vfs-7.4.lookup
[01/12] fs/namei.c: use trailing_slashes()
https://git.kernel.org/vfs/vfs/c/614d27b66efd
[02/12] vfs: prepare vfs_creat|mkdir_no_perm for reuse in lookup_open()
https://git.kernel.org/vfs/vfs/c/2eacbb993f65
[03/12] vfs: lookup_open(): move setting FMODE_CREATED down
https://git.kernel.org/vfs/vfs/c/db4a1ea6a55e
[04/12] vfs: move ->create check in lookup_open() to before try_break_deleg()
https://git.kernel.org/vfs/vfs/c/27ab83a2c268
[05/12] vfs: lookup_open(): use vfs_create_no_perm()
https://git.kernel.org/vfs/vfs/c/900d6abddd54
[06/12] vfs: lookup_open(): lock the parent as I_MUTEX_PARENT
https://git.kernel.org/vfs/vfs/c/0567a62f5ea1
[07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
https://git.kernel.org/vfs/vfs/c/596e3a392d61
[08/12] vfs: change ->create/->mkdir operations unavailable errno
https://git.kernel.org/vfs/vfs/c/dc6315c50724
[09/12] vfs: move O_IS_MKDIR check from lookup_open() into individual filesystems
https://git.kernel.org/vfs/vfs/c/885bf5afa273
[10/12] vfs: refuse O_CREAT for directories through a dangling symlink
https://git.kernel.org/vfs/vfs/c/c7a93442a9b2
[11/12] vfs: short-circuit MAY_WRITE access for O_DIRECTORY opens
https://git.kernel.org/vfs/vfs/c/093f13c7a954
[12/12] selftest: add tests for open*(O_CREAT|O_DIRECTORY)
https://git.kernel.org/vfs/vfs/c/76208ee704c2
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno
2026-09-13 18:50 ` [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno Jori Koolstra
@ 2026-09-18 8:04 ` Christian Brauner
2026-09-25 22:55 ` Jori Koolstra
0 siblings, 1 reply; 40+ messages in thread
From: Christian Brauner @ 2026-09-18 8:04 UTC (permalink / raw)
To: Jori Koolstra
Cc: Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown, Amir Goldstein,
Jan Kara, linux-fsdevel, linux-kernel
On Sun, Sep 13, 2026 at 08:50:12PM +0200, Jori Koolstra wrote:
> Currently filesystems return -EACCES for missing ->create and -EPERM for
> missing ->mkdir in vfs_create/vfs_mkdir. Instead of duplicating these
> dubious and inconsistent error codes to lookup_open(), change all to
> -EOPNOTSUPP.
>
> Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
> ---
I suspect that still has regression potential but we can try.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-13 18:50 ` [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
@ 2026-09-18 8:17 ` Christian Brauner
2026-09-18 10:07 ` NeilBrown
2026-09-25 23:13 ` Jori Koolstra
0 siblings, 2 replies; 40+ messages in thread
From: Christian Brauner @ 2026-09-18 8:17 UTC (permalink / raw)
To: Jori Koolstra, NeilBrown
Cc: Jeff Layton, Al Viro, Aleksa Sarai, Amir Goldstein, Jan Kara,
linux-fsdevel, linux-kernel
On Sun, Sep 13, 2026 at 08:50:11PM +0200, Jori Koolstra wrote:
> Currently there is no way to race-freely create and open a directory.
> For regular files we have open(O_CREAT) for creating a new file inode,
> and returning a pinning fd to it. The lack of such functionality for
> directories means that when populating a directory tree there's always
> a race involved: the inodes first need to be created, and then opened
> to adjust their permissions/ownership/labels/timestamps/acls/xattrs/...,
> but in the time window between the creation and the opening they might
> be replaced by something else.
>
> Addressing this race without a proper API is only partially possible:
> the caller can immediately fstat() what was opened to verify that it
> has the expected inode type, owner and mode. But besides being easy to
> get wrong, this cannot establish who created the directory: a directory
> created by another process with identical credentials is
> indistinguishable from one the caller created itself, so the caller
> cannot tell whether the directory is its own to manage.
>
> Historically, the O_CREAT|O_DIRECTORY behaviour was to return ENOTDIR if
> a regular file exists at the open path; EISDIR if a directory exists at
> the path; and to create a regular file if no file exists at the path.
> This behaviour changed accidentally with
> commit 973d4b73fbaf ("do_last(): rejoin the common path even earlier in
> FMODE_{OPENED,CREATED} case") causing ENOTDIR to return in the last case
> while still creating the file. As this change was not detected for a
> long time, Brauner proposed to adopt the more consistent NetBSD
> behaviour, i.e. to return EINVAL on the O_CREAT|O_DIRECTORY combination.
> This change was applied in commit 43b450632676 ("open: return EINVAL for
> O_DIRECTORY | O_CREAT") in March, 2023. As the EINVAL behaviour has been
> in the kernel for about 3 years now, no rollback is expected as a result
> of userspace reliance on old behaviour, leaving us free to reassign the
> O_CREAT|O_DIRECTORY semantics.
>
> O_CREAT|O_DIRECTORY is made to reduce to a lookup on ->atomic_open()
> filesystems. These filesystems currently cannot handle
> O_CREAT|O_DIRECTORY without protocol extensions and therefore are forced
> into a fallback mode by stripping the O_CREAT bit. This causes existing
> directories to be successfully opened, while for targets that should
> have been created, -ENOENT is returned. This -ENOENT is then converted
> to -EOPNOTSUPP in later atomic_open(). The simple option of just
> returning -EOPNOTSUPP directly leads to inconsistent behaviour: before
> ->atomic_open() is called in lookup_open(), the dcache is queried. So
> returning -EOPNOTSUPP immediately would make O_CREAT|O_DIRECTORY
> dependent on the cache state of the dentry.
>
> There is no separate sysctl for directory creation implemented currently.
> Therefore, for the S_ISDIR case, disabling sysctl_protected_regular is
> not enough to allow creating a directory in a sticky folder, because that
> may surprise users not expecting that O_CREAT|O_DIRECTORY is possible on
> newer kernels.
>
> This feature idea (and some of its description) is taken from the
> UAPI group:
> https://github.com/uapi-group/kernel-features?tab=readme-ov-file#race-free-creation-and-opening-of-non-file-inodes
>
> Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
> ---
> fs/namei.c | 116 +++++++++++++++++++++++++++++++++++-------
> fs/open.c | 25 +++++----
> include/linux/fcntl.h | 6 +++
> 3 files changed, 117 insertions(+), 30 deletions(-)
>
> diff --git a/fs/namei.c b/fs/namei.c
> index 0efd395a1a65..6ff0a3c04f02 100644
> --- a/fs/namei.c
> +++ b/fs/namei.c
> @@ -1382,13 +1382,13 @@ int may_linkat(struct mnt_idmap *idmap, const struct path *link)
>
> /**
> * may_create_in_sticky - Check whether an O_CREAT open in a sticky directory
> - * should be allowed, or not, on files that already
> - * exist.
> + * should be allowed, or not, on files/directories that
> + * already exist.
> * @idmap: idmap of the mount the inode was found from
> * @nd: nameidata pathwalk data
> * @inode: the inode of the file to open
> *
> - * Block an O_CREAT open of a FIFO (or a regular file) when:
> + * Block an O_CREAT open of a FIFO (or a regular file/directory) when:
> * - sysctl_protected_fifos (or sysctl_protected_regular) is enabled
> * - the file already exists
> * - we are in a sticky directory
> @@ -1416,6 +1416,14 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> if (likely(!(dir_mode & S_ISVTX)))
> return 0;
>
> + /*
> + * There is no separate sysctl for directory creation in sticky
> + * folders. Therefore, for the S_ISDIR case, disabling
> + * sysctl_protected_regular is not enough to allow creating a
> + * directory in a sticky folder, because that may surprise users
> + * not expecting that O_CREAT|O_DIRECTORY is possible on newer
> + * kernels.
> + */
> if (S_ISREG(inode->i_mode) && !sysctl_protected_regular)
> return 0;
>
> @@ -1447,6 +1455,12 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> "sticky_create_regular");
> return -EACCES;
> }
> +
> + if (S_ISDIR(inode->i_mode)) {
> + audit_log_path_denied(AUDIT_ANOM_CREAT,
> + "sticky_create_dir");
> + return -EACCES;
> + }
> }
>
> return 0;
> @@ -4334,21 +4348,43 @@ static inline int open_to_namei_flags(int flag)
>
> static int may_o_create(struct mnt_idmap *idmap,
> const struct path *dir, struct dentry *dentry,
> - umode_t mode)
> + int open_flag, umode_t mode)
> {
> - int error = security_path_mknod(dir, dentry, mode, 0);
> + struct inode *dir_inode = dir->dentry->d_inode;
> + bool create_dir = O_IS_MKDIR(open_flag);
> + int error;
> +
> + WARN_ON_ONCE(create_dir && !(mode & S_IFDIR));
> +
> + if (create_dir)
> + error = security_path_mkdir(dir, dentry, mode);
> + else
> + error = security_path_mknod(dir, dentry, mode, 0);
> if (error)
> return error;
>
> if (!fsuidgid_has_mapping(dir->dentry->d_sb, idmap))
> return -EOVERFLOW;
>
> - error = inode_permission(idmap, dir->dentry->d_inode,
> - MAY_WRITE | MAY_EXEC);
> + error = inode_permission(idmap, dir_inode, MAY_WRITE | MAY_EXEC);
> if (error)
> return error;
>
> - return security_inode_create(dir->dentry->d_inode, dentry, mode);
> + if (create_dir)
> + error = security_inode_mkdir(dir_inode, dentry, mode);
> + else
> + error = security_inode_create(dir_inode, dentry, mode);
> +
> + return error;
> +}
> +
> +static inline umode_t o_create_mode(struct mnt_idmap *idmap,
> + const struct inode *dir, int open_flag, umode_t mode)
> +{
> + if (O_IS_MKDIR(open_flag))
> + return vfs_prepare_mode(idmap, dir, mode, S_IRWXUGO | S_ISVTX, S_IFDIR);
> + else
> + return vfs_prepare_mode(idmap, dir, mode, S_IALLUGO, S_IFREG);
> }
>
> /**
> @@ -4384,8 +4420,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
>
> file->__f_path.dentry = DENTRY_NOT_SET;
> file->__f_path.mnt = path->mnt;
> +
> error = dir_inode->i_op->atomic_open(dir_inode, dentry, file,
> - open_to_namei_flags(open_flag), mode);
> + open_to_namei_flags(open_flag), mode);
> d_lookup_done(dentry);
>
> if (!error) {
> @@ -4427,12 +4464,32 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> */
> audit_inode_child(dir_inode, dentry, AUDIT_TYPE_CHILD_CREATE);
> error = create_error;
> + } else if (O_IS_MKDIR(open_flag) && error == -ENOENT) {
> + /*
> + * If the underlying filesystem does not implement
> + * O_CREAT|O_DIRECTORY, it strips the O_CREAT bit and
> + * continues as a lookup. We can't simply return
> + * -EOPNOTSUPP from unsupported ->atomic_open()
> + * implementations because the dentry might be in the
> + * dcache. In that case, lookup_open() returns before
> + * reaching ->atomic_open(), and hence whether you get
> + * -EOPNOTSUPP on O_CREAT|O_DIRECTORY would not only
> + * depend on the underlying filesystem, but also on
> + * the state of the dcache. Still, we must make an
> + * effort to differentiate a regular -ENOENT from the
> + * unsupported O_CREAT|O_DIRECTORY case.
> + */
> + error = -EOPNOTSUPP;
> }
> dput(dentry);
> dentry = ERR_PTR(error);
> } else {
> - if (file->f_mode & FMODE_CREATED)
> - fsnotify_create(dir_inode, dentry);
> + if (file->f_mode & FMODE_CREATED) {
> + if (d_is_dir(dentry))
> + fsnotify_mkdir(dir_inode, dentry);
> + else
> + fsnotify_create(dir_inode, dentry);
> + }
> if (file->f_mode & FMODE_OPENED)
> fsnotify_open(file);
> }
> @@ -4441,6 +4498,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> return dentry;
> }
>
> +static inline
> +struct dentry *vfs_mkdir_no_perm(struct mnt_idmap *, struct inode *, struct dentry *,
> + umode_t, struct delegated_inode *);
> /*
> * Look up and maybe create and open the last component.
> *
> @@ -4462,6 +4522,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> struct mnt_idmap *idmap;
> struct dentry *dir = nd->path.dentry;
> struct inode *dir_inode = dir->d_inode;
> + bool create_dir = O_IS_MKDIR(op->open_flag);
> int open_flag;
> struct dentry *dentry;
> int error, create_error;
> @@ -4474,6 +4535,9 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> mode = op->mode;
> create_error = 0;
>
> + if (create_dir && dir_inode->i_op->atomic_open)
> + open_flag &= ~O_CREAT;
> +
> if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) {
> got_write = !mnt_want_write(nd->path.mnt);
> /*
> @@ -4534,10 +4598,10 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> if (open_flag & O_CREAT) {
> if (open_flag & O_EXCL)
> open_flag &= ~O_TRUNC;
> - mode = vfs_prepare_mode(idmap, dir_inode, mode, mode, mode);
> + mode = o_create_mode(idmap, dir_inode, open_flag, mode);
> if (likely(got_write))
> create_error = may_o_create(idmap, &nd->path,
> - dentry, mode);
> + dentry, open_flag, mode);
> else
> create_error = -EROFS;
> }
> @@ -4582,12 +4646,25 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> goto out_dput;
> }
>
> - if (!dir_inode->i_op->create) {
> + /* mimic operation missing errnos of vfs_mkdir/vfs_create */
> + if (create_dir && !dir_inode->i_op->mkdir) {
> + error = -EPERM;
> + goto out_dput;
> + }
> + if (!create_dir && !dir_inode->i_op->create) {
> error = -EACCES;
> goto out_dput;
> }
>
> - error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
> + if (create_dir) {
> + struct dentry *res = vfs_mkdir_no_perm(idmap, dir_inode, dentry, mode,
> + &delegated_inode);
So, I think this is broken. Whatever vfs_mkdir_no_perm() returns is
passed to do_open(). Kernfs makes that buggy.
cgroup, cgroup2, and resctrl are all implemented on top of kernfs. And
kernfs ->mkdir:: iop never instantiates the dentry.
So that means e.g.,
openat(cgroup_dir, "subdir", O_CREAT|O_DIRECTORY) creates a cgroup
and then fails with ENOTDIR.
So the negative dentry gets handed out and now userspace holds an fd
with that negative dentry. So say userspace does fchown() to 1000 and
then fchmod() with the sticky bit and then you get a NULL deref. I have
reproduced this.
Neil can correct me but the fix might be to check whether the dentry is
negative in lookup_open() and re-lookup nd->last with the parent still locked.
I think that's what nfsd_create_locked() and cachefiles_get_directory() do
after vfs_mkdir().
Here's the callchain:
Common path: create the cgroup, come back with a negative dentry
__x64_sys_openat
do_sys_openat2 fs/open.c
build_open_flags O_IS_MKDIR(flags) -> op->mode = mode | S_IFDIR
O_DIRECTORY -> lookup_flags |= LOOKUP_DIRECTORY
do_file_open fs/namei.c
lookup_fast_for_open dcache miss -> NULL
lookup_open
mnt_want_write
inode_lock_nested(dir, I_MUTEX_PARENT)
d_lookup -> NULL
d_alloc_parallel in-lookup dentry
o_create_mode S_IFDIR | 0755 & ~umask
may_o_create
security_path_mkdir
inode_permission(MAY_WRITE|MAY_EXEC)
kernfs_iop_permission
generic_permission owner of a 0755 dir: allowed
security_inode_mkdir
dir_inode->i_op->lookup
kernfs_iop_lookup
kernfs_find_ns -> NULL node does not exist yet
d_splice_alias(NULL, dentry) -> __d_add(): hashed, NEGATIVE
try_break_deleg
dir->i_op->mkdir
kernfs_iop_mkdir
scops->mkdir
cgroup_mkdir no capable() check
cgroup_create owner = current_fsuid()
css_populate_dir
kernfs_activate cgroup now exists
return ERR_PTR(0) == NULL, dentry untouched
de == NULL -> keep original dentry
fsnotify_mkdir(dir, dentry) fine with a negative dentry
dentry = res still negative
file->f_mode |= FMODE_CREATED
inode_unlock(dir); mnt_drop_write
FMODE_CREATED set ->
dput(nd->path.dentry)
nd->path.dentry = dentry negative dentry becomes the "opened" path
return NULL
do_open
open_flag & O_CREAT ->
may_create_in_sticky(idmap, nd, d_backing_inode(nd->path.dentry))
inode == NULL
Scenario 1, parent not sticky: cgroup created, ENOTDIR returned
may_create_in_sticky
if (!(dir_mode & S_ISVTX)) return 0; inode never touched
(nd->flags & LOOKUP_DIRECTORY) && !d_can_lookup(nd->path.dentry)
DCACHE_MISS_TYPE -> false
return -ENOTDIR
terminate_walk
fput_close(file)
return ERR_PTR(-ENOTDIR) "child" cgroup stays behind
A second openat of the same name works because kernfs_dop_revalidate() sees the parent's revision changed, d_invalidate()s the stale negative dentry, and the fresh kernfs_iop_lookup() now
finds the node.
Scenario 2, sticky parent: NULL dereference
Setup, one-time, as root (what systemd's Delegate=yes does for user.slice/user-1000.slice/user@1000.service):
fchown(pfd, 1000, 1000)
chown_common -> notify_change -> kernfs_iop_setattr -> __kernfs_setattr
Then as uid 1000:
fchmod(pfd, 01755)
chmod_common
newattrs.ia_mode = (mode & S_IALLUGO) | ... S_ISVTX is inside S_IALLUGO
notify_change
setattr_prepare
inode_owner_or_capable owner -> ok
kernfs_iop_setattr
__kernfs_setattr kn->mode = ia_mode, no masking
setattr_copy inode->i_mode = 01755
and the open, same chain as above until may_create_in_sticky(), now with nd->dir_mode = 01755:
may_create_in_sticky
if (!(dir_mode & S_ISVTX)) return 0; not taken
if (S_ISREG(inode->i_mode) && ...) inode == NULL
-> KASAN null-ptr-deref
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-17 10:38 ` [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Christian Brauner
@ 2026-09-18 8:18 ` Christian Brauner
0 siblings, 0 replies; 40+ messages in thread
From: Christian Brauner @ 2026-09-18 8:18 UTC (permalink / raw)
To: Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown, Amir Goldstein,
Jan Kara, linux-fsdevel, linux-kernel, Jori Koolstra
On Thu, Sep 17, 2026 at 12:38:33PM +0200, Christian Brauner wrote:
> On Sun, 13 Sep 2026 20:50:04 +0200, Jori Koolstra wrote:
> > Finally got back from holiday. I know Neil is also attempting to make
> > changes to the same code paths as the O_CREAT|O_DIRECTORY series touch,
> > so I hope, now that I have a bit more time, that I can push this
> > forwards before I have to do more rebasing.
> >
> > This series implements new semantics for the O_CREAT|O_DIRECTORY flag
> > combination for open*(2): perform a mkdir and open the resulting
> > directory; return a pinning fd (which mkdir does not).
> >
> > [...]
>
> Applied to the vfs-7.4.lookup branch of the vfs/vfs.git tree.
> Patches in the vfs-7.4.lookup branch should appear in linux-next soon.
Dropped after I reviewed it once more. But we're close, I think.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-18 8:17 ` Christian Brauner
@ 2026-09-18 10:07 ` NeilBrown
2026-09-19 0:01 ` NeilBrown
2026-09-25 23:13 ` Jori Koolstra
1 sibling, 1 reply; 40+ messages in thread
From: NeilBrown @ 2026-09-18 10:07 UTC (permalink / raw)
To: Christian Brauner
Cc: Jori Koolstra, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
On Fri, 18 Sep 2026, Christian Brauner wrote:
> On Sun, Sep 13, 2026 at 08:50:11PM +0200, Jori Koolstra wrote:
> > Currently there is no way to race-freely create and open a directory.
> > For regular files we have open(O_CREAT) for creating a new file inode,
> > and returning a pinning fd to it. The lack of such functionality for
> > directories means that when populating a directory tree there's always
> > a race involved: the inodes first need to be created, and then opened
> > to adjust their permissions/ownership/labels/timestamps/acls/xattrs/...,
> > but in the time window between the creation and the opening they might
> > be replaced by something else.
> >
> > Addressing this race without a proper API is only partially possible:
> > the caller can immediately fstat() what was opened to verify that it
> > has the expected inode type, owner and mode. But besides being easy to
> > get wrong, this cannot establish who created the directory: a directory
> > created by another process with identical credentials is
> > indistinguishable from one the caller created itself, so the caller
> > cannot tell whether the directory is its own to manage.
> >
> > Historically, the O_CREAT|O_DIRECTORY behaviour was to return ENOTDIR if
> > a regular file exists at the open path; EISDIR if a directory exists at
> > the path; and to create a regular file if no file exists at the path.
> > This behaviour changed accidentally with
> > commit 973d4b73fbaf ("do_last(): rejoin the common path even earlier in
> > FMODE_{OPENED,CREATED} case") causing ENOTDIR to return in the last case
> > while still creating the file. As this change was not detected for a
> > long time, Brauner proposed to adopt the more consistent NetBSD
> > behaviour, i.e. to return EINVAL on the O_CREAT|O_DIRECTORY combination.
> > This change was applied in commit 43b450632676 ("open: return EINVAL for
> > O_DIRECTORY | O_CREAT") in March, 2023. As the EINVAL behaviour has been
> > in the kernel for about 3 years now, no rollback is expected as a result
> > of userspace reliance on old behaviour, leaving us free to reassign the
> > O_CREAT|O_DIRECTORY semantics.
> >
> > O_CREAT|O_DIRECTORY is made to reduce to a lookup on ->atomic_open()
> > filesystems. These filesystems currently cannot handle
> > O_CREAT|O_DIRECTORY without protocol extensions and therefore are forced
> > into a fallback mode by stripping the O_CREAT bit. This causes existing
> > directories to be successfully opened, while for targets that should
> > have been created, -ENOENT is returned. This -ENOENT is then converted
> > to -EOPNOTSUPP in later atomic_open(). The simple option of just
> > returning -EOPNOTSUPP directly leads to inconsistent behaviour: before
> > ->atomic_open() is called in lookup_open(), the dcache is queried. So
> > returning -EOPNOTSUPP immediately would make O_CREAT|O_DIRECTORY
> > dependent on the cache state of the dentry.
> >
> > There is no separate sysctl for directory creation implemented currently.
> > Therefore, for the S_ISDIR case, disabling sysctl_protected_regular is
> > not enough to allow creating a directory in a sticky folder, because that
> > may surprise users not expecting that O_CREAT|O_DIRECTORY is possible on
> > newer kernels.
> >
> > This feature idea (and some of its description) is taken from the
> > UAPI group:
> > https://github.com/uapi-group/kernel-features?tab=readme-ov-file#race-free-creation-and-opening-of-non-file-inodes
> >
> > Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
> > ---
> > fs/namei.c | 116 +++++++++++++++++++++++++++++++++++-------
> > fs/open.c | 25 +++++----
> > include/linux/fcntl.h | 6 +++
> > 3 files changed, 117 insertions(+), 30 deletions(-)
> >
> > diff --git a/fs/namei.c b/fs/namei.c
> > index 0efd395a1a65..6ff0a3c04f02 100644
> > --- a/fs/namei.c
> > +++ b/fs/namei.c
> > @@ -1382,13 +1382,13 @@ int may_linkat(struct mnt_idmap *idmap, const struct path *link)
> >
> > /**
> > * may_create_in_sticky - Check whether an O_CREAT open in a sticky directory
> > - * should be allowed, or not, on files that already
> > - * exist.
> > + * should be allowed, or not, on files/directories that
> > + * already exist.
> > * @idmap: idmap of the mount the inode was found from
> > * @nd: nameidata pathwalk data
> > * @inode: the inode of the file to open
> > *
> > - * Block an O_CREAT open of a FIFO (or a regular file) when:
> > + * Block an O_CREAT open of a FIFO (or a regular file/directory) when:
> > * - sysctl_protected_fifos (or sysctl_protected_regular) is enabled
> > * - the file already exists
> > * - we are in a sticky directory
> > @@ -1416,6 +1416,14 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> > if (likely(!(dir_mode & S_ISVTX)))
> > return 0;
> >
> > + /*
> > + * There is no separate sysctl for directory creation in sticky
> > + * folders. Therefore, for the S_ISDIR case, disabling
> > + * sysctl_protected_regular is not enough to allow creating a
> > + * directory in a sticky folder, because that may surprise users
> > + * not expecting that O_CREAT|O_DIRECTORY is possible on newer
> > + * kernels.
> > + */
> > if (S_ISREG(inode->i_mode) && !sysctl_protected_regular)
> > return 0;
> >
> > @@ -1447,6 +1455,12 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> > "sticky_create_regular");
> > return -EACCES;
> > }
> > +
> > + if (S_ISDIR(inode->i_mode)) {
> > + audit_log_path_denied(AUDIT_ANOM_CREAT,
> > + "sticky_create_dir");
> > + return -EACCES;
> > + }
> > }
> >
> > return 0;
> > @@ -4334,21 +4348,43 @@ static inline int open_to_namei_flags(int flag)
> >
> > static int may_o_create(struct mnt_idmap *idmap,
> > const struct path *dir, struct dentry *dentry,
> > - umode_t mode)
> > + int open_flag, umode_t mode)
> > {
> > - int error = security_path_mknod(dir, dentry, mode, 0);
> > + struct inode *dir_inode = dir->dentry->d_inode;
> > + bool create_dir = O_IS_MKDIR(open_flag);
> > + int error;
> > +
> > + WARN_ON_ONCE(create_dir && !(mode & S_IFDIR));
> > +
> > + if (create_dir)
> > + error = security_path_mkdir(dir, dentry, mode);
> > + else
> > + error = security_path_mknod(dir, dentry, mode, 0);
> > if (error)
> > return error;
> >
> > if (!fsuidgid_has_mapping(dir->dentry->d_sb, idmap))
> > return -EOVERFLOW;
> >
> > - error = inode_permission(idmap, dir->dentry->d_inode,
> > - MAY_WRITE | MAY_EXEC);
> > + error = inode_permission(idmap, dir_inode, MAY_WRITE | MAY_EXEC);
> > if (error)
> > return error;
> >
> > - return security_inode_create(dir->dentry->d_inode, dentry, mode);
> > + if (create_dir)
> > + error = security_inode_mkdir(dir_inode, dentry, mode);
> > + else
> > + error = security_inode_create(dir_inode, dentry, mode);
> > +
> > + return error;
> > +}
> > +
> > +static inline umode_t o_create_mode(struct mnt_idmap *idmap,
> > + const struct inode *dir, int open_flag, umode_t mode)
> > +{
> > + if (O_IS_MKDIR(open_flag))
> > + return vfs_prepare_mode(idmap, dir, mode, S_IRWXUGO | S_ISVTX, S_IFDIR);
> > + else
> > + return vfs_prepare_mode(idmap, dir, mode, S_IALLUGO, S_IFREG);
> > }
> >
> > /**
> > @@ -4384,8 +4420,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> >
> > file->__f_path.dentry = DENTRY_NOT_SET;
> > file->__f_path.mnt = path->mnt;
> > +
> > error = dir_inode->i_op->atomic_open(dir_inode, dentry, file,
> > - open_to_namei_flags(open_flag), mode);
> > + open_to_namei_flags(open_flag), mode);
> > d_lookup_done(dentry);
> >
> > if (!error) {
> > @@ -4427,12 +4464,32 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > */
> > audit_inode_child(dir_inode, dentry, AUDIT_TYPE_CHILD_CREATE);
> > error = create_error;
> > + } else if (O_IS_MKDIR(open_flag) && error == -ENOENT) {
> > + /*
> > + * If the underlying filesystem does not implement
> > + * O_CREAT|O_DIRECTORY, it strips the O_CREAT bit and
> > + * continues as a lookup. We can't simply return
> > + * -EOPNOTSUPP from unsupported ->atomic_open()
> > + * implementations because the dentry might be in the
> > + * dcache. In that case, lookup_open() returns before
> > + * reaching ->atomic_open(), and hence whether you get
> > + * -EOPNOTSUPP on O_CREAT|O_DIRECTORY would not only
> > + * depend on the underlying filesystem, but also on
> > + * the state of the dcache. Still, we must make an
> > + * effort to differentiate a regular -ENOENT from the
> > + * unsupported O_CREAT|O_DIRECTORY case.
> > + */
> > + error = -EOPNOTSUPP;
> > }
> > dput(dentry);
> > dentry = ERR_PTR(error);
> > } else {
> > - if (file->f_mode & FMODE_CREATED)
> > - fsnotify_create(dir_inode, dentry);
> > + if (file->f_mode & FMODE_CREATED) {
> > + if (d_is_dir(dentry))
> > + fsnotify_mkdir(dir_inode, dentry);
> > + else
> > + fsnotify_create(dir_inode, dentry);
> > + }
> > if (file->f_mode & FMODE_OPENED)
> > fsnotify_open(file);
> > }
> > @@ -4441,6 +4498,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > return dentry;
> > }
> >
> > +static inline
> > +struct dentry *vfs_mkdir_no_perm(struct mnt_idmap *, struct inode *, struct dentry *,
> > + umode_t, struct delegated_inode *);
> > /*
> > * Look up and maybe create and open the last component.
> > *
> > @@ -4462,6 +4522,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > struct mnt_idmap *idmap;
> > struct dentry *dir = nd->path.dentry;
> > struct inode *dir_inode = dir->d_inode;
> > + bool create_dir = O_IS_MKDIR(op->open_flag);
> > int open_flag;
> > struct dentry *dentry;
> > int error, create_error;
> > @@ -4474,6 +4535,9 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > mode = op->mode;
> > create_error = 0;
> >
> > + if (create_dir && dir_inode->i_op->atomic_open)
> > + open_flag &= ~O_CREAT;
> > +
> > if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) {
> > got_write = !mnt_want_write(nd->path.mnt);
> > /*
> > @@ -4534,10 +4598,10 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > if (open_flag & O_CREAT) {
> > if (open_flag & O_EXCL)
> > open_flag &= ~O_TRUNC;
> > - mode = vfs_prepare_mode(idmap, dir_inode, mode, mode, mode);
> > + mode = o_create_mode(idmap, dir_inode, open_flag, mode);
> > if (likely(got_write))
> > create_error = may_o_create(idmap, &nd->path,
> > - dentry, mode);
> > + dentry, open_flag, mode);
> > else
> > create_error = -EROFS;
> > }
> > @@ -4582,12 +4646,25 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > goto out_dput;
> > }
> >
> > - if (!dir_inode->i_op->create) {
> > + /* mimic operation missing errnos of vfs_mkdir/vfs_create */
> > + if (create_dir && !dir_inode->i_op->mkdir) {
> > + error = -EPERM;
> > + goto out_dput;
> > + }
> > + if (!create_dir && !dir_inode->i_op->create) {
> > error = -EACCES;
> > goto out_dput;
> > }
> >
> > - error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
> > + if (create_dir) {
> > + struct dentry *res = vfs_mkdir_no_perm(idmap, dir_inode, dentry, mode,
> > + &delegated_inode);
>
> So, I think this is broken. Whatever vfs_mkdir_no_perm() returns is
> passed to do_open(). Kernfs makes that buggy.
>
> cgroup, cgroup2, and resctrl are all implemented on top of kernfs. And
> kernfs ->mkdir:: iop never instantiates the dentry.
>
> So that means e.g.,
>
> openat(cgroup_dir, "subdir", O_CREAT|O_DIRECTORY) creates a cgroup
> and then fails with ENOTDIR.
>
> So the negative dentry gets handed out and now userspace holds an fd
> with that negative dentry. So say userspace does fchown() to 1000 and
> then fchmod() with the sticky bit and then you get a NULL deref. I have
> reproduced this.
>
> Neil can correct me but the fix might be to check whether the dentry is
> negative in lookup_open() and re-lookup nd->last with the parent still locked.
> I think that's what nfsd_create_locked() and cachefiles_get_directory() do
> after vfs_mkdir().
nfsd_create_locked() used to do that before vfs_mkdir() could return a
dentry, but it doesn't any more. The reason was because
d_splice_alias() on might return a different dentry.
In this case we want the same dentry, but we need to do a lookup on it.
I'd rather fix this in kernfs, but maybe that is a longer-term goal.
The comment in kernfs_dop_revalidate() suggests the we should d_drop()
the negative dentry and d_alloc_parallel() a new one and ->lookup that.
I'm not certain that is needed if we keep the parent locked, but we
would need to be certain.
We at least need to d_drop() the dentry before ->lookup as ->lookup
cannot handle hashed dentries and a hashed-negative dentry is passed
to ->mkdir.
I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
probably not.
Summary: I think that if vfs_mkdir() returns NULL (success) but the
dentry is negative, we need to d_drop() and call ->lookup with a big
comment about kernfs. But we need to double-check that this will do the
right thing with ->d_time (I think it will).
We also need to think carefully about races with
kernfs_dop_revalidate(), which could happen concurrently with the
->lookup.
NeilBrown
>
> Here's the callchain:
>
> Common path: create the cgroup, come back with a negative dentry
>
> __x64_sys_openat
> do_sys_openat2 fs/open.c
> build_open_flags O_IS_MKDIR(flags) -> op->mode = mode | S_IFDIR
> O_DIRECTORY -> lookup_flags |= LOOKUP_DIRECTORY
> do_file_open fs/namei.c
> lookup_fast_for_open dcache miss -> NULL
> lookup_open
> mnt_want_write
> inode_lock_nested(dir, I_MUTEX_PARENT)
> d_lookup -> NULL
> d_alloc_parallel in-lookup dentry
> o_create_mode S_IFDIR | 0755 & ~umask
> may_o_create
> security_path_mkdir
> inode_permission(MAY_WRITE|MAY_EXEC)
> kernfs_iop_permission
> generic_permission owner of a 0755 dir: allowed
> security_inode_mkdir
> dir_inode->i_op->lookup
> kernfs_iop_lookup
> kernfs_find_ns -> NULL node does not exist yet
> d_splice_alias(NULL, dentry) -> __d_add(): hashed, NEGATIVE
> try_break_deleg
> dir->i_op->mkdir
> kernfs_iop_mkdir
> scops->mkdir
> cgroup_mkdir no capable() check
> cgroup_create owner = current_fsuid()
> css_populate_dir
> kernfs_activate cgroup now exists
> return ERR_PTR(0) == NULL, dentry untouched
> de == NULL -> keep original dentry
> fsnotify_mkdir(dir, dentry) fine with a negative dentry
> dentry = res still negative
> file->f_mode |= FMODE_CREATED
> inode_unlock(dir); mnt_drop_write
> FMODE_CREATED set ->
> dput(nd->path.dentry)
> nd->path.dentry = dentry negative dentry becomes the "opened" path
> return NULL
> do_open
> open_flag & O_CREAT ->
> may_create_in_sticky(idmap, nd, d_backing_inode(nd->path.dentry))
> inode == NULL
>
> Scenario 1, parent not sticky: cgroup created, ENOTDIR returned
>
> may_create_in_sticky
> if (!(dir_mode & S_ISVTX)) return 0; inode never touched
> (nd->flags & LOOKUP_DIRECTORY) && !d_can_lookup(nd->path.dentry)
> DCACHE_MISS_TYPE -> false
> return -ENOTDIR
> terminate_walk
> fput_close(file)
> return ERR_PTR(-ENOTDIR) "child" cgroup stays behind
>
> A second openat of the same name works because kernfs_dop_revalidate() sees the parent's revision changed, d_invalidate()s the stale negative dentry, and the fresh kernfs_iop_lookup() now
> finds the node.
>
> Scenario 2, sticky parent: NULL dereference
>
> Setup, one-time, as root (what systemd's Delegate=yes does for user.slice/user-1000.slice/user@1000.service):
>
> fchown(pfd, 1000, 1000)
> chown_common -> notify_change -> kernfs_iop_setattr -> __kernfs_setattr
>
> Then as uid 1000:
>
> fchmod(pfd, 01755)
> chmod_common
> newattrs.ia_mode = (mode & S_IALLUGO) | ... S_ISVTX is inside S_IALLUGO
> notify_change
> setattr_prepare
> inode_owner_or_capable owner -> ok
> kernfs_iop_setattr
> __kernfs_setattr kn->mode = ia_mode, no masking
> setattr_copy inode->i_mode = 01755
>
> and the open, same chain as above until may_create_in_sticky(), now with nd->dir_mode = 01755:
>
> may_create_in_sticky
> if (!(dir_mode & S_ISVTX)) return 0; not taken
> if (S_ISREG(inode->i_mode) && ...) inode == NULL
> -> KASAN null-ptr-deref
>
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-18 10:07 ` NeilBrown
@ 2026-09-19 0:01 ` NeilBrown
2026-09-25 13:57 ` Christian Brauner
2026-09-29 11:54 ` Jori Koolstra
0 siblings, 2 replies; 40+ messages in thread
From: NeilBrown @ 2026-09-19 0:01 UTC (permalink / raw)
To: Christian Brauner
Cc: Jori Koolstra, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
On Fri, 18 Sep 2026, NeilBrown wrote:
> On Fri, 18 Sep 2026, Christian Brauner wrote:
> > On Sun, Sep 13, 2026 at 08:50:11PM +0200, Jori Koolstra wrote:
> > > Currently there is no way to race-freely create and open a directory.
> > > For regular files we have open(O_CREAT) for creating a new file inode,
> > > and returning a pinning fd to it. The lack of such functionality for
> > > directories means that when populating a directory tree there's always
> > > a race involved: the inodes first need to be created, and then opened
> > > to adjust their permissions/ownership/labels/timestamps/acls/xattrs/...,
> > > but in the time window between the creation and the opening they might
> > > be replaced by something else.
> > >
> > > Addressing this race without a proper API is only partially possible:
> > > the caller can immediately fstat() what was opened to verify that it
> > > has the expected inode type, owner and mode. But besides being easy to
> > > get wrong, this cannot establish who created the directory: a directory
> > > created by another process with identical credentials is
> > > indistinguishable from one the caller created itself, so the caller
> > > cannot tell whether the directory is its own to manage.
> > >
> > > Historically, the O_CREAT|O_DIRECTORY behaviour was to return ENOTDIR if
> > > a regular file exists at the open path; EISDIR if a directory exists at
> > > the path; and to create a regular file if no file exists at the path.
> > > This behaviour changed accidentally with
> > > commit 973d4b73fbaf ("do_last(): rejoin the common path even earlier in
> > > FMODE_{OPENED,CREATED} case") causing ENOTDIR to return in the last case
> > > while still creating the file. As this change was not detected for a
> > > long time, Brauner proposed to adopt the more consistent NetBSD
> > > behaviour, i.e. to return EINVAL on the O_CREAT|O_DIRECTORY combination.
> > > This change was applied in commit 43b450632676 ("open: return EINVAL for
> > > O_DIRECTORY | O_CREAT") in March, 2023. As the EINVAL behaviour has been
> > > in the kernel for about 3 years now, no rollback is expected as a result
> > > of userspace reliance on old behaviour, leaving us free to reassign the
> > > O_CREAT|O_DIRECTORY semantics.
> > >
> > > O_CREAT|O_DIRECTORY is made to reduce to a lookup on ->atomic_open()
> > > filesystems. These filesystems currently cannot handle
> > > O_CREAT|O_DIRECTORY without protocol extensions and therefore are forced
> > > into a fallback mode by stripping the O_CREAT bit. This causes existing
> > > directories to be successfully opened, while for targets that should
> > > have been created, -ENOENT is returned. This -ENOENT is then converted
> > > to -EOPNOTSUPP in later atomic_open(). The simple option of just
> > > returning -EOPNOTSUPP directly leads to inconsistent behaviour: before
> > > ->atomic_open() is called in lookup_open(), the dcache is queried. So
> > > returning -EOPNOTSUPP immediately would make O_CREAT|O_DIRECTORY
> > > dependent on the cache state of the dentry.
> > >
> > > There is no separate sysctl for directory creation implemented currently.
> > > Therefore, for the S_ISDIR case, disabling sysctl_protected_regular is
> > > not enough to allow creating a directory in a sticky folder, because that
> > > may surprise users not expecting that O_CREAT|O_DIRECTORY is possible on
> > > newer kernels.
> > >
> > > This feature idea (and some of its description) is taken from the
> > > UAPI group:
> > > https://github.com/uapi-group/kernel-features?tab=readme-ov-file#race-free-creation-and-opening-of-non-file-inodes
> > >
> > > Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
> > > ---
> > > fs/namei.c | 116 +++++++++++++++++++++++++++++++++++-------
> > > fs/open.c | 25 +++++----
> > > include/linux/fcntl.h | 6 +++
> > > 3 files changed, 117 insertions(+), 30 deletions(-)
> > >
> > > diff --git a/fs/namei.c b/fs/namei.c
> > > index 0efd395a1a65..6ff0a3c04f02 100644
> > > --- a/fs/namei.c
> > > +++ b/fs/namei.c
> > > @@ -1382,13 +1382,13 @@ int may_linkat(struct mnt_idmap *idmap, const struct path *link)
> > >
> > > /**
> > > * may_create_in_sticky - Check whether an O_CREAT open in a sticky directory
> > > - * should be allowed, or not, on files that already
> > > - * exist.
> > > + * should be allowed, or not, on files/directories that
> > > + * already exist.
> > > * @idmap: idmap of the mount the inode was found from
> > > * @nd: nameidata pathwalk data
> > > * @inode: the inode of the file to open
> > > *
> > > - * Block an O_CREAT open of a FIFO (or a regular file) when:
> > > + * Block an O_CREAT open of a FIFO (or a regular file/directory) when:
> > > * - sysctl_protected_fifos (or sysctl_protected_regular) is enabled
> > > * - the file already exists
> > > * - we are in a sticky directory
> > > @@ -1416,6 +1416,14 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> > > if (likely(!(dir_mode & S_ISVTX)))
> > > return 0;
> > >
> > > + /*
> > > + * There is no separate sysctl for directory creation in sticky
> > > + * folders. Therefore, for the S_ISDIR case, disabling
> > > + * sysctl_protected_regular is not enough to allow creating a
> > > + * directory in a sticky folder, because that may surprise users
> > > + * not expecting that O_CREAT|O_DIRECTORY is possible on newer
> > > + * kernels.
> > > + */
> > > if (S_ISREG(inode->i_mode) && !sysctl_protected_regular)
> > > return 0;
> > >
> > > @@ -1447,6 +1455,12 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> > > "sticky_create_regular");
> > > return -EACCES;
> > > }
> > > +
> > > + if (S_ISDIR(inode->i_mode)) {
> > > + audit_log_path_denied(AUDIT_ANOM_CREAT,
> > > + "sticky_create_dir");
> > > + return -EACCES;
> > > + }
> > > }
> > >
> > > return 0;
> > > @@ -4334,21 +4348,43 @@ static inline int open_to_namei_flags(int flag)
> > >
> > > static int may_o_create(struct mnt_idmap *idmap,
> > > const struct path *dir, struct dentry *dentry,
> > > - umode_t mode)
> > > + int open_flag, umode_t mode)
> > > {
> > > - int error = security_path_mknod(dir, dentry, mode, 0);
> > > + struct inode *dir_inode = dir->dentry->d_inode;
> > > + bool create_dir = O_IS_MKDIR(open_flag);
> > > + int error;
> > > +
> > > + WARN_ON_ONCE(create_dir && !(mode & S_IFDIR));
> > > +
> > > + if (create_dir)
> > > + error = security_path_mkdir(dir, dentry, mode);
> > > + else
> > > + error = security_path_mknod(dir, dentry, mode, 0);
> > > if (error)
> > > return error;
> > >
> > > if (!fsuidgid_has_mapping(dir->dentry->d_sb, idmap))
> > > return -EOVERFLOW;
> > >
> > > - error = inode_permission(idmap, dir->dentry->d_inode,
> > > - MAY_WRITE | MAY_EXEC);
> > > + error = inode_permission(idmap, dir_inode, MAY_WRITE | MAY_EXEC);
> > > if (error)
> > > return error;
> > >
> > > - return security_inode_create(dir->dentry->d_inode, dentry, mode);
> > > + if (create_dir)
> > > + error = security_inode_mkdir(dir_inode, dentry, mode);
> > > + else
> > > + error = security_inode_create(dir_inode, dentry, mode);
> > > +
> > > + return error;
> > > +}
> > > +
> > > +static inline umode_t o_create_mode(struct mnt_idmap *idmap,
> > > + const struct inode *dir, int open_flag, umode_t mode)
> > > +{
> > > + if (O_IS_MKDIR(open_flag))
> > > + return vfs_prepare_mode(idmap, dir, mode, S_IRWXUGO | S_ISVTX, S_IFDIR);
> > > + else
> > > + return vfs_prepare_mode(idmap, dir, mode, S_IALLUGO, S_IFREG);
> > > }
> > >
> > > /**
> > > @@ -4384,8 +4420,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > >
> > > file->__f_path.dentry = DENTRY_NOT_SET;
> > > file->__f_path.mnt = path->mnt;
> > > +
> > > error = dir_inode->i_op->atomic_open(dir_inode, dentry, file,
> > > - open_to_namei_flags(open_flag), mode);
> > > + open_to_namei_flags(open_flag), mode);
> > > d_lookup_done(dentry);
> > >
> > > if (!error) {
> > > @@ -4427,12 +4464,32 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > > */
> > > audit_inode_child(dir_inode, dentry, AUDIT_TYPE_CHILD_CREATE);
> > > error = create_error;
> > > + } else if (O_IS_MKDIR(open_flag) && error == -ENOENT) {
> > > + /*
> > > + * If the underlying filesystem does not implement
> > > + * O_CREAT|O_DIRECTORY, it strips the O_CREAT bit and
> > > + * continues as a lookup. We can't simply return
> > > + * -EOPNOTSUPP from unsupported ->atomic_open()
> > > + * implementations because the dentry might be in the
> > > + * dcache. In that case, lookup_open() returns before
> > > + * reaching ->atomic_open(), and hence whether you get
> > > + * -EOPNOTSUPP on O_CREAT|O_DIRECTORY would not only
> > > + * depend on the underlying filesystem, but also on
> > > + * the state of the dcache. Still, we must make an
> > > + * effort to differentiate a regular -ENOENT from the
> > > + * unsupported O_CREAT|O_DIRECTORY case.
> > > + */
> > > + error = -EOPNOTSUPP;
> > > }
> > > dput(dentry);
> > > dentry = ERR_PTR(error);
> > > } else {
> > > - if (file->f_mode & FMODE_CREATED)
> > > - fsnotify_create(dir_inode, dentry);
> > > + if (file->f_mode & FMODE_CREATED) {
> > > + if (d_is_dir(dentry))
> > > + fsnotify_mkdir(dir_inode, dentry);
> > > + else
> > > + fsnotify_create(dir_inode, dentry);
> > > + }
> > > if (file->f_mode & FMODE_OPENED)
> > > fsnotify_open(file);
> > > }
> > > @@ -4441,6 +4498,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > > return dentry;
> > > }
> > >
> > > +static inline
> > > +struct dentry *vfs_mkdir_no_perm(struct mnt_idmap *, struct inode *, struct dentry *,
> > > + umode_t, struct delegated_inode *);
> > > /*
> > > * Look up and maybe create and open the last component.
> > > *
> > > @@ -4462,6 +4522,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > struct mnt_idmap *idmap;
> > > struct dentry *dir = nd->path.dentry;
> > > struct inode *dir_inode = dir->d_inode;
> > > + bool create_dir = O_IS_MKDIR(op->open_flag);
> > > int open_flag;
> > > struct dentry *dentry;
> > > int error, create_error;
> > > @@ -4474,6 +4535,9 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > mode = op->mode;
> > > create_error = 0;
> > >
> > > + if (create_dir && dir_inode->i_op->atomic_open)
> > > + open_flag &= ~O_CREAT;
> > > +
> > > if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) {
> > > got_write = !mnt_want_write(nd->path.mnt);
> > > /*
> > > @@ -4534,10 +4598,10 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > if (open_flag & O_CREAT) {
> > > if (open_flag & O_EXCL)
> > > open_flag &= ~O_TRUNC;
> > > - mode = vfs_prepare_mode(idmap, dir_inode, mode, mode, mode);
> > > + mode = o_create_mode(idmap, dir_inode, open_flag, mode);
> > > if (likely(got_write))
> > > create_error = may_o_create(idmap, &nd->path,
> > > - dentry, mode);
> > > + dentry, open_flag, mode);
> > > else
> > > create_error = -EROFS;
> > > }
> > > @@ -4582,12 +4646,25 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > goto out_dput;
> > > }
> > >
> > > - if (!dir_inode->i_op->create) {
> > > + /* mimic operation missing errnos of vfs_mkdir/vfs_create */
> > > + if (create_dir && !dir_inode->i_op->mkdir) {
> > > + error = -EPERM;
> > > + goto out_dput;
> > > + }
> > > + if (!create_dir && !dir_inode->i_op->create) {
> > > error = -EACCES;
> > > goto out_dput;
> > > }
> > >
> > > - error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
> > > + if (create_dir) {
> > > + struct dentry *res = vfs_mkdir_no_perm(idmap, dir_inode, dentry, mode,
> > > + &delegated_inode);
> >
> > So, I think this is broken. Whatever vfs_mkdir_no_perm() returns is
> > passed to do_open(). Kernfs makes that buggy.
> >
> > cgroup, cgroup2, and resctrl are all implemented on top of kernfs. And
> > kernfs ->mkdir:: iop never instantiates the dentry.
> >
> > So that means e.g.,
> >
> > openat(cgroup_dir, "subdir", O_CREAT|O_DIRECTORY) creates a cgroup
> > and then fails with ENOTDIR.
> >
> > So the negative dentry gets handed out and now userspace holds an fd
> > with that negative dentry. So say userspace does fchown() to 1000 and
> > then fchmod() with the sticky bit and then you get a NULL deref. I have
> > reproduced this.
> >
> > Neil can correct me but the fix might be to check whether the dentry is
> > negative in lookup_open() and re-lookup nd->last with the parent still locked.
> > I think that's what nfsd_create_locked() and cachefiles_get_directory() do
> > after vfs_mkdir().
>
> nfsd_create_locked() used to do that before vfs_mkdir() could return a
> dentry, but it doesn't any more. The reason was because
> d_splice_alias() on might return a different dentry.
> In this case we want the same dentry, but we need to do a lookup on it.
>
> I'd rather fix this in kernfs, but maybe that is a longer-term goal.
>
> The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> I'm not certain that is needed if we keep the parent locked, but we
> would need to be certain.
> We at least need to d_drop() the dentry before ->lookup as ->lookup
> cannot handle hashed dentries and a hashed-negative dentry is passed
> to ->mkdir.
>
> I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> probably not.
>
> Summary: I think that if vfs_mkdir() returns NULL (success) but the
> dentry is negative, we need to d_drop() and call ->lookup with a big
> comment about kernfs. But we need to double-check that this will do the
> right thing with ->d_time (I think it will).
> We also need to think carefully about races with
> kernfs_dop_revalidate(), which could happen concurrently with the
> ->lookup.
I've thought a bit more about this ... I think that doing a lookup after
the vfs_mkdir() results in a negative is a bit ugly. It assumes things
about the fs that I would rather not assume.
I would rather have the current proposed code check for a negative
dentry, and fail with -EIO or similar.
We could then "fix" kernfs by providing an atomic_open which does the
mkdir and then the lookup, and provides the dentry to
finish_no_lookup().
That way we don't need to change kernfs mkdir.
Note that tracefs_syscall_mkdir() has the same behaviour as
kernfs_iop_mkdir, and could have the same fix.
NeilBrown
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-19 0:01 ` NeilBrown
@ 2026-09-25 13:57 ` Christian Brauner
2026-09-25 21:26 ` NeilBrown
2026-09-29 11:54 ` Jori Koolstra
1 sibling, 1 reply; 40+ messages in thread
From: Christian Brauner @ 2026-09-25 13:57 UTC (permalink / raw)
To: NeilBrown
Cc: Jori Koolstra, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
On Sat, Sep 19, 2026 at 10:01:34AM +1000, NeilBrown wrote:
> On Fri, 18 Sep 2026, NeilBrown wrote:
> > On Fri, 18 Sep 2026, Christian Brauner wrote:
> > > On Sun, Sep 13, 2026 at 08:50:11PM +0200, Jori Koolstra wrote:
> > > > Currently there is no way to race-freely create and open a directory.
> > > > For regular files we have open(O_CREAT) for creating a new file inode,
> > > > and returning a pinning fd to it. The lack of such functionality for
> > > > directories means that when populating a directory tree there's always
> > > > a race involved: the inodes first need to be created, and then opened
> > > > to adjust their permissions/ownership/labels/timestamps/acls/xattrs/...,
> > > > but in the time window between the creation and the opening they might
> > > > be replaced by something else.
> > > >
> > > > Addressing this race without a proper API is only partially possible:
> > > > the caller can immediately fstat() what was opened to verify that it
> > > > has the expected inode type, owner and mode. But besides being easy to
> > > > get wrong, this cannot establish who created the directory: a directory
> > > > created by another process with identical credentials is
> > > > indistinguishable from one the caller created itself, so the caller
> > > > cannot tell whether the directory is its own to manage.
> > > >
> > > > Historically, the O_CREAT|O_DIRECTORY behaviour was to return ENOTDIR if
> > > > a regular file exists at the open path; EISDIR if a directory exists at
> > > > the path; and to create a regular file if no file exists at the path.
> > > > This behaviour changed accidentally with
> > > > commit 973d4b73fbaf ("do_last(): rejoin the common path even earlier in
> > > > FMODE_{OPENED,CREATED} case") causing ENOTDIR to return in the last case
> > > > while still creating the file. As this change was not detected for a
> > > > long time, Brauner proposed to adopt the more consistent NetBSD
> > > > behaviour, i.e. to return EINVAL on the O_CREAT|O_DIRECTORY combination.
> > > > This change was applied in commit 43b450632676 ("open: return EINVAL for
> > > > O_DIRECTORY | O_CREAT") in March, 2023. As the EINVAL behaviour has been
> > > > in the kernel for about 3 years now, no rollback is expected as a result
> > > > of userspace reliance on old behaviour, leaving us free to reassign the
> > > > O_CREAT|O_DIRECTORY semantics.
> > > >
> > > > O_CREAT|O_DIRECTORY is made to reduce to a lookup on ->atomic_open()
> > > > filesystems. These filesystems currently cannot handle
> > > > O_CREAT|O_DIRECTORY without protocol extensions and therefore are forced
> > > > into a fallback mode by stripping the O_CREAT bit. This causes existing
> > > > directories to be successfully opened, while for targets that should
> > > > have been created, -ENOENT is returned. This -ENOENT is then converted
> > > > to -EOPNOTSUPP in later atomic_open(). The simple option of just
> > > > returning -EOPNOTSUPP directly leads to inconsistent behaviour: before
> > > > ->atomic_open() is called in lookup_open(), the dcache is queried. So
> > > > returning -EOPNOTSUPP immediately would make O_CREAT|O_DIRECTORY
> > > > dependent on the cache state of the dentry.
> > > >
> > > > There is no separate sysctl for directory creation implemented currently.
> > > > Therefore, for the S_ISDIR case, disabling sysctl_protected_regular is
> > > > not enough to allow creating a directory in a sticky folder, because that
> > > > may surprise users not expecting that O_CREAT|O_DIRECTORY is possible on
> > > > newer kernels.
> > > >
> > > > This feature idea (and some of its description) is taken from the
> > > > UAPI group:
> > > > https://github.com/uapi-group/kernel-features?tab=readme-ov-file#race-free-creation-and-opening-of-non-file-inodes
> > > >
> > > > Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
> > > > ---
> > > > fs/namei.c | 116 +++++++++++++++++++++++++++++++++++-------
> > > > fs/open.c | 25 +++++----
> > > > include/linux/fcntl.h | 6 +++
> > > > 3 files changed, 117 insertions(+), 30 deletions(-)
> > > >
> > > > diff --git a/fs/namei.c b/fs/namei.c
> > > > index 0efd395a1a65..6ff0a3c04f02 100644
> > > > --- a/fs/namei.c
> > > > +++ b/fs/namei.c
> > > > @@ -1382,13 +1382,13 @@ int may_linkat(struct mnt_idmap *idmap, const struct path *link)
> > > >
> > > > /**
> > > > * may_create_in_sticky - Check whether an O_CREAT open in a sticky directory
> > > > - * should be allowed, or not, on files that already
> > > > - * exist.
> > > > + * should be allowed, or not, on files/directories that
> > > > + * already exist.
> > > > * @idmap: idmap of the mount the inode was found from
> > > > * @nd: nameidata pathwalk data
> > > > * @inode: the inode of the file to open
> > > > *
> > > > - * Block an O_CREAT open of a FIFO (or a regular file) when:
> > > > + * Block an O_CREAT open of a FIFO (or a regular file/directory) when:
> > > > * - sysctl_protected_fifos (or sysctl_protected_regular) is enabled
> > > > * - the file already exists
> > > > * - we are in a sticky directory
> > > > @@ -1416,6 +1416,14 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> > > > if (likely(!(dir_mode & S_ISVTX)))
> > > > return 0;
> > > >
> > > > + /*
> > > > + * There is no separate sysctl for directory creation in sticky
> > > > + * folders. Therefore, for the S_ISDIR case, disabling
> > > > + * sysctl_protected_regular is not enough to allow creating a
> > > > + * directory in a sticky folder, because that may surprise users
> > > > + * not expecting that O_CREAT|O_DIRECTORY is possible on newer
> > > > + * kernels.
> > > > + */
> > > > if (S_ISREG(inode->i_mode) && !sysctl_protected_regular)
> > > > return 0;
> > > >
> > > > @@ -1447,6 +1455,12 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> > > > "sticky_create_regular");
> > > > return -EACCES;
> > > > }
> > > > +
> > > > + if (S_ISDIR(inode->i_mode)) {
> > > > + audit_log_path_denied(AUDIT_ANOM_CREAT,
> > > > + "sticky_create_dir");
> > > > + return -EACCES;
> > > > + }
> > > > }
> > > >
> > > > return 0;
> > > > @@ -4334,21 +4348,43 @@ static inline int open_to_namei_flags(int flag)
> > > >
> > > > static int may_o_create(struct mnt_idmap *idmap,
> > > > const struct path *dir, struct dentry *dentry,
> > > > - umode_t mode)
> > > > + int open_flag, umode_t mode)
> > > > {
> > > > - int error = security_path_mknod(dir, dentry, mode, 0);
> > > > + struct inode *dir_inode = dir->dentry->d_inode;
> > > > + bool create_dir = O_IS_MKDIR(open_flag);
> > > > + int error;
> > > > +
> > > > + WARN_ON_ONCE(create_dir && !(mode & S_IFDIR));
> > > > +
> > > > + if (create_dir)
> > > > + error = security_path_mkdir(dir, dentry, mode);
> > > > + else
> > > > + error = security_path_mknod(dir, dentry, mode, 0);
> > > > if (error)
> > > > return error;
> > > >
> > > > if (!fsuidgid_has_mapping(dir->dentry->d_sb, idmap))
> > > > return -EOVERFLOW;
> > > >
> > > > - error = inode_permission(idmap, dir->dentry->d_inode,
> > > > - MAY_WRITE | MAY_EXEC);
> > > > + error = inode_permission(idmap, dir_inode, MAY_WRITE | MAY_EXEC);
> > > > if (error)
> > > > return error;
> > > >
> > > > - return security_inode_create(dir->dentry->d_inode, dentry, mode);
> > > > + if (create_dir)
> > > > + error = security_inode_mkdir(dir_inode, dentry, mode);
> > > > + else
> > > > + error = security_inode_create(dir_inode, dentry, mode);
> > > > +
> > > > + return error;
> > > > +}
> > > > +
> > > > +static inline umode_t o_create_mode(struct mnt_idmap *idmap,
> > > > + const struct inode *dir, int open_flag, umode_t mode)
> > > > +{
> > > > + if (O_IS_MKDIR(open_flag))
> > > > + return vfs_prepare_mode(idmap, dir, mode, S_IRWXUGO | S_ISVTX, S_IFDIR);
> > > > + else
> > > > + return vfs_prepare_mode(idmap, dir, mode, S_IALLUGO, S_IFREG);
> > > > }
> > > >
> > > > /**
> > > > @@ -4384,8 +4420,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > > >
> > > > file->__f_path.dentry = DENTRY_NOT_SET;
> > > > file->__f_path.mnt = path->mnt;
> > > > +
> > > > error = dir_inode->i_op->atomic_open(dir_inode, dentry, file,
> > > > - open_to_namei_flags(open_flag), mode);
> > > > + open_to_namei_flags(open_flag), mode);
> > > > d_lookup_done(dentry);
> > > >
> > > > if (!error) {
> > > > @@ -4427,12 +4464,32 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > > > */
> > > > audit_inode_child(dir_inode, dentry, AUDIT_TYPE_CHILD_CREATE);
> > > > error = create_error;
> > > > + } else if (O_IS_MKDIR(open_flag) && error == -ENOENT) {
> > > > + /*
> > > > + * If the underlying filesystem does not implement
> > > > + * O_CREAT|O_DIRECTORY, it strips the O_CREAT bit and
> > > > + * continues as a lookup. We can't simply return
> > > > + * -EOPNOTSUPP from unsupported ->atomic_open()
> > > > + * implementations because the dentry might be in the
> > > > + * dcache. In that case, lookup_open() returns before
> > > > + * reaching ->atomic_open(), and hence whether you get
> > > > + * -EOPNOTSUPP on O_CREAT|O_DIRECTORY would not only
> > > > + * depend on the underlying filesystem, but also on
> > > > + * the state of the dcache. Still, we must make an
> > > > + * effort to differentiate a regular -ENOENT from the
> > > > + * unsupported O_CREAT|O_DIRECTORY case.
> > > > + */
> > > > + error = -EOPNOTSUPP;
> > > > }
> > > > dput(dentry);
> > > > dentry = ERR_PTR(error);
> > > > } else {
> > > > - if (file->f_mode & FMODE_CREATED)
> > > > - fsnotify_create(dir_inode, dentry);
> > > > + if (file->f_mode & FMODE_CREATED) {
> > > > + if (d_is_dir(dentry))
> > > > + fsnotify_mkdir(dir_inode, dentry);
> > > > + else
> > > > + fsnotify_create(dir_inode, dentry);
> > > > + }
> > > > if (file->f_mode & FMODE_OPENED)
> > > > fsnotify_open(file);
> > > > }
> > > > @@ -4441,6 +4498,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > > > return dentry;
> > > > }
> > > >
> > > > +static inline
> > > > +struct dentry *vfs_mkdir_no_perm(struct mnt_idmap *, struct inode *, struct dentry *,
> > > > + umode_t, struct delegated_inode *);
> > > > /*
> > > > * Look up and maybe create and open the last component.
> > > > *
> > > > @@ -4462,6 +4522,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > > struct mnt_idmap *idmap;
> > > > struct dentry *dir = nd->path.dentry;
> > > > struct inode *dir_inode = dir->d_inode;
> > > > + bool create_dir = O_IS_MKDIR(op->open_flag);
> > > > int open_flag;
> > > > struct dentry *dentry;
> > > > int error, create_error;
> > > > @@ -4474,6 +4535,9 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > > mode = op->mode;
> > > > create_error = 0;
> > > >
> > > > + if (create_dir && dir_inode->i_op->atomic_open)
> > > > + open_flag &= ~O_CREAT;
> > > > +
> > > > if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) {
> > > > got_write = !mnt_want_write(nd->path.mnt);
> > > > /*
> > > > @@ -4534,10 +4598,10 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > > if (open_flag & O_CREAT) {
> > > > if (open_flag & O_EXCL)
> > > > open_flag &= ~O_TRUNC;
> > > > - mode = vfs_prepare_mode(idmap, dir_inode, mode, mode, mode);
> > > > + mode = o_create_mode(idmap, dir_inode, open_flag, mode);
> > > > if (likely(got_write))
> > > > create_error = may_o_create(idmap, &nd->path,
> > > > - dentry, mode);
> > > > + dentry, open_flag, mode);
> > > > else
> > > > create_error = -EROFS;
> > > > }
> > > > @@ -4582,12 +4646,25 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > > goto out_dput;
> > > > }
> > > >
> > > > - if (!dir_inode->i_op->create) {
> > > > + /* mimic operation missing errnos of vfs_mkdir/vfs_create */
> > > > + if (create_dir && !dir_inode->i_op->mkdir) {
> > > > + error = -EPERM;
> > > > + goto out_dput;
> > > > + }
> > > > + if (!create_dir && !dir_inode->i_op->create) {
> > > > error = -EACCES;
> > > > goto out_dput;
> > > > }
> > > >
> > > > - error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
> > > > + if (create_dir) {
> > > > + struct dentry *res = vfs_mkdir_no_perm(idmap, dir_inode, dentry, mode,
> > > > + &delegated_inode);
> > >
> > > So, I think this is broken. Whatever vfs_mkdir_no_perm() returns is
> > > passed to do_open(). Kernfs makes that buggy.
> > >
> > > cgroup, cgroup2, and resctrl are all implemented on top of kernfs. And
> > > kernfs ->mkdir:: iop never instantiates the dentry.
> > >
> > > So that means e.g.,
> > >
> > > openat(cgroup_dir, "subdir", O_CREAT|O_DIRECTORY) creates a cgroup
> > > and then fails with ENOTDIR.
> > >
> > > So the negative dentry gets handed out and now userspace holds an fd
> > > with that negative dentry. So say userspace does fchown() to 1000 and
> > > then fchmod() with the sticky bit and then you get a NULL deref. I have
> > > reproduced this.
> > >
> > > Neil can correct me but the fix might be to check whether the dentry is
> > > negative in lookup_open() and re-lookup nd->last with the parent still locked.
> > > I think that's what nfsd_create_locked() and cachefiles_get_directory() do
> > > after vfs_mkdir().
> >
> > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > dentry, but it doesn't any more. The reason was because
> > d_splice_alias() on might return a different dentry.
> > In this case we want the same dentry, but we need to do a lookup on it.
> >
> > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> >
> > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > I'm not certain that is needed if we keep the parent locked, but we
> > would need to be certain.
> > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > cannot handle hashed dentries and a hashed-negative dentry is passed
> > to ->mkdir.
> >
> > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > probably not.
> >
> > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > dentry is negative, we need to d_drop() and call ->lookup with a big
> > comment about kernfs. But we need to double-check that this will do the
> > right thing with ->d_time (I think it will).
> > We also need to think carefully about races with
> > kernfs_dop_revalidate(), which could happen concurrently with the
> > ->lookup.
>
> I've thought a bit more about this ... I think that doing a lookup after
> the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> about the fs that I would rather not assume.
>
> I would rather have the current proposed code check for a negative
> dentry, and fail with -EIO or similar.
>
> We could then "fix" kernfs by providing an atomic_open which does the
> mkdir and then the lookup, and provides the dentry to
> finish_no_lookup().
>
> That way we don't need to change kernfs mkdir.
>
> Note that tracefs_syscall_mkdir() has the same behaviour as
> kernfs_iop_mkdir, and could have the same fix.
Sounds good to me. But should we do this in one single release so
userspace doesn't have a 90% working thing?
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-25 13:57 ` Christian Brauner
@ 2026-09-25 21:26 ` NeilBrown
2026-09-25 22:53 ` Jori Koolstra
0 siblings, 1 reply; 40+ messages in thread
From: NeilBrown @ 2026-09-25 21:26 UTC (permalink / raw)
To: Christian Brauner
Cc: Jori Koolstra, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
On Fri, 25 Sep 2026, Christian Brauner wrote:
> On Sat, Sep 19, 2026 at 10:01:34AM +1000, NeilBrown wrote:
> > On Fri, 18 Sep 2026, NeilBrown wrote:
> > > On Fri, 18 Sep 2026, Christian Brauner wrote:
> > > > On Sun, Sep 13, 2026 at 08:50:11PM +0200, Jori Koolstra wrote:
> > > > > Currently there is no way to race-freely create and open a directory.
> > > > > For regular files we have open(O_CREAT) for creating a new file inode,
> > > > > and returning a pinning fd to it. The lack of such functionality for
> > > > > directories means that when populating a directory tree there's always
> > > > > a race involved: the inodes first need to be created, and then opened
> > > > > to adjust their permissions/ownership/labels/timestamps/acls/xattrs/...,
> > > > > but in the time window between the creation and the opening they might
> > > > > be replaced by something else.
> > > > >
> > > > > Addressing this race without a proper API is only partially possible:
> > > > > the caller can immediately fstat() what was opened to verify that it
> > > > > has the expected inode type, owner and mode. But besides being easy to
> > > > > get wrong, this cannot establish who created the directory: a directory
> > > > > created by another process with identical credentials is
> > > > > indistinguishable from one the caller created itself, so the caller
> > > > > cannot tell whether the directory is its own to manage.
> > > > >
> > > > > Historically, the O_CREAT|O_DIRECTORY behaviour was to return ENOTDIR if
> > > > > a regular file exists at the open path; EISDIR if a directory exists at
> > > > > the path; and to create a regular file if no file exists at the path.
> > > > > This behaviour changed accidentally with
> > > > > commit 973d4b73fbaf ("do_last(): rejoin the common path even earlier in
> > > > > FMODE_{OPENED,CREATED} case") causing ENOTDIR to return in the last case
> > > > > while still creating the file. As this change was not detected for a
> > > > > long time, Brauner proposed to adopt the more consistent NetBSD
> > > > > behaviour, i.e. to return EINVAL on the O_CREAT|O_DIRECTORY combination.
> > > > > This change was applied in commit 43b450632676 ("open: return EINVAL for
> > > > > O_DIRECTORY | O_CREAT") in March, 2023. As the EINVAL behaviour has been
> > > > > in the kernel for about 3 years now, no rollback is expected as a result
> > > > > of userspace reliance on old behaviour, leaving us free to reassign the
> > > > > O_CREAT|O_DIRECTORY semantics.
> > > > >
> > > > > O_CREAT|O_DIRECTORY is made to reduce to a lookup on ->atomic_open()
> > > > > filesystems. These filesystems currently cannot handle
> > > > > O_CREAT|O_DIRECTORY without protocol extensions and therefore are forced
> > > > > into a fallback mode by stripping the O_CREAT bit. This causes existing
> > > > > directories to be successfully opened, while for targets that should
> > > > > have been created, -ENOENT is returned. This -ENOENT is then converted
> > > > > to -EOPNOTSUPP in later atomic_open(). The simple option of just
> > > > > returning -EOPNOTSUPP directly leads to inconsistent behaviour: before
> > > > > ->atomic_open() is called in lookup_open(), the dcache is queried. So
> > > > > returning -EOPNOTSUPP immediately would make O_CREAT|O_DIRECTORY
> > > > > dependent on the cache state of the dentry.
> > > > >
> > > > > There is no separate sysctl for directory creation implemented currently.
> > > > > Therefore, for the S_ISDIR case, disabling sysctl_protected_regular is
> > > > > not enough to allow creating a directory in a sticky folder, because that
> > > > > may surprise users not expecting that O_CREAT|O_DIRECTORY is possible on
> > > > > newer kernels.
> > > > >
> > > > > This feature idea (and some of its description) is taken from the
> > > > > UAPI group:
> > > > > https://github.com/uapi-group/kernel-features?tab=readme-ov-file#race-free-creation-and-opening-of-non-file-inodes
> > > > >
> > > > > Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
> > > > > ---
> > > > > fs/namei.c | 116 +++++++++++++++++++++++++++++++++++-------
> > > > > fs/open.c | 25 +++++----
> > > > > include/linux/fcntl.h | 6 +++
> > > > > 3 files changed, 117 insertions(+), 30 deletions(-)
> > > > >
> > > > > diff --git a/fs/namei.c b/fs/namei.c
> > > > > index 0efd395a1a65..6ff0a3c04f02 100644
> > > > > --- a/fs/namei.c
> > > > > +++ b/fs/namei.c
> > > > > @@ -1382,13 +1382,13 @@ int may_linkat(struct mnt_idmap *idmap, const struct path *link)
> > > > >
> > > > > /**
> > > > > * may_create_in_sticky - Check whether an O_CREAT open in a sticky directory
> > > > > - * should be allowed, or not, on files that already
> > > > > - * exist.
> > > > > + * should be allowed, or not, on files/directories that
> > > > > + * already exist.
> > > > > * @idmap: idmap of the mount the inode was found from
> > > > > * @nd: nameidata pathwalk data
> > > > > * @inode: the inode of the file to open
> > > > > *
> > > > > - * Block an O_CREAT open of a FIFO (or a regular file) when:
> > > > > + * Block an O_CREAT open of a FIFO (or a regular file/directory) when:
> > > > > * - sysctl_protected_fifos (or sysctl_protected_regular) is enabled
> > > > > * - the file already exists
> > > > > * - we are in a sticky directory
> > > > > @@ -1416,6 +1416,14 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> > > > > if (likely(!(dir_mode & S_ISVTX)))
> > > > > return 0;
> > > > >
> > > > > + /*
> > > > > + * There is no separate sysctl for directory creation in sticky
> > > > > + * folders. Therefore, for the S_ISDIR case, disabling
> > > > > + * sysctl_protected_regular is not enough to allow creating a
> > > > > + * directory in a sticky folder, because that may surprise users
> > > > > + * not expecting that O_CREAT|O_DIRECTORY is possible on newer
> > > > > + * kernels.
> > > > > + */
> > > > > if (S_ISREG(inode->i_mode) && !sysctl_protected_regular)
> > > > > return 0;
> > > > >
> > > > > @@ -1447,6 +1455,12 @@ static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd,
> > > > > "sticky_create_regular");
> > > > > return -EACCES;
> > > > > }
> > > > > +
> > > > > + if (S_ISDIR(inode->i_mode)) {
> > > > > + audit_log_path_denied(AUDIT_ANOM_CREAT,
> > > > > + "sticky_create_dir");
> > > > > + return -EACCES;
> > > > > + }
> > > > > }
> > > > >
> > > > > return 0;
> > > > > @@ -4334,21 +4348,43 @@ static inline int open_to_namei_flags(int flag)
> > > > >
> > > > > static int may_o_create(struct mnt_idmap *idmap,
> > > > > const struct path *dir, struct dentry *dentry,
> > > > > - umode_t mode)
> > > > > + int open_flag, umode_t mode)
> > > > > {
> > > > > - int error = security_path_mknod(dir, dentry, mode, 0);
> > > > > + struct inode *dir_inode = dir->dentry->d_inode;
> > > > > + bool create_dir = O_IS_MKDIR(open_flag);
> > > > > + int error;
> > > > > +
> > > > > + WARN_ON_ONCE(create_dir && !(mode & S_IFDIR));
> > > > > +
> > > > > + if (create_dir)
> > > > > + error = security_path_mkdir(dir, dentry, mode);
> > > > > + else
> > > > > + error = security_path_mknod(dir, dentry, mode, 0);
> > > > > if (error)
> > > > > return error;
> > > > >
> > > > > if (!fsuidgid_has_mapping(dir->dentry->d_sb, idmap))
> > > > > return -EOVERFLOW;
> > > > >
> > > > > - error = inode_permission(idmap, dir->dentry->d_inode,
> > > > > - MAY_WRITE | MAY_EXEC);
> > > > > + error = inode_permission(idmap, dir_inode, MAY_WRITE | MAY_EXEC);
> > > > > if (error)
> > > > > return error;
> > > > >
> > > > > - return security_inode_create(dir->dentry->d_inode, dentry, mode);
> > > > > + if (create_dir)
> > > > > + error = security_inode_mkdir(dir_inode, dentry, mode);
> > > > > + else
> > > > > + error = security_inode_create(dir_inode, dentry, mode);
> > > > > +
> > > > > + return error;
> > > > > +}
> > > > > +
> > > > > +static inline umode_t o_create_mode(struct mnt_idmap *idmap,
> > > > > + const struct inode *dir, int open_flag, umode_t mode)
> > > > > +{
> > > > > + if (O_IS_MKDIR(open_flag))
> > > > > + return vfs_prepare_mode(idmap, dir, mode, S_IRWXUGO | S_ISVTX, S_IFDIR);
> > > > > + else
> > > > > + return vfs_prepare_mode(idmap, dir, mode, S_IALLUGO, S_IFREG);
> > > > > }
> > > > >
> > > > > /**
> > > > > @@ -4384,8 +4420,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > > > >
> > > > > file->__f_path.dentry = DENTRY_NOT_SET;
> > > > > file->__f_path.mnt = path->mnt;
> > > > > +
> > > > > error = dir_inode->i_op->atomic_open(dir_inode, dentry, file,
> > > > > - open_to_namei_flags(open_flag), mode);
> > > > > + open_to_namei_flags(open_flag), mode);
> > > > > d_lookup_done(dentry);
> > > > >
> > > > > if (!error) {
> > > > > @@ -4427,12 +4464,32 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > > > > */
> > > > > audit_inode_child(dir_inode, dentry, AUDIT_TYPE_CHILD_CREATE);
> > > > > error = create_error;
> > > > > + } else if (O_IS_MKDIR(open_flag) && error == -ENOENT) {
> > > > > + /*
> > > > > + * If the underlying filesystem does not implement
> > > > > + * O_CREAT|O_DIRECTORY, it strips the O_CREAT bit and
> > > > > + * continues as a lookup. We can't simply return
> > > > > + * -EOPNOTSUPP from unsupported ->atomic_open()
> > > > > + * implementations because the dentry might be in the
> > > > > + * dcache. In that case, lookup_open() returns before
> > > > > + * reaching ->atomic_open(), and hence whether you get
> > > > > + * -EOPNOTSUPP on O_CREAT|O_DIRECTORY would not only
> > > > > + * depend on the underlying filesystem, but also on
> > > > > + * the state of the dcache. Still, we must make an
> > > > > + * effort to differentiate a regular -ENOENT from the
> > > > > + * unsupported O_CREAT|O_DIRECTORY case.
> > > > > + */
> > > > > + error = -EOPNOTSUPP;
> > > > > }
> > > > > dput(dentry);
> > > > > dentry = ERR_PTR(error);
> > > > > } else {
> > > > > - if (file->f_mode & FMODE_CREATED)
> > > > > - fsnotify_create(dir_inode, dentry);
> > > > > + if (file->f_mode & FMODE_CREATED) {
> > > > > + if (d_is_dir(dentry))
> > > > > + fsnotify_mkdir(dir_inode, dentry);
> > > > > + else
> > > > > + fsnotify_create(dir_inode, dentry);
> > > > > + }
> > > > > if (file->f_mode & FMODE_OPENED)
> > > > > fsnotify_open(file);
> > > > > }
> > > > > @@ -4441,6 +4498,9 @@ static struct dentry *atomic_open(const struct path *path, struct dentry *dentry
> > > > > return dentry;
> > > > > }
> > > > >
> > > > > +static inline
> > > > > +struct dentry *vfs_mkdir_no_perm(struct mnt_idmap *, struct inode *, struct dentry *,
> > > > > + umode_t, struct delegated_inode *);
> > > > > /*
> > > > > * Look up and maybe create and open the last component.
> > > > > *
> > > > > @@ -4462,6 +4522,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > > > struct mnt_idmap *idmap;
> > > > > struct dentry *dir = nd->path.dentry;
> > > > > struct inode *dir_inode = dir->d_inode;
> > > > > + bool create_dir = O_IS_MKDIR(op->open_flag);
> > > > > int open_flag;
> > > > > struct dentry *dentry;
> > > > > int error, create_error;
> > > > > @@ -4474,6 +4535,9 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > > > mode = op->mode;
> > > > > create_error = 0;
> > > > >
> > > > > + if (create_dir && dir_inode->i_op->atomic_open)
> > > > > + open_flag &= ~O_CREAT;
> > > > > +
> > > > > if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) {
> > > > > got_write = !mnt_want_write(nd->path.mnt);
> > > > > /*
> > > > > @@ -4534,10 +4598,10 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > > > if (open_flag & O_CREAT) {
> > > > > if (open_flag & O_EXCL)
> > > > > open_flag &= ~O_TRUNC;
> > > > > - mode = vfs_prepare_mode(idmap, dir_inode, mode, mode, mode);
> > > > > + mode = o_create_mode(idmap, dir_inode, open_flag, mode);
> > > > > if (likely(got_write))
> > > > > create_error = may_o_create(idmap, &nd->path,
> > > > > - dentry, mode);
> > > > > + dentry, open_flag, mode);
> > > > > else
> > > > > create_error = -EROFS;
> > > > > }
> > > > > @@ -4582,12 +4646,25 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > > > > goto out_dput;
> > > > > }
> > > > >
> > > > > - if (!dir_inode->i_op->create) {
> > > > > + /* mimic operation missing errnos of vfs_mkdir/vfs_create */
> > > > > + if (create_dir && !dir_inode->i_op->mkdir) {
> > > > > + error = -EPERM;
> > > > > + goto out_dput;
> > > > > + }
> > > > > + if (!create_dir && !dir_inode->i_op->create) {
> > > > > error = -EACCES;
> > > > > goto out_dput;
> > > > > }
> > > > >
> > > > > - error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
> > > > > + if (create_dir) {
> > > > > + struct dentry *res = vfs_mkdir_no_perm(idmap, dir_inode, dentry, mode,
> > > > > + &delegated_inode);
> > > >
> > > > So, I think this is broken. Whatever vfs_mkdir_no_perm() returns is
> > > > passed to do_open(). Kernfs makes that buggy.
> > > >
> > > > cgroup, cgroup2, and resctrl are all implemented on top of kernfs. And
> > > > kernfs ->mkdir:: iop never instantiates the dentry.
> > > >
> > > > So that means e.g.,
> > > >
> > > > openat(cgroup_dir, "subdir", O_CREAT|O_DIRECTORY) creates a cgroup
> > > > and then fails with ENOTDIR.
> > > >
> > > > So the negative dentry gets handed out and now userspace holds an fd
> > > > with that negative dentry. So say userspace does fchown() to 1000 and
> > > > then fchmod() with the sticky bit and then you get a NULL deref. I have
> > > > reproduced this.
> > > >
> > > > Neil can correct me but the fix might be to check whether the dentry is
> > > > negative in lookup_open() and re-lookup nd->last with the parent still locked.
> > > > I think that's what nfsd_create_locked() and cachefiles_get_directory() do
> > > > after vfs_mkdir().
> > >
> > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > dentry, but it doesn't any more. The reason was because
> > > d_splice_alias() on might return a different dentry.
> > > In this case we want the same dentry, but we need to do a lookup on it.
> > >
> > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > >
> > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > I'm not certain that is needed if we keep the parent locked, but we
> > > would need to be certain.
> > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > to ->mkdir.
> > >
> > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > probably not.
> > >
> > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > comment about kernfs. But we need to double-check that this will do the
> > > right thing with ->d_time (I think it will).
> > > We also need to think carefully about races with
> > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > ->lookup.
> >
> > I've thought a bit more about this ... I think that doing a lookup after
> > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > about the fs that I would rather not assume.
> >
> > I would rather have the current proposed code check for a negative
> > dentry, and fail with -EIO or similar.
> >
> > We could then "fix" kernfs by providing an atomic_open which does the
> > mkdir and then the lookup, and provides the dentry to
> > finish_no_lookup().
> >
> > That way we don't need to change kernfs mkdir.
> >
> > Note that tracefs_syscall_mkdir() has the same behaviour as
> > kernfs_iop_mkdir, and could have the same fix.
>
> Sounds good to me. But should we do this in one single release so
> userspace doesn't have a 90% working thing?
>
The current proposal leaves all ->atomic_open using filesystems as not
supporting O_CREAT|O_DIRECTORY and I think that is reasonable. Getting
them all done in the one release is probably unrealistic.
But that is a slightly different issue to kernfs/tracefs.
I wonder *why* those two filesystems don't do the lookup to instantiate
the dentry after mkdir. If it was just "unnecessary" then we can safely
change it. If it was "there is a common use-case where mkdir isn't
followed by a lookup, and we can avoid cluttering the icache/dcache",
then we probably don't want to.
In the first case, adding a ->lookup() at the end of mkdir and returning
the result would suffice. In the second case adding ->atomic_open
would be better. I cannot find any evidence in git history of anything
beyond "unnecessary".
Part of the point of O_CREAT|O_DIRECTORY is that it is atomic - the
thing opened is the thing created. tracefs drops the parent ->i_rwsem
while actually performing the creation. It isn't immediately clear what
that means for atomicity but it does raise questions.
At this stage I think I would lean towards O_CREAT|O_DIRECTORY not
working on these two filesystems. Neither support a .create
inode_operation, so providing a .atomic_open would be quite easy: just
do a ->lookup and pass the result to finish_no_open, but return an
appropriate error if O_CREAT was requested but no inode was found. That
would then treat these like other atomic_open filesystems in that
O_CREAT|O_DIRECTORY wouldn't work until the fs maintainer accepted a
patch for it.
NeilBrown
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-25 21:26 ` NeilBrown
@ 2026-09-25 22:53 ` Jori Koolstra
0 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-25 22:53 UTC (permalink / raw)
To: NeilBrown, NeilBrown, Christian Brauner
Cc: Jeff Layton, Al Viro, Aleksa Sarai, Amir Goldstein, Jan Kara,
linux-fsdevel, linux-kernel
> Op 25-09-2026 17:26 EDT schreef NeilBrown <neilb@ownmail.net>:
>
> The current proposal leaves all ->atomic_open using filesystems as not
> supporting O_CREAT|O_DIRECTORY and I think that is reasonable. Getting
> them all done in the one release is probably unrealistic.
> But that is a slightly different issue to kernfs/tracefs.
>
Exactly, we're not going to get 100% support in a single release anyway.
> I wonder *why* those two filesystems don't do the lookup to instantiate
> the dentry after mkdir. If it was just "unnecessary" then we can safely
> change it. If it was "there is a common use-case where mkdir isn't
> followed by a lookup, and we can avoid cluttering the icache/dcache",
> then we probably don't want to.
>
Yes, I'd rather fix this there if possible, but we first need to know the idea
behind the curious choice not to instantiate the dentry. Otherwise, we can go
the ->atomic_open route as you suggested.
>
> At this stage I think I would lean towards O_CREAT|O_DIRECTORY not
> working on these two filesystems. Neither support a .create
> inode_operation, so providing a .atomic_open would be quite easy: just
Wait, so an O_CREAT opens already return an error on kernf/tracefs? They
only support directories? I have to read up on these a bit over the weekend.
> do a ->lookup and pass the result to finish_no_open, but return an
> appropriate error if O_CREAT was requested but no inode was found. That
> would then treat these like other atomic_open filesystems in that
> O_CREAT|O_DIRECTORY wouldn't work until the fs maintainer accepted a
> patch for it.
>
This sounds good to me.
Best,
Jori.
PS. @Neil, will you be at Plumbers by chance?
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno
2026-09-18 8:04 ` Christian Brauner
@ 2026-09-25 22:55 ` Jori Koolstra
0 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-25 22:55 UTC (permalink / raw)
To: Christian Brauner
Cc: Jeff Layton, Al Viro, Aleksa Sarai, NeilBrown, Amir Goldstein,
Jan Kara, linux-fsdevel, linux-kernel
> Op 18-09-2026 04:04 EDT schreef Christian Brauner <brauner@kernel.org>:
>
>
> On Sun, Sep 13, 2026 at 08:50:12PM +0200, Jori Koolstra wrote:
> > Currently filesystems return -EACCES for missing ->create and -EPERM for
> > missing ->mkdir in vfs_create/vfs_mkdir. Instead of duplicating these
> > dubious and inconsistent error codes to lookup_open(), change all to
> > -EOPNOTSUPP.
> >
> > Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
> > ---
>
> I suspect that still has regression potential but we can try.
I know, but the one that showed up in next wasn't related to this. We can
drop it if there's issues, but it would be really really ugly. That's why
I suggest we try this lighter errno correction.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-18 8:17 ` Christian Brauner
2026-09-18 10:07 ` NeilBrown
@ 2026-09-25 23:13 ` Jori Koolstra
1 sibling, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-09-25 23:13 UTC (permalink / raw)
To: Christian Brauner, NeilBrown
Cc: Jeff Layton, Al Viro, Aleksa Sarai, Amir Goldstein, Jan Kara,
linux-fsdevel, linux-kernel
> Op 18-09-2026 04:17 EDT schreef Christian Brauner <brauner@kernel.org>:
>
>
> > @@ -4582,12 +4646,25 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
> > goto out_dput;
> > }
> >
> > - if (!dir_inode->i_op->create) {
> > + /* mimic operation missing errnos of vfs_mkdir/vfs_create */
> > + if (create_dir && !dir_inode->i_op->mkdir) {
> > + error = -EPERM;
> > + goto out_dput;
> > + }
> > + if (!create_dir && !dir_inode->i_op->create) {
> > error = -EACCES;
> > goto out_dput;
> > }
> >
> > - error = vfs_create_no_perm(idmap, dentry, mode, &delegated_inode);
> > + if (create_dir) {
> > + struct dentry *res = vfs_mkdir_no_perm(idmap, dir_inode, dentry, mode,
> > + &delegated_inode);
>
> So, I think this is broken. Whatever vfs_mkdir_no_perm() returns is
> passed to do_open(). Kernfs makes that buggy.
>
> cgroup, cgroup2, and resctrl are all implemented on top of kernfs. And
> kernfs ->mkdir:: iop never instantiates the dentry.
>
> So that means e.g.,
>
> openat(cgroup_dir, "subdir", O_CREAT|O_DIRECTORY) creates a cgroup
> and then fails with ENOTDIR.
>
> So the negative dentry gets handed out and now userspace holds an fd
> with that negative dentry. So say userspace does fchown() to 1000 and
> then fchmod() with the sticky bit and then you get a NULL deref. I have
> reproduced this.
>
Damn, nice catch. None of the LLMs caught this, only our flesh and blood
maintainer ;) I must admit I didn't consider a ->mkdir that left the dentry
negative.
I haven't looked into kernfs yet, but does it make sense you can set the
sticky bit there? Not that it is a viable solution to just mask that out,
as anyone would expect positive dentry at the time of do_open(). Curiously,
with a very quick glance, it does seem that only may_create_in_sticky() relies
on the inode being available.
Thanks for looking into this, Christian. See you at Plumbers, I guess?
Best wishes,
Jori.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-19 0:01 ` NeilBrown
2026-09-25 13:57 ` Christian Brauner
@ 2026-09-29 11:54 ` Jori Koolstra
2026-09-29 22:15 ` NeilBrown
1 sibling, 1 reply; 40+ messages in thread
From: Jori Koolstra @ 2026-09-29 11:54 UTC (permalink / raw)
To: NeilBrown, NeilBrown, Christian Brauner
Cc: Jeff Layton, Al Viro, Aleksa Sarai, Amir Goldstein, Jan Kara,
linux-fsdevel, linux-kernel
> Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@ownmail.net>:
>
> >
> > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > dentry, but it doesn't any more. The reason was because
> > d_splice_alias() on might return a different dentry.
> > In this case we want the same dentry, but we need to do a lookup on it.
> >
> > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> >
> > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > I'm not certain that is needed if we keep the parent locked, but we
> > would need to be certain.
> > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > cannot handle hashed dentries and a hashed-negative dentry is passed
> > to ->mkdir.
> >
> > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > probably not.
> >
> > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > dentry is negative, we need to d_drop() and call ->lookup with a big
> > comment about kernfs. But we need to double-check that this will do the
> > right thing with ->d_time (I think it will).
> > We also need to think carefully about races with
> > kernfs_dop_revalidate(), which could happen concurrently with the
> > ->lookup.
>
> I've thought a bit more about this ... I think that doing a lookup after
> the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> about the fs that I would rather not assume.
>
> I would rather have the current proposed code check for a negative
> dentry, and fail with -EIO or similar.
>
I just noticed that there's precedent for this in overlayfs in super.c:
/* Weird filesystem returning with hashed negative (kernfs)? */
err = -EINVAL;
if (d_really_is_negative(work))
goto out_dput;
Shall we just do this for current kernel release, then we can add support
later if wanted.
(But let's do EOPNOTSUPP instead of EINVAL)
What do you think?
Thanks,
Jori.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-29 11:54 ` Jori Koolstra
@ 2026-09-29 22:15 ` NeilBrown
2026-09-30 9:45 ` Amir Goldstein
2026-09-30 22:33 ` Jori Koolstra
0 siblings, 2 replies; 40+ messages in thread
From: NeilBrown @ 2026-09-29 22:15 UTC (permalink / raw)
To: Jori Koolstra
Cc: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@ownmail.net>:
> >
> > >
> > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > dentry, but it doesn't any more. The reason was because
> > > d_splice_alias() on might return a different dentry.
> > > In this case we want the same dentry, but we need to do a lookup on it.
> > >
> > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > >
> > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > I'm not certain that is needed if we keep the parent locked, but we
> > > would need to be certain.
> > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > to ->mkdir.
> > >
> > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > probably not.
> > >
> > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > comment about kernfs. But we need to double-check that this will do the
> > > right thing with ->d_time (I think it will).
> > > We also need to think carefully about races with
> > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > ->lookup.
> >
> > I've thought a bit more about this ... I think that doing a lookup after
> > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > about the fs that I would rather not assume.
> >
> > I would rather have the current proposed code check for a negative
> > dentry, and fail with -EIO or similar.
> >
>
> I just noticed that there's precedent for this in overlayfs in super.c:
>
> /* Weird filesystem returning with hashed negative (kernfs)? */
> err = -EINVAL;
> if (d_really_is_negative(work))
> goto out_dput;
>
> Shall we just do this for current kernel release, then we can add support
> later if wanted.
>
> (But let's do EOPNOTSUPP instead of EINVAL)
>
> What do you think?
The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
might create the directory, then return -EOPNOTSUPP. This is weird and
I'd rather it not be visible.
Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
remain a -EINVAL on any filesystem which doesn't completely support
the functionality.
To do that we need some way to detect kernfs and tracefs. I think
the only way we can do that is to make some change to those two
filesystems.
Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
NeilBrown
>
> Thanks,
> Jori.
>
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-29 22:15 ` NeilBrown
@ 2026-09-30 9:45 ` Amir Goldstein
2026-09-30 22:16 ` Jori Koolstra
2026-09-30 22:33 ` Jori Koolstra
1 sibling, 1 reply; 40+ messages in thread
From: Amir Goldstein @ 2026-09-30 9:45 UTC (permalink / raw)
To: NeilBrown, Jori Koolstra
Cc: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, Jan Kara,
linux-fsdevel, linux-kernel, Theodore Tso
On Wed, Sep 30, 2026 at 12:15 AM NeilBrown <neilb@ownmail.net> wrote:
>
> On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@ownmail.net>:
> > >
> > > >
> > > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > > dentry, but it doesn't any more. The reason was because
> > > > d_splice_alias() on might return a different dentry.
> > > > In this case we want the same dentry, but we need to do a lookup on it.
> > > >
> > > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > > >
> > > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > > I'm not certain that is needed if we keep the parent locked, but we
> > > > would need to be certain.
> > > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > > to ->mkdir.
> > > >
> > > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > > probably not.
> > > >
> > > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > > comment about kernfs. But we need to double-check that this will do the
> > > > right thing with ->d_time (I think it will).
> > > > We also need to think carefully about races with
> > > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > > ->lookup.
> > >
> > > I've thought a bit more about this ... I think that doing a lookup after
> > > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > > about the fs that I would rather not assume.
> > >
> > > I would rather have the current proposed code check for a negative
> > > dentry, and fail with -EIO or similar.
> > >
> >
> > I just noticed that there's precedent for this in overlayfs in super.c:
> >
> > /* Weird filesystem returning with hashed negative (kernfs)? */
> > err = -EINVAL;
> > if (d_really_is_negative(work))
> > goto out_dput;
> >
> > Shall we just do this for current kernel release, then we can add support
> > later if wanted.
> >
> > (But let's do EOPNOTSUPP instead of EINVAL)
> >
> > What do you think?
>
> The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> might create the directory, then return -EOPNOTSUPP. This is weird and
> I'd rather it not be visible.
>
> Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> remain a -EINVAL on any filesystem which doesn't completely support
> the functionality.
>
Joining late to this party so apologies in advance if my questions
have already been addressed.
I agree with Neil's statement above, but IMO, the atomic_open() fs
match the description of "doesn't completely support the functionality."
Therefore, I think that rather than success if directory exists, they
should also return -EINVAL/-EOPNOTSUPP consistently (see below).
> To do that we need some way to detect kernfs and tracefs. I think
> the only way we can do that is to make some change to those two
> filesystems.
> Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
>
Maybe, but then I think we should also exclude all the atomic_open fs.
Food for thought:
If you agree that tracefs/kernfs and <network>fs should have consistent
"API not supported" behavior, then gating the flag combination on
sb->s_d_flags & DCACHE_OP_REVALIDATE covers all the fs that we
want to exclude, plus some fs that we shouldn't really care about.
My agent tells me that the list is:
afs, coda, ecryptfs, exfat, hfs, jfs, ocfs2, orangefs, proc, ubifs, vfat
Two exceptions that we MAY care about:
1. overlayfs DCACHE_OP_REVALIDATE is derived from underlying
layers per dentry (If any of them have the flag) but always has it
in sb default_d_op, but it could technically derive sb->s_d_flags
layers at sb fill time, without any behavior change
2. encrypted/casefolded directories with fscrypt_d_revalidate (ext4/f2fs)
with CONFIG_FS_ENCRYPTION=y
ecrypted/casefold is the non common configuration, technically
sb->s_d_flags could be set according to sb features at sb fill time
without any behavior change
Thanks,
Amir.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-30 9:45 ` Amir Goldstein
@ 2026-09-30 22:16 ` Jori Koolstra
2026-10-01 9:29 ` Amir Goldstein
0 siblings, 1 reply; 40+ messages in thread
From: Jori Koolstra @ 2026-09-30 22:16 UTC (permalink / raw)
To: Amir Goldstein, NeilBrown
Cc: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai, Jan Kara,
linux-fsdevel, linux-kernel, Theodore Tso
> Op 30-09-2026 05:45 EDT schreef Amir Goldstein <amir73il@gmail.com>:
>
>
> On Wed, Sep 30, 2026 at 12:15 AM NeilBrown <neilb@ownmail.net> wrote:
> >
> > On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > > > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@ownmail.net>:
> > > >
> > > > >
> > > > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > > > dentry, but it doesn't any more. The reason was because
> > > > > d_splice_alias() on might return a different dentry.
> > > > > In this case we want the same dentry, but we need to do a lookup on it.
> > > > >
> > > > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > > > >
> > > > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > > > I'm not certain that is needed if we keep the parent locked, but we
> > > > > would need to be certain.
> > > > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > > > to ->mkdir.
> > > > >
> > > > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > > > probably not.
> > > > >
> > > > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > > > comment about kernfs. But we need to double-check that this will do the
> > > > > right thing with ->d_time (I think it will).
> > > > > We also need to think carefully about races with
> > > > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > > > ->lookup.
> > > >
> > > > I've thought a bit more about this ... I think that doing a lookup after
> > > > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > > > about the fs that I would rather not assume.
> > > >
> > > > I would rather have the current proposed code check for a negative
> > > > dentry, and fail with -EIO or similar.
> > > >
> > >
> > > I just noticed that there's precedent for this in overlayfs in super.c:
> > >
> > > /* Weird filesystem returning with hashed negative (kernfs)? */
> > > err = -EINVAL;
> > > if (d_really_is_negative(work))
> > > goto out_dput;
> > >
> > > Shall we just do this for current kernel release, then we can add support
> > > later if wanted.
> > >
> > > (But let's do EOPNOTSUPP instead of EINVAL)
> > >
> > > What do you think?
> >
> > The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> > might create the directory, then return -EOPNOTSUPP. This is weird and
> > I'd rather it not be visible.
> >
> > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > remain a -EINVAL on any filesystem which doesn't completely support
> > the functionality.
> >
>
> Joining late to this party so apologies in advance if my questions
> have already been addressed.
>
> I agree with Neil's statement above, but IMO, the atomic_open() fs
> match the description of "doesn't completely support the functionality."
> Therefore, I think that rather than success if directory exists, they
> should also return -EINVAL/-EOPNOTSUPP consistently (see below).
>
The issue with this is that if you want per fs atomic_open() opt-in (in
contrast to either implementing all instances in one release or disabling
all), you have the issue that your lookup now depends on the caching status
of the directory dentry.
If that dentry is in cache, and positive, d_lookup() earlier in lookup_open()
makes it return early:
if (dentry->d_inode) {
/* Cached positive dentry: will open in do_open(). */
goto out;
}
So you get your lookup. But if the same dentry is not in cache, now you
suddenly get -EINVAL. I thought that behavior was more unwanted then
what I eventually settled on, namely to strip the O_CREAT bit.
Best,
Jori.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-29 22:15 ` NeilBrown
2026-09-30 9:45 ` Amir Goldstein
@ 2026-09-30 22:33 ` Jori Koolstra
2026-09-30 22:56 ` NeilBrown
1 sibling, 1 reply; 40+ messages in thread
From: Jori Koolstra @ 2026-09-30 22:33 UTC (permalink / raw)
To: NeilBrown, NeilBrown
Cc: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
> Op 29-09-2026 18:15 EDT schreef NeilBrown <neilb@ownmail.net>:
>
>
> On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@ownmail.net>:
> > >
> > > >
> > > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > > dentry, but it doesn't any more. The reason was because
> > > > d_splice_alias() on might return a different dentry.
> > > > In this case we want the same dentry, but we need to do a lookup on it.
> > > >
> > > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > > >
> > > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > > I'm not certain that is needed if we keep the parent locked, but we
> > > > would need to be certain.
> > > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > > to ->mkdir.
> > > >
> > > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > > probably not.
> > > >
> > > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > > comment about kernfs. But we need to double-check that this will do the
> > > > right thing with ->d_time (I think it will).
> > > > We also need to think carefully about races with
> > > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > > ->lookup.
> > >
> > > I've thought a bit more about this ... I think that doing a lookup after
> > > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > > about the fs that I would rather not assume.
> > >
> > > I would rather have the current proposed code check for a negative
> > > dentry, and fail with -EIO or similar.
> > >
> >
> > I just noticed that there's precedent for this in overlayfs in super.c:
> >
> > /* Weird filesystem returning with hashed negative (kernfs)? */
> > err = -EINVAL;
> > if (d_really_is_negative(work))
> > goto out_dput;
> >
> > Shall we just do this for current kernel release, then we can add support
> > later if wanted.
> >
> > (But let's do EOPNOTSUPP instead of EINVAL)
> >
> > What do you think?
>
> The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> might create the directory, then return -EOPNOTSUPP. This is weird and
> I'd rather it not be visible.
>
Err, *derp*, what a stupid suggestion of mine.
> Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> remain a -EINVAL on any filesystem which doesn't completely support
> the functionality.
>
I don't think that works for the reason I just wrote in my email to Amir:
it would make lookup dependent on the dentry cache. If it's in-cache, you get
your dir, otherwise suddenly -EINVAL.
> To do that we need some way to detect kernfs and tracefs. I think
> the only way we can do that is to make some change to those two
> filesystems.
> Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
>
We can just implement atomic_open() for kernfs/tracefs, do a lookup there,
and if negative with O_CREAT return maybe -ENOENT (or really we need a new
error that says "the requested create could not be serviced," like -ENOCREATE,
or whatever). And if it is positive we do finish_no_open().
It's a bit of a hack because it does not really have anything to do with
atomicity, but it does short-circuit the mkdir call in lookup_open(). I guess
that would work. Maybe I am confused, but wasn't that what you proposed here
earlier?
Best,
Jori.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-30 22:33 ` Jori Koolstra
@ 2026-09-30 22:56 ` NeilBrown
2026-09-30 23:25 ` Jori Koolstra
0 siblings, 1 reply; 40+ messages in thread
From: NeilBrown @ 2026-09-30 22:56 UTC (permalink / raw)
To: Jori Koolstra
Cc: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
On Thu, 01 Oct 2026, Jori Koolstra wrote:
> > Op 29-09-2026 18:15 EDT schreef NeilBrown <neilb@ownmail.net>:
> >
> >
> > On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > > > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@ownmail.net>:
> > > >
> > > > >
> > > > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > > > dentry, but it doesn't any more. The reason was because
> > > > > d_splice_alias() on might return a different dentry.
> > > > > In this case we want the same dentry, but we need to do a lookup on it.
> > > > >
> > > > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > > > >
> > > > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > > > I'm not certain that is needed if we keep the parent locked, but we
> > > > > would need to be certain.
> > > > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > > > to ->mkdir.
> > > > >
> > > > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > > > probably not.
> > > > >
> > > > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > > > comment about kernfs. But we need to double-check that this will do the
> > > > > right thing with ->d_time (I think it will).
> > > > > We also need to think carefully about races with
> > > > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > > > ->lookup.
> > > >
> > > > I've thought a bit more about this ... I think that doing a lookup after
> > > > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > > > about the fs that I would rather not assume.
> > > >
> > > > I would rather have the current proposed code check for a negative
> > > > dentry, and fail with -EIO or similar.
> > > >
> > >
> > > I just noticed that there's precedent for this in overlayfs in super.c:
> > >
> > > /* Weird filesystem returning with hashed negative (kernfs)? */
> > > err = -EINVAL;
> > > if (d_really_is_negative(work))
> > > goto out_dput;
> > >
> > > Shall we just do this for current kernel release, then we can add support
> > > later if wanted.
> > >
> > > (But let's do EOPNOTSUPP instead of EINVAL)
> > >
> > > What do you think?
> >
> > The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> > might create the directory, then return -EOPNOTSUPP. This is weird and
> > I'd rather it not be visible.
> >
>
> Err, *derp*, what a stupid suggestion of mine.
>
> > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > remain a -EINVAL on any filesystem which doesn't completely support
> > the functionality.
> >
>
> I don't think that works for the reason I just wrote in my email to Amir:
> it would make lookup dependent on the dentry cache. If it's in-cache, you get
> your dir, otherwise suddenly -EINVAL.
>
> > To do that we need some way to detect kernfs and tracefs. I think
> > the only way we can do that is to make some change to those two
> > filesystems.
> > Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
> >
>
> We can just implement atomic_open() for kernfs/tracefs, do a lookup there,
> and if negative with O_CREAT return maybe -ENOENT (or really we need a new
> error that says "the requested create could not be serviced," like -ENOCREATE,
> or whatever). And if it is positive we do finish_no_open().
>
> It's a bit of a hack because it does not really have anything to do with
> atomicity, but it does short-circuit the mkdir call in lookup_open(). I guess
> that would work. Maybe I am confused, but wasn't that what you proposed here
> earlier?
Yes, it is what I proposed earlier. But I think it would require more
review and probably make it unrealistic to land this cycle. But I'm not
thinking it is unlikely to be ready this cycle any way.
I'm now wondering if we should keep ->atomic_open out of the loop and
always use ->mkdir to create a directory.
Based on your justification you probably always want O_EXCL and I would
be inclined to require that.
So if the dentry is in-lookup we call ->atomic_open(O_DIRECTORY). If
that succeeds - good. If it reports ENOENT or a negative dentry, then
we cal ->mkdir. If that succeeds with a positive dentry, we call
through to call ->open.
If ->mkdir succeeds with a negative dentry - we have the problem of
kernfs and tracefs. I'm leaning towards fixing those to do the lookup.
I don't think any filesystems *can* combine mkdir with open, so not
using ->atomic_open for the mkdir doesn't actually lose anything.
NeilBrown
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-30 22:56 ` NeilBrown
@ 2026-09-30 23:25 ` Jori Koolstra
2026-10-01 1:00 ` NeilBrown
0 siblings, 1 reply; 40+ messages in thread
From: Jori Koolstra @ 2026-09-30 23:25 UTC (permalink / raw)
To: NeilBrown, NeilBrown
Cc: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
> Op 30-09-2026 18:56 EDT schreef NeilBrown <neilb@ownmail.net>:
>
>
> On Thu, 01 Oct 2026, Jori Koolstra wrote:
> >
> > Err, *derp*, what a stupid suggestion of mine.
> >
> > > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > > remain a -EINVAL on any filesystem which doesn't completely support
> > > the functionality.
> > >
> >
> > I don't think that works for the reason I just wrote in my email to Amir:
> > it would make lookup dependent on the dentry cache. If it's in-cache, you get
> > your dir, otherwise suddenly -EINVAL.
> >
> > > To do that we need some way to detect kernfs and tracefs. I think
> > > the only way we can do that is to make some change to those two
> > > filesystems.
> > > Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
> > >
> >
> > We can just implement atomic_open() for kernfs/tracefs, do a lookup there,
> > and if negative with O_CREAT return maybe -ENOENT (or really we need a new
> > error that says "the requested create could not be serviced," like -ENOCREATE,
> > or whatever). And if it is positive we do finish_no_open().
> >
> > It's a bit of a hack because it does not really have anything to do with
> > atomicity, but it does short-circuit the mkdir call in lookup_open(). I guess
> > that would work. Maybe I am confused, but wasn't that what you proposed here
> > earlier?
>
> Yes, it is what I proposed earlier. But I think it would require more
> review and probably make it unrealistic to land this cycle. But I'm not
> thinking it is unlikely to be ready this cycle any way.
>
Not unlikely, so likely that is :)
> I'm now wondering if we should keep ->atomic_open out of the loop and
> always use ->mkdir to create a directory.
> Based on your justification you probably always want O_EXCL and I would
> be inclined to require that.
>
I wanted it to work as regular O_CREAT, so no forced O_EXCL per se.
> So if the dentry is in-lookup we call ->atomic_open(O_DIRECTORY). If
> that succeeds - good. If it reports ENOENT or a negative dentry, then
Can ->atomic_open() return a negative dentry?
> we cal ->mkdir. If that succeeds with a positive dentry, we call
> through to call ->open.
> If ->mkdir succeeds with a negative dentry - we have the problem of
> kernfs and tracefs. I'm leaning towards fixing those to do the lookup.
Yes, I follow this. But we do need to ask their maintainers then why that
was chosen. Plus, we need to ban negative dentry returns also for future
fses, which may be limiting.
>
> I don't think any filesystems *can* combine mkdir with open, so not
> using ->atomic_open for the mkdir doesn't actually lose anything.
>
Yeah, I don't quite understand the specifics of this as I am not familiar
with NFS and the likes. I presume because of things like network traffic
->atomic_open() was needed as a single call into the underlying fs. But
I have no idea why or why not that would make sense for directories. If
we do it the way you suggest now, we can get rid of all the O_CREAT stripping
> NeilBrown
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-30 23:25 ` Jori Koolstra
@ 2026-10-01 1:00 ` NeilBrown
2026-10-01 14:07 ` Jori Koolstra
0 siblings, 1 reply; 40+ messages in thread
From: NeilBrown @ 2026-10-01 1:00 UTC (permalink / raw)
To: Jori Koolstra
Cc: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
On Thu, 01 Oct 2026, Jori Koolstra wrote:
> > Op 30-09-2026 18:56 EDT schreef NeilBrown <neilb@ownmail.net>:
> >
> >
> > On Thu, 01 Oct 2026, Jori Koolstra wrote:
> > >
> > > Err, *derp*, what a stupid suggestion of mine.
> > >
> > > > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > > > remain a -EINVAL on any filesystem which doesn't completely support
> > > > the functionality.
> > > >
> > >
> > > I don't think that works for the reason I just wrote in my email to Amir:
> > > it would make lookup dependent on the dentry cache. If it's in-cache, you get
> > > your dir, otherwise suddenly -EINVAL.
> > >
> > > > To do that we need some way to detect kernfs and tracefs. I think
> > > > the only way we can do that is to make some change to those two
> > > > filesystems.
> > > > Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
> > > >
> > >
> > > We can just implement atomic_open() for kernfs/tracefs, do a lookup there,
> > > and if negative with O_CREAT return maybe -ENOENT (or really we need a new
> > > error that says "the requested create could not be serviced," like -ENOCREATE,
> > > or whatever). And if it is positive we do finish_no_open().
> > >
> > > It's a bit of a hack because it does not really have anything to do with
> > > atomicity, but it does short-circuit the mkdir call in lookup_open(). I guess
> > > that would work. Maybe I am confused, but wasn't that what you proposed here
> > > earlier?
> >
> > Yes, it is what I proposed earlier. But I think it would require more
> > review and probably make it unrealistic to land this cycle. But I'm not
> > thinking it is unlikely to be ready this cycle any way.
> >
>
> Not unlikely, so likely that is :)
>
> > I'm now wondering if we should keep ->atomic_open out of the loop and
> > always use ->mkdir to create a directory.
> > Based on your justification you probably always want O_EXCL and I would
> > be inclined to require that.
> >
>
> I wanted it to work as regular O_CREAT, so no forced O_EXCL per se.
Why? The whole point is atomicity, and without O_EXCL there is no
atomicity.
I guess it doesn't hurt to not require O_EXCL, but it is just another
combination to test...
>
> > So if the dentry is in-lookup we call ->atomic_open(O_DIRECTORY). If
> > that succeeds - good. If it reports ENOENT or a negative dentry, then
>
> Can ->atomic_open() return a negative dentry?
No reason why not. It is mapped to -ENOENT
if (unlikely(d_is_negative(dentry)))
error = -ENOENT;
I'm in favour of ->atomic_open() calling finish_no_open() is all
no-error cases where it didn't do something about opening the file.
But I don't object to an explicit -ENOENT return.
>
> > we cal ->mkdir. If that succeeds with a positive dentry, we call
> > through to call ->open.
> > If ->mkdir succeeds with a negative dentry - we have the problem of
> > kernfs and tracefs. I'm leaning towards fixing those to do the lookup.
>
> Yes, I follow this. But we do need to ask their maintainers then why that
> was chosen. Plus, we need to ban negative dentry returns also for future
> fses, which may be limiting.
Odds a good that there wasn' a specific reason why leaving the dentry
negative was chosen. I think the better question is "do you have a
problem with this patch" where the patch does a lookup.
Putting a WARN_ON(d_really_is_negative()) somewhere in vfs_mkdir would be
easy enough. I cannot imagine it really being a burden for any fs. If
that does happen were can solve the problem then. Being prepared of all
possible eventuality is a no-win proposition. I've tried - I gave up.
>
> >
> > I don't think any filesystems *can* combine mkdir with open, so not
> > using ->atomic_open for the mkdir doesn't actually lose anything.
> >
>
> Yeah, I don't quite understand the specifics of this as I am not familiar
> with NFS and the likes. I presume because of things like network traffic
> ->atomic_open() was needed as a single call into the underlying fs. But
> I have no idea why or why not that would make sense for directories. If
> we do it the way you suggest now, we can get rid of all the O_CREAT stripping
->atomic_open was created specifically for NFSv4 which has a combined
lookup/create/open request. You can do everything with a combination of
->lookup and (exclusive) ->create (providing create returns a
non-negative dentry) and ->truncate and open. Using ->atomic_open means
fewer round-trips to the server so less latency.
NFS doesn't have a concept of "open" for directories. CIFS probably
does. FUSE probably doesn't but could possibly add one.
But directories are not opened nearly as often as files, so combining
things isn't so important. And O_TRUNC is not meaningful for
directories so that is one fewer thing that could be combined.
NeilBrown
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-09-30 22:16 ` Jori Koolstra
@ 2026-10-01 9:29 ` Amir Goldstein
2026-10-01 10:08 ` NeilBrown
2026-10-01 15:42 ` Jori Koolstra
0 siblings, 2 replies; 40+ messages in thread
From: Amir Goldstein @ 2026-10-01 9:29 UTC (permalink / raw)
To: Jori Koolstra
Cc: NeilBrown, Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Jan Kara, linux-fsdevel, linux-kernel, Theodore Tso
On Thu, Oct 1, 2026 at 12:16 AM Jori Koolstra <jkoolstra@xs4all.nl> wrote:
>
>
> > Op 30-09-2026 05:45 EDT schreef Amir Goldstein <amir73il@gmail.com>:
> >
> >
> > On Wed, Sep 30, 2026 at 12:15 AM NeilBrown <neilb@ownmail.net> wrote:
> > >
...
> > > The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> > > might create the directory, then return -EOPNOTSUPP. This is weird and
> > > I'd rather it not be visible.
> > >
> > > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > > remain a -EINVAL on any filesystem which doesn't completely support
> > > the functionality.
> > >
> >
> > Joining late to this party so apologies in advance if my questions
> > have already been addressed.
> >
> > I agree with Neil's statement above, but IMO, the atomic_open() fs
> > match the description of "doesn't completely support the functionality."
> > Therefore, I think that rather than success if directory exists, they
> > should also return -EINVAL/-EOPNOTSUPP consistently (see below).
> >
>
> The issue with this is that if you want per fs atomic_open() opt-in (in
> contrast to either implementing all instances in one release or disabling
> all), you have the issue that your lookup now depends on the caching status
> of the directory dentry.
>
> If that dentry is in cache, and positive, d_lookup() earlier in lookup_open()
> makes it return early:
>
> if (dentry->d_inode) {
> /* Cached positive dentry: will open in do_open(). */
> goto out;
> }
>
> So you get your lookup. But if the same dentry is not in cache, now you
> suddenly get -EINVAL. I thought that behavior was more unwanted then
> what I eventually settled on, namely to strip the O_CREAT bit.
>
Maybe I am missing something, but I think you misunderstand me.
What I mean is - if directory inode has a ->atomic_open() op,
bail early with -EINVAL/-EOPNOTSUPP, because this is a network
filesystem that does not support atomic O_CREATE|O_DIRECTORY
and in most likelihood never will support it.
This gating criteria is not dependent on cache state,
which is what we wanted.
The justification of using ->d_revalidate() as another opt-out
is that existence of ->d_revalidate() means that the state known
to dcache is only semi-reliable, so making atomic create/open
promises is problematic (O_EXCL for example).
The problem is that some fs (overlayfs/ext4/f2fs) register
a mostly-noop ->d_revalidate() so I proposed how to deal with those.
Thanks,
Amir.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-10-01 9:29 ` Amir Goldstein
@ 2026-10-01 10:08 ` NeilBrown
2026-10-01 11:29 ` Amir Goldstein
2026-10-01 15:42 ` Jori Koolstra
1 sibling, 1 reply; 40+ messages in thread
From: NeilBrown @ 2026-10-01 10:08 UTC (permalink / raw)
To: Amir Goldstein
Cc: Jori Koolstra, Christian Brauner, Jeff Layton, Al Viro,
Aleksa Sarai, Jan Kara, linux-fsdevel, linux-kernel,
Theodore Tso
On Thu, 01 Oct 2026, Amir Goldstein wrote:
> On Thu, Oct 1, 2026 at 12:16 AM Jori Koolstra <jkoolstra@xs4all.nl> wrote:
> >
> >
> > > Op 30-09-2026 05:45 EDT schreef Amir Goldstein <amir73il@gmail.com>:
> > >
> > >
> > > On Wed, Sep 30, 2026 at 12:15 AM NeilBrown <neilb@ownmail.net> wrote:
> > > >
> ...
> > > > The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> > > > might create the directory, then return -EOPNOTSUPP. This is weird and
> > > > I'd rather it not be visible.
> > > >
> > > > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > > > remain a -EINVAL on any filesystem which doesn't completely support
> > > > the functionality.
> > > >
> > >
> > > Joining late to this party so apologies in advance if my questions
> > > have already been addressed.
> > >
> > > I agree with Neil's statement above, but IMO, the atomic_open() fs
> > > match the description of "doesn't completely support the functionality."
> > > Therefore, I think that rather than success if directory exists, they
> > > should also return -EINVAL/-EOPNOTSUPP consistently (see below).
> > >
> >
> > The issue with this is that if you want per fs atomic_open() opt-in (in
> > contrast to either implementing all instances in one release or disabling
> > all), you have the issue that your lookup now depends on the caching status
> > of the directory dentry.
> >
> > If that dentry is in cache, and positive, d_lookup() earlier in lookup_open()
> > makes it return early:
> >
> > if (dentry->d_inode) {
> > /* Cached positive dentry: will open in do_open(). */
> > goto out;
> > }
> >
> > So you get your lookup. But if the same dentry is not in cache, now you
> > suddenly get -EINVAL. I thought that behavior was more unwanted then
> > what I eventually settled on, namely to strip the O_CREAT bit.
> >
>
> Maybe I am missing something, but I think you misunderstand me.
> What I mean is - if directory inode has a ->atomic_open() op,
> bail early with -EINVAL/-EOPNOTSUPP, because this is a network
> filesystem that does not support atomic O_CREATE|O_DIRECTORY
> and in most likelihood never will support it.
NFS can certainly support O_CREATE|O_DIRECTORY. The MKDIR request
creates a directory and returns the file-handle of the directory that
was created. I think other network filesystems return the identity of
the created thing - or fail if it already existed. That is enough for
full support.
The only cases where I think I think there is any doubt of support is
kernfs and tracefs because they don't return the inode. tracefs is
interesting because it drops and retakes the parent lock so it isn't
immediately clear what atomicity is available, though it doesn't support
rename at all so maybe there is no interesting race.
>
> This gating criteria is not dependent on cache state,
> which is what we wanted.
>
> The justification of using ->d_revalidate() as another opt-out
> is that existence of ->d_revalidate() means that the state known
> to dcache is only semi-reliable, so making atomic create/open
> promises is problematic (O_EXCL for example).
The idea of using ->d_revalidate as a gate is certainly interesting.
Apart from the interaction with case-insensitivity, if we initially only
supported filesystems that don't have ->d_revalidate, I think we would get
coverage for enough filesystems to be interesting. We could then take a
bit more time to think through the rest of the picture.
I don't think O_EXCL is at all problematic. ->mkdir() is already
required to return -EEXIST if the directory already exists.
>
> The problem is that some fs (overlayfs/ext4/f2fs) register
> a mostly-noop ->d_revalidate() so I proposed how to deal with those.
Presumably the fs would provide two dentry_operations structures and
choose which to pass to set_default_d_op() when creating the superblock.
Doing that would even provide slightly better performance in the
case where no d_revalidate is needed.
Thanks,
NeilBrown
>
> Thanks,
> Amir.
>
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-10-01 10:08 ` NeilBrown
@ 2026-10-01 11:29 ` Amir Goldstein
0 siblings, 0 replies; 40+ messages in thread
From: Amir Goldstein @ 2026-10-01 11:29 UTC (permalink / raw)
To: NeilBrown
Cc: Jori Koolstra, Christian Brauner, Jeff Layton, Al Viro,
Aleksa Sarai, Jan Kara, linux-fsdevel, linux-kernel,
Theodore Tso
On Thu, Oct 1, 2026 at 12:08 PM NeilBrown <neilb@ownmail.net> wrote:
>
> On Thu, 01 Oct 2026, Amir Goldstein wrote:
> > On Thu, Oct 1, 2026 at 12:16 AM Jori Koolstra <jkoolstra@xs4all.nl> wrote:
> > >
> > >
> > > > Op 30-09-2026 05:45 EDT schreef Amir Goldstein <amir73il@gmail.com>:
> > > >
> > > >
> > > > On Wed, Sep 30, 2026 at 12:15 AM NeilBrown <neilb@ownmail.net> wrote:
> > > > >
> > ...
> > > > > The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> > > > > might create the directory, then return -EOPNOTSUPP. This is weird and
> > > > > I'd rather it not be visible.
> > > > >
> > > > > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > > > > remain a -EINVAL on any filesystem which doesn't completely support
> > > > > the functionality.
> > > > >
> > > >
> > > > Joining late to this party so apologies in advance if my questions
> > > > have already been addressed.
> > > >
> > > > I agree with Neil's statement above, but IMO, the atomic_open() fs
> > > > match the description of "doesn't completely support the functionality."
> > > > Therefore, I think that rather than success if directory exists, they
> > > > should also return -EINVAL/-EOPNOTSUPP consistently (see below).
> > > >
> > >
> > > The issue with this is that if you want per fs atomic_open() opt-in (in
> > > contrast to either implementing all instances in one release or disabling
> > > all), you have the issue that your lookup now depends on the caching status
> > > of the directory dentry.
> > >
> > > If that dentry is in cache, and positive, d_lookup() earlier in lookup_open()
> > > makes it return early:
> > >
> > > if (dentry->d_inode) {
> > > /* Cached positive dentry: will open in do_open(). */
> > > goto out;
> > > }
> > >
> > > So you get your lookup. But if the same dentry is not in cache, now you
> > > suddenly get -EINVAL. I thought that behavior was more unwanted then
> > > what I eventually settled on, namely to strip the O_CREAT bit.
> > >
> >
> > Maybe I am missing something, but I think you misunderstand me.
> > What I mean is - if directory inode has a ->atomic_open() op,
> > bail early with -EINVAL/-EOPNOTSUPP, because this is a network
> > filesystem that does not support atomic O_CREATE|O_DIRECTORY
> > and in most likelihood never will support it.
>
> NFS can certainly support O_CREATE|O_DIRECTORY. The MKDIR request
> creates a directory and returns the file-handle of the directory that
> was created. I think other network filesystems return the identity of
> the created thing - or fail if it already existed. That is enough for
> full support.
>
> The only cases where I think I think there is any doubt of support is
> kernfs and tracefs because they don't return the inode. tracefs is
> interesting because it drops and retakes the parent lock so it isn't
> immediately clear what atomicity is available, though it doesn't support
> rename at all so maybe there is no interesting race.
>
> >
> > This gating criteria is not dependent on cache state,
> > which is what we wanted.
> >
> > The justification of using ->d_revalidate() as another opt-out
> > is that existence of ->d_revalidate() means that the state known
> > to dcache is only semi-reliable, so making atomic create/open
> > promises is problematic (O_EXCL for example).
>
> The idea of using ->d_revalidate as a gate is certainly interesting.
> Apart from the interaction with case-insensitivity, if we initially only
> supported filesystems that don't have ->d_revalidate, I think we would get
> coverage for enough filesystems to be interesting. We could then take a
> bit more time to think through the rest of the picture.
>
> I don't think O_EXCL is at all problematic. ->mkdir() is already
> required to return -EEXIST if the directory already exists.
>
> >
> > The problem is that some fs (overlayfs/ext4/f2fs) register
> > a mostly-noop ->d_revalidate() so I proposed how to deal with those.
>
> Presumably the fs would provide two dentry_operations structures and
> choose which to pass to set_default_d_op() when creating the superblock.
Yes, either that or clear sb->s_d_flags & DCACHE_OP_REVALIDATE
when d_revalidate is not going to be used.
> Doing that would even provide slightly better performance in the
> case where no d_revalidate is needed.
>
Very slightly, so this should not be a reason, but advertising corrects
contracts to vfs.
For example, if overlayfs would want to install:
static const struct dentry_operations ovl_dentry_real_operations = {
.d_real = ovl_d_real,
};
It needs to know that none of the underlying dentries is expected to have
a d_revalidate op at super fill time.
Thanks,
Amir.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-10-01 1:00 ` NeilBrown
@ 2026-10-01 14:07 ` Jori Koolstra
0 siblings, 0 replies; 40+ messages in thread
From: Jori Koolstra @ 2026-10-01 14:07 UTC (permalink / raw)
To: NeilBrown, NeilBrown
Cc: Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Amir Goldstein, Jan Kara, linux-fsdevel, linux-kernel
> Op 01-10-2026 03:00 CEST schreef NeilBrown <neilb@ownmail.net>:
>
>
> On Thu, 01 Oct 2026, Jori Koolstra wrote:
> > > Op 30-09-2026 18:56 EDT schreef NeilBrown <neilb@ownmail.net>:
> > >
> > >
> > > On Thu, 01 Oct 2026, Jori Koolstra wrote:
> > > >
> > > > Err, *derp*, what a stupid suggestion of mine.
> > > >
> > > > > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > > > > remain a -EINVAL on any filesystem which doesn't completely support
> > > > > the functionality.
> > > > >
> > > >
> > > > I don't think that works for the reason I just wrote in my email to Amir:
> > > > it would make lookup dependent on the dentry cache. If it's in-cache, you get
> > > > your dir, otherwise suddenly -EINVAL.
> > > >
> > > > > To do that we need some way to detect kernfs and tracefs. I think
> > > > > the only way we can do that is to make some change to those two
> > > > > filesystems.
> > > > > Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
> > > > >
> > > >
> > > > We can just implement atomic_open() for kernfs/tracefs, do a lookup there,
> > > > and if negative with O_CREAT return maybe -ENOENT (or really we need a new
> > > > error that says "the requested create could not be serviced," like -ENOCREATE,
> > > > or whatever). And if it is positive we do finish_no_open().
> > > >
> > > > It's a bit of a hack because it does not really have anything to do with
> > > > atomicity, but it does short-circuit the mkdir call in lookup_open(). I guess
> > > > that would work. Maybe I am confused, but wasn't that what you proposed here
> > > > earlier?
> > >
> > > Yes, it is what I proposed earlier. But I think it would require more
> > > review and probably make it unrealistic to land this cycle. But I'm not
> > > thinking it is unlikely to be ready this cycle any way.
> > >
> >
> > Not unlikely, so likely that is :)
> >
> > > I'm now wondering if we should keep ->atomic_open out of the loop and
> > > always use ->mkdir to create a directory.
> > > Based on your justification you probably always want O_EXCL and I would
> > > be inclined to require that.
> > >
> >
> > I wanted it to work as regular O_CREAT, so no forced O_EXCL per se.
>
> Why? The whole point is atomicity, and without O_EXCL there is no
> atomicity.
> I guess it doesn't hurt to not require O_EXCL, but it is just another
> combination to test...
>
Let me rephrase your point: if you do an O_CREAT without O_EXCL you still
don't know whether it's your file because it might have been there already.
So the fact that you immediately get an fd still does not tell you everything.
Right?
But, at least you'll know it's a directory and you still save a syscall.
On the other hand, maybe that just breeds incorrect use, and if needed the
O_EXCL-less variant can always be added, so we might as well force its use
for now. Then again, it might be confusing to do open(O_CREAT|O_DIRECTORY)
and get an -EINVAL or -EEXISTS, given current O_CREAT semantics for regular
files.
> >
> > > So if the dentry is in-lookup we call ->atomic_open(O_DIRECTORY). If
> > > that succeeds - good. If it reports ENOENT or a negative dentry, then
> >
> > Can ->atomic_open() return a negative dentry?
>
> No reason why not. It is mapped to -ENOENT
>
> if (unlikely(d_is_negative(dentry)))
> error = -ENOENT;
>
> I'm in favour of ->atomic_open() calling finish_no_open() is all
> no-error cases where it didn't do something about opening the file.
> But I don't object to an explicit -ENOENT return.
>
I don't have LLM access right now: are there any ->atomic_open() implementations
currently that do return negative dentries? Seems to be a bit strange, just like
with ->mkdir(). What would be the purpose?
> >
> > > we cal ->mkdir. If that succeeds with a positive dentry, we call
> > > through to call ->open.
> > > If ->mkdir succeeds with a negative dentry - we have the problem of
> > > kernfs and tracefs. I'm leaning towards fixing those to do the lookup.
> >
> > Yes, I follow this. But we do need to ask their maintainers then why that
> > was chosen. Plus, we need to ban negative dentry returns also for future
> > fses, which may be limiting.
>
> Odds a good that there wasn' a specific reason why leaving the dentry
> negative was chosen. I think the better question is "do you have a
> problem with this patch" where the patch does a lookup.
>
> Putting a WARN_ON(d_really_is_negative()) somewhere in vfs_mkdir would be
> easy enough. I cannot imagine it really being a burden for any fs. If
> that does happen were can solve the problem then. Being prepared of all
> possible eventuality is a no-win proposition. I've tried - I gave up.
>
Fair enough. I'll spin this idea up then, to add the lookup to kern/tracefs
->mkdir() and add that warning to vfs_mkdir(). And I'll CC those maintainers
on that version.
> >
> > >
> > > I don't think any filesystems *can* combine mkdir with open, so not
> > > using ->atomic_open for the mkdir doesn't actually lose anything.
> > >
> >
> > Yeah, I don't quite understand the specifics of this as I am not familiar
> > with NFS and the likes. I presume because of things like network traffic
> > ->atomic_open() was needed as a single call into the underlying fs. But
> > I have no idea why or why not that would make sense for directories. If
> > we do it the way you suggest now, we can get rid of all the O_CREAT stripping
>
> ->atomic_open was created specifically for NFSv4 which has a combined
> lookup/create/open request. You can do everything with a combination of
> ->lookup and (exclusive) ->create (providing create returns a
> non-negative dentry) and ->truncate and open. Using ->atomic_open means
> fewer round-trips to the server so less latency.
>
> NFS doesn't have a concept of "open" for directories. CIFS probably
> does. FUSE probably doesn't but could possibly add one.
> But directories are not opened nearly as often as files, so combining
> things isn't so important. And O_TRUNC is not meaningful for
> directories so that is one fewer thing that could be combined.
>
OK, if it isn't likely that any ->atomic_open() wants O_CREAT|O_DIRECTORY
support we can get rid of some of the complexity. Summarize:
1. use ->atomic_open(O_DIRECTORY) for lookup, no O_CREAT bit passed in.
2. if positive then done, otherwise (on -ENOENT) we call ->mkdir()
3. open that dentry in the regular way, WARN_ON negative dentry return
Best,
Jori.
PS. Maybe I asked already before (in case apologies), but will you be
at LPC?
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-10-01 9:29 ` Amir Goldstein
2026-10-01 10:08 ` NeilBrown
@ 2026-10-01 15:42 ` Jori Koolstra
2026-10-01 15:59 ` Amir Goldstein
1 sibling, 1 reply; 40+ messages in thread
From: Jori Koolstra @ 2026-10-01 15:42 UTC (permalink / raw)
To: Amir Goldstein
Cc: NeilBrown, Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Jan Kara, linux-fsdevel, linux-kernel, Theodore Tso
> Op 01-10-2026 11:29 CEST schreef Amir Goldstein <amir73il@gmail.com>:
>
>
> On Thu, Oct 1, 2026 at 12:16 AM Jori Koolstra <jkoolstra@xs4all.nl> wrote:
> >
> >
> > > Op 30-09-2026 05:45 EDT schreef Amir Goldstein <amir73il@gmail.com>:
> > >
> > >
> > > On Wed, Sep 30, 2026 at 12:15 AM NeilBrown <neilb@ownmail.net> wrote:
> > > >
> ...
> > > > The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> > > > might create the directory, then return -EOPNOTSUPP. This is weird and
> > > > I'd rather it not be visible.
> > > >
> > > > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > > > remain a -EINVAL on any filesystem which doesn't completely support
> > > > the functionality.
> > > >
> > >
> > > Joining late to this party so apologies in advance if my questions
> > > have already been addressed.
> > >
> > > I agree with Neil's statement above, but IMO, the atomic_open() fs
> > > match the description of "doesn't completely support the functionality."
> > > Therefore, I think that rather than success if directory exists, they
> > > should also return -EINVAL/-EOPNOTSUPP consistently (see below).
> > >
> >
> > The issue with this is that if you want per fs atomic_open() opt-in (in
> > contrast to either implementing all instances in one release or disabling
> > all), you have the issue that your lookup now depends on the caching status
> > of the directory dentry.
> >
> > If that dentry is in cache, and positive, d_lookup() earlier in lookup_open()
> > makes it return early:
> >
> > if (dentry->d_inode) {
> > /* Cached positive dentry: will open in do_open(). */
> > goto out;
> > }
> >
> > So you get your lookup. But if the same dentry is not in cache, now you
> > suddenly get -EINVAL. I thought that behavior was more unwanted then
> > what I eventually settled on, namely to strip the O_CREAT bit.
> >
>
> Maybe I am missing something, but I think you misunderstand me.
> What I mean is - if directory inode has a ->atomic_open() op,
> bail early with -EINVAL/-EOPNOTSUPP, because this is a network
> filesystem that does not support atomic O_CREATE|O_DIRECTORY
> and in most likelihood never will support it.
>
> This gating criteria is not dependent on cache state,
> which is what we wanted.
>
No in that case I think I've understood you (or maybe still not?) My
point is that we can't do that if we want to be able to individually
support O_CREAT|O_DIRECTORY for some ->atomic_open fs. If we bail early
the how can say only NFS support it at some point? You get into the
situation where everybody needs to have support or no one.
Does that make sense, or do I still misunderstand you point?
> The justification of using ->d_revalidate() as another opt-out
> is that existence of ->d_revalidate() means that the state known
> to dcache is only semi-reliable, so making atomic create/open
> promises is problematic (O_EXCL for example).
>
I must admit I don't know too much about what the revalidate step is for.
What does semi-reliable mean here? That the dcache might have returned a
dentry that has timed out in some sense and does not reflect the actual
state of the fs?
But if we force O_EXCL like Neil suggested, is that a problem? You'd only
get an fd if you actually created the thing.
I should really start to learn and contribute to an actual fs that is used
instead of just staying at the VFS level. It would make understanding all
these subtleties easier...
> The problem is that some fs (overlayfs/ext4/f2fs) register
> a mostly-noop ->d_revalidate() so I proposed how to deal with those.
>
> Thanks,
> Amir.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-10-01 15:42 ` Jori Koolstra
@ 2026-10-01 15:59 ` Amir Goldstein
2026-10-01 16:23 ` Jori Koolstra
0 siblings, 1 reply; 40+ messages in thread
From: Amir Goldstein @ 2026-10-01 15:59 UTC (permalink / raw)
To: Jori Koolstra
Cc: NeilBrown, Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Jan Kara, linux-fsdevel, linux-kernel, Theodore Tso
> > Maybe I am missing something, but I think you misunderstand me.
> > What I mean is - if directory inode has a ->atomic_open() op,
> > bail early with -EINVAL/-EOPNOTSUPP, because this is a network
> > filesystem that does not support atomic O_CREATE|O_DIRECTORY
> > and in most likelihood never will support it.
> >
> > This gating criteria is not dependent on cache state,
> > which is what we wanted.
> >
>
> No in that case I think I've understood you (or maybe still not?) My
> point is that we can't do that if we want to be able to individually
> support O_CREAT|O_DIRECTORY for some ->atomic_open fs. If we bail early
> the how can say only NFS support it at some point? You get into the
> situation where everybody needs to have support or no one.
>
> Does that make sense, or do I still misunderstand you point?
>
If we gate on an existing criteria now like ->atomic_open()
and/or ->d_revalidate(), we keep semantics simple and it
does not limit us to change the criteria in the future.
But I don't insist on this criteria - it's just a suggestion.
Thanks,
Amir.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-10-01 15:59 ` Amir Goldstein
@ 2026-10-01 16:23 ` Jori Koolstra
2026-10-01 17:53 ` Amir Goldstein
0 siblings, 1 reply; 40+ messages in thread
From: Jori Koolstra @ 2026-10-01 16:23 UTC (permalink / raw)
To: Amir Goldstein
Cc: NeilBrown, Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Jan Kara, linux-fsdevel, linux-kernel, Theodore Tso
> Op 01-10-2026 17:59 CEST schreef Amir Goldstein <amir73il@gmail.com>:
>
>
> > > Maybe I am missing something, but I think you misunderstand me.
> > > What I mean is - if directory inode has a ->atomic_open() op,
> > > bail early with -EINVAL/-EOPNOTSUPP, because this is a network
> > > filesystem that does not support atomic O_CREATE|O_DIRECTORY
> > > and in most likelihood never will support it.
> > >
> > > This gating criteria is not dependent on cache state,
> > > which is what we wanted.
> > >
> >
> > No in that case I think I've understood you (or maybe still not?) My
> > point is that we can't do that if we want to be able to individually
> > support O_CREAT|O_DIRECTORY for some ->atomic_open fs. If we bail early
> > the how can say only NFS support it at some point? You get into the
> > situation where everybody needs to have support or no one.
> >
> > Does that make sense, or do I still misunderstand you point?
> >
>
> If we gate on an existing criteria now like ->atomic_open()
> and/or ->d_revalidate(), we keep semantics simple and it
> does not limit us to change the criteria in the future.
>
> But I don't insist on this criteria - it's just a suggestion.
>
Oh no, I don't think it's a bad idea. I was just trying to understand.
And like I mentioned, I don't understand the detail of ->d_revalidate()
usage.
What do you think about Neil's suggestion to keep ->atomic_open out
of directory creation for the time being with:
1. use ->atomic_open(O_DIRECTORY) for lookup, no O_CREAT bit passed in.
2. if positive then done, otherwise (on -ENOENT from atomic_open()) we
call ->mkdir
3. open that dentry in the regular way, WARN_ON negative dentry return
But I don't know if the kind of atomicity I care about is affected by
->d_revalidate() filesystems. With O_CREAT|O_DIRECTORY we just want provide
the mechanism to ensure that the process actually did the create you get
the fd for.
Best,
Jori.
^ permalink raw reply [flat|nested] 40+ messages in thread
* Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
2026-10-01 16:23 ` Jori Koolstra
@ 2026-10-01 17:53 ` Amir Goldstein
0 siblings, 0 replies; 40+ messages in thread
From: Amir Goldstein @ 2026-10-01 17:53 UTC (permalink / raw)
To: Jori Koolstra
Cc: NeilBrown, Christian Brauner, Jeff Layton, Al Viro, Aleksa Sarai,
Jan Kara, linux-fsdevel, linux-kernel, Theodore Tso
On Thu, Oct 1, 2026 at 6:23 PM Jori Koolstra <jkoolstra@xs4all.nl> wrote:
>
>
> > Op 01-10-2026 17:59 CEST schreef Amir Goldstein <amir73il@gmail.com>:
> >
> >
> > > > Maybe I am missing something, but I think you misunderstand me.
> > > > What I mean is - if directory inode has a ->atomic_open() op,
> > > > bail early with -EINVAL/-EOPNOTSUPP, because this is a network
> > > > filesystem that does not support atomic O_CREATE|O_DIRECTORY
> > > > and in most likelihood never will support it.
> > > >
> > > > This gating criteria is not dependent on cache state,
> > > > which is what we wanted.
> > > >
> > >
> > > No in that case I think I've understood you (or maybe still not?) My
> > > point is that we can't do that if we want to be able to individually
> > > support O_CREAT|O_DIRECTORY for some ->atomic_open fs. If we bail early
> > > the how can say only NFS support it at some point? You get into the
> > > situation where everybody needs to have support or no one.
> > >
> > > Does that make sense, or do I still misunderstand you point?
> > >
> >
> > If we gate on an existing criteria now like ->atomic_open()
> > and/or ->d_revalidate(), we keep semantics simple and it
> > does not limit us to change the criteria in the future.
> >
> > But I don't insist on this criteria - it's just a suggestion.
> >
>
> Oh no, I don't think it's a bad idea.
It probably is though..
> I was just trying to understand.
> And like I mentioned, I don't understand the detail of ->d_revalidate()
> usage.
>
> What do you think about Neil's suggestion to keep ->atomic_open out
> of directory creation for the time being with:
>
> 1. use ->atomic_open(O_DIRECTORY) for lookup, no O_CREAT bit passed in.
> 2. if positive then done, otherwise (on -ENOENT from atomic_open()) we
> call ->mkdir
> 3. open that dentry in the regular way, WARN_ON negative dentry return
>
I don't know enough about network filesystems to say, but it sounds fishy.
I think we want consistent defined behavior for O_CREAT|O_DIRECTORY
for filesystems that support it and my understanding is that this
suggestion is not going in that direction?
> But I don't know if the kind of atomicity I care about is affected by
> ->d_revalidate() filesystems.
It is not directly related. It is circumstantially related.
->d_revalidate() *could* mean, in the worst case, that holding the ref
on the dentry gives no guarantee on keeping the object behind it alive.
In the worst case, maybe a new object with same inode number could
be created in the same name.
I am not sure this can happen in many filesystems for real, but I
am pretty sure that it can happen with FUSE.
> With O_CREAT|O_DIRECTORY we just want provide
> the mechanism to ensure that the process actually did the create you get
> the fd for.
Yes. And I don't know if you can get this guarantee from all network
filesystems. Requires an audit.
Thanks,
Amir.
^ permalink raw reply [flat|nested] 40+ messages in thread
end of thread, other threads:[~2026-10-01 17:53 UTC | newest]
Thread overview: 40+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-13 18:50 [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 01/12] fs/namei.c: use trailing_slashes() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 02/12] vfs: prepare vfs_creat|mkdir_no_perm for reuse in lookup_open() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 03/12] vfs: lookup_open(): move setting FMODE_CREATED down Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 04/12] vfs: move ->create check in lookup_open() to before try_break_deleg() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 05/12] vfs: lookup_open(): use vfs_create_no_perm() Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 06/12] vfs: lookup_open(): lock the parent as I_MUTEX_PARENT Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Jori Koolstra
2026-09-18 8:17 ` Christian Brauner
2026-09-18 10:07 ` NeilBrown
2026-09-19 0:01 ` NeilBrown
2026-09-25 13:57 ` Christian Brauner
2026-09-25 21:26 ` NeilBrown
2026-09-25 22:53 ` Jori Koolstra
2026-09-29 11:54 ` Jori Koolstra
2026-09-29 22:15 ` NeilBrown
2026-09-30 9:45 ` Amir Goldstein
2026-09-30 22:16 ` Jori Koolstra
2026-10-01 9:29 ` Amir Goldstein
2026-10-01 10:08 ` NeilBrown
2026-10-01 11:29 ` Amir Goldstein
2026-10-01 15:42 ` Jori Koolstra
2026-10-01 15:59 ` Amir Goldstein
2026-10-01 16:23 ` Jori Koolstra
2026-10-01 17:53 ` Amir Goldstein
2026-09-30 22:33 ` Jori Koolstra
2026-09-30 22:56 ` NeilBrown
2026-09-30 23:25 ` Jori Koolstra
2026-10-01 1:00 ` NeilBrown
2026-10-01 14:07 ` Jori Koolstra
2026-09-25 23:13 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 08/12] vfs: change ->create/->mkdir operations unavailable errno Jori Koolstra
2026-09-18 8:04 ` Christian Brauner
2026-09-25 22:55 ` Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 09/12] vfs: move O_IS_MKDIR check from lookup_open() into individual filesystems Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 10/12] vfs: refuse O_CREAT for directories through a dangling symlink Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 11/12] vfs: short-circuit MAY_WRITE access for O_DIRECTORY opens Jori Koolstra
2026-09-13 18:50 ` [PATCH v6 12/12] selftest: add tests for open*(O_CREAT|O_DIRECTORY) Jori Koolstra
2026-09-17 10:38 ` [PATCH v6 00/12] vfs: add O_CREAT|O_DIRECTORY to open*(2) Christian Brauner
2026-09-18 8:18 ` Christian Brauner
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®