mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v2 0/4] md: add a control device for array management
@ 2026-10-01 12:56 Abd-Alrhman Masalkhi
  2026-10-01 12:56 ` [RFC PATCH v2 1/4] md: add uapi definitions for the md control device Abd-Alrhman Masalkhi
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Abd-Alrhman Masalkhi @ 2026-10-01 12:56 UTC (permalink / raw)
  To: song, yukuai, chengzhihao1, magiclinan, xiao
  Cc: linux-raid, linux-kernel, Abd-Alrhman Masalkhi

Hi,

This series adds a new /dev/md-control misc character device for
managing md arrays without opening the corresponding md block device.

Currently, md management ioctls are issued on the md block device
itself, which requires the caller to hold the array open while
configuring or stopping it.

This is a problem for STOP_ARRAY and STOP_ARRAY_RO. Before the array is
stopped, the page cache must be flushed, and no other task may have the
device open or be writing to it. The issue arises when several tasks
share the same file descriptor table, as they count as a single opener.
Consequently, one task may still be writing while another task flushes
the page cache and stops the array. A write within this window can race
with the stop operation.

The new control device is not tied to any md array. Each request
identifies the target array by name, UUID, or device number, and
carries a flags field. Unknown flags are rejected with -ENOTTY. The new
commands mirror the existing md block-device ioctls, with STOP_ARRAY_RO
represented by MD_STOP_ARRAY with MD_RO_FLAG set.

New 64-bit structures, including mdu_ioctl, are introduced
(mdu_array_info64, mdu_disk_info64, mdu_param64, mdu_bitmap_file64,
and mdu_version64) to resolve padding and overflow issues in fields
such as size, ctime, and utime.

The existing md block-device ioctl interface remains unchanged.
A warning is emitted when it is used to recommend upgrading mdadm to
use the new control interface.

The mdadm has been modified correspondingly:
Link: https://lore.kernel.org/linux-raid/20260928211849.3602414-1-abd.masalkhi@gmail.com

This is an RFC because the new UAPI (struct mdu_ioctl, mdu_array_info64,
mdu_disk_info64, mdu_param64, mdu_bitmap_file64, and mdu_version64 and
the new command set). I am thinking about adding a new command
MD_NEW_ARRAY or MD_CREATE_ARRAY to create a new array. Feedback on the
interface is welcome.

I am aslo considering adding a command to create a new array
MD_NEW_ARRAY/MD_CREATE_ARRAY.

I am thinking of adding a new command MD_NEW_ARRAY or MD_CREATE_ARRAY
to create a new array.

Changes in v2:
 - Move the code from md-ctl.c into md.c and remove md-ctl.c.
 - Handle the new commands directly in mdctl_ioctl() instead of
   converting them to the old ioctls.
 - Use a dynamic misc minor number for /dev/md-control.
 - Drop the devname module alias, which only works with a fixed minor.
 - Fix the issues reported by sashiko-bot.
 - Take disks_mutex around the lookups, so that they never see an
   mddev that md_alloc() has not finished.
 - Clear hold_active only when a command succeeds.
 - Link-v1: https://lore.kernel.org/linux-raid/20260928201224.3602262-1-abd.masalkhi@gmail.com

Abd-Alrhman Masalkhi

Abd-Alrhman Masalkhi (4):
  md: add uapi definitions for the md control device
  md: pass struct mdu_disk_info64 to md_add_new_disk()
  md: use struct mdu_array_info64 for SET_ARRAY_INFO
  md: add a control device for array management

 drivers/md/md-autodetect.c     |   4 +-
 drivers/md/md.c                | 700 ++++++++++++++++++++++++++++++++-
 drivers/md/md.h                |  17 +-
 include/uapi/linux/raid/md_u.h | 119 ++++++
 4 files changed, 820 insertions(+), 20 deletions(-)


base-commit: 5c2f4115051d064ad79b0b4edac70a2b62528d9d
-- 
2.43.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC PATCH v2 1/4] md: add uapi definitions for the md control device
  2026-10-01 12:56 [RFC PATCH v2 0/4] md: add a control device for array management Abd-Alrhman Masalkhi
@ 2026-10-01 12:56 ` Abd-Alrhman Masalkhi
  2026-10-01 12:56 ` [RFC PATCH v2 2/4] md: pass struct mdu_disk_info64 to md_add_new_disk() Abd-Alrhman Masalkhi
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 5+ messages in thread
From: Abd-Alrhman Masalkhi @ 2026-10-01 12:56 UTC (permalink / raw)
  To: song, yukuai, chengzhihao1, magiclinan, xiao
  Cc: linux-raid, linux-kernel, Abd-Alrhman Masalkhi

Add the ioctl commands and structures for /dev/md-control, a misc device
that manages md arrays without opening the md block device. The device
itself is added in the next patch.

Every command except MD_RAID_VERSION takes a struct mdu_ioctl, which
names the array by name, UUID or device number, carries a flags field,
and points to the command's payload. The commands mirror the existing
block device ioctls and use the MD_MAJOR ioctl type, with numbers
starting at 0x40, grouped like the old ones: status at 0x40,
configuration at 0x50 and usage at 0x60.

The payload structures are new versions of the old ones with fixed-size
types and explicit padding, so they have the same layout on 32-bit and
64-bit systems. Fields that overflow in the old structures are 64-bit
now: size, ctime and utime in struct mdu_array_info64.
struct mdu_array_info64 also reports the number of journal disks.

STOP_ARRAY_RO has no separate command: MD_STOP_ARRAY with MD_RO_FLAG
set switches the array to read-only mode instead of stopping it.

Signed-off-by: Abd-Alrhman Masalkhi <abd.masalkhi@gmail.com>
---
 include/uapi/linux/raid/md_u.h | 119 +++++++++++++++++++++++++++++++++
 1 file changed, 119 insertions(+)

diff --git a/include/uapi/linux/raid/md_u.h b/include/uapi/linux/raid/md_u.h
index a893010735fb..adc62b5e50ec 100644
--- a/include/uapi/linux/raid/md_u.h
+++ b/include/uapi/linux/raid/md_u.h
@@ -12,6 +12,8 @@
 #ifndef _UAPI_MD_U_H
 #define _UAPI_MD_U_H
 
+#include <linux/types.h>
+
 /*
  * Different major versions are not compatible.
  * Different minor versions are only downward compatible.
@@ -30,6 +32,10 @@
  */
 #define MD_PATCHLEVEL_VERSION           3
 
+#define MD_NAME_LEN                     32
+#define MD_UUID_LEN                     16
+#define MD_CTL_NODE                     "md-control"
+
 /* ioctls */
 
 /* status */
@@ -61,6 +67,29 @@
 #define RESTART_ARRAY_RW	_IO (MD_MAJOR, 0x34)
 #define CLUSTERED_DISK_NACK	_IO (MD_MAJOR, 0x35)
 
+/* ioctl commands for the MD misc control driver */
+/* status */
+#define MD_RAID_VERSION		_IOR(MD_MAJOR, 0x40, struct mdu_version64)
+#define MD_GET_ARRAY_INFO	_IOWR(MD_MAJOR, 0x41, struct mdu_ioctl)
+#define MD_GET_DISK_INFO	_IOWR(MD_MAJOR, 0x42, struct mdu_ioctl)
+#define MD_RAID_AUTORUN		_IOWR(MD_MAJOR, 0x43, struct mdu_ioctl)
+#define MD_GET_BITMAP_FILE	_IOWR(MD_MAJOR, 0x44, struct mdu_ioctl)
+
+/* configuration */
+#define MD_ADD_NEW_DISK		_IOWR(MD_MAJOR, 0x50, struct mdu_ioctl)
+#define MD_HOT_ADD_DISK		_IOWR(MD_MAJOR, 0x51, struct mdu_ioctl)
+#define MD_HOT_REMOVE_DISK	_IOWR(MD_MAJOR, 0x52, struct mdu_ioctl)
+#define MD_SET_ARRAY_INFO	_IOWR(MD_MAJOR, 0x53, struct mdu_ioctl)
+#define MD_SET_DISK_INFO	_IOWR(MD_MAJOR, 0x54, struct mdu_ioctl)
+#define MD_SET_DISK_FAULTY	_IOWR(MD_MAJOR, 0x55, struct mdu_ioctl)
+#define MD_SET_BITMAP_FILE	_IOWR(MD_MAJOR, 0x56, struct mdu_ioctl)
+
+/* usage */
+#define MD_RUN_ARRAY		_IOWR(MD_MAJOR, 0x60, struct mdu_ioctl)
+#define MD_STOP_ARRAY		_IOWR(MD_MAJOR, 0x61, struct mdu_ioctl)
+#define MD_RESTART_ARRAY_RW	_IOWR(MD_MAJOR, 0x62, struct mdu_ioctl)
+#define MD_CLUSTERED_DISK_NACK	_IOWR(MD_MAJOR, 0x63, struct mdu_ioctl)
+
 /* 63 partitions with the alternate major number (mdp) */
 #define MdpMinorShift 6
 
@@ -146,4 +175,94 @@ typedef struct mdu_param_s
 	int			max_fault;	/* unused for now */
 } mdu_param_t;
 
+/* structures for the MD misc control driver */
+struct mdu_version64 {
+	__u32 major;
+	__u32 minor;
+	__u32 patchlevel;
+	__u32 padding;
+};
+
+struct mdu_array_info64 {
+	/*
+	 * Generic constant information
+	 */
+	__u32 major_version;
+	__u32 minor_version;
+	__u32 patch_version;
+	__s32 level;
+
+	__u64 ctime;
+	__s64 size;
+
+	__u32 nr_disks;
+	__u32 raid_disks;
+	__u32 md_minor;
+	__u32 not_persistent;
+
+	/*
+	 * Generic state information
+	 */
+	__u64 utime;		/*  0 Superblock update time		      */
+
+	__u32 state;		/*  1 State bits (clean, ...)		      */
+	__u32 active_disks;	/*  2 Number of currently active disks	      */
+	__u32 working_disks;	/*  3 Number of working disks		      */
+	__u32 failed_disks;	/*  4 Number of failed disks		      */
+	__u32 spare_disks;	/*  5 Number of spare disks		      */
+	__u32 journal_disks;    /*  6 Number of disks used for journaling     */
+
+	/*
+	 * Personality information
+	 */
+	__s32 layout;		/*  0 the array's physical layout	      */
+	__u32 chunk_size;	/*  1 chunk size in bytes		      */
+
+};
+
+struct mdu_disk_info64 {
+	/*
+	 * configuration/status of one particular disk
+	 */
+	__u32 state;
+	__u32 major;
+	__u32 minor;
+	__u32 padding;
+
+	__s32 number;
+	__s32 raid_disk;
+};
+
+struct mdu_bitmap_file64 {
+	char pathname[4096];
+};
+
+struct mdu_param64 {
+	__s32 personality;	/* 1,2,3,4 */
+	__u32 max_fault;	/* unused for now */
+	__u32 chunk_size;	/* in bytes */
+	__u32 padding;
+};
+
+struct mdu_ioctl {
+	char name[MD_NAME_LEN];
+	char uuid[MD_UUID_LEN];
+	__u64 dev;
+	__u32 flags;
+	__u32 padding;
+	union {
+		__u64 array;
+		__u64 disk;
+		__u64 param;
+		__u64 bitmap_file;
+		__u64 arg;
+	};
+};
+
+/*
+ * If set, MD_STOP_ARRAY will switch the array to read-only mode instead
+ * of fully stopping it.
+ */
+#define MD_RO_FLAG              (1u << 0)
+
 #endif /* _UAPI_MD_U_H */
-- 
2.43.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC PATCH v2 2/4] md: pass struct mdu_disk_info64 to md_add_new_disk()
  2026-10-01 12:56 [RFC PATCH v2 0/4] md: add a control device for array management Abd-Alrhman Masalkhi
  2026-10-01 12:56 ` [RFC PATCH v2 1/4] md: add uapi definitions for the md control device Abd-Alrhman Masalkhi
@ 2026-10-01 12:56 ` Abd-Alrhman Masalkhi
  2026-10-01 12:56 ` [RFC PATCH v2 3/4] md: use struct mdu_array_info64 for SET_ARRAY_INFO Abd-Alrhman Masalkhi
  2026-10-01 12:56 ` [RFC PATCH v2 4/4] md: add a control device for array management Abd-Alrhman Masalkhi
  3 siblings, 0 replies; 5+ messages in thread
From: Abd-Alrhman Masalkhi @ 2026-10-01 12:56 UTC (permalink / raw)
  To: song, yukuai, chengzhihao1, magiclinan, xiao
  Cc: linux-raid, linux-kernel, Abd-Alrhman Masalkhi

The md control device passes disk information as struct mdu_disk_info64.
Make md_add_new_disk() take this structure, so that the control device
and the old ADD_NEW_DISK ioctl can share it.

The ADD_NEW_DISK ioctl converts its mdu_disk_info_t with the new helper
convert_to_disk_info64(), and md_setup_drive() uses a struct
mdu_disk_info64 directly.

No functional change.

Signed-off-by: Abd-Alrhman Masalkhi <abd.masalkhi@gmail.com>
---
 drivers/md/md-autodetect.c |  2 +-
 drivers/md/md.c            | 26 +++++++++++++++++++++-----
 drivers/md/md.h            |  4 ++--
 3 files changed, 24 insertions(+), 8 deletions(-)

diff --git a/drivers/md/md-autodetect.c b/drivers/md/md-autodetect.c
index 4b80165afd23..9ba061f1a628 100644
--- a/drivers/md/md-autodetect.c
+++ b/drivers/md/md-autodetect.c
@@ -201,7 +201,7 @@ static void __init md_setup_drive(struct md_setup_args *args)
 	err = md_set_array_info(mddev, &ainfo);
 
 	for (i = 0; i <= MD_SB_DISKS && devices[i]; i++) {
-		struct mdu_disk_info_s dinfo = {
+		struct mdu_disk_info64 dinfo = {
 			.major	= MAJOR(devices[i]),
 			.minor	= MINOR(devices[i]),
 		};
diff --git a/drivers/md/md.c b/drivers/md/md.c
index 154ee5a65cb7..be1634ae1d09 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -7486,7 +7486,7 @@ static int get_disk_info(struct mddev *mddev, void __user * arg)
 	return 0;
 }
 
-int md_add_new_disk(struct mddev *mddev, struct mdu_disk_info_s *info)
+int md_add_new_disk(struct mddev *mddev, struct mdu_disk_info64 *info)
 {
 	struct md_rdev *rdev;
 	dev_t dev = MKDEV(info->major,info->minor);
@@ -8305,6 +8305,16 @@ static bool md_ioctl_need_suspend(unsigned int cmd)
 	}
 }
 
+static void convert_to_disk_info64(struct mdu_disk_info64 *info64,
+				   mdu_disk_info_t *info)
+{
+	info64->number = info->number;
+	info64->major = info->major;
+	info64->minor = info->minor;
+	info64->raid_disk = info->raid_disk;
+	info64->state = info->state;
+}
+
 static int __md_set_array_info(struct mddev *mddev, void __user *argp)
 {
 	mdu_array_info_t info;
@@ -8452,13 +8462,16 @@ static int md_ioctl(struct block_device *bdev, blk_mode_t mode,
 		 */
 		if (mddev->pers) {
 			mdu_disk_info_t info;
+			struct mdu_disk_info64 info64 = {0};
 			if (copy_from_user(&info, argp, sizeof(info)))
 				err = -EFAULT;
 			else if (!(info.state & (1<<MD_DISK_SYNC)))
 				/* Need to clear read-only for this */
 				break;
-			else
-				err = md_add_new_disk(mddev, &info);
+			else {
+				convert_to_disk_info64(&info64, &info);
+				err = md_add_new_disk(mddev, &info64);
+			}
 			goto unlock;
 		}
 		break;
@@ -8493,10 +8506,13 @@ static int md_ioctl(struct block_device *bdev, blk_mode_t mode,
 	case ADD_NEW_DISK:
 	{
 		mdu_disk_info_t info;
+		struct mdu_disk_info64 info64 = {0};
 		if (copy_from_user(&info, argp, sizeof(info)))
 			err = -EFAULT;
-		else
-			err = md_add_new_disk(mddev, &info);
+		else {
+			convert_to_disk_info64(&info64, &info);
+			err = md_add_new_disk(mddev, &info64);
+		}
 		goto unlock;
 	}
 
diff --git a/drivers/md/md.h b/drivers/md/md.h
index 6440da292105..a84c8b97fa2c 100644
--- a/drivers/md/md.h
+++ b/drivers/md/md.h
@@ -1043,12 +1043,12 @@ static inline void mddev_unlock_and_resume(struct mddev *mddev)
 }
 
 struct mdu_array_info_s;
-struct mdu_disk_info_s;
+struct mdu_disk_info64;
 
 extern int mdp_major;
 void md_autostart_arrays(int part);
 int md_set_array_info(struct mddev *mddev, struct mdu_array_info_s *info);
-int md_add_new_disk(struct mddev *mddev, struct mdu_disk_info_s *info);
+int md_add_new_disk(struct mddev *mddev, struct mdu_disk_info64 *info);
 int do_md_run(struct mddev *mddev);
 #define MDDEV_STACK_INTEGRITY	(1u << 0)
 int mddev_stack_rdev_limits(struct mddev *mddev, struct queue_limits *lim,
-- 
2.43.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC PATCH v2 3/4] md: use struct mdu_array_info64 for SET_ARRAY_INFO
  2026-10-01 12:56 [RFC PATCH v2 0/4] md: add a control device for array management Abd-Alrhman Masalkhi
  2026-10-01 12:56 ` [RFC PATCH v2 1/4] md: add uapi definitions for the md control device Abd-Alrhman Masalkhi
  2026-10-01 12:56 ` [RFC PATCH v2 2/4] md: pass struct mdu_disk_info64 to md_add_new_disk() Abd-Alrhman Masalkhi
@ 2026-10-01 12:56 ` Abd-Alrhman Masalkhi
  2026-10-01 12:56 ` [RFC PATCH v2 4/4] md: add a control device for array management Abd-Alrhman Masalkhi
  3 siblings, 0 replies; 5+ messages in thread
From: Abd-Alrhman Masalkhi @ 2026-10-01 12:56 UTC (permalink / raw)
  To: song, yukuai, chengzhihao1, magiclinan, xiao
  Cc: linux-raid, linux-kernel, Abd-Alrhman Masalkhi

The md control device passes array information as struct
mdu_array_info64. Make md_set_array_info() and update_array_info()
take the new struct mdu_array_info64, so that the md control device
and the old SET_ARRAY_INFO ioctl can share them.

md_set_array_info() no longer checks for a negative major_version:
the field is unsigned now, and the existing check against the size
of super_types still rejects such values.

md_setup_drive() uses a struct mdu_array_info64 directly.

Signed-off-by: Abd-Alrhman Masalkhi <abd.masalkhi@gmail.com>
---
 drivers/md/md-autodetect.c |  2 +-
 drivers/md/md.c            | 37 +++++++++++++++++++++++++++++++------
 drivers/md/md.h            |  4 ++--
 3 files changed, 34 insertions(+), 9 deletions(-)

diff --git a/drivers/md/md-autodetect.c b/drivers/md/md-autodetect.c
index 9ba061f1a628..227cd4a688a5 100644
--- a/drivers/md/md-autodetect.c
+++ b/drivers/md/md-autodetect.c
@@ -124,7 +124,7 @@ static void __init md_setup_drive(struct md_setup_args *args)
 {
 	char *devname = args->device_names;
 	dev_t devices[MD_SB_DISKS + 1], mdev;
-	struct mdu_array_info_s ainfo = { };
+	struct mdu_array_info64 ainfo = { };
 	struct mddev *mddev;
 	int err = 0, i;
 	char name[16];
diff --git a/drivers/md/md.c b/drivers/md/md.c
index be1634ae1d09..63b9468285b0 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -7904,12 +7904,11 @@ static int set_bitmap_file(struct mddev *mddev, int fd)
  *  The minor and patch _version numbers are also kept incase the
  *  super_block handler wishes to interpret them.
  */
-int md_set_array_info(struct mddev *mddev, struct mdu_array_info_s *info)
+int md_set_array_info(struct mddev *mddev, struct mdu_array_info64 *info)
 {
 	if (info->raid_disks == 0) {
 		/* just setting version number for superblock loading */
-		if (info->major_version < 0 ||
-		    info->major_version >= ARRAY_SIZE(super_types) ||
+		if (info->major_version >= ARRAY_SIZE(super_types) ||
 		    super_types[info->major_version].name == NULL) {
 			/* maybe try to auto-load a module? */
 			pr_warn("md: superblock version %d not known\n",
@@ -8101,7 +8100,7 @@ static void put_cluster_ops(struct mddev *mddev)
  * Any differences that cannot be handled will cause an error.
  * Normally, only one change can be managed at a time.
  */
-static int update_array_info(struct mddev *mddev, mdu_array_info_t *info)
+static int update_array_info(struct mddev *mddev, struct mdu_array_info64 *info)
 {
 	int rv = 0;
 	int cnt = 0;
@@ -8315,9 +8314,33 @@ static void convert_to_disk_info64(struct mdu_disk_info64 *info64,
 	info64->state = info->state;
 }
 
+static void convert_to_array_info64(struct mdu_array_info64 *info64,
+				    mdu_array_info_t *info)
+{
+	info64->major_version = info->major_version;
+	info64->minor_version = info->minor_version;
+	info64->patch_version = info->patch_version;
+	info64->level = info->level;
+	info64->ctime = info->ctime;
+	info64->size = info->size;
+	info64->nr_disks = info->nr_disks;
+	info64->raid_disks = info->raid_disks;
+	info64->md_minor = info->md_minor;
+	info64->not_persistent = info->not_persistent;
+	info64->utime = info->utime;
+	info64->state = info->state;
+	info64->active_disks = info->active_disks;
+	info64->working_disks = info->working_disks;
+	info64->failed_disks = info->failed_disks;
+	info64->spare_disks = info->spare_disks;
+	info64->layout = info->layout;
+	info64->chunk_size = info->chunk_size;
+}
+
 static int __md_set_array_info(struct mddev *mddev, void __user *argp)
 {
 	mdu_array_info_t info;
+	struct mdu_array_info64 info64 = {0};
 	int err;
 
 	if (!argp)
@@ -8325,8 +8348,10 @@ static int __md_set_array_info(struct mddev *mddev, void __user *argp)
 	else if (copy_from_user(&info, argp, sizeof(info)))
 		return -EFAULT;
 
+	convert_to_array_info64(&info64, &info);
+
 	if (mddev->pers) {
-		err = update_array_info(mddev, &info);
+		err = update_array_info(mddev, &info64);
 		if (err)
 			pr_warn("md: couldn't update array info. %d\n", err);
 		return err;
@@ -8342,7 +8367,7 @@ static int __md_set_array_info(struct mddev *mddev, void __user *argp)
 		return -EBUSY;
 	}
 
-	err = md_set_array_info(mddev, &info);
+	err = md_set_array_info(mddev, &info64);
 	if (err)
 		pr_warn("md: couldn't set array info. %d\n", err);
 
diff --git a/drivers/md/md.h b/drivers/md/md.h
index a84c8b97fa2c..27de34a8bdf9 100644
--- a/drivers/md/md.h
+++ b/drivers/md/md.h
@@ -1042,12 +1042,12 @@ static inline void mddev_unlock_and_resume(struct mddev *mddev)
 	mddev_resume(mddev);
 }
 
-struct mdu_array_info_s;
+struct mdu_array_info64;
 struct mdu_disk_info64;
 
 extern int mdp_major;
 void md_autostart_arrays(int part);
-int md_set_array_info(struct mddev *mddev, struct mdu_array_info_s *info);
+int md_set_array_info(struct mddev *mddev, struct mdu_array_info64 *info);
 int md_add_new_disk(struct mddev *mddev, struct mdu_disk_info64 *info);
 int do_md_run(struct mddev *mddev);
 #define MDDEV_STACK_INTEGRITY	(1u << 0)
-- 
2.43.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC PATCH v2 4/4] md: add a control device for array management
  2026-10-01 12:56 [RFC PATCH v2 0/4] md: add a control device for array management Abd-Alrhman Masalkhi
                   ` (2 preceding siblings ...)
  2026-10-01 12:56 ` [RFC PATCH v2 3/4] md: use struct mdu_array_info64 for SET_ARRAY_INFO Abd-Alrhman Masalkhi
@ 2026-10-01 12:56 ` Abd-Alrhman Masalkhi
  3 siblings, 0 replies; 5+ messages in thread
From: Abd-Alrhman Masalkhi @ 2026-10-01 12:56 UTC (permalink / raw)
  To: song, yukuai, chengzhihao1, magiclinan, xiao
  Cc: linux-raid, linux-kernel, Abd-Alrhman Masalkhi

All md management ioctls (SET_ARRAY_INFO, ADD_NEW_DISK, RUN_ARRAY,
STOP_ARRAY, ...) are issued on the md block device itself, so the
caller must hold the array open while it configures or stops it.

This is a problem for STOP_ARRAY and STOP_ARRAY_RO. Before the array is
stopped, the page cache must be flushed, and no other task may have the
device open or be writing to it. The issue arises when several tasks
share the same file descriptor table, as they count as a single opener.
Consequently, one task may still be writing while another task flushes
the page cache and stops the array. A write within this window can race
with the stop operation.

Add a misc character device, /dev/md-control. The control device is
not tied to any md device, each request specifies the array in the
payload, either by name, by UUID or by device number, without the need
to open the md block device. Every request structure has a flags field.
Unknown flags are rejected with -ENOTTY.

Each new command matches an existing block device ioctl with only one
exception. The STOP_ARRAY_RO has no separate command, it is handled
via MD_STOP_ARRAY with the MD_RO_FLAG set.

md_alloc() adds a new mddev to all_mddevs before its gendisk exists,
and frees the mddev directly if creation fails. Move disks_mutex out
of md_alloc() and take it around the control device lookups, so that
they never take a reference to an mddev that md_alloc() has not
finished.

The existing ioctl on the md block device is unchanged, but a warning
message will be printed to recommend upgrading mdadm.

Suggested-by: Yu Kuai <yukuai@fygo.io>
Suggested-by: Zhihao Cheng <chengzhihao1@huawei.com>
Signed-off-by: Abd-Alrhman Masalkhi <abd.masalkhi@gmail.com>
---
 drivers/md/md.c | 637 +++++++++++++++++++++++++++++++++++++++++++++++-
 drivers/md/md.h |   9 +
 2 files changed, 643 insertions(+), 3 deletions(-)

diff --git a/drivers/md/md.c b/drivers/md/md.c
index 63b9468285b0..ef137662760d 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -64,6 +64,7 @@
 #include <linux/slab.h>
 #include <linux/percpu-refcount.h>
 #include <linux/part_stat.h>
+#include <linux/miscdevice.h>
 
 #include "md.h"
 #include "md-bitmap.h"
@@ -367,6 +368,8 @@ EXPORT_SYMBOL_GPL(md_new_event);
 static LIST_HEAD(all_mddevs);
 static DEFINE_SPINLOCK(all_mddevs_lock);
 
+static DEFINE_MUTEX(disks_mutex);
+
 static bool is_md_suspended(struct mddev *mddev)
 {
 	return percpu_ref_is_dying(&mddev->active_io);
@@ -6319,7 +6322,6 @@ struct mddev *md_alloc(dev_t dev, char *name)
 	 * If "name" is not NULL, the device is being created by
 	 * writing to /sys/module/md_mod/parameters/new_array.
 	 */
-	static DEFINE_MUTEX(disks_mutex);
 	struct mddev *mddev;
 	struct gendisk *disk;
 	int partitioned;
@@ -6386,8 +6388,10 @@ struct mddev *md_alloc(dev_t dev, char *name)
 	disk->events |= DISK_EVENT_MEDIA_CHANGE;
 	mddev->gendisk = disk;
 	error = add_disk(disk);
-	if (error)
+	if (error) {
+		mddev->gendisk = NULL;
 		goto out_put_disk;
+	}
 
 	kobject_init(&mddev->kobj, &md_ktype);
 	error = kobject_add(&mddev->kobj, &disk_to_dev(disk)->kobj, "%s", "md");
@@ -8375,7 +8379,7 @@ static int __md_set_array_info(struct mddev *mddev, void __user *argp)
 }
 
 static int md_ioctl(struct block_device *bdev, blk_mode_t mode,
-			unsigned int cmd, unsigned long arg)
+		    unsigned int cmd, unsigned long arg)
 {
 	int err = 0;
 	unsigned int noio_flags = 0;
@@ -8387,6 +8391,8 @@ static int md_ioctl(struct block_device *bdev, blk_mode_t mode,
 	if (err)
 		return err;
 
+	pr_warn_once("md: ioctl is deprecated and will be removed in future, please upgrade to mdadm-4.6+\n");
+
 	/*
 	 * Commands dealing with the RAID driver but not any
 	 * particular array:
@@ -8602,6 +8608,623 @@ static int md_compat_ioctl(struct block_device *bdev, blk_mode_t mode,
 }
 #endif /* CONFIG_COMPAT */
 
+static struct mddev *mdctl_get_mddev_dev(dev_t dev)
+{
+	struct mddev *mddev, *ret = ERR_PTR(-EINVAL);
+
+	if (MINOR(dev) >= (1 << MINORBITS))
+		return ret;
+
+	if (MAJOR(dev) != MD_MAJOR)
+		dev &= ~((1 << MdpMinorShift) - 1);
+
+	ret = ERR_PTR(-ENODEV);
+	spin_lock(&all_mddevs_lock);
+	list_for_each_entry(mddev, &all_mddevs, all_mddevs) {
+		if (mddev->unit == dev) {
+			ret = mddev_get(mddev);
+			if (!ret)
+				ret = ERR_PTR(-ENODEV);
+			break;
+		}
+	}
+	spin_unlock(&all_mddevs_lock);
+
+	return ret;
+}
+
+static struct mddev *mdctl_get_mddev_name(const char *name)
+{
+	unsigned long minor;
+	struct mddev *mddev, *ret = ERR_PTR(-EINVAL);
+
+	if (strncmp(name, "md", 2) || name[MD_NAME_LEN - 1])
+		return ret;
+
+	if (name[2] != '_' && (!isdigit(name[2]) ||
+			       kstrtoul(&name[2], 10, &minor) ||
+			       minor > MINORMASK))
+		return ret;
+
+	ret = ERR_PTR(-ENODEV);
+	spin_lock(&all_mddevs_lock);
+	list_for_each_entry(mddev, &all_mddevs, all_mddevs) {
+		if (mddev->gendisk &&
+		    strcmp(mddev->gendisk->disk_name, name) == 0) {
+			ret = mddev_get(mddev);
+			if (!ret)
+				ret = ERR_PTR(-ENODEV);
+			break;
+		}
+	}
+	spin_unlock(&all_mddevs_lock);
+
+	return ret;
+}
+
+static struct mddev *mdctl_get_mddev_uuid(const char *uuid)
+{
+	struct mddev *mddev, *ret = ERR_PTR(-ENODEV);
+
+	spin_lock(&all_mddevs_lock);
+	list_for_each_entry(mddev, &all_mddevs, all_mddevs) {
+		if (memcmp(mddev->uuid, uuid, sizeof(mddev->uuid)) == 0) {
+			ret = mddev_get(mddev);
+			if (!ret)
+				ret = ERR_PTR(-ENODEV);
+			break;
+		}
+	}
+	spin_unlock(&all_mddevs_lock);
+
+	return ret;
+}
+
+static struct mddev *mdctl_find_mddev(struct mdu_ioctl *md_ctl)
+{
+	struct mddev *ret = ERR_PTR(-EINVAL);
+
+	/*
+	 * md_alloc() adds a new mddev to all_mddevs before it creates the
+	 * gendisk. If creation fails. it frees the mddev directly, without
+	 * looking at ->active. All of this happens under disks_mutex.
+	 * Holding disks_mutex for the lookup means we only see mddevs that
+	 * md_alloc() has finished, so taking a reference is safe and every
+	 * mddev we find has a gendisk.
+	 */
+	mutex_lock(&disks_mutex);
+
+	if (memchr_inv(md_ctl->uuid, 0, MD_UUID_LEN)) {
+		if (*md_ctl->name || md_ctl->dev) {
+			pr_err_ratelimited("md: only supply one of name, uuid or dev number\n");
+			goto out;
+		}
+		ret = mdctl_get_mddev_uuid(md_ctl->uuid);
+	} else if (*md_ctl->name) {
+		if (md_ctl->dev) {
+			pr_err_ratelimited("md: only supply one of name, uuid or dev number\n");
+			goto out;
+		}
+		ret = mdctl_get_mddev_name(md_ctl->name);
+	} else if (md_ctl->dev) {
+		ret = mdctl_get_mddev_dev(new_decode_dev(md_ctl->dev));
+	}
+
+out:
+	mutex_unlock(&disks_mutex);
+	return ret;
+}
+
+static int mdctl_get_raid_version(struct mdu_version64 __user *user)
+{
+	struct mdu_version64 v = {0};
+
+	v.major = MD_MAJOR_VERSION;
+	v.minor = MD_MINOR_VERSION;
+	v.patchlevel = MD_PATCHLEVEL_VERSION;
+
+	if (copy_to_user(user, &v, sizeof(v)))
+		return -EFAULT;
+
+	return 0;
+}
+
+static int mdctl_get_array_info(struct mddev *mddev, void __user *arg)
+{
+	struct mdu_array_info64 info = {0};
+	int nr, working, insync, failed, spare, journal;
+	struct md_rdev *rdev;
+
+	nr = working = insync = failed = spare = journal = 0;
+	rcu_read_lock();
+	rdev_for_each_rcu(rdev, mddev) {
+		nr++;
+		if (test_bit(Faulty, &rdev->flags)) {
+			failed++;
+		} else {
+			working++;
+			if (test_bit(In_sync, &rdev->flags))
+				insync++;
+			else if (test_bit(Journal, &rdev->flags))
+				journal++;
+			else
+				spare++;
+		}
+	}
+	rcu_read_unlock();
+
+	info.major_version = mddev->major_version;
+	info.minor_version = mddev->minor_version;
+	info.patch_version = MD_PATCHLEVEL_VERSION;
+	info.ctime         = mddev->ctime;
+	info.level         = mddev->level;
+	info.size          = mddev->dev_sectors / 2;
+	info.nr_disks      = nr;
+	info.raid_disks    = mddev->raid_disks;
+	info.md_minor      = mddev->md_minor;
+	info.not_persistent = !mddev->persistent;
+
+	info.utime         = mddev->utime;
+	info.state         = 0;
+	if (mddev->in_sync)
+		info.state = (1 << MD_SB_CLEAN);
+	if (mddev->bitmap && mddev->bitmap_info.offset)
+		info.state |= (1 << MD_SB_BITMAP_PRESENT);
+	if (mddev_is_clustered(mddev))
+		info.state |= (1 << MD_SB_CLUSTERED);
+	info.active_disks  = insync;
+	info.working_disks = working;
+	info.failed_disks  = failed;
+	info.spare_disks   = spare;
+	info.journal_disks = journal;
+
+	info.layout        = mddev->layout;
+	info.chunk_size    = mddev->chunk_sectors << 9;
+
+	if (copy_to_user(arg, &info, sizeof(info)))
+		return -EFAULT;
+
+	return 0;
+}
+
+static int mdctl_get_disk_info(struct mddev *mddev, void __user *arg)
+{
+	struct mdu_disk_info64 info = {0};
+	struct md_rdev *rdev;
+
+	if (copy_from_user(&info, arg, sizeof(info)))
+		return -EFAULT;
+
+	rcu_read_lock();
+	rdev = md_find_rdev_nr_rcu(mddev, info.number);
+	if (rdev) {
+		info.major = MAJOR(rdev->bdev->bd_dev);
+		info.minor = MINOR(rdev->bdev->bd_dev);
+		info.raid_disk = rdev->raid_disk;
+		info.state = 0;
+		if (test_bit(Faulty, &rdev->flags))
+			info.state |= (1 << MD_DISK_FAULTY);
+		else if (test_bit(In_sync, &rdev->flags)) {
+			info.state |= (1 << MD_DISK_ACTIVE);
+			info.state |= (1 << MD_DISK_SYNC);
+		}
+		if (test_bit(Journal, &rdev->flags))
+			info.state |= (1 << MD_DISK_JOURNAL);
+		if (test_bit(WriteMostly, &rdev->flags))
+			info.state |= (1 << MD_DISK_WRITEMOSTLY);
+		if (test_bit(FailFast, &rdev->flags))
+			info.state |= (1 << MD_DISK_FAILFAST);
+	} else {
+		info.major = info.minor = 0;
+		info.raid_disk = -1;
+		info.state = (1 << MD_DISK_REMOVED);
+	}
+	rcu_read_unlock();
+
+	if (copy_to_user(arg, &info, sizeof(info)))
+		return -EFAULT;
+
+	return 0;
+}
+
+static int mdctl_get_bitmap_file(struct mddev *mddev, void __user *arg)
+{
+	struct mdu_bitmap_file64 *file; /* too big for stack allocation */
+	char *ptr;
+	int err;
+
+	file = kzalloc_obj(*file, GFP_NOIO);
+	if (!file)
+		return -ENOMEM;
+
+	err = 0;
+	spin_lock(&mddev->lock);
+	/* bitmap enabled */
+	if (mddev->bitmap_info.file) {
+		ptr = file_path(mddev->bitmap_info.file, file->pathname,
+				sizeof(file->pathname));
+		if (IS_ERR(ptr))
+			err = PTR_ERR(ptr);
+		else
+			memmove(file->pathname, ptr,
+				sizeof(file->pathname)-(ptr-file->pathname));
+	}
+	spin_unlock(&mddev->lock);
+
+	if (err == 0 &&
+	    copy_to_user(arg, file, sizeof(*file)))
+		err = -EFAULT;
+
+	kfree(file);
+	return err;
+}
+
+static int mdctl_set_array_info(struct mddev *mddev,
+				struct mdu_array_info64 __user *user)
+{
+	struct mdu_array_info64 info;
+	int err;
+
+	if (!user)
+		memset(&info, 0, sizeof(info));
+	else if (copy_from_user(&info, user, sizeof(info)))
+		return -EFAULT;
+
+	if (mddev->pers) {
+		err = update_array_info(mddev, &info);
+		if (err)
+			pr_warn("md: couldn't update array info. %d\n", err);
+		return err;
+	}
+
+	if (!list_empty(&mddev->disks)) {
+		pr_warn("md: array %s already has disks!\n", mdname(mddev));
+		return -EBUSY;
+	}
+
+	if (mddev->raid_disks) {
+		pr_warn("md: array %s already initialised!\n", mdname(mddev));
+		return -EBUSY;
+	}
+
+	err = md_set_array_info(mddev, &info);
+	if (err)
+		pr_warn("md: couldn't set array info. %d\n", err);
+
+	return err;
+}
+
+/*
+ * Return true if @cmd needs a configured array, false otherwise.
+ */
+static bool mdctl_ioctl_need_config(unsigned int cmd, u32 flags)
+{
+	switch (cmd) {
+	case MD_RAID_AUTORUN:
+	case MD_HOT_ADD_DISK:
+	case MD_HOT_REMOVE_DISK:
+	case MD_RESTART_ARRAY_RW:
+	case MD_CLUSTERED_DISK_NACK:
+		return true;
+
+	case MD_STOP_ARRAY:
+		/* switching to read-only needs an array, a full stop does not */
+		return !!(flags & MD_RO_FLAG);
+
+	default:
+		return false;
+	}
+}
+
+static bool mdctl_ioctl_need_suspend(unsigned int cmd)
+{
+	switch (cmd) {
+	case MD_ADD_NEW_DISK:
+	case MD_HOT_ADD_DISK:
+	case MD_HOT_REMOVE_DISK:
+	case MD_SET_BITMAP_FILE:
+	case MD_SET_ARRAY_INFO:
+		return true;
+	default:
+		return false;
+	}
+}
+
+static int mdctl_ioctl_valid(unsigned int cmd)
+{
+	switch (cmd) {
+	case MD_GET_ARRAY_INFO:
+	case MD_GET_DISK_INFO:
+	case MD_RAID_VERSION:
+		return 0;
+
+	case MD_RAID_AUTORUN:
+	case MD_SET_DISK_FAULTY:
+	case MD_GET_BITMAP_FILE:
+	case MD_ADD_NEW_DISK:
+	case MD_HOT_ADD_DISK:
+	case MD_HOT_REMOVE_DISK:
+	case MD_RESTART_ARRAY_RW:
+	case MD_RUN_ARRAY:
+	case MD_SET_ARRAY_INFO:
+	case MD_SET_BITMAP_FILE:
+	case MD_STOP_ARRAY:
+	case MD_CLUSTERED_DISK_NACK:
+		if (!capable(CAP_SYS_ADMIN))
+			return -EACCES;
+		return 0;
+	default:
+		return -ENOTTY;
+	}
+}
+
+static int mdctl_flags_valid(unsigned int cmd, u32 flags)
+{
+	switch (cmd) {
+	case MD_GET_ARRAY_INFO:
+	case MD_GET_DISK_INFO:
+	case MD_RAID_VERSION:
+	case MD_RAID_AUTORUN:
+	case MD_GET_BITMAP_FILE:
+	case MD_ADD_NEW_DISK:
+	case MD_HOT_ADD_DISK:
+	case MD_HOT_REMOVE_DISK:
+	case MD_RESTART_ARRAY_RW:
+	case MD_RUN_ARRAY:
+	case MD_SET_ARRAY_INFO:
+	case MD_SET_BITMAP_FILE:
+	case MD_SET_DISK_FAULTY:
+	case MD_CLUSTERED_DISK_NACK:
+		if (flags)
+			return -ENOTTY;
+		return 0;
+	case MD_STOP_ARRAY:
+		if (flags && flags != MD_RO_FLAG)
+			return -ENOTTY;
+		return 0;
+	default:
+		return -ENOTTY;
+	}
+}
+
+static long mdctl_ioctl(struct file *filp, uint cmd, ulong u)
+{
+	int err;
+	unsigned int noio_flags = 0;
+	struct mddev *mddev;
+	void __user *arg;
+	struct mdu_ioctl md_ctl;
+	bool suspend;
+	u32 flags;
+
+	err = mdctl_ioctl_valid(cmd);
+	if (err)
+		return err;
+
+	/*
+	 * Commands dealing with the RAID driver but not any
+	 * particular array:
+	 */
+	if (cmd == MD_RAID_VERSION)
+		return mdctl_get_raid_version((struct mdu_version64 __user *)u);
+
+	if (copy_from_user(&md_ctl, (void __user *)u, sizeof(md_ctl)))
+		return -EFAULT;
+
+	flags = md_ctl.flags;
+
+	/*
+	 * for now, Only MD_STOP_ARRAY takes flags; other commands need
+	 * flags == 0.
+	 */
+	err = mdctl_flags_valid(cmd, flags);
+	if (err)
+		return err;
+
+	mddev = mdctl_find_mddev(&md_ctl);
+	if (IS_ERR(mddev))
+		return PTR_ERR(mddev);
+
+	arg = u64_to_user_ptr(md_ctl.arg);
+
+	/* These commands do not need reconfig_mutex. */
+	err = -ENODEV;
+	switch (cmd) {
+	case MD_GET_ARRAY_INFO:
+		if (mddev_configured(mddev))
+			err = mdctl_get_array_info(mddev, arg);
+		goto out;
+
+	case MD_GET_DISK_INFO:
+		if (mddev_configured(mddev))
+			err = mdctl_get_disk_info(mddev, arg);
+		goto out;
+
+	case MD_GET_BITMAP_FILE:
+		err = mdctl_get_bitmap_file(mddev, arg);
+		goto out;
+
+	case MD_SET_DISK_FAULTY:
+		err = set_disk_faulty(mddev, new_decode_dev(md_ctl.arg));
+		goto out;
+	}
+
+	if (cmd == MD_STOP_ARRAY) {
+		/* Need to flush page cache, and ensure no openers */
+		err = mddev_set_closing_and_sync_blockdev(mddev, 0);
+		if (err)
+			goto out;
+	}
+
+	if (!md_is_rdwr(mddev))
+		flush_work(&mddev->sync_work);
+
+	suspend = mdctl_ioctl_need_suspend(cmd);
+	err = suspend ? mddev_suspend_and_lock(mddev) : mddev_lock(mddev);
+	if (err) {
+		pr_debug("md: ioctl lock interrupted, reason %d, cmd %d\n",
+			 err, cmd);
+		goto out;
+	}
+	if (suspend)
+		noio_flags = memalloc_noio_save();
+
+	/*
+	 * Commands querying/configuring an existing array.
+	 *
+	 * If the array is not initialised yet, check of the command requires
+	 * the array to be initialised.
+	 */
+	if (!mddev_configured(mddev) && mdctl_ioctl_need_config(cmd, flags)) {
+		err = -ENODEV;
+		goto unlock;
+	}
+
+	/*
+	 * Commands even a read-only array can execute:
+	 */
+	switch (cmd) {
+	case MD_SET_ARRAY_INFO:
+		err = mdctl_set_array_info(mddev, arg);
+		goto unlock;
+
+	case MD_RESTART_ARRAY_RW:
+		err = restart_array(mddev);
+		goto unlock;
+
+	case MD_STOP_ARRAY:
+		err = -EINVAL;
+		if (!flags)
+			err = do_md_stop(mddev, 0);
+		else if (flags == MD_RO_FLAG && mddev->pers)
+			err = md_set_readonly(mddev);
+		goto unlock;
+
+	case MD_HOT_REMOVE_DISK:
+		err = hot_remove_disk(mddev, new_decode_dev(md_ctl.arg));
+		goto unlock;
+
+	case MD_ADD_NEW_DISK:
+		/* We can support ADD_NEW_DISK on read-only arrays
+		 * only if we are re-adding a preexisting device.
+		 * So require mddev->pers and MD_DISK_SYNC.
+		 */
+		if (mddev->pers) {
+			struct mdu_disk_info64 info;
+
+			if (copy_from_user(&info, arg, sizeof(info)))
+				err = -EFAULT;
+			else if (!(info.state & (1 << MD_DISK_SYNC)))
+				/* Need to clear read-only for this */
+				break;
+			else
+				err = md_add_new_disk(mddev, &info);
+			goto unlock;
+		}
+		break;
+	}
+
+	/*
+	 * The remaining ioctls are changing the state of the
+	 * superblock, so we do not allow them on read-only arrays.
+	 */
+	if (!md_is_rdwr(mddev) && mddev->pers) {
+		if (mddev->ro != MD_AUTO_READ) {
+			err = -EROFS;
+			goto unlock;
+		}
+		mddev->ro = MD_RDWR;
+		sysfs_notify_dirent_safe(mddev->sysfs_state);
+		set_bit(MD_RECOVERY_NEEDED, &mddev->recovery);
+		/* mddev_unlock will wake thread */
+		/* If a device failed while we were read-only, we
+		 * need to make sure the metadata is updated now.
+		 */
+		if (test_bit(MD_SB_CHANGE_DEVS, &mddev->sb_flags)) {
+			mddev_unlock(mddev);
+			wait_event(mddev->sb_wait,
+				   !test_bit(MD_SB_CHANGE_DEVS, &mddev->sb_flags) &&
+				   !test_bit(MD_SB_CHANGE_PENDING, &mddev->sb_flags));
+			mddev_lock_nointr(mddev);
+		}
+	}
+
+	switch (cmd) {
+	case MD_ADD_NEW_DISK:
+	{
+		struct mdu_disk_info64 info;
+
+		if (copy_from_user(&info, arg, sizeof(info)))
+			err = -EFAULT;
+		else
+			err = md_add_new_disk(mddev, &info);
+
+		goto unlock;
+	}
+
+	case MD_CLUSTERED_DISK_NACK:
+		if (mddev_is_clustered(mddev))
+			mddev->cluster_ops->new_disk_ack(mddev, false);
+		else
+			err = -EINVAL;
+		goto unlock;
+
+	case MD_HOT_ADD_DISK:
+		err = hot_add_disk(mddev, new_decode_dev(md_ctl.arg));
+		goto unlock;
+
+	case MD_RUN_ARRAY:
+		err = do_md_run(mddev);
+		goto unlock;
+
+	case MD_SET_BITMAP_FILE:
+		err = set_bitmap_file(mddev, md_ctl.arg);
+		goto unlock;
+
+	default:
+		err = -EINVAL;
+		goto unlock;
+	}
+
+unlock:
+	if (!err && mddev->hold_active == UNTIL_IOCTL)
+		mddev->hold_active = 0;
+
+	if (suspend) {
+		memalloc_noio_restore(noio_flags);
+		mddev_unlock_and_resume(mddev);
+	} else {
+		mddev_unlock(mddev);
+	}
+out:
+	mddev_put(mddev);
+	return err;
+}
+
+static const struct file_operations mdctl_file_ops = {
+	.open = nonseekable_open,
+	.unlocked_ioctl = mdctl_ioctl,
+	.compat_ioctl = compat_ptr_ioctl,
+	.owner = THIS_MODULE,
+	.llseek  = noop_llseek,
+};
+
+static struct miscdevice mdctl_misc = {
+	.name = MD_CTL_NODE,
+	.minor = MISC_DYNAMIC_MINOR,
+	.fops = &mdctl_file_ops,
+};
+
+static int __init mdctl_init(void)
+{
+	return misc_register(&mdctl_misc);
+}
+
+static void __exit mdctl_exit(void)
+{
+	misc_deregister(&mdctl_misc);
+}
+
 static int md_set_read_only(struct block_device *bdev, bool ro)
 {
 	struct mddev *mddev = bdev->bd_disk->private_data;
@@ -10815,12 +11438,19 @@ static int __init md_init(void)
 		goto err_mdp;
 	mdp_major = ret;
 
+	ret = mdctl_init();
+	if (ret < 0)
+		goto err_ctl;
+
 	register_reboot_notifier(&md_notifier);
 	raid_table_header = register_sysctl("dev/raid", raid_table);
 
 	md_geninit();
 	return 0;
 
+err_ctl:
+	unregister_blkdev(mdp_major, "mdp");
+	mdp_major = 0;
 err_mdp:
 	unregister_blkdev(MD_MAJOR, "md");
 err_md:
@@ -11098,6 +11728,7 @@ static __exit void md_exit(void)
 	struct mddev *mddev;
 	int delay = 1;
 
+	mdctl_exit();
 	unregister_blkdev(MD_MAJOR,"md");
 	unregister_blkdev(mdp_major, "mdp");
 	unregister_reboot_notifier(&md_notifier);
diff --git a/drivers/md/md.h b/drivers/md/md.h
index 27de34a8bdf9..0dfae7b86830 100644
--- a/drivers/md/md.h
+++ b/drivers/md/md.h
@@ -996,6 +996,15 @@ static inline void rdev_record_write_error(struct md_rdev *rdev)
 		set_bit(MD_RECOVERY_NEEDED, &rdev->mddev->recovery);
 }
 
+/*
+ * Return true if array is configured or its metadata is managed by user space,
+ * false otherwise.
+ */
+static inline bool mddev_configured(struct mddev *mddev)
+{
+	return mddev->raid_disks || mddev->external;
+}
+
 static inline int mddev_is_clustered(struct mddev *mddev)
 {
 	return mddev->cluster_info && mddev->bitmap_info.nodes > 1;
-- 
2.43.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-01 12:56 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 12:56 [RFC PATCH v2 0/4] md: add a control device for array management Abd-Alrhman Masalkhi
2026-10-01 12:56 ` [RFC PATCH v2 1/4] md: add uapi definitions for the md control device Abd-Alrhman Masalkhi
2026-10-01 12:56 ` [RFC PATCH v2 2/4] md: pass struct mdu_disk_info64 to md_add_new_disk() Abd-Alrhman Masalkhi
2026-10-01 12:56 ` [RFC PATCH v2 3/4] md: use struct mdu_array_info64 for SET_ARRAY_INFO Abd-Alrhman Masalkhi
2026-10-01 12:56 ` [RFC PATCH v2 4/4] md: add a control device for array management Abd-Alrhman Masalkhi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®