mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload
@ 2026-10-06 14:57 Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
                   ` (8 more replies)
  0 siblings, 9 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

Hi,

This series adds progress diagnostics to perf, and a workload that makes
false sharing visible to data type profiling.

The changes are:

  - move perf_config__set_variable() to util/config.c and serialize config
    parser and read-modify-write state, so non-builtin perf code can persist
    configuration changes safely;

  - add 'perf report --progress' for stdio users, showing the current phase,
    percentage, and counts while a session is processed;

  - wire up 'perf report --no-progress', the counterpart of the option
    above, for the TUI and GTK browsers, whose progress there is no other
    way to turn off;

  - add 'perf test -w false_sharing', a synthetic TCP-shaped workload with
    identity and packet counters sharing a cacheline, and include it in the
    data type profiling shell test.

Follow up work: the DO_ONCE() one-time init primitive added here mirrors
what the eight pre-existing raw pthread_once() users in util/ (annotate.c,
callchain.c, comm.c, dso.c, fncache.c, intel-tpebs.c, libbfd.c, pmus.c)
need, conversions that will also exercise the DEFINE_MUTEX() static
initializer; converting them is left for after this series.
Best regards,

- Arnaldo

What changed from v11:

  - The 'perf-stuck' patch is removed for the time being, to be submitted
    separately later on;

What changed from v10:

  - tools/perf/scripts/perf-stuck.sh: pin the watched process by its start
    time, so a recycled PID can't get samples or gdb; Sashiko, v10;

  - tools/perf/scripts/perf-stuck.sh: drop the control characters from the
    shown progress updates, no terminal escape injection; Sashiko, v10;

  - tools/perf/tests/shell/perf_stuck.sh: test the script, from option
    validation to the gdb DIE chain runs on stand-in workloads;

What changed from v9:

  - tools/perf/ui/stdio/progress.c: progress updates now work well with
    the pager, printed to /dev/tty so they update in place;

  - tools/perf/scripts/perf-stuck.sh: take the last progress update
    from a bounded tail window, dropping the \r's.  Sashiko, v9;

What changed from v8:

  - tools/perf/util/mutex.h: DEFINE_MUTEX() statics keep mutex_init()'s
    !NDEBUG errorcheck type;

  - patch 4/9 commit log: clarify the acquire/release pairing, what the
    dropped static key used to guarantee.

What changed from v7:

  - tools/perf/util/mutex.h: the DO_ONCE() lockless fast path pairs its
    __ATOMIC_ACQUIRE load with the __ATOMIC_RELEASE store.  sashiko, v7;

What changed from v6:

  - patch 1/5 split into four: the move, perf_etc_perfconfig() never
    returning NULL, the DEFINE_MUTEX() prep and the serialization.
    Namhyung Kim asked, v6;

  - tools/perf/util/{config.c,mutex.h}: perf's own mutex type, adding
    the DEFINE_MUTEX() initializer it lacked.  Namhyung Kim, v6;

  - tools/perf/util/mutex.h: add DO_ONCE() one-time init, from the
    kernel's once.h, moving the lazy inits in config.c to it;

  - tools/perf/builtin-report.c: drop the option negation and stdio
    hook comments Namhyung Kim asked to drop, reviewing v6;

  - tools/perf/scripts/perf-stuck.gdb: break the long printf() line.
    Namhyung Kim, reviewing v6;

  - the false_sharing cset now carries the data type profiling output
    with cacheline info.  Namhyung Kim asked for it, reviewing v6.

What changed from v5:

  - tools/perf/util/config.c: drop the config_file_name read in
    bad_config(), locking would buy a consistent NULL.  Sashiko, v5;

  - tools/perf/util/config.c: new config_set_mutex around the shared
    set's init/teardown, home init via pthread_once.  Sashiko, v5;

  - tools/perf/builtin-report.c, perf-report.txt: --no-progress is
    --progress's auto negation, last wins.  Namhyung Kim, reviewing v5.

What changed from v4:

  - tools/perf/builtin-report.c: mark --progress PARSE_OPT_NOAUTONEG,
    parse_long_opt() claimed --no-progress first.  Sashiko, v4;

  - tools/perf/util/config.c: format the path buffer inside the critical
    section.  Sashiko pointed out the window, reviewing v4;

  - tools/perf/tests/workloads/false_sharing.c: walk every bit of the
    affinity mask, not _SC_NPROCESSORS_CONF.  Sashiko, reviewing v4.

What changed from v3:

  - tools/perf/util/config.c: keep the buffer handed to the parser in
    static storage.  Sashiko, reviewing v3;

  - tools/perf/scripts/perf-stuck.gdb: perf-dso walks each dso candidate
    until one evaluates, for REFCNT_CHECKING builds.  Sashiko, v3;

  - tools/perf/tests/workloads/false_sharing.c: sum before cpu in struct
    fs_reader, the padding after cpu rounded it to 128.  Sashiko, v3;

  - tools/perf/builtin-report.c: --quiet wins over --progress, the phases
    stay uncounted, documented next to it.  Namhyung Kim, v3;

  - tools/perf/builtin-report.c, tools/perf/ui/progress.c: --no-progress
    installs the no-op ops.  Suggested by Namhyung Kim, reviewing v3;

  - tools/perf/scripts/perf-stuck.sh: check gdb is there before
    watching.  Namhyung Kim, reviewing v3.

What changed from v2:

  - tools/perf/util/config.c: make perf_etc_perfconfig() total, falling
    back to the unresolved path.  Sashiko, reviewing v2;

  - tools/perf/scripts/perf-stuck.gdb: don't deref map_symbol.sym without
    a NULL check.  Sashiko, reviewing v2;

  - tools/perf/scripts/perf-stuck.gdb: note the REFCNT_CHECKING proxy
    next to the structure walks;

  - tools/perf/scripts/perf-stuck.sh: count samples with no progress to
    look at, an empty log never fired -g;

  - tools/perf/scripts/perf-stuck.sh: `--` for pgrep and tail against
    names starting with a hyphen;

  - tools/perf/scripts/perf-stuck.sh: bound the gdb run with
    `timeout --signal=INT 30`.

What changed from v1:

  - avoid calling CPU_SET() with -1 when false_sharing runs with only one
    CPU available in its affinity mask.

  tools/perf/Documentation/perf-report.txt      |  19 ++
  tools/perf/builtin-config.c                   |  70 +----
  tools/perf/builtin-report.c                   |   9 +
  tools/perf/tests/builtin-test.c               |   1 +
  tools/perf/tests/shell/data_type_profiling.sh |   9 +-
  tools/perf/tests/tests.h                      |   1 +
  tools/perf/tests/workloads/Build              |   2 +
  tools/perf/tests/workloads/false_sharing.c    | 251 ++++++++++++++++++
  tools/perf/ui/Build                           |   1 +
  tools/perf/ui/progress.c                      |   6 +
  tools/perf/ui/progress.h                      |   4 +
  tools/perf/ui/stdio/progress.c                | 170 ++++++++++++
  tools/perf/util/config.c                      | 166 ++++++++++--
  tools/perf/util/config.h                      |   2 +
  tools/perf/util/mutex.h                       |  36 +++
  tools/perf/util/ordered-events.c              |  16 +-
  tools/perf/util/session.c                     |  12 +-
  17 files changed, 675 insertions(+), 100 deletions(-)
  create mode 100644 tools/perf/tests/workloads/false_sharing.c
  create mode 100644 tools/perf/ui/stdio/progress.c
base-commit: 705da5b15ab89ba9
v1-head: 45d7917f7e05e8a29828ed5f0bbdc94fc938f79f
v2-head: d4f84e4de8890194924ccd897a8e6773e7d4240b
v3-head: 485532296710225862daf8ffb19802ae327efe2a
v4-head: 38193635508e0a5f04a6ebf70db27534b68b2d5b
v5-head: e99bc48f6c72b9cc67b4e4143a0364acac5523bb
v6-head: 7dc99dbb6504f199c3c39487d91370d1bbb7a827
v7-head: b357a76666d08f74eb99c755cca41855a96f6169
v8-head: 257b5ea6115ad5126d2c881523ea6c660cf5d28c
v9-head: 70235104c6cf7d18d91f46de1a233abbeb65893a
v10-head: 1ecf850c2694d20ca9cd0014f2a42528ef4dbfec
v11-head: a3ddbaab18aa24fba520927f71d0ea66edfc5593
--

^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 2/8] perf config: Make perf_etc_perfconfig() never return NULL Arnaldo Carvalho de Melo
                   ` (7 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Move perf_config__set_variable() out of the 'perf config' builtin so
that opt-in features can persist their choice from outside it, e.g.
util/debuginfo.c writing core.debuginfod=false when the user disables
debuginfod for the rest of the session.  The set_config() body becomes
perf_config_set__write(), with the system_config choice as an argument,
as the builtin's use_system_config/use_user_config statics are not
available outside it.

perf_config_set__write() checked fopen() but none of the fprintf()s or
fclose(), so a write failure after truncating the file was reported as
success.  Harmless for the interactive 'perf config' this came from,
but this makes it an entry point a background feature can call with no
other feedback, so propagate those errors too.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/builtin-config.c | 70 +---------------------------------
 tools/perf/util/config.c    | 76 +++++++++++++++++++++++++++++++++++++
 tools/perf/util/config.h    |  2 +
 3 files changed, 79 insertions(+), 69 deletions(-)

diff --git a/tools/perf/builtin-config.c b/tools/perf/builtin-config.c
index cefd042e4f853466..3b074aca8d344539 100644
--- a/tools/perf/builtin-config.c
+++ b/tools/perf/builtin-config.c
@@ -41,37 +41,7 @@ static struct option config_options[] = {
 
 static int set_config(struct perf_config_set *set, const char *file_name)
 {
-	struct perf_config_section *section = NULL;
-	struct perf_config_item *item = NULL;
-	const char *first_line = "# this file is auto-generated.";
-	FILE *fp;
-
-	if (set == NULL)
-		return -1;
-
-	fp = fopen(file_name, "w");
-	if (!fp)
-		return -1;
-
-	fprintf(fp, "%s\n", first_line);
-
-	/* overwrite configvariables */
-	perf_config_items__for_each_entry(&set->sections, section) {
-		if (!use_system_config && section->from_system_config)
-			continue;
-		fprintf(fp, "[%s]\n", section->name);
-
-		perf_config_items__for_each_entry(&section->items, item) {
-			if (!use_system_config && item->from_system_config)
-				continue;
-			if (item->value)
-				fprintf(fp, "\t%s = %s\n",
-					item->name, item->value);
-		}
-	}
-	fclose(fp);
-
-	return 0;
+	return perf_config_set__write(set, file_name, use_system_config);
 }
 
 static int show_spec_config(struct perf_config_set *set, const char *var)
@@ -158,44 +128,6 @@ static int parse_config_arg(char *arg, char **var, char **value)
 	return 0;
 }
 
-int perf_config__set_variable(const char *var, const char *value)
-{
-	char path[PATH_MAX];
-	char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
-	const char *config_filename;
-	struct perf_config_set *set;
-	int ret = -1;
-
-	if (use_system_config)
-		config_exclusive_filename = perf_etc_perfconfig();
-	else if (use_user_config)
-		config_exclusive_filename = user_config;
-
-	if (!config_exclusive_filename)
-		config_filename = user_config;
-	else
-		config_filename = config_exclusive_filename;
-
-	set = perf_config_set__new();
-	if (!set)
-		goto out_err;
-
-	if (perf_config_set__collect(set, config_filename, var, value) < 0) {
-		pr_err("Failed to add '%s=%s'\n", var, value);
-		goto out_err;
-	}
-
-	if (set_config(set, config_filename) < 0) {
-		pr_err("Failed to set the configs on %s\n", config_filename);
-		goto out_err;
-	}
-
-	ret = 0;
-out_err:
-	perf_config_set__delete(set);
-	return ret;
-}
-
 int cmd_config(int argc, const char **argv)
 {
 	int i, ret = -1;
diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
index 8fe43b032e9af88a..85e50d0a25580da0 100644
--- a/tools/perf/util/config.c
+++ b/tools/perf/util/config.c
@@ -12,6 +12,7 @@
 #include "config.h"
 
 #include <errno.h>
+#include <limits.h>
 #include <stdbool.h>
 #include <stdio.h>
 #include <stdlib.h>
@@ -883,6 +884,81 @@ void perf_config__exit(void)
 	config_set = NULL;
 }
 
+int perf_config_set__write(struct perf_config_set *set,
+			   const char *file_name, bool system_config)
+{
+	struct perf_config_section *section = NULL;
+	struct perf_config_item *item = NULL;
+	int ret = 0;
+	FILE *fp;
+
+	fp = fopen(file_name, "w");
+	if (!fp)
+		return -1;
+
+	if (fprintf(fp, "# this file is auto-generated.\n") < 0)
+		ret = -1;
+
+	/* overwrite configvariables */
+	perf_config_sections__for_each_entry(&set->sections, section) {
+		if (!system_config && section->from_system_config)
+			continue;
+		if (fprintf(fp, "[%s]\n", section->name) < 0)
+			ret = -1;
+
+		perf_config_items__for_each_entry(&section->items, item) {
+			if (!system_config && item->from_system_config)
+				continue;
+			if (item->value &&
+			    fprintf(fp, "\t%s = %s\n", item->name, item->value) < 0)
+				ret = -1;
+		}
+	}
+	if (fclose(fp) != 0)
+		ret = -1;
+
+	return ret;
+}
+
+/*
+ * Set @var=@value in the config file perf is using: ~/.perfconfig or the
+ * file named by PERF_CONFIG.  Same rewrite 'perf config' does, comments
+ * are not preserved.
+ */
+int perf_config__set_variable(const char *var, const char *value)
+{
+	char path[PATH_MAX];
+	char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
+	const char *config_filename;
+	bool system_config;
+	struct perf_config_set *set;
+	int ret = -1;
+
+	config_filename = config_exclusive_filename ?: user_config;
+
+	/* Rewriting the system wide file keeps its entries, or it is truncated. */
+	system_config = strcmp(config_filename, perf_etc_perfconfig()) == 0;
+
+	set = perf_config_set__new();
+	if (!set)
+		goto out_err;
+
+	if (perf_config_set__collect(set, config_filename, var, value) < 0) {
+		pr_err("Failed to add '%s=%s'\n", var, value);
+		goto out_err;
+	}
+
+	if (perf_config_set__write(set, config_filename, system_config) < 0) {
+		pr_err("Failed to set the configs on %s\n", config_filename);
+		goto out_err;
+	}
+
+	ret = 0;
+out_err:
+	perf_config_set__delete(set);
+	return ret;
+}
+
 static void perf_config_item__delete(struct perf_config_item *item)
 {
 	zfree(&item->name);
diff --git a/tools/perf/util/config.h b/tools/perf/util/config.h
index 987b47cf54c350ba..9098f8a045850c97 100644
--- a/tools/perf/util/config.h
+++ b/tools/perf/util/config.h
@@ -33,6 +33,8 @@ int perf_config_scan(const char *name, const char *fmt, ...) __scanf(2, 3);
 const char *perf_config_get(const char *name);
 int perf_config_set(struct perf_config_set *set,
 		    config_fn_t fn, void *data);
+int perf_config_set__write(struct perf_config_set *set,
+			   const char *file_name, bool system_config);
 int perf_config_int(int *dest, const char *, const char *);
 int perf_config_u8(u8 *dest, const char *name, const char *value);
 int perf_config_u64(u64 *dest, const char *, const char *);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 2/8] perf config: Make perf_etc_perfconfig() never return NULL
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 3/8] perf mutex: Add DEFINE_MUTEX() static initializer Arnaldo Carvalho de Melo
                   ` (6 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

perf_etc_perfconfig() returns what system_path() gives, and that
allocates, so on failure every caller dereferenced NULL, the
pre-existing ones in perf_config_set__init() and in the daemon
included.  For the usual absolute ETC_PERFCONFIG system_path() just
returns a strdup of it, so use ETC_PERFCONFIG itself when that fails.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/util/config.c | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
index 85e50d0a25580da0..6e0ff8a9140ccf96 100644
--- a/tools/perf/util/config.c
+++ b/tools/perf/util/config.c
@@ -571,8 +571,14 @@ static int perf_config_from_file(config_fn_t fn, const char *filename, void *dat
 const char *perf_etc_perfconfig(void)
 {
 	static const char *system_wide;
+
 	if (!system_wide)
-		system_wide = system_path(ETC_PERFCONFIG);
+		/*
+		 * ETC_PERFCONFIG is absolute, so its unresolved path is
+		 * the same string, better than the callers crashing.
+		 */
+		system_wide = system_path(ETC_PERFCONFIG) ?: ETC_PERFCONFIG;
+
 	return system_wide;
 }
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 3/8] perf mutex: Add DEFINE_MUTEX() static initializer
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 2/8] perf config: Make perf_etc_perfconfig() never return NULL Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 4/8] perf mutex: Add DO_ONCE() for one-time initialization Arnaldo Carvalho de Melo
                   ` (5 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

For file scope mutexes, so that they don't need a constructor function
to initialize them at startup, mirroring PTHREAD_MUTEX_INITIALIZER
while keeping the clang -Wthread-safety annotations of struct mutex.

DEBUG=1 builds have mutex_init() set PTHREAD_MUTEX_ERRORCHECK, making
the CHECK_ERR() paths in mutex_lock()/mutex_unlock() report relocking a
held mutex and unlocking one that isn't held.  PTHREAD_MUTEX_INITIALIZER
gives a default type mutex, so statically initialized ones would silently
lose that: use PTHREAD_ERRORCHECK_MUTEX_INITIALIZER_NP, a glibc
extension, falling back to the default type in libcs that lack it.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/util/mutex.h | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/tools/perf/util/mutex.h b/tools/perf/util/mutex.h
index 70232d8d094f8bfc..fe04db21d19c5b0b 100644
--- a/tools/perf/util/mutex.h
+++ b/tools/perf/util/mutex.h
@@ -92,6 +92,18 @@ struct LOCKABLE mutex {
 	pthread_mutex_t lock;
 };
 
+/*
+ * Statically initialized mutex, for the file scope ones. Error checking
+ * in !NDEBUG builds, like mutex_init(); the glibc-only errorcheck
+ * initializer falls back to the default type elsewhere.
+ */
+#if !defined(NDEBUG) && defined(PTHREAD_ERRORCHECK_MUTEX_INITIALIZER_NP)
+#define __PERF_MUTEX_INITIALIZER PTHREAD_ERRORCHECK_MUTEX_INITIALIZER_NP
+#else
+#define __PERF_MUTEX_INITIALIZER PTHREAD_MUTEX_INITIALIZER
+#endif
+#define DEFINE_MUTEX(name) struct mutex name = { .lock = __PERF_MUTEX_INITIALIZER }
+
 /* A wrapper around the condition variable implementation. */
 struct cond {
 	pthread_cond_t cond;
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 4/8] perf mutex: Add DO_ONCE() for one-time initialization
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
                   ` (2 preceding siblings ...)
  2026-10-06 14:57 ` [PATCH 3/8] perf mutex: Add DEFINE_MUTEX() static initializer Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 5/8] perf config: Serialize config file access with a mutex Arnaldo Carvalho de Melo
                   ` (4 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

For lazy init that more than one thread can race into, without
pthread_once() so that it stays within perf's locking primitives and
clang's -Wthread-safety keeps an eye on the mutexes involved.

Mirrors the kernel's include/linux/once.h DO_ONCE(), dropping its
static key fast path, unnecessary at these cold init paths: a per call
site static bool plus a statically initialized mutex, double checked.
As there, code reachable from more than one call site must go through
a common helper.

The lockless fast path reads ___done with an acquire, paired with a
release store after fn(), so that fn()'s side effects are visible to
threads taking the fast path, which the dropped static key used to
guarantee.

The __atomic builtins are used rather than smp_load_acquire()/
smp_store_release() because ThreadSanitizer understands them,
avoiding reports on ___done and on what fn() initialized.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/util/mutex.h | 24 ++++++++++++++++++++++++
 1 file changed, 24 insertions(+)

diff --git a/tools/perf/util/mutex.h b/tools/perf/util/mutex.h
index fe04db21d19c5b0b..5554937675886ef3 100644
--- a/tools/perf/util/mutex.h
+++ b/tools/perf/util/mutex.h
@@ -104,6 +104,30 @@ struct LOCKABLE mutex {
 #endif
 #define DEFINE_MUTEX(name) struct mutex name = { .lock = __PERF_MUTEX_INITIALIZER }
 
+/*
+ * Runs fn() exactly once however many threads race to it, the rest wait
+ * for it to finish; per call site state, so multiple sites wanting the
+ * same once use a common helper, as in the kernel's DO_ONCE().
+ */
+#define DO_ONCE(fn, ...)						       \
+	({								       \
+		static bool ___done;					       \
+		static DEFINE_MUTEX(___once_lock);			       \
+		bool ___ret = false;					       \
+									       \
+		if (!__atomic_load_n(&___done, __ATOMIC_ACQUIRE)) {	       \
+			mutex_lock(&___once_lock);			       \
+			if (!___done) {					       \
+				fn(__VA_ARGS__);			       \
+				__atomic_store_n(&___done, true,	       \
+						 __ATOMIC_RELEASE);	       \
+				___ret = true;				       \
+			}						       \
+			mutex_unlock(&___once_lock);			       \
+		}							       \
+		___ret;							       \
+	})
+
 /* A wrapper around the condition variable implementation. */
 struct cond {
 	pthread_cond_t cond;
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 5/8] perf config: Serialize config file access with a mutex
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
                   ` (3 preceding siblings ...)
  2026-10-06 14:57 ` [PATCH 4/8] perf mutex: Add DO_ONCE() for one-time initialization Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 6/8] perf report: Add --progress option Arnaldo Carvalho de Melo
                   ` (3 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Config file access shares the static parser state and, since
perf_config__set_variable() became an entry point for background
features, can run on more than one thread: a feature writing the
config while perf top's display thread reads it.  Serialize parsing
and rewriting with config_mutex and the whole read-modify-write of
perf_config__set_variable() with config_update_mutex; config_set_mutex
comes before it, as building the set parses the config files.

perf has its own mutex type, util/mutex.h; the file scope mutexes
here use the DEFINE_MUTEX() static initializer added in the previous
patch.  The two lazy inits that system_path() and home_perfconfig()
do would leak all but one of the strings racing on two threads, so
they move to DO_ONCE(), added in the previous patch as well.
bad_config() runs on the dispatching thread, outside config_mutex, so
it stops reading config_file_name, which the parsing thread owns.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/util/config.c | 106 +++++++++++++++++++++++++++------------
 1 file changed, 73 insertions(+), 33 deletions(-)

diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
index 6e0ff8a9140ccf96..287401f159e66e8c 100644
--- a/tools/perf/util/config.c
+++ b/tools/perf/util/config.c
@@ -31,6 +31,7 @@
 #include "callchain.h"
 #include "debug.h"
 #include "header.h"
+#include "mutex.h"
 #include "path.h"
 #include "srcline.h"
 #include "unwind.h"
@@ -371,10 +372,8 @@ static int perf_parse_long(const char *value, long *ret)
 
 static void bad_config(const char *name)
 {
-	if (config_file_name)
-		pr_warning("bad config value for '%s' in %s, ignoring...\n", name, config_file_name);
-	else
-		pr_warning("bad config value for '%s', ignoring...\n", name);
+	/* config_file_name is owned by the parsing thread, under config_mutex. */
+	pr_warning("bad config value for '%s', ignoring...\n", name);
 }
 
 int perf_config_u64(u64 *dest, const char *name, const char *value)
@@ -550,11 +549,18 @@ int perf_default_config(const char *var, const char *value,
 	return 0;
 }
 
+/* Parsing and rewriting share the static parser state. */
+static DEFINE_MUTEX(config_mutex);
+
+/* Serializes whole perf_config__set_variable() updates. */
+static DEFINE_MUTEX(config_update_mutex);
+
 static int perf_config_from_file(config_fn_t fn, const char *filename, void *data)
 {
 	int ret;
 	FILE *f = fopen(filename, "r");
 
+	mutex_lock(&config_mutex);
 	ret = -1;
 	if (f) {
 		config_file = f;
@@ -565,21 +571,24 @@ static int perf_config_from_file(config_fn_t fn, const char *filename, void *dat
 		fclose(f);
 		config_file_name = NULL;
 	}
+	mutex_unlock(&config_mutex);
 	return ret;
 }
 
-const char *perf_etc_perfconfig(void)
-{
-	static const char *system_wide;
+/* system_path() allocates, so it is computed once. */
+static const char *etc_perfconfig;
 
-	if (!system_wide)
-		/*
-		 * ETC_PERFCONFIG is absolute, so its unresolved path is
-		 * the same string, better than the callers crashing.
-		 */
-		system_wide = system_path(ETC_PERFCONFIG) ?: ETC_PERFCONFIG;
+static void perf_etc_perfconfig__init(void)
+{
+	etc_perfconfig = system_path(ETC_PERFCONFIG);
+	if (!etc_perfconfig)
+		etc_perfconfig = ETC_PERFCONFIG;
+}
 
-	return system_wide;
+const char *perf_etc_perfconfig(void)
+{
+	DO_ONCE(perf_etc_perfconfig__init);
+	return etc_perfconfig;
 }
 
 static int perf_env_bool(const char *k, int def)
@@ -637,19 +646,18 @@ static char *home_perfconfig(void)
 	return NULL;
 }
 
-const char *perf_home_perfconfig(void)
-{
-	static const char *config;
-	static bool failed;
-
-	if (failed || config)
-		return config;
+/* home_perfconfig() allocates and warns, so it is computed once. */
+static const char *home_config;
 
-	config = home_perfconfig();
-	if (!config)
-		failed = true;
+static void perf_home_perfconfig__init(void)
+{
+	home_config = home_perfconfig();
+}
 
-	return config;
+const char *perf_home_perfconfig(void)
+{
+	DO_ONCE(perf_home_perfconfig__init);
+	return home_config;
 }
 
 static struct perf_config_section *find_section(struct list_head *sections,
@@ -790,8 +798,15 @@ static int collect_config(const char *var, const char *value,
 int perf_config_set__collect(struct perf_config_set *set, const char *file_name,
 			     const char *var, const char *value)
 {
+	int ret;
+
+	mutex_lock(&config_mutex);
 	config_file_name = file_name;
-	return collect_config(var, value, set);
+	ret = collect_config(var, value, set);
+	/* Don't leave the static parser state pointing at the caller's buffer. */
+	config_file_name = NULL;
+	mutex_unlock(&config_mutex);
+	return ret;
 }
 
 static int perf_config_set__init(struct perf_config_set *set)
@@ -838,6 +853,10 @@ struct perf_config_set *perf_config_set__load_file(const char *file)
 	return set;
 }
 
+/* Not config_mutex: building the set parses the config files, which takes it. */
+static DEFINE_MUTEX(config_set_mutex);
+
+/* Called with config_set_mutex held. */
 static int perf_config__init(void)
 {
 	if (config_set == NULL)
@@ -878,16 +897,26 @@ int perf_config_set(struct perf_config_set *set,
 
 int perf_config(config_fn_t fn, void *data)
 {
-	if (config_set == NULL && perf_config__init())
+	struct perf_config_set *set;
+
+	/* Not held across the dispatch: a callback can call perf_config() again. */
+	mutex_lock(&config_set_mutex);
+	if (perf_config__init()) {
+		mutex_unlock(&config_set_mutex);
 		return -1;
+	}
+	set = config_set;
+	mutex_unlock(&config_set_mutex);
 
-	return perf_config_set(config_set, fn, data);
+	return perf_config_set(set, fn, data);
 }
 
 void perf_config__exit(void)
 {
+	mutex_lock(&config_set_mutex);
 	perf_config_set__delete(config_set);
 	config_set = NULL;
+	mutex_unlock(&config_set_mutex);
 }
 
 int perf_config_set__write(struct perf_config_set *set,
@@ -898,9 +927,12 @@ int perf_config_set__write(struct perf_config_set *set,
 	int ret = 0;
 	FILE *fp;
 
+	mutex_lock(&config_mutex);
 	fp = fopen(file_name, "w");
-	if (!fp)
+	if (!fp) {
+		mutex_unlock(&config_mutex);
 		return -1;
+	}
 
 	if (fprintf(fp, "# this file is auto-generated.\n") < 0)
 		ret = -1;
@@ -922,6 +954,7 @@ int perf_config_set__write(struct perf_config_set *set,
 	}
 	if (fclose(fp) != 0)
 		ret = -1;
+	mutex_unlock(&config_mutex);
 
 	return ret;
 }
@@ -933,14 +966,20 @@ int perf_config_set__write(struct perf_config_set *set,
  */
 int perf_config__set_variable(const char *var, const char *value)
 {
-	char path[PATH_MAX];
-	char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
 	const char *config_filename;
 	bool system_config;
-	struct perf_config_set *set;
+	struct perf_config_set *set = NULL;
 	int ret = -1;
 
-	config_filename = config_exclusive_filename ?: user_config;
+	mutex_lock(&config_update_mutex);
+
+	/* Static: the parser publishes it as config_file_name. */
+	{
+		static char path[PATH_MAX];
+		char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
+
+		config_filename = config_exclusive_filename ?: user_config;
+	}
 
 	/* Rewriting the system wide file keeps its entries, or it is truncated. */
 	system_config = strcmp(config_filename, perf_etc_perfconfig()) == 0;
@@ -962,6 +1001,7 @@ int perf_config__set_variable(const char *var, const char *value)
 	ret = 0;
 out_err:
 	perf_config_set__delete(set);
+	mutex_unlock(&config_update_mutex);
 	return ret;
 }
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 6/8] perf report: Add --progress option
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
                   ` (4 preceding siblings ...)
  2026-10-06 14:57 ` [PATCH 5/8] perf config: Serialize config file access with a mutex Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 7/8] perf report: Add --no-progress option Arnaldo Carvalho de Melo
                   ` (2 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Processing a large session with stdio output gives no feedback about
which phase perf is in or how far along it is: ui_progress updates are
only shown by the TUI.  Add --progress, installing a stdio backend
(ui/stdio/progress.c) that prints the phase title, percentage and
counts:

  Processing events... [ 42.3%] 317M / 746M

Phases can be nested, so the backend tracks the ones started to
complete the right one on ui_progress__finish(); that requires
init()/finish() pairs, fixed in ordered-events.c and the pipe and
directory event processing.

With a pager both stdout and stderr lead to it, so isatty(stderr)
turns false and the updates end up printed one per line in the
pager's output: use /dev/tty in that case, so that the updates keep
updating in place, untouched by the pager.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/Documentation/perf-report.txt |  14 ++
 tools/perf/builtin-report.c              |   6 +
 tools/perf/ui/Build                      |   1 +
 tools/perf/ui/progress.h                 |   2 +
 tools/perf/ui/stdio/progress.c           | 170 +++++++++++++++++++++++
 tools/perf/util/ordered-events.c         |  16 ++-
 tools/perf/util/session.c                |  12 +-
 7 files changed, 214 insertions(+), 7 deletions(-)
 create mode 100644 tools/perf/ui/stdio/progress.c

diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index 3718ebd297ce0673..a7429a30ec28f903 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -29,6 +29,20 @@ OPTIONS
 --quiet::
 	Do not show any warnings or messages.  (Suppress -v)
 
+--progress::
+	Show progress for each of the processing phases, printing the
+	percentage done and the current/total counts: the first phase
+	counts the bytes of the perf.data file processed so far, the
+	merge and sort phases count hist entries, e.g.:
+
+	Processing events... [ 42.3%] 4G / 10G
+	Merging related events... [  7.1%] 1024 / 14387
+	Sorting events for output... [ 98.2%] 14132 / 14387
+
+	It is a no-op when using the TUI or GTK browsers, that already
+	present progress information, or when --quiet is used, that asks
+	for no messages at all.
+
 -n::
 --show-nr-samples::
 	Show the number of samples for each symbol
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 57225bc87731d074..81e13fb8194976c9 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -87,6 +87,7 @@ struct report {
 	bool			use_gtk;
 #endif
 	bool			use_stdio;
+	bool			progress;
 	bool			show_full_info;
 	bool			show_threads;
 	bool			inverted_callchain;
@@ -1384,6 +1385,8 @@ int cmd_report(int argc, const char **argv)
 		    "Use the stdio interface"),
 	OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
 			"Show or hide weight columns in annotation. Default show if non-zero."),
+	OPT_BOOLEAN(0, "progress", &report.progress,
+		    "Show progress while processing the perf.data file"),
 	OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
 	OPT_BOOLEAN(0, "header-only", &report.header_only,
 		    "Show only data header."),
@@ -1790,6 +1793,9 @@ int cmd_report(int argc, const char **argv)
 	else
 		use_browser = 0;
 
+	if (report.progress && !quiet && use_browser == 0)
+		stdio_progress__init();
+
 	if (report.data_type && use_browser == 1) {
 		symbol_conf.annotate_data_member = true;
 		symbol_conf.annotate_data_sample = true;
diff --git a/tools/perf/ui/Build b/tools/perf/ui/Build
index 6005f813c9e3990c..a7b1740d51c80f23 100644
--- a/tools/perf/ui/Build
+++ b/tools/perf/ui/Build
@@ -4,6 +4,7 @@ perf-ui-y += progress.o
 perf-ui-y += util.o
 perf-ui-y += hist.o
 perf-ui-y += stdio/hist.o
+perf-ui-y += stdio/progress.o
 
 CFLAGS_setup.o += -DLIBDIR="BUILD_STR($(LIBDIR))"
 
diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
index 4f52c37b2f099a82..03f1a8bb260ba076 100644
--- a/tools/perf/ui/progress.h
+++ b/tools/perf/ui/progress.h
@@ -23,6 +23,8 @@ void __ui_progress__init(struct ui_progress *p, u64 total,
 
 void ui_progress__update(struct ui_progress *p, u64 adv);
 
+void stdio_progress__init(void);
+
 struct ui_progress_ops {
 	void (*init)(struct ui_progress *p);
 	void (*update)(struct ui_progress *p);
diff --git a/tools/perf/ui/stdio/progress.c b/tools/perf/ui/stdio/progress.c
new file mode 100644
index 0000000000000000..57f615c455b0ca7d
--- /dev/null
+++ b/tools/perf/ui/stdio/progress.c
@@ -0,0 +1,170 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Progress feedback for the stdio (non-TUI/GTK) case, enabled with
+ * 'perf report --progress'.
+ */
+#include <inttypes.h>
+#include <stdio.h>
+#include <unistd.h>
+#include <linux/kernel.h>
+#include <subcmd/pager.h>
+#include "../../util/debug.h"
+#include "../../util/units.h"
+#include "../progress.h"
+
+/*
+ * Phases can be nested, so keep track of the ones started so far to be
+ * able to complete the right one on ui_progress__finish(), which gets
+ * no arguments.
+ */
+#define STDIO_PROGRESS__MAX_DEPTH 8
+
+struct stdio_progress_phase {
+	struct ui_progress	*p;
+	u64			last_printed;
+	size_t			last_len;
+};
+
+static struct stdio_progress_phase stdio_progress__stack[STDIO_PROGRESS__MAX_DEPTH];
+static int stdio_progress__depth;
+static FILE *stdio_progress__out;
+static bool stdio_progress__is_tty;
+/* Phases that didn't fit on the stack are not shown. */
+static int stdio_progress__dropped;
+
+static void stdio_progress__print_phase(struct stdio_progress_phase *phase,
+					u64 curr)
+{
+	struct ui_progress *p = phase->p;
+	char buf_cur[20], buf_tot[20], buf[128];
+	double percent = p->total ? 100.0 * (double)curr / (double)p->total : 0.0;
+	size_t len;
+
+	/*
+	 * Only the completion line shows 100.0%: a 99.99% progress would round
+	 * up to it and look like a duplicate at finish time.
+	 */
+	if (curr < p->total && percent > 99.9)
+		percent = 99.9;
+
+	if (p->size) {
+		unit_number__scnprintf(buf_cur, sizeof(buf_cur), curr);
+		unit_number__scnprintf(buf_tot, sizeof(buf_tot), p->total);
+		len = scnprintf(buf, sizeof(buf), "%s [%5.1f%%] %s / %s",
+				p->title, percent, buf_cur, buf_tot);
+	} else {
+		len = scnprintf(buf, sizeof(buf), "%s [%5.1f%%] %" PRIu64 " / %" PRIu64,
+				p->title, percent, curr, p->total);
+	}
+
+	if (!stdio_progress__is_tty) {
+		fprintf(stdio_progress__out, "%s\n", buf);
+		goto out;
+	}
+
+	/* Pad to the length of the previous line to erase its leftovers. */
+	fprintf(stdio_progress__out, "\r%s%*s", buf,
+		(int)(len < phase->last_len ? phase->last_len - len : 0), "");
+	phase->last_len = len;
+out:
+	phase->last_printed = curr;
+	fflush(stdio_progress__out);
+}
+
+static void __stdio_progress__init(struct ui_progress *p)
+{
+	/* The default step is meant for the TUI bar, use 1% steps for stdio. */
+	p->next = p->step = p->total / 100 ?: 1;
+
+	if (stdio_progress__depth == STDIO_PROGRESS__MAX_DEPTH) {
+		/*
+		 * Out of room: don't start this phase, its finish() is
+		 * swallowed and its updates ignored below.
+		 */
+		pr_warning("progress phases nested deeper than %d, not showing progress for %s\n",
+			   STDIO_PROGRESS__MAX_DEPTH, p->title);
+		stdio_progress__dropped++;
+		return;
+	}
+
+	/* Start a nested phase in a line of its own. */
+	if (stdio_progress__depth && stdio_progress__is_tty)
+		fputc('\n', stdio_progress__out);
+
+	stdio_progress__stack[stdio_progress__depth++] =
+		(struct stdio_progress_phase) {
+			.p = p,
+			.last_printed = 0,
+			.last_len = 0,
+		};
+
+	stdio_progress__print_phase(&stdio_progress__stack[stdio_progress__depth - 1],
+				    p->curr);
+}
+
+static void stdio_progress__update(struct ui_progress *p)
+{
+	/*
+	 * An update that doesn't match the innermost phase means something
+	 * is out of sync: print nothing rather than another phase's
+	 * numbers, or read past the stack.
+	 */
+	if (!stdio_progress__depth ||
+	    stdio_progress__stack[stdio_progress__depth - 1].p != p)
+		return;
+
+	stdio_progress__print_phase(&stdio_progress__stack[stdio_progress__depth - 1],
+				    p->curr);
+}
+
+static void stdio_progress__finish(void)
+{
+	struct stdio_progress_phase *phase;
+
+	/*
+	 * Being the innermost phase, its finish() comes first: swallow it, or
+	 * it would complete the phase that encloses it.
+	 */
+	if (stdio_progress__dropped) {
+		stdio_progress__dropped--;
+		return;
+	}
+
+	if (!stdio_progress__depth)
+		return;
+
+	phase = &stdio_progress__stack[--stdio_progress__depth];
+
+	/*
+	 * The last line may have stopped short of the total, close this phase
+	 * showing it as complete unless that was already printed.
+	 */
+	if (phase->last_printed != phase->p->total)
+		stdio_progress__print_phase(phase, phase->p->total);
+
+	phase->last_printed	= 0;
+	phase->last_len		= 0;
+
+	if (stdio_progress__is_tty)
+		fputc('\n', stdio_progress__out);
+
+	fflush(stdio_progress__out);
+}
+
+static struct ui_progress_ops stdio_progress__ops = {
+	.init	= __stdio_progress__init,
+	.update	= stdio_progress__update,
+	.finish	= stdio_progress__finish,
+};
+
+void stdio_progress__init(void)
+{
+	if (pager_in_use()) {
+		/* stderr leads to the pager, the updates go to the terminal. */
+		stdio_progress__out = fopen("/dev/tty", "w");
+	}
+	if (!stdio_progress__out)
+		stdio_progress__out = stderr;
+	stdio_progress__is_tty = isatty(fileno(stdio_progress__out)) == 1;
+	ui_progress__ops = &stdio_progress__ops;
+}
diff --git a/tools/perf/util/ordered-events.c b/tools/perf/util/ordered-events.c
index a5857f9f5af2d3de..54c85663e733be4d 100644
--- a/tools/perf/util/ordered-events.c
+++ b/tools/perf/util/ordered-events.c
@@ -237,14 +237,16 @@ static int do_flush(struct ordered_events *oe, bool show_progress)
 		ui_progress__init(&prog, oe->nr_events, "Processing time ordered events...");
 
 	list_for_each_entry_safe(iter, tmp, head, list) {
-		if (session_done())
-			return 0;
+		if (session_done()) {
+			ret = 0;
+			goto out_progress;
+		}
 
 		if (iter->timestamp > limit)
 			break;
 		ret = oe->deliver(oe, iter);
 		if (ret < 0)
-			return ret;
+			goto out_progress;
 
 		ordered_events__delete(oe, iter);
 		oe->last_flush = iter->timestamp;
@@ -258,10 +260,16 @@ static int do_flush(struct ordered_events *oe, bool show_progress)
 	else if (last_ts <= limit)
 		oe->last = list_entry(head->prev, struct ordered_event, list);
 
+	ret = 0;
+out_progress:
+	/*
+	 * Always pair ui_progress__init() with ui_progress__finish(), the
+	 * stdio backend tracks the phases on a stack.
+	 */
 	if (show_progress)
 		ui_progress__finish();
 
-	return 0;
+	return ret;
 }
 
 static int __ordered_events__flush(struct ordered_events *oe, enum oe_flush how,
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index 7fea9e72726c936c..c4b4c7fb3b864589 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -3253,8 +3253,12 @@ static int __perf_session__process_pipe_events(struct perf_session *session)
 	cur_size = sizeof(union perf_event);
 
 	buf = malloc(cur_size);
-	if (!buf)
-		return -errno;
+	if (!buf) {
+		err = -errno;
+		if (update_prog)
+			ui_progress__finish();
+		return err;
+	}
 	ordered_events__set_copy_on_queue(oe, true);
 more:
 	event = buf;
@@ -3753,8 +3757,10 @@ static int __perf_session__process_dir_events(struct perf_session *session)
 	}
 
 	rd = calloc(nr_readers, sizeof(struct reader));
-	if (!rd)
+	if (!rd) {
+		ui_progress__finish();
 		return -ENOMEM;
+	}
 
 	rd[0] = (struct reader) {
 		.fd		 = perf_data__fd(session->data),
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 7/8] perf report: Add --no-progress option
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
                   ` (5 preceding siblings ...)
  2026-10-06 14:57 ` [PATCH 6/8] perf report: Add --progress option Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
  2026-10-06 14:57 ` [PATCH 8/8] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
  2026-10-07  0:07 ` [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Namhyung Kim
  8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Now that --progress is being added for the stdio case, wire up its
counterpart for the browsers: the TUI and GTK ones present progress of
their own and there is no way to turn it off.  Install the no-op
ui_progress ops, the ones already used until a backend sets theirs,
after setup_browser() installed the ones of the browser in use.  The
phases are still counted, nothing is shown for them, and no second
option is needed for it: parse-options provides --no-progress as the
negation of --progress, report.progress_set saying that it was asked
for, report.progress being false both when nothing was asked for and
when --no-progress was.

Suggested-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/Documentation/perf-report.txt | 5 +++++
 tools/perf/builtin-report.c              | 9 ++++++---
 tools/perf/ui/progress.c                 | 6 ++++++
 tools/perf/ui/progress.h                 | 2 ++
 4 files changed, 19 insertions(+), 3 deletions(-)

diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index a7429a30ec28f903..e145124d6c9a1097 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -43,6 +43,11 @@ OPTIONS
 	present progress information, or when --quiet is used, that asks
 	for no messages at all.
 
+--no-progress::
+	Do not show progress while processing the perf.data file.  It
+	also turns off the progress the TUI and GTK browsers present,
+	which is their own.
+
 -n::
 --show-nr-samples::
 	Show the number of samples for each symbol
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 81e13fb8194976c9..29085fdc91e58884 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -88,6 +88,7 @@ struct report {
 #endif
 	bool			use_stdio;
 	bool			progress;
+	bool			progress_set;
 	bool			show_full_info;
 	bool			show_threads;
 	bool			inverted_callchain;
@@ -1385,8 +1386,8 @@ int cmd_report(int argc, const char **argv)
 		    "Use the stdio interface"),
 	OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
 			"Show or hide weight columns in annotation. Default show if non-zero."),
-	OPT_BOOLEAN(0, "progress", &report.progress,
-		    "Show progress while processing the perf.data file"),
+	OPT_BOOLEAN_SET(0, "progress", &report.progress, &report.progress_set,
+			"Show progress while processing the perf.data file"),
 	OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
 	OPT_BOOLEAN(0, "header-only", &report.header_only,
 		    "Show only data header."),
@@ -1793,7 +1794,9 @@ int cmd_report(int argc, const char **argv)
 	else
 		use_browser = 0;
 
-	if (report.progress && !quiet && use_browser == 0)
+	if (report.progress_set && !report.progress)
+		ui_progress__noop_init();
+	else if (report.progress && !quiet && use_browser == 0)
 		stdio_progress__init();
 
 	if (report.data_type && use_browser == 1) {
diff --git a/tools/perf/ui/progress.c b/tools/perf/ui/progress.c
index 99d60223c74b2957..362680989ace606a 100644
--- a/tools/perf/ui/progress.c
+++ b/tools/perf/ui/progress.c
@@ -13,6 +13,12 @@ static struct ui_progress_ops null_progress__ops =
 
 struct ui_progress_ops *ui_progress__ops = &null_progress__ops;
 
+/* Everything counts but nothing is shown, the way it starts out. */
+void ui_progress__noop_init(void)
+{
+	ui_progress__ops = &null_progress__ops;
+}
+
 void ui_progress__update(struct ui_progress *p, u64 adv)
 {
 	u64 last = p->curr;
diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
index 03f1a8bb260ba076..e8c4f9f768aaf12b 100644
--- a/tools/perf/ui/progress.h
+++ b/tools/perf/ui/progress.h
@@ -25,6 +25,8 @@ void ui_progress__update(struct ui_progress *p, u64 adv);
 
 void stdio_progress__init(void);
 
+void ui_progress__noop_init(void);
+
 struct ui_progress_ops {
 	void (*init)(struct ui_progress *p);
 	void (*update)(struct ui_progress *p);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 8/8] perf test: Add false_sharing workload exhibiting cross-CPU false sharing
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
                   ` (6 preceding siblings ...)
  2026-10-06 14:57 ` [PATCH 7/8] perf report: Add --no-progress option Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
  2026-10-07  0:07 ` [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Namhyung Kim
  8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Add a 'perf test -w false_sharing' workload that hammers one shared
struct from several CPUs, shaped as a TCP connection: a read-mostly
identity (five-tuple) shares a cacheline with per-packet rx counters
(the false-sharing line), a second line has packet-path private tx and
congestion control counters, and a third the connection config.

The packet path runs in the main thread and up to four lookup threads
sum the five-tuple and pull the config, reading one volatile shared
instance directly so the accesses are PC-relative and resolvable by the
data type profiler.

Data type profiling of the workload with cacheline info:

  $ perf mem record -- perf test -r5 -w false_sharing
  [ perf record: Woken up 19 times to write data ]
  [ perf record: Captured and wrote 5.612 MB perf.data (72454 samples) ]
  $ perf report -s type,typecln -H --group --stdio
  # Total Lost Samples: 0
  #
  # Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
  # Event count (approx.): 1261814616
  #
  #            Overhead  Data Type / Data Type Cacheline
  # ...................  ...............................
  #
      88.99%  35.88%     struct net_conn
         56.43%  17.50%     struct net_conn: cache-line 0
         32.55%   0.00%     struct net_conn: cache-line 2
          0.01%  18.38%     struct net_conn: cache-line 1
      10.08%  63.03%     (unknown)
         10.08%  63.03%     (unknown): cache-line 0
       0.82%   0.13%     int
          0.82%   0.13%     int: cache-line 0
       0.01%   0.02%     struct folio
          0.01%   0.02%     struct folio: cache-line 0
       0.01%   0.03%     Elf64_Addr
          0.01%   0.03%     Elf64_Addr: cache-line 0
       0.01%   0.03%     struct sched_entity
          0.01%   0.03%     struct sched_entity: cache-line 1
          0.00%   0.00%     struct sched_entity: cache-line 4
          0.00%   0.00%     struct sched_entity: cache-line 2
          0.00%   0.00%     struct sched_entity: cache-line 3
          0.00%   0.00%     struct sched_entity: cache-line 0
       0.01%   0.00%     struct task_group
          0.01%   0.00%     struct task_group: cache-line 4
          0.00%   0.00%     struct task_group: cache-line 5
       0.00%   0.00%     struct css_rstat_cpu
          0.00%   0.00%     struct css_rstat_cpu: cache-line 0
  $ perf report -s type,typecln,typeoff -H --group --stdio
  # Total Lost Samples: 0
  #
  # Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
  # Event count (approx.): 1261814616
  #
  #               Overhead  Data Type / Data Type Cacheline / Data Type Offset
  # ......................  ..................................................
  #
      88.99%  35.88%        struct net_conn
         56.43%  17.50%        struct net_conn: cache-line 0
             9.81%   0.00%        struct net_conn +0xa (dport)
             9.54%   0.00%        struct net_conn +0x8 (sport)
             9.53%   0.00%        struct net_conn +0xc (state)
             9.48%   0.00%        struct net_conn +0xd (protocol)
             9.27%   0.00%        struct net_conn +0x4 (daddr)
             8.79%   0.00%        struct net_conn +0 (saddr)
             0.00%  12.08%        struct net_conn +0x10 (bytes_rx)
             0.00%   1.71%        struct net_conn +0x18 (packets_rx)
             0.00%   2.80%        struct net_conn +0x20 (rx_queue)
             0.00%   0.91%        struct net_conn +0x3c (last_ack)
         32.55%   0.00%        struct net_conn: cache-line 2
             5.70%   0.00%        struct net_conn +0x83 (rcv_wscale)
             5.46%   0.00%        struct net_conn +0x84 (keepalive_int)
             5.43%   0.00%        struct net_conn +0x82 (snd_wscale)
             5.39%   0.00%        struct net_conn +0x80 (mss)
             5.34%   0.00%        struct net_conn +0x88 (mark)
             5.23%   0.00%        struct net_conn +0x8c (priority)
          0.01%  18.38%        struct net_conn: cache-line 1
             0.01%   0.05%        struct net_conn +0x5c (retrans)
             0.00%   4.83%        struct net_conn +0x40 (bytes_tx)
             0.00%  10.63%        struct net_conn +0x48 (packets_tx)
             0.00%   1.54%        struct net_conn +0x58 (rtt_us)
             0.00%   1.07%        struct net_conn +0x50 (cwnd)
             0.00%   0.25%        struct net_conn +0x54 (ssthresh)
      10.08%  63.03%        (unknown)
         10.08%  63.03%        (unknown): cache-line 0
            10.08%  63.03%        (unknown)
       0.82%   0.13%        int
          0.82%   0.13%        int: cache-line 0
             0.82%   0.13%        int +0 (no field)
       0.01%   0.02%        struct folio
          0.01%   0.02%        struct folio: cache-line 0
             0.01%   0.00%        struct folio +0 (flags.f)
             0.00%   0.01%        struct folio +0x34 (_refcount.counter)
             0.00%   0.00%        struct folio +0x18 (mapping)

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/tests/builtin-test.c               |   1 +
 tools/perf/tests/shell/data_type_profiling.sh |   9 +-
 tools/perf/tests/tests.h                      |   1 +
 tools/perf/tests/workloads/Build              |   2 +
 tools/perf/tests/workloads/false_sharing.c    | 251 ++++++++++++++++++
 5 files changed, 262 insertions(+), 2 deletions(-)
 create mode 100644 tools/perf/tests/workloads/false_sharing.c

diff --git a/tools/perf/tests/builtin-test.c b/tools/perf/tests/builtin-test.c
index d2f594921e25bda9..98134b1c74cf80f4 100644
--- a/tools/perf/tests/builtin-test.c
+++ b/tools/perf/tests/builtin-test.c
@@ -175,6 +175,7 @@ static struct test_workload *workloads[] = {
 	&workload__context_switch_loop,
 	&workload__deterministic,
 	&workload__callchain,
+	&workload__false_sharing,
 
 #ifdef HAVE_RUST_SUPPORT
 	&workload__code_with_type,
diff --git a/tools/perf/tests/shell/data_type_profiling.sh b/tools/perf/tests/shell/data_type_profiling.sh
index a916c410274aa888..57203e0e5f857540 100755
--- a/tools/perf/tests/shell/data_type_profiling.sh
+++ b/tools/perf/tests/shell/data_type_profiling.sh
@@ -8,8 +8,8 @@ set -e
 # data type profiling manifestation
 
 # Values in testtypes and testprogs should match
-testtypes=("# data-type: struct Buf" "# data-type: struct buf")
-testprogs=("perf test -w code_with_type" "perf test -w datasym")
+testtypes=("# data-type: struct Buf" "# data-type: struct buf" "# data-type: struct net_conn")
+testprogs=("perf test -w code_with_type" "perf test -w datasym" "perf test -w false_sharing")
 
 err=0
 perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX)
@@ -59,6 +59,9 @@ test_basic_annotate() {
 
     "xC")
     index=1 ;;
+
+    "xFS")
+    index=2 ;;
   esac
 
   # Under 'set -e' a bare failing command aborts the script through the EXIT
@@ -114,6 +117,8 @@ test_basic_annotate Basic Rust
 test_basic_annotate Pipe Rust
 test_basic_annotate Basic C
 test_basic_annotate Pipe C
+test_basic_annotate Basic FS
+test_basic_annotate Pipe FS
 
 cleanup
 exit $err
diff --git a/tools/perf/tests/tests.h b/tools/perf/tests/tests.h
index 9c96f33483d14356..f379620069d34f3a 100644
--- a/tools/perf/tests/tests.h
+++ b/tools/perf/tests/tests.h
@@ -250,6 +250,7 @@ DECLARE_WORKLOAD(jitdump);
 DECLARE_WORKLOAD(context_switch_loop);
 DECLARE_WORKLOAD(deterministic);
 DECLARE_WORKLOAD(callchain);
+DECLARE_WORKLOAD(false_sharing);
 
 #ifdef HAVE_RUST_SUPPORT
 DECLARE_WORKLOAD(code_with_type);
diff --git a/tools/perf/tests/workloads/Build b/tools/perf/tests/workloads/Build
index 048e371eb63e3164..18fceddccb56c013 100644
--- a/tools/perf/tests/workloads/Build
+++ b/tools/perf/tests/workloads/Build
@@ -14,6 +14,7 @@ perf-test-y += jitdump.o
 perf-test-y += context_switch_loop.o
 perf-test-y += deterministic.o
 perf-test-y += callchain.o
+perf-test-y += false_sharing.o
 
 ifeq ($(CONFIG_RUST_SUPPORT),y)
     perf-test-y += code_with_type.o
@@ -29,3 +30,4 @@ CFLAGS_inlineloop.o       = -g -O2
 CFLAGS_deterministic.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
 CFLAGS_named_threads.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
 CFLAGS_callchain.o        = -g -O0 -fno-inline -U_FORTIFY_SOURCE
+CFLAGS_false_sharing.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
diff --git a/tools/perf/tests/workloads/false_sharing.c b/tools/perf/tests/workloads/false_sharing.c
new file mode 100644
index 0000000000000000..e949bb36646a97a5
--- /dev/null
+++ b/tools/perf/tests/workloads/false_sharing.c
@@ -0,0 +1,251 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * False-sharing demo for data type profiling, shaped as a TCP
+ * connection: a read-mostly identity shares a cacheline with per-packet
+ * rx counters (the false-sharing line), a second line has packet-path
+ * private tx and congestion control counters, and a third the connection
+ * config.
+ *
+ * 'perf mem record' of this workload followed by 'perf report -s type'
+ * (see tests/shell/data_type_profiling.sh) resolves the accesses to
+ * struct net_conn members, showing the rx counters and the read-mostly
+ * identity sharing cacheline 0.
+ */
+#include <pthread.h>
+#include <sched.h>
+#include <stdint.h>
+#include <stdlib.h>
+#include <stdio.h>
+#include <signal.h>
+#include <unistd.h>
+#include <linux/compiler.h>
+#include "../tests.h"
+
+struct net_conn {
+	/* cacheline 0: identity (read-mostly) + rx counters (per packet) */
+	uint32_t	saddr;		/*   0 */
+	uint32_t	daddr;		/*   4 */
+	uint16_t	sport;		/*   8 */
+	uint16_t	dport;		/*  10 */
+	uint8_t		state;		/*  12: 1 == ESTABLISHED */
+	uint8_t		protocol;	/*  13: 6 == TCP */
+	uint16_t	__pad0;		/*  14 */
+	uint64_t	bytes_rx;	/*  16: every packet */
+	uint64_t	packets_rx;	/*  24: every packet */
+	uint32_t	rx_queue;	/*  32: backlog depth, fluctuates */
+	uint8_t		__pad1[24];	/*  36..59 */
+	uint32_t	last_ack;	/*  60: written per ACK */
+	/* cacheline 1: tx + congestion control (packet-path private) */
+	uint64_t	bytes_tx;	/*  64: every packet */
+	uint64_t	packets_tx;	/*  72: every packet */
+	uint32_t	cwnd;		/*  80: on every ACK */
+	uint32_t	ssthresh;	/*  84: on loss */
+	uint32_t	rtt_us;		/*  88: on every ACK */
+	uint32_t	retrans;	/*  92: on timeout */
+	uint32_t	__pad2[8];	/*  96..127 */
+	/* cacheline 2: config, set at setup, read by everybody */
+	uint16_t	mss;		/* 128 */
+	uint8_t		snd_wscale;	/* 130 */
+	uint8_t		rcv_wscale;	/* 131 */
+	uint32_t	keepalive_int;	/* 132 */
+	uint32_t	mark;		/* 136: firewall mark */
+	uint32_t	priority;	/* 140: traffic class */
+	uint32_t	__pad3[12];	/* 144..191 */
+} __attribute__((aligned(64)));
+
+/* Volatile so every iteration really loads and stores. */
+static volatile struct net_conn conn;
+/* Keeps the reader checksums alive after the threads join. */
+static volatile unsigned long fs_sink;
+
+static volatile sig_atomic_t done;
+
+/*
+ * One cacheline each: sum before cpu, or the implicit padding after cpu
+ * pushes the struct past 64 bytes, and aligning it to a cacheline then
+ * rounds it up to 128.
+ */
+struct fs_reader {
+	pthread_t	thread;
+	unsigned long	sum;
+	int		cpu;
+	char		__pad[64 - sizeof(pthread_t) - sizeof(unsigned long) - sizeof(int)];
+} __attribute__((aligned(64)));
+
+static void sighandler(int sig __maybe_unused)
+{
+	done = 1;
+}
+
+static void pin_to_cpu(int cpu)
+{
+	cpu_set_t set;
+
+	/* There may be no second CPU in a restricted cpuset. */
+	if (cpu < 0)
+		return;
+
+	CPU_ZERO(&set);
+	CPU_SET(cpu, &set);
+	/* Best effort: in a restricted cpuset this fails and the thread runs unpinned. */
+	pthread_setaffinity_np(pthread_self(), sizeof(set), &set);
+}
+
+/*
+ * Connection lookup, as a load balancer or 'ss' scrape would do it: the
+ * reads go straight to the global, this file is built -O0 and a local
+ * pointer would be reloaded from a stack slot, a form the data type
+ * resolver does not track.
+ */
+static void *reader_fn(void *arg)
+{
+	struct fs_reader *r = arg;
+	unsigned long sum = 0;
+
+	pthread_setname_np(pthread_self(), "fs-reader");
+	pin_to_cpu(r->cpu);
+
+	while (!done) {
+		sum += conn.saddr + conn.daddr + conn.sport + conn.dport +
+		       conn.state + conn.protocol;
+		sum += conn.mss + conn.snd_wscale + conn.rcv_wscale +
+		       conn.keepalive_int + conn.mark + conn.priority;
+	}
+	r->sum = sum;
+	return NULL;
+}
+
+static int false_sharing(int argc, const char **argv)
+{
+	double sec = 2.0;
+	int nreaders = 0, nr_allowed = 0, err = 1;
+	int *allowed = NULL, nallowed = 0;
+	cpu_set_t set;
+	int nr_mask_bits = sizeof(set) * 8 < CPU_SETSIZE ? sizeof(set) * 8 : CPU_SETSIZE;
+	struct fs_reader *readers = NULL;
+	int i, writer_cpu;
+	unsigned long n = 0;
+
+	pthread_setname_np(pthread_self(), "fs-writer");
+	if (argc > 0)
+		sec = atof(argv[0]);
+	if (!(sec > 0.0)) {
+		fprintf(stderr, "Error: seconds (%f) must be > 0\n", sec);
+		return 1;
+	}
+	if (argc > 1)
+		nreaders = atoi(argv[1]);
+
+	/*
+	 * A connection that just got established: identity and config fixed
+	 * from here on, counters at zero.
+	 */
+	conn.saddr = 0x0a000001;	/* 10.0.0.1 */
+	conn.daddr = 0x0a000002;	/* 10.0.0.2 */
+	conn.sport = 54321;
+	conn.dport = 443;
+	conn.state = 1;			/* ESTABLISHED */
+	conn.protocol = 6;		/* TCP */
+	conn.mss = 1448;
+	conn.snd_wscale = 7;
+	conn.rcv_wscale = 7;
+	conn.keepalive_int = 7200;
+	conn.cwnd = 10;
+	conn.ssthresh = 65535;
+	conn.rtt_us = 50;
+
+	/*
+	 * Pin against the allowed set, restricted cpusets still spread the threads.
+	 * The whole mask is looked at, not the CPU count: the count can be lower
+	 * than the highest ID in it, as when a cpuset allows only high numbered
+	 * CPUs, and then no allowed CPU would be found at all.
+	 */
+	if (sched_getaffinity(0, sizeof(set), &set) == 0) {
+		for (i = 0; i < nr_mask_bits; i++) {
+			if (!CPU_ISSET(i, &set))
+				continue;
+			nr_allowed++;
+		}
+		allowed = malloc(nr_allowed * sizeof(int));
+		if (allowed == NULL) {
+			fprintf(stderr, "Error: malloc failed for CPU list\n");
+			return 1;
+		}
+		for (i = 0; i < nr_mask_bits; i++) {
+			if (CPU_ISSET(i, &set))
+				allowed[nallowed++] = i;
+		}
+	}
+	if (nreaders <= 0) {
+		/* By default leave one CPU for the packet path, up to 4 readers. */
+		nreaders = nallowed > 1 ? nallowed - 1 : 1;
+		if (nreaders > 4)
+			nreaders = 4;
+	}
+
+	signal(SIGINT, sighandler);
+	signal(SIGALRM, sighandler);
+
+	readers = calloc(nreaders, sizeof(*readers));
+	if (readers == NULL) {
+		fprintf(stderr, "Error: calloc failed for %d readers\n", nreaders);
+		goto out;
+	}
+	for (i = 0; i < nreaders; i++) {
+		int cpu = nallowed > 1 ? allowed[(i + 1) % nallowed] : -1;
+
+		readers[i].cpu = cpu;
+		if (pthread_create(&readers[i].thread, NULL, reader_fn, &readers[i])) {
+			fprintf(stderr, "Error: failed to create reader %d\n", i);
+			done = 1; // Ensure started threads terminate.
+			nreaders = i;
+			goto out_join;
+		}
+	}
+	writer_cpu = nallowed > 0 ? allowed[0] : -1;
+	if (nallowed == 1)
+		fprintf(stderr, "Warning: single CPU allowed, no cross-CPU traffic expected\n");
+	if (writer_cpu >= 0)
+		pin_to_cpu(writer_cpu);
+
+	/*
+	 * The packet path: receive, acknowledge, transmit, repeat; every 64th
+	 * packet simulates a loss so ssthresh/retrans get sampled too.
+	 */
+	if (sec < 1.0) {
+		useconds_t usecs = (useconds_t)(sec * 1000000.0);
+
+		ualarm(usecs > 0 ? usecs : 1, 0);
+	} else
+		alarm((unsigned int)sec);
+	while (!done) {
+		conn.bytes_rx += conn.mss;
+		conn.packets_rx++;
+		conn.rx_queue = (uint32_t)(n & 0x3f);
+		conn.last_ack = (uint32_t)n;
+		conn.bytes_tx += conn.mss;
+		conn.packets_tx++;
+		conn.cwnd = 10 + (n & 15);
+		conn.rtt_us = 50 + (n & 7);
+		if ((n & 63) == 0) {
+			conn.ssthresh = conn.cwnd / 2;
+			conn.retrans++;
+		}
+		n++;
+	}
+	err = 0;
+out_join:
+	for (i = 0; i < nreaders; i++) {
+		if (readers[i].thread) {
+			pthread_join(readers[i].thread, /*retval=*/NULL);
+			fs_sink += readers[i].sum;
+		}
+	}
+	fs_sink += (unsigned long)(conn.bytes_rx + conn.bytes_tx + n);
+	free(readers);
+out:
+	free(allowed);
+	return err;
+}
+
+DEFINE_WORKLOAD(false_sharing);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload
  2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
                   ` (7 preceding siblings ...)
  2026-10-06 14:57 ` [PATCH 8/8] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
@ 2026-10-07  0:07 ` Namhyung Kim
  8 siblings, 0 replies; 10+ messages in thread
From: Namhyung Kim @ 2026-10-07  0:07 UTC (permalink / raw)
  To: Arnaldo Carvalho de Melo
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users

On Tue, Oct 06, 2026 at 04:57:21PM +0200, Arnaldo Carvalho de Melo wrote:
> Hi,
> 
> This series adds progress diagnostics to perf, and a workload that makes
> false sharing visible to data type profiling.
> 
> The changes are:
> 
>   - move perf_config__set_variable() to util/config.c and serialize config
>     parser and read-modify-write state, so non-builtin perf code can persist
>     configuration changes safely;
> 
>   - add 'perf report --progress' for stdio users, showing the current phase,
>     percentage, and counts while a session is processed;
> 
>   - wire up 'perf report --no-progress', the counterpart of the option
>     above, for the TUI and GTK browsers, whose progress there is no other
>     way to turn off;
> 
>   - add 'perf test -w false_sharing', a synthetic TCP-shaped workload with
>     identity and packet counters sharing a cacheline, and include it in the
>     data type profiling shell test.
> 
> Follow up work: the DO_ONCE() one-time init primitive added here mirrors
> what the eight pre-existing raw pthread_once() users in util/ (annotate.c,
> callchain.c, comm.c, dso.c, fncache.c, intel-tpebs.c, libbfd.c, pmus.c)
> need, conversions that will also exercise the DEFINE_MUTEX() static
> initializer; converting them is left for after this series.
> Best regards,

Reviewed-by: Namhyung Kim <namhyung@kernel.org>

Thanks,
Namhyung

> 
> - Arnaldo
> 
> What changed from v11:
> 
>   - The 'perf-stuck' patch is removed for the time being, to be submitted
>     separately later on;
> 
> What changed from v10:
> 
>   - tools/perf/scripts/perf-stuck.sh: pin the watched process by its start
>     time, so a recycled PID can't get samples or gdb; Sashiko, v10;
> 
>   - tools/perf/scripts/perf-stuck.sh: drop the control characters from the
>     shown progress updates, no terminal escape injection; Sashiko, v10;
> 
>   - tools/perf/tests/shell/perf_stuck.sh: test the script, from option
>     validation to the gdb DIE chain runs on stand-in workloads;
> 
> What changed from v9:
> 
>   - tools/perf/ui/stdio/progress.c: progress updates now work well with
>     the pager, printed to /dev/tty so they update in place;
> 
>   - tools/perf/scripts/perf-stuck.sh: take the last progress update
>     from a bounded tail window, dropping the \r's.  Sashiko, v9;
> 
> What changed from v8:
> 
>   - tools/perf/util/mutex.h: DEFINE_MUTEX() statics keep mutex_init()'s
>     !NDEBUG errorcheck type;
> 
>   - patch 4/9 commit log: clarify the acquire/release pairing, what the
>     dropped static key used to guarantee.
> 
> What changed from v7:
> 
>   - tools/perf/util/mutex.h: the DO_ONCE() lockless fast path pairs its
>     __ATOMIC_ACQUIRE load with the __ATOMIC_RELEASE store.  sashiko, v7;
> 
> What changed from v6:
> 
>   - patch 1/5 split into four: the move, perf_etc_perfconfig() never
>     returning NULL, the DEFINE_MUTEX() prep and the serialization.
>     Namhyung Kim asked, v6;
> 
>   - tools/perf/util/{config.c,mutex.h}: perf's own mutex type, adding
>     the DEFINE_MUTEX() initializer it lacked.  Namhyung Kim, v6;
> 
>   - tools/perf/util/mutex.h: add DO_ONCE() one-time init, from the
>     kernel's once.h, moving the lazy inits in config.c to it;
> 
>   - tools/perf/builtin-report.c: drop the option negation and stdio
>     hook comments Namhyung Kim asked to drop, reviewing v6;
> 
>   - tools/perf/scripts/perf-stuck.gdb: break the long printf() line.
>     Namhyung Kim, reviewing v6;
> 
>   - the false_sharing cset now carries the data type profiling output
>     with cacheline info.  Namhyung Kim asked for it, reviewing v6.
> 
> What changed from v5:
> 
>   - tools/perf/util/config.c: drop the config_file_name read in
>     bad_config(), locking would buy a consistent NULL.  Sashiko, v5;
> 
>   - tools/perf/util/config.c: new config_set_mutex around the shared
>     set's init/teardown, home init via pthread_once.  Sashiko, v5;
> 
>   - tools/perf/builtin-report.c, perf-report.txt: --no-progress is
>     --progress's auto negation, last wins.  Namhyung Kim, reviewing v5.
> 
> What changed from v4:
> 
>   - tools/perf/builtin-report.c: mark --progress PARSE_OPT_NOAUTONEG,
>     parse_long_opt() claimed --no-progress first.  Sashiko, v4;
> 
>   - tools/perf/util/config.c: format the path buffer inside the critical
>     section.  Sashiko pointed out the window, reviewing v4;
> 
>   - tools/perf/tests/workloads/false_sharing.c: walk every bit of the
>     affinity mask, not _SC_NPROCESSORS_CONF.  Sashiko, reviewing v4.
> 
> What changed from v3:
> 
>   - tools/perf/util/config.c: keep the buffer handed to the parser in
>     static storage.  Sashiko, reviewing v3;
> 
>   - tools/perf/scripts/perf-stuck.gdb: perf-dso walks each dso candidate
>     until one evaluates, for REFCNT_CHECKING builds.  Sashiko, v3;
> 
>   - tools/perf/tests/workloads/false_sharing.c: sum before cpu in struct
>     fs_reader, the padding after cpu rounded it to 128.  Sashiko, v3;
> 
>   - tools/perf/builtin-report.c: --quiet wins over --progress, the phases
>     stay uncounted, documented next to it.  Namhyung Kim, v3;
> 
>   - tools/perf/builtin-report.c, tools/perf/ui/progress.c: --no-progress
>     installs the no-op ops.  Suggested by Namhyung Kim, reviewing v3;
> 
>   - tools/perf/scripts/perf-stuck.sh: check gdb is there before
>     watching.  Namhyung Kim, reviewing v3.
> 
> What changed from v2:
> 
>   - tools/perf/util/config.c: make perf_etc_perfconfig() total, falling
>     back to the unresolved path.  Sashiko, reviewing v2;
> 
>   - tools/perf/scripts/perf-stuck.gdb: don't deref map_symbol.sym without
>     a NULL check.  Sashiko, reviewing v2;
> 
>   - tools/perf/scripts/perf-stuck.gdb: note the REFCNT_CHECKING proxy
>     next to the structure walks;
> 
>   - tools/perf/scripts/perf-stuck.sh: count samples with no progress to
>     look at, an empty log never fired -g;
> 
>   - tools/perf/scripts/perf-stuck.sh: `--` for pgrep and tail against
>     names starting with a hyphen;
> 
>   - tools/perf/scripts/perf-stuck.sh: bound the gdb run with
>     `timeout --signal=INT 30`.
> 
> What changed from v1:
> 
>   - avoid calling CPU_SET() with -1 when false_sharing runs with only one
>     CPU available in its affinity mask.
> 
>   tools/perf/Documentation/perf-report.txt      |  19 ++
>   tools/perf/builtin-config.c                   |  70 +----
>   tools/perf/builtin-report.c                   |   9 +
>   tools/perf/tests/builtin-test.c               |   1 +
>   tools/perf/tests/shell/data_type_profiling.sh |   9 +-
>   tools/perf/tests/tests.h                      |   1 +
>   tools/perf/tests/workloads/Build              |   2 +
>   tools/perf/tests/workloads/false_sharing.c    | 251 ++++++++++++++++++
>   tools/perf/ui/Build                           |   1 +
>   tools/perf/ui/progress.c                      |   6 +
>   tools/perf/ui/progress.h                      |   4 +
>   tools/perf/ui/stdio/progress.c                | 170 ++++++++++++
>   tools/perf/util/config.c                      | 166 ++++++++++--
>   tools/perf/util/config.h                      |   2 +
>   tools/perf/util/mutex.h                       |  36 +++
>   tools/perf/util/ordered-events.c              |  16 +-
>   tools/perf/util/session.c                     |  12 +-
>   17 files changed, 675 insertions(+), 100 deletions(-)
>   create mode 100644 tools/perf/tests/workloads/false_sharing.c
>   create mode 100644 tools/perf/ui/stdio/progress.c
> base-commit: 705da5b15ab89ba9
> v1-head: 45d7917f7e05e8a29828ed5f0bbdc94fc938f79f
> v2-head: d4f84e4de8890194924ccd897a8e6773e7d4240b
> v3-head: 485532296710225862daf8ffb19802ae327efe2a
> v4-head: 38193635508e0a5f04a6ebf70db27534b68b2d5b
> v5-head: e99bc48f6c72b9cc67b4e4143a0364acac5523bb
> v6-head: 7dc99dbb6504f199c3c39487d91370d1bbb7a827
> v7-head: b357a76666d08f74eb99c755cca41855a96f6169
> v8-head: 257b5ea6115ad5126d2c881523ea6c660cf5d28c
> v9-head: 70235104c6cf7d18d91f46de1a233abbeb65893a
> v10-head: 1ecf850c2694d20ca9cd0014f2a42528ef4dbfec
> v11-head: a3ddbaab18aa24fba520927f71d0ea66edfc5593
> --

^ permalink raw reply	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-10-07  0:07 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 2/8] perf config: Make perf_etc_perfconfig() never return NULL Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 3/8] perf mutex: Add DEFINE_MUTEX() static initializer Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 4/8] perf mutex: Add DO_ONCE() for one-time initialization Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 5/8] perf config: Serialize config file access with a mutex Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 6/8] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 7/8] perf report: Add --no-progress option Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 8/8] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
2026-10-07  0:07 ` [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Namhyung Kim

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®