* [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload
@ 2026-10-06 14:57 Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
` (8 more replies)
0 siblings, 9 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
Hi,
This series adds progress diagnostics to perf, and a workload that makes
false sharing visible to data type profiling.
The changes are:
- move perf_config__set_variable() to util/config.c and serialize config
parser and read-modify-write state, so non-builtin perf code can persist
configuration changes safely;
- add 'perf report --progress' for stdio users, showing the current phase,
percentage, and counts while a session is processed;
- wire up 'perf report --no-progress', the counterpart of the option
above, for the TUI and GTK browsers, whose progress there is no other
way to turn off;
- add 'perf test -w false_sharing', a synthetic TCP-shaped workload with
identity and packet counters sharing a cacheline, and include it in the
data type profiling shell test.
Follow up work: the DO_ONCE() one-time init primitive added here mirrors
what the eight pre-existing raw pthread_once() users in util/ (annotate.c,
callchain.c, comm.c, dso.c, fncache.c, intel-tpebs.c, libbfd.c, pmus.c)
need, conversions that will also exercise the DEFINE_MUTEX() static
initializer; converting them is left for after this series.
Best regards,
- Arnaldo
What changed from v11:
- The 'perf-stuck' patch is removed for the time being, to be submitted
separately later on;
What changed from v10:
- tools/perf/scripts/perf-stuck.sh: pin the watched process by its start
time, so a recycled PID can't get samples or gdb; Sashiko, v10;
- tools/perf/scripts/perf-stuck.sh: drop the control characters from the
shown progress updates, no terminal escape injection; Sashiko, v10;
- tools/perf/tests/shell/perf_stuck.sh: test the script, from option
validation to the gdb DIE chain runs on stand-in workloads;
What changed from v9:
- tools/perf/ui/stdio/progress.c: progress updates now work well with
the pager, printed to /dev/tty so they update in place;
- tools/perf/scripts/perf-stuck.sh: take the last progress update
from a bounded tail window, dropping the \r's. Sashiko, v9;
What changed from v8:
- tools/perf/util/mutex.h: DEFINE_MUTEX() statics keep mutex_init()'s
!NDEBUG errorcheck type;
- patch 4/9 commit log: clarify the acquire/release pairing, what the
dropped static key used to guarantee.
What changed from v7:
- tools/perf/util/mutex.h: the DO_ONCE() lockless fast path pairs its
__ATOMIC_ACQUIRE load with the __ATOMIC_RELEASE store. sashiko, v7;
What changed from v6:
- patch 1/5 split into four: the move, perf_etc_perfconfig() never
returning NULL, the DEFINE_MUTEX() prep and the serialization.
Namhyung Kim asked, v6;
- tools/perf/util/{config.c,mutex.h}: perf's own mutex type, adding
the DEFINE_MUTEX() initializer it lacked. Namhyung Kim, v6;
- tools/perf/util/mutex.h: add DO_ONCE() one-time init, from the
kernel's once.h, moving the lazy inits in config.c to it;
- tools/perf/builtin-report.c: drop the option negation and stdio
hook comments Namhyung Kim asked to drop, reviewing v6;
- tools/perf/scripts/perf-stuck.gdb: break the long printf() line.
Namhyung Kim, reviewing v6;
- the false_sharing cset now carries the data type profiling output
with cacheline info. Namhyung Kim asked for it, reviewing v6.
What changed from v5:
- tools/perf/util/config.c: drop the config_file_name read in
bad_config(), locking would buy a consistent NULL. Sashiko, v5;
- tools/perf/util/config.c: new config_set_mutex around the shared
set's init/teardown, home init via pthread_once. Sashiko, v5;
- tools/perf/builtin-report.c, perf-report.txt: --no-progress is
--progress's auto negation, last wins. Namhyung Kim, reviewing v5.
What changed from v4:
- tools/perf/builtin-report.c: mark --progress PARSE_OPT_NOAUTONEG,
parse_long_opt() claimed --no-progress first. Sashiko, v4;
- tools/perf/util/config.c: format the path buffer inside the critical
section. Sashiko pointed out the window, reviewing v4;
- tools/perf/tests/workloads/false_sharing.c: walk every bit of the
affinity mask, not _SC_NPROCESSORS_CONF. Sashiko, reviewing v4.
What changed from v3:
- tools/perf/util/config.c: keep the buffer handed to the parser in
static storage. Sashiko, reviewing v3;
- tools/perf/scripts/perf-stuck.gdb: perf-dso walks each dso candidate
until one evaluates, for REFCNT_CHECKING builds. Sashiko, v3;
- tools/perf/tests/workloads/false_sharing.c: sum before cpu in struct
fs_reader, the padding after cpu rounded it to 128. Sashiko, v3;
- tools/perf/builtin-report.c: --quiet wins over --progress, the phases
stay uncounted, documented next to it. Namhyung Kim, v3;
- tools/perf/builtin-report.c, tools/perf/ui/progress.c: --no-progress
installs the no-op ops. Suggested by Namhyung Kim, reviewing v3;
- tools/perf/scripts/perf-stuck.sh: check gdb is there before
watching. Namhyung Kim, reviewing v3.
What changed from v2:
- tools/perf/util/config.c: make perf_etc_perfconfig() total, falling
back to the unresolved path. Sashiko, reviewing v2;
- tools/perf/scripts/perf-stuck.gdb: don't deref map_symbol.sym without
a NULL check. Sashiko, reviewing v2;
- tools/perf/scripts/perf-stuck.gdb: note the REFCNT_CHECKING proxy
next to the structure walks;
- tools/perf/scripts/perf-stuck.sh: count samples with no progress to
look at, an empty log never fired -g;
- tools/perf/scripts/perf-stuck.sh: `--` for pgrep and tail against
names starting with a hyphen;
- tools/perf/scripts/perf-stuck.sh: bound the gdb run with
`timeout --signal=INT 30`.
What changed from v1:
- avoid calling CPU_SET() with -1 when false_sharing runs with only one
CPU available in its affinity mask.
tools/perf/Documentation/perf-report.txt | 19 ++
tools/perf/builtin-config.c | 70 +----
tools/perf/builtin-report.c | 9 +
tools/perf/tests/builtin-test.c | 1 +
tools/perf/tests/shell/data_type_profiling.sh | 9 +-
tools/perf/tests/tests.h | 1 +
tools/perf/tests/workloads/Build | 2 +
tools/perf/tests/workloads/false_sharing.c | 251 ++++++++++++++++++
tools/perf/ui/Build | 1 +
tools/perf/ui/progress.c | 6 +
tools/perf/ui/progress.h | 4 +
tools/perf/ui/stdio/progress.c | 170 ++++++++++++
tools/perf/util/config.c | 166 ++++++++++--
tools/perf/util/config.h | 2 +
tools/perf/util/mutex.h | 36 +++
tools/perf/util/ordered-events.c | 16 +-
tools/perf/util/session.c | 12 +-
17 files changed, 675 insertions(+), 100 deletions(-)
create mode 100644 tools/perf/tests/workloads/false_sharing.c
create mode 100644 tools/perf/ui/stdio/progress.c
base-commit: 705da5b15ab89ba9
v1-head: 45d7917f7e05e8a29828ed5f0bbdc94fc938f79f
v2-head: d4f84e4de8890194924ccd897a8e6773e7d4240b
v3-head: 485532296710225862daf8ffb19802ae327efe2a
v4-head: 38193635508e0a5f04a6ebf70db27534b68b2d5b
v5-head: e99bc48f6c72b9cc67b4e4143a0364acac5523bb
v6-head: 7dc99dbb6504f199c3c39487d91370d1bbb7a827
v7-head: b357a76666d08f74eb99c755cca41855a96f6169
v8-head: 257b5ea6115ad5126d2c881523ea6c660cf5d28c
v9-head: 70235104c6cf7d18d91f46de1a233abbeb65893a
v10-head: 1ecf850c2694d20ca9cd0014f2a42528ef4dbfec
v11-head: a3ddbaab18aa24fba520927f71d0ea66edfc5593
--
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 2/8] perf config: Make perf_etc_perfconfig() never return NULL Arnaldo Carvalho de Melo
` (7 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
Move perf_config__set_variable() out of the 'perf config' builtin so
that opt-in features can persist their choice from outside it, e.g.
util/debuginfo.c writing core.debuginfod=false when the user disables
debuginfod for the rest of the session. The set_config() body becomes
perf_config_set__write(), with the system_config choice as an argument,
as the builtin's use_system_config/use_user_config statics are not
available outside it.
perf_config_set__write() checked fopen() but none of the fprintf()s or
fclose(), so a write failure after truncating the file was reported as
success. Harmless for the interactive 'perf config' this came from,
but this makes it an entry point a background feature can call with no
other feedback, so propagate those errors too.
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/builtin-config.c | 70 +---------------------------------
tools/perf/util/config.c | 76 +++++++++++++++++++++++++++++++++++++
tools/perf/util/config.h | 2 +
3 files changed, 79 insertions(+), 69 deletions(-)
diff --git a/tools/perf/builtin-config.c b/tools/perf/builtin-config.c
index cefd042e4f853466..3b074aca8d344539 100644
--- a/tools/perf/builtin-config.c
+++ b/tools/perf/builtin-config.c
@@ -41,37 +41,7 @@ static struct option config_options[] = {
static int set_config(struct perf_config_set *set, const char *file_name)
{
- struct perf_config_section *section = NULL;
- struct perf_config_item *item = NULL;
- const char *first_line = "# this file is auto-generated.";
- FILE *fp;
-
- if (set == NULL)
- return -1;
-
- fp = fopen(file_name, "w");
- if (!fp)
- return -1;
-
- fprintf(fp, "%s\n", first_line);
-
- /* overwrite configvariables */
- perf_config_items__for_each_entry(&set->sections, section) {
- if (!use_system_config && section->from_system_config)
- continue;
- fprintf(fp, "[%s]\n", section->name);
-
- perf_config_items__for_each_entry(§ion->items, item) {
- if (!use_system_config && item->from_system_config)
- continue;
- if (item->value)
- fprintf(fp, "\t%s = %s\n",
- item->name, item->value);
- }
- }
- fclose(fp);
-
- return 0;
+ return perf_config_set__write(set, file_name, use_system_config);
}
static int show_spec_config(struct perf_config_set *set, const char *var)
@@ -158,44 +128,6 @@ static int parse_config_arg(char *arg, char **var, char **value)
return 0;
}
-int perf_config__set_variable(const char *var, const char *value)
-{
- char path[PATH_MAX];
- char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
- const char *config_filename;
- struct perf_config_set *set;
- int ret = -1;
-
- if (use_system_config)
- config_exclusive_filename = perf_etc_perfconfig();
- else if (use_user_config)
- config_exclusive_filename = user_config;
-
- if (!config_exclusive_filename)
- config_filename = user_config;
- else
- config_filename = config_exclusive_filename;
-
- set = perf_config_set__new();
- if (!set)
- goto out_err;
-
- if (perf_config_set__collect(set, config_filename, var, value) < 0) {
- pr_err("Failed to add '%s=%s'\n", var, value);
- goto out_err;
- }
-
- if (set_config(set, config_filename) < 0) {
- pr_err("Failed to set the configs on %s\n", config_filename);
- goto out_err;
- }
-
- ret = 0;
-out_err:
- perf_config_set__delete(set);
- return ret;
-}
-
int cmd_config(int argc, const char **argv)
{
int i, ret = -1;
diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
index 8fe43b032e9af88a..85e50d0a25580da0 100644
--- a/tools/perf/util/config.c
+++ b/tools/perf/util/config.c
@@ -12,6 +12,7 @@
#include "config.h"
#include <errno.h>
+#include <limits.h>
#include <stdbool.h>
#include <stdio.h>
#include <stdlib.h>
@@ -883,6 +884,81 @@ void perf_config__exit(void)
config_set = NULL;
}
+int perf_config_set__write(struct perf_config_set *set,
+ const char *file_name, bool system_config)
+{
+ struct perf_config_section *section = NULL;
+ struct perf_config_item *item = NULL;
+ int ret = 0;
+ FILE *fp;
+
+ fp = fopen(file_name, "w");
+ if (!fp)
+ return -1;
+
+ if (fprintf(fp, "# this file is auto-generated.\n") < 0)
+ ret = -1;
+
+ /* overwrite configvariables */
+ perf_config_sections__for_each_entry(&set->sections, section) {
+ if (!system_config && section->from_system_config)
+ continue;
+ if (fprintf(fp, "[%s]\n", section->name) < 0)
+ ret = -1;
+
+ perf_config_items__for_each_entry(§ion->items, item) {
+ if (!system_config && item->from_system_config)
+ continue;
+ if (item->value &&
+ fprintf(fp, "\t%s = %s\n", item->name, item->value) < 0)
+ ret = -1;
+ }
+ }
+ if (fclose(fp) != 0)
+ ret = -1;
+
+ return ret;
+}
+
+/*
+ * Set @var=@value in the config file perf is using: ~/.perfconfig or the
+ * file named by PERF_CONFIG. Same rewrite 'perf config' does, comments
+ * are not preserved.
+ */
+int perf_config__set_variable(const char *var, const char *value)
+{
+ char path[PATH_MAX];
+ char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
+ const char *config_filename;
+ bool system_config;
+ struct perf_config_set *set;
+ int ret = -1;
+
+ config_filename = config_exclusive_filename ?: user_config;
+
+ /* Rewriting the system wide file keeps its entries, or it is truncated. */
+ system_config = strcmp(config_filename, perf_etc_perfconfig()) == 0;
+
+ set = perf_config_set__new();
+ if (!set)
+ goto out_err;
+
+ if (perf_config_set__collect(set, config_filename, var, value) < 0) {
+ pr_err("Failed to add '%s=%s'\n", var, value);
+ goto out_err;
+ }
+
+ if (perf_config_set__write(set, config_filename, system_config) < 0) {
+ pr_err("Failed to set the configs on %s\n", config_filename);
+ goto out_err;
+ }
+
+ ret = 0;
+out_err:
+ perf_config_set__delete(set);
+ return ret;
+}
+
static void perf_config_item__delete(struct perf_config_item *item)
{
zfree(&item->name);
diff --git a/tools/perf/util/config.h b/tools/perf/util/config.h
index 987b47cf54c350ba..9098f8a045850c97 100644
--- a/tools/perf/util/config.h
+++ b/tools/perf/util/config.h
@@ -33,6 +33,8 @@ int perf_config_scan(const char *name, const char *fmt, ...) __scanf(2, 3);
const char *perf_config_get(const char *name);
int perf_config_set(struct perf_config_set *set,
config_fn_t fn, void *data);
+int perf_config_set__write(struct perf_config_set *set,
+ const char *file_name, bool system_config);
int perf_config_int(int *dest, const char *, const char *);
int perf_config_u8(u8 *dest, const char *name, const char *value);
int perf_config_u64(u64 *dest, const char *, const char *);
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 2/8] perf config: Make perf_etc_perfconfig() never return NULL
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 3/8] perf mutex: Add DEFINE_MUTEX() static initializer Arnaldo Carvalho de Melo
` (6 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
perf_etc_perfconfig() returns what system_path() gives, and that
allocates, so on failure every caller dereferenced NULL, the
pre-existing ones in perf_config_set__init() and in the daemon
included. For the usual absolute ETC_PERFCONFIG system_path() just
returns a strdup of it, so use ETC_PERFCONFIG itself when that fails.
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/util/config.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
index 85e50d0a25580da0..6e0ff8a9140ccf96 100644
--- a/tools/perf/util/config.c
+++ b/tools/perf/util/config.c
@@ -571,8 +571,14 @@ static int perf_config_from_file(config_fn_t fn, const char *filename, void *dat
const char *perf_etc_perfconfig(void)
{
static const char *system_wide;
+
if (!system_wide)
- system_wide = system_path(ETC_PERFCONFIG);
+ /*
+ * ETC_PERFCONFIG is absolute, so its unresolved path is
+ * the same string, better than the callers crashing.
+ */
+ system_wide = system_path(ETC_PERFCONFIG) ?: ETC_PERFCONFIG;
+
return system_wide;
}
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 3/8] perf mutex: Add DEFINE_MUTEX() static initializer
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 2/8] perf config: Make perf_etc_perfconfig() never return NULL Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 4/8] perf mutex: Add DO_ONCE() for one-time initialization Arnaldo Carvalho de Melo
` (5 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
For file scope mutexes, so that they don't need a constructor function
to initialize them at startup, mirroring PTHREAD_MUTEX_INITIALIZER
while keeping the clang -Wthread-safety annotations of struct mutex.
DEBUG=1 builds have mutex_init() set PTHREAD_MUTEX_ERRORCHECK, making
the CHECK_ERR() paths in mutex_lock()/mutex_unlock() report relocking a
held mutex and unlocking one that isn't held. PTHREAD_MUTEX_INITIALIZER
gives a default type mutex, so statically initialized ones would silently
lose that: use PTHREAD_ERRORCHECK_MUTEX_INITIALIZER_NP, a glibc
extension, falling back to the default type in libcs that lack it.
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/util/mutex.h | 12 ++++++++++++
1 file changed, 12 insertions(+)
diff --git a/tools/perf/util/mutex.h b/tools/perf/util/mutex.h
index 70232d8d094f8bfc..fe04db21d19c5b0b 100644
--- a/tools/perf/util/mutex.h
+++ b/tools/perf/util/mutex.h
@@ -92,6 +92,18 @@ struct LOCKABLE mutex {
pthread_mutex_t lock;
};
+/*
+ * Statically initialized mutex, for the file scope ones. Error checking
+ * in !NDEBUG builds, like mutex_init(); the glibc-only errorcheck
+ * initializer falls back to the default type elsewhere.
+ */
+#if !defined(NDEBUG) && defined(PTHREAD_ERRORCHECK_MUTEX_INITIALIZER_NP)
+#define __PERF_MUTEX_INITIALIZER PTHREAD_ERRORCHECK_MUTEX_INITIALIZER_NP
+#else
+#define __PERF_MUTEX_INITIALIZER PTHREAD_MUTEX_INITIALIZER
+#endif
+#define DEFINE_MUTEX(name) struct mutex name = { .lock = __PERF_MUTEX_INITIALIZER }
+
/* A wrapper around the condition variable implementation. */
struct cond {
pthread_cond_t cond;
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 4/8] perf mutex: Add DO_ONCE() for one-time initialization
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
` (2 preceding siblings ...)
2026-10-06 14:57 ` [PATCH 3/8] perf mutex: Add DEFINE_MUTEX() static initializer Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 5/8] perf config: Serialize config file access with a mutex Arnaldo Carvalho de Melo
` (4 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
For lazy init that more than one thread can race into, without
pthread_once() so that it stays within perf's locking primitives and
clang's -Wthread-safety keeps an eye on the mutexes involved.
Mirrors the kernel's include/linux/once.h DO_ONCE(), dropping its
static key fast path, unnecessary at these cold init paths: a per call
site static bool plus a statically initialized mutex, double checked.
As there, code reachable from more than one call site must go through
a common helper.
The lockless fast path reads ___done with an acquire, paired with a
release store after fn(), so that fn()'s side effects are visible to
threads taking the fast path, which the dropped static key used to
guarantee.
The __atomic builtins are used rather than smp_load_acquire()/
smp_store_release() because ThreadSanitizer understands them,
avoiding reports on ___done and on what fn() initialized.
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/util/mutex.h | 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
diff --git a/tools/perf/util/mutex.h b/tools/perf/util/mutex.h
index fe04db21d19c5b0b..5554937675886ef3 100644
--- a/tools/perf/util/mutex.h
+++ b/tools/perf/util/mutex.h
@@ -104,6 +104,30 @@ struct LOCKABLE mutex {
#endif
#define DEFINE_MUTEX(name) struct mutex name = { .lock = __PERF_MUTEX_INITIALIZER }
+/*
+ * Runs fn() exactly once however many threads race to it, the rest wait
+ * for it to finish; per call site state, so multiple sites wanting the
+ * same once use a common helper, as in the kernel's DO_ONCE().
+ */
+#define DO_ONCE(fn, ...) \
+ ({ \
+ static bool ___done; \
+ static DEFINE_MUTEX(___once_lock); \
+ bool ___ret = false; \
+ \
+ if (!__atomic_load_n(&___done, __ATOMIC_ACQUIRE)) { \
+ mutex_lock(&___once_lock); \
+ if (!___done) { \
+ fn(__VA_ARGS__); \
+ __atomic_store_n(&___done, true, \
+ __ATOMIC_RELEASE); \
+ ___ret = true; \
+ } \
+ mutex_unlock(&___once_lock); \
+ } \
+ ___ret; \
+ })
+
/* A wrapper around the condition variable implementation. */
struct cond {
pthread_cond_t cond;
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 5/8] perf config: Serialize config file access with a mutex
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
` (3 preceding siblings ...)
2026-10-06 14:57 ` [PATCH 4/8] perf mutex: Add DO_ONCE() for one-time initialization Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 6/8] perf report: Add --progress option Arnaldo Carvalho de Melo
` (3 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
Config file access shares the static parser state and, since
perf_config__set_variable() became an entry point for background
features, can run on more than one thread: a feature writing the
config while perf top's display thread reads it. Serialize parsing
and rewriting with config_mutex and the whole read-modify-write of
perf_config__set_variable() with config_update_mutex; config_set_mutex
comes before it, as building the set parses the config files.
perf has its own mutex type, util/mutex.h; the file scope mutexes
here use the DEFINE_MUTEX() static initializer added in the previous
patch. The two lazy inits that system_path() and home_perfconfig()
do would leak all but one of the strings racing on two threads, so
they move to DO_ONCE(), added in the previous patch as well.
bad_config() runs on the dispatching thread, outside config_mutex, so
it stops reading config_file_name, which the parsing thread owns.
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/util/config.c | 106 +++++++++++++++++++++++++++------------
1 file changed, 73 insertions(+), 33 deletions(-)
diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
index 6e0ff8a9140ccf96..287401f159e66e8c 100644
--- a/tools/perf/util/config.c
+++ b/tools/perf/util/config.c
@@ -31,6 +31,7 @@
#include "callchain.h"
#include "debug.h"
#include "header.h"
+#include "mutex.h"
#include "path.h"
#include "srcline.h"
#include "unwind.h"
@@ -371,10 +372,8 @@ static int perf_parse_long(const char *value, long *ret)
static void bad_config(const char *name)
{
- if (config_file_name)
- pr_warning("bad config value for '%s' in %s, ignoring...\n", name, config_file_name);
- else
- pr_warning("bad config value for '%s', ignoring...\n", name);
+ /* config_file_name is owned by the parsing thread, under config_mutex. */
+ pr_warning("bad config value for '%s', ignoring...\n", name);
}
int perf_config_u64(u64 *dest, const char *name, const char *value)
@@ -550,11 +549,18 @@ int perf_default_config(const char *var, const char *value,
return 0;
}
+/* Parsing and rewriting share the static parser state. */
+static DEFINE_MUTEX(config_mutex);
+
+/* Serializes whole perf_config__set_variable() updates. */
+static DEFINE_MUTEX(config_update_mutex);
+
static int perf_config_from_file(config_fn_t fn, const char *filename, void *data)
{
int ret;
FILE *f = fopen(filename, "r");
+ mutex_lock(&config_mutex);
ret = -1;
if (f) {
config_file = f;
@@ -565,21 +571,24 @@ static int perf_config_from_file(config_fn_t fn, const char *filename, void *dat
fclose(f);
config_file_name = NULL;
}
+ mutex_unlock(&config_mutex);
return ret;
}
-const char *perf_etc_perfconfig(void)
-{
- static const char *system_wide;
+/* system_path() allocates, so it is computed once. */
+static const char *etc_perfconfig;
- if (!system_wide)
- /*
- * ETC_PERFCONFIG is absolute, so its unresolved path is
- * the same string, better than the callers crashing.
- */
- system_wide = system_path(ETC_PERFCONFIG) ?: ETC_PERFCONFIG;
+static void perf_etc_perfconfig__init(void)
+{
+ etc_perfconfig = system_path(ETC_PERFCONFIG);
+ if (!etc_perfconfig)
+ etc_perfconfig = ETC_PERFCONFIG;
+}
- return system_wide;
+const char *perf_etc_perfconfig(void)
+{
+ DO_ONCE(perf_etc_perfconfig__init);
+ return etc_perfconfig;
}
static int perf_env_bool(const char *k, int def)
@@ -637,19 +646,18 @@ static char *home_perfconfig(void)
return NULL;
}
-const char *perf_home_perfconfig(void)
-{
- static const char *config;
- static bool failed;
-
- if (failed || config)
- return config;
+/* home_perfconfig() allocates and warns, so it is computed once. */
+static const char *home_config;
- config = home_perfconfig();
- if (!config)
- failed = true;
+static void perf_home_perfconfig__init(void)
+{
+ home_config = home_perfconfig();
+}
- return config;
+const char *perf_home_perfconfig(void)
+{
+ DO_ONCE(perf_home_perfconfig__init);
+ return home_config;
}
static struct perf_config_section *find_section(struct list_head *sections,
@@ -790,8 +798,15 @@ static int collect_config(const char *var, const char *value,
int perf_config_set__collect(struct perf_config_set *set, const char *file_name,
const char *var, const char *value)
{
+ int ret;
+
+ mutex_lock(&config_mutex);
config_file_name = file_name;
- return collect_config(var, value, set);
+ ret = collect_config(var, value, set);
+ /* Don't leave the static parser state pointing at the caller's buffer. */
+ config_file_name = NULL;
+ mutex_unlock(&config_mutex);
+ return ret;
}
static int perf_config_set__init(struct perf_config_set *set)
@@ -838,6 +853,10 @@ struct perf_config_set *perf_config_set__load_file(const char *file)
return set;
}
+/* Not config_mutex: building the set parses the config files, which takes it. */
+static DEFINE_MUTEX(config_set_mutex);
+
+/* Called with config_set_mutex held. */
static int perf_config__init(void)
{
if (config_set == NULL)
@@ -878,16 +897,26 @@ int perf_config_set(struct perf_config_set *set,
int perf_config(config_fn_t fn, void *data)
{
- if (config_set == NULL && perf_config__init())
+ struct perf_config_set *set;
+
+ /* Not held across the dispatch: a callback can call perf_config() again. */
+ mutex_lock(&config_set_mutex);
+ if (perf_config__init()) {
+ mutex_unlock(&config_set_mutex);
return -1;
+ }
+ set = config_set;
+ mutex_unlock(&config_set_mutex);
- return perf_config_set(config_set, fn, data);
+ return perf_config_set(set, fn, data);
}
void perf_config__exit(void)
{
+ mutex_lock(&config_set_mutex);
perf_config_set__delete(config_set);
config_set = NULL;
+ mutex_unlock(&config_set_mutex);
}
int perf_config_set__write(struct perf_config_set *set,
@@ -898,9 +927,12 @@ int perf_config_set__write(struct perf_config_set *set,
int ret = 0;
FILE *fp;
+ mutex_lock(&config_mutex);
fp = fopen(file_name, "w");
- if (!fp)
+ if (!fp) {
+ mutex_unlock(&config_mutex);
return -1;
+ }
if (fprintf(fp, "# this file is auto-generated.\n") < 0)
ret = -1;
@@ -922,6 +954,7 @@ int perf_config_set__write(struct perf_config_set *set,
}
if (fclose(fp) != 0)
ret = -1;
+ mutex_unlock(&config_mutex);
return ret;
}
@@ -933,14 +966,20 @@ int perf_config_set__write(struct perf_config_set *set,
*/
int perf_config__set_variable(const char *var, const char *value)
{
- char path[PATH_MAX];
- char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
const char *config_filename;
bool system_config;
- struct perf_config_set *set;
+ struct perf_config_set *set = NULL;
int ret = -1;
- config_filename = config_exclusive_filename ?: user_config;
+ mutex_lock(&config_update_mutex);
+
+ /* Static: the parser publishes it as config_file_name. */
+ {
+ static char path[PATH_MAX];
+ char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
+
+ config_filename = config_exclusive_filename ?: user_config;
+ }
/* Rewriting the system wide file keeps its entries, or it is truncated. */
system_config = strcmp(config_filename, perf_etc_perfconfig()) == 0;
@@ -962,6 +1001,7 @@ int perf_config__set_variable(const char *var, const char *value)
ret = 0;
out_err:
perf_config_set__delete(set);
+ mutex_unlock(&config_update_mutex);
return ret;
}
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 6/8] perf report: Add --progress option
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
` (4 preceding siblings ...)
2026-10-06 14:57 ` [PATCH 5/8] perf config: Serialize config file access with a mutex Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 7/8] perf report: Add --no-progress option Arnaldo Carvalho de Melo
` (2 subsequent siblings)
8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
Processing a large session with stdio output gives no feedback about
which phase perf is in or how far along it is: ui_progress updates are
only shown by the TUI. Add --progress, installing a stdio backend
(ui/stdio/progress.c) that prints the phase title, percentage and
counts:
Processing events... [ 42.3%] 317M / 746M
Phases can be nested, so the backend tracks the ones started to
complete the right one on ui_progress__finish(); that requires
init()/finish() pairs, fixed in ordered-events.c and the pipe and
directory event processing.
With a pager both stdout and stderr lead to it, so isatty(stderr)
turns false and the updates end up printed one per line in the
pager's output: use /dev/tty in that case, so that the updates keep
updating in place, untouched by the pager.
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/Documentation/perf-report.txt | 14 ++
tools/perf/builtin-report.c | 6 +
tools/perf/ui/Build | 1 +
tools/perf/ui/progress.h | 2 +
tools/perf/ui/stdio/progress.c | 170 +++++++++++++++++++++++
tools/perf/util/ordered-events.c | 16 ++-
tools/perf/util/session.c | 12 +-
7 files changed, 214 insertions(+), 7 deletions(-)
create mode 100644 tools/perf/ui/stdio/progress.c
diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index 3718ebd297ce0673..a7429a30ec28f903 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -29,6 +29,20 @@ OPTIONS
--quiet::
Do not show any warnings or messages. (Suppress -v)
+--progress::
+ Show progress for each of the processing phases, printing the
+ percentage done and the current/total counts: the first phase
+ counts the bytes of the perf.data file processed so far, the
+ merge and sort phases count hist entries, e.g.:
+
+ Processing events... [ 42.3%] 4G / 10G
+ Merging related events... [ 7.1%] 1024 / 14387
+ Sorting events for output... [ 98.2%] 14132 / 14387
+
+ It is a no-op when using the TUI or GTK browsers, that already
+ present progress information, or when --quiet is used, that asks
+ for no messages at all.
+
-n::
--show-nr-samples::
Show the number of samples for each symbol
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 57225bc87731d074..81e13fb8194976c9 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -87,6 +87,7 @@ struct report {
bool use_gtk;
#endif
bool use_stdio;
+ bool progress;
bool show_full_info;
bool show_threads;
bool inverted_callchain;
@@ -1384,6 +1385,8 @@ int cmd_report(int argc, const char **argv)
"Use the stdio interface"),
OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
"Show or hide weight columns in annotation. Default show if non-zero."),
+ OPT_BOOLEAN(0, "progress", &report.progress,
+ "Show progress while processing the perf.data file"),
OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
OPT_BOOLEAN(0, "header-only", &report.header_only,
"Show only data header."),
@@ -1790,6 +1793,9 @@ int cmd_report(int argc, const char **argv)
else
use_browser = 0;
+ if (report.progress && !quiet && use_browser == 0)
+ stdio_progress__init();
+
if (report.data_type && use_browser == 1) {
symbol_conf.annotate_data_member = true;
symbol_conf.annotate_data_sample = true;
diff --git a/tools/perf/ui/Build b/tools/perf/ui/Build
index 6005f813c9e3990c..a7b1740d51c80f23 100644
--- a/tools/perf/ui/Build
+++ b/tools/perf/ui/Build
@@ -4,6 +4,7 @@ perf-ui-y += progress.o
perf-ui-y += util.o
perf-ui-y += hist.o
perf-ui-y += stdio/hist.o
+perf-ui-y += stdio/progress.o
CFLAGS_setup.o += -DLIBDIR="BUILD_STR($(LIBDIR))"
diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
index 4f52c37b2f099a82..03f1a8bb260ba076 100644
--- a/tools/perf/ui/progress.h
+++ b/tools/perf/ui/progress.h
@@ -23,6 +23,8 @@ void __ui_progress__init(struct ui_progress *p, u64 total,
void ui_progress__update(struct ui_progress *p, u64 adv);
+void stdio_progress__init(void);
+
struct ui_progress_ops {
void (*init)(struct ui_progress *p);
void (*update)(struct ui_progress *p);
diff --git a/tools/perf/ui/stdio/progress.c b/tools/perf/ui/stdio/progress.c
new file mode 100644
index 0000000000000000..57f615c455b0ca7d
--- /dev/null
+++ b/tools/perf/ui/stdio/progress.c
@@ -0,0 +1,170 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Progress feedback for the stdio (non-TUI/GTK) case, enabled with
+ * 'perf report --progress'.
+ */
+#include <inttypes.h>
+#include <stdio.h>
+#include <unistd.h>
+#include <linux/kernel.h>
+#include <subcmd/pager.h>
+#include "../../util/debug.h"
+#include "../../util/units.h"
+#include "../progress.h"
+
+/*
+ * Phases can be nested, so keep track of the ones started so far to be
+ * able to complete the right one on ui_progress__finish(), which gets
+ * no arguments.
+ */
+#define STDIO_PROGRESS__MAX_DEPTH 8
+
+struct stdio_progress_phase {
+ struct ui_progress *p;
+ u64 last_printed;
+ size_t last_len;
+};
+
+static struct stdio_progress_phase stdio_progress__stack[STDIO_PROGRESS__MAX_DEPTH];
+static int stdio_progress__depth;
+static FILE *stdio_progress__out;
+static bool stdio_progress__is_tty;
+/* Phases that didn't fit on the stack are not shown. */
+static int stdio_progress__dropped;
+
+static void stdio_progress__print_phase(struct stdio_progress_phase *phase,
+ u64 curr)
+{
+ struct ui_progress *p = phase->p;
+ char buf_cur[20], buf_tot[20], buf[128];
+ double percent = p->total ? 100.0 * (double)curr / (double)p->total : 0.0;
+ size_t len;
+
+ /*
+ * Only the completion line shows 100.0%: a 99.99% progress would round
+ * up to it and look like a duplicate at finish time.
+ */
+ if (curr < p->total && percent > 99.9)
+ percent = 99.9;
+
+ if (p->size) {
+ unit_number__scnprintf(buf_cur, sizeof(buf_cur), curr);
+ unit_number__scnprintf(buf_tot, sizeof(buf_tot), p->total);
+ len = scnprintf(buf, sizeof(buf), "%s [%5.1f%%] %s / %s",
+ p->title, percent, buf_cur, buf_tot);
+ } else {
+ len = scnprintf(buf, sizeof(buf), "%s [%5.1f%%] %" PRIu64 " / %" PRIu64,
+ p->title, percent, curr, p->total);
+ }
+
+ if (!stdio_progress__is_tty) {
+ fprintf(stdio_progress__out, "%s\n", buf);
+ goto out;
+ }
+
+ /* Pad to the length of the previous line to erase its leftovers. */
+ fprintf(stdio_progress__out, "\r%s%*s", buf,
+ (int)(len < phase->last_len ? phase->last_len - len : 0), "");
+ phase->last_len = len;
+out:
+ phase->last_printed = curr;
+ fflush(stdio_progress__out);
+}
+
+static void __stdio_progress__init(struct ui_progress *p)
+{
+ /* The default step is meant for the TUI bar, use 1% steps for stdio. */
+ p->next = p->step = p->total / 100 ?: 1;
+
+ if (stdio_progress__depth == STDIO_PROGRESS__MAX_DEPTH) {
+ /*
+ * Out of room: don't start this phase, its finish() is
+ * swallowed and its updates ignored below.
+ */
+ pr_warning("progress phases nested deeper than %d, not showing progress for %s\n",
+ STDIO_PROGRESS__MAX_DEPTH, p->title);
+ stdio_progress__dropped++;
+ return;
+ }
+
+ /* Start a nested phase in a line of its own. */
+ if (stdio_progress__depth && stdio_progress__is_tty)
+ fputc('\n', stdio_progress__out);
+
+ stdio_progress__stack[stdio_progress__depth++] =
+ (struct stdio_progress_phase) {
+ .p = p,
+ .last_printed = 0,
+ .last_len = 0,
+ };
+
+ stdio_progress__print_phase(&stdio_progress__stack[stdio_progress__depth - 1],
+ p->curr);
+}
+
+static void stdio_progress__update(struct ui_progress *p)
+{
+ /*
+ * An update that doesn't match the innermost phase means something
+ * is out of sync: print nothing rather than another phase's
+ * numbers, or read past the stack.
+ */
+ if (!stdio_progress__depth ||
+ stdio_progress__stack[stdio_progress__depth - 1].p != p)
+ return;
+
+ stdio_progress__print_phase(&stdio_progress__stack[stdio_progress__depth - 1],
+ p->curr);
+}
+
+static void stdio_progress__finish(void)
+{
+ struct stdio_progress_phase *phase;
+
+ /*
+ * Being the innermost phase, its finish() comes first: swallow it, or
+ * it would complete the phase that encloses it.
+ */
+ if (stdio_progress__dropped) {
+ stdio_progress__dropped--;
+ return;
+ }
+
+ if (!stdio_progress__depth)
+ return;
+
+ phase = &stdio_progress__stack[--stdio_progress__depth];
+
+ /*
+ * The last line may have stopped short of the total, close this phase
+ * showing it as complete unless that was already printed.
+ */
+ if (phase->last_printed != phase->p->total)
+ stdio_progress__print_phase(phase, phase->p->total);
+
+ phase->last_printed = 0;
+ phase->last_len = 0;
+
+ if (stdio_progress__is_tty)
+ fputc('\n', stdio_progress__out);
+
+ fflush(stdio_progress__out);
+}
+
+static struct ui_progress_ops stdio_progress__ops = {
+ .init = __stdio_progress__init,
+ .update = stdio_progress__update,
+ .finish = stdio_progress__finish,
+};
+
+void stdio_progress__init(void)
+{
+ if (pager_in_use()) {
+ /* stderr leads to the pager, the updates go to the terminal. */
+ stdio_progress__out = fopen("/dev/tty", "w");
+ }
+ if (!stdio_progress__out)
+ stdio_progress__out = stderr;
+ stdio_progress__is_tty = isatty(fileno(stdio_progress__out)) == 1;
+ ui_progress__ops = &stdio_progress__ops;
+}
diff --git a/tools/perf/util/ordered-events.c b/tools/perf/util/ordered-events.c
index a5857f9f5af2d3de..54c85663e733be4d 100644
--- a/tools/perf/util/ordered-events.c
+++ b/tools/perf/util/ordered-events.c
@@ -237,14 +237,16 @@ static int do_flush(struct ordered_events *oe, bool show_progress)
ui_progress__init(&prog, oe->nr_events, "Processing time ordered events...");
list_for_each_entry_safe(iter, tmp, head, list) {
- if (session_done())
- return 0;
+ if (session_done()) {
+ ret = 0;
+ goto out_progress;
+ }
if (iter->timestamp > limit)
break;
ret = oe->deliver(oe, iter);
if (ret < 0)
- return ret;
+ goto out_progress;
ordered_events__delete(oe, iter);
oe->last_flush = iter->timestamp;
@@ -258,10 +260,16 @@ static int do_flush(struct ordered_events *oe, bool show_progress)
else if (last_ts <= limit)
oe->last = list_entry(head->prev, struct ordered_event, list);
+ ret = 0;
+out_progress:
+ /*
+ * Always pair ui_progress__init() with ui_progress__finish(), the
+ * stdio backend tracks the phases on a stack.
+ */
if (show_progress)
ui_progress__finish();
- return 0;
+ return ret;
}
static int __ordered_events__flush(struct ordered_events *oe, enum oe_flush how,
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index 7fea9e72726c936c..c4b4c7fb3b864589 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -3253,8 +3253,12 @@ static int __perf_session__process_pipe_events(struct perf_session *session)
cur_size = sizeof(union perf_event);
buf = malloc(cur_size);
- if (!buf)
- return -errno;
+ if (!buf) {
+ err = -errno;
+ if (update_prog)
+ ui_progress__finish();
+ return err;
+ }
ordered_events__set_copy_on_queue(oe, true);
more:
event = buf;
@@ -3753,8 +3757,10 @@ static int __perf_session__process_dir_events(struct perf_session *session)
}
rd = calloc(nr_readers, sizeof(struct reader));
- if (!rd)
+ if (!rd) {
+ ui_progress__finish();
return -ENOMEM;
+ }
rd[0] = (struct reader) {
.fd = perf_data__fd(session->data),
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 7/8] perf report: Add --no-progress option
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
` (5 preceding siblings ...)
2026-10-06 14:57 ` [PATCH 6/8] perf report: Add --progress option Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 8/8] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
2026-10-07 0:07 ` [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Namhyung Kim
8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
Now that --progress is being added for the stdio case, wire up its
counterpart for the browsers: the TUI and GTK ones present progress of
their own and there is no way to turn it off. Install the no-op
ui_progress ops, the ones already used until a backend sets theirs,
after setup_browser() installed the ones of the browser in use. The
phases are still counted, nothing is shown for them, and no second
option is needed for it: parse-options provides --no-progress as the
negation of --progress, report.progress_set saying that it was asked
for, report.progress being false both when nothing was asked for and
when --no-progress was.
Suggested-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/Documentation/perf-report.txt | 5 +++++
tools/perf/builtin-report.c | 9 ++++++---
tools/perf/ui/progress.c | 6 ++++++
tools/perf/ui/progress.h | 2 ++
4 files changed, 19 insertions(+), 3 deletions(-)
diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index a7429a30ec28f903..e145124d6c9a1097 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -43,6 +43,11 @@ OPTIONS
present progress information, or when --quiet is used, that asks
for no messages at all.
+--no-progress::
+ Do not show progress while processing the perf.data file. It
+ also turns off the progress the TUI and GTK browsers present,
+ which is their own.
+
-n::
--show-nr-samples::
Show the number of samples for each symbol
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 81e13fb8194976c9..29085fdc91e58884 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -88,6 +88,7 @@ struct report {
#endif
bool use_stdio;
bool progress;
+ bool progress_set;
bool show_full_info;
bool show_threads;
bool inverted_callchain;
@@ -1385,8 +1386,8 @@ int cmd_report(int argc, const char **argv)
"Use the stdio interface"),
OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
"Show or hide weight columns in annotation. Default show if non-zero."),
- OPT_BOOLEAN(0, "progress", &report.progress,
- "Show progress while processing the perf.data file"),
+ OPT_BOOLEAN_SET(0, "progress", &report.progress, &report.progress_set,
+ "Show progress while processing the perf.data file"),
OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
OPT_BOOLEAN(0, "header-only", &report.header_only,
"Show only data header."),
@@ -1793,7 +1794,9 @@ int cmd_report(int argc, const char **argv)
else
use_browser = 0;
- if (report.progress && !quiet && use_browser == 0)
+ if (report.progress_set && !report.progress)
+ ui_progress__noop_init();
+ else if (report.progress && !quiet && use_browser == 0)
stdio_progress__init();
if (report.data_type && use_browser == 1) {
diff --git a/tools/perf/ui/progress.c b/tools/perf/ui/progress.c
index 99d60223c74b2957..362680989ace606a 100644
--- a/tools/perf/ui/progress.c
+++ b/tools/perf/ui/progress.c
@@ -13,6 +13,12 @@ static struct ui_progress_ops null_progress__ops =
struct ui_progress_ops *ui_progress__ops = &null_progress__ops;
+/* Everything counts but nothing is shown, the way it starts out. */
+void ui_progress__noop_init(void)
+{
+ ui_progress__ops = &null_progress__ops;
+}
+
void ui_progress__update(struct ui_progress *p, u64 adv)
{
u64 last = p->curr;
diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
index 03f1a8bb260ba076..e8c4f9f768aaf12b 100644
--- a/tools/perf/ui/progress.h
+++ b/tools/perf/ui/progress.h
@@ -25,6 +25,8 @@ void ui_progress__update(struct ui_progress *p, u64 adv);
void stdio_progress__init(void);
+void ui_progress__noop_init(void);
+
struct ui_progress_ops {
void (*init)(struct ui_progress *p);
void (*update)(struct ui_progress *p);
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 8/8] perf test: Add false_sharing workload exhibiting cross-CPU false sharing
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
` (6 preceding siblings ...)
2026-10-06 14:57 ` [PATCH 7/8] perf report: Add --no-progress option Arnaldo Carvalho de Melo
@ 2026-10-06 14:57 ` Arnaldo Carvalho de Melo
2026-10-07 0:07 ` [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Namhyung Kim
8 siblings, 0 replies; 10+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-06 14:57 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
Add a 'perf test -w false_sharing' workload that hammers one shared
struct from several CPUs, shaped as a TCP connection: a read-mostly
identity (five-tuple) shares a cacheline with per-packet rx counters
(the false-sharing line), a second line has packet-path private tx and
congestion control counters, and a third the connection config.
The packet path runs in the main thread and up to four lookup threads
sum the five-tuple and pull the config, reading one volatile shared
instance directly so the accesses are PC-relative and resolvable by the
data type profiler.
Data type profiling of the workload with cacheline info:
$ perf mem record -- perf test -r5 -w false_sharing
[ perf record: Woken up 19 times to write data ]
[ perf record: Captured and wrote 5.612 MB perf.data (72454 samples) ]
$ perf report -s type,typecln -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
# Overhead Data Type / Data Type Cacheline
# ................... ...............................
#
88.99% 35.88% struct net_conn
56.43% 17.50% struct net_conn: cache-line 0
32.55% 0.00% struct net_conn: cache-line 2
0.01% 18.38% struct net_conn: cache-line 1
10.08% 63.03% (unknown)
10.08% 63.03% (unknown): cache-line 0
0.82% 0.13% int
0.82% 0.13% int: cache-line 0
0.01% 0.02% struct folio
0.01% 0.02% struct folio: cache-line 0
0.01% 0.03% Elf64_Addr
0.01% 0.03% Elf64_Addr: cache-line 0
0.01% 0.03% struct sched_entity
0.01% 0.03% struct sched_entity: cache-line 1
0.00% 0.00% struct sched_entity: cache-line 4
0.00% 0.00% struct sched_entity: cache-line 2
0.00% 0.00% struct sched_entity: cache-line 3
0.00% 0.00% struct sched_entity: cache-line 0
0.01% 0.00% struct task_group
0.01% 0.00% struct task_group: cache-line 4
0.00% 0.00% struct task_group: cache-line 5
0.00% 0.00% struct css_rstat_cpu
0.00% 0.00% struct css_rstat_cpu: cache-line 0
$ perf report -s type,typecln,typeoff -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
# Overhead Data Type / Data Type Cacheline / Data Type Offset
# ...................... ..................................................
#
88.99% 35.88% struct net_conn
56.43% 17.50% struct net_conn: cache-line 0
9.81% 0.00% struct net_conn +0xa (dport)
9.54% 0.00% struct net_conn +0x8 (sport)
9.53% 0.00% struct net_conn +0xc (state)
9.48% 0.00% struct net_conn +0xd (protocol)
9.27% 0.00% struct net_conn +0x4 (daddr)
8.79% 0.00% struct net_conn +0 (saddr)
0.00% 12.08% struct net_conn +0x10 (bytes_rx)
0.00% 1.71% struct net_conn +0x18 (packets_rx)
0.00% 2.80% struct net_conn +0x20 (rx_queue)
0.00% 0.91% struct net_conn +0x3c (last_ack)
32.55% 0.00% struct net_conn: cache-line 2
5.70% 0.00% struct net_conn +0x83 (rcv_wscale)
5.46% 0.00% struct net_conn +0x84 (keepalive_int)
5.43% 0.00% struct net_conn +0x82 (snd_wscale)
5.39% 0.00% struct net_conn +0x80 (mss)
5.34% 0.00% struct net_conn +0x88 (mark)
5.23% 0.00% struct net_conn +0x8c (priority)
0.01% 18.38% struct net_conn: cache-line 1
0.01% 0.05% struct net_conn +0x5c (retrans)
0.00% 4.83% struct net_conn +0x40 (bytes_tx)
0.00% 10.63% struct net_conn +0x48 (packets_tx)
0.00% 1.54% struct net_conn +0x58 (rtt_us)
0.00% 1.07% struct net_conn +0x50 (cwnd)
0.00% 0.25% struct net_conn +0x54 (ssthresh)
10.08% 63.03% (unknown)
10.08% 63.03% (unknown): cache-line 0
10.08% 63.03% (unknown)
0.82% 0.13% int
0.82% 0.13% int: cache-line 0
0.82% 0.13% int +0 (no field)
0.01% 0.02% struct folio
0.01% 0.02% struct folio: cache-line 0
0.01% 0.00% struct folio +0 (flags.f)
0.00% 0.01% struct folio +0x34 (_refcount.counter)
0.00% 0.00% struct folio +0x18 (mapping)
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/tests/builtin-test.c | 1 +
tools/perf/tests/shell/data_type_profiling.sh | 9 +-
tools/perf/tests/tests.h | 1 +
tools/perf/tests/workloads/Build | 2 +
tools/perf/tests/workloads/false_sharing.c | 251 ++++++++++++++++++
5 files changed, 262 insertions(+), 2 deletions(-)
create mode 100644 tools/perf/tests/workloads/false_sharing.c
diff --git a/tools/perf/tests/builtin-test.c b/tools/perf/tests/builtin-test.c
index d2f594921e25bda9..98134b1c74cf80f4 100644
--- a/tools/perf/tests/builtin-test.c
+++ b/tools/perf/tests/builtin-test.c
@@ -175,6 +175,7 @@ static struct test_workload *workloads[] = {
&workload__context_switch_loop,
&workload__deterministic,
&workload__callchain,
+ &workload__false_sharing,
#ifdef HAVE_RUST_SUPPORT
&workload__code_with_type,
diff --git a/tools/perf/tests/shell/data_type_profiling.sh b/tools/perf/tests/shell/data_type_profiling.sh
index a916c410274aa888..57203e0e5f857540 100755
--- a/tools/perf/tests/shell/data_type_profiling.sh
+++ b/tools/perf/tests/shell/data_type_profiling.sh
@@ -8,8 +8,8 @@ set -e
# data type profiling manifestation
# Values in testtypes and testprogs should match
-testtypes=("# data-type: struct Buf" "# data-type: struct buf")
-testprogs=("perf test -w code_with_type" "perf test -w datasym")
+testtypes=("# data-type: struct Buf" "# data-type: struct buf" "# data-type: struct net_conn")
+testprogs=("perf test -w code_with_type" "perf test -w datasym" "perf test -w false_sharing")
err=0
perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX)
@@ -59,6 +59,9 @@ test_basic_annotate() {
"xC")
index=1 ;;
+
+ "xFS")
+ index=2 ;;
esac
# Under 'set -e' a bare failing command aborts the script through the EXIT
@@ -114,6 +117,8 @@ test_basic_annotate Basic Rust
test_basic_annotate Pipe Rust
test_basic_annotate Basic C
test_basic_annotate Pipe C
+test_basic_annotate Basic FS
+test_basic_annotate Pipe FS
cleanup
exit $err
diff --git a/tools/perf/tests/tests.h b/tools/perf/tests/tests.h
index 9c96f33483d14356..f379620069d34f3a 100644
--- a/tools/perf/tests/tests.h
+++ b/tools/perf/tests/tests.h
@@ -250,6 +250,7 @@ DECLARE_WORKLOAD(jitdump);
DECLARE_WORKLOAD(context_switch_loop);
DECLARE_WORKLOAD(deterministic);
DECLARE_WORKLOAD(callchain);
+DECLARE_WORKLOAD(false_sharing);
#ifdef HAVE_RUST_SUPPORT
DECLARE_WORKLOAD(code_with_type);
diff --git a/tools/perf/tests/workloads/Build b/tools/perf/tests/workloads/Build
index 048e371eb63e3164..18fceddccb56c013 100644
--- a/tools/perf/tests/workloads/Build
+++ b/tools/perf/tests/workloads/Build
@@ -14,6 +14,7 @@ perf-test-y += jitdump.o
perf-test-y += context_switch_loop.o
perf-test-y += deterministic.o
perf-test-y += callchain.o
+perf-test-y += false_sharing.o
ifeq ($(CONFIG_RUST_SUPPORT),y)
perf-test-y += code_with_type.o
@@ -29,3 +30,4 @@ CFLAGS_inlineloop.o = -g -O2
CFLAGS_deterministic.o = -g -O0 -fno-inline -U_FORTIFY_SOURCE
CFLAGS_named_threads.o = -g -O0 -fno-inline -U_FORTIFY_SOURCE
CFLAGS_callchain.o = -g -O0 -fno-inline -U_FORTIFY_SOURCE
+CFLAGS_false_sharing.o = -g -O0 -fno-inline -U_FORTIFY_SOURCE
diff --git a/tools/perf/tests/workloads/false_sharing.c b/tools/perf/tests/workloads/false_sharing.c
new file mode 100644
index 0000000000000000..e949bb36646a97a5
--- /dev/null
+++ b/tools/perf/tests/workloads/false_sharing.c
@@ -0,0 +1,251 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * False-sharing demo for data type profiling, shaped as a TCP
+ * connection: a read-mostly identity shares a cacheline with per-packet
+ * rx counters (the false-sharing line), a second line has packet-path
+ * private tx and congestion control counters, and a third the connection
+ * config.
+ *
+ * 'perf mem record' of this workload followed by 'perf report -s type'
+ * (see tests/shell/data_type_profiling.sh) resolves the accesses to
+ * struct net_conn members, showing the rx counters and the read-mostly
+ * identity sharing cacheline 0.
+ */
+#include <pthread.h>
+#include <sched.h>
+#include <stdint.h>
+#include <stdlib.h>
+#include <stdio.h>
+#include <signal.h>
+#include <unistd.h>
+#include <linux/compiler.h>
+#include "../tests.h"
+
+struct net_conn {
+ /* cacheline 0: identity (read-mostly) + rx counters (per packet) */
+ uint32_t saddr; /* 0 */
+ uint32_t daddr; /* 4 */
+ uint16_t sport; /* 8 */
+ uint16_t dport; /* 10 */
+ uint8_t state; /* 12: 1 == ESTABLISHED */
+ uint8_t protocol; /* 13: 6 == TCP */
+ uint16_t __pad0; /* 14 */
+ uint64_t bytes_rx; /* 16: every packet */
+ uint64_t packets_rx; /* 24: every packet */
+ uint32_t rx_queue; /* 32: backlog depth, fluctuates */
+ uint8_t __pad1[24]; /* 36..59 */
+ uint32_t last_ack; /* 60: written per ACK */
+ /* cacheline 1: tx + congestion control (packet-path private) */
+ uint64_t bytes_tx; /* 64: every packet */
+ uint64_t packets_tx; /* 72: every packet */
+ uint32_t cwnd; /* 80: on every ACK */
+ uint32_t ssthresh; /* 84: on loss */
+ uint32_t rtt_us; /* 88: on every ACK */
+ uint32_t retrans; /* 92: on timeout */
+ uint32_t __pad2[8]; /* 96..127 */
+ /* cacheline 2: config, set at setup, read by everybody */
+ uint16_t mss; /* 128 */
+ uint8_t snd_wscale; /* 130 */
+ uint8_t rcv_wscale; /* 131 */
+ uint32_t keepalive_int; /* 132 */
+ uint32_t mark; /* 136: firewall mark */
+ uint32_t priority; /* 140: traffic class */
+ uint32_t __pad3[12]; /* 144..191 */
+} __attribute__((aligned(64)));
+
+/* Volatile so every iteration really loads and stores. */
+static volatile struct net_conn conn;
+/* Keeps the reader checksums alive after the threads join. */
+static volatile unsigned long fs_sink;
+
+static volatile sig_atomic_t done;
+
+/*
+ * One cacheline each: sum before cpu, or the implicit padding after cpu
+ * pushes the struct past 64 bytes, and aligning it to a cacheline then
+ * rounds it up to 128.
+ */
+struct fs_reader {
+ pthread_t thread;
+ unsigned long sum;
+ int cpu;
+ char __pad[64 - sizeof(pthread_t) - sizeof(unsigned long) - sizeof(int)];
+} __attribute__((aligned(64)));
+
+static void sighandler(int sig __maybe_unused)
+{
+ done = 1;
+}
+
+static void pin_to_cpu(int cpu)
+{
+ cpu_set_t set;
+
+ /* There may be no second CPU in a restricted cpuset. */
+ if (cpu < 0)
+ return;
+
+ CPU_ZERO(&set);
+ CPU_SET(cpu, &set);
+ /* Best effort: in a restricted cpuset this fails and the thread runs unpinned. */
+ pthread_setaffinity_np(pthread_self(), sizeof(set), &set);
+}
+
+/*
+ * Connection lookup, as a load balancer or 'ss' scrape would do it: the
+ * reads go straight to the global, this file is built -O0 and a local
+ * pointer would be reloaded from a stack slot, a form the data type
+ * resolver does not track.
+ */
+static void *reader_fn(void *arg)
+{
+ struct fs_reader *r = arg;
+ unsigned long sum = 0;
+
+ pthread_setname_np(pthread_self(), "fs-reader");
+ pin_to_cpu(r->cpu);
+
+ while (!done) {
+ sum += conn.saddr + conn.daddr + conn.sport + conn.dport +
+ conn.state + conn.protocol;
+ sum += conn.mss + conn.snd_wscale + conn.rcv_wscale +
+ conn.keepalive_int + conn.mark + conn.priority;
+ }
+ r->sum = sum;
+ return NULL;
+}
+
+static int false_sharing(int argc, const char **argv)
+{
+ double sec = 2.0;
+ int nreaders = 0, nr_allowed = 0, err = 1;
+ int *allowed = NULL, nallowed = 0;
+ cpu_set_t set;
+ int nr_mask_bits = sizeof(set) * 8 < CPU_SETSIZE ? sizeof(set) * 8 : CPU_SETSIZE;
+ struct fs_reader *readers = NULL;
+ int i, writer_cpu;
+ unsigned long n = 0;
+
+ pthread_setname_np(pthread_self(), "fs-writer");
+ if (argc > 0)
+ sec = atof(argv[0]);
+ if (!(sec > 0.0)) {
+ fprintf(stderr, "Error: seconds (%f) must be > 0\n", sec);
+ return 1;
+ }
+ if (argc > 1)
+ nreaders = atoi(argv[1]);
+
+ /*
+ * A connection that just got established: identity and config fixed
+ * from here on, counters at zero.
+ */
+ conn.saddr = 0x0a000001; /* 10.0.0.1 */
+ conn.daddr = 0x0a000002; /* 10.0.0.2 */
+ conn.sport = 54321;
+ conn.dport = 443;
+ conn.state = 1; /* ESTABLISHED */
+ conn.protocol = 6; /* TCP */
+ conn.mss = 1448;
+ conn.snd_wscale = 7;
+ conn.rcv_wscale = 7;
+ conn.keepalive_int = 7200;
+ conn.cwnd = 10;
+ conn.ssthresh = 65535;
+ conn.rtt_us = 50;
+
+ /*
+ * Pin against the allowed set, restricted cpusets still spread the threads.
+ * The whole mask is looked at, not the CPU count: the count can be lower
+ * than the highest ID in it, as when a cpuset allows only high numbered
+ * CPUs, and then no allowed CPU would be found at all.
+ */
+ if (sched_getaffinity(0, sizeof(set), &set) == 0) {
+ for (i = 0; i < nr_mask_bits; i++) {
+ if (!CPU_ISSET(i, &set))
+ continue;
+ nr_allowed++;
+ }
+ allowed = malloc(nr_allowed * sizeof(int));
+ if (allowed == NULL) {
+ fprintf(stderr, "Error: malloc failed for CPU list\n");
+ return 1;
+ }
+ for (i = 0; i < nr_mask_bits; i++) {
+ if (CPU_ISSET(i, &set))
+ allowed[nallowed++] = i;
+ }
+ }
+ if (nreaders <= 0) {
+ /* By default leave one CPU for the packet path, up to 4 readers. */
+ nreaders = nallowed > 1 ? nallowed - 1 : 1;
+ if (nreaders > 4)
+ nreaders = 4;
+ }
+
+ signal(SIGINT, sighandler);
+ signal(SIGALRM, sighandler);
+
+ readers = calloc(nreaders, sizeof(*readers));
+ if (readers == NULL) {
+ fprintf(stderr, "Error: calloc failed for %d readers\n", nreaders);
+ goto out;
+ }
+ for (i = 0; i < nreaders; i++) {
+ int cpu = nallowed > 1 ? allowed[(i + 1) % nallowed] : -1;
+
+ readers[i].cpu = cpu;
+ if (pthread_create(&readers[i].thread, NULL, reader_fn, &readers[i])) {
+ fprintf(stderr, "Error: failed to create reader %d\n", i);
+ done = 1; // Ensure started threads terminate.
+ nreaders = i;
+ goto out_join;
+ }
+ }
+ writer_cpu = nallowed > 0 ? allowed[0] : -1;
+ if (nallowed == 1)
+ fprintf(stderr, "Warning: single CPU allowed, no cross-CPU traffic expected\n");
+ if (writer_cpu >= 0)
+ pin_to_cpu(writer_cpu);
+
+ /*
+ * The packet path: receive, acknowledge, transmit, repeat; every 64th
+ * packet simulates a loss so ssthresh/retrans get sampled too.
+ */
+ if (sec < 1.0) {
+ useconds_t usecs = (useconds_t)(sec * 1000000.0);
+
+ ualarm(usecs > 0 ? usecs : 1, 0);
+ } else
+ alarm((unsigned int)sec);
+ while (!done) {
+ conn.bytes_rx += conn.mss;
+ conn.packets_rx++;
+ conn.rx_queue = (uint32_t)(n & 0x3f);
+ conn.last_ack = (uint32_t)n;
+ conn.bytes_tx += conn.mss;
+ conn.packets_tx++;
+ conn.cwnd = 10 + (n & 15);
+ conn.rtt_us = 50 + (n & 7);
+ if ((n & 63) == 0) {
+ conn.ssthresh = conn.cwnd / 2;
+ conn.retrans++;
+ }
+ n++;
+ }
+ err = 0;
+out_join:
+ for (i = 0; i < nreaders; i++) {
+ if (readers[i].thread) {
+ pthread_join(readers[i].thread, /*retval=*/NULL);
+ fs_sink += readers[i].sum;
+ }
+ }
+ fs_sink += (unsigned long)(conn.bytes_rx + conn.bytes_tx + n);
+ free(readers);
+out:
+ free(allowed);
+ return err;
+}
+
+DEFINE_WORKLOAD(false_sharing);
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
` (7 preceding siblings ...)
2026-10-06 14:57 ` [PATCH 8/8] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
@ 2026-10-07 0:07 ` Namhyung Kim
8 siblings, 0 replies; 10+ messages in thread
From: Namhyung Kim @ 2026-10-07 0:07 UTC (permalink / raw)
To: Arnaldo Carvalho de Melo
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users
On Tue, Oct 06, 2026 at 04:57:21PM +0200, Arnaldo Carvalho de Melo wrote:
> Hi,
>
> This series adds progress diagnostics to perf, and a workload that makes
> false sharing visible to data type profiling.
>
> The changes are:
>
> - move perf_config__set_variable() to util/config.c and serialize config
> parser and read-modify-write state, so non-builtin perf code can persist
> configuration changes safely;
>
> - add 'perf report --progress' for stdio users, showing the current phase,
> percentage, and counts while a session is processed;
>
> - wire up 'perf report --no-progress', the counterpart of the option
> above, for the TUI and GTK browsers, whose progress there is no other
> way to turn off;
>
> - add 'perf test -w false_sharing', a synthetic TCP-shaped workload with
> identity and packet counters sharing a cacheline, and include it in the
> data type profiling shell test.
>
> Follow up work: the DO_ONCE() one-time init primitive added here mirrors
> what the eight pre-existing raw pthread_once() users in util/ (annotate.c,
> callchain.c, comm.c, dso.c, fncache.c, intel-tpebs.c, libbfd.c, pmus.c)
> need, conversions that will also exercise the DEFINE_MUTEX() static
> initializer; converting them is left for after this series.
> Best regards,
Reviewed-by: Namhyung Kim <namhyung@kernel.org>
Thanks,
Namhyung
>
> - Arnaldo
>
> What changed from v11:
>
> - The 'perf-stuck' patch is removed for the time being, to be submitted
> separately later on;
>
> What changed from v10:
>
> - tools/perf/scripts/perf-stuck.sh: pin the watched process by its start
> time, so a recycled PID can't get samples or gdb; Sashiko, v10;
>
> - tools/perf/scripts/perf-stuck.sh: drop the control characters from the
> shown progress updates, no terminal escape injection; Sashiko, v10;
>
> - tools/perf/tests/shell/perf_stuck.sh: test the script, from option
> validation to the gdb DIE chain runs on stand-in workloads;
>
> What changed from v9:
>
> - tools/perf/ui/stdio/progress.c: progress updates now work well with
> the pager, printed to /dev/tty so they update in place;
>
> - tools/perf/scripts/perf-stuck.sh: take the last progress update
> from a bounded tail window, dropping the \r's. Sashiko, v9;
>
> What changed from v8:
>
> - tools/perf/util/mutex.h: DEFINE_MUTEX() statics keep mutex_init()'s
> !NDEBUG errorcheck type;
>
> - patch 4/9 commit log: clarify the acquire/release pairing, what the
> dropped static key used to guarantee.
>
> What changed from v7:
>
> - tools/perf/util/mutex.h: the DO_ONCE() lockless fast path pairs its
> __ATOMIC_ACQUIRE load with the __ATOMIC_RELEASE store. sashiko, v7;
>
> What changed from v6:
>
> - patch 1/5 split into four: the move, perf_etc_perfconfig() never
> returning NULL, the DEFINE_MUTEX() prep and the serialization.
> Namhyung Kim asked, v6;
>
> - tools/perf/util/{config.c,mutex.h}: perf's own mutex type, adding
> the DEFINE_MUTEX() initializer it lacked. Namhyung Kim, v6;
>
> - tools/perf/util/mutex.h: add DO_ONCE() one-time init, from the
> kernel's once.h, moving the lazy inits in config.c to it;
>
> - tools/perf/builtin-report.c: drop the option negation and stdio
> hook comments Namhyung Kim asked to drop, reviewing v6;
>
> - tools/perf/scripts/perf-stuck.gdb: break the long printf() line.
> Namhyung Kim, reviewing v6;
>
> - the false_sharing cset now carries the data type profiling output
> with cacheline info. Namhyung Kim asked for it, reviewing v6.
>
> What changed from v5:
>
> - tools/perf/util/config.c: drop the config_file_name read in
> bad_config(), locking would buy a consistent NULL. Sashiko, v5;
>
> - tools/perf/util/config.c: new config_set_mutex around the shared
> set's init/teardown, home init via pthread_once. Sashiko, v5;
>
> - tools/perf/builtin-report.c, perf-report.txt: --no-progress is
> --progress's auto negation, last wins. Namhyung Kim, reviewing v5.
>
> What changed from v4:
>
> - tools/perf/builtin-report.c: mark --progress PARSE_OPT_NOAUTONEG,
> parse_long_opt() claimed --no-progress first. Sashiko, v4;
>
> - tools/perf/util/config.c: format the path buffer inside the critical
> section. Sashiko pointed out the window, reviewing v4;
>
> - tools/perf/tests/workloads/false_sharing.c: walk every bit of the
> affinity mask, not _SC_NPROCESSORS_CONF. Sashiko, reviewing v4.
>
> What changed from v3:
>
> - tools/perf/util/config.c: keep the buffer handed to the parser in
> static storage. Sashiko, reviewing v3;
>
> - tools/perf/scripts/perf-stuck.gdb: perf-dso walks each dso candidate
> until one evaluates, for REFCNT_CHECKING builds. Sashiko, v3;
>
> - tools/perf/tests/workloads/false_sharing.c: sum before cpu in struct
> fs_reader, the padding after cpu rounded it to 128. Sashiko, v3;
>
> - tools/perf/builtin-report.c: --quiet wins over --progress, the phases
> stay uncounted, documented next to it. Namhyung Kim, v3;
>
> - tools/perf/builtin-report.c, tools/perf/ui/progress.c: --no-progress
> installs the no-op ops. Suggested by Namhyung Kim, reviewing v3;
>
> - tools/perf/scripts/perf-stuck.sh: check gdb is there before
> watching. Namhyung Kim, reviewing v3.
>
> What changed from v2:
>
> - tools/perf/util/config.c: make perf_etc_perfconfig() total, falling
> back to the unresolved path. Sashiko, reviewing v2;
>
> - tools/perf/scripts/perf-stuck.gdb: don't deref map_symbol.sym without
> a NULL check. Sashiko, reviewing v2;
>
> - tools/perf/scripts/perf-stuck.gdb: note the REFCNT_CHECKING proxy
> next to the structure walks;
>
> - tools/perf/scripts/perf-stuck.sh: count samples with no progress to
> look at, an empty log never fired -g;
>
> - tools/perf/scripts/perf-stuck.sh: `--` for pgrep and tail against
> names starting with a hyphen;
>
> - tools/perf/scripts/perf-stuck.sh: bound the gdb run with
> `timeout --signal=INT 30`.
>
> What changed from v1:
>
> - avoid calling CPU_SET() with -1 when false_sharing runs with only one
> CPU available in its affinity mask.
>
> tools/perf/Documentation/perf-report.txt | 19 ++
> tools/perf/builtin-config.c | 70 +----
> tools/perf/builtin-report.c | 9 +
> tools/perf/tests/builtin-test.c | 1 +
> tools/perf/tests/shell/data_type_profiling.sh | 9 +-
> tools/perf/tests/tests.h | 1 +
> tools/perf/tests/workloads/Build | 2 +
> tools/perf/tests/workloads/false_sharing.c | 251 ++++++++++++++++++
> tools/perf/ui/Build | 1 +
> tools/perf/ui/progress.c | 6 +
> tools/perf/ui/progress.h | 4 +
> tools/perf/ui/stdio/progress.c | 170 ++++++++++++
> tools/perf/util/config.c | 166 ++++++++++--
> tools/perf/util/config.h | 2 +
> tools/perf/util/mutex.h | 36 +++
> tools/perf/util/ordered-events.c | 16 +-
> tools/perf/util/session.c | 12 +-
> 17 files changed, 675 insertions(+), 100 deletions(-)
> create mode 100644 tools/perf/tests/workloads/false_sharing.c
> create mode 100644 tools/perf/ui/stdio/progress.c
> base-commit: 705da5b15ab89ba9
> v1-head: 45d7917f7e05e8a29828ed5f0bbdc94fc938f79f
> v2-head: d4f84e4de8890194924ccd897a8e6773e7d4240b
> v3-head: 485532296710225862daf8ffb19802ae327efe2a
> v4-head: 38193635508e0a5f04a6ebf70db27534b68b2d5b
> v5-head: e99bc48f6c72b9cc67b4e4143a0364acac5523bb
> v6-head: 7dc99dbb6504f199c3c39487d91370d1bbb7a827
> v7-head: b357a76666d08f74eb99c755cca41855a96f6169
> v8-head: 257b5ea6115ad5126d2c881523ea6c660cf5d28c
> v9-head: 70235104c6cf7d18d91f46de1a233abbeb65893a
> v10-head: 1ecf850c2694d20ca9cd0014f2a42528ef4dbfec
> v11-head: a3ddbaab18aa24fba520927f71d0ea66edfc5593
> --
^ permalink raw reply [flat|nested] 10+ messages in thread
end of thread, other threads:[~2026-10-07 0:07 UTC | newest]
Thread overview: 10+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 14:57 [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 1/8] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 2/8] perf config: Make perf_etc_perfconfig() never return NULL Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 3/8] perf mutex: Add DEFINE_MUTEX() static initializer Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 4/8] perf mutex: Add DO_ONCE() for one-time initialization Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 5/8] perf config: Serialize config file access with a mutex Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 6/8] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 7/8] perf report: Add --no-progress option Arnaldo Carvalho de Melo
2026-10-06 14:57 ` [PATCH 8/8] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
2026-10-07 0:07 ` [PATCH v12 0/8] perf tools: Add progress diagnostics and a false-sharing workload Namhyung Kim
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®