From: Changbin Du <changbin.du@gmail.com>
To: Namhyung Kim <namhyung@kernel.org>
Cc: Changbin Du <changbin.du@gmail.com>,
Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
Arnaldo Carvalho de Melo <acme@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
Alexander Shishkin <alexander.shishkin@linux.intel.com>,
Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
Adrian Hunter <adrian.hunter@intel.com>,
James Clark <james.clark@linaro.org>,
linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org
Subject: Re: [PATCH] perf bench: Add atomic CAS benchmark
Date: Wed, 7 Oct 2026 13:27:36 +0800 [thread overview]
Message-ID: <asXXZJzEPoEsMSni@mail.google.com> (raw)
In-Reply-To: <asQlFvclL-vj5xlh@google.com>
Hello,
On Mon, Oct 05, 2026 at 03:30:46PM -0700, Namhyung Kim wrote:
> Hello,
>
> On Wed, Sep 30, 2026 at 05:16:17PM +0800, Changbin Du wrote:
> > Add a new 'atomic' collection to perf bench for benchmarking
> > compare-and-swap (CAS) atomic operations with multi-threaded
> > contention testing.
> >
> > The benchmark tests __atomic_compare_exchange_n operations
> > with configurable thread count and iteration count to measure
> > atomic contention effects.
> >
> > Why this benchmark is needed:
> > - CAS operations are fundamental to lock-free algorithms and data
> > structures. Understanding their performance characteristics under
> > contention is critical for designing high-performance concurrent
> > applications.
> > - The benchmark helps identify atomic operation latency and
> > scalability issues across different thread counts, revealing
> > contention patterns that are not visible in single-threaded tests.
> > - Useful for evaluating atomic implementation quality on different
> > architectures and for regression testing after changes to atomic
> > primitives or memory ordering.
> >
> > Measurement methodology:
> > - Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
> > after synchronizing on a pthread_barrier, ensuring all threads
> > begin simultaneously.
> > - Each thread performs a hot loop of atomic compare-and-swap on a
> > shared u64 counter, incrementing from 0 to iterations.
> > - The shared counter is cache-line aligned (64 bytes) to isolate
> > contention to the target cache line and avoid false sharing.
> > - The wall-clock time is measured as the max of all per-thread
> > runtimes (the time for the slowest thread to finish).
> > - The first repeat is excluded from statistics as a warmup phase
> > to avoid cache-cold effects.
> > Example usage:
> > $ perf bench atomic cas --threads 2
> > # Running 'atomic/cas' benchmark:
> >
> > Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
> > Avg wall-clock time: 7365.480 msec (stddev 66.014 msec)
> > Total ops: 200,000,000
> > Throughput total: 27,153,697 ops/sec
> > Per-thread times and throughput (last repeat):
> > fastest: 7510.031 msec (13315525 ops/sec)
> > slowest: 7581.678 msec (13189692 ops/sec)
> > avg: 7545.854 msec (13252310 ops/sec)
> >
> > Output fields explained:
> > - Threads: number of contending threads
> > - iterations/thread: CAS operations each thread performs
> > - repeats: number of test runs (first is warmup)
> > - Avg wall-clock time: mean time for all threads to complete
> > - stddev: standard deviation across repeats
> > - Total ops: threads x iterations/thread
> > - Throughput total: aggregate ops/sec across all threads
> > - Per-thread times: fastest/slowest/avg thread completion time
> > - Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
>
> Thanks for the contribution! I think it's very useful.
> Just a few suggestions.
>
> 1. it'd be nice to add simple atomic_inc benchmark too.
> 2. it'd be nice to have an option to try other ordering requirements
> than "relaxed".
>
> Thanks,
> Namhyung
Thanks for the review!
1. Done in v2. The collection now provides atomic inc alongside
cas, measuring contended __atomic_fetch_add() throughput with the
same skeleton (barrier-synchronized start, wall-clock taken from the
slowest thread, warmup repeat). Note that its ops/sec is not
instruction-level comparable with cas — cas counts successful
compare-and-swaps only, not the retries and loads in between — which
the documentation now points out.
2. Adding memory-order variants is not necessary because ordering
has no semantic role in this benchmark. Memory ordering exists to
constrain the visibility of other memory operations relative to an
atomic access; its cost and effect only become meaningful in an
algorithm with additional accesses to order — for example, ordering
the initialization stores of a new node before publishing a pointer
with a release CAS. This benchmark, however, measures a single shared
counter with no other memory operations in the loop, so a stronger
ordering would not order anything of consequence: it would only
measure the marginal cost of the stronger instruction itself. Those
numbers would say nothing about how orderings behave in real
workloads, and would therefore add a configuration knob that invites
misleading comparisons rather than useful measurement.
prev parent reply other threads:[~2026-10-07 5:28 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 9:16 Changbin Du
2026-10-05 22:30 ` Namhyung Kim
2026-10-07 5:27 ` Changbin Du [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asXXZJzEPoEsMSni@mail.google.com \
--to=changbin.du@gmail.com \
--cc=acme@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=alexander.shishkin@linux.intel.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=mingo@redhat.com \
--cc=namhyung@kernel.org \
--cc=peterz@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®