On Fri, 2026-10-02 at 14:44 +0100, David Woodhouse wrote: > > Hah, I should stop predicting results before I have them. I'm clearly > not very prescient. Spinning on the MMIO (the GPIO wait_for_edge() > method idea) makes more of a difference than I thought (~30% off every > column brings us down to only 3× the entry.S capture): > >   ┌──────────────────────────────┬─────┬──────┬──────┬──────┬─────┐ >   │ capture method               │ p50 │  p95 │  p99 │  max │   σ │ >   ├──────────────────────────────┼─────┼──────┼──────┼──────┼─────┤ >   │ pps-gpio (IRQ)               │  84 │ 1206 │ 4013 │ 4889 │ 688 │ >   │ polling, stamp after edge    │ 206 │  536 │  665 │ 1106 │ 286 │ >   │ polling, bracketed counter   │ 206 │  545 │  647 │  811 │ 286 │ >   │ polling, bracketed, raw MMIO │ 135 │  359 │  434 │  564 │ 188 │ >   │ entry.S counter capture      │  46 │  109 │  135 │  206 │  58 │ >   └──────────────────────────────┴─────┴──────┴──────┴──────┴─────┘ > > (NB: We should be careful not to forget the common-mode hardware > latency which could be different for the IRQ paths vs. the polling > paths. On my list is a test where one CPU polls while the other takes > the interrupt, so we can compare on the *same* pulse.) I did this test, on a live system rather than a tweaked-for-idleness initramfs environment which was going to flatter the IRQ path massively. It compares the cycle count from the GPIO spin on one CPU, with the entry.S (and, eventually, pps_gpio_irq_hardirq()) on the other. Full results at https://david.woodhou.se/ntptest-r64/capture-latency/ Running idle-ish, with schedutil disabled, the poll captures the edge sooner (-200ns), but with higher jitter (σ 150 ns vs. 115 ns for entry.S). Running under load, we see a lot more jitter even on the entry.S capture (σ 7.6µs). And some captures are seen *before* the polled pulse, presumably because something else asserted the IRQ line first and then the GPIO interrupt was present by the time it was checked. Eliminating those, σ 5.8µs under load. All of which brings me back to the case for better filtering. The CDF suggests there are enough *good* pulses — even under load the median entry.S capture is within ~300ns and 95% are within 1.5µs. Ultimately, the common mode offset is unmeasurable and includes things like the antenna length. The *jitter* has a hard floor for the polling method, based on the time it takes to actually read the GPIO (gpio_read @600ns → σ286ns, direct MMIO @400ns → σ188ns as shown in the table cited above, which is slightly different because it's phase offset against what hardpps was tracking). There's no such obvious hard floor for the entry.S capture, but it does require filtering to throw away the outliers.  Intelligently and in *retrospect* building a best fit line for today's data (no hardpps here; just counter captures), we get σ 150-180ns for polling, and 50-280ns for entry.S capture, the latter correlated with load. Polling under load basically does its *own* filtering, because if it misses a pulse when its hrtimer doesn't get to run, that datum just doesn't exist in the first place. So its jitter *is* its hard floor. (Modulo the bus taking a bit longer to do the GPIO read, perhaps). While interrupt capture — however early you do it — is worse either way when the load is high. (My friend made me add "when the load is high" in reviewing for statistical accuracy and defensibility against the results, but it's tautological — the variability is high, when you don't exclude the outliers.) If we need to cope with load (and I think we do), the interrupt capture didn't survive the test. Which is just as well really, because nobody actually wanted to implement it that way. But I think it was worth finding out, while I had the test reg set up to do so. Still looking forward to hardware which captures the counter for itself when the edge occurs.