← Tepna preprints

We were measuring the wrong error: a wearable clock estimator recovers an injected rate to 0.03 %, while one missed beat in a thousand costs 20.8 % of rMSSD

Michal Planicka  ·  corresponding author — Tepna Project

Clock Contract §7 · capture-host · Tepna physiological-signal suite

Draft v4 · August 2026 · Analysis tools: tools/known-clock-recovery.mjs, tools/beat-error-recovery.mjs, tools/beat-capture-recapture.mjs, tools/beat-injection-recovery.mjs · Preregistration: experiments/known-clock/acceptance.json (sha256 b061d279…, hashed before any run) · Substrate: 395 arrival sidecars, 21 nights, 3 devices

Abstract

Background. Consumer wearables carry crystals wrong by parts-per-million, and pipelines that fuse two devices correct for it by comparing each device's counter against a capture host's clock. Every such correction in this suite was validated only against itself: the device-vs-host estimator, the stability curve, and the three-source closure check are mutually consistent but were never compared to a defect known in advance. A shared error would leave all of them green. Methods. We injected exactly-known defects into real recordings and asked the shipped pipeline to recover them, against acceptance criteria written and hashed before any run. Two substrates: 61 Bluetooth arrival streams (395 sidecars, 21 nights, Polar H10 · Polar Verity Sense · Wellue O2Ring) for clock defects — constant offset, frequency error over ±1…±500 ppm, random-walk wander, contiguous and interleaved packet loss, timestamp steps — and 101 real H10 RR trains for beat-detection defects. All perturbations are deterministic; every figure re-runs to the bit. Results. The clock estimator is accurate: frequency recovery is unbiased and linear across three and a half orders of magnitude (median relative error 0.03–0.09 %, worst 1.06 %); a constant offset recovers as exactly 0.000 ppm, being unidentifiable by construction — and a second, aperiodic route to it also fails on this corpus (§3.2b), where the obvious validation for such methods turns out to be vacuous; contiguous packet loss of 50 % moves the rate by exactly 0.0000 ppm. Its blind spots are a 200 ms timestamp step (invisible) and realistic frequency wander (buried under Bluetooth phase noise). But σy(τ) at each recording's own span exceeds the rate being quoted for 100 % of streams under 300 s and 42 % beyond an hour — so the number is usually smaller than its own noise. The larger finding is that this whole family is the wrong thing to worry about. On the beat substrate, one missed beat in a thousand inflates rMSSD by 20.8 % while mean RR moves 0.06 %; a 30 ppm clock error over the same night moves rMSSD by ~0.003 %. Detector error and clock error differ by four orders of magnitude in their effect on the metric that reaches a user, and only the smaller had been measured. A common-mode test across simultaneously captured device pairs puts the host's contribution to the timing residual at r = 0.036 — the host is not the problem either. Measuring the detector. Two no-ground-truth methods were applied to the beat detector itself. Capture–recapture across three independent detectors is unidentifiable on clean recordings — the three agreed on 935 of 986 beats, and perfect agreement is equally consistent with nothing missed and with all three missing the same beats — so the estimator refuses. An artificial-star test does work: planting the subject's own beat into the raw waveform gives a completeness curve with a threshold near 20–30× local noise and no induced false positives, placing on the order of 1 % of real beats at risk. Attenuating real beats in situ then showed that planting into gaps overstates the detector by ~3×, driven by the adaptive threshold rather than by searchback — a correction in the opposite direction to the one we predicted, and one that surfaced a template-subtraction artefact in our own amplitude axis. Conclusion. A pipeline validated only against itself can be accurate, uninformative, and pointed at the wrong failure mode simultaneously. We also report two estimator defects found by the same run, one claim of our own that this work withdrew, and an independent validation of our Allan implementation against allantools to 4.8 × 10⁻¹⁴.

Keywords: clock synchronisation · Allan deviation · wearable sensors · Bluetooth Low Energy · known-answer testing · preregistration · frequency stability · measurement uncertainty

0. Layman overview (delete before submission)

Two wearables recording the same night do not agree on what time it is. Each has a cheap crystal that runs slightly fast or slow, so software measures that error — in parts per million — and corrects for it. We wanted to know whether that measurement is any good.

The honest way to find out is to break the data on purpose by an amount you write down first, then see what the software reports. We did that: we took real overnight recordings, shifted the clocks by exact amounts, deleted packets, inserted jumps — and asked the software to tell us what we had done.

It is very good at it. Told to find a rate error, it finds it to within about a thousandth of its size. Deleting half the data changes its answer by nothing at all.

But the number it produces is usually too small to be real. Every measurement has a noise floor, and for these devices the floor drops the longer you record. On recordings shorter than an hour, the "crystal error" being reported is smaller than the noise it was measured through. Only on recordings longer than an hour, and even then only about half the time, does the number mean what it appears to mean.

Then we did the same thing to heartbeats, and that is where the paper turned. If the software misses just one heartbeat in a thousand — a rate any beat detector would call excellent — the headline variability number that people actually read moves by a fifth. The average heart rate barely flinches, so nothing looks wrong. Meanwhile the clock error we had spent all this effort measuring moves that same number by about a thousandth of a percent.

So: the tool is not broken, the habit of quoting it without saying how long the recording was is — and both of those are small next to the fact that we had built careful instruments for the error that barely matters, because it was the one we could see.

1. Introduction

Multi-device physiological recording rests on an assumption that is rarely tested: that the devices can be placed on a common timebase. A chest strap and an armband sampling the same heart produce beat trains that can only be compared if the relationship between their clocks is known, and each device's clock is a low-cost crystal with a frequency error of tens of parts per million. The standard remedy is to record, alongside every packet, the capture host's own timestamp, and to estimate the device-vs-host rate from the divergence between the two columns.

The estimator used here is described in the suite's Clock Contract §7: a running median of the host-minus-device residual, interpolated between anchors and held flat outside them, with refusals rather than corrections when the implied rate is implausible. It was built against real recordings and is internally consistent with two neighbouring instruments — an overlapping Allan deviation of the same residual, and a three-source closure check across devices.

Internal consistency is the problem. Each of those three instruments is validated against the others; none has been compared with a defect known in advance. An error shared between them — a sign convention, a units factor, a subtraction that removes the quantity of interest — would leave all three agreeing and all three wrong. This is not hypothetical for this suite: an earlier claim about one of these constants was published, challenged, and withdrawn, with the reviewer noting that the evidence was one file and a marginal estimator.

We therefore did the one experiment that can fail in a way the others cannot: inject a known defect, run the production pipeline blind, and compare what comes back with what went in — with the acceptance criteria fixed and hashed beforehand so that a disappointing result cannot be reinterpreted afterwards.

2. Methods

2.1 Substrate

A bedside capture host (Raspberry Pi, chrony-disciplined; source-stratum 1 via a local stratum-1 reference, chrony skew 0.004–0.011 ppm, RMS offset 19.5 µs — the host is rate-trusted by the capture daemon, a correction to draft v1 which reported it as stratum 2 and failing that bar) holds Bluetooth Low Energy links to a Polar H10 chest strap, a Polar Verity Sense armband and a Wellue O2Ring pulse oximeter, and writes one row per received packet carrying both the host stamp and the device's own sensor counter. That pair is the two clocks. We used 395 such sidecars spanning 21 nights; 61 streams carried ≥ 500 packets and entered the analysis, of which 41 had a span exceeding 300 s.

A counter reset inside a session is a different clock, not a clock error, so each stream is reduced to its longest monotonic run before analysis rather than "corrected". One O2Ring session steps backwards by 27 463 s and is split on that basis.

2.2 Injected defects

Six families, all deterministic — no call to a random number generator or a wall clock appears in the tool, so every figure below reproduces exactly. Wander, the one family requiring noise, uses a seeded linear congruential generator.

TargetInjectionExpected recovery
Null controlnonebaseline, exactly
Constant offsethost + 1 / 5 / 60 s0 ppm — see §3.2
Constant frequencydevice × (1 + f), f ∈ ±{1, 10, 100, 500} ppmf
Frequency wanderrandom walk, 0.5 and 5 ppm per stepAllan slope → τ+1/2
Packet loss10 / 30 / 50 %, contiguous and interleavedrate unchanged
Timestamp jump200 and 2000 ms step at mid-recordinglocalised to max step

The sign of the frequency target is worth stating explicitly, because it is the classic route to a confident wrong answer. The injection makes the device run fast by +f; the estimator reports host-minus-device; so the correct recovery is f. Our own first self-test asserted against +f and failed at −198.7 %, which is exactly this sign, doubled. The expected value is now carried in the output record rather than in a comment.

2.3 Independent validation of the stability instrument

Allan deviation underlies §3.5 and §3.6, so its implementation was validated against allantools 2024.06, the reference implementation following Riley (NIST SP 1065). Overlapping ADEV was computed from identical phase series of four synthesised noise types. Agreement is exact to floating-point: worst per-τ relative difference 4.78 × 10⁻¹⁴, and slopes identical to four decimals (white PM −1.0004, white FM −0.5235, random-walk FM +0.3905, drift +1.0000). This rules out a units or convention error in the instrument used to make the paper's central claim.

3. Results

3.1 Null control and determinism

Across all 61 streams the unperturbed re-run reproduces the baseline with a maximum absolute error of 0.00 × 10⁰ ppm — not "within tolerance", but bit-identical, on every stream. A pipeline that recovered a perturbation from an unperturbed recording would have failed here, and nothing else in the design would have caught it.

3.2 A constant offset is unidentifiable, not merely difficult

Offsets of 1, 5 and 60 s produce a rate change of exactly 0.000 ppm on all 61 streams. This is a structural property: the estimator measures divergence relative to its first anchor, so a constant offset is removed before the rate is formed. The consequence is not a limitation to be improved but a statement about what the output means — a parts-per-million figure carries no information about absolute alignment, and any downstream consumer treating a small ppm as evidence that two devices are aligned is reading a quantity that is blind to the question. Recovering exactly zero is the correct behaviour, and it is only visible as a result because the expectation was written down first.

3.2b The second route to a constant offset also fails — and the obvious test for it cannot

§3.2 shows this estimator is blind to constant offset by construction, not that offset is unrecoverable in principle. The standard alternative is to align an aperiodic feature common to both devices — beat trains cannot serve, since two trains one inter-beat interval apart are indistinguishable from two aligned ones. Published designs supply such a feature deliberately (a commanded event, a tap artefact). We asked whether one is already present: a subject turning over in bed produces a transient in both a chest and an arm accelerometer at the same instant.

It is not usable. On the paired overnight recording (chest ECG strap versus arm optical band, 4.75 h of overlap, ~868,000 accelerometer samples per device), cross-correlation of the aperiodic envelopes gives a peak prominence of 0.0017–0.018 against a posture-only null control of 0.002 — indistinguishable. The decisive diagnostic is that the peak moves with the search range:

search rangepeak lagprominence
±4 s3850 ms0.0152
±6 s5750 ms0.0181
±9 s9000 ms (at the boundary)0.0083
±12 / 20 / 30 s9000 ms0.0017

A genuine lock does not move when the window widens. Two accelerometers on different body segments do not register the same impulse closely enough, over a night that is mostly still, to align the recordings.

The methodological result is the transferable one, and it is a warning about how such methods are validated. The natural test — inject a known constant offset and check that it is recovered — cannot fail. A constant time shift translates the entire cross-correlation surface rigidly, so the peak moves by exactly the injected amount whether or not the peak locates anything. On the recording above, whose prominence is 0.0017:

injectedexpectedrecoverederror
−3000 ms600060000
−1000 ms800080000
+1000 ms10000100000
+3000 ms12000120000

Four injections, error exactly zero in each, from a method that locates nothing. Reported alone these numbers are indistinguishable from a validated alignment, and they are more persuasive than a noisy correct result would be. An injected constant offset is therefore not evidence for a correlation-based offset estimator; peak prominence against a null control, and invariance of the peak to the search range, are. Both are asserted in the accompanying tool's self-test.

3.3 Frequency recovery is unbiased across three and a half orders of magnitude

Injected (ppm)nmedian |rel. err|p90worst
−500410.039 %0.451 %0.97 %
−100410.029 %0.490 %1.00 %
−10410.038 %0.500 %1.01 %
−1410.039 %0.503 %1.01 %
+1410.039 %0.504 %1.01 %
+10410.040 %0.502 %1.02 %
+100410.054 %0.510 %1.02 %
+500410.094 %0.550 %1.06 %

The error is flat in magnitude and symmetric in sign — the estimator is linear and unbiased over the whole range, including at ±1 ppm, an injection far below the devices' own rate errors. On the face of it this is a strong endorsement. §3.6 explains why it is not.

3.4 Packet loss: the realistic case costs nothing, the unrealistic case costs everything

LossFractionmedian |Δppm|p90max
contiguous10 %0.00000.0000.00
30 %0.00000.0000.00
50 %0.00000.0000.00
interleaved10 %0.000013.06304.5
30 %0.108251.13943.4
50 %0.940966.991115.1

Removing half of every stream in one contiguous block changes the recovered rate by exactly nothing, on all 61 streams. Removing the same quantity by decimation changes it by up to 1115 ppm. Since a Bluetooth dropout is a contiguous outage — the link fails for an interval — the estimator is robust to the loss that actually occurs and fragile to a pattern that does not. We report both because the preregistered criterion was written against the decimated case and failed; the mechanism was only identified by testing the alternative afterwards. Had the criterion been written after the fact, this would have been recorded as a pass with no insight attached.

3.5 Two sensitivity floors

Timestamp steps. A 200 ms injected step produces a median max-step ratio of 1.0× — it is indistinguishable from the ordinary Bluetooth delivery jitter it hides in, on every stream. A 2000 ms step is localised (median 13.0×) but exceeds the 10× detection criterion on only 51 % of streams. The estimator cannot see a clock step smaller than roughly its own jitter, which for this transport is of order a second.

Frequency wander. Injected random-walk wander of 0.5 and 5 ppm per step moves the median Allan slope from −0.990 to −0.968 and −0.937 respectively — a τ−1 phase-noise process throughout, never approaching the τ+1/2 signature of the injected walk. The noise classifier declines to name a type on an increasing share of streams as the injection grows (42/45 named at 0.5 ppm, 31/45 at 5 ppm), which is the honest response to a slope whose standard error spans two mechanisms. At realistic magnitudes, wander is not recoverable from this substrate: it is buried beneath the phase jitter of the link.

3.6 The central result: the rate is usually smaller than its own uncertainty

σy(τ) evaluated at each recording's own span is the natural uncertainty on a rate quoted from that recording. Comparing it with the rate itself:

Spannmedian |ppm|median σy (ppm)|ppm| / σyfraction resolved
< 300 s4105.92137.50.0500 %
300–3600 s1750.4489.80.1686 %
> 3600 s2428.222.01.31658 %

Per device, restricted to spans over one hour: H10 median |ppm| 21.7 against σy 10.5 (resolved, ratio 2.1); Verity 33.8 against 22.0 (marginal, 1.5); O2Ring 2611.7 against 908.5 — a large ratio that means nothing, because that device's axis is synthesised rather than measured (§3.7).

The estimator's own accuracy (§3.3, 0.03 %) and the quantity's resolvability are independent questions, and only the first has previously been asked. The tool recovers an injected rate almost perfectly because the injection is large and coherent; the device's own rate is neither, and on most recordings it does not rise above the floor. This is the empirical content behind the Clock Contract's existing prose rule that a ppm must never be quoted without its span — a rule previously derived by hand from two anecdotal figures, now given a curve.

3.6b A claim we made from §3.6 and then had to withdraw

A draft of this paper argued that because one node admits a rate correction above a 2400 s span, and because only 6 % of streams in the 300–3600 s band have a rate exceeding their own σy, that threshold was roughly an order of magnitude too permissive. That inference was wrong, and the measurement that overturns it is reported here rather than quietly removed.

Two questions were conflated. Is the rate resolved? — can it be quoted as a measurement distinguishable from zero — is answered by |ppm| vs σy, and the answer under an hour is no. Does applying the correction reduce error? is a different question, and a span gate governs the second, not the first. A point estimate can sit below its own noise floor and still be closer to the truth than assuming zero.

We tested it directly. Taking the full-span rate as truth on streams where that full-span estimate is itself resolved, then truncating the same stream and asking whether the short-span correction lands nearer the truth than no correction at all:

Spannmedian error, correctedmedian error, uncorrectedcorrection helped
2400 s (shipped gate)118.41 ppm22.27 ppm82 %
4800 s114.71 ppm22.27 ppm100 %
9600 s91.93 ppm22.19 ppm100 %

The gate is doing net good where it stands. At 2400 s the correction more than halves the median error and helps on 82 % of cases; the residual 18 % are modest (worst observed: a Verity whose true rate is −34.3 ppm estimated at −77.3 ppm, an error of 43.0 ppm against 34.3 uncorrected). Raising the gate to 4800 s would remove the harm cases at the cost of refusing correction on fragments it currently improves — a defensible tightening, but a much smaller claim than the one withdrawn. n is small (9–11 streams qualify, since truth requires a resolved full-span estimate) and that bounds how hard this should be pushed.

This project had previously made a claim of the same shape and withdrawn it, with the reviewer noting it was marginal, not wrong. We re-opened it on stronger evidence, and independently reproduced the reason it was withdrawn. The withdrawal was correct. The §3.6 result stands as stated — a rate quoted from a short fragment is below its own noise floor — but it licenses a caveat on quoting, not a change to a gate that governs applying.

3.7 Two estimator defects surfaced by the same run

A synthesised counter passes as an independent clock. The flag distinguishing "two genuinely independent clocks" from "a host column derived from the device" thresholded the residual spread. A device counter constructed as sample-index × an assumed rate, at 1 s granularity, produces an enormous spread — so the coarser the fabrication, the more independent it appeared. On a real O2Ring segment the estimator returned a confident 2765.5 ppm with the independence flag set, for a device with no oscillator at all. The whole-file run did refuse, but only because a counter reset made the span negative and tripped an unrelated plausibility bound.

The separating quantity is the concentration of the device's own inter-sample deltas. Over 381 sidecars, the 356 real clock streams reach at most 56.00 % modal-delta share while the 25 drawn streams start at 79.04 %, with nothing in between; a threshold of 0.67 gives 0/25 missed and 0/356 false positives. A pre-existing detector set at ≥ 99 % misses 5 of the 25, because a fabricated axis is only ~100 % concentrated when nothing interrupts it — a single counter reset or repeated stamp is enough. We also find the Verity's inter-pulse-interval stream to be 100 % drawn, previously unflagged.

3.8 The error that actually matters is four orders of magnitude larger

Everything above concerns the time axis. None of it can express the failure that reaches a user: a heartbeat the detector missed or invented. That needs a second substrate — 101 real H10 RR trains, median 1440 beats, median true rMSSD 44.6 ms — and a second readout, the HRV metric itself. We injected missed beats (two intervals merge), false positives (one interval splits) and detector jitter at known rates.

Injectionraw rMSSD errorECG-Malik (0.20)PPG-Malik (0.30)beats corrected
miss 0.1 %+20.8 %−22.0 %−18.9 %3
miss 0.5 %+114.4 %−22.0 %−19.6 %9
miss 2 %+387.0 %−22.0 %−19.4 %31
miss 5 %+614.5 %−22.2 %−17.8 %73
false positive 0.1 %+23.0 %−22.0 %−18.8 %8
jitter 2 ms+0.1 %−22.0 %−18.8 %3
jitter 30 ms+15.3 %−5.4 %−3.9 %3
null (clean)0.00 %−21.95 %−18.85 %3

One missed beat in a thousand inflates rMSSD by 20.8 %. Mean RR moves 0.06 % for the same injection, so a summary heart rate looks untouched while the variability metric is a fifth wrong. For comparison, a 30 ppm rate error — larger than any device measured here — moves rMSSD by roughly 0.003 %, because rMSSD is a first difference and a smooth rate error cancels between adjacent intervals, whereas a missed beat fuses two intervals into one of about twice the length and rMSSD is quadratic in that difference. The two failure families differ by four orders of magnitude in their effect on the delivered number, and the pipeline had instrumentation for the smaller one only.

The artefact corrector removes the injected damage — +614 % becomes −22 %, a real repair — but lands on the same value regardless of what was injected, including on an unperturbed train. It converges on its own flattened estimate rather than on truth.

What this cannot decide. Our "truth" is the as-recorded device RR train, which already contains genuine ectopy and genuine detector artefacts. The null row's −21.95 % therefore admits two readings this design cannot separate: a corrector bias, or the correct removal of about 22 % of real artefact inflation. We do not claim the former. What the experiment does establish is magnitude — correcting a median of 3 intervals out of 1440 moves rMSSD by 22 %, because the metric is quadratic and its extremes dominate it. Separating the two requires independently adjudicated R-peaks, which this corpus does not contain.

Corrector choice is itself a measurable error. The suite deliberately applies a stricter Malik bound to ECG and pulse-interval data than to optical photoplethysmography (0.20 versus 0.30). An earlier run of this analysis applied the optical corrector to ECG-derived intervals — the wrong one — and both legs are now reported side by side for that reason. On a synthetic train with one planted merged beat, the ECG-tuned corrector recovers rMSSD to 8.1 ms against a truth of 8.1 ms, correcting exactly one interval.

3.9 The capture host contributes almost nothing to the residual

§2.1 argued from a ratio that the host could not be the limiting term. That is an argument, not a measurement, so we tested it directly. A host-side disturbance is common-mode: it displaces every device's host-minus-device residual at the same instant. A device- or transport-side disturbance is not. On nights where two devices were captured simultaneously by the same host (4 pairs, median 4.8 h of genuine overlap), we detrended each residual and correlated them on a shared one-minute grid.

Median correlation r = 0.036 (range −0.057 to 0.105). Roughly 3.6 % of residual variance is common-mode; the remainder is device or transport. Residual standard deviations are 42.6 ms (H10) and 48.9 ms (Verity), against a host RMS offset of 19.5 µs — 0.046 % of the residual it would have to explain. The host is not the term to improve, and this is now measured rather than inferred. With only 4 qualifying pairs the interval on r is wide; the conclusion is robust because the effect size is not marginal.

3.10 The noise is stream-specific, and one stream is an order of magnitude worse

Attributing noise to a sensor is normally confounded with which radio adapter carried it. That confound is escapable by comparing two streams from the same device, on the same adapter, on the same night, at matched span — any difference between them cannot be an adapter term.

Streamnmedian σy (ppm)median Allan slopemedian residual spread
H10 / ecg434.5−0.993525 ms
H10 / acc443.6−0.992534 ms
Verity / ppg642.0−0.990734 ms
Verity / acc5322.8−0.9712030 ms

At a matched 4800 s span the Verity accelerometer stream is 7.7× noisier than the same device's optical stream. Because both come from one device over one link, that gap is a property of the stream — its packet cadence and buffering — and not of the device, the adapter or the recording length. Every stream shows a clean τ−1 slope, so all are phase-noise dominated; the accelerometer is simply far worse, and its slope is also the least steep (−0.971), meaning it averages down slightly less well.

3.11 Measuring the detector's miss rate without a ground truth — two instruments, one of which refuses

§3.8 establishes that a missed beat is four orders of magnitude more damaging than the clock error this paper began with, but it injects errors into an already-detected interval series and therefore assumes a miss rate rather than measuring one. Measuring it requires knowing which beats a detector missed, and no recording here carries adjudicated R-peaks. Two fields have solved that without ground truth, and we applied both.

Capture–recapture, from epidemiology, does not work here — and the reason is informative. Given several imperfect detectors, the size of the unobserved class is estimable from their overlap. Three independent instruments were available on one recording (chest ECG, arm photoplethysmography, finger photoplethysmography), which is the minimum admitting a dependence correction: two lists leave the model saturated, and detectors that fail together — motion, poor perfusion and apnea degrade optical and electrical channels at once — bias a two-list estimate downward. On a clean 18.85-min window the three detectors agreed on 935 of 986 beats, and the estimator's informative cells fell to single digits; the closed form returned 701 unseen beats, i.e. 41 % of the recording, which is absurd. It now refuses, reporting that the data cannot identify the undercount — explicitly not that nothing was missed. Ranking windows by detector disagreement did not rescue it: in seven of nine windows the three detectors agreed on every beat. Perfect agreement is equally consistent with nothing missed and with all three detectors missing the same beats, so on quiet recordings the quantity is unidentifiable by construction rather than merely imprecise.

The artificial-star test, from observational astronomy, does work. A survey's true source count is likewise unknown, and the standard remedy is to plant synthetic sources of known brightness into the real image, re-run the real pipeline, and measure the recovered fraction — a completeness curve, requiring no ground truth, no second detector and no independence assumption. Here the image is the raw waveform and the star is a beat: the subject's own averaged beat, planted at positions at least a refractory period from any existing beat so that correct refractory rejection is never scored as a miss.

amplitude (× local noise)010203040≥60
completeness0 %0 %8.7 %94.7 %100 %100 %
spurious detections000000

A sharp threshold between 20 and 30, and no induced false positives at any amplitude. Real beats on the same recording have median amplitude 48.9× local noise, with 0.61 % below 20 and 2.21 % below 30; convolving the distribution with the curve gives on the order of 1 % of beats at risk.

3.12 Planting beats in gaps overstates the detector — a correction to §3.11 in the opposite direction to the one we predicted

We first reported the §3.11 figure as an upper bound, reasoning that a planted beat sits in a gap and cannot benefit from the searchback that recovers a faint beat arriving in rhythm. A second experiment, attenuating real beats in situ so that position and neighbouring intervals are preserved exactly, shows the reasoning was right about the existence of a bias and wrong about its sign.

At matched amplitude, in-rhythm recovery is lower than gap-planted — 2.6 % against 94.7 % near the threshold — and convolved over the same beats on a common axis the in-rhythm curve gives ~3× more misses (1.56 % versus 0.53 %). Gap-planting overstates what the detector achieves, so the bound moves up rather than down.

The mechanism is the detector's adaptive threshold, not searchback, and it was isolated rather than assumed: attenuating a beat together with its neighbours lifts its recovery from 1.2 % to roughly 74 % at the same amplitude, because full-amplitude neighbours hold the running threshold high while a planted beat faces no such local competition. Searchback contributes nothing measurable. An alternative explanation — that the attenuation construction distorts beat morphology rather than merely reducing amplitude — was tested by a second construction preserving morphology, and the two agree to about 11 %, so it is excluded rather than merely unconsidered.

A measurement artefact of our own was uncovered in the process, and it had been hiding the effect threefold. Subtracting a template does not leave silence: it leaves a beat-shaped, beat-positioned residual, so the nominal amplitude axis reported zero where the detector still saw signal. Detections at nominal-zero amplitude proved to sit a median 15.4 ms from the true beat position — about two samples — rather than at the ~507 ms interval midpoint an inference would produce, which identified them as real residuals rather than fabrications and forced the axis to be re-derived from measured rather than nominal amplitude.

Consequence for the delivered metric. A missed beat and an invented one damage variability in opposite directions but offset in mean interval, so the summary heart rate cancels the error while the variability metric compounds it — the same asymmetry §3.8 measured from the other side, where a 0.1 % miss rate moved mean interval by 0.06 % and variability by 20.8 %. The clustered regime is milder than that framing suggests, however: where several consecutive beats degrade together, the failure is late or imprecise detection (15 ms of scatter) rather than a merged or split interval (500–1000 ms), and only the latter is a large variability error. Which regime a given recording occupies is not measured here.

4. Discussion

Three results here point the same way. The estimator is accurate and its output is usually unresolvable (§3.3, §3.6); the host, long suspected of dominating the residual, contributes 3.6 % of it (§3.9); and the error that moves the delivered metric by a fifth lives in a component this analysis did not originally include (§3.8). Each was invisible for the same reason — the pipeline was validated against its own parts, and every part agreed.

The practical recommendation is not to improve the clock estimator. It is to instrument beat-detection error at the same standard, because a detector operating at a 0.1 % miss rate — a figure most would call excellent — is already delivering a 20 % error in rMSSD, and nothing downstream would show it. The corrector that exists for this partly masks it: it removes the injected damage but converges on a fixed offset regardless of input, so a metric that looks stable after correction is not thereby correct.

4.1 Limitations

This is a single-subject, single-host study; the crystals characterised are three specific devices. Injection occurs after capture — at the arrival sidecar and at the RR train — so the capture daemon itself is outside the loop and a defect introduced during acquisition would not be seen. The beat-substrate "truth" is an as-recorded device train rather than adjudicated R-peaks, which is what prevents §3.8 from separating corrector bias from correct repair. Several sub-analyses rest on small n: 9–11 streams for the span-gate test in §3.6b, 4 device pairs in §3.9, 4–6 streams per row in §3.10. Where n is small we report the effect size and say so rather than testing significance on it. The beat-detector completeness work (§3.11–§3.12) is single-recording and single-subject, its capture–recapture arm is reported as a refusal rather than an estimate, and whether real recordings sit in the isolated or the clustered failure regime is unmeasured. Finally, criteria were preregistered and hashed, but the same operator injected and analysed; §3.6b is a live demonstration of why that matters, since the error it retracts was ours and was found only by running one more measurement.

4.2 Conclusion

Injecting known defects into real recordings shows that the clock-rate estimator under test is accurate, unbiased, linear over three and a half orders of magnitude, exactly robust to the packet loss that actually occurs, and structurally blind to constant offset. It also shows that the quantity it produces is, on most wearable recordings, smaller than its own noise floor; that two of its auxiliary judgements were wrong in ways self-consistent validation could not reveal; and that one inference we drew from our own headline result was itself unsound and had to be withdrawn.

The most useful result is the one that came from asking the same question of a different substrate. A missed beat in one thousand costs 20.8 % of rMSSD; the entire clock-error family costs about 0.003 %. A great deal of engineering attention — including this paper's first two thirds — had gone to the smaller term, because it was the one the pipeline had instruments for. The host, likewise, turns out to contribute 3.6 % of the timing residual it was suspected of dominating, and the worst noise in the corpus belongs not to a device but to one stream on a device whose other stream is fine.

Measuring the larger error proved harder than measuring the smaller one, and instructive in the same way. One imported method could not answer the question at all on clean data and says so; another answers it, and its first answer was wrong by a factor of three in the flattering direction because our own amplitude axis reported silence where signal remained. Every correction in this paper moved the number the uncomfortable way, and none was found by re-reading code.

The general lesson is not about clocks. A pipeline validated only against itself can be accurate, uninformative, and pointed at the wrong failure mode all at once — and only a truth fixed in advance can tell the three apart.

5. Reproducibility

Preregistered criteria: experiments/known-clock/acceptance.json, sha256 b061d2792c1ff8d605ec82ff9fd298d56ca40915877ab274aa1785f2baff1586, committed before the first run. Regenerate every figure with:

node tools/known-clock-recovery.mjs --self-test
node tools/known-clock-recovery.mjs --root <captures> --out results.json
node tools/beat-error-recovery.mjs --self-test
node tools/beat-error-recovery.mjs --root <captures> --out beats.json
node tools/beat-capture-recapture.mjs --self-test
node tools/beat-injection-recovery.mjs --self-test
python3 experiments/known-clock/validate_allan.py

All perturbations are deterministic; no figure depends on a seed drawn at run time.

6. References

  1. Riley WJ. Handbook of Frequency Stability Analysis. NIST Special Publication 1065, 2008. — overlapping Allan deviation, noise-type identification by log-log slope, and the divergence of the standard deviation for the noise types measured here.
  2. Allan DW. Statistics of atomic frequency standards. Proceedings of the IEEE 54(2):221–230, 1966. — the two-sample variance underlying §2.3 and §3.5. doi:10.1109/PROC.1966.4634
  3. Wallin AEE, Price DC, Carson CG, Meynadier JM. allantools — Allan deviation tools in Python, version 2024.06. Reference implementation used for the independent validation in §2.3.
  4. Malik M, Bigger JT, Camm AJ, et al. Heart rate variability: standards of measurement, physiological interpretation, and clinical use. Task Force of the European Society of Cardiology and the North American Society of Pacing and Electrophysiology. European Heart Journal 17(3):354–381, 1996. — rMSSD and SDNN definitions, and the ectopic-interval correction convention applied in §3.8.
  5. Citations here carry no DOI links by design: the two works cited are not in audits/CITATION-VERIFICATION-2026-08-05.json, and this analysis has no network to verify attribution against Crossref. Adding an unverified ledger entry to satisfy the citation gate is the one edit that hides a real defect, so the full bibliographic record is given instead — matching how the suite already cites Task Force 1996 throughout.
  6. Chao A. Estimating the population size for capture–recapture data with unequal catchability. Biometrics 43:783–791, 1987 — the lower bound underlying the undercount estimator in §3.11, and Böhning et al.'s later one-inflation–robust variant used as its contamination diagnostic.
  7. Artificial-star completeness testing, standard in observational astronomy: synthetic sources of known brightness are injected into real imaging and the recovered fraction is measured as a function of brightness. §3.11 applies the same design to a beat detector.
  8. Tepna Project. briefs/SEARCHBACK-AWARE-INJECTION-2026-08-15-BRIEF.md — the in-situ attenuation experiment of §3.12, its adaptive-threshold mechanism, and the measured-versus-nominal amplitude correction. Performed independently and reported here with its numbers unchanged.
  9. Tepna Project. Clock Contract §7 — the host-disciplined time axis. CLAUDE.md, 2026. — the estimator under test.
v2.8.0