PpgDex (raw PPG) · ECGDex (raw ECG) nodes, Tepna physiological-signal suite
Background. Draft v1 of this paper compared optical HRV against chest ECG on four overnight sessions and reported rMSSD agreement of −0.6% on the two nights the pipeline's own quality gate trusted. That result was correct but under-powered and selected: it described the two best nights of four. Data. One subject wore a Polar H10 chest ECG (~130 Hz) and a Polar Verity Sense upper-arm PPG (3 green LED pairs, ~176 Hz) simultaneously for 26 nights; 20 yielded a usable paired analysis (1,477 five-minute epochs, 121.5 h). All beats derive from raw signal by the production detectors. Methods. A single global affine time map between the two independent BLE clocks was found to be insufficient — it produced a physiologically impossible negative pulse-transit time and an apparent beat F1 of 0.26 — because the devices do not drift at a constant rate and pulse-transit time itself wanders. Alignment is therefore re-estimated per epoch. Results. Beat detection is near-perfect: median sensitivity 1.0000 (IQR 0.928–1.000) and PPV 1.0000 (IQR 0.937–1.000). Pulse-interval timing carries essentially zero bias (+0.01 ms) but a jitter of 5.92 ms (IQR 3.98–10.61). rMSSD runs +4.24% high (per-night median; IQR +1.83 to +6.85; 18 of 20 nights within ±10%). The central result is mechanistic: rMSSD error is fiducial jitter and essentially nothing else — the model rMSSD²ppg = rMSSD²ecg + k·σ² fits with per-night R² = 0.964 (IQR 0.937–0.989) and k = 2.39 (IQR 1.79–2.85), against a theoretical k ∈ [2,6]. Catastrophic epochs (|ΔrMSSD| > 25%, 10.6% of the corpus) are simply high-jitter epochs: median jitter 35.0 ms versus 5.5 ms elsewhere, with 91% exceeding 15 ms against 7% of the rest. Conclusion. Optical rMSSD accuracy reduces to a single engineering quantity, and the design budget follows in closed form: σ ≤ 3.51 ms buys 1% bias, ≤ 4.98 ms buys 2%, ≤ 6.11 ms buys 3%. This is a single-subject method study, not a population accuracy rating.
Keywords: heart-rate variability · rMSSD · photoplethysmography · chest ECG · pulse-rate variability · method comparison · fiducial jitter · error budget · inter-device clock drift · consumer wearables
An armband reads your pulse with light; a chest strap reads the heart's electricity directly and is the trusted “truth.” We wore both to bed for twenty nights and compared them beat by beat. The light sensor finds essentially every heartbeat — it almost never misses one and almost never invents one. But it places each beat with a small timing wobble of about six thousandths of a second, because it is watching a pressure wave arrive rather than the electrical spark that caused it. That wobble turns out to be the whole story: a common heart-rhythm score comes out about 4% too high, and the wobble explains almost all of it — the fit is tight enough (96%) that knowing the wobble lets you predict the error. That converts a vague “light sensors are noisier” into an engineering target: cut the wobble to 3.5 thousandths of a second and the error falls to 1%. The earlier version of this paper looked at four nights and reported the two best ones, which agreed almost perfectly; with twenty nights and no picking, the honest number is 4%, with two genuinely bad nights that are bad for a reason we can now name. One person, twenty nights — a method result, not a product rating.
Heart-rate variability from a photoplethysmogram (PPG) is attractive because the hardware is cheap and unobtrusive, but optical pulse timing is a mechanical, downstream proxy for the electrical heartbeat an ECG records directly. The literature's summary of that gap is largely qualitative — pulse-rate variability is “unbiased but noisier” than heart-rate variability — and the engineering consequence is left implicit. The most recent quantitative synthesis (Xu et al. 2026, a meta-analysis of PRV–HRV agreement on RMSSD and SDNN) pools only resting, controlled-condition studies and explicitly warns that its estimates “should not be generalized to sleep … or free-living settings” — which is precisely the setting this paper measures: overnight sleep, on consumer hardware, every epoch reported. This paper closes that gap on real overnight data by asking not whether optical rMSSD is biased but what quantity produces the bias, and then inverting the relationship into a design budget.
What changed since Draft v1. The first draft used four nights and reported rMSSD agreement of −0.6% on the two that the pipeline's reference-free gate trusted. Nothing in it was fabricated, but two things were weak: n = 2 for the headline number, and those two nights were selected by a gate whose behaviour was then evaluated on the same four nights. This draft uses 20 paired nights and reports every epoch that aligns, selected or not. The headline agreement figure consequently gets worse (+4.24% rather than −0.6%), and that is the honest direction of travel: v1 described the best case, v2 describes the corpus.
One subject wore a Polar H10 chest strap (single-lead ECG, measured 129.94–130.03 Hz) and a Polar Verity Sense optical armband (measured 176.35–176.42 Hz) simultaneously across 26 nights, captured with Polar Sensor Logger. The Verity's three PPG channels are three diametrically-opposed green LED pairs time-multiplexed onto one shared photodiode, plus a fourth LEDs-off ambient slot — not three wavelengths and not three photodiodes (FCC INW4J internal photographs; four Polar staff statements in the vendor SDK issue tracker). Streaming raw PPG forces SDK mode, which disables the device's own HR and PPI outputs, so both vendor files are empty by design and every optical beat here is derived from the raw waveform.
ECG R-peaks come from a Pan–Tompkins implementation with sub-sample parabolic refinement, validated against template cross-correlation at 0.244 ms residual jitter (median night) versus the 2.213 ms the bare 7.69 ms sample grid would impose, and cross-checked against the H10's own RR file at 0.47 ms median interval agreement. Optical beats come from the production PPGDSP path: per-channel band-pass, TERMA detection with a cadence-adaptive refractory, three-LED consensus (a beat is kept only where ≥2 of 3 channels agree within ±50 ms), and intersecting-tangent foot timing on the reference channel.
This is the methodological trap in the study, and it is worth stating plainly because the naive analysis produces confident nonsense. The two devices are independent BLE peripherals with independent crystals, logged by one phone; the phone timestamp is derived from each device's own sensor clock, so same-phone logging does not synchronise them. Fitting one affine map tppg = a·tecg + b over a whole night — offset plus a constant drift rate — yielded a median pulse-transit time of −1 ms (physiologically impossible; the pulse cannot precede the R-wave) and an apparent beat F1 of 0.26, while two nights of the same corpus scored 0.99 under the identical procedure. The failure is not in the signal: restricted to a 20-minute window the same data matches at 1,244 of 1,245 beats with a residual standard deviation of 8 ms.
The cause is that neither the relative clock rate nor the pulse-transit time is constant across seven hours. Alignment is therefore estimated per five-minute epoch: a coarse lag from cross-correlating the two instantaneous-heart-rate series (a slowly-varying signal with no periodicity at the beat interval, so it cannot alias by a whole beat), then a local refinement, then one-to-one beat matching within ±75 ms. Epochs whose match rate falls below 50% are recorded as alignment failures and excluded (30 of 1,488, 2.0%). Because each epoch's lag absorbs the local transit time, the reported pulse-interval bias is near zero by construction and is not evidence about transit time.
If each optical fiducial carries timing noise of standard deviation σ, that noise propagates into successive-interval differences and inflates rMSSD in quadrature. Writing the measured optical rMSSD as a function of the ECG truth and the jitter gives rMSSD²ppg = rMSSD²ecg + k·σ², with k set by how much fiducial noise the successive-difference operator sees. Fully independent per-beat noise gives k = 6; noise shared between adjacent intervals pulls it toward 2, which is what a foot-to-foot construction predicts since each foot enters two intervals. k is fitted per night by least squares on the epoch series and reported with its distribution, never assumed.
Epochs within a night are not independent. Every inferential statement below is therefore computed per night and summarised across the 20 nights; pooled epoch-level figures appear only as descriptive and are labelled as such. This distinction is not cosmetic — a companion paper (Pooling epochs inverts gate rankings) shows that on this same corpus the pooled and night-clustered analyses rank signal-quality gates in exactly opposite order.
Across 20 nights the optical detector finds essentially every ECG beat and invents essentially none: median sensitivity 1.0000 (IQR 0.928–1.000) and median PPV 1.0000 (IQR 0.937–1.000). Pulse-interval bias is +0.01 ms. What the optical arm loses is not beats but timing precision: median jitter 5.92 ms (IQR 3.98–10.61, range 3.02–48.80 across nights).
| Quantity | Median | IQR | Range across nights |
|---|---|---|---|
| Beat sensitivity | 1.0000 | 0.928 – 1.000 | 0.609 – 1.000 |
| Beat PPV | 1.0000 | 0.937 – 1.000 | 0.610 – 1.000 |
| PPI jitter (sd) | 5.92 ms | 3.98 – 10.61 | 3.02 – 48.80 |
| PPI mean absolute error | 3.94 ms | 3.00 – 6.00 | 2.25 – 39.95 |
| PPI bias | +0.01 ms | — | −0.62 – +0.37 |
| rMSSD bias | +4.24% | +1.83 – +6.85 | −1.1 – +118.9 |
| SDNN bias | +2.46% | +1.15 – +3.76 | +0.14 – +63.0 |
The rMSSD distribution is strongly right-skewed by two nights: 18 of 20 nights fall within ±10% and 10 within ±5%, while two sit at +72.1% and +118.9%. The mean across nights (+13.09%, sd 29.34) is therefore not a useful summary and the median (+4.24%) is reported as primary. A night-level Wilcoxon signed-rank test rejects zero bias (p < 0.0001). The night-level 95% limits of agreement are wide — −44% to +71% — and that width is dominated entirely by the two outlier nights, which §3.3 accounts for.
Fitting rMSSD²ppg = rMSSD²ecg + k·σ² per night gives a median R² of 0.964 (IQR 0.937–0.989; minimum 0.670 across the 20 nights) with k = 2.39 (IQR 1.79–2.85). The fitted k lies inside the theoretical band [2,6] and near its lower end, indicating that fiducial noise is partially shared between adjacent intervals rather than fully independent — exactly what a foot-to-foot interval construction predicts. Pooled across all 1,488 epochs the same model gives k = 2.29 and R² = 0.909 (descriptive only).
Because a single quantity governs the error, the design budget inverts in closed form. Holding the corpus's median ECG rMSSD fixed and solving for σ:
| Target median rMSSD bias | Required jitter σ | Status |
|---|---|---|
| ≤ 1% | 3.51 ms | not reached — needs a 41% jitter reduction |
| ≤ 2% | 4.98 ms | not reached — needs 16% |
| ≤ 3% | 6.11 ms | marginally reached on the median night |
| ≤ 5% | 7.93 ms | reached on 15 of 20 nights |
158 of 1,488 epochs (10.6%) carry |ΔrMSSD| > 25%. They are not a distinct pathology: their median jitter is 35.0 ms against 5.5 ms elsewhere, and 91% of them exceed 15 ms jitter against 7% of the remainder. The model that explains the well-behaved epochs explains these too; they simply sit far out along the same curve. Practically this means the two outlier nights in §3.1 need no separate explanation, and that any intervention which reduces jitter — or any gate that excludes high-jitter epochs — addresses the whole error distribution at once.
Draft v1 reported that the pipeline's reference-free gate was “a near-perfect self-classifier,” having labelled two of four nights trusted and two flagged in exact correspondence with the ECG-measured error. That observation stands and reproduces here — but it is a claim about night-level triage resting on a single binary split of four nights, which is consistent with chance. It is not evidence about epoch-level gating, and the two questions have different answers. At epoch level, evaluated within nights, the shipped per-beat SQI discriminates catastrophic epochs with a median AUC of 0.763 (IQR 0.645–0.791; above chance on 9 of 10 evaluable nights, Wilcoxon p = 0.014) — good, but not “near-perfect,” and materially short of what the v1 phrasing implies. The full gate comparison, including a methodological hazard that inverts the ranking entirely, is the subject of the companion paper.
Three conclusions follow. (1) The optical arm's weakness on real overnight data is not beat detection — sensitivity and PPV are both at ceiling — but fiducial timing precision. Effort spent on finding more beats is misdirected; effort spent on locating them more precisely is not. (2) Because rMSSD error is jitter and essentially nothing else (per-night R² = 0.964), the accuracy question collapses to a single scalar with a closed-form budget. That is more useful than an agreement interval, because it tells a designer what to change and by how much: reaching 1% bias requires cutting jitter by 41%, a concrete and falsifiable target rather than an aspiration. (3) The corpus-honest agreement figure (+4.24%) is roughly seven times worse than the selected-best-nights figure Draft v1 reported (−0.6%). Both are correct about what they measured; reporting the second without the first overstates what a wearable delivers in practice, and this revision exists partly to correct that.
A negative result worth recording: the two obvious routes to lower jitter were tested on this hardware and both failed. Combining the three LED channels into one waveform — by simple averaging, by the first principal component, or by a maximum-SNR generalised-eigenvalue beamformer — loses to selecting the best single channel, out of sample, on both a spectral criterion and on foot-to-foot interval dispersion; the channels correlate at ρ ≈ 0.99, so there is little independent noise to average away. And substituting the middle-amplitude fiducial that the tilt-table literature ranks above the intersecting-tangent point makes timing worse here (jitter 5.02 ms vs 4.44 ms on a matched segment), consistent with both source papers' own caution that the ranking is morphology- and site-dependent. The jitter floor on this device is therefore not trivially removable, which is precisely why quantifying it is useful.
Polar_Sense_*_PPG.txt (raw 3-channel + ambient, ~300 MB/night) and Polar_H10_*_ECG.txt, 2026-06-10 → 2026-07-13; 20 yield a paired analysis. Raw archives are gitignored for size; the derived per-epoch table is the committed artifact.ecgdex-dsp.js / ppgdex-dsp.js, run unmodified in a co-loaded realm in PpgDex.src.html script order, so the optical arm is byte-identical to the shipped app (~28 s and ~1.1 GB peak per night).*_RR.txt at 0.47 ms median.Draft v1 projected that within-subject limits of agreement would stabilise at roughly 25–30 trusted paired nights and that the budget should then shift from depth to breadth. At 20 nights this draft sits inside that window and the projection broadly holds: the night-level bias interval is now informative (±12.9 percentage points at 95%) while the limits of agreement remain wide because two outlier nights dominate the spread. The more useful observation is that the jitter model reached its asymptote far earlier than the agreement interval did — per-night R² is already 0.964 with a tight IQR, so the mechanistic claim is well supported at this n even though the descriptive agreement claim is not. Additional single-subject nights will narrow the agreement interval slowly and will not meaningfully improve the model fit. The remaining value is therefore in breadth: more subjects, and above all more motion regimes, since every night here is quiet sleep and the jitter distribution under motion is exactly what a wearable-grade claim would need.
papers/ppg-quality-gate-pooling.html — pooled vs night-clustered evaluation of signal-quality gates on this corpus.papers/rmssd-equivalence.html — three-detector rMSSD equivalence on shared synthetic beats.papers/sigma-no-reference.html — reference-free per-device σ on the same device corpus.papers/wearable-clock-drift.html — inter-device timing drift between two consumer wearables logged by one phone.CLAUDE.md (Clock Contract, evidence-grade system), Tepna suite.