← Tepna preprints

Measuring a device's σ without a canonical reference: heart-rate error across the O2Ring, Polar H10 and Verity Sense by repeatability, transfer-standard agreement, and the three-cornered hat

Michal Planicka  ·  corresponding author — Tepna Project

OxyDex (oximetry) · PulseDex / ECGDex (cardiac) nodes, Tepna physiological-signal suite

Draft v2 · June 2026, revised August 2026 · Analysis tool: sigma-no-reference-analysis.html · Inputs: real co-recorded device files · 100% local, reproducible

v2 (2026-08-15) corrects limitation (x): the claim that this corpus's raw capture was unavailable to re-derive was false, and the window-length sensitivity is now measured on these nights. v2.1 (2026-08-15, later) re-runs that sweep on the full 52-night committed corpus (3.1× the headline sample) — which withdraws v2's "the corners reorder" reading as a 24-night small-sample effect, widens the non-monotonicity finding from one corner to two, and shows the O2Ring corner reproducing across both analysis pipelines once window length is matched. No headline σ changes in either revision.

Abstract

Problem. How do you state a consumer sensor's measurement uncertainty (σ) when you own no laboratory-calibrated reference? Approach. A three-rung recipe that needs no certified instrument: (1) repeatability — random scatter from repeated reads, estimated with no reference; (2) transfer-standard agreement — promote the best device on hand (a chest-strap ECG) to a working reference and report Bland–Altman bias, 95% limits of agreement and the accuracy root-mean-square (Arms), recovering the test device's random σ by variance subtraction; (3) the three-cornered hat — with three devices measuring the same quantity at once, solve each one's individual variance with none assumed canonical. Data. The pulse-oximeter heart-rate channel (Wellue O2Ring) against a Polar H10 ECG strap, co-recorded over twenty-six overnight sessions (291,561 paired 1-Hz seconds), plus a Polar Verity Sense armband as the third corner; the analysis tool folder-ingests a night's raw captures and derives every corner from raw signal. Results. The O2Ring pulse carries a negligible mean bias (−0.20 bpm) but a second-by-second random σ of ≈3.2 bpm on the clean nights (one motion-corrupted night, 06-12, reached 9.9 bpm), Arms 3.3 bpm. With all three corners reduced from raw signal — H10 HR from raw ECG (Pan-Tompkins QRS), Verity HR from raw PPG (PPGDSP), O2Ring native pulse — a fused-weight artifact-robust three-cornered hat over the full twenty-six-night folder corpus (291,561 s, production SQI-gated Verity) returns σ 2.41 / 1.28 / 1.42 bpm (O2Ring / H10 / Verity), each an across-night 95% bootstrap CI. The plain hat's variance decomposition has breakdown point zero — a single transient artifact detonates it (a spurious-QRS burst on one night, 06-12, inflated the raw per-window σH10 to ≈9.6 bpm) — so each per-second pairwise difference is weighted by a per-corner beat-confidence (window-relative beat-density × signal-quality, self-calibrating and arrhythmia-safe) under a gentle cross-sensor consensus floor; the burst is then down-weighted in place and 06-12 stays in the corpus as a clean ≈2.6-bpm H10 corner rather than being discarded. The raw-ECG H10 corner is the quietest (1.28, resolving positive in 21/26 nights), confirming it as the gold leg; the O2Ring pulse is the noisiest (2.41) and the raw-PPG Verity sits between (1.42). (An earlier deep hat over six hand-selected clean windows had instead put Verity noisiest at 3.50 — a small-sample artifact of one motion-contaminated Verity window; the twenty-six-night corpus, with a cleaner production PPGDSP Verity derivation and more O2Ring motion nights, redistributes the coupled variance and settles the ordering.) None is assumed canonical; the ranking reflects error structure, not fidelity (the O2Ring's firmware emits an internally-smoothed integer pulse, visible as diagonal Bland–Altman banding, that deflates its variance). The ECG-derived H10 matches Polar's onboard RR to −0.04 bpm, confirming the reference leg. Channel limit. Only the HR channel is testable — the ECG and PPG references carry no SpO₂, so the O2Ring's oxygen-saturation trueness cannot be established here. Capture lesson. The Verity Sense's onboard HR/PPI streams were empty (all-zero / header-only), yet its HR was fully recoverable from the raw photodiode signal at SQI ≈ 1.0 — raw-signal logging, not the device's firmware estimate, is what made the third corner possible. This is a single-subject methods pilot — it demonstrates the apparatus, not a population accuracy claim.

Keywords: measurement uncertainty · metrology without a standard · three-cornered hat · Gray–Allan variance · Bland–Altman · pulse oximetry · heart rate · raw PPG / ECG · QRS detection · consumer wearables

0. Layman overview (delete before submission)

Every gadget that reports your heart rate is a little bit wrong. The usual way to measure how wrong is to compare it against a lab-grade “truth” instrument — but ordinary people don't own one. So how do you put a real error number on a fitness ring without a reference? This paper uses three tricks that need no certified equipment: (1) measure the same thing repeatedly and look at the scatter (that gives precision); (2) borrow the best device you do own — a chest-strap ECG — as a stand-in reference; and (3) the clever one: wear three devices at once and use a bit of algebra (borrowed from atomic-clock testing) to solve for each device's own error with none of them assumed perfect.

What we found, on one person across several nights: the O2Ring's average heart rate is essentially correct (off by a third of a beat), but any single second can be wrong by ~4 beats — fine for an overnight summary, not for instant readings. The most useful surprise was about capture, not accuracy: one armband's built-in heart-rate output was completely empty (its firmware never locked on), yet the raw light-sensor signal it recorded was perfect — we reconstructed a clean heart rate from it ourselves. Lesson: always log the raw waveform, not just the device's own number. This is a real-data methods demonstration on one subject — it shows the recipe works, not a population accuracy rating for these products.

1. Introduction

A number reported by a sensor is only as useful as the uncertainty attached to it. The textbook way to obtain that uncertainty is to compare against a reference traceable to a national standard — a calibrated instrument the consumer device manufacturer, but rarely the end user, can access. This note asks the practical question that arises when no such reference is on the bench: given two or three imperfect devices, what can be said rigorously about each one's σ? The answer separates cleanly into two quantities that are routinely conflated. Precision (repeatability) is random scatter and needs no reference at all — only a stable thing to measure repeatedly. Trueness (bias) is systematic offset and fundamentally requires a comparison. We treat them separately and add a third tool — the three-cornered hat from frequency metrology — that recovers each device's own variance when three measure the same quantity simultaneously, assuming none is perfect.

One mismatch to declare up front. The O2Ring reports two signals — SpO₂ and pulse rate — but the references here (an ECG strap and a PPG armband) measure only heart rate. Everything below concerns the HR channel. The O2Ring's SpO₂ trueness is not obtainable from this data: establishing it requires arterial blood-gas CO-oximetry, the actual gold standard. SpO₂ would admit only a repeatability figure and cross-oximeter agreement, never absolute accuracy, without that reference.

2. Methods

2.1 Devices and channels

Three devices were worn together: a Polar H10 chest strap (single-lead ECG, the clinical reference for inter-beat timing), a Wellue O2Ring ring oximeter (SpO₂ + photoplethysmographic pulse rate at 1 Hz), and a Polar Verity Sense arm-band (PPG heart rate). Each device exposes both a firmware-computed heart rate and, via Polar Sensor Logger, its raw waveform (H10 raw ECG at ~130 Hz; Verity raw PPG at ~176 Hz). Where the analysis benefits — and wherever a device's firmware HR failed — we work from the raw signal with the suite's own detectors (ECGDSP Pan-Tompkins QRS; PPGDSP optical beat detection) rather than the vendor's estimate. The H10 is the most precise heart-rate source (an electrical R-wave is sharper than any optical pulse), so it serves as the transfer standard for the two optical devices.

2.2 Alignment — the Clock Contract does the work

The devices share no clock and no timezone, but each stamps local civil time. Under the suite's time model every record is stored as UTC-normalized floating wall-clock milliseconds (tMs = Date.UTC(y,mo−1,d,h,mi,s)), so two devices that record the same wall-clock second produce the same tMs by construction — alignment is then exact-second intersection with no zone negotiation. H10 R-R intervals are converted to instantaneous heart rate (60000/RR) and averaged into 1-Hz buckets; the O2Ring pulse is natively 1 Hz. We keep only seconds present in both streams, discarding the O2Ring's -- contact-loss rows and any out-of-range (<30 or >220 bpm) sample.

2.3 The three rungs

(1) Repeatability σ. Short-term precision with no reference: the 1-Hz residual after a 7-point rolling median, summarized by a robust SD (1.4826·MAD). On the H10 this is a genuine precision figure; on the O2Ring it is degenerate because the device's reported pulse is internally smoothed (successive seconds are frequently identical), so the O2Ring's precision is read instead from its scatter against the H10.

(2) Transfer-standard agreement. With the H10 promoted to a working reference, the per-second difference d = O2Ring − H10 gives the mean bias , the SD of differences s_d, the 95% limits of agreement d̄ ± 1.96·s_d, and the accuracy root-mean-square Arms = √(d̄² + s_d²) — the single number used to grade pulse devices. Because the reference is itself slightly noisy, the test device's random uncertainty is recovered by variance subtraction:

σ_device = √( s_d² − σ_ref² ) , σ_ref = H10 repeatability ≈ 0.86 bpm

(3) Three-cornered hat (Gray–Allan). With three devices A, B, C measuring the same true HR, and pairwise difference variances V_AB, V_AC, V_BC, each device's own variance follows with none assumed canonical:

σ²_A = ½(V_AB + V_AC − V_BC) , σ²_B = ½(V_AB + V_BC − V_AC) , σ²_C = ½(V_AC + V_BC − V_AB)

Shared physiological variation cancels in the differences, leaving device error; a negative output flags correlated errors that break the independence assumption. The estimator is implemented in the tool and runs on every window where all three usable streams coexist — here, the H10 ECG, the raw-recovered Verity PPG, and the O2Ring pulse.

(3′) The fused-weight, artifact-robust hat. The variance decomposition above is a difference-of-variances, and an ordinary variance has breakdown point zero: a single transient artifact — one device's detector mis-firing for a few dozen seconds — detonates the whole estimate. This is not hypothetical here. On the night of 06-12 a spurious-QRS burst in the raw-ECG derivation inflated the raw per-window σH10 from ≈1.5 to ≈9.6 bpm, an artifact large enough to reorder the corners on its own. Rather than discard the night (throwing away ~15,000 good seconds to remove a few bad ones), the tool solves a confidence-weighted variance for each pairwise difference. Every corner emits a per-second beat-confidence c = density_trust × quality_trust ∈ [0,1]: density_trust is a window-relative Hampel score on the local beat rate (a burst of fabricated beats spikes the density far above the record's own median), quality_trust is the detector's own signal-quality index (a burst also depresses SQI). Both cues must fire for c to fall (min-combined), which is what keeps it arrhythmia-safe — genuine ectopy or AF is high-variability but clean, so its density is irregular yet its SQI stays high and it is preserved. Each difference series is then weighted by the product of its two corners' confidences under a gentle cross-sensor consensus floor (a Tukey biweight, C=30 bpm, that only trims a second where all three devices grossly disagree). The estimator is self-calibrating (every threshold is a within-record median/MAD, no corpus-tuned constant), O(n), and reduces exactly to the plain hat when every c=1. Under it the 06-12 burst is down-weighted in place: the night stays in the corpus and contributes a clean σH10 ≈ 2.6 bpm instead of ≈9.6.

Why weight, not gate. A hard threshold ("drop any second whose confidence < τ") reintroduces exactly the corpus-tuned constant the suite avoids, and it is brittle at the boundary. A continuous weight degrades gracefully — a marginally noisy second contributes proportionally less rather than being cliff-edged in or out — and it is the standard robust-statistics move (weighted / M-estimated variance) for an estimator with breakdown point zero. The same fragility afflicts every variance-family metric downstream (RMSSD, SDNN, CV, MAGE), so the confidence channel is signal-agnostic and reused fleet-wide (see the follow-up brief).

3. Results

Three-device overlay — H10 ECG, O2Ring pulse and Verity raw-PPG HR, first ~600 aligned seconds of a clean window
Figure 1a. A clean three-device window (2026-06-14), first ~600 aligned seconds — H10 ECG, O2Ring pulse and Verity raw-PPG HR all track the same physiology; the spread between the three is device error. Live output of sigma-no-reference-analysis.html.
Bland-Altman against H10: O2Ring-H10 and Verity-H10 clouds over the three-device overlap, with O2Ring diagonal banding
Figure 1b. Bland–Altman against the H10 reference over the twenty-six-night three-device overlap: O2Ring−H10 (Arms 3.93 bpm) and Verity−H10 (Arms ≈3.6 bpm), drawn on the same simultaneous seconds so the two are directly comparable. Both bias lines sit near zero; over the broad corpus the O2Ring cloud is the wider of the two (Arms 3.93 vs ≈3.6) — the noisy-corner ordering (§3.2) showing up directly in the raw differences. The pronounced diagonal banding in the O2Ring cloud is the fingerprint of the device's firmware smoothing: its pulse is reported as held, internally-smoothed integers, so for each integer value the difference-vs-mean points fall on a line of slope −2 — one parallel band per beat-per-minute.
Per-night standard deviation of O2Ring minus H10, twenty-six nights
Figure 1c. Per-night SD of (O2Ring − H10): twenty-five clean nights span ≈2–4 bpm; night 06-12 (red) is auto-flagged as motion-corrupted at 9.9 bpm (dashed line = the 1.8× median threshold).

3.1 The O2Ring pulse is unbiased but noisy against ECG

Table 1. O2Ring pulse vs Polar H10 ECG, per co-recorded night over the twenty-six-night broad corpus (1-Hz paired seconds; representative nights shown, pooled rows span all 26). Reference-free σ = √(SD² − σ²H10), σH10 = 0.82 bpm. ⚑ 06-12 is auto-flagged motion-corrupted (SD > 1.8× median). † 06-18 is a short (~30 min) session — the tightest night.
NightPaired sBiasSD95% LoAArmsPearson rσ ref-free
06-1017,125−0.113.29±6.53.300.593.19
06-1120,203−0.113.63±7.13.640.613.54
06-12 ⚑15,415−1.319.87±19.39.960.389.83
06-149,606−0.012.60±5.12.600.572.47
06-1514,930−0.323.57±7.03.580.713.47
06-167,588+0.012.91±5.72.910.522.79
06-18 †1,829+0.032.24±4.42.240.782.08
06-2414,367−0.063.61±7.13.610.723.52
06-2714,507−0.363.62±7.13.640.683.52
06-286,563+0.213.03±5.93.040.652.92
Pooled (clean, 25 nt)276,146−0.133.28±6.43.283.17
Pooled (all, 26 nt)291,561−0.203.92±7.73.930.743.84

The O2Ring's pulse rate is essentially unbiased against ECG — the pooled clean offset is −0.13 bpm, well inside a single quantization step. Its random uncertainty, however, is substantial: a reference-free σ of ≈3.2 bpm on clean nights, i.e. a typical second's pulse can be wrong by several beats even when the average is right. The accuracy figure Arms ≈ 3.3 bpm sits at the boundary of the ±5 bpm / 10% tolerance commonly cited for wrist/ring HR. Including the one motion-corrupted night (06-12) inflates every pooled scatter statistic — σ rises to 3.8 bpm and the LoA widen to ±7.7 — which is exactly why a per-night quality flag, not a single grand pool, is the honest summary.

Why the correlation is only ≈0.6 despite a near-zero bias. These are resting nights: true HR barely moves, so its variance is small relative to the device's beat-to-beat noise (range compression). Pearson r divides signal variance by signal-plus-noise variance and therefore looks poor whenever the subject is stable — it is the wrong agreement metric here. Bias, SD/LoA and Arms describe the error directly and are unaffected by how much the underlying HR happened to vary.

The H10's own short-term repeatability is 0.82 bpm — small enough that subtracting it changes the recovered O2Ring σ by under 0.15 bpm (3.28 → 3.17), which both justifies treating the strap as a transfer standard and shows the result is insensitive to that assumption.

The O2Ring's firmware smoothing is visible in the data — and it matters for §3.2. Two signatures expose it. First, in the Bland–Altman plot (Fig 1b) the reported pulse is whole-integer and held across runs, so the difference-vs-mean cloud collapses onto a lattice of diagonal bands — one per integer bpm, each of slope −2 (eliminate the held value O from x = (O+H)/2, y = O−H and the locus is the line y = −2x + 2O) — rather than the formless scatter a continuously-valued device leaves. Second, successive seconds are frequently identical, so the device's own second-to-second variance is small because the firmware averages real beat-to-beat variation away, not because it tracks the instantaneous rate better. This is why a same-device self-residual is a degenerate precision proxy for the O2Ring (we read its precision from the scatter against the H10 instead) and — decisively — why the three-cornered hat in §3.2 places the O2Ring's σ below the 130 Hz ECG's: smoothness deflates variance. That ranking is error structure, not sensor fidelity.

3.2 The third corner, recovered from raw signal — a real three-cornered hat

The Verity Sense was meant to be the third corner, and its onboard outputs nearly defeated it: across every night the heart-rate stream is all-zero and the pulse-interval export is a header with no rows — the band's firmware never locked a pulse. But Polar Sensor Logger also recorded the raw photodiode signal, and the embedded algorithm's failure does not mean the signal is gone. Running the raw PPG through the suite's own optical pipeline (PPGDSP: 0.5–8 Hz band-pass, channel selection, systolic-foot detection, per-beat SQI) recovers a clean heart rate where the firmware reported nothing. We did the same for the reference leg: rather than trust Polar's onboard R-R, we re-derived the H10 heart rate from its raw ECG with the suite's Pan-Tompkins detector (ECGDSP).

Table 2. Onboard firmware output vs raw-signal recovery, Verity Sense. The device's own HR/PPI are empty; the raw PPG is excellent.
SourceOnboard usableRaw-recovered (PPGDSP)Mean SQIClean beats
Verity HR stream (all nights)0 samples
Verity PPI stream (06-16)0 beats (header only)
Verity raw PPG (06-16/17)n/afull HR series, ~176 Hz0.95–1.0096–100%

The first window came on the night of 06-16: the Verity raw PPG runs ~6 hours and an H10 ECG recording begins at 01:06, while the O2Ring records throughout — so for ~2 hours (01:06–03:04, 7,057 paired seconds) all three devices observe the same heart at once. With the three streams aligned on the shared floating tMs grid, the three-cornered hat recovers each device's own σ with none assumed canonical. The tool runs this as multi-window machinery (an array of windows, each solved by the same kernel; see §5), reporting a per-device σ distribution — median, a 95% CI, the window count and the total simultaneous seconds — rather than a bare point. The analysis tool folder-ingests a whole night's raw captures — auto-detecting the O2Ring, H10 and Verity files, deriving each corner in parallel (Web Workers) and applying a decorrelation quality gate that excludes a night whose Verity extraction failed — and solves the three-cornered hat across the full twenty-six-night corpus: of 58 indexed nights, 31 are trio-eligible and 26 resolve a clean across-night hat (291,561 simultaneous seconds, each retained window ≥1,000 s and SQI-gated), reported in Table 3 as a per-device σ distribution — median, a 95% bootstrap CI over the nights, the night count and the total simultaneous seconds. (An earlier deep hat over six hand-selected clean windows, 63,231 s, was reported previously and is retained only as a historical note in §3.2 below; the folder-ingest corpus supersedes it.) Each retained night is solved by the fused-weight robust hat (§2.3.3′), so a night with a transient single-corner artifact is kept and down-weighted in place rather than discarded: 06-12 — whose spurious-QRS burst would have detonated the plain hat to σH10 ≈ 9.6 — remains in the corpus and contributes a clean σH10 ≈ 2.6. The five non-solved nights of the 31 eligible are excluded for structural, not artifact, reasons — three for insufficient three-way overlap (06-13, 06-17, 06-19, each < 100 s after intersection) and two for a missing Verity HR (06-30, 07-05) — and are surfaced, not hidden. The H10↔O2Ring leg is still carried as a built-in control, and any negative TCH variance (five nights, where H10's variance resolved below zero under the correlated-error test) is surfaced rather than clamped.

Fused-weight three-cornered-hat per-device sigma over the 26-night corpus: O2Ring 2.41, H10 1.28, Verity 1.42 bpm with 95% CI whiskers and per-night dots
Figure 2. Fused-weight three-cornered-hat reference-free σ (live tool output) over the full twenty-six-night folder corpus (291,561 s) — O2Ring 2.41, H10 1.28, Verity 1.42 bpm. Bars are the aggregate (across-night median) per-device σ with 95% CI whiskers (bootstrap over the nights) and faint per-night dots. H10 HR from raw ECG (Pan-Tompkins), Verity HR from raw PPG (PPGDSP), O2Ring native pulse. Each bar is that device's individual variance solved from the three confidence-weighted pairwise difference variances (§2.3.3′) — no device assumed perfect: the raw-ECG H10 is the quietest corner (the gold leg, variance resolving positive in 21/26 nights), the O2Ring pulse the noisiest, Verity between. The 06-12 QRS-burst night is now down-weighted in place to an ordinary σH10 ≈ 2.6 dot (it detonated to ≈9.6 under the plain hat) — no longer a lone high outlier. Same kernel, confidence-weighted, none canonical.
Table 3. Fused-weight three-cornered-hat σ (none canonical) over the full N=26 night folder corpus (291,561 s, production SQI-gated Verity) on the raw-ECG gold leg — each an across-night median with 95% bootstrap CI. Each night solved by the confidence-weighted robust hat (§2.3.3′), so transient single-corner artifacts (06-12) are down-weighted in place, not excluded. Pairwise rows are the raw (unweighted) pooled Bland–Altman over the corpus. H10 = ECG-derived (its σ is the median over the 21/26 nights where its variance resolved positive — the remaining five went negative under the correlated-error test and are surfaced, not hidden).
Device / pairσ — 26-night hat [95% CI]biasSD95% LoAArmsr
O2Ring (pulse)2.41 [2.18–2.61]— device σ, reference-free —
H10 (ECG-derived)1.28 [0.96–1.65]— device σ, reference-free —
Verity Sense (PPG)1.42 [0.96–1.88]— device σ, reference-free —
H10 − O2Ring (control)+0.203.92±7.73.93
H10 − Verity+0.023.62±7.13.62
Verity − O2Ring+0.183.63±7.13.64

The H10↔O2Ring leg doubles as the hat's built-in control leg: per night it holds bias≈0 with an SD of ≈2–4 bpm (the O2Ring's own noise), so a night whose whole alignment drifts far outside that band is flagged as mis-aligned rather than trusted. A within-night transient — 06-12's spurious-QRS burst — is a different failure, and the fused-weight hat now handles it directly: the burst seconds are down-weighted by their collapsed beat-confidence, so the night is retained and its σH10 lands at ≈2.6 (it detonated to ≈9.6 under the plain hat) rather than the whole night being thrown away. The twenty-six-night fused hat ranks σ as H10 (1.28) < Verity (1.42) < O2Ring (2.41) bpm: the raw-ECG H10 is the quietest corner and the O2Ring pulse the noisiest, with Verity's CI overlapping the H10's (the two are statistically close). Historically, an earlier deep hat over six hand-selected clean windows (63,231 s) had instead ranked O2Ring (1.83) ≈ H10 (2.04) < Verity (3.50) — but that reflected a single bad Verity window (σ 6.2 bpm on 06-16/17) dominating a tiny sample; the full corpus, with a cleaner production SQI-gated PPGDSP Verity derivation and more O2Ring motion nights, settles the ordering and shows that the apparent Verity noisiness — not the O2Ring's quiet — was the small-sample artefact. The robust fact is the gold leg: the raw-ECG H10 is the tightest corner. Its hat σ (1.28 bpm) is not in tension with the ≈0.8 bpm repeatability quoted in §2.3/§3.1 — the two measure different quantities: §2.3's 0.8 bpm is the H10's short-term precision (the 1-Hz residual left after a rolling median, i.e. tracking jitter once slow trend is removed), whereas the three-cornered-hat 1.28 bpm is the H10's total reference-free variance over the overlap, which additionally absorbs 1-Hz bucketing and the beat-to-beat granularity of an instantaneous ECG HR against the smoothed peers. A sanity check anchors the reference: the ECG-derived H10 HR matches Polar's onboard R-R to a −0.04 bpm bias (SD 2.27, r 0.85, n 7,224) — two independent QRS algorithms on the same heart agree, so the H10 leg is sound whichever way it is computed.

Read the ranking as error structure, not a plain accuracy order. The three-cornered hat assumes the three devices' errors are uncorrelated, and two physical facts bend that: the O2Ring's reported pulse is internally smoothed (successive seconds rarely change), which on a quiet night deflates its apparent σ below its true instantaneous error; and the raw-recovered Verity HR is instantaneous per beat, so it carries genuine beat-to-beat variability (true HRV) that the smoothed devices average away, inflating its apparent σ. Over the full twenty-six-night corpus a third factor dominates the O2Ring: its ring-PPG is the most motion-susceptible corner, so its σ (2.41) is lifted to the top by the motion nights that the hand-picked deep windows had excluded — the firmware smoothing that flattered it on a few clean windows no longer hides that. The hat still does what no pairwise method can — separate three devices with none assumed canonical — but the honest reading is about error structure (smoothing, true HRV, motion susceptibility), not that one device is simply "more accurate." The decisive, assumption-light lesson is the capture one: log the raw waveform, not the firmware's HR — it turned a dead third corner into a clean one at SQI ≈ 1.0. Concretely: the raw-ECG H10 (1.28) and raw-PPG Verity (1.42) sit close together as the two quiet corners with overlapping CIs, while the 1-Hz smoothed O2Ring pulse (2.41) is the noisiest across the corpus.

4. Discussion

The recipe answers the opening question concretely. Without any certified instrument we can state that the O2Ring's pulse rate is unbiased to within a beat and carries a random uncertainty near 3.2 bpm at rest — a number good enough to trust an overnight average while distrusting any single second, and good enough to flag a night where motion has destroyed the signal. The separation of precision from trueness is what makes this defensible: repeatability needed no reference; bias needed one, and a chest-strap ECG — itself uncertified but an order of magnitude more precise — is a sound stand-in whose residual noise we measured and subtracted. The three-cornered hat then removed even that mild assumption where all three devices co-recorded, putting a number on each device with none held canonical — solved by folder-ingesting each night's raw captures across a twenty-six-night corpus (291,561 s) with a fused-weight robust hat that down-weights transient artifacts in place: the raw-ECG H10 is the quietest corner (1.28 bpm, the gold leg), the O2Ring pulse the noisiest (2.41), and the raw-PPG Verity between (1.42), its CI overlapping the H10's. An earlier six-window deep hat had instead put Verity noisiest (3.50) — a small-sample artefact of one bad Verity window that the full corpus corrects, a reminder that a σ ranking from too few windows is provisional. The sharpest practical lesson was unexpected: the third device's firmware produced nothing, yet its raw signal was pristine — so the right thing to log is the waveform, and the right thing to run on it is a detector you control.

Limitations. (i) Single subject, twenty-six co-recorded three-device nights — a methods demonstration, not a population accuracy statement; the σ is this O2Ring on this finger, at rest. (ii) HR channel only; the O2Ring's SpO₂ trueness is untouched and would need arterial CO-oximetry. (iii) The reference is a consumer ECG strap, not a clinical monitor — fit for a transfer standard, not a regulatory claim. (iv) Resting data compresses the dynamic range, so these figures bound rest-HR error and may understate error during rapid HR change; an exercise protocol would probe the high-slew regime. (v) The three-cornered hat assumes uncorrelated errors; the O2Ring's internal smoothing and the Verity's instantaneous derivation bend that assumption (Table 3 callout), so its σ ranking describes error structure, not a clean accuracy order. (vi) Verity HR is recoverable only where raw PPG was logged; the tool folder-ingests the full twenty-six-night corpus, of which 26 nights clear the ≥1,000-s SQI-gated bar for the hat. One (06-18/19) had a contact-flat H10 ECG, so its reference leg falls back to the strap's onboard R-R rather than ECG-derived QRS — a footnoted substitution, defensible because the two H10 derivations agree to a few hundredths of a percent (PulseDex's RR↔ECG comparator). The five eligible nights that do not resolve are excluded for structural reasons — three for < 100 s of three-way overlap (06-13, 06-17, 06-19) and two for a missing Verity HR (06-30, 07-05) — not for artifact: 06-12, whose QRS burst would have failed a plain hat, is retained and down-weighted in place by the fused-weight estimator (§2.3.3′). The deep and broad hats agree on the H10 gold leg but reorder the noisy corner (Verity on six windows, O2Ring on ten) — a coupled small-sample effect, so a device-σ ranking is reported as provisional until it is stable across the two derivations. With the recommended 5–10 windows (§6) now far exceeded, the reported CI is an across-night bootstrap and the uncorrelated-error assumption is testable; spanning non-resting states remains a capture task. (viii) The fused weighting itself is an assumption: down-weighting confirmed-artifact seconds cleans the estimate but, applied corpus-wide, lowers every corner modestly relative to the plain hat (O2Ring 2.60→2.41, H10 1.50→1.28, Verity 1.56→1.42 here) — the H10 collapse is the intended artifact removal (its per-night spread fell 9.3→2.4 bpm), the O2Ring/Verity drops are the estimator declining to count scattered artifact seconds as device noise. The weights are self-calibrating and arrhythmia-safe by construction (both a density and a quality cue must fire, so clean high-variability rhythm survives), but the numbers should be read as the artifact-robust σ, slightly below the plain-hat σ by design. (vii) The hat cannot be extended to a fourth corner using the H10's onboard R-R: it is the same electrodes and signal as the ECG-derived leg (the two agree to ≈ 0.04% via PulseDex's comparator, error correlation ≈ 1), so it is not an independent measurement. Fed as a separate corner it violates the uncorrelated-error assumption and degenerates the solve (a near-zero mutual variance, typically a negative output); a genuine fourth corner needs a physically separate sensor, which is why the onboard R-R serves only as a same-device cross-check and the 06-18/19 fallback leg. (ix) Every σ reported here is variance-only, and the estimator itself has never been validated against an external truth. The three-cornered hat carries no bias term at all — it partitions scatter among the corners and is structurally blind to any offset. No per-device bias can currently be quoted here at all, and the attempt to quote one is instructive. A cross-node comparison per 5-min epoch against the raw-ECG Pan–Tompkins leg appears to show the O2Ring under-reading by −0.269 bpm (n = 3,136 epochs over 40 nights, 11.0 σ from zero). That figure is an estimator confound, not a device bias. The two nodes summarise an epoch differently — ECGDex reports 60000/mean(RR), the rate of the mean interval, while OxyDex reports the median of the 1 Hz rate — and pushing one real series through both statistics, with no device involved, yields median(rate) − 60000/mean(RR) = −0.299 bpm (SD 0.49) over 1,670 300-beat blocks: the confound is the size of the finding. Measured directly instead, against the ring's own photoplethysmogram over 20 nights and 237 windows (tools/o2ring-finger-validate-batch.mjs), the ring's firmware HR is statistically indistinguishable from the chest ECG (−0.027 bpm, 0.6 σ). The mechanism behind the −0.299 has since been isolated: it is the shape of the interval distribution, dominated by its variability. Regressing the per-block gap on each block's own R-R statistics over 1,670 real 300-beat blocks gives gap ≈ 0.2989 − 8.7175·CV + 0.2121·skew (R² = 0.601, residual SD 0.309 bpm against a raw 0.489), with r(gap, CV) = −0.719 and r(gap, skew) = +0.690 while heart-rate level itself is negligible (r = −0.134). Real overnight R-R has mean CV 0.0522 and mean skew −0.671; substituting those returns −0.298 against the measured −0.299. That also accounts for the synthetic series quoted above — smooth, pause-injected and trended runs give +0.03, +0.54 and −0.03 because each carries a different CV and skew, not because the effect is capricious. The gap is therefore a property of the interval distribution the two statistics are applied to, and of nothing about either device — which is precisely why it cannot be attributed to one. No per-device HR bias may be read off cross-node epoch HR until the nodes agree on one statistic. So the hat's blindness to bias is not merely structural here: there is at present no validated bias figure for any corner to be blind to. Nor can this corpus test the assumption its σ rests on: because the chest-ECG is one of the three corners, the hat's σECG² is algebraically identical to the covariance of the other two corners' reference-relative errors (verified numerically to 7×10⁻¹⁴ on the committed corpus), so an independence or accuracy check run against that reference returns its own input and has exactly zero power — any "excess correlation" reported that way would be fabricated. A hat cannot be validated using one of its own corners as the reference; it requires a genuinely external Nth device, chosen so the truth leg is not mechanistically twinned with a corner (the ResMed oximeter's 1 Hz pulse is the candidate here, though it shares photoplethysmography with the O2Ring). The σ values reported above are therefore precision estimates, not accuracy statements. (x) σ here is a σ at one window length, and that parameter is not stated anywhere above. The hat is fed 291,561 simultaneous seconds over twenty-six nights — a median of roughly 11,200 s per night — but window length is treated throughout as a consequence of what the recordings happened to contain, not as an analysis choice. It is one. Re-running this same fused-weight estimator on an independent seventeen-night corpus (the same three devices, box-captured, host-axis corrected) while varying only how many simultaneous seconds reach the hat gives a monotonic rise in every corner: σVerity 2.36 → 3.51 (+49 %), σO2Ring 2.34 → 2.99 (+28 %), σH10 1.41 → 1.78 (+26 %) from a one-hour window to a whole night (tools/tch-window-sensitivity.mjs). The mechanism is not mysterious — a longer window admits more of the restless margins of the night, where the corners disagree most — but the consequence is that two analysts with the same devices, the same nights and the same estimator can publish σ differing by half again, and neither is wrong. The figures above should therefore be read as σ at ≈11.2 ks per night, and are not directly comparable to a σ derived from shorter or longer windows. Two honest caveats were attached to this caveat, and one of them was false. It read: "the sensitivity has not been characterised on this paper's own corpus, whose raw capture is no longer available to re-derive." The raw capture is available — it is the phone-captured tree these nights were logged into, still on disk, with all three streams present for every night named in this section. The claim was never checked; it was inferred. It has now been characterised on this paper's own nights (2026-08-15, 24 three-way-eligible nights re-folded by tools/trio-batch.mjs, swept by tools/tch-window-sensitivity.mjs), and the result both vindicates and corrects the surviving caveat:
cornerfull committed corpus (52 nights, 2026-08-15)this paper's own nights (24 nights)as published, box corpus (17 nights)
O2Ring2.18 → 2.84 (+30 %)1.98 → 2.57 (+30 %)2.34 → 2.99 (+28 %)
H100.78 → 1.13 (+45 %)0.61 → 0.92 (+51 %)1.41 → 1.78 (+26 %)
Verity0.58 → 0.72 (+24 %)0.45 → 0.47 (+4 %)2.36 → 3.51 (+49 %)

The caveat that the magnitude may not transfer was right, and the dependence replicates everywhere: every corner's σ is larger over a whole night than over an hour, on all three corpora. Beyond that, the corpus has now grown enough to correct the reading this section carried on 2026-08-15, and the correction is a lesson about sample size rather than about the devices. That reading called the per-corner magnitudes a reorder — Verity largest on the box corpus (+49 %) and smallest here (+4 %), H10 the reverse. Re-run across the full committed corpus — 52 nights and 903,265 simultaneous seconds, a 3.1× larger sample than the 26 nights and 291,561 s behind this paper's headline — the Verity corner moves from +4 % to +24 %, i.e. most of that dramatic contrast was the 24-night subset, not a property of the capture path. What survives corpus expansion is narrower and more useful: the O2Ring corner transfers (+28 / +30 / +30 % across three independent corpora), the H10 corner genuinely does differ by capture path (+26 % on the box against +51 % and +45 % on the two phone-captured samples), and the Verity corner is simply unstable across samples (+49 / +4 / +24 %) — which is itself the finding, since it is the corner whose sensitivity a practitioner would most want to borrow. (xi) σ here is also a σ at one DSP generation — that parameter is not stated either, and it is not per-corner. The recipe derives every corner from raw signal, so each σ is a function of the code that produced the intervals, and that code moves: ppgdex-dsp.js changed 20 times in the three weeks after 2026-08-08, including a filtfilt running unpadded from zero state (a DC-sized transient at both record ends), the frequency domain computed over correctRR's substituted intervals, and a crystal axis running backward that was hiding real dropouts. Measured on one night (2026-08-04, 22.5 ks, H10 and O2Ring corners held byte-identical, only the PPG generation varied): σVerity 2.14 → 4.25 bpm. And because the hat is coupled, the same swap moved σH10 1.64 → 1.85 with the ECG code unchanged — so a per-corner σ from this estimator is a function of the generation of all three DSPs, not of its own. ⚠️ The generation behind this paper's headline 2.41 / 1.28 / 1.42 is not recorded anywhere, so unlike the corpus (26 nights) and the sample (291,561 s), which are stated, those figures are not presently re-derivable. A reference-free σ needs three things named to be reproducible — corpus and n, window length, and the generation of every corner's DSP — and this paper states the first only.

The second published statement — that the rise is monotonic in every corner — fails, and fails wider than first reported. On the full corpus two corners are non-monotonic, not one: Verity runs 0.58 → 0.54 → 0.56 → 0.67 → 0.73 → 0.67 → 0.67 → 0.72 across the sweep (peaking at 11.2 ks, the paper's own window length, then falling back), and the H10 dips 0.78 → 0.74 before rising to 1.13. Only the O2Ring is monotonic throughout. "Monotonic in every corner" describes the box corpus; it does not describe the phenomenon.

What this comparison is and is not. The sweep runs through the node-export path, not the raw-ingest fused-weight hat that produced this paper's headline σ (1.28 / 1.42 / 2.41 bpm) — the two paths were known not to agree in absolute terms (tools/tch-fused-corpus.mjs re-fit the published hat on node-exports and did not reproduce it). So the absolute σ in the table above must not be read as a restatement of the headline figures; only the relative window-length sensitivity is the comparable quantity, and only that is claimed above.

That blanket "the paths disagree" is now too coarse, and controlling window length is what sharpens it. Truncating the 52-night node-export sweep to this paper's own 11,214 s median window — rather than comparing a whole-night σ against an 11.2 ks σ, which is precisely the error limitation (x) exists to name — gives O2Ring 2.44 [2.28–2.77], H10 0.93 [0.68–1.13], Verity 0.72 [0.47–0.91] bpm against the published 2.41 / 1.28 / 1.42. The O2Ring corner reproduces across the two pipelines (2.44 vs 2.41, within 1.2 %, the published value sitting inside the CI); the H10 and Verity corners do not — each published value falls outside its interval, both in the direction of the node-export path reporting a quieter corner. So the path disagreement is not global: it falls on the two corners the node-export path summarises most aggressively, and it was partly masquerading as a path effect while being a window-length effect. ⚠️ It must not, however, be read as evidence about those two corners' own summarisation. The hat is linear in the pairwise variances — σ²H10 = ½(VHV + VHO − VVO) — and two of those three contain the Verity corner, so a change confined to one corner's processing moves all three recovered σ. Measured 2026-08-27 on one night (2026-08-04, 22.5 ks) with the H10 and O2Ring corners held byte-identical and only the PPG DSP generation varied: σVerity 2.14 → 4.25 and σH10 1.64 → 1.85, with the ECG code unchanged. “Corner X's σ moved, therefore corner X's processing differs” is therefore not a valid inference from this estimator, and the per-corner reconciliation owed below must hold the other two corners fixed to mean anything. A per-corner reconciliation of the two pipelines is owed and is not yet done; characterising the sensitivity through the raw-ingest path on these same nights remains possible and likewise not yet done. This is the same discipline this project already applies to clock rate, where a ppm figure may never be quoted without the span it was measured over.

5. Reproducibility

6. Sample size & statistical power

Unlike the simulation pilots, this paper runs on real captured data, so “power” means having enough co-recorded time — and enough simultaneous three-device time — rather than synthetic patients. Two different sample sizes matter: paired seconds set the precision of the bias/SD/Arms figures (their SE falls as ~1/√n_seconds, and one night already supplies >20,000), while co-recorded nights/sessions set how well we separate a stable device-σ from night-specific artifacts, and three-device overlap windows are the binding constraint on the three-cornered hat.

Table 4. Data-sufficiency guidance for reference-free σ (real co-recorded captures).
QuantityMinimum (acceptable)RecommendedDiminishing returns
Bias / SD / Arms vs transfer standard1 clean night (~20k paired s) → SD to ≈±0.05 bpm5–7 clean nights — separates a stable σ from night artifacts; lets you flag (not average over) a bad night> ~15 nights: per-night σ already stable; extra nights mainly characterize night-to-night spread, not the central σ
Three-cornered-hat per-device σachieved: 26 nights, 291,561 simultaneous s (each ≥1,000 s, SQI-gated) — tool reports an across-night 95% bootstrap CI5–10 overlap windows on different nights — tightens the across-window CI and probes the uncorrelated-error assumption furtheronce windows span varied HR/motion states; more identical resting windows add little
HR dynamic range probedresting only (this run)+1 exercise/recovery session — bounds high-slew error the rest data can't see
Subjects1 (methods demonstration)≥10–20 for any population accuracy statementset by the claim, not by σ precision

Practical reading for this dataset: the per-night HR-error σ is already well-determined (twenty-six nights, 292k paired seconds), so additional captures buy the most where we are currently thinnest — more simultaneous three-device windows (every session that logs Verity raw PPG alongside the H10 adds a corner), and at least one non-resting session to bound error during rapid heart-rate change. More resting single-device nights add the least. This directly shapes what to prioritize as further data is added: capture all three devices together, keep the raw waveforms, and include some movement.

References

  1. D. W. Allan, “Statistics of atomic frequency standards,” Proc. IEEE 54(2):221–230, 1966 — the variance framework underlying frequency-stability estimation. doi:10.1109/PROC.1966.4634
  2. J. E. Gray & D. W. Allan, “A method for estimating the frequency stability of an individual oscillator,” Proc. 28th Ann. Symp. Frequency Control, 1974 — the three-cornered-hat variance decomposition. doi:10.1109/FREQ.1974.200027
  3. A. Premoli & P. Tavella, “A revisited three-cornered hat method for estimating frequency standard instability,” IEEE Trans. Instrum. Meas. 42(1):7–13, 1993 — handling negative and correlated variance estimates. doi:10.1109/19.206671
  4. F. Torcaso, C. R. Ekstrom, E. A. Burt & D. N. Matsakis, “Estimating frequency stability and cross-correlations,” Proc. 30th PTTI Meeting, 1998 — the N-cornered hat with correlated sources (why a non-independent corner degenerates the solve).
  5. W. J. Riley, “Handbook of Frequency Stability Analysis,” NIST Special Publication 1065, 2008 — practical multi-source variance estimation.
  6. J. M. Bland & D. G. Altman, “Statistical methods for assessing agreement between two methods of clinical measurement,” Lancet 1(8476):307–310, 1986 — bias and 95% limits of agreement. doi:10.1016/S0140-6736(86)90837-8
  7. J. M. Bland & D. G. Altman, “Measuring agreement in method comparison studies,” Stat. Methods Med. Res. 8(2):135–160, 1999 — repeated-measures limits of agreement. doi:10.1177/096228029900800204
  8. D. Giavarina, “Understanding Bland Altman analysis,” Biochem. Med. 25(2):141–151, 2015 — interpretation of difference-vs-mean plots, including discretization/banding artifacts. doi:10.11613/BM.2015.015
  9. J. Pan & W. J. Tompkins, “A real-time QRS detection algorithm,” IEEE Trans. Biomed. Eng. 32(3):230–236, 1985 — the R-peak detector applied to the raw H10 ECG. doi:10.1109/TBME.1985.325532
  10. J. Allen, “Photoplethysmography and its application in clinical physiological measurement,” Physiol. Meas. 28(3):R1–R39, 2007 — PPG signal fundamentals. doi:10.1088/0967-3334/28/3/R01
  11. A. Schäfer & J. Vagedes, “How accurate is pulse rate variability as an estimate of heart rate variability?” Int. J. Cardiol. 166(1):15–29, 2013 — PPG- vs ECG-derived heart rate. doi:10.1016/j.ijcard.2012.03.119
  12. M. Gilgen-Ammann, T. Schweizer & T. Wyss, “RR interval signal quality of a heart rate monitor and an ECG Holter at rest and during exercise,” Eur. J. Appl. Physiol. 119:1525–1532, 2019 — Polar H10 validation against ECG.
  13. B. Bent, B. A. Goldstein, W. A. Kibbe & J. P. Dunn, “Investigating sources of inaccuracy in wearable optical heart rate sensors,” npj Digit. Med. 3:18, 2020 — error sources in PPG heart rate (motion, perfusion).
  14. A. Jubran, “Pulse oximetry,” Crit. Care 19:272, 2015 — oximetry principles and accuracy limits (why SpO₂ trueness is not testable here).
  15. ISO 80601-2-61, “Medical electrical equipment — particular requirements for basic safety and essential performance of pulse oximeter equipment” — the accuracy root-mean-square (Arms) convention; to be cited precisely at submission.
  16. Project documentation: CLAUDE.md (Clock Contract, capture provenance), Tepna suite; analysis apparatus sigma-no-reference-analysis.html; detectors ecgdex-dsp.js (ECGDSP) and ppgdex-dsp.js (PPGDSP); same-device RR↔ECG cross-derivation comparator in pulsedex-app.js.
T © 2026 Michal Planicka ·Tepna v1.0.0 ·Apache-2.0 ·◈ Asheville, NC ·not a medical device
v2.8.0