← Tepna preprints

A map of the walls: negative results in consumer physiological-signal analysis

Michal Planicka  ·  corresponding author — Tepna Project

Cross-node synthesis, Tepna physiological-signal suite

Draft v2 (wall 2.7 corrected, wall 2.8 added — 2026-07-29) · Synthesis paper — every figure regenerates from the cited source paper's own named analysis tool · 100% local

Abstract

Motivation. A suite of physiological-signal analyses accumulates as much knowledge in what it has ruled out as in what it has confirmed, yet negative results are almost never collected into a citable artifact. This paper is that artifact for the Tepna suite: a structured map of the walls — the analyses that looked promising and did not survive scrutiny, each with the evidence that felled it and its current disposition (fixed, flagged, or fundamental). Contents. Eight results: (1) a single-value optical PRV is not a drop-in for ECG/RR HRV — unbiased in the mean but carrying ~22× wider limits of agreement; (2) a per-beat signal-quality index stays green while optical beat-yield fails, inflating rMSSD a median +83%; (3) daily glycemic variability (CGM-CV) has near-zero test-retest reliability (ICC₁≈0) — a state, not a one-number trait; (4) a glucose↔HRV "coupling" collapses from −0.11 to +0.01 once a shared apnea driver is partialled out; (5) a rolling-baseline ODI-4 systematically under-counts events in severe apnea (slope 0.23 vs truth) — since traced to a trailing-mean baseline artifact and corrected at the detector level; (6) black-box vendor wellness composites emit a fabricated 0 when their inputs are absent; (7) beat-level cross-device wearable PAT is unrecoverable (0 of 54 pairings) — but not, as this paper's own first version claimed, because of ~48 ppm clock drift: that "~1147 ms/night drift" was one cardiac cycle admitted by an unenforced coupling window, the true inter-device rate is ~1.5 ppm rather than 48 (v1's rate is excluded on 51 of 54 nights), the prescribed single-clock remedy was applied and does not help — that path caps accumulated drift at 8.6 ms even at the claimed rate, and still couples 18.8 % vs 19.0 % — and — corrected 2026-08-13 — the ~96 ms of "peripheral beat-to-beat scatter" that v2 named as the real limit is itself an artifact: it is the standard deviation of a fixed 450 ms acceptance window (450/√12 = 129.90 ms), measured through an ECG rate rounded to nominal (46–126 ppm) on a corpus with no second clock. The wall's cause is now open, not settled; (8) a negative result from an uncalibrated search is not evidence — widening a cross-correlation to ±50 min drives the chance-maximum correlation (≈0.41) above the acceptance threshold it was lowered to (0.35), so the search reports ≈0 whatever the truth is; demonstrated by a known-answer control that recovers a planted −39 min offset on 1 night in 13. Value. The contribution is framing and honesty, not new compute — a public account of where the field's intuitions break, so others need not re-walk into the same walls. Wall 8 is the one this paper walked into itself, and §2.7 carries the correction record.

Keywords: negative results · reproducibility · consumer wearables · heart-rate variability · pulse oximetry · continuous glucose monitoring · signal quality · confounding · shared-driver effects · pulse arrival time · cross-device clock drift · scientific honesty

0. Layman overview (delete before submission)

Most research papers tell you what worked. This one tells you what didn't — on purpose, because the failures are just as useful and almost nobody writes them down. Eight of them, from analyzing everyday health-gadget data: the wrist blood-flow sensor can't stand in for a chest strap for the fine heart-rhythm score (it's right on average but wildly variable); a "signal looks good" light stays green even while that same sensor is quietly losing heartbeats and inflating the score; a day's blood-sugar-variability number tells you almost nothing stable about a person; an apparent glucose–heart link turns out to be a third thing (breathing pauses) driving both; a cheap sleep-apnea count under-reports the worst cases (we found why and fixed it); and some gadgets' “wellness scores” silently print a zero when they actually have no data; and two health wearables recording the same night through one phone can't be lined up finely enough to time the tiny delay between heartbeat and pulse. That last one is where we caught ourselves being wrong. We had blamed the gadgets' clocks drifting about a second apart over a night, and said the fix was to record both sensors on one device. Neither was true. Our software was allowed to match a heartbeat to a pulse up to two seconds later — longer than the gap between two heartbeats — so whenever it missed a pulse it grabbed the next one, and the "one second of clock drift" was simply one heartbeat. The clocks were fine all along: the real disagreement is about thirty times smaller than we claimed. And we did try recording through one little computer — it keeps time to a few millionths of a second and re-checks both gadgets' clocks every few minutes, so they cannot drift apart by more than a few thousandths of a second. It made no difference at all. What actually defeats it is that the pulse itself arrives a tenth of a second early or late from beat to beat. So the eighth wall is the general lesson: if you conclude that something can't be found, first check your tool could have found it — we tested ours on a case where we already knew the answer, and it failed there too. None of this is a discovery in the usual sense — it's a map of the dead ends, so the next person doesn't waste the trip, including one trip we wasted ourselves.

1. Introduction

The Tepna suite's central discipline is to distrust its own outputs until a deterministic test earns them. That discipline produces a steady stream of negative results — analyses that a reasonable person would expect to work and that, on inspection, do not. Scattered across a dozen preprints, each negative result is a footnote to a positive one; collected, they form something more useful: a map of where consumer physiological-signal analysis breaks, which is exactly the knowledge a practitioner needs before trusting any of these signals. This paper assembles eight such walls. Each entry states the intuition, the evidence against it, and — critically — its current disposition, because a wall that has since been patched is a different object from one that is fundamental. We distinguish three dispositions: fundamental (a real property of the physics or statistics, not fixable), flagged (surfaced honestly to the consumer rather than corrected), and fixed (traced to a mechanism and repaired). No number here is new; each regenerates from the cited source paper's own named analysis tool.

A collection of negative results carries an obligation the positive literature can shirk: a null is only as strong as the demonstration that the instrument could have returned a positive. We failed that obligation once, in this paper, and §2.7 is the record. Wall 2.7 was published with the right conclusion and the wrong mechanism, and the refutation was already printed in its own table — visible to one correlation. Rather than quietly restate it, the entry retains the original claim, shows what falsified it, and the failure mode is promoted to wall 2.8 as a finding in its own right. A map of the walls that concealed the one wall its own authors walked into would be worth less than this.

2. The walls

2.1 Optical PRV is not a drop-in for ECG/RR HRV — fundamental (variance, not bias)

The intuition is that a wrist/arm optical pulse gives you the same heart-rate-variability number as a chest ECG. On shared synthetic beats scored by the real detectors (rmssd-equivalence), ECG- and RR-derived rMSSD are genuinely interchangeable (bias −0.02 ms, Pearson r = 0.9999), but the optical arm is unbiased in the mean (+0.3 ms, +0.7%) while carrying ~22× wider limits of agreement (r = 0.93). The honest failure is therefore not a bias to subtract but a precision penalty: a single optical rMSSD reading cannot be trusted as a point estimate the way an ECG reading can, even though a large pooled average of optical readings converges to the truth. (Note: an earlier draft's headline "+32% optical bias" did not survive re-running on a broadband RR texture — the divergence was a texture artifact, not a real bias; correcting that is itself a negative result about the first analysis.)

External confirmation (added 2026-08-23): Kantrowitz AB et al. 2025, Pulse rate variability is not the same as heart rate variability, Front Physiol 16:1630032 (doi: 10.3389/fphys.2025.1630032) — beat-to-beat comparison in 931 adults: PPG-PRV significantly underestimates SDNN, rMSSD and pNN50 versus ECG-HRV across chronic-disease groups. Independent, larger-N support for this wall’s direction; their design (chest+bicep, rest) differs from this suite’s overnight wrist/arm corpus, so magnitudes are not transferable.

2.2 A per-beat SQI stays green while optical beat-yield fails — flagged

The intuition is that a signal-quality index guards the beat train — if SQI is high, the beats are good. In the FULL-lane yield harness (qrs-yield, ≈363 windows/arm, ≈188k true beats), electrical QRS recall was ≈100% in both clean and apnea segments, but the optical arm recovered only ~96% and over-detected, and its mean SQI fell only from 0.90 to 0.78 under apnea — not low enough to mark the bad beats unusable. The downstream cost is large: rMSSD reconstructed identically from detected vs true beat-times was inflated a median +83% (+34 ms). SQI reports morphology plausibility, not beat completeness, so it cannot be the sole gate on HRV; the suite's response is to down-weight optical HRV on the event channel rather than trust SQI alone.

2.3 Daily glycemic variability is a state, not a trait — fundamental

The intuition is that a person's CGM coefficient-of-variation is a stable personal number. In a 6,000-subject stable cohort (nights-icc), daily CGM-CV had a single-occasion ICC₁ ≈ 0.00 — essentially no between-subject variance — against rMSSD's 0.93 and ODI-4's 0.75 on the same harness. Stable people's days look alike in CV while each person's day-to-day swings dominate, so no number of averaged days makes a single CGM-CV a reliable individual trait (Spearman–Brown minimum occasions → ∞). Reporting it as a personal metric is the wall; it is a state descriptor.

2.4 A glucose↔HRV coupling is a shared driver, not a link — fundamental (confound)

The intuition is that a measured within-person glucose–HRV correlation reflects a direct physiological link. In the coupling analysis (cgm-hrv-coupling), partialling out the planted apnea burden collapsed the within-patient coupling from −0.11 to +0.01 — effectively zero — while the two driver legs (apnea→glucose, apnea→HRV) were strong and opposite-signed. The apparent coupling was almost entirely a third variable acting on both. The lesson generalizes: any cross-signal correlation in this domain must be tested against the obvious shared drivers before it is called a link.

2.5 Rolling-baseline ODI-4 under-counts severe apnea — fixed

The intuition is that an oxygen-desaturation-index event counter is severity-neutral. It is not: with a trailing-mean SpO₂ baseline (odi4-ahi-bias) ODI-4 was strongly linear in reference AHI but with slope 0.23 (R² 0.93) — recovering only ~a quarter of scored events, worst in severe OSA. The mechanism was self-suppression: dense event clusters drag a trailing-mean baseline down into the desaturations, so subsequent drops no longer clear the 4% threshold. This one is fixed — a ceiling (rolling-max) baseline raised the cohort slope from 0.42 to 0.69 without inflating the no-apnea stratum. It earns its place here as a documented wall and its repair — the model of what to do with a negative result.

2.6 Black-box vendor composites fabricate a zero — flagged (ingest guard)

The intuition is that a vendor "wellness/readiness" composite degrades gracefully when an input is missing. Some do not: when their black-box constituent inputs are absent they emit a fabricated 0 that is indistinguishable, downstream, from a true zero score. This is the same class of error as the timestamp parsers that return now() on a miss — a fabricated value where a null belongs. The suite's disposition is to treat such composites as untrusted at ingest and reject the sentinel 0, consistent with the never-fabricate-on-miss rule established in the Clock Contract benchmark.

2.7 Cross-device wearable PAT is drift-dominated limited by peripheral beat-to-beat scatter unrecoverable — cause OPEN — fundamental (both capture paths); a single acquisition clock does NOT fix it

The wall is real: beat-level cross-device PAT is not recoverable — 0 of 54 pairings across two corpora clear a coupling ≥ 55 % / IQR ≤ 60 ms gate. Two wearables logged by one phone (a chest H10 ECG + a peripheral Verity PPG, Polar Sensor Logger) are the same heartbeats to 0.02 % beat parity, and the "phone timestamp" is not a shared clock — it equals each device's own sensor clock to within ms over 6.6 h, because the logger writes start + device-elapsed. Both of those observations stand. What the wall is made of does not.

Corrected 2026-07-29 — this wall was mis-attributed for a year. The original entry read: "the R→foot lag drifts ~1.1 s/night at a fixed ~48 ppm — ordinary quartz-crystal tolerance… so it needs a single acquisition clock (host-side SDK capture)", with 89 % coupling and 48 ms beat-to-beat as supporting figures. All of that is now refuted:

Corrected 2026-08-13, carried into this section 2026-08-17 — the "~96 ms" below is an ARTIFACT, and the disposition drawn from it is withdrawn. The abstract has carried this correction since 2026-08-13; this paragraph did not, so a reader met the retracted mechanism in the very section that states the finding, and the retraction only in the abstract — the same ordering hazard this paper files as wall 2.8. The figure is the standard deviation of a fixed 450 ms acceptance window (450/√12 = 129.90 ms), measured through an ECG rate rounded to nominal (46–126 ppm) on a corpus with no second clock: it is a property of the search window, not of the peripheral pulse. Every PAT SD measured through that window inherits it. What replaces the number is deliberately not a single figure: within-bin σ is ~68 ms (46–94) in one analysis, and 10–23 ms on three of six nights on box captures with the axis fixed in another — neither gate-backed, and measured under different conditions, which is precisely why no new “true limit” is named here. The wall stands — 0 of 54 pairings clear the gate — but its cause is OPEN, not settled, and no disposition may be drawn from the number below. Sibling record: One phone is not one clock (v3 withdrawal).

Superseded (v2, retained for the record): What actually limits it: ~96 ms of beat-to-beat scatter in the peripheral pulse-foot time, against a 60 ms requirement — roughly twice too much, and not a timekeeping quantity. The R→foot interval is stable in its centre and loose in its detail. The disposition changes from "needs better hardware sync" to "needs better peripheral foot timing" ← withdrawn with the number it rested on, and the outstanding measurement is a beat-correspondence audit: matching beat counts to 0.02 % refutes net dropout but not local insertion/deletion pairs, which preserve the total while scrambling which foot belongs to which beat. Full correction record: One phone is not one clock — but the clock was never the problem §7.

Update 2026-08-15 — the wall stays open, and the evidence base for re-opening it is thinner than the night count suggests. A re-fold of the whole capture corpus under current code applies three-source closure to every night rather than to the contested ones (One phone is not one clock §3.6): of 52 nights carrying a pairwise H10↔Verity rate, 17 close consistently, 10 fail closure outright, and 25 cannot be tested for want of a third source — so 67 % of the per-night rates in this corpus are not measurements, and the three populations are indistinguishable from their ppm values alone. This does not change this wall's disposition, because the wall was refuted on mechanism (beat-slip, then the 450 ms acceptance-window artifact) rather than on a rate. It changes what a future re-opening may lean on: any timing claim brought against this wall has, at present, a 17-night screened base rather than a 52-night one, and pooling the unscreened remainder is actively misleading — the pooled corpus median (+5.0 ppm) carries the opposite sign to the closed-only median (−3.0 ppm). Filed here because this is the pattern wall 2.8 names: the size of a dataset is not the size of its evidence.

Why this wall is the most instructive one in the set. Every other entry here was caught by a check we ran. This one we got wrong, published, and then caught — and the refutation was sitting in our own printed table the whole time, requiring one correlation to see. It is entered as wall 2.8 below in its own right.

Literature parallel (added 2026-08-23): the synchronization prerequisite this wall keeps hitting is documented at scale by Liang Y et al. 2019, How effective is pulse arrival time for evaluating blood pressure?, J Clin Med 8(3):337 (doi: 10.3390/jcm8030337): MIMIC waveform channels are widely assumed synchronized and are not, undermining PAT-based BP evaluation built on them. Same failure class, independent corpus.

2.8 A negative result from an uncalibrated search is not evidence — methodological; fixed by a known-answer control

The generalisation of §2.7's failure, and the reason it is listed as a wall rather than an erratum. Three of this suite's negative findings were reached by a search that was never asked to recover a known answer, and two of them were wrong.

The mechanism is a multiple-comparisons blowout that a lowered acceptance threshold conceals. Widening a cross-correlation search from ±1.6 s to ±50 min at 250 ms resolution takes the candidate-lag count from 65 to 24,001, while a ±15 s correlation window supplies only 121 samples. For ~121 quasi-independent samples sd(r) ≈ 0.091, so the expected best-of-24,001 spurious correlation is ≈ 0.091·√(2 ln 24001) ≈ 0.41. Any threshold below that — and the wide search had been lowered to 0.35 precisely because it was "finding nothing" — is cleared by chance, and the median of the resulting noise regresses to the centre of the search window. So the search reports ≈ 0 whatever the truth is, and a null looks like a measurement.

Demonstrated on a pair whose answer is known by construction (two accelerometers, one acquisition clock, true offset 0): the wide search recovers a planted −39 min offset on 1 of 13 nights, attenuated to −14.3 min; the same code at its design range aligns 13/13. Lengthening the correlation window restores the gate monotonically — ±60 s → −27.3 min (3/13), ±240 s → −31.4 min (6/13) — confirming the mechanism, though it never converges. Two published conclusions rested on searches in this regime: §2.7's accelerometer re-sync failure (withdrawn), and a CPAP-clock fit from flow-vs-movement coincidence (void — the question is reopened, not answered).

Disposition — fixed, and cheap. Calibrate the search before believing its null: run it on a case with a known answer, and check that the acceptance threshold exceeds the chance maximum for the window length and lag count in use. The control is committed as tools/acc-acc-control.mjs. Where a positive claim invokes a physical mechanism, test one prediction that mechanism makes and its rivals do not — for §2.7 that was proportionality to recording duration, one correlation over an already-published table.

2.9 A static analysis that reports zero has not said the code is clean — methodological; caught by a known-answer control

§2.8 in a second domain, and the reason it is a wall rather than a tooling note: the same failure recurs wherever a search can return a null, and a static null is more persuasive than a statistical one because nothing looks noisy.

A test-quality analyser was written to find assertions that compute their own answer and never reach the module under test — a real family, one instance of which had just been removed from the suite. Reading the tests lexically, it reported 132 such assertions, then 96, 72, 5 and finally 0 as five successive false-positive classes were identified by hand and fixed (reassignment, accumulator flags set inside loops, in-place mutation, nested-bracket index assignment, out-parameters). At zero it also reported its own reach — 4,901 assertions examined, 99 % resolved to the module — which reads as a clean bill of health for a 32,895-line suite.

It was not. Planting one assertion of exactly the shape being hunted — an array built in the test, compared with itself, never calling the module — the analyser failed to flag it, and counted it among the 99 % that "reached the module". The mechanism is name resolution without scope: the declaration map is file-wide, and in a file this size short identifiers collide en masse — s merges 261 distinct assignments, a 223, m 143, across 2,950 names. Almost any expression using a short local therefore "reaches" the module through an unrelated declaration hundreds of lines away. The 99 % measured collisions, not coverage, and every intermediate count above is untrustworthy in the opposite direction.

Disposition — abandoned, and the tool deleted rather than shipped. A zero that means "could not resolve" is indistinguishable, to a reader, from one that means "clean" — the precise failure the analyser existed to expose, reproduced by the analyser. The control that settled it cost one planted case and ran in seconds, after five rounds of hand-tuning had not. Written up as briefs/JS-SEALED-ASSERTION-DEAD-END-2026-08-05-BRIEF.md, which records what a working version would need (an AST with scope resolution, control dependence, out-parameters) so the attempt is not repeated. The equivalent tool on the Python side, where files are small and single-scope, works and is validated by full-suite mutation.

3. Discussion

Three patterns run through the first six walls. First, unbiased is not the same as usable: the optical-PRV and reference-free-σ results both show a metric can be centered on truth yet too imprecise to act on from a single reading (2.1). Second, a quality signal must measure the failure that actually occurs: SQI guards morphology but not completeness, so it is silent on the yield failure that dominates the HRV error (2.2). Third, a number is only a trait if its variance separates people, and correlation in this domain is routinely a shared-driver artifact (2.3, 2.4). Two further walls are about honesty of missingness — a fabricated 0 or a fabricated now() is worse than an explicit null (2.5's mechanism, 2.6) — the through-line to the suite's timestamp and provenance discipline.

The last two walls add a fourth pattern, and it is the sharpest one in the collection: a measurement can be wrong in a way its own companion metric is structurally unable to report. Wall 2.7's coupling window was unenforced, admitting a whole extra cardiac cycle, and the beat-to-beat interquartile spread that sat beside the headline number cannot move under that failure — ten slipped beats in sixty do not shift a quartile. So the error produced a large, stable, physically-plausible number with a healthy-looking check beside it. Wall 2.8 generalises the same shape to search: lower an acceptance threshold below the chance-maximum correlation and the search returns a confident value that is pure noise, centred on the middle of its own search window. In both cases the diagnostic that felt like independent confirmation was blind by construction. The remedy in both cases is the same and costs almost nothing: ask what the check could report, and where a mechanism is claimed, test one prediction it makes that its rivals do not — for 2.7, whether a "clock drift" scales with recording duration (it did not: r = +0.17).

The dispositions matter as much as the findings: one wall is fixed at the detector, two are flagged to the consumer, one is fixed methodologically (2.8, by a committed known-answer control), and four are fundamental. Wall 2.7 changed disposition in this revision — from "fixable at capture time with a hardware sync" to genuinely fundamental for this sensor pair, since the quantity that fix removes was finally measured (~1.5 ppm) and is ~30× smaller than the fault it was invoked against. That reclassification is the single most consequential edit in the paper's history, and it moved in the direction that costs us something: a wall we thought we could climb, we cannot. Publishing the fundamental ones is the point — they are the boundary of what consumer signals can honestly claim, and a boundary drawn one wall too generously is worse than no map.

Limitations. This is a synthesis paper: it introduces no new data or compute, and each wall inherits the limitations of its source paper — most are synthetic ground truth scored by the real detectors, so the magnitudes (the +83% inflation, the ~22× LoA ratio, the 0.23 slope) are properties of the generator's noise models as much as of the physics, and the real-data arms (the σ and PPG-vs-ECG papers) are single-subject pilots. The map is therefore of the walls this suite has walked into, not a complete census of the field's dead ends; new devices and analyses will add entries. Three walls (2.1's superseded +32% figure, 2.5's fixed baseline, and 2.7's withdrawn drift mechanism) are explicitly included because they moved — a negative result about an earlier analysis is still a negative result worth recording. 2.7 is the strongest case for that policy and the least comfortable: it moved because this paper asserted a cause it had not established, so the entry now carries both the original claim and its refutation. A limitation specific to it: the known-answer control that withdrew its anatomical explanation pairs a chest accelerometer with an arm one, while 2.7's corpus placed the PPG on the ankle, so that explanation is unsupported rather than disproven.

4. Reproducibility

5. References

  1. papers/rmssd-equivalence.html — optical rMSSD: unbiased mean, ~22× wider LoA (wall 2.1).
  2. papers/qrs-yield.html — modality-asymmetric beat yield; SQI vs recall (wall 2.2).
  3. papers/nights-icc.html — test-retest reliability; CGM-CV ICC₁≈0 (wall 2.3).
  4. papers/cgm-hrv-coupling.html — shared-driver collapse of the glucose↔HRV coupling (wall 2.4).
  5. papers/odi4-ahi-bias.html — rolling-baseline ODI-4 under-count and its detector-level fix (wall 2.5).
  6. papers/timestamp-pathology.html — the never-fabricate-on-miss rule (wall 2.6's disposition).
  7. PAT Feasibility.html + PAT-FEASIBILITY-2026-07-08-BRIEF.md — cross-device PAT drift, real single-night H10+Verity (wall 2.7); unblock via POLAR-SDK-CAPTURE-2026-07-07-BRIEF.md.
  8. Project documentation: CLAUDE.md, SIGNAL-ADAPTER-AND-FRONTIER-2026-06-23-BRIEF.md §8 (the machine-readable graveyard registry), Tepna suite.
v2.8.0