PeakU · Guide

Garmin vs Oura vs Whoop: which recovery score should you believe?

Body Battery, Readiness and Recovery are built from similar signals and disagree anyway. What each one measures, why they diverge, and what to do when they do.

By the upeak team4 min read2 sources
72 Combined Three devices, three scores, one decision to make

Key takeaways

  • The three scores answer different questions — energy remaining, sleep quality, and readiness for load — so they can all be right while disagreeing.
  • Sleep duration and consistency are reliable across devices; stage breakdowns and any single day's score are much weaker.
  • Trust trends and agreement between devices over any one morning's number, and trust the raw inputs over the score built on them.

If you wear more than one device you have already had the morning where the ring says 84 and the strap says 41. Both are measuring you correctly. The disagreement is not a malfunction, and understanding where it comes from tells you which number to trust for what.

What each score is actually built from

All three start from overnight autonomic signals — heart rate, heart-rate variability, sleep — and then weight them differently according to what the company thinks recovery means.

Garmin Body Battery is a running energy budget across the whole day, draining with stress and activity and charging with rest, derived largely from HRV-based stress. It answers "how much have you got left right now".

Oura Readiness leans hardest on sleep and on overnight physiology, including body temperature deviation, and compares them to your own baselines. It answers "how well did your night go".

Whoop Recovery weights HRV and resting heart rate during slow-wave sleep against your baseline, explicitly in the context of accumulated strain. It answers "how ready are you to take on load".

Those are three different questions. A morning where you slept well after a very hard week can genuinely produce a high Oura score and a low Whoop score, because both are right about the thing they are measuring.

Why the numbers diverge even when the inputs agree

Three reasons, in order of size.

Different weightings and different baselines. Each score is normalised against your history on that device. Swap devices and your own score resets; the number is not portable and was never meant to be.

Different measurement windows. A score built from the last three hours of sleep and one built from the whole night will disagree on a night with a rough start and a calm finish.

Sensor conditions. Optical heart-rate accuracy is not uniform. Bent and colleagues, testing six consumer and medical wearables, found significant differences between devices and between activity types, with absolute error during activity averaging about 30% higher than at rest — and notably no significant difference across skin tones. A finger ring and a wrist strap are not in the same measurement conditions overnight.

Sleep staging is shakier still. Chinoy's laboratory comparison of seven consumer devices against polysomnography found most were decent at detecting sleep versus wake and considerably weaker at staging it. Treat duration and consistency as solid; treat a deep-sleep percentage as an estimate.

Which one to believe

None of them, individually, as an instruction. All of them, collectively, as evidence.

The useful rules: trust trends over single days, because every one of these scores is noisy day to day and informative week to week. Trust agreement — when two independent devices both say you are down, that is a real signal. Trust the inputs over the score: if resting heart rate is up five beats and HRV is well under baseline, that is a fact, whereas the score built on top of it is an opinion about what the fact means.

And when they disagree sharply, ask which question you actually have. Deciding whether to do intervals is a Whoop-shaped question. Deciding whether you are getting ill is an Oura-shaped one.

What PeakU does with this

PeakU does not add an eighth opinion, and it does not ask you to pick a winner. It reads whichever of these you already have through Health Connect, plus the training your wearable never sees — lifts from Hevy, runs from Strava — and computes one Peak Score over the combination.

Two things follow from that. It can see the disagreement and say so, rather than silently picking one. And it shows the contributions: which signal pushed the number up, which pulled it down, and by how much, so you can check its reasoning against what you know about your own week. Your watch's score is an input to the answer, not a rival to it.

Want this answered from your own data?

That is what PeakU does. A coach that reads everything you already track. In private beta — the waitlist is free.

Join the beta About PeakU

Questions

Which recovery score is the most accurate?
There is no single answer, because they are not measuring the same construct. Judge them on the underlying signals — resting heart rate and HRV against your own baseline — which are comparable, rather than on the proprietary score, which is not.
Can I use two recovery devices at once?
Yes, and the disagreement is informative rather than a problem. Agreement between two independent devices is a much stronger signal than either one alone.
Does my score reset if I change watches?
Effectively, yes. Each score is normalised against your history on that device, so a new device needs two to three weeks to rebuild a baseline. A layer that reads all of them keeps a continuous history across the switch.

Sources

  1. Bent B. et al. (2020). Investigating sources of inaccuracy in wearable optical heart rate sensors. npj Digital Medicine. https://www.nature.com/articles/s41746-020-0226-6
  2. Chinoy E.D. et al. (2021). Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep. https://academic.oup.com/sleep/article/44/5/zsaa291/6055610