Best devices to reach your goals

Can a headband tell when you are dreaming?

A wrist device guesses your sleep stages from heart rate and movement. A sleep lab reads them from your brain. In between sits a newer category: headbands and masks with electrodes on your forehead, promising lab-style staging at home for a few hundred dollars.

We spent a day reading the published validation studies while choosing a reference recorder for our own development work. This is not a ranking and it names no products. It is what the evidence says about the category, and the questions it taught us to ask.

How a lab scores sleep

In a sleep lab, a technician glues electrodes to the top and sides of your head, beside your eyes and under your chin. The night is cut into 30-second slices, and a trained scorer labels each slice as awake, light sleep (two grades, called N1 and N2), deep sleep (N3) or REM — the stage in which most vivid dreaming happens. That labelled record is called a hypnogram.

The first honest number in this whole topic: two human experts scoring the same night disagree on roughly one slice in six. Agreement is usually reported as a statistic called kappa, where 1.0 is perfect and 0 is chance; expert-versus-expert is commonly reported in the 0.7s. Any device is being compared against a reference that is itself a little blurry. Keep that in mind when a box says “90% accurate”.

What a forehead can see

Home headbands cannot put electrodes on the top of your head; nobody would sleep that way. They read from the forehead, sometimes with a second contact behind the ear. The natural worry is that the forehead is the wrong place. The studies say the picture is more specific than that.

  • Deep sleep travels well. The large, slow waves that define N3 are strongest at the front of the head. Forehead recordings scored with an algorithm built for that placement reach per-stage agreement for deep sleep in the low-to-high 80s (percent, or an equivalent F1 score) in independent studies.
  • REM travels well too, for a different reason. REM is named for the rapid eye movements that come with it, and electrodes on the forehead sit close enough to the eyes to pick those up. Independent studies report REM agreement in the 80s to low 90s — typically the best-detected stage after wake.
  • Light sleep is where it gets vague. The signature of N2, a brief burst called a spindle, is strongest at the centre of the scalp and weaker at the front. N1, the drowsy transition, is barely a stage at all: agreement for N1 is around 40 to 55 percent on every device we read about, and not much better between two humans.

So the two stages people most want to know about — deep and REM — are the two a forehead sensor is best placed to see. The blurriness lives in the light stages, and you should expect any “light sleep” figure to be the softest number on the screen.

The algorithm matters as much as the sensor

This was the finding that changed how we read the spec sheets. The electrodes only record a signal; a scoring model turns it into stages. Most published scoring models were trained on lab recordings from the top of the head, where the waveforms have a different shape and size. Run one of those on forehead data and it can fall apart.

One widely used open model, excellent on lab recordings, dropped to a kappa of under 0.6 when pointed at a forehead device unchanged — and recovered to about 0.7 only after being retrained on nights from that device. The same device, scored by a model built for its own electrode placement, did as well or better from the start. Different device, same lesson: the maker’s own scoring software scored markedly worse than an independent open model trained on the same recordings.

The question to ask is therefore not “is this an EEG headband?” but “was the staging validated on this exact device, against a lab, on people like me?” Sensor and scorer are a pair. Validation of one is not validation of the other.

How to read a validation study

Makers increasingly publish or cite one. A few minutes with the paper tell you most of what you need.

  • Who wrote it. Authors employed by the maker is not disqualifying, but independent authors with the maker blinded to the lab result is stronger. Some of the best studies we found lent the hardware and stepped back.
  • Which hardware. Check the model number. More than once, the study covered a previous generation, and the version on sale today had no published data at all. A new sensor layout is a new device.
  • How many people, and who. Ten healthy 25-year-olds is a start. Fifty adults across ages, including poor sleepers, is evidence. Look for whether anyone with sleep apnoea or insomnia was included if that describes you.
  • How many nights were thrown away. One careful study excluded 16 percent of nights for poor signal before computing accuracy. That is honest reporting, and it is also your likely experience: expect one night in six to ten to be degraded by a shifted contact, hair, or side sleeping.
  • Per-stage numbers, not one headline. A single “accuracy” figure is dominated by the easy stages. Ask for the sensitivity or F1 for deep sleep and REM separately. If the paper does not report them, the headline is telling you very little.
  • Peer-reviewed, preprint, or press release. All three exist in this category. A press release figure with no paper behind it is a claim, not a result. We found more than one accuracy percentage that no journal index could locate.

Beyond accuracy

The questions from choosing a tracking device apply with extra force here, because these devices are more intimate and more expensive.

  • Can you get the recording out? A raw overnight file in a standard format (EDF is the lab convention) means you can re-score it with a better model in two years. A stage chart in an app is a rental. Several devices in this category export nothing at all, and at least one gates the hypnogram itself behind a subscription.
  • Does it work without the cloud? Some record onto the device and never need a phone or an account. Others stream to a phone all night, or upload to the maker’s servers where the scoring happens. Both are common. Ask which, and ask what happens when the servers go away.
  • Platform support can be withdrawn. During our reading, one headband maker ended its Android app outright, leaving Android owners with a band that cannot log in. Check that your phone is supported now and read the maker’s update history.
  • Consumables. Gel electrode patches give a better signal than dry contacts and cost money every few nights. Over a year that can exceed the price of the device. Dry contacts cost nothing but lose contact more easily.
  • Comfort and contact points. A forehead-only band is one thing; a band that also needs contact behind the ears will fight with hair, with side sleeping, and with a sleep mask’s strap. Try it for a week and keep the receipt.

What this means for us

SoluneFit estimates sleep stages from wrist and band signals: heart rhythm and movement, not brain activity. Those are estimates, and the app will always say so and show a confidence level. To make the estimates as good as they can be, we plan to check them against a forehead EEG recorder scored with an open, independently validated model — the same standard this article asks you to hold a product to. If a night’s estimate cannot be trusted, we would rather tell you than draw a confident chart.

The takeaway

  • Two sleep experts disagree on about one 30-second slice in six. No device beats the reference it is scored against.
  • Forehead sensors see deep sleep and REM well. Light sleep, especially N1, is vague on every device.
  • The scoring algorithm matters as much as the electrodes. Ask whether it was validated on this exact model.
  • Read the study: who wrote it, which hardware, how many people, how many nights were discarded, per-stage numbers.
  • Raw export, offline operation, consumable cost and platform support decide whether the device is still useful in a year.
  • Whatever you wear, treat home sleep stages as estimates with a confidence level, not measurements.

Sources

The studies behind the numbers above, so you can read them yourself. Listing a paper is not an endorsement of the device it tested.

  • Lanthier et al. (2026). SLEEP Advances 7(1), zpaf089 — consumer forehead headband vs level-1 polysomnography, n=47.
  • Esfahani et al. (2023). bioRxiv, doi 10.1101/2023.08.18.553744 — two-channel forehead wearable, 135 nights, open scoring pipeline (preprint).
  • Salfi et al. (2026). Journal of Sleep Research, doi 10.1111/jsr.70282 — deep-learning ensembles on forehead recordings, n=10.
  • Arnal et al. (2020). Sleep 43(11), zsaa097 — dry-electrode headband vs polysomnography, n=25.
  • Vallat & Walker (2021). eLife 10, e70092 — an open automated sleep-staging tool, trained on central-scalp recordings.
  • Perslev et al. (2021). npj Digital Medicine 4, 72 — a general-purpose open staging model.
  • Berent et al. (2026). Bioelectronic Medicine, doi 10.1186/s42234-026-00207-x — in-ear EEG vs polysomnography, n=16.
  • Rosenberg & Van Hout (2013). Journal of Clinical Sleep Medicine 9(1), 81–87 — inter-scorer agreement among human experts.

← All articles

This article is general wellness education, not medical advice. SoluneFit is a wellness tool, not a medical device. It does not diagnose or treat any condition. Talk to a clinician about health concerns.