How to Compare Fall Detection Accuracy
Why headline fall detection accuracy claims mislead, the four numbers that actually compare across platforms, and a two-log pilot that produces real.
Thirty days in one monitored room is 43,200 minutes. If the resident falls once that month, a sensor that never fires is right for 43,199 of them — 99.99 percent accurate — while missing the only minute that mattered. That arithmetic sits under most fall detection accuracy slides, and it is why the headline percentage tells you almost nothing. This guide gives you the four numbers that do compare across platforms, and a two-log pilot that produces them on your own floors.
Why headline fall detection accuracy is not comparable
Start with what the word blends. One vendor counts a fall as an impact above a force threshold. Another counts any transition to the floor lasting longer than a set time — including the resident who lowers herself to reach a dropped remote. A third scored its algorithm entirely on staged drops. Same word, three different experiments. Until you know the event definition, the test population, the denominator and who established ground truth, 99.4 and 96.8 are not two points on one scale.
This protocol is written for nursing homes, assisted living and memory care — settings where many residents cannot report their own falls, so the system's misses stay invisible unless you go looking. The care-facility overview covers the operational side; this page stays on the evidence.
Detection has exactly two ways to fail: miss a real fall, or fire when nothing happened. The dials pull against each other. Turn sensitivity up and you catch the marginal events — and hand the night shift a handset that never stops. Turn it down and the shift goes quiet, including during the fall. A blended score hides both directions at once, because thousands of uneventful minutes swamp a handful of events.
Then there is the test set. A systematic review of fall-detection devices found almost no long-term real-world evaluations; the bulk of published numbers come from scripted falls — volunteers dropping onto padding on cue. Your incident log looks nothing like that. It holds a resident easing down a wardrobe door while gripping the handle, a knee buckling halfway through a transfer, a slide off the bath bench. No impact spike, no clean signature. So the first question for any vendor: how many unscripted falls by older adults are inside your percentage?
What four numbers should you demand from any vendor?
Write these four into the RFP before you sit through a single demo, and hold every platform to the same event definition and denominator.
- 1. True-positive rate on real-world falls. Of the falls that actually happened in deployed buildings, what share raised an alert? The denominator has to come from a source the system doesn't control — an incident log, not the vendor dashboard. A dashboard only knows about the falls it caught.
- 2. False alerts per occupied room-week. Pick the unit your rota understands and make every vendor convert to it. Then ask what an ordinary Tuesday night looks like across 40 rooms: how often the handset goes off, and who walks to check.
- 3. Latency, split into stages. Fall, alert on the handset, acknowledgement, a person at the door — four timestamps, not one. Tinetti's New England Journal of Medicine work found that most older people who fall cannot get up without help, and that lying unhelped for more than an hour sharply worsens outcomes. Every stage of that chain spends from the same clock, and a fast notification means little if acknowledgement is where the time goes.
- 4. A coverage map of your actual building. Where is the system blind — physically and by the hour? Bathrooms are the classic gap; so are the corridor between the lounge and the lift, two-bed rooms the product doesn't support, and every hour a wearable sits on a charger. Have the vendor mark your floor plan room by room. The blank spaces are where falls go unwitnessed.
Total alerts sent, customer logos and satisfaction scores answer none of this. They measure adoption. They cannot surface a missed fall or the workload a false alarm creates on a short-staffed night.
How does the detection technology change the numbers?
Each sensing modality fails in its own characteristic way, and knowing the pattern tells your pilot team where to look. Our overview of fall detection devices walks the whole market; here is the short version for the two families you'll meet in procurement.
Accelerometer wearables — pendants, and watches like the Apple Watch — look for an impact plus an orientation change. That signature fits a stunt fall better than a slow collapse, and the device only works while it's on the body. One study found 97 percent of worn emergency buttons went unpressed during real falls — the person froze, or couldn't reach, or didn't register the fall as an emergency. Add the pendant left on the nightstand during a shower, in the room with the hardest floor, and the coverage math changes fast. Score automatic detection separately from manual SOS presses, and count unworn hours as uncovered hours.
Radar sensing reads reflected radio waves instead of an image, so it works in a dark bathroom and asks nothing of the resident — no wearing, no charging, no pressing. Its limits are spatial: a room without a sensor is a room without coverage, and layout, occupancy and pets shape what the model sees. Put your two-bed rooms and the facility cat in the pilot, not in the assumptions. We've written a plain-language explainer on how radar fall detection works if you want the mechanics before a vendor call.
There is no league table across modalities. You are choosing a failure pattern — missed slow falls and dead batteries versus per-room blind spots — and the right one is whichever your building and night rota absorb best. Then you measure it locally.
How do you run a pilot properly?
A pilot exists to produce the four numbers from your own logs. Be honest about which ones it delivers. The CDC counts one in four adults 65 and over falling each year — yet spread across one unit and a couple of months, that still means only a handful of events in your trial window. A pilot measures alarm burden and latency well and sensitivity poorly; report it that way.
Before the pilot, fix the baseline. Pull a representative stretch of your existing incident log and extract fall count, time, location, witnessed-or-found status, and time-to-discovery where notes let you reconstruct it. Freeze those definitions in writing. The same baseline later anchors the return-on-investment conversation in your records instead of a vendor's model.
During the pilot, run two logs. Normal incident reporting continues exactly as before, independent of the new system — staff record every fall whether or not an alert fired. In parallel, log every alert with its timestamps and an agreed outcome label. Then watch the humans, not just the hardware. A director of nursing we spoke with crossed the logs at the end of a wearable trial and found stretches of night alerts acknowledged within seconds and investigated never — the volume had trained her team to tap the handset quiet. No detection rate survives that.
After the pilot, cross the logs. A fall in the incident log with no matching alert is a miss. An alert with no matching event is a false alarm. Divide by the exposure unit you fixed up front, read acknowledgement and arrival times separately, and review the results with the people who carried the handsets — the night shift first.
What are the red flags in a vendor deck?
The fastest tell is a lone percentage with no event count, denominator or setting attached. The rest of the checklist:
- Lab-only validation. Every figure traces back to volunteers and crash mats, and the deck goes quiet when you ask for deployed-building results.
- No false-alarm number anywhere. A vendor that measured it publishes it. Silence here usually is the number.
- Roadmap instead of evidence. "The next model update" doesn't answer how the version you'd install performs today.
- Latency as a single figure. Detection-to-notification hides who acknowledged the alert and when a human reached the room.
- Coverage gaps in a footnote. A bathroom excluded for privacy is a defensible design choice — drawn on the map, not buried under an asterisk.
- Case studies with no denominator. "214 falls detected" is not a result. Out of how many? Found how — by the system, or by the morning shift?
- Resistance to your independent log. If the vendor dashboard is the only referee, misses are unfindable by design.
- Detection sold as prevention. Fall detection shortens how long someone lies unhelped. It stops nothing. A deck that blurs that line will blur others.
One weak answer doesn't sink a product. Repeated refusal to define the event, denominator and pilot method sinks the comparison — you're left grading a number nobody defined.
Apply the same standard to OdeCare
Here is our own entry, in the format this page demands. OdeCare puts a small radar sensor in each monitored room — no cameras, nothing worn, nothing to charge, nothing a resident has to remember to press.
Three things we don't do. We don't publish a single blended accuracy percentage — for every reason above. We don't claim coverage in rooms without a sensor; the blind spots go on your map before you sign, bathrooms included. And we don't ask you to grade us from our own dashboard: bring your incident log, run the two-log pilot from this page against our alerts, and let the crossing decide.