CAP-FOG-PD: Calibrated, Anticipatory, and Personalized Freezing-of-Gait Detection in Parkinson's Disease: A Real-Time, Edge-Deployable Framework on Open-Access Datasets
Main Article Content
Abstract
Freezing of gait (FOG) is a disabling and unpredictable symptom of PD that increases fall risk and severely impairs quality of life. The majority of available wearable-based systems are reactive, based on population-level data, and tested as if the video-based ground-truth labels are perfect, despite the fact that there are annotator-disagreement flags in many publicly available datasets and it is known that there is a small tolerance of a few seconds for the video annotation in this domain. This work addresses that gap by treating label uncertainty as a calibrated, quantifiable noise term integrated into both model training and evaluation, rather than an unmodeled limitation, which allows genuine model error to be distinguished from irreducible annotation ambiguity, a distinction absent from prior open-dataset FOG research. In this work, we present a framework that anticipates FOG onset from 0.5 to 3 seconds in advance using inertial sensing, validated exclusively on open-access datasets: Daphnet (ankle, thigh, and trunk accelerometers, 64 Hz) and the Michael J. Fox Foundation's tdcsfog (lab, single lower-back sensor, 128 Hz) and defog (home, single lower-back sensor, approximately 100 Hz) datasets. Performance is reported as a complete sensitivity, precision, and false-positive-per-hour curve over four anticipation horizons, decomposed into model-attributable and annotation-noise-attributable error. The differences in sensor placement, channel count, and sampling rate between the Daphnet and tdcsfog/defog datasets allow for the cross-dataset evaluation of the same architecture, a different-input-configuration type of comparison across populations. Stratified sampling using class-weighted loss is used to deal with class imbalance. Two claims are evaluated under a paired, subject-quarantined, leave-one-subject-out scheme: (1) label-uncertainty-calibrated anticipatory detection performance across the four horizons, and (2) improvement from few-shot personalized anticipatory detection on-device over a population-learned baseline. Both models are benchmarked for real-time feasibility by applying 8-bit quantization for inference on sub-100ms latency on embedded ARM Cortex-class hardware. Directions for future work include fully harmonized cross-site generalization, model explainability, and multimodal sensing. All datasets, annotation fields, and thresholds employed are public and independent and can be easily reproduced to report the results.
