Binned Data Provide Better Imputation of Missing Time Series Data from Wearables.

Chakrabarti, Shweta; Biswas, Nupur; Karnani, Khushi; Padul, Vijay; Jones, Lawrence D; Kesari, Santosh; Ashili, Shashaanka

Chakrabarti, Shweta; Biswas, Nupur; Karnani, Khushi; Padul, Vijay; Jones, Lawrence D; Kesari, Santosh; Ashili, Shashaanka.

Afiliação

Chakrabarti S; Rhenix Lifesciences, Hyderabad 500038, India.
Biswas N; Rhenix Lifesciences, Hyderabad 500038, India.
Karnani K; Department of BioSciences and BioEngineering, Indian Institute of Technology, Guwahati 781039, India.
Padul V; Rhenix Lifesciences, Hyderabad 500038, India.
Jones LD; CureScience, 5820 Oberlin Dr, 202, San Diego, CA 92121, USA.
Kesari S; Department of Translational Neurosciences, Pacific Neuroscience Institute and Saint John's Cancer Institute at Providence Saint John's Health Center, Santa Monica, CA 90404, USA.
Ashili S; CureScience, 5820 Oberlin Dr, 202, San Diego, CA 92121, USA.

Sensors (Basel) ; 23(3)2023 Jan 28.

Article em En | MEDLINE | ID: mdl-36772494

ABSTRACT

ABSTRACT

The presence of missing values in a time-series dataset is a very common and well-known problem. Various statistical and machine learning methods have been developed to overcome this problem, with the aim of filling in the missing values in the data. However, the performances of these methods vary widely, showing a high dependence on the type of data and correlations within the data. In our study, we performed some of the well-known imputation methods, such as expectation maximization, k-nearest neighbor, iterative imputer, random forest, and simple imputer, to impute missing data obtained from smart, wearable health trackers. In this manuscript, we proposed the use of data binning for imputation. We showed that the use of data binned around the missing time interval provides a better imputation than the use of a whole dataset. Imputation was performed for 15 min and 1 h of continuous missing data. We used a dataset with different bin sizes, such as 15 min, 30 min, 45 min, and 1 h, and we carried out evaluations using root mean square error (RMSE) values. We observed that the expectation maximization algorithm worked best for the use of binned data. This was followed by the simple imputer, iterative imputer, and k-nearest neighbor, whereas the random forest method had no effect on data binning during imputation. Moreover, the smallest bin sizes of 15 min and 1 h were observed to provide the lowest RMSE values for the majority of the time frames during the imputation of 15 min and 1 h of missing data, respectively. Although applicable to digital health data, we think that this method will also find applicability in other domains.

Assuntos

Algoritmos; Dispositivos Eletrônicos Vestíveis; Fatores de Tempo; Algoritmo Florestas Aleatórias

Palavras-chave

binning; imputation; missing data; time-series data; wearables

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Assunto principal: Algoritmos / Dispositivos Eletrônicos Vestíveis Idioma: En Revista: Sensors (Basel) Ano de publicação: 2023 Tipo de documento: Article

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google