Your browser doesn't support javascript.
loading
Data-driven mathematical and visualization approaches for removing rare features for Compositional Data Analysis (CoDA).
Ortiz-Velez, Adrian; Kelley, Scott T.
Afiliación
  • Ortiz-Velez A; Biological and Medical Informatics Program, San Diego State University, San Diego, CA 92182, USA.
  • Kelley ST; Department of Biology, San Diego State University, San Diego, CA 92182, USA.
NAR Genom Bioinform ; 6(1): lqad110, 2024 Mar.
Article en En | MEDLINE | ID: mdl-38187087
ABSTRACT
Sparse feature tables, in which many features are present in very few samples, are common in big biological data (e.g. metagenomics). Ignoring issues of zero-laden datasets can result in biased statistical estimates and decreased power in downstream analyses. Zeros are also a particular issue for compositional data analysis using log-ratios since the log of zero is undefined. Researchers typically deal with this issue by removing low frequency features, but the thresholds for removal differ markedly between studies with little or no justification. Here, we present CurvCut, an unsupervised data-driven approach with human confirmation for rare-feature removal. CurvCut implements two distinct approaches for determining natural breaks in the feature distributions a method based on curvature analysis borrowed from thermodynamics and the Fisher-Jenks statistical method. Our results show that CurvCut rapidly identifies data-specific breaks in these distributions that can be used as cutoff points for low-frequency feature removal that maximizes feature retention. We show that CurvCut works across different biological data types and rapidly generates clear visual results that allow researchers to confirm and apply feature removal cutoffs to individual datasets.

Texto completo: 1 Colección: 01-internacional Base de datos: MEDLINE Idioma: En Revista: NAR Genom Bioinform Año: 2024 Tipo del documento: Article País de afiliación: Estados Unidos

Texto completo: 1 Colección: 01-internacional Base de datos: MEDLINE Idioma: En Revista: NAR Genom Bioinform Año: 2024 Tipo del documento: Article País de afiliación: Estados Unidos