Your browser doesn't support javascript.
loading
Integrative analysis of individual-level data and high-dimensional summary statistics.
Fu, Sheng; Deng, Lu; Zhang, Han; Qin, Jing; Yu, Kai.
Afiliação
  • Fu S; Division of Cancer Epidemiology and Genetics, National Cancer Institute, Bethesda, MD 20892, USA.
  • Deng L; School of Statistics and Data Science, Nankai University, Tianjin 300071, China.
  • Zhang H; Information Management Services, Inc, Bethesda, MD 20892, USA.
  • Qin J; National Institute of Allergy and Infectious Diseases, National Institutes of Health, Bethesda, MD 20892, USA.
  • Yu K; Division of Cancer Epidemiology and Genetics, National Cancer Institute, Bethesda, MD 20892, USA.
Bioinformatics ; 39(4)2023 04 03.
Article em En | MEDLINE | ID: mdl-36964712
ABSTRACT
MOTIVATION Researchers usually conduct statistical analyses based on models built on raw data collected from individual participants (individual-level data). There is a growing interest in enhancing inference efficiency by incorporating aggregated summary information from other sources, such as summary statistics on genetic markers' marginal associations with a given trait generated from genome-wide association studies. However, combining high-dimensional summary data with individual-level data using existing integrative procedures can be challenging due to various numeric issues in optimizing an objective function over a large number of unknown parameters.

RESULTS:

We develop a procedure to improve the fitting of a targeted statistical model by leveraging external summary data for more efficient statistical inference (both effect estimation and hypothesis testing). To make this procedure scalable to high-dimensional summary data, we propose a divide-and-conquer strategy by breaking the task into easier parallel jobs, each fitting the targeted model by integrating the individual-level data with a small proportion of summary data. We obtain the final estimates of model parameters by pooling results from multiple fitted models through the minimum distance estimation procedure. We improve the procedure for a general class of additive models commonly encountered in genetic studies. We further expand these two approaches to integrate individual-level and high-dimensional summary data from different study populations. We demonstrate the advantage of the proposed methods through simulations and an application to the study of the effect on pancreatic cancer risk by the polygenic risk score defined by BMI-associated genetic markers. AVAILABILITY AND IMPLEMENTATION R package is available at https//github.com/fushengstat/MetaGIM.
Assuntos

Texto completo: 1 Base de dados: MEDLINE Assunto principal: Polimorfismo de Nucleotídeo Único / Estudo de Associação Genômica Ampla Tipo de estudo: Prognostic_studies Limite: Humans Idioma: En Revista: Bioinformatics Assunto da revista: INFORMATICA MEDICA Ano de publicação: 2023 Tipo de documento: Article País de afiliação: Estados Unidos

Texto completo: 1 Base de dados: MEDLINE Assunto principal: Polimorfismo de Nucleotídeo Único / Estudo de Associação Genômica Ampla Tipo de estudo: Prognostic_studies Limite: Humans Idioma: En Revista: Bioinformatics Assunto da revista: INFORMATICA MEDICA Ano de publicação: 2023 Tipo de documento: Article País de afiliação: Estados Unidos