Efficient data integration under prior probability shift.

Huang, Ming-Yueh; Qin, Jing; Huang, Chiung-Yu

Huang, Ming-Yueh; Qin, Jing; Huang, Chiung-Yu.

Afiliação

Huang MY; Institute of Statistical Science, Academia Sinica, Taipei 11529, Taiwan.
Qin J; Biostatistics Research Branch, National Institute of Allergy and Infectious Diseases, National Institutes of Health,Bethesda, MD 20892, United States.
Huang CY; Department of Epidemiology and Biostatistics, University of California, San Francisco, CA 94158, United States.

Biometrics ; 80(2)2024 Mar 27.

Article em En | MEDLINE | ID: mdl-38768225

ABSTRACT

ABSTRACT

Conventional supervised learning usually operates under the premise that data are collected from the same underlying population. However, challenges may arise when integrating new data from different populations, resulting in a phenomenon known as dataset shift. This paper focuses on prior probability shift, where the distribution of the outcome varies across datasets but the conditional distribution of features given the outcome remains the same. To tackle the challenges posed by such shift, we propose an estimation algorithm that can efficiently combine information from multiple sources. Unlike existing methods that are restricted to discrete outcomes, the proposed approach accommodates both discrete and continuous outcomes. It also handles high-dimensional covariate vectors through variable selection using an adaptive least absolute shrinkage and selection operator penalty, producing efficient estimates that possess the oracle property. Moreover, a novel semiparametric likelihood ratio test is proposed to check the validity of prior probability shift assumptions by embedding the null conditional density function into Neyman's smooth alternatives (Neyman, 1937) and testing study-specific parameters. We demonstrate the effectiveness of our proposed method through extensive simulations and a real data example. The proposed methods serve as a useful addition to the repertoire of tools for dealing dataset shifts.

Assuntos

Algoritmos; Simulação por Computador; Modelos Estatísticos; Probabilidade; Humanos; Funções Verossimilhança; Biometria/métodos; Interpretação Estatística de Dados; Aprendizado de Máquina Supervisionado

Palavras-chave

dataset shift; penalized likelihood; profile likelihood; semiparametric efficiency

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Assunto principal: Algoritmos / Simulação por Computador / Probabilidade / Modelos Estatísticos Limite: Humans Idioma: En Ano de publicação: 2024 Tipo de documento: Article

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google