HostSeq: a Canadian whole genome sequencing and clinical data resource.
BMC Genom Data
; 24(1): 26, 2023 05 02.
Article
en En
| MEDLINE
| ID: mdl-37131148
ABSTRACT
HostSeq was launched in April 2020 as a national initiative to integrate whole genome sequencing data from 10,000 Canadians infected with SARS-CoV-2 with clinical information related to their disease experience. The mandate of HostSeq is to support the Canadian and international research communities in their efforts to understand the risk factors for disease and associated health outcomes and support the development of interventions such as vaccines and therapeutics. HostSeq is a collaboration among 13 independent epidemiological studies of SARS-CoV-2 across five provinces in Canada. Aggregated data collected by HostSeq are made available to the public through two data portals a phenotype portal showing summaries of major variables and their distributions, and a variant search portal enabling queries in a genomic region. Individual-level data is available to the global research community for health research through a Data Access Agreement and Data Access Compliance Office approval. Here we provide an overview of the collective project design along with summary level information for HostSeq. We highlight several statistical considerations for researchers using the HostSeq platform regarding data aggregation, sampling mechanism, covariate adjustment, and X chromosome analysis. In addition to serving as a rich data source, the diversity of study designs, sample sizes, and research objectives among the participating studies provides unique opportunities for the research community.
Palabras clave
Texto completo:
1
Base de datos:
MEDLINE
Asunto principal:
SARS-CoV-2
/
COVID-19
Tipo de estudio:
Risk_factors_studies
País/Región como asunto:
America do norte
Idioma:
En
Revista:
BMC Genom Data
Año:
2023
Tipo del documento:
Article