Impacts of Terraces on Phylogenetic Inference.

Sanderson, Michael J; McMahon, Michelle M; Stamatakis, Alexandros; Zwickl, Derrick J; Steel, Mike

Sanderson, Michael J; McMahon, Michelle M; Stamatakis, Alexandros; Zwickl, Derrick J; Steel, Mike.

Afiliación

Sanderson MJ; Department of Ecology and Evolutionary Biology, University of Arizona, Tucson, AZ 85721, USA; sanderm@email.arizona.edu.
McMahon MM; Department of Ecology and Evolutionary Biology, University of Arizona, Tucson, AZ 85721, USA;
Stamatakis A; Department of Ecology and Evolutionary Biology, University of Arizona, Tucson, AZ 85721, USA; Scientific Computing Group, Heidelberg Institute for Theoretical Studies, Heidelberg 69118, Germany; Institute of Theoretical Informatics, Karlsruhe Institute of Technology, Karlsruhe 76131, Germany;
Zwickl DJ; Department of Ecology and Evolutionary Biology, University of Arizona, Tucson, AZ 85721, USA;
Steel M; Department of Ecology and Evolutionary Biology, University of Arizona, Tucson, AZ 85721, USA;

Syst Biol ; 64(5): 709-26, 2015 Sep.

Article en En | MEDLINE | ID: mdl-25999395

ABSTRACT

ABSTRACT

Terraces are sets of trees with precisely the same likelihood or parsimony score, which can be induced by missing sequences in partitioned multi-locus phylogenetic data matrices. The potentially large set of trees on a terrace can be characterized by enumeration algorithms or consensus methods that exploit the pattern of partial taxon coverage in the data, independent of the sequence data themselves. Terraces can add ambiguity and complexity to phylogenetic inference, particularly in settings where inference is already challenging data sets with many taxa and relatively few loci. In this article we present five new findings about terraces and their impacts on phylogenetic inference. First, we clarify assumptions about partitioning scheme model parameters that are necessary for the existence of terraces. Second, we explore the dependence of terrace size on partitioning scheme and indicate how to find the partitioning scheme associated with the largest terrace containing a given tree. Third, we highlight the impact of terrace size on bootstrap estimates of confidence limits in clades, and characterize the surprising result that the bootstrap proportion for a clade, as it is usually calculated, can be entirely determined by the frequency of bipartitions on a terrace, with some bipartitions receiving high support even when incorrect. Fourth, we dissect some effects of prior distributions of edge lengths on the computed posterior probabilities of clades on terraces, to understand an example in which long edges "attract" each other in Bayesian inference. Fifth, we describe how assuming relationships between edge-lengths of different loci, as an attempt to avoid terraces, can also be problematic when taxon coverage is partial, specifically when heterotachy is present. Finally, we discuss strategies for remediation of some of these problems. One promising approach finds a minimal set of taxa which, when deleted from the data matrix, reduces the size of a terrace to a single tree.

Asunto(s)

Clasificación/métodos; Simulación por Computador/normas; Filogenia; Modelos Genéticos

Palabras clave

Bootstrap; partitioned model; phylogenetics; posterior probability; terrace

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google

Texto completo: 1 Colección: 01-internacional Banco de datos: MEDLINE Asunto principal: Filogenia / Simulación por Computador / Clasificación Tipo de estudio: Prognostic_studies Idioma: En Revista: Syst Biol Asunto de la revista: BIOLOGIA Año: 2015 Tipo del documento: Article

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google