Back to Publications

Peer-reviewed Article

How should gaps be treated in parsimony? A comparison of approaches using simulation

Abstract

Simulation with indels was used to produce alignments where true site homologies in DNA sequences were known; the gaps from these datasets were removed and the sequences were then aligned to produce hypothesized alignments. Both alignments were then analyzed under three widely used methods of treating gaps during tree reconstruction under the maximum parsimony principle. With the true alignments, for many cases (82%), there was no difference in topological accuracy for the different methods of gap coding. However, in cases where a difference was present, coding gaps as a fifth state character or as separate presence/absence characters outperformed treating gaps as unknown/missing data nearly 90% of the time. For the hypothesized alignments, on average, all gap treatment approaches performed equally well. Data sets with higher sequence divergence and more pectinate tree shapes with variable branch lengths are more affected by gap coding than datasets associated with shallower non-pectinate tree shapes.

Full Citation

Ogden, T.H., and M.S. Rosenberg (2007) How should gaps be treated in parsimony? A comparison of approaches using simulation. Molecular Phylogenetics and Evolution 42(3):817–826. [Erratum: 2008, 46(2):807-808.]

DOI

10.1016/j.ympev.2006.07.021

Associated Software

IndelCoder

Associated Data

Direct Download

Download PDF

PubMed Record

PMID: 17011794

Google Scholar Data

Google Scholar Record

104 citations as of 2024-02-19

Altmetrics