Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference


Accurate phylogenies are critical to taxonomy as well as studies of speciation processes and other evolutionary patterns. Accurate branch lengths in phylogenies are critical for dating and rate measurements. Such accuracy may be jeopardized by unacknowledged sequencing error. We use simulated data to test a correction for DNA sequencing error in maximum likelihood phylogeny inference. Over a wide range of data polymorphism and true error rate, we found that correcting for sequencing error improves recovery of the branch lengths, even if the assumed error rate is up to twice the true error rate. Low error rates have little effect on recovery of the topology. When error is high, correction improves topological inference; however, when error is extremely high, using an assumed error rate greater than the true error rate leads to poor recovery of both topology and branch lengths. The error correction approach tested here was proposed in 2004 but has not been widely used, perhaps because researchers do not want to commit to an estimate of the error rate. This study shows that correction with an approximate error rate is generally preferable to ignoring the issue.

DOI: 10.1534/g3.114.014365

Extracted Key Phrases

6 Figures and Tables

Citations per Year

Citation Velocity: 16

Averaging 16 citations per year over the last 2 years.

Learn more about how we calculate this metric in our FAQ.

Cite this paper

@inproceedings{Kuhner2014CorrectingFS, title={Correcting for Sequencing Error in Maximum Likelihood Phylogeny Inference}, author={Mary K. Kuhner and James R. McGill}, booktitle={G3}, year={2014} }