Link to bioRxiv paper: http://biorxiv.org/cgi/content/short/2022.10.26.513936v1?rss=1

Authors: Wedell, E., Shen, C., Warnow, T.

Abstract: Phylogenetic placement, the problem of placing sequences into phylogenetic trees, has been limited either by the number of sequences placed in a single run or by the size of the placement tree. The most accurate scalable phylogenetic placement method with respect to the number of query sequences placed, EPA-ng (Barbera et al., 2019), has a runtime that scales sublinearly to the number of query sequences. However, larger phylogenetic trees cause an increase in the memory usage for EPA-ng, limiting the method to placement trees of up to 10,000 sequences. Our recently designed SCAMPP (Wedell et al., 2021) framework has been shown to scale EPA-ng to larger placement trees of up to 200,000 sequences by building a subtree for the placement of each query sequence. The approach of SCAMPP does not take advantage of the parallel efficiency in EPA-ng since it only places a single query for each run of EPA-ng. Here we present BATCH-SCAMPP, a new technique that overcomes this barrier and enables EPA-ng and other phylogenetic placement methods to scale to ultra-large backbone trees and many query sequences. BATCH-SCAMPP is freely available in GitHub.

Copy rights belong to original authors. Visit the link for more info

Podcast created by Paper Player, LLC