[en] The fast-growing number of available prokaryotic genomes, along with their uneven taxonomic distribution, is a problem when trying to assemble broadly sampled genome sets for phylogenomics and comparative genomics. Indeed, most of the new genomes belong to the same subset of hyper-sampled phyla, such as Proteobacteria and Firmicutes, or even to single species, such as Escherichia coli (>3000 genomes as of March 2017), while the continuous flow of newly discovered phyla prompts for regular updates of in-house databases. This situation makes it difficult to maintain sets of representative genomes combining lesser known phyla, for which only few species are available, and sound subsets of highly abundant phyla. An automated method is required but none are publicly available. In this work, the kmer composition of DNA sequences, in conjunction with quality metrics for publicly available assemblies, was used to develop an automated approach for selecting a high-quality subset of
representative genomes without redundancy by using our hybrid divide-and-conquer / greedy clustering method.
Disciplines :
Biochemistry, biophysics & molecular biology
Author, co-author :
Léonard, Raphaël ; Université de Liège > Département des sciences de la vie > Cristallographie des macromolécules biologiques
Sirjacobs, Damien ; Université de Liège > Département des sciences de la vie > Phylogénomique des eucaryotes
Sauvage, Eric ; Université de Liège > Département des sciences de la vie > Centre d'ingénierie des protéines
Kerff, Frédéric ; Université de Liège > Département des sciences de la vie > Centre d'ingénierie des protéines
Baurain, Denis ; Université de Liège > Département des sciences de la vie > Phylogénomique des eucaryotes
Language :
English
Title :
ToRQuEMaDA: Tool for retrieving queried eubacteria, metadata and dereplicating assemblie
Publication date :
11 May 2017
Number of pages :
A0
Event name :
PhD Student Day 2017
Event organizer :
ULg - Université de Liège
Event place :
Liège, Belgium
Event date :
11 mai 2017
Funders :
FRIA - Fonds pour la Formation à la Recherche dans l'Industrie et dans l'Agriculture
This website uses cookies to improve user experience. Read more
Save & Close
Accept all
Decline all
Show detailsHide details
Cookie declaration
About cookies
Strictly necessary
Performance
Strictly necessary cookies allow core website functionality such as user login and account management. The website cannot be used properly without strictly necessary cookies.
This cookie is used by Cookie-Script.com service to remember visitor cookie consent preferences. It is necessary for Cookie-Script.com cookie banner to work properly.
Performance cookies are used to see how visitors use the website, eg. analytics cookies. Those cookies cannot be used to directly identify a certain visitor.
Used to store the attribution information, the referrer initially used to visit the website
Cookies are small text files that are placed on your computer by websites that you visit. Websites use cookies to help users navigate efficiently and perform certain functions. Cookies that are required for the website to operate properly are allowed to be set without your permission. All other cookies need to be approved before they can be set in the browser.
You can change your consent to cookie usage at any time on our Privacy Policy page.