학술논문

A multi-source domain annotation pipeline for quantitative metagenomic and metatranscriptomic functional profiling
Document Type
Report
Source
Microbiome. August 28, 2018, Vol. 6 Issue 1
Subject
Research
Marine ecosystems -- Research
Microbial colonies -- Research
Genomes -- Research
Language
English
ISSN
2049-2618
Abstract
Author(s): Ari Ugarte[sup.1] , Riccardo Vicedomini[sup.1,2] , Juliana Bernardes[sup.1] and Alessandra Carbone[sup.1,3] Background Ecosystem changes are often correlated with the presence of new communities disturbing their stability by importing new [...]
Background Biochemical and regulatory pathways have until recently been thought and modelled within one cell type, one organism and one species. This vision is being dramatically changed by the advent of whole microbiome sequencing studies, revealing the role of symbiotic microbial populations in fundamental biochemical functions. The new landscape we face requires the reconstruction of biochemical and regulatory pathways at the community level in a given environment. In order to understand how environmental factors affect the genetic material and the dynamics of the expression from one environment to another, we want to evaluate the quantity of gene protein sequences or transcripts associated to a given pathway by precisely estimating the abundance of protein domains, their weak presence or absence in environmental samples. Results MetaCLADE is a novel profile-based domain annotation pipeline based on a multi-source domain annotation strategy. It applies directly to reads and improves identification of the catalog of functions in microbiomes. MetaCLADE is applied to simulated data and to more than ten metagenomic and metatranscriptomic datasets from different environments where it outperforms InterProScan in the number of annotated domains. It is compared to the state-of-the-art non-profile-based and profile-based methods, UProC and HMM-GRASPx, showing complementary predictions to UProC. A combination of MetaCLADE and UProC improves even further the functional annotation of environmental samples. Conclusions Learning about the functional activity of environmental microbial communities is a crucial step to understand microbial interactions and large-scale environmental impact. MetaCLADE has been explicitly designed for metagenomic and metatranscriptomic data and allows for the discovery of patterns in divergent sequences, thanks to its multi-source strategy. MetaCLADE highly improves current domain annotation methods and reaches a fine degree of accuracy in annotation of very different environments such as soil and marine ecosystems, ancient metagenomes and human tissues. Keywords: Domain annotation, Metagenomic, Metatranscriptomic, Functional annotation, Probabilistic model, Environment, Motif