학술논문

An expanded evaluation of protein function prediction methods shows an improvement in accuracy

Document Type

Working Paper

Author

Jiang, Yuxiang; Oron, Tal Ronnen; Clark, Wyatt T; Bankapur, Asma R; D'Andrea, Daniel; Lepore, Rosalba; Funk, Christopher S; Kahanda, Indika; Verspoor, Karin M; Ben-Hur, Asa; Koo, Emily; Penfold-Brown, Duncan; Shasha, Dennis; Youngs, Noah; Bonneau, Richard; Lin, Alexandra; Sahraeian, Sayed ME; Martelli, Pier Luigi; Profiti, Giuseppe; Casadio, Rita; Cao, Renzhi; Zhong, Zhaolong; Cheng, Jianlin; Altenhoff, Adrian; Skunca, Nives; Dessimoz, Christophe; Dogan, Tunca; Hakala, Kai; Kaewphan, Suwisa; Mehryary, Farrokh; Salakoski, Tapio; Ginter, Filip; Fang, Hai; Smithers, Ben; Oates, Matt; Gough, Julian; Törönen, Petri; Koskinen, Patrik; Holm, Liisa; Chen, Ching-Tai; Hsu, Wen-Lian; Bryson, Kevin; Cozzetto, Domenico; Minneci, Federico; Jones, David T; Chapman, Samuel; C., Dukka B K.; Khan, Ishita K; Kihara, Daisuke; Ofer, Dan; Rappoport, Nadav; Stern, Amos; Cibrian-Uhalte, Elena; Denny, Paul; Foulger, Rebecca E; Hieta, Reija; Legge, Duncan; Lovering, Ruth C; Magrane, Michele; Melidoni, Anna N; Mutowo-Meullenet, Prudence; Pichler, Klemens; Shypitsyna, Aleksandra; Li, Biao; Zakeri, Pooya; ElShal, Sarah; Tranchevent, Léon-Charles; Das, Sayoni; Dawson, Natalie L; Lee, David; Lees, Jonathan G; Sillitoe, Ian; Bhat, Prajwal; Nepusz, Tamás; Romero, Alfonso E; Sasidharan, Rajkumar; Yang, Haixuan; Paccanaro, Alberto; Gillis, Jesse; Sedeño-Cortés, Adriana E; Pavlidis, Paul; Feng, Shou; Cejuela, Juan M; Goldberg, Tatyana; Hamp, Tobias; Richter, Lothar; Salamov, Asaf; Gabaldon, Toni; Marcet-Houben, Marina; Supek, Fran; Gong, Qingtian; Ning, Wei; Zhou, Yuanpeng; Tian, Weidong; Falda, Marco; Fontana, Paolo; Lavezzo, Enrico; Toppo, Stefano; Ferrari, Carlo; Giollo, Manuel; Piovesan, Damiano; Tosatto, Silvio; del Pozo, Angela; Fernández, José M; Maietta, Paolo; Valencia, Alfonso; Tress, Michael L; Benso, Alfredo; Di Carlo, Stefano; Politano, Gianfranco; Savino, Alessandro; Rehman, Hafeez Ur; Re, Matteo; Mesiti, Marco; Valentini, Giorgio; Bargsten, Joachim W; van Dijk, Aalt DJ; Gemovic, Branislava; Glisic, Sanja; Perovic, Vladmir; Veljkovic, Veljko; Veljkovic, Nevena; Almeida-e-Silva, Danillo C; Vencio, Ricardo ZN; Sharan, Malvika; Vogel, Jörg; Kansakar, Lakesh; Zhang, Shanshan; Vucetic, Slobodan; Wang, Zheng; Sternberg, Michael JE; Wass, Mark N; Huntley, Rachael P; Martin, Maria J; O'Donovan, Claire; Robinson, Peter N; Moreau, Yves; Tramontano, Anna; Babbitt, Patricia C; Brenner, Steven E; Linial, Michal; Orengo, Christine A; Rost, Burkhard; Greene, Casey S; Mooney, Sean D; Friedberg, Iddo; Radivojac, Predrag

Source

Subject

Quantitative Biology - Quantitative Methods

Language

Abstract

Background: The increasing volume and variety of genotypic and phenotypic data is a major defining characteristic of modern biomedical sciences. At the same time, the limitations in technology for generating data and the inherently stochastic nature of biomolecular events have led to the discrepancy between the volume of data and the amount of knowledge gleaned from it. A major bottleneck in our ability to understand the molecular underpinnings of life is the assignment of function to biological macromolecules, especially proteins. While molecular experiments provide the most reliable annotation of proteins, their relatively low throughput and restricted purview have led to an increasing role for computational function prediction. However, accurately assessing methods for protein function prediction and tracking progress in the field remain challenging. Methodology: We have conducted the second Critical Assessment of Functional Annotation (CAFA), a timed challenge to assess computational methods that automatically assign protein function. One hundred twenty-six methods from 56 research groups were evaluated for their ability to predict biological functions using the Gene Ontology and gene-disease associations using the Human Phenotype Ontology on a set of 3,681 proteins from 18 species. CAFA2 featured significantly expanded analysis compared with CAFA1, with regards to data set size, variety, and assessment metrics. To review progress in the field, the analysis also compared the best methods participating in CAFA1 to those of CAFA2. Conclusions: The top performing methods in CAFA2 outperformed the best methods from CAFA1, demonstrating that computational function prediction is improving. This increased accuracy can be attributed to the combined effect of the growing number of experimental annotations and improved methods for function prediction.
Comment: Submitted to Genome Biology

Online Access

Open Access (Arxiv) Find it@PNU

이메일

부산대학교 도서관

Online Access

메일 발송