Protein contact prediction by joint evolutionary coupling analysis across multiple families

نویسندگان

  • Jianzhu Ma
  • Sheng Wang
  • Jinbo Xu
چکیده

Protein contacts contain important information for protein structure and functional study, but contact prediction is very challenging especially for protein families without many sequence homologs. Recently evolutionary coupling (EC) analysis, which predicts contacts by analyzing residue co-evolution in a single target family, has made good progress due to better statistical and optimization techniques. Different from these single-family EC methods, this paper presents a joint multi-family EC analysis method that predicts contacts of one target family by jointly modeling residue co-evolution in itself and also (distantly) related families with divergent sequences but similar folds, and enforcing their co-evolution pattern consistency based upon their evolutionary distance. To implement this multi-family strategy, we use a set of correlated multivariate Gaussian distributions to model related families, the inverse covariance matrix of each distribution encoding the contact pattern of one family. Then we co-estimate the inverse covariance matrices subject to the constraint that they shall share similar patterns. Experiments show that joint multi-family EC analysis can reveal many more native contacts than single-family analysis even for a target family with 4000 sequence homologs, which makes many more protein families amenable to co-evolution-based structure and function prediction.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Protein Contact Prediction by Integrating Joint Evolutionary Coupling Analysis and Supervised Learning

MOTIVATION Protein contact prediction is important for protein structure and functional study. Both evolutionary coupling (EC) analysis and supervised machine learning methods have been developed, making use of different information sources. However, contact prediction is still challenging especially for proteins without a large number of sequence homologs. RESULTS This article presents a gro...

متن کامل

CoinFold: a web server for protein contact prediction and contact-assisted protein folding

CoinFold (http://raptorx2.uchicago.edu/ContactMap/) is a web server for protein contact prediction and contact-assisted de novo structure prediction. CoinFold predicts contacts by integrating joint multi-family evolutionary coupling (EC) analysis and supervised machine learning. This joint EC analysis is unique in that it not only uses residue coevolution information in the target protein famil...

متن کامل

Simultaneous identification of specifically interacting paralogs and interprotein contacts by direct coupling analysis.

Understanding protein-protein interactions is central to our understanding of almost all complex biological processes. Computational tools exploiting rapidly growing genomic databases to characterize protein-protein interactions are urgently needed. Such methods should connect multiple scales from evolutionary conserved interactions between families of homologous proteins, over the identificati...

متن کامل

Large-scale structure prediction by improved contact predictions and model quality assessment

Motivation Accurate contact predictions can be used for predicting the structure of proteins. Until recently these methods were limited to very big protein families, decreasing their utility. However, recent progress by combining direct coupling analysis with machine learning methods has made it possible to predict accurate contact maps for smaller families. To what extent these predictions can...

متن کامل

An Evolutionary Relationship Between Stearoyl-CoA Desaturase (SCD) Protein Sequences Involved in Fatty Acid Metabolism

Background: Stearoyl-CoA desaturase (SCD) is a key enzyme that converts saturated fatty acids (SFAs) to monounsaturated fatty acids (MUFAs) in fat biosynthesis. Despite being crucial for interpreting SCDs’ roles across species, the evolutionary relationship of SCD proteins across species has yet to be elucidated. This study aims to present this evolutionary relationship based on amino aci...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:
  • CoRR

دوره abs/1312.2988  شماره 

صفحات  -

تاریخ انتشار 2013