Search PubMed⌕ Search

Biomedical subjects

Kuang Lin

Publications and source records attributed to Kuang Lin.

8 recordsLinked to original sources

Scooby-domain: prediction of globular domains in protein sequence.

Scooby-domain (sequence hydrophobicity predicts domains) is a fast and simple method to identify globular domains in protein sequence, based on the observed lengths and hydrophobicities of domains from proteins with known tertiary structure. The prediction method successfully identifies sequence regions that will form a globular structure and those that are likely to be unstructured. The method does not rely on homology searches and, therefore, can identify previously unknown domains for structural elucidation. Scooby-domain is available as a Java applet at http://ibivu.cs.vu.nl/programs/scoobywww. It may be used to visualize local properties within a protein sequence, such as average hydrophobicity, secondary structure propensity and domain boundaries, as well as being a method for fast domain assignment of large sequence sets.

Hydrophobic and Hydrophilic Interactions↗

A simple and fast secondary structure prediction method using hidden neural networks.

MOTIVATION: In this paper, we present a secondary structure prediction method YASPIN that unlike the current state-of-the-art methods utilizes a single neural network for predicting the secondary structure elements in a 7-state local structure scheme and then optimizes the output using a hidden Markov model, which results in providing more information for the prediction. RESULTS: YASPIN was compared with the current top-performing secondary structure prediction methods, such as PHDpsi, PROFsec, SSPro2, JNET and PSIPRED. The overall prediction accuracy on the independent EVA5 sequence set is comparable with that of the top performers, according to the Q3, SOV and Matthew's correlations accuracy measures. YASPIN shows the highest accuracy in terms of Q3 and SOV scores for strand prediction. AVAILABILITY: YASPIN is available on-line at the Centre for Integrative Bioinformatics website (http://ibivu.cs.vu.nl/programs/yaspinwww/) at the Vrije University in Amsterdam and will soon be mirrored on the Mathematical Biology website (http://www.mathbio.nimr.mrc.ac.uk) at the NIMR in London. CONTACT: kxlin@nimr.mrc.ac.uk

Algorithms↗

Contact-based sequence alignment.

This paper introduces the novel method of contact-based protein sequence alignment, where structural information in the form of contact mutation probabilities is incorporated into an alignment routine using contact-mutation matrices (CAO: Contact Accepted mutatiOn). The contact-based alignment routine optimizes the score of matched contacts, which involves four (two per contact) instead of two residues per match in pairwise alignments. The first contact refers to a real side-chain contact in a template sequence with known structure, and the second contact is the equivalent putative contact of a homologous query sequence with unknown structure. An algorithm has been devised to perform a pairwise sequence alignment based on contact information. The contact scores were combined with PAM-type (Point Accepted Mutation) substitution scores after parameterization of gap penalties and score weights by means of a genetic algorithm. We show that owing to the structural information contained in the CAO matrices, significantly improved alignments of distantly related sequences can be obtained. This has allowed us to annotate eight putative Drosophila IGF sequences. Contact-based sequence alignment should therefore prove useful in comparative modelling and fold recognition.

Algorithms↗

A knot or not a knot? SETting the record 'straight' on proteins.

A novel knot found in the SET domain is examined in the light of five recent crystal structures and their descriptions in the literature. Using the algorithm of Taylor it was established that the backbone chain does not form a true knot. However, only two crosslinks corresponding to hydrogen-bonds were needed to form a knotted structure. Such loosely knotted structures formed by hydrogen-bonded crosslinks were assessed as lying between covalent crosslinks (such as disulphide bonds) and threaded-loops which are formed by close (unbonded) contacts between different parts of the chain. The term pseudo-knot was introduced (from the RNA field) to distinguish hydrogen-bonded 'knots'.

Algorithms↗

Testing homology with Contact Accepted mutatiOn (CAO): a contact-based Markov model of protein evolution.

Point Accepted Mutation (PAM) is the Markov model of amino acid replacements in proteins introduced by Dayhoff and her co-workers (Dayhoff et al., 1978). The PAM matrices and other matrices based on the PAM model have been widely accepted as the standard scoring system of protein sequence similarity in protein sequence alignment tools. Here, we present Contact Accepted mutatiOn (CAO), a Markov model of protein residue contact mutations. The CAO model simulates the interchanging of structurally defined side-chain contacts, and introduces additional structural information into protein sequence alignments. Therefore, similarities between structurally conserved sequences can be detected even without apparent sequence similarity. CAO has been benchmarked on the HOMSTRAD database and a subset of the CATH database, by comparing sequence alignments with reference alignments derived from structural superposition. CAO yields scores that reflect coherently the structural quality of sequence alignments, which has implications particularly for homology modelling and threading techniques.

Amino Acid Sequence↗

Amino acid encoding schemes from protein structure alignments: multi-dimensional vectors to describe residue types.

Bioinformatic software has used various numerical encoding schemes to describe amino acid sequences. Orthogonal encoding, employing 20 numbers to describe the amino acid type of one protein residue, is often used with artificial neural network (ANN) models. However, this can increase the model complexity, thus leading to difficulty in implementation and poor performance. Here, we use ANNs to derive encoding schemes for the amino acid types from protein three-dimensional structure alignments. Each of the 20 amino acid types is characterized with a few real numbers. Our schemes are tested on the simulation of amino acid substitution matrices. These simplified schemes outperform the orthogonal encoding on small data sets. Using one of these encoding schemes, we generate a colouring scheme for the amino acids in which comparable amino acids are in similar colours. We expect it to be useful for visual inspection and manual editing of protein multiple sequence alignments.

Algorithms↗

Threading using neural nEtwork (TUNE): the measure of protein sequence-structure compatibility.

MOTIVATION: Fold recognition programs align a probe protein sequence onto protein three-dimensional (3D) structure templates. The alignment between the probe sequence and the most suitable template can be used to predict the 3D structure and often biological function of the probe. Here we present a new threading scoring function of protein sequence-structure compatibility. An artificial neural network model is trained to predict compatibility of amino acid side-chains with structural environments. Log-odds scores of predicted probabilities from this model can then be used to construct protein sequence-structure alignments. RESULTS: Our model is tested on discrimination of native and decoy protein 3D structures. With a residue level structural description, its performance is comparable to those of pseudo-energy functions with atom level structural descriptions, better than the two functions with residue level structural descriptions. AVAILABILITY: The C++ source code of our neural network model is available at http://mathbio.nimr.mrc.ac.uk/~kxlin.

Amino Acid Sequence↗