Search PubMed⌕ Search

Biomedical subjects

Thomas W H Lui

Publications and source records attributed to Thomas W H Lui.

3 recordsLinked to original sources

A multiple-pattern biosequence analysis method for diverse source association mining.

BACKGROUND: In order to understand the intricacy of biomolecules more comprehensively, significant patterns extracted from related data collected from diverse sources must be integrated. These data sources may be local or distributed, possibly with different representation schemes. Often, related data from different sources correspond only with respect to some of their values. METHODS: In biological sequence analysis, a goal is to identify new, previously unknown, relevant patterns, to obtain additional insights into the biomolecule. This is known as a pattern discovery task, rather than a pattern matching task. In this research, we present a method to tackle this problem typically found in molecular sequence analysis when the alignment of the sequences is represented as a relation. In this article, we propose an information measure to select attribute values that reflect multiple patterns of significant interdependence information. Based on these selected values, the patterns are evaluated with data values from other sources. RESULTS: In the experiments, a cancer-suppressor gene known as TP53 (encoding tumour protein p53) is analysed with the mutation records of patients. The experiments identify previously unknown points in the molecule that have patterns negatively associated with the occurrence of cancer. CONCLUSION: Since the evaluated interdependence pattern is a global property of the molecule, we conjecture that the identified points might also be a reflection of the molecule's cancer-suppressor characteristics. The experiments also confirm the usefulness of the proposed method.

Amino Acid Sequence↗

Empirical models for substitution in ribosomal RNA.

Empirical models of substitution are often used in protein sequence analysis because the large alphabet of amino acids requires that many parameters be estimated in all but the simplest parametric models. When information about structure is used in the analysis of substitutions in structured RNA, a similar situation occurs. The number of parameters necessary to adequately describe the substitution process increases in order to model the substitution of paired bases. We have developed a method to obtain substitution rate matrices empirically from RNA alignments that include structural information in the form of base pairs. Our data consisted of alignments from the European Ribosomal RNA Database of Bacterial and Eukaryotic Small Subunit and Large Subunit Ribosomal RNA ( Wuyts et al. 2001. Nucleic Acids Res. 29:175-177; Wuyts et al. 2002. Nucleic Acids Res. 30:183-185). Using secondary structural information, we converted each sequence in the alignments into a sequence over a 20-symbol code: one symbol for each of the four individual bases, and one symbol for each of the 16 ordered pairs. Substitutions in the coded sequences are defined in the natural way, as observed changes between two sequences at any particular site. For given ranges (windows) of sequence divergence, we obtained substitution frequency matrices for the coded sequences. Using a technique originally developed for modeling amino acid substitutions ( Veerassamy, Smith, and Tillier. 2003. J. Comput. Biol. 10:997-1010), we were able to estimate the actual evolutionary distance for each window. The actual evolutionary distances were used to derive instantaneous rate matrices, and from these we selected a universal rate matrix. The universal rate matrices were incorporated into the Phylip Software package ( Felsenstein 2002. http://evolution.genetics.washington.edu/phylip.html), and we analyzed the ribosomal RNA alignments using both distance and maximum likelihood methods. The empirical substitution models performed well on simulated data, and produced reasonable evolutionary trees for 16S ribosomal RNA sequences from sequenced Bacterial genomes. Empirical models have the advantage of being easily implemented, and the fact that the code consists of 20 symbols makes the models easily incorporated into existing programs for protein sequence analysis. In addition, the models are useful for simulating the evolution of RNA sequence and structure simultaneously.

Amino Acid Substitution↗

Using multiple interdependency to separate functional from phylogenetic correlations in protein alignments.

MOTIVATION: Multiple sequence alignments of homologous proteins are useful for inferring their phylogenetic history and to reveal functionally important regions in the proteins. Functional constraints may lead to co-variation of two or more amino acids in the sequence, such that a substitution at one site is accompanied by compensatory substitutions at another site. It is not sufficient to find the statistical correlations between sites in the alignment because these may be the result of several undetermined causes. In particular, phylogenetic clustering will lead to many strong correlations. RESULTS: A procedure is developed to detect statistical correlations stemming from functional interaction by removing the strong phylogenetic signal that leads to the correlations of each site with many others in the sequence. Our method relies upon the accuracy of the alignment but it does not require any assumptions about the phylogeny or the substitution process. The effectiveness of the method was verified using computer simulations and then applied to predict functional interactions between amino acids in the Pfam database of alignments.

Algorithms↗