Search PubMed⌕ Search

Biomedical subjects

T P Flores

Publications and source records attributed to T P Flores.

12 recordsLinked to original sources

Protein structural topology: Automated analysis and diagrammatic representation.

The topology of a protein structure is a highly simplified description of its fold including only the sequence of secondary structure elements, and their relative spatial positions and approximate orientations. This information can be embodied in a two-dimensional diagram of protein topology, called a TOPS cartoon. These cartoons are useful for the understanding of particular folds and making comparisons between folds. Here we describe a new algorithm for the production of TOPS cartoons, which is more robust than those previously available, and has a much higher success rate. This algorithm has been used to produce a database of protein topology cartoons that covers most of the data bank of known protein structures.

Algorithms↗

Novel techniques for visualising biological information.

The major challenge facing the bioinformatics community is the continuing increase in the number, size and complexity of biological databases with which it must contend. The goal of the research discussed herein is the development and utilisation of techniques that allow researchers to extract new and useful information from these burgeoning information resources using advanced visualisation methods and paradigms, coupled with distributed object technologies that allow communications between applications and remote databases. Visualisation has roles not only in analysis, but also in building more user-friendly interfaces, implementing methods to navigate large information spaces intuitively and powerful techniques to browse and query data. By using platform-independent object-oriented programming languages, these resources may be developed as reusable pieces of software componentry with their methods and interfaces defined fully, and then distributed through organisations such as the bioWidget Consortium. The widget and object-oriented approach is a powerful paradigm in developing new applications from existing components. Development time is reduced and greater time is spent on analysing these data, rather than in the writing of monolithic applications. More powerful applications can be constructed from components interacting in concert and offers the opportunity of a new generation of bioinformatics tools.

Base Composition↗

The European Bioinformatics Institute (EBI) databases.

The European Bioinformatics Institute (EBI) maintains and distributes the EMBL Nucleotide Sequence database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence database, in collaboration with Amos Bairoch of the University of Geneva. Over fifty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists are available. The EBI network services include database searching and sequence similarity searching facilities.

Amino Acid Sequence↗

Submission of nucleotide sequence data to EMBL/GenBank/DDBJ.

This review outlines the various methods available for submitting sequence data to the EMBL Nucleotide Sequence Database. Depending on the type of sequence data and the facilities available to the submitter, one method may be more suitable than another. Recent developments have been the World Wide Web submission tool, procedures for bulk submissions and genome projects.

Base Sequence↗

Multiple protein structure alignment.

A method was developed to compare protein structures and to combine them into a multiple structure consensus. Previous methods of multiple structure comparison have only concatenated pairwise alignments or produced a consensus structure by averaging coordinate sets. The current method is a fusion of the fast structure comparison program SSAP and the multiple sequence alignment program MULTAL. As in MULTAL, structures are progressively combined, producing intermediate consensus structures that are compared directly to each other and all remaining single structures. This leads to a hierarchic "condensation," continually evaluated in the light of the emerging conserved core regions. Following the SSAP approach, all interatomic vectors were retained with well-conserved regions distinguished by coherent vector bundles (the structural equivalent of a conserved sequence position). Each bundle of vectors is summarized by a resultant, whereas vector coherence is captured in an error term, which is the only distinction between conserved and variable positions. Resultant vectors are used directly in the comparison, which is weighted by their error values, giving greater importance to the matching of conserved positions. The resultant vectors and their errors can also be used directly in molecular modeling. Applications of the method were assessed by the quality of the resulting sequence alignments, phylogenetic tree construction, and databank scanning with the consensus. Visual assessment of the structural superpositions and consensus structure for various well-characterized families confirmed that the consensus had identified a reasonable core.

Amino Acid Sequence↗

An algorithm for automatically generating protein topology cartoons.

An algorithm is described for automatically generating protein topology cartoons. This algorithm optimally places circles and triangles depicting alpha-helices and beta-strands respectively giving a pictorial topological summary of any protein structure. beta-Sheets, sandwiches and barrels are automatically identified and represented using special templates. The output from this algorithm may be controlled by adjustment of variable weights during the optimization step giving a preferred result. The rules for generating protein toplogy cartoons, including consideration of the handedness of local structure motifs, are discussed. The design of this algorithm is completely general and is easily adapted to include further rules that dictate the generation of the cartoons.

Algorithms↗

Comparison of conformational characteristics in structurally similar protein pairs.

Although it is known that three-dimensional structure is well conserved during the evolutionary development of proteins, there have been few studies that consider other parameters apart from divergence of the main-chain coordinates. In this study, we align the structures of 90 pairs of homologous proteins having sequence identities ranging from 5 to 100%. Their structures are compared as a function of sequence identity, including not only consideration of C alpha coordinates but also accessibility, Ooi numbers, secondary structure, and side-chain angles. We discuss how these properties change as the sequences become less similar. This will be of practical use in homology modeling, especially for modeling very distantly related or analogous proteins. We also consider how the average size and number of insertions and deletions vary as sequences diverge. This study presents further quantitative evidence that structure is remarkably well conserved in detail, as well as at the topological level, even when the sequences do not show similarity that is significant statistically.

Algorithms↗

Identification and classification of protein fold families.

We have developed a method for identifying fold families in the protein structure data bank. Pairwise sequence alignments are first performed to extract families of homologous proteins having 35% or more sequence identity. Representatives are selected with the best resolution and R-factor to give a nonhomologous data set. Subsequent structure comparisons between all members of this set detect homologous folds with low sequence identity but highly conserved structures. By softening the requirement on structural similarity, families of analogous proteins are obtained that have related folds but more diverse structures. Representatives are selected to give a non-analogous data set. Starting with 1410 chains from the Brookhaven Data Bank, we generate a set of 150 nonhomologous folds and a set of 112 non-analogous folds. Analysis of sequence and structure conservation within the larger families shows the globins to be the most highly conserved family and the TIM barrels the most weakly conserved.

Classification↗

Prediction of beta-turns in proteins using neural networks.

The use of neural networks to improve empirical secondary structure prediction is explored with regard to the identification of the position and conformational class of beta-turns, a four-residue chain reversal. Recently an algorithm was developed for beta-turn predictions based on the empirical approach of Chou and Fasman using different parameters for three classes (I, II and non-specific) of beta-turns. In this paper, using the same data, an alternative approach to derive an empirical prediction method is used based on neural networks which is a general learning algorithm extensively used in artificial intelligence. Thus the results of the two approaches can be compared. The most severe test of prediction accuracy is the percentage of turn predictions that are correct and the neural network gives an overall improvement from 20.6% to 26.0%. The proportion of correctly predicted residues is 71%, compared to a chance level of about 58%. Thus neural networks provide a method of obtaining more accurate predictions from empirical data than a simpler method of deriving propensities.

Algorithms↗