Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

The nematode leucine-rich repeat-containing, G protein-coupled receptor (LGR) protein homologous to vertebrate gonadotropin and thyrotropin receptors is constitutively active in mammalian cells.

The receptors for LH, FSH, and TSH belong to the large G protein-coupled, seven-transmembrane protein family and are unique in having a large N-terminal extracellular (ecto-) domain containing leucine-rich repeats important for interactions with the large glycoprotein hormone ligands. Recent studies indicated the evolution of an expanding family of homologous leucine-rich repeat-containing, G protein-coupled receptors (LGRs), including the three known glycoprotein hormone receptors; mammalian LGR4 and LGR5; and LGRs in sea anemone, fly, and snail. We isolated nematode LGR cDNA and characterized its gene from the Caenorhabditis elegans genome. This receptor cDNA encodes 929 amino acids consisting of a signal peptide for membrane insertion, an ectodomain with nine leucine-rich repeats, a seven-TM region, and a long C-terminal tail. The nematode LGR has five potential N-linked glycosylation sites in its ectodomain and multiple consensus phosphorylation sites for protein kinase A and C in the cytoplasmic loop and C tail. The nematode receptor gene has 13 exons; its TM region and C tail, unlike mammalian glycoprotein hormone receptors, are encoded by multiple exons. Sequence alignments showed that the TM region of the nematode receptor has 30% identity and 50% similarity to the same region in mammalian glycoprotein hormone receptors. Although human 293T cells expressing the nematode LGR protein do not respond to human glycoprotein hormones, these cells exhibited major increases in basal cAMP production in the absence of ligand stimulation, reaching levels comparable to those in cells expressing a constitutively activated mutant human LH receptor found in patients with familial male-limited precocious puberty. Analysis of cAMP production mediated by chimeric receptors further indicated that the ectodomain and TM region of the nematode LGR and human LH receptor are interchangeable and the TM region of the nematode LGR is responsible for constitutive receptor activation. Thus, the identification and characterization of the nematode receptor provides the basis for understanding the evolutionary relationship of diverse LGRs and for future analysis of mechanisms underlying the activation of glycoprotein hormone receptors and related LGRs.

Amino Acid Motifs↗

Creating lipoxygenases with new positional specificities by site-directed mutagenesis.

In order to analyse the amino acid determinants which alter the positional specificity of plant lipoxygenases (LOXs), multiple LOX sequence alignments and structural modelling of the enzyme-substrate interactions were carried out. These alignments suggested three amino acid residues as the primary determinants of positional specificity. Here we show the generation of two plant LOXs with new positional specificities, a gamma-linoleneate 6-LOX and an arachidonate 11-LOX, by altering only one of these determinants within the active site of two plant LOXs. In the past, site-directed-mutagenesis studies have mainly been carried out with mammalian lipoxygenases (LOXs) [1]. In these experiments two regions have been identified in the primary structure containing sequence determinants for positional specificity. Amino acids aligning with the Sloane determinants [2] are highly conserved among plant LOXs. In contrast, there is amino acid heterogeneity among plant LOXs at the position that aligns with P353 of the rabbit reticulocyte 15-LOX (Borngräber determinants) [3].

Amino Acid Substitution↗

Comparative analysis of the secondary structural motifs of P450BM-3 and the regions located upstream of the calmodulin-binding domain in the nitric oxide synthases.

The proposed method of multiple alignment of secondary structures makes it possible to estimate the multidomain enzymes' structural similarity in those cases where, following alignment, the identity appears to be inadequate or too insignificant to estimate protein relationship in spite of their functional analogy. Multiple alignment of sequences, representing the secondary structural elements of the nitric oxide synthase (NOS) N-terminal "tails", localized upstream of the calmodulin (CAM)-binding domain, and the P450 superfamily representative P450BM-3 has revealed the existence of a common secondary structural motif, standing of 60% identity between the polypeptide chain packing of NOS and the full sequence of P450BM3. This fact may point to the existence of their common ancestor. Presence of the linker olygopeptides within NOS permitted us to determine the boundary of the cytochrome-like part of NOS.

Amino Acid Sequence↗

Differences between pair-wise and multi-sequence alignment methods affect vertebrate genome comparisons.

Producing complete and accurate alignments of multiple genomic sequences is complex and prone to errors, especially with sequences generated from highly diverged species. In this article, we show that multi-sequence (as opposed to pair-wise) alignment methods are substantially better at aligning (or 'capturing') all of the available orthologous sequence from phylogenetically diverse vertebrates (i.e. those separated by relatively long branch lengths). Maximum gains are obtained only when sequences from many species are aligned. Such multi-sequence alignments contain significant amounts of exonic and highly conserved non-exonic sequences that are not captured in pair-wise alignments, thus illustrating the importance of the alignment method used for performing comparative genome analyses.

Animals↗

Predicting reliable regions in protein alignments from sequence profiles.

For applications such as comparative modelling one major issue is the reliability of sequence alignments. Reliable regions in alignments can be predicted using sub-optimal alignments of the same pair of sequences. Here we show that reliable regions in alignments can also be predicted from multiple sequence profile information alone. Alignments were created for a set of remotely related pairs of proteins using five different test methods. Structural alignments were used to assess the quality of the alignments and the aligned positions were scored using information from the observed frequencies of amino acid residues in sequence profiles pre-generated for each template structure. High-scoring regions of these profile-derived alignment scores were a good predictor of reliably aligned regions. These profile-derived alignment scores are easy to obtain and are applicable to any alignment method. They can be used to detect those regions of alignments that are reliably aligned and to help predict the quality of an alignment. For those residues within secondary structure elements, the regions predicted as reliably aligned agreed with the structural alignments for between 92% and 97.4% of the residues. In loop regions just under 92% of the residues predicted to be reliable agreed with the structural alignments. The percentage of residues predicted as reliable ranged from 32.1% for helix residues to 52.8% for strand residues. This information could also be used to help predict conserved binding sites from sequence alignments. Residues in the template that were identified as binding sites, that aligned to an identical amino acid residue and where the sequence alignment agreed with the structural alignment were in highly conserved, high scoring regions over 80% of the time. This suggests that many binding sites that are present in both target and template sequences are in sequence-conserved regions and that there is the possibility of translating reliability to binding site prediction.

Algorithms↗

A 3D structural model of memapsin 2 protease generated from theoretical study.

AIM: To build a 3D structural model of memapsin 2 (M2) protease for theoretical study and drug design. METHODS: Structural alignment was performed based on multiple and pairwise sequence alignment of three templates. After the initial model was generated, energy minimization was completed by applying molecular mechanics method. Molecular dynamics (MD) technique was used to do further structural optimization. RESULTS: The 3D structural model of memapsin 2 was constructed. The model is reasonable according to several validation criteria. The active-site motifs of M2 are structurally supported by a beta-sheet rich domain and linked together with this domain through alpha helices. Tyr132 contained in beta-hairpin is a general characteristic of aspartic protease. The Calpha atom superimposing result is a direct verification that M2 is structurally unique but still belongs to the aspartic protease superfamily. CONCLUSION: The 3D-structure model from our study is informative to guide future molecular biology study about M2 and drug design based on database searching.

Amino Acid Sequence↗

Generalized affine gap costs for protein sequence alignment.

Based on the observation that a single mutational event can delete or insert multiple residues, affine gap costs for sequence alignment charge a penalty for the existence of a gap, and a further length-dependent penalty. From structural or multiple alignments of distantly related proteins, it has been observed that conserved residues frequently fall into ungapped blocks separated by relatively nonconserved regions. To take advantage of this structure, a simple generalization of affine gap costs is proposed that allows nonconserved regions to be effectively ignored. The distribution of scores from local alignments using these generalized gap costs is shown empirically to follow an extreme value distribution. Examples are presented for which generalized affine gap costs yield superior alignments from the standpoints both of statistical significance and of alignment accuracy. Guidelines for selecting generalized affine gap costs are discussed, as is their possible application to multiple alignment.

Algorithms↗

Comparative modeling in CASP6 using consensus approach to template selection, sequence-structure alignment, and structure assessment.

Along with over 150 other groups we have tested our template-based protein structure prediction approach by submitting models for 30 target proteins to the sixth round of the Critical Assessment of Protein Structure Prediction Methods (CASP6, http://predictioncenter.org). Most of our modeled proteins fall into the comparative or homology modeling (CM) category, and some are fold recognition (FR) targets. The key feature of our structure prediction strategy in CASP6 was an attempt to optimally select structural templates and to make accurate sequence-structure alignments. Template selection was based mainly on consensus results of multiple sequence searches. Likewise, the consensus of multiple alignment variants (or lack of it) was used to initially delineate reliable and unreliable alignment regions. Structure evaluation approaches were then used to identify the correct sequence-structure mapping. Our results suggest that in many cases use of multiple templates is advantageous. Selecting correct alignments even within the context of a three-dimensional structure remains a challenge. Together with more effective energy evaluation methods the simultaneous relaxation/refinement of a "frozen" backbone inherited from the template is likely needed to see a clear progress in tackling this problem. Our analysis also suggests that human input has little to contribute to automatic methods in modeling high homology targets. On the other hand, human expertise can be very valuable in modeling distantly related proteins and critical in cases of unexpected evolutionary changes in protein structure.

Algorithms↗

ClustalW-MPI: ClustalW analysis using distributed and parallel computing.

ClustalW is a tool for aligning multiple protein or nucleotide sequences. The alignment is achieved via three steps: pairwise alignment, guide-tree generation and progressive alignment. ClustalW-MPI is a distributed and parallel implementation of ClustalW. All three steps have been parallelized to reduce the execution time. The software uses a message-passing library called MPI (Message Passing Interface) and runs on distributed workstation clusters as well as on traditional parallel computers.

Amino Acid Sequence↗

Fast algorithms for large-scale genome alignment and comparison.

We describe a suffix-tree algorithm that can align the entire genome sequences of eukaryotic and prokaryotic organisms with minimal use of computer time and memory. The new system, MUMmer 2, runs three times faster while using one-third as much memory as the original MUMmer system. It has been used successfully to align the entire human and mouse genomes to each other, and to align numerous smaller eukaryotic and prokaryotic genomes. A new module permits the alignment of multiple DNA sequence fragments, which has proven valuable in the comparison of incomplete genome sequences. We also describe a method to align more distantly related genomes by detecting protein sequence homology. This extension to MUMmer aligns two genomes after translating the sequence in all six reading frames, extracts all matching protein sequences and then clusters together matches. This method has been applied to both incomplete and complete genome sequences in order to detect regions of conserved synteny, in which multiple proteins from one organism are found in the same order and orientation in another. The system code is being made freely available by the authors.

Algorithms↗

Automated de novo identification of repeat sequence families in sequenced genomes.

Repetitive sequences make up a major part of eukaryotic genomes. We have developed an approach for the de novo identification and classification of repeat sequence families that is based on extensions to the usual approach of single linkage clustering of local pairwise alignments between genomic sequences. Our extensions use multiple alignment information to define the boundaries of individual copies of the repeats and to distinguish homologous but distinct repeat element families. When tested on the human genome, our approach was able to properly identify and group known transposable elements. The program, should be useful for first-pass automatic classification of repeats in newly sequenced genomes.

Algorithms↗

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BACKGROUND: Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. RESULTS: Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. CONCLUSION: MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

Algorithms↗

Combining phylogenetic data with co-regulated genes to identify regulatory motifs.

MOTIVATION: Discovery of regulatory motifs in unaligned DNA sequences remains a fundamental problem in computational biology. Two categories of algorithms have been developed to identify common motifs from a set of DNA sequences. The first can be called a 'multiple genes, single species' approach. It proposes that a degenerate motif is embedded in some or all of the otherwise unrelated input sequences and tries to describe a consensus motif and identify its occurrences. It is often used for co-regulated genes identified through experimental approaches. The second approach can be called 'single gene, multiple species'. It requires orthologous input sequences and tries to identify unusually well conserved regions by phylogenetic footprinting. Both approaches perform well, but each has some limitations. It is tempting to combine the knowledge of co-regulation among different genes and conservation among orthologous genes to improve our ability to identify motifs. RESULTS: Based on the Consensus algorithm previously established by our group, we introduce a new algorithm called PhyloCon (Phylogenetic Consensus) that takes into account both conservation among orthologous genes and co-regulation of genes within a species. This algorithm first aligns conserved regions of orthologous sequences into multiple sequence alignments, or profiles, then compares profiles representing non-orthologous sequences. Motifs emerge as common regions in these profiles. Here we present a novel statistic to compare profiles of DNA sequences and a greedy approach to search for common subprofiles. We demonstrate that PhyloCon performs well on both synthetic and biological data. AVAILABILITY: Software available upon request from the authors. http://ural.wustl.edu/softwares.html

Algorithms↗

Saccharomyces Genome Database (SGD) provides tools to identify and analyze sequences from Saccharomyces cerevisiae and related sequences from other organisms.

The Saccharomyces Genome Database (SGD; http://www.yeastgenome.org/), a scientific database of the molecular biology and genetics of the yeast Saccharomyces cerevisiae, has recently developed several new resources that allow the comparison and integration of information on a genome-wide scale, enabling the user not only to find detailed information about individual genes, but also to make connections across groups of genes with common features and across different species. The Fungal Alignment Viewer displays alignments of sequences from multiple fungal genomes, while the Sequence Similarity Query tool displays PSI-BLAST alignments of each S.cerevisiae protein with similar proteins from any species whose sequences are contained in the non-redundant (nr) protein data set at NCBI. The Yeast Biochemical Pathways tool integrates groups of genes by their common roles in metabolism and displays the metabolic pathways in a graphical form. Finally, the Find Chromosomal Features search interface provides a versatile tool for querying multiple types of information in SGD.

Amino Acid Sequence↗

A space-efficient algorithm for aligning large genomic sequences.

SUMMARY: In the segment-by-segment approach to sequence alignment, pairwise and multiple alignments are generated by comparing gap-free segments of the sequences under study. This method is particularly efficient in detecting local homologies, and it has been used to identify functional regions in large genomic sequences. Herein, an algorithm is outlined that calculates optimal pairwise segment-by-segment alignments in essentially linear space. AVAILABILTIY: The program is available at the Bielefeld Bioinformatics Server (BiBiServ) at http://bibiserv.techfak. uni-bielefeld.de/dialign/

Algorithms↗

Assessment of cry1 gene contents of Bacillus thuringiensis strains by use of DNA microarrays.

A single Bacillus thuringiensis strain can harbor numerous different insecticidal crystal protein (cry) genes from 46 known classes or primary ranks. The cry1 primary rank is the best known and contains the highest number of cry genes which currently totals over 130. We have designed an oligonucleotide-based DNA microarray (cryArray) to test the feasibility of using microarrays to identify the cry gene content of B. thuringiensis strains. Specific 50-mer oligonucleotide probes representing the cry1 primary and tertiary ranks were designed based on multiple cry gene sequence alignments. To minimize false-positive results, a consentaneous approach was adopted in which multiple probes against a specific gene must unanimously produce positive hybridization signals to confirm the presence of a particular gene. In order to validate the cryArray, several well-characterized B. thuringiensis strains including isolates from a Mexican strain collection were tested. With few exceptions, our probes performed in agreement with known or PCR-validated results. In one case, hybridization of primary- but not tertiary-ranked cry1I probes indicated the presence of a novel cry1I gene. Amplification and partial sequencing of the cry1I gene in strains IB360 and IB429 revealed the presence of a cry1Ia gene variant. Since a single microarray hybridization can replace hundreds of individual PCRs, DNA microarrays should become an excellent tool for the fast screening of new B. thuringiensis isolates presenting interesting insecticidal activity.

Bacillus thuringiensis↗