Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

The Adaptive Evolution Database (TAED).

BACKGROUND: Developing an understanding of the molecular basis for the divergence of species lies at the heart of biology. The Adaptive Evolution Database (TAED) serves as a starting point to link events that occur at the same time in the evolutionary history (tree of life) of species, based upon coding sequence evolution analyzed with the Master Catalog. The Master Catalog is a collection of evolutionary models, including multiple sequence alignments, phylogenetic trees, and reconstructed ancestral sequences, for all independently evolving protein sequence modules encoded by genes in GenBank [1]. RESULTS: We have estimated from these models the ratio of nonsynonymous to synonymous nucleotide substitution (Ka/Ks), for each branch in their respective evolutionary trees of every subtree containing only chordata or only embryophyta proteins. Branches with high Ka/Ks values represent candidate episodes in the history of the family where the protein may have undergone positive selection, a phenomenon in molecular evolution where the mutant form of a gene must have conferred more fitness than the ancestral form. Such episodes are frequently associated with change in function. We have found that an unexpectedly large number of families (between 10 and 20% of those families examined) have at least one branch with a notably high Ka/Ks value (putative adaptive evolution). As a resource for biologists wishing to understand the interaction between protein sequences and the Darwinian processes that shape these sequences, we have collected these into The Adaptive Evolution Database (TAED). CONCLUSIONS: Placed in a phylogenetic perspective, candidate genes that are undergoing evolution at the same time in the same lineage can be viewed together. This framework based upon coding sequence evolution can be readily expanded to include other types of evolution. In its present form, TAED provides a resource for bioinformaticists interested in data mining and for experimental evolutionists seeking candidate examples of adaptive evolution for further experimental study.

Animals↗

The adaptive evolution database (TAED).

BACKGROUND: The Master Catalog is a collection of evolutionary families, including multiple sequence alignments, phylogenetic trees and reconstructed ancestral sequences, for all protein-sequence modules encoded by genes in GenBank. It can therefore support large-scale genomic surveys, of which we present here The Adaptive Evolution Database (TAED). In TAED, potential examples of positive adaptation are identified by high values for the normalized ratio of nonsynonymous to synonymous nucleotide substitution rates (KA/KS values) on branches of an evolutionary tree between nodes representing reconstructed ancestral sequences. RESULTS: Evolutionary trees and reconstructed ancestral sequences were extracted from the Master Catalog for every subtree containing proteins from the Chordata only or the Embryophyta only. Branches with high KA/KS values were identified. These represent candidate episodes in the history of the protein family when the protein may have undergone positive selection, where the mutant form conferred more fitness than the ancestral form. Such episodes are frequently associated with change in function. An unexpectedly large number of families (between 10% and 20% of those families examined) were found to have at least one branch with high KA/KS values above arbitrarily chosen cut-offs (1 and 0.6). Most of these survived a robustness test and were collected into TAED. CONCLUSIONS: TAED is a raw resource for bioinformaticists interested in data mining and for experimental evolutionists seeking candidate examples of adaptive evolution for further experimental study. It can be expanded to include other evolutionary information (for example changes in gene regulation or splicing) placed in a phylogenetic perspective.

Adaptation, Physiological↗

JEvTrace: refinement and variations of the evolutionary trace in JAVA.

BACKGROUND: Details of functional speciation within gene families can be difficult to identify using standard multiple sequence alignment (MSA) methods. The evolutionary trace (ET) was developed as a visualization tool to combine MSA, phylogenetic and structural data for identification of functional sites in proteins. The method has been successful in extracting evolutionary details of functional surfaces in a number of biological systems and modifications of the method are useful in creating hypotheses about the function of previously unannotated genes. We wish to facilitate the graphical interpretation of disparate data types through the creation of flexible software implementations. RESULTS: We have implemented the ET method in a JAVA graphical interface, JEvTrace. Users can analyze and visualize ET input and output with respect to protein phylogeny, sequence and structure. Function discovery with JEvTrace is demonstrated on two proteins with recently determined crystal structures: YlxR from Streptococcus pneumoniae with a predicted RNA-binding function, and a Haemophilus influenzae protein of unknown function, YbaK. To facilitate analysis and storage of results we propose a MSA coloring data structure. The sequence coloring format readily captures evolutionary, biological, functional and structural features of MSAs. CONCLUSIONS: Protein families and phylogeny represent complex data with statistical outliers and special cases. The JEvTrace implementation of the ET method allows detailed mining and graphical visualization of evolutionary sequence relationships.

Amino Acid Sequence↗

Phylogenetics of Theileria species in small ruminants.

Our study is based on the collection of blood and ticks from sheep in Iran and Italy. Polymerase chain reaction (PCR) testing was performed to target the 18S rRNA gene and RLB was performed using previously published probes. In Italy and Iran 78.7% and 76.0% of the sheep were PCR positive, which after sequencing and RLB showed that they were Theileria ovis and Theileria lestoquardi, respectively. Phylogenetic analysis was performed using the Clustal W multiple sequence alignment program and our sequences were compared with more than 50 others already published in the EMBL database. Our T. lestoquardi sequences linked with other T. lestoquardi sequences from Iran, Tanzania, and Sudan and Theileria annulata showed the importance of having species-specific probes between these two species. However, distinctive clades were found between T. lestoquardi ticks and those found in sheep blood. Italian T. ovis seemed to be closer to Theileria spp. from Namibia and Iran than with other T. ovis from Spain, Turkey, Tanzania, and Sudan adding some information to the controversy about this species. However, some confusion was found on the existing database where the location of pathogens, years, and species names was inaccurate and when available sequences were not always appropriately used. This article will discuss our results and some comparisons with other phylogenetic approaches.

Animals↗

Thermostable esterase from a thermoacidophilic archaeon: purification and characterization for enzymatic resolution of a chiral compound.

Homolog to lipolytic enzymes having the consensus sequence Gly-X-Ser-X-Gly, from the Sulfolobus solfataricus P2 genome, were identified by multiple sequence alignments. Among three potential candidate sequences, one (Est3), which displayed higher activity than the other enzymes on the indicate plates, was characterized. The gene (est 3) was expressed in Escherichia coli, and the recombinant protein (Est3) was purified by chromatographic separation. The enzyme is a trimeric protein and has a molecular weight of 32 kDa in monomer form in its native structure. The optimal pH and temperature of the esterase were 7.4 and 80 degrees C respectively. The enzyme showed broad substrate specificities toward various p-nitrophenyl esters ranging from C2 to C16. The catalytic activity of the Est3 esterase was strongly inhibited by phenylmethylsulfonyl fluoride (PMSF) and diethyl p-nitrophenyl phosphate. Based on substrate specificity and the action of inhibitors, the Est3 enzyme was estimated to be a carboxylesterase (EC 3.1.1.1). The enzyme with methyl (+/-)-2-(3-benzoylphenyl)propionate-hydrolyzing activity to (-)-2-(3-benzoylphenyl)propionic acid displayed a moderate degree of enantioselectivity. The product, (-)-2-(3-benzoylphenyl)propionic acid, rather than its methyl ester, was obtained in 80% enantiomeric excess (e.e.(p)) at 20% conversion at 60 degrees C after a 32-h reaction. This result indicates that S. solfataricus esterase can be used for application in the synthesis of chiral compounds.

Amino Acid Sequence↗

Identification of a parathyroid hormone in the fish Fugu rubripes.

UNLABELLED: A PTH gene has been isolated from the fish Fugu rubripes. The encoded protein of 80 amino acid has the lowest homology with any of the PTH family members. Fugu PTH(1-34) had 5-fold lower potency than human PTH(1-34) in a mammalian cell system. INTRODUCTION: Parathyroid hormone (PTH) is the major hypercalcemic hormone in higher vertebrates. Fish lack parathyroid glands, but there have numerous attempts to identify and isolate PTH from fish. MATERIALS AND METHODS: Polymerase chain reaction (PCR) was performed with primers based on preliminary data from the Joint Genome Institute database. PCR amplification was performed on genomic DNA isolated from Fugu rubripes. PCR products were purified and DNA was sequenced. All sequence was confirmed from more than one independently amplified PCR product. Multiple sequence alignments were carried out, and the percentage of identities and similarities were calculated. An unrooted phylogenetic tree, using all the known PTH and PTH-related protein (PTHrP) amino acid sequences, was determined. Synthetic peptides were tested in a biological assay that measured cyclic adenosine 3',5'-monophosphate formation in UMR106.1 cells. Rabbit polyclonal antisera specific for N-terminal human PTHrP and one rabbit polyclonal antiserum specific for N terminus hPTH were used to test the cross-reactivity with fPTH(1-34) in immunoblots.

Amino Acid Sequence↗

Evolutionary trace analysis of eukaryotic DNA topoisomerase I superfamily: identification of novel antitumor drug binding site.

The studies of novel inhibitors of DNA topoisomerase I (Topo I) have already become very promising in cancer chemotherapy. Identifying the new drug-binding residues is playing an important role in the design and optimization of Topo I inhibitors. The designed compounds may have novel scaffolds, thus will be helpful to overcome the toxicities of current camptothecin (CPT) drugs and may provide a solution to cross resistance with these drugs. Multiple sequence alignments were performed on eukaryotic DNA topoisomerase I superfamily and thus the evolutionary tree was constructed. The Evolutionary Trace method was applied to identify functionally important residues of human Topo I. It has been demonstrated that class-specific hydrophobic residues Ala351, Met428, Pro431 are located around the 7,9-position of CPT, indicating suitable substitution of hydrophobic group on CPT will increase antitumor activity. The conservative residue Lys436 in the superfamily is of particular interest and new CPT derivatives designed based on this residue may greatly increase water solubility of such drugs. It has also been demonstrated that the residues Asn352 and Arg364 were conservative in the superfamily, whose mutation will render CPT resistance. As our molecular docking studies demonstrated they did not make any direct interaction with CPT, they are important drug-binding site residues for future design of novel non-camptothecin lead compounds. This work provided a strong basis for the design and synthesis of novel highly potent CPT derivatives and virtual screening for novel lead compounds.

Amino Acid Sequence↗

A guild of 45 CRISPR-associated (Cas) protein families and multiple CRISPR/Cas subtypes exist in prokaryotic genomes.

Clustered regularly interspaced short palindromic repeats (CRISPRs) are a family of DNA direct repeats found in many prokaryotic genomes. Repeats of 21-37 bp typically show weak dyad symmetry and are separated by regularly sized, nonrepetitive spacer sequences. Four CRISPR-associated (Cas) protein families, designated Cas1 to Cas4, are strictly associated with CRISPR elements and always occur near a repeat cluster. Some spacers originate from mobile genetic elements and are thought to confer "immunity" against the elements that harbor these sequences. In the present study, we have systematically investigated uncharacterized proteins encoded in the vicinity of these CRISPRs and found many additional protein families that are strictly associated with CRISPR loci across multiple prokaryotic species. Multiple sequence alignments and hidden Markov models have been built for 45 Cas protein families. These models identify family members with high sensitivity and selectivity and classify key regulators of development, DevR and DevS, in Myxococcus xanthus as Cas proteins. These identifications show that CRISPR/cas gene regions can be quite large, with up to 20 different, tandem-arranged cas genes next to a repeat cluster or filling the region between two repeat clusters. Distinctive subsets of the collection of Cas proteins recur in phylogenetically distant species and correlate with characteristic repeat periodicity. The analyses presented here support initial proposals of mobility of these units, along with the likelihood that loci of different subtypes interact with one another as well as with host cell defensive, replicative, and regulatory systems. It is evident from this analysis that CRISPR/cas loci are larger, more complex, and more heterogeneous than previously appreciated.

Genes, Archaeal↗

Large-scale turnover of functional transcription factor binding sites in Drosophila.

The gain and loss of functional transcription factor binding sites has been proposed as a major source of evolutionary change in cis-regulatory DNA and gene expression. We have developed an evolutionary model to study binding-site turnover that uses multiple sequence alignments to assess the evolutionary constraint on individual binding sites, and to map gain and loss events along a phylogenetic tree. We apply this model to study the evolutionary dynamics of binding sites of the Drosophila melanogaster transcription factor Zeste, using genome-wide in vivo (ChIP-chip) binding data to identify functional Zeste binding sites, and the genome sequences of D. melanogaster, D. simulans, D. erecta, and D. yakuba to study their evolution. We estimate that more than 5% of functional Zeste binding sites in D. melanogaster were gained along the D. melanogaster lineage or lost along one of the other lineages. We find that Zeste-bound regions have a reduced rate of binding-site loss and an increased rate of binding-site gain relative to flanking sequences. Finally, we show that binding-site gains and losses are asymmetrically distributed with respect to D. melanogaster, consistent with lineage-specific acquisition and loss of Zeste-responsive regulatory elements.

Animals↗

Entropy and oligomerization in GPCRs.

Evolutionary trace (ET) and entropy are two related methods for analyzing a multiple sequence alignment to determine functionally important residues in proteins. In this article, these methods have been enhanced with a view to reinvestigate the issue ofGPCR dimerization and oligomerization. In particular, cluster analysis has replaced the subjective visual analysis element of the original ET method. Previous applications of the ET method predicted two dimerization interfaces on the external transmembrane lipid-facing region of GPCRs; these were discussed in terms of dimerization and linear oligomers. Removing the subjective element of the ET method gives rise to the prediction of functionally important residues on the external face of each transmembrane helix for a large number of class A GPCRs. These results are consistent with a growing body of experimental information that, taken over many receptor subtypes, has implicated each transmembrane helix in dimeric interactions. In this application, entropy gave superior results to those obtained from the ET method in that its use gives rise to higher z-scores and fewer instances of z-scores below 3.

Animals↗

Bioinformatics methods to predict protein structure and function. A practical approach.

Protein structure prediction by using bioinformatics can involve sequence similarity searches, multiple sequence alignments, identification and characterization of domains, secondary structure prediction, solvent accessibility prediction, automatic protein fold recognition, constructing three-dimensional models to atomic detail, and model validation. Not all protein structure prediction projects involve the use of all these techniques. A central part of a typical protein structure prediction is the identification of a suitable structural target from which to extrapolate three-dimensional information for a query sequence. The way in which this is done defines three types of projects. The first involves the use of standard and well-understood techniques. If a structural template remains elusive, a second approach using nontrivial methods is required. If a target fold cannot be reliably identified because inconsistent results have been obtained from nontrivial data analyses, the project falls into the third type of project and will be virtually impossible to complete with any degree of reliability. In this article, a set of protocols to predict protein structure from sequence is presented and distinctions among the three types of project are given. These methods, if used appropriately, can provide valuable indicators of protein structure and function.

Algorithms↗

Molecular dynamics simulations of E. coli MsbA transmembrane domain: formation of a semipore structure.

The human P-glycoprotein (MDR1/P-gp) is an ATP-binding cassette (ABC) transporter involved in cellular response to chemical stress and failures of anticancer chemotherapy. In the absence of a high-resolution structure for P-gp, we were interested in the closest P-gp homolog for which a crystal structure is available: the bacterial ABC transporter MsbA. Here we present the molecular dynamics simulations performed on the transmembrane domain of the open-state MsbA in a bilayer composed of palmitoyl oleoyl phosphatidylethanolamine lipids. The system studied contained more than 90,000 atoms and was simulated for 50 ns. This simulation shows that the open-state structure of MsbA can be stable in a membrane environment and provides invaluable insights into the structural relationships between the protein and its surrounding lipids. This study reveals the formation of a semipore-like structure stabilized by two key phospholipids which interact with the hinge region of the protein during the entire simulation. Multiple sequence alignments of ABC transporters reveal that one of the residues involved in the interaction with these two phospholipids are under a strong selection pressure specifically applied on the bacterial homologs of MsbA. Hence, comparison of molecular dynamics simulation and phylogenetic data appears as a powerful approach to investigate the functional relevance of molecular events occurring during simulations.

ATP-Binding Cassette Transporters↗

Clinical and biochemical description of a novel CYP21A2 gene mutation 962_963insA using a new 3D model for the P450c21 protein.

OBJECTIVE: A severely virilized 46, XX newborn girl was referred to our center for evaluation and treatment of congenital adrenal hyperplasia (CAH) because of highly elevated 17alpha-hydroxyprogesterone levels at newborn screening; biochemical tests confirmed the diagnosis of salt-wasting CAH. Genetic analysis revealed that the girl was compound heterozygote for a previously reported Q318X mutation in exon 8 and a novel insertion of an adenine between nucleotides 962 and 963 in exon 4 of the CYP21A2 gene. This 962_963insA mutation created a frameshift leading to a stop codon at amino acid 161 of the P450c21 protein. AIM AND METHODS: To better understand structure-function relationships of mutant P450c21 proteins, we performed multiple sequence alignments of P450c21 with three mammalian P450s (P450 2C8, 2C9 and 2B4) with known structures as well as with human P450c17. Comparative molecular modeling of human P450c21 was then performed by MODELLER using the X-ray crystal structure of rabbit P450 2B4 as a template. RESULTS: The new three dimensional model of human P450c21 and the sequence alignment were found to be helpful in predicting the role of various amino acids in P450c21, especially those involved in heme binding and interaction with P450 oxidoreductase, the obligate electron donor. CONCLUSION: Our model will help in analyzing the genotype-phenotype relationship of P450c21 mutations which have not been tested for their functional activity in an in vitro assay.

Adrenal Hyperplasia, Congenital↗

A Dbf4p BRCA1 C-terminal-like domain required for the response to replication fork arrest in budding yeast.

Dbf4p is an essential regulatory subunit of the Cdc7p kinase required for the initiation of DNA replication. Cdc7p and Dbf4p orthologs have also been shown to function in the response to DNA damage. A previous Dbf4p multiple sequence alignment identified a conserved approximately 40-residue N-terminal region with similarity to the BRCA1 C-terminal (BRCT) motif called "motif N." BRCT motifs encode approximately 100-amino-acid domains involved in the DNA damage response. We have identified an expanded and conserved approximately 100-residue N-terminal region of Dbf4p that includes motif N but is capable of encoding a single BRCT-like domain. Dbf4p orthologs diverge from the BRCT motif at the C terminus but may encode a similar secondary structure in this region. We have therefore called this the BRCT and DBF4 similarity (BRDF) motif. The principal role of this Dbf4p motif was in the response to replication fork (RF) arrest; however, it was not required for cell cycle progression, activation of Cdc7p kinase activity, or interaction with the origin recognition complex (ORC) postulated to recruit Cdc7p-Dbf4p to origins. Rad53p likely directly phosphorylated Dbf4p in response to RF arrest and Dbf4p was required for Rad53p abundance. Rad53p and Dbf4p therefore cooperated to coordinate a robust cellular response to RF arrest.

Amino Acid Motifs↗

An oligonucleotide probe derived from kDNA minirepeats is specific for Leishmania (Viannia).

Sequence analysis of Leishmania (Viannia) kDNA minicircles and analysis of multiple sequence alignments of the conserved region (minirepeats) of five distinct minicircles from L. (V.) braziliensis species with corresponding sequences derived from other dermotropic leishmanias indicated the presence of a sub-genus specific sequence. An oligonucleotide bearing this sequence was designed and used as a molecular probe, being able to recognize solely the sub-genus Viannia species in hybridization experiments. A dendrogram reflecting the homologies among the minirepeat sequences was constructed. Sequence clustering was obtained corresponding to the traditional classification based on similarity of biochemical, biological and parasitological characteristics of these Leishmania species, distinguishing the Old World dermotropic leishmanias, the New World dermotropic leishmanias of the sub-genus Leishmania and of the sub-genus Viannia.

Animals↗

Acidic ribosomal proteins and histone H3 from Leishmania present a high rate of divergence.

Another additional peculiarity in Leishmania will be discussed about of the amino acid divergence rate of three structural proteins: acidic ribosomal P1 and P2b proteins, and histone H3 by using multiple sequence alignment and dendrograms. These structural proteins present a high rate of divergence regarding to their homologous protein in Trypanosoma cruzi. At this regard, L. (V.) peruviana P1 and T. cruzi P1 showed 57.4% of divergence rate. Likewise, L. (V.) braziliensis histone H3 and acidic ribosomal P2 protein exhibited 31.8% and 41.7% respectively of rate of divergence in comparison with their homologous in T. cruzi.

Amino Acids↗

Molecular cloning of the mouse homologue of Rab3c.

Small GTP-binding proteins of the Rab subfamily are key regulators of intracellular vesicle transport. Here we report the isolation of a cDNA clone encoding the complete Rab3c isoform from mouse embryo using a degenerative PCR-based approach. Multiple sequence alignment revealed that the predicted amino acid sequence was identical to the previously identified rat Rab3c isoform and 98% identical to the published bovine Rab3c GTPase from brain. Furthermore by in situ hybridisation, Rab3c mRNA was detectable within various regions of the brain, cartilage and highly enriched within intestinal villi of foetal tissues. Chondrocytes in the hypertrophic zone, but not reserve or proliferative zones, expressed high levels of Rab3c. This pattern of expression corresponds with the genesis of matrix vesicles during endochondral ossification. In all, our results suggest that in addition to its functional role during regulated secretion in brain, Rab3c may play a part in matrix vesicle trafficking during skeletal development.

Amino Acid Sequence↗

GWFASTA: server for FASTA search in eukaryotic and microbial genomes.

Similarity searches are a powerful method for solving important biological problems such as database scanning, evolutionary studies, gene prediction, and protein structure prediction. FASTA is a widely used sequence comparison tool for rapid database scanning. Here we describe the GWFASTA server that was developed to assist the FASTA user in similarity searches against partially and/or completely sequenced genomes. GWFASTA consists of more than 60 microbial genomes, eight eukaryote genomes, and proteomes of annotatedgenomes. Infact, it provides the maximum number of databases for similarity searching from a single platform. GWFASTA allows the submission of more than one sequence as a single query for a FASTA search. It also provides integrated post-processing of FASTA output, including compositional analysis of proteins, multiple sequences alignment, and phylogenetic analysis. Furthermore, it summarizes the search results organism-wise for prokaryotes and chromosome-wise for eukaryotes. Thus, the integration of different tools for sequence analyses makes GWFASTA a powerful toolfor biologists.

Computer Systems↗