Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

Characterization of the nuclear gene encoding mitochondrial aconitase in the marine red alga Gracilaria verrucosa.

We have cloned a nuclear gene from the marine red alga Gracilaria verrucosa that encodes the complete 779 amino-acid mitochondrial aconitase (m-ACN), the first characterized from a photosynthetic organism. The N-terminal 28 deduced amino acids are predicted to constitute the mitochondrial transit peptide, the first described from a red alga. Putative transcriptional cis-acting elements were identified in the upstream untranslated region. The G. verrucosa m-ACN gene (m-ACN) is present in a single copy and is located ca. 1.5 kb upstream from the single-copy polyubiquitin gene. The single spliceosomal intron is located near the 5' end of the region encoding the mature m-ACN in precisely the same location and phase as intron 2 in Caenorhabditis elegans m-ACN; sequences at its 3' and 5' splice junctions and at the predicted lariat branch point conform well to the eukaryote consensus sequences. Multiple protein-sequence alignment of m-ACN, bacterial aconitase (b-ACN) and iron-responsive element-binding protein (IRE-BP), and phylogenetic analyses, revealed that m-ACN does not share a recent common ancestry with either b-ACN or IRE-BP.

Aconitate Hydratase↗

Evolutionary motif and its biological and structural significance.

We developed a method for multiple alignment of protein sequences. The main feature of this method is that it takes the evolutionary relationships of the proteins in question into account repeatedly for execution, until the relationships and alignment results are in agreement. We then applied this method to the data of the international DNA sequence databases, which are the most comprehensive and updated DNA databases in the world, in order to estimate the "evolutionary motif" by extensive use of a supercomputer. Though a few problems needed to be solved, we could estimate the length of the motifs in the range of 20 to 200 amino acids, with about 60 the most frequent length. We then discussed their biological and structural significance. We believe that we are now in a position to analyze DNA and protein not only in vivo and in vitro but also in silico.

Amino Acid Sequence↗

Identification and characteristics of a novel testis-specific gene, Tsc21, in mice and human.

Testis-specific genes are essential for spermatogenesis in mammalian male reproduction. We have identified a novel gene, Tsc21, exclusively expressed in mice and human testes from the results of the Affymetrix Genechip analysis in the six developmental stages of testis of postnatal Balb/C mice. The full cDNA length of Tsc21 was 810 bp, with a 543 bp open reading frame encoding a 180 amino acids protein with a predicted molecular weight of 21.040 kDa. A Blast search in the mouse genome database localized the Tsc21 gene to mice chromosome 6C3. Multiple amino acid sequence alignment of human, mouse, and rat homologous genes showed that mice Tsc21 protein was highly homologous with the human Tsc21 gene (70%) and rat Tsc21 gene (86%). The results of reverse transcriptase-polymerase chain reaction analysis showed that the mice Tsc21 is exclusively expressed in the testis and epididymis of mice, and its expression is only detected after the mice is 35 days old. Human Tsc21 is also exclusively expressed in testis of human. Considering the expression profile Tsc21 in mice and human, we propose that Tsc21 may play a role during mammalian male spermatogenesis. Our study should be a basis for function characterization of the Tsc21 gene, leading to the elucidation of the molecular events underlying mammalian male reproduction.

Amino Acid Sequence↗

Discovery of novel conserved peptide domains by ortholog comparison within plant multi-protein families.

Assigning individual functions to the proteins encoded by the genome of the dicotyledonous reference species Arabidopsis thaliana is one of the major challenges in current plant molecular biology. Frequently, Arabidopsis protein families are biocomputationally analyzed by multiple amino acid sequence alignments of the respective family members for detection of conserved peptide motifs that might be of functional relevance. Mere sequence alignment of paralogous sequences may obscure amino acid patches that are highly conserved amongst orthologs and thus potentially relevant for isoform-specific protein function(s). Here I exemplarily illustrate this potential pitfall by amino acid sequence alignments of the heptahelical MLO proteins using either the suite of 15 isoforms (paralogs) encoded by the Arabidopsis genome or a collection of 13 ortholog sequences derived from a set of both monocotyledonous and dicotyledonous plant species. The findings are corroborated by an analogous analysis of the distinct plant multi-protein family of CONSTANS-like transcription regulators. The data reveal that the generally higher sequence similarity of orthologs versus paralogs is not uniformly distributed among the amino acid positions of the orthologs but at least partially clustered in distinct sites/domains, suggesting conservation of isoform-specific functional modules across taxa.

Amino Acid Sequence↗

Expansion of the mammalian 3 beta-hydroxysteroid dehydrogenase/plant dihydroflavonol reductase superfamily to include a bacterial cholesterol dehydrogenase, a bacterial UDP-galactose-4-epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus.

Mammalian 3 beta-hydroxysteroid dehydrogenase and plant dihydroflavonol reductases are descended from a common ancestor. Here we present evidence that Nocardia cholesterol dehydrogenase, E. coli UDP-galactose-4 epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus are homologous to 3 beta-hydroxysteroid dehydrogenase and dihydroflavonol reductase. Analysis of a multiple alignment of these sequences indicates that viral ORFs are most closely related to the mammalian 3 beta-hydroxysteroid dehydrogenases. The ancestral protein of this superfamily is likely to be one that metabolized sugar nucleotides. The sequence similarity between 3 beta-hydroxysteroid dehydrogenase and the viral ORFs is sufficient to suggest that these ORFs have an activity that is similar to 3 beta-hydroxysteroid dehydrogenase or cholesterol dehydrogenase, although the putative substrates are not yet known.

3-Hydroxysteroid Dehydrogenases↗

Phylogenetic comparison of the serotype-specific VP2 protein of bluetongue and related orbiviruses.

Regions of the VP2 gene from various bluetongue virus serotypes were sequenced and phylogenetic comparisons were performed. The sequences were characteristic for each BTV serotype and isolates of the same serotype could be grouped geographically, mimicking the topotyping characteristics of BTV VP3 gene sequences. PCR amplification and sequence analysis were used to show the close relationship between Caribbean BTV isolates and South African BTV isolates of the same serotype. Similarly, Australian BTV isolates showed a close genetic relationship with Asian BTV isolates of the same serotype. A multiple amino acid sequence alignment of fifteen BTV serotypes and other orbiviruses over a proposed major neutralization site showed this region (317 335 aa.) was highly variable and nucleotide sequences showed that BTV serotypes could be grouped into nucleotypes, or related serotypes, in broad agreement with the inter-relationships postulated by Erasmus (1990), using plaque-reduction tests.

Amino Acid Sequence↗

Computer-assisted dissection of rolling circle DNA replication.

A comparative analysis of the proteins involved in initiation and termination of rolling circle replication (RCR) was performed using computer-assisted methods of data based screening, motif search and multiple amino acid sequence alignment. Two vast classes of such proteins were delineated, one of these being associated with RCR proper, and the other with mobilization (conjugal transfer) of plasmid DNA. The common denominator of the two classes was found to be a conserved amino acid motif that consists of the sequence HisUHisUUU (U--bulky hydrophobic residue; hereafter HUH motif). Based on analogies with metalloenzymes, it is hypothesized that the two conserved His residues this motif may be involved in metal ion coordination required for the activity of the RCR and mobilization proteins. The proteins of the replication (Rep) class contained two additional conserved motifs, with the motif around the Tyr residue(s) forming the covalent link with nicked DNA being located C-proximally of the HUH motif. This class further split into two large superfamilies and several smaller families, with the proteins belonging to a single but not to different (super)families demonstrating statistically significant similarity to each other. Superfamily I, prototyped by the gene A proteins of small isometric single-stranded (ss) DNA bacteriophages, included also Rep proteins of P2-related double-stranded (ds) DNA bacteriophages, the small phage-plasmid hybrid phasyl, and several cyanobacterial and archaebacterial plasmids. These proteins contained two invariant Tyr residues separated by three partially conserved amino acids, suggesting that they all may share the cleavage-ligation mechanism proposed for phi X174 A protein and involving alternate covalent binding of both tyrosines to DNA (Van Mansfeld, A.D., Van Teeffelen, H.A., Baas, P.D., Jansz, H.S., 1986. Nucl. Acids Res. 14, 4229-4238). Superfamily II included Rep proteins of a number of ssDNA plasmids replicating mainly in gram-positive bacteria that unexpectedly were shown to be related to the Rep proteins of plant geminiviruses. Conservation of the "HUH" motif and a motif around the putative DNA-linking Tyr residue was observed also in the Rep proteins of animal parvoviruses containing linear ssDNA with a terminal hairpin and replicating via the rolling hairpin mechanism. The class of plasmid mobilization (Mob) proteins was characterized by the opposite orientation of the conserved motifs, with the (putative) DNA-linking Tyr being located N-proximally of the "HUH" motif.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

A new D-2-hydroxyacid dehydrogenase with dual coenzyme-specificity from Haloferax mediterranei, sequence analysis and heterologous overexpression.

A gene encoding a new D-2-hydroxyacid dehydrogenase (E.C. 1.1.1.) from the halophilic Archaeon Haloferax mediterranei has been sequenced, cloned and expressed in Escherichia coli cells with the inducible expression plasmid pET3a. The nucleotide sequence analysis showed an open reading frame of 927 bp which encodes a 308 amino acid protein. Multiple amino acid sequence alignments of the D-2-hydroxyacid dehydrogenase from H. mediterranei showed high homology with D-2-hydroxyacid dehydrogenases from different organisms and other enzymes of this family. Analysis of the amino acid sequence showed catalytic residues conserved in hydroxyacid dehydrogenases with d-stereospecificity. In the reductive reaction, the enzyme showed broad substrate specificity, although alpha-ketoisoleucine was the most favourable of all alpha-ketocarboxylic acids tested. Kinetic data revealed that this new D-2-hydroxyacid dehydrogenase from H. mediterranei exhibits dual coenzyme-specificity, using both NADPH and NADH as coenzymes. To date, all D-2-hydroxyacid dehydrogenases have been found to be NADH-dependent. Here, we report the first example of a D-2-hydroxyacid dehydrogenase with dual coenzyme-specificity.

Alcohol Oxidoreductases↗

Characterization of the cheY genes from Leptospira interrogans and their effects on the behavior of Escherichia coli.

The motility and chemotaxis system are critical for the virulence of pathogenic leptospire, which enable them to penetrate host tissue barriers during infection. The completed genome sequence of a representative virulent serovar type strain (Lai) of Leptospira interrogans serogroups Icterohaemorrhagiae (L. interrogans strain Lai) suggested that there were multiple copies of putative chemotaxis homologues located at its large chromosome. In order to verify the function of these proteins, the putative cheY genes were cloned into pQE31 vector and then expressed, respectively, in wild-type Escherichia coli strain RP437 and cheY defective strain RP5232. The results showed that all the five cheYs could restore the swarming of RP5232 strain to some extend. Overexpression of CheYs in RP437 showed inhibited swarming of RP437. To investigate the mechanism of chemotaxis signaling in L. interrogans strain Lai, certain aspartates (Asp-53, Asp-61, Asp-70, Asp-62, and Asp-66 for L. interrogans strain Lai CheY1, CheY2, CheY3, CheY4, and CheY5, respectively) were mutated. Expression of these mutated cheYs manifested neither restoration of the swarming ability of RP5232 nor inhibition on swarming ability of RP437. Multiple amino acid sequence alignment predicted ternary structures and the result of mutation experiment suggested that these conserved aspartate residues of L. interrogans were analogous to that in E. coli CheY in function and structure. So, L. interrogans and E. coli may have similar mechanisms of activation of the chemotaxis phosphorelay pathway, but there are differences in their control by signal terminator.

Amino Acid Sequence↗

Comparative analysis of base biases around the stop codons in six eukaryotes.

Using full-length cDNA sequences, a comparative analysis of sequence patterns around the stop codons in six eukaryotes was performed. Here, it was showed that the codon immediately before and after the stop codons (defined as -1 codon and +1 codon, respectively) were much more biased than other examined positions, especially at the second position of -1 codons and the first position of +1 codons which were rich in As/Us and purines, respectively, for most species. The author speculated that strongly biased sequence pattern from position -2 to +4 might act as an extended translation termination signal. Translation termination was catalyzed by release factors that recognized the stop codons. The multiple amino acid sequence alignment of eukaryotic release factor 1 (eRF1) of 20 species showed that there were 16 residue sites that were strictly conserved, especially the invariant amino acids Ile70 and Lys71. Accordingly, it could be inferred that those candidate amino acids might involve in the recognition process. Moreover, the possible stop signal recognition hypothesis was also discussed herein.

Amino Acid Sequence↗

Molecular cloning and functional analysis of cytochrome P450 1A2 from Japanese monkey liver: comparison with marmoset cytochrome P450 1A2.

A cDNA encoding a novel cytochrome P450 1A2 (CYP1A2) was cloned from the liver of an adult female Japanese monkey. The CYP1A2 protein was expressed in yeast cells and its enzymatic properties were compared with those of marmoset CYP1A2 using ethoxyresorufin (ER) and phenacetin (PN) as substrates. The nucleotide sequence of Japanese monkey CYP1A2 revealed 94.7, 99.5 and 93.5% identities to those of human, cynomolgus monkey and marmoset monkey CYP1A2, respectively. Multiple amino acid sequence alignment of Japanese monkey CYP1A2 with CYP1A2 of humans, cynomolgus monkeys and marmosets showed that Japanese monkey CYP1A2 had 92.4, 99.0 and 91.9% identities to the human, cynomolgus monkey and marmoset enzymes, respectively. Kinetic studies demonstrated that the enzymatic properties as ER and PN O-deethylases were considerably different between the Japanese monkey and the marmoset CYP1A2. Furthermore, both of these reactions in liver microsomal fractions from the Japanese monkey and marmoset showed biphasic kinetics. On the basis of the kinetic parameters, it is suggested that Japanese monkey CYP1A2 is a high-K(m) enzyme in both ER and PN O-deethylations, whereas marmoset CYP1A2 is a high-K(m) and low-K(m) enzyme in ER and PN O-deethylations, respectively. alpha-Naphthoflavone, an inhibitor of human CYP1A1 and CYP1A2, did not completely inhibit the liver microsomal oxidations of ER and PN even at the highest concentration (50muM), supporting the notion that CYP1A2 enzymes are not the sole ER or PN O-deethylase in Japanese monkey and marmoset liver microsomes. Inhibitory effects of furafylline, an inhibitor of human CYP1A2, on ER O-deethylation by recombinant CYP1A2 enzymes were much lower than those of alpha-naphthoflavone, but marmoset CYP1A2 was more sensitive to furafylline than Japanese monkey CYP1A2. These results indicate that the properties of Japanese monkey CYP1A2 are considerably different from those of marmoset CYP1A2.

Adult↗

Restricted variable residues in the C-terminal segment of HIV-1 V3 loop regulate the molecular anatomy of CCR5 utilization.

The V3 loop of the HIV-1 envelope glycoprotein (Env) is the major determinant for coreceptor utilization, but the structural basis for this specificity remains to be defined. By characterizing a set of naturally occurring R5 Env variants, we demonstrate that Asp324 in the conserved IIGDIR motif of the V3 loop (CTRPN(300)NNTRKSIHIGP(311)GRAFYTTGEIIGD(324)IRQAHC) C-terminal segment regulates the molecular anatomy of CCR5 utilization. Whereas gp120 subunits with Asp or Asn at position 324 were fusogenic with coreceptor chimeras containing either the N-terminal domain or the body of CCR5, substitution of charged (Glu, Lys) or small hydrophobic (Gly, Ala) residues resulted in complete loss of fusogenic activity with the N terminus and markedly reduced utilization of the body of CCR5, although their ability to use wild-type CCR5 was unchanged. This phenotypic conversion was confirmed in both gain and loss of function experiments using Env from multiple subtypes. Alignment of sequences of R5 V3 loops (n=599) from the HIV database revealed that the mutation of Asp324 in the conserved IIGDIR motif is restricted to Asn324, with proportions of 71.5% and 28%, respectively. Infection of primary CD4(+)T cells demonstrated that Env bearing Asp324 was less sensitive to RANTES, suggesting that Asp or Asn in this position may be crucial for viral fitness. The CD4-dependent gp120 binding to CCR5 was decreased when Asp324 was replaced with a charged or hydrophobic residue, but unchanged when replaced with Asn. Molecular modeling analyses predicted that Asp/Asn324 forms a critical H-bond with Asn300. These findings indicate that Asp or Asn at position 324 of the V3 stem stabilizes the conformation of V3 loop and hence influences the intensities of interaction between CD4-activated gp120 and CCR5 which results in viral entry.

Amino Acid Sequence↗

A conserved insertion in protein-primed DNA polymerases is involved in primer terminus stabilisation.

Protein-primed DNA polymerases form a subgroup of the eukaryotic-type DNA polymerases family, also called family B or alpha-like. A multiple amino acid sequence alignment of this subgroup of DNA polymerases led to the identification of two insertions, TPR-1 and TPR-2, in the polymerisation domain. We showed previously that Asp332 of the TPR-1 insertion of phi29 DNA polymerase is involved in the correct orientation of the terminal protein (TP) for the initiation of replication. In this work, the functional role of two other conserved residues from TPR-1, Lys305 and Tyr315, has been analysed. The four mutant derivatives constructed, K305I, K305R, Y315A and Y315F, displayed a wild-type 3'-5' exonuclease activity on single-stranded DNA. However, when assayed on double-stranded DNA such activity was higher than that of the wild-type enzyme. This activity led to a reduced pol/exo ratio, suggesting a defect in stabilising the primer terminus at the polymerase active site. On the other hand, although mutant polymerases K305I and Y315A were able to couple processive DNA polymerisation to strand displacement, they were severely impaired in phi29 TP-DNA replication. The possible role of the TPR-1 insertion in the set of interactions with the nascent chain during the first steps of TP-DNA replication is discussed.

Amino Acid Sequence↗

Role of transmembrane domains in the functions of B- and T-cell receptors.

The antigen receptors on the surface of B- and T-lymphocytes are complexes of several integral membrane proteins, essential for their proper expression and function. Recent studies demonstrated that transmembrane (TM) domains of the components of these receptors play a critical role in their association and function. It was specifically demonstrated that in many cases point mutations in the TM domains can partially or completely disrupt the receptor surface expression and function. Here we review studies of the TM domains of B- and T-cell receptors. Furthermore, we use a novel method, PHDtopology, to provide estimates of the exact locations and lengths of the TM domains of the subunit components of these receptors. Most previous studies used single residue hydrophobicity as a criterion for determining the position and length of the TM domains. In contrast, PHDtopology utilizes a system of neural networks and the evolutionary information contained in multiple alignments of related sequences to predict the location, length, and orientation of transmembrane helices. Present results significantly differ from most published estimates of the TM domains of the B- and T-cell receptor components, primarily in the length of the TM domains. These results may lead to modification of putative TM motifs and re-interpretation of the results of studies using mutated TM domains. The availability of PHDtopology on the Internet would make it a valuable tool in the future studies of the TM domains of integral membrane proteins.

Amino Acid Sequence↗

Transmembrane domains in the functions of Fc receptors.

In the present study, we use a novel method, PHDhtm, to predict the exact locations and extents of the transmembrane (TM) domains of multisubunit immunoglobulin Fc-receptors. Whereas most previous studies have used single residue hydrophobicity plots for characterizing of these domains, PHDhtm utilizes a system of neural networks and the evolutionary information contained in multiple alignments of related sequences to predict the above. Present PHDhtm application predicts TM domains of immunoglobulin Fc-receptors that in many cases differ significantly from those derived by using earlier methods. Comparisons of helical wheel projections of the presently derived TM domains from PHDhtm with those produced earlier reveal different hydrophobic moments as well as hydrophobic and hydrophilic surfaces. These differences probably alter the character of subunit association within the receptor complexes. This new algorithm can also be used for other membrane protein complexes and may advance both understanding the principles underlying such complexes formation and design of peptides that can interfere with such TM domain association so as to modulate specific cellular responses.

Amino Acid Sequence↗

Mitochondrial DNA polymerases from yeast to man: a new family of polymerases.

We report the sequence of a 4.5-kb cDNA clone isolated from a human melanoma library which bears high amino acid sequence identity to the yeast mitochondrial (mt) DNA polymerase (Mip1p). This cDNA contains a 3720-bp open reading frame encoding a predicted 140-kDa polypeptide that is 43% identical to Mip1p. The N-terminal part of the sequence contains a 13 glutamine stretch encoded by a CAG trinucleotide repeat which is not found in the other DNA polymerases gamma (Pol gamma). Multiple amino acid sequence alignments with Pol gamma from Saccharomyces cerevisiae, Schizosaccharomyces pombe, Pichia pastoris, Drosophila melanogaster, Xenopus laevis and Mus musculus show that these DNA polymerases form a family strongly conserved from yeast to man and are only loosely related to the Family A DNA polymerases.

Amino Acid Sequence↗

Identification of 167 polymorphisms in 88 genes from candidate neurodegeneration pathways.

Catalogs of intra-gene polymorphisms are needed to facilitate wide-ranging candidate gene-based association studies in common complex diseases. With this in mind, we have scanned multiple alignments of expressed sequence tags and of genomic DNA sequences (PCR products from four to eight unrelated individuals) to find polymorphisms in 195 genes putatively involved in neurodegenerative illness (including components of oxidative stress, excitotoxicity, inflammation, apoptosis and aging). This led to the discovery of 167 polymorphisms in 88 genes. These comprised 163 single nucleotide polymorphisms, one insertion/deletion, and three other variations involving more than one base pair. The polymorphisms were distributed in the exons (87), introns (70), and gene flanking regions (10). Of the exonic polymorphisms, 17 would give rise to non-synonymous amino acid substitutions. These findings now provide a valuable resource for association studies in neurodegenerative disorders such as Alzheimer's disease and Parkinson's disease.

Evolution, Molecular↗

Structural analysis of the PsbQ protein of photosystem II by Fourier transform infrared and circular dichroic spectroscopy and by bioinformatic methods.

The structure of PsbQ, one of the three main extrinsic proteins associated with the oxygen-evolving complex (OEC) of higher plants and green algae, is examined by Fourier transform infrared (FTIR) and circular dichroic (CD) spectroscopy and by computational structural prediction methods. This protein, together with two other lumenally bound extrinsic proteins, PsbO and PsbP, is essential for the stability and full activity of the OEC in plants. The FTIR spectra obtained in both H(2)O and D(2)O suggest a mainly alpha-helix structure on the basis of the relative areas of the constituents of the amide I and I' bands. The FTIR quantitative analyses indicate that PsbQ contains about 53% alpha-helix, 7% turns, 14% nonordered structure, and 24% beta-strand plus other beta-type extended structures. CD analyses indicate that PsbQ is a mainly alpha-helix protein (about 64%), presenting a small percentage assigned to beta-strand ( approximately 7%) and a larger amount assigned to turns and nonregular structures ( approximately 29%). Independent of the spectroscopic analyses, computational methods for protein structure prediction of PsbQ were utilized. First, a multiple alignment of 12 sequences of PsbQ was obtained after an extensive search in the public databases for protein and EST sequences. Based on this alignment, computational prediction of the secondary structure and the solvent accessibility suggest the presence of two different structural domains in PsbQ: a major C-terminal domain containing four alpha-helices and a minor N-terminal domain with a poorly defined secondary structure enriched in proline and glycine residues. The search for PsbQ analogues by fold recognition methods, not based on the secondary structure, also indicates that PsbQ is a four alpha-helix protein, most probably folding as an up-down bundle. The results obtained by both the spectroscopic and computational methods are in agreement, all indicating that PsbQ is mainly an alpha protein, and show the value of using both methodologies for protein structure investigation.

Amides↗