Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

The binding site for UCH-L3 on ubiquitin: mutagenesis and NMR studies on the complex between ubiquitin and UCH-L3.

The ubiquitin fold is a versatile and widely used targeting signal that is added post-translationally to a variety of proteins. Covalent attachment of one or more ubiquitin domains results in localization of the target protein to the proteasome, the nucleus, the cytoskeleton or the endocytotic machinery. Recognition of the ubiquitin domain by a variety of enzymes and receptors is vital to the targeting function of ubiquitin. Several parallel pathways exist and these must be able to distinguish among ubiquitin, several different types of polymeric ubiquitin, and the various ubiquitin-like domains. Here we report the first molecular description of the binding site on ubiquitin for ubiquitin C-terminal hydrolase L3 (UCH-L3). The site on ubiquitin was experimentally determined using solution NMR, and site-directed mutagenesis. The site on UCH-L3 was modeled based on X-ray crystallography, multiple sequence alignments, and computer-aided docking. Basic residues located on ubiquitin (K6, K11, R72, and R74) are postulated to contact acidic residues on UCH-L3 (E10, E14, D33, E219). These putative interactions are testable and fully explain the selectivity of ubiquitin domain binding to this enzyme.

Allosteric Site↗

Effective use of sequence correlation and conservation in fold recognition.

Protein families are a rich source of information; sequence conservation and sequence correlation are two of the main properties that can be derived from the analysis of multiple sequence alignments. Sequence conservation is related to the direct evolutionary pressure to retain the chemical characteristics of some positions in order to maintain a given function. Sequence correlation is attributed to the small sequence adjustments needed to maintain protein stability against constant mutational drift. Here, we showed that sequence conservation and correlation were each frequently informative enough to detect incorrectly folded proteins. Furthermore, combining conservation, correlation, and polarity, we achieved an almost perfect discrimination between native and incorrectly folded proteins. Thus, we made use of this information for threading by evaluating the models suggested by a threading method according to the degree of proximity of the corresponding correlated, conserved, and apolar residues. The results showed that the fold recognition capacity of a given threading approach could be improved almost fourfold by selecting the alignments that score best under the three different sequence-based approaches.

Amino Acid Sequence↗

Structural clues in the sequences of the aquaporins.

The large number of sequences available for the aquaporin family represents a valuable source of information to incorporate into three-dimensional structure determination. Phylogenetic analysis was used to define type sequences to avoid extreme over-representation of some subfamilies, and as a measure of the quality of multiple sequence alignment. Inspection of the sequence alignment suggested eight conserved segments that define the core architecture of six transmembrane helices and two functional loops, B and E, projecting into the plane of the membrane. The sum of the core segments and the minimum lengths of the interlinking loops constitute the 208 residues necessary to satisfy the aquaporin architecture. Analysis of hydrophobic and conservation periodicity and of correlated mutations across the alignment indicated the likely assignment and orientation of the helices in the bilayer. This assignment is examined with respect to the structure of the erythrocyte aquaporin 1 determined by electron crystallography. The aquaporin 1 tetramer is described as three rings of helices, each ring with a different exposure to the lipid environment. The sequence analysis clearly suggests that two helices are exposed along their whole lengths, two helices are exposed only at their N termini, and two helices are not exposed to lipid. It is further proposed that, besides loops B and E, the highly conserved motifs on helices 1 and 4, ExxxTxxF/L, could line the water channel.

Amino Acid Sequence↗

Co-evolution of proteins with their interaction partners.

The divergent evolution of proteins in cellular signaling pathways requires ligands and their receptors to co-evolve, creating new pathways when a new receptor is activated by a new ligand. However, information about the evolution of binding specificity in ligand-receptor systems is difficult to glean from sequences alone. We have used phosphoglycerate kinase (PGK), an enzyme that forms its active site between its two domains, to develop a standard for measuring the co-evolution of interacting proteins. The N-terminal and C-terminal domains of PGK form the active site at their interface and are covalently linked. Therefore, they must have co-evolved to preserve enzyme function. By building two phylogenetic trees from multiple sequence alignments of each of the two domains of PGK, we have calculated a correlation coefficient for the two trees that quantifies the co-evolution of the two domains. The correlation coefficient for the trees of the two domains of PGK is 0. 79, which establishes an upper bound for the co-evolution of a protein domain with its binding partner. The analysis is extended to ligands and their receptors, using the chemokines as a model. We show that the correlation between the chemokine ligand and receptor trees' distances is 0.57. The chemokine family of protein ligands and their G-protein coupled receptors have co-evolved so that each subgroup of chemokine ligands has a matching subgroup of chemokine receptors. The matching subfamilies of ligands and their receptors create a framework within which the ligands of orphan chemokine receptors can be more easily determined. This approach can be applied to a variety of ligand and receptor systems.

Chemokines↗

An aspartic acid residue in TPR-1, a specific region of protein-priming DNA polymerases, is required for the functional interaction with primer terminal protein.

A multiple sequence alignment of eukaryotic-type DNA polymerases led to the identification of two regions of amino acid residues that are only present in the group of DNA polymerases that make use of terminal proteins. (TPs) as primers to initiate DNA replication of linear genomes. These amino acid regions (named terminal region (TPR protein-1 and TPR-2) are inserted between the generally conserved motifs Dx(2)SLYP and Kx(3)NSxYG (TPR-1) and motifs Kx(3)NSxYG and YxDTDS (TPR-2) of the eukaryotic-type family of DNA polymerases. We carried out site-directed mutagenesis in two of the most conserved residues of phi29 DNA polymerase TPR-1 to study the possible role of this specific region. Two mutant DNA polymerases, in conserved residues AsP332 and Leu342, were purified and subjected to a detailed biochemical analysis of their enzymatic activities. Both mutant DNA polymerases were essentially normal when assayed for synthetic activities in DNA-primed reactions. However, mutant D332Y was drastically affected in phi29 TP-DNA replication as a consequence of a large reduction in the catalytic efficiency of the protein-primed reactions. The molecular basis of this defect is a non-functional interaction with TP that strongly reduces the activity of the DNA polymerase/TP heterodimer.

Amino Acid Motifs↗

ConSurf: an algorithmic tool for the identification of functional regions in proteins by surface mapping of phylogenetic information.

Experimental approaches for the identification of functionally important regions on the surface of a protein involve mutagenesis, in which exposed residues are replaced one after another while the change in binding to other proteins or changes in activity are recorded. However, practical considerations limit the use of these methods to small-scale studies, precluding a full mapping of all the functionally important residues on the surface of a protein. We present here an alternative approach involving the use of evolutionary data in the form of multiple-sequence alignment for a protein family to identify hot spots and surface patches that are likely to be in contact with other proteins, domains, peptides, DNA, RNA or ligands. The underlying assumption in this approach is that key residues that are important for binding should be conserved throughout evolution, just like residues that are crucial for maintaining the protein fold, i.e. buried residues. A main limitation in the implementation of this approach is that the sequence space of a protein family may be unevenly sampled, e.g. mammals may be overly represented. Thus, a seemingly conserved position in the alignment may reflect a taxonomically uneven sampling, rather than being indicative of structural or functional importance. To avoid this problem, we present here a novel methodology based on evolutionary relations among proteins as revealed by inferred phylogenetic trees, and demonstrate its capabilities for mapping binding sites in SH2 and PTB signaling domains. A computer program that implements these ideas is available freely at: http://ashtoret.tau.ac.il/ approximately rony

Algorithms↗

Three-dimensional cluster analysis identifies interfaces and functional residue clusters in proteins.

Three-dimensional cluster analysis offers a method for the prediction of functional residue clusters in proteins. This method requires a representative structure and a multiple sequence alignment as input data. Individual residues are represented in terms of regional alignments that reflect both their structural environment and their evolutionary variation, as defined by the alignment of homologous sequences. From the overall (global) and the residue-specific (regional) alignments, we calculate the global and regional similarity matrices, containing scores for all pairwise sequence comparisons in the respective alignments. Comparing the matrices yields two scores for each residue. The regional conservation score (C(R)(x)) defines the conservation of each residue x and its neighbors in 3D space relative to the protein as a whole. The similarity deviation score (S(x)) detects residue clusters with sequence similarities that deviate from the similarities suggested by the full-length sequences. We evaluated 3D cluster analysis on a set of 35 families of proteins with available cocrystal structures, showing small ligand interfaces, nucleic acid interfaces and two types of protein-protein interfaces (transient and stable). We present two examples in detail: fructose-1,6-bisphosphate aldolase and the mitogen-activated protein kinase ERK2. We found that the regional conservation score (C(R)(x)) identifies functional residue clusters better than a scoring scheme that does not take 3D information into account. C(R)(x) is particularly useful for the prediction of poorly conserved, transient protein-protein interfaces. Many of the proteins studied contained residue clusters with elevated similarity deviation scores. These residue clusters correlate with specificity-conferring regions: 3D cluster analysis therefore represents an easily applied method for the prediction of functionally relevant spatial clusters of residues in proteins.

Adenosine Triphosphate↗

A novel method of protein secondary structure prediction with high segment overlap measure: support vector machine approach.

We have introduced a new method of protein secondary structure prediction which is based on the theory of support vector machine (SVM). SVM represents a new approach to supervised pattern classification which has been successfully applied to a wide range of pattern recognition problems, including object recognition, speaker identification, gene function prediction with microarray expression profile, etc. In these cases, the performance of SVM either matches or is significantly better than that of traditional machine learning approaches, including neural networks.The first use of the SVM approach to predict protein secondary structure is described here. Unlike the previous studies, we first constructed several binary classifiers, then assembled a tertiary classifier for three secondary structure states (helix, sheet and coil) based on these binary classifiers. The SVM method achieved a good performance of segment overlap accuracy SOV=76.2 % through sevenfold cross validation on a database of 513 non-homologous protein chains with multiple sequence alignments, which out-performs existing methods. Meanwhile three-state overall per-residue accuracy Q(3) achieved 73.5 %, which is at least comparable to existing single prediction methods. Furthermore a useful "reliability index" for the predictions was developed. In addition, SVM has many attractive features, including effective avoidance of overfitting, the ability to handle large feature spaces, information condensing of the given data set, etc. The SVM method is conveniently applied to many other pattern classification tasks in biology.

Computer Simulation↗

Molecular cloning of the arylsulfate sulfotransferase gene and characterization of its product from Enterobacter amnigenus AR-37.

The gene encoding the Enterobacter amnigenus AR-37 arylsulfate sulfotransferase (ASST) was cloned, sequenced, and expressed in Escherichia coli NM522. Sequencing led to the identification of three contiguous open reading frames (ORFs) on the same strand. Based on amino acid sequence homology, ORF1, ORF2, and ORF3 are designated astA, dsbA, and dsbB, respectively. A multiple sequence alignment revealed conserved regions in ASST. An N-terminal amino acid sequence analysis of the purified ASST from E. coli NM522 (pEAST72) showed that it is subject to N-terminal processing. The specific activity of purified ASST is 436.5 U/mg of protein. The enzyme is a monomeric protein with a molecular mass of 64 kDa. Using phenol as an acceptor substrate, 4-methylumbelliferyl sulfate is the best donor substrate, followed by beta-naphthyl sulfate, p-nitrophenyl sulfate (PNS), and alpha-naphthyl sulfate. For PNS, alpha-naphthol is the best acceptor substrate, followed by phenol, resorcinol, p-acetaminophen, tyramine, and tyrosine. The enzyme has a different acceptor specificity than the enzyme purified from Eubacterium A-44. It is similar to Klebsiella K-36 and Haemophilus K-12. The apparent K(m) values for PNS using phenol as an acceptor and for phenol using PNS as a donor are 0.163 and 0.314 mM, respectively. The pI and optimum pH are 6.1 and 9.0, respectively.

Amino Acid Sequence↗

Molecular cloning, expression, purification, and characterization of fructose-1,6-bisphosphate aldolase from Thermus aquaticus.

Fructose-1,6-bisphosphate aldolase from the thermophilic eubacteria, Thermus aquaticus YT-1, was cloned and sequenced. Nucleotide-sequence analysis revealed an open reading frame coding for a 33-kDa protein of 305 amino acids having amino acid sequence typical of thermophilic adaptation. Multiple sequence alignment classifies the enzyme as a class II B aldolase that shares similarity with aldolases from other extremophiles: Thermotoga maritima, Aquifex aeolicus, and Helicobacter pylori (49--54% identity, 76--81% homology). Taq FBP aldolase was overexpressed under tac promoter control in Escherichia coli and purified to homogeneity using heat treatment followed by two chromatographic steps. Yields of 40--50 mg of monodisperse protein were obtained per liter of culture. The quaternary structure is that of a homotetramer stabilized by an apparent 21-amino-acid insertion sequence. The recombinant protein is thermostable for at least 45 min at 80 degrees C with little residual activity below 60 degrees C. Kinetic characterization at 70 degrees C, the optimal growth temperature for T. aquaticus, indicates extreme negative subunit cooperativity (h = 0.32) with a limiting K(m) of 305 microM. The maximal specific activity (V(max)) is 46 U/mg at 70 degrees C.

Amino Acid Sequence↗

Adelaide river rhabdovirus expresses consecutive glycoprotein genes as polycistronic mRNAs: new evidence of gene duplication as an evolutionary process.

A 3914 nucleotide region of the Adelaide River virus (ARV) genome, located immediately downstream of the M2 gene, has been cloned and sequenced. The region contains two long open reading frames (ORFs). The first encodes a protein comprising 660 amino acids which shares extensive sequence homology with the virion G protein of bovine ephemeral fever virus (BEFV) and less but significant homology with other rhabdovirus glycoproteins. The size and structural characteristics of the product indicate that it represents the 90-kDa ARV virion G protein. The second ORF encodes a polypeptide of 609 residues with nine potential glycosylation sites which is most closely related to the BEFV non-structural glycoprotein (GNS). In infected mammalian cells, the ARV G and GNS genes are transcribed primarily as a polycistronic mRNA which appears to extend from the consensus sequence (AACAG) at the start of the G gene to the next recognized polyadenylation signal (CATG[A]7) located 697 nucleotides downstream of the GNS protein termination codon. Less abundant mRNAs which appeared to initiate at consensus sequences immediately preceding and following the GNS ORF and terminate at the same polyadenylation signal were also detected. Polyadenylation-like sequences at the end of each ORF do not appear to be recognized as transcription stop signals. Multiple sequence alignments and phylogenetic analyses indicated that the ARV G and GNS glycoproteins, like those of BEFV, are structurally related and appear to have evolved at different rates from a common ancestral gene. A copy-choice mechanism, involving upstream relocation of the polymerase during replication, is proposed to account for the evolution of the tandem glycoprotein genes.

Amino Acid Sequence↗

Evolutionary relationships among putative RNA-dependent RNA polymerases encoded by a mitochondrial virus-like RNA in the Dutch elm disease fungus, Ophiostoma novo-ulmi, by other viruses and virus-like RNAs and by the Arabidopsis mitochondrial genome.

The nucleotide sequence (2617 nucleotides) of virus-like double-stranded (ds) RNA 3a in a diseased isolate, Log1/3-8d2 (Ld), of the ascomycete fungus Ophiostoma novo-ulmi has been determined. One strand of the dsRNA contains an open reading frame (ORF) with the potential to encode a protein of 718 amino acids, and the complementary strand contains two smaller ORFs with the potential to encode proteins of 178 and 182 amino acids, respectively. The large ORF contains 12 UGA codons which code for tryptophan in ascomycete mitochondria and has a codon bias typical of mitochondrial genes, consistent with the localization of Ld dsRNAs within the mitochondria. The amino acid sequence contains motifs characteristic of RNA-dependent RNA polymerases (RdRps). This putative RdRp was shown to be related to putative RdRps of mitochondrial dsRNAs of another ascomycete and a basidiomycete fungus and also to a putative RdRp encoded by the mitochondrial genome of Arabidopsis thaliana. In multiple sequence alignments, the fungal mitochondrial dsRNA-encoded RdRp-like proteins formed a cluster, ancestrally related to the RdRps of the yeast 20S and 23S RNA replicons and of the positive-stranded RNA bacteriophages of the Leviviridae family, but distinct from RdRps of other families and genera of fungal RNA viruses and related plant and animal RNA viruses. Northern blot analysis with RNA 3a strand-specific probes indicated that nucleic acid extracts of Ld contain more single-stranded (positive-stranded) RNA than dsRNA, consistent with an evolutionary relationship between RNA 3a and positive-stranded RNA phages.

Amino Acid Sequence↗

Bioinformatics in protein analysis.

The chapter gives an overview of bioinformatic techniques of importance in protein analysis. These include database searches, sequence comparisons and structural predictions. Links to useful World Wide Web (WWW) pages are given in relation to each topic. Databases with biological information are reviewed with emphasis on databases for nucleotide sequences (EMBL, GenBank, DDBJ), genomes, amino acid sequences (Swissprot, PIR, TrEMBL, GenePept), and three-dimensional structures (PDB). Integrated user interfaces for databases (SRS and Entrez) are described. An introduction to databases of sequence patterns and protein families is also given (Prosite, Pfam, Blocks). Furthermore, the chapter describes the widespread methods for sequence comparisons, FASTA and BLAST, and the corresponding WWW services. The techniques involving multiple sequence alignments are also reviewed: alignment creation with the Clustal programs, phylogenetic tree calculation with the Clustal or Phylip packages and tree display using Drawtree, njplot or phylo_win. Finally, the chapter also treats the issue of structural prediction. Different methods for secondary structure predictions are described (Chou-Fasman, Garnier-Osguthorpe-Robson, Predator, PHD). Techniques for predicting membrane proteins, antigenic sites and postranslational modifications are also reviewed.

Computational Biology↗

Protein engineering in the alpha-amylase family: catalytic mechanism, substrate specificity, and stability.

Most starch hydrolases and related enzymes belong to the alpha-amylase family which contains a characteristic catalytic (beta/alpha)8-barrel domain. Currently known primary structures that have sequence similarities represent 18 different specificities, including starch branching enzyme. Crystal structures have been reported in three of these enzyme classes: the alpha-amylases, the cyclodextrin glucanotransferases, and the oligo-1,6-glucosidases. Throughout the alpha-amylase family, only eight amino acid residues are invariant, seven at the active site and a glycine in a short turn. However, comparison of three-dimensional models with a multiple sequence alignment suggests that the diversity in specificity arises by variation in substrate binding at the beta-->alpha loops. Designed mutations thus have enhanced transferase activity and altered the oligosaccharide product patterns of alpha-amylases, changed the distribution of alpha-, beta- and gamma-cyclodextrin production by cyclodextrin glucanotransferases, and shifted the relative alpha-1,4:alpha-1,6 dual-bond specificity of neopullulanase. Barley alpha-amylase isozyme hybrids and Bacillus alpha-amylases demonstrate the impact of a small domain B protruding from the (beta/alpha)8-scaffold on the function and stability. Prospects for rational engineering in this family include important members of plant origin, such as alpha-amylase, starch branching and debranching enzymes, and amylomaltase.

Amino Acid Sequence↗

Response regulators of bacterial signal transduction systems: selective domain shuffling during evolution.

Response regulators of bacterial sensory transduction systems generally consist of receiver module domains covalently linked to effector domains. The effector domains include DNA binding and/or catalytic units that are regulated by sensor kinase-catalyzed aspartyl phosphorylation within their receiver modules. Most receiver modules are associated with three distinct families of DNA binding domains, but some are associated with other types of DNA binding domains, with methylated chemotaxis protein (MCP) demethylases, or with sensor kinases. A few exist as independent entities which regulate their target systems by noncovalent interactions. In this study the molecular phylogenies of the receiver modules and effector domains of 49 fully sequenced response regulators and their homologues were determined. The three major, evolutionarily distinct, DNA binding domains found in response regulators were evaluated for their phylogenetic relatedness, and the phylogenetic trees obtained for these domains were compared with those for the receiver modules. Members of one family (family 1) of DNA binding domains are linked to large ATPase domains which usually function cooperatively in the activation of E. coli sigma 54-dependent promoters or their equivalents in other bacteria. Members of a second family (family 2) always function in conjunction with the E. coli sigma 70 or its equivalent in other bacteria. A third family of DNA binding domains (family 3) functions by an uncharacterized mechanism involving more than one sigma factor. These three domain families utilize distinct helix-turn-helix motifs for DNA binding. The phylogenetic tree of the receiver modules revealed three major and several minor clusters of these domains. The three major receiver module clusters (clusters 1, 2, and 3) generally function with the three major families of DNA binding domains (families 1, 2, and 3, respectively) to comprise three classes of response regulators (classes 1, 2, and 3), although several exceptions exist. The minor clusters of receiver modules were usually, but not always, associated with other types of effector domains. Finally, several receiver modules did not fit into a cluster. It was concluded that receiver modules usually diverged from common ancestral protein domains together with the corresponding effector domains, although domain shuffling, due to intragenic splicing and fusion, must have occurred during the evolution of some of these proteins. Multiple sequence alignments of the 49 receiver modules and their various types of effector domains, together with other homologous domains, allowed definition of regions of striking sequence similarity and degrees of conservation of specific residues. Sequence data were correlated with structure/function when such information was available.(ABSTRACT TRUNCATED AT 250 WORDS)

Adenosine Triphosphatases↗

The evolutionary divergence of neurotransmitter receptors and second-messenger pathways.

Members of the superfamily of G-protein-coupled neurotransmitter receptors have a conserved secondary structure, a moderate and reasonably steady rate of sequence change, and usually lack introns within the coding sequence. These properties are advantageous for evolutionary studies. The duplication and divergence of the genes in this gene family led to the formation of distinct neurotransmitter pathways and may have facilitated the evolution of complex nervous systems. I have analyzed this evolutionary divergence by quantitative multiple sequence alignment, bootstrap resampling, and statistical analysis of 49 adrenergic, muscarinic cholinergic, dopamine, and octopamine receptor sequences from 12 animal species. The results indicate that the first event to occur within this gene family was the divergence of the catecholamine receptors from the muscarinic acetylcholine receptors, which occurred prior to the divergence of the arthropod and vertebrate lineages. Subsequently, the ability to activate specific second-messenger pathways diverged independently in both the muscarinic and the catecholamine receptors. This appears to have occurred after the divergence of the arthropod and vertebrate lineages but before the divergence of the avian and mammalian lineages. However, the second-messenger pathways activated by adrenergic and dopamine receptors did not diverge independently. Rather, the ability of the catecholamine receptors to bind to specific ligands, such as epinephrine, norepinephrine, dopamine, or octopamine, was repeatedly modified in evolutionary history, and in some cases was modified after the divergence of the second-messenger pathways.

Adenylyl Cyclases↗

Comparison of lantibiotic gene clusters and encoded proteins.

Lantibiotics form a group of modified peptides with unique structures, containing post-translationally modified amino acids such as dehydrated and lanthionine residues. In the gram-positive bacteria that secrete these lantibiotics, the gene clusters flanking the structural genes for various linear (type A) lantibiotics have recently been characterized. The best studied representatives are those of nisin (nis), subtilin (spa), epidermin (epi), Pep5 (pep), cytolysin (cyl), lactocin S (las) and lacticin 481 (lct). Comparison of the lantibiotic gene clusters shows that they contain conserved genes that probably encode similar functions. The nis, spa, epi and pep clusters contain lanB and lanC genes that are presumed to code for two types of enzymes that have been implicated in the modification reactions characteristic of all lantibiotics, i.e. dehydration and thio-ether ring formation. The cyl, las and lct gene clusters have no homologue of the lanB gene, but they do contain a much larger lanM gene that is the lanC gene homologue. Most lantibiotic gene clusters contain a lanP gene encoding a serine protease that is presumably involved in the proteolytic processing of the prelantibiotics. All clusters contain a lanT gene encoding an ABC transporter likely to be involved in the export of (precursors of) the lantibiotics. The lanE, lanF and lanG genes in the nis, spa and epi clusters encode another transport system that is possibly involved in self-protection. In the nisin and subtilin gene clusters two tandem genes, lanR and lanK, have been located that code for a two-component regulatory system. Finally, non-homologous genes are found in some lantibiotic gene clusters. The nisI and spaI genes encode lipoproteins that are involved in immunity, the pepI gene encodes a membrane-located immunity protein, and epiD encodes an enzyme involved in a post-translational modification found only in the C-terminus of epidermin. Several genes of unknown function are also found in the las gene cluster. A database has been assembled for all putative gene products of type A lantibiotic gene clusters. Database searches, multiple sequence alignment and secondary structure prediction have been used to identify conserved sequence segments in the LanB, LanC, LanE, LanF, LanG, LanK, LanM, LanP, LanR and LanT gene products that may be essential for structure and function. This database allows for a rapid screening of newly determined sequences in lantibiotic gene clusters.

ATP-Binding Cassette Transporters↗

Identification of novel homologues of three low molecular weight subunits of the mitochondrial bc1 complex.

Large-scale random cDNA sequencing projects have been started for several organisms and are a valuable tool for the analysis of quantitative and qualitative aspects of gene expression. However, the reliability of the obtained data is limited as most of the clones are only partially analysed on one strand. As a consequence the sequence entries derived from random cDNA sequencing projects usually comprise incomplete open reading frames. They nevertheless define complete and reliable coding sequences, if two prerequisites are fulfilled: (i) the clones encode very small proteins, and (ii) the clones have a high frequency in the cDNA-banks. The present study describes the use of cDNA databases for the identification of homologues of three low-molecular-weight subunits of the mitochondrial bc1 complex, termed the QCR6, QCR9 and QCR10 proteins. These polypeptides are only characterized for a small number of organisms, have a scarcely defined function and exhibit a low degree of structural conservation if compared between different species. Several clones were identified for each polypeptide by searches with TBLASTN using the known sequences as probes. Most of the database entries contain complete open reading frames and sequencing queries could be excluded due to the abundancy of the clones. Multiple sequence alignments are presented for all three polypeptides and consensus sequences are given which may provide a basis for the investigation of the proteins by site-directed mutagenesis.

Animals↗