Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Network of dynamically important residues in the open/closed transition in polymerases is strongly conserved.

The open/closed transition in polymerases is a crucial event in DNA replication and transcription. We hypothesize that the residues that transmit the signal for the open/closed transition are also strongly conserved. To identify the dynamically relevant residues, we use an elastic network model of polymerases and probe the residue-specific response to a local perturbation. In a variety of DNA/RNA polymerases, a network of residues spanning the fingers and palm domains is involved in the open/closed transition. The similarity in the network of residues responsible for large-scale domain movements supports the notion of a common induced-fit mechanism in the polymerase families for the formation of a closed ternary complex. Multiple sequence alignment shows that many of these residues are also strongly conserved. Residues with the largest sensitivity to local perturbations include those that are not so obviously involved in the polymerase catalysis. Our results suggest that mutations of the mechanical "hot spots" can compromise the efficiency of the enzyme.

Animals↗

Classification of spider neurotoxins using structural motifs by primary structure features. Single residue distribution analysis and pattern analysis techniques.

In recent years the data on the novel structures of spider toxins have been greatly increasing. The sequence data should be classified. We introduced two primary structure analysis techniques-single residue distribution analysis (SRDA) and pattern analysis for classifying spider polypeptide toxins with molecular weight less than 10kDa. For multiple sequence alignment, we also introduced three novel sequence representation formats named as a simple record, motif record and a pattern record, which can be useful for large-scale analysis of structures. About 300 sequences of spider toxins were analyzed and nine primary structure motifs were identified. New classification of spider toxins was proposed on the basis of previously described principal structural motif (PSM) and extra structural motif (ESM) [Kozlov, S.A., Malyavka, A.A., McCutchen, B., Lu, A., Schepers, E., Herrmann, R., Grishin, E.V., 2005. A novel strategy for the identification of toxin-like structures in spider venom. Proteins 59 (1), 131-140]. Five main structural classes were revealed, and for putative ion channel inhibitors from the most numerous classes 1, 2, and 3, five-digital personal ID numbers were introduced. A reference table with simple, motif and pattern representation sequence formats was created for all analyzed structures.

Amino Acid Motifs↗

Molecular cloning of novel serine proteases and phospholipases A2 from green pit viper (Trimeresurus albolabris) venom gland cDNA library.

Green pit viper (Trimeresurus albolabris) is the most common venomous snake responsible for bites in Bangkok. It causes local edema and systemic hypofibrinogenemia resulted from the thrombin-like, as well as the fibrinolytic effects of the venom. However, the amino acid sequences of these venom proteins have never been reported. In this study, we have cloned five novel serine proteases from the Thai T. albolabris venom gland cDNA library. They were all closely homologous to the corresponding serine proteases from Chinese green viper (Trimeresurus stejnegeri), suggesting the evolutionary proximity of the two species. In addition, their functional activities could be deduced. There were predicted to be two thrombin-like enzymes (GPV-TL1 and GPV-TL-2), two isoforms of a fibrinogenolytic enzyme (albofibrase) and a plasminogen activator (GPV-PA), suggesting that defibrination syndrome in patients is a combination of these enzymatic effects. By multiple sequence alignment, no conserved residue or motif responsible for distinct functions of snake venom serine proteases could be observed. Moreover, one Lys 49 and one Asn 49 phospholipase A2 (PLA2) genes were cloned. Lys 49 PLA2 was predicted to devoid of catalytic activity, but showed a carboxy terminal cytotoxic region. No Asp 49 PLA2 was found in 150 clones screened. This explains the marked limb edema but no hemolysis in patients. These novel serine proteases have potentials to be therapeutic anti-thrombotic and thrombolytic agents in the future.

Animals↗

Neurovirulence of four encephalitogenic dengue 3 virus strains isolated in Malaysia (1992-1994) is not attributed to their envelope protein.

The amino acid sequences of the envelope (E) protein of four encephalitogenic and five non-encephalitogenic dengue 3 virus strains isolated in Malaysia were determined and compared. Multiple sequence alignment revealed a high degree of similarity in the E protein of the strains suggesting that neurovirulence of these four encephalitogenic strains is not attributed to this protein.

Amino Acid Sequence↗

High throughput sequence analysis reveals hitherto unreported recombination in the genus Norovirus.

Viruses of the Norovirus genus (Caliciviridae family) are a major cause of human gastroenteritis. In some viruses, recombination is an important evolutionary process and therefore we should try to discover the quantity and characteristics of such events in Noroviruses. In order to identify recombination events, multiple sequence alignments were assembled from publicly available strains, and were tested using RAT, a recently developed software tool. Strains identified by RAT as putative recombinants were tested further, using a phylogenetic approach, the LARD software, and a Monte Carlo method, to gain additional support for their status. The identification of two previously described recombinants, WUG1 and Snow Mountain, was made. Furthermore, three instances of hitherto unreported recombination implicating Norovirus strains MD 145-12, Gifu'96 and Saitama U4 were found, with good statistical support for the latter two of these cases. Lordsdale-like viruses were highlighted as major contributors to recombination events during Norovirus evolution. Finally, the relevance of recombinants to the worldwide transmission of Norovirus is discussed.

Caliciviridae Infections↗

Complete nucleotide sequence of the hirame rhabdovirus, a pathogen of marine fish.

Reverse transcription-polymerase chain reaction (RT-PCR) derived clones were constructed for the hirame rhabdovirus (HIRRV) CA-9703 strain from Korea, and the DNA was sequenced. The 3'-end of genomic RNA was cloned by poly(A)-tailing of the genomic RNA before reverse transcription, and the 5'-end of the genome was cloned by poly(G)- or poly(C)-tailing of the first strand, followed by PCR. The remainder of the genomic DNA was cloned by reverse transcription-polymerase chain reaction using primers that were based on the published rhabdovirus sequences. The complete genome of HIRRV CA-9703 strain comprises 11,034 nucleotides and encodes six genes in the order of: 3'-leader, N, P, M, G, NV, L, and 5'-trailer. These genes are separated by conserved sequences or gene junctions, with one-nucleotide gene spacers. The first 16 of the 19 nucleotides at the ends of the HIRRV genome are complementary, and the first four nucleotides at the 3'-ends of the HIRRV, infectious hematopoietic necrosis virus (IHNV), viral hemorrhagic septicemia virus (VHSV), and snakehead rhabdovirus (SHRV) genomes are identical. The HIRRV proteins share the highest amino acid sequence homology (ranging from 72% to 92%) with the proteins of IHNV, of all the known fish rhabdoviruses, and the highest sequence homology with respect to the L protein was shared among HIRRV, IHNV, VHSV, and SHRV. Although there were differences in the degrees of relatedness, phylogenetic trees that were derived from multiple sequence alignments of the rhabdovirus proteins showed similar patterns of relationship among these viruses, in which fish Novirhabdoviruses formed a separate clade from spring viremia of carp virus (SVCV), unassigned fish rhabdovirus that was closer to mammalian rhabdoviruses.

Amino Acid Sequence↗

Molecular cloning and characterization of FSH and LH receptors in Atlantic salmon (Salmo salar L.).

Two cDNAs encoding the FSH receptor (FSHR) and the LH receptor (LHR) from Atlantic salmon (Salmo salar) were cloned and characterized. The predicted protein sequence for FSHR comprises a mature protein of 635 amino acids (aa) and a signal peptide of 23aa, and for LHR a mature protein of 701aa and a signal peptide of 27aa. Multiple sequence alignment of Atlantic salmon FSHR and LHR with gonadotropin receptor sequences of available teleosts and representative vertebrates revealed high sequence homology with other salmonids (97-98% for both receptors); amino acid identities ranged from 59 to 67% for FSHR and 47-79% for LHR compared with other teleosts, and between 50 and 52% compared with other vertebrates. The salmon FSHR and LHR showed the typical characteristics of glycoprotein receptors, including a long N-terminal extracellular domain (ECD), seven transmembrane domains and a short C-terminal intracellular domain. The ECD of the Atlantic salmon FSHR and LHR were composed of nine imperfect leucine-rich repeats forming the potential recognition sites for the corresponding hormone. The comparative analysis of the recognition sites in the Atlantic salmon gonadotropin receptors with the corresponding sites in the human receptors showed that the nature of the residues involved in the key contacts with the glycoprotein alpha-subunit were highly conserved. In contrast the recognition sites for the specific beta-subunits showed clear differences between the two salmon gonadotropin receptors and the human receptors. In the salmon LHR the recognition sites for the LH beta-subunit were relatively conserved, while the recognition sites for the FSH beta-subunit in the salmon FSHR showed a higher divergence, suggesting different evolution rates for the two teleost gonadotropin receptors. Both FSHR and LHR were mainly expressed in the ovary and testis, but were also detected at low abundance in extra-gonadal tissues such as gills, brain, liver and heart.

Amino Acid Sequence↗

Molecular phylogenetic analysis of methylenetetrahydrofolate reductase family of proteins.

Methylenetetrahydrofolate reductase (MTHFR) family of proteins catalyze the conversion of 5,10-methylenetetrahydrofolate to 5-methyltetrahydrofolate. They contain a flavin adenine dinucleotide (FAD) as the cofactor and the enzyme in eukaryotes, except in yeast, is known to be allosterically regulated by S-adenosylmethionine. Some cardiovascular diseases, neural tube defects, neuropsychiatric diseases and certain type of cancers in humans are associated with certain polymorphisms of MTHFR. Here, we analyzed 57 of MTHFR polypeptide sequences by multiple sequence alignment and determined previously unrecognized conserved residues that may have a functional or structural importance. A previously unrecognized ATP synthase motif was found in all of the examined plant MTHFRs, suggesting a different functional capability to the plant MTHFRs in addition to the known function. On a phylogenetic tree built, eukaryotic MTHFR proteins formed a clear cluster separated from prokaryotic and archeal relatives. The sequence identities among the eukaryotic MTHFRs were less divergent than the bacterial MTHFRs.

Amino Acid Sequence↗

Sequence analysis of cytochrome bd oxidase suggests a revised topology for subunit I.

Numerous sequences of the cytochrome bd quinol oxidase (cytochrome bd) have recently become available for analysis. The analysis has revealed a small number of conserved residues, a new topology for subunit I and a phylogenetic tree involving extensive horizontal gene transfer. There are 20 conserved residues in subunit I and two in subunit II. Algorithms utilizing multiple sequence alignments predicted a revised topology for cytochrome bd, adding two transmembrane helices to subunit I to the seven that were previously indicated by the analysis of the sequence of the oxidase from E. coli. This revised topology has the effect of relocating the N-terminus and C-terminus to the periplasmic and cytoplasmic sides of the membrane, respectively. The new topology repositions I-H19, the putative ligand for heme b595, close to the periplasmic edge of the membrane, which suggests that the heme b595/heme d active site of the oxidase is located near the outer (periplasmic) surface of the membrane. The most highly conserved region of the sequence of subunit I contains the sequence GRQPW and is located in a predicted periplasmic loop connecting the eighth and ninth transmembrane helices. The potential importance of this region of the protein was previously unsuspected, and it may participate in the binding of either quinol or heme d. There are two very highly conserved glutamates in subunit I, E99 and E107, within the third transmembrane helix (E. coli cytochrome bd-I numbering). It is speculated that these glutamates may be part of a proton channel leading from the cytoplasmic side of the membrane to the heme d oxygen-reactive site, now placed near the periplasmic surface. The revised topology and newly revealed conserved residues provide a clear basis for further experimental tests of these hypotheses. Phylogenetic analysis of the new sequences of cytochrome bd reveals considerable deviation from the 16sRNA tree, suggesting that a large amount of horizontal gene transfer has occurred in the evolution of cytochrome bd.

Amino Acid Sequence↗

A new multidrug resistance protein at the blood-brain barrier.

Porcine brain capillary endothelial cells (PBCEC) cultured in serum-free and hydrocortisone supplemented medium are characterised by high transendothelial electrical resistances and low cell monolayer permeabilities for small solutes very similar to the blood-brain barrier (BBB) in vivo. Differential screening of a subtracted cDNA library disclosed a 2.1-kb mRNA that is overexpressed in hydrocortisone treated PBCEC relative to untreated cells. The mRNA encodes a 656-aa member of the ATP-binding cassette (ABC) superfamily of transporters that we named brain multidrug resistance protein (BMDP). Phylogenetic analysis and multiple sequence alignment showed that porcine BMDP is most related to the human and mouse breast cancer resistance protein (BCRP). Northern blot analysis revealed that BMDP is expressed in brain tissue in vivo and was predominantly localised within the endothelial cells isolated from brain capillaries. Thus, we identified a new transport protein at the BBB that might play an important role in the exclusion of xenobiotics from the brain.

ATP Binding Cassette Transporter, Subfamily B↗

Roles of conserved methionine residues in tobacco acetolactate synthase.

Acetolactate synthase (ALS) catalyzes the first common step in the biosynthesis of valine, leucine, and isoleucine. ALS is the target of several classes of herbicides, including the sulfonylureas, the imidazolinones, and the triazolopyrimidines. The conserved methionine residues of ALS from plants were identified by multiple sequence alignment using ClustalW. The alignment of 17 ALS sequences from plants revealed 149 identical residues, seven of which were methionine residues. The roles of three well-conserved methionine residues (M350, M512, and M569) in tobacco ALS were determined using site-directed mutagenesis. The mutation of M350V, M512V, and M569V inactivated the enzyme and abolished the binding affinity for cofactor FAD. Nevertheless, the secondary structure of each of the mutants determined by CD spectrum was not affected significantly by the mutation. Both M350C and M569C mutants were strongly resistant to three classes of herbicides, Londax (a sulfonylurea), Cadre (an imidazolinone), and TP (a triazolopyrimidine), while M512C mutant did not show a significant resistance to the herbicides. The mutant M350C was more sensitive to pH change, while the mutant M569C showed a profile for pH dependence activity similar to that of wild type. These results suggest that M512 residue is likely located at or near the active site, and that M350 and M569 residues are probably located at the overlapping region between the active site and a common herbicide binding site.

Acetolactate Synthase↗

Evolutionary relationship between K(+) channels and symporters.

The hypothesis is presented that at least four families of putative K(+) symporter proteins, Trk and KtrAB from prokaryotes, Trk1,2 from fungi, and HKT1 from wheat, evolved from bacterial K(+) channel proteins. Details of this hypothesis are organized around the recently determined crystal structure of a bacterial K(+) channel: i. e., KcsA from Streptomyces lividans. Each of the four identical subunits of this channel has two fully transmembrane helices (designated M1 and M2), plus an intervening hairpin segment that determines the ion selectivity (designated P). The symporter sequences appear to contain four sequential M1-P-M2 motifs (MPM), which are likely to have arisen from gene duplication and fusion of the single MPM motif of a bacterial K(+) channel subunit. The homology of MPM motifs is supported by a statistical comparison of the numerical profiles derived from multiple sequence alignments formed for each protein family. Furthermore, these quantitative results indicate that the KtrAB family of symporters has remained closest to the single-MPM ancestor protein. Strong sequence evidence is also found for homology between the cytoplasmic C-terminus of numerous bacterial K(+) channels and the cytoplasm-resident TrkA and KtrA subunits of the Trk and KtrAB symporters, which in turn are homologous to known dinucleotide-binding domains of other proteins. The case for homology between bacterial K(+) channels and the four families of K(+) symporters is further supported by the accompanying manuscript, in which the patterns of residue conservation are demonstrated to be similar to each other and consistent with the known 3D structure of the KcsA K(+) channel.

Amino Acid Sequence↗

Pseudomonas fluorescens mannitol 2-dehydrogenase and the family of polyol-specific long-chain dehydrogenases/reductases: sequence-based classification and analysis of structure-function relationships.

Multiple sequence alignment and analysis of evolutionary relationships have been used to characterize a family of polyol-specific long-chain dehydrogenases/reductases (PSLDRs). At the present time, 66 known and putative NAD(P)H-dependent oxidoreductases of mainly prokaryotic origin and between 357 and 544 amino acids in length constitute this family. The family is shown to include D-mannitol 2-dehydrogenase, D-mannonate 5-oxidoreductase, D-altronate 5-oxidoreductase, D-arabinitol 4-dehydrogenase, and D-mannitol-1-phosphate 5-dehydrogenase which form individual sub-families (defined by internal sequence identity of >/=30%) having distant origin and divergent substrate specificity but clearly displaying entire-chain relationship. When all forms are aligned, only three residues, Gly-33, Asp-230, and Lys-295 (in the numbering of Pseudomonas fluorescens D-mannitol 2-dehydrogenase (PsM2DH)) are strictly conserved. By combining sequence alignment with the known structure of PsM2DH and results from site-directed mutagenesis, we have developed a structure/function analysis for the family. Gly-33 is in the N-terminal coenzyme-binding domain and part of a nucleotide fingerprint region for the family, and Asp-230 and Lys-295 are at an interdomain segment contributing to the active site in which the lysine likely functions as the catalytic general acid/base. PSLDRs do not require a metal cofactor for activity and are specific for transferring the 4-pro-S hydrogen from NAD(P)H. Comparisons reveal that the core part of the two-domain fold has been conserved throughout all family members, perhaps reflecting the recruitment of a stable oxidoreductase structure and extensive trimming thereof to acquire functional properties specific to each sub-family. They also identify interactions that define the chemical mechanism of oxidoreduction and likely contribute to substrate and co-substrate specificities and are thus relevant for protein engineering.

Amino Acid Sequence↗

The molecular peculiarities of catalase-peroxidases.

In developing ideas of how protein structure modifies haem reactivity, the activity of Class I of the plant peroxidase superfamily (including cytochrome c peroxidase, ascorbate peroxidase and catalase-peroxidases (KatGs)) is an exciting field of research. Despite striking sequence homologies, there are dramatic differences in catalytic activity and substrate specificity with KatGs being the only member with substantial catalase activity. Based on multiple sequence alignment performed for Class I peroxidases, we present a hypothesis for the pronounced catalase activity of KatGs. In their catalytic domains KatGs are shown to possess three large insertions, two of them are typical for KatGs showing highly conserved sequence patterns. Besides an extra C-terminal copy of the ancestral hydroperoxidase gene resulting from gene duplication, these two large loops are likely to control the orientation of both the haem group and of essential residues in the active site. They seem to modulate the access of substrates to the prosthetic group at the distal side as well as the flexibility and character of the bond between the proximal histidine and the ferric iron. The hypothesis presented opens new possibilities in the rational engineering of peroxidases.

Amino Acid Sequence↗

A refined structure of human aquaporin-1.

A refined structure of the human water channel aquaporin-1 is presented. The model rests on the high resolution X-ray structure of the homologous bacterial glycerol transporter GlpF, electron crystallographic data at 3.8 A resolution and a multiple sequence alignment of the aquaporin superfamily. The crystallographic R and free R values (36.7% and 37.8%) for the refined structure are significantly lower than for previous models. Improved geometry and enhanced stability in molecular dynamics simulations demonstrate a significant improvement of the aquaporin-1 structure. Comparison with previous aquaporin-1 models shows significant differences, not only in the loop regions, but also in the core of the water channel.

Aquaporin 1↗

BTK-2, a new inhibitor of the Kv1.1 potassium channel purified from Indian scorpion Buthus tamulus.

A novel inhibitor of voltage-gated potassium channel was isolated and purified to homogeneity from the venom of the red scorpion Buthus tamulus. The primary sequence of this toxin, named BTK-2, as determined by peptide sequencing shows that it has 32 amino acid residues with six conserved cysteines. The molecular weight of the toxin was found to be 3452 Da. It was found to block the human potassium channel hKv1.1 (IC(50)=4.6 microM). BTK-2 shows 40-70% sequence similarity to the family of the short-chain toxins that specifically block potassium channels. Multiple sequence alignment helps to categorize the toxin in the ninth subfamily of the K+ channel blockers. The modeled structure of BTK-2 shows an alpha/beta scaffold similar to those of the other short scorpion toxins. Comparative analysis of the structure with those of the other toxins helps to identify the possible structure-function relationship that leads to the difference in the specificity of BTK-2 from that of the other scorpion toxins. The toxin can also be used to study the assembly of the hKv1.1 channel.

Amino Acid Sequence↗

Diversification and evolution of L-myo-inositol 1-phosphate synthase.

L-myo-Inositol 1-phosphate synthase (MIPS, EC 5.5.1.4), the key enzyme in the inositol and phosphoinositide biosynthetic pathway, is present throughout evolutionarily diverse organisms and is considered an ancient protein/gene. Analysis by multiple sequence alignment, phylogenetic tree generation and comparison of newly determined crystal structures provides new insight into the origin and evolutionary relationships among the various MIPS proteins/genes. The evolution of the MIPS protein/gene among the prokaryotes seems more diverse and complex than amongst the eukaryotes. However, conservation of a 'core catalytic structure' among the MIPS proteins implies an essential function of the enzyme in cellular metabolism throughout the biological kingdom.

Amino Acid Sequence↗

Structural localization of disease-associated sequence variations in the NACHT and LRR domains of PYPAF1 and NOD2.

Several autoinflammatory diseases with distinct clinical manifestations have been associated with sequence variations in the gene products PYPAF1/CIAS1 and NOD2/CARD15. Both proteins belong to the PYD/CARD-containing family of apoptosis regulators and activators of pro-inflammatory caspases. To gain insight into the dysfunctional role of sequence alterations, we assembled a structure-based multiple sequence alignment of family members and related proteins. This allowed us to analyze the putative effect of the alterations on the function of nucleotide-binding (NACHT) and leucine-rich repeat (LRR) domains shared by the family members. In support of this analysis, we carefully selected template structures for the NACHT and LRR domains and mapped the genetic variations onto 3D domain models. Additionally, we propose a model of the NACHT and LRR domain complex. Our study revealed that many of the disease-associated sequence variants are located close to highly conserved sequence regions of functional relevance and are spatially adjacent in the predicted 3D structure. The implications on the domain functions such as NTP-hydrolysis or oligomerization are discussed.

Amino Acid Sequence↗