Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Secator: a program for inferring protein subfamilies from phylogenetic trees.

With the huge increase of protein data, an important problem is to estimate, within a large protein family, the number of sensible subsets for subsequent in-depth structural, functional, and evolutionary analyses. To tackle this problem, we developed a new program, Secator, which implements the principle of an ascending hierarchical method using a distance matrix based on a multiple alignment of protein sequences. Dissimilarity values assigned to the nodes of a deduced phylogenetic tree are partitioned by a new stopping rule introduced to automatically determine the significant dissimilarity values. The quality of the clusters obtained by Secator is verified by a separate Jackknife study. The method is demonstrated on 24 large protein families covering a wide spectrum of structural and sequence conservation and its usefulness and accuracy with real biological data is illustrated on two well-studied protein families (the Sm proteins and the nuclear receptors).

Animals↗

Plant lanosterol synthase: divergence of the sterol and triterpene biosynthetic pathways in eukaryotes.

Sterols, essential eukaryotic constituents, are biosynthesized through either cyclic triterpenes, lanosterol (fungi and animals) or cycloartenol (plants). The cDNA for OSC7 of Lotus japonicus was shown to encode lanosterol synthase (LAS) by the complementation of a LAS-deficient mutant yeast and structural identification of the accumulated lanosterol. A double site-directed mutant of OSC7, in which amino acid residues crucial for the reaction specificity were changed to the cycloartenol synthase (CAS) type, produced parkeol and cycloartenol. The multiple amino acid sequence alignment of a conserved region suggests that the LAS of different eukaryotic lineages emerged from the ancestral CAS by convergent evolution.

Amino Acid Sequence↗

RNA sequence of potato virus X strain HB.

The genomic RNA of the potato virus X (PVX) strain HB, isolated in Bolivia and able to overcome all known resistance genes, has been cloned and sequenced. The PVXHB RNA sequence is 6432 nucleotides long and contains, similarly to the RNAs of other PVX strains, five open reading frames encoding proteins of M(r)s 165.1K, 24.5K, 12.4K, 7.6K and 25.1K (coat protein), respectively. Multiple amino acid sequence alignments of the coat proteins of four PVX strains identified eight amino acid residues unique for PVXHB. Structural prediction comparisons of the coat proteins of PVXHB and of the other strains suggest a general structural similarity. However, two of the eight amino acid residues unique for strain HB gave rise to a change in the predicted coat protein structure, suggesting a possible involvement in the resistance-breaking activity of PVXHB.

Amino Acid Sequence↗

Experimental validation of predicted mammalian erythroid cis-regulatory modules.

Multiple alignments of genome sequences are helpful guides to functional analysis, but predicting cis-regulatory modules (CRMs) accurately from such alignments remains an elusive goal. We predict CRMs for mammalian genes expressed in red blood cells by combining two properties gleaned from aligned, noncoding genome sequences: a positive regulatory potential (RP) score, which detects similarity to patterns in alignments distinctive for regulatory regions, and conservation of a binding site motif for the essential erythroid transcription factor GATA-1. Within eight target loci, we tested 75 noncoding segments by reporter gene assays in transiently transfected human K562 cells and/or after site-directed integration into murine erythroleukemia cells. Segments with a high RP score and a conserved exact match to the binding site consensus are validated at a good rate (50%-100%, with rates increasing at higher RP), whereas segments with lower RP scores or nonconsensus binding motifs tend to be inactive. Active DNA segments were shown to be occupied by GATA-1 protein by chromatin immunoprecipitation, whereas sites predicted to be inactive were not occupied. We verify four previously known erythroid CRMs and identify 28 novel ones. Thus, high RP in combination with another feature of a CRM, such as a conserved transcription factor binding site, is a good predictor of functional CRMs. Genome-wide predictions based on RP and a large set of well-defined transcription factor binding sites are available through servers at http://www.bx.psu.edu/.

Amino Acid Motifs↗

FoldMiner: structural motif discovery using an improved superposition algorithm.

We report an unsupervised structural motif discovery algorithm, FoldMiner, which is able to detect global and local motifs in a database of proteins without the need for multiple structure or sequence alignments and without relying on prior classification of proteins into families. Motifs, which are discovered from pairwise superpositions of a query structure to a database of targets, are described probabilistically in terms of the conservation of each secondary structure element's position and are used to improve detection of distant structural relationships. During each iteration of the algorithm, the motif is defined from the current set of homologs and is used both to recruit additional homologous structures and to discard false positives. FoldMiner thus achieves high specificity and sensitivity by distinguishing between homologous and nonhomologous structures by the regions of the query to which they align. We find that when two proteins of the same fold are aligned, highly conserved secondary structure elements in one protein tend to align to highly conserved elements in the second protein, suggesting that FoldMiner consistently identifies the same motif in members of a fold. Structural alignments are performed by an improved superposition algorithm, LOCK 2, which detects distant structural relationships by placing increased emphasis on the alignment of secondary structure elements. LOCK 2 obeys several properties essential in automated analysis of protein structure: It is symmetric, its alignments of secondary structure elements are transitive, its alignments of residues display a high degree of transitivity, and its scoring system is empirically found to behave as a metric.

Algorithms↗

Optimization of the catalytic properties of Aspergillus fumigatus phytase based on the three-dimensional structure.

Previously, we determined the DNA and amino acid sequences as well as biochemical and biophysical properties of a series of fungal phytases. The amino acid sequences displayed 49-68% identity between species, and the catalytic properties differed widely in terms of specific activity, substrate specificity, and pH optima. With the ultimate goal to combine the most favorable properties of all phytases in a single protein, we attempted, in the present investigation, to increase the specific activity of Aspergillus fumigatus phytase. The crystal structure of Aspergillus niger NRRL 3135 phytase known at 2.5 A resolution served to specify all active site residues. A multiple amino acid sequence alignment was then used to identify nonconserved active site residues that might correlate with a given favorable property of interest. Using this approach, Gln27 of A. fumigatus phytase (amino acid numbering according to A. niger phytase) was identified as likely to be involved in substrate binding and/or release and, possibly, to be responsible for the considerably lower specific activity (26.5 vs. 196 U x [mg protein](-1) at pH 5.0) of A. fumigatus phytase when compared to Aspergillus terreus phytase, which has a Leu at the equivalent position. Site-directed mutagenesis of Gln27 of A. fumigatus phytase to Leu in fact increased the specific activity to 92.1 U x (mg protein)(-1), and this and other mutations at position 27 yielded an interesting array of pH activity profiles and substrate specificities. Analysis of computer models of enzyme-substrate complexes suggested that Gln27 of wild-type A. fumigatus phytase forms a hydrogen bond with the 6-phosphate group of myo-inositol hexakisphosphate, which is weakened or lost with the amino acid substitutions tested. If this hydrogen bond were indeed responsible for the differences in specific activity, this would suggest product release as the rate-limiting step of the A. fumigatus wild-type phytase reaction.

6-Phytase↗

Protein engineering studies of dichloromethane dehalogenase/glutathione S-transferase from Methylophilus sp. strain DM11. Ser12 but not Tyr6 is required for enzyme activity.

The structural gene for dichloromethane dehalogenase/glutathione S-transferase (GST, EC 2.5.1.18) from Methylophilus sp. strain DM11 was subcloned into a multicopy plasmid under the control of the T7 polymerase promoter, allowing expression in Escherichia coli and easy purification of the enzyme in good yield. Several point mutations leading to amino acid changes at residues Tyr6, His8 and Ser12 of the protein were introduced in this gene. Mutations at Tyr6, the N-terminal tyrosine known to be essential for enzymatic activity in glutathione S-transferases of the alpha, mu, and pi classes, had little effect on the activity of dichloromethane dehalogenase. The same applied for mutations at residue His8, which from multiple alignments of GST sequences may also correspond to the conserved N-terminal tyrosine residue of GST enzymes. The higher turnover rate of the wild-type enzyme with dibromomethane compared with dichloromethane was lost in mutants with amino acid replacements at residue His8, but retained in mutant proteins at Tyr6. Mutations at Ser12 led to mutants with drastically reduced enzymatic activity, pinpointing this residue as an essential determinant of catalytic efficiency.

Amino Acid Sequence↗

Genetic characterization of 2 novel feline caliciviruses isolated from cats with idiopathic lower urinary tract disease.

Feline caliciviruses (FCVs) are potential etiologic agents in feline idiopathic lower urinary tract disease (I-LUTD). By means of a modified virus isolation method, we examined urine obtained from 28 male and female cats with nonobstructive I-LUTD, 12 male cats with obstructive I-LUTD, and 18 clinically healthy male and female cats. All cats had been routinely vaccinated for FCV. Two FCVs were isolated; I (FCV-U1) from a female cat with nonobstructive I-LUTD, and another (FCV-U2) from a male cat with obstructive I-LUTD. To determine the genetic relationship of FCV-U1 and FCV-U2 to other FCVs. capsid protein gene RNA was reverse transcribed into cDNA, amplified, and sequenced. Multiple amino acid sequence alignments and phylogenetic trees were constructed for the entire capsid protein, hypervariable region E, and the more conserved (nonhypervariable) regions A, B, D, and F. When compared to 23 other FCV isolates with known biotypes, the overall amino acid sequence identity of the capsid protein of FCV-U1 and FCV-U2 ranged from 83 to 96%; identity of hypervariable regions C and E ranged from 58 to 85%. Phylogenetically, FCV-U1 clearly separated from other FCV strains in phenograms based on nonhypervariable regions. In contrast, FCV-U2 consistently segregated with the Urbana strain in all phenograms. Clustering of isolates by geographic origin was most apparent in phenograms based on nonhypervariable regions. No clustering of isolates by biotype was apparent in any phenograms. Our results indicate that FCV-UI and FCV-U2 are genetically distinct from other known vaccine and field strains of FCV.

Animals↗

Molecular cloning and expression of adenosine kinase from Leishmania donovani: identification of unconventional P-loop motif.

The unique catalytic characteristics of adenosine kinase (Adk) and its stage-specific differential activity pattern have made this enzyme a prospective target for chemotherapeutic manipulation in the purine-auxotrophic parasitic protozoan Leishmania donovani. However, nothing is known about the structure of the parasite Adk. We report here the cloning of its gene and the characterization of the gene product. The encoded protein, consisting of 345 amino acid residues with a calculated molecular mass of 37173 Da, shares limited but significant similarity with sugar kinases and inosine-guanosine kinase of microbial origin, supporting the notion that these enzymes might have the same ancestral origin. The identity of the parasite enzyme with the corresponding enzyme from two other sources so far described was only 40%. Furthermore, 5' RNA mapping studies indicated that the Adk gene transcript is matured post-transcriptionally with the trans-splicing of the mini-exon (spliced leader) occurring at nt -160 from the predicted translation initiation site. The biochemical properties of the recombinant enzyme were similar to those of the enzyme isolated from leishmanial cells. The intrinsic tryptophan fluorescence of the enzyme was substrate-sensitive. On the basis of a multiple protein-alignment sequence comparison and ATP-induced fluorescence quenching in the presence or the absence of KI and acrylamide, the docking site for ATP has been provisionally identified and shown to have marked divergence from the consensus P-loop motif reported for ATP- or GTP-binding proteins from other sources.

Acrylamide↗

Mutation analysis of exon 9 of the LDL receptor gene in Thai subjects with primary hypercholesterolemia.

The low density lipoprotein (LDL) receptor plays an important role in cholesterol homeostasis. A mutation in this gene causes an autosomal codominant disorder, namely familial hypercholesterolemia (FH). In this study, single strand conformation polymorphism (SSCP) analysis was used to screen for mutations in exon 9 of the LDL receptor gene in a group of 45 Thai patients (11 males and 34 females) with primary hypercholesterolemia. The peptide encoded by exon 9 belongs to the epidermal growth factor (EGF) precursor homology domain which is highly conserved in the LDL receptor protein. An abnormal SSCP pattern was observed in one female patient. The same screening strategy was also used to screen DNA samples from 33 normolipidemic subjects. All of these samples showed normal SSCP pattern. By direct DNA sequencing, the underlying mutation in the DNA with abnormal SSCP pattern was identified. The index subject was heterozygous for a T to C transition at nucleotide 1235. This transition would cause a nonconservative substitution of a nonpolar side chain amino acid "methionine" at codon 391, with an uncharged polar side chain amino acid "threonine", note M391T. From multiple amino acid sequence alignment in six species, the amino acid at codon 391 and the others nearby are completely conserved. Such nonconservative substitution of an amino acid residue in a highly conserved region could consequently result in a functional and/or structural defect in the receptor protein. In conclusion, we propose that M391T is likely to be the cause of hypercholesterolemia in this index subject.

DNA Mutational Analysis↗

[Development and preparation of recombinant gD antigen of the herpes simplex type 1 (HSV-1) virus].

The most potent antigen among HSV-1 proteins are glycoproteins gB(UL27) and gD(US6). Multiple amino acid sequence alignment of these proteins shows that gD protein is the most specific for HSV-1. Analysis of gD protein epitopes detected the main antigenic determinants not cross-reactive with antigens of other viruses. Virus was isolated and genome DNA was prepared from morphological elements of a patient with herpes simplex infection. US6 gene fragment was cloned in pUC19 vector. Cloning in bacterial expression vectors helped obtain beta-galactosidase-fused recombinant HSV-1 gD protein with 6-histidines affine target for high-performance chromatography purification. ELISA with a set of HSV-1-positive and negative donor sera and a commercial panel of HSV-1 sera (Vektor-Best) showed that recombinant gD can be used as an antigen to HSV-1-specific IgG.

Amino Acid Sequence↗

[Molecular screening of MC4R gene and association with fat traits in pig resource family].

Melanocortin-4 Receptor (MC4R) plays an important role in the regulation of human obesity. It can cooperate with leptin, neuropeptide Y(NPY) and melanocyte-stimulating hormone (MSH) to regulate body weight and feeding. Inactivation of this receptor by gene targeting in mice results in a maturity onset obesity syndrome associated with hyperphagic, hyperinsulinemia, hyperglycemia, as well as decreased linear growth and adult obesity. Multiple alignments of the sequences from individuals of several pig lines identified a single nucleotide substitution(G-->A) at position 298 of the seventh transmembrane domain. In present study, polymorphism distribution of MC4R gene fragment in resource population was studied using PCR-RFLP method based on the enzyme Taq I. The genotype was analyzed with the phenotype of the slaughtered individuals. The results showed that the frequencies of MC4R genotype varied in different breeds. The correlation analysis demonstrated the genotype of MC4R was in significant relation with back-fat thickness on thorax-waist, buttock and the average back-fat thickness, as well as with the width and area of longissmus dorsi (LD), and the percentage of skin. MC4R gene plays a role mainly in the pattern of dominant effect, and all the additive effects were not significant.

Animals↗

Molecular dissection of the beta subunit of F1-ATPase into peptide fragments.

Partial digestion of the native beta subunit of F1-ATPase from the thermophilic Bacillus strain PS3 by three different proteases produced a limited number of peptide fragments. In most cases, the peptides remained associated, and the gross structure of the beta subunit was not destroyed. Furthermore, most peptides were able to reassociate into the form of the beta subunit after denaturating urea treatment. Therefore, the cleaved sites are most likely located in water-exposed loop regions in the tertiary structure of the protein. Almost all peptides were analyzed, and 17 cleaved sites were determined. From the analysis of the distribution of cleaved sites and deletions or insertions in the multiple amino acid sequence alignment of proteins homologous to the beta subunit, locations of five loops and four candidate loops in the beta subunit are suggested. There are two large loops in the central region of the beta subunit sequence, and dicyclohexylcarbodiimide-reactive Glu190 is located in one of them. Tyr341, involved in putative catalytic ATP binding, is also found in one of the loops. Then, taking cleaved sites as a reference, two kinds of expression plasmids, each of which carried genes of two complementary peptide fragments, 1-193 and 198-473 or 1-284 and 285-473, were constructed and expressed in Escherichia coli. For each plasmid, two peptides were coexpressed, associated into a stable beta subunit form in E. coli cells, and purified without dissociation. When these beta subunits were denatured by urea and applied to polyacrylamide gel without denaturant, a protein band with the same mobility as that of the beta subunit appeared, indicating that reassociation of peptide fragments into the form of the beta subunit occurred upon removal of urea. These beta subunits retained the ability to reconstitute the alpha 3 beta 3 gamma complexes even though the efficiency of reconstitution and the recovered ATPase activities were decreased. These complexes were stable at high or low temperature, and ATPase activities were sensitive to inhibition by N3-.

Amino Acid Sequence↗

[Use of structural MNA descriptors for designing profiles of protein families].

A new approach to constructing the profiles of protein families is proposed, which uses only structural similarity of amino acid residues. We derived multiple alignments of protein sequences from 3D superpositions of the protein structures and constructed protein family profiles using structural molecular MNA descriptors. MNA (Multilevel Neighborhoods of Atoms) descriptors were developed earlier and are successfully applied for predicting the biological activity in drug-like compounds. In our approach, each aligned position was described by a set of MNA descriptors calculated for each amino acid residue in the alignment column. In this study, we constructed MNA profiles for trypsin, subtilase, and cytochrome P450 protein families and scanned SWISSPROT with some fragments of these profiles. We also calculated the Independence Accuracy of Prediction for each profile fragment. It was shown that the approach developed could be applied to predict protein function.

Amino Acid Sequence↗

[Cloning and characterization of a full-length HIV-1 genome of a prevalent subtype B-Thai strain in Henan Province].

OBJECTIVE: To clone, identify and phylogenetically characterize a clade B-Thai HIV isolate representing the most prevalent virus in Henan province. METHODS: Peripheral blood mononuclear cells (PBMCs) from an HIV-1 infected patient in Henan Province were separated, and co-cultivated with phytohemagglutinin-stimulated healthy donor PBMCs. Proviral DNA was extracted from productively infected PBMCs. The full-length HIV-1 genome was amplified by using the LA Tag long template PCR system. Primers were positioned in conserved regions within the HIV-1 long terminal repeats. Purified PCR products were T-A ligated into a pWSK29-T vector(CNHN 24 clone). Three recombinant clones containing virtually full-length HIV-1 genome were identified by PCR. The full-length genome was sequenced by using the primer-walking approach. Nucleotide sequence similarities were calculated by the local-homology algorithm. Phylogenetic trees of gag, pol and env reading frames were constructed using the Phylip software. RESULTS: HIV-1 C3V4 sequences indicate that the epidemic in this area was B-Thai subtype. V3 loop multiple amino acid sequence alignments showed amino acid alterations at nine positions. The 9,010 bp genomic sequence derived from isolate CNHN 24 contained all known structural and regulatory genes of an HIV-1 genome. No major deletions, insertions, or rearrangements were found. The highest homologies of the gag, pol, vpr, and vif reading frames to the corresponding clade B-Thai RL 42 sequences were 95.42%-97.08%. Phylogenetic trees showed the closest relationship of CNHN 24 and RL 42. CONCLUSION: The cloning and characterization of a virtually full-length HIV-1 B-Thai subtype in central China was completed in our laboratory. The data should be helpful to future studies on the genetic diversity of HIV-1.

Amino Acid Sequence↗

PCR-based detection of Mycoplasma species.

In this study, we describe our newly-developed sensitive two-stage PCR procedure for the detection of 13 common mycoplasmal contaminants (M. arthritidis, M. bovis, M. fermentans, M. genitalium, M. hominis, M. hyorhinis, M. neurolyticum, M. orale, M. pirum, M. pneumoniae, M. pulmonis, M. salivarium, U. urealyticum). For primary amplification, the DNA regions encompassing the 16S and 23S rRNA genes of 13 species were targeted using general mycoplasma primers. The primary PCR products were then subjected to secondary nested PCR, using two different primer pair sets, designed via the multiple alignment of nucleotide sequences obtained from the 13 mycoplasmal species. The nested PCR, which generated DNA fragments of 165-353 bp, was found to be able to detect 1-2 copies of the target DNA, and evidenced no cross-reactivity with the genomic DNA of related microorganisms or of human cell lines, thereby confirming the sensitivity and specificity of the primers used. The identification of contaminated species was achieved via the performance of restriction fragment length polymorphism (RFLP) coupled with Sau3AI digestion. The results obtained in this study furnish evidence suggesting that the employed assay system constitutes an effective tool for the diagnosis of mycoplasmal contamination in cell culture systems.

Animals↗

Porcine aromatases: studies on tissue-specific, functionally distinct isozymes from a single gene?

Aromatase cytochrome P450 (P450arom) is expressed in a variety of tissues. Pigs express P450arom as bilaminar blastocysts in utero, and thereafter in the gonads, adrenal glands and placenta. Our studies also demonstrate the existence of porcine isozymes of P450arom which differ substantially in their amino acid composition and function. The placental isoform, most similar to P450arom in other mammals, consists of 503 amino acids. The ovarian isoform, expressed in both theca and granulosa cells, is a 501 amino acid protein exhibiting less than 20% of the activity of the placental isozyme. Furthermore, it is inhibited not only by CGS16949A but also by etomidate which does not inhibit the placental P450arom. Partial sequences generated by the rapid amplification of the cDNA ends (RACE) procedure indicate that the expression of a third isoform in the blastocyst is switched to the placental isozyme during differentiation of the fetal membranes. In addition, these transcripts, and others from the theca, granulosa, testes, adrenal glands and placenta demonstrate differences in the 5'-untranslated region (putative exon I) suggestive of tissue-specific alternative splicing. An identical 5'-untranslated sequence was obtained from transcripts expressed in the theca and granulosa. Testes and adrenal transcripts also have identical 5' ends, which differ substantially from the ovarian sequence. Blastocyst and placenta 5'-untranslated sequences differ from each other and from those expressed in the gonads and adrenals. Several tissue-specific transcripts thus encode porcine P450arom. Interestingly, distinct 5' sequences exist for ovarian and testes P450arom mRNAs, suggesting different promoters and therefore regulation in the male and female gonads. The molecular origins of the functional isoforms and the tissue-specific transcripts are uncertain, however partial genomic sequence and other genetic analyses suggest the existence of multiple genes. However, sequence alignment of the placental and ovarian isoforms indicates complete conservation of putative exon III, so that complex splicing remains a possibility. Clearly, the regulation of P450arom expression is more complex in the pig than in other vertebrates investigated to date.

Amino Acid Sequence↗

Malate dehydrogenase: distribution, function and properties.

Malate dehydrogenase (MDH) (EC 1.1.1.37) catalyzes the conversion of oxaloacetate and malate. This reaction is important in cellular metabolism, and it is coupled with easily detectable cofactor oxidation/reduction. It is a rather ubiquitous enzyme, for which several isoforms have been identified, differing in their subcellular localization and their specificity for the cofactor NAD or NADP. The nucleotide binding characteristics can be altered by a single amino acid change. Multiple amino acid sequence alignments of MDH show that there is a low degree of primary structural similarity, apart from several positions crucial for catalysis, cofactor binding and the subunit interface. Despite the low amino acids sequence identity their 3-dimensional structures are very similar. MDH is a group of multimeric enzymes consisting of identical subunits usually organized as either dimer or tetramers with subunit molecular weights of 30-35 kDa. MDH has been isolated from different sources including archaea, eubacteria, fungi, plant and mammals.

Animals↗