Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Sequence variation of the glycoprotein gene identifies three distinct lineages within field isolates of viral haemorrhagic septicaemia virus, a fish rhabdovirus.

To evaluate the genetic diversity of viral haemorrhagic septicaemia virus (VHSV), the sequence of the glycoprotein genes (G) of 11 North American and European isolates were determined. Comparison with the G protein of representative members of the family Rhabdoviridae suggested that VHSV was a different virus species from infectious haemorrhagic necrosis virus (IHNV) and Hirame rhabdovirus (HIRRV). At a higher taxonomic level, VHSV, IHNV and HIRRV formed a group which was genetically closest to the genus Lyssavirus. Compared with each other, the G genes of VHSV displayed a dissimilar overall genetic diversity which correlated with differences in geographical origin. The multiple sequence alignment of the complete G protein, showed that the divergent positions were not uniformly distributed along the sequence. A central region (amino acid position 245-300) accumulated substitutions and appeared to be highly variable. The genetic heterogeneity within a single isolate was high, with an apparent internal mutation frequency of 1.2 x 10(-3) per nucleotide site, attesting the quasispecies nature of the viral population. The phylogeny separated VHSV strains according to the major geographical area of isolation: genotype I for continental Europe, genotype II for the British Isles, and genotype III for North America. Isolates from continental Europe exhibited the highest genetic variability, with sub-groups correlated partially with the serological classification. Neither neutralizing polyclonal sera, nor monoclonal antibodies, were able to discriminate between the genotypes. The overall structure of the phylogenetic tree suggests that VHSV genetic diversity and evolution fit within the model of random change and positive selection operating on quasispecies.

Amino Acid Sequence↗

Antigenic and molecular analyses of the variability of bovine respiratory syncytial virus G glycoprotein.

Antigenic variation among eight bovine respiratory syncytial virus (BRSV) isolates was determined using monoclonal antibodies (MAbs) specific for the attachment (G) protein. Two major (and one intermediate) subgroups were identified, as well as one strain that did not fit any pattern. The subgroups could also be differentiated on the basis of the Mr of the F protein cleavage product, F2. The nucleotide sequence of the G gene of seven of the BRSV strains was determined and compared with published G gene sequences. Subgroups A and A/B were more closely related in protein sequence than subgroups A and B or subgroups A/B and B. These results could not be correlated with those obtained by the determination of the Mr of the F2 polypeptide. Multiple sequence alignments showed a high level of amino acid identity at the inter-subgroup level (85% identity between subgroup A and subgroup B strains), similar to the intra-subgroup human (H)RSV identity, suggesting that the BRSV isolates form a continuum rather than distinct subgroups. However, unusual variability was observed within the immunodominant domain (amino acids 174-188) in contrast with the situation in HRSV strains belonging to the same subgroup.

Amino Acid Sequence↗

Analysis of the sequence diversity of the P1, HC, P3, NIb and CP genomic regions of several yam mosaic potyvirus isolates: implications for the intraspecies molecular diversity of potyviruses.

Partial sequences from serologically characterized yam mosaic potyvirus (YMV) isolates were determined in conserved (helper-component proteinase, HC; nuclear inclusion b, NIb) and variable (first protein, P1; third protein, P3; and coat protein, CP) regions of the potyviral genome in order to investigate the intraspecies molecular diversity of YMV. Multiple sequence alignments and pairwise comparisons were used to quantify the sequence polymorphism in these regions. Two levels of diversity were observed among YMV isolates: above 90% nucleotide (nt) sequence identities were found between YMV isolates of the same group (intragroup) regardless of the region considered, whereas identities between isolates from different groups (intergroup) were lower and depended upon the protein chosen. For instance, the average intergroup nt sequence identity between YMV isolates was about 65% in the P1 protein and the N terminus of the CP while there was more than 80% nt identity in the HC, P3 and NIb proteins. Thus P3 appeared to be conserved between YMV isolates even though this region was variable between potyvirus species. Similar analysis of the intraspecies molecular diversity of other potyviruses (potato virus Y, zucchini yellow mosaic virus, plum pox virus, pea seed-borne mosaic virus) led to the same results: (i) two levels of intraspecies molecular diversity were found (intragroup and intergroup); (ii) intraspecies molecular diversity differed from interspecies molecular diversity in the P3, P1 and N-terminal regions.

Amino Acid Sequence↗

Molecular epidemiology of rabbit haemorrhagic disease virus outbreaks in France during 1988 to 1995.

In order to evaluate genetic variation between rabbit haemorrhagic disease virus (RHDV) isolates and to derive phylogenetic relationships, 56 virus isolates collected from various parts of France over a 7 year period (1988 to 1995) were examined. Analyses were carried out by direct nucleotide sequencing of PCR fragments of three genomic regions encoding the capsid protein (VP60) (regions A and B) and a non-structural protein (region C). Multiple sequence alignments revealed maximum nucleotide divergence of 7.6, 9.4 and 8.7% for regions A, B and C, respectively, indicating a high level of conservation between isolates. Irrespective of the genomic region analysed, phylogenetic analyses carried out using various methods allowed the identification of three genogroups; distribution of isolates within these genogroups appears to be more related to the year of their collection than to their geographical origin. The possible evolution of RHDV is discussed.

Animals↗

Taxonomic characteristics of fijiviruses based on nucleotide sequences of the oat sterile dwarf virus genome.

Sequence determination of full-length cDNA clones of genomic segments 7-10 (S7-S10) of oat sterile dwarf fijivirus (OSDV) revealed that the 5' and 3' ends of the plus strands of these segments had the same conserved terminal sequences, 5' AACGAAAAA and UUUUUUUAGUC 3'. These sequences are similar, but not identical, to the conserved terminal nucleotide sequences of the genomic segments of rice black streaked dwarf fijivirus (RBSDV) and maize rough dwarf fijivirus (MRDV). The coding strands of S7 and S10 each contained two large nonoverlapping open reading frames (ORFs), as do RBSDV S7 and S9, MRDV S6 and S8 and Nilaparvata lugens reovirus (NLRV; a putative member of Fijivirus) S9. These results strongly suggest that the dicistronic nature of certain genomic segments is characteristic of fijiviruses. Computer analyses revealed sequence homology between RBSDV S7 ORF2, MRDV S6 ORF2 and OSDV S7 ORF2, suggesting that this protein is conserved among plant fijiviruses. No counterparts were found in the genome of NLRV, which is a nonphytopathogenic insect reovirus. Furthermore, phylogenetic trees derived from multiple sequence alignments of each of the homologous proteins from OSDV, RBSDV, MRDV and NLRV suggest that NLRV did not evolve from either Fijivirus group 2 (RBSDV and MRDV) or group 3 (OSDV).

Amino Acid Sequence↗

An acyl-coenzyme A carboxylase encoding gene associated with jadomycin biosynthesis in Streptomyces venezuelae ISP5230.

Analysis of a region of chromosomal DNA lying between jadR1 and jadI in the gene cluster for jadomycin biosynthesis in Streptomyces venezuelae ISP5230 detected an ORF encoding 584 amino acids similar in sequence to the biotin carboxylase (BC) and biotin carboxyl carrier protein (BCCP) components of acyl-coenzyme A carboxylases. Multiple sequence alignments of the deduced Jad protein with acyl-coenzyme A carboxylases from various sources located the BC and BCCP components in the N- and C-terminal regions, respectively, of the deduced polypeptides. The organization and amino acid sequence of the deduced polypeptide most closely resembled those in other Gram-positive bacteria broadly classified as actinomycetes. Disrupting the gene, designated jadJ, severely reduced but did not eliminate jadomycin production. The disruption had no effect on growth or morphology of the organism, implying that the product of jadJ is not essential for fatty acid biosynthesis. It is concluded that jadJ supplies malonyl-coenzyme A for biosynthesis of the polyketide intermediate that is eventually processed to form the antibiotic jadomycin B.

Amino Acid Sequence↗

Rapid detection of polyhydroxyalkanoate-accumulating bacteria isolated from the environment by colony PCR.

Colony PCR and semi-nested PCR techniques were employed for screening polyhydroxyalkanoate (PHA) producers isolated from the environment. Three degenerate primers were designed based on multiple sequence alignment results and were used as PCR primers to detect PHA synthase genes. Optimized colony PCR conditions were achieved by adding 3% DMSO combined with 1 M betaine to the reaction mixture. The sensitivity limit of the colony PCR was 1x 10(5) viable cells for Ralstonia eutropha. Nineteen PHA-positive bacteria were used to evaluate this PCR protocol; fifteen of the nineteen could be detected by colony PCR, and the other four could be detected by applying semi-nested PCR detection following colony PCR. In a preliminary screening project, 38 PHA-positive strains were isolated from environmental samples by applying the PCR protocol, and their phenotype was further confirmed by Nile blue A staining assay. By combining the colony PCR and semi-nested PCR techniques, a rapid, reliable and highly accurate detection method has been developed for detecting PHA producers. This protocol is suitable for screening large numbers of environmental isolates. The PHA accumulation ability of well-separated colonies isolated from environmental samples can be directly validated by PCR with no further culturing or chromosomal DNA extraction procedures. In addition to its application to the screening of wild-type isolates, the individual PCR-amplified product is also suitable as a specific probe for PHA operon cloning. The results suggest that the application of this PCR protocol for rapid detection of PHA producers from the environment is plausible.

Acyltransferases↗

RNA polymerase beta-subunit-based phylogeny of Ehrlichia spp., Anaplasma spp., Neorickettsia spp. and Wolbachia pipientis.

Sequence analysis of rpoB, the gene encoding the beta-subunit of RNA polymerase, was used in a phylogenetic investigation of nine species from the genera Ehrlichia, Neorickettsia, Wolbachia and Anaplasma. The complete nucleotide sequences obtained for Anaplasma phagocytophilum (HGE agent), Ehrlichia chaffeensis, Neorickettsia sennetsu, Neorickettsia risticii, Anaplasma marginale and Wolbachia pipientis were amongst the longest rpoB sequences in GenBank and ranged from 4074 bp for N. sennetsu to 4311 bp for W. pipientis. Additional partial rpoB sequences were obtained for Ehrlichia canis, Ehrlichia ruminantium and Ehrlichia muris. Identical phylogenetic trees were inferred from multiple sequence alignments of the nucleotide sequences and the derived amino acid sequences using either distance, maximum-likelihood or parsimony methods. This study confirms the phylogeny previously inferred from sequence analyses of the 16S rRNA gene, groESL and gltA and allows the confirmation of four monophyletic clades. The rpoB nucleotide sequences were more variable than the 16S rRNA gene and groESL sequences at the species level.

Anaplasma↗

Only one catalase, katG, is detectable in Rhizobium etli, and is encoded along with the regulator OxyR on a plasmid replicon.

The plasmid-borne Rhizobium etli katG gene encodes a dual-function catalase-peroxidase (KatG) (EC 1.11.1.7) that is inducible and heat-labile. In contrast to other rhizobia, katG was shown to be solely responsible for catalase and peroxidase activity in R. etli. An R. etli mutant that did not express catalase activity exhibited increased sensitivity to hydrogen peroxide (H(2)O(2)). Pre-exposure to a sublethal concentration of H(2)O(2) allowed R. etli to adapt and survive subsequent exposure to higher concentrations of H(2)O(2). Based on a multiple sequence alignment with other catalase-peroxidases, it was found that the catalytic domains of the R. etli KatG protein had three large insertions, two of which were typical of KatG proteins. Like the katG gene of Escherichia coli, the R. etli katG gene was induced by H(2)O(2) and was important in sustaining the exponential growth rate. In R. etli, KatG catalase-peroxidase activity is induced eightfold in minimal medium during stationary phase. It was shown that KatG catalase-peroxidase is not essential for nodulation and nitrogen fixation in symbiosis with Phaseolus vulgaris, although bacteroid proteome analysis indicated an alternative compensatory mechanism for the oxidative protection of R. etli in symbiosis. Next to, and divergently transcribed from the catalase promoter, an ORF encoding the regulator OxyR was found; this is the first plasmid-encoded oxyR gene described so far. Additionally, the katG promoter region contained sequence motifs characteristic of OxyR binding sites, suggesting a possible regulatory mechanism for katG expression.

Amino Acid Sequence↗

PorH, a new channel-forming protein present in the cell wall of Corynebacterium efficiens and Corynebacterium callunae.

Corynebacterium callunae and Corynebacterium efficiens are close relatives of the glutamate-producing mycolata species Corynebacterium glutamicum. The properties of the pore-forming proteins, extracted by organic solvents, were studied. The cell extracts contained channel-forming proteins that formed ion-permeable channels with a single-channel conductance of about 2 to 3 nS in 1 M KCl in a lipid bilayer assay. The corresponding proteins from both corynebacteria were purified to homogeneity and were named PorH(C.call) and PorH(C.eff). Electrophysiological studies of the channels suggested that they are wide and water-filled. Channels formed by PorH(C.call) are cation-selective, whereas PorH(C.eff) forms slightly anion-selective channels. Both proteins were partially sequenced. A multiple sequence alignment search within the known chromosome of C. efficiens demonstrated that it contains a gene that fits the partial amino acid sequence of PorH(C.eff). PorH(C.call) shows high homology to PorH(C.eff). PorH(C.eff) is encoded in the bacterial chromosome by a gene that is localized within the vicinity of the porA gene of C. efficiens. PorH(C.eff) has no signal sequence at the N terminus, which means that it is not exported by the Sec-secretion pathway. The structure of PorH in the cell wall of the corynebacteria is discussed.

Bacterial Outer Membrane Proteins↗

Genomic sequence analysis of Fugu rubripes CFTR and flanking genes in a 60 kb region conserving synteny with 800 kb of human chromosome 7.

To define control elements that regulate tissue-specific expression of the cystic fibrosis transmembrane regulator (CFTR), we have sequenced 60 kb of genomic DNA from the puffer fish Fugu rubripes (Fugu) that includes the CFTR gene. This region of the Fugu genome shows conservation of synteny with 800-kb sequence of the human genome encompassing the WNT2, CFTR, Z43555, and CBP90 genes. Additionally, the genomic structure of each gene is conserved. In a multiple sequence alignment of human, mouse, and Fugu, the putative WNT2 promoter sequence is shown to contain highly conserved elements that may be transcription factor or other regulatory binding sites. We have found two putative ankyrin repeat-containing genes that flank the CFTR gene. Overall sequence analysis suggests conservation of intron/exon boundaries between Fugu and human CFTR and revealed extensive homology between functional protein domains. However, the immediate 5' regions of human and Fugu CFTR are highly divergent with few conserved sequences apart from those resembling diminished cAMP response elements (CRE) and CAAT box elements. Interestingly, the polymorphic polyT tract located upstream of exon 9 is present in human and Fugu but absent in mouse. Similarly, an intron 1 and intron 9 element common to human and Fugu is absent in mouse. The euryhaline killifish CFTR coding sequence is highly homologous to the Fugu sequence, suggesting that upregulation of CFTR in that species in response to salinity may be regulated transcriptionally.

Amino Acid Sequence↗

Quantitative estimates of sequence divergence for comparative analyses of mammalian genomes.

Comparative sequence analyses on a collection of carefully chosen mammalian genomes could facilitate identification of functional elements within the human genome and allow quantification of evolutionary constraint at the single nucleotide level. High-resolution quantification would be informative for determining the distribution of important positions within functional elements and for evaluating the relative importance of nucleotide sites that carry single nucleotide polymorphisms (SNPs). Because the level of resolution in comparative sequence analyses is a direct function of sequence diversity, we propose that the information content of a candidate mammalian genome be defined as the sequence divergence it would add relative to already-sequenced genomes. We show that reliable estimates of genomic sequence divergence can be obtained from small genomic regions. On the basis of a multiple sequence alignment of approximately 1.4 megabases each from eight mammals, we generate such estimates for five unsequenced mammals. Estimates of the neutral divergence in these data suggest that a small number of diverse mammalian genomes in addition to human, mouse, and rat would allow single nucleotide resolution in comparative sequence analyses.

Animals↗

Exploration of novel motifs derived from mouse cDNA sequences.

We performed a systematic maximum density subgraph (MDS) detection of conserved sequence regions to discover new, biologically relevant motifs from a set of 21,050 conceptually translated mouse cDNA (FANTOM1) sequences. A total of 3202 candidate sequences, which shared similar regions over >20 amino acid residues, were screened against known conserved regions listed in Pfam, ProDom, and InterPro. The filtering procedure resulted in 139 FANTOM1 sequences belonging to 49 new motif candidates. Using annotations and multiple sequence alignment information, we removed by visual inspection 42 candidates whose members were found to be false positives because of sequence redundancy, alternative splicing, low complexity, transcribed retroviral repeat elements contained in the region of the predicted open reading frame, and reports in the literature. The remaining seven motifs have been expanded by hidden Markov model (HMM) profile searches of SWISS-PROT/TrEMBL from 28 FANTOM1 sequences to 164 members and analyzed in detail on sequence and structure level to elucidate the possible functions of motifs and members. The novel and conserved motif MDS00105 is specific for the mammalian inhibitor of growth (ING) family. Three submotifs MDS00105.1-3 are specific for ING1/ING1L, ING1-homolog, and ING3 subfamilies. The motif MDS00105 together with a PHD finger domain constitutes a module for ING proteins. Structural motif MDS00113 represents a leucine zipper-like motif. Conserved motif MDS00145 is a novel 1-acyl-SN-glycerol-3-phosphate acyltransferase (AGPAT) submotif containing a transmembrane domain that distinguishes AGPAT3 and AGPAT4 from all other acyltransferase domain-containing proteins. Functional motif MDS00148 overlaps with the kazal-type serine protease inhibitor domain but has been detected only in an extracellular loop region of solute carrier 21 (SLC21) (organic anion transporters) family members, which may regulate the specificity of anion uptake. Our motif discovery not only aided in the functional characterization of new mouse orthologs for potential drug targets but also allowed us to predict that at least 16 other new motifs are waiting to be discovered from the current SWISS-PROT/TrEMBL database.

1-Acylglycerol-3-Phosphate O-Acyltransferase↗

PANTHER: a library of protein families and subfamilies indexed by function.

In the genomic era, one of the fundamental goals is to characterize the function of proteins on a large scale. We describe a method, PANTHER, for relating protein sequence relationships to function relationships in a robust and accurate way. PANTHER is composed of two main components: the PANTHER library (PANTHER/LIB) and the PANTHER index (PANTHER/X). PANTHER/LIB is a collection of "books," each representing a protein family as a multiple sequence alignment, a Hidden Markov Model (HMM), and a family tree. Functional divergence within the family is represented by dividing the tree into subtrees based on shared function, and by subtree HMMs. PANTHER/X is an abbreviated ontology for summarizing and navigating molecular functions and biological processes associated with the families and subfamilies. We apply PANTHER to three areas of active research. First, we report the size and sequence diversity of the families and subfamilies, characterizing the relationship between sequence divergence and functional divergence across a wide range of protein families. Second, we use the PANTHER/X ontology to give a high-level representation of gene function across the human and mouse genomes. Third, we use the family HMMs to rank missense single nucleotide polymorphisms (SNPs), on a database-wide scale, according to their likelihood of affecting protein function.

Algorithms↗

WebLogo: a sequence logo generator.

WebLogo generates sequence logos, graphical representations of the patterns within a multiple sequence alignment. Sequence logos provide a richer and more precise description of sequence similarity than consensus sequences and can rapidly reveal significant features of the alignment otherwise difficult to perceive. Each logo consists of stacks of letters, one stack for each position in the sequence. The overall height of each stack indicates the sequence conservation at that position (measured in bits), whereas the height of symbols within the stack reflects the relative frequency of the corresponding amino or nucleic acid at that position. WebLogo has been enhanced recently with additional features and options, to provide a convenient and highly configurable sequence logo generator. A command line interface and the complete, open WebLogo source code are available for local installation and customization.

Amino Acid Sequence↗

CAP3: A DNA sequence assembly program.

We describe the third generation of the CAP sequence assembly program. The CAP3 program includes a number of improvements and new features. The program has a capability to clip 5' and 3' low-quality regions of reads. It uses base quality values in computation of overlaps between reads, construction of multiple sequence alignments of reads, and generation of consensus sequences. The program also uses forward-reverse constraints to correct assembly errors and link contigs. Results of CAP3 on four BAC data sets are presented. The performance of CAP3 was compared with that of PHRAP on a number of BAC data sets. PHRAP often produces longer contigs than CAP3 whereas CAP3 often produces fewer errors in consensus sequences than PHRAP. It is easier to construct scaffolds with CAP3 than with PHRAP on low-pass data with forward-reverse constraints.

Algorithms↗

Molecular cloning and biological activity of alpha-, beta-, and gamma-megaspermin, three elicitins secreted by Phytophthora megasperma H20.

We report on the molecular cloning of the Phytophthora megasperma H20 (PmH20) glycoprotein shown previously as an inducer of the hypersensitive response, of localized acquired resistance and of systemic acquired resistance in tobacco (Nicotiana tabacum), and of the PmH20 alpha- and beta-megaspermin, two elicitins of class I-A and I-B, respectively. The structure of the glycoprotein shows a signal peptide of 20 amino acids followed by the typical elicitin 98-amino acid-long domain and a 77-amino acid-long C-terminal domain carrying an O-glycosylated moiety. The molecular mass deduced from the translated cDNA sequence is 14,920 and 18,676 D as determined by mass spectrometry. This structure together with multiple sequence alignments and phylogenetic analyses indicate that the glycoprotein belongs to class III elicitins. It is the first class III elicitin protein characterized, which we named gamma-megaspermin. We compared the biological activity of the three PmH20 elicitins when applied to tobacco cv Samsun NN plants. Although alpha- and gamma-megaspermin were similarly active, beta-megaspermin was the most active in inducing the hypersensitive response and localized acquired resistance, which was assessed by measuring the levels of acidic and basic pathogenesis-related proteins and of the antioxidant phytoalexin scopoletin. The three elicitins induced similar levels of systemic acquired resistance measured as the expression of acidic PR proteins and is increased resistance to challenge tobacco mosaic virus infection.

Algal Proteins↗

Arabidopsis proteins containing similarity to the universal stress protein domain of bacteria.

We have collected a set of 44 Arabidopsis proteins with similarity to the USPA (universal stress protein A of Escherichia coli) domain of bacteria. The USPA domain is found either in small proteins, or it makes up the N-terminal portion of a larger protein, usually a protein kinase. Phylogenetic tree analysis based upon a multiple sequence alignment of the USPA domains shows that these domains of protein kinases 1.3.1 and 1.3.2 form distinct groups, as do the protein kinases 1.4.1. This indicates that their USPA domain structures have diverged appreciably and suggests that they may subserve distinct cellular functions. Two USPA fold classes have been proposed: one based on Methanococcus jannaschii MJ0577 (1MJH) that binds ATP, and the other based on the Haemophilus influenzae universal stress protein (1JMV), highly similar to E. coli UspA, which does not bind ATP. A set of common residues involved in ATP binding in 1MJH and conserved in similar bacterial sequences is also found in a distinct cluster of Arabidopsis sequences. Threading analysis, which examines aspects of secondary and tertiary structure, confirms this Arabidopsis sequence cluster as highly similar to 1MJH. This structural approach can distinguish between the characteristic fold differences of 1MJH-like and 1JMV-like bacterial proteins and was used to assign the complete set of candidate Arabidopsis proteins to one of these fold classes. It is clear that all the plant sequences have arisen from a 1MJH-like ancestor.

Adenosine Triphosphate↗