Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Cross-packaging of human immunodeficiency virus type 1 vector RNA by spleen necrosis virus proteins: construction of a new generation of spleen necrosis virus-derived retroviral vectors.

The ability of the nonlentiviral retrovirus spleen necrosis virus (SNV) to cross-package the genomic RNA of the distantly related human immunodeficiency virus type 1 (HIV-1) and vice versa was analyzed. Such a model may allow us to further study HIV-1 replication and pathogenesis, as well as to develop safe gene therapy vectors. Our results suggest that SNV can cross-package HIV-1 genomic RNA but with lower efficiency than HIV-1 proteins. However, HIV-1-specific proteins were unable to cross-package SNV RNA. We also constructed SNV-based gag-pol chimeric variants by replacing the SNV integrase with the HIV-1 integrase, based on multiple sequence alignments and domain analyses. These analyses revealed that there are conserved domains in all retroviral integrase open reading frames (orf), despite the divergence in the primary sequences. The transcomplementation assays suggested that SNV proteins recognized one of the chimeric variants. This demonstrated that HIV-1 integrase is functional in the SNV gag-pol orf with a lower transduction efficiency, utilizing homologous (SNV) RNA, as well as the heterologous vector RNA of HIV-1. These findings suggest that homology in the conserved sequences of the integrase protein may not be fully competent in the replacement of protein(s) from one retrovirus to another, and there are likely several other factors involved in each of the steps related to replication, integration, and infection. However, further studies to dissect the gag-pol region will be critical for understanding the mechanisms involved in the cleavage of reverse transcriptase, RNase H, and integrase. These studies should provide further insight into the design and development of novel molecular approaches to block HIV-1 replication and to construct a new generation of SNV-based vectors.

Amino Acid Sequence↗

Analysis of missense variation in human BRCA1 in the context of interspecific sequence variation.

INTRODUCTION: Interpretation of results from mutation screening of tumour suppressor genes known to harbour high risk susceptibility mutations, such as APC, BRCA1, BRCA2, MLH1, MSH2, TP53, and PTEN, is becoming an increasingly important part of clinical practice. Interpretation of truncating mutations, gene rearrangements, and obvious splice junction mutations, is generally straightforward. However, classification of missense variants often presents a difficult problem. From a series of 20,000 full sequence tests of BRCA1 carried out at Myriad Genetic Laboratories, a total of 314 different missense changes and eight in-frame deletions were observed. Before this study, only 21 of these missense changes were classified as deleterious or suspected deleterious and 14 as neutral or of little clinical significance. METHODS: We have used a combination of a multiple sequence alignment of orthologous BRCA1 sequences and a measure of the chemical difference between the amino acids present at individual residues in the sequence alignment to classify missense variants and in-frame deletions detected during mutation screening of BRCA1. RESULTS: In the present analysis we were able to classify an additional 50 missense variants and two in-frame deletions as probably deleterious and 92 missense variants as probably neutral. Thus we have tentatively classified about 50% of the unclassified missense variants observed during clinical testing of BRCA1. DISCUSSION: An internal test of the analysis is consistent with our classification of the variants designated probably deleterious; however, we must stress that this classification is tentative and does not have sufficient independent confirmation to serve as a clinically applicable stand alone method.

Amino Acid Sequence↗

Comprehensive statistical study of 452 BRCA1 missense substitutions with classification of eight recurrent substitutions as neutral.

BACKGROUND: Genetic testing for hereditary cancer syndromes contributes to the medical management of patients who may be at increased risk of one or more cancers. BRCA1 and BRCA2 testing for hereditary breast and ovarian cancer is one such widely used test. However, clinical testing methods with high sensitivity for deleterious mutations in these genes also detect many unclassified variants, primarily missense substitutions. METHODS: We developed an extension of the Grantham difference, called A-GVGD, to score missense substitutions against the range of variation present at their position in a multiple sequence alignment. Combining two methods, co-occurrence of unclassified variants with clearly deleterious mutations and A-GVGD, we analysed most of the missense substitutions observed in BRCA1. RESULTS: A-GVGD was able to resolve known neutral and deleterious missense substitutions into distinct sets. Additionally, eight previously unclassified BRCA1 missense substitutions observed in trans with one or more deleterious mutations, and within the cross-species range of variation observed at their position in the protein, are now classified as neutral. DISCUSSION: The methods combined here can classify as neutral about 50% of missense substitutions that have been observed with two or more clearly deleterious mutations. Furthermore, odds ratios estimated for sets of substitutions grouped by A-GVGD scores are consistent with the hypothesis that most unclassified substitutions that are within the cross-species range of variation at their position in BRCA1 are also neutral. For most of these, clinical reclassification will require integrated application of other methods such as pooled family histories, segregation analysis, or validated functional assay.

Amino Acid Sequence↗

Identification of novel mutations in the SEMA4A gene associated with retinal degenerative diseases.

Semaphorins are a large family of transmembrane proteins. The gene for SEMA4A encodes a transmembrane protein comprising 760 amino acids. To investigate its association with human retinal degeneration, mutation screening of the SEMA4A gene was carried out on 190 unrelated patients suffering from a variety of eye diseases. We report the first observation of the involvement of SEMA4A gene mutations causing retinitis pigmentosa (RP) and cone rod dystrophy (CRD). We screened the DNA of 135 patients with RP, 25 patients with CRD, and 30 with LCA using SSCP and direct DNA sequencing for mutations in the SEMA4A gene. Two mutations, p.D345H and p.F350C, were observed only in affected patients; they were not observed in any of the normal members or the 100 control subjects. Both mutations identified occur in the conserved semaphorin domain. Multiple sequence alignments using Clustal analysis showed that R713Q is a conserved substitution and D345H is a semi-conserved substitution. We conclude that these mutations are a cause of various retinal degenerations.

Blindness↗

Identification by 16S ribosomal RNA gene sequencing of an Enterobacteriaceae species from a bone marrow transplant recipient.

AIMS: To ascertain the clinical relevance of a strain of Enterobacteriaceae isolated from the stool of a bone marrow transplant recipient with diarrhoea. The isolate could not be identified to the genus level by conventional phenotypic methods and required 16S ribosomal RNA (rRNA) gene sequencing for full identification. METHODS: The isolate was investigated phenotypically by standard biochemical methods using conventional biochemical tests and two commercially available systems, the Vitek (GNI+) and API (20E) systems. Genotypically, the 16S bacterial rRNA gene was amplified by the polymerase chain reaction (PCR) and sequenced. The sequence of the PCR product was compared with known 16S rRNA gene sequences in the GenBank database by multiple sequence alignment. RESULTS: Conventional biochemical tests did not reveal a pattern resembling any known member of the Enterobacteriaceae family. The isolate was identified as Salmonella arizonae (73%) and Escherichia coli (76%) by the Vitek (GNI+) and API (20E) systems, respectively. 16S rRNA sequencing showed that there was only one base difference between the isolate and E coli K-12, but 48 and 47 base differences between the isolate and S typhimurium (NCTC 8391) and S typhi (St111), respectively, showing that it was an E coli strain. The patient did not require any specific treatment and the diarrhoea subsided spontaneously. CONCLUSIONS: 16S rRNA gene sequencing was useful in ascertaining the clinical relevance of the strain of Enterobacteriaceae isolated from the stool of the bone marrow transplant recipient with diarrhoea.

Adult↗

The molecular basis of allergenicity: comparative analysis of the three dimensional structures of diverse allergens reveals a common structural motif.

BACKGROUND: Although a large number of allergens have been characterised, the structural, functional, and biochemical features that these molecules have in common, and that could explain their ability to elicit powerful IgE antibody responses, are still uncertain. Recently, there has been considerable interest in the role of the cysteine protease activity of the house dust mite allergen Der p 1 in biasing the immune response in favour of IgE production. AIMS: To search for remote homologues of Der p 1 with sequences similar to the 30 conserved amino acids surrounding the catalytic cysteine residue (Cys34). METHODS: Potential homologues were analysed by examining their three dimensional structures and multiple sequence alignments using the programs PROPSEARCH, ClustalW, GeneDoc, and Swiss Pdb Viewer. RESULTS: Diverse allergens (for example, the plant cysteine protease papain, the transport protein lipocalin Mus m 1, and the ragweed allergen Amb a 5) have a similar structural motif; namely, a groove resembling the substrate binding groove of Der p 1. The groove is located inside an alpha-beta motif, between an alpha helix on one side and an antiparallel beta sheet on the other side. A similar common motif (a cysteine stabilised alpha-beta fold) can also be found in some toxins and defensins. CONCLUSION: Allergens of diverse sources have a common structural motif, namely a groove located inside an alpha-beta motif, which could potentially serve as a ligand binding site.

Allergens↗

Identification of slide coagulase positive, tube coagulase negative Staphylococcus aureus by 16S ribosomal RNA gene sequencing.

AIMS: To ascertain the clinical importance of a strain of slide coagulase positive but tube coagulase negative Staphylococcus species isolated from the blood culture of a 43 year old patient with refractory anaemia with excessive blasts in transformation who had neutropenic fever. METHODS: The isolate was investigated phenotypically by standard biochemical methods using conventional biochemical tests and two commercially available systems, the Vitek (GPI) and API (Staph) systems. Genotypically, the 16S ribosomal RNA (rRNA) gene of the bacteria was amplified by the polymerase chain reaction (PCR) and sequenced. The sequence of the PCR product was compared with known 16S rRNA gene sequences in the GenBank by multiple sequence alignment. RESULTS: Conventional biochemical tests did not reveal a pattern resembling a known Staphylococcus species. The Vitek system (GPI) showed that it was 94% S. simulans and 3% S. haemolyticus, whereas the API system (Staph) showed that it was 86.8% S. aureus and 5.1% S. warneri. 16S rRNA gene sequencing showed that there was a 0 base difference between the isolate and S. aureus, 28 base difference between the isolate and S. lugdunensis, 39 base difference between the isolate and S. schleiferi, 21 base difference between the isolate and S. haemolyticus, 41 base difference between the isolate and S. simulans, and 23 base difference between the isolate and S. warneri, indicating that the isolate was a strain of S. aureus. Vancomycin was subsequently prescribed and blood cultures taken four days after the start of treatment were negative. CONCLUSIONS: 16S rRNA gene sequencing was useful in ascertaining the clinical importance of the strain of slide coagulase positive but tube coagulase negative Staphylococcus species isolated from blood culture and allowing appropriate management.

Adult↗

Identification by 16S ribosomal RNA gene sequencing of Arcobacter butzleri bacteraemia in a patient with acute gangrenous appendicitis.

AIMS: To identify a strain of Gram negative facultative anaerobic curved bacillus, concomitantly isolated with Escherichia coli and Streptococcus milleri, from the blood culture of a 69 year old woman with acute gangrenous appendicitis. The literature on arcobacter bacteraemia and arcobacter infections associated with appendicitis was reviewed. METHODS: The isolate was phenotypically investigated by standard biochemical methods using conventional biochemical tests. Genotypically, the 16S ribosomal RNA (rRNA) gene of the bacterium was amplified by the polymerase chain reaction (PCR) and sequenced. The sequence of the PCR product was compared with known 16S rRNA gene sequences in the GenBank by multiple sequence alignment. Literature review was performed by MEDLINE search (1966-2000). RESULTS: The bacterium grew on blood agar, chocolate agar, and MacConkey agar to sizes of 1 mm in diameter after 24 hours of incubation at 37 degrees C in 5% CO2. It grew at 15 degrees C, 25 degrees C, and 37 degrees C; it also grew in a microaerophilic environment, and was cytochrome oxidase positive and motile, typically a member of the genus arcobacter. Furthermore, phenotypic testing showed that the biochemical profile of the isolate did not fit into the pattern of any of the known arcobacter species. 16S rRNA gene sequencing showed one to two base differences between the isolate and A butzleri, but 35 to 39 base differences between the isolate and A cryaerophilus, indicating that the isolate was a strain of A butzleri. Only three cases of arcobacter bacteraemia with detailed clinical characteristics were found in the English literature. The sources of the arcobacter species in the three cases were largely unknown, although the gastrointestinal tract is probably the portal of entry of the A butzleri isolated from the present case because the two concomitant isolates (E coli and S milleri) in the blood culture were common flora of the gastrointestinal tract. In addition, A butzleri has previously been isolated from the abdominal contents or peritoneal fluid of three patients with acute appendicitis. CONCLUSIONS: 16S rRNA gene sequencing was useful in the identification of the strain of A butzleri isolated from the blood culture of a patient with acute gangrenous appendicitis. Arcobacter bacteraemia is rare. Further studies using selective medium for the delineation of the association between A butzleri and acute appendicitis are warranted.

Acute Disease↗

The genes encoding granule-bound starch synthases at the waxy loci of the A, B, and D progenitors of common wheat.

Three genes encoding granule-bound starch synthase (wx-TmA, wx-TsB, and wx-TtD) have been isolated from Triticum monococcum (AA), and Triticum speltoides (BB), by the polymerase chain reaction (PCR) approach, and from Triticum tauschii (DD), by screening a genomic DNA library. Multiple sequence alignment indicated that the wx-TmA, wx-TsB, and wx-TtD genes had the same extron and (or) intron structure as the previously reported waxy gene from barley. The lengths of the three wx-TmA, wx-TsB, and wx-TtD genes were 2834 bp, 2826 bp, and 2893 bp, respectively, each covering 31 bp in the untranslated leader and the entire coding region consisting of 11 exons and 10 introns. The three genes had identical lengths of exons, except exonl, and shared over 95% identity with each other within the exon regions. The majority of introns were significantly variable in length and sequence, differing mainly in length (1-57 bp) as a result of insertion and (or) deletion events. The deduced amino acid sequence from these three genes indicated that the mature WX-TMA, -TSB, and -TTD proteins contained the same number of amino acids, but differed in predicted molecular weight and isoelectric point (pI) due to amino acid substitutions (13-18). The predicted physical characteristics of the WX proteins matched the respective proteins in wheat very closely, but the match was not perfect. Furthermore the exon5 sequences of the wx-TmA, wx-TsB, and wx-TtD genes were different from a cDNA encoding a waxy gene of common wheat previously reported. The striking difference was that an insertion of 11 amino acids occurred in the cDNA sequence that could not be observed in the exons of the A, B, and D genes. It was noted, however, that the 3' end of intron4 of these genes could account for the additional 11 amino acids. The sequence information from the available waxy genes identified the intron4-exon5-intron5 region as being diagnostic for sequence variation in waxy. The sequence variation in the waxy genes provides the basis for primer design to distinguish the respective genes in common wheat, and its progenitors, using PCR.

Amino Acid Sequence↗

Mutational analysis of conserved glycines 42 and 256 in Cephalosporium acremonium isopenicillin N synthase.

Isopenicillin N synthase (IPNS) is critical for the catalytic conversion of delta-(L-alpha-aminoadipoyl)-L-cysteinyl-D-valine to isopenicillin N in the penicillin and cephalosporin biosynthetic pathway. Two conserved glycine residues in Cephalosporium acremonium IPNS (cIPNS), namely glycine-42 and glycine-256, were identified by multiple sequence alignment and investigated by site-directed mutagenesis to study the effect of the substitution on catalysis. Our study showed that both the mutations from glycine to alanine or to serine reduced the catalytic activity of cIPNS and affected its soluble expression in a heterologous host at 37 degrees C. Soluble expression was restored at a reduced temperature of 25 degrees C, and thus, it is possible that these glycine residues may have a role in maintaining the local protein structure and are critical for the soluble expression of cIPNS.

Acremonium↗

Molecular characterization of DNA encoding 16S-23S rRNA intergenic spacer regions and 16S rRNA of pectolytic Erwinia species.

Sequences of 16S rDNAs and the intergenic spacer (IGS) regions between the 16S and 23S rDNA of bacterial strains from genus Erwinia were determined. Comparison of 16S rDNA sequences from different species and subspecies clearly revealed intraspecies-subspecies homology and interspecies heterogeneity. Phylogenetic analyses of 16S rDNA sequence data revealed that Erwinia spp. formed a discrete monophyletic clade with moderate to high bootstrap values. PCR amplification of the 16S-23S rDNA regions using primers complementary to the 3' end of 16S and 5' end of 23S rRNA genes generated two DNA fragments. The small 16S-23S rDNA IGS regions of Erwinia spp. examined in this study varied considerably in size and nucleotide sequence. Multiple sequence alignment and phylogenetic analysis of small IGS sequence data showed a consistent relationship among the test strains that was roughly in agreement with the 16S rDNA data that reflected the accepted species and subspecies structure of the taxon. Sequence data derived from the large IGS resolved the strains into coherent groups; however, the sequence information would not allow any phylogenetic conclusion, because it failed to reflect the accepted species structure of the test strains.

Base Sequence↗

ANN-Spec: a method for discovering transcription factor binding sites with improved specificity.

This work describes ANN-Spec, a machine learning algorithm and its application to discovering un-gapped patterns in DNA sequence. The approach makes use of an Artificial Neural Network and a Gibbs sampling method to define the Specificity of a DNA-binding protein. ANN-Spec searches for the parameters of a simple network (or weight matrix) that will maximize the specificity for binding sequences of a positive set compared to a background sequence set. Binding sites in the positive data set are found with the resulting weight matrix and these sites are then used to define a local multiple sequence alignment. Training complexity is O(lN) where l is the width of the pattern and N is the size of the positive training data. A quantitative comparison of ANN-Spec and a few related programs is presented. The comparison shows that ANN-Spec finds patterns of higher specificity when training with a background data set. The program and documentation are available from the authors for UNIX systems.

Algorithms↗

Identification and analysis of proton-translocating pyrophosphatases in the methanogenic archaeon Methansarcina mazei.

Analysis of genome sequence data from the methanogenic archaeon Methanosarcina mazei Gö1 revealed the existence of two open reading frames encoding proton-translocating pyrophosphatases (PPases). These open reading frames are linked by a 750-bp intergenic region containing TC-rich stretches and are transcribed in opposite directions. The corresponding polypeptides are referred to as Mvp1 and Mvp2 and consist of 671 and 676 amino acids, respectively. Both enzymes represent extremely hydrophobic, integral membrane proteins with 15 predicted transmembrane segments and an overall amino acid sequence similarity of 50.1%. Multiple sequence alignments revealed that Mvp1 is closely related to eukaryotic PPases, whereas Mvp2 shows highest homologies to bacterial PPases. Northern blot experiments with RNA from methanol-grown cells harvested in the mid-log growth phase indicated that only Mvp2 was produced under these conditions. Analysis of washed membranes showed that Mvp2 had a specific activity of 0.34 U mg (protein)(-1). Proton translocation experiments with inverted membrane vesicles prepared from methanol-grown cells showed that hydrolysis of 1 mol of pyrophosphate was coupled to the translocation of about 1 mol of protons across the cytoplasmic membrane. Appropriate conditions for mvp1 expression could not be determined yet. The pyrophosphatases of M. mazei Gö1 represent the first examples of this enzyme class in methanogenic archaea and may be part of their energy-conserving system.

Amino Acid Sequence↗

Identification of SmtB/ArsR cis elements and proteins in archaea using the Prokaryotic InterGenic Exploration Database (PIGED).

Microbial genome sequencing projects have revealed an apparently wide distribution of SmtB/ArsR metal-responsive transcriptional regulators among prokaryotes. Using a position-dependent weight matrix approach, prokaryotic genome sequences were screened for SmtB/ArsR DNA binding sites using data derived from intergenic sequences upstream of orthologous genes encoding these regulators. Sixty SmtB/ArsR operators linked to metal detoxification genes, including nine among various archaeal species, are predicted among 230 annotated and draft prokaryotic genome sequences. Independent multiple sequence alignments of putative operator sites and corresponding winged helix-turn-helix motifs define sequence signatures for the DNA binding activity of this SmtB/ArsR subfamily. Prediction of an archaeal SmtB/ArsR based upon these signature sequences is confirmed using purified Methanosarcina acetivorans C2A protein and electrophoretic mobility shift assays. Tools used in this study have been incorporated into a web application, the Prokaryotic InterGenic Exploration Database (PIGED; http://bioinformatics.uwp.edu/~PIGED/home.htm), facilitating comparable studies. Use of this tool and establishment of orthology based on DNA binding signatures holds promise for deciphering potential cellular roles of various archaeal winged helix-turn-helix transcriptional regulators.

Archaea↗

Comparison of the 5.8S rRNA gene and internal transcribed spacer regions of trichomonadid protozoa recovered from the bovine preputial cavity.

Sequence analysis of the 5.8S rRNA gene and the internal transcribed spacer regions (ITSRs) was used to compare trichomonadid protozoa (n = 39) of varying morphologies isolated from the bovine preputial cavity. A multiple sequence alignment was performed with bovine isolate sequences and other trichomonadid protozoa sequences available in GenBank. As a group, Tritrichomonasfoetus isolates (n = 7) had nearly complete homology. A similarity matrix showed low homology between the T. foetus isolates and other trichomonads recovered from cattle (<70%). Two clusters of trichomonads other than T. foetus were identified. Eighteen isolates comprised 1 group. These isolates shared >99% homology among themselves and with Pentatrichomonas hominis. The other non-T. foetus cluster (n = 14) did not exhibit a high degree of homology (<87%) with other bovine isolates or any of the trichomonad sequences available in GenBank. The sequence homology among isolates in that cluster was >99%, except for 1 isolate that varied from the others in both ITSRs (approximately 2% dissimilarity). Sequence analysis of the 5.8S rRNA gene and ITSRs was useful for comparing trichomonadid protozoa isolated from the bovine preputial cavity and demonstrated that 2 distinct types of trichomonads constituted the non-T. foetus isolates recovered from the bovine preputial cavity.

Animals↗

Study on the phylogenetic tree of human platelet glycoproteins.

Platelet glycoprotein is an important group of glycoproteins on platelets. Several types of platelet glycoproteins have been studied for their functions in the hemostasis system. A bioinformatic analysis was performed to find out how the platelet glycoproteins' genes are related to each other. A multiple sequence alignment phylogenetic tree was performed to present the family tree of the human platelet glycoproteins recorded in the genomic database, ExPASY. These derived sequences from the database were processed by ClustalW and subsequently used for preparation of the distance matrix by Phylip protdist. The final generated phylogenetic tree of human platelet glycoproteins was presented and discussed.

Blood Platelets↗

Identification and characterization of subfamily-specific signatures in a large protein superfamily by a hidden Markov model approach.

BACKGROUND: Most profile and motif databases strive to classify protein sequences into a broad spectrum of protein families. The next step of such database studies should include the development of classification systems capable of distinguishing between subfamilies within a structurally and functionally diverse superfamily. This would be helpful in elucidating sequence-structure-function relationships of proteins. RESULTS: Here, we present a method to diagnose sequences into subfamilies by employing hidden Markov models (HMMs) to find windows of residues that are distinct among subfamilies (called signatures). The method starts with a multiple sequence alignment (MSA) of the subfamily. Then, we build a HMM database representing all sliding windows of the MSA of a fixed size. Finally, we construct a HMM histogram of the matches of each sliding window in the entire superfamily. To illustrate the efficacy of the method, we have applied the analysis to find subfamily signatures in two well-studied superfamilies: the cadherin and the EF-hand protein superfamilies. As a corollary, the HMM histograms of the analyzed subfamilies revealed information about their Ca2+ binding sites and loops. CONCLUSIONS: The method is used to create HMM databases to diagnose subfamilies of protein superfamilies that complement broad profile and motif databases such as BLOCKS, PROSITE, Pfam, SMART, PRINTS and InterPro.

Binding Sites↗

A comprehensive comparison of comparative RNA structure prediction approaches.

BACKGROUND: An increasing number of researchers have released novel RNA structure analysis and prediction algorithms for comparative approaches to structure prediction. Yet, independent benchmarking of these algorithms is rarely performed as is now common practice for protein-folding, gene-finding and multiple-sequence-alignment algorithms. RESULTS: Here we evaluate a number of RNA folding algorithms using reliable RNA data-sets and compare their relative performance. CONCLUSIONS: We conclude that comparative data can enhance structure prediction but structure-prediction-algorithms vary widely in terms of both sensitivity and selectivity across different lengths and homologies. Furthermore, we outline some directions for future research.

Algorithms↗