Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequence Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Sequence analysis of expressed sequence tags from an ABA-treated cDNA library identifies stress response genes in the moss Physcomitrella patens.

Partial cDNA sequencing was used to obtain 169 expressed sequence tags (ESTs) in the moss, Physcomitrella patens. The source of ESTs was a random cDNA library constructed from 7 day-old protonemata following treatment with 10(-4) M abscisic acid (ABA). Analysis of the ESTs identified 69% with homology to known sequences, 61% of which had significant homology to sequences of plant origin. More importantly, at least 11 ESTs had significant similarities to genes which are implicated in plant stress-responses, including responses which may involve ABA. These included a cDNA associated with desiccation tolerance, two heat shock protein genes, one cold acclimation protein cDNA and five others that may be involved in either oxidative or chemical stress or both, i.e., Zn/Cu-superoxide dismutase, NADPH protochlorophyllide oxidoreductase (PorB), selenium binding protein, glutathione peroxidase and glutathione S transferase. Analysis of codon usage between P. patens and seed plants indicated that although mosses and higher plants are to a large extent similar, minor variations also exists that may represent the distinctiveness of each group.

Abscisic Acid↗

A model of the statistical power of comparative genome sequence analysis.

Comparative genome sequence analysis is powerful, but sequencing genomes is expensive. It is desirable to be able to predict how many genomes are needed for comparative genomics, and at what evolutionary distances. Here I describe a simple mathematical model for the common problem of identifying conserved sequences. The model leads to some useful rules of thumb. For a given evolutionary distance, the number of comparative genomes needed for a constant level of statistical stringency in identifying conserved regions scales inversely with the size of the conserved feature to be detected. At short evolutionary distances, the number of comparative genomes required also scales inversely with distance. These scaling behaviors provide some intuition for future comparative genome sequencing needs, such as the proposed use of "phylogenetic shadowing" methods using closely related comparative genomes, and the feasibility of high-resolution detection of small conserved features.

Animals↗

Rigorous pattern-recognition methods for DNA sequences. Analysis of promoter sequences from Escherichia coli.

The basic nature of the sequence features that define a promoter sequence for Escherichia coli RNA polymerase have been established by a variety of biochemical and genetic methods. We have developed rigorous analytical methods for finding unknown patterns that occur imperfectly in a set of several sequences, and have used them to examine a set of bacterial promoters. The algorithm easily discovers the "consensus" sequences for the -10 and -35 regions, which are essentially identical to the results of previous analyses, but requires no prior assumptions about the common patterns. By explicitly specifying the nature of the search for consensus sequences, we give a rigorous definition to this concept that should be widely applicable. We also have provided estimates for the statistical significance of common patterns discovered in sets of sequences. In addition to providing a rigorous basis for defining known consensus regions, we have found additional features in these promoters that may have functional significance. These added features were located on either side of the -35 region. The pattern 5', or upstream, from the -35 region was found using the standard alphabet (A, G, C and T), but the pattern between the -10 and the -35 regions was detectable only in a sub-alphabet. Recent results relating DNA sequence to helix conformation suggest that the former (upstream) pattern may have a functional significance. Possible roles in promoter function are discussed in this light, and an observation of altered promoter function involving the upstream region is reported that appears to support the suggestion of function in at least one case.

Base Sequence↗

Comparative nucleotide and amino acid sequence analysis of the sequence-specific RNA-binding rotavirus nonstructural protein NSP3.

NSP3, an acidic nonstructural protein, encoded by gene 7 has been implicated as the key player in the assembly of the 11 viral plus-strand RNAs into the early replication intermediates during rotavirus morphogenesis. To date, the sequence of NSP3 from only three animal rotaviruses (SA11, SA114F, and bovine UK) has been determined and that from a human strain has not been reported. To determine the genetic diversity among gene 7 alleles from group A rotaviruses, the nucleotide sequence of the NSP3 gene from 13 strains belonging to nine different G serotypes, from both humans and animals, has been determined. Based on the amino acid sequence identity as well as phylogenetic analysis, NSP3 from group A rotaviruses falls into three evolutionarily related groups, i.e., the SA11 group, the Wa group, and the S2 group. The SA11/SA114F gene appears to have a distant ancestral origin from that of the others and codes for a polypeptide of 315 amino acids (aa) in length. NSP3 from all other group A rotaviruses is only 313 aa in length because of a 2-amino-acid deletion near the carboxy-terminus. While the SA114F gene has the longest 3' untranslated region (UTR) of 132 nucleotides, that from other strains suffered deletions of varying lengths at two positions downstream of the translational termination codon. In spite of the divergence of the nucleotide (nt) sequence in the protein coding region, a stretch of about 80 nt in the 3' UTR is highly conserved in the NSP3 gene from all the strains. This conserved sequence in the 3' UTR might play an important role in the regulation of expression of the NSP3 gene.

Amino Acid Sequence↗

Cloning, sequence analysis and confirmation of derived gene sequences for three epitope-mapped monoclonal antibodies against human phagocyte flavocytochrome b.

The integral membrane protein flavocytochrome b (Cyt b) is the catalytic core of the NADPH oxidase complex, a multicomponent enzyme system that initiates a cascade of reactive oxygen species that play a critical role in innate immunity and vascular physiology. Epitope-mapped, monoclonal antibodies (mAb) that recognize the large (gp91phox) and small (p22phox) subunits of Cyt b provide valuable reagents that have been used to examine structural and mechanistic aspects of oxidase function. In the present study, the heavy and light chain variable region genes of the Cyt b-specific mAbs 44.1, NS5, and NL7 have been amplified by RT-PCR, cloned and subject to DNA sequence analysis. Since the 5' degenerate primer sets used for mAb gene amplification were observed to introduce extensive heterogeneity into the heavy and light chain FR1 regions, N-terminal protein sequence analysis was also conducted to obtain the correct amino acid sequence of this region. In order to confirm the identity of the cloned genes, intact mAbs were resolved by two-dimensional electrophoresis and subject to in-gel tryptic digestion for analysis by both MALDI and nanospray LC-MS/MS. Databases searches using the derived mAb sequences predicted residues comprising CDR loops, identified candidate germline genes, and showed the respective germline genes to accurately predict the N-terminal amino acid residues for each variable region. The above studies report the amino acid sequence of Cyt b-specific mAb variable region genes with high confidence and provide essential information for future efforts at Cyt b structure analysis by resonance energy transfer and X-ray crystallography.

Amino Acid Sequence↗

Sulfide-quinone reductase from Rhodobacter capsulatus: requirement for growth, periplasmic localization, and extension of gene sequence analysis.

The entire sequence of the 3.5-kb fragment of genomic DNA from Rhodobacter capsulatus which contains the sqr gene and a second complete and two further partial open reading frames has been determined. A correction of the previously published sqr gene sequence (M. Schütz, Y. Shahak, E. Padan, and G. Hauska, J. Biol. Chem. 272:9890-9894, 1997) which in the deduced primary structure of the sulfide-quinone reductase changes four positive into four negative charges and the number of amino acids from 425 to 427 was necessary. The correction has no further bearing on the former sequence analysis. Deletion and interruption strains document that sulfide-quinone reductase is essential for photoautotrophic growth on sulfide. The sulfide-oxidizing enzyme is involved in energy conversion, not in detoxification. Studies with an alkaline phosphatase fusion protein reveal a periplasmic localization of the enzyme. Exonuclease treatment of the fusion construct demonstrated that the C-terminal 38 amino acids of sulfide-quinone reductase were required for translocation. An N-terminal signal peptide for translocation was not found in the primary structure of the enzyme. The possibility that the neighboring open reading frame, which contains a double arginine motif, may be involved in translocation has been excluded by gene deletion (rather, the product of this gene functions in an ATP-binding cassette transporter system, together with the product of one of the other open reading frames). The results lead to the conclusion that the sulfide-quinone reductase of R. capsulatus functions at the periplasmic surface of the cytoplasmic membrane and that this flavoprotein is translocated by a hitherto-unknown mechanism.

Alkaline Phosphatase↗

AGenDA: gene prediction by comparative sequence analysis.

UNLABELLED: Comparative sequence analysis is a powerful approach to identify functional elements in genomic sequences. Herein, we describe AGenDA (Alignment-based GENe Detection Algorithm), a novel method for gene prediction that is based on long-range alignment of syntenic regions in eukaryotic genome sequences. Local sequence homologies identified by the DIALIGN program are searched for conserved splice signals to define potential protein-coding exons; these candidate exons are then used to assemble complete gene structures. The performance of our method was tested on a set of 105 human-mouse sequence pairs. These test runs showed that sensitivity and specificity of AGenDA are comparable with the best gene- prediction program that is currently available. However, since our method is based on a completely different type of input information, it can detect genes that are not detectable by standard methods and vice versa. Thus, our approach seems to be a useful addition to existing gene-prediction programs. AVAILABILITY: DIALIGN is available through the Bielefeld Bioinformatics Server (BiBiServ) at http://bibiserv.techfak.uni-bielefeld.de/dialign/ The gene-prediction program AGenDA described in this paper will be available through the BiBiServ or MIPS web server at http://mips.gsf.de.

Algorithms↗

Phylogenetic relationships of seven palearctic members of the maculipennis complex inferred from ITS2 sequence analysis.

The sequences of the second internal transcribed spacer (ITS2) of ribosomal DNA (rDNA) were determined from seven palearctic mosquitoes species belonging to the Anopheles maculipennis species complex, namely An. atroparvus, An. labranchiae, An. maculipennis, An. messeae, An. melanoon, An. sacharovi and An. martinius. The length of the ITS2 ranged from 280 to 300 bp, with a GC content of 49.4-54.1%. With the exception of An. messeae, negligible levels of intraspecific polymorphism and no intrapopulation variation were observed. The phylogenetic relationships among the members of the maculipennis complex were inferred by maximum-parsimony analysis of the PAUP program and the neighbour-joining and maximum-likelihood analysis of the PHYLIP program. All of the trees obtained were almost identical in topology, although the relationships among three species, i.e. An. maculipennis, An. messeae and An. melanoon, remained unresolved. The phylogenies were in good agreement with the previous gene-enzyme and polytene chromosome banding pattern studies.

Animals↗

[The study of insertion sequence IS2. II. Polarity effect of the mutant and DNA sequence analysis].

Insertion Sequence IS2 brought about polarity effect when it was inserted into a transcription. In study I, we found that the into polarity effects were different whether IS2 was inserted into the left side or into the right side. In this study, further research also showed that the polarity effects were very different when different IS2 was inserted into the left side. The analysis on IS2 DNA sequence showed that when the direction of IS2 ORF was the same as the transcription direction of the inserted transcription, known as a left insertion, the polarity effect was weaker. When the direction of IS2 ORF was opposite to the transcription direction of the inserted transcription, known as a right insertion, the polarity effect was stronger. When a stop codon was produced in IS2 due to base mutation, the polarity effect of a left insertion was increased.

Base Sequence↗

PEPPLOT, a protein secondary structure analysis program for the UWGCG sequence analysis software package.

We describe a program for the analysis of protein secondary structure that operates with the Sequence Analysis Software Package of the University of Wisconsin Genetics Computer Group (UWGCG). The program produces both graphic and printed output. Structure prediction using the Chou and Fasman and Robson et al methods, and hydropathy analysis by the method of Kyte and Doolittle are included along with a simplified method of hydrophobic moment analysis. The power of the program is the coordinated presentation of many different kinds of structural information on the same plot.

Amino Acid Sequence↗

Protein sequence analysis using Hewlett-Packard biphasic sequencing cartridges in an applied biosystems 473A protein sequencer.

Protein sequence analysis using an adsorptive biphasic sequencing cartridge, a set of two coupled columns introduced by Hewlett-Packard for protein sequencing by Edman degradation, in an Applied Biosystems 473A protein sequencer has been demonstrated. Samples containing salts, detergents, excipients, etc. (e.g., formulated protein drugs) can be easily analyzed using the ABI sequencer. Simple modifications to the ABI sequencer to accommodate the cartridge extend its utility in the analysis of difficult samples. The ABI sequencer solvents and reagents were compatible with the HP cartridge for sequencing. Sequence information up to ten residues can be easily generated by this nonoptimized procedure, and it is sufficient for identifying proteins by database search and for preparing a DNA probe for cloning novel proteins.

Immunoglobulin G↗

AnaBench: a Web/CORBA-based workbench for biomolecular sequence analysis.

BACKGROUND: Sequence data analyses such as gene identification, structure modeling or phylogenetic tree inference involve a variety of bioinformatics software tools. Due to the heterogeneity of bioinformatics tools in usage and data requirements, scientists spend much effort on technical issues including data format, storage and management of input and output, and memorization of numerous parameters and multi-step analysis procedures. RESULTS: In this paper, we present the design and implementation of AnaBench, an interactive, Web-based bioinformatics Analysis workBench allowing streamlined data analysis. Our philosophy was to minimize the technical effort not only for the scientist who uses this environment to analyze data, but also for the administrator who manages and maintains the workbench. With new bioinformatics tools published daily, AnaBench permits easy incorporation of additional tools. This flexibility is achieved by employing a three-tier distributed architecture and recent technologies including CORBA middleware, Java, JDBC, and JSP. A CORBA server permits transparent access to a workbench management database, which stores information about the users, their data, as well as the description of all bioinformatics applications that can be launched from the workbench. CONCLUSION: AnaBench is an efficient and intuitive interactive bioinformatics environment, which offers scientists application-driven, data-driven and protocol-driven analysis approaches. The prototype of AnaBench, managed by a team at the Université de Montréal, is accessible on-line at: http://malawimonas.bcm.umontreal.ca:8091/anabench. Please contact the authors for details about setting up a local-network AnaBench site elsewhere.

Computational Biology↗

Identifying constraints on the higher-order structure of RNA: continued development and application of comparative sequence analysis methods.

Comparative sequence analysis addresses the problem of RNA folding and RNA structural diversity, and is responsible for determining the folding of many RNA molecules, including 5S, 16S, and 23S rRNAs, tRNA, RNAse P RNA, and Group I and II introns. Initially this method was utilized to fold these sequences into their secondary structures. More recently, this method has revealed numerous tertiary correlations, elucidating novel RNA structural motifs, several of which have been experimentally tested and verified, substantiating the general application of this approach. As successful as the comparative methods have been in elucidating higher-order structure, it is clear that additional structure constraints remain to be found. Deciphering such constraints requires more sensitive and rigorous protocols, in addition to RNA sequence datasets that contain additional phylogenetic diversity and an overall increase in the number of sequences. Various RNA databases, including the tRNA and rRNA sequence datasets, continue to grow in number as well as diversity. Described herein is the development of more rigorous comparative analysis protocols. Our initial development and applications on different RNA datasets have been very encouraging. Such analyses on tRNA, 16S and 23S rRNA are substantiating previously proposed associations and are now beginning to reveal additional constraints on these molecules. A subset of these involve several positions that correlate simultaneously with one another, implying units larger than a basepair can be under a phylogenetic constraint.

Base Sequence↗

Direct cloning and sequence analysis of enzymatically amplified genomic sequences.

A method is described for directly cloning enzymatically amplified segments of genomic DNA into an M13 vector for sequence analysis. A 110-base pair fragment of the human beta-globin gene and a 242-base pair fragment of the human leukocyte antigen DQ alpha locus were amplified by the polymerase chain reaction method, a procedure based on repeated cycles of denaturation, primer annealing, and extension by DNA polymerase I. Oligonucleotide primers with restriction endonuclease sites added to their 5' ends were used to facilitate the cloning of the amplified DNA. The analysis of cloned products allowed the quantitative evaluation of the amplification method's specificity and fidelity. Given the low frequency of sequence errors observed, this approach promises to be a rapid method for obtaining reliable genomic sequences from nanogram amounts of DNA.

Base Sequence↗

ADSP--a new package for computational sequence analysis.

A new protein sequence analysis package, ADSP, is described, of which the SOMAP Screen-Oriented Multiple Alignment Procedure forms an integral part. ADSP (Algorithms and Data Structures for Protein sequence analysis) incorporates facilities to generate potent pattern-recognition discriminators and offers four algorithms with which to scan any NBRF format sequence database: the package has been designed, in particular, to interface with the OWL composite sequence database, one of the largest, distributed non-redundant sources of sequence data of its kind. The system incorporates a powerful method for compound feature analysis, which provides the basis for characterizing and predicting the occurrence of complete protein superfamilies and for pinpointing the emergence of related sub-families. Used iteratively, the approach allows diagnostic performance to be rigorously refined and its efficacy to be assessed both qualitatively and quantitatively, and results in the generation of refined structural or functional features suitable for entry into a database: this compilation of characteristic signatures is distinct from, but complementary to, widely used compendia of pattern templates such as PROSITE.

Algorithms↗

Molecular cloning and sequence analysis of highly repetitive DNA sequences contained in the eliminated genome of Ascaris lumbricoides.

High molecular weight DNA from germ line and somatic cells of the DNA eliminating nematode Ascaris lumbricoides has has been isolated and digested with different restriction enzymes. The resulting DNA fragments were separated by agarose gel electrophoresis. Germ line but not somatic DNA shows a prominent band about 120 bp long as well as multiples of that length. These fragments are shown to be monomers and multimers of a highly repetitive satellite DNA, which is eliminated mostly but not completely during the process of chromatin diminution. Restriction digests, hybridization experiments and sequence analysis revealed that this eliminated satellite is composed of a whole set of different but related variant classes, all of them showing the same repeating unit length of about 120 bp. Members of the same variant class are tandemly linked and therefore physically separated from other variant classes. All satellite sequences can be derived from the same common ancestor sequence, differing only by base substitutions, insertions and deletions. There is no evidence for transcription of satellite DNA at any stage and tissues analyzed.

Animals↗