Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequence Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

High-speed conversion of cytosine to uracil in bisulfite genomic sequencing analysis of DNA methylation.

Bisulfite genomic sequencing is a widely used technique for analyzing cytosine-methylation of DNA. By treating DNA with bisulfite, cytosine residues are deaminated to uracil, while leaving 5-methylcytosine largely intact. Subsequent PCR and nucleotide sequence analysis permit unequivocal determination of the methylation status at cytosine residues. A major caveat associated with the currently practiced procedure is that it takes 16-20 hr for completion of the conversion of cytosine to uracil. Here we report that a complete deamination of cytosine to uracil can be achieved in shorter periods by using a highly concentrated bisulfite solution at an elevated temperature. Time course experiments demonstrated that treating DNA with 9 M bisulfite for 20 min at 90 degrees C or 40 min at 70 degrees C all cytosine residues in the DNA were converted to uracil. Under these conditions, the majority of 5-methylcytosines remained intact. When a high molecular weight DNA derived from a cell line (containing a number of genes whose methylation status was known) was treated with bisulfite under the above conditions and amplified and sequenced, the results obtained were consistent with those reported in the literature. Although some degradation of DNA occurred during this process, the amount of treated DNA required for the amplification was nearly equal to that required for the conventional bisulfite genomic sequencing procedure. The increased speed of DNA methylation analysis with this novel procedure is expected to advance various aspects of DNA sciences.

Base Sequence↗

[DNA sequence analysis for the promoter of pyruvate oxidase gene from Streptococcus oralis].

OBJECTIVE: To elucidate the molecular structure of pyruvate oxidase gene promoter. METHODS: The 1.30 kb fragment with promoter activity, amplified from upstream of Streptococcus oralis pyruvate oxidase gene (Sopox), was cloned into vector PBK-CMV. The positive transformed E. coli JM109 clone was selected, the recombinant plasmid was further identified with restriction mapping analysis. The positive recombinant plasmid was studied with sequence analysis. RESULTS: After digesting the recombinant plasmid with Hind III, 1% agarose electrophoresis showed 1.30 kb fragment, which was consistent with predicted size. Sequence analysis revealed 1,350 bp. CONCLUSION: The Sopox promoter region is sequenced. Further characterization of the Sopox promoter region will elucidate the molecular mechanism of H2O2 production of streptococcus oralis.

Base Sequence↗

Two-step high resolution sequence-based HLA-DRB typing of exon 2 DNA with taxonomy-based sequence analysis allele assignment.

A two-step high resolution sequence-based DRB typing method was developed. The system needs only one polymerase chain reaction (PCR) to type all functional DRB alleles of a given individual. It uses a pair of generic PCR primers to amplify exon 2 DNA of all functional DRB genes and a first-step taxonomy-based sequence analysis (FSTBSA) method to assign allele groups after sequencing the PCR products with a generic primer. In the second step, group-specific primers are used to sequence the same PCR products and a taxonomy-based sequence analysis (TBSA) is used to assign alleles. Thus, both low and high resolution DRB typing can be done with PCR amplified exon 2 DNA from a single PCR reaction. Correct allele group assignment by FSTBSA was confirmed by sequencing the PCR products with group-specific primers and correctly assigned all 158 DNA samples including 34 samples pre-typed by PCR-sequence-specific primer or PCR-sequence-specific oligonucleotide probe. FSTBSA correctly assigned 116 heterozygous combinations of 81 DRB1-DRB3/4/5 haplotypes. Sixty-seven DRB1, 6 DRB3, 1 DRB4, and 3 DRB5 alleles were identified in this study. TBSA successfully resolved all heterozygous allele combinations including 31 heterozygous combinations of 33 alleles of DRB1*03, 08, 11, 12, 13, and 14 allele groups, and six heterozygous combinations of six DRB3 alleles.

Alleles↗

Sequence analysis of ARS elements in fission yeast.

Chromosomal DNA of Schizosaccharomyces pombe contains sequences with properties analogous to ARS elements of Saccharomyces cerevisiae. Following Sau3A fragmentation of the S. pombe genome we have recovered a number of such fragments in an M13-based shuttle vector, suitable for subsequent sequence analysis. The complete nucleotide sequence has been obtained for eight ARS+ inserts derived from the Sau3A cloning and for the ARS present in pFL20 isolated previously by Losson and Lacroute (Cell, 32, 371-377, 1983). The Sau3A clones are single fragments between 0.8 and 1.8 kb. No ARS+ clones smaller than this were recovered even though the average size Sau3A fragment in S. pombe is approximately 200-300 bp. The sequence analysis revealed that all clones are AT-rich (69-75% A + T residues), and all contain a particularly AT-rich 11 bp core element represented by the consensus sequence 5' (A/T)PuTT-TATTTA(A/T) 3'. Deletion mapping indicates that the consensus in all cases is in the vicinity of a functional ARS domain. However precise excision of the consensus by in vitro mutagenesis has little effect on ARS activity as judged by the transformation assay. We argue that the association of the consensus with the ARS domain occurs too reproducibly to be explained by chance alone. We suggest that although it may not be essential for the extrachromosomal maintenance of plasmids in S. pombe, the consensus does have a function in situ in the chromosome and thus is always present as a cryptic sequence in the isolated ARS element.

Base Sequence↗

Differentiation of Mycobacterium ulcerans, M. marinum, and M. haemophilum: mapping of their relationships to M. tuberculosis by fatty acid profile analysis, DNA-DNA hybridization, and 16S rRNA gene sequence analysis.

Although Mycobacterium ulcerans, M. marinum, and M. haemophilum are closely related, their exact taxonomic placements have not been determined. We performed gas chromatography of fatty acids and alcohols, as well as DNA-DNA hybridization and 16S rRNA gene sequence analysis, to clarify their relationships to each other and to M. tuberculosis. M. ulcerans and M. marinum were most closely related to one another, and each displayed very strong genetic affinities to M. tuberculosis; they are actually the two mycobacterial species outside the M. tuberculosis complex most closely related to M. tuberculosis. M. haemophilum was more distinct from M. ulcerans and M. marinum, and it appeared to be as related to these two species as to M. tuberculosis. These results are important with regard to the development of diagnostic and epidemiological tools such as species-specific DNA probes and PCR assays for M. ulcerans, M. marinum, and M. haemophilum. In addition, the finding that M. ulcerans and M. marinum are more closely related to M. tuberculosis than are other pathogenic mycobacterial species suggests that they may be evaluated as useful models for studying the pathogenesis of M. tuberculosis. M. marinum may be particularly useful in this regard since strains of this species grow much more rapidly than M. tuberculosis and yet can cause systemic disease in immunocompromised hosts.

Chromatography, Gas↗

Prediction of human rotavirus serotype by nucleotide sequence analysis of the VP7 protein gene.

Human rotavirus field isolates were characterized by direct sequence analysis of the gene encoding the serotype-specific major neutralization protein (VP7). Single-stranded RNA transcripts were prepared from virus particles obtained directly from stool specimens or after two or three passages in MA-104 cells. Two regions of the gene (nucleotides 307 through 351 and 670 through 711) which had previously been shown to contain regions of sequence divergence among rotavirus serotypes were sequenced by the dideoxynucleotide method with two different synthetic oligonucleotide primers. The resulting nucleotide sequences were compared with the corresponding sequences from rotaviruses of known serotype (serotype 1, 2, 3, or 4). A total of 25 field isolates and 10 laboratory strains examined by this method exhibited marked sequence identity in both areas of the gene with the corresponding regions of 1 of the 4 reference strains. In addition, the predicted serotype from the sequence analysis correlated in each case with the serotype determined when the rotaviruses were examined by plaque reduction neutralization or reactivity with serotype-specific monoclonal antibodies. These data suggest that as a result of the high degree of sequence conservation observed among rotaviruses of the same serotype, it is possible to predict the serotype of a rotavirus isolate by direct sequence analysis of its VP7 gene.

Amino Acid Sequence↗

The Staden sequence analysis package.

I describe the current version of the sequence analysis package developed at the MRC Laboratory of Molecular Biology, which has come to be known as the "Staden Package." The package covers most of the standard sequence analysis tasks such as restriction site searching, translation, pattern searching, comparison, gene finding, and secondary structure prediction, and provides powerful tools for DNA sequence determination. Currently the programs are only available for computers running the UNIX operating system. Detailed information about the package is available from our WWW site: http:@www.mrc-lmb.cam.ac.uk/pubseq/.

Database Management Systems↗

Sequence codes for extended conformation: a neighbor-dependent sequence analysis of loops in proteins.

We performed an extensive sequence analysis on the loops of proteins. By dividing a loop databank derived from the Protein Data Bank into groups, we analyzed the chemical characteristics and the sequence preferences of loops of different lengths and loops connecting different secondary structures in proteins. We found that a large population of loops in our loop databank (94.4%) is either partially or completely surface-exposed. A majority of surface loops in proteins are hydrophilic, whereas the chemical characteristics of interior loops are relatively neutral according to Eisenberg's consensus hydrophobicity scale. As a first step in investigating the intrinsic sequence-structure relationship of loop sequences in proteins, we performed a neighbor-dependent sequence analysis that calculated the effect of the neighboring amino acid type on the loop propensity of residues in loops. This method enhances the statistical significance of residue propensity, thus allowing us to explore the positional preference of amino acids in loops. Our analysis yielded a series of amino acid dyads that showed high preference for loop conformation. The data presented in this study should prove useful for developing potential codes in recognizing loop sequences in proteins.

Amino Acids↗

Comprehensive sequence analysis of the 182 predicted open reading frames of yeast chromosome III.

With the completion of the first phase of the European yeast genome sequencing project, the complete DNA sequence of chromosome III of Saccharomyces cerevisiae has become available (Oliver, S. G., et al., 1992, Nature 357, 38-46). We have tested the predictive power of computer sequence analysis of the 176 probable protein products of this chromosome, after exclusion of six problem cases. When the results of database similarity searches are pooled with prior knowledge, a likely function can be assigned to 42% of the proteins, and a predicted three-dimensional structure to a third of these (14% of the total). The function of the remaining 58% remains to be determined. Of these, about one-third have one or more probable transmembrane segments. Among the most interesting proteins with predicted functions are a new member of the type X polymerase family, a transcription factor with an N-terminal DNA-binding domain related to GAL4, a "fork head" DNA-binding domain previously known only in Drosophila and in mammals, and a putative methyltransferase. Our analysis increased the number of known significant sequence similarities on chromosome III by 13, to now 67. Although the near 40% success rate of identifying unknown protein function by sequence analysis is surprisingly high, the information gap between known protein sequences and unknown function is expected to widen and become a major bottleneck of genome projects in the near future. Based on the experience gained in this test study, we suggest that the development of an automated computer workbench for protein sequence analysis must be an important item in genome projects.

Acetolactate Synthase↗

Sequence analysis of genes and genomes.

A major step towards understanding of the genetic basis of an organism is the complete sequence determination of all genes in its genome. The development of powerful techniques for DNA sequencing has enabled sequencing of large amounts of gene fragments and even complete genomes. Important new techniques for physical mapping, DNA sequencing and sequence analysis have been developed. To increase the throughput, automated procedures for sample preparation and new software for sequence analysis have been applied. This review describes the development of new sequencing methods and the optimisation of sequencing strategies for whole genome and cDNA analysis, as well as discusses issues regarding sequence analysis and annotation.

Animals↗

Compact protein sequencer for the C-terminal sequence analysis of peptides and proteins.

We describe the construction of a compact protein sequencer designed specifically for the C-terminal sequence analysis of peptides and proteins. This sequencer has a vertical flow path and is equipped with a continuous flow reactor (CFR). The flow paths for the various reagents and solvents have been minimized. A unique feature of this instrument is the design of a quadrate valve (quad valve) which permits the delivery of four solvents or reagents to the conversion flask (CF). Combination of two of these quad valves in series permits the delivery of eight solvents and reagents to the CFR. The CF contains three inputs from the top, one for transfer of the contents of the CFR, one which is used as a vent, and one for input of solvents or reagents from the CF quad valve. The CF drains from the bottom, connecting to a switching valve which allows delivery either to a waste bottle or to an on-line HPLC. Another unique feature of this instrument is the design of an optical flow detector which permits injection of approximately 90% of the contents of the CF for HPLC analysis. The overall size of the instrument (11 w x 16.5 h x 23.5 d in.) is smaller than commercially available instruments for protein sequencing and represents the first time an instrument has been constructed specifically for C-terminal sequence analysis. The utility of this instrument is demonstrated with the C-terminal sequence analysis of protein samples noncovalently applied to Zitex strips and with a peptide covalently attached to carboxylic acid-modified polyethylene film.

Amino Acids↗

SynBrowse: a synteny browser for comparative sequence analysis.

MOTIVATION: The recent efforts of various sequence projects to sequence deeply into various phylogenies provide great resources for comparative sequence analysis. A generic and portable tool is essential for scientists to visualize and analyze sequence comparisons. RESULTS: We have developed SynBrowse, a synteny browser for visualizing and analyzing genome alignments both within and between species. It is intended to help scientists study macrosynteny, microsynteny and homologous genes between sequences. It can also aid with the identification of uncharacterized genes, putative regulatory elements and novel structural features of a species. SynBrowse is a GBrowse (the Generic Genome Browser) family software tool that runs on top of the open source BioPerl modules. It consists of two components: a web-based front end and a set of relational database back ends. Each database stores pre-computed alignments from a focus sequence to reference sequences in addition to the genome annotations of the focus sequence. The user interface lets end users select a key comparative alignment type and search for syntenic blocks between two sequences and zoom in to view the relationships among the corresponding genome annotations in detail. SynBrowse is portable with simple installation, flexible configuration, convenient data input and easy integration with other components of a model organism system. AVAILABILITY: The software is available at http://www.gmod.org CONTACT: vbrendel@iastate.edu

Algorithms↗

Mitochondrial DNA sequence analysis of human skeletal remains: identification of remains from the Vietnam War.

Deoxyribonucleic acid (DNA) sequence analysis of the control region of the mitochondrial DNA (mtDNA) genome was used to identify human skeletal remains returned to the United States government by the Vietnamese government in 1984. The postmortem interval was thought to be 24 years at the time of testing, and the remains presumed to be an American service member. DNA typing methods using nuclear genomic DNA, HLA-DQ alpha and the variable number of tandem repeat (VNTR) locus D1S80, were unsuccessful using the polymerase chain reaction (PCR). Amplification of a portion of the mtDNA control region was performed, and the resulting PCR product subjected to DNA sequence analysis. The DNA sequence generated from the skeletal remains was identical to the maternal reference sequence, as well as the sequence generated from two siblings (sisters). The sequence was unique when compared to more than 650 DNA sequences found both in the literature and provided by personal communications. The individual sequence polymorphisms were present in only 23 of the more than 1300 nucleotide positions analyzed. These results support the observation that in cases where conventional DNA typing is unavailable, mtDNA sequencing can be used for human remains identification.

Anthropology, Physical↗

Identification of medically important yeast species by sequence analysis of the internal transcribed spacer regions.

Infections caused by yeasts have increased in previous decades due primarily to the increasing population of immunocompromised patients. In addition, infections caused by less common species such as Pichia, Rhodotorula, Trichosporon, and Saccharomyces spp. have been widely reported. This study extensively evaluated the feasibility of sequence analysis of the rRNA gene internal transcribed spacer (ITS) regions for the identification of yeasts of clinical relevance. Both the ITS1 and ITS2 regions of 373 strains (86 species), including 299 reference strains and 74 clinical isolates, were amplified by PCR and sequenced. The sequences were compared to reference data available at the GenBank database by using BLAST (basic local alignment search tool) to determine if species identification was possible by ITS sequencing. Since the GenBank database currently lacks ITS sequence entries for some yeasts, the ITS sequences of type (or reference) strains of 15 species were submitted to GenBank to facilitate identification of these species. Strains producing discrepant identifications between the conventional methods and ITS sequence analysis were further analyzed by sequencing of the D1-D2 domain of the large-subunit rRNA gene for species clarification. The rates of correct identification by ITS1 and ITS2 sequence analysis were 96.8% (361/373) and 99.7% (372/373), respectively. Of the 373 strains tested, only 1 strain (Rhodotorula glutinis BCRC 20576) could not be identified by ITS2 sequence analysis. In conclusion, identification of medically important yeasts by ITS sequencing, especially using the ITS2 region, is reliable and can be used as an accurate alternative to conventional identification methods.

Ascomycota↗

Molecular cloning and sequence analysis of the porcine precursor of endothelin-2.

The amino acid sequences of two of the three endothelin (ET) family peptides, ET-1 and ET-3, are identical among mammals, whereas for the other family member, ET-2 or vasoactive intestinal contractor (VIC), the mouse and rat sequences differ from the human counterpart ET-2 by one amino acid residue. To examine more deeply the structural diversity among ET-2/VIC orthologs (EDN2), we screened porcine ET-2/VIC-like cDNAs using the 5' rapid amplification of cDNA ends (RACE) method with degenerate primers based on ET-2/VIC mature peptides. Sequence analysis of the cDNAs showed that ET-2 is present in pig. The full-length cDNA sequence, produced by combining 5' RACE and 3' RACE products, revealed the porcine precursor protein of ET-2 (PPET-2). Porcine PPET-2, made up of 214 amino acids, includes a 26-residue putative signal sequence, big ET-2, mature ET-2, and ET-2-like peptide. The percent sequence identity of porcine PPET-2 with human PPET-2, and rat or mouse precursor protein of VIC runs between approximately 70% and 74%. ET-2, although expressed in intestine, has no anti-microbial activity.

Amino Acid Sequence↗

Rapid p53 sequence analysis in primary lung cancer using an oligonucleotide probe array.

The p53 gene was sequenced in 100 primary human lung cancers by using direct dideoxynucleotide cycle sequencing and compared with sequence analysis by using the p53 GeneChip assay. Differences in sequence analysis between the two techniques were further evaluated to determine the accuracy and limitations of each method. p53 mutations were either detected by using both techniques or, if only detected by one technique, were confirmed by using mutation-specific oligonucleotide hybridization. Dideoxynucleotide sequencing of the conserved regions of the p53 gene (exons 5-9) detected 76% of the mutations within this region of the gene. The GeneChip p53 assay detected 81% of all (exons 2-11) mutations, including 80% of the mutations within the conserved regions of the gene. The GeneChip assay detected 46 of 52 missense mutations (88%), but 0 of 5 frameshift mutations. The specificity of direct sequencing and of the p53 GeneChip assay at detecting p53 mutations were 100% and 98%, respectively. The GeneChip p53 assay is a rapid and reasonably accurate approach for detecting p53 mutations; however, neither direct sequencing nor the p53 GeneChip are infallible at p53 mutation detection.

Humans↗

Sequence analysis of the AAA protein family.

The AAA protein family, a recently recognized group of Walker-type ATPases, has been subjected to an extensive sequence analysis. Multiple sequence alignments revealed the existence of a region of sequence similarity, the so-called AAA cassette. The borders of this cassette were localized and within it, three boxes of a high degree of conservation were identified. Two of these boxes could be assigned to substantial parts of the ATP binding site (namely, to Walker motifs A and B); the third may be a portion of the catalytic center. Phylogenetic trees were calculated to obtain insights into the evolutionary history of the family. Subfamilies with varying degrees of intra-relatedness could be discriminated; these relationships are also supported by analysis of sequences outside the canonical AAA boxes: within the cassette are regions that are strongly conserved within each subfamily, whereas little or even no similarity between different subfamilies can be observed. These regions are well suited to define fingerprints for subfamilies. A secondary structure prediction utilizing all available sequence information was performed and the result was fitted to the general 3D structure of a Walker A/GTPase. The agreement was unexpectedly high and strongly supports the conclusion that the AAA family belongs to the Walker superfamily of A/GTPases.

Adenosine Triphosphatases↗

Identification of phosphotyrosine residues during protein sequence analysis.

Synthetic tyrosine-phosphorylated peptides were subjected to protein sequence analysis using a gas-phase sequencer and on-line phenylthiohydantoin (PTH) amino acid analysis. Our data show that phosphotyrosine is stable to the gas-phase sequencing chemistry and can be detected as its PTH-derivative during routine sequence analysis without the need of prior tyrosine radiolabeling.

Amino Acid Sequence↗