Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

How accurately can we discriminate G-protein-coupled receptors as 7-tms TM protein sequences from other sequences?

The group of 2502 transmembrane (TM) protein sequences with seven TM segments (7-tms) registered in SWISS-PROT 46.0 contains 2200 G-protein-coupled receptors (GPCRs), indicating that GPCR candidates can be detected with a reliability of 87.9% in the eukaryotic genomes merely by correctly predicting the number of TM segments as 7-tms. The predictive accuracies of TM topology-prediction methods proposed so far are not as high as expected; even the best method, HMMTOP 2.0, can only achieve a capture rate of 7-tms sequences of 77.6%. It is necessary to improve this performance as much as possible, even if by only a few percentage points, in order to identify as many novel GPCR candidate genes as possible among the increasing number of newly sequenced genomes. In this study, we propose a simple but useful prediction method for detecting as many 7-tms TM protein sequences as GPCR candidates in eukaryotic genomes as possible. This is achieved by employing a two-step prediction procedure. The first step involves collecting 7-tms sequences by the best prediction method (HMMTOP 2.0), and the second involves picking up the remaining 7-tms sequences by the second-best method (TMHMM 2.0). By this procedure, the capture rate of 7-tms TM protein sequences in SWISS-PROT can be improved considerably from 77.6% to 84.5%, and the number of GPCR candidate sequences predicted as 7-tms in the human genome (Build 35) is increased from 790 (by HMMTOP 2.0) to 903. These 790 and 903 candidate sequences include, respectively, 587 and 636 of the known human GPCRs of the 717 registered in SWISS-PROT 46.0, demonstrating that the proposed combinatorial method is effective in detecting GPCR candidate genes in eukaryotic genomes.

Amino Acid Sequence↗

Calculating sequence-dependent melting stability of duplex DNA oligomers and multiplex sequence analysis by graphs.

The analytical methods for characterizing DNA sequence-dependent thermodynamic stability have been reviewed. A set of n-n sequence stability parameters is presented. Examples in which these values are used to calculate the thermodynamic stability of short duplex DNA oligomers are presented. The problem of determining sets of isothermal sequences is addressed by representing DNA sequences as graphs. Representing DNA sequences by a graph descriptor with special mathematical properties minimizes the computational difficulty of determining the number of DNA sequences with identical predicted thermodynamic stability. This is achieved by replacement of a whole set of sequences by a single representative. Applications of this concept were demonstrated for sequences assembled from individual bases and sequences assembled from oligomeric blocks.

Base Sequence↗

Universal minicircle sequence-binding protein, a sequence-specific DNA-binding protein that recognizes the two replication origins of the kinetoplast DNA minicircle.

Replication of the kinetoplast DNA minicircle lagging (heavy (H))-strand initiates at, or near, a unique hexameric sequence (5'-ACGCCC-3') that is conserved in the minicircles of trypanosomatid species. A protein from the trypanosomatid Crithidia fasciculata binds specifically a 14-mer sequence, consisting of the complementary strand hexamer and eight flanking nucleotides at the H-strand replication origin. This protein was identified as the previously described universal minicircle sequence (UMS)-binding protein (UMSBP) (Tzfati, Y., Abeliovich, H., Avrahami, D., and Shlomai, J. (1995) J. Biol. Chem. 270, 21339-21345). This CCHC-type zinc finger protein binds the single-stranded form of both the 12-mer (UMS) and 14-mer sequences, at the replication origins of the minicircle L-strand and H-strand, respectively. The attribution of the two different DNA binding activities to the same protein relies on their co-purification from C. fasciculata cell extracts and on the high affinity of recombinant UMSBP to the two origin-associated sequences. Both the conserved H-strand hexamer and its flanking nucleotides at the replication origin are required for binding. Neither the hexameric sequence per se nor this sequence flanked by different sequences could support the generation of specific nucleoprotein complexes. Stoichiometry analysis indicates that each UMSBP molecule binds either of the two origin-associated sequences in the nucleoprotein complex but not both simultaneously.

Animals↗

Sequencing from compomers: using mass spectrometry for DNA de novo sequencing of 200+ nt.

One of the main endeavors in today's life science remains the efficient sequencing of long DNA molecules. Today, most de novo sequencing of DNA is still performed using the electrophoresis-based Sanger concept of 1977, in spite of certain restrictions of this method. Methods using mass spectrometry to acquire the Sanger sequencing data are limited by short sequencing lengths of 15-25 nt. We propose a new method for DNA sequencing using base-specific cleavage and mass spectrometry that appears to be a promising alternative to classical DNA sequencing approaches. A single stranded DNA or RNA molecule is cleaved by a base-specific (bio-)chemical reaction using, for example, RNAses. The cleavage reaction is modified such that not all, but only a certain percentage of bases are cleaved. The resulting mixture of fragments is then analyzed using MALDI-TOF mass spectrometry, whereby we acquire the molecular masses of fragments. For every peak in the mass spectrum, we calculate those base compositions that will potentially create a peak of the observed mass and, repeating the cleavage reaction for all four bases, finally try to uniquely reconstruct the underlying sequence from these observed spectra. This leads us to the combinatorial problem of sequencing from compomers and, finally, to the graph-theoretical problem of finding a walk in a subgraph of the de Bruijn graph. Application of this method to simulated data indicates that it might be capable of sequencing DNA molecules with 200+ nt.

Algorithms↗

The nucleotide sequence of the ubiquitous repetitive DNA sequence B1 complementary to the most abundant class of mouse fold-back RNA.

Three copies of a highly repetitive DNA sequence B1 which is complementary to the most abundant class of mouse fold-back RNA have been cloned in pBR322 plasmid and sequenced by the method of Maxam and Gilbert. All the three have a length of about 130 base pairs and are very similar in their base sequence. The deviation from the average sequence is equal to 4% and the overall mismatch between each two is not higher than 8%. One of the recombinant clones used contained two copies of B1 oriented in the same direction. All of the B1 copies are flanked with sequences which possess nonidentical but very similar structure. They consist of a number of AmCn blocks (where m varies from 2 to 8 and n equals 1-2). These peculiar sequences in all cases are separated from B1 by non-homologous DNA stretches of 2-8 residues. In one case, a long polypurine stretch is located next to such a block. It consists of 74 residues most of which represent a reiteration of the basic sequence AAAAG. We have found two regions within the B1 sequence which are homologous to the intron-exon junctions, especially to those present in the large intron of the mouse beta-globin gene. It may indicate the involvement of the B1 sequence in pre-mRNA splicing.

Animals↗

Detecting and analyzing DNA sequencing errors: toward a higher quality of the Bacillus subtilis genome sequence.

During the determination of a DNA sequence, the introduction of artifactual frameshifts and/or in-frame stop codons in putative genes can lead to misprediction of gene products. Detection of such errors with a method based on protein similarity matching is only possible when related sequences are available in databases. Here, we present a method to detect frameshift errors in DNA sequences that is based on the intrinsic properties of the coding sequences. It combines the results of two analyses, the search for translational initiation/termination sites and the prediction of coding regions. This method was used to screen the complete Bacillus subtilis genome sequence and the regions flanking putative errors were resequenced for verification. This procedure allowed us to correct the sequence and to analyze in detail the nature of the errors. Interestingly, in several cases in-frame termination codons or frameshifts were not sequencing errors but confirmed to be present in the chromosome, indicating that the genes are either nonfunctional (pseudogenes) or subject to regulatory processes such as programmed translational frameshifts. The method can be used for checking the quality of the sequences produced by any prokaryotic genome sequencing project.

Bacillus subtilis↗

Sequence organisation in nuclear DNA from Physarum polycephalum. Interspersion of repetitive and single-copy sequences.

Nuclear DNA from Physarum polycephalum is shown to contain three sequence components by reassociation kinetic analysis; a foldback component consisting of 6% of the DNA, a component with the properties of repetitive sequences comprising 31% of the DNA, and a majority component containing 63% of the DNA which reassociates with the kinetics characteristic of single-copy sequences. The complement of repetitive sequences is comprised of about 80 families of repeated elements, each containing approximately 1800 repeats per family. On average, these sequences are 6.4% richer in guanine and cytosine than total Physarum nuclear DNA. The repetitive sequences within a single family appear not to be identical, since on denaturation and annealing they give rise to collections of heteroduplexes less stable than native DNA. It is calculated that these duplexes are about 10% mismatched on average. Hydroxyapatite binding of DNA fragments of different sizes containing reassociated repeated elements demonstrates that these sequences are interspersed with single-copy sequences in a large portion of the Physarum genome. These observations are confirmed by direct examination of reassociated DNA using the electron microscope. In this manner it is shown that repetitive sequence elements possess a wide spectrum of lengths averaging 590 nucleotide residues, and they are separated by intervening segments of DNA about 930 residues in length.

Base Sequence↗

Usefulness of rpoB gene sequencing for identification of Afipia and Bosea species, including a strategy for choosing discriminative partial sequences.

Bacteria belonging to the genera Afipia and Bosea are amoeba-resisting bacteria that have been recently reported to colonize hospital water supplies and are suspected of being responsible for intensive care unit-acquired pneumonia. Identification of these bacteria is now based on determination of the 16S ribosomal DNA sequence. However, the 16S rRNA gene is not polymorphic enough to ensure discrimination of species defined by DNA-DNA relatedness. The complete rpoB sequences of 20 strains were first determined by both PCR and genome walking methods. The percentage of homology between different species ranged from 83 to 97% and was in all cases lower than that observed with the 16S rRNA gene; this was true even for species that differed in only one position. The taxonomy of Bosea and Afipia is discussed in light of these results. For strain identification that does not require the complete rpoB sequence (4,113 to 4,137 bp), we propose a simple computerized method that allows determination of nucleotide positions of high variability in the sequence that are bordered by conserved sequences and that could be useful for design of universal primers. A fragment of 740 to 752 bp that contained the most highly variable area (positions 408 to 420) was amplified and sequenced with these universal primers for 47 strains. The variability of this sequence allowed identification of all strains and correlated well with results of DNA-DNA relatedness. In the future, this method could be also used for the determination of variability "hot spots" in sets of housekeeping genes, not only for identification purposes but also for increasing the discriminatory power of sequence typing techniques such as multilocus sequence typing.

Afipia↗

Characterizatiion of rat genetic sequences of Kirsten sarcoma virus: distinct class of endogenous rat type C viral sequences.

The nucleic acid sequences found in DNA and RNA from rat cells which are homologous to Kirsten sarcoma virus have been characterized. The homologous sequences are present in multiple copies per diploid rat cellular genome in a variety of different rat cellular dna's. In certain cells that constitutively express only low levels of sequences homologous to Kirsten sarcoma virus, bromodeoxyuridine treatment leads to the expression of high levels of these sequences in RNA. Supernatants from cell lines producing the sequences homologous to Kirsten sarcoma virus contain high levels of these sequences which are purified to the same degree as the previously known rat type C viral nucleic acid sequences by type C particles being released from such cells. The results indicate that the sequences in rat cells homologous to Kisten sarcoma virus have three characteristics of known mammalian type C viruses, and suggest that at least part of Kirsten sarcoma virus rat-derived sequences represent a distinct class of endogenous rat type C virus that has no detectable homology to the other known class of endogenous rat type C virus.

Animals↗

A computer simulation analysis of the accuracy of partial genome sequencing and restriction fragment analysis in estimating genetic relationships: an application to papillomavirus DNA sequences.

BACKGROUND: Determination of genetic relatedness among microorganisms provides information necessary for making inferences regarding phylogeny. However, there is little information available on how well the genetic relationships inferred from different genotyping methods agree with true genetic relationships. In this report, two genotyping methods - restriction fragment analysis (RFA) and partial genome DNA sequencing - were each compared to complete DNA sequencing as the definitive standard for classification. RESULTS: Using the Genbank database, 16 different types or subtypes of papillomavirus were selected as study samples, because numerous complete genome sequences were available. RFA was achieved by computer-simulated digestion. The genetic similarity of samples, based on RFA, was determined from the proportion of fragments that matched in size. DNA sequences of four specific genes (E1, E6, E7, and L1), representing partial genome sequencing, were also selected for comparison to complete genome sequencing. Laboratory error was not taken into account. Evaluation of the correlation between genetic similarity matrices (Mantel's r) and comparisons of the structure of the derived dendrograms (partition metric) indicated that partial genome sequencing (for single genes) had higher agreement with complete genome sequencing, achieving a maximum Mantel's r = 0.97 and a minimum partition metric = 10. RFA had lower agreement, with a maximum Mantel's r = 0.60 and a minimum partition metric = 18. CONCLUSIONS: This simulation indicated that for smaller genomes, such as papillomavirus, partial genome sequencing is superior to restriction fragment analysis in representing genetic relatedness among isolates. The generalizability of these results to larger genomes, as well as the impact of laboratory error, remains to be demonstrated.

Animals↗

Evaluation of the performance of a p53 sequencing microarray chip using 140 previously sequenced bladder tumor samples.

BACKGROUND: Testing for mutations of the TP53 gene in tumors is a valuable predictor for disease outcome in certain cancers, but the time and cost of conventional sequencing limit its use. The present study compares traditional sequencing with the much faster microarray sequencing on a commercially available chip and describes a method to increase the specificity of the chip. METHODS: DNA from 140 human bladder tumors was extracted and subjected to a multiplex-PCR before loading onto the p53 GeneChip from Affymetrix. The same samples were previously sequenced by manual dideoxy sequencing. In addition, two cell lines with two different homozygous mutations at the TP53 gene locus were analyzed. RESULTS: Of 1464 gene chip positions, each of which corresponded to an analyzed nucleotide in the sequence, 251 had background signals that were not attributable to mutations, causing the specificity of mutation calling without mathematical correction to be low. This problem was solved by regarding each chip position as a separate entity with its own noise and threshold characteristics. The use of background plus 2 SD as the cutoff improved the specificity from 0.34 to 0.86 at the cost of a reduced sensitivity, from 0.92 to 0.84, leading to a much better concordance (92%) with results obtained by traditional sequencing. The chip method detected as little as 1% mutated DNA. CONCLUSIONS: Microarray-based sequencing is a novel option to assess TP53 mutations, representing a fast and inexpensive method compared with conventional sequencing.

Humans↗

Automated cycle sequencing with Taquenase: protocols for internal labeling, dye primer and "doublex" simultaneous sequencing.

This paper describes automated cycle sequencing protocols for internal labeling, dye primer and "doublex" simultaneous sequencing using Taquenase, a new genetically modified DNA polymerase with increased thermostability. Sequencing performance both with labeled and unlabeled primer yields uniform unambiguous signals up to the resolution limit of the sequencing gels. Primer walking with internal labeling was successfully performed on Pl-derived artificial chromosome (PAC) constructs with 130-kb inserts. Taquenase, a commercially available modified thermostable sequencing enzyme (delta 280, F667Y Taq DNA polymerase), incorporates a variety of fluorescent dNTPs carrying fluorescein isothiocyanate, TexasRed or Cy5 labels during the cycle-sequencing process with higher efficiency than other thermostable DNA polymerases. Comparison to other modified Taq DNA polymerases suggests that the particular N-terminal deletion of Taquenase rather than the presence of the F667Y mutation is responsible for the efficient incorporation and extension of labeled dNTPs. Taquenase makes feasible highly accurate "doublex" simultaneous cylce sequencing on both strands of template DNA with two internal labels or two dye-labeled primers in combination with the EMBL-2-dye DNA sequencing system, ARAKIS, or with two commercial DNA sequencers. It allows up to 2000 bases at > 99% accuracy to be determined in a single reaction.

DNA↗

Purification and partial amino acid sequence of bovine adrenal phenylethanolamine N-methyltransferase: a comparison of nucleic acid and protein sequence data.

Recently, we have reported the isolation and characterization of a putative genomic DNA clone encoding bovine adrenal phenylethanolamine N-methyltransferase (PNMT) (Batter et al., 1988). However, the lack of primary amino-acid sequence data for this enzyme precluded a definitive proof of the authenticity of this clone. To establish identity, the amino acid sequence of several peptides generated by chemical and enzymatic hydrolysis of purified PNMT was compared to that predicted from the nucleotide sequence of the exons of the putative PNMT gene. Bovine adrenomedullary PNMT was purified by ammonium sulfate precipitation, gel filtration, and ion exchange chromatography. Treatment with 70% formic acid cleaved the protein at a single Asp-Pro bond near the N-terminus. The purified protein was also cleaved at a single methionine residue near the C-terminus by treatment with cyanogen bromide. N-terminal amino acid sequence analysis identified 8 and 10 amino acid residues, respectively, following each of the scissile peptide bonds. Four tryptic peptides, generated by complete enzymatic digestion, were isolated by reverse-phase HPLC and subjected to sequence analysis. Combined, the amino acid sequences of these six peptides represent 20% of the PNMT protein. These amino acid sequences matched exactly the sequences predicted from the exons of the putative PNMT genomic clone.

Adrenal Medulla↗

On evolutionarily conserved simple repetitive DNA sequences: do "sex-specific" satellite components serve any sequence dependent function?

The nuclear genomes of eukaryotes contain DNA of varying degrees of repetition. Highly repetitious DNA and simple repetitive sequences as a fraction thereof appear to be distributed in a non-random fashion in the genome. There are arguments for and against functional roles of simple repetitive sequences, and the reasons for their evolutionary conservation are not at all clear. In order to learn more about the biologic role of simple repetitive sequences in the context of their evolutionary history, we report here the following results from studies of sex-specific snake satellite DNA: 1) The snake simple repeat sequence is 5'-GATAGACA-3' and it is strictly conserved throughout vertebrate evolution. 2) The simple repeat sequence is intimately interspersed with single-copy DNA throughout the mouse genome. 3) The simple repeat is transcribed into RNA in several animal systems and it is translatable in bacterial test systems. 4) The simple repeat sequence is sex-specifically arranged in vertebrates. 5) In snake DNA, the simple repeat is adjacent to a single-copy sequence which singles out a male-specific putative mRNA in mouse polysomal poly (A)+ RNA. Thus even if this snake simple repetitive sequence is not involved in a basic cellular function such as sex-determination, it is nevertheless a valuable tool to approach those problems.

Animals↗

The protein sequence homology of gamma-crystallins among major vertebrate classes and their DNA sequence homology to heat-shock protein genes.

A systematic characterization of lens crystallins from five major classes of vertebrates was carried out by exclusion gel filtration, cation-exchange chromatography and N-terminal sequence determination. All crystallin fractions except that of gamma-crystallin were found to be N-terminally blocked. gamma-Crystallin is present in major classes of vertebrates except the bird, showing none, or decreased amounts, of this protein in chicken and duck lenses, respectively. N-Terminal sequence analysis of the purified gamma-crystallin polypeptides showed extensive homology between different classes of vertebrates, supporting the close relatedness of this family of crystallin even from the evolutionarily distant species. Comparison of nucleotide sequences and their predicted amino acid sequences between gamma-crystallins of carp and rat lenses and heat-shock proteins demonstrated partial sequence homology of the encoded polypeptides and striking homology at the gene level. The unexpected strong homology of complementary DNA (cDNA) lies in the regions coding for 40 N-terminal residues of carp gamma-II, rat gamma 2-1, and the middle segments of 23,000- and 70,000-Mr heat-shock proteins. The optimal alignment of DNA sequences along these two segments shows about 50% homology. The percentage of protein sequence identity for the corresponding aligned segments is only 20%. The weak sequence homology at the protein level is also found between the invertebrate squid crystallin and rat gamma-crystallin polypeptides. These results pointed to the possibility of unifying three major classes of vertebrate crystallins into one alpha/beta/gamma superfamily and corroborated the previous supposition that the existing crystallins in the animal kingdom are probably mutually interrelated, sharing a common ancestry.

Amino Acid Sequence↗

Evaluation of complete genome sequences and sequences of individual gene products for the classification of hepatitis C viruses.

Comparisons of genome and polyprotein sequences of hepatitis C virus (HCV) isolates world-wide has led to the identification of nine major genotypes and many subtypes. This classification is based on either complete genome/polyprotein sequences or sequence data from the 5' noncoding region, core, E1, NS3 or NS5B genes. The relative merit of different gene segments as taxonomic markers and the validity of the resulting assignments is not clear at this stage. To resolve the taxonomy of HCV genotypes and subtypes, we have compared the complete genome and polyprotein sequences of 19 HCV isolates available in the databases as well as sequences of individual genes and gene products of these isolates. Based on the correlation between sequence relationships and taxonomic assignments of other RNA viruses, we show that the nine major genotypes of HCV represent nine distinct virus species and their subtypes subspecies. Our sequence comparison of the 5' noncoding regions and the individual gene products suggests that E2, NS2, NS5B, E1, NS4A, NS4B and NS5A (in that order) are the most appropriate regions for the discrimination between species, subspecies and strains of HCV. The 5' noncoding, core and NS3 regions are less effective in distinguishing between species, subspecies and strains. Based on a comparison of the polymerase sequence identities of HCVs, pestiviruses and flaviviruses as well as the recent information on the size and morphology of HCV virions, we propose that HCVs, pestiviruses and flaviviruses should be classified into three separate families, named Hepciviridae, Pestiviridae and Flaviviridae, respectively rather than three genera of the Flaviviridae as currently classified. We also propose "Hepcivirus" as the genus name for HCVs.

Base Sequence↗

Rat liver NAD(P)H:quinone reductase: isolation of a quinone reductase structural gene and prediction of the NH2 terminal sequence of the protein by double-stranded sequencing of exons 1 and 2.

Recently two reports [J. A. Robertson et al. (1986) J. Biol. Chem. 261, 15794-15799 and R. M. Bayney et al. (1987) J. Biol. Chem. 262, 572-575] have appeared concerning the nucleotide sequence of quinone reductase cDNA clones. Although the cDNA clones are virtually identical, they diverge in the 5' region that encodes the NH2 terminus of the protein. In order to clarify the sequence of this region, we have isolated quinone reductase clones from a rat genomic library using a cDNA clone, pDTD55, isolated and characterized by our laboratory. We have determined the sequence of exons 1 and 2 of the structural gene by double-stranded sequencing using oligonucleotide primers. The sequence of exons 1 and 2 of the quinone reductase structural gene along with our previous nucleotide sequence analysis of pDTD55 as well as conventional amino acid sequence analysis of the purified protein indicates that quinone reductase is composed of 274 amino acids with a molecular weight of 30,946. These data agree with the published sequence of lambda NMOR1 reported by Robertson et al.

Amino Acid Sequence↗

cDNA sequence of zebrafish (Brachydanio rerio) translation elongation factor-1 alpha: molecular phylogeny of eukaryotes based on elongation factor-1 alpha protein sequences.

We have isolated and determined the nucleotide sequence of a cDNA clone containing the complete coding region for elongation factor-1 alpha (EF-1 alpha) from an embryonic zebrafish cDNA library. A secondary structure model based on all known EF-1 alpha and EF-Tu protein sequences is presented and the presence of conserved putative protein kinase C phosphorylation sites in loop regions of eukaryotic EF-1 alpha is demonstrated. Using distance matrix and maximum parsimony methods we constructed multi-kingdom phylogenetic trees containing 22 different eukaryotic sequences. Strikingly, both tree constructions show Fungi to be the closest relative of Animalia among eukaryotic kingdoms. A 12 amino acid stretch present in all animal and fungal sequences known to date was found to be absent from all plant, protist an archaebacterial EF-1 alpha sequences suggesting that this sequence was inserted following the separation of plants from the lineage leading to fungi and animals. In contrast to our results, molecular phylogenies based on small subunit ribosomal RNA sequences as well as other protein sequences have failed to yield consistent results regarding the branching order among the kingdoms Plantae, Fungi and Animalia. The slow evolutionary rate and universal occurrence of EF-1 alpha (EF-Tu in eubacteria) makes this protein a particularly interesting tool for probing distant evolutionary relationships.

Animals↗