Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Molecular cloning and nucleotide sequence of rat brain argininosuccinate lyase cDNA with an extremely long 5'-untranslated sequence: evidence for the identity of the brain and liver enzymes.

Argininosuccinate lyase (EC 4.3.2.1) is an enzyme present in the brain of ureotelic animals. Using as a probe rat liver argininosuccinate lyase cDNA, already isolated and sequenced (Amaya, Y., Matsubasa, T., Takiguchi, M., Kobayashi, K., Saheki, T., Kawamoto, S. and Mori, M., J. Biochem., 103 (1988) 177-181), we screened a rat brain cDNA library constructed in the lambda gt11 expression vector and obtained a single cDNA clone. This cDNA clone contained an open reading frame encoding a polypeptide of 461 amino acid residues (predicted Mr = 51,390), a 5'-untranslated sequence of 967 bp and a 3'-untranslated sequence of 74 bp. The length of the 5'-non-coding region of the cDNA seems to be one of the longest among the cDNAs heretofore isolated. A comparison of the brain cDNA sequence (2424 bp) with the corresponding region of the liver cDNA (1574 bp) revealed differences in 5 nucleotides. The brain clone contained A----G and C----G base differences from the hepatic sequence, resulting in amino acid changes from Tyr and Arg in the liver clone, to Cys and Gly in the brain clone, respectively. The other 3 nucleotide differences are silent with respect to the amino acid sequence of the protein. Therefore, the amino acid sequence of the brain argininosuccinate lyase, as deduced from the nucleotide sequence of its cDNA clone, was identical with that of the liver protein, except for two amino acid residues. These minor changes may reflect a microheterogeneity of the argininosuccinate lyase gene. The brain and liver enzymes seem to be encoded by the same structural gene.

Amino Acid Sequence↗

Sequence confirmation of synthetic phosphorothioate oligodeoxynucleotides using Sanger sequencing reactions in combination with mass spectrometry.

A protocol relying on Sanger sequencing reactions in combination with mass spectrometry (MS) for sequence confirmation of antisense phosphorothioate oligodeoxynucleotides is described. In this procedure, synthetic phosphorothioate oligodeoxynucleotides are used as reverse primers for extension of matched templates with enough length (approximately 150-300 bp) for well-established Sanger sequencing. Because the complementary strand of modified primer is used directly for sequencing primer extension, the base order shown in the sequencing result is reversely complementary to phosphorothioate oligodeoxynucleotide. This sequencing method can be applied not only to phosphorothioate oligodeoxynucleotides with different lengths (13-21 mer) and base composition but also to sequences with bases' switch, deletion, or insertion. In addition, modified primers incorporate the 5' end of polymerase chain reaction (PCR) products conveying the characters of phosphorothioate modification. The method requires only common reagents and instruments and so is better suited to routine sequence analysis in quality control of phosphorothioate antisense drugs.

Base Sequence↗

De novo sequencing, peptide composition analysis, and composition-based sequencing: a new strategy employing accurate mass determination by fourier transform ion cyclotron resonance mass spectrometry.

A new strategy is described for the determination of amino acid sequences of unknown peptides. Different from the well-known but often inefficient de novo sequencing approach, the new method is based on a two-step process. In the first step the amino acid composition of an unknown peptide is determined on the basis of accurate mass values of the peptide precursor ion and a small number of accurate fragment ion mass values, and, as in de novo sequencing, without employing protein database information or other pre-information. In the second step the sequence of the found amino acids of the peptide is determined by scoring the agreement between expected and observed fragment ion signals of the permuted sequences. It was found that the new approach is highly efficient if accurate mass values are available and that it easily outstrips common approaches of de novo sequencing being based on lower accuracies and detailed knowledge of fragmentation behavior. Simple permutation and calculation of all possible amino acid sequences, however, is only efficient if the composition is known or if possible compositions are at least reduced to a small list. The latter requires the highest possible instrumental mass accuracy, which is currently provided only by fourier transform ion cyclotron resonance mass spectrometry. The connection between mass accuracy and peptide composition variability is described and an example of peptide compositioning and composition-based sequencing is presented.

Amino Acid Sequence↗

Highly conserved amino-acid sequence between murine STAT3 and a revised human STAT3 sequence.

Signal transducer and activator of transcription 3 (STAT3) is an important mediator of cytokine signaling, whose cDNA and protein sequences have been fully characterized. We sequenced the whole human STAT3 cDNA isolated from HepG2 cells. The new sequence determined contains 43 nucleotide changes overall, corresponding to six modifications at the amino-acid level. The revised amino-acid sequence of human STAT3 is now completely identical to the mouse sequence, except for a single amino-acid change at position 760. Thus STAT3 now results as one of the most evolutionarily conserved among known proteins. By using specific RT-PCR we could discriminate between the original sequence and the new variant. Amplification of regions within the src-homology domain 2 (SH2) of STAT3, from the RNAs of 11 different tissues or cells, revealed only the expression of the new SH2 variant. Besides, only this SH2 variant was amplified from human genomic DNA. We conclude that the new sequence we have determined in this study represents a revised sequence of hSTAT3 or, less likely, a new predominant allele.

Amino Acid Sequence↗

cDNA nucleotide sequence encoding the ZPC protein of Australian hydromyine rodents: a novel sequence of the putative sperm-combining site within the family Muridae.

This comparative study of the cDNA sequence of the zona pellucida C (ZPC) glycoprotein in murid rodents focuses on the nucleotide and amino acid sequence of the putative sperm-combining site. We ask the question: Has divergence evolved in the nucleotide sequence of ZPC in the murid rodents of Australia? Using RT-PCR and (RACE) PCR, the complete cDNA coding region of ZPC in the Australian hydromyine rodents Notomys alexis and Pseudomys australis, and a partial cDNA sequence from a third hydromyine rodent, Hydromys chrysogaster, has been determined. Comparison between the cDNA sequences of the hydromyine rodents reveals that the level of amino acid sequence identity between N. alexis and P. australis is 96%, whereas that between the two species of hydromyine rodents and M. musculus and R. norvegicus is 88% and 87% respectively. Despite being reproductively isolated from each other, the three species of hydromyine rodents have a 100% level of amino acid sequence identity at the putative sperm-combining site. This finding does not support the view that this site is under positive selective pressure. The sequence data obtained in this study may have important conservation implications for the dissemination of immunocontraception directed against M. musculus using ZPC antibodies.

Amino Acid Sequence↗

Evolutionary and genetic implications of sequence variation in two nonallelic HLA-DR beta-chain cDNA sequences.

Most HLA haplotypes carry two expressed DR beta-chain genes; in the DR4 haplotype, the polymorphic locus has been called DR beta 1 and the apparently nonpolymorphic locus has been called DR beta 2. We have isolated nearly full-length DR beta-chain cDNA clones representing each of these two loci from a cell line homozygous for DR4 and Dw4. The clones have been sequenced and the sequences compared with published DR beta cDNA sequences derived from other haplotypes. A comparison of our sequences with other published cDNA sequences did not allow assignment of these other sequences to either the beta 1 or beta 2 locus. Comparison of our DR4 beta 1 sequence with DR beta 1 sequences isolated from other DR4-positive cells suggests that the alleles of DR4 beta 1 may have recently diverged from a common ancestor. The apparent lack of polymorphism of DR beta 2 may in part be a reflection of this recent divergence.

Alleles↗

Amino acid sequence of S-adenosyl-L-homocysteine hydrolase from rat liver as derived from the cDNA sequence.

Rat liver cDNA libraries constructed in lambda gt11 were screened for reactivity with polyclonal antibodies to native S-adenosyl-L-homocysteine (AdoHcy) hydrolase (adenosylhomocysteinase; EC 3.3.1.1). Five clones were isolated and sequenced. The amino acid sequence, deduced from the cDNA sequence, contained the sequence of eight peptides obtained by tryptic and cyanogen bromide fragmentation of rat liver AdoHcy hydrolase. Identification of the amino- and carboxyl-terminal peptides in the amino acid sequence showed that the complete sequence was obtained. A "fingerprint" sequence was found that is characteristic of dinucleotide-binding domains of many proteins. For AdoHcy hydrolase, this region from the lysine at position 213 to the aspartate at position 244, containing the sequence Gly-Xaa-Gly-Xaa-Xaa-Gly at positions 219-224, is presumably the site of binding for NAD+, which is required for the activity of the enzyme.

Adenosylhomocysteinase↗

Direct cloning of a target gene from a pool of homologous sequences: complete cDNA sequence of a weak neurotoxin from cobra Naja kaouthia.

Selective cloning of the cDNA coding for a weak neurotoxin (WTX) from cobra N. kaouthia including the 5'- and 3'-non-translated regions (NTR) is described. The known amino acid sequence of WTX was used together with the nucleotide sequence of a weak neurotoxin NNAM2 from cobra Naja atra, to design WTX-specific primers for direct amplification of an internal WTX cDNA fragment by RT- PCR. The sequence of the complete WTX cDNA was determined in sequencing runs on internal PCR products, cloned 3'- and 5'-RACE-fragments and several full-length cDNA clones. The cDNA coding sequence is in excellent agreement with the previously determined WTX amino acid sequence, has a high homology with other known weak toxin cDNAs, whereas even higher homology (up to 96%) with several classes of 3-finger toxins was detected in the 59 bp 3'-NTR consensus sequence. A possible function of the highly conserved nucleotide sequence elements is discussed.

Amino Acid Sequence↗

Nucleotide sequence and deduced amino acid sequence of Escherichia coli pyruvate oxidase, a lipid-activated flavoprotein.

The entire nucleotide sequence of the poxB (pyruvate oxidase) gene of Escherichia coli K-12 has been determined by the dideoxynucleotide (Sanger) sequencing of fragments of the gene cloned into a phage M13 vector. The gene is 1716 nucleotides in length and has an open reading frame which encodes a protein of Mr 62,018. This open reading frame was shown to encode pyruvate oxidase by alignment of the amino acid sequences deduced for the amino and carboxy termini and several internal segments of the mature protein with sequences obtained by amino acid sequence analysis. The deduced amino acid sequence of the oxidase was not unusually rich in hydrophobic sequences despite the peripheral membrane location and lipid binding properties of the protein. The codon usage of the oxidase gene was typical of a moderately expressed protein. The deduced amino acid sequence shares homology with the large subunits of the acetohydroxy acid synthase isozymes I, II, and III, encoded by the ilvB, ilvG, and ilvI genes of E. coli.

Acetolactate Synthase↗

Sequence logos: a new way to display consensus sequences.

A graphical method is presented for displaying the patterns in a set of aligned sequences. The characters representing the sequence are stacked on top of each other for each position in the aligned sequences. The height of each letter is made proportional to its frequency, and the letters are sorted so the most common one is on top. The height of the entire stack is then adjusted to signify the information content of the sequences at that position. From these 'sequence logos', one can determine not only the consensus sequence but also the relative frequency of bases and the information content (measured in bits) at every position in a site or sequence. The logo displays both significant residues and subtle sequence patterns.

Amino Acid Sequence↗

Repeated sequence sets in mitochondrial DNA molecules of root knot nematodes (Meloidogyne): nucleotide sequences, genome location and potential for host-race identification.

Within a 7 kb segment of the mtDNA molecule of the root knot nematode, Meloidogyne javanica, that lacks standard mitochondrial genes, are three sets of strictly tandemly arranged, direct repeat sequences: approximately 36 copies of a 102 ntp sequence that contains a TaqI site; 11 copies of a 63 ntp sequence, and 5 copies of an 8 ntp sequence. The 7 kb repeat-containing segment is bounded by putative tRNAasp and tRNAf-met genes and the arrangement of sequences within this segment is: the tRNAasp gene; a unique 1,528 ntp segment that contains two highly stable hairpin-forming sequences; the 102 ntp repeat set; the 8 ntp repeat set; a unique 1,068 ntp segment; the 63 ntp repeat set; and the tRNAf-met gene. The nucleotide sequences of the 102 ntp copies and the 63 ntp copies have been conserved among the species examined. Data from Southern hybridization experiments indicate that 102 ntp and 63 ntp repeats occur in the mtDNAs of three, two and two races of M.incognita, M.hapla and M.arenaria, respectively. Nucleotide sequences of the M.incognita Race-3 102 ntp repeat were found to be either identical or highly similar to those of the M.javanica 102 ntp repeat. Differences in migration distance and number of 102 ntp repeat-containing bands seen in Southern hybridization autoradiographs of restriction-digested mtDNAs of M.javanica and the different host races of M.incognita, M.hapla and M.arenaria are sufficient to distinguish the different host races of each species.

Animals↗

A subtelomeric DNA sequence is required for correct processing of the macronuclear DNA sequences during macronuclear development in the hypotrichous ciliate Stylonychia lemnae.

During macronuclear differentiation in ciliated protozoa a series of programed DNA reorganization processes occur. These include the elimination of micronuclear-specific DNA sequences, the specific fragmentation of the genome into small gene-sized DNA molecules, the de novo addition of telomeric sequences to these DNA molecules and the specific amplification of the remaining DNA molecules. Recently we constructed a vector containing the modified micronuclear version of macronuclear destined DNA sequences that was correctly fragmented and telomeres were added de novo after injection into the developing macronucleus. It therefore must contain all the cis- acting sequences required for these processes. We made a series of vectors deleting different sequences from the original vector. It could be shown that at least in the case studied here no micronuclear-specific sequences are required for specific fragmentation of the genome and telomere addition. However, a short subtelomeric sequence at the 3[prime]-end is essential for these processes, whereas no specific cut seems to occur at the 5[prime]-end. In addition, we can show that the processing activity is restricted to a short period of time during macronuclear differentiation and that a preceding transcription is required for correct processing of macronuclear-destined DNA sequences. Possible mechanisms of these processes will be discussed.

Animals↗

Whole genome sequence-enabled prediction of sequences performed for random PCR products of Escherichia coli.

The sequence of an unknown PCR product generated by random (and conventional) PCR could be determined without sequencing when it is provided with the template DNA sequence. Theoretically, this was based on formerly established ideas which assert that the amount of random PCR product mainly depends on the stability of the primer-binding structures and that the dynamic solution structure of DNA is essentially governed by the Watson-Crick base pairing. However, it has not been clear whether this holds true for larger genomes of mega- to gigabase size, beside the lambda phage genome (of 50 kb) used previously, nor has it been ascertained to uniquely specify the sequence of a random PCR product. Here, we jointly use two computer programs together with experimental data from Genome Profiling (i.e. TGGE analysis of random PCR products). The first procedure carried out by a newly remodeled computer program (PCRAna-A1) was shown to be competent to calculate a set of random PCR products from Escherichia coli genome DNA (4.7 Mb). The other procedure performed with another program (Poland-H) played a critical role in determining the final candidate sequence by theoretically offering the initial melting temperature and the melting pattern of unspecified candidate sequences. The success attained here not only proved our method to be useful for sequence prediction but also confirmed the above-mentioned ideas as rational. We believe that this is the first case to computer-utilize a genome sequence as a whole.

Base Sequence↗

Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.

The similarity of two nucleotide sequences is often expressed in terms of evolutionary distance, a measure of the amount of change needed to transform one sequence into the other. Given two sequences with a small distance between them, can their similarity be explained by their base composition alone? The nucleotide order of these sequences contributes to their similarity if the distance is much smaller than their average permutation distance, which is obtained by calculating the distances for many random permutations of these sequences. To determine whether their similarity can be explained by their dinucleotide and codon usage, random sequences must be chosen from the set of permuted sequences that preserve dinucleotide and codon usage. The problem of choosing random dinucleotide and codon-preserving permutations can be expressed in the language of graph theory as the problem of generating random Eulerian walks on a directed multigraph. An efficient algorithm for generating such walks is described. This algorithm can be used to choose random sequence permutations that preserve (1) dinucleotide usage, (2) dinucleotide and trinucleotide usage, or (3) dinucleotide and codon usage. For example, the similarity of two 60-nucleotide DNA segments from the human beta-1 interferon gene (nucleotides 196-255 and 499-558) is not just the result of their nonrandom dinucleotide and codon usage.

Base Sequence↗

Nucleotide sequence and transcription of the fbc operon from Rhodopseudomonas sphaeroides. Evaluation of the deduced amino acid sequences of the FeS protein, cytochrome b and cytochrome c1.

The fbc operon from Rhodopseudomonas sphaeroides encodes the three redox carriers of the ubiquinol-cytochrome-c reductase (b/c1 complex): FeS protein, cytochrome b and cytochrome c1 [Gabellini, N. et al. (1985) EMBO J.2, 549-553]. The nucleotide sequence of 3874 bp of cloned R. sphaeroides chromosomal DNA, including the three structural genes fbcF, fbcB and fbcC has been determined. The reading frames of the fbc genes could be identified readily since the encoded amino acid sequences are highly homologous with the sequences of the corresponding mitochondrial polypeptides. Initiation and termination points for transcription have been investigated by S1 nuclease protection analysis. The transcription of the fbc operon starts approximately 240 base pairs upstream from the start codon of the fbcF gene and terminates 120 base pairs downstream from the stop codon of the fbcC gene. Nucleotide sequences resembling recognition signals for the binding and release of the RNA polymerase were identified. The N-terminal amino acid sequence of the mature cytochrome c1 was obtained by automated Edman degradation of the isolated subunit, confirming the fbcC reading frame and indicating that the bacterial preapocytochrome c1 has a transient leader sequence including 21 residues. The N-terminal sequence of one hydrophilic peptide of the FeS protein has been also obtained confirming the fbcF reading frame. The deduced amino acid sequences are discussed in relation to the known primary structures of the homologous proteins from mitochondria and chloroplasts. The primary structures of the polypeptides are evaluated with respect to their topology in the membrane, their biogenesis, the structure of the catalytic sites and subunit interactions.

Amino Acid Sequence↗

Herpes simplex virus type 1 variant a sequence generated by recombination and breakage of the a sequence in defined regions, including the one involved in recombination.

A herpes simplex virus type 1 clone, GN29, having exclusively the variant a sequence was isolated. This a sequence was composed of unique (U) and directly repeated (DR) elements DR1, Ub, (DR2)14, Ucd, Ubd, (DR2)5, DR4n2, and Uc and was assumed to be generated by recombination between sites in Ub and Uc. Unusual DNA fragments containing parts of the a sequence, present in the DNA preparations of GN29, were molecularly cloned. Almost all termini of the cloned unusual DNA fragments were situated in defined regions assumed to be recombinogenic: (i) a site in the inverted repeat of the L component, (ii) DR1, (iii) DR2, (iv) the DR4 stretch, and (v) the novel recombination stretch in the variant a sequence of GN29. The termini of unusual DNA fragments, possibly produced by strand breaks, can serve as free DNA ends to initiate recombination of the a sequence. These results support the model of double-strand-break repair for recombination of the a sequence. Sequence-specific enhancement of the recombination of the a sequence probably depends on the presence of recombinogenic elements apt to break, such as DR2 repeats and the DR4 stretch.

Base Sequence↗

The complete sequence of the mouse skeletal alpha-actin gene reveals several conserved and inverted repeat sequences outside of the protein-coding region.

The complete nucleotide sequence of a genomic clone encoding the mouse skeletal alpha-actin gene has been determined. This single-copy gene codes for a protein identical in primary sequence to the rabbit skeletal alpha-actin. It has a large intron in the 5'-untranslated region 12 nucleotides upstream from the initiator ATG and five small introns in the coding region at codons specifying amino acids 41/42, 150, 204, 267, and 327/328. These intron positions are identical to those for the corresponding genes of chickens and rats. Similar to other skeletal alpha-actin genes, the nucleotide sequence codes for two amino acids, Met-Cys, preceding the known N-terminal Asp of the mature protein. Comparison of the nucleotide sequences of rat, mouse, chicken, and human skeletal muscle alpha-actin genes reveals conserved sequences (some not previously noted) outside of the protein-coding region. Furthermore, several inverted repeat sequences, partially within these conserved regions, have been identified. These sequences are not present in the vertebrate cytoskeletal beta-actin genes. The strong conservation of the inverted repeat sequences suggests that they may have a role in the tissue-specific expression of skeletal alpha-actin genes.

Actins↗

Highly informative nature of inter simple sequence repeat (ISSR) sequences amplified using tri- and tetra-nucleotide primers from DNA of cauliflower (Brassica oleracea var. botrytis L.).

Inter simple sequence repeat (ISSR) sequences as molecular markers can lead to the detection of polymorphism and also be a new approach to the study of SSR distribution and frequency. In this study, ISSR amplification with nonanchored primer was performed in closely related cauliflower lines. Fourty-four different amplified fragments were sequenced. Sequences of PCR products are delimited by the expected motifs and number of repeats, which validates the ISSR nonanchored primer amplification technique. DNA and amino acids homology search between internal sequences and databases (i) show that the majority of the internal regions of ISSR had homologies with known sequences, mainly with genes coding for proteins implicated in DNA interaction or gene expression, which reflected the significance of amplified ISSR sequences and (ii) display long and numerous homologies with the Arabidopsis thaliana genome. ISSR amplifications revealed a high conservation of these sequences between Arabidopsis thaliana and Brassica oleracea var. botrytis. Thirty-four of the 44 ISSRs had one or several perfect or imperfect internal microsatellites. Such distribution indicates the presence in genomes of highly concentrated regions of SSR, or "SSR hot spots." Among the four nonanchored primers used in this study, trinucleotide repeats, and especially (CAA)5, were the most powerful primers for ISSR amplifications regarding the number of amplified bands, level of polymorphism, and their nature.

Base Sequence↗