Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Comparative sequence analysis of the PRKAG3 region between human and pig: evolution of repetitive sequences and potential new exons.

The PRKAG3 gene encodes the gamma3 chain of AMP-activated protein kinase (AMPK). A non-conservative missense mutation in the PRKAG3 gene causes a dominant phenotype involving abnormally high glycogen content in pig skeletal muscle. We have determined >126 kb (in 13 contigs) of porcine genomic sequence surrounding the PRKAG3 gene and the corresponding mouse region covering the gene. A comparison of these PRKAG3 sequences and the human sequence was conducted and used to predict evolutionarily conserved regions, including regulatory regions. A comparison of the human genomic sequence and a porcine BAC sequence containing the PRKAG3 gene, revealed a conserved organization and the presence of three additional genes, CYP27A1 (cytochrome P450, family 27, subfamily A, polypeptide 1), STK36 (Serine Threonine Kinase 36), and the homolog of the unidentified human mRNA KIAA0173. Interspersed repetitive elements constituted 51.4 and 38.6% of this genomic region in human and pig, respectively. We were able to reliably align 12.6 kb of orthologous repeats shared between pig and human and these showed an average sequence identity of 72.4%. Our analysis revealed that the human KIAA0173 gene harbors alternative 5' untranslated exons originating from repetitive elements. This provides an obvious example how transposable elements may affect gene evolution.

5' Untranslated Regions↗

Determination of the amino acid sequence of rabbit, human, and wheat germ protein synthesis factor eIF-4C by cloning and chemical sequencing.

The small eukaryotic initiation factor (eIF)-4C is implicated in the initiation pathway, where it enhances ribosome dissociation into subunits and stabilizes the binding of the initiator Met-tRNA(i) to 40 S ribosomal subunits. In order to elucidate the function of eIF-4C, its structure has been further characterized. The amino acid sequence of many peptides from rabbit reticulocyte and wheat germ eIF-4C have been determined chemically. From the chemical sequencing of the rabbit protein, it was noted that at least two different eIF-4C molecules were present which differed by conservative substitutions at three positions (2 aspartic acid for glutamic acid switches and 1 valine for isoleucine switch). By the use of unique sequences with low codon degeneracy, primers were used to obtain a polymerase chain reaction product of appropriate size and sequence. This product was then used to isolate full-length coding sequence cDNA clones for human eIF-4C. A similar strategy was used to design PCR primers and then isolate a wheat cDNA clone which lacked the coding region for the first 23 amino acids, but contained a complete 3'-untranslated region. The protein amino acid sequence of wheat germ eIF-4C is 68% identical with the mammalian protein, and, allowing for the most conservative substitutions, the proteins are 76% similar. Both the mammalian and wheat germ proteins are 143 amino acids in length and have molecular weights of about 16,400. A unique feature of eIF-4C is its apparent "polarity" as 9 of the first 15 amino acids are basic while 13 of the last 20 amino acids are acidic. This dipole nature may enable the protein to interact with both the ribosome (perhaps via the rRNA) and other translation initiation factors.

Amino Acid Sequence↗

C4d DNA sequences of two infrequent human allotypes (C4A13 and C4B12) and the presence of signal sequences enhancing recombination.

The DNA sequences of the polymorphic region (C4d) that belong to the infrequent complement C4 allotypes C4A13 and C4B12 have been obtained. In addition, C4A4 and C4B2 C4d sequences have been completed. C4A13 shows a new combination of amino acids at the following polymorphic positions: Asp1054, Pro1101 Cys1102, Leu1105, Asp1106, Asn1157, Ala1188, and Arg1191. These amino acids conform to the antigenic determinants Chido 1 and Rodgers 3; thus C4A13 is the only allele described thus far that carries both Ags. C4A13 and C4A4 carry the motif "ggctc*" (* means "deletion") at positions 14 to 19 in their intron 28; this motif had previously been reported only in C4B alleles. The C4B12 nucleotide sequence is analogous to C4B1b and C4B3 sequences, except for codon 1076, which is GCC in C4B1b and C4B3 and GGA in C4B12, which is coding for glycine in both cases. A recombination model for the generation of C4 alleles is formulated based on the analysis of these new sequences. One recombination would take place between positions 1157 and 1186 and would give rise to C4A13 and C4B5 or C4A3 (or C4A6) and C4B2; another one would occur between positions 1054 and 1076 and would generate C4A3 (or C4A6) and C4B12 or C4A2 and C4Bnew. Analysis of 1157 to 1186 and 1054 to 1076 fragments reveals the presence of putative sequence signals for recombination (similar to Escherichia coli chi recombination signal); the accumulation of such signals in fragments 1054 to 1076 supports the notion that a recombination hot spot for the C4 gene may exist and it also enhances new allele generation and intraspecies C4 gene homogenization.

Alleles↗

The GOR47-1 sequence in human DNA encoding for a potential autoantigen in connection with hepatitis C--a sequence not only reserved for humans.

The sequence 'GOR47-1' is a consistent part of human DNA; the expressed polypeptide of it 'GOR' is accepted to be an autoantigen, and the anti-GOR an autoantibody. However, GOR47-1 was originally isolated through a cDNA clone from blood of a chimpanzee. This animal belonged to a series of chimpanzees, in which human plasma of a patient with non-A, non-B hepatitis had been passaged. To date, nothing is known how it is that this 'sequence GOR47-1' without recognizable self-replicating properties and allocated to the human genome could be isolated from a chimpanzee plasma. The aim of this study was to detect by polymerase chain reaction GOR47-1 sequences in healthy, anti-HCV-negative humans, HCV-positive patients, chimpanzee, snake, and in maize and tobacco plants. The GOR47-1 sequence is present not only in human DNA but also with a high degree of homology in chimpanzee DNA. Essential parts of this sequence are also present in DNA of a snake and the two plants listed above. Our findings reveal that the GOR47-1 sequence isolated from a chimpanzee was probably of the chimpanzee origin. This fact has not yet been considered up until now, when discussing the role of GOR/anti-GOR in humans particularly suffering from chronic hepatitis C.

Animals↗

Regulated expression of repetitive sequences including the identifier sequence during myotube formation in culture.

We have isolated and characterized a cDNA of 1183 bp, pL6-411, from rat L6 muscle cells. This cDNA contains repetitive sequences - including two inverted copies of the previously described identifier sequence - as shown by sequence analysis. Repetitive sequences from pL6-411 characterize a family of RNAs which is specifically induced during L6 myotube formation. Another part of the pL6-411 sequence, existing at low-copy number per haploid rat genome, hybridized to two RNAs of 5 kb and 2 kb from L6 myoblasts as well as from L6 myotubes. A third pL6-411-related RNA of 150 bases was detected which hybridized with the repetitive sequence but did not hybridize with the low-copy number part of pL6-411. It appears that the 'identifier' sequence in this population of small RNAs is complementary to one of the 'identifier' copies in the pL6-411-related RNA. Finally, we identified on cDNA pL6-411 the recognition site for the TGGCA-binding protein and in both orientations a total of four putative promoters for RNA polymerase III.

Animals↗

The human major histocompatibility complex: 42,221 bp of genomic sequence, high-density sequence-tagged site map, evolution, and polymorphism for HLA class I.

We report the isolation and characterization of newly identified yeast artificial chromosome (YAC) and bacterial artificial chromosome (BAC) clones spanning the HLA class I region between HLA-C and HLA-E and of YACs extending telomeric of HLA-F. When included with previously characterized HLA class I YACs, a contiguous stretch of over 2.4 Mb pairs including the entire class I region has been isolated as a series of overlapping YAC and BAC clones. Evidence that the cloned DNA faithfully represents the source genomic DNA was obtained by extensive characterization of the YACs and by independent isolation of two or more overlapping YACs or BACs spanning the entire region. As a result of this work, over 80 unique sequence probes were identified, the majority of which were sequenced to yield 42,221 bp of new major histocompatibility complex (MHC)-derived sequence. Some of these data were reduced to sequenced tagged site primer sets, facilitating the isolation of all or nearly all of HLA class I from a variety of genomic libraries. The sequence data were analyzed for protein coding capacity and homology to existing expressed tagged sites and tested for conservation of sequences in other mammalian genomes. These results indicated that large portions of the HLA class I region are conserved among mammals. Measurements of polymorphism within non-HLA class I loci generated additional data pointing toward information of potential relevance to MHC-associated diseases. The combined data and clones presented here set the stage for the determination of the complete nucleotide sequence of HLA class I.

Animals↗

Human DNA sequences isolated with an immunoglobulin switch region probe: sequence, chromosomal localization, and restriction fragment length polymorphisms.

We have screened a human genomic DNA library with an immunoglobulin (Ig) derived switch (S) region specific probe for homologous sequences. Five Ig independent phage clones were isolated and characterized. The S sequence homologous DNA fragments are short compared to the S region sequences. Ig independent S sequences are flanked by highly repetitive DNA elements and perfect inverted repeats can be demonstrated in their close vicinity. Using subclones of S homologous sequences restriction fragment length polymorphisms were shown within DNA of different T cell leukemias, Burkitt lymphomas, lymphoblastoid cell lines, and DNA of healthy individuals. One of the five clones isolated with the S region probe was evidently localized to chromosome 2 and/or 10 and showed a complex hybridisation pattern with several different human DNAs. S homologous sequences of another clone are most likely localized on chromosome 1. It is possible that these Ig independent S sequences have arisen by amplification and transposition and that they are involved in genetic recombination.

Base Sequence↗

Phylogenetic analysis of Oryza species, based on simple sequence repeats and their flanking nucleotide sequences from the mitochondrial and chloroplast genomes.

Simple sequence repeats (SSR) and their flanking regions in the mitochondrial and chloroplast genomes were sequenced in order to reveal DNA sequence variation. This information was used to gain new insights into phylogenetic relationships among species in the genus Oryza. Seven mitochondrial and five chloroplast SSR loci equal to or longer than ten mononucleotide repeats were chosen from known rice mitochondrial and chloroplast genome sequences. A total of 50 accessions of Oryza that represented six different diploid genomes and three different allopolyploid genomes of Oryza species were analyzed. Many base substitutions and deletions/insertions were identified in the SSR loci as well as their flanking regions. Of mononucleotide SSR, G (or C) repeats were more variable than A (or T) repeats. Results obtained by chloroplast and mitochondrial SSR analyses showed similar phylogenetic relationships among species, although chloroplast SSR were more informative because of their higher sequence diversity. The CC genome is suggested to be the maternal parent for the two BBCC genome species (O. punctata and O. minuta) and the CCDD species O. latifolia, based on the high level of sequence conservation between the diploid CC genome species and these allotetraploid species. This is the first report of phylogenetic analysis among plant species, based on mitochondrial and chloroplast SSR and their flanking sequences.

Base Sequence↗

Occurrence of reiterated sequences in an untranslated region of Simian virus 40 DNA determined by nucleotide sequence analysis.

An earlier report (Subramanian, Dhar, and Weissman, 1977c) presented the nucleotide sequence of Eco RII-G fragment of SV40 DNA, which contains the origin of DNA replication. The nucleotide sequence of Eco RII-N fragment located next to Eco RII-G on the physical map of SV40 DNA is presented in this report. Eco RII-N is found to be a tandem duplication of the last 55 nucleotides of Eco RII-G. This tandem repeat is immediately preceded by two other reiterated sequences occurring within Eco RII-G, one of them being a tandem repeat of 21 nucleotides and the other a nontandem repeat of 10 nucleotides. These repetitive sequences occur in close proximity to the origin of DNA replication which is known to contain other specialized sequences such as a few palindromes (one of which is 27 long and possesses a perfect 2-fold axis of symmetry), one "true" palindrome, and a long A/T-rich cluster. The repeats (and the replication origin) occur within an untranslated region of SV40 DNA flanked by (the few) structural genes coding for the "late" proteins on the one side and that (those) coding for the "early" protein(s) on the other side. The reiterated sequences are comparable in some respects to repetitive sequences occurring in eucaryotic DNAs. Possible biological functions of the repeats are discussed.

Base Sequence↗

Sequence of inverted terminal repetitions from different adenoviruses: demonstration of conserved sequences and homology between SA7 termini and SV40 DNA.

We have established the nucleotide sequence for the inverted terminal repetition of human adenovirus type 3, a subgroup B adenovirus. The repetition, which is 136 bp long, shows a high degree of homology with the known sequence for the inverted repetition of adenovirus type 5 (Steenbergh et al., 1977) a subgroup C adenovirus. Partial sequence information convering 120 bp of the inverted terminal repetitions of human serotype 12, a subgroup A member, and of simian adenovirus type 7 has also been obtained. A comparison of the established sequences shows that the terminal repetitions, in particular the first 50 bp from the ends, contain sequences that have been well conserved in adenovirus evolution. For instance, only six mismatched base pairs were detected among the first 50 bp in the repetitions of simian adenovirus type 7 and human adenovirus type 5, although the homology between simian adenovirus 7 and human subgroup C adenoviruses was estimated to be only 30%. A 14 bp sequence located 9-22 nucleotides from the ends is present in DNAs from all the human serotypes examined as well as in simian adenovirus 7 DNA. Furthermore, the simian adenovirus 7 repetition contains a 21 bp sequence which is present in SV40 DNA, close to the origin of DNA replication.

Adenoviridae↗

The study of DNA sequences by their sequences of twist angles.

To study the properties of DNA sequences we have transformed the sequences of bases into the sequences of twist angles along the chain of DNA double helix by using the Dickerson sum function. The Fourier transform and the auto-correlation function of the twist angles sequences have been used to study the periodicity and randomness of the original DNA sequences. Basing on the correlation coefficient, a "distance" between two DNA fragments has been defined and used to compare some realistic DNA sequences. It is hoped that the techniques developed here could be used to analyze more realistic DNA sequences.

Base Sequence↗

Homologous nucleotide sequences between prokaryotic and eukaryotic mRNAs: the 5'-end sequence of the mRNA of the lipoprotein of the Escherichia coli outer membrane.

The sequence of the first 89 nucleotides at the 5' end of the mRNA for the lipoprotein of the Escherichia coli outer membrane is: GCUACAUGGAGAUUAACUCAAUCU-AGAGGGUAUUAAUAAUGAAAGCUACUAAACUGGUACU-GGGCGCGGUAAUCCUGGGUUCUACUCUG. The sequence of the first 72 nucleotides was established by direct sequencing methods and was extended to 89 residues on the basis of the known sequences of oligonucleotides obtained from complete digestion of the mRNA by ribonuclease T1 or A and the known amino acid sequence of the prolipoprotein. The mRNA has an untranslated region of 38 residues before the initiation codon, AUG. A unique feature of the 5'-end sequence of the mRNA is that the sequence of 12 nucleotides (GUAUUAAUAAUG) prior to, and including, the initiation codon is the same as that found at the ribosome-binding site for 80S ribosomes in brome mosaic virus RNA4, a eukaryotic mRNA [Dasgupta, R., Shih, D., Saris, C. & Kaesberg, P. (1975) Nature 256, 624-628].

Bacterial Proteins↗

NH2-terminal amino acid sequence and peptide mapping of purified human beta-lipotropin: comparison with previously proposed sequences.

Beta-Lipotropin was purified from human pituitary glands to a purity of greater than 90%. The amino acid compositions of beta-lipotropin and its three cyanogen bromide cleavage peptide fragments were in agreement with the structure proposed by Li and Chung [Li, C.H. & Chung, D. (1981) Int. J. Pept. Protein Res. 17, 131-142]. However, the amino acid sequence of its NH2-terminal 46 amino acid residues established here differs both from the sequence derived from the direct sequence analysis of the peptide reported by Li and Chung and from that predicted on the basis of the nucleotide sequence of the human pro-opiolipomelanocortin gene proposed by Chang et al. [Chang, A.C.Y., Cochet, M. & Cohen, S.W. (1980) Proc. Natl. Acad. Sci. USA 77,4890-4894] but agrees with the structure recently derived by direct sequence analysis by Hsi et al. [Hsi, K.L., Seidah, N.G., Lu, C.L. & Chrétien, M. (1981) Biochem. Biophys. Res. Commun. 103, 1329-1335] and predicted on the basis of nucleotide sequence analysis by Takahashi et al. [Takahashi, H., Teranishi, Y., Nakanishi, S. & Numa, S. (1981) FEBS Lett. 135, 97-102]. These discrepancies, found from residues 9 to 25 of beta-lipotropin, could result from pro-opiolipomelanocortin gene polymorphism, from the existence of multiple genes for pro-opiolipomelanocortin, or, more probably, from minor errors in nucleotide and amino acid sequence analyses.

Amino Acid Sequence↗

Nucleotide sequence of the insertion sequence found in the T-DNA region of mutant Ti plasmid pTiA66 and distribution of its homologues in octopine Ti plasmid.

The octopine tumor-inducing (Ti) plasmid pTiA66 has an insertion mutation in its T region (the DNA region incorporated into the plant genome) that results in the slow growth of crown gall tumors. These tumors exhibit hormonal autonomy different from that of the crown gall tumors caused by wild-type Ti plasmids. In the present study, the nucleotide sequences of both the DNA segment inserted into pTiA66 and its target site have been determined. The inserted segment is 2548 base pairs long and has 20-base-pair terminal inverted repeats. An 8-base-pair sequence at the target site is duplicated at both integration junctions. These structural features of the insert suggest that it is a bacterial insertion sequence (IS) element, which we have named IS66. Blot-hybridization analyses using IS66 probes revealed that genomes of octopine Ti plasmids contain at least three sequences homologous to IS66: two homologues are located in the virulence region and one is located between the left-hand (TL-DNA) and right-hand (TR-DNA) portions of T-DNA. The chromosome of Agrobacterium tumefaciens A66 also contains two sequences highly homologous to IS66. These results suggest that the mutant pTiA66 plasmid was generated by translocation of one of the sequences showing homology with IS66 into the T region. The fact that a sequence homologous to IS66 is present between TL-DNA and TR-DNA also suggests that the octopine T region was split into two portions, TL-DNA and TR-DNA, by translocation of IS66 or its relatives. Thus, IS66 may cause genetic and structural variations of the T region and the vir region of the octopine Ti plasmids.

Arginine↗

GTPase of bovine rod outer segments: the amino acid sequence of the alpha subunit as derived from the cDNA sequence.

The sequence of the 350 amino acids in the alpha subunit of GTPase of bovine rod outer segments has been determined. Enriched GTPase mRNA was used to prepare a cDNA library in the expression vector lambda gt11 and several overlapping cDNA clones corresponding to the alpha subunit of the GTPase were identified. The cDNA sequence determined contains 93 nucleotides upstream of the 5' end of the coding region, 1050 nucleotides that specify the amino acid sequence, and 45 nucleotides downstream from the 3' end. The previously described partial amino acid sequences and the sequences at the ADP-ribosylation sites for cholera and pertussis toxins are all confirmed and fitted into the present complete sequence. Homologies are found between the sequence of the alpha subunit and those of other guanine nucleotide-binding proteins, the ras proteins, peptide chain elongation factors EF-Tu and EF-G, and the initiation factor IF2.

Amino Acid Sequence↗

Bifunctionality of the AcMNPV homologous region sequence (hr1): enhancer and ori functions have different sequence requirements.

The Autographa californica multinucleocapsid nuclear polyhedrosis virus (AcMNPV) homologous region sequence hr1 is a putative origin of replication (ori) sequence and can also function as a transcriptional enhancer for delayed-early genes. We demonstrate that this 750-bp sequence, carrying five 28-bp core palin-dromes, enhances expression from the very late polyhedrin promoter up to 11-fold in a classical enhancer fashion in transient expression assays. Enhancement is at the level of transcription, as evident from RNase protection assay analysis. It is mediated by an alpha-amanitin-insensitive RNA polymerase from the authentic polyhedrin promoter transcription start site and follows the temporal activation profile characteristic of the polyhedrin promoter. Three lines of evidence conclusively demonstrated that hr1 acts typically as an enhancer of polyhedrin gene transcription independent of its role as an ori: (i) linearized hr1-reporter plasmids, incapable of replicating in the host cell, could enhance transcription from the promoter; (ii) reporter plasmid copy number was not affected by the presence of aphidicolin during transfection; (iii) reporter plasmid DNA recovered from Sf9 cells was sensitive to Dpn I confirming its unreplicated state in the transfection regime followed by us. Molecular dissection of the hr1 sequence elements revealed that a core palindrome alone can function as an ori sequence whereas a palindrome along with flanking sequences is essential for the enhancer activity. Enhancement of luciferase expression from the polyhedrin promoter is a function of the number of core palindromes and flanking sequences. Our results demonstrate that hr1, which has several motifs for enhancer binding proteins and transcription factors, has a dual role associated with both DNA replication and transcriptional enhancement.

Animals↗

Prediction of the coding sequences of unidentified human genes. XIII. The complete sequences of 100 new cDNA clones from brain which code for large proteins in vitro.

As a part of our cDNA project for deducing the coding sequence of unidentified human genes, we newly determined the sequences of 100 cDNA clones from a set of size-fractionated human brain cDNA libraries, and predicted the coding sequences of the corresponding genes, named KIAA0919 to KIAA1018. The sequencing of these clones revealed that the average sizes of the inserts and corresponding open reading frames were 4.9 kb and 2.6 kb (882 amino acid residues), respectively. A computer search of the sequences against the public databases indicated that predicted coding sequences of 87 genes contained sequences similar to known genes, 53% of which (46 genes) were categorized as proteins relating to cell signaling/communication, cell structure/motility and nucleic acid management. The chromosomal locations of the genes were determined by using human-rodent hybrid panels unless their mapping data were already available in the public databases. The expression profiles of all the genes among 10 human tissues, 8 brain regions (amygdala, corpus callosum, cerebellum, caudate nucleus, hippocampus, substania nigra, subthalamic nucleus, and thalamus), spinal cord, fetal brain and fetal liver were also examined by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Animals↗

Fine mapping of satellite DNA sequences along the Y chromosome of Drosophila melanogaster: relationships between satellite sequences and fertility factors.

The entirely heterochromatic Y chromosome of Drosophila melanogaster contains a series of simple sequence satellite DNAs which together account for about 80% of its length. Molecular cloning of the three simple sequence satellite DNAs of D. melanogaster (1.672, 1.686 and 1.705 g/ml) revealed that each satellite comprises several distinct repeat sequences. Together 11 related sequences were identified and 9 of them were shown to be located on the Y chromosome. In the present study we have finely mapped 8 of these sequences along the Y by in situ hybridization on mitotic chromosome preparations. The hybridization experiments were performed on a series of cytologically determined rearrangements involving the Y chromosome. The breakpoints of these rearrangements provided an array of landmarks along the Y which have been used to localize each sequence on the various heterochromatic blocks defined by Hoechst and N-banding techniques. The results of this analysis indicate a good correlation between the N-banded regions and 1.705 repeats and between the Hoechst-bright regions and the 1.672 repeats. However, the molecular basis for banding does not appear to depend exclusively on DNA content, since heterochromatic blocks showing identical banding patterns often contain different combinations of satellite repeats. The distribution of satellite repeats has also been analyzed with respect to the male fertility factors of the Y chromosome. Both loop-forming (kl-5, kl-3 and ks-1) and non-loop-forming (kl-2 and ks-2) fertility genes contain substantial amounts of satellite DNAs. Moreover, each fertility region is characterized by a specific combination of satellite sequences rather than by an homogeneous array of a single type of repeat.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗