Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Fewer genes, more noncoding RNA.

Recent studies showing that most "messenger" RNAs do not encode proteins finally explain the long-standing discrepancy between the small number of protein-coding genes found in vertebrate genomes and the much larger and ever-increasing number of polyadenylated transcripts identified by tag-sampling or microarray-based methods. Exploring the role and diversity of these numerous noncoding RNAs now constitutes a main challenge in transcription research.

Animals↗

Characterization of the hamster DDT-1 cell aFGF/HGBF-I gene and cDNA and its modulation by steroids.

Syrian hamster DDT-1 cells are derived from smooth muscle of the ductus deferens. DDT-1 cell growth is increased by the addition of testosterone (T). Acidic fibroblast growth factor (aFGF) or basic fibroblast growth factor (bFGF) also known as heparin binding growth factor I and II (HBGF-I and HBGF-II) can replace T in the stimulation of growth in these cells. This phenomenon is correlated with testosterone's ability to elevate aFGF/HBGF-I mRNA. The increase steady-state levels of aFGF/HBGF-I mRNA were documented by northern blots and by in situ hybridization. Using a 520 bp human aFGF/HBGF-I cDNA probe, a genomic clone with a 38 kb DNA insert was isolated from a cosmid library. By restriction enzyme analysis and southern hybridization, it was determined that there are three coding exons. DNA sequence analysis showed all of the coding region and 3' noncoding sequences were on this clone. A 5' noncoding exon not in the 38 kb insert is indicated, based on the cDNA sequences and genomic sequences of aFGF/HBGF-I's from hamster DDT-1 cells and several other species. The cDNA for hamster aFGF/HBGF-I was isolated from a DDT-1 lambda gt11 library and sequenced. Comparison of the coding region of aFGF/HBGF-I from four species shows a greater than 90% conservation of amino acid sequence.

Amino Acid Sequence↗

Comparative genomics and transcriptional analysis of prophages identified in the genomes of Lactobacillus gasseri, Lactobacillus salivarius, and Lactobacillus casei.

Lactobacillus gasseri ATCC 33323, Lactobacillus salivarius subsp. salivarius UCC 118, and Lactobacillus casei ATCC 334 contain one (LgaI), four (Sal1, Sal2, Sal3, Sal4), and one (Lca1) distinguishable prophage sequences, respectively. Sequence analysis revealed that LgaI, Lca1, Sal1, and Sal2 prophages belong to the group of Sfi11-like pac site and cos site Siphoviridae, respectively. Phylogenetic investigation of these newly described prophage sequences revealed that they have not followed an evolutionary development similar to that of their bacterial hosts and that they show a high degree of diversity, even within a species. The attachment sites were determined for all these prophage elements; LgaI as well as Sal1 integrates in tRNA genes, while prophage Sal2 integrates in a predicted arginino-succinate lyase-encoding gene. In contrast, Lca1 and the Sal3 and Sal4 prophage remnants are integrated in noncoding regions in the L. casei ATCC 334 and L. salivarius UCC 118 genomes. Northern analysis showed that large parts of the prophage genomes are transcriptionally silent and that transcription is limited to genome segments located near the attachment site. Finally, pulsed-field gel electrophoresis followed by Southern blot hybridization with specific prophage probes indicates that these prophage sequences are narrowly distributed within lactobacilli.

Animals↗

The complete nucleotide sequence of a variant of Coxsackievirus A24, an agent causing acute hemorrhagic conjunctivitis.

The complete nucleotide sequence was determined for the cDNAs that represent the RNA genome of the standard strain of a variant of coxsackievirus A24, the EH24/70, one of the agents causing acute hemorrhagic conjunctivitis. The genome is 7461 nucleotide long and is polyadenylated at the 3'-end terminus. Following a 750-nucleotide 5'-noncoding region, there was a long open reading frame of 6642 nucleotides, which serve to encode a viral polyprotein consisting of 2214 amino acids. Comparison of the deduced amino acid sequence of the polyprotein with those of known enteroviruses allowed us to predict the possible cleavage sites. The overall structure and the organization of the RNA genome is typical for an enterovirus. Based on the similarity of the nucleotide sequence of the 5' and 3' noncoding regions, together with the amino-acid sequence of the encoded proteins, EH24/70 appeared to be closely related to polioviruses and coxsackievirus A21.

Amino Acid Sequence↗

Complete genome sequence of Rickettsia typhi and comparison with sequences of other rickettsiae.

Rickettsia typhi, the causative agent of murine typhus, is an obligate intracellular bacterium with a life cycle involving both vertebrate and invertebrate hosts. Here we present the complete genome sequence of R. typhi (1,111,496 bp) and compare it to the two published rickettsial genome sequences: R. prowazekii and R. conorii. We identified 877 genes in R. typhi encoding 3 rRNAs, 33 tRNAs, 3 noncoding RNAs, and 838 proteins, 3 of which are frameshifts. In addition, we discovered more than 40 pseudogenes, including the entire cytochrome c oxidase system. The three rickettsial genomes share 775 genes: 23 are found only in R. prowazekii and R. typhi, 15 are found only in R. conorii and R. typhi, and 24 are unique to R. typhi. Although most of the genes are colinear, there is a 35-kb inversion in gene order, which is close to the replication terminus, in R. typhi, compared to R. prowazekii and R. conorii. In addition, we found a 124-kb R. typhi-specific inversion, starting 19 kb from the origin of replication, compared to R. prowazekii and R. conorii. Inversions in this region are also seen in the unpublished genome sequences of R. sibirica and R. rickettsii, indicating that this region is a hot spot for rearrangements. Genome comparisons also revealed a 12-kb insertion in the R. prowazekii genome, relative to R. typhi and R. conorii, which appears to have occurred after the typhus (R. prowazekii and R. typhi) and spotted fever (R. conorii) groups diverged. The three-way comparison allowed further in silico analysis of the SpoT split genes, leading us to propose that the stringent response system is still functional in these rickettsiae.

Chromosome Inversion↗

New insights into the function of noncoding RNA and its potential role in disease pathogenesis.

All polyadenylated RNAs expressed in mammalian tissues are assumed to be transported to the cytoplasm where they direct the synthesis of a protein product. This mainstream view of the function of polyadenylated transcripts is currently being challenged by the identification of a novel class of genes which, although they encode polyadenylated RNA, do not make a translated protein. Many of these noncoding RNAs are developmentally regulated or show highly restricted patterns of gene expression, and their functions are providing important insight into RNA-based mechanisms of gene expression, genomic imprinting, cell cycle progression, and differentiation. The purpose of this review is to discuss the current understanding of mammalian noncoding RNAs, and to highlight their potential for identifying new pathways of human disease.

Animals↗

The compositional properties of human genes.

The present work represents the first attempt to study in greater detail previously proposed compositional correlations in genomes, based on a body of additional data relating to gene localizations as well as to extended flanking sequences extracted from gene banks. We have investigated the correlations that exist between (1) the GC levels of exons of human genes, and (2) the GC levels of either intergenic sequences or introns associated with the genes under consideration. In both cases, linear relationships with slopes close to unity were found. The similarity of the linear relationships indicates similar GC levels in intergenic sequences and introns located in the same isochores. Moreover, both intergenic sequences and introns showed GC levels 5-10% lower than the corresponding exons. The above findings considerably strengthen the previously drawn conclusion that coding and noncoding sequences (both inter- and intragenic) from the same isochores of the human genome are compositionally correlated. In addition, we find linear correlations between the GC levels of codon positions and of the intergenic sequences or introns associated with the corresponding genes, as well as among the GC levels of codon positions of genes.

Base Composition↗

Reduced variation in Drosophila simulans mitochondrial DNA.

We investigated the evolutionary dynamics of infection of a Drosophila simulans population by a maternally inherited insect bacterial parasite, Wolbachia, by analyzing nucleotide variability in three regions of the mitochondrial genome in four infected and 35 uninfected lines. Mitochondrial variability is significantly reduced compared to a noncoding region of a nuclear-encoded gene in both uninfected and pooled samples of flies, indicating a sweep of genetic variation. The selective sweep of mitochondrial DNA may have been generated by the fixation of an advantageous mitochondrial gene mutation in the mitochondrial genome. Alternatively, the dramatic reduction in mitochondrial diversity may be related to Wolbachia.

Animals↗

Average mutual information of coding and noncoding DNA.

One basic problem in the analysis of DNA sequences is the recognition of protein-coding genes. Computer algorithms to facilitate gene identification have become important as genome sequencing projects have turned from mapping to large-scale sequencing, resulting in an exponentially growing number of sequenced nucleotides that await their annotation. Many statistical patterns have been discovered that are different in coding and noncoding DNA, but most of them vary from species to species, and hence require prior training on organism-specific data sets. Here, we investigate if there exist species-independent statistical patterns that are different in coding and noncoding DNA. We introduce an information-theoretic quantity, the average mutual information (AMI), and we find that the probability distribution functions of the AMI are significantly different in coding and noncoding DNA, while they are almost identical for different species. This finding suggests that the AMI might be useful for the recognition of protein-coding regions in genomes for which training sets do not exist.

Algorithms↗

Evolving genomic metaphors: a new look at the language of DNA.

Recent genome-sequencing efforts have confirmed that traditional "good-citizen" genes (those that encode functional RNA and protein molecules of obvious benefit to the organism) constitute only a small fraction of the genomic populace in humans and other multicellular creatures. The rest of the DNA sequence includes an astonishing collection of noncoding regions, regulatory modules, deadbeat pseudogenes, legions of repetitive elements, and hosts of oft-shifty, self-interested nomads, renegades, and immigrants. To help visualize functional operations in such intracellular genomic societies and to better encapsulate the evolutionary origins of complex genomes, new and evocative metaphors may be both entertaining and research-stimulating.

Animals↗

Rice mitochondrial genome contains a rearranged chloroplast gene cluster.

We have previously reported the isolation and partial sequence analysis of a rice mitochondrial DNA fragment (6.9 kb) which contains a transferred copy of a chloroplast gene cluster coding for the large subunit of ribulose-1,5-bisphosphate carboxylase (rbcL), beta and epsilon subunits of ATPase (atpB and atpE), methionine tRNA (trnM) and valine tRNA (trnV). We have now completely sequenced this 6.9 kb fragment and found it to also contain a sequence homologous to the chloroplast gene coding for the ribosomal protein L2 (rpl2), beginning at a site 430 bp downstream from the termination codon of rbcL. In the chloroplast genome, two copies of rpl2 are located at distances of 20 kb and 40 kb, respectively, from rbcL. We have sequenced these two copies of rice chloroplast rpl2 and found their sequences to be identical. In addition, a 151 bp sequence located upstream of the chloroplast rpl2 coding region is also found in the 3' noncoding region of chloroplast rbcL and other as yet undefined locations in the rice chloroplast genome. Hybridization analysis revealed that this 151 bp repeat sequence identified in rice is also present in several copies in 11 other plant species we have examined. Findings from these studies suggest that the translocation of rpl2 to the rbcL gene cluster found in the rice mitochondrial genome might have occurred through homologous recombination between the 151 bp repeat sequence present in both rpl2 and rbcL.

Base Sequence↗

Genome size and chromatin condensation in vertebrates.

Cell membrane-dependent chromatin condensation was studied by flow cytometry in erythrocytes of 36 species from six classes of vertebrates. A positive relationship was found between the degree of condensation and genome size. The distribution of variances among taxonomic levels is similar for both parameters. However, chromatin condensation varied relatively more at the lower taxonomic levels, which suggests that the degree of DNA packaging might serve for fine-tuning the 'skeletal' and/or 'buffering' function of noncoding DNA (although the range of this fine-tuning is smaller than the range of genome size changes). For two closely related amphibian species differing in genome size, change in chromatin condensation under the action of elevated extracellular salinity was investigated. Condensation was steadier and its reaction to changes in solvent composition was more inertial in the species with a larger genome, which is in agreement with the buffering function postulated for redundant DNA. The uppermost genome size in vertebrates (and in living beings in general) was updated using flow cytometry and was found to be about 80 pg (78,400 Mb). The widespread opinion that the largest genome occurs in unicellular organisms is rejected as being based on artifacts.

Animals↗

Nucleotide sequence of the 5' noncoding region and part of the gag gene of Rous sarcoma virus.

Several functions of the retrovirus genome involve structural features in the vicinity of its 5' terminus. In an effort to further elucidate the relationship between structure and function in retrovirus RNA, we have determined the sequence of the first 1,010 nucleotides at the 5' end of the genome of Rous sarcoma virus by using the Maxam-Gilbert method to sequence suitable domains in cloned Rous sarcoma virus DNA. The results (i) locate the initiation codon for the gag gene of Rous sarcoma virus 372 nucleotides from the 5' end of viral RNA; (ii) demonstrate that this codon is preceded by three methionine codons that are apparently not used in translation; (iii) sustain previous conclusions that the principal site to which ribosomes bind on the Rous sarcoma virus genome in vitro does not contain the initiation codon for gag; (iv) permit deduction of the amino acid sequence of a viral structural protein, p19; (v) confirm the amino-terminal sequence of Pr76gag; and (vi) substantiate the identification of a splice donor site described in the accompanying manuscript (Hackett et al., J. Virol., 41:527-534, 1982).

Avian Sarcoma Viruses↗

Nucleotide sequence and deduced amino acid sequence of the nonstructural proteins of dengue type 2 virus, Jamaica genotype: comparative analysis of the full-length genome.

The sequence of the 5'-end of the genome of dengue 2 (Jamaica genotype) virus has been previously reported (V. Deubel, R. M. Kinney, and D. W. Trent, 1986, Virology 155, 365-377). We have now cloned and sequenced the remaining 75% of the genomic RNA that encodes the nonstructural proteins. The complete genome is 10,723 bases in length with a single open reading frame extending from nucleotides 97 to 10,269 encoding 3391 amino acids. The 3'-noncoding extremity presents a stem- and loop-structure and contains a repeated oligonucleotide sequence. Comparisons of the nucleotide sequences of the genomes of dengue 2 viruses of different topotypes reveal 90-95% similarity, with 64-66% similarity evident between dengue viruses of different serotypes. The amino acid sequence of the polyprotein of dengue 2 Jamaica virus shows 97, 68, 50, and 44% similarity with those of other dengue 2, dengue 1, or dengue 4, West Nile, and yellow fever viruses, respectively. Despite amino acid sequence divergence, the hydrophobic profile of the flavivirus proteins is highly conserved. Proteins NS1, NS3, and NS5 are the most conserved. Conserved amino acid stretches present in all flavivirus proteins may be involved in common essential biological functions.

Amino Acid Sequence↗

Identification of specific nucleotide sequences within the conserved 3'-SL in the dengue type 2 virus genome required for replication.

The flavivirus genome is a positive-stranded approximately 11-kb RNA including 5' and 3' noncoding regions (NCR) of approximately 100 and 400 to 600 nucleotides (nt), respectively. The 3' NCR contains adjacent, thermodynamically stable, conserved short and long stem-and-loop structures (the 3'-SL), formed by the 3'-terminal approximately 100 nt. The nucleotide sequences within the 3'-SL are not well conserved among species. We examined the requirement for the 3'-SL in the context of dengue virus type 2 (DEN2) replication by mutagenesis of an infectious cDNA copy of a DEN2 genome. Genomic full-length RNA was transcribed in vitro and used to transfect monkey kidney cells. A substitution mutation, in which the 3'-terminal 93 nt constituting the wild-type (wt) DEN2 3'-SL sequence were replaced by the 96-nt sequence of the West Nile virus (WN) 3'-SL, was sublethal for virus replication. An analysis of the growth phenotypes of additional mutant viruses derived from RNAs containing DEN2-WN chimeric 3'-SL structures suggested that the wt DEN2 nucleotide sequence forming the bottom half of the long stem and loop in the 3'-SL was required for viability. One 7-bp substitution mutation in this domain resulted in a mutant virus that grew well in monkey kidney cells but was severely restricted in cultured mosquito cells. In contrast, transpositions of and/or substitutions in the wt DEN2 nucleotide sequence in the top half of the long stem and in the short stem and loop were relatively well tolerated, provided the stem-loop secondary structure was conserved.

Animals↗

The analysis of nucleotide substitutions, gaps, and recombination events between RHD and RHCE genes through complete sequencing.

We determined the entire nucleotide sequences of all introns within the RHD and RHCE genes by amplifying genomic DNA using long PCR methods. The RHD and RHCE genes were 57,295 and 57,831 bp in length, respectively. Aligning both genes revealed 138 gaps (insertions and deletions) below 100 bp, 1116 substitutions in all introns and all exons (coding region), and 5 gaps of over 100 bp. Homologies (%) between the RH genes were 93.8% over all introns and coding exons and 91.7% over all exons and introns. Various short tandem repeats (STRs) and many interspersed nuclear elements were identified in both genes. The proportions of Alu sequences in the RHD and RHCE genes were 25.9 and 25.7%, respectively and these Alu sequences were concentrated in several regions. We confirmed multiple recombinations in introns 1 and 2. Such multiple recombination, which probably arose due to the concentrations of Alu sequences and the high level of the homology (%), is one of most important factors in the formation and evolution of RH gene. The variability of the Rh system may be generated because of these features of RH genes. Apparent mutational hotspots and regions with low of K values (the numbers of substitutions per nucleotide site) caused by recombinations as well as true mutational hotspots may be found in human genome. Accordingly, in searching for and identifying single nucleotide polymorphisms (SNPs) especially in noncoding regions, apparent mutational hotspots and areas of low K values by recombination should be noted since the unequal distribution of SNPs will reduce the power of SNPs as genetic maker. Combining the complete sequences' data of both RH genes with serological findings will provide beneficial information with which to elucidate the mechanism of recombination, mutation, polymorphism, and evolution of other genes containing the RH gene as well as to analyze Rh variants and develop new methods of Rh genotyping.

Base Sequence↗

Molecular cloning and sequence determination of DA strain of Theiler's murine encephalomyelitis viruses.

Theiler's murine encephalomyelitis viruses (TMEV) belong to the Picornaviridae, and are divided into two subgroups. TO subgroup strains produce a persistent demyelinating central nervous system infection in mice, while GDVII subgroup strains cause acute polioencephalomyelitis. We generated three overlapping clones of the genome of DA strain, a member of TO subgroup. Sequence analysis revealed that the genome is 8093 nucleotides long with a poly(A) tail. The 5' noncoding region stretches from nucleotides 1 to 1065 and lacks a poly(C) tract. The open reading frame stretches from 1066 to 7968 and encodes 2301 amino acids. DA strain sequence is more closely related to members of the Cardiovirus genus than to members of other Picornavirus genera. Comparison with sequence of BeAn strain, another TO subgroup strain, showed that the P1 area has the greatest number of differences, while the noncoding regions are more well-conserved. The three overlapping clones will be important in recombinant infectious cDNA studies between strains of both subgroups.

Amino Acid Sequence↗

Characterization of the murine gene encoding the hyaluronan receptor RHAMM.

We describe the isolation and characterization of the murine gene encoding RHAMM, a hyaluronan receptor which regulates focal adhesion turnover, is required for cell locomotion and is a critical downstream regulator of ras transformation. The RHAMM gene spans at least 20 kb and comprises 14 exons ranging in size from 75 to 1099 bp. Primer extension studies indicate that the major transcription start point is in position -31, relative to the start Met. Northern blot analysis of mouse fibroblast RNA identified two hybridizing species of 4.2 and 1.7 kb. Comparison of cDNA clones and RT-PCR products with the genomic clones identified alternately spliced exons in both the coding and 5' noncoding regions of RHAMM. In the coding region exon 4 is alternately spliced. The major RHAMM transcript (RHAMM1) in 3T3 fibroblasts does not contain exon 4 and encodes a protein of 70 kDa. A minor transcript containing exon 4, namely RHAMM v4, encodes a 73-kDa protein, as demonstrated by isoform-specific antibodies. Western analysis demonstrated both a major 70-kDa (RHAMM 1) and minor 73-kDa RHAMM protein (v4) in 3T3 murine fibroblast cell lysates. The functional significance of these two isoforms is currently being investigated.

Amino Acid Sequence↗