Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Comparative genomic sequence analysis of the FXR gene family: FMR1, FXR1, and FXR2.

Mutations in the X-linked gene FMR1 cause fragile X syndrome, the leading cause of inherited mental retardation. Two autosomal paralogs of FMR1 have been identified, and are known as FXR1 and FXR2. Here we describe and compare the genomic structures of the mouse and human genes FMR1, FXR1, and FXR2. All three genes are very well conserved from mouse to human, with identical exon sizes for all but two FXR2 exons. In addition, the three genes share a conserved gene structure, suggesting they are derived from a common ancestral gene. As a first step towards exploring this hypothesis, we reexamined the Drosophila melanogaster gene Fmr1, and found it to have several of the same intron/exon junctions as the mammalian FXRs. Finally, we noted several regions of mouse/human homology in the noncoding portions of FMR1 and FXR1. Knowledge of the genomic structure and sequence of the FXR family of genes will facilitate further studies into the function of these proteins.

3' Untranslated Regions↗

Nucleotide sequence of noncoding regions in Rous-associated virus-2: comparisons delineate conserved regions important in replication and oncogenesis.

The nucleotide sequence of the regions flanking the long terminal repeat of Rous-associated virus-2 has been determined. The region analyzed spans the ends of the viral genome and includes the terminus of the env gene, the 3' noncoding region, the 5' noncoding region, and the beginning of the gag gene. These data have been compared with sequences available from other avian retroviruses. The comparisons reveal sections which are highly conserved and others which are quite variable. Sequence homologies within the conserved regions suggest details concerning the mode of origin of the src-transducing viruses. Included in the variable section is a region (XSR) found only in certain strains of Rous-derived virus. Its absence from other oncogenic viruses indicates that these sequences are not required to elicit disease.

Amino Acid Sequence↗

rVISTA 2.0: evolutionary analysis of transcription factor binding sites.

Identifying and characterizing the transcription factor binding site (TFBS) patterns of cis-regulatory elements represents a challenge, but holds promise to reveal the regulatory language the genome uses to dictate transcriptional dynamics. Several studies have demonstrated that regulatory modules are under positive selection and, therefore, are often conserved between related species. Using this evolutionary principle, we have created a comparative tool, rVISTA, for analyzing the regulatory potential of noncoding sequences. Our ability to experimentally identify functional noncoding sequences is extremely limited, therefore, rVISTA attempts to fill this great gap in genomic analysis by offering a powerful approach for eliminating TFBSs least likely to be biologically relevant. The rVISTA tool combines TFBS predictions, sequence comparisons and cluster analysis to identify noncoding DNA regions that are evolutionarily conserved and present in a specific configuration within genomic sequences. Here, we present the newly developed version 2.0 of the rVISTA tool, which can process alignments generated by both the zPicture and blastz alignment programs or use pre-computed pairwise alignments of several vertebrate genomes available from the ECR Browser and GALA database. The rVISTA web server is closely interconnected with the TRANSFAC database, allowing users to either search for matrices present in the TRANSFAC library collection or search for user-defined consensus sequences. The rVISTA tool is publicly available at http://rvista.dcode.org/.

Algorithms↗

Comparative physical mapping reveals features of microsynteny between Glycine max, Medicago truncatula, and Arabidopsis thaliana.

To gain insight into genomic relationships between soybean (Glycine max) and Medicago truncatula, eight groups of bacterial artificial chromosome (BAC) contigs, together spanning 2.60 million base pairs (Mb) in G. max and 1.56 Mb in M. truncatula, were compared through high-resolution physical mapping combined with sequence and hybridization analysis of low-copy BAC ends. Cross-hybridization among G. max and M. truncatula contigs uncovered microsynteny in six of the contig groups and extensive microsynteny in three. Between G. max homoeologous (within genome duplicate) contigs, 85% of coding and 75% of noncoding sequences were conserved at the level of cross-hybridization. By contrast, only 29% of sequences were conserved between G. max and M. truncatula, and some kilobase-scale rearrangements were also observed. Detailed restriction maps were constructed for 11 contigs from the three highly microsyntenic groups, and these maps suggested that sequence order was highly conserved between G. max duplicates and generally conserved between G. max and M. truncatula. One instance of homoeologous BAC contigs in M. truncatula was also observed and examined in detail. A sequence similarity search against the Arabidopsis thaliana genome sequence identified up to three microsyntenic regions in A. thaliana for each of two of the legume BAC contig groups. Together, these results confirm previous predictions of one recent genome-wide duplication in G. max and suggest that M. truncatula also experienced ancient large-scale genome duplications.

Arabidopsis↗

Genetic variation in yellow fever virus: duplication in the 3' noncoding region of strains from Africa.

The nucleotide sequences of three regions of the genomes of 13 yellow fever (YF) virus isolates were determined to define genetic variation and evolution of the virus. Phylogenetic trees generated from sequences of either the 5' terminal 1320 nucleotides of the genome, 754 nucleotides from the NS4A and NS4B genes, or the 3' terminal 511 nucleotides were very similar and contained minor differences. Overall, these results suggested that there were at least four major genotypes of YF virus, including one in Central/East Africa, one in West Africa, and two in South America. Examination of the 3' noncoding region (3'NCR) showed that only West African strains had a 3'NCR of 511 nucleotides while strains from Central/East Africa and South America had shorter 3'NCRs (443-469 nucleotides) due to the absence of YF specific repeat sequences (RYFs). Central/East African strains have two RYFs and West African strains have three RYFs while South American strains only have one copy of the RYF. It is speculated that duplication of the RYF took place in West Africa subsequent to the presumed introduction of YF virus into South America. Thus, both tick-borne (Mandl et al., J. Virol. 65, 4070-4077, 1991 and Wallner et al., Virology 213, 169-178, 1995) and mosquito-borne flaviviruses have variable 3'NCRs.

Africa↗

The hepatitis C virus: overview.

Our knowledge of hepatitis C virus (HCV) dates only from 1975, when non-A, non-B hepatitis was first recognized. It was not until 1989 that the genome of the virus was first cloned and sequenced, and expressed viral antigens used to develop serological assays for screening and diagnosis. HCV is in a separate genus of the virus family Flaviviridae. It is a spherical enveloped virus of approximately 50 nm in diameter. Its genome is a single-stranded linear RNA molecule of positive sense and consists of a 5' noncoding region, a single large open reading frame, and a 3' noncoding region. The open reading frame encodes at least three structural and six nonstructural proteins. The genome is characterized by significant genetic heterogeneity, based on which HCV isolates can be classified into six major genotypes and more than 50 subtypes. Even individual isolates of HCV are genetically heterogeneous (quasispecies diversity). Genetic heterogeneity of HCV is greatest in the amino-terminal end of the second envelope protein (hypervariable region 1). This region may represent a neutralization epitope that is under selective pressure from the host's humoral immune response. Infection with HCV proceeds to chronicity in more than 80% of cases, and even recovery does not protect against subsequent re-exposure to the virus. The development of a broadly protective vaccine against HCV will therefore require a better understanding of the molecular biology and immune response to this virus.

Animals↗

Influenza B viruses with site-specific mutations introduced into the HA gene.

We have succeeded in engineering changes into the genome of influenza B virus. First, model RNAs containing the chloramphenicol acetyltransferase gene flanked by the noncoding sequences of the HA or NS genes of influenza B virus were transfected into cells which were previously infected with an influenza B helper virus. Like those of the influenza A viruses, the termini of influenza B virus genes contain cis-acting signals which are sufficient to direct replication, expression, and packaging of the RNA. Next, a full-length copy of the HA gene from influenza B/Maryland/59 virus was cloned. Following transfection of this RNA, we rescued transfectant influenza B viruses which contain a point mutation introduced into the original cDNA. A series of mutants which bear deletions or changes in the 5' noncoding region of the influenza B/Maryland/59 virus HA gene were constructed. We were able to rescue viruses which contained deletions of 10 or 33 nucleotides at the 5' noncoding region of the HA gene. The viability of these viruses implies that this region of the genome is flexible in sequence and length.

Amino Acid Sequence↗

Molecular characterization of the 3' terminus of the simian hemorrhagic fever virus genome.

The 3' end of the simian hemorrhagic fever virus (SHFV) single-stranded RNA genome was cloned and sequenced. Adjacent to the 3' poly(A) tract, we identified a 76-nucleotide noncoding region preceded by two overlapping reading frames (ORFs). The ultimate 3' ORF of the viral genome encodes the capsid protein, and the penultimate ORF encodes the smallest SHFV envelope protein. These two ORFs overlap each other by 26 nucleotides. Northern (RNA) blot hybridization analyses of cytoplasmic RNA extracts from SHFV-infected MA-104 cells with gene-specific probes revealed the presence of full-length genomic RNA as well as six subgenomic SHFV-specific mRNA species. The subgenomic mRNAs are 3' coterminal. In its virion morphology and size, genome structure and length, and replication strategy, SHFV is most similar to lactate dehydrogenase-elevating virus, equine arteritis virus, and porcine reproductive and respiratory syndrome virus.

Amino Acid Sequence↗

Applications of recursive segmentation to the analysis of DNA sequences.

Recursive segmentation is a procedure that partitions a DNA sequence into domains with a homogeneous composition of the four nucleotides A, C, G and T. This procedure can also be applied to any sequence converted from a DNA sequence, such as to a binary strong(G + C)/weak(A + T) sequence, to a binary sequence indicating the presence or absence of the dinucleotide CpG, or to a sequence indicating both the base and the codon position information. We apply various conversion schemes in order to address the following five DNA sequence analysis problems: isochore mapping, CpG island detection, locating the origin and terminus of replication in bacterial genomes, finding complex repeats in telomere sequences, and delineating coding and noncoding regions. We find that the recursive segmentation procedure can successfully detect isochore borders, CpG islands, and the origin and terminus of replication, but it needs improvement for detecting complex repeats as well as borders between coding and noncoding regions.

Algorithms↗

A distant upstream enhancer at the maize domestication gene tb1 has pleiotropic effects on plant and inflorescent architecture.

Although quantitative trait locus (QTL) mapping has been successful in describing the genetic architecture of complex traits, the molecular basis of quantitative variation is less well understood, especially in plants such as maize that have large genome sizes. Regulatory changes at the teosinte branched1 (tb1) gene have been proposed to underlie QTLs of large effect for morphological differences that distinguish maize (Zea mays ssp. mays) from its wild ancestors, the teosintes (Z. mays ssp. parviglumis and mexicana). We used a fine mapping approach to show that intergenic sequences approximately 58-69 kb 5' to the tb1 cDNA confer pleiotropic effects on Z. mays morphology. Moreover, using an allele-specific expression assay, we found that sequences >41 kb upstream of tb1 act in cis to alter tb1 transcription. Our findings show that the large stretches of noncoding DNA that comprise the majority of many plant genomes can be a source of variation affecting gene expression and quantitative phenotypes.

Genes, Plant↗

High-throughput localization of functional elements by quantitative chromatin profiling.

Identification of functional, noncoding elements that regulate transcription in the context of complex genomes is a major goal of modern biology. Localization of functionality to specific sequences is a requirement for genetic and computational studies. Here, we describe a generic approach, quantitative chromatin profiling, that uses quantitative analysis of in vivo chromatin structure over entire gene loci to rapidly and precisely localize cis-regulatory sequences and other functional modalities encoded by DNase I hypersensitive sites. To demonstrate the accuracy of this approach, we analyzed approximately 300 kilobases of human genome sequence from diverse gene loci and cleanly delineated functional elements corresponding to a spectrum of classical cis-regulatory activities including enhancers, promoters, locus control regions and insulators as well as novel elements. Systematic, high-throughput identification of functional elements coinciding with DNase I hypersensitive sites will substantially expand our knowledge of transcriptional regulation and should simplify the search for noncoding genetic variation with phenotypic consequences.

Algorithms↗

Detecting conserved regulatory elements with the model genome of the Japanese puffer fish, Fugu rubripes.

Comparative vertebrate genome sequencing offers a powerful method for detecting conserved regulatory sequences. We propose that the compact genome of the teleost Fugu rubripes is well suited for this purpose. The evolutionary distance of teleosts from other vertebrates offers the maximum stringency for such evolutionary comparisons. To illustrate the comparative genome approach for F. rubripes, we use sequence comparisons between mouse and Fugu Hoxb-4 noncoding regions to identify conserved sequence blocks. We have used two approaches to test the function of these conserved blocks. In the first, homologous sequences were deleted from a mouse enhancer, resulting in a tissue-specific loss of activity when assayed in transgenic mice. In the second approach, Fugu DNA sequences showing homology to mouse sequences were tested for enhancer activity in transgenic mice. This strategy identified a neural element that mediates a subset of Hoxb-4 expression that is conserved between mammals and teleosts. The comparison of noncoding vertebrate sequences with those of Fugu, coupled to a transgenic bioassay, represents a general approach suitable for many genome projects.

Animals↗

Human papillomaviruses in Buschke-Löwenstein tumors: physical state of the DNA and identification of a tandem duplication in the noncoding region of a human papillomavirus 6 subtype.

Six Buschke-Löwenstein tumors, i.e., highly differentiated squamous cell tumors of the genital region, were shown to contain human papillomavirus 6 (HPV 6) or HPV 11 genomes. The viral DNA was found in an episomal state, including a very small fraction of circular oligomers. HPV 6a and HPV 6d genomes were cloned from two of the tumors. Comparison with HPV 6b, cloned from a benign genital wart (E. -M. de Villiers, L. Gissmann, and H. zur Hausen, J. Virol. 40:932-935, 1981) by restriction mapping and partial sequence analysis, revealed a very high degree of homology with the different HPV 6 subtypes. A tandem duplication of 459 base pairs within the noncoding region of the genome was found in the new subtype HPV 6d. This structural rearrangement in a region containing the putative control elements for early gene transcription might influence the biological potential of that virus. No evidence for rearrangement of this region was found in the HPV DNA from the five other tumors.

Base Sequence↗

C-ski cDNAs are encoded by eight exons, six of which are closely linked within the chicken genome.

The c-ski locus extends a minimum of 65 kb in the chicken genome and is expressed as multiple mRNAs resulting from alternative exon usage. Four exons comprising approximately 1.5 kb of cDNA sequence have been mapped within the chicken c-ski locus. However, c-ski cDNAs include almost 3 kb of sequence for which the exon structure was not defined. From our studies using the polymerase chain reaction and templates of RNA and genomic DNA, it is clear that c-ski cDNAs are encoded by a minimum of eight exons. A long 3' untranslated region is contiguous in the genome with the distal portion of the ski open reading frame such that exon 8 is composed of both coding and noncoding sequences. Exons 2 and 3 are separated by more than 25 kb of genomic sequence. In contrast, exons 3 through 8, representing more than half the length of c-ski cDNA sequences, are closely linked within 10 kb in the chicken genome.

Animals↗

Genome-wide analyses of two families of snoRNA genes from Drosophila melanogaster, demonstrating the extensive utilization of introns for coding of snoRNAs.

Small nucleolar RNAs (snoRNAs) are an abundant group of noncoding RNAs mainly involved in the post-transcriptional modifications of rRNAs in eukaryotes. In this study, a large-scale genome-wide analysis of the two major families of snoRNA genes in the fruit fly Drosophila melanogaster has been performed using experimental and computational RNomics methods. Two hundred and twelve gene variants, encoding 56 box H/ACA and 63 box C/D snoRNAs, were identified, of which 57 novel snoRNAs have been reported for the first time. These snoRNAs were predicted to guide a total of 147 methylations and pseudouridylations on rRNAs and snRNAs, showing a more comprehensive pattern of rRNA modification in the fruit fly. With the exception of nine, all the snoRNAs identified to date in D. melanogaster are intron encoded. Remarkably, the genomic organization of the snoRNAs is characteristic of 8 dUhg genes and 17 intronic gene clusters, demonstrating that distinct organizations dominate the expression of the two families of snoRNAs in the fruit fly. Of the 267 introns in the host genes, more than half have been identified as host introns for coding of snoRNAs. In contrast to mammals, the variation in size of the host introns is mainly due to differences in the number of snoRNAs they contain. These results demonstrate the extensive utilization of introns for coding of snoRNAs in the host genes and shed light on further research of other noncoding RNA genes in the large introns of the Drosophila genome.

Animals↗

Genome-wide analysis of C/D and H/ACA-like small nucleolar RNAs in Leishmania major indicates conservation among trypanosomatids in the repertoire and in their rRNA targets.

Small nucleolar RNAs (snoRNAs) are a large group of noncoding RNAs that exist in eukaryotes and archaea and guide modifications such as 2'-O-ribose methylations and pseudouridylation on rRNAs and snRNAs. Recently, we described a genome-wide screening approach with Trypanosoma brucei that revealed over 90 guide RNAs. In this study, we extended this approach to analyze the repertoire of the closely related human pathogen Leishmania major. We describe 23 clusters that encode 62 C/Ds that can potentially guide 79 methylations and 37 H/ACA-like RNAs that can potentially guide 30 pseudouridylation reactions. Like T. brucei, Leishmania also contains many modifications and guide RNAs relative to its genome size. This study describes 10 H/ACAs and 14 C/Ds that were not found in T. brucei. Mapping of 2'-O-methylations in rRNA regions rich in modifications suggests the existence of trypanosomatid-specific modifications conserved in T. brucei and Leishmania. Structural features of C/D snoRNAs, such as copy number, conservation of boxes, K turns, and intragenic and extragenic base pairing, were examined to elucidate the great variation in snoRNA abundance. This study highlights the power of comparative genomics for determining conserved features of noncoding RNAs.

Animals↗

Retrotransposon-derived elements in the mammalian genome: a potential source of disease.

The plethora of genomic information gathered by the sequencing of the human and mouse genomes has paved the way for a new era of genetics. While in the past we focused mainly on the small percentage of DNA that codes for proteins, we can now concentrate on the remainder, i.e. the noncoding sequences that interrupt and separate genes. This portion of the genome is made up, in most part, of repetitive DNA sequences including DNA transposons, long terminal repeat (LTR) retrotransposons, LINEs (long interspersed nuclear elements) and SINEs (short interspersed nuclear elements). Some of these elements are transcriptionally active and can transpose or retrotranspose around the genome, resulting in insertional mutagenesis that can cause disease. In these cases, insertions have occurred in the coding sequence. However, recent evidence suggests that the main effect of these elements is their ability to influence transcription of neighbouring genes. The elements themselves contain promoters that can initiate transcription of flanking genomic DNA. Furthermore, they are susceptible to epigenetic silencing, which is often stochastic and incomplete, resulting in complex patterns of transcription. This review discusses some diseases in both human and mouse that are caused by these repetitive elements.

Animals↗

Long noncoding RNAs and diabetic retinopathy: Current understanding, future directions and challenges.

Diabetic retinopathy remains as the leading cause of preventable blindness in working-aged people. The pathophysiology of this sight-threatening disease is complex and involves intricate interactions among metabolic, hemodynamic and epigenetic pathways, leading to molecular, structural, functional and genomic abnormalities in retinal vascular and nonvascular cells. Diabetes also results in differential expressions of several noncoding RNAs, including micro RNAs (miRNAs) and long noncoding RNAs (LncRNAs). Compared to about 2000 miRNA identified in human genome thus far, more than 30,000 LncRNA transcripts have been already identified, but the function and mechanism of action of most of the LncRNAs is still not fully characterized, and there remains a possibility that some LncRNAs could have diverse functions under different contexts. LncRNAs are mainly noncoding, but they have many regulatory functions, and regulate gene expression by interacting with DNA, RNA and protein. Aberrant expressions of several LncRNAs including MALAT1, MEG3, HOTAIR, MIAT1 and H19, is associated with metabolic abnormalities implicated in the pathogenesis of diabetic retinopathy. LncRNAs are also released into circulation and show high organ and cell specificity, and greater disease-associated differences compared to disease-associated mRNAs. Furthermore, LncRNAs maintain stable expression in the plasma and can be isolated from total RNA present in biological samples including blood, which makes them promising and reliable candidates for diagnostic or prognostic markers and therapeutic targets for various diseases. With continued improvement in innovative RNA modifications and delivery modalities, use of LncRNAs as possible biomarkers, and of LncRNA-based therapeutics, for diabetic retinopathy appears promising.

Biomarkers↗