Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Conserved expression domains for genes upstream and within the HoxA and HoxD clusters suggests a long-range enhancer existed before cluster duplication.

The posterior HoxA and HoxD genes are essential in appendicular development. Studies have demonstrated that a "distal limb enhancer," remotely located upstream of the HoxD complex, is required to drive embryonic autopod expression of the posterior Hox genes as well as the two additional non-Hox genes in the region: Evx2 and Lnp. Our work demonstrates a similar mode of regulation for Hoxa13 and four upstream genes: Evx1, Hibadh, Tax1bp, and Jaz1. These genes all show embryonic (E11.5-E13.5) distal limb and genital bud expression, suggesting the existence of a nearby enhancer influencing the expression of a domain of genes. Comparative sequence analysis between homologous human and mouse genomic sequence upstream of Hoxa13 revealed a remote 2.25-kb conserved noncoding sequence (mmA13CNS) within the fourth intron of the Hibadh gene. mmA13CNS shares a common 131-bp core identity within a conserved noncoding sequence upstream of Hoxd13, which is located within the previously identified distal limb enhancer critical region. To test the function of this conserved sequence, we created mmA13CNS-Hsp86-lacZ transgenic mice. mmA13CNS directed a wide range of tissue expression, including the central nervous system, developing olfactory tissue, limb, and genital bud. Limb and genital bud expression directed by mmA13CNS is not identical to the patterns exhibited by Hoxa13/Evx1/Hibadh/Tax1bp1/Jaz1, suggesting that mmA13CNS is not sufficient to fully recapitulate their expression in those tissues. The Evx1- and Evx2-like central nervous system expression observed in these mice suggests that the long-range regulatory element(s) for the Hox cluster existed before the cluster duplication.

Animals↗

Nested polymerase chain reaction for high-sensitivity detection of enteroviral RNA in biological samples.

A method based on nested polymerase chain reaction was developed for the detection of enteroviral genomes in biological samples. By taking advantage of the conserved 5' noncoding region of the enteroviral RNA, two sets of primers were utilized, enabling the detection either of a broad range of enteroviruses or of group B coxsackieviruses only. The sensitivity of the method is close to the detection of single molecules of viral RNA in as much as 1 mg of tissue sample. A preliminary study showed the usefulness of this technique for the analysis of endomyocardial biopsy samples from patients with idiopathic dilated cardiomyopathy and myocarditis.

Base Sequence↗

Encephalomyocarditis virus 3C protease: efficient cell-free expression from clones which link viral 5' noncoding sequences to the P3 region.

All picornaviral peptides are derived by progressive posttranslational cleavage of a giant precursor polyprotein. Translation of encephalomyocarditis virus (EMC) RNA in rabbit reticulocyte extracts produces active viral peptides, including protease 3C, which is responsible for many cleavage reactions within the processing cascade. DNA plasmids containing 5' noncoding sequences of EMC linked to other portions of the viral genome were constructed and transcribed into RNA. Like virion RNA, the clone-derived transcripts directed efficient protein translation in vitro. The 5'-linked constructions may represent examples of a general method for cell-free expression of any cloned gene segment. One construction produced a self-cleaving P3 region precursor, which contained active 3C protease. A genetically engineered insertion within the 3C sequences eliminated endogenous self-cleavage activity without altering the ability of the P3 peptide to serve as substrate in bimolecular reactions with added 3C. Another plasmid encoding the L-VP0 portion of the capsid region was used to demonstrate that scission between the leader peptide (L) and capsid protein VP0 can be catalyzed by 3C. The enzyme responsible for this step was previously unidentified. A rapid purification scheme for isolation of 3C from EMC-infected HeLa cells is also presented.

Base Sequence↗

Induction of lymphomas by the hamster papovavirus correlates with massive replication of nonrandomly deleted extrachromosomal viral genomes.

The hamster papovavirus isolated from skin epithelioma can induce lymphomas and leukemias after subcutaneous inoculation into newborn hamsters. The lymphoma cells are virus free but contain large amounts of extrachromosomal hamster papovavirus DNA. We have cloned and partly sequenced some of these DNA molecules from independent tumors. These genomes displayed overlapping deletions consistently sharing a common end within the noncoding regulatory sequences; the other end was variable but always extended into the sequence coding for the N-terminal part of the viral capsid VP2. This unique in vivo interaction between a polyomavirus and its cellular host, the genesis of these variant molecules, and their role in the lymphoma formation are discussed.

Animals↗

Inhibition of Rous sarcoma virus replication by antisense RNA.

Previous results have indicated that Rous sarcoma virus env gene expression is specifically inhibited by antisense RNA (L.-J. Chang and C. M. Stoltzfus, Mol. Cell. Biol. 5:2341-2348, 1985). In this study, we compare the extents of inhibition by antisense RNA derived from different parts of the Rous sarcoma virus genome, and we show that antisense constructs containing the 3'-end noncoding region inhibit env expression to a similar extent as those containing the 5'-end noncoding region or coding region. Furthermore, we show that antisense RNA inhibits virus replication at other levels in addition to translation.

Avian Sarcoma Viruses↗

Kaposi's sarcoma-associated herpesvirus lytic origin (ori-Lyt)-dependent DNA replication: identification of the ori-Lyt and association of K8 bZip protein with the origin.

Herpesviruses utilize different origins of replication during lytic versus latent infection. Latent DNA replication depends on host cellular DNA replication machinery, whereas lytic cycle DNA replication requires virally encoded replication proteins. In lytic DNA replication, the lytic origin (ori-Lyt) is bound by a virus-specified origin binding protein (OBP) that recruits the core replication machinery. In this report, we demonstrated that DNA sequences in two noncoding regions of the Kaposi's sarcoma-associated herpesvirus (KSHV) genome, between open reading frames (ORFs) K4.2 and K5 and between K12 and ORF71, are able to serve as origins for lytic cycle-specific DNA replication. The two ori-Lyt domains share an almost identical 1,153-bp sequence and a 600-bp downstream GC-rich repeat sequence, and the 1.7-kb DNA sequences are sufficient to act as a cis signal for replication. We also showed that an AT-palindromic sequence in the ori-Lyt domain is essential for the DNA replication. In addition, a virally encoded bZip protein, namely K8, was found to bind to a DNA sequence within the ori-Lyt by using a DNA binding site selection assay. The binding of K8 to this region was confirmed in cells by using a chromatin immunoprecipitation method. Further analysis revealed that K8 binds to an extended region, and the entire region is 100% conserved between two KSHV ori-Lyt's. K8 protein displays significant similarity to the Zta protein of Epstein-Barr virus (EBV), which is a known OBP of EBV. This notion, together with the ability of K8 to bind to the KSHV ori-Lyt, suggests that K8 may function as an OBP in KSHV.

Base Sequence↗

Assembly of severe acute respiratory syndrome coronavirus RNA packaging signal into virus-like particles is nucleocapsid dependent.

The severe acute respiratory syndrome coronavirus (SARS-CoV) was recently identified as the etiology of SARS. The virus particle consists of four structural proteins: spike (S), small envelope (E), membrane (M), and nucleocapsid (N). Recognition of a specific sequence, termed the packaging signal (PS), by a virus N protein is often the first step in the assembly of viral RNA, but the molecular mechanisms involved in the assembly of SARS-CoV RNA are not clear. In this study, Vero E6 cells were cotransfected with plasmids encoding the four structural proteins of SARS-CoV. This generated virus-like particles (VLPs) of SARS-CoV that can be partially purified on a discontinuous sucrose gradient from the culture medium. The VLPs bearing all four of the structural proteins have a density of about 1.132 g/cm(3). Western blot analysis of the culture medium from transfection experiments revealed that both E and M expressed alone could be released in sedimentable particles and that E and M proteins are likely to form VLPs when they are coexpressed. To examine the assembly of the viral genomic RNA, a plasmid representing the GFP-PS580 cDNA fragment encompassing the viral genomic RNA from nucleotides 19715 to 20294 inserted into the 3' noncoding region of the green fluorescent protein (GFP) gene was constructed and applied to the cotransfection experiments with the four structural proteins. The SARS-CoV VLPs thus produced were designated VLP(GFP-PS580). Expression of GFP was detected in Vero E6 cells infected with the VLP(GFP-PS580), indicating that GFP-PS580 RNA can be assembled into the VLPs. Nevertheless, when Vero E6 cells were infected with VLPs produced in the absence of the viral N protein, no green fluorescence was visualized. These results indicate that N protein has an essential role in the packaging of SARS-CoV RNA. A filter binding assay and competition analysis further demonstrated that the N-terminal and C-terminal regions of the SARS-CoV N protein each contain a binding activity specific to the viral RNA. Deletions that presumably disrupt the structure of the N-terminal domain diminished its RNA-binding activity. The GFP-PS-containing SARS-CoV VLPs are powerful tools for investigating the tissue tropism and pathogenesis of SARS-CoV.

Animals↗

Organization and expression of the chicken N-myc gene.

We cloned the chicken N-myc gene and analyzed its structure and expression. We found that it consisted of three exons with coding regions in exons 2 and 3. Comparison to mammalian N-myc genomic sequence indicated that nucleotide sequences of the 5'-flanking region, noncoding exon 1, and introns were not conserved, but coding and 3' noncoding sequences showed significant homology to mammalian N-myc. Alignment of deduced amino acid sequences of chicken and mammalian N-myc proteins revealed nine conserved domains interrupted by different lengths of nonhomologous sequences. Two of the domains were specific to N-myc proteins, and the other seven were common to c-myc proteins. Northern blot (immunoblot) and in situ hybridization analyses of 3.5-day-old chicken embryos revealed that high-level expression of the N-myc gene was confirmed to certain tissues, e.g., the central nervous system, neural crest derivatives, and mesenchyme of limb buds. In the beak and limb primordia, N-myc expression in the mesenchyme was higher toward the distal end, suggesting possible involvement in positional assignment of the tissue within the rudimentary structures.

Amino Acid Sequence↗

Molecular evolution and phylogeny of the atpB-rbcL spacer of chloroplast DNA in the true mosses.

The nucleotide variation of a noncoding region between the atpB and rbcL genes of the chloroplast genome was used to estimate the phylogeny of 11 species of true mosses (subclass Bryidae). The A+T rich (82.6%) spacer sequence is conserved with 48% of bases showing no variation between the ingroup and outgroup. Rooted at liverworts, Marchantia and Bazzania, the monophyly of true mosses was supported cladistically and statistically. A nonparametric Wilcoxon Signed-Ranks test Ts statistic for testing the taxonomic congruence showed no significant differences between gene trees and organism trees as well as between parsimony trees and neighbor-joining trees. The reconstructed phylogeny based on the atpB-rbcL spacer sequences indicated the validity of the division of acrocarpous and pleurocarpous mosses. The size of the chloroplast spacer in mosses fits into an evolutionary trend of increasing spacer length from liverworts through ferns to seed plants. According to the relative rate tests, the hypothesis of a molecular clock was supported in all species except for Thuidium, which evolved relatively fast. The evolutionary rate of the chloroplast DNA spacer in mosses was estimated to be (1.12 +/- 0.019) x 10(-10) nucleotides per site per year, which is close to the nonsynonymous substitution rates of the rbcL gene in the vascular plants. The constrained molecular evolution (total nucleotide substitutions, K approximately 0.0248) of the chloroplast DNA spacer is consistent with the slow evolution in morphological traits of mosses. Based on the calibrated evolutionary rate, the time of the divergence of true mosses was estimated to have been as early as 220 million years ago.

AT Rich Sequence↗

Natural antisense transcripts: sound or silence?

Antisense RNA was a rather uncommon term in a physiology environment until short interfering RNAs emerged as the tool of choice to knock down the expression of specific genes. As a consequence, the concept of RNA having regulatory potential became widely accepted. Yet, there is more to come. Computational studies suggest that between 15 and 25% of mammalian genes overlap, giving rise to pairs of sense and antisense RNAs. The resulting transcripts potentially interfere with each other's processing, thus representing examples of RNA-mediated gene regulation by endogenous, naturally occurring antisense transcripts. Concerns that the large-scale antisense transcription may represent transcriptional noise rather than a gene regulatory mechanism are strongly opposed by recent reports. A relatively small, well-defined group of antisense or noncoding transcripts is linked to monoallelic gene expression as observed in genomic imprinting, X chromosome inactivation, and clonal expression of B and T leukocytes. For the remaining, much larger group of bidirectionally transcribed genes, however, the physiological consequences of antisense transcription as well as the cellular mechanism(s) involved remain largely speculative.

Alleles↗

A probabilistic model for the evolution of RNA structure.

BACKGROUND: For the purposes of finding and aligning noncoding RNA gene- and cis-regulatory elements in multiple-genome datasets, it is useful to be able to derive multi-sequence stochastic grammars (and hence multiple alignment algorithms) systematically, starting from hypotheses about the various kinds of random mutation event and their rates. RESULTS: Here, we consider a highly simplified evolutionary model for RNA, called "The TKF91 Structure Tree" (following Thorne, Kishino and Felsenstein's 1991 model of sequence evolution with indels), which we have implemented for pairwise alignment as proof of principle for such an approach. The model, its strengths and its weaknesses are discussed with reference to four examples of functional ncRNA sequences: a riboswitch (guanine), a zipcode (nanos), a splicing factor (U4) and a ribozyme (RNase P). As shown by our visualisations of posterior probability matrices, the selected examples illustrate three different signatures of natural selection that are highly characteristic of ncRNA: (i) co-ordinated basepair substitutions, (ii) co-ordinated basepair indels and (iii) whole-stem indels. CONCLUSIONS: Although all three types of mutation "event" are built into our model, events of type (i) and (ii) are found to be better modeled than events of type (iii). Nevertheless, we hypothesise from the model's performance on pairwise alignments that it would form an adequate basis for a prototype multiple alignment and genefinding tool.

Evolution, Molecular↗

Simple repetitive sequences in the genome: structure and functional significance.

The current explosion of DNA sequence information has generated increasing evidence for the claim that noncoding repetitive DNA sequences present within and around different genes could play an important role in genetic control processes, although the precise role and mechanism by which these sequences function are poorly understood. Several of the simple repetitive sequences which occur in a large number of loci throughout the human and other eukaryotic genomes satisfy the sequence criteria for forming non-B DNA structures in vitro. We have summarized some of the features of three different types of simple repeats that highlight the importance of repetitive DNA in the control of gene expression and chromatin organization. (i) (TG/CA)n repeats are widespread and conserved in many loci. These sequences are associated with nucleosomes of varying linker length and may play a role in chromatin organization. These Z-potential sequences can help absorb superhelical stress during transcription and aid in recombination. (ii) Human telomeric repeat (TTAGGG)n adopts a novel quadruplex structure and exhibits unusual chromatin organization. This unusual structural motif could explain chromosome pairing and stability. (iii) Intragenic amplification of (CTG)n/(CAG)n trinucleotide repeat, which is now known to be associated with several genetic disorders, could down-regulate gene expression in vivo. The overall implications of these findings vis-à-vis repetitive sequences in the genome are summarized.

Animals↗

Massively parallel approaches for characterizing noncoding functional variation in human evolution.

The genetic differences underlying unique phenotypes in humans compared to our closest primate relatives have long remained a mystery. Similarly, the genetic basis of adaptations between human groups during our expansion across the globe is poorly characterized. Uncovering the downstream phenotypic consequences of these genetic variants has been difficult, as a substantial portion lies in noncoding regions, such as cis-regulatory elements (CREs). Here, we review recent high-throughput approaches to measure the functions of CREs and the impact of variation within them. CRISPR screens can directly perturb CREs in the genome to understand downstream impacts on gene expression and phenotypes, while massively parallel reporter assays can decipher the regulatory impact of sequence variants. Machine learning has begun to be able to predict regulatory function from sequence alone, further scaling our ability to characterize genome function. Applying these tools across diverse phenotypes, model systems, and ancestries is beginning to revolutionize our understanding of noncoding variation underlying human evolution.

Humans↗

Genome-wide analysis of chromosomal features repressing human immunodeficiency virus transcription.

We have investigated regulatory sequences in noncoding human DNA that are associated with repression of an integrated human immunodeficiency virus type 1 (HIV-1) promoter. HIV-1 integration results in the formation of precise and homogeneous junctions between viral and host DNA, but integration takes place at many locations. Thus, the variation in HIV-1 gene expression at different integration sites reports the activity of regulatory sequences at nearby chromosomal positions. Negative regulation of HIV transcription is of particular interest because of its association with maintaining HIV in a latent state in cells from infected patients. To identify chromosomal regulators of HIV transcription, we infected Jurkat T cells with an HIV-based vector transducing green fluorescent protein (GFP) and separated cells into populations containing well-expressed (GFP-positive) or poorly expressed (GFP-negative) proviruses. We then determined the chromosomal locations of the two classes by sequencing 971 junctions between viral and cellular DNA. Possible effects of endogenous cellular transcription were characterized by transcriptional profiling. Low-level GFP expression correlated with integration in (i) gene deserts, (ii) centromeric heterochromatin, and (iii) very highly expressed cellular genes. These data provide a genome-wide picture of chromosomal features that repress transcription and suggest models for transcriptional latency in cells from HIV-infected patients.

Base Sequence↗

RNA-protein interactions directed by the 3' end of human rhinovirus genomic RNA.

The replication of a picornavirus genomic RNA is a template-specific process involving the recognition of viral RNAs as target replication templates for the membrane-bound viral replication initiation complex. The virus-encoded RNA-dependent RNA polymerase, 3Dpol, is a major component of the replication complex; however, when supplied with a primed template, 3Dpol is capable of copying polyadenylated RNAs which are not of viral origin. Therefore, there must be some other molecular mechanism to direct the specific assembly of the replication initiation complex at the 3' end of viral genomic RNAs, presumably involving cis-acting binding determinants within the 3' noncoding region (3' NCR). This report describes the use of an in vitro UV cross-linking assay to identify proteins which interact with the 3' NCR of human rhinovirus 14 RNA. A cellular protein(s) was identified in cytoplasmic extracts from human rhinovirus 14-infected cells which had a marked binding preference for RNAs containing the rhinovirus 3' NCR sequence. This protein(s) showed reduced cross-linking efficiency for a 3' NCR with an engineered deletion. Virus recovered from RNA transfections with in vitro transcribed RNA containing the same 3' NCR deletion demonstrated a defective replication phenotype in vivo. Cross-linking experiments with RNAs containing the poliovirus 3' NCR and cytoplasmic extracts from poliovirus-infected cells produced an RNA-protein complex with indistinguishable electrophoretic properties, suggesting that the appearance of the cellular protein(s) may be a common phenomenon of picornavirus infection. We suggest that the observed cellular protein(s) is sequestered or modified as a result of rhinovirus or poliovirus infection and is utilized in viral RNA replication, perhaps by binding to the 3' NCR as a prerequisite for replication complex assembly at the 3' end of the viral genomic RNA.

Base Sequence↗

Use of conserved sequences from hepatitis C virus for the detection of viral RNA in infected sera by polymerase chain reaction.

Three oligonucleotide primer combinations selected from the 5' noncoding, the nucleocapsid and the putative nonstructural regions of the hepatitis C virus genome were compared in a nested polymerase chain reaction assay with respect to sensitivity and specificity for the detection of viral RNA in chimpanzee-infected and human-infected sera. Sera from both the acute and the chronic phase of the infection were obtained from 13 animals inoculated with five different non-A, non-B hepatitis strains and from seven cardiac surgery patients who had non-A, non-B hepatitis develop after transfusion and who had been tested in parallel for the presence of hepatitis C virus RNA and anti-C 100-3. A total of 90% of the acute-phase and 100% of the chronic-phase sera tested positive for hepatitis C virus RNA when the 5' noncoding-derived or the nucleocapsid-derived combinations were used; only 58% and 56%, respectively, gave positive results with the putative nonstructural primers, whereas 33% and 71%, respectively, scored positive for C100-3. Thus polymerase chain reaction primers selected from either the highly conserved 5' noncoding or nucleocapsid-regions appear to provide the sensitivity and the specificity necessary to detect low levels of hepatitis C virus RNA in both chimpanzee-infected and human-infected sera.

Animals↗

A genome-wide analysis of C/D and H/ACA-like small nucleolar RNAs in Trypanosoma brucei reveals a trypanosome-specific pattern of rRNA modification.

Small nucleolar RNAs (snoRNAs) constitute newly discovered noncoding small RNAs, most of which function in guiding modifications such as 2'-O-ribose methylation and pseudouridylation on rRNAs and snRNAs. To investigate the genome organization of Trypanosoma brucei snoRNAs and the pattern of rRNA modifications, we used a whole-genome approach to identify the repertoire of these guide RNAs. Twenty-one clusters encoding for 57 C/D snoRNAs and 34 H/ACA-like RNAs, which have the potential to direct 84 methylations and 32 pseudouridines, respectively, were identified. The number of 2'-O-methyls (Nms) identified on rRNA represent 80% of the expected modifications. The modifications guided by these RNAs suggest that trypanosomes contain many modifications and guide RNAs relative to their genome size. Interestingly, approximately 40% of the Nms are species-specific modifications that do not exist in yeast, humans, or plants, and 40% of the species-specific predicted modifications are located in unique positions outside the highly conserved domains. Although most of the guide RNAs were found in reiterated clusters, a few single-copy genes were identified. The large repertoire of modifications and guide RNAs in trypanosomes suggests that these modifications possibly play a central role in these parasites.

Animals↗

Methods in comparative genomics: genome correspondence, gene identification and regulatory motif discovery.

In Kellis et al. (2003), we reported the genome sequences of S. paradoxus, S. mikatae, and S. bayanus and compared these three yeast species to their close relative, S. cerevisiae. Genomewide comparative analysis allowed the identification of functionally important sequences, both coding and noncoding. In this companion paper we describe the mathematical and algorithmic results underpinning the analysis of these genomes. (1) We present methods for the automatic determination of genome correspondence. The algorithms enabled the automatic identification of orthologs for more than 90% of genes and intergenic regions across the four species despite the large number of duplicated genes in the yeast genome. The remaining ambiguities in the gene correspondence revealed recent gene family expansions in regions of rapid genomic change. (2) We present methods for the identification of protein-coding genes based on their patterns of nucleotide conservation across related species. We observed the pressure to conserve the reading frame of functional proteins and developed a test for gene identification with high sensitivity and specificity. We used this test to revisit the genome of S. cerevisiae, reducing the overall gene count by 500 genes (10% of previously annotated genes) and refining the gene structure of hundreds of genes. (3) We present novel methods for the systematic de novo identification of regulatory motifs. The methods do not rely on previous knowledge of gene function and in that way differ from the current literature on computational motif discovery. Based on genomewide conservation patterns of known motifs, we developed three conservation criteria that we used to discover novel motifs. We used an enumeration approach to select strongly conserved motif cores, which we extended and collapsed into a small number of candidate regulatory motifs. These include most previously known regulatory motifs as well as several noteworthy novel motifs. The majority of discovered motifs are enriched in functionally related genes, allowing us to infer a candidate function for novel motifs. Our results demonstrate the power of comparative genomics to further our understanding of any species. Our methods are validated by the extensive experimental knowledge in yeast and will be invaluable in the study of complex genomes like that of the human.

Algorithms↗