Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

A Fourier characteristic of coding sequences: origins and a non-Fourier approximation.

The 3-base periodicity, identified as a pronounced peak at the frequency N/3 (N is the length of the DNA sequence) of the Fourier power spectrum of protein coding regions, is used as a marker in gene-finding algorithms to distinguish protein coding regions (exons) and noncoding regions (introns) of genomes. In this paper, we reveal the explanation of this phenomenon which results from a nonuniform distribution of nucleotides in the three coding positions. There is a linear correlation between the nucleotide distributions in the three codon positions and the power spectrum at the frequency N/3. Furthermore, this study indicates the relationship between the length of a DNA sequence and the variance of nucleotide distributions and the average Fourier power spectrum, which is the noise signal in gene-finding methods. The results presented in this paper provide an efficient way to compute the Fourier power spectrum at N/3 and the noise signal in gene-finding methods by calculating the nucleotide distributions in the three codon positions.

Animals↗

The mitochondrial genome of the hemichordate Balanoglossus carnosus and the evolution of deuterostome mitochondria.

The complete nucleotide sequence of the mitochondrial genome of the hemichordate Balanoglossus carnosus (acorn worm) was determined. The arrangement of the genes encoding 13 protein, 22 tRNA, and 2 rRNA genes is essentially the same as in vertebrates, indicating that the vertebrate and hemichordate mitochondrial gene arrangement is close to that of their common ancestor, and, thus, that it has been conserved for more than 600 million years, whereas that of echinoderms has been rearranged extensively. The genetic code of hemichordate mitochondria is similar to that of echinoderms in that ATA encodes isoleucine and AGA serine, whereas the codons AAA and AGG, whose amino acid assignments also differ between echinoderms and vertebrates, are absent from the B. carnosus mitochondrial genome. There are three noncoding regions of length 277, 41, and 32 bp: the larger one is likely to be equivalent to the control region of other deuterostomes, while the two others may contain transcriptional promoters for genes encoded on the minor coding strand. Phylogenetic trees estimated from the inferred protein sequences indicate that hemichordates are a sister group of echinoderms.

Animals↗

Phylogenetic estimation of context-dependent substitution rates by maximum likelihood.

Nucleotide substitution in both coding and noncoding regions is context-dependent, in the sense that substitution rates depend on the identity of neighboring bases. Context-dependent substitution has been modeled in the case of two sequences and an unrooted phylogenetic tree, but it has only been accommodated in limited ways with more general phylogenies. In this article, extensions are presented to standard phylogenetic models that allow for better handling of context-dependent substitution, yet still permit exact inference at reasonable computational cost. The new models improve goodness of fit substantially for both coding and noncoding data. Considering context dependence leads to much larger improvements than does using a richer substitution model or allowing for rate variation across sites, under the assumption of site independence. The observed improvements appear to derive from three separate properties of the models: their explicit characterization of context-dependent substitution within N-tuples of adjacent sites, their ability to accommodate overlapping N-tuples, and their rich parameterization of the substitution process. Parameter estimation is accomplished using an expectation maximization algorithm, with a quasi-Newton algorithm for the maximization step; this approach is shown to be preferable to ordinary Newton methods for parameter-rich models. Overlapping tuples are efficiently handled by assuming Markov dependence of the observed bases at each site on those at the N - 1 preceding sites, and the required conditional probabilities are computed with an extension of Felsenstein's algorithm. Estimated substitution rates based on a data set of about 160,000 noncoding sites in mammalian genomes indicate a pronounced CpG effect, but they also suggest a complex overall pattern of context-dependent substitution, comprising a variety of subtle effects. Estimates based on about 3 million sites in coding regions demonstrate that amino acid substitution rates can be learned at the nucleotide level, and suggest that context effects across codon boundaries are significant.

Algorithms↗

Reversion to neurovirulence of the live-attenuated Sabin type 3 oral poliovirus vaccine.

The complete nucleotide sequence has been determined of a strain of poliovirus type 3, P3/119, isolated from the central nervous system of a victim of fatal vaccine-associated poliomyelitis. Comparison of this sequence with those obtained previously for the Sabin type 3 vaccine, P3/Leon 12a1b and its neurovirulent progenitor, P3/Leon/37, reveals that these three strains are on a direct geneaological lineage and therefore that P3/119 is a bona fide revertant of the vaccine. P3/119 differs in sequence from its attenuated vaccine parent at just seven positions. Only one of these differences, a mutation from U to C at position 472 in the presumed noncoding region of the genome, is a back mutation to the wild type sequence. Of the six other differences, three give rise to coding changes in virus structural proteins, two are silent changes in the major open reading frame of the genome and one affects the 3'-terminus just prior to the poly A tract. These differences indicate that there are three possible types of molecular change which could, singly or collectively, result in attenuation and reversion to neurovirulence of the Sabin type 3 vaccine.

Amino Acid Sequence↗

Human microsatellites applicable for analysis of genetic variation in apes and Old World monkeys.

In studies of the genetics and social structure of primate populations there is a need to develop highly variable genetic markers for characterizing mating success and the nature of population movement or change through time. Because of their highly polymorphic nature, relatively simple amplification and typing, and the possibility of noninvasive sampling, microsatellites have become the molecular tool of choice in such studies. However, until recently it was assumed that many microsatellite loci, which are primarily situated in noncoding regions of the genome, evolve too rapidly to be applicable in evolutionarily divergent species. This has often resulted in the time-consuming process of cloning and sequencing microsatellites in new species. Here we describe the application of 11 human microsatellite primer pairs to a large group of primate species. The loci described are informative in all major groups of apes and Old World monkeys, although levels of allelic variability and heterozygosity differ across species. We confirm that with the use of appropriate universally applicable PCR conditions, a subset of human microsatellites are informative genetic markers in a wide range of divergent primate taxa.

Animals↗

What is so special about oskar wild?

The amazing world of regulatory noncoding RNA has been at the center of biologists' attention in many different fields, from structural biology to transcriptional regulation and cell signaling. The latest example comes from developmental biology. A mutation in the Drosophila gene Oskar reveals a novel developmental function for the 3' untranslated region (UTR) of the oscar mRNA. This study further suggests that, when transcribed, the noncoding parts of the genome may well carry essential regulatory functions fundamental for the coordinated gene expression and development of multicellular organisms.

3' Untranslated Regions↗

Genotyping human papillomavirus type 16 isolates from persistently infected promiscuous individuals and cervical neoplasia patients.

Nucleotide sequence variation in the noncoding region of the genome of human papillomavirus type 16 (HPV16) was determined by direct sequencing and single-strand conformation polymorphism analysis of DNA fragments amplified by PCR. Individuals of diverse sexual promiscuity and/or cervicopathology were studied. In a group of 14 healthy, monogamous HPV16-positive females, only two HPV16 sequence variants could be documented. Among 17 females and 3 males with multiple sex partners and living in the same geographical region, nine sequence variants were found, whereas among 7 patients with cervical neoplasia from another region, five variants were detected. Although numbers are limited, in the group of individuals at high risk of acquiring a sexually transmitted disease or with cervical neoplasia, a larger number of HPV16 sequence variants was encountered (two types among 14 individuals versus nine types among 20; Fisher's exact test, P = 0.07). Seven of the individuals were sampled repeatedly over time. For these persistently infected women, no differences in HPV16 sequences were detected, irrespective of promiscuity, and persistence of a single viral variant, spread over multiple anatomic sites, for more than 2 years could be demonstrated. This indicates that viral persistence may be a common feature and that successful superinfection with a new variant may be rare, despite a potentially high frequency of viral reinoculation.

Base Sequence↗

Provirus of M7 baboon endogenous virus: nucleotide sequence of the gag-pol region.

A 3,023-base nucleotide sequence of the M7 baboon endogenous virus genome, spanning the 5' noncoding region as well as the entire gag gene and part of the pol gene, is reported. Within the 562-base 5' noncoding region, a 21-base sequence complementary to the OH terminus of tRNApro is located immediately downstream from the long terminal repeat. Amino acid sequences were deduced from the 1,596 nucleotides comprising the gag gene, and the four structural gag polypeptides, p12, p15, p30, and p10, appeared to be coded contiguously. Only one termination codon interrupted the M7 gag and pol genes. The data suggest that 55 additional amino acids may be attached to the NH2 terminus of the gag precursor protein. However, such a sequence was not detected in virions or in virus-infected cells. With the exception of the p15 region, nucleotide and amino acid sequences of the gag and pol regions of M7 virus exhibited strong homologies to those of Moloney leukemia virus.

Amino Acid Sequence↗

Properties of intracellular bovine papillomavirus chromatin.

Episomal nucleoprotein complexes of bovine papillomavirus type 1 (BPV-1) in transformed cells were exposed to DNase I treatment to localize hypersensitive regions. Such regions, which are indicative for gene expression, were found within the noncoding part of the genome, coinciding with the origin of replication and the 5' ends of most of the early mRNAs. However, there were also regions of hypersensitivity within the structural genes. These intragenic perturbations of the chromatin structure coincide with regulatory sequences at the DNA level. One of these regions maps in close proximity to a Z-DNA antibody-binding site which is located near the putative BPV-1 enhancer sequence.

Animals↗

Conservation of RNA-protein interactions among picornaviruses.

Picornavirus genomes encode unique 5' noncoding regions (5' NCRs) which are approximately 600 to 1,300 nucleotides in length, contain multiple upstream AUG codons, and display the ability to form extensive secondary structures. A number of recent reports have shown that picornavirus 5' NCRs are able to facilitate cap-independent internal initiation of translation. This mechanism of translation occurs in the absence of viral gene products, suggesting that the host cell contains the necessary components for the cap-independent internal initiation of translation of picornavirus RNAs as well as cellular mRNAs. In an attempt to identify some of the perhaps novel cellular proteins involved in this newly discovered mechanism of translation, we utilized RNA mobility shifts assays to identify and characterize interactions that occur between the 5'NCR of poliovirus type 1 (PV1) and cellular proteins. In this report, we describe two separate interactions between RNA structures from the 5' NCR of PV1 and proteins present in extracts from HeLa cells as well as other cell types. We describe the interaction between nucleotides 186 to 220 (stem-loop D) and a cellular protein(s) present in HeLa cell extracts. Mutational analysis of this stem-loop structure suggests that maintenance of a base-paired structure in the lower stem is necessary to present the sequences which directly interact with the protein(s). We also describe the interaction between nucleotides 220 to 460 (stem-loop E) and a cellular protein present in HeLa cell extracts. This RNA binding activity fractionates to a specific ammonium sulfate fraction (A cut) of a ribosomal salt wash. Mutational analysis of the stem-loop E structure suggests that the preservation of an extensive RNA structure is necessary for a strong interaction with the cellular protein(s), although smaller RNAs derived from this region of the 5' NCR can interact to lesser extents. Finally, we show that both of these RNA-protein interactions are conserved among the closely related enteroviruses PV1 and coxsackievirus type B3, human rhinovirus type 14, and the more distantly related cardiovirus Theiler's murine encephalomyelitis virus, suggesting that such RNA-protein interactions serve basic functions which are conserved and utilized by each of these picornaviruses.

Base Sequence↗

Temporal and spatial analysis of Sin Nombre virus quasispecies in naturally infected rodents.

Sin Nombre virus (SNV) is thought to establish a persistent infection in its natural reservoir, the deer mouse (Peromyscus maniculatus), despite a strong host immune response. SNV-specific neutralizing antibodies were routinely detected in deer mice which maintained virus RNA in the blood and lungs. To determine whether viral diversity played a role in SNV persistence and immune escape in deer mice, we measured the prevalence of virus quasispecies in infected rodents over time in a natural setting. Mark-recapture studies provided serial blood samples from naturally infected deer mice, which were sequentially analyzed for SNV diversity. Viral RNA was detected over a period of months in these rodents in the presence of circulating antibodies specific for SNV. Nucleotide and amino acid substitutions were observed in viral clones from all time points analyzed, including changes in the immunodominant domain of glycoprotein 1 and the 3' small segment noncoding region of the genome. Viral RNA was also detected in seven different organs of sacrificed deer mice. Analysis of organ-specific viral clones revealed major disparities in the level of viral diversity between organs, specifically between the spleen (high diversity) and the lung and liver (low diversity). These results demonstrate the ability of SNV to mutate and generate quasispecies in vivo, which may have implications for viral persistence and possible escape from the host immune system.

Amino Acid Sequence↗

Absence of internal ribosome entry site-mediated tissue specificity in the translation of a bicistronic transgene.

The 5' noncoding regions of the genomes of picornaviruses form a complex structure that directs cap-independent initiation of translation. This structure has been termed the internal ribosome entry site (IRES). The efficiency of translation initiation was shown, in vitro, to be influenced by the binding of cellular factors to the IRES. Hence, we hypothesized that the IRES might control picornavirus tropism. In order to test this possibility, we made a bicistronic construct in which translation of the luciferase gene is controlled by the IRES of Theiler's murine encephalomyelitis virus. In vitro, we observed that the IRES functions in various cell types and in macrophages, irrespective of their activation state. In vivo, we observed that the IRES is functional in different tissues of transgenic mice. Thus, it seems that the IRES is not an essential determinant of Theiler's virus tropism. On the other hand, the age of the mouse could be critical for IRES function. Indeed, the IRES was found to be more efficient in young mice. Picornavirus IRESs are becoming popular tools in transgenesis technology, since they allow the expression of two genes from the same transcription unit. Our results show that the Theiler's virus IRES is functional in cells of different origins and that it is thus a broad-spectrum tool. The possible age dependency of the IRES function, however, could be a drawback for gene expression in adult mice.

Animals↗

Sequence requirements for viral RNA replication and VPg uridylylation directed by the internal cis-acting replication element (cre) of human rhinovirus type 14.

Until recently, the cis-acting signals required for replication of picornaviral RNAs were believed to be restricted to the 5' and 3' noncoding regions of the genome. However, an RNA stem-loop in the VP1-coding sequence of human rhinovirus type 14 (HRV-14) is essential for viral minus-strand RNA synthesis (K. L. McKnight and S. M. Lemon, RNA 4:1569-1584, 1998). The nucleotide sequence of the apical loop of this internal cis-acting replication element (cre) was critical for RNA synthesis, while secondary RNA structure, but not primary sequence, was shown to be important within the duplex stem. Similar cres have since been identified in other picornaviral genomes. These RNA segments appear to serve as template for the uridylylation of the genome-linked protein, VPg, providing the VPg-pUpU primer required for viral RNA transcription (A. V. Paul et al., J. Virol. 74:10359-10370, 2000). Here, we show that the minimal functional HRV-14 cre resides within a 33-nucleotide (nt) RNA segment that is predicted to form a simple stem-loop with a 14-nt loop sequence. An extensive mutational analysis involving every possible base substitution at each position within the loop segment defined the sequence that is required within this loop for efficient replication of subgenomic HRV-14 replicon RNAs. These results indicate that three consecutive adenosine residues (nt 2367 to 2369) within the 5' half of this loop are critically important for cre function and suggest that a common RNNNAARNNNNNNR loop motif exists among the cre sequences of enteroviruses and rhinoviruses. We found a direct, positive correlation between the capacity of mutated cres to support RNA replication and their ability to function as template in an in vitro VPg uridylylation reaction, suggesting that these functions are intimately linked. These data thus define more precisely the sequence and structural requirements of the HRV-14 cre and provide additional support for a model in which the role of the cre in RNA replication is to act as template for VPg uridylylation.

Base Sequence↗

Poliovirus internal ribosome entry segment structure alterations that specifically affect function in neuronal cells: molecular genetic analysis.

Translation of poliovirus RNA is driven by an internal ribosome entry segment (IRES) present in the 5' noncoding region of the genomic RNA. This IRES is structured into several domains, including domain V, which contains a large lateral bulge-loop whose predicted secondary structure is unclear. The primary sequence of this bulge-loop is strongly conserved within enteroviruses and rhinoviruses: it encompasses two GNAA motifs which could participate in intrabulge base pairing or (in one case) could be presented as a GNRA tetraloop. We have begun to address the question of the significance of the sequence conservation observed among enterovirus reference strains and field isolates by using a comprehensive site-directed mutagenesis program targeted to these two GNAA motifs. Mutants were analyzed functionally in terms of (i) viability and growth kinetics in both HeLa and neuronal cell lines, (ii) structural analyses by biochemical probing of the RNA, and (iii) translation initiation efficiencies in vitro in rabbit reticulocyte lysates supplemented with HeLa or neuronal cell extracts. Phenotypic analyses showed that only viruses with both GNAA motifs destroyed were significantly affected in their growth capacities, which correlated with in vitro translation defects. The phenotypic defects were strongly exacerbated in neuronal cells, where a temperature-sensitive phenotype could be revealed at between 37 and 39.5 degrees C. Biochemical probing of mutated domain V, compared to the wild type, demonstrated that such mutations lead to significant structural perturbations. Interestingly, revertant viruses possessed compensatory mutations which were distant from the primary mutations in terms of sequence and secondary structure, suggesting that intradomain tertiary interactions could exist within domain V of the IRES.

Amino Acid Motifs↗

Characterization of mouse cellular deoxyribonucleic acid homologous to Abelson murine leukemia virus-specific sequences.

The genome of Abelson murine leukemia virus (A-MuLV) consists of sequences derived from both BALB/c mouse deoxyribonucleic acid and the genome of Moloney murine leukemia virus. Using deoxyribonucleic acid linear intermediates as a source of retroviral deoxyribonucleic acid, we isolated a recombinant plasmid which contained 1.9 kilobases of the 3.5-kilobase mouse-derived sequences found in A-MuLV (A-MuLV-specific sequences). We used this clone, designated pSA-17, as a probe restriction enzyme and Southern blot analyses to examine the arrangement of homologous sequences in BALB/c deoxyribonucleic acid (endogenous Abelson sequences). The endogenous Abelson sequences within the mouse genome were interrupted by noncoding regions, suggesting that a rearrangement of the cell sequences was required to produce the sequence found in the virus. Endogenous Abelson sequences were arranged similarly in mice that were susceptible to A-MuLV tumors and in mice that were resistant to A-MuLV tumors. An examination of three BALB/c plasmacytomas and a BALB/c early B-cell tumor likewise revealed no alteration in the arrangement of the endogenous Abelson sequences. Homology to pSA-17 was also observed in deoxyribonucleic acids prepared from rat, hamster, chicken, and human cells. An isolate of A-MuLV which encoded a 160,000-dalton transforming protein (P160) contained 700 more base pairs of mouse sequences than the standard A-MuLV isolate, which encoded a 120,000-dalton transforming protein (P120).

3T3 Cells↗

Intron 1 sequences are required for pancreatic expression of the human proglucagon gene.

The mammalian proglucagon gene is expressed in pancreatic islet A-cells, intestinal L-cells, and select neurons of the brain, where posttranslational processing results in the liberation of a unique profile of peptides. Despite the importance of proglucagon-derived peptides in human biology, little is known about the regulation of the human gene, as the rat gene has been the preferred model for understanding the regulation of proglucagon gene expression. Previously, we have shown that although the immediate promoter region of the rat proglucagon gene is sufficient for expression in pancreatic islet cells, the homologous human proglucagon promoter sequences are not sufficient. We have now used a comparative genomic approach to identify noncoding sequences near the human proglucagon gene that are conserved among mammals, and thus potentially are regulatory sequences. Our alignments identified three evolutionarily conserved noncoding regions (ECR), one is the immediate promoter region (ECR1), the second is about 5 kb 5' to the mRNA start site (ECR2), and the third is near the 3' end of the first intron (ECR3). Our in vitro transient transfection assays with reporter gene constructs that include the human ECR3 support expression in rodent islet cell lines. Complementary studies with transgenic mice possessing a reporter gene regulated by a human proglucagon gene promoter-intron 1 (including ECR3) sequences express the reporter gene in the pancreas, as well as the intestine and selected neurons. These studies suggest that conserved sequences within intron 1 of the human proglucagon gene are important for expression in the pancreas.

Animals↗

Intronic microRNA (miRNA).

Nearly 97% of the human genome is composed of noncoding DNA, which varies from one species to another. Changes in these sequences often manifest themselves in clinical and circumstantial malfunction. Numerous genes in these non-protein-coding regions encode microRNAs, which are responsible for RNA-mediated gene silencing through RNA interference (RNAi)-like pathways. MicroRNAs (miRNAs), small single-stranded regulatory RNAs capable of interfering with intracellular messenger RNAs (mRNAs) with complete or partial complementarity, are useful for the design of new therapies against cancer polymorphisms and viral mutations. Currently, many varieties of miRNA are widely reported in plants, animals, and even microbes. Intron-derived microRNA (Id-miRNA) is a new class of miRNA derived from the processing of gene introns. The intronic miRNA requires type-II RNA polymerases (Pol-II) and spliceosomal components for their biogenesis. Several kinds of Id-miRNA have been identified in C elegans, mouse, and human cells; however, neither function nor application has been reported. Here, we show for the first time that intron-derived miRNAs are able to induce RNA interference in not only human and mouse cells, but in also zebrafish, chicken embryos, and adult mice, demonstrating the evolutionary preservation of intron-mediated gene silencing via functional miRNA in cell and in vivo. These findings suggest an intracellular miRNA-mediated gene regulatory system, fine-tuning the degradation of protein-coding messenger RNAs.

Journal Article↗

A systematic search for new mammalian noncoding RNAs indicates little conserved intergenic transcription.

BACKGROUND: Systematic identification and functional characterization of novel types of noncoding (nc)RNA in genomes is more difficult than it is for protein coding mRNAs, since ncRNAs typically do not possess sequence features such as splicing or translation signals, or long open reading frames. Recent "tiling" microarray studies have reported that a surprisingly larger proportion of mammalian genomes is transcribed than was previously anticipated. However, these non-genic transcripts often appear to be low in abundance, and their functional significance is not known. RESULTS: To systematically search for functional ncRNAs, we designed microarrays to detect 3,478 intergenic and intronic sequences that are conserved between the human, mouse, and rat genomes, and that score highly by other criteria that characterize ncRNAs. We probed these arrays with total RNA isolated from 16 wild-type mouse tissues. Among 55 candidates for highly-expressed novel ncRNAs tested by northern blotting, eight were confirmed as small, highly-and ubiquitously-expressed RNAs in mouse. Of the eight, five were also detected in rat tissues, but none were detected at appreciable levels in human tissues or cultured cells. CONCLUSION: Since the sequence and expression of most known coding transcripts and functional ncRNAs is conserved between human and mouse, the lack of northern-detectable expression in human cells and tissues of the novel mouse and rat ncRNAs that we identified suggests that they are not functional or possibly have rodent-specific functions. Our results confirm that relatively little of the intergenic sequence conserved between human, mouse and rat is transcribed at high levels in mammalian tissues, possibly suggesting a limited role for transcribed intergenic and intronic sequences as independent functional elements.

Animals↗