Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

[Comparative approach to correction of annotated gene starts in complete bacterial genomes].

A method for refining the beginnings of genes and a search for shifts of the reading frame is proposed. The method is based on a comparison of nucleotide and amino acid sequences of homologous genes of related organisms. The algorithm is based on the fact that the rate of changes in the protein-coding regions of the genome is substantially lower than that of noncoding regions. A modification of the Smith-Waterman algorithm is proposed, which makes it possible to align the amino acid sequences obtained by formal translation of the starting nucleotide sequences by taking into account a possible shift of the reading frame. The algorithm has been implemented in the package of ORTOLOGATOR-GeneCorrector programs. Testing the program showed that the approach enables one to detect a wrong annotation of the beginnings in 1% of genes (even in well-studied organisms such as Escherichia coli) and identify several (approximately 10) shifts of the open reading frame. Thus, the algorithm can be used at both the initial and final stages of analysis of the genome.

Algorithms↗

Nucleotide sequence of segment S9 of the genome of rice gall dwarf virus.

DNA complementary to the ninth largest (S9) of the 12 genome segments of rice gall dwarf virus (RGDV) was cloned and its sequence was determined. It is 1202 nucleotides in length and contains one open reading frame which extends for 969 nucleotides from nucleotide 26. It encodes a polypeptide of 323 amino acids with an Mr of 35,560. The dinucleotide sequence at the 5' end and the trinucleotide sequence at the 3' end of the plus strand, 5' GG--GAU 3', which are present in the RNA of both wound tumour virus (WTV) and rice dwarf virus (RDV), were also found in RGDV genome segment S9. The nucleotide sequences in the noncoding region at the 5' terminus and in the 15 nucleotides at the 3' terminus, which form an imperfect inverted repeat of 10 bp together with the 5' terminus, are approximately 70% homologous with those of the WTV genome segment S9, but only 30% and 50% homologous with the respective termini of RDV S9.

Amino Acid Sequence↗

[Nucleotide sequence of the noncoding regions of measles virus stain CC-47 and comparison with other measles viruses].

BACKGROUND: To determine the nucleotide sequence of noncoding regions of Measles virus strain Changchun-47 (CC-47) and to compare them with other measles virus for revealing some vaccine-related information. METHODS: Six overlapped fragments, that covers complete genome of CC-47 were amplified by using RT-PCR, all of the amplified fragments have been cloned and sequenced. RESULTS: Seven noncoding regions lie in the genome sequence of CC-47 separated by six structure genes. The noncoding regions of CC-47 contain 1791 nucleotides totally. Comparing of the noncoding regions of CC-47 with other wild type strains and five Edmonston-derived vaccine strains, four nucleotide substitutions were shared by nearly all vaccine strains. Two of these were in the genomic 3 terminal transcriptional control region: position 26 (U->G), 42 (U->G); the other were in the F mRNA 5 -untranslated region of M/F intergenic region. These site substitutions may influence the efficiency of mRNA synthesis, processing, and translation, as wel l as genome replication and encapsidation. CONCLUSIONS: The nucleotide substitutions shared by different genotype measles virus vaccine strains may be related to virus growth in semipermissive cell or process of attenuation, which provide useful data for molecular epidemiology of measles virus. The nucleotide substitutions for attenuation were involved in several regions other than one definite region.

3' Untranslated Regions↗

Analysis of conserved noncoding DNA in Drosophila reveals similar constraints in intergenic and intronic sequences.

Comparative genomic approaches to gene and cis-regulatory prediction are based on the principle that differential DNA sequence conservation reflects variation in functional constraint. Using this principle, we analyze noncoding sequence conservation in Drosophila for 40 loci with known or suspected cis-regulatory function encompassing >100 kb of DNA. We estimate the fraction of noncoding DNA conserved in both intergenic and intronic regions and describe the length distribution of ungapped conserved noncoding blocks. On average, 22%-26% of noncoding sequences surveyed are conserved in Drosophila, with median block length approximately 19 bp. We show that point substitution in conserved noncoding blocks exhibits transition bias as well as lineage effects in base composition, and occurs more than an order of magnitude more frequently than insertion/deletion (indel) substitution. Overall, patterns of noncoding DNA structure and evolution differ remarkably little between intergenic and intronic conserved blocks, suggesting that the effects of transcription per se contribute minimally to the constraints operating on these sequences. The results of this study have implications for the development of alignment and prediction algorithms specific to noncoding DNA, as well as for models of cis-regulatory DNA sequence evolution.

Animals↗

Novel genes derived from noncoding DNA in Drosophila melanogaster are frequently X-linked and exhibit testis-biased expression.

Descriptions of recently evolved genes suggest several mechanisms of origin including exon shuffling, gene fission/fusion, retrotransposition, duplication-divergence, and lateral gene transfer, all of which involve recruitment of preexisting genes or genetic elements into new function. The importance of noncoding DNA in the origin of novel genes remains an open question. We used the well annotated genome of the genetic model system Drosophila melanogaster and genome sequences of related species to carry out a whole-genome search for new D. melanogaster genes that are derived from noncoding DNA. Here, we describe five such genes, four of which are X-linked. Our RT-PCR experiments show that all five putative novel genes are expressed predominantly in testes. These data support the idea that these novel genes are derived from ancestral noncoding sequence and that new, favored genes are likely to invade populations under selective pressures relating to male reproduction.

Animals↗

Exploring the Effect of Whole-Genome Duplication on Salmonid LincRNA Repertoire.

Long intergenic non-coding RNAs (lincRNAs) are key epigenetic regulators of genome function, yet their evolutionary dynamics following whole-genome duplication (WGD) events remain poorly understood. Salmonids, which underwent a lineage-specific autotetraploidization (salmonid-specific WGD, ~88-100 million years ago), provide an excellent model to investigate the retention, divergence, and functional potential of recently duplicated non-coding elements. LincRNA repertoires were compared across five genome-annotated salmonids (Oncorhynchus tshawytscha, O. kisutch, O. mykiss, Salmo salar, and S. trutta) and their closest non-duplicated relative, northern pike (Esox lucius). LincRNAs represented ~5-7% of annotated genes in all salmonids except S. salar (18%). Sequence conservation was low relative to coding genes, with only 11-68 highly similar (e-value < 1 &#xd7; 10-30; similarity > 70% and alignments > 100 nucleotides) putative orthologues shared between salmonids and northern pike, and 161-338 among salmonids alone. Synteny conservation was modest in lincRNAs, with lower conservation in putative orthologues (8-16%) compared to putative ohnologues (8-33%). Secondary structure conservation was associated with sequence similarity (&#x3c1; = -0.45; p = 2.2 &#xd7; 10-16), and the association was stronger among WGD ohnologues than orthologues. In S. salar and O. mykiss, lincRNA putative ohnologues showed weaker expression correlations than coding genes, suggesting widespread regulatory divergence, possibly through neo- and subfunctionalisation. Conserved salmonid lincRNAs showed enriched predicted interactions with miRNAs involved in tumour suppression, brain, bone, and muscle development (e.g., miR-455, miR-365, miR124, miR-133a, miR-140, and miR-9), a finding supported by limited transcriptomic data. Although salmonid WGD expanded lincRNA repertoires, lincRNAs have undergone rapid sequence and transcriptional divergence, with limited conservation across species based on sequence similarity, chromosomal position, synteny, and secondary structure. A subset of conserved lincRNAs retains structural features and regulatory signatures consistent with roles as miRNA sponges in brain, skeletal, and muscle development and tumour suppression, potentially acting within conserved regulatory networks. These findings provide new insights into lincRNA evolution following genome duplication and highlight the need for experimental validation of their regulatory functions.

Animals↗

Generation of infectious pancreatic necrosis virus from cloned cDNA.

We developed a reverse genetics system for infectious pancreatic necrosis virus (IPNV), a prototype virus of the Birnaviridae family, with the use of plus-stranded RNA transcripts derived from cloned cDNA. Full-length cDNA clones of the IPNV genome that contained the entire coding and noncoding regions of RNA segments A and B were constructed. Segment A encodes a 106-kDa precursor protein which is cleaved to yield mature VP2, nonstructural protease, and VP3 proteins, whereas segment B encodes the RNA-dependent RNA polymerase VP1. Plus-sense RNA transcripts of both segments were prepared by in vitro transcription of linearized plasmids with T7 RNA polymerase. Transfection of chinook salmon embryo (CHSE) cells with combined transcripts of segments A and B generated infectious IPNV particles 10 days posttransfection. Furthermore, a transfectant virus containing a genetically tagged sequence was generated to confirm the feasibility of this system. The presence and specificity of the recovered virus were ascertained by immunofluorescence staining of infected CHSE cells with rabbit anti-IPNV serum and by nucleotide sequence analysis. In addition, 3'-terminal sequence analysis of RNA from the recovered virus showed that extraneous nucleotides synthesized at the 3' end during in vitro transcription were precisely trimmed or excluded during replication, and hence these were not incorporated into the genome. An attempt was made to determine if RNA-dependent RNA polymerase of IPNV and infectious bursal disease virus (IBDV), another birnavirus, can support virus rescue in heterologous combinations. Thus, CHSE cells were transfected with transcripts derived from IPNV segment A and IBDV segment B and Vero cells were transfected with transcripts derived from IBDV segment A and IPNV segment B. In either case, no infectious IPNV or IBDV particles were generated even after a third passage in cell culture, suggesting that viral RNA-dependent RNA polymerase is species specific. However, the reverse genetics system for IPNV that we developed will greatly facilitate studies of viral replication and pathogenesis and the design of a new generation of live attenuated vaccines.

Animals↗

Characterization of the promoter region of the rat neprilysin gene.

The neprilysin gene is composed of three distinct 5' noncoding exons which can be joined to the first coding exon to generate multiple mRNA species, all encoding the same protein. Genomic fragments containing upstream sequences of each of these three noncoding exons from the rat neprilysin gene were subcloned in the promoterless vector pXp1, which contains the luciferase reporter gene. Expression was compared between a neprilysin positive human spinal cord cell line, HSC-2, and neprilysin negative lines MCF-7, a human breast adenocarcinoma cell line and Hep G2, a human liver carcinoma cell line. The first and second promoter regions showed high activity in the positive cell line, but low activity in the negative cell lines. An analysis of the exon 1 promoter region showed that the proximal 85 nucleotides exhibited basal promoter activity. An enhancer-like sequence was found to be located within a 22-bp fragment located at -136 to -115. Scanning mutagenesis of a 29-bp fragment containing the enhancer-like sequence showed that changes in each 5- or 6-bp segment throughout this fragment decreased activity; however, mutations of the segment encompassing positions 19 to 24 eliminated >98% of the promoter activity. Binding of nuclear proteins from HSC-2 cells to this 29-bp fragment was observed by gel shift analysis. The ability of mutations within the 29-bp fragment to affect enhancer activity correlated with the ability of these mutant oligonucleotides to compete for the wild-type sequence in gel shift assays.

Animals↗

The phosphoenolpyruvate carboxylase gene family of Sorghum: promoter structures, amino acid sequences and expression of genes.

Two different members of the phosphoenolpyruvate carboxylase(PEPC)-encoding multigene family (clones lambda CP21 and lambda CP46) have been isolated from a Sorghum vulgare lambda EMBL4 genomic library. The use of the 3'-noncoding regions to probe Northern blots of RNA from roots, etiolated leaves and green leaves indicated that lambda CP21 and lambda CP46 encode the C3- and C4-type leaf PEPC isoforms, respectively. The lambda CP21 clone is expressed in the three tissues and is not light-regulated, whereas lambda CP46 is only expressed in greening leaves. The nucleotide sequence of the 5'-flanking DNA (520 bp) has been determined for both genes. For lambda CP46, several direct repeats were located in this region with similarities to sequences found in other light-regulated genes, but not in lambda CP21. The deduced amino acid sequences of the two S. vulgare PEPC proteins are 75% identical.

Amino Acid Sequence↗

Structural and evolutionary studies on sterol 14-demethylase P450 (CYP51), the most conserved P450 monooxygenase: I. Structural analyses of the gene and multiple sizes of mRNA.

The structure of rat CYP51 gene encoding sterol 14-demethylase was examined. The CYP51 gene spanned about 18 kb and contained 10 exons. The copy number of CYP51 in the rat genome was determined to be one. In addition, one CYP51 processed (intron-less) pseudogene covering the coding and ca. 600-bp 3'-noncoding sequences of CYP51 cDNA was found in the rat genome. Multiple transcription initiation sites were predicted by primer extension and 5'-RACE methods using poly(A)+ RNA from liver, ovary, and testis, and the major ones were located at 126 and 123 nucleotides upstream from the initiation ATG codon. The primer extension also showed several minor sites around the major ones. In addition to these sites, other minor initiation sites were also predicted at around 330 and 460 nucleotides upstream from the initiation ATG codon. No TATA box was found in the putative promoter region, but multiple GC boxes were found around the cap sites, supporting the previously inferred housekeeping nature of CYP51 gene and the existence of the multiple transcription initiation sites. A few consensus transcription regulatory elements such as CRE were found in the 5'-flanking region. Four polyadenylation signals were found in the 3'-noncoding region by the 3'-RACE method. Three of them were used to generate 3.1-, 2.7-, and 2.3-kb mRNAs in liver and ovary. The remaining one was used only in testis to generate 1.9-kb mRNA having an unusually short trailer sequence, suggesting a specific regulatory mechanism for generating CYP51 mRNA in testis different from that in liver and ovary.

Amino Acid Sequence↗

Transgenic plants that express genes including the 3' untranslated region of the turnip yellow mosaic virus (TYMV) genome are partially protected against TYMV infection.

In order to evaluate new possibilities for protecting plants against virus infection by interference with viral replication, two chimeric genes were constructed in which the (+) strand 3'-terminal 100 nucleotides (nt) of the noncoding region of the turnip yellow mosaic virus (TYMV) genome were placed downstream from the sense or antisense cat coding region. The two chimeric genes were then introduced into the genome of rapeseed (Brassica napus) using an Agrobacterium rhizogenes vector system. Plants expressing high levels of either chimeric gene showed partial protection against infection by TYMV RNA or virions. One interesting feature of the protection is that a proportion of the inoculated transgenic plants does not become infected. Protection was overcome when the inoculum concentration was increased. RNA complementary to the initial transcript was detected after infection.

Base Sequence↗

Sequence and translation of the murine coronavirus 5'-end genomic RNA reveals the N-terminal structure of the putative RNA polymerase.

A 28-kilodalton protein has been suggested to be the amino-terminal protein cleavage product of the putative coronavirus RNA polymerase (gene A) (M.R. Denison and S. Perlman, Virology 157:565-568, 1987). To elucidate the structure and mechanism of synthesis of this protein, the nucleotide sequence of the 5' 2.0 kilobases of the coronavirus mouse hepatitis virus strain JHM genome was determined. This sequence contains a single, long open reading frame and predicts a highly basic amino-terminal region. Cell-free translation of RNAs transcribed in vitro from DNAs containing gene A sequences in pT7 vectors yielded proteins initiated from the 5'-most optimal initiation codon at position 215 from the 5' end of the genome. The sequence preceding this initiation codon predicts the presence of a stable hairpin loop structure. The presence of an RNA secondary structure at the 5' end of the RNA genome is supported by the observation that gene A sequences were more efficiently translated in vitro when upstream noncoding sequences were removed. By comparing the translation products of virion genomic RNA and in vitro transcribed RNAs, we established that our clones encompassing the 5'-end mouse hepatitis virus genomic RNA encode the 28-kilodalton N-terminal cleavage product of the gene A protein. Possible cleavage sites for this protein are proposed.

Amino Acid Sequence↗

CYP3A variation and the evolution of salt-sensitivity variants.

Members of the cytochrome P450 3A subfamily catalyze the metabolism of endogenous substrates, environmental carcinogens, and clinically important exogenous compounds, such as prescription drugs and therapeutic agents. In particular, the CYP3A4 and CYP3A5 genes play an especially important role in pharmacogenetics, since they metabolize >50% of the drugs on the market. However, known genetic variants at these two loci are not sufficient to account for the observed phenotypic variability in drug response. We used a comparative genomics approach to identify conserved coding and noncoding regions at these genes and resequenced them in three ethnically diverse human populations. We show that remarkable interpopulation differences exist with regard to frequency spectrum and haplotype structure. The non-African samples are characterized by a marked excess of rare variants and the presence of a homogeneous group of long-range haplotypes at high frequency. The CYP3A5*1/*3 polymorphism, which is likely to influence salt and water retention and risk for salt-sensitive hypertension, was genotyped in >1,000 individuals from 52 worldwide population samples. The results reveal an unusual geographic pattern whereby the CYP3A5*3 frequency shows extreme variation across human populations and is significantly correlated with distance from the equator. Furthermore, we show that an unlinked variant, AGT M235T, previously implicated in hypertension and pre-eclampsia, exhibits a similar geographic distribution and is significantly correlated in frequency with CYP3A5*1/*3. Taken together, these results suggest that variants that influence salt homeostasis were the targets of a shared selective pressure that resulted from an environmental variable correlated with latitude.

Black or African American↗

Evolution of genome size: multilevel selection, mutation bias or dynamical chaos?

In the past two years, new data on conceptual aspects of the evolution of eukaryotic genome size have appeared, including the adaptivity of genome enlargement, the mechanisms of genome size change and the relation of genome size to organismal complexity. New data on the hypotheses of "selfish DNA" and "mutational equilibrium" have been recently obtained. A relationship is emerging between the intragenomic distribution of noncoding DNA and differential gene expression, which suggests that noncoding DNA is involved in epigenetic organization of the genome and organismal complexity. The standpoint of dynamical chaos, which integrates multilevel selection and mutation biases, may provide a framework for studying the evolution of genome size.

Animals↗

Novel mitochondrial gene content and gene arrangement indicate illegitimate inter-mtDNA recombination in the chigger mite, Leptotrombidium pallidum.

To better understand the evolution of mitochondrial (mt) genomes in the Acari (mites and ticks), we sequenced the mt genome of the chigger mite, Leptotrombidium pallidum (Arthropoda: Acari: Acariformes). This genome is highly rearranged relative to that of the hypothetical ancestor of the arthropods and the other species of Acari studied. The mt genome of L. pallidum has two genes for large subunit rRNA, a pseudogene for small subunit rRNA, and four nearly identical large noncoding regions. Nineteen of the 22 tRNAs encoded by this genome apparently lack either a T-arm or a D-arm. Further, the mt genome of L. pallidum has two distantly separated sections with identical sequences but opposite orientations of transcription. This arrangement cannot be accounted for by homologous recombination or by previously known mechanisms of mt gene rearrangement. The most plausible explanation for the origin of this arrangement is illegitimate inter-mtDNA recombination, which has not been reported previously in animals. In light of the evidence from previous experiments on recombination in nuclear and mt genomes of animals, we propose a model of illegitimate inter-mtDNA recombination to account for the novel gene content and gene arrangement in the mt genome of L. pallidum.

Animals↗

Noncoding RNA synthesis and loss of Polycomb group repression accompanies the colinear activation of the human HOXA cluster.

The ratio of noncoding to protein coding DNA rises with the complexity of the organism, culminating in nearly 99% of nonprotein coding DNA in humans. Nevertheless, a large portion of these regions is transcribed, creating the alleged paradox that noncoding RNA (ncRNA) represents the largest output of the human genome. Such a complex scenario may include epigenetic mechanisms where ncRNAs would be involved in chromatin regulation. We have investigated the intergenic, noncoding transcriptomes of mammalian HOX clusters. We show that "opposite strand transcription" from the intergenic spacer regions in the human HOXA cluster correlates with the activity state of adjacent HOXA genes. This noncoding transcription is regulated by the retinoic acid morphogen and follows the colinear activation pattern of the cluster. Opening of the cluster at sites of activation of intergenic transcripts is accompanied by changes in histone modifications and a loss of interaction with Polycomb group (PcG) repressive complexes. We propose that noncoding transcription is of fundamental importance for the opening and maintenance of the active state of HOX clusters.

Animals↗

Gene-rich and gene-poor chromosomal regions have different locations in the interphase nuclei of cold-blooded vertebrates.

In situ hybridizations of single-copy GC-rich, gene-rich and GC-poor, gene-poor chicken DNA allowed us to localize the gene-rich and the gene-poor chromosomal regions in interphase nuclei of cold-blooded vertebrates. Our results showed that the gene-rich regions from amphibians (Rana esculenta) and reptiles (Podarcis sicula) occupy the more internal part of the nuclei, whereas the gene-poor regions occupy the periphery. This finding is similar to that previously reported in warm-blooded vertebrates, in spite of the lower GC levels of the gene-rich regions of cold-blooded vertebrates. This suggests that this similarity extends to chromatin structure, which is more open in the gene-rich regions of both mammals and birds and more compact in the gene-poor regions. In turn, this may explain why the compositional transition undergone by the genome at the emergence of homeothermy did not involve the entire ancestral genome but only a small part of it, and why it involved both coding and noncoding sequences. Indeed, the GC level increased only in that part of the genome that needed a thermodynamic stabilization, namely in the more open gene-rich chromatin of the nuclear interior, whereas the gene-poor chromatin of the periphery was stabilized by its own compact structure.

Animals↗

Amplification of noncoding chloroplast DNA for phylogenetic studies in lycophytes and monilophytes with a comparative example of relative phylogenetic utility from Ophioglossaceae.

Noncoding DNA sequences from numerous regions of the chloroplast genome have provided a significant source of characters for phylogenetic studies in seed plants. In lycophytes and monilophytes (leptosporangiate ferns, eusporangiate ferns, Psilotaceae, and Equisetaceae), on the other hand, relatively few noncoding chloroplast DNA regions have been explored. We screened 30 lycophyte and monilophyte species to determine the potential utility of PCR amplification primers for 18 noncoding chloroplast DNA regions that have previously been used in seed plant studies. Of these primer sets eight appear to be nearly universally capable of amplifying lycophyte and monilophyte DNAs, and an additional six are useful in at least some groups. To further explore the application of noncoding chloroplast DNA, we analyzed the relative phylogenetic utility of five cpDNA regions for resolving relationships in Botrychium s.l. (Ophioglossaceae). Previous studies have evaluated both the gene rbcL and the trnL(UAA)-trnF(GAA) intergenic spacer in this group. To these published data we added sequences of the trnS(GCU)-trnG(UUC) intergenic spacer + the trnG(UUC) intron region, the trnS(GGA)-rpS4 intergenic spacer+rpS4 gene, and the rpL16 intron. Both the trnS(GCU)-trnG(UUC) and rpL16 regions are highly variable in angiosperms and the trnS(GGA)-rpS4 region has been widely used in monilophyte phylogenetic studies. Phylogenetic resolution was equivalent across regions, but the strength of support for the phylogenies varied among regions. Of the five sampled regions the trnS(GCU)-trnG(UUC) spacer+trnG(UUC) intron region provided the strongest support for the inferred phylogeny.

DNA, Chloroplast↗