Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Sequence and gene organization of mouse mitochondrial DNA.

The complete sequence of the 16,295 bp mouse L cell mitochondrial DNA genome has been determined. Genes for the 12S and 16S ribosomal RNAs; 22 tRNAs; cytochrome c oxidase subunits I, II and III; ATPase subunit 6; cytochrome b; and eight unidentified proteins have been located. The genome displays exceptional economy of organization, with tRNA genes interspersed between rRNA and protein-coding genes with zero or few noncoding nucleotides between coding sequences. Only two significant portions of the genome, the 879 nucleotide displacement-loop region containing the origin of heavy-strand replication and the 32 nucleotide origin of light-strand replication, do not encode a functional RNA species. All of the remaining nucleotide sequence serves as a defined coding function, with the exception of 32 nucleotides, of which 18 occur at the 5' ends of open reading frames. Mouse mitochondrial DNA is unique in that the translational start codon is AUN, with any of the four nucleotides in the third position, whereas the only translational stop codon is the orthodox UAA. The mouse mitochondrial DNA genome is highly homologous in overall sequence and in gene organization to human mitochondrial DNA, with the descending order of conserved regions being tRNA genes; origin of light-strand replication; rRNA genes; known protein-coding genes; unidentified protein-coding genes; displacement-loop region.

Animals↗

Spontaneous and engineered deletions in the 3' noncoding region of tick-borne encephalitis virus: construction of highly attenuated mutants of a flavivirus.

The flavivirus genome is a positive-strand RNA molecule containing a single long open reading frame flanked by noncoding regions (NCR) that mediate crucial processes of the viral life cycle. The 3' NCR of tick-borne encephalitis (TBE) virus can be divided into a variable region that is highly heterogeneous in length among strains of TBE virus and in certain cases includes an internal poly(A) tract and a 3'-terminal conserved core element that is believed to fold as a whole into a well-defined secondary structure. We have now investigated the genetic stability of the TBE virus 3' NCR and its influence on viral growth properties and virulence. We observed spontaneous deletions in the variable region during growth of TBE virus in cell culture and in mice. These deletions varied in size and location but always included the internal poly(A) element of the TBE virus 3' NCR and never extended into the conserved 3'-terminal core element. Subsequently, we constructed specific deletion mutants by using infectious cDNA clones with the entire variable region and increasing segments of the core element removed. A virus mutant lacking the entire variable region was indistinguishable from wild-type virus with respect to cell culture growth properties and virulence in the mouse model. In contrast, even small extensions of the deletion into the core element led to significant biological effects. Deletions extending to nucleotides 10826, 10847, and 10870 caused distinct attenuation in mice without measurable reduction of cell culture growth properties, which, however, were significantly restricted when the deletion was extended to nucleotide 10919. An even larger deletion (to nucleotide 10994) abolished viral viability. In spite of their high degree of attenuation, these mutants efficiently induced protective immune responses even at low inoculation doses. Thus, 3'-NCR deletions represent a useful technique for achieving stable attenuation of flaviviruses that can be included in the rational design of novel flavivirus live vaccines.

Animals↗

Unraveling transcription regulatory networks by protein-DNA and protein-protein interaction mapping.

Metazoan genomes contain thousands of protein-coding and noncoding RNA genes, most of which are differentially expressed, i.e., at different locations or at different times during development, function, or pathology of the organism. Differential gene expression is achieved in part by the action of regulatory transcription factors (TFs) that bind to cis-regulatory elements that are often located in or near their target genes. Each TF likely regulates many targets in the context of intricate transcription regulatory networks. Up to 10% of a genome may encode TFs, but only a handful of these have been studied in detail. Here, I will discuss the different steps involved in the mapping and analysis of transcription regulatory networks, including the identification of network nodes (TFs and their target sequences) and edges (TF-TF dimers and TF-DNA target interactions), integration with other data types, and network properties and emerging principles that provide insights into differential gene expression.

Animals↗

Comparative sequence analysis of four complete primary structures of plum pox virus strains.

The complete nucleotide sequence of plum pox virus (PPV) strain SK 68 was determined from a series of overlapping cDNA clones. The exact 5' terminus was determined by direct RNA sequencing. The RNA sequence was 9786 nucleotides in length, excluding a 3' terminal poly(A) sequence. The large open reading frame starts at nucleotide position 147 and is terminated at position 9568. Comparison of cistrons from other plum pox virus strains with those predicted for the SK 68 strain indicated the same genomic organizations. Comparison of sequences leads to the following conclusions: (1) The genetic organization of all four PPV strains is identical, containing one large polyprotein gene and two noncoding regions at the 5' and 3' ends; (2) pairwise comparison of the genomic sequence of PPV SK 68 with other PPV strains shows 11% alteration. Sequence differences among strains are spread in a uniform manner upon the genome, except for the P1, HC-pro, and two noncoding regions, which are more conserved (with a 4% and 6.6% change). The stability of the noncoding regions is probably linked to their role in replication. The sequence variation has little effect on the amino acid sequence of the corresponding polypeptides, as changes occur preferentially in the third position of the reading frame triplets, except in the case of the 5' end of the coat protein gene (2.7% average difference in amino acid level, while in the case of coat protein it is 7.7%).(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Mutation pressure and the evolution of organelle genomic architecture.

The nuclear genomes of multicellular animals and plants contain large amounts of noncoding DNA, the disadvantages of which can be too weak to be effectively countered by selection in lineages with reduced effective population sizes. In contrast, the organelle genomes of these two lineages evolved to opposite ends of the spectrum of genomic complexity, despite similar effective population sizes. This pattern and other puzzling aspects of organelle evolution appear to be consequences of differences in organelle mutation rates. These observations provide support for the hypothesis that the fundamental features of genome evolution are largely defined by the relative power of two nonadaptive forces: random genetic drift and mutation pressure.

Animals↗

Noncoding DNA, isochores and gene expression: nucleosome formation potential.

The nucleosome formation potential of introns, intergenic spacers and exons of human genes is shown here to negatively correlate with among-tissues breadth of gene expression. The nucleosome formation potential is also found to negatively correlate with the GC content of genomic sequences; the slope of regression line is steeper in exons compared with noncoding DNA (introns and intergenic spacers). The correlation with GC content is independent of sequence length; in turn, the nucleosome formation potential of introns and intergenic spacers positively (albeit weakly) correlates with sequence length independently of GC content. These findings help explain the functional significance of the isochores (regions differing in GC content) in the human genome as a result of optimization of genomic structure for epigenetic complexity and support the notion that noncoding DNA is important for orderly chromatin condensation and chromatin-mediated suppression of tissue-specific genes.

DNA, Intergenic↗

Evolution of programmed DNA rearrangements in a scrambled gene.

Gene unscrambling in spirotrichous ciliates involves massive genome-wide DNA deletion and rearrangement events during development. During each sexual cycle, the somatic nucleus (macronucleus) regenerates from the germ line nucleus (micronucleus). Development of the polyploid somatic genome requires programmed DNA deletion of micronuclear-limited intragenic noncoding sequences and permutation and amplification of the protein-coding regions. Recent studies suggest that, despite novel insertions of endogenous transposon or foreign DNA into the germ line genome, ciliates possess a whole-genome surveillance system that guides the recapitulation of a functional somatic genome. This renders the germ line genome an extremely dynamic structure over evolutionary time. Here we describe the germ line and somatic architectures of the gene encoding alpha-telomere-binding protein in three early-diverging species (Holosticha sp., Uroleptus sp., and Paraurostyla weissei) and trace the natural history of DNA rearrangements in this gene in six species, including three previously studied oxytrichids. Comparisons of homologous coding regions between earlier and later diverging species provide evidence for fusion of scrambled germ line fragments as small as 24 bp during evolution, as well as simultaneous fragmentation and scrambling of the germ line locus and shifting of the boundaries between coding and noncoding DNA, leading to distinct gene architectures in each species. We infer an evolutionary recombination pathway that passes through identified intermediate species and gives rise to the observed patterns in all known species, capitalizing on their unique DNA rearrangement machinery and germ line flexibility.

Animals↗

Characterization of the cod (Gadus morhua) steroidogenic acute regulatory protein (StAR) sheds light on StAR gene structure in fish.

The full-length cDNA for the cod (Gadus morhua) StAR was cloned by RT-PCR and library screening using ovarian RNA. From the library screening, 2 size classes of cDNA were obtained; a 1577 bp cDNA (cStAR1) and a 2851 bp cDNA (cStAR2). The cStAR1 cDNA presumably encodes a protein of 286 amino acids. The cStAR2 cDNA was composed of 6 separated sequences that contained all of the coding regions of cStAR1 when added together, but also contained 5 noncoding regions not observed in cStAR1. Polymerase chain reactions of cod genomic DNA produced products slightly larger than cStAR2. The sequence of these products were the same as cStAR2 but revealed one additional noncoding region (intron). Thus, the fish StAR gene contains the same number of exons (7) and introns (6) as observed in mammals, but is approximately half the size of the mammalian gene. Using Northern analysis and RT-PCR, cStAR1 expression was observed only in testes, ovaries and head kidneys. Polymerase chain reaction products were also observed using cDNA from steroidogenic tissues and primers designed to regions specific for cStAR2, indicating that cStAR2 is expressed in tissues and may account for the presence of larger transcripts observed on Northern blots.

Animals↗

Analysis of a 103 kbp cluster homology region from the left end of Saccharomyces cerevisiae chromosome I.

The DNA sequence and preliminary functional analysis of a 103-kbp section of the left arm of yeast chromosome I is presented. This region, from the left telomere to the LTE1 gene, can be divided into two distinct portions. One portion, the telomeric 29 kbp, has a very low gene density (only five potential genes and 21 kbp of noncoding sequence), does not encode any "functionally important" genes, and is rich in sequences repeated several times within the yeast genome. The other portion, with 37 genes and only 14.5 kbp of noncoding sequence, is gene rich and codes for at least 16 "functionally important" genes. The entire gene-rich portion is apparently duplicated on chromosome XV as an extensive region of partial gene synteney called a cluster homology region. A function can be assigned with varying degrees of precision to 23 of the 42 potential genes in this region; however, the precise function is know for only eight genes. Nineteen genes encode products presently novel to yeast, although five of these have homologs elsewhere in the yeast genome.

Base Sequence↗

Cloning of the aldehyde reductase gene from a red yeast, Sporobolomyces salmonicolor, and characterization of the gene and its product.

An NADPH-dependent aldehyde reductase (ALR) isolated from a red yeast, Sporobolomyces salmonicolor, catalyzes the reduction of a variety of carbonyl compounds. To investigate its primary structure, we cloned and sequenced the cDNA coding for ALR. The aldehyde reductase gene (ALR) comprises 969 bp and encodes a polypeptide of 35,232 Da. The deduced amino acid sequence showed a high degree of similarity to other members of the aldo-keto reductase superfamily. Analysis of the genomic DNA sequence indicated that the ALR gene was interrupted by six introns (two in the 5' noncoding region and four in the coding region). Southern hybridization analysis of the genomic DNA from S. salmonicolor indicated that there was one copy of the gene. The ALR gene was expressed in Escherichia coli under the control of the tac promoter. The enzyme expressed in E. coli was purified to homogeneity and showed the same catalytic properties as did the enzyme from S. salmonicolor.

Alcohol Oxidoreductases↗

The mitochondrial genome of the wine yeast Hanseniaspora uvarum: a unique genome organization among yeast/fungal counterparts.

The complete sequence of the apiculate wine yeast Hanseniaspora uvarum mtDNA has been determined and analysed. It is an extremely compact linear molecule containing the shortest functional region ever found in fungi (11 094 bp long), flanked by Type 2 telomeric inverted repeats. The latter contained a 2704-bp-long subterminal region and tandem repeats of 839-bp units. In consequence, a population of mtDNA molecules that differed at the number of their telomeric reiterations was detected. The functional region of the mitochondrial genome coded for 32 genes, which included seven subunits of respiratory complexes and ATP synthase (the genes encoding for NADH oxidoreductase subunits were absent), two rRNAs and 23 tRNA genes which recognized codons for all amino acids. A single intron interrupted the cytochrome oxidase subunit 1 gene. A number of reasons contributed towards its strikingly small size, namely: (1) the remarkable size reduction (by >40%) of the rns and rnl genes; (2) that most tRNA genes and five of the seven protein-coding genes were the shortest among known yeast homologs; and (3) that the noncoding regions were restricted to 5.1% of the genome. In addition, the genome showed multiple changes in the orientation of transcription and the gene order differed drastically from other yeasts. When all protein coding gene sequences were considered as one unit and were compared with the corresponding molecules from all other complete mtDNAs of yeasts, the phylogenetic trees constructed robustly supported its placement basal to the yeast species of the 'Saccharomyces complex', demonstrating the advantage of this approach over single-gene or multigene approaches of unlinked genes.

Base Sequence↗

Comparisons of the genomic cis-elements and coding regions in RNA beta components of the hordeiviruses barley stripe mosaic virus, lychnis ringspot virus, and poa semilatent virus.

Nucleotide sequences of the genomic RNA beta components of hordeiviruses poa semilatent virus (PSLV) and lychnis ringspot virus (LRSV) were determined. PSLV and LRSV closely resemble barley stripe mosaic virus (BSMV), type hordeivirus, in the gene arrangement of their RNAs beta, comprising 5'-proximal beta a (coat protein) gene and downstream triple gene block (TGB) coding for the beta b, beta c, and beta d putative transport proteins. The beta a, beta b, beta c, and beta d proteins of the three hordeiviruses showed significant sequence similarity, with the respective proteins of PSLV and BSMV being closer to each other than to their counterparts of LSRV. Comparisons of the TGB-encoded proteins of hordeiviruses, potexviruses, carlaviruses, and furoviruses indicate that the first and second TGB genes belong to the monophyletic groups, whereas the third gene may have multiple ancestry. LRSV, PSLV, and BSMV showed remarkable variation in the 3'-untranslated regions of their genomic RNAs. Among the three hordeiviruses, LRSV has the shortest 3'-noncoding region that lacks tentative pseudoknot-forming elements conserved upstream of the 3'-tRNA-like structure in the BSMV and PSLV genomes. On the other hand, LRSV RNA beta, like that of BSMV, contained the internal poly(A) sequence that is absent from PSLV RNA.

Adaptation, Physiological↗

Mapping of mutations associated with neurovirulence in monkeys infected with Sabin 1 poliovirus revertants selected at high temperature.

Poliovirus type 1 neurovirulence is difficult to analyze because of the 56 mutations which differentiate the neurovirulent Mahoney strain from the attenuated Sabin strain. We have isolated four neurovirulent mutants which differ from the temperature-sensitive parental Sabin 1 strain by only a few mutations, using selection for temperature resistance: mutant S(1)37C1 was isolated at 37.5 degrees C, S(1)38C5 was isolated at 38.5 degrees C, and S(1)39C6 and S(1)39C10 were isolated at 39.5 degrees C. All four mutants had a positive reproductive capacity at supraoptimal temperature (Rct+ phenotype). Mutant S(1)37C1 induced paralysis in two of four cynomolgus monkeys, and the three other mutants induced paralysis in four of four monkeys. The lesion score increased from the S(1)37C1 mutant to the S(1)39 mutants. To map the mutations associated with thermoresistance and neurovirulence, we sequenced all regions in which the Sabin 1 genome differs from the Mahoney genome. The S(1)37C1 mutant had one mutation in the 5' noncoding region and another in the 3' noncoding region. Mutant S(1)38C5 had these mutations plus another mutation in the 3D polymerase gene. The S(1)39 mutants had three additional mutations in the capsid protein region. The mutations were located at positions at which the Sabin 1 and Mahoney genomes differ, except for the mutation in the 5' noncoding region. The noncoding-region mutations apparently confer a low degree of neurovirulence. The 3D polymerase mutation, which distinguishes S(1)38C5 and S(1)39 mutants from S(1)37C1, is probably responsible for the high neurovirulence of S(1)38C5 and S(1)39 mutants. The capsid region mutations may contribute to the neurovirulence of the S(1)39 mutants, which was the highest among the mutants.

Animals↗

Evolution of DNA organization in hypotrichous ciliates.

Profound changes have been introduced into the germline (micronuclear) and somatic (macronuclear) genomes of hypotrichous ciliates during evolution. First, multiple, short, unique, noncoding sequences, called IESs, have been inserted into micronuclear genes. IESs are spliced out of each gene, and the gene segments, called MDSs, are ligated during conversion of the micronuclear genome to a macronuclear genome after cell mating. The IESs in a particular gene can change dramatically in number, position, length, and sequence during speciation. Once inserted, IESs can shift along the DNA of a gene 1 or 2 bp at a time by a mutational mechanism that does not alter the coding sequence of the gene. Second, the MDSs in the same genes have been rearranged into random or nonrandom, scrambled disorder. The origin of nonrandom scrambling patterns can be explained by a model of simultaneous insertion of multiple IESs into a germline gene during evolution. Subsequent recombinations among the IESs may change the MDS arrangement from a nonrandomly scrambled to a randomly scrambled pattern, including inversions of MDSs. Third, during the conversion of a micronucleus to a macronucleus after cell mating, MDSs are ligated in the unscrambled, orthodox order in association with IES excision, and the genes are then removed from the chromosomes as individual, short DNA molecules. These gene-size molecules are amplified many-fold to produce a mature macronucleus. All of these phenomena attest to a remarkable fluidity of the hypotrich genome both over evolutionary time and in the conversion of a germline genome to a somatic genome. The significance of this fluidity for the life and evolution of these organisms is still obscure. Recombination among IESs could shuffle MDSs and facilitate faster evolution of new genes.

Animals↗

Comprehensive genome sequence analysis of a breast cancer amplicon.

Gene amplification occurs in most solid tumors and is associated with poor prognosis. Amplification of 20q13.2 is common to several tumor types including breast cancer. The 1 Mb of sequence spanning the 20q13.2 breast cancer amplicon is one of the most exhaustively studied segments of the human genome. These studies have included amplicon mapping by comparative genomic hybridization (CGH), fluorescent in-situ hybridization (FISH), array-CGH, quantitative microsatellite analysis (QUMA), and functional genomic studies. Together these studies revealed a complex amplicon structure suggesting the presence of at least two driver genes in some tumors. One of these, ZNF217, is capable of immortalizing human mammary epithelial cells (HMEC) when overexpressed. In addition, we now report the sequencing of this region in human and mouse, and on quantitative expression studies in tumors. Amplicon localization now is straightforward and the availability of human and mouse genomic sequence facilitates their functional analysis. However, comprehensive annotation of megabase-scale regions requires integration of vast amounts of information. We present a system for integrative analysis and demonstrate its utility on 1.2 Mb of sequence spanning the 20q13.2 breast cancer amplicon and 865 kb of syntenic murine sequence. We integrate tumor genome copy number measurements with exhaustive genome landscape mapping, showing that amplicon boundaries are associated with maxima in repetitive element density and a region of evolutionary instability. This integration of comprehensive sequence annotation, quantitative expression analysis, and tumor amplicon boundaries provide evidence for an additional driver gene prefoldin 4 (PFDN4), coregulated genes, conserved noncoding regions, and associate repetitive elements with regions of genomic instability at this locus.

Animals↗

Noncoding RNA genes.

Some genes produce RNAs that are functional instead of encoding proteins. Noncoding RNA genes are surprisingly numerous. Recently, active research areas include small nucleolar RNAs, antisense riboregulator RNAs, and RNAs involved in X-dosage compensation. Genome sequences and new algorithms have begun to make systematic computational screens for noncoding RNA genes possible.

Animals↗

A subtype of human papillomavirus 5 (HPV-5b) and its subgenomic segment amplified in a carcinoma: nucleotide sequences and genomic organizations.

A subtype of human papillomavirus 5 (HPV-5b) is closely associated with carcinomas in the disease epidermodysplasia verruciformis (EV). The complete genome was cloned from virus particles in benign lesions of a patient with EV and sequenced: it was 7779 nucleotides long and consisted of six open reading frames (ORFs) (E6, E7, E1, E2, E4, and E5) in the early region, three ORFs (L2, L3, and L1) in the late region, and a noncoding region, all existing on one DNA strand. The 40% segment of the HPV-5b genome specifically amplified in carcinomas was cloned from a primary carcinoma of the same EV patient and sequenced: it was 3143 nucleotides long and corresponded to a segment of the original HPV-5b genome containing the entire sequences of E6, E7, and the noncoding region and portions of E1 and L1. Compared to the whole genomic DNA, no mutations were detected in this probable malignancy-associated viral subgenomic segment cloned from carcinoma. These results suggest that amplification of the viral segment containing E6, E7, and the noncoding region may play a role in the malignant conversion of HPV-5b-infected benign lesions and that mutations in these genes or regions are not necessarily required.

Base Sequence↗

Comparative analysis of noncoding regions of 77 orthologous mouse and human gene pairs.

A data set of 77 genomic mouse/human gene pairs has been compiled from the EMBL nucleotide database, and their corresponding features determined. This set was used to analyze the degree of conservation of noncoding sequences between mouse and human. A new alignment algorithm was developed to cope with the fact that large parts of noncoding sequences are not alignable in a meaningful way because of genetic drift. This new algorithm, DNA Block Aligner (DBA), finds colinear-conserved blocks that are flanked by nonconserved sequences of varying lengths. The noncoding regions of the data set were aligned with DBA. The proportion of the noncoding regions covered by blocks >60% identical was 36% for upstream regions, 50% for 5' UTRs, 23% for introns, and 56% for 3' UTRs. These blocks of high identity were more or less evenly distributed across the length of the features, except for upstream regions in which the first 100 bp upstream of the transcription start site was covered in up to 70% of the gene pairs. This data set complements earlier sets on the basis of cDNA sequences and will be useful for further comparative studies. [This paper contains supplementary data that can be found at http://www.genome.org [corrected]].

3' Untranslated Regions↗