Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Intron-genome size relationship on a large evolutionary scale.

The intron-genome size relationship was studied across a wide evolutionary range (from slime mold and yeast to human and maize), as well as the relationship between genome size and the ratio of intervening/coding sequence size. The average intron size is scaled to genome size with a slope of about one-fourth for the log-transformed values; i.e., on the global scale its increase in evolution is lower than the increase in genome size by four orders of magnitude. There are exceptions to the general trend. In baker's yeast introns are extraordinarily long for its genome size. Tetrapods also have longer introns than expected for their genome sizes. In teleost fish the mean intron size does not differ significantly, notwithstanding the differences in genome size. In contrast to previous reports, avian introns were not found to be significantly shorter than introns of mammals, although avian genomes are smaller than genomes of mammals on average by about a factor of 2.5. The extra-/intragenic ratio of noncoding DNA can be higher in fungi than in animals, notwithstanding the smaller fungal genomes. In vertebrates and invertebrates taken separately, this ratio is increasing as the increase in genome size. Two hypotheses are proposed to explain the variation in the extra-/intragenic ratio of noncoding DNA in organisms with similar numbers of genes: transition (dynamic) and equilibrium (static). According to the transition model, this variation arises with the rapid shift of genome size because the bulk of extragenic DNA can be changed more rapidly than the finely interspersed intron sequences. The equilibrium model assumes that this variation is a result of selective adjustment of genome size with constraints imposed on the intron size due to its putative link to chromatin structure (and constraints of the splicing machinery).

Animals↗

Second-site suppressor mutations assist in studying the function of the 3' noncoding region of turnip yellow mosaic virus RNA.

The 3' noncoding region of turnip yellow mosaic virus RNA includes an 82-nucleotide-long tRNA-like structure domain and a short upstream region that includes a potential pseudoknot overlapping the coat protein termination codon. Genomic RNAs with point mutations in the 3' noncoding region that result in poor replication in protoplasts and no systemic symptoms in planta were inoculated onto Chinese cabbage plants in an effort to obtain second-site suppressor mutations. Putative second-site suppressor mutations were identified by RNase protection and sequencing and were then introduced into genomic cDNA clones to permit their characterization. A C-57----U mutation in the tRNA-like structure was a strong suppressor of the C-55----A mutation which prevented both systemic infection and in vitro valylation of the viral RNA. Both of these phenotypes were rescued in the double mutant. An A-107----C mutation was a strong second-site suppressor of the U-96----G mutation, permitting the double mutant to establish systemic infection. The C-107 and G-96 mutations are located on opposite strands of one helix of a potential pseudoknot, and the results support a functional role for the pseudoknot structure. A mutation near the 5' end of the genome (G + 92----A), at position -3 relative to the initiation codon of the essential open reading frame 206, was found to be a general potentiator of viral replication, probably as a result of enhanced expression of open reading frame 206. The A + 92 mutation enhanced the replication of mutant TYMC-G96 in protoplasts but was not a sufficiently potent suppressor to permit systemic spread of the A + 92/G-96 double mutant in plants.

Anticodon↗

Putting More Genetics into Genetic Algorithms.

The majority of current genetic algorithms (GAs), while inspired by natural evolutionary systems, are seldom viewed as biologically plausible models. This is not a criticism of GAs, but rather a reflection of choices made regarding the level of abstraction at which biological mechanisms are modeled, and a reflection of the more engineering-oriented goals of the evolutionary computation community. Understanding better and reducing this gap between GAs and genetics has been a central issue in an interdisciplinary project whose goal is to build GA-based computational models of viral evolution. The result is a system called Virtual Virus (VIV). The VIV incorporates a number of more biologically plausible mechanisms, including a more flexible genotype-to-phenotype mapping. In VIV the genes are independent of position, and genomes can vary in length and may contain noncoding regions, as well as duplicative or competing genes. Initial computational studies with VIV have already revealed several emergent phenomena of both biological and computational interest. In the absence of any penalty based on genome length, VIV develops individuals with long genomes and also performs more poorly (from a problem-solving viewpoint) than when a length penalty is used. With a fixed linear length penalty, genome length tends to increase dramatically in the early phases of evolution and then decrease to a level based on the mutation rate. The plateau genome length (i.e., the average length of individuals in the final population) generally increases in response to an increase in the base mutation rate. When VIV converges, there tend to be many copies of good alternative genes within the individuals. We observed many instances of switching between active and inactive genes during the entire evolutionary process. These observations support the conclusion that noncoding regions serve a positive step in understanding how GAs might exploit more of the power and flexibility of biological evolution while simultaneously providing better tools for understanding evolving biological systems.

Journal Article↗

Molecular characterization of the 11th RNA segment from human group C rotavirus.

The complete nucleotide sequence of genome segment 11 from the noncultivatable, human group C rotavirus (Bristol strain) was determined. Comparison of the nucleotide sequence of the segment termini with the consensus 5' and 3' terminal noncoding sequences of the human group C rotavirus genome revealed characteristic 5' and 3' sequences. Human group C rotavirus genome segment 11 is 613 bp long and encodes a single open reading frame of 450 nucleotides (150 amino acids) starting at nucleotide 39 and terminating at nucleotide 489, leaving a long 3' untranslated region of 124 nucleotides. The predicted translation product has a calculated molecular weight of 17.7 kD and contains four potential N-linked glycosylation sites. No significant homologies to other viral proteins were found in database searches. Hydropathy analysis predicted the human group C rotavirus genome segment 11 translation product has a hydrophilic carboxy terminus (amino acids 54-150) and a hydrophobic amino terminus (amino acids 1-53) that can be further subdivided into three short hydrophobic sequences--H1, H2, and H3. These features are analogous to the integral membrane glycoprotein NSP4 encoded by group A rotavirus gene 10.

Amino Acid Sequence↗

Complete mtDNA sequences of two millipedes suggest a new model for mitochondrial gene rearrangements: duplication and nonrandom loss.

We determined the complete mitochondrial DNA (mtDNA) sequences of the millipedes Narceus annularus and Thyropygus sp. (Arthropoda: Diplopoda) and identified, in both genomes, all 37 genes typical for metazoan mtDNA. The arrangement of these genes is identical in the two millipedes, but differs from others found in arthropod mtDNAs in the location of at least four genes or gene blocks. This novel gene arrangement is unusual for animal mtDNA in that genes with identical transcriptional polarities are clustered in the genome, and the two clusters are separated by two noncoding regions. The only exception to this pattern is the gene for cysteine tRNA, which is located in the part of the genome that otherwise contains all genes with the opposite transcriptional polarity. We suggest that a mechanism involving complete mtDNA duplication followed by the loss of genes, predetermined by their transcriptional polarity and location in the genome, could generate this gene arrangement from the one ancestral for arthropods. The proposed mechanism has important implications for phylogenetic inferences that are drawn on the basis of gene arrangement comparisons.

Animals↗

Are noncoding sequences of Rickettsia prowazekii remnants of "neutralized" genes?

It has been hypothesized that a large fraction of 24% noncoding DNA in R. prowazekii consists of degraded genes. This hypothesis has been based on the relatively high G+C content of noncoding DNA. However, a comparison with other genomes also having a low overall G+C content shows that this argument would also apply to other bacteria. To test this hypothesis, we study the coding potential in sets of genes, pseudogenes, and intergenic regions. We find that the correlation function and the chi(2)-measure are clearly indicative of the coding function of genes and pseudogenes. However, both coding potentials make almost no indication of a preexisting reading frame in the remaining 23% of noncoding DNA. We simulate the degradation of genes due to single-nucleotide substitutions and insertions/deletions and quantify the number of mutations required to remove indications of the reading frame. We discuss a reduced selection pressure as another possible origin of this comparatively large fraction of noncoding sequences.

DNA, Intergenic↗

A novel gene family NBPF: intricate structure generated by gene duplications during primate evolution.

Partial and complete genome duplications occurred during evolution and resulted in the creation of new genes and gene families. We identified a novel and intricate human gene family located primarily in regions of segmental duplications on human chromosome 1. We named it NBPF, for neuroblastoma breakpoint family, because one of its members is disrupted by a chromosomal translocation in a neuroblastoma patient. The NBPF genes have a repetitive structure with high intragenic and intergenic sequence similarity in both coding and noncoding regions. These similarities might expose these genomic regions to illegitimate recombination, resulting in structural variation in the NBPF genes. The encoded proteins contain a highly conserved domain of unknown function, which we have named the NBPF repeat. In silico analysis combined with the isolation of multiple full-length cDNA clones showed that several members of this gene family are abundantly expressed in a large variety of tissues and cell lines. Strikingly, no discernable orthologues could be identified in the completed genomes of fruit fly, nematode, mouse, or rat, but sequences with low homology could be isolated from the draft canine and bovine genomes. Interestingly, this gene family shows primate-specific duplications that result in species-specific arrays of NBPF homologous sequences. Overall, this novel NBPF family reflects the continuous evolution of primate genomes that resulted in large physiological differences, and its potential role in this process is discussed.

Amino Acid Sequence↗

Evidence for recombination of mtDNA in the marine mussel Mytilus trossulus from the Baltic.

A number of studies have claimed that recombination occurs in animal mtDNA, although this evidence is controversial. Ladoukakis and Zouros (2001) provided strong evidence for mtDNA recombination in the COIII gene in gonadal tissue in the marine mussel Mytilus galloprovincialis from the Black Sea. The recombinant molecules they reported had not however become established in the population from which experimental animals were sampled. In the present study, we provide further evidence of the generality of mtDNA recombination in Mytilus by reporting recombinant mtDNA molecules in a related mussel species, Mytilus trossulus, from the Baltic. The mtDNA region studied begins in the 16S rRNA gene and terminates in the cytochrome b gene and includes a major noncoding region that may be analogous to the D-loop region observed in other animals. Many bivalve species, including some Mytilus species, are unusual in that they have two mtDNA genomes, one of which is inherited maternally (F genome) the other inherited paternally (M genome). Two recombinant variants reported in the present study have population frequencies of 5% and 36% and appear to be mosaic for F-like and M-like sequences. However, both variants have the noncoding region from the M genome, and both are transmitted to sperm like the M genome. We speculate that acquisition of the noncoding region by the recombinant molecules has conferred a paternal role on mtDNA genomes that otherwise resemble the F genome in sequence.

Animals↗

The noncoding RNA taurine upregulated gene 1 is required for differentiation of the murine retina.

BACKGROUND: With the advent of genome-wide analyses, it is becoming evident that a large number of noncoding RNAs (ncRNAs) are expressed in vertebrates. However, of the thousands of ncRNAs identified, the functions of relatively few have been established. RESULTS: In a screen for genes upregulated by taurine in developing retinal cells, we identified a gene that appears to be a ncRNA. Taurine Upregulated Gene 1 (TUG1) is a spliced, polyadenylated RNA that does not encode any open reading frame greater than 82 amino acids in its full-length, 6.7 kilobase (kb) RNA sequence. Analyses of Northern blots and in situ hybridization revealed that TUG1 is expressed in the developing retina and brain, as well as in adult tissues. In the newborn retina, knockdown of TUG1 with RNA interference (RNAi) resulted in malformed or nonexistent outer segments of transfected photoreceptors. Immunofluorescent staining and microarray analyses suggested that this loss of proper photoreceptor differentiation is a result of the disregulation of photoreceptor gene expression. CONCLUSIONS: A function for a newly identified ncRNA, TUG1, has been established. TUG1 is necessary for the proper formation of photoreceptors in the developing rodent retina.

Amino Acid Sequence↗

Active conservation of noncoding sequences revealed by three-way species comparisons.

Human and mouse genomic sequence comparisons are being increasingly used to search for evolutionarily conserved gene regulatory elements. Large-scale human-mouse DNA comparison studies have discovered numerous conserved noncoding sequences of which only a fraction has been functionally investigated A question therefore remains as to whether most of these noncoding sequences are conserved because of functional constraints or are the result of a lack of divergence time.

Animals↗

Genome exposure and regulation in mammalian cells.

A method of measurement of exposed DNA (i.e. hypersensitive to DNase I hydrolysis) as opposed to sequestered (hydrolysis resistant) DNA in isolated nuclei of mammalian cells is described. While cell cultures exhibit some differences in behavior from day to day, the general pattern of exposed and sequestered DNA is satisfactorily reproducible and agrees with results previously obtained by other methods. The general pattern of DNA hydrolysis exhibited by all cells tested consists of a curve which at first rises sharply with increasing DNase I, and then becomes almost horizontal, indicating that roughly about half of the nuclear DNA is highly sequestered. In 4 cases where transformed cells (Raszip6, CHO, HL60 and PC12) were compared, each with its more normal homolog (3T3, and the reverse transformed versions of CHO, HL60 and PC12, achieved by dibutyryl cyclic AMP [DBcAMP], retinoic acid, and nerve growth factor [NGF] respectively), the transformed form displayed less genome exposure than the nontransformed form at every DNase I dose tested. When Ca++ was excluded from the hydrolysis medium in both the Raszip6-3T3 and the CHO-DBcAMP systems, the normal cell forms lost their increased exposure reverting to that of the transformed forms. Therefore Ca++ appears necessary for maintenance of the DNA in the more highly exposed state characteristic of the nontransformed phenotype. LiCl increases the DNA exposure of all transformed cells tested. Dextran sulfate and heparin each can increase the DNA exposure of several different cancers. Colcemid prevents the increase of exposure of CHO by DBcAMP but it must be administered before or simultaneously with the latter compound. Measurements on mouse biopsies reveal large differences in exposure in different normal tissues. Thus, the exposure from adult liver cells was greater than that of adult brain, but both fetal liver and fetal brain had significantly greater exposure than their adult counterparts. Exposure in normal human fibroblasts as revealed by in situ nick translation reveals a nuclear distribution pattern around the periphery, around the nucleoli and in punctate positions in the nuclear interior in parts of both S and G1 phases of the cell cycle. The same exposure pattern is duplicated by the pattern of DNA synthesis in S cells. It would appear that these nuclear regions represent positions of special activity. The previously proposed theory of genome regulation in mammalian cells is supported by these findings. The theory proposes that: a) gene activity requires exposure of the given locus followed by action of transcription factors on the exposed genes; b) the fiber system of the cell (cytoskeleton, nuclear fibers, and extracellular fibers) are required for normal exposure; c) active sites for gene expression and replication consist of the nuclear periphery where differentiation genes particularly are exposed; the nucleoli where at least some housekeeping genes are exposed; and possibly also punctate regions in the interior; d) noncoding sequences play a critical role in genome regulation, possibly including the transport of loci to be activated to appropriate exposure transcriptional and replicating locations. Cancer cells have lost specific differentiation gene activities, at least sometimes because of mutation of appropriate exposure genes; at least some protooncogenes and tumor suppressor genes are responsible for exposure and transport of specific differentiation gene loci to their appropriate exposure sites in the nucleus and for inducing exposure.

3T3 Cells↗

The evolution of word composition in metazoan promoter sequence.

The field of molecular evolution provides many examples of the principle that molecular differences between species contain information about evolutionary history. One surprising case can be found in the frequency of short words in DNA: more closely related species have more similar word compositions. Interest in this has often focused on its utility in deducing phylogenetic relationships. However, it is also of interest because of the opportunity it provides for studying the evolution of genome function. Word-frequency differences between species change too slowly to be purely the result of random mutational drift. Rather, their slow pattern of change reflects the direct or indirect action of purifying selection and the presence of functional constraints. Many such constraints are likely to exist, and an important challenge is to distinguish them. Here we develop a method to do so by isolating the effects acting at different word sizes. We apply our method to 2-, 4-, and 8-base-pair (bp) words across several classes of noncoding sequence. Our major result is that similarities in 8-bp word frequencies scale with evolutionary time for regions immediately upstream of genes. This association is present although weaker in intronic sequence, but cannot be detected in intergenic sequence using our method. In contrast, 2-bp and 4-bp word frequencies scale with time in all classes of noncoding sequence. These results suggest that different genomic processes are involved at different word sizes. The pattern in 2-bp and 4-bp words may be due to evolutionary changes in processes such as DNA replication and repair, as has been suggested before. The pattern in 8-bp words may reflect evolutionary changes in gene-regulatory machinery, such as changes in the frequencies of transcription-factor binding sites, or in the affinity of transcription factors for particular sequences.

Amino Acids↗

Mitochondrial diversity of early-branching metazoa is revealed by the complete mt genome of a haplosclerid demosponge.

The first mitochondrial (mt) genomes of demosponges have recently been sequenced and appear to be markedly different from published eumetazoan mt genomes. Here we show that the mt genome of the haplosclerid demosponge Amphimedon queenslandica has features that it shares with both demosponges and eumetazoans. Although the A. queenslandica mt genome has typical demosponge features, including size, long noncoding regions, and bacterialike rRNA genes, it lacks atp9, which is found in the other demosponges sequenced to date. We found strong evidence of a recent transposon-mediated transfer of atp9 to the nuclear genome. In addition, A. queenslandica bears an incomplete tRNA set, unusual amino acid deletion patterns, and a putative control region. Furthermore, the arrangement of mt rRNA genes differs from that of other demosponges. These genes evolve at significantly higher rates than observed in other demosponges, similar to previously observed nuclear rRNA gene rates in other haplosclerid demosponges.

Animals↗

Comparative genomic analysis of three strains of Ehrlichia ruminantium reveals an active process of genome size plasticity.

Ehrlichia ruminantium is the causative agent of heartwater, a major tick-borne disease of livestock in Africa that has been introduced in the Caribbean and is threatening to emerge and spread on the American mainland. We sequenced the complete genomes of two strains of E. ruminantium of differing phenotypes, strains Gardel (Erga; 1,499,920 bp), from the island of Guadeloupe, and Welgevonden (Erwe; 1,512,977 bp), originating in South Africa and maintained in Guadeloupe in a different cell environment. Comparative genomic analysis of these two strains was performed with the recently published parent strain of Erwe (Erwo) and other Rickettsiales (Anaplasma, Wolbachia, and Rickettsia spp.). Gene order is highly conserved between the E. ruminantium strains and with A. marginale. In contrast, there is very little conservation of gene order with members of the Rickettsiaceae. However, gene order may be locally conserved, as illustrated by the tuf operons. Eighteen truncated protein-encoding sequences (CDSs) differentiate Erga from Erwe/Erwo, whereas four other truncated CDSs differentiate Erwe from Erwo. Moreover, E. ruminantium displays the lowest coding ratio observed among bacteria due to unusually long intergenic regions. This is related to an active process of genome expansion/contraction targeted at tandem repeats in noncoding regions and based on the addition or removal of ca. 150-bp tandem units. This process seems to be specific to E. ruminantium and is not observed in the other Rickettsiales.

Conserved Sequence↗

A genetic map of Cottus gobio (Pisces, Teleostei) based on microsatellites can be linked to the physical map of Tetraodon nigroviridis.

To initiate QTL studies in the nonmodel fish Cottus gobio we constructed a genetic map based on 171 microsatellite markers. The mapping panel consisted of F1 intercrosses between two divergent Cottus lineages from the River Rhine System. Basic local alignment search tool (BLAST) searches with the flanking sequences of the microsatellite markers yielded a significant (e < 10(-5)) hit with the Tetraodon nigroviridis genomic sequence for 45% of the Cottus loci. Remarkably, most of these hits were due to short highly conserved noncoding stretches. These have an average length of 40 bp and are on average 92% conserved. Comparison of the map locations between the two genomes revealed extensive conserved synteny, suggesting that the Tetraodon genomic sequence will serve as an excellent genomic reference for at least the Acanthopterygii, which include evolutionarily interesting fish groups such as guppies (Poecilia), cichlids (Tilapia) or Xiphophorus (Platy). The apparent high density of short conserved noncoding stretches in these fish genomes will highly facilitate the identification of genes that have been identified in QTL mapping strategies of evolutionary relevant traits.

Animals↗

RNA expression in a cartilaginous fish cell line reveals ancient 3' noncoding regions highly conserved in vertebrates.

We have established a cartilaginous fish cell line [Squalus acanthias embryo cell line (SAE)], a mesenchymal stem cell line derived from the embryo of an elasmobranch, the spiny dogfish shark S. acanthias. Elasmobranchs (sharks and rays) first appeared >400 million years ago, and existing species provide useful models for comparative vertebrate cell biology, physiology, and genomics. Comparative vertebrate genomics among evolutionarily distant organisms can provide sequence conservation information that facilitates identification of critical coding and noncoding regions. Although these genomic analyses are informative, experimental verification of functions of genomic sequences depends heavily on cell culture approaches. Using ESTs defining mRNAs derived from the SAE cell line, we identified lengthy and highly conserved gene-specific nucleotide sequences in the noncoding 3' UTRs of eight genes involved in the regulation of cell growth and proliferation. Conserved noncoding 3' mRNA regions detected by using the shark nucleotide sequences as a starting point were found in a range of other vertebrate orders, including bony fish, birds, amphibians, and mammals. Nucleotide identity of shark and human in these regions was remarkably well conserved. Our results indicate that highly conserved gene sequences dating from the appearance of jawed vertebrates and representing potential cis-regulatory elements can be identified through the use of cartilaginous fish as a baseline. Because the expression of genes in the SAE cell line was prerequisite for their identification, this cartilaginous fish culture system also provides a physiologically valid tool to test functional hypotheses on the role of these ancient conserved sequences in comparative cell biology.

3' Untranslated Regions↗

Genomic restriction endonuclease analysis and mapping of murine guanylate cyclase-A/atrial natriuretic factor receptor gene.

The membrane-bound form of guanylate cyclase represents a biologically active atrial natriuretic factor receptor (GC/ANF-R). We have constructed genomic map of murine GC-A/ANF-R gene using 17 different restriction endonucleases. The restriction mapping results indicated that murine GC-A/ANF-R gene is approximately 20 kb single copy with multiple smaller exons and bigger introns. The Kpn I and Sfu I restriction digests produced 27 kb and 35 kb fragments, respectively, which hybridized with 5'- and 3'-flanking cDNA probes. Both of these fragments should cover the entire murine GC-A/ANF-R genomic sequences. The southern blot hybridization of genomic DNA from human, rat and mouse, using murine 5'-flanking cDNA probe indicated the presence of higher variant sequences in the 5'-flanking region of GC-A/ANF-R gene among different species. The noncoding 5'-flanking probe (350 bp) hybridized only to mouse genomic DNA but not to the human or rat DNA. These sequence variations located in the noncoding 5'-flanking region of GC-A/ANF-R gene may explain the divergent evolutionary development among different species. This is the first demonstration of the restriction endonuclease digestion and genomic mapping of murine GC-A/ANF-R gene which should be valuable to the understanding of its regulation and function.

Animals↗

Genomic organization and mapping of the human HEP-COP gene (COPA) to 1q.

In eukaryotic cells, protein transport between the endoplasmic reticulum and Golgi compartments is mediated in part by non-clathrin-coated vesicular coat proteins (COP). Seven COP subunits have been recognized, and represent components of a complex known as coatomer. We have previously isolated the cDNA of the human homolog of alpha-COP, designated HEP-COP and given the official gene symbol COPA. Here we report the genomic organization of COPA, which contains 33 exons ranging in size from 67 to 611 bp. Mapped by PCR and cycle sequencing, all the exon-intron junctions conformed with the GT-AG rule, the 32 introns ranging from about 80 bp to 4 kbp, with the genomic DNA of COPA estimated to span approximately 37 kb. Southern blot analysis of genomic DNAs of nine eukaryotic species, from human to yeast, revealed identical signals totaling 36 kb each for man and monkey only. Using 5' RACE and primer extension analysis, the putative transcriptional start site was localized to 466 nucleotides upstream of the translation initiation codon. Comprising a 126-nucleotide 5' untranscribed genomic sequence and a 466-nucleotide 5' noncoding cDNA sequence, the 592-nucleotide 5' CpG island lacked TATA and CAAT boxes but displayed a high G+C content, was enriched for CpG dinucleotides, and contained a potential Sp1-binding site, i.e., features compatible with a housekeeping gene. COPA was mapped by fluorescence in situ hybridization to chromosome region 1q23-->q25.

Base Sequence↗