Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

[The paradox of genome size and the problem of redundant DNA].

Data on the relationship between genome size and phenotypic traits of animals and plants are reviewed. Different concepts explaining accumulation of redundant (noncoding) DNA in genomes of eukaryotes are discussed: regulatory (repressory), neutralistic (permissive), skeletal, bodyguard, and buffering, respectively. The relationship between genome size and metabolic rate is presumably the primary one. Possible molecular mechanisms responsible for this relationship are considered. It is concluded that the problem of redundant DNA (genome size paradox) is hardly to solve by studies on molecular level only, since the genome size is a parameter related to molecular, cellular and organismal levels. Redundant DNA seems to be of ecophysiological significance, and therefore the interdisciplinary field dealing with genome size paradox can be called "genome cytoecology".

Animals↗

CSTminer: a web tool for the identification of coding and noncoding conserved sequence tags through cross-species genome comparison.

The identification and characterization of genome tracts that are highly conserved across species during evolution may contribute significantly to the functional annotation of whole-genome sequences. Indeed, such sequences are likely to correspond to known or unknown coding exons or regulatory motifs. Here, we present a web server implementing a previously developed algorithm that, by comparing user-submitted genome sequences, is able to identify statistically significant conserved blocks and assess their coding or noncoding nature through the measure of a coding potential score. The web tool, available at http://www.caspur.it/CSTminer/, is dynamically interconnected with the Ensembl genome resources and produces a graphical output showing a map of detected conserved sequences and annotated gene features.

Animals↗

A protective function for noncoding, or secondary DNA.

The genome of many eukaryotic organisms contains a large amount of noncoding, or secondary DNA. I propose that secondary DNA functions as a sink for the integration of viral and nonviral inserting elements, thereby protecting coding sequences from insertional damage.

DNA↗

Functional L polymerase of La Crosse virus allows in vivo reconstitution of recombinant nucleocapsids.

La Crosse virus (LACV), a member of the family Bunyaviridae, is the primary cause of paediatric encephalitis in the United States. In this study, a functional RNA polymerase (L) gene of LACV was cloned and a reverse genetics system established. A reporter minireplicon mimicking the viral genome was constructed by flanking the Renilla luciferase gene with the 3' and 5' noncoding regions of the genomic M segment. These noncoding regions serve as promoters for the viral polymerase. Both L and nucleocapsid (N) genes were expressed by means of T7 RNA polymerase, which was provided by the recombinant T7-expressing modified vaccinia virus Ankara. Renilla reporter activity in transfected cells reflected reconstitution of recombinant nucleocapsids by functional L and N gene products. Time-course experiments revealed a rapid increase in minireplicon activity from 10 to 18 h after the onset of L and N expression. Minireplicon activity was found to be dependent on the correct ratio of L to N plasmids, with too much of either construct resulting in downregulation. Furthermore, a specific inhibitory effect of LACV NSs protein on minireplicon activity was found. In passaging experiments using parental helper virions, it was demonstrated that the recombinant nucleocapsids are a useful model for transcription, replication and packaging of LACV.

Animals↗

Major human epididymis-specific gene product, HE3, is the first representative of a novel gene family.

Differential screening of a human epididymal cDNA library led to the isolation and characterization of a major epididymis-specific cDNA clone family, referred to as HE3. More detailed sequence and PCR analysis identified two different but homologous gene transcripts, HE3 alpha and HE3 beta. The former represents an mRNA of ca. 1 kb, encoding a putative small secretory polypeptide of 14903 MW. The HE3 beta transcript was only found as incomplete 3' fragments. Analysis of human genomic DNA by Southern blotting suggested the presence in the human genome of at least three independent HE3-related genes. Isolation of genomic clones for the HE3 alpha gene showed this to contain a single intron of 1.4 kb in the 5' noncoding region. Although genomic clones corresponding to HE3 beta could not be found, a third highly homologous gene, HE3 gamma, was identified as a potential pseudogene. Neither nucleotide nor encoded amino acid sequences of the HE3 gene family are related to any other known sequence in the central databases, and thus represents a novel human gene family, with at least three nonallelic members. Northern hybridization analysis showed that HE3 gene products are specifically expressed in the human epididymis, and not in any other tissue examined. Furthermore, except for the pig, no other nonprimate species has been identified to express homologous sequences in the epididymis. RNase protection assays showed that both the HE3 alpha and HE3 beta, but not the HE3 gamma genes, are expressed in the human epididymis.

Amino Acid Sequence↗

Correlation of detectability of hepatitis C virus genome in saliva of elderly Japanese symptomatic HCV carriers with their hepatic function.

The hepatitis C virus (HCV) genome was sought in the saliva of 76 chronic HCV carriers (mean age nearly 60 years) in a rural Japanese town, who had high serum titers of c-100 and anti-core second generation antibodies. In 27 samples (27 cases, 36%), the HCV-RNA genome was detected by the reverse transcriptase - polymerase chain reaction with either of two sets of primers covering two regions of the HCV genome: the 5'noncoding region and the region encompassing the putative envelope (E1). Transaminase values at the time of sampling were higher in the patients with than in those without detectable HCV RNA in saliva (p = 0.04 for alanine aminotransferase, p = 0.04 for aspartate aminotransferase; Wilcoxon test). The prevalence of the positivity was higher by 5'noncoding primers (14/59 vs. 15/68). Our data show that the severity and duration of hepatic dysfunction influence the detectability of the HCV genome in the saliva. This has been a controversial point among investigators.

Adult↗

Determination of the 5' and 3' terminal noncoding sequences of the bi-segmented genome of the avibirnavirus infectious bursal disease virus.

Terminal sequences of the bi-segmented dsRNA genome of 3 different strains of infectious bursal disease virus (IBDV) were analyzed by the rapid amplification of cDNA 5' ends (5'RACE) procedure. Both segments are 85% homologous in a 32-nucleotide sequence comprising the 5' end, whereas the 3' end has a conserved pentamer. Comparison to published terminal sequences of other IBDV strains revealed high conservation between the two segments but more serotype-specific nucleotide changes (5 on segment A and 3 on segment B) in the 5' noncoding region compared to the 3' noncoding region (none on segment A and 1 on segment B).

Animals↗

Mutational analysis of the 5' noncoding region of human immunodeficiency virus type 1 genome.

Retrovirus particles are released by budding from the membranes of infected cells. In the course of virus production, particularly during the late stage, viral genomic RNA is incorporated specifically into virion particles. This specific incorporation of the genomic RNA requires a packaging signal sequence. A region that functions as the packaging signal was mapped to a location upstream of the gag open reading frame on the HIV-1 viral genome. In addition of this packaging signal, other cis-acting elements that are scattered throughout the genome are also required for efficient packaging. The region upstream of the splice donor site is probably important for dimer formation. Therefore, we focused on one region located between the 3' end of the primer binding site and the 5' splice donor site of HIV-1. Experiments were conducted to investigate how deletions or point mutations in this region affect both dimerization in vitro and the production of infectious virus particles. A series of RNAs of varying lengths containing the 5' noncoding region were generated, and genomic dimerization of the altered viral RNA was analyzed in vitro. One RNA construct which consisted of 112 nucleotides (nt) from nt 639 to nt 750 formed a heterodimeric complex with the RNA which consisted of 200 nucleotides from nt 551 to nt 750. We then constructed proviruses with mutations in the 639 to 750 nt region and assayed for virus production. Several mutants that lacked the complementarity necessary to form a possible stem-loop structure in this region showed decreased production of infectious virus particles. Moreover, both deletion of this region and randomization of its nucleotide sequence completely impaired infectious virus production. Thus, the way that this region affects infectious virus production may be through its RNA secondary structure.

Animals↗

Microsatellites: genomic distribution, putative functions and mutational mechanisms: a review.

Microsatellites, or tandem simple sequence repeats (SSR), are abundant across genomes and show high levels of polymorphism. SSR genetic and evolutionary mechanisms remain controversial. Here we attempt to summarize the available data related to SSR distribution in coding and noncoding regions of genomes and SSR functional importance. Numerous lines of evidence demonstrate that SSR genomic distribution is nonrandom. Random expansions or contractions appear to be selected against for at least part of SSR loci, presumably because of their effect on chromatin organization, regulation of gene activity, recombination, DNA replication, cell cycle, mismatch repair system, etc. This review also discusses the role of two putative mutational mechanisms, replication slippage and recombination, and their interaction in SSR variation.

Animals↗

Progress in rickettsial genome analysis from pioneering of Rickettsia prowazekii to the recent Rickettsia typhi.

Three rickettsial genomes have been sequenced and annotated. Rickettsia prowazekii and R. typhi have similar gene order and content. The few differences between R. prowazekii and R. typhi include a 12-kb insertion in R. prowazekii, a large inversion close to the origin of replication in R. typhi, and loss of the complete cytochrome c oxidase system by R. typhi. R. prowazekii, R. typhi, and R. conorii have 13, 24, and 560 unique genes, respectively, and share 775 genes, most likely their essential genes. The small genomes contain many pseudogenes and much noncoding DNA, reflecting the process of genome decay. R. typhi contains the largest number of pseudogenes (41), and R. conorii the fewest, in accordance with its larger number of genes and smaller proportion of noncoding DNA. Conversely, typhus rickettsiae contain fewer repetitive sequences. These genomes portray the key themes of rickettsial intracellular survival: lack of enzymes for sugar metabolism, lipid biosynthesis, nucleotide synthesis, and amino acid metabolism, suggesting that rickettsiae depend on the host for nutrition and building blocks; enzymes for the complete TCA cycle and several copies of ATP/ADP translocase genes, suggesting independent synthesis of ATP and acquisition of host ATP; and type IV secretion system. All rickettsiae share two outer membrane proteins (OmpB and Sca 4) and LPS biosynthesis machinery. RickA, unique to spotted fever rickettsiae, plays a role in induction of actin polymerization in R. conorii, but not in R. prowazekii or R. typhi. The genome of R. typhi contains four potentially membranolytic genes (tlyA, tlyC, pldA, and pat-1) and five autotransporter genes, sca 1, sca 2, sca 3, ompA, and ompB. The presence of six 50-amino acid repeat units in Sca 2 suggests function as an adhesin. The high laboratory passage of the sequenced strains raises the issue of the occurrence of laboratory mutations in genes not required for growth in cell culture or eggs. Resequencing revealed that eight annotated pseudogenes of E strain are actually intact genes. Comparative genomics of virulent and avirulent strains of rickettsial species may reveal their virulence factors.

Genome, Bacterial↗

Organization of the mouse ghrelin gene and promoter: occurrence of a short noncoding first exon.

Ghrelin is a growth hormone-releasing peptide recently discovered in the stomach of rat and human as an endogenous ligand for growth hormone-secretagogue receptor. In the present study, a full-length cDNA for mouse ghrelin has been cloned from the stomach using the oligo-capping and rapid amplification methods, and the organization of its gene and promoter has been analyzed. The mouse ghrelin cDNA was 521 bp long, consisting of 44 bp 5'-noncoding region, 354 bp coding region encoding a pre-proghrelin composed of 117 amino acid residues and 123 bp 3'-noncoding region. The genomic sequence analysis has revealed that the mouse ghrelin gene consists of 5 exons and 4 introns. The first exon was revealed to be only 19 bp long presented at the noncoding region of cDNA. The identical 19 bp sequence was also found as the first exon at the 5'-end of full-length rat ghrelin cDNA obtained from the stomach. A TATA box-like sequence, TATATAA was localized 24 bp upstream of the transcription start site of the mouse ghrelin gene. The sequence of the 5'-promoter region of mouse ghrelin gene including the TATA-like sequence and short exon 1 was highly homologous to that of reported human ghrelin gene. These findings suggest that the structure of the promoter region including the short noncoding first exon and its transcriptional regulation are conserved among the mammalian ghrelin genes.

Animals↗

"Word" preference in the genomic text and genome evolution: different modes of n-tuplet usage in coding and noncoding sequences.

Extensive work on n-tuplet occurrence in genomic sequences has revealed the correlation of their usage with sequence origin. Parallel to that, there exist different restrictions in the nucleotide composition of coding and noncoding sequences that may result in distinct modes of usage of n-tuplets. The relatively simple approaches described herein focus on such differences. They are based on simple summation measures of n-tuplet frequencies, computed after filtering the background nucleotide composition. Among the main targets of this work is to draw some conclusions on the qualitative differences in the composition of genomic sequences depending on their functionality. Moreover, an evolutionary model is formulated, including simple forms of ubiquitous events of genome dynamics: genomic fusions, genome shuffling due to transpositions, replication slippage, and point mutations. This model is shown to be able to reproduce all the statistical features of genomic sequences discussed herein.

Base Sequence↗

NfCR1, the first non-LTR retrotransposon characterized in the Australian lungfish genome, Neoceratodus forsteri, shows similarities to CR1-like elements.

The genomes of lungfish, together with those of some urodele amphibians, are the largest of all vertebrate genomes. It has been assumed that the bulk of the DNA making up these large genomes has been derived from repeat elements, like the noncoding DNA of those genomes that have been sequenced (e.g., human). In an attempt to characterize repeat sequences in the lungfish genome, we have isolated, by restriction enzyme digestion of genomic DNA, sequences of a repeat element in Neoceratodus forsteri, the most primitive of the living lungfishes. The fragments sequenced from the EcoRI and BglII digests were used to perform genome walking PCR in order to obtain the full sequence of the repeat element. This element shares homology with the non-LTR (LINE) element, Chicken Repeat 1 (CR1), described for several vertebrates and some invertebrates; we have called it N. forsteri CR1 (NfCR1). NfCR1 shares all the domains of other CR1 elements but it also has several unique features that suggest it may no longer be active in the lungfish genome. It occurs in both full-length and 5'-truncated versions and in its present "inactive" form represents approximately 0.05% of the lungfish genome.

Amino Acid Sequence↗

Gene expression, synteny, and local similarity in human noncoding mutation rates.

The human genome is organized with regard to many features such as isochores, Giemsa bands, clusters of genes with similar expression patterns, and contiguous regions with shared evolutionary histories (synteny blocks). In addition to these genomic features, it is clear that mutation rates also vary across the human genome. To address how mutation rates and genomic features are related, we analyzed substitution rates at three classes of putatively neutral noncoding sites (nongenic, intronic, and ancestral repeats) in approximately 14 Mb of human-chimpanzee alignments covering human chromosome 7. Patterns of mutation rate variation inferred from substitution rate variation differ among the three site classes. In particular, we find that intronic mutation rates are strongly affected by the breadth of expression of the genes in which they reside, with broadly expressed genes exhibiting low mutation rates, probably as a consequence of the transcription-coupled repair process acting in the germ line. All site classes show significant local similarities in mutation rate at the megabase scale, and regional similarities in nongenic mutation rate covary with blocks of synteny between the human and mouse genomes, indicating that the evolutionary history of a genomic region is an important determinant of mutation rate.

Animals↗

The 3'-nucleotides of flavivirus genomic RNA form a conserved secondary structure.

The terminal noncoding regions of viral RNA genomes are presumed to contain signal sequences and sometimes also secondary structures involved in regulating viral RNA synthesis. Such signals would be expected to be highly conserved among related viruses. In order to identify replication signal features for flaviviruses we have compared the 3'-terminal nucleotide sequences of West Nile virus (WNV), Saint Louis encephalitis (SLE) virus, and yellow fever virus (YFV) genome RNAs. The existence of a stable 3'-terminal secondary structure was previously predicted by a cDNA sequence obtained from YFV genome RNA. We have confirmed the existence of this structure by direct RNA sequencing methods. Even though the size and shape of the 3'-terminal secondary structure is highly conserved, sequence conservation is restricted to the loop regions of the secondary structure and to 27 nucleotides immediately adjacent to the 5' side of the structure. The regions of conserved sequence represent likely signals for viral polymerase recognition and binding. However, the preservation of the configuration of the secondary structure by a means other than sequence conservation indicate that this structure is important for the survival of the virus. A WNV mutant, which replicates progeny genome RNA more efficiently than parental WNV, was found to have a 3'-genomic sequence identical to that of its parent virus. The sequence change conferring the phenotype of this mutant is therefore located in another region of the genome.

Base Sequence↗

Genomic organization, expression, and comparative analysis of noncoding region of the rat Ndrg4 gene.

Rat Ndrg4 is a member of the NDRG gene family and has been suggested to relate to brain development. The structure of the rat Ndrg4 gene was studied to understand the mechanism for the expression of multiple forms of Ndrg4 protein, which were revealed in the brain. Subcloning and DNA sequencing analysis of a bacterial artificial chromosome (BAC) clone, together with analysis of a transcriptional start site by a cap-site hunting, indicated that the Ndrg4 gene spans about 39 kilobases (kb) and consists of 19 exons, in which the first and second exons were first found in rat. An alternative promoter usage at different transcriptional start sites may produce three types of messages, Ndrg4-A, Ndrg4-B, and Ndrg4-C, and there is a variant that lacks exon 18 for each type of transcript. Thereby, Ndrg4-A1, Ndrg4-A2, Ndrg4-B1, Ndrg4-B2, Ndrg4-C1, and Ndrg4-C2 were identified to be expressed. These six variants might explain the heterogeneity of the Ndrg4 protein in the brain. The variants without exon 18 were revealed in the embryonic and early postnatal brains while those with exon 18 were detected in the maturing and adult brains. Radiation hybrid mapping suggests that the rat Ndrg4 gene is located on chromosome 19 at 90.6 centirays (cR) from the top. Comparison of the noncoding sequence of the rat Ndrg4 gene to those of the orthologous mouse and human genes suggests that the AP-1 binding site is a candidate regulatory element.

Alternative Splicing↗

Chromosomal integration pattern of a helper-dependent minimal adenovirus vector with a selectable marker inserted into a 27.4-kilobase genomic stuffer.

Helper-dependent minimal adenovirus vectors are promising tools for gene transfer and therapy because of their high capacity and the absence of immunostimulatory or cytotoxic viral genes. In order to characterize this new vector system with respect to its integrative properties, the integration pattern of a minimal adenovirus vector with a neo(r) gene inserted centrally into a noncoding 27.4-kb genomic stuffer element derived from the human X chromosome after infection of a sex chromosome aneuploid (X0) human glioblastoma cell line was studied. Our results indicate that even extensive homologies and abundant chromosomal repeat elements present in the vector did not lead to integration of the vector via homologous or homology-mediated mechanisms. Instead, integration occurred primarily by insertion of a monomer with no or little loss of sequences at the vector ends, apparently at random sites, which is very similar to E1 deletion adenovirus vectors. It is therefore unlikely that the incorporation of stuffer elements derived from human genomic DNA, which were shown to allow long-term transgene expression in vivo in a number of studies, leads to an enhanced risk of insertional mutagenesis. Furthermore, our findings indicate that the potential of minimal adenovirus vectors as tools for targeted insertion and gene targeting is limited despite the possibility of incorporating long stretches of homologous sequences. However, we found an enhanced efficiency of stable neo(r) transduction of the minimal adenovirus vector compared to an E1 deletion adenovirus vector, possibly caused by the absence of potential growth-inhibitory viral genes. Complete integration of the vector and tolerance of the integrated vector sequences by the cell might indicate a potential use of these vectors as tools for stable transfer of (large) genes.

Adenoviruses, Human↗

Sequence, organization, and evolution of Rh50 glycoprotein genes in nonhuman primates.

The human RHAG locus encodes Rh50 glycoprotein, a polytopic protein that modulates expression of Rh antigens carried by Rh30 polypeptides. Rh50 is almost invariant, whereas Rh30 shows high polymorphism. To assess the relative conservation and phylogenetic relationship of RHAG genes, we characterized their protein expression, transcript structure, genomic organization, and noncoding regions (promoter and introns) in seven nonhuman primate species. Western blot showed that only ape Rh50 glycoproteins are recognized by the antibody 2D10 specific for the human counterpart. Analysis of RHAG gene and its transcript showed a high degree of sequence identity and features of interspecific diversity. The nonhuman primate RHAG genes are highly similar in promoter region and identical in exon-intron organization. Genomic sequencing identified one retro-transposon-like element in intron 2 and three types of Alu elements in intron 4 and 9, with varying copies of minisatellites. Reconstruction of coding and noncoding sequence trees revealed concordances and discordances with regard to the branching of RHAG-like genes in higher primates. A joined tree of Rh50 glycoproteins and Rh30 polypeptides shows that the former evolved at a rate about two times slower than the latter. Statistical tests demonstrated that at least a portion of the RHAG gene was subjected to a positive selection during evolution of anthropoids.

Amino Acid Sequence↗