Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Unbiased mapping of transcription factor binding sites along human chromosomes 21 and 22 points to widespread regulation of noncoding RNAs.

Using high-density oligonucleotide arrays representing essentially all nonrepetitive sequences on human chromosomes 21 and 22, we map the binding sites in vivo for three DNA binding transcription factors, Sp1, cMyc, and p53, in an unbiased manner. This mapping reveals an unexpectedly large number of transcription factor binding site (TFBS) regions, with a minimal estimate of 12,000 for Sp1, 25,000 for cMyc, and 1600 for p53 when extrapolated to the full genome. Only 22% of these TFBS regions are located at the 5' termini of protein-coding genes while 36% lie within or immediately 3' to well-characterized genes and are significantly correlated with noncoding RNAs. A significant number of these noncoding RNAs are regulated in response to retinoic acid, and overlapping pairs of protein-coding and noncoding RNAs are often coregulated. Thus, the human genome contains roughly comparable numbers of protein-coding and noncoding genes that are bound by common transcription factors and regulated by common environmental signals.

Amino Acid Motifs↗

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article↗

Replicative intermediates of human papillomavirus type 11 in laryngeal papillomas: site of replication initiation and direction of replication.

We have examined the structures of replication intermediates from the human papillomavirus type 11 genome in DNA extracted from papilloma lesions (laryngeal papillomas). The sites of replication initiation and termination utilized in vivo were mapped by using neutral/neutral and neutral/alkaline two-dimensional agarose gel electrophoresis methods. Initiation of replication was detected in or very close to the upstream regulatory region (URR; the noncoding, regulatory sequences upstream of the open reading frames in the papillomavirus genome). We also show that replication forks proceed bidirectionally from the origin and converge 180 degrees opposite the URR. These results demonstrate the feasibility of analysis of replication of viral genomes directly from infected tissue.

DNA Replication↗

Analysis of the set of GABA(A) receptor genes in the human genome.

The genes of the ionotropic gamma-aminobutyric acid receptor (GABR) subunits have shown an unusual chromosomal clustering, but only now can this be fully specified by analyses of the human genome. We have characterized the genes encoding the 18 known human GABR subunits, plus one now located here, for their precise locations, sizes, and exon/intron structures. Clusters of 17 of the 19, distributed between five chromosomes, are specified in detail, and their possible significance is considered. By applying search algorithms designed to recognize sequences of all known GABR-type subunits in species from man down to nematodes, we found no new GABR subunit is detectable in the human genome. However, the sequence of the human orthologue of the rat GABR rho3 receptor subunit was uncovered by these algorithms, and its gene could be analyzed. Consistent with those search results, orthologues of the beta4 and gamma4 subunits from the chicken, not cloned from mammals, were not detectable in the human genome by specific searches for them. The relationships are consistent with the mammalian subunit being derived from the beta line and epsilon from the gamma line, with mammalian loss of beta4 and gamma4. In their structures the human GABR genes show a basic pattern of nine coding exons, with six different genomic mechanisms for the alternative splicing found in various subunits. Additional noncoding exons occur for certain subunits, which can be regulatory. A dicysteine loop and its exon show remarkable constancy between all GABR subunits and species, of deduced functional significance.

Algorithms↗

Cloning and sequencing of potato virus Y (Hungarian isolate) genomic RNA.

A sequence of 9703 nucleotides (nt) is reported for the genomic RNA of potato virus Y (Hungarian isolate, PVY-H), which causes necrotic rings around the buds on the tubers and mottling of leaves. The sequence contains one large open reading frame of 3061 amino acids (aa), a noncoding region of 189 nt at the 5' end and a 330-nt 3' nontranslated region. The nt sequence and the predicted aa sequence of the polyprotein of PVY-H were analysed pairwise with the only available complete sequence of PVY strain N (PVYn) and with the partial sequences of different PVY strains, as well as with other potyviruses and potyvirus-related plant viruses. The overall relationship between PVY-H and PVYn shows a nt sequence identity of 88.5% and an aa sequence identity of 94.2%. The lowest degree of homology was detected at the 5' terminus of the genome, including the 5' noncoding region (70.3%) and the 275-aa P1 protein (78%). A fivefold sequence repeat block of 5'-UUUCA was found in the 5' noncoding region of PVY-H, which seems to be characteristic of PVY strains.

Amino Acid Sequence↗

Probable reassortment of genomic elements among elongated RNA-containing plant viruses.

The relationships of genome organization among elongated (rod-shaped and filamentous) plant viruses have been analyzed. Sequences in coding and noncoding regions of barley stripe mosaic virus (BSMV) RNAs 1, 2, and 3 were compared with those of the monopartite RNA genomes of potato virus X (PVX), white clover mosaic virus (WClMV), and tobacco mosaic virus, the bipartite genome of tobacco rattle virus (TRV), the quadripartite genome of beet necrotic yellow vein virus (BNYVV), and icosahedral tricornaviruses. These plant viruses belong to a supergroup having 5'-capped genomic RNAs. The results suggest that the genomic elements in each BSMV RNA are phylogenetically related to those of different plant RNA viruses. RNA 1 resembles the corresponding RNA 1 of tricornaviruses. The putative proteins encoded in BSMV RNA 2 are related to the products of BNYVV RNA 2, PVX RNA, and WClMV RNA. Amino acid sequence comparisons suggest that BSMV RNA 3 resembles TRV RNA 1. Also, it can be proposed that in the case of monopartite genomes, as a rule, every gene or block of genes retains phylogenetic relationships that are independent of adjacent genomic elements of the same RNA. Such differential evolution of individual elements of one and the same viral genome implies a prominent role for gene reassortment in the formation of viral genetic systems.

Amino Acid Sequence↗

Translation of hepatitis C virus genome.

Translation of the human hepatitis C virus (HCV) RNA genome occurs by internal ribosome entry through the 5' end (5' noncoding region) in a cap-independent fashion. The relatively long stretch of this noncoding region contains multiple initiation codons that are apparently not used for translation. Translation of the HCV polyprotein is initiated instead from an AUG located at nt 342. Using computer-assisted analysis (and subsequently substantiated by enzymatic probing), a complex secondary and tertiary structure of the 5' noncoding region (5'NCR) has been predicted. Based on an RNA folding model proposed by Brown et al. (1992), a detailed mutational analysis carried out identified the key secondary structural regions that are of functional significance in translational control. Maintenance of a helical structural element relevant to an oligopyrimidine tract is essential for internal initiation. A putative coaxial stacking or a pseudoknot structure upstream of the initiator AUG seems to be central to an internal ribosome entry site (IRES)-mediated translation of the HCV RNA genome.

Hepacivirus↗

Cell culture adaptation of Puumala hantavirus changes the infectivity for its natural reservoir, Clethrionomys glareolus, and leads to accumulation of mutants with altered genomic RNA S segment.

This paper reports the establishment of a model for hantavirus host adaptation. Wild-type (wt) (bank vole-passaged) and Vero E6 cell-cultured variants of Puumala virus strain Kazan were analyzed for their virologic and genetic properties. The wt variant was well adapted for reproduction in bank voles but not in cell culture, while the Vero E6 strains replicated to much higher efficiency in cell culture but did not reproducibly infect bank voles. Comparison of the consensus sequences of the respective viral genomes revealed no differences in the coding region of the S gene. However, the noncoding regions of the S gene were found to be different at positions 26 and 1577. In one additional and independent adaptation experiment, all analyzed cDNA clones from the Vero E6-adapted variant were found to carry substitutions at position 1580 of the S segment, just 3 nucleotides downstream of the mutation observed in the first adaptation. No differences were found in the consensus sequences of the entire M segments from the wt and the Vero E6-adapted variants. The results indicated different impacts of the S and the M genomic segments for the adaptation process and selective advantages for the variants that carried altered noncoding sequences of the S segment. We conclude that the isolation in cell culture resulted in a phenotypically and genotypically altered hantavirus.

Adaptation, Physiological↗

Architectural logic of the 3D genome: mechanisms of dysregulation and emerging cancer therapeutics.

The three-dimensional (3D) genome provides an essential layer of organization that shapes genome function in space and time. Chromatin compartments and topologically associating domains (TADs) arise from the interplay between intrinsic properties of chromatin and architectural factors, including cohesin and CTCF. Despite substantial progress in defining these structural features, whether 3D genome architecture plays a causal role in regulating processes such as transcription, DNA replication, and DNA repair, or instead reflects underlying regulatory activity, remains unresolved. Here, we use the distinction between chromatin-intrinsic features and architectural factors as a framework to evaluate evidence for causality in genome structure-function relationships. We extend this framework to cancer, where both intrinsic alterations (including noncoding mutations, structural variants, and changes in chromatin state) and architectural factor perturbations (such as mutations in architectural proteins and dysregulation of transcriptional machinery) disrupt genome organization and contribute to disease progression. These findings suggest that alterations in genome structure can, in some contexts, actively reshape oncogenic programs. A major limitation in applying 3D genome insights to cancer biology is the cost and complexity of omics assays. Recent advances in artificial intelligence (AI) and machine learning (ML) enable inference and prediction of 3D genome organization from sequence and epigenomic features, providing insight into the extent to which genome folding is encoded intrinsically versus dynamically regulated in architectural factors. This perspective provides a unified view of how genome structure is established, how it relates to function, and how its disruption contributes to tumorigenesis.

3D genome↗

Reverse transcription-PCR detection of hepatitis G virus.

Hepatitis G virus (HGV) was recently identified as a new member of the Flaviviridae, but its clinical significance is still unclear. Since no immunoassay for the diagnosis of HGV is available, we developed a sensitive reverse transcription-PCR (RT-PCR) assay to facilitate the detection of the viral genome by mass screening in the clinical laboratory. Sequences within the 5'-noncoding region and within the putative NS5a region are independently amplified in the presence of digoxigenin-11-dUTP and are detected by hybridization with biotinylated capture probes binding to a streptavidin-coated matrix. Semiquantitative Enzymun-Test DNA detection via chemiluminescence can be performed either in a microtiter plate format or on fully automated ES 300 machines. We were able to detect at least 8 x 10(2) genome equivalents per ml of serum using both primer pairs. HGV was shown to be present in 43 of 130 (33%) serum samples from intravenous drug abusers with a high risk of parenteral exposure. However, only two of the patients were positive when the NS5a primers only were used, and only one patient was positive when only the 5'-noncoding region primers were used, demonstrating the increased sensitivity of HGV detection with two sets of primers. Among these patients, there was no obvious correlation with other viral infections like hepatitis B virus, hepatitis C virus, or human immunodeficiency virus. Within a blood donor panel, 3 of 92 (3%) samples were found to be HGV positive, suggesting that donated blood may need to be screened for HGV.

Base Sequence↗

The 3' stem loop of the West Nile virus genomic RNA can suppress translation of chimeric mRNAs.

Cis-acting elements that regulate translation have been identified in the 3' noncoding regions (NCRs) of cellular and viral mRNAs. As one means of analyzing the effect on translation of the conserved 3' terminal RNA structure of the West Nile virus (WNV) genome, the translation efficiencies of chimeric mRNAs composed of a CAT reporter gene flanked by viral or nonviral 5' and 3' terminal sequences were compared. In vitro, the WNV 3'(+) stem loop (SL) RNA reduced the translation efficiencies of chimeric mRNAs with either viral or nonviral 5' NCRs, suggesting that a specific 3'-5' RNA-RNA interaction was not involved. In contrast, the 3' terminal sequence of a togavirus, rubella virus, enhanced translation efficiency. The WNV 3'(+)SL reduced translation efficiency both in cis and in trans and of both capped and uncapped chimeric mRNAs. We have previously reported that three cellular proteins bind specifically to the WNV 3'(+)SL RNA. Competition between the WNV 3'(+)SL and the 5' terminus of the chimeric mRNAs for proteins involved in translation initiation could explain the translation inhibition observed.

Animals↗

Genomic structure and analysis of transcriptional regulation of the mouse zinc-fingers and homeoboxes 1 (ZHX1) gene.

The mouse zinc-fingers and homeoboxes 1 (ZHX1) gene was cloned and its transcriptional regulatory mechanism analysed. The mouse ZHX1 gene spans approximately 29 kb and consists of five exons. Exons 1-3 contain the nucleotide sequence of the 5'-noncoding region of mouse ZHX1 cDNA, exon 4 contains a part of the 5'-noncoding region, an entire coding sequence, and a part of the 3'-noncoding sequence, and exon 5 contains the resulting 3'-noncoding sequence. The ZHX1 gene exists as one copy in the haploid mouse genome. Two species of ZHX1 mRNA with or without the nucleotide sequence of the third exon are produced by an alternative splicing. To investigate the regulatory elements involved in the transcription of the ZHX1 gene, transient DNA transfection experiments with ZHX1/firefly luciferase reporter genes were performed using a lipofection method. Functional analyses of a series of 5'- and 3'-deletion constructs of the reporter genes revealed that the nucleotide sequence between -59 and +50 is required for full promoter activity in mouse embryonal carcinoma F9 cells. Two positive regulatory cis-acting elements in the region were identified. These elements, designated as Box A and Box B, are located between nucleotides -47 and -42 and +22 and +27, respectively, and synergistically stimulate transcription of the mouse ZHX1 gene. Electrophoretic mobility shift assays with specific competitors and antibodies show that PEA3 and Yin and Yang 1 (YY1) bind to Box A and Box B, respectively. Thus, we conclude that PEA3 and YY1 synergistically stimulate the transcription of the ZHX1 gene.

Alternative Splicing↗

Isolation and partial characterization of the gene for goose fatty acid synthase.

Fatty acid synthase is regulated by diet and hormones, with regulation being primarily transcriptional. In chick embryo hepatocytes in culture, triiodothyronine stimulates accumulation of enzyme and transcription of the gene. Since the 5'-flanking region of this gene is likely involved in hormonal regulation of its expression, we have isolated and partially characterized an avian fatty acid synthase gene. A genomic DNA library was constructed in a cosmid vector and screened with cDNA clones that contained sequence complementary to the 3' end of goose fatty acid synthase mRNA. A genomic clone (approximately 35 kilobase pairs (kb] was isolated, and a 6.5-kb EcoRI fragment thereof contained DNA complementary to the 3' noncoding region of fatty acid synthase mRNA. Additional cosmid libraries were screened with 5' fragments of previously isolated genomic clones, resulting in the isolation of five overlapping cosmid DNAs. The entire region of cloned DNA spans approximately 105 kb. Exon-containing fragments were identified by hybridization with end-labeled poly(A)+ RNA and by hybridization of labeled exon-containing genomic DNA fragments to fatty acid synthase mRNA. A new set of cDNA clones spanning approximately 3.2 kb was isolated from a lambda-ZAP goose liver cDNA library using the 5'-most exon-containing fragment of the 5'-most genomic DNA clone. This region of mRNA contains a 5'-untranslated sequence and a continuous open reading frame which includes a region that codes for the essential cysteine of the beta-ketoacyl synthase domain. The entire fatty acid synthase gene spans about 50 kb. The 5' 15 kb of the gene contain 7 exons. S1 nuclease and primer extension analyses were used to identify a single site for initiation of transcription, 174 nucleotides upstream from the putative translation initiation codon. Putative "TATA" and "CCAAT" boxes are located 28 and 60 base pairs (bp), respectively, upstream of the site of initiation of transcription. The 5'-flanking 597 bp of DNA contains G/C-rich sequences including several "GC" boxes corresponding to binding sites for the nuclear transcription factor Sp1. Putative sites for AP-2, C/EBP, and the triiodothyronine and glucocorticoid receptors also were found in this region. A chimeric DNA, containing approximately 1.6 kb of 5'-flanking sequence and 139 bp of untranslated sequence of the goose fatty acid synthase gene ligated to the bacterial chloramphenicol acetyl-transferase (CAT) gene, was transfected into chick embryo hepatocytes in culture. Cells treated with triiodothyronine contained increased chloramphenicol acetyltransferase and fatty acid synthase activities.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

The mosaic genome of warm-blooded vertebrates.

Most of the nuclear genome of warm-blooded vertebrates is a mosaic of very long (much greater than 200 kilobases) DNA segments, the isochores; these isochores are fairly homogeneous in base composition and belong to a small number of major classes distinguished by differences in guanine-cytosine (GC) content. The families of DNA molecules derived from such classes can be separated and used to study the genome distribution of any sequence which can be probed. This approach has revealed (i) that the distribution of genes, integrated viral sequences, and interspersed repeats is highly nonuniform in the genome, and (ii) that the base composition and ratio of CpG to GpC in both coding and noncoding sequences, as well as codon usage, mainly depend on the GC content of the isochores harboring the sequences. The compositional compartmentalization of the genome of warm-blooded vertebrates is discussed with respect to its evolutionary origin, its causes, and its effects on chromosome structure and function.

Animals↗

Silence of the fathers: early X inactivation.

X chromosome inactivation is the mammalian answer to the dilemma of dosage compensation between males and females. The study of this fascinating form of chromosome-wide gene regulation has yielded surprising insights into early development and cellular memory. In the past few months, three papers reported unexpected findings about the paternal X chromosome (X(p)). All three studies agree that the X(p) is imprinted to become inactive earlier than ever suspected during embryonic development. Although apparently incomplete, this early form of inactivation insures dosage compensation throughout development. Silencing of the X(p) persists in cells of extraembryonic tissues, but it is erased and followed by random X inactivation in cells of the embryo proper. These findings challenge several aspects of the current view of X inactivation during early development and may have profound impact on studies of pluripotency and epigenetics.

Animals↗

Outbreak of jaundice and hemorrhagic fever in the Southeast of Brazil in 2001: detection and molecular characterization of yellow fever virus.

Between January and March 2001, an outbreak of jaundice and hemorrhagic fever occurred in the state of Minas Gerais, Southeast region of Brazil, in which a mortality rate of 53% was reported. Seroconversion, virus isolation, histopathological and immunohistochemical findings, and reverse transcription-polymerase chain reaction (RT-PCR) identified yellow fever virus (YFV) as the etiological agent responsible for the outbreak. Partial nucleotide sequence analysis from a fragment of the YFV genome spanning parts of nonstructural (NS) 5 gene and 3' noncoding region (3' UTR) showed that the YFV involved in this outbreak belongs to South American genotype I and differs from the Brazilian virus identified in 1996.

Amino Acid Sequence↗

The sequence of the reovirus serotype 3 L3 genome segment which encodes the major core protein lambda 1.

We present the sequence of reovirus serotype 3 (strain Dearing) genome segment L3 which encodes protein lambda 1, one of the two major components of the core shell. The genome segment is 3896 nucleotides long, with 5'- and 3'-noncoding regions of 13 and 181 nucleotides, respectively. Protein lambda 1 is 1233 amino acids long. It is a slightly acidic protein, with a predicted alpha-helix and beta-sheet content of 23.6 and 28.3%, respectively. Its rather low predicted alpha-helix contact is consistent with its being a structural protein. The 123 amino acid long region at its amino terminus is very hydrophilic and contains three alpha-helical regions, one being 26 amino acids long. Protein lambda 1 contains two functional motifs. The first is a nucleotide binding site -TKGKSSG- starting at residue 8, the other is a "zinc finger" motif centered around amino acid residue 194. This suggests that protein lambda 1 functions during the transcription of either dsRNA into plus strands or of plus strands into minus strands, or during both. It displays no significant sequence similarity to any protein sequence in the GenBank data base.

Amino Acid Sequence↗