Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans↗

Experimental validation of predicted mammalian erythroid cis-regulatory modules.

Multiple alignments of genome sequences are helpful guides to functional analysis, but predicting cis-regulatory modules (CRMs) accurately from such alignments remains an elusive goal. We predict CRMs for mammalian genes expressed in red blood cells by combining two properties gleaned from aligned, noncoding genome sequences: a positive regulatory potential (RP) score, which detects similarity to patterns in alignments distinctive for regulatory regions, and conservation of a binding site motif for the essential erythroid transcription factor GATA-1. Within eight target loci, we tested 75 noncoding segments by reporter gene assays in transiently transfected human K562 cells and/or after site-directed integration into murine erythroleukemia cells. Segments with a high RP score and a conserved exact match to the binding site consensus are validated at a good rate (50%-100%, with rates increasing at higher RP), whereas segments with lower RP scores or nonconsensus binding motifs tend to be inactive. Active DNA segments were shown to be occupied by GATA-1 protein by chromatin immunoprecipitation, whereas sites predicted to be inactive were not occupied. We verify four previously known erythroid CRMs and identify 28 novel ones. Thus, high RP in combination with another feature of a CRM, such as a conserved transcription factor binding site, is a good predictor of functional CRMs. Genome-wide predictions based on RP and a large set of well-defined transcription factor binding sites are available through servers at http://www.bx.psu.edu/.

Amino Acid Motifs↗

Rates and patterns of molecular evolution in inbred and outbred Arabidopsis.

The evolution of self-fertilization is associated with a large reduction in the effective rate of recombination and a corresponding decline in effective population size. If many spontaneous mutations are slightly deleterious, this shift in the breeding system is expected to lead to a reduced efficacy of natural selection and genome-wide changes in the rates of molecular evolution. Here, we investigate the effects of the breeding system on molecular evolution in the highly self-fertilizing plant Arabidopsis thaliana by comparing its coding and noncoding genomic regions with those of its close outcrossing relative, the self-incompatible A. lyrata. More distantly related species in the Brassicaceae are used as outgroups to polarize the substitutions along each lineage. In contrast to expectations, no significant difference in the rates of protein evolution is observed between selfing and outcrossing Arabidopsis species. Similarly, no consistent overall difference in codon bias is observed between the species, although for low-biased genes A. lyrata shows significantly higher major codon usage. There is also evidence of intron size evolution in A. thaliana, which has consistently smaller introns than its outcrossing congener, potentially reflecting directional selection on intron size. The results are discussed in the context of heterogeneity in selection coefficients across loci and the effects of life history and population structure on rates of molecular evolution. Using estimates of substitution rates in coding regions and approximate estimates of divergence and generation times, the genomic deleterious mutation rate (U) for amino acid substitutions in Arabidopsis is estimated to be approximately 0.2-0.6 per generation.

Arabidopsis↗

Analysis of multiple restriction fragment length polymorphisms of the gene for the human complement receptor type I. Duplication of genomic sequences occurs in association with a high molecular mass receptor allotype.

Human CR1 exhibits an unusual form of polymorphism in which allotypic variants differ in the molecular weight of their respective polypeptide chains. To address mechanisms involved in the generation of the CR1 allotypes, DNA from individuals having the F allotype (250,000 Mr), the S allotype (290,000 Mr), and the F' allotype (210,000 Mr) was digested by restriction enzymes, and Southern blots were hybridized with CR1 cDNA and genomic probes. With the use of Bam HI and Sac I, an additional restriction fragment was observed in 20 of 21 individuals having the S allotype with no associated loss of other restriction fragments. Southern blot analysis with a noncoding genomic probe derived from the S allotype-specific Bam HI fragment showed hybridization to this fragment and to two other fragments that were also present in FF individuals. Thus, an intervening sequence may be repeated twice in the F allele and three times in the S allele. A restriction fragment length polymorphism (RFLP) unique to two individuals expressing the F' allotype was seen with Eco RV, but the absence of persons homozygous for this rare allotype prevented further comparisons with the F and S allotypes. Analysis of the CR1 transcripts associated with the three CR1 allotypes indicated that these differed by 1.3-1.5 kb and had the same rank order as the corresponding allotypes. Taken together, these findings suggest that the S allele was generated from the F allele by the acquisition of additional sequences, the coding portion of which may correspond to a long homologous repeat of approximately 1.4 kb that has been identified in CR1 cDNA. We saw two other RFLPs with Hind III and Pvu II that were in linkage dysequilibrium with the Bam HI-Sac I RFLPs associated with the S allotype, and a third polymorphism was seen with Eco RI that was not in linkage dysequilibrium with the other polymorphisms. Thus, 10 commonly occurring CR1 alleles can be defined, making this locus a useful marker for the long arm of chromosome 1 to which the CR1 gene maps.

Alleles↗

Canavan disease: genomic organization and localization of human ASPA to 17p13-ter and conservation of the ASPA gene during evolution.

Canavan disease, or spongy degeneration of the brain, is a severe leukodystrophy caused by the deficiency of aspartoacylase (ASPA). Recently, a missense mutation was identified in human ASPA coding sequence from patients with Canavan disease. The human ASPA gene has been cloned and found to span 29 kb of the genome. Human aspartoacylase is coded by six exons intervened by five introns. The exons vary from 94 (exon III) to 514 (exon VI) bases. The exon/intron splice junction sites follow the gt/ag consensus sequence rule. Southern blot analysis of genomic DNA from human/mouse somatic cell hybrid cell lines localized ASPA to human chromosome 17. The human ASPA locus was further mapped in the 17p13-ter region by fluorescence in situ hybridization. The bovine aspa gene has also been cloned, and its exon/intron organization is identical to that of the human gene. The 500-base sequence upstream of the initiator ATG codon in the human gene and that in the bovine gene are 77% identical. Human ASPA coding sequences cross-hybridize with genomic DNA from yeast, chicken, rabbit, cow, dog, mouse, rat, and monkey. The specificity of cross-species hybridization of coding sequences suggests that aspartoacylase has been conserved during evolution. It should now be possible to identify mutations in the noncoding genomic sequences that lead to Canavan disease and to study the regulation of ASPA.

Amidohydrolases↗

Functional Annotation of the Major Histocompatibility Complex Locus.

The human major histocompatibility complex (MHC) locus has the greatest density of disease-associations in the human genome, including links to over 100 polygenic disorders. Its complex haplotype structure, rich gene density, and high degree of linkage disequilibrium combine to make deciphering the gene regulatory logic of the MHC locus extremely challenging. Employing complementary high-throughput CRISPR interference (CRISPRi) and activation (CRISPRa) epigenetic screens coupled with single-cell transcriptome profiling across three distinct human cell types, we identified hundreds of new connections between cis -regulatory elements (CREs) and their target genes in this locus. These CRE-gene links are largely cell type-specific and act as enhancers. Additionally, some CREs have complex features, including harboring both active and repressive histone marks, lacking chromatin accessibility, targeting multiple genes, or acting as silencers. Computational methods fail to predict a majority of these CRE-gene connections. These findings emphasize the potential for functional perturbation experiments to dissect complex loci and reveal shared and cell type-specific regulatory mechanisms relevant to genomics of complex diseases. Collectively, this study provides a unique resource for understanding the complex regulatory landscape within the MHC locus and supports the need for creating new models that encompass CRE-gene interactions, cell type-specific gene expression, and disease genetics in the noncoding genome.

Journal Article↗

Application of genome sequence information to the classification of bovine enteroviruses: the importance of 5'- and 3'-nontranslated regions.

Comparative genomics of viruses in evolutionary and phylogenetic studies is well established. Previous nucleic acid sequence analyses have demonstrated that enteroviruses and rhinoviruses of the family Picornaviridae exhibit a similar structure of the 5'-nontranslated region (NTR) differing significantly from the 5'-NTR of cardiovirus, aphthovirus, hepatovirus, and echovirus 22 (provisionally parechovirus 1). Available nucleotide sequence information of the 5'- and 3'-nontranslated regions of more than 70 serotypes of enteroviruses, bovine enteroviruses and rhinoviruses has been compared and correlated with previous findings obtained after analysis of the coding and noncoding genome regions. As a result, the 5'- and 3'-NTRs of all three virus groups are characterized by group-specific nucleotide sequences. Focusing on bovine enterovirus (BEV) serotypes, unique characteristics in all secondary structures of the NTRs were observed. These features clearly separate the BEVs from the human enteroviruses and rhinoviruses. Concerning the 5'-NTR, the most remarkable property is an insertion of about 110 nucleotides between the putative cloverleaf structure at the very 5'-end of the viral genome and the IRES element. This insertion was demonstrated for BEV 1 and 2 and has a predicted folding pattern which is very similar to the 5'-cloverleaf structure. One stem-loop of this second cloverleaf is almost identical to the 3CDpro-binding domain of rhinoviral 5'-cloverleafs. It was also demonstrated that the IRES elements and the 3'-NTRs of both, enteroviruses and rhinoviruses, have group-specific features which differ significantly from the corresponding genome regions of BEV. These results suggest that bovine enteroviruses hold an exceptional taxonomic position besides the established genera Enterovirus and Rhinovirus. Within the Enterovirus and Rhinovirus genera, the existence of virus clusters representing subgenera was previously proposed. Whereas the 5'-NTRs of the four human enterovirus clusters fall into two groups, all four clusters have characteristic secondary structures at the 3'-NTR supporting the concept of enterovirus clusters. For rhinoviruses, the existence of two virus clusters was confirmed.

Animals↗

Emergence of Talanin protein associated with human uric acid nephrolithiasis in the Hominidae lineage.

Recently, we identified a susceptibility locus for human uric acid nephrolithiasis (UAN) on 10q21-q22 and demonstrated that a novel gene (ZNF365) included in this region produces through alternative splicing several transcripts coding for four protein isoforms. Mutation analysis showed that one of them (Talanin) is associated with UAN. We examined the evolutionary conservation of ZNF365 gene through a comparative genomic approach. Searching for mouse homologs of ZNF365 transcripts, we identified a highly conserved mouse ortholog of ZNF365A transcript, expressed specifically in brain. We did not found a mouse homolog for ZNF365D transcript encoding the Talanin protein, even if we were able to identify the corresponding genomic region in mouse and rat not yet organized in canonical gene structure suggesting that ZNF365D was originated after the branching of hominoid from rodent lineage. In mouse and in most mammals, a functional uricase degrades the uric acid to allantoin, but uricase activity was lost during the Miocene epoch in hominoids. Searching for the presence of Talanin in Primates, we found a canonical intron-exon structure with several stop codons preventing protein production in Old World and New World monkeys. In humans, we observe expression and we have evidence that ZNF365D transcript produces a functional protein. It seems therefore that ZNF365D transcript emerged during primate evolution from a noncoding genomic sequence that evolved in a standard gene structure and assumed its role in parallel with the disappearance of uricase, probably against a disadvantageous excessive hyperuricemia.

Alternative Splicing↗

A 1.1-Mb transcript map of the hereditary hemochromatosis locus.

In the process of positionally cloning a candidate gene responsible for hereditary hemochromatosis (HH), we constructed a 1.1-Mb transcript map of the region of human chromosome 6p that lies 4.5 Mb telomeric to HLA-A. A combination of three gene-finding techniques, direct cDNA selection, exon trapping, and sample sequencing, were used initially for a saturation screening of the 1.1-Mb region for expressed sequence fragments. As genetic analysis further narrowed the HH candidate locus, we sequenced completely 0.25 Mb of genomic DNA as a final measure to identify all genes. Besides the novel MHC class 1-like HH candidate gene HLA-H, we identified a family of five butyrophilin-related sequences, two genes with structural similarity to a type 1 sodium phosphate transporter, 12 novel histone genes, and a gene we named RoRet based on its strong similarity to the 52-kD Ro/SSA lupus and Sjogren's syndrome auto-antigen and the RET finger protein. Several members of the butyrophilin family and the RoRet gene share an exon of common evolutionary origin called B30-2. The B30-2 exon was originally isolated from the HLA class 1 region, yet has apparently "shuffled" into several genes along the chromosome telomeric to the MHC. The conservation of the B30-2 exon in several novel genes and the previously described amino acid homology of HLA-H to MHC class 1 molecules provide further support that this gene-rich region of 6p21.3 is related to the MHC. Finally, we performed an analysis of the four approaches for gene finding and conclude that direct selection provides the most effective probes for cDNA screening, and that as much as 30% of ESTs in this 1.1-Mb region may be derived from noncoding genomic DNA.

Amino Acid Sequence↗

Functional RNA microarrays for high-throughput screening of antiprotein aptamers.

High-throughput methods for generating aptamer microarrays are described. As a proof-of-principle, the microarrays were used to screen the affinity and specificity of a pool of robotically selected antilysozyme RNA aptamers. Aptamers were transcribed in vitro in reactions supplemented with biotinyl-guanosine 5'-monophosphate, which led to the specific addition of a 5' biotin moiety, and then spotted on streptavidin-coated microarray slides. The aptamers captured target protein in a dose-dependent manner, with linear signal response ranges that covered seven orders of magnitude and a lower limit of detection of 1 pg/mL (70 fM). Aptamers on the microarray retained their specificity for target protein in the presence of a 10,000-fold (w/w) excess of T-4 cell lysate protein. The RNA aptamer microarrays performed comparably to current antibody microarrays and within the clinically relevant ranges of many disease biomarkers. These methods should also prove useful for generating other functional RNA microarrays, including arrays for genomic noncoding RNAs that bind proteins. Integrating RNA aptamer microarray production with the maturing technology for automated in vitro selection of antiprotein aptamers should result in the high-throughput production of proteome chips.

Base Sequence↗

DNA sequence evolution with neighbor-dependent mutation.

We introduce a model of DNA sequence evolution which can account for biases in mutation rates that depend on the identity of the neighboring bases. An analytic solution for this class of models is developed by adopting well-known methods of nonlinear dynamics. Results are presented for the CpG-methylation-deamination process, which dominates point substitutions in vertebrates. The dinucleotide frequencies generated by the model (using empirically obtained mutation rates) match the overall pattern observed in noncoding DNA. A web-based tool has been constructed to compute single- and dinucleotide frequencies for arbitrary neighbor-dependent mutation rates. Also provided is the backward procedure to infer the mutation rates using maximum likelihood analysis given the observed single- and dinucleotide frequencies. Reasonable estimates of the mutation rates can be obtained very efficiently, using generic noncoding DNA sequences as input, after masking out long homonucleotide subsequences. Our method is much more convenient and versatile to use than the traditional method of deducing mutation rates by counting mutation events in carefully chosen sequences. More generally, our approach provides a more realistic but still tractable description of noncoding genomic DNA and may be used as a null model for various sequence analysis applications.

DNA↗

Identification of scaffold/matrix attachment region in recurrent site of woodchuck hepatitis virus integration.

Scaffold or matrix attachment regions (S/MARs) are noncoding genomic DNA sequences displaying in vitro selective binding affinity for nuclear scaffold. They have been reported to be involved in the physical attachment of genomic DNA to the nuclear scaffold, and thus in the organization of the chromatin in functional loops or domains, and in the regulation of gene expression. In this work, we report the identification of an S/MAR in a woodchuck chromosomal locus, named b3n, previously described as a recurrent site of woodchuck hepatitis virus (WHV) DNA integration in woodchuck hepatocellular carcinoma (HCC). The 4.3-kb sequence of this locus contains several Alu-like repeats and a gag-like coding region with frameshift mutations. Computer analysis revealed the presence of a region with unusually high AT content, typical of most S/MARs, and of specific motifs (A boxes, T boxes, topoisomerase II sites, and unwinding elements) overlapping or in proximity to the region with high AT content, predicting that b3n might contain an S/MAR. Fragments of the b3n locus were isolated by conventional and inverse PCR techniques. In in vitro binding experiments with both heterologous and autologous scaffold preparations, a 592-bp fragment spanning the region rich in S/MAR features showed marked scaffold affinity, which was specific when autologous scaffolds were used. The presence of an S/MAR at the b3n locus and its nature as a recurrent WHV integration site in HCC suggest the involvement of S/MAR elements in some of the mechanisms leading to liver oncogenesis.

Amino Acid Sequence↗

Detection and visualization of compositionally similar cis-regulatory element clusters in orthologous and coordinately controlled genes.

Evolutionarily conserved noncoding genomic sequences represent a potentially rich source for the discovery of gene regulatory regions. However, detecting and visualizing compositionally similar cis-element clusters in the context of conserved sequences is challenging. We have explored potential solutions and developed an algorithm and visualization method that combines the results of conserved sequence analyses (BLASTZ) with those of transcription factor binding site analyses (MatInspector) (http://trafac.chmcc.org). We define hits as the density of co-occurring cis-element transcription factor (TF)-binding sites measured within a 200-bp moving average window through phylogenetically conserved regions. The results are depicted as a Regulogram, in which the hit count is plotted as a function of position within each of the two genomic regions of the aligned orthologs. Within a high-scoring region, the relative arrangement of shared cis-elements within compositionally similar TF-binding site clusters is depicted in a Trafacgram. On the basis of analyses of several training data sets, the approach also allows for the detection of similarities in composition and relative arrangement of cis-element clusters within nonorthologous genes, promoters, and enhancers that exhibit coordinate regulatory properties. Known functional regulatory regions of nonorthologous and less-conserved orthologous genes frequently showed cis-element shuffling, demonstrating that compositional similarity can be more sensitive than sequence similarity. These results show that combining sequence similarity with cis-element compositional similarity provides a powerful aid for the identification of potential control regions.

Animals↗

Autoantibodies and hepatitis C virus genotypes in chronic hepatitis C patients in Estonia.

AIM: To determine the prevalence of several autoantibodies in chronic hepatitis C patients, and to find out whether the pattern of autoantibodies was associated with hepatitis C virus (HCV) genotypes. METHODS: Sera from 90 consecutive patients with chronic hepatitis C were investigated on the presence of anti-nuclear (ANA), anti-mitochondrial (AMA), anti-smooth muscle (SMA), anti-liver-kidney microsomal type 1 (LKMA1), anti-parietal cell (PCA), anti-thyroid microsomal (TMA), and anti-reticulin (ARA) autoantibodies. The autoantibodies were identified by indirect immunofluorescence. HCV genotypes were determined by a restriction fragment length polymorphism analysis of the amplified 5' noncoding genome region. RESULTS: Forty-six (51.1%) patients were positive for at least one autoantibody. Various antibodies were presented as follows: ANA in 13 (14.4%) patients, SMA in 39 (43.3%), TMA in 2 (2.2%), and ARA in 1 (1.1%) patients. In 9 cases, sera were positive for two autoantibodies (ANA and SMA). AMA, PCA and LKMA1 were not detected in the observed sera. HCV genotypes were distributed as follows: 1b in 66 (73.3%) patients, 3a in 18 (20.0%), and 2a in 6 (6.7%) patients. CONCLUSION: A high prevalence of ANA and SMA can be found in chronic hepatitis C patients. Autoantibodies are present at low titre (1:10) in most of the cases. Distribution of the autoantibodies show no differences in the sex groups and between patients infected with different HCV genotypes.

Autoantibodies↗

Microsatellite instability is uncommon in breast cancer.

In some tumors, defects in mismatch repair enzymes lead to errors in the replication of simple nucleotide repeat segments. This condition is commonly known as microsatellite instability (MSI) because of the frequent mutations of microsatellite sequences. Although the MSI phenotype is well recognized in some colon, gastric, pancreatic, and endometrial cancers, reports of MSI in breast cancer are inconsistent. We report here our experience with >10,000 amplifications of simple nucleotide repeats in noncoding genomic regions using DNA from 267 cases of breast cancer, including cases that represent all major histological types of breast cancer. We rarely (10 reactions) found unexpected bands in amplifications of tumor DNA that were not present in amplifications of normal DNA. Moreover, repeats of these reactions did not confirm microsatellite instability in a single case. We also evaluated the simple nucleotide repeats in the transforming growth factor type II receptor, insulin-like growth factor type II receptor, BAX, and E2F-4 genes, which are frequently mutated in tumors with microsatellite instability. No mutations of these genes were found in any of the 30 breast cancer cell lines and 61 primary breast cancer samples examined. These results indicate that mismatch repair errors characteristic of the MSI phenotype are uncommon in human breast cancer.

Breast Neoplasms↗

[Nuclear genetic material as an initial substrate for animal aging].

General properties of aging in animals are considered on the basis of the literature evidence and the results obtained by the authors of this paper. The existence of a specific aging mechanism is inferred. The operation of this mechanism is controlled not only by genes but also by particular noncoding genomic sequences with variable structure. The beginning of senescence in animals is determined by DNA lesions located in neural cells and probably in a minor genomic fraction. The authors refute the narrow concept of aging as a mechanism increasing the probability of death. Mortality as a continuous process occurring with the probability of 100 percent is an integral attribute of living organisms on the Earth.

Aging↗

The deletion of 41 proximal nucleotides reverts a poliovirus mutant containing a temperature-sensitive lesion in the 5' noncoding region of genomic RNA.

We generated a number of small deletions and insertions in the 5' noncoding region of an infectious cDNA copy of the poliovirus RNA genome. Transfection of these mutated cDNAs into COS-1 cells produced the following phenotypic categories: (i) wild-type mutations, (ii) lethal mutations, (iii) mutations exhibiting slow growth or low-titer properties, and (iv) temperature-sensitive (ts) mutations. The deletion of nucleotides 221 to 224 produced a ts virus, 220D1. Mutant 220D1 was found to have a dramatic reduction in growth, virus-specific protein and RNA synthesis, and the shutoff of host cell protein synthesis at 37 or 39 degrees C compared with 33 degrees C. Temperature shift experiments showed that the mutant viral RNA is not an effective template for protein or RNA synthesis at 39 degrees C and suggested a decreased stability of the 220D1 RNA at 39 degrees C. Selection for a non-ts revertant of 220D1 yielded the virus R2, which was no longer ts for growth or viral protein and RNA synthesis. Sequencing the 5' noncoding region of the genomic RNA from R2 revealed the deletion of 41 proximal nucleotides for an overall deletion of nucleotides 184 to 228. These data suggest that the deleted sequences are nonessential to the poliovirus life cycle during growth in HeLa cells. According to computer-predicted RNA secondary structures of the 5' noncoding region of poliovirus RNA, the R2 revertant virus has deleted an entire predicted stem-loop structure.

Animals↗

[A full-size DNA copy of the tick-borne encephalitis virus genome. I. Analysis of the 5'- and 3'-terminal noncoding regions of the genome].

Using reverse transcription and the polymerase chain reaction, cDNA fragments of noncoding regions of the tick-borne encephalitis virus (TBEV) genome were obtained. These fragments were cloned into a pGEM3 vector, and their nucleotide sequences were determined. The heterogeneity of the 3'-terminal untranslated region of the TBEV RNA was revealed. To create a stable full-size DNA copy of the TBEV genome, four cDNA variants differing in length and structure of the 3'-terminal fragment of the viral RNA were cloned into a pBR322-derived vector.

Base Sequence↗