Search PubMedSearch

SEARCH · Search PubMed

Results for “Databases, Nucleic Acid”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

17 recordsLinked to original sources

eccDNABase: A Comprehensive and High-Quality Database for Extrachromosomal Circular DNA.

Extrachromosomal circular DNA (eccDNA) refers to small, circular DNA molecules that originate from chromosomal sequences and are prevalent across nearly all eukaryotic organisms. In humans, eccDNAs are widely distributed in normal tissues, cancerous tissues, and body fluids, where they play important roles in tumorigenesis and are often associated with poor clinical outcomes. Given their biological and clinical significance, a well-integrated and high-quality database is essential for advancing eccDNA-related research. To address this need, we developed eccDNABase, a comprehensive and curated resource for browsing, searching, and analyzing eccDNAs across multiple species. The database systematically catalogs eccDNA-disease associations from diverse tissues and organisms. Currently, eccDNABase contains 1,875,452 eccDNA-disease associations, encompassing 8,398 ecDNA entries across nine species, 63 diseases, and healthy individuals. Each entry provides detailed information, including eccDNA ID, type, chromosomal localization, species, tissue or cell line source, disease name and Disease Ontology ID, overlap length and percentage with genes, oncogene overlap, detection method, and links to literature and source databases. Given its extensive and curated datasets, eccDNABase serves as a valuable resource for both basic and translational research, offering deeper insights into the role of eccDNA in health and disease. The database is publicly accessible at http://cgga.org.cn/eccDNABase/.

Humans

Reporting and representation of population descriptors in public RNA-seq databases.

Diverse and globally representative datasets are essential to genomic science and medicine. Here, we analyzed population descriptor metadata from RNA sequencing (RNA-seq) studies in two major public repositories: the Sequence Read Archive (SRA) and the Database of Genotypes and Phenotypes. We examined geographic and economic characteristics of institutions depositing the data and compared SRA-deposited descriptors to empirical estimates of genetic ancestry and to those reported in publications, analyzing trends over time. We found that 55% of RNA-seq samples were deposited by United States (US) institutions and 90% by institutions in high-income countries. Only 3% of SRA samples were associated with population descriptors, and among those with US Census terms, 69% were labeled as White. Among samples with continental descriptors, 56% were labeled as European. Our analyses emphasize widespread bias in the composition of public RNA-seq datasets and, more generally, a lack of consistent and careful reporting of population descriptors needing urgent improvement.

Humans

circASbase: A Comprehensive Database of Alternative Splicing Events in circRNAs.

Although extensive evidence has underscored the critical role of alternative splicing (AS) in generating mature circular RNA (circRNA) isoforms and augmenting their functional diversity, a significant gap remains in the availability of specialized databases housing circRNA AS events. To bridge this gap, we develop circASbase, a pioneering and comprehensive database that catalogs 452,129 AS events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNA. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of AS events, highlighting the unique regulatory landscape of circRNAs. These special splicing events result in functional differences of circRNAs by affecting internal ribosome entry sites, N6-methyladenosine sites, open reading frames, protein features, microRNA targets, and more. In summary, circASbase not only meets the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of AS events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is freely available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/.

Alternative Splicing

cfMethDB: A Comprehensive cfDNA Methylation Data Resource for Cancer Biomarkers.

Cancer is a major global health threat, and early detection is crucial for improving patient outcomes. DNA methylation in circulating cell-free DNA (cfDNA) has emerged as a promising biomarker for non-invasive cancer diagnosis. However, the integration and utilization of existing cfDNA methylation data have been limited, hindering comprehensive research efforts, particularly in the discovery of cfDNA methylation biomarkers. To address this challenge, we introduced cfMethDB, a comprehensive database dedicated to cfDNA methylation in cancer that encompasses 4828 publicly available datasets. Through standardized analysis, we identified 1,048,770 differentially methylated cytosines (DMCs) as candidate biomarkers across seven cancer types. With cfMethDB, we not only identified known cfDNA methylation biomarkers, but also discovered several genes, such as ZIC4, that could be novel biomarkers. Moreover, cfMethDB offers a suite of user-friendly tools, including biomarker evaluation, pan-cancer search, and end motif analysis. We hope that cfMethDB will serve as a valuable platform for the discovery of novel cancer cfDNA methylation biomarkers and facilitate cancer research and clinical applications. cfMethDB is publicly available at https://cfmethdb.hzau.edu.cn/home.

Humans

The INSDC specifications-foundations for a FAIR and global INSDC.

Members of the International Nucleotide Sequence Database Collaboration (INSDC; https://www.insdc.org/) collect, exchange, and preserve comprehensive open nucleotide sequence information and provide tools for its access. The INSDC has stated its commitment to welcoming new members into the collaboration to be more representative of the global community of data and users. To reach this goal, a comprehensive definition of the INSDC data model and minimum requirements for data acceptance have been established. Here we describe the processes used to arrive upon these INSDC Specifications and lay out strategies for their continued upkeep to remain current and relevant. Database URL:  https://www.insdc.org/.

Databases, Nucleic Acid

KERIS: kaleidoscope of gene responses to inflammation between species.

A cornerstone of modern biomedical research is the use of animal models to study disease mechanisms and to develop new therapeutic approaches. In order to help the research community to better explore the similarities and differences of genomic response between human inflammatory diseases and murine models, we developed KERIS: kaleidoscope of gene responses to inflammation between species (available at http://www.igenomed.org/keris/). As of June 2016, KERIS includes comparisons of the genomic response of six human inflammatory diseases (burns, trauma, infection, sepsis, endotoxin and acute respiratory distress syndrome) and matched mouse models, using 2257 curated samples from the Inflammation and the Host Response to Injury Glue Grant studies and other representative studies in Gene Expression Omnibus. A researcher can browse, query, visualize and compare the response patterns of genes, pathways and functional modules across different diseases and corresponding murine models. The database is expected to help biologists choosing models when studying the mechanisms of particular genes and pathways in a disease and prioritizing the translation of findings from disease models into clinical studies.

Animals

Apollo: a sequence annotation editor.

The well-established inaccuracy of purely computational methods for annotating genome sequences necessitates an interactive tool to allow biological experts to refine these approximations by viewing and independently evaluating the data supporting each annotation. Apollo was developed to meet this need, enabling curators to inspect genome annotations closely and edit them. FlyBase biologists successfully used Apollo to annotate the Drosophila melanogaster genome and it is increasingly being used as a starting point for the development of customized annotation editing tools for other genome projects.

Animals

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface

GenBank mining reveals novel insights into Rhizobium phylogeny: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct 'microbial h-index'.

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were 'citations.' The more one finds the same query sequence, the more that sequence can be considered biologically 'cited', i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index = 201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon's sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

RNA, Ribosomal, 16S

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein

Hematological diseases-related mucormycosis: A retrospective single center study.

BACKGROUND AND AIM: Mucormycosis is a life-threatening invasive fungal infection. This study aimed to analyze the clinical characteristics of patients with hematologic malignancies complicated with mucormycosis. METHODS: This retrospective study investigated the clinical characteristics, epidemiological features, treatment, and prognosis of 46 patients with hematological diseases and Mucor infection as indicated by mNGS from August 28, 2020 to September 11, 2023. Metagenomic next-generation sequencing (mNGS) refers to the application of high-throughput sequencing technology for the comprehensive analysis of nucleic acid content in patient samples, facilitating the detection and characterization of microbial DNA and/or RNA, and then comparing and analyzing the results with an information database to determine the types of pathogenic microorganisms present in the sample. RESULTS: The median age of admission for the included patients was 49 years (9-78). Multivariate analysis identified age over 60 years (p&#x2009;=&#x2009;0.006&#x2009;<&#x2009;0.05), high-dose corticosteroids (p&#x2009;=&#x2009;0.001&#x2009;<&#x2009;0.05), neutropenia lasting more than 10 days (p&#x2009;=&#x2009;0.041&#x2009;<&#x2009;0.05), and two or more Mucor infections (p&#x2009;=&#x2009;0.004&#x2009;<&#x2009;0.05) were independent risk factors for OS in patients with hematological diseases. Moreover, differences between groups were analyzed using the Fisher exact probability method, and no significant difference was observed in the efficacy of various types of antifungal therapies. CONCLUSION: Patients with hematologic malignancies benefit greatly from early diagnosis and treatment when suspected of Mucor infection. mNGS is an important supplementary method for early diagnosis of Mucor infection. Moderated use of corticosteroids, reducing the duration of neutropenia, and enhancing autologous immune function are important measures to reduce patient mortality rate.

Retrospective Studies

Detecting known neoepitopes, gene fusions, transposable elements, and circular RNAs in cell-free RNA.

MOTIVATION: Cancer is the second leading cause of death worldwide, and although there have been advances in treatments, including immunotherapies, these often require biopsies which can be costly and invasive to obtain. Due to lack of pre-emptive cancer detection methods, many cases of cancer are detected at a late stage when the definitive symptoms appear. Plasma samples are relatively easy to obtain, and they can be used to monitor the molecular signatures of ongoing processes in the body. Profiling cell-free DNA is a popular method for monitoring cancer, but only a few studies have explored the use of cell-free RNA (cfRNA), which shows the recent footprint of systemic transcription. RESULTS: Here, we developed FastNeo, a computational method for detecting known neoepitopes in human cfRNA. We show that neoepitopes and other biomarkers detected in cfRNA can discern Hepatocellular carcinoma patients from the healthy patients with a sensitivity of 0.84 and a specificity of 0.79. For colorectal cancer we achieve a sensitivity of 0.87 and a specificity of 0.8. An important advantage of our cfRNA based approach is that it also reports putative neoepitopes which are important for therapeutic purposes. AVAILABILITY AND IMPLEMENTATION: The FastNeo package is available at https://github.com/yashumayank/FastNeo and https://zenodo.org/records/11521368. The benchmark pipelines to detect Immune Epitope database and Tumor-Specific Neoantigen database neoepitopes using HaplotypeCaller, bcftools, and Lofreq, and to run FastNeo with STAR instead of Bowtie2 are also available in the above github repository.

Humans

Serum, cell-free, HPV-human DNA junction detection and HPV typing for predicting and monitoring cervical cancer recurrence.

Almost all cervical cancers are caused by human papillomaviruses (HPVs). In most cases, HPV DNA is integrated into the human genome. We found that tumor-specific, HPV-human DNA junctions are detectable in serum cell-free DNA of a fraction of cervical cancer patients at the time of initial treatment and/or at 6 months following treatment. Retrospective analysis revealed these junctions were more frequently detectable in women in whom the cancer later recurred. We also found that cervical cancers caused by HPV types outside of phylogenetic clade &#x3b1;9 had a higher recurrence frequency than those caused by &#x3b1;9 types in both our study and The Cancer Genome Atlas cervical cancer database, despite the higher prevalence of&#x3b1;9 types, including HPV16, in cervical cancer. Thus, HPV-human DNA junction detection in serum cell-free DNA and HPV type determination in tumor tissue may help predict recurrence risk. Screening serum cell-free DNA for junctions may also offer an unambiguous non-invasive means to monitor absence of recurrence following treatment.

Humans

G4STAB: a multi-input deep learning model to predict G-quadruplex thermodynamic stability based on sequence and salt concentration.

MOTIVATION: G-quadruplexes (G4s) are non-canonical nucleic acid structures formed in guanine-rich regions that modulate gene regulation and genomic stability. The thermodynamic stability of G4s directly influences their biological functions and potential as therapeutic targets. However, current quantitative frameworks for predicting G4 stability rely on predetermined structural features, limiting their effectiveness for diverse G4 topologies, and fail to account for environmental factors such as ion concentration and pH that significantly modulate G4 stability in cellular contexts. RESULTS: We present G4STAB, a multi-input deep learning neural network that accurately predicts DNA G4 melting temperatures based on sequence features, salt concentration, and pH. Trained on 2382 diverse DNA G4 sequences, our model achieves high accuracy (R&#x2002;2=0.8) without relying on predetermined G4 structural features. G4STAB successfully captures established G4 stability determinants and proposes previously unobserved sequence-stability relationships. Analysis of 391&#xa0;502 experimentally validated G4s reveals that cancer-like ionic environments alter G4 stability profiles, with a 13.5-fold increase in the number of structures exhibiting physiological melting temperatures (36-42&#xb0;C). These findings suggest systematic genomic patterns in G4 stability responses across chromosomes and gene types. AVAILABILITY AND IMPLEMENTATION: G4STAB is available at https://github.com/donn-liew/G4STAB; G4STAB web database interface is available at https://donn-liew.github.io/g4stab-web-database/.

G-Quadruplexes

Visualization using NIPTviewer support the clinical interpretation of noninvasive prenatal testing results.

BACKGROUND: Noninvasive prenatal testing (NIPT) is increasingly used to screen for fetal chromosomal aneuploidy by analyzing cell-free DNA (cfDNA) in peripheral maternal blood. The method provides an opportunity for early detection of large genetic abnormalities without an increased risk of miscarriage due to invasive procedures. Commercial applications for use at clinical laboratories often take advantage of DNA sequencing technologies and include the bioinformatic workup of the sequence data. The interpretation of the test results and the clinical report writing, however, remains the responsibility of the diagnostic laboratory. In order to facilitate this step, we developed NIPTviewer, a web-based application to visualize and guide the interpretation of NIPT data results. RESULTS: NIPTviewer has a database functionality to store the NIPT results and a web interface for user interaction and visualization. The application has been implemented as part of a novel analysis pipeline for NIPT in a diagnostic laboratory at Uppsala University Hospital. The validation data set included 84 previously analyzed plasma samples with known results regarding chromosomes 13, 18, 21, X and Y. They were sequenced in six different experiments, uploaded to NIPTviewer and assigned to a clinical laboratory geneticist for interpretation. The results of all previously analyzed samples were replicated. CONCLUSION: NIPTviewer facilitates NIPT results interpretation and has been implemented as part of a NIPT analysis routine that was accredited by the national accreditation body for Sweden (Swedac).

Humans

Structural Features of DNA in TATA-Containing and TATA-Less Core Promoters of RNA Polymerase II Differ.

Nucleotide motifs in the core promoters of eukaryotic protein-coding genes transcribed by RNA polymerase II (Pol II) play an important role in the transcription process. We analyzed the role of an octanucleotide located in the TATA box position. Depending on whether this octanucleotide can form a complex with the TATA-binding protein (TBP), the promoter is classified as either TATA-containing or TATA-less. We analyzed the differences in the primary and spatial structures, as well as their dynamics, in TATA-containing and TATA-less promoters of mammals and plants. We divided the complete promoter sets of six organisms (H. sapiens, M. musculus, C. familiaris, A. thaliana, Z. mays, and H. vulgare) from the EPDnew database into TATA-containing and TATA-less fractions. The sizes of the TATA-containing promoter fractions are significantly smaller than those of the TATA-less fractions in all studied organisms, except in A. thaliana, where the sizes of both fractions are approximately equal. We characterized promoter architecture using variation profiles of various base-pair step parameters, minor-groove width, and the conformational dynamics of native DNA. The architectures of TATA-containing and TATA-less promoters differ significantly. The possible mechanistic influence of DNA structural features on the formation of the pre-initiation complex (PIC) in both types of promoters is discussed.

Promoter Regions, Genetic