Search PubMedSearch

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Chromosome-level genome assembly of the longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae).

The longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae) is a widely distributed wood-boring pest of conifers. Here, we assembled a chromosome-level genome of A. rusticus using Illumina, Oxford Nanopore, and Hi-C sequencing technologies. The assembled genome is 1180.40 Mb, with a scaffold N50 of 125.01 Mb, and BUSCO completeness of 93.6%. All contigs were assembled into ten pseudo-chromosomes. The genome contains 69.87% repeat sequences. We identify 18, 377 protein-coding genes in the genome, of which 11,368 were functionally annotated. This genome provides a valuable resource for understanding the ecology, genetics, and evolution of A. rusticus, as well as for controlling wood-boring pests.

Animals

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55 Gb, with a scaffold N50 of 93.38 Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant

Near-complete reference genome assembly of Hoya carnosa.

Hoya R. Br. is the largest genus in the tribe Marsdenieae (Apocynaceae), comprising 350-450 species. Hoya species are popular in horticulture for their distinctive floral traits and fragrances, primarily sourced from domestication and mutation breeding. However, the lack of molecular analysis for floral morphological traits has limited their cultivation and application. In this study, we assembled a near-complete reference genome for H. carnosa, the model species of the genus, using PacBio HiFi reads and Hi-C method. The genome size was approximately 465.7 Mb with a contig N50 of 39.3 Mb. 99.7% of the sequences were anchored to 11 pseudochromosomes, and the assembly achieved a BUSCO score of 98.5%. We predicted 24,309 protein-coding genes, of which 90.2% (21,927) were functionally annotated. This high-quality genome provides a valuable reference for the research of evolution, conservation and molecular breeding in Hoya.

Genome, Plant

Chromosome-level genome assembly of bivalve mollusk, Xishishe Coelomactra antiquata.

Coelomactra antiquata, a significant marine economic shellfish in China, is experiencing a natural population decline due to habitat destruction and overfishing, making the restoration and conservation of its natural resources an urgent priority. This study provides a high - quality chromosome - level genome assembly for C. antiquata, created by PacBio and Hi - C sequencing and resulting in a 19 - chromosome map. The assembly encompasses a genome size of 807.31 Mb, with a contig N50 of 17.35 Mb and a scaffold N50 of 42.90 Mb. A total of 28,070 protein - coding genes were identified, 25,959 of which were functionally annotated. Overall, this study offers a chromosome - level genome for C. antiquata that is highly continuous and complete, providing an indispensable resource for subsequent molecular and genetic studies of this species.

Animals

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38 Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55 Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87 Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals

Chromosome-level genome assembly of a cosmopolitan marine harmful algal bloom diatom species Chaetoceros socialis (Chaetocerotaceae).

Chaetoceros socialis is a cosmopolitan diatom species that is crucial for maintaining marine ecosystem structure and driving elemental cycles. C. socialis can form harmful algal blooms (HABs) that may cause a negative impact on the marine ecosystems. Whole-genome information for C. socialis is still unavailable, which may hinder more targeted studies on its ecological adaptive responses and evolutionary drivers. To address this gap, we employed cutting-edge genomic technologies including PacBio single-molecule real-time (SMRT) sequencing and high-throughput chromatin conformation capture (Hi-C) to achieve the first chromosome-level genome assembly of C. socialis. The assembled genome is 60.22 Mb in size with a scaffold N50 of 7.81 Mb and has been anchored to eight pseudochromosomes. A total of 13,378 protein-coding genes were predicted, of which 12,069 (90.22%) were functionally annotated. This high-quality genomic resource provides a fundamental data platform for systematically elucidating the ecological adaptation mechanisms of C. socialis.

Chromosomes

A high-quality chromosome-level genome assembly of apple of Peru (Nicandra physalodes).

Nicandra physalodes, a member of the Solanaceae family, is known for its medicinal potential and strong natural insect-repellent properties, which are mainly attributed to its bioactive withanolides and alkaloids. Despite its ecological and pharmacological significance, genomic information for this species has remained limited. Here, we generated a chromosome-level reference genome for N. physalodes based on PacBio high-fidelity (HiFi) long-read sequencing and Hi-C scaffolding. The assembled genome is 933.97 Mb in size, with a contig N50 of 87.37 Mb, and 99.95% (933.54 Mb) of the sequences anchored to 10 pseudochromosomes. Repetitive elements account for 73.06% of the genome, and 27,925 protein-coding genes were predicted, 97.81% of which were functionally annotated. This genomic resource provides a valuable foundation for investigating the genetic basis of specialized metabolite biosynthesis, insect resistance, and environmental adaptation in N. physalodes, as well as for comparative studies within the Solanaceae family.

Genome, Plant

A chromosome-level assembly of the alpine snow alga Chloromonas typhlos.

Chloromonas typhlos is a cosmopolitan alpine snow alga distributed across continents, and its blooming accelerates snow melting by decreasing the amount of snow albedo. To elucidate the genetic traits underlying the adaptation of C. typhlos to the alpine habitat, we combined PacBio sequencing and Hi-C to generate a high-quality chromosome-level genome assembly (contig N50: 1.29 Mb; scaffold N50: 7.23 Mb) with 31 chromosomes and a genome size of 200.86 Mb. Repetitive elements constituted 11.05% of the genome, and 16,133 protein-coding genes were predicted, of which 82% were functionally annotated. This study provides a set of omics resources both for snow algae and the genus Chloromonas.

Snow

Genomic and phenotypic characterization of Klebsiella pneumoniae phage KP Ø1: a novel lytic Slopekvirus targeting uropathogenic multidrug-resistant Klebsiella pneumoniae.

The rise of multidrug-resistant (MDR) uropathogenic gram-negative bacteria (GNB) necessitates the development of alternative therapeutic strategies. This study aimed to isolate, phenotypically characterize, and perform whole-genome sequencing of the bacteriophage demonstrating the broadest host range against MDR uropathogens. Fifty MDR GNB isolates were screened for lytic phages. The most promising candidate, Klebsiella pneumoniae phage KP Ø1, was characterized using plaque assay, Transmission Electron Microscopy (TEM), and pH/thermal stability testing. Genomic characterization was performed via whole-genome sequencing (WGS), with functional annotation and lifestyle prediction using PhaBOX and PhageScope software. Klebsiella pneumoniae was the most prevalent MDR uropathogen. Klebsiella pneumoniae phage KP Ø1 exhibited a 50% host range and high lytic titer (10⁸ PFU/mL). TEM revealed an icosahedral head and short contractile tail. Genomic characterization by WGS revealed that Klebsiella pneumoniae phage KP Ø1 possesses a 174,591 bp double-stranded deoxyribonucleic acid (dsDNA) genome containing 274 predicted open reading frames (ORFs). No lysogeny-related genes, toxins, or antibiotic resistance markers were detected, confirming its strictly lytic nature and supporting its potential as a candidate for phage therapy applications. The phage remained stable (10⁸ PFU/mL) across temperatures of - 20 °C to 50 °C; supporting its suitability for long-term biobanking and suggesting potential activity at physiological temperature, and across a pH range of 7-9. Klebsiella pneumoniae phage KP Ø1 is a novel, obligately lytic Slopekvirus whose genomic architecture, stability profile, and absence of lysogeny-associated, virulence, and antimicrobial resistance genes ( AMR) collectively support its candidacy for further preclinical evaluation as a phage therapy agent against uropathogenic MDR Klebsiella pneumoniae.

Klebsiella pneumoniae

Discussion on the mechanism of Lingguizhugan Decoction in treating hypertension based on network pharmacology and molecular simulation technology.

To explore the mechanism of Lingguizhugan Decoction in treating hypertension based on network pharmacology and molecular simulation. The active ingredients and potential targets were screened by the Systematic Pharmacological Analysis Platform of Traditional Chinese Medicine (TCMSP). Hypertension-related targets were obtained from OMIM and GeneCards databases. Common targets between drug and hypertension were screened in the Venny platform. A protein-protein interaction (PPI) network was constructed in the STRING database using intersection targets. Key targets in PPI network were analyzed by Cytoscape. R language program was used for Gene Ontology (GO) functional annotation and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Finally, the binding abilities of the main active ingredients to critical targets were verified by molecular simulation. Naringenin, quercetin, kaempferol, and β-sitosterol in Lingguizhugan Decoction, and potential targets such as STAT3, AKT1, TNF, IL6, JUN, PTGS2, MMP9, CASP3, TP53, and MAPK3, were screened out. KEGG Enrichment analysis revealed that the common targets of Lingguizhugan Decoction and hypertension are mainly involved in the lipid and atherosclerosis signaling pathway, AGE-RAGE signaling pathway in diabetic complications, fluid shear stress and atherosclerosis, and IL17 signaling pathway. The molecular simulation results showed that naringenin-MAPK3, quercetin-MMP9, quercetin-PTGS2, and quercetin-TP53 were the top four in the docking scores. Naringenin-MAPK3 and quercetin-MMP9 were stable, with binding free energies of -27.97 ± 1.41 kcal/mol and -21.15 ± 3.17 kcal/mol, respectively. The possible mechanism of Lingguizhugan Decoction in treating hypertension is characterized of multi-component, multi-target, and multi-pathway.Communicated by Ramaswamy H. Sarma.

Network Pharmacology

Exploring biosynthetic potential of the endophytic Penicillium turbatum BLH34 using whole-genome sequence analysis and molecular networking.

An in-depth genomic and metabolomic investigation was conducted on the endophytic fungus Penicillium turbatum BLH34, isolated from Macleaya cordata. Hybrid sequencing (Illumina-Nanopore) generated a high-quality 27.9 Mb genome (GC 48.6%) encoding 9798 proteins, with functional annotation linking 5350 genes to the NCBI non-redundant database and 3404 to KEGG pathways. AntiSMASH analysis uncovered 35 biosynthetic gene clusters (BGCs), 23 of which lacked homology to known pathways, highlighting BLH34's potential for novel metabolite discovery. Molecular networking (GNPS) and LC-MS/MS identified 19 specialised metabolites, including antimicrobial polyketides. Bioassays demonstrated potent inhibition against Staphylococcus aureus (36 mm), Bacillus subtilis (28 mm) and Escherichia coli (24 mm), underscoring its pharmaceutical relevance.

Penicillium

Blood-based DNA methylation markers for autism spectrum disorder identification using machine learning.

BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder lacking objective biomarkers for early diagnosis. DNA methylation is a promising epigenetic marker, and machine learning offers a data-driven classification approach. However, few studies have examined whole-blood, genome-wide DNA methylation profiles for ASD diagnosis in school-aged children. METHODS: We analyzed genome-wide DNA methylation data from GEO dataset GSE113967, including 52 children with ASD and 48 typically developing (TD) controls. Differentially methylated positions (DMPs) were identified, and feature selection was performed using support vector machine-recursive feature elimination with cross-validation (SVM-RFECV). Classification models were developed using random forest (RF), extreme gradient boosting (XGBoost), and decision tree (DT) classifiers. A nomogram visualized feature contributions. RESULTS: A total of 138 DMPs differentiated ASD from TD children. Eleven CpG sites selected by SVM-RFECV formed the basis for model construction. RF and XGBoost achieved the highest accuracy (75%), with DT reaching 70%. Functional annotation indicated enrichment in cell adhesion and immune-related pathways. CONCLUSIONS: This exploratory study demonstrates the feasibility of integrating peripheral blood DNA methylation data with machine learning to distinguish children with ASD. While limited by sample size and moderate accuracy, this study provides methodological insights into the feasibility of integrating epigenetic and computational approaches for ASD-related biomarker exploration.

Humans

Changes of DNA methylation and gene expression profile in placental villi and chorioamniotic membranes under preeclampsia.

BACKGROUND: Preeclampsia (PE) is a serious pregnancy complication with elusive pathogenesis. Although epigenetic dysregulation is implicated, its layer-specific placental roles are poorly defined. This study aimed to identify shared and layer-specific epigenetic alterations in PE by profiling DNA methylation and gene expression in placental villi (PV) and chorioamniotic membranes (CAM). RESEARCH DESIGN AND METHODS: PV and CAM samples were collected from 7 normal and 8 PE pregnancies, and three public DNA methylation datasets (GSE98224, GSE44667, GSE75196) were integrated. Differentially methylated genes (DMGs) and differentially expressed genes (DEGs) were identified based on whole-genome methylation and transcriptome sequencing. Layer-specific and shared gene sets were identified by cross-analysis, with functional annotation using Gene Ontology (GO). RESULTS: EM-seq revealed a hypermethylation-dominant, tissue-specific methylation landscape in PE placentas. Cross-tissue comparison identified shared DMGs between the two layers, including nine key genes consistently altered in public datasets. Integrated analysis in PV further identified 22 co-dysregulated genes, enriched in thermoregulation, maternal-fetal immunity, signal transduction, and cell differentiation. CONCLUSIONS: This study elucidates the shared and layer-specific dysregulation of gene networks at methylomic and transcriptomic levels in PE placenta. Comparing PV and CAM highlights placental epigenetic heterogeneity and dysfunction, offering novel clues for mechanistic research and layer-targeted therapies.

Humans

X chromosome-wide association studies for quantitative trait loci based on the mixture of general pedigrees and additional unrelated individuals.

Genome-wide association studies have successfully identified many genetic variants associated with complex traits. However, most existing methods target autosomes rather than X chromosome, and several existing X chromosome-wide association studies (XWAS) at quantitative trait loci (QTL) largely focus on unrelated individuals, with limited attention to general pedigrees or mixture of general pedigrees and additional unrelated individuals (called the mixed data for brevity). In this study, we propose nine novel methods for XWAS at QTL in the mixed data (${\mathrm{MQX}}_{\mathrm{cat}}$, ${\mathrm{MQZ}}_{\mathrm{max}}$, ${\mathrm{MT}}_{\mathrm{plinkw}}$, ${\mathrm{MT}}_{\mathrm{chenw}}$, $\mathrm{MwM}3\mathrm{VNA}$, ${\mathrm{MQMVX}}_{\mathrm{cat}}$, ${\mathrm{MQMVZ}}_{\mathrm{max}}$, $\mathrm{MpMV}$, and $\mathrm{McMV}$), also applicable to general pedigrees alone. The first four methods test for mean differences across genotypes; the latter four test for differences in both means and variances; $\mathrm{MwM}3\mathrm{VNA}$ tests for variance differences only. All mean-based and mean-variance-based methods incorporate X chromosome inactivation information, and all nine methods consider genetic relatedness in pedigrees. Simulation studies confirm well-controlled type I error rates, and inclusion of pedigrees significantly improves statistical power. Note that there has been no study focusing on X chromosome for the mixed data or general pedigrees from UK Biobank database, so we apply our proposed methods to this dataset, which identify five total cholesterol (TC)-associated and 13 low-density lipoprotein cholesterol (LDL-C)-associated single nucleotide polymorphisms (SNPs). Linkage disequilibrium (LD) analysis reveals that these SNPs fall into three distinct LD blocks. Functional annotation and gene ontology enrichment analysis reveal 16 and 28 enriched pathways for TC-associated and LDL-C-associated genes, respectively. These methods provide robust and powerful tools for XWAS at QTL in both mixed data and general pedigrees.

Quantitative Trait Loci

Automatic evaluation of protein sequence functional patterns.

A procedure that automatically provides an evaluation of the diagnostic ability of a protein sequence functional pattern is described. The procedure relies on the identification of the closest definable set in terms of a (protein sequence) database functional annotation to the set of database instances containing a given pattern. Assuming annotation correctness and completeness in the protein sequence database, the degree of statistical association between these sets provides an appropriate measure of the diagnostic ability of the pattern. An experimental implementation of the procedure, using the NBRF/PIR protein database, has been applied to a diverse collection of published sequence patterns. Results obtained reveal that frequently it is not possible to define (in NBRF/PIR database terminology) the set of database instances containing a given pattern, suggesting either lack of pattern diagnostic ability or protein database annotation incompleteness and/or inconsistencies.

Algorithms

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

Fantastic microbes and where to find them: evaluating learning-by-doing outcomes in a crowdfunded metagenomics workshop.

Metagenomics offers a powerful framework for authentic, interdisciplinary learning, yet it remains underrepresented in undergraduate education due to technical and infrastructural barriers. We hypothesized that a research-based, learning-by-doing metagenomics workshop supported by accessible bioinformatics tools could enhance students' perceived skills, self-efficacy, and conceptual understanding of metagenomic analysis. To test this hypothesis, we designed and evaluated a hybrid hands-on workshop in which undergraduate and postgraduate students analyzed real environmental shotgun metagenomic datasets generated from soil samples collected during a citizen science initiative. Using the graphical workflow platform KBase, participants completed an end-to-end metagenomic analysis, from quality control and assembly to genome reconstruction, taxonomic classification, functional annotation, and scientific presentation of results. Educational outcomes were assessed through validated retrospective pre-post questionnaires, self-efficacy scales, and an open-ended conceptual understanding task. Participants showed significant increases in perceived metagenomic skills and confidence in performing metagenomic analyses, while gains in perceived learning showed a positive trend. Conceptual understanding improved across educational levels, particularly among participants with limited prior experience. Together, these findings demonstrate that authentic, data-driven metagenomics activities can effectively lower barriers to computational biology and foster meaningful learning through hands-on research experiences.

Metagenomics