Search PubMedSearch

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A Network Pharmacology and Molecular Docking Study of TongBi Formula for Osteoarthritis.

This study applied network pharmacology combined with molecular docking to predict the potential therapeutic targets and molecular mechanisms of TongBi Formula (TBF) in osteoarthritis (OA). Active components and corresponding targets of TBF were retrieved from the traditional Chinese medicine Systems Pharmacology Database and Analysis Platform, while OA-related targets were collected from Online Mendelian Inheritance in Man, GeneCards, DrugBank, and Therapeutic Target Database. A network visualization and analysis software was used to construct compound-target and protein-protein interaction (PPI) networks. Gene Ontology functional annotation and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed using the Database for Annotation, Visualization and Integrated Discovery platform. Molecular docking analysis was conducted using a molecular docking software to evaluate the predicted binding affinity between key active compounds and core target proteins. A total of 47 overlapping targets between TBF and OA were identified. PPI network analysis highlighted JUN, RELA, IL6, MAPK1, and IL10 as potential hub targets. Enrichment analysis suggested that TBF may regulate inflammation, lipid metabolism, and multiple intracellular signaling pathways associated with OA progression. Molecular docking results demonstrated favorable predicted binding affinities between core active compounds and key OA-related protein targets. These findings provide a computational framework for understanding the potential mechanisms of TBF against OA and support further experimental validation.

Molecular Docking Simulation

A comparative genomic analysis of left- and right-sided colon cancer using real-world data from the AACR project GENIE BPC dataset.

Left- and Right-sided colon cancers (LCC and RCC) are increasingly recognized as distinct clinicopathological and molecular subtypes with divergent prognoses and therapeutic responses. Leveraging a large, multi-institutional cohort from the AACR Project Genomics Evidence Neoplasia Information Exchange (GENIE) Biopharma Collaborative (BPC) (n = 750; LCC: 363 vs. RCC: 387), we conducted a comprehensive analysis of mutational profiles, tumor mutation burden (TMB), and survival outcomes. Our findings revealed a markedly higher TMB in RCC compared to LCC (6.65 &#xb1; 11.3 vs. 3.17 &#xb1; 4.35; adjusted P = 3.12&#xd7;10-32), suggesting greater genomic instability in RCC. After applying functional annotation filters (PolyPhen > 0.85, SIFT < 0.05), RCC tumors were significantly enriched for mutations in BRAF (23.1% vs. 6.7%), KMT2D (8.6% vs. 3.2%), and SMAD4 (13.1% vs. 7.3%), while TP53 mutations predominated in LCC (40.6% vs. 31.8%). Multivariate Cox regression analysis identified RCC as an independent predictor of poorer overall survival (OS) relative to LCC (HR: 1.30, 95% CI: 1.02-1.66, P = 0.033). Notably, KRAS mutations were associated with significantly worse OS in LCC (HR: 1.68, 95% CI: 1.06-2.70, P = 0.027), while BRAF mutations predicted adverse outcomes in RCC (HR: 1.58, 95% CI: 1.05-2.37, P = 0.028). These results underscore the prognostic value of tumor sidedness and specific genetic alterations in colon adenocarcinoma. Our study highlights the need for sidedness-specific molecular profiling to inform precision oncology strategies in colon cancer management.

BRAF

Multi-Ancestry Survival GWAS of Substance Use Initiation in the ABCD Study.

BACKGROUND: Substance use initiation in adolescence is influenced by both genetic and environmental factors; however, large-scale genetic studies often treat initiation as a binary outcome and underuse longitudinal timing information. METHODS: We conducted time-to-event (survival) genome-wide association analyses (GWAS) of initiation for four outcomes-alcohol, nicotine, cannabis, and any substance use-using longitudinal follow-up data from the Adolescent Brain Cognitive Development (ABCD) Study. We performed ancestry-stratified GWAS within European (EUR), African (AFR), and Hispanic (HISP) groups, applying consistent quality control and covariate adjustment. Summary statistics were harmonized across ancestries and meta-analyzed using inverse-variance weighted fixed-effects and DerSimonian-Laird random-effects models. We evaluated genomic inflation and heterogeneity (Cochran's Q and I 2), identified independent lead variants at genome-wide and suggestive significance thresholds, and assessed cross-trait overlap of associated loci. RESULTS: In the multi-ancestry meta-analysis, we observed suggestive association signals across traits (minimum p-values: alcohol ~ 1 &#xd7; 10-7, any ~ 1 &#xd7; 10-7, cannabis ~ 5 &#xd7; 10-8, nicotine ~ 1 &#xd7; 10-8). Nicotine initiation showed one genome-wide significant variant in both fixed- and random-effects meta-analyses (p < 5 &#xd7; 10-8). Across traits, suggestive loci demonstrated limited overlap, with the strongest concordance between alcohol and any substance use, consistent with shared liability. Heterogeneity statistics indicated that some loci exhibited cross-ancestry variation in effect estimates. CONCLUSIONS: Survival GWAS leveraging initiation timing can identify genetic signals that may be missed by binary designs and enables principled multi-ancestry synthesis. Our results highlight both shared and trait-specific genetic contributions to early substance initiation and provide a foundation for downstream functional annotation and integrative modeling with environmental risk factors. These findings demonstrate the value of incorporating developmental timing into genetic discovery and provide a framework for integrating longitudinal risk modeling with genomic analyses.

ABCD

GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.

BACKGROUND: The function of a novel gene product is typically predicted by transitive assignment of annotation from similar sequences. We describe a novel method, GOtcha, for predicting gene product function by annotation with Gene Ontology (GO) terms. GOtcha predicts GO term associations with term-specific probability (P-score) measures of confidence. Term-specific probabilities are a novel feature of GOtcha and allow the identification of conflicts or uncertainty in annotation. RESULTS: The GOtcha method was applied to the recently sequenced genome for Plasmodium falciparum and six other genomes. GOtcha was compared quantitatively for retrieval of assigned GO terms against direct transitive assignment from the highest scoring annotated BLAST search hit (TOPBLAST). GOtcha exploits information deep into the 'twilight zone' of similarity search matches, making use of much information that is otherwise discarded by more simplistic approaches. At a P-score cutoff of 50%, GOtcha provided 60% better recovery of annotation terms and 20% higher selectivity than annotation with TOPBLAST at an E-value cutoff of 10(-4). CONCLUSIONS: The GOtcha method is a useful tool for genome annotators. It has identified both errors and omissions in the original Plasmodium falciparum annotation and is being adopted by many other genome sequencing projects.

Animals

The Annotated Blueprint: Integrated Functional Genomic Resources for a model Tetraploid Wheat Triticum turgidum cv. Kronos.

Triticum turgidum cv. Kronos is a tetraploid wheat cultivar that underpins one of the richest community platforms for functional genomics. Over the past decade, about 3,000 exome- and promoter-capture datasets, linked to mutagenized seed stocks, and transcriptomic and phenotypic resources have accumulated, yet the absence of a reference genome has constrained their impact. Here, we present a chromosome-scale reference genome of Kronos with high-confidence annotations, including manual curation of over 1,000 disease resistance (NLR) genes. This reference revealed previously hidden NLR diversity and clarified their genomic organization at chromosomal ends. Re-analysis of exome- and promoter-capture datasets enabled high-resolution mutation discovery in genes and regulatory regions that were previously inaccessible, uncovering the full standing variation present in Kronos mutant lines. We further re-curated transcriptomic and small RNA datasets, generating improved, genome-wide maps of microRNAs and phasiRNAs important for wheat development. Collectively, these resources elevate Kronos to reference quality and establish it as a versatile platform for functional and translational wheat research.

Journal Article

Molecular Characterization of Listeria monocytogenes Isolated from Retail Yak Meat in Nyingchi, Xizang, China.

Listeria monocytogenes is a Gram-positive zoonotic pathogen responsible for listeriosis, a severe foodborne disease with high mortality in humans and animals. This study aimed to investigate the molecular epidemiology and genomic characteristics of L. monocytogenes isolated from raw yak meat in Nyingchi, Xizang, China. A total of 231 yak-related samples were collected in Nyingchi, consisting of 214 retail raw yak meat samples, 14 farm environmental samples, and 3 nearby water source samples. L. monocytogenes isolates were identified and characterized using culture-based methods, PCR serotyping, and whole-genome sequencing (WGS). Bioinformatic analyses were performed for virulence, antimicrobial resistance, and functional gene annotation using KEGG and COG databases. The overall contamination rate of Lm was 13.08% (28/214) for retail raw yak meat samples, whereas no isolates were recovered from 14 farm environmental samples (0.00%, 0/14) and 3 nearby water source samples (0.00%, 0/3). The serotypes of isolates were 1/2a (9/28, 32.14%), 1/2b (7/28, 25.00%), and 1/2c (12/28, 42.86%). These 28 isolates exhibited varied antimicrobial resistance profiles, with universal resistance to trimethoprim-sulfamethoxazole, high resistance to erythromycin and clindamycin, and low resistance to vancomycin. MLST analysis revealed seven sequence types (STs): ST9 (12/28, 42.86%), ST619 (6/28, 21.43%), ST8 (6/28, 21.43%), ST7 (1/28, 3.57%), ST87 (1/28, 3.57%), ST121 (1/28, 3.57%), ST297 (1/28, 3.57%). ST619 isolates harbored multiple virulence genes, including those located on Listeria pathogenicity islands LIPI-1, LIPI-3, and LIPI-4, indicating high genomic potential for virulence. Representative isolate Y2 (ST619) possessed a 3,009,858 bp genome with 3036 coding genes, four genomic islands, and two prophages. Functional annotation revealed enrichment of genes involved in carbohydrate transport and metabolism and amino acid biosynthesis pathways. Our findings provide the first genomic insight into L. monocytogenes contamination in yak meat from Nyingchi, Xizang, China, highlighting the urgent need to strengthen food safety monitoring and hygiene management in this region.

Listeria monocytogenes

Mycoplasma genes: a case for reflective annotation.

Although function can be assigned to genome sequence by homology at a macroscopic level, this can be misleading in the absence of data on enzyme activities. Together, such data can reveal whether open reading frames are expressed, identify multienzyme function and point to 'orphan' function. Because of their small size and small genomes, the genome sequences of some Mycoplasma spp. are very amenable to detailed analyses.

Genes, Bacterial

Large language models improve annotation of prokaryotic viral proteins.

Viral genomes are poorly annotated in metagenomic samples, representing an obstacle to understanding viral diversity and function. Current annotation approaches rely on alignment-based sequence homology methods, which are limited by the paucity of characterized viral proteins and divergence among viral sequences. Here we show that protein language models can capture prokaryotic viral protein function, enabling new portions of viral sequence space to be assigned biologically meaningful labels. When applied to global ocean virome data, our classifier expanded the annotated fraction of viral protein families by 29%. Among previously unannotated sequences, we highlight the identification of an integrase defining a mobile element in marine picocyanobacteria and a capsid protein that anchors globally widespread viral elements. Furthermore, improved high-level functional annotation provides a means to characterize similarities in genomic organization among diverse viral sequences. Protein language models thus enhance remote homology detection of viral proteins, serving as a useful complement to existing approaches.

Viral Proteins

Predicting function: from genes to genomes and back.

Predicting function from sequence using computational tools is a highly complicated procedure that is generally done for each gene individually. This review focuses on the added value that is provided by completely sequenced genomes in function prediction. Various levels of sequence annotation and function prediction are discussed, ranging from genomic sequence to that of complex cellular processes. Protein function is currently best described in the context of molecular interactions. In the near future it will be possible to predict protein function in the context of higher order processes such as the regulation of gene expression, metabolic pathways and signalling cascades. The analysis of such higher levels of function description uses, besides the information from completely sequenced genomes, also the additional information from proteomics and expression data. The final goal will be to elucidate the mapping between genotype and phenotype.

Bacterial Proteins

EUCLID: automatic classification of proteins in functional classes by their database annotations.

UNLABELLED: A tool is described for the automatic classification of sequences in functional classes using their database annotations. The Euclid system is based on a simple learning procedure from examples provided by human experts. AVAILABILITY: Euclid is freely available for academics at http://www.gredos.cnb.uam.es/EUCLID, with the corresponding dictionaries for the generation of three, eight and 14 functional classes. CONTACT: E-mail: valencia@cnb.uam.es SUPPLEMENTARY INFORMATION: The results of the EUCLID classification of different genomes are available at http://www.sander.ebi.ac. uk/genequiz/. A detailed description of the different applications mentioned in the text is available at http://www.gredos.cnb.uam. es/EUCLID/Full_Paper

Computational Biology

The comparative metabolism of the mollicutes (Mycoplasmas): the utility for taxonomic classification and the relationship of putative gene annotation and phylogeny to enzymatic function in the smallest free-living cells.

Mollicutes or mycoplasmas are a class of wall-less bacteria descended from low G + C% Gram-positive bacteria. Some are exceedingly small, about 0.2 micron in diameter, and are examples of the smallest free-living cells known. Their genomes are equally small; the smallest in Mycoplasma genitalium is sequenced and is 0.58 mb with 475 ORFs, compared with 4.639 mb and 4288 ORFs for Escherichia coli. Because of their size and apparently limited metabolic potential, Mollicutes are models for describing the minimal metabolism necessary to sustain independent life. Mollicutes have no cytochromes or the TCA cycle except for malate dehydrogenase activity. Some uniquely require cholesterol for growth, some require urea and some are anaerobic. They fix CO2 in anaplerotic or replenishing reactions. Some require pyrophosphate not ATP as an energy source for reactions, including the rate-limiting step of glycolysis: 6-phosphofructokinase. They scavenge for nucleic acid precursors and apparently do not synthesize pyrimidines or purines de novo. Some genera uniquely lack dUTPase activity and some species also lack uracil-DNA glycosylase. The absence of the latter two reactions that limit the incorporation of uracil or remove it from DNA may be related to the marked mutability of the Mollicutes and their tachytelic or rapid evolution. Approximately 150 cytoplasmic activities have been identified in these organisms, 225 to 250 are presumed to be present. About 100 of the core reactions are graphically linked in a metabolic map, including glycolysis, pentose phosphate pathway, arginine dihydrolase pathway, transamination, and purine, pyrimidine, and lipid metabolism. Reaction sequences or loci of particular importance are also described: phosphofructokinases, NADH oxidase, thioredoxin complex, deoxyribose-5-phosphate aldolase, and lactate, malate, and glutamate dehydrogenases. Enzymatic activities of the Mollicutes are grouped according to metabolic similarities that are taxonomically discriminating. The arrangements attempt to follow phylogenetic relationships. The relationships of putative gene assignments and enzymatic function in My. genitalium, My. pneumoniae, and My. capricolum subsp. capricolum are specially analyzed. The data are arranged in four tables. One associates gene annotations with congruent reports of the enzymatic activity in these same Mollicutes, and hence confirms the annotations. Another associates putative annotations with reports of the enzyme activity but from different Mollicutes. A third identifies the discrepancies represented by those enzymatic activities found in Mollicutes with sequenced genomes but without any similarly annotated ORF. This suggests that the gene sequence is significantly different from those already deposited in the databanks and putatively annotated with the same function. Another comparison lists those enzymatic activities that are both undetected in Mollicutes and not associated with any ORF. Evidence is presented supporting the theory that there are relatively small gene sequences that code for functional centers of multiple enzymatic activity. This property is seemingly advantageous for an organism with a small genome and perhaps under some coding restraint. The data suggest that a concept of "remnant" or "useless genes" or "useless enzymes" should be considered when examining the relationship of gene annotation and enzymatic function. It also suggests that genes in addition to representing what cells are doing or what they may do, may also identify what they once might have done and may never do again.

Adenosine Triphosphate

Comprehensive profiling of antibiotic resistance genes and functional clusters of orthologous groups annotation of gut microbiota in Indonesian Kedu chickens.

Antibiotic resistance is a growing global health concern, with poultry systems acting as important reservoirs of antibiotic resistance genes (ARGs). However, resistome and functional profiles of indigenous chickens raised under traditional systems remain underexplored. This study aimed to characterize the antibiotic resistome, virulence factor genes, and metabolic potential of gut microbiota in Indonesian Kedu chickens using a shotgun metagenomic approach. Digesta samples from five gastrointestinal segments of 21 healthy adult chickens were analyzed through high-throughput sequencing. ARGs were identified using the Comprehensive Antibiotic Resistance Database (CARD) and Antibiotic Resistance Genes Databases (ARDB), while virulence factors and functional genes were annotated using Virulence Factor Database (VFDB), Clusters of Orthologous Groups (COG), and Carbohydrate-Active EnZymes (CAZy) databases. Results revealed a diverse resistome dominated by multidrug resistance and efflux pump mechanisms, with prominent genes associated with fluoroquinolone, tetracycline, &#x3b2;-lactam, and glycopeptide resistance. The detection of clinically relevant ARGs suggests that genetic determinants associated with antimicrobial resistance are present in the gut microbiota of traditionally raised Kedu chickens, although metagenomic data alone cannot determine whether these genes are actively expressed or confer phenotypic resistance. Virulence factor analysis showed functions related to adherence, immune evasion, iron acquisition, quorum sensing, and efflux activity, reflecting strong microbial adaptability. Functional profiling demonstrated enrichment in translation, carbohydrate and amino acid metabolism, genome maintenance, and cell envelope biogenesis. Additionally, CAZyme analysis indicated a high capacity for complex polysaccharide degradation, supporting efficient utilization of fiber-rich traditional diets. In conclusion, this study provides a comprehensive metagenomic overview of antibiotic resistance and functional potential in Kedu chicken gut microbiota, emphasizing the importance of incorporating indigenous poultry into antimicrobial resistance surveillance within a One Health framework.

Antibiotic resistance genes

Microbial diversity and metabolic pathways linked to benzene degradation in petrochemical-polluted groundwater.

The rapid advance in shotgun metagenome sequencing has enabled us to identify uncultivated functional microorganisms in polluted environments. While aerobic petrochemical-degrading pathways have been extensively studied, the anaerobic mechanisms remain less explored. Here, we conducted a study at a petrochemical-polluted groundwater site in Henan Province, Central China. A total of twelve groundwater monitoring wells were installed to collect groundwater samples. Benzene appeared to be the predominant pollutant, detected in 10 out of 12 samples, with concentrations ranging from 1.4&#xa0;&#x3bc;g/L to 5,280&#xa0;&#x3bc;g/L. Due to the low aquifer permeability, pollutant migration occurred slowly, resulting in relatively low benzene concentrations downstream within the heavily polluted area. Deep metagenome sequencing revealed Proteobacteria as the dominant phylum, accounting for over 63&#xa0;% of total abundances. Microbial &#x3b1;-diversity was low in heavily polluted samples, with community compositions substantially differing from those in lightly polluted samples. dmpK encoding the phenol/toluene 2-monooxygenase was detected across all samples, while the dioxygenase bedC1 was not detected, suggesting that aerobic benzene degradation might occur through monooxygenation. Sequence assembly and binning yielded 350 high-quality metagenome-assembled genomes (MAGs), with 30 MAGs harboring functional genes associated with aerobic or anaerobic benzene degradation. About 80&#xa0;% of MAGs harboring functional genes associated with anaerobic benzene degradation remained taxonomically unclassified at the genus level, suggesting that our current database coverage of anaerobic benzene-degrading microorganisms is very limited. Furthermore, two genes integral to anaerobic benzene metabolism, i.e, benzoyl-CoA reductase (bamB) and glutaryl-CoA dehydrogenase (acd), were not annotated by metagenome functional analyses but were identified within the MAGs, signifying the importance of integrating both contig-based and MAG-based approaches. Together, our efforts of functional annotation and metagenome binning generate a robust blueprint of microbial functional potentials in petrochemical-polluted groundwater, which is crucial for designing proficient bioremediation strategies.

Groundwater

PhyloNaP: a user-friendly database of phylogeny for natural product-producing enzymes.

SUMMARY: Phylogenetic analysis is widely used to predict enzyme function, yet building annotated and reusable trees is labor-intensive and requires extensive knowledge about the specific enzymes. Existing resources rarely cover biosynthetic enzymes and lack the context needed for meaningful analysis. We present PhyloNaP, the first large-scale resource dedicated to phylogenies of biosynthetic enzymes. PhyloNaP provides &#x223c;51&#x2009;000 annotated and interactive trees enriched with chemical, functional, and taxonomic information. Users can classify their own sequences via phylogenetic placement, enabling functional inference in an evolutionary context. A contribution portal allows the community to submit curated trees. By combining scale, breadth of annotation, and interactive functionality, PhyloNaP fills a major gap in bioinformatics resources for enzyme discovery and annotation, with immediate applications to secondary metabolism and beyond. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at&#xa0;https://phylonap.cs.uni-tuebingen.de.

Phylogeny

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural