Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Huntingtin interacts with the receptor sorting family protein GASP2.

Protein interaction networks are useful resources for the functional annotation of proteins. Recently, we have generated a highly connected protein-protein interaction network for Huntington's disease (HD) by automated yeast two-hybrid (Y2H) screening (Goehler et al., 2004). The network included several novel direct interaction partners for the disease protein huntingtin (htt). Some of these interactions, however, have not been validated by independent methods. Here we describe the verification of the interaction between htt and GASP2 (G protein-coupled receptor associated sorting protein 2), a protein involved in membrane receptor degradation. Using membrane-based and classical coimmunoprecipitation assays we demonstrate that htt and GASP2 form a complex in cotransfected mammalian cells. Moreover, we show that the two proteins colocalize in SH-SY5Y cells, raising the possibility that htt and GASP2 interact in neurons. As the GASP protein family plays a role in G protein-coupled receptor sorting, our data suggest that htt might influence receptor trafficking via the interaction with GASP2.

Animals↗

Activity-based proteomics: enzymatic activity profiling in complex proteomes.

In the postgenomic era new technologies are emerging for global analysis of protein function. The introduction of active site-directed chemical probes for enzymatic activity profiling in complex mixtures, known as activity-based proteomics has greatly accelerated functional annotation of proteins. Here we review probe design for different enzyme classes including serine hydrolases, cysteine proteases, tyrosine phosphatases, glycosidases, and others. These probes are usually detected by their fluorescent, radioactive or affinity tags and their protein targets are analyzed using established proteomics techniques. Recent developments, such as the design of probes for in vivo analysis of proteomes, as well as microarray technologies for higher throughput screenings of protein specificity and the application of activity-based probes for drug screening are highlighted. We focus on biological applications of activity-based probes for target and inhibitor discovery and discuss challenges for future development of this field.

Animals↗

Mining functional information associated with expression arrays.

Deciphering the networks of interactions between molecules in biological systems has gained momentum with the monitoring of gene expression patterns at the genomic scale. Expression array experiments provide vast amounts of experimental data about these networks, the analysis of which requires new computational methods. In particular, issues related to the extraction of biological information are key for the end users. We propose here a strategy, implemented in a system called GEISHA (gene expression information system for human analysis) and able to detect biological terms significantly associated to different gene expression clusters by mining collections of Medline abstracts. GEISHA is based on a comparison of the frequency of abstracts linked to different gene clusters and containing a given term. Interpretation by the end user of the biological meaning of the terms is facilitated by embedding them in the corresponding significant sentences and abstracts and by establishing relations with other, equally significant terms. The information provided by GEISHA for the available yeast expression data compares favorably with the functional annotations provided by human experts, demonstrating the potential value of GEISHA as an assistant for the analysis of expression array experiments.

Animals↗

Genomic and polyphasic characterization of six novel Hymenobacter species isolated from soil in Korea.

Six bacterial strains (BT523T, BT559T, 15J16-1T3BT, BT730T, DG25AT, and DG25BT) were isolated from soil samples in Korea and assigned to the family Hymenobacteraceae (order Cytophagales, class Cytophagia). Phylogenetic analysis based on 16S rRNA gene sequences showed that the strains formed distinct lineages within the genus Hymenobacter. Strains BT523T, BT559T, and 15J16-1T3BT exhibited highest sequence similarities to Hymenobacter armeniacus BT189T (97.6%), Hymenobacter polaris RP-2-7 T (97.8%), and Hymenobacter paludis KBP-30 T (98.3%), respectively, while strains BT730T, DG25AT, and DG25BT were most closely related to Hymenobacter tibetensis XTM003T, with similarities of 96.3-96.6%. All strains were Gram-negative, aerobic, rod-shaped, and formed red to pink pigmented colonies. Whole-genome analysis revealed genome sizes ranging from 3.78 to 6.33 Mb with G + C contents of 55.5-65.0%. Average nucleotide identity (ANI) and digital DNA-DNA hybridization (dDDH) values between the strains and their closest relatives were below the accepted thresholds for species delineation, supporting their classification as novel species. Functional annotation indicated the presence of genes associated with core metabolism, stress response, and pigment biosynthesis, reflecting adaptation to soil environments. Secondary metabolite analysis further revealed the presence of biosynthetic gene clusters, including terpene and siderophore pathways. Based on polyphasic taxonomic evidence, the six strains are proposed to represent six novel species of the genus Hymenobacter, for which the names Hymenobacter miniatus sp. nov., Hymenobacter madidus sp. nov., Hymenobacter convexus sp. nov., Hymenobacter rubellus sp. nov., Hymenobacter erythromyxa sp. nov., and Hymenobacter radioresistens sp. nov. are proposed. The type strains are BT523T (= KCTC 72341 T = NBRC 114851 T), BT559T (= KACC 21821 T = NBRC 114852 T), 15J16-1T3BT (= KCTC 42995 T = NBRC 112818 T), BT730T (= KACC 22457 T = NBRC 116482 T), DG25AT (= KCTC 32451 T = TBRC 19856 T), and DG25BT (= KCTC 32452 T = JCM 19444 T).

Soil Microbiology↗

Gene-dosage effect on chromosome 21 transcriptome in trisomy 21: implication in Down syndrome cognitive disorders.

In the era of human functional genomics, the chromosome 21 has represented a prototype for pioneering global biotechnologies. Its relatively low gene content enabled studying Down syndrome at the chromosomal scale, for which the last years have seen intense research activity aiming at genotype-phenotype correlations. The global gene-dose dependent upregulation of gene expression seen in the context of trisomy and preliminary functional annotation of chromosome 21 genes points towards candidate genes and molecular pathways potentially associated with the cognitive defects observed in Down syndrome.

Chromosome Mapping↗

NMR solution structure of Thermotoga maritima protein TM1509 reveals a Zn-metalloprotease-like tertiary structure.

The 150-residue protein TM1509 is encoded in gene YF09_THEMA of Thermotoga maritima. TM1509 has so far no functional annotation and belongs to protein family UPF0054 (PFAM accession number: PF02130) which contains at least 146 members. The NMR structure of TM1509 reveals an alpha+beta fold comprising a four stranded beta-sheet with topology A( upward arrow), B( upward arrow), D( upward arrow), C( downward arrow) as well as five alpha-helices I-V. The structures of most members of family PF02130 can be reliably constructed using the TM1509 NMR structure, demonstrating high leverage for exploration of fold space. A multiple sequence alignment of TM1509 with homologues of family UPF0054 shows that three polypeptide segments, as well as a putative zinc-binding consensus motif HGXLHLXGYDH located at the C-terminal end of alpha-helix IV, are highly conserved. The spatial arrangement of the three His residues of this UPF0054 consensus motif is similar to the arrangement found for the His residues in the HEXXHXXGXXH zinc-binding consensus motif of matrix metallo-proteases (MMPs). Moreover, the other conserved polypeptide segments form a large cavity which encloses the putative Zn-binding pocket and might confer specificity during catalysis. However, TM1509 and the other members of the UPF0054 family do not have the crucial Glu residue in position 2 of the MMP consensus motif. Intriguingly, the TM1509 structure indicates that the Asp in the UPF0054 consensus motif (Asp 111 in TM1509) may overtake the catalytic role of the Glu. This suggests that protein family UPF0054 might contain members of a hitherto uncharacterized class of metalloproteases.

Amino Acid Sequence↗

Whole-genome analysis of Brevibacterium sanguinis AZMABM HM27: a bacterial isolate from the sea anemone Radianthus magnifica and exhibiting promising multi-therapeutic properties.

BACKGROUND: The marine anemone Radianthus magnifica harbors symbiotic microbes with promising biomedical potential, yet their diversity and therapeutic properties remain underexplored. This study aimed to characterize a symbiotic bacterium isolated from R. magnifica collected from Samalona Island, Indonesia, and to evaluate its multi-therapeutic potential. METHODS: Strain AZMABM HM27 was characterized using whole-genome sequencing, functional annotation, biosynthetic gene cluster prediction, molecular docking, and in vitro bioactivity assays. RESULTS: Phylogenetic and genome-based analyses confirmed AZMABM HM27 as Brevibacterium sanguinis, with an OrthoANI value of 97.37% and a dDDH value of 76.50% against the type strain. The genome comprises a 3,834,082 bp chromosome encoding 3,362 protein-coding genes, including 95 genes involved in secondary metabolite biosynthesis. Five biosynthetic gene clusters were predicted, including those associated with ectoine, terpene, and siderophore production. The crude extract demonstrated antioxidant activity (IC₅₀ = 0.87 mg/mL), anti-inflammatory activity (up to 60% inhibition), antidiabetic activity through α-glucosidase inhibition (up to 40% inhibition), and dose-dependent antiproliferative activity against MCF-7 breast cancer cells (74.10% viability at 1 mg/mL). Molecular docking identified a lead compound, 8,9,9,10,10,11-hexafluoro-4,4-dimethyl-3,5-dioxatetracyclo [5.4.1.0(2,6)0.0(8,11)] dodecane, with strong binding affinities to selected therapeutic targets. CONCLUSIONS: B. sanguinis AZMABM HM27 represents a marine symbiotic strain associated with R. magnifica and a promising source of bioactive compounds with antioxidant, anti-inflammatory, antidiabetic, and antiproliferative potential. Further purification, structural elucidation, and in vivo studies are warranted to validate its therapeutic potential.

Animals↗

Endosperm-preferred expression of maize genes as revealed by transcriptome-wide analysis of expressed sequence tags.

The transcriptome-wide endosperm-preferred expression of maize genes was addressed by analyzing a large database of expressed sequence tags (ESTs). We generated 30,531 high quality sequence-reads from the 5'-ends of cDNA libraries from maize endosperm harvested at 10, 15, and 20 days after pollination. A further 196,900 maize sequence-reads retrieved from public databases were added to this endosperm collection to generate MAIZEST, a database with tools for data storage and analysis. MAIZEST contains 227,431 ESTs, one third of which represents developing endosperm and the remaining two-thirds represent transcripts from 49 cDNA libraries constructed from different organs and tissues. Assembling the MAIZEST ESTs generated 29,206 putative transcripts, of which a set of 4032 assembled sequences was composed exclusively of sequences derived from endosperm cDNA libraries. After sequence analysis using overlapping parameters, a sub-set of 2403 assembled sequences was functionally annotated and revealed a wide variety of putative new genes involved in endosperm development and metabolism.

Expressed Sequence Tags↗

Integrated Genome Mining and Bioactivity-Guided Isolation of Antimicrobial Peptides from Bacillus amyloliquefaciens BS4.

Bacterial resistance remains a critical global health challenge, driving the continuous search for novel antimicrobial agents. Bacillus amyloliquefaciens is a recognized repository of bioactive metabolites; however, its full biosynthetic potential requires integrated genomic and experimental validation. This study characterized the antimicrobial profile of B. amyloliquefaciens BS4 through a hybrid pipeline. Genome sequencing and de novo assembly revealed a 3.9 Mb chromosome with a G + C content of 46.14%. Functional annotation identified 3,887 coding sequences, including pathways for siderophore biosynthesis and a complete bacilysin biosynthetic cluster. BGC analysis using antiSMASH v7.1.0 and BAGEL4 identified 18 biosynthetic gene clusters, while similarity network analysis via BiG-SCAPE highlighted unique singleton BGCs, indicating untapped biosynthetic diversity. Although in silico screening via Macrel predicted two putative cationic antimicrobial peptides (AMPs), bioactivity-guided purification utilizing sequential RP-HPLC, and de novo sequencing revealed a distinct set of four active peptides. Notably, three of these sequences were identified as fragments derived from the BclA exosporium protein family, highlighting the structural proteome as a non-canonical source of antimicrobials. The purified fractions exhibited activity against M. luteus and E. coli, while displaying no significant hemolytic activity or cytotoxicity, even above the MIC values. Molecular docking further supported the interaction of these candidates with bacterial targets. Overall, this hybrid strategy effectively uncovers the antimicrobial complexity of BS4, revealing 'cryptic' peptide candidates with therapeutic potential.

Bacillus amyloliquefaciens BS4↗

Genomic diversity, inbreeding, and selection signatures in duroc, landrace, and yorkshire pigs from a long-term closed breeding system.

Duroc (DD), Landrace (LL), and Yorkshire (YY) are among the most widely used commercial pig breeds, having undergone intense long-term selection within closed breeding systems. This study presents a comprehensive genomic analysis of genetic diversity, inbreeding patterns, and selection signatures in DD, LL, and YY populations that have been subject to close breeding for over 15 years. Genomic and pedigree data were available for 1,088 animals (DD = 348, LL = 276, YY = 464), genotyped using the GenoBaits® Porcine 100 K SNP panel. Principal component analysis and genetic diversity metrics revealed distinct population structures among the three breeds. Pairwise genetic differentiation supported this pattern, with DD showing the greatest divergence from LL (0.34 ± 0.24) and YY (0.33 ± 0.24), while LL and YY were more closely related (FST = 0.22 ± 0.19). Linkage disequilibrium (LD) analysis further confirmed these differences, as DD exhibited the highest average r² (0.34), followed by LL (0.28) and YY (0.25). Within-breed genetic diversity metrics, including observed heterozygosity (HO: 0.37 in DD, 0.39 in LL, 0.38 in YY), expected heterozygosity (HE: 0.36 in DD, 0.37 in LL, 0.38 in YY), and minor allele frequency (MAF: 0.27 in DD, 0.28 in LL, 0.29 in YY), indicated greater genetic variability in LL and YY compared to DD. Runs of homozygosity (ROH) analyses revealed different patterns of autozygosity, with DD exhibiting more long ROH indicative of recent inbreeding, while YY harbored a higher number of short ROH, suggestive of more ancient demographic events. ROH-based inbreeding coefficients (FROH) consistently exceeded pedigree-based estimates (FPED) across all breeds, highlighting the presence of recent or unrecorded inbreeding that pedigree data may not fully capture. According to Generation Proxy Selection Mapping (GPSM), 17, 1, and 12 significant SNPs were detected in DD, LL, and YY, respectively. Functional annotation of ROH islands and GPSM-significant loci revealed both breed-specific and overlapping QTLs related to traits such as growth, reproduction, and carcass. In general, the findings of this study contribute to a deeper understanding of the genomic consequences of long-term closed breeding and provide reference information to support consideration of breeding strategies that balance continued selection for productivity with the maintenance of genetic diversity in modern commercial pig populations.

Animals↗

Mining metagenomes from extremophiles as a resource for novel glycoside hydrolases for industrial applications.

The exploration of metagenomes from extremophiles has emerged as a promising approach for discovering novel glycoside hydrolases (GHs) with potential industrial applications. Extremophiles, which thrive in harsh conditions such as high salinity, extreme temperatures, and acidic or alkaline environments, produce enzymes naturally adapted to function under these conditions. This unique adaptability makes them highly desirable for industrial processes requiring robust and efficient biocatalysts. These biocatalysts reduce reliance on harsh chemicals and energy-intensive processes, contributing to greener industrial operations. This review underscores the power of metagenomics in bypassing the need to culture large libraries of extremophiles in the lab. High-throughput sequencing and bioinformatics enable the identification of novel GH-encoding genes directly from environmental DNA. While metagenomic mining has yielded promising results, challenges such as the expression of extremophile-derived genes in mesophilic hosts, low activity yields, and scalability remain. Advances in synthetic biology and protein engineering could address these bottlenecks, enabling more efficient utilization of GHs. Additionally, integrating machine learning for predictive functional annotation may accelerate the identification of high-value candidates.

Glycoside Hydrolases↗

Novel HLA class I and II insights into the pathogenesis of systemic sclerosis-associated interstitial lung disease.

OBJECTIVES: Systemic sclerosis-associated interstitial lung disease (SSc-ILD) is the leading cause of mortality in systemic sclerosis (SSc), yet its genetic architecture remains incompletely understood. Therefore, given the key role of the major histocompatibility complex (MHC) in SSc, we aimed to perform a comprehensive MHC-wide association study in the largest SSc-ILD cohort to date. METHODS: We analysed 2412 patients with SSc-ILD⁺, 3550 patients with SSc-ILD⁻, and 15,076 controls of European ancestry from 10 international cohorts. After quality control, the MHC region was imputed, and inverse variance weighted meta-analysis was performed. Subsequently, conditional stepwise analyses, adjustment for antitopoisomerase autoantibody (ATA) status, and functional annotation of significant single-nucleotide polymorphisms were performed. Finally, we constructed a composite score combining genetic, clinical, and demographic variables to predict SSc-ILD. RESULTS: After conditional analysis, we detected 12 significant associations within class I and class II human leukocyte antigen (HLA) genes. ATA adjustment reduced the significance of class II HLA variants, whereas class I HLA variants remained unaffected. Finally, the built composite score had an area under the curve of 0.754, significantly outperforming the models including any of the variables alone. CONCLUSIONS: In this study, we identify genetic mechanisms underlying SSc-ILD that support the potential implication of CD8+ T cells and ATAs in its pathogenesis. Moreover, we also demonstrate the enhanced efficacy of integrating genetic information into predictive models to detect patients at high risk of SSc-ILD. These findings provide new insights into disease pathogenesis and suggest potential biomarkers and therapeutic targets for improved patient management.

Humans↗

Shared genetic architecture and therapeutic targets across paediatric immune-mediated diseases.

OBJECTIVES: Paediatric-onset immune-mediated inflammatory diseases (IMIDs), including juvenile idiopathic arthritis and related rheumatic diseases, remain genetically undercharacterised. We aimed to define shared and category-specific genetic architecture across paediatric IMIDs, compare signals with adult IMIDs, and identify therapeutic opportunities. METHODS: We analysed 24 paediatric IMIDs classified as autoimmune, polygenic-autoinflammatory, mixed-pattern, or allergic. Genome-wide association analyses included 18,086 cases and 131,019 controls of European ancestry. We estimated single nucleotide polymorphism (SNP)-based heritability, genetic correlations, and polygenic overlap; performed subset-based meta-analysis; and conducted functional annotation, gene prioritisation, pathway and protein network analyses, adult-IMID comparison, and drug-target prioritisation. RESULTS: SNP-based heritability ranged from 28.9% for allergic IMIDs to 61.9% for autoimmune IMIDs. Genetic correlation and polygenic modelling supported partial sharing across categories with category-specific components. Meta-analysis identified 39 genome-wide significant loci outside the Major Histocompatibility Complex (MHC) region, including 15 previously unreported loci; 19 loci were shared between categories. Gene-prioritisation and protein interaction analyses identified a core MHC-centred antigen-presentation network, with category-enriched modules involving complement, innate/barrier pathways, epithelial biology, and type 2 immunity. Enriched pathways included nuclear factor κB signalling, T helper 17 related pathways, Janus kinase-signal transducer and activator of transcription signalling, programmed cell death protein 1/programmed death‑ligand 1, cytotoxic T‑lymphocyte associated protein 4 regulation, and osteoclast differentiation, several of which are relevant to rheumatic diseases. Paediatric IMIDs shared broad polygenic architecture with adult IMIDs, whereas top-ranked genes converged strongly with adult rheumatic diseases. Priority Index analysis identified 178 high-scoring genes, including 43 approved or investigational IMID drug targets. CONCLUSIONS: Paediatric-onset IMIDs share core pathways with adult forms but exhibit distinct genetic architecture shaped by age-specific immune and neurodevelopmental biology. These findings provide a genomic framework for paediatric precision medicine, guiding classification, risk prediction, and therapeutic development.

Humans↗

Rad51-related changes in global gene expression.

High expression of Rad51, the catalytic component in homologous recombination, has been reported to contribute to genomic instability. To elucidate biological processes related to Rad51, we performed global gene expression analysis on human fibrosarcoma cells induced to express variable Rad51 levels. The results indicate that Rad51 overexpression mediates late rather than early transcriptional responses. Using Gene Ontology analysis, we extracted functional annotations for Rad51-related changes in gene expression that were independent of general cell culture effects. High Rad51 levels conferred increased expression of genes involved in actin remodelling. These changes were accompanied by alterations in cell morphology. Moreover, core components of the mismatch repair (MMR) machinery were down-regulated in response to increased Rad51 expression. Given the role of MMR in the correction of DNA mismatches during replication and recombination, a concurrent increase in Rad51 levels and decrease in the expression of MMR genes could conceivably act synergistically towards genomic instability.

Actins↗

Regulation of gene expression in RAW 264.7 macrophage cell line by interferon-gamma.

Macrophages play an important role in immune responses and in inflammatory disease states such as atherosclerosis. Interferon-gamma (IFN-gamma) is a major cytokine involved in the activation of macrophages. To elucidate the primary response of various genes and biological pathways regulated by IFN-gamma in macrophage, we analyzed the gene expression profile in RAW 264.7 macrophage cells treated with IFN-gamma for 4h. Microarray analysis revealed that about 400 genes were differentially expressed, of which about 250 genes were up-regulated and 150 were down-regulated. Functional organization of the transcriptome revealed that induced genes are involved in antimicrobial and antiviral responses, antigen presentation, chemokine and cytokine signaling, and inhibition of cell growth. We also found that expression of genes involved in cell-cycle control, DNA repair, and lipid metabolism was suppressed by IFN-gamma. We also identified induction of multiple transcription factors by IFN-gamma in RAW 264.7 cells. Functional annotation of genes regulated by IFN-gamma in RAW 264.7 cells may provide novel insights into the role of macrophages in immunity and in inflammatory disease.

Animals↗

Can Psychiatric Genetics Advance Without Incorporating a Life Course Perspective?

Psychiatric disorders unfold over the life course; however, genomic studies of these conditions overwhelmingly rely on phenotypes collected at a single time point, often in adulthood. Therefore, genome-wide association studies (GWASs) of psychiatric conditions may miss genetic variants with time-varying relevance to etiology, prevention, and treatment, such as those that influence trajectories of symptoms and behaviors, age at onset, course of treatment response, and the co-evolution of comorbidities. With recent advances in longitudinal biobanks and analytic tools, we posit that incorporating a life course perspective in psychiatric genetics will enable critically relevant insights into each of these areas of investigation. We propose that the current inconsistent portability of polygenic scores across age groups can be reconciled through the design of carefully considered longitudinal GWASs in age-diverse samples. Pioneering longitudinal GWASs in psychiatry have revealed novel genomic signals associated with time-dependent phenotypes that are distinct from those influencing lifetime diagnosis, suggesting that the study of longitudinal phenotypes will complement cross-sectional approaches and empower biological and therapeutic discoveries. Advances in post-GWAS functional annotation resources and analytic approaches now enable us to contextualize the genetic contributions to psychiatric disorders as dynamic age- and exposure-dependent processes. Although longitudinal GWASs pose unique challenges with regard to data availability, selection bias, and missing data, integrating temporality into psychiatric genetics at scale is now attainable and promises to reveal novel biology and therapeutic opportunities for psychiatric conditions.

Cohort study↗

A theory of optimal differential gene expression.

We investigate a model of optimal regulation, intended to describe large-scale differential gene expression. Relations between the optimal expression patterns and the function of genes are deduced from an optimality principle: the regulators have to maximise a fitness function which they influence directly via a cost term, and indirectly via their control on important cell variables, such as metabolic fluxes. According to the model, the optimal linear response to small perturbations reflects the regulators' functions, namely their linear influences on the cell variables. The optimal behaviour can be realised by a linear feedback mechanism. Known or assumed properties of response coefficients lead to predictions about regulation patterns. A symmetry relation predicted for deletion experiments is verified with gene expression data. Where the optimality assumption is valid, our results justify the use of expression data for functional annotation and for pathway reconstruction and suggest the use of linear factor models for the analysis of gene expression data.

Adaptation, Physiological↗

Synonymous codon usage and gene function are strongly related in Oryza sativa.

The relationship between codon usage and gene function was investigated while considering a dataset of 2106 nuclear genes of Oryza sativa. The results of standard chi(2) test and F-statistic showed that for every 59 synonymous codons, a strongly significant association with gene functional categories existed in rice, indicating that codon usage was generally coordinated with gene function whether it was at the level of individual amino acids or at the level of nucleotides. However, it could not be directly said that the use of every codons differed significantly between any two functional categories. Notably, there existed large difference both in selection for biased codons or selection intensity among functional categories. Therefore, we identified at least two classes of genes: one group of genes, mainly belonging to the "METABOLISM" category, was tended to use G- and/or C-ending codons while the other was more biased to choose codons ending with A and/or U. The latter group contained genes of various functions, especially those genes classified into the "Nuclear Structure" category. These observations will be more important for molecular genetic engineering and genome functional annotation.

Chromosome Mapping↗