Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

The SBASE protein domain library, release 7.0: a collection of annotated protein sequence segments.

SBASE 7.0 is the seventh release of the SBASE protein domain library sequences that contains 237 937 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to all major sequence databases and sequence pattern collections. The entries are clustered into over 1811 groups and are provided with two WWW-based search facilities for on-line use. SBASE 7.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb. trieste.it. Automated searching of SBASE with BLAST can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/and http://sbase.abc.hu/sbase/

Amino Acid Sequence↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

The SBASE protein domain library, release 8.0: a collection of annotated protein sequence segments.

SBASE 8.0 is the eighth release of the SBASE library of protein domain sequences that contains 294 898 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to most major sequence databases and sequence pattern collections. The entries are clustered into over 2005 statistically validated domain groups (SBASE-A) and 595 non-validated groups (SBASE-B), provided with several WWW-based search and browsing facilities for online use. A domain-search facility was developed, based on non-parametric pattern recognition methods, including artificial neural networks. SBASE 8.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb.trieste.it. Automated searching of SBASE can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/ and http://sbase.abc. hu/sbase/.

Binding Sites↗

The SBASE protein domain library, release 9.0: an online resource for protein domain identification.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource of protein domain sequences designed to facilitate detection of domain homologies based on a simple database search. The ninth release of the SBASE library of protein domain sequences contains 320 000 annotated structural, functional, ligand-binding and topogenic segments of proteins clustered into over 3481 domain groups and 483 protein families. Domain identification and functional prediction are based on a comparison of BLAST search outputs with a knowledge base of within-group ('self') and out-of-group ('non-self') similarities of the known domain groups. This is a memory-based approach wherein class-specific similarity functions are automatically learned from the database [Stanfill,C. and Waltz,D. (1986) COMMUN: ACM, 29, 1213-1228].

Animals↗

SubtiList: the reference database for the Bacillus subtilis genome.

SubtiList is the reference database dedicated to the genome of Bacillus subtilis 168, the paradigm of Gram-positive endospore-forming bacteria. Developed in the framework of the B.subtilis genome project, SubtiList provides a curated dataset of DNA and protein sequences, combined with the relevant annotations and functional assignments. Information about gene functions and products is continuously updated by linking relevant bibliographic references. Recently, sequence corrections arising from both systematic verifications and submissions by individual scientists were included in the reference genome sequence. SubtiList is based on a generic relational data schema and a World Wide Web interface developed for the handling of bacterial genomes, called GenoList. The World Wide Web interface was designed to allow users to easily browse through genome data and retrieve information according to common biological queries. SubtiList also provides more elaborate tools, such as pattern searching, which are tightly connected to the overall browsing system. SubtiList is accessible at http://genolist.pasteur.fr/SubtiList/. Similar bacterial databases are accessible at http://genolist.pasteur.fr/.

Bacillus subtilis↗

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome↗

Gene and pathway analysis of genome-wide genetic associations of bladder cancer.

BACKGROUND: Although genetic variants associated with bladder cancer (BCa) risk have been identified through hypothesis-driven and genome-wide association studies, a systematic understanding of BCa genetic susceptibility at the gene and pathway levels remains to be achieved. MATERIALS AND METHODS: In this 2-stage functional genomics study, we used 5 independent tools for genome-wide gene mapping and ranking based on BCa genome-wide association studies summary statistics, followed by a meta-analysis of gene-level significance p values, to obtain a consensus gene ranking in terms of association with BCa. Subsequently, we performed preranked gene-set enrichment analysis to identify the functional pathways involved in BCa genetic susceptibility. Joint analysis with gene-set enrichment analysis, based on somatic alteration frequency, was performed to explore the pathway-level relationships between genetic susceptibility and somatic alterations in BCa. RESULTS: Other than the well-known BCa genes (such as FGFR3, MYC, TERT, CCNE1, and TP63), we additionally prioritized a set of novel genes likely to be genetically implicated in BCa development, including SETD2, a possible tumor suppressor gene involved in chromatin remodeling. We further demonstrated convergence between genetic associations and somatic alterations at both the gene (eg, FGFR3 and TERT) and pathway levels (eg, cell cycle and chromatin modification), as well as functional ontologies specifically implicated in germline predisposition to BCa (eg, CD8/TCR signaling, immune checkpoints, and cytokine signaling). CONCLUSIONS: We identified several novel genes associated with BCa and demonstrated that genetic variants contribute to the development of BCa by affecting antitumor immunity, response to toxic exposure, and RNA and protein homeostasis and synergizing with somatic alterations in various cancer-related pathways.

Bladder cancer↗

Gene expression in autumn leaves.

Two cDNA libraries were prepared, one from leaves of a field-grown aspen (Populus tremula) tree, harvested just before any visible sign of leaf senescence in the autumn, and one from young but fully expanded leaves of greenhouse-grown aspen (Populus tremula x tremuloides). Expressed sequence tags (ESTs; 5,128 and 4,841, respectively) were obtained from the two libraries. A semiautomatic method of annotation and functional classification of the ESTs, according to a modified Munich Institute of Protein Sequences classification scheme, was developed, utilizing information from three different databases. The patterns of gene expression in the two libraries were strikingly different. In the autumn leaf library, ESTs encoding metallothionein, early light-inducible proteins, and cysteine proteases were most abundant. Clones encoding other proteases and proteins involved in respiration and breakdown of lipids and pigments, as well as stress-related genes, were also well represented. We identified homologs to many known senescence-associated genes, as well as seven different genes encoding cysteine proteases, two encoding aspartic proteases, five encoding metallothioneins, and 35 additional genes that were up-regulated in autumn leaves. We also indirectly estimated the rate of plastid protein synthesis in the autumn leaves to be less that 10% of that in young leaves.

Arabidopsis Proteins↗

A comparison of position-specific score matrices based on sequence and structure alignments.

Sequence comparison methods based on position-specific score matrices (PSSMs) have proven a useful tool for recognition of the divergent members of a protein family and for annotation of functional sites. Here we investigate one of the factors that affects overall performance of PSSMs in a PSI-BLAST search, the algorithm used to construct the seed alignment upon which the PSSM is based. We compare PSSMs based on alignments constructed by global sequence similarity (ClustalW and ClustalW-pairwise), local sequence similarity (BLAST), and local structure similarity (VAST). To assess performance with respect to identification of conserved functional or structural sites, we examine the accuracy of the three-dimensional molecular models predicted by PSSM-sequence alignments. Using the known structures of those sequences as the standard of truth, we find that model accuracy varies with the algorithm used for seed alignment construction in the pattern local-structure (VAST) > local-sequence (BLAST) > global-sequence (ClustalW). Using structural similarity of query and database proteins as the standard of truth, we find that PSSM recognition sensitivity depends primarily on the diversity of the sequences included in the alignment, with an optimum around 30-50% average pairwise identity. We discuss these observations, and suggest a strategy for constructing seed alignments that optimize PSSM-sequence alignment accuracy and recognition sensitivity.

Algorithms↗

Moderate expression and activity of flocculins underlie the characteristic flocculation phenotype of Saccharomyces pastorianus.

Flocculation is a key technological trait in lager brewing, governing fermentation performance, yeast recovery, and beer quality. In the allo-aneuploid hybrid yeast Saccharomyces pastorianus, the genetic basis of flocculation remains poorly resolved due to its complex dual sub-genome architecture. Here, we systematically re-annotated and functionally characterized the complete FLO gene repertoire of the Group II strain CBS 1483. Thirteen FLO genes were identified, including allelic variants and a previously uncharacterized adhesin, Flo12, containing a Hyphal_reg_CWP domain instead of the canonical PA14 lectin-binding domain. Structural modeling revealed strong conservation of Ca²+-binding residues in PA14 domains, alongside repeat-region diversification likely contributing to functional variability. Using optogenetic expression in a FLO-null background, we demonstrated that SpcI-FLO9-1 and SpcI-FLO9-2_1 are the strongest drivers of flocculation, exhibiting NewFlo-like sugar sensitivity. Transcriptomic analysis during 17°P wort fermentation showed dynamic induction of these genes coinciding with flocculation onset. Surprisingly, deletion of both loci in CBS 1483 did not abolish but only delayed sedimentation in wort, accompanied by improved maltose utilization and attenuation. These findings reveal functional redundancy and compensatory mechanisms within the FLO network of lager yeast, highlighting the genetic complexity underlying flocculation, and providing a molecular framework to inform yeast selection, strain development, and optimization of the lager fermentation processes.IMPORTANCEFlocculation, the process by which yeast cells aggregate and settle, is essential for producing clear, high-quality lager beer, and for efficient yeast recovery during brewing. However, the genetic basis of this trait in lager yeast has remained poorly understood because these strains possess unusually complex hybrid genomes. In this study, we systematically identified and characterized the complete set of flocculation genes in the industrial lager yeast Saccharomyces pastorianus CBS 1483. We demonstrated that lager yeast flocculation is not controlled by a single dominant gene, but instead emerges from the combined action of several moderately active adhesion proteins that are expressed at low levels during fermentation. Surprisingly, deleting the two strongest candidate genes only delayed, rather than eliminated, sedimentation, revealing a robust compensatory network that preserves brewing performance. These findings refine the current understanding of yeast flocculation and provide a molecular framework for developing brewing strains with improved fermentation efficiency, product consistency, and flavor quality.

Saccharomyces pastorianus↗

Identification of temporal patterns of gene expression in the uteri of immature, ovariectomized mice following exposure to ethynylestradiol.

Estrogen induction of uterine wet weight provides an excellent model to investigate relationships between changes in global gene expression and well-characterized physiological responses. In this study, time course microarray GeneChip data were analyzed using a novel approach to identify temporal changes in uterine gene expression following treatment of immature ovariectomized C57BL/6 mice with 0.1 mg/kg 17alpha-ethynylestradiol. Functional gene annotation information from public databases facilitated the association of changes in gene expression with physiological outcomes, which allowed detailed mechanistic inferences to be drawn regarding cell cycle control and proliferation, transcription and translation, structural tissue remodeling, and immunologic responses. These systematic approaches confirm previously established responses, identify novel estrogen-regulated transcriptional effects, and disclose the coordinated activation of multiple modes of action that support the uterotrophic response elicited by estrogen. In particular, it was possible to elucidate the physiological significance of the dramatic induction of arginase, a classic estrogenic response, by elucidating its mechanistic relevance and delineating the role of arginine and ornithine utilization in the estrogen-stimulated induction of uterine wet weight.

Animals↗

FunSpec: a web-based cluster interpreter for yeast.

BACKGROUND: For effective exposition of biological information, especially with regard to analysis of large-scale data types, researchers need immediate access to multiple categorical knowledge bases and need summary information presented to them on collections of genes, as opposed to the typical one gene at a time. RESULTS: We present here a web-based tool (FunSpec) for statistical evaluation of groups of genes and proteins (e.g. co-regulated genes, protein complexes, genetic interactors) with respect to existing annotations (e.g. functional roles, biochemical properties, localization). FunSpec is available online at http://funspec.med.utoronto.ca CONCLUSION: FunSpec is helpful for interpretation of any data type that generates groups of related genes and proteins, such as gene expression clustering and protein complexes, and is useful for predictive methods employing "guilt-by-association."

Cluster Analysis↗

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50 K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40 kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals↗

Genomic distribution characteristics and interspecific differences of microsatellite landscapes in Felidae.

BACKGROUND: Microsatellites within genomes play crucial roles in regulating gene expression, DNA replication, and chromosomal structure and function. Analyzing the composition and distribution patterns of microsatellites in closely related species not only reveals their evolutionary dynamics and adaptive mechanisms but also provides essential technical support for applications in genetic breeding, species conservation, and disease research. As one of the world's most captivating animal groups, the landscape patterns of microsatellites across feline genomes remain to be systematically characterized. RESULTS: This study utilized high-quality genomic data to conduct a systematic comparative analysis of microsatellite landscape distribution patterns across the genomes of 13 felid species. The findings revealed that microsatellite abundance and distribution exhibit species-specific characteristics, with a non-random genomic distribution and a negative correlation between microsatellite abundance and repeat length. The predominant distribution pattern followed the sequence: single > double > quadruple > triple > quintuple > sextuple nucleotide repeats. Microsatellite abundance peaked in intergenic regions, whereas trinucleotide repeats were more prevalent within exons. Coding regions showed a marked preference for trinucleotide and hexanucleotide repeats. Enrichment analysis of GO and KEGG pathways indicated that coding sequences containing microsatellites were primarily involved in transcription and translation processes. CONCLUSIONS: Our study elucidates the distribution patterns and characteristics of microsatellites across diverse feline species, providing significant insights into their evolutionary mechanisms and functional roles. Furthermore, these findings establish a valuable reference and foundational dataset for the future development of high-quality, species-specific microsatellite markers in felids.

Animals↗

Similarities and differences in genome-wide expression data of six organisms.

Comparing genomic properties of different organisms is of fundamental importance in the study of biological and evolutionary principles. Although differences among organisms are often attributed to differential gene expression, genome-wide comparative analysis thus far has been based primarily on genomic sequence information. We present a comparative study of large datasets of expression profiles from six evolutionarily distant organisms: S. cerevisiae, C. elegans, E. coli, A. thaliana, D. melanogaster, and H. sapiens. We use genomic sequence information to connect these data and compare global and modular properties of the transcription programs. Linking genes whose expression profiles are similar, we find that for all organisms the connectivity distribution follows a power-law, highly connected genes tend to be essential and conserved, and the expression program is highly modular. We reveal the modular structure by decomposing each set of expression data into coexpressed modules. Functionally related sets of genes are frequently coexpressed in multiple organisms. Yet their relative importance to the transcription program and their regulatory relationships vary among organisms. Our results demonstrate the potential of combining sequence and expression data for improving functional gene annotation and expanding our understanding of how gene expression and diversity evolved.

Animals↗

eQTM (expression quantitative trait methylation) Atlas: a comprehensive resource of over 11 million DNA methylation-gene expression associations through across 11 tissues and 4 diseases.

MOTIVATION: Epigenome-wide association studies (EWAS) have identified numerous DNA methylation (DNAm) CpG sites associated with complex traits and diseases, but interpretation of those CpG sites remains challenging because in EWAS, CpGs are mostly linked to nearby genes based only on genomic proximity. Expression quantitative trait methylation (eQTM) analyses connect DNAm CpGs with statistically associated gene expression levels. However, a comprehensive, searchable resource integrating eQTMs across diverse tissues and disease contexts has been lacking. RESULTS: We developed the eQTM Atlas, a web-based resource that manually curates more than 11 million DNAm-gene expression associations from eight cohorts, covering 11 tissue types, four broad disease contexts, 173,886 unique CpG probes and 20,231 unique genes. The Atlas supports gene- or CpG- searches by tissue or disease type and finding associated CpG or genes, visualization of cis- and trans-eQTMs through genome browser, heatmap interfaces across various tissues, and cohort-level data downloads. By integrating eQTM results with EWAS resources, the eQTM Atlas enables users to connect disease- or trait-associated CpGs to statistically associated genes rather than relying solely on proximity-based gene annotation, supporting functional interpretation of EWAS findings and generation of disease-specific regulatory hypotheses. AVAILABILITY AND IMPLEMENTATION: The eQTM Atlas is freely available at https://shiny.crc.pitt.edu/eqtm_browser/. The web interface is implemented in R Shiny and hosted through the University of Pittsburgh Center for Research Computing (CRC). Source code is available at https://github.com/ads303/eQTM-Atlas.

DNA methylation↗

[Application of DNA chip technology to biomedical research].

The completion of Human Genome Project enabled us to access to the information on nucleotide sequences of whole human genome. One of the most valuable information on human genome would be the list of approximately 35,000 genes. Although 35% of them are still needed to annotate their functions, we can genome-widely approach to various conditions including disease states. To analyze bunch of information at once, we need high-throughput technology containing most of genes. DNA chip successfully provide a stable platform technology for the massive screening of genomes. Microarrays can be used to obtain genome-wide fingerprint on transcriptional changes in various physiological and pathological conditions, leading to the mining novel genes related to those specific states. We can check the multiple molecular markers for diagnosis, prediction or prognosis of specific diseases. Data from microarray will provide huge amounts of experssion profile, which might induce the transformation of biomedical research.

Base Sequence↗

A bioinformatics-based approach for the prediction and identification of novel proteins potentially involved in phosphorylation signalling pathways.

Together with the explosion in the availability of genome data of a number of organisms including human and mouse, various methods and programs for computational prediction of protein-coding genes and annotation of functional proteins have dramatically increased. For the last decade there has been intense interest in the role of protein phosphorylation which is involved in post-translation modification mechanisms critically regulating inter/intracellular communication, patho/physiological responses and homeostasis during many biological processes. In the present study a total of 202 functionally uncharacterized human full-coding cDNA sequences were investigated using a bioinformatics-based approach. Ten novel potential substrates for protein kinases have been identified which may play multiple roles in regulating intracellular phosphorylation signalling pathways. In addition, 5 of those may be involved in the human-only post-translation mechanism regulated by specific protein kinases. The data presented here therefore would greatly contribute toward the understanding of human molecular basis and cellular signalling networks.

Amino Acid Motifs↗