Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Annotating nucleic acid-binding function based on protein structure.

Many of the targets of structural genomics will be proteins with little or no structural similarity to those currently in the database. Therefore, novel function prediction methods that do not rely on sequence or fold similarity to other known proteins are needed. We present an automated approach to predict nucleic-acid-binding (NA-binding) proteins, specifically DNA-binding proteins. The method is based on characterizing the structural and sequence properties of large, positively charged electrostatic patches on DNA-binding protein surfaces, which typically coincide with the DNA-binding-sites. Using an ensemble of features extracted from these electrostatic patches, we predict DNA-binding proteins with high accuracy. We show that our method does not rely on sequence or structure homology and is capable of predicting proteins of novel-binding motifs and protein structures solved in an unbound state. Our method can also distinguish NA-binding proteins from other proteins that have similar, large positive electrostatic patches on their surfaces, but that do not bind nucleic acids.

Amino Acid Motifs↗

Genomic structure and evolutionary context of the human feline leukemia virus subgroup C receptor (hFLVCR) gene: evidence for block duplications and de novo gene formation within duplicons of the hFLVCR locus.

In this paper we sought to analyze the genomic structure and context of human feline leukemia virus subgroup C receptor (hFLVCR), a human glucarate transporter-like gene at chromosome 1q31, and compare it to that of a paralog (FLVCR14q) at chromosome 14q24. Splicing, polyadenylation, and expression patterns, as estimated by in silico analysis, differed between the two FLVCR genes despite their similar genomic structures, suggesting active and independent evolution of transcriptional and messenger RNA processing patterns after gene duplication. Promoter activity was bi-directional for hFLVCR, but not for its 14q paralog. The upstream 1q transcribed sequences were determined to comprise a novel gene of unknown function, LQK1. Annotation of contigs centered at hFLVCR and FLVCRL14q also revealed highly conserved gene clusters on chromosomes 1 and 14, inferred to result from a duplication. The clusters contained members of the FLVCR, Angel (KIAA0759), JDP, p21SNFT, and TGF- families, as well as two uncharacterized families. The genome-wide locations of both previously recognized and four de novo in silico predicted genes belonging to these seven families were determined. Phylogenetic analyses of these families were consistent with the hypothesis that the 1q/14q duplication occurred early within, or immediately prior to the vertebrate divergence, after the protostome-deuterostome divergence but before the amniote-amphibian divergence.

3' Untranslated Regions↗

Fulfilling the promise: drug discovery in the post-genomic era.

The genomic era has brought with it a basic change in experimentation, enabling researchers to look more comprehensively at biological systems. The sequencing of the human genome coupled with advances in automation and parallelization technologies have afforded a fundamental transformation in the drug target discovery paradigm, towards systematic whole genome and proteome analyses. In conjunction with novel proteomic techniques, genome-wide annotation of function in cellular models is possible. Overlaying data derived from whole genome sequence, expression and functional analysis will facilitate the identification of causal genes in disease and significantly streamline the target validation process. Moreover, several parallel technological advances in small molecule screening have resulted in the development of expeditious and powerful platforms for elucidating inhibitors of protein or pathway function. Conversely, high-throughput and automated systems are currently being used to identify targets of orphan small molecules. The consolidation of these emerging functional genomics and drug discovery technologies promises to reap the fruits of the genomic revolution.

Animals↗

Selection system for genes encoding nuclear-targeted proteins.

Nuclear proteins have essential roles in cell proliferation and differentiation. We have developed a yeast selection system-the nuclear transportation trap (NTT)-to identify genes encoding nuclear transport signals. Both unknown and previously identified nuclear localization signals were identified from a human fetal brain cDNA library. The majority (75%) of the unknown proteins examined were exclusively localized to the nucleus in COS-7 cells. We propose that NTT is an efficient method for isolating cDNAs that encode nuclear targeted proteins that can be applied to the retrieval of novel nuclear proteins and to annotate gene function.

Amino Acid Sequence↗

OpTiles: an R package for adaptive tiling and methylation variability profiling.

SUMMARY: OpTiles is an R package that dynamically defines tiling windows based on the distribution of sequenced CpGs, addressing the limitations of traditional fixed-tiling approaches in targeted methylation datasets. By integrating CpG density with intra-region methylation variability, it provides a reliability metric and extended functionality for annotating, prioritizing, and interpreting complex methylation data. AVAILABILITY AND IMPLEMENTATION: OpTiles is implemented in R and source code is freely available at https://github.com/fhaive/OpTiles. Data are available on Zenodo at https://doi.org/10.5281/zenodo.16961292.

DNA Methylation↗

Prediction of the coding sequences of unidentified human genes. XXI. The complete sequences of 60 new cDNA clones from brain which code for large proteins.

As an extension of a sequencing project of human cDNA clones which encode large proteins of unidentified genes, we herein present the entire sequences of 60 cDNA clones for the genes named KIAA1879-KIAA1938. The cDNA clones were isolated from size-fractionated cDNA libraries derived from human fetal brain, adult whole brain and amygdala, and their protein-coding sequences were predicted. Thirty-seven cDNA clones entirely sequenced in this study were selected as cDNAs which have coding potentiality by in vitro transcription/translation experiments, and the remaining 23 cDNA clones were chosen by computer-assisted analysis of terminal sequences of cDNAs. The average sizes of the inserts and corresponding open reading frames of cDNA clones analyzed here were 4.5 kb and 2.2 kb (733 amino acid residues), respectively. Sequence analyses against the public databases enabled us to annotate the functions of the predicted products of the 25 genes; 84% of these predicted gene products (21 gene products) were classified into proteins related to cell signaling/communication, nucleic acid management, and cell structure/motility. In addition to the sequence information about these 60 genes, their expression profiles were also studied in some human tissues including brain regions by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Adult↗

The SBASE protein domain library, release 7.0: a collection of annotated protein sequence segments.

SBASE 7.0 is the seventh release of the SBASE protein domain library sequences that contains 237 937 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to all major sequence databases and sequence pattern collections. The entries are clustered into over 1811 groups and are provided with two WWW-based search facilities for on-line use. SBASE 7.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb. trieste.it. Automated searching of SBASE with BLAST can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/and http://sbase.abc.hu/sbase/

Amino Acid Sequence↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

The SBASE protein domain library, release 8.0: a collection of annotated protein sequence segments.

SBASE 8.0 is the eighth release of the SBASE library of protein domain sequences that contains 294 898 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to most major sequence databases and sequence pattern collections. The entries are clustered into over 2005 statistically validated domain groups (SBASE-A) and 595 non-validated groups (SBASE-B), provided with several WWW-based search and browsing facilities for online use. A domain-search facility was developed, based on non-parametric pattern recognition methods, including artificial neural networks. SBASE 8.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb.trieste.it. Automated searching of SBASE can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/ and http://sbase.abc. hu/sbase/.

Binding Sites↗

The SBASE protein domain library, release 9.0: an online resource for protein domain identification.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource of protein domain sequences designed to facilitate detection of domain homologies based on a simple database search. The ninth release of the SBASE library of protein domain sequences contains 320 000 annotated structural, functional, ligand-binding and topogenic segments of proteins clustered into over 3481 domain groups and 483 protein families. Domain identification and functional prediction are based on a comparison of BLAST search outputs with a knowledge base of within-group ('self') and out-of-group ('non-self') similarities of the known domain groups. This is a memory-based approach wherein class-specific similarity functions are automatically learned from the database [Stanfill,C. and Waltz,D. (1986) COMMUN: ACM, 29, 1213-1228].

Animals↗

SubtiList: the reference database for the Bacillus subtilis genome.

SubtiList is the reference database dedicated to the genome of Bacillus subtilis 168, the paradigm of Gram-positive endospore-forming bacteria. Developed in the framework of the B.subtilis genome project, SubtiList provides a curated dataset of DNA and protein sequences, combined with the relevant annotations and functional assignments. Information about gene functions and products is continuously updated by linking relevant bibliographic references. Recently, sequence corrections arising from both systematic verifications and submissions by individual scientists were included in the reference genome sequence. SubtiList is based on a generic relational data schema and a World Wide Web interface developed for the handling of bacterial genomes, called GenoList. The World Wide Web interface was designed to allow users to easily browse through genome data and retrieve information according to common biological queries. SubtiList also provides more elaborate tools, such as pattern searching, which are tightly connected to the overall browsing system. SubtiList is accessible at http://genolist.pasteur.fr/SubtiList/. Similar bacterial databases are accessible at http://genolist.pasteur.fr/.

Bacillus subtilis↗

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome↗

Gene and pathway analysis of genome-wide genetic associations of bladder cancer.

BACKGROUND: Although genetic variants associated with bladder cancer (BCa) risk have been identified through hypothesis-driven and genome-wide association studies, a systematic understanding of BCa genetic susceptibility at the gene and pathway levels remains to be achieved. MATERIALS AND METHODS: In this 2-stage functional genomics study, we used 5 independent tools for genome-wide gene mapping and ranking based on BCa genome-wide association studies summary statistics, followed by a meta-analysis of gene-level significance p values, to obtain a consensus gene ranking in terms of association with BCa. Subsequently, we performed preranked gene-set enrichment analysis to identify the functional pathways involved in BCa genetic susceptibility. Joint analysis with gene-set enrichment analysis, based on somatic alteration frequency, was performed to explore the pathway-level relationships between genetic susceptibility and somatic alterations in BCa. RESULTS: Other than the well-known BCa genes (such as FGFR3, MYC, TERT, CCNE1, and TP63), we additionally prioritized a set of novel genes likely to be genetically implicated in BCa development, including SETD2, a possible tumor suppressor gene involved in chromatin remodeling. We further demonstrated convergence between genetic associations and somatic alterations at both the gene (eg, FGFR3 and TERT) and pathway levels (eg, cell cycle and chromatin modification), as well as functional ontologies specifically implicated in germline predisposition to BCa (eg, CD8/TCR signaling, immune checkpoints, and cytokine signaling). CONCLUSIONS: We identified several novel genes associated with BCa and demonstrated that genetic variants contribute to the development of BCa by affecting antitumor immunity, response to toxic exposure, and RNA and protein homeostasis and synergizing with somatic alterations in various cancer-related pathways.

Bladder cancer↗

Gene expression in autumn leaves.

Two cDNA libraries were prepared, one from leaves of a field-grown aspen (Populus tremula) tree, harvested just before any visible sign of leaf senescence in the autumn, and one from young but fully expanded leaves of greenhouse-grown aspen (Populus tremula x tremuloides). Expressed sequence tags (ESTs; 5,128 and 4,841, respectively) were obtained from the two libraries. A semiautomatic method of annotation and functional classification of the ESTs, according to a modified Munich Institute of Protein Sequences classification scheme, was developed, utilizing information from three different databases. The patterns of gene expression in the two libraries were strikingly different. In the autumn leaf library, ESTs encoding metallothionein, early light-inducible proteins, and cysteine proteases were most abundant. Clones encoding other proteases and proteins involved in respiration and breakdown of lipids and pigments, as well as stress-related genes, were also well represented. We identified homologs to many known senescence-associated genes, as well as seven different genes encoding cysteine proteases, two encoding aspartic proteases, five encoding metallothioneins, and 35 additional genes that were up-regulated in autumn leaves. We also indirectly estimated the rate of plastid protein synthesis in the autumn leaves to be less that 10% of that in young leaves.

Arabidopsis Proteins↗

A comparison of position-specific score matrices based on sequence and structure alignments.

Sequence comparison methods based on position-specific score matrices (PSSMs) have proven a useful tool for recognition of the divergent members of a protein family and for annotation of functional sites. Here we investigate one of the factors that affects overall performance of PSSMs in a PSI-BLAST search, the algorithm used to construct the seed alignment upon which the PSSM is based. We compare PSSMs based on alignments constructed by global sequence similarity (ClustalW and ClustalW-pairwise), local sequence similarity (BLAST), and local structure similarity (VAST). To assess performance with respect to identification of conserved functional or structural sites, we examine the accuracy of the three-dimensional molecular models predicted by PSSM-sequence alignments. Using the known structures of those sequences as the standard of truth, we find that model accuracy varies with the algorithm used for seed alignment construction in the pattern local-structure (VAST) > local-sequence (BLAST) > global-sequence (ClustalW). Using structural similarity of query and database proteins as the standard of truth, we find that PSSM recognition sensitivity depends primarily on the diversity of the sequences included in the alignment, with an optimum around 30-50% average pairwise identity. We discuss these observations, and suggest a strategy for constructing seed alignments that optimize PSSM-sequence alignment accuracy and recognition sensitivity.

Algorithms↗

Moderate expression and activity of flocculins underlie the characteristic flocculation phenotype of Saccharomyces pastorianus.

Flocculation is a key technological trait in lager brewing, governing fermentation performance, yeast recovery, and beer quality. In the allo-aneuploid hybrid yeast Saccharomyces pastorianus, the genetic basis of flocculation remains poorly resolved due to its complex dual sub-genome architecture. Here, we systematically re-annotated and functionally characterized the complete FLO gene repertoire of the Group II strain CBS 1483. Thirteen FLO genes were identified, including allelic variants and a previously uncharacterized adhesin, Flo12, containing a Hyphal_reg_CWP domain instead of the canonical PA14 lectin-binding domain. Structural modeling revealed strong conservation of Ca²+-binding residues in PA14 domains, alongside repeat-region diversification likely contributing to functional variability. Using optogenetic expression in a FLO-null background, we demonstrated that SpcI-FLO9-1 and SpcI-FLO9-2_1 are the strongest drivers of flocculation, exhibiting NewFlo-like sugar sensitivity. Transcriptomic analysis during 17°P wort fermentation showed dynamic induction of these genes coinciding with flocculation onset. Surprisingly, deletion of both loci in CBS 1483 did not abolish but only delayed sedimentation in wort, accompanied by improved maltose utilization and attenuation. These findings reveal functional redundancy and compensatory mechanisms within the FLO network of lager yeast, highlighting the genetic complexity underlying flocculation, and providing a molecular framework to inform yeast selection, strain development, and optimization of the lager fermentation processes.IMPORTANCEFlocculation, the process by which yeast cells aggregate and settle, is essential for producing clear, high-quality lager beer, and for efficient yeast recovery during brewing. However, the genetic basis of this trait in lager yeast has remained poorly understood because these strains possess unusually complex hybrid genomes. In this study, we systematically identified and characterized the complete set of flocculation genes in the industrial lager yeast Saccharomyces pastorianus CBS 1483. We demonstrated that lager yeast flocculation is not controlled by a single dominant gene, but instead emerges from the combined action of several moderately active adhesion proteins that are expressed at low levels during fermentation. Surprisingly, deleting the two strongest candidate genes only delayed, rather than eliminated, sedimentation, revealing a robust compensatory network that preserves brewing performance. These findings refine the current understanding of yeast flocculation and provide a molecular framework for developing brewing strains with improved fermentation efficiency, product consistency, and flavor quality.

Saccharomyces pastorianus↗

Identification of temporal patterns of gene expression in the uteri of immature, ovariectomized mice following exposure to ethynylestradiol.

Estrogen induction of uterine wet weight provides an excellent model to investigate relationships between changes in global gene expression and well-characterized physiological responses. In this study, time course microarray GeneChip data were analyzed using a novel approach to identify temporal changes in uterine gene expression following treatment of immature ovariectomized C57BL/6 mice with 0.1 mg/kg 17alpha-ethynylestradiol. Functional gene annotation information from public databases facilitated the association of changes in gene expression with physiological outcomes, which allowed detailed mechanistic inferences to be drawn regarding cell cycle control and proliferation, transcription and translation, structural tissue remodeling, and immunologic responses. These systematic approaches confirm previously established responses, identify novel estrogen-regulated transcriptional effects, and disclose the coordinated activation of multiple modes of action that support the uterotrophic response elicited by estrogen. In particular, it was possible to elucidate the physiological significance of the dramatic induction of arginase, a classic estrogenic response, by elucidating its mechanistic relevance and delineating the role of arginine and ornithine utilization in the estrogen-stimulated induction of uterine wet weight.

Animals↗

FunSpec: a web-based cluster interpreter for yeast.

BACKGROUND: For effective exposition of biological information, especially with regard to analysis of large-scale data types, researchers need immediate access to multiple categorical knowledge bases and need summary information presented to them on collections of genes, as opposed to the typical one gene at a time. RESULTS: We present here a web-based tool (FunSpec) for statistical evaluation of groups of genes and proteins (e.g. co-regulated genes, protein complexes, genetic interactors) with respect to existing annotations (e.g. functional roles, biochemical properties, localization). FunSpec is available online at http://funspec.med.utoronto.ca CONCLUSION: FunSpec is helpful for interpretation of any data type that generates groups of related genes and proteins, such as gene expression clustering and protein complexes, and is useful for predictive methods employing "guilt-by-association."

Cluster Analysis↗