Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Functional genomic studies of aldo-keto reductases.

Aldose reductase (AR) is considered a potential mediator of diabetic complications and is a drug target for inhibitors of diabetic retinopathy and neuropathy in clinical trials. However, the physiological role of this enzyme still has not been established. Since effective inhibition of diabetic complications will require early intervention, it is important to delineate whether AR fulfills a physiological role that cannot be compensated by an alternate aldo-keto reductase. Functional genomics provides a variety of powerful new tools to probe the physiological roles of individual genes, especially those comprising gene families. Several eucaryotic genomes have been sequenced and annotated, including yeast, nematode and fly. To probe the function of AR, we have chosen to utilize the budding yeast Saccharomyces cerevisiae as a potential model system. Unlike Caenorhabditis elegans and D. melanogaster, yeast provides a more desirable system for our studies because its genome is manipulated more readily and is able to sustain multiple gene deletions in the presence of either drug or auxotrophic selectable markers. Using BLAST searches against the human AR gene sequence, we identified six genes in the complete S. cerevisiae genome with strong homology to AR. In all cases, amino acids thought to play important catalytic roles in human AR are conserved in the yeast AR-like genes. All six yeast AR-like open reading frames (ORFs) have been cloned into plasmid expression vectors. Substrate and AR inhibitor specificities have been surveyed on four of the enzyme forms to identify, which are the most functionally similar to human AR. Our data reveal that two of the enzymes (YDR368Wp and YHR104Wp) are notable for their similarity to human AR in terms of activity with aldoses and substituted aromatic aldehydes. Ongoing studies are aimed at characterizing the phenotypes of yeast strains containing single and multiple knockouts of the AR-like genes.

Alcohol Oxidoreductases↗

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny↗

Annotating eukaryote genomes.

The Genome Annotation Assessment Project tested current methods of gene identification, including a critical assessment of the accuracy of different methods. Two new databases have provided new resources for gene annotation: these are the InterPro database of protein domains and motifs, and the Gene Ontology database for terms that describe the molecular functions and biological roles of gene products. Efforts in genome annotation are most often based upon advances in computer systems that are specifically designed to deal with the tremendous amounts of data being generated by current sequencing projects. These efforts in analysis are being linked to new ways of visualizing computationally annotated genomes.

Animals↗

Malaria and the red blood cell membrane.

Malaria is the most serious and widespread parasitic disease of humans and is arguably the commonest disease of red blood cells (RBCs). Malaria has exerted a powerful effect on human evolution and selection for resistance has led to the appearance and persistence of a number of inherited diseases. After parasite invasion, RBCs are progressively and dramatically modified. New structures appear inside the RBC and novel parasite proteins are exported to the erythrocyte cytoplasm and membrane skeleton. Radical biochemical, morphological, and rheological alterations manifest as increased membrane rigidity, reduced cell deformability, and greater adhesiveness for the vascular endothelium and other blood cells. Numerous protein-protein interactions between the malaria-parasite and the host RBC are important for many aspects of parasite biology and the pathogenesis of malaria. In addition, there are many other parasite proteins located within the infected red cell and at the membrane skeleton, for which no precise functional roles have yet been elucidated. Sequencing and annotation of the complete genome of Plasmodium falciparum, the production of proteomic and transcriptomic profiles of parasites, and the development of a transfection system for the asexual stage of the parasite are all recent achievements that should advance understanding of the molecular mechanisms that underlie the parasite-induced functional alterations in red cells.

Animals↗

Annotation of bacterial genomes using improved phylogenomic profiles.

MOTIVATION: Phylogenomic profiling is a large-scale comparative genomic method used to infer protein function from evolutionary information first described in a binary form by Pellegrini et al. (1999). Here, we propose improvements of this approach including the use of normalized Blastp bit scores, a normalization of the matrix of profiles to take into account the evolutionary distances between bacteria, the definition of a phylogenomic neighborhood based on continuous pairwise distances between genes and an original annotation procedure including the computation of a p-value for each functional assignment. RESULTS: The method presented here increases the number of Ecocyc enzymes identified as being evolutionarily related by about 25% with respect to the original binary form (absent/present) method. The fraction of 'false' positives is shown to be smaller than 20%. Based on their phylogenomic relationships, genes of unknown function can then be automatically related to annotated genes. Each gene annotation predicted is associated with a p-value, i.e. its probability to be obtained by chance. The validity of this method was extensively tested on a large set of genes of known function using the MultiFun database. We find that 50% of 3122 function attributions that can be made at a p-value level of 10(-11) correspond to the actual gene annotation. The method can be readily applied to any newly sequenced microbial genome. In contrast to earlier work on the same topic, our approach avoids the use of arbitrary cut-off values, and provides a reliability estimate of the functional predictions in form of p-values.

Algorithms↗

Small genes/gene-products in Escherichia coli K-12.

Forty-two protein spots of observed M(r) 6-15 kDa were resolved by two-dimensional gel electrophoresis, stained by Coomassie blue and subjected to Edman microsequencing. All of the proteins could be related back to their encoding open reading frames, thereby vindicating the bioinformatic tools currently utilised in their identification. However, only 14/42 gene-products were expressed as annotated. Translation was confirmed for 14 open reading frames with no attributed function (EcoGene Y-entries), while N-terminal sequence allowed the start codon to be accurately annotated for the genes yigF, yccU, yqiC, ynfD, and yeeX. The methionine start codon was cleaved in 11 gene-products (AtpE, Hns, RpoZ, RplL, CspC, YccJ, YggX, YjgF, HimA, InfA, RpsQ) and a further five showed loss of a signal peptide (PspE, HdeB, HdeA, YnfD, YkfE). Internal (Tig, AtpA, TufA) and N-terminal fragmentation (CspD, RpsF, AtcU) of much larger proteins was also detected, which may have resulted from physiological or translational processes. M(r) and pI isoforms were detected respectively for PtsH and GatB, each being phosphoproteins, as well as RplY which manifested differences with respect to predicted M(r) and pI. In addition, YjgF was shown to belong to a small gene family of unknown function with ancient conserved regions across procaryotes and eucaryotes. YgiN was revealed to have a paralogue and orthologues in Bacillus subtilis, Synechocystis sp., Mycobacterium tuberculosis, Neisseria gonorrhoea, and Rhodococcus erythropolis. Orthologues are also reported for YihD, YccU and YeeX. Of the 14 Y-genes, only YkfE possessed no detectable orthologues. These results highlight the need to complement genomic analysis with detailed proteomics in order to gain a better understanding of cellular molecular biology, while the confirmation of the open reading frame start codon using Edman degradation protein microsequencing has yet to be superseded by recent advances in mass spectrometry.

Amino Acid Sequence↗

Multiple alignment of complete sequences (MACS) in the post-genomic era.

Multiple alignment, since its introduction in the early seventies, has become a cornerstone of modern molecular biology. It has traditionally been used to deduce structure / function by homology, to detect conserved motifs and in phylogenetic studies. There has recently been some renewed interest in the development of multiple alignment techniques, with current opinion moving away from a single all-encompassing algorithm to iterative and / or co-operative strategies. The exploitation of multiple alignments in genome annotation projects represents a qualitative leap in the functional analysis process, opening the way to the study of the co-evolution of validated sets of proteins and to reliable phylogenomic analysis. However, the alignment of the highly complex proteins detected by today's advanced database search methods is a daunting task. In addition, with the explosion of the sequence databases and with the establishment of numerous specialized biological databases, multiple alignment programs must evolve if they are to successfully rise to the new challenges of the post-genomic era. The way forward is clearly an integrated system bringing together sequence data, knowledge-based systems and prediction methods with their inherent unreliability. The incorporation of such heterogeneous, often non-consistent, data will require major changes to the fundamental alignment algorithms used to date. Such an integrated multiple alignment system will provide an ideal workbench for the validation, propagation and presentation of this information in a format that is concise, clear and intuitive.

Amino Acid Sequence↗

Real spherical harmonic expansion coefficients as 3D shape descriptors for protein binding pocket and ligand comparisons.

MOTIVATION: An increasing number of protein structures are being determined for which no biochemical characterization is available. The analysis of protein structure and function assignment is becoming an unexpected challenge and a major bottleneck towards the goal of well-annotated genomes. As shape plays a crucial role in biomolecular recognition and function, the examination and development of shape description and comparison techniques is likely to be of prime importance for understanding protein structure-function relationships. RESULTS: A novel technique is presented for the comparison of protein binding pockets. The method uses the coefficients of a real spherical harmonics expansion to describe the shape of a protein's binding pocket. Shape similarity is computed as the L2 distance in coefficient space. Such comparisons in several thousands per second can be carried out on a standard linux PC. Other properties such as the electrostatic potential fit seamlessly into the same framework. The method can also be used directly for describing the shape of proteins and other molecules. AVAILABILITY: A limited version of the software for the real spherical harmonics expansion of a set of points in PDB format is freely available upon request from the authors. Binding pocket comparisons and ligand prediction will be made available through the protein structure annotation pipeline Profunc (written by Roman Laskowski) which will be accessible from the EBI website shortly.

Algorithms↗

Visualizing the genome: techniques for presenting human genome data and annotations.

BACKGROUND: In order to take full advantage of the newly available public human genome sequence data and associated annotations, biologists require visualization tools ("genome browsers") that can accommodate the high frequency of alternative splicing in human genes and other complexities. RESULTS: In this article, we describe visualization techniques for presenting human genomic sequence data and annotations in an interactive, graphical format. These techniques include: one-dimensional, semantic zooming to show sequence data alongside gene structures; color-coding exons to indicate frame of translation; adjustable, moveable tiers to permit easier inspection of a genomic scene; and display of protein annotations alongside gene structures to show how alternative splicing impacts protein structure and function. These techniques are illustrated using examples from two genome browser applications: the Neomorphic GeneViewer annotation tool and ProtAnnot, a prototype viewer which shows protein annotations in the context of genomic sequence. CONCLUSION: By presenting techniques for visualizing genomic data, we hope to provide interested software developers with a guide to what features are most likely to meet the needs of biologists as they seek to make sense of the rapidly expanding body of public genomic data and annotations.

Alternative Splicing↗

PIRSF: family classification system at the Protein Information Resource.

The Protein Information Resource (PIR) is an integrated public resource of protein informatics. To facilitate the sensible propagation and standardization of protein annotation and the systematic detection of annotation errors, PIR has extended its superfamily concept and developed the SuperFamily (PIRSF) classification system. Based on the evolutionary relationships of whole proteins, this classification system allows annotation of both specific biological and generic biochemical functions. The system adopts a network structure for protein classification from superfamily to subfamily levels. Protein family members are homologous (sharing common ancestry) and homeomorphic (sharing full-length sequence similarity with common domain architecture). The PIRSF database consists of two data sets, preliminary clusters and curated families. The curated families include family name, protein membership, parent-child relationship, domain architecture, and optional description and bibliography. PIRSF is accessible from the website at http://pir.georgetown.edu/pirsf/ for report retrieval and sequence classification. The report presents family annotation, membership statistics, cross-references to other databases, graphical display of domain architecture, and links to multiple sequence alignments and phylogenetic trees for curated families. PIRSF can be utilized to analyze phylogenetic profiles, to reveal functional convergence and divergence, and to identify interesting relationships between homeomorphic families, domains and structural classes.

Amino Acid Motifs↗

BASys: a web server for automated bacterial genome annotation.

BASys (Bacterial Annotation System) is a web server that supports automated, in-depth annotation of bacterial genomic (chromosomal and plasmid) sequences. It accepts raw DNA sequence data and an optional list of gene identification information and provides extensive textual annotation and hyperlinked image output. BASys uses >30 programs to determine approximately 60 annotation subfields for each gene, including gene/protein name, GO function, COG function, possible paralogues and orthologues, molecular weight, isoelectric point, operon structure, subcellular localization, signal peptides, transmembrane regions, secondary structure, 3D structure, reactions and pathways. The depth and detail of a BASys annotation matches or exceeds that found in a standard SwissProt entry. BASys also generates colorful, clickable and fully zoomable maps of each query chromosome to permit rapid navigation and detailed visual analysis of all resulting gene annotations. The textual annotations and images that are provided by BASys can be generated in approximately 24 h for an average bacterial chromosome (5 Mb). BASys annotations may be viewed and downloaded anonymously or through a password protected access system. The BASys server and databases can also be downloaded and run locally. BASys is accessible at http://wishart.biology.ualberta.ca/basys.

Chromosomes, Bacterial↗

Identifying Co-Expressed lncRNAs Correlated With Traits of Interest in an Animal Model for Metabolic Diseases in Humans.

Nutrigenomics investigates how nutrients modulate gene expression. Among them, fatty acids (FA) play important roles in regulating gene transcription, while long non-coding RNAs (lncRNAs) may be associated with gene regulation and metabolic diseases. This study aimed to analyze the hepatic transcriptome of pigs, a species frequently used as a model for nutrigenomic studies, to identify novel lncRNAs and their potential target genes in response to diets containing different sources of FA. Seventy-two pigs were fed four diets supplemented with 1.5% soybean oil (control), 3% canola oil, 3% fish oil, and 3% soybean oil. RNA sequencing of liver samples was performed to identify novel lncRNAs. Weighted Gene Co-expression Network Analysis (WGCNA) was used to identify modules associated with phenotypic traits related to lipid metabolism and inflammation. Functional enrichment analyses were then conducted to annotate genes within these modules using Gene Ontology (GO) terms and to assess overlap with Quantitative Trait Loci (QTL). The results revealed 106 novel lncRNAs potentially regulating genes associated with lipid metabolism and immune responses in pigs fed diets with different FA sources. These findings enhance understanding of the regulatory role of lncRNAs in pigs and reinforce their relevance as models for human metabolic diseases.

Animals↗

Mutational data integration in gene-oriented files of the Hermansky-Pudlak Syndrome database.

Hermansky-Pudlak Syndrome (HPS) is a genetically heterogeneous disorder characterized by oculocutaneous albinism and prolonged bleeding due to abnormal vesicle trafficking to lysosomes and related organelles such as melanosomes and platelet dense granules. This HPS database (HPSD; http://liweilab.genetics.ac.cn/HPSD/) provides integrated, annotatory, and curative data that is distributed in a variety of public databases or predicted by bioinformatics servers for the recently cloned human and mouse HPS genes, as well as for the genes responsible for HPSrelated syndromes, such as ChediakHigashi Syndrome (CHS), Griscelli syndrome (GS), oculocutaneous albinism (OCA), Usher syndrome type 1B (USH1B), and ocular albinism (OA). The HPSD is designed by using a unique GeneOriented File (GOF) format. Seven blocks (genomic, transcript, protein, function, mutation, phenotype, and reference) are carefully annotated in each userfriendly GOF entry. The HPSD emphasizes paired human and mouse GOF entries. The genes included in this database (currently 58 in total) are arbitrarily divided into four categories: 1) Human and Mouse HPS, 2) Mouse HPS Only, 3) Putative Mouse or Human HPS, and 4) HPS Related Syndromes. All the mutations in these genes are integrated in the GOFs. We expect that these very informative and peerreviewed GOFs will be shortcuts to utilize the webbased information for the emerging interdisciplinary studies of HPS.

Animals↗

Proteomic analysis on metastasis-associated proteins of human hepatocellular carcinoma tissues.

PURPOSE: A comparative proteomic approach was used to identify and analyze proteins related to metastasis of hepatocellular carcinoma (HCC). METHODS: Proteins extracted from 12 HCC tissue specimens (six with metastases and six without) were separated by two-dimensional gel electrophoresis (2-DE). The protein spots exhibiting statistical alternations between the two groups through computerized image analysis were then identified by mass spectrometry. In addition immunohistochemistry (IHC), Western blotting and RT-PCR were performed to verify the expression of certain candidate proteins. RESULTS: 16 proteins including HSP27, S100A11, CK18 were annotated by mass spectrometry, relevant to chaperone function, cell mobility, cytoskeletal architecture, respectively. Most were previously unconnected with metastasis of HCC. Of these HSP27 was found overexpressed consistently in 2-DE patterns of all metastatic HCC tissues compared with nonmetastatic ones. IHC and Western blotting of HCC tissues confirmed this difference while RT-PCR did not. CONCLUSION: There are various proteins joined together in HCC metastasis. The overexpression of HSP27 may serve as a biomarker for early detection and therapeutic targets unique to the metastatic phenotype of HCC.

Actins↗

LigProf: a simple tool for in silico prediction of ligand-binding sites.

With the increasing amount of data provided by both high-throughput sequencing and structural genomics studies, there is a growing need for tools to augment functional predictions for protein sequences. Broad descriptions of function can be provided by establishing the presence of protein domains associated with a particular function. To extend the domain-based annotation, LigProf provides predictions of potential ligands that bind to a protein, as well as critical residues that stabilize ligands. A P-value statistic for estimating the significance of motif occurrence is provided for all sites. Although the usefulness of the method will rise with increasing numbers of crystallographically solved molecules deposited in the PDB database, we show that it can already be applied successfully to the highly represented ligand-bound protein kinase domains of viral and human origin. The LigProf webserver is freely available at: http://www.cropnet.pl/ligprof . At present, LigProf descriptors annotate and extend major protein families from the PfamA database.

Binding Sites↗

Global analysis of bacterial transcription factors to predict cellular target processes.

Whole-genome sequences are now available for >100 bacterial species, giving unprecedented power to comparative genomics approaches. We have applied genome-context methods to predict target processes that are regulated by transcription factors (TFs). Of 128 orthologous groups of proteins annotated as TFs, to date, 36 are functionally uncharacterized; in our analysis we predict a probable cellular target process or biochemical pathway for half of these functionally uncharacterized TFs.

Bacteria↗

Assigning function to yeast proteins by integration of technologies.

Interpreting genome sequences requires the functional analysis of thousands of predicted proteins, many of which are uncharacterized and without obvious homologs. To assess whether the roles of large sets of uncharacterized genes can be assigned by targeted application of a suite of technologies, we used four complementary protein-based methods to analyze a set of 100 uncharacterized but essential open reading frames (ORFs) of the yeast Saccharomyces cerevisiae. These proteins were subjected to affinity purification and mass spectrometry analysis to identify copurifying proteins, two-hybrid analysis to identify interacting proteins, fluorescence microscopy to localize the proteins, and structure prediction methodology to predict structural domains or identify remote homologies. Integration of the data assigned function to 48 ORFs using at least two of the Gene Ontology (GO) categories of biological process, molecular function, and cellular component; 77 ORFs were annotated by at least one method. This combination of technologies, coupled with annotation using GO, is a powerful approach to classifying genes.

Computational Biology↗

Genetic heterogeneity affects the risk of incident depression, comorbidity, and response to environment: A prospective trajectory study.

BACKGROUND: Depression exhibits significant heterogeneity in its genetic underpinnings. The role of genetic components in the development of depression and its comorbidities remains insufficiently explored. METHODS: First, depression risk loci from a large-scale genome-wide meta-analysis were annotated to Gene Ontology (GO) terms by functional enrichment. GO-based polygenic risk scores (GO-PRS) were then calculated for individuals in the UK Biobank. Principal component analysis (PCA) was applied for dimensionality reduction, followed by cluster analysis to identify genetic subtypes of depression. Multistate models were applied to assess the impact of genetic patterns on the trajectory from healthy status to incident depression, and depression to 26 subsequent diseases, as well as the associations between environmental factors and disease trajectories across genetic subtypes. RESULTS: Participants were categorized into three genetic subtypes: immune-dominant, neuro-dominant, and comprehensive-risk. Significant differences in risk of depression and subsequent diseases, and susceptibility to environmental factors were observed across subtypes. Comprehensive-risk subtype showed higher risks of depression compared to immune-dominant (HR: 1.10, 95% CI: 1.05-1.15) and neuro-dominant subtype (HR: 1.12, 95% CI: 1.08-1.16). Comprehensive-risk subtype exhibited higher risks of transition from depression to subsequent diseases, such as anemia compared to immune-dominant subtype, and diseases of the digestive system compared to neuro-dominant subtype. Environmental factors were more strongly associated with the transition from depression to subsequent diseases in immune-dominant and comprehensive-risk subtypes, including cardiovascular, respiratory, and metabolic diseases. CONCLUSIONS: Our findings highlight the genetic heterogeneity of depression and comorbidities, and shed light on how genetic components modulate responses to environmental factors.

Humans↗