Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

aCHEdb: the database system for ESTHER, the alpha/beta fold family of proteins and the Cholinesterase gene server.

Acetylcholinesterase belongs to a family of proteins, the alpha/beta hydrolase fold family, whose constituents evolutionarily diverged from a common ancestor and share a similar structure of a central beta sheet surrounded by alpha helices. These proteins fulfil a wide range of physiological functions (hydrolases, adhesion molecules, hormone precursors) [Krejci,E., Duval,N., Chatonnet,A., Vincens,P. and Massoulié,J. (1991) Proc. Natl. Acad. Sci. USA , 88, 6647-6651]. ESTHER (for esterases, alpha/beta hydrolase enzymes and relatives) is a database aimed at collecting in one information system, sequence data together with biological annotations and experimental biochemical results related to the structure-function analysis of the enzymes of the family. The major upgrade of the database comes from the use of a new database management system: aCHEdb which uses the ACeDB program designed by Richard Durbin and Jean Thierry-Mieg. It can be found at http://www.ensam.inra.fr/cholinesterase

Animals↗

Chemosensory proteins in the honey bee: Insights from the annotated genome, comparative analyses and expressional profiling.

Small chemosensory proteins (CSPs) belong to a conserved, but poorly understood protein family that has been implicated in transporting chemical stimuli within insect sensilla. However, their expression patterns suggest that these molecules are also critical for other functions including early development. Here we used both bioinformatics and experimental approaches to characterize the CSP gene family in a social insect, the Western honey bee Apis mellifera, and then compared its members to CSPs in other arthropods. The number of CSPs in the honey bee genome (six) is similar to that found in the sequenced dipteran species (four-seven), but is much lower than the number of CSPs in the moth or in the beetle (around 20 each). These differences seem to be the result of lineage specific expansions. Our analysis of CSPs in a number of arthropods reveals a conserved gene family found in both Mandibulates and Chelicerates. Expressional profiling in diverse tissues and throughout development reveals broader than expected patterns of expression with none of the CSPs restricted to the antennae and one found only in the queen ovaries and in embryos. We conclude that CSPs are multifunctional context-dependent proteins involved in diverse cellular processes ranging from embryonic development to chemosensory signal transduction. Some CSPs may function in cuticle synthesis, consistent with their evolutionary origins in the arthropods.

Amino Acid Sequence↗

A microdeletion in Xp11.3 accounts for co-segregation of retinitis pigmentosa and mental retardation in a large kindred.

In a previous report, Aldred et al. [1994] described a 5-generation family in which severe retinitis pigmentosa (RP) co-segregates with mild-moderate mental retardation as an X-linked recessive phenotype mapping to the broad interval between Xp21-q21. We re-examined this family, initially analyzing RP2, a gene in the disease interval that was identified as a cause of RP after the initial report of this family. We found that the male propositus lacked the 5' three exons of RP2 and that RP2 marks the centromeric boundary of a 1.27 Mb deletion that includes two other annotated genes (SLC9A7, CHST7), one predicted transcript encoding a zinc finger protein (FLJ20344) and two highly conserved miRNAs (mir221, mir222). We conclude that this family is segregating a contiguous gene deletion and that the absence of a functional RP2 accounts, at least in part, for the retinal degeneration while deletion of one or more of the other genes is likely responsible for the mental retardation phenotype.

Antiporters↗

Circular RNAs orchestrate integrated post-transcriptional responses to combined heat and drought stress in rice.

Circular RNAs (circRNAs) are emerging post-transcriptional regulators, yet their landscape and functional roles in rice under combined abiotic stress remain largely unexplored. Here, we systematically reanalyzed strand-specific RNA-seq data to characterize circRNAs responsive to simultaneous heat and drought stress. Following quality control, read mapping, and dual-algorithm prediction using CIRI2 and CIRCexplorer2, we identified 208 high-confidence circRNAs distributed across all 12 chromosomes. Comparative profiling revealed 83 circRNAs uniquely expressed in control samples, 51 in stressed samples, and 74 shared between conditions, indicating stress-dependent circularization. Junction-read analysis highlighted a spectrum of circularization strength, ranging from highly abundant circRNAs with dominant junction reads to low-confidence candidates masked by linear transcript background. Genomic annotation showed that circRNAs primarily originated from exonic and intergenic regions, with a pronounced negative-strand bias; several genes generated multiple circRNA isoforms via alternative back-splicing. Functional enrichment of host genes suggested involvement in protein folding, nutrient reservoir activity, RNA degradation, and branched-chain amino acid catabolism, implicating roles in stress adaptation and metabolic regulation. Differential expression analysis identified seven circRNAs specifically induced under combined stress conditions. Network topology analysis pinpointed key miRNAs-including osa-miR414, osa-miR1439, and osa-miR2919-as candidate topological hubs within the predicted network. Their predicted target genes, such as those encoding stress-responsive transcription factors and signaling proteins, suggest potential roles in coordinating post-transcriptional responses to combined stress. Network topology analysis pinpointed key miRNAs-including osa-miR414, osa-miR1439, and osa-miR2919-as candidate topological hubs within the predicted network. Their predicted target genes, such as those encoding stress-responsive transcription factors and signaling proteins, suggest potential roles in coordinating post-transcriptional responses to combined stress. Overall, this study provides a comprehensive map of circRNAs in rice under combined heat and drought stress, suggests their potential as ceRNAs based on predictive analysis, and lays a foundation for future experimental validation of circRNA-mediated regulation.

Oryza↗

Comparative transcriptome analysis provides insights into dorso-ventral color pattern formation of Holothuria edulis.

Animal body color patterns are highly diverse and play critical roles in camouflage, intraspecific communication, and environmental adaptation. Holothuria edulis, an important echinoderm inhabiting tropical waters, exhibits a typical dorsoventral dichromatism. This unique body color difference represents a key phenotypic trait for its habitat adaptation; however, the core differential genes regulating this trait remain to be elucidated. In this study, comparative transcriptome sequencing was performed on the dorsal and ventral body wall tissues of H. edulis, leading to the identification of a number of differentially expressed genes (DEGs), followed by GO functional annotation and KEGG pathway enrichment analysis. GO enrichment analysis indicated that the DEGs were significantly enriched in functional categories such as extracellular region, peptidase inhibitor activity, and tetrapyrrole binding. KEGG pathway analysis further revealed significant enrichment of protein digestion and absorption, the TNF signaling pathway, and cholesterol metabolism. Notably, the pigmentation-related gene FMO2 was highly expressed in the dorsal body wall tissue, whereas cyp1a1, ZIC1, Slc7a11, WNT-1, and ADAMTS20 were highly expressed in the ventral body wall tissue. This study identified DEGs and enriched pathways associated with dorsoventral body color differences in H. edulis, providing new insights into the molecular regulatory mechanisms underlying body color pattern formation. From the perspective of aquaculture applications, body color is one of the important traits affecting the quality and market value of sea cucumber products. Elucidating the molecular mechanisms of body color variation can provide a scientific basis for molecular marker-assisted breeding of superior sea cucumber variety.

Animals↗

Proteome analyses reveal endoplasmic reticulum stress-induced changes in protein abundance associated with Ube2j2 deficiency in human cell culture.

The unfolded protein response (UPR) helps reinstate cellular proteostasis upon an accumulation of misfolded proteins in the endoplasmic reticulum (ER), in part through ER-associated degradation (ERAD). Ube2j2 is an ER-localized E2 ubiquitin-conjugating enzyme that participates in ERAD. We used mass spectrometry analysis of cultured U2OS cells to investigate how the loss of Ube2j2 affects the cellular proteome in response to tunicamycin-induced ER stress. We constructed a network of twelve statistically distinct modules of protein abundance profiles across conditions. We describe the gene ontology annotations for each module along with the "hub gene" proteins whose abundance levels most closely adhere to each module's protein abundance profile. Our analysis identifies known Ube2j2-associated pathways (eg the UPR and ERAD) and cellular functions that were previously unassociated with Ube2j2 (eg RNA metabolism, ER-Golgi transport, and cell-cycle progression). These data are available via ProteomeXchange with identifier PXD076153 and provide avenues for further investigation into the cellular functions of Ube2j2 under basal and ER-stressed conditions.

Humans↗

MutDB services: interactive structural analysis of mutation data.

Non-synonymous single nucleotide polymorphisms (SNPs) and mutations have been associated with human phenotypes and disease. As more and more SNPs are mapped to phenotypes, understanding how these variations affect the function and expression of genes and gene products becomes an important endeavor. We have developed a set of tools to aid in the understanding of how amino acid substitutions affect protein structures. To do this, we have annotated SNPs in dbSNP and amino acid substitutions in Swiss-Prot with protein structural information, if available. We then developed a novel web interface to this data that allows for visualization of the location of these substitutions. We have also developed a web service interface to the dataset and developed interactive plugins for UCSF's Chimera structural modeling tool and PyMOL that integrate our annotations with these sophisticated structural visualization and modeling tools. The web services portal and plugins can be downloaded from http://www.lifescienceweb.org/ and the web interface is at http://www.mutdb.org/.

Amino Acid Substitution↗

k-mer-based Upstream Preprocessing of long reads for Isoform Discovery.

Eukaryotic genes can encode multiple protein isoforms based on alternative splicing of their transcribed regions. Most modern novel isoform discovery methods function by identifying and assembling exon splice junctions from an RNA-seq sample. However, splice junctions can only be accurately annotated with time-intensive dynamic programming alignment. This manuscript introduces KuPID, a method for preprocessing long RNA-seq reads with the goal of better identifying novel isoform transcripts. KuPID utilizes k-mer sketching as a prefilter to quickly pseudo-align reads to known reference isoforms. Full alignment need only then be applied to reads that are most relevant to isoform discovery. Not only does KuPID speed up the discovery pipeline, it also increases downstream accuracy by filtering out extraneous reads. KuPID preprocessing simultaneously increases the f1 accuracy of isoform discovery pipelines by up to 11.6 points while decreasing the runtime by a factor of 2-3×;. An optional mode permits a KuPID sample to be paired with both isoform discovery and transcript quantification.

Journal Article↗

Gene expression in the brain and kidney of rainbow trout in response to handling stress.

BACKGROUND: Microarray technologies are rapidly becoming available for new species including teleost fishes. We constructed a rainbow trout cDNA microarray targeted at the identification of genes which are differentially expressed in response to environmental stressors. This platform included clones from normalized and subtracted libraries and genes selected through functional annotation. Present study focused on time-course comparisons of stress responses in the brain and kidney and the identification of a set of genes which are diagnostic for stress response. RESULTS: Fish were stressed with handling and samples were collected 1, 3 and 5 days after the first exposure. Gene expression profiles were analysed in terms of Gene Ontology categories. Stress affected different functional groups of genes in the tissues studied. Mitochondria, extracellular matrix and endopeptidases (especially collagenases) were the major targets in kidney. Stress response in brain was characterized with dramatic temporal alterations. Metal ion binding proteins, glycolytic enzymes and motor proteins were induced transiently, whereas expression of genes involved in stress and immune response, cell proliferation and growth, signal transduction and apoptosis, protein biosynthesis and folding changed in a reciprocal fashion. Despite dramatic difference between tissues and time-points, we were able to identify a group of 48 genes that showed strong correlation of expression profiles (Pearson r > /0.65/) in 35 microarray experiments being regulated by stress. We evaluated performance of the clone sets used for preparation of microarray. Overall, the number of differentially expressed genes was markedly higher in EST than in genes selected through Gene Ontology annotations, however 63% of stress-responsive genes were from this group. CONCLUSIONS: 1. Stress responses in fish brain and kidney are different in function and time-course. 2. Identification of stress-regulated genes provides the possibility for measuring stress responses in various conditions and further search for the functionally related genes.

Animals↗

MIPS: analysis and annotation of proteins from whole genomes in 2005.

The Munich Information Center for Protein Sequences (MIPS at the GSF), Neuherberg, Germany, provides resources related to genome information. Manually curated databases for several reference organisms are maintained. Several of these databases are described elsewhere in this and other recent NAR database issues. In a complementary effort, a comprehensive set of >400 genomes automatically annotated with the PEDANT system are maintained. The main goal of our current work on creating and maintaining genome databases is to extend gene centered information to information on interactions within a generic comprehensive framework. We have concentrated our efforts along three lines (i) the development of suitable comprehensive data structures and database technology, communication and query tools to include a wide range of different types of information enabling the representation of complex information such as functional modules or networks Genome Research Environment System, (ii) the development of databases covering computable information such as the basic evolutionary relations among all genes, namely SIMAP, the sequence similarity matrix and the CABiNet network analysis framework and (iii) the compilation and manual annotation of information related to interactions such as protein-protein interactions or other types of relations (e.g. MPCDB, MPPI, CYGD). All databases described and the detailed descriptions of our projects can be accessed through the MIPS WWW server (http://mips.gsf.de).

Animals↗

Prediction and overview of the RpoN-regulon in closely related species of the Rhizobiales.

BACKGROUND: In the rhizobia, a group of symbiotic Gram-negative soil bacteria, RpoN (sigma54, sigmaN, NtrA) is best known as the sigma factor enabling transcription of the nitrogen fixation genes. Recent reports, however, demonstrate the involvement of RpoN in other symbiotic functions, although no large-scale effort has yet been undertaken to unravel the RpoN-regulon in rhizobia. We screened two complete rhizobial genomes (Mesorhizobium loti, Sinorhizobium meliloti) and four symbiotic regions (Rhizobium etli, Rhizobium sp. NGR234, Bradyrhizobium japonicum, M. loti) for the presence of the highly conserved RpoN-binding sites. A comparison was also made with two closely related non-symbiotic members of the Rhizobiales (Agrobacterium tumefaciens, Brucella melitensis). RESULTS: A highly specific weight-matrix-based screening method was applied to predict members of the RpoN-regulon, which were stored in a highly annotated and manually curated dataset. Possible enhancer-binding proteins (EBPs) controlling the expression of RpoN-dependent genes were predicted with a profile hidden Markov model. CONCLUSIONS: The methodology used to predict RpoN-binding sites proved highly effective as nearly all known RpoN-controlled genes were identified. In addition, many new RpoN-dependent functions were found. The dependency of several of these diverse functions on RpoN seems species-specific. Around 30% of the identified genes are hypothetical. Rhizobia appear to have recruited RpoN for symbiotic processes, whereas the role of RpoN in A. tumefaciens and B. melitensis remains largely to be elucidated. All species screened possess at least one uncharacterized EBP as well as the usual ones. Lastly, RpoN could significantly broaden its working range by direct interfering with the binding of regulatory proteins to the promoter DNA.

Bacterial Proteins↗

PAHdb: a locus-specific knowledgebase.

PAHdb is an online relational locus-specific "mutation database" (http://www.mcgill.ca/pahdb) for the human phenylalanine hydroxylase gene (symbol PAH) and its associated phenotypes (protein, metabolic, clinical). When combined with associated information (population distribution of allele, haplotype association, etc.) PAHdb functions as a knowledgebase. From the outset, and in the absence of raw data (e.g., sequence gels), PAHdb has instead been an annotated repository of information about mutations maintained by a team of curators. It is also disease-oriented, being focused on a variant phenotype (hyperphenylalaninemia (HPA) and its most important form of disease, phenylketonuria (PKU)) resulting from primary dysfunction of the PAH enzyme (EC 1.14.16.1); it is "patient friendly" in that it contains information for those personally involved with HPA/PKU (MIM# 261600). PAHdb also serves its community through direct interaction.

Alleles↗

The Human Genome Project--an overview.

The human genome sequence will underpin human biology and medicine in the next century, providing a single, essential reference to all genetic information. The international program to determine the complete DNA sequence (3,000 million bases) is well underway. As of January 2000, 50% of the sequence is available in the public domain. A comprehensive working draft is expected this year, and the entire sequence is projected to be finished in 2003. DNA sequencing is carried out on mapped, overlapping bacterial clones of 150-200 kb. The working draft comprises assembled unfinished sequence and is released immediately in the public domain. The draft sequence of each clone is then completed, by closing any remaining gaps and resolving any ambiguities, before the entire sequence is checked, annotated, and submitted to the public databases. The sequence of each clone is finished to an accuracy of >99.99%. The availability of a reference sequence of the genome provides the basis for studying the nature of sequence variation, particularly single nucleotide polymorphisms (SNPs), in human populations. SNP typing is a powerful tool for genetic analysis, and will enable us to uncover the association of loci at specific sites in the genome with many disease traits. SNPs occur at a frequency of approximately 1 SNP/kb throughout the genome when the sequence of any two individuals is compared. Programs to detect and map SNPs in the human genome are underway with the aim of establishing a SNP map of the genome during the next two years. The human genome sequence will provide a complete description of all the genes. Annotation of the sequence with the gene structures is achieved by a combination of computational analysis (predictive and homology-based) and experimental confirmation by cDNA sequencing. Detecting homologies between newly defined gene products and proteins of known function helps to postulate biochemical functions for them, which can then be tested. Establishing the association of specific genes with disease phenotypes by mutation screening, particularly for monogenic disorders, provides further assistance in defining the functions of some gene products, as well as helping to establish the cause of the disease. As our knowledge of gene sequences and sequence variation in populations increases, we will pinpoint more and more of the genes and proteins that are important in common, complex diseases. A more detailed understanding of the function of the human genome will be achieved as we identify sequences that control gene expression. Given the availability of gene sequences, the expression status of genes in particular tissues can be monitored in parallel. By comparing corresponding genomic sequences in different species (for example: man, mouse, chicken, and zebrafish), regions that have been highly conserved during evolution can be identified, many of which reflect conserved functions such as gene regulation. These approaches promise to greatly accelerate our interpretation of the human genome sequence.

Human Genome Project↗

Comparative analysis of expressed sequence tags from different organs of Vitis vinifera L.

Expressed sequence tags (ESTs) are providing a valuable approach to sampling organism-expressed genomes, especially when studying large genomes such as those of many plants. We report on the comparison of 8,647 ESTs generated from six different grape (Vitis vinifera L.) organs: berry, root, leaf, bud, shoot and inflorescence. Clustering and assembly of these ESTs resulted in 4,203 unique sequences and revealed that at this level of EST sampling, each organ shares a low percentage of transcripts with the others. To define organ relationships based on EST counts, we calculated a distance matrix of pairwise correlation coefficients between the libraries which indicated bud, inflorescence and shoot as a group distinct from the other organs considered in this study. A putative function was identified for about 85% of the unique sequences. By assigning them to specific functional classes, we were able to highlight strong differences between organs in the metabolism, protein biosynthesis and photosynthesis categories. This grape EST collection has also proven to be a valuable source for the development of 'functional' simple sequence repeats (SSRs) markers: a total of 405 SSRs have been identified. EST sequences and annotation results have been organised in the IASMA-grape database, freely available at the address http://genomics.iasma.it.

DNA, Complementary↗

Comparative analysis of the expressed genome of the infective juvenile entomopathogenic nematode, Heterorhabditis bacteriophora.

We report the first cDNA-sequencing project of the entomopathogenic nematode, Heterorhabditis bacteriophora. A total of 1246 expressed sequence tags (ESTs) were generated by random sequencing of clones from a cDNA library of the infective juvenile stage. The ESTs were annotated resulting in 1072 useful ESTs that were categorized into functional categories according to Kyoto Encyclopedia of Genes and Genomes. Approximately 459 of 1072 ESTs (43%) had significant similarities to annotated sequences in GenBank. Of these, 417 had significant similarities to the free-living nematode Caenorhanditis elegans proteins. Most ESTs (18%) belonged to the genetic information processing category followed by metabolism (15% ESTs) and environmental information processing (15%) pathways. Several interesting ESTs were found that may have roles in the infectivity and survival of infective juveniles. These included proteases, dauer pathway genes (akt-1, pdk-1 & daf-7) and aging and stress resistance genes such as superoxide dismutase (sod-4), heat shock genes (hsp-4 & hsp-6), and eat genes, and signaling proteins like G-protein coupled receptors, regulators of G-protein signaling (rgs), and serine/threonine kinases. Other interesting ESTs include systemic RNAi defective protein (sid-1), ribonuclease III family members (rnh-2 &rnc) and transposase gene (Tc3A). About 67% of the ESTs did not find matches in any of the searched databases suggesting potentially novel genes in this enomopathogenic nematode. Note: Sequences described in this paper have been deposited in Genbank under the accessions DN 152655-DN 152999, and DN 153000-DN 153726.

3-Phosphoinositide-Dependent Protein Kinases↗

Re-annotation of the genome sequence of Mycobacterium tuberculosis H37Rv.

Original genome annotations need to be regularly updated if the information they contain is to remain accurate and relevant. Here the complete re-annotation of the genome sequence of Mycobacterium tuberculosis strain H37Rv is presented almost 4 years after the first submission. Eighty-two new protein-coding sequences (CDS) have been included and 22 of these have a predicted function. The majority were identified by manual or automated re-analysis of the genome and most of them were shorter than the 100 codon cut-off used in the initial genome analysis. The functional classification of 643 CDS has been changed based principally on recent sequence comparisons and new experimental data from the literature. More than 300 gene names and over 1000 targeted citations have been added and the lengths of 60 genes have been modified. Presently, it is possible to assign a function to 2058 proteins (52% of the 3995 proteins predicted) and only 376 putative proteins share no homology with known proteins and thus could be unique to M. tuberculosis.

Bacterial Proteins↗

Novel transcription factors in human CD34 antigen-positive hematopoietic cells.

Transcription factors (TFs) and the regulatory proteins that control them play key roles in hematopoiesis, controlling basic processes of cell growth and differentiation; disruption of these processes may lead to leukemogenesis. Here we attempt to identify functionally novel and partially characterized TFs/regulatory proteins that are expressed in undifferentiated hematopoietic tissue. We surveyed our database of 15 970 genes/expressed sequence tags (ESTs) representing the normal human CD34(+) cells transcriptosome (http://westsun.hema.uic.edu/cd34.html), using the UniGene annotation text descriptor, to identify genes with motifs consistent with transcriptional regulators; 285 genes were identified. We also extracted the human homologues of the TFs reported in the murine stem cell database (SCdb; http://stemcell.princeton.edu/), selecting an additional 45 genes/ESTs. An exhaustive literature search of each of these 330 unique genes was performed to determine if any had been previously reported and to obtain additional characterizing information. Of the resulting gene list, 106 were considered to be potential TFs. Overall, the transcriptional regulator dataset consists of 165 novel or poorly characterized genes, including 25 that appeared to be TFs. Among these novel and poorly characterized genes are a cell growth regulatory with ring finger domain protein (CGR19, Hs.59106), an RB-associated CRAB repressor (RBAK, Hs.7222), a death-associated transcription factor 1 (DATF1, Hs.155313), and a p38-interacting protein (P38IP, Hs. 171185). The identification of these novel and partially characterized potential transcriptional regulators adds a wealth of information to understanding the molecular aspects of hematopoiesis and hematopoietic disorders.

Amino Acid Motifs↗

An infrastructure for comparative genomics to functionally characterize genes and proteins.

Current genome projects are resulting in a flood of sequence data. The interpretation of these sequences is lagging, and optimized data analysis strategies need to be developed. Much can be learned from comparing different genomes, as genomes of distant organisms may still encode proteins with high sequence similarity. The order of genes (co linearity) in genomes may also be conserved to some extend. We have employed both these observations to create a multi-functional, computational analysis system (genomeSCOUT) which allows for rapid identification and functional characterization of genes and proteins through genome comparison. With a number of independent algorithms, information about different levels of protein homology (concerning e.g. paralogs, orthologs and clusters of orthologous groups, COGs) and gene order is collected and stored in several value added databases. These databases are then used for interactive comparison of genomes and subsequent analysis. The application is based on the well established data integration system SRS. This ensures (1) fast handling of large genomic data sets, (2) straightforward access to a multitude of biological databases, (3) unique linking functions between these databases, (4) highly efficient collection of information on genes and proteins, and 5. fully integrated and user friendly graphical representations of search results. This application can be used for projects as diverse as the correct annotation of genomes, the optimization of (micro) organisms for industrial production, or the identification of drug targets.

Computational Biology↗