Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

The REFOLD database: a tool for the optimization of protein expression and refolding.

A large proportion of proteins expressed in Escherichia coli form inclusion bodies and thus require renaturation to attain a functional conformation for analysis. In this process, identifying and optimizing the refolding conditions and methodology is often rate limiting. In order to address this problem, we have developed REFOLD, a web-accessible relational database containing the published methods employed in the refolding of recombinant proteins. Currently, REFOLD contains >300 entries, which are heavily annotated such that the database can be searched via multiple parameters. We anticipate that REFOLD will continue to grow and eventually become a powerful tool for the optimization of protein renaturation. REFOLD is freely available at http://refold.med.monash.edu.au.

Databases, Protein↗

Average mutual information of coding and noncoding DNA.

One basic problem in the analysis of DNA sequences is the recognition of protein-coding genes. Computer algorithms to facilitate gene identification have become important as genome sequencing projects have turned from mapping to large-scale sequencing, resulting in an exponentially growing number of sequenced nucleotides that await their annotation. Many statistical patterns have been discovered that are different in coding and noncoding DNA, but most of them vary from species to species, and hence require prior training on organism-specific data sets. Here, we investigate if there exist species-independent statistical patterns that are different in coding and noncoding DNA. We introduce an information-theoretic quantity, the average mutual information (AMI), and we find that the probability distribution functions of the AMI are significantly different in coding and noncoding DNA, while they are almost identical for different species. This finding suggests that the AMI might be useful for the recognition of protein-coding regions in genomes for which training sets do not exist.

Algorithms↗

A gene expression map for the euchromatic genome of Drosophila melanogaster.

We used a maskless photolithography method to produce DNA oligonucleotide microarrays with unique probe sequences tiled throughout the genome of Drosophila melanogaster and across predicted splice junctions. RNA expression of protein coding and nonprotein coding sequences was determined for each major stage of the life cycle, including adult males and females. We detected transcriptional activity for 93% of annotated genes and RNA expression for 41% of the probes in intronic and intergenic sequences. Comparison to genome-wide RNA interference data and to gene annotations revealed distinguishable levels of expression for different classes of genes and higher levels of expression for genes with essential cellular functions. Differential splicing was observed in about 40% of predicted genes, and 5440 previously unknown splice forms were detected. Genes within conserved regions of synteny with D. pseudoobscura had highly correlated expression; these regions ranged in length from 10 to 900 kilobase pairs. The expressed intergenic and intronic sequences are more likely to be evolutionarily conserved than nonexpressed ones, and about 15% of them appear to be developmentally regulated. Our results provide a draft expression map for the entire nonrepetitive genome, which reveals a much more extensive and diverse set of expressed sequences than was previously predicted.

Algorithms↗

Whole-genome analysis of Brevibacterium sanguinis AZMABM HM27: a bacterial isolate from the sea anemone Radianthus magnifica and exhibiting promising multi-therapeutic properties.

BACKGROUND: The marine anemone Radianthus magnifica harbors symbiotic microbes with promising biomedical potential, yet their diversity and therapeutic properties remain underexplored. This study aimed to characterize a symbiotic bacterium isolated from R. magnifica collected from Samalona Island, Indonesia, and to evaluate its multi-therapeutic potential. METHODS: Strain AZMABM HM27 was characterized using whole-genome sequencing, functional annotation, biosynthetic gene cluster prediction, molecular docking, and in vitro bioactivity assays. RESULTS: Phylogenetic and genome-based analyses confirmed AZMABM HM27 as Brevibacterium sanguinis, with an OrthoANI value of 97.37% and a dDDH value of 76.50% against the type strain. The genome comprises a 3,834,082 bp chromosome encoding 3,362 protein-coding genes, including 95 genes involved in secondary metabolite biosynthesis. Five biosynthetic gene clusters were predicted, including those associated with ectoine, terpene, and siderophore production. The crude extract demonstrated antioxidant activity (IC₅₀ = 0.87 mg/mL), anti-inflammatory activity (up to 60% inhibition), antidiabetic activity through α-glucosidase inhibition (up to 40% inhibition), and dose-dependent antiproliferative activity against MCF-7 breast cancer cells (74.10% viability at 1 mg/mL). Molecular docking identified a lead compound, 8,9,9,10,10,11-hexafluoro-4,4-dimethyl-3,5-dioxatetracyclo [5.4.1.0(2,6)0.0(8,11)] dodecane, with strong binding affinities to selected therapeutic targets. CONCLUSIONS: B. sanguinis AZMABM HM27 represents a marine symbiotic strain associated with R. magnifica and a promising source of bioactive compounds with antioxidant, anti-inflammatory, antidiabetic, and antiproliferative potential. Further purification, structural elucidation, and in vivo studies are warranted to validate its therapeutic potential.

Animals↗

Human proteomic databases: a powerful resource for functional genomics in health and disease.

Decoding of the genome information in terms of regulation and function will be the next great challenge in the life sciences in this millennium and indeed, today we are experiencing a rapid explosion of technology for the high throughput expression analysis of genes and their products (functional genomics). In particular, the field of proteomics is booming as proteins are often the functional molecules and represent important targets for the pharmaceutical industry. The proteomic technology is complex, and comprises a plethora of state-of-the-art techniques to resolve, identify and detect their interacting partners, as well as to store and communicate protein information in comprehensive two-dimensional polyacrylamide gel electrophoresis (2D PAGE) databases. Besides annotating the genome, these databases will offer a global approach to the study of gene expression both in health and disease. Here, we review the current status of human 2D PAGE databases that we are systematically constructing for the study of bladder cancer and skin ageing.

Databases, Protein↗

Altered expression of novel genes in the cerebral cortex following experimental brain injury.

Damage to the cerebral cortex results in neurological impairments such as motor, attention, memory and executive dysfunctions. To examine the molecular mechanisms contributing to these deficits, mRNA expression was profiled using high-density cDNA microarray hybridization after experimental cortical impact injury in mice. The mRNA levels at 2 h, 6 h, 24 h, 3 days and 14 days after injury were compared with those of control animals. This revealed 86 annotated genes and 24 expression sequence tags (ESTs) as being differentially expressed with a 1.5-fold or greater change. Quantitative real-time PCR analysis was used to independently verify these results for selected genes. Seven functional classes of genes were found to be altered following injury, including transcription factors, signal transduction genes and inflammatory proteins. While a few of these genes have been previously reported to be differentially regulated following injury, the most of the genes have not been previously implicated in traumatic brain injury (TBI) pathophysiology. For example, consistent with previous reports, the transcription factor c-jun and the neurotrophic factor bdnf mRNA levels were altered as a result of TBI. Among the novel genes, the mRNA levels for the high mobility group protein 1 (hmg-1), the regulator of G-protein signaling 2 (rgs-2), the transforming growth factor beta inducible early growth response (tieg), the inhibitor of DNA binding 3 (id3), and the heterogeneous nuclear ribonucleoprotein H (hnrnp h) were changed following injury. The functional significance of these genes in neurite outgrowth, neuronal regeneration, and plasticity following injury are discussed.

Animals↗

Prediction of the coupling specificity of G protein coupled receptors to their G proteins.

UNLABELLED: G protein coupled receptors (GPCRs) are found in great numbers in most eukaryotic genomes. They are responsible for sensing a staggering variety of structurally diverse ligands, with their activation resulting in the initiation of a variety of cellular signalling cascades. The physiological response that is observed following receptor activation is governed by the guanine nucleotide-binding proteins (G proteins) to which a particular receptor chooses to couple. Previous investigations have demonstrated that the specificity of the receptor-G protein interaction is governed by the intracellular domains of the receptor. Despite many studies it has proven very difficult to predict de novo, from the receptor sequence alone, the G proteins to which a GPCR is most likely to couple. We have used a data-mining approach, combining pattern discovery with membrane topology prediction, to find patterns of amino acid residues in the intracellular domains of GPCR sequences that are specific for coupling to a particular functional class of G proteins. A prediction system was then built, being based upon these discovered patterns. We can report this approach was successful in the prediction of G protein coupling specificity of unknown sequences. Such predictions should be of great use in providing in silico characterisation of newly cloned receptor sequences and for improving the annotation of GPCRs stored in protein sequence databases. AVAILABILITY: http://www.ebi.ac.uk/~croning/coupling.html.

Amino Acid Sequence↗

HUNT: launch of a full-length cDNA database from the Helix Research Institute.

The Helix Research Institute (HRI) in Japan is releasing 4356 HUman Novel Transcripts and related information in the newly established HUNT database. The institute is a joint research project principally funded by the Japanese Ministry of International Trade and Industry, and the clones were sequenced in the governmental New Energy and Industrial Technology Development Organization (NEDO) Human cDNA Sequencing Project. The HUNT database contains an extensive amount of annotation from advanced analysis and represents an essential bioinformatics contribution towards understanding of the gene function. The HRI human cDNA clones were obtained from full-length enriched cDNA libraries constructed with the oligo-capping method and have resulted in novel full-length cDNA sequences. A large fraction has little similarity to any proteins of known function and to obtain clues about possible function we have developed original analysis procedures. Any putative function deduced here can be validated or refuted by complementary analysis results. The user can also extract information from specific categories like PROSITE patterns, PFAM domains, PSORT localization, transmembrane helices and clones with GENIUS structure assignments. The HUNT database can be accessed at http://www.hri.co.jp/HUNT.

Computational Biology↗

GABI-Kat SimpleSearch: an Arabidopsis thaliana T-DNA mutant database with detailed information for confirmed insertions.

Insertional mutagenesis approaches, especially by T-DNA, play important roles in gene function studies of the model plant Arabidopsis thaliana. GABI-Kat SimpleSearch (http://www.GABI-Kat.de) is a Flanking Sequence Tag (FST)-based database for T-DNA insertion mutants generated by the GABI-Kat project. Currently, the database contains >108,000 mapped FSTs from approximately 64,000 lines which cover 64% of all annotated A.thaliana protein-coding genes. The web interface allows searching for relevant insertions by gene code, keyword, line identifier, GenBank accession number of the FST, and also by BLAST. A graphic display of the genome region around the gene or the FST assists users to select insertion lines of their interests. About 3500 insertions were confirmed in the offspring of the plant from which the original FST was generated, and the seeds of these lines are available from the Nottingham Arabidopsis Stock Centre. The database now also contains additional information such as segregation data, gene-specific primers and confirmation sequences. This information not only helps users to evaluate the usefulness of the mutant lines, but also covers a big part of the molecular characterization of the insertion alleles.

Arabidopsis↗

Transcriptome profiling of vertical stem segments provides insights into the genetic regulation of secondary growth in hybrid aspen trees.

In order to better understand the genetic regulation of secondary growth in hybrid aspen (Populus tremula L.xP. alba L.), we carried out a series of cDNA-amplified fragment length polymorphism (AFLP)-based transcriptome analyses in vertical stem segments that represent a gradient of developmental stages with regard to secondary growth. This approach allowed us to screen >80% of the transcriptome expressed in six samples and identify genes differentially expressed with the progress of secondary growth, in a tissue-specific manner. Of the 76,800 transcript-derived fragments (TDFs) analyzed, 271 TDFs were selected and sequenced based on their differential expression patterns. Many of the xylem-up-regulated genes were involved in cell wall and lignin biosynthesis, while the bark-up-regulated genes had diverse functional roles. About 25% of the xylem-up-regulated TDFs analyzed were involved in the phenylpropanoid biosynthesis pathway, which produces the cell wall polymer lignin and various wood extractives. In addition, many of the TDFs showing secondary xylem-specific expression were annotated as genes not previously reported in Populus, including novel cell death proteins, cytoskeleton-interacting proteins, transporters and putative transcription factors.

Blotting, Northern↗

Functional genomics by integrated analysis of metabolome and transcriptome of Arabidopsis plants over-expressing an MYB transcription factor.

The integration of metabolomics and transcriptomics can provide precise information on gene-to-metabolite networks for identifying the function of unknown genes unless there has been a post-transcriptional modification. Here, we report a comprehensive analysis of the metabolome and transcriptome of Arabidopsis thaliana over-expressing the PAP1 gene encoding an MYB transcription factor, for the identification of novel gene functions involved in flavonoid biosynthesis. For metabolome analysis, we performed flavonoid-targeted analysis by high-performance liquid chromatography-mass spectrometry and non-targeted analysis by Fourier-transform ion-cyclotron mass spectrometry with an ultrahigh-resolution capacity. This combined analysis revealed the specific accumulation of cyanidin and quercetin derivatives, and identified eight novel anthocyanins from an array of putative 1800 metabolites in PAP1 over-expressing plants. The transcriptome analysis of 22,810 genes on a DNA microarray revealed the induction of 38 genes by ectopic PAP1 over-expression. In addition to well-known genes involved in anthocyanin production, several genes with unidentified functions or annotated with putative functions, encoding putative glycosyltransferase, acyltransferase, glutathione S-transferase, sugar transporters and transcription factors, were induced by PAP1. Two putative glycosyltransferase genes (At5g17050 and At4g14090) induced by PAP1 expression were confirmed to encode flavonoid 3-O-glucosyltransferase and anthocyanin 5-O-glucosyltransferase, respectively, from the enzymatic activity of their recombinant proteins in vitro and results of the analysis of anthocyanins in the respective T-DNA-inserted mutants. The functional genomics approach through the integration of metabolomics and transcriptomics presented here provides an innovative means of identifying novel gene functions involved in plant metabolism.

Arabidopsis↗

Possible linking and treatment between Parkinson's disease and inflammatory bowel disease: a study of Mendelian randomization based on gut-brain axis.

BACKGROUND: Mounting evidence suggests that Parkinson's disease (PD) and inflammatory bowel disease (IBD) are closely associated and becoming global health burdens. However, the causal relationships and common pathogeneses between them are uncertain. Furthermore, they are uncurable. Thus, we aimed to identify the causal relationships and novel therapeutic targets shared between them based on their common pathophysiological mechanisms in gut-brain-axis (GBA). METHODS: A meta-analysis on bidirectional Mendelian randomization (MR) utilizing various datasets was performed to estimate their causal relationship. Then, pleiotropic analysis under the composite null hypothesis (PLACO) with functional mapping combined with annotation of genetic associations (FUMA) analysis were conducted to identify pleiotropic genes. Next, blood, brain and intestine expression quantitative trait locus (eQTL) were taken to perform drug-target MR finding common causal genes in two diseases. Colocalization analysis ensured the eQTLs of corresponding gene colocalized with disease. Enrichment analysis and protein‒protein interaction (PPI) network were done to explore common pathogenesis pathways. Genes passed all analysis were regarded as drug targets. RESULTS: Our MR meta-analysis revealed the bidirectional causal relationship between diseases, with combined ORs for PD on IBD, CD, UC (1.050 [95% CI 1.014-1.086], 1.044 [95% CI 0.995-1.095], 1.063 [95% CI 1.016-1.120]); for IBD, CD, UC on PD (1.003 [95% CI 0.973-1.034], 1.035 [95% CI 1.004-1.067], 1.008 [95% CI 0.977-1.040]). Overall, 277, 216 and 201 genes were identified as pleiotropic genes between PD and IBD, CD, UC. Total of 733 genes were classified as tier 3 (found in only one tissue) druggable targets, 57 as tier 2 (found in two tissues, 51 protein-coding genes) and 9 as tier 3 (found in three tissues). Among 60 protein-coding druggable targets over tier 2, 18 overlapped with pleiotropic genes and enriched in mitochondria, antigen presentation, processing and immune cell regulation pathways. Three druggable genes (LRRK2, RAB29 and HLA-DQA2) passed colocalization analysis. LRRK2 and RAB29 were reported to be pleiotropic genes, and RAB29 and HLA-DQA2 were reported for the first time as potential drug targets. CONCLUSIONS: This study established a reliable causal relationship, possible shared drug targets and common pathogenesis pathways of two diseases, which had important implications for intervention and treatment of two diseases simultaneously.

Humans↗

The proteins of human chromosome 21.

Recent genomic sequence annotation suggests that the long arm of human chromosome 21 encodes more than 400 genes. Because there is no evidence to exclude any significant segment of 21 q from containing genes relevant to the Down syndrome (DS) cognitive phenotype, all genes in this entire set must be considered as candidates. Only a subset, however, is likely to make critical contributions. Determining which these are is both a major focus in biology and a critical step in efficient development of therapeutics. The subtle molecular abnormality in DS, the 50% increase in chromosome 21 gene expression, presents significant challenges for researchers in detection and quantitation. Another challenge is the current limitation in understanding gene functions and in interpreting biological characteristics. Here, we review information on chromosome 21-encoded proteins compiled from the literature and from genomics and proteomics databases. For each protein, we summarize their evolutionary conservation, the complexity of their known protein interactions and their level of expression in brain, and discuss the implications and limitations of these data. For a subset, we discuss neurologically relevant phenotypes of mouse models that include knockouts, mutations, or overexpression. Lastly, we highlight a small number of genes for which recent evidence suggests a function in biochemical/cellular pathways that are relevant to cognition. Until knowledge deficits are overcome, we suggest that effective development of gene-phenotype correlations in DS requires a serious and continuous effort to assimilate broad categories of information on chromosome 21 genes, plus the creation of more versatile mouse models.

Animals↗

Large-scale statistical analysis of secondary xylem ESTs in pine.

A computational analysis of pine transcripts was conducted to contribute to the functional annotation of conifer sequences. A statistical analysis of expressed sequential tags(ESTs) belonging the 7732 contigs in the TIGR Pinus Gene Index (PGI1.0) identified 260 differentially represented gene sequences across six cDNA libraries from loblolly pine secondary xylem. Cluster analysis of this subset of contigs resulted in five groups representing genes preferentially represented in one of the xylem samples (compression wood, plannings, root xylem, latewood) and one group containing mostly genes simultaneously present in compression and side wood libraries. To complement the sequence annotation, 27 cDNA clones representing selected transcripts were completely sequenced. Several genes were identified that could represent putative markers for xylem from different organs, at different stages of development. Several sequences encoding regulatory proteins were over-represented in root xylem as opposed to the other xylem samples. Some of them belonged to known families of plant transcription factors, but two genes were previously uncharacterized in plants. One transcript was homologous to the gene encoding the Smad4 interacting factor, a key co-activator in TGFbeta (transforming growth factor) signalling in animals. Thus, the digital analysis of pine ESTs highlighted a putative gene function of potentially broad interest but that has yet to be investigated in plants. More generally, this study showed that the application of numerical approaches to EST databases should be helpful in establishing priorities among genes to consider for targeted functional studies. Thus, we illustrated the potential of extracting information from conifer sequences already accessible through well-structured public databases.

Amino Acid Sequence↗

IntNetDB v1.0: an integrated protein-protein interaction network database generated by a probabilistic model.

BACKGROUND: Although protein-protein interaction (PPI) networks have been explored by various experimental methods, the maps so built are still limited in coverage and accuracy. To further expand the PPI network and to extract more accurate information from existing maps, studies have been carried out to integrate various types of functional relationship data. A frequently updated database of computationally analyzed potential PPIs to provide biological researchers with rapid and easy access to analyze original data as a biological network is still lacking. RESULTS: By applying a probabilistic model, we integrated 27 heterogeneous genomic, proteomic and functional annotation datasets to predict PPI networks in human. In addition to previously studied data types, we show that phenotypic distances and genetic interactions can also be integrated to predict PPIs. We further built an easy-to-use, updatable integrated PPI database, the Integrated Network Database (IntNetDB) online, to provide automatic prediction and visualization of PPI network among genes of interest. The networks can be visualized in SVG (Scalable Vector Graphics) format for zooming in or out. IntNetDB also provides a tool to extract topologically highly connected network neighborhoods from a specific network for further exploration and research. Using the MCODE (Molecular Complex Detections) algorithm, 190 such neighborhoods were detected among all the predicted interactions. The predicted PPIs can also be mapped to worm, fly and mouse interologs. CONCLUSION: IntNetDB includes 180,010 predicted protein-protein interactions among 9,901 human proteins and represents a useful resource for the research community. Our study has increased prediction coverage by five-fold. IntNetDB also provides easy-to-use network visualization and analysis tools that allow biological researchers unfamiliar with computational biology to access and analyze data over the internet. The web interface of IntNetDB is freely accessible at http://hanlab.genetics.ac.cn/IntNetDB.htm. Visualization requires Mozilla version 1.8 (or higher) or Internet Explorer with installation of SVGviewer.

Algorithms↗

The mighty microproteins: from versatile cellular regulators to precision medicine therapeutics.

Microproteins, are tiny proteins encoded by small open reading frame (sORF), translation of these non-canonical open reading frames (ncORFs) has been implicated in diverse biological processes and diseases. This review summarizes recent developments in the discovery, biogenesis, and functional characterization of microproteins, and their involvement in various disease, with special focus on their roles in cancer, cardiovascular, metabolic, neurodegenerative and immune-related disorders. We emphasize the regulation of key cellular pathways by microproteins, including mitochondrial homeostasis, apoptosis, metabolic reprogramming, and immune signaling, all of which affect disease initiation and progression. Emerging evidence also supports their potential as disease biomarkers and therapeutic candidates for precision medicine. Finally, the review critically discusses the current challenges including discrepancies in microprotein annotation, the limitations of ribosome profiling and proteogenomic approaches, the gap between computationally predicted and experimentally validated microproteins, and the need for rigorous orthogonal validation by means of CRISPR-based genome editing, ribosome release assays, mutational analysis, high-resolution mass spectrometry, and functional studies. Finally, we review recent development of AI-assisted ORF prediction, single-cell translatomics, spatial proteomics, and integrated multi-omics as emerging technologies reshaping. Microprotein discovery and functional annotation. Finally, we discuss the translational potential of microproteins and highlight the remaining challenges to clinical application, including peptide stability, pharmacokinetics, tissue-specific delivery, immunogenicity, and the need for rigorous preclinical and clinical validation. Together, this review provides an updated and critical overview of the rapidly evolving microprotein field and highlights future research priorities for translating these molecules into clinically useful biomarkers and precision therapeutics.

Microproteins↗

Neuroendocrinology of protochordates: insights from Ciona genomics.

The genome for two species of Ciona is available making these tunicates excellent models for studies on the evolution of the chordates. In this review most of the data is from Ciona intestinalis, as the annotation of the C. savignyi genome is not yet available. The phylogenetic position of tunicates at the origin of the chordates and the nature of the genome before expansion in vertebrates allows tunicates to be used as a touchstone for understanding genes that either preceded or arose in vertebrates. A comparison of Ciona, a sea squirt, to other model organisms such as a nematode, fruit fly, zebrafish, frog, chicken and mouse shows that Ciona has many useful traits including accessibility for embryological, lineage tracing, forward genetics, and loss- or gain-of-function experiments. For neuroendocrine studies, these traits are important for determining gene function, whereas the availability of the genome is critical for identification of ligands, receptors, transcription factors and signaling pathways. Four major neurohormones and their receptors have been identified by cloning and to some extent by function in Ciona: gonadotropin-releasing hormone, insulin, insulin-like growth factor, and cionin, a member of the CCK/gastrin family. The simplicity of tunicates should be an advantage in searching for novel functions for these hormones. Other neuroendocrine components that have been annotated in the genome are a multitude of receptors, which are available for cloning, expression and functional studies.

Animals↗

FlyBase: anatomical data, images and queries.

FlyBase (http://flybase.org/) is a database of genetic and genomic data on the model organism Drosophila melanogaster and the entire insect family Drosophilidae. The FlyBase Consortium curates, annotates, integrates and maintains a wide variety of data within this domain. Access to the data is provided through graphical and textual user interfaces tailored to particular types of data. FlyBase data types include maps at the cytological, genetic and sequence levels, genes and alleles including their products, functions, expression patterns, mutant phenotypes and genetic interactions as well as aberrant chromosomes, annotated genomes, genetic stock collections, transposons, transgene constructs and insertions, anatomy and images, bibliographic data, and community contact information.

Animals↗