Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Super paramagnetic clustering of protein sequences.

BACKGROUND: Detection of sequence homologues represents a challenging task that is important for the discovery of protein families and the reliable application of automatic annotation methods. The presence of domains in protein families of diverse function, inhomogeneity and different sizes of protein families create considerable difficulties for the application of published clustering methods. RESULTS: Our work analyses the Super Paramagnetic Clustering (SPC) and its extension, global SPC (gSPC) algorithm. These algorithms cluster input data based on a method that is analogous to the treatment of an inhomogeneous ferromagnet in physics. For the SwissProt and SCOP databases we show that the gSPC improves the specificity and sensitivity of clustering over the original SPC and Markov Cluster algorithm (TRIBE-MCL) up to 30%. The three algorithms provided similar results for the MIPS FunCat 1.3 annotation of four bacterial genomes, Bacillus subtilis, Helicobacter pylori, Listeria innocua and Listeria monocytogenes. However, the gSPC covered about 12% more sequences compared to the other methods. The SPC algorithm was programmed in house using C++ and it is available at http://mips.gsf.de/proj/spc. The FunCat annotation is available at http://mips.gsf.de. CONCLUSION: The gSPC calculated to a higher accuracy or covered a larger number of sequences than the TRIBE-MCL algorithm. Thus it is a useful approach for automatic detection of protein families and unsupervised annotation of full genomes.

Algorithms↗

The OtsAB pathway is essential for trehalose biosynthesis in Mycobacterium tuberculosis.

The disaccharide trehalose is the major free sugar in the cytoplasm of mycobacteria; it is a constituent of cell wall glycolipids, and it plays a role in mycolic acid transport during cell wall biogenesis. The pleiotropic role of trehalose in the biology of Mycobacterium tuberculosis and its absence from mammalian cells suggests that its biosynthesis may provide a useful target for novel drugs. However, there are three potential pathways for trehalose biosynthesis in M. tuberculosis, and the aim of the present study was to introduce mutations into each of the pathways to determine whether or not they are functionally redundant. The results show that the OtsAB pathway, which generates trehalose from glucose and glucose-6-phosphate, is the dominant pathway required for M. tuberculosis growth in laboratory culture and for virulence in a mouse model. Of the two otsB homologues annotated in the genome sequence of M. tuberculosis, only OtsB2 (Rv3372) has a functional role in the pathway. OtsB2, trehalose-6-phosphate phosphatase, is strictly essential for growth and provides a tractable target for high throughput screening. Inactivation of the TreYZ pathway, which can generate trehalose from alpha-1,4-linked glucose polymers, had no effect on the growth of M. tuberculosis in vitro or in mice. Deletion of the treS gene altered the late stages of pathogenesis of M. tuberculosis in mice, significantly increasing the time to death in a chronic infection model. Because the TreS enzyme catalyzes the interconversion of trehalose and maltose, the mouse phenotype could reflect either a requirement for synthesis of additional trehalose or, conversely, a requirement for breakdown of stored trehalose to liberate free glucose.

Animals↗

Total sequence decomposition distinguishes functional modules, "molegos" in apurinic/apyrimidinic endonucleases.

BACKGROUND: Total sequence decomposition, using the web-based MASIA tool, identifies areas of conservation in aligned protein sequences. By structurally annotating these motifs, the sequence can be parsed into individual building blocks, molecular legos ("molegos"), that can eventually be related to function. Here, the approach is applied to the apurinic/apyrimidinic endonuclease (APE) DNA repair proteins, essential enzymes that have been highly conserved throughout evolution. The APEs, DNase-1 and inositol 5'-polyphosphate phosphatases (IPP) form a superfamily that catalyze metal ion based phosphorolysis, but recognize different substrates. RESULTS: MASIA decomposition of APE yielded 12 sequence motifs, 10 of which are also structurally conserved within the family and are designated as molegos. The 12 motifs include all the residues known to be essential for DNA cleavage by APE. Five of these molegos are sequentially and structurally conserved in DNase-1 and the IPP family. Correcting the sequence alignment to match the residues at the ends of two of the molegos that are absolutely conserved in each of the three families greatly improved the local structural alignment of APEs, DNase-1 and synaptojanin. Comparing substrate/product binding of molegos common to DNase-1 showed that those distinctive for APEs are not directly involved in cleavage, but establish protein-DNA interactions 3' to the abasic site. These additional bonds enhance both specific binding to damaged DNA and the processivity of APE1. CONCLUSION: A modular approach can improve structurally predictive alignments of homologous proteins with low sequence identity and reveal residues peripheral to the traditional "active site" that control the specificity of enzymatic activity.

Amino Acid Motifs↗

Inter-genomic displacement via lateral gene transfer of bacterial trp operons in an overall context of vertical genealogy.

BACKGROUND: The growing conviction that lateral gene transfer plays a significant role in prokaryote genealogy opens up a need for comprehensive evaluations of gene-enzyme systems on a case-by-case basis. Genes of tryptophan biosynthesis are frequently organized as whole-pathway operons, an attribute that is expected to facilitate multi-gene transfer in a single step. We have asked whether events of lateral gene transfer are sufficient to have obscured our ability to track the vertical genealogy that underpins tryptophan biosynthesis. RESULTS: In 47 complete-genome Bacteria, the genes encoding the seven catalytic domains that participate in primary tryptophan biosynthesis were distinguished from any paralogs or xenologs engaged in other specialized functions. A reliable list of orthologs with carefully ascertained functional roles has thus been assembled and should be valuable as an annotation resource. The protein domains associated with primary tryptophan biosynthesis were then concatenated, yielding single amino-acid sequence strings that represent the entire tryptophan pathway. Lateral gene transfer of several whole-pathway trp operons was demonstrated by use of phylogenetic analysis. Lateral gene transfer of partial-pathway trp operons was also shown, with newly recruited genes functioning either in primary biosynthesis (rarely) or specialized metabolism (more frequently). CONCLUSIONS: (i) Concatenated tryptophan protein trees are congruent with 16S rRNA subtrees provided that the genomes represented are of sufficiently close phylogenetic spacing. There are currently seven tryptophan congruency groups in the Bacteria. Recognition of a succession of others can be expected in the near future, but ultimately these should coalesce to a single grouping that parallels the 16S rRNA tree (except for cases of lateral gene transfer). (ii) The vertical trace of evolution for tryptophan biosynthesis can be deduced. The daunting complexities engendered by paralogy, xenology, and idiosyncrasies of nomenclature at this point in time have necessitated an expert-assisted manual effort to achieve a correct analysis. Once recognized and sorted out, paralogy and xenology can be viewed as features that enrich evolutionary histories.

Bacteria↗

KOBAS server: a web-based platform for automated annotation and pathway identification.

There is an increasing need to automatically annotate a set of genes or proteins (from genome sequencing, DNA microarray analysis or protein 2D gel experiments) using controlled vocabularies and identify the pathways involved, especially the statistically enriched pathways. We have previously demonstrated the KEGG Orthology (KO) as an effective alternative controlled vocabulary and developed a standalone KO-Based Annotation System (KOBAS). Here we report a KOBAS server with a friendly web-based user interface and enhanced functionalities. The server can support input by nucleotide or amino acid sequences or by sequence identifiers in popular databases and can annotate the input with KO terms and KEGG pathways by BLAST sequence similarity or directly ID mapping to genes with known annotations. The server can then identify both frequent and statistically enriched pathways, offering the choices of four statistical tests and the option of multiple testing correction. The server also has a 'User Space' in which frequent users may store and manage their data and results online. We demonstrate the usability of the server by finding statistically enriched pathways in a set of upregulated genes in Alzheimer's Disease (AD) hippocampal cornu ammonis 1 (CA1). KOBAS server can be accessed at http://kobas.cbi.pku.edu.cn.

Alzheimer Disease↗

Prediction of orthologous relationship by functionally important sites.

Making accurate functional predictions plays an important role in the era of proteomics. Reliable functional information can be extracted from orthologs in other species when annotating an unknown gene. Here a site-based approach called PORFIS is proposed to predict orthologous relationship. When applied to the bacterial transcription factor PurR/LacI family and the protein kinase AGC family, our method was able to identify, with few false positives, the important sites that agree with those verified by biological experiments. We also tested it on the alpha-proteasome family, the glycoprotein hormone family and the growth hormone family to demonstrate its ability to predict orthologous relationship. Compared with other prediction methods based on phylogenetic analysis or hidden Markov models, PORFIS not only has competitive prediction accuracy, but also provides valuable biological information of functionally important sites associated with orthologs which can be further studied in biological experiments.

Humans↗

Two-dimensional gel electrophoresis maps of the proteome and phosphoproteome of primitively cultured rat mesangial cells.

Mesangial cells (MC) play an important role in maintaining the structure and function of the glomerulus. The proliferation of MC is a prominent feature of many kinds of glomerular disease. The first reference 2-DE maps of rat mesangial cells (RMC), stained with silver staining or Pro-Q Diamond dye, have been established here to describe the proteome and phosphoproteome of RMC, respectively. A total of 157 selected protein spots, corresponding to 118 unique proteins, have been identified by MALDI-TOF-MS or LC-ESI-IT-MS/MS, in which 37 protein spots representing 28 unique proteins have also been stained with Pro-Q Diamond, indicating that they are in phosphorylated forms. All the identified proteins were bioinformatically annotated in detail according to their physiochemical characteristics, subcellular location, and function. Most of the separated or identified protein spots are distributed in the area of mass 10-70 kDa and pI 5.0-8.0. The identified proteins include mainly cytoplasmic and nuclear proteins and some mitochondrial, endoplasmic reticulum, and membrane proteins. These proteins are classified into different functional groups such as structure and mobility proteins (21.2%), metabolic enzymes (16.9%), protein folding and metabolism proteins (13.6%), signaling proteins (14.4%), heat-shock proteins (7.6%), and other functional proteins (12.7%). While structure and mobility proteins are mostly represented by protein spots with high abundance, signaling proteins are mostly represented by protein spots with relatively low abundance. Such a 2-DE database for RMC, especially with many signaling proteins and phosphoproteins characterized, will provide a valuable resource for comparative proteomics analysis of normal and pathologic conditions affecting MC function or pathologic progress.

Animals↗

Analysis of expressed sequence tags from Musa acuminata ssp. burmannicoides, var. Calcutta 4 (AA) leaves submitted to temperature stresses.

In order to discover genes expressed in leaves of Musa acuminata ssp. burmannicoides var. Calcutta 4 (AA), from plants submitted to temperature stress, we produced and characterized two full-length enriched cDNA libraries. Total RNA from plants subjected to temperatures ranging from 5 degrees C to 25 degrees C and from 25 degrees C to 45 degrees C was used to produce a COLD and a HOT cDNA library, respectively. We sequenced 1,440 clones from each library. Following quality analysis and vector trimming, we assembled 2,286 sequences from both libraries into 1,019 putative transcripts, consisting of 217 clusters and 802 singletons, which we denoted Musa acuminata assembled expressed sequence tagged (EST) sequences (MaAES). Of these MaAES, 22.87% showed no matches with existing sequences in public databases. A global analysis of the MaAES data set indicated that 10% of the sequenced cDNAs are present in both cDNA libraries, while 42% and 48% are present only in the COLD or in the HOT libraries, respectively. Annotation of the MaAES data set categorized them into 22 functional classes. Of the 2,286 high-quality sequences, 715 (31.28%) originated from full-length cDNA clones and resulted in a set of 149 genes.

Base Sequence↗

Sequencing of three lambda clones from the genome of alkaliphilic Bacillus sp. strain C-125.

The nucleotide sequences of three independent fragments (designated no. 3, 4, and 9; each 15-20 kb in size) of the genome of alkaliphilic Bacillus sp. C-125 cloned in a lambda phage vector have been determined. Thirteen putative open reading frames (ORFs) were identified in sequenced fragment no. 3 and 11 ORFs were identified in no. 4. Twenty ORFs were also identified in fragment no. 9. All putative ORFs were analyzed in comparison with the BSORF database and non-redundant protein databases. The functions of 5 ORFs in fragment no. 3 and 3 ORFs in fragment no. 4 were suggested by their significant similarities to known proteins in the database. Among the 20 ORFs in fragment no. 9, the functions of 11 ORFs were similarly suggested. Most of the annotated ORFs in the DNA fragments of the genome of alkaliphilic Bacillus sp. C-125 were conserved in the Bacillus subtilis genome. The organization of ORFs in the genome of strain C-125 was found to differ from the order of genes in the chromosome of B. subtilis, although some gene clusters (ydh, yqi, yer, and yts) were conserved as operon units the same as in B. subtilis.

Bacillus↗

Biological fingerprinting analysis of the interactome of a kinase inhibitor in human plasma by a chemiproteomic approach.

In this study, a gel free chemiproteomic method based on chromatography was developed and applied for the biological fingerprinting analysis of complex biological system. p-Aminobenzamidine (ABA), an inhibitor of trypsin-like serine proteases, was immobilized for characterizing their interacting proteins in human plasma. By the proteomic analysis method, 214 proteins were identified with obvious affinity to the immobilized ABA. By searching the sequences of above proteins with consensus patterns of the two active sites, seven proteins belong to trypsin-like serine protease group were found. Based on the Gene Ontology annotation, the identified trypsin-like serine proteases have the function of catalytic activity and calcium ion binding, and are mainly involved in the biological process of blood coagulation. Eight more other proteins related to calcium ion binding and blood coagulation were found. Nearly all of these proteins cannot be identified by directly analyzing the plasma sample demonstrating the chemiproteomics a useful approach to characterize interacting proteins in the low abundance range.

Adult↗

PLET1 (C11orf34), a highly expressed and processed novel gene in pig and mouse placenta, is transcribed but poorly spliced in human.

Sequencing of porcine cDNAs identified a novel EST with high frequency in placenta tissue. Full-length PLET1 (placenta-expressed transcript 1, also called C11orf34) matched a mouse cDNA and many bovine and mouse ESTs but no human transcripts or ESTs. However, the porcine cDNA matched several putative exons within a human genomic DNA fragment on chromosome 11. This human locus is in a region of conserved synteny with pig chromosome 9, to which the porcine gene was subsequently mapped. RNA blot hybridization showed that this gene had high expression in porcine and mouse conceptus and throughout placenta development. In situ hybridization using mouse placenta showed PLET1 expression in trophoblast cells of the labyrinth, as well as in spongiotrophoblast and glycogen trophoblast cells. However, no expression of PLET1 was detected by RNA blot analysis of human placenta, although RT-PCR analysis detected very small amounts of partially spliced RNA that were significantly less abundant than the RNA levels in mouse placenta. Donor and acceptor splicing site sequences in the exons of the human gene are poorly conserved and may be the cause of inefficient splicing found specifically in human tissue. Our data correct GenomeScan annotation of this region of the human genome and describe functional gene discovery in mammals not recognized in human EST projects.

Amino Acid Sequence↗

Taking advantage of sophisticated pacemaker diagnostics.

The ever-increasing complexity of pacing systems, combined with functions that vary from one manufacturer to another, can pose challenges during analysis of device function. Standard pacemaker diagnostics are measured data, electrogram telemetry, maker annotations and event counters, albeit with their current limitations. New diagnostic features discussed include time-based diagnostics, histograms of sensed amplitudes, pacing thresholds, or impedance trending. Mode-switching algorithms, combined with diagnostic features, facilitate the use of dual-chamber devices in patients with paroxysmal atrial tachyarrhythmias. The introduction of electrogram storage into pacemakers further improves diagnostic capabilities and allows a permanent validation and optimization of diagnostic and therapeutic algorithms. External diagnostic devices, which provide Holter recordings with continuous marker annotations and patient-triggered diagnostics, are additional features that will become increasingly important.

Arrhythmias, Cardiac↗

Genome-wide transcription profiling of Corynebacterium glutamicum after heat shock and during growth on acetate and glucose.

To monitor the global gene expression of Corynebacterium glutamicum we established two formats of DNA-arrays on nylon membranes. We produced an ordered DNA-array of PCR fragments from a shotgun library of C. glutamicum representing a threefold coverage of the genome. With this format we studied genome-wide transcriptional changes after heat shock. Sequence and subsequent BLAST analysis of PCR fragments with elevated expression after heat shock revealed PCR fragments harboring genes that encode several proteins of the heat shock family, proteins of the oxidative stress response and proteins with unknown function. DNA-arrays based on PCR fragments representing 2804 annotated ORFs of C. glutamicum were used to monitor the transcript levels during growth on acetate and glucose. We determined minimal detectable ratios and compared labeling approaches with random hexamers and ORF-specific primers. ORF-based DNA-array analysis with different labeling approaches showed similar results: e.g. increased mRNA levels of the pta-ack operon, aceA, aceB and genes encoding phosphoenolpyruvate carboxykinase and enzymes of the citric acid cycle during growth on acetate and elevated mRNA levels of some enzymes of the glycolytic pathway and lactate dehydrogenase upon growth on glucose. These results demonstrate that shotgun DNA-arrays and ORF-based DNA-arrays are appropriate tools to study physiology of microorganism.

Acetates↗

Characterization and comparative analysis of the EGLN gene family.

Rat Sm-20 is a homologue of the Caenorhabditis elegans gene egl-9 and has been implicated in the regulation of growth, differentiation and apoptosis in muscle and nerve cells. Null mutants in egl-9 result in a complete tolerance to an otherwise lethal toxin produced by Pseudomonas aeruginosa. This study describes the conserved Egl-Nine (EGLN) gene family of which rat SM-20 and C. elegans Egl-9 are members and characterizes the mouse and human homologues. Each of the human genes (EGLN1, EGLN2 and EGLN3) are of a conserved genomic structure consisting of five coding exons. Phylogenetic analysis and domain organization show that EGLN1 represents the ancestral form of the gene family and that EGLN3 is the human orthologue of rat Sm-20. The previously observed mitochondrial targeting of rat SM-20 is unlikely to be a general feature of the protein family and may be a feature specific to rats. An EGLN gene is unexpectedly found in the genome of P. aeruginosa, a bacterium known to produce a toxin that acts through the Egl-9 protein. The pathogenic bacterium Vibrio cholerae is also shown to have an EGLN gene suggesting that it is an important pathogenicity factor. These results provide new insights into host-pathogen interactions and a basis for further functional characterization of the gene family and resolve discrepancies in annotation between gene family members.

Amino Acid Sequence↗

Expression of hypothetical proteins in human fetal brain: increased expression of hypothetical protein 28.5kDa in Down syndrome, a clue for its tentative role.

Major advances have been made in annotation of sequences of the human genome, although elucidating the functions of these newly discovered genes remains to be a strong challenge. In an effort to give insight into how triplication of chromosome 21 leads to mental retardation in Down syndrome, we have constructed a two-dimensional protein map from control and Down syndrome fetal brain and identified hypothetical proteins with no known functions. Subsequent quantitative analysis of these proteins revealed no apparent change in expression of hypothetical proteins DKZp564P0562.1 (fragment), 16.6, 21.4, 39.5, and 40kDa as well as putative 55kDa protein between controls and Down syndrome fetuses. By contrast, hypothetical protein 28.5kDa was significantly elevated (P<0.05) in fetal Down syndrome. This finding offers an important clue that a hypothetical protein might be involved in the pathomechanisms of brain abnormality in Down syndrome.

Brain↗

In silico analysis of SH3BP2 genomic alterations and expression profiles in CRC.

AIM: Colorectal cancer (CRC) is a widespread health issue that attains high mortality. The adaptor protein SH3BP2 amplification results in metabolic changes, oxidative stress, NK cell activity, and inflammation. The NK cells are capable of destroying tumor cells without prior activation, help prevent metastasis, and have prognostic value. Targeting SH3BP2 to regulate NK cell activity in the TME could enhance CRC-based immunotherapy. MATERIALS AND METHODS: The cancer hallmark tool helps in understanding SH3BP2&#xa0;hallmark annotation. Utilizing the STRING tool and the KEGG pathway, protein functional enrichment and PPI networking were analyzed. TIMER 2.0 was used for immune cell infiltration correlation analysis, and UALCAN was used for CPTAC-based protein expression profiling. RESULTS AND CONCLUSIONS: The GEO (GSE9348) dataset showed SH3BP2 is upregulated in CRC (log2 fold change&#x2009;=&#x2009;1.18). GEO, TCGA, and cBioPortal revealed SH3BP2 alterations in CRC cases, potentially aiding immune evasion. Mutations in SH3BP2 influence cancer growth, suppressing tumors or promoting them by activating NF-&#x3ba;B and affecting immune responses through WNT/&#x3b2;-catenin, PI3K, MAPK, and JAK-STAT pathways. Overall, SH3BP2 plays a key role in cancer growth and immune regulation, making it a promising target for CRC therapy. Further experimental validation is needed to demonstrate its diagnostic and therapeutic potency.

Humans↗

Gene finding with a hidden Markov model of genome structure and evolution.

MOTIVATION: A growing number of genomes are sequenced. The differences in evolutionary pattern between functional regions can thus be observed genome-wide in a whole set of organisms. The diverse evolutionary pattern of different functional regions can be exploited in the process of genomic annotation. The modelling of evolution by the existing comparative gene finders leaves room for improvement. RESULTS: A probabilistic model of both genome structure and evolution is designed. This type of model is called an Evolutionary Hidden Markov Model (EHMM), being composed of an HMM and a set of region-specific evolutionary models based on a phylogenetic tree. All parameters can be estimated by maximum likelihood, including the phylogenetic tree. It can handle any number of aligned genomes, using their phylogenetic tree to model the evolutionary correlations. The time complexity of all algorithms used for handling the model are linear in alignment length and genome number. The model is applied to the problem of gene finding. The benefit of modelling sequence evolution is demonstrated both in a range of simulations and on a set of orthologous human/mouse gene pairs. AVAILABILITY: Free availability over the Internet on www server: http://www.birc.dk/Software/evogene.

Algorithms↗

The non-redundant Bacillus subtilis (NRSub) database: update 1998.

The non-redundant Bacillus subtilis database (NRSub) has been developed in the context of the sequencing project devoted to this bacterium. As this project has reached completion, the whole genome is now available as a single contig. Thanks to the ACNUC database management system and its associated retrieval system Query_win, each functional region of the genome can be accessed individually. Extra annotations have been added such as accession numbers for the genes, locations on the genetic map, codon adaptation index values, as well as cross-references with other collections. NRSub is distributed through anonymous FTP as a text file in EMBL format and as an ACNUC database. It is also possible to access NRSub through two dedicated World Wide Web servers located in France (http://acnuc. univ-lyon1.fr/nrsub/nrsub.html ) and in Japan (http://ddbjs4h.genes. nig.ac.jp/ ).

Bacillus subtilis↗