Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Application of string kernels in protein sequence classification.

INTRODUCTION: The production of biological information has become much greater than its consumption. The key issue now is how to organise and manage the huge amount of novel information to facilitate access to this useful and important biological information. One core problem in classifying biological information is the annotation of new protein sequences with structural and functional features. METHOD: This article introduces the application of string kernels in classifying protein sequences into homogeneous families. A string kernel approach used in conjunction with support vector machines has been shown to achieve good performance in text categorisation tasks. We evaluated and analysed the performance of this approach, and we present experimental results on three selected families from the SCOP (Structural Classification of Proteins) database. We then compared the overall performance of this method with the existing protein classification methods on benchmark SCOP datasets. RESULTS: According to the F1 performance measure and the rate of false positive (RFP) measure, the string kernel method performs well in classifying protein sequences. The method outperformed all the generative-based methods and is comparable with the SVM-Fisher method. DISCUSSION: Although the string kernel approach makes no use of prior biological knowledge, it still captures sufficient biological information to enable it to outperform some of the state-of-the-art methods.

Algorithms↗

Analysis of expressed sequence tags from Brassica rapa L. ssp. pekinensis.

Non-redundant expressed sequence tags (ESTs) were generated from six different organs at various developmental stages of Chinese cabbage, Brassica rapa L. ssp. pekinensis. Of the 1,295 ESTs, 915 (71%) showed significantly high homology in nucleotide or deduced amino acid sequences with other sequences deposited in databases, while 380 did not show similarity to any sequences. Briefly, 598 ESTs matched with proteins of identified biological function, 177 with hypothetical proteins or non-annotated Arabidopsis genome sequences, and 140 with other ESTs. About 82% of the top-scored matching sequences were from Arabidopsis or Brassica, but overall 558 (43%) ESTs matched with Arabidopsis ESTs at the nucleotide sequence level. This observation strongly supports the idea that gene-expression profiles of Chinese cabbage differ from that of Arabidopsis, despite their genome structures being similar to each other. Moreover, sequence analyses of 21 Brassica ESTs revealed that their primary structure is different from those of corresponding annotated sequences of Arabidopsis genes. Our data suggest that direct prediction of Brassica gene expression pattern based on the information from Arabidopsis genome research has some limitations. Thus, information obtained from the Brassica EST study is useful not only for understanding of unique developmental processes of the plant, but also for the study of Arabidopsis genome structure.

Arabidopsis↗

Mantle cell lymphomas express a distinct genetic signature affecting lymphocyte trafficking and growth regulation as compared with subpopulations of normal human B cells.

Differential gene expression analysis, using high-density microarray chips, demonstrated 300-400 genes to be deregulated in mantle cell lymphomas (MCLs) compared with normal B-cell populations. To investigate the significance of this genetic signature in lymphoma etiology and diagnostics, we selected 90 annotated genes involved in a number of cellular functions for further analysis. Our findings demonstrated a normal gene expression of CCR7, which indicated a normal homing to primary follicles, which was in contrast to other receptors for B-cell trafficking, such as a significant down-regulation for CXCR5 and CCR6, as well as down-regulation of IL4R involved in differentiation. This indicated that the malignant transformation of a normal B cell could have appeared during the transition of a primary follicle to a germinal center, i.e., after an initial B-cell activation. Genes involved in blockage of antiproliferative signals in normal cells were also deregulated, e.g., gene expression of TGF beta 2 and Smad3 was suppressed in MCLs. Furthermore, lymphoproliferative signal pathways were active in MCLs compared with normal B cells, because genes encoding, e.g., IL10R alpha and IL18 were up-regulated, as were oncogenes like Bcl-2 and MERTK. Genes encoding receptors for different neurotransmitters mediating B-cell stimulation, such as norepinephrine and cannabinoids were also up-regulated, again illustrating deregulation of a complex network of genes involved in growth and differentiation. Furthermore, hierarchical cluster analysis revealed two subpopulations of MCLs, which indicates that despite the homogeneous and strong overexpression of cyclin D1, further subtyping might be possible.

Apoptosis↗

LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources.

MOTIVATION: The NCBI dbSNP database lists over 9 million single nucleotide polymorphisms (SNPs) in the human genome, but currently contains limited annotation information. SNPs that result in amino acid residue changes (nsSNPs) are of critical importance in variation between individuals, including disease and drug sensitivity. RESULTS: We have developed LS-SNP, a genomic scale software pipeline to annotate nsSNPs. LS-SNP comprehensively maps nsSNPs onto protein sequences, functional pathways and comparative protein structure models, and predicts positions where nsSNPs destabilize proteins, interfere with the formation of domain-domain interfaces, have an effect on protein-ligand binding or severely impact human health. It currently annotates 28,043 validated SNPs that produce amino acid residue substitutions in human proteins from the SwissProt/TrEMBL database. Annotations can be viewed via a web interface either in the context of a genomic region or by selecting sets of SNPs, genes, proteins or pathways. These results are useful for identifying candidate functional SNPs within a gene, haplotype or pathway and in probing molecular mechanisms responsible for functional impacts of nsSNPs. AVAILABILITY: http://www.salilab.org/LS-SNP CONTACT: rachelk@salilab.org SUPPLEMENTARY INFORMATION: http://salilab.org/LS-SNP/supp-info.pdf.

Algorithms↗

Comprehensive post-genomic data analysis approaches integrating biochemical pathway maps.

Post-genomic era research is focusing on studies to attribute functions to genes and their encoded proteins, and to describe the regulatory networks controlling metabolic, protein synthesis and signal transduction pathways. To facilitate the analysis of experiments using post-genomic technologies, new concepts for linking the vast amount of raw data to a biological context have to be developed. Visual representations of pathways help biologists to understand the complex relationships between components of metabolic networks, and provide an invaluable resource for the integration of transcriptomics, proteomics and metabolomics data sets. Besides providing an overview of currently available bioinformatic tools for plant scientists, we introduce BioPathAt, a newly developed visual interface that allows the knowledge-based analysis of genome-scale data by integrating biochemical pathway maps (BioPathAtMAPS module) with a manually scrutinized gene-function database (BioPathAtDB) for the model plant Arabidopsis thaliana. In addition, we discuss approaches for generating a biochemical pathway knowledge database for A. thaliana that includes, in addition to accurate annotation, condensed experimental information regarding in vitro and in vivo gene/protein function.

Computational Biology↗

Microenvironment analysis and identification of magnesium binding sites in RNA.

Interactions with magnesium (Mg2+) ions are essential for RNA folding and function. The locations and function of bound Mg2+ ions are difficult to characterize both experimentally and computationally. In particular, the P456 domain of the Tetrahymena thermophila group I intron, and a 58 nt 23s rRNA from Escherichia coli have been important systems for studying the role of Mg2+ binding in RNA, but characteristics of all the binding sites remain unclear. We therefore investigated the Mg2+ binding capabilities of these RNA systems using a computational approach to identify and further characterize their Mg2+ binding sites. The approach is based on the FEATURE algorithm, reported previously for microenvironment analysis of protein functional sites. We have determined novel physicochemical descriptions of site-bound and diffusely bound Mg2+ ions in RNA that are useful for prediction. Electrostatic calculations using the Non-Linear Poisson Boltzmann (NLPB) equation provided further evidence for the locations of site-bound ions. We confirmed the locations of experimentally determined sites and further differentiated between classes of ion binding. We also identified potentially important, high scoring sites in the group I intron that are not currently annotated as Mg2+ binding sites. We note their potential function and believe they deserve experimental follow-up.

Algorithms↗

Cancer gene expression database (CGED): a database for gene expression profiling with accompanying clinical information of human cancer tissues.

Gene expression profiling of cancer tissues is expected to contribute to our understanding of cancer biology as well as developments of new methods of diagnosis and therapy. Our collaborative efforts in Japan have been mainly focused on solid tumors such as breast, colorectal and hepatocellular cancers. The expression data are obtained by a high-throughput RT-PCR technique, and patients are recruited mainly from a single hospital. In the cancer gene expression database (CGED), the expression and clinical data are presented in a way useful for scientists interested in specific genes or biological functions. The data can be retrieved either by gene identifiers or by functional categories defined by Gene Ontology terms or the Swiss-Prot annotation. Expression patterns of multiple genes, selected by names or similarity search of the patterns, can be compared. Visual presentation of the data with sorting function enables users to easily recognize of relationships between gene expression and clinical parameters. Data for other cancers such as lung and thyroid cancers will be added in the near future. The URL of CGED is http://cged.hgc.jp.

Databases, Genetic↗

The Cerefy Neuroradiology Atlas: a Talairach-Tournoux atlas-based tool for analysis of neuroimages available over the internet.

The article introduces an atlas-assisted method and a tool called the Cerefy Neuroradiology Atlas (CNA), available over the Internet for neuroradiology and human brain mapping. The CNA contains an enhanced, extended, and fully segmented and labeled electronic version of the Talairach-Tournoux brain atlas, including parcelated gyri and Brodmann's areas. To our best knowledge, this is the first online, publicly available application with the Talairach-Tournoux atlas. The process of atlas-assisted neuroimage analysis is done in five steps: image data loading, Talairach landmark setting, atlas normalization, image data exploration and analysis, and result saving. Neuroimage analysis is supported by a near-real-time, atlas-to-data warping based on the Talairach transformation. The CNA runs on multiple platforms; is able to process simultaneously multiple anatomical and functional data sets; and provides functions for a rapid atlas-to-data registration, interactive structure labeling and annotating, and mensuration. It is also empowered with several unique features, including interactive atlas warping facilitating fine tuning of atlas-to-data fit, navigation on the triplanar formed by the image data and the atlas, multiple-images-in-one display with interactive atlas-anatomy-function blending, multiple label display, and saving of labeled and annotated image data. The CNA is useful for fast atlas-assisted analysis of neuroimage data sets. It increases accuracy and reduces time in localization analysis of activation regions; facilitates to communicate the information on the interpreted scans from the neuroradiologist to other clinicians and medical students; increases the neuroradiologist's confidence in terms of anatomy and spatial relationships; and serves as a user-friendly, public domain tool for neuroeducation. At present, more than 700 users from five continents have subscribed to the CNA.

Atlases as Topic↗

Gene expression profiling of bovine in vitro adipogenesis using a cDNA microarray.

The gene expression profile of bovine bone marrow stromal cells undergoing adipogenesis was established using a custom cDNA microarray. Cells that were treated with adipogenic stimulants and those that were not were collected at each of the six time points, and gene expression differences between the treated and untreated samples within each time point were compared using a microarray. Statistical analyses revealed that 158 genes showed a minimum fold change of 2 in at least one of the five post-differentiation time points. These genes are involved in various cellular pathways and functions, including lipogenesis, glycolysis, cytoskeleton remodelling, extracellular matrix, transcription as well as various signalling pathways such as insulin, calcium and wingless signalling. The experiment also identified 17 differentially expressed (DE) microarray elements with no assigned function. Quantitative real-time PCR was employed to validate eight DE genes, and the PCR data were found to reproduce the microarray data for these eight genes. Subsequent gene ontology annotation was able to provide a global overview of the molecular function of DE genes during adipogenesis. This analysis was able to indicate the importance of different gene categories at various stages of adipogenic conversion, thereby providing further insights into the molecular changes during bovine adipogenesis.

Adipogenesis↗

RNAdb 2.0--an expanded database of mammalian non-coding RNAs.

RNAdb is a comprehensive database of mammalian non-protein-coding RNAs (ncRNAs). There is increasing recognition that ncRNAs play important regulatory roles in multicellular organisms, and there is an expanding rate of discovery of novel ncRNAs as well as an increasing allocation of function. In this update to RNAdb, we provide nucleotide sequences and annotations for tens of thousands of non-housekeeping ncRNAs, including a wide range of mammalian microRNAs, small nucleolar RNAs and larger mRNA-like ncRNAs. Some of these have documented functions and/or expression patterns, but the majority remain of unclear significance, and include PIWI-interacting RNAs, ncRNAs identified from the latest rounds of large-scale cDNA sequencing projects, putative antisense transcripts, as well as ncRNAs predicted on the basis of structural features and alignments. Improvements to the database comprise not only new and updated ncRNA datasets, but also provision of microarray-based expression data and closer interface with more specialized ncRNA resources such as miRBase and snoRNA-LBME-db. To access RNAdb, visit http://research.imb.uq.edu.au/RNAdb.

Animals↗

A tunable, ultrasensitive threshold in enzymatic activity governs the DNA methylation landscape.

DNA methylation is a widely studied epigenetic mark, affecting gene expression and cellular function at multiple levels. DNA methylation in the mammalian genome occurs primarily at cytosine-phosphate-guanine (CpG) dinucleotides, and patterning of the methylation landscape (i.e., the presence or absence of CpG methylation at a given genomic location) exhibits a generally bimodal distribution. Although much is known about the enzymatic writers and erasers of CpG methylation, it is not fully understood how these enzymes, along with genetic, chromatin, and regulatory factors, control the genome-wide methylation landscape. In this study, methylation is analyzed at annotated CpG islands (CGIs) and independent CpGs as a function of their proximity to other CpG substrates. Analysis is aided by a computationally efficient stochastic mathematical model of methylation dynamics, enabling parameterization from data. We find that methylation exhibits a switch-like dependence on local CpG density with a threshold of 7-8 CpGs per 100 bp and a Hill coefficient of 4-5. The threshold and steepness of the switch is modified in cell lines in which key enzymes are knocked out. Modeling further elucidates how enzymatic parameters, including catalytic rates and lengthscales of inter-CpG interaction, tune the properties of the switch. Together, the results support a model in which competition between opposing TET1-3 demethylating enzymes and DNA methyltransferases (DNMT3A/B) results in an ultrasensitive switch, analogous to the protein phosphorylation switch (termed "zero-order ultrasensitivity"). Our study provides insight to the mechanisms underlying establishment and maintenance of bimodal DNA methylation landscapes, and further provides a flexible pipeline for gleaning molecular insights to the cellular methylation machinery across cell-specific, epigenomic data sets.

DNA Methylation↗

Evolutionary origins of the endocannabinoid system.

Endocannabinoid system evolution was estimated by searching for functional orthologs in the genomes of twelve phylogenetically diverse organisms: Homo sapiens, Mus musculus, Takifugu rubripes, Ciona intestinalis, Caenorhabditis elegans, Drosophila melanogaster, Saccharomyces cerevisiae, Arabidopsis thaliana, Plasmodium falciparum, Tetrahymena thermophila, Archaeoglobus fulgidus, and Mycobacterium tuberculosis. Sequences similar to human endocannabinoid exon sequences were derived from filtered BLAST searches, and subjected to phylogenetic testing with ClustalX and tree building programs. Monophyletic clades that agreed with broader phylogenetic evidence (i.e., gene trees displaying topographical congruence with species trees) were considered orthologs. The capacity of orthologs to function as endocannabinoid proteins was predicted with pattern profilers (Pfam, Prosite, TMHMM, and pSORT), and by examining queried sequences for amino acid motifs known to serve critical roles in endocannabinoid protein function (obtained from a database of site-directed mutagenesis studies). This novel transfer of functional information onto gene trees enabled us to better predict the functional origins of the endocannabinoid system. Within this limited number of twelve organisms, the endocannabinoid genes exhibited heterogeneous evolutionary trajectories, with functional orthologs limited to mammals (TRPV1 and GPR55), or vertebrates (CB2 and DAGLbeta), or chordates (MAGL and COX2), or animals (DAGLalpha and CB1-like receptors), or opisthokonta (animals and fungi, NAPE-PLD), or eukaryotes (FAAH). Our methods identified fewer orthologs than did automated annotation systems, such as HomoloGene. Phylogenetic profiles, nonorthologous gene displacement, functional convergence, and coevolution are discussed.

Animals↗

QPath: a method for querying pathways in a protein-protein interaction network.

BACKGROUND: Sequence comparison is one of the most prominent tools in biological research, and is instrumental in studying gene function and evolution. The rapid development of high-throughput technologies for measuring protein interactions calls for extending this fundamental operation to the level of pathways in protein networks. RESULTS: We present a comprehensive framework for protein network searches using pathway queries. Given a linear query pathway and a network of interest, our algorithm, QPath, efficiently searches the network for homologous pathways, allowing both insertions and deletions of proteins in the identified pathways. Matched pathways are automatically scored according to their variation from the query pathway in terms of the protein insertions and deletions they employ, the sequence similarity of their constituent proteins to the query proteins, and the reliability of their constituent interactions. We applied QPath to systematically infer protein pathways in fly using an extensive collection of 271 putative pathways from yeast. QPath identified 69 conserved pathways whose members were both functionally enriched and coherently expressed. The resulting pathways tended to preserve the function of the original query pathways, allowing us to derive a first annotated map of conserved protein pathways in fly. CONCLUSION: Pathway homology searches using QPath provide a powerful approach for identifying biologically significant pathways and inferring their function. The growing amounts of protein interactions in public databases underscore the importance of our network querying framework for mining protein network data.

Algorithms↗

In vivo enhancer analysis of human conserved non-coding sequences.

Identifying the sequences that direct the spatial and temporal expression of genes and defining their function in vivo remains a significant challenge in the annotation of vertebrate genomes. One major obstacle is the lack of experimentally validated training sets. In this study, we made use of extreme evolutionary sequence conservation as a filter to identify putative gene regulatory elements, and characterized the in vivo enhancer activity of a large group of non-coding elements in the human genome that are conserved in human-pufferfish, Takifugu (Fugu) rubripes, or ultraconserved in human-mouse-rat. We tested 167 of these extremely conserved sequences in a transgenic mouse enhancer assay. Here we report that 45% of these sequences functioned reproducibly as tissue-specific enhancers of gene expression at embryonic day 11.5. While directing expression in a broad range of anatomical structures in the embryo, the majority of the 75 enhancers directed expression to various regions of the developing nervous system. We identified sequence signatures enriched in a subset of these elements that targeted forebrain expression, and used these features to rank all approximately 3,100 non-coding elements in the human genome that are conserved between human and Fugu. The testing of the top predictions in transgenic mice resulted in a threefold enrichment for sequences with forebrain enhancer activity. These data dramatically expand the catalogue of human gene enhancers that have been characterized in vivo, and illustrate the utility of such training sets for a variety of biological applications, including decoding the regulatory vocabulary of the human genome.

Animals↗

A Web-based classification system of DNA-binding protein families.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The family of DNA-binding proteins is one of the most populated and studied amongst the various genomes of bacteria, archaea and eukaryotes and the Web-based system presented here is an approach to their classification. The DnaProt resource is an annotated and searchable collection of protein sequences for the families of DNA-binding proteins. The database contains 3238 full-length sequences (retrieved from the SWISS-PROT database, release 38) that include, at least, a DNA-binding domain. Sequence entries are organized into families defined by PROSITE patterns, PRINTS motifs and de novo excised signatures. Combining global similarities and functional motifs into a single classification scheme, DNA-binding proteins are classified into 33 unique classes, which helps to reveal comprehensive family relationships. To maximize family information retrieval, DnaProt contains a collection of multiple alignments for each DNA-binding family while the recognized motifs can be used as diagnostically functional fingerprints. All available structural class representatives have been referenced. The resource was developed as a Web-based management system for online free access of customized data sets. Entries are fully hyperlinked to facilitate easy retrieval of the original records from the source databases while functional and phylogenetic annotation will be applied to newly sequenced genomes. The database is freely available for online search of a library containing specific patterns of the identified DNA-binding protein classes and retrieval of individual entries from our WWW server (http://kronos.biol.uoa.gr/~mariak/dbDNA.html).

Amino Acid Motifs↗

Identification of promoter regions in the human genome by using a retroviral plasmid library-based functional reporter gene assay.

Attempts to identify regulatory sequences in the human genome have involved experimental and computational methods such as cross-species sequence comparisons and the detection of transcription factor binding-site motifs in coexpressed genes. Although these strategies provide information on which genomic regions are likely to be involved in gene regulation, they do not give information on their functions. We have developed a functional selection for promoter regions in the human genome that uses a retroviral plasmid library-based system. This approach enriches for and detects promoter function of isolated DNA fragments in an in vitro cell culture assay. By using this method, we have discovered likely promoters of known and predicted genes, as well as many other putative promoter regions based on the presence of features such as CpG islands. Comparison of sequences of 858 plasmid clones selected by this assay with the human genome draft sequence indicates that a significantly higher percentage of sequences align to the 500-bp segment upstream of the transcription start sites of known genes than would be expected from random genomic sequences. We also observed enrichment for putative promoter regions of genes predicted in at least two annotation databases and for clones overlapping with CpG islands. Functional validation of randomly selected clones enriched by this method showed that a large fraction of these putative promoters can drive the expression of a reporter gene in transient transfection experiments. This method promises to be a useful genome-wide function-based approach that can complement existing methods to look for promoters.

3T3 Cells↗

aCHEdb: the database system for ESTHER, the alpha/beta fold family of proteins and the Cholinesterase gene server.

Acetylcholinesterase belongs to a family of proteins, the alpha/beta hydrolase fold family, whose constituents evolutionarily diverged from a common ancestor and share a similar structure of a central beta sheet surrounded by alpha helices. These proteins fulfil a wide range of physiological functions (hydrolases, adhesion molecules, hormone precursors) [Krejci,E., Duval,N., Chatonnet,A., Vincens,P. and Massoulié,J. (1991) Proc. Natl. Acad. Sci. USA , 88, 6647-6651]. ESTHER (for esterases, alpha/beta hydrolase enzymes and relatives) is a database aimed at collecting in one information system, sequence data together with biological annotations and experimental biochemical results related to the structure-function analysis of the enzymes of the family. The major upgrade of the database comes from the use of a new database management system: aCHEdb which uses the ACeDB program designed by Richard Durbin and Jean Thierry-Mieg. It can be found at http://www.ensam.inra.fr/cholinesterase

Animals↗

Confirmation of data mining based predictions of protein function.

MOTIVATION: A central problem in bioinformatics is the assignment of function to sequenced open reading frames (ORFs). The most common approach is based on inferred homology using a statistically based sequence similarity (SIM) method, e.g. PSI-BLAST. Alternative non-SIM based bioinformatic methods are becoming popular. One such method is Data Mining Prediction (DMP). This is based on combining evidence from amino-acid attributes, predicted structure and phylogenic patterns; and uses a combination of Inductive Logic Programming data mining, and decision trees to produce prediction rules for functional class. DMP predictions are more general than is possible using homology. In 2000/1, DMP was used to make public predictions of the function of 1309 Escherichia coli ORFs. Since then biological knowledge has advanced allowing us to test our predictions. RESULTS: We examined the updated (20.02.02) Riley group genome annotation, and examined the scientific literature for direct experimental derivations of ORF function. Both tests confirmed the DMP predictions. Accuracy varied between rules, and with the detail of prediction, but they were generally significantly better than random. For voting rules, accuracies of 75-100% were obtained. Twenty-one of these DMP predictions have been confirmed by direct experimentation. The DMP rules also have interesting biological explanations. DMP is, to the best of our knowledge, the first non-SIM based prediction method to have been tested directly on new data. AVAILABILITY: We have designed the "Genepredictions" database for protein functional predictions. This is intended to act as an open repository for predictions for any organism and can be accessed at http://www.genepredictions.org

Abstracting and Indexing↗