Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Motor skill and mobility recovery outcomes of children and youth with traumatic brain injury.

Traumatic brain injury (TBI) is a major cause of disability in children. Along with other neurological clinical sequelae, children often exhibit motor skill impairment and limitations in functional mobility following TBI. The purpose of this annotated bibliography is to: (1) familiarize therapists with the literature available regarding motor skill and mobility recovery outcomes for children and adolescents with TBI; (2) assist therapists in the selection of motor skill and mobility outcome assessments for use in clinical practice; and (3) provide therapists with comparisons of outcomes for external benchmarking. A number of reports document motor and mobility recovery outcomes as well as recovery in other domains. Studies vary, however, in design, sample size, number and type of outcome assessments used, time since injury at assessment(s), and the consideration of correlating factors such as age at time of injury and injury severity. Further research is needed to describe clinical, satisfaction and resource utilization outcomes, determine outcome predictors, and provide evidence for therapeutic intervention effectiveness.

Brain Injuries↗

MAASE: an alternative splicing database designed for supporting splicing microarray applications.

Alternative splicing is a prominent feature of higher eukaryotes. Understanding of the function of mRNA isoforms and the regulation of alternative splicing is a major challenge in the post-genomic era. The development of mRNA isoform sensitive microarrays, which requires precise splice-junction sequence information, is a promising approach. Despite the availability of a large number of mRNAs and ESTs in various databases and the efforts made to align transcript sequences to genomic sequences, existing alternative splicing databases do not offer adequate information in an appropriate format to aid in splicing array design. Here we describe our effort in constructing the Manually Annotated Alternatively Spliced Events (MAASE) database system, which is specifically designed to support splicing microarray applications. MAASE comprises two components: (1) a manual/computational annotation tool for the efficient extraction of critical sequence and functional information for alternative splicing events and (2) a user-friendly database of annotated events that allows convenient export of information to aid in microarray design and data analysis. We provide a detailed introduction and a step-by-step user guide to the MAASE database system to facilitate future large-scale annotation efforts, integration with other alternative splicing databases, and splicing array fabrication.

Alternative Splicing↗

The proteomes of neurotransmitter receptor complexes form modular networks with distributed functionality underlying plasticity and behaviour.

Neuronal synapses play fundamental roles in information processing, behaviour and disease. Neurotransmitter receptor complexes, such as the mammalian N-methyl-D-aspartate receptor complex (NRC/MASC) comprising 186 proteins, are major components of the synapse proteome. Here we investigate the organisation and function of NRC/MASC using a systems biology approach. Systematic annotation showed that the complex contained proteins implicated in a wide range of cognitive processes, synaptic plasticity and psychiatric diseases. Protein domains were evolutionarily conserved from yeast, but enriched with signalling domains associated with the emergence of multicellularity. Mapping of protein-protein interactions to create a network representation of the complex revealed that simple principles underlie the functional organisation of both proteins and their clusters, with modularity reflecting functional specialisation. The known functional roles of NRC/MASC proteins suggest the complex co-ordinates signalling to diverse effector pathways underlying neuronal plasticity. Importantly, using quantitative data from synaptic plasticity experiments, our model correctly predicts robustness to mutations and drug interference. These studies of synapse proteome organisation suggest that molecular networks with simple design principles underpin synaptic signalling properties with important roles in physiology, behaviour and disease.

Animals↗

Prediction of functional sites by analysis of sequence and structure conservation.

We present a method for prediction of functional sites in a set of aligned protein sequences. The method selects sites which are both well conserved and clustered together in space, as inferred from the 3D structures of proteins included in the alignment. We tested the method using 86 alignments from the NCBI CDD database, where the sites of experimentally determined ligand and/or macromolecular interactions are annotated. In agreement with earlier investigations, we found that functional site predictions are most successful when overall background sequence conservation is low, such that sites under evolutionary constraint become apparent. In addition, we found that averaging of conservation values across spatially clustered sites improves predictions under certain conditions: that is, when overall conservation is relatively high and when the site in question involves a large macromolecular binding interface. Under these conditions it is better to look for clusters of conserved sites than to look for particular conserved sites.

Algorithms↗

Genome-wide DNA polymorphism analyses using VariScan.

BACKGROUND: DNA sequence polymorphisms analysis can provide valuable information on the evolutionary forces shaping nucleotide variation, and provides an insight into the functional significance of genomic regions. The recent ongoing genome projects will radically improve our capabilities to detect specific genomic regions shaped by natural selection. Current available methods and software, however, are unsatisfactory for such genome-wide analysis. RESULTS: We have developed methods for the analysis of DNA sequence polymorphisms at the genome-wide scale. These methods, which have been tested on a coalescent-simulated and actual data files from mouse and human, have been implemented in the VariScan software package version 2.0. Additionally, we have also incorporated a graphical-user interface. The main features of this software are: i) exhaustive population-genetic analyses including those based on the coalescent theory; ii) analysis adapted to the shallow data generated by the high-throughput genome projects; iii) use of genome annotations to conduct a comprehensive analyses separately for different functional regions; iv) identification of relevant genomic regions by the sliding-window and wavelet-multiresolution approaches; v) visualization of the results integrated with current genome annotations in commonly available genome browsers. CONCLUSION: VariScan is a powerful and flexible suite of software for the analysis of DNA polymorphisms. The current version implements new algorithms, methods, and capabilities, providing an important tool for an exhaustive exploratory analysis of genome-wide DNA polymorphism data.

Algorithms↗

LC-MS/MS based proteomic analysis and functional inference of hypothetical proteins in Desulfovibrio vulgaris.

High efficiency capillary liquid chromatography-tandem mass spectrometry (LC-MS/MS) was used to examine the proteins extracted from Desulfovibrio vulgaris cells across six treatment conditions. While our previous study provided a proteomic overview of the cellular metabolism based on proteins with known functions [W. Zhang, M.A. Gritsenko, R.J. Moore, D.E. Culley, L. Nie, K. Petritis, E.F. Strittmatter, D.G. Camp II, R.D. Smith, F.J. Brockman, A proteomic view of the metabolism in Desulfovibrio vulgaris determined by liquid chromatography coupled with tandem mass spectrometry, Proteomics 6 (2006) 4286-4299], this study describes the global detection and functional inference for hypothetical D. vulgaris proteins. Using criteria that a given peptide of a protein is identified from at least two out of three independent LC-MS/MS measurements and that for any protein at least two different peptides are identified among the three measurements, 129 open reading frames (ORFs) originally annotated as hypothetical proteins were found to encode expressed proteins. Functional inference for the conserved hypothetical proteins was performed by a combination of several non-homology based methods: genomic context analysis, phylogenomic profiling, and analysis of a combination of experimental information, including peptide detection in cells grown under specific culture conditions and cellular location of the proteins. Using this approach we were able to assign possible functions to 20 conserved hypothetical proteins. This study demonstrated that a combination of proteomics and bioinformatics methodologies can provide verification of the expression of hypothetical proteins and improve genome annotation.

Amino Acid Sequence↗

Extension and integration of the gene ontology (GO): combining GO vocabularies with external vocabularies.

Structured vocabulary development enhances the management of information in biological databases. As information grows, handling the complexity of vocabularies becomes difficult. Defined methods are needed to manipulate, expand and integrate complex vocabularies. The Gene Ontology (GO) project provides the scientific community with a set of structured vocabularies to describe domains of molecular biology. The vocabularies are used for annotation of gene products and for computational annotation of sequence data sets. The vocabularies focus on three concepts universal to living systems, biological process, molecular function and cellular component. As the vocabularies expand to incorporate terms needed by diverse annotation communities, species-specific terms become problematic. In particular, the use of species-specific anatomical concepts remains unresolved. We present a method for expansion of GO into areas outside of the three original universal concept domains. We combine concepts from two orthogonal vocabularies to generate a larger, more specific vocabulary. The example of mammalian heart development is presented because it addresses two issues that challenge GO; inclusion of organism-specific anatomical terms, and proliferation of terms and relationships. The combination of concepts from orthogonal vocabularies provides a robust representation of relevant terms and an opportunity for evaluation of hypothetical concepts.

Animals↗

The search for essential genes.

The bacterial genomic era began with the publication of the chromosomal sequence of Haemophilus influenzae. As few of the observed genes had been examined experimentally, functional assignments were made by comparative analysis and for many genes no annotation could be made. This mini-review briefly describes the genomic-scale experimental approaches being used to identify genes required for the growth of microorganisms. Identifying 'essential genes', the simplest possible annotation for the unknown open reading frames, is important for antibacterial and antifungal research and is a first step to defining the minimum functional requirement for autonomous growth.

Bacteria↗

Neuroendocrinology of protochordates: insights from Ciona genomics.

The genome for two species of Ciona is available making these tunicates excellent models for studies on the evolution of the chordates. In this review most of the data is from Ciona intestinalis, as the annotation of the C. savignyi genome is not yet available. The phylogenetic position of tunicates at the origin of the chordates and the nature of the genome before expansion in vertebrates allows tunicates to be used as a touchstone for understanding genes that either preceded or arose in vertebrates. A comparison of Ciona, a sea squirt, to other model organisms such as a nematode, fruit fly, zebrafish, frog, chicken and mouse shows that Ciona has many useful traits including accessibility for embryological, lineage tracing, forward genetics, and loss- or gain-of-function experiments. For neuroendocrine studies, these traits are important for determining gene function, whereas the availability of the genome is critical for identification of ligands, receptors, transcription factors and signaling pathways. Four major neurohormones and their receptors have been identified by cloning and to some extent by function in Ciona: gonadotropin-releasing hormone, insulin, insulin-like growth factor, and cionin, a member of the CCK/gastrin family. The simplicity of tunicates should be an advantage in searching for novel functions for these hormones. Other neuroendocrine components that have been annotated in the genome are a multitude of receptors, which are available for cloning, expression and functional studies.

Animals↗

Cloning, expression and characterization of the murine Efemp1, a gene mutated in Doyne-Honeycomb retinal dystrophy.

Development of the bone and cartilage structures is one of the best-studied systems for epithelial-mesenchymal interaction as well as proliferation and differentiation. In a screen for genes differentially expressed in mice deficient for transcription factor AP-2alpha, we have identified a gene which, based on its homology to the human EFEMP-1 gene was designated Efemp1. It encodes for six repeats similar to the domain of the epidermal growth factor. Sequence comparison with EFEMP1 genes of human and rat revealed that the three proteins share a high amino acid identity (92%), suggesting a conserved function during vertebrate development. However, there is no EFEMP1ortholog annotated in sequence databases of other non-mammalian species indicating that it might have evolved in higher vertebrates only. Analysis of the murine genomic locus revealed that the gene is encoded by 11 exons, which are spread over 80 kb of distance on murine chromosome 11A4. The multidomain protein structure may indicate that Efemp1 protein interacts with extracellular matrix components and serves to connect and integrate the function of multiple partner molecules. The gene is expressed in the embryo proper starting from day 9.5 to day 18.5 of murine development. In situ analyses showed that Efemp1 is found in condensing mesenchyme, giving rise to bone and cartilage as well as in developing bone structures of the cranial and the axial skeleton. These results will help in further defining the role of Efemp1 during murine embryogenesis.

Amino Acid Sequence↗

Recognition of human genes by stochastic parsing.

A gene finding system, GeneDecoder, based on a parsing technique using a stochastic grammar and dictionary of genetic words is introduced. The structure of human genes are expressed by a stochastic grammar and a dictionary, whose components are the genetic words consisting of genetic phonemes, built as hidden Markov models (HMMs). The HMMs represent the nucleotide acid bases, the codons, and the amino acids. The genetic words in the dictionary are described by the sequence of these HMMs and represent exons, introns, intergenic regions, tRNA regions and signals in DNA sequences. The statistics between these regions are expressed by the grammar, which is a stochastic network of the genetic words. Using the same kind of technique of speech recognition by HMMs with a word dictionary and a grammar, the stochastic network of genetic words enables the motif dictionary to be used during the parsing of the DNA sequences. At the same time, stochastic features of donor/acceptor sites, information of the di-codon statistics, and other important features are integrated into stochastic scores during the parsing. As a result, while the system parses DNA sequences and finds the exon/intron structures, the protein motifs are automatically annotated in the regions. It helps to identify the functions of the genes and reduces the cost of homology search for each hypothetical coding regions. This method is different from simply using the information of homology search. This method uses the information of the motif patterns during the parsing process, but searching the motif patterns after/before finding the coding regions cannot directly affect the parsing process itself. Experimental results have shown that this method reasonably finds and annotates the motifs in the exons in the DNA sequence of human.

Amino Acid Sequence↗

HCVDB: hepatitis C virus sequences database.

UNLABELLED: To date, more than 30 000 hepatitis C virus (HCV) sequences have been deposited in the generalist databases DNA Data Bank of Japan (DDBJ), EMBL Nucleotide Sequence Database (EMBL) and GenBank. The main difficulties with HCV sequences in these databases are their retrieval, annotation and analyses. To help HCV researchers face the increasing needs of HCV sequence analyses, we developed a specialised database of computer-annotated HCV sequences, called HCVDB. HCVDB is re-built every month from an up-to-date EMBL database by an automated process. HCVDB provides key data about the HCV sequences (e.g. genotype, genomic region, protein names and functions, known 3-dimensional structures) and ensures consistency of the annotations, which enables reliable keyword queries. The database is highly integrated with sequence and structure analysis tools and the SRS (LION bioscience) keywords query system. Thus, any user can extract subsets of sequences matching particular criteria or enter their own sequences and analyse them with various bioinformatics programs available on the same server. AVAILABILITY: HCVDB is available from http://hepatitis.ibcp.fr.

Amino Acid Sequence↗

The proteome of Mannheimia succiniciproducens, a capnophilic rumen bacterium.

Mannheimia succiniciproducens MBEL55E isolated from bovine rumen is an industrially important bacterium as an efficient succinic acid producer. Recently, its full genome sequence was determined. In the present study, we analyzed the M. succiniciproducens proteome based on the genome information using 2-DE and MS. We established proteome reference map of M. succiniciproducens by analyzing whole cellular proteins, membrane proteins, and secreted proteins. More than 200 proteins were identified and characterized by MS/MS supported by various bioinformatic tools. The presence of proteins previously annotated as hypothetical proteins or proteins having putative functions were also confirmed. Based on the proteome reference map, cells in the different growth phases were analyzed at the proteome level. Comparative proteome profiling revealed valuable information to understand physiological changes during growth, and subsequently suggested target genes to be manipulated for the strain improvement.

Animals↗

PRED-CLASS: cascading neural networks for generalized protein classification and genome-wide applications.

A cascading system of hierarchical, artificial neural networks (named PRED-CLASS) is presented for the generalized classification of proteins into four distinct classes-transmembrane, fibrous, globular, and mixed-from information solely encoded in their amino acid sequences. The architecture of the individual component networks is kept very simple, reducing the number of free parameters (network synaptic weights) for faster training, improved generalization, and the avoidance of data overfitting. Capturing information from as few as 50 protein sequences spread among the four target classes (6 transmembrane, 10 fibrous, 13 globular, and 17 mixed), PRED-CLASS was able to obtain 371 correct predictions out of a set of 387 proteins (success rate approximately 96%) unambiguously assigned into one of the target classes. The application of PRED-CLASS to several test sets and complete proteomes of several organisms demonstrates that such a method could serve as a valuable tool in the annotation of genomic open reading frames with no functional assignment or as a preliminary step in fold recognition and ab initio structure prediction methods. Detailed results obtained for various data sets and completed genomes, along with a web sever running the PRED-CLASS algorithm, can be accessed over the World Wide Web at http://o2.biol.uoa.gr/PRED-CLASS.

Computational Biology↗

Transcriptome characterization of the dimorphic and pathogenic fungus Paracoccidioides brasiliensis by EST analysis.

Paracoccidioides brasiliensis is a pathogenic fungus that undergoes a temperature-dependent cell morphology change from mycelium (22 degrees C) to yeast (36 degrees C). It is assumed that this morphological transition correlates with the infection of the human host. Our goal was to identify genes expressed in the mycelium (M) and yeast (Y) forms by EST sequencing in order to generate a partial map of the fungus transcriptome. Individual EST sequences were clustered by the CAP3 program and annotated using Blastx similarity analysis and InterPro Scan. Three different databases, GenBank nr, COG (clusters of orthologous groups) and GO (gene ontology) were used for annotation. A total of 3,938 (Y = 1,654 and M = 2,274) ESTs were sequenced and clustered into 597 contigs and 1,563 singlets, making up a total of 2,160 genes, which possibly represent one-quarter of the complete gene repertoire in P. brasiliensis. From this total, 1,040 were successfully annotated and 894 could be classified in 18 functional COG categories as follows: cellular metabolism (44%); information storage and processing (25%); cellular processes-cell division, posttranslational modifications, among others (19%); and genes of unknown functions (12%). Computer analysis enabled us to identify some genes potentially involved in the dimorphic transition and drug resistance. Furthermore, computer subtraction analysis revealed several genes possibly expressed in stage-specific forms of P. brasiliensis. Further analysis of these genes may provide new insights into the pathology and differentiation of P. brasiliensis.

Base Sequence↗

A putative novel alpha/beta hydrolase ORFan family in Bacillus.

A large number of sequences in each newly sequenced genome correspond to lineage and species-specific proteins, also known as ORFans. Amongst these ORFans, a large number are sequences with unknown structures and functions. We have identified a family of sequences, annotated as hypothetical proteins, which are specific to Bacillus and have carried out a computational study aimed at characterizing this family. Fold-recognition methods predict that these sequences belong to the alpha/beta hydrolase fold. We suggest possible catalytic triads for the ORFans and propose a hypothesis regarding the possible families within the alpha/beta hydrolase superfamily to which they may belong.

Amino Acid Sequence↗

Tropinone reductases, enzymes at the branch point of tropane alkaloid metabolism.

Two stereospecific oxidoreductases constitute a branch point in tropane alkaloid metabolism. Products of tropane metabolism are the alkaloids hyoscyamine, scopolamine, cocaine, and polyhydroxylated nortropane alkaloids, the calystegines. Both tropinone reductases reduce the precursor tropinone to yield either tropine or pseudotropine. In Solanaceae, tropine is incorporated into hyoscyamine and scopolamine; pseudotropine is the first specific metabolite on the way to the calystegines. Isolation, cloning and heterologous expression of both tropinone reductases enabled kinetic characterisation, protein crystallisation, and structure elucidation. Stereospecificity of reduction is achieved by binding tropinone in the respective enzyme active centre in opposite orientation. Immunolocalisation of both enzyme proteins in cultured roots revealed a tissue-specific protein accumulation. Metabolite flux through both arms of the tropane alkaloid pathway appears to be regulated by the activity of both enzymes and by their access to the precursor tropinone. Both tropinone reductases are NADPH-dependent short-chain dehydrogenases with amino acid sequence similarity of more than 50% suggesting their descent from a common ancestor. Putative tropinone reductase sequences annotated in plant genomes other that Solanaceae await functional characterisation.

Alcohol Oxidoreductases↗

Searchlight on domains.

In this issue of Structure, examine in detail the functions of selected domains within proteins both when they are alone and when in combination with others. Domain function is relevant to molecular evolution and to annotation of proteins known only by sequence.

Enzymes↗