Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

PicSNP: a browsable catalog of nonsynonymous single nucleotide polymorphisms in the human genome.

Recent progress in identification and mapping of single nucleotide polymorphisms (SNPs) in the human genome generates an unprecedented opportunity to explore cause-effect relationships between genetic variations and susceptibility to common diseases. For this purpose, one promising strategy would be to select a set of SNPs that potentially alter the function of proteins involved in the pathogenesis of the diseases and compare their frequencies in the affected individuals and the healthy population. In this respect, SNPs that change amino acid sequences (nonsynonymous SNPs; nsSNPs) are of particular interest, since they are more likely to affect protein functions. In this study, we have constructed a catalog of nsSNPs (PicSNP), whose unique features are (i) nsSNPs are classified according to the functions of the affected genes and are searchable under the guidance of hierarchical lists of protein functions and (ii) nsSNPs that lead to amino acid changes in the known functional sites and domains of proteins are highlighted. Out of 1,190,295 SNPs extracted from public database, we identified 3793 nsSNPs and classified them in 1247 categories of protein functions. 495 sites and domains annotated in the Swiss-Prot database were found to include nsSNPs, including 2 nsSNPs in disulfide-binding sites and 38 nsSNPs in transmembrane regions. PicSNP is available via the World Wide Web (http://picsnp.org) and would support research questing for SNPs involved in common diseases.

Databases, Factual↗

Assessment of the metabolic capabilities of Haemophilus influenzae Rd through a genome-scale pathway analysis.

The annotated full DNA sequence is becoming available for a growing number of organisms. This information along with additional biochemical and strain-specific data can be used to define metabolic genotypes and reconstruct cellular metabolic networks. The first free-living organism for which the entire genomic sequence was established was Haemophilus influenzae. Its metabolic network is reconstructed herein and contains 461 reactions operating on 367 intracellular and 84 extracellular metabolites. With the metabolic reaction network established, it becomes necessary to determine its underlying pathway structure as defined by the set of extreme pathways. The H. influenzae metabolic network was subdivided into six subsystems and the extreme pathways determined for each subsystem based on stoichiometric, thermodynamic, and systems-specific constraints. Positive linear combinations of these pathways can be taken to determine the extreme pathways for the complete system. Since these pathways span the capabilities of the full system, they could be used to address a number of important physiological questions. First, they were used to reconcile and curate the sequence annotation by identifying reactions whose function was not supported in any of the extreme pathways. Second, they were used to predict gene products that should be co-regulated and perhaps co-expressed. Third, they were used to determine the composition of the minimal substrate requirements needed to support the production of 51 required metabolic products such as amino acids, nucleotides, phospholipids, etc. Fourth, sets of critical gene deletions from core metabolism were determined in the presence of the minimal substrate conditions and in more complete conditions reflecting the environmental niche of H. influenzae in the human host. In the former case, 11 genes were determined to be critical while six remained critical under the latter conditions. This study represents an important milestone in theoretical biology, namely the establishment of the first extreme pathway structure of a whole genome.

Genome, Bacterial↗

The Macrostomum lignano EST database as a molecular resource for studying platyhelminth development and phylogeny.

We report the development of an Expressed Sequence Tag (EST) resource for the flatworm Macrostomum lignano. This taxon is of interest due to its basal placement within the flatworms. As such, it provides a useful comparative model for understanding the development of neural and sensory organization. It was anticipated on the basis of previous studies [e.g., Sánchez-Alvarado et al., Development, 129:5659-5665, (2002)] that a wide range of developmental markers would be expressed in later-stage macrostomids, and this proved to be the case, permitting recovery of a range of gene sequences important in development. To this end, an adult Macrostomum cDNA library was generated and 7,680 Macrostomum ESTs were sequenced from the 5' end. In addition, 1,536 of these aforementioned sequences were sequenced from the 3' end. Of the roughly 5,416 non-redundant sequences identified, 68% are similar to previously reported genes of known function. In addition, nearly 100 specific clones were obtained with potential neural and sensory function. From these data, an annotated searchable database of the Macrostomum EST collection has been made available on the web. A major objective was to obtain genes that would allow reconstruction of embryogenesis, and in particular neurogenesis, in a basal platyhelminth. The sequences recovered will serve as probes with which the origin and morphogenesis of lineages and tissues can be followed. To this end, we demonstrate a protocol for combined immunohistochemistry and in situ hybridization labeling in juvenile Macrostomum, employing homologs of lin11/lim1 and six3/optix. Expression of these genes is shown in the context of the neuropile/muscle system.

Animals↗

Identification and characterization of genes associated with the induction of embryogenic competence in leaf-protoplast-derived alfalfa cells.

Alfalfa leaf protoplast-derived cells can develop into somatic embryos depending on the concentration of 2,4-dichlorophenoxyacetic acid (2,4-D) in the initial culture medium. In order to reveal gene expression changes during the establishment of embryogenic competence, we compared the cell types developed in the presence of 1 and 10 microM 2,4-D, respectively, at the time of their first cell divisions (fourth day of culture) using a PCR-based cDNA subtraction approach. Although the subtraction efficiency was relatively low, applying an additional differential screening step allowed the identification of 38 10 microM 2,4-D up-regulated transcripts. The corresponding genes/proteins were annotated and representatives of various functional groups were selected for more detailed gene expression analysis. Real-time quantitative PCR (RT-QPCR) analysis was used to determine relative expression of the selected genes in 2,4-D-treated leaves as well as during the whole process of somatic embryogenesis. Gene expression patterns confirmed 2,4-D inducibility for all but one of the 11 investigated genes as well as for the positive control leafy cotyledon1 (MsLEC1) gene. The characterized genes exhibited differential expression patterns during the early induction phase and the late embryo differentiation phase of somatic embryogenesis. Genes coding for a GST-transferase, a PR10 pathogenesis-related protein, a cell division-related ribosomal (S3a) protein, an ARF-type small GTPase and the nucleosome assembly factor family SET protein exhibited higher relative expression not only during the induction of somatic embryogenesis but at the time of somatic embryo differentiation as well. This may indicate that the expression of these genes is associated with developmental transitions (differentiation as well as de-differentiation) during the process of somatic embryogenesis.

2,4-Dichlorophenoxyacetic Acid↗

Toxicogenomic approach for assessing toxicant-related disease.

The problems of identifying environmental factors involved in the etiology of human disease and performing safety and risk assessments of drugs and chemicals have long been formidable issues. Three principal components for predicting potential human health risks are: (1) the diverse structure and properties of thousands of chemicals and other stressors in the environment; (2) the time and dose parameters that define the relationship between exposure and disease; and (3) the genetic diversity of organisms used as surrogates to determine adverse chemical effects. The global techniques evolving from successful genomics efforts are providing new exciting tools with which to address these intractable problems of environmental health and toxicology. In order to exploit the scientific opportunities, the National Institute of Environmental Health Sciences has created the National Center for Toxicogenomics (NCT). The primary mission of the NCT is to use gene expression technology, proteomics and metabolite profiling to create a reference knowledge base that will allow scientists to understand mechanisms of toxicity and to be able to predict the potential toxicity of new chemical entities and drugs. A principal scientific objective underpinning the use of microarray analysis of chemical exposures is to demonstrate the utility of signature profiling of the action of drugs or chemicals and to utilize microarray methodologies to determine biomarkers of exposure and potential adverse effects. The initial approach of the NCT is to utilize proof-of-principle experiments in an effort to "phenotypically anchor" the altered patterns of gene expression to conventional parameters of toxicity and to define dose and time relationships in which the expression of such signature genes may precede the development of overt toxicity. The microarray approach is used in conjunction with proteomic techniques to identify specific proteins that may serve as signature biomarkers. The longer-range goal of these efforts is to develop a reference relational database of chemical effects in biological systems (CEBS) that can be used to define common mechanisms of toxicity, chemical and drug actions, to define cellular pathways of response, injury and, ultimately, disease. In order to implement this strategy, the NCT has created a consortium of research organizations and private sector companies to actively collaborative in populating the database with high quality primary data. The evolution of discrete databases to a knowledge base of toxicogenomics will be accomplished through establishing relational interfaces with other sources of information on the structure and activity of chemicals such as that of the National Toxicology Program (NTP) and with databases annotating gene identity, sequence, and function.

Animals↗

Yeast Protein database (YPD): a database for the complete proteome of Saccharomyces cerevisiae.

The Yeast Protein Database (YPD) is a database for the proteins of the budding yeast,Saccharomyces cerevisiae. YPD is the first annotated database for the complete proteome of any organism. Now that the complete genome sequence of yeast is available, YPD contains entries for each of the characterized proteins and for each of the uncharacterized proteins predicted from the sequence. Contained in YPD are the calculated properties of each protein such as molecular weight and isoelectric point, experimentally determined properties such as subcellular localization and post-translational modifications, and extensive annotations from the yeast literature. YPD contains 25 000 lines of textual annotation that describe the known functions, mutant phenotypes, interactions, and other properties for the approximately 6000 proteins in the yeast proteome. The information in YPD is updated daily, and it is available on the World Wide Web at http://www.proteome.com/YPDhome.html .

Amino Acid Sequence↗

3MATRIX and 3MOTIF: a protein structure visualization system for conserved sequence motifs.

Computational methods such as sequence alignment and motif construction are useful in grouping related proteins into families, as well as helping to annotate new proteins of unknown function. These methods identify conserved amino acids in protein sequences, but cannot determine the specific functional or structural roles of conserved amino acids without additional study. In this work, we present 3MATRIX (http://3matrix.stanford.edu) and 3MOTIF (http://3motif.stanford.edu), a web-based sequence motif visualization system that displays sequence motif information in its appropriate three-dimensional (3D) context. This system is flexible in that users can enter sequences, keywords, structures or sequence motifs to generate visualizations. In 3MOTIF, users can search using discrete sequence motifs such as PROSITE patterns, eMOTIFs, or any other regular expression-like motif. Similarly, 3MATRIX accepts an eMATRIX position-specific scoring matrix, or will convert a multiple sequence alignment block into an eMATRIX for visualization. Each query motif is used to search the protein structure database for matches, in which the motif is then visually highlighted in three dimensions. Important properties of motifs such as sequence conservation and solvent accessible surface area are also displayed in the visualizations, using carefully chosen color shading schemes.

Amino Acid Motifs↗

New challenges in gene expression data analysis and the extended GEPAS.

Since the first papers published in the late nineties, including, for the first time, a comprehensive analysis of microarray data, the number of questions that have been addressed through this technique have both increased and diversified. Initially, interest focussed on genes coexpressing across sets of experimental conditions, implying, essentially, the use of clustering techniques. Recently, however, interest has focussed more on finding genes differentially expressed among distinct classes of experiments, or correlated to diverse clinical outcomes, as well as in building predictors. In addition to this, the availability of accurate genomic data and the recent implementation of CGH arrays has made mapping expression and genomic data on the chromosomes possible. There is also a clear demand for methods that allow the automatic transfer of biological information to the results of microarray experiments. Different initiatives, such as the Gene Ontology (GO) consortium, pathways databases, protein functional motifs, etc., provide curated annotations for genes. Whereas many resources on the web focus mainly on clustering methods, GEPAS has evolved to cope with the aforementioned new challenges that have recently arisen in the field of microarray data analysis. The web-based pipeline for microarray gene expression data, GEPAS, is available at http://gepas.bioinfo.cnio.es.

Gene Expression Profiling↗

Over 20% of human transcripts might form sense-antisense pairs.

The major challenge to identifying natural sense- antisense (SA) transcripts from public databases is how to determine the correct orientation for an expressed sequence, especially an expressed sequence tag sequence. In this study, we established a set of very stringent criteria to identify the correct orientation of each human transcript. We used these orientation-reliable transcripts to create 26 741 transcription clusters in the human genome. Our analysis shows that 22% (5880) of the human transcription clusters form SA pairs, higher than any previous estimates. Our orientation-specific RT-PCR results along with the comparison of experimental data from previous studies confirm that our SA data set is reliable. This study not only demonstrates that our criteria for the prediction of SA transcripts are efficient, but also provides additional convincing data to support the view that antisense transcription is quite pervasive in the human genome. In-depth analyses show that SA transcripts have some significant differences compared with other types of transcripts, with regard to chromosomal distribution and Gene Ontology-annotated categories of physiological roles, functions and spatial localizations of gene products.

Base Pairing↗

VOMBAT: prediction of transcription factor binding sites using variable order Bayesian trees.

Variable order Markov models and variable order Bayesian trees have been proposed for the recognition of transcription factor binding sites, and it could be demonstrated that they outperform traditional models, such as position weight matrices, Markov models and Bayesian trees. We develop a web server for the recognition of DNA binding sites based on variable order Markov models and variable order Bayesian trees offering the following functionality: (i) given datasets with annotated binding sites and genomic background sequences, variable order Markov models and variable order Bayesian trees can be trained; (ii) given a set of trained models, putative DNA binding sites can be predicted in a given set of genomic sequences and (iii) given a dataset with annotated binding sites and a dataset with genomic background sequences, cross-validation experiments for different model combinations with different parameter settings can be performed. Several of the offered services are computationally demanding, such as genome-wide predictions of DNA binding sites in mammalian genomes or sets of 10(4)-fold cross-validation experiments for different model combinations based on problem-specific data sets. In order to execute these jobs, and in order to serve multiple users at the same time, the web server is attached to a Linux cluster with 150 processors. VOMBAT is available at http://pdw-24.ipk-gatersleben.de:8080/VOMBAT/.

Algorithms↗

Nuclease activity of the MutS homologue MutS2 from Thermus thermophilus is confined to the Smr domain.

MutS homologues are highly conserved enzymes engaged in DNA mismatch repair (MMR), meiotic recombination and other DNA modifications. Genome sequencing projects have revealed that bacteria and plants possess a MutS homologue, MutS2. MutS2 lacks the mismatch-recognition domain of MutS, but contains an extra C-terminal region called the small MutS-related (Smr) domain. Sequences homologous to the Smr domain are annotated as 'proteins of unknown function' in various organisms ranging from bacteria to human. Although recent in vivo studies indicate that MutS2 plays an important role in recombinational events, there had been only limited characterization of the biochemical function of MutS2 and the Smr domain. We previously established that Thermus thermophilus MutS2 (ttMutS2) possesses endonuclease activity. In this study, we report that a Smr-deleted ttMutS2 mutant retains the dimerization, ATPase and DNA-binding activities, but has no endonuclease activity. Furthermore, the Smr domain alone was stable and functional in binding and incising DNA. It is noteworthy that an endonuclease activity is associated with a MutS homologue, which is generally thought to recognize specific DNA structures.

Adenosine Triphosphatases↗

The National Microbial Pathogen Database Resource (NMPDR): a genomics platform based on subsystem annotation.

The National Microbial Pathogen Data Resource (NMPDR) (http://www.nmpdr.org) is a National Institute of Allergy and Infections Disease (NIAID)-funded Bioinformatics Resource Center that supports research in selected Category B pathogens. NMPDR contains the complete genomes of approximately 50 strains of pathogenic bacteria that are the focus of our curators, as well as >400 other genomes that provide a broad context for comparative analysis across the three phylogenetic Domains. NMPDR integrates complete, public genomes with expertly curated biological subsystems to provide the most consistent genome annotations. Subsystems are sets of functional roles related by a biologically meaningful organizing principle, which are built over large collections of genomes; they provide researchers with consistent functional assignments in a biologically structured context. Investigators can browse subsystems and reactions to develop accurate reconstructions of the metabolic networks of any sequenced organism. NMPDR provides a comprehensive bioinformatics platform, with tools and viewers for genome analysis. Results of precomputed gene clustering analyses can be retrieved in tabular or graphic format with one-click tools. NMPDR tools include Signature Genes, which finds the set of genes in common or that differentiates two groups of organisms. Essentiality data collated from genome-wide studies have been curated. Drug target identification and high-throughput, in silico, compound screening are in development.

Bacteria↗

Gene discovery and gene expression in the rice blast fungus, Magnaporthe grisea: analysis of expressed sequence tags.

Over 28,000 expressed sequence tags (ESTs) were produced from cDNA libraries representing a variety of growth conditions and cell types. Several Magnaporthe grisea strains were used to produce the libraries, including a nonpathogenic strain bearing a mutation in the PMK1 mitogen-activated protein kinase. Approximately 23,000 of the ESTs could be clustered into 3,050 contigs, leaving 5,127 singleton sequences. The estimate of 8,177 unique sequences indicates that over half of the genes of the fungus are represented in the ESTs. Analysis of EST frequency reveals growth and cell type-specific patterns of gene expression. This analysis establishes criteria for identification of fungal genes involved in pathogenesis. A large fraction of the genes represented by ESTs have no known function or described homologs. Manual annotation of the most abundant cDNAs with no known homologs allowed us to identify a family of metallothionein proteins present in M. grisea, Neurospora crassa, and Fusarium graminearum. In addition, multiply represented ESTs permitted the identification of alternatively spliced mRNA species. Alternative splicing was rare, and in most cases, the alternate mRNA forms were unspliced, although alternative 5' splice sites were also observed.

Expressed Sequence Tags↗

Identification of psl, a locus encoding a potential exopolysaccharide that is essential for Pseudomonas aeruginosa PAO1 biofilm formation.

Bacteria inhabiting biofilms usually produce one or more polysaccharides that provide a hydrated scaffolding to stabilize and reinforce the structure of the biofilm, mediate cell-cell and cell-surface interactions, and provide protection from biocides and antimicrobial agents. Historically, alginate has been considered the major exopolysaccharide of the Pseudomonas aeruginosa biofilm matrix, with minimal regard to the different functions polysaccharides execute. Recent chemical and genetic studies have demonstrated that alginate is not involved in the initiation of biofilm formation in P. aeruginosa strains PAO1 and PA14. We hypothesized that there is at least one other polysaccharide gene cluster involved in biofilm development. Two separate clusters of genes with homology to exopolysaccharide biosynthetic functions were identified from the annotated PAO1 genome. Reverse genetics was employed to generate mutations in genes from these clusters. We discovered that one group of genes, designated psl, are important for biofilm initiation. A PAO1 strain with a disruption of the first two genes of the psl cluster (PA2231 and PA2232) was severely compromised in biofilm initiation, as confirmed by static microtiter and continuous culture flow cell and tubing biofilm assays. This impaired biofilm phenotype could be complemented with the wild-type psl sequences and was not due to defects in motility or lipopolysaccharide biosynthesis. These results implicate an as yet unknown exopolysaccharide as being required for the formation of the biofilm matrix. Understanding psl-encoded exopolysaccharide expression and protection in biofilms will provide insight into the pathogenesis of P. aeruginosa in cystic fibrosis and other infections involving biofilms.

Bacterial Proteins↗

Bioinformatics: use in bacterial vaccine discovery.

Bioinformatics has now become a common laboratory name for groups studying genomic sequences. It is composed of many different, yet interrelated scientific fields such as genomics, proteomics, and transcriptional profiling. The availability of complete genomic sequences, especially prokaryotic organisms, allows one to rapidly identify, analyze, and clone genes of interest. For bacterial vaccine discovery, one can "mine" the genomic sequence for potential surface targets using various algorithms, characterize these gene targets, and produce primers for cloning, all before one enters the wet laboratory. This review will focus on various genomic mining tools/algorithms available for predicting open reading frames and their associated annotation (if known), physical and functional characterization, and cellular localization. Finally, examples are given of how all of this is being used for the identification of potential bacterial vaccine candidates.

Animals↗

Interpreting experimental results using gene ontologies.

High-throughput experimental techniques, such as microarrays, produce large amounts of data and knowledge about gene expression levels. However, interpretation of these data and turning it into biologically meaningful knowledge can be challenging. Frequently the output of such an analysis is a list of significant genes or a ranked list of genes. In the case of DNA microarray studies, data analysis often leads to lists of hundreds of differentially expressed genes. Also, clustering of gene expression data may lead to clusters of tens to hundreds of genes. These data are of little use if one is not able to interpret the results in a biological context. The Gene Ontology Consortium provides a controlled vocabulary to annotate the biological knowledge we have or that is predicted for a given gene. The Gene Ontologies (GOs) are organized as a hierarchy of annotation terms that facilitate an analysis and interpretation at different levels. The top-level ontologies are molecular function, biological process, and cellular component. Several annotation databases for genes of different organisms exist. This chapter describes how to use GO in order to help biologically interpret the lists of genes resulting from high-throughput experiments. It describes some statistical methods to find significantly over- or underrepresented GO terms within a list of genes and describes some tools and how to use them in order to do such an analysis. This chapter focuses primarily on the tool GOstat (http://gostat.wehi.edu.au). Other tools exist that enable similar analyses, but are not described in detail here.

Animals↗

CLENCH: a program for calculating Cluster ENriCHment using the Gene Ontology.

SUMMARY: Analysis of microarray data most often produces lists of genes with similar expression patterns, which are then subdivided into functional categories for biological interpretation. Such functional categorization is most commonly accomplished using Gene Ontology (GO) categories. Although there are several programs that identify and analyze functional categories for human, mouse and yeast genes, none of them accept Arabidopsis thaliana data. In order to address this need for A.thaliana community, we have developed a program that retrieves GO annotations for A.thaliana genes and performs functional category analysis for lists of genes selected by the user. AVAILABILITY: http://www.personal.psu.edu/nhs109/Clench

Abstracting and Indexing↗

The untranslated regions of eukaryotic mRNAs: structure, function, evolution and bioinformatic tools for their analysis.

The crucial role of the non-coding portion of genomes is now widely acknowledged. In particular, mRNA untranslated regions are involved in many post-transcriptional regulatory pathways that control mRNA localisation, stability and translation efficiency. A review is given of the most recent research works on the functional characterisation of eukaryotic mRNA untranslated regions. In order to make possible a systematic and detailed sequence analysis of mRNA untranslated regions (UTRs), a non-redundant database of metazoan mRNA untranslated sequences annotated for the occurrence of specific functional elements, UTRdb, was devised. These elements, whose consensus structure has been devised on the basis of experimental assays and of comparative analyses, have been collected in the UTRsite database. A suitable pattern-matching software has been devised to search UTRsite patterns in user-submitted sequences, also assessing their statistical significance. Structural, compositional and evolutionary features of untranslated sequences of metazoan mRNAs have been investigated showing peculiar intra- and interspecific patterns.

Animals↗