Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Microbial diversity and metabolic pathways linked to benzene degradation in petrochemical-polluted groundwater.

The rapid advance in shotgun metagenome sequencing has enabled us to identify uncultivated functional microorganisms in polluted environments. While aerobic petrochemical-degrading pathways have been extensively studied, the anaerobic mechanisms remain less explored. Here, we conducted a study at a petrochemical-polluted groundwater site in Henan Province, Central China. A total of twelve groundwater monitoring wells were installed to collect groundwater samples. Benzene appeared to be the predominant pollutant, detected in 10 out of 12 samples, with concentrations ranging from 1.4 μg/L to 5,280 μg/L. Due to the low aquifer permeability, pollutant migration occurred slowly, resulting in relatively low benzene concentrations downstream within the heavily polluted area. Deep metagenome sequencing revealed Proteobacteria as the dominant phylum, accounting for over 63 % of total abundances. Microbial α-diversity was low in heavily polluted samples, with community compositions substantially differing from those in lightly polluted samples. dmpK encoding the phenol/toluene 2-monooxygenase was detected across all samples, while the dioxygenase bedC1 was not detected, suggesting that aerobic benzene degradation might occur through monooxygenation. Sequence assembly and binning yielded 350 high-quality metagenome-assembled genomes (MAGs), with 30 MAGs harboring functional genes associated with aerobic or anaerobic benzene degradation. About 80 % of MAGs harboring functional genes associated with anaerobic benzene degradation remained taxonomically unclassified at the genus level, suggesting that our current database coverage of anaerobic benzene-degrading microorganisms is very limited. Furthermore, two genes integral to anaerobic benzene metabolism, i.e, benzoyl-CoA reductase (bamB) and glutaryl-CoA dehydrogenase (acd), were not annotated by metagenome functional analyses but were identified within the MAGs, signifying the importance of integrating both contig-based and MAG-based approaches. Together, our efforts of functional annotation and metagenome binning generate a robust blueprint of microbial functional potentials in petrochemical-polluted groundwater, which is crucial for designing proficient bioremediation strategies.

Groundwater↗

'PACLIMS': a component LIM system for high-throughput functional genomic analysis.

BACKGROUND: Recent advances in sequencing techniques leading to cost reduction have resulted in the generation of a growing number of sequenced eukaryotic genomes. Computational tools greatly assist in defining open reading frames and assigning tentative annotations. However, gene functions cannot be asserted without biological support through, among other things, mutational analysis. In taking a genome-wide approach to functionally annotate an entire organism, in this application the approximately 11,000 predicted genes in the rice blast fungus (Magnaporthe grisea), an effective platform for tracking and storing both the biological materials created and the data produced across several participating institutions was required. RESULTS: The platform designed, named PACLIMS, was built to support our high throughput pipeline for generating 50,000 random insertion mutants of Magnaporthe grisea. To be a useful tool for materials and data tracking and storage, PACLIMS was designed to be simple to use, modifiable to accommodate refinement of research protocols, and cost-efficient. Data entry into PACLIMS was simplified through the use of barcodes and scanners, thus reducing the potential human error, time constraints, and labor. This platform was designed in concert with our experimental protocol so that it leads the researchers through each step of the process from mutant generation through phenotypic assays, thus ensuring that every mutant produced is handled in an identical manner and all necessary data is captured. CONCLUSION: Many sequenced eukaryotes have reached the point where computational analyses are no longer sufficient and require biological support for their predicted genes. Consequently, there is an increasing need for platforms that support high throughput genome-wide mutational analyses. While PACLIMS was designed specifically for this project, the source and ideas present in its implementation can be used as a model for other high throughput mutational endeavors.

Algorithms↗

Protein variety and functional diversity: Swiss-Prot annotation in its biological context.

We all know that the dogma 'one gene, one protein' is obsolete. A functional protein and, likewise, a protein's ultimate function depend not only on the underlying genetic information but also on the ongoing conditions of the cellular system. Frequently the transcript, like the polypeptide, is processed in multiple ways, but only one or a few out of a multitude of possible variants are produced at a time. An overview on processes that can lead to sequence variety and structural diversity in eukaryotes is given. The UniProtKB/Swiss-Prot protein knowledgebase provides a wealth of information regarding protein variety, function and associated disorders. Examples for such annotation are shown and further ones are available at http://www.expasy.org/sprot/tutorial/examples_CRB.

Amino Acid Sequence↗

Ontology annotation: mapping genomic regions to biological function.

With numerous whole genomes now in hand, and experimental data about genes and biological pathways on the increase, a systems approach to biological research is becoming essential. Ontologies provide a formal representation of knowledge that is amenable to computational as well as human analysis, an obvious underpinning of systems biology. Mapping function to gene products in the genome consists of two, somewhat intertwined enterprises: ontology building and ontology annotation. Ontology building is the formal representation of a domain of knowledge; ontology annotation is association of specific genomic regions (which we refer to simply as 'genes', including genes and their regulatory elements and products such as proteins and functional RNAs) to parts of the ontology. We consider two complementary representations of gene function: the Gene Ontology (GO) and pathway ontologies. GO represents function from the gene's eye view, in relation to a large and growing context of biological knowledge at all levels. Pathway ontologies represent function from the point of view of biochemical reactions and interactions, which are ordered into networks and causal cascades. The more mature GO provides an example of ontology annotation: how conclusions from the scientific literature and from evolutionary relationships are converted into formal statements about gene function. Annotations are made using a variety of different types of evidence, which can be used to estimate the relative reliability of different annotations.

Animals↗

Comprehensive gene expression analysis by transcript profiling.

After the completion of the genomic sequence of Arabidopsis thaliana, it is now a priority to identify all the genes, their patterns of expression and functions. Transcript profiling is playing a substantial role in annotating and determining gene functions, having advanced from one-gene-at-a-time methods to technologies that provide a holistic view of the genome. In this review, comprehensive transcript profiling methodologies are described, including two that are used extensively by the authors, cDNA-AFLP and cDNA microarraying. Both these technologies illustrate the requirement to integrate molecular biology, automation, LIMS and data analysis. With so much uncharted territory in the Arabidopsis genome, and the desire to tackle complex biological traits, such integrated systems will provide a rich source of data for the correlative, functional annotation of genes.

Gene Expression Profiling↗

Inference of protein function from protein structure.

Structural genomics has brought us three-dimensional structures of proteins with unknown functions. To shed light on such structures, we have developed ProKnow (http://www.doe-mbi.ucla.edu/Services/ProKnow/), which annotates proteins with Gene Ontology functional terms. The method extracts features from the protein such as 3D fold, sequence, motif, and functional linkages and relates them to function via the ProKnow knowledgebase of features, which links features to annotated functions via annotation profiles. Bayes' theorem is used to compute weights of the functions assigned, using likelihoods based on the extracted features. The description level of the assigned function is quantified by the ontology depth (from 1 = general to 9 = specific). Jackknife tests show approximately 89% correct assignments at ontology depth 1 and 40% at depth 9, with 93% coverage of 1507 distinct folded proteins. Overall, about 70% of the assignments were inferred correctly. This level of performance suggests that ProKnow is a useful resource in functional assessments of novel proteins.

Amino Acid Motifs↗

A defect in a novel Nek-family kinase causes cystic kidney disease in the mouse and in zebrafish.

The murine autosomal recessive juvenile cystic kidney (jck) mutation results in polycystic kidney disease. We have identified in jck mice a mutation in Nek8, a novel and highly conserved member of the Nek kinase family. In vitro expression of mutated Nek8 results in enlarged, multinucleated cells with an abnormal actin cytoskeleton. To confirm that a defect in the Nek8 gene can cause cystic disease, we performed a cross-species analysis: injection of zebrafish embryos with a morpholino anti-sense oligonucleotide corresponding to the ortholog of Nek8 resulted in the formation of pronephric cysts. These results demonstrate that comparative analysis of gene function in different model systems represents a powerful means to annotate gene function.

Amino Acid Sequence↗

Protein sequence analysis in silico: application of structure-based bioinformatics to genomic initiatives.

The current pace of high-throughput genome sequencing programs coupled with high-throughput functional genomic screens has provided researchers with a bewildering array of sequence and biological data to contend with. Identification of proteins of interest from a particular biological study requires the application of bioinformatic tools to process and prioritise the data. From a protein function standpoint, transfer of annotation from known proteins to a novel target is currently the only practical way to convert vast quantities of raw sequence data into meaningful information. New bioinformatics tools now provide more sophisticated methods to transfer functional annotation, integrating sequence, family profile and structural search methodology. The importance of these approaches to medical research is increasing as we move to annotate the proteome through functional and structural genomic efforts.

Animals↗

A methodology and implementation for annotating digital images for context-appropriate use in an academic health care environment.

Use of digital medical images has become common over the last several years, coincident with the release of inexpensive, mega-pixel quality digital cameras and the transition to digital radiology operation by hospitals. One problem that clinicians, medical educators, and basic scientists encounter when handling images is the difficulty of using business and graphic arts commercial-off-the-shelf (COTS) software in multicontext authoring and interactive teaching environments. The authors investigated and developed software-supported methodologies to help clinicians, medical educators, and basic scientists become more efficient and effective in their digital imaging environments. The software that the authors developed provides the ability to annotate images based on a multispecialty methodology for annotation and visual knowledge representation. This annotation methodology is designed by consensus, with contributions from the authors and physicians, medical educators, and basic scientists in the Departments of Radiology, Neurobiology and Anatomy, Dermatology, and Ophthalmology at the University of Utah. The annotation methodology functions as a foundation for creating, using, reusing, and extending dynamic annotations in a context-appropriate, interactive digital environment. The annotation methodology supports the authoring process as well as output and presentation mechanisms. The annotation methodology is the foundation for a Windows implementation that allows annotated elements to be represented as structured eXtensible Markup Language and stored separate from the image(s).

Academic Medical Centers↗

PhyloNaP: a user-friendly database of phylogeny for natural product-producing enzymes.

SUMMARY: Phylogenetic analysis is widely used to predict enzyme function, yet building annotated and reusable trees is labor-intensive and requires extensive knowledge about the specific enzymes. Existing resources rarely cover biosynthetic enzymes and lack the context needed for meaningful analysis. We present PhyloNaP, the first large-scale resource dedicated to phylogenies of biosynthetic enzymes. PhyloNaP provides ∼51 000 annotated and interactive trees enriched with chemical, functional, and taxonomic information. Users can classify their own sequences via phylogenetic placement, enabling functional inference in an evolutionary context. A contribution portal allows the community to submit curated trees. By combining scale, breadth of annotation, and interactive functionality, PhyloNaP fills a major gap in bioinformatics resources for enzyme discovery and annotation, with immediate applications to secondary metabolism and beyond. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://phylonap.cs.uni-tuebingen.de.

Phylogeny↗

FREP: a database of functional repeats in mouse cDNAs.

The FREP database (http://facts.gsc.riken.go.jp/FREP/) contains 31 396 RepeatMasker-identified non-redundant variant repeat sequences derived from 16,527 mouse cDNAs with protein-coding potential. The repeats were computationally associated with potential effects on transcriptional variation, translation, protein function or involvement in disease to identify Functional REPeats (FREPs). FREPs are defined by the (i) occurrence of exon-exon boundaries in repeats, (ii) presence of polyadenylation sites in 3'UTR-located repeats, (iii) effect on translation, (iv) position in the protein- coding region or protein domains or (v) conditional association with disease MeSH terms. Currently the database contains 9261 (29.5%) inferred FREPs derived from 6861 (41.5%) mouse cDNAs. Integrated evidence of the functional assignments and dynamically generated sequence similarity search results support the exploration and annotation of functional, ancestral or taxon-specific repeats. Keyword and pre-selected feature searches (e.g. coding sequence-repeat or splice site-repeat relations) support intuitive database querying as well as the retrieval of repeat sequences. Integrated sequence search and alignment tools allow the analysis of known or identification of new functional repeat candidates. FREP is a unique resource for illuminating the role of transposons and repetitive sequences in shaping the coding part of the mouse transcriptome and for selecting the appropriate experimental model to study diseases with suspected repeat etiology contributions.

Animals↗

Annotation: the cognitive neuroscience of face recognition: implications for developmental disorders.

Face recognition is often considered to be a modular (encapsulated) function. This annotation supports the proposal that faces are special, but suggests that their identification makes use of general-purpose cortical systems that are implicated in high-level vision and also in memory and learning more generally. These systems can be considered to function within two distinct cortical streams: a medial stream (for learning and salience of faces encountered) and a lateral stream (for distributed representations of visual properties and identities of faces). Function in the lateral stream, especially, may be critically dependent on the normal development of magnocellular vision. The relevance of face recognition anomalies in three developmental syndromes (Autism, Williams syndrome, and Turner syndrome) and the two-route model sketched above is considered.

Autistic Disorder↗

Gene functional similarity search tool (GFSST).

BACKGROUND: With the completion of the genome sequences of human, mouse, and other species and the advent of high throughput functional genomic research technologies such as biomicroarray chips, more and more genes and their products have been discovered and their functions have begun to be understood. Increasing amounts of data about genes, gene products and their functions have been stored in databases. To facilitate selection of candidate genes for gene-disease research, genetic association studies, biomarker and drug target selection, and animal models of human diseases, it is essential to have search engines that can retrieve genes by their functions from proteome databases. In recent years, the development of Gene Ontology (GO) has established structured, controlled vocabularies describing gene functions, which makes it possible to develop novel tools to search genes by functional similarity. RESULTS: By using a statistical model to measure the functional similarity of genes based on the Gene Ontology directed acyclic graph, we developed a novel Gene Functional Similarity Search Tool (GFSST) to identify genes with related functions from annotated proteome databases. This search engine lets users design their search targets by gene functions. CONCLUSION: An implementation of GFSST which works on the UniProt (Universal Protein Resource) for the human and mouse proteomes is available at GFSST Web Server. GFSST provides functions not only for similar gene retrieval but also for gene search by one or more GO terms. This represents a powerful new approach for selecting similar genes and gene products from proteome databases according to their functions.

Chromosome Mapping↗

Prediction of protein subcellular locations by support vector machines using compositions of amino acids and amino acid pairs.

MOTIVATION: The subcellular location of a protein is closely correlated to its function. Thus, computational prediction of subcellular locations from the amino acid sequence information would help annotation and functional prediction of protein coding genes in complete genomes. We have developed a method based on support vector machines (SVMs). RESULTS: We considered 12 subcellular locations in eukaryotic cells: chloroplast, cytoplasm, cytoskeleton, endoplasmic reticulum, extracellular medium, Golgi apparatus, lysosome, mitochondrion, nucleus, peroxisome, plasma membrane, and vacuole. We constructed a data set of proteins with known locations from the SWISS-PROT database. A set of SVMs was trained to predict the subcellular location of a given protein based on its amino acid, amino acid pair, and gapped amino acid pair compositions. The predictors based on these different compositions were then combined using a voting scheme. Results obtained through 5-fold cross-validation tests showed an improvement in prediction accuracy over the algorithm based on the amino acid composition only. This prediction method is available via the Internet.

Algorithms↗

Blast2GO: a universal tool for annotation, visualization and analysis in functional genomics research.

SUMMARY: We present here Blast2GO (B2G), a research tool designed with the main purpose of enabling Gene Ontology (GO) based data mining on sequence data for which no GO annotation is yet available. B2G joints in one application GO annotation based on similarity searches with statistical analysis and highlighted visualization on directed acyclic graphs. This tool offers a suitable platform for functional genomics research in non-model species. B2G is an intuitive and interactive desktop application that allows monitoring and comprehension of the whole annotation and analysis process. AVAILABILITY: Blast2GO is freely available via Java Web Start at http://www.blast2go.de. SUPPLEMENTARY MATERIAL: http://www.blast2go.de -> Evaluation.

Algorithms↗

Extracting gene networks for low-dose radiation using graph theoretical algorithms.

Genes with common functions often exhibit correlated expression levels, which can be used to identify sets of interacting genes from microarray data. Microarrays typically measure expression across genomic space, creating a massive matrix of co-expression that must be mined to extract only the most relevant gene interactions. We describe a graph theoretical approach to extracting co-expressed sets of genes, based on the computation of cliques. Unlike the results of traditional clustering algorithms, cliques are not disjoint and allow genes to be assigned to multiple sets of interacting partners, consistent with biological reality. A graph is created by thresholding the correlation matrix to include only the correlations most likely to signify functional relationships. Cliques computed from the graph correspond to sets of genes for which significant edges are present between all members of the set, representing potential members of common or interacting pathways. Clique membership can be used to infer function about poorly annotated genes, based on the known functions of better-annotated genes with which they share clique membership (i.e., "guilt-by-association"). We illustrate our method by applying it to microarray data collected from the spleens of mice exposed to low-dose ionizing radiation. Differential analysis is used to identify sets of genes whose interactions are impacted by radiation exposure. The correlation graph is also queried independently of clique to extract edges that are impacted by radiation. We present several examples of multiple gene interactions that are altered by radiation exposure and thus represent potential molecular pathways that mediate the radiation response.

Algorithms↗