Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

New local potential useful for genome annotation and 3D modeling.

A new potential energy function representing the conformational preferences of sequentially local regions of a protein backbone is presented. This potential is derived from secondary structure probabilities such as those produced by neural network-based prediction methods. The potential is applied to the problem of remote homolog identification, in combination with a distance-dependent inter-residue potential and position-based scoring matrices. This fold recognition jury is implemented in a Java application called JThread. These methods are benchmarked on several test sets, including one released entirely after development and parameterization of JThread. In benchmark tests to identify known folds structurally similar to (but not identical with) the native structure of a sequence, JThread performs significantly better than PSI-BLAST, with 10% more structures identified correctly as the most likely structural match in a fold library, and 20% more structures correctly narrowed down to a set of five possible candidates. JThread also improves the average sequence alignment accuracy significantly, from 53% to 62% of residues aligned correctly. Reliable fold assignments and alignments are identified, making the method useful for genome annotation. JThread is applied to predicted open reading frames (ORFs) from the genomes of Mycoplasma genitalium and Drosophila melanogaster, identifying 20 new structural annotations in the former and 801 in the latter.

Animals↗

Prenatal black carbon exposure and DNA methylation in umbilical cord blood.

BACKGROUND/OBJECTIVES: Prenatal exposure to ambient air pollution is associated with adverse cardiometabolic outcomes in childhood. We previously observed that prenatal black carbon (BC) was inversely associated with adiponectin, a hormone secreted by adipocytes, in early childhood. Changes to DNA methylation have been proposed as a potential mediator linking in utero exposures to lasting health impacts. METHODS: Among 532 mother-child pairs enrolled in the Colorado-based Healthy Start study, we performed an epigenome-wide association study of the relationship between prenatal exposure to a component of air pollution, BC, and DNA methylation in cord blood. Average pregnancy ambient BC was estimated at the mother's residence using a spatiotemporal prediction model. DNA methylation was measured using the Illumina 450K array. We used multiple linear regression to estimate associations between prenatal ambient BC and 429,246 cysteine-phosphate-guanine sites (CpGs), adjusting for potential confounders. We identified differentially methylated regions (DMRs) using DMRff and ENmix-combp. In a subset of participants (n = 243), we investigated DNA methylation as a potential mediator of the association between prenatal ambient BC and lower adiponectin in childhood. RESULTS: We identified 44 CpGs associated with average prenatal ambient BC after correcting for multiple testing. Several genes annotated to the top CpGs had reported functions in the immune system. There were 24 DMRs identified by both DMRff and ENmix-combp. One CpG (cg01123250), located on chromosome 2 and annotated to the UNC80 gene, was found to mediate approximately 20% of the effect of prenatal BC on childhood adiponectin, though the confidence interval was wide (95% CI: 3, 84). CONCLUSIONS: Prenatal BC was associated with DNA methylation in cord blood at several sites and regions in the genome. DNA methylation may partially mediate associations between prenatal BC and childhood cardiometabolic outcomes.

Humans↗

Microarray gene expression profiling and analysis in renal cell carcinoma.

BACKGROUND: Renal cell carcinoma (RCC) is the most common cancer in adult kidney. The accuracy of current diagnosis and prognosis of the disease and the effectiveness of the treatment for the disease are limited by the poor understanding of the disease at the molecular level. To better understand the genetics and biology of RCC, we profiled the expression of 7,129 genes in both clear cell RCC tissue and cell lines using oligonucleotide arrays. METHODS: Total RNAs isolated from renal cell tumors, adjacent normal tissue and metastatic RCC cell lines were hybridized to affymatrix HuFL oligonucleotide arrays. Genes were categorized into different functional groups based on the description of the Gene Ontology Consortium and analyzed based on the gene expression levels. Gene expression profiles of the tissue and cell line samples were visualized and classified by singular value decomposition. Reverse transcription polymerase chain reaction was performed to confirm the expression alterations of selected genes in RCC. RESULTS: Selected genes were annotated based on biological processes and clustered into functional groups. The expression levels of genes in each group were also analyzed. Seventy-four commonly differentially expressed genes with more than five-fold changes in RCC tissues were identified. The expression alterations of selected genes from these seventy-four genes were further verified using reverse transcription polymerase chain reaction (RT-PCR). Detailed comparison of gene expression patterns in RCC tissue and RCC cell lines shows significant differences between the two types of samples, but many important expression patterns were preserved. CONCLUSIONS: This is one of the initial studies that examine the functional ontology of a large number of genes in RCC. Extensive annotation, clustering and analysis of a large number of genes based on the gene functional ontology revealed many interesting gene expression patterns in RCC. Most notably, genes involved in cell adhesion were dominantly up-regulated whereas genes involved in transport were dominantly down-regulated. This study reveals significant gene expression alterations in key biological pathways and provides potential insights into understanding the molecular mechanism of renal cell carcinogenesis.

Adenocarcinoma, Clear Cell↗

Domain-based small molecule binding site annotation.

BACKGROUND: Accurate small molecule binding site information for a protein can facilitate studies in drug docking, drug discovery and function prediction, but small molecule binding site protein sequence annotation is sparse. The Small Molecule Interaction Database (SMID), a database of protein domain-small molecule interactions, was created using structural data from the Protein Data Bank (PDB). More importantly it provides a means to predict small molecule binding sites on proteins with a known or unknown structure and unlike prior approaches, removes large numbers of false positive hits arising from transitive alignment errors, non-biologically significant small molecules and crystallographic conditions that overpredict ion binding sites. DESCRIPTION: Using a set of co-crystallized protein-small molecule structures as a starting point, SMID interactions were generated by identifying protein domains that bind to small molecules, using NCBI's Reverse Position Specific BLAST (RPS-BLAST) algorithm. SMID records are available for viewing at http://smid.blueprint.org. The SMID-BLAST tool provides accurate transitive annotation of small-molecule binding sites for proteins not found in the PDB. Given a protein sequence, SMID-BLAST identifies domains using RPS-BLAST and then lists potential small molecule ligands based on SMID records, as well as their aligned binding sites. A heuristic ligand score is calculated based on E-value, ligand residue identity and domain entropy to assign a level of confidence to hits found. SMID-BLAST predictions were validated against a set of 793 experimental small molecule interactions from the PDB, of which 472 (60%) of predicted interactions identically matched the experimental small molecule and of these, 344 had greater than 80% of the binding site residues correctly identified. Further, we estimate that 45% of predictions which were not observed in the PDB validation set may be true positives. CONCLUSION: By focusing on protein domain-small molecule interactions, SMID is able to cluster similar interactions and detect subtle binding patterns that would not otherwise be obvious. Using SMID-BLAST, small molecule targets can be predicted for any protein sequence, with the only limitation being that the small molecule must exist in the PDB. Validation results and specific examples within illustrate that SMID-BLAST has a high degree of accuracy in terms of predicting both the small molecule ligand and binding site residue positions for a query protein.

Binding Sites↗

A functional profile of gene expression in ARPE-19 cells.

BACKGROUND: Retinal pigment epithelium cells play an important role in the pathogenesis of age related macular degeneration. Their morphological, molecular and functional phenotype changes in response to various stresses. Functional profiling of genes can provide useful information about the physiological state of cells and how this state changes in response to disease or treatment. In this study, we have constructed a functional profile of the genes expressed by the ARPE-19 cell line of retinal pigment epithelium. METHODS: Using Affymetrix MAS 5.0 microarray analysis, genes expressed by ARPE-19 cells were identified. Using GeneChip annotations, these genes were classified according to their known functions to generate a functional gene expression profile. RESULTS: We have determined that of approximately 19,044 unique gene sequences represented on the HG-U133A GeneChip, 6,438 were expressed in ARPE-19 cells irrespective of the substrate on which they were grown (plastic, fibronectin, collagen, or Matrigel). Rather than focus our subsequent analysis on the identity or level of expression of each individual gene in this large data set, we examined the number of genes expressed within 130 functional categories. These categories were selected from a library of HG-U133A GeneChip annotations linked to the Affymetrix MAS 5.0 data sets. Using this functional classification scheme, we were able to categorize about 70% of the expressed genes and condense the original data set of over 6,000 data points into a format with 130 data points. The resulting ARPE-19 Functional Gene Expression Profile is displayed as a percentage of ARPE-19-expressed genes. CONCLUSION: The Profile can readily be compared with equivalent microarray data from other appropriate samples in order to highlight cell-specific attributes or treatment-induced changes in gene expression. The usefulness of these analyses is based on the assumption that the numbers of genes expressed within a functional category provide an indicator of the overall level of activity within that particular functional pathway.

Cell Line↗

An organism-specific method to rank predicted coding regions in Trypanosoma brucei.

Genome annotation in differently evolved organisms presents challenges because the lack of sequence-based homology limits the ability to determine the function of putative coding regions. To provide an alternative to annotation by sequence homology, we developed a method that takes advantage of unusual trypanosomatid biology and skews in nucleotide composition between coding regions and upstream regions to rank putative open reading frames based on the likelihood of coding. The method is 93% accurate when tested on known genes. We have applied our method to the full complement of open reading frames on Chromosome I of Trypanosoma brucei, and we can predict with high confidence that 226 putative coding regions are likely to be functional. Methods such as the one described here for discriminating true coding regions are critical for genome annotation when other sources of evidence for function are limited.

Animals↗

Structural genomics of minimal organisms and protein fold space.

The initial aim of the Berkeley Structural Genomics Center is to obtain a near-complete structural complement of two minimal organisms, closely related pathogens Mycoplasma genitalium and M. pneumoniae. The former has fewer than 500 genes and the latter fewer than 700 genes. To achieve this goal, the current protein targets have been selected starting with those predicted to be most tractable and likely to yield new structural and functional information. During the past 3 years, the semi-automated structural genomics pipeline has been set up from cloning, expression, purification, and ultimately to structural determination. The results from the pipeline substantially increased the coverage of the protein fold space of M. pneumoniae and M. genitalium. Furthermore, about 1/2 of the structures of 'unique' protein sequences revealed new and novel folds, and over 2/3 of the structures of previously annotated 'hypothetical proteins' inferred their molecular functions.

Bacterial Proteins↗

Will we ever understand? The undescribable diversity of the prokaryotes.

This communication will summarize recent estimations on prokaryotic cell numbers, technical aspects of the exploration of the hidden diversity of the as-yet-uncultured prokaryotes and their function in the environment, recognition of novel major lines of descent, elucidation of novel metabolic pathways and attempts to improve the definition of the taxon species. These are some of recent highlights that will reflect only incompletely recent advances of microbiologist and other aspects of diversity at the genomic level, including the tremendous influence of whole genome sequences on the development of DNA macro- and microarrays for rapid identification of genes and specimen, the annotation of gene sequences to gene function by proteomics and the recognition of the extent of lateral gene transfer in the evolution of the genome, hence of contemporary organisms. As an example of modern trends in systematics three new families will be described for recently described genera, namely Thermoleophilaceae, Solirubrobacteraceae and Conexibacteraceae, which are phylogenetically positioned among environmental clone sequences at the root of the class Actinobacteria.

Bacteria↗

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis↗

RiceGAAS: an automated annotation system and database for rice genome sequence.

An extensive effort of the International Rice Genome Sequencing Project (IRGSP) has resulted in rapid accumulation of genome sequence, and >137 Mb has already been made available to the public domain as of August 2001. This requires a high-throughput annotation scheme to extract biologically useful and timely information from the sequence data on a regular basis. A new automated annotation system and database called Rice Genome Automated Annotation System (RiceGAAS) has been developed to execute a reliable and up-to-date analysis of the genome sequence as well as to store and retrieve the results of annotation. The system has the following functional features: (i) collection of rice genome sequences from GenBank; (ii) execution of gene prediction and homology search programs; (iii) integration of results from various analyses and automatic interpretation of coding regions; (iv) re-execution of analysis, integration and automatic interpretation with the latest entries in reference databases; (v) integrated visualization of the stored data using web-based graphical view. RiceGAAS also has a data submission mechanism that allows public users to perform fully automated annotation of their own sequences. The system can be accessed at http://RiceGAAS.dna.affrc.go.jp/.

Automation↗

The Adult Mouse Anatomical Dictionary: a tool for annotating and integrating data.

We have developed an ontology to provide standardized nomenclature for anatomical terms in the postnatal mouse. The Adult Mouse Anatomical Dictionary is structured as a directed acyclic graph, and is organized hierarchically both spatially and functionally. The ontology will be used to annotate and integrate different types of data pertinent to anatomy, such as gene expression patterns and phenotype information, which will contribute to an integrated description of biological phenomena in the mouse.

Animals↗

Development and application of a salmonid EST database and cDNA microarray: data mining and interspecific hybridization characteristics.

We report 80,388 ESTs from 23 Atlantic salmon (Salmo salar) cDNA libraries (61,819 ESTs), 6 rainbow trout (Oncorhynchus mykiss) cDNA libraries (14,544 ESTs), 2 chinook salmon (Oncorhynchus tshawytscha) cDNA libraries (1317 ESTs), 2 sockeye salmon (Oncorhynchus nerka) cDNA libraries (1243 ESTs), and 2 lake whitefish (Coregonus clupeaformis) cDNA libraries (1465 ESTs). The majority of these are 3' sequences, allowing discrimination between paralogs arising from a recent genome duplication in the salmonid lineage. Sequence assembly reveals 28,710 different S. salar, 8981 O. mykiss, 1085 O. tshawytscha, 520 O. nerka, and 1176 C. clupeaformis putative transcripts. We annotate the submitted portion of our EST database by molecular function. Higher- and lower-molecular-weight fractions of libraries are shown to contain distinct gene sets, and higher rates of gene discovery are associated with higher-molecular weight libraries. Pyloric caecum library group annotations indicate this organ may function in redox control and as a barrier against systemic uptake of xenobiotics. A microarray is described, containing 7356 salmonid elements representing 3557 different cDNAs. Analyses of cross-species hybridizations to this cDNA microarray indicate that this resource may be used for studies involving all salmonids.

Animals↗

Whole metagenome sequencing: not deep enough for complete microbial function recovery.

BACKGROUND: Whole metagenome shotgun sequencing (WMS) is widely used to profile microbial function. However, technical variability in sequencing and analysis often obscures true biological patterns. Large-scale studies are particularly susceptible to batch effects, such as differences in sequencing depth and platform and annotation strategies, as well as sample-to-flow-cell assignments. However, the relative effects of these factors on functional inference in such studies have yet to be systematically evaluated. We analyzed oral-rinse WMS data from 671 Nigerian youths aged 9-18, sequenced on two Illumina platforms. Microbial molecular functionality encoded in these data was annotated using the mi-faser/Fusion pipeline, to capture the broad functional repertoire, and HUMAnN 3/EC numbers pipeline to characterize curated enzymatic activities. We then quantified how technical factors and batch effects shaped the recovery of microbial functionality. RESULTS: Three findings of our work were most salient. First, we observed that the choice of annotation strategy traded off between breadth and specificity of functional coverage. Second, we found that low-prevalence functions were disproportionately lost at shallow sequencing depths, indicating that in, e.g., case-control studies with few representatives of the minor class, sequencing depth could critically impact study resolution. Finally, using our newly developed model relating sequencing depth to functional recovery, we demonstrated that increasing sequencing depth does not directly or proportionally improve functional recall. That is, at as little as 10% of this study's sequencing depth, 30% of the estimated complete microbiome functional repertoire was detectable. However, even at the full depth used in this study, we were only able to recover an estimated 60% of that complete functional repertoire. We further showed that despite biomes differences in functional diversity and host contamination levels (e.g., soil, fecal), incomplete functional recovery at commonly used sequencing depths was consistently observed. CONCLUSIONS: Together, these findings and our depth-to-function mapping framework provide practical guidelines for the design and interpretation of WMS studies. Coordinating sequencing depth planning with annotation strategy, experimental design, and rigorous batch control is thus essential for robust detection of microbial functions and for ensuring reproducible microbiome insights. Video Abstract.

Humans↗

Gene annotation and network inference by phylogenetic profiling.

BACKGROUND: Phylogenetic analysis is emerging as one of the most informative computational methods for the annotation of genes and identification of evolutionary modules of functionally related genes. The effectiveness with which phylogenetic profiles can be utilized to assign genes to pathways depends on an appropriate measure of correlation between gene profiles, and an effective decision rule to use the correlate. Current methods, though useful, perform at a level well below what is possible, largely because performance of the latter deteriorates rapidly as coverage increases. RESULTS: We introduce, test and apply a new decision rule, correlation enrichment (CE), for assigning genes to functional categories at various levels of resolution. Among the results are: (1) CE performs better than standard guilt by association (SGA, assignment to a functional category when a simple correlate exceeds a pre-specified threshold) irrespective of the number of genes assigned (i.e. coverage); improvement is greatest at high coverage where precision (positive predictive value) of CE is approximately 6-fold higher than that of SGA. (2) CE is estimated to allocate each of the 2918 unannotated orthologs to KEGG pathways with an average precision of 49% (approximately 7-fold higher than SGA) (3) An estimated 94% of the 1846 unannotated orthologs in the COG ontology can be assigned a function with an average precision of 0.4 or greater. (4) Dozens of functional and evolutionarily conserved cliques or quasi-cliques can be identified, many having previously unannotated genes. CONCLUSION: The method serves as a general computational tool for annotating large numbers of unknown genes, uncovering evolutionary and functional modules. It appears to perform substantially better than extant stand alone high throughout methods.

Algorithms↗

Fast and accurate method for identifying high-quality protein-interaction modules by clique merging and its application to yeast.

Molecular networks in cells are organized into functional modules, where genes in the same module interact densely with each other and participate in the same biological process. Thus, identification of modules from molecular networks is an important step toward a better understanding of how cells function through the molecular networks. Here, we propose a simple, automatic method, called MC(2), to identify functional modules by enumerating and merging cliques in the protein-interaction data from large-scale experiments. Application of MC(2) to the S. cerevisiae protein-interaction data produces 84 modules, whose sizes range from 4 to 69 genes. The majority of the discovered modules are significantly enriched with a highly specific process term (at least 4 levels below root) and a specific cellular component in Gene Ontology (GO) tree. The average fraction of genes with the most enriched GO term for all modules is 82% for specific biological processes and 78% for specific cellular components. In addition, the predicted modules are enriched with coexpressed proteins. These modules are found to be useful for annotating unknown genes and uncovering novel functions of known genes. MC(2) is efficient, and takes only about 5 min to identify modules from the current yeast gene interaction network with a typical PC (Intel Xeon 2.5 GHz CPU and 512 MB memory). The CPU time of MC(2) is affordable (12 h) even when the number of interactions is increased by a factor of 10. MC(2) and its results are publicly available on http://theory.med.buffalo.edu/MC2.

Algorithms↗

PartsList: a web-based system for dynamically ranking protein folds based on disparate attributes, including whole-genome expression and interaction information.

As the number of protein folds is quite limited, a mode of analysis that will be increasingly common in the future, especially with the advent of structural genomics, is to survey and re-survey the finite parts list of folds from an expanding number of perspectives. We have developed a new resource, called PartsList, that lets one dynamically perform these comparative fold surveys. It is available on the web at http://bioinfo.mbb.yale.edu/partslist and http://www.partslist.org. The system is based on the existing fold classifications and functions as a form of companion annotation for them, providing 'global views' of many already completed fold surveys. The central idea in the system is that of comparison through ranking; PartsList will rank the approximately 420 folds based on more than 180 attributes. These include: (i) occurrence in a number of completely sequenced genomes (e.g. it will show the most common folds in the worm versus yeast); (ii) occurrence in the structure databank (e.g. most common folds in the PDB); (iii) both absolute and relative gene expression information (e.g. most changing folds in expression over the cell cycle); (iv) protein-protein interactions, based on experimental data in yeast and comprehensive PDB surveys (e.g. most interacting fold); (v) sensitivity to inserted transposons; (vi) the number of functions associated with the fold (e.g. most multi-functional folds); (vii) amino acid composition (e.g. most Cys-rich folds); (viii) protein motions (e.g. most mobile folds); and (ix) the level of similarity based on a comprehensive set of structural alignments (e.g. most structurally variable folds). The integration of whole-genome expression and protein-protein interaction data with structural information is a particularly novel feature of our system. We provide three ways of visualizing the rankings: a profiler emphasizing the progression of high and low ranks across many pre-selected attributes, a dynamic comparer for custom comparisons and a numerical rankings correlator. These allow one to directly compare very different attributes of a fold (e.g. expression level, genome occurrence and maximum motion) in the uniform numerical format of ranks. This uniform framework, in turn, highlights the way that the frequency of many of the attributes falls off with approximate power-law behavior (i.e. according to V(-b), for attribute value V and constant exponent b), with a few folds having large values and most having small values.

Cysteine↗

ASAP, a systematic annotation package for community analysis of genomes.

ASAP (a systematic annotation package for community analysis of genomes) is a relational database and web interface developed to store, update and distribute genome sequence data and functional characterization (https://asap.ahabs.wisc.edu/annotation/php/ASAP1.htm). ASAP facilitates ongoing community annotation of genomes and tracking of information as genome projects move from preliminary data collection through post-sequencing functional analysis. The ASAP database includes multiple genome sequences at various stages of analysis, corresponding experimental data and access to collections of related genome resources. ASAP supports three levels of users: public viewers, annotators and curators. Public viewers can currently browse updated annotation information for Escherichia coli K-12 strain MG1655, genome-wide transcript profiles from more than 50 microarray experiments and an extensive collection of mutant strains and associated phenotypic data. Annotators worldwide are currently using ASAP to participate in a community annotation project for the Erwinia chrysanthemi strain 3937 genome. Curation of the E. chrysanthemi genome annotation as well as those of additional published enterobacterial genomes is underway and will be publicly accessible in the near future.

Databases, Genetic↗

On methods for gene function scoring as a means of facilitating the interpretation of microarray results.

As gene annotation databases continue to evolve and improve, it has become feasible to incorporate the functional and pathway information about genes, available in these databases into the analysis of gene expression data, for a better understanding of the underlying mechanisms. A few methods have been proposed in the literature to formally convert individual gene results into gene function results. In this paper, we will compare the various methods, propose and examine some new ones, and offer a structured approach to incorporating gene function or pathway information into the analysis of expression data. We study the performance of the various methods and also compare them on real data, using a case study from the toxicogenomics area. Our results show that the approaches based on gene function scores yield a different, and functionally more interpretable, array of genes than methods that rely solely on individual gene scores. They also suggest that functional class scoring methods appear to perform better and more consistently than overrepresentation analysis and distributional score methods.

Databases, Genetic↗