Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Clinical phenotype and gene expression profile in Crohn's disease.

The clinical course varies significantly among patients with Crohn's disease (CD). This study investigated whether gene expression profiles generated by DNA microarray technology might predict disease progression. Biopsies from the descending colon were obtained colonoscopically from 40 CD patients. Gene profiling analyses were performed using a Human Genome U133 Plus 2.0 GeneChip Array, and summarization into a single expression measure for each probe set was performed using the robust multiple array procedure. Principal component analysis demonstrated that three components explain two-thirds of the total variation. The most important parameters for the determination of the colonic gene expression patterns were the presence of disease (CD) and presence of inflammation. Superimposition of clinical phenotype data revealed a grouping of the samples from patients with stenosis toward negative values on the axis of the second principal component. The functional annotation analysis suggested that the expression of genes involved in intracellular transport and cytoskeletal organization might influence the development of stenosis. In conclusion, even though most variation in the colonic gene expression patterns is due to presence or absence of CD and inflammation status, the development of stenosis is a parameter that affects colonic gene expression to some extent.

Adolescent↗

Microarray analyses identify molecular biomarkers of Atlantic salmon macrophage and hematopoietic kidney response to Piscirickettsia salmonis infection.

Piscirickettsia salmonis is the intracellular bacterium that causes salmonid rickettsial septicemia, an infectious disease that kills millions of farmed fish each year. The mechanisms used by P. salmonis to survive and replicate within host cells are not known. Piscirickettsiosis causes severe necrosis of hematopoietic kidney. Microarray-based experiments with QPCR validation were used to identify Atlantic salmon macrophage and hematopoietic kidney genes differentially transcribed in response to P. salmonis infection. Infections were confirmed by microscopy and RT-PCR with pathogen-specific primers. In infected salmon macrophages, 71 different transcripts were upregulated and 31 different transcripts were downregulated. In infected hematopoietic kidney, 30 different transcripts were upregulated and 39 different transcripts were downregulated. Ten antioxidant genes, including glutathione S-transferase, glutathione reductase, glutathione peroxidase, and cytochrome b558 alpha- and beta-subunits, were upregulated in infected macrophages but not in infected hematopoietic kidney. Changes in redox status of infected macrophages may allow these cells to tolerate P. salmonis infection, raising the possibility that treatment with antioxidants may reduce hematopoietic tissue damage caused by this rickettsial infection. The downregulation of transcripts involved in adaptive immune responses (e.g., T cell receptor alpha-chain and C-C chemokine receptor 7) in infected hematopoietic kidney but not in infected macrophages may contribute to infection-induced kidney tissue damage. Molecular biomarkers of P. salmonis infection, characterized by immune-relevant functional annotations and high fold differences in expression between infected and noninfected samples, may aid in the development of anti-piscirickettsial vaccines and therapeutics.

Animals↗

Quantitative trait locus dissection in congenic strains of the Goto-Kakizaki rat identifies a region conserved with diabetes loci in human chromosome 1q.

Genetic studies in human populations and rodent models have identified regions of human chromosome 1q21-25 and rat chromosome 2 showing evidence of significant and replicated linkage to diabetes-related phenotypes. To investigate the relationship between the human and rat diabetes loci, we fine mapped the rat locus Nidd/gk2 linked to hyperinsulinemia in an F2 cross derived from the diabetic (type 2) Goto-Kakizaki (GK) rat and the Brown Norway (BN) control rat, and carried out its genetic and pathophysiological characterization in BN.GK congenic strains. Evidence of glucose intolerance and enhanced insulin secretion in a congenic strain allowed us to localize the underlying diabetes gene(s) in a rat chromosomal interval of approximately 3-6 cM conserved with an 11-Mb region of human 1q21-23. Positional diabetes candidate genes were tested for transcriptional changes between congenics and controls and sequence variations in a panel of inbred rat strains. Congenic strains of the GK rats represent powerful novel models for accurately defining the pathophysiological impact of diabetes gene(s) at the locus Nidd/gk2 and improving functional annotations of diabetes candidates in human 1q21-23.

Animals↗

Conventional and Shared Genetic Association Analysis Between Diabetes Mellitus and Sensorineural Hearing Loss.

PURPOSE: This study aims to investigate the epidemiological and genetic associations between diabetes mellitus (DM) and sensorineural hearing loss (SNHL) across different subtypes. METHODS: We analyzed 502,490 participants from the UK Biobank using multivariate logistic regression to examine the association between DM and SNHL, considering gender, age, and HbA1c levels. Genetic correlations and causality were examined by linkage disequilibrium score regression and bidirectional Mendelian randomization. Cross-trait meta-analyses identified shared loci between DM and SNHL, followed by gene annotation, functional analysis, and drug candidate exploration for the shared traits. RESULTS: Observational analysis revealed significant associations between DM and SNHL, consistent in subgroups based on age, sex, and certain HbA1c levels. A positive genetic correlation was found between type 2 diabetes mellitus (T2D) and SNHL (Rg = 0.0982, p = 0.0095) between T2D and SNHL, and four loci were identified, with ARHGEF28 and TCF7L2 prioritized as credible pleiotropic genes. Enrichment was indicated in glucose metabolism and organogenesis, with shared heritability in metabolic tissues and outer hair cells. Metformin was identified as potential drug candidates for the T2D-SNHL comorbidity. CONCLUSION: These findings progress our understanding of the epidemiological association, shared genetic basis, and potential therapeutic targets between T2D and SNHL, which might contribute to the management of their comorbidity.

Humans↗

Identification of a PAX-FKHR gene expression signature that defines molecular classes and determines the prognosis of alveolar rhabdomyosarcomas.

Alveolar rhabdomyosarcomas (ARMS) are aggressive soft-tissue sarcomas affecting children and young adults. Most ARMS tumors express the PAX3-FKHR or PAX7-FKHR (PAX-FKHR) fusion genes resulting from the t(2;13) or t(1;13) chromosomal translocations, respectively. However, up to 25% of ARMS tumors are fusion negative, making it unclear whether ARMS represent a single disease or multiple clinical and biological entities with a common phenotype. To test to what extent PAX-FKHR determine class and behavior of ARMS, we used oligonucleotide microarray expression profiling on 139 primary rhabdomyosarcoma tumors and an in vitro model. We found that ARMS tumors expressing either PAX-FKHR gene share a common expression profile distinct from fusion-negative ARMS and from the other rhabdomyosarcoma variants. We also observed that PAX-FKHR expression above a minimum level is necessary for the detection of this expression profile. Using an ectopic PAX3-FKHR and PAX7-FKHR expression model, we identified an expression signature regulated by PAX-FKHR that is specific to PAX-FKHR-positive ARMS tumors. Data mining for functional annotations of signature genes suggested a role for PAX-FKHR in regulating ARMS proliferation and differentiation. Cox regression modeling identified a subset of genes within the PAX-FKHR expression signature that segregated ARMS patients into three risk groups with 5-year overall survival estimates of 7%, 48%, and 93%. These prognostic classes were independent of conventional clinical risk factors. Our results show that PAX-FKHR dictate a specific expression signature that helps define the molecular phenotype of PAX-FKHR-positive ARMS tumors and, because it is linked with disease outcome in ARMS patients, determine tumor behavior.

Biomarkers, Tumor↗

Gene expression profiling reveals a massive, aneuploidy-dependent transcriptional deregulation and distinct differences between lymph node-negative and lymph node-positive colon carcinomas.

To characterize patterns of global transcriptional deregulation in primary colon carcinomas, we did gene expression profiling of 73 tumors [Unio Internationale Contra Cancrum stage II (n = 33) and stage III (n = 40)] using oligonucleotide microarrays. For 30 of the tumors, expression profiles were compared with those from matched normal mucosa samples. We identified a set of 1,950 genes with highly significant deregulation between tumors and mucosa samples (P < 1e-7). A significant proportion of these genes mapped to chromosome 20 (P = 0.01). Seventeen genes had a >5-fold average expression difference between normal colon mucosa and carcinomas, including up-regulation of MYC and of HMGA1, a putative oncogene. Furthermore, we identified 68 genes that were significantly differentially expressed between lymph node-negative and lymph node-positive tumors (P < 0.001), the functional annotation of which revealed a preponderance of genes that play a role in cellular immune response and surveillance. The microarray-derived gene expression levels of 20 deregulated genes were validated using quantitative real-time reverse transcription-PCR in >40 tumor and normal mucosa samples with good concordance between the techniques. Finally, we established a relationship between specific genomic imbalances, which were mapped for 32 of the analyzed colon tumors by comparative genomic hybridization, and alterations of global transcriptional activity. Previously, we had conducted a similar analysis of primary rectal carcinomas. The systematic comparison of colon and rectal carcinomas revealed a significant overlap of genomic imbalances and transcriptional deregulation, including activation of the Wnt/beta-catenin signaling cascade, suggesting similar pathogenic pathways.

Adenocarcinoma↗

Sequence analysis of a rainbow trout cDNA library and creation of a gene index.

Expressed sequence tag (EST) projects have produced extremely valuable resources for identifying genes affecting phenotypes of interest. A large-scale EST sequencing project for rainbow trout was initiated to identify and functionally annotate as many unique transcripts as possible. Over 45,000 5' ESTs were obtained by sequencing clones from a single normalized library constructed using mRNA from six tissues. The production of this sequence data and creation of a rainbow trout Gene Index eliminating redundancy and providing annotation for these sequences will facilitate research in this species.

Animals↗

MiCoViTo: a tool for gene-centric comparison and visualization of yeast transcriptome states.

BACKGROUND: Information obtained by DNA microarray technology gives a rough snapshot of the transcriptome state, i.e., the expression level of all the genes expressed in a cell population at any given time. One of the challenging questions raised by the tremendous amount of microarray data is to identify groups of co-regulated genes and to understand their role in cell functions. RESULTS: MiCoViTo (Microarray Comparison Visualization Tool) is a set of biologists' tools for exploring, comparing and visualizing changes in the yeast transcriptome by a gene-centric approach. A relational database includes data linked to genome expression and graphical output makes it easy to visualize clusters of co-expressed genes in the context of available biological information. To this aim, upload of personal data is possible and microarray data from fifty publications dedicated to S. cerevisiae are provided on-line. A web interface guides the biologist during the usage of this tool and is freely accessible at http://www.transcriptome.ens.fr/micovito/. CONCLUSIONS: MiCoViTo offers an easy-to-read picture of local transcriptional changes connected to current biological knowledge. This should help biologists to mine yeast microarray data and better understand the underlying biology. We plan to add functional annotations from other organisms. That would allow inter-species comparison of transcriptomes via orthology tables.

Cluster Analysis↗

Iterative Group Analysis (iGA): a simple tool to enhance sensitivity and facilitate interpretation of microarray experiments.

BACKGROUND: The biological interpretation of even a simple microarray experiment can be a challenging and highly complex task. Here we present a new method (Iterative Group Analysis) to facilitate, improve, and accelerate this process. RESULTS: Our Iterative Group Analysis approach (iGA) uses elementary statistics to identify those functional classes of genes that are significantly changed in an experiment and at the same time determines which of the class members are most likely to be differentially expressed. iGA does not require that all members of a class change and is therefore robust against imperfect class assignments, which can be derived from public sources (e.g. GeneOntologies) or automated processes (e.g. key word extraction from gene names). In contrast to previous non-iterative approaches, iGA does not depend on the availability of fixed lists of differentially expressed genes, and thus can be used to increase the sensitivity of gene detection especially in very noisy or small data sets. In the extreme, iGA can even produce statistically meaningful results without any experimental replication. The automated functional annotation provided by iGA greatly reduces the complexity of microarray results and facilitates the interpretation process. In addition, iGA can be used as a fast and efficient tool for the platform-independent comparison of a microarray experiment to the vast number of published results, automatically highlighting shared genes of potential interest. CONCLUSIONS: By applying iGA to a wide variety of data from diverse organisms and platforms we show that this approach enhances and accelerates the interpretation of microarray experiments.

Animals↗

Ontological visualization of protein-protein interactions.

BACKGROUND: Cellular processes require the interaction of many proteins across several cellular compartments. Determining the collective network of such interactions is an important aspect of understanding the role and regulation of individual proteins. The Gene Ontology (GO) is used by model organism databases and other bioinformatics resources to provide functional annotation of proteins. The annotation process provides a mechanism to document the binding of one protein with another. We have constructed protein interaction networks for mouse proteins utilizing the information encoded in the GO annotations. The work reported here presents a methodology for integrating and visualizing information on protein-protein interactions. RESULTS: GO annotation at Mouse Genome Informatics (MGI) captures 1318 curated, documented interactions. These include 129 binary interactions and 125 interaction involving three or more gene products. Three networks involve over 30 partners, the largest involving 109 proteins. Several tools are available at MGI to visualize and analyze these data. CONCLUSIONS: Curators at the MGI database annotate protein-protein interaction data from experimental reports from the literature. Integration of these data with the other types of data curated at MGI places protein binding data into the larger context of mouse biology and facilitates the generation of new biological hypotheses based on physical interactions among gene products.

Animals↗

Atlas - a data warehouse for integrative bioinformatics.

BACKGROUND: We present a biological data warehouse called Atlas that locally stores and integrates biological sequences, molecular interactions, homology information, functional annotations of genes, and biological ontologies. The goal of the system is to provide data, as well as a software infrastructure for bioinformatics research and development. DESCRIPTION: The Atlas system is based on relational data models that we developed for each of the source data types. Data stored within these relational models are managed through Structured Query Language (SQL) calls that are implemented in a set of Application Programming Interfaces (APIs). The APIs include three languages: C++, Java, and Perl. The methods in these API libraries are used to construct a set of loader applications, which parse and load the source datasets into the Atlas database, and a set of toolbox applications which facilitate data retrieval. Atlas stores and integrates local instances of GenBank, RefSeq, UniProt, Human Protein Reference Database (HPRD), Biomolecular Interaction Network Database (BIND), Database of Interacting Proteins (DIP), Molecular Interactions Database (MINT), IntAct, NCBI Taxonomy, Gene Ontology (GO), Online Mendelian Inheritance in Man (OMIM), LocusLink, Entrez Gene and HomoloGene. The retrieval APIs and toolbox applications are critical components that offer end-users flexible, easy, integrated access to this data. We present use cases that use Atlas to integrate these sources for genome annotation, inference of molecular interactions across species, and gene-disease associations. CONCLUSION: The Atlas biological data warehouse serves as data infrastructure for bioinformatics research and development. It forms the backbone of the research activities in our laboratory and facilitates the integration of disparate, heterogeneous biological sources of data enabling new scientific inferences. Atlas achieves integration of diverse data sets at two levels. First, Atlas stores data of similar types using common data models, enforcing the relationships between data types. Second, integration is achieved through a combination of APIs, ontology, and tools. The Atlas software is freely available under the GNU General Public License at: http://bioinformatics.ubc.ca/atlas/

Computational Biology↗

Speeding disease gene discovery by sequence based candidate prioritization.

BACKGROUND: Regions of interest identified through genetic linkage studies regularly exceed 30 centimorgans in size and can contain hundreds of genes. Traditionally this number is reduced by matching functional annotation to knowledge of the disease or phenotype in question. However, here we show that disease genes share patterns of sequence-based features that can provide a good basis for automatic prioritization of candidates by machine learning. RESULTS: We examined a variety of sequence-based features and found that for many of them there are significant differences between the sets of genes known to be involved in human hereditary disease and those not known to be involved in disease. We have created an automatic classifier called PROSPECTR based on those features using the alternating decision tree algorithm which ranks genes in the order of likelihood of involvement in disease. On average, PROSPECTR enriches lists for disease genes two-fold 77% of the time, five-fold 37% of the time and twenty-fold 11% of the time. CONCLUSION: PROSPECTR is a simple and effective way to identify genes involved in Mendelian and oligogenic disorders. It performs markedly better than the single existing sequence-based classifier on novel data. PROSPECTR could save investigators looking at large regions of interest time and effort by prioritizing positional candidate genes for mutation detection and case-control association studies.

Algorithms↗

INTEGRATOR: interactive graphical search of large protein interactomes over the Web.

BACKGROUND: The rapid growth of protein interactome data has elevated the necessity and importance of network analysis tools. However, unlike pure text data, network search spaces are of exponential complexity. This poses special challenges for storing, searching, and navigating this data efficiently. Moreover, development of effective web interfaces has been difficult. RESULTS: We present Integrator, a web-integrated graphical search tool for protein-protein interaction networks across 50+ genomes. CONCLUSION: Integrator provides single and multiple protein searches of the Bioverse database containing experimentally-derived and predicted protein-protein interactions. The interface provides animated local network views, rapid subgraph manipulation, and cross-referencing of functional annotations. Integrator is available at http://bioverse.compbio.washington.edu/integrator.

Algorithms↗

An integrated approach to the prediction of domain-domain interactions.

BACKGROUND: The development of high-throughput technologies has produced several large scale protein interaction data sets for multiple species, and significant efforts have been made to analyze the data sets in order to understand protein activities. Considering that the basic units of protein interactions are domain interactions, it is crucial to understand protein interactions at the level of the domains. The availability of many diverse biological data sets provides an opportunity to discover the underlying domain interactions within protein interactions through an integration of these biological data sets. RESULTS: We combine protein interaction data sets from multiple species, molecular sequences, and gene ontology to construct a set of high-confidence domain-domain interactions. First, we propose a new measure, the expected number of interactions for each pair of domains, to score domain interactions based on protein interaction data in one species and show that it has similar performance as the E-value defined by Riley et al. Our new measure is applied to the protein interaction data sets from yeast, worm, fruitfly and humans. Second, information on pairs of domains that coexist in known proteins and on pairs of domains with the same gene ontology function annotations are incorporated to construct a high-confidence set of domain-domain interactions using a Bayesian approach. Finally, we evaluate the set of domain-domain interactions by comparing predicted domain interactions with those defined in iPfam database that were derived based on protein structures. The accuracy of predicted domain interactions are also confirmed by comparing with experimentally obtained domain interactions from H. pylori. As a result, a total of 2,391 high-confidence domain interactions are obtained and these domain interactions are used to unravel detailed protein and domain interactions in several protein complexes. CONCLUSION: Our study shows that integration of multiple biological data sets based on the Bayesian approach provides a reliable framework to predict domain interactions. By integrating multiple data sources, the coverage and accuracy of predicted domain interactions can be significantly increased.

Algorithms↗

On single and multiple models of protein families for the detection of remote sequence relationships.

BACKGROUND: The detection of relationships between a protein sequence of unknown function and a sequence whose function has been characterised enables the transfer of functional annotation. However in many cases these relationships can not be identified easily from direct comparison of the two sequences. Methods which compare sequence profiles have been shown to improve the detection of these remote sequence relationships. However, the best method for building a profile of a known set of sequences has not been established. Here we examine how the type of profile built affects its performance, both in detecting remote homologs and in the resulting alignment accuracy. In particular, we consider whether it is better to model a protein superfamily using a single structure-based alignment that is representative of all known cases of the superfamily, or to use multiple sequence-based profiles each representing an individual member of the superfamily. RESULTS: Using profile-profile methods for remote homolog detection we benchmark the performance of single structure-based superfamily models and multiple domain models. On average, over all superfamilies, using a truncated receiver operator characteristic (ROC5) we find that multiple domain models outperform single superfamily models, except at low error rates where the two models behave in a similar way. However there is a wide range of performance depending on the superfamily. For 12% of all superfamilies the ROC5 value for superfamily models is greater than 0.2 above the domain models and for 10% of superfamilies the domain models show a similar improvement in performance over the superfamily models. CONCLUSION: Using a sensitive profile-profile method we have investigated the performance of single structure-based models and multiple sequence models (domain models) in detecting remote superfamily members. We find that overall, multiple models perform better in recognition although single structure-based models display better alignment accuracy.

Amino Acid Sequence↗

IntNetDB v1.0: an integrated protein-protein interaction network database generated by a probabilistic model.

BACKGROUND: Although protein-protein interaction (PPI) networks have been explored by various experimental methods, the maps so built are still limited in coverage and accuracy. To further expand the PPI network and to extract more accurate information from existing maps, studies have been carried out to integrate various types of functional relationship data. A frequently updated database of computationally analyzed potential PPIs to provide biological researchers with rapid and easy access to analyze original data as a biological network is still lacking. RESULTS: By applying a probabilistic model, we integrated 27 heterogeneous genomic, proteomic and functional annotation datasets to predict PPI networks in human. In addition to previously studied data types, we show that phenotypic distances and genetic interactions can also be integrated to predict PPIs. We further built an easy-to-use, updatable integrated PPI database, the Integrated Network Database (IntNetDB) online, to provide automatic prediction and visualization of PPI network among genes of interest. The networks can be visualized in SVG (Scalable Vector Graphics) format for zooming in or out. IntNetDB also provides a tool to extract topologically highly connected network neighborhoods from a specific network for further exploration and research. Using the MCODE (Molecular Complex Detections) algorithm, 190 such neighborhoods were detected among all the predicted interactions. The predicted PPIs can also be mapped to worm, fly and mouse interologs. CONCLUSION: IntNetDB includes 180,010 predicted protein-protein interactions among 9,901 human proteins and represents a useful resource for the research community. Our study has increased prediction coverage by five-fold. IntNetDB also provides easy-to-use network visualization and analysis tools that allow biological researchers unfamiliar with computational biology to access and analyze data over the internet. The web interface of IntNetDB is freely accessible at http://hanlab.genetics.ac.cn/IntNetDB.htm. Visualization requires Mozilla version 1.8 (or higher) or Internet Explorer with installation of SVGviewer.

Algorithms↗

Biclustering of gene expression data by Non-smooth Non-negative Matrix Factorization.

BACKGROUND: The extended use of microarray technologies has enabled the generation and accumulation of gene expression datasets that contain expression levels of thousands of genes across tens or hundreds of different experimental conditions. One of the major challenges in the analysis of such datasets is to discover local structures composed by sets of genes that show coherent expression patterns across subsets of experimental conditions. These patterns may provide clues about the main biological processes associated to different physiological states. RESULTS: In this work we present a methodology able to cluster genes and conditions highly related in sub-portions of the data. Our approach is based on a new data mining technique, Non-smooth Non-Negative Matrix Factorization (nsNMF), able to identify localized patterns in large datasets. We assessed the potential of this methodology analyzing several synthetic datasets as well as two large and heterogeneous sets of gene expression profiles. In all cases the method was able to identify localized features related to sets of genes that show consistent expression patterns across subsets of experimental conditions. The uncovered structures showed a clear biological meaning in terms of relationships among functional annotations of genes and the phenotypes or physiological states of the associated conditions. CONCLUSION: The proposed approach can be a useful tool to analyze large and heterogeneous gene expression datasets. The method is able to identify complex relationships among genes and conditions that are difficult to identify by standard clustering algorithms.

Algorithms↗

Prognostic meta-signature of breast cancer developed by two-stage mixture modeling of microarray data.

BACKGROUND: An increasing number of studies have profiled tumor specimens using distinct microarray platforms and analysis techniques. With the accumulating amount of microarray data, one of the most intriguing yet challenging tasks is to develop robust statistical models to integrate the findings. RESULTS: By applying a two-stage Bayesian mixture modeling strategy, we were able to assimilate and analyze four independent microarray studies to derive an inter-study validated "meta-signature" associated with breast cancer prognosis. Combining multiple studies (n = 305 samples) on a common probability scale, we developed a 90-gene meta-signature, which strongly associated with survival in breast cancer patients. Given the set of independent studies using different microarray platforms which included spotted cDNAs, Affymetrix GeneChip, and inkjet oligonucleotides, the individually identified classifiers yielded gene sets predictive of survival in each study cohort. The study-specific gene signatures, however, had minimal overlap with each other, and performed poorly in pairwise cross-validation. The meta-signature, on the other hand, accommodated such heterogeneity and achieved comparable or better prognostic performance when compared with the individual signatures. Further by comparing to a global standardization method, the mixture model based data transformation demonstrated superior properties for data integration and provided solid basis for building classifiers at the second stage. Functional annotation revealed that genes involved in cell cycle and signal transduction activities were over-represented in the meta-signature. CONCLUSION: The mixture modeling approach unifies disparate gene expression data on a common probability scale allowing for robust, inter-study validated prognostic signatures to be obtained. With the emerging utility of microarrays for cancer prognosis, it will be important to establish paradigms to meta-analyze disparate gene expression data for prognostic signatures of potential clinical use.

Bayes Theorem↗