Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

A new method of representing DNA sequences which combines ease of visual analysis with machine readability.

A new method of representing DNA sequences has been devised which is termed stave projection. Compared with other formats for showing the base sequences of DNA, this method greatly enhances the ease of visual analysis of the sequences of bases and it is also in a machine readable form. Using this method it is possible to identify and annotate all of the functional features found in DNA sequences.

Base Sequence↗

CDD: a curated Entrez database of conserved domain alignments.

The Conserved Domain Database (CDD) is now indexed as a separate database within the Entrez system and linked to other Entrez databases such as MEDLINE(R). This allows users to search for domain types by name, for example, or to view the domain architecture of any protein in Entrez's sequence database. CDD can be accessed on the WorldWideWeb at http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=cdd. Users may also employ the CD-Search service to identify conserved domains in new sequences, at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi. CD-Search results, and pre-computed links from Entrez's protein database, are calculated using the RPS-BLAST algorithm and Position Specific Score Matrices (PSSMs) derived from CDD alignments. CD-Searches are also run by default for protein-protein queries submitted to BLAST(R) at http://www.ncbi.nlm.nih.gov/BLAST. CDD mirrors the publicly available domain alignment collections SMART and PFAM, and now also contains alignment models curated at NCBI. Structure information is used to identify the core substructure likely to be present in all family members, and to produce sequence alignments consistent with structure conservation. This alignment model allows NCBI curators to annotate 'columns' corresponding to functional sites conserved among family members.

Amino Acid Sequence↗

PDBSiteScan: a program for searching for active, binding and posttranslational modification sites in the 3D structures of proteins.

PDBSiteScan is a web-accessible program designed for searching three-dimensional (3D) protein fragments similar in structure to known active, binding and posttranslational modification sites. A collection of known sites we designated as PDBSite was set up by automated processing of the PDB database using the data on site localization in the SITE field. Additionally, protein-protein interaction sites were generated by analysis of atom coordinates in heterocomplexes. The total number of collected sites was more than 8100; they were assigned to more than 80 functional groups. PDBSiteScan provides automated search of the 3D protein fragments whose maximum distance mismatch (MDM) between N, Calpha and C atoms in a fragment and a functional site is not larger than the MDM threshold defined by the user. PDBSiteScan requires perfect matching of amino acids. PDBSiteScan enables recognition of functional sites in tertiary structures of proteins and allows proteins with functional information to be annotated. The program PDBSiteScan is available at http://wwwmgs.bionet.nsc.ru/mgs/systems/fastprot/pdbsitescan.html.

Binding Sites↗

Computational identification of transcriptional regulatory elements in DNA sequence.

Identification and annotation of all the functional elements in the genome, including genes and the regulatory sequences, is a fundamental challenge in genomics and computational biology. Since regulatory elements are frequently short and variable, their identification and discovery using computational algorithms is difficult. However, significant advances have been made in the computational methods for modeling and detection of DNA regulatory elements. The availability of complete genome sequence from multiple organisms, as well as mRNA profiling and high-throughput experimental methods for mapping protein-binding sites in DNA, have contributed to the development of methods that utilize these auxiliary data to inform the detection of transcriptional regulatory elements. Progress is also being made in the identification of cis-regulatory modules and higher order structures of the regulatory sequences, which is essential to the understanding of transcription regulation in the metazoan genomes. This article reviews the computational approaches for modeling and identification of genomic regulatory elements, with an emphasis on the recent developments, and current challenges.

Animals↗

KBERG: KnowledgeBase for Estrogen Responsive Genes.

Estrogen has a profound impact on human physiology affecting transcription of numerous genes. To decipher functional characteristics of estrogen responsive genes, we developed KnowledgeBase for Estrogen Responsive Genes (KBERG). Genes in KBERG were derived from Estrogen Responsive Gene Database (ERGDB) and were analyzed from multiple aspects. We explored the possible transcription regulation mechanism by capturing highly conserved promoter motifs across orthologous genes, using promoter regions that cover the range of [-1200, +500] relative to the transcription start sites. The motif detection is based on ab initio discovery of common cis-elements from the orthologous gene cluster from human, mouse and rat, thus reflecting a degree of promoter sequence preservation during evolution. The identified motifs are linked to transcription factor binding sites based on the TRANSFAC database. In addition, KBERG uses two established ontology systems, GO and eVOC, to associate genes with their function. Users may assess gene functionality through the description terms in GO. Alternatively, they can gain gene co-expression information through evidence from human EST libraries via eVOC. KBERG is a user-friendly system that provides links to other relevant resources such as ERGDB, UniGene, Entrez Gene, HomoloGene, GO, eVOC and GenBank, and thus offers a platform for functional exploration and potential annotation of genes responsive to estrogen. KBERG database can be accessed at http://research.i2r.a-star.edu.sg/kberg.

Animals↗

WormBase: new content and better access.

WormBase (http://wormbase.org), a model organism database for Caenorhabditis elegans and other related nematodes, continues to evolve and expand. Over the past year WormBase has added new data on C.elegans, including data on classical genetics, cell biology and functional genomics; expanded the annotation of closely related nematodes with a new genome browser for Caenorhabditis remanei; and deployed new hardware for stronger performance. Several existing datasets including phenotype descriptions and RNAi experiments have seen a large increase in new content. New datasets such as the C.remanei draft assembly and annotations, the Vancouver Fosmid library and TEC-RED 5' end sites are now available as well. Access to and searching WormBase has become more dependable and flexible via multiple mirror sites and indexing through Google.

Animals↗

Comprehensive genome sequence analysis of a breast cancer amplicon.

Gene amplification occurs in most solid tumors and is associated with poor prognosis. Amplification of 20q13.2 is common to several tumor types including breast cancer. The 1 Mb of sequence spanning the 20q13.2 breast cancer amplicon is one of the most exhaustively studied segments of the human genome. These studies have included amplicon mapping by comparative genomic hybridization (CGH), fluorescent in-situ hybridization (FISH), array-CGH, quantitative microsatellite analysis (QUMA), and functional genomic studies. Together these studies revealed a complex amplicon structure suggesting the presence of at least two driver genes in some tumors. One of these, ZNF217, is capable of immortalizing human mammary epithelial cells (HMEC) when overexpressed. In addition, we now report the sequencing of this region in human and mouse, and on quantitative expression studies in tumors. Amplicon localization now is straightforward and the availability of human and mouse genomic sequence facilitates their functional analysis. However, comprehensive annotation of megabase-scale regions requires integration of vast amounts of information. We present a system for integrative analysis and demonstrate its utility on 1.2 Mb of sequence spanning the 20q13.2 breast cancer amplicon and 865 kb of syntenic murine sequence. We integrate tumor genome copy number measurements with exhaustive genome landscape mapping, showing that amplicon boundaries are associated with maxima in repetitive element density and a region of evolutionary instability. This integration of comprehensive sequence annotation, quantitative expression analysis, and tumor amplicon boundaries provide evidence for an additional driver gene prefoldin 4 (PFDN4), coregulated genes, conserved noncoding regions, and associate repetitive elements with regions of genomic instability at this locus.

Animals↗

Modeling gene expression networks using fuzzy logic.

Gene regulatory networks model regulation in living organisms. Fuzzy logic can effectively model gene regulation and interaction to accurately reflect the underlying biology. A new multiscale fuzzy clustering method allows genes to interact between regulatory pathways and across different conditions at different levels of detail. Fuzzy cluster centers can be used to quickly discover causal relationships between groups of coregulated genes. Fuzzy measures weight expert knowledge and help quantify uncertainty about the functions of genes using annotations and the gene ontology database to confirm some of the interactions. The method is illustrated using gene expression data from an experiment on carbohydrate metabolism in the model plant Arabidopsis thaliana. Key gene regulatory relationships were evaluated using information from the gene ontology database. A new regulatory relationship concerning trehalose regulation of carbohydrate metabolism was also discovered in the extracted network.

Animals↗

Transcriptional divergence of the duplicated oxidative stress-responsive genes in the Arabidopsis genome.

Previous studies have indicated that Arabidopsis thaliana experienced a genome-wide duplication event shortly before its divergence from Brassica followed by extensive chromosomal rearrangements and deletions. While a large number of the duplicated genes have significantly diverged or lost their sister genes, we found 4222 pairs that are still highly conserved, and as a result had similar functional assignments during the annotation of the genome sequence. Using whole-genome DNA microarrays, we identified 906 duplicated gene pairs in which at least one member exhibited a significant response to oxidative stress. Among these, only 117 pairs were up- or down-regulated in both pairs and many of these exhibited dissimilar patterns of expression. Examination of the expression patterns of PAL1 and PAL2, ACD1 and ACD2, genes coding for two Hsp20s, various P450s, and electron transfer flavoproteins suggests Arabidopsis evolved a number of distinct oxidative stress response mechanisms using similar gene sets following the duplication of its genome.

Arabidopsis↗

C-terminal WxL domain mediates cell wall binding in Enterococcus faecalis and other gram-positive bacteria.

Analysis of the genome sequence of Enterococcus faecalis clinical isolate V583 revealed novel genes encoding surface proteins. Twenty-seven of these proteins, annotated as having unknown functions, possess a putative N-terminal signal peptide and a conserved C-terminal region characterized by a novel conserved domain designated WxL. Proteins having similar characteristics were also detected in other low-G+C-content gram-positive bacteria. We hypothesized that the WxL region might be a determinant of bacterial cell location. This hypothesis was tested by generating protein fusions between the C-terminal regions of two WxL proteins in E. faecalis and a nuclease reporter protein. We demonstrated that the C-terminal regions of both proteins conferred a cell surface localization to the reporter fusions in E. faecalis. This localization was eliminated by introducing specific deletions into the domains. Interestingly, exogenously added protein fusions displayed binding to whole cells of various gram-positive bacteria. We also showed that the peptidoglycan was a binding ligand for WxL domain attachment to the cell surface and that neither proteins nor carbohydrates were necessary for binding. Based on our findings, we propose that the WxL region is a novel cell wall binding domain in E. faecalis and other gram-positive bacteria.

Amino Acid Sequence↗

Linking molecular imaging terminology to the gene ontology (GO).

The rapidly developing domain of molecular imaging represents the merging of current advances in the fields of molecular biology and imaging research. Despite this merger, an information gap continues to exist between the scientists who discover new gene products and the imaging scientists who can exploit this information. The Gene Ontology (GO) Consortium seeks to provide a set of structured terminologies for the conceptual annotation of gene product function, process and location in databases. However, no such structured set of concept-oriented terminology exists for the molecular imaging domain. Since the purpose of GO is to capture the information about the role of gene products, we propose that the mapping of GO's established ontological concepts to a molecular imaging terminology will provide the necessary bridge to fill the information gap between the two fields. We have extracted terms and definitions from an already published molecular imaging glossary as well as molecular imaging research articles, and developed molecular imaging concepts. We then mapped our molecular imaging concepts to the existing gene ontology concepts as a method to comprehensively represent molecular imaging.

Computational Biology↗

NovelFam3000--uncharacterized human protein domains conserved across model organisms.

BACKGROUND: Despite significant efforts from the research community, an extensive portion of the proteins encoded by human genes lack an assigned cellular function. Most metazoan proteins are composed of structural and/or functional domains, of which many appear in multiple proteins. Once a domain is characterized in one protein, the presence of a similar sequence in an uncharacterized protein serves as a basis for inference of function. Thus knowledge of a domain's function, or the protein within which it arises, can facilitate the analysis of an entire set of proteins. DESCRIPTION: From the Pfam domain database, we extracted uncharacterized protein domains represented in proteins from humans, worms, and flies. A data centre was created to facilitate the analysis of the uncharacterized domain-containing proteins. The centre both provides researchers with links to dispersed internet resources containing gene-specific experimental data and enables them to post relevant experimental results or comments. For each human gene in the system, a characterization score is posted, allowing users to track the progress of characterization over time or to identify for study uncharacterized domains in well-characterized genes. As a test of the system, a subset of 39 domains was selected for analysis and the experimental results posted to the NovelFam3000 system. For 25 human protein members of these 39 domain families, detailed sub-cellular localizations were determined. Specific observations are presented based on the analysis of the integrated information provided through the online NovelFam3000 system. CONCLUSION: Consistent experimental results between multiple members of a domain family allow for inferences of the domain's functional role. We unite bioinformatics resources and experimental data in order to accelerate the functional characterization of scarcely annotated domain families.

Animals↗

Possible linking and treatment between Parkinson's disease and inflammatory bowel disease: a study of Mendelian randomization based on gut-brain axis.

BACKGROUND: Mounting evidence suggests that Parkinson's disease (PD) and inflammatory bowel disease (IBD) are closely associated and becoming global health burdens. However, the causal relationships and common pathogeneses between them are uncertain. Furthermore, they are uncurable. Thus, we aimed to identify the causal relationships and novel therapeutic targets shared between them based on their common pathophysiological mechanisms in gut-brain-axis (GBA). METHODS: A meta-analysis on bidirectional Mendelian randomization (MR) utilizing various datasets was performed to estimate their causal relationship. Then, pleiotropic analysis under the composite null hypothesis (PLACO) with functional mapping combined with annotation of genetic associations (FUMA) analysis were conducted to identify pleiotropic genes. Next, blood, brain and intestine expression quantitative trait locus (eQTL) were taken to perform drug-target MR finding common causal genes in two diseases. Colocalization analysis ensured the eQTLs of corresponding gene colocalized with disease. Enrichment analysis and protein‒protein interaction (PPI) network were done to explore common pathogenesis pathways. Genes passed all analysis were regarded as drug targets. RESULTS: Our MR meta-analysis revealed the bidirectional causal relationship between diseases, with combined ORs for PD on IBD, CD, UC (1.050 [95% CI 1.014-1.086], 1.044 [95% CI 0.995-1.095], 1.063 [95% CI 1.016-1.120]); for IBD, CD, UC on PD (1.003 [95% CI 0.973-1.034], 1.035 [95% CI 1.004-1.067], 1.008 [95% CI 0.977-1.040]). Overall, 277, 216 and 201 genes were identified as pleiotropic genes between PD and IBD, CD, UC. Total of 733 genes were classified as tier 3 (found in only one tissue) druggable targets, 57 as tier 2 (found in two tissues, 51 protein-coding genes) and 9 as tier 3 (found in three tissues). Among 60 protein-coding druggable targets over tier 2, 18 overlapped with pleiotropic genes and enriched in mitochondria, antigen presentation, processing and immune cell regulation pathways. Three druggable genes (LRRK2, RAB29 and HLA-DQA2) passed colocalization analysis. LRRK2 and RAB29 were reported to be pleiotropic genes, and RAB29 and HLA-DQA2 were reported for the first time as potential drug targets. CONCLUSIONS: This study established a reliable causal relationship, possible shared drug targets and common pathogenesis pathways of two diseases, which had important implications for intervention and treatment of two diseases simultaneously.

Humans↗

Transcription profiling of renal cell carcinoma.

AIMS: Our aim was to prepare a comprehensive catalogue of the changes in gene expression accompanying the development and progression of renal cell carcinoma, and to correlate these with histo-pathological, cytogenetic and clinical findings. METHODS: mRNA samples from paired neoplastic and non-cancerous human kidney tissue were labeled and hybridized in duplicate against high-density cDNA arrays. Two array technologies were used: 31,500-element transcriptome-wide nylon arrays for hybridization with 37 radioactively labelled sample pairs, and 4200-element kidney- and cancer-specific glass microarrays for hybridization with 19 fluorescently labelled sample pairs. RESULTS: We identified more than 1700 cDNA clones that show differential transcription levels in kidney tumor tissue compared to normal kidney tissue. The functional classification of 389 annotated genes provided views of the changes in the activities of specific biological processes in renal cancer. Among the biological processes with a large proportion of up-regulated genes we found cell adhesion, signal transduction, and nucleotide metabolism. Down-regulated processes included small molecule transport, ion homeostasis, and oxygen and radical metabolism. Furthermore, we explored the feasibility of molecular diagnosis for renal cell tumors using cDNA microarrays on glass slides, investigating the association of transcription levels with tumor type, progression, and a putative prognostic variable. The experimental data is available from the GEO gene expression database (http://www.ncbi.nlm.nih.gov/geo; accession no. GSE3), and a comprehensive presentation of the results is available in the web supplement (http://www.dkfz-heidelberg.de/abt0840/whuber/rcc). CONCLUSION: Transcription profiling using high-density cDNA arrays is a powerful method with the potential to improve cancer diagnosis and prognosis. The identification and classification of differentially transcribed genes, as described in our study, is the beginning of a more complete understanding of kidney cancer.

Carcinoma, Renal Cell↗

goCluster integrates statistical analysis and functional interpretation of microarray expression data.

MOTIVATION: Several tools that facilitate the interpretation of transcriptional profiles using gene annotation data are available but most of them combine a particular statistical analysis strategy with functional information. goCluster extends this concept by providing a modular framework that facilitates integration of statistical and functional microarray data analysis with data interpretation. RESULTS: goCluster enables scientists to employ annotation information, clustering algorithms and visualization tools in their array data analysis and interpretation strategy. The package provides four clustering algorithms and GeneOntology terms as prototype annotation data. The functional analysis is based on the hypergeometric distribution whereby the Bonferroni correction or the false discovery rate can be used to correct for multiple testing. The approach implemented in goCluster was successfully applied to interpret the results of complex mammalian and yeast expression data obtained with high density oligonucleotide microarrays (GeneChips). AVAILABILITY: goCluster is available via the BioConductor portal at www.bioconductor.org. The software package, detailed documentation, user- and developer guides as well as other background information are also accessible via a web portal at http://www.bioz.unibas.ch/gocluster CONTACT: michael.primig@unibas.ch

Algorithms↗

Protein interaction mapping on a functional shotgun sequence of Rickettsia sibirica.

Protein interaction maps can reveal novel pathways and functional complexes, allowing 'guilt by association' annotation of uncharacterized proteins. To address the need for large-scale protein interaction analyses, a bacterial two-hybrid system was coupled with a whole genome shotgun sequencing approach for microbial genome analysis. We report the first large-scale proteomics study using this system, integrating de novo genome sequencing with functional interaction mapping and annotation in a high-throughput format. We apply the approach by shotgun sequencing and annotating the genome of Rickettsia sibirica strain 246, an obligate intracellular human pathogen among the Spotted Fever Group rickettsiae. The bacteria invade endothelial cells and cause lysis after large amounts of progeny have accumulated. Little is known about specific Rickettsial virulence factors and their mode of pathogenicity. Analysis of the combined genomic sequence and protein-protein interaction data for a set of virulence related Type IV secretion system (T4SS) proteins revealed over 250 interactions and will provide insight into the mechanism of Rickettsial pathogenicity.

Bacterial Proteins↗

SIGNATURE: a single-particle selection system for molecular electron microscopy.

SIGNATURE is a particle selection system for molecular electron microscopy. It applies a hierarchical screening procedure to identify molecular particles in EM micrographs. The user interface of the program provides versatile functions to facilitate image data visualization, particle annotation and particle quality inspection. The system design emphasizes both functionality and usability. This software has been released to the EM community and has been successfully applied to macromolecular structural analyses.

Algorithms↗

A description scheme of biological processes based on elementary bricks of action.

With the fast growth of high-throughput strategies in Biology, there is a strong need to accelerate knowledge acquisition and organization of molecular functions. Unfortunately, although we know that there is a correlation between protein molecules and their functions, we are unable to clearly identify this link. Here, we revisit the current views of protein functions as well as their annotation, and we show that they are incompatible with unambiguous interpretations and the use of this knowledge. We describe herein a description scheme for biological processes based on elementary bricks of action that may be associated with biological molecules. To retrieve the descriptive quality found in annotations of other kinds of biological data, it was decided to develop a scheme involving four levels of abstraction: Basic Elements of Action, Biological Activities, Biological Functionalities and Biological Roles. This multi-level organization is a generic method; it allows for a description of biological processes by using a limited number of elementary bricks of action. Moreover, by using this description of biological processes, it should now be possible to clearly identify unambiguous relationships between the organization of biological processes and the structural or functional organizations of biological molecules.

Algorithms↗