Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Functional annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

GOstat: find statistically overrepresented Gene Ontologies within a group of genes.

SUMMARY: Modern experimental techniques, as for example DNA microarrays, as a result usually produce a long list of genes, which are potentially interesting in the analyzed process. In order to gain biological understanding from this type of data, it is necessary to analyze the functional annotations of all genes in this list. The Gene-Ontology (GO) database provides a useful tool to annotate and analyze the functions of a large number of genes. Here, we introduce a tool that utilizes this information to obtain an understanding of which annotations are typical for the analyzed list of genes. This program automatically obtains the GO annotations from a database and generates statistics of which annotations are overrepresented in the analyzed list of genes. This results in a list of GO terms sorted by their specificity. AVAILABILITY: Our program GOstat is accessible via the Internet at http://gostat.wehi.edu.au

Abstracting and Indexing↗

Common genetic variants associated with urinary phthalate levels in children: A genome-wide study.

INTRODUCTION: Phthalates, or dieters of phthalic acid, are a ubiquitous type of plasticizer used in a variety of common consumer and industrial products. They act as endocrine disruptors and are associated with increased risk for several diseases. Once in the body, phthalates are metabolized through partially known mechanisms, involving phase I and phase II enzymes. OBJECTIVE: In this study we aimed to identify common single nucleotide polymorphisms (SNPs) and copy number variants (CNVs) associated with the metabolism of phthalate compounds in children through genome-wide association studies (GWAS). METHODS: The study used data from 1,044 children with European ancestry from the Human Early Life Exposome (HELIX) cohort. Ten phthalate metabolites were assessed in a two-void pooled urine collected at the mean age of 8&#xa0;years. Six ratios between secondary and primary phthalate metabolites were calculated. Genome-wide genotyping was done with the Infinium Global Screening Array (GSA) and imputation with the Haplotype Reference Consortium (HRC) panel. PennCNV was used to estimate copy number variants (CNVs) and CNVRanger to identify consensus regions. GWAS of SNPs and CNVs were conducted using PLINK and SNPassoc, respectively. Subsequently, functional annotation of suggestive SNPs (p-value&#xa0;<&#xa0;1E-05) was done with the FUMA web-tool. RESULTS: We identified four genome-wide significant (p-value&#xa0;<&#xa0;5E-08) loci at chromosome (chr) 3 (FECHP1 for oxo-MiNP_oh-MiNP ratio), chr6 (SLC17A1 for MECPP_MEHHP ratio), chr9 (RAPGEF1 for MBzP), and chr10 (CYP2C9 for MECPP_MEHHP ratio). Moreover, 115 additional loci were found at suggestive significance (p-value&#xa0;<&#xa0;1E-05). Two CNVs located at chr11 (MRGPRX1 for oh-MiNP and SLC35F2 for MEP) were also identified. Functional annotation pointed to genes involved in phase I and phase II detoxification, molecular transfer across membranes, and renal excretion. CONCLUSION: Through genome-wide screenings we identified known and novel loci implicated in phthalate metabolism in children. Genes annotated to these loci participate in detoxification, transmembrane transfer, and renal excretion.

Humans↗

EGassembler: online bioinformatics service for large-scale processing, clustering and assembling ESTs and genomic DNA fragments.

Expressed sequence tag (EST) sequencing has proven to be an economically feasible alternative for gene discovery in species lacking a draft genome sequence. Ongoing large-scale EST sequencing projects feel the need for bioinformatics tools to facilitate uniform EST handling. This brings about a renewed importance for a universal tool for processing and functional annotation of large sets of ESTs. EGassembler (http://egassembler.hgc.jp/) is a web server, which provides an automated as well as a user-customized analysis tool for cleaning, repeat masking, vector trimming, organelle masking, clustering and assembling of ESTs and genomic fragments. The web server is publicly available and provides the community a unique all-in-one online application web service for large-scale ESTs and genomic DNA clustering and assembling. Running on a Sun Fire 15K supercomputer, a significantly large volume of data can be processed in a short period of time. The results can be used to functionally annotate genes, to facilitate splice alignment analysis, to link the transcripts to genetic and physical maps, design microarray chips, to perform transcriptome analysis and to map to KEGG metabolic pathways. The service provides an excellent bioinformatics tool to research groups in wet-lab as well as an all-in-one-tool for sequence handling to bioinformatics researchers.

Computational Biology↗

Survey of current protein family databases and their application in comparative, structural and functional genomics.

The last two decades have witnessed significant expansions in the databases storing information on the sequences and structures of proteins. This has led to the creation of many excellent protein family resources, which classify proteins according to their evolutionary relationship. These have allowed extensive insights into evolution and particularly how protein function mutates and evolves over time. Such analyses have greatly assisted the inheritance of functional annotations between experimentally characterised and uncharacterised genes. Moreover, the development of bioinformatics tools acts as a companion to the new technologies emerging in biology, such as transcriptomics and proteomics. The latter enable researchers to analyse gene expression profiles and interactions on a genome-wide scale, generating vast datasets of proteins, many of which include experimentally uncharacterised proteins. Protein family/function databases can be used to help interpret this data and allow us to benefit more fully from these technologies. This review aims to summarise the most popular sequence- and structure-based protein family databases. We also cover their application to comparative genomics and the functional annotation of the genomes.

Biological Evolution↗

Shared genetic architecture between ADHD and intelligence varies across ADHD subtypes.

BACKGROUND: Attention-deficit/hyperactivity disorder (ADHD) is a heterogeneous neurodevelopmental condition frequently accompanied by cognitive difficulties. Although previous genetic studies have demonstrated substantial overlap between ADHD and intelligence, most have treated ADHD as a single phenotype. However, whether this shared genetic architecture differs across ADHD subtypes remains unclear. METHODS: We conducted a genome-wide cross-trait analysis integrating large-scale genome-wide association study (GWAS) datasets of overall ADHD, its subtypes-childhood ADHD, persistent ADHD, and late-diagnosed ADHD-and intelligence (total N&#x2009;>&#x2009;300,000). Genome-wide genetic correlations, polygenic overlap, local genetic correlations, and variant-level associations between ADHD phenotypes and intelligence were evaluated to characterize their shared genetic architecture. Shared variants were identified through cross-trait enrichment analyses and subsequently mapped to genes for functional annotation and gene-set enrichment. Bidirectional associations were evaluated using two-sample Mendelian randomization with sensitivity analyses. Additional GWAS datasets were used to validate the robustness of shared loci by assessing the consistency of effect directions. RESULTS: All ADHD phenotypes showed significant negative genetic correlations with intelligence (rg ranging from -0.3442 to -0.4205). Despite these modest genome-wide correlations, cross-trait analyses revealed substantial genetic overlap, including polygenic overlap, local genetic correlations, and variant-level associations. We identified 184 loci jointly associated with ADHD traits and intelligence, including 64 novel loci, whereas no shared loci were detected for persistent ADHD under the current analysis. Functional annotation revealed biologically distinct enrichment patterns across subtypes: childhood ADHD loci were linked to early neurodevelopmental processes, while late-diagnosed ADHD loci were enriched in synapse-related and neuronal signaling pathways. Mendelian randomization analyses suggested bidirectional associations, with stronger evidence supporting a directional association from intelligence to ADHD risk. Furthermore, these shared loci showed largely consistent effect directions across additional GWAS datasets, providing support for the robustness of the findings. CONCLUSIONS: The shared genetic architecture between ADHD and intelligence varies across ADHD subtypes, highlighting distinct biological pathways underlying cognitive heterogeneity in ADHD. These findings suggest that the relationship between ADHD liability and general cognitive ability is not uniform across ADHD subtypes and may inform future research on risk stratification and early identification in child and adolescent psychiatry.

Humans↗

Expression profiling of human idiopathic dilated cardiomyopathy.

OBJECTIVE: To investigate the global changes accompanying human dilated cardiomyopathy (DCM) we performed a large-scale expression screen using myocardial biopsies from a group of DCM patients with moderate heart failure. By hierarchical clustering and functional annotation of the deregulated genes we examined extensive changes in the cellular and molecular processes associated to DCM. METHODS: The expression profiles were obtained using a whole genome covering library (UniGene RZPD1) comprising 30336 cDNA clones and amplified RNA from myocardiac biopsies from 10 DCM patients in comparison to tissue samples from four non-failing, healthy donors. RESULTS: By setting stringent selection criteria 364 differentially expressed, sequence-verified non-redundant transcripts were identified with a false discovery rate of <0.001. Numerous genes and ESTs were identified representing previously recognised, as well as novel DCM-associated transcripts. Many of them were found to be upregulated and involved in cardiomyocyte energetics, muscle contraction or signalling. Two hundred and twenty-two deregulated transcripts were functionally annotated and hierarchically clustered providing an insight into the pathophysiology of DCM. Data was validated using the MLP-deficient mouse, in which several differentially expressed transcripts identified in the human DCM biopsies could be confirmed. CONCLUSIONS: We report the first genome-wide expression profile analysis using cardiac biopsies from DCM patients at various stages of the disease. Although there is a diversity of links between the cytoskeleton and the initiation of DCM, we speculate that genes implicated in intracellular signalling and in muscle contraction are associated with early stages of the disease. Altogether this study represents the most comprehensive and inclusive molecular portrait of human cardiomyopathy to date.

Adult↗

Genomics and proteomics of bone cancer.

Although the control of bone metastasis has been the focus of intensive investigation, relatively little is known about the molecular mechanisms that regulate or predict the process, even though widespread skeletal dissemination is an important step in the progression of many tumors. As a result, understanding the complex interactions contributing to the metastatic behavior of tumor cells is essential for the development of effective therapies. Using a state-of-the-art combination of gene expression profiling and functional annotation of human tumor cells, and surface-enhanced laser desorption/ionization time-of-flight mass spectrometry of patient serum, we have shown that changes in tumor biochemistry correlate with disease progression and help to define the aggressive tumor phenotype. Based on these approaches, it is apparent that the metastatic phenotype of tumor cells is extremely complex. The identification of the phenotype of tumor cells has benefited greatly from the application of gene expression profiling (microarray analysis). This technology has been used by many investigators to identify changes in gene expression and cytokine and growth factor elaboration (such as interleukin 8). The tumor phenotype(s) presumably also include changes in the cell surface carbohydrate profile (via altered glycosyltransferase expression) and heparan sulfate expression (via increased heparanase activity), to name but a few. These specific alterations in gene expression, identified by functional annotation of accumulated microarray data, have been validated using a variety of approaches. Collectively, the data described here suggest that each of these activities is associated with distinct aspects of the aggressive tumor cell phenotype. Collectively, the data suggest that multiple factors constitute the complex phenotype of metastatic tumor cells. In particular, the differences observed in gene expression profiles and serum protein biomarkers play a critical role in defining the mechanisms responsible for bone-specific colonization and growth of tumors in bone. Future studies will identify the mechanisms that participate in the formation of secondary tumor growths of cancers in bone.

Biomarkers, Tumor↗

COCO-CL: hierarchical clustering of homology relations based on evolutionary correlations.

MOTIVATION: Determining orthology relations among genes across multiple genomes is an important problem in the post-genomic era. Identifying orthologous genes can not only help predict functional annotations for newly sequenced or poorly characterized genomes, but can also help predict new protein-protein interactions. Unfortunately, determining orthology relation through computational methods is not straightforward due to the presence of paralogs. Traditional approaches have relied on pairwise sequence comparisons to construct graphs, which were then partitioned into putative clusters of orthologous groups. These methods do not attempt to preserve the non-transitivity and hierarchic nature of the orthology relation. RESULTS: We propose a new method, COCO-CL, for hierarchical clustering of homology relations and identification of orthologous groups of genes. Unlike previous approaches, which are based on pairwise sequence comparisons, our method explores the correlation of evolutionary histories of individual genes in a more global context. COCO-CL can be used as a semi-independent method to delineate the orthology/paralogy relation for a refined set of homologous proteins obtained using a less-conservative clustering approach, or as a refiner that removes putative out-paralogs from clusters computed using a more inclusive approach. We analyze our clustering results manually, with support from literature and functional annotations. Since our orthology determination procedure does not employ a species tree to infer duplication events, it can be used in situations when the species tree is unknown or uncertain. CONTACT: jothi@mail.nih.gov, przytyck@mail.nih.gov SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online.

Algorithms↗

SNPSplicer: systematic analysis of SNP-dependent splicing in genotyped cDNAs.

Functional annotation of SNPs (as generated by HapMap (http://www.hapmap.org) for instance) is a major challenge. SNPs that lead to single amino acid substitutions, stop codons, or frameshift mutations can be readily interpreted, but these represent only a fraction of known SNPs. Many SNPs are located in sequences of splicing relevance-the canonical splice site consensus sequences, exonic and intronic splice enhancers or silencers (exonic splice enhancer [ESE], intronic splice enhancer [ISE], exonic splicing silencer [ESS], and intronic splicing silencer [ISS]), and others. We propose using sets of matching DNA and complementary DNA (cDNA) as a screening method to investigate the potential splice effects of SNPs in RT-PCR experiments with tissue material from genotyped sources. We have developed a software solution (SNPSplicer; http://www.ikmb.uni-kiel.de/snpsplicer) that aids in the rapid interpretation of such screening experiments. The utility of the approach is illustrated for SNPs affecting the donor splice sites (rs2076530:A>G, rs3816989:G>A) leading to the use of a cryptic splice site and exon skipping, respectively, and an exonic splice enhancer SNP (rs2274987:C/T), leading to inclusion of a new exon. We anticipate that this methodology may help in the functional annotation of SNPs in a more high-throughput fashion.

Alternative Splicing↗

Structural insights into adeno-associated virus serotype 5.

The adeno-associated viruses (AAVs) display differential cell binding, transduction, and antigenic characteristics specified by their capsid viral protein (VP) composition. Toward structure-function annotation, the crystal structure of AAV5, one of the most sequence diverse AAV serotypes, was determined to 3.45-&#xc5; resolution. The AAV5 VP and capsid conserve topological features previously described for other AAVs but uniquely differ in the surface-exposed HI loop between &#x3b2;H and &#x3b2;I of the core &#x3b2;-barrel motif and have pronounced conformational differences in two of the AAV surface variable regions (VRs), VR-IV and VR-VII. The HI loop is structurally conserved in other AAVs despite amino acid differences but is smaller in AAV5 due to an amino acid deletion. This HI loop is adjacent to VR-VII, which is largest in AAV5. The VR-IV, which forms the larger outermost finger-like loop contributing to the protrusions surrounding the icosahedral 3-fold axes of the AAVs, is shorter in AAV5, creating a smoother capsid surface topology. The HI loop plays a role in AAV capsid assembly and genome packaging, and VR-IV and VR-VII are associated with transduction and antigenic differences, respectively, between the AAVs. A comparison of interior capsid surface charge and volume of AAV5 to AAV2 and AAV4 showed a higher propensity of acidic residues but similar volumes, consistent with comparable DNA packaging capacities. This structure provided a three-dimensional (3D) template for functional annotation of the AAV5 capsid with respect to regions that confer assembly efficiency, dictate cellular transduction phenotypes, and control antigenicity.

Capsid Proteins↗

The COG database: an updated version includes eukaryotes.

BACKGROUND: The availability of multiple, essentially complete genome sequences of prokaryotes and eukaryotes spurred both the demand and the opportunity for the construction of an evolutionary classification of genes from these genomes. Such a classification system based on orthologous relationships between genes appears to be a natural framework for comparative genomics and should facilitate both functional annotation of genomes and large-scale evolutionary studies. RESULTS: We describe here a major update of the previously developed system for delineation of Clusters of Orthologous Groups of proteins (COGs) from the sequenced genomes of prokaryotes and unicellular eukaryotes and the construction of clusters of predicted orthologs for 7 eukaryotic genomes, which we named KOGs after eukaryotic orthologous groups. The COG collection currently consists of 138,458 proteins, which form 4873 COGs and comprise 75% of the 185,505 (predicted) proteins encoded in 66 genomes of unicellular organisms. The eukaryotic orthologous groups (KOGs) include proteins from 7 eukaryotic genomes: three animals (the nematode Caenorhabditis elegans, the fruit fly Drosophila melanogaster and Homo sapiens), one plant, Arabidopsis thaliana, two fungi (Saccharomyces cerevisiae and Schizosaccharomyces pombe), and the intracellular microsporidian parasite Encephalitozoon cuniculi. The current KOG set consists of 4852 clusters of orthologs, which include 59,838 proteins, or approximately 54% of the analyzed eukaryotic 110,655 gene products. Compared to the coverage of the prokaryotic genomes with COGs, a considerably smaller fraction of eukaryotic genes could be included into the KOGs; addition of new eukaryotic genomes is expected to result in substantial increase in the coverage of eukaryotic genomes with KOGs. Examination of the phyletic patterns of KOGs reveals a conserved core represented in all analyzed species and consisting of approximately 20% of the KOG set. This conserved portion of the KOG set is much greater than the ubiquitous portion of the COG set (approximately 1% of the COGs). In part, this difference is probably due to the small number of included eukaryotic genomes, but it could also reflect the relative compactness of eukaryotes as a clade and the greater evolutionary stability of eukaryotic genomes. CONCLUSION: The updated collection of orthologous protein sets for prokaryotes and eukaryotes is expected to be a useful platform for functional annotation of newly sequenced genomes, including those of complex eukaryotes, and genome-wide evolutionary studies.

Animals↗

The extracellular vesicle transcriptome provides tissue-specific functional genomic annotation relevant to disease susceptibility in obesity.

We characterized circulating extracellular vesicles (EVs) in obese and lean humans, identifying transcriptional cargo differentially expressed in obesity (277 unique genes; false discovery rate < 10%). Since circulating EVs may have broad origin, we compared this obesity EV transcriptome with expression from human visceral-adipose-tissue-derived EVs from freshly collected and cultured biopsies from the same obese individuals, observing high concordance. Using a comprehensive set of adipose-specific epigenomic and chromatin conformation assays, we found that the differentially expressed transcripts from the EVs were those regulated in adipose by body mass index-associated SNPs (p < 5 &#xd7; 10-8) from a large-scale genome-wide association study (GWAS). Using a phenome-wide association study of the regulatory SNPs for the EV-derived transcripts, we identified a substantial enrichment for inflammatory phenotypes, including type 2 diabetes. Collectively, these findings represent the convergence of the GWAS (genetics), epigenomics (transcript regulation), and EV (liquid biopsy) fields, enabling powerful future genomic studies of complex diseases.

Humans↗

Exploring trafficking GTPase function by mRNA expression profiling: use of the SymAtlas web-application and the Membrome datasets.

Despite complete sequencing of the human and mouse genomes, functional annotation of novel gene function still remains a major challenge in mammalian biology. Emerging strategies to help elucidate unknown gene function include the analysis of tissue-specific patterns of mRNA expression. A recent study investigated the steady-state mRNA expression profiling of the vast majority of protein-encoding human and mouse genes across a panel of 79 human and 61 mouse nonredundant tissues. The microarray data from this study constitutes the Genomics Institute of Novartis Foundation (GNF) Human and Mouse Gene Atlases and is publicly available for exploration through the SymAtlas web-application (http://symatlas.gnf.org/). We have recently reported the use of these data and hierarchical clustering algorithms to generate a global overview of the distribution of Rabs, SNAREs, and coat machinery components, as well as their respective adaptors, effectors, and regulators. This systems biology approach led us to propose Rab-centric protein activity hubs as a framework for an integrated coding system, the membrome network, which orchestrates the dynamics of specialized membrane architecture of differentiated cells. Here, we describe the use of the SymAtlas web-application and the Membrome datasets to help explore trafficking GTPase function. The human and mouse membrome datasets are available through the Membrome homepage (http://www.membrome.org/) and correspond to subsets of the SymAtlas content restricted to known membrane trafficking components. Considering the fragmentary nature of the current reductionist approaches in elucidating trafficking component functions, the membrome datasets provide a more focused systems biology perspective that not only complements our current understanding of transport in complex tissues but also provides an integrated perspective of Rab activity in controlling membrane architecture.

Animals↗

The Universal Protein Resource (UniProt): an expanding universe of protein information.

The Universal Protein Resource (UniProt) provides a central resource on protein sequences and functional annotation with three database components, each addressing a key need in protein bioinformatics. The UniProt Knowledgebase (UniProtKB), comprising the manually annotated UniProtKB/Swiss-Prot section and the automatically annotated UniProtKB/TrEMBL section, is the preeminent storehouse of protein annotation. The extensive cross-references, functional and feature annotations and literature-based evidence attribution enable scientists to analyse proteins and query across databases. The UniProt Reference Clusters (UniRef) speed similarity searches via sequence space compression by merging sequences that are 100% (UniRef100), 90% (UniRef90) or 50% (UniRef50) identical. Finally, the UniProt Archive (UniParc) stores all publicly available protein sequences, containing the history of sequence data with links to the source databases. UniProt databases continue to grow in size and in availability of information. Recent and upcoming changes to database contents, formats, controlled vocabularies and services are described. New download availability includes all major releases of UniProtKB, sequence collections by taxonomic division and complete proteomes. A bibliography mapping service has been added, and an ID mapping service will be available soon. UniProt databases can be accessed online at http://www.uniprot.org or downloaded at ftp://ftp.uniprot.org/pub/databases/.

Databases, Protein↗

Annotations and functional analyses of the rice WRKY gene superfamily reveal positive and negative regulators of abscisic acid signaling in aleurone cells.

The WRKY proteins are a superfamily of regulators that control diverse developmental and physiological processes. This family was believed to be plant specific until the recent identification of WRKY genes in nonphotosynthetic eukaryotes. We have undertaken a comprehensive computational analysis of the rice (Oryza sativa) genomic sequences and predicted the structures of 81 OsWRKY genes, 48 of which are supported by full-length cDNA sequences. Eleven OsWRKY proteins contain two conserved WRKY domains, while the rest have only one. Phylogenetic analyses of the WRKY domain sequences provide support for the hypothesis that gene duplication of single- and two-domain WRKY genes, and loss of the WRKY domain, occurred in the evolutionary history of this gene family in rice. The phylogeny deduced from the WRKY domain peptide sequences is further supported by the position and phase of the intron in the regions encoding the WRKY domains. Analyses for chromosomal distributions reveal that 26% of the predicted OsWRKY genes are located on chromosome 1. Among the dozen genes tested, OsWRKY24, -51, -71, and -72 are induced by abscisic acid (ABA) in aleurone cells. Using a transient expression system, we have demonstrated that OsWRKY24 and -45 repress ABA induction of the HVA22 promoter-beta-glucuronidase construct, while OsWRKY72 and -77 synergistically interact with ABA to activate this reporter construct. This study provides a solid base for functional genomics studies of this important superfamily of regulatory genes in monocotyledonous plants and reveals a novel function for WRKY genes, i.e. mediating plant responses to ABA.

Abscisic Acid↗

Protein binding microarrays (PBMs) for rapid, high-throughput characterization of the sequence specificities of DNA binding proteins.

DNA binding proteins play a number of key roles in cells, in processes including transcriptional regulation, recombination, genome rearrangements, and DNA replication, repair, and modification. Of particular interest are the interactions between transcription factors and their DNA binding sites, as they are an integral part of the transcriptional regulatory networks that control gene expression. Despite their importance, the DNA binding specificities of most DNA binding proteins remain unknown, as earlier technologies aimed at characterizing DNA-protein interactions have been time consuming and not highly scalable. We have developed a new DNA microarray-based technology, termed protein binding microarrays (PBMs), that allows rapid, high-throughput characterization of the in vitro DNA binding site sequence specificities of transcription factors in a single day. The resulting DNA binding site data can be used in a number of ways, including for the prediction of the genes regulated by a given transcription factor, annotation of transcription factor function, and functional annotation of the predicted target genes.

Base Sequence↗

ELISA: structure-function inferences based on statistically significant and evolutionarily inspired observations.

UNLABELLED: The problem of functional annotation based on homology modeling is primary to current bioinformatics research. Researchers have noted regularities in sequence, structure and even chromosome organization that allow valid functional cross-annotation. However, these methods provide a lot of false negatives due to limited specificity inherent in the system. We want to create an evolutionarily inspired organization of data that would approach the issue of structure-function correlation from a new, probabilistic perspective. Such organization has possible applications in phylogeny, modeling of functional evolution and structural determination. ELISA (Evolutionary Lineage Inferred from Structural Analysis, http://romi.bu.edu/elisa) is an online database that combines functional annotation with structure and sequence homology modeling to place proteins into sequence-structure-function "neighborhoods". The atomic unit of the database is a set of sequences and structural templates that those sequences encode. A graph that is built from the structural comparison of these templates is called PDUG (protein domain universe graph). We introduce a method of functional inference through a probabilistic calculation done on an arbitrary set of PDUG nodes. Further, all PDUG structures are mapped onto all fully sequenced proteomes allowing an easy interface for evolutionary analysis and research into comparative proteomics. ELISA is the first database with applicability to evolutionary structural genomics explicitly in mind. AVAILABILITY: The database is available at http://romi.bu.edu/elisa.

Amino Acid Sequence↗

Statistically rigorous automated protein annotation.

MOTIVATION: Assignment of putative protein functional annotation by comparative analysis using pre-defined experimental annotations is performed routinely by molecular biologists. The number and statistical significance of these assignments remains a challenge in this era of high-throughput proteomics. A combined statistical method that enables robust, automated protein annotation by reliably expanding existing annotation sets is described. An existing clustering scheme, based on relevant experimental information (e.g. sequence identity, keywords or gene expression data) is required. The method assigns new proteins to these clusters with a measure of reliability. It can also provide human reviewers with a reliability score for both new and previously classified proteins. RESULTS: A dataset of 27 000 annotated Protein Data Bank (PDB) polypeptide chains (of 36 000 chains currently in the PDB) was generated from 23 000 chains classified a priori. AVAILABILITY: PDB annotations and sample software implementation are freely accessible on the Web at http://pmr.sdsc.edu/go

Abstracting and Indexing↗