Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Methods for evaluating clustering algorithms for gene expression data using a reference set of functional classes.

BACKGROUND: A cluster analysis is the most commonly performed procedure (often regarded as a first step) on a set of gene expression profiles. In most cases, a post hoc analysis is done to see if the genes in the same clusters can be functionally correlated. While past successes of such analyses have often been reported in a number of microarray studies (most of which used the standard hierarchical clustering, UPGMA, with one minus the Pearson's correlation coefficient as a measure of dissimilarity), often times such groupings could be misleading. More importantly, a systematic evaluation of the entire set of clusters produced by such unsupervised procedures is necessary since they also contain genes that are seemingly unrelated or may have more than one common function. Here we quantify the performance of a given unsupervised clustering algorithm applied to a given microarray study in terms of its ability to produce biologically meaningful clusters using a reference set of functional classes. Such a reference set may come from prior biological knowledge specific to a microarray study or may be formed using the growing databases of gene ontologies (GO) for the annotated genes of the relevant species. RESULTS: In this paper, we introduce two performance measures for evaluating the results of a clustering algorithm in its ability to produce biologically meaningful clusters. The first measure is a biological homogeneity index (BHI). As the name suggests, it is a measure of how biologically homogeneous the clusters are. This can be used to quantify the performance of a given clustering algorithm such as UPGMA in grouping genes for a particular data set and also for comparing the performance of a number of competing clustering algorithms applied to the same data set. The second performance measure is called a biological stability index (BSI). For a given clustering algorithm and an expression data set, it measures the consistency of the clustering algorithm's ability to produce biologically meaningful clusters when applied repeatedly to similar data sets. A good clustering algorithm should have high BHI and moderate to high BSI. We evaluated the performance of ten well known clustering algorithms on two gene expression data sets and identified the optimal algorithm in each case. The first data set deals with SAGE profiles of differentially expressed tags between normal and ductal carcinoma in situ samples of breast cancer patients. The second data set contains the expression profiles over time of positively expressed genes (ORF's) during sporulation of budding yeast. Two separate choices of the functional classes were used for this data set and the results were compared for consistency. CONCLUSION: Functional information of annotated genes available from various GO databases mined using ontology tools can be used to systematically judge the results of an unsupervised clustering algorithm as applied to a gene expression data set in clustering genes. This information could be used to select the right algorithm from a class of clustering algorithms for the given data set.

Algorithms↗

Identification of a novel protein with guanylyl cyclase activity in Arabidopsis thaliana.

Guanylyl cyclases (GCs) catalyze the formation of the second messenger guanosine 3',5'-cyclic monophosphate (cGMP) from guanosine 5'-triphosphate (GTP). While many cGMP-mediated processes in plants have been reported, no plant molecule with GC activity has been identified. When the Arabidopsis thaliana genome is queried with GC sequences from cyanobacteria, lower and higher eukaryotes no unassigned proteins with significant similarity are found. However, a motif search of the A. thaliana genome based on conserved and functionally assigned amino acids in the catalytic center of annotated GCs returns one candidate that also contains the adjacent glycine-rich domain typical for GCs. In this molecule, termed AtGC1, the catalytic domain is in the N-terminal part. AtGC1 contains the arginine or lysine that participates in hydrogen bonding with guanine and the cysteine that confers substrate specificity for GTP. When AtGC1 is expressed in Escherichia coli, cell extracts yield >2.5 times more cGMP than control extracts and this increase is not nitric oxide dependent. Furthermore, purified recombinant AtGC1 has Mg(2+)-dependent GC activity in vitro and >3 times less adenylyl cyclase activity when assayed with ATP as substrate in the absence of GTP. Catalytic activity in vitro proves that AtGC1 can function either as a monomer or homo-oligomer. AtGC1 is thus not only the first functional plant GC but also, due to its unusual domain organization, a member of a new class of GCs.

Amino Acid Sequence↗

[Annotation to the mitochondrial genome].

After a brief explanation of the mitochondrial function, especially in the relation to the inner-cell coordination, the study analyzed the mitochondrial hypertroph-dilatative cardiomyopathy, myopathy and scrapie which were recently tied to the "D-loop fragment" of the mtDNA. Any primary connection between viral unconventional slow infections and the mitochondrial genome seems unlikely. It is argued in the study that this category of diseases can be much better explained through the transfer of the so-called mobile retroelements.

Animals↗

Web-based analysis of the mouse transcriptome using Genevestigator.

BACKGROUND: Gene function analysis often requires a complex and laborious sequence of laboratory and computer-based experiments. Choosing an effective experimental design generally results from hypotheses derived from prior knowledge or experimentation. Knowledge obtained from meta-analyzing compendia of expression data with annotation libraries can provide significant clues in understanding gene and network function, resulting in better hypotheses that can be tested in the laboratory. DESCRIPTION: Genevestigator is a microarray database and analysis system allowing context-driven queries. Simple but powerful tools allow biologists with little computational background to retrieve information about when, where and how genes are expressed. We manually curated and quality-controlled 3110 mouse Affymetrix arrays from public repositories. Data queries can be run against an annotation library comprising 160 anatomy categories, 12 developmental stage groups, 80 stimuli, and 182 genetic backgrounds or modifications. The quality of results obtained through Genevestigator is illustrated by a number of biological scenarios that are substantiated by other types of experimentation in the literature. CONCLUSION: The Genevestigator-Mouse database effectively provides biologically meaningful results and can be accessed at https://www.genevestigator.ethz.ch.

Animals↗

Mutant laboratory mice with abnormalities in pigmentation: annotated tables.

Mammalian pigment cell research has recently entered a phase of significantly increased activity due largely to the exploitation of the many mutant mouse stocks that are coming on stream. Numerous transgenic, targeted mutagenesis (so-called 'knockouts'), conditional (so-called 'gene switch') and spontaneous mutant mice develop abnormal coat color phenotypes. The number of mice that exhibit such abnormalities is increasing exponentially as genetic engineering methods become routine. Since defined abnormalities in such mutant mice provide important clues to the as yet often poorly understood functional roles of many gene products, this overview includes a corresponding, annotated table of mutant mice with pigmentation alterations. These range from early developmental defects via a large array of coat color abnormalities to a melanoma metastasis model. This overview should provide helpful pointers to investigators who are looking for mouse models to explore or to compare functional activities of genes of interest and for comparing coat color phenotypes of spontaneous or genetically engineered mouse mutants with novel ones. Secondly, this review includes a table of mouse models of specific human diseases with genetically defined pigmentation abnormalities. In summary, this annotated table should serve as a useful reference for anyone interested in the molecular controls of pigmentation.

Animals↗

Experimental data of a single promoter can be used for in silico detection of genes with related regulation in the absence of sequence similarity.

Gene expression is presently a major focus in genome analysis, and the experimental data on regulatory mechanisms and functional transcription factor binding sites are steadily growing. However, the annotation of transcriptional regulation of sequences cannot keep pace with the exponential growth of sequence databases. Employing detailed experimental data of a single promoter or enhancer to predict genes with similar regulation would provide a powerful method to link the literature about transcriptional regulation and sequence databases. To this end, we used information on individual functional transcription factor binding sites to compose in silico promoter and enhancer models of muscle-specific genes and to analyze the rodents section of EMBL with these models. Exhaustive evaluation of all hits revealed every second to third match to be a muscle-associated gene. Moreover, functionally related regulatory regions were detected by our model-based approach even in the absence of sequence similarity. We believe that this new approach is a substanial extension to database analysis by BLAST or FASTA, which are restricted to sequence similarity.

Animals↗

Intraspecies sequence comparisons for annotating genomes.

Analysis of sequence variation among members of a single species offers a potential approach to identify functional DNA elements responsible for biological features unique to that species. Due to its high rate of allelic polymorphism and ease of genetic manipulability, we chose the sea squirt, Ciona intestinalis, to explore intraspecies sequence comparisons for genome annotation. A large number of C. intestinalis specimens were collected from four continents, and a set of genomic intervals were amplified, resequenced, and analyzed to determine the mutation rates at each nucleotide in the sequence. We found that regions with low mutation rates efficiently demarcated functionally constrained sequences: these include a set of noncoding elements, which we showed in C. intestinalis transgenic assays to act as tissue-specific enhancers, as well as the location of coding sequences. This illustrates that comparisons of multiple members of a species can be used for genome annotation, suggesting a path for the annotation of the sequenced genomes of organisms occupying uncharacterized phylogenetic branches of the animal kingdom. It also raises the possibility that the resequencing of a large number of Homo sapiens individuals might be used to annotate the human genome and identify sequences defining traits unique to our species.

Animals↗

The PAS fold. A redefinition of the PAS domain based upon structural prediction.

In the postgenomic era it is essential that protein sequences are annotated correctly in order to help in the assignment of their putative functions. Over 1300 proteins in current protein sequence databases are predicted to contain a PAS domain based upon amino acid sequence alignments. One of the problems with the current annotation of the PAS domain is that this domain exhibits limited similarity at the amino acid sequence level. It is therefore essential, when using proteins with low-sequence similarities, to apply profile hidden Markov model searches for the PAS domain-containing proteins, as for the PFAM database. From recent 3D X-ray and NMR structures, however, PAS domains appear to have a conserved 3D fold as shown here by structural alignment of the six representative 3D-structures from the PDB database. Large-scale modelling of the PAS sequences from the PFAM database against the 3D-structures of these six structural prototypes was performed. All 3D models generated (> 5700) were evaluated using prosaii. We conclude from our large-scale modelling studies that the PAS and PAC motifs (which are separately defined in the PFAM database) are directly linked and that these two motifs form the PAS fold. The existing subdivision in PAS and PAC motifs, as used by the PFAM and SMART databases, appears to be caused by major differences in sequences in the region connecting these two motifs. This region, as has been shown by Gardner and coworkers for human PAS kinase (Amezcua, C.A., Harper, S.M., Rutter, J. & Gardner, K.H. (2002) Structure 10, 1349-1361, [1]), is very flexible and adopts different conformations depending on the bound ligand. Some PAS sequences present in the PFAM database did not produce a good structural model, even after realignment using a structure-based alignment method, suggesting that these representatives are unlikely to have a fold resembling any of the structural prototypes of the PAS domain superfamily.

Amino Acid Sequence↗

Improving the precision of the structure-function relationship by considering phylogenetic context.

Understanding the relationship between protein structure and function is one of the foremost challenges in post-genomic biology. Higher conservation of structure could, in principle, allow researchers to extend current limitations of annotation. However, despite significant research in the area, a precise and quantitative relationship between biochemical function and protein structure has been elusive. Attempts to draw an unambiguous link have often been complicated by pleiotropy, variable transcriptional control, and adaptations to genomic context, all of which adversely affect simple definitions of function. In this paper, I report that integrating genomic information can be used to clarify the link between protein structure and function. First, I present a novel measure of functional proximity between protein structures (F-score). Then, using F-score and other entirely automatic methods measuring structure and phylogenetic similarity, I present a three-dimensional landscape describing their inter-relationship. The result is a "well-shaped" landscape that demonstrates the added value of considering genomic context in inferring function from structural homology. A generalization of methodology presented in this paper can be used to improve the precision of annotation of genes in current and newly sequenced genomes.

Journal Article↗

Engineered viruses to select genes encoding secreted and membrane-bound proteins in mammalian cells.

We have developed a functional genomics tool to identify the subset of cDNAs encoding secreted and membrane-bound proteins within a library (the 'secretome'). A Sindbis virus replicon was engineered such that the envelope protein precursor no longer enters the secretory pathway. cDNA fragments were fused to the mutant precursor and expression screened for their ability to restore membrane localization of envelope proteins. In this way, recombinant replicons were released within infectious viral particles only if the cDNA fragment they contain encodes a secretory signal. By using engineered viral replicons to selectively export cDNAs of interest in the culture medium, the methodology reported here efficiently filters genetic information in mammalian cells without the need to select individual clones. This adaptation of the 'signal trap' strategy is highly sensitive (1/200 000) and efficient. Indeed, of the 2546 inserts that were retrieved after screening various libraries, more than 97% contained a putative signal peptide. These 2473 clones encoded 419 unique cDNAs, of which 77% were previously annotated. Of the 94 cDNAs encoding proteins of unknown function, 24% either had no match in databases or contained a secretory signal that could not be predicted from electronic data.

Animals↗

Assessment of the total number of human transcription units.

Variation in the estimates of the number of genes encoded by the human genome (28,000-120,000) attests to the difficulty of systematically identifying human genes. Sequencing of human chromosome 22 (Chr22) provided the first comprehensive, unbiased view of an entire human chromosome, and intensive analysis of this sequence identified 545 genes and 134 pseudogenes that had similarity or identity to known proteins and/or ESTs and which were listed in the gene annotation (http://www.sanger.ac.uk/HGP/Chr22). This analysis yielded an estimate of approximately 36,000 functional expressed genes in the human genome (and 9000 pseudogenes). However, a key uncertainty in this estimate was that hundreds of additional genes beyond those annotated in the Chr22 sequence are predicted by the gene prediction program Genscan, an unknown number of which might represent additional expressed genes. To determine what fraction of these "predicted novel genes" (PNGs) represents expressed human genes, we used a sensitive RT-PCR assay to detect predicted transcripts in 17 tissues and one cell line. Our results indicate that at least 5000-9000 additional human genes which lack similarity to known genes or proteins exist in the human genome, increasing baseline gene estimates to approximately 41,000-45,000.

Chromosomes, Human, Pair 22↗

Tailored gene array databases: applications in mechanistic toxicology.

MOTIVATION: The development of an annotated global database suitable for a wide range of investigations is a challenging and labor-intensive task. Thus, the development of databases tailored for specific applications remains necessary. For example, in the field of toxicology, no annotated gene array databases are now available that may assist in the correlation of changes in gene activity to cellular functions and processes associated with the toxic response. RESULTS: As an example of a tailored annotated database, an attempt was made to systematize available biological information on genes present on the Affymetrix Rat Toxicology U34 GeneChip, with a focus on how the gene products relate to liver cells and their response to chemical toxins. The information collected was imbedded in a local relational database to analyze data obtained in toxicological gene array experiments with hydrazine-exposed hepatocytes. The advantages and benefits of the tailored database in the biological interpretation of the results are demonstrated.

Abstracting and Indexing↗

Evidence of a large-scale functional organization of mammalian chromosomes.

Evidence from inbred strains of mice indicates that a quarter or more of the mammalian genome consists of chromosome regions containing clusters of functionally related genes. The intense selection pressures during inbreeding favor the coinheritance of optimal sets of alleles among these genetically linked, functionally related genes, resulting in extensive domains of linkage disequilibrium (LD) among a set of 60 genetically diverse inbred strains. Recombination that disrupts the preferred combinations of alleles reduces the ability of offspring to survive further inbreeding. LD is also seen between markers on separate chromosomes, forming networks with scale-free architecture. Combining LD data with pathway and genome annotation databases, we have been able to identify the biological functions underlying several domains and networks. Given the strong conservation of gene order among mammals, the domains and networks we find in mice probably characterize all mammals, including humans.

Animals↗

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software↗

Recent improvements to the SMART domain-based sequence annotation resource.

SMART (Simple Modular Architecture Research Tool, http://smart.embl-heidelberg.de) is a web-based resource used for the annotation of protein domains and the analysis of domain architectures, with particular emphasis on mobile eukaryotic domains. Extensive annotation for each domain family is available, providing information relating to function, subcellular localization, phyletic distribution and tertiary structure. The January 2002 release has added more than 200 hand-curated domain models. This brings the total to over 600 domain families that are widely represented among nuclear, signalling and extracellular proteins. Annotation now includes links to the Online Mendelian Inheritance in Man (OMIM) database in cases where a human disease is associated with one or more mutations in a particular domain. We have implemented new analysis methods and updated others. New advanced queries provide direct access to the SMART relational database using SQL. This database now contains information on intrinsic sequence features such as transmembrane regions, coiled-coils, signal peptides and internal repeats. SMART output can now be easily included in users' documents. A SMART mirror has been created at http://smart.ox.ac.uk.

Animals↗

Argument-predicate distance as a filter for enhancing precision in extracting predications on the genetic etiology of disease.

BACKGROUND: Genomic functional information is valuable for biomedical research. However, such information frequently needs to be extracted from the scientific literature and structured in order to be exploited by automatic systems. Natural language processing is increasingly used for this purpose although it inherently involves errors. A postprocessing strategy that selects relations most likely to be correct is proposed and evaluated on the output of SemGen, a system that extracts semantic predications on the etiology of genetic diseases. Based on the number of intervening phrases between an argument and its predicate, we defined a heuristic strategy to filter the extracted semantic relations according to their likelihood of being correct. We also applied this strategy to relations identified with co-occurrence processing. Finally, we exploited postprocessed SemGen predications to investigate the genetic basis of Parkinson's disease. RESULTS: The filtering procedure for increased precision is based on the intuition that arguments which occur close to their predicate are easier to identify than those at a distance. For example, if gene-gene relations are filtered for arguments at a distance of 1 phrase from the predicate, precision increases from 41.95% (baseline) to 70.75%. Since this proximity filtering is based on syntactic structure, applying it to the results of co-occurrence processing is useful, but not as effective as when applied to the output of natural language processing. In an effort to exploit SemGen predications on the etiology of disease after increasing precision with postprocessing, a gene list was derived from extracted information enhanced with postprocessing filtering and was automatically annotated with GFINDer, a Web application that dynamically retrieves functional and phenotypic information from structured biomolecular resources. Two of the genes in this list are likely relevant to Parkinson's disease but are not associated with this disease in several important databases on genetic disorders. CONCLUSION: Information based on the proximity postprocessing method we suggest is of sufficient quality to be profitably used for subsequent applications aimed at uncovering new biomedical knowledge. Although proximity filtering is only marginally effective for enhancing the precision of relations extracted with co-occurrence processing, it is likely to benefit methods based, even partially, on syntactic structure, regardless of the relation.

Genetic Diseases, Inborn↗

Identification and comparative analysis of components from the signal recognition particle in protozoa and fungi.

BACKGROUND: The signal recognition particle (SRP) is a ribonucleoprotein complex responsible for targeting proteins to the ER membrane. The SRP of metazoans is well characterized and composed of an RNA molecule and six polypeptides. The particle is organized into the S and Alu domains. The Alu domain has a translational arrest function and consists of the SRP9 and SRP14 proteins bound to the terminal regions of the SRP RNA. So far, our understanding of the SRP and its evolution in lower eukaryotes such as protozoa and yeasts has been limited. However, genome sequences of such organisms have recently become available, and we have now analyzed this information with respect to genes encoding SRP components. RESULTS: A number of SRP RNA and SRP protein genes were identified by an analysis of genomes of protozoa and fungi. The sequences and secondary structures of the Alu portion of the RNA were found to be highly variable. Furthermore, proteins SRP9/14 appeared to be absent in certain species. Comparative analysis of the SRP RNAs from different Saccharomyces species resulted in models which contain features shared between all SRP RNAs, but also a new secondary structure element in SRP RNA helix 5. Protein SRP21, previously thought to be present only in Saccharomyces, was shown to be a constituent of additional fungal genomes. Furthermore, SRP21 was found to be related to metazoan and plant SRP9, suggesting that the two proteins are functionally related. CONCLUSIONS: Analysis of a number of not previously annotated SRP components show that the SRP Alu domain is subject to a more rapid evolution than the other parts of the molecule. For instance, the RNA portion is highly variable and the protein SRP9 seems to have evolved into the SRP21 protein in fungi. In addition, we identified a secondary structure element in the Saccharomyces RNA that has been inserted close to the Alu region. Together, these results provide important clues as to the structure, function and evolution of SRP.

Amino Acid Sequence↗

Structure- and function-based characterization of a new phosphoglycolate phosphatase from Thermoplasma acidophilum.

The protein TA0175 has a large number of sequence homologues, most of which are annotated as unknown and a few as belonging to the haloacid dehalogenase superfamily, but has no known biological function. Using a combination of amino acid sequence analysis, three-dimensional crystal structure information, and kinetic analysis, we have characterized TA0175 as phosphoglycolate phosphatase from Thermoplasma acidophilum. The crystal structure of TA0175 revealed two distinct domains, a larger core domain and a smaller cap domain. The large domain is composed of a centrally located five-stranded parallel beta-sheet with strand order S10, S9, S8, S1, S2 and a small beta-hairpin, strands S3 and S4. This central sheet is flanked by a set of three alpha-helices on one side and two helices on the other. The smaller domain is composed of an open faced beta-sandwich represented by three antiparallel beta-strands, S5, S6, and S7, flanked by two oppositely oriented alpha-helices, H3 and H4. The topology of the large domain is conserved; however, structural variation is observed in the smaller domain among the different functional classes of the haloacid dehalogenase superfamily. Enzymatic assays on TA0175 revealed that this enzyme catalyzed the dephosphorylation of phosphoglycolate in vitro with similar kinetic properties seen for eukaryotic phosphoglycolate phosphatase. Activation by divalent cations, especially Mg2+, and competitive inhibition behavior with Cl- ions are similar between TA0175 and phosphoglycolate phosphatase. The experimental evidence presented for TA0175 is indicative of phosphoglycolate phosphatase.

Amino Acid Sequence↗