Search PubMed⌕ Search

Biomedical subjects

Liisa Holm

Publications and source records attributed to Liisa Holm.

18 recordsLinked to original sources

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article↗

Loss of neurturin in frog--comparative genomics study of GDNF family ligand-receptor pairs.

Four different GDNF family ligand (GFL)-receptor (GFRalpha) binding pairs exist in mammals, and they all signal via the RET receptor tyrosine kinase. However, the evolution of these molecules is poorly understood. We identified orthologs of all four GFRalpha receptors and GRAL (GDNF Receptor Alpha-Like) in all vertebrate classes, and a predicted GFR-like protein in several invertebrates. In addition, Gas1 (growth arrest-specific 1), a distant member of the GFR-superfamily, is present in both vertebrates and invertebrates. Analysis of exon structures suggests a common origin of GFR-superfamily proteins and early divergence of Gas1 from the common ancestor. Bony fishes have orthologs of all four mammalian GFLs, consistent with genome duplications in early vertebrates. Surprisingly, the clawed frog and chicken have only three GFLs: synteny analysis indicates loss of neurturin in frog and of persephin in chicken. Evolutionary trace analysis and protein structure homology modeling points at GDNF as the endogenous ligand of frog GFRalpha2.

Amino Acid Sequence↗

Identifying functional gene sets from hierarchically clustered expression data: map of abiotic stress regulated genes in Arabidopsis thaliana.

We present MultiGO, a web-enabled tool for the identification of biologically relevant gene sets from hierarchically clustered gene expression trees (http://ekhidna.biocenter.helsinki.fi/poxo/multigo). High-throughput gene expression measuring techniques, such as microarrays, are nowadays often used to monitor the expression of thousands of genes. Since these experiments can produce overwhelming amounts of data, computational methods that assist the data analysis and interpretation are essential. MultiGO is a tool that automatically extracts the biological information for multiple clusters and determines their biological relevance, and hence facilitates the interpretation of the data. Since the entire expression tree is analysed, MultiGO is guaranteed to report all clusters that share a common enriched biological function, as defined by Gene Ontology annotations. The tool also identifies a plausible cluster set, which represents the key biological functions affected by the experiment. The performance is demonstrated by analysing drought-, cold- and abscisic acid-related expression data sets from Arabidopsis thaliana. The analysis not only identified known biological functions, but also brought into focus the less established connections to defense-related gene clusters. Thus, in comparison to analyses of manually selected gene lists, the systematic analysis of every cluster can reveal unexpected biological phenomena and produce much more comprehensive biological insights to the experiment of interest.

Abscisic Acid↗

Bayesian search of functionally divergent protein subgroups and their function specific residues.

MOTIVATION: The rapid increase in the amount of protein sequence data has created a need for an automated identification of evolutionarily related subgroups from large datasets. The existing methods typically require a priori specification of the number of putative groups, which defines the resolution of the classification solution. RESULTS: We introduce a Bayesian model-based approach to simultaneous identification of evolutionary groups and conserved parts of the protein sequences. The model-based approach provides an intuitive and efficient way of determining the number of groups from the sequence data, in contrast to the ad hoc methods often exploited for similar purposes. Our model recognizes the areas in the sequences that are relevant for the clustering and regards other areas as noise. We have implemented the method using a fast stochastic optimization algorithm which yields a clustering associated with the estimated maximum posterior probability. The method has been shown to have high specificity and sensitivity in simulated and real clustering tasks. With real datasets the method also highlights the residues close to the active site. AVAILABILITY: Software 'kPax' is available at http://www.rni.helsinki.fi/jic/softa.html

Algorithms↗

POXO: a web-enabled tool series to discover transcription factor binding sites.

We present POXO, a comprehensive tool series to discover transcription factor binding sites from co-expressed genes (www.bioinfo.biocenter.helsinki.fi/poxo). POXO manages tasks such as functional evaluation and grouping of genes, sequence retrieval, pattern discovery and pattern verification. It also allows users to tailor analytical pipelines from these tools, with single mouse clicks. One typical pipeline of POXO begins by examining the biological functions that a set of co-expressed genes are involved in. In this examination, the functional coherence of the gene set is evaluated and representative functions are associated with the gene set. This examination can also be used to group genes into functionally similar subsets, if several biological processes are affected in the experiment. The next step in the pipeline is then to discover over-represented nucleotide patterns from the upstream sequences of the selected gene sets. This enables to investigate the possibility that the genes are co-regulated by common cis-elements. If over-represented patterns are found, similar ones can then be clustered together and be verified. The performance of POXO is demonstrated by analysing expression data from pathogen treated Arabidopsis thaliana. In this example, POXO detected activated gene sets and suggested transcription factors responsible for their regulation.

Arabidopsis↗

From sequences to a functional unit.

Functional insights at the gene product level would help the drug discovery industry to effectively tap targets for therapeutics and biomedical applications. A complete functional unit can be multidomain, and it is the co-occurrence and interaction of these multiple domains that determine the function and functional diversity of their gene products. With at least 10% of genes from complete genomes existing in fused form, identifying gene fusion events helps us categorize the protein universe into distinct functional units with only sequence information.

Animals↗

Oligomerization of hantavirus nucleocapsid protein: analysis of the N-terminal coiled-coil domain.

Hantaviruses constitute a genus in the family Bunyaviridae. They are enveloped negative-strand RNA viruses with a tripartite genome encoding the nucleocapsid (N) protein, the two surface glycoproteins Gn and Gc, and an RNA-dependent RNA polymerase. The N protein is the most abundant component of the virion; it encapsidates genomic RNA segments forming ribonucleoproteins and participates in genome transcription and replication as well as virus assembly. In the course of RNA encapsidation, N protein forms intermediate trimers via head-to-head and tail-to-tail interactions. We analyzed the amino-terminal trimerization domain (amino acid residues 1 to 77) of Tula hantavirus using computer modeling, mammalian two-hybrid assay, and immunofluorescence assay. The results obtained were consistent with the existence of an antiparallel coiled-coil stabilized by interactions between hydrophobic residues. Residues L44, V51, and L58 were important for the N-N interaction; other residues, e.g., L25 and V32, also made a contribution, albeit a modest one. Our alignments of the N-terminal domain of the hantaviral N proteins suggest the coiled-coil structure, and hence the mode of N-protein oligomerization, is conserved among hantaviruses.

Amino Acid Sequence↗

Evolution of the GDNF family ligands and receptors.

Four different ligand-receptor binding pairs of the GDNF (glial cell line-derived neurotrophic factor) family exist in mammals, and they all signal via the transmembrane RET receptor tyrosine kinase. In addition, GRAL (GDNF Receptor Alpha-Like) protein of unknown function and Gas1 (growth arrest specific 1) have GDNF family receptor (GFR)-like domains. Orthologs of the four GFRalpha receptors, GRAL and Gas1 are present in all vertebrate classes. In contrast, although bony fishes have orthologs of all four GDNF family ligands (GFLs), one of the ligands, neurturin, is absent in clawed frog and another, persephin, is absent in the chicken genome. Frog GFRalpha2 has selectively evolved possibly to accommodate GDNF as a ligand. The key role of GDNF and its receptor GFRalpha1 in enteric nervous system development is conserved from zebrafish to humans. The role of neurturin, signaling via GFRalpha2, for parasympathetic neuron development is conserved between chicken and mice. The role of artemin and persephin that signal via GFRalpha3 and GFRalpha4, respectively, is unknown in non-mammals. The presence of RET- and GFR-like genes in insects suggests that a ProtoGFR and a ProtoRET arose early in the evolution of bilaterian animals, but when the ProtoGFL diverged from existing transforming growth factor (TGFbeta)-like proteins remains unclear. The four GFLs and GFRalphas were presumably generated by genome duplications at the origin of vertebrates. Loss of neurturin in frog and persephin in chicken suggests functional redundancy in early tetrapods. Functions of non-mammalian GFLs and prechordate RET and GFR-like proteins remain to be explored.

Animals↗

POCO: discovery of regulatory patterns from promoters of oppositely expressed gene sets.

Functionally associated genes tend to be co-expressed, which indicates that they could also be co-regulated. Since co-regulation is usually governed by transcription factors via their specific binding elements, putative regulators can be identified from promoter sets of (co-expressed) genes by screening for over-represented nucleotide patterns. Here, we present a program, POCO, which discovers such over-represented patterns from either one or two promoter sets. Typical microarray experiments yield up- and down-regulated gene sets that may represent, for example, distinct defense pathways. Assuming that a functional transcription factor cannot simultaneously both up- and down-regulate the gene sets, its binding element should respectively be over- and under-represented in the corresponding promoter sets. This idea is implemented in POCO, which tests the hypothesis that the distributions of a pattern differ among three sets of promoters: up-regulated, down-regulated and randomly-chosen. In the program, pattern discovery is based on explicit enumeration of all possible patterns on the alphabet (A, C, G, T and N). The mean occurrences and SDs of the patterns are estimated using bootstrapping and their significance is assessed using ANOVA F-statistics, Tukey's honestly significantly difference test and P-values. The program is freely available at http://ekhidna.biocenter.helsinki.fi/poco.

Binding Sites↗

PSIbase: a database of Protein Structural Interactome map (PSIMAP).

UNLABELLED: Protein Structural Interactome map (PSIMAP) is a global interaction map that describes domain-domain and protein-protein interaction information for known Protein Data Bank structures. It calculates the Euclidean distance to determine interactions between possible pairs of structural domains in proteins. PSIbase is a database and file server for protein structural interaction information calculated by the PSIMAP algorithm. PSIbase also provides an easy-to-use protein domain assignment module, interaction navigation and visual tools. Users can retrieve possible interaction partners of their proteins of interests if a significant homology assignment is made with their query sequences. AVAILABILITY: http://psimap.org and http://psibase.kaist.ac.kr/

Binding Sites↗

ADDA: a domain database with global coverage of the protein universe.

We used the Automatic Domain Decomposition Algorithm (ADDA) to generate a database of protein domain families with complete coverage of all protein sequences. Sequences are split into domains and domains are grouped into protein domain families in a completely automated process. The current database contains domains for more than 1.5 million sequences in more than 40,000 domain families. In particular, there are 3828 novel domain families that do not overlap with the curated domain databases Pfam, SCOP and InterPro. The data are freely available for downloading and querying via a web interface (http://ekhidna.biocenter.helsinki.fi:9801/sqgraph/pairsdb).

Algorithms↗

POBO, transcription factor binding site verification with bootstrapping.

Transcription factors can either activate or repress target genes by binding onto short nucleotide sequence motifs in the promoter regions of these genes. Here, we present POBO, a promoter bootstrapping program, for gene expression data. POBO can be used to detect, compare and verify predetermined transcription factor binding site motifs in the promoters of one or two clusters of co-regulated genes. The program calculates the frequencies of the motif in the input promoter sets. A bootstrap analysis detects significantly over- or underrepresented motifs. The output of the program presents bootstrapped results in picture and text formats. The program was tested with published data from transgenic WRKY70 microarray experiments. Intriguingly, motifs recognized by the WRKY transcription factors of plant defense pathways are similarly enriched in both up- and downregulated clusters. POBO analysis suggests slightly modified hypothetical motifs that discriminate between up- and downregulated clusters. In conclusion, POBO allows easy, fast and accurate verification of putative regulatory motifs. The statistical tests implemented in POBO can be useful in eliminating false positives from the results of pattern discovery programs and increasing the reliability of true positives. POBO is freely available from http://ekhidna.biocenter.helsinki.fi:9801/pobo.

Binding Sites↗

Accurate detection of very sparse sequence motifs.

Protein sequence alignments are more reliable the shorter the evolutionary distance. Here, we align distantly related proteins using many closely spaced intermediate sequences as stepping stones. Such transitive alignments can be generated between any two proteins in a connected set, whether they are direct or indirect sequence neighbors in the underlying library of pairwise alignments. We have implemented a greedy algorithm, MaxFlow, using a novel consistency score to estimate the relative likelihood of alternative paths of transitive alignment. In contrast to traditional profile models of amino acid preferences, MaxFlow models the probability that two positions are structurally equivalent and retains high information content across large distances in sequence space. Thus, MaxFlow is able to identify sparse and narrow active-site sequence signatures which are embedded in high-entropy sequence segments in the structure based multiple alignment of large diverse enzyme superfamilies. In a challenging benchmark based on the urease superfamily, MaxFlow yields better reliability and double coverage compared to available sequence alignment software. This promises to increase information returns from functional and structural genomics, where reliable sequence alignment is a bottleneck to transferring the functional or structural characterization of model proteins to entire protein superfamilies.

Actins↗

Unraveling protein interaction networks with near-optimal efficiency.

The functional characterization of genes and their gene products is the main challenge of the genomic era. Examining interaction information for every gene product is a direct way to assemble the jigsaw puzzle of proteins into a functional map. Here we demonstrate a method in which the information gained from pull-down experiments, in which single proteins act as baits to detect interactions with other proteins, is maximized by using a network-based strategy to select the baits. Because of the scale-free distribution of protein interaction networks, we were able to obtain fast coverage by focusing on highly connected nodes (hubs) first. Unfortunately, locating hubs requires prior global information about the network one is trying to unravel. Here, we present an optimized 'pay-as-you-go' strategy that identifies highly connected nodes using only local information that is collected as successive pull-down experiments are performed. Using this strategy, we estimate that 90% of the human interactome can be covered by 10,000 pull-down experiments, with 50% of the interactions confirmed by reciprocal pull-down experiments.

Algorithms↗

Exhaustive enumeration of protein domain families.

Domains are considered as the basic units of protein folding, evolution, and function. Decomposing each protein into modular domains is thus a basic prerequisite for accurate functional classification of biological molecules. Here, we present ADDA, an automatic algorithm for domain decomposition and clustering of all protein domain families. We use alignments derived from an all-on-all sequence comparison to define domains within protein sequences based on a global maximum likelihood model. In all, 90% of domain boundaries are predicted within 10% of domain size when compared with the manual domain definitions given in the SCOP database. A representative database of 249,264 protein sequences were decomposed into 450,462 domains. These domains were clustered on the basis of sequence similarities into 33,879 domain families containing at least two members with less than 40% sequence identity. Validation against family definitions in the manually curated databases SCOP and PFAM indicates almost perfect unification of various large domain families while contamination by unrelated sequences remains at a low level. The global survey of protein-domain space by ADDA confirms that most large and universal domain families are already described in PFAM and/or SMART. However, a survey of the complete set of mobile modules leads to the identification of 1479 new interesting domain families which shuffle around in multi-domain proteins. The data are publicly available at ftp://ftp.ebi.ac.uk/pub/contrib/heger/adda.

Algorithms↗

Sensitive pattern discovery with 'fuzzy' alignments of distantly related proteins.

MOTIVATION: Evolutionary comparison leads to efficient functional characterisation of hypothetical proteins. Here, our goal is to map specific sequence patterns to putative functional classes. The evolutionary signal stands out most clearly in a maximally diverse set of homologues. This diversity, however, leads to a number of technical difficulties. The targeted patterns-as gleaned from structure comparisons-are too sparse for statistically significant signals of sequence similarity and accurate multiple sequence alignment. RESULTS: We address this problem by a fuzzy alignment model, which probabilistically assigns residues to structurally equivalent positions (attributes) of the proteins. We then apply multivariate analysis to the 'attributes x proteins' matrix. The dimensionality of the space is reduced using non-negative matrix factorization. The method is general, fully automatic and works without assumptions about pattern density, minimum support, explicit multiple alignments, phylogenetic trees, etc. We demonstrate the discovery of biologically meaningful patterns in an extremely diverse superfamily related to urease.

Algorithms↗

A theoretical model for the regulation of Sex-lethal, a gene that controls sex determination and dosage compensation in Drosophila melanogaster.

Cell fate commitment relies upon making a choice between different developmental pathways and subsequently remembering that choice. Experimental studies have thoroughly investigated this central theme in biology for sex determination. In the somatic cells of Drosophila melanogaster, Sex-lethal (Sxl) is the master regulatory gene that specifies sexual identity. We have developed a theoretical model for the initial sex-specific regulation of Sxl expression. The model is based on the well-documented molecular details of the system and uses a stochastic formulation of transcription. Numerical simulations allow quantitative assessment of the role of different regulatory mechanisms in achieving a robust switch. We establish on a formal basis that the autoregulatory loop involved in the alternative splicing of Sxl primary transcripts generates an all-or-none bistable behavior and constitutes an efficient stabilization and memorization device. The model indicates that production of a small amount of early Sxl proteins leaves the autoregulatory loop in its off state. Numerical simulations of mutant genotypes enable us to reproduce and explain the phenotypic effects of perturbations induced in the dosage of genes whose products participate in the early Sxl promoter activation.

Animals↗

Automated detection of remote homology.

The classification of a newly identified protein as a member of a superfamily is important for focusing experiments on its most likely functions. Such classification, often performed by hand, has now been fully automated. This sophisticated new approach takes into account not only alignment scores but also a number of other computable attributes, such as functional sites deduced from sequence conservation patterns.

Automation↗