Search PubMed⌕ Search

Biomedical subjects

Martijn A Huynen

Publications and source records attributed to Martijn A Huynen.

At least 19 recordsLinked to original sources

Benchmarking ortholog identification methods using functional genomics data.

BACKGROUND: The transfer of functional annotations from model organism proteins to human proteins is one of the main applications of comparative genomics. Various methods are used to analyze cross-species orthologous relationships according to an operational definition of orthology. Often the definition of orthology is incorrectly interpreted as a prediction of proteins that are functionally equivalent across species, while in fact it only defines the existence of a common ancestor for a gene in different species. However, it has been demonstrated that orthologs often reveal significant functional similarity. Therefore, the quality of the orthology prediction is an important factor in the transfer of functional annotations (and other related information). To identify protein pairs with the highest possible functional similarity, it is important to qualify ortholog identification methods. RESULTS: To measure the similarity in function of proteins from different species we used functional genomics data, such as expression data and protein interaction data. We tested several of the most popular ortholog identification methods. In general, we observed a sensitivity/selectivity trade-off: the functional similarity scores per orthologous pair of sequences become higher when the number of proteins included in the ortholog groups decreases. CONCLUSION: By combining the sensitivity and the selectivity into an overall score, we show that the InParanoid program is the best ortholog identification method in terms of identifying functionally equivalent proteins.

Algorithms↗

Deciphering the evolution and metabolism of an anammox bacterium from a community genome.

Anaerobic ammonium oxidation (anammox) has become a main focus in oceanography and wastewater treatment. It is also the nitrogen cycle's major remaining biochemical enigma. Among its features, the occurrence of hydrazine as a free intermediate of catabolism, the biosynthesis of ladderane lipids and the role of cytoplasm differentiation are unique in biology. Here we use environmental genomics--the reconstruction of genomic data directly from the environment--to assemble the genome of the uncultured anammox bacterium Kuenenia stuttgartiensis from a complex bioreactor community. The genome data illuminate the evolutionary history of the Planctomycetes and allow us to expose the genetic blueprint of the organism's special properties. Most significantly, we identified candidate genes responsible for ladderane biosynthesis and biological hydrazine metabolism, and discovered unexpected metabolic versatility.

Anaerobiosis↗

Origin and evolution of the peroxisomal proteome.

BACKGROUND: Peroxisomes are ubiquitous eukaryotic organelles involved in various oxidative reactions. Their enzymatic content varies between species, but the presence of common protein import and organelle biogenesis systems support a single evolutionary origin. The precise scenario for this origin remains however to be established. The ability of peroxisomes to divide and import proteins post-translationally, just like mitochondria and chloroplasts, supports an endosymbiotic origin. However, this view has been challenged by recent discoveries that mutant, peroxisome-less cells restore peroxisomes upon introduction of the wild-type gene, and that peroxisomes are formed from the Endoplasmic Reticulum. The lack of a peroxisomal genome precludes the use of classical analyses, as those performed with mitochondria or chloroplasts, to settle the debate. We therefore conducted large-scale phylogenetic analyses of the yeast and rat peroxisomal proteomes. RESULTS: Our results show that most peroxisomal proteins (39-58%) are of eukaryotic origin, comprising all proteins involved in organelle biogenesis or maintenance. A significant fraction (13-18%), consisting mainly of enzymes, has an alpha-proteobacterial origin and appears to be the result of the recruitment of proteins originally targeted to mitochondria. Consistent with the findings that peroxisomes are formed in the Endoplasmic Reticulum, we find that the most universally conserved Peroxisome biogenesis and maintenance proteins are homologous to proteins from the Endoplasmic Reticulum Assisted Decay pathway. CONCLUSION: Altogether our results indicate that the peroxisome does not have an endosymbiotic origin and that its proteins were recruited from pools existing within the primitive eukaryote. Moreover the reconstruction of primitive peroxisomal proteomes suggests that ontogenetically as well as phylogenetically, peroxisomes stem from the Endoplasmic Reticulum. REVIEWERS: This article was reviewed by Arcady Mushegian, Gáspár Jékely and John Logsdon. OPEN PEER REVIEW: Reviewed by Arcady Mushegian, Gáspar Jékely and John Logsdon. For the full reviews, please go to the Reviewers' comments section.

Journal Article↗

Horizontal gene transfer from Bacteria to rumen Ciliates indicates adaptation to their anaerobic, carbohydrates-rich environment.

BACKGROUND: The horizontal transfer of expressed genes from Bacteria into Ciliates which live in close contact with each other in the rumen (the foregut of ruminants) was studied using ciliate Expressed Sequence Tags (ESTs). More than 4000 ESTs were sequenced from representatives of the two major groups of rumen Cilates: the order Entodiniomorphida (Entodinium simplex, Entodinium caudatum, Eudiplodinium maggii, Metadinium medium, Diploplastron affine, Polyplastron multivesiculatum and Epidinium ecaudatum) and the order Vestibuliferida, previously called Holotricha (Isotricha prostoma, Isotricha intestinalis and Dasytricha ruminantium). RESULTS: A comparison of the sequences with the completely sequenced genomes of Eukaryotes and Prokaryotes, followed by large-scale construction and analysis of phylogenies, identified 148 ciliate genes that specifically cluster with genes from the Bacteria and Archaea. The phylogenetic clustering with bacterial genes, coupled with the absence of close relatives of these genes in the Ciliate Tetrahymena thermophila, indicates that they have been acquired via Horizontal Gene Transfer (HGT) after the colonization of the gut by the rumen Ciliates. CONCLUSION: Among the HGT candidates, we found an over-representation (>75%) of genes involved in metabolism, specifically in the catabolism of complex carbohydrates, a rich food source in the rumen. We propose that the acquisition of these genes has greatly facilitated the Ciliates' colonization of the rumen providing evidence for the role of HGT in the adaptation to new niches.

Adaptation, Physiological↗

A global definition of expression context is conserved between orthologs, but does not correlate with sequence conservation.

BACKGROUND: The massive scale of microarray derived gene expression data allows for a global view of cellular function. Thus far, comparative studies of gene expression between species have been based on the level of expression of the gene across corresponding tissues, or on the co-expression of the gene with another gene. RESULTS: To compare gene expression between distant species on a global scale, we introduce the "expression context". The expression context of a gene is based on the co-expression with all other genes that have unambiguous counterparts in both genomes. Employing this new measure, we show 1) that the expression context is largely conserved between orthologs, and 2) that sequence identity shows little correlation with expression context conservation after gene duplication and speciation. CONCLUSION: This means that the degree of sequence identity has a limited predictive quality for differential expression context conservation between orthologs, and thus presumably also for other facets of gene function.

Animals↗

Combinatorial gene regulation in Plasmodium falciparum.

The malaria parasite Plasmodium falciparum has a complicated life cycle with large variations in its gene expression pattern, but it contains relatively few specific transcriptional regulators. To elucidate this paradox, we identified regulatory sequences, using an approach that integrates the sequence conservation among species and the correlation in mRNA expression within a species. Our analysis identified several DNA sequence motifs that are associated with mRNA expression, two of which were previously determined experimentally. We found more putative regulatory sequences per gene in P. falciparum than in other eukaryotes, such as yeast. We propose that Plasmodium uses the few regulatory proteins it has in a combinatorial approach for gene regulation, explaining the relative paucity in regulatory proteins.

Animals↗

Correlation between sequence conservation and the genomic context after gene duplication.

A key complication in comparative genomics for reliable gene function prediction is the existence of duplicated genes. To study the effect of gene duplication on function prediction, we analyze orthologs between pairs of genomes where in one genome the orthologous gene has duplicated after the speciation of the two genomes (i.e. inparalogs). For these duplicated genes we investigate whether the gene that is most similar on the sequence level is also the gene that has retained the ancestral gene-neighborhood. Although the majority of investigated cases show a consistent pattern between sequence similarity and gene-neighborhood conservation, a substantial fraction, 29-38%, is inconsistent. The observation of inconsistency is not the result of a chance outcome owing to a lack of divergence time between inparalogs, but rather it seems to be the result of a chance outcome caused by very similar rates of sequence evolution of both inparalogs relative to their ortholog. If one-to-one orthologous relationships are required, it is advisable to combine contextual information (i.e. gene-neighborhood in prokaryotes and co-expression in eukaryotes) with protein sequence information to predict the most probable functional equivalent ortholog in the presence of inparalogs.

Artifacts↗

Lineage-specific gene loss following mitochondrial endosymbiosis and its potential for function prediction in eukaryotes.

MOTIVATION: The endosymbiotic origin of mitochondria has resulted in a massive horizontal transfer of genetic material from an alpha-proteobacterium to the early eukaryotes. Using large-scale phylogenetic analysis we have previously identified 630 orthologous groups of proteins derived from this event. Here we show that this proto-mitochondrial protein set has undergone extensive lineage-specific gene loss in the eukaryotes, with an average of three losses per orthologous group in a phylogeny of nine species. This gene loss has resulted in a high variability of the alphaproteobacterial-derived gene content of present-day eukaryotic genomes that might reflect functional adaptation to different environments. Proteins functioning in the same biochemical pathway tend to have a similar history of gene loss events, and we use this property to predict functional interactions among proteins in our set.

Animals↗

Tracing the evolution of a large protein complex in the eukaryotes, NADH:ubiquinone oxidoreductase (Complex I).

The increasing availability of sequenced genomes enables the reconstruction of the evolutionary history of large protein complexes. Here, we trace the evolution of NADH:ubiquinone oxidoreductase (Complex I), which has increased in size, by so-called supernumary subunits, from 14 subunits in the bacteria to 30 in the plants and algae, 37 in the fungi and 46 in the mammals. Using a combination of pair-wise and profile-based sequence comparisons at the levels of proteins and the DNA of the sequenced eukaryotic genomes, combined with phylogenetic analyses to establish orthology relationships, we were able to (1) trace the origin of six of the supernumerary subunits to the alpha-proteobacterial ancestor of the mitochondria, (2) detect previously unidentified homology relations between subunits from fungi and mammals, (3) detect previously unidentified subunits in the genomes of several species and (4) document several cases of gene duplications among supernumerary subunits in the eukaryotes. One of these, a duplication of N7BM (B17.2), is particularly interesting as it has been lost from genomes that have also lost Complex I proteins, making it a candidate for a Complex I interacting protein. A parsimonious reconstruction of eukaryotic Complex I evolution shows an initial increase in size that predates the separation of plants, fungi and metazoa, followed by a gradual adding and incidental losses of subunits in the various evolutionary lineages. This evolutionary scenario is in contrast to that for Complex I in the prokaryotes, for which the combination of several separate, and previously independently functioning modules into a single complex has been proposed.

Amino Acid Sequence↗

Variation and evolution of biomolecular systems: searching for functional relevance.

The availability of genome sequences and functional genomics data from multiple species enables us to compare the composition of biomolecular systems like biochemical pathways and protein complexes between species. Here, we review small- and large-scale, "genomics-based" approaches to biomolecular systems variation. In general, caution is required when comparing the results of bioinformatics analyses of genomes or of functional genomics data between species. Limitations to the sensitivity of sequence analysis tools and the noisy nature of genomics data tend to lead to systematic overestimates of the amount of variation. Nevertheless, the results from detailed manual analyses, and of large-scale analyses that filter out systematic biases, point to a large amount of variation in the composition of biomolecular systems. Such observations challenge our understanding of the function of the systems and their individual components and can potentially facilitate the identification and functional characterization of sub-systems within a system. Mapping the inter-species variation of complex biomolecular systems on a phylogenetic species tree allows one to reconstruct their evolution.

Evolution, Molecular↗

The properties of protein family space depend on experimental design.

MOTIVATION: Databases of protein families often exhibit drastically different properties of the protein family space. RESULTS: We compared the properties of protein family space as reflected by exhaustive protein family databases and databases with predefined families. We used TRIBES, Protomap, ProDom and COGs as representatives of the exhaustive databases, and Pfam-A and Superfamily as databases that predefine families. We observe a power-law distribution of family sizes in all these databases, albeit in predefined databases the power-law line collapses before reaching smaller sized families. We discuss the future trends of this power-law distribution and suggest that saturation in the sampling of protein family space will result in a distortion of the power law in small family sizes. For larger genome sizes, predefined databases show logarithmic growth of the number of families per genome, whereas exhaustive databases exhibit a virtually linear relationship. All databases consistently differ in the proportion of protein families shared between taxa. Predefined databases have a larger number of protein families shared between the three domains of life, while exhaustive databases show a much more fragmented distribution. We argue that these discrepancies reflect alternative approaches to the trade-off issue of sensitivity versus specificity in the detection of homologous proteins. We conclude that these properties are complementary rather than contradictory, while describing the protein universe from different perspectives.

Algorithms↗

An anaerobic mitochondrion that produces hydrogen.

Hydrogenosomes are organelles that produce ATP and hydrogen, and are found in various unrelated eukaryotes, such as anaerobic flagellates, chytridiomycete fungi and ciliates. Although all of these organelles generate hydrogen, the hydrogenosomes from these organisms are structurally and metabolically quite different, just like mitochondria where large differences also exist. These differences have led to a continuing debate about the evolutionary origin of hydrogenosomes. Here we show that the hydrogenosomes of the anaerobic ciliate Nyctotherus ovalis, which thrives in the hindgut of cockroaches, have retained a rudimentary genome encoding components of a mitochondrial electron transport chain. Phylogenetic analyses reveal that those proteins cluster with their homologues from aerobic ciliates. In addition, several nucleus-encoded components of the mitochondrial proteome, such as pyruvate dehydrogenase and complex II, were identified. The N. ovalis hydrogenosome is sensitive to inhibitors of mitochondrial complex I and produces succinate as a major metabolic end product--biochemical traits typical of anaerobic mitochondria. The production of hydrogen, together with the presence of a genome encoding respiratory chain components, and biochemical features characteristic of anaerobic mitochondria, identify the N. ovalis organelle as a missing link between mitochondria and hydrogenosomes.

Anaerobiosis↗

Combining data from genomes, Y2H and 3D structure indicates that BolA is a reductase interacting with a glutaredoxin.

Genomes, functional genomics data and 3D structure reflect different aspects of protein function. Here, we combine these data to predict that BolA, a widely distributed protein family with unknown function, is a reductase that interacts with a glutaredoxin. Comparisons at the 3D structure level as well as at the sequence profile level indicate homology between BolA and OsmC, an enzyme that reduces organic peroxides. Complementary to this, comparative analyses of genomes and genomics data provide strong evidence of an interaction between BolA and the mono-thiol glutaredoxin family. The interaction between BolA and a mono-thiol glutaredoxin is of particular interest because BolA does not, in contrast to its homolog OsmC, have evolutionarily conserved cysteines to provide it with reducing equivalents. We propose that BolA uses the mono-thiol glutaredoxin as the source for these.

Genome↗

STRING: known and predicted protein-protein associations, integrated and transferred across organisms.

A full description of a protein's function requires knowledge of all partner proteins with which it specifically associates. From a functional perspective, 'association' can mean direct physical binding, but can also mean indirect interaction such as participation in the same metabolic pathway or cellular process. Currently, information about protein association is scattered over a wide variety of resources and model organisms. STRING aims to simplify access to this information by providing a comprehensive, yet quality-controlled collection of protein-protein associations for a large number of organisms. The associations are derived from high-throughput experimental data, from the mining of databases and literature, and from predictions based on genomic context analysis. STRING integrates and ranks these associations by benchmarking them against a common reference set, and presents evidence in a consistent and intuitive web interface. Importantly, the associations are extended beyond the organism in which they were originally described, by automatic transfer to orthologous protein pairs in other organisms, where applicable. STRING currently holds 730,000 proteins in 180 fully sequenced organisms, and is available at http://string.embl.de/.

Databases, Protein↗

Genome trees and the nature of genome evolution.

Genome trees are a means to capture the overwhelming amount of phylogenetic information that is present in genomes. Different formalisms have been introduced to reconstruct genome trees on the basis of various aspects of the genome. On the basis of these aspects, we separate genome trees into five classes: (a) alignment-free trees based on statistic properties of the genome, (b) gene content trees based on the presence and absence of genes, (c) trees based on chromosomal gene order, (d) trees based on average sequence similarity, and (e) phylogenomics-based genome trees. Despite their recent development, genome tree methods have already had some impact on the phylogenetic classification of bacterial species. However, their main impact so far has been on our understanding of the nature of genome evolution and the role of horizontal gene transfer therein. An ideal genome tree method should be capable of using all gene families, including those containing paralogs, in a phylogenomics framework capitalizing on existing methods in conventional phylogenetic reconstruction. We expect such sophisticated methods to help us resolve the branching order between the main bacterial phyla.

Evolution, Molecular↗

Shaping the mitochondrial proteome.

Mitochondria are eukaryotic organelles that originated from a single bacterial endosymbiosis some 2 billion years ago. The transition from the ancestral endosymbiont to the modern mitochondrion has been accompanied by major changes in its protein content, the so-called proteome. These changes included complete loss of some bacterial pathways, amelioration of others and gain of completely new complexes of eukaryotic origin such as the ATP/ADP translocase and most of the mitochondrial protein import machinery. This renewal of proteins has been so extensive that only 14-16% of modern mitochondrial proteome has an origin that can be traced back to the bacterial endosymbiont. The rest consists of proteins of diverse origin that were eventually recruited to function in the organelle. This shaping of the proteome content reflects the transformation of mitochondria into a highly specialized organelle that, besides ATP production, comprises a variety of functions within the eukaryotic metabolism. Here we review recent advances in the fields of comparative genomics and proteomics that are throwing light on the origin and evolution of the mitochondrial proteome.

Animals↗

Gene co-regulation is highly conserved in the evolution of eukaryotes and prokaryotes.

Differences between species have been suggested to largely reside in the network of connections among the genes. Nevertheless, the rate at which these connections evolve has not been properly quantified. Here, we measure the extent to which co-regulation between pairs of genes is conserved over large phylogenetic distances; between two eukaryotes Caenorhabditis elegans and Saccharomyces cerevisiae, and between two prokaryotes Escherichia coli and Bacillus subtilis. We first construct a reliable set of co-regulated genes by combining various functional genomics data from yeast, and subsequently determine conservation of co-regulation in worm from the distribution of co-expression values. For B.subtilis and E.coli, we use known operons and regulons. We find that between 76 and 80% of the co-regulatory connections are conserved between orthologous pairs of genes, which is very high compared with previous estimates and expectations regarding network evolution. We show that in the case of gene duplication after speciation, one of the two inparalogous genes tends to retain its original co-regulatory relationship, while the other loses this link and is presumably free for differentiation or sub-functionalization. The high level of co-regulation conservation implies that reliably predicted functional relationships from functional genomics data in one species can be transferred with high accuracy to another species when that species also harbours the associated genes.

Animals↗

The yeast coexpression network has a small-world, scale-free architecture and can be explained by a simple model.

We investigated the gene coexpression network in Saccharomyces cerevisiae, in which genes are linked when they are coregulated. This network is shown to have a scale-free, small-world architecture. Such architecture is typical of biological networks in which the nodes are connected when they are involved in the same biological process. Current models for the evolution of intracellular networks do not adequately reproduce the features that we observe in the network. We therefore derive a new model for its evolution based on the observation that there is a positive correlation between the sequence similarity of paralogues and their probability of coexpression or sharing of transcription factor binding sites (TFBSs). The simple, neutralist's model consists of (1) coduplication of genes with their TFBSs, (2) deletion and duplication of individual TFBSs and (3) gene loss. A network is constructed by connecting genes that share multiple TFBSs. Our model reproduces the scale-free, small-world architecture of the coregulation network and the homology relations between coregulated genes without the need for selection either at the level of the network structure or at the level of gene regulation.

Binding Sites↗