Search PubMed⌕ Search

Biomedical subjects

Edward M Marcotte

Publications and source records attributed to Edward M Marcotte.

At least 19 recordsLinked to original sources

Chromatographic alignment of ESI-LC-MS proteomics data sets by ordered bijective interpolated warping.

Mass spectrometry proteomics typically relies upon analyzing outcomes of single analyses; however, comparing raw data across multiple experiments should enhance both peptide/protein identification and quantitation. In the absence of convincing tandem MS identifications, comparing peptide quantities between experiments (or fractions) requires the chromatographic alignment of MS signals. An extension of dynamic time warping (DTW), termed ordered bijective interpolated warping (OBI-Warp), is presented and used to align a variety of electrospray ionization liquid chromatography mass spectrometry (ESI-LC-MS) proteomics data sets. An algorithm to produce a bijective (one-to-one) function from DTW output is coupled with piecewise cubic hermite interpolation to produce a smooth warping function. Data sets were chosen to represent a broad selection of ESI-LC-MS alignment cases. High confidence, overlapping tandem mass spectra are used as standards to optimize and compare alignment parameters. We determine that Pearson's correlation coefficient as a measure of spectra similarity outperforms covariance, dot product, and Euclidean distance in its ability to produce correct alignments with optimal and suboptimal alignment parameters. We demonstrate the importance of penalizing gaps for best alignments. Using optimized parameters, we show that OBI-Warp produces alignments consistent with time standards across these data sets. The source and executables are released under MIT style license at http://obi-warp.sourceforge.net/.

Chromatography, Liquid↗

Global metabolic changes following loss of a feedback loop reveal dynamic steady states of the yeast metabolome.

Metabolic enzymes control cellular metabolite concentrations dynamically in response to changing environmental and intracellular conditions. Such real-time feedback regulation suggests the global metabolome may sample distinct dynamic steady states, forming "basins of stability" in the energy landscape of possible metabolite concentrations and enzymatic activities. Using metabolite, protein and transcriptional profiling, we characterize three dynamic steady states of the yeast metabolome that form by perturbing synthesis of the universal methyl donor S-adenosylmethionine (AdoMet). Conversion between these states is driven by replacement of serine with glycine+formate in the media, loss of feedback inhibition control by the metabolic enzyme Met13, or both. The latter causes hyperaccumulation of methionine and AdoMet, and dramatic global compensatory changes in the metabolome, including differences in amino acid and sugar metabolism, and possibly in the global nitrogen balance, ultimately leading to a G1/S phase cell cycle delay. Global metabolic changes are not necessarily accompanied by global transcriptional changes, and metabolite-controlled post-transcriptional regulation of metabolic enzymes is clearly evident.

Feedback, Physiological↗

A fast coarse filtering method for peptide identification by mass spectrometry.

MOTIVATION: We reformulate the problem of comparing mass-spectra by mapping spectra to a vector space model. Our search method leverages a metric space indexing algorithm to produce an initial candidate set, which can be followed by any fine ranking scheme. RESULTS: We consider three distance measures integrated into a multi-vantage point index structure. Of these, a semi-metric fuzzy-cosine distance using peptide precursor mass constraints performs the best. The index acts as a coarse, lossless filter with respect to the SEQUEST and ProFound scoring schemes, reducing the number of distance computations and returned candidates for fine filtering to about 0.5% and 0.02% of the database respectively. The fuzzy cosine distance term improves specificity over a peptide precursor mass filter, reducing the number of returned candidates by an order of magnitude. Run time measurements suggest proportional speedups in overall search times. Using an implementation of ProFound's Bayesian score as an example of a fine filter on a test set of Escherichia coli protein fragmentation spectra, the top results of our sample system are consistent with that of SEQUEST.

Algorithms↗

Systematic profiling of cellular phenotypes with spotted cell microarrays reveals mating-pheromone response genes.

We have developed spotted cell microarrays for measuring cellular phenotypes on a large scale. Collections of cells are printed, stained for subcellular features, then imaged via automated, high-throughput microscopy, allowing systematic phenotypic characterization. We used this technology to identify genes involved in the response of yeast to mating pheromone. Besides morphology assays, cell microarrays should be valuable for high-throughput in situ hybridization and immunoassays, enabling new classes of genetic assays based on cell imaging.

Gene Expression Profiling↗

Synthetic biology: engineering Escherichia coli to see light.

We have designed a bacterial system that is switched between different states by red light. The system consists of a synthetic sensor kinase that allows a lawn of bacteria to function as a biological film, such that the projection of a pattern of light on to the bacteria produces a high-definition (about 100 megapixels per square inch), two-dimensional chemical image. This spatial control of bacterial gene expression could be used to 'print' complex biological materials, for example, and to investigate signalling pathways through precise spatial and temporal control of their phosphorylation steps.

Agar↗

Consolidating the set of known human protein-protein interactions in preparation for large-scale mapping of the human interactome.

BACKGROUND: Extensive protein interaction maps are being constructed for yeast, worm, and fly to ask how the proteins organize into pathways and systems, but no such genome-wide interaction map yet exists for the set of human proteins. To prepare for studies in humans, we wished to establish tests for the accuracy of future interaction assays and to consolidate the known interactions among human proteins. RESULTS: We established two tests of the accuracy of human protein interaction datasets and measured the relative accuracy of the available data. We then developed and applied natural language processing and literature-mining algorithms to recover from Medline abstracts 6,580 interactions among 3,737 human proteins. A three-part algorithm was used: first, human protein names were identified in Medline abstracts using a discriminator based on conditional random fields, then interactions were identified by the co-occurrence of protein names across the set of Medline abstracts, filtering the interactions with a Bayesian classifier to enrich for legitimate physical interactions. These mined interactions were combined with existing interaction data to obtain a network of 31,609 interactions among 7,748 human proteins, accurate to the same degree as the existing datasets. CONCLUSION: These interactions and the accuracy benchmarks will aid interpretation of current functional genomics data and provide a basis for determining the quality of future large-scale human protein interaction assays. Projecting from the approximately 15 interactions per protein in the best-sampled interaction set to the estimated 25,000 human genes implies more than 375,000 interactions in the complete human protein interaction network. This set therefore represents no more than 10% of the complete network.

Algorithms↗

Protein function prediction using the Protein Link EXplorer (PLEX).

UNLABELLED: We introduce the Protein Link EXplorer (PLEX), a web-based environment that allows the construction of a phylogenetic profile for any given amino acid sequence, and its comparison with profiles of approximately 350,000 predicted genes from 89 genomes, as a means of interactively identifying functionally linked genes and predicting protein function. PLEX can be searched iteratively and also enables searches for chromosomal gene neighbors and Rosetta Stone linkages. PLEX search results are accompanied by quantitative estimates of linkage confidence, enabling users to take advantage of coinheritance, operon and gene fusion-based methods for inferring gene function and reconstructing cellular systems and pathways. AVAILABILITY: http://bioinformatics.icmb.utexas.edu/plex

Algorithms↗

Comparative experiments on learning information extractors for proteins and their interactions.

OBJECTIVE: Automatically extracting information from biomedical text holds the promise of easily consolidating large amounts of biological knowledge in computer-accessible form. This strategy is particularly attractive for extracting data relevant to genes of the human genome from the 11 million abstracts in Medline. However, extraction efforts have been frustrated by the lack of conventions for describing human genes and proteins. We have developed and evaluated a variety of learned information extraction systems for identifying human protein names in Medline abstracts and subsequently extracting information on interactions between the proteins. METHODS AND MATERIAL: We used a variety of machine learning methods to automatically develop information extraction systems for extracting information on gene/protein name, function and interactions from Medline abstracts. We present cross-validated results on identifying human proteins and their interactions by training and testing on a set of approximately 1000 manually-annotated Medline abstracts that discuss human genes/proteins. RESULTS: We demonstrate that machine learning approaches using support vector machines and maximum entropy are able to identify human proteins with higher accuracy than several previous approaches. We also demonstrate that various rule induction methods are able to identify protein interactions with higher precision than manually-developed rules. CONCLUSION: Our results show that it is promising to use machine learning to automatically build systems for extracting information from biomedical text. The results also give a broad picture of the relative strengths of a wide variety of methods when tested on a reasonably large human-annotated corpus.

Algorithms↗

Mass spectrometry of the M. smegmatis proteome: protein expression levels correlate with function, operons, and codon bias.

The fast-growing bacterium Mycobacterium smegmatis is a model mycobacterial system, a nonpathogenic soil bacterium that nonetheless shares many features with the pathogenic Mycobacterium tuberculosis, the causative agent of tuberculosis. The study of M. smegmatis is expected to shed light on mechanisms of mycobacterial growth and complex lipid metabolism, and provides a tractable system for antimycobacterial drug development. Although the M. smegmatis genome sequence is not yet completed, we used multidimensional chromatography and tandem mass spectrometry, in combination with the partially completed genome sequence, to detect and identify a total of 901 distinct proteins from M. smegmatis over the course of 25 growth conditions, providing experimental annotation for many predicted genes with an approximately 5% false-positive identification rate. We observed numerous proteins involved in energy production (9.8% of expressed proteins), protein translation (8.7%), and lipid biosynthesis (5.4%); 33% of the 901 proteins are of unknown function. Protein expression levels were estimated from the number of observations of each protein, allowing measurement of differential expression of complete operons, and the comparison of the stationary and exponential phase proteomes. Expression levels are correlated with proteins' codon biases and mRNA expression levels, as measured by comparison with codon adaptation indices, principle component analysis of codon frequencies, and DNA microarray data. This observation is consistent with notions that either (1) prokaryotic protein expression levels are largely preset by codon choice, or (2) codon choice is optimized for consistency with average expression levels regardless of the mechanism of regulating expression.

Bacterial Proteins↗

A probabilistic functional network of yeast genes.

A conceptual framework for integrating diverse functional genomics data was developed by reinterpreting experiments to provide numerical likelihoods that genes are functionally linked. This allows direct comparison and integration of different classes of data. The resulting probabilistic gene network estimates the functional coupling between genes. Within this framework, we reconstructed an extensive, high-quality functional gene network for Saccharomyces cerevisiae, consisting of 4681 (approximately 81%) of the known yeast genes linked by approximately 34,000 probabilistic linkages comparable in accuracy to small-scale interaction assays. The integrated linkages distinguish true from false-positive interactions in earlier data sets; new interactions emerge from genes' network contexts, as shown for genes in chromatin modification and ribosome biogenesis.

Bayes Theorem↗

LGL: creating a map of protein function with an algorithm for visualizing very large biological networks.

Networks are proving to be central to the study of gene function, protein-protein interaction, and biochemical pathway data. Visualization of networks is important for their study, but visualization tools are often inadequate for working with very large biological networks. Here, we present an algorithm, called large graph layout (LGL), which can be used to dynamically visualize large networks on the order of hundreds of thousands of vertices and millions of edges. LGL applies a force-directed iterative layout guided by a minimal spanning tree of the network in order to generate coordinates for the vertices in two or three dimensions, which are subsequently visualized and interactively navigated with companion programs. We demonstrate the use of LGL in visualizing an extensive protein map summarizing the results of approximately 21 billion sequence comparisons between 145579 proteins from 50 genomes. Proteins are positioned in the map according to sequence homology and gene fusions, with the map ultimately serving as a theoretical framework that integrates inferences about gene function derived from sequence homology, remote homology, gene fusions, and higher-order fusions. We confirm that protein neighbors in the resulting map are functionally related, and that distinct map regions correspond to distinct cellular systems, enabling a computational strategy for discovering proteins' functions on the basis of the proteins' map positions. Using the map produced by LGL, we infer general functions for 23 uncharacterized protein families.

Algorithms↗

Development through the eyes of functional genomics.

In many of the model organisms used to study development, it is becoming relatively routine to carry out global analyses of gene function. These analyses take many forms, from microarray analyses to the construction of physical interaction maps to the systematic analyses of loss-of-function phenotypes. Such large-scale datasets can be integrated to generate complex gene networks, and we explore how these gene networks can contribute to an understanding of developmental pathways. In particular, we examine how combining large-scale expression experiments and gene networks may move us towards a molecular description of the events of development, embodied in a succession of stage-specific subnetworks sampled from an organism's overall gene network.

Animals↗

Protein interaction networks from yeast to human.

Protein interaction networks summarize large amounts of protein-protein interaction data, both from individual, small-scale experiments and from automated high-throughput screens. The past year has seen a flood of new experimental data, especially on metazoans, as well as an increasing number of analyses designed to reveal aspects of network topology, modularity and evolution. As only minimal progress has been made in mapping the human proteome using high-throughput screens, the transfer of interaction information within and across species has become increasingly important. With more and more heterogeneous raw data becoming available, proper data integration and quality control have become essential for reliable protein network reconstruction, and will be especially important for reconstructing the human protein interaction network.

Animals↗

A probabilistic view of gene function.

Cells are controlled by the complex and dynamic actions of thousands of genes. With the sequencing of many genomes, the key problem has shifted from identifying genes to knowing what the genes do; we need a framework for expressing that knowledge. Even the most rigorous attempts to construct ontological frameworks describing gene function (e.g., the Gene Ontology project) ultimately rely on manual curation and are thus labor-intensive and subjective. But an alternative exists: the field of functional genomics is piecing together networks of gene interactions, and although these data are currently incomplete and error-prone, they provide a glimpse of a new, probabilistic view of gene function. We outline such a framework, which revolves around a statistical description of gene interactions derived from large, systematically compiled data sets. In this probabilistic view, pleiotropy is implicit, all data have errors and the definition of gene function is an iterative process that ultimately converges on the correct functions. The relationships between the genes are defined by the data, not by hand. Even this comprehensive view fails to capture key aspects of gene function, not least their dynamics in time and space, showing that there are limitations to the model that must ultimately be addressed.

Animals↗

Diametrical clustering for identifying anti-correlated gene clusters.

MOTIVATION: Clustering genes based upon their expression patterns allows us to predict gene function. Most existing clustering algorithms cluster genes together when their expression patterns show high positive correlation. However, it has been observed that genes whose expression patterns are strongly anti-correlated can also be functionally similar. Biologically, this is not unintuitive-genes responding to the same stimuli, regardless of the nature of the response, are more likely to operate in the same pathways. RESULTS: We present a new diametrical clustering algorithm that explicitly identifies anti-correlated clusters of genes. Our algorithm proceeds by iteratively (i). re-partitioning the genes and (ii). computing the dominant singular vector of each gene cluster; each singular vector serving as the prototype of a 'diametric' cluster. We empirically show the effectiveness of the algorithm in identifying diametrical or anti-correlated clusters. Testing the algorithm on yeast cell cycle data, fibroblast gene expression data, and DNA microarray data from yeast mutants reveals that opposed cellular pathways can be discovered with this method. We present systems whose mRNA expression patterns, and likely their functions, oppose the yeast ribosome and proteosome, along with evidence for the inverse transcriptional regulation of a number of cellular systems.

Algorithms↗

Expression deconvolution: a reinterpretation of DNA microarray data reveals dynamic changes in cell populations.

Cells grow in dynamically evolving populations, yet this aspect of experiments often goes unmeasured. A method is proposed for measuring the population dynamics of cells on the basis of their mRNA expression patterns. The population's expression pattern is modeled as the linear combination of mRNA expression from pure samples of cells, allowing reconstruction of the relative proportions of pure cell types in the population. Application of the method, termed expression deconvolution, to yeast grown under varying conditions reveals the population dynamics of the cells during the cell cycle, during the arrest of cells induced by DNA damage and the release of arrest in a cell cycle checkpoint mutant, during sporulation, and following environmental stress. Using expression deconvolution, cell cycle defects are detected and temporally ordered in 146 yeast deletion mutants; six of these defects are independently experimentally validated. Expression deconvolution allows a reinterpretation of the cell cycle dynamics underlying all previous microarray experiments and can be more generally applied to study most forms of cell population dynamics.

Cell Cycle↗

Discovery of uncharacterized cellular systems by genome-wide analysis of functional linkages.

We introduce a general computational method, applicable on a genome-wide scale, for the systematic discovery of uncharacterized cellular systems. Quantitative analysis of the coinheritance of pairs of genes among different organisms, calculated using phylogenetic profiles, allows the prediction of thousands of functional linkages between the corresponding proteins. A comparison of these functional linkages to known pathways reveals that calculated linkages are comparable in accuracy to genome-wide yeast two-hybrid screens or mass spectrometry interaction assays. In aggregate, these linkages describe the structure of large-scale networks, with the resulting yeast network composed of 3,875 linkages among 804 proteins, and the resulting pathogenic Escherichia coli network composed of 2,043 linkages among 828 proteins. The search of such networks for groups of uncharacterized, linked proteins led to the identification of 27 novel cellular systems from one nonpathogenic and three pathogenic bacterial genomes.

Algorithms↗