Search PubMed⌕ Search

Biomedical subjects

Eoin Fahy

Publications and source records attributed to Eoin Fahy.

14 recordsLinked to original sources

Structure-centric searching enables global mapping of the public metabolome.

Searching and learning from aggregated public metabolomics data spanning thousands of studies remained largely inaccessible. Here we present StructureMASST, a web-based application enabling scalable, structure-centric searches across public metabolomics repositories using molecule names or chemical representations. It queries a precomputed knowledgebase of 2.19 billion spectral matches and 420 million metadata links, supports modification-tolerant and mass-shift searches, and maps chemical structures across taxonomy, biological context and environmental conditions to accelerate discovery.

Journal Article↗

The Data Distillery: A Graph Framework for Semantic Integration and Querying of Biomedical Data.

The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.

Journal Article↗

A comprehensive classification system for lipids.

Lipids are produced, transported, and recognized by the concerted actions of numerous enzymes, binding proteins, and receptors. A comprehensive analysis of lipid molecules, "lipidomics," in the context of genomics and proteomics is crucial to understanding cellular physiology and pathology; consequently, lipid biology has become a major research target of the postgenomic revolution and systems biology. To facilitate international communication about lipids, a comprehensive classification of lipids with a common platform that is compatible with informatics requirements has been developed to deal with the massive amounts of data that will be generated by our lipid community. As an initial step in this development, we divide lipids into eight categories (fatty acyls, glycerolipids, glycerophospholipids, sphingolipids, sterol lipids, prenol lipids, saccharolipids, and polyketides) containing distinct classes and subclasses of molecules, devise a common manner of representing the chemical structures of individual lipids and their derivatives, and provide a 12 digit identifier for each unique lipid molecule. The lipid classification scheme is chemically based and driven by the distinct hydrophobic and hydrophilic elements that compose the lipid. This structured vocabulary will facilitate the systematization of lipid biology and enable the cataloging of lipids and their properties in a way that is compatible with other macromolecular databases.

Database Management Systems↗

MITOPRED: a web server for the prediction of mitochondrial proteins.

MITOPRED web server enables prediction of nucleus-encoded mitochondrial proteins in all eukaryotic species. Predictions are made using a new algorithm based primarily on Pfam domain occurrence patterns in mitochondrial and non-mitochondrial locations. Pre-calculated predictions are instantly accessible for proteomes of Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila, Homo sapiens, Mus musculus and Arabidopsis species as well as all the eukaryotic sequences in the Swiss-Prot and TrEMBL databases. Queries, at different confidence levels, can be made through four distinct options: (i) entering Swiss-Prot/TrEMBL accession numbers; (ii) uploading a local file with such accession numbers; (iii) entering protein sequences; (iv) uploading a local file containing protein sequences in FASTA format. Automated updates are scheduled for the pre-calculated prediction database so as to provide access to the most current data. The server, its documentation and the data are available from http://mitopred.sdsc.edu.

Algorithms↗

MITOPRED: a genome-scale method for prediction of nucleus-encoded mitochondrial proteins.

MOTIVATION: Currently available methods for the prediction of subcellular location of mitochondrial proteins rely largely on the presence of mitochondrial targeting signals in the protein sequences. However, a large fraction of mitochondrial proteins lack such signals, making those tools ineffective for genome-scale prediction of mitochondria-targeted proteins. Here, we propose a method for genome-scale prediction of nucleus-encoded mitochondrial proteins. The new method, MITOPRED, is based on the Pfam domain occurrence patterns and the amino acid compositional differences between mitochondrial and non-mitochondrial proteins. RESULTS: MITOPRED could predict mitochondrial proteins with 100% specificity at a 44% sensitivity rate and with 67% specificity at 99% sensitivity. Additionally, it was sufficiently robust to predict mitochondrial proteins across different eukaryotic species with similar accuracy. Based on Matthews correlation coefficient measure, the prediction performance of MITOPRED is clearly superior (0.73) to those of the two popular methods TargetP (0.51) and PSORT (0.53). Using this method, we predicted the nucleus-encoded mitochondrial proteins from six complete genomes (three invertebrate, two vertebrate and one plant species) and estimated the total number in each genome. In human, our method estimated the existence of 1362 mitochondrial proteins corresponding to 4.8% of the total proteome. AVAILABILITY: MITOPRED program is freely accessible at http://mitopred.sdsc.edu. Source code is available on request from the authors. SUPPLEMENTARY INFORMATION: Training data sets are also available at http://mitopred.sdsc.edu

Algorithms↗

MitoProteome: mitochondrial protein sequence database and annotation system.

MitoProteome is an object-relational mitochondrial protein sequence database and annotation system. The initial release contains 847 human mitochondrial protein sequences, derived from public sequence databases and mass spectrometric analysis of highly purified human heart mitochondria. Each sequence is manually annotated with primary function, subfunction and subcellular location, and extensively annotated in an automated process with data extracted from external databases, including gene information from LocusLink and Ensembl; disease information from OMIM; protein-protein interaction data from MINT and DIP; functional domain information from Pfam; protein fingerprints from PRINTS; protein family and family-specific signatures from InterPro; structure data from PDB; mutation data from PMD; BLAST homology data from NCBI NR; and proteins found to be related based on LocusLink and SWISS-PROT references and sequence and taxonomy data. By highly automating the processes of maintaining the MitoProteome Protein List and extracting relevant data from external databases, we are able to present a dynamic database, updated frequently to reflect changes in public resources. The MitoProteome database is publicly available at http://www. mitoproteome.org/. Users may browse and search MitoProteome, and access a complete compilation of data relevant to each protein of interest, cross-linked to external databases.

Computational Biology↗

Oxidative post-translational modification of tryptophan residues in cardiac mitochondrial proteins.

We examined the distribution of N-formylkynurenine, a product of the dioxidation of tryptophan residues in proteins, throughout the human heart mitochondrial proteome. This oxidized amino acid is associated with a distinct subset of proteins, including an over-representation of complex I subunits as well as complex V subunits and enzymes involved in redox metabolism. No relationship was observed between the tryptophan modification and methionine oxidation, a known artifact of sample handling. As the mitochondria were isolated from normal human heart tissue and not subject to any artificially induced oxidative stress, we suggest that the susceptible tryptophan residues in this group of proteins are "hot spots" for oxidation in close proximity to a source of reactive oxygen species in respiring mitochondria.

Amino Acid Sequence↗

The subunit composition of the human NADH dehydrogenase obtained by rapid one-step immunopurification.

Defects of the NADH dehydrogenase complex are predominantly manifested in mitochondrial diseases and are significantly associated with the development of many late onset neurological disorders such as Parkinson's disease. Here we describe an immunocapture procedure for isolating this multisubunit membrane-bound complex from human tissue. Using small amounts of immunoisolated protein, one-dimensional and two-dimensional gel electrophoresis, matrix-assisted laser desorption ionization time-of-flight (MALDI-TOF) peptide mass finger printing (PMF), and nanoflow liquid chromatography mass spectrometry/mass spectrometry (LC-MS/MS), we can resolve and identify the human homologues of 42 polypeptides detected so far in the more extensively studied beef heart complex I. These polypeptides include the GRIM-19 protein, which is claimed to be involved in apoptosis, a polypeptide first identified by gene screening as a neuronal protein, as well as a protein thought to be in differentiation linked processes. The concordance of data from human and bovine complex I isolated by different procedures adds to the certainty that these novel proteins of seemingly diverse function are a part of complex I.

Animals↗

Characterization of the human heart mitochondrial proteome.

To gain a better understanding of the critical role of mitochondria in cell function, we have compiled an extensive catalogue of the mitochondrial proteome using highly purified mitochondria from normal human heart tissue. Sucrose gradient centrifugation was employed to partially resolve protein complexes whose individual protein components were separated by one-dimensional PAGE. Total in-gel processing and subsequent detection by mass spectrometry and rigorous bioinformatic analysis yielded a total of 615 distinct protein identifications. All protein pI values, molecular weight ranges, and hydrophobicities were represented. The coverage of the known subunits of the oxidative phosphorylation machinery within the inner mitochondrial membrane was >90%. A significant proportion of identified proteins are involved in signaling, RNA, DNA, and protein synthesis, ion transport, and lipid metabolism. The biochemical roles of 19% of the identified proteins have not been defined. This database of proteins provides a comprehensive resource for the discovery of novel mitochondrial functions and pathways.

Adolescent↗

Global organellar proteomics.

Cataloging the proteomes of single-celled microorganisms, cells, biological fluids, tissue and whole organisms is being undertaken at a rapid pace as advances are made in protein and peptide separation, detection and identification. For metazoans, subcellular organelles represent attractive targets for global proteome analysis because they represent discrete functional units, their complexity in protein composition is reduced relative to whole cells and, when abundant cytoskeletal proteins are removed, lower abundance proteins specific to the organelle are revealed. Here, we review recent literature on the global analysis of subcellular organelles and briefly discuss how that information is being used to elucidate basic biological processes that range from cellular signaling pathways through protein-protein interactions to differential expression of proteins in response to external stimuli. We assess the relative merits of the different methods used and discuss issues and future directions in the field.

Amino Acid Sequence↗

Reduced-median-network analysis of complete mitochondrial DNA coding-region sequences for the major African, Asian, and European haplogroups.

The evolution of the human mitochondrial genome is characterized by the emergence of ethnically distinct lineages or haplogroups. Nine European, seven Asian (including Native American), and three African mitochondrial DNA (mtDNA) haplogroups have been identified previously on the basis of the presence or absence of a relatively small number of restriction-enzyme recognition sites or on the basis of nucleotide sequences of the D-loop region. We have used reduced-median-network approaches to analyze 560 complete European, Asian, and African mtDNA coding-region sequences from unrelated individuals to develop a more complete understanding of sequence diversity both within and between haplogroups. A total of 497 haplogroup-associated polymorphisms were identified, 323 (65%) of which were associated with one haplogroup and 174 (35%) of which were associated with two or more haplogroups. Approximately one-half of these polymorphisms are reported for the first time here. Our results confirm and substantially extend the phylogenetic relationships among mitochondrial genomes described elsewhere from the major human ethnic groups. Another important result is that there were numerous instances both of parallel mutations at the same site and of reversion (i.e., homoplasy). It is likely that homoplasy in the coding region will confound evolutionary analysis of small sequence sets. By a linkage-disequilibrium approach, additional evidence for the absence of human mtDNA recombination is presented here.

Africa↗

Identification of protein associations in organelles, using mass spectrometry-based proteomics.

Recent literature that highlights the power of using mass spectrometry (MS) for protein identification from preparations of highly purified organelles and other large subcellular structures is covered in this review with an emphasis on techniques that preserve the integrity of the functional protein complexes. Recent advances in distinguishing contaminant proteins from "bonafide" organelle-localized proteins and the affinity capture of protein complexes are reviewed, as well as bioinformatic strategies to predict protein organellar localization and to integrate protein-protein interaction maps obtained from MS-affinity capture methods with data obtained from other techniques. Those developments demonstrate that a revolution in cellular biology, fueled by technical advances in MS-based proteomic techniques, is well underway.

Animals↗

An alternative strategy to determine the mitochondrial proteome using sucrose gradient fractionation and 1D PAGE on highly purified human heart mitochondria.

An alternative strategy for mitochondrial proteomics is described that is complementary to previous investigations using 2D PAGE techniques. The strategy involves (a) obtaining highly purified preparations of human heart mitochondria using metrizamide gradients to remove cytosolic and other subcellular contaminant proteins; (b) separation of mitochondrial protein complexes using sucrose density gradients after solubilization with n-dodecyl-beta-D-maltoside; (c) 1D electrophoresis of the sucrose gradient fractions; (d) high-throughput proteomics using robotic gel band excision, in-gel digestion, MALDI target spotting and automated spectral acquisition; and (e) protein identification from mixtures of tryptic peptides by high-precision peptide mass fingerprinting. Using this approach, we rapidly identified 82 bona fide or potential mitochondrial proteins, 40 of which have not been previously reported using 2D PAGE techniques. These proteins include small complex I and complex IV subunits, as well as very basic and hydrophobic transmembrane proteins such as the adenine nucleotide translocase that are not recovered in 2D gels. The technique described here should also be useful for the identification of new protein-protein associations as exemplified by the validation of a recently discovered complex that involves proteins belonging to the prohibitin family.

Blotting, Western↗

Expanded coverage of the human heart mitochondrial proteome using multidimensional liquid chromatography coupled with tandem mass spectrometry.

Recent evidence suggests that mitochondria are closely linked with the aging process and degenerative disorders such as Alzheimer's disease and Parkinson's disease. Thus, there has been increasing interest in cataloging mitochondrial proteomes to identify potential diagnostic and therapeutic targets. We have previously reported results of a one-dimensional electrophoresis/liquid chromatography MS/MS study to characterize the proteome of normal human heart mitochondria (Taylor et al. Nat. Biotechnol. 2003, 21, 281-286). We now report two subsequent studies where multidimensional liquid chromatography MS/MS was investigated as an alternative means for characterizing the same sample.

Chromatography, Liquid↗