Search PubMed⌕ Search

Biomedical subjects

James O McInerney

Publications and source records attributed to James O McInerney.

At least 19 recordsLinked to original sources

PanForest: predicting genes in genomes using random forests.

MOTIVATION: The presence or absence of some genes in a genome can influence whether other genes are likely to be present or absent. Understanding these gene co-occurrence and avoidance patterns reveals fundamental principles of genome organization, with applications ranging from evolutionary reconstruction to rational design of synthetic genomes. RESULTS: PanForest, presented here, uses random forest classifiers to predict the presence and absence of genes in genomes from the set of other genes present. Performance statistics output by PanForest reveal how predictable each gene's presence or absence is, based on the presence or absence of other genes in the genome. Further, PanForest produces statistics indicating the importance of each gene in predicting the presence or absence of each other gene. The PanForest software can run serially or in parallel, thereby facilitating the analysis of pangenomes at Network of Life scale.A pangenome of 12 741 accessory genes in 1000 Escherichia coli genomes was analysed in around 5 h using eight processors. To demonstrate PanForest's utility, we present a case study and show that certain genes associated with resistance to antimicrobial drugs reliably predict the presence or absence of other genes associated with resistance to the same drug. Further, we highlight several associations between those genes and others not known to be associated with antimicrobial resistance (AMR), or associated with resistance to other drugs. We envisage PanForest's use in studies from multiple disciplines concerning the dynamics of gene distributions in pangenomes ranging from biomedical science and synthetic biology to molecular ecology. AVAILABILITY AND IMPLEMENTATION: The software if freely available with a full manual and can be found with at www.github.com/alanbeavan/PanForest DOI: https://doi.org/10.5281/zenodo.17865482.

Software↗

The molecular phylogeny of a nematode-specific clade of heterotrimeric G-protein alpha-subunit genes.

In animal olfactory systems, odorant molecules are detected by olfactory receptors (ORs). ORs are part of the G-protein-coupled receptor (GPCR) superfamily. Heterotrimeric guanine nucleotide binding G-proteins (G-proteins) relay signals from GPCRs to intracellular effectors. G-proteins are comprised of three peptides. The G-protein alpha subunit confers functional specificity to G-proteins. Vertebrate and insect Galpha-subunit genes are divided into four subfamilies based on functional and sequence attributes. The nematode Caenorhabditis elegans contains 21 Galpha genes, 14 of which are exclusively expressed in sensory neurons. Most individual mammalian cells express multiple distinct GPCR gene products, however, individual mammalian and insect olfactory neurons express only one functional odorant OR. By contrast C. elegans expresses multiple ORs and multiple Galpha subunits within each olfactory neuron. Here we show that, in addition to having at least one member of each of the four mammalian Galpha gene classes, C. elegans and other nematodes also possess two lineage-specific Galpha gene expansions, homologues of which are not found in any other organisms examined. We hypothesize that these novel nematode-specific Galpha genes increase the functional complexity of individual chemosensory neurons, enabling them to integrate odor signals from the multiple distinct ORs expressed on their membranes. This neuronal gene expansion most likely occurred in nematodes to enable them to compensate for the small number of chemosensory cells and the limited emphasis on cephalization during nematode evolution.

Animals↗

The causes of protein evolutionary rate variation.

The rate of protein evolution varies more than 1000-fold and, for the past 30 years, it was thought that the rate was determined by protein function. Drummond and co-workers have now shown that a single factor underlying mRNA expression, protein abundance and synonymous codon usage is the chief causal agent of protein evolutionary rate in yeast. It will be interesting to see whether this is shown to be a universal rule for all biological systems.

Evolution, Molecular↗

Gene evolution and drug discovery.

Mutation and selection are the principle forces governing gene and protein sequence. Mutation is the major source of variation, and selection removes variation. Although many mutations are likely to be neutral with respect to natural selection, much of the extant sequence that is functionally important has experienced selective pressures in the past. By examining the history of DNA sequences, we can infer the functional importance of particular residues and the selective pressures that have influenced their evolution. In this chapter, we review the most interesting approaches for inferring the evolutionary history of DNA and protein sequences and indicate how these analyses can be useful in the drug discovery process.

Animals↗

Gamma chain receptor interleukins: evidence for positive selection driving the evolution of cell-to-cell communicators in the mammalian immune system.

The interleukin-2 receptor (IL-2R) gamma chain, or common gamma chain (gammac), is the hub of a protein interaction network in the mammalia that is central to defense against disease. It is the indispensable subunit of the functional receptor complexes for a group of interleukins known as the gamma-chain-dependent interleukins (IL-2, IL-4, -7, -9, -15, and -21). The gammac links these proteins through their interaction with it and their competition for its recruitment. The gammac-dependent interleukins also interact with each other to either enhance or suppress expression through manipulation of expression of receptor subunits. Given the influence of protein-protein interactions on evolution, such as those documented for many genes including the reproductive proteins of the sperm and egg coat, here we have asked whether there is a common thread in the evolution of these interleukins. Our findings indicate that positive selection has acted by fixing a large number of amino acid replacement mutations in every single one of these interleukins, this adaptive evolution is also observed in a lineage-specific manner. Crucially, however, there does not appear to have ever been an instance of adaptive evolution in the gammac chain itself, thereby providing an insight into the evolution of this hub protein. These findings highlight the importance of adaptive evolutionary events in the evolution of this central network in the immune system and suggest underlying causes for differences in defense responses in the mammalia.

Animals↗

Adaptive evolution of the human fatty acid synthase gene: support for the cancer selection and fat utilization hypotheses?

Cancer may act as the etiological agent for natural selection in some genes. This selective pressure would act to reduce the success of neoplastic lineages over normal cell lineages in individuals of reproductive age. In addition, human's relatively larger brain and longer lifespan may have also acted as a selective force requiring new genotypes. One of the most important proteins in both processes is the fatty acid synthase (FAS) gene involved in fatty acid biosynthesis. A variety of other proteins, including PTEN, MAPK1, SREBP1, SREBP2 and PI are also involved in the regulation of fatty acid biosynthesis. We have specifically analysed variability in selective pressure across all these genes in human, mouse and other vertebrates. We have found that the FAS gene alone has signatures indicative of adaptive evolution. We did not find any signatures of adaptive evolution in any of the other proteins. In the FAS gene, we have detected an excess of non-synonymous over synonymous substitutions in approximately 6% of sites in the human lineage. Contrastingly, the substitution process at these sites in other available vertebrates and mammals indicates strong purifying selection. This is likely to reflect a functional shift in human FAS and correlates well with previously observed changes in FAS biochemical activities. We speculate that the role played by FAS either in cancer development or in human brain development has created this selective pressure, although we cannot rule out the various other functions of FAS.

Biological Evolution↗

Genome phylogenies indicate a meaningful alpha-proteobacterial phylogeny and support a grouping of the mitochondria with the Rickettsiales.

Placement of the mitochondrial branch on the tree of life has been problematic. Sparse sampling, the uncertainty of how lateral gene transfer might overwrite phylogenetic signals, and the uncertainty of phylogenetic inference have all contributed to the issue. Here we address this issue using a supertree approach and completed genomic sequences. We first determine that a sensible alpha-proteobacterial phylogenetic tree exists and that it can confidently be inferred using orthologous genes. We show that congruence across these orthologous gene trees is significantly better than might be expected by random chance. There is some evidence of horizontal gene transfer within the alpha-proteobacteria, but it appears to be restricted to a minority of genes ( approximately 23%) most of whom ( approximately 74%) can be categorized as operational. This means that placement of the mitochondrion should not be excessively hampered by interspecies gene transfer. We then show that there is a consistently strong signal for placement of the mitochondrion on this tree and that this placement is relatively insensitive to methodological approach or data set. A concatenated alignment was created consisting of 15 mitochondrion-encoded proteins that are unlikely to have undergone any lateral gene transfer in the timeline under consideration. This alignment infers that the sister group of the mitochondria, for the taxa that have been sampled, is the order Rickettsiales.

Alphaproteobacteria↗

Evidence of positive Darwinian selection in putative meningococcal vaccine antigens.

Meningococcal meningitidis is a life-threatening disease. In Europe and the United States the majority of cases are caused by virulent meningococcal strains belonging to serogroup B. Presently there is no effective vaccine against serogroup B strains, as traditional vaccine antigens such as polysaccharide capsules are unusable as they lead to autoimmunity. The year 2000 saw the publication of the complete genome of Neisseria meningitidis MC58, a virulent serogroup B bacterium. Working in conjunction with the sequencing project, researchers endeavored to locate highly conserved membrane-associated proteins that elicit an immune response. It is hoped that these proteins will provide a basis for novel vaccines against serogroup B strains. A number of potential vaccine antigens have been located and are presently in phase I clinical trials. Recently many reports pertaining to the evidence of positive Darwinian selection in membrane proteins of pathogens have been reported. This study utilized in silico methods to test for evidence of historical positive Darwinian selection in seven such vaccine candidates. We found that two of these proteins show signatures of adaptive evolution, while the remaining proteins show evidence of strong purifying selection. This has significant implications for the design of a vaccine against serogroup B strains, as it has been shown that vaccines that target epitopes that are under strong purifying selection are better than those that target variable epitopes.

Amino Acid Sequence↗

The Opisthokonta and the Ecdysozoa may not be clades: stronger support for the grouping of plant and animal than for animal and fungi and stronger support for the Coelomata than Ecdysozoa.

In considering the best possible solutions for answering phylogenetic questions from genomic sequences, we have chosen a strategy that we suggest is superior to others that have gone previously. We have ignored multigene families and instead have used single-gene families. This minimizes the inadvertent analysis of paralogs. We have employed strict data controls and have reasoned that if a protein is not capable of recovering the uncontroversial parts of a phylogenetic tree, then why should we use it for the more controversial parts? We have sliced and diced the data in as many ways as possible in order to uncover the signals in that data. Using this strategy, we have tested two controversial hypotheses concerning eukaryotic phylogenetic relationships: the placement of arthropoda and nematodes and the relationships of animals, plants, and fungi. We have constructed phylogenetic trees from 780 single-gene families from 10 completed genomes and amalgamated these into a single supertree. We have also carried out a total evidence analysis on the only universally distributed protein families that can accurately reconstruct the uncontroversial parts of the phylogenetic tree: a total of five families. In doing so, we ignore the majority of single-gene families that are universally distributed as they do not have the appropriate signals to recover the uncontroversial parts of the tree. We have also ignored every protein that has ever been used previously to address this issue, simply because none of them meet our strict criteria. Using these data controls, site stripping, and multiple analyses, 24 out of 26 analyses strongly support the grouping of vertebrates with arthropods (Coelomata hypothesis) and plants with animals. In the other two analyses, the data were ambivalent. The latter finding overturns an 11-year theory of Eukaryotic evolution; the first confirms what has already been said by others. In the light of this new tree, we re-analyze the evolution of intron gain and loss in the rpL14 gene and find that it is much more compatible with the hypothesis presented here than with the Opisthokonta hypothesis.

Animal Population Groups↗

Evidence of positive Darwinian selection in Omp85, a highly conserved bacterial outer membrane protein essential for cell viability.

Omp85 is a highly conserved outer membrane protein found in all gram-negative bacteria. It is essential for bacterial cell viability and plays an integral function in the positioning and folding of other outer membrane proteins into the bacterial outer membrane. We have employed a maximum likelihood and a maximum parsimony approach to detect evidence of positive Darwinian selection in Omp85 homologues from 10 delta-proteobacteria and have identified 14 amino acid sites that show evidence of being under the influence of adaptive evolution. Interestingly all sites bar one are concentrated within surface loops of the protein that most likely interact with host immune response or the surrounding environment. Alternatively amino acids within membrane-spanning regions of the protein are found to be under purifying selection most likely as a result of structural constraints.

Bacterial Outer Membrane Proteins↗

New methods ring changes for the tree of life.

Relationships among prokaryotes and the origin of eukaryotes have both proven controversial, with results depending upon the gene sequences and methods used. Extensive horizontal gene transfer is one possible reason why inferring such deep phylogenetic relationships is difficult. In two recent papers, Lake and Rivera introduce new methods that can be used to reconstruct the genomic tree in the presence of horizontal gene transfers, but which suggest that a ring rather than a tree is a better representation of some parts of the history of life on Earth.

Journal Article↗

The shape of supertrees to come: tree shape related properties of fourteen supertree methods.

Using a simple example and simulations, we explore the impact of input tree shape upon a broad range of supertree methods. We find that input tree shape can affect how conflict is resolved by several supertree methods and that input tree shape effects may be substantial. Standard and irreversible matrix representation with parsimony (MRP), MinFlip, duplication-only Gene Tree Parsimony (GTP), and an implementation of the average consensus method have a tendency to resolve conflict in favor of relationships in unbalanced trees. Purvis MRP and the average dendrogram method appear to have an opposite tendency. Biases with respect to tree shape are correlated with objective functions that are based upon unusual asymmetric tree-to-tree distance or fit measures. Split, quartet, and triplet fit, most similar supertree, and MinCut methods (provided the latter are interpreted as Adams consensus-like supertrees) are not revealed to have any bias with respect to tree shape by our example, but whether this holds more generally is an open problem. Future development and evaluation of supertree methods should consider explicitly the undesirable biases and other properties that we highlight. In the meantime, use of a single, arbitrarily chosen supertree method is discouraged. Use of multiple methods and/or weighting schemes may allow practical assessment of the extent to which inferences from real data depend upon methodological biases with respect to input tree shape or size.

Classification↗

Evidence for heterogeneous selective pressures in the evolution of the env gene in different human immunodeficiency virus type 1 subtypes.

Recent studies have demonstrated the emergence of human immunodeficiency virus type 1 (HIV-1) subtypes with various levels of fitness. Using heterogeneous maximum-likelihood models of adaptive evolution implemented in the PAML software package, with env sequences representing each HIV-1 group M subtype, we examined the various intersubtype selective pressures operating across the env gene. We found heterogeneity of evolutionary mechanisms between the different subtypes with a category of amino acid sites observed that had undergone positive selection for subtypes C, F1, and G, while these sites had undergone purifying selection in all other subtypes. Also, amino acid sites within subtypes A and K that had undergone purifying selection were observed, while these sites had undergone positive selection in all other subtypes. The presence of such sites indicates heterogeneity of selective pressures within HIV-1 group M subtype evolution that may account for the various levels of fitness of the subtypes.

Amino Acid Sequence↗

Does a tree-like phylogeny only exist at the tips in the prokaryotes?

The extent to which prokaryotic evolution has been influenced by horizontal gene transfer (HGT) and therefore might be more of a network than a tree is unclear. Here we use supertree methods to ask whether a definitive prokaryotic phylogenetic tree exists and whether it can be confidently inferred using orthologous genes. We analysed an 11-taxon dataset spanning the deepest divisions of prokaryotic relationships, a 10-taxon dataset spanning the relatively recent gamma-proteobacteria and a 61-taxon dataset spanning both, using species for which complete genomes are available. Congruence among gene trees spanning deep relationships is not better than random. By contrast, a strong, almost perfect phylogenetic signal exists in gamma-proteobacterial genes. Deep-level prokaryotic relationships are difficult to infer because of signal erosion, systematic bias, hidden paralogy and/or HGT. Our results do not preclude levels of HGT that would be inconsistent with the notion of a prokaryotic phylogeny. This approach will help decide the extent to which we can say that there is a prokaryotic phylogeny and where in the phylogeny a cohesive genomic signal exists.

Bacteria↗

Analysis of gene expression in the bovine corpus luteum through generation and characterisation of 960 ESTs.

To gain new insights into gene identity and gene expression in the bovine corpus luteum (CL) a directionally cloned CL cDNA library was constructed, screened with a total CL cDNA probe and clones representing abundant and rare mRNA transcripts isolated. The 5'-terminal DNA sequence of 960 cDNA clones, composed of 192 abundant and 768 rare mRNA transcripts was determined and clustered into 351 non-redundant expressed sequence tag (EST) groups. Bioinformatic analysis revealed that 309 (88%) of the ESTs showed significant homology to existing sequences in the protein and nucleotide public databases. Several previously unidentified bovine genes encoding proteins associated with key aspects of CL function including extracellular matrix remodelling, lipid metabolism/steroid biosynthesis and apoptosis, were identified. Forty-two (12%) of the ESTs showed homology with human or with other uncharacterised ESTs, some of these were abundantly expressed and may therefore play an important role in primary CL function. Tissue-specificity and temporal CL gene expression of selected clones previously unidentified in bovine CL tissue was also examined. The most interesting finds indicated that mRNA encoding squalene epoxidase was constitutively expressed in CL tissue throughout the oestrous cycle and 7-fold down-regulated (P < 0.05) in late luteal tissue, concomitant with the disappearance of systemic progesterone, suggesting that de novo cholesterol biosynthesis plays an important role in steroidogenesis. The mRNA encoding the growth factor, insulin-like growth factor-binding protein-related protein 1 (IGFBP-rP1), remained constant during the oestrous cycle and was 1.8-fold up-regulated (P < 0.05) in late luteal tissue implying a role in CL regression.

Animals↗

Gene prediction using the Self-Organizing Map: automatic generation of multiple gene models.

BACKGROUND: Many current gene prediction methods use only one model to represent protein-coding regions in a genome, and so are less likely to predict the location of genes that have an atypical sequence composition. It is likely that future improvements in gene finding will involve the development of methods that can adequately deal with intra-genomic compositional variation. RESULTS: This work explores a new approach to gene-prediction, based on the Self-Organizing Map, which has the ability to automatically identify multiple gene models within a genome. The current implementation, named RescueNet, uses relative synonymous codon usage as the indicator of protein-coding potential. CONCLUSIONS: While its raw accuracy rate can be less than other methods, RescueNet consistently identifies some genes that other methods do not, and should therefore be of interest to gene-prediction software developers and genome annotation teams alike. RescueNet is recommended for use in conjunction with, or as a complement to, other gene prediction methods.

Chromosome Mapping↗