Search PubMed⌕ Search

Biomedical subjects

J J Rowland

Publications and source records attributed to J J Rowland.

10 recordsLinked to original sources

Model selection methodology in supervised learning with evolutionary computation.

The expressive power, powerful search capability, and the explicit nature of the resulting models make evolutionary methods very attractive for supervised learning applications in bioinformatics. However, their characteristics also make them highly susceptible to overtraining or to discovering chance relationships in the data. Identification of appropriate criteria for terminating evolution and for selecting an appropriately validated model is vital. Some approaches that are commonly applied to other modelling methods are not necessarily applicable in a straightforward manner to evolutionary methods. An approach to model selection is presented that is not unduly computationally intensive. To illustrate the issues and the technique two bioinformatic datasets are used, one relating to metabolite determination and the other to disease prediction from gene expression data.

Algorithms↗

Primary and secondary metabolism, and post-translational protein modifications, as portrayed by proteomic analysis of Streptomyces coelicolor.

The newly sequenced genome of Streptomyces coelicolor is estimated to encode 7825 theoretical proteins. We have mapped approximately 10% of the theoretical proteome experimentally using two-dimensional gel electrophoresis and matrix-assisted laser desorption ionization time-of-flight (MALDI-TOF) mass spectrometry. Products from 770 different genes were identified, and the types of proteins represented are discussed in terms of their annotated functional classes. An average of 1.2 proteins per gene was observed, indicating extensive post-translational regulation. Examples of modification by N-acetylation, adenylylation and proteolytic processing were characterized using mass spectrometry. Proteins from both primary and certain secondary metabolic pathways are strongly represented on the map, and a number of these enzymes were identified at more than one two-dimensional gel location. Post-translational modification mechanisms may therefore play a significant role in the regulation of these pathways. Unexpectedly, one of the enzymes for synthesis of the actinorhodin polyketide antibiotic appears to be located outside the cytoplasmic compartment, within the cell wall matrix. Of 20 gene clusters encoding enzymes characteristic of secondary metabolism, eight are represented on the proteome map, including three that specify the production of novel metabolites. This information will be valuable in the characterization of the new metabolites.

Acetylation↗

A functional genomics strategy that uses metabolome data to reveal the phenotype of silent mutations.

A large proportion of the 6,000 genes present in the genome of Saccharomyces cerevisiae, and of those sequenced in other organisms, encode proteins of unknown function. Many of these genes are "silent, " that is, they show no overt phenotype, in terms of growth rate or other fluxes, when they are deleted from the genome. We demonstrate how the intracellular concentrations of metabolites can reveal phenotypes for proteins active in metabolic regulation. Quantification of the change of several metabolite concentrations relative to the concentration change of one selected metabolite can reveal the site of action, in the metabolic network, of a silent gene. In the same way, comprehensive analyses of metabolite concentrations in mutants, providing "metabolic snapshots," can reveal functions when snapshots from strains deleted for unstudied genes are compared to those deleted for known genes. This approach to functional analysis, using comparative metabolomics, we call FANCY-an abbreviation for functional analysis by co-responses in yeast.

Adenine Nucleotides↗

Rapid analysis of high-dimensional bioprocesses using multivariate spectroscopies and advanced chemometrics.

There are an increasing number of instrumental methods for obtaining data from biochemical processes, many of which now provide information on many (indeed many hundreds) of variables simultaneously. The wealth of data that these methods provide, however, is useless without the means to extract the required information. As instruments advance, and the quantity of data produced increases, the fields of bioinformatics and chemometrics have consequently grown greatly in importance. The chemometric methods nowadays available are both powerful and dangerous, and there are many issues to be considered when using statistical analyses on data for which there are numerous measurements (which often exceed the number of samples). It is not difficult to carry out statistical analysis on multivariate data in such a way that the results appear much more impressive than they really are. The authors present some of the methods that we have developed and exploited in Aberystwyth for gathering highly multivariate data from bioprocesses, and some techniques of sound multivariate statistical analyses (and of related methods based on neural and evolutionary computing) which can ensure that the results will stand up to the most rigorous scrutiny.

Algorithms↗

Genome-scale cloning and expression of individual open reading frames using topoisomerase I-mediated ligation.

The in vitro cloning of DNA molecules traditionally uses PCR amplification or site-specific restriction endonucleases to generate linear DNA inserts with defined termini and requires DNA ligase to covalently join those inserts to vectors with the corresponding ends. We have used the properties of Vaccinia DNA topoisomerase I to develop a ligase-free technology for the covalent joining of DNA fragments to suitable plasmid vectors. This system is much more efficient than cloning methods that require ligase because the rapid DNA rejoining activity of Vaccinia topoisomerase I allows ligation in only 5 min at room temperature, whereas the enzyme's high substrate specificity ensures a low rate of vector-alone transformants. We have used this topoisomerase I-mediated cloning technology to develop a process for accelerated cloning and expression of individual ORFs. Its suitability for genome-scale molecular cloning and expression is demonstrated in this report.

Animals↗

Quantification of microbial productivity via multi-angle light scattering and supervised learning.

This article describes the use of chemometric methods for prediction of biological parameters of cell suspensions on the basis of their light scattering profiles. Laser light is directed into a vial or flow cell containing media from the suspension. The intensity of the scattered light is recorded at 18 angles. Supervised learning methods are then used to calibrate a model relating the parameter of interest to the intensity values. Using such models opens up the possibility of estimating the biological properties of fermentor broths extremely rapidly (typically every 4 sec), and, using the flow cell, without user interaction. Our work has demonstrated the usefulness of this approach for estimation of yeast cell counts over a wide range of values (10(5)-10(9) cells mL-1), although it was less successful in predicting cell viability in such suspensions.

Biotechnology↗

The deconvolution of pyrolysis mass spectra using genetic programming: application to the identification of some Eubacterium species.

Pyrolysis mass spectrometry was used to produce complex biochemical fingerprints of Eubacterium exiguum, E. infirmum, E. tardum and E. timidum. To examine the relationship between these organisms the spectra were clustered by canonical variates analysis, and four clusters, one for each species, were observed. In an earlier study we trained artificial neural networks to identify these clinical isolates successfully; however, the information used by the neural network was not accessible from this so-called 'black box' technique. To allow the deconvolution of such complex spectra (in terms of which masses were important for discrimination) it was necessary to develop a system that itself produces 'rules' that are readily comprehensible. We here exploit the evolutionary computational technique of genetic programming; this rapidly and automatically produced simple mathematical functions that were also able to classify organisms to each of the four bacterial groups correctly and unambiguously. Since the rules used only a very limited set of masses, from a search space some 50 orders of magnitude greater than the dimensionality actually necessary, visual discrimination of the organisms on the basis of these spectral masses alone was also then possible.

Artificial Intelligence↗

Rapid identification of Streptococcus and Enterococcus species using diffuse reflectance-absorbance Fourier transform infrared spectroscopy and artificial neural networks.

Diffuse reflectance-absorbance Fourier transform infrared spectroscopy (FT-IR) was used to analyse 19 hospital isolates which had been identified by conventional means to one Enterococcus faecalis, E. faecium, Streptococcus bovis, S. mitis, S. pneumoniae, or S. pyogenes. Principal components analysis of the FT-IR spectra showed that this 'unsupervised' learning method failed to form six separable clusters (one of each species) and thus could not be used to identify these bacteria base on their FT-IR spectra. By contrast, artificial neural networks (ANNs) could be trained by 'supervised' learning (using the back-propagation algorithm) with the principal components scores of derivatised spectra to recognise the strains from their FT-IR spectra. These results demonstrate that the combination of FT-IR and ANNs provides a rapid, novel and accurate bacterial identification technique.

Algorithms↗

Two novel Streptomyces protein protease inhibitors. Purification, activity, cloning, and expression.

In contrast to the Gram-negative bacteria, Gram-positive bacteria such as Streptomyces lack a mucopolysaccharide cell wall which allows them to produce and secrete a variety of proteins directly into their environment. In an effort to understand and eventually exploit the synthesis and secretion of proteins by Streptomyces, we identified and characterized two naturally occurring abundantly produced proteins in culture supernatants of Streptomyces lividans and Streptomyces longisporus. We purified these 10-kDa proteins and obtained partial amino acid sequence information which was then used to design oligonucleotide probes in order to clone their genes. Analysis of the sequence data indicated that these proteins were related to each other and to several other previously characterized Streptomyces protein protease inhibitors. We demonstrate that both proteins are protein protease inhibitors with specificity for trypsin-like enzymes. The presumptive signal peptidase cleavage sites and subsequent aminopeptidase products of each protein are characterized. Finally, we show that the cloned genes contain all of the information necessary to direct synthesis and secretion of the proteins by Streptomyces spp. or Escherichia coli.

Amino Acid Sequence↗