Search PubMed⌕ Search

Biomedical subjects

Carsten Peterson

Publications and source records attributed to Carsten Peterson.

8 recordsLinked to original sources

Potential for dramatic improvement in sequence alignment against structures of remote homologous proteins by extracting structural information from multiple structure alignment.

A novel method has been developed for acquiring the correct alignment of a query sequence against remotely homologous proteins by extracting structural information from profiles of multiple structure alignment. A systematic search algorithm combined with a group of score functions based on sequence information and structural information has been introduced in this procedure. A limited number of top solutions (15,000) with high scores were selected as candidates for further examination. On a test-set comprising 301 proteins from 75 protein families with sequence identity less than 30%, the proportion of proteins with completely correct alignment as first candidate was improved to 39.8% by our method, whereas the typical performance of existing sequence-based alignment methods was only between 16.1% and 22.7%. Furthermore, multiple candidates for possible alignment were provided in our approach, which dramatically increased the possibility of finding correct alignment, such that completely correct alignments were found amongst the top-ranked 1000 candidates in 88.3% of the proteins. With the assistance of a sequence database, completely correct alignment solutions were achieved amongst the top 1000 candidates in 94.3% of the proteins. From such a limited number of candidates, it would become possible to identify more correct alignment using a more time-consuming but more powerful method with more detailed structural information, such as side-chain packing and energy minimization, etc. The results indicate that the novel alignment strategy could be helpful for extending the application of highly reliable methods for fold identification and homology modeling to a huge number of homologous proteins of low sequence similarity. Details of the methods, together with the results and implications for future development are presented.

Algorithms↗

Analyzing tumor gene expression profiles.

A brief introduction to high throughput technologies for measuring and analyzing gene expression is given. Various supervised and unsupervised data mining methods for analyzing the produced high-dimensional data are discussed. The main emphasis is on supervised machine learning methods for classification and prediction of tumor gene expression profiles. Furthermore, methods to rank the genes according to their importance for the classification are explored. The approaches are illustrated by exploratory studies using two examples of retrospective clinical data from routine tests; diagnostic prediction of small round blue cell tumors (SRBCT) of childhood and determining the estrogen receptor (ER) status of sporadic breast cancer. The classification performance is gauged using blind tests. These studies demonstrate the feasibility of machine learning-based molecular cancer classification.

Adult↗

RNA analysis of B cell lines arrested at defined stages of differentiation allows for an approximation of gene expression patterns during B cell development.

The development of a mature B lymphocyte from a bone marrow stem cell is a highly ordered process involving stages with defined features and gene expression patterns. To obtain a deeper understanding of the molecular genetics of this process, we have performed RNA expression analysis of a set of mouse B lineage cell lines representing defined stages of B cell development using Affymetrix microarrays. The cells were grouped based on their previously defined phenotypic features, and a gene expression pattern for each group of cell lines was established. The data indicated that the cell lines representing a defined stage generally presented a high similarity in overall expression profiles. Numerous genes could be identified as expressed with a restricted pattern using dCHIP-based, quantitative comparisons or presence/absence-based, probabilistic state analysis. These experiments provide a model for gene expression during B cell development, and the correctly identified expression patterns of a number of control genes suggest that a series of cell lines can be useful tools in the elucidation of the molecular genetics of a complex differentiation process.

Animals↗

Microarray-based cancer diagnosis with artificial neural networks.

In recent years, the advent of experimental methods to probe gene expression profiles of cancer on a genome-wide scale has led to widespread use of supervised machine learning algorithms to characterize these profiles. The main applications of these analysis methods range from assigning functional classes of previously uncharacterized genes to classification and prediction of different cancer tissues. This article surveys the application of machine learning algorithms to classification and diagnosis of cancer based on expression profiles. To exemplify the important issues of the classification procedure, the emphasis of this article is on one such method, namely artificial neural networks. In addition, methods to extract genes that are important for the performance of a classifier, as well as the influence of sample selection on prediction results are discussed.

Algorithms↗

Expression profiling to predict outcome in breast cancer: the influence of sample selection.

Gene expression profiling of tumors using DNA microarrays is a promising method for predicting prognosis and treatment response in cancer patients. It was recently reported that expression profiles of sporadic breast cancers could be used to predict disease recurrence better than currently available clinical and histopathological prognostic factors. Having observed an overlap in those data between the genes that predict outcome and those that predict estrogen receptor-alpha status, we examined their predictive power in an independent data set. We conclude that it may be important to define prognostic expression profiles separately for estrogen receptor-alpha-positive and estrogen receptor-alpha-negative tumors.

Breast Neoplasms↗

BioArray Software Environment (BASE): a platform for comprehensive management and analysis of microarray data.

The microarray technique requires the organization and analysis of vast amounts of data. These data include information about the samples hybridized, the hybridization images and their extracted data matrices, and information about the physical array, the features and reporter molecules. We present a web-based customizable bioinformatics solution called BioArray Software Environment (BASE) for the management and analysis of all areas of microarray experimentation. All software necessary to run a local server is freely available.

Database Management Systems↗

Analyzing array data using supervised methods.

Pharmacogenomics is the application of genomic technologies to drug discovery and development, as well as for the elucidation of the mechanisms of drug action on cells and organisms. DNA microarrays measure genome-wide gene expression patterns and are an important tool for pharmacogenomic applications, such as the identification of molecular targets for drugs, toxicological studies and molecular diagnostics. Genome-wide investigations generate vast amounts of data and there is a need for computational methods to manage and analyze this information. Recently, several supervised methods, in which other information is utilized together with gene expression data, have been used to characterize genes and samples. The choice of analysis methods will influence the results and their interpretation, therefore it is important to be familiar with each method, its scope and limitations. Here, methods with special reference to applications for pharmacogenomics are reviewed.

Artificial Intelligence↗