Search PubMed⌕ Search

Biomedical subjects

Brian P Suomela

Publications and source records attributed to Brian P Suomela.

3 recordsLinked to original sources

Large-scale genomic correlations in Arabidopsis thaliana relate to chromosomal structure.

BACKGROUND: The chromosomes of the plant Arabidopsis thaliana contain various genomic elements, distributed with appreciable spatial heterogeneity. Clustering of and/or correlations between these elements presumably should reflect underlying functional or structural factors. We studied the positional density fluctuations and correlations between genes, indels, single nucleotide polymorphisms (SNPs), retrotransposons, 180 bp tandem repeats, and conserved centromeric sequences (CCSs) in Arabidopsis in order to elucidate any patterns and possible responsible factors for their genomic distributions. RESULTS: The spatial distributions of all these elements obeyed a common pattern: the density profiles of each element within chromosomes exhibited low-frequency fluctuations indicative of regional clustering, and the individual density profiles tended to correlate with each other at large measurement scales. This pattern could be attributed to the influence of major chromosomal structures, such as centromeres. At smaller scales the correlations tended to weaken -- evidence that localized cis-interactions between the different elements had a comparatively minor, if any, influence on their placement. CONCLUSION: The conventional notion that retrotransposon insertion sites are strongly influenced by cis-interactions was not supported by these observations. Moreover, we would propose that large-scale chromosomal structure has a dominant influence on the intrachromosomal distributions of genomic elements, and provides for an additional shared hierarchy of genomic organization within Arabidopsis.

Arabidopsis↗

Ranking the whole MEDLINE database according to a large training set using text indexing.

BACKGROUND: The MEDLINE database contains over 12 million references to scientific literature, with about 3/4 of recent articles including an abstract of the publication. Retrieval of entries using queries with keywords is useful for human users that need to obtain small selections. However, particular analyses of the literature or database developments may need the complete ranking of all the references in the MEDLINE database as to their relevance to a topic of interest. This report describes a method that does this ranking using the differences in word content between MEDLINE entries related to a topic and the whole of MEDLINE, in a computational time appropriate for an article search query engine. RESULTS: We tested the capabilities of our system to retrieve MEDLINE references which are relevant to the subject of stem cells. We took advantage of the existing annotation of references with terms from the MeSH hierarchical vocabulary (Medical Subject Headings, developed at the National Library of Medicine). A training set of 81,416 references was constructed by selecting entries annotated with the MeSH term stem cells or some child in its sub tree. Frequencies of all nouns, verbs, and adjectives in the training set were computed and the ratios of word frequencies in the training set to those in the entire MEDLINE were used to score references. Self-consistency of the algorithm, benchmarked with a test set containing the training set and an equal number of references randomly selected from MEDLINE was better using nouns (79%) than adjectives (73%) or verbs (70%). The evaluation of the system with 6,923 references not used for training, containing 204 articles relevant to stem cells according to a human expert, indicated a recall of 65% for a precision of 65%. CONCLUSION: This strategy appears to be useful for predicting the relevance of MEDLINE references to a given concept. The method is simple and can be used with any user-defined training set. Choice of the part of speech of the words used for classification has important effects on performance. Lists of words, scripts, and additional information are available from the web address http://www.ogic.ca/projects/ks2004/.

Abstracting and Indexing↗

Study of stem cell function using microarray experiments.

DNA Microarrays are used to simultaneously measure the levels of thousands of mRNAs in a sample. We illustrate here that a collection of such measurements in different cell types and states is a sound source of functional predictions, provided the microarray experiments are analogous and the cell samples are appropriately diverse. We have used this approach to study stem cells, whose identity and mechanisms of control are not well understood, generating Affymetrix microarray data from more than 200 samples, including stem cells and their derivatives, from human and mouse. The data can be accessed online (StemBase; http://www.scgp.ca:8080/StemBase/).

Animals↗