Search PubMed⌕ Search

Biomedical subjects

R Scott Winters

Publications and source records attributed to R Scott Winters.

4 recordsLinked to original sources

An automated procedure to identify biomedical articles that contain cancer-associated gene variants.

The proliferation of biomedical literature makes it increasingly difficult for researchers to find and manage relevant information. However, identifying research articles containing mutation data, a requisite first step in integrating large and complex mutation data sets, is currently tedious, time-consuming and imprecise. More effective mechanisms for identifying articles containing mutation information would be beneficial both for the curation of mutation databases and for individual researchers. We developed an automated method that uses information extraction, classifier, and relevance ranking techniques to determine the likelihood of MEDLINE abstracts containing information regarding genomic variation data suitable for inclusion in mutation databases. We targeted the CDKN2A (p16) gene and the procedure for document identification currently used by CDKN2A Database curators as a measure of feasibility. A set of abstracts was manually identified from a MEDLINE search as potentially containing specific CDKN2A mutation events. A subset of these abstracts was used as a training set for a maximum entropy classifier to identify text features distinguishing "relevant" from "not relevant" abstracts. Each document was represented as a set of indicative word, word pair, and entity tagger-derived genomic variation features. When applied to a test set of 200 candidate abstracts, the classifier predicted 88 articles as being relevant; of these, 29 of 32 manuscripts in which manual curation found CDKN2A sequence variants were positively predicted. Thus, the set of potentially useful articles that a manual curator would have to review was reduced by 56%, maintaining 91% recall (sensitivity) and more than doubling precision (positive predictive value). Subsequent expansion of the training set to 494 articles yielded similar precision and recall rates, and comparison of the original and expanded trials demonstrated that the average precision improved with the larger data set. Our results show that automated systems can effectively identify article subsets relevant to a given task and may prove to be powerful tools for the broader research community. This procedure can be readily adapted to any or all genes, organisms, or sets of documents.

Computational Biology↗

An entity tagger for recognizing acquired genomic variations in cancer literature.

VTag is an application for identifying the type, genomic location and genomic state-change of acquired genomic aberrations described in text. The application uses a machine learning technique called conditional random fields. VTag was tested with 345 training and 200 evaluation documents pertaining to cancer genetics. Our experiments resulted in 0.8541 precision, 0.7870 recall and 0.8192 F-measure on the evaluation set.

Abstracting and Indexing↗

Fine mapping of the Schnyder's crystalline corneal dystrophy locus.

Schnyder's crystalline corneal dystrophy (SCCD) is a rare autosomal dominant eye disease with a spectrum of clinical manifestations that may include bilateral corneal clouding, arcus lipoides, and anterior corneal crystalline cholesterol deposition. We have previously performed a genome-wide linkage analysis on two large Swede-Finn families and mapped the SCCD locus to a 16-cM interval between markers D1S2633 and D1S228 on chromosome 1p36. We have collected 11 additional families from Finland, Germany, Turkey, and USA to narrow the critical region for SCCD. Here, we have used haplotype analysis with densely spaced microsatellite markers in a total of 13 families to refine the candidate interval. A common disease haplotype was observed among the four Swede-Finn families indicating the presence of a founder effect. Recombination results from all 13 families refined the SCCD locus to 2.32 Mbp between markers D1S1160 and D1S1635. Within this interval, identity-by-state was present in all 13 families for two markers D1S244 and D1S3153, further refining the candidate region to 1.58 Mbp.

Chromosome Mapping↗

me-PCR: a refined ultrafast algorithm for identifying sequence-defined genomic elements.

We have adapted the originally described electronic PCR (e-PCR) algorithm to perform string searches more accurately and much more rapidly than previously possible. Our implementation [multithreaded e-PCR (me-PCR)] runs sufficiently fast to allow even desktop machines to query quickly large genomes with very large genomic element sets. In addition, me-PCR is multithreaded, interprets all IUPAC nucleotide symbols, allows searches with elements specified by long sequences (such as SNPs), accepts ranges in the expected PCR size input field, requires substantially less memory for analysis of large sequences and corrects a number of minor flaws causing misreporting of hits in exceptional cases. Thus, me-PCR provides increased annotation capabilities for complex genomes to non-expert laboratories.

Algorithms↗