Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “classifier”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Bayesian automatic relevance determination algorithms for classifying gene expression data.

MOTIVATION: We investigate two new Bayesian classification algorithms incorporating feature selection. These algorithms are applied to the classification of gene expression data derived from cDNA microarrays. RESULTS: We demonstrate the effectiveness of the algorithms on three gene expression datasets for cancer, showing they compare well with alternative kernel-based techniques. By automatically incorporating feature selection, accurate classifiers can be constructed utilizing very few features and with minimal hand-tuning. We argue that the feature selection is meaningful and some of the highlighted genes appear to be medically important.

Algorithms↗

Comparison of classic statistical methods and machine learning approaches to classify readiness.

MOTIVATION: Predicting physical and cognitive readiness in warfighters is critical for mission success. These predictions can be improved by identifying key biomarkers using multiple omics modalities. The MASTR-E study conducted by McKetney and colleagues is one of the most comprehensive multi-omics studies of saliva samples collected from warfighters, which also applied classic linear statistical (CLS) techniques to discover key biomarkers of readiness. Aligning with McKetney et al.'s assumptions, we operationalize readiness as a binary proxy, where pre-mission samples are labeled as "ready" to reflect a rested, unstressed physiological baseline, while post-mission samples are labeled "not ready" to reflect cumulative physical and cognitive load from the mission. As such, readiness here is not a direct biological or physiological construct, but an inferred state likely dominated by stress-related physiological changes. This assumption and definition is discussed further in the Introduction and Limitations sections. Here, we apply machine learning (ML) analyses to better assess generalizability, consider hidden interactions, and identify nonlinear patterns in the data. We investigated whether ML approaches could predict readiness and identify relevant biomarkers. ML models were trained on proteomics-only or metabolomics-only datasets to classify participants as ready or not ready and important model features were considered as putative biomarkers. Training and testing datasets were curated for two objectives: (i) recognize biomolecular signatures indicative of readiness within the same donor and (ii) assess generalizability across warfighters by withholding donors for testing. RESULTS: Proteomics-based models achieved AUCs of 0.907 ± 0.034 and 0.860 ± 0.063 for Objectives 1 and 2, respectively. Metabolomics-based models achieved Objective 1 AUC of 0.994 ± 0.007 and Objective 2 AUC of 0.993 ± 0.010. Comparative analysis with existing literature validates the model's feature importances, but the identified putative biomarkers significantly differ from those discovered through CLS analyses, as only one ML-identified biomarker overlapping with those identified through CLS methods. We show that these ML models and identified features are more robust to noise and generalizable across participants than those identified using CLS methods. AVAILABILITY: The analysis pipelines are provided as Jupyter notebooks, including all code and documentation, and are available publicly on GitHub at {https://github.com/netrias/ReadinessClassification}.

Machine Learning↗

Discovery of significant rules for classifying cancer diagnosis data.

METHODS AND RESULTS: We introduce a new method to discover many diversified and significant rules from high dimensional profiling data. We also propose to aggregate the discriminating power of these rules for reliable predictions. The discovered rules are found to contain low-ranked features; these features are found to be sometimes necessary for classifiers to achieve perfect accuracy. The use of low-ranked but essential features in our method is in contrast to the prevailing use of an ad-hoc number of only top-ranked features. On a wide range of data sets, our method displayed highly competitive accuracy compared to the best performance of other kinds of classification models. In addition to accuracy, our method also provides comprehensible rules to help elucidate the translation between raw data and useful knowledge.

Algorithms↗

Predicting subcellular localization of proteins using machine-learned classifiers.

MOTIVATION: Identifying the destination or localization of proteins is key to understanding their function and facilitating their purification. A number of existing computational prediction methods are based on sequence analysis. However, these methods are limited in scope, accuracy and most particularly breadth of coverage. Rather than using sequence information alone, we have explored the use of database text annotations from homologs and machine learning to substantially improve the prediction of subcellular location. RESULTS: We have constructed five machine-learning classifiers for predicting subcellular localization of proteins from animals, plants, fungi, Gram-negative bacteria and Gram-positive bacteria, which are 81% accurate for fungi and 92-94% accurate for the other four categories. These are the most accurate subcellular predictors across the widest set of organisms ever published. Our predictors are part of the Proteome Analyst web-service.

Algorithms↗

Standardization and denoising algorithms for mass spectra to classify whole-organism bacterial specimens.

MOTIVATION: Application of mass spectrometry in proteomics is a breakthrough in high-throughput analyses. Early applications have focused on protein expression profiles to differentiate among various types of tissue samples (e.g. normal versus tumor). Here our goal is to use mass spectra to differentiate bacterial species using whole-organism samples. The raw spectra are similar to spectra of tissue samples, raising some of the same statistical issues (e.g. non-uniform baselines and higher noise associated with higher baseline), but are substantially noisier. As a result, new preprocessing procedures are required before these spectra can be used for statistical classification. RESULTS: In this study, we introduce novel preprocessing steps that can be used with any mass spectra. These comprise a standardization step and a denoising step. The noise level for each spectrum is determined using only data from that spectrum. Only spectral features that exceed a threshold defined by the noise level are subsequently used for classification. Using this approach, we trained the Random Forest program to classify 240 mass spectra into four bacterial types. The method resulted in zero prediction errors in the training samples and in two test datasets having 240 and 300 spectra, respectively.

Algorithms↗

Autoregressive modeling of analytical sensor data can yield classifiers in the predictor coefficient parameter space.

SUMMARY: The analysis of chromatographic data resulting from complex chemical mixtures is challenging. Components may co-elute, causing their signals to overlap. An algorithm that will increase the signal-to-noise ratio so compounds present in low abundance can be better distinguished from noise is useful in this type of analysis. The autoregressive (AR) filter offers the advantage of smoothing chromatograms to increase this ratio, while also offering data compression and increased resolution. Furthermore, this filter can be useful for classification, as the roots of the predictor coefficient vectors represent features present in the data and can therefore be used for pattern recognition. In this paper, we present a novel method for applying AR filtering to chromatogram data. We show that the AR filter outperforms the Savitzky-Golay filter for smoothing noise while retaining important information within chromatograms, and also that AR correlation coefficients have the potential to be used to classify chromatogram data into groups. CONTACT: cdavis@draper.com.

Algorithms↗

ROCR: visualizing classifier performance in R.

UNLABELLED: ROCR is a package for evaluating and visualizing the performance of scoring classifiers in the statistical language R. It features over 25 performance measures that can be freely combined to create two-dimensional performance curves. Standard methods for investigating trade-offs between specific performance measures are available within a uniform framework, including receiver operating characteristic (ROC) graphs, precision/recall plots, lift charts and cost curves. ROCR integrates tightly with R's powerful graphics capabilities, thus allowing for highly adjustable plots. Being equipped with only three commands and reasonable default values for optional parameters, ROCR combines flexibility with ease of usage. AVAILABILITY: http://rocr.bioinf.mpi-sb.mpg.de. ROCR can be used under the terms of the GNU General Public License. Running within R, it is platform-independent. CONTACT: tobias.sing@mpi-sb.mpg.de.

Computer Graphics↗

ClaNC: point-and-click software for classifying microarrays to nearest centroids.

SUMMARY: ClaNC (classification to nearest centroids) is a simple and an accurate method for classifying microarrays. This document introduces a point-and-click interface to the ClaNC methodology. The software is available as an R package. AVAILABILITY: ClaNC is freely available from http://students.washington.edu/adabney/clanc

Algorithms↗

Automated classification of alternative splicing and transcriptional initiation and construction of visual database of classified patterns.

MOTIVATION: Large-scale detection and classification of alternative splicing and transcriptional initiation (ASTI) is the first step towards detailed studies of the functional implication and mechanisms of these phenomena. RESULTS: We have developed an algorithm that classifies all observed units of ASTI into an extendable set of distinct types (e.g. cassette type) by converting a collection of alignments between a genomic DNA sequence and cDNA sequences into binary description. This description system can uniquely and compactly encode not only typical patterns but also any rare patterns that are usually collectively assigned to 'others.' More than 150 distinct ASTI types were found when this system was applied to genome-wide detection of ASTI units in human and five other eukaryotes. AVAILABILITY: The data detected by this system are available through ASTRA (http://alterna.cbrc.jp/), a database equipped with a Java-based browser that can interactively reorganize the order of displayed splicing patterns on demand.

Algorithms↗

Combining multi-species genomic data for microRNA identification using a Naive Bayes classifier.

MOTIVATION: Most computational methodologies for microRNA gene prediction utilize techniques based on sequence conservation and/or structural similarity. In this study we describe a new technique, which is applicable across several species, for predicting miRNA genes. This technique is based on machine learning, using the Naive Bayes classifier. It automatically generates a model from the training data, which consists of sequence and structure information of known miRNAs from a variety of species. RESULTS: Our study shows that the application of machine learning techniques, along with the integration of data from multiple species is a useful and general approach for miRNA gene prediction. Based on our experiments, we believe that this new technique is applicable to an extensive range of eukaryotes' genomes. Specific structure and sequence features are first used to identify miRNAs followed by a comparative analysis to decrease the number of false positives (FPs). The resulting algorithm exhibits higher specificity and similar sensitivity compared to currently used algorithms that rely on conserved genomic regions to decrease the rate of FPs.

Algorithms↗

Analysis of mutual information content for EEG responses to odor stimulation for subjects classified by occupation.

To investigate the changes of cortico-cortical connectivity during odor stimulation of subjects classified by occupation, the mutual information content of EEGs was examined for general workers, perfume salespersons and professional perfume researchers. Analysis of the averaged-cross mutual information content (A-CMI) from the EEGs revealed that among the professional perfume researchers changes in the A-CMI values during odor stimulation were more apparent in the frontal region of the brain, while for the general workers and perfume salespersons such changes were more conspicuous in the overall posterior temporal, parietal and frontal regions. These results indicate that the brains of professional perfume researchers respond to odors mainly in the frontal region, reflecting the function of the orbitofrontal cortex (OFC) due to the occupational requirement of these subjects to discriminate or identify odors. During odor stimulation, the perfume salespersons, although relatively more exposed to odors than the general workers, showed similar changes to the general workers. The A-CMI value is in inverse proportion to psychological preferences of the professional perfume researchers and perfume salespersons, though this is not the case with the general workers. This result suggests that functional coupling for people who are occupationally exposed to odors may be related to psychological preference.

Adult↗

Antibiotic resistance: effect of different criteria for classifying isolates as duplicates on apparent resistance frequencies.

OBJECTIVE: To investigate the effect of screening specimens and different criteria for exclusion of duplicate isolates when surveillance of antimicrobial resistances is performed. MATERIALS AND METHODS: Trends in resistance were analysed for recent isolates of selected organisms from Guy's and St Thomas' Hospitals with the use of various criteria for the exclusion of duplicates, including time since the last isolate and antibiogram pattern, and the effect of excluding screening specimens. RESULTS: There was a significant difference of about 8% in the apparent frequency of methicillin resistance in Staphylococcus aureus in inpatients if the time limit for duplicates was set at 5 rather than 30 days; it was about 10% if a 5 day limit was compared with a 365 day limit. There was also a significant difference, of 6-10%, in apparent resistance frequencies if isolates from screening specimens were excluded. Apparent gentamicin resistance rates in Klebsiella spp. varied between 11% and 28%, and the number of apparent patient isolates of gentamicinresistant organisms varied by up to 35%, depending on the duplicate exclusion criteria chosen. Effects were smaller, though still significant, for vancomycin resistance in Enterococcus spp. There was little effect for amoxicillin or cefuroxime resistance in Escherichia coli isolates from general practitioners, where the proportion of duplicates was small. CONCLUSION: Improved surveillance of antibiotic resistance is needed. However, care needs to be taken in setting the criteria for classifying isolates as duplicates and in comparing results where these criteria may be different or unknown.

Bacteria↗

Classifying genealogical origins in hybrid populations using dominant markers.

In hybrid studies, potential for error is high when classifying genealogical origins of individuals (e.g., parental, F1, F2) based on their genotypic arrays. For codominant markers, previous researchers have considered the probability of misclassification by genotypic inspection and proposed alternative maximum-likelihood approaches to estimating genealogical class frequencies. Recently developed dominant marker systems may significantly increase the number of diagnostic loci available for hybrid studies. I examine probabilities of classification error based on the number of dominant loci. As in earlier studies, I assume that only parental and first- and second-generation hybrid crosses between two taxa potentially exist. Thirteen loci with dominant expression from each parental taxon (i.e., 26 total loci) are needed to reduce classification error below 5% for F2 individuals, compared to 13 codominant loci for the same error rate. Use of loci in similar numbers from both taxa most efficiently increases power to characterize all genealogical classes. In contrast, classification of backcrosses to one parental taxon is wholly dependent on loci from the other taxon. Use of dominant diagnostic markers may increase the power and expand the use of maximum-likelihood methods for evaluating hybrid mixtures.

Chimera↗

Cancer registry problems in classifying invasive bladder cancer.

A slide review of diagnostic pathologic tissue obtained from 364 bladder cancer cases, identified through the Iowa Surveillance, Epidemiology, and End Results (SEER) Program in 1983, classified 97 (26.6%) of these cases as invasive bladder cancers. These findings contrasted sharply with the Iowa SEER Program classification that coded 289 (79.4%) of these cases as invasive bladder cancers. These results were validated further by the hazard ratio of 4.54 (95% confidence interval, 2.57 to 8.03) among invasive relative to noninvasive bladder cancer cases when the slide review findings were used. In contrast, the hazard ratio was only 1.70 (95% confidence interval, 0.76 to 3.79) when the Iowa SEER Program findings were used. The traditional method used by the National Cancer Institute's SEER Program to deal with this problem is described and its implications are discussed.

Humans↗

Towards standardization of RNA quality assessment using user-independent classifiers of microcapillary electrophoresis traces.

While it is universally accepted that intact RNA constitutes the best representation of the steady-state of transcription, there is no gold standard to define RNA quality prior to gene expression analysis. In this report, we evaluated the reliability of conventional methods for RNA quality assessment including UV spectroscopy and 28S:18S area ratios, and demonstrated their inconsistency. We then used two new freely available classifiers, the Degradometer and RIN systems, to produce user-independent RNA quality metrics, based on analysis of microcapillary electrophoresis traces. Both provided highly informative and valuable data and the results were found highly correlated, while the RIN system gave more reliable data. The relevance of the RNA quality metrics for assessment of gene expression differences was tested by Q-PCR, revealing a significant decline of the relative expression of genes in RNA samples of disparate quality, while samples of similar, even poor integrity were found highly comparable. We discuss the consequences of these observations to minimize artifactual detection of false positive and negative differential expression due to RNA integrity differences, and propose a scheme for the development of a standard operational procedure, with optional registration of RNA integrity metrics in public repositories of gene expression data.

Cell Line↗

A method for classifying patients according to the nosocomial infection risks associated with diagnoses and surgical procedures.

To compare validly the nosocomial infection rates (NIRs) in groups of patients studied from different time periods and/or different hospitals, one must control for the important factors that influence a patient's susceptibility to infection. The authors developed a method for assessing one component of nosocomial infection risk, based on patients' diagnoses and surgical procedures. This method classifies patients according to their risk of developing a nosocomial infection at each of four infection sites and at all four sites combined. Applying the method to data collected on 136,516 patients from 276 hospitals studied in the SENIC Project (Study on the Efficacy of Nosocomial Infection Control), the authors found that NIRs increased according to the predicted ranking of risk categories, even when the analyses were stratified individually by age, sex, hospital service and exposure to urinary catheterization or continuous ventilatory support. Depending on the site of infection, the rate increased as much as 100-fold from low-risk to high-risk categories. The data indicate that infection risk as assessed with this classification method will account for some of the variation in NIRs due to differences in patients' clinical conditions. Further analyses using multivariate techniques must be performed to explore in detail the relative importance of this risk classification in comparison with other risk factors and to determine which factors must be controlled in SENIC analyses.

Adolescent↗

Regression modeling of consumption or exposure variables classified by type.

Consumption or exposure variables, as potential risk factors, are commonly measured and related to health effects. The measurements may be continuous or discrete, may be grouped into categories and may, in addition, be classified by type. Data analyses utilizing regression methods for the assessment of these risk factors present many problems of modeling and interpretation. Various models are proposed and evaluated, and recommendations are made. Use of the models is illustrated with Cox regression analyses of coronary heart disease mortality after 24 years of follow-up of subjects in the Framingham Study, with the focus being on alcohol consumption among these subjects.

Alcohol Drinking↗

Classifying extraintestinal non-typhoid Salmonella infections.

Non-typhoid Salmonella infection in man has been divided into five clinical groups: gastroenteritis, enteric fever, bacteraemia, chronic carrier state and localized infection. This classification has neither pathogenic nor prognostic significance. We retrospectively reviewed the charts of 183 patients with extraintestinal salmonellosis who presented to our institution during a period of 32 years. Patients were classified into four groups: primary bacteraemia (PB), enteritis-associated bacteraemia (secondary bacteraemia) (SB), digestive focal infection (DI) and non-digestive focal infection (NDI). Sex, age, acquisition, underlying disease and outcome were compared between patients with bacteraemia and diseases with focal infection. The differences found between PB and SB were: community acquisition (66% in PB and 85% in SB, p = 0.06) severe immunosuppression (53% in PB and 15% in SB, p < 0.001) and mortality (37% in PB and 3% in SB, p < 0.001). The differences found between NDI and DI were: age over 60 years (45% in NDI and 18% in DI, p < 0.05), severe immunosuppression (51% in NDI and 12% DI, p < 0.001) and associated bacteraemia (38% in NDI and 6% in DI, p < 0.001). This classification of extraintestinal salmonellosis may have pathogenic and prognostic implications, and could help us to understand the clinical significance of this disease.

Aged↗