Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “classifier”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Classifying children with heavy prenatal alcohol exposure using measures of attention.

Deficits in attention are a hallmark of the effects of heavy prenatal alcohol exposure but although such deficits have been described in the literature, no attempt to use measures of attention to classify children with such exposure has been described. Thus, the current study attempted to classify children with heavy prenatal alcohol exposure (ALC) and non-exposed controls (CON), using four measures of attentional functioning: the Freedom from Distractibility index from the Wechsler Intelligence Scale for Children-Third Edition (WISC-III), the Attention Problems scale from the Child Behavior Checklist (CBCL), and omission and commission error scores from the Test of Variables of Attention (TOVA). Data from two groups of children were analyzed: children with heavy prenatal alcohol exposure and non-exposed controls. Children in the alcohol-exposed group included both children with or without fetal alcohol syndrome. Groups were matched on age, sex, ethnicity, and social class. Data were analyzed using backward logistic regression. The final model included the Freedom from Distractibility index from the WISC-III and the Attention Problems scale from the CBCL. The TOVA variables were not retained in the final model. Classification accuracy was 91.7% overall. Specifically, 93.3% of the alcohol-exposed children and 90% of the control children were accurately classified. These data indicate that children with heavy prenatal alcohol exposure can be distinguished from non-exposed controls with a high degree of accuracy using 2 commonly used measures of attention.

Attention↗

Predicting the genotoxicity of polycyclic aromatic compounds from molecular structure with different classifiers.

Classification models were developed to provide accurate prediction of genotoxicity of 277 polycyclic aromatic compounds (PACs) directly from their molecular structures. Numerical descriptors encoding the topological, geometric, electronic, and polar surface area properties of the compounds were calculated to represent the structural information. Each compound's genotoxicity was represented with IMAX (maximal SOS induction factor) values measured by the SOS Chromotest in the presence and absence of S9 rat liver homogenate. The compounds' class identity was determined by a cutoff IMAX value of 1.25-compounds with IMAX > 1.25 in either test were classified as genotoxic, and the ones with IMAX < or = 1.25 were nongenotoxic. Several binary classification models were generated to predict genotoxicity: k-nearest neighbor (k-NN), linear discriminant analysis, and probabilistic neural network. The study showed k-NN to provide the highest predictive ability among the three classifiers with a training set classification rate of 93.5%. A consensus model was also developed that incorporated the three classifiers and correctly predicted 81.2% of the 277 compounds. It also provided a higher prediction rate on the genotoxic class than any other single model.

Animals↗

Mapping the public health workforce I: a tool for classifying the public health workforce.

We aimed to develop a tool to identify members of the public health workforce and classify them using categories developed for the Chief Medical Officer's project to strengthen the public health function. The tool was developed to gain a picture of London's public health workforce, and needed to be reliable and easy to use in many settings inside and outside the health service. We needed to be able to classify posts from brief information without interrogation of postholders, so that the entire workforces of large organisations could be classified from information provided by only a few key informants. Key questions and decision rules were defined by presenting interviewees in public health with brief information on nine jobs and discussing with them the process by which they determined whether each post was in the public health workforce, and if so, in which category. The questions and decision rules were refined into a classification tool which was presented as a flow diagram and a questionnaire. Application of the tool revealed that it was understood by key informants and resulted in classifications which were accepted by the researchers. The tool has now been applied extensively in London and yielded useful results. Many other applications in public health workforce planning and development are anticipated.

Humans↗

The value of classifying interstitial pneumonitis in childhood according to defined histological patterns.

AIMS: Interstitial pneumonitis in children is very rare and most cases have been classified according to their counterparts in adults, although the term 'chronic pneumonitis of infancy' has recently been proposed for a particular pattern of interstitial lung disease in infants. We reviewed our paediatric cases of interstitial pneumonitis, first, to look at the spectrum of histological patterns found in this age group and, second, to determine whether the classification of such cases in childhood is both appropriate and worthwhile. METHODS AND RESULTS: Twenty-five of 38 open lung biopsies showed an overlapping spectrum of interstitial pneumonitis, including three cases that fulfilled the histological criteria for chronic pneumonitis of infancy. There were 11 cases of reactive pulmonary lymphoid hyperplasia (either lymphoid interstitial pneumonitis or follicular bronchiolitis), five of which were associated with abnormalities of the immune system. Four cases were classified as desquamative interstitial pneumonitis and the remaining seven cases were classified as nonspecific interstitial pneumonitis. There were no cases with the histological features of usual interstitial pneumonitis. Most patients responded to steroids but tended to have a residual deficit in lung function. Mortality appeared to be associated with presentation at a young age. CONCLUSION: Classification of interstitial pneumonitis according to their adult counterparts is appropriate for this younger age group and can provide valuable information for the clinician. The term 'chronic pneumonitis of infancy' refers to a specific histological pattern, but whether it represents a separate disease or a reflection of pulmonary immaturity remains to be proven.

Adolescent↗

Urodynamic variables cannot be used to classify the severity of detrusor instability.

OBJECTIVE: To explore the relationship between subjective severity of symptoms of detrusor instability (DI) on presentation, outcome after treatment for DI and initial diagnostic urodynamic variables, with the aim of identifying a urodynamic variable which might, by predicting a favourable outcome from treatment, classify the severity of DI. PATIENTS AND METHODS: Women with a urodynamically proven diagnosis of DI were recruited prospectively for the study. Data on disease symptoms and variables from their diagnostic cystometrogram were collected. All women were then treated and their outcome at 6 weeks after treatment compared with the initial urodynamic variables. Data on severity of symptoms were compared with initial urodynamic variables to explore any differences in these variables attributable to symptom severity. RESULTS: Of 300 women studied (mean age 54 years, SD 16), 290 were treated with oxybutynin and bladder retraining. At 6 weeks, 82 women had their treatment outcome classified as worse/no change; 218 women had improved. When good or poor outcome was compared with the urodynamic results, there was no significant difference between the groups. Likewise, the severity of symptoms did not relate to the values of urodynamic variables. CONCLUSIONS: There was no statistically significant relationship between reported severity of symptoms and urodynamic variables, and no relationship between the urodynamic variables used and response to treatment. Therefore, using these values it is not possible to predict a favourable outcome from treatment or to use them to classify disease severity.

Female↗

Factors associated with an increased risk of prevalent and incident grade III cervical intraepithelial neoplasia and invasive cervical cancer among women with Papanicolaou tests classified as grades I or II cervical intraepithelial neoplasia.

OBJECTIVE: Women with Papanicolaou tests classified as cervical intraepithelial neoplasia grade I or II are treated conservatively in many countries. However, these women are at an increased risk of having underlying prevalent and incident grade III cervical intraepithelial neoplasia and invasive cancer. This study was undertaken to identify factors that could predict these clinically important disease states. STUDY DESIGN: Five hundred women with Papanicolaou tests classified as persistent grade I or II cervical intraepithelial neoplasia underwent a repeat test, human papillomavirus testing with Hybrid Capture assay (Digene, Silver Spring, Md) and polymerase chain reaction, and colposcopy with histologic assessment. One hundred fifty-seven women with histologically proven grade I or II cervical intraepithelial neoplasia were monitored conservatively for a minimum of 9 months to assess predictors of incident grade III cervical intraepithelial neoplasia. RESULTS: One hundred fifty-one women with prevalent grade III cervical intraepithelial neoplasia and 5 women with prevalent invasive cancer were identified at the first colposcopy. A repeated Papanicolaou test classified as higher than grade II cervical intraepithelial neoplasia and detection of oncogenic human papillomavirus types were significant predictors of underlying grade III cervical intraepithelial neoplasia and cancer in the multivariate analysis. Seventeen of 157 women (10.8%) with grade I or II cervical intraepithelial neoplasia progressed to grade III cervical intraepithelial neoplasia. Age >30 years and detection of oncogenic human papillomavirus were significantly correlated with progression in the multivariate analysis. No progression was observed in women who were negative for human papillomavirus. CONCLUSION: The high rate of underlying prevalent grade III cervical intraepithelial neoplasia and cancer found in our study (31.2%) indicates that conservative management of women with persistent grade I or II cervical intraepithelial neoplasia should be discouraged. Colposcopy with histologic assessment should be recommended as the standard of care. However, for women with histologically proven grade I or II cervical intraepithelial neoplasia, subsequent conservative management was safe in our study for those who were negative for human papillomavirus by type-specific polymerase chain reaction.

Adolescent↗

The effects of distinctiveness in recognising and classifying faces.

In an earlier study it was found that distinctive familiar faces were recognised faster than typical familiar faces in a familiarity decision task. In the first experiment reported here this effect was replicated with the use of celebrities' faces rather than personally familiar faces. In the second and third experiments the effect of distinctiveness was found to reverse if the task was to distinguish between faces and jumbled faces. Subjects took longer to classify distinctive faces as faces than they did to classify typical faces. Thus distinctive faces were recognised faster, but were classified as faces more slowly than were typical faces, both when personally familiar faces and when famous faces were used as stimuli. These results are interpeted as evidence that faces are encoded by reference to a general face prototype.

Decision Making↗

Hepatitis C virus variants from Vietnam are classifiable into the seventh, eighth, and ninth major genetic groups.

Thirty-four (41%) of 83 hepatitis C virus (HCV) isolates from commercial blood donors in Vietnam were not classifiable into genotype I/1a, II/1b, III/2a, IV/2b, or V/3a; for 15 of them, the sequence was determined for 1.6 kb in the 5'-terminal region and 1.1 kb in the 3'-terminal region. Comparison of the 15 Vietnamese isolates among themselves and with reported full or partial HCV genomic sequences indicated that they were classifiable into four major groups (groups 6-9) divided into six genotypes (6a, 7a, 7b, 8a, 8b, and 9a). Vietnamese HCV isolates of genotypes 7a, 7b, 8a, 8b, and 9a were significantly different from those classified into groups 4, 5, and 6 based on divergence within partial sequences; those of genotype 6a were homologous to a Hong Kong isolate (HK2) of genotype 6a. Phylogenetic trees based on the envelope 1 (E1) gene (576 bp) of 55 isolates and a part of the nonstructural 5 (NS5) region (1093 bp) of 43 isolates revealed at least nine major groups, three of which (groups 7, 8, and 9) were identified only in Vietnamese blood donors. With a prospect that many more HCV isolates with significant sequence divergence will be reported from all over the world, the domain of the HCV genome to be compared and criteria for grouping/typing and genotyping/subtyping will have to be determined, so that they may be correlated with virological, epidemiological, and clinical characteristics.

Base Sequence↗

Use of an artificial neural network (ANN) for classifying nursing care needed, using incomplete input data.

BACKGROUND: In German nursing insurance, the act of classifying the client into four categories of disability is based on legally defined distinct criteria. When classifying deceased persons it is often impossible to collect all the required information. PRIMARY OBJECTIVE: We aimed to determine the ability of an artificial neural network (ANN) to calculate the category of disability, to investigate the response of the ANN to input items of different nature, quantity and data quality, and to estimate the minimum number of training data required. RESEARCH DESIGN: The investigation was conducted as a retrospective observational study. METHODS AND PROCEDURES: The analysis was based on routine records of 14000 adult clients of the nursing insurance. Several ANNs were trained, varying nature, number and quality of the input items as well as the size of the training data set. Each ANN's classification competence was tested on independent validation data, judging the ANN's conformance to the result of the individual expert assessment, using kappa statistics. MAIN RESULTS: Fed with all 30 input items available, the net classified 80% of cases correctly (weighted kappa = 0.78). Using three input items, weighted kappa was 0.63. Severe misclassification (deviation by more than one category in either direction) ranged between 0.2% (all 30 input items) and 3.7% (3/30 items). The less complete the individual input items were, the less accurate was the net's estimate. A 20% rate of missing values was well tolerated. A training set comprising 500 cases was adequate. CONCLUSIONS: The input item set inherits redundancy. The ANN's ability to correctly respond to subsets of input items makes it a powerful tool in quality control. In the categorization of deceased persons when only an incomplete input item set is available, the ANN can achieve satisfactory results.

Activities of Daily Living↗

Joint classifier and feature optimization for comprehensive cancer diagnosis using gene expression data.

Recent research has demonstrated quite convincingly that accurate cancer diagnosis can be achieved by constructing classifiers that are designed to compare the gene expression profile of a tissue of unknown cancer status to a database of stored expression profiles from tissues of known cancer status. This paper introduces the JCFO, a novel algorithm that uses a sparse Bayesian approach to jointly identify both the optimal nonlinear classifier for diagnosis and the optimal set of genes on which to base that diagnosis. We show that the diagnostic classification accuracy of the proposed algorithm is superior to a number of current state-of-the-art methods in a full leave-one-out cross-validation study of five widely used benchmark datasets. In addition to its superior classification accuracy, the algorithm is designed to automatically identify a small subset of genes (typically around twenty in our experiments) that are capable of providing complete discriminatory information for diagnosis. Focusing attention on a small subset of genes is useful not only because it produces a classifier with good generalization capacity, but also because this set of genes may provide insights into the mechanisms responsible for the disease itself. A number of the genes identified by the JCFO in our experiments are already in use as clinical markers for cancer diagnosis; some of the remaining genes may be excellent candidates for further clinical investigation. If it is possible to identify a small set of genes that is indeed capable of providing complete discrimination, inexpensive diagnostic assays might be widely deployable in clinical settings.

Algorithms↗

CD5- small B-cell leukemias are rarely classifiable as chronic lymphocytic leukemia.

Expression of the CD5 antigen by neoplastic cells often is considered a diagnostic criterion for B-cell chronic lymphocytic leukemia (B-CLL). However, published series frequently include a number of CD5- cases. We studied the spectrum of CD5- B-cell lymphoproliferative disorders presenting with leukemia involvement and reassessed the prevalence of CD5- B-CLL. We immunophenotyped 192 cases of clonal, small lymphocytic, B-cell disorders involving peripheral blood or bone marrow. Of these, 41 CD5- cases were further analyzed, correlating the immunophenotypic findings with pathologic material and clinical data. Only 3 CD5- cases were classified as CD5- B-CLL. These 3 cases had features unusual for B-CLL, including bright surface immunoglobulin expression, bright CD20 expression, and absence of CD23 expression (2 cases) or Richter syndrome (1 case). The remainder of the CD5- cases consisted of hairy cell leukemia, hairy cell variant, prolymphocytic leukemia, follicular center cell lymphoma, lymphoplasmacytic lymphoma, splenic marginal zone lymphoma (SMZL), small lymphocytic lymphoma with marrow fibrosis, and lymphoma, not further classified. Eight cases remained unclassified, but some displayed features of SMZL. CD5- lymphoproliferative disorders of peripheral blood or bone marrow are unlikely to be CLL and often are classified more appropriately as non-Hodgkin lymphoma in the leukemia phase.

Adult↗

Recurrent detoxifications are associated with craving in patients classified as type 1 according to Lesch's typology.

AIMS: Recurrent detoxifications have been suggested to be associated with elevated alcohol craving. The aim of this investigation was to study the influence of preceding detoxifications on craving in patients with alcoholism classified according to Lesch's typology. METHODS: We examined 192 patients (154 men, 38 women) after admission for detoxification treatment. Craving was assessed using the Obsessive Compulsive Drinking Scale, and patients were classified into one of the four subgroups of Lesch's typology. The number of preceding detoxifications was assessed with a structured interview. RESULTS: Lesch's typology type 4 patients showed significantly higher craving scores than type 1-3 patients (Mann-Whitney U-Test; P < 0.05). With respect to the influence of recurrent detoxifications, we found a significant correlation between the number of preceding detoxifications and the extent of craving for the whole population (Spearman's rho r = 0.241, P = 0.001, N = 192), particularly for patients of Lesch's type 1 (Spearman's rho r = 0.534, P = 0.001, N = 37). No significant association was found for patients of the other subgroups (Lesch's type 2-4). CONCLUSION: The influence of recurrent detoxifications on craving is especially important in patients with Lesch's type 1. Our results underline the importance of the kindling effect particularly in this group of patients, possibly mediated by an increase of glutamatergic neurotransmission. Furthermore, our results emphasize the need to classify patients with alcohol-dependency in addiction research.

Adult↗

A neural network classifier capable of recognizing the patterns of all major subcellular structures in fluorescence microscope images of HeLa cells.

MOTIVATION: Assessment of protein subcellular location is crucial to proteomics efforts since localization information provides a context for a protein's sequence, structure, and function. The work described below is the first to address the subcellular localization of proteins in a quantitative, comprehensive manner. RESULTS: Images for ten different subcellular patterns (including all major organelles) were collected using fluorescence microscopy. The patterns were described using a variety of numeric features, including Zernike moments, Haralick texture features, and a set of new features developed specifically for this purpose. To test the usefulness of these features, they were used to train a neural network classifier. The classifier was able to correctly recognize an average of 83% of previously unseen cells showing one of the ten patterns. The same classifier was then used to recognize previously unseen sets of homogeneously prepared cells with 98% accuracy. AVAILABILITY: Algorithms were implemented using the commercial products Matlab, S-Plus, and SAS, as well as some functions written in C. The scripts and source code generated for this work are available at http://murphylab.web.cmu.edu/software. CONTACT: murphy@cmu.edu

Algorithms↗

HallmarkGraph: a cancer hallmark informed graph neural network for classifying hierarchical tumor subtypes.

MOTIVATION: Accurate tumor subtype diagnosis is crucial for precision oncology, yet current methodologies face significant challenges. These include balancing model accuracy with interpretability and the high costs of generating multi-omics data in clinical settings. Moreover, there is a lack of validated models capable of classifying hierarchical tumor subtypes across a comprehensive pan-cancer cohort. RESULTS: We present a graph neural network, HallmarkGraph, the first biologically informed model developed to classify hierarchical tumor subtypes in human cancer. Inspired by cancer hallmarks, the model's architecture integrates transcriptome profiles and gene regulatory interactions to perform multi-label classification. We evaluate the model on a comprehensive pan-cancer cohort comprising 11&#xa0;476 samples from 26 primary cancers with 405 subtypes up to eight levels. The model demonstrates exceptional performance, achieving 5-fold cross-validation accuracy between 85% and 99% for tumor subtypes labeled with increasing details of genomic information. It also shows good generalizability on a validation dataset of 887 samples, assessed using three metrics that consider tumor subtypes at individual, combined, and sample levels. Benchmarking and ablation experiments show that hallmark-based embeddings slightly influence model performance, while the integrated multilayer perceptron plays a significant role in determining classifier accuracy. Additionally, we use the SHAP method to link cancer hallmarks with genes, identifying key features that influence model decisions. Our findings present a biologically informed machine learning framework capable of tracking tumor transcriptomic trajectories and distinguishing inter- and intra-tumor heterogeneity in pan-cancer. This approach holds promise for enhancing cancer diagnostics. AVAILABILITY AND IMPLEMENTATION: HallmarkGraph is accessible at https://github.com/laixn/HallmarkGraph.

Humans↗

How many samples are needed to build a classifier: a general sequential approach.

MOTIVATION: The standard paradigm for a classifier design is to obtain a sample of feature-label pairs and then to apply a classification rule to derive a classifier from the sample data. Typically in laboratory situations the sample size is limited by cost, time or availability of sample material. Thus, an investigator may wish to consider a sequential approach in which there is a sufficient number of patients to train a classifier in order to make a sound decision for diagnosis while at the same time keeping the number of patients as small as possible to make the studies affordable. RESULTS: A sequential classification procedure is studied via the martingale central limit theorem. It updates the classification rule at each step and provides stopping criteria to ensure with a certain confidence that at stopping a future subject will have misclassification probability smaller than a predetermined threshold. Simulation studies and applications to microarray data analysis are provided. The procedure possesses several attractive properties: (1) it updates the classification rule sequentially and thus does not rely on distributions of primary measurements from other studies; (2) it assesses the stopping criteria at each sequential step and thus can substantially reduce cost via early stopping; and (3) it is not restricted to any particular classification rule and therefore applies to any parametric or non-parametric method, including feature selection or extraction. AVAILABILITY: R-code for the sequential stopping rule is available at http://stat.tamu.edu/~wfu/microarray/sequential/R-code.html

Algorithms↗

A two-stage classifier for identification of protein-protein interface residues.

MOTIVATION: The ability to identify protein-protein interaction sites and to detect specific amino acid residues that contribute to the specificity and affinity of protein interactions has important implications for problems ranging from rational drug design to analysis of metabolic and signal transduction networks. RESULTS: We have developed a two-stage method consisting of a support vector machine (SVM) and a Bayesian classifier for predicting surface residues of a protein that participate in protein-protein interactions. This approach exploits the fact that interface residues tend to form clusters in the primary amino acid sequence. Our results show that the proposed two-stage classifier outperforms previously published sequence-based methods for predicting interface residues. We also present results obtained using the two-stage classifier on an independent test set of seven CAPRI (Critical Assessment of PRedicted Interactions) targets. The success of the predictions is validated by examining the predictions in the context of the three-dimensional structures of protein complexes.

Amino Acid Sequence↗

Classifying noisy protein sequence data: a case study of immunoglobulin light chains.

SUMMARY: The classification of protein sequences obtained from patients with various immunoglobulin-related conformational diseases may provide insight into structural correlates of pathogenicity. However, clinical data are very sparse and, in the case of antibody-related proteins, the collected sequences have large variability with only a small subset of variations relevant to the protein pathogenicity (function). On this basis, these sequences represent a model system for development of strategies to recognize the small subset of function-determining variations among the much larger number of primary structure diversifications introduced during evolution. Under such conditions, most protein classification algorithms have limited accuracy. To address this problem, we propose a support vector machine (SVM)-based classifier that combines sequence and 3D structural averaging information. Each amino acid in the sequence is represented by a set of six physicochemical properties: hydrophobicity, hydrophilicity, volume, surface area, bulkiness and refractivity. Each position in the sequence is described by the properties of the amino acid at that position and the properties of its neighbors in 3D space or in the sequence. A structure template is selected to determine neighbors in 3D space and a window size is used to determine the neighbors in the sequence. The test data consist of 209 proteins of human antibody immunoglobulin light chains, each represented by aligned sequences of 120 amino acids. The methodology is applied to the classification of protein sequences collected from patients with and without amyloidosis, and indicates that the proposed modified classifiers are more robust to sequence variability than standard SVM classifiers, improving classification error between 5 and 25% and sensitivity between 9 and 17%. The classification results might also suggest possible mechanisms for the propensity of immunoglobulin light chains to amyloid formation.

Algorithms↗

Small, fuzzy and interpretable gene expression based classifiers.

MOTIVATION: Interpretation of classification models derived from gene-expression data is usually not simple, yet it is an important aspect in the analytical process. We investigate the performance of small rule-based classifiers based on fuzzy logic in five datasets that are different in size, laboratory origin and biomedical domain. RESULTS: The classifiers resulted in rules that can be readily examined by biomedical researchers. The fuzzy-logic-based classifiers compare favorably with logistic regression in all datasets. AVAILABILITY: Prototype available upon request.

Algorithms↗