Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “classifier”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Rapid identification of Gram-positive anaerobic coccal species originally classified in the genus Peptostreptococcus by multiplex PCR assays using genus- and species-specific primers.

Here, a rapid and reliable two-step multiplex PCR assay for identifying 14 Gram-positive anaerobic cocci (GPAC) species originally classified in the genus Peptostreptococcus (Anaerococcus hydrogenalis, Anaerococcus lactolyticus, Anaerococcus octavius, Anaerococcus prevotii, Anaerococcus tetradius, Anaerococcus vaginalis, Finegoldia magna, Micromonas micros, Peptostreptococcus anaerobius, Peptoniphilus asaccharolyticus, Peptoniphilus harei, Peptoniphilus indolicus, Peptoniphilus ivorii and Peptoniphilus lacrimalis) is reported. Fourteen type strains representing 14 GPAC species were first identified to the genus level by multiplex PCR (multiplex PCR-G). Since three of these genera (Finegoldia, Micromonas and Peptostreptococcus) contain only a single species, F. magna, M. micros and P. anaerobius, respectively, these organisms were identified to the species level directly by using the multiplex PCR-G. Then six species of the genus Anaerococcus (A. hydrogenalis, A. lactolyticus, A. octavius, A. prevotii, A. vaginalis and A. tetradius) were further identified to the species level using multiplex PCR assays (multiplex PCR-Ia and multiplex PCR-Ib). Similarly, five species of the genus Peptoniphilus (Pn. asaccharolyticus, Pn. harei, Pn. indolicus, Pn. ivorii and Pn. lacrimalis) were identified to the species level using multiplex PCR-IIa and multiplex PCR-IIb. The established two-step multiplex PCR identification scheme was applied to the identification of 190 clinical isolates of GPAC species that had been identified previously to the species level by 16S rRNA sequencing and phenotypic tests. The identification obtained from multiplex PCR assays showed 100 % agreement with 16S rDNA sequencing identification, but only 65 % (123/190) agreement with the identification obtained by phenotypic tests. The multiplex PCR scheme established in this study is a simple, rapid and reliable method for the identification of GPAC species. It will permit a more accurate assessment of the role of various GPAC species in infection and of the degree of antimicrobial resistance in each of the group members.

Base Sequence↗

Genotype C of hepatitis B virus can be classified into at least two subgroups.

A genomic characterization of hepatitis B virus (HBV) was done for 56 pre-S1/pre-S2 genes and 10 full-length HBV genotype C isolates from five Asian countries. Phylogenetic analysis of the pre-S1/pre-S2 genes revealed two major groups within genotype C: one for isolates from southeast Asia including Vietnam, Myanmar and Thailand (named HBV/C1) and the other for isolates from Far East Asia including Japan, Korea and China (named HBV/C2). This finding was confirmed by phylogenetic analysis based on the full-length sequence of 32 HBV genotype C isolates, including 22 from database entries. Two isolates from Okinawa, the island off the southern end of Japan, formed a different branch. Specific amino acid sequence changes were identified in the large S protein (amino acids 51, 54, 60, 62 and 73) and P protein (amino acids 231, 233, 236, 248, 252 and 304). Our results indicate that genotype C of HBV can be classified into at least two subgroups.

Asia↗

Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE): a cloud-based platform for curating and classifying germline variants.

Variant interpretation in the era of massively parallel sequencing is challenging. Although many resources and guidelines are available to assist with this task, few integrated end-to-end tools exist. Here, we present the Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE), a web- and cloud-based platform for annotation, identification, and classification of variations in known or putative disease genes. Starting from a set of variants in variant call format (VCF), variants are annotated, ranked by putative pathogenicity, and presented for formal classification using a decision-support interface based on published guidelines from the American College of Medical Genetics and Genomics (ACMG). The system can accept files containing millions of variants and handle single-nucleotide variants (SNVs), simple insertions/deletions (indels), multiple-nucleotide variants (MNVs), and complex substitutions. PeCanPIE has been applied to classify variant pathogenicity in cancer predisposition genes in two large-scale investigations involving >4000 pediatric cancer patients and serves as a repository for the expert-reviewed results. PeCanPIE was originally developed for pediatric cancer but can be easily extended for use for nonpediatric cancers and noncancer genetic diseases. Although PeCanPIE's web-based interface was designed to be accessible to non-bioinformaticians, its back-end pipelines may also be run independently on the cloud, facilitating direct integration and broader adoption. PeCanPIE is publicly available and free for research use.

Child↗

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype↗

Classifying N-qubit entanglement via Bell's inequalities.

All the states of N qubits can be classified into N-1 entanglement classes from 2-entangled to N-entangled (fully entangled) states. Each class of entangled states is characterized by an entanglement index that depends on the partition of N. The larger the entanglement index of a state, the more entangled or the less separable is the state in the sense that a larger maximal violation of Bell's inequality is attainable for this class of state.

Journal Article↗

Strange attractors are classified by bounding Tori.

There is at present a doubly discrete classification for strange attractors of low dimension, d(L)<3. A branched manifold describes the stretching and squeezing processes that generate the strange attractor, and a basis set of orbits describes the complete set of unstable periodic orbits in the attractor. To this we add a third discrete classification level. Strange attractors are organized by the boundary of an open set surrounding their branched manifold. The boundary is a torus with g holes that is dressed by a surface flow with 2(g-1) singular points. All known strange attractors in R3 are classified by genus, g, and flow type.

Journal Article↗

Classifying a protein in the CATH database of domain structures.

The CATH database of protein domain structures classifies structures according to their (C)lass, (A)rchitecture, (T)opology or fold and (H)omologous family (http://www.biochem.ucl.ac.uk/bsm/cath). Although the protocol used is mostly automatic, manual inspection is used to check assignments at some critical stages, such as the detection of very distantly related homologues and anologues and the assignment of novel architectures. Described in this article is a recently established facility to search the database with the coordinates of a newly determined structure. The CATH server first locates domain boundaries and then uses automatic sequence and structure comparison methods to assign this new structure to one or more of the domain families within CATH. Diagnostic reports are generated, together with multiple structural alignments for close relatives. The Server can be accessed over the World Wide Web (WWW) and mirror sites are planned to improve access.

Amino Acid Sequence↗

Linear and neural models for classifying breast masses.

Computational methods can be used to provide an initial screening or a second opinion in medical settings and may improve the sensitivity and specificity of diagnoses. In the current study, linear discriminant models and artificial neural networks are trained to detect breast cancer in suspicious masses using radiographic features and patient age. Results on 139 suspicious breast masses (79 malignant, 60 benign, biopsy proven) indicate that a significant probability of detecting malignancies can be achieved at the risk of a small percentage of false positives. Receiver operating characteristic (ROC) analysis favors the use of linear models, however, a new measure related to the area under the ROC curve (AZ) suggests a possible benefit from hybridizing linear and nonlinear classifiers.

Breast Neoplasms↗

Classifying mammographic mass shapes using the wavelet transform modulus-maxima method.

In this article, multiresolution analysis, specifically the discrete wavelet transform modulus-maxima (mod-max) method, is utilized for the extraction of mammographic mass shape features. These shape features are used in a classification system to classify masses as round, nodular, or stellate. The multiresolution shape features are compared with traditional uniresolution shape features for their class discriminating abilities. The study involved 60 digitized mammographic images. The masses were segmented manually by radiologists, prior to introduction to the classification system. The uniresolution and multiresolution shape features were calculated using the radial distance measure of the mass boundaries. The discriminating power of the shape features were analyzed via linear discriminant analysis (LDA). The classification system utilized a simple Euclidean metric to determine class membership. The system was tested using the apparent and leave-one-out test methods. The classification system when using the multiresolution and uniresolution shape features resulted in classification rates of 83% and 80% for the apparent and leave-one-out test methods, respectively. In comparison, when only the uniresolution shape features were used, the classification rates were 72 and 68% for the apparent and leave-one-out test methods, respectively.

Breast Neoplasms↗

Monocular precrash vehicle detection: features and classifiers.

Robust and reliable vehicle detection from images acquired by a moving vehicle (i.e., on-road vehicle detection) is an important problem with applications to driver assistance systems and autonomous, self-guided vehicles. The focus of this work is on the issues of feature extraction and classification for rear-view vehicle detection. Specifically, by treating the problem of vehicle detection as a two-class classification problem, we have investigated several different feature extraction methods such as principal component analysis, wavelets, and Gabor filters. To evaluate the extracted features, we have experimented with two popular classifiers, neural networks and support vector machines (SVMs). Based on our evaluation results, we have developed an on-board real-time monocular vehicle detection system that is capable of acquiring grey-scale images, using Ford's proprietary low-light camera, achieving an average detection rate of 10 Hz. Our vehicle detection algorithm consists of two main steps: a multiscale driven hypothesis generation step and an appearance-based hypothesis verification step. During the hypothesis generation step, image locations where vehicles might be present are extracted. This step uses multiscale techniques not only to speed up detection, but also to improve system robustness. The appearance-based hypothesis verification step verifies the hypotheses using Gabor features and SVMs. The system has been tested in Ford's concept vehicle under different traffic conditions (e.g., structured highway, complex urban streets, and varying weather conditions), illustrating good performance.

Accidents, Traffic↗

A support vectors classifier approach to predicting the risk of progression of adolescent idiopathic scoliosis.

A support vector classifier (SVC) approach was employed in predicting the risk of progression of adolescent idiopathic scoliosis (AIS), a condition that causes visible trunk asymmetries. As the aetiology of AIS is unknown, its risk of progression can only be predicted from measured indicators. Previous studies suggest that individual indicators of AIS do not reliably predict its risk of progression. Complex indicators with better predictive values have been developed but are unsuitable for clinical use as obtaining their values is often onerous, involving much skill and repeated measurements taken over time. Based on the hypothesis that combining common indicators of AIS using an SVC approach would produce better prediction results more quickly, we conducted a study using three datasets comprising a total of 44 moderate AIS patients (30 observed, 14 treated with brace). Of the 44 patients, 13 progressed less than 5 degrees and 31 progressed more than 5 degrees. One dataset comprised all the patients. A second dataset comprised all the observed patients and a third comprised all the brace-treated patients. Twenty-one radiographic and clinical indicators were obtained for each patient. The result of testing on the three datasets showed that the system achieved 100% accuracy in training and 65%-80% accuracy in testing. It outperformed a "statistically equivalent" logistic regression model and a stepwise linear regression model on the said datasets. It took less than 20 min per patient to measure the indicators, input their values into the system, and produce the needed results, making the system viable for use in a clinical environment.

Adolescent↗

Implementation of a real-time human movement classifier using a triaxial accelerometer for ambulatory monitoring.

The real-time monitoring of human movement can provide valuable information regarding an individual's degree of functional ability and general level of activity. This paper presents the implementation of a real-time classification system for the types of human movement associated with the data acquired from a single, waist-mounted triaxial accelerometer unit. The major advance proposed by the system is to perform the vast majority of signal processing onboard the wearable unit using embedded intelligence. In this way, the system distinguishes between periods of activity and rest, recognizes the postural orientation of the wearer, detects events such as walking and falls, and provides an estimation of metabolic energy expenditure. A laboratory-based trial involving six subjects was undertaken, with results indicating an overall accuracy of 90.8% across a series of 12 tasks (283 tests) involving a variety of movements related to normal daily activities. Distinction between activity and rest was performed without error; recognition of postural orientation was carried out with 94.1% accuracy, classification of walking was achieved with less certainty (83.3% accuracy), and detection of possible falls was made with 95.6% accuracy. Results demonstrate the feasibility of implementing an accelerometry-based, real-time movement classifier using embedded intelligence.

Acceleration↗

Classifiability-based omnivariate decision trees.

Top-down induction of decision trees is a simple and powerful method of pattern classification. In a decision tree, each node partitions the available patterns into two or more sets. New nodes are created to handle each of the resulting partitions and the process continues. A node is considered terminal if it satisfies some stopping criteria (for example, purity, i.e., all patterns at the node are from a single class). Decision trees may be univariate, linear multivariate, or nonlinear multivariate depending on whether a single attribute, a linear function of all the attributes, or a nonlinear function of all the attributes is used for the partitioning at each node of the decision tree. Though nonlinear multivariate decision trees are the most powerful, they are more susceptible to the risks of overfitting. In this paper, we propose to perform model selection at each decision node to build omnivariate decision trees. The model selection is done using a novel classifiability measure that captures the possible sources of misclassification with relative ease and is able to accurately reflect the complexity of the subproblem at each node. The proposed approach is fast and does not suffer from as high a computational burden as that incurred by typical model selection algorithms. Empirical results over 26 data sets indicate that our approach is faster and achieves better classification accuracy compared to statistical model select algorithms.

Algorithms↗

On using prototype reduction schemes and classifier fusion strategies to optimize kernel-based nonlinear subspace methods.

In Kernel-based Nonlinear Subspace (KNS) methods, the length of the projections onto the principal component directions in the feature space, is computed using a kernel matrix, K, whose dimension is equivalent to the number of sample data points. Clearly this is problematic, especially, for large data sets. In this paper, we solve this problem by subdividing the data into smaller subsets, and utilizing a Prototype Reduction Scheme (PRS) as a preprocessing module, to yield more refined representative prototypes. Thereafter, a Classifier Fusion Strategy (CFS) is invoked as a postprocessing module, to combine the individual KNS classification results to derive a consensus decision. Essentially, the PRS is used to yield computational advantage, and the CFS, in turn, is used to compensate for the decreased efficiency caused by the data set division. Our experimental results demonstrate that the proposed mechanism significantly reduces the prototype extraction time as well as the computation time without sacrificing the classification accuracy. The results especially demonstrate a significant computational advantage for large data sets within a parallel processing philosophy.

Algorithms↗

Assessing classifiers from two independent data sets using ROC analysis: a nonparametric approach.

This paper considers binary classification. We assess a classifier in terms of the Area Under the ROC Curve (AUC). We estimate three important parameters, the conditional AUC (conditional on a particular training set) and the mean and variance of this AUC. We derive, as well, a closed form expression of the variance of the estimator of the AUC. This expression exhibits several components of variance that facilitate an understanding for the sources of uncertainty of that estimate. In addition, we estimate this variance, i.e., the variance of the conditional AUC estimator. Our approach is nonparametric and based on general methods from U-statistics; it addresses the case where the data distribution is neither known nor modeled and where there are only two available data sets, the training and testing sets. Finally, we illustrate some simulation results for these estimators.

Algorithms↗

An empirical risk functional to improve learning in a neuro-fuzzy classifier.

The paper proposes a new Empirical Risk Functional as cost function for training neuro-fuzzy classifiers. This cost function, called Approximate Differentiable Empirical Risk Functional (ADERF), provides a differentiable approximation of the misclassification rate so that the Empirical Risk Minimization Principle formulated in Vapnik's Statistical Learning Theory can be applied. Also, based on the proposed ADERF, a learning algorithm is formulated. Experimental results on a number of benchmark classification tasks are provided and comparison to alternative approaches given.

Journal Article↗

Class decomposition for GA-based classifier agents--a Pitt approach.

This paper proposes a class decomposition approach to improve the performance of GA-based classifier agents. This approach partitions a classification problem into several class modules in the output domain, and each module is responsible for solving a fraction of the original problem. These modules are trained in parallel and independently, and results obtained from them are integrated to form the final solution by resolving conflicts. Benchmark classification data sets are used to evaluate the proposed approaches. The experiment results show that class decomposition can help achieve higher classification rate with training time reduced.

Journal Article↗

An empirical comparison of nine pattern classifiers.

There are many learning algorithms available in the field of pattern classification and people are still discovering new algorithms that they hope will work better. Any new learning algorithm, beside its theoretical foundation, needs to be justified in many aspects including accuracy and efficiency when applied to real life problems. In this paper, we report the empirical comparison of a recent algorithm RM, its new extensions and three classical classifiers in different aspects including classification accuracy, computational time and storage requirement. The comparison is performed in a standardized way and we believe that this would give a good insight into the algorithm RM and its extension. The experiments also show that nominal attributes do have an impact on the performance of those compared learning algorithms.

Algorithms↗