Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “CLASSIFICATION”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Optimization models for cancer classification: extracting gene interaction information from microarray expression data.

MOTIVATION: Microarray data appear particularly useful to investigate mechanisms in cancer biology and represent one of the most powerful tools to uncover the genetic mechanisms causing loss of cell cycle control. Recently, several different methods to employ microarray data as a diagnostic tool in cancer classification have been proposed. These procedures take changes in the expression of particular genes into account but do not consider disruptions in certain gene interactions caused by the tumor. It is probable that some genes participating in tumor development do not change their expression level dramatically. Thus, they cannot be detected by simple classification approaches used previously. For these reasons, a classification procedure exploiting information related to changes in gene interactions is needed. RESULTS: We propose a MAximal MArgin Linear Programming (MAMA) method for the classification of tumor samples based on microarray data. This procedure detects groups of genes and constructs models (features) that strongly correlate with particular tumor types. The detected features include genes whose functional relations are changed for particular cancer types. The proposed method was tested on two publicly available datasets and demonstrated a prediction ability superior to previously employed classification schemes. AVAILABILITY: The MAMA system was developed using the linear programming system LINDO http://www.lindo.com. A Perl script that specifies the optimization problem for this software is available upon request from the authors.

Algorithms↗

Classification using partial least squares with penalized logistic regression.

MOTIVATION: One important aspect of data-mining of microarray data is to discover the molecular variation among cancers. In microarray studies, the number n of samples is relatively small compared to the number p of genes per sample (usually in thousands). It is known that standard statistical methods in classification are efficient (i.e. in the present case, yield successful classifiers) particularly when n is (far) larger than p. This naturally calls for the use of a dimension reduction procedure together with the classification one. RESULTS: In this paper, the question of classification in such a high-dimensional setting is addressed. We view the classification problem as a regression one with few observations and many predictor variables. We propose a new method combining partial least squares (PLS) and Ridge penalized logistic regression. We review the existing methods based on PLS and/or penalized likelihood techniques, outline their interest in some cases and theoretically explain their sometimes poor behavior. Our procedure is compared with these other classifiers. The predictive performance of the resulting classification rule is illustrated on three data sets: Leukemia, Colon and Prostate.

Algorithms↗

HykGene: a hybrid approach for selecting marker genes for phenotype classification using microarray gene expression data.

MOTIVATION: Recent studies have shown that microarray gene expression data are useful for phenotype classification of many diseases. A major problem in this classification is that the number of features (genes) greatly exceeds the number of instances (tissue samples). It has been shown that selecting a small set of informative genes can lead to improved classification accuracy. Many approaches have been proposed for this gene selection problem. Most of the previous gene ranking methods typically select 50-200 top-ranked genes and these genes are often highly correlated. Our goal is to select a small set of non-redundant marker genes that are most relevant for the classification task. RESULTS: To achieve this goal, we developed a novel hybrid approach that combines gene ranking and clustering analysis. In this approach, we first applied feature filtering algorithms to select a set of top-ranked genes, and then applied hierarchical clustering on these genes to generate a dendrogram. Finally, the dendrogram was analyzed by a sweep-line algorithm and marker genes are selected by collapsing dense clusters. Empirical study using three public datasets shows that our approach is capable of selecting relatively few marker genes while offering the same or better leave-one-out cross-validation accuracy compared with approaches that use top-ranked genes directly for classification. AVAILABILITY: The HykGene software is freely available at http://www.cs.dartmouth.edu/~wyh/software.htm CONTACT: wyh@cs.dartmouth.edu SUPPLEMENTARY INFORMATION: Supplementary material is available from http://www.cs.dartmouth.edu/~wyh/hykgene/supplement/index.htm.

Algorithms↗

Classification of oligonucleotide fingerprints: application for microbial community and gene expression analyses.

MOTIVATION: Oligonucleotide fingerprinting of ribosomal RNA genes (OFRG) is a procedure that sorts rRNA gene (rDNA) clones into taxonomic groups through a series of hybridization experiments. The hybridization signals are classified into three discrete values 0, 1 and N, where 0 and 1, respectively, specify negative and positive hybridization events and N designates an uncertain assignment. This study examined various approaches for classifying the values including Bayesian classification with normally distributed signal data, Bayesian classification with the exponentially distributed data, and with gamma distributed data, along with tree-based classification. All classification data were clustered using the unweighted pair group method with arithmetic mean. RESULTS: The performance of each classification/clustering procedure was compared with results from known reference data. Comparisons indicated that the approach using the Bayesian classification with normal densities followed by tree clustering out-performed all others. The paper includes a discussion of how this Bayesian approach may be useful for the analysis of gene expression data.

Algorithms↗

The conversion between ICPC and ICD-10. Requirements for a family of classification systems in the next decade.

The International Classification of Primary Care (ICPC) was developed to order medical concepts into classes that have been chosen for their relevance for family medicine. Family physicians use this to label the most prevalent conditions in their practice as well as their patients' symptoms and complaints. At the same time they do not want to be divorced from the needs of the medical community at large as these are reflected in the most recent medical nomenclature: the Tenth Revision of the International Classification of Diseases (ICD-10). A full conversion between all classes in the first and seventh component of ICPC (n = 646) with those of ICD-10 (n = 1983), with the exception of the chapter on external causes, has been prepared. It was concluded that ICD-10 at the three-digit level cannot function as a core classification for an international primary care system. Of the three-digit ICD-10 rubrics only 120 are compatible on a one to one basis with an ICPC rubric. A total of 114 three-digit ICD-10 rubrics have to be broken open into four-digit rubrics to allow at least one compatible conversion to one or more ICPC rubrics. On this basis only 25% of the diagnostic classes in ICPC can be converted to a single three- or four-digit ICD-10 rubric without lumping. The rest of ICD-10, either on the three- or on the four-digit level, has to be grouped into combinations of classes (lumping) to allow compatible conversion to the remaining rubrics of ICPC. Even though ICD-10 cannot serve as a core classification for primary care, a technical conversion between ICPC and ICD-10 is practically always possible which allows primary care physicians to implement ICD-10 as a contemporary nomenclature within the classification structure of ICPC.

Humans↗

Beyond the clinical classification of azoospermia: opinion.

There is an ongoing debate regarding the appropriate classification of azoospermia. This manuscript reviews the rationale for the current classification of azoospermia and how to effect a change if there is a need to do so. The current classification of azoospermia into obstructive and non-obstructive is because azoospermia due to ejaculatory duct dysfunction and hypogonadotrophism are extremely rare. Though the use of clinical protocols (defective spermatogenesis, genital tract obstruction, ejaculatory duct dysfunction, hypogonadotrophism or pre-testicular, testicular and post-testicular) may be useful in selecting patients for appropriate treatment, no study has shown that they provide a better method of classification of azoospermia than the current approach. There is increasing evidence of a genetic basis of male infertility as well as the evidence that men's fertility potential may be classified genetically. Moreover, genetic disorders may be transmitted to the offspring and their presence in infertile couples may affect treatment outcome. It is therefore useful to explore a genetic classification of azoospermia.

Clinical Protocols↗

Classification differences and maternal mortality: a European study. MOMS Group. MOthers' Mortality and Severe morbidity.

OBJECTIVES: To compare the ways maternal deaths are classified in national statistical offices in Europe and to evaluate the ways classification affects published rates. METHODS: Data on pregnancy-associated deaths were collected in 13 European countries. Cases were classified by a European panel of experts into obstetric or non-obstetric causes. An ICD-9 code (International Classification of Diseases) was attributed to each case. These were compared to the codes given in each country. Correction indices were calculated, giving new estimates of maternal mortality rates. SUBJECTS: There were sufficient data to complete reclassification of 359 or 82% of the 437 cases for which data were collected. RESULTS: Compared with the statistical offices, the European panel attributed more deaths to obstetric causes. The overall number of deaths attributed to obstetric causes increased from 229 to 260. This change was substantial in three countries (P < 0.05) where statistical offices appeared to attribute fewer deaths to obstetric causes. In the other countries, no differences were detected. According to official published data, the aggregated maternal mortality rate for participating countries was 7.7 per 100,000 live births, but it increased to 8.7 after classification by the European panel (P < 0.001). CONCLUSION: The classification of pregnancy-associated deaths differs between European countries. These differences in coding contribute to variations in the reported numbers of maternal deaths and consequently affect maternal mortality rates. Differences in classification of death must be taken into account when comparing maternal mortality rates, as well as differences in obstetric care, underreporting of maternal deaths and other factors such as the age distribution of mothers.

Cause of Death↗

ProtoMap: automatic classification of protein sequences and hierarchy of protein families.

The ProtoMap site offers an exhaustive classification of all proteins in the SWISS-PROT database, into groups of related proteins. The classification is based on analysis of all pairwise similarities among protein sequences. The analysis makes essential use of transitivity to identify homologies among proteins. Within each group of the classification, every two members are either directly or transitively related. However, transitivity is applied restrictively in order to prevent unrelated proteins from clustering together. The classification is done at different levels of confidence, and yields a hierarchical organization of all proteins. The resulting classification splits the protein space into well-defined groups of proteins, which are closely correlated with natural biological families and superfamilies. Many clusters contain protein sequences that are not classified by other databases. The hierarchical organization suggested by our analysis may help in detecting finer subfamilies in families of known proteins. In addition it brings forth interesting relationships between protein families, upon which local maps for the neighborhood of protein families can be sketched. The ProtoMap web server can be accessed at http://www.protomap.cs.huji.ac.il

Computer Graphics↗

SVM-Prot: Web-based support vector machine software for functional classification of a protein from its primary sequence.

Prediction of protein function is of significance in studying biological processes. One approach for function prediction is to classify a protein into functional family. Support vector machine (SVM) is a useful method for such classification, which may involve proteins with diverse sequence distribution. We have developed a web-based software, SVMProt, for SVM classification of a protein into functional family from its primary sequence. SVMProt classification system is trained from representative proteins of a number of functional families and seed proteins of Pfam curated protein families. It currently covers 54 functional families and additional families will be added in the near future. The computed accuracy for protein family classification is found to be in the range of 69.1-99.6%. SVMProt shows a certain degree of capability for the classification of distantly related proteins and homologous proteins of different function and thus may be used as a protein function prediction tool that complements sequence alignment methods. SVMProt can be accessed at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi.

Amino Acid Sequence↗

STCDB: Signal Transduction Classification Database.

The Signal Transduction Classification Database (STCDB) is a database of information relative to the classification of signal transduction. It is based primarily on a proposed classification of signal transduction and it describes each type of characterized signal transduction for which a unique ST number has been provided. This document presents, in its first version, the classification of signal transduction in eukaryotic cells. Approved classifications are available for web browsing at http://www.techfak.uni-bielefeld.de/~ mchen/STCDB.

Animals↗

ACLAME: a CLAssification of Mobile genetic Elements.

The ACLAME database (http://aclame.ulb.ac.be) is a collection and classification of prokaryotic mobile genetic elements (MGEs) from various sources, comprising all known phage genomes, plasmids and transposons. In addition to providing information on the full genomes and genetic entities, it aims to build a comprehensive classification of the functional modules of MGEs at the protein, gene and higher levels. This first version contains a comprehensive classification of 5069 proteins from 119 DNA bacteriophages into over 400 functional families. This classification was produced automatically using TRIBE-MCL, a graph-theory-based Markov clustering algorithm that uses sequence measures as input, and then manually curated. Manual curation was aided by consulting annotations available in public databases retrieved through additional sequence similarity searches using Psi-Blast and Hidden Markov Models. The database is publicly accessible and open to expert volunteers willing to participate in its curation. Its web interface allows browsing as well as querying the classification. The main objectives are to collect and organize in a rational way the complexity inherent to MGEs, to extend and improve the inadequate annotation currently associated with MGEs and to screen known genomes for the validation and discovery of new MGEs.

Bacteriophages↗

TCDB: the Transporter Classification Database for membrane transport protein analyses and information.

The Transporter Classification Database (TCDB) is a web accessible, curated, relational database containing sequence, classification, structural, functional and evolutionary information about transport systems from a variety of living organisms. TCDB is a curated repository for factual information compiled from >10,000 references, encompassing approximately 3000 representative transporters and putative transporters, classified into >400 families. The transporter classification (TC) system is an International Union of Biochemistry and Molecular Biology approved system of nomenclature for transport protein classification. TCDB is freely accessible at http://www.tcdb.org. The web interface provides several different methods for accessing the data, including step-by-step access to hierarchical classification, direct search by sequence or TC number and full-text searching. The functional ontology that underlies the database structure facilitates powerful query searches that yield valuable data in a quick and easy way. The TCDB website also offers several tools specifically designed for analyzing the unique characteristics of transport proteins. TCDB not only provides curated information and a tool for classifying newly identified membrane proteins, but also serves as a genome transporter-annotation tool.

Databases, Protein↗

IgA nephropathy: prognostic classification of end-stage renal failure. L'Association des Néphrologues de l'Est.

BACKGROUND: As yet, no clinical or morphological prognostic classification of IgA nephropathy (IgAN) has been generally accepted. The objective of our study was to quantify the risk of developing end-stage renal failure (ESRF) in IgAN. METHODS: We report a prospective longitudinal study of 210 patients with IgAN confirmed by biopsy between 1987 and 1991. Thirty-two (15.2%) patients were lost to follow-up. Mean follow-up after renal biopsy was 5.6 (SD = 2.6) years. The variables included age, gender, illnesses prior to discovery of IgAN, clinical features at IgAN discovery, 24-h proteinuria, serum creatinine, IgA level, and antihypertensive drugs taken at the time of renal biopsy. Sixty-six renal biopsies were classified by light-microscopy according to Lee's morphological classification. The end-point was ESRF. Survival was analysed by a backward and forward stepwise procedure using the Cox model. The most accurate determination of relative risk was obtained by assessing collinearity of the variables. RESULTS: Thirty-three patients (15.7%) (31 men) developed ESRF. The five univariately significant variables: gender, gross haematuria, 24-h proteinuria (24-P), serum creatinine (SC), and antihypertensive treatment, were candidates for multivariate analysis. The final model used SC (< or = 100, 100-150, > 150 mumol/l), 24-P (< 1, > or = 1 g/day) and gender (female vs male) as independent variables (relative risk and 95% confidence interval were 3.5 (2.1, 5.9) for SC; 5.1 (1.9, 13.6) for 24-P; and 3.5 (0.9, 15) for gender). These estimates were used to construct a prognostic classification of ERSF for men with IgAN: stage 1 (SC < or = 150 mumol/l and 24-P < 1 g/day), stage 2 ((SC > 150 mumol/l and 24-P < 1 g/day) or (SC < or = 150 mumol/l and 24-P > or = 1 g/day)); stage 3 (SC > 150 mumol/l and 24-P > or = 1 g/day). The ESRF-free survival was estimated with Kaplan-Meier analysis. It was 98.5% for stage 1, 86.6% for stage 2, 21.3% for stage 3 (P < 0.001), 7 years after histological diagnosis. The validity of Lee's prognostic classification was confirmed using an independent sample. CONCLUSIONS: These classifications identify groups at high risk of ESRF. Therapeutic studies should focus on these groups.

Adult↗

A UK-wide trial of the Banff classification of renal transplant pathology in routine diagnostic practice.

BACKGROUND: The Banff classification of renal transplant pathology has gained wide support since its introduction in 1993. There have been several studies which have tested its usefulness in the context of research-oriented centres. We sought to evaluate its use in a wider context. METHODS: We recruited pathologists from all but one of the renal transplant centres in the UK. Sections were circulated from 21 selected, 'difficult' cases, in all of which the clinical question was confirmation or exclusion of acute rejection, and in all of which a definite diagnosis had been obvious from the subsequent clinical course. Participants were asked first to diagnose or exclude acute rejection by their usual approach, then to apply the Banff classification. No clinical information was given beyond the time since engraftment, in order to confine the evaluation to the morphological features present in the sections. At the end of the study the subjective impressions of the participants were sought using a structured questionnaire. RESULTS: Using the Banff classification produced no detectable difference in the number of 'correct' diagnoses when compared with a conventional approach, irrespective of whether the 'correct' diagnosis is based on retrospective clinical information or on the consensus opinion of the pathologists involved, and irrespective of where in the Banff schema one applies a 'cut-off' for the diagnosis of acute rejection. However, the reproducibility of the diagnoses was improved. The results suggest that in the Banff classification the best 'cut-off' for the diagnosis of acute rejection is between Banff category 3 and category 4, although in this difficult area we found a large improvement in diagnostic accuracy if input of clinical information occurs. CONCLUSIONS: The improved reproducibility justifies the use of the Banff classification to harmonise approaches between centres, especially in research projects. While there are good reasons also to adopt it in routine diagnostic practice, further refinement is necessary before an improvement in the accuracy of diagnosis can be demonstrated.

Diagnostic Errors↗

Evaluation of the Sørensen diagnostic criteria in the classification of systemic vasculitis.

OBJECTIVES: To evaluate the use of the diagnostic criteria for Wegener's granulomatosis (WG) and microscopic polyangiitis (mPA) proposed by Sørensen et al. in the classification of primary systemic vasculitis (PSV). METHODS: We applied to our cohort of PSV patients the American College of Rheumatology (ACR) criteria for WG, Churg-Strauss syndrome (CSS) and polyarteritis nodosa (PAN), the Chapel Hill Consensus Conference (CHCC) definitions for WG, mPA and CSS, the Hammersmith criteria for CSS and the Sørensen criteria for WG and mPA. RESULTS: Ninety-nine PSV cases were identified. Fifty-six fulfilled criteria for WG (ACR), 60 for PAN (ACR) and 15 for CSS (ACR). Four fulfilled the Hammersmith criteria for CSS. Thirty-nine were defined as mPA (CHCC). Fifty-three patients fulfilled the Sørensen criteria for WG and three for mPA. Five of six patients classified as WG (ACR) who did not meet the Sørensen criteria were excluded by eosinophilia. Six patients who did not fulfil WG (ACR) met the Sørensen criteria for WG. CONCLUSION: The classification of systemic vasculitis is complicated and many cases fulfil more than one set of criteria. The Sørensen criteria for WG is limited by its exclusion of eosinophilia despite reports of an association. We recommend that tissue eosinophilia or peripheral eosinophilia of <1.5x10(9)/l should not exclude a diagnosis of WG. With this modification, the Sørensen criteria for WG may be a useful method of classification, especially in confirming the classification of WG in patients who fulfil both WG (ACR) and mPA (CHCC). Few patients fulfilled the Sørensen criteria for mPA which suggests that they are not of value in classification.

Churg-Strauss Syndrome↗

Distinguishing neurocognitive functions in schizophrenia using partially ordered classification models.

Current methods for statistical analysis of neuropsychological test data in schizophrenia are inherently insufficient for revealing valid cognitive impairment profiles. While neuropsychological tests aim to selectively sample discrete cognitive domains, test performance often requires several cognitive operations or "attributes." Conventional statistical approaches assign each neuropsychological score of interest to a single attribute or "domain" (e.g., attention, executive, etc.), and scores are calculated for each. This can yield misleading information about underlying cognitive impairments. We report findings applying a new method for examining neuropsychological test data in schizophrenia, based on finite partially ordered sets (posets) as classification models. A total of 220 schizophrenia outpatients were administered the Positive and Negative Symptom Scale (PANSS) and a neuropsychological test battery. Selected tests were submitted to cognitive attribute analysis a priori by two neuropsychologists. Applying Bayesian classification methods (posets), each patient was classified with respect to proficiency on the underlying attributes, based upon his or her individual test performance pattern. Twelve cognitive "classes" are described in the sample. Resulting classification models provided detailed "diagnoses" into "attribute-based" profiles of cognitive strength/weakness, mimicking expert clinician judgment. Classification was efficient, requiring few measures to achieve accurate classification. Attributes were associated with PANSS factors in the expected manner (only the negative and cognition factors were associated with the attributes), and a double dissociation was observed in which divergent thinking was selectively associated with negative symptoms, possibly reflecting a manifestation of Kraepelin's hypothesis regarding the impact of volitional disturbances on thought. Using posets for extracting more precise cognitive information from neuropsychological data may reveal more valid cognitive endophenotypes, while dramatically reducing the amount of testing required.

Adult↗

Comparability of sleep disorders diagnoses using DSM-IV and ICSD classifications with adolescents.

STUDY OBJECTIVES: The use of diagnostic classifications to define sleep disorders is still unusual in epidemiological studies assessing the prevalence of sleep disorders in an adolescent population. DESIGN: Cross-sectional study. Representative samples of general populations in United Kingdom, Germany and Italy were selected and interviewed by telephone about their sleep habits, sleep and mental disorder diagnoses. Overall, 724 adolescents ages 15-18 years and 1447 young adults ages 19 to 24 years were interviewed. ICSD-90 and DSM-IV diagnoses provided by the Sleep-EVAL expert system were used for the comparisons. SETTING: N/A. PARTICIPANTS: N/A. INTERVENTIONS: N/A. MEASUREMENTS AND RESULTS: 8% of the adolescents and 12.6% of the young adults had ICSD dyssomnia or sleep disturbances associated with a mental disorder. According to the DSM-IV classification, 5.7% of the adolescents and 8.1% of the young adults had a dyssomnia diagnosis. The comparison between the two classifications show that 73.2% of adolescents and young adults with a DSM-IV dyssomnia diagnosis also had similar ICSD diagnosis. The reverse comparison, ICSD vs. DSM-IV, shows that 39.8% of the subjects with an ICSD diagnosis had a DSM-IV diagnosis. DSM-IV primary insomnia was the most frequent diagnosis. Subjects with such a diagnosis were found in about 10 different ICSD diagnoses, mainly inadequate sleep hygiene, psychophysiological or idiopathic insomnia and insufficient sleep syndrome. CONCLUSIONS: ICSD-90 classification provided higher prevalence of sleep disorder diagnoses than the DSM-IV classification. In adolescents and young adults, DSM-IV primary insomnia is two times more often associated with ICSD inadequate sleep hygiene than with ICSD psychophysiological or idiopathic insomnia.

Adolescent↗

The use of segmental anatomy for an operative classification of liver injuries.

There is no universally accepted standard classification for liver injuries, and thus accurate comparison of reports on the subject is impossible. Most published reports on liver trauma suggest that both morbidity and mortality have a linear correlation with not only the amount of liver parenchyma injured but also with the magnitude of the surgical intervention. The exceptions are retrohepatic vein injuries, which have a mortality independent of associated parenchymal injury but should be integrated in any classification of liver injury. The classification proposed is based on the segmental anatomy of the liver (as defined by Couinaud): Grade I--Injuries requiring no operative intervention, or any injury that requires operative intervention limited to a segment or less. Grade II--Any injury that requires operative intervention involving two or more segments. Grade III--Any injury with an associated juxta- or retrohepatic vein injury. We reviewed all patients with isolated liver injuries during the past 5 years and prospectively reviewed all patients for the 6-month period from January to June 1988 and applied this classification. Sixty-nine patients had grade I injuries, with one death (1%); thirteen patients had grade II injuries, with six deaths (46%); and 13 patients had grade III injuries with nine deaths (69%). Postoperative morbidity was 7% for grade I, 57% for grade II, and 50% for grade III. This study supports the conclusion that morbidity and mortality from liver injury are directly related to the volume of parenchyma involved, and that segmental anatomy can be applied to define this volume. Mortality from retrohepatic vein injuries is independent of associated parenchymal injury. We believe that this proposed classification will provide a simple, reproducible, and accurate means for reporting and comparing liver injuries.

Adult↗