Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Predicting carcinoid heart disease with the noisy-threshold classifier.

OBJECTIVE: To predict the development of carcinoid heart disease (CHD), which is a life-threatening complication of certain neuroendocrine tumors. To this end, a novel type of Bayesian classifier, known as the noisy-threshold classifier, is applied. MATERIALS AND METHODS: Fifty-four cases of patients that suffered from a low-grade midgut carcinoid tumor, of which 22 patients developed CHD, were obtained from the Netherlands Cancer Institute (NKI). Eleven attributes that are known at admission have been used to classify whether the patient develops CHD. Classification accuracy and area under the receiver operating characteristics (ROC) curve of the noisy-threshold classifier are compared with those of the naive-Bayes classifier, logistic regression, the decision-tree learning algorithm C4.5, and a decision rule, as formulated by an expert physician. RESULTS: The noisy-threshold classifier showed the best classification accuracy of 72% correctly classified cases, although differences were significant only for logistic regression and C4.5. An area under the ROC curve of 0.66 was attained for the noisy-threshold classifier, and equaled that of the physician's decision-rule. CONCLUSIONS: The noisy-threshold classifier performed favorably to other state-of-the-art classification algorithms, and equally well as a decision-rule that was formulated by the physician. Furthermore, the semantics of the noisy-threshold classifier make it a useful machine learning technique in domains where multiple causes influence a common effect.

Algorithms↗

Long-term depression as a memory process in the cerebellum.

When details of neuronal network structures of the cerebellum were uncovered in the 1960's, a hope emerged that functions of the cerebellum would eventually be explained in terms of operation of the cerebellar neuronal network. While various network models were proposed, involvement of synaptic plasticity in the cerebellar neuronal network as a memory process became a focus of discussion. The characteristic dual inputs to Purkinje cells, one from parallel fibers (axons of granule cells) and the other from climbing fibers, were suggested to represent such synaptic plasticity, and under this assumption, the cerebellar cortex was envisaged as a learning machine for pattern recognition. Despite these theoretical suggestions, earlier efforts to reveal the postulated synaptic plasticity in the cerebellar cortex were unsuccessful. It had then to wait for a decade before long-term depression (LTD) was finally found as its possible substrate. LTD is a long-lasting depression of parallel fiber-to-Purkinje cell transmission that occurs following conjunctive activation of parallel fibers and a climbing fiber both converging onto one and the same Purkinje cell. LTD has now been established by means of various testing methods, and recent efforts have been directed toward its molecular mechanisms. Efforts have also been devoted to demonstrate roles of LTD in motor learning through studies of adaptation of the vestibulo-ocular reflex, adaptive adjustment of hand movement, and more recently eyelid blink conditioned reflex. This article reviews recent efforts to characterize the LTD as a memory process, presumably the major, in the cerebellum.

Animals↗

Prediction of rodent carcinogenicity bioassays from molecular structure using inductive logic programming.

The machine learning program Progol was applied to the problem of forming the structure-activity relationship (SAR) for a set of compounds tested for carcinogenicity in rodent bioassays by the U.S. National Toxicology Program (NTP). Progol is the first inductive logic programming (ILP) algorithm to use a fully relational method for describing chemical structure in SARs, based on using atoms and their bond connectivities. Progol is well suited to forming SARs for carcinogenicity as it is designed to produce easily understandable rules (structural alerts) for sets of noncongeneric compounds. The Progol SAR method was tested by prediction of a set of compounds that have been widely predicted by other SAR methods (the compounds used in the NTP's first round of carcinogenesis predictions). For these compounds no method (human or machine) was significantly more accurate than Progol. Progol was the most accurate method that did not use data from biological tests on rodents (however, the difference in accuracy is not significant). The Progol predictions were based solely on chemical structure and the results of tests for Salmonella mutagenicity. Using the full NTP database, the prediction accuracy of Progol was estimated to be 63% (+/- 3%) using 5-fold cross validation. A set of structural alerts for carcinogenesis was automatically generated and the chemical rationale for them investigated- these structural alerts are statistically independent of the Salmonella mutagenicity. Carcinogenicity is predicted for the compounds used in the NTP's second round of carcinogenesis predictions. The results for prediction of carcinogenesis, taken together with the previous successful applications of predicting mutagenicity in nitroaromatic compounds, and inhibition of angiogenesis by suramin analogues, show that Progol has a role to play in understanding the SARs of cancer-related compounds.

Animals↗

Precision Genomics: A Reality Having Universal Impact in a New Era of Psychiatry - Lessons Learned, Past and Present.

Addiction neuroscience explores the complex interplay between genetic, neurobiological, environmental, and socio-spiritual factors underlying substance and behavioral addictions. Over the past three decades, research in this domain has identified critical molecular and epigenetic mechanisms-particularly those affecting dopaminergic signaling and reward pathways-that contribute to both vulnerability and resilience to addictive behaviors. Central to this understanding is the concept of reward deficiency syndrome (RDS), first introduced by Kenneth Blum, which posits that hypodopaminergic functioning predisposes individuals to seek maladaptive rewards. Advances in neurogenetics, including the identification of key polymorphisms such as the DRD2 A1 allele, have paved the way for precision tools like the genetic addiction risk severity (GARS®) test. This test, alongside pro-dopaminergic nutraceutical interventions like KB220, demonstrates the potential for early detection and individualized treatment of "pre-addiction" risk states. Despite ongoing reliance on opioids for opioid use disorder (OUD), emerging paradigms advocate for dopamine homeostasis through non-addictive, integrative approaches. Furthermore, the integration of whole genome sequencing data can be used for Genome-Wide Association Studies (GWAS), multi-omics, and machine learning into clinical practice holds promise for advancing personalized medicine in addiction treatment. As the field progresses, addressing health equity and improving genomic representation across populations remain critical goals. This evolving framework underscores the importance of leveraging genomic insights to prevent, predict, and personalize interventions for addiction and mental illness at scale.

Disorder↗

From prediction to mechanism: Explainable AI uncovers plasma and CSF proteomic signatures of Alzheimer's disease.

Alzheimer's disease (AD) plasma and cerebrospinal fluid (CSF) proteomics can distinguish AD from cognitively normal controls, but the generalizability of machine learning performance and the recurrence of biological signals across datasets require cautious interpretation. We developed an explainable artificial intelligence framework spanning two fluids and four ADNI proteomic datasets, covering 2082 modality specific samples, all analysed internally within ADNI. Phase 1 analysed plasma using a 119 analyte NULISA and targeted UPENN panel (n&#xa0;=&#xa0;727; 216&#xa0;CE, 511 controls). Phase 2 extended the analysis to CSF using SOMAscan7k, TMT-MS and targeted SET2, with Elecsys A&#x3b2;42, A&#x3b2;40, total tau and p-tau181 as anchor biomarkers. Only SOMAscan was subject-independent relative to Phase 1 plasma; TMT-MS and SET2 overlapped with Phase 1 for 96.0% and 97.7% of subjects and therefore are not independent replication cohorts. Under subject-level splits with fold internal preprocessing, we compared Elastic Net, Explainable Boosting Machines and gradient boosted trees with SHAP-based explanations. Among the candidate pipelines, we selected the pipeline with the highest held-out test ROC AUC for each platform; the selected values were 0.927 in plasma and 0.954-0.973 across the three CSF datasets. Because the same held out test performance was used for pipeline selection and headline reporting, these are optimistically selected single-holdout estimates, not unbiased estimates of generalizable or clinical performance. Explanations identified five recurring biological axes within ADNI: cholinergic (ACHE), tau/14-3-3 (YWHAG, YWHAZ, YWHAB, YWHAE), neuro-axonal (NEFL, NEFH), microglial/complement (CHIT1, SMOC1, CHI3L1, C7, CFH) and synaptic (NPTXR, NPTX2, DLG4, SYT5, VSNL1, ELAVL2). CSF analyses showed synaptic vesicle-cycle enrichment (q&#xa0;=&#xa0;2&#xa0;&#xd7;&#xa0;10-6), and CSF YWHAG correlated strongly with total tau (&#x3c1;&#xa0;=&#xa0;0.87). Cross-fluid directional concordance was modest overall (54-57%) but increased to 73-80% among mapped analyte/protein rows reaching q&#xa0;<&#xa0;0.05 in CSF. These findings provide hypothesis-generating, internally supported evidence within ADNI. Independent external cohorts with locked pipelines are required to evaluate generalizable performance and biological reproducibility; the overlapping TMT-MS and SET2 analyses should not be interpreted as independent replication.

Alzheimer Disease↗

Prediction of protein structural classes using support vector machines.

The support vector machine, a machine-learning method, is used to predict the four structural classes, i.e. mainly alpha, mainly beta, alpha-beta and fss, from the topology-level of CATH protein structure database. For the binary classification, any two structural classes which do not share any secondary structure such as alpha and beta elements could be classified with as high as 90% accuracy. The accuracy, however, will decrease to less than 70% if the structural classes to be classified contain structure elements in common. Our study also shows that the dimensions of feature space 20(2) = 400 (for dipeptide) and 20(3) = 8 000 (for tripeptide) give nearly the same prediction accuracy. Among these 4 structural classes, multi-class classification gives an overall accuracy of about 52%, indicating that the multi-class classification technique in support of vector machines may still need to be further improved in future investigation.

Algorithms↗

Analysis of end-stage renal disease mediated by cuproptosis-related genes.

OBJECTIVE: The complex pathophysiological mechanism of end-stage renal disease (ESRD) has not been fully understood. Cuproptosis is a newly discovered type of programmed cell death. Therefore, this study attempts to clarify the relationship between cuproptosis-related genes (CRGs) and the phenotype of ESRD. MATERIALS AND METHODS: The National Center for Biological Information Gene Expression Omnibus database was applied to obtain the GSE37171 dataset comprising whole-genome microarray analysis of peripheral blood samples. A 3&#xa0;:&#xa0;1 case-control design was employed with 75 ESRD patients and 20 healthy controls who were frequency-matched for age, sex, and ethnicity. Based on differentially expressed genes (DEGs) and genes related to cuproptosis, CRGs were identified. Thereafter, we explored two different subpopulations based on the cuproptosis gene and analyzed their expression and immune infiltration. Genes specific to the CRG cluster were identified through the weighted gene co-expression network analysis algorithm, and the best prediction model was determined and verified by four machine learning methods. RESULTS: The study identified 14 differentially expressed CRGs, among which ATP7B, SLC31A1, LIAS, LIPT1, DLD, MTF1, CDKN2A, DBT, and DLST had relatively high expression levels in the ESRD samples. Compared with the control group, expression levels of FDX1, DLAT, PDHA1, PDHB, and GLS were significantly lower in the ESRD group, and CRGs played a key role in the regulation of immune infiltration in ESRD. Two cuproptosis-related molecular clusters were identified in the ESRD samples. Cluster2 was more correlated with the immune infiltration of ESRD. By analyzing the intersection points between CRG cluster and key genes of ESRD, a total of 888 specific DEGs were identified. Functional differences related to specific DEGs were further explored using gene set variation analysis. Five significant genes (SMC5, USP47, USP53, AGA, and DMXL1) were identified by the support vector machine model as key predictors for ESRD disease risk, achieving an area under the curve (AUC) of 1.00 in internal validation. However, external validation in independent cohorts is required prior to clinical application. Individual gene analysis showed an AUC >&#xa0;0.81 in discriminating ESRD patients from healthy controls, and the expression of all 5 genes in ESRD patients was significantly lower than in the control group. CONCLUSION: This study clarified the relationship between CRGs and the phenotype of ESRD, analyzed their specific roles in the immune microenvironment, and obtained a predictive model, providing new insights for the study of its potential therapeutic targets.

Humans↗

Will my protein crystallize? A sequence-based predictor.

We propose a machine-learning approach to sequence-based prediction of protein crystallizability in which we exploit subtle differences between proteins whose structures were solved by X-ray analysis [or by both X-ray and nuclear magnetic resonance (NMR) spectroscopy] and those proteins whose structures were solved by NMR spectroscopy alone. Because the NMR technique is usually applied on relatively small proteins, sequence length distributions of the X-ray and NMR datasets were adjusted to avoid predictions biased by protein size. As feature space for classification, we used frequencies of mono-, di-, and tripeptides represented by the original 20-letter amino acid alphabet as well as by several reduced alphabets in which amino acids were grouped by their physicochemical and structural properties. The classification algorithm was constructed as a two-layered structure in which the output of primary support vector machine classifiers operating on peptide frequencies was combined by a second-level Naive Bayes classifier. Due to the application of metamethods for cost sensitivity, our method is able to handle real datasets with unbalanced class representation. An overall prediction accuracy of 67% [65% on the positive (crystallizable) and 69% on the negative (noncrystallizable) class] was achieved in a 10-fold cross-validation experiment, indicating that the proposed algorithm may be a valuable tool for more efficient target selection in structural genomics. A Web server for protein crystallizability prediction called SECRET is available at http://webclu.bio.wzw.tum.de:8080/secret.

Amino Acid Sequence↗

Prediction of protein solvent accessibility using support vector machines.

A Support Vector Machine learning system has been trained to predict protein solvent accessibility from the primary structure. Different kernel functions and sliding window sizes have been explored to find how they affect the prediction performance. Using a cut-off threshold of 15% that splits the dataset evenly (an equal number of exposed and buried residues), this method was able to achieve a prediction accuracy of 70.1% for single sequence input and 73.9% for multiple alignment sequence input, respectively. The prediction of three and more states of solvent accessibility was also studied and compared with other methods. The prediction accuracies are better than, or comparable to, those obtained by other methods such as neural networks, Bayesian classification, multiple linear regression, and information theory. In addition, our results further suggest that this system may be combined with other prediction methods to achieve more reliable results, and that the Support Vector Machine method is a very useful tool for biological sequence analysis.

Bayes Theorem↗

The potential of latent semantic analysis for machine grading of clinical case summaries.

OBJECTIVE: This paper introduces latent semantic analysis (LSA), a machine learning method for representing the meaning of words, sentences, and texts. LSA induces a high-dimensional semantic space from reading a very large amount of texts. The meaning of words and texts can be represented as vectors in this space and hence can be compared automatically and objectively. PSYCHOLOGICAL THEORY: A generative theory of the mental lexicon based on LSA is described. The word vectors LSA constructs are context free, and each word, irrespective of how many meanings or senses it has, is represented by a single vector. However, when a word is used in different contexts, context appropriate word senses emerge. CURRENT APPLICATIONS: Several applications of LSA to educational software are described, involving the ability of LSA to quickly compare the content of texts, such as an essay written by a student and a target essay. POTENTIAL MEDICAL APPLICATIONS: An LSA-based software tool is sketched for machine grading of clinical case summaries written by medical students.

Artificial Intelligence↗

Alzheimer's subtypes A supervised, unsupervised, multimodal, multilayered embedded recursive (SUMMER) AI study.

Since Alzheimer's disease (AD) is a heterogeneous disease, different subtypes may have distinct biological, genetic, and clinical characteristics, requiring tailored interventions. While several proposed subtypes of AD exist, there is still no clear consensus on a definitive classification. By leveraging complementary AI approaches, including supervised and unsupervised learning, within a recursive pipeline (SUMMER) that integrates multimodal datasets encompassing MRI measurements, phenotypes, and genetic data, our goal was to generate robust scientific evidence for identifying AD subtypes. Data was downloaded from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database and included neuroimaging data (MRI), genetics (SNPs), clinical diagnosis, and demographics. 1133 European American participants' images, aged 55-95, were included in this study. The analysis was multi-fold, where the first step involved applying an unsupervised application to a subset of the MRI sample (AD + cognitively normal (CN) aged matched groups, 100 men aged 68-85 years, and 76 women aged 68-85 years). The MRI brain gray matter was segmented into 44 regions of interest (ROIs) according to a standard atlas, and 618 features were extracted, including ROI voxel intensity measurements such as minimum, maximum, and histogram variables. Results identified a cluster of subtype AD men and a cluster of subtype AD women that were distinct from the rest of their respective samples. In the next step, the integrity of the identified subtype AD clusters was investigated using the XGBoost supervised machine learning application with genetic features (SNPs, N=36,724) and labels: the identified subtype AD cluster vs. the rest of the sample, stratified by sex. A significant AD subtype men model (accuracy=0.85, F1=0.72, AUC=0.83) and a significant women AD subtype model (accuracy=0.81, F1=0.81, AUC=0.81) were built, confirming the homogeneity of the isolated AD subtype clusters. Discriminative biomarkers were extracted from the significant models, including selected ROIs and SNPs. Finally, the subtype models were tested on an unseen subset of ADNI data. The genetic-based models identified clusters of AD subtype participants consisting of 34% of the men AD group and 47% of the women AD group. Phenotypic analysis indicates that lower body weight was associated with the women's AD subtype. Complex diseases like AD demand a sophisticated, multimodal approach for precise diagnosis. Effectively identifying disease subtypes enhances the potential for personalized treatment, ultimately improving patient outcomes.

Journal Article↗

Parameter selection for and implementation of a web-based decision-support tool to predict extubation outcome in premature infants.

BACKGROUND: Approximately 30% of intubated preterm infants with respiratory distress syndrome (RDS) will fail attempted extubation, requiring reintubation and mechanical ventilation. Although ventilator technology and monitoring of premature infants have improved over time, optimal extubation remains challenging. Furthermore, extubation decisions for premature infants require complex informational processing, techniques implicitly learned through clinical practice. Computer-aided decision-support tools would benefit inexperienced clinicians, especially during peak neonatal intensive care unit (NICU) census. METHODS: A five-step procedure was developed to identify predictive variables. Clinical expert (CE) thought processes comprised one model. Variables from that model were used to develop two mathematical models for the decision-support tool: an artificial neural network (ANN) and a multivariate logistic regression model (MLR). The ranking of the variables in the three models was compared using the Wilcoxon Signed Rank Test. The best performing model was used in a web-based decision-support tool with a user interface implemented in Hypertext Markup Language (HTML) and the mathematical model employing the ANN. RESULTS: CEs identified 51 potentially predictive variables for extubation decisions for an infant on mechanical ventilation. Comparisons of the three models showed a significant difference between the ANN and the CE (p = 0.0006). Of the original 51 potentially predictive variables, the 13 most predictive variables were used to develop an ANN as a web-based decision-tool. The ANN processes user-provided data and returns the prediction 0-1 score and a novelty index. The user then selects the most appropriate threshold for categorizing the prediction as a success or failure. Furthermore, the novelty index, indicating the similarity of the test case to the training case, allows the user to assess the confidence level of the prediction with regard to how much the new data differ from the data originally used for the development of the prediction tool. CONCLUSION: State-of-the-art, machine-learning methods can be employed for the development of sophisticated tools to aid clinicians' decisions. We identified numerous variables considered relevant for extubation decisions for mechanically ventilated premature infants with RDS. We then developed a web-based decision-support tool for clinicians which can be made widely available and potentially improve patient care world wide.

Birth Weight↗

Robust Bayesian clustering.

A new variational Bayesian learning algorithm for Student-t mixture models is introduced. This algorithm leads to (i) robust density estimation, (ii) robust clustering and (iii) robust automatic model selection. Gaussian mixture models are learning machines which are based on a divide-and-conquer approach. They are commonly used for density estimation and clustering tasks, but are sensitive to outliers. The Student-t distribution has heavier tails than the Gaussian distribution and is therefore less sensitive to any departure of the empirical distribution from Gaussianity. As a consequence, the Student-t distribution is suitable for constructing robust mixture models. In this work, we formalize the Bayesian Student-t mixture model as a latent variable model in a different way from Svensén and Bishop [Svensén, M., & Bishop, C. M. (2005). Robust Bayesian mixture modelling. Neurocomputing, 64, 235-252]. The main difference resides in the fact that it is not necessary to assume a factorized approximation of the posterior distribution on the latent indicator variables and the latent scale variables in order to obtain a tractable solution. Not neglecting the correlations between these unobserved random variables leads to a Bayesian model having an increased robustness. Furthermore, it is expected that the lower bound on the log-evidence is tighter. Based on this bound, the model complexity, i.e. the number of components in the mixture, can be inferred with a higher confidence.

Algorithms↗

Computer-automated dementia screening using a touch-tone telephone.

BACKGROUND: This study investigated the sensitivity and specificity of a computer-automated telephone system to evaluate cognitive impairment in elderly callers to identify signs of early dementia. METHODS: The Clinical Dementia Rating Scale was used to assess 155 subjects aged 56 to 93 years (n = 74, 27, 42, and 12, with a Clinical Dementia Rating Scale score of 0, 0.5, 1, and 2, respectively). These subjects performed a battery of tests administered by an interactive voice response system using standard Touch-Tone telephones. Seventy-four collateral informants also completed an interactive voice response version of the Symptoms of Dementia Screener. RESULTS: Sixteen cognitively impaired subjects were unable to complete the telephone call. Performances on 6 of 8 tasks were significantly influenced by Clinical Dementia Rating Scale status. The mean (SD) call length was 12 minutes 27 seconds (2 minutes 32 seconds). A subsample (n = 116) was analyzed using machine-learning methods, producing a scoring algorithm that combined performances across 4 tasks. Results indicated a potential sensitivity of 82.0% and specificity of 85.5%. The scoring model generalized to a validation subsample (n = 39), producing 85.0% sensitivity and 78.9% specificity. The kappa agreement between predicted and actual group membership was 0.64 (P<.001). Of the 16 subjects unable to complete the call, 11 provided sufficient information to permit us to classify them as impaired. Standard scoring of the interactive voice response-administered Symptoms of Dementia Screener (completed by informants) produced a screening sensitivity of 63.5% and 100% specificity. A lower criterion found a 90.4% sensitivity, without lowering specificity. CONCLUSIONS: Computer-automated telephone screening for early dementia using either informant or direct assessment is feasible. Such systems could provide wide-scale, cost-effective screening, education, and referral services to patients and caregivers.

Aged↗

Computerized breast cancer diagnosis and prognosis from fine-needle aspirates.

OBJECTIVE: To use digital image analysis and machine learning to (1) improve breast mass diagnosis based on fine-needle aspirates and (2) improve breast cancer prognostic estimations. DESIGN: An interactive computer system evaluates, diagnoses, and determines prognosis based on cytologic features derived from a digital scan of fine-needle aspirate slides. SETTING: The University of Wisconsin (Madison) Departments of Computer Science and Surgery and the University of Wisconsin Hospital and Clinics. PATIENTS: Five hundred sixty-nine consecutive patients (212 with cancer and 357 with benign masses) provided the data for the diagnostic algorithm, and an additional 118 (31 with malignant masses and 87 with benign masses) consecutive, new patients tested the algorithm. One hundred ninety of these patients with invasive cancer and without distant metastases were used for prognosis. INTERVENTIONS: Surgical biopsy specimens were taken from all cancers and some benign masses. The remaining cytologically benign masses were followed up for a year and surgical biopsy specimens were taken if they changed in size or character. Patients with cancer received standard treatment. OUTCOME MEASURES: Cross validation was used to project the accuracy of the diagnostic algorithm and to determine the importance of prognostic features. In addition, the mean errors were calculated between the actual times of distant disease occurrence and the times predicted using various prognostic features. Statistical analyses were also done. RESULTS: The predicted diagnostic accuracy was 97% and the actual diagnostic accuracy on 118 new samples was 100%. Tumor size and lymph node status were weak prognosticators compared with nuclear features, in particular those measuring nuclear size. Compared with the actual time for recurrence, the mean error of predicted times for recurrence with the nuclear features was 17.9 months and was 20.1 months with tumor size and lymph node status (P = .11). CONCLUSION: Computer technology will improve breast fine-needle aspirate accuracy and prognostic estimations.

Biopsy, Needle↗

Decision-tree approach to the immunophenotype-based prognosis of the B-cell chronic lymphocytic leukemia.

Use of a nonlinear prediction method, such as machine learning, is a valuable choice in predicting progression rate of disease when applied to the highly variable and correlated biological data such as those in patients with chronic lymphocytic leukemia (CLL). In this work, decision-tree approach to cell phenotype-based prognosis of CLL was adopted. The panel of 33 (32 different phenotypic features and serum concentration of sCD23) parameters was simultaneously presented to the C4.5 decision tree which extracted the most informative of them and subsequently performed classification of CLL patients against the modified Rai staging system. It has been shown that substantial correlation between the percentage of expression of the CD23 molecule on CD19+ B-cells, the level of sCD23, the percentage of CD45RA+, and the absolute number of CD4CD45RA+RO+ T-cells and the clinical stages, exists. The prediction vector, composed of their concatenated values, was able to correctly associate 83% of the cases in the low-risk group (Rai stage 0), 100% of the cases in the intermediate-risk group (Rai stage I and II), and 89% of the cases in the high-risk group (Rai stage III and IV) of CLL patients. Predictivity of this vector was 100%, 95%, and 89%, respectively. In conclusion, from the described analysis, it may be inferred that two processes play important roles in the progression rate of CLL: 1.deregulated function of the CD23 gene in B-cells accompanied by the appearance of its cleaved product sCD23 in the sera; and 2. functionally impaired and imbalanced CD4 T-cell subpopulations found in the peripheral blood of CLL patients.

Aged↗

Pharmacophore features of potential drugs.

Drug discovery efforts rely increasingly on the identification of quality lead compounds through high-throughput synthesis and screening. However, large-scale random libraries have yielded only a low number of quality lead molecules. To address this shortcoming researchers have paid more attention to the concept of "drug-likeness" of molecules in combinatorial and screening libraries. Database profiling and analysis methods have been employed to identify the structural features of known drug molecules. Neural networks and machine learning methods help to distinguish between drugs and nondrugs. More recently, database-independent pharmacophore filters have been introduced that provide simple intuitive rules to classify potential drugs.

Combinatorial Chemistry Techniques↗

Discovery of new rheumatoid arthritis biomarkers using the surface-enhanced laser desorption/ionization time-of-flight mass spectrometry ProteinChip approach.

OBJECTIVE: To identify serum protein biomarkers specific for rheumatoid arthritis (RA), using surface-enhanced laser desorption/ionization time-of-flight mass spectrometry (SELDI-TOF-MS) technology. METHODS: A total of 103 serum samples from patients and healthy controls were analyzed. Thirty-four of the patients had a diagnosis of RA, based on the American College of Rheumatology criteria. The inflammation control group comprised 20 patients with psoriatic arthritis (PsA), 9 with asthma, and 10 with Crohn's disease. The noninflammation control group comprised 14 patients with knee osteoarthritis and 16 healthy control subjects. Serum protein profiles were obtained by SELDI-TOF-MS and compared in order to identify new biomarkers specific for RA. Data were analyzed by a machine learning algorithm called decision tree boosting, according to different preprocessing steps. RESULTS: The most discriminative mass/charge (m/z) values serving as potential biomarkers for RA were identified on arrays for both patients with RA versus controls and patients with RA versus patients with PsA. From among several candidates, the following peaks were highlighted: m/z values of 2,924 (RA versus controls on H4 arrays), 10,832 and 11,632 (RA versus controls on CM10 arrays), 4,824 (RA versus PsA on H4 arrays), and 4,666 (RA versus PsA on CM10 arrays). Positive results of proteomic analysis were associated with positive results of the anti-cyclic citrullinated peptide test. Our observations suggested that the 10,832 peak could represent myeloid-related protein 8. CONCLUSION: SELDI-TOF-MS technology allows rapid analysis of many serum samples, and use of decision tree boosting analysis as the main statistical method allowed us to propose a pattern of protein peaks specific for RA.

Adult↗