Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Plasma Exosome Metabolomics Reveal Stage-Specific Alterations in Elderly Women With Premetabolic and Metabolic Syndrome.

BACKGROUND: Metabolic syndrome (MetS) is a chronic disorder that poses a major threat to global health. Exosomes have emerged as promising biomarkers for diagnosing and monitoring chronic diseases. However, stage-specific alterations in the exosomal metabolome during MetS development remain poorly understood. This study aimed to characterize the plasma exosomal metabolome and explore candidate exosomal biomarkers in individuals with MetS. METHODS: This study included 20 patients with MetS, 23 individuals with pre-MetS, and 45 healthy controls. Plasma exosomes were isolated and analyzed using untargeted liquid chromatography-mass spectrometry-based metabolomics. Differential metabolites were defined by a dual-threshold, that is, p&#x2009;<&#x2009;0.05 from t-test and variable importance in projection >&#x2009;1 from partial least squares discriminant analysis, with fold change indicating their expression changes. Further, we employed machine learning algorithms to predict MetS status. RESULTS: We identified 27 differential metabolites between the pre-MetS and control groups, mainly enriched in histidine metabolism and the tricarboxylic acid cycle. Of these, 12 metabolites were upregulated, and 15 were downregulated, with 1-methylhistidine and isocitrate playing central regulatory roles. Comparison between the MetS and control groups revealed 45 differentially expressed metabolites, mainly enriched in thiamine metabolism, including 13 upregulated and 32 downregulated. In the pre-MetS group, cladribine showed the highest area under the curve (AUC) (0.743, p&#x2009;<&#x2009;0.05), whereas 3-methylxanthine yielded the largest AUC (0.714, p&#x2009;<&#x2009;0.05) in the MetS group. CONCLUSION: Our study characterized stage-dependent alterations in the plasma exosome-derived metabolome in MetS and suggests that exosomal metabolomics may provide complementary molecular information on early MetS metabolic perturbations.

exosomal features↗

Karyotyping of comparative genomic hybridization human metaphases using kernel nearest-neighbor algorithm.

BACKGROUND: Comparative genomic hybridization (CGH) is a relatively new molecular cytogenetic method that detects chromosomal imbalances. Automatic karyotyping is an important step in CGH analysis because the precise position of the chromosome abnormality must be located and manual karyotyping is tedious and time-consuming. In the past, computer-aided karyotyping was done by using the 4',6-diamidino-2-phenylindole, dihydrochloride (DAPI)-inverse images, which required complex image enhancement procedures. METHODS: An innovative method, kernel nearest-neighbor (K-NN) algorithm, is proposed to accomplish automatic karyotyping. The algorithm is an application of the "kernel approach," which offers an alternative solution to linear learning machines by mapping data into a high dimensional feature space. By implicitly calculating Euclidean or Mahalanobis distance in a high dimensional image feature space, two kinds of K-NN algorithms are obtained. New feature extraction methods concerning multicolor information in CGH images are used for the first time. RESULTS: Experiment results show that the feature extraction method of using multicolor information in CGH images improves greatly the classification success rate. A high success rate of about 91.5% has been achieved, which shows that the K-NN classifier efficiently accomplishes automatic chromosome classification from relatively few samples. CONCLUSIONS: The feature extraction method proposed here and K-NN classifiers offer a promising computerized intelligent system for automatic karyotyping of CGH human chromosomes.

Algorithms↗

Melanie II--a third-generation software package for analysis of two-dimensional electrophoresis images: II. Algorithms.

After two generations of software systems for the analysis of two-dimensional electrophoresis (2-DE) images, a third generation of such software packages has recently emerged that combines state-of-the-art graphical user interfaces with comprehensive spot data analysis capabilities. A key characteristic common to most of these software packages is that many of their tools are implementations of algorithms that resulted from research areas such as image processing, vision, artificial intelligence or machine learning. This article presents the main algorithms implemented in the Melanie II 2-D PAGE software package. The applications of these algorithms, embodied as the feature of the program, are explained in an accompanying article (R. D. Appel et al.; Electrophoresis 1997, 18, 2724-2734).

Algorithms↗

Adaptive classification of two-dimensional gel electrophoretic spot patterns by neural networks and cluster analysis.

The interpretation of two-dimensional gel electrophoresis spot profiles can be facilitated by statistical and machine learning programs. Two different approaches to classification of spot profiles - cluster analysis and neural networks - are discussed. Neural networks for two different model patterns were designed and an algorithm for training of the net for the classification was developed. It was shown that the performance of neural networks is higher compared to cluster and principal component analysis. The possibility of combining both approaches into one process can increase reliability and speed of classification. Artificially created training sets with added random noise can be used for network training. The analysis was applied on the Streptomyces coelicolor developmental two-dimensional (2-D) gel database.

Cluster Analysis↗

Algorithms for inferring haplotypes.

Haplotype phase information in diploid organisms provides valuable information on human evolutionary history and may lead to the development of more efficient strategies to identify genetic variants that increase susceptibility to human diseases. Molecular haplotyping methods are labor-intensive, low-throughput, and very costly. Therefore, algorithms based on formal statistical theories were shown to be very effective and cost-efficient for haplotype reconstruction. This review covers 1) population-based haplotype inference methods: Clark's algorithm, expectation-maximization (EM) algorithm, coalescence-based algorithms (pseudo-Gibbs sampler and perfect/imperfect phylogeny), and partition-ligation algorithm implemented by a fully Bayesian model (Haplotyper) or by EM (PLEM); 2) family-based haplotype inference methods; 3) the handling of genotype scoring uncertainties (i.e., genotyping errors and raw two-dimensional genotype scatterplots) in inferring haplotypes; and 4) haplotype inference methods for pooled DNA samples. The advantages and limitations of each algorithm are discussed. By using simulations based on empirical data on the G6PD gene and TNFRSF5 gene, I demonstrate that different algorithms have different degrees of sensitivity to various extents of population diversities and genotyping error rates. Future development of statistical algorithms for addressing haplotype reconstruction will resort more and more to ideas based on combinatorial mathematics, graphical models, and machine learning, and they will have profound impacts on population genetics and genetic epidemiology with the advent of the human HapMap.

Algorithms↗

Data mining.

Group 14 used data-mining strategies to evaluate a number of issues, including appropriate diagnosis, haplotype estimation, genetic linkage and association studies, and type I error. Methods ranged from exploratory analyses, to machine learning strategies (neural networks, supervised learning, and tree-based methods), to false discovery rate control of type I errors. The general motivations were to find the "story" in the data and to summarize information from a multitude of measures. Several methods illustrated strategies for better trait definition, using summarization of related traits. In the few studies that sought to identify genes for alcoholism, there was little agreement among the different strategies, likely reflecting the complexities of the disease. Nevertheless, Group 14 found that these methods offered strategies to gain a better understanding of the complex pathways by which disease develops.

Alcoholism↗

Cancer-associated molecular signature in the tissue samples of patients with cirrhosis.

Several types of aggressive cancers, including hepatocellular carcinoma (HCC), often arise as a multifocal primary tumor. This suggests a high rate of premalignant changes in noncancerous tissue before the formation of a solitary tumor. Examination of the messenger RNA expression profiles of tissue samples derived from patients with cirrhosis of various etiologies by complementary DNA (cDNA) microarray indicated that they can be grossly separated into two main groups. One group included hepatitis B and C virus infections, hemochromatosis, and Wilson's disease. The other group contained mainly alcoholic liver disease, autoimmune hepatitis, and primary biliary cirrhosis. Analysis of these two groups by the cross-validated leave-one-out machine-learning algorithms revealed a molecular signature containing 556 discriminative genes (P <.001). It is noteworthy that 273 genes in this signature (49%) were also significantly altered in HCC (P <.001). Many genes were previously known to be related to HCC. The 273-gene signature was validated as cancer-associated genes by matching this set to additional independent tumor tissue samples from 163 patients with HCC, 56 patients with lung carcinoma, and 38 patients with breast carcinoma. From this signature, 30 genes were altered most significantly in tissue samples from high-risk individuals with cirrhosis and from patients with HCC. Among them, 12 genes encoded secretory proteins found in sera. In conclusion, we identified a unique gene signature in the tissue samples of patients with cirrhosis, which may be used as candidate markers for diagnosing the early onset of HCC in high-risk populations and may guide new strategies for chemoprevention. Supplementary material for this article can be found on the HEPATOLOGY website (http://interscience.wiley.com/jpages/0270-9139/suppmat/index.html).

Antigens, Neoplasm↗

Selection criteria for drug-like compounds.

The fast identification of quality lead compounds in the pharmaceutical industry through a combination of high throughput synthesis and screening has become more challenging in recent years. Although the number of available compounds for high throughput screening (HTS) has dramatically increased, large-scale random combinatorial libraries have contributed proportionally less to identify novel leads for drug discovery projects. Therefore, the concept of 'drug-likeness' of compound selections has become a focus in recent years. In parallel, the low success rate of converting lead compounds into drugs often due to unfavorable pharmacokinetic parameters has sparked a renewed interest in understanding more clearly what makes a compound drug-like. Various approaches have been devised to address the drug-likeness of molecules employing retrospective analyses of known drug collections as well as attempting to capture 'chemical wisdom' in algorithms. For example, simple property counting schemes, machine learning methods, regression models, and clustering methods have been employed to distinguish between drugs and non-drugs. Here we review computational techniques to address the drug-likeness of compound selections and offer an outlook for the further development of the field.

Algorithms↗

Application of Omics Technologies for Cowpea Improvement.

Cowpea (Vigna unguiculata) is a vital crop for food security, nutrition, and climate resilience in sub-Saharan African and other semi-arid regions. However, its improvement is constrained by the complexity of polygenic traits such as drought tolerance, pest resistance, and seed quality. Conventional breeding, while foundational, remains insufficient to address these challenges at the required pace. Recent advances in multi-omics technologies, including genomics, transcriptomics, proteomics, and metabolomics, provide new opportunities to dissect complex traits, identify candidate genes, and accelerate the development of resilient, high-yielding cultivars. This review presents a critical synthesis of current applications of omics technologies in cowpea improvement, highlighting their contributions to stress adaptation, nutritional enhancement, and precision breeding. The review also examines key technical and institutional constraints limiting the adoption of omics-assisted breeding in cowpea, including inadequate research infrastructure, challenges in multi-omics data integration, and limited technical capacity across breeding programs in sub-Saharan Africa. It discusses strategies to address these barriers through regional collaboration, investment in bioinformatics capacity, and the integration of computational approaches into breeding pipelines. Overall, the review concludes that combining multi-omics technologies with artificial intelligence and machine learning has strong potential to improve genotype-phenotype prediction, accelerate breeding decisions, and support the development of climate-resilient and nutritionally enhanced cowpea cultivars.

cowpea↗

Proteomic signature of human cancer cells.

We assessed proteomic profiles as biomarkers for monitoring cell phenotypes. Protein expression profiles were obtained by fluorescence two-dimensional difference gel electrophoresis (2-D-DIGE), in which quantitative ability is improved by labeling proteins with fluorescent dyes prior to electrophoresis. Integrated protein spot intensities were analyzed by a statistical approach. The proteomic data of two groups of cell lines: (1) adenocarcinoma (AC) cell lines derived from lung, pancreas and colon tissues and (2) lung cancer cell lines with different histological backgrounds, including AC, squamous cell carcinoma and small cell carcinoma, were assessed on the basis of prior biological information. Hierarchical clustering analysis and principal component analysis were used to divide the cell lines into subgroups on the basis of similarities between their protein expression profiles. The majority of cell lines were grouped according to their organ of origin or histological background. A machine-learning algorithm selected 32 protein spots that were responsible for the classification. The results indicate that proteomic data generated by 2-D-DIGE can provide a signature of essential cell phenotypes, suggesting that it might be possible to apply this technique to developing tumor markers that could identify the organ of origin of metastatic tumors and contribute to the differential diagnosis of lung cancer.

Cell Line, Tumor↗

Normalization and analysis of residual variation in two-dimensional gel electrophoresis for quantitative differential proteomics.

Although two-dimensional gel electrophoresis (2-DE) has long been a favorite experimental method to screen proteomes, its reproducibility is seldom analyzed with the assistance of quantitative error models. The lack of models of residual distributions that can be used to assign likelihood to differential expression reflects the difficulty in tackling the combined effect of variability in spot intensity and uncertain recognition of the same spot in different gels. In this report we have analyzed a series of four triplicate two-dimensional gels of chicken embryo heart samples at two distinct development stages to produce such a model of residual distribution. In order to achieve this reference error model, a nonparametric procedure for consistent spot intensity normalization had to be established, and is also reported here. In addition to variability in normalized intensity due to various sources, the residual variation between replicates was observed to be compounded by failure to identify the spot itself (gel alignment). The mixed effect is reflected by variably skewed bimodal density distributions of residuals. The extraction of a global error model that accommodated such distribution was achieved empirically by machine learning, specifically by bootstrapped artificial neural networks. The model described is being used to assign confidence values to observed variations in arbitrary 2-DE gels in order to quantify the degree of over-expression and under-expression of protein spots.

Animals↗

A robust meta-classification strategy for cancer detection from MS data.

We propose a novel method for phenotype identification involving a stringent noise analysis and filtering procedure followed by combining the results of several machine learning tools to produce a robust predictor. We illustrate our method on SELDI-TOF MS prostate cancer data (http://home.ccr.cancer.gov/ncifdaproteomics/ppatterns.asp). Our method identified 11 proteomic biomarkers and gave significantly improved predictions over previous analyses with these data. We were able to distinguish cancer from non-cancer cases with a sensitivity of 90.31% and a specificity of 98.81%. The proposed method can be generalized to multi-phenotype prediction and other types of data (e.g., microarray data).

Biomarkers, Tumor↗

Improving the reliability and throughput of mass spectrometry-based proteomics by spectrum quality filtering.

In contemporary peptide-centric or non-gel proteome studies, vast amounts of peptide fragmentation data are generated of which only a small part leads to peptide or protein identification. This motivates the development and use of a filtering algorithm that removes spectra that contribute little to protein identification. Removal of unidentifiable spectra reduced both the amount of computational and human time spent on analyzing spectra as well as the chances of obtaining false identifications. Thorough testing on various proteome datasets from different instruments showed that the best suggested machine-learning classifier is, on average, able to recognize half of the unidentified spectra as bad spectra. Further analyses showed that several unidentified spectra classified as good were derived from peptides carrying unanticipated amino acid modifications or contained sequence tags that allowed peptide identification using homology searches. The implementation of the classifiers is available under the GNU General Public License at http://www.bioinfo.no/software/spectrumquality.

Adult↗

Significance of structural changes in proteins: expected errors in refined protein structures.

A quantitative expression key to evaluating significant structural differences or induced shifts between any two protein structures is derived. Because crystallography leads to reports of a single (or sometimes dual) position for each atom, the significance of any structural change based on comparison of two structures depends critically on knowing the expected precision of each median atomic position reported, and on extracting it for each atom, from the information provided in the Protein Data Bank and in the publication. The differences between structures of protein molecules that should be identical, and that are normally distributed, indicating that they are not affected by crystal contacts, were analyzed with respect to many potential indicators of structure precision, so as to extract, essentially by "machine learning" principles, a generally applicable expression involving the highest correlates. Eighteen refined crystal structures from the Protein Data Bank, in which there are multiple molecules in the crystallographic asymmetric unit, were selected and compared. The thermal B factor, the connectivity of the atom, and the ratio of the number of reflections to the number of atoms used in refinement correlate best with the magnitude of the positional differences between regions of the structures that otherwise would be expected to be the same. These results are embodied in a six-parameter equation that can be applied to any crystallographically refined structure to estimate the expected uncertainty in position of each atom. Structure change in a macromolecule can thus be referenced to the expected uncertainty in atomic position as reflected in the variance between otherwise identical structures with the observed values of correlated parameters.

Crystallization↗

Normal mode analysis of macromolecular motions in a database framework: developing mode concentration as a useful classifying statistic.

We investigated protein motions using normal modes within a database framework, determining on a large sample the degree to which normal modes anticipate the direction of the observed motion and were useful for motions classification. As a starting point for our analysis, we identified a large number of examples of protein flexibility from a comprehensive set of structural alignments of the proteins in the PDB. Each example consisted of a pair of proteins that were considerably different in structure given their sequence similarity. On each pair, we performed geometric comparisons and adiabatic-mapping interpolations in a high-throughput pipeline, arriving at a final list of 3,814 putative motions and standardized statistics for each. We then computed the normal modes of each motion in this list, determining the linear combination of modes that best approximated the direction of the observed motion. We integrated our new motions and normal mode calculations in the Macromolecular Motions Database, through a new ranking interface at http://molmovdb.org. Based on the normal mode calculations and the interpolations, we identified a new statistic, mode concentration, related to the mathematical concept of information content, which describes the degree to which the direction of the observed motion can be summarized by a few modes. Using this statistic, we were able to determine the fraction of the 3,814 motions where one could anticipate the direction of the actual motion from only a few modes. We also investigated mode concentration in comparison to related statistics on combinations of normal modes and correlated it with quantities characterizing protein flexibility (e.g., maximum backbone displacement or number of mobile atoms). Finally, we evaluated the ability of mode concentration to automatically classify motions into a variety of simple categories (e.g., whether or not they are "fragment-like"), in comparison to motion statistics. This involved the application of decision trees and feature selection (particular machine-learning techniques) to training and testing sets derived from merging the "list" of motions with manually classified ones.

Databases, Protein↗

Prediction of alpha-turns in proteins using PSI-BLAST profiles and secondary structure information.

In this paper a systematic attempt has been made to develop a better method for predicting alpha-turns in proteins. Most of the commonly used approaches in the field of protein structure prediction have been tried in this study, which includes statistical approach "Sequence Coupled Model" and machine learning approaches; i) artificial neural network (ANN); ii) Weka (Waikato Environment for Knowledge Analysis) Classifiers and iii) Parallel Exemplar Based Learning (PEBLS). We have also used multiple sequence alignment obtained from PSIBLAST and secondary structure information predicted by PSIPRED. The training and testing of all methods has been performed on a data set of 193 non-homologous protein X-ray structures using five-fold cross-validation. It has been observed that ANN with multiple sequence alignment and predicted secondary structure information outperforms other methods. Based on our observations we have developed an ANN-based method for predicting alpha-turns in proteins. The main components of the method are two feed-forward back-propagation networks with a single hidden layer. The first sequence-structure network is trained with the multiple sequence alignment in the form of PSI-BLAST-generated position specific scoring matrices. The initial predictions obtained from the first network and PSIPRED predicted secondary structure are used as input to the second structure-structure network to refine the predictions obtained from the first net. The final network yields an overall prediction accuracy of 78.0% and MCC of 0.16. A web server AlphaPred (http://www.imtech.res.in/raghava/alphapred/) has been developed based on this approach.

Amino Acid Sequence↗