Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

CpGene: a web application for epigenetic signature identification from DNA methylation arrays.

MOTIVATION: DNA methylation (DNAme) is the best studied epigenetic mechanism that plays pivotal role in tissue differentiation and epigenetic disruption has been correlated to diverse disease types (e.g. cancer, metabolic disorders). While various DNAme array platforms have been discovered, data analysis remains a challenging task which often requires in-depth bioinformatic expertise. Here, we developed a user-friendly web-based application for data analysis and visualization that accommodates users ranging from early-career basic/translational researchers to experienced bioinformaticians. RESULTS: CpGene is a web application for analyzing DNA methylation array data. It supports Illumina 450K, EPIC, and EPICv2 methylation array platforms and processes .idat files with integrated preprocessing, normalization, and quality control. Biomarker discovery is available through either classic differential methylation point analysis or machine learning-based feature selection as well as gene enrichment analysis. Results are summarized with clear visualizations, to aid interpretation. By combining these functions in a unified interface, CpGene streamlines methylation analysis and helps identify CpG sites and genes with biological and clinical relevance. AVAILABILITY AND IMPLEMENTATION: CpGene is openly accessible as a web service through http://cpgene.duckdns.org:8001/ and it's source code is available on https://github.com/kostaslazaros/cpgenene.

DNA Methylation↗

RankGene: identification of diagnostic genes based on expression data.

RankGene is a program for analyzing gene expression data and computing diagnostic genes based on their predictive power in distinguishing between different types of samples. The program integrates into one system a variety of popular ranking criteria, ranging from the traditional t-statistic to one-dimensional support vector machines. This flexibility makes RankGene a useful tool in gene expression analysis and feature selection.

Algorithms↗

BagBoosting for tumor classification with gene expression data.

MOTIVATION: Microarray experiments are expected to contribute significantly to the progress in cancer treatment by enabling a precise and early diagnosis. They create a need for class prediction tools, which can deal with a large number of highly correlated input variables, perform feature selection and provide class probability estimates that serve as a quantification of the predictive uncertainty. A very promising solution is to combine the two ensemble schemes bagging and boosting to a novel algorithm called BagBoosting. RESULTS: When bagging is used as a module in boosting, the resulting classifier consistently improves the predictive performance and the probability estimates of both bagging and boosting on real and simulated gene expression data. This quasi-guaranteed improvement can be obtained by simply making a bigger computing effort. The advantageous predictive potential is also confirmed by comparing BagBoosting to several established class prediction tools for microarray data. AVAILABILITY: Software for the modified boosting algorithms, for benchmark studies and for the simulation of microarray data are available as an R package under GNU public license at http://stat.ethz.ch/~dettling/bagboost.html.

Algorithms↗

How many samples are needed to build a classifier: a general sequential approach.

MOTIVATION: The standard paradigm for a classifier design is to obtain a sample of feature-label pairs and then to apply a classification rule to derive a classifier from the sample data. Typically in laboratory situations the sample size is limited by cost, time or availability of sample material. Thus, an investigator may wish to consider a sequential approach in which there is a sufficient number of patients to train a classifier in order to make a sound decision for diagnosis while at the same time keeping the number of patients as small as possible to make the studies affordable. RESULTS: A sequential classification procedure is studied via the martingale central limit theorem. It updates the classification rule at each step and provides stopping criteria to ensure with a certain confidence that at stopping a future subject will have misclassification probability smaller than a predetermined threshold. Simulation studies and applications to microarray data analysis are provided. The procedure possesses several attractive properties: (1) it updates the classification rule sequentially and thus does not rely on distributions of primary measurements from other studies; (2) it assesses the stopping criteria at each sequential step and thus can substantially reduce cost via early stopping; and (3) it is not restricted to any particular classification rule and therefore applies to any parametric or non-parametric method, including feature selection or extraction. AVAILABILITY: R-code for the sequential stopping rule is available at http://stat.tamu.edu/~wfu/microarray/sequential/R-code.html

Algorithms↗

A causal inference approach for constructing transcriptional regulatory networks.

MOTIVATION: Transcriptional regulatory networks specify the interactions among regulatory genes and between regulatory genes and their target genes. Discovering transcriptional regulatory networks helps us to understand the underlying mechanism of complex cellular processes and responses. METHOD: This paper describes a causal inference approach for constructing transcriptional regulatory networks using gene expression data, promoter sequences and information on transcription factor (TF) binding sites. The method first identifies active TFs in each individual experiment using a feature selection approach. TFs are viewed as "treatments" and gene expression levels as "responses". For every TF and gene pair, a marginal structural model is built to estimate the causal effect of the TF on the expression level of the gene. The model parameters can be estimated using the G-computation procedure or the IPTW estimator. The P-value associated with the causal parameter in each of these models is used to measure how strongly a TF regulates a gene. These results are further used to infer the overall regulatory network structures. RESULTS: Our analysis of yeast data suggests that the method is capable of identifying significant transcriptional regulatory interactions and the corresponding regulatory networks. AVAILABILITY: The software is under development.

Algorithms↗

COMPAM :visualization of combining pairwise alignments for multiple genomes.

UNLABELLED: COMPAM is a tool for visualizing relationships among multiple whole genomes by combining all pairwise genome alignments. It displays shared conserved regions (blocks) and where these blocks occur (edges) as block relation graphs which can be explored interactively. An unannotated genome, e.g. can then be explored using information from well-annotated genomes, COG-based genome annotation and genes. COMPAM can run either as a stand-alone application or through an applet that is provided as service to PLATCOM, a toolset for whole genome comparative analysis, where a wide variety of genomes can be easily selected. Features provided by COMPAM include the ability to export genome relationship information into file formats that can be used by other existing tools. AVAILABILITY: http://bio.informatics.indiana.edu/projects/compam/

Algorithms↗

Predicting the prognosis of breast cancer by integrating clinical and microarray data with Bayesian networks.

MOTIVATION: Clinical data, such as patient history, laboratory analysis, ultrasound parameters--which are the basis of day-to-day clinical decision support--are often underused to guide the clinical management of cancer in the presence of microarray data. We propose a strategy based on Bayesian networks to treat clinical and microarray data on an equal footing. The main advantage of this probabilistic model is that it allows to integrate these data sources in several ways and that it allows to investigate and understand the model structure and parameters. Furthermore using the concept of a Markov Blanket we can identify all the variables that shield off the class variable from the influence of the remaining network. Therefore Bayesian networks automatically perform feature selection by identifying the (in)dependency relationships with the class variable. RESULTS: We evaluated three methods for integrating clinical and microarray data: decision integration, partial integration and full integration and used them to classify publicly available data on breast cancer patients into a poor and a good prognosis group. The partial integration method is most promising and has an independent test set area under the ROC curve of 0.845. After choosing an operating point the classification performance is better than frequently used indices.

Bayes Theorem↗

Integration of gel-based proteome data with pProRep.

UNLABELLED: pProRep is a web application integrating electrophoretic and mass spectral data from proteome analyses into a relational database. The graphical web-interface allows users to upload, analyse and share experimental proteome data. It offers researchers the possibility to query all previously analysed datasets and can visualize selected features, such as the presence of a certain set of ions in a peptide mass spectrum, on the level of the two-dimensional gel. AVAILABILITY: The pProRep package and instructions for its use can be downloaded from http://www.ptools.ua.ac.be/pProRep. The application requires a web server that runs PHP 5 (http://www.php.net) and MySQL. Some (non-essential) extensions need additional freely available libraries: details are described in the installation instructions.

Computational Biology↗

Model-based multifacet clustering with high-dimensional omics applications.

High-dimensional omics data often contain intricate and multifaceted information, resulting in the coexistence of multiple plausible sample partitions based on different subsets of selected features. Conventional clustering methods typically yield only one clustering solution, limiting their capacity to fully capture all facets of cluster structures in high-dimensional data. To address this challenge, we propose a model-based multifacet clustering (MFClust) method based on a mixture of Gaussian mixture models, where the former mixture achieves facet assignment for gene features and the latter mixture determines cluster assignment of samples. We demonstrate superior facet and cluster assignment accuracy of MFClust through simulation studies. The proposed method is applied to three transcriptomic applications from postmortem brain and lung disease studies. The result captures multifacet clustering structures associated with critical clinical variables and provides intriguing biological insights for further hypothesis generation and discovery.

Humans↗

Prescribing rates for psychotropic medication amongst east London general practices: low rates where Asian populations are greatest.

OBJECTIVES: The aim of this study was to examine the contribution of Asian ethnicity to the variation in rates of practice prescribing for antidepressant and anxiolytic medication, taking into account other population and practice organizational factors. METHODS: A practice-based cross-sectional survey was carried out of the prescribing of antidepressants and anxiolytics (daily defined dosages) in 164 general practices. The study was set in East London and the City Health Authority, which includes the multiethnic inner London boroughs of Hackney, Tower Hamlets, Newham and the City of London. The main outcome measures were the annual prescribing rates for each group of drugs, calculated as the total annual daily defined dosages divided by the practice population, and the ratio of antidepressant/ anxiolytic annual prescribing rates. RESULTS: Prescribing rates for antidepressants showed a 25-fold variation between practices; this was greater for anxiolytics. The median annual prescribing rate for all antidepressants combined was 4.13 (interquartile range 2.50-5.88). For all anxiolytics and hypnotics combined the median annual prescribing rate was 3.55 (interquartile range 1.71-6.36). Univariate analysis showed that Asian ethnicity alone accounted for 28% of the variation in antidepressant prescribing and 20.5% of the variation in the anxiolytic prescribing. A backwards multiple regression model using 10 explanatory practice and population variables accounted for 47.7% of the variance in antidepressant prescribing and 34% of the variance in the anxiolytic prescribing. CONCLUSION: In practices where the proportion of Asian patients is high, both antidepressant and anxiolytic prescribing is low. This is important for understanding interpractice prescribing variation and for setting levels of drug budgets. This study confirms that the low rates of non-psychotic disorders presented by Asian populations is not a selective feature of access to secondary care, but is evident in the prescribing behaviour of GPs. Uncertainty remains as to how much this is due to a lower prevalence rate, "culture-bound syndromes" or practical difficulties in diagnosis and management within the general practice setting.

Anxiety Disorders↗

Dietary cadmium, zinc and copper: effects on chick lung morphology and elastin cross-linking.

Day-old White Leghorn cockerels were divided into seven dietary groups and fed one of the following diets: 1) a casein-based basal diet; 2) a casein-based diet supplemented with 10 mg/kg cadmium, 3) 100 mg/kg cadmium, 4) or 800 mg/kg zinc; 5) a casein-based diet pair-fed to the 100 mg/kg Cd group; 6) a spray-dried nonfat milk-based diet with no added copper, or 7) a spray-dried nonfat milk-based diet supplemented with 5 mg/kg copper. At termination (5 weeks), the birds were killed, and the effects of the diets on selected features of lung composition and morphology were assessed. Body weights were reduced in the 100 mg/kg Cd, pair-fed, and Cu-deficient groups when compared to their controls (casein-based or milk-based copper-supplemented diets). There were no differences in lung weights (expressed relative to metabolic body size) among the groups, although copper deficiency did result in a slight decrease in the dry to wet weight ratio of lung. Lung elastin content and the desmosine content in elastin were significantly lower in the Cu-deficient group and tended to be lower in the group fed 800 ppm Zn. Significant alterations (enlargement of the tertiary bronchial lumen) in morphology were also observed in lungs from both the 100 mg/kg Cd and Cu-deficient groups. Alteration in lung morphology observed in the 100 mg/kg Cd group could not be explained by changes in the elastin content of lung.

Animals↗

Serum protein MALDI profiling to distinguish upper aerodigestive tract cancer patients from control subjects.

BACKGROUND: There are no reliable blood markers for the early detection and monitoring of aerodigestive tract tumors. Recent studies have suggested that serum protein patterns may be able to distinguish cancer patients from control subjects. METHODS: We used matrix-assisted laser desorption and ionization (MALDI) mass spectroscopy to obtain serum protein patterns from patients with head and neck cancer (n = 99) or lung cancer (n = 92) and from control subjects (n = 143) at risk for the development of these cancers. From the mass spectra, we predicted the cancer status of patients using a simple classification procedure based on a t test feature selection and linear discriminant analysis (LDA). We cross-validated the data with 200 random data simulations to establish a range of the LDA tuning parameter, which was used to construct receiver operating characteristic (ROC) curves. RESULTS: Average total protein levels were higher in case patients than in control subjects, although the differences were not statistically significant. Ten individual m/z peaks, from 5 to 111 kd, appeared frequently in head and neck cancer patients but not in control subjects. Using the 45 top predictors, selected by spectral mass and LDA, we observed that ROC curves differed from those expected under the null hypothesis, suggesting that spectral profiles from the sera of patients with head and neck cancer statistically significantly differed from the sera of control subjects. The model developed on head and neck cancer patients could also be used to identify patients with lung cancer. CONCLUSIONS: The pattern of protein spectra in total serum reliably distinguished cancer case patients from control subjects. Incorporation of MALDI assays into prospective longitudinal trials to assess the true predictive values of protein spectra in cancer detection is needed.

Adult↗

Operon prediction using both genome-specific and general genomic information.

We have carried out a systematic analysis of the contribution of a set of selected features that include three new features to the accuracy of operon prediction. Our analyses have led to a number of new insights about operon prediction, including that (i) different features have different levels of discerning power when used on adjacent gene pairs with different ranges of intergenic distance, (ii) certain features are universally useful for operon prediction while others are more genome-specific and (iii) the prediction reliability of operons is dependent on intergenic distances. Based on these new insights, our newly developed operon-prediction program achieves more accurate operon prediction than the previous ones, and it uses features that are most readily available from genomic sequences. Our prediction results indicate that our (non-linear) decision tree-based classifier can predict operons in a prokaryotic genome very accurately when a substantial number of operons in the genome are already known. For example, the prediction accuracy of our program can reach 90.2 and 93.7% on Bacillus subtilis and Escherichia coli genomes, respectively. When no such information is available, our (linear) logistic function-based classifier can reach the prediction accuracy at 84.6 and 83.3% for E.coli and B.subtilis, respectively.

Bacillus subtilis↗

PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.

Peroxisomes are essential organelles of eukaryotic origin, ubiquitously distributed in cells and organisms, playing key roles in lipid and antioxidant metabolism. Loss or malfunction of peroxisomes causes more than 20 fatal inherited conditions. We have created a peroxisomal database (http://www.peroxisomeDB.org) that includes the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae, by gathering, updating and integrating the available genetic and functional information on peroxisomal genes. PeroxisomeDB is structured in interrelated sections 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases', that include hyperlinks to selected features of NCBI, ENSEMBL and UCSC databases. We have designed graphical depictions of the main peroxisomal metabolic routes and have included updated flow charts for diagnosis. Precomputed BLAST, PSI-BLAST, multiple sequence alignment (MUSCLE) and phylogenetic trees are provided to assist in direct multispecies comparison to study evolutionary conserved functions and pathways. Highlights of the PeroxisomeDB include new tools developed for facilitating (i) identification of novel peroxisomal proteins, by means of identifying proteins carrying peroxisome targeting signal (PTS) motifs, (ii) detection of peroxisomes in silico, particularly useful for screening the deluge of newly sequenced genomes. PeroxisomeDB should contribute to the systematic characterization of the peroxisomal proteome and facilitate system biology approaches on the organelle.

Animals↗

What is 'nephrosclerosis'? lessons from the US, Japan, and Mexico.

BACKGROUND: Selected features of 'nephrosclerosis' can be quantitated morphometrically in renal histology at autopsy. Specimens are available from Japan, Mexico, and the US (blacks and whites). METHODS: Autopsies of men and women aged 15-79 years provided renal samples for paraffin sectioning. These were assembled in New Orleans for objective evaluation after standardized staining with PAS-Alcian blue and interspersion with each other. Obsolescence of glomeruli, interstitial fibrosis, fibroplastic intimal thickenings of arteries, and arteriolar hyalinization, as operationally defined, were measured by objective morphometry. RESULTS: Obsolescence of glomeruli and interstitial fibrosis displayed the expected correlation with arterial intimal fibroplasia, but failed to confirm any direct association with arteriolar hyalinization. Some of the variation of 'nephrosclerosis', within and between populations, cannot be fully explained by microvascular defects. CONCLUSIONS: Arterial intimal fibroplasia appeared to promote 'nephrosclerosis', in the sense of fibrous replacement of atrophied nephrons, but arteriolar hyalinization did not. Hyaline deposits in arterioles may offer little or no threat to the integrity of the affected nephrons. 'Nephrosclerosis' appears to be multifactorial; it may be, in part, a consequence of fibroplasia in microscopic arteries causing ischaemic injury to scattered nephrons, but may also be a confluence of basically separate conditions, only some of which are known.

Adolescent↗

Spontaneous direct carotid-cavernous fistula in childhood.

We report the occurrence, surgical treatment and long-term follow-up of a spontaneous, direct carotid-cavernous fistula in a child. It is the third angiographically documented, spontaneously occurring fistula to be reported in this age group and the first to arise directly from the internal carotid artery based on our review of the literature. Although our patient required fistula closure, other reported fistulae were nonprogressive and did not require treatment. The communication between the internal carotid artery and the cavernous sinus was left sided while the contralateral eye was proptotic. The clinical features, selected hemodynamic characteristics, and treatment of carotid-cavernous fistulas are reviewed.

Arteriovenous Fistula↗

Automated decision tree classification of corneal shape.

PURPOSE: The volume and complexity of data produced during videokeratography examinations present a challenge of interpretation. As a consequence, results are often analyzed qualitatively by subjective pattern recognition or reduced to comparisons of summary indices. We describe the application of decision tree induction, an automated machine learning classification method, to discriminate between normal and keratoconic corneal shapes in an objective and quantitative way. We then compared this method with other known classification methods. METHODS: The corneal surface was modeled with a seventh-order Zernike polynomial for 132 normal eyes of 92 subjects and 112 eyes of 71 subjects diagnosed with keratoconus. A decision tree classifier was induced using the C4.5 algorithm, and its classification performance was compared with the modified Rabinowitz-McDonnell index, Schwiegerling's Z3 index (Z3), Keratoconus Prediction Index (KPI), KISA%, and Cone Location and Magnitude Index using recommended classification thresholds for each method. We also evaluated the area under the receiver operator characteristic (ROC) curve for each classification method. RESULTS: Our decision tree classifier performed equal to or better than the other classifiers tested: accuracy was 92% and the area under the ROC curve was 0.97. Our decision tree classifier reduced the information needed to distinguish between normal and keratoconus eyes using four of 36 Zernike polynomial coefficients. The four surface features selected as classification attributes by the decision tree method were inferior elevation, greater sagittal depth, oblique toricity, and trefoil. CONCLUSION: Automated decision tree classification of corneal shape through Zernike polynomials is an accurate quantitative method of classification that is interpretable and can be generated from any instrument platform capable of raw elevation data output. This method of pattern classification is extendable to other classification problems.

Cornea↗

Proteomic Immune Signatures of Severe HIV-Associated Tuberculosis in Sub-Saharan Africa: A Prospective, Multicenter Analysis From Uganda.

OBJECTIVES: Severe tuberculosis (TB) is a major cause of critical illness and death in people living with HIV (PLWH) worldwide. Despite this, the immunopathology of severe HIV-associated TB (HIV/TB) is poorly understood. We aimed to identify an immunopathologic signature of severe HIV/TB in sub-Saharan Africa. DESIGN AND SETTING: We analyzed proteomic data from two prospective observational cohorts of adults hospitalized with severe undifferentiated infection in Uganda: an urban discovery cohort (Entebbe, n = 241) and a rural validation cohort (Tororo, n = 253). PATIENTS: Adults (age ≥ 18 yr) hospitalized with severe febrile illness. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: Across both cohorts, severe HIV/TB was common, affecting 18% of participants in the discovery cohort and 21% in the validation cohort. Overall mortality was significant (30-d mortality of 22% in the discovery cohort and 60-d mortality of 26% in the validation cohort). Participants were stratified into three HIV/TB phenotypes: HIV-negative without TB, PLWH without TB, and PLWH with microbiologically diagnosed TB. We applied ordinal random forest models in the discovery cohort as a supervised feature-selection approach to identify proteins associated with progressive HIV/TB phenotype. In both cohorts, PLWH with microbiologically diagnosed TB were at highest risk of critical illness and death (30-d mortality of 42% in the discovery cohort and 60-d mortality of 52% in the validation cohort). An eight-protein signature reliably distinguished this phenotype, reflecting mediators of macrophage/dendritic cell activation (lysosome-associated membrane glycoprotein 3), natural killer cell and T-cell stimulation and cytotoxicity (cluster of differentiation 70, class I-restricted T-cell-associated molecule), B-cell activation (immunoglobulin lambda constant 2), protease-mediated tissue injury (protease, serine 2 [trypsin-2]), dysregulated coagulation (serpin peptidase inhibitor, clade A [alpha-1 antitrypsin], member 5), extracellular matrix remodeling (epidermal growth factor-containing fibulin-like extracellular matrix protein 1), and growth hormone/insulin-like growth factor axis dysregulation (insulin-like growth factor binding protein 3). CONCLUSIONS: We identified an immunologic signature of severe HIV/TB defined by mediators of macrophage/dendritic cell and cytotoxic lymphocyte activation, extracellular matrix remodeling, and dysregulated coagulation. These findings offer new insight into HIV/TB pathobiology and highlight potential targets for host-directed therapies in this high-risk population.

Humans↗