Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Modeling cellular processes with variational Bayesian cooperative vector quantizer.

Gene expression of a cell is controlled by sophisticated cellular processes. The capability of inferring the states of these cellular processes would provide insight into the mechanism of gene expression control system. In this paper, we propose and investigate the cooperative vector quantizer (CVQ) model for analysis of microarray data. The CVQ model could be capable of decomposing observed microarray data into many different regulatory subprocesses. To make the CVQ analysis tractable we develop and apply variational approximations. Bayesian model selection is employed in the model, so that the optimal number processes is determined purely from observed micro-array data. We test the model and algorithms on two datasets: (1) simulated gene-expression data and (2) real-world yeast cell-cycle microarray data. The results illustrate the ability of the CVQ approach to recover and characterize regulatory gene expression subprocesses, indicating a potential for advanced gene expression data analysis.

Algorithms↗

MHC class I-Ly49 interactions shape the Ly49 repertoire on murine NK cells.

This study aims to determine how the interaction of Ly49 receptors with MHC class I molecules shapes the development of the Ly49 repertoire. We have examined the percentage of NK cells that expressed Ly49A, Ly49G2, and Ly49D in single and double Ly49A/C-transgenic mice on four different MHC backgrounds, H-2(b), H-2(d), H-2(b/d), and beta(2)-microglobulin(-/-). The results show that the total numbers of NK cells were not different among the strains. The prior expression of a Ly49 receptor capable of binding to self MHC class I altered the percentage of NK cells expressing endogenous Ly49A, Ly49G2, and Ly49D even in mice in which no MHC ligand was present for the latter receptors. The NK cells in the Ly49-transgenic mice expressed the same level of endogenous Ly49 receptors as wild-type mice of a similar MHC background. In contrast, the number of NK T cells was reduced in mice in which the Ly49 transgene could bind to a MHC class I molecule. The onset of Ly49 receptor expression on NK cells during ontogeny was not altered in the presence of transgenic Ly49 receptors. These data support a sequential model and argue against a selection model for Ly49 repertoire development on NK cells.

Animals↗

A nonlinear tissue-binding model for creatine: estimation of creatine turnover time and creatinine production rate.

Intravenous bolus injections of 14C-labeled creatine were made in rabbits. Plasma and urine concentrations were measured. Plasma specific activity and urinary cumulative radioactivity data were examined by several means including the use of NONLIN. Utilization of urine and blood data suggested that a nonlinear model was more appropriate. The selected model has two nonlinear binding parameters and an elimination rate constant. The creatine turnover time was calculated to be about 22 minutes. The model allows the predication of creatine production rate that is in general agreement with the literature values.

Animals↗

Capture-recapture and multiple-record systems estimation I: History and theoretical development. International Working Group for Disease Monitoring and Forecasting.

This paper reviews the historical background and the theoretical development of models for the analysis of data from capture-recapture or multiple-record systems for estimating the size of closed populations. The models and methods were originally developed for use in fisheries and wildlife biology and were later adapted for use in connection with human populations. Application to epidemiology came much later. The simplest capture-recapture model involves two lists or samples and has four key assumptions: that the population is closed, that individuals can be matched from capture to recapture, that capture in the second sample is independent of capture in the first sample, and that the capture probabilities are homogeneous across all individuals in the population. Log-linear models provide a convenient representation for this basic capture-recapture model and its extensions to K lists. The paper provides an overview for these models and illustrates how they allow for dependency among the lists and heterogeneity in the population. The use of log-linear models for estimation in the presence of both dependence and heterogeneity is illustrated on a four-list example involving ascertainment of diabetes using data gathered in 1988 from residents of Casale Monferrato, Italy. The final section of the paper discusses techniques for model selection in the context of models for estimating the size of populations.

Bias↗

Minimalist molecular model for nanopore selectivity.

Using a simple model it is shown that the cost of constraining a hydrated potassium ion inside a narrow nanopore is smaller than the cost of constraining the smaller hydrated sodium ion. The former allows for a greater distortion of its hydration shell and can therefore maintain a better coordination. We propose that in this way the larger ion can go through narrow pores more easily. This is relevant to the molecular basis of ion selective nanopores and since this mechanism does not depend on the molecular details of the pore, it could also operate in all sorts of nanotubes, from biological to synthetic.

Cations, Monovalent↗

The molecular similarity landscape of preclinical cancer models to patient tumors.

Selecting appropriate preclinical models is fundamental for translational oncology, yet a large-scale, multi-omic quantitative comparison of their similarity to primary human tumors is lacking. To address this, we integrated transcriptomic, proteomic, and genomic profiles from over 10,000 primary tumors from The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC), alongside 4,000 preclinical models. Using a robust computational framework, we revealed a clear hierarchy of transcriptomic and proteomic similarity to patient tumors: with patient-dervied xenografts (PDXs) having greater transcriptomic and proteomic similarity to patient tumors (>) compared with patient-derived organoids (PDOs), which are equal in hierarchy to that of PDX-dervied organoids (PDXOs) > cell lines. We also quantified high molecular conservation (Pearson correlation coefficient = 0.96) across paired in vitro to in vivo platform (organoids to PDX) transitions. Furthermore, genomic analysis demonstrated that whole-exome sequencing (WES) outperforms RNA-seq in detecting DNA variants, and it identified a clonal complexity hierarchy (cell lines > PDXOs > PDXs > PDOs) reflecting the effect of passaging history on intratumor heterogeneity. Ultimately, this study delivers a comprehensive quantitative benchmark, establishing a population-level hierarchy of molecular similarity between preclinical models and primary tumors and providing a data-driven reference for model selection. These findings offer a data-driven framework for selecting models that balance biological representativeness with experimental practicality.

Humans↗

A model of kin selection for an altruistic trait considered as a quantitative character.

Conditions for natural selection to favor increase of a quantitative character are derived for a model in which individuals associate in groups of size n. It is assumed that the logarithm of the fitness of an individual is the sum of two parts, one proportional to the individual's own phenotype, and the other to the mean phenotype in its group. The resulting conditions for the trait to increase under natural selection are analogous to the results found previously in single-locus kin selection models.

Altruism↗

Dynamics of genetic variability in two-locus models of stabilizing selection.

We study a two locus model, with additive contributions to the phenotype, to explore the dynamics of different phenotypic characteristics under stabilizing selection and recombination. We demonstrate that the interaction of selection and recombination results in constraints on the mode of phenotypic evolution. Let Vg be the genic variance of the trait and CL be the contribution of linkage disequilibrium to the genotypic variance. We demonstrate that, independent of the initial conditions, the dynamics of the system on the plane (Vg, CL) are typically characterized by a quick approach to a straight line with slow evolution along this line afterward. We analyze how the mode and the rate of phenotypic evolution depend on the strength of selection relative to recombination, on the form of fitness function, and the difference in allelic effect. We argue that if selection is not extremely weak relative to recombination, linkage disequilibrium generated by stabilizing selection influences the dynamics significantly. We demonstrate that under these conditions, which are plausible in nature and certainly the case in artificial stabilizing selection experiments, the model can have a polymorphic equilibrium with positive linkage disequilibrium that is stable simultaneously with monomorphic equilibria.

Animals↗

Dynamical stability conditions for recurrent neural networks with unsaturating piecewise linear transfer functions.

We establish two conditions that ensure the nondivergence of additive recurrent networks with unsaturating piecewise linear transfer functions, also called linear threshold or semilinear transfer functions. As Hahnloser, Sarpeshkar, Mahowald, Douglas, and Seung (2000) showed, networks of this type can be efficiently built in silicon and exhibit the coexistence of digital selection and analog amplification in a single circuit. To obtain this behavior, the network must be multistable and nondivergent, and our conditions allow determining the regimes where this can be achieved with maximal recurrent amplification. The first condition can be applied to nonsymmetric networks and has a simple interpretation of requiring that the strength of local inhibition match the sum over excitatory weights converging onto a neuron. The second condition is restricted to symmetric networks, but can also take into account the stabilizing effect of nonlocal inhibitory interactions. We demonstrate the application of the conditions on a simple example and the orientation-selectivity model of Ben-Yishai, Lev Bar-Or, and Sompolinsky (1995). We show that the conditions can be used to identify in their model regions of maximal orientation-selective amplification and symmetry breaking.

Animals↗

Asymptotic properties of mathematical models of excitability.

We analyse small parameters in selected models of biological excitability, including Hodgkin-Huxley (Hodgkin & Huxley 1952 J. Physiol.117, 500-544) model of nerve axon, Noble (Noble 1962 J. Physiol.160, 317-352) model of heart Purkinje fibres and Courtemanche et al. (Courtemanche et al. 1998 Am. J. Physiol.275, H301-H321) model of human atrial cells. Some of the small parameters are responsible for differences in the characteristic time-scales of dynamic variables, as in the traditional singular perturbation approaches. Others appear in a way which makes the standard approaches inapplicable. We apply this analysis to study the behaviour of fronts of excitation waves in spatially extended cardiac models. Suppressing the excitability of the tissue leads to a decrease in the propagation speed, but only to a certain limit; further suppression blocks active propagation and leads to a passive diffusive spread of voltage. Such a dissipation may happen if a front propagates into a tissue recovering after a previous wave, e.g. re-entry. A dissipated front does not recover even when the excitability restores. This has no analogy in FitzHugh-Nagumo model and its variants, where fronts can stop and then start again. In two spatial dimensions, dissipation accounts for breakups and self-termination of re-entrant waves in excitable media with Courtemanche et al. kinetics.

Action Potentials↗

Reconstructing bifurcation diagrams from noisy time series using nonlinear autoregressive models.

We introduce a formalism for the reconstruction of bifurcation diagrams from noisy time series. The method consists in finding a parametrized predictor function whose bifurcation structure is similar to that of the given system. The reconstruction algorithm is composed of two stages: model selection and bifurcation parameter identification. In the first stage, an appropriate model that best represents all the given time series is selected. A nonlinear autoregressive model with polynomial terms is employed in this study. The identification of the bifurcation parameters from among the many model parameters is done in the second stage. The algorithm works well even for a limited number of time series.

Journal Article↗

Simultaneous gene clustering and subset selection for sample classification via MDL.

MOTIVATION: The microarray technology allows for the simultaneous monitoring of thousands of genes for each sample. The high-dimensional gene expression data can be used to study similarities of gene expression profiles across different samples to form a gene clustering. The clusters may be indicative of genetic pathways. Parallel to gene clustering is the important application of sample classification based on all or selected gene expressions. The gene clustering and sample classification are often undertaken separately, or in a directional manner (one as an aid for the other). However, such separation of these two tasks may occlude informative structure in the data. Here we present an algorithm for the simultaneous clustering of genes and subset selection of gene clusters for sample classification. We develop a new model selection criterion based on Rissanen's MDL (minimum description length) principle. For the first time, an MDL code length is given for both explanatory variables (genes) and response variables (sample class labels). The final output of the proposed algorithm is a sparse and interpretable classification rule based on cluster centroids or the closest genes to the centroids. RESULTS: Our algorithm for simultaneous gene clustering and subset selection for classification is applied to three publicly available data sets. For all three data sets, we obtain sparse and interpretable classification models based on centroids of clusters. At the same time, these models give competitive test error rates as the best reported methods. Compared with classification models based on single gene selections, our rules are stable in the sense that the number of clusters has a small variability and the centroids of the clusters are well correlated (or consistent) across different cross validation samples. We also discuss models where the centroids of clusters are replaced with the genes closest to the centroids. These models show comparable test error rates to models based on single gene selection, but are more sparse as well as more stable. Moreover, we comment on how the inclusion of a classification criterion affects the gene clustering, bringing out class informative structure in the data. AVAILABILITY: The methods presented in this paper have been implemented in the R language. The source code is available from the first author.

Algorithms↗

Intestinal glucose transport using perfused rat jejunum in vivo: model analysis and derivation of corrected kinetic constants.

1. The transport model that best describes intestinal glucose transport in vivo remains unsettled. Three models have been proposed: (1) a single carrier, (2) a single carrier plus passive diffusion, and (3) a two-carrier system. The objectives of the current studies were to define the transport model that best fits experimental data and to devise methods to obtain the kinetic constants, corrected for diffusion barrier resistance, with this model. 2. Intestinal glucose uptake was measured during perfusion of rat jejunum in vivo over a wide range of perfusate concentrations and diffusion barrier resistance was determined under identical experimental conditions. The data were fitted to the transport equations that describe the three models with appropriate diffusion barrier corrections, and the kinetic constants were derived by non-linear regression techniques. The fit of each model to the data was assessed using six statistical tests, five of which favoured a model described by a single carrier and passive diffusion. 3. The main conclusions of these studies are: (1) kinetic constants uncorrected for diffusion barrier resistance are seriously in error; (2) values for the derived kinetic constants are strongly dependent on the transport model selected for the data analysis which underscores the need for rigorous model analysis; (3) corrected kinetic constants may be obtained by either non-linear regression or by a simpler graphical analysis once the correct transport model has been selected and diffusion barrier resistance determined; (4) only corrected kinetic constants should be used for inter-species comparisons or to study the effect of specific interventions on intestinal glucose transport.

Animals↗

Hierarchical Bayesian spatial models for alcohol availability, drug "hot spots" and violent crime.

BACKGROUND: Ecologic studies have shown a relationship between alcohol outlet densities, illicit drug use and violence. The present study examined this relationship in the City of Houston, Texas, using a sample of 439 census tracts. Neighborhood sociostructural covariates, alcohol outlet density, drug crime density and violent crime data were collected for the year 2000, and analyzed using hierarchical Bayesian models. Model selection was accomplished by applying the Deviance Information Criterion. RESULTS: The counts of violent crime in each census tract were modelled as having a conditional Poisson distribution. Four neighbourhood explanatory variables were identified using principal component analysis. The best fitted model was selected as the one considering both unstructured and spatial dependence random effects. The results showed that drug-law violation explained a greater amount of variance in violent crime rates than alcohol outlet densities. The relative risk for drug-law violation was 2.49 and that for alcohol outlet density was 1.16. Of the neighbourhood sociostructural covariates, males of age 15 to 24 showed an effect on violence, with a 16% decrease in relative risk for each increase the size of its standard deviation. Both unstructured heterogeneity random effect and spatial dependence need to be included in the model. CONCLUSION: The analysis presented suggests that activity around illicit drug markets is more strongly associated with violent crime than is alcohol outlet density. Unique among the ecological studies in this field, the present study not only shows the direction and magnitude of impact of neighbourhood sociostructural covariates as well as alcohol and illicit drug activities in a neighbourhood, it also reveals the importance of applying hierarchical Bayesian models in this research field as both spatial dependence and heterogeneity random effects need to be considered simultaneously.

Alcohol Drinking↗

SLAM: a connectionist model for attention in visual selection tasks.

SLAM, the SeLective Attention Model, performs visual selective attention tasks, an analysis of which shows that two processes, object and attribute selection, are both necessary and sufficient. It is based upon the McClelland and Rumelhart (1981) model for visual word recognition, with the addition of a response selection and evaluation mechanism. The responses may be correct or incorrect and, in particular conditions, SLAM may not make a response at all. Moreover, it allows for the generation of specific responses in time. SLAM's main characteristics are parallelism restricted by competition within modules, heterarchical processing in a hierarchical structure, and generation of responses as a result of relaxation given the conjoint constraints of stimulation, object, and attribute selection. The model is considered to represent an individual subject performing filtering tasks and demonstrates appropriate selective behavior. It is also tested quantitatively using a single tentative set of model parameters. The study reports simulations of four different filtering experiments, modeling response latencies, and error proportions. Specifications are made to take account of instructions, previous trials, and the effect of a barmarker cue and of asynchronies in stimulus and cue onsets. The model is then extended in order to provide simulations of a number of Stroop experiments, which can be regarded as filtering tasks with nonequivalent stimuli. The extension required for Stroop simulations is the addition of direct connections between compatible stimulus and response aspects. The direct connections do not affect the simulation of simpler filtering tasks. A variety of different experiments carried out by different authors is simulated. The model is discussed in terms of how modular architecture and the interaction of excitation and inhibition generate facilitation or inhibition of response latencies.

Arousal↗

A genetic model of interpopulation variation and covariation of quantitative characters.

Evolutionary consequences of natural selection, migration, genotype-environment interaction, and random genetic drift on interpopulation variation and covariation of quantitative characters are analysed in terms of a selection model that partitions natural selection into directional and stabilizing components. Without migration, interpopulation variation and covariation depend mainly on the pattern and intensities of selection among populations and the harmonic mean of effective population sizes. Both transient and equilibrium covariance structures are formulated with suitable approximations. Migration reduces the differentiation among populations, but its effect is less with genotype-environment interaction. In some special cases of genotype-environment interaction, the equilibrium interpopulation variation and covariation is independent of migration.

Environment↗

An analysis paradigm for investigating multi-locus effects in complex disease: examination of three GABA receptor subunit genes on 15q11-q13 as risk factors for autistic disorder.

Gene-gene interactions are likely involved in many complex genetic disorders and new statistical approaches for detecting such interactions are needed. We propose a multi-analytic paradigm, relying on convergence of evidence across multiple analysis tools. Our paradigm tests for main and interactive effects, through allele, genotype and haplotype association. We applied our paradigm to genotype data from three GABAA receptor subunit genes (GABRB3, GABRA5, and GABRG3) on chromosome 15 in 470 Caucasian autism families. Previously implicated in autism, we hypothesized these genes interact to contribute to risk. We detected no evidence of main effects by allelic (PDT, FBAT) or genotypic (genotype-PDT) association at individual markers. However, three two-marker haplotypes in GABRG3 were significant (HBAT). We detected no significant multi-locus associations using genotype-PDT analysis or the EMDR data reduction program. However, consistent with the haplotype findings, the best single locus EMDR model selected a GABRG3 marker. Further, the best pairwise genotype-PDT result involved GABRB3 and GABRG3, and all multi-locus EMDR models also selected GABRB3 and GABRG3 markers. GABA receptor subunit genes do not significantly interact to contribute to autism risk in our overall data set. However, the consistency of results across analyses suggests that we have defined a useful framework for evaluating gene-gene interactions.

Autistic Disorder↗

Tg.AC genetically altered mouse: assay working group overview of available data.

In a Government/Industry/Academic partnership to evaluate alternative approaches to carcinogenicity testing, 21 pharmaceutical agents representing a variety of chemical and pharmacological classes and possessing known human and or rodent carcinogenic potential were selected for study in several rodent models. The studies from this partnership project, coordinated by the International Life Sciences Institute, provide additional data to better understand the models' limitations and sensitivity in identifying carcinogens. The results of these alternative model studies were reviewed by members of Assay Working Groups (AWG) composed of scientists from government and industry with expertise in toxicology, genetics, statistics, and pathology. The Tg.AC genetically manipulated mouse was one of the models selected for this project based on previous studies indicating that Tg.AC mice seem to respond to topical application of either mutagenic or nonmutagenic carcinogens with papilloma formation at the site of application. This communication describes the results and AWG interpretations of studies conducted on 14 chemicals administered by the topical and oral (gavage and/or diet) routes to Tg.AC genetically manipulated mice. Cyclosporin A, an immunosuppresant human carcinogen, ethinyl estradiol and diethylstilbestrol (human hormone carcinogens) and clofibrate, an hepatocarcinogenic peroxisome proliferator in rodents, were considered clearly positive in the topical studies. In the oral studies, ethinyl estradiol and diethylstilbestrol were negative, cyclosporin was considered equivocal, and results were not available for the clofibrate study. Of the 3 genotoxic human carcinogens (phenacetin, melphalan, and cyclophosphamide), phenacetin was negative by both the topical and oral routes. Melphalan and cyclophosphamide are, respectively, direct and indirect DNA alkylating agents and topical administration of both caused equivocal responses. With the exception of clofibrate, Tg.AC mice did not exhibit tumor responses to the rodent carcinogens that were putative human noncarcinogens, (di(2-ethylhexyl) phthalate, methapyraline HCl, phenobarbital Na, reserpine, sulfamethoxazole or WY-14643, or the nongenotoxic, noncarcinogen, sulfisoxazole) regardless of route of administration. Based on the observed responses in these studies, it was concluded by the AWG that the Tg.AC model was not overly sensitive and possesses utility as an adjunct to the battery of toxicity studies used to establish human carcinogenic risk.

Animal Testing Alternatives↗