Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Regression”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Characterization of low-intensity lesions in the peripheral zone of prostate on pre-biopsy endorectal coil MR imaging.

The aim of this study was to determine which morphological features of low-intensity lesions in the peripheral zone of the prostate are predictable of prostate cancer on pre-biopsy T2-weighted integrated endorectal phased-array MR images. The MR examinations were performed in 69 consecutive patients with elevated level of prostate-specific antigen (>4 ng/ml) and/or a positive digital rectal examination before transperineal 12-site biopsy. Two radiologists evaluated presence of lesions, their morphological features, and possibility of malignancy in divided into four sections of the peripheral zone. Imaging analysis findings were compared with biopsy results. Discriminative features were selected by stepwise logistic regression. Descriptive statistics and receiver operating characteristics (ROC) curves were also calculated. Sixty-eight benign lesions and 23 malignant lesions were found. Wedge shape and diffuse extensions without mass effect were significantly associated with benignity ( P=0.0105 and 0.002, respectively). Lesion size was significantly associated with malignancy ( P=0.0001). For evaluating probability of malignancy for lesions, regression model showed a comparable accuracy with the total impression for the readers in ROC analysis (Az 0.9095 vs 0.9266, respectively). Wedge shape, diffuse extension without mass effect, and size are the morphological features of low-intensity lesions in the peripheral zone on pre-biopsy T2-weighted MR images that give the best prediction of malignancy.

Aged↗

A Bayesian analysis of regression models with continuous errors with application to longitudinal studies.

We employ a regression model with errors that follow a continuous autoregressive process to analyse longitudinal studies. In this way, unequally spaced observations do not present a problem in the analysis. We employ a Bayesian approach, where our inferences are based on a direct resampling process that generates values from the posterior distribution of the parameters of the model. We illustrate these Bayesian inferences with an analysis of a longitudinal study that involves the regression of foetal head circumference on menstrual age. Using these same data, we contrast the Bayesian approach with a maximum likelihood technique.

Bayes Theorem↗

Estimating missing data: an iterative regression approach.

The problem of missing data is common in all fields of science. Various methods of estimating missing values in a dataset exist, such as deletion of cases, insertion of sample mean, and linear regression. Each approach presents problems inherent in the method itself or in the nature of the pattern of missing data. We report a method that (1) is more general in application and (2) provides better estimates than traditional approaches, such as one-step regression. The model is general in that it may be applied to singular matrices, such as small datasets or those that contain dummy or index variables. The strength of the model is that it builds a regression equation iteratively, using a bootstrap method. The precision of the regressed estimates of a variable increases as regressed estimates of the predictor variables improve. We illustrate this method with a set of measurements of European Upper Paleolithic and Mesolithic human postcranial remains, as well as a set of primate anthropometric data. First, simulation tests using the primate data set involved randomly turning 20% of the values to "missing". In each case, the first iteration produced significantly better estimates than other estimating techniques. Second, we applied our method to the incomplete set of human postcranial measurements. MISDAT estimates always perform better than replacement of missing data by means and better than classical multiple regression. As with classical multiple regression, MISDAT performs when squared multiple correlation values approach the reliability of the measurement to be estimated, e.g., above about 0. 8.

Animals↗

Subset selection with additional order information.

Traditional subset selection procedures were developed without assuming any order information about the response variable. However, in some applications there is additional, even though incomplete, order information about the treatment effects at increasing treatment levels. One important example is the up-then-down umbrella ordering with an unknown peak. This type of additional order information is utilized explicitly in this paper to construct subset selection procedures for several settings studied in the literature where only order restricted tests are known to exist. This paper also proposes a straightforward algorithm to compute the isotonic regression with respect to umbrella orderings, which can be used to carry out the proposed procedures. Examples are given to illustrate the procedures and algorithm.

Algorithms↗

Regression-based variable clustering for data reduction.

In many studies it is of interest to cluster states, counties or other small regions in order to obtain improved estimates of disease rates or other summary measures, and a more parsimonious representation of the country as a whole. This may be the case if there are too many to summarize concisely, and/or many regions with a small number of cases. By merging the regions into larger geographic areas, we obtain more cases within each area (and hence lower standard errors for parameter estimates), as well as fewer areas to summarize in terms of disease rates. The resulting clusters should be such that regions within the same cluster are similar in terms of their disease rates. In this paper we present a clustering algorithm which uses data at the subject-specific level in order to cluster the original regions into a reduced set of larger areas. The proposed clustering algorithm expresses the clustering goals in terms of a regression framework. This formulation of the problem allows the regions to be clustered in terms of their association with the response, and confounding variables measured at the subject-specific level may be easily incorporated during the clustering process. Additionally, this framework allows estimation and testing of the association between the areas and the response. The statistical properties and performance of the algorithm were evaluated via simulation studies, and the results are promising. Additional simulations illustrate the importance of controlling for confounding variables during the clustering process, rather than after the clusters are determined. The algorithm is illustrated with data from the Cardiovascular Health Study. Although developed with a specific application in mind, the method is applicable to a wide range of problems.

Aged↗

Efficient size control of amphiphilic cyclodextrin nanoparticles through a statistical mixture design methodology.

PURPOSE: the aim of the study was to investigate size control of amphiphilic beta-cyclodextrin nanoparticles obtained by solvent displacement technique. METHODS: An experimental design methodology for mixture design was undertaken using D-optimal approach with the following technique variables: water fraction X1 (40-70% v/v), acetone fraction X2 (0-60% v/v) and ethanol fraction X3 (0-60% v/v). RESULTS: The resulting quadratic model obtained after logarithmic transformation of data and partial least-square regression was statistically validated and experimentally checked. Also, the morphology of the colloidal nanoparticles from selected experiments was observed by cryo-transmission electron microscopy. CONCLUSIONS: This experimental design approach allowed to produce interesting amphiphilic beta-cyclodextrin nanoparticles with a predicted mean size varying from 60 to 400 nm.

Cyclodextrins↗

Determining confidence limits for drug potency in immunoassay.

Nonlinear models have frequently been used to characterize dose-response data obtained from biological assays. The effect of a bioactive agent is observed and the model allows prediction of the dose required to obtain the observed effect (the 'inverse prediction'). The precision of this estimate is important in potency determination. Here, a general method is presented for calculating the inverse confidence intervals for estimates of dose potencies obtained from nonlinear models often used to describe these tests. The approach is demonstrated with application of data sets to the negative exponential and four-parameter logistic regression models. Necessary theory is presented and followed by detailed discussion in which estimation strategies are explained and intermediate quantities calculated.

Biological Assay↗

Tobit, fixed effects, and cohort analyses of the relationship between severity and duration of rheumatoid arthritis.

Three methodological problems are commonly faced by researchers investigating relationships between severity and duration of illness among patients with rheumatoid arthritis (RA). (1) Linear regression techniques yield biased estimates when measures of severity are continuous but range between and include limiting values such as 0 and 3. (2) Data from the same patient over time are typically pooled together with data from different patients at the same time and over time. Models are then used that do not account for the statistical problems that can result from pooling. (3) Persons with varying years of duration of disease are typically combined and analyzed without any special attention to cohort effects. Changes in severity over time for cohorts of patients with fewer than 10 years of duration may be different from changes in severity of patients with more than 20 years of duration from the onset of the disease. In this study, severity is measured by the 0-3 disability scale in the Stanford Health Assessment Questionnaire (HAQ). Duration is measured by self-report of the onset of symptoms by subjects. Popular techniques are borrowed from econometrics--Tobit, Fixed Effects, and dummy variables for Cohort Models--that were developed to address three analogous problems in economic data. The three economic techniques are applied separately and together using data collected by Arthritis, Rheumatism, and Aging Medical Information System (ARAMIS) on 330 RA patients in 1981 who were followed until 1989. Although the Tobit technique does not appear to be especially useful with these data, Fixed Effects and Cohort Models do appear to be useful.(ABSTRACT TRUNCATED AT 250 WORDS)

Activities of Daily Living↗

On the relation between initial value and slope.

Suppose measurements of a particular feature are collected at baseline and at a number of subsequent time points and that for each individual there is a roughly linear trend in time. This paper takes three approaches to testing whether there is a relation between the initial value and the slope. It also considers whether the initial value for an individual is a useful predictor of the slope for that individual. The problems are formulated in terms of regression models with random coefficients. The solutions are illustrated using data from an observational study of clinical correlates of disability and progression in Huntington's disease.

Biometry↗

Estimating haplotype effects on dichotomous outcome for unphased genotype data using a weighted penalized log-likelihood approach.

OBJECTIVE: To develop a method to estimate haplotype effects on dichotomous outcomes when phase is unknown, that can also estimate reliable effects of rare haplotypes. METHODS: In short, the method uses a logistic regression approach, with weights attached to all possible haplotype combinations of an individual. An EM-algorithm was used: in the E-step the weights are estimated, and the M-step consists of maximizing the joint log-likelihood. When rare haplotypes were present, a penalty function was introduced. We compared four different penalties. To investigate statistical properties of our method, we performed a simulation study for different scenarios. The evaluation criteria are the mean bias of the parameter estimates, the root of the mean squared error, the coverage probability, power, Type I error rate and the false discovery rate. RESULTS: For the unpenalized approach, mean bias was small, coverage probabilities were approximately 95%, power ranged from 15.2 to 44.7% depending on haplotype frequency, and Type I error rate was around 5%. All penalty functions reduced the standard errors of the rare haplotypes, but introduced bias. This trade-off decreased power. CONCLUSION: The unpenalized weighted log-likelihood approach performs well. A penalty function can help to estimate an effect for rare haplotypes.

Algorithms↗

Parameter estimation from incomplete data in binomial regression when the missing data mechanism is nonignorable.

We propose a method for estimating parameters in binomial regression models when the response variable is missing and the missing data mechanism is nonignorable. We assume throughout that the covariates are fully observed. Using a logit model for the missing data mechanism, we show how parameter estimation can be accomplished using the EM algorithm by the method of weights proposed in Ibrahim (1990, Journal of the American Statistical Association 85, 765-769). An example from the Six Cities Study (Ware et al., 1984, American Review of Respiratory Diseases 129, 366-374) is presented to illustrate the method.

Air Pollution↗

Cutpoint selection for categorizing a continuous predictor.

This article presents a new approach for choosing the number of categories and the location of category cutpoints when a continuous exposure variable needs to be categorized to obtain tabular summaries of the exposure effect. The optimum categorization is defined as the partition that minimizes a measure of distance between the true expected value of the outcome for each subject and the estimated average outcome among subjects in the same exposure category. To estimate the optimum partition, an efficient nonparametric estimate of the unknown regression function is substituted into a formula for the asymptotically optimum categorization. This new approach is easy to implement and it outperforms existing cutpoint selection methods.

Age Factors↗

Risk factors for invasive carcinoma of the uterine cervix in Latin America.

A study of 759 cervical cancer patients, 1,430 controls, and 689 sex partners in four Latin American countries has made it possible to assess the influence of multiple factors upon the risk of invasive cervical cancer. The principal risk factors identified were the woman's age at first coitus, the number of her steady sex partners, her number of live births, the presence of DNA from human papillomavirus (HPV) types 16 or 18, a history of venereal disease, nonparticipation in early detection programs, and low socioeconomic status. There is good reason to believe that extensive detection programs directed mainly at high-risk groups in the Americas can reduce the high incidence of cervical cancer in this Region.

Adult↗

General methods for analysing repeated measures.

We present an overview of issues and methods for analysing repeated measures of a continuous random variable. We discuss modelling the mean vector and covariance structure, statistical efficiency, regression diagnostics, and discrepancies between longitudinal and cross-sectional methods. We illustrate key points with examples and discuss areas requiring further development.

Adult↗

Robust regression-based analysis of drug-nucleic acid binding.

Outlier detection can be very important in analyzing data from Scatchard plots. In this study, a robust (outlier-resistant) regression procedure was used in conjunction with a Scatchard plot to study the binding of the methylphenazinium cation with double-stranded DNA. The procedures, their results, and their advantages are discussed.

DNA↗