Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Regression”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Estimation of parameters and missing values under a regression model with non-normally distributed and non-randomly incomplete data.

We carried out a simulation study to compare the performance of three algorithms (complete cases, ALLVALUE, and expectation maximization, EM) in estimating regression parameters and missing values for situations that have varying amounts of missing data, distributions (normal, mixture of normals and lognormal), patterns of incomplete data (random, related and censored), and degrees of correlational structure among the dependent and independent variables. We found that the EM and complete cases algorithms performed equally well regardless of the correlational structure, when the percentage of incomplete data was only 5 per cent. When this percentage increased to 25 per cent, the EM algorithm was generally best for estimation, but the complete cases algorithm was safe and conservative. This finding may be attributed to the study design, which required that the slopes be the same in the population of all cases, and in the population of complete cases. In addition, the one-step imputing method (ALLVALUE) was competitive only for situations with weak correlational structure and/or little missing data. In that situation the bias caused with use of all available information was less than that caused with use of only complete cases. On the other hand, for imputation, the EM algorithm performed optimally, even in situations of censored or log-normally distributed data.

Algorithms↗

Random regression analyses using B-splines to model growth of Australian Angus cattle.

Regression on the basis function of B-splines has been advocated as an alternative to orthogonal polynomials in random regression analyses. Basic theory of splines in mixed model analyses is reviewed, and estimates from analyses of weights of Australian Angus cattle from birth to 820 days of age are presented. Data comprised 84533 records on 20731 animals in 43 herds, with a high proportion of animals with 4 or more weights recorded. Changes in weights with age were modelled through B-splines of age at recording. A total of thirteen analyses, considering different combinations of linear, quadratic and cubic B-splines and up to six knots, were carried out. Results showed good agreement for all ages with many records, but fluctuated where data were sparse. On the whole, analyses using B-splines appeared more robust against "end-of-range" problems and yielded more consistent and accurate estimates of the first eigenfunctions than previous, polynomial analyses. A model fitting quadratic B-splines, with knots at 0, 200, 400, 600 and 821 days and a total of 91 covariance components, appeared to be a good compromise between detailedness of the model, number of parameters to be estimated, plausibility of results, and fit, measured as residual mean square error.

Algorithms↗

Communities mobilizing for change on alcohol: outcomes from a randomized community trial.

OBJECTIVE: Communities Mobilizing for Change on Alcohol (CMCA) was a randomized 15-community trial of a community organizing intervention designed to reduce the accessibility of alcoholic beverages to youths under the legal drinking age. METHOD: Data were collected at baseline before random assignment of communities to intervention or control condition, and again at follow-up after a 2.5-year intervention. Data collection included in-school surveys of twelfth graders, telephone surveys of 18- to 20-year-olds and alcohol merchants, and direct testing of the propensity of alcohol outlets to sell to young buyers. Analyses were based on mixed-model regression, used the community as the unit of assignment, took into account the nesting of individual respondents or alcohol outlets within each community, and controlled for relevant covariates. RESULTS: Results show that the CMCA intervention significantly and favorably affected both the behavior of 18- to 20-year-olds (effect size = 0.76, p<.01) and the practices of on-sale alcohol establishments (effect size = 1.18, p<.05), may have favorably affected the practices of off-sale alcohol establishments (effect size = 0.32, p = .08), but had little effect on younger adolescents. Alcohol merchants appear to have increased age-identification checking and reduced propensity to sell to minors. Eighteen- to 20-year-olds reduced their propensity to provide alcohol to other teens and were less likely to try to buy alcohol, drink in a bar or consume alcohol. CONCLUSIONS: Community organizing is a useful intervention approach for mobilizing communities for institutional and policy change to improve the health of the population.

Adolescent↗

Experience with a test-day model.

The Canadian Test-Day Model is a 12-trait random regression animal model in which traits are milk, fat, and protein test-day yields, and somatic cell scores on test days within each of first three lactations. Test-day records from later lactations are not used. Random regressions (genetic and permanent environmental) were based on Wilmink's three parameter function that includes an intercept, regression on days in milk, and regression on an exponential function to the power -0.05 times days in milk. The model was applied to over 22 million test-day records of over 1.4 million cows in seven dairy breeds for cows first calving since 1988. A theoretical comparison of test-day model to 305-d complete lactation animal model is given. Each animal in an analysis receives 36 additive genetic solutions (12 traits by three regression coefficients), and these are combined to give one estimated breeding value (EBV) for each of milk, fat, and protein yields, average daily somatic cell score and milk yield persistency (for bulls only). Correlation of yield EBV with previous 305-d lactation model EBV for bulls was 0.97 and for cows was 0.93 (Holsteins). A question is whether EBV for yield traits for each lactation should be combined into one overall EBV, and if so, what method to combine them. Implementation required development of new methods for approximation of reliabilities of EBV, inclusion of cows without test day records in analysis, but which were still alive and had progeny with test-day records, adjustments for heterogeneous herd-test date variances, and international comparisons. Efforts to inform the dairy industry about changes in EBV due to the model and recovering information needed to explain changes in specific animals' EBV are significant challenges. The Canadian dairy industry will require a year or more to become comfortable with the test-day model and to realize the impact it could have on selection decisions.

Analysis of Variance↗

Comprehensive survey of the relationship between serum concentration and therapeutic effect of amitriptyline in depression.

The relationship between serum concentration (C(s)) of amitriptyline and its therapeutic effect in depression has been investigated frequently over the last 3 decades; however, the results were controversial and no consensus was reached. Therefore, we have performed a comprehensive survey and meta-analysis of the subject. All relevant literature was included, and the design of studies on the serum concentration-therapeutic effect relationship (SCTER) of amitriptyline was evaluated. Pooled original data from SCTER studies with adequate design were analysed by various statistical methods: regression analysis of therapeutic effect and C(s); comparison of the mean therapeutic effect in various ranges of C(s); dichotomisation of outcome and analysis according to sensitivity of receiver operation curves; frequency of responders and nonresponders in ranges determined by points of sensitivity; analysis of the distribution of C(s) in responders and nonresponders; logistic regression of responders and nonresponders with C(s) and other independent variables; calculation of effect size (g) and mean effect size (g(m)). Forty-five SCTER studies of amitriptyline were identified, and 27 studies met the minimum criteria of adequate study design. Inadequate study design predicted the finding of no SCTER. Analysis of the pooled data from studies with adequate design confirmed a therapeutic window of the sum of C(s) of amitriptyline and its active metabolite nortriptyline of about 80 to 200 microg/L. A moderate and significant positive g(m) (0.538, 95% confidence interval 0.167 to 0.909) was calculated for treatment with C(s) within the therapeutic window in comparison with treatment with C(s) outside the therapeutic window (19 studies with adequate design and original data available, n = 583). In conclusion, the evidence for a biphasic SCTER of amitriptyline in depression is considerably improved, and the results may help to find a consensus in the future. However, the clinical benefit of therapeutic drug monitoring of amitriptyline can only be demonstrated in a controlled and randomised study. Furthermore, the results provide further evidence that antidepressants at optimum C(s) are superior to placebo in the treatment of depression.

Amitriptyline↗

Self-reported long-term smoking cessation in patients with respiratory disease: prediction of success and perception of health effects.

Smoking status of 372 patients with respiratory disease, who had been advised to quit smoking by a respiratory specialist, was assessed six months after the advice. A multiple logistic regression model was developed for prediction of successful abstinence. The patients were again followed four to seven years later. Questionnaires were returned by 160 patients (43.0%). Of the remaining patients, 27 (7.3%) had died, 12 (3.2%) refused to participate, 53 (14.2%) had no current address available and 120 (32.3%) did not return questionnaires mailed to them. Among the respondents, 31.9% reported at least one year of abstinence from cigarettes, 63.1% were still smoking and 5.0% had quit smoking for periods of less than one year. While the original logistic model was not very useful for predicting long-term success (69.7% accuracy of classification), a model that included, as predictors, six-month smoking status and reasons for smoking other than addiction, was more useful (78.9% accuracy). At follow-up, successful abstainers reported improvement in their respiratory condition but no differences were found in reported symptoms or emotional well-being when they were compared to those who continued to smoke. Treatment implications of these results are discussed and include offers of alternative treatments if short-term abstinence is not achieved following physician advice.

Attitude to Health↗

Parametric survival models for interval-censored data with time-dependent covariates.

We present a parametric family of regression models for interval-censored event-time (survival) data that accomodates both fixed (e.g. baseline) and time-dependent covariates. The model employs a three-parameter family of survival distributions that includes the Weibull, negative binomial, and log-logistic distributions as special cases, and can be applied to data with left, right, interval, or non-censored event times. Standard methods, such as Newton-Raphson, can be employed to estimate the model and the resulting estimates have an asymptotically normal distribution about the true values with a covariance matrix that is consistently estimated by the information function. The deviance function is described to assess model fit and a robust sandwich estimate of the covariance may also be employed to provide asymptotically robust inferences when the model assumptions do not apply. Spline functions may also be employed to allow for non-linear covariates. The model is applied to data from a long-term study of type 1 diabetes to describe the effects of longitudinal measures of glycemia (HbA1c) over time (the time-dependent covariate) on the risk of progression of diabetic retinopathy (eye disease), an interval-censored event-time outcome.

Biometry↗

Genetic parameters for milk production and persistency for Danish Holsteins estimated in random regression models using REML.

(Co)variance components for milk, fat, and protein yield of 8075 first-parity Danish Holsteins (DH) were estimated in random regression models by REML. For all analyses, the fixed part of the model was held constant, whereas four different functions were applied to model the additive genetic effect and the permanent environment effect. Homogeneous residual variance was assumed throughout lactation. Univariate models were compared using a minimum of -2 ln(restricted likelihood) as the criterion for best fit. Heritabilities as a function of time were calculated from the estimated curve parameters from univariate analyses. Independent of the function applied and the trait in question, heritabilities were lowest in the beginning of the lactation. Heritabilities for persistency of fat yield were slightly higher than heritabilities for persistency of milk and protein yield. Genetic correlations between persistency and 305-d production were higher for protein and milk yield than for fat yield. Bivariate analyses between the production traits were carried out in sire models using the models with the best 3-parameter curve fit in the univariate analyses. Correlations between traits were calculated from covariance components for curve parameters estimated in bivariate analyses. Genetic correlations between milk and protein yield were higher than between milk and fat yield.

Algorithms↗

Risk stratification for cardiac valve replacement. National Cardiac Surgery Database. Database Committee of The Society of Thoracic Surgeons.

BACKGROUND: The Society of Thoracic Surgeons National Database Committee is committed to risk stratification and assessment as integral elements in the practice of cardiac operations. The National Cardiac Surgery Database was created to analyze data from subscribing institutions across the country. We analyzed the database for valve replacement procedures with and without coronary artery bypass grafting to determine trends in risk stratification. METHODS: The database contains complete records of 86,580 patients who had valve replacement procedures at the participating institutions between 1986 and 1995, inclusive. The 1995 harvest of data was conducted in late 1996 and available for evaluation in 1997. These records were used to conduct an in-depth analysis of risk factors associated with valve replacement and to provide prediction of operative death by using regression analysis. Regression models were made for six subgroups. RESULTS: Adverse patient risk factors, including diabetes, hypertension and reoperation, but not ventricular function, increased over time. There were trends with regard to increasing age of the various population subsets. The types of prostheses used remained similar over time, with more mechanical prostheses than bioprostheses used for both aortic and mitral valve replacement. There was a trend toward increased use of bioprostheses in aortic replacements and decreased use in mitral replacements between 1991 and 1995 than between 1986 and 1990. The mortality rate was determined by patient subset for primary operation and reoperation and by urgency status. The modeling showed that the predicted and observed mortality correlated for all age groups and within patient subsets. CONCLUSIONS: Risk modeling is a valuable tool for predicting the probability of operative death in any individual patient. This large, multiinstitutional database is capable of determining modern operative risk and should provide standards for acceptable care. The study illustrates the importance of risk stratification for early death both for the patient and the surgeon.

Adult↗

Analyses of growth curves of nellore cattle by multiple-trait and random regression models.

The purpose of this study was to compare estimates of genetic parameters for sequential growth of beef cattle using two models and two data sets. Growth curves of Nellore cattle were analyzed using body weights measured at ages 1 (birth weight) to 733 d. Two data samples were created, one with 71,867 records sampled from all herds (MISS), and the other with 74,601 records sampled from herds with no missing traits (NMISS). Records preadjusted to a fixed age were analyzed by a multiple-trait model (MTM), which included the effects of contemporary group, age of dam class, additive direct, additive maternal, and maternal permanent environment. Analyses were by REML, with five traits at a time. The random regression model (RRM) included the effects of age of animal, contemporary group, age of dam class, additive direct, additive maternal, permanent environment, and maternal permanent environment. All effects were modeled as cubic Legendre polynomials. These analyses were also by REML. Shapes of estimates of variances by MTM were mostly similar for both data sets for all except late ages, where estimates for MISS were less regular, and for birth weight with MISS. Genetic correlations among ages for the direct and maternal effects were less smooth with MISS. Genetic correlations between direct and maternal effects were more negative for NMISS, where few sires were maternal grandsires. Parameter estimates with RRM were similar to MTM cept that estimates of variances showed more artifacts for MISS; the estimates of additive direct-maternal correlations were more negative with both data sets and approached -1.0 for some ages with NMISS. When parameters of a growth model obtained by used for genetic evaluation, these parameters should be examined for consistency with parameters from MTM and prior information, and adjustments may be required to eliminate artifacts.

Age Factors↗

Prediction of length of hospital stay in neonatal units for very low birth weight infants.

OBJECTIVE: To develop models for estimating the length of hospital stay (LOS) of very low birth weight infants (VLBW), based on perinatal risk factors present during the first week of life and during the entire hospitalization period. STUDY DESIGN: The files of 155 VLBW were analyzed, and the influence of individual risk factors were initially evaluated by univariate analysis, using multiple-regression. Two mathematical models were built to estimate the LOS. RESULTS: The first model, using risk factors present during the first 3 days of life, is as follows: LOS = -0.074A + 22.06B + 22.85C - 16.78D - 2.07E + 10.51F + 203.12 (R2 = 0.63). (The letters are added to show what each number represents: A: birth weight; B: occurrence of respiratory distress syndrome; C: endotracheal intubation during resuscitation; D: 1-minute Apgar score; E: gestational age; F: presence of complications during delivery.) The second model, using factors present during the entire hospitalization period, is: LOS = 0.61G + 29.19H + 24.68I + 14.21J + 23.56K + 9.54L + 7.41M + 20.43 (R2 = 0.82). (G: age receiving nutritional support of > or = 120 kcal/kg per day; H: occurrence of systemic candidiasis; I: birth weight < 1000 gm; J: presence of delivery complication; K: occurrence of bronchopulmonary dysplasia; L: birth weight > or = 1000 gm and < or = 1249 gm; M: occurrence of anemia). CONCLUSION: Both models are applicable for estimating the hospitalization period, and the addition of variables present during the entire hospitalization period improved the accuracy of the model.

Brazil↗

The contribution of foods to the dietary lipid profile of a Spanish population.

OBJECTIVE: To identify the food that has the greatest effect on the variation in the percentage of energy intake derived from fat and saturated fatty acids for the consumption of a Spanish population. DESIGN: A cross-sectional study of food consumption, using the 24-hour recall method for three non-consecutive days, one of which was a non-working day. Subjects were interviewed by trained interviewers in the subjects' homes. We used multiple linear regression for statistical analysis. SETTING: The citizens of Reus. SUBJECTS: One thousand and sixty subjects over five years old, randomly selected from the population census of Reus. RESULTS: In both sexes, the foods that mainly determine a high consumption of fat are oil and red meat while those that determine a lower consumption of fat are bread, savoury cereals and fruit. The foods that mainly determine a high consumption of saturated fatty acids are red meat and whole-fat dairy products while those that determine a low consumption are bread, savoury cereals and fruit. CONCLUSIONS: In our population, feasible variations in the intake of some foods - less than one portion - would reduce the estimated percentage of energy intake derived from fat and saturated fatty acids by a quantity considered important for cardiovascular disease prevention. The periodic identification and quantification of the food that most affects the dietary fat profile will help in drawing up dietary guidelines with more reasonable strategies for consuming a healthier diet and decreasing the risk of developing nutritional disorders.

Adolescent↗

Controlling for socioeconomic confounding using regression methods.

STUDY OBJECTIVE: To describe the advantages of using Poisson regression methods as an alternative to standardisation when computing expected numbers of disease occurrences adjusted for possible confounding factors. The problem of assessing the adequacy of model fit when the expectations are small is addressed by analytical calculations and by simulation. The method is illustrated with data from the national register of childhood tumours. DESIGN: The tumour data are recorded in a national register. SETTING: England, Scotland, and Wales. SUBJECTS: The cases considered are all children registered with leukaemia or non-Hodgkin lymphoma under the age of 15 years between 1966-87. MAIN RESULTS: The methods show a significant variation of leukaemia incidence in relation to the Register General's standard region and a negative association with socioeconomic deprivation, as measured by the Townsend index. After allowing for these variables, the incidence seems to be reasonably homogeneous throughout the population, in the sense that the residual deviance does not seem to be much larger than would be expected by chance. CONCLUSIONS: The methods described have major advantages over standardisation in controlling for confounding, both in terms of flexibility of factor selection and assessment and also in the ability to determine whether there is residual variability of incidence after allowing for these factors.

Child, Preschool↗

Mapping and exclusion mapping of genomic imprinting effects in mouse F2 families.

Parent-of-origin effects were mapped by multimarker regression analysis in a cross between a high body weight selected line (DU6) and a control line (DUKs). The difference between F(2) progeny being heterozygous Qq and qQ (first allele is paternally derived) for grandpaternal Q and grandmaternal q alleles was genome-wide significant for the traits liver weight and spleen weight with a paternal imprinting effect at 1 cM on proximal chromosome 11. Suggestive imprinting effects (chromosome-wide error probability less than 0.05) were found for the traits body weight, liver weight, and kidney weight, and were located on chromosome 14 at 25 cM, 23 cM, and 32 cM, respectively. A genome-wide significant quantitative trait locus (QTL) for spleen weight at 26 cM slightly failed the suggestive significance level for imprinting. The effect was consistently maternal for all these traits on chromosome 14. Further suggestive imprinting effects were found for abdominal fat percentage on chromosome 3, for spleen weight on chromosome 5, and for liver weight on chromosome X. Our results are supported by a likely imprinting in a human genome region with homology to mouse chromosome 14 and agree well with the known imprinting of proximal chromosome 11 in the mouse.

Animals↗

Effects of systematic errors in blood pressure measurements on the diagnosis of hypertension.

OBJECTIVE: To estimate the effects of systematic errors in measurements of blood pressure on the diagnosis of hypertension. METHODS: We fitted regression curves to distributions of diastolic and systolic BP from recent Canadian and UK surveys and calculated the effect of systematic measurement errors on changes in the numbers of patients who would be classified hypertensive at thresholds of 85, 90 and 95 mmHg diastolic and 140 and 160 mmHg systolic pressure respectively. RESULTS: Overestimation of diastolic BP by 5 mmHg increases the number of patients whose diastolic BP exceeds 85, 90 and 95 mmHg by 102, 132 and 166% respectively. Equivalent underestimation causes 57, 62 and 67% respectively of hypertensive patients to be missed. If systematic error in diastolic pressure is limited to +/-1 mmHg the diagnosis errors are between -15 and +23%. Overestimation of systolic BP by 3 and 5 mmHg increases the number classified as hypertensive by 24 and 43% respectively. Equivalent underestimation causes 19 and 30% of patients with systolic hypertension to be missed. CONCLUSIONS: Small systematic errors in BP measurements may cause large variations in the proportion of patients diagnosed as hypertensive. To limit over- or under-diagnosis of diastolic hypertension to approximately 20%, systematic errors in diastolic BP measurements should be limited to 1 mmHg. An uncertainty of 3 mmHg may be adequate for detecting systolic hypertension.

Blood Pressure↗

A semimechanistic-physiologic population pharmacokinetic/pharmacodynamic model for neutropenia following pemetrexed therapy.

PURPOSE: The objectives of these analyses were to (1) develop a semimechanistic-physiologic population pharmacokinetic/pharmacodynamic (PK/PD) model to describe neutropenic response to pemetrexed and to (2) identify influential covariates with respect to pharmacodynamic response. PATIENTS AND METHODS: Data from 279 patients who received 1,136 treatment cycles without folic acid or vitamin B12 supplementation during participation in one of eight phase II cancer trials were available for analysis. Starting doses were 500 or 600 mg pemetrexed per m2 body surface area (BSA), administered as 10-min intravenous infusions every 21 days (1 cycle). The primary analyses included 105 patients (279 cycles) for which selected covariates-including vitamin deficiency marker data (i.e., homocysteine, cystathionine, methylmalonic acid, and methylcitrate [I, II, and total] plasma concentrations)-were available. Classical statistical multivariate regression analyses and a semimechanistic-physiologic population PK/PD model were used to evaluate neutropenic response to single-agent pemetrexed administration. RESULTS: The timecourse of neutropenia following single-agent pemetrexed administration was adequately described by a semimechanistic-physiologic model. Population estimates for system-based model parameters (i.e., baseline neutrophil count, mean transit time, and the feedback parameter), which mathematically represent current understanding of the process and physiology of hematopoiesis, were consistent with previously reported values. The population PK/PD model included homocysteine, cystathionine, albumin, total protein, and BSA as covariates relative to neutropenic response. CONCLUSION: These results support the programmatic decision to introduce folic acid and vitamin B12 supplementation during pemetrexed clinical development as a means of normalizing patient homocysteine levels, thereby managing the risk of severe neutropenia secondary to pemetrexed administration. The current results also suggest that the addition of vitamin B6 supplementation to normalize patient cystathionine levels may further decrease the incidence of grade 4 neutropenia following pemetrexed administration. The results also suggest the use of folic acid as a means of lessening hematologic toxicity following administration of cytotoxic agents other than antifolates.

Adult↗

Risk factors for tubal infertility. Influence of history of prior pelvic inflammatory disease.

In order to explore possible etiologic differences between tubal infertility in women who had been physician-diagnosed as having pelvic inflammatory disease ("overt" PID) and in women who had not ("silent" pelvic inflammatory disease), we made use of self-reported data from a large, population-based, case-control study of infertility in King County, Washington. Responses from 33 infertile women with no history of physician-reported PID and 129 infertile women with such a history were compared to those of 501 fertile women. No cultures or blood for antibody titers were obtained. Logistic regression was used to compute the relative risks for silent and overt PID-related tubal dysfunction associated with various lifestyle and contraceptive habits in an effort to identify practices that potentially affect these outcomes. In general, practices associated with an increased risk of overt tubal disease, such as use of Dalkon Shield and other types of intrauterine devices, were also associated with an increased risk of silent tubal disease, but to a lesser extent. Women who used oral contraceptives for longer than three years had a decreased risk for silent disease (relative risk = 0.5, 95% confidence interval = 0.3-0.8), but their risk for overt disease did not decrease to the same extent (relative risk = 0.9, 95% confidence interval = 0.3-2.5). These results suggest that silent and overt tubal disease share many common lifestyle risk factors.

Adult↗

A covariance function for feed intake, live weight, and milk yield estimated using a random regression model.

To enable investigation of genetic variation during early lactation in heifers, multitrait covariance functions were used to describe genetic covariances among feed intake, live weight, and milk yield during the first 15 wk of lactation (n = 628). Random regression models were used to estimate covariance functions for the additive genetic and permanent environmental effects. Fixed effects were date of the week that records were collected, a group effect, and week of lactation. Second or third order polynomials were sufficient to describe the additive genetic variation for milk yield, dry matter intake, and live weight during the first 15 wk of lactation. Estimates for the genetic covariance function demonstrated that a high milk yield is only moderately correlated with high feed intake (0.21) but is very strongly correlated to an increase of intake and a loss of live weight during the first 15 wk of lactation. Levels of weight and intake were correlated strongly (0.81). The reduced fit covariance function was used to estimate genetic correlations between traits at different lactation stages. Estimates for the genetic correlations between wk 1 and 15 were 0.62, 0.24, and 0.79 for milk yield, dry matter intake, and live weight, respectively. Feed intake during early lactation was negatively correlated with milk yield, but feed intake during the later weeks was positively correlated with milk yield. The implication is that when selection is for a linear combination of milk yield, feed intake, and live weight (i.e., energy balance or efficiency), it is important to consider when each trait is measured during lactation.

Analysis of Variance↗