Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Statistical analysis of skin tumor data from Tg.AC mouse bioassays.

New strategies for identifying chemical carcinogens and assessing risk have been proposed based on the Tg.AC (zetaglobin promoted v-Ha-ras) transgenic mouse. Preliminary studies suggest that the Tg. AC mouse bioassay may be an effective means of quickly evaluating the carcinogenic potential of a test agent. The skin of the Tg.AC mouse is genetically initiated, and the induction of epidermal papillomas in response to dermal or oral exposure to a chemical agent acts as a reporter phenotype of the activity of the test chemical. In Tg.AC mouse bioassays, the test agent is typically applied topically for up to 26 weeks, and the number of papillomas in the treated area is counted weekly. Statistical analyses are complicated by within-animal and serial dependency in the papilloma counts, survival differences between animals, and missing data. In this paper, we describe a statistical model for the analysis of skin tumor data from a Tg.AC mouse bioassay. The model separates effects on papilloma latency and multiplicity and accommodates important features of the data, including variability in expression of the transgene and dependency in the tumor counts. Methods are described for carcinogenicity testing and risk assessment. We illustrate our approach using data from a study of the effect of 2,3,7, 8-tetrachlorodibenzo-p-dioxin (TCDD) exposure on tumorigenesis.

Administration, Topical↗

The Pediatric Cancer Quality of Life Inventory: a modular approach to measuring health-related quality of life in children with cancer.

Measurement of pediatric cancer patients' health-related quality of life (HRQL) in phase III randomized, controlled clinical trials is being recognized increasingly as an essential component in evaluating the comprehensive health outcomes of modern anti-neoplastic treatment protocols. Use of a brief core measure of HRQL plus disease-specific symptom modules is a way to assess specific HRQL outcomes with a minimum of subject burden. Demonstrating a measure's feasibility, reliability and validity also represents children's ability to provide reliable and valid responses to HRQL questions. The Pediatric Cancer Quality of Life Inventory (PCQL) Modular Approach consists of a 15-item core measure of HRQL and 2 specific symptom modules: pain and nausea. To validate a patient-report form and a parent-report form, the PCQL was administered to 291 pediatric cancer patients and to their parents. Feasibility and range of measurement, as well as patient-parent concordance, were assessed. Internal consistency reliability was assessed via Cronbach's alpha. Validity was determined by the known-groups approach and by correlating PCQL scores with days missed from school. Patients had minimal missing data, and the range of measurement for the items was good. Patient-parent concordance was large but not perfect. For both patient and parent forms, internal consistency reliability of the PCQL core scale (0.83 and 0. 86, respectively) was strong. The internal consistency reliabilities of the 2 symptom modules for both patient and parent forms were in the acceptable range for group comparisons. Regarding clinical validity, the core scale and the 2 symptom modules distinguished between patients on and off treatment for both patient and parent reports. Further, both patient and parent reports correlated with days of missed school in the past 6 and 12 months. The PCQL Modular Approach has demonstrated acceptable internal consistency reliability and clinical validity for both patient-report and parent-report forms. By implication, children are capable of providing reliable and valid responses to these HRQL questions.

Child↗

Case-control single-marker and haplotypic association analysis of pedigree data.

Related individuals collected for use in linkage studies may be used in case-control linkage disequilibrium analysis, provided one takes into account correlations between individuals due to identity-by-descent (IBD) sharing. We account for these correlations by calculating a weight for each individual. The weights are used in constructing a composite likelihood, which is maximized iteratively to form likelihood ratio tests for single-marker and haplotypic associations. The method scales well with increasing pedigree size and complexity, and is applicable to both autosomal and X chromosomes. We apply the approach to an analysis of association between type 2 diabetes and single-nucleotide polymorphism markers in the PPAR-gamma gene. Simulated data are used to check validity of the test and examine power. Analysis of related cases has better power than analysis of population-based cases because of the increased frequencies of disease-susceptibility alleles in pedigrees with multiple cases compared to the frequencies of these alleles in population-based cases. Also, utilizing all cases in a pedigree rather than just one per pedigree improves power by increasing the effective sample size. We demonstrate that our method has power at least as great as that of several competing methods, while offering advantages in the ability to handle missing data and perform haplotypic analysis.

Alleles↗

Dementia, cognitive impairment and mortality in persons aged 65 and over living in the community: a systematic review of the literature.

BACKGROUND: No recent attempt has been made to synthesise information on mortality and dementia despite the theoretical and practical interest in the topic. Our objective was to estimate the influence on mortality of cognitive impairment and dementia. METHODS: Data sources were Medline, Embase, personal files and colleagues' records. Studies were considered if they included a majority of persons aged 65 and over at baseline either drawn from a total community sample or drawn from a random sample from the community. Samples from health care facilities were excluded. The search located 68 community studies. Effect sizes were extracted from the studies and if they were not included in the published studies, effect sizes were calculated where possible: this was possible for 23 studies of cognitive impairment and 32 of dementia. No attempt was made to contact authors for missing data. RESULTS: For the studies of cognitive impairment Fisher's method (a vote counting method), gave a p-value (from eight studies) of 0.00001. For studies of dementia, age-adjusted confidence intervals (CI) were pooled (odds ratio (OR) 2.63 with 95% CI 2.17 to 3.21 from six studies). CONCLUSIONS: Levels of cognitive impairment commonly found in community studies give rise to an increased risk of mortality, and this appears to be true even for quite mild levels of impairment. The analysis confirms the increased risk of mortality for dementia, but reveals a dearth of information on the causes of the excess mortality and on possible effect modification by age, dementia subtype or other variables.

Age Distribution↗

A comparison of two measures of stage of change for smoking cessation.

AIMS: To compare two questionnaires used to identify the stages of change of current and former smokers: a conventional five-item questionnaire and an alternative one-item questionnaire. DESIGN AND SETTING: Mail surveys of 1167 ever smokers in Geneva (Switzerland), conducted in 1997. Participants were classified into five stages: precontemplation, contemplation, preparation, action and maintenance. Other questions covered smoking-related behaviours, attitudes and self-efficacy. FINDINGS: Only 62% of participants were classified as being in the same stage by the two questionnaires (weighted kappa = 0.69). The five-item questionnaire produced more missing data (8%) than the one-item questionnaire (2%, p < 0.001). Using the conventional questionnaire, the precontemplation stage included a group of smokers who had absolutely no intention of stopping smoking and a group more prone to change, and the preparation stage included only 43% of people who had made a "firm decision" to quit smoking in the next 30 days. Using the alternative questionnaire, the contemplation stage was also quite heterogeneous. The action stage included over 35% of people who were still smoking occasionally, whichever questionnaire was used. CONCLUSIONS: The single-item questionnaire was better at avoiding missing responses. However, both staging questionnaires classified smokers in heterogeneous groups, and both misclassified many occasional smokers and ex-smokers, which suggests that a discrete five-stage model does not fit reality well. This may reflect underlying conceptual issues, notably that the classic definition of stages incorporates separate dimensions, albeit incompletely (current behaviour, quit attempts, intention to change, time). Both theoretical and methodological developments are needed to overcome these problems.

Adaptation, Psychological↗

Clinical and psychometric validation of an EORTC questionnaire module, the EORTC QLQ-OES18, to assess quality of life in patients with oesophageal cancer.

Quality of life (QOL) assessment requires clinically relevant questionnaires that yield accurate data. This study defined measurement properties and the clinical validity of the European Organisation for Research and Treatment of Cancer (EORTC) questionnaire module to assess QOL in oesophageal cancer. The oesophageal module the QLQ-OES24 and core questionnaire, the Quality of Life-Core 30 questionnaire (QLQ-C30) was administered patients undergoing treatment with curative (n=267) or palliative intent (n=224) and second assessments performed 3 months or 3 weeks later respectively. Psychometric tests examined scales and measurement properties of the module. Questionnaires were well accepted, compliance rates were high and less than 2% of items had missing data. Multi-trait scaling analyses and face validity refined the module to four scales and six single items (QLQ-OES18). Selective scales distinguished between clinically distinct groups of patients and demonstrated treatment-induced changes over time. The EORTC QLQ-OES18 demonstrates good psychometric and clinical validity. It is recommended for use with the core questionnaire, the QLQ-C30, to assess QOL in patients with oesophageal cancer.

Adenocarcinoma↗

Beverage intake among preschool children and its effect on weight status.

OBJECTIVE: The obesity epidemic in the United States continues to increase. Because obesity tends to track over time, the increase in overweight among young children is of significant concern. A number of eating patterns have been associated with overweight among preschool-aged children. Recently, 100% fruit juice and sweetened fruit drinks have received considerable attention as potential sources of high-energy beverages that could be related to the prevalence of obesity among young children. Our aim was to evaluate the beverage intake among preschool children who participated in the National Health and Nutrition Examination Survey 1999-2002 and investigate associations between types and amounts of beverages consumed and weight status in preschool-aged children. METHODS: We performed a secondary analysis of the data from the National Health and Nutrition Examination Survey 1999-2002, which is a continuous, cross-sectional survey of a nationally representative sample of the noninstitutionalized population of the United States. It included the collection of parent reported demographic descriptors, a 24-hour dietary recall, a measure of physical activity, and a standardized physical examination. The 24-hour dietary recall was obtained in person by a trained interviewer and reflected the foods and beverages that were consumed by the participant the previous day. The National Health and Nutrition Examination Survey food groups were classified on the basis of the US Department of Agriculture's Food and Nutrient Database for Dietary Studies. We reviewed the main food descriptors used and classified all beverages listed. One hundred percent fruit juice was classified as only beverages that contained 100% fruit juice, without sweetener. Fruit drinks included any sweetened fruit juice, fruit-flavored drink (natural or artificial), or drink that contained fruit juice in part. Milk included any type of cow milk and then was subcategorized by percentage of milk fat. Any sweetened soft drink, caffeinated or uncaffeinated, was categorized as soda. Diet drinks included any fruit drink, tea, or soda that was sweetened by low-calorie sweetener. Several beverages were removed from the analysis because of low frequency of consumption among the sample. Water was not included in the analysis because it is not part of the US Department of Agriculture's Food and Nutrient Database categories. For the purposes of this analysis, the beverages were converted and reported as ounces, rather than grams, as reported by the National Health and Nutrition Examination Survey, to make it more clinically relevant. The child's BMI percentile for age and gender were calculated on the basis of Centers for Disease Control and Prevention criteria and used to identify children's weight status as underweight (< 5%), normal weight (5% to < 85%), at risk for overweight (85% to < 95%), or overweight (> or = 95%). Because of the small number of children in the underweight category, they were included in the normal-weight category for this analysis. Data were analyzed using SUDAAN 9.0.1 statistical software programs. SUDAAN allows for improved accuracy and validity of results by calculating test statistics for the stratified, multistage probability design of the National Health and Nutrition Examination Survey. Sample weights were applied to all analyses to account for unequal probability of selection from oversampling low-income children and black and Mexican American children. Descriptive and chi2 analyses and analysis of covariance, adjusting for age, gender, ethnicity, household income, energy intake, and physical activity, were conducted. RESULTS: All children who were aged 2 to 5 years were identified (N = 1572). Those with missing data were removed from additional analysis, resulting in a final sample of 1160 preschool children. Of the 1160 children analyzed, 579 (49.9%) were male. White children represented 35%, black children represented 28.3%, and Hispanic children represented 36.7% of the sample. Twenty-four percent of the children were overweight or at risk for overweight (BMI > or = 85%), and 10.7% were overweight (BMI > or = 95%). There were no statistically significant differences in BMI between boys and girls or among the ethnicities. Overweight children tended to be older (mean age: 3.83 years) compared with the normal-weight children (mean age: 3.48 years). Eighty-three percent of children drank milk, 48% drank 100% fruit juice, 44% drank fruit drink, and 39% drank soda. Whole milk was consumed by 46.5% of the children, and 3.1% and 5.5% of the children consumed skim milk and 1% milk, respectively. Preschool children consumed a mean total beverage volume of 26.93 oz/day, which included 12.32 oz of milk, 4.70 oz of 100% fruit juice, 4.98 oz of fruit drinks, and 3.25 oz of soda. Weight status of the child had no association with the amount of total beverages, milk, 100% fruit juice, fruit drink, or soda consumed. There was no clinically significant association between the types of milk (percentage of fat) consumed and weight status. In analysis of covariance, daily total energy intake increased with increased consumption of milk, 100% fruit juice, fruit drinks, and soda. However, there was not a statistically significant increase in BMI on the basis of quantity of milk, 100% fruit juice, fruit drink, or soda consumed. CONCLUSIONS: On average, preschool children drank less milk than the 2005 Dietary Guidelines for Americans recommendation of 16 oz/day. Only 8.6% drank low-fat or skim milk, as recommended for children who are older than 2 years. On average, preschool children drank < 6 oz/day 100% fruit juice. Increased beverage consumption was associated with an increase in the total energy intake of the children but not with their BMI. Prospectively studying preschool children beyond 2 to 5 years of age, through their adiposity rebound (approximately 5.5-6 years) to determine whether there is a trajectory increase in their BMI, may help to clarify the role of beverage consumption in total energy intake and weight status.

Beverages↗

Neurosis and mortality in persons aged 65 and over living in the community: a systematic review of the literature.

BACKGROUND: No previous attempt has been made to synthesise information on mortality and neurosis in older people. Our objective was to estimate the influence on mortality of various types of neurosis in the older population. METHODS: Data sources were: Medline; Embase; and personal files. Studies were considered if they included a majority of persons aged 65 and over at baseline either drawn from a total community sample or drawn from a random sample from the community. Studies which sampled from a larger age range were also included if it was possible to retrieve results about those aged 65 and over. Samples from health care facilities were excluded. Effect sizes were extracted from the papers and if they were not included in the published papers effect sizes were calculated if possible. No attempt was made to contact authors for missing data. RESULTS: We found seven reports (six of which used a neurosis diagnosis and one which used a symptom scale). Using Fisher's method we found an increase in mortality which was not significant (p = 0.08). CONCLUSION: There have been few studies, and the evidence is weakly in favour of an increased mortality risk.

Aged↗

Overview of the national spinal cord injury statistical center database.

OBJECTIVE: An evaluation of the history, design, and status of the database of the National Spinal Cord Injury Statistical Center (NSCISC) was undertaken to identify its continued relevance. RESEARCH DESIGN: A systematic review was conducted of goals, content, and quality control procedures, as well as its suitability and public availability for conducting future epidemiologic and health services research. RESULTS: The NSCISC database contains information on approximately 29,000 persons injured since 1973 and treated at any regional model spinal cord injury system within 1 year of injury. The NSCISC database is structured longitudinally with data collected at discharge, 1 year after injury, 5 years after injury, and every 5 years thereafter. The database includes information on demographics, injury severity, medical complications, surgical procedures, types and amounts of therapy, length of stay, charges, and both short-term and long-term treatment outcomes. Strengths include large sample size, use of valid and reliable measures, geographic and patient diversity, comprehensiveness, availability of long-term prospective follow-up information, good case identification, and rigorous quality control procedures. Limitations include lack of population basis, inclusion of only model system patients, losses to follow-up, and other missing data. Recent content additions include detailed information on each treatment phase, depression, substance abuse, environmental barriers to community integration, and patient identifying information. A process exists for researchers to gain access to the data. CONCLUSIONS: The database remains a valuable resource. Future plans include linkage to other databases to enhance research capability, a published research compendium, and development of a user's guide to facilitate database usage.

Databases as Topic↗

[Prognostic factors in operated non-small cell cancer of the lung. Study from a randomized therapeutic trial].

This article presents the results of a prognostic study of primary resected lung cancer (non-small cell). The data result from a randomised clinical trial of immunotherapy with a non-specific adjuvant; the follow-up was between four to seven years. Thirty-five clinical, biological and anatomo-pathological parameters were gathered at the time of inclusion in the trial. The response criteria used were survival without recurrence and total survival. A multivariate analysis using the Cox's model was carried out for each criterion. At the reference date of the 1st April 1985, 125 relapses and 132 deaths were counted amongst 219 patients; there was only one patient lost to follow-up and only 39 missing data were observed. The negative therapeutic results of the immunotherapy used were confirmed by this new intermediate analysis. The rate of survival without recurrence at 5 years was 43% and the overall survival at five years was 42%. The use of Cox's model to show the prognostic information at the 5% level for survival without recurrence could be summarised by five factors: main staging (the prognostic factor), leucocytosis, the cutaneous reaction to proteus, Karnofsky index and presence of physical signs. For stages I and II the outcome was identical and no factor was predictive at the 5% level. For stage III the cutaneous reaction to proteus and leucocytosis were prognostic. For overall survival, the prognostic information at the 5% level could be summarised by five factors: staging (main prognostic factor), leucocytosis, Karnofsky index, presence of physical signs and lymphocytosis. For stages I and II whose outcome was identical only Karnofsky index and lymphocytosis were predictive at the 5% level.(ABSTRACT TRUNCATED AT 250 WORDS)

Carcinoma, Non-Small-Cell Lung↗

Multipoint linkage-disequilibrium mapping with haplotype-block structure.

The HapMap Project is providing a great deal of new information on high-resolution haplotype structure in various human populations. This information has the potential to greatly increase the power of association mapping for a fixed amount of genotyping. A number of methods have been proposed for the identification of haplotype blocks, common haplotypes, and tagging single-nucleotide polymorphisms. Here, we build on this work by developing novel methods for case-control multipoint linkage-disequilibrium (LD) mapping that gain power and speed by making explicit use of the inferred block structure. Specifically, we developed a virtual-variant approach that uses the haplotype-block information to greatly increase power for detection of untyped common variants associated with a trait. Because full multipoint LD mapping can be slow, we exploited the haplotype-block information to develop a fast single-block multipoint mapping method. Our methods are appropriate for genotype data and take into account the uncertainty in phase. We describe the methods in the context of case-parents trios, although they are also applicable to unrelated cases and controls. Our simulations indicate that the most important gains from taking into account the haplotype-block structure at the analysis stage of multipoint LD mapping come from (1) greatly increased power to detect association with untyped variants and (2) greatly improved localization of untyped variants associated with the trait. More-modest gains are obtained in improving power to detect association with a variant that is typed with a moderate amount of missing data. The methods are applied to a Crohn disease data set.

Algorithms↗

The validity of explicit indicators of prescribing appropriateness.

OBJECTIVE: To assess, from the perspective of UK hospital doctors, the content validity and operational validity of a set of 14 previously developed explicit indicators of the appropriateness of long-term prescribing started during a hospital admission. METHOD: A combination of data extraction from medical records and qualitative interviews with a maximum variability sample of hospital doctors. PARTICIPANTS: The indicators were applied to 132 new prescriptions, intended for long-term use, prescribed for 61 patients; 36 doctors, of various grades, were purposively selected for interview. RESULTS: Appropriate prescribing was viewed as prescribing that was indicated, necessary, evidence based (using a broad meaning of 'evidence') and of acceptable cost and risk-benefit ratio. These concepts applied to individual drugs for individual patients, rather than at a more general, public health level. Where drugs had failed an indicator, rationales were explored. Often, it was missing data in the medical notes that had resulted in the drug failing the indicator. CONCLUSIONS: The 14 indicators were considered to have content validity, reflecting all aspects of appropriate prescribing discussed by the doctors. Their operational validity was less clear-cut, due to the lack of necessary data in the medical notes. This has implications for the use of explicit indicators for assessing prescribing appropriateness, as these hospital doctors did not consider that the data required for objective, systematic assessment of prescribing would ever be recorded in hospital medical notes.

Attitude of Health Personnel↗

[Cervical cancer screening for high risk women: is it possible? Results of a cervical cancer screening program in three suburban districts of Lyon].

Between november 1993 and october 1996, a cervical screening program was proposed for women 25-65-year-old who tend to have little or no medical supervision, in three suburban districts of Lyon. The data and results of the two last Pap-smears have been collected together with details of gynecological follow-up. Both general practitioners and gynaecologists were actively involved. A total of 3,792 women (12.3% of the target) were registered, with a larger proportion of women over 60 (17.7%). According to the "Consensus of Lille", only 403 women (34.4%) had adequate screening (over 50 y: 25.8%, 35-49 y: 39.4%, 25-35 y: 36.5%) and 2,489 women had inappropriate gynaecological follow-up: no smear for 185 women (4.9%) and inadequate schedule of follow-up visits for 476 others (12.5%). Missing data (date or results of Pap smear) were noted for 1,828 patients (48.2%). The screening procedure for women over 50 years was carried out mainly by general practitioner. Of 3,127 registered smears, 62 positive results were found (2.1%). Of these women, 9 were lost to follow-up and 4 did not have appropriate tests. Others results were: 27 negative further investigations, 9 CIN1, 7 CIN2, 3 CIN3, 1 in situ carcinoma and 2 invasive carcinoma. Despite low participation, this pilot study indicates that a procedure can be established to integrate high risk women in cervical cancer screening programme. Active participation of general practitioners is essential.

Adult↗

Intentionally incomplete longitudinal designs: I. Methodology and comparison of some full span designs.

Longitudinal designs are important in medical research and in many other disciplines. Complete longitudinal studies, in which each subject is evaluated at each measurement occasion, are often very expensive and motivate a search for more efficient designs. Recently developed statistical methods foster the use of intentionally incomplete longitudinal designs that have the potential to be more efficient than complete designs. Mixed models provide appropriate data analysis tools. Fixed effect hypotheses can be tested via a recently developed test statistic, FH. An accurate approximation of the statistic's small sample non-central distribution makes power computations feasible. After reviewing some longitudinal design terminology and mixed model notation, this paper summarizes the computation of FH and approximate power from its non-central distribution. These methods are applied to obtain a large number of intentionally incomplete full-span designs that are more powerful and/or less costly alternatives to a complete design. The source of the greater efficiency of incomplete designs and potential fragility of incomplete designs to randomly missing data are discussed.

Longitudinal Studies↗

Predictive factors for sacral neuromodulation in chronic lower urinary tract dysfunction.

OBJECTIVES: To investigate data from 211 patients who underwent a trial stimulation (percutaneous nerve evaluation [PNE]) to determine the clinical parameters that can enhance the prediction of PNE success. The advantageous effect of sacral neuromodulation depends on the accurate identification of suitable candidates during the preimplantation PNE. METHODS: A total of 211 patients (161 women and 50 men), with refractory urge incontinence, urgency-frequency syndrome, and urinary retention, underwent a PNE. Patient data (demographics, medical history, urologic investigations, and diagnosis) were collected. The PNE results were evaluated from a voiding diary and patient history. More than 50% improvement of voiding parameters was considered a successful PNE, and those patients were selected for implantation. Logistic regression analysis was performed. The factors tested for predicting the test result were sex, patient age, diagnosis, previous surgery, neurogenic bladder dysfunction, duration of complaints, and previous treatments. RESULTS: The PNEs were positive in 85 patients (40.3%) and negative in 105 patients (49.8%). In 18 patients (8.5%), the test electrode had migrated; 3 more patients were not assessable and were also excluded. Missing data on the variable "duration of complaints" reduced the number of patients in the analyses from 190 to 174 patients. CONCLUSIONS: Intervertebral disk prolapse, duration of complaints, neurogenic bladder dysfunction, and urge incontinence were found to be significant predictive factors. However, a PNE remains necessary to evaluate a patient's chance of implant success objectively.

Adult↗

Interpretation of data on dietary intake.

Although this discussion has focused on the interpretation of dietary data, assuming that it is representative of actual and usual intake and that the nutrient analysis based on it involved the use of up-to-date food composition tables, the readers should be sensitive to other potential sources of error or bias in obtaining information on food and/or nutrient intake. These include errors due to irregularity of food consumption, under- or overreporting of intake, errors in reporting either the amount or the description of the food consumed, recording errors on the part of the interviewer or coding errors on the part of the coder, limitations in the tables of food composition due to missing data for certain nutrients in certain foods or to biologic variability in the same foods from different sources or in those marketed under different conditions, imputed values, the unknown composition of formulated foods or foods prepared from home or commercial recipes, differences in bioavailability of nutrients as a function of the diet, or the use of abridged tables of food composition. In spite of the many unresolved issues relating to dietary standards and the interpretation of dietary intake data, we are still able to make a reasonable assessment of dietary adequacy of groups and individuals with our current system, which is viewed as a unique federal resource. It is hoped that the eventual passage of the National Nutrition Monitoring and Related Research Act will provide both the impetus and the resources to permit us to develop a more sophisticated system for assessing both food intake and nutritional status.

Data Interpretation, Statistical↗

Some methodologic issues in analyzing data from a randomized adolescent tobacco and alcohol use prevention trial.

Three issues concerning the design and analysis of randomized behavioral intervention studies are illustrated and discussed within the framework of a tobacco and alcohol prevention trial among migrant Latino adolescents. The first issue arises when subjects are randomized in clusters rather than individually. Because subject observations cannot be assumed to be independent, information pertaining to the degree of clustering must be reported, and analyses must take the clustering into account. The second issue concerns the impact of compliance to the intervention and the importance of measuring compliance in the experimental and attention-control groups. A compliance analysis should control for participant contact with study personnel. Investigators must consider ways of constructing a compliance measure that is common to both conditions. Third, because outcomes are measured repeatedly over time, we illustrate the importance of assessing the impact of missing-data patterns on outcomes and the extent to which the patterns may modify the treatment effect.

Adolescent↗

Clustering microarray gene expression data using weighted Chinese restaurant process.

MOTIVATION: Clustering microarray gene expression data is a powerful tool for elucidating co-regulatory relationships among genes. Many different clustering techniques have been successfully applied and the results are promising. However, substantial fluctuation contained in microarray data, lack of knowledge on the number of clusters and complex regulatory mechanisms underlying biological systems make the clustering problems tremendously challenging. RESULTS: We devised an improved model-based Bayesian approach to cluster microarray gene expression data. Cluster assignment is carried out by an iterative weighted Chinese restaurant seating scheme such that the optimal number of clusters can be determined simultaneously with cluster assignment. The predictive updating technique was applied to improve the efficiency of the Gibbs sampler. An additional step is added during reassignment to allow genes that display complex correlation relationships such as time-shifted and/or inverted to be clustered together. Analysis done on a real dataset showed that as much as 30% of significant genes clustered in the same group display complex relationships with the consensus pattern of the cluster. Other notable features including automatic handling of missing data, quantitative measures of cluster strength and assignment confidence. Synthetic and real microarray gene expression datasets were analyzed to demonstrate its performance. AVAILABILITY: A computer program named Chinese restaurant cluster (CRC) has been developed based on this algorithm. The program can be downloaded at http://www.sph.umich.edu/csg/qin/CRC/.

Algorithms↗