Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Design and analysis of survival data under an integrated type-II interval censoring scheme.

This article considers the planning of Type-II interval censored life tests for exponentially distributed lifetimes. In particular, we explore two criteria that are based on the minimization of the asymptotic variance of maximum likelihood estimator of the model parameter and the associated cost of the experiment. A numerical study is conducted to evaluate the "optimal" plans based on these two criteria. The results of this study should provide practitioners useful insight in planning a Type-II interval censoring life test.

Algorithms↗

Composition-on-composition regression analysis for multi-omics integration of metagenomic data.

MOTIVATION: Compositional data are frequently encountered in many disciplines, such as in next-generation sequencing experiments widely used in biomedical studies. Regression analysis with compositional data as either responses or predictors has been well studied. However, when both responses and predictors are compositional, the inventory of analysis tools is surprisingly limited, especially in the high-dimensional setting. Among the few existing methods, most of them rely on a log-ratio transformation to move compositional data from the simplex to real numbers. Yet, a serious weakness of these methods is their failure to handle the substantial fraction of zeroes observed in data collected from next-generation sequencing experiments. RESULTS: To investigate associations between two high-dimensional multi-omics compositions, we propose a composition-on-composition (COC) regression analysis method which does not require log-ratio transformations and hence can handle zeroes in the data. To account for high dimensionality, we estimate regression coefficients using a penalized estimation equation approach. Finally, inference procedures for COC regression are also proposed. Superior performance of COC is demonstrated through both comprehensive numerical simulations and case studies. AVAILABILITY AND IMPLEMENTATION: Source R codes to implement COC method is available at https://github.com/nrios4/COC.

Regression Analysis↗

NOODAI: a webserver for network-oriented multi-omics data analysis and integration pipeline.

SUMMARY: Omics profiling has proven of great use for unbiased and comprehensive identification of key features that define biological phenotypes and underlie medical conditions. While each omics profile assists characterization of specific molecular components relevant for the studied phenotype, their joint evaluation can offer deeper insights into the overall mechanistic functioning of biological systems. Here, we introduce an approach where, starting from representative traits (e.g. differentially expressed elements) obtained for each omics profile, we construct and analyze joint interaction networks. The resulting networks rely on the existing knowledge of confident interactions among biological entities. We use these maps to identify and describe central elements, which connect multiple entities characteristic of the studied phenotypes and we leverage MONET network decomposition tool in order to highlight functionally connected network modules. In order to enable broad usage of this approach, we developed the NOODAI software platform, which enables integrative omics analysis through a user-friendly interface. The analysis outcomes are presented both as raw output tables as well as informative summary plots and written reports. Since the MONET tool enables the use of algorithms with strong performance in identifying disease-relevant modules, NOODAI software platform can be of a high value for analyzing clinical multi-omics datasets. AVAILABILITY AND IMPLEMENTATION: NOODAI is freely accessible at https://omics-oracle.com. Source code is available under GPL3 at: https://github.com/TotuTiberiu/NOODAI with the DOI: 10.5281/zenodo.17203984.

Software↗

PROSPECT improves cis-acting regulatory element prediction by integrating expression profile data with consensus pattern searches.

Consensus pattern and matrix-based searches designed to predict cis-acting transcriptional regulatory sequences have historically been subject to large numbers of false positives. We sought to decrease false positives by incorporating expression profile data into a consensus pattern-based search method. We have systematically analyzed the expression phenotypes of over 6000 yeast genes, across 121 expression profile experiments, and correlated them with the distribution of 14 known regulatory elements over sequences upstream of the genes. Our method is based on a metric we term probabilistic element assessment (PEA), which is a ranking of potential sites based on sequence similarity in the upstream regions of genes with similar expression phenotypes. For eight of the 14 known elements that we examined, our method had a much higher selectivity than a naïve consensus pattern search. Based on our analysis, we have developed a web-based tool called PROSPECT, which allows consensus pattern-based searching of gene clusters obtained from microarray data.

5' Untranslated Regions↗

Individual monitoring for internal exposure in Europe and the integration of dosimetric data.

The European Radiation Dosimetry Group, EURADOS, established a working group consisting of experts whose aim is to assist in the process of harmonisation of individual monitoring as part of the protection of occupationally exposed workers. A catalogue of facilities and internal dosimetric techniques related to individual monitoring in Europe has been completed as a result of this EURADOS study. A questionnaire was sent in 2002 to services requesting information on various topics including type of exposures, techniques used for direct and indirect measurements including calibration and sensitivity data and the methods employed for the assessment of internal doses. Information relating to Quality Control procedures for direct and indirect measurements, Quality Assurance Programmes in the facilities and legal requirements for "approved dosimetric services" were also considered. A total of 71 completed questionnaires were returned by internal dosimetry facilities in 26 countries. This results in an overview of the actual status of the processes used in internal exposure estimation in Europe. In many ways harmonisation is a reality in internal dose assessments, especially when taking into account the measurements of the activity retained or excreted from the body. However, a future study detailing the estimation of minimum detectable activity in the laboratories is highly recommended. Points to focus on in future harmonisation activities are as follows: the process of calculation of doses from measured activity, establishment of guidelines, similar dosimetric tools and application of the same ICRP recommendations. This would lead to a better and more harmonised approach to the estimation of internal exposures in all European facilities.

Body Burden↗

Do participants' reports of symptom prevalence or severity vary by interviewer gender?

BACKGROUND: Although interviews are commonly used to gather research data, the integrity of interview data can be threatened by discrepancies between interviewers and respondents on such characteristics as race, gender, or age. OBJECTIVES: To determine if participants' reports of the prevalence and severity of 30 symptoms varied as a function of interviewer gender. Symptoms that were assessed included general physical symptoms (diarrhea, headaches), psychological symptoms (feel blue and depressed, worry about nervous breakdown), and menopausal symptoms (hot flashes, vaginal dryness). METHOD: Structured telephone interviews were completed by 137 women who were a mean of 56.5 years old (SD = 11.1, range 36-83) and a mean of 38.8 months (SD = 23.6) postdiagnosis of breast cancer. Interviewers included two women and two men. RESULTS: Symptom prevalence and severity did not vary as a function of interviewer gender. CONCLUSIONS: Findings suggest that both male and female interviewers can be used successfully to assess participants' reports of physical, psychological, or menopausal symptoms.

Adult↗

Identification of expression patterns associated with hemorrhage and resuscitation: integrated approach to data analysis.

BACKGROUND: Although transcriptional profiling is a well-established technique, its application to systematic studying of various biological phenomena is still limited because of problems with high-volume data analysis and interpretation. This research project's objective was to create a comprehensive summary of changes in gene expression after hemorrhagic shock (HS), reliant and impartial of multiple variables, such as resuscitation treatments, organ analyzed, and time after impact. METHODS: Rat model of severe (40% total blood loss) HS was employed. Hemorrhagic shock was treated with 6 different resuscitation strategies: (1) racemic lactated Ringer's (DL-LR); (2) L-lactated Ringer's (L-LR); (3) ketone Ringer's (KR); (4) pyruvate Ringer's (PR); (5) 6% hetastarch (Hex); (6) 7.5% hypertonic saline (HTS). Nonresuscitated and nonhemorrhaged rats served as controls. Ketone and pyruvate Ringer solutions were identical to the lactated Ringer's solution except for equimolar substitution of lactate with beta-hydroxybutyrate and sodium pyruvate, respectively. Total RNA from liver, lung, and spleen was isolated immediately (0 hour) and 24 hour postresuscitation. Each organ, time point and treatment was profiled using individual cDNA array (1,200 genes), to produce 183 separate data files. Methods of analysis included one-way and unbalanced factorial ANOVA, Sokal-Michener average linkage clustering and contextual mapping. RESULTS: : Unresuscitated HS produced the highest number (56) of upregulated expressions in spleen and lungs. HEX and HTS affected mostly pulmonary genes (22 and 9). Fourteen genes changed in response to combination of all three factors: treatment, organ, and time. Eighteen genes were identified as treatment-specific. Fifteen genes adjusted expression 24 hour post-treatment. The largest number of genes with altered expression (168) responded differently in all three organs. In this study 15 gene clusters were pinpointed. Contextual mapping identified novel and confirmed known pathways contributing to hemorrhage/resuscitation. CONCLUSIONS: We have reliably identified genes and pathways that are affected by HS and are responsive to resuscitation. Gene expression in various organs is affected differentially by HS, which can be further modulated by the choice of resuscitation strategy.

Analysis of Variance↗

Promoter analysis and transcription profiling: Integration of genetic data enhances understanding of gene expression.

It is increasingly evident that transcription control might be conserved among organisms. For this reason, genome sequencing and gene expression profiling methods, which have yielded a plethora of data in different organisms, may be applied in species where genomic sequence is limited to mostly expression array and EST data. The identification of transcription factors and promoters associated with gene expression profiles and ESTs could therefore contribute to elucidate and predict complex regulatory events in plants.

Journal Article↗

Large-scale comparative genomics meta-analysis of Campylobacter jejuni isolates reveals low level of genome plasticity.

We have used comparative genomic hybridization (CGH) on a full-genome Campylobacter jejuni microarray to examine genome-wide gene conservation patterns among 51 strains isolated from food and clinical sources. These data have been integrated with data from three previous C. jejuni CGH studies to perform a meta-analysis that included 97 strains from the four separate data sets. Although many genes were found to be divergent across multiple strains (n = 350), many genes (n = 249) were uniquely variable in single strains. Thus, the strains in each data set comprise strains with a unique genetic diversity not found in the strains in the other data sets. Despite the large increase in the collective number of variable C. jejuni genes (n = 599) found in the meta-analysis data set, nearly half of these (n = 276) mapped to previously defined variable loci, and it therefore appears that large regions of the C. jejuni genome are genetically stable. A detailed analysis of the microarray data revealed that divergent genes could be differentiated on the basis of the amplitudes of their differential microarray signals. Of 599 variable genes, 122 could be classified as highly divergent on the basis of CGH data. Nearly all highly divergent genes (117 of 122) had divergent neighbors and showed high levels of intraspecies variability. The approach outlined here has enabled us to distinguish global trends of gene conservation in C. jejuni and has enabled us to define this group of genes as a robust set of variable markers that can become the cornerstone of a new generation of genotyping methods that use genome-wide C. jejuni gene variability data.

Animals↗

Integration of clinical data, pathology, and cDNA microarrays in influenza virus-infected pigtailed macaques (Macaca nemestrina).

For most severe viral pandemics such as influenza and AIDS, the exact contribution of individual viral genes to pathogenicity is still largely unknown. A necessary step toward that understanding is a systematic comparison of different influenza virus strains at the level of transcriptional regulation in the host as a whole and interpretation of these complex genetic changes in the context of multifactorial clinical outcomes and pathology. We conducted a study by infecting pigtailed macaques (Macaca nemestrina) with a genetically reconstructed strain of human influenza H1N1 A/Texas/36/91 virus and hypothesized not only that these animals would respond to the virus similarly to humans, but that gene expression patterns in the lungs and tracheobronchial lymph nodes would fit into a coherent and complete picture of the host-virus interactions during infection. The disease observed in infected macaques simulated uncomplicated influenza in humans. Clinical signs and an antibody response appeared with induction of interferon and B-cell activation pathways, respectively. Transcriptional activation of inflammatory cells and apoptotic pathways coincided with gross and histopathological signs of inflammation, with tissue damage and concurrent signs of repair. Additionally, cDNA microarrays offered new evidence of the importance of cytotoxic T cells and natural killer cells throughout infection. With this experiment, we confirmed the suitability of the nonhuman primate model in the quest for understanding the individual and joint contributions of viral genes to influenza virus pathogenesis by using cDNA microarray technology and a reverse genetics approach.

Animals↗

Polyphasic taxonomy, a consensus approach to bacterial systematics.

Over the last 25 years, a much broader range of taxonomic studies of bacteria has gradually replaced the former reliance upon morphological, physiological, and biochemical characterization. This polyphasic taxonomy takes into account all available phenotypic and genotypic data and integrates them in a consensus type of classification, framed in a general phylogeny derived from 16S rRNA sequence analysis. In some cases, the consensus classification is a compromise containing a minimum of contradictions. It is thought that the more parameters that will become available in the future, the more polyphasic classification will gain stability. In this review, the practice of polyphasic taxonomy is discussed for four groups of bacteria chosen for their relevance, complexity, or both: the genera Xanthomonas and Campylobacter, the lactic acid bacteria, and the family Comamonadaceae. An evaluation of our present insights, the conclusions derived from it, and the perspectives of polyphasic taxonomy are discussed, emphasizing the keystone role of the species. Taxonomists did not succeed in standardizing species delimitation by using percent DNA hybridization values. Together with the absence of another "gold standard" for species definition, this has an enormous repercussion on bacterial taxonomy. This problem is faced in polyphasic taxonomy, which does not depend on a theory, a hypothesis, or a set of rules, presenting a pragmatic approach to a consensus type of taxonomy, integrating all available data maximally. In the future, polyphasic taxonomy will have to cope with (i) enormous amounts of data, (ii) large numbers of strains, and (iii) data fusion (data aggregation), which will demand efficient and centralized data storage. In the future, taxonomic studies will require collaborative efforts by specialized laboratories even more than now is the case. Whether these future developments will guarantee a more stable consensus classification remains an open question.

Amino Acid Sequence↗

Integrating performance measure data into the Joint Commission accreditation process.

This article describes the Joint Commission's implementation plans, experience, and results to date of incorporating performance measurement data into the accreditation process. These plans have evolved in response to changes in the health care environment, feedback from accredited organizations, and both technical and political obstacles encountered. During the late 1980s, the Joint Commission developed a national performance measurement system, the IMSystem, to incorporate information about the process and outcomes of care into the accreditation process. In 1995, the ORYX initiative was introduced to offer health care organizations significant flexibility in selecting a measurement system and measures while promoting organizational self-improvement and accountability. Recently, the plans have evolved to incorporate standardized core measures that are known to be valid and reliable. These initiatives have moved the field much closer to the day when quality assessment will reflect a comprehensive view of organizational performance, based, in part, on performance measurement data.

Joint Commission on Accreditation of Healthcare Or↗

OrthoParaMap: distinguishing orthologs from paralogs by integrating comparative genome data and gene phylogenies.

BACKGROUND: In eukaryotic genomes, most genes are members of gene families. When comparing genes from two species, therefore, most genes in one species will be homologous to multiple genes in the second. This often makes it difficult to distinguish orthologs (separated through speciation) from paralogs (separated by other types of gene duplication). Combining phylogenetic relationships and genomic position in both genomes helps to distinguish between these scenarios. This kind of comparison can also help to describe how gene families have evolved within a single genome that has undergone polyploidy or other large-scale duplications, as in the case of Arabidopsis thaliana - and probably most plant genomes. RESULTS: We describe a suite of programs called OrthoParaMap (OPM) that makes genomic comparisons, identifies syntenic regions, determines whether sets of genes in a gene family are related through speciation or internal chromosomal duplications, maps this information onto phylogenetic trees, and infers internal nodes within the phylogenetic tree that may represent local - as opposed to speciation or segmental - duplication. We describe the application of the software using three examples: the melanoma-associated antigen (MAGE) gene family on the X chromosomes of mouse and human; the 20S proteasome subunit gene family in Arabidopsis, and the major latex protein gene family in Arabidopsis. CONCLUSION: OPM combines comparative genomic positional information and phylogenetic reconstructions to identify which gene duplications are likely to have arisen through internal genomic duplications (such as polyploidy), through speciation, or through local duplications (such as unequal crossing-over). The software is freely available at http://www.tc.umn.edu/~cann0010/.

Animals↗

Accelerometer use in physical activity: best practices and research recommendations.

Researchers are increasingly interested in the potential of accelerometers to improve our ability to measure and understand the health impacts of physical activity. Although accelerometers have been available commercially for more than 25 yr, broad consensus about how to use these tools has not been established. At a scientific conference in December 2004, a number of scientists were invited to present papers, serve as reactors or moderators to papers, present posters of original research, or serve as members of an audience knowledgeable about the use of accelerometers. During 2 1/2 d, information about best practices of accelerometer use was presented and suggestions for future research were made. From the collective experience of papers presented and discussions held, five areas of accelerometer use were described. This paper summarizes the best practices and future research needs from those five areas: monitor selection, quality, and dependability; monitor use protocols; monitor calibration; analysis of accelerometer data; and integration with other data sources. Suggestions for reporting standards for journal articles also are presented.

Acceleration↗

Sensation seeking and drug use by adolescents and their friends: models for marijuana and alcohol.

OBJECTIVE: To investigate the prospective influence of individual adolescents' sensation seeking tendency and the sensation seeking tendency of named peers on the use of alcohol and marijuana, controlling for a variety of interpersonal and attitudinal risk and protective factors. METHOD: Data were collected from a cohort of adolescents (N = 428; 60% female) at three points in time, starting in the eighth grade. Respondents provided information about sensation seeking, the positivity of family relations, attitudes toward alcohol and drug use, perceptions of their friends' use of alcohol and marijuana, perceptions of influence by their friends to use alcohol and marijuana, and their own use of alcohol and marijuana. In addition, they named up to three peers, whose sensation seeking and use data were integrated with respondents' data to allow for tests of hypotheses about peer clustering and substance use. RESULTS: Structural equation modeling analyses revealed direct effects of peers' sensation seeking on adolescents' own use of both marijuana and alcohol 2 years later. An unexpected finding was that the individual's own sensation seeking had indirect (not direct) effects on drug use 2 years later. CONCLUSIONS: These findings indicate the potential importance of sensation seeking as a characteristic on which adolescent peers cluster. Furthermore, the findings indicate that, beyond the influence of a variety of other risk factors, peer sensation seeking contributes to adolescents' substance use.

Adolescent↗