Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Piecewise exponential survival curves with smooth transitions.

Several models of a population survival curve composed of two piecewise exponential distributions are developed. In one formulation the hazard rate changes at a point that is an unobservable random variable that varies between individuals. The population hazard function may decrease with age even when all individuals' hazards are increasing. In a second formulation, the population hazard function is modeled directly. Several models are fit to the survival history of a cohort of 5751 highly inbred male Drosophila melanogaster and the British coal mining disaster data.

Animals↗

Traveling the tau pathway: a personal account.

Studies of the tau protein and its pathological fate as a neurofibrillary tangle have been a pillar of Alzheimer's disease research. The understanding of the fundamental position that tau occupies in the disease cascade is a tribute to an international group of scientists who brought rigor, candid assessments of data, and critical thinking to the problem. The tau pathway winds its way from astute clinical observations to pathological correlations, from molecular and cellular experiments to mining informatic data, and from animal behavior to the biophysics of protein structure. For most the vindication of this tireless effort will come from tau-based therapies; but for others the remarkable biology revealed by the Alzheimer disease process has been its own reward.

Alzheimer Disease↗

Expression profiling by microarrays in colorectal cancer (Review).

Genome-wide gene profiling studies using microarrays have the potential to improve diagnosis and treatment of human cancers. Microarrays have identified many genes that are deregulated in colorectal cancer compared to normal tissue. Groups of genes that are predictive of tumor stage or presence of metastases, hence putatively associated with cancer progression have also been revealed. Microarray studies have identified genes whose expression are impacted by chemotherapies for colorectal cancer, thus could potentially be used to predict response to treatments. Unique gene expression profiles have also been used to classify metastases of uncertain origin. The wide application of microarrays generates exciting prospects in translational research. However, to date overlaps of candidate gene lists associated with specific clinical/biological phenotypes remain disturbingly poor between studies. Overfitting, bias, reporting of only the best results, and fidelity of probe annotations could present limitations for the interpretation of results shown in microarray publications. Making raw data from these microarray experiments publicly available for analysis by other investigators using different analytical algorithms or for in silico studies may facilitate the most thorough mining of data from these expensive studies. Validations of the results using other more precise techniques and at the biological level represent critical follow-up goals for microarray studies.

Adenoma↗

Identifying potential tumor markers and antigens by database mining and rapid expression screening.

Genes expressed specifically in malignant tissue may have potential as therapeutic targets but have been difficult to locate for most cancers. The information hidden within certain public databases can reveal RNA transcripts specifically expressed in transformed tissue. To be useful, database information must be verified and a more complete pattern of tissue expression must be demonstrated. We tested database mining plus rapid screening by fluorescent-PCR expression comparison (F-PEC) as an approach to locate candidate brain tumor antigens. Cancer Genome Anatomy Project (CGAP) data was mined for genes highly expressed in glioblastoma multiforme. From 13 mined genes, seven showed potential as possible tumor markers or antigens as determined by further expression profiling. Now that large-scale expression information is readily available for many of the commonly occurring cancers, other candidate tumor markers or antigens could be located and evaluated with this approach.

Algorithms↗

Lung carcinoma by histologic type in coal miners.

Histologic types of lung carcinoma were studied in 171 coal miners in the National Coal Workers' Autopsy Study. These miners had an average underground mining tenure of 29 +/- 14 years and an average smoking history of 31 +/- 23 pack-years. The proportion of carcinomas by cell type were: squamous cell carcinoma, 30%; adenocarcinoma, 27%; small-cell undifferentiated carcinoma, 26%; large-cell undifferentiated carcinoma, 9%; and other carcinomas, 8%. More tumors were observed in the right lung and in the upper lobes of both lungs than in the left lung and in the lower lobes of both lungs, respectively. The majority of the tumors were centered on cartilaginous airways (81%) as compared with the peripheral regions of the lung (19%). Squamous cell carcinomas predominated in the older miners and in larger airways. Adenocarcinomas were more common in the peripheral lung. No significant interaction was demonstrated between cell type and years of underground mining. The data indicate that lung carcinoma in coal miners differs little in its pathologic features from men in the general population who smoke cigarettes. No effect of coal mine dust exposure on lung carcinoma histogenesis was demonstrated.

Adenocarcinoma↗

Antimony accumulation in Achillea ageratum, Plantago lanceolata and Silene vulgaris growing in an old Sb-mining area.

Preliminary data of a biogeochemical survey concerning antimony transfer from soil to plants in an abandoned Sb-mining area are presented. Achillea ageratum, Plantago lanceolata and Silene vulgaris can strongly accumulate antimony when its extractable fraction in the soil is high (139-793 mg/kg). A. ageratum accumulates in basal leaves (1367 mg/kg) and inflorescences (1105 mg/kg), P. lanceolata in roots (1150 mg/kg) and S. vulgaris in shoots (1164 mg/kg). In these plant species, the efficiency of antimony accumulation decreases when the antimony availability in the soil is high. In A. ageratum and S. vulgaris, the death of the epigeal target part at the end of the growing season contributes to a reduction of the antimony load in the plant. A study to test the use of these species as bioindicators of antimony availability in soil is suggested by our results.

Journal Article↗

Mining statistically significant associations for exploratory analysis of human sleep data.

We introduce a specialized association rule mining technique that can extract patterns from complex sleep data comprising polysomnographic recordings, clinical summaries, and sleep questionnaire responses. The rules mined can describe associations among temporally annotated events and questionnaire or summary data; e.g., the likelihood that an occurrence of a rapid eye movement (REM) sleep stage during the second 100 sleep epochs of the night is associated with moderate caffeine intake. We use chi2 analysis to ensure statistical significance of the mined rules at the level P < 0.05. Our results, obtained by mining sleep-related data from 242 human subjects, reveal clinically interesting associations among the polysomnographic and summary variables. Our experience suggests that association mining may also be useful for selection of variables prior to using logistic regression.

Algorithms↗

A prospective study evaluating early rehabilitation in preventing back pain chronicity in mine workers.

STUDY DESIGN: This was a prospective study. OBJECTIVES: To evaluate the results of education and early rehabilitation in the prevention of back pain chronicity in coal mine workers. SUMMARY OF BACKGROUND DATA: A new mine was established in central Queensland, Australia. Preventing chronicity is important in the treatment of back pain in the industrial setting, because back pain is often refractory to treatment. Back pain patients also constitute the majority of compensation claims. METHODS: A back pain program was instituted that comprised work force education, early injury reporting, first aid at the mine, and changing workplace psychosocial perceptions. Management employees were actively involved. The time off work, number of claims per hundred workers, and costs per claim were compared with another mine in the area. RESULTS: The median time to return to work was 10 days. In the study group the number of claims and costs per claim were significantly less (P < 0.01) as compared to the control group. CONCLUSIONS: This was an easy to institute, inexpensive back pain program, which succeeded in preventing back pain chronicity in the studied group of mine workers, with no worker being off work for more than 60 days.

Back Pain↗

An atlas of forecasted molecular data. 2. Vibration frequencies of main-group and transition-metal neutral gas-phase diatomic molecules in the ground state.

This atlas of diatomic-molecular vibration frequencies parallels the previously offered Atlas of Internuclear Separations. The Atlas was produced by mining the data from Huber and Herzberg and training neural network software to forecast new data. New protocols were employed with the powerful software, which was originally designed for forecasting the financial markets. The Atlas presents 1920 additional vibration frequencies for use until critical tables are available to fill the needs more precisely. The precision of the predictions is characterized by the average fractional 1% confidence limit, that is, 10.66%. The accuracies of the predictions are determined in two ways. First, 221 of the 224 Huber and Herzberg data values used for training and validation fall within the prediction confidence limits or fall outside by less than 10% of the Huber and Herzberg values, and 181 values agree (within the limits). Second, 87 of 101 comparison data values, consisting of literature data and some additional Huber and Herzberg values, fall within the prediction confidence limits or fall outside by less than half the prediction values, and 44 of the 101 values agree (within the limits).

Journal Article↗

Genomics. University company to exploit heart data.

This month, Boston University, which directs the Framingham Heart Study, a massive government effort begun in 1948 to monitor the cardiovascular health of more than 10,000 residents of this suburb of Boston, announced plans to form a bioinformatics company that will mine the data. The university will own 20% of Framingham Genomic Medicine Inc., which hopes to raise $21 million to begin modernizing the immense database and packaging it in a format that will be valuable to the pharmaceutical industry. The plan raises a host of difficult ethical issues, including patient privacy, academic conflicts of interest, and reciprocal value to the Framingham residents whose medical data will form the basis for the new enterprise.

Bioethics↗

Global analysis of cellular transcription following infection with an HIV-based vector.

We have examined the changes in cellular transcription resulting from infection with HIV-based vectors. Previous work suggested that the incoming viral genome may under some circumstances be detected as DNA damage, so to explore this possibility, we compared the transcriptional response to infection with an HIV-based vector to the response to treatment with the DNA-damaging agent etoposide. Expression levels of about 12,000 cellular RNA transcripts were determined in a human B-cell line at different times after either treatment. Statistical analysis revealed that the infection with the lentivirus vector resulted in quite modest changes in gene expression. Treatment with etoposide, in contrast, caused drastic changes in expression of genes known or inferred to be involved in apoptosis. Statistically significant though subtle parallels in the cellular transcriptional responses to etoposide treatment and HIV-vector infection could be detected. Several further data sets analyzing infections with HIV-based vectors or wild-type HIV-1 showed similar modest effects on cellular transcription and very modest parallels among different data sets. These findings establish that HIV-vector or HIV-1 infection has remarkably little effect on cellular transcription. The statistical methods described here may be of wide use in mining microarray data sets. Our observations support the idea that gene therapy with HIV-based vectors should not be particularly toxic to cells due to disruption of cellular transcription.

B-Lymphocytes↗

[Evaluation of effects of dust prevention in the principal tungsten mines in Jiangxi].

An evaluation of the effects of dust prevention in the Jiangxi tungsten mines has been carried out. The rate of silicosis morbidity in most mines was under 1%. Up to 1983, the rate in individual mines is 1.95%. According to the data from those mines, the forecasting of cumulative probability of morbidity of mine workers having been in contact with dust for 30 years is up to 7.5%. From those data, the authors suggest that the maximum permissible concentration of dust should be 1.0 mg/m3 in the mines with concentration of silicon dioxide dust over 70%.

Dust↗

MADCAP: isolation of novel nAb-na&#xef;ve AAV capsids from metagenomic data.

UNLABELLED: Gene therapy using adeno-associated virus (AAV) vectors offers promising treatment for genetic disorders, but significant limitations restrict clinical application. Current AAV serotypes exhibit strong liver tropism and require high doses for extra-hepatic targeting, and pre-existing antibodies (NAbs) exclude up to 50% of potential patients. Evolutionarily distant isolates can evade neutralization but typically transduce human tissues poorly and require extensive engineering. We developed MADCAP (Metagenomic AAV Discovery and Capsid Annotation Pipeline) to systematically mine metagenomic data for functional, clinically relevant AAV capsids. We hypothesized that these sources might contain capsids that do not circulate widely in humans, can transduce human cells, and avoid neutralization. We screened 4.2 million metagenomic samples and identified 139 novel AAV capsid isolates which were tested for viral capsid assembly, viability, neutralization evasion, and tissue transduction in non-human primates. While natural serotypes (AAV1, AAV2, AAV9) were neutralized at low dilutions of pooled human immunoglobulin (IVIG), 68% of tested MADCAP capsids exhibited minimal to undetectable neutralization even at supra-physiological IVIG concentrations. Systemically delivered MADCAP capsids effectively transduced multiple clinically relevant tissues in non-human primates. Two capsids, MC46 and MC55, demonstrated improved CNS tropism compared to AAV9 while maintaining comparable production yields. In passive transfer studies, MC46 retained full transduction efficiency in the presence of human antibodies, while AAV9 transduction was completely lost. This work establishes metagenomic mining as a powerful tool for accelerating AAV capsid discovery, identifying isolates with favorable tissue tropisms and resistance to broadly neutralizing antibodies. IMPORTANCE: This work provides proof of concept that potentially clinically relevant AAVs can be isolated from metagenomic data. Our findings lay the groundwork for accelerated discovery of AAV capsids which could potentially increase the accessibility and effectiveness of AAV gene therapy.

AAV↗

Integration and mining of malaria molecular, functional and pharmacological data: how far are we from a chemogenomic knowledge space?

The organization and mining of malaria genomic and post-genomic data is important to significantly increase the knowledge of the biology of its causative agents, and is motivated, on a longer term, by the necessity to predict and characterize new biological targets and new drugs. Biological targets are sought in a biological space designed from the genomic data from Plasmodium falciparum, but using also the millions of genomic data from other species. Drug candidates are sought in a chemical space containing the millions of small molecules stored in public and private chemolibraries. Data management should, therefore, be as reliable and versatile as possible. In this context, five aspects of the organization and mining of malaria genomic and post-genomic data were examined: 1) the comparison of protein sequences including compositionally atypical malaria sequences, 2) the high throughput reconstruction of molecular phylogenies, 3) the representation of biological processes, particularly metabolic pathways, 4) the versatile methods to integrate genomic data, biological representations and functional profiling obtained from X-omic experiments after drug treatments and 5) the determination and prediction of protein structures and their molecular docking with drug candidate structures. Recent progress towards a grid-enabled chemogenomic knowledge space is discussed.

Animals↗

PROPHECY--a yeast phenome database, update 2006.

Connecting genotype to phenotype is fundamental in biomedical research and in our understanding of disease. Phenomics--the large-scale quantitative phenotypic analysis of genotypes on a genome-wide scale--connects automated data generation with the development of novel tools for phenotype data integration, mining and visualization. Our yeast phenomics database PROPHECY is available at http://prophecy.lundberg.gu.se. Via phenotyping of 984 heterozygous diploids for all essential genes the genotypes analysed and presented in PROPHECY have been extended and now include all genes in the yeast genome. Further, phenotypic data from gene overexpression of 574 membrane spanning proteins has recently been included. To facilitate the interpretation of quantitative phenotypic data we have developed a new phenotype display option, the Comparative Growth Curve Display, where growth curve differences for a large number of mutants compared with the wild type are easily revealed. In addition, PROPHECY now offers a more informative and intuitive first-sight display of its phenotypic data via its new summary page. We have also extended the arsenal of data analysis tools to include dynamic visualization of phenotypes along individual chromosomes. PROPHECY is an initiative to enhance the growing field of phenome bioinformatics.

Chromosomes, Fungal↗

Growth of children in two economically diverse Peruvian high-altitude communities.

The growth of children living in two high-altitude communities associated with an active copper mine in southern Peru was examined. In the community directly associated with mining operations, nutritional and health conditions were believed to be relatively favorable as a result of the substantial mine-related infrastructure that had developed over the previous 12 years. In contrast, few such benefits were available in the other community, which provides limited part-time labor at the mine. Anthropometric data, including measurements of height, weight, skinfold thicknesses, upper arm circumference, and chest dimensions, and determination of bone age, were collected from a total of 880 children between the ages of 4 and 18 years. There were significant differences between the two communities, with those in the mining community exhibiting significantly greater height and weight, a higher level of body fat, and more rapid skeletal development. Among children over the age of 12 years, a plateau in height was seen, suggesting that the benefits to growth resulting from mining-related development were more noticeable in younger children. Compared with Peruvian high-altitude populations examined during the 1960s, both samples from the present study were substantially taller and heavier, suggesting that despite local differences in socioeconomic conditions between the communities studied, overall conditions for growth are generally more favorable than those that existed among Peruvian high-altitude populations surveyed in the 1960s.

Adolescent↗

Prevalence of pneumonoconiosis among coal and heavy metal miners in Zimbabwe.

No prevalence data on pneumonoconiosis among Zimbabwe's 30,000 miners have been available. Passage of a 1984 law requiring examination of all miners has provided a data base to assess this, but the records had not been previously evaluated or stored in a manner to facilitate this. In 1988 we developed a strategy to utilize the existing records to estimate cross-sectional rates rapidly. In this report, we describe the approach and demonstrate high rates of simple pneumonoconiosis among long-term workers in the coal, nickel, copper, and gold mines. These data, though limited, provide a rationale for more detailed investigations in these workforces and an impetus to establish an ongoing surveillance plan for the nation's miners.

Analysis of Variance↗

Bioinformatic analysis of primary endothelial cell gene array data illustrated by the analysis of transcriptome changes in endothelial cells exposed to VEGF-A and PlGF.

We recently published a review in this journal describing the design, hybridisation and basic data processing required to use gene arrays to investigate vascular biology (Evans et al. Angiogenesis 2003; 6: 93-104). Here, we build on this review by describing a set of powerful and robust methods for the analysis and interpretation of gene array data derived from primary vascular cell cultures. First, we describe the evaluation of transcriptome heterogeneity between primary cultures derived from different individuals, and estimation of the false discovery rate introduced by this heterogeneity and by experimental noise. Then, we discuss the appropriate use of Bayesian t-tests, clustering and independent component analysis to mine the data. We illustrate these principles by analysis of a previously unpublished set of gene array data in which human umbilical vein endothelial cells (HUVEC) cultured in either rich or low-serum media were exposed to vascular endothelial growth factor (VEGF)-A165 or placental growth factor (PlGF)-1(131). We have used Affymetrix U95A gene arrays to map the effects of these factors on the HUVEC transcriptome. These experiments followed a paired design and were biologically replicated three times. In addition, one experiment was repeated using serial analysis of gene expression (SAGE). In contrast to some previous studies, we found that VEGF-A and PlGF consistently regulated only small, non-overlapping and culture media-dependant sets of HUVEC transcripts, despite causing significant cell biological changes.

Cells, Cultured↗