Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Development and application of a salmonid EST database and cDNA microarray: data mining and interspecific hybridization characteristics.

We report 80,388 ESTs from 23 Atlantic salmon (Salmo salar) cDNA libraries (61,819 ESTs), 6 rainbow trout (Oncorhynchus mykiss) cDNA libraries (14,544 ESTs), 2 chinook salmon (Oncorhynchus tshawytscha) cDNA libraries (1317 ESTs), 2 sockeye salmon (Oncorhynchus nerka) cDNA libraries (1243 ESTs), and 2 lake whitefish (Coregonus clupeaformis) cDNA libraries (1465 ESTs). The majority of these are 3' sequences, allowing discrimination between paralogs arising from a recent genome duplication in the salmonid lineage. Sequence assembly reveals 28,710 different S. salar, 8981 O. mykiss, 1085 O. tshawytscha, 520 O. nerka, and 1176 C. clupeaformis putative transcripts. We annotate the submitted portion of our EST database by molecular function. Higher- and lower-molecular-weight fractions of libraries are shown to contain distinct gene sets, and higher rates of gene discovery are associated with higher-molecular weight libraries. Pyloric caecum library group annotations indicate this organ may function in redox control and as a barrier against systemic uptake of xenobiotics. A microarray is described, containing 7356 salmonid elements representing 3557 different cDNAs. Analyses of cross-species hybridizations to this cDNA microarray indicate that this resource may be used for studies involving all salmonids.

Animals↗

Data mining and characterization of a novel pediocin-like bacteriocin system from the genome of Pediococcus pentosaceus ATCC 25745.

The genome of Pediococcus pentosaceus ATCC 25745 contains a gene cluster that resembles a regulated bacteriocin system. The gene cluster has an operon-like structure consisting of a putative pediocin-like bacteriocin gene (termed penA) and a potential immunity gene (termed peiA). Genetic determinants involved in bacteriocin transport and regulation are also found in proximity to penA and peiA but the so-called accessory gene involved in transport and the inducer gene involved in regulation are missing. Consequently, this bacterium is a poor bacteriocin producer. To analyse the potency of the putative bacteriocin operon, the two genes penA-peiA were heterologously expressed in a Lactobacillus sakei host that contains the complete apparatus for gene activation, maturation and externalization of bacteriocins. It was demonstrated that the heterologous host expressing penA and peiA produced a strong bacteriocin activity; in addition, the host became immune to its own bacteriocin, identifying the gene pair penA-peiA as a potent bacteriocin system. The novel pediocin-like bacteriocin, termed penocin A, has an isotopic mass [M+H]+ of 4684.6 Da as determined by mass spectrometry; this value corresponds well to the expected size of the mature 42 aa peptide containing a disulfide bridge. The bacteriocin is heat-stable but protease-sensitive and has a calculated pI of 9.45. Penocin A has a relatively broad inhibition spectrum, including pathogenic Listeria and Clostridium species. Immediately upstream of the regulatory genes reside some features that resemble remnants of a disrupted inducer gene. This degenerate gene was restored and shown to encode a double-glycine leader-containing peptide. Furthermore, expression of the restored gene triggered high bacteriocin production in P. pentosaceus ATCC 25745, thus confirming its role as an inducer in the pen regulon.

Amino Acid Sequence↗

Data mining: Efficiency of using sequence databases for polymorphism discovery.

An open question in research on Single Nucleotide Polymorphisms (SNPs) is, what is the percentage of true SNPs found by in silico pre-screening? To this end, we selected 13 genes, and determined the complete collection of "true" polymorphisms, or polymorphisms experimentally detected, existing in these genes in our laboratory using Denaturing High Performance Liquid Chromatography (DHPLC) and fluorescent sequencing, or in other laboratories using comparable methods. The genes studied by our group were PTGS2, IGFBP1, IGFBP3, and CYP19. GenBank sequence information was then aligned using two methods, and sequence differences termed "candidate" polymorphisms. We then compared the series of SNPs obtained experimentally and in silico and we have found that in silico methods are relatively specific (up to 55% of candidate SNPs found by SNPFinder have been discovered by experimental procedure) but have low sensitivity (not more than 27% of true SNPs are found by in silico methods).

Aromatase↗

Data mining of public SNP databases for the selection of intragenic SNPs.

Different strategies to search public single nucleotide polymorphism (SNP) databases for intragenic SNPs were evaluated. First, we assembled a strategy to annotate SNPs onto candidate genes based on a BLAST search of public SNP databases (Intragenic SNP Annotation by BLAST, ISAB). Only BLAST hits that complied with stringent criteria according to 1) percentage identity (minimum 98%), 2) BLAST hit length (the hit covers at least 98% of the length of the SNP entry in the database, or the hit is longer than 250 base pairs), and 3) location in non-repetitive DNA, were considered as valid SNPs. We assessed the intragenic context and redundancy of these SNPs, and demonstrated that the SNP content of the dbSNP and HGBASE/HGVbase databases are highly complementary but also overlap significantly. Second, we assessed the validity of intragenic SNP annotation available on the dbSNP and HGVbase websites by comparison with the results of the ISAB strategy. Only a minority of all annotated SNPs was found in common between the respective public SNP database websites and the ISAB annotation strategy. A detailed analysis was performed aiming to explain this discrepancy. As a conclusion, we recommend the application of an independent strategy (such as ISAB) to annotate intragenic SNPs, complementary to the annotation provided at the dbSNP and HGVbase websites. Such an approach might be useful in the selection process of intragenic SNPs for genotyping in genetic studies. Hum Mutat 20:162-173, 2002.

DNA↗

Mining data from potato pedigrees: tracking the origin of susceptibility and resistance to Verticillium dahliae in North American cultivars through molecular marker analysis.

Potato ( Solanum tuberosum L.) cultivated in North America is an autotetraploid species with a narrow genetic base. Most of the popular commercial cultivars are susceptible to Verticillium dahliae, a fungal pathogen causing Verticillium wilt disease, though some cultivars with relatively high resistance also exist. We have used the available pedigree information to track the origin of susceptibility and resistance to Verticillium wilt present in cultivated potatoes. One hundred thirty-nine potato cultivars and breeding selections were analyzed for resistance to the pathogen and for the presence of the microsatellite marker allele STM1051-193 that is closely linked to the resistance quantitative trait locus located on the short arm of chromosome 9. We detected an unusually high frequency of susceptible genotypes in the progeny descending from the breeding selection USDA X96-56. Molecular analysis revealed that USDA X96-56 does not have the STM1051-193 allele. Most of the first-generation progeny of this breeding selection also lack the allele. On the other hand, pedigree analysis indicated that breeding selection USDA 41956 often transfers V. dahliae resistance to its progeny. Molecular analysis detected presence of (at least) three STM1051-193 alleles in this breeding selection. These two genotypes (USDA X96-56 and USDA 41956) appear to have contributed greatly to the susceptibility or resistance, respectively, found in present commercial cultivars. Our results also indicate that the maturity class substantially affects the plant resistance response. In the intermediate to very late maturing class, the presence of the STM1051-193 allele significantly increases the resistance. Early to very early potatoes are usually more susceptible to the disease regardless of the allelic status, though the pattern of the allele effect is always the same. The results indicate that the STM1051-193 allele can be used for marker-assisted selection, but the potato maturity class also needs to be considered when making the final decision about the plant resistance level.

Crosses, Genetic↗

Molecular characterization of the developmental gene in eyes: through data-mining on integrated transcriptome databases.

OBJECTIVES: Our aim was to utilize publicly available and proprietary sources to discover candidate genes important for ocular development. DESIGN AND METHODS: The collated information on our 5092 non-redundant clusters was grouped and functional annotation was conducted using gene ontology (FatiGO) for categorizing them with respect to molecular function. The web-based viewer technological platform (H-InvDB) was employed for transcription analyses of in-house high quality fetal eye Expressed Sequence Tags (ESTs). Eye-specific ESTs were also analyzed across species by using EMBEST. RESULTS: According to adult eye cDNA libraries, nucleic acid binding and cell structure/cytoskeletal protein genes were the most abundant among the ESTs of fetal eyes. Using cDNA assembly in H-InvDB, 20 (80%) of the 25 most commonly expressed genes in the human eye are also expressed in extraocular tissues. The crystalline gamma S gene is highly expressed in the eye, but not in other tissues. We used EMBEST to compare human fetal eye and octopus eye ESTs and the expression similarity was low (1.6%). This indicated that our fetal eye library contains genes necessary for the developmental process and biological function of the eye, which may not be expressed in the fully developed octopus eyes. The human fetal eye cDNA library also contained highly abundant eye tissue genes, including alphaA-crystallin, eukaryotic translation elongation factor 1 alpha 1 (EEF1A1), bestrophin (VMD2), cystatin C, and transforming growth factor, beta-induced (BIGH3). CONCLUSIONS: Our annotated EST set provides a valuable resource for gene discovery and functional genomic analysis. This display will help to appreciate the strengths and weaknesses of the different technological platforms, so that in future studies the maximum amount of beneficial information can be derived from the appropriate use of each method.

Animals↗

Identification and classification of toe-walkers based on ankle kinematics, using a data-mining method.

A database of 1,736 patients and 2,511 gait analyses was reviewed to identify for trials where the first rocker was absent. A fuzzy c-means algorithm was used to identify sagittal ankle kinematic patterns and three groups were identified. The first showed a progressive dorsiflexion during the stance phase, while the second had a short-lived dorsiflexion, followed by a progressive plantarflexion. The third group exhibited a double bump pattern, moving successively from a short-lived dorsiflexion to a short-lived plantarflexion and then returning to a further short-lived dorsiflexion before ending with plantarflexion until toe-off. The three patterns were linked to different neurological conditions. Myopathy, neuropathy and arthogryposis essentially revealed group 1 patterns, whereas idiopathic toe-walkers mainly displayed group 2 patterns. Cerebral palsy patients, however, were relatively homogeneously distributed amongst the three groups. Able-bodied subjects walking on their toes showed a high proportion of unclassifiable ankle patterns, due to a variable gait whilst toe walking. Despite the variety of neurological conditions included in this meta-analysis repeatable biomechanical patterns appeared that could influence therapeutic management.

Adolescent↗

Oligonucleotide microarray data mining: search for age-dependent gene expression.

Information on gene expression in colon tumors versus normal human colon was recently generated by an oligonucleotide microarray study. We used the associated database to search for genes that display age-dependent variations in expression. Statistically significant evidence was obtained that such genes are present in both the tumor and normal tissue databases. Besides the analysis of all genes included in the database, three subsets of genes were analyzed separately: genes controlled by p53, and genes coding for ribosomal proteins and for nuclear-encoded mitochondrial proteins. Among the genes controlled by p53 some show an age-dependent change in expression in tumor tissues, in the sense compatible with an activation of p53 at higher age. A decreased expression of some ribosomal genes at advanced age was detected both in tumor and normal tissues. No significant age-dependent expression could be detected for genes encoding mitochondrial proteins.

Adult↗

A data mining approach for the elucidation of the action of putative etiological agents: application to the non-genotoxic carcinogenicity of genistein.

A procedure designated "the virtual similarity index" (VSI) is described to determine the probability that two or more toxicants are related mechanistically. The approach is structure-activity relationship (SAR) based and generates the virtual toxicological profiles of the chemicals under investigation. It also determines the similarities between them. That commonality is compared to the frequency with which it is found among a population of 10,000 chemicals representing the "universe of chemicals". The similarities between the candidate chemicals and chemicals known to act by other recognized mechanisms are also determined. If the similarities between the candidate chemicals are significantly greater than for the non-related ones, the chemicals are assumed to act by a common mechanism. In that context, the putative non-genotoxic mechanism responsible for the carcinogenicity of genistein (GEN) and its relationship to the action of diethylstilbestrol is examined.

Animals↗

Data mining for indicators of early mortality in a database of clinical records.

This paper describes the analysis of a database of diabetic patients' clinical records and death certificates. The objective of the study was to find rules that describe associations between observations made of patients at their first visit to the hospital and early mortality.Pre-processing was carried out and a knowledge discovery in databases (KDD) package, developed by the Lanner Group and the University of East Anglia, was used for rule induction using simulated annealing.The most significant discovered rules describe an association that was not generally known or accepted by the medical community, however, recent independent studies confirm their validity.

Aged↗