Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Characterization and quality control of recombinant adenovirus vectors for gene therapy.

Highly purified recombinant adenovirus undergoes routine quality controls for identity, potency and purity prior to its use as a gene therapy vector. Quantitative characterization of infectivity is measurable by the expression of the DNA binding protein, an early adenoviral protein, in an immunofluorescence bioassay on permissive cells as a potency determinant. The specific particle count, a key quality indicator, is the total number of intact particles present compared to the number of infectious units. Electron microscopic analysis using negative staining gives a qualitative biophysical analysis of the particles eluted from anion-exchange HPLC. One purity assessment is accomplished via the documented presence and relative ratios of component adenoviral proteins as well as potential contaminants by reversed-phase HPLC of the intact virus followed by protein peak identification using MALDI-TOF mass spectrometry and subsequent data mining. Verification of the viral genome is performed and expression of the transgene is evaluated in in vitro systems for identity. Production lots are also evaluated for replication-competent adenovirus prior to human use. For adenovirus carrying the human IL-2 transgene, quantitative IL-2 expression is demonstrated by ELISA and cytokine potency by cytotoxic T lymphocyte assay following infection of permissive cells. Both quantitative and qualitative analyses show good batch to batch reproducibility under routine test conditions using validated methods.

Adenoviridae↗

Identification of novel tropomyosin 1 genes of pufferfish (Fugu rubripes) on genomic sequences and tissue distribution of their transcripts.

Fugu genome database enabled us to identify two novel tropomyosin 1 (TPM1) genes through in silico data mining and isolation of their corresponding cDNAs in vivo. The duplicate TPM1 genes in Japanese pufferfish Fugu rubripes suggest that additional an ancient segmental duplication or whole genome duplication occurred in fish lineage, which, like many other reported Fugu genes, showed reduction in genomic size in comparison with their human homologue. Computer analysis predicted that the coiled-coil probabilities, that were thought to be the most major function of TPM, were the same between the two TPM1 isoforms. We confirmed that the tissue expression profiles of the two TPM1 genes differed from each other, which implied that changes in expression pattern could fix duplicated TPM1 genes although the two TPM1 isoforms appear to have similar function.

Animals↗

Human genomic deletions mediated by recombination between Alu elements.

Recombination between Alu elements results in genomic deletions associated with many human genetic disorders. Here, we compare the reference human and chimpanzee genomes to determine the magnitude of this recombination process in the human lineage since the human-chimpanzee divergence approximately 6 million years ago. Combining computational data mining and wet-bench experimental verification, we identified 492 human-specific deletions (for a total of approximately 400 kb) attributable to this process, a significant component of the insertion/deletion spectrum of the human genome. The majority of the deletions (295 of 492) coincide with known or predicted genes (including 3 that deleted functional exons, as compared with orthologous chimpanzee genes), which implicates this process in creating a substantial portion of the genomic differences between humans and chimpanzees. Overall, we found that Alu recombination-mediated genomic deletion has had a much higher impact than was inferred from previously identified isolated events and that it continues to contribute to the dynamic nature of the human genome.

Animals↗

[Molecular cloning and character analysis of the mouse zinc finger protein gene Zfp474 exclusively expressed in testis and ovary].

A novel mouse zinc finger protein gene that contains C2HC/C3H domain was first isolated by using a data-mining tool called Digital Differential Display (DDD) from the National Center for Biotechnology Information. The full-length cDNA of this transcript was deduced and further confirmed by reverse transcriptase polymerase chain reaction (RT-PCR). Total four exons of the mouse gene spaning a 29 869 bp genomic DNA sequence was mapped to chromosome 18D1. The cDNA encodes a novel protein of 347 amino acids and the protein contains four C2HC/C3H domains. Northern blot analyses revealed that Zfp474 mRNA was exclusively expressed in testis and ovary and had one transcript. We hypothesize that Zfp474 functions as a germ cell specific transcription factor that plays important roles in spermatid differentiation and oocyte development.

Amino Acid Sequence↗

Homologs of the alpha- and beta-subunits of mammalian brain platelet-activating factor acetylhydrolase Ib in the Drosophila melanogaster genome.

The mammalian intracellular brain platelet-activating factor acetylhydrolase, implicated in the development of cerebral cortex, is a member of the phospholipase A2 superfamily. It is made up of a homodimer of the 45 kDa LIS1 protein (a product of the causative gene for type I lissencephaly) and a pair of homologous 26-kDa alpha-subunits which account for all the catalytic activity. LIS1 is hypothesized to regulate nuclear movement in migrating neurons through interactions with the cytoskeleton, while the alpha-subunits, whose structure is known, contain a trypsin-like triad within the framework of a unique tertiary fold. The physiological significance of the association of the two types of subunits is not known. In an effort to better understand the function of the complex we turned to genomic data mining in search of related proteins in lower eukaryotes. We found that the Drosophila melanogaster genome contains homologs of both alpha- and beta-subunits, and we cloned both genes. The alpha-subunit homolog has been overexpressed, purified and crystallized. It lacks two of the three active-site residues and, consequently, is catalytically inactive against PAF-AH (Ib) substrates. Our study shows that the beta-subunit homolog is highly conserved from Drosophila to mammals and is able to interact with the mammalian alpha-subunits but is unable to interact with the Drosophila alpha-subunit. Proteins 2000;39:1-8.

1-Alkyl-2-acetylglycerophosphocholine Esterase↗

Understanding spatial concentrations of road accidents using frequent item sets.

This paper aims at understanding why road accidents tend to cluster in specific road segments. More particularly, it aims at analyzing which are the characteristics of the accidents occurring in "black" zones compared to those scattered all over the road. A technique of frequent item sets (data mining) is applied for automatically identifying accident circumstances that frequently occur together, for accidents located in and outside "black" zones. A Belgian periurban region is used as case study. Results show that accidents occurring in "black" zones are characterized by left-turns at signalized intersections, collisions with pedestrians, loss control of the vehicle (run-off-roadway) and rainy weather conditions. Accidents occurring outside "black" zones (scattered in space) are characterized by left turns on intersections with traffic signs, head-on collisions and drunken road user(s). Furthermore, parallel collisions and accidents on highways or roads with separated lanes, occurring at night or during the weekend are frequently occurring accident patterns for all accident locations. These exploratory results show the potentiality of the frequent item set method in addition to more classical statistical techniques, but also suggest that there is no unique countermeasure for reducing the number of accidents.

Accidents, Traffic↗

A family of 12 human genes containing oxysterol-binding domains.

Oxysterol-binding proteins (OSBPs) have been described in a wide range of eukaryotes, and are often found to be part of a multi-gene family. We have used bioinformatics and data mining as a starting point for identifying new family members in humans based on the presence of the OSBP signature EQVSHHPP. In addition to OSBP and the recently reported OSBP2, we have found 10 other genes encoding oxysterol-binding domains. Here, we report cDNA and deduced peptide sequences of the previously unknown OSBPs and compare the peptides and genes. All of the genes encode a pleckstrin homology domain, except OSBPL2. However, two of the peptides, OSBPL2 and OSBPL1A, consist of the OSBP domain only. A second OSBPL1 transcript (OSBPL1B) contains 15 additional upstream exons, with a deduced peptide containing a pleckstrin homology domain. Cladistic analysis divides the human OSBP genes into five groups, whose members share similarities in sequence and gene structure; RT-PCR analysis indicates that expression patterns among group members vary widely.

Amino Acid Sequence↗

Evolution of matrix and bone gamma-carboxyglutamic acid proteins in vertebrates.

The evolution of calcified tissues is a defining feature in vertebrate evolution. Investigating the evolution of proteins involved in tissue calcification should help elucidate how calcified tissues have evolved. The purpose of this study was to collect and compare sequences of matrix and bone gamma-carboxyglutamic acid proteins (MGP and BGP, respectively) to identify common features and determine the evolutionary relationship between MGP and BGP. Thirteen cDNAs and genes were cloned using standard methods or reconstructed through the use of comparative genomics and data mining. These sequences were compared with available annotated sequences (a total of 48 complete or nearly complete sequences, 28 BGPs and 20 MGPs) have been identified across 32 different species (representing most classes of vertebrates), and evolutionarily conserved features in both MGP and BGP were analyzed using bioinformatic tools and the Tree-Puzzle software. We propose that: 1) MGP and BGP genes originated from two genome duplications that occurred around 500 and 400 million years ago before jawless and jawed fish evolved, respectively; 2) MGP appeared first concomitantly with the emergence of cartilaginous structures, and BGP appeared thereafter along with bony structures; and 3) BGP derives from MGP. We also propose a highly specific pattern definition for the Gla domain of BGP and MGP.

Amino Acid Sequence↗

Sulf-2, a proangiogenic heparan sulfate endosulfatase, is upregulated in breast cancer.

Sulf-2 is an endosulfatase with activity against glucosamine-6-sulfate modifications within subregions of intact heparin. The enzyme has the potential to modify the sulfation status of extracellular heparan sulfate proteoglycan (HSPG) glycosaminoglycan chains and thereby to regulate interactions with HSPG-binding proteins. In the present investigation, data mining from published studies was employed to establish Sulf-2 mRNA upregulation in human breast cancer. We further found that cultured breast carcinoma cells expressed Sulf-2 mRNA and released enzymatically active proteins into conditioned medium. In two mouse models of mammary carcinoma, Sulf-2 mRNA was upregulated in comparison to its expression in normal mammary gland. Although mRNA was present in normal tissues, Sulf-2 protein was undetectable; it was, however, detected in some premalignant lesions and in tumors. The protein was localized to the epithelial cells of the tumors. In support of the possible mechanistic relevance of Sulf-2 upregulation in tumors, purified recombinant Sulf-2 promoted angiogenesis in the chick chorioallantoic membrane assay.

Allantois↗

Serpins in the Caenorhabditis elegans genome.

Data mining in genome sequences can identify distant homologues of known protein families, and is most powerful if solved structures are available to reveal the three-dimensional implications of very dissimilar sequences. Here we describe putative serpin sequences identified with very high statistical significance in the Caenorhabditis elegans genome. When mapped onto vertebrate serpins such as alpha1-antitrypsin, they suggest novel structural features. Some appear complete, some show extensive deletions, and others appear to contain only the C-terminal part of the known serpin fold, probably in partnership with N-terminal regions that have conformations unlike those of known serpins. The observation of such striking sequence similarity, in proteins that must have significantly different overall structures, substantially extends the structural characteristics of the serpin family of proteins.

Amino Acid Sequence↗

Identification and analysis of expressed resistance gene sequences in wheat.

Forty-eight resistance (R) genes conferring resistance to various types of pests have been cloned from 12 plant species. Irrespective of the host or the pest type, most R genes share a strong protein sequence similarity especially for domains and motifs. The objective of this study was to identify expressed R genes of wheat, the fraction of which is expected to be very low in the genome. Using modified RNA fingerprinting and data mining approaches we identified 220 expressed R-gene candidates. Of these, 125 sequences structurally resembled known R genes. In addition to 25-87% protein sequence similarity with the known R genes, the sequence, order, and distribution of the domains and motifs were also the same. Among the remaining 95, 17 were probable R-related, 21 were a new class of nucleotide-binding kinases, 21 were probable kinases, and 36 were p-loop-containing unknown sequences. About 76% were rare including 73 novel sequences. Three new R-gene specific motifs were also identified. Physical mapping of the 164 best R-gene candidates on 339 deletion lines localized 121 mappable R-gene candidates to 26 small chromosomal regions encompassing about 16% of the genome. About 90 of the 110 phenotypically characterized wheat R genes corresponding to 18 different pests also mapped in these regions.

Amino Acid Sequence↗

The conserved tyrosine residues 401 and 1044 in ATP sites of human P-glycoprotein are critical for ATP binding and hydrolysis: evidence for a conserved subdomain, the A-loop in the ATP-binding cassette.

Each nucleotide-binding domain (NBD) of mammalian P-glycoproteins (Pgps) and human ATP-binding cassette (ABC) B subfamily members contains a tyrosine residue approximately 25 residues upstream of the Walker A domain. To assess the role of the conserved Y401 and Y1044 residues of human Pgp, we substituted these residues with F, W, C, or A either singly or together. The mutant proteins were expressed in a Vaccinia virus-based transient expression system as well as in baculovirus-infected HighFive insect cells. The Y401F, Y401W, Y1044F, Y1044W, or Y401F/Y1004F mutants transported fluorescent substrates similar to the wild-type protein. On the other hand, Y401L and Y401C exhibited partial (30-50%) function, and transport was completely abolished in Y401A, Y1044A, and Y401A/Y1044A mutant Pgps. Similarly, in Y401A, Y1044A, and Y401A/Y1044A mutants, TNP-ATP binding, vanadate-induced trapping of nucleotide, and ATP hydrolysis were completely abolished. Thus, an aromatic residue upstream of the Walker A motif in ABC transporters is critical for binding of ATP. Additionally, the crystal structures of several NBDs in the nucleotide-bound form, data mining, and alignment of 18,514 ABC domains with the consensus conserved sequence in a database of all nonredundant proteins indicate that an aromatic residue is highly conserved in approximately 85% of ABC proteins. Although the role of this aromatic residue has previously been studied in a few ABC proteins, we provide evidence for a near-universal structural and functional role for this residue and recognize its presence as a conserved subdomain approximately 25 amino acids upstream of the Walker A motif that is critical for ATP binding. We named this subdomain the "A-loop" (aromatic residue interacting with the adenine ring of ATP).

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Magic roundabout is a new member of the roundabout receptor family that is endothelial specific and expressed at sites of active angiogenesis.

We have used bioinformatic data mining to identify a novel, endothelial-specific gene encoding a protein with homology to the axon guidance protein roundabout (ROBO1). The new gene has been called magic roundabout (ROBO4; GenBank acc. no. AF361473) and is smaller than other members of the roundabout gene family. Thus, in the extracellular region, magic roundabout has only two of the five immunoglobulin and two of the three fibronectin domains present in other roundabout genes. Expression of magic roundabout in vitro was detected in only endothelial cells and was greater in cells exposed to hypoxia. In situ hybridization and immunohistochemistry validated the bioinformatic prediction that magic roundabout expression would be endothelial specific in vivo. Magic roundabout expression in the adult was restricted exclusively to sites of active angiogenesis, notably tumor vessels. The identification of magic roundabout shows that the roundabout gene family extends beyond neuronal tissue and that roundabout/slit interactions are likely to have a role in angiogenesis.

Amino Acid Sequence↗

Characterization of isoforms and genomic organization of mouse calumenin.

Calumenin is a multiple EF-hand protein located in endo/sarcoplasmic reticulum of mammalian heart and other tissues [J. Biol. Chem. 272 (1997) 18232; Genomics 49 (1998) 331; Biochim. Biophys. Acta 1386 (1998) 121]. In the present study, a new isoform of mouse calumenin (mouse calumenin 2) was cloned by RT-PCR and genomic DNA PCR. The deduced amino acid sequence of mouse calumenin 2 is 315 aa long with the calculated MW of 37,064 and pI of 4.26. It has 92% aa sequence identity to previously identified mouse calumenin [J. Biol. Chem. 272 (1997) 18232] (mouse calumenin 1). The difference in the aa sequence was restricted to the first two EF-hand regions (residues 74-138). Northern blot analysis shows that mouse calumenin 2 is highly expressed in heart, lung, testis and unpregnant uterus. The expression of mouse calumenin 2 appears to decrease when fetal development is progressed. Genomic DNA PCR, sequencing and data mining of mouse genome database were utilized to examine the exon-intron boundaries of mouse calumenin genes. Both mouse calumenin 1 and 2 genes encompass six exons, and five of them (Exon1, 3, 4, 5 and 6) are identical. However, mouse calumenin 1 contains Exon2-1, whereas mouse calumenin 2 contains a neighboring Exon2-2. The calumenin genes are localized on mouse chromosome 6 having conserved synteny with human chromosome 7q32. For comparison, the genomic organization of human calumenin was also examined using the published human genome database (UCSC Genome Bioinformatics at ). Like mouse calumenin genes, two human calumenin genes also consist of five identical exons (Exon1, 3, 4, 5 and 6) and a different Exon2. The present study suggests that the genomic organization of calumenin genes is well conserved between human and mouse.

Amino Acid Sequence↗

WebGestalt: an integrated system for exploring gene sets in various biological contexts.

High-throughput technologies have led to the rapid generation of large-scale datasets about genes and gene products. These technologies have also shifted our research focus from 'single genes' to 'gene sets'. We have developed a web-based integrated data mining system, WebGestalt (http://genereg.ornl.gov/webgestalt/), to help biologists in exploring large sets of genes. WebGestalt is composed of four modules: gene set management, information retrieval, organization/visualization, and statistics. The management module uploads, saves, retrieves and deletes gene sets, as well as performs Boolean operations to generate the unions, intersections or differences between different gene sets. The information retrieval module currently retrieves information for up to 20 attributes for all genes in a gene set. The organization/visualization module organizes and visualizes gene sets in various biological contexts, including Gene Ontology, tissue expression pattern, chromosome distribution, metabolic and signaling pathways, protein domain information and publications. The statistics module recommends and performs statistical tests to suggest biological areas that are important to a gene set and warrant further investigation. In order to demonstrate the use of WebGestalt, we have generated 48 gene sets with genes over-represented in various human tissue types. Exploration of all the 48 gene sets using WebGestalt is available for the public at http://genereg.ornl.gov/webgestalt/wg_enrich.php.

Computer Graphics↗

Differentiation of regions with atypical oligonucleotide composition in bacterial genomes.

BACKGROUND: Complete sequencing of bacterial genomes has become a common technique of present day microbiology. Thereafter, data mining in the complete sequence is an essential step. New in silico methods are needed that rapidly identify the major features of genome organization and facilitate the prediction of the functional class of ORFs. We tested the usefulness of local oligonucleotide usage (OU) patterns to recognize and differentiate types of atypical oligonucleotide composition in DNA sequences of bacterial genomes. RESULTS: A total of 163 bacterial genomes of eubacteria and archaea published in the NCBI database were analyzed. Local OU patterns exhibit substantial intrachromosomal variation in bacteria. Loci with alternative OU patterns were parts of horizontally acquired gene islands or ancient regions such as genes for ribosomal proteins and RNAs. OU statistical parameters, such as local pattern deviation (D), pattern skew (PS) and OU variance (OUV) enabled the detection and visualization of gene islands of different functional classes. CONCLUSION: A set of approaches has been designed for the statistical analysis of nucleotide sequences of bacterial genomes. These methods are useful for the visualization and differentiation of regions with atypical oligonucleotide composition prior to or accompanying gene annotation.

DNA, Bacterial↗

Structural details (kinks and non-alpha conformations) in transmembrane helices are intrahelically determined and can be predicted by sequence pattern descriptors.

One of the promising methods of protein structure prediction involves the use of amino acid sequence-derived patterns. Here we report on the creation of non-degenerate motif descriptors derived through data mining of training sets of residues taken from the transmembrane-spanning segments of polytopic proteins. These residues correspond to short regions in which there is a deviation from the regular alpha-helical character (i.e. pi-helices, 3(10)-helices and kinks). A 'search engine' derived from these motif descriptors correctly identifies, and discriminates amongst instances of the above 'non-canonical' helical motifs contained in the SwissProt/TrEMBL database of protein primary structures. Our results suggest that deviations from alpha-helicity are encoded locally in sequence patterns only about 7-9 residues long and can be determined in silico directly from the amino acid sequence. Delineation of such variations in helical habit is critical to understanding the complex structure-function relationships of polytopic proteins and for drug discovery. The success of our current methodology foretells development of similar prediction tools capable of identifying other structural motifs from sequence alone. The method described here has been implemented and is available on the World Wide Web at http://cbcsrv.watson.ibm.com/Ttkw.html.

Amino Acid Sequence↗

Cloning and characterization of human CAGLP gene encoding a novel EF-hand protein.

The EF-hand proteins, containing conserved Ca2+ binding motifs, play important roles in many biological processes. Through data mining, a novel human gene, CAGLP (calglandulin-like protein) was predicted and subsequently isolated from human skeleton muscle. The open reading frame of CAGLP is 543 bp in length, coding a putative Ca2+ binding protein with four EF-hand motifs. The deduced amino acid sequence of CAGLP displays high similarity with Bothrops insularis snake protein calglandulin (80%). The results of PCR amplification using cDNA from 17 human tissues indicated that human CAGLP is expressed in prostate, thymus, heart, skeleton muscle, bone marrow and ovary. Functional CAGLP::EGFP (enhanced green fluorescent protein) fusion protein revealed that CAGLP accumulated through-out Hela cells. Western blot using anti-EGFP antibodies indicated that the CAGLP protein has a molecular weight of about 19 kD. A phylogenetic tree showed that CAGLP and calglandulin may be orthologous proteins representing a distinct group in the EF-hand proteins.

Base Sequence↗