Opinion: Mapping context and content: the BrainMap model.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
We present an approach to evaluate the support for candidate genes as quantitative trait loci (QTLs) within the context of genome-wide map-based cloning strategies. To establish candidacy, a bacterial artificial chromosome (BAC) clone containing a putative candidate gene is physically assigned to an anchored linkage map to localise the gene relative to an identified QTL effect. Microsatellite loci derived from BAC clones containing an established candidate gene are integrated into the linkage map facilitating the evaluation by interval analysis of the statistical support for QTL identity. Permutation analysis is employed to determine experiment-wise statistical support. The approach is illustrated for the growth hormone 1 (GH1) gene and growth and carcass phenotypes in cattle. Polymerase chain reaction (PCR) primers which amplify a 441 bp fragment of GH1 were used to systematically screen a bovine BAC library comprising 60,000 clones and with a 95% probability of containing a single copy sequence. The presence of GH1 in BAC-110R2C3 was confirmed by sequence analysis of the PCR product from this clone and by the physical assignment of BAC110R2C3 to bovine chromosome 19 (BTA19) band 22 by fluorescence in situ hybridisation (FISH). Microsatellite KHGH1 was isolated from BAC110R2C3 and scored in 529 reciprocal backcross and F2 fullsib progeny from 41 resource families derived from Angus (Bos taurus) and Brahman (Bos indicus). The microsatellite KHGH1 was incorporated into a framework genetic map of BTA19 comprising 12 microsatellite loci, the erythrocyte antigen T and a GH1-TaqI restriction fragment length polymorphism (RFLP). Interval analysis localised effects of taurus vs. indicus alleles on subcutaneous fat and the percentage of either extractable fat from the Iongissimus dorsi muscle to the region of BTA19 harbouring GH1.
The goal of this study was to identify and map genes expressed during the elongation phase of embryogenesis in swine. Expressed sequence tags were analysed from a previously described porcine cDNA library prepared from elongating swine embryos. Average insert length of randomly selected clones was approximately 600 bp, with a range from < 100 to > 2500 bp. Single-pass, coding strand sequences from 1132 independent clones were compared with the GenBank non-redundant (nr) database via BLASTN analysis to identify potential porcine homologous of known genes. Among these sequences, 781 (69%) showed significant (score > 300) homology to non- mitochondrial sequences previously deposited in GenBank. Sequences matching interleucin 1 beta and thymosin beta 10 were most frequently observed (24 and 18 clones, respectively), in addition to matches with 310 other distinct genes. No significant match in the GenBank nr database was obtained for 303 sequences. Analysis demonstrated that 151 (50%) had open reading frames (ORF) extending at least 50 codons from the first base of the clone insert. Genetic markers were developed and used to map a subset of 17 genes, selected on the basis of function or of the ability to design primers that successfully amplified porcine genomic DNA, to 10 different porcine chromosomes, providing a set of mapped markers corresponding to genes expressed during conceptus elongation.
The mechanism of attachment of acetylcholinesterase (AChE) to neuronal membranes in interneuronal synapses is poorly understood. We have isolated, sequenced, and cloned a hydrophobic protein that copurifies with AChE from human caudate nucleus and that we propose forms a part of a complex of membrane proteins attached to this enzyme. It is a short protein of 136 amino acids and has a molecular mass of 18 kDa. The sequence contains stretches of both hydrophobic and hydrophilic amino acids and two cysteine residues. Analysis of the genomic sequence reveals that the coding region is divided among five short exons. Fluorescence in situ hybridization localizes the gene to chromosome 6p21.32-p21.2. Northern blot analysis shows that this gene is widely expressed in the brain with an expression pattern that parallels that of AChE.
We present evidence for the existence of a novel chromosome 2q32 locus involved in the pathogenesis of isolated cleft palate. We have studied two unrelated patients with strikingly similar clinical features, in whom there are apparently balanced, de novo cytogenetic rearrangements involving the same region of chromosome 2q. Both children have cleft palate, facial dysmorphism, and mild learning disability. Their karyotypes were originally reported as 46, XX, t(2;7)(q33;p21) and 46, XX, t(2;11)(q33;p14). However, our molecular cytogenetic analyses localize both translocation breakpoints to a small region between markers D2S311 and D2S116. This suggests that the true location of these breakpoints is 2q32 rather than 2q33. To obtain independent support for the existence of a cleft-palate locus in 2q32, we performed a detailed statistical analysis for all cases in the human cytogenetics database of nonmosaic, single, contiguous autosomal deletions associated with orofacial clefting. This revealed 2q32 to be one of only three chromosomal regions in which haploinsufficiency is significantly associated with isolated cleft palate. In combination, our data provide strong evidence for the location at 2q32 of a gene that is critical to the development of the secondary palate. The close proximity of these two translocation breakpoints should also allow rapid progress toward the positional cloning of this cleft-palate gene.
Detection of genotyping errors and integration of such errors in statistical analysis are relatively neglected topics, given their importance in gene mapping. A few inopportunely placed errors, if ignored, can tremendously affect evidence for linkage. The present study takes a fresh look at the calculation of pedigree likelihoods in the presence of genotyping error. To accommodate genotyping error, we present extensions to the Lander-Green-Kruglyak deterministic algorithm for small pedigrees and to the Markov-chain Monte Carlo stochastic algorithm for large pedigrees. These extensions can accommodate a variety of error models and refrain from simplifying assumptions, such as allowing, at most, one error per pedigree. In principle, almost any statistical genetic analysis can be performed taking errors into account, without actually correcting or deleting suspect genotypes. Three examples illustrate the possibilities. These examples make use of the full pedigree data, multiple linked markers, and a prior error model. The first example is the estimation of genotyping error rates from pedigree data. The second-and currently most useful-example is the computation of posterior mistyping probabilities. These probabilities cover both Mendelian-consistent and Mendelian-inconsistent errors. The third example is the selection of the true pedigree structure connecting a group of people from among several competing pedigree structures. Paternity testing and twin zygosity testing are typical applications.
Large exploratory studies, including candidate-gene-association testing, genomewide linkage-disequilibrium scans, and array-expression experiments, are becoming increasingly common. A serious problem for such studies is that statistical power is compromised by the need to control the false-positive rate for a large family of tests. Because multiple true associations are anticipated, methods have been proposed that combine evidence from the most significant tests, as a more powerful alternative to individually adjusted tests. The practical application of these methods is currently limited by a reliance on permutation testing to account for the correlated nature of single-nucleotide polymorphism (SNP)-association data. On a genomewide scale, this is both very time-consuming and impractical for repeated explorations with standard marker panels. Here, we alleviate these problems by fitting analytic distributions to the empirical distribution of combined evidence. We fit extreme-value distributions for fixed lengths of combined evidence and a beta distribution for the most significant length. An initial phase of permutation sampling is required to fit these distributions, but it can be completed more quickly than a simple permutation test and need be done only once for each panel of tests, after which the fitted parameters give a reusable calibration of the panel. Our approach is also a more efficient alternative to a standard permutation test. We demonstrate the accuracy of our approach and compare its efficiency with that of permutation tests on genomewide SNP data released by the International HapMap Consortium. The estimation of analytic distributions for combined evidence will allow these powerful methods to be applied more widely in large exploratory studies.
Knowledge on interactions between molecules in living cells is indispensable for theoretical analysis and practical applications in modern genomics and molecular biology. Building such networks relies on the assumption that the correct molecular interactions are known or can be identified by reading a few research articles. However, this assumption does not necessarily hold, as truth is rather an emerging property based on many potentially conflicting facts. This paper explores the processes of knowledge generation and publishing in the molecular biology literature using modelling and analysis of real molecular interaction data. The data analysed in this article were automatically extracted from 50000 research articles in molecular biology using a computer system called GeneWays containing a natural language processing module. The paper indicates that truthfulness of statements is associated in the minds of scientists with the relative importance (connectedness) of substances under study, revealing a potential selection bias in the reporting of research results. Aiming at understanding the statistical properties of the life cycle of biological facts reported in research articles, we formulate a stochastic model describing generation and propagation of knowledge about molecular interactions through scientific publications. We hope that in the future such a model can be useful for automatically producing consensus views of molecular interaction data.
MOTIVATION: The living cell is a complex machine that depends on the proper functioning of its numerous parts, including proteins. Understanding protein functions and how they modify and regulate each other is the next great challenge for life-sciences researchers. The collective knowledge about protein functions and pathways is scattered throughout numerous publications in scientific journals. Bringing the relevant information together becomes a bottleneck in a research and discovery process. The volume of such information grows exponentially, which renders manual curation impractical. As a viable alternative, automated literature processing tools could be employed to extract and organize biological data into a knowledge base, making it amenable to computational analysis and data mining. RESULTS: We present MedScan, a completely automated natural language processing-based information extraction system. We have used MedScan to extract 2976 interactions between human proteins from MEDLINE abstracts dated after 1988. The precision of the extracted information was found to be 91%. Comparison with the existing protein interaction databases BIND and DIP revealed that 96% of extracted information is novel. The recall rate of MedScan was found to be 21%. Additional experiments with MedScan suggest that MEDLINE is a unique source of diverse protein function information, which can be extracted in a completely automated way with a reasonably high precision. Further directions of the MedScan technology improvement are discussed. AVAILABILITY: MedScan is available for commercial licensing from Ariadne Genomics, Inc.
Information on molecular networks, such as networks of interacting proteins, comes from diverse sources that contain remarkable differences in distribution and quantity of errors. Here, we introduce a probabilistic model useful for predicting protein interactions from heterogeneous data sources. The model describes stochastic generation of protein-protein interaction networks with real-world properties, as well as generation of two heterogeneous sources of protein-interaction information: research results automatically extracted from the literature and yeast two-hybrid experiments. Based on the domain composition of proteins, we use the model to predict protein interactions for pairs of proteins for which no experimental data are available. We further explore the prediction limits, given experimental data that cover only part of the underlying protein networks. This approach can be extended naturally to include other types of biological data sources.
UNLABELLED: Repair-FunMap is a functional database of the DNA repair systems. This database contains not only the proteins directly involved in DNA repair, but also the proteins that interact with the DNA repair proteins. A protein interaction network associated with the human DNA repair processes was established according to the functional relationship between proteins in the database. This network represents the current knowledge on the intrinsic signaling pathways related to DNA repair. The Repair-FunMap could become an essential resource center for cancer research, providing clues to understanding the inter-relationship between proteins in the network, and to building scientific models of the DNA repair processes. AVAILABILITY: http://astro.temple.edu/~feng/Servers/BioinformaticServers.htm
MOTIVATION: Although there are several databases storing protein-protein interactions, most such data still exist only in the scientific literature. They are scattered in scientific literature written in natural languages, defying data mining efforts. Much time and labor have to be spent on extracting protein pathways from literature. Our aim is to develop a robust and powerful methodology to mine protein-protein interactions from biomedical texts. RESULTS: We present a novel and robust approach for extracting protein-protein interactions from literature. Our method uses a dynamic programming algorithm to compute distinguishing patterns by aligning relevant sentences and key verbs that describe protein interactions. A matching algorithm is designed to extract the interactions between proteins. Equipped only with a dictionary of protein names, our system achieves a recall rate of 80.0% and precision rate of 80.5%. AVAILABILITY: The program is available on request from the authors.
SUMMARY: PDZBase is a database that aims to contain all known PDZ-domain-mediated protein-protein interactions. Currently, PDZBase contains approximately 300 such interactions, which have been manually extracted from > 200 articles. The database can be queried through both sequence motif and keyword-based searches, and the sequences of interacting proteins can be visually inspected through alignments (for the comparison of several interactions), or as residue-based diagrams including schematic secondary structure information (for individual complexes).
SUMMARY: The MIPS mammalian protein-protein interaction database (MPPI) is a new resource of high-quality experimental protein interaction data in mammals. The content is based on published experimental evidence that has been processed by human expert curators. We provide the full dataset for download and a flexible and powerful web interface for users with various requirements.
MOTIVATION: The advent of high-throughput experiments in molecular biology creates a need for methods to efficiently extract and use information for large numbers of genes. Recently, the associative concept space (ACS) has been developed for the representation of information extracted from biomedical literature. The ACS is a Euclidean space in which thesaurus concepts are positioned and the distances between concepts indicates their relatedness. The ACS uses co-occurrence of concepts as a source of information. In this paper we evaluate how well the system can retrieve functionally related genes and we compare its performance with a simple gene co-occurrence method. RESULTS: To assess the performance of the ACS we composed a test set of five groups of functionally related genes. With the ACS good scores were obtained for four of the five groups. When compared to the gene co-occurrence method, the ACS is capable of revealing more functional biological relations and can achieve results with less literature available per gene. Hierarchical clustering was performed on the ACS output, as a potential aid to users, and was found to provide useful clusters. Our results suggest that the algorithm can be of value for researchers studying large numbers of genes. AVAILABILITY: The ACS program is available upon request from the authors.
MOTIVATION: An enormous number of protein-protein interaction relationships are buried in millions of research articles published over the years, and the number is growing. Rediscovering them automatically is a challenging bioinformatics task. Solutions to this problem also reach far beyond bioinformatics. RESULTS: We study a new approach that involves automatically discovering English expression patterns, optimizing them and using them to extract protein-protein interactions. In a sister paper, we described how to generate English expression patterns related to protein-protein interactions, and this approach alone has already achieved precision and recall rates significantly higher than those of other automatic systems. This paper continues to present our theory, focusing on how to improve the patterns. A minimum description length (MDL)-based pattern-optimization algorithm is designed to reduce and merge patterns. This has significantly increased generalization power, and hence the recall and precision rates, as confirmed by our experiments. AVAILABILITY: http://spies.cs.tsinghua.edu.cn.
UNLABELLED: A number of freely available text mining tools have been put together to extract highly reliable Drosophila gene interaction data from text. The system has been tested with The Interactive Fly, showing low recall (27-34%), but very high precision (93-97%). AVAILABILITY: The extracted data and a web interface for submission of texts to GIFT analysis are available at http://gift.cryst.bbk.ac.uk/gift CONTACT: n.domedel_puig@cryst.bbk.ac.uk SUPPLEMENTARY INFORMATION: Additional documentation, such as the dictionaries and the reference sets, are available at the GIFT website.