Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

CCG repeats in cDNAs from human brain.

Expansion mutations of trinucleotide repeats and other units of unstable DNA have been proposed to account for at least some of the genetic susceptibility to a number of neuropsychiatric disorders, including bipolar affective disorder, schizophrenia, autism, and panic disorder. To generate additional candidate genes for these and other disorders, cDNA libraries from human brain were probed at high stringency for clones containing CCG, CGC, GCC, CGG, GCG, and GGC repeats (referred to collectively as CCG repeats). Some 18 cDNAs containing previously unpublished or uncharacterized repeats were characterized for chromosomal locus, repeat length polymorphism, and similarity to genes of known function. The cDNAs were also compared with the 37 human genes with eight or more consecutive CCG triplets in GenBank. The repeats were mapped to a number of loci, including 1p34, 2p11.2, 2q30-32, 3p21, 3p22, 4q35, 6q22, 7qter, 13p13, 17q24, 18p11, 19p13.3, 20q12, 20q13.3, and 22q12. Length polymorphism was detected in 50% of the repeats. The newly cloned cDNAs include a complete transcript of human neurexin-1B, a portion of BCNG-1 (a newly described brain-specific ion channel), a previously unreported polymorphic repeat located in the 5' UTR region of the guanine nucleotide-binding protein (G-protein) beta2 subunit, and a human version of the mouse proline-rich protein 7. This list of cDNAs should expedite the search for expansion mutations associated with diseases of the central nervous system.

Brain Chemistry↗

Separation technologies for glycomics.

Progress in genome projects has provided us with fundamentals on genetic information; however, the functions of a large number of genes remain to be elucidated. To understand the in vivo functions of eukaryotic genes, it is essential to grasp the features of their post-translational modifications. Among them, protein glycosylation is a central issue to be discussed, considering the predominant roles of glycoproteins in cell-cell and cell-substratum recognition events in multicellular organisms. In this context, it is necessary to establish a core strategy for analyzing glycosylated proteins under the concept of the "glycome" [Trends Glycosci. Glycotechnol. 12 (2000) 1]. Though the term glycome should be defined, in analogy to the genome and proteome, as "a whole set of glycans produced in a single organism", here we propose a glycome project specifically focusing on glycoproteins. Principal objectives in the project are to identify: (1) which genes encode glycoproteins (i.e. genome information); (2) which sites among potential glycosylation sites are actually glycosylated (i.e. glycosylation site information); (3) what are the structures of glycans (i.e. structural information); and (4) what are the effects (functions) of glycosylation (functional information). For these purposes, two affinity technologies have been introduced. One is named the "glyco-catch method" to identify genes encoding glycoproteins [Proteomics 1 (2001) 295], and the other is the recently reinforced "frontal affinity chromatography" [J. Chromatogr. A 890 (2000) 261]. By the former method, genes that encode glycoproteins as well as glycosylation sites are systematically identified by the efficient combination of conventional lectin-affinity chromatography and contemporary in silico database searching. The following three actions have been devised for rapid and systematic characterization of glycans: (1) mass spectrometry to acquire exact mass information; (2) 2-D/3-D mapping to obtain refined chemical information; and (3) reinforced frontal affinity chromatography to determine affinity constants (K(a)-values) for a set of lectins. Pyridylaminated glycans are used throughout the characterization processes. In this review, the concept and strategy of glycomic approaches are described referring to the on-going glycome project focused on the nematode Caenorhabditis elegans.

Glycoproteins↗

222 base pairs in NS5B region and the determination of hepatitis C virus genotype 6.

OBJECTIVE: The present study was performed to genotype hepatitis C virus (HCV) by direct sequencing of a 222-bp nucleotide in the NS5B region and comparing the results with those of direct sequencing in the core region. We investigated a new region for HCV genotyping which gave the best performance to discriminate HCV genotype 6a, the unique genotype found in Southeast Asia. METHODS: Plasma samples taken from 57 HCV-infected blood donors were used in this study. RT-PCR products were amplified using primers located in the NS5B region. The 222-bp PCR products were purified and sequenced. The genotype of HCV isolates were obtained by phylogenetic analysis and compared with HCV reference strains stored in the GenBank database. The HCV sequences clustering in the same node were considered to be of the same genotype. RESULTS: Thirty-one, 22 and 4 samples of HCV genotype 3a, 1a and 1b, respectively, were analyzed by this method. Upon comparison with genotyping in the core region, 86 and 14% of the samples yielded concordant and discordant genotype results, respectively. The majority of discordant results (63%; 5 of 8) was observed with HCV genotype 6a which yielded 6a upon core sequencing as opposed to 1a or 3a upon NS5B sequencing. CONCLUSION: HCV genotype 6a obtained by direct sequencing in the core region could not be unequivocally arrived at by sequencing 222 bp in the NS5B region. Hence, sequencing in the core region is preferable for genotyping our specimens, even though longer PCR products are required as this method enables discrimination between genotype 6a and the remaining genotypes.

Base Pairing↗

Analysis of high-resolution HapMap of DTNBP1 (Dysbindin) suggests no consistency between reported common variant associations and schizophrenia.

DTNBP1 was first identified as a putative schizophrenia-susceptibility gene in Irish pedigrees, with a report of association to common genetic variation. Several replication studies have reported confirmation of an association to DTNBP1 in independent European samples; however, reported risk alleles and haplotypes appear to differ between studies, and comparison among studies has been confounded because different marker sets were employed by each group. To facilitate evaluation of existing evidence of association and further work, we supplemented the extensive genotype data, available through the International HapMap Project (HapMap), about DTNBP1 by specifically typing all associated single-nucleotide polymorphisms reported in each of the studies of the Centre d'Etude du Polymorphisme Humain (CEPH)-derived HapMap sample (CEU). Using this high-density reference map, we compared the putative disease-associated haplotype from each study and found that the association studies are inconsistent with regard to the identity of the disease-associated haplotype at DTNBP1. Specifically, all five "replication" studies define a positively associated haplotype that is different from the association originally reported. We further demonstrate that, in all six studies, the European-derived populations studied have haplotype patterns and frequencies that are consistent with HapMap CEU samples (and each other). Thus, it is unlikely that population differences are creating the inconsistency of the association studies. Evidence of association is, at present, equivocal and unsatisfactory. The new dense map of the region may be valuable in more-comprehensive follow-up studies.

Alleles↗

Isolation and characterization of cell-specific cDNA clones from a subtractive library of the ocular ciliary body of a single normal human donor: transcription and synthesis of plasma proteins.

A subtractive cDNA library was developed for the purpose of identifying cell-specific genes expressed within the human ocular ciliary body, a tissue responsible for regulating aqueous humor secretion and intraocular pressure. Partial DNA sequence of a large number of cDNA clones and homology searches of nucleic acid and protein databases revealed significant homologies to at least 90 independently known genes. A group of biologically significant genes that were previously not known to have transcriptional expression in the ciliary body, complement component C4; alpha 2 macroglobulin; selenoprotein-P; and apolipoprotein D, were further demonstrated by Northern hybridization. Antibodies to these and other proteins (i.e., tyrosinase-related protein and pigment epithelium-derived factor) confirmed their cell-restricted expression in ciliary epithelial cells (pigmented, nonpigmented), or vascular endothelial cells. We provide evidence that two human plasma proteins, complement component C4 and alpha 2-macroglobulin are metabolically labeled with [35S]methionine in ciliary processes explants, suggesting that the ciliary body is an organ of synthesis and secretion of plasma proteins present in aqueous humor. These results challenge the notion that plasma proteins in aqueous humor are imported from outside of the eye. The subtractive cDNA library reported in this work should be very useful for identifying potential candidate genes in ocular abnormalities affecting the ciliary body, or involved in the regulation of intraocular pressure.

Adult↗

The association of Alu repeats with the generation of potential AU-rich elements (ARE) at 3' untranslated regions.

BACKGROUND: A significant portion (about 8% in the human genome) of mammalian mRNA sequences contains AU (Adenine and Uracil) rich elements or AREs at their 3' untranslated regions (UTR). These mRNA sequences are usually stable. However, an increasing number of observations have been made of unstable species, possibly depending on certain elements such as Alu repeats. ARE motifs are repeats of the tetramer AUUU and a monomer A at the end of the repeats ((AUUU)nA). The importance of AREs in biology is that they make certain mRNA unstable. Proto-oncogene, such as c-fos, c-myc, and c-jun in humans, are associated with AREs. Although it has been known that the increased number of ARE motifs caused the decrease of the half-life of mRNA containing ARE repeats, the exact mechanism is as of yet unknown. We analyzed the occurrences of AREs and Alu and propose a possible mechanism for how human mRNA could acquire and keep AREs at its 3' UTR originating from Alu repeats. RESULTS: Interspersed in the human genome, Alu repeats occupy 5% of the 3' UTR of mRNA sequences. Alu has poly-adenine (poly-A) regions at its end, which lead to poly-thymine (poly-T) regions at the end of its complementary Alu. It has been found that AREs are present at the poly-T regions. From the 3' UTR of the NCBI's reference mRNA sequence database, we found nearly 40% (38.5%) of ARE (Class I) were associated with Alu sequences (Table 1) within one mismatch allowance in ARE sequences. Other ARE classes had statistically significant associations as well. This is far from a random occurrence given their limited quantity. At each ARE class, random distribution was simulated 1,000 times, and it was shown that there is a special relationship between ARE patterns and the Alu repeats. CONCLUSION: AREs are mediating sequence elements affecting the stabilization or degradation of mRNA at the 3' untranslated regions. However, AREs' mechanism and origins are unknown. We report that Alu is a source of ARE. We found that half of the longest AREs were derived from the poly-T regions of the complementary Alu.

3' Untranslated Regions↗

Identification of genes expressed in primate primordial oocytes.

BACKGROUND: The factors involved in oocyte survival and transition from quiescence to the growing phenotype remain unknown. Herein we report genes that are differentially expressed in the primordial oocyte revealed by DNA arrays. METHODS: Primordial oocytes were captured selectively in rhesus monkey ovary sections using laser capture microdissection. The RNA was extracted and amplified in two rounds by T7-based linear RNA amplification, fluorescence labelled and then hybridized to human cDNA arrays containing 7680 elements. RNA from human placenta served as a reference sample. RESULTS: Ninety-five genes were found to be consistently expressed at a higher level in primordial oocytes. Expression of several of these genes in the oocyte has been reported before, e.g. deleted in azoospermia (DAZ), prohibitin and transglutaminase 2. Oocyte expression of several novel transcripts revealed on array hybridization, such as gene 33, ubiquitin-conjugating enzyme E2A, G1 to S phase transition 1, growth arrest and DNA damage-inducible (GADD), and dendritic cell-derived ubiquitin-like protein (DC-UbP) was confirmed by in situ hybridization. Some array-identified gene products [integrin beta3, alpha-tubulin, regulatory telomere elongation protein (RAP1) and cellular repressor of EIA-stimulated genes (CREG protein)] were detected in human oocytes by immunofluorescence. Bioinformatic analysis of the oocyte-enriched transcripts reveals a functional profile summarized as follows: cell cycle (14%); transporter (13%); signal transduction (10%); cytoskeletal (7%); transcription factor (5%); immune response (5%); apoptosis-related (5%); RNA processing (5%); and the remainder of miscellaneous categories. CONCLUSIONS: These observations may contribute to the elucidation of molecular pathways involved in oocyte survival and maturation.

Animals↗

Long-term nutritional outcome after pediatric intestinal transplantation.

BACKGROUND/PURPOSE: The aim of this study was to describe the long-term nutritional status of a large population of children after intestinal transplantation and to identify factors associated with nutritional outcomes. METHODS: Longitudinal anthropometric data are maintained in a database registry for all patients referred to our Intestinal Care Center (ICC). Z-scores for weight and height were calculated biannually over a maximum of 2 years, and associations between baseline and follow-up laboratory measures and growth were evaluated for patients greater than 6 months post intestinal transplant. RESULTS: Since the inception of the ICC in December 1996, 24 pediatric patients (18 boys, 18 white) received an isolated small bowel or small bowel/liver transplant (median age, 3.2 years). The majority of cases (75%) had been diagnosed with surgical short bowel syndrome and were dependent on total parenteral nutrition (TPN) at the time of transplant. Of the 23 patients who survived the initial postoperative period, 87% were weaned from TPN to an amino-acid or peptide-based enteral formula or solid food within 3 months. A positive trend in z-scores for weight and height/length was observed in only 30% and 26% of patients, respectively, during the follow-up period. Although mean albumin levels increased significantly from 2.8 to 3.1 mg/dl by 6 months posttransplant (P <.01) no difference in alkaline phosphatase was found over time. Steroid doses were weaned within 3 to 4 months after transplantation but not discontinued. The cumulative survival rate was 91% at 1 year and 86% at 2 years posttransplant, whereas those weaned from TPN achieved 100% and 94% survival, respectively. CONCLUSIONS: Attainment of positive linear growth remains a challenge in the pediatric transplant population despite successful liberation from TPN, protein anabolism, and high survival rates. Further investigation into alternative methods of nutritional evaluation and manipulation as well as the use of growth factors to enhance the growth process need to be investigated.

Body Height↗

Starvation-induced expression of SspA and SspB: the effects of a null mutation in sspA on Escherichia coli protein synthesis and survival during growth and prolonged starvation.

Maxicell labelling and two-dimensional gel electrophoresis (2-D PAGE) have identified the proteins encoded by sspA and sspB (SspA, SspB) as proteins D27.1 and A25.8, respectively, in the Escherichia coli gene-protein database. SspA expression increases with decreasing growth rate and is induced by glucose, nitrogen, phosphate or amino acid starvation. The promoter, Pssp, is similar to gearbox promoters. Inactivation of SspA (sspA::neo) blocks sspB expression. [35S]-methionine-labelled proteins synthesized during growth and during stationary phase are different in delta sspA strains compared to sspA+ strains. This difference is enhanced during extended stationary phase (24-72 h). Long-term (10 d) viability of arginine-starved isogenic strains shows that sspA+ cultures remain viable significantly longer than delta sspA mutants. 2-D PAGE of proteins expressed during exponential growth shows that expression of at least 11 proteins is altered in delta sspA strains. A functional relA gene is required for sspA to affect protein synthesis.

Adaptation, Biological↗

Quality assessment of NMR structures: a statistical survey.

A statistical analysis is reported of experimental data and coordinates of a set of 97 NMR structures deposited in the PDB. The aim is to assess the quality of these structures in relation to the amount of experimental information. Experimental restraints were analysed using the program AQUA. Many nomenclature inconsistencies between deposited restraint and coordinate files were observed. The experimental restraint files were found to contain a high proportion of redundant restraints. Procedures for analysing and correcting the inconsistencies and restraint counts are described. The analysis of NOE restraint violations (using AQUA) and of a wide variety of geometrical quality indicators (using PROCHECK-NMR and WHAT IF) provides a reference for other NMR structure determinations. The extent of NOE violations is anti-correlated with the quality of the Ramachandran map. The precision as measured by the circular variance of backbone dihedral angles, does increase with the amount of experimental data, as expected, but is sometimes overestimated. Bond lengths, bond angles and planarity of groups can deviate considerably from ideal values. Outliers appear to cluster per laboratory, indicating that the results depend on particulars of refinement protocols and/or software. We have identified a problem of atom overlap in a number of refined structures.We recommend adhering to the standard nomenclature as put forward by an IUPAC Task Group, to ensure consistency between restraints and coordinates, and to omit redundant restraints from the deposition. The results obtained from this analysis and the AQUA program are available through the World Wide Web.

Databases, Factual↗

Phylogenetic analysis of rabbit haemorrhagic disease virus in France between 1993 and 2000, and the characterisation of RHDV antigenic variants.

The first molecular epidemiological study of Rabbit haemorrhagic disease virus undertaken in France between 1988 and 1995, identified three genogroups, two of which (G1, G2) disappeared quickly. We used immunocapture-RT-PCR and sequencing to analyse 104 new RHDV isolates collected between 1993 and 2000. One isolate was obtained in 2000 from a French overseas territory, the Reunion Island. The nucleotide sequences of these isolates were aligned with those of some French RHDV isolates representative of the three genogroups previously identified, of some reference strains and German and American RHDV antigenic variants. Despite the low degree of nucleotide sequence variation, three new genogroups (G4 to G6) were identified with significant bootstrap values. Two of these genogroups (G4 and G5) were related to the year in which the RHDV isolates were collected. Genogroup G4 emerged from genogroup G3, which has now disappeared. Genogroup G5 is a new independent group. The genogroup G6 contained an isolate collected in mainland France in 1999 and the isolate collected from the Reunion Island, as well as German and American RHDV variants. Multiple sequence alignments of the VP60 gene and antigenic analysis with monoclonal antibodies demonstrated that these French isolates are two new isolates of the RHDV variant.

Amino Acid Sequence↗

GEM: a Gaussian Evolutionary Method for predicting protein side-chain conformations.

We have developed an evolutionary approach to predicting protein side-chain conformations. This approach, referred to as the Gaussian Evolutionary Method (GEM), combines both discrete and continuous global search mechanisms. The former helps speed up convergence by reducing the size of rotamer space, whereas the latter, integrating decreasing-based Gaussian mutations and self-adaptive Gaussian mutations, continuously adapts dihedrals to optimal conformations. We tested our approach on 38 proteins ranging in size from 46 to 325 residues and showed that the results were comparable to those using other methods. The average accuracies of our predictions were 80% for chi(1), 66% for chi(1 + 2), and 1.36 A for the root mean square deviation of side-chain positions. We found that if our scoring function was perfect, the prediction accuracy was also essentially perfect. However, perfect prediction could not be achieved if only a discrete search mechanism was applied. These results suggest that GEM is robust and can be used to examine the factors limiting the accuracy of protein side-chain prediction methods. Furthermore, it can be used to systematically evaluate and thus improve scoring functions.

Algorithms↗

Cluster analysis of consensus water sites in thrombin and trypsin shows conservation between serine proteases and contributions to ligand specificity.

Cluster analysis is presented as a technique for analyzing the conservation and chemistry of water sites from independent protein structures, and applied to thrombin, trypsin, and bovine pancreatic trypsin inhibitor (BPTI) to locate shared water sites, as well as those contributing to specificity. When several protein structures are superimposed, complete linkage cluster analysis provides an objective technique for resolving the continuum of overlaps between water sites into a set of maximally dense microclusters of overlapping water molecules, and also avoids reliance on any one structure as a reference. Water sites were clustered for ten superimposed thrombin structures, three trypsin structures, and four BPTI structures. For thrombin, 19% of the 708 microclusters, representing unique water sites, contained water molecules from at least half of the structures, and 4% contained waters from all 10. For trypsin, 77% of the 106 microclusters contained water sites from at least half of the structures, and 57% contained waters from all three. Water site conservation correlated with several environmental features: highly conserved microclusters generally had more protein atom neighbors, were in a more hydrophilic environment, made more hydrogen bonds to the protein, and were less mobile. There were significant overlaps between thrombin and trypsin conserved water sites, which did not localize to their similar active sites, but were concentrated in buried regions including the solvent channel surrounding the Na+ site in thrombin, which is associated with ligand selectivity. Cluster analysis also identified water sites conserved in thrombin but not trypsin, and vice versa, providing a list of water sites that may contribute to ligand discrimination. Thus, in addition to facilitating the analysis of water sites from multiple structures, cluster analysis provides a useful tool for distinguishing between conserved features within a protein family and those conferring specificity.

Animals↗

Bacterial taxonomics: finding the wood through the phylogenetic trees.

Bacterial taxonomy comprises systematics (theory of classification), nomenclature (formal process of naming), and identification. There are two basic approaches to classification. Similarities may be derived between microorganisms by numerical taxonomic methods based on a range of present-day observable characteristics (phenetics), drawing in particular on conventional morphological and physiological test characters as well as chemotaxonomic markers such as whole-cell protein profiles, mol% G+C content, and DNA-DNA homologies. By contrast, phylogenetics, the process of reconstructing possible evolutionary relationships, uses nucleotide sequences from conserved genes that act as molecular chronometers. A combination of both phenetics and phylogenetics is referred to as polyphasic taxonomy, and is the recommended strategy in description of new species and genera. Numerical analysis of small-subunit ribosomal RNA genes (rDNA) leading to the construction of branching trees representing the distance of divergence from a common ancestor has provided the mainstay of microbial phylogenetics. The approach has some limitations, particularly in the discrimination of closely related taxa, and there is a growing interest in the use of alternative loci as molecular chronometers, such as gyrA and RNAase P sequences. Comparison of the degree of congruence between phylogenetic trees derived from different genes provides a valuable test of the extent they represent gene trees or species trees. Rapid expansion in genome sequences will provide a rich source of data for future taxonomic analysis that should take into account population structure of taxa and novel methods for analysis of nonclonal bacterial populations.

Bacteria↗

Independent mutational events are rare in the ATM gene: haplotype prescreening enhances mutation detection rate.

Mutations in the ATM gene are responsible for the autosomal recessive disorder ataxia-telangiectasia (A-T). Many different mutations have been identified using various techniques, with detection efficiencies ranging from 57 to 85%. In this study, we employed short tandem repeat (STR) haplotypes to enhance mutation identification in 55 unrelated A-T families of Iberian origin (20 Spanish, 17 Brazilian, and 18 Hispanic-American); we were able to identify 95% of the expected mutations. Allelic sizes were standardized based on a reference sample (CEPH 1347-2). Subsequent mutation screening was performed by PTT, SSCP, and DHPLC, and abnormal regions were sequenced. Many STR haplotypes were found within each population and six haplotypes were observed across several of these populations. Single nucleotide polymorphism (SNP) haplotypes further suggested that most of these common mutations are ancestrally related, and not hot spots. However, two mutations (8977C>T and 8264_8268delATAAG) may indeed be recurring mutational events. Common haplotypes were present in 13 of 20 Spanish A-T families (65%), in 11 of 17 Brazilian A-T families (65%), and, in contrast, in only eight of 18 Hispanic-American families (44%). Three mutations were identified that would be missed by conventional screening strategies. In all, 62 different mutations (28 not previously reported) were identified and their associated haplotypes defined, thereby establishing a new database for Iberian A-T families, and extending the spectrum of worldwide ATM mutations.

Ataxia Telangiectasia↗

Characterization of human RNA polymerase III identifies orthologues for Saccharomyces cerevisiae RNA polymerase III subunits.

Unlike Saccharomyces cerevisiae RNA polymerase III, human RNA polymerase III has not been entirely characterized. Orthologues of the yeast RNA polymerase III subunits C128 and C37 remain unidentified, and for many of the other subunits, the available information is limited to database sequences with various degrees of similarity to the yeast subunits. We have purified an RNA polymerase III complex and identified its components. We found that two RNA polymerase III subunits, referred to as RPC8 and RPC9, displayed sequence similarity to the RNA polymerase II RPB7 and RPB4 subunits, respectively. RPC8 and RPC9 associated with each other, paralleling the association of the RNA polymerase II subunits, and were thus paralogues of RPB7 and RPB4. Furthermore, the complex contained a prominent 80-kDa polypeptide, which we called RPC5 and which corresponded to the human orthologue of the yeast C37 subunit despite limited sequence similarity. RPC5 associated with RPC53, the human orthologue of S. cerevisiae C53, paralleling the association of the S. cerevisiae C37 and C53 subunits, and was required for transcription from the type 2 VAI and type 3 human U6 promoters. Our results provide a characterization of human RNA polymerase III and show that the RPC5 subunit is essential for transcription.

Amino Acid Sequence↗

Combining evidence, biomedical literature and statistical dependence: new insights for functional annotation of gene sets.

BACKGROUND: Large-scale genomic studies based on transcriptome technologies provide clusters of genes that need to be functionally annotated. The Gene Ontology (GO) implements a controlled vocabulary organised into three hierarchies: cellular components, molecular functions and biological processes. This terminology allows a coherent and consistent description of the knowledge about gene functions. The GO terms related to genes come primarily from semi-automatic annotations made by trained biologists (annotation based on evidence) or text-mining of the published scientific literature (literature profiling). RESULTS: We report an original functional annotation method based on a combination of evidence and literature that overcomes the weaknesses and the limitations of each approach. It relies on the Gene Ontology Annotation database (GOA Human) and the PubGene biomedical literature index. We support these annotations with statistically associated GO terms and retrieve associative relations across the three GO hierarchies to emphasise the major pathways involved by a gene cluster. Both annotation methods and associative relations were quantitatively evaluated with a reference set of 7397 genes and a multi-cluster study of 14 clusters. We also validated the biological appropriateness of our hybrid method with the annotation of a single gene (cdc2) and that of a down-regulated cluster of 37 genes identified by a transcriptome study of an in vitro enterocyte differentiation model (CaCo-2 cells). CONCLUSION: The combination of both approaches is more informative than either separate approach: literature mining can enrich an annotation based only on evidence. Text-mining of the literature can also find valuable associated MEDLINE references that confirm the relevance of the annotation. Eventually, GO terms networks can be built with associative relations in order to highlight cooperative and competitive pathways and their connected molecular functions.

Algorithms↗

The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease.

To pursue a systematic approach to the discovery of functional connections among diseases, genetic perturbation, and drug action, we have created the first installment of a reference collection of gene-expression profiles from cultured human cells treated with bioactive small molecules, together with pattern-matching software to mine these data. We demonstrate that this "Connectivity Map" resource can be used to find connections among small molecules sharing a mechanism of action, chemicals and physiological processes, and diseases and drugs. These results indicate the feasibility of the approach and suggest the value of a large-scale community Connectivity Map project.

Alzheimer Disease↗