Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Use of proteomics methodology to evaluate inflammatory protein expression in tendinitis.

In previous studies we established a rat model of acute tendinitis including functional and mechanical measures of healing. Achilles' tendinitis was induced by injection of collagenase, an enzyme that produces localized fiber digestion and edema formation. As quantitative measures of tissue inflammation, hypercellularity and edema were evaluated in injured tendons in comparison with controls. Using the rat tendinitis model, we have applied isotope-coded affinity tag analysis (ICAT) methodology to indicate localized tendon healing by quantitating protein expression. This novel proteomics method allows detection of subtle differences in protein levels that provide a detailed picture of tendinitis healing. The method involves a new class of chemical linkers used to differentially label cysteine residues from similar peptides in control and treated protein samples with heavy (deuterium off of backbone) and light (hydrogen off of backbone) ICAT reagents that are otherwise chemically identical. Proteins were extracted under liquid nitrogen from control untreated or injured Achilles' tendons 72 hours after collagenase-injection. These proteins were digested with endoproteinase Glu-C and trypsin and the resulting peptide mixtures were evaluated using reverse-phase C18 HPLC and Tristricine SDS-polyacrylamide gel electrophoresis. The two ICAT-modified peptide populations were mixed, affinity-purified and analyzed using microcapillary liquid chromatography and electrospray ionization tandem mass-spectroscopy. The process resulted in relative abundance and charge-to-mass ratio data used in conjunction with database searching to identify proteins expressed differentially in the two treatment groups. By analyzing different time periods in the healing process, an accurate model of the healing rat tendon can be made.

Achilles Tendon↗

The cobZ gene of Methanosarcina mazei Go1 encodes the nonorthologous replacement of the alpha-ribazole-5'-phosphate phosphatase (CobC) enzyme of Salmonella enterica.

Open reading frame (ORF) Mm2058 of the methanogenic archaeon Methanosarcina mazei strain Gö1 was shown in vivo and in vitro to encode the nonorthologous replacement of the alpha-ribazole-phosphate phosphatase (CobC; EC 3.1.3.73) enzyme of Salmonella enterica serovar Typhimurium LT2. Bioinformatics analysis of sequences available in databases tentatively identified ORF Mm2058, which was cloned under the control of an inducible promoter and was used to support growth of an S. enterica strain under conditions that demanded CobC-like activity. The Mm2058 protein was expressed with a decahistidine tag at its N terminus and was purified to homogeneity using nickel affinity chromatography. High-performance liquid chromatography followed by electrospray ionization mass spectrometry showed that the Mm2058 protein had phosphatase activity that converted alpha-ribazole-5'-phosphate to alpha-ribazole, as reported for the bacterial CobC enzyme. On the basis of the data reported here, we refer to ORF Mm2058 as cobZ. We tested the prediction by Rodionov et al. (D. A. Rodionov, A. G. Vitreschak, A. A. Mironov, and M. S. Gelfand, J. Biol. Chem. 278:41148-41159, 2003) that ORF HSL01294 (also called Vng1577) encoded the nonorthologous replacement of the bacterial CobC enzyme in the extremely halophilic archaeon Halobacterium sp. strain NRC-1. A strain of the latter carrying an in-frame deletion of ORF Vng1577 was not a cobalamin auxotroph, suggesting that either there is redundancy of this function in Halobacterium or the gene was misannotated.

Bacterial Proteins↗

Nutritional status of endurance athletes: what is the available information?

Nutritional status is a critical determinant of athletic performance. We question whether currently available studies can give adequate information on nutritional status of endurance athletes. This paper is a critical review of articles published from 1989 to 2003 that investigate nutritional status of endurance athletes. The terms, "nutrition", "diet", or "nutrient", were combined with "endurance athletes" to perform Medline and Pubmed electronic database searches. Two inclusion criteria were considered: (a) study subjects should be adults and (b) articles should report gender-specific values for total energy expenditure and intake of energy, macro and micronutrient from food. Only seven studies fulfilled inclusion criteria. In general, the conclusions of these studies are that endurance athletes have negative energy balance, low intake of carbohydrate, adequate to high intake of protein, and high intake of fat. A critical discussion of the articles' data on vitamins, minerals and trace elements adequacy is conducted using insights and methodology proposed by the newly published assessment and interpretation of Dietary Reference Intakes (DRIs). The studies evaluated give an inappropriate evaluation of the prevalence of adequacy/inadequacy of micronutrient intake among endurance athletes. In this work we indicate potential limitations of existing nutritional data, which reflects the misconceptions found in published literature on nutritional group evaluation. This review stresses the need for a comprehensive and well-conducted nutrition assessment planning to fulfill the existing gap in reliable information about micronutrient adequacy of endurance athletes.

Adult↗

Nutritional supplementation for hip fracture aftercare in older people.

BACKGROUND: Older people with hip fractures are often malnourished at the time of fracture, and have poor food intake subsequently. OBJECTIVES: To review the effects of nutritional interventions in older people recovering from hip fracture. SEARCH STRATEGY: We searched the Cochrane Bone, Joint and Muscle Trauma Group Specialised Register (December 2005), the Cochrane Central Register of Controlled Trials (The Cochrane Library 2006, Issue 1), MEDLINE, six other databases and reference lists. We contacted investigators and handsearched journals. SELECTION CRITERIA: Randomised and quasi-randomised controlled trials of nutritional interventions for people aged over 65 years with hip fracture. DATA COLLECTION AND ANALYSIS: Both authors independently selected trials, extracted data and assessed trial quality. We sought additional information from trialists, and pooled data for primary outcomes. MAIN RESULTS: Twenty-one randomised trials involving 1727 participants were included. Overall trial quality was poor, specifically regarding allocation concealment, assessor blinding and intention-to-treat analysis, and limited availability of outcome data. Eight trials evaluated oral multinutrient feeds: providing non-protein energy, protein, some vitamins and minerals. Oral feeds had no statistically significant effect on mortality (15/161 versus 17/176; relative risk (RR) 0.89, 95% confidence interval (CI) 0.47 to 1.68) but may reduce 'unfavourable outcome' (combined outcome of mortality and survivors with medical complications) (14/66 versus 26/73; RR 0.52, 95% CI 0.32 to 0.84). Four trials examining nasogastric multinutrient feeding showed no evidence of an effect on mortality (RR 0.99, 95% CI 0.50 to 1.97) but the studies were heterogeneous regarding case mix. Nasogastric feeding was poorly tolerated. There was insufficient information for other outcomes. Increasing protein intake in an oral feed was tested in four trials. There was no evidence for an effect on mortality (RR 1.42, 95% CI 0.85 to 2.37). Protein supplementation may have reduced the number of long term medical complications. Two trials, testing intravenous vitamin B1 and other water soluble vitamins, or 1-alpha-hydroxycholecalciferol (an active form of vitamin D) respectively, produced no evidence of effect for either supplement. One trial, evaluating dietetic assistants to help with feeding, showed a trend for a reduction in mortality (RR 0.57, 99% CI 0.29 to 1.11). AUTHORS' CONCLUSIONS: Some evidence exists for the effectiveness of oral protein and energy feeds, but overall the evidence for the effectiveness of nutritional supplementation remains weak. Adequately sized trials are required which overcome the methodological defects of the reviewed studies. In particular, the role of dietetic assistants requires further evaluation.

Aftercare↗

alpha-Helix region prediction with stochastic rule learning.

We propose a new method, based on the theory of stochastic rule learning, for predicting alpha-helix regions in a given protein sequence. Our method (hereafter referred to as the SR method) produces stochastic rules, each of which assigns, to any region in an amino acid sequence, the probability that it is an alpha-helix region. When learning a stochastic rule from a particular alpha-helix region, our method makes use of positive training examples obtained from a number of regions that are homologous to that region. Each stochastic rule is optimized using the minimum description length (MDL) principle, and such optimized stochastic rules are used to predict alpha-helix regions of any given protein sequence. In our experiments, using 25 proteins selected from the HSSP database as training examples, we applied the SR method to the problem of predicting alpha-helix regions in test examples, which consisted of > 5000 residues with 38% alpha-helix content. Each of these test examples possesses < 25% homology to any proteins in the training and other test examples. Our method achieved 81% average prediction accuracy for the test examples; this compares favorably to Qian and Sejnowski's method, which attains no more than 75% average accuracy, and further which compares to Rost and Sander's method which has proven to be one of the best secondary structure prediction methods.

Amino Acid Sequence↗

Clustering protein sequences--structure prediction by transitive homology.

MOTIVATION: It is widely believed that for two proteins Aand Ba sequence identity above some threshold implies structural similarity due to a common evolutionary ancestor. Since this is only a sufficient, but not a necessary condition for structural similarity, the question remains what other criteria can be used to identify remote homologues. Transitivity refers to the concept of deducing a structural similarity between proteins A and C from the existence of a third protein B, such that A and B as well as B and C are homologues, as ascertained if the sequence identity between A and B as well as that between B and C is above the aforementioned threshold. It is not fully understood if transitivity always holds and whether transitivity can be extended ad infinitum. RESULTS: We developed a graph-based clustering approach, where transitivity plays a crucial role. We determined all pair-wise similarities for the sequences in the SwissProt database using the Smith-Waterman local alignment algorithm. This data was transformed into a directed graph, where protein sequences constitute vertices. A directed edge was drawn from vertex A to vertex B if the sequences A and B showed similarity, scaled with respect to the self-similarity of A, above a fixed threshold. Transitivity was important in the clustering process, as intermediate sequences were used, limited though by the requirement of having directed paths in both directions between proteins linked over such sequences. The length dependency-implied by the self-similarity-of the scaling of the alignment scores appears to be an effective criterion to avoid clustering errors due to multi-domain proteins. To deal with the resulting large graphs we have developed an efficient library. Methods include the novel graph-based clustering algorithm capable of handling multi-domain proteins and cluster comparison algorithms. Structural Classification of Proteins (SCOP) was used as an evaluation data set for our method, yielding a 24% improvement over pair-wise comparisons in terms of detecting remote homologues. AVAILABILITY: The software is available to academic users on request from the authors. CONTACT: e.bolten@science-factory.com; schliep@zpr.uni-koeln.de; s.schneckener@science-factory.com; d.schomburg@uni-koeln.de; schrader@zpr.uni-koeln.de. SUPPLEMENTARY INFORMATION: http://www.zaik.uni-koeln.de/~schliep/ProtClust.html.

Algorithms↗

Numerical classification of coding sequences.

DNA sequences coding for protein may be represented by counts of nucleotides or codons. A complete reading frame may be abbreviated by its base count, e.g. A76C158G121T74, or with the corresponding codon table, e.g. (AAA)0(AAC)1(AAG)9 ... (TTT)0. We propose that these numerical designations be used to augment current methods of sequence annotation. Because base counts and codon tables do not require revision as knowledge of function evolves, they are well-suited to act as cross-references, for example to identify redundant GenBank entries. These descriptors may be compared, in place of DNA sequences, to extract homologous genes from large databases. This approach permits rapid searching with good selectivity.

Animals↗

Re-annotation of genome microbial coding-sequences: finding new genes and inaccurately annotated genes.

BACKGROUND: Analysis of any newly sequenced bacterial genome starts with the identification of protein-coding genes. Despite the accumulation of multiple complete genome sequences, which provide useful comparisons with close relatives among other organisms during the annotation process, accurate gene prediction remains quite difficult. A major reason for this situation is that genes are tightly packed in prokaryotes, resulting in frequent overlap. Thus, detection of translation initiation sites and/or selection of the correct coding regions remain difficult unless appropriate biological knowledge (about the structure of a gene) is imbedded in the approach. RESULTS: We have developed a new program that automatically identifies biologically significant candidate genes in a bacterial genome. Twenty-six complete prokaryotic genomes were analyzed using this tool, and the accuracy of gene finding was assessed by comparison with existing annotations. This analysis revealed that, despite the enormous effort of genome program annotators, a small but not negligible number of genes annotated within the framework of sequencing projects are likely to be partially inaccurate or plainly wrong. Moreover, the analysis of several putative new genes shows that, as expected, many short genes have escaped annotation. In most cases, these new genes revealed frameshifts that could be either artifacts or genuine frameshifts. Some entirely unexpected new genes have also been identified. This allowed us to get a more complete picture of prokaryotic genomes. The results of this procedure are progressively integrated into the SWISS-PROT reference databank. CONCLUSIONS: The results described in the present study show that our procedure is very satisfactory in terms of gene finding accuracy. Except in few cases, discrepancies between our results and annotations provided by individual authors can be accounted for by the nature of each annotation process or by specific characteristics of some genomes. This stresses that close cooperation between scientists, regular update and curation of the findings in databases are clearly required to reduce the level of errors in genome annotation (and hence in reducing the unfortunate spreading of errors through centralized data libraries).

Computational Biology↗

Identification and characterization of a new family of guanine nucleotide exchange factors for the ras-related GTPase Ral.

Guanine nucleotide exchange factors (GEFs) are responsible for coupling cell surface receptors to Ras protein activation. Here we describe the characterization of a novel family of differentially expressed GEFs, identified by database sequence homology searching. These molecules share the core catalytic domain of other Ras family GEFs but lack the catalytic non-conserved (conserved non-catalytic/Ras exchange motif/structurally conserved region 0) domain that is believed to contribute to Sos1 integrity. In vitro binding and in vivo nucleotide exchange assays indicate that these GEFs specifically catalyze the GTP loading of the Ral GTPase when overexpressed in 293T cells. A central proline-rich motif associated with the Src homology (SH)2/SH3-containing adapter proteins Grb2 and Nck in vivo, whereas a pleckstrin homology (PH) domain was located at the GEF C terminus. We refer to these GEFs as RalGPS 1A, 1B, and 2 (Ral GEFs with PH domain and SH3 binding motif). The PH domain was required for in vivo GEF activity and could be functionally replaced by the Ki-Ras C terminus, suggesting a role in membrane targeting. In the absence of the PH domain RalGPS 1B cooperated with Grb2 to promote Ral activation, indicating that SH3 domain interaction also contributes to RalGPS regulation. In contrast to the Ral guanine nucleotide dissociation stimulator family of Ral GEFs, the RalGPS proteins do not possess a Ras-GTP-binding domain, suggesting that they are activated in a Ras-independent manner.

Amino Acid Sequence↗

The Nordic Reference Interval Project 2000: recommended reference intervals for 25 common biochemical properties.

Each of 102 Nordic routine clinical biochemistry laboratories collected blood samples from at least 25 healthy reference individuals evenly distributed for gender and age, and analysed 25 of the most commonly requested serum/plasma components from each reference individual. A reference material (control) consisting of a fresh frozen liquid pool of serum with values traceable to reference methods (used as the project "calibrator" for non-enzymes to correct reference values) was analysed together with other serum pool controls in the same series as the project samples. Analytical data, method data and data describing the reference individuals were submitted to a central database for evaluation and calculation of reference intervals intended for common use in the Nordic countries. In parallel to the main project, measurements of commonly requested haematology properties on EDTA samples were also carried out, mainly by laboratories in Finland and Sweden. Aliquots from reference samples were submitted to storage in a central bio-bank for future establishment of reference intervals for other properties. The 25 components were, in alphabetical order: alanine transaminase, albumin, alkaline phosphatase, amylase, amylase pancreatic, aspartate transaminase, bilirubins, calcium, carbamide, cholesterol, creatine kinase, creatininium, gamma-glutamyltransferase, glucose, HDL-cholesterol, iron, iron binding capacity, lactate dehydrogenase, magnesium, phosphate, potassium, protein, sodium, triglyceride and urate.

Biomarkers↗

Enhanced functional and structural domain assignments using remote similarity detection procedures for proteins encoded in the genome of Mycobacterium tuberculosis H37Rv.

The sequencing of the Mycobacterium tuberculosis (MTB) H37Rv genome has facilitated deeper insights into the biology of MTB, yet the functions of many MTB proteins are unknown. We have used sensitive profile-based search procedures to assign functional and structural domains to infer functions of gene products encoded in MTB. These domain assignments have been made using a compendium of sequence and structural domain families. Functions are predicted for 78 % of the encoded gene products. For 69 % of these, functions can be inferred by domain assignments. The functions for the rest are deduced from their homology to proteins of known function. Superfamily relationships between families of unknown and known structures have increased structural information by approximately 11%. Remote similarity detection methods have enabled domain assignments for 1325 'hypothetical proteins'. The most populated families in MTB are involved in lipid metabolism, entry and survival of the bacillus in host. Interestingly, for 353 proteins, which we refer to as MTB-specific, no homologues have been identified. Numerous, previously unannotated, hypothetical proteins have been assigned domains and some of these could perhaps be the possible chemotherapeutic targets. MTB-specific proteins might include factors responsible for virulence. Importantly, these assignments could be valuable for experimental endeavors. The detailed results are publicly available at http://hodgkin.mbu.iisc.ernet.in/~dots.

Amino Acid Sequence↗

The PAN module: the N-terminal domains of plasminogen and hepatocyte growth factor are homologous with the apple domains of the prekallikrein family and with a novel domain found in numerous nematode proteins.

Based on homology search and structure prediction methods we show that (1) the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, (2) the apple domains of the plasma prekallikrein/coagulation factor XI family, and (3) domains of various nematode proteins belong to the same module superfamily, hereafter referred to as the PAN module. The patterns of conserved residues correspond to secondary structural elements of the known three-dimensional structure of hepatocyte growth factor N domain, therefore we predict a similar fold for all members of this superfamily. Based on available functional informations on apple domains and N domains, it is clear that PAN modules have significant functional versatility, they fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.

Amino Acid Sequence↗

The organization of the keratin I and II gene clusters in placental mammals and marsupials show a striking similarity.

The genomic database for a marsupial, the opossum Monodelphis domestica, is highly advanced. This allowed a complete analysis of the keratin I and keratin II gene cluster with some 30 genes in each cluster as well as a comparison with the human keratin clusters. Human and marsupial keratin gene clusters have an astonishingly similar organization. As placental mammals and marsupials are sister groups a corresponding organization is also expected for the archetype mammal. Since hair is a mammalian acquisition the following features of the cluster refer to its origin. In both clusters hair keratin genes arose at an interior position. While we do not know from which epithelial keratin genes the first hair keratins type-I and -II genes evolved, subsequent gene duplications gave rise to a subdomain of the clusters with many neighboring hair keratin genes. A second subdomain accounts in both clusters for 4 neighboring genes encoding the keratins of the inner root sheath (irs) keratins. Finally the hair keratin gene subdomain in the type-I gene cluster is interrupted after the second gene by a region encoding numerous genes for the high/ultrahigh sulfur hair keratin-associated proteins (KAPs). We also propose a tentative synteny relation of opossum and human genes based on maximal sequence conservation of the encoded keratins. The keratin gene clusters of the opossum seem to lack pseudogenes and display a slightly increased number of genes. Opossum keratin genes are usually longer than their human counterparts and also show longer intergenic distances.

Animals↗

Identification of PEX5p-related novel peroxisome-targeting signal 1 (PTS1)-binding proteins in mammals.

Based on peroxin protein 5 (Pex5p) homology searches in the expressed sequence tag database and sequencing of large full-length cDNA inserts, three novel and related human cDNAs were identified. The brain-derived cDNAs coded for two related proteins that differ only slightly at their N-terminus, and exhibit 39.8% identity to human PEX5p. The shorter liver-derived cDNA coded for the C-terminal tetratricopeptide repeat-containing domain of the brain cDNA-encoded proteins. Since these three proteins specifically bind to various C-terminal peroxisome-targeting signals in a manner indistinguishable from Pex5p and effectively compete with Pex5p in an in vitro peroxisome-targeting signal 1 (PTS1)-binding assay, we refer to them as 'Pex5p-related proteins' (Pex5Rp). In contrast to Pex5p, however, human PEX5Rp did not bind to Pex14p or to the RING finger motif of Pex12p, and could not restore PTS1 protein import in Pex5(-/-) mouse fibroblasts. Immunofluorescence analysis of epitope-tagged PEX5Rp in Chinese hamster ovary cells suggested an exclusively cytosolic localization. Northern-blot analysis showed that the PEX5R gene, which is localized to chromosome 3q26.2--3q27, is expressed preferentially in brain. Mouse PEX5Rp was also delineated. In addition, experimental evidence established that the closest-related yeast homologue, YMR018wp, did not bind PTS1. Based on its subcellular localization and binding properties, Pex5Rp may function as a regulator in an early step of the PTS1 protein import process.

Amino Acid Sequence↗

The Molecular Biology Database Collection: 2005 update.

The Nucleic Acids Research Molecular Biology Database Collection is a public online resource that lists the databases described in this and previous issues of Nucleic Acids Research together with other databases of value to the biologist and available throughout the world. All databases included in this Collection are freely available to the public. The 2005 update includes 719 databases, 171 more than the 2004 one. The databases are organized in a hierarchical classification that simplifies the process of finding the right database for any given task. The growing number of databases related to immunology, plant and organelle research have been accommodated by separating them into three new categories. The database summaries provide brief descriptions of the databases, contact details, appropriate references and acknowledgements. The online summaries also serve as a venue for the maintainers of each database to introduce database updates and other improvements in the scope and tools. These updates are particularly important for those databases that have not been described in print in the recent past. The database list and summaries are available online at the Nucleic Acids Research web site, http://nar.oupjournals.org/.

Allergy and Immunology↗

PHEXdb, a locus-specific database for mutations causing X-linked hypophosphatemia.

X-linked hypophosphatemia (XLH) is a dominant disorder of phosphate (Pi) homeostasis characterized by growth retardation, rachitic and osteomalacic bone disease, hypophosphatemia, and renal defects in Pi reabsorption and vitamin D metabolism. The gene responsible for XLH was identified by positional cloning and designated PHEX (formerly PEX) to depict a Phosphate regulating gene with homology to Endopeptidases on the X chromosome. To date, 131 mutations in the PHEX gene have been reported. We undertook to centralize information on mutations in the PHEX gene by establishing a database search tool, PHEXdb (http://data.mch.mcgill.ca/phexdb). This site is dedicated to the collection and distribution of information on PHEX mutations, and is accessible to the scientific community. PHEXdb provides a submission form to allow the addition of newly identified mutations in the PHEX gene. Users can search the database by mutation, phenotype, and authors who have published or submitted mutations. The PHEXdb home page includes links to information pages, which refer to recent publications on PHEX, XLH, and murine Hyp and Gy homologues, and to other web pages relevant to XLH. This resource will facilitate the identification of PHEX structure-function relationships and phenotype-genotype correlations.

DNA Mutational Analysis↗

BAliBASE (Benchmark Alignment dataBASE): enhancements for repeats, transmembrane sequences and circular permutations.

BAliBASE is specifically designed to serve as an evaluation resource to address all the problems encountered when aligning complete sequences. The database contains high quality, manually constructed multiple sequence alignments together with detailed annotations. The alignments are all based on three-dimensional structural superpositions, with the exception of the transmembrane sequences. The first release provided sets of reference alignments dealing with the problems of high variability, unequal repartition and large N/C-terminal extensions and internal insertions. Here we describe version 2.0 of the database, which incorporates three new reference sets of alignments containing structural repeats, trans-membrane sequences and circular permutations to evaluate the accuracy of detection/prediction and alignment of these complex sequences. BAliBASE can be viewed at the web site http://www-igbmc.u-strasbg. fr/BioInfo/BAliBASE2/index.html or can be downloaded from ftp://ftp-igbmc.u-strasbg.fr/pub/BAliBASE2 /.

Algorithms↗

The computational analysis of scientific literature to define and recognize gene expression clusters.

A limitation of many gene expression analytic approaches is that they do not incorporate comprehensive background knowledge about the genes into the analysis. We present a computational method that leverages the peer-reviewed literature in the automatic analysis of gene expression data sets. Including the literature in the analysis of gene expression data offers an opportunity to incorporate functional information about the genes when defining expression clusters. We have created a method that associates gene expression profiles with known biological functions. Our method has two steps. First, we apply hierarchical clustering to the given gene expression data set. Secondly, we use text from abstracts about genes to (i) resolve hierarchical cluster boundaries to optimize the functional coherence of the clusters and (ii) recognize those clusters that are most functionally coherent. In the case where a gene has not been investigated and therefore lacks primary literature, articles about well-studied homologous genes are added as references. We apply our method to two large gene expression data sets with different properties. The first contains measurements for a subset of well-studied Saccharomyces cerevisiae genes with multiple literature references, and the second contains newly discovered genes in Drosophila melanogaster; many have no literature references at all. In both cases, we are able to rapidly define and identify the biologically relevant gene expression profiles without manual intervention. In both cases, we identified novel clusters that were not noted by the original investigators.

Animals↗