Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

High-throughput functional affinity purification of mannose binding proteins from Oryza sativa.

We have used affinity chromatography in combination with mass spectrometry to isolate, identify, and assign a preliminary functional annotation to a large number of both known and novel proteins from rice. Rice (Oryza sativa) leaf, root, and seed tissue extracts were fractionated by column affinity chromatography using alpha-D-mannose as the ligand. Bound fractions were eluted and subjected to one-dimensional electrophoresis, followed by high-performance liquid chromatography-tandem mass spectrometric analysis of separated proteins. This multiplexed technology resulted in the isolation and identification of 136 distinct mannose binding proteins from rice. A comparative analysis demonstrates very little overlap of identified proteins between the respective tissues, and confirms the correctly compartmentalized presence of a significant number of proteins from largely tissue-specific biochemical pathways. Over 30% of the identified proteins with a previously annotated function are directly involved in sugar metabolism, including several highly expressed known rice lectins. Direct comparison of the peptide sequences identified in this study to those peptides identified in the most comprehensive survey of the rice proteome to date indicates that our current data represents a significant enrichment of proteins unique to this dataset. Nearly 15% of the identified proteins, identified on the basis of exact peptide matching to sequences in the rice genomic database, represent proteins without a previously known functional annotation, indicating the potential of this combined chromatographic approach to assign a preliminary function to novel proteins in a high-throughput fashion.

Binding, Competitive↗

The Fanconi anemia gene network is conserved from zebrafish to human.

Fanconi anemia (FA) is a complex disease involving nine identified and two unidentified loci that define a network essential for maintaining genomic stability. To test the hypothesis that the FA network is conserved in vertebrate genomes, we cloned and sequenced zebrafish (Danio rerio) cDNAs and/or genomic BAC clones orthologous to all nine cloned FA genes (FANCA, FANCB, FANCC, FANCD1, FANCD2, FANCE, FANCF, FANCG, and FANCL), and identified orthologs in the genome database for the pufferfish Tetraodon nigroviridis. Genomic organization of exons and introns was nearly identical between zebrafish and human for all genes examined. Hydrophobicity plots revealed conservation of FA protein structure. Evolutionarily conserved regions identified functionally important domains, since many amino acid residues mutated in human disease alleles or shown to be critical in targeted mutagenesis studies are identical in zebrafish and human. Comparative genomic analysis demonstrated conserved syntenies for all FA genes. We conclude that the FA gene network has remained intact since the last common ancestor of zebrafish and human lineages. The application of powerful genetic, cellular, and embryological methodologies make zebrafish a useful model for discovering FA gene functions, identifying new genes in the network, and identifying therapeutic compounds.

Amino Acid Sequence↗

Genome-wide bioinformatic and molecular analysis of introns in Saccharomyces cerevisiae.

Introns have typically been discovered in an ad hoc fashion: introns are found as a gene is characterized for other reasons. As complete eukaryotic genome sequences become available, better methods for predicting RNA processing signals in raw sequence will be necessary in order to discover genes and predict their expression. Here we present a catalog of 228 yeast introns, arrived at through a combination of bioinformatic and molecular analysis. Introns annotated in the Saccharomyces Genome Database (SGD) were evaluated, questionable introns were removed after failing a test for splicing in vivo, and known introns absent from the SGD annotation were added. A novel branchpoint sequence, AAUUAAC, was identified within an annotated intron that lacks a six-of-seven match to the highly conserved branchpoint consensus UACUAAC. Analysis of the database corroborates many conclusions about pre-mRNA substrate requirements for splicing derived from experimental studies, but indicates that splicing in yeast may not be as rigidly determined by splice-site conservation as had previously been thought. Using this database and a molecular technique that directly displays the lariat intron products of spliced transcripts (intron display), we suggest that the current set of 228 introns is still not complete, and that additional intron-containing genes remain to be discovered in yeast. The database can be accessed at http://www.cse.ucsc.edu/research/compbi o/yeast_introns.html.

Computational Biology↗

Mitochondrial DNA ligases of Trypanosoma brucei.

The mitochondrial DNA of Trypanosoma brucei, termed kinetoplast DNA or kDNA, consists of thousands of minicircles and a small number of maxicircles catenated into a single network organized as a nucleoprotein disk at the base of the flagellum. Minicircles are replicated free of the network but still contain nicks and gaps after rejoining to the network. Covalent closure of remaining discontinuities in newly replicated minicircles after their rejoining to the network is delayed until all minicircles have been replicated. The DNA ligase involved in this terminal step in minicircle replication has not been identified. A search of kinetoplastid genome databases has identified two putative DNA ligase genes in tandem. These genes (LIG k alpha and LIG k beta) are highly diverged from mitochondrial and nuclear DNA ligase genes of higher eukaryotes. Expression of epitope-tagged versions of these genes shows that both LIG k alpha and LIG k beta are mitochondrial DNA ligases. Epitope-tagged LIG k alpha localizes throughout the kDNA, whereas LIG k beta shows an antipodal localization close to, but not overlapping, that of topoisomerase II, suggesting that these proteins may be contained in distinct structures or protein complexes. Knockdown of the LIG k alpha mRNA by RNA interference led to a cessation of the release of minicircles from the network and resulted in a reduction in size of the kDNA networks and rapid loss of the kDNA from the cell. Closely related pairs of mitochondrial DNA ligase genes were also identified in Leishmania major and Crithidia fasciculata.

Amino Acid Sequence↗

The predicted impact of coding single nucleotide polymorphisms database.

Nonsynonymous single nucleotide polymorphisms (nsSNP) have the potential to affect the structure or function of expressed proteins and are, therefore, likely to represent modifiers of inherited susceptibility. We have classified and catalogued the predicted functionality of nsSNPs in genes relevant to the biology of cancer to facilitate sequence-based association studies. Candidate genes were identified using targeted search terms and pathways to interrogate the Gene Ontology Consortium database, Kyoto Encyclopedia of Genes and Genomes database, Iobion's Interaction Explorer PathwayAssist Program, National Center for Biotechnology Information Entrez Gene database, and CancerGene database. A total of 9,537 validated nsSNPs located within annotated genes were retrieved from National Center for Biotechnology Information dbSNP Build 123. Filtering this list and linking it to 7,080 candidate genes yielded 3,666 validated nsSNPs with minor allele frequencies > or =0.01 in Caucasian populations. The functional effect of nsSNPs in genes with a single mRNA transcript was predicted using three computational tools-Grantham matrix, Polymorphism Phenotyping, and Sorting Intolerant from Tolerant algorithms. The resultant pool of 3,009 fully annotated nsSNPs is accessible from the Predicted Impact of Coding SNPs database at http://www.icr.ac.uk/cancgen/molgen/MolPopGen_PICS_database.htm. Predicted Impact of Coding SNPs is an ongoing project that will continue to curate and release data on the putative functionality of coding SNPs.

Algorithms↗

Effect of Chang'an decoction on ulcerative colitis by regulating T helper 17 cells and regulatory T cellsRab27 in the p53/high mobility group box 1 pathway.

OBJECTIVE: To explore the effect of Chang'an decoction (, CAD) of ameliorating the immune imbalances in ulcerative colitis (UC) by regulating Rab27 in the P53/high mobility group box 1 pathway. METHODS: The functions and important signaling pathways of the Rab27- and UC-related genes were analyzed viathe use of microarray data from the gene expression omnibus database, gene ontology database, Kyoto encyclopedia of genes and genomes database and gene set enrichment analysis. Dextran sulfate sodium salt-induced colitis mouse model was used to verify the bioinformatics results. Colon length, body weight, and disease activity index were measured. Hematoxylin and eosin staining was applied to validate the histopathology. Tight junction proteins were detected by immunohistochemistry. The proportions of T helper 17 cells (Th17) and regulatory T cells (Treg) in mesenteric lymph nodes were measured viaflow cytometry. Proinflammatory cytokines like interleukin (IL) 17 (IL-17), IL-21 and IL-22 and anti-inflammatory cytokines like transforming growth factor β and IL-10 in the serum and colon of mice were detected by enzyme-linked immunosorbent assay and quantitative real-time polymerase chain reaction, respectively. The expression levels of high mobility group box 1 (HMGB1), P53 and phospho- P53 (P-P53) in colonic tissues were detected by immunofluorescence and Western blotting. RESULTS: Bioinformatics analysis revealed that compared with normal tissues, the expression of Rab27 was significantly increased in UC tissues. Receiver operating characteristic curve showed that Rab27 has the potential to be used as a biomarker for the diagnosis of disease activity. Enrichment analysis showed that UC and Rab27 were mainly associated with small molecule transport, nutrient metabolism, transmembrane transport and the downstream pathway of P53. According to animal experiments, the expression of Rab27 was increased in UC tissues, which aggravated the colonic pathological damage, activated the expression of HMGB1, and also leaded to the imbalance of Th17 and Treg cells. After CAD intervention, Rab27 overexpression, weight loss, colon shortening, and pathological damage were substantial reduced, the expression of tight junction proteins, zona occludens 1 and Occludin were increased. The effect of CAD at high-dose was more obvious. In addition, CAD upgraded the number of Treg cells and the production of TGF-β and IL-10, while decreasing the number of Th17 cells and the expression of inflammatory cytokines (IL-17, IL-21, and IL-22). Moreover, colon inflammation was alleviated by CAD, as indicated by the regulation of HMGB1 and P-P53 expression. CONCLUSION: The expression of Rab27, HMGB1 and P-P53 could be decreased by CAD, and the balance of Th17 and Treg cells as well as their related cytokines could be regulated by CAD.

Animals↗

[In silico analysis of the restriction fragments length distribution in the human genome].

The Restriction On Computer (ROC) program (freely available at http://www.mcb.harvard.edu/gilbert/ROC) was developed and used to analyze the restriction fragment length distribution in the human genome. In contrast to other programs searching for restriction sites, ROC simultaneously analyzes several long nucleotide sequences, such as the entire genomes, and in essence simulates electrophoretic analysis of DNA restriction fragments. In addition, this program extracts and analyzes DNA repeats that account for peaks in the restriction fragment length distribution. The ROC analysis data are consistent with the experimental data obtained via in vitro restriction enzyme analysis (taxonomic printing). A difference between the in vitro and in silico results is explained by underrepresentation of tandem DNA repeats in genomic databases. The ROC analysis of individual genome fragments elucidated the nature of several DNA markers, which were earlier revealed by taxonomic printing, and showed that L1 and Alu repeats are nonrandomly distributed in various chromosomes. Another advantage is that the ROC procedure makes it possible to analyze the nonrandom character of a genomic distribution of short DNA sequences. The ROC analysis showed that a low poly(G) frequency is characteristic of the entire human genome, rather than of only coding sequences. The method was proposed for a more complex in silico analysis of the genome. For instance, it is possible to simulate DNA restriction together with blot hybridization and then to analyze the nature of markers revealed.

Base Sequence↗

IMGT, the international ImMunoGeneTics information system.

The international ImMunoGeneTics information system (IMGT) (http://imgt.cines.fr), created in 1989, by the Laboratoire d'ImmunoGenetique Moleculaire LIGM (Universite Montpellier II and CNRS) at Montpellier, France, is a high-quality integrated knowledge resource specializing in the immunoglobulins (IGs), T cell receptors (TRs), major histocompatibility complex (MHC) of human and other vertebrates, and related proteins of the immune systems (RPI) that belong to the immunoglobulin superfamily (IgSF) and to the MHC superfamily (MhcSF). IMGT includes several sequence databases (IMGT/LIGM-DB, IMGT/PRIMER-DB, IMGT/PROTEIN-DB and IMGT/MHC-DB), one genome database (IMGT/GENE-DB) and one three-dimensional (3D) structure database (IMGT/3Dstructure-DB), Web resources comprising 8000 HTML pages (IMGT Marie-Paule page), and interactive tools. IMGT data are expertly annotated according to the rules of the IMGT Scientific chart, based on the IMGT-ONTOLOGY concepts. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in normal physiological and pathological situations. IMGT is used in medical research (autoimmune diseases, infectious diseases, AIDS, leukemias, lymphomas, myelomas), veterinary research, biotechnology related to antibody engineering (phage displays, combinatorial libraries, chimeric, humanized and human antibodies), diagnostics (clonalities, detection and follow up of residual diseases) and therapeutical approaches (graft, immunotherapy and vaccinology). IMGT is freely available at http://imgt.cines.fr.

Animals↗

Identification of functional, endogenous programmed -1 ribosomal frameshift signals in the genome of Saccharomyces cerevisiae.

In viruses, programmed -1 ribosomal frameshifting (-1 PRF) signals direct the translation of alternative proteins from a single mRNA. Given that many basic regulatory mechanisms were first discovered in viral systems, the current study endeavored to: (i) identify -1 PRF signals in genomic databases, (ii) apply the protocol to the yeast genome and (iii) test selected candidates at the bench. Computational analyses revealed the presence of 10 340 consensus -1 PRF signals in the yeast genome. Of the 6353 yeast ORFs, 1275 contain at least one strong and statistically significant -1 PRF signal. Eight out of nine selected sequences promoted efficient levels of PRF in vivo. These findings provide a robust platform for high throughput computational and laboratory studies and demonstrate that functional -1 PRF signals are widespread in the genome of Saccharomyces cerevisiae. The data generated by this study have been deposited into a publicly available database called the PRFdb. The presence of stable mRNA pseudoknot structures in these -1 PRF signals, and the observation that the predicted outcomes of nearly all of these genomic frameshift signals would direct ribosomes to premature termination codons, suggest two possible mRNA destabilization pathways through which -1 PRF signals could post-transcriptionally regulate mRNA abundance.

Base Sequence↗

Genome-based predictions of metabolic preferences and substrate phenotypes in psychrotrophic bacteria from permafrost environments.

Genomes reveal vast functional potential, but harbor genomic noise that obscures prediction of metabolic and environmental preferences. Genomic databases are skewed towards clinically relevant and easily cultivated bacteria, limiting predictions for diverse and underrepresented environmental taxa. Psychrotrophic bacteria, which can survive and grow in cold, nutrient-limited, dry, and saline environments, are especially underrepresented despite their relevance for understanding microbial responses to changing cold environments and potential biotechnological value given growth at low temperatures. Assembling complete genomes of 48 isolates from Alaskan permafrost, seasonally frozen active layer soils, and terrestrial ice, we used Kyoto Encyclopedia of Genes and Genomes (KEGG) ortholog annotations to evaluate the predictability of metabolic resource-use traits observed using phenotypic tests. Genome-predicted values for glycolytic versus gluconeogenic catabolic preference index, or sugar-acid preference (SAP), explained over 50% of the variance in empirically observed SAP. SAP was inversely correlated to genomic GC content, which follows phylum-level trends, indicating that coarse metabolic preference covaries with phylogeny. Regularized elastic net models offered a more granular view, linking KEGG genes to specific substrate utilization and sensitivity phenotypes and yielding moderate but reproducible accuracy (AUC 0.70-0.79) for 11 substrates, demonstrating that specific substrate responses may be predictable from relatively small subsets of KO genes. These results extend recent advances, such as the SAP metric, and highlight associations among genomic GC content, phylum, and broad metabolic strategy. Linking genomic content to phenotype using isolates is a necessary step toward predictive models of microbial function in environmental communities, and this work can be used for hypothesis generation, with applications towards more expansive data sets.IMPORTANCECold region soils and ice host psychrotrophic bacteria with metabolic traits and adaptations that enable persistence in harsh, resource-limited environments. However, these taxa are underrepresented in genomic reference databases dominated by well-studied, mesophilic organisms. This gap limits inference of ecological strategies and our ability to predict how these microbes may influence the large, thaw-vulnerable carbon reservoirs in permafrost. Here, we show that genomic GC content is associated with the sugar-versus-acid catabolic preference (SAP) of isolates across major phyla, suggesting that broad genomic features may provide a coarse signal of metabolic strategy. We demonstrate that a modified SAP metric, using binary (positive/negative) substrate utilization rather than detailed growth rate measurements, is moderately predictive, thus extending its application to slow-growing or difficult-to-culture taxa. Together, these advances broaden the toolkit for linking genome content to resource-use traits (phenotype) in poorly characterized, cold-adapted bacteria and offer a tractable entry point to broad prediction and hypothesis generation.

Genome, Bacterial↗

RatMap--rat genome tools and data.

The rat genome database RatMap (http://ratmap.org or http://ratmap.gen.gu.se) has been one of the main resources for rat genome information since 1994. The database is maintained by CMB-Genetics at Goteborg University in Sweden and provides information on rat genes, polymorphic rat DNA-markers and rat quantitative trait loci (QTLs), all curated at RatMap. The database is under the supervision of the Rat Gene and Nomenclature Committee (RGNC); thus much attention is paid to rat gene nomenclature. RatMap presents information on rat idiograms, karyotypes and provides a unified presentation of the rat genome sequence and integrated rat linkage maps. A set of tools is also available to facilitate the identification and characterization of rat QTLs, as well as the estimation of exon/intron number and sizes in individual rat genes. Furthermore, comparative gene maps of rat in regard to mouse and human are provided.

Animals↗

The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.

For most proteins in the genome databases, function is predicted via sequence comparison. In spite of the popularity of this approach, the extent to which it can be reliably applied is unknown. We address this issue by systematically investigating the relationship between protein function and structure. We focus initially on enzymes functionally classified by the Enzyme Commission (EC) and relate these to by structurally classified domains the SCOP database. We find that the major SCOP fold classes have different propensities to carry out certain broad categories of functions. For instance, alpha/beta folds are disproportionately associated with enzymes, especially transferases and hydrolases, and all-alpha and small folds with non-enzymes, while alpha+beta folds have an equal tendency either way. These observations for the database overall are largely true for specific genomes. We focus, in particular, on yeast, analyzing it with many classifications in addition to SCOP and EC (i.e. COGs, CATH, MIPS), and find clear tendencies for fold-function association, across a broad spectrum of functions. Analysis with the COGs scheme also suggests that the functions of the most ancient proteins are more evenly distributed among different structural classes than those of more modern ones. For the database overall, we identify the most versatile functions, i.e. those that are associated with the most folds, and the most versatile folds, associated with the most functions. The two most versatile enzymatic functions (hydro-lyases and O-glycosyl glucosidases) are associated with seven folds each. The five most versatile folds (TIM-barrel, Rossmann, ferredoxin, alpha-beta hydrolase, and P-loop NTP hydrolase) are all mixed alpha-beta structures. They stand out as generic scaffolds, accommodating from six to as many as 16 functions (for the exceptional TIM-barrel). At the conclusion of our analysis we are able to construct a graph giving the chance that a functional annotation can be reliably transferred at different degrees of sequence and structural similarity. Supplemental information is available from http://bioinfo.mbb.yale.edu/genome/foldfunc++ +.

Enzymes↗

MagnaportheDB: a federated solution for integrating physical and genetic map data with BAC end derived sequences for the rice blast fungus Magnaporthe grisea.

We have created a federated database for genome studies of Magnaporthe grisea, the causal agent of rice blast disease, by integrating end sequence data from BAC clones, genetic marker data and BAC contig assembly data. A library of 9216 BAC clones providing >25-fold coverage of the entire genome was end sequenced and fingerprinted by HindIII digestion. The Image/FPC software package was then used to generate an assembly of 188 contigs covering >95% of the genome. The database contains the results of this assembly integrated with hybridization data of genetic markers to the BAC library. AceDB was used for the core database engine and a MySQL relational database, populated with numerical representations of BAC clones within FPC contigs, was used to create appropriately scaled images. The database is being used to facilitate sequencing efforts. The database also allows researchers mapping known genes or other sequences of interest, rapid and easy access to the fundamental organization of the M.grisea genome. This database, MagnaportheDB, can be accessed on the web at http://www.cals.ncsu.edu/fungal_genomics/mgdatabase/int.htm.

Base Sequence↗

The PEPR GeneChip data warehouse, and implementation of a dynamic time series query tool (SGQT) with graphical interface.

Publicly accessible DNA databases (genome browsers) are rapidly accelerating post-genomic research (see http://www.genome.ucsc.edu/), with integrated genomic DNA, gene structure, EST/ splicing and cross-species ortholog data. DNA databases have relatively low dimensionality; the genome is a linear code that anchors all associated data. In contrast, RNA expression and protein databases need to be able to handle very high dimensional data, with time, tissue, cell type and genes, as interrelated variables. The high dimensionality of microarray expression profile data, and the lack of a standard experimental platform have complicated the development of web-accessible databases and analytical tools. We have designed and implemented a public resource of expression profile data containing 1024 human, mouse and rat Affymetrix GeneChip expression profiles, generated in the same laboratory, and subject to the same quality and procedural controls (Public Expression Profiling Resource; PEPR). Our Oracle-based PEPR data warehouse includes a novel time series query analysis tool (SGQT), enabling dynamic generation of graphs and spreadsheets showing the action of any transcript of interest over time. In this report, we demonstrate the utility of this tool using a 27 time point, in vivo muscle regeneration series. This data warehouse and associated analysis tools provides access to multidimensional microarray data through web-based interfaces, both for download of all types of raw data for independent analysis, and also for straightforward gene-based queries. Planned implementations of PEPR will include web-based remote entry of projects adhering to quality control and standard operating procedure (QC/SOP) criteria, and automated output of alternative probe set algorithms for each project (see http://microarray.cnmcresearch.org/pgadatatable.asp).

Algorithms↗

Sequence analysis by additive scales: DNA structure for sequences and repeats of all lengths.

MOTIVATION: DNA structure plays an important role in a variety of biological processes. Different di- and tri-nucleotide scales have been proposed to capture various aspects of DNA structure including base stacking energy, propeller twist angle, protein deformability, bendability, and position preference. Yet, a general framework for the computational analysis and prediction of DNA structure is still lacking. Such a framework should in particular address the following issues: (1) construction of sequences with extremal properties; (2) quantitative evaluation of sequences with respect to a given genomic background; (3) automatic extraction of extremal sequences and profiles from genomic databases; (4) distribution and asymptotic behavior as the length N of the sequences increases; and (5) complete analysis of correlations between scales. RESULTS: We develop a general framework for sequence analysis based on additive scales, structural or other, that addresses all these issues. We show how to construct extremal sequences and calibrate scores for automatic genomic and database extraction. We show that distributions rapidly converge to normality as Nincreases. Pairwise correlations between scales depend both on background distribution and sequence length and rapidly converge to an analytically predictable asymptotic value. For di- and tri-nucleotide scales, normal behavior and asymptotic correlation values are attained over a characteristic window length of about 10-15 bp. With a uniform background distribution, pairwise correlations between empirically-derived scales remain relatively small and roughly constant at all lengths, except for propeller twist and protein deformability which are positively correlated. There is a positive (resp. negative) correlation between dinucleotide base stacking (resp. propeller twist and protein deformability) and AT-content that increases in magnitude with length. The framework is applied to the analysis of various DNA tandem repeats. We derive exact expressions for counting the number of repeat unit classes at all lengths. Tandem repeats are likely to result from a variety of different mechanisms, a fraction of which is likely to depend on profiles characterized by extreme structural features.

Animals↗

Multigenic families and proteomics: extended protein characterization as a tool for paralog gene identification.

In classical proteomic studies, the searches in protein databases lead mostly to the identification of protein functions by homology due to the non-exhaustiveness of the protein databases. The quality of the identification depends on the studied organism, its complexity and its representation in the protein databases. Nevertheless, this basic function identification is insufficient for certain applications namely for the development of RNA-based gene-silencing strategies, commonly termed RNA interference (RNAi) in animals and post-transcriptional gene silencing (PTGS) in plants, that require an unambiguous identification of the targeted gene sequence. A PTGS strategy was considered in the study of the infection of Oryza sativa by the Rice Yellow Mottle Virus (RYMV). It is suspected that the RYMV recruits host proteins after its entry into plant cells to form a complex facilitating virus multiplication and spreading. The protein partners of this complex were identified by a classical proteomic approach, nano liquid chromatography tandem mass spectrometry. Among the identified proteins, several were retained for a PTGS strategy. Nevertheless most of the protein candidates appear to be members of multigenic families for which all paralog genes are not present in protein databases. Thus the identification of the real expressed paralog gene with classical protein database searches is impossible. Consequently, as the genome contains all genes and thus all paralog genes, a whole genome search strategy was developed to determine the specific expressed paralog gene. With this approach, the identification of peptides matching only a single gene, called discriminant peptides, allows definitive proof of the expression of this identified gene. This strategy has several requirements: (i) a genome completely sequenced and accessible; (ii) high protein sequence coverage. In the present work, through three examples, we report and validate for the first time a genome database search strategy to specifically identify paralog genes belonging to multigenic families expressed under specific conditions.

Chaperonin 60↗

Update of the Human MitBASE database.

Human MitBASE is a database collecting human mtDNA variants. This database is part of a greater mitochondrial genome database (MitBASE) funded within the EU Biotech Program. The present paper reports the recent improvements in data structure, data quality and data quantity. As far as the database structure is concerned it is now fully designed and implemented. Based on the previously described structure some changes have been made to optimise both data input and data quality. Cross-references with other bio-databases (EMBL, OMIM, MEDLINE) have been implemented. Human MitBASE data can be queried with the MitBASE Simple Query System (http://www.ebi.ac.uk/htbin/Mitbase/mit base.pl) and with SRS at the EBI under the 'Mutation' section (http://srs.ebi.ac.uk/srs5/). At present the HumanMitBASE node contains approximately 5000 variants related to studies investigating population polymorphisms and pathologies.

Animals↗

Characterization of blood pressure and morphological traits in cardiovascular-related organs in 13 different inbred mouse strains.

To better understand the contributions of various genetic backgrounds to complex quantitative phenotypes, we have measured several quantitative traits of cardiovascular interest [i.e., systolic blood pressure, weight (corrected by body weight) of several cardiac compartments and adrenals and kidneys, and histological correlates for kidneys and adrenals] in male and female mice from 13 different inbred strains. We selected strains so that each major genealogical group would be represented and to conform to priorities set by the Mouse Phenome Database project. Interstrain comparisons of phenotypes made it possible to identify strains that displayed values that belonged to either the low or the high end of the interstrain variance for quantitative traits, such as systolic blood pressure, body weight, left ventricular weight, and/or adrenocortical structure. For instance, both male and female C3H/HeJ and A/J mice displayed either low systolic blood pressure or low cardiac ventricular mass, respectively, and male C57BL6/J displayed low adrenal weight. Likewise, intersex comparisons made it possible to identify phenotypic values that were sexually dimorphic for some of the same traits. For instance, female AKR/J mice had relatively higher body weight and systolic blood pressure values than their male counterparts, perhaps constituting an animal model of the metabolic X syndrome. These strain- and sex-specific features will be of value both for future genetic and/or developmental studies and for the development of new animal models that will help in the generation of mechanistic hypotheses. All data have been deposited to the Mouse Phenome Database for future integration with the Mouse Genome Database and can be further analyzed and compared with tools available on the site.

Adrenal Glands↗