Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

MaizeGDB, the community database for maize genetics and genomics.

The Maize Genetics and Genomics Database (MaizeGDB) is a central repository for maize sequence, stock, phenotype, genotypic and karyotypic variation, and chromosomal mapping data. In addition, MaizeGDB provides contact information for over 2400 maize cooperative researchers, facilitating interactions between members of the rapidly expanding maize community. MaizeGDB represents the synthesis of all data available previously from ZmDB and from MaizeDB-databases that have been superseded by MaizeGDB. MaizeGDB provides web-based tools for ordering maize stocks from several organizations including the Maize Genetics Cooperation Stock Center and the North Central Regional Plant Introduction Station (NCRPIS). Sequence searches yield records displayed with embedded links to facilitate ordering cloned sequences from various groups including the Maize Gene Discovery Project and the Clemson University Genomics Institute. An intuitive web interface is implemented to facilitate navigation between related data, and analytical tools are embedded within data displays. Web-based curation tools for both designated experts and general researchers are currently under development. MaizeGDB can be accessed at http://www.maizegdb.org/.

Computational Biology↗

GPCEG-A database for genomic polymorphism of Chinese ethnic groups.

This paper reports the construction of the database for Genomic Polymorphism of Chinese Ethnic Groups (GPCEG). GPCEG contains denomination and basic information of Chinese 56 ethnic groups, with introduction of their in geographic distribution, population quantity, spoken and written language, religious belief and physical characteristics. GPCEG collects the data of genomic polymorphism, cell lines, reference and links of other international related databases. The visualization, query and update system were also available. GPCEG laid the foundations of establishing a national database with Chinese characteristics.

Cell Line↗

Organization and structure of hox gene loci in medaka genome and comparison with those of pufferfish and zebrafish genomes.

We isolated BAC clones that cover the entire hox gene loci in the medaka fish Oryzias latipes. The BAC clones were characterized by the Southern hybridization with many hox gene probes isolated in our previous study and by PCR using primers designed for selective amplification of respective hox genes. Then, the BAC clones have been subjected to shotgun sequencing. The results revealed the organization of the entire hox gene loci. Forty-six hox genes in total are encoded in seven clusters as follows: 10 hox genes in Aa cluster; 5 in Ab; 9 in Ba; 4 in Bb; 10 in Ca; 6 in Da; and 2 in Db. Together with the information on the hox gene loci registered in the Fugu genome database and in the Danio genome database, the physical maps of three fish genomes were constructed and compared one another. Not only numbers of hox genes but also the distances between the neighboring hox genes are highly similar between medaka and fugu. As for six clusters, Aa, Ab, Ba, Bb, Ca and Da that are commonly present in the three fishes, only few or no differences were found in each cluster. Thus, the hox gene sets should have been well conserved once they had been established in respective species.

Animals↗

The Genome Sequence DataBase (GSDB): improving data quality and data access.

In 1997 the primary focus of the Genome Sequence DataBase (GSDB; www. ncgr.org/gsdb ) located at the National Center for Genome Resources was to improve data quality and accessibility. Efforts to increase the quality of data within the database included two major projects; one to identify and remove all vector contamination from sequences in the database and one to create premier sequence sets (including both alignments and discontiguous sequences). Data accessibility was improved during the course of the last year in several ways. First, a graphical database sequence viewer was made available to researchers. Second, an update process was implemented for the web-based query tool, Maestro. Third, a web-based tool, Excerpt, was developed to retrieve selected regions of any sequence in the database. And lastly, a GSDB flatfile that contains annotation unique to GSDB (e.g., sequence analysis and alignment data) was developed. Additionally, the GSDB web site provides a tool for the detection of matrix attachment regions (MARs), which can be used to identify regions of high coding potential. The ultimate goal of this work is to make GSDB a more useful resource for genomic comparison studies and gene level studies by improving data quality and by providing data access capabilities that are consistent with the needs of both types of studies.

Base Sequence↗

Genomic structure and promoter analysis of PKC-delta.

Protein kinase C-delta (PKC-delta) is a ubiquitously expressed kinase involved in a variety of cellular signaling pathways including cell growth, differentiation, apoptosis, tumor promotion, and carcinogenesis. While signaling pathways downstream of PKC-delta are well studied, the regulation of the gene has not been extensively analyzed. A mouse genomic DNA fragment containing the PKC-delta gene was sequenced by the primer-walking method, and the subsequent DNA sequence data were used as a query to clone Caenorhabditis elegans and human genomic homologs from the publicly available genomic databases. The genomic structures of C. elegans, mouse, rat, and human PKC-delta were analyzed, and the result revealed that PKC-delta genes comprise 12, 18, 19, and 18 exons for C. elegans, mouse, rat, and human, respectively. The translation start methionine resides in the second exon in mouse and human and in the third exon in rat. The first intron between the first exon and the exon with the translation start methionine in mammalian genes represents a very large gap, as long as 17 kb in human, indicating a complexity involved in gene splicing. Overall exon-intron genomic structure is highly conserved among mammals, while significantly diverged in C. elegans. Putative transcription factor binding sites on the 1.7-kb promoter region of the mouse gene suggest that PKC-delta might be involved in spermatogenesis, embryogenesis, development, brain generation, immune response, oxidative environment, and oncogenesis. Studies on the promoter and subsequent biological testing on mouse keratinocytes indicate that tumor necrosis factor (TNF)-alpha increases the expression of PKC-delta, and this correlates with the time of NFkappaB nuclear translocation and activation. This TNF-alpha-mediated upregulation of PKC-delta is repressed in keratinocytes that are preinfected with IkappaB superrepressor adenovirus, suggesting that NFkappaB is involved directly in PKC-delta expression.

Animals↗

Genetics and genomics in infectious disease susceptibility.

The past decade has witnessed a rapid transition from the first positional cloning of an infectious disease susceptibility gene (Slc11a1, also called Nramp1) in the mouse to genome-wide scans in human multicase families and the identification of potential disease-causing genes by simple inspection of the public human genome databases. Pathogen genome projects have facilitated multilocus sequence typing of pathogen isolates and studies of ecological fitness and virulence patterns in disease-causing isolates. Comparative sequence analysis of pathogen strains and functional genomics studies are now underway, hopefully providing new insight into infectious disease susceptibility.

Cloning, Molecular↗

[Modeling of all genome and database].

We have developed the protein modeling software FAMS (Full automatic protein modeling system), and using the FAMS the proteins coded in the all the genes were modeled. And we developed web browsing software. We had participated in the CAFASP2 contest of the CASP4 which is the competition of the protein structure prediction. We won almost best server in the CAFASP2 which is the contest of full automatic protein modeling. Accordingly the database quality made by using the FAMS program will be very good. The FAMS modeling web service is available in http://physchem.pharm.kitasato-u.ac.jp/. FAMSBASE is seen in the web site of http://famsbase.bio.nagoya-u.ac.jp/.

Databases, Genetic↗

Genome cluster database. A sequence family analysis platform for Arabidopsis and rice.

The genome-wide protein sequences from Arabidopsis (Arabidopsis thaliana) and rice (Oryza sativa) spp. japonica were clustered into families using sequence similarity and domain-based clustering. The two fundamentally different methods resulted in separate cluster sets with complementary properties to compensate the limitations for accurate family analysis. Functional names for the identified families were assigned with an efficient computational approach that uses the description of the most common molecular function gene ontology node within each cluster. Subsequently, multiple alignments and phylogenetic trees were calculated for the assembled families. All clustering results and their underlying sequences were organized in the Web-accessible Genome Cluster Database (http://bioinfo.ucr.edu/projects/GCD) with rich interactive and user-friendly sequence family mining tools to facilitate the analysis of any given family of interest for the plant science community. An automated clustering pipeline ensures current information for future updates in the annotations of the two genomes and clustering improvements. The analysis allowed the first systematic identification of family and singlet proteins present in both organisms as well as those restricted to one of them. In addition, the established Web resources for mining these data provide a road map for future studies of the composition and structure of protein families between the two species.

Algorithms↗

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

Amino Acid Sequence↗

Improvements to GALA and dbERGE II: databases featuring genomic sequence alignment, annotation and experimental results.

We describe improvements to two databases that give access to information on genomic sequence similarities, functional elements in DNA and experimental results that demonstrate those functions. GALA, the database of Genome ALignments and Annotations, is now a set of interlinked relational databases for five vertebrate species, human, chimpanzee, mouse, rat and chicken. For each species, GALA records pairwise and multiple sequence alignments, scores derived from those alignments that reflect the likelihood of being under purifying selection or being a regulatory element, and extensive annotations such as genes, gene expression patterns and transcription factor binding sites. The user interface supports simple and complex queries, including operations such as subtraction and intersections as well as clustering and finding elements in proximity to features. dbERGE II, the database of Experimental Results on Gene Expression, contains experimental data from a variety of functional assays. Both databases are now run on the DB2 database management system. Improved hardware and tuning has reduced response times and increased querying capacity, while simplified query interfaces will help direct new users through the querying process. Links are available at http://www.bx.psu.edu/.

Animals↗

In silico discovery of gene-coding variants in murine quantitative trait loci using strain-specific genome sequence databases.

BACKGROUND: The identification of genes underlying complex traits has been aided by quantitative trait locus (QTL) mapping approaches, which in turn have benefited from advances in mammalian genome research. Most recently, whole-genome draft sequences and assemblies have been generated for mouse strains that have been used for a large fraction of QTL mapping studies. Here we show how such strain-specific mouse genome sequence databases can be used as part of a high-throughput pipeline for the in silico discovery of gene-coding variations within murine QTLs. As a test of this approach we focused on two QTLs on mouse chromosomes 1 and 13 that are involved in physical dependence on alcohol. RESULTS: Interstrain alignment of sequences derived from the relevant mouse strain genome sequence databases for 199 QTL-localized genes spanning 210,020 base-pairs of coding sequence identified 21 genes with different coding sequences for the progenitor strains. Several of these genes, including four that exhibit strong phenotypic links to chronic alcohol withdrawal, are promising candidates to underlie these QTLs. CONCLUSIONS: This approach has wide general utility, and should be applicable to any of the several hundred mouse QTLs, encompassing over 60 different complex traits, that have been identified using strains for which relatively complete genome sequences are available.

Alcohol Withdrawal Seizures↗

A statistical basis for testing the significance of mass spectrometric protein identification results.

A method for testing the significance of mass spectrometric (MS) protein identification results is presented. MS proteolytic peptide mapping and genome database searching provide a rapid, sensitive, and potentially accurate means for identifying proteins. Database search algorithms detect the matching between proteolytic peptide masses from an MS peptide map and theoretical proteolytic peptide masses of the proteins in a genome database. The number of masses that matches is used to compute a score, S, for each protein, and the protein that yields the best score is assumed as the identification result. There is a risk of obtaining a false result, because masses determined by MS are not unique; i.e., each mass in a peptide map can match randomly one or several proteins in a genome database. A false result is obtained when the score, S, due to random matching cannot be discerned from the score due to matching with a real protein in the sample. We therefore introduce the frequency function, f(S), for false (random) identification results as a basis for testing at what significance level, alpha, one can reject a null hypothesis, H0: "the result is false". The significance is tested by comparing an experimental score, S(E), with a critical score, S(C), required for a significant result at the level alpha. If S(E) > or = S(C), H0 is rejected. f(S) and S(C) were obtained by simulations utilizing random tryptic peptide maps generated from a genome database. The critical score, S(C), was studied as a function of the number of masses in the peptide map, the mass accuracy, the degree of incomplete enzymatic cleavage, the protein mass range, and the size of the genome. With S(C) known for a variety of experimental constraints, significance testing can be fully automated and integrated with database searching software used for protein identification.

Genome↗

FGDB: a comprehensive fungal genome resource on the plant pathogen Fusarium graminearum.

The MIPS Fusarium graminearum Genome Database (FGDB) is a comprehensive genome database on one of the most devastating fungal plant pathogens of wheat and barley. FGDB provides information on two gene sets independently derived by automated annotation of the F.graminearum genome sequence. A complete manually revised gene set will be completed within the near future. The initial results of systematic manual correction of gene calls are already part of the current gene set. The database can be accessed to retrieve information from bioinformatics analyses and functional classifications of the proteins. The data are also organized in the well established MIPS catalogs and novel query techniques are available to search the data. The comprehensive set of gene calls was also used for the design of an Affymetrix GeneChip. The resource is accessible on http://mips.gsf.de/genre/proj/fusarium/.

Databases, Genetic↗

GALA, a database for genomic sequence alignments and annotations.

We have developed a relational database to contain whole genome sequence alignments between human and mouse with extensive annotations of the human sequence. Complex queries are supported on recorded features, both directly and on proximity among them. Searches can reveal a wide variety of relationships, such as finding all genes expressed in a designated tissue that have a highly conserved noncoding sequence 5' to the start site. Other examples are finding single nucleotide polymorphisms that occur in conserved noncoding regions upstream of genes and identifying CpG islands that overlap the 5' ends of divergently transcribed genes. The database is available online at http://globin.cse.psu.edu/ and http://bio.cse.psu.edu/.

5' Untranslated Regions↗

The Vertebrate Genome Annotation (Vega) database.

The Vertebrate Genome Annotation (Vega) database (http://vega.sanger.ac.uk) has been designed to be a community resource for browsing manual annotation of finished sequences from a variety of vertebrate genomes. Its core database is based on an Ensembl-style schema, extended to incorporate curation-specific metadata. In collaboration with the genome sequencing centres, Vega attempts to present consistent high-quality annotation of the published human chromosome sequences. In addition, it is also possible to view various finished regions from other vertebrates, including mouse and zebrafish. Vega displays only manually annotated gene structures built using transcriptional evidence, which can be examined in the browser. Attempts have been made to standardize the annotation procedure across each vertebrate genome, which should aid comparative analysis of orthologues across the different finished regions.

Animals↗

PoMaMo--a comprehensive database for potato genome data.

A database for potato genome data (PoMaMo, Potato Maps and More) was established. The database contains molecular maps of all twelve potato chromosomes with about 1000 mapped elements, sequence data, putative gene functions, results from BLAST analysis, SNP and InDel information from different diploid and tetraploid potato genotypes, publication references, links to other public databases like GenBank (http://www.ncbi.nlm.nih.gov/) or SGN (Solanaceae Genomics Network, http://www.sgn.cornell.edu/), etc. Flexible search and data visualization interfaces enable easy access to the data via internet (https://gabi.rzpd.de/PoMaMo.html). The Java servlet tool YAMB (Yet Another Map Browser) was designed to interactively display chromosomal maps. Maps can be zoomed in and out, and detailed information about mapped elements can be obtained by clicking on an element of interest. The GreenCards interface allows a text-based data search by marker-, sequence- or genotype name, by sequence accession number, gene function, BLAST Hit or publication reference. The PoMaMo database is a comprehensive database for different potato genome data, and to date the only database containing SNP and InDel data from diploid and tetraploid potato genotypes.

Chromosome Mapping↗

Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database.

Incomplete penetrance, or absence of disease phenotype in an individual with a disease-associated variant, is a major challenge in variant interpretation. Studying individuals with apparent incomplete penetrance can shed light on underlying drivers of altered phenotype penetrance. Here, we investigate clinically relevant variants from ClinVar in 807,162 individuals from the Genome Aggregation Database (gnomAD), demonstrating improved representation in gnomAD version 4. We then conduct a comprehensive case-by-case assessment of 734 predicted loss of function variants in 77 genes associated with severe, early-onset, highly penetrant haploinsufficient disease. Here, we identify explanations for the presumed lack of disease manifestation in 701 of 734 variants (95%). Individuals with unexplained lack of disease manifestation in this set of disorders are rare, underscoring the need and power of deep case-by-case assessment presented here to minimize false assignments of disease risk, particularly in unaffected individuals with higher rates of secondary properties that result in rescue.

Humans↗

A strategy for rapid, high-confidence protein identification.

A procedure is described for rapid, high-confidence identification of proteins using matrix-assisted laser desorption/ionization tandem ion trap mass spectrometry in conjunction with a genome database searching strategy. The procedure involves excision of copper-stained bands or spots from electrophoretic gels, in-gel trypsin digestion of the proteins, single-stage mass spectrometric analysis of the resultant mixture of tryptic peptides, followed by tandem ion trap mass spectrometric analysis of selected individual peptides, and database searching of the relevant genomic database using the program PepFrag. The scheme provides sensitive, real-time protein identification as well as facile identification of modifications. A single operator can unambiguously identify 5-10 proteins/day from an organism whose genome is known at a level of > 0.5 pmol of protein loaded on a gel. The utility of the technique was demonstrated by the identification and characterization of a band from a human HTLV-I preparation and 11 different proteins from a yeast RNA polymerase II C-terminal repeat domain-affinity preparation. The technology has great potential for postgenome biological science, where it promises to facilitate the dissection and anatomy of macromolecular assemblages, the definition of disease state markers, and the investigation of protein targets in biological processes such as the cell cycle and signal transduction.

Amino Acid Sequence↗