Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Transitive functional annotation by shortest-path analysis of gene expression data.

Current methods for the functional analysis of microarray gene expression data make the implicit assumption that genes with similar expression profiles have similar functions in cells. However, among genes involved in the same biological pathway, not all gene pairs show high expression similarity. Here, we propose that transitive expression similarity among genes can be used as an important attribute to link genes of the same biological pathway. Based on large-scale yeast microarray expression data, we use the shortest-path analysis to identify transitive genes between two given genes from the same biological process. We find that not only functionally related genes with correlated expression profiles are identified but also those without. In the latter case, we compare our method to hierarchical clustering, and show that our method can reveal functional relationships among genes in a more precise manner. Finally, we show that our method can be used to reliably predict the function of unknown genes from known genes lying on the same shortest path. We assigned functions for 146 yeast genes that are considered as unknown by the Saccharomyces Genome Database and by the Yeast Proteome Database. These genes constitute around 5% of the unknown yeast ORFome.

Cell Nucleus↗

PAINT: a promoter analysis and interaction network generation tool for gene regulatory network identification.

We have developed a bioinformatics tool named PAINT that automates the promoter analysis of a given set of genes for the presence of transcription factor binding sites. Based on coincidence of regulatory sites, this tool produces an interaction matrix that represents a candidate transcriptional regulatory network. This tool currently consists of (1) a database of promoter sequences of known or predicted genes in the Ensembl annotated mouse genome database, (2) various modules that can retrieve and process the promoter sequences for binding sites of known transcription factors, and (3) modules for visualization and analysis of the resulting set of candidate network connections. This information provides a substantially pruned list of genes and transcription factors that can be examined in detail in further experimental studies on gene regulation. Also, the candidate network can be incorporated into network identification methods in the form of constraints on feasible structures in order to render the algorithms tractable for large-scale systems. The tool can also produce output in various formats suitable for use in external visualization and analysis software. In this manuscript, PAINT is demonstrated in two case studies involving analysis of differentially regulated genes chosen from two microarray data sets. The first set is from a neuroblastoma N1E-115 cell differentiation experiment, and the second set is from neuroblastoma N1E-115 cells at different time intervals following exposure to neuropeptide angiotensin II. PAINT is available for use as an agent in BioSPICE simulation and analysis framework (www.biospice.org), and can also be accessed via a WWW interface at www.dbi.tju.edu/dbi/tools/paint/.

Angiotensin II↗

Genomic data visualization on the Web.

UNLABELLED: Many types of genomic data can be represented in matrix format, with rows corresponding to genes and columns corresponding to gene features. The heat map is a popular technique for visualizing such data, plotting the data on a two-dimensional grid and using a color scale to represent the magnitude of each matrix entry. Prism is a Web-based software tool for generating annotated heat map visualizations of genome-wide data quickly. The tool provides a selection of genome-specific annotation catalogs as well as a catalog upload capability. The heat maps generated are clickable, allowing the user to drill down to examine specific matrix entries, and gene annotations are linked to relevant genomic databases. AVAILABILITY: http://noble.gs.washington.edu/prism

Computer Graphics↗

Whole-genome experimental identification of insertion/deletion polymorphisms of interspersed repeats by a new general approach.

A new experimental technique for genome-wide detection of integration sites of polymorphic retroelements (REs) is described. The technique allows one to reveal the absence of a retroelement in an individual genome provided that this retroelement is present in at least one of several other genomes under comparison. Since quite a number of genomes are compared simultaneously, the search for polymorphic REs insertions is very efficient. The technique includes two whole-genome selective PCR amplifications of sequences flanking REs: one for a particular genome and another one for a mixture of ten different genomes. A subsequent subtractive hybridization of the obtained amplicons with DNA of a particular genome as driver results in isolation of polymorphic insertions. The technique was successfully applied for identification of 41 new polymorphic human AluYa5/Ya8 insertions. Among them, 18 individual Alu elements first sequenced in this work were not found in the available human genome databases. This result suggests that significant part of polymorphic REs were not identified during genome sequencing and remain to be detected and characterized. The proposed method does not depend on preliminary knowledge of evolutionary history of retroelements and can be applied for identification of insertion/deletion polymorphic markers in genomes of different species.

Alu Elements↗

Designability, aggregation propensity and duplication of disease-associated proteins.

Over 2000 proteins in the Ensembl human genome database have been linked with disease information from OMIM. In comparison with all human proteins, we find that disease-associated proteins tend to have less designable folds in terms of their SCOP family counts, suggesting that they are intrinsically less robust to mutation and environmental stress. Disease proteins also tend to have isoelectric points closer to neutrality and more alternating hydrophilic-hydrophobic amino acid stretches compared with the average human protein. These results suggest that protein aggregation is a significant phenomenon associated with diseases. Another finding in this work is that many disease proteins are highly sequence similar to other disease proteins, suggesting that gene duplication has contributed to the expansion of disease-prone protein families.

Amino Acid Substitution↗

The maize genome contains a helitron insertion.

The maize mutation sh2-7527 was isolated in a conventional maize breeding program in the 1970s. Although the mutant contains foreign sequences within the gene, the mutation is not attributable to an interchromosomal exchange or to a chromosomal inversion. Hence, the mutation was caused by an insertion. Sequences at the two Sh2 borders have not been scrambled or mutated, suggesting that the insertion is not caused by a catastrophic reshuffling of the maize genome. The insertion is large, at least 12 kb, and is highly repetitive in maize. As judged by hybridization, sorghum contains only one or a few copies of the element, whereas no hybridization was seen to the Arabidopsis genome. The insertion acts from a distance to alter the splicing of the sh2 pre-mRNA. Three distinct intron-bearing maize genes were found in the insertion. Of most significance, the insertion bears striking similarity to the recently described DNA helicase-bearing transposable elements termed HELITRONS: Like Helitrons, the inserted sequence of sh2-7527 is large, lacks terminal repeats, does not duplicate host sequences, and was inserted between a host dinucleotide AT. Like Helitrons, the maize element contains 5' TC and 3' CTRR termini as well as two short palindromic sequences near the 3' terminus that potentially can form a 20-bp hairpin. Although the maize element lacks sequence information for a DNA helicase, it does contain four exons with similarity to a plant DEAD box RNA helicase. A second Helitron insertion was found in the maize genomic database. These data strongly suggest an active Helitron in the present-day maize genome.

Alternative Splicing↗

Indigo: a World-Wide-Web review of genomes and gene functions.

The present article describes a genome database reviewing gene-related knowledge of two model bacteria, Bacillus subtilis and Escherichia coli. The database, Indigo, is open through the World-Wide Web (http://indigo.genetique.uvsq.fr). The concept used for organising the data, the concept of neighbourhood, allows one to explore the database content in an efficient although somewhat unusual way. Here, genes are related to each other by a variety of neighbourhoods, including proximity in the chromosome, phylogenetic kinship, participation in a common metabolic pathway, common presence in an article of the literature, or similar use of the genetic code. Several examples illustrate how this concept of neighbourhood permits one to review the available knowledge about a given gene or gene family, and elaborate unexpected, but revealing, analyses about gene functions.

Bacillus subtilis↗

The connexin gene family in mammals.

Unannotated mammalian genome databases (dog, cow, opossum) were searched for candidate connexin genes, using sequences from annotated genomes (man, mouse). All 18 'multi-species' connexin genes, i.e., orthologs of connexin26 , 29/31.3 (duplicated in opossum), 30, 30.2/31.9, 30.3, 31, 31.1, 32, 36, 37, 39/40.1, 40, 43, 45, 44/46, 47, 50, and 57/62 , were found in dog, cow and opossum. Connexin25 and 58 have been considered specific for man, but evident orthologs of connexin25 were found in dog, cow and opossum, and orthologs of connexin58 were found in dog and cow. Moreover, a connexin43 -like sequence (approx. 80% identical to connexin43 ) was found in man, chimpanzee, dog and cow. In the three former species, the sequences were located on the X chromosome. In man, chimpanzee and cow, there were stop codons in all reading frames; these sequences are therefore judged as pseudogenes, called here Cx43pX . In the dog, the sequence contained an open reading frame for a protein of 35.7 kDa (connexin35.7). We suggest that these sequences are orthologs of connexin33 , previously considered as a rodent-specific connexin gene. Thus, connexin25 , 33 and 58 are not species-specific genes. However, the opossum may possess a candidate, connexin39.2 , without obvious orthologs in other mammals. Furthermore, pseudogenes of primate connexin31.3 and opossum connexin35 (one of the two orthologs of primate connexin31.3) were detected. These results suggest that the structure of the mammalian connexin gene family should be revised, especially with regard to the so-called 'species-specific' connexins .

Animals↗

The TRIPLES database: a community resource for yeast molecular biology.

TRIPLES is a web-accessible database of TRansposon-Insertion Phenotypes, Localization and Expression in Saccharomyces cerevisiae-a relational database housing nearly half a million data points generated from an ongoing study using large-scale transposon mutagenesis to characterize gene function in yeast. At present, TRIPLES contains three principal data sets (i.e. phenotypic data, protein localization data and expression data) for over 3500 annotated yeast genes as well as several hundred non-annotated open reading frames. In addition, the TRIPLES web site provides online order forms linked to each data set so that users may request any strain or reagent generated from this project free of charge. In response to user requests, the TRIPLES web site has undergone several recent modifications. Our localization data have been supplemented with approximately 500 fluorescent micrographs depicting actual staining patterns observed upon indirect immunofluorescence analysis of indicated epitope-tagged proteins. These localization data, as well as all other data sets within TRIPLES, are now available in full as tab-delimited text. To accommodate increased reagent requests, all orders are now cataloged in a separate database, and users are notified immediately of order receipt and shipment. Also, TRIPLES is one of five sites incorporated into the new functional analysis tool Function Junction provided by the Saccharomyces Genome Database. TRIPLES may be accessed from the Yale Genome Analysis Center (YGAC) homepage at http://ygac.med.yale.edu.

Computer Graphics↗

N-Terminal amino acid side-chain cleavage of chemically modified peptides in the gas phase: a mass spectrometry technique for N-terminus identification.

Although genome databases have become the key for proteomic analyses, de novo sequencing remains essential for the study of organisms whose genomes have not been completed. In addition, post-translational modifications present a challenge in database searching. Recognition of the b or y-ion series in a peptide MS/MS spectrum as well as identification of the b1 - and yn-1 -ions can facilitate de novo analyses. Therefore, it is valuable to identify either amino-acid terminus. In previous work, we have demonstrated that peptides modified at the epsilon-amino group of lysine as a t-butyl peroxycarbamate derivative undergo free radical promoted peptide backbone fragmentation under low-energy collision-induced dissociation (CID) conditions. Here we explore the chemistry of the N-terminal amino group modified as a t-butyl peroxycarbamate. The conversion of N-terminal amines to peroxycarbamates of simple amino acids and peptides was studied with aryl t-butyl peroxycarbonates. ESI-MS/MS analysis of the peroxycarbamate adducts gave evidence of a product ion corresponding to the neutral loss of the N-terminal side chain (R), thus identifying this residue. Further fragmentation (MS3) of product ions formed by N-terminal residue side-chain loss (-R) exhibited an m/z shift of the b-ions equal to the neutral loss of R, therefore labeling the b-ion series. The study was extended to the analysis of a protein tryptic digest where the SALSA algorithm was used to identify spectra containing these neutral losses. The method for N-terminus identification presented here has the potential for improvement of de novo analyses as well as in constraining peptide mass mapping database searches.

Amines↗

COMBO-FISH: specific labeling of nondenatured chromatin targets by computer-selected DNA oligonucleotide probe combinations.

Here we present the principle of fluorescence in situ hybridization (FISH) with combinatorial oligonucleotide (COMBO) probes as a new approach for the specific labeling of genomic sites. COMBO-FISH takes advantage of homopurine/homopyrimidine oligonucleotides that form triple helices with intact duplex genomic DNA, without the need for prior denaturation of the target sequence that is usually applied for probe binding in standard FISH protocols. An analysis of human genome databases has shown that homopurine/homopyrimidine sequences longer than 14 bp are nearly homogeneously distributed over the genome, and they represent from 1% to 2% of the entire genome. Because the observation volume in a confocal laser-scanning microscope equipped with a high numerical aperture lens typically corresponds to an approximate 250-kb chromatin domain in a normal mammalian cell nucleus, this volume should contain 150-200 homopurine/homopyrimidine stretches. Using DNA database information, one can configure a set of distinct, uniformly labeled oligonucleotide probes from these stretches that is expected to exclusively co-localize within a 250-kb chromatin domain. Due to the diffraction-limited resolution of a microscope, the fluorescence signals of the configured oligonucleotide probe set merge into a typical, nearly homogenous FISH spot. Using a set of 32 homopyrimidine probes, we performed experiments in the Abelson murine leukemia region of human chromosome 9 as some of the very first proofs-of-principle of COMBO-FISH. Although the experimental protocol currently contains several steps that are incompatible with living cell conditions, the theoretical approach may be the first methodological advance toward the long-term but still elusive goal of carrying out specific FISH in high-resolution fluorescence microscopy of vital cells.

Chromatin↗

Real-world estrogen receptor alpha 1 (ESR1) testing patterns and results for ER+/HER2- metastatic breast cancer in the United States, 2018-2024.

PURPOSE: To understand historical and recent ESR1 testing rates, and when ESR1 mutations emerge during first-line (1 L) treatment. METHODS: This retrospective, observational cohort study used the Flatiron Health Research Database (FHRD) and the Flatiron Health-Foundation Medicine metastatic breast cancer (mBC) Clinico-Genomic Database (CGDB). Adult patients with a confirmed diagnosis of hormone receptor-positive/human epidermal growth factor receptor 2-negative mBC from 1/1/2018 to 6/30/2024 were included. ESR1 testing patterns and test results were descriptively analyzed. RESULTS: Among 7772 patients with mBC in the FHRD who initiated 1 L therapy, tumor ESR1 mutation status was evaluated for 222 (3%) patients at baseline (≤ 90 days before 1 L) and 1355 (17%) during 1 L. The percentage of patients who had an ESR1 test result reported during 1 L increased over time (11% in 2018-19, 19% in 2020-21, 22% in 2022-24). Median time from 1 L start to first ESR1 test was 7.4 months (mos) among tested patients. A positive test result was reported for 29/222 (13%) patients tested at baseline and 240/1355 (18%) tested during 1 L. Most (60%) tests during 1 L used tissue specimens, while the remaining 40% were liquid biopsies, and the median time from specimen collection to result reporting in 1 L was 28 (IQR:10-84) days. Focusing on time periods wherein specimens were provided, ESR1 test positivity was 6.7% (76/1,127) for specimens provided at baseline, 23% (15/65) for specimens provided 9 to 12 months into 1L therapy, 38% (26/69) for those provided 15 to 18 months into 1L, and 40% (38/94) for those provided from 18 to 24 months into 1L. CONCLUSIONS: ESR1 mutations can be detected at any time interval during 1L. CLINICAL TRIAL NUMBER: Not applicable.

Adult↗

A metadata approach to query interoperation between molecular biology databases.

MOTIVATION: Molecular biology databases have been proliferating rapidly. Their heterogeneity and complexity pose a great challenge to efforts in database interoperation. To minimize the efforts of interoperating heterogeneous databases, it is useful to develop a system that lets a user of a particular genomic database access another related database as if the latter is structurally similar to the former. RESULTS: We extend a structurally simple model-the entity-attribute-value (EAV) model-to describe uniformly metadata relating to individual databases. Such metadata, which are necessary for performing database comparisons, include descriptions of primitive database objects (including entities, attributes, domain values and entity relationships) and specification of correspondences among the database objects. We show how to decompose SQL queries and map them from one database to another based on the EAV representation of the basic database objects. A prototype system is implemented to demonstrate query interoperation between two chromosome map databases. AVAILABILITY: Freely available (Cold Fusion source code and an Access database containing the mapping knowledge) upon request from the author. CONTACT: kei.cheung@yale.edu

Chromosome Mapping↗

Evolution of the integral membrane desaturase gene family in moths and flies.

Lepidopteran insects use sex pheromones derived from fatty acids in their species-specific mate recognition system. Desaturases play a particularly prominent role in the generation of structural diversity in lepidopteran pheromone biosynthesis as a result of the diverse enzymatic properties they have evolved. These enzymes are homologous to the integral membrane desaturases, which play a primary role in cold adaptation in eukaryotic cells. In this investigation, we screened for desaturase-encoding sequences in pheromone glands of adult females of eight lepidopteran species. We found, on average, six unique desaturase-encoding sequences in moth pheromone glands, the same number as is found in the genome database of the fly, Drosophila melanogaster, vs. only one to three in other characterized eukaryotic genomes. The latter observation suggests the expansion of this gene family in insects before the divergence of lepidopteran and dipteran lineages. We present the inferred homology relationships among these sequences, analyze nonsynonymous and synonymous substitution rates for evidence of positive selection, identify sequence and structural correlates of three lineages containing characterized enzymatically distinct desaturases, and discuss the evolution of this sequence family in insects.

Amino Acid Motifs↗

Isolation of a cDNA for a novel human RING finger protein gene, RNF18, by the virtual transcribed sequence (VTS) approach(1).

We have recently developed a novel database system, designated as the virtual transcribed sequence (VTS) which efficiently extracts many genes from public human genome databases, and tested the feasibility of this novel computational approach (N. Miyajima, C. Burge, T. Saito, Biochem. Biophys. Res. Commun. 272 (2000) 801; http://host45.maze.co.jp/vts/). In this study, using the VTS approach, we isolated a cDNA for a novel human gene with RING finger motif (C(3)HC(4)), which is not deposited in public EST databases. The isolated cDNA clone is 2163 bp in length, and contains an open reading frame of 452 amino acids. We designated the novel gene as RNF18. A database search showed that the RNF18 gene had the moderate similarity to SS-A/Ro52 protein, which is a ribonucleoprotein reactive with autoantibodies in patients with Sjögren's syndrome and systemic lupus erythematosus. Tissue distribution analyses by Northern blot and RT-PCR methods demonstrated that the RNF18 messenger RNA was preferentially expressed in testis. The exon-intron boundaries of RNF18 gene were determined by aligning the cDNA sequence with the corresponding genome sequence. The isolated cDNA consists of eight exons that span about 11 kb of the genome DNA. The precise chromosomal location of the RNF18 gene was determined by PCR-based radiation hybrid mapping, and the gene was located to centromere region of chromosome 11 between markers NIB1900 and D11S1350. Taken together, the VTS approach should provide a novel cDNA cloning strategy for isolating unidentified genes, which are not found even in EST databases but are detectable computationally.

Amino Acid Sequence↗

AnoBase: a genetic and biological database of anophelines.

AnoBase (http://www.anobase.org) is an integrated, relational database of basic biological and genetic data on anopheline species, with a particular emphasis on Anopheles gambiae. It has been designed as an information source and research support tool for the broad vector biology community. Although AnoBase is not a primary genomic database that develops and provides tools to access the genome of the malaria mosquito, it nevertheless contains several sections that offer data of genomic interest such as in situ hybridization images, an integrated gene tool and direct online access to AnoXcel, the proteomic database of An. gambiae. Moreover, AnoBase also contains information on non-gambiae mosquito species and a novel section on studies related to insecticide resistance.

Animals↗

Analysis of proteins and proteomes by mass spectrometry.

A decade after the discovery of electrospray and matrix-assisted laser desorption ionization (MALDI), methods that finally allowed gentle ionization of large biomolecules, mass spectrometry has become a powerful tool in protein analysis and the key technology in the emerging field of proteomics. The success of mass spectrometry is driven both by innovative instrumentation designs, especially those operating on the time-of-flight or ion-trapping principles, and by large-scale biochemical strategies, which use mass spectrometry to detect the isolated proteins. Any human protein can now be identified directly from genome databases on the basis of minimal data derived by mass spectrometry. As has already happened in genomics, increased automation of sample handling, analysis, and the interpretation of results will generate an avalanche of qualitative and quantitative proteomic data. Protein-protein interactions can be analyzed directly by precipitation of a tagged bait followed by mass spectrometric identification of its binding partners. By these and similar strategies, entire protein complexes, signaling pathways, and whole organelles are being characterized. Posttranslational modifications remain difficult to analyze but are starting to yield to generic strategies.

Chromatography, Liquid↗

Proteomic analysis of the Arabidopsis thaliana cell wall.

With the completion of the Arabidopsis genome, many hypothetical proteins have been predicted without any information on their expression, subcellular localisation and function. We have performed proteomic analysis of proteins sequentially extracted from enriched Arabidopsis cell wall fractions and separated by two-dimensional gel electrophoresis (2-DE). The proteins were identified by peptide mass fingerprinting using matrix-assisted laser desorption/ionisation-time of flight (MALDI-TOF) mass spectrometry and genomic database searches. This is part of a targeted exercise to establish the entire Arabidopsis secretome database. We report evidence for new proteins of unknown function whose existence had been predicted from genomic sequences and, furthermore, localise them to the cell wall. In addition, we observed an unexpected presence in the cell wall preparations of proteins whose known biochemical activity has never been associated with this compartment hitherto. We discuss the implications of these findings and present results suggesting a possible involvement of cell wall kinases in plant responses to pathogen attack.

Arabidopsis↗