Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Popitam: towards new heuristic strategies to improve protein identification from tandem mass spectrometry data.

In recent years, proteomics research has gained importance due to increasingly powerful techniques in protein purification, mass spectrometry and identification, and due to the development of extensive protein and DNA databases from various organisms. Nevertheless, current identification methods from spectrometric data have difficulties in handling modifications or mutations in the source peptide. Moreover, they have low performance when run on large databases (such as genomic databases), or with low quality data, for example due to bad calibration or low fragmentation of the source peptide. We present a new algorithm dedicated to automated protein identification from tandem mass spectrometry (MS/MS) data by searching a peptide sequence database. Our identification approach shows promising properties for solving the specific difficulties enumerated above. It consists of matching theoretical peptide sequences issued from a database with a structured representation of the source MS/MS spectrum. The representation is similar to the spectrum graphs commonly used by de novo sequencing software. The identification process involves the parsing of the graph in order to emphasize relevant sections for each theoretical sequence, and leads to a list of peptides ranked by a correlation score. The parsing of the graph, which can be a highly combinatorial task, is performed by a bio-inspired algorithm called Ant Colony Optimization algorithm.

Algorithms↗

Mass spectrometric analysis of expression of ATPase subunits encoded by duplicated genes in the 19S regulatory particle of rice 26S proteasome.

The 26S proteasome consisting of a 20S proteasome and a pair of 19S regulatory particles (RP) plays important roles in degradation of the ubiquitinated protein in eukaryotic cells. The RP consists of six different ATPase subunits and, at least, 11 non-ATPase subunits. In rice, we previously identified duplicated genes encoding four ATPase subunits, OsRpt1, OsRpt2, OsRpt4, and OsRpt5. In this study, the genomic sequences of all rice ATPase subunits were identified from the rice genome database and the genomic structure of ATPase subunit genes was determined. The rice RP was purified, and the ATPase subunit isoforms encoded by three pairs of duplicated genes, OsRpt2a/OsRpt2b, OsRpt4a/OsRpt4b, and OsRpt5a/OsRpt5b, were identified in RP by using electrospray ionization quadrupole time-of-flight mass spectrometry. The relative amounts and the expression patterns of these ATPase subunit isoforms in the bran were found to be different from those of the callus, suggesting the presence of multiform 19S regulatory particles engaged in the tissue-specific protein metabolism.

Adenosine Triphosphatases↗

Role of ctDNA Tumor Fraction in Selecting Immunotherapy-Based Regimens in Advanced Non-Small Cell Lung Cancer.

PURPOSE: Immune checkpoint blockers (ICB) have transformed advanced non-small cell lung cancer (aNSCLC) treatment, but identifying patients who benefit from adding chemotherapy remains challenging, especially in PD-L1 &#x2265; 50%. PD-L1 is an imperfect biomarker, highlighting the need for better selection tools. EXPERIMENTAL DESIGN: Liquid biopsy (LBx) assessment was performed using hybrid capture-based next-generation sequencing of plasma cell-free DNA. LBx data, molecular profile, and clinicopathologic data were collected. The predictive and prognostic values of tumor fraction (TF) were assessed using a deidentified nationwide (US-based) NSCLC clinicogenomic database [Clinico-Genomic Database (CGDB)]. An independent cohort with aNSCLC from Gustave Roussy was used to validate the findings and to study the correlation of circulating tumor DNA (ctDNA) TF and total metabolic tumor volume and its molecular correlates. RESULTS: In the CGDB database (n = 965), elevated ctDNA TF was prognostic for worse outcomes on ICBs and, when &#x2265;5%, predictive of benefit from ICB + chemotherapy [HR for real-world progression-free survival 0.58 (0.41-0.82); P = 0.002]. The 5% cutoff for TF was validated in an independent cohort from Gustave Roussy. In 283 patients with paired PET scans, ctDNA TF correlated with metabolic tumor volume (rho = 0.46; P < 0.001) and was influenced by TP53/RB1 mutations. CONCLUSIONS: ctDNA TF integrates disease burden and biology. Patients with high ctDNA TF derive greater benefit from chemoimmunotherapy, supporting its use as a biomarker to guide treatment intensification.

Humans↗

G-InforBIO: integrated system for microbial genomics.

BACKGROUND: Genome databases contain diverse kinds of information, including gene annotations and nucleotide and amino acid sequences. It is not easy to integrate such information for genomic study. There are few tools for integrated analyses of genomic data, therefore, we developed software that enables users to handle, manipulate, and analyze genome data with a variety of sequence analysis programs. RESULTS: The G-InforBIO system is a novel tool for genome data management and sequence analysis. The system can import genome data encoded as eXtensible Markup Language documents as formatted text documents, including annotations and sequences, from DNA Data Bank of Japan and GenBank encoded as flat files. The genome database is constructed automatically after importing, and the database can be exported as documents formatted with eXtensible Markup Language or tab-deliminated text. Users can retrieve data from the database by keyword searches, edit annotation data of genes, and process data with G-InforBIO. In addition, information in the G-InforBIO database can be analyzed seamlessly with nine different software programs, including programs for clustering and homology analyses. CONCLUSION: The G-InforBIO system simplifies genome analyses by integrating several available software programs to allow efficient handling and manipulation of genome data. G-InforBIO is freely available from the download site.

Algorithms↗

Matching peptide mass spectra to EST and genomic DNA databases.

The use of mass spectrometry data to search molecular sequence databases is a well-established method for protein identification. The technique can be extended to searching raw genomic sequences, providing experimental confirmation or correction of predicted coding sequences, and has the potential to identify novel genes and elucidate splicing patterns.

Amino Acid Sequence↗

Toward genomic identification of beta-barrel membrane proteins: composition and architecture of known structures.

The amino acid composition and architecture of all beta-barrel membrane proteins of known three-dimensional structure have been examined to generate information that will be useful in identifying beta-barrels in genome databases. The database consists of 15 nonredundant structures, including several novel, recent structures. Known structures include monomeric, dimeric, and trimeric beta-barrels with between 8 and 22 membrane-spanning beta-strands each. For this analysis the membrane-interacting surfaces of the beta-barrels were identified with an experimentally derived, whole-residue hydrophobicity scale, and then the barrels were aligned normal to the bilayer and the position of the bilayer midplane was determined for each protein from the hydrophobicity profile. The abundance of each amino acid, relative to the genomic abundance, was calculated for the barrel exterior and interior. The architecture and diversity of known beta-barrels was also examined. For example, the distribution of rise-per-residue values perpendicular to the bilayer plane was found to be 2.7 +/- 0.25 A per residue, or about 10 +/- 1 residues across the membrane. Also, as noted by other authors, nearly every known membrane-spanning beta-barrel strand was found to have a short loop of seven residues or less connecting it to at least one adjacent strand. Using this information we have begun to generate rapid screening algorithms for the identification of beta-barrel membrane proteins in genomic databases. Application of one algorithm to the genomes of Escherichia coli and Pseudomonas aeruginosa confirms its ability to identify beta-barrels, and reveals dozens of unidentified open reading frames that potentially code for beta-barrel outer membrane proteins.

Bacterial Outer Membrane Proteins↗

Screening of putative oxygenase genes in the Fusarium graminearum genome sequence database for their role in trichothecene biosynthesis.

In the biosynthesis of type B trichothecenes, four oxygenation steps remain to have genes functionally assigned to them. On the basis of the complete genome sequence of Fusarium graminearum, expression patterns of all oxygenase genes were investigated in Fusarium asiaticum (F. graminearum lineage 6). As a result, we identified five cytochrome P450 monooxygenase (CYP) genes that are specifically expressed under trichothecene-producing conditions and are unique to the toxin-producing strains. The entire coding regions of four of these genes were identified in F. asiaticum. When expressed in Saccharomyces cerevisiae, none of the oxygenases were able to transform trichodiene-11-one to expected products. However, one of the oxygenases catalyzed the 2beta-hydroxylation rather than the expected 2alpha-hydroxylation. Targeted disruption of the five CYP genes did not alter the trichothecene profiles of F. asiaticum. The results are discussed in relation to the presence of as-yet-unidentified oxygenation genes that are necessary for the biosynthesis of trichothecenes.

DNA, Fungal↗

Mapping the oral microbiome opens links to periodontitis.

Many microbiome analysis techniques can only detect the microbes present in the reference genome database used. In this issue of Cell Host & Microbe, Cha et al. establish an improved genome database of the human oral microbiome, which they use to discover a connection between periodontitis and an enigmatic bacterial phylum.

Humans↗

Accurate mass multiplexed tandem mass spectrometry for high-throughput polypeptide identification from mixtures.

We report a new tandem mass spectrometric approach for the improved identification of polypeptides from mixtures (e.g., using genomic databases). The approach involves the dissociation of several species simultaneously in a single experiment and provides both increased speed and sensitivity. The data analysis makes use of the known fragmentation pathways for polypeptides and highly accurate mass measurements for both the set of parent polypeptides and their fragments. The accurate mass information makes it possible to attribute most fragments to a specific parent species. We provide an initial demonstration of this multiplexed tandem MS approach using an FTICR mass spectrometer with a mixture of seven polypeptides dissociated using infrared irradiation from a CO2 laser. The peptides were added to, and then successfully identified from, the largest genomic database yet available (C. elegans), which is equivalent in complexity to that for a specific differentiated mammalian cell type. Additionally, since only a few enzymatic fragments are necessary to unambiguously identify a protein from an appropriate database, it is anticipated that the multiplexed MS/MS method will allow the more rapid identification of complex protein mixtures with on-line separation of their enzymatically produced polypeptides.

Amino Acid Sequence↗

A novel mycobacterial antigen relevant to cellular immunity belongs to a family of secreted lipoproteins.

The gene sequence of a novel 24.1 kDa Mycobacterium tuberculosis protein was identified within the Sanger Centre (UK) M. tuberculosis genome database (cosmid MTCY24G1) by searching with a 126 bp DNA sequence isolated from a genomic M. leprae lambda gt11 library with M. leprae reactive human T cell clones as probes. The 24.1 kDa antigen is common to the vaccine strain Mycobacterium bovis BCG, as well as Mycobacterium leprae. The 699 bp open reading frame encodes a 233 amino acid long precursor protein with a signal peptide sequence for secretion and a consensus motif for lipid conjugation, which suggests that the mature protein is an exported lipoprotein antigen. The molecular mass of the mature protein antigen from M. leprae sonicate was shown to correspond to the deduced size of the M. tuberculosis protein by T cell Western analysis. Homology searches revealed two other similarly sized hypothetical secreted mycobacterial lipoproteins within the M. tuberculosis genome database.

Amino Acid Sequence↗

[Methods of statistical genetics and use of database for genome information].

Knowledge and technology of bioinformatics have become inevitable for gene and genome research. Education and research in this field of science are not sufficient in Japan. There are two different approaches to trait mapping, the way by which traits are mapped on the genome. Thus, the knowledge-based approach uses functions of molecules while the statistics-based approach uses polymorphisms. Statistics-based approach uses two different methods, linkage analysis and analysis based on linkage disequilibrium. Various phenotypes are efficiently mapped on the genome using such methods. Recently, bioinformatic data base search is mostly performed using internet. Anyone can perform sequence-search, homology-search and SNP-search. Since such data bases change quickly, readers should access the databases themselves and be used to the procedures for them.

Computational Biology↗

Immune-related, lectin-like receptors are differentially expressed in the myeloid and lymphoid lineages of zebrafish.

The identification of C-type lectin (Group V) natural killer (NK) cell receptors in bony fish has remained elusive. Analyses of the Fugu rubripes genome database failed to identify Group V C-type lectin domains (Zelensky and Gready, BMC Genomics 5:51, 2004) suggesting that bony fish, in general, may lack such receptors. Numerous Group II C-type lectin receptors, which are structurally similar to Group V (NK) receptors, have been characterized in bony fish. By searching the zebrafish genome database we have identified a multi-gene family of Group II immune-related, lectin-like receptors (illrs) whose members possess inhibiting and/or activating signaling motifs typical of Group V NK receptors. Illr genes are differentially expressed in the myeloid and lymphoid lineages, suggesting that they may play important roles in the immune functions of multiple hematopoietic cell lineages.

Amino Acid Sequence↗

Comparative genome analysis of the yellow fever mosquito Aedes aegypti with Drosophila melanogaster and the malaria vector mosquito Anopheles gambiae.

An in silico comparative genomics approach was used to identify putative orthologs to genetically mapped genes from the mosquito, Aedes aegypti, in the Drosophila melanogaster and Anopheles gambiae genome databases. Comparative chromosome positions of 73 D. melanogaster orthologs indicated significant deviations from a random distribution across each of the five A. aegypti chromosomal regions, suggesting that some ancestral chromosome elements have been conserved. However, the two genomes also reflect extensive reshuffling within and between chromosomal regions. Comparative chromosome positions of A. gambiae orthologs indicate unequivocally that A. aegypti chromosome regions share extensive homology to the five A. gambiae chromosome arms. Whole-arm or near-whole-arm homology was contradicted with only two genes among the 75 A. aegypti genes for which orthologs to A. gambiae were identified. The two genomes contain large conserved chromosome segments that generally correspond to break/fusion events and a reciprocal translocation with extensive paracentric inversions evident within. Only very tightly linked genes are likely to retain conserved linear orders within chromosome segments. The D. melanogaster and A. gambiae genome databases therefore offer limited potential for comparative positional gene determinations among even closely related dipterans, indicating the necessity for additional genome sequencing projects with other dipteran species.

Aedes↗

Refined physical map of the human PAX2/HOX11/NFKB2 cancer gene region at 10q24 and relocalization of the HPV6AI1 viral integration site to 14q13.3-q21.1.

BACKGROUND: Chromosome band 10q24 is a gene-rich domain and host to a number of cancer, developmental, and neurological genes. Recurring translocations, deletions and mutations involving this chromosome band have been observed in different human cancers and other disease conditions, but the precise identification of breakpoint sites, and detailed characterization of the genetic basis and mechanisms which underlie many of these rearrangements has yet to be resolved. Towards this end it is vital to establish a definitive genetic map of this region, which to date has shown considerable volatility through time in published works of scientific journals, within different builds of the same international genomic database, and across the differently constructed databases. RESULTS: Using a combination of chromosome and interphase fluorescent in situ hybridization (FISH), BAC end-sequencing and genomic database analysis we present a physical map showing that the order and chromosomal orientation of selected genes within 10q24 is CEN-CYP2C9-PAX2-HOX11-NFKB2-TEL. Our analysis has resolved the orientation of an otherwise dynamically evolving assembly of larger contigs upstream of this region, and in so doing verifies the order and orientation of a further 9 cancer-related genes and GOT1. This study further shows that the previously reported human papillomavirus type 6a DNA integration site HPV6AI1 does not map to 10q24, but that it maps at the interface of chromosome bands 14q13.3-q21.1. CONCLUSIONS: This revised map will allow more precise localization of chromosome rearrangements involving chromosome band 10q24, and will serve as a useful baseline to better understand the molecular aetiology of chromosomal instability in this region. In particular, the relocation of HPV6AI1 is important to report because this HPV6a integration site, originally isolated from a tonsillar carcinoma, was shown to be rearranged in other HPV6a-related malignancies, including 2 of 25 genital condylomas, and 2 of 7 head and neck tumors tested. Our finding shifts the focus of this genomic interest from 10q24 to the chromosome 14 site.

Chromosomes, Artificial, Bacterial↗

Congenital and idiopathic scoliosis: clinical and genetic aspects.

OBJECTIVE: Genetic and environmental factors influencing spinal development in lower vertebrates are likely to play a role in the abnormalities associated with human congenital scoliosis (CS) and idiopathic scoliosis (IS). An overview of the molecular embryology of spinal development and the clinical and genetic aspects of CS and IS are presented. Utilizing synteny analysis of the mouse and human genetic databases, likely candidate genes for human CS and IS were identified. DESIGN: Review and synteny analysis. METHODS: A search of the Mouse Genome Database was performed for "genes," "markers" and "phenotypes" in the categories Neurological and neuromuscular, Skeleton, and Tail and other appendages. The Online Mendelian Inheritance in Man was used to determine whether each mouse locus had a known human homologue. If so, the human homologue was assigned candidate gene status. Linkage maps of the chromosomes carrying loci with possibly relevant phenotypes, but without known human homologues, were examined and regions of documented synteny between the mouse and human genomes were identified. RESULTS: Searching the Mouse Genome Database by phenotypic category yielded 100 mutants of which 66 had been mapped. The descriptions of each of these 66 loci were retrieved to determine which among these included phenotypes of scoliosis, kinky or bent tails, other vertebral abnormalities, or disturbances of axial skeletal development. Forty-five loci of interest remained, and for 27 of these the comparative linkage maps of mouse and human were used to identify human syntenic regions to which plausible candidate genes had been mapped. CONCLUSION: Synteny analysis of mouse candidate genes for CS and IS holds promise due to the close evolutionary relationship between mice and human beings. With the identification of additional genes in animal model systems that contribute to different stages of spine development, the list of candidate genes for CS and IS will continue to grow.

Animals↗

Identifying transcription factor binding sites through Markov chain optimization.

Even though every cell in an organism contains the same genetic material, each cell does not express the same cohort of genes. Therefore, one of the major problems facing genomic research today is to determine not only which genes are differentially expressed and under what conditions, but also how the expression of those genes is regulated. The first step in determining differential gene expression is the binding of sequence-specific DNA binding proteins (i.e. transcription factors) to regulatory regions of the genes (i.e. promoters and enhancers). An important aspect to understanding how a given transcription factor functions is to know the entire gamut of binding sites and subsequently potential target genes that the factor may bind/regulate. In this study, we have developed a computer algorithm to scan genomic databases for transcription factor binding sites, based on a novel Markov chain optimization method, and used it to scan the human genome for sites that bind to hepatocyte nuclear factor 4 alpha (HNF4alpha). A list of 71 known HNF4alpha binding sites from the literature were used to train our Markov chain model. By looking at the window of 600 nucleotides around the transcription start site of each confirmed gene on the human genome, we identified 849 sites with varying binding potential and experimentally tested 109 of those sites for binding to HNF4alpha. Our results show that the program was very successful in identifying 77 new HNF4alpha binding sites with varying binding affinities (i.e. a 71% success rate). Therefore, this computational method for searching genomic databases for potential transcription factor binding sites is a powerful tool for investigating mechanisms of differential gene regulation.

Algorithms↗